EP4681155A1 - Semantic-based image copying - Google Patents
Semantic-based image copyingInfo
- Publication number
- EP4681155A1 EP4681155A1 EP24739314.3A EP24739314A EP4681155A1 EP 4681155 A1 EP4681155 A1 EP 4681155A1 EP 24739314 A EP24739314 A EP 24739314A EP 4681155 A1 EP4681155 A1 EP 4681155A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- model
- generative
- generating
- source image
- image
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
Definitions
- the present disclosure relates to techniques for copying images, using generative artificial intelligence, in a manner that creatively changes the images while maintaining semantic context and visual qualities of the source images.
- generative Al “copies” often fail to strike a good balance between: (1) maintaining visual properties of the source image (e.g., maintaining a look and feel of the source image); (2) maintaining the semantic context of the source image (e.g., maintaining what is shown, at a descriptive/conceptual level, in the source image); and yet (3) differing from the source image in a meaningful way (e.g., such that the copy creatively differs from the source image and is not merely a trivial variation of the source image).
- generative Al copies of an original creative may fail to capture a look and feel associated with a particular advertiser (e.g., as represented in the source image by colors, styles, etc.), may fail to represent an advertised product accurately, or may fail to differ from the original creative enough to significantly differ in perform (e.g., have a different level of customer/user engagement) as compared to the original creative.
- a system generates new images by semantically copying source images.
- copying a source image generally refers to generating a new image that is in some way derived from the source image, but without being an exact copy of the source image, and without necessarily (but possibly) reusing or reproducing any portion/pixels of the source image.
- the system of the present disclosure generates a new image from (i.e., generates a semantic copy of) a source image by: (1) generating a text prompt by applying the source image to a first generative artificial intelligence (Al) model that generates a descriptive caption for the source image; (2) generating a visual embedding based on the source image (e.g., a visual embedding of either the source image or a processed version of the source image); and (3) generating the new image using a second generative Al model and based on both the text prompt and the visual embedding.
- Al generative artificial intelligence
- the system By generating the new image based on a text prompt that is itself a descriptive caption generated directly from the source image (or at least, a text prompt that is derived from such a caption), the system provides a “semantic” copy of the source image. That is, the new image adheres well to the source image at a conceptual/descriptive level. Moreover, because the system also uses a visual embedding of the source image to generate the new image, the new image adheres well to visual properties of the source image.
- a text prompt derived from the source image itself is less likely to conflict with a visual embedding of the source image (e.g., as compared to a text prompt that is manually generated), and thus less likely to produce a strange, confusing, or unappealing new image.
- the visual embedding helps the second (image generation) generative Al model avoid hallucinations with respect to what should be shown at a conceptual level. For instance, if the text prompt “a cell phone mounted on a car” is derived from a source image that shows a cell phone mounted on a dashboard of a car, the concurrent use of the visual embedding of the source image may prevent the generative Al model from generating, for example, a new image in which a giant cell phone is mounted on top of and external to a car.
- the visual embedding derived from the source image helps to ground the generative Al model.
- the text prompt and the visual embedding can help ensure that generated images maintain the key concept underlying the source image, while also introducing varieties that may have been introduced via text prompt mutation and/or image processing steps prior to visual embedding (e.g., cropping)
- the disclosed techniques also provide other technical advantages.
- One advantage stems from the fact that the text prompt contains information that is not, in a strict sense, within the source image itself. As a result, the text prompt, despite being derived from the source image, provides a higher level of entropy than would exist if using visual embeddings without the text prompt. That is, the text prompt provides the second generative Al model with a broader creative/imaginary space in which to create the new image.
- converting the source image to a text prompt using the first (captioning) generative Al model, facilitates changes to the source image, as compared to making such changes based solely on a visual embedding or through other image processing means.
- generative Al models e.g., large language models
- modifications e.g., adding, removing, or changing the state or position of objects
- the disclosed system uses a third generative Al model to make such modifications, by modifying or “mutating” the text prompt before the system applies the text prompt (along with the visual embedding) to the second generative Al model.
- a method of semantic-based image copying includes generating, by one or more processors, a text prompt.
- Generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image.
- the method also includes generating, by the one or more processors, a visual embedding based on the source image, and generating, by the one or more processors, a new image using a second generative Al model and based on the text prompt and the visual embedding.
- a system includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to: (1) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image; (2) generate a visual embedding based on the source image; and (3) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
- one or more non-transitory, computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to: (1) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image; (2) generate a visual embedding based on the source image; and (3) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
- FIG. 1 is a block diagram of an example system in which techniques for semanticbased copying of images can be implemented.
- FIG. 2 depicts an example process for semantic-based copying of a source image, which may be implemented by the computing system of FIG. 1.
- FIGs. 3A and 3B depict example processes for semantic-based copying of a source image with feedback, which may be implemented by the computing system of FIG. 1.
- FIG. 4 depicts examples of semantic copies that may be produced by the computing system of FIG.1, using the process of FIG. 2, 3A, or 3B.
- FIG. 5 depicts examples of copies that may be produced by an alternative implementation in which the new image is generated using a text prompt but no visual embedding.
- FIG. 6 is a flow diagram of an example method for semantic-based copying of a source image.
- FIG. 1 is a block diagram of an example system 100 in which techniques for semantic-based image copying/generation can be implemented.
- the example system 100 includes a computing system 102, a client device 104, a content provider 106 (e.g., a server of a content provider), and a network 110.
- the computing system 102 is remote from the client device 104 and content provider 106, and is communicatively coupled to the client device 104 and content provider 106 via the network 110.
- the system 100 does not include client device 104 and/or content provider 106.
- the network 110 may be a single communication network (e.g., the Internet), and in some implementations also includes one or more additional networks.
- the network 110 may include a cellular network, the Internet, and a server- side local area network (LAN). While FIG. 1 shows only a single client device 104 and single content provider 106, it is understood that the computing system 102 may also be in communication with a number (e.g., millions) of other client devices that are generally similar to the client device 104, and/or in communication with a number (e.g., thousands) of other content providers that are generally similar to content provider 106.
- a number e.g., millions
- other client devices that are generally similar to the client device 104
- a number e.g., thousands
- computing system 102 can perform image copying/generation services (e.g., for providers such as content provider 106).
- image copying is generally used herein to refer to generating a new image that is in some way derived from the source image, without being an exact copy of the source image, and without necessarily (but possibly) reusing or reproducing any portion/pixels of the source image.
- computing system 102 may use existing images from content providers such as content provider 106 to generate new images that the content provider can use in additional digital advertising.
- the new/additional images can be used to provide a greater diversity of images/advertisements, the performance of which can then be measured (e.g., based on click- through rate, conversion rate, etc.) to determine which images/advertisements are most effective.
- the new/additional images may have aspect ratios different from the original image, making the new images better suited to ad slots (e.g., in a web page or mobile application) that have different aspect ratio constraints.
- the techniques described herein e.g., in connection with FIGs. 2, 3A, 3B, and 6) can change the aspect ratio of the source image in a more seamless manner than conventional techniques (e.g., salient region detection plus cropping).
- computing system 102 may generate new images/copies that are intended to facilitate viewer understanding (e.g., images for instructional materials), where performance is measured by way of determining what proportion of viewers take the correct actions upon viewing the images.
- Other contexts are also possible. For ease and consistency of explanation, however, this disclosure primarily uses examples that are related to a digital advertising implementation/context.
- the client device 104 is generally configured to access information resources (e.g., web pages and/or user interfaces of mobile applications or other applications) that can present the images generated by computing system 102.
- computing system 102 may generate digital advertisements that include (or consist entirely of) the semantic copies discussed herein.
- Computing system 102 or another computing system may then serve the digital advertisements to users of client device 104 and/or other similar client devices using suitable techniques, such as conducting auctions (e.g., auctions based on keyword bids by advertisers, relevancy metrics, etc.).
- the digital advertisements may be served in slots of web pages visited by the users, and/or slots of application user interfaces displayed to the users, etc.
- the content provider 106 generally may commission or request that computing system 102 generate one or more images, and/or may provide the source image(s) upon which the image generation is based.
- content provider 106 may be a digital advertiser who provides a digital advertisement image for each of a number of offered products or services, as part of one or more advertising campaigns owned or managed by content provider 106.
- the source image may be a screenshot of a web page hosted by content provider 106, a screenshot of a mobile application that content provider 106 offers, and so on.
- the computing system 102 includes a network interface 120, a processor 122, and memory 124.
- the network interface 120 includes hardware, firmware, and/or software configured to enable the computing system 102 to exchange electronic data with the client device 104 and other, similar client devices (and possibly content provider 106, etc.) via the network 110.
- the network interface 120 may include a wired or wireless router and a modem.
- the processor 122 may be a single processor (e.g., a central processing unit (CPU)), or may include multiple processors (e.g., multiple CPUs, or one or more CPUs and one or more graphics processing units (GPUs)).
- Computing system 102 may be a single computing device (e.g., server) at a single location, or may include multiple, coordinating computing devices that are either co-located or remotely distributed.
- the memory 124 is a computer-readable, non-transitory storage unit or device, or collection of such units/devices, that may include persistent and/or non-persistent memory components.
- the memory 124 stores instructions executable by processor 122 to perform various operations, including the instructions of various software applications and the data generated and/or used by such applications.
- memory 124 stores the instructions of a semantic copy generator 130, which includes a captioner 140, a prompt mutator 142, an image processor 144, a visual embedder 146, and an image generator 148.
- Memory 124 can also store generative artificial intelligence (Al) models.
- memory 124 stores a first generative Al model 150 used by captioner 140, a second generative Al model 152 used by image generator 148, and a third generative Al model 154 used by prompt mutator 142.
- third generative Al model 154 is not included in system 100.
- memory 124 may omit one or more modules/elements shown in FIG. 1, such as prompt mutator 142 and/or image processor 144. It is also understood that, in some implementations, memory 124 may include one or more additional modules/elements not shown in FIG.
- first generative Al model 150, second generative Al model 152, and/or third generative Al model 154 are not stored in memory 124, and instead are stored in one or more remote servers or other computing systems.
- one or more of models 150, 152, and 154 may be remotely accessed (e.g., as a cloud service) by semantic copy generator 130.
- the client device 104 may be or include any stationary, mobile, or portable computing device with wired and/or wireless communication capability (e.g., a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart wearable device such as smart glasses or a smart watch, a vehicle head unit computer, etc.).
- client device 104 includes a network interface 160, a processor 162, memory 164, and a display 166.
- the processor 162 may be a single processor, or may include multiple processors.
- Memory 164 includes one or more computer-readable, non-transitory storage units or devices, which may include persistent and/or non-persistent memory components.
- the memory 164 stores instructions that are executable by processor 162 to perform various operations, including the instructions of various software applications and the data generated and/or used by such applications.
- memory 164 stores at least an application 170.
- application 170 is executed by processor 162 to provide one or more user interfaces via display 166, where the user interface(s) enable a user to access information resources that can include images (semantic copies) generated by computing system 102.
- application 170 may be a web browser application, and images generated by computing system 102 may be included in content slots of web pages visited by the user and presented on display 166.
- the images may be digital advertisements that are generated by computing system 102, and then selected and provided to client device 104 by computing system 102 (or by another computing system) for insertion in the content slots.
- application 170 is a dedicated application (e.g., a “mobile app”), and images generated by computing system 102 are included in content slots of user interfaces that are presented by the application 170 on display 166.
- the display 166 includes hardware, firmware, and/or software configured to enable a user to view visual outputs of the client device 104, and may use any suitable display technology (e.g., LED, OLED, LCD, etc.). In some implementations, the display 166 is incorporated in a touchscreen having both display and manual input capabilities. Moreover, in some implementations where the client device 104 is a wearable device, the display 166 is a transparent viewing component (e.g., lenses of smart glasses) with integrated electronic components. For example, the display 166 may include micro-LED or OLED electronics embedded in lenses of smart glasses.
- any suitable display technology e.g., LED, OLED, LCD, etc.
- the display 166 is incorporated in a touchscreen having both display and manual input capabilities.
- the display 166 is a transparent viewing component (e.g., lenses of smart glasses) with integrated electronic components.
- the display 166 may include micro-LED or OLED electronics embedded in lenses of smart glasses.
- the network interface 160 includes hardware, firmware, and/or software configured to enable the client device 104 to exchange electronic data with the computing system 102 via the network 110.
- the network interface 160 may include a cellular communication transceiver, a WiFi transceiver, and/or transceivers for one or more other wired and/or wireless communication technologies.
- FIG. 1 shows client device 104 as a single component communicating directly (i.e., via network 110) with the computing system 102
- the subcomponents of client device 104 shown in FIG. 1 are instead divided among two or more user-side devices.
- a pair of smart glasses may include the processor 162, the memory 164, and the display 166
- a smartphone may include another processing unit, another memory, another display, and the network interface 160.
- the smart glasses may then communicate as needed with the smartphone (e.g., via Bluetooth) to enable the operations described herein.
- the semantic copy generator 130 generally operates by obtaining a source image (e.g., by accessing a database 180, or received directly from content provider 106, etc.), and generating both semantic (textual) information and visual information based on the same source image.
- the semantic information is generated by captioner 140, and possibly prompt mutator 142, while the visual information is generated by visual embedder 146, and possibly image processor 144, as discussed in greater detail below.
- the image generator 148 then generates the new image (semantic copy) based on both the semantic information and the visual information, as is also discussed in further detail below.
- the captioner 140 generates the semantic information (caption) using the first generative Al model 150, and the image generator 148 generates the new image (semantic copy) using the second generative Al model 152.
- the prompt mutator 142 mutates the caption (or other text prompt derived from the caption) using the third generative Al model 154.
- the first generative Al model 150 is a multimodal large language model (LLM)
- the third generative Al model 154 is a finetuned LLM
- the second generative Al model 152 is a pixel diffusion model.
- the second generative Al model 152 is a latent diffusion model, a regular (non-latent) diffusion model, or another suitable type of image generation model.
- the semantic copy generator 130 utilizes reinforcement learning and/or other feedback mechanisms to improve/finetune the operation of the first generator Al model 150, the second generator Al model 152, and/or the third generator Al model 154.
- the semantic copy generator 130 obtains quality data.
- the computing system 102 may generate the quality data or obtain (e.g., receive) the quality data from another system or device, depending on the implementation.
- the quality data is stored in a quality database 184.
- the quality data may be of any format or type that is suitable to indicate performance in the desired context.
- the quality data/indicators may include manually generated scores (e.g., based on human review of images), scores generated by the computing system 102 or another system or device (e.g., based on predictive machine learning model(s)), or measured or predicted performance metrics such as click-through rate (CTR) or conversion rate (CVR), etc.
- the quality data/indicators may include, for each image, a set scores (whether manually or computer generated) that include an aesthetic score (e.g., how “professional” an image looks), a performance score (e.g., how well the image performs in the desired context), and a relevance score (e.g., how relevant the image is to information that an advertiser wishes to promote).
- Prompt mutator 142 if present in semantic copy generator 130, automatically mutates the text prompt (i.e., the caption generated by captioner 140, or other text derived from or otherwise based on the caption) using the third generative Al model 154. While there is a degree of randomness in the way the third generative Al model 154 will mutate any given text prompt, the feedback techniques discussed below in connection with FIG. 3A can greatly improve the ability of the third generative Al model 154 to modify text prompts in useful or helpful ways (i.e., in ways more likely to generate new images with good/high quality indicators).
- Image processor 144 if present in semantic copy generator 130, automatically modifies the source image in some way before the visual embedder 146 operates on the source image to create a vectorized visual embedding.
- image processor 144 may use a machine learning model to identify one or more salient regions of the source image, and crop out all other portions of the source image.
- image processor 144 may remove background objects (in order to allow the second generative Al model 152 to generate an image with a better type(s) or variety of background objects), may remove unwanted overlays such as text or buttons/controls, and/or may crop out regions of the source image to focus more on the central subject of the source image.
- Visual embedder 146 may comprise one or more embedding layers of a machine learning model (e.g., layer(s) of the second generative Al model 152 that is used to generate the new image).
- the image generator 148 operates on the visual embedding and the (possibly mutated) text prompt by applying both as inputs to the second generative Al model 152, which outputs the new image (i.e., a semantic copy of the source image).
- computing system 102 (or another system) then provides the new image to a client device (e.g., client device 104) for presentation on display via a user interface (e.g., for presentation in a user interface of application 170, on display 166).
- client device e.g., client device 104
- a user interface e.g., for presentation in a user interface of application 170, on display 166.
- the new image may be provided and presented within the digital advertising context as discussed above.
- FIG. 2 depicts an example process 200 for semantic-based copying that produces a new image 202 based on a source image 204 (e.g., from database 180).
- the process 200 may be implemented by the computing system 102 of FIG. 1 (e.g., by software instructions of semantic copy generator 130 as executed by processor 122), or by another suitable application and/or computing system.
- the process 200 is explained below with reference to elements of the example system 100 of FIG. 1.
- the captioner 140 generates descriptive text (i.e., a caption) for the source image 204, by applying the source image 204 as an input to the first generative Al model 150 (e.g., a multimodal LLM that outputs the caption for the source image 204).
- the caption output by the first generative Al model 150 is itself used as a text prompt 212 for subsequent processing, while in other implementations the text prompt 212 results from processing the caption in some way (e.g., annotating the caption with predetermined language, removing restricted language, and so on).
- the prompt mutator 142 modifies/mutates the text prompt 212 by applying the text prompt 212 as an input to the third generative Al model 154 (e.g., an LLM finetuned for prompt mutation in the desired context, such as digital advertising), which then outputs the mutated text prompt.
- the process 200 omits stage 214.
- the image processor 144 processes the source image 204 to generate a processed image 222.
- Stage 220 may include identifying salient region(s) of the source image 204, cropping salient or non-salient regions of the source image 204, and so on.
- the visual embedder 146 generates a visual embedding of the processed image 222 using one or more embedding layers. In other implementations, the process 200 omits stage 220, and the visual embedder 146 instead generates a visual embedding of the original source image 204.
- the image generator 148 generates the new image 202 by applying the text prompt 212 (with or without the mutation at stage 214, depending on the implementation), and the visual embedding from stage 224, as inputs to the second generative Al model 152, which creates/outputs the new image 202.
- the second generative Al model 152 performs both stage 224 and stage 230.
- semantic copy generator 130 may apply the source image 204 (after any post-processing at stage 220) as an input to embedding layer(s) of the second generative Al model 152, and then apply both (1) the output of the embedding layer(s) and (2) the text prompt 212 (or a mutated version of text prompt 212) as inputs to subsequent layers of the second generative Al model 152.
- FIGs. 3A and 3B depict example processes 300 and 320, respectively, that are similar to the process 200 of FIG. 2, but also employ feedback to improve performance over time.
- the processes 300 and 320 may be implemented by the computing system 102 of FIG. 1 (e.g., by software instructions of semantic copy generator 130 as executed by processor 122), or by another suitable application and/or computing system. Again, for ease of explanation, the processes 300 and 320 are explained below with reference to elements of the example system 100 of FIG. 1.
- elements with labels the same as elements in FIG. 2 may be similar to, or the same as, the like-labeled elements of FIG. 2.
- prompt mutation at stage 314 incorporates feedback based on quality indicators of new images (semantic copies) that were generated by earlier iterations of the process 300.
- semantic copy generator 130 generates or obtains a quality indicator for a given new image 202 at stage 316.
- a “quality indicator” may be a single metric or value (e.g., a CTR, or a performance score, etc.), or a set of multiple metrics or values (e.g., a CTR and a CVR, or relevancy, aesthetics, and performance scores, etc.).
- Stage 316 may include, for example, obtaining the quality indicator from another server that measures or calculates one or more components of the quality indicator. Alternatively or additionally, stage 316 may include the computing system 102 itself measuring or calculating one or more components of the quality indicator. In some implementations, the quality indicator includes one or more predicted scores and/or other values. For example, semantic copy generator 130 (or software of another computing system or device) may utilize a machine learning model, separate and distinct from the models 150, 152, 154 shown in FIG. 1, to predict a performance score for the new image 202 if the new image 202 were to be used in a digital advertisement.
- semantic copy generator 130 provides feedback to stage 314 (specifically, to finetune the third generative Al model 154) based on the quality indicator.
- semantic copy generator 130 (or a component of another computing system) applies the feedback as a part of a reinforcement learning technique with rewards and/or penalties.
- the quality indicator may include aesthetic, performance, and relevancy scores, with each being applied as a reward component.
- semantic copy generator 130 (or a component of another computing system) applies the feedback as additional samples on which to finetune the third generative Al model 154. Other types of feedback to improve the performance of the third generative Al model 154 are also possible.
- the process 300 can improve the quality or usefulness of the prompt mutations at stage 314 in future iterations, thereby improving the quality of new images output by stage 230 in those future iterations.
- the captioning at stage 310 incorporates feedback based on quality indicators of new images (semantic copies) that were generated by earlier iterations of the process 320.
- semantic copy generator 130 generates or obtains a quality indicator for a given new image 202 at stage 316.
- the quality indicator, and the manner of generating or obtaining the quality indicator may be as described above with reference to FIG. 3A, for example.
- semantic copy generator 130 provides feedback to stage 310 (specifically, to the first generative Al model 150) based on the quality indicator.
- semantic copy generator 130 (or a component of another computing system) applies the feedback as a part of a reinforcement learning technique with rewards and/or penalties.
- the quality indicator may include aesthetic, performance, and relevancy scores, with each being applied as a reward component.
- semantic copy generator 130 (or a component of another computing system) applies the feedback as additional samples on which to finetune the first generative Al model 150. Other types of feedback to improve the performance of the first generative Al model 150 are also possible.
- the process 300 can improve the quality or usefulness of the captioning at stage 310 in future iterations, thereby improving the quality of new images output by stage 230 in those future iterations.
- semantic copy generator 130 employs the feedback techniques of both process 300 and process 320 concurrently, and/or employs one or more other feedback techniques.
- semantic copy generator 130 may instead, or also, employ a similar feedback technique to improve the performance of the second generative Al model 152 used at stage 230 for image generation.
- the image generation stage 230 e.g., in process 200, 300, or 320
- the image generation stage 230 includes providing, as an input to the second generative Al model 152, an indication of a desired aspect ratio for the new image 202.
- the desired aspect ratio is input to the second generative Al model 152 separate from the (possibly mutated) text prompt, e.g., as a condition.
- semantic copy generator 130 automatically annotates or otherwise modifies the text prompt with the desired aspect ratio.
- the second generative Al model 152 can more seamlessly change the aspect ratio of the source image (e.g., source image 204).
- the aspect ratio may seamlessly be changed from portrait to landscape or vice versa (e.g., without positioning objects in an aesthetically displeasing way due to the format change, and/or without stressing or minimizing features of the new image in a way that makes the new image perform poorly, etc.).
- FIG. 4 depicts example semantic copies (e.g., corresponding to new image 202) that may be produced by the computing system 102 (e.g., by semantic copy generator 130), using process 200 of FIG. 2, process 300 of FIG. 3A, or process 320 of FIG. 3B, for example.
- the source image is shown on the left, and its corresponding new image is shown on the right.
- each new image modifies its corresponding source image in significant respects, while nonetheless adhering closely to certain visual qualities of the source image.
- a semantic copy of a source image that is a digital advertisement for a company may maintain visual qualities (style, brand colors, etc.) that are associated with that company or its advertisements.
- FIG. 5 depicts other examples semantic copies that may be produced by the computing system 102 (e.g., by semantic copy generator 130), in an alternative implementation in which the new image is generated using a text prompt but no visual embedding (e.g., with the top path of process 200, 300, or 320, but not the respective bottom path).
- the new image may be inferior to that of process 200, 300, or 320, with the new image failing to maintain key visual similarities with the source image.
- the top right semantic copy in FIG. 5 is thematically similar to the top left source image (with both showing a home with some yard and sky), the top right image has a different overall tone, and depicts a very different style of home, etc.
- the bottom right semantic copy in FIG. 5 is thematically similar to the bottom left source image (with both showing elements of a bathroom), the bottom right image has a different overall tone, and depicts very different elements within a bathroom, etc.
- FIG. 6 is a flow diagram of an example method 600 for semantic-based copying of a source image.
- the method 600 may be implemented by the computing system 102 (e.g., semantic copy generator 130) of FIG. 1, for example.
- Block 602 includes applying a source image to a first generative Al model (e.g., first generative Al model 150) to generate a descriptive caption for the source image.
- a first generative Al model e.g., first generative Al model 150
- Block 602 may be similar to stage 210 of process 200 or 300, or stage 310 of process 320, for example.
- Block 604 a visual embedding is generated based on the source image.
- Block 604 may be similar to stage 224 of process 200, 300, or 320, for example.
- a new image is generated using a second generative Al model (e.g., second generative Al model 152) and based on the text prompt and the visual embedding.
- Block 606 may be similar to stage 230 of process 200, 300, or 320, for example.
- block 606 includes generating a mutated text prompt by applying the text prompt to a third generative Al model (e.g., third generative Al model 154), and then applying the mutated text prompt and the visual embedding to the second generative Al model.
- block 606 may be similar to the combination of stages 214 and 230 in FIG. 2 or 3B, or the combination of stages 314 and 230 in FIG. 3A.
- the method 600 may include iterations for multiple new images, and may include feedback mechanisms.
- the method 600 may include generating a plurality of new images using the second generative Al model and based on the same text prompt and visual embedding, and generating these new images may include generating mutated text prompts by applying the text prompt to the third generative Al model, and generating each of the new images by applying a respective one of the mutated text prompts to the second generative Al model.
- the method 600 may include one or more additional blocks not shown in FIG. 6.
- the method 600 may include a first additional block in which quality indicators are generated or obtained, with each quality indicator corresponding to a respective one of the new images mentioned above, and a second additional block in which the third generative Al model is finetuned (e.g., with reinforcement learning or other techniques) based on the quality indicators.
- the blocks of FIG. 6 need not be performed strictly in the order shown. For example, block 604 may occur before, or in parallel with, block 602.
- Artificial intelligence is a segment of computer science that focuses on the creation of models that can perform tasks with little to no human intervention.
- Artificial intelligence systems can utilize, for example, machine learning, natural language processing, and computer vision.
- Natural language processing focuses on analyzing and generating human language.
- Computer vision focuses on analyzing and interpreting images and videos.
- Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and/or other content, in response to input prompts and/or based on other information.
- Example machine-learned models include neural networks or other multi-layer nonlinear models.
- Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.
- Some example machine-learned models can leverage an attention mechanism such as self-attention.
- some machine-learned models can include multi-headed self-attention models (e.g., transformer models).
- the model(s) can be trained using various training or learning techniques.
- the training can implement supervised learning, unsupervised learning, reinforcement learning, etc.
- the training can use techniques such as, for example, backwards propagation of errors.
- a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function).
- Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and/or various other loss functions.
- Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.
- a number of generalization techniques e.g., weight decays, dropouts
- the model(s) can be pre-trained before domain- specific alignment. For instance, a model can be pretrained over a general corpus of training data and finetuned on a more targeted corpus of training data. A model can be aligned using prompts that are designed to elicit domain- specific outputs. Prompts can be designed to include learned prompt values (e.g., soft prompts).
- the trained model(s) may be validated prior to their use using input data other than the training data, and may be further updated or refined during their use based on additional feedback/inputs.
- the computing system 102 may use one or more of the machine learning models or techniques noted above to perform any one or more of the operations discussed herein in connection with machine learning.
- the computing system 102 may use one or more such machine learning techniques to pre-train and/or finetune the first generative Al model 150, the second generative Al model 152, and/or the third generative Al model 154, and possibly to pre-train and/or finetune a model that predicts performance of an image (e.g., to generate quality indicators as discussed above), etc.
- Example 1 A method of semantic-based image copying, the method comprising: generating, by one or more processors, a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image; generating, by the one or more processors, a visual embedding based on the source image; and generating, by the one or more processors, a new image using a second generative Al model and based on the text prompt and the visual embedding.
- Al generative artificial intelligence
- Example 2 The method of Example 1, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
- Example 3. The method of Example 2, comprising: generating, by the one or more processors, a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
- Example 4 The method of Example 3, comprising: generating or obtaining, by the one or more processors, a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetuning, by the one or more processors, the third generative Al model based on the plurality of quality indicators.
- Example 5 The method of Example 1, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
- Example 6 The method of any one of Examples 1-5, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
- Example 7 The method of any one of Examples 1-6, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
- Example 8 The method of Example 7, wherein processing the source image includes: cropping out a portion of the source image.
- Example 9 The method of any one of Examples 1-8, wherein the text prompt is the descriptive caption.
- Example 10 The method of any one of Examples 1-9, wherein the second generative Al model is a pixel diffusion model.
- Example 11 A system comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to (i) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image, (ii) generate a visual embedding based on the source image, and (iii) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
- Example 12 The system of Example 11, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
- Example 13 The system of Example 12, wherein the instructions cause the one or more processors to: generate a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
- Example 14 The system of Example 13, wherein the instructions cause the one or more processors to: generate or obtain a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetune the third generative Al model based on the plurality of quality indicators.
- Example 15 The system of Example 11, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
- Example 17 The system of any one of Examples 11-16, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
- Example 18 The system of any one of Examples 11-17, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
- Example 19 The system of Example 18, wherein processing the source image includes: cropping out a portion of the source image.
- Example 20 The system of any one of Examples 11-19, wherein the text prompt is the descriptive caption.
- Example 21 The system of any one of Examples 11-20, wherein the second generative Al model is a pixel diffusion model.
- Example 22 One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image; generate a visual embedding based on the source image; and generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
- Al generative artificial intelligence
- Example 23 The one or more non-transitory, computer-readable media of Example 22, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
- Example 24 The one or more non-transitory, computer-readable media of Example 23, wherein the instructions cause the one or more processors to: generate a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
- Example 25 The one or more non-transitory, computer-readable media of Example 24, wherein the instructions cause the one or more processors to: generate or obtain a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetune the third generative Al model based on the plurality of quality indicators.
- Example 26 The one or more non-transitory, computer-readable media of Example 22, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
- Example 27 The one or more non-transitory, computer-readable media of any one of Examples 22-26, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
- Example 28 The one or more non-transitory, computer-readable media of any one of Examples 22-27, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
- Example 29 The one or more non-transitory, computer-readable media of Example 28, wherein processing the source image includes: cropping out a portion of the source image.
- Example 30 The one or more non-transitory, computer-readable media of any one of Examples 22-29, wherein the second generative Al model is a pixel diffusion model.
- Example 31 The one or more non-transitory, computer-readable media of any one of Examples 22-30, wherein the text prompt is the descriptive caption.
- “generating, by one or more processors, X; and generating, by the one or more processors, Y” can encompass: (1) implementations in which a first set of one or more processors (e.g., in a first computing device) generates X and a distinct, second set of one or more processors (e.g., in a different, second computing device) independently generates Y ; (2) implementations in which all processors in the set of one or more processors (e.g., all in the same device, or distributed among multiple devices) contribute to the generation of both X and Y ; and (3) other variations.
- “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
- a machine e.g., a computer
- memories e.g., volatile memory, non-volatile memory, or a combination thereof
- registers e.g., volatile memory, non-volatile memory, or a combination thereof
- any reference to “one implementation” or “an implementation” means that a particular element, feature, structure, or characteristic described in connection with the implementation is included in at least one implementation or implementation.
- the appearances of the phrase “in one implementation” in various places in the specification are not necessarily all referring to the same implementation.
- the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion.
- a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus.
- “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
Landscapes
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Image Processing (AREA)
Abstract
A method of semantic-based image copying includes generating a text prompt. Generating the text prompt includes by applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image. The method also includes generating a visual embedding based on the source image, and generating a new image using a second generative Al model and based on the text prompt and the visual embedding.
Description
SEMANTIC-BASED IMAGE COPYING
FIELD OF TECHNOLOGY
[0001] The present disclosure relates to techniques for copying images, using generative artificial intelligence, in a manner that creatively changes the images while maintaining semantic context and visual qualities of the source images.
BACKGROUND
[0002] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventor(s), to the extent it is described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.
[0003] In recent years, significant progress has been made in the field of image generation and image modification. In particular, generative artificial intelligence (Al) models have begun to find widespread use in both personal and commercial domains. In some cases, it is desirable to generate a non-exact “copy” of a source image, i.e., an image that is in some ways similar to or inspired by the source image, but creatively different than the source image. In digital advertising, for example, an advertiser may wish to expand from a single creative/asset/image to a number of new creatives/assets/images, e.g., in order to expand the advertiser’s catalog of available advertisements and/or improve the performance of existing advertisements. However, generative Al “copies” often fail to strike a good balance between: (1) maintaining visual properties of the source image (e.g., maintaining a look and feel of the source image); (2) maintaining the semantic context of the source image (e.g., maintaining what is shown, at a descriptive/conceptual level, in the source image); and yet (3) differing from the source image in a meaningful way (e.g., such that the copy creatively differs from the source image and is not merely a trivial variation of the source image). In the digital advertising context, for example, generative Al copies of an original creative may fail to capture a look and feel associated with a particular advertiser (e.g., as represented in the source image by colors, styles, etc.), may fail to represent an advertised product accurately, or may fail to differ from the original creative enough to significantly differ in perform (e.g., have a different level of customer/user engagement) as compared to the original creative.
SUMMARY
[0004] In the disclosed techniques, a system generates new images by semantically copying source images. As the term is used herein, “copying” a source image generally refers to generating a new image that is in some way derived from the source image, but without being an exact copy of the source image, and without necessarily (but possibly) reusing or reproducing any portion/pixels of the source image. The system of the present disclosure generates a new image from (i.e., generates a semantic copy of) a source image by: (1) generating a text prompt by applying the source image to a first generative artificial intelligence (Al) model that generates a descriptive caption for the source image; (2) generating a visual embedding based on the source image (e.g., a visual embedding of either the source image or a processed version of the source image); and (3) generating the new image using a second generative Al model and based on both the text prompt and the visual embedding.
[0005] By generating the new image based on a text prompt that is itself a descriptive caption generated directly from the source image (or at least, a text prompt that is derived from such a caption), the system provides a “semantic” copy of the source image. That is, the new image adheres well to the source image at a conceptual/descriptive level. Moreover, because the system also uses a visual embedding of the source image to generate the new image, the new image adheres well to visual properties of the source image. Furthermore, a text prompt derived from the source image itself is less likely to conflict with a visual embedding of the source image (e.g., as compared to a text prompt that is manually generated), and thus less likely to produce a bizarre, confusing, or unappealing new image.
[0006] Furthermore, concurrent use of the visual embedding, rather than the text prompt alone, makes it more likely that the new image will depict things that are truly relevant to the source image. That is, the visual embedding helps the second (image generation) generative Al model avoid hallucinations with respect to what should be shown at a conceptual level. For instance, if the text prompt “a cell phone mounted on a car” is derived from a source image that shows a cell phone mounted on a dashboard of a car, the concurrent use of the visual embedding of the source image may prevent the generative Al model from generating, for example, a new image in which a giant cell phone is mounted on top of and external to a car. Thus, the visual embedding derived from the source image helps to ground the generative Al model. Used concurrently, the text prompt and the visual embedding can help
ensure that generated images maintain the key concept underlying the source image, while also introducing varieties that may have been introduced via text prompt mutation and/or image processing steps prior to visual embedding (e.g., cropping)
[0007] The disclosed techniques also provide other technical advantages. One advantage stems from the fact that the text prompt contains information that is not, in a strict sense, within the source image itself. As a result, the text prompt, despite being derived from the source image, provides a higher level of entropy than would exist if using visual embeddings without the text prompt. That is, the text prompt provides the second generative Al model with a broader creative/imaginary space in which to create the new image.
[0008] As another example, converting the source image to a text prompt, using the first (captioning) generative Al model, facilitates changes to the source image, as compared to making such changes based solely on a visual embedding or through other image processing means. Because generative Al models (e.g., large language models) are relatively adept at making modifications (e.g., adding, removing, or changing the state or position of objects) in the semantic space, and because the disclosed techniques convert the source image to a text description/caption, such modifications can be made efficiently and with relative ease. In some implementations, the disclosed system uses a third generative Al model to make such modifications, by modifying or “mutating” the text prompt before the system applies the text prompt (along with the visual embedding) to the second generative Al model.
[0009] Other advantages will also become apparent to one of ordinary skill in the art upon reading this disclosure and viewing the corresponding drawings.
[0010] In one aspect, a method of semantic-based image copying includes generating, by one or more processors, a text prompt. Generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image. The method also includes generating, by the one or more processors, a visual embedding based on the source image, and generating, by the one or more processors, a new image using a second generative Al model and based on the text prompt and the visual embedding.
[0011] In another aspect, a system includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to: (1) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image; (2) generate a visual embedding based on the source image; and
(3) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
[0012] In another aspect, one or more non-transitory, computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to: (1) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative Al model to generate a descriptive caption for the source image; (2) generate a visual embedding based on the source image; and (3) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIG. 1 is a block diagram of an example system in which techniques for semanticbased copying of images can be implemented.
[0014] FIG. 2 depicts an example process for semantic-based copying of a source image, which may be implemented by the computing system of FIG. 1.
[0015] FIGs. 3A and 3B depict example processes for semantic-based copying of a source image with feedback, which may be implemented by the computing system of FIG. 1.
[0016] FIG. 4 depicts examples of semantic copies that may be produced by the computing system of FIG.1, using the process of FIG. 2, 3A, or 3B.
[0017] FIG. 5 depicts examples of copies that may be produced by an alternative implementation in which the new image is generated using a text prompt but no visual embedding.
[0018] FIG. 6 is a flow diagram of an example method for semantic-based copying of a source image.
DETAILED DESCRIPTION OF THE DRAWINGS
[0019] FIG. 1 is a block diagram of an example system 100 in which techniques for semantic-based image copying/generation can be implemented. The example system 100 includes a computing system 102, a client device 104, a content provider 106 (e.g., a server of a content provider), and a network 110. The computing system 102 is remote from the client device 104 and content provider 106, and is communicatively coupled to the client device 104 and content provider 106 via the network 110. In some implementations, the system 100 does not include client device 104 and/or content provider 106.
[0020] The network 110 may be a single communication network (e.g., the Internet), and in some implementations also includes one or more additional networks. As just one example, the network 110 may include a cellular network, the Internet, and a server- side local area network (LAN). While FIG. 1 shows only a single client device 104 and single content provider 106, it is understood that the computing system 102 may also be in communication with a number (e.g., millions) of other client devices that are generally similar to the client device 104, and/or in communication with a number (e.g., thousands) of other content providers that are generally similar to content provider 106.
[0021] Generally, computing system 102 can perform image copying/generation services (e.g., for providers such as content provider 106). As noted above, the term “copying” is generally used herein to refer to generating a new image that is in some way derived from the source image, without being an exact copy of the source image, and without necessarily (but possibly) reusing or reproducing any portion/pixels of the source image.
[0022] In a digital advertising or marketing context, for example, computing system 102 may use existing images from content providers such as content provider 106 to generate new images that the content provider can use in additional digital advertising. In one such example, the new/additional images can be used to provide a greater diversity of images/advertisements, the performance of which can then be measured (e.g., based on click- through rate, conversion rate, etc.) to determine which images/advertisements are most effective. As another digital advertising example, the new/additional images may have aspect ratios different from the original image, making the new images better suited to ad slots (e.g., in a web page or mobile application) that have different aspect ratio constraints. Notably, the techniques described herein (e.g., in connection with FIGs. 2, 3A, 3B, and 6) can change the aspect ratio of the source image in a more seamless manner than conventional techniques (e.g., salient region detection plus cropping).
[0023] As another example, computing system 102 may generate new images/copies that are intended to facilitate viewer understanding (e.g., images for instructional materials), where performance is measured by way of determining what proportion of viewers take the correct actions upon viewing the images. Other contexts are also possible. For ease and consistency of explanation, however, this disclosure primarily uses examples that are related to a digital advertising implementation/context.
[0024] The client device 104 is generally configured to access information resources (e.g., web pages and/or user interfaces of mobile applications or other applications) that can present the images generated by computing system 102. For example, computing system 102 may generate digital advertisements that include (or consist entirely of) the semantic copies discussed herein. Computing system 102 or another computing system may then serve the digital advertisements to users of client device 104 and/or other similar client devices using suitable techniques, such as conducting auctions (e.g., auctions based on keyword bids by advertisers, relevancy metrics, etc.). The digital advertisements may be served in slots of web pages visited by the users, and/or slots of application user interfaces displayed to the users, etc.
[0025] The content provider 106 generally may commission or request that computing system 102 generate one or more images, and/or may provide the source image(s) upon which the image generation is based. For example, content provider 106 may be a digital advertiser who provides a digital advertisement image for each of a number of offered products or services, as part of one or more advertising campaigns owned or managed by content provider 106. As other examples, the source image may be a screenshot of a web page hosted by content provider 106, a screenshot of a mobile application that content provider 106 offers, and so on.
[0026] The computing system 102 includes a network interface 120, a processor 122, and memory 124. The network interface 120 includes hardware, firmware, and/or software configured to enable the computing system 102 to exchange electronic data with the client device 104 and other, similar client devices (and possibly content provider 106, etc.) via the network 110. For example, the network interface 120 may include a wired or wireless router and a modem. The processor 122 may be a single processor (e.g., a central processing unit (CPU)), or may include multiple processors (e.g., multiple CPUs, or one or more CPUs and one or more graphics processing units (GPUs)). Computing system 102 may be a single computing device (e.g., server) at a single location, or may include multiple, coordinating computing devices that are either co-located or remotely distributed.
[0027] The memory 124 is a computer-readable, non-transitory storage unit or device, or collection of such units/devices, that may include persistent and/or non-persistent memory components. The memory 124 stores instructions executable by processor 122 to perform various operations, including the instructions of various software applications and the data
generated and/or used by such applications. In the example system 100 of FIG. 1, memory 124 stores the instructions of a semantic copy generator 130, which includes a captioner 140, a prompt mutator 142, an image processor 144, a visual embedder 146, and an image generator 148.
[0028] Memory 124 can also store generative artificial intelligence (Al) models. In particular, in the example system 100 of FIG. 1, memory 124 stores a first generative Al model 150 used by captioner 140, a second generative Al model 152 used by image generator 148, and a third generative Al model 154 used by prompt mutator 142. In other implementations, third generative Al model 154 is not included in system 100. More generally, it is understood that, in some implementations, memory 124 may omit one or more modules/elements shown in FIG. 1, such as prompt mutator 142 and/or image processor 144. It is also understood that, in some implementations, memory 124 may include one or more additional modules/elements not shown in FIG. 1, such as modules that facilitate serving images (e.g., digital advertisements) to users of devices such as client device 104. In some implementations, first generative Al model 150, second generative Al model 152, and/or third generative Al model 154 are not stored in memory 124, and instead are stored in one or more remote servers or other computing systems. For example, one or more of models 150, 152, and 154 may be remotely accessed (e.g., as a cloud service) by semantic copy generator 130.
[0029] The client device 104 may be or include any stationary, mobile, or portable computing device with wired and/or wireless communication capability (e.g., a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart wearable device such as smart glasses or a smart watch, a vehicle head unit computer, etc.). In the example implementation of FIG. 1, client device 104 includes a network interface 160, a processor 162, memory 164, and a display 166. The processor 162 may be a single processor, or may include multiple processors.
[0030] Memory 164 includes one or more computer-readable, non-transitory storage units or devices, which may include persistent and/or non-persistent memory components. The memory 164 stores instructions that are executable by processor 162 to perform various operations, including the instructions of various software applications and the data generated and/or used by such applications.
[0031] In the example system 100 of FIG. 1, memory 164 stores at least an application 170. Generally, application 170 is executed by processor 162 to provide one or more user interfaces via display 166, where the user interface(s) enable a user to access information resources that can include images (semantic copies) generated by computing system 102. For example, application 170 may be a web browser application, and images generated by computing system 102 may be included in content slots of web pages visited by the user and presented on display 166. As a more specific example, the images may be digital advertisements that are generated by computing system 102, and then selected and provided to client device 104 by computing system 102 (or by another computing system) for insertion in the content slots. In other implementations, application 170 is a dedicated application (e.g., a “mobile app”), and images generated by computing system 102 are included in content slots of user interfaces that are presented by the application 170 on display 166.
[0032] The display 166 includes hardware, firmware, and/or software configured to enable a user to view visual outputs of the client device 104, and may use any suitable display technology (e.g., LED, OLED, LCD, etc.). In some implementations, the display 166 is incorporated in a touchscreen having both display and manual input capabilities. Moreover, in some implementations where the client device 104 is a wearable device, the display 166 is a transparent viewing component (e.g., lenses of smart glasses) with integrated electronic components. For example, the display 166 may include micro-LED or OLED electronics embedded in lenses of smart glasses.
[0033] The network interface 160 includes hardware, firmware, and/or software configured to enable the client device 104 to exchange electronic data with the computing system 102 via the network 110. For example, the network interface 160 may include a cellular communication transceiver, a WiFi transceiver, and/or transceivers for one or more other wired and/or wireless communication technologies.
[0034] While FIG. 1 shows client device 104 as a single component communicating directly (i.e., via network 110) with the computing system 102, in some implementations the subcomponents of client device 104 shown in FIG. 1 are instead divided among two or more user-side devices. As just one example, a pair of smart glasses may include the processor 162, the memory 164, and the display 166, while a smartphone may include another processing unit, another memory, another display, and the network interface 160. The smart
glasses may then communicate as needed with the smartphone (e.g., via Bluetooth) to enable the operations described herein.
[0035] Returning to the computing system 102, the semantic copy generator 130 generally operates by obtaining a source image (e.g., by accessing a database 180, or received directly from content provider 106, etc.), and generating both semantic (textual) information and visual information based on the same source image. The semantic information is generated by captioner 140, and possibly prompt mutator 142, while the visual information is generated by visual embedder 146, and possibly image processor 144, as discussed in greater detail below. The image generator 148 then generates the new image (semantic copy) based on both the semantic information and the visual information, as is also discussed in further detail below. The captioner 140 generates the semantic information (caption) using the first generative Al model 150, and the image generator 148 generates the new image (semantic copy) using the second generative Al model 152. In implementations that support prompt mutation, the prompt mutator 142 mutates the caption (or other text prompt derived from the caption) using the third generative Al model 154. In some implementations, the first generative Al model 150 is a multimodal large language model (LLM), the third generative Al model 154 is a finetuned LLM, and the second generative Al model 152 is a pixel diffusion model. In other implementations, the second generative Al model 152 is a latent diffusion model, a regular (non-latent) diffusion model, or another suitable type of image generation model.
[0036] In some implementations, as discussed in further detail below, the semantic copy generator 130 utilizes reinforcement learning and/or other feedback mechanisms to improve/finetune the operation of the first generator Al model 150, the second generator Al model 152, and/or the third generator Al model 154. In such implementations, the semantic copy generator 130 obtains quality data. The computing system 102 may generate the quality data or obtain (e.g., receive) the quality data from another system or device, depending on the implementation. In the example system 100 of FIG. 1, the quality data is stored in a quality database 184. The quality data may be of any format or type that is suitable to indicate performance in the desired context. In a digital advertising context, for example, the quality data/indicators may include manually generated scores (e.g., based on human review of images), scores generated by the computing system 102 or another system or device (e.g., based on predictive machine learning model(s)), or measured or predicted performance metrics such as click-through rate (CTR) or conversion rate (CVR), etc. As a more specific
example, the quality data/indicators may include, for each image, a set scores (whether manually or computer generated) that include an aesthetic score (e.g., how “professional” an image looks), a performance score (e.g., how well the image performs in the desired context), and a relevance score (e.g., how relevant the image is to information that an advertiser wishes to promote).
[0037] Prompt mutator 142, if present in semantic copy generator 130, automatically mutates the text prompt (i.e., the caption generated by captioner 140, or other text derived from or otherwise based on the caption) using the third generative Al model 154. While there is a degree of randomness in the way the third generative Al model 154 will mutate any given text prompt, the feedback techniques discussed below in connection with FIG. 3A can greatly improve the ability of the third generative Al model 154 to modify text prompts in useful or helpful ways (i.e., in ways more likely to generate new images with good/high quality indicators).
[0038] Image processor 144, if present in semantic copy generator 130, automatically modifies the source image in some way before the visual embedder 146 operates on the source image to create a vectorized visual embedding. As one example, image processor 144 may use a machine learning model to identify one or more salient regions of the source image, and crop out all other portions of the source image. For example, image processor 144 may remove background objects (in order to allow the second generative Al model 152 to generate an image with a better type(s) or variety of background objects), may remove unwanted overlays such as text or buttons/controls, and/or may crop out regions of the source image to focus more on the central subject of the source image. Visual embedder 146 may comprise one or more embedding layers of a machine learning model (e.g., layer(s) of the second generative Al model 152 that is used to generate the new image).
[0039] The image generator 148 operates on the visual embedding and the (possibly mutated) text prompt by applying both as inputs to the second generative Al model 152, which outputs the new image (i.e., a semantic copy of the source image). In some implementations, computing system 102 (or another system) then provides the new image to a client device (e.g., client device 104) for presentation on display via a user interface (e.g., for presentation in a user interface of application 170, on display 166). For example, the new image may be provided and presented within the digital advertising context as discussed above.
[0040] FIG. 2 depicts an example process 200 for semantic-based copying that produces a new image 202 based on a source image 204 (e.g., from database 180). The process 200 may be implemented by the computing system 102 of FIG. 1 (e.g., by software instructions of semantic copy generator 130 as executed by processor 122), or by another suitable application and/or computing system. For ease of explanation, the process 200 is explained below with reference to elements of the example system 100 of FIG. 1.
[0041] At stage 210, in one path of the process 200, the captioner 140 generates descriptive text (i.e., a caption) for the source image 204, by applying the source image 204 as an input to the first generative Al model 150 (e.g., a multimodal LLM that outputs the caption for the source image 204). In some implementations, the caption output by the first generative Al model 150 is itself used as a text prompt 212 for subsequent processing, while in other implementations the text prompt 212 results from processing the caption in some way (e.g., annotating the caption with predetermined language, removing restricted language, and so on).
[0042] At stage 214, the prompt mutator 142 modifies/mutates the text prompt 212 by applying the text prompt 212 as an input to the third generative Al model 154 (e.g., an LLM finetuned for prompt mutation in the desired context, such as digital advertising), which then outputs the mutated text prompt. In alternative implementations, the process 200 omits stage 214.
[0043] At stage 220, in another path of the process 200, the image processor 144 processes the source image 204 to generate a processed image 222. Stage 220 may include identifying salient region(s) of the source image 204, cropping salient or non-salient regions of the source image 204, and so on. At stage 224, the visual embedder 146 generates a visual embedding of the processed image 222 using one or more embedding layers. In other implementations, the process 200 omits stage 220, and the visual embedder 146 instead generates a visual embedding of the original source image 204.
[0044] At stage 230, the image generator 148 generates the new image 202 by applying the text prompt 212 (with or without the mutation at stage 214, depending on the implementation), and the visual embedding from stage 224, as inputs to the second generative Al model 152, which creates/outputs the new image 202. In some implementations, the second generative Al model 152 performs both stage 224 and stage 230. For example, semantic copy generator 130 may apply the source image 204 (after any post-processing at
stage 220) as an input to embedding layer(s) of the second generative Al model 152, and then apply both (1) the output of the embedding layer(s) and (2) the text prompt 212 (or a mutated version of text prompt 212) as inputs to subsequent layers of the second generative Al model 152.
[0045] FIGs. 3A and 3B depict example processes 300 and 320, respectively, that are similar to the process 200 of FIG. 2, but also employ feedback to improve performance over time. Like the process 200 of FIG. 2, the processes 300 and 320 may be implemented by the computing system 102 of FIG. 1 (e.g., by software instructions of semantic copy generator 130 as executed by processor 122), or by another suitable application and/or computing system. Again, for ease of explanation, the processes 300 and 320 are explained below with reference to elements of the example system 100 of FIG. 1.
[0046] In FIGs. 3A and 3B, elements with labels the same as elements in FIG. 2 (e.g., images 202 and/or 204, and/or stages 210, 220, etc.) may be similar to, or the same as, the like-labeled elements of FIG. 2. In FIG. 3A, however, prompt mutation at stage 314 incorporates feedback based on quality indicators of new images (semantic copies) that were generated by earlier iterations of the process 300. In particular, in the process 300, semantic copy generator 130 generates or obtains a quality indicator for a given new image 202 at stage 316. A “quality indicator” may be a single metric or value (e.g., a CTR, or a performance score, etc.), or a set of multiple metrics or values (e.g., a CTR and a CVR, or relevancy, aesthetics, and performance scores, etc.).
[0047] Stage 316 may include, for example, obtaining the quality indicator from another server that measures or calculates one or more components of the quality indicator. Alternatively or additionally, stage 316 may include the computing system 102 itself measuring or calculating one or more components of the quality indicator. In some implementations, the quality indicator includes one or more predicted scores and/or other values. For example, semantic copy generator 130 (or software of another computing system or device) may utilize a machine learning model, separate and distinct from the models 150, 152, 154 shown in FIG. 1, to predict a performance score for the new image 202 if the new image 202 were to be used in a digital advertisement.
[0048] In any event, semantic copy generator 130 provides feedback to stage 314 (specifically, to finetune the third generative Al model 154) based on the quality indicator. In some implementations, semantic copy generator 130 (or a component of another computing
system) applies the feedback as a part of a reinforcement learning technique with rewards and/or penalties. For example, the quality indicator may include aesthetic, performance, and relevancy scores, with each being applied as a reward component. In other implementations, semantic copy generator 130 (or a component of another computing system) applies the feedback as additional samples on which to finetune the third generative Al model 154. Other types of feedback to improve the performance of the third generative Al model 154 are also possible. By providing feedback as shown in FIG. 3A, the process 300 can improve the quality or usefulness of the prompt mutations at stage 314 in future iterations, thereby improving the quality of new images output by stage 230 in those future iterations.
[0049] In process 320 of FIG. 3B, the captioning at stage 310 incorporates feedback based on quality indicators of new images (semantic copies) that were generated by earlier iterations of the process 320. In particular, in the process 320, semantic copy generator 130 generates or obtains a quality indicator for a given new image 202 at stage 316. The quality indicator, and the manner of generating or obtaining the quality indicator, may be as described above with reference to FIG. 3A, for example. In the process 320, semantic copy generator 130 provides feedback to stage 310 (specifically, to the first generative Al model 150) based on the quality indicator. In some implementations, semantic copy generator 130 (or a component of another computing system) applies the feedback as a part of a reinforcement learning technique with rewards and/or penalties. For example, the quality indicator may include aesthetic, performance, and relevancy scores, with each being applied as a reward component. In other implementations, semantic copy generator 130 (or a component of another computing system) applies the feedback as additional samples on which to finetune the first generative Al model 150. Other types of feedback to improve the performance of the first generative Al model 150 are also possible. By providing feedback as shown in FIG. 3B, the process 300 can improve the quality or usefulness of the captioning at stage 310 in future iterations, thereby improving the quality of new images output by stage 230 in those future iterations.
[0050] In some implementations, semantic copy generator 130 employs the feedback techniques of both process 300 and process 320 concurrently, and/or employs one or more other feedback techniques. For example, semantic copy generator 130 may instead, or also, employ a similar feedback technique to improve the performance of the second generative Al model 152 used at stage 230 for image generation.
[0051] In some implementations, the image generation stage 230 (e.g., in process 200, 300, or 320) includes providing, as an input to the second generative Al model 152, an indication of a desired aspect ratio for the new image 202. In some of these implementations, the desired aspect ratio is input to the second generative Al model 152 separate from the (possibly mutated) text prompt, e.g., as a condition. In other implementations, semantic copy generator 130 automatically annotates or otherwise modifies the text prompt with the desired aspect ratio. In either case, by using the disclosed techniques (and in particular by jointly using both a semantic/text path and a visual embedding path), the second generative Al model 152 can more seamlessly change the aspect ratio of the source image (e.g., source image 204). For example, the aspect ratio may seamlessly be changed from portrait to landscape or vice versa (e.g., without positioning objects in an aesthetically displeasing way due to the format change, and/or without stressing or minimizing features of the new image in a way that makes the new image perform poorly, etc.).
[0052] FIG. 4 depicts example semantic copies (e.g., corresponding to new image 202) that may be produced by the computing system 102 (e.g., by semantic copy generator 130), using process 200 of FIG. 2, process 300 of FIG. 3A, or process 320 of FIG. 3B, for example. In FIG. 4, the source image is shown on the left, and its corresponding new image is shown on the right. As seen in FIG. 4, each new image modifies its corresponding source image in significant respects, while nonetheless adhering closely to certain visual qualities of the source image. Thus, for example, a semantic copy of a source image that is a digital advertisement for a company may maintain visual qualities (style, brand colors, etc.) that are associated with that company or its advertisements.
[0053] FIG. 5 depicts other examples semantic copies that may be produced by the computing system 102 (e.g., by semantic copy generator 130), in an alternative implementation in which the new image is generated using a text prompt but no visual embedding (e.g., with the top path of process 200, 300, or 320, but not the respective bottom path). As seen in FIG. 5, however, such an approach may be inferior to that of process 200, 300, or 320, with the new image failing to maintain key visual similarities with the source image. For example, while the top right semantic copy in FIG. 5 is thematically similar to the top left source image (with both showing a home with some yard and sky), the top right image has a different overall tone, and depicts a very different style of home, etc. Similarly, while the bottom right semantic copy in FIG. 5 is thematically similar to the bottom left
source image (with both showing elements of a bathroom), the bottom right image has a different overall tone, and depicts very different elements within a bathroom, etc.
[0054] FIG. 6 is a flow diagram of an example method 600 for semantic-based copying of a source image. The method 600 may be implemented by the computing system 102 (e.g., semantic copy generator 130) of FIG. 1, for example.
[0055] At block 602, a text prompt is generated. Block 602 includes applying a source image to a first generative Al model (e.g., first generative Al model 150) to generate a descriptive caption for the source image. Block 602 may be similar to stage 210 of process 200 or 300, or stage 310 of process 320, for example.
[0056] At block 604, a visual embedding is generated based on the source image. Block 604 may be similar to stage 224 of process 200, 300, or 320, for example.
[0057] At block 606, a new image is generated using a second generative Al model (e.g., second generative Al model 152) and based on the text prompt and the visual embedding. Block 606 may be similar to stage 230 of process 200, 300, or 320, for example. In some implementations, block 606 includes generating a mutated text prompt by applying the text prompt to a third generative Al model (e.g., third generative Al model 154), and then applying the mutated text prompt and the visual embedding to the second generative Al model. For example, block 606 may be similar to the combination of stages 214 and 230 in FIG. 2 or 3B, or the combination of stages 314 and 230 in FIG. 3A.
[0058] The method 600 may include iterations for multiple new images, and may include feedback mechanisms. For example, the method 600 may include generating a plurality of new images using the second generative Al model and based on the same text prompt and visual embedding, and generating these new images may include generating mutated text prompts by applying the text prompt to the third generative Al model, and generating each of the new images by applying a respective one of the mutated text prompts to the second generative Al model.
[0059] The method 600 may include one or more additional blocks not shown in FIG. 6. For example, the method 600 may include a first additional block in which quality indicators are generated or obtained, with each quality indicator corresponding to a respective one of the new images mentioned above, and a second additional block in which the third generative Al model is finetuned (e.g., with reinforcement learning or other techniques) based on the quality indicators.
[0060] It is understood that the blocks of FIG. 6 need not be performed strictly in the order shown. For example, block 604 may occur before, or in parallel with, block 602.
[0061] As is apparent from the above description, techniques disclosed herein use artificial intelligence to generate high-performing images. Artificial intelligence (Al) is a segment of computer science that focuses on the creation of models that can perform tasks with little to no human intervention. Artificial intelligence systems can utilize, for example, machine learning, natural language processing, and computer vision. Machine learning, and its subsets, such as deep learning, focus on developing models that can infer outputs from data. The outputs can include, for example, predictions and/or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and/or other content, in response to input prompts and/or based on other information.
[0062] Example machine-learned models include neural networks or other multi-layer nonlinear models. Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some machine-learned models can include multi-headed self-attention models (e.g., transformer models).
[0063] The model(s) can be trained using various training or learning techniques. The training can implement supervised learning, unsupervised learning, reinforcement learning, etc. The training can use techniques such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and/or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations. A number of generalization techniques (e.g., weight decays, dropouts) can be used to improve the generalization capability of the models being trained.
[0064] The model(s) can be pre-trained before domain- specific alignment. For instance, a model can be pretrained over a general corpus of training data and finetuned on a more targeted corpus of training data. A model can be aligned using prompts that are designed to
elicit domain- specific outputs. Prompts can be designed to include learned prompt values (e.g., soft prompts). The trained model(s) may be validated prior to their use using input data other than the training data, and may be further updated or refined during their use based on additional feedback/inputs.
[0065] In some implementations, the computing system 102 may use one or more of the machine learning models or techniques noted above to perform any one or more of the operations discussed herein in connection with machine learning. For example, the computing system 102 may use one or more such machine learning techniques to pre-train and/or finetune the first generative Al model 150, the second generative Al model 152, and/or the third generative Al model 154, and possibly to pre-train and/or finetune a model that predicts performance of an image (e.g., to generate quality indicators as discussed above), etc.
[0066] Although the foregoing text sets forth a detailed description of numerous different aspects and implementations of the invention, it should be understood that the scope of the patent is defined by the words of the claims set forth at the end of this patent. The detailed description is to be construed as exemplary only and does not describe every possible implementation because describing every possible implementation would be impractical, if not impossible. Numerous alternative implementations could be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims. The disclosure herein contemplates at least the following examples:
[0067] Example 1. A method A method of semantic-based image copying, the method comprising: generating, by one or more processors, a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image; generating, by the one or more processors, a visual embedding based on the source image; and generating, by the one or more processors, a new image using a second generative Al model and based on the text prompt and the visual embedding.
[0068] Example 2. The method of Example 1, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
[0069] Example 3. The method of Example 2, comprising: generating, by the one or more processors, a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
[0070] Example 4. The method of Example 3, comprising: generating or obtaining, by the one or more processors, a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetuning, by the one or more processors, the third generative Al model based on the plurality of quality indicators.
[0071] Example 5. The method of Example 1, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
[0072] Example 6. The method of any one of Examples 1-5, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
[0073] Example 7. The method of any one of Examples 1-6, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
[0074] Example 8. The method of Example 7, wherein processing the source image includes: cropping out a portion of the source image.
[0075] Example 9. The method of any one of Examples 1-8, wherein the text prompt is the descriptive caption.
[0076] Example 10. The method of any one of Examples 1-9, wherein the second generative Al model is a pixel diffusion model.
[0077] Example 11. A system comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to (i) generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image, (ii) generate a visual embedding based on the source image, and (iii) generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
[0078] Example 12. The system of Example 11, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
[0079] Example 13. The system of Example 12, wherein the instructions cause the one or more processors to: generate a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
[0080] Example 14. The system of Example 13, wherein the instructions cause the one or more processors to: generate or obtain a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetune the third generative Al model based on the plurality of quality indicators.
[0081] Example 15. The system of Example 11, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
[0082] Example 17. The system of any one of Examples 11-16, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
[0083] Example 18. The system of any one of Examples 11-17, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
[0084] Example 19. The system of Example 18, wherein processing the source image includes: cropping out a portion of the source image.
[0085] Example 20. The system of any one of Examples 11-19, wherein the text prompt is the descriptive caption.
[0086] Example 21. The system of any one of Examples 11-20, wherein the second generative Al model is a pixel diffusion model.
[0087] Example 22. One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors
to: generate a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image; generate a visual embedding based on the source image; and generate a new image using a second generative Al model and based on the text prompt and the visual embedding.
[0088] Example 23. The one or more non-transitory, computer-readable media of Example 22, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
[0089] Example 24. The one or more non-transitory, computer-readable media of Example 23, wherein the instructions cause the one or more processors to: generate a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and (ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
[0090] Example 25. The one or more non-transitory, computer-readable media of Example 24, wherein the instructions cause the one or more processors to: generate or obtain a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetune the third generative Al model based on the plurality of quality indicators.
[0091] Example 26. The one or more non-transitory, computer-readable media of Example 22, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
[0092] Example 27. The one or more non-transitory, computer-readable media of any one of Examples 22-26, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
[0093] Example 28. The one or more non-transitory, computer-readable media of any one of Examples 22-27, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
[0094] Example 29. The one or more non-transitory, computer-readable media of Example 28, wherein processing the source image includes: cropping out a portion of the source image.
[0095] Example 30. The one or more non-transitory, computer-readable media of any one of Examples 22-29, wherein the second generative Al model is a pixel diffusion model.
[0096] Example 31. The one or more non-transitory, computer-readable media of any one of Examples 22-30, wherein the text prompt is the descriptive caption.
[0097] The following additional considerations apply to the foregoing discussion and the appended claims. Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter of the present disclosure.
[0098] Unless otherwise apparent from the context of use, reference in the present disclosure to a same set of “one or more processors” (or a same “plurality of processors,” etc.) performing multiple operations can encompass implementations in which performance of the operations is divided among the processor(s) in any suitable way. For example, “generating, by one or more processors, X; and generating, by the one or more processors, Y” can encompass: (1) implementations in which a first set of one or more processors (e.g., in a first computing device) generates X and a distinct, second set of one or more processors (e.g., in a different, second computing device) independently generates Y ; (2) implementations in which all processors in the set of one or more processors (e.g., all in the same device, or distributed among multiple devices) contribute to the generation of both X and Y ; and (3) other variations.
[0099] Unless specifically stated otherwise, discussions in the present disclosure using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,”
“displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that
manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[00100] As used in the present disclosure any reference to “one implementation” or “an implementation” means that a particular element, feature, structure, or characteristic described in connection with the implementation is included in at least one implementation or implementation. The appearances of the phrase “in one implementation” in various places in the specification are not necessarily all referring to the same implementation.
[00101] As used in the present disclosure, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[00102] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs through the principles described herein. Thus, while particular implementations and applications have been illustrated and described, it is to be understood that the disclosed implementations are not limited to the precise construction and components disclosed in the present disclosure. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of the method and apparatus disclosed in the present disclosure without departing from the spirit and scope defined in the appended claims.
Claims
1. A method of semantic-based image copying, the method comprising: generating, by one or more processors, a text prompt, wherein generating the text prompt includes applying a source image to a first generative artificial intelligence (Al) model to generate a descriptive caption for the source image; generating, by the one or more processors, a visual embedding based on the source image; and generating, by the one or more processors, a new image using a second generative Al model and based on the text prompt and the visual embedding.
2. The method of claim 1, wherein generating the new image includes: generating a mutated text prompt by applying the text prompt to a third generative Al model; and applying the mutated text prompt and the visual embedding to the second generative Al model.
3. The method of claim 2, comprising: generating, by the one or more processors, a plurality of new images using the second generative Al model and based on the text prompt and the visual embedding, at least in part by
(i) generating a plurality of mutated text prompts by applying the text prompt to the third generative Al model, and
(ii) generating each of the plurality of new images by applying a respective one of the plurality of mutated text prompts to the second generative Al model.
4. The method of claim 3, comprising: generating or obtaining, by the one or more processors, a plurality of quality indicators each corresponding to a respective one of the plurality of new images; and finetuning, by the one or more processors, the third generative Al model based on the plurality of quality indicators.
5. The method of claim 1, wherein generating the new image includes: applying the text prompt and the visual embedding to the second generative Al model.
6. The method of any one of claims 1-5, wherein generating the new image based on the text prompt and the visual embedding includes: generating the new image to have a different aspect ratio than the source image.
7. The method of any one of claims 1-6, wherein generating the visual embedding based on the source image includes: processing the source image; and generating the visual embedding based on the processed source image.
8. The method of claim 7, wherein processing the source image includes: cropping out a portion of the source image.
9. The method of any one of claims 1-8, wherein the text prompt is the descriptive caption.
10. The method of any one of claims 1-9, wherein the second generative Al model is a pixel diffusion model.
11. A system comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1- 10.
12. One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1-10.
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/US2024/030170 WO2025244625A1 (en) | 2024-05-20 | 2024-05-20 | Semantic-based image copying |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4681155A1 true EP4681155A1 (en) | 2026-01-21 |
Family
ID=91829171
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24739314.3A Pending EP4681155A1 (en) | 2024-05-20 | 2024-05-20 | Semantic-based image copying |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4681155A1 (en) |
| CN (1) | CN121399669A (en) |
| WO (1) | WO2025244625A1 (en) |
-
2024
- 2024-05-20 EP EP24739314.3A patent/EP4681155A1/en active Pending
- 2024-05-20 WO PCT/US2024/030170 patent/WO2025244625A1/en active Pending
- 2024-05-20 CN CN202480003718.8A patent/CN121399669A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN121399669A (en) | 2026-01-23 |
| WO2025244625A1 (en) | 2025-11-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20230013199A1 (en) | Digital Media Environment for Analysis of Audience Segments | |
| US10873782B2 (en) | Generating user embedding representations that capture a history of changes to user trait data | |
| US10789610B2 (en) | Utilizing a machine learning model to predict performance and generate improved digital design assets | |
| US10334328B1 (en) | Automatic video generation using auto-adaptive video story models | |
| WO2019079390A1 (en) | In-application advertisement placement | |
| US20110161479A1 (en) | Systems and Methods for Presenting Content | |
| CA2825814C (en) | Method and system for searching, and monitoring assessment of, original content | |
| CN111602152A (en) | A machine learning model for ranking different content | |
| US10290040B1 (en) | Discovering cross-category latent features | |
| JP2020024674A (en) | Method and apparatus for pushing information | |
| AU2017205232A1 (en) | Webinterface generation and testing using artificial neural networks | |
| JP2019125313A (en) | Learning device, learning method, and learning program | |
| CN102598039A (en) | Multimode online advertisements and online advertisement exchanges | |
| US20110153542A1 (en) | Opinion aggregation system | |
| Oosterhuis et al. | Ranking for relevance and display preferences in complex presentation layouts | |
| JP2023162154A (en) | Method, computer device, and computer program for providing recommendation information based on a local knowledge graph | |
| US20170052926A1 (en) | System, method, and computer program product for recommending content to users | |
| AlDahoul et al. | Towards a World Wide Web powered by generative AI | |
| US20250133273A1 (en) | Machine learning assisted and template guided video synthesis | |
| CN110162714A (en) | Content delivery method, calculates equipment and computer readable storage medium at device | |
| Ziegler et al. | Interactive recommendation systems | |
| US20250004797A1 (en) | Training Pipeline for Training Machine-Learned User Interface Customization Models | |
| EP4681155A1 (en) | Semantic-based image copying | |
| CN120525586A (en) | Information recommendation method, device, electronic device, computer-readable storage medium, and computer program product | |
| Feng et al. | Impact of textual and visual complexity and consistency on review usefulness: moderated by commodity price |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250731 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |