EP4699014A1 - Resource efficient image content verification - Google Patents

Resource efficient image content verification

Info

Publication number
EP4699014A1
EP4699014A1 EP23848678.1A EP23848678A EP4699014A1 EP 4699014 A1 EP4699014 A1 EP 4699014A1 EP 23848678 A EP23848678 A EP 23848678A EP 4699014 A1 EP4699014 A1 EP 4699014A1
Authority
EP
European Patent Office
Prior art keywords
image
new
entity
new image
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23848678.1A
Other languages
German (de)
French (fr)
Inventor
Sikun LIN
Xiaohang Li
Nathan P. Lucash
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Google LLC
Original Assignee
Google LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Google LLC filed Critical Google LLC
Publication of EP4699014A1 publication Critical patent/EP4699014A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/50Information retrieval; Database structures therefor; File system structures therefor of still image data
    • G06F16/58Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/583Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • G06F16/5854Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using shape and object relationship
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/50Information retrieval; Database structures therefor; File system structures therefor of still image data
    • G06F16/55Clustering; Classification

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Library & Information Science (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for determining whether a new image satisfies one or more image criteria. In one aspect, a method includes maintaining reference images that each satisfy one or more criteria and, for each of the reference images, corresponding reference entity data that specifies reference entities depicted in the reference image and reference relationships among the reference entities. New entity data is generated for a new image. The new image data specifies new entities depicted in the new image and new relationships among the new entities. A reference image is identified for the new image reference image. A difference between the new entity data for the new image and the corresponding reference entity data for the identified reference image is determined. A portion of the new image that represents a modification relative to the identified reference image is identified.

Description

RESOURCE EFFICIENT IMAGE CONTENT VERIFICATION
BACKGROUND
[0001] In a computer networked environment such as the Internet, third-party content providers provide third-party' content items for display on end-user computing devices. These third-party content items, for example, digital images and video, can be display ed on client devices in the environment. Digital images and video can be used, for example, on the Internet, for remote meetings via video conferencing, high-definition video entertainment, and/or sharing of user-generated content.
[0002] Recent developments in artificial intelligence and in particular, generative artificial intelligence have caused user-produced visual content such as digital images to become ubiquitous. For example, various types of images can be generated by using a text-to- image models based on text prompts.
[0003] In order to verify whether the content of these digital images is appropriate for sharing on digital platforms or some other purposes after having been generated, a verification process needs to be repeatedly performed. For example, this verification process is required each time a new image is obtained using such a text-to-image model. Repeatedly performing this verification process is computationally intensive and consumes a significant amount of computational resources, especially when the image is a high resolution image that has a large data size, or when there exists a large number of images that need to be verified, or both.
SUMMARY
[0004] This specification describes an image verification system implemented as computer programs on one or more computers in one or more locations that obtains new images, processes the new images to generate verification results, and takes any of a number of different actions on the new images based on the verification results. The verification result indicates whether the new image satisfies image criteria.
[0005] In general, one innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions of maintaining multiple reference images that each satisfy one or more criteria and, for each of the multiple reference images, corresponding reference entity data that specifies (i) multiple reference entities depicted in the reference image and (ii) reference relationships among the multiple reference entities; obtaining a new image; generating new entity’ data for the new image that specifies (i) multiple new entities depicted in the new image and (ii) new relationships among the multiple new entities depicted in the new image; identifying, for the new image, a reference image from the multiple reference images; determining a difference between (i) the new entity data for the new image and (ii) the corresponding reference entity data for the identified reference image; identifying, based on the difference, a portion of the new image that represents a modification relative to the identified reference image; generating, based on the portion of the new image, a verification result that indicates whether the new image satisfies the one or more criteria; and performing an action with respect to the new image based on the verification result. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.
[0006] In some aspects, determining the difference includes determining that the new entity7 data represents a new entity that is not represented by the reference entity data and identifying the portion of the new image includes identifying a portion of the new image that depicts at least the new entity. Identifying the portion of the new image can include identify ing the portion of the new image that depicts the new entity and also one or more entities that are related to the new entity and that are represented by the reference entity data. [0007] In some aspects, determining the difference includes determining that the new entity data represents a new relationship between two existing entities, where the two existing entities are represented by the reference entity data but the new relationship is not represented by the reference entity data. Identifying the portion of the new image can include identifying a portion of the new image that depicts at least the two existing entities.
[0008] In some aspects, for each of the multiple reference images, the corresponding reference entity data includes a graph that includes nodes representing the multiple reference entities and edges representing the reference relationship among the plurality of reference entities.
[0009] In some aspects, the reference relationship includes one or more of a spatial relationship, a temporal relationship, a semantic relationship. The spatial relationship can represent an object in physical contact with or in proximity to another object. The semantic relationship can represent a human pointing toward or leading an eye to another human or another object.
[0010] In some aspects, obtaining the new image includes receiving an existing image and a text prompt and generating the new image from the input image and the text prompt by using a generative machine learning model. Identifying, for the new image, the reference image from the multiple reference images can include identifying, as the reference image, the existing image from the multiple reference images. Generating the new entity data for the new image can include generating the new entity data based on the text prompt.
[0011] In some aspects, generating the new entity data for the new image includes performing image segmentation on the new image to partition the new7 image into multiple segments each depicting a new entity and generating the new entity data based on the multiple segments.
[0012] In some aspects, determining whether the new7 image satisfies the one or more criteria includes determining, based on the portion of the new7 image and the new7 entity data, whether the new image satisfies the one or more criteria.
[0013] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages. The specification describes an image verification system that is configured to verify content of new images in a comprehensive and yet, resource efficient manner. The verification can be used to ensure that an image, e.g.. an image generated using artificial intelligence, satisfies a set of criteria before allowing the image to be distributed or otherwise published. To improve the efficiency and speed at which the verification is performed, the verification system can compare entity data for a new image to entity7 data for a reference image that has been verified or for which it is otherwise known that the image satisfies the image criteria to identify a smaller portion of the image that has changed from the reference image to the new image. By using entity7 data that represents the entities depicted in an image and relationships between the entities depicted in the image to identify a smaller portion of the image to be further processed for content verification, the described image verification system reduces the amount of computational resources consumed by the verification process and the time involved in verifying the images because processing the image in its entirety is no longer required. Instead, only a portion of the image that includes a relatively' smaller number of pixels compared to the entire image is processed to determine the verification result for the image. The identified portion of the image corresponds to a difference between the entity data generated for the new image and the entity data generated for the reference image. Meanwhile, the image verification system can identify the portion of the new7 image in a manner that is consistent with relationships between the entities as represented by the entity7 data to ensure the comprehensiveness of the content verification process. Thus, the image verification system can perform image content verification with reduced latency and reduced consumption of computational resources relative to conventional techniques that require full image processing, while still maintaining verification accuracy.
[0014] In addition, using entity data to identify the smaller portion of the image for evaluation reduces the computational resources and latency involved in identifying the smaller portion of the image relative to techniques that perform comparisons of the images themselves, e.g., by comparing the properties of the pixels between the images or other types of vision analysis. Using entity data and reducing the evaluation of the new image to those portions that have changed provides a synergistic effect of substantially reducing the consumption of computational resources and latency involved in evaluating images. Aggregated over millions of images, e.g., per day. as the use of generative artificial intelligence increases, this results in enormous computing resource consumption savings and time savings, which can reduce the amount of hardware required for a system that performs the evaluations.
[0015] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0016] FIG. 1 is a block diagram illustrating an example environment in which an image verification system verifies images.
[0017] FIG. 2A is an illustration of an example process for identifying a portion of a new image that represents a modification relative to a reference image.
[0018] FIG. 2B is an illustration of an example process for identifying a portion of a new image that represents a modification relative to a reference image.
[0019] FIG. 3 illustrates an example process for verifying an image by determining whether a new image satisfies one or more image criteria.
[0020] FIG. 4 shows a block diagram of an example computer.
[0021] Like reference numbers and designations in the various drawings indicate like elements.
DETAILED DESCRIPTION
[0022] FIG. 1 is a block diagram illustrating an example environment 101 in which an image verification system 100 for verifies images. The image verification system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations that obtains new images 102, processes the new images 102 to generate verification results 1 13, and takes any of a number of different actions, e.g., on the new images 102, based on the verification results 113.
[0023] The example environment 101 includes a network 107, e.g., a local area network (LAN), a w ide area network (WAN), a mobile network, the Internet, or a combination thereof. The network 107 connects the image verification system 100, content provider computing device(s) 103, and client device(s) 105. The example environment 101 may include many different content provider computing devices 103 and client devices 105.
[0024] A client device 105 is an electronic device capable of requesting and receiving online resources over the network 107. Example client devices 105 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over the network 107. A client device 105 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 107, but native applications (other than browsers) executed by the client device 105 can also facilitate the sending and receiving of data over the network 107.
[0025] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally , or execute a gaming application that is at least partly stored and/or served by a cloud server (e.g., online gaming applications). Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.
[0026] Digital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e.g., by way of a wireless or wired connection).
Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.
[0027] The content provider computing device 103 can include servers or other computing devices operated by a content provider entity to generate and provide content, including new images 102, for display at client devices 105. In some cases, the content provided by the content provider computing device 103 can include third-party content items, e.g., th i r -party images, for display on electronic documents that include primary content, e.g., content provided by electronic document servers operated by a content publisher.
[0028] An electronic document is data that presents a set of content at the client device 105. Examples of electronic documents include webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps” and/or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents.
[0029] As used herein, the term "image’' can mean a digital image, such as a two- dimensional image or a three-dimensional image, or even consecutive frames of video. An image can include multiple pixels.
[0030] In some cases, the content provider computing device 103 can generate new images 102 in response to received requests, e.g., from the content provider entity or the client device 105. In some cases, the content provider computing device 103 can generate digital components, which in turn include one or more new images 102.
[0031] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g.. a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a phy sical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.
[0032] To generate new images 102 that are creative and factual, in a large volume, or both, the content provider computing device 103 can use artificial intelligence, e.g., generative artificial intelligence. Artificial intelligence (Al) is a segment of computer science that focuses on the creation of intelligent agents that can leam and act autonomously (e.g., without human intervention). Any of a variety' of machine learning models/algorithms can be used by the content provider computing device 103 to generate new images 102.
[0033] In one example, the content provider computing device 103 can use an Imagen model to generate new images conditioned on the provided text prompts. The Imagen model (arXiv:2205. 11487) is a text-to-image model with a high degree of photorealism and deep language representations. The Imagen model includes multiple diffusion models, which start by generating a small image and progressively increase its resolution.
[0034] In another example, the content provider computing device 103 can use a Pathways Autoregressive Text-to-Image (Parti) model. The Parti model (arXiv: 2206. 10789) is an autoregressive text-to-image generation model that achieves high-fidelity photorealistic image generation and supports content-rich synthesis involving complex compositions and world knowledge.
[0035] In yet another example, the content provider computing device 103 can use a Muse model. The Muse model (arXiv:2301.00704) is a text-to-image Transformer model that achieves strong image generation performance while being significantly more efficient than diffusion or autoregressive models. It’s trained on a masked modeling task in a discrete visual token space, conditioned on a text embedding extracted from a pre-trained large language model (LLM). Muse also supports text-guided and mask-based image editing.
[0036] It is to be appreciated that these are merely examples for illustrative purposes, and the content provider computing device 103 can use other suitable machine learning models/algorithms to generate images either unconditioned, or conditioned on a context input which includes input text, or input image, or both input text and input image and possibly other data.
[0037] As a particular example for generating images conditioned on both input text and an input image, the content provider computing device 103 can modify an existing image to generate a new image 102. That is, the new image 102 is a modified version of the existing image. To do this, the content provider computing device 103 can use one of the text-to- image models mentioned above or other known text-to-image models.
[0038] For example, consider an existing image of sunflowers in a vase. In this example, the content provider computing device 103 receives a request to modify the existing image to remove the sunflowers, e.g., receives the following text prompt from the content provider entity or another user of the computing device: “make the sunflowers disappear.’’ Correspondingly, the content provider computing device 103 processes the existing image conditioned on the text prompt to generate a new image 102 that is an image of an empty- vase that is the same as or similar to the vase that holds the sunflowers.
[0039] As another example, consider existing image of an urban skyline with buildings. In this example, the content provider computing device 103 receives a request to modify the existing image to additionally depict fireworks in the sky-, e.g., receives the following text prompt from the content provider entity or another user of the computing device: “add fireworks to the sky/’ Correspondingly, the content provider computing device 103 processes the existing image conditioned on the text prompt to generate a new image 102 that is an image of an urban skyline with the same or similar buildings as the existing image, and additionally depicts fireworks in the sky above these buildings.
[0040] As yet another example, consider an existing image of a man wearing a down jacket. In this example, the content provider computing device 103 receives a request to modify- the existing image to depict the man instead wearing a leather jacket, e.g., receives the following text prompt from the content provider entity' or another user of the computing device: “swap his down jacket with a leather one.-’ Correspondingly, the content provider computing device 103 processes the existing image conditioned on the text prompt to generate a new image 102 that is an image of the man wearing a leather jacket.
[0041] After having generated the new image 102, the content provider computing device 103 provides the new image 102 to the image verification system 100 and the image verification system 100 processes the new image 102 using the image content verification techniques, as will be described below-, to generate a verification result 1 13 for the new image 102. The verification result indicates whether the new image 102 satisfies image criteria. In some cases, providing the new image includes the content provider computing device 103 uploading the new image 102 over the network 107 to the image verification system 100. In some cases, providing the new image includes the content provider computing device 103 providing to the image verification system 100 data that specifies a name or a netw ork location, e.g., a Uniform Resource Identifier (URI) or a Uniform Resource Locator (URL) of a sen' er from which the image verification system 100 can obtain the new image 102.
[0042] The image verification system 100 includes or has access to image criteria 115. which is referred to as image criteria 115 for ease of subsequent description. Data defining the image criteria 115 can either be uploaded by a user of the image verification system 100, or may alternatively be determined by the image verification system 100 itself based on the known requirements of a digital platform on which images may be shared. The image criteria 115 may be stored in a datastore in associated with different digital platforms. The image criteria 115 may be periodically corrected and updated as needed.
[0043] In various cases, the image criteria 1 15 may include criteria specifying particular image characteristics, e.g., visual quality characteristics related to the visual appearance of the images, e.g., images should have at least a threshold resolution in order to be presented on the client devices 105, particular image content criteria, e.g., if some types of image content features are designated to not be presented on the client devices 105 or if certain types of content are designated to not be shown with certain other types of content, and/or other characteristics.
[0044] As one example, the image criteria 115 can specify that an image should not include certain graphic content. For example, the criteria can specify that an image should not include any graphic content that is obscene, pornographic, illegal, defamatory, negative, offensive, etc. As another example, the image criteria 115 can specify that an image should not include any confidential or private information. For example, the criteria can specify that an image should not include any watermark. A watermark can be a digital signature, a user identifier, a logo, or similar value that can sufficiently identify the image as to its origin or ownership. As another example, the criteria can specify that an image should not include content related to topic A and content related to topic B. The image criteria 115 can vary for different types of images or images that depict different types of content.
[0045] The image verification system 100 includes or has access to an image database 1 16. The image database 1 16 stores reference images 120 that have each satisfied the image criteria 115. For example, the images 120 can be the images 120 that the image verification system 100 has verified successfully. The image database 116 can store any type of image, such as a static image (e.g., a photograph, a drawing, etc.), or an image that includes motion (e.g., a cinemagraph, a live photo, a video segment, clip, or snippet, etc.) that satisfies the image criteria 115.
[0046] In some cases, the image verification system 100 receives images captured by a scanner, a camera, a specially-adapted sensor array (e g., CCD array), a microscope, a smartphone camera, or a video camera, and evaluates the received images against the image criteria 115. When the received images satisfy' the image criteria 115, the image verification system 100 can store the images in the image database 116 as reference images 120 in a digital data format, such a viewable file format, a compressed image file format, a raw image file format, or the like. [0047] In some cases, the image verification system 100 receives images generated using Al. e.g., by the content provider computing device 103 or some other computing devices that utilized Al to generate images, and evaluates the received images using the image criteria 115. Likewise, when the received images meet the image criteria 115, the image verification system 100 stores the images in the image database 116 as reference images 120.
[0048] In particular, if an image received by the image verification system 100 does not meet the image criteria 115, it will not be stored in the image database 116. Instead, only images that are “verified,” namely, images that satisfy the image criteria 115 will be stored in the image database 116 as reference images 120. Thus, the image database 116 may not include any images that have not been successfully verified.
[0049] The image database 116 also stores corresponding reference entity data 122 in association with each of the reference images 120. For each reference image 120, the corresponding reference entity data 122 is a collection of data, e.g., nodes, representing entities depicted in the reference image 120 and relationships, e.g., edges, between entities represented by the nodes. An image can include multiple pixels, and the entities can correspond to different respective subsets of pixels included in the image (possibly overlapping subsets) that depict the entities, which may be spatially displaced relative to each other. The entities can include, for example, humans or parts of a human, e.g., a hand or a head of a human, objects or parts of an object, animals or parts of an animal, and so on. [0050] In some cases, for a given image, the corresponding reference entity’ data that represents entities depicted in the given image and relationships between the entities can be represented by a graph. The graph can include a node for each entity. For each pair of related entities, the graph can include an edge that connects the entities and that specifies one or more relationships between the entities.
[0051] The image verification system 100 can generate the data that represents the entities in a given reference image in any appropriate way. For example, the image verification system 100 can use an object detection machine learning model, e.g., a convolutional neural network or an attention neural network, to identity’ and/or recognize the entities in the given reference image. In this example, the machine learning model can be configured, e.g., trained, to process the reference image to generate, as output, data that defines multiple bounding boxes with reference to the given reference image and, for each of the multiple bounding boxes, a respective likelihood that an entity' belonging to an entity category from a set of possible entity categories is present in the region of the given reference image circumscribed by the bounding box. [0052] As another example, the image verification system 100 can use an image segmentation machine learning model, e.g., a convolutional neural network or an attention neural network, to partition the given reference image into multiple segments that each likely depict an entity. In this example, the image segmentation machine learning model can be configured, e.g., trained, to process the given reference image to generate, as output, data that defines, for each pixel included in the given reference image, to which of multiple entity categories mentioned above that the pixel belongs.
[0053] As another example, the image verification system 100 can use instance segmentation techniques or other segmentation techniques such as semantic segmentation techniques to identify the pixel boundaries of each of the entities in the given reference image.
[0054] Likewise, the image verification system 100 can generate the data that represents the relationships between the entities a given reference image in any appropriate way. For example, the relationships between the entities can include one or more of: a spatial relationship, a temporal relationship, or a semantic relationship between two or more entities in the given reference image. The image verification system 100 can use any known computer vision/machine learning techniques to identify and/or recognize such relationships between the entities in the given reference image.
[0055] For example, the image verification system 100 can use one of the entity relationship recognition techniques described in Yang Wang, Huilin Peng, Yiwei Xiong. Haitao Song, “Spatial relationship recognition via heterogeneous representation: A review,” Neurocomputing, Volume 533, 2023, Pages 116-140, and in Cui, Wei, Fei Wang, Xin He, Dongyou Zhang, Xuxiang Xu, Meng Yao. Ziwei Wang, and Jiejun Huang. 2019. “MultiScale Semantic Segmentation and Spatial Relationship Recognition of Remote Sensing Images Based on an Attention Model,” Remote Sensing 11, no. 9: 1044.
[0056] Other techniques may also be used to determine the entities and relationships between entities when generating the reference entity data 122 in association with each of the reference images 120.
[0057] To generate the verification result 113 for the new image 102. the image verification system 100 generates new entity data 112 for the new image 102. Like the reference entity data 122, the new entity data 112 is a collection of data, e.g., nodes, representing entities depicted in the new image 102 and relationships, e.g., labeled edges, between entities, e.g., represented by the nodes. In other words, the same type of graph representation can be used to express the entities identified in reference and new images and to express the relationships between the entities.
[0058] In some cases, the new entity data 112 can be generated using similar techniques to those mentioned above. In some other cases, when the new image 102 is a modified version of an existing image, the new entity data 112 that corresponds to the new image 102 can be generated based on the text prompt that was used to generate the new image 102. Continuing with example of the urban skyline image described above, the image verification system 100 can generate the new entity data 112 that includes data representing a new entity (“fireworks”) based on parsing the received text prompt “add fireworks to the sky” to identity “fireworks” as the new entity.
[0059] The image verification system 100 uses the new entity data 112 and the reference entity’ data 122 to identify a smaller portion of the new image 102 that is less than all pixels of the new image 102. Rather than evaluate the entire new image, the image verification system 100 can evaluate the identified portion of the new image 102 against the image criteria 115 to generate the verification result 113 for the new image 102. The identified portion of the new image 102 corresponds to a portion of the new image that may represent a modification relative to the identified reference image, and may include a proper subset of the set of pixels included the new image 102. A proper subset of a set is one that includes at least one member of the set but does not include all members of the set. Example techniques for identifying the smaller portion of the new image 102 is described further below with reference to FIGS. 2A-B.
[0060] A number of actions can then be taken by the image verification system 100, e.g., on the new image 102, after the verification result 113 for the new image 102 has been generated. In some cases, the image verification system 100 can output the verification result 1 13 as a response to the content provider computing device 103 that generated the new image 102. For example, the image verification system 100 can output a response to the content provider computing device 103 which indicates that the new image 102 has been verified. Additionally or alternatively, the image verification system 100 can store the new image 102 in the image database 116 as a reference image 120. or in another local storage device for some other purposes.
[0061] In some cases, when the verification result 113 indicates that the new image 102 satisfies the image criteria 115, the image verification system 100 can provide the new image 102 to one or more client devices 105 for display thereon. That is, the image verification system 100 can provide the new image 102 over the network 107 to the client device 105 and the client device 105 in turn presents the new image 102.
[0062] In some of these cases, the image verification system 100 can provide the new image 102 as a standalone image while in others of these cases, the image verification system 100 can provide the new image 102 as a, or as part of, a digital component. In the latter cases, the image verification system 100 can also provide event data specifying other event features, such as an electronic document and characteristics of locations of the electronic document at which the digital component that includes the new image 102 can be presented. For example, event data specifying a reference, e.g., URL, to an electronic document, e.g., webpage, in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and/or media types that are eligible for presentation in the locations can be provided to the client device 105.
[0063] In some cases, when the verification result 113 indicates that the new image 102 fails to satisfy the image criteria 115, the image verification system 100 can provide the new image 102 to other components of the environment 101 for further processing. For example, the image verification system 100 can use one of the text-to-image models mentioned above, or another image edit tool, to modify, e.g., edit, or otherwise manipulate the content of the new image 102, such that the modified image would satisfy the image criteria 115. The mage verification system 100 can then provide the modified image to one or more client devices 105, e g., if the modified image satisfies the image criteria 1 15.
[0064] FIG. 2A is an illustration of an example process for identifying a portion of a new image that represents a modification relative to a reference image. On the left side, FIG. 2A illustrates an example of a reference image 210 and reference entity data 220 that corresponds to the reference image 210. The reference image 210 is a verified image that has passed the image criteria. The reference image 210 includes multiple pixels and depicts a total of four entities (entity A, entity B, entity C, and entity D) at four different respective subsets of the pixels included the reference image 210.
[0065] The reference entity data 220 is a collection of data representing the four entities (entity A, entity B, entity C, and entity D) and relationships between the four entities. The data is logically described, illustrated, and/or stored as a graph, in which each distinct entity is represented by a respective node and each relationship between a pair of entities is represented by an edge or link between the nodes. Each edge may be either directed, or undirected. Each edge specifies a relationship, and the existence of the edge represents that the specified relationship exists between the nodes connected by the edge. [0066] In the reference image 210, entity A is related to entity B, entity B is related to entity C, and entity' C is related to entity D, thus FIG. 2A illustrates that the reference entity data 220 includes a node A representing entity A connected by an edge 221 to node B representing entity B, which is connected by another edge 222 to node C representing entity C, which is connected by another edge 223 to node D representing entity D.
[0067] For example, entity’ A may be spatially related to entity B, and the edge 221 connecting node A and node B represents a spatial relationship between node A and node B. For example, assuming entity' A and entity B each correspond to a respective object, they may be spatially related, e.g., when entity A is in physical contact with or in proximity to entity’ B. In this example, the special relationship can specify that entity’ A is in physical contact with entity B or that entity A is in proximity’ to, e.g., within a defined distance of, entity B.
[0068] As another example, entity A may be temporally related to entity B, and the edge 221 connecting node A and node B represents a temporal relationship between node A and node B. For example, assuming entity A and entity B each correspond to a respective object, they may be temporally related, e.g., when entity A and entity' B are present in a same scene at the same time.
[0069] As another example, entity A may be semantically related to entity B. and the edge 221 connecting node A and node B represents a semantic relationship between node A and node B. For example, assuming entity’ A and entity B each correspond to a respective human, or assuming entity A and entity B correspond to a human and an object, respectively, they may be semantically related, e.g.. when entity A is pointing toward or leading the eye to entity B.
[0070] The reference entity data 220 can alternatively be stored in any of a variety of data structures other than a graph. For example, the entity data can be stored as triples that each represent two entities in order and a relationship from the first to the second entity; for example, [entity A, entity B. is related to], or [entity A. is related to, entity B], are alternative ways of representing the same fact. Other example data structures include linked lists, adjacency lists (in which the adjacency information includes relationship information), and so on.
[0071] In some cases, an edge can have a corresponding label which identifies the type of relationship represented by the edge, e.g., can have one of: a spatial relationship, a temporal relationship, or a semantic relationship label. In some cases, the reference entity data 220 can include different edges, or different labels for the same edge, to represent different relationship between a pair of nodes, e.g., can include multiple edges for different relationships between the same pair of nodes. In other cases, the reference entity data 220 can include a single edge to represent different relationships between the pair of nodes. The label can describe the relationship, e.g., the objects are in contact, one object is above another object, one object is in the hand of a person, etc. for spatial relationships.
[0072] On the right side, FIG. 2A illustrates an example of a new image 230 and new entity7 data 240 that corresponds to the new image 230. The new image 230 is a modified version of the reference image 210 that includes an additional entity (entity E). The new image 230 thus includes multiple pixels and depicts a total of five entities (entity A, entity B, entity’ C, entity D, and entity E) at five different respective subsets of the pixels included the new image 230.
[0073] The new entity data 240 is a collection of data representing the five entities (entity A, entity B, entity C. entity D, and entity E) and relationships between the five entities. In the new image 230, entity A is related (e.g.. spatially, or temporally, semantically, or otherwise related) to entity B, entity B is related to entity C, and entity C is related to both entity D and entity E, thus FIG. 2A illustrates that the new entity data 240 includes a node A (representing entity A) connected by an edge to node B (representing entity B), which is connected by another edge to node C (representing entity C), which is connected by respective edges 223 and 224 to both node D (representing entity D) and node E (representing entity E).
[0074] The image verification system compares the reference entity' data 220 and the new entity data 240 to determine differences between the reference entity data 220 and the new entity data 240. In the example of FIG. 2A, because an additional node (node E, representing entity E) relative to the reference entity data 220 has been added in the new entity data 240, the image verification system can determine that the difference resides in the part of the new entity data 240 that includes node E. Such a difference is illustrated in FIG. 2A by the dashed ellipse 225.
[0075] Alternatively, because there exists a relationship between node C and node E (as represented by the edge that connects nodes C and E), the image verification system may determine that the difference resides in the part of the new entity data 240 that includes both node C and node E. Such a difference is illustrated in FIG. 2A by the dashed ellipse 226. In either case, the part of the new entity data 240 in which the difference resides is smaller than the entirety of the new entity data 240. [0076] The image verification system identifies a portion of the new image 230 based on the determined difference. An example of such a portion of the new image 230 is illustrated as the shaded region within the new image 230, which depicts both entity C and entity E. Such a portion of the new image 230, which can be identified by the image verification system based on the difference illustrated in FIG. 2A by the larger dashed ellipse 226, represents a modification in the new image 230 relative to the reference image 210.
[0077] Correspondingly, the image verification system evaluates the identified portion of the new image 230 against the image criteria to generate a verification result for the new image 230. The verification result indicates whether the new image 230 satisfies the image criteria.
[0078] Thus, rather than evaluating the new image 230 in its entirety against the image criteria, only a portion of the new image 230 (the portion corresponding to the shaded region — or an even smaller region — within the new image 230) is used, e.g., processed, by the image verification system to make this determination, while the remaining portion of the new image 230 that remains unmodified (the portion corresponding to the transparent region — or an even larger region — within the new image 230) is unused, e.g., not processed, by the image verification system.
[0079] Moreover, the image verification system can identify the portion of the new image 230 for evaluation against the image criteria in a manner that is consistent with relationships between the entities as represented by the entity data, thereby improving the comprehensiveness of the identified portion in terms of encapsulating the modification relative to the reference image which, in turn, improves the accuracy of the verification result.
[0080] FIG. 2B is an illustration of an example process for identifying a portion of a new image that represents a modification relative to a reference image. On the left side, FIG. 2B illustrates the example of the reference image 210 and the reference entity data 220 that corresponds to the reference image 210.
[0081] On the right side. FIG. 2B illustrates an example of a new image 250 and new entity data 260 that corresponds to the new image 250. The new image 250 is a modified version of the reference image 210 in which an existing entity (entity D) in the reference image 210 has been relocated closer to another existing entity (entity B). In other words, the new image 250 includes multiple pixels and depicts the same entities (entity A, entity B, entity C, and entity D) as the reference image 210; however, in the new image 250, entity D is depicted in a different subset of the pixels than in the reference image 210. [0082] The new entity data 260 thus includes an additional relationship between the two existing entities in the reference image 210 (entity B and entity D), as represented by the additional edge 232 that connects node B and node D (representing entity B and entity D, respectively).
[0083] The image verification system compares the reference entity data 220 and the new entity data 260 to determine differences between the reference entity data 220 and the new entity data 260. In the example of FIG. 2B, because an additional edge (connecting node B and node D) relative to the reference entity data 220 has been added in the new entity data 260, the image verification system may determine that the difference resides in the part of the new entity data 240 that includes nodes B and D that are connected by the additional edge. Such a difference is illustrated in FIG. 2B by the dashed concave contour 235.
[0084] Alternatively, because there exists a relationship between node B and node C, and also a relationship between node C and node D (as represented by the edges that connect nodes B and C. and nodes C and D, respectively), the image verification system may determine that the difference resides in the part of the new entity data 260 that includes node B, node C. and node D. Such a difference is illustrated in FIG. 2B by the dashed ellipse 236. Like the example of FIG. 2A, in either case, the part of the new entity data 260 in which the difference resides is smaller than the entirety of the new entity data 240.
[0085] Correspondingly, the image verification system evaluates the identified portion of the new image 250 against the image criteria to generate a verification result for the new image 250. The verification result indicates whether the new image 250 satisfies the image criteria. Like the example of FIG. 2A, in FIG. 2B, rather evaluating the new image 250 in its entirety against the image criteria, only a portion of the new image 250 (the portion corresponding to the shaded region — or an even smaller region — within the new image 250) is used, e.g., processed, by the image verification system to make this determination, while the remaining portion of the new image 250 that remains unmodified (the portion corresponding to the transparent region — or an even larger region — within the new image 250) is unused, e.g., not processed, by the image verification system.
[0086] Although entity B is described above as being additionally related to entity D because of its relocation, as shown in the new image 250, it will be appreciated that entity B may be additionally related to entity7 D because of other reasons. For example, the relationship between node B and node D represented by the new entity data 260 may be a semantic relationship, and assuming entity B corresponds to a human, then entity B may be
Y1 semantically related to entity D because entity B is pointing toward or leading the eye to entity D.
[0087] FIG. 3 is a flow diagram of an example process 300 for verifying an image by determining whether a new image satisfies one or more image criteria. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, an image verification system, e.g., the image verification system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300. Operations of the process 300 can also be implemented as instructions stored on one or more computer-readable media, which may be non-transitory, and execution of the instructions by one or more data processing apparatus can cause the one or more data processing apparatus to perform the operations of the process 300.
[0088] The system maintains multiple reference images and corresponding reference entity data (step 302). Each reference image satisfies image criteria. In various cases, the image criteria may include criteria specifying particular image characteristics, e.g.. visual quality characteristics related to the visual appearance of the images, particular image content criteria, or other characteristics.
[0089] For each of the multiple reference images, the corresponding reference entity7 data specifies multiple reference entities depicted in the reference image and reference relationships among the multiple reference entities. For example, the reference entity data for a given reference image can specify one or more of: spatial relationships, temporal relationships, semantic relationships between different reference entities depicted in the given reference image, or other ty pes of relationships between entities.
[0090] The system obtains a new image (step 304). In some cases, the new image is an image captured by a sensor, e.g., a still or a video camera. In some other cases, the new image is generated by utilizing (generative) artificial intelligence. For example, the new image is an image generated by a generative machine learning model, e.g., a text-to-image model, conditioned on a text prompt. As another example, the new image is a modified version of an existing image that is generated by a generative machine learning model, e.g., a text-to-image model, based on the existing image and a text prompt that describes how the existing image should be modified. The existing image can for example be one of the multiple reference images that are already maintained by the system. In this example, to generate the new image, a content provider computing device can provide, as input, an existing image and a text prompt to the generative machine learning model, and process the input using the generative machine learning model to generate, as output, the modified image.
[0091] The system generates new entity data for the new image (step 306). The new entity data specifies multiple new entities depicted in the new image and new relationships among the multiple new entities depicted in the new image. The system can do this by using similar techniques to those used to generate the reference entity data. Additionally or alternatively, when the new image is a modified version of an existing image, the new entity data that corresponds to the new image can be generated based on the text prompt that was used to generate the new image.
[0092] The system identifies, for the new image, a reference image from the multiple reference images (step 308). In some cases, the identification is based on similarity scores between the new image and the multiple reference images. In some other cases, when the new image is a modified version of an existing image, the system can identify, as the reference image, the existing image based on which the new image is generated.
[0093] In the prior cases, a similarity score between the new image and a given reference image may for example depend on a vector distance in n-dimensional embedding space between a first embedding generated by a neural network from the new image and a second embedding generated by the neural network from the given reference image. In these cases, after calculating the similarity scores, the reference image that has the highest similarity scores from among the multiple reference images, or the reference image that has a similarity score that is greater than a threshold similarity' score can be selected. The similarity scores may also be computed in other ways.
[0094] The system compares the new entity data and the corresponding reference entity data for the identified reference image to determine one or more differences between the new entity data for the new image and the corresponding reference entity data for the identified reference image (step 310). The differences may vary depending on what is actually depicted in the new image and in the identified reference image.
[0095] For example, as a result of the comparison, the system can determine that the new entity data and the corresponding reference entity data represent different entities. For example, assuming the reference entity data represents a reference entity, and the new entity data represents a substitute entity in place of the reference entity, the system can determine that the differences include the substitute entity. As another example, assuming the new entity data represents a new entity that is not represented by the reference entity data, the system can determine that the differences include the new entity. [0096] As another example, as a result of the comparison, the system can determine that the new entity data and the corresponding reference entity data represent different relationships among the same entities. For example, assuming the new entity data represents anew relationship, e.g., that is not represented by the reference entity data, between two existing entities, the system can determine that the differences include the two existing entities that are connected by the new relationship.
[0097] The system identifies, based on the one or more differences, a portion of the new image that represents a modification relative to the identified reference image (step 312). In particular, the new image includes multiple pixels, and the identified portion of the new image includes a proper subset of the multiple pixels. The proper subset includes at least one pixel in the new image, but less than all of the pixels in the new image.
[0098] In the example above where the system has determined that the differences include the substitute entity, the system can identify a portion of the new image that depicts at least the substitute entity, or may alternatively identify' a portion of the new image that depicts at least the new entity and any entities that are related to the substitute entity.
[0099] In the example above where the system has determined that the differences include the new entity, the system can identify' a portion of the new image that depicts at least the new entity, or can alternatively identify' a portion of the new image that depicts at least the new entity and any entities that are related to the new entity.
[00100] In the example above where the system has determined that the differences include the two existing entities has the new relationship, the system can identify a portion of the new image that depicts at least the two existing entities.
[00101] The system generates, based on the identified portion of the new image, a verification result that specifies whether the new image satisfies the image criteria (step 314). In some implementations, this involves the system processing the identified portion of the new image and not the remaining portion of the new image to determine whether the content depicted in the identified portion of the new image satisfies the image criteria.
[00102] Optionally, in some implementations, the system also processes the new entity data when making this determination. For example, the relationships between the entities in the new image as represented by the new entity data may allow the system to more accurately generate the verification result. Because the new entity data is structured, this adds minimal computational costs.
[00103] In particular, the system can determine that the new image to satisfies the image criteria when the content depicted in the identified portion of the new image salsifies the image criteria. Alternatively, the system can determine that the new image fails to satisfy the image criteria when the content depicted in the identified portion of the new image does not meet the image criteria.
[00104] The system performs an action with respect to the new image based on the verification result (step 316). In some cases, the system can output the verification result as a response to a content provider computing device that generated the new image. In some cases, when the verification result indicates that the new image satisfies the image criteria, the system can provide the new image to a client device for display thereon. In some cases, when the verification result indicates that the new image fails to satisfy the image criteria, the system can modify (e.g., edit) or otherwise manipulate the content of the new image, such that the modified image would satisfy the image criteria.
[00105] FIG. 4 is a block diagram of an example computer system 400 that can be used to perform operations described above. The system 400 includes a processor 410, a memory 420, a storage device 430, and an input/output device 440. Each of the components 410, 420, 430, and 440 can be interconnected, for example, using a system bus 450. The processor 410 is capable of processing instructions for execution within the system 400. In one implementation, the processor 410 is a single-threaded processor. In another implementation, the processor 410 is a multi -threaded processor. The processor 410 is capable of processing instructions stored in the memory 420 or on the storage device 430.
[00106] The memory 420 stores information within the system 400. In one implementation, the memory 420 is a computer-readable medium. In one implementation, the memory^ 420 is a volatile memory unit. In another implementation, the memory7 420 is a non-volatile memory unit.
[00107] The storage device 430 is capable of providing mass storage for the system 400. In one implementation, the storage device 430 is a computer-readable medium. In various different implementations, the storage device 430 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity7 storage device. [00108] The input/output device 440 provides input/output operations for the system 400. In one implementation, the input/output device 440 can include one or more of a network interface devices, e.g., an Ethernet card, a serial communication device, e.g., and RS-232 port, and/or a wireless interface device, e.g.. and 802. 11 card. In another implementation, the input/output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 460. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[00109] Although an example processing system has been described in FIG. 4, implementations of the subject matter and the functional operations described in this specification can be implemented in other ty pes of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[00110] An electronic document (which for brevity will simply be referred to as a document) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.
[00111] For situations in which the systems discussed here collect and/or use personal information about users, the users may be provided with an opportunity to enable/disable or control programs or features that may collect and/or use personal information (e.g., information about a user’s social network, social actions or activities, a user’s preferences, or a user’s current location). In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user’s identity7 may be anonymized so that the no personally identifiable information can be determined for the user, or a user's geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.
[00112] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be. or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[00113] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[00114] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[00115] This document refers to a service apparatus. As used herein, a service apparatus is one or more data processing apparatus that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in block diagrams. However, while the service apparatus could be a single device or single set of devices, this disclosure contemplates that the service apparatus could also be a group of devices, or even multiple different systems that communicate in order to provide various content to client devices. For example, the service apparatus could encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other serv ice.
[00116] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may. but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[00117] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[00118] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g.. magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory', media and memory devices, including by7 way of example semiconductor memory7 devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks: magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by. or incorporated in, special purpose logic circuitry.
[00119] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well: for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser. [00120] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g.. a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN'’) and a wide area network (“WAN"), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[00121] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g.. for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[00122] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[00123] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[00124] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
[00125] What is claimed is:

Claims

1. A computer-implemented method comprising: maintaining a plurality of reference images that each satisfy one or more criteria and, for each of the plurality of reference images, corresponding reference entity data that specifies (i) a plurality of reference entities depicted in the reference image and (ii) reference relationships among the plurality of reference entities; obtaining a new image; generating new entity data for the new image that specifies (i) a plurality of new entities depicted in the new image and (ii) new relationships among the plurality of new entities depicted in the new image; identifying, for the new image, a reference image from the plurality of reference images; determining a difference between (i) the new entity data for the new image and (ii) the corresponding reference entity data for the identified reference image; identifying, based on the difference, a portion of the new image that represents a modification relative to the identified reference image; generating, based on the portion of the new image, a verification result that indicates whether the new image satisfies the one or more criteria; and performing an action with respect to the new image based on the verification result.
2. The method of claim 1, wherein: determining the difference comprises determining that the new entity data represents a new entity that is not represented by the reference entity' data; and identifying the portion of the new image comprises identifying a portion of the new- image that depicts at least the new entity.
3. The method of claim 2, wherein identifying the portion of the new image comprises: identify ing the portion of the new image that depicts the new entity and also one or more entities that are related to the new' entity' and that are represented by the reference entity data.
4. The method of claim 1, wherein: determining the difference comprises determining that the new entity data represents a new relationship between tw o existing entities, wherein the two existing entities are represented by the reference entity data but the new relationship is not represented by the reference entity data: and identifying the portion of the new image comprises identifying a portion of the new image that depicts at least the two existing entities.
5. The method of any one of claims 1-4, wherein for each of the plurality of reference images, the corresponding reference entity’ data comprises a graph that includes nodes representing the plurality of reference entities and edges representing the reference relationship among the plurality of reference entities.
6. The method of any one of claims 1-5, wherein the reference relationship comprises one or more of: a spatial relationship, a temporal relationship, a semantic relationship.
7. The method of claim 6, wherein the spatial relationship represents an object in physical contact with or in proximity’ to another object.
8. The method of claim 6, wherein the semantic relationship represents a human pointing toward or leading an eye to another human or another object.
9. The method of any one of claims 1-8, wherein obtaining the new image comprises: receiving an existing image and a text prompt; and generating the new image from the input image and the text prompt by using a generative machine learning model.
10. The method of claim 9, wherein identifying, for the new image, the reference image from the plurality of reference images comprises: identifying, as the reference image, the existing image from the plurality of reference images.
11. The method of any one of claims 9-10, wherein generating the new entity' data for the new image comprises: generating the new entity data based on the text prompt.
12. The method of any one of claims 1-10, wherein generating the new entity data for the new image comprises: performing image segmentation on the new image to partition the new image into a plurality of segments each depicting a new entity; and generating the new entity data based on the plurality of segments.
13. The method of any one of claims 1-12, determining whether the new image satisfies the one or more criteria comprises: determining, based on the portion of the new image and the new entity data, whether the new image satisfies the one or more criteria.
14. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the operations of the respective method of any one of claims 1- 13.
15. A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1-13.
EP23848678.1A 2023-12-29 2023-12-29 Resource efficient image content verification Pending EP4699014A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2023/086347 WO2025144411A1 (en) 2023-12-29 2023-12-29 Resource efficient image content verification

Publications (1)

Publication Number Publication Date
EP4699014A1 true EP4699014A1 (en) 2026-02-25

Family

ID=89853693

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23848678.1A Pending EP4699014A1 (en) 2023-12-29 2023-12-29 Resource efficient image content verification

Country Status (2)

Country Link
EP (1) EP4699014A1 (en)
WO (1) WO2025144411A1 (en)

Also Published As

Publication number Publication date
WO2025144411A1 (en) 2025-07-03

Similar Documents

Publication Publication Date Title
CN113704531B (en) Image processing method, device, electronic equipment and computer readable storage medium
US11120093B1 (en) System and method for providing a content item based on computer vision processing of images
CN108446390B (en) Method and device for pushing information
US10593085B2 (en) Combining faces from source images with target images based on search queries
US20100312608A1 (en) Content advertisements for video
KR20250114569A (en) Messaging system with trend analysis of content
EP3267333A1 (en) Local processing of biometric data for a content selection system
CN106202574A (en) The appraisal procedure recommended towards microblog topic and device
CN113569888A (en) Image labeling method, device, equipment and medium
US20230067628A1 (en) Systems and methods for automatically detecting and ameliorating bias in social multimedia
JP2020536332A (en) Keyframe scheduling methods and equipment, electronics, programs and media
US20240020336A1 (en) Search using generative model synthesized images
JP2021510216A (en) How to classify and match videos, equipment and selection engine
WO2025031067A1 (en) Image processing method and apparatus, device, and computer readable storage medium
CN107818160A (en) Expression label updates and realized method, equipment and the system that expression obtains
CN112104914A (en) Video recommendation method and device
US20240331107A1 (en) Automated radial blurring based on saliency and co-saliency
JP2025540007A (en) Video processing method, device, electronic device and storage medium
CN115604510B (en) A video recommendation method, apparatus, computer device, and storage medium
US20250315986A1 (en) Generative artificial intelligence
CN118152609B (en) Image generation method, device and computer equipment
CN118071867B (en) Method and device for converting text data into image data
CN115935049A (en) Artificial intelligence-based recommendation processing method, device and electronic equipment
EP4699014A1 (en) Resource efficient image content verification
CN114817697A (en) Method and device for determining label information, electronic equipment and storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251120

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR