EP4634827A1 - Multi-attribute combined embedding model - Google Patents

Multi-attribute combined embedding model

Info

Publication number
EP4634827A1
EP4634827A1 EP24714344.9A EP24714344A EP4634827A1 EP 4634827 A1 EP4634827 A1 EP 4634827A1 EP 24714344 A EP24714344 A EP 24714344A EP 4634827 A1 EP4634827 A1 EP 4634827A1
Authority
EP
European Patent Office
Prior art keywords
combination
embedding
embeddings
layout
query
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP24714344.9A
Other languages
German (de)
French (fr)
Inventor
Yu Chen
Xiaohang Li
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Google LLC
Original Assignee
Google LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Google LLC filed Critical Google LLC
Publication of EP4634827A1 publication Critical patent/EP4634827A1/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/71Indexing; Data structures therefor; Storage structures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/41Indexing; Data structures therefor; Storage structures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/48Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/483Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/50Information retrieval; Database structures therefor; File system structures therefor of still image data
    • G06F16/51Indexing; Data structures therefor; Storage structures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/50Information retrieval; Database structures therefor; File system structures therefor of still image data
    • G06F16/55Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/50Information retrieval; Database structures therefor; File system structures therefor of still image data
    • G06F16/58Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/583Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/73Querying
    • G06F16/732Query formulation
    • G06F16/7328Query by example, e.g. a complete video frame or video sequence
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/73Querying
    • G06F16/735Filtering based on additional data, e.g. user or group profiles
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/75Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/783Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/783Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • G06F16/7844Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using original textual content or text extracted from visual content or transcript of audio data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/783Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • G06F16/7847Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using low-level visual features of the video content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0499Feedforward networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/762Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks

Definitions

  • This specification relates to data processing, and the generation of combination embeddings corresponding to a set of two or more attributes of digital components.
  • Machine learning models that are implemented to evaluate content (e.g., images, text, video, etc.), or otherwise operate in a real-time serving environment (e.g., where content is identified and served to a client device within less than a specified number of milliseconds) require significant computing resources to evaluate the content, or otherwise generate a prediction, within a specified time constraint.
  • content e.g., images, text, video, etc.
  • real-time serving environment e.g., where content is identified and served to a client device within less than a specified number of milliseconds
  • machine learning models to generate the prediction machine learning models generally require a significant amount of energy, computer storage, computer memory, and processing power, particularly in a real-time serving environment. Methods which reduce such resource consumption are therefore desirable.
  • each digital component is one of text, an image, or a video
  • each layout specifies one or more formatting attributes that are applicable to the set of digital components
  • creating a set of layout-digital component combinations wherein each layout-digital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components
  • creating a training embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination
  • training a machine learning model based on the training embeddings and the set of layout-digital component combinations to obtain a combination embedding model
  • identifying a set of creatives each comprising (i)
  • Methods can include filtering the set of combination embeddings based on one or more metrics of each creative having a combination embedding in the set of combination embeddings.
  • Methods can include receiving a query; identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
  • Methods can include identifying a query embedding representing a combination of multiple different features of the query; and submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query-specific ranking based on the one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
  • Methods can include obtaining a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding.
  • Methods can include creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per- combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
  • Methods can include generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings.
  • a system can include memory and one or more processors, wherein the memory stores computer program instructions which, upon execution by the processor, cause the processor to perform operations of one or more methods.
  • the operations can include receiving a set of digital components, wherein each digital component is one of text, an image, or a video; receiving a set of layouts, wherein each layout specifies one or more formatting attributes that are applicable to the set of digital components; generating a set of layout-digital component combinations, wherein each layout-digital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components; creating a combination embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination; training a machine learning model based on the combination embeddings and the set of layout-digital component combinations; after training the machine learning model, identifying a set of creatives each comprising (i) a combination of digital components and (ii) a given layout; generating, by the machine learning model, a set of combination embedd
  • the operations can include receiving a query; identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
  • the operations can include identifying a query embedding representing a combination of multiple different features of the query; and submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query-specific ranking based on the one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
  • the operations can include obtaining a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding.
  • the operations can include creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per- combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
  • the operations can include generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings.
  • the operations discussed above can also be implemented as instructions stored on a computer-readable medium. Execution of the instructions can cause one or more data processing apparatus to perform the operations.
  • the techniques discussed throughout this specification can reduce the amount of time required to evaluate/select a portion of content (e.g., creative) that is ultimately served to a user in response to a content request.
  • content e.g., creative
  • current systems utilize a separate machine learning model for each different attribute of the content, which each separately generate embeddings of the attribute, and then use the separate embeddings to generate an evaluation (e.g., prediction) based on the separate embeddings.
  • the techniques discussed herein train a model to generate a combination embedding that represents multiple different attributes (e.g., a combination of two or more attributes) of the content to be evaluated. In this way, one machine learning model can be used in the place of two or more different models, while still enabling the evaluation (e.g., prediction) to be based on multiple different attributes of the content.
  • a model can be trained to create a single combination embedding representing a combination of text included in a given portion of content, an image/video included in the given portion of content, and a layout of the given portion of content.
  • existing systems would use three separate models to separately create different embeddings for each of the text, image/video, and layout, resulting in the creation of at least three embeddings by three different models.
  • only one model is required to be used to obtain the single combination embedding, thereby reducing the amount of processing power, memory, and time required to obtain an embedding representing the multiple different attributes of the given portion of content. This makes systems implementing the techniques discussed herein more efficient relative to existing systems.
  • the efficiencies realized by the present techniques are particularly important when evaluating content for distribution to users in response to a request for content because there are often millions of available portions of content that can be served in response to a request for content, such that multiplying the number of embeddings that need to be created by separate models and evaluated multiplies the amount of processing required to be performed.
  • the models required to create embeddings require a large amount of memory and processor usage. As such, being able to create a single embedding that represents multiple different attributes of a given portion of content using a single model substantially reduces the memory and processor usage that is required to represent the given portion of content using embeddings.
  • the ability to create the combination embedding using a single model is achieved by training the model using a training set of content and the separate embeddings for the different attributes of portions of content in the training set of content. For example, separate embeddings can be created for each attribute of the content to be evaluated. The separate embeddings can then be concatenated to create a combination embedding that represents the multiple attributes of the content, and the model can be trained to directly generate combination embeddings based on the combination of attributes of the content. Once trained, the model can accept, as input, a given portion of content and create, as output, the combination embedding representing the combination of attributes of the given portion of content. As such, a single model can create a single combination embedding, which can be used to evaluate the given portion of content rather than having to generate multiple different embeddings using different models.
  • Another advantage provided by the techniques discussed herein is a solution to the “cold start” problem of content evaluation in existing systems. More specifically, existing systems assign random identifiers (e.g., numbers) to different digital components (e.g., text portions, images, videos) and/or combination of digital components (e.g., a combination of text and an image). Because these identifiers are randomly assigned, they do not carry any semantic information, such that two given portions of content (e.g., combinations of text and images) that have the same attributes are not linked, or otherwise identified as semantically similar. Thus, data collected/generated with respect to one of the two given portions of content is not available for use when evaluating the second of the two given portions of content.
  • identifiers e.g., numbers
  • a new given portion of content e.g., combination of digital components having a certain layout
  • the existing systems must independently collect data for purposes of evaluating the given portion of content, rather than leveraging data collected for similar portions of content. This adds delay in the ability of existing systems to be able to accurately evaluate new portions of content, and also requires duplication of operations that have already been performed with respect to the similar portions of content, resulting in wasted processing/memory resources.
  • systems implemented using the techniques discussed herein need not store redundant data about semantically similar portions of content. Rather, systems implemented using the techniques discussed herein can store a single instance of data and associate (e.g., index) that single instance of stored data with all of the semantically similar portions of content (e g., having similar text/image/layout).
  • the combination embeddings of semantically similar portions of content will be similar (e.g., within a specified distance in space), such that the portions of content can be grouped/clustered together as being semantically similar.
  • each similar portion of content can be assigned, associated with, indexed, or otherwise linked to the data collected for each member of the group/cluster.
  • the new portion of content can be immediately assigned to a group of semantically similar portions of content, and the data that has already been collected for the semantically similar portions of content can be used for the new portion of content, thereby solving the “cold start” problem.
  • newly collected data for the new portion of content can also be stored in association with the other content in the cluster/group, such that the data can be shared among the members of the cluster/group, thereby eliminating the need to store separate instances of data for each of the group members. This is a significant memory savings, particularly in the context of online content where there can be millions/billions of different instances of content for which data needs to be stored.
  • FIG. 1 illustrates an example environment in which a combination embedding can be used in the context of serving content to client devices.
  • FIG. 2 illustrates an example of two digital components displayed according to different layouts.
  • FIG. 3 illustrates a flow chart of an example process for creating combination embeddings and a machine learning model for producing new combination embeddings.
  • FIG. 4 illustrates an example diagram showing details of the generation of a combination embedding.
  • FIG. 5 illustrates how including a combination embedding with a query embedding can lead to generating a metric.
  • FIG. 6 is a flow chart of an example process for using combination embeddings to distribute creatives.
  • FIG. 7 is a flow chart of an example process for using combination embeddings to distribute creatives.
  • FIG. 8 illustrates an example of a computing device.
  • ID identifier
  • these ID tags are assigned to digital components when uploaded to a server, and each uploaded digital component is assigned a different ID tag.
  • this ID tag does not carry any semantic information about the content of the digital component, and similar digital components are not identifiable using anu information in the ID tag.
  • each digital component could be assigned an embedding which captures or includes information about that digital component, such that similar digital components would be identifiable based on the embeddings.
  • combination embeddings can be used to represent the content of creatives, which include multiple digital components that are formatted according to a layout, such that similar creatives would be identifiable on the basis of the combination embedding. Including such information in an embedding assigned to the combination of digital components (e.g., creatives) with a particular layout scheme would enable a more efficient use of computer resources to optimize such displays as well as to enable more effective use of the information by capturing similarities between similarly displayed combinations of digital components and also to optimize the particular layout which may be most effective.
  • the present specification describes training machine learning models to assign combination embeddings to creatives that include multiple digital components formatted according to a layout.
  • the trained machine learning models accept a creative as an input and output the combination embedding representing the content of that creative.
  • a new combination embedding can be easily and efficiently created by a single model, rather than having to process the creative, or digital components, using multiple models as currently required.
  • These combination embeddings may be combined with an embedding of a query to generate a metric prediction, which may be used to rank the efficacy of a particular combination of digital components with a layout.
  • digital component refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content).
  • a digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.
  • a combination of digital components can be referred to as a creative, which can be formatted according to a layout.
  • FIG. 1 is a block diagram of an example environment 100 in which a combination embedding can be used in the context of context of serving content to client devices.
  • a combination embedding can be used in the context of context of serving content to client devices.
  • using the combination embeddings in the environment would help reduce computer resource demands and enable the evaluation of more digital components or more layouts or result in new insights into how users perceive the display of digital components displayed in different arrangements.
  • the example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof.
  • the network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and a service apparatus 110.
  • the example environment 100 may include many different electronic document servers 104, client devices 106, and digital component servers 108.
  • a client device 106 is an electronic device capable of requesting and receiving online resources over the network 102.
  • Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, wearable computing devices, and other devices that can send and receive data over the network 102.
  • a client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102, but native applications (other than browsers), which are also referred to as “apps”, that are executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.
  • a gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application.
  • a gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application.
  • the gaming device can store and execute the gaming application locally or execute a gaming application that is at least partly stored and/or served by a cloud server (e.g., online gaming applications).
  • the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device.
  • the gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.
  • Di gital assistant devices include devices that include a microphone and a speaker.
  • Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information.
  • digital assistant devices also include a visual display or are in communication with a visual display (e g., by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present.
  • digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.
  • the client device 106 is presenting an electronic document 150.
  • An electronic document is data that presents a set of content at a client device 106.
  • Examples of electronic documents include webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources.
  • Native applications e.g., “apps” and/or gaming applications
  • Electronic documents can be provided to client devices 106 by the electronic document servers 104 (“Electronic Doc Servers”).
  • the electronic document servers 104 can include servers that host publisher websites.
  • the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate presentation of the given webpage at the client device 106.
  • the electronic document servers 104 can include application servers (e.g., native app servers) from which client devices 106 can download apps.
  • the client device 106 can download files required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device).
  • the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server.
  • the cloud server can execute the application and stream a user interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app and communicate any user interactions with the user interface back to the cloud server for processing.
  • Electronic documents can include a variety of content.
  • an electronic document 150 can include native content 152 that is within the electronic document 150 itself and/or does not change over time.
  • Electronic documents can also include dynamic content that may change over time or on a per-request basis.
  • a publisher of a given electronic document e.g., electronic document 150
  • the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component or creative) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server).
  • the client device 106 integrates the content (e.g., digital component or creative) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source.
  • content e.g., digital component or creative
  • the following discussion of selecting content responsive to a request uses digital components for purposes of example, but the discussion is equally applicable to selecting creatives for distribution.
  • a given electronic document e.g., electronic document 150
  • can include a digital component script e.g., script 154
  • the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106.
  • Execution of the digital component script configures the client device 106 to generate a request for digital components 112 (referred to as a “component request” or “submitted user request”), which is transmitted over the network 102 to the service apparatus 110.
  • the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data.
  • the component request 112 can include event data specifying features such as a name (or network location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e.g., the client device 106), and/or information that the service apparatus 110 can use to select one or more digital components, or other content, provided in response to the request.
  • the component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.
  • the component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which digital component can be presented.
  • event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and/or media types that are eligible for presentation in the locations can be provided to the service apparatus 110.
  • event data specifying keywords associated with the electronic document (“document keywords”) or entities (e.g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provided to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document.
  • the event data can also include a search query that was submitted from the client device 106 to obtain a search results page.
  • Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device).
  • Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can be formatted as packetized data having a header and payload data.
  • the header can specify a destination of the packet and the payload data can include any of the information discussed above.
  • the service apparatus 110 chooses digital components (e.g., third-party content, such as video files, audio files, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and/or using information included in the component request 112.
  • digital components e.g., third-party content, such as video files, audio files, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content
  • a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delays in providing digital components and layouts in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.
  • the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital component in response to requests 112.
  • the set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a corpus of millions of available digital components (DCi - DC X ).
  • the millions of available digital components can be indexed, for example, in a digital component database 116.
  • Each digital component index entry can reference the corresponding digital component and/or include distribution parameters (DPi-DP x ) that contribute to (e.g., trigger, condition, or limit) the distribution/transmission of the corresponding digital component.
  • distribution parameters DPi-DP x
  • the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.
  • a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.
  • the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and/or consumption details (e.g., page viewport, user scrolling speed, or other information about the consumption of data).
  • the distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and/or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation.
  • the distribution parameters can also specify an eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components).
  • the identification of the eligible digital component can be segmented into multiple tasks 117a-l 17c that are then assigned among computing devices within the set of multiple computing devices 114.
  • different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match information included in the component request 112.
  • each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a- 118c of the analysis back to the service apparatus 110.
  • the results 118a-l 18c provided by each of the computing devices in the set 114 may identify a subset of digital components that are eligible for distribution in response to the component request and/or a subset of the digital component that have certain distribution parameters.
  • the identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying the subset of digital components having distribution parameters that match at least some features of the event data.
  • the service apparatus 110 aggregates the results 118a- 118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components/ creatives) based on the outcome of one or more content and/or layout evaluation processes, as discussed below.
  • the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.
  • reply data 120 e.g., digital data representing a reply
  • the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.
  • the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components and layouts from one or more digital component servers 108.
  • the instructions in the reply data 120 can include a network location (e g., a Uniform Resource Locator (URL)) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component and layout combination from the digital component server 108.
  • a network location e g., a Uniform Resource Locator (URL)
  • SR server request
  • the client device 106 When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the layout or by a script 154 from the client device.
  • the script 154 can create a walled garden environment, such as a frame, that is presented within, (e.g., beside), the native content 152 of the electronic document 150.
  • the digital component is overlay ed over (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify the presentation layout within the electronic document 150 in the reply 120.
  • the native content 152 includes video content
  • the service apparatus 110 can specify a layout within the scene depicted in the video content over which the digital component is to be presented.
  • the service apparatus 110 can also include an artificial intelligence system 160 configured to autonomously generate digital components and layouts, either prior to a request 112 (e.g., offline) and/or in response to a request 112 (e.g., online or real-time).
  • the artificial intelligence (“Al”) system 160 can collect online content about a specific entity (e.g., digital component provider or another entity) and summarize the collected online content using one or more language models 170, which can include large language models.
  • LLM large language model
  • LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.
  • the language model 170 can be any appropriate language model or neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto-regressively generates an output sequence made up of text tokens from the vocabulary.
  • the language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model.
  • the language model 170 can be referred to as an autoregressive neural network when the neural network used to implement the language model 170 auto-regressively generates an output sequence of tokens.
  • the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and a context input that provides context for the output sequence.
  • the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence.
  • the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence.
  • the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
  • the language model 170 can be an auto-regressive Transformer-based neural network that includes (i) a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.
  • the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence.
  • the attention block then updates each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens.
  • the input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.
  • the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.
  • the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model 170 to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.
  • the service apparatus 110 pre-trains the language model 170 (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data.
  • the language model 170 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e g., text that is publicly available from the Internet or another text corpus.
  • the Al system 160 can generate a prompt 172 that is submitted to the language model 170 and causes the language model 170 to generate the output sequences 174, also referred to as passages or simply as “output”.
  • the Al system 160 can generate the prompt in a manner (e.g., having a structure) that identifies a list of online sources of information, such as a list of websites or data repositories, and specifying a set of constraints the language model 160 must use to generate a summary of information found at the online sources specified in the prompt 172.
  • the Al system 160 submits the prompt 172 to the one or more language models 170, which use the prompt 172 to evaluate the information found at the online sources specified in the prompt 172 and generate the output 174 that summarizes the information according to the constraints specified in the prompt 172.
  • the Al system 160 can use the generated summary as part of another prompt 172 that is sent to the language model 170.
  • the Al system 160 can insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary) that is submitted to the language model 170 as a constraint for generating clauses for use in digital components being generated by the Al system 160.
  • the Al system 160 is generating a digital component to provide in response to the request 112, which includes a keyword/query.
  • the Al system 160 can generate the additional prompt 172 to include the query and a set of constraints including the summary received in the prior output 174.
  • the set of constraints of the additional prompt 172 can also include instructions regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, semantically styled, among other things (e.g., specifying content that should be excluded from the clauses, such as granular details, such as numbers).
  • the additional prompt 172 could take the following form:
  • the Al system 160 can use the generated summary as part of another prompt 172 that is sent to the language model 170.
  • the Al system 160 can insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary) that is submitted to the language model 170 as a constraint for generating clauses for use in digital components being generated by the Al system 160.
  • additional prompt 172 e.g., a prompt generated after receiving the summary
  • the Al system 160 is generating a digital component and a layout to provide in response to the request 112, which includes a keyword/query.
  • the Al system 160 can generate the additional prompt 172 to include the query and a set of constraints including the summary received in the prior output 174.
  • the set of constraints of the additional prompt 172 can also include instructions (e.g., layout instructions) regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, semantically styled, among other things (e.g., specifying content that should be excluded from the clauses, such as granular details, such as numbers).
  • the additional prompt 172 could take the following form:
  • good output must be in bullet-point format, good output must have exactly 3 bullet-points. Each bullet-point must be less than 90 characters, good output must have no nested bullets. good_output must be catchy and show valueprop. good output must be useful and informative, and must avoid boring details like numbers.
  • the Al system 160 is providing the language model
  • a query constraint specifies the query “10G Network” to which the output clauses should be relevant.
  • entity constraint specifies “example_network_provider” as the entity name to use in the output clauses.
  • a summary constraint specifies the content summary to use during clause generation, i.e., “200 Mbps internet with WiFi...”.
  • submission of this additional prompt 172 to the language model 170 causes the language model to generate an additional output 174, which includes multiple sets of clauses generated according to the query and constraints, which is communicated electronically to the Al system 160.
  • the Al system 160 receives the clauses of the additional output 174 and generates multiple candidate digital components and layouts that could be provided in response to the request 112.
  • each different candidate digital component and layout includes a different combination of the clauses received from the language model 170 in the additional output 174.
  • the Al system 160 could also create the candidate digital components using a set of different links to online content (e.g., second level domain links to web pages discussing a topic of the candidate digital components, phone numbers, etc.), which can continue to exponentially increase the number of different candidate digital components that the Al system 160 can create using the clauses of the additional output 174 of the language model 170.
  • the Al system 160 can perform one or more post-processing operations that evaluate one or more characteristics of the multiple candidate digital components and layouts.
  • the post-processing operations can include generating a grounding score for each clause from the additional output 174.
  • the grounding score is a value specifying a likelihood that the clause is factual.
  • the language model 170 (or more generally the Al system 160) can be configured to generate an output using a combination of input text alone, or using a combination of text and images, audio, video, or other audio/visual content.
  • the language model can accept the input of an image of a brown dog with the text “generate an image of the dog with white spots.”
  • the language model 170 can process the image/text input, create a new image of the brown dog with white spots, and output the new image in response to the input.
  • the language model 170 can be used to generate a large volume of modified images and/or text in a very short period of time.
  • FIG. 2 illustrates how a combination of multiple digital components are arranged together according to a layout.
  • Digital components may include different content types, such as image digital components, video digital components, text digital components, audio components and the like. Each of these digital components may be presented in a different fashion depending on the specified layout (e.g., dimensions and/or arrangement of digital components in a resulting creative or combination of digital components).
  • an image digital component 202 (“Image DC”) may be combined with a text digital component 204 (“Text DC”) in multiple ways depending on the constraints specified by the layout 206.
  • the layout 206 specifies how the image DC 202 and the text DC 204 are arranged when presented at a client device. In the example illustrated in FIG.
  • the text DC 204 is the phrase “Men’s & Women’s Clothes” and the image DC 202 is a standard image of a stylized man and a stylized woman. Combinations of these two DCs may be presented in a variety of ways such as a first combination 208-1 in which the text DC 204 is presented above the image DC 202, a second combination 208-2 in which the text DC 204 is presented below the image DC 202, a third combination 208-3 in which the text DC 204 is presented to the left and to the right of the image DC 202. Other combinations of layouts and DCs (208-N) may also be used and the number is not limited to those displayed in this figure.
  • the layout 206 could specify different outer dimensions (or aspect ratios) for the resulting creative (e.g., combination of digital components), which can result in different fonts/font sizes being used and/or different arrangements of the text DC 202 and/or image DC 204.
  • the image DC 204 may need to be resized, cropped, or reformatted to fit within the outer boundaries of the resulting creative.
  • the layout 206 may include instructions of where to place a certain visual DC 202 in relation to a text DC 204 but is not limited thereto.
  • the layout 206 may also provide the timing of display of certain DCs, such as when an audio DC might play over a speaker of the user’s electronic device.
  • the layout 206 may provide the font or details of the text DC 204 to be displayed, such as the font typeface and whether the text should be displayed using a bold or an italic style. In general, the layout 206 describes the formatting details, arrangement, size, and location of each of the digital components.
  • the layout 206 my include physical arrangements of each DC in relation to every other DC, but can also include temporal aspects amongst the digital components such as when a video DC may play and whether to include an audio DC while the video DC is playing or whether to wait until a user has indicated an interest in listening to the audio DC.
  • a visual DC 202 or a text DC 204 may also be located in different locations on the display of a user’s electronic device and not just in relation to each other.
  • the layout 206 may indicate an absolute location on the display rather than a location relative only to the other digital components.
  • a video DC may be presented to a user in one of several locations such as bottom left, bottom right, upper left, upper right, and middle of the display of the user’s electronic device.
  • a combination of one or more digital components DCi and a layout 206 is referred to as a creative.
  • the layout may include an arrangement of a first digital component to a second digital component.
  • the layout 206 may also be much more complicated if the number of digital components is larger and each needs to be specified in relation to every other digital component.
  • Each layout 206 may be assigned an identification (ID) number.
  • the layout ID may be assigned at random, though without permitting duplicate IDs.
  • the layout ID may be assigned based on when the layout was added to the set of layouts.
  • the layout 206 may also be assigned an ID which is indicative of the actual arrangement of the digital components DCi.
  • the layout ID may be an embedding of the layout instructions or an embedding of the arrangement of the various digital components.
  • the training of a combination embedding model can be performed using multiple different combinations of digital components.
  • the system can create every combination of stored digital components according to each stored layout. This provides a robust training set that will enable the trained model to accurately assign a combination embedding to a large variety of newly obtained/identified creatives.
  • the digital components can also be received from a third-party corpus of digital components.
  • a third-party corpus of digital components For example, an entity implementing the process 300 can obtain access to a corpus of digital components and/or layouts that are maintained by a third party.
  • each layout-digital component combination includes a layout from the set of layouts that was obtained and at least one digital component from the set of digital components that was obtained.
  • layout-digital component combinations can include multiple different digital components.
  • Each of the layoutdigital component combinations can be considered a separate creative/portion of content.
  • each possible combination of digital components and layouts can be created to provide a robust training set for training the combination embedding model.
  • the system can iteratively create different combinations of digital components and format each of the different combinations according to a different layout until all possible combinations of digital components and layouts has been created.
  • fewer than all possible combinations can be created, for example, depending on timing constraints, model training evaluation, etc. For instance, if the combination embedding model accuracy reaches an acceptable level without having to create every possible layout-digital component combination, the creation of layout-digital component combination can be halted.
  • a target number of layout-digital component combination can be created to perform a specified number of training samples, and this target number can be fewer than all possible combinations.
  • an embedding of the layout of the given layout-digital component combination can be obtained from a database or generated using a layout encoder in a similar fashion as the text embedding and the image embedding.
  • the layout identifier need not be an embedding of the layout. For example, if each layout is identifiable by an assigned layout number and similar layouts have similar numbers, the assigned layout number can be used as the layout identifier rather than creating an embedding of the layout. Creation of embeddings is discussed in more detail with reference to FIG. 4.
  • the digital component embeddings and the layout identifier for a given layout-digital component combination can be concatenated, thereby creating a training embedding.
  • the obtained text embedding is ABCD
  • the obtained image embedding is EFGH
  • the obtained layout identifier is 1234.
  • the resulting training embedding corresponding to the given layout-digital component combination can be ABCDEFGH1234, which is a concatenation of the digital component embeddings and the layout identifier and can be used as training data for training the combination embedding model. Similar training embeddings can be created for each given layout-digital component combination.
  • a machine learning model is trained based on the training embeddings and the set of layout-digital component combinations (308).
  • a given layout-digital component combination and its corresponding training embedding can be considered a training pair that is used to train the machine learning model to generate a combination embedding output based on the input of a creative (e.g., combination of digital components formatted according to a layout).
  • a creative e.g., combination of digital components formatted according to a layout.
  • Various different machine learning techniques can be used as part of the machine learning model process. For example, a multi-axis vision transformer architecture or another attention model architecture can be used.
  • Vision transformers may be included in the model which may receive an image and produce, for example, text which describes the image.
  • the descriptive text may then be used to produce an embedding, for example, with the help of a large language model.
  • the model may then be trained to produce combination embeddings from a new combination of digital components and layouts.
  • the trained model can be referred to as a combination embedding model because it is trained to output a combination embedding based on the input of a creative (e.g., combination of digital components formatted according to a layout).
  • the training can be performed using unsupervised machine learning techniques.
  • a set of creatives is identified after the training (310).
  • each of the creatives in the set of creatives includes (i) a combination of digital components and (ii) a given layout.
  • the creatives can be uploaded by a user of the service apparatus 110 of FIG. 1.
  • a user who uses the service apparatus 110 to distribute creatives through the network 102 can upload an already assembled creative that includes text and an image that is formatted according to a layout.
  • the creative can be dynamically generated by the service apparatus 110 based on a set of digital components uploaded by the user.
  • the service apparatus 110 can use information obtained from a content request to generate a new creative (e.g., using the Al system 160). In these situations, the creative may not have existed prior to the creation of the creative by the service apparatus 110. Creatives for a large number of users of the service apparatus 110 can be identified for potential evaluation.
  • a set of combination embeddings are generated for the set of creatives (312).
  • the combination embeddings are generated by the combination embedding model, which was trained to generate combination embeddings using the creatives as input. That is, a combination embedding can be generated for each creative in the set of combination embeddings.
  • the combination embedding model generates a single combination embedding that represents the combination of digital components that are included in the creative as well as the formatting of the digital components/creative according to a layout. This is in contrast to existing systems, which use different models to create different embeddings for each different digital component type that might be included in a creative.
  • the use of the combination embedding model results in only one embedding being generated by one model, such that the system is more efficient relative to systems that use multiple different models to generate multiple different embeddings.
  • the set of combination embeddings can be filtered based on one or more metrics (314). For example, for each given creative having a combination embedding in the set of combination embeddings the given creative can be evaluated and filtered based on the one or metrics, as discussed below.
  • the number of combination embeddings required will be much greater than the number of embeddings produced by multiple models so that filtering may be required in some instances to reduce the overall number of combination embeddings being evaluated.
  • the filtering is an optional operation in some instances.
  • the filtering may be implemented in situations where the number of creatives to be evaluated is larger than the number of creatives that can be evaluated within a specified time constraint and/or using a specified amount of limited processing resources (e.g., cores, processors, memory, etc.).
  • a specified amount of limited processing resources e.g., cores, processors, memory, etc.
  • the number of queries e.g., search queries, keywords, or topics
  • the evaluation of all combinations of digital components for all queries would result in 10 billion evaluations being required.
  • filtering may be performed to reduce the number of evaluations that are required to be performed.
  • the purpose of the filtering is to select a subset of the creatives generated using the combinations of digital components that are most likely to be the best candidates for distribution prior to performing the full evaluation of the creatives.
  • the combination embeddings have been generated for each of the creatives generated using combinations of the digital components. These combination embeddings can therefore be used to perform the filtering.
  • each given combination embedding among the combination embeddings can be used to search a database of performance metrics for matching combination embeddings that match (e.g., are identical or within a specified distance in a vector space) the given combination embedding of one of the creatives/combinations of digital components.
  • the performance metrics for the matching combination embedding can be used as the predicted performance metrics of the given combination embedding.
  • a simple lookup operation can be performed to identify a predicted performance metric for the given combination embedding rather than having to use a predictive model, which is much more computationally intensive and time consuming than a lookup operation, to predict the performance of the creative associated with the given combination embedding.
  • the performance metric used for the fdtering is not critical to the techniques discussed herein and can be selected by an administrator of the system based on the goal to be achieved.
  • simple performance metrics that can be used include a user interaction rate of the creative (e.g., click through rate, video view through rate, survey response rate, conversion rate, etc.)
  • the metric being used to perform the fdtering is a click-through-rate.
  • a certain creative including a combination of digital components and having a given layout has been used and is known to have a certain CTR
  • a new creative with a combination embedding similar to that of the certain creative may be expected to have a similar CTR.
  • the new creatives can then be filtered based on the expected or predicted CTR.
  • the top X number or top Y% of combination embeddings according to the predicted CTR (or other metric) can be selected as combination embeddings to keep, and creatives that do not have combination embeddings matching those selected combination embeddings can be filtered out, or otherwise removed from consideration such that processing resources and time are not spent fully evaluating those creatives.
  • a set of 10 billion creatives can be filtered down to a more manageable set of creatives, such as 100 thousand creatives, or another specified number.
  • the filtering described above can be performed in an offline state.
  • an offline state refers to a state in which operations are not being performed in response to a query (or other content request) that was submitted by a user.
  • the offline processing is performed independently of (e.g., before) the system being requested to provide a response to a specific user request.
  • this processing can be utilized to reduce the number of creatives that are further evaluated for potential distribution to users, at least responsive to a given query.
  • Offline filtering may also be performed in initially and then followed by online filtering by using a querydependent model to further reduce the number of candidate creatives.
  • Another form of filtering can be performed in an online state.
  • the online state is a state in which the system is being requested to provide a response to a user query (or other content request).
  • the online processing is performed after a specific user has submitted a specific content request, such that the user is awaiting a response to the specific content request.
  • the filtering in the online state can utilize predictive models to further reduce the number of creatives that are further analyzed for potential distribution by a final selection evaluation.
  • the predictive models can utilize various attributes/features of the creatives available for evaluation (e.g., after the offline filtering) at the time the specific content request and/or a portion of event data included in the users request. Because a user is awaiting a response form the system during the online filtering, the evaluation is performed within a specified time constraint (e.g., within 500 milliseconds).
  • the features/attributes of the creatives are submitted to a predictive performance model (e g., a predictive CTR model), which returns a predicted performance measure for each of the creatives. Using these predicted performance measures, the system can select a specified number, or percentage, of creatives having the highest predicted performance as those creatives that will be submitted to the final selection evaluation, while the unselected creatives will be filtered out, or otherwise removed from further consideration.
  • One or more creatives are distributed based, at least in part, on the set of generated combination embeddings (316).
  • the distribution of the one or more creatives can be performed, for example, based on the outcome of the final selection evaluation of those creatives remaining after one or more of the filtering operations discussed above.
  • the remaining creatives can be submitted for evaluation based on distribution criteria of the creatives in combination with specific characteristics of the environment in which the distributed creative will be presented.
  • the final selection evaluation can adjust the previously created predicted performance measures for the remaining creatives to arrive at a selection score.
  • the predicted performance measures can be adjusted based on distribution value (e.g., bid) or some other information associated with the creatives.
  • the selection evaluation, selection, and distribution of the remaining creatives can be performed, for example, as discussed above with respect to FIG. 1.
  • FIG. 4 illustrates an example 400 of how embeddings for various digital components may be incorporated into a combination embedding (e.g., as performed when creating the training embeddings).
  • Some example digital components include a visual DC 202, a text DC 204, video DCs, audio DCs, or others.
  • a layout 206 may also be included as may other features 408.
  • Each of the digital components (or a set of digital components) may be fed singly or as a group into an encoder such as an image encoder 412 or a text encoder 414.
  • the layout 206 may be fed into a layout encoder 416, or may be assigned an ID based on other information, as discussed elsewhere in this specification.
  • An advantage of using a layout encoder 416 is that similar layouts may be assigned highly similar embeddings allowing an easier selection of particularly good layouts, when the combinations of layouts and digital components are evaluated.
  • other features may be included through the use of a feature encoder 418 to create an embedding. Examples of additional features may include such items as the day of the week, calendar day, time of day, URL domain, and the like.
  • Each of the encoders 412, 414, 416, and 418 (or additional encoders as the case may be) produces an embedding such as a visual embedding 422, a text embedding 424, a layout embedding 426, or a feature embedding 428.
  • the digital component encoders 412, 414 may create embeddings (e.g., vector representations) representing the digital components (e.g., visual DC 202 or text DC 204) and/or other corresponding features. All these embeddings may be fed into a combination embedding generator 430 to create a combination embedding 440.
  • the combination embedding generator 430 may create the combination embedding 440 by concatenating the individual embeddings 412, 414, 416, and 418.
  • the combination embedding generator 430 may also be more complicated such as a trained machine learning model which uses other methods to generate a combined embedding 440.
  • An example of a visual DC encoder 412 may be a visual transformer. These are a type of self-supervised learning which extracts a deeper meaning or additional information from an image.
  • all the digital components, the layout, and the other features may be fed into a single encoder to produce the combination embedding.
  • the digital component embeddings 422, 424 may be combined in additional steps in the combination embedding generator 430.
  • concatenated feature embeddings e.g., concatenation of the visual DC embedding 422, the text embedding 424, the layout embedding 426, and the other feature embedding 428, may be fed into various functions such as a squeeze and excitation module, a feature interaction module, or a multi-head self-attenuation module.
  • the combination embedding generator 430 takes as a minimum input a digital component and a layout and outputs a combined embedding.
  • the generator 430 may take additional inputs such as additional digital components and may also take as additional inputs other features such features related to events.
  • the combination embedding generator 430 may also take other forms such a using a factorized neural network model trained on training data.
  • FIG. 5 illustrates how a combination embedding 440 may be used along with a query embedding to generate an output metric.
  • a query may be defined by query features 502 (e.g., terms, term expansions, query time, query origin location, query language, browser type, etc ).
  • the query features may be fed into a query embedding generator 530 to generate a query embedding 540.
  • a visual DC embedding 422, a text embedding 424, a layout embedding 426, and other feature embedding 428 may be fed into a combination embedding generator 430 to produce a combination embedding 440 of digital components with a layout 206.
  • the query embedding 540 and the combination embedding 440 may be used to determine an output metric 550.
  • a dot product of the combination embedding 440 and the query embedding 540 may be used to determine an output metric 550.
  • the combination embedding 440 and the query embedding may both be input into another model to generate the metric 550.
  • a trained combination embedding model can be used to generate a combination embedding directly from a creative that includes a combination of digital components arranged according to a particular layout, as discussed above.
  • the combination embedding output by the trained combination embedding model can be used with the query embedding in a manner similar to that described with reference to FIG. 5.
  • the system may generate a query embedding 540 and determine the set of highest ranked combination embeddings 440 for the query associated with the highest metrics 550 generated for the pair of the combination embedding 440 with the query embedding 540.
  • the system may cut off all combination embeddings 440 below a certain threshold value for the metric 550.
  • the system may rank all the combination embeddings 440 and select only a set number of the combination embeddings 440 for consideration.
  • the system may use other methods to fdter or limit the total number of combinations of digital components and layouts.
  • FIG. 6 is a flow chart of an example process 600 for using combination embeddings to distribute creatives.
  • the process 600 can be implemented, for example, by the service apparatus 110 of FIG. 1, or one or more other data processing apparatuses.
  • Operations of the process 600 can be implemented by instructions stored on a computer readable medium (e.g., non-transitory), and execution of the instructions can cause one or more processors, or a computing device, to perform operations of the process 600.
  • a query is received (602).
  • the query can be a search query submitted by a client device, or another content request.
  • the query can include keywords or other terms with which content, such as creatives, can be identified for distribution to the client device in response to the query.
  • a set of highest ranked combination embeddings for the query are identified (604).
  • the set of highest ranked combination embeddings are identified based on one or more metrics that are stored in association with (e.g., indexed to) stored combination embeddings.
  • the set of highest ranked combination embeddings includes fewer than all stored combination embeddings.
  • the set of highest ranked combination embeddings can include a specified number, or percentage, of combination embeddings having the highest value for the one or more metrics. More specifically, the set of highest ranked combination embeddings can include the N combination embeddings having the highest values for user interaction (e.g., CTR).
  • the set of highest ranked combination embeddings can be a specified percentage of the stored combination embeddings with the highest values for user interaction.
  • Another option is to include all combination embeddings having at least a specified value for the one or more metrics in the set of highest ranked combination embeddings.
  • the set of highest ranked combination embeddings are submitted to a distribution evaluation apparatus (606).
  • the set of highest ranked combination embeddings are submitted to the distribution evaluation apparatus rather than submitting all the stored combination embeddings. In this way, the amount of processing needing to be performed by the distribution evaluation apparatus can be reduced relative to having to evaluate all of the stored combination embeddings. This can be helpful in situations, such as those in which a user has submitted a query or another content requests, where the system must respond to the user in less than a second.
  • the distribution evaluation apparatus can perform predictive performance calculations for the set of highest ranked combination embeddings and/or score the combination embeddings in the set of highest ranked combination embeddings based on information related to the query.
  • a query-specific ranking of the set of highest ranked combination embeddings is created (608).
  • the query-specific ranking can indicate a relative likelihood that creatives represented by each of the highest ranked combination embeddings are appropriate for distribution in response to receipt of the query.
  • the highest ranked combination embedding in the query-specific ranking can represent one or more creatives that are most likely to be appropriate for distribution in response to the query.
  • the query-specific ranking can be based on the predicted user interaction rate (e.g., CTR) of creatives represented by the combination embeddings if distributed in response to the query.
  • the predicted user interaction rate can be generated by a machine learning model trained to make the prediction based on specifics regarding the query (e.g., query content, time of query, country of origin from which the query originated, language of the query, and/or topics of interest of the user who submitted the query).
  • specifics regarding the query e.g., query content, time of query, country of origin from which the query originated, language of the query, and/or topics of interest of the user who submitted the query).
  • the query can be characterized/represented by a query embedding.
  • the query embedding can represent a combination of multiple different features of the query, as discussed above.
  • the query embedding can also be submitted to the distribution evaluation apparatus and used in the creation of the query-specific ranking.
  • the query-specific ranking can be based on one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
  • a given creative is identified based on the query-specific combination embedding ranking (610).
  • the given creative is identified because it has a combination embedding that matches a given combination embedding (e.g., highest ranked combination embedding) among the query-specific combination embedding rankings. For example, using the highest ranked combination embedding in the query-specific combination embedding ranking, a determination can be made that the given creative is represented by a combination embedding that matches the highest ranked combination embedding (e.g., by being the same or sufficiently similar). More specifically, the highest ranked combination embedding according to the query-specific combination embedding ranking can be used as a search key to search the stored combination embeddings for a match. When the match is detected, a given creative (or one or more creatives) indexed to the matching combination embedding can be selected for distribution. In turn, the system can distribute the given creative as a response to the query (612).
  • FIG. 7 is a flow chart of an example process 700 for using combination embeddings to distribute creatives.
  • the process 700 can be implemented, for example, by the service apparatus 110 of FIG. 1, or one or more other data processing apparatuses.
  • Operations of the process 700 can be implemented by instructions stored on a computer readable medium (e.g., non-transitory), and execution of the instructions can cause one or more processors, or a computing device, to perform operations of the process 700.
  • a new creative is obtained (702).
  • the new creative can be received from a content producer, or generated by generative Al systems, as discussed above.
  • the new creative can include one or more digital components, or at least two digital components, that are formatted according to a layout, as discussed throughout this specification.
  • a new combination embedding representing the new creative is created (704).
  • the new combination embedding can be created using a machine learning model that has been trained to accept a creative as input and output a combination embedding representing the attributes of the input creative, as previously discussed. For brevity, that training is not discussed again.
  • a set of stored combination embeddings are searched for a matching combination embedding (706).
  • the search for a matching combination embedding includes comparing the stored combination embeddings for an exact match or a combination embedding that is within a specified distance of a stored combination embedding in semantic space.
  • the matching combination embedding can be included in one of multiple combination embedding clusters.
  • combination embeddings can be grouped together (e.g., clustered) when a similarity measure (e.g., cosine distance) of multiple combination embeddings are within a specified similarity threshold (e.g., distance). This can result in the creation of multiple combination embedding clusters from which a matching combination embedding can be identified.
  • a similarity measure e.g., cosine distance
  • one or more metrics can be aggregated on a per-combination-embedding-cluster basis. This enables the metrics of similar creatives to be grouped together and associated with a single set of metrics, rather than having to store separate instances of the metrics for the individual creatives or the separate instances of similar combination embeddings.
  • the search of the stored combination embeddings can be a search for a combination embedding cluster that includes the matching combination embedding.
  • the one or more metrics stored in association with the matching combination embedding are identified/retrieved (708).
  • the one or more metrics can include one or more measures of user interaction (e.g., CTR) with creatives represented by the matching combination embedding.
  • CTR measures of user interaction
  • the matching combination embedding is included in a given combination embedding cluster (among multiple different combination embedding clusters)
  • the aggregated one or more metrics for the combination embeddings in the given combination embedding cluster can be identified/retrieved. In this way, the “cold start” problem can be avoided, as discussed elsewhere in the specification.
  • Distribution of the new creative is modified based on the one or more metrics (710).
  • the distribution of the new creative is modified based on further evaluation of the new creative and/or the one or more metrics by the distribution evaluation apparatus.
  • the new creative can be distributed in response to certain queries based on evaluation by the distribution evaluation apparatus.
  • the distribution of the new creative can be halted or prevented.
  • the rate of distribution of the new creative can also be modified based on the one or more metrics, and/or other information.
  • FIG. 8 is a block diagram of an example computer system 800 that can be used to perform operations described above.
  • the system 800 includes a processor 810, a memory 820, a storage device 830, and an input/output device 840.
  • Each of the components 810, 820, 830, and 840 can be interconnected, for example, using a system bus 850.
  • the processor 810 is capable of processing instructions for execution within the system 800.
  • the processor 810 is a single-threaded processor.
  • the processor 810 is a multi -threaded processor.
  • the processor 810 is capable of processing instructions stored in the memory 820 or on the storage device 830.
  • the memory 820 stores information within the system 800.
  • the memory 820 is a computer-readable medium.
  • the memory 820 is a volatile memory unit.
  • the memory 820 is a non-volatile memory unit.
  • the storage device 830 is capable of providing mass storage for the system 800.
  • the storage device 830 is a computer-readable medium.
  • the storage device 830 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.
  • the input/output device 840 provides input/output operations for the system 800.
  • the input/output device 840 can include one or more of a network interface device, e.g., an Ethernet card, a serial communication device, e.g., and RS-232 port, and/or a wireless interface device, e.g., and 802.11 card.
  • the input/output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 860.
  • Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
  • Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus.
  • the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
  • a computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them.
  • a computer storage medium is not a propagated signal
  • a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal.
  • the computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
  • the term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing.
  • the apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
  • the apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them.
  • the apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
  • Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both.
  • the essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data.
  • a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.
  • mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.
  • a computer need not have such devices.
  • a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
  • PDA personal digital assistant
  • GPS Global Positioning System
  • USB universal serial bus
  • Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
  • semiconductor memory devices e.g., EPROM, EEPROM, and flash memory devices
  • magnetic disks e.g., internal hard disks or removable disks
  • magneto-optical disks e.g., CD-ROM and DVD-ROM disks.
  • the processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
  • a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer.
  • a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor
  • a keyboard and a pointing device e.g., a mouse or a trackball
  • Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
  • a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to
  • Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components.
  • the components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network.
  • Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
  • LAN local area network
  • WAN wide area network
  • inter-network e.g., the Internet
  • peer-to-peer networks e.g., ad hoc peer-to-peer networks.
  • the computing system can include clients and servers.
  • a client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
  • a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device).
  • client device e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device.
  • Data generated at the client device e.g., a result of the user interaction

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Databases & Information Systems (AREA)
  • Library & Information Science (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Editing Of Facsimile Originals (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

This specification relates to generating a combination embedding of a digital component or a set of digital components along with a layout. The digital components may be such items as images, text, audio, or video. The layout may specify in time and/or space when and where each of the digital components is displayed. In addition, the system creates an embedding of the combination of the digital components and the layout and creates a generator for generating new combination embeddings when presented with new digital components or new layouts. These combination embeddings may be used to improve the efficiency of operation of a system for evaluating content and may be used also to reduce computer resources used for content evaluation.

Description

MULTI-ATTRIBUTE COMBINED EMBEDDING MODEL
TECHNICAL FIELD
[0001] This specification relates to data processing, and the generation of combination embeddings corresponding to a set of two or more attributes of digital components.
BACKGROUND
[0002] Machine learning models that are implemented to evaluate content (e.g., images, text, video, etc.), or otherwise operate in a real-time serving environment (e.g., where content is identified and served to a client device within less than a specified number of milliseconds) require significant computing resources to evaluate the content, or otherwise generate a prediction, within a specified time constraint. For example, to generate the prediction machine learning models generally require a significant amount of energy, computer storage, computer memory, and processing power, particularly in a real-time serving environment. Methods which reduce such resource consumption are therefore desirable.
SUMMARY
[0003] In general, one innovative aspect of the subject matter described in this specification can be embodied in methods including the operations of obtaining a set of digital components, wherein each digital component is one of text, an image, or a video; obtaining a set of layouts, wherein each layout specifies one or more formatting attributes that are applicable to the set of digital components; creating a set of layout-digital component combinations, wherein each layout-digital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components; creating a training embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination; training a machine learning model based on the training embeddings and the set of layout-digital component combinations to obtain a combination embedding model; after the training, identifying a set of creatives each comprising (i) a combination of digital components and (ii) a given layout; generating, by the combination embedding model, a set of combination embeddings based on the set of creatives; and distributing one or more creatives among the set of creatives based on the set of combination embeddings.
[0004] Methods can include filtering the set of combination embeddings based on one or more metrics of each creative having a combination embedding in the set of combination embeddings.
[0005] Methods can include receiving a query; identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
[0006] Methods can include identifying a query embedding representing a combination of multiple different features of the query; and submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query-specific ranking based on the one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
[0007] Methods can include obtaining a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding. [0008] Methods can include creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per- combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
[0009] Methods can include generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings. [0010] A system can include memory and one or more processors, wherein the memory stores computer program instructions which, upon execution by the processor, cause the processor to perform operations of one or more methods. The operations can include receiving a set of digital components, wherein each digital component is one of text, an image, or a video; receiving a set of layouts, wherein each layout specifies one or more formatting attributes that are applicable to the set of digital components; generating a set of layout-digital component combinations, wherein each layout-digital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components; creating a combination embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination; training a machine learning model based on the combination embeddings and the set of layout-digital component combinations; after training the machine learning model, identifying a set of creatives each comprising (i) a combination of digital components and (ii) a given layout; generating, by the machine learning model, a set of combination embeddings based on the set of creatives; and distributing one or more creatives among the set of creatives based on the set of combination embeddings [0011] The operations can include filtering the set of combination embeddings based on one or more metrics of each creative having a combination embedding in the set of combination embeddings.
[0012] The operations can include receiving a query; identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
[0013] The operations can include identifying a query embedding representing a combination of multiple different features of the query; and submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query-specific ranking based on the one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
[0014] The operations can include obtaining a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding.
[0015] The operations can include creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per- combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
[0016] The operations can include generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings. [0017] The operations discussed above can also be implemented as instructions stored on a computer-readable medium. Execution of the instructions can cause one or more data processing apparatus to perform the operations.
[0018] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. For example, the techniques discussed throughout this specification can reduce the amount of time required to evaluate/select a portion of content (e.g., creative) that is ultimately served to a user in response to a content request. For example, current systems utilize a separate machine learning model for each different attribute of the content, which each separately generate embeddings of the attribute, and then use the separate embeddings to generate an evaluation (e.g., prediction) based on the separate embeddings. In contrast, the techniques discussed herein train a model to generate a combination embedding that represents multiple different attributes (e.g., a combination of two or more attributes) of the content to be evaluated. In this way, one machine learning model can be used in the place of two or more different models, while still enabling the evaluation (e.g., prediction) to be based on multiple different attributes of the content.
[0019] In a particular example, a model can be trained to create a single combination embedding representing a combination of text included in a given portion of content, an image/video included in the given portion of content, and a layout of the given portion of content. Meanwhile, existing systems would use three separate models to separately create different embeddings for each of the text, image/video, and layout, resulting in the creation of at least three embeddings by three different models. Using the present techniques, only one model is required to be used to obtain the single combination embedding, thereby reducing the amount of processing power, memory, and time required to obtain an embedding representing the multiple different attributes of the given portion of content. This makes systems implementing the techniques discussed herein more efficient relative to existing systems.
[0020] The efficiencies realized by the present techniques are particularly important when evaluating content for distribution to users in response to a request for content because there are often millions of available portions of content that can be served in response to a request for content, such that multiplying the number of embeddings that need to be created by separate models and evaluated multiplies the amount of processing required to be performed. As those of ordinary skill in the art appreciate, the models required to create embeddings require a large amount of memory and processor usage. As such, being able to create a single embedding that represents multiple different attributes of a given portion of content using a single model substantially reduces the memory and processor usage that is required to represent the given portion of content using embeddings.
[0021] As discussed in detail throughout the specification, the ability to create the combination embedding using a single model is achieved by training the model using a training set of content and the separate embeddings for the different attributes of portions of content in the training set of content. For example, separate embeddings can be created for each attribute of the content to be evaluated. The separate embeddings can then be concatenated to create a combination embedding that represents the multiple attributes of the content, and the model can be trained to directly generate combination embeddings based on the combination of attributes of the content. Once trained, the model can accept, as input, a given portion of content and create, as output, the combination embedding representing the combination of attributes of the given portion of content. As such, a single model can create a single combination embedding, which can be used to evaluate the given portion of content rather than having to generate multiple different embeddings using different models.
[0022] Another advantage provided by the techniques discussed herein is a solution to the “cold start” problem of content evaluation in existing systems. More specifically, existing systems assign random identifiers (e.g., numbers) to different digital components (e.g., text portions, images, videos) and/or combination of digital components (e.g., a combination of text and an image). Because these identifiers are randomly assigned, they do not carry any semantic information, such that two given portions of content (e.g., combinations of text and images) that have the same attributes are not linked, or otherwise identified as semantically similar. Thus, data collected/generated with respect to one of the two given portions of content is not available for use when evaluating the second of the two given portions of content. As such, when a new given portion of content (e.g., combination of digital components having a certain layout) is obtained by existing systems, the existing systems must independently collect data for purposes of evaluating the given portion of content, rather than leveraging data collected for similar portions of content. This adds delay in the ability of existing systems to be able to accurately evaluate new portions of content, and also requires duplication of operations that have already been performed with respect to the similar portions of content, resulting in wasted processing/memory resources.
[0023] For example, because similar portions of content are not linked in existing systems, the existing systems must store a large amount of redundant data, whereas systems implemented using the techniques discussed herein need not store redundant data about semantically similar portions of content. Rather, systems implemented using the techniques discussed herein can store a single instance of data and associate (e.g., index) that single instance of stored data with all of the semantically similar portions of content (e g., having similar text/image/layout). More specifically, when a model is trained to generate a combination embedding using the combination of attributes of a portion of content as input, the combination embeddings of semantically similar portions of content will be similar (e.g., within a specified distance in space), such that the portions of content can be grouped/clustered together as being semantically similar.
[0024] Furthermore, each similar portion of content can be assigned, associated with, indexed, or otherwise linked to the data collected for each member of the group/cluster. In this way, when the model creates a combination embedding for a new portion of content (e.g., newly uploaded or discovered), the new portion of content can be immediately assigned to a group of semantically similar portions of content, and the data that has already been collected for the semantically similar portions of content can be used for the new portion of content, thereby solving the “cold start” problem. Also, newly collected data for the new portion of content can also be stored in association with the other content in the cluster/group, such that the data can be shared among the members of the cluster/group, thereby eliminating the need to store separate instances of data for each of the group members. This is a significant memory savings, particularly in the context of online content where there can be millions/billions of different instances of content for which data needs to be stored.
[0025] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims. DESCRIPTION OF DRAWINGS
[0026] FIG. 1 illustrates an example environment in which a combination embedding can be used in the context of serving content to client devices.
[0027] FIG. 2 illustrates an example of two digital components displayed according to different layouts.
[0028] FIG. 3 illustrates a flow chart of an example process for creating combination embeddings and a machine learning model for producing new combination embeddings.
[0029] FIG. 4 illustrates an example diagram showing details of the generation of a combination embedding.
[0030] FIG. 5 illustrates how including a combination embedding with a query embedding can lead to generating a metric.
[0031] FIG. 6 is a flow chart of an example process for using combination embeddings to distribute creatives.
[0032] FIG. 7 is a flow chart of an example process for using combination embeddings to distribute creatives.
[0033] FIG. 8 illustrates an example of a computing device.
[0034] Like reference symbols in the various drawings indicate like elements.
DETAILED DESCRIPTION
[0035] In current systems, some content distribution systems use a randomly assigned identifier (“ID”) tag to identify each digital component (e.g., each unit of text or image). Generally speaking, these ID tags are assigned to digital components when uploaded to a server, and each uploaded digital component is assigned a different ID tag. However, this ID tag does not carry any semantic information about the content of the digital component, and similar digital components are not identifiable using anu information in the ID tag.
[0036] In an improvement of such a system, each digital component could be assigned an embedding which captures or includes information about that digital component, such that similar digital components would be identifiable based on the embeddings. Additionally, combination embeddings can be used to represent the content of creatives, which include multiple digital components that are formatted according to a layout, such that similar creatives would be identifiable on the basis of the combination embedding. Including such information in an embedding assigned to the combination of digital components (e.g., creatives) with a particular layout scheme would enable a more efficient use of computer resources to optimize such displays as well as to enable more effective use of the information by capturing similarities between similarly displayed combinations of digital components and also to optimize the particular layout which may be most effective.
[0037] The present specification describes training machine learning models to assign combination embeddings to creatives that include multiple digital components formatted according to a layout. In operation, the trained machine learning models accept a creative as an input and output the combination embedding representing the content of that creative. As such, when given a new digital component or a new layout, or both, a new combination embedding can be easily and efficiently created by a single model, rather than having to process the creative, or digital components, using multiple models as currently required. These combination embeddings may be combined with an embedding of a query to generate a metric prediction, which may be used to rank the efficacy of a particular combination of digital components with a layout.
[0038] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component. A combination of digital components can be referred to as a creative, which can be formatted according to a layout.
[0039] FIG. 1 is a block diagram of an example environment 100 in which a combination embedding can be used in the context of context of serving content to client devices. For example, using the combination embeddings in the environment would help reduce computer resource demands and enable the evaluation of more digital components or more layouts or result in new insights into how users perceive the display of digital components displayed in different arrangements.
[0040] The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and a service apparatus 110. The example environment 100 may include many different electronic document servers 104, client devices 106, and digital component servers 108.
[0041] A client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, wearable computing devices, and other devices that can send and receive data over the network 102. A client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102, but native applications (other than browsers), which are also referred to as “apps”, that are executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.
[0042] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally or execute a gaming application that is at least partly stored and/or served by a cloud server (e.g., online gaming applications). Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.
[0043] Di gital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e g., by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.
[0044] As illustrated, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content at a client device 106. Examples of electronic documents include webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps” and/or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to client devices 106 by the electronic document servers 104 (“Electronic Doc Servers”).
[0045] For example, the electronic document servers 104 can include servers that host publisher websites. In this example, the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate presentation of the given webpage at the client device 106.
[0046] In another example, the electronic document servers 104 can include application servers (e.g., native app servers) from which client devices 106 can download apps. In this example, the client device 106 can download files required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively, or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server. In response to receiving the request, the cloud server can execute the application and stream a user interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app and communicate any user interactions with the user interface back to the cloud server for processing.
[0047] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and/or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component or creative) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital component or creative) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source. The following discussion of selecting content responsive to a request uses digital components for purposes of example, but the discussion is equally applicable to selecting creatives for distribution. [0048] In some situations, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references the service apparatus 110, or a particular service provided by the service apparatus 110. In these situations, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for digital components 112 (referred to as a “component request” or “submitted user request”), which is transmitted over the network 102 to the service apparatus 110. For example, the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data. The component request 112 can include event data specifying features such as a name (or network location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e.g., the client device 106), and/or information that the service apparatus 110 can use to select one or more digital components, or other content, provided in response to the request. The component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.
[0049] The component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which digital component can be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and/or media types that are eligible for presentation in the locations can be provided to the service apparatus 110. Similarly, event data specifying keywords associated with the electronic document (“document keywords”) or entities (e.g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provided to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document. The event data can also include a search query that was submitted from the client device 106 to obtain a search results page.
[0050] Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device). Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can be formatted as packetized data having a header and payload data. The header can specify a destination of the packet and the payload data can include any of the information discussed above.
[0051] The service apparatus 110 chooses digital components (e.g., third-party content, such as video files, audio files, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and/or using information included in the component request 112.
[0052] In some implementations, a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delays in providing digital components and layouts in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.
[0053] The techniques discussed, and in particular, the use of a model to produce a combination embedding of digital components with layouts, reduce the likelihood that errors are caused by delayed selection of digital components/creatives and by their display on the client device. For example, by using the combination embedding model, a much smaller subset of digital components and layouts (e.g., tens of layouts and tens of digital components or fewer, rather than hundreds, thousands, or even millions of digital components and thousands of layouts) are required to be evaluated before serving one or more digital components and a layout in response to a request.
[0054] Also, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negatively impacting a user's experience with the electronic document. Further, delays in providing the digital component can result in a failed delivery of the digital component, for example, if the electronic document is no longer presented at the client device 106 when the digital component is provided. Use of the techniques discussed herein also reduces the likelihood of these types of errors for reasons similar to those discussed above.
[0055] In some implementations, the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital component in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a corpus of millions of available digital components (DCi - DCX). The millions of available digital components can be indexed, for example, in a digital component database 116. Each digital component index entry can reference the corresponding digital component and/or include distribution parameters (DPi-DPx) that contribute to (e.g., trigger, condition, or limit) the distribution/transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component. In addition, by using the combination embeddings of digital components with a layout, described elsewhere in this specification, it will be possible to identify similar features of digital components and/or layouts which may be particularly appealing to users. By more easily identifying such features, the use of computer resources to identify them will be reduced.
[0056] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and/or consumption details (e.g., page viewport, user scrolling speed, or other information about the consumption of data). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and/or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify an eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components).
[0057] The identification of the eligible digital component can be segmented into multiple tasks 117a-l 17c that are then assigned among computing devices within the set of multiple computing devices 114. For example, different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a- 118c of the analysis back to the service apparatus 110. For example, the results 118a-l 18c provided by each of the computing devices in the set 114 may identify a subset of digital components that are eligible for distribution in response to the component request and/or a subset of the digital component that have certain distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying the subset of digital components having distribution parameters that match at least some features of the event data.
[0058] The service apparatus 110 aggregates the results 118a- 118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components/ creatives) based on the outcome of one or more content and/or layout evaluation processes, as discussed below. In turn, the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.
[0059] In some implementations, the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components and layouts from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e g., a Uniform Resource Locator (URL)) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component and layout combination from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component and layout pair specified in the server request 121 (e.g., within a database storing multiple digital components and layouts) and transmit, to the client device 106, digital component data (DC Data) 122 that presents the given winning digital component in the electronic document at the client device 106.
[0060] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the layout or by a script 154 from the client device. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within, (e.g., beside), the native content 152 of the electronic document 150. In some implementations, the digital component is overlay ed over (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify the presentation layout within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service apparatus 110 can specify a layout within the scene depicted in the video content over which the digital component is to be presented.
[0061] The service apparatus 110 can also include an artificial intelligence system 160 configured to autonomously generate digital components and layouts, either prior to a request 112 (e.g., offline) and/or in response to a request 112 (e.g., online or real-time). The artificial intelligence (“Al”) system 160 can collect online content about a specific entity (e.g., digital component provider or another entity) and summarize the collected online content using one or more language models 170, which can include large language models.
[0062] A large language model (“LLM”) is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.
[0063] The language model 170 can be any appropriate language model or neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto-regressively generates an output sequence made up of text tokens from the vocabulary. For example, the language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model. [0064] In some situations, the language model 170 can be referred to as an autoregressive neural network when the neural network used to implement the language model 170 auto-regressively generates an output sequence of tokens. More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and a context input that provides context for the output sequence.
[0065] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
[0066] More specifically, to generate a particular token at a particular position within an output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The neural network of the language model 170 can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network of the language model 170 can greedily select the highest- scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.
[0067] As a particular example, the language model 170 can be an auto-regressive Transformer-based neural network that includes (i) a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.
[0068] Generally, however, the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.
[0069] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.
[0070] Generally, because the language model can be auto-regressive, the service apparatus 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from score distributions generated by the language model 170, using a Sample- and-Rank decoding strategy, by using different random seeds for the pseudo-random number generator that’s used in sampling for different runs through the language model 170 or using another decoding strategy that leverages the auto-regressive nature of the language model.
[0071] In some implementations, the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model 170 to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.
[0072] For example, the service apparatus 110 (e.g., Al system 160), or a separate training system, pre-trains the language model 170 (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. As a particular example, the language model 170 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e g., text that is publicly available from the Internet or another text corpus.
[0073] In some implementations, the Al system 160 can generate a prompt 172 that is submitted to the language model 170 and causes the language model 170 to generate the output sequences 174, also referred to as passages or simply as “output”. The Al system 160 can generate the prompt in a manner (e.g., having a structure) that identifies a list of online sources of information, such as a list of websites or data repositories, and specifying a set of constraints the language model 160 must use to generate a summary of information found at the online sources specified in the prompt 172. To initiate creation of the output sequences 174, the Al system 160 submits the prompt 172 to the one or more language models 170, which use the prompt 172 to evaluate the information found at the online sources specified in the prompt 172 and generate the output 174 that summarizes the information according to the constraints specified in the prompt 172. [0074] The Al system 160 can use the generated summary as part of another prompt 172 that is sent to the language model 170. For example, the Al system 160 can insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary) that is submitted to the language model 170 as a constraint for generating clauses for use in digital components being generated by the Al system 160. More specifically, assume that the Al system 160 is generating a digital component to provide in response to the request 112, which includes a keyword/query. In this example, the Al system 160 can generate the additional prompt 172 to include the query and a set of constraints including the summary received in the prior output 174. The set of constraints of the additional prompt 172 can also include instructions regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, semantically styled, among other things (e.g., specifying content that should be excluded from the clauses, such as granular details, such as numbers). For example, the additional prompt 172 could take the following form:
[0075] The Al system 160 can use the generated summary as part of another prompt 172 that is sent to the language model 170. For example, the Al system 160 can insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary) that is submitted to the language model 170 as a constraint for generating clauses for use in digital components being generated by the Al system 160. More specifically, assume that the Al system 160 is generating a digital component and a layout to provide in response to the request 112, which includes a keyword/query. In this example, the Al system 160 can generate the additional prompt 172 to include the query and a set of constraints including the summary received in the prior output 174. The set of constraints of the additional prompt 172 can also include instructions (e.g., layout instructions) regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, semantically styled, among other things (e.g., specifying content that should be excluded from the clauses, such as granular details, such as numbers). For example, the additional prompt 172 could take the following form:
Write a good output - a search digital component where the query is "10g network" and entity is " exampl e_network_provider". good_output is based only on the key parts of this summary: "200 Mbps internet with WiFi equipment on the example_network_provider 10G Network — $50/mo for 2 years. example_network_provider 10G Network is getting faster and more reliable every day. example_network_provider gives you all this and then some. exampl e_network_provider Mobile is the fastest mobile service and the best price for 2 lines of Unlimited. Stream the latest season wherever you go. Storm-Ready WiFi that's backed by a wireless connection, unlimited data, and a battery." good output must be in bullet-point format, good output must have exactly 3 bullet-points. Each bullet-point must be less than 90 characters, good output must have no nested bullets. good_output must be catchy and show valueprop. good output must be useful and informative, and must avoid boring details like numbers.
[0076] In this example prompt, the Al system 160 is providing the language model
170 with the following constraints:
- A query constraint specifies the query “10G Network” to which the output clauses should be relevant.
- A entity constraint specifies “example_network_provider” as the entity name to use in the output clauses.
- A summary constraint specifies the content summary to use during clause generation, i.e., “200 Mbps internet with WiFi...”.
- Styling constraints of “must be in bullet-point format ... must have exactly 3 bullet-points. Each bullet-point must be less than 90 characters ... must have no nested bullets” specify the format the output clauses must use. - Semantic/Tone Control constraints of “must be catchy and show value-prop. good_output must be useful and informative and must avoid boring details like numbers” defines the tone and content of the output clauses generated using the prompt.
[0077] Submission of this additional prompt 172 to the language model 170 causes the language model to generate an additional output 174, which includes multiple sets of clauses generated according to the query and constraints, which is communicated electronically to the Al system 160. The Al system 160 receives the clauses of the additional output 174 and generates multiple candidate digital components and layouts that could be provided in response to the request 112. In some implementations, each different candidate digital component and layout includes a different combination of the clauses received from the language model 170 in the additional output 174. For example, assume that the additional output 174 includes 12 different clauses, and that the formatting of the digital components being generated by the Al system 160 each include space for three different clauses, the Al system 160 could make many different candidate digital components using 3 different clauses in each of the candidate digital components (e g., 121/(3 ! (12-3)!)=220). In some situations, the Al system 160 could also create the candidate digital components using a set of different links to online content (e.g., second level domain links to web pages discussing a topic of the candidate digital components, phone numbers, etc.), which can continue to exponentially increase the number of different candidate digital components that the Al system 160 can create using the clauses of the additional output 174 of the language model 170.
[0078] The Al system 160 can perform one or more post-processing operations that evaluate one or more characteristics of the multiple candidate digital components and layouts. In some implementations, the post-processing operations can include generating a grounding score for each clause from the additional output 174. The grounding score is a value specifying a likelihood that the clause is factual.
[0079] The language model 170 (or more generally the Al system 160) can be configured to generate an output using a combination of input text alone, or using a combination of text and images, audio, video, or other audio/visual content. For example, the language model can accept the input of an image of a brown dog with the text “generate an image of the dog with white spots.” In this example, the language model 170 can process the image/text input, create a new image of the brown dog with white spots, and output the new image in response to the input. Given the fact that the language model 170 can create such output in a very short period of time, the language model 170 can be used to generate a large volume of modified images and/or text in a very short period of time. This capability to generate such a large volume of images and variations of images has led to a drastic increase in the number of digital components that are available for evaluation by the service apparatus 110. For example, the number of new digital components to be evaluated can grow by the millions in a very short period of time (e.g., seconds, minutes, or hours).
[0080] The drastic increase in the number of digital components to be evaluated by the service apparatus 110 (e.g., through expansion of human created and/or uploaded digital components and/or the use of language models) makes it less feasible to exclusively use online models for the evaluation of digital components and also less feasible to vary the arrangement of the digital components according to multiple layout schemes. For example, the amount of time required to evaluate all of the new digital components would require many more computing resources than was previously used to evaluate digital components. Furthermore, the new digital components do not have any historical serving data (e.g., performance or feedback data) regarding the appropriateness of the new digital components for serving in response to specific user queries. As such, the effectiveness of the online model for evaluating the new digital components will be lower than when evaluating digital components that have been widely served in response to various queries. By using a combination embedding of a creative (e.g., digital components with a layout) it will be easier to evaluate the creatives because the combination embedding will represent semantic information about the creative such that information obtained with respect to similar creatives can be shared.
[0081] FIG. 2 illustrates how a combination of multiple digital components are arranged together according to a layout. Digital components may include different content types, such as image digital components, video digital components, text digital components, audio components and the like. Each of these digital components may be presented in a different fashion depending on the specified layout (e.g., dimensions and/or arrangement of digital components in a resulting creative or combination of digital components). For example, an image digital component 202 (“Image DC”) may be combined with a text digital component 204 (“Text DC”) in multiple ways depending on the constraints specified by the layout 206. The layout 206 specifies how the image DC 202 and the text DC 204 are arranged when presented at a client device. In the example illustrated in FIG. 2, the text DC 204 is the phrase “Men’s & Women’s Clothes” and the image DC 202 is a standard image of a stylized man and a stylized woman. Combinations of these two DCs may be presented in a variety of ways such as a first combination 208-1 in which the text DC 204 is presented above the image DC 202, a second combination 208-2 in which the text DC 204 is presented below the image DC 202, a third combination 208-3 in which the text DC 204 is presented to the left and to the right of the image DC 202. Other combinations of layouts and DCs (208-N) may also be used and the number is not limited to those displayed in this figure. For example, the layout 206 could specify different outer dimensions (or aspect ratios) for the resulting creative (e.g., combination of digital components), which can result in different fonts/font sizes being used and/or different arrangements of the text DC 202 and/or image DC 204. In a specific example, the image DC 204 may need to be resized, cropped, or reformatted to fit within the outer boundaries of the resulting creative.
[0082] The layout 206 may include instructions of where to place a certain visual DC 202 in relation to a text DC 204 but is not limited thereto. The layout 206 may also provide the timing of display of certain DCs, such as when an audio DC might play over a speaker of the user’s electronic device. The layout 206 may provide the font or details of the text DC 204 to be displayed, such as the font typeface and whether the text should be displayed using a bold or an italic style. In general, the layout 206 describes the formatting details, arrangement, size, and location of each of the digital components. For more complicated combinations such as multiple visual DCs 202, multiple text DCs 204, multiple audio DCs, or multiple video DCs, the layout 206 my include physical arrangements of each DC in relation to every other DC, but can also include temporal aspects amongst the digital components such as when a video DC may play and whether to include an audio DC while the video DC is playing or whether to wait until a user has indicated an interest in listening to the audio DC.
[0083] A visual DC 202 or a text DC 204 may also be located in different locations on the display of a user’s electronic device and not just in relation to each other. In such instances, the layout 206 may indicate an absolute location on the display rather than a location relative only to the other digital components. In an example, a video DC may be presented to a user in one of several locations such as bottom left, bottom right, upper left, upper right, and middle of the display of the user’s electronic device. For brevity, a combination of one or more digital components DCi and a layout 206 is referred to as a creative.
[0084] When discussing a layout 206, there may be multiple ways the layout can be identified and stored. As noted above, the layout may include an arrangement of a first digital component to a second digital component. The layout 206 may also be much more complicated if the number of digital components is larger and each needs to be specified in relation to every other digital component. Each layout 206 may be assigned an identification (ID) number. The layout ID may be assigned at random, though without permitting duplicate IDs. The layout ID may be assigned based on when the layout was added to the set of layouts. The layout 206 may also be assigned an ID which is indicative of the actual arrangement of the digital components DCi. In another example, the layout ID may be an embedding of the layout instructions or an embedding of the arrangement of the various digital components.
[0085] As discussed in more detail below with reference to FIG. 3, the training of a combination embedding model can be performed using multiple different combinations of digital components. For example, the system can create every combination of stored digital components according to each stored layout. This provides a robust training set that will enable the trained model to accurately assign a combination embedding to a large variety of newly obtained/identified creatives.
[0086] FIG. 3 is a flow chart of an example process 300 for training and deploying/using a combination embedding model. Operations (e.g., steps) of the process 300 can be implemented, for example, by the service apparatus 110 of FIG. 1 or another data processing apparatus (e g., computing device). Operations of the process 300 can also be implemented as instructions stored on a non-transitory computer readable medium, and execution of the instructions can cause one or more data processing apparatus to perform operations of the process 300.
[0087] Digital components and layouts are obtained (302). The digital components can include text digital components, image digital components, video digital components, audio digital components, or other digital components (e.g., animations). As discussed above, the layouts specify one or more formatting attributes that are applicable to a combination of digital components. For example, the layouts can each specify fonts, font sizes, text locations, image locations, temporal presentation information, dimensions of the resulting creative and/or other formatting attributes that control/constrain the presentation of the digital components at a client device.
[0088] The digital components can be obtained, for example, from a database storing the digital components. For example, the digital components can be uploaded by users of the service apparatus 110 of FIG. 1 and stored in the digital component database 116.
The digital components can also be received from a third-party corpus of digital components. For example, an entity implementing the process 300 can obtain access to a corpus of digital components and/or layouts that are maintained by a third party.
[0089] A set of layout-digital component combinations are created (304). In some implementations, each layout-digital component combination includes a layout from the set of layouts that was obtained and at least one digital component from the set of digital components that was obtained. In some situations, layout-digital component combinations can include multiple different digital components. Each of the layoutdigital component combinations can be considered a separate creative/portion of content. Some example digital component combinations 208 with different layouts 206 were illustrated in FIG. 2.
[0090] In some implementations, each possible combination of digital components and layouts can be created to provide a robust training set for training the combination embedding model. For example, the system can iteratively create different combinations of digital components and format each of the different combinations according to a different layout until all possible combinations of digital components and layouts has been created. In some implementations, fewer than all possible combinations can be created, for example, depending on timing constraints, model training evaluation, etc. For instance, if the combination embedding model accuracy reaches an acceptable level without having to create every possible layout-digital component combination, the creation of layout-digital component combination can be halted. Similarly, a target number of layout-digital component combination can be created to perform a specified number of training samples, and this target number can be fewer than all possible combinations.
[0091] A training embedding is created for each given layout-digital component combination (306). In some implementations, each training embedding is created based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination. To create each training embedding, a separate embedding for each digital component in the given layout-digital component combination is obtained. For example, a text embedding of a text digital component in the given layout-digital component combination can be obtained from a database or generated using a text encoder. Similarly, an image embedding of an image digital component in the given layout-digital component combination can be obtained from a database or generated using an image encoder.
[0092] With respect to the layout identifier, an embedding of the layout of the given layout-digital component combination can be obtained from a database or generated using a layout encoder in a similar fashion as the text embedding and the image embedding. In some implementations, the layout identifier need not be an embedding of the layout. For example, if each layout is identifiable by an assigned layout number and similar layouts have similar numbers, the assigned layout number can be used as the layout identifier rather than creating an embedding of the layout. Creation of embeddings is discussed in more detail with reference to FIG. 4.
[0093] Once the digital component embeddings and the layout identifier for a given layout-digital component combination are obtained, the digital component embeddings and the layout identifier can be concatenated, thereby creating a training embedding. For example, assume that, for a given layout-digital component combination, the obtained text embedding is ABCD, the obtained image embedding is EFGH, and the obtained layout identifier is 1234. In this example, the resulting training embedding corresponding to the given layout-digital component combination can be ABCDEFGH1234, which is a concatenation of the digital component embeddings and the layout identifier and can be used as training data for training the combination embedding model. Similar training embeddings can be created for each given layout-digital component combination.
[0094] A machine learning model is trained based on the training embeddings and the set of layout-digital component combinations (308). For example, a given layout-digital component combination and its corresponding training embedding can be considered a training pair that is used to train the machine learning model to generate a combination embedding output based on the input of a creative (e.g., combination of digital components formatted according to a layout). Various different machine learning techniques can be used as part of the machine learning model process. For example, a multi-axis vision transformer architecture or another attention model architecture can be used.
[0095] Vision transformers may be included in the model which may receive an image and produce, for example, text which describes the image. The descriptive text may then be used to produce an embedding, for example, with the help of a large language model. The model may then be trained to produce combination embeddings from a new combination of digital components and layouts. The trained model can be referred to as a combination embedding model because it is trained to output a combination embedding based on the input of a creative (e.g., combination of digital components formatted according to a layout). In some implementations, the training can be performed using unsupervised machine learning techniques.
[0096] A set of creatives is identified after the training (310). In some implementations, each of the creatives in the set of creatives includes (i) a combination of digital components and (ii) a given layout. In some implementations, the creatives can be uploaded by a user of the service apparatus 110 of FIG. 1. For example, a user who uses the service apparatus 110 to distribute creatives through the network 102 can upload an already assembled creative that includes text and an image that is formatted according to a layout. Alternatively, the creative can be dynamically generated by the service apparatus 110 based on a set of digital components uploaded by the user. For example, in some implementations, the service apparatus 110 can use information obtained from a content request to generate a new creative (e.g., using the Al system 160). In these situations, the creative may not have existed prior to the creation of the creative by the service apparatus 110. Creatives for a large number of users of the service apparatus 110 can be identified for potential evaluation.
[0097] A set of combination embeddings are generated for the set of creatives (312). The combination embeddings are generated by the combination embedding model, which was trained to generate combination embeddings using the creatives as input. That is, a combination embedding can be generated for each creative in the set of combination embeddings. As previously discussed, the combination embedding model generates a single combination embedding that represents the combination of digital components that are included in the creative as well as the formatting of the digital components/creative according to a layout. This is in contrast to existing systems, which use different models to create different embeddings for each different digital component type that might be included in a creative. The use of the combination embedding model results in only one embedding being generated by one model, such that the system is more efficient relative to systems that use multiple different models to generate multiple different embeddings. [0098] In some implementations, the set of combination embeddings can be filtered based on one or more metrics (314). For example, for each given creative having a combination embedding in the set of combination embeddings the given creative can be evaluated and filtered based on the one or metrics, as discussed below. In some implementations, the number of combination embeddings required will be much greater than the number of embeddings produced by multiple models so that filtering may be required in some instances to reduce the overall number of combination embeddings being evaluated. [0099] The filtering is an optional operation in some instances. For example, the filtering may be implemented in situations where the number of creatives to be evaluated is larger than the number of creatives that can be evaluated within a specified time constraint and/or using a specified amount of limited processing resources (e.g., cores, processors, memory, etc.). In a specific example, assume that there are 10 text digital components, 10 image digital components, and 10 layouts that can be used to generate creatives. In this example, the total number of text/image/layout combinations will be 1,000 (e.g., 10x10x10=1000). Further assume that the number of queries (e.g., search queries, keywords, or topics) for which the different combinations are to be evaluated is 10 million. In this example, the evaluation of all combinations of digital components for all queries would result in 10 billion evaluations being required. In this situation, time constraints, processing resource constraints, or other restraints may require that fewer than all of the combinations of digital components be evaluated. Therefore, filtering may be performed to reduce the number of evaluations that are required to be performed. [00100] Generally speaking, the purpose of the filtering is to select a subset of the creatives generated using the combinations of digital components that are most likely to be the best candidates for distribution prior to performing the full evaluation of the creatives. At this point in the process 300, the combination embeddings have been generated for each of the creatives generated using combinations of the digital components. These combination embeddings can therefore be used to perform the filtering.
[00101] For example, information stored in association with the combination embeddings or a similar combination (e.g., for other instances of similar creatives) are likely to exist in a database. Therefore, each given combination embedding among the combination embeddings can be used to search a database of performance metrics for matching combination embeddings that match (e.g., are identical or within a specified distance in a vector space) the given combination embedding of one of the creatives/combinations of digital components. Upon finding a match, the performance metrics for the matching combination embedding can be used as the predicted performance metrics of the given combination embedding. In this way, a simple lookup operation can be performed to identify a predicted performance metric for the given combination embedding rather than having to use a predictive model, which is much more computationally intensive and time consuming than a lookup operation, to predict the performance of the creative associated with the given combination embedding. The performance metric used for the fdtering is not critical to the techniques discussed herein and can be selected by an administrator of the system based on the goal to be achieved. However, simple performance metrics that can be used include a user interaction rate of the creative (e.g., click through rate, video view through rate, survey response rate, conversion rate, etc.)
[00102] For purposes of an example, assume that the metric being used to perform the fdtering is a click-through-rate. In this example, if a certain creative including a combination of digital components and having a given layout has been used and is known to have a certain CTR, then a new creative with a combination embedding similar to that of the certain creative may be expected to have a similar CTR. The new creatives can then be filtered based on the expected or predicted CTR. For example, the top X number or top Y% of combination embeddings according to the predicted CTR (or other metric) can be selected as combination embeddings to keep, and creatives that do not have combination embeddings matching those selected combination embeddings can be filtered out, or otherwise removed from consideration such that processing resources and time are not spent fully evaluating those creatives. In some situations, a set of 10 billion creatives can be filtered down to a more manageable set of creatives, such as 100 thousand creatives, or another specified number.
[00103] In some implementations, the filtering described above can be performed in an offline state. As used herein, an offline state refers to a state in which operations are not being performed in response to a query (or other content request) that was submitted by a user. In other words, the offline processing is performed independently of (e.g., before) the system being requested to provide a response to a specific user request. As such, this processing can be utilized to reduce the number of creatives that are further evaluated for potential distribution to users, at least responsive to a given query. Offline filtering may also be performed in initially and then followed by online filtering by using a querydependent model to further reduce the number of candidate creatives.
[001041 Another form of filtering can be performed in an online state. In contrast to the offline state, the online state is a state in which the system is being requested to provide a response to a user query (or other content request). In other words, the online processing is performed after a specific user has submitted a specific content request, such that the user is awaiting a response to the specific content request.
[00105] In some implementations, the filtering in the online state can utilize predictive models to further reduce the number of creatives that are further analyzed for potential distribution by a final selection evaluation. The predictive models can utilize various attributes/features of the creatives available for evaluation (e.g., after the offline filtering) at the time the specific content request and/or a portion of event data included in the users request. Because a user is awaiting a response form the system during the online filtering, the evaluation is performed within a specified time constraint (e.g., within 500 milliseconds). In some implementations, the features/attributes of the creatives are submitted to a predictive performance model (e g., a predictive CTR model), which returns a predicted performance measure for each of the creatives. Using these predicted performance measures, the system can select a specified number, or percentage, of creatives having the highest predicted performance as those creatives that will be submitted to the final selection evaluation, while the unselected creatives will be filtered out, or otherwise removed from further consideration.
[00106] One or more creatives are distributed based, at least in part, on the set of generated combination embeddings (316). In some implementations, the distribution of the one or more creatives can be performed, for example, based on the outcome of the final selection evaluation of those creatives remaining after one or more of the filtering operations discussed above. For example, the remaining creatives can be submitted for evaluation based on distribution criteria of the creatives in combination with specific characteristics of the environment in which the distributed creative will be presented. For example, the final selection evaluation can adjust the previously created predicted performance measures for the remaining creatives to arrive at a selection score. In some situations, the predicted performance measures can be adjusted based on distribution value (e.g., bid) or some other information associated with the creatives. The selection evaluation, selection, and distribution of the remaining creatives can be performed, for example, as discussed above with respect to FIG. 1.
[00107] FIG. 4 illustrates an example 400 of how embeddings for various digital components may be incorporated into a combination embedding (e.g., as performed when creating the training embeddings). Some example digital components include a visual DC 202, a text DC 204, video DCs, audio DCs, or others. A layout 206 may also be included as may other features 408. Each of the digital components (or a set of digital components) may be fed singly or as a group into an encoder such as an image encoder 412 or a text encoder 414. In addition, the layout 206 may be fed into a layout encoder 416, or may be assigned an ID based on other information, as discussed elsewhere in this specification. An advantage of using a layout encoder 416 is that similar layouts may be assigned highly similar embeddings allowing an easier selection of particularly good layouts, when the combinations of layouts and digital components are evaluated. In addition, other features may be included through the use of a feature encoder 418 to create an embedding. Examples of additional features may include such items as the day of the week, calendar day, time of day, URL domain, and the like. Each of the encoders 412, 414, 416, and 418 (or additional encoders as the case may be) produces an embedding such as a visual embedding 422, a text embedding 424, a layout embedding 426, or a feature embedding 428. The digital component encoders 412, 414 may create embeddings (e.g., vector representations) representing the digital components (e.g., visual DC 202 or text DC 204) and/or other corresponding features. All these embeddings may be fed into a combination embedding generator 430 to create a combination embedding 440. The combination embedding generator 430 may create the combination embedding 440 by concatenating the individual embeddings 412, 414, 416, and 418. The combination embedding generator 430 may also be more complicated such as a trained machine learning model which uses other methods to generate a combined embedding 440. An example of a visual DC encoder 412 may be a visual transformer. These are a type of self-supervised learning which extracts a deeper meaning or additional information from an image. In another example, all the digital components, the layout, and the other features may be fed into a single encoder to produce the combination embedding.
[00108] The digital component embeddings 422, 424 (and including other digital components, as needed) may be combined in additional steps in the combination embedding generator 430. For example, concatenated feature embeddings (e.g., concatenation of the visual DC embedding 422, the text embedding 424, the layout embedding 426, and the other feature embedding 428) may be fed into various functions such as a squeeze and excitation module, a feature interaction module, or a multi-head self-attenuation module. The combination embedding generator 430 takes as a minimum input a digital component and a layout and outputs a combined embedding. The generator 430 may take additional inputs such as additional digital components and may also take as additional inputs other features such features related to events. The combination embedding generator 430 may also take other forms such a using a factorized neural network model trained on training data.
[00109] FIG. 5 illustrates how a combination embedding 440 may be used along with a query embedding to generate an output metric. A query may be defined by query features 502 (e.g., terms, term expansions, query time, query origin location, query language, browser type, etc ). The query features may be fed into a query embedding generator 530 to generate a query embedding 540. As shown in FIG. 4, a visual DC embedding 422, a text embedding 424, a layout embedding 426, and other feature embedding 428 may be fed into a combination embedding generator 430 to produce a combination embedding 440 of digital components with a layout 206. The query embedding 540 and the combination embedding 440 may be used to determine an output metric 550. In an example, a dot product of the combination embedding 440 and the query embedding 540 may be used to determine an output metric 550. In another example, the combination embedding 440 and the query embedding may both be input into another model to generate the metric 550.
[00110] Instead of using the combination embedding generator 430, a trained combination embedding model can be used to generate a combination embedding directly from a creative that includes a combination of digital components arranged according to a particular layout, as discussed above. The combination embedding output by the trained combination embedding model can be used with the query embedding in a manner similar to that described with reference to FIG. 5.
[00111] When a query is received, the system may generate a query embedding 540 and determine the set of highest ranked combination embeddings 440 for the query associated with the highest metrics 550 generated for the pair of the combination embedding 440 with the query embedding 540. The system may cut off all combination embeddings 440 below a certain threshold value for the metric 550. The system may rank all the combination embeddings 440 and select only a set number of the combination embeddings 440 for consideration. The system may use other methods to fdter or limit the total number of combinations of digital components and layouts. The system may also create a query-specific ranking of the highest ranked set of creatives and identify a given creative (combination of digital components and layout) to be output as discussed above. [00112] FIG. 6 is a flow chart of an example process 600 for using combination embeddings to distribute creatives. The process 600 can be implemented, for example, by the service apparatus 110 of FIG. 1, or one or more other data processing apparatuses. Operations of the process 600 can be implemented by instructions stored on a computer readable medium (e.g., non-transitory), and execution of the instructions can cause one or more processors, or a computing device, to perform operations of the process 600.
[00113] A query is received (602). In some implementations, the query can be a search query submitted by a client device, or another content request. The query can include keywords or other terms with which content, such as creatives, can be identified for distribution to the client device in response to the query.
[00114] A set of highest ranked combination embeddings for the query are identified (604). In some implementations, the set of highest ranked combination embeddings are identified based on one or more metrics that are stored in association with (e.g., indexed to) stored combination embeddings. The set of highest ranked combination embeddings includes fewer than all stored combination embeddings. For example, the set of highest ranked combination embeddings can include a specified number, or percentage, of combination embeddings having the highest value for the one or more metrics. More specifically, the set of highest ranked combination embeddings can include the N combination embeddings having the highest values for user interaction (e.g., CTR). Alternatively, the set of highest ranked combination embeddings can be a specified percentage of the stored combination embeddings with the highest values for user interaction. Another option is to include all combination embeddings having at least a specified value for the one or more metrics in the set of highest ranked combination embeddings.
[00115] The set of highest ranked combination embeddings are submitted to a distribution evaluation apparatus (606). In some implementations, the set of highest ranked combination embeddings are submitted to the distribution evaluation apparatus rather than submitting all the stored combination embeddings. In this way, the amount of processing needing to be performed by the distribution evaluation apparatus can be reduced relative to having to evaluate all of the stored combination embeddings. This can be helpful in situations, such as those in which a user has submitted a query or another content requests, where the system must respond to the user in less than a second. In some situations, the distribution evaluation apparatus can perform predictive performance calculations for the set of highest ranked combination embeddings and/or score the combination embeddings in the set of highest ranked combination embeddings based on information related to the query.
[00116] A query-specific ranking of the set of highest ranked combination embeddings is created (608). The query-specific ranking can indicate a relative likelihood that creatives represented by each of the highest ranked combination embeddings are appropriate for distribution in response to receipt of the query. For example, the highest ranked combination embedding in the query-specific ranking can represent one or more creatives that are most likely to be appropriate for distribution in response to the query. In some implementations, the query-specific ranking can be based on the predicted user interaction rate (e.g., CTR) of creatives represented by the combination embeddings if distributed in response to the query. The predicted user interaction rate can be generated by a machine learning model trained to make the prediction based on specifics regarding the query (e.g., query content, time of query, country of origin from which the query originated, language of the query, and/or topics of interest of the user who submitted the query).
[00117] In some implementations, the query can be characterized/represented by a query embedding. For example, the query embedding can represent a combination of multiple different features of the query, as discussed above. In these situations, the query embedding can also be submitted to the distribution evaluation apparatus and used in the creation of the query-specific ranking. For example, the query-specific ranking can be based on one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
[00118] A given creative is identified based on the query-specific combination embedding ranking (610). In some implementations, the given creative is identified because it has a combination embedding that matches a given combination embedding (e.g., highest ranked combination embedding) among the query-specific combination embedding rankings. For example, using the highest ranked combination embedding in the query-specific combination embedding ranking, a determination can be made that the given creative is represented by a combination embedding that matches the highest ranked combination embedding (e.g., by being the same or sufficiently similar). More specifically, the highest ranked combination embedding according to the query-specific combination embedding ranking can be used as a search key to search the stored combination embeddings for a match. When the match is detected, a given creative (or one or more creatives) indexed to the matching combination embedding can be selected for distribution. In turn, the system can distribute the given creative as a response to the query (612).
[00119] FIG. 7 is a flow chart of an example process 700 for using combination embeddings to distribute creatives. The process 700 can be implemented, for example, by the service apparatus 110 of FIG. 1, or one or more other data processing apparatuses. Operations of the process 700 can be implemented by instructions stored on a computer readable medium (e.g., non-transitory), and execution of the instructions can cause one or more processors, or a computing device, to perform operations of the process 700. [00120] A new creative is obtained (702). In some implementations, the new creative can be received from a content producer, or generated by generative Al systems, as discussed above. In either case, the new creative can include one or more digital components, or at least two digital components, that are formatted according to a layout, as discussed throughout this specification.
[00121] A new combination embedding representing the new creative is created (704). In some implementations, the new combination embedding can be created using a machine learning model that has been trained to accept a creative as input and output a combination embedding representing the attributes of the input creative, as previously discussed. For brevity, that training is not discussed again.
[00122] A set of stored combination embeddings are searched for a matching combination embedding (706). In some implementations, the search for a matching combination embedding includes comparing the stored combination embeddings for an exact match or a combination embedding that is within a specified distance of a stored combination embedding in semantic space.
[00123] In some situations, the matching combination embedding can be included in one of multiple combination embedding clusters. For example, combination embeddings can be grouped together (e.g., clustered) when a similarity measure (e.g., cosine distance) of multiple combination embeddings are within a specified similarity threshold (e.g., distance). This can result in the creation of multiple combination embedding clusters from which a matching combination embedding can be identified.
[00124] In situations where multiple combination embedding clusters are created, one or more metrics can be aggregated on a per-combination-embedding-cluster basis. This enables the metrics of similar creatives to be grouped together and associated with a single set of metrics, rather than having to store separate instances of the metrics for the individual creatives or the separate instances of similar combination embeddings. As such, the search of the stored combination embeddings can be a search for a combination embedding cluster that includes the matching combination embedding.
[00125] When a matching combination embedding is found, one or more metrics stored in association with the matching combination embedding are identified/retrieved (708). In some implementations, the one or more metrics can include one or more measures of user interaction (e.g., CTR) with creatives represented by the matching combination embedding. When the matching combination embedding is included in a given combination embedding cluster (among multiple different combination embedding clusters), the aggregated one or more metrics for the combination embeddings in the given combination embedding cluster can be identified/retrieved. In this way, the “cold start” problem can be avoided, as discussed elsewhere in the specification.
[00126] Distribution of the new creative is modified based on the one or more metrics (710). In some implementations, the distribution of the new creative is modified based on further evaluation of the new creative and/or the one or more metrics by the distribution evaluation apparatus. For example, the new creative can be distributed in response to certain queries based on evaluation by the distribution evaluation apparatus. Similarly, if the new creative is deemed inappropriate for distribution, for example, based on having a low expected user interaction rate, the distribution of the new creative can be halted or prevented. The rate of distribution of the new creative can also be modified based on the one or more metrics, and/or other information.
[00127] FIG. 8 is a block diagram of an example computer system 800 that can be used to perform operations described above. The system 800 includes a processor 810, a memory 820, a storage device 830, and an input/output device 840. Each of the components 810, 820, 830, and 840 can be interconnected, for example, using a system bus 850. The processor 810 is capable of processing instructions for execution within the system 800. In one implementation, the processor 810 is a single-threaded processor. In another implementation, the processor 810 is a multi -threaded processor. The processor 810 is capable of processing instructions stored in the memory 820 or on the storage device 830.
[00128] The memory 820 stores information within the system 800. In one implementation, the memory 820 is a computer-readable medium. In one implementation, the memory 820 is a volatile memory unit. In another implementation, the memory 820 is a non-volatile memory unit. [00129] The storage device 830 is capable of providing mass storage for the system 800. In one implementation, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.
[00130] The input/output device 840 provides input/output operations for the system 800. In one implementation, the input/output device 840 can include one or more of a network interface device, e.g., an Ethernet card, a serial communication device, e.g., and RS-232 port, and/or a wireless interface device, e.g., and 802.11 card. In another implementation, the input/output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 860. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[00131] Although an example processing system has been described in FIG. 8, implementations of the subject matter and the functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. [00132] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[00133] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[00134] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[00135] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). [00136] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[00137] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[00138] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[00139] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[00140] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[00141] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[00142] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

WHAT IS CLAIMED IS
1. A method comprising: obtaining a set of digital components, wherein each digital component is one of text, an image, or a video; obtaining a set of layouts, wherein each layout specifies one or more formatting attributes that are applicable to the set of digital components; creating a set of layout-digital component combinations, wherein each layoutdigital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components; creating a training embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination; training a machine learning model based on the training embeddings and the set of layout-digital component combinations to obtain a combination embedding model; after the training, identifying a set of creatives each comprising (i) a combination of digital components and (ii) a given layout; generating, by the combination embedding model, a set of combination embeddings based on the set of creatives; and distributing one or more creatives among the set of creatives based on the set of combination embeddings.
2. The method of claim 1, further comprising filtering the set of combination embeddings based on one or more metrics of each creative having a combination embedding in the set of combination embeddings.
3. The method of claim 2, further comprising: receiving a query; identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
4. The method of claim 3, further comprising: identifying a query embedding representing a combination of multiple different features of the query; and submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query-specific ranking based on the one or more metrics stored in association with (i) the query embedding and (ii) the set of highest ranked combination embeddings.
5. The method of claim 1, further comprising: obtaining a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding.
6. The method of claim 5, further comprising: creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per-combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
7. The method of claim 1, further comprising: generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings.
8. A system comprising: one or more memory devices; and one or more processors configured to interact with the one or more memory devices, wherein the one or more memory devices store computer program instructions that, upon execution by the one or more processors cause the one or more porcessors to perform operations comprising: receiving a set of digital components, wherein each digital component is one of text, an image, or a video; receiving a set of layouts, wherein each layout specifies one or more formatting attributes that are applicable to the set of digital components; generating a set of layout-digital component combinations, wherein each layout-digital component combination comprises a layout from the set of layouts and at least one digital component from the set of digital components; creating a combination embedding for each given layout-digital component combination based on a concatenation of (i) separate embeddings of each digital component of the given layout-digital component combination and (ii) a layout identifier of the layout of the given layout-digital component combination; training a machine learning model based on the combination embeddings and the set of layout-digital component combinations; after training the machine learning model, identifying a set of creatives each comprising (i) a combination of digital components and (ii) a given layout; generating, by the machine learning model, a set of combination embeddings based on the set of creatives; and distributing one or more creatives among the set of creatives based on the set of combination embeddings.
9. The system of claim 8, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising filtering the set of combination embeddings based on one or more metrics of each creative having a combination embedding in the set of combination embeddings.
10. The system of claim 9, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising: receiving a query identifying a set of highest ranked combination embeddings for the query according to the one or more metrics, wherein the set of highest ranked combination embeddings includes fewer than all stored combination embeddings; submitting, to a distribution evaluation apparatus, the set of highest ranked combination embeddings rather than all stored combination embeddings; creating, by the distribution evaluation apparatus, a query-specific ranking of the set of highest ranked combination embeddings; and identifying a given creative having an embedding matching a given combination embedding among the set of highest ranked combination embeddings, wherein distributing one or more creatives comprises distributing the given creative as a response to the query.
11. The system of claim 10, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising: identifying a query embedding representing a combination of multiple different features of the query; and a submitting, to the distribution evaluation apparatus, the query embedding, wherein creating a query-specific ranking comprises creating the query specific ranking based on the one or more metrics stored in association with the query embedding and the set of highest ranked combination embeddings.
12. The system of claim 8, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising: receiving a new creative comprising one or more digital components and a layout; creating, by the machine learning model, a new combination embedding representing the new creative; searching a set of stored combination embeddings for a matching combination embedding that matches the new combination embedding; identifying one or more metrics stored in association with the matching combination embedding; and modifying a distribution of the new creative based on the one or more metrics stored in association with the matching combination embedding.
13. The system of claim 12, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising: creating multiple combination embedding clusters that each include a group of combination embeddings having a similarity measure that is within a specified similarity threshold; and aggregating the one or more metrics on a per-combination-embedding-cluster basis, wherein: searching a set of stored combination embeddings comprises searching the multiple combination embedding clusters for the matching combination embedding; and identifying one or more metrics stored in association with the matching combination embedding comprises identifying the aggregated one or more metrics for a given combination embedding cluster that includes the matching combination embedding.
14. The system of claim 8, wherein execution of the computer program instructions cause the one or more processors to perform operations further comprising: generating a separate embedding for each digital component, wherein generating the separate embedding comprises using different embedding models based on a type of the digital component for which the embedding is being generated, the different embedding models comprising at least two of a text embedding model, an image embedding model, or a video embedding model, wherein generating, by the machine learning model, a set of combination embeddings comprises using a single machine learning model to generate the set of combination embeddings.
15. A non-transitory computer readable medium storing instructions that, upon execution by one or more data processing apparatus, cause the one or more data processing apparatus to perform operations of any of claims 1-7.
EP24714344.9A 2024-02-22 2024-02-22 Multi-attribute combined embedding model Withdrawn EP4634827A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2024/016820 WO2025178622A1 (en) 2024-02-22 2024-02-22 Multi-attribute combined embedding model

Publications (1)

Publication Number Publication Date
EP4634827A1 true EP4634827A1 (en) 2025-10-22

Family

ID=90473512

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24714344.9A Withdrawn EP4634827A1 (en) 2024-02-22 2024-02-22 Multi-attribute combined embedding model

Country Status (4)

Country Link
EP (1) EP4634827A1 (en)
JP (1) JP2026509620A (en)
CN (1) CN121195266A (en)
WO (1) WO2025178622A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11636291B1 (en) * 2020-04-06 2023-04-25 Amazon Technologies, Inc. Content similarity determination
CN115223182A (en) * 2022-07-14 2022-10-21 河南中原消费金融股份有限公司 Document layout identification method and related device

Also Published As

Publication number Publication date
CN121195266A (en) 2025-12-23
WO2025178622A1 (en) 2025-08-28
JP2026509620A (en) 2026-03-23

Similar Documents

Publication Publication Date Title
AU2014201827B2 (en) Scoring concept terms using a deep network
CN113609374A (en) Data processing method, device and equipment based on content push and storage medium
CN111279332A (en) Generating request-agnostic interaction scores for electronic communications using machine-learning models and utilizing request-agnostic interaction scores
US20250124264A1 (en) Generating customized content descriptions using artificial intelligence
US20250315463A1 (en) Deep linking using generative artificial intelligence
EP4565996A1 (en) Specificity aware teacher model and student model based on large language model
EP4587939A1 (en) Generative artificial intelligence
WO2025178622A1 (en) Multi-attribute combined embedding model
WO2025029248A1 (en) Language model for predicting digital component selection data
EP4623371A1 (en) Specificity aware teacher model and student model based on large language model
US20260050772A1 (en) Generative ai techniques guided by network signals
US20250322214A1 (en) Self-criticizing artificial intelligence system
EP4699061A1 (en) Efficient estimation of online machine learning model results using an offline machine learning model
JP7223164B2 (en) Data integrity optimization
US20240005040A1 (en) Cardinality models for privacy-sensitive assessment of digital component transmission reach
EP4710257A1 (en) Augmenting machine learning models
WO2025030115A1 (en) Image generation using prompt chains
WO2025116909A1 (en) Efficient utilization of generative artificial intelligence
WO2024249391A1 (en) Retrieval token generation from queries using language model
WO2025101294A1 (en) Artificial intelligence for efficient image editing
WO2025085179A1 (en) Efficient response generation using refinement queries and artificial intelligence
EP4377813A1 (en) Privacy sensitive estimation of digital resource access frequency
WO2026049733A1 (en) Artificial intelligence model training for image generation
EP4581516A1 (en) Using intermediate embeddings of language model neural networks to select digital components
EP4662593A1 (en) Self-improving artificial intelligence system

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250521

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: EXAMINATION IS IN PROGRESS

17Q First examination report despatched

Effective date: 20251112

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN WITHDRAWN

18W Application withdrawn

Effective date: 20260309