EP4713828A1 - Generative digital component creation - Google Patents
Generative digital component creationInfo
- Publication number
- EP4713828A1 EP4713828A1 EP24827491.2A EP24827491A EP4713828A1 EP 4713828 A1 EP4713828 A1 EP 4713828A1 EP 24827491 A EP24827491 A EP 24827491A EP 4713828 A1 EP4713828 A1 EP 4713828A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- digital
- language model
- digital components
- generated
- clauses
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
- G06F40/35—Discourse or dialogue representation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/34—Browsing; Visualisation therefor
- G06F16/345—Summarisation for human users
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/40—Processing or translation of natural language
- G06F40/42—Data-driven translation
- G06F40/44—Statistical methods, e.g. probability models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/40—Processing or translation of natural language
- G06F40/55—Rule-based translation
- G06F40/56—Natural language generation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/335—Filtering based on additional data, e.g. user or group profiles
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Probability & Statistics with Applications (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating digital components. In one aspect, a method includes generating, by an artificial intelligence system, prompts that each includes a query for generating digital components and a text summary of a specified source of online content for use in generating the plurality of digital components by a first language model. The AI system obtains digital components as generated by the first language model based on the prompts. Training data including the digital components and the text summary of the specified source of online content to train a second language model. The second language model is trained to generate clauses to be used to generate new digital components that meet a grounding threshold for the second language model to classify the clauses as grounded. The AI system provides at least one generated digital component.
Description
GENERATIVE DIGITAL COMPONENT CREATION
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63/616,159, filed on December 29, 2023. The disclosure of the prior application is considered part of and are incorporated by reference in the disclosure of this application.
BACKGROUND
[0002] This specification relates to data processing, artificial intelligence, and generating digital content using artificial intelligence.
[0003] Advances in machine learning are enabling artificial intelligence to be implemented in more applications. For example, large language models have been implemented to allow for a conversational interaction with computers using natural language rather than a restricted set of prompts. This allows for a more natural interaction with the computer.
SUMMARY
[0004] In general, a first innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions of generating, by an artificial intelligence system, a plurality of prompts that each includes a query for generating a plurality of digital components and a text summary of a specified source of online content for use in generating the plurality of digital components by a first language model; obtaining, by the artificial intelligence system, the plurality of digital components as generated by the first language model based on the plurality of prompts; providing training data including (i) the plurality of digital components and (ii) the text summary of the specified source of online content to train a second language model; training the second language model to generate clauses to be used to generate new digital components that meet a grounding threshold for the second language model to classify the clauses as grounded; and providing, by the artificial intelligence system, at least one digital component generated by the second language model based on the generated clauses. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.
[0005] These and other embodiments can each optionally include one or more of the following features. Some aspects include evaluating each clause of the generated clauses to determine grounding measures for each clause that specifies a likelihood that content of the clause is factual as being present in the text summary of the specified source of online content. Each of the clauses is classified as factual or not based on comparing the determined grounding measure and the grounding threshold.
[0006] In some aspects, providing the at least one digital component generated by the second language model includes selecting the at least one digital component based on a selection criterion including the grounding threshold defining a likelihood threshold for determining that the at least one digital component is factual and verifiable at a text corpus provided as context for generating the at least one digital components.
[0007] In some aspects, providing the training data includes reducing the text summary of the specified source of online content to include a first portion of the text summary that is limited to a specified character count to generate the clauses. Reducing the text summary can include determining a second portion of the text summary of the specified source to include relevant information at least based on positioning of the second portion of the text within a display of the online content at a web browser previewing the specified source and defining the first portion of the text summary included in the reduced text to include at least some of the second portion of the text within a threshold content size.
[0008] Some aspects include deploying the second language model to serve requests to generate clauses based on provided details for a source of online content.
[0009] Some aspects include generating, by the second language model, the at least one digital component based on generating clauses for a provided online content via a request received at the artificial intelligence system.
[0010] Some aspects include determining the plurality of the digital components by validating the digital components generated by the first language model, wherein each of the digital components of the plurality of digital components is classified as faithful according to a digital component faithfulness classifier model executed to classify the plurality of digital components.
[0011] In some aspects, providing the training data includes determining the text summary based on identifying portions of the online content that are relevant for the query according to
a relevance criterion. Identifying the portions of the online content can include monitoring user behavior and/or click interaction with the online content when presented on a display screen; identifying a first set of portions presented at locations associated with areas of the display screen that are associated with viewing and interaction that are above a threshold level; and providing the first set of portions for determining the text summary.
[0012] In some aspects, providing the at least one digital component includes providing a request for the at least one digital component by invoking the second language model to provide clauses, wherein the request provides a source of online content; and generating a new clause, by the second language model, that is grounded in the text of the online content of the source.
[0013] In some aspects, providing the at least one digital component includes measuring a plurality of digital components generated by the second language model; and based on the measuring, selecting the at least one digital component that meets a threshold criterion for a measure determined for at least one digital component. The measuring can be performed based on a trained deep network model that determines a likelihood that a digital component generated by the second language model is to be interacted with by a user.
[0014] In some aspects, a clause generated by the second language model points to a token defined for an input text of a requested source.
[0015] In some aspects, the text summary includes text content generated based on processing text content and/or image content extracted from the specified source of online content.
[0016] In some instances, a second innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions of obtaining, by an artificial intelligence system, a set of digital components to be used to train a first language model; providing training data including (i) the set of digital components as generated by the first language model and (ii) a text summary of a specified source of online content to train the first language model to generate clauses to be used to generate new digital components that meet a grounding threshold to classify the clauses as factual and grounded; and providing, by the artificial intelligence system, at least one digital component generated by the first language model based on the generated clauses. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.
[0017] Similar operations and processes may be performed in a system comprising at least one processor and a memory communicatively coupled to the at least one processor where the memory stores instructions that when executed cause the at least one processor to perform the operations. Further, a computer-readable medium, which can be non-transitory, storing instructions which, when executed, cause at least one processor to perform the operations may also be contemplated. In other words, while generally described as computer implemented software embodied on tangible, non-transitory media that processes and transforms the respective data, some or all of the aspects may be computer implemented methods or further included in respective systems or other devices for performing this described functionality.
[0018] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The techniques discussed in this specification enable artificial intelligence (Al) to be used to generate customized content, e.g., customized digital components, based on data related to the digital components and/or one or more queries received from a client device. Digital components can be created in real-time (e.g., in response to a request or a query and without delay) and/or offline based on generative model techniques that support improved relevance of newly generated digital components with pre-existing digital components while remaining faithful to content of an online source and a context for using the digital component. For example, the context can be based on a query or user, through the environment where the query is received or from where the query is received (e.g., a context in which the digital component will be presented), and/or through pre-defined configurations and set-ups based on rules.
[0019] Using the described techniques in the present disclosure support efficient utilization of computer resources based on the ability to obtain input, create intent information using a trained language model, and use the intent information to generate digital components that can be evaluated and/or filtered to select an output digital component as a response. Such processes can reduce the network calls between the system and a device of a user if the processes involve interaction with a user to arrive at the generation of the customized content. Further, the Al system can generate customized content that accurately shows an item that is the subject of a digital component in a context that is relevant to the user and/or the user’s informational needs. [0020] In some implementations, generative responses can be provided at query time using the intent information and a subset of pre-created content, which enables the distribution of many
more different assets than those stored in memory without having to store all of the different variations, which would otherwise occupy data storage resources that can be used for other purposes. In some examples, created content can be pre-stored without pre-generating and pre-storing digital components based on combinations of the content. Such techniques can facilitate fast generation of digital components at query time without having to pre-store them in memory and thus optimize the memory usage. The described techniques enable the system to provide content in real-time that can be used to create a distribution plan and enable content publishers to obtain usable results faster, with fewer user interactions, and more computationally efficiently.
[0021] The Al models can be trained using grounding information such that new or modified clauses output by the models for use in new digital components (e.g., presented by the digital components) satisfy a grounding threshold. This ensures that the text that is displayed on a digital component is error free and does not include hallucinations or other incorrect information. As the input to current generative models often include substantial amounts of information and/or other content often results in hallucinations and errors, training the models as described herein improves the accuracy and quality of digital components generated using the models. This allows for digital components to be created offline or in real-time (e.g., in response to a query or component request) without spending time and resources to check the quality of the digital components before distributing the digital components for presentation to users.
[0022] In some instances, the use of grounded clauses for newly generated digital components can reduce the time needed for the digital components’ generation while yet providing results that meet a grounding criterion. By using a trained model, multiple grounded clauses based on one or multiple resources provided as input for the clause generation can be generated, where subsequent digital component generation that can be based on various combination of the generated grounded clauses. The digital component generation is performed faster since the generation of the clauses can be performed as a pre-phase of the component generation, which allows for a larger number of components be generated based on various combinations of the generated clauses. The digital component generation efficiently uses the computing resources so that clause generation can support a faster component generation that also meets the grounding criterion. In some cases, digital components can be created offline or in real-
time (e.g., in response to a query or component request) without spending time and resources to provide a basis for generating digital components that use one or multiple of the generated clauses (e.g., in various combinations). As such, the generated digital components can be associated with a quality of the digital components that meets grounding criteria, where the performance of the quality check can be performed faster and with fewer resources, as the clauses used for the digital components are grounded as generated. The generated digital components based on such pre-generated clauses (whether online or offline) can be distributed for presentation to users.
[0023] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0024] FIG. 1 is a block diagram of an example environment in which generative artificial intelligence can be implemented to generate digital components.
[0025] FIG. 2A is a block diagram illustrating interactions between an artificial intelligence system, a language model, and a client device.
[0026] FIG. 2B is a flow chart of an example process for generating digital components in real-time using artificial intelligence.
[0027] FIG. 3 is a flow chart of an example process for training a language model and generating digital components based on the trained language model.
[0028] FIG. 4A is a block diagram of an example architecture of components interacting to generate digital components based on a trained language model.
[0029] FIG. 4B is a block diagram of an example environment for customization of generated digital components in real-time.
[0030] FIG. 5 is a block diagram of an example process of digital components creation based on training data generated by a language model.
[0031] FIG. 6 a block diagram of an example computer.
[0032] Like reference numbers and designations in the various drawings indicate like elements.
DETAILED DESCRIPTION
[0033] This document describes techniques for enabling artificial intelligence to generate new digital components in faster and more efficient ways. Artificial intelligence (Al) is a segment of computer science that focuses on the creation of models that can perform tasks and actions autonomously, e g., with little to no human intervention. Al systems can utilize, for example, one or more of machine learning, natural language processing, or computer vision. Machine learning, and its subsets, such as deep learning, focus on developing models that can infer outputs from data. The outputs can include, for example, predictions and/or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and/or other content, in response to input prompts and/or based on other information.
[0034] The techniques described throughout this specification enable Al to generate large numbers of new digital components using various combinations of text, images, and/or or types of content or resources. For example, an Al system can gather information from various sources, such as online sources (e.g., web pages or other trusted online sources), and use this information in different ways to create different digital components, e.g., to respond to user requests (e.g., generated at real-time (or on-the-fly) or as a prepared set to be used upon request). Generally, speaking, the system can use an input prompt to a language model, such as a large language model (LLM), that outputs multiple clauses that can be presented by the digital components.
[0035] In some implementations, when digital components are created (e.g., automatically based on generative language models), the digital components can be evaluated to determine whether they are eligible to be provided, e.g., immediately, in response received requests. The system can use the clauses to create multiple different digital components, and can also perform post-processing techniques to select, from among the different digital components, a result set of one or more digital components. The post-processing can be based on clause evaluation and/or fdtering those of the clauses that match criteria for selection. The criteria for selection of a component can include a grounding threshold to evaluate classifications and determine that the clauses are factual and grounded. At least some digital components can be determined as grounded as having a high likelihood that the digital components are factual. Ungrounded can be considered to include information that cannot be verified in a particular corpus, such as the text of one or more resources related to the digital component, e.g., one or more web pages related to a subject of the digital component. The pages can include a landing page linked to by the digital component, which can also be referred to as a digital component page.
[0036] In some implementations, a classifier can be implemented to determine the likelihood that a generated digital component is valid based on the source text (e.g., text of a digital component page and/or other digital components that can be provided as input) used for generating the digital component. In some instances, the classifier can be trained based on human collected data (e.g., trained on collated training data with human labels) and can be used to determine whether at least a portion of the text in the digital component is a hallucination by identifying whether that portion of the text can be attributed to a specific portion of the source text, e.g., a particular portion of the digital component page linked to by
the digital component. When a hallucination is determined, it can mean that the text in the digital component is not faithful with respect to the digital component page provided for the digital component generation. Hallucinations can be determined when a generated digital component includes inaccurate information that cannot be validated on the whole corpus of the source text associated with the digital component page. When at least a portion of the text of the digital component is determined to be a hallucination, then the digital component may not be considered as grounded.
[0037] A selection of clauses that are grounded can be made and those clauses can be used in the digital component generation. Ungrounded clauses can be filtered from the generation process. The selection of clauses can also include hallucination control to improve the quality of generated digital components.
[0038] In some implementations, the system can rely on grounding techniques to produce results that are accurate and relevant for a particular request. A large language model can be provided with information that is specific and not used for training the model as input when requesting an output to obtain accurate and relevant output to a particular input used to ground the model in that specific use-case.
[0039] In some implementations, once digital components are generated in accordance with the present implementations, post-processing operations can then be used to evaluate the generated digital components against each other to determine which candidate digital components have higher quality than other candidate digital components (e.g., given the current context), and one or more of the higher quality digital components can be output to a computing device (e.g., a client device of a user).
[0040] In some implementations, digital components can be model -generated. Models used for generating digital components can be trained on relevant input such as queries, web content (e.g., page content of a page linked to by a digital component, portions of other web pages, summarized content from one or more web pages, etc.), and/or other digital components. For example, the models can be trained on input such as pages including page content, such as pages linked to by digital components, which can also be referred to as landing pages or digital component pages. In some implementations, the models used for generating the digital components are not trained on other digital components, and new digital components are generated based on input to the model that can include an online source. For example, the
input can be provided as a Uniform Resource Locator (URL) to a digital component page, for example, associated with an organization) and other digital components (e.g., associated with existing digital components and/or a user). In some other examples, the input can be provided as content of a digital component page linked to by a digital component.
[0041] In some implementations, the new digital components created based on a trained model (or combination of models) can be digital component variations or rewrites (e.g., edited versions) of existing digital components. For example, the new digital components can be digital components that are relevant to a received request providing context for the component generation (e.g., online content, query, other digital components previously generated). The new digital components can remain faithful to a source of online content provided with the request, e.g., by using the evaluations of clauses described herein.
[0042] In some implementations, the generation of the digital components can be performed based on a received query (e.g., can be referred to as query or user aware digital components creations). In some implementations, the generation of the digital components can be performed as adjustments to perceptible features of existing digital components (text, images, graphics, combinations thereof at real-time) while processing a request for digital components. In some instances, digital components can be created at real-time and on-demand (at query time) to improve relevance of newly generated digital components with pre-existing digital components while remaining faithful to content of an online source such as a digital component page of a party relevant to the digital components (e.g., an organization that publishes the digital component) and a context for using the digital components (e.g., context provided through the query or user, through the environment where the query is received, or through pre-defined configurations and set-ups based on rules). In some implementations an online source can be an electronic document as described in further detail in relation to FIG. 1.
[0043] In some instances, a generative model used for creation of digital components at realtime can be trained to stay factual and not lose context while creating a new digital component. In some instances, the used generative model can be associated with an error rate that can be determined during evaluation of the generative model. The evaluation of the error rate can be performed with reference to a type of digital components that are to be created (e.g., type related to technology of the online source or to the subject of the digital component).
[0044] In some implementations, an on-demand digital component (e.g., generated at realtime) can be generated by deriving content from a digital component page that is relevant for the component creation and a query, where the query can be used to prioritize which content or portion of the content is most relevant. In some implementations, the component generation can be performed based on a language model that receives a query and content associated with a digital component page and generates clauses that are to be used in new digital components. The clauses can be created as rewrites where the model can provide techniques for substitution (rephrasing or paraphrasing) of words or phrases to provide components that are faithful to the digital component page.
[0045] In some implementations, generated clauses can be evaluated to determine whether they comply with policy and quality criteria. In some implementations, the policy criteria can include different rules with which the text in the digital component has to comply. These rules can include trademark violation rules or brand violation rules. For example, the rules can specify that if a generated digital component includes a trademark or a brand that is not affiliated with the digital component, the digital component violates a policy and cannot be distributed.
[0046] In some instances, the rules can be specified for sensitive text within the digital component that can relate to different categories such as health related, age restrictions, content specifics, law and governmental specific. In some instances, a machine learning model can be trained to determine whether a digital component meets a policy criterion by looking into different rules defined for compliance of text in the digital component. In some instances, multiple models can be used including a trademark validator, or a sensitive text classifier, among other examples of models trained to evaluate digital components based on defined compliance rules to determine whether the digital components meet a threshold level of policy compliance for providing the digital components as output. In some instances, the quality criteria for evaluating generated clauses can be associated with compliance with grammatical rules, which may be specific rules defined for clauses. For example, the clause may not be a full sentence and may miss a punctuation mark, e.g., a full stop at the end, which can comply with a grammatical rule that is a linguistic rule can be different from the natural language grammar rules. In some instances, a custom set of rules can be defined for the quality criteria to evaluate generated clauses and determine whether those clauses can be considered as
complying with a quality criterion related to grammar that can be specifically defined and tailored to the digital components to be generated. In some instances, the quality criteria can include rules for truncation compliance considering which truncations are contextually correct. In some instances, the quality criteria can be customized to adapt to expected standards or syntax of presenting text within a digital component that would not be misleading or miss the context of the source text associated with the digital component page as provided.
[0047] In some implementations, a faithfulness classifier can be implemented to detect hallucination and grounding errors. The faithfulness classifier can be trained to automatically measure faithfulness of digital components, which can be used to filter the unfaithful generated components or can be used to improve the generation process of the digital components by feeding the measure as input to the generation process. In some instances, the faithfulness classifier be generated to avoid bias to the labels which can lead to bias toward token matching algorithms. The training of the model can rely on an expert rated label set where the labels of the training data are trusted and can be used for the evaluation. In some implementations, if such expert rated data is unavailable, the faithfulness classifier can be generated as either a noisier classifier or as a biased one.
[0048] In some instances, generated digital components can include truncation errors that may render the text as meaningless or incomplete. Such truncation issues may arise due to truncating the source text. Table 1 shows examples of truncation errors that can be identified by a quality classifier as grammar rules can include to avoid a pattern of ending a phrase with a preposition.
Table 1
[0049] Table 2 shows examples of truncation errors that may not be identified as they may comply with the grammar quality checks but still be incomplete in their meaning.
Table 2
[0050] In some implementations, a model is trained on positive examples that can be synthetically generated based on patterns of truncation, such as:
• Pattern 1 : Truncate such that main word is dropped.
• Pattern 2: Truncate such that the text ends with personal pronoun.
• Pattern 3 : Truncate such that the text ends with stop word.
[0051] The training data can be generated based on truncating the digital component page based on one or more of the above patterns or other patterns associated with identified truncation errors from historical data. By training the model to reduce truncation errors based on positive examples as explained above, a truncation error classifier can be provided to improve the accuracy in determining digital components that include such errors and reduce these errors in general for the final provided digital components.
[0052] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.
[0053] FIG. l is a block diagram of an example environment 100 in which generative artificial intelligence can be implemented to generate digital components. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, client devices 106, digital component servers 108, and a service apparatus 110. The example environment 100 may include many different electronic document servers 104, client devices 106, and digital component servers 108.
[0054] A client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over the network 102. A client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102, but native applications (other than browsers) executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.
[0055] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally or execute a gaming application that is at least partly stored and/or provided by a cloud server (e.g., online gaming applications). Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.
[0056] Digital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e.g., by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.
[0057] As illustrated, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content at a client device 106. Examples of electronic documents include online resources, webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed
sources. Native applications (e.g., “apps” and/or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to client devices 106 by electronic document servers 104 (“Electronic Doc Servers”).
[0058] For example, the electronic document servers 104 can include servers that host publisher websites. In this example, the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate presentation of the given webpage at the client device 106.
[0059] In another example, the electronic document servers 104 can include app servers from which client devices 106 can download apps. In this example, the client device 106 can download fdes required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively, or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server. In response to receiving the request, the cloud server can execute the application and stream a user interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app and communicate any user interactions with the user interface back to the cloud server for processing.
[0060] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and/or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital component) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source.
[0061] In some situations, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references the service apparatus 110, or a particular service provided by the service apparatus 110. In these situations, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for digital components 112 (referred to as a “component request”), which is transmitted over the network 102 to the service apparatus 110. For example, the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data. The component request 112 can include event data specifying features such as a name (or network location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e g., the client device 106), and/or information that the service apparatus 110 can use to select one or more digital components, or other content, provided in response to the request. The component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.
[0062] The component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which digital components can be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and/or media types that are eligible for presentation in the locations can be provided to the service apparatus 110. Similarly, event data specifying keywords associated with the electronic document (“document keywords”) or entities (e g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provided to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document. The event data can also include a query that was submitted from the client device 106 to obtain a search results page.
[0063] Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that
provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device). Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can be formatted as packetized data having a header and payload data. The header can specify a destination of the packet and the payload data can include any of the information discussed above.
[0064] The service apparatus 110 chooses digital components (e.g., third-party content, such as video fdes, audio fdes, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and/or using information included in the component request 112.
[0065] In some implementations, a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delays in providing digital components in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.
[0066] Also, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negatively impacting a user's experience with the electronic document. Further, delays in providing the digital component can result in a failed delivery of the digital component, for example, if the electronic document is no longer presented at the client device 106 when the digital component is provided.
[0067] In some implementations, the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital component in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a
corpus of millions of available digital components (DCi.x). The millions of available digital components can be indexed, for example, in a digital component database 116. Each digital component index entry can reference the corresponding digital component and/or include distribution parameters (DPi-DPx) that contribute to (e.g., trigger, condition, or limit) the distribution/transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.
[0068] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and/or consumption details (e.g., page viewport, user scrolling speed, or other information about the consumption of data). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and/or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify an eligibility value (e.g., ranking measure, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution/transmission (e.g., among other available digital components).
[0069] The identification of the eligible digital component can be segmented into multiple tasks 117a- 117c that are then assigned among computing devices within the set of multiple computing devices 114. For example, different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a-l 18c of the analysis back to the service apparatus 110. For example, the results 118a-
118c provided by each of the computing devices in the set 114 may identify a subset of digital components that are eligible for distribution in response to the component request and/or a subset of the digital components that have certain distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying the subset of digital components having distribution parameters that match at least some features of the event data.
[0070] The service apparatus 110 aggregates the results 118a-l 18c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components) based on the outcome of one or more content evaluation processes, as discussed below. In turn, the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.
[0071] In some implementations, the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g., a URL) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and transmit, to the client device 106, digital component data (DC Data) 122 that presents the given winning digital component in the electronic document at the client device 106.
[0072] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the script 154. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within, e.g., besides, the native
content 152 of the electronic document 150. In some implementations, the digital component is overlayed over (or adj acent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service apparatus 110 can specify a location or object within the scene depicted in the video content over which the digital component is to be presented.
[0073] The service apparatus 110 can also include an artificial intelligence (“Al”) system 160 configured to autonomously generate digital components, either prior to a request 112 (e.g., offline) and/or in response to a request 112 (e.g., online or real-time). As described in more detail throughout this specification, the Al system 160 can collect online content about a specific entity (e.g., digital component provider or another entity) and summarize the collected online content using one or more language models 170, which can include large language models.
[0074] A large language model (“LLM”) is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.
[0075] The language model 170 can be any appropriate language model neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto- regressively generates an output sequence made up of text tokens from the vocabulary. For example, the language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model.
[0076] In some situations, the language model 170 can be referred to as an auto-regressive neural network when the neural network used to implement the language model 170 auto- regressively generates an output sequence of tokens. More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any
previous positions in the output sequence that precede the particular position of the particular token, and a context input that provides context for the output sequence.
[0077] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
[0078] More specifically, to generate a particular token at a particular position within an output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The neural network of the language model 170 can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network of the language model 170 can greedily select the highest-scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.
[0079] As a particular example, the language model 170 can be an auto-regressive Transformer-based neural network that includes (i) a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.
[0080] The language model 170 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A.Wu, E. Eisen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifire, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D.
Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs/2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
[0081] Generally, however, the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.
[0082] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.
[0083] Generally, because the language model 170 is auto-regressive, the service apparatus 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from score distributions generated by the language model 170, using a Sample-and-Rank decoding strategy, by using different random seeds for the pseudo-random number generator that is used
in sampling for different runs through the language model 170 or using another decoding strategy that leverages the auto-regressive nature of the language model.
[0084] In some implementations, the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model 170 to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.
[0085] For example, the service apparatus 110 (e.g., Al system 160), or a separate training system, pre-trains the language model 170 (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. As a particular example, the language model 170 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus.
[0086] In some implementations, the Al system 160 can generate a prompt 172 that is submitted to the language model 170, and causes the language model 170 to generate the output sequences 174, also referred to simply as “output”. The Al system 160 can generate the prompt in a manner (e.g., having a structure) that identifies an online source (or more) of information, such as a website or data repository. In some implementations, the prompt can specify a set of constraints the language model 160 must use to generate a summary of information found at the online sources specified in the prompt 172, where the summary would be submitted to the language model 170. To initiate creation of the output sequences 174, the Al system 160 submits the prompt 172 to the one or more language models 170, which use the prompt 172 to evaluate the information found at the online sources specified in the prompt 172, and generate the output 174 that summarizes the information according to the constraints specified in the prompt 172.
[0087] The Al system 160 can use the generated summary as part of another prompt 172 that is sent to the language model 170. For example, the Al system 160 can insert the generated summary into an additional prompt 172 (e.g., a prompt generated after receiving the summary) that is submitted to the language model 170 as a constraint for generating clauses for use when generating digital components by the Al system 160. In a particular example, assume that the Al system 160 is generating a digital component to provide in response to the request 112,
which includes a keyword/query. In this example, the Al system 160 can generate the additional prompt 172 to include the query and a summary of online content (or constraints to generate a summary based on a provided link to an online resource) received in the prior output 174. The additional prompt 172 can also include instructions regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, or semantically styled, among other example configurations for the output clauses (e.g., specifying content that should be excluded from the clauses, such as granular details, such as numbers).
[0088] In some implementations, submission of this additional prompt 172 to the language model 170 can cause the language model 170 to generate an additional output 174, which can include multiple sets of clauses generated according to the query and constraints. The additional output with the clauses is communicated electronically to the Al system 160. The Al system 160 receives the clauses of the additional output 174, and generates multiple candidate digital components that are candidates for distribution to the client device 106 in response to the request 112. In some implementations, each different candidate digital component includes a different combination of the clauses received from the language model 170 in the additional output 174. For example, assume that the additional output 174 includes 12 different clauses, and that the formatting of the digital components being generated by the Al system 160 each include space for three different clauses, the Al system 160 could make 220 different candidate digital components using 3 different clauses in each of the candidate digital components (e.g., 121/(3 ! (12-3)1 ) 220).
[0089] In general, the Al system 160 can generate new digital components based on one or more existing digital components related to a subject, e.g., that depict content related to the subject, and data or other content related to the one or more existing digital components. The Al system 160 can use one or more Al models, e.g., language models 170, to generate new clauses for the digital components based on the existing digital component(s) and related data and/or content. The related data and/or content can include content of one or more electronic documents related to the existing digital component(s), e.g., content of digital component page(s) linked to by the existing digital component(s). The new digital components can include one or more of the clauses, e.g., one or more clauses that meet one or more grounding criteria.
[0090] In some instances, the number of candidate digital components that can be generated can include all or some of all possible combination of available clauses. In some instances, the available clauses can include model generated clauses and/or user-defined clauses provided as input. In some implementations, a prediction can be made for the possible combinations to determine the relevance of the respective digital components from the respective combinations to the digital component page. In some cases, based on the determined relevance, a set of candidate digital components can be as associated with relevance above a threshold level of relevance and the set can be generated. For example, the relevance can be measured based on relevance criteria defining scoring of clauses to be included in a candidate digital component based on a received query for generating a digital component. In some cases, the relevance can be determined based on evaluation of relevance associated with a quality criterion (e.g., user defined), a query element, or a usefulness criterion. In some instances, the set of candidate digital components include clauses that are evaluated with high relevance as meeting a relevance threshold for one or more relevance criteria as explained before. In some implementations, the determination of the number of components to be generated as the set of candidate digital components can be determined based on optimization as a trade-off between resources needed to generate all the combinations and relevance of the combinations to be generated.
[0091] In some implementations, the Al system 160 can also create the candidate digital components using a set of different links to online content (e g., second level domain links to web pages discussing a topic of the candidate digital components, phone numbers, etc.), which can continue to exponentially increase the number of different candidate digital components that the Al system 160 can create using the clauses of the additional output 174 of the language model 170.
[0092] The Al system 160 can perform one or more post-processing operations that evaluate one or more characteristics of the multiple candidate digital components. In some implementations, the post-processing operations can include generating a grounding measure for each clause from the additional output 174. The grounding measure can be determined based on matching of text or clauses generated for the digital component with a digital component page (e.g., content of a digital component page linked to by the digital component) and/or other pages or electronic documents related to the digital component, and/or by applying
an Al model (e.g., a language model 170) to the clauses and/or content of the digital component page to determine whether the generated component is factual. This model can be a machine learning model that is trained to classify text as being factual and/or to output a score that indicates the likelihood that the text is factual. This output can be compared to a threshold, e.g., a grounding threshold, to determine whether the clause is factual. If so, the clause can be used to generate a digital component, e.g., by including the clause in the digital component.
[0093] In some implementations, a digital component can be generated in real-time, e.g., in response to a request, e.g., based on contextual data such as query context, user context (e.g., data about a user to which the digital component will be presented), device context (e.g., data about the type of client device and/or technical requirements or limitations of a user interface where the digital component is to be rendered, or operating system version), among other example contextual or other types of input. The generation of the digital component in realtime can be based on a language model that includes for example, two layers, and multiple parameters but smaller as a number compared to other language models discussed herein, such as for example language model 170. For example, the language model can be optimized to perform the digital component generation based on the input such as that described above in a short period of time and with optimized resource expenditures (e.g., minimal CPU cycles). The language model, even if including fewer parameters than other models, can be configured to generate digital components that meet defined quality criteria while generating the digital components that include fewer words or phrases (e g., within a threshold range number of words). Such a language model is optimized for query-to-text relevance and generated grounded components based on all of the input.
[0094] In some instances, the language model with fewer parameters can be optimized on the number of matching terms from the digital component to tokens to provide an end-result that is grounded with relation to the digital component page (e.g., the page linked to by the digital component). In some instances, the number of tokens can be evaluated to determine whether some tokens can be excluded from the model consideration while still being able to generate digital components that meet the requirements (e.g., quality criteria). Since the language model with fewer parameters is used in the context of a query, the provided digital components can be optimized to be relevant to the query, which can include query terms and/or contextual data
related to a query, as described above. The query can be a search query or a component request 112.
[0095] In some implementations, the language model with fewer number of parameters can be trained on specifically collected training data that may include synthetically generated portion that is used to optimize on an objective function for the result while using minimum number of parameters. In some instances, the generated synthetic portion of the training data can include customized text generated for the query that is grounded in the digital component (e.g., in the electronic documents related to the digital component) used for the training. In some instances, the training data can be labeled based on human input. In some implementations, the generated digital components as output can be processed to determine their quality by running classifiers that focus on grammatical rules defined for the component, relevance to the query, and truncation, jumble, or slang criteria in text of the digital components.
[0096] The post-processing operations can include an evaluation of the relevance of the clauses to the query constraint, a level of completeness of the clauses relative to content located at the link included in the candidate digital component, and/or an evaluation of the tone (e.g., positive or negative) of the clause. As discussed in more detail with reference to FIGS. 2A&B, 3, 4A&B, and 5, the post-processing operations can be used to measure or evaluate, or otherwise assign a level of priority to, each of the candidate digital components so that the Al system 160 can evaluate (or filter) the multiple candidate digital components relative to each other, and ultimately provide one or more of the highest ranking candidate digital components as output digital components as a reply 120 to the request 112. Note that, although the operations of the Al system 160 and language model 170 are described above as being performed responsive to receipt of the request 112, at least some of the operations can be performed prior to receipt of the request 112, as described in more detail below, e.g., with reference to FIG. 2A.
[0097] Furthermore, although a single language model 170 is shown in FIG. 1, different language models can be specially trained to process different prompts at different stages of the processing pipeline. For example, a more general (e.g., larger) language model can be used to generate the summaries of online content as an offline process (e.g., independent of receipt of the request 112), which can then be inserted into prompts that are input to a more specialized and faster language model in an online process (e.g., real-time in response to receiving the
request 112). Additionally, the Al system 160 can generate a set of candidate digital components as an offline process (e.g., prior to receiving the request 112), and store the set of candidate digital components in a database. In this scenario, when the Al system 160 receives the request 112, the Al system 160 can further evaluate the stored candidate digital components, e g., for selection, based on additional information included in the request 112 and/or other contextual data (e.g., time of day, day of week, weather conditions, etc.).
[0098] FIG. 2A is a block diagram 200 illustrating interactions between an Al 160, a language model 250, and a client device. In some implementations, the language model 250 can be substantially similar to the language model 170 of FIG. 1 and trained to generate digital components (offline or in real-time) as described, e.g., at FIGS. 3 and 5. The language model 250 can be trained to generate digital components in response to input including at least one of a previously generated digital component, page content (e.g., digital component page content of a web page provided as a reference link or as direct content), or other query input such as a product, service, location, time period (e.g., a holiday period). In some implementations, a digital component is requested to be generated in relation to an object or an item (e.g., those can be linked to a digital component page), and the artificial intelligence system 160 can process such query to generate clauses that can be used to create multiple candidate digital components. In that example, one or more of the candidate digital components can be provided in response to the query.
[0099] In digital component generation in real-time, the digital components are generated within a flow of a received query including one or more terms. An example process for realtime generation is illustrated in FIG. 3 and described below. In offline generation, digital components can be generated and stored before they are eligible to be provided in response a received request. An example process for offline generation is illustrated in FIG. 5 and described below.
[00100] The Al system 160 is configured to autonomously generate digital components, either prior to a request (e.g., offline) and/or in response to a request (e.g., real-time or on-the- fly) as described in relation to FIG. 1. For example, the request can be received from a client device 265 that can be the same as or substantially similar to the client device 106 of FIG. 1. The request can include a component request and/or search query. In accordance with the present disclosure, the Al system 160 can collect content, e.g., content that can be accessed on
the Internet, about a specific entity (e.g., digital component provider or another entity such as an organization that publishes a digital component). The Al system 160 can collect the content in response to a received request or as a pre-process step performed for provided references (e.g., links, URLs, etc.).
[00101] The Al system 160 includes a summary apparatus 220 that is configured to summarize the collected content (e.g., according to summary definition logic for example, using a language model, which can include large language models as described in relation to FIG. 1, using a set of rules, or otherwise based on text analysis and/or image analysis to detect words or phrases in objects part of the content of a web page (text and/or visual)). In some instances, the summary apparatus 220 can include logic to select a portion of the online content and provide such portion as a summary that can be used as input to the language model 250 to generate clauses as described.
[00102] In some implementations, determining the portion of the content that is to be used as input by identifying a first set of portions presented at locations associated with areas of a display screen that are associated with most viewing. For example, when a web page is displayed on a display screen, it can be determined which areas of the screen are mostly viewed (and or interacted with) based on monitoring user behavior (e.g., scrolling to various parts of the page) and/or click interaction (e.g., which parts of the page receive user clicks). Usually, such areas of the web page are used for positioning the most important information for targeting content consumers. Further, the determination of the portion of the content can also include identifying a second set of portions associated with latest updated data as part of the content of the digital component page. In some implementations, it can be configured that the determined portion of the content would be either the first set of the portions or the second set of the portions. Further criteria for determining which portions of the content are relevant and can be used for generating summarized text can be used.
[00103] The Al system 160 includes a prompt apparatus 230 configured to generate prompts (such as prompt 172 of FIG. 1) including a prompt 245 that can be submitted to the first language model 250, which can cause the language model 250 to generate output that is provided as response 255 to the Al system 160. The Al system 160 can generate the prompt 245 in a manner (e.g., having a structure) that identifies a source (or more) of information, such as a website or data repository. In some implementations, the language model 250 can be
trained to generate clauses that are to be used for new digital component generation to meet a grounding threshold so that output clauses are classified as factual and grounded. In some implementations, the language model 250 can be trained based on digital components obtained for the training and/or content of electronic documents, e.g., digital component pages for the digital components. The digital components used for training the language model 250 can include, for example, generated digital components. In some examples, the digital components used for the training of the language model can be existing components generated by the language model at previous requests or can be components generated through other means (e.g., other language models or other techniques). The response 255 can include generated clauses that can be used to generate a new digital component 260 at the Al system 160 that can be provided to the client device 265. In some implementations, the Al system 160 can use the clauses to create multiple different digital components, and then performs post-processing to select, from among the different candidate digital components, a set of output digital components.
[00104] The Al system 160 can obtain the response 255 and can include logic to perform post training analysis at the post training apparatus 240 to evaluate the obtained clauses from the language model 250 and to fdter those of the clauses that are grounded in the content of the respective online source. In some implementations, generated clauses can be evaluated based on quality error rating criteria that can measure the performance of generated clauses with respect to different quality characteristics. For example, the quality characteristics can be related to linguistic aspects of the generated causes (e.g., grammar and spelling errors) and/or semantics and context (e.g., not supported by the digital component page provided with a request for digital component generation, awkward wording, or else). In some implementations, automatically created digital components can be determined immediately as eligible to be provided upon generation in response to requests. In some cases, if it is determined that a digital component does not meet quality criteria for generation and/or determined as not used after generations, can be automatically removed from the set of digital assets stored for providing in response to requests.
[00105] The post training apparatus 240 can generate digital components by using one or more of the created clauses and can provide at least one digital component 260 to the client device 265 in response to the query 270. In some implementations, the generation of the digital
components can be performed after or before receipt of the query 270. In the latter case, the Al system 160 can pre-store digital components that are generated for different online sources (e.g., digital component pages) and stored at a digital components 285 storage at the memory structure 275. In some implementations, generated clauses from the language model 250 as obtained through the response 255, can be stored by the Al system 160 at the memory structure 275, for example, at a clause data 280 storage.
[00106] The post-training apparatus 240 can also include logic to evaluate the relevance of generated clauses to query constraints received with requests for generating digital components. The relevance can be determined as a level of completeness of one or more clauses to content located in a candidate digital component, and/or an evaluation of the tone (e.g., positive or negative) of the clause.
[00107] FIG. 2B is a flow chart of an example process 290 for generating grounded digital components at real-time using artificial intelligence. Operations of the process 290 can be performed, for example, by the service apparatus 110 of FIG. 1, or another data processing apparatus. The operations of the process 290 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 290.
[00108] In some implementations, the process 290 is executed on-the-fly in response to a query request. The example process 290 is executed to generate digital components, where the digital components as generated at real-time are relevant to the digital component page(s) associated with the request. In some implementations, the digital component execution can be performed without pre-storing of previously generated clauses or other digital components. In some implementations, clauses can be generated and prestored, where those clauses can be provided at query time for digital component generation to create multiple candidate digital components that can be provided in response to requests from client devices such as client device 265. In those cases, from multiple pre-stored clauses, a portion of the clauses can be used to meet the constraints of a query to generate digital components. The digital components as generated remain faithful to the context of the query (e.g., context of the query) related to the purpose of use of the digital components (e.g., generating a digital component which subject is providing context for an item or object such as ski shoes, summer event, other) once
generated and a digital component page linked to by the digital component. Such faithfulness of the digital components is important for fact grounding and hallucination control of the generated digital components in the context of requesting and using the digital component (e.g., constraints of the requests, user-awareness, item or object awareness, other) and/or the digital component page associated with the subject for requesting the query (e.g., entity organization).
[00109] At 291, a query is received, for example from a client device (e.g., the client device 265 of FIG. 2A) via a communication interface. The query is received from a user that interacts through the communication interface on the client device with an Al system that is used for processing the query. In some implementations, the Al system can be substantially similar to the Al system 160 of FIGS. 1 and/or 2A. For example, the user can enter a search query into a search interface that is in communication with the service apparatus 110 that includes the Al system 160. In another example, the client device of the user can submit a query in the form of a component request, e.g., when the client device loads an electronic document.
[00110] At 292, a set of resources relevant to the query are determined. For example, the set of resources can include web pages, other types of electronic documents as described in relation FIG. 1, or other forms of online content, among other examples as discussed throughout this specification. In some examples, the set of resources can include electronic documents, e.g., web pages, corresponding to candidate digital components that are identified based on the content of the received query. For example, the candidate digital components can be identified as described above with reference to FIG. 1.
[00111] At 293, a digital component page of a third party (as an online source) is determined. The digital component page is identified as being relevant to the query received at 292. For example, a digital component or the digital component page for the digital component can be determined as relevant to a term of the query. For example, if the query is defined to surface content related to “sport shoes”, a digital component page of a sport shoe provider can be determined at 293, e.g., based on a digital component related to sports shoes being identified. In some implementations, the digital component page can be determined as related to another digital component generated for the same provider or for items determined
to match the query. A digital component when generated can include a link to a digital component page used for the generation.
[00112] At 294, the digital component page is provided to a trained language model as input to request the generation of at least one clause to be used for generating a digital component. The digital component that is generated at the trained model can be selected from a set of candidate digital components, where the digital component is a component that meets a grounding threshold. In some implementations, the clauses generated by the trained language model can be classified based on measuring to determine which clauses can be classified as factual and grounded in content of the digital component page. In some implementations, the classifying of clauses can be performed as described in relation to grounding measure computations for clauses as described in relation to FIG. 1.
[00113] The trained language model can obtain as input only a portion of content of the digital component page. For example, the portion of content can be determined as a summary of the content as described at FIG. 2A in relation to the summary apparatus 220. In some implementations, the summary can be a text summary generated based on provided constraints for the length of the text, e g., not more than 120 symbols, 60 words, or otherwise. By constraining the amount of data related to the page that is provided as input to the Al model, the occurrence of hallucinations is reduced.
[00114] At 295, at least one digital component relevant to the query and to the digital component page resource of the third party is generated based on the generated at least one clause.
[00115] In some instances, the at least one digital component (or its generated clauses) is classified as relevant (or faithful) according to a digital component classifier model executed on digital components generated in response to the query and the digital component page. In some implementations, to classify digital components, the clauses data and the query data can be evaluated to identify matching concepts and/or mismatching concepts. Concepts do not have to be an exact match to be considered a match. For example, a concept related to the query data can be “winter” while a concept related to the clause data (and a respective digital component) can be “ski shoes.” While the words are not an exact match, it can be determined that they match based on the concepts being similar (e.g., having a threshold similarity measure) or being related to the same higher-level concept. A mismatching concept is a
concept for either the query data or the clause data that does not match a concept for the other of the query data or the clause data. In some implementations, mismatching concepts or mismatching concepts that do not have at least a threshold level of importance can be filtered from all generated digital components and not provided for digital component generation and/or for display on a client device in response to the received query.
[00116] The generated digital component at 295 in response to a determined resource relevant to the query (as part of the determined resources 292) is provided for display along with the set of search results determined in response to the query. In some instances, the digital components can also be provided otherwise, for example, displayed at a separate interface.
[00117] At 296, the at least one digital component is provided for display on a client device’s display, such as a display of the client device 265.
[00118] Example digital components that can be generated based on the process 290 of FIG. 2B are presented in Table 1 below. In the example of Table 1, a query (such as the query 291 of FIG. 2B) is received for generating a digital component for an adjustable desk of Brand X. A digital component page associated with Brand X can be either provided as part of the query or determined as relevant for the query based on other digital components for Brand X referencing the digital component page. In some cases, a previously generated digital asset for the same query and digital component page may be available. That previous digital component can be considered as initial digital component which was created based on an initial clause text as shown in Table 1. By generating a new clause for generating a digital component, the initial clause is in practice rewrite based on the trained language model (as trained at 294 of FIG. 2b) to be more relevant to the unique context of the query and the digital component page.
[00119] The new digital component can include the new clause, e.g., in place of the initial clause. The process 290 can be used to generate the new clause based on the digital component that includes the initial clause, content of the digital component page linked to by the digital component, and the query. The Al model can ensure that the new clause satisfies a
ground criterion with relation to the digital component page and, if so, the clause can be presented by the new digital component. In this way, accurate clauses can be generated based on query context and used to generate many different digital components that are queryspecific.
[00120] FIG. 3 is a flow chart of an example process 300 for training a language model and generating digital components based on the trained language model. In some implementations, the process 300 can be executed at an Al system such as the Al system described in relation to FIG. 1 and 2A and 2B. Operations of the process 300 can be performed, for example, by the service apparatus 110 of FIG. 1, or another data processing apparatus. The operations of the process 300 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 300.
[00121] At 310, a set of digital components is obtained by an Al system. The set of digital components are obtained to be used to train a first language model. In some implementations, the first language model is trained to generate clauses that are to be used to generate new digital components. The set of digital components can be digital components that have been previously generated and used in relation to an online source (e.g., a digital component page of an entity).
[00122] At 320, training data including (i) the set of digital components as generated by the first language model and (ii) a text summary of a specified source of online content are provided to train the first language model. The text summary of the specified source of online content can be generated as a summary of the content of the online source (e.g., a digital component page such as a web page). For example, the text summary can be generated as described in relation to the summary apparatus 220 of FIG. 2A.
[00123] At 330, the first language model is trained to generate clauses that are used to respond to queries for digital component generations. The generated clauses can be used to generate digital components that are grounded. In some instances, the generated clauses can be classified based on determining whether they meet a grounding threshold (e.g., by determining a grounding measure for the clauses and comparing it with the grounding threshold). In some cases, digital components can be generated where at least a portion of
those that are associated with clauses that are classified as grounded and those grounded clauses can be provided for use for generating digital components.
[00124] At 340, in response to sending a query for generating a digital component for an item or object, at least one digital component generated by the first language model based on generated clauses for the query is obtained. The provided digital component is grounded in the content of the specified source.
[00125] In some instances, there may be no digital components that are pre-generated that can be used to train the first language model as described at step 320. In those cases, digital components can be generated based on another large language model, and those can be used as training data for the first language model.
[00126] In some implementations, the training data that is provided at 320 can be filtered data obtained from existing digital components and content of online sources. For example, the text part of the content of online sources can be filtered to remove personal identifiable data, other types of sensitive data, and/or to reduce content to minimize the amount of data for processing by the first language model and thus to improve the performance of the digital component generation. The amount of training data used for training the language model has a direct effect on the processing time for the training and the resources that are to be used. The more training data is provided, the longer the training process can take and more resources for processing each set of training data would be used. By reducing the amount of training data while also keeping the training data relevant, the quality of the training may not be affected and accurate output results can be provided and yet the computational resources can be saved. [00127] FIG. 4A is a block diagram of an example architecture 400 of components interacting to generate digital components based on a trained language model. In some implementations, the trained language model can be substantially similar to the trained language model of FIG. 2A. In some instances, the architecture 400 can be implemented in the context of model generated digital components, where the architecture 400 can include a language model 415 that can be the same or substantially similar to the second language model. The architecture 400 can be implemented to support generative digital components based on a trained model based on training data including digital components (e.g., as generated by another language model or previously generated by the same language model) and a text summary of a source of content. The language model 41 can be trained to generate new
digital components that meet a grounding threshold to classify the clauses as factual and grounded.
[00128] In some implementations, the Al system can include pre-processing, processing, and post-processing logic to provide digital components that are faithful and grounded in the text of a provided online source (e.g., a digital component page). The preprocessing can include operations for preparing training data for training a language model, and the post-processing can include operations for evaluating results from the processing of the trained language model to provide results that meet threshold criteria associated with the quality of the provided digital components.
[00129] An input source 405 can be provided to a trained large language model, LLM 415 to generate faithful digital components 435 and to provide a report 440 to a user (or customer) that had requested the digital component generation. The LLM 415 is a language model trained to generate digital components, where the LLM can be trained as described in relation to FIGS. 1, 2A&2B, and 3. In some implementations, the LLM 415 is trained based on provided digital components that can be either pre-generated components or can be components requested to be generated for the purpose of creating training data (e g., generated by another language model as described in more detail in relation to FIG. 5).
[00130] The input source 405 can include text (or other image content) from content of a digital component page when presented in a web browser or other application. In some instances, the input source 405 can include text generated based on image content from the digital component page. For example, the image can be processed, e.g., using OCR text recognition, to identify text depicted by images of the digital component page.
[00131] The text can be pre-processed at 410 to remove portions of the text that are not visible when displayed in a web browser. The pre-processing can include removal of nonsalient text from the text content of the digital component page. The pre-processing can also include filtering of content to remove personal identifiable data, other types of sensitive data, and/or to reduce content to minimize the amount of data for processing by the language model and thus to improve the performance of the digital component generation as discussed in relation to FIG. 3. The LLM 415 can obtain the pre-processed input and generate clauses to be used to generate new digital components. When the clauses are generated by the LLM 415, post-processing steps can be invoked and a signal can be send to a computing device at 420.
The computing device can have logic to trigger clause evaluation and a request for classification of the generated digital components can be sent to the post-gen classifier 425. The digital components can be evaluated based on evaluation of the clause(s) used for the components’ generation as discussed in relation to FIG. 2A and 2B. The clauses can be classified by generating a classification metric for the grounding of the clauses in the text of the provided input (from 410). The clauses can be evaluated to determine whether the clauses meet a grounding threshold (e.g., predefined for the post-processing logic, or provided as dynamic input from a client device or other related device or service, among other examples) to classify the clauses as factual and grounded. The clauses that meet the grounding criteria can be provided to generate digital components that are faithful digital components 435. For example, a prediction regarding the likelihood that a particular digital component is grounded (e g., high likelihood that the digital component is factual, e.g., as determined based on the content of the digital component page) or ungrounded (e.g., includes information that cannot be verified in the text of the digital component page).
[00132] The post-processing operations can include, for example, evaluating the candidate digital components based on various criteria, and measuring each of the candidate digital components based on the evaluation. For example, one post-processing operation can perform a prediction regarding the likelihood that a particular candidate digital component is ungrounded (e.g., includes information that cannot be verified in a specified corpus). In some implementations, the likelihood can be determined on the scale of 0 to 1 or 1 to 100, and a threshold value to use as a reference point and provide those of the candidate digital components that are above the threshold value (e.g., 0.8 or 70). Other values can also be used. In some instances, the used threshold value can be dynamically determined or re-evaluated based on determining the number of candidate digital components that would have a likelihood above the threshold. The threshold value can be modified to tune the number of created candidate digital components that are to be considered, while taking into consideration the available computation resources for their generation and/or any rules for the resource usage (e.g., restrict the number to particular number of candidate components, or processing resources or time). Using this type of a post-processing operation allows for looser constraints in the construction of the specialized prompt, which can allow the language model to generate more creative candidate digital components, while still ensuring that the output digital
component has at least a baseline level of truthfulness. The post-processing operations can also use various heuristics (e.g., heuristic filters defined based on measurements (such as hallucination rate, egregious rate, other) for the relationships between digital components and digital component pages) to evaluate different characteristics of each of the candidate digital components, and the measures can be assigned based on the various heuristics. In some implementations, the measures are weighted and aggregated to create a final measure, which is used to measure the candidate digital components and determine output digital components. [00133] In some implementations, a heuristic filter can be applied to evaluate words from the digital components and match them with portions of the text from the digital component page to determine dispersion of the output text throughout the text of the digital component page. In some instances, if a digital component is associated with a high dispersion rate (e.g., words from the digital component are mapped to word in distant locations within the text (e.g., according to a distance criterion)), then the digital component can be considered to have a high likelihood of being a hallucination. In some instances, such heuristic filter can be implemented as part of the quality criteria for evaluating generated digital components to determine an output set of digital components. The output digital components can meet grounding criteria or other quality criteria to provide digital components that are faithful as described in relation to FIG. 4A. Additionally, or alternatively, a machine learning model can be trained to measure digital component quality based on heuristics and those quality measures can be used to determine measurements for the candidate digital components. One or more of the highest-evaluated candidate digital components are then selected for providing as output digital components at 435.
[00134] The post-processing can include explainability analysis executed at an explainability module 430. The generated clauses can be evaluated, and a report can be provided to for display at a user interface at 440. Quality can be evaluated using post-providing human evaluation to rate the quality of generated digital components relative to digital component benchmarks related to the context of providing the digital components (e.g., advertisement).
[00135] FIG. 4B is a block diagram of an example environment 450 for customization of generated digital components in real-time. In some implementations, when a digital component is generated at real-time (in real time) as described in relation to FIG. 2B, the
generation of the digital components can include further post-generation logic to fdter at least some of the created digital components and to provide at least one digital component that is high-quality and grounded in the content of the digital component page.
[00136] In some implementations, a generative model 465 can be used for generating digital components in real-time, where the generative model 465 can receive as input a query (e.g., including terms, constraints, or digital component requirements, among other query parameters). The generative model can be a large language model as described throughout this specification. The generated components can be provided for further evaluation at hallucination detectors 470 and for real-time policy check 475 to generate new digital components (e.g., headlines). The real-time policy check 475 can include a real-time check based on policy or quality criteria as defined, as well as ad-hoc human evaluation for grading against benchmarks (e.g., externally obtained).
[00137] FIG. 5 is a block diagram of an example process 500 of digital components creation based on training data generated by a language model. Operations of the process 500 can be performed, for example, by the service apparatus 110 of FIG. 1, or another data processing apparatus. The operations of the process 500 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 500. In some implementations, the example process 500 can be executed when previously generated digital components related to online sources that can be used as training data for training a language model are not readily available, in contrast to the process 300 where such training data is available.
[00138] At 510, prompts that each include a query and a text summary of a specified source of online content are generated to create digital components. The prompt generation can be executed by an Al system such as the Al system 160 of FIG. 1 and 2A. The digital components are generated by a language model that is different from the language model used to generatively create digital components as described at FIG. 3.
[00139] At 520, the digital components as generated by the language model based on the obtained plurality of prompts are obtained by the Al system. In some implementations, the plurality of the digital components that can be generated can be validated. Each of the digital
components of the plurality of digital components is classified as faithful according to a digital component faithfulness classifier model executed to classify the digital components.
[00140] At 530, training data is defined to include (i) the plurality of digital components and (ii) the text summary of the specified source of online content. The training data is provided to train another language model to generate clauses to be used to generate new digital components that meet a grounding threshold to classify the clauses as factual and grounded. In some instances, the training can be done in a substantially similar way as described in relation to step 320 of FIG. 3. In some instances, the digital components are obtained from a storage including generated digital components or the digital components can be provided in other ways, e.g., through user input. In some implementations the online content of a specified source (e.g., digital component page of an entity or a third party) can include text which can be processed to generate a text summary. For example, the summary can be generated as discussed throughout the present disclosure and for example at FIG. 2A. In some cases, the text summary of the specified source of online content can be reduced to include a first portion of the text summary that is limited to a specified character count to generate the clauses. By limiting the size of the text used for the training, the training can be executed faster and also resulting clauses from the trained language model can be of shorter size (and e.g., a smaller number of clause variations) that can further optimize the process.
[00141] In some implementations, reducing the text summary can include determining a second portion of the text summary of the specified source to include relevant information at least based on positioning of the second portion of the text within a display of the online content at a web browser previewing the specified source. In some instances, the reduced text summary can include a title of the content displayed at the web browser when previewing the specified source. In some instances, the reduced text summary can be defined based on portions of the text or part of the content as previewed at particular predefined locations, e.g., upper right comer, middle of a display screen, etc., or with predefined visual characteristics (e g., text with a particular font size or with a font size above a threshold size, or font style such as bold but not italic, among other example visual characteristics of the text). In some implementations, the definition of the reduced text summary includes determining the first portion of the text summary that includes at least some of the second portion of the text within a threshold content size. In some implementations, the text summary can be based on
identifying portions of the online content that are relevant for a query that can be defined as a relevance criterion.
[00142] In some implementations, each clause of the generated clauses is evaluated to determine a grounding measure that specifies a likelihood that content of the clause is present in the text summary of the specified source of online content. The clauses can be classified as factual or not based on comparing the determined grounding measures and the grounding threshold.
[00143] In some implementations, at least one digital component is generated by the second language model based on the generated clauses. The at least one digital component is processed by the artificial intelligence system, for example, to a data storage in the context of offline digital component generation or in response to a received real-time request.
[00144] At 540, the second language model is deployed to serve requests to generate clauses based on provided details for a source of online content (e.g., provided online content via a request at the Al system). The request for generating clauses can include a reference to an online source relevant for the generation. In some instances, a clause generated by the second language model can point to a token defined for an input text (e.g., the text summary or the reduced text summary) of the requested source (e.g., digital component page).
[00145] In some implementations, digital components can be generated based on clauses provided by the second language model.
[00146] At 550, the at least one digital component is provided in response to a received request. The providing of the digital component as an output can also include evaluating (560) digital components generated by the second language model as candidate digital components. The evaluation at 550 can be same or substantially similar to the post-processing and classification as described in relation to the post-generation classifier 425 of FIG. 4A, or the described operations executed at the post training apparatus 240 of FIG. 2A. At 570, based on the measuring, the at least one digital component that meets a threshold criterion for the measure of the at least one digital component is selected to be provided.
[00147] In some implementations, the measuring is performed based on a trained deep network model that determines a likelihood that a digital component generated by the second language model is to be interacted with by a user.
[00148] At 580, the at least one digital component is provided to a client device. For example, the client device may have requested digital components that triggered the start of method 500 and be associated with the generated prompts at 510. The client device may display the received digital component on a display screen and provide it for user interaction. [00149] FIG. 6 is a block diagram of an example computer system 600 that can be used to perform operations described above. The system 600 includes a processor 610, a memory 620, a storage device 630, and an input/output device 640. Each of the components 610, 620, 630, and 640 can be interconnected, for example, using a system bus 650. The processor 610 is capable of processing instructions for execution within the system 600. In one implementation, the processor 610 is a single-threaded processor. In another implementation, the processor 610 is a multi -threaded processor. The processor 610 is capable of processing instructions stored in the memory 620 or on the storage device 630.
[00150] The memory 620 stores information within the system 600. In one implementation, the memory 620 is a computer-readable medium. In one implementation, the memory 620 is a volatile memory unit. In another implementation, the memory 620 is a nonvolatile memory unit.
[00151] The storage device 630 is capable of providing mass storage for the system 600. In one implementation, the storage device 630 is a computer-readable medium. In various different implementations, the storage device 630 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.
[00152] The input/output device 640 provides input/output operations for the system 600. In one implementation, the input/output device 640 can include one or more of a network interface devices, e.g., an Ethernet card, a serial communication device, e g., and RS-232 port, and/or a wireless interface device, e.g., and 802.11 card. In another implementation, the input/output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 660. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[00153] Although an example processing system has been described in FIG. 6, implementations of the subject matter and the functional operations described in this
specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[00154] An electronic document (which for brevity will simply be referred to as a document) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.
[00155] For situations in which the systems discussed here collect and/or use personal information about users, the users may be provided with an opportunity to enable/disable or control programs or features that may collect and/or use personal information (e.g., information about a user’s social network, social actions or activities, a user’s preferences, or a user’s current location). In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user’s identity may be anonymized so that the no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.
[00156] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially- generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage
medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[00157] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[00158] The term “data processing apparatus” encompasses all kinds of apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the aforementioned. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[00159] This document refers to a service apparatus. As used herein, a service apparatus is one or more data processing apparatuses that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in block diagrams. However, while the service apparatus could be a single device or single set of devices, this disclosure contemplates that the service apparatus could also be a group of devices, or even multiple different systems that communicate in order to provide various content to client devices. For example, the service apparatus could encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.
[00160] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other
unit suitable for use in a computing environment. A computer program may, but need not, correspond to a fde in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[00161] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[00162] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[00163] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’ s client device in response to requests received from the web browser.
[00164] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an internetwork (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[00165] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[00166] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be
claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a sub-combination.
[00167] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. [00168] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
1. A computer-implemented method for generating digital components, the method comprising: generating, by an artificial intelligence system, a plurality of prompts that each includes a query for generating a plurality of digital components and a text summary of a specified source of online content for use in generating the plurality of digital components by a first language model; obtaining, by the artificial intelligence system, the plurality of digital components as generated by the first language model based on the plurality of prompts; providing training data including (i) the plurality of digital components and (ii) the text summary of the specified source of online content to train a second language model; training the second language model to generate clauses to be used to generate new digital components that meet a grounding threshold for the second language model to classify the clauses as grounded; and providing, by the artificial intelligence system, at least one digital component generated by the second language model based on the generated clauses.
2. The method of claim 1, further comprising: evaluating each clause of the generated clauses to determine grounding measures for each clause that specifies a likelihood that content of the clause is factual as being present in the text summary of the specified source of online content, wherein each of the clauses is classified as factual or not based on comparing the determined grounding measure and the grounding threshold.
3. The method of claim 1, wherein providing the at least one digital component generated by the second language model comprises: selecting the at least one digital component based on a selection criterion including the grounding threshold defining a likelihood threshold for determining that the at least one digital
component is factual and verifiable at a text corpus provided as context for generating the at least one digital components.
4. The method of any one of the preceding claims, wherein providing the training data comprises: reducing the text summary of the specified source of online content to include a first portion of the text summary that is limited to a specified character count to generate the clauses.
5. The method of claim 4, wherein reducing the text summary comprises: determining a second portion of the text summary of the specified source to include relevant information at least based on positioning of the second portion of the text within a display of the online content at a web browser previewing the specified source; and defining the first portion of the text summary included in the reduced text to include at least some of the second portion of the text within a threshold content size.
6. The method of any one of the preceding claims, further comprising: deploying the second language model to serve requests to generate clauses based on provided details for a source of online content.
7. The method of any one of the preceding claims, further comprising: generating, by the second language model, the at least one digital component based on generating clauses for a provided online content via a request received at the artificial intelligence system.
8. The method of any one of the preceding claims, comprising: determining the plurality of the digital components by validating the digital components generated by the first language model, wherein each of the digital components of the plurality of digital components is classified as faithful according to a digital component faithfulness classifier model executed to classify the plurality of digital components.
9. The method of any one of the preceding claims, wherein providing the training data comprises: determining the text summary based on identifying portions of the online content that are relevant for the query according to a relevance criterion.
10. The method of claim 9, wherein identifying the portions of the online content comprises: monitoring user behavior and/or click interaction with the online content when presented on a display screen; identifying a first set of portions presented at locations associated with areas of the display screen that are associated with viewing and interaction that are above a threshold level; and providing the first set of portions for determining the text summary.
11. The method of any one of the preceding claims, wherein providing the at least one digital component comprises: providing a request for the at least one digital component by invoking the second language model to provide clauses, wherein the request provides a source of online content; and generating a new clause, by the second language model, that is grounded in the text of the online content of the source.
12. The method of any one of the preceding claims, wherein providing the at least one digital component comprises: measuring a plurality of digital components generated by the second language model; and based on the measuring, selecting the at least one digital component that meets a threshold criterion for a measure determined for at least one digital component.
13. The method of claim 12, wherein the measuring is performed based on a trained deep network model that determines a likelihood that a digital component generated by the second language model is to be interacted with by a user.
14. The method of any one of the preceding claims, wherein a clause generated by the second language model points to a token defined for an input text of a requested source.
15. The method of any one of the preceding claims, wherein the text summary includes text content generated based on processing text content and/or image content extracted from the specified source of online content.
16. A computer-implemented method for generating digital components, the method comprising: obtaining, by an artificial intelligence system, a set of digital components to be used to train a first language model; providing training data including (i) the set of digital components as generated by the first language model and (ii) a text summary of a specified source of online content to train the first language model to generate clauses to be used to generate new digital components that meet a grounding threshold to classify the clauses as factual and grounded; and providing, by the artificial intelligence system, at least one digital component generated by the first language model based on the generated clauses.
17. A system comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to carry out the method of any preceding claim.
18. A computer readable storage medium carrying instructions that, when executed by one or more processors, cause the one or more processors to carry out the method of any one of claims 1 to 16.
19. A computer program product comprising instructions which, when executed by one or more computers, cause the one or more computers to carry out the steps of the method of any of claims 1 to 16.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363616159P | 2023-12-29 | 2023-12-29 | |
| PCT/US2024/057459 WO2025144540A1 (en) | 2023-12-29 | 2024-11-26 | Generative digital component creation |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4713828A1 true EP4713828A1 (en) | 2026-03-25 |
Family
ID=93925029
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24827491.2A Pending EP4713828A1 (en) | 2023-12-29 | 2024-11-26 | Generative digital component creation |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4713828A1 (en) |
| WO (1) | WO2025144540A1 (en) |
-
2024
- 2024-11-26 WO PCT/US2024/057459 patent/WO2025144540A1/en active Pending
- 2024-11-26 EP EP24827491.2A patent/EP4713828A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025144540A1 (en) | 2025-07-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8788442B1 (en) | Compliance model training to classify landing page content that violates content item distribution guidelines | |
| US11756059B2 (en) | Discovery of new business openings using web content analysis | |
| US8374983B1 (en) | Distributed object classification | |
| US10102246B2 (en) | Natural language consumer segmentation | |
| US8984414B2 (en) | Function extension for browsers or documents | |
| US12332965B1 (en) | Website personalization and interactive assistant | |
| CN112805715B (en) | Identifying entity-attribute relationships | |
| CN115280314B (en) | Pattern-based classification | |
| US20250124264A1 (en) | Generating customized content descriptions using artificial intelligence | |
| US20250139385A1 (en) | Efficient image generation using artificial intelligence | |
| US20250315463A1 (en) | Deep linking using generative artificial intelligence | |
| US20250086434A1 (en) | Artificial intelligence for evaluating attributes over multiple iterations | |
| US20220245345A1 (en) | Article topic alignment | |
| WO2025264203A2 (en) | Image generation using enhanced prompts for artificial intelligence models | |
| US20250148364A1 (en) | Generative artificial intelligence for generating responses based on predicted trajectories | |
| EP4587939A1 (en) | Generative artificial intelligence | |
| EP4699097A2 (en) | Generative artificial intelligence | |
| EP4713828A1 (en) | Generative digital component creation | |
| EP4713797A1 (en) | Efficient real-time digital component creation | |
| EP4581501A1 (en) | Language model for predicting digital component selection data | |
| US20260050772A1 (en) | Generative ai techniques guided by network signals | |
| US20250322214A1 (en) | Self-criticizing artificial intelligence system | |
| EP4710227A1 (en) | Image generation using prompt chains | |
| WO2024249391A1 (en) | Retrieval token generation from queries using language model | |
| WO2025018995A1 (en) | User clustering and prompt generation tools for refining outputs of language models |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251219 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |