WO2025259680A1 - Generating data items based on a multidimensional reward model - Google Patents
Generating data items based on a multidimensional reward modelInfo
- Publication number
- WO2025259680A1 WO2025259680A1 PCT/US2025/033016 US2025033016W WO2025259680A1 WO 2025259680 A1 WO2025259680 A1 WO 2025259680A1 US 2025033016 W US2025033016 W US 2025033016W WO 2025259680 A1 WO2025259680 A1 WO 2025259680A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data item
- vector
- neural network
- generated
- training
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/092—Reinforcement learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/004—Artificial life, i.e. computing arrangements simulating life
- G06N3/006—Artificial life, i.e. computing arrangements simulating life based on simulated virtual individual or collective life forms, e.g. social simulations or particle swarm optimisation [PSO]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/047—Probabilistic or stochastic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/088—Non-supervised learning, e.g. competitive learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
- G06N7/01—Probabilistic graphical models, e.g. probabilistic networks
Definitions
- Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. [0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
- a generative neural network system processes a conditioning input to generate a data item.
- a generative neural network system may include a language model, a vision- language model or the like which processes an input query, comprising a sequence of tokens and/or other data, to generate output data which may represent a piece of text, an image, an audio waveform, a video or the like which is a sensible response to the input query, e.g. an answer to a question posed by the input query.
- This specification describes systems and methods implemented as computer programs on one or more computers in one or more locations that enable steering a pre-trained generative neural network system towards responses that exhibit one or more desirable aspects. This specification also describes the training of the generative neural network system.
- the generative neural network system comprises a data item generator neural network configured to process a conditioning input to generate a data item according to a data item generation policy.
- the data item represents a response to the conditioning input, where the conditioning input may be, e.g. a “prompt” for the data item generator neural network.
- the data item generator neural network may be pre-trained (e.g. on a large, unlabeled dataset by self-supervised learning).
- the generative neural network system further may comprise a feature neural network subsystem comprising a feature neural network configured to process i) the conditioning input, and ii) the generated data item, to generate a vector of features; and to determine a reward value 38072935-1 of the data item from the vector of features and a vector of weights that characterizes one or more aspects of the generated data item.
- the vector of weights may be associated with a user or a (small) group of users.
- the group of users may comprise different users (e.g. users that share a preference or use case) and/or the same user at different times, places, contexts, or the like.
- the feature neural network may be configured to determine the reward value by determining an inner product of the vector of features and the vector of weights. For example, denoting the conditioning input ⁇ ( ⁇ ⁇ ⁇ and the generated data item ⁇ ( ⁇ ⁇ ⁇ , the feature neural network may generate the vector of features ⁇ ⁇ ⁇ , ⁇ using a learned function ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ R ⁇ (where ⁇ denotes learnable parameters of the feature neural network and ⁇ ⁇ 1).
- the vector of weights (denoted ⁇ , ⁇ may have the same dimensionality as the vector of features ⁇ , ⁇ (i.e. ⁇ ⁇ R ⁇ ).
- each dimension ⁇ , ⁇ may be considered to represent a criterion used by a population of users (e.g. human users) to express their preferences with respect to the one or more aspects of the generated data items.
- the ⁇ -th element ⁇ ⁇ of the vector of weights ⁇ may be considered to represent how much the user (or group of users) associated with the vector of weights ⁇ values (or does not value) criterion ⁇ ⁇ , ⁇ .
- the feature neural network enables deriving a reward function ⁇ ⁇ , ⁇ ⁇ , ⁇ , ⁇ (where ⁇ ... ⁇ denotes the inner product) that is specialized to a user (or group of users).
- Reward values generated using the user-specific reward function ⁇ ⁇ , ⁇ may be used to steer the pre-trained generative neural network system towards responses that are aligned with said user (as described in more detail below). Because the reward function ⁇ ⁇ , ⁇ processes (in addition to the conditioning input ⁇ and the generated data item ⁇ ) the vector of weights ⁇ having a plurality of elements ( ⁇ ⁇ R ⁇ with ⁇ ⁇ 1), the described system may be considered to implement a “multidimensional model”. [0009] The method comprises obtaining the vector of weights ⁇ for a data item to be generated (e.g. from a user-annotated preference dataset, as described further below). The obtained vector of weights ⁇ may characterize one or more aspects of the to-be-generated data item.
- the vector of weights ⁇ may be associated with a specific user (or group of users) and may specify the user’s preferences (the vector of weights ⁇ associated with a user h (or a group of users) is denoted ⁇ ⁇ hereafter).
- the method further comprises obtaining the conditioning input ⁇ for the data item to be generated, e.g. from a user.
- the user providing the conditioning input ⁇ may be the same or a different user than the user associated with the vector of weights ⁇ ⁇ .
- the method further comprises using the data item generator neural network to process the conditioning input ⁇ to generate the data item ⁇ based on a reward value ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ , ⁇ ⁇ of the data item determined from the vector of features ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ for the generated data item and from the vector of weights ⁇ for the data item t ⁇ ⁇ o be generated (e.g. reward may be determined by determining an inner product of the vector of features ⁇ ⁇ ⁇ , ⁇ and the vector of weights ⁇ ⁇ ).
- reward may be determined by determining an inner product of the vector of features ⁇ ⁇ ⁇ , ⁇ and the vector of weights ⁇ ⁇ ).
- the reward values can be readily adapted to specific user(s), e.g. by simply providing the appropriate vector of weights ⁇ ⁇ . Examples of how the vector of weights ⁇ ⁇ can be determined are described further below. In particular, this enables adapting the generator neural network system to preferences of “new” users, e.g. where the generator neural network system has not been trained on data annotated by these users. [0012] Many ways exist to generate the data item based on a reward value ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ . As one example, the reward value ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ may be used to fine-tune (i.e.
- the generative neural network system further train) the generative neural network system to implement an adapted data item generation policy (e.g. using known reinforcement techniques), and to process the conditioning input to generate the data item according to the adapted data item generation policy.
- the reward value ⁇ , ⁇ , ⁇ may be used to update a plurality of learnable parameters, e.g. weights, of the data item generator neural network based on the reward value.
- the reward value ⁇ , ⁇ , ⁇ may be used to generate the data item without resource-intensive (i.e. computationally and data intensive) re-training of the generative neural network system. This can be advantageous since it enables steering the output characteristic of the trained generative neural network system on computing devices with limited computational power and limited storage capacity, e.g.
- the conditioning input data ⁇ may be by processed using the data item generator neural network for each of a plurality of data generation episodes to generate a plurality of different versions of the data item according to the data item generation policy.
- the conditioning input and each respective version of the generated data item may then be processed to generate the vector of features for each respective version of the generated data item.
- a respective reward value of each version of the data item 38072935-1 can be determined from an inner product of the vector of features for the version of the generated data item and from the vector of weights for the data item to be generated.
- the respective reward values of the versions of the generated data item may then be used to select one of the versions as the generated data item, e.g. a version of the generated data having the highest reward value may be selected as the generated data item.
- the generative neural network system may comprise a plurality of pre-trained data item generator neural networks each configured to process the conditioning input to generate the data item according to a different respective data item generation policy (the policies are denoted ⁇ ⁇ ⁇ ).
- the number of data item generator neural networks may be equal to the number of elements in the vector of weights ⁇ ⁇ (i.e. the number of data item generator neural networks may be “d” when ⁇ ⁇ ⁇ ⁇ R with ⁇ ⁇ 1).
- each policy ⁇ ⁇ may be a solution from a reinforcement learning problem resulting from using ⁇ ⁇ , ⁇ reward function, but any suitable way of selecting the policies ⁇ ⁇ may be used.
- the data item generator neural networks may be pre- trained on the same training dataset.
- the method may further comprise processing the conditioning input data using each of the item generator neural networks to further generate at least one version of the data item according to each of the different respective data item generation policies ⁇ ⁇ .
- Using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item may comprise using the respective reward values of the versions of the generated data item to select one of the data item generator neural networks.
- the policy with the largest expected performance under the obtained vector of weight ⁇ is selected, i.e. a policy ⁇ ⁇ ⁇ may be selected according to ⁇ ⁇ arg m ⁇ ⁇ a ⁇ x ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ , ⁇ ⁇ .
- the version of the generated data item having the highest reward value may be selected as the generated data item, i.e. ⁇ ⁇ arg m ⁇ ax ⁇ ⁇ , ⁇ , ⁇
- a (human) user may not be able to clearly articulate their criteria and/or preferences in terms of 38072935-1 values of the vector of weights ⁇ ⁇ .
- the user may rank different versions of the data item and the weights are determined from the values of the vector of weights ⁇ ⁇ are determined from the ranking.
- obtaining the vector of weights may comprise processing an example generating conditioning input data using the data item generator neural network to generate a plurality of different versions of a corresponding example data item from the example generating conditioning input data, obtaining a ranking of the different versions of the example data item from a user, or a group of users, and determining the vector of weights from the ranking.
- a “user” includes a group of users, e.g. a plurality of different users, and can also refer to the same user in different circumstance, e.g. at different places, times, contexts, and so forth.
- the user may rank pairs of versions of the example data item (i.e.
- the user may indicate whether a first or a second version of the example data item is a preferred response given the example generating conditioning input data.
- the vector of weights ⁇ ⁇ may be determined from the ranking by processing the example generating conditioning input data and each version of the example data item to generate the vector of features for each version of the example data item, and by determining the vector of weights from the ranking and the vectors of features for the different versions of the example data item.
- a candidate vector of weights may be used to determine respective reward values for the first and second version.
- the values of the candidate vector of weights may be optimized to maximize a difference between the reward values, or equivalently to minimize a loss function that includes a term that measure the difference between the reward values of the first and second version.
- a reward value of the first version may be determined from the vector of features for the first version of the example data item and the candidate vector of weights
- a respective reward value of the second version may be determined from the vector of features for the second version of the example data item and the candidate vector of weights.
- the weights of the candidate vector of weights may be optimized to maximize a difference between the reward value of the first version of the example data item and the reward value of the second version of the example data item (assuming the user indicated that the first 38072935-1 version is preferred over the second version).
- the optimized vector of weights may then be determined as the obtained the vector of weights ⁇ ⁇ .
- the values of vector of weights ⁇ ⁇ may be determined by minimizing the loss function ⁇ ⁇ , ⁇ ⁇ log ⁇ , ⁇ where ⁇ ⁇ ⁇ ⁇ , ⁇ - ⁇ ⁇ , ⁇ ′ ⁇ when the user indicated that the first version y is preferred over the second version y’, and ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ ′ ⁇ - ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ when the user indicated that the second version y’ is preferred over the first version y.
- the vector of weights ⁇ can be determined from a small number of ranked samples.
- the candidate vector of weights may be initialized with values corresponding to the “average population preference” (as described further below, these values are naturally determined during training of the feature neural network). Initialization with the “average population preference” may further reduce the number of ranked samples needed to obtain the vector of weights ⁇ ⁇ .
- the generative neural network system comprises a data item generator neural network that generates an output token sequence from an input token sequence including the conditioning input. The data item generator neural network may then be configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens, that is used to select an output token for the output token sequence.
- the generative neural network system can have more than a billion parameters and can require substantial computing resources, power, and time to process a network input to train the model. Sometimes such models can have can more than 10 billion or more than 100 billion parameters.
- a digital assistant device e.g., a mobile device
- a computing system that includes a back end component, in particular a data server, in communication with the digital assistant device over a data communications network such as the Internet.
- a data communications network such as the Internet
- the described techniques facilitate a reduced a computational load, and improved load distribution, 38072935-1 e.g. when the generative neural network system is implemented in a multitasking and parallel processing computer system, distributed across multiple sites and interconnected by a data communication network.
- the described techniques enable a beneficial distribution of computing load between a local, mobile computing device and a remote back-end server in a network.
- the system may be implemented on a digital assistant device such as a mobile device.
- the generative neural network system can be implemented the generative neural network system (wholly) on the mobile device.
- the mobile device generally has less working memory than the back-end data server, less computational capacity than the back-end data server, or both – which makes (re-)training of the generative neural network system on the local, mobile device undesirable if not unfeasible.
- Computational capacity can be measured in computing operations per second, e.g. FLOPS (floating point operations per second).
- FLOPS floating point operations per second
- the remote server may pre-train the generative neural network system and send the values of a plurality of learned parameters defining the trained generative neural network system to the mobile device via the communication network.
- the user of the mobile device may provide the conditioning input for the data item (by inputting the conditioning input into the mobile device).
- a computer-implemented method of training a feature neural network e.g. the above described the feature neural network.
- the method may involve obtaining a plurality of training items ⁇ ⁇ ⁇ , ⁇ , ⁇ ′ ⁇ , ⁇ ,h ⁇ ⁇ ⁇ ⁇ , each training item comprising a respective training conditioning input data ⁇ ⁇ , a plurality of different versions ⁇ ⁇ , ⁇ ′ ⁇ of a corresponding training data item generated from the training conditioning input data using the data item generator neural network, a ranking ⁇ ⁇ of the respective different versions of the training data item obtained from a respective user, and user information h ⁇ specifying said user from a group of users (e.g. a group of human raters), and training the feature neural network using the plurality of training items.
- a group of users e.g. a group of human raters
- the method may involve obtaining a plurality of training items ⁇ ⁇ ⁇ , ⁇ , ⁇ ′ ⁇ , ⁇ , h ⁇ ⁇ ⁇ ⁇ , each training item comprising a respective training conditioning input data ⁇ ⁇ , a plurality of different versions ⁇ ⁇ , ⁇ ′ ⁇ of a corresponding training data item generated from the training conditioning input data using the data item generator 38072935-1 neural network, a ranking ⁇ ⁇ of the respective different versions of the training data item, and information h ⁇ specifying an origin of the ranking, and training the feature neural network using the plurality of training items.
- the feature neural network can have any suitable architecture and can include, e.g., one or more feed forward neural network layers, one or more recurrent neural network layers, one or more convolutional neural network layers, one or more attention neural network layers, or one or more normalization layers.
- the feature neural network may comprise a Transformer neural network that is configured to process the conditioning input (which may comprise an input sequence), and the generated data item (which may comprise an output sequence), to generate the vector of features.
- the feature neural network can have fewer learnable or learned parameters than the data item generator neural network.
- the feature neural network is trained in a plurality of training iterations.
- a respective training reward value ⁇ ⁇ ⁇ , ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ′ ⁇ of each version of the respective training data item may be determined from the vector of features for the respective version of the training data item and from a candidate vector of weights associated with the user specified in the respective user information.
- a plurality of learnable parameters e.g.
- weights, of the feature neural network and the weights of candidate vectors of weights may be updated based on the training reward values. This generally involves backpropagating gradients of an objective function to update the learnable parameters using any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm.
- gradient descent optimization algorithm e.g. Adam or another optimization algorithm.
- the plurality of learnable parameters ⁇ of the feature neural network and the weights of candidate vectors of weights may be updated to minimize the loss function ⁇ ⁇
- this minimization results in updates for the plurality of learnable parameters of the feature neural network and the matrix ⁇ .
- the so-obtained matrix ⁇ can be used to initialize the above described candidate vector of weights when a ranking of the 38072935-1 different versions of example data item is processed to determine the vector of weights.
- the candidate vector of weights may be initialized as the average over the rows of ⁇ which may be considered to correspond to initialising the model with the “average human preference”.
- the method For each of a plurality of the training data items the method obtains a vector of weights that characterizes one or more aspects of data items to be generated, processes the training conditioning input and the training data item using a feature neural network subsystem, e.g. as described herein, to generate a vector of features, and determining a reward value for the data item from the vector of features and the vector of weights.
- the method trains the data item generator neural network using the reward values, using any suitable objective function, e.g. a maximum likelihood objective, a cross-entropy objective, and so forth. In general the training can involve backpropagating gradients of the objective to update learnable parameters of the data item generator neural network.
- a system that includes one or more computers and one or more storage devices communicatively coupled to the one or more computers and storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the previously described method.
- one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the previously described method.
- this is achieved without the need of re-training the generative neural network system.
- This makes the described techniques particularly suitable for implementations in which the generative neural network system is pre-trained on a remote 38072935-1 server and then implemented on a mobile device which lacks the computational power to fully re-train the system.
- the mobile device can use the described techniques to locally adapt the data item generation of the generative neural network system without having to receive and process large amounts of training data (reducing network traffic and computation cost).
- the vector of weights that characterizes the one or more desired aspects of the data items can conveniently be found by processing only a few user-annotated samples.
- FIG.1 shows an example system for generating data items and associated rewards.
- FIG.2 is a flow diagram of an example process for generating data items and associated rewards generating data items and associated rewards.
- FIG.3 shows an example training system for a feature neural network.
- FIGS.1 to 8 show experimental results generated by the described techniques.
- LM language model neural networks
- an input query comprising a sequence of tokens and/or other data such as a media element
- an output representing a sensible response to the input query, e.g. an answer to a question posed by the input query.
- a LM is first pre-trained on a large, unlabelled (text) dataset (e.g. Web data) and then fine-tuned for downstream tasks.
- the pre-training may be performed by self-supervised learning (also known as predictive learning) in which the LM learns to correctly predict continuations of received samples of a text database, such as a large, publically-available database of natural language. 38072935-1 [0052] Fine-tuning is performed because self-supervised LMs often exhibit factual errors, biases, and other undesirable behavior. Thus, the pre-trained LM may be fine-tuned to align with human expectations and values. Fine-tuning can be performed with a human-annotated preference dataset, e.g. via reinforcement learning from human feedback (RLHF) which performs alignment by first learning a scalar-valued reward model, that mimics human judgment, and then employs reinforcement learning to optimize the LM against this reward.
- RLHF human feedback
- RLHF models preferences using a reward model that does not distinguish between people. For example, human feedback may be collected by asking humans to rank examples of a LM’s behavior/output, and then all of the feedback is integrated to derive a single reward function that reflects the preferences of the population of interest. However, this approach is not effective when there is considerable disagreement across the population (of human raters). This is likely the case in the training of LMs and other generative neural networks. [0054] As an illustration, given a pair of alternative responses to a subjective question, 51% of the target audience may prefer the first option while the remaining 49% may prefer the second.
- a user-annotated preference dataset may be generated as follows.
- a human rater is sampled, h ⁇ H where ⁇ H ⁇ ⁇ H ⁇ is a distribution over a set of users H. Then, a context is sampled, ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ ⁇ is a distribution over ⁇ that may depend on h.
- Two 38072935-1 responses are then sampled ⁇ , ⁇ ′ ⁇
- the two sampled responses ⁇ and ⁇ ′ are ranked by the human rater h, who ranks them by sampling ⁇ , ⁇ , ⁇ ,h ⁇ , where ⁇ ⁇ ⁇ 0,1 ⁇ is a Bernoulli distribution whose mean is ⁇ ⁇ ⁇
- intra-user generalisation ⁇ is used to model ⁇ ⁇ ⁇
- inter-user generalisation the same data is used to model ⁇ ⁇ ⁇
- Generalization across all users in H may require a distinct reward function ⁇ , ⁇ per user h ⁇ H.
- Another way to accomplish this is by employing a separate parametric function ⁇ per user, but it may be hard to obtain inter-user generalization in this way.
- Another way is to use two disjoint sets of parameters: i) a set of common parameters ⁇ ⁇ R ⁇ that is shared among all users h ⁇ H, including those h ⁇ H ⁇ , and ii) parameters ⁇ ⁇ ⁇ ⁇ R that are specific to user h. As further described below with reference to FIG.3, this of the parameters may induce a division of the training procedure, i.e.
- the data in ⁇ ⁇ can be used to learn ⁇ and ⁇ ⁇ for a rater h ⁇ ⁇ H ⁇ , and additional data ⁇ ⁇ ⁇ , ⁇ , ⁇ ′ ⁇ , ⁇ ⁇ ⁇ (with ⁇ , ⁇ , ⁇ ⁇ , h ⁇ ) cont feedback provided by a can be used to learn the parameters ⁇ ⁇ .
- the number of shared parameters may be much larger than the number of parameters that are specific to a given individual (i.e. ⁇ ⁇ ⁇ ). This may be advantageous to achieve quick adaptation of the model to a specific user h ⁇ H ⁇ H ⁇ .
- FIG. 1 shows an example computer system 100.
- the computer system 100 is an example of a system, implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below are implemented.
- the computer system 100 is configured to generate data items aligned with preferences of a specific user or a group of users.
- the computer system 100 employs the above-described disjoint sets of parameters ⁇ ⁇ R ⁇ (shared among all users) and user-specific parameters ⁇ ⁇ R ⁇ (specific to user or a group of users) to implement a user-specific reward function. Reward values generated by the computer system 100 are then used to steer/modulate the output of a generative neural network.
- the computer system 100 comprises a generative neural network system 110 comprising a data item generator neural network 112 and a feature neural network 114.
- the data item generator neural network 112 configured to process a conditioning input 116 (denoted ⁇ ) to generate a data item 118 (denoted ⁇ ) according to a data item generation policy.
- the conditioning input 116 defines a query or task for the data item generator neural network 112, and may be obtained from a user (e.g. via user interface), from another software application, or the like.
- the conditioning input 116 can include one or more modalities including text, image, audio, video, or a combination of one or more such modalities.
- the network input can include multiple modalities depending on the specific architecture of the data item generator neural network 112.
- the data item 118 represents a response to the conditioning input 116.
- the data item 118 may comprise text, image and/or audio tokens.
- the generative neural network system 110 is further configured to obtain (e.g. as further input) a vector of weights 120, and to process the conditioning input 116, data item 118 and the vector of weights 120 to generate a reward value 122.
- the vector of weights 120 is specific to a user and denoted ⁇ ⁇ R ⁇ for user h.
- ⁇ for user h may be expressed as ⁇ , ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ where ⁇ , ⁇ is a function that maps a conditioning input ⁇ and a data item ⁇ 38072935-1 to a vector of “reward features” ⁇ , ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ R ⁇ , and ⁇ , ⁇ denotes inner product.
- a personalized version of the Bradley-Terry model may then be expressed as: ⁇ ⁇ ⁇ ⁇
- the feature neural network 114 is configured to process i) the conditioning input ⁇ 116, and ii) the generated data item ⁇ 118, to generate a vector of features ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ 124.
- the generative neural network system 110 is further configured to generate the reward value 122 based on the vector of features ⁇ ⁇ ⁇ , ⁇ 124 and the vector of weights ⁇ 120, e.g. based on ⁇ , ⁇ , ⁇ ⁇ ⁇ , ⁇ , ⁇ .
- the feature neural network can have any suitable architecture and can include, e.g., one or more feed forward neural network layers, one or more recurrent neural network layers, one or more convolutional neural network layers, one or more attention neural network layers, or one or more normalization layers.
- the feature neural network may comprise a Transformer neural network that is configured to process the conditioning input (which may comprise an input sequence), and the generated data item (which may comprise an output sequence), to generate the vector of features.
- the feature neural network can have fewer learnable or learned parameters than the data item generator neural network.
- the so-obtained the reward value 122 is used to modulate the output of the data item generator neural network 112 so as to align the output with the preferences of the user h with respect to one or more aspects of the generated data item.
- the generative neural network system 110 is configured to adopt ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ⁇ as a criterion to select responses generated by the data item generator neural network 112. More specifically, the generative neural network system 110 is configured to implement a “best-of-n” approach, i.e.
- “n” candidate generated data items ⁇ ⁇ ⁇ x ⁇ (where ⁇ x ⁇ denotes the data item generation policy of the data item generation neural network 112) are generated and the generative neural 38072935-1 network system 110 selects, as out data item 118, the data item ⁇ ⁇ associated with the highest reward value, i.e. based on arg max ⁇ ⁇ , ⁇ , ⁇ .
- the vector of weights ⁇ ⁇ may be used to directly modulate the output of the data item generator neural network.
- the vector of weights ⁇ ⁇ may be processed by one of these methods to (immediately) obtain a policy that is specialised to the corresponding reward function.
- the generative neural network system may comprise a plurality of pre-trained data item generator neural networks each configured to process the conditioning input to generate the data item according to a different respective data item generation policy (the policies are denoted ⁇ ⁇ ⁇ ⁇ ).
- Each data item generator neural network may be trained induced by different linear combinations of reward features ⁇ ⁇ , and the outputs of the data item generator neural networks may be combined based on the vector of weights ⁇ ⁇ .
- the number of data item generator neural networks may be equal to the number of elements in the vector of weights ⁇ ⁇ (i.e. the number of data item generator neural networks may be “d” when ⁇ ⁇ ⁇ ⁇ R with ⁇ ⁇ 1).
- Each policy ⁇ may be a solution from a reinforcement learning from using ⁇ ⁇ , ⁇ as the reward function, but any suitable way of selecting the policies ⁇ ⁇ may be used.
- the data item generator neural networks may be pre-trained on the same training dataset.
- the method may further comprise processing the conditioning input data using each of the item generator neural networks to further generate at least one version of the data item according to each of the different respective data item generation policies ⁇ ⁇ .
- the plurality of pre-trained data item generator neural networks may be used to implement a form of generalised policy evaluation termed “successor features” in the field of “generalised policy improvement” (GPI), for example as described in A. Barreto et al., “Successor features for transfer in reinforcement learning”, Advances in Neural Information Processing Systems (NIPS), pages 4055–4065. Curran Associates, Inc., 2017.
- GPI generalised policy improvement
- the method may comprise computing successor features ⁇ ⁇ ⁇ of the policies ⁇ ⁇ ⁇ as ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇
- the method may further comprise computing a policy for user as GPI as ⁇ ⁇ arg m ⁇ ⁇ a ⁇ x ⁇ m ⁇ ⁇ a ⁇ x ⁇ ⁇ ⁇ , ⁇ ⁇ .
- This illustrates how to quickly compute a policy for an user represented by its preference vector ⁇ ⁇ .
- using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item may comprise using the respective reward values of the versions of the generated data item to select one of the data item generator neural networks. For example, the policy with the largest expected performance under the obtained vector of weight ⁇ ⁇ ⁇ is selected, i.e.
- a policy ⁇ ⁇ ⁇ may be selected according to ⁇ ⁇ arg max ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ .
- the version of the generated data item having the highest reward value may be selected as the generated data item, i.e. ⁇ ⁇ arg m ⁇ ax ⁇ ⁇ , ⁇ , ⁇ ⁇
- a data item generator neural network that generates an output token sequence from an input token sequence including the conditioning input.
- the data item generator neural network may then be configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens, that is used to select an output token for the output token sequence.
- the tokens can represent text, e.g., words, wordpieces or characters, in a natural or computer language.
- text may be received, e.g., as a series of encoded characters, e.g. UTF-8 encoded characters; such “characters” can include Chinese and other similar characters, as well as logograms, syllabograms and the like.
- a text encoder i.e.
- a tokenizer can process a sequence of text to represent the text as a series of text tokens from a vocabulary of text tokens, e.g. that each represent words, wordpieces or characters in a natural or computer language.
- the computer language may be any formal language used to communicate with a computer, e.g. a markup language, or a command or configuration language, or a data exchange language such as JSON, or a programming language.
- the tokenizer can, e.g., implement BPE (Byte Pair Encoding) or Wordpiece 38072935-1 tokenization.
- the text can be obtained from audio data representing speech; the output tokens may be converted into audio data that represent speech corresponding to the text.
- the tokens may represent an image.
- a set (sequence) of input or output tokens can represent an image.
- Each image token may comprise a block encoding of values of the pixels in a different region of an image that maps a set of values of the pixels to a respective image token.
- the block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network.
- the tokens may represent an audio waveform.
- a set (sequence) of input or output tokens can represent audio data representing an waveform e.g. instantaneous audio amplitude values or time-frequency audio data.
- Each image token may comprise a block encoding of the audio waveform in a different time segment of the audio that maps a set of values representing the audio waveform to a respective image token.
- the block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network.
- audio data or an image may be flagged by a start-of-audio token or start-of-image token.
- the generative neural network system can also or instead comprise a data item generator neural network that is a diffusion model neural network.
- a diffusion model neural network can be a neural network that has been trained to process a diffusion input comprising a current noisy data item and data specifying a current time to generate a diffusion output that defines an estimate (given the current time) of either a noise component of the current noisy data item, i.e. an estimate of the noise that has been added to an original data item to generate the current noisy data item; or of a de-noised version of the current noisy data item.
- the generative neural network system can be a multimodal system that is configured to process a conditioning input comprising one or more of text data, audio data defining an audio signal (e.g.
- the conditioning input may comprise text and the data item may comprise an image or an audio signal that represents speech an image generated in response to the text, e.g. described by the text.
- the conditioning input may comprise an audio signal that represents speech, or an image
- the data item may comprise text, e.g. that describes 38072935-1 the conditioning input.
- the conditioning input may comprise one or more of text, audio, video, or image data
- the data item may comprise one or more of text, audio, video or image data.
- the conditioning input may comprise an observation, e.g. of a real world environment, e.g. from sensor such as a camera or other image sensor; and optionally additional information such as information defining a particular task to be deformed.
- the output data item may comprise agent control data that defines one or more actions to be performed by an agent, e.g. by a mechanical agent such as a robot or autonomous vehicle, to perform a task.
- the reward model(s) may, e.g., define a preferred trajectory of motion of the mechanical agent in the (real-world) environment.
- the generative neural network system may comprise a language and/or image generation neural network system, that may have been trained before being fine-tuned by the above described method.
- the conditioning input may comprise a prompt, e.g. a natural or computer language prompt for the generative neural network system.
- the generated data item may comprise a natural or computer language and/or image response to the prompt.
- the generative neural network system can have any appropriate architecture for processing the conditioning input to generate the data item.
- the generative neural network system may comprise an auto- regressive generative model (e.g., a Transformer, a recurrent neural network, etc.) that can auto-regressively generate an output sequence as the data item based on the conditioning input.
- the generative model can, for example, comprise a large language model (LLM) that can auto- regressively generate tokenized representations of text data, a vision-language model (VLM) that can auto-regressively generate tokenized representations of image or video data, e.g. in response to a text conditioning input or that can auto-regressively generate tokenized representations of text, e.g.
- the generative neural network system may comprise a diffusion model (e.g., a denoising diffusion model, a score-based diffusion model, a latent diffusion model, etc.) that can generate the data item by repeatedly transforming samples from a noise distribution (e.g., a Gaussian distribution) based on the conditioning input over a sequence of 38072935-1 iterations.
- a diffusion model e.g., a denoising diffusion model, a score-based diffusion model, a latent diffusion model, etc.
- a noise distribution e.g., a Gaussian distribution
- the generative neural network system may comprise a diffusion model that transforms samples from the noise distribution using a denoising neural network with any appropriate architecture (e.g., a convolutional neural network, a recurrent neural network, etc.). Such a diffusion model may be used to generate, e.g., a still or moving (video) image.
- the generative neural network system may comprise a neural network that can generate the data item by transforming samples from a noise distribution (e.g., a Gaussian distribution).
- the generative neural network system may comprise, e.g., a generator network of a generative adversarial network, a decoder of a variational auto-encoder, a normalizing flow, and so on.
- an image may be any still or moving image, i.e. the image may be part of a video, in 2D or 3D, and may be a monochrome, color or hyperspectral image, i.e. comprising monochrome or color pixels.
- an “image” includes a point cloud e.g. from a LIDAR system, and a “pixel” includes a point of the point cloud.
- An image may have been captured by a camera or other image sensor from the real world; and objects in the image may comprise physical objects, represented by the image.
- FIG. 2 is a flow diagram of an example process 200 of using a generative neural network system to generate a data item. The process 200 of FIG.
- the process 200 may be implemented by one or more computers in one or more locations.
- the process 200 may be implemented by the system of FIG. 1, and for convenience the process is described with reference to FIG.1.
- the vector of weights 120 is obtained from a user.
- the corresponding the vector of weights ⁇ ⁇ can be readily retrieved.
- a “new” or “unseen” user i.e. a user that did not acted as a rater for the training dataset used to train the feature neural network 114, may rank different versions of data items and the values of the vector of weights ⁇ ⁇ are determined from the ranking.
- an example generating conditioning input data may be processed using the data item generator neural network to generate a plurality of different versions of a corresponding example data item from the example generating conditioning input data, and a ranking of the different versions of the example data item may be obtained from the user, or a group of users.
- the user may rank pairs of versions of the example data item (i.e. the user may indicate whether a first or a second version of the example data item is a preferred response given the example generating conditioning input data.
- the aforementioned dataset ⁇ ⁇ ⁇ 38072935-1 ⁇ , ⁇ , ⁇ ′ ⁇ , ⁇ ⁇ ⁇ ⁇ is obtained for a particular user h (“ ⁇ ” denotes the number of examples ranked by the user h).
- generating conditioning input data and each version of the example data item may be processed to generate a corresponding vector of features for each version of the example data item, and a candidate vector of weights may be used to determine respective reward values for the first and second version.
- the values of the candidate vector of weights may be optimized to maximize a difference between the reward values, or equivalently to minimize a loss function that includes a term that measure the difference between the reward values of the first and second version. More specifically, a reward value of the first version may be determined from the vector of features for the first version of the example data item and the candidate vector of weights, and a respective reward value of the second version may be determined from the vector of features for the second version of the example data item and the candidate vector of weights.
- the weights of the candidate vector of weights may be optimized to maximize a difference between the reward value of the first version of the example data item and the reward value of the second version of the example data item (assuming the user indicated that the first version is preferred over the second version).
- the optimized vector of weights may then be determined as the obtained the vector of weights ⁇ ⁇ . More specifically, the vector of weights ⁇ ⁇ may be determined by maximizing a log likelihood function with respect to a new set of coefficients for ⁇ ⁇ , i.e. max ⁇ ⁇ log ⁇ ⁇ ⁇ , ⁇ , ⁇ , ⁇ , ⁇ ⁇ ⁇ Eq.
- the vector of weights ⁇ can be determined from a small number of ranked samples.
- determining a vector of weights ⁇ ⁇ for a new user may be expressed as a logistic regression problem since the parameters ⁇ of ⁇ ⁇ can be frozen. This process of determining the values of the vector of weights ⁇ ⁇ for a user may also be referred to as “adaptation”.
- the conditioning input ⁇ 116 is obtained based on input from a user (e.g. via a user interface) or from a software application in communication with the generative neural 38072935-1 network system 110.
- the feature neural network 114 processes the conditioning input ⁇ 116, and the generated data item ⁇ 118, to generate the vector of features ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ 124.
- the reward value 122 is generated based on the vector of features ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ 124 a nd the vector of weights ⁇ 120, e.g. based on ⁇ , ⁇ , ⁇ ⁇ ⁇ , ⁇ , ⁇ .
- the reward value 122 is used obtain the output data item 118. As noted above, this can be implemented in many ways, e.g. by performing steps 206 and 208 for a plurality of times to generate a plurality of candidate data items and corresponding reward values (for the same conditioning input), and using the reward value 122 to select one of the candidate data items as the output data item 118.
- FIG.3 shows a training system 300 for training the feature neural network 114 of Figure 1.
- the training system comprises training data 310 (comprising a plurality of training items ⁇ ⁇ ⁇ ⁇ , ⁇ , ⁇ ′ ⁇ , ⁇ , h ⁇ ⁇ ⁇ ⁇ ) and a training engine 320.
- Each training data item 312 conditioning input data ⁇ ⁇ 314, a plurality of different versions a training data item generated from the training conditioning input data ⁇ ⁇ using the data item generator neural network, a ranking ⁇ ⁇ of the respective different versions of the training data item obtained from a respective user, and user information h ⁇ specifying said user from a group of users (e.g. a group of human raters).
- the training system 300 trains the feature neural network 114 (i.e.
- a training item is processed. More specifically, the feature neural network 114 processes the training generating conditioning input data ⁇ ⁇ 314 and each version ⁇ ⁇ , ⁇ ⁇ ′ 316 of the corresponding training data item to generate the corresponding vectors of f eatures 324 (i.e. ⁇ ⁇ , ⁇ , and ⁇ ⁇ , ⁇ ′ ⁇ ). Further, training reward values 326 (i.e. ⁇ , ⁇ , and ⁇ ⁇ ⁇ , ⁇ ′ ⁇ ) are determined from the vectors of features 324 and from the candidate vector of weights 328 associated with the user specified in the respective user information 318.
- a plurality of learnable parameters, e.g. weights, of the feature neural network 114 and the weights of candidate vectors of weights 328 are updated by the training engine 320 based on the training reward values. This generally involves backpropagating gradients of an objective function to update the learnable parameters using any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm.
- An appropriate optimization objective for the training phase of feature neural network m ay derived as follows.
- a likelihood of ⁇ and ⁇ can be defined with respect to ⁇ : 38072935-1 L ⁇ , ⁇ ⁇ ⁇ , ⁇ , ⁇ ,h ⁇ ; ⁇ , ⁇ ⁇ ⁇ , ⁇ , ⁇ , ⁇ , ⁇ with ⁇ ⁇ , ⁇ , ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ , ⁇ ′ ⁇ if ⁇ ⁇ 1, and ⁇ ⁇ , ⁇ , ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇ ′ ⁇ ⁇ ⁇ ⁇ , ⁇ otherwise.
- the so-obtained matrix ⁇ can be used to initialize the above described candidate vector of weights for a “new” user when a ranking of the different versions of example data item is processed to determine the vector of weights.
- the candidate vector of weights may be initialized as the average over the rows of ⁇ which may be considered to correspond to initialising the model with the “average human preference”.
- the pre-trained large language model is fine-tuned using a reward model that neither distinguishes between raters nor performs adaptation. It was trained 38072935-1 using gradient ascent to solve max ⁇ log ⁇ ⁇ ⁇ , ⁇ , ⁇ , ⁇ ⁇ starting from the pre-trained ⁇ parameters of Gemma. is obtained by optimizing max ⁇ log ⁇ ⁇ ⁇ , ⁇ , ⁇ , ⁇ ⁇ to adapt the parameters of the non- ⁇ ⁇ ⁇ baseline to user h.
- the adaptive linear baseline is similar to the adaptive baseline, but layer is adapted.
- UltraFeedback data is not rater-annotated, it does come with the features used to compute preferences. Associated with each ( ⁇ ⁇ , ⁇ ⁇ ) and each ( ⁇ ⁇ , ⁇ ′ ⁇ ) in the dataset, we ⁇ have four features ⁇ ⁇ ⁇ ⁇ , ⁇ ⁇ ⁇ ⁇ indicating the level of helpfulness, honesty, instruction following, and ⁇ with respect to context ⁇ . These features were computed using OpenAI’s GPT-4. The preferences in the original dataset were defined based on a single feature ⁇ ⁇ selected per example. We used instead the more general notion of a linear c ombination of features.
- training raters sampled from four Gaussian distributions whose means are the “one-hot raters”: ⁇ ⁇ ⁇ 1,0,0,0 ⁇ , ⁇ ⁇ ⁇ 0,1,0,0 ⁇ , ⁇ ⁇ ⁇ 0,0,1,0 ⁇ , and ⁇ ⁇ ⁇ 0,0,0,1 ⁇ .
- the covariance of all four normal distributions was set to 0.3 ⁇ , where ⁇ is the 4x4 identify matrix. That is ⁇ ⁇ ⁇ , 0.3 ⁇ for ⁇ ⁇ 1, 2, 3, 4.
- preference vectors ⁇ are sampled from each distribution, resulting in a set ⁇ ⁇ ⁇ w ⁇ ith 120
- This way of generating preference vectors ⁇ ⁇ ⁇ results in raters h that mostly care about a single assess RFM’s performance using different numbers of reward features.
- the notation RFM(d) indicates that a d- dimensional feature vector ⁇ ⁇ R ⁇ was used.
- RFM did not have access to the real features ⁇ underlying the to derive ⁇ ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇ ′ ⁇ , ⁇ . Instead, RFM learned feature functions ⁇ by ⁇ ⁇ R ⁇ as described above.
- FIG. 4 shows the intra-user test accuracy obtained the baseline model and RFM throughout training (two independent training runs for each model). This corresponds to the fraction of examples ⁇ , ⁇ , ⁇ ′ ⁇ , h ⁇ in the test set for which the models can correctly predict the preference ⁇ ⁇ ⁇ 0,1 ⁇ .
- the reference numerals 400, 402, 404, and 406 respectively indicate experimental data for RFM(128), RFM(32), RFM(8) and the baseline. That is, it is an estimate of the models’ intra-user generalisation. It can be seen that RFM significantly outperforms the rater-agnostic baseline model.
- an initial dataset ⁇ ⁇ is defined by sampling 10 examples uniformly at random from ⁇ and labelling them using ( ⁇ ⁇ ⁇ ⁇ , ⁇ , ⁇ ⁇ ⁇ ⁇ , ⁇ ′ ⁇ , ⁇ ) with the corresponding ⁇ ⁇ .
- Eq. (1) is then solved using this data and its accuracy is assessed on the test set (also properly relabelled). This process is iterated until a dataset ⁇ ⁇ with 90 examples is obtained.
- the two feature functions ⁇ ⁇ computed during training i.e. Eq. (2)
- each ⁇ ⁇ ⁇ 8, 32, 128 ⁇ the aforementioned steps are repeated 5 times.
- the dashed line 500 indicates the baseline model
- the reference numerals 502, 504, 506, 508, 510 respectively indicate experimental data for RFM(8), RFM(32), RFM(128), the adaptive baseline, and adaptive linear baseline.
- the top row panels of FIG. 5 show the performance of the baselines and RFM in predicting the preferences of the held-out users.
- a second scenario used a similar protocol to the above-described first scenario, but the training raters were sampled from different distributions. In particular, the raters are sampled to represent a situation where the human raters disagree significantly.
- a mean vector for each possible instantiation of a 4-dimensional vector is defined with one element equal to 1, one element equal to ⁇ 1, and the remaining elements equal to zero, that is: ⁇ ⁇ ⁇ 1, ⁇ 1,0,0 ⁇ , ⁇ ⁇ ⁇ 1,0, ⁇ 1,0 ⁇ , ⁇ ⁇ ⁇ 1,0,0, ⁇ 1 ⁇ , ⁇ ⁇ ⁇ 1,1,0,0 ⁇ , ... , ⁇ ⁇ ⁇ 0,0, ⁇ 1,1 ⁇ .
- FIG. 5 show the experimental results for the second scenario. Because the baseline cannot distinguish between raters during training, opposing preferences become contradictory learning signals, making it difficult to capture any trends in the data. This explains why the non-adaptive baseline’s performance reduces to chance. It can be seen that all RFM versions outperform the comparative examples. Notably, RFM performs on held-out user h ⁇ , whose preferences are based on the first feature ⁇ ⁇ only. This suggests that the RFM can capture the reward features 38072935-1 underlying the raters’ preferences even when all the raters are based on combinations of such features (that is, even when the effect of features is not observed in isolation). [0103] FIG.
- FIG. 6 shows experimental results for the accuracy in predicting the preferences of held-out users on test set under the second scenario for the non-adaptive baseline, RFM, and adaptive baselines. More specifically, two highly capable models, Gemini 1.5 Pro and GPT- 4o, performing “in-context” adaptation are used to implement the adaptive baselines. To assess the LLMs’ prediction accuracy for held-out user h ⁇ , ⁇ ⁇ 10 training examples ranked by h ⁇ are provided together with the test example to be ranked.
- FIG.6 shows the results of the non- adaptive baseline 600, RFM 602, Gemini 1.5 Pro and GPT-4o 608 (for reference, the shot” performance of Gemini obtained with ⁇ ⁇ 0 is also shown in FIG. 6 as indicated by reference numeral 604).
- the bottom row of FIG.5 shows how well the baselines and RFM can predict the preferences of the held- out users using the new features ⁇ ′.
- RFM’s prediction accuracy lies between 80% and 90%, a significant improvement over the results shown in the middle row of FIG. 5 referring to the second scenario. This indicates that, when the features underlying the data can be computed with RFM’s architecture, training via Eq. (2) does indeed recover them.
- FIG. 7 shows results when best-of-n is applied with increasing “n”.
- Reference numerals 700, and 702 show respectively the win-rate of RFM and the non-adaptive baseline (reference numeral 704 indicates draws).
- the candidate responses used with best-of-n are all qualitatively similar (and thus not easy to distinguish), and they originate from a distribution that is different from the one used for training.
- FIG. 8 shows results for experiments in which publicly available reward models have been used as the raters h ⁇ .
- results are shown in FIG. 8. It can be seen that given enough (but still few) adaptation examples, RFM’s performance 800 either matches or significantly surpasses that of the non-adaptive baseline 802 (dashed line). This indicates that RFM is useful in real scenarios. For example, it can work as a form of “safety net” to make sure that minority preferences are also properly represented.
- Example hardware implementations 38072935-1 [0112]
- the generative neural network system e.g. a language model or a visual language model, is stored on a user computing device, i.e. a device local to the user, such as a mobile device e.g. a mobile phone, or a smart speaker.
- the input mechanism may comprise a system configured to input audio data characterizing a speech waveform of speech representing the input from the user in a natural language, and configured to convert the audio data into tokens representing the speech in the natural language, e.g. representing a transcription of the spoken input.
- the output mechanism may comprise a system configured to receive tokens representing the output for the user in the or another natural language and a system configured to convert the received tokens into audio data representing a waveform of speech representing the output to the user in the natural language, i.e. representing spoken words.
- the trained system can be deployed in an environment that enables a user to provide a request for the system, e.g.
- a users can provide the request, e.g., by way of a user interface or through an application programming interface (API).
- the request can be transmitted from a user device, e.g., over a data communications network such as the internet, to one or more computers implementing the system, e.g., in a data center.
- the system can generate a data item and then transmit the data item to a user device over a data communications network.
- the conditioning input may comprise a description and/or image of one 38072935-1 or more observations of the mechanical or computing system, e.g. of operation of the system, optionally obtained from one or more sensors sensing a condition or operation of the system.
- An image observation may be converted into a text description e.g. using an image captioning system or in other ways.
- the generated data item may comprise an image, audio, or text that identifies (described) a likely cause of the fault or undesired behavior. This may be used to repair the fault or correct the behavior.
- the reward model can define relatively more useful types of output for repairing the fault or correcting the behavior.
- the (trained) generative neural network system can be used for controlling a mechanical agent such as a robot or vehicle.
- the conditioning input may comprise a description of a task to be performed
- the generated data item may comprise a list of sub- tasks to be performed by the mechanical agent (trained to perform such sub-tasks), in order to perform the task.
- the reward model can define relatively more preferable or useful types of sub-task.
- Example multimodal applications [0120]
- the generative neural network system may comprise a multimodal machine learning system such as a visual language model (VLM). That is implementations of the generative neural network system can perform a multimodal task in which the conditioning input and data item, collectively, comprise data of multiple different types.
- VLM visual language model
- the generative neural network system can be trained on multiple natural and/or computer languages and the prompt may then specify a language to use.
- the tasks described below may be tasks that require 38072935-1 spatial awareness or other context from the image or video.
- a prompt may ask “What is the object in the top left corner?”.
- the system can have been trained or fine-tuned on examples of the input and output for the task.
- the system can have been trained using still or moving images containing one or more objects or actions, and corresponding sequences of text or other data e.g. describing or classifying the images.
- the computer language in the generated data item may comprise computer language for invoking a function or calling one or more external APIs.
- a data item may comprise data formatted as a JSON object.
- the conditioning input may define the task to be performed and may also include an image in relation to which the task is to be performed.
- the task can involves manipulation of particular types of data that may benefit from access to an API such as mathematical data, date/time related data, scientific data, recent data that may post-date training of the system (that may be accessed by a search function or API), and so forth; and the generated data item may comprise text in a computer language for performing the task.
- the method may then include using the text in the computer language to perform the task.
- the generated data item comprises text this may be converted to speech representing the text, and an audio (speech) output provided.
- the task comprises an agent control task in which the agent interacts with an environment to perform the agent control task.
- the conditioning input can include an observation characterizing the environment.
- the conditioning input can include a sequence of text that defines the task to be performed by the agent and the image can represent an observation of the environment, e.g. captured by a camera or other imaging device from a real-world environment.
- the generated data item can comprise an action selection output, e.g. including text, that is used to select one or more actions to be performed by the agent in the environment in response to the observation.
- the generated data item may define an action as text such as “A: 132114128525 156”, that can be converted into a control signal for a mechanical agent, such as a robot, e.g. 38072935-1 “ ⁇ ⁇ ⁇ 0.1, ⁇ 0.2,0 ⁇ ⁇ ⁇ ⁇ 10 ⁇ , 25 ⁇ , ⁇ 7 ⁇ ”.
- the action selection output may also or instead define one or more low-level skills, e.g. from a vocabulary of previously learnt skills.
- the sequence of text in the conditioning input to the system may describe the task to be performed, e.g. “What action should the robot take to [perform task]”.
- Examples of systems for controlling an agent that may be fine tuned as described herein can include PaLM-E (Driess et al. arXiv:2303.03378), RT-1 (Brohan et al. arXiv:2212.06817), and RT-2 (Brohan et al. arXiv:2307.15818).
- the environment is a real-world environment and the agent is a mechanical agent interacting with the real-world environment, e.g., a robot or an autonomous or semi-autonomous land, air, or sea vehicle operating in or navigating through the environment, and the actions are actions taken by the mechanical agent in the real- world environment to perform the task.
- the agent may be a robot or other mechanical agent interacting with the environment to accomplish a specific task, e.g., to locate or manipulate an object of interest in the environment or to move an object of interest to a specified location in the environment or to navigate to a specified destination in the environment.
- the observations may include, e.g., one or more of: images, object position data, and sensor data to capture observations as the agent interacts with the environment.
- the actions may define control signals to control the robot or other mechanical agent, e.g., positions, torques, or other control signals for the parts of the mechanical agent, or higher-level control commands.
- the agent may be a human agent and the environment may be a real-world environment.
- the agent can be a human user of a digital assistant such as a smart speaker, smart display, or some other device that is used to instruct the user to perform actions.
- the task may be any real-world task that the user wishes to perform.
- the observations may be obtained from an observation capture subsystem, e.g.
- a monitoring system such as a video camera or sound capture system, to capture visual observations of the user performing the task.
- the actions may comprise instructions in the form of, e.g., text, image, video, or audio data such as speech, that guide the user in performing the task.
- the term "configured” is used in relation to computing systems and environments, as well as computer program components.
- a computing system or environment is considered “configured” to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those 38072935-1 operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities.
- one or more computer programs are "configured” to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.
- the embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof.
- the subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by or to control the operation of a computing device or hardware.
- the storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these.
- the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware.
- implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.
- the term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose.
- Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
- a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed.
- GPGPU General-Purpose computing on Graphics Processing Units
- a computer program also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment.
- a program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments).
- a computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network.
- the specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.
- engine broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions.
- An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of AI and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results.
- GPUs graphics processing units
- TPUs tensor processing units
- other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks.
- a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both.
- ROM read-only memory
- RAM random access memory
- the essential elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data.
- the specific configuration of processing units and memory will depend on factors like the complexity of the AI model, the volume of data being processed, and the desired performance and latency requirements.
- Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities.
- the system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.
- Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), 38072935-1 and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs.
- ROM read-only memory
- SSDs solid-state drives
- HDDs hard disk drives
- optical discs such as CDs, DVDs, and Blu-ray discs.
- embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user.
- a display device such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display
- Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application.
- Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback.
- computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms.
- the selection of input and output modalities will depend on the specific application and the desired form of user interaction.
- Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.
- Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Mathematical Physics (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Molecular Biology (AREA)
- General Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Probability & Statistics with Applications (AREA)
- Mathematical Analysis (AREA)
- Computational Mathematics (AREA)
- Algebra (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
There is provided a method of generating a data item using a generative neural network system comprising a data item generator neural network configured to process a conditioning input to generate a data item, and a feature neural network subsystem comprising a feature neural network configured to process i) the conditioning input, and ii) the generated data item, to generate a vector of features, and to determine a reward value for the data item from the vector of features and from a vector of weights that characterizes one or more aspects of the generated data item. The method comprises obtaining the vector of weights and the conditioning input for the data item to be generated, and using the data item generator neural network to process the conditioning input to generate the data item based on a reward value determined from the vector of features and the vector of weights.
Description
GENERATING DATA ITEMS BASED ON A MULTIDIMENSIONAL REWARD MODEL BACKGROUND [0001] This specification relates to processing data using machine learning models. [0002] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. [0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output. SUMMARY [0004] A generative neural network system processes a conditioning input to generate a data item. For example, a generative neural network system may include a language model, a vision- language model or the like which processes an input query, comprising a sequence of tokens and/or other data, to generate output data which may represent a piece of text, an image, an audio waveform, a video or the like which is a sensible response to the input query, e.g. an answer to a question posed by the input query. [0005] This specification describes systems and methods implemented as computer programs on one or more computers in one or more locations that enable steering a pre-trained generative neural network system towards responses that exhibit one or more desirable aspects. This specification also describes the training of the generative neural network system. [0006] According to a first aspect there is provided a computer-implemented method of generating a data item using a (trained) generative neural network system. The generative neural network system comprises a data item generator neural network configured to process a conditioning input to generate a data item according to a data item generation policy. Generally, the data item represents a response to the conditioning input, where the conditioning input may be, e.g. a “prompt” for the data item generator neural network. The data item generator neural network may be pre-trained (e.g. on a large, unlabeled dataset by self-supervised learning). [0007] The generative neural network system further may comprise a feature neural network subsystem comprising a feature neural network configured to process i) the conditioning input, and ii) the generated data item, to generate a vector of features; and to determine a reward value 38072935-1
of the data item from the vector of features and a vector of weights that characterizes one or more aspects of the generated data item. The vector of weights may be associated with a user or a (small) group of users. In some implementations the group of users may comprise different users (e.g. users that share a preference or use case) and/or the same user at different times, places, contexts, or the like. [0008] The feature neural network may be configured to determine the reward value by determining an inner product of the vector of features and the vector of weights. For example, denoting the conditioning input ^^ (^^ ∈ ^^^ and the generated data item ^^ (^^ ∈ ^^^, the feature neural network may generate the vector of features ^^ఏ^^^,^^^ using a learned function ^^ఏ ∶ ^^ ൈ ^^ → ℝௗ (where ^^ denotes learnable parameters of the feature neural network and ^^ ^ 1). The vector of weights (denoted ^^^^, ^ may have the same dimensionality as the vector of features ^^ఏ^^^,^^^ (i.e. ^^ ∈ ℝௗ). In broad terms, each dimension ^^ఏ^^^,^^^ (denoted ^^ఏ,^) may be considered to represent a criterion used by a population of users (e.g. human users) to express their preferences with respect to the one or more aspects of the generated data items. The ^^-th element ^^^ of the vector of weights ^^ may be considered to represent how much the user (or group of users) associated with the vector of weights ^^ values (or does not value) criterion ^^ఏ,^. Advantageously the feature neural network enables deriving a reward function ^^ ^^^,^^^ 〈^^ఏ^^^,^^^,^^〉 (where 〈… 〉 denotes the inner product) that is specialized to a user (or group of users). Reward values generated using the user-specific reward function ^^ ^^^, ^^^ may be used to steer the pre-trained generative neural network system towards responses that are aligned with said user (as described in more detail below). Because the reward function ^^ ^^^,^^^ processes (in addition to the conditioning input ^^ and the generated data item ^^) the vector of weights ^^ having a plurality of elements (^^ ∈ ℝௗ with ^^ ^ 1), the described system may be considered to implement a “multidimensional
model”. [0009] The method comprises obtaining the vector of weights ^^ for a data item to be generated (e.g. from a user-annotated preference dataset, as described further below). The obtained vector of weights ^^ may characterize one or more aspects of the to-be-generated data item. For example, the vector of weights ^^ may be associated with a specific user (or group of users) and may specify the user’s preferences (the vector of weights ^^ associated with a user ℎ (or a group of users) is denoted ^^^^ hereafter). The method further comprises obtaining the conditioning input ^^ for the data item to be generated, e.g. from a user. [0010] The user providing the conditioning input ^^ may be the same or a different user than the user associated with the vector of weights ^^^^. 38072935-1
[0011] The method further comprises using the data item generator neural network to process the conditioning input ^^ to generate the data item ^^ based on a reward value 〈 ^^ఏ ^ ^^,^^ ^ ,^^^^ 〉 of the data item determined from the vector of features ^^ఏ ^^^,^^^ for the generated data item and from the vector of weights ^^ for the data item t
^^ o be generated (e.g. reward may be determined by determining an inner product of the vector of features ^^ఏ^^^,^^^ and the vector of weights ^^^^). Thus the described method derives reward values which can be tailored to the preferences of a specific user(s). It is an advantage of the described method that the reward values can be readily adapted to specific user(s), e.g. by simply providing the appropriate vector of weights ^^^^. Examples of how the vector of weights ^^^^ can be determined are described further below. In particular, this enables adapting the generator neural network system to preferences of “new” users, e.g. where the generator neural network system has not been trained on data annotated by these users. [0012] Many ways exist to generate the data item based on a reward value 〈^^ఏ^^^,^^^,^^^^〉. As one example, the reward value 〈^^ఏ^^^,^^^,^^^^〉 may be used to fine-tune (i.e. further train) the generative neural network system to implement an adapted data item generation policy (e.g. using known reinforcement techniques), and to process the conditioning input to generate the data item according to the adapted data item generation policy. In this case, the reward value 〈^^ఏ^^^, ^^^,^^^^〉 may be used to update a plurality of learnable parameters, e.g. weights, of the data item generator neural network based on the reward value. As another example, the reward value 〈^^ఏ^^^, ^^^,^^^^〉 may be used to generate the data item without resource-intensive (i.e. computationally and data intensive) re-training of the generative neural network system. This can be advantageous since it enables steering the output characteristic of the trained generative neural network system on computing devices with limited computational power and limited storage capacity, e.g. smart phones, tablet computer, laptops, and the like. [0013] One possibility to use the reward value 〈^^ఏ^^^,^^^,^^^^〉 to generate the data item without re-training of the generative neural network system is by implementing a “best-of-n” approach, i.e. given a conditioning input, “n” candidate responses are sampled and the candidate response with the highest reward is selected as output. More specifically, the conditioning input data ^^ may be by processed using the data item generator neural network for each of a plurality of data generation episodes to generate a plurality of different versions of the data item according to the data item generation policy. The conditioning input and each respective version of the generated data item may then be processed to generate the vector of features for each respective version of the generated data item. A respective reward value of each version of the data item 38072935-1
can be determined from an inner product of the vector of features for the version of the generated data item and from the vector of weights for the data item to be generated. The respective reward values of the versions of the generated data item may then be used to select one of the versions as the generated data item, e.g. a version of the generated data having the highest reward value may be selected as the generated data item. [0014] In some implementations, the generative neural network system may comprise a plurality of pre-trained data item generator neural networks each configured to process the conditioning input to generate the data item according to a different respective data item generation policy (the policies are denoted ^^^ ∈ Π). In this case, a “best-of-n” approach may be implemented with respect to samples generated by the plurality of pre-trained data item generator neural networks. [0015] In some implementations, the number of data item generator neural networks may be equal to the number of elements in the vector of weights ^^^^ (i.e. the number of data item generator neural networks may be “d” when ^^ ௗ ^^ ∈ ℝ with ^^ ^ 1). [0016] In some implementations, each policy ^^^ may be a solution from a reinforcement learning problem resulting from using ^^ఏ,^ reward function, but any suitable way of
selecting the policies ^^^ may be used. The data item generator neural networks may be pre- trained on the same training dataset. The method may further comprise processing the conditioning input data using each of the item generator neural networks to further generate at least one version of the data item according to each of the different respective data item generation policies ^^^. [0017] Using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item may comprise using the respective reward values of the versions of the generated data item to select one of the data item generator neural networks. For example, the policy with the largest expected performance under the obtained vector of weight ^^^^ is selected, i.e. a policy ^^ᇱ ∈ Π may be selected according to ^^ᇱ^^^^ ← arg m గ ∈a ஈx ^^^ ~గ 〈 ^^ఏ ^ ^^,^^ ^ ,^^^^ 〉 . Next, from the versions of the generated data item generated by the item generator neural network the version of the generated data item having the highest reward value may be selected as the generated data item, i.e. ^^ᇱᇱ^^^^ ← arg m ^ax ^〈^^ఏ ^^^,^^^^,^^^^〉 |^^^~^^ᇱ^^^^, ^^ ൌ 1, 2, … ,^^^. ^
is obtained from a user. However, a (human) user may not be able to clearly articulate their criteria and/or preferences in terms of 38072935-1
values of the vector of weights ^^^^. Thus, in some implementations, the user may rank different versions of the data item and the weights are determined from the values of the vector of weights ^^^^ are determined from the ranking. More specifically, in some implementations, obtaining the vector of weights may comprise processing an example generating conditioning input data using the data item generator neural network to generate a plurality of different versions of a corresponding example data item from the example generating conditioning input data, obtaining a ranking of the different versions of the example data item from a user, or a group of users, and determining the vector of weights from the ranking. [0019] As used herein a “user” includes a group of users, e.g. a plurality of different users, and can also refer to the same user in different circumstance, e.g. at different places, times, contexts, and so forth. [0020] In some implementations, the user may rank pairs of versions of the example data item (i.e. the user may indicate whether a first or a second version of the example data item is a preferred response given the example generating conditioning input data. [0021] In some implementations, the vector of weights ^^^^ may be determined from the ranking by processing the example generating conditioning input data and each version of the example data item to generate the vector of features for each version of the example data item, and by determining the vector of weights from the ranking and the vectors of features for the different versions of the example data item. [0022] More specifically, in the case when the user ranks pairs of versions of the example data item (i.e. a first and a second version), a candidate vector of weights may be used to determine respective reward values for the first and second version. The values of the candidate vector of weights may be optimized to maximize a difference between the reward values, or equivalently to minimize a loss function that includes a term that measure the difference between the reward values of the first and second version. [0023] More specifically, a reward value of the first version may be determined from the vector of features for the first version of the example data item and the candidate vector of weights, and a respective reward value of the second version may be determined from the vector of features for the second version of the example data item and the candidate vector of weights. [0024] The weights of the candidate vector of weights may be optimized to maximize a difference between the reward value of the first version of the example data item and the reward value of the second version of the example data item (assuming the user indicated that the first 38072935-1
version is preferred over the second version). The optimized vector of weights may then be determined as the obtained the vector of weights ^^^^. [0025] As an example, the values of vector of weights ^^^^ may be determined by minimizing the loss function ^^ ^^^,^^^ ൌ log^^^〈^^ఏ,^^〉^ where ^^ఏ ൌ ^^ఏ ^^^,^^^ - ^^ఏ ^^^,^^′^ when the user indicated that the first version y is preferred over the second version y’, and ^^ఏ ൌ ^^ఏ ^^^,^^′^ - ^^ఏ ^^^,^^^ when the user indicated that the second version y’ is preferred over the first version y. In this way, the vector of weights ^^ can be determined from a small number of ranked samples. [0026] In some implementations, the candidate vector of weights may be initialized with values corresponding to the “average population preference” (as described further below, these values are naturally determined during training of the feature neural network). Initialization with the “average population preference” may further reduce the number of ranked samples needed to obtain the vector of weights ^^^^. [0027] In some implementations the generative neural network system comprises a data item generator neural network that generates an output token sequence from an input token sequence including the conditioning input. The data item generator neural network may then be configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens, that is used to select an output token for the output token sequence. [0028] In some implementations of the generative neural network system, particularly when the generative neural network system comprises one or more Transformer-based models, can have more than a billion parameters and can require substantial computing resources, power, and time to process a network input to train the model. Sometimes such models can have can more than 10 billion or more than 100 billion parameters. When the generative neural network system is implemented on a digital assistant device, e.g., a mobile device, implemented in a computing system that includes a back end component, in particular a data server, in communication with the digital assistant device over a data communications network such as the Internet. There is then a need to optimize the computing load between the digital assistant device and the back end component. This need can be particularly acute with a large-scale language model because of its substantial memory and computing requirements compared with those typically found on a mobile device. [0029] The techniques described herein address these problems. In some implementations the described techniques facilitate a reduced a computational load, and improved load distribution, 38072935-1
e.g. when the generative neural network system is implemented in a multitasking and parallel processing computer system, distributed across multiple sites and interconnected by a data communication network. [0030] In some implementations the described techniques enable a beneficial distribution of computing load between a local, mobile computing device and a remote back-end server in a network. For example, the system may be implemented on a digital assistant device such as a mobile device. In such implementations the generative neural network system can be implemented the generative neural network system (wholly) on the mobile device. The mobile device generally has less working memory than the back-end data server, less computational capacity than the back-end data server, or both – which makes (re-)training of the generative neural network system on the local, mobile device undesirable if not unfeasible. Computational capacity can be measured in computing operations per second, e.g. FLOPS (floating point operations per second). The described techniques enable changing one or more aspects of the data items generated by the generative neural network system running on the mobile device without having to retrain, i.e. without the need for receiving and processing large amounts of training data. Thus, in this case, the remote server may pre-train the generative neural network system and send the values of a plurality of learned parameters defining the trained generative neural network system to the mobile device via the communication network. It is to be understood that in this case, the user of the mobile device may provide the conditioning input for the data item (by inputting the conditioning input into the mobile device). [0031] There is also described a computer-implemented method of training a feature neural network (e.g. the above described the feature neural network). In some implementations, the method may involve obtaining a plurality of training items ^^ା ≡ ^^^^^ ,^^^ ,^^′^ , ^^^ ,ℎ^^^ ^ ^ୀ^ , each training item comprising a respective training conditioning input data ^^^, a plurality of different versions ^^^ ,^^′^ of a corresponding training data item generated from the training conditioning input data using the data item generator neural network, a ranking ^^^ of the respective different versions of the training data item obtained from a respective user, and user information ℎ^ specifying said user from a group of users (e.g. a group of human raters), and training the feature neural network using the plurality of training items. [0032] In some implementations, the method may involve obtaining a plurality of training items ^^ା ≡ ^^^^^ , ^^^ ,^^′^ , ^^^ , ℎ^^^ ^ ^ୀ^ , each training item comprising a respective training conditioning input data ^^^, a plurality of different versions ^^^ ,^^′^ of a corresponding training data item generated from the training conditioning input data using the data item generator 38072935-1
neural network, a ranking ^^^ of the respective different versions of the training data item, and information ℎ^ specifying an origin of the ranking, and training the feature neural network using the plurality of training items. [0033] In general the feature neural network can have any suitable architecture and can include, e.g., one or more feed forward neural network layers, one or more recurrent neural network layers, one or more convolutional neural network layers, one or more attention neural network layers, or one or more normalization layers. Merely as an example the feature neural network may comprise a Transformer neural network that is configured to process the conditioning input (which may comprise an input sequence), and the generated data item (which may comprise an output sequence), to generate the vector of features. In general the feature neural network can have fewer learnable or learned parameters than the data item generator neural network. [0034] In some implementations, the feature neural network is trained in a plurality of training iterations. In each iteration, using the feature neural network processes, for each training item, the respective training generating conditioning input data ^^^ and each version ^^^ ,^^^′ of the corresponding training data item to generate the vector of features ^^ఏ ^^^, ^^^^, ^^ఏ ^^^,^^^′^ for each version of the corresponding training data item. Further, a respective training reward value ^^^ ^ ^^,^^ ^ , ^^^ ^ ^^,^^′ ^ of each version of the respective training data item may be determined from the vector of features for the respective version of the training data item and from a candidate vector of weights associated with the user specified in the respective user information. Further, a plurality of learnable parameters, e.g. weights, of the feature neural network and the weights of candidate vectors of weights may be updated based on the training reward values. This generally involves backpropagating gradients of an objective function to update the learnable parameters using any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm. [0035] For example, the plurality of learnable parameters ^^ of the feature neural network and the weights of candidate vectors of weights may be updated to minimize the loss function ^^ ఏ |ு|௫ ^శ ^^^,^^^ ൌ െ ∑^ log ^^^〈^^ ,^^〉^ where ^^ ∈ ℝ ௗ is the matrix formed by stacking all |^^| candidate vectors of weights ^^^^, and
^^ఏ ^^^,^^^ - ^^ఏ ^^^, ^^′^ if ^^ ൌ 1 (the rater indicates that version y is
over version y’), and ^^ఏ ൌ ^^ఏ ^^^,^^′^ - ^^ఏ ^^^,^^^ otherwise. Thus, this minimization results in updates for the plurality of learnable parameters of the feature neural network and the matrix ^^. Advantageously the so-obtained matrix ^^ can be used to initialize the above described candidate vector of weights when a ranking of the 38072935-1
different versions of example data item is processed to determine the vector of weights. For example, the candidate vector of weights may be initialized as the average over the rows of ^^ which may be considered to correspond to initialising the model with the “average human preference”. [0036] According to another aspect, there is described a computer-implemented method of training a data item generator neural network as described herein. [0037] The method involves obtaining dataset of training examples, each training example comprising a training conditioning input and a corresponding training data item. In some implementations the corresponding training data item is obtained by processing the training conditioning input using the data item generator neural network. [0038] For each of a plurality of the training data items the method obtains a vector of weights that characterizes one or more aspects of data items to be generated, processes the training conditioning input and the training data item using a feature neural network subsystem, e.g. as described herein, to generate a vector of features, and determining a reward value for the data item from the vector of features and the vector of weights. [0039] The method trains the data item generator neural network using the reward values, using any suitable objective function, e.g. a maximum likelihood objective, a cross-entropy objective, and so forth. In general the training can involve backpropagating gradients of the objective to update learnable parameters of the data item generator neural network. [0040] According to another aspect, there is provided a system that includes one or more computers and one or more storage devices communicatively coupled to the one or more computers and storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the previously described method. [0041] According to another aspect, there is provided one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the previously described method. [0042] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. [0043] Implementations of the described techniques enable adapting the data item generation of the generative neural network system in a convenient manner so as to exhibit one or more desired aspects. In some implementations this is achieved without the need of re-training the generative neural network system. This makes the described techniques particularly suitable for implementations in which the generative neural network system is pre-trained on a remote 38072935-1
server and then implemented on a mobile device which lacks the computational power to fully re-train the system. In this case the mobile device can use the described techniques to locally adapt the data item generation of the generative neural network system without having to receive and process large amounts of training data (reducing network traffic and computation cost). [0044] Further, in some implementations, the vector of weights that characterizes the one or more desired aspects of the data items can conveniently be found by processing only a few user-annotated samples. In implementations only a few annotated samples are necessary because, in broad terms, the trained feature neural network has learned properties of the distribution of preferences in the population so as to effectively create a relatively low dimensional representation of the preferences. [0045] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS [0046] FIG.1 shows an example system for generating data items and associated rewards. [0047] FIG.2 is a flow diagram of an example process for generating data items and associated rewards generating data items and associated rewards. [0048] FIG.3 shows an example training system for a feature neural network. [0049] FIGS.1 to 8 show experimental results generated by the described techniques. [0050] Like reference numbers and designations in the various drawings indicate like elements. DESCRIPTION [0051] Generative neural networks, such as language model neural networks (LM), are neural networks which process an input query, comprising a sequence of tokens and/or other data such as a media element, to generate an output representing a sensible response to the input query, e.g. an answer to a question posed by the input query. Sometimes, a LM is first pre-trained on a large, unlabelled (text) dataset (e.g. Web data) and then fine-tuned for downstream tasks. The pre-training may be performed by self-supervised learning (also known as predictive learning) in which the LM learns to correctly predict continuations of received samples of a text database, such as a large, publically-available database of natural language. 38072935-1
[0052] Fine-tuning is performed because self-supervised LMs often exhibit factual errors, biases, and other undesirable behavior. Thus, the pre-trained LM may be fine-tuned to align with human expectations and values. Fine-tuning can be performed with a human-annotated preference dataset, e.g. via reinforcement learning from human feedback (RLHF) which performs alignment by first learning a scalar-valued reward model, that mimics human judgment, and then employs reinforcement learning to optimize the LM against this reward. [0053] Sometimes RLHF models preferences using a reward model that does not distinguish between people. For example, human feedback may be collected by asking humans to rank examples of a LM’s behavior/output, and then all of the feedback is integrated to derive a single reward function that reflects the preferences of the population of interest. However, this approach is not effective when there is considerable disagreement across the population (of human raters). This is likely the case in the training of LMs and other generative neural networks. [0054] As an illustration, given a pair of alternative responses to a subjective question, 51% of the target audience may prefer the first option while the remaining 49% may prefer the second. Without distinguishing between users, there are two options: either the preferred answer is picked and 49% of the users are left unhappy 100% of the time, or the answers are sampled proportionally to how often they are preferred which leaves 100% of the users unhappy approximately half of the time. Both options are unsatisfactory. It is therefore desirable to have reward models that can be specialized to individual users or groups of users. However, it can be hard to capture subjective criteria that underlie human preferences since humans often cannot fully articulate the reasons why they prefer one behavior over the other. [0055] For the sake of clarity, terminology of LMs is adopted in the following although it is to be understood that the proposed techniques are also applicable to other types of generative neural networks. More specially, ^^ denotes a (finite) set of symbols (e.g. tokens), ^^ ≡ ⋃ ^ ^^ ୀ^ ^^^ denotes the “context space” where ^^ ^^ ௫ is the maximum context length, and ^^ ≡ ⋃ ^ ^ୀ^ ^^ denotes the “response space”. Π ≡ ^^^|^^:^^ → ∆^^^^^ denotes a “policy space” where ∆^^^^ is a set of (discrete) probability distributions over ^^. A policy ^^ ∈ Π can be considered to represent a distribution over responses ^^ conditioned on context ^^ that is encoded in a LM. [0056] A user-annotated preference dataset may be generated as follows. A human rater is sampled, ℎ~^^ℋ where ^^ℋ ∈ ∆^ℋ^ is a distribution over a set of users ℋ. Then, a context is sampled, ^^ ∈ ^^^ ^ ^ ,
^^ ^ ^^ ∈ ∆^^^^ is a distribution over ^^ that may depend on ℎ. Two
38072935-1
responses are then sampled ^^,^^′~^^^∙ |^^^, where ^^ ∈ Π is an appropriately defined exploration policy (possibly implemented by an LM). The two sampled responses ^^ and ^^′ are ranked by the human rater ℎ, who ranks them by sampling ^^~^^^^^,^^,^^ᇱ,ℎ^, where ^^ ∈ ∆^^0,1^^ is a Bernoulli distribution whose mean is ^^^^^ ≻ ^^ᇱ|^^, ℎ^, i.e. the probability of response ^^ to be preferred over ^^′ given the context ^^ (the notation “^^ ≻ ^^ᇱ” is used to indicate that response ^^ to be preferred over ^^′). This procedure results in a dataset ^^ା ≡ ^^^^^ ,^^^ ,^^′^ , ^^^ ,ℎ^^^ ^ ^ୀ^ , where ℎ^ is an index indicating which rater defined ^^^ based on ^^^ ,^^^ , and ^^′^. [0057] The objective in the reward personalisation problem is not to model the dataset ^^ା itself, but rather to use the data to generalise over the sets of interest: ℋ, ^^, and ^^. Denoting ℋ^ ⊆ ℋ the set of human raters ℎ^ who provided feedback for the ା
of ^^ , two types of generalisations can be distinguished: intra-user generalisation and inter-user generalisation. In intra-user generalisation ^^ା is used to model ^^^^^ ≻ ^^ᇱ|^^, ℎ^^ over ^^ ൈ ^^ ൈ ℋ^ only. This may enable predicting preferences over contexts ^^ and responses ^^ beyond the training data but not extrapolating to users not in the set of raters ℋ^ . In inter-user generalisation the same data is used to model ^^^^^ ≻ ^^ᇱ|^^,ℎ^ over the entire set ^^ ൈ ^^ ൈ ℋ. This may enable predicting the preferences of any user in ℋ over the entire sets ^^ and ^^. At least some of the embodiments described below enable inter-user generalisation. [0058] Generalization across all users in ℋ may require a distinct reward function ^^^^^^, ^^^ per user ℎ ∈ ℋ. One way to accomplish this is by employing a separate parametric function ^^ఏ^ per user, but it may be hard to obtain inter-user generalization in this way. Another way is to use two disjoint sets of parameters: i) a set of common parameters ^^ ∈ ℝ^^ that is shared among all users ℎ ∈ ℋ, including those ℎ ∉ ℋ^ , and ii) parameters ^^ ^^ ^ ∈ ℝ that are specific to user ℎ. As further described below with reference to FIG.3, this
of the parameters may induce a division of the training procedure, i.e. the data in ^^ା can be used to learn ^^ and ^^^^ for a rater ℎ^ ∈ ℋ^ , and additional data ^^^ ≡ ^^^^^ , ^^^ ,^^′^ , ^^^^^^ ᇱ ^ୀ^ (with ^^^~^^൫^^^ ,^^^ ,^^ ^ , ℎ൯) cont feedback provided by a
can be used to learn the parameters ^^^. [0059] In some implementations, the number of shared parameters may be much larger than the number of parameters that are specific to a given individual (i.e. ^^ ≫ ^^). This may be advantageous to achieve quick adaptation of the model to a specific user ℎ ∈ ℋ െ ℋ^ . Further, having a large number of shared parameters has a further practical
that this allows for a distributed learning architecture in which the learning of ^^ can make use of a centralised, 38072935-1
powerful computational infrastructure, while each ^^^ can be learned “locally” using less resources. [0060] FIG. 1 shows an example computer system 100. The computer system 100 is an example of a system, implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below are implemented. In broad terms, the computer system 100 is configured to generate data items aligned with preferences of a specific user or a group of users. This involves generating reward values according to a reward function/model that can be adapted to the specific individual user (or group of users), including people outside of the group of raters who provided feedback for training. To this end, the computer system 100 employs the above-described disjoint sets of parameters ^^ ∈ ℝ^^ (shared among all users) and user-specific parameters ^^^ ∈ ℝ ^^ (specific to user or a group of users) to implement a user-specific reward function. Reward values generated by the computer system 100 are then used to steer/modulate the output of a generative neural network. [0061] More specifically, the computer system 100 comprises a generative neural network system 110 comprising a data item generator neural network 112 and a feature neural network 114. [0062] The data item generator neural network 112 configured to process a conditioning input 116 (denoted ^^) to generate a data item 118 (denoted ^^) according to a data item generation policy. The conditioning input 116 defines a query or task for the data item generator neural network 112, and may be obtained from a user (e.g. via user interface), from another software application, or the like. The conditioning input 116 can include one or more modalities including text, image, audio, video, or a combination of one or more such modalities. In general, the network input can include multiple modalities depending on the specific architecture of the data item generator neural network 112. The data item 118 represents a response to the conditioning input 116. Similar to the conditioning input 116, the data item 118 may comprise text, image and/or audio tokens. [0063] The generative neural network system 110 is further configured to obtain (e.g. as further input) a vector of weights 120, and to process the conditioning input 116, data item 118 and the vector of weights 120 to generate a reward value 122. The vector of weights 120 is specific to a user and denoted ^^^^ ∈ ℝ ^^ for user ℎ. [0064] In broad
^^^^^^,^^^ for user ℎ may be expressed as ^^^^^^, ^^^ ൌ 〈^^^^^,^^^,^^^^〉 where ^^^^^,^^^ is a function that maps a conditioning input ^^ and a data item ^^ 38072935-1
to a vector of “reward features” ^^^^^,^^^ ∶ ^^ ൈ ^^ → ℝௗ, and 〈∙,∙〉 denotes inner product. A personalized version of the Bradley-Terry model may then be expressed as: ^^^^^ ≻ ^^ᇱ|^^, ℎ^ ൌ ^^൫^^^^^^,^^^ െ
, where
to implement a parametrized form of the reward ^^^^^^,^^^ which can be defined as ^^ఏ,௪^^^^,^^^ ൌ 〈^^ ^^^, ^^^ ,^^ 〉 with the parametrized function ^^ ^^^,^^^ ∶ ^^ ൈ ^^ → ℝௗ ௗ ఏ ^^ ఏ with ^^ ∈ ℝ and the user specific vector ^^^^ ∈ ℝௗ. More specifically, the feature neural network 114 is configured to process i) the conditioning input ^^ 116, and ii) the generated data item ^^ 118, to generate a vector of features ^^ఏ ^^^,^^^ 124. The generative neural network system 110 is further configured to generate the reward value 122 based on the vector of features ^^ఏ^^^,^^^ 124 and the vector of weights ^^^^ 120, e.g. based on ^^ఏ,௪^^^^,^^^ ൌ 〈^^ఏ^^^, ^^^,^^^^〉. [0066] The feature neural network can have any suitable architecture and can include, e.g., one or more feed forward neural network layers, one or more recurrent neural network layers, one or more convolutional neural network layers, one or more attention neural network layers, or one or more normalization layers. Merely as an example the feature neural network may comprise a Transformer neural network that is configured to process the conditioning input (which may comprise an input sequence), and the generated data item (which may comprise an output sequence), to generate the vector of features. In general the feature neural network can have fewer learnable or learned parameters than the data item generator neural network. [0067] The so-obtained the reward value 122 is used to modulate the output of the data item generator neural network 112 so as to align the output with the preferences of the user ℎ with respect to one or more aspects of the generated data item. This can be achieved in many different ways. In the example of FIG. 1, the generative neural network system 110 is configured to adopt ^^ఏ,௪^ ^^^,^^^ as a criterion to select responses generated by the data item generator neural network 112. More specifically, the generative neural network system 110 is configured to implement a “best-of-n” approach, i.e. given the the conditioning input ^^ 116, “n” candidate generated data items ^^^ ∼ π^x^ (where π^x^ denotes the data item generation policy of the data item generation neural network 112) are generated and the generative neural 38072935-1
network system 110 selects, as out data item 118, the data item ^^^ associated with the highest reward value, i.e. based on arg max௬^ ^^ఏ,௪^^^^,^^^^. [0068] In other implementations, the vector of weights ^^^^ may be used to directly modulate the output of the data item generator neural network. In general, several methods exist that are designed to synthesise a policy that performs well under a linear combination of features whose coefficients are only provided at deployment time. These methods may be employed for obtaining an adapted data generation policy based on the vector of weights ^^^^. For example, the vector of weights ^^^^ may be processed by one of these methods to (immediately) obtain a policy that is specialised to the corresponding reward function. More specifically, the generative neural network system may comprise a plurality of pre-trained data item generator neural networks each configured to process the conditioning input to generate the data item according to a different respective data item generation policy (the policies are denoted ^^^ ∈ Π). Each data item generator neural network may be trained induced by different linear combinations of reward features ^^ఏ, and the outputs of the data item generator neural networks may be combined based on the vector of weights ^^^^. [0069] Further in these implementations, the number of data item generator neural networks may be equal to the number of elements in the vector of weights ^^^^ (i.e. the number of data item generator neural networks may be “d” when ^^ ௗ ^^ ∈ ℝ with ^^ ^ 1). Each policy ^^^ may be a solution from a reinforcement learning
from using ^^ఏ,^ as the reward function, but any suitable way of selecting the policies ^^^ may be used. The data item generator neural networks may be pre-trained on the same training dataset. The method may further comprise processing the conditioning input data using each of the item generator neural networks to further generate at least one version of the data item according to each of the different respective data item generation policies ^^^. [0070] In one possibility, the plurality of pre-trained data item generator neural networks may be used to implement a form of generalised policy evaluation termed “successor features” in the field of “generalised policy improvement” (GPI), for example as described in A. Barreto et al., “Successor features for transfer in reinforcement learning”, Advances in Neural Information Processing Systems (NIPS), pages 4055–4065. Curran Associates, Inc., 2017. [0071] As an example the method may comprise computing successor features ^^ఎഏ ൌ of the policies ^^ ∈ Π as ^^ ^^^ ^ ଶ ఎഏ ,^^ ൌ ^^గ^^^ఏ,௧ା^ ^ ^^^^ఏ,௧ାଶ ^ ^^ ^^ఏ,௧ାଷ | ^^௧ ൌ ^^,^^௧ ൌ ^^, ൧, where ^^ ∈
38072935-1
^0, 1^ is the discount factor. Since ^^ఎഏ is a d-dimensional value function, ^^గ can be computed using reinforcement learning techniques, e.g. temporal difference learning. An action-value function of ^^ under the corresponding reward ^^ఏ,௪^^^^,^^^ ൌ 〈^^ఏ^^^,^^^,^^^^〉 can computed as ^^గ^^^, ^^^ ൌ 〈^^ఎഏ^^^, ^^^,^^^^〉. Thus, the method may further comprise computing a policy for user as GPI as ^^ᇱ^^^^ ← arg m ௬ ∈a ^x^m గ ∈a ஈx〈 ^^గ ^^^,^^^ 〉 . This illustrates how to quickly
compute a policy for an user represented by its preference vector ^^^^. [0072] In another possibility, using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item may comprise using the respective reward values of the versions of the generated data item to select one of the data item generator neural networks. For example, the policy with the largest expected performance under the obtained vector of weight ^^ ᇱ ^^ is selected, i.e. a policy ^^ ∈ Π may be selected according to ^^ᇱ^^^^ ← arg max ^^^ ~గ〈^^ఏ ^^^,^^^,^^^ 〉. Next, from the versions of the generated గ ∈ ஈ ^ data item generated by the
item generator neural network the version of the generated data item having the highest reward value may be selected as the generated data item, i.e. ^^ᇱᇱ^^^^ ← arg m ^ax ^〈^^ ^^^,^^^,^^ 〉 |^^~^^ᇱ^^^^, ^^ ൌ 1, 2, … ,^^^. ^ ఏ ^ ^^ ^ [0073] In some
system comprises a data item generator neural network that generates an output token sequence from an input token sequence including the conditioning input. The data item generator neural network may then be configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens, that is used to select an output token for the output token sequence. [0074] In some implementations the tokens can represent text, e.g., words, wordpieces or characters, in a natural or computer language. For example, text may be received, e.g., as a series of encoded characters, e.g. UTF-8 encoded characters; such “characters” can include Chinese and other similar characters, as well as logograms, syllabograms and the like. A text encoder, i.e. a tokenizer, can process a sequence of text to represent the text as a series of text tokens from a vocabulary of text tokens, e.g. that each represent words, wordpieces or characters in a natural or computer language. The computer language may be any formal language used to communicate with a computer, e.g. a markup language, or a command or configuration language, or a data exchange language such as JSON, or a programming language. The tokenizer can, e.g., implement BPE (Byte Pair Encoding) or Wordpiece 38072935-1
tokenization. Optionally the text can be obtained from audio data representing speech; the output tokens may be converted into audio data that represent speech corresponding to the text. [0075] Also or instead the tokens may represent an image. For example a set (sequence) of input or output tokens can represent an image. Each image token may comprise a block encoding of values of the pixels in a different region of an image that maps a set of values of the pixels to a respective image token. The block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network. [0076] Also or instead the tokens may represent an audio waveform. For example a set (sequence) of input or output tokens can represent audio data representing an waveform e.g. instantaneous audio amplitude values or time-frequency audio data. Each image token may comprise a block encoding of the audio waveform in a different time segment of the audio that maps a set of values representing the audio waveform to a respective image token. The block encoder may comprise a neural network, e.g. having one or more (self-)attention layers, such as a Transformer neural network. In a multimodal system audio data or an image may be flagged by a start-of-audio token or start-of-image token. [0077] In some implementations the generative neural network system can also or instead comprise a data item generator neural network that is a diffusion model neural network. In general a diffusion model neural network can be a neural network that has been trained to process a diffusion input comprising a current noisy data item and data specifying a current time to generate a diffusion output that defines an estimate (given the current time) of either a noise component of the current noisy data item, i.e. an estimate of the noise that has been added to an original data item to generate the current noisy data item; or of a de-noised version of the current noisy data item. [0078] That is, in some implementations the generative neural network system can be a multimodal system that is configured to process a conditioning input comprising one or more of text data, audio data defining an audio signal (e.g. as amplitude values of the audio signal or as a time-frequency representation of the audio signal), or a still or moving image (e.g. as image pixel values), to generate a data item that can similarly comprise text data, audio data, or a still or moving image. [0079] For example the conditioning input may comprise text and the data item may comprise an image or an audio signal that represents speech an image generated in response to the text, e.g. described by the text. Also or instead the conditioning input may comprise an audio signal that represents speech, or an image, and the data item may comprise text, e.g. that describes 38072935-1
the conditioning input. In another example, the conditioning input may comprise one or more of text, audio, video, or image data, and the data item may comprise one or more of text, audio, video or image data. [0080] As another example the conditioning input may comprise an observation, e.g. of a real world environment, e.g. from sensor such as a camera or other image sensor; and optionally additional information such as information defining a particular task to be deformed. The output data item may comprise agent control data that defines one or more actions to be performed by an agent, e.g. by a mechanical agent such as a robot or autonomous vehicle, to perform a task. The reward model(s) may, e.g., define a preferred trajectory of motion of the mechanical agent in the (real-world) environment. [0081] In some implementations the generative neural network system may comprise a language and/or image generation neural network system, that may have been trained before being fine-tuned by the above described method. The conditioning input may comprise a prompt, e.g. a natural or computer language prompt for the generative neural network system. The generated data item may comprise a natural or computer language and/or image response to the prompt. [0082] In general the generative neural network system can have any appropriate architecture for processing the conditioning input to generate the data item. [0083] As one example, the generative neural network system may comprise an auto- regressive generative model (e.g., a Transformer, a recurrent neural network, etc.) that can auto-regressively generate an output sequence as the data item based on the conditioning input. The generative model can, for example, comprise a large language model (LLM) that can auto- regressively generate tokenized representations of text data, a vision-language model (VLM) that can auto-regressively generate tokenized representations of image or video data, e.g. in response to a text conditioning input or that can auto-regressively generate tokenized representations of text, e.g. in response to an image conditioning input, an audio language model that can auto-regressively generate tokenized representations of text data, or a multimodal model that can that can generate tokens representing any of text, image or audio, e.g. in response to a conditioning input comprising any of text, image or audio, and so forth. [0084] As another example, the generative neural network system may comprise a diffusion model (e.g., a denoising diffusion model, a score-based diffusion model, a latent diffusion model, etc.) that can generate the data item by repeatedly transforming samples from a noise distribution (e.g., a Gaussian distribution) based on the conditioning input over a sequence of 38072935-1
iterations. For example, the generative neural network system may comprise a diffusion model that transforms samples from the noise distribution using a denoising neural network with any appropriate architecture (e.g., a convolutional neural network, a recurrent neural network, etc.). Such a diffusion model may be used to generate, e.g., a still or moving (video) image. [0085] As another example, the generative neural network system may comprise a neural network that can generate the data item by transforming samples from a noise distribution (e.g., a Gaussian distribution). The generative neural network system may comprise, e.g., a generator network of a generative adversarial network, a decoder of a variational auto-encoder, a normalizing flow, and so on. [0086] As used herein an image may be any still or moving image, i.e. the image may be part of a video, in 2D or 3D, and may be a monochrome, color or hyperspectral image, i.e. comprising monochrome or color pixels. As defined herein an “image” includes a point cloud e.g. from a LIDAR system, and a “pixel” includes a point of the point cloud. An image may have been captured by a camera or other image sensor from the real world; and objects in the image may comprise physical objects, represented by the image. [0087] FIG. 2 is a flow diagram of an example process 200 of using a generative neural network system to generate a data item. The process 200 of FIG. 2 may be implemented by one or more computers in one or more locations. For example, the process 200 may be implemented by the system of FIG. 1, and for convenience the process is described with reference to FIG.1. [0088] At step 202, the vector of weights 120 is obtained from a user. As described below with reference to FIG. 3, for a user that acted as a rater for the training dataset used to train the feature neural network 114, the corresponding the vector of weights ^^^^ can be readily retrieved. A “new” or “unseen” user, i.e. a user that did not acted as a rater for the training dataset used to train the feature neural network 114, may rank different versions of data items and the values of the vector of weights ^^^^ are determined from the ranking. To this end, an example generating conditioning input data may be processed using the data item generator neural network to generate a plurality of different versions of a corresponding example data item from the example generating conditioning input data, and a ranking of the different versions of the example data item may be obtained from the user, or a group of users. The user may rank pairs of versions of the example data item (i.e. the user may indicate whether a first or a second version of the example data item is a preferred response given the example generating conditioning input data. In this way, the aforementioned dataset ^^^ ≡ 38072935-1
^^^^^ ,^^^ ,^^′^ , ^^^^^ ^ ^ୀ^ is obtained for a particular user ℎ (“^^” denotes the number of examples ranked by the user ℎ).
generating conditioning input data and each version of the example data item may be processed to generate a corresponding vector of features for each version of the example data item, and a candidate vector of weights may be used to determine respective reward values for the first and second version. The values of the candidate vector of weights may be optimized to maximize a difference between the reward values, or equivalently to minimize a loss function that includes a term that measure the difference between the reward values of the first and second version. More specifically, a reward value of the first version may be determined from the vector of features for the first version of the example data item and the candidate vector of weights, and a respective reward value of the second version may be determined from the vector of features for the second version of the example data item and the candidate vector of weights. [0090] The weights of the candidate vector of weights may be optimized to maximize a difference between the reward value of the first version of the example data item and the reward value of the second version of the example data item (assuming the user indicated that the first version is preferred over the second version). The optimized vector of weights may then be determined as the obtained the vector of weights ^^^^. More specifically, the vector of weights ^^^^ may be determined by maximizing a log likelihood function with respect to a new set of coefficients for ^^^^, i.e. max ^^ ∑^∈^^ log ^^^〈^^௫ ఏ ^,௬^,௬ᇱ^,௭^ ,^^ 〉 ^ Eq. (1) where ^^ ఏ ௫^,௬^,௬ᇱ^,௭^ ൌ ^^ఏ ^^^^ ,^^^^ - ^^ఏ ^^^^
that the first version y is
y’ (i.e. ^^^ ൌ 1), and ^^ ఏ ௫^,௬^,௬ᇱ^,௭^ ൌ ^^ఏ ^^^^ ,^^′^^ - ^^ఏ ^^^^ ,^^^^ when the user indicated that the second version y’ is
way, the vector of weights ^^ can be determined from a small number of ranked samples. Thus, determining a vector of weights ^^^^ for a new user may be expressed as a logistic regression problem since the parameters θ of ^^ఏ can be frozen. This process of determining the values of the vector of weights ^^^^ for a
user may also be referred to as “adaptation”. Notably, a reward value that is linear in ^^^^ may give rise to a convex “adaptation” problem enabling reliable adaptation to new users even in a low-data regime. [0091] At step 204, the conditioning input ^^ 116 is obtained based on input from a user (e.g. via a user interface) or from a software application in communication with the generative neural 38072935-1
network system 110. At step 206, the feature neural network 114 processes the conditioning input ^^ 116, and the generated data item ^^ 118, to generate the vector of features ^^ఏ ^^^,^^^ 124. At step 206, the reward value 122 is generated based on the vector of features ^^ఏ ^^^,^^^ 124 and the vector of weights ^^^^ 120, e.g. based on ^^ఏ,௪^^^^,^^^ ൌ 〈^^ఏ^^^,^^^,^^^^〉. At step 210, the reward value 122 is used obtain the output data item 118. As noted above, this can be implemented in many ways, e.g. by performing steps 206 and 208 for a plurality of times to generate a plurality of candidate data items and corresponding reward values (for the same conditioning input), and using the reward value 122 to select one of the candidate data items as the output data item 118. [0092] FIG.3 shows a training system 300 for training the feature neural network 114 of Figure 1. The training system comprises training data 310 (comprising a plurality of training items ^^ା ≡ ^^^^^ ,^^^ ,^^′^ , ^^^ , ℎ^^^ ^ ^ୀ^ ) and a training engine 320. Each training data item 312 conditioning input data ^^^ 314, a plurality of different versions
a training data item generated from the training conditioning input data ^^^ using the data item generator neural network, a ranking ^^^ of the respective different versions of the training data item obtained from a respective user, and user information ℎ^ specifying said user from a group of users (e.g. a group of human raters). The training system 300 trains the feature neural network 114 (i.e. adjusts the learnable parameters ^^ of the feature neural network 114 over a plurality of training iterations. [0093] In each iteration, a training item is processed. More specifically, the feature neural network 114 processes the training generating conditioning input data ^^^ 314 and each version ^^^ ,^^^′ 316 of the corresponding training data item to generate the corresponding vectors of features 324 (i.e. ^^ఏ ^^^, ^^^^, and ^^ఏ ^^^,^^^′^). Further, training reward values 326 (i.e. ^^^^^^,^^^, and ^^^^^^,^^′^) are determined from the vectors of features 324 and from the candidate vector of weights 328 associated with the user specified in the respective user information 318. Further, a plurality of learnable parameters, e.g. weights, of the feature neural network 114 and the weights of candidate vectors of weights 328 are updated by the training engine 320 based on the training reward values. This generally involves backpropagating gradients of an objective function to update the learnable parameters using any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm. [0094] An appropriate optimization objective for the training phase of feature neural network may derived as follows. Denoting ^^ ∈ ℝ|ு|௫ ௗ the matrix formed by stacking all |^^| candidate vectors of weights ^^^^, a likelihood of ^^ and ^^ can be defined with respect to ^^: 38072935-1
ℒ^^^,^^^ ൌ ^^൫^^ห^^,^^,^^ᇱ,ℎ^;^^,^^൯ ൌ 〈^^௫ఏ ,௬,௬ᇱ,௭ ,^^^^〉 with ^^௫ ఏ ,௬,௬ᇱ,௭ ൌ ^^ఏ ^^^,^^^ െ ^^ఏ ^^^, ^^′^ if ^^^ ൌ 1, and ^^ఏ ௫,௬,௬ᇱ,௭ ൌ ^^ఏ ^^^,^^′^ െ ^^ఏ ^^^,^^^ otherwise. Assuming the training data items ^^ା ≡ ^^^^^ ,^^^ ,^^′^ , ^^^ , ℎ^^^ ^ ^ୀ^ are i.i.d, one can obtain: ℒ^^^, ^^|^^ା^ ൌ ∏^∈^శ ^^^〈^^௫ ఏ ^,௬^,௬ᇱ^,௭^ ,^^^^^〉^ . The problem of learning of ^^ and ^^ can be formalized as a maximum-likelihood estimation: max ∑^∈^శ log^^^〈^^௫ ఏ ^,௬^,௬ᇱ^,௭ ,^^^^ 〉^ Eq. (2). ఏ,^^ ^ ^ [0095] Thus, this of learnable parameters of the feature neural
used to obtain the vector weights for any user ℎ^ that rated examples in the training dataset ^^ା (intra-user generalisation). This because ℎ ^ ^ can be used as an index to select the appropriate row of ^^ used to compute the reward of that particular user. Each dimension of ^^ఏ, ^^ఏ,^ can be interpreted as a
criterion humans to express their preferences. The i-th element of ^^^^ , ^^^^,^ represents how much human rater ℎ^ values (or does not value) criterion ^^ .
the features ^^ఏ does not depend on raters being able to articulate the criteria underlying their preferences; ranked pairs of candidate responses are sufficient. [0096] Advantageously the so-obtained matrix ^^ can be used to initialize the above described candidate vector of weights for a “new” user when a ranking of the different versions of example data item is processed to determine the vector of weights. For example, the candidate vector of weights may be initialized as the average over the rows of ^^ which may be considered to correspond to initialising the model with the “average human preference”. [0097] The performance of example implementations of the system 100 of FIG. 1 has been experimentally investigated. In general, the experiments described below with reference to FIGS. 4 to 8 employ a pre-trained large language model (Gemma 1.12B model, “Gemma: Open models based on Gemini research and technology”, 2024, arxiv.org/abs/2403.08295) which has been fine-tuned as described above with reference to FIGS.1 to 3. The final layer of the Gemma model has been replaced by “^^” counterparts corresponding to the reward features ^^ఏ. Training was carried out using gradient ascent to solve the aforementioned optimization objectives Eq. (1) and (2). This trained model is referred to as “RFM” (reward feature model) in the following. The experiments also involve three comparative examples referred to as “non-adaptive baseline”, “adaptive baseline” and “adaptive linear baseline”. For the non-adaptive baseline model, the pre-trained large language model is fine-tuned using a reward model that neither distinguishes between raters nor performs adaptation. It was trained 38072935-1
using gradient ascent to solve max ∑^∈^ log^^^〈^^௫ ఏ ^,௬^,௬ᇱ^,௭^ 〉^ starting from the pre-trained ఏ parameters of Gemma. is obtained by optimizing max ∑^∈^ log^^^〈^^௫ ఏ ^,௬^,௬ᇱ ,௭ 〉^
to adapt the parameters of the non- ఏ ^ ^ baseline to user ℎ. The adaptive linear baseline is similar to the adaptive baseline, but
layer is adapted. All models were trained using the “UltraFeedback” dataset (Cui et al., “UltraFeedback: Boosting language models with high-quality feedback”, 2023, arxiv.org/abs/2310.01377) which is a dataset carefully curated to ensure the quality and diversity of the responses. The version of the dataset we adopted has a training set with 60829 examples and a test set with 985 examples. The maximum context and response lengths were set to ^^௫ ൌ ^^௬ ൌ 1525 tokens. Training and adaptation were carried out using gradient ascent with a learning rate of 10e-5. We evaluated the models using the weights computed at the end of training (with the exception of the experiments with reward models as raters (of FIG.8): in this case a random 90%-10% split of the training set was carried out and the error in the smaller subset is used as a criterion to select the model to undergo adaptation). Training was carried out for 6000 parameter updates with a batch size of 32. This means that the training procedure went over the entire UltraFeedback training set approximately three times. Each time the example (^^^ , ^^^ ,^^′^) was encountered a new rater ^^^^ೖ was sampled uniformly at random from ^ ^ ^, as further explained below. [0098] Although UltraFeedback’s data is not rater-annotated, it does come with the features used to compute preferences. Associated with each (^^^ ,^^^) and each (^^^ ,^^′^) in the dataset, we ସ have four features ^^^^^^^^ ,^^^^^^ୀ^ indicating the level of helpfulness, honesty, instruction following, and
^^ with respect to context ^^. These features were computed using OpenAI’s GPT-4. The preferences in the original dataset were defined based on a single feature ^^^ selected per example. We used instead the more general notion of a linear combination of features. Specifically, given ^^^^ ,^^^ ,^^′^^ and ^^^^ ∈ ℝ ସ , ^^^ ൌ ^^^〈^^ ^^^^ ,^^^^,^^^^〉 ^ 〈^^ ^^^^ ,^^′^^,^^^^〉^ is set where ^^^∙^ is the
function. Three
using different sets ^^^ defined through the raters’ preference coefficients ^^^^. When an example was used during training, a rater ℎ^ was sampled uniformly at random from ^^^ and concatenated to the example together with the corresponding preference ^^^ computed as described above. 38072935-1
[0099] For a first scenario, training raters sampled from four Gaussian distributions whose means are the “one-hot raters”: ^^^^ ൌ ^1,0,0,0^, ^^^^ ൌ ^0,1,0,0^, ^^^^ ൌ ^0,0,1,0^, and ^^^^ ൌ ^0,0,0,1^. The covariance of all four normal distributions was set to 0.3^^, where ^^ is the 4x4 identify matrix. That is
^^^ ≡ ^^൫^^^^, 0.3^^൯ for ^^ ൌ 1, 2, 3, 4. 30 preference vectors ^^^^ are sampled from each distribution, resulting in a set ^ ^ ^ w
^ ith 120 This way of generating preference vectors ^^ ^
^^ results in raters ℎ that mostly care about a single assess RFM’s performance using different numbers of reward features. The notation RFM(d) indicates that a d- dimensional feature vector ^^ఏ ∈ ℝௗ was used. Notably RFM did not have access to the real features ^^ underlying the to derive ^^^ ൌ ^^^〈^^ ^^^^ ,^^^^,^^^^〉 ^ 〈^^ ^^^^ ,^^′^^,^^^^〉^.
Instead, RFM learned feature functions ^^ఏ by ^^ ∈ ℝௗ as described above. FIG.
4 shows the intra-user test accuracy obtained the baseline model and RFM throughout training (two independent training runs for each model). This corresponds to the fraction of examples ^^^^ ,^^^ ,^^′^ , ℎ^^^ in the test set for which the models can correctly predict the preference ^^^ ∈ ^0,1^. The reference numerals 400, 402, 404, and 406 respectively indicate experimental data for RFM(128), RFM(32), RFM(8) and the baseline. That is, it is an estimate of the models’ intra-user generalisation. It can be seen that RFM significantly outperforms the rater-agnostic baseline model. [0100] To assess the ability of the models to adapt to unseen individuals, three held-out users are employedthe following preference weights: ^^^భ ൌ ^1, 0, 0, 0^, ^^^మ ൌ ^1, 0,0,1^, and ^^^య ൌ ^1, 0,െ1, 1^. These users were defined to
profiles in the population ^^. Comparing the held-out users with the aforementioned distributions ^^^, it can be seen that the held-out users ^^^^are increasingly “out of distribution” that is, they become increasingly less likely to be sampled from any of the distributions ^^^. The RFM’s inter-user generalisation is assessed using different numbers of examples “m” to do the adaptation. For each held-out user ℎ^, an initial dataset ^^^^ is defined by sampling 10 examples uniformly at random from ^^ା and labelling them using (^^^ ൌ ^^^〈^^ ^^^^ ,^^^^,^^^^〉 ^ 〈^^ ^^^^ ,^^′^^,^^^^〉^) with the corresponding ^^^^. Eq. (1) is then solved using this data and its accuracy is assessed on the test set (also properly relabelled). This process is iterated until a dataset ^^^^with 90 examples is obtained. For each of the two feature functions ^^ఏ computed during training (i.e. Eq. (2)), and each ^^ ∈ ^8, 32, 128^, the aforementioned steps are repeated 5 times. For each ^^ ∈ ^10, 30, 50, 70, 90^, 5 independent adaptation runs were carried out with each of the 6 feature 38072935-1
functions computed during training. 1000 parameter updates were carried out and the inter- user test accuracy was estimated using the resulting weights. [0101] In FIG.5, the dashed line 500 indicates the baseline model, and the reference numerals 502, 504, 506, 508, 510 respectively indicate experimental data for RFM(8), RFM(32), RFM(128), the adaptive baseline, and adaptive linear baseline. The top row panels of FIG. 5 show the performance of the baselines and RFM in predicting the preferences of the held-out users. It can be seen that the preferences of held-out users ℎ^ and ℎଶ are close to the average preference among the raters ^^^, as indicated by the non- baseline’s results 500 on these
raters. It can further be seen that for held-out user ℎଷ, most “out-of-distribution” of the three, adaptation provides a clear increase in performance. [0102] A second scenario used a similar protocol to the above-described first scenario, but the training raters were sampled from different distributions. In particular, the raters are sampled to represent a situation where the human raters disagree significantly. To simulate this, the normal distributions used in the first scenario are replaced with 12 counterparts whose means are all possible permutations of the vector ^1,െ1,0,0^, i.e. a mean vector for each possible instantiation of a 4-dimensional vector is defined with one element equal to 1, one element equal to −1, and the remaining elements equal to zero, that is: ^^^^ ൌ ^1,െ1,0,0^, ^^^^ ൌ ^1,0,െ1,0^, ^^^^ ൌ ^1,0,0,െ1^, ^^^^ ൌ ^െ1,1,0,0^, … , ^^^^^^ ൌ ^0,0,െ1,1^. 12 normal distributions ^^^ ≡ ^^൫^^^^, 0.3^^൯ are defined and 10 training raters ^^^^ೖ~ ^^൫^^^^, 0.3^^൯ are sampled
^^^, totalling again 120
forming a set ^^^ଶ. An aspect of this experiment is to see how the methods perform with training raters whose preferences take more than one feature into account, i.e. with a more complex set of training raters than in the first scenario. Further, the means ^^ are selected to mimic disagreement among the resulting raters. The experiments are performed as described for the first scenario but with ^ ^ ^ଶ instead of ^ ^ ^^. That is in the second scenario the preferences of the same three held-out users as before are
but now using a different set of raters. The middle row panels in FIG. 5 show the experimental results for the second scenario. Because the baseline cannot distinguish between raters during training, opposing preferences become contradictory learning signals, making it difficult to capture any trends in the data. This explains why the non-adaptive baseline’s performance reduces to chance. It can be seen that all RFM versions outperform the comparative examples. Notably, RFM performs on held-out user ℎ^, whose preferences are based on the first feature ^^^ only. This suggests that the RFM can capture the reward features 38072935-1
underlying the raters’ preferences even when all the raters are based on combinations of such features (that is, even when the effect of features is not observed in isolation). [0103] FIG. 6 shows experimental results for the accuracy in predicting the preferences of held-out users on test set under the second scenario for the non-adaptive baseline, RFM, and adaptive baselines. More specifically, two highly capable models, Gemini 1.5 Pro and GPT- 4o, performing “in-context” adaptation are used to implement the adaptive baselines. To assess the LLMs’ prediction accuracy for held-out user ℎ^, ^^ ൌ 10 training examples ranked by ℎ^ are provided together with the test example to be ranked. FIG.6 shows the results of the non- adaptive baseline 600, RFM 602, Gemini 1.5 Pro
and GPT-4o 608 (for reference, the
shot” performance of Gemini obtained with ^^ ൌ 0 is also shown in FIG. 6 as indicated by reference numeral 604). It can be seen that RFM outperforms its in-context counterparts in predicting the preferences of held-out users ℎ^ and ℎଶ, and essentially matches their performance on ℎଷ.
[0104] In a third scenario, it is assessed how the good performance of RFM in predicting preferences transfers to the scenario where it is used to steer the behaviour of an LLM. A “best- of-n” over ^^ ൌ 40 responses to each context ^^ in the test set. The responses were generated by Gemma 29B and Gemma 227B (20 responses each). To be able to easily score the responses, the UltraFeedback features ^^^^^,^^^ are replaced with simple functions ^^′^^^,^^^ of the context ^^ and the response ^^. Specifically, four functions are defined: [0105] ^^^ ᇱ^^^,^^^: the length of ^^; [0106] ^^ ᇱ ଶ ^^^, ^^^: the number of adjectives in ^^; [0107] ^^ ᇱ ଷ ^^^, ^^^: the number of alliterations in ^^; [0108] ^^ସ ᇱ^^^,^^^: the number of words in ^^ that also occur in ^^. [0109] All the features were normalised to fall in the interval [0, 1]. These features were defined as simple illustrations of possibly conflicting subjective criteria. They were also kept simple to ensure that they can be computed with the architecture adopted for RFM. The bottom row of FIG.5 shows how well the baselines and RFM can predict the preferences of the held- out users using the new features ϕ′. Notably, RFM’s prediction accuracy lies between 80% and 90%, a significant improvement over the results shown in the middle row of FIG. 5 referring to the second scenario. This indicates that, when the features underlying the data can be computed with RFM’s architecture, training via Eq. (2) does indeed recover them. [0110] For each context ^^ in the test set, all 40 responses are scored using the non-adaptive baseline and RFM(32) adapted with ^^ ൌ 30 examples. Next, the best-of-n response to context 38072935-1
^^ is selected according to the baseline, ^^^, and RFM, ^^^. Then, 〈^^′^^^,^^^^,^^^^〉 and 〈^^′^^^,^^^^,^^^〉 are compared for each ^ℎ^^ ଷ ^ୀ^ , a winning model or a draw is declared. In words, which model was the winner is decided by comparing the “ground-truth” score of their selected responses. FIG. 7 shows results when best-of-n is applied with increasing “n”. Reference numerals 700, and 702 show respectively the win-rate of RFM and the non-adaptive baseline (reference numeral 704 indicates draws). The candidate responses used with best-of-n are all qualitatively similar (and thus not easy to distinguish), and they originate from a distribution that is different from the one used for training. Yet, RFM consistently outperforms the baseline for all held-out users, and the fraction of times its selected response is preferred grows with “n”. For example, ignoring ties, this means that by using RFM instead of the non-adaptive baseline, the response delivered to, say, ℎଶ would be an improvement 66% of the time, in expectation, according to ℎଶ’s own preference. [0111] FIG. 8 shows results for experiments in which publicly available reward models have been used as the raters ℎ^. Specifically, 8 reward models are used to score and rank all the test- set examples (the 8 employed model are: OpenAssistant reward-model-deberta-v3-large-v2, weqweasdas RM-Mistral-7B, OpenAssistant oasst-rm-2.1-pythia-1.4b-epoch-2.5, Ray2333_GRM-Gemma-2B-sftreg, Ray2333 reward-model-Mistral-7B-instruct-Unified- Feedback, weqweasdas RM-Gemma-7B, internlm internlm2-7b-reward, and openbmb Eurus- RM-7b). Notably, these reward models were trained with human preference data, and hence they reflect the opinion of real people. However, as these models do not distinguish between raters (much like the non-adaptive baseline), they reflect an average over preferences, and thus tend to “agree” considerably with each other. To avoid such overlapping, which may render the distinction of raters unnecessary, examples in the training and test sets in which 2 or fewer raters disagreed with the majority have been filtered out. This resulted in 23614 training examples and 401 test examples. A “leave-one-out” cross-validation is carried out. The cross- validation is composed of 8 rounds in which 7 of the models played the role of the raters ^^^ and the remaining model played the role of the held-out user. Results are shown in FIG. 8. It can be seen that given enough (but still few) adaptation examples, RFM’s performance 800 either matches or significantly surpasses that of the non-adaptive baseline 802 (dashed line). This indicates that RFM is useful in real scenarios. For example, it can work as a form of “safety net” to make sure that minority preferences are also properly represented. Example hardware implementations 38072935-1
[0112] In some implementations the generative neural network system, e.g. a language model or a visual language model, is stored on a user computing device, i.e. a device local to the user, such as a mobile device e.g. a mobile phone, or a smart speaker. [0113] In some implementations the generative neural network system is implemented on a remove server in communication with a user computing device over a wired or wireless network communications link between the user computing device and the server. [0114] The user computing device may be provided with an input mechanism, such as a text or voice interface, that enables user input from the user in a natural language. The user computing device may be provided with an output mechanism that provides a system output for the user in the or another natural language e.g. as speech or text; or in some other way, e.g. by displaying an image. The input and output mechanism may comprise, e.g., a keyboard, microphone, speaker, display, and/or camera. [0115] As an example the input mechanism may comprise a system configured to input audio data characterizing a speech waveform of speech representing the input from the user in a natural language, and configured to convert the audio data into tokens representing the speech in the natural language, e.g. representing a transcription of the spoken input. The output mechanism may comprise a system configured to receive tokens representing the output for the user in the or another natural language and a system configured to convert the received tokens into audio data representing a waveform of speech representing the output to the user in the natural language, i.e. representing spoken words. [0116] As a further example, the trained system can be deployed in an environment that enables a user to provide a request for the system, e.g. to process a multimodal conditioning input to generate a corresponding data item output. A users can provide the request, e.g., by way of a user interface or through an application programming interface (API). The request can be transmitted from a user device, e.g., over a data communications network such as the internet, to one or more computers implementing the system, e.g., in a data center. The system can generate a data item and then transmit the data item to a user device over a data communications network. Some example applications [0117] The (trained) generative neural network system can be used for diagnosing a fault, or for correcting undesired behavior, in a mechanical or computing system operating in the real world environment. The conditioning input may comprise a description and/or image of one 38072935-1
or more observations of the mechanical or computing system, e.g. of operation of the system, optionally obtained from one or more sensors sensing a condition or operation of the system. An image observation may be converted into a text description e.g. using an image captioning system or in other ways. The generated data item may comprise an image, audio, or text that identifies (described) a likely cause of the fault or undesired behavior. This may be used to repair the fault or correct the behavior. The reward model can define relatively more useful types of output for repairing the fault or correcting the behavior. [0118] The (trained) generative neural network system can be used for controlling a mechanical agent such as a robot or vehicle. For example the conditioning input may comprise a description of a task to be performed, and the generated data item may comprise a list of sub- tasks to be performed by the mechanical agent (trained to perform such sub-tasks), in order to perform the task. The reward model can define relatively more preferable or useful types of sub-task. [0119] Example multimodal applications [0120] The generative neural network system may comprise a multimodal machine learning system such as a visual language model (VLM). That is implementations of the generative neural network system can perform a multimodal task in which the conditioning input and data item, collectively, comprise data of multiple different types. As used herein text can include numbers, punctuation, special symbols, and so forth. [0121] In some implementations, after training, a particular task that is to be performed by the generative neural network system can be described by part or all of a sequence of text in the conditioning input to the system. For example in a conditioning input that includes an image such a prompt might specify “Generate a caption”, “Generate a description”, “Answer the following question: [about the image or video]”, or “Detect a person”. Where the system is used for an agent control task a prompt may define “Take the knife out of the drawer”, or “Q: What action should the robot take to take the knife out of the drawer?”. Also or instead such a prompt may give one or more examples of a task to be performed. The generative neural network system can be trained on multiple natural and/or computer languages and the prompt may then specify a language to use. [0122] A few further examples of some machine learning tasks that can be performed by a system trained as described herein follow. The tasks described below may be tasks that require 38072935-1
spatial awareness or other context from the image or video. For example, a prompt may ask “What is the object in the top left corner?”. [0123] In general for the tasks below the system can have been trained or fine-tuned on examples of the input and output for the task. For example the system can have been trained using still or moving images containing one or more objects or actions, and corresponding sequences of text or other data e.g. describing or classifying the images. However large, “foundation” models can, in general, perform some tasks zero-shot, i.e. without having been specifically trained on those tasks. [0124] As one example the task may comprise an object or action detection task. For example the generated data item may comprise or represent text that describes or otherwise labels detected object(s) or action(s) in a conditioning input comprising an image or audio, and may include coordinates such as bounding-box coordinates for the detected object(s) or action(s), e.g. "102090100 cat 2030100100 dog”. [0125] As another example the task may comprise a classification task, e.g. an object or action classification task. The generated data item may comprise data, e.g. text, that classifies the object(s) or action(s) in represented in the conditioning data, e.g. in an image or audio, into one of a plurality of classes, or that otherwise classify object(s) or action(s) represented in the conditioning data. [0126] As another example the task may comprise a still or moving image describing task, e.g. a captioning task (which, as used here, includes an audio description task to explain what is happening in an image). The generated data item may comprise data, e.g. text, describing an image or video in the conditioning data. For example the generated data item may provide a caption or description or it may count objects in the image or video, or it may provide some other form of description. [0127] As another example the task may comprise a still or moving image question-answering task. The generated data item may comprise data, e.g. text, that answers a question about the conditioning input, e.g. an image or audio, where the question is also specified in the conditioning input, e.g. as sequence of text. This may be used, e.g., to answer questions about visual plots and charts or about sounds. [0128] As another example the task may comprise a character or word recognition task, e.g. an OCR (optical character recognition) task. The conditioning input may comprise a still or moving image and the generated data item may comprise text that represents characters or words in the conditioning input, e.g. in a natural language. 38072935-1
[0129] As another example the task may comprise a still or moving image generation task. The generated data item may comprise image data defining values for pixels of a still or moving image, and the conditioning input, e.g. a sequence of text, may describe or characterize the image to be generated. Merely as an example, an image of a plot or chart may be generated to represent the conditioning input, e.g. comprising text. [0130] As another example the task may comprise a computer language text generation task. The conditioning data may comprise a natural language description of a task to be performed, and optionally an image (if the task is to be performed on or in relation to an image), and the generated data item may comprise text in a computer language to perform the task, e.g. a task of analyzing the content of the image to provide a result of the analysis or to search for information relating to the content of the image. [0131] As a particular example the computer language in the generated data item may comprise computer language for invoking a function or calling one or more external APIs. Merely as one example, such a data item may comprise data formatted as a JSON object. As previously, the conditioning input may define the task to be performed and may also include an image in relation to which the task is to be performed. In general the task can involves manipulation of particular types of data that may benefit from access to an API such as mathematical data, date/time related data, scientific data, recent data that may post-date training of the system (that may be accessed by a search function or API), and so forth; and the generated data item may comprise text in a computer language for performing the task. The method may then include using the text in the computer language to perform the task. [0132] In general where the generated data item comprises text this may be converted to speech representing the text, and an audio (speech) output provided. [0133] In some implementations the task comprises an agent control task in which the agent interacts with an environment to perform the agent control task. In these implementations the conditioning input can include an observation characterizing the environment. For example the conditioning input can include a sequence of text that defines the task to be performed by the agent and the image can represent an observation of the environment, e.g. captured by a camera or other imaging device from a real-world environment. The generated data item can comprise an action selection output, e.g. including text, that is used to select one or more actions to be performed by the agent in the environment in response to the observation. As an illustration the generated data item may define an action as text such as “A: 132114128525 156”, that can be converted into a control signal for a mechanical agent, such as a robot, e.g. 38072935-1
“Δ^^ ൌ ^0.1,െ0.2,0^ Δ^^ ൌ ^10^, 25^,െ7^^”. The action selection output may also or instead define one or more low-level skills, e.g. from a vocabulary of previously learnt skills. As before, the sequence of text in the conditioning input to the system may describe the task to be performed, e.g. “What action should the robot take to [perform task]”. Examples of systems for controlling an agent that may be fine tuned as described herein can include PaLM-E (Driess et al. arXiv:2303.03378), RT-1 (Brohan et al. arXiv:2212.06817), and RT-2 (Brohan et al. arXiv:2307.15818). [0134] In some agent control implementations, the environment is a real-world environment and the agent is a mechanical agent interacting with the real-world environment, e.g., a robot or an autonomous or semi-autonomous land, air, or sea vehicle operating in or navigating through the environment, and the actions are actions taken by the mechanical agent in the real- world environment to perform the task. For example, the agent may be a robot or other mechanical agent interacting with the environment to accomplish a specific task, e.g., to locate or manipulate an object of interest in the environment or to move an object of interest to a specified location in the environment or to navigate to a specified destination in the environment. In these implementations, the observations may include, e.g., one or more of: images, object position data, and sensor data to capture observations as the agent interacts with the environment. The actions may define control signals to control the robot or other mechanical agent, e.g., positions, torques, or other control signals for the parts of the mechanical agent, or higher-level control commands. [0135] In some agent control implementations the agent may be a human agent and the environment may be a real-world environment. For example the agent can be a human user of a digital assistant such as a smart speaker, smart display, or some other device that is used to instruct the user to perform actions. The task may be any real-world task that the user wishes to perform. The observations may be obtained from an observation capture subsystem, e.g. a monitoring system such as a video camera or sound capture system, to capture visual observations of the user performing the task. The actions may comprise instructions in the form of, e.g., text, image, video, or audio data such as speech, that guide the user in performing the task. [0136] In this specification, the term "configured" is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those 38072935-1
operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions. [0137] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure. [0138] The term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. 38072935-1
Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics. [0139] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics. [0140] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of AI and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors. 38072935-1
[0141] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in AI and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases. [0142] Computers capable of executing a computer program can be based on general-purpose microprocessors, special-purpose microprocessors, or a combination of both. They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The essential elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the AI model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage. [0143] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), 38072935-1
and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence. [0144] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction. [0145] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models. [0146] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend 38072935-1
on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience. [0147] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP/IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities. [0148] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. [0149] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program 38072935-1
components and systems can generally be integrated together in a single software product or packaged into multiple software products. [0150] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. [0151] What is claimed is: 38072935-1
Claims
CLAIMS 1. A computer-implemented method of generating a data item using a generative neural network system, wherein the generative neural network system comprises: a data item generator neural network configured to process a conditioning input to generate a data item according to a data item generation policy; and a feature neural network subsystem comprising a feature neural network configured to process i) the conditioning input, and ii) the generated data item, to generate a vector of features; and to determine a reward value of the data item from the vector of features and from a vector of weights that characterizes one or more aspects of the generated data item; the method comprising: obtaining the vector of weights for a data item to be generated; obtaining the conditioning input for the data item to be generated; and using the data item generator neural network to process the conditioning input to generate the data item based on a reward value of the data item determined from the vector of features for the generated data item and the vector of weights for the data item to be generated.
2. The method of claim 1, wherein the feature neural network is configured to determine the reward value for the data item from the vector of features and from the vector of weights that characterizes one or more aspects of the generated data item by determining an inner product of the vector of features and a vector of weights that characterizes one or more aspects of the generated data item.
3. The method of claim 1 or 2, wherein using the data item generator neural network to process the conditioning input to generate the data item based on a reward value of the data item determined from the vector of features for the generated data item and from the vector of weights for the data item to be generated comprises: processing the conditioning input data using the data item generator neural network for each of a plurality of data generation episodes to generate a plurality of different versions of the data item according to the data item generation policy; processing the conditioning input and each respective version of the generated data item to generate the vector of features for each respective version of the generated data item; 38072935-1
determining a respective reward value of each version of the data item from the vector of features for the version of the generated data item and from the vector of weights for the data item to be generated; and using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item.
4. The method of any preceding claim, wherein the generative neural network system comprises a plurality of data item generator neural networks each configured to process the conditioning input to generate the data item according to a different respective data item generation policy; and further comprising: processing the conditioning input data using each of the item generator neural networks to further generate at least one version of the data item according to each of the different respective data item generation policies, wherein using the data item generator neural network to process the conditioning input to generate the data item based on a reward value of the data item determined from the vector of features for the generated data item and from the vector of weights for the data item to be generated comprises: processing the conditioning input and each respective version of the generated data item to generate the vector of features for each respective version of the generated data item; determining a respective reward value of each version of the data item from the vector of features for the version of the generated data item and from the vector of weights for the data item to be generated; and using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item.
5. The method of claim 3 or 4, wherein using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item comprises selecting, as the generated data item, the version of the generated data item having the highest reward value.
6. The method of claim 4, wherein using the respective reward values of the versions of the generated data item to select one of the versions as the generated data item comprises: 38072935-1
using the respective reward values of the versions of the generated data item to select one of the data item generator neural networks, and selecting, from the versions of the generated data item generated by the selected data item generator neural network, the version of the generated data item having the highest reward value.
7. The method of claim 1 or 2, wherein using the data item generator neural network to process the conditioning input to generate the data item based on a reward value of the data item determined from the vector of features for the generated data item and from the vector of weights for the data item to be generated comprises: updating a plurality of learnable parameters of the data item generator neural network based on the reward value.
8. The method of any preceding claim, wherein obtaining the vector of weights comprises: processing an example conditioning input data using the data item generator neural network to generate a plurality of different versions of a corresponding example data item from the example conditioning input data; obtaining a ranking of the different versions of the example data item from a user; and determining the vector of weights from the ranking.
9. The method of claim 8, wherein obtaining a ranking of the different versions of the example data item from a user comprises: obtaining a ranking of the different versions of the example data item from a group of users.
10. The method of claim 8 or 9, wherein determining the vector of weights from the ranking comprises: processing the example conditioning input data and each version of the example data item to generate the vector of features for each version of the example data item, and determining the vector of weights from the ranking and the vectors of features for the different versions of the example data item. 38072935-1
11. The method of claim 10, wherein the different versions of the example data item comprise a first and a second version, wherein the first version has been ranked higher than the second version by the user, and determining the vector of weights from the ranking and the vectors of features for the different versions of the example data item comprises: determining a reward value of the first version of the example data item from the vector of features for the first version of the example data item and from a candidate vector of weights, determining a reward value of the second version of the example data item from the vector of features for the second version of the example data item and from the candidate vector of weights, optimizing the weights of the candidate vector of weights to maximise a difference between the reward value of the first version of the example data item and the reward value of the second version of the example data item, and determining the optimized candidate vector of weights as the vector of weights.
12. The method of any preceding claim, the method further comprises an initial step of receiving, by a first computer system, over a data communications network, from a second computer system having one or both of a larger working memory and higher computational capacity than the first computer system, values of a plurality of learned parameters defining the trained generative neural network system, and wherein the first computer system implements the trained generative neural network system based on the received values of the plurality of learned parameters to perform said steps of i) obtaining the vector of weights for a data item to be generated, ii) obtaining the conditioning input for the data item to be generated, and iii) using the data item generator neural network to process the conditioning input to generate the data item based on a reward value of the data item determined from the vector of features for the generated data item and from the vector of weights for the data item to be generated.
13. The method of any preceding claim, wherein the second computer system has performed the training of the generative neural network system.
14. The method of claim 12 or 13, wherein the second computer system is a user device, and the second computer system is a remote server. 38072935-1
15. The method of any one of claims 12 to 14, wherein obtaining the conditioning input for the data item to be generated comprises receiving user input specifying the conditioning input for the data item to be generated from a user of the first computer system.
16. A method of training a feature neural network, the method comprising: obtaining a plurality of training items, each training item comprising a respective training conditioning input data, a plurality of different versions of a corresponding training data item generated from the training conditioning input data using the data item generator neural network, a ranking of the respective different versions of the training data item obtained from a respective user, and user information specifying said user from a group of users, and training the feature neural network using the plurality of training items.
17. The method of claim 16, wherein training the feature neural network using the plurality of training items comprises, in each of a plurality of training iterations: for each training item, processing, using the feature neural network, the respective training conditioning input data and each version the corresponding training data item to generate a vector of features for each version of the corresponding training data item; for each training item, determining a respective training reward value for each version of the respective training data item from the vector of features for the respective version of the training data item and from a candidate vector of weights associated with the user specified in the respective user information, and updating a plurality of learnable parameters of the feature neural network and the weights of candidate vectors of weights based on the training reward values.
18. The method of claim 16 or 17, wherein the feature neural network is the feature neural network of any one of claims 1 to 15.
19. A computer-implemented method of training a data item generator neural network, wherein a generative neural network system comprises the data item generator neural network configured to process a conditioning input to generate a data item according to a data item generation policy, the method comprising: 38072935-1
obtaining dataset of training examples, each training example comprising a training conditioning input and a corresponding training data item and, for each of a plurality of the training data items: obtaining a vector of weights that characterizes one or more aspects of data items to be generated; processing the training conditioning input and the training data item using a feature neural network subsystem to generate a vector of features; determining a reward value for the data item from the vector of features and the vector of weights; and training the data item generator neural network using the reward values.
20. The method of any of claims 1-19, wherein the data item generator neural network generates an output token sequence from an input token sequence including the conditioning input, and wherein the data item generator neural network is configured to process the input token sequence to generate for each position in the output token sequence, a respective score for each token in a vocabulary of output tokens.
21. The method of any of claims 1-20, wherein the generative neural network system comprises a language and/or image generation neural network system, wherein the conditioning input comprises a prompt for the generative neural network system, and wherein the data item comprises a language and/or image response to the prompt.
22. The method of any preceding claim, wherein the data item generator neural network is used for controlling an agent acting in an environment to perform a task, and wherein the conditioning input comprises an observation of the environment and the data item generated from the conditioning input specifies an action to be performed by the agent.
23. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1-22. 38072935-1
24. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1-22. 38072935-1
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463658098P | 2024-06-10 | 2024-06-10 | |
| US63/658,098 | 2024-06-10 | ||
| US202563751634P | 2025-01-30 | 2025-01-30 | |
| US63/751,634 | 2025-01-30 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025259680A1 true WO2025259680A1 (en) | 2025-12-18 |
Family
ID=98051662
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/033016 Pending WO2025259680A1 (en) | 2024-06-10 | 2025-06-10 | Generating data items based on a multidimensional reward model |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025259680A1 (en) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200293497A1 (en) * | 2019-03-13 | 2020-09-17 | Deepmind Technologies Limited | Compressed sensing using neural networks |
| US20200372370A1 (en) * | 2019-05-23 | 2020-11-26 | Deepmind Technologies Limited | Large scale generative neural network model with inference for representation learning using adversial training |
| US20210397819A1 (en) * | 2020-06-19 | 2021-12-23 | Samsung Electronics Co., Ltd. | Object recognition method and object recognition apparatus |
| US20220101186A1 (en) * | 2020-09-29 | 2022-03-31 | International Business Machines Corporation | Machine-learning model retraining detection |
| US20240127058A1 (en) * | 2017-10-27 | 2024-04-18 | Google Llc | Training neural networks using priority queues |
-
2025
- 2025-06-10 WO PCT/US2025/033016 patent/WO2025259680A1/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20240127058A1 (en) * | 2017-10-27 | 2024-04-18 | Google Llc | Training neural networks using priority queues |
| US20200293497A1 (en) * | 2019-03-13 | 2020-09-17 | Deepmind Technologies Limited | Compressed sensing using neural networks |
| US20200372370A1 (en) * | 2019-05-23 | 2020-11-26 | Deepmind Technologies Limited | Large scale generative neural network model with inference for representation learning using adversial training |
| US20210397819A1 (en) * | 2020-06-19 | 2021-12-23 | Samsung Electronics Co., Ltd. | Object recognition method and object recognition apparatus |
| US20220101186A1 (en) * | 2020-09-29 | 2022-03-31 | International Business Machines Corporation | Machine-learning model retraining detection |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111279362B (en) | capsule neural network | |
| US20240029436A1 (en) | Action classification in video clips using attention-based neural networks | |
| US11429860B2 (en) | Learning student DNN via output distribution | |
| Le | A tutorial on deep learning part 1: Nonlinear classifiers and the backpropagation algorithm | |
| JP7820624B2 (en) | Neural Networks with Adaptive Gradient Clipping | |
| US20200327450A1 (en) | Addressing a loss-metric mismatch with adaptive loss alignment | |
| JP7512416B2 (en) | A Cross-Transform Neural Network System for Few-Shot Similarity Determination and Classification | |
| US20250131694A1 (en) | Learning with Neighbor Consistency for Noisy Labels | |
| WO2025104314A1 (en) | Training image processing neural networks using cross-modal alignment | |
| WO2025166256A1 (en) | Generation of an output token sequence from an input token sequence using two language model neural networks | |
| US20240378869A1 (en) | Decoupled Encoder-Decoder Networks for Image Simulation and Modification | |
| US20250284971A1 (en) | Training neural networks through reinforcement learning using multi-objective reward neural networks | |
| US20250252309A1 (en) | Hardware-friendly and parameter-efficient tuning of neural networks | |
| WO2025166364A1 (en) | Generating outputs using a trained model and a task-specific model | |
| US20240256865A1 (en) | Training neural networks using learned optimizers | |
| WO2025068601A1 (en) | Evolving prompts for neural networks | |
| WO2024236063A1 (en) | Performing image processing tasks based on demonstration examples | |
| WO2025259680A1 (en) | Generating data items based on a multidimensional reward model | |
| CN116868203A (en) | Neural network using adaptive gradient clipping | |
| US20260031092A1 (en) | Training audio encoder neural networks using denoising losses | |
| US20260093982A1 (en) | Efficient decoding of output sequences using parameter sharing | |
| US20260093990A1 (en) | Alignment of neural networks using architectural modifications and training examples | |
| US20250363337A1 (en) | Training generative neural networks using soft preferences | |
| US20250384663A1 (en) | Influential data selection for neural network training | |
| WO2025104214A1 (en) | Training machine learning models using online data selection techniques |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25822074 Country of ref document: EP Kind code of ref document: A1 |