EP4681093A1 - Search query generation using machine-learning model - Google Patents
Search query generation using machine-learning modelInfo
- Publication number
- EP4681093A1 EP4681093A1 EP25728594.0A EP25728594A EP4681093A1 EP 4681093 A1 EP4681093 A1 EP 4681093A1 EP 25728594 A EP25728594 A EP 25728594A EP 4681093 A1 EP4681093 A1 EP 4681093A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- media items
- templates
- user
- template
- attributes
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/40—Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
- G06F16/43—Querying
- G06F16/432—Query formulation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/40—Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
- G06F16/41—Indexing; Data structures therefor; Storage structures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/40—Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
- G06F16/43—Querying
- G06F16/438—Presentation of query results
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
- G06F40/186—Templates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
Definitions
- a computer-implemented method includes identifying attributes from a group of media items associated with a user account.
- the method further includes generating one or more templates with the attributes.
- the method further includes generating templates with combinations of the attributes.
- the method further includes scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes.
- the method further includes selecting one or more of the templates based on corresponding scores.
- the method further includes providing the one or more templates as input to a large language model.
- the method further includes outputting, with the large language model, descriptive text based on the one or more templates.
- the method further includes providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
- the method further includes receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the method further includes determining a number of times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items.
- the method further includes determining a number of the media items that include the corresponding attributes; and for each template, calculating a minimum threshold value by dividing a number of categories by the number of the media items in the group of media items, where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and where selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
- the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
- the one or more templates are provided to the large language model with a level of randomness to be associated with the descriptive text.
- the one or more templates are provided to the large language model with a level of complexity to be associated with the descriptive text.
- the method further includes comprising precomputing the attributes that are part of the group of media items associated with the user account.
- the method further includes receiving a request from a user for media items that include one or more particular attributes and providing the group of media items to the user based on the group of media items including the one or more particular attributes, where the suggested search query is provided with the group of media items.
- the media items include one or more screenshots captured by a user device associated with the user account.
- a system includes one or more processors and a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the processor, cause the one or more processors to perform operations.
- the operations include identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
- the operations further include receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the operations further include determining a number of times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items.
- the operations further include determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
- the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
- a non-transitory computer-readable medium includes instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations.
- the operations include identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
- the operations further include receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the operations further include determining a number of times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items.
- the operations further include determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
- the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
- Figure 1 is a block diagram of an example network environment, according to some embodiments described herein.
- Figure 2 is a block diagram of an example computing device, according to some embodiments described herein.
- Figure 3 illustrates an example user interface that includes a group of media items, according to some embodiments described herein.
- Figure 5 illustrates two example user interfaces that include suggested search queries, according to some embodiments described herein.
- Figure 6 illustrates an example user interface that includes suggested search queries based on an initial search request from a user, according to some embodiments described herein.
- Figure 8 illustrates a flowchart of another example method to generate suggested search queries, according to some embodiments described herein.
- the media items may comprise any kind of media items, e.g., image(s), documents), video file(s)/data, audio file(s)/data, and/or video and audio file(s)/data.
- Users search for media items such as photos and videos using basic queries like “March 2024” and “John”; however, due to the overwhelming number of media items that are stored for a user, the search results may return too many results for users to find what they are looking for. Searches can be improved by adding details to the images (e.g., adding descriptions, tags, etc.), but users may not remember enough about the images to improve their search queries.
- the technology described below advantageously uses a large language model (LLM) to output suggested search queries for users for searching media items that guides the users to implement more complex search queries.
- the complex search queries can result specific results that match the user’s information need. Further, as a result of teaching users to implement more complex search queries, the users become more adept at crafting search queries. Instead of performing multiple searches to find a particular image, the suggested search queries result in using fewer computer resources to obtain their desired search results.
- Figure 1 illustrates a block diagram of an example network environment 100.
- the network environment 100 includes a media server 101, and user devices 115 coupled to a network 105. Users 125a, 125n may be associated with respective user devices 115a, 115n.
- the network environment 100 may include other servers or devices not shown in Figure 1.
- a letter after a reference number e.g., “115a,” represents a reference to the element having that particular reference number.
- a reference number in the text without a following letter, e.g., “115,” represents a general reference to embodiments of the element bearing that reference number.
- the media server 101 may include a processor, a memory, and network communication hardware.
- the media server 101 is a hardware server.
- the media server 101 is communicatively coupled to the network 105 via signal line 102.
- Signal line 102 may be a wired connection, such as Ethernet, coaxial cable, fiber-optic cable, etc., or a wireless connection, such as Wi-Fi®, Bluetooth®, or other wireless technology.
- the media server 101 sends and receives data to and from one or more of the user devices 115a, 115n via the network 105.
- the media server 101 may include a media application 103a, a machine-learning model 120, and a database 199. Although the machine-learning model 120 is illustrated as being stored on the same media server 101 as the media application 103a, in some embodiments the machine-learning model 120 is stored on a separate server.
- the machine-learning model 120 is trained to provide text in response to a query.
- the machine-learning model 120 may be a large language model (LLM) that is designed for natural language processing tasks such as language generation.
- LLM is trained to receive one or more templates and a prompt, and output descriptive text based on the one or more templates.
- the machine-learning model 120 receives the one or more templates with attributes from a media application 103 (e.g., from the media application 103a on the media server, or a media application 103b, 103c stored on a user device 115a, 115b.
- the machine-learning model 120 outputs descriptive text.
- the trained machine-learning model 120 may include one or more model forms or structures.
- model forms or structures can include any type of neural-network, such as a linear network, a deep-leaming neural network that implements a plurality of layers (e.g., “hidden layers” between an input layer and an output layer, with each layer being a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural- network layers, and aggregates the results from the processing of each tile), a sequence-to- sequence neural network (e.g., a network that receives as input sequential data, such as words in a sentence, frames in a video, etc. and produces as output a result sequence), etc.
- a convolutional neural network e.g., a network that splits or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural- network layers, and aggregates the results from the processing of each tile
- the model form or structure may specify connectivity between various nodes and organization of nodes into layers.
- nodes of a first layer e.g., an input layer
- Such data can include, for example, one or more pixels per node, e.g., when the trained model is used for analysis, e.g., of an initial image.
- Subsequent intermediate layers may receive as input, output of nodes of a previous layer per the connectivity specified in the model form or structure.
- These layers may also be referred to as hidden layers.
- a first layer may output a segmentation between a foreground and a background.
- a final layer (e.g., output layer) produces an output of the machine-learning model.
- model form or structure also specifies a number and/ or type of nodes in each layer.
- the trained model can include one or more models.
- One or more of the models may include a plurality of nodes, arranged into layers per the model structure or form.
- the nodes may be computational nodes with no memory, e.g., configured to process one unit of input to produce one unit of output.
- Computation performed by a node may include, for example, multiplying each of a plurality of node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output.
- the computation performed by a node may also include applying a step/activation function to the adjusted weighted sum.
- the step/activation function may be a nonlinear function.
- such computation may include operations such as matrix multiplication.
- computations by the plurality of nodes may be performed in parallel, e.g., using multiple processors cores of a multicore processor, using individual processing units of a graphics processing unit (GPU), or special-purpose neural circuitry.
- nodes may include memory, e.g., may be able to store and use one or more earlier inputs in processing a subsequent input.
- nodes with memory may include long short-term memory (LSTM) nodes.
- LSTM nodes may use the memory to maintain “state” that permits the node to act like a finite state machine (FSM).
- FSM finite state machine
- the trained model may include embeddings or weights for individual nodes. For example, a model may be initiated as a plurality of nodes organized into layers as specified by the model form or structure. At initialization, a respective weight may be applied to a connection between each pair of nodes that are connected per the model form, e.g., nodes in successive layers of the neural network. For example, the respective weights may be randomly assigned, or initialized to default values. The model may then be trained, e.g., using training data, to produce a result.
- Training may include applying supervised learning techniques.
- the training data can include a plurality of inputs (e.g., templates) and a corresponding groundtruth output for each input (e.g., corresponding descriptive text).
- values of the weights are automatically adjusted, e.g., in a manner that increases a probability that the model produces the groundtruth output for the image.
- a trained model includes a set of weights, or embeddings, corresponding to the model structure.
- the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.
- a trained model includes a set of weights, or embeddings, corresponding to the model structure.
- the trained model may be is based on prior training, e.g., by a developer, by a third-party, etc.
- the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.
- the database 199 may store machine-learning models, training data sets, media items, etc.
- the database 199 may also store social network data associated with users 125, user preferences for the users 125, etc.
- the user device 115 may be a computing device that includes a memory coupled to a hardware processor.
- the user device 115 may include a mobile device, a tablet computer, a mobile telephone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing a network 105.
- user device 115a is coupled to the network 105 via signal line 108 and user device 115n is coupled to the network 105 via signal line 110.
- the media application 103 may be stored as media application 103b on the user device 115a and/or media application 103c on the user device 115n.
- Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, fiber-optic cable, etc., or wireless connections, such as Wi-Fi®, Bluetooth®, or other wireless technology.
- User devices 115a, 115n are accessed by users 125a, 125n, respectively.
- the user devices 115a, 115n in Figure 1 are used by way of example. While Figure 1 illustrates two user devices, 115a and 115n, the disclosure applies to a system architecture having one or more user devices 115.
- the media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some may be performed on the user device 115.
- Operations are performed in accordance with user settings.
- the user 125a may specify settings that images and/or other data of the user is to be stored only locally on a user device 115a and not on the media server 101.
- Transmission of user data (e.g., templates, attributes, etc.) to the media server 101, any temporary or permanent storage of such data by the media server 101, and performance of operations on such data by the media server 101 are performed only if the user has agreed to transmission, storage, and performance of operations by the media server 101.
- User data from media is not used in advertisements, responses are not reviewed by humans unless the user provides feedback or to address abuses or harm, and the user data is not used to train machine-learning models outside of the machine-learning models used to provide suggested search queries. Users are provided with options to change the settings at any time, e.g., such that they can enable or disable the use of the media server 101.
- machine learning models e.g., neural networks or other types of models
- Server-side models are used only if permitted by the user.
- a trained model may be provided for use on a user device 115.
- Updated model parameters may be transmitted to the media server 101 if permitted by the user 115, e.g., to enable federated learning. Model parameters do not include any user data.
- the media application 103 identifies attributes from a group of media items associated with a user account. Attributes of media items (e.g., image(s), documents), video file(s)/data, audio file(s)/data, and/or video and audio file(s)/data ) may comprise, for example, one or more of objects, people, landmarks, actions, etc. The attributes may be identified by performing independent component analysis.
- Attributes of media items e.g., image(s), documents), video file(s)/data, audio file(s)/data, and/or video and audio file(s)/data
- the attributes may be identified by performing independent component analysis.
- a template may be a set of attributes.
- a first template may include ⁇ Sara, John, climbing ⁇ and a second template may include ⁇ John, eating ⁇ .
- the media application 103 scores each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes.
- the number of attribute categories for the first template is three (i.e., Sara, John, and climbing).
- generating the score may include calculating a minimum threshold value based on the number of attribute categories in the corresponding template and multiplying the number of media items by a corresponding minimum threshold value.
- the media application 103 selects one or more of the templates based on the corresponding scores.
- the score of 2.64 results in needing at least three of the media items to include Sara, John, and climbing.
- the score of 0.5 results in needing at least one of the media items to include John and eating. Because only 2 images include Sara, John, and climbing and one of the images includes John and eating, the media application 103 selects the second template.
- the media application 103 provides one or more templates as input to the machinelearning model 120.
- the machine-learning model 120 outputs descriptive text based on the one or more templates. For example, the machine-learning model 120 may output “John is eating.”
- the media application 103 provides the descriptive text to the user as a suggested search query.
- the queries can be complex in various cases, such as “John eating a hamburger on a sidewalk in Manhattan,” “John eating shrimp noodles at a Japanese restaurant,” “John eating a taco and holding a soda in his other hand, with a taco stand in the background,” etc.
- the media application 103 may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), machine learning processor/ coprocessor, any other type of processor, or a combination thereof.
- the media application 103a may be implemented using a combination of hardware and software.
- FIG. 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein.
- Computing device 200 can be any suitable computer system, server, or other electronic or hardware device.
- computing device 200 is media server 101 used to implement the media application 103a.
- computing device 200 is a user device 115.
- computing device 200 includes a processor 235, a memory 237, an input/output (I/O) interface 239, a display 241, a camera 243, and a storage device 245 all coupled via a bus 218.
- the processor 235 may be coupled to the bus 218 via signal line 222
- the memory 237 may be coupled to the bus 218 via signal line 224
- the I/O interface 239 may be coupled to the bus 218 via signal line 226
- the display 241 may be coupled to the bus 218 via signal line 228,
- the camera 243 may be coupled to the bus 218 via signal line 230
- the storage device 245 may be coupled to the bus 218 via signal line 232.
- Processor 235 can be one or more processors and/or processing circuits to execute program code and control basic operations of the computing device 200.
- a “processor” includes any suitable hardware system, mechanism or component that processes data, signals or other information.
- a processor may include a system with a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for achieving functionality, a special-purpose processor to implement neural network modelbased processing, neural circuits, processors optimized for matrix computations (e.g., matrix multiplication), or other systems.
- CPU general-purpose central processing unit
- cores e.g., in a single-core, dual-core, or multi-core configuration
- multiple processing units
- processor 235 may include one or more co-processors that implement neural-network processing.
- processor 235 may be a processor that processes data to produce probabilistic output, e.g., the output produced by processor 235 may be imprecise or may be accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have temporal limitations. For example, a processor may perform its functions in real-time, offline, in a batch mode, etc. Portions of processing may be performed at different times and at different locations, by different (or the same) processing systems.
- a computer may be any processor in communication with a memory.
- Memory 237 is typically provided in computing device 200 for access by the processor 235, and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), Electrical Erasable Read-only Memory (EEPROM), Flash memory, etc., suitable for storing instructions for execution by the processor or sets of processors, and located separate from processor 235 and/or integrated therewith.
- Memory 237 can store software operating on the computing device 200 by the processor 235, including a media application 103.
- the memory 237 may include an operating system 262, other applications 264, and application data 266.
- Other applications 264 can include, e.g., an image library application, an image management application, an image gallery application, communication applications, web hosting engines or applications, media sharing applications, etc.
- One or more methods disclosed herein can operate in several environments and platforms, e.g., as a stand-alone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app") run on a mobile computing device, etc.
- the application data 266 may be data generated by the other applications 264 or hardware of the computing device 200.
- the application data 266 may include images used by the image library application and user actions identified by the other applications 264 (e.g., a social networking application), etc.
- I/O interface 239 can provide functions to enable interfacing the computing device 200 with other systems and devices. Interfaced devices can be included as part of the computing device 200 or can be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and/or storage device 245), and input/output devices can communicate via I/O interface 239. In some embodiments, the I/O interface 239 can connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensors, etc.) and/or output devices (display devices, speaker devices, printers, monitors, etc.).
- input devices keyboard, pointing device, touchscreen, microphone, scanner, sensors, etc.
- output devices display devices, speaker devices, printers, monitors, etc.
- Some examples of interfaced devices that can connect to I/O interface 239 can include a display 241 that can be used to display content, e.g., images, video, and/or a user interface of an output application as described herein, and to receive touch (or gesture) input from a user.
- display 241 may be utilized to display a user interface that includes a graphical guide on a viewfinder.
- Display 241 can include any suitable display device such as a liquid crystal display (LCD), light emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device.
- display 241 can be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen for a computer device.
- Camera 243 may be any type of image capture device that can capture images and/or video. In some embodiments, the camera 243 captures images or video that the I/O interface 239 transmits to the media application 103.
- the storage device 245 is a database that stores data related to the media application 103.
- the storage device 245 may store a training data set that includes labeled images, a machine-learning model, output from the machine-learning model, media items, etc.
- Figure 2 illustrates an example media application 103, stored in memory 237.
- the media application 103 includes a user interface module 202, an event module 204, an attribute module 206, and a template module 208.
- the user interface module 202 generates a user interface that displays information about media items.
- the media items may include images, videos, screenshots, etc.
- the media items may be captured by the camera 243 or retrieved from another source.
- the media items may be displayed chronologically, according to groups created by the event module 204, etc.
- the user interface module 202 generates a user interface that includes questions about identifying people and/or pets that appear in media items. For example, the user interface may ask a user to identify a name of a person in an image and a relationship of the person to the user. The user interface module 202 may ask for the identification of the people and pets based on the people and/or pets being in a predetermined number of the images (e.g., 10%). The user interface module 202 asks permission to make use of user data before asking for an identity of people and/or pets in the media items, unless such permission has previously been provided by the user.
- a predetermined number of the images e.g. 10%
- the user interface includes an option for performing a search of media items.
- the user interface generates suggested search queries for a user.
- the user interface module 202 requests user permission before generating the suggested search queries. If the user does not provide permission for the use of their user data, the media application 103 does not generate suggested search queries.
- the event module 204 generates groups of media items.
- the groups may be based on different factors, such as periodic grouping of media items (e.g., monthly, weekly, yearly, etc.) based on dates associated with the media items (e.g., capture data, last modified date).
- the event module 204 may generate groups of media items based on different events. For example, the event module 204 may generate a group of media items based on a trip (e.g., trip to London), an event (e.g., a wedding), a person or pet (e.g., a user’s child through the years), a theme (e.g., beach adventures, running trails during the seasons, evolution of a home renovation projection, etc.).
- a trip e.g., trip to London
- an event e.g., a wedding
- a person or pet e.g., a user’s child through the years
- a theme e.g., beach adventures, running trails during the seasons, evolution of
- the event module 204 outputs a group of media items based on a user request. For example, a user may request media items associated with a particular person. The event module 204 generates a group that includes all media items that include the particular person (e.g., images or videos that depict the person).
- the event module 204 includes a machine-learning model that is trained to receive media items as input and outputs groups of media items based on different events.
- the machine-learning model may include different types of machinelearning models, such as the examples mentioned above with reference to the machine- learning model in Figure 1.
- the event machine-learning model is a classifier that receives media items along with information about the media items and uses the information to output an event signal that indicates a likelihood that an event occurred.
- the information may include the results of performing optical character recognition on the images to identify text within the image that is indicative of a particular event. For example, an image of a menu captured at dinner may include the term “wedding” on it.
- the information may include the results of performing object recognition to identify objects associated with an event. For example, media items captured at a baby shower may have presents associated with babies.
- Metadata may include user-permitted factors such as a location and/or a time of capture of a video; whether a video was shared via a social network, an image sharing application, a messaging application, etc.; depth information associated with one or more video frames; sensor values of one or more sensors of a camera that captured the video, e.g., accelerometer, gyroscope, light sensor, or other sensors; an identity of the user (if user consent has been obtained), etc. For example, if the video was captured at night in an outdoor location with the camera pointing upwards, such metadata may indicate that the camera was pointed to the sky at the time of capture of the video, and therefore, is associated with an astronomical event.
- the event module 204 trains a machine-learning model to identify groups of images based on prediction of an event.
- the event machine-learning model may use a combination of metadata, optical character recognition, and other signals as inputs to the event machine-learning model.
- the metadata may indicate that a person is using a firecracker and the date is July 4 th , which results in the event machine-learning model outputting an event signal that corresponds to Independence Day.
- the event machine-learning model also outputs an event type for one or more of the events. Continuing with the example above, the event machine-learning model outputs the event signal and a likelihood that the event is Independence Day or a holiday.
- the attribute module 206 identifies attributes from a group of media items associated with a user account.
- the attributes may include people and pets in the media item (e.g., depicted in pixels of an image media item, described in text or attributes of a text media item, etc.), a location where the media item was captured, a time when the media item was taken, objects in the media item, text in the media item, events or activities associated with the media item, and/or an action performed in the media item (e.g., eating, swimming, running, etc.).
- one or more attributes are determined from labels provided by a user, such as when a user identifies different people in an image.
- the attribute module 206 may use the attributes in the media items to infer an identification of other attributes. For example, an image of a user in front of a restaurant with the name of the restaurant may be used by the attribute module 206 to identify an event that includes eating or an action that includes eating and a setting of a restaurant.
- the attribute module 206 performs independent component analysis of the media items to identify objects. Independent component analysis is a technique used to separate mixed signals into their independent components. In some embodiments, the attribute module 206 performs object recognition to identify objects in media items.
- the attribute module 206 may precompute attributes for a group of media items. For example, the attribute module 206 may precompute attributes for a media item responsive to a user creating a photo album, responsive to the attribute module 206 generating a particular group of media items (e.g., a group of media items for a theme, such as hiking, a vacation, an event, etc.), after every month, etc. Precomputing attributes advantageously reduces the time between a user requesting suggested search queries and receiving the suggested search queries generated by media application 103.
- the attribute module 206 generates a histogram that includes an identification of the attributes in each media item and a number of times an attribute appears in the group of media items.
- Table 1 includes an example of the attributes identified from six media items and Table 2 includes a count of the number of times each attribute appears in all six media items.
- the template module 208 generates templates with combinations of the attributes.
- the template module 208 generates the templates using a fixed set of attribute categories.
- the attribute categories may include different combinations of attribute categories, such as PERSON, PLACE, DATE, EVENT, ACTIVITY, and SCENE.
- the template module 208 may combine the attribute categories in particular combinations that are derived from common user search queries.
- the template may include “PERSON, PERSON, ACTIVITY,” “PERSON, PERSON, EVENT,” “EVENT, DATE,” “PERSON, ACTIVITY, DATE,” etc.
- the template module 208 generates a template with the attributes by populating a template of ⁇ PERSON, PERSON, EVENT ⁇ to ⁇ “personl”, “person2”, “skiing” ⁇ .
- the template module 208 calculates a minimum threshold value based on a number of attribute categories in a corresponding template. For example, a minimum threshold for two attribute category categories is 0.5, a minimum threshold for three attribute category categories is 0.66, and a minimum threshold for four attribute category categories is 0.75.
- the minimum threshold value is calculated using the following equation:
- N is the number of categories and N > 2.
- the equation for the minimum threshold uses the pigeonhole principle to ensure that there will be at least one media item that matches the generated combination. In some embodiments, 2 ⁇ N ⁇ 4.
- the template module 208 scores each template based on the number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. If the number of media items that include the corresponding attributes exceed the corresponding score, the template module 208 selects the template. In some embodiments, if multiple people are included in a template, the multiple people do not need to be included in the same media item to be considered in the number of media items with an attribute.
- Figure 3 illustrates an example user interface 300 of a group of media items, according to some embodiments described herein.
- the group of media items include a first image 305 of person 1 skiing, a second image 307 of person 1 and person 2 skiing, a third image 309 of person 2 climbing, a fourth image 311 of person 2 skiing while person 1 is present along with two other spectators, a fifth image 313 of person 1 kayaking, and a sixth image 315 of person 1 and person 2 skiing.
- the machine-learning model 120 receives ⁇ “personl”, “person2”, “skiing” ⁇ as input and provides “personl and person2 skiing together” as output.
- ⁇ “person2”, “climbing”, “December,” “2023 ⁇ is provided as input and the machine-learning model 120 outputs “person2 climbing in December 2023”
- ⁇ “personl”, “person2”, “wedding” ⁇ is provided as input and the machinelearning model 120 outputs “personl and person2 during their wedding”
- ⁇ “personl”, “person2”, “Christmas” ⁇ is provided as input and the machine-learning model 120 outputs “personl celebrating Christmas with person2.”
- the template module 208 provides the one or more templates along with a prompt where the prompt is based on at least one feature selected from a group of: emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, and delimiters.
- Figure 4 illustrates an example prompt 400 that is submitted with one or more templates to the machine-learning model 120, according to some embodiments described herein.
- the template module 208 provides the one or more templates with a level of randomness to the machine-learning model 120.
- the level of randomness may be on a scale of numbers (e.g., 0 for no randomness, 5 for medium randomness, and 10 for most random), a scale of words (e.g., no randomness, a little randomness, medium randomness, very random, etc.), or other paradigm.
- the level of randomness is used by the machine-learning model 120 to modify the descriptive text.
- a low level of randomness may result in concatenating the search results in a natural language style, such as the example above of “personl celebrating Christmas with person2.”
- the search results may include more creativity, such as inserting emojis in the descriptive text to result in “personl celebrating vl ⁇ with person 2.”
- the user may specify the level of randomness via a user interface generated by the user interface module 202.
- the template module 208 provides the one or more templates with a specified level of complexity to the machine-learning model 120.
- the level of complexity may be on a scale of numbers (e.g., 0 for most simple, 5 for medium complexity, and 10 for most complex), a scale of words (e.g., most simple, slightly complex, medium complexity, very complex, etc.), or other paradigm.
- the level of complexity is used by the machine-learning model 120 to modify the descriptive text. For example, a low level of complexity may include descriptive text with fewer words added to the template than for higher levels of complexity.
- the level of complexity and the level of randomness may be specified based on user preference (e.g., input via a user interface generated by the user interface module 202), feedback from search results, etc.
- the template module 208 may set the level of complexity to be at the level where the user is most likely to accept the suggested search results.
- the machine-learning model 120 outputs descriptive text.
- the user interface module 202 may provide the descriptive text to the user as a suggested search query in a user interface for the media items associated with the user account, use the descriptive text to evaluate search quality, etc.
- the user interface module 202 receives a selection of the suggested search query from the user via the user interface and provides the user with search results that match the descriptive text from a database associated with the media items associated with the user account, such as a database that is part of the storage device 245.
- the user may provide feedback that is used to refine the user’s preferences for suggested search queries. For example, if a user modifies the suggested search query, the modification may be used by the machine-learning model 120 and/or the template module 208 to improve suggested search queries for the user in the future.
- Figure 5 illustrates two example user interfaces 500, 525 that include suggested search queries, according to some embodiments described herein.
- the user interface module 202 In the first user interface 500, the user interface module 202 generates a list of suggested search queries 502 based on a group of media items 510. The user may enter their search query in the text field 505 based on being inspired by the list of suggested search queries 502, or may choose one of the suggested search queries 502.
- the user interface module 202 In the second user interface 525, the user interface module 202 generates a list of suggestions 531 based on two different groups of media items 527, 529. Each suggestion is associated with a corresponding search button 533, 535, 537 for selecting a particular search suggestion.
- Figure 6 illustrates an example user interface 600 that includes suggested search queries based on an initial search request from a user, according to some embodiments described herein.
- a user provides a request in a text field 602 for media items that include one or more particular attributes.
- the particular attributes are for media items that include the user’s daughter Ava.
- the user interface module 202 provides suggested search queries 605, corresponding search buttons 612, 614 for further refinement of the search results, and search results 610.
- the list of suggestions 605 is based on attributes identified in media items that correspond to the initial search for media items that include the user’s daughter.
- Figure 7 illustrates a flowchart of an example method 700 to generate suggested search queries.
- the method 700 may be performed by the computing device 200 in Figure 2.
- the method 700 is performed by the user device 115, the media server 101, or in part on the user device 115 and in part on the media server 101.
- groups of media items 705 are used by the media application 103 to extract and aggregate attributes 710. Histograms of attributes 715 are created from the attributes and used to generate 725 a template.
- the generated templates 725 (based on various groups of media items) are used by a template generator 730 that provides the templates 725 to a large language model 735.
- the large language model 735 returns natural language style search queries 740.
- the natural language style search queries 740 may be used to evaluate search quality 745. For example, they may be used as model examples of search quality that are used to teach users how to craft search results.
- the natural language style search queries 740 may also be provided as search suggestions to a user 750.
- Figure 8 illustrates a flowchart of another example method 800 to generate suggested search queries.
- the method 800 may be performed by the computing device 200 in Figure 2.
- the flowchart 800 is performed by the user device 115, the media server 101 , or in part on the user device 115 and in part on the media server 101.
- the method 800 of Figure 8 may begin at block 802.
- a request for a suggested search query of a group of media items associated with a user account is received.
- the media items include one or more screenshots captured by a user device associated with the user account.
- Block 802 may be followed by block 804.
- a permission interface element is caused to be displayed.
- a media application 103 may display the permission interface element before satisfying the request for the suggested search query.
- the permission interface element is displayed before the media application displays suggested search queries.
- Block 804 may be followed by block 806.
- block 806 it is determined whether permission of user data is granted by the user. If permission is not granted by the user, block 806 is follow'ed by block 808 where the method 800 ends. If permission is granted, block 806 may be followed by block 810.
- Block 810 atributes from a group of media items associated with a user are identified. In some embodiments, the attributes that are part of the group of media items associated with the user account are precomputed. Block 810 may be followed by block 812.
- Block 812 templates with combinations of the attributes are generated. Block 812 may be followed by block 814.
- each template is scored based on a number of atribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes.
- the method 800 further includes determining a number of times each attribute appears in the media items of the group of media items and scoring the templates is based on the number of times each corresponding attribute appears in the media items.
- the method 800 further includes determining a number of the media items that include the corresponding atributes and for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
- Block 814 may be followed by block 816.
- Block 816 one or more of the templates are selected based on corresponding scores. Block 816 may be followed by block 818.
- one or more of the selected templates are provided as input to a large language model.
- the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
- the one or more templates are provided to the large language model with a level of randomness to be associated with the descriptive text.
- the one or more templates are provided to the large language model with a level of complexity to be associated with the descriptive text. Block 818 may be followed by block 820.
- the large language model outputs descriptive text based on the one or more templates.
- Block 820 may be followed by block 822.
- the descriptive text is provided to the user as a suggested search query in a user interface for the media items associated with the user account.
- the method 800 further includes receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account.
- the method 800 further includes receiving a request from a user for media items that include one or more particular attributes and providing the group of media items to the user based on the group of media items including the one or more particular attributes, where the suggested search query is provided with the group of media items.
- a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server.
- user information e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location
- certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed.
- a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.
- location information such as to a city, ZIP code, or state level
- the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
- the embodiments of the specification can also relate to a processor for performing one or more steps of the methods described above.
- the processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer.
- a computer program may be stored in a non-transitory computer- readable storage medium, including, but not limited to, any type of disk including optical disks, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memories including USB keys with non-volatile memory, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.
- the specification can take the form of some entirely hardware embodiments, some entirely software embodiments or some embodiments containing both hardware and software elements.
- the specification is implemented in software, which includes, but is not limited to, firmware, resident software, microcode, etc.
- the description can take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system.
- a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
- a data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements through a system bus.
- the memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporarystorage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Software Systems (AREA)
- Mathematical Physics (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Medical Informatics (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
A computer-implemented method includes identifying attributes from a group of media items associated with a user account. The method further includes generating one or more templates with the attributes. The method further includes generating templates with combinations of the attributes. The method further includes scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. The method further includes selecting one or more of the templates based on corresponding scores. The method further includes providing the one or more templates as input to a large language model. The method further includes outputting, with the large language model, descriptive text based on the one or more templates. The method further includes providing the descriptive text to the user as a suggested search query in a user interface.
Description
SEARCH QUERY GENERATION USING MACHINE-LEARNING MODEL
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to U.S. Provisional Patent Application No. 63/640,690, filed April 30, 2024 and titled “Search Query Generation Using Machine- Learning Model,” which is incorporated herein in its entirety.
BACKGROUND
[0002] With the proliferation of smartphones, consumers often have thousands of media items stored on their mobile devices or backed up in cloud storage. Users search for photos and videos using basic queries like “March 2024” and “John”; however, due to the overwhelming number of media items that are stored for a user, the search may return too many results for users to find what they are looking for.
[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.
SUMMARY
[0004] A computer-implemented method includes identifying attributes from a group of media items associated with a user account. The method further includes generating one or more templates with the attributes. The method further includes generating templates with combinations of the attributes. The method further includes scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. The method further includes
selecting one or more of the templates based on corresponding scores. The method further includes providing the one or more templates as input to a large language model. The method further includes outputting, with the large language model, descriptive text based on the one or more templates. The method further includes providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
[0005] In some embodiments, the method further includes receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the method further includes determining a number of times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items. In some embodiments, the method further includes determining a number of the media items that include the corresponding attributes; and for each template, calculating a minimum threshold value by dividing a number of categories by the number of the media items in the group of media items, where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and where selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
[0006] In some embodiments, the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof. In some embodiments, the one or more templates are provided to the large language model with a
level of randomness to be associated with the descriptive text. In some embodiments, the one or more templates are provided to the large language model with a level of complexity to be associated with the descriptive text. In some embodiments, the method further includes comprising precomputing the attributes that are part of the group of media items associated with the user account. In some embodiments, the method further includes receiving a request from a user for media items that include one or more particular attributes and providing the group of media items to the user based on the group of media items including the one or more particular attributes, where the suggested search query is provided with the group of media items. In some embodiments, the media items include one or more screenshots captured by a user device associated with the user account.
[0007] A system includes one or more processors and a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the processor, cause the one or more processors to perform operations. The operations include identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
[0008] In some embodiments, the operations further include receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the operations further include determining a number of
times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items. In some embodiments, the operations further include determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score. In some embodiments, the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
[0009] A non-transitory computer-readable medium includes instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations. The operations include identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
[0010] In some embodiments, the operations further include receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the operations further include determining a number of times each attribute appears in the media items of the group of media items, where scoring the templates is based on the number of times each corresponding attribute appears in the media items. In some embodiments, the operations further include determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score. In some embodiments, the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram of an example network environment, according to some embodiments described herein.
[0012] Figure 2 is a block diagram of an example computing device, according to some embodiments described herein.
[0013] Figure 3 illustrates an example user interface that includes a group of media items, according to some embodiments described herein.
[0014] Figure 4 illustrates an example prompt that is submitted with one or more templates to a machine-learning model, according to some embodiments described herein.
[0015] Figure 5 illustrates two example user interfaces that include suggested search queries, according to some embodiments described herein.
[0016] Figure 6 illustrates an example user interface that includes suggested search queries based on an initial search request from a user, according to some embodiments described herein.
[0017] Figure 7 illustrates a flowchart of an example method to generate suggested search queries, according to some embodiments described herein.
[0018] Figure 8 illustrates a flowchart of another example method to generate suggested search queries, according to some embodiments described herein.
DETAILED DESCRIPTION
Overview
[0019] Mobile devices capture photos and videos using increasingly sophisticated techniques.
Users often have thousands of media items stored on their mobile devices or backed up in cloud storage. The media items may comprise any kind of media items, e.g., image(s), documents), video file(s)/data, audio file(s)/data, and/or video and audio file(s)/data. Users search for media items such as photos and videos using basic queries like “March 2024” and “John”; however, due to the overwhelming number of media items that are stored for a user, the search results may return too many results for users to find what they are looking for.
Searches can be improved by adding details to the images (e.g., adding descriptions, tags, etc.), but users may not remember enough about the images to improve their search queries. [0020] The technology described below advantageously uses a large language model (LLM) to output suggested search queries for users for searching media items that guides the users to implement more complex search queries. The complex search queries can result specific results that match the user’s information need. Further, as a result of teaching users to implement more complex search queries, the users become more adept at crafting search queries. Instead of performing multiple searches to find a particular image, the suggested search queries result in using fewer computer resources to obtain their desired search results.
Network Environment
[0021] Figure 1 illustrates a block diagram of an example network environment 100. In some embodiments, the network environment 100 includes a media server 101, and user devices 115 coupled to a network 105. Users 125a, 125n may be associated with respective user devices 115a, 115n. In some embodiments, the network environment 100 may include other servers or devices not shown in Figure 1. In Figure 1 and the remaining figures, a letter after a reference number, e.g., “115a,” represents a reference to the element having that particular reference number. A reference number in the text without a following letter, e.g., “115,” represents a general reference to embodiments of the element bearing that reference number.
[0022] The media server 101 may include a processor, a memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to the network 105 via signal line 102. Signal line 102 may be a wired connection, such as Ethernet, coaxial cable, fiber-optic cable, etc., or a wireless connection, such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more of the user devices 115a, 115n via the network 105.
[0023] The media server 101 may include a media application 103a, a machine-learning model 120, and a database 199. Although the machine-learning model 120 is illustrated as being stored on the same media server 101 as the media application 103a, in some embodiments the machine-learning model 120 is stored on a separate server.
[0024] The machine-learning model 120 is trained to provide text in response to a query. For example, the machine-learning model 120 may be a large language model (LLM) that is designed for natural language processing tasks such as language generation. The LLM is trained to receive one or more templates and a prompt, and output descriptive text based on the one or more templates. The machine-learning model 120 receives the one or more templates with attributes from a media application 103 (e.g., from the media application 103a on the media server, or a media application 103b, 103c stored on a user device 115a, 115b. The machine-learning model 120 outputs descriptive text.
[0025] The trained machine-learning model 120 may include one or more model forms or structures. For example, model forms or structures can include any type of neural-network, such as a linear network, a deep-leaming neural network that implements a plurality of layers (e.g., “hidden layers” between an input layer and an output layer, with each layer being a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural- network layers, and aggregates the results from the processing of each tile), a sequence-to- sequence neural network (e.g., a network that receives as input sequential data, such as words in a sentence, frames in a video, etc. and produces as output a result sequence), etc.
[0026] The model form or structure may specify connectivity between various nodes and organization of nodes into layers. For example, nodes of a first layer (e.g., an input layer) may receive data as input data or application data. Such data can include, for example, one or
more pixels per node, e.g., when the trained model is used for analysis, e.g., of an initial image. Subsequent intermediate layers may receive as input, output of nodes of a previous layer per the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. For example, a first layer may output a segmentation between a foreground and a background. A final layer (e.g., output layer) produces an output of the machine-learning model. In some embodiments, model form or structure also specifies a number and/ or type of nodes in each layer.
[0027] In different embodiments, the trained model can include one or more models. One or more of the models may include a plurality of nodes, arranged into layers per the model structure or form. In some embodiments, the nodes may be computational nodes with no memory, e.g., configured to process one unit of input to produce one unit of output.
Computation performed by a node may include, for example, multiplying each of a plurality of node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output. In some embodiments, the computation performed by a node may also include applying a step/activation function to the adjusted weighted sum. In some embodiments, the step/activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by the plurality of nodes may be performed in parallel, e.g., using multiple processors cores of a multicore processor, using individual processing units of a graphics processing unit (GPU), or special-purpose neural circuitry. In some embodiments, nodes may include memory, e.g., may be able to store and use one or more earlier inputs in processing a subsequent input. For example, nodes with memory may include long short-term memory (LSTM) nodes. LSTM nodes may use the memory to maintain “state” that permits the node to act like a finite state machine (FSM).
[0028] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, a model may be initiated as a plurality of nodes organized into layers as specified by the model form or structure. At initialization, a respective weight may be applied to a connection between each pair of nodes that are connected per the model form, e.g., nodes in successive layers of the neural network. For example, the respective weights may be randomly assigned, or initialized to default values. The model may then be trained, e.g., using training data, to produce a result.
[0029] Training may include applying supervised learning techniques. In supervised learning, the training data can include a plurality of inputs (e.g., templates) and a corresponding groundtruth output for each input (e.g., corresponding descriptive text). Based on a comparison of the output of the model with the groundtruth output, values of the weights are automatically adjusted, e.g., in a manner that increases a probability that the model produces the groundtruth output for the image.
[0030] In various embodiments, a trained model includes a set of weights, or embeddings, corresponding to the model structure. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights. In various embodiments, a trained model includes a set of weights, or embeddings, corresponding to the model structure. In embodiments where data is omitted, the trained model may be is based on prior training, e.g., by a developer, by a third-party, etc. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.
[0031] The database 199 may store machine-learning models, training data sets, media items, etc. The database 199 may also store social network data associated with users 125, user preferences for the users 125, etc.
[0032] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile telephone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing a network 105.
[0033] In the illustrated implementation, user device 115a is coupled to the network 105 via signal line 108 and user device 115n is coupled to the network 105 via signal line 110. The media application 103 may be stored as media application 103b on the user device 115a and/or media application 103c on the user device 115n. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, fiber-optic cable, etc., or wireless connections, such as Wi-Fi®, Bluetooth®, or other wireless technology. User devices 115a, 115n are accessed by users 125a, 125n, respectively. The user devices 115a, 115n in Figure 1 are used by way of example. While Figure 1 illustrates two user devices, 115a and 115n, the disclosure applies to a system architecture having one or more user devices 115.
[0034] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some may be performed on the user device 115.
[0035] Operations are performed in accordance with user settings. For example, the user 125a may specify settings that images and/or other data of the user is to be stored only locally on a user device 115a and not on the media server 101.
[0036] Transmission of user data (e.g., templates, attributes, etc.) to the media server 101, any temporary or permanent storage of such data by the media server 101, and performance of operations on such data by the media server 101 are performed only if the user has agreed
to transmission, storage, and performance of operations by the media server 101. User data from media is not used in advertisements, responses are not reviewed by humans unless the user provides feedback or to address abuses or harm, and the user data is not used to train machine-learning models outside of the machine-learning models used to provide suggested search queries. Users are provided with options to change the settings at any time, e.g., such that they can enable or disable the use of the media server 101.
[0037] In some embodiments, machine learning models (e.g., neural networks or other types of models), if utilized for one or more operations, are stored and utilized locally on a user device 115, with specific user permission. Server-side models are used only if permitted by the user. Further, a trained model may be provided for use on a user device 115. During such use, if permitted by the user 125, on-device training of the model may be performed. Updated model parameters may be transmitted to the media server 101 if permitted by the user 115, e.g., to enable federated learning. Model parameters do not include any user data.
[0038] The media application 103 identifies attributes from a group of media items associated with a user account. Attributes of media items (e.g., image(s), documents), video file(s)/data, audio file(s)/data, and/or video and audio file(s)/data ) may comprise, for example, one or more of objects, people, landmarks, actions, etc. The attributes may be identified by performing independent component analysis.
[0039] The media application 103 generates templates with combinations of the attributes. A template may be a set of attributes. For example, a first template may include {Sara, John, climbing} and a second template may include {John, eating} .
[0040] The media application 103 scores each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. For example, the number of attribute categories
for the first template is three (i.e., Sara, John, and climbing). In one embodiment, generating the score may include calculating a minimum threshold value based on the number of attribute categories in the corresponding template and multiplying the number of media items by a corresponding minimum threshold value. For example, for the first template, the minimum threshold value may be 0.66 based on the number of attribute categories being 3 and the score is a product of the minimum threshold value and the number of media items, which is 0.66*4 = 2.64. For the second template, the number of attribute categories for the second template is two (i.e., John and eating) and the minimum threshold value is 0.5, which results in a score of the product of the minimum threshold value and the number of media items: 0.5*1 = 0.5.
[0041] The media application 103 selects one or more of the templates based on the corresponding scores. In the first example, the score of 2.64 results in needing at least three of the media items to include Sara, John, and climbing. In the second example, the score of 0.5 results in needing at least one of the media items to include John and eating. Because only 2 images include Sara, John, and climbing and one of the images includes John and eating, the media application 103 selects the second template.
[0042] The media application 103 provides one or more templates as input to the machinelearning model 120. The machine-learning model 120 outputs descriptive text based on the one or more templates. For example, the machine-learning model 120 may output “John is eating.” The media application 103 provides the descriptive text to the user as a suggested search query.
[0043] While the foregoing example illustrates a simple query, the queries can be complex in various cases, such as “John eating a hamburger on a sidewalk in Manhattan,” “John eating
shrimp noodles at a Japanese restaurant,” “John eating a taco and holding a soda in his other hand, with a taco stand in the background,” etc.
[0044] In some embodiments, the media application 103 may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), machine learning processor/ coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.
Computing Device
[0045] Figure 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is media server 101 used to implement the media application 103a. In another example, computing device 200 is a user device 115.
[0046] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input/output (I/O) interface 239, a display 241, a camera 243, and a storage device 245 all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via signal line 222, the memory 237 may be coupled to the bus 218 via signal line 224, the I/O interface 239 may be coupled to the bus 218 via signal line 226, the display 241 may be coupled to the bus 218 via signal line 228, the camera 243 may be coupled to the bus 218 via signal line 230, and the storage device 245 may be coupled to the bus 218 via signal line 232.
[0047] Processor 235 can be one or more processors and/or processing circuits to execute program code and control basic operations of the computing device 200. A “processor” includes any suitable hardware system, mechanism or component that processes data, signals or other information. A processor may include a system with a general-purpose central
processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for achieving functionality, a special-purpose processor to implement neural network modelbased processing, neural circuits, processors optimized for matrix computations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more co-processors that implement neural-network processing. In some embodiments, processor 235 may be a processor that processes data to produce probabilistic output, e.g., the output produced by processor 235 may be imprecise or may be accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have temporal limitations. For example, a processor may perform its functions in real-time, offline, in a batch mode, etc. Portions of processing may be performed at different times and at different locations, by different (or the same) processing systems. A computer may be any processor in communication with a memory.
[0048] Memory 237 is typically provided in computing device 200 for access by the processor 235, and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), Electrical Erasable Read-only Memory (EEPROM), Flash memory, etc., suitable for storing instructions for execution by the processor or sets of processors, and located separate from processor 235 and/or integrated therewith. Memory 237 can store software operating on the computing device 200 by the processor 235, including a media application 103.
[0049] The memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 can include, e.g., an image library application, an image management application, an image gallery application, communication applications,
web hosting engines or applications, media sharing applications, etc. One or more methods disclosed herein can operate in several environments and platforms, e.g., as a stand-alone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("app") run on a mobile computing device, etc.
[0050] The application data 266 may be data generated by the other applications 264 or hardware of the computing device 200. For example, the application data 266 may include images used by the image library application and user actions identified by the other applications 264 (e.g., a social networking application), etc.
[0051] I/O interface 239 can provide functions to enable interfacing the computing device 200 with other systems and devices. Interfaced devices can be included as part of the computing device 200 or can be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and/or storage device 245), and input/output devices can communicate via I/O interface 239. In some embodiments, the I/O interface 239 can connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensors, etc.) and/or output devices (display devices, speaker devices, printers, monitors, etc.).
[0052] Some examples of interfaced devices that can connect to I/O interface 239 can include a display 241 that can be used to display content, e.g., images, video, and/or a user interface of an output application as described herein, and to receive touch (or gesture) input from a user. For example, display 241 may be utilized to display a user interface that includes a graphical guide on a viewfinder. Display 241 can include any suitable display device such as a liquid crystal display (LCD), light emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device. For example, display 241 can be a flat display screen provided on a
mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen for a computer device.
[0053] Camera 243 may be any type of image capture device that can capture images and/or video. In some embodiments, the camera 243 captures images or video that the I/O interface 239 transmits to the media application 103.
[0054] The storage device 245 is a database that stores data related to the media application 103. For example, the storage device 245 may store a training data set that includes labeled images, a machine-learning model, output from the machine-learning model, media items, etc.
Media Application
[0055] Figure 2 illustrates an example media application 103, stored in memory 237. In some embodiments, the media application 103 includes a user interface module 202, an event module 204, an attribute module 206, and a template module 208.
[0056] The user interface module 202 generates a user interface that displays information about media items. The media items may include images, videos, screenshots, etc. The media items may be captured by the camera 243 or retrieved from another source. The media items may be displayed chronologically, according to groups created by the event module 204, etc.
[0057] In some embodiments, the user interface module 202 generates a user interface that includes questions about identifying people and/or pets that appear in media items. For example, the user interface may ask a user to identify a name of a person in an image and a relationship of the person to the user. The user interface module 202 may ask for the identification of the people and pets based on the people and/or pets being in a predetermined
number of the images (e.g., 10%). The user interface module 202 asks permission to make use of user data before asking for an identity of people and/or pets in the media items, unless such permission has previously been provided by the user.
[0058] The user interface includes an option for performing a search of media items. In some embodiments, the user interface generates suggested search queries for a user. The user interface module 202 requests user permission before generating the suggested search queries. If the user does not provide permission for the use of their user data, the media application 103 does not generate suggested search queries.
[0059] The event module 204 generates groups of media items. The groups may be based on different factors, such as periodic grouping of media items (e.g., monthly, weekly, yearly, etc.) based on dates associated with the media items (e.g., capture data, last modified date). The event module 204 may generate groups of media items based on different events. For example, the event module 204 may generate a group of media items based on a trip (e.g., trip to London), an event (e.g., a wedding), a person or pet (e.g., a user’s child through the years), a theme (e.g., beach adventures, running trails during the seasons, evolution of a home renovation projection, etc.).
[0060] In some embodiments, the event module 204 outputs a group of media items based on a user request. For example, a user may request media items associated with a particular person. The event module 204 generates a group that includes all media items that include the particular person (e.g., images or videos that depict the person).
[0061] In some embodiments, the event module 204 includes a machine-learning model that is trained to receive media items as input and outputs groups of media items based on different events. The machine-learning model may include different types of machinelearning models, such as the examples mentioned above with reference to the machine-
learning model in Figure 1. In some embodiments, the event machine-learning model is a classifier that receives media items along with information about the media items and uses the information to output an event signal that indicates a likelihood that an event occurred. The information may include the results of performing optical character recognition on the images to identify text within the image that is indicative of a particular event. For example, an image of a menu captured at dinner may include the term “wedding” on it. In some embodiments, the information may include the results of performing object recognition to identify objects associated with an event. For example, media items captured at a baby shower may have presents associated with babies.
[0062] In some embodiments, metadata associated with a media item may be provided as additional input to the machine-learning model, if the user permits such use of metadata. Metadata may include user-permitted factors such as a location and/or a time of capture of a video; whether a video was shared via a social network, an image sharing application, a messaging application, etc.; depth information associated with one or more video frames; sensor values of one or more sensors of a camera that captured the video, e.g., accelerometer, gyroscope, light sensor, or other sensors; an identity of the user (if user consent has been obtained), etc. For example, if the video was captured at night in an outdoor location with the camera pointing upwards, such metadata may indicate that the camera was pointed to the sky at the time of capture of the video, and therefore, is associated with an astronomical event.
[0063] In some embodiments, the event module 204 trains a machine-learning model to identify groups of images based on prediction of an event. In some embodiments, the event machine-learning model may use a combination of metadata, optical character recognition, and other signals as inputs to the event machine-learning model. For example, the metadata may indicate that a person is using a firecracker and the date is July 4th, which results in the
event machine-learning model outputting an event signal that corresponds to Independence Day. In some embodiments, the event machine-learning model also outputs an event type for one or more of the events. Continuing with the example above, the event machine-learning model outputs the event signal and a likelihood that the event is Independence Day or a holiday.
[0064] The attribute module 206 identifies attributes from a group of media items associated with a user account. The attributes may include people and pets in the media item (e.g., depicted in pixels of an image media item, described in text or attributes of a text media item, etc.), a location where the media item was captured, a time when the media item was taken, objects in the media item, text in the media item, events or activities associated with the media item, and/or an action performed in the media item (e.g., eating, swimming, running, etc.). In some embodiments, one or more attributes are determined from labels provided by a user, such as when a user identifies different people in an image. The attribute module 206 may use the attributes in the media items to infer an identification of other attributes. For example, an image of a user in front of a restaurant with the name of the restaurant may be used by the attribute module 206 to identify an event that includes eating or an action that includes eating and a setting of a restaurant.
[0065] In some embodiments, the attribute module 206 performs independent component analysis of the media items to identify objects. Independent component analysis is a technique used to separate mixed signals into their independent components. In some embodiments, the attribute module 206 performs object recognition to identify objects in media items.
[0066] Responsive to obtaining user consent, the attribute module 206 may precompute attributes for a group of media items. For example, the attribute module 206 may
precompute attributes for a media item responsive to a user creating a photo album, responsive to the attribute module 206 generating a particular group of media items (e.g., a group of media items for a theme, such as hiking, a vacation, an event, etc.), after every month, etc. Precomputing attributes advantageously reduces the time between a user requesting suggested search queries and receiving the suggested search queries generated by media application 103.
[0067] In some embodiments, the attribute module 206 generates a histogram that includes an identification of the attributes in each media item and a number of times an attribute appears in the group of media items. Table 1 includes an example of the attributes identified from six media items and Table 2 includes a count of the number of times each attribute appears in all six media items.
[0068] Table 1: Attributes Extracted from a Group of Media Items
[0069] Table 2: Count of the Attributes in the Group of Media Items
[0070] The template module 208 generates templates with combinations of the attributes. In some embodiments, the template module 208 generates the templates using a fixed set of attribute categories. The attribute categories may include different combinations of attribute categories, such as PERSON, PLACE, DATE, EVENT, ACTIVITY, and SCENE. The template module 208 may combine the attribute categories in particular combinations that are derived from common user search queries. For example, the template may include “PERSON, PERSON, ACTIVITY,” “PERSON, PERSON, EVENT,” “EVENT, DATE,” “PERSON, ACTIVITY, DATE,” etc. Continuing with the example from the tables above, the template module 208 generates a template with the attributes by populating a template of {PERSON, PERSON, EVENT} to {“personl”, “person2”, “skiing”}.
[0071] In some embodiments, the template module 208 calculates a minimum threshold value based on a number of attribute categories in a corresponding template. For example, a minimum threshold for two attribute category categories is 0.5, a minimum threshold for three attribute category categories is 0.66, and a minimum threshold for four attribute category categories is 0.75.
[0072] In some embodiments, the minimum threshold value is calculated using the following equation:
[0074] Where N is the number of categories and N > 2. The equation for the minimum threshold uses the pigeonhole principle to ensure that there will be at least one media item that matches the generated combination. In some embodiments, 2 < N < 4.
[0075] The template module 208 scores each template based on the number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. If the number of media items that include the corresponding attributes exceed the corresponding score, the template module 208 selects the template. In some embodiments, if multiple people are included in a template, the multiple people do not need to be included in the same media item to be considered in the number of media items with an attribute.
[0076] Figure 3 illustrates an example user interface 300 of a group of media items, according to some embodiments described herein. The group of media items include a first image 305 of person 1 skiing, a second image 307 of person 1 and person 2 skiing, a third image 309 of person 2 climbing, a fourth image 311 of person 2 skiing while person 1 is present along with two other spectators, a fifth image 313 of person 1 kayaking, and a sixth image 315 of person 1 and person 2 skiing.
[0077] Continuing with the example above, for the template “PERSON, PERSON, ACTIVITIES,” there are three attribute categories, which corresponds to the minimum threshold 0.66. For a group of six media items, the attributes need to appear in at least four media items because 6*0.66 = 4 media items. Since personl, person2, and skiing appear in four media items (305, 307, 311 , 315), they are selected as a template {“personl ”, “person2”, “skiing”}. Since climbing and kayaking only appear in one media item, {“personl”, “climbing”} and {“person2”, “kayaking”} are not selected as templates.
[0078] The template module 208 provides the one or more templates as input to the machinelearning model 120. Continuing with the example above, the machine-learning model 120 receives {“personl”, “person2”, “skiing”} as input and provides “personl and person2 skiing together” as output. In other examples, {“person2”, “climbing”, “December,” “2023} is provided as input and the machine-learning model 120 outputs “person2 climbing in December 2023”; {“personl”, “person2”, “wedding”} is provided as input and the machinelearning model 120 outputs “personl and person2 during their wedding”; and {“personl”, “person2”, “Christmas”} is provided as input and the machine-learning model 120 outputs “personl celebrating Christmas with person2.”
[0079] In some embodiments, the template module 208 provides the one or more templates along with a prompt where the prompt is based on at least one feature selected from a group of: emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, and delimiters. Figure 4 illustrates an example prompt 400 that is submitted with one or more templates to the machine-learning model 120, according to some embodiments described herein.
[0080] In some embodiments, the template module 208 provides the one or more templates with a level of randomness to the machine-learning model 120. The level of randomness may be on a scale of numbers (e.g., 0 for no randomness, 5 for medium randomness, and 10 for most random), a scale of words (e.g., no randomness, a little randomness, medium randomness, very random, etc.), or other paradigm. The level of randomness is used by the machine-learning model 120 to modify the descriptive text. For example, a low level of randomness may result in concatenating the search results in a natural language style, such as the example above of “personl celebrating Christmas with person2.” In an example with a high level of randomness, the search results may include more creativity, such as inserting emojis in the descriptive text to result in “personl celebrating vl < with person 2.” In
some embodiments, the user may specify the level of randomness via a user interface generated by the user interface module 202.
[0081] In some embodiments, the template module 208 provides the one or more templates with a specified level of complexity to the machine-learning model 120. The level of complexity may be on a scale of numbers (e.g., 0 for most simple, 5 for medium complexity, and 10 for most complex), a scale of words (e.g., most simple, slightly complex, medium complexity, very complex, etc.), or other paradigm. The level of complexity is used by the machine-learning model 120 to modify the descriptive text. For example, a low level of complexity may include descriptive text with fewer words added to the template than for higher levels of complexity. The level of complexity and the level of randomness may be specified based on user preference (e.g., input via a user interface generated by the user interface module 202), feedback from search results, etc. In some embodiments, if a user rejects suggested search results with high complexity, but accepts suggested search results with lower complexity, the template module 208 may set the level of complexity to be at the level where the user is most likely to accept the suggested search results.
[0082] The machine-learning model 120 outputs descriptive text. The user interface module 202 may provide the descriptive text to the user as a suggested search query in a user interface for the media items associated with the user account, use the descriptive text to evaluate search quality, etc. In some embodiments, the user interface module 202 receives a selection of the suggested search query from the user via the user interface and provides the user with search results that match the descriptive text from a database associated with the media items associated with the user account, such as a database that is part of the storage device 245.
[0083] In some embodiments, the user may provide feedback that is used to refine the user’s preferences for suggested search queries. For example, if a user modifies the suggested search query, the modification may be used by the machine-learning model 120 and/or the template module 208 to improve suggested search queries for the user in the future.
Example User Interfaces
[0084] Figure 5 illustrates two example user interfaces 500, 525 that include suggested search queries, according to some embodiments described herein. In the first user interface 500, the user interface module 202 generates a list of suggested search queries 502 based on a group of media items 510. The user may enter their search query in the text field 505 based on being inspired by the list of suggested search queries 502, or may choose one of the suggested search queries 502.
[0085] In the second user interface 525, the user interface module 202 generates a list of suggestions 531 based on two different groups of media items 527, 529. Each suggestion is associated with a corresponding search button 533, 535, 537 for selecting a particular search suggestion.
[0086] Figure 6 illustrates an example user interface 600 that includes suggested search queries based on an initial search request from a user, according to some embodiments described herein. In this example, a user provides a request in a text field 602 for media items that include one or more particular attributes. In this case the particular attributes are for media items that include the user’s daughter Ava.
[0087] The user interface module 202 provides suggested search queries 605, corresponding search buttons 612, 614 for further refinement of the search results, and search results 610.
The list of suggestions 605 is based on attributes identified in media items that correspond to the initial search for media items that include the user’s daughter.
Example Methods
[0088] Figure 7 illustrates a flowchart of an example method 700 to generate suggested search queries. The method 700 may be performed by the computing device 200 in Figure 2. In various embodiments, the method 700 is performed by the user device 115, the media server 101, or in part on the user device 115 and in part on the media server 101.
[0089] In Figure 7, groups of media items 705 are used by the media application 103 to extract and aggregate attributes 710. Histograms of attributes 715 are created from the attributes and used to generate 725 a template. The generated templates 725 (based on various groups of media items) are used by a template generator 730 that provides the templates 725 to a large language model 735. The large language model 735 returns natural language style search queries 740. The natural language style search queries 740 may be used to evaluate search quality 745. For example, they may be used as model examples of search quality that are used to teach users how to craft search results. The natural language style search queries 740 may also be provided as search suggestions to a user 750.
[0090] Figure 8 illustrates a flowchart of another example method 800 to generate suggested search queries. The method 800 may be performed by the computing device 200 in Figure 2. In various embodiments, the flowchart 800 is performed by the user device 115, the media server 101 , or in part on the user device 115 and in part on the media server 101.
[0091] The method 800 of Figure 8 may begin at block 802. At block 802, a request for a suggested search query of a group of media items associated with a user account is received. In some embodiments, the media items include one or more screenshots captured by a user device associated with the user account. Block 802 may be followed by block 804.
[0092] At block 804, a permission interface element is caused to be displayed. For example, a media application 103 may display the permission interface element before satisfying the
request for the suggested search query. In some embodiments where the suggested search query is provided without a user requesting the suggested search query, the permission interface element is displayed before the media application displays suggested search queries. Block 804 may be followed by block 806.
[0093] At block 806 it is determined whether permission of user data is granted by the user. If permission is not granted by the user, block 806 is follow'ed by block 808 where the method 800 ends. If permission is granted, block 806 may be followed by block 810.
[0094] At block 810, atributes from a group of media items associated with a user are identified. In some embodiments, the attributes that are part of the group of media items associated with the user account are precomputed. Block 810 may be followed by block 812.
[0095] At block 812, templates with combinations of the attributes are generated. Block 812 may be followed by block 814.
[0096] At block 814, each template is scored based on a number of atribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes. Tn some embodiments, the method 800 further includes determining a number of times each attribute appears in the media items of the group of media items and scoring the templates is based on the number of times each corresponding attribute appears in the media items. In some embodiments, the method 800 further includes determining a number of the media items that include the corresponding atributes and for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; where scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value and selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media
items that include the corresponding attributes exceeding a corresponding score. Block 814 may be followed by block 816.
[0097] At block 816, one or more of the templates are selected based on corresponding scores. Block 816 may be followed by block 818.
[0098] At block 818, one or more of the selected templates are provided as input to a large language model. In some embodiments, the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof. In some embodiments, the one or more templates are provided to the large language model with a level of randomness to be associated with the descriptive text. In some embodiments, the one or more templates are provided to the large language model with a level of complexity to be associated with the descriptive text. Block 818 may be followed by block 820.
[0099] At block 820, the large language model outputs descriptive text based on the one or more templates. Block 820 may be followed by block 822.
[00100] At block 822, the descriptive text is provided to the user as a suggested search query in a user interface for the media items associated with the user account. In some embodiments, the method 800 further includes receiving a selection of the suggested search query from the user via the user interface and providing the user with search results that match the descriptive text from a database of the media items associated with the user account. In some embodiments, the method 800 further includes receiving a request from a user for media items that include one or more particular attributes and providing the group of media items to the user based on the group of media items including the one or more
particular attributes, where the suggested search query is provided with the group of media items.
[00101] Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[00102] In the above description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. It will be apparent, however, to one skilled in the art that the disclosure can be practiced without these specific details. In some instances, structures and devices are shown in block diagram form in order to avoid obscuring the description. For example, the embodiments can be described above primarily with reference to user interfaces and particular hardware. However, the embodiments can apply to any type of computing device that can receive data and commands, and any peripheral devices providing sendees.
[00103] Reference in the specification to “some embodiments” or “some instances” means that a particular feature, structure, or characteristic described in connection with the
embodiments or instances can be included in at least one implementation of the description. The appearances of the phrase “in some embodiments” in various places in the specification are not necessarily all referring to the same embodiments.
[00104] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing art s to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.
[00105 ] I It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms including “processing” or “computing” or “calculating” or “determining” or “displaying” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission, or display devices.
[00106] The embodiments of the specification can also relate to a processor for performing one or more steps of the methods described above. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer- readable storage medium, including, but not limited to, any type of disk including optical disks, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memories including USB keys with non-volatile memory, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.
[00107] The specification can take the form of some entirely hardware embodiments, some entirely software embodiments or some embodiments containing both hardware and software elements. In some embodiments, the specification is implemented in software, which includes, but is not limited to, firmware, resident software, microcode, etc.
[00108] Furthermore, the description can take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[00109] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporarystorage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution.
Claims
1. A computer-implemented method comprising: identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
2. The method of claim 1, further comprising: receiving a selection of the suggested search query from the user via the user interface; and providing the user with search results that match the descriptive text from a database of the media items associated with the user account.
3. The method of claim 1, further comprising: determining a number of times each attribute appears in the media items of the group of media items;
wherein scoring the templates is based on the number of times each corresponding attribute appears in the media items.
4. The method of claim 1, further comprising: determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; wherein scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value; and wherein selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
5. The method of claim 1 , wherein the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
6. The method of claim 1, wherein the one or more templates are provided to the large language model with a level of randomness to be associated with the descriptive text.
7. The method of claim 1, wherein the one or more templates are provided to the large language model with a level of complexity to be associated with the descriptive text.
8. The method of claim 1, further comprising precomputing the attributes that are part of the group of media items associated with the user account.
9. The method of claim 1, further comprising: receiving a request from a user for media items that include one or more particular attributes; and providing the group of media items to the user based on the group of media items including the one or more particular attributes, wherein the suggested search query is provided with the group of media items.
10. The method of claim 1, wherein the media items include one or more screenshots captured by a user device associated with the user account.
11. A system, comprising: one or more processors; and a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the processor, cause the one or more processors to perform operations comprising: identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model;
outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
12. The system of claim 11 , wherein the operations further include: receiving a selection of the suggested search query from the user via the user interface; and providing the user with search results that match the descriptive text from a database of the media items associated with the user account.
13. The system of claim 11 , wherein the operations further include: determining a number of times each attribute appears in the media items of the group of media items; wherein scoring the templates is based on the number of times each corresponding attribute appears in the media items.
14. The system of claim 11 , wherein the operations further include: determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; wherein scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value; and wherein selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number
of the media items that include the corresponding attributes exceeding a corresponding score.
15. The system of claim 11 , wherein the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
16. A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising: identifying attributes from a group of media items associated with a user account; populating templates with combinations of the attributes; scoring each template based on a number of attribute categories in a corresponding template and a number of media items in the group of media items that include corresponding attributes; selecting one or more of the templates based on corresponding scores; providing the one or more templates as input to a large language model; outputting, with the large language model, descriptive text based on the one or more templates; and providing the descriptive text to a user as a suggested search query in a user interface for the media items associated with the user account.
17. The computer-readable medium of claim 16, wherein the operations further comprise: receiving a selection of the suggested search query from the user via the user interface; and providing the user with search results that match the descriptive text from a database of the media items associated with the user account.
18. The computer-readable medium of claim 16, wherein the operations further include: determining a number of times each attribute appears in the media items of the group of media items; wherein scoring the templates is based on the number of times each corresponding attribute appears in the media items.
19. The computer-readable medium of claim 16, wherein the operations further include: determining a number of the media items that include the corresponding attributes; for each template, calculating a minimum threshold value based on the number of attribute categories in the corresponding template; wherein scoring each template includes, for each template, multiplying the number of media items by a corresponding minimum threshold value; and wherein selecting the one or more of the templates based on the corresponding scores includes selecting the one or more of the templates responsive to the number of the media items that include the corresponding attributes exceeding a corresponding score.
20. The computer-readable medium of claim 16, wherein the one or more templates are provided with a prompt to the large language model and the prompt is based on a feature
selected from a group of emphasis through capitalization, emphasis through repetition, iterative prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463640690P | 2024-04-30 | 2024-04-30 | |
| PCT/US2025/027092 WO2025231129A1 (en) | 2024-04-30 | 2025-04-30 | Search query generation using machine-learning model |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4681093A1 true EP4681093A1 (en) | 2026-01-21 |
Family
ID=95895717
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP25728594.0A Pending EP4681093A1 (en) | 2024-04-30 | 2025-04-30 | Search query generation using machine-learning model |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4681093A1 (en) |
| KR (1) | KR20250169271A (en) |
| CN (1) | CN121241339A (en) |
| WO (1) | WO2025231129A1 (en) |
-
2025
- 2025-04-30 WO PCT/US2025/027092 patent/WO2025231129A1/en active Pending
- 2025-04-30 EP EP25728594.0A patent/EP4681093A1/en active Pending
- 2025-04-30 KR KR1020257036890A patent/KR20250169271A/en active Pending
- 2025-04-30 CN CN202580002518.5A patent/CN121241339A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| KR20250169271A (en) | 2025-12-02 |
| WO2025231129A1 (en) | 2025-11-06 |
| CN121241339A (en) | 2025-12-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11537884B2 (en) | Machine learning model training method and device, and expression image classification method and device | |
| US12437358B2 (en) | Performing segmentation of objects in media items based on user input | |
| CN112632403B (en) | Recommendation model training method, recommendation method, device, equipment and medium | |
| KR102858127B1 (en) | Electronic apparatus and controlling method thereof | |
| CN114564666B (en) | Encyclopedia information display method, device, equipment and medium | |
| US20180181592A1 (en) | Multi-modal image ranking using neural networks | |
| US10248847B2 (en) | Profile information identification | |
| US20140280093A1 (en) | Social entity previews in query formulation | |
| CN108292309A (en) | Identify content items using deep learning models | |
| CN115098644A (en) | Image and text matching method and device, electronic equipment and storage medium | |
| CN115244527A (en) | Cross example SOFTMAX and/or cross example negative mining | |
| US12340584B2 (en) | Automatic generation of events using a machine-learning model | |
| WO2024233814A1 (en) | Prompt-driven image editing using machine learning | |
| WO2023069330A1 (en) | Searching for products through a social media platform | |
| JP7843895B2 (en) | Determining the visual theme within the media item collection | |
| US12079290B2 (en) | Systems and methods for a customized search platform | |
| US20220405813A1 (en) | Price comparison and adjustment application | |
| EP4681093A1 (en) | Search query generation using machine-learning model | |
| US12008057B2 (en) | Determining a visual theme in a collection of media items | |
| US20240273155A1 (en) | Photo location destinations systems, methods, and computer readable media | |
| KR20250029021A (en) | Generating images for video communication sessions | |
| WO2022240443A1 (en) | Automatic generation of events using a machine-learning model |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251014 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |