WO2022214616A1 - Personalizing audio-visual content based on user's interest - Google Patents
Personalizing audio-visual content based on user's interest Download PDFInfo
- Publication number
- WO2022214616A1 WO2022214616A1 PCT/EP2022/059319 EP2022059319W WO2022214616A1 WO 2022214616 A1 WO2022214616 A1 WO 2022214616A1 EP 2022059319 W EP2022059319 W EP 2022059319W WO 2022214616 A1 WO2022214616 A1 WO 2022214616A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- features
- task
- feature
- audio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/167—Audio in a user interface, e.g. using voice commands for navigating, audio feedback
Definitions
- the present disclosure generally relates to Artificial Intelligence and more particularly to techniques to provide an audio-visual content based on user’s interest and availability of corresponding computing resources.
- loT uses embedded sensors, software, user devices and exchange of data over the Internet to allow for the development of home automation devices.
- Home automation devices also known as domotics, enable automation building for a home, a car, a television set or the like. These automations allow for what is come to be known as smart homes, smart cars and the like. In these scenarios the automation system monitors and/or controls an environment’s attributes such as lighting, climate, entertainment systems, and appliances. It may also include security systems that provide access control and alarm systems. When connected with the Internet, home devices are an important constituent of the loT.
- a method and device are provided for receiving a user input for performing a task.
- one or more features associated with the task are obtained according to at least one feature associated with the user input or associated with a previous user history. Resource availability is then determined for the performance of the task and its completion. The task is then performed and one or more outputs are provided associated with the task completion to the user.
- FIG. 1 is a traditional AVQA system
- FIG. 2 is an AVQA system having personalized content as per one embodiment
- FIG. 3 is an illustration of personalized framework intelligent network according to one embodiment
- FIG. 4 is an illustration of an application transformation example according to one embodiment
- FIG. 5 is an illustration of a three different computational implementation scenarios according to one embodiment
- FIG. 6 is an illustration of a flow chart according to one embodiment
- FIG. 7 is a schematic illustration of a general overview of an encoding and decoding system according to one or more embodiments;
- FIG. 8 is another illustration of personalized framework intelligent network according to one embodiment;
- FIG. 9 is a further illustration of personalized framework intelligent network according to one embodiment.
- FIG. 10 is a further illustration of personalized framework intelligent network according to one embodiment.
- FIG. 11 is an illustration of a flow chart according to one embodiment.
- FIG 1 is an illustration of a traditional audio-visual question and answer enabled system (AVQA) used in many of the prior art.
- AVQA audio-visual question and answer enabled system
- These systems often allow the user to interact with search engines through the cloud or similar networks.
- the user A1 interacts orally through his/her device A2 with a network such as the cloud A4 and the information is gathered through a set of questions or demands/commands A3 initiated by the user.
- the interaction may be limited to speech or may include other or additional forms such as text, video and the like in different settings.
- the user communication is then sent through the network and/or the cloud to be analyzed as shown at A5. This can entail a speech or visual analysis step and another search and recovery stage for the relevant response to be to be returned to the user.
- the response A6, is often sent to the user device A2 to be shown to the user A1 and often matches the same form of communication (text, video, speech etc.) as used originally by the user A1.
- the request may then be processed in two different manners.
- the request may be processed in the usual manner as discussed in relation with Figure 1 .
- additional analysis may be performed at B6 to detect interest or other components that may provide further information, for example, about true intent of the user. This may include determining certain features, like user’s previous history, user’s intonation, previous requests, which may determine the context of new user request and the like in determining additional information as related to the request as shown at B7.
- the input from B3 and B7 may be then taken together to provide for a more personalized parameters for a search at B4 and the output of the search shown at B5 may reflect the more personalized output that may (e.g., ultimately) be provided to the user.
- Figures 3 and 8 are an illustration of one framework as per one embodiment. The embodiments of Figures 3 and 8 are provided as an example with the understanding that alternate embodiments can be provided as appreciated by those skilled in the art.
- the personalization module in Figures 3 and 8 may comprise (e.g., consists of) several machine learning models and may take as input(s) any of: the original response content D, the feature set (F) and the user interest level (I), and the available processing resources (R P ) to transform the audio-visual content so as to increase user’s interest.
- the level of global user interest may be tracked, for example, based on the analysis of speech and/or video (block 40); and (3) Machine learning models (for speech and/or avatar personalization) may take into account any of: the list of features and user interests (F, I), the available resources for personalization device (Rp) to transform the responding data (D) (block 2000).
- Figure 3 shows two separate sides - shown by numerals 340 and 350 respectively.
- computational features relating to user interest detection is provided on the left side referenced by numerals 340.
- the right side, denoted by numerals 350 further provide content personalization.
- machine learning tools ML can be further used as models to extract more information to aid further personalization.
- Block 100 - this is an interaction device and can include range of devices such as smart TVs, smart phones and other smart devices like Amazon Echo or Alexa and Google Home devices. These devices, at the minimum can record user’s speech/video and/or display/illustrate and/or show the content in form of text, speech, audio of other sorts or video to a user. [0027] On the left right side of Figure 3, at 340, these elements provide:
- Block 10 - this device may be a computation device that may be used to compute the set of features utilized for the content personalization and/or to detect the user’s interest level.
- This device can include any interactive device (block 100), or can include other more sophisticated devices such as an artificial (Al) hub, or a network or cloud service.
- Al artificial
- Block 30 This block may provide a functional analyses component. This can include a number of available resources, including memory and/or components that provide processing power. This components (e.g., often) may be in charge of communicating the results and can include the feature extraction part and/or the personalization part. These resources can be (e.g., easily) checked by available tool/function in each device.
- Such features can be e.g. speech ascent, timbre (for speech); face landmark, face emotion (for face); gesture (for hand, body).
- timbre for speech
- face landmark for face
- gesture for hand, body.
- Block 40 This function may analyze the recorded speech and/or video from the user to output the interest level associated with each feature.
- An example of implementation is a machine learning model trained in a supervised or weakly supervised manner with an annotated audio/video dataset containing various user’s audio/video documents and the corresponding interest level. Once the model is trained, given an input audio/video recording of the user, the ML model will predict the corresponding interest level.
- Another example of implementation may be a simple manual setting of rules by experts: e.g. when the user face is detected as sad and speech is slow, the interest is set to low.
- a variety of tools can be used for human emotion and interest detection function that may be based on audio-visual features.
- a global interest level can be also inferred based on an aggregation technique.
- feature fusion technique known in the prior art such as max-pooling, averaging, weighted-sum, self- attention mechanism, etc.
- Block 20 Personalization device is used to perform content transformation using machine learning models. This device can be the interactive device (block 100), Al hub (for example in block 10), or cloud service (for example in block 200).
- Block 2000 This block may contain one or several ML models and/or signal processing (SP) algorithms which may perform a certain type of speech and/or video style transformation/manipulation, for example, in order to personalize the respond content (D) coming from the block 200.
- SP signal processing
- this block may be a list of rules, for example, manually set by experts: e.g. when the user is detected to be happy and interested, the avatar with more motion, and exciting/faster speech can be used to communicate with the user.
- Block 200 may provide the response (audio/visual) content)
- An example of this may be a transfer style model (originally applied to images, but can also be applied to video and/or speech as shown by many recent works in the domain). Speech style transfer will make the machine’s speech containing speaking style/ascent of the user’s speech (aka personalized text-to-speech function). This may make the user more interested to communicating with the device.
- Style transfer approach exploits a deep neural network (DNN) to generate the output content (e.g. video or speech) D’ , for example, by minimizing the loss containing two term: content loss and style loss :
- the L content may be a function of the original content (D) and the transformed content (D’), which may guarantee the semantic similarity between the transformed content with the original content (e.g. spoken words must be the same).
- L_style may be a function of the feature fi in the feature list F extracted from the user (in the block 1000) and the corresponding feature fi’ extracted from D’.
- Lambda may be a trade-off parameter, the higher lambda, the more style can be transferred in general.
- lambda can be a function of the user’s interest (I) as: starting with a small lambda, if the user interest level increase, lambda can also be increased to transform more style in the content.
- this can be accomplished in a variety of ways. For example, in one scenario, an implementation of audio style transfer using different types of audio features can be used. Similarly, another example can be one of a scenario that uses image style transfer, avatar face manipulation for visual modality.
- Block 300 As speech and visual contents can be modified in the block 2000 separately, they may (e.g., need to) be synchronized and/or combined before being presented to the user. This is a basic signal processing function in the prior art, but it may be used (e.g., necessary) to complete the system. When speech and visual contents are modified jointly in the block 2000, such synchronization processing may not be used (e.g., needed).
- Figure 4 is an illustration of an example provided to aid understanding.
- Figure 4 shows a transformation model where there are more resources available that can help in analysis.
- resources Rp
- the first only feature f1 e.g. speech ascent
- feature f5 e.g. face emotion
- body gesture feature (f6) 430 can also be mimicked (435).
- Figure 5 provides some examples for ease of understanding. These are just a few examples and those skilled in the art appreciate that alternate embodiments provide other and different scenarios and examples.
- a variety of computations of different functions can be shared between one or more edge devices, the cloud, and/or an Al hub.
- feature computation device block 10 from Figures 3 and 8) is the device and the personalization device (block 20 in Figures 3 and 8) is the Al Hub.
- the personalization device block 20 in Figures 3 and 8) is the Al Hub.
- scenario (b) denoted by 520, a case is provided where without Al hub, both response content (D), feature and personalization are done in the cloud.
- Al Hub can be a novel function which can be implemented in home devices like smart TV, gateway, STB, smart assistant, and/or stand-alone devices.
- FIG. 6 is a flowchart illustration of one embodiment.
- a user input for performing a task may be received.
- any features associated with this task performance may be obtained according to at least one feature associated with the user input and/or associated with previous user history.
- resource availability may be determined for performance of the task to optimize its performance completion.
- the task may be performed and any outputs may be provided that are associated to its performance to the user.
- a device can be provided in one embodiment that personalizes the audio visual or other type of content based on the received features and their associated user’s interest.
- the available computational resources for the content adaptation may also be taken into consideration.
- the personalization order may be based on either (a) a pre-defined list of features (e.g. speech modification first, face adaptation second, or the other way around), or (b) the interest level associated with each type of features (e.g. if audio feature-based interest level is higher than emotion-based interest level, then the personalization is done first for the audio).
- the device may generate and/or transmit a list of features and/or their associated interest levels.
- the features and their associated interest levels can be computed locally with audio/video content in captured device or remotely in Al hub or in cloud.
- the device can be an audio/video captured device (like TV, intelligent assistants), Al hub in a gateway, or server in the cloud.
- the user’s interest can be a measure of the engagement, emotion, behavior, or even their combination.
- Figure 7 schematically illustrates a general overview of an encoding and decoding system according to one or more embodiments.
- the system of Figure 7 is configured to perform one or more functions and can have a pre-processing module 700 to prepare a received content (including one more images or videos) for encoding by an encoding device 740.
- the pre-processing module 730 may perform multi-image acquisition, merging of the acquired multiple images in a common space and the like, acquiring of an omnidirectional video in a particular format and other functions to allow preparation of a format more suitable for encoding.
- Another implementation might combine the multiple images into a common space having a point cloud representation.
- Encoding device 740 packages the content in a form suitable for transmission and/or storage for recovery by a compatible decoding device 770.
- the encoding device 740 provides a degree of compression, allowing the common space to be represented more efficiently (i.e., using less memory for storage and/or less bandwidth required for transmission.
- the data is sent to a network interface 750, which may be typically implemented in any network interface, for instance present in a gateway.
- the data can be then transmitted through a communication network 750, such as internet but any other network may be foreseen.
- the data received via network interface 760 may be implemented in a gateway, in a device.
- the data are sent to a decoding device 770.
- Decoded data are then processed by the device 780 that can be also in communication with sensors or users input data.
- the decoder 770 and the device 780 may be integrated in a single device (e.g., a smartphone, a game console, a STB, a tablet, a computer, etc.).
- a rendering device 790 may also be incorporated.
- Figures 9 and 10 provide some examples for ease of understanding. These are just a few examples and those skilled in the art appreciate that alternate embodiments provide other and different scenarios and examples.
- FIG. 11 is a flowchart illustration of one embodiment implemented by a device, for example comprised in any of: (1) an audio video device, (2) an artificial intelligent hub, (3) a gateway and (4) a cloud network.
- a user input for performing a task may be received.
- any features associated with this task performance may be obtained according to at least one feature associated with the user input and/or associated with previous user history.
- resource availability may be determined for performance of the task.
- the task may be performed by at least one machine learning model, based on the one or more available resources and the one or more features associated with the task performance.
- the one or more outputs associated with the task performed by the at least one machine learning model may be transmitted.
- a plurality of features are associated with said task performance and said features are prioritized.
- said prioritization is performed according to a pre-defined list of features.
- said user input has at least one of an audio component and/or video component.
- a user interest level is determined by analyzing said user input and/or said previous user history and said features are prioritized according to said user interest level.
- said user interest level is associated with each type of a feature when a plurality of features is obtained including at least speech and video components.
- a feature of the plurality of features at least includes an audio based feature and an emotion based feature.
- a list of features and associated interest levels of a user are generated, said generated list of features being suitable for storage.
- said generated list of features is transmitted to a user or a different device.
- said user interest level includes at least one of user engagement, user emotion or user behavior.
- said user interest level is determined from one or more of user engagement, user emotion and/or user behavior.
- said resource availability includes amount of available processing power and/or memory availability.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
A method and device are provided for receiving a user input for performing a task. In one embodiment, one or more features associated with the task are obtained according to at least one feature associated with the user input or associated with a previous user history. Resource availability is then determined for the performance of the task and its completion. The task is then performed and one or more outputs are provided associated with the task completion to the user.
Description
PERSONALIZING AUDIO-VISUAL CONTENT BASED ON USER’S INTEREST
TECHNICAL FIELD
[0001] The present disclosure generally relates to Artificial Intelligence and more particularly to techniques to provide an audio-visual content based on user’s interest and availability of corresponding computing resources.
BACKGROUND
[0002] New developing technologies have recently converged to further enable real- time analytics, machine learning, commodity sensors, artificial intelligence, Internet of things (loT) and other techniques that are affecting the quality of everyday life. For example, loT uses embedded sensors, software, user devices and exchange of data over the Internet to allow for the development of home automation devices.
[0003] Home automation devices, also known as domotics, enable automation building for a home, a car, a television set or the like. These automations allow for what is come to be known as smart homes, smart cars and the like. In these scenarios the automation system monitors and/or controls an environment’s attributes such as lighting, climate, entertainment systems, and appliances. It may also include security systems that provide access control and alarm systems. When connected with the Internet, home devices are an important constituent of the loT.
[0004] The desire for having Smart Homes, Smart Cars and devices like Smart TVs have contributed to the popularity of emerging technologies such Google Home and Amazon Echo and Alexa type devices. These devices now allow users to interact with other consumer these devices seamlessly via their voice and gesture, and expect the devices responding to their need (e.g. answering the questions, playing contents)
smartly. In Audio-visual question-answering (AVQA) systems, the device currently offers a fixed avatar with a neutral voice to talk to the user. Information recorded from the user (speech, video) is often sent to the cloud to be processed and a relevant response is sent back to local device and presented to the user via a spoken avatar (audio-visual) or speech (audio only). This poses difficult challenges in understanding the intent of the user. Consequently, an improved system may be needed that can take user feedback into consideration and detect real intent and interest of the user when direct feedback is ambiguous or unclear. SUMMARY
[0005] A method and device are provided for receiving a user input for performing a task. In one embodiment, one or more features associated with the task are obtained according to at least one feature associated with the user input or associated with a previous user history. Resource availability is then determined for the performance of the task and its completion. The task is then performed and one or more outputs are provided associated with the task completion to the user.
BRIEF DESCRIPTION OF THE DRAWINGS [0006] The teachings of the present disclosure can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, in which:
[0007] FIG. 1 is a traditional AVQA system;
[0008] FIG. 2 is an AVQA system having personalized content as per one embodiment;
[0009] FIG. 3 is an illustration of personalized framework intelligent network according to one embodiment;
[0010] FIG. 4 is an illustration of an application transformation example according to one embodiment; [0011] FIG. 5 is an illustration of a three different computational implementation scenarios according to one embodiment;
[0012] FIG. 6 is an illustration of a flow chart according to one embodiment;
[0013] FIG. 7 is a schematic illustration of a general overview of an encoding and decoding system according to one or more embodiments; [0014] FIG. 8 is another illustration of personalized framework intelligent network according to one embodiment;
[0015] FIG. 9 is a further illustration of personalized framework intelligent network according to one embodiment;
[0016] FIG. 10 is a further illustration of personalized framework intelligent network according to one embodiment; and
[0017] FIG. 11 is an illustration of a flow chart according to one embodiment.
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0018] Figure 1 is an illustration of a traditional audio-visual question and answer enabled system (AVQA) used in many of the prior art. These systems often allow the user to interact with search engines through the cloud or similar networks. In the example shown in Figure 1 , the user A1 interacts orally through his/her device A2 with a network such as the cloud A4 and the information is gathered through a set of
questions or demands/commands A3 initiated by the user. The interaction may be limited to speech or may include other or additional forms such as text, video and the like in different settings. The user communication is then sent through the network and/or the cloud to be analyzed as shown at A5. This can entail a speech or visual analysis step and another search and recovery stage for the relevant response to be to be returned to the user. The response A6, is often sent to the user device A2 to be shown to the user A1 and often matches the same form of communication (text, video, speech etc.) as used originally by the user A1.
[0019] The problem with the systems such as the one shown and discussed in Figure 1 is that they do not appreciate user’s feedback entirely. Ambiguities in speech, search requests that may be too vague or too broad and the like can affect a poor response back from the system. Therefore, one challenge is to provide an improved system that takes user feedback based on user’s true intent. This can include deciphering user’s interest based on audio-visual content. In other words, a more personalized user interaction is needed. Another challenge is to efficiently leverage available computing resources to improve the user experience and achieve these improved results when interacting with the user and the user’s smart devices. Figures 2 and 3 provide an improved technique that address the resolution to some of these challenges as per one embodiment. [0020] In Figure 2, once the content of the user request is captured at B1 , it may then be processed in two different manners. In B2 and B3, the request may be processed in the usual manner as discussed in relation with Figure 1 . As shown in an alternate path, additional analysis may be performed at B6 to detect interest or other components that may provide further information, for example, about true intent of the user. This may include determining certain features, like user’s previous history, user’s
intonation, previous requests, which may determine the context of new user request and the like in determining additional information as related to the request as shown at B7.
[0021] The input from B3 and B7 may be then taken together to provide for a more personalized parameters for a search at B4 and the output of the search shown at B5 may reflect the more personalized output that may (e.g., ultimately) be provided to the user.
[0022] Figures 3 and 8 are an illustration of one framework as per one embodiment. The embodiments of Figures 3 and 8 are provided as an example with the understanding that alternate embodiments can be provided as appreciated by those skilled in the art.
[0023] As shown generally, the personalization module in Figures 3 and 8, as will be discussed presently in more detail, may comprise (e.g., consists of) several machine learning models and may take as input(s) any of: the original response content D, the feature set (F) and the user interest level (I), and the available processing resources (RP) to transform the audio-visual content so as to increase user’s interest.
[0024] Before discussing each element more specifically, it should be indicated that in the particular embodiment of Figures 3 and 8:
(1) A list of features F={f1 , f2,...,fn}, for example, used for speech/video transformation may be computed depending on the available resources in the feature device (memory, processing power, etc..), and/or the associated user interest level I={i1 , i2,..., global interest} inferred by a machine learning technique from those features (block 1000);
(2) The level of global user interest may be tracked, for example, based on the analysis of speech and/or video (block 40); and
(3) Machine learning models (for speech and/or avatar personalization) may take into account any of: the list of features and user interests (F, I), the available resources for personalization device (Rp) to transform the responding data (D) (block 2000).
[0025] Now returning to Figure 3, the components can be discussed in more detail. Figure 3 shows two separate sides - shown by numerals 340 and 350 respectively. On the left side referenced by numerals 340, computational features relating to user interest detection is provided. The right side, denoted by numerals 350 further provide content personalization. In the example provided, machine learning tools (ML) can be further used as models to extract more information to aid further personalization. [0026] Looking at Figures 3 and 8 the following blocks can be defined as below:
Block 100 - this is an interaction device and can include range of devices such as smart TVs, smart phones and other smart devices like Amazon Echo or Alexa and Google Home devices. These devices, at the minimum can record user’s speech/video and/or display/illustrate and/or show the content in form of text, speech, audio of other sorts or video to a user. [0027] On the left right side of Figure 3, at 340, these elements provide:
Block 10 - this device may be a computation device that may be used to compute the set of features utilized for the content personalization and/or to detect the user’s interest level. This device can include any interactive device (block 100), or can include other more sophisticated devices such as an artificial (Al) hub, or a network or cloud service.
[0028] Block 30 : This block may provide a functional analyses component. This can include a number of available resources, including memory and/or components that provide processing power. This components (e.g., often) may be in charge of
communicating the results and can include the feature extraction part and/or the personalization part. These resources can be (e.g., easily) checked by available tool/function in each device.
[0029] Block 1000: This block may contain a list of machine learning models and/or signal processing functions which may allow to extract a list of different features F={f1 , f 2 , ... } concerning the user’s speech/face/gesture characteristics. Such features can be e.g. speech ascent, timbre (for speech); face landmark, face emotion (for face); gesture (for hand, body). Prior art approaches in affective computing, multimedia processing, audio processing, computer vision exist for such feature extraction functions.
[0030] Block 40: This function may analyze the recorded speech and/or video from the user to output the interest level associated with each feature. An example of implementation is a machine learning model trained in a supervised or weakly supervised manner with an annotated audio/video dataset containing various user’s audio/video documents and the corresponding interest level. Once the model is trained, given an input audio/video recording of the user, the ML model will predict the corresponding interest level. Another example of implementation may be a simple manual setting of rules by experts: e.g. when the user face is detected as sad and speech is slow, the interest is set to low. [0031] As known to those skilled in the art, a variety of tools can be used for human emotion and interest detection function that may be based on audio-visual features. Besides the local interest (ii) associated with each feature (fi), a global interest level can be also inferred based on an aggregation technique. Such feature fusion technique known in the prior art such as max-pooling, averaging, weighted-sum, self- attention mechanism, etc. The output could be in the form (F, I) = ((f1 , i 1 ), (f2, i2), (f4,
N/A), (N/A, global interest)). N/A may appear when an interest associated with a feature fi is not detected.
[0032] Now looking at the right side of Figure 3, in the 350 part, the following functions can be described as follows: [0033] Block 20: Personalization device is used to perform content transformation using machine learning models. This device can be the interactive device (block 100), Al hub (for example in block 10), or cloud service (for example in block 200).
[0034] Block 2000: This block may contain one or several ML models and/or signal processing (SP) algorithms which may perform a certain type of speech and/or video style transformation/manipulation, for example, in order to personalize the respond content (D) coming from the block 200. In another example of implementation, this block may be a list of rules, for example, manually set by experts: e.g. when the user is detected to be happy and interested, the avatar with more motion, and exciting/faster speech can be used to communicate with the user. [0035] Block 200 : may provide the response (audio/visual) content)
[0036] Each ML model (Mi) or SP algorithm may potentially take as input any of: the original content (D) to be transformed, the available resources (Rp), the corresponding feature fi in the list F={fi}i, and the interest level (ii, global interest), for example, in order to transform the content D so as it has some characteristics similarly to the feature fi. An example of this, may be a transfer style model (originally applied to images, but can also be applied to video and/or speech as shown by many recent works in the domain). Speech style transfer will make the machine’s speech containing speaking style/ascent of the user’s speech (aka personalized text-to-speech function). This may make the user more interested to communicating with the device. Style transfer approach exploits a deep neural network (DNN) to generate the output content
(e.g. video or speech) D’ , for example, by minimizing the loss containing two term: content loss and style loss :
L = L content + lambda*L style
[0037] The L content may be a function of the original content (D) and the transformed content (D’), which may guarantee the semantic similarity between the transformed content with the original content (e.g. spoken words must be the same). L_style may be a function of the feature fi in the feature list F extracted from the user (in the block 1000) and the corresponding feature fi’ extracted from D’. Lambda may be a trade-off parameter, the higher lambda, the more style can be transferred in general.
[0038] In one embodiment, lambda can be a function of the user’s interest (I) as: starting with a small lambda, if the user interest level increase, lambda can also be increased to transform more style in the content. As can be appreciated by those skilled in the art, this can be accomplished in a variety of ways. For example, in one scenario, an implementation of audio style transfer using different types of audio features can be used. Similarly, another example can be one of a scenario that uses image style transfer, avatar face manipulation for visual modality.
[0039] Block 300 : As speech and visual contents can be modified in the block 2000 separately, they may (e.g., need to) be synchronized and/or combined before being presented to the user. This is a basic signal processing function in the prior art, but it may be used (e.g., necessary) to complete the system. When speech and visual contents are modified jointly in the block 2000, such synchronization processing may not be used (e.g., needed).
[0040] Figure 4 is an illustration of an example provided to aid understanding. Figure 4 shows a transformation model where there are more resources available that can
help in analysis. When more resources (Rp) are available, there will be more features that can be transformed as shown. In this example, shown at 410, the first only feature f1 (e.g. speech ascent) may be transformed to generate the content 415 at D’i. When more resources are available, feature f5 (e.g. face emotion) 420 may be transferred in addition to the output content 425 D’2. If there are yet more resources, body gesture feature (f6) 430 can also be mimicked (435).
[0041] Figure 5 provides some examples for ease of understanding. These are just a few examples and those skilled in the art appreciate that alternate embodiments provide other and different scenarios and examples. [0042] In the examples of Figure 5, a variety of computations of different functions can be shared between one or more edge devices, the cloud, and/or an Al hub. In scenario (a), enumerated by 510, feature computation device (block 10 from Figures 3 and 8) is the device and the personalization device (block 20 in Figures 3 and 8) is the Al Hub. [0043] In scenario (b), denoted by 520, a case is provided where without Al hub, both response content (D), feature and personalization are done in the cloud.
[0044] In scenario (c), denoted by 530; feature and personalization are done in the Al hub while response content (D) is from the cloud. Here Al Hub can be a novel function which can be implemented in home devices like smart TV, gateway, STB, smart assistant, and/or stand-alone devices.
[0045] Figure 6 is a flowchart illustration of one embodiment. In block 610 a user input for performing a task may be received. In block 620 any features associated with this task performance may be obtained according to at least one feature associated with the user input and/or associated with previous user history. In block 630, resource availability may be determined for performance of the task to optimize its performance
completion. In block 640 the task may be performed and any outputs may be provided that are associated to its performance to the user.
[0046] In conjunction with Figure 6, a device can be provided in one embodiment that personalizes the audio visual or other type of content based on the received features and their associated user’s interest. In one embodiment the available computational resources for the content adaptation may also be taken into consideration. In one embodiment, the personalization order may be based on either (a) a pre-defined list of features (e.g. speech modification first, face adaptation second, or the other way around), or (b) the interest level associated with each type of features (e.g. if audio feature-based interest level is higher than emotion-based interest level, then the personalization is done first for the audio).
[0047] In one embodiment, the device may generate and/or transmit a list of features and/or their associated interest levels. In addition, the features and their associated interest levels can be computed locally with audio/video content in captured device or remotely in Al hub or in cloud. The device can be an audio/video captured device (like TV, intelligent assistants), Al hub in a gateway, or server in the cloud. In one embodiment, the user’s interest can be a measure of the engagement, emotion, behavior, or even their combination.
[0048] Figure 7 schematically illustrates a general overview of an encoding and decoding system according to one or more embodiments. The system of Figure 7 is configured to perform one or more functions and can have a pre-processing module 700 to prepare a received content (including one more images or videos) for encoding by an encoding device 740. The pre-processing module 730 may perform multi-image acquisition, merging of the acquired multiple images in a common space and the like, acquiring of an omnidirectional video in a particular format and other functions to allow
preparation of a format more suitable for encoding. Another implementation might combine the multiple images into a common space having a point cloud representation. Encoding device 740 packages the content in a form suitable for transmission and/or storage for recovery by a compatible decoding device 770. In general, though not strictly required, the encoding device 740 provides a degree of compression, allowing the common space to be represented more efficiently (i.e., using less memory for storage and/or less bandwidth required for transmission. After being encoded, the data, is sent to a network interface 750, which may be typically implemented in any network interface, for instance present in a gateway. The data can be then transmitted through a communication network 750, such as internet but any other network may be foreseen. Then the data received via network interface 760 may be implemented in a gateway, in a device. After reception, the data are sent to a decoding device 770. Decoded data are then processed by the device 780 that can be also in communication with sensors or users input data. The decoder 770 and the device 780 may be integrated in a single device (e.g., a smartphone, a game console, a STB, a tablet, a computer, etc.). In another embodiment, a rendering device 790 may also be incorporated.
[0049] Figures 9 and 10 provide some examples for ease of understanding. These are just a few examples and those skilled in the art appreciate that alternate embodiments provide other and different scenarios and examples.
[0050] In the examples of Figures 9 and 10, a variety of computations of different functions can be shared between one or more edge devices, the cloud, and/or an Al hub. In Figure 9, feature computation device (block 10 from Figures 3 and 8) and the personalization device (block 20 in Figures 3 and 8) is the Al Hub. In Figure 10, feature
computation device (block 10 from Figures 3 and 8) is the Al Hub and the personalization device (block 20 in Figures 3 and 8) is the cloud server.
[0051] Figure 11 is a flowchart illustration of one embodiment implemented by a device, for example comprised in any of: (1) an audio video device, (2) an artificial intelligent hub, (3) a gateway and (4) a cloud network.. In block 1110 a user input for performing a task may be received. In block 1120 any features associated with this task performance may be obtained according to at least one feature associated with the user input and/or associated with previous user history. In block 1130, resource availability may be determined for performance of the task. In block 1140, the task may be performed by at least one machine learning model, based on the one or more available resources and the one or more features associated with the task performance. In block 1150, the one or more outputs associated with the task performed by the at least one machine learning model may be transmitted.
[0052] According to an embodiment, a plurality of features are associated with said task performance and said features are prioritized. According to an embodiment, wherein said prioritization is performed according to a pre-defined list of features. According to an embodiment, said user input has at least one of an audio component and/or video component. According to an embodiment, a user interest level is determined by analyzing said user input and/or said previous user history and said features are prioritized according to said user interest level. According to an embodiment, said user interest level is associated with each type of a feature when a plurality of features is obtained including at least speech and video components. According to an embodiment, a feature of the plurality of features at least includes an audio based feature and an emotion based feature. According to an embodiment, a list of features and associated interest levels of a user are generated, said generated
list of features being suitable for storage. According to an embodiment, said generated list of features is transmitted to a user or a different device. According to an embodiment, said user interest level includes at least one of user engagement, user emotion or user behavior. According to an embodiment, said user interest level is determined from one or more of user engagement, user emotion and/or user behavior. According to an embodiment, said resource availability includes amount of available processing power and/or memory availability.
[0053] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) as the implementations disclosed. Accordingly, these and other implementations are contemplated by this application.
Claims
1. A method comprising: receiving a user input for performing a task; obtaining one or more features associated with the task performance, wherein at least one feature is associated with the user input or associated with a previous user history; determining one or more resources available for performing the task; performing, by at least one machine learning model, the task based on the one or more available resources and the one or more features associated with the task performance; and transmitting one or more outputs associated with the task performed by the at least one machine learning model.
2. The method of claim 1 , wherein a plurality of features are associated with said task performance and said features are prioritized.
3. The method of claim 2, wherein said prioritization is performed according to a pre-defined list of features.
4. The method of any of claims 1-3, wherein said user input has at least one of an audio component and/or video component.
5. The method of any of claims 1-4, wherein a user interest level is determined by analyzing said user input and/or said previous user history and said features are prioritized according to said user interest level.
6. The method of claim 5, wherein said user interest level is associated with each type of a feature when a plurality of features is obtained including at least speech and video components.
7. The method of claim 6, wherein a feature of the plurality of features at least includes an audio based feature and an emotion based feature.
8. The method of any of claims 1-7, wherein a list of features and associated interest levels of a user are generated, said generated list of features being suitable for storage.
9. The method of claim 8, wherein said generated list of features is transmitted to a user or a different device.
10. The method of any of claims 5-9, wherein said user interest level includes at least one of user engagement, user emotion or user behavior.
11. The method of any of claims 5-10, wherein said user interest level is determined from one or more of user engagement, user emotion and/or user behavior.
12. The method of any of claims 1-11 , wherein said resource availability includes amount of available processing power and/or memory availability.
13. A device comprising at least one processor configured to : receive a user input for performing a task; obtain one or more features associated with the task performance, wherein at least one feature is associated with the user input or associated with a previous user history; determine one or more resources available for performing the task perform, by at least one machine learning model, the based on the one or more available resources and the one or more features associated with the task performance; and to transmit one or more outputs associated with the task performed by the at least one machine learning model.
14. The device of claim 13, wherein the device is comprised in any of: (1) an audio video device, (2) an artificial intelligent hub, (3) a gateway and (4) a cloud network.
15. A computer program comprising software code instructions for performing the method according to any one of claims 1 to 12, when the computer program is executed by a processor.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP21305473.7 | 2021-04-09 | ||
| EP21305473 | 2021-04-09 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022214616A1 true WO2022214616A1 (en) | 2022-10-13 |
Family
ID=75690221
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2022/059319 Ceased WO2022214616A1 (en) | 2021-04-09 | 2022-04-07 | Personalizing audio-visual content based on user's interest |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2022214616A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2016089929A1 (en) * | 2014-12-04 | 2016-06-09 | Microsoft Technology Licensing, Llc | Emotion type classification for interactive dialog system |
| US20180285752A1 (en) * | 2017-03-31 | 2018-10-04 | Samsung Electronics Co., Ltd. | Method for providing information and electronic device supporting the same |
| US20210005187A1 (en) * | 2019-07-05 | 2021-01-07 | Korea Electronics Technology Institute | User adaptive conversation apparatus and method based on monitoring of emotional and ethical states |
-
2022
- 2022-04-07 WO PCT/EP2022/059319 patent/WO2022214616A1/en not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2016089929A1 (en) * | 2014-12-04 | 2016-06-09 | Microsoft Technology Licensing, Llc | Emotion type classification for interactive dialog system |
| US20180285752A1 (en) * | 2017-03-31 | 2018-10-04 | Samsung Electronics Co., Ltd. | Method for providing information and electronic device supporting the same |
| US20210005187A1 (en) * | 2019-07-05 | 2021-01-07 | Korea Electronics Technology Institute | User adaptive conversation apparatus and method based on monitoring of emotional and ethical states |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20240395028A1 (en) | Latent diffusion model autodecoders | |
| EP3885966B1 (en) | Method and device for generating natural language description information | |
| CN114282055B (en) | Video feature extraction method, device, equipment and computer storage medium | |
| CN114860187B (en) | Intelligent voice device control method, device, computer device and storage medium | |
| CN111885398B (en) | Interaction method, device and system based on three-dimensional model, electronic equipment and storage medium | |
| CN116721334A (en) | Training methods, devices, equipment and storage media for image generation models | |
| US12353822B2 (en) | Methods and systems for generating alternative content | |
| CN118096924B (en) | Image processing method, device, equipment and storage medium | |
| CN116185191A (en) | A method for interacting with server, display device and virtual digital human | |
| WO2024249060A1 (en) | Text-to-image diffusion model rearchitecture | |
| CN114580425B (en) | Named entity recognition method and device, electronic equipment and storage medium | |
| CN114419661B (en) | Methods, devices, media, and computer equipment for capturing hand motion in live streaming | |
| WO2024249181A1 (en) | Latent diffusion model autodecoders | |
| CN116932788A (en) | Cover image extraction method, device, equipment and computer storage medium | |
| Zhang et al. | Multimodal llm integrated semantic communications for 6g immersive experiences | |
| WO2023185257A1 (en) | Data processing method, and device and computer-readable storage medium | |
| CN115359220A (en) | Virtual image updating method and device of virtual world | |
| CN118674812A (en) | Image processing and model training method, device, equipment and storage medium | |
| CN119204220A (en) | Multimodal question answering method and device in customer service application scenario | |
| CN117579889A (en) | Image generation method, device, electronic equipment and storage medium | |
| CN116962600A (en) | Subtitle content display methods, devices, equipment, media and program products | |
| CN114255169A (en) | Video generation method and device | |
| CN114417875B (en) | Data processing method, device, equipment, readable storage medium and program product | |
| US20260105735A1 (en) | Computing System with Multi-Layered and Unified Machine-Learning Model | |
| US20260067421A1 (en) | Feature cache-based generative video editing for dynamic frame generation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22721351 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22721351 Country of ref document: EP Kind code of ref document: A1 |