WO2018127627A1 - Method and apparatus for automatic video summarisation - Google Patents
Method and apparatus for automatic video summarisation Download PDFInfo
- Publication number
- WO2018127627A1 WO2018127627A1 PCT/FI2018/050001 FI2018050001W WO2018127627A1 WO 2018127627 A1 WO2018127627 A1 WO 2018127627A1 FI 2018050001 W FI2018050001 W FI 2018050001W WO 2018127627 A1 WO2018127627 A1 WO 2018127627A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video
- attention
- temporal
- attention map
- text description
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/854—Content authoring
- H04N21/8549—Creating video summaries, e.g. movie trailer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/738—Presentation of query results
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/738—Presentation of query results
- G06F16/739—Presentation of query results in form of a video summary, e.g. the video summary being a video sequence, a composite still image or having synthesized frames
Definitions
- This specification generally relates to automatic video summarisation.
- Video summarisation includes producing a video which is smaller in size. Temporal video summarisation includes producing a shorter video. Spatial video summarisation includes producing a video which has less spatial extent that the original. Video summarisation may include detecting events in the video which are relatively more interesting than other events in the video.
- the specification describes a method comprising: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
- the attention map may be a temporal attention map, wherein the locations correspond to temporal locations of the attention map having the highest attention value.
- the attention map may be a spatial attention map, wherein the locations correspond to spatial locations of the attention map having the highest attention value.
- the attention map may be a spatio-temporal attention map, wherein the locations correspond to spatial and temporal locations of the attention map having the highest attention value.
- the method may further comprise outputting a video summary having video portions corresponding to the locations of the attention map having the highest attention value.
- the method may further comprise selecting a video portion based on the temporal location of the attention map having the highest attention value and surrounding temporal locations having an attention value above a threshold attention value.
- the method may further comprise converting the input video to the text description.
- the method may further comprise converting the text description and input question respectively to a text description summary vector and a question summary vector.
- the method may further comprise providing the text description summary vector and the question summary vector to the neural network.
- the specification describes a computer program comprising machine readable instructions that, when executed by computing apparatus, causes it to perform any method as described with reference to the first aspect.
- the specification describes an apparatus configured to perform any method as described with reference to the first aspect.
- the specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to perform a method comprising: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
- the attention map may be a temporal attention map, wherein the locations correspond to temporal locations of the attention map having the highest attention value.
- the attention map may be a spatial attention map, wherein the locations correspond to spatial locations of the attention map having the highest attention value.
- the attention map may be a spatio-temporal attention map, wherein the locations correspond to spatial and temporal locations of the attention map having the highest attention value.
- the computer program code when executed, may cause the apparatus to perform:
- the computer program code when executed, may cause the apparatus to perform:
- the computer program code when executed, may cause the apparatus to perform:
- the computer program code when executed, may cause the apparatus to perform:
- the computer program code when executed, may cause the apparatus to perform:
- the specification describes a computer-readable medium having computer-readable code stored thereon, the computer-readable code, when executed by at least one processor, causes performance of at least: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
- an apparatus comprising means for:
- FIG. 1 is a schematic illustration of an automatic video summariser, according to embodiments of this specification.
- FIG. 2 is a schematic illustration of temporal video summarisation according to embodiments of this specification
- Figure 3 is a schematic illustration of spatial video summarisation according to embodiments of this specification
- Figure 4 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to convert video to a text description according to embodiments of this specification;
- Figure 5 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to produce a video summary based on a user's question according to embodiments of this specification;
- Figure 6 is a flow chart illustrating operations which may be performed by the automatic video summariser in order to produce a spatio-temporal attention map according to embodiments of this specification;
- Figure 7 illustrates an example of a spatio-temporal attention map produced by the automatic video summariser according to embodiments of this specification
- Figure 8 is a schematic illustration of an example configuration of the automatic video summariser according to embodiments of this specification.
- Figure 9 is a computer-readable memory medium upon which computer-readable code may be stored, according to embodiments of this specification.
- FIG. 1 is a schematic illustration of an automatic video summariser 10.
- the automatic video summariser 10 described herein make use of neural networks in order to produce spatio-temporal summaries including visual information relevant to a user's question or request. In this way, the events in the video which are considered to be relevant to the user's question are determined and video portions showing these events can be output as a spatio-temporal summary for the user.
- the automatic video summariser 10 comprises a video-to-text module 20, an artificial intelligence (AI) attention module 30, a user interface 40 for receiving a user input, and an output 50, which may be a display, for example.
- the AI attention module may use deep learning methods such as attention mechanisms, neural attention mechanisms, or one or more neural networks outputting attention weights.
- FIG. 2 is a schematic illustration of temporal video summarisation.
- temporal summarisation the size of an input video 100 made up of video frames looa-iooi is reduced in size in terms of content by producing a video summary with a shorter time duration.
- a number of frames may be extracted from a video 100 formed of frames 100a- i.
- frames iooa,b,e,f,g,i may be extracted and joined temporally one after the other, maintaining the temporal order intact.
- the output video summary would comprise video portion 101 made up of frames looa-b, video portion 102 made up of frames looe-f, and video portion 103 made up of frames looh-i. Accordingly, the summary will be a video having fewer frames than the input video.
- the portions may be made up of any number of frames.
- the portions may contain different frame numbers to the other portions.
- the temporal portions may be determined based on events occurring in the video. For example, a temporal portion may relate to one specific event occurring in the video. Selection of the temporal portions of the video may be performed as described in more detail with reference to Figures 4 to 7.
- the video 100 may be a virtual reality video, for example a 360 degree video shot by a camera having a 360 degree field of view, such as the Nokia OZO camera.
- An example of a frame 110 from a virtual reality video can be seen in figure 3.
- the video may include multiple events in different spatial sectors of the video.
- the video may therefore be spatially summarised.
- a spatial video summary is a video comprising video crops, i.e. spatial video portions extracted from the original video by cropping spatially.
- Figure 3 illustrates spatial crops 111, 112, and 113.
- the size of the video crops may be the same for all crops.
- a resizing step may be applied to increase the resolution of at least one video crop.
- Increasing the resolution may be performed, for example, by upsampling with or without interpolation. Increasing the resolution may also be performed by using neural super- resolution methods. Alternatively, the resizing step may involve decreasing the resolution of at least one video crop. Decreasing the resolution may be performed, for example, by down-sampling of the video crop. Selection of the spatial portions of the video may be performed as described in more detail with reference to Figures 4 to 7.
- the video 100 may be a full length 360 degree movie.
- the movie may include multiple events temporally and multiple events spatially.
- Figure 4 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to convert video to a text description. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 4.
- the automatic video summariser may receive an input video from a video source.
- the video may be a video extract, or it may be a full length movie.
- the video may be provided from any suitable video source.
- the video may be stored on a storage medium such as a DVD, Blue-Ray, hard drive, or any other suitable storage medium.
- the video may be obtained via streaming or download from an external server.
- the feature extraction module may comprise a Convolutional Neural Network (CNN).
- a CNN is an artificial neural network which represents currently the state-of-the-art for performing feature extraction from images and videos.
- a CNN consists of a sequence of computation layers, where the input is the data (a video frame or an image) and the output is a feature vector, i.e., a vector describing the input image.
- a convolutional layer performs a convolution operation on its input, but using a set of convolution kernels.
- Other types of computation layers present in a CNN may be pooling layers, non-linear activation function layers, batch-normalization layers, etc. However, the present invention is not limited to a CNN and other feature extraction methodologies may be utilized.
- the features extracted in operation Siooo may be input to a temporal neural network.
- the temporal neural network may comprise a Recurrent Neural Network (RNN).
- RNN Recurrent Neural Network
- a suitable RNN may be, for example, a Long Short-Term Memory network (LSTM).
- the temporal network outputs a "frame-description" vector, for each input video-frame.
- the frame description vector corresponds to a description of the video-frame.
- the frame-description vector may be used for generating a sentence or phrase describing the video frame, represented by a vector of real numbers.
- the frame description vectors may be analysed by a second RNN.
- the second RNN may also be a LSTM network, or any other suitable temporal neural network.
- the second RNN generates a set of characters, or words, describing the input video-frame. As such a vector comprising a set of sentences describing the whole video is output.
- a softmax function is applied to the vector output by the second RNN as a result of operation S1400. This indicates the distribution of the words corresponding to the extracted features throughout the video.
- the vector which is output may be referred to as a "text description vector”.
- an index synchronisation is performed.
- the text description is synchronised with the video. This includes associating each word or character with a certain video frame. A word or character may be associated with several adjacent video frames.
- the association of the words or characters with corresponding video frames can be achieved by outputting a video-frame index for each word or character, corresponding to the index of the frame which is described by those words or characters.
- a word may be associated with multiple adjacent frames.
- the automatic video summariser outputs a text description of the video associated with corresponding time indexes.
- Figure 5 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to produce a spatio-temporal summary of an input video. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 5.
- an input video is received.
- the video is converted to text, for example as described with reference to Figure 4. However, it will be understood that any suitable video to text conversion may be used.
- the automatic video summariser outputs text descriptions of the video.
- the automatic video summariser receives a user question or request.
- the question or request is input, or converted into, a text format.
- the question or request may relate to information the user would like to know about the input video. For example, the user may wish to find out whether there are any car crashes in the video. Therefore, the user may input a question such as "was there any car crashes in this movie?", or a request such as "would you summarise all the romantic scenes from the movie".
- the interface may be configured such that the user can input the question or request through user interface 40, for example by typing on a keyboard or on a touchscreen device connected to the automatic video summariser.
- the question or request may be verbally output by a user and received by voice recognition software to convert the question into text.
- the text question (or request) and the text descriptions of the video are input into an artificial intelligence (AI) attention module 30, which may comprise one or more neural networks, for example attention neural networks, and/or other operations which produce an "attention vector" .
- AI artificial intelligence
- the text question and text description are analysed by the AI attention module.
- the question may be analysed before being input into the AI attention module 30. An example of how the question may be analysed is described in more detail with reference to Figure 6.
- the AI attention module 30 produces a spatio-temporal attention map representing the attention-intensity that a neural network has put at that point in time and spatial region when trying to answer the user's questions.
- the automatic video summariser retrieves the spatial and temporal portions of the input video corresponding to the temporal and spatial locations of the spatio-temporal attention map having the highest attention-intensity values.
- step S2700 the automatic video summariser outputs the selected video portions as a spatio-temporal video summary.
- Figure 6 is a flow chart illustrating in more detail the steps involved in producing the spatio-temporal attention map used in order to produce the spatio-temporal video summarisation. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 6.
- operation S3000 the text descriptions output as a result of operation S1700 of Figure 4 are input to a word-embedding module.
- the word-embedding module converts the text descriptions to a set of dense vectors.
- Each of the dense vectors may represent a single word with a plurality of real numbers.
- the words in the text description are each converted from a vocabulary representation to a vector of real numbers.
- the vector of real numbers may be of lower dimensionality than the input vector of vocabulary entries, for example a vector with less dimensions or axes.
- the new representation is a point in an "embedding space", where words with similar semantics are nearby.
- the word-embedding module may be implemented by a multi-layer perceptron network or alternatively a single fully-connected layer. In general, the word-embedding module may transform an input into a more convenient output representation. For example, words may be transformed into a new representation for which similar words lie close to each other in the new representation space.
- the text description vectors are input to an RNN where the vectors are analysed.
- the RNN outputs a single output vector, which will be referred to herein as a text description summary vector.
- the RNN may be an LSTM.
- the word-embedding module converts the question to a set of dense vectors.
- the words in the question are each converted from a vocabulary representation to a vector of real numbers with lower dimensionality, in a similar way to the text
- the question vectors are input into an RNN where the vectors are analysed.
- the RNN outputs a single output vector which summarises the question, which will be referred to herein as a question summary vector.
- the RNN may be an LSTM.
- the text description summary vector and question summary vector are combined.
- the combination operation may be a concatenation in one of the dimensions of the input vectors, or an element-wise addition (if the input vectors have same dimensionalities). However, any suitable combination operation may be used at this step.
- the concatenated summary vectors are provided to a multi-layer perceptron (MLP) neural network.
- MLP multi-layer perceptron
- the MLP neural network may be referred to as an "attention neural network”.
- the MLP is a neural network comprising a set of dense (i.e. fully connected) layers, followed by a softmax layer.
- the dense layers of the MLP learn how to map the concatenated word-embedded text descriptions and user questions to an attention vector.
- the mapping is learned from data via a training process which happens offline, and which happens end-to-end for the whole model proposed in this invention.
- the input data is videos and a set of questions for each video, and the ground-truth output is the video segments which form the target video summary.
- the attention vector is in practice a set of attention weights (i.e., real numbers), summing up to 1, where each attention weight is associated to a certain temporal location of the video.
- the softmax layer will output a probability distribution over "temporal attention weights" w.
- the size of the output vector (i.e. the number of weights w) is the number of temporal locations, which is the number of words in the text describing the input video.
- the size of the output vector is less than the number of words in the video description, and thus an attention weight can refer to more than one word. This would be a case where the attention is "quantized".
- the weights represent a l-dimensional "temporal attention map" (TAM), having bins which each correspond to a temporal location, and having a value of the value of the attention weight associated to that temporal location.
- TAM value at location f, TAM[f] represents the attention-intensity that the attention neural network has put at that point in time when trying to answer the user's question.
- the temporal locations associated with each bin correspond to temporal location of the input video.
- the attention weights output by the MLP are a vector of N bins, where N is the number of total temporal locations of the video. Therefore, the attention weights correspond to words of the text description and are arranged in the same temporal order as the temporal order of the words of the text description of the video.
- temporal synchronisation is achieved based on the temporal location of the attention weights and the corresponding words of the text description.
- the dimensionality of the vector output by the MLP is determined automatically based on the number of words of the text description created by the video-to-text module 20.
- the attention neural network outputs the probability distribution over attention weights which can be represented as the temporal attention map.
- a temporal location t* of the TAM corresponding to the highest attention value in the TAM indicates the temporal location of the video which answers the user's question.
- the temporal extent of the video portion is determined based on the temporal extent of the attention values around f *.
- a threshold value of the attention weight values may determine the temporal boundaries of the video portion to extract. That is, the video portion is selected based on temporal locations of attention weights above a given threshold.
- the temporal extent of the video portion may be selected in any other suitable way.
- FIG. 7 illustrates an example of a spatio-temporal attention map (STAM) produced by the automatic video summariser.
- STAM represents the attention weights
- the TAM is extended to the spatial domain by analysing the video separately in the spatial dimension.
- the video may be divided into a given number of angular sectors.
- Each sector is analysed separately by several attention networks.
- the joint output of the attention network is a 2-dimensional attention map, or "spatio-temporal attention map" (STAM).
- the STAM is output as a matrix indexed using two indices, one for the time (f), and one for the space (the angular sector s).
- the time runs along the x-axis of the map, and the space runs along the y-axis.
- the video portion i.e. the particular temporal location and extent, and the spatial crop
- the video portion is based on the temporal location t* and angular sector s* having the highest attention value in the STAM matrix.
- the video may be divided into a number of predetermined sectors.
- the division of the video may be dynamic.
- the automatic video summariser may divide the video by means of "spatial scene cut detection".
- Spatial scene cut detection may be achieved by analysing the video with deep learning or multimedia analysis techniques to detect objects, actions and activities, and then virtually cutting the scene to include the object, action and activity spatially. Therefore, the amount of data needed to analyse and summarize a spatial virtual reality video may be reduced.
- Spatial summary may be applicable to 360 degree videos in order to convert a 360 degree video to a standard size video. Spatial summary may also be performed without any temporal summarisation, if this is desired.
- the indices may be used to extract the corresponding temporal and spatial portions of the video for output to a user as a spatio-temporal summary of the video, based on the user's question.
- the user is provided with two video portions.
- the first video portion corresponds to the temporal location of the video indicated by indices ti, si.
- the temporal extent of the video portion may be determined as described above, for example by setting a threshold attention value for the values temporally adjacent to ti.
- the second video portion corresponds to the temporal location of the video indicated by indices t2, s2.
- the automatic video summariser is able to output a video summary which is determined to be the most relevant to the user's question or request.
- the video portions may be output through the output 50.
- the output may be a display which forms part of the automatic video summariser 10.
- the automatic video summariser 10 may be configured to output the video portions to a display which does not form part of the automatic video summariser 10, such as a display of a TV or PC, etc.
- the automatic video summariser may be located on a server which is separate to the display through which the video portions are output.
- the automatic video summariser may be configured to output indicators of temporal locations of a video to be played in a video summary.
- FIG 8 is a schematic block diagram of an example configuration of an automatic video summariser such as that described with reference to Figures 1 to 7.
- the video summariser may comprise memory and processing circuitry.
- the memory 11 may comprise any combination of different types of memory.
- the memory comprises one or more read-only memory (ROM) media 13 and one or more random access memory (RAM) memory media 12.
- the processing circuitry 14 may be configured to process an input video and user question as described with reference to Figures 1 to 7.
- the memory described with reference to Figure 8 may have computer readable instructions stored thereon 13A, which when executed by the processing circuitry 14 causes the processing circuitry 14 to cause performance of various ones of the operations described above.
- the processing circuitry 14 described above with reference to Figure 8 may be of any suitable composition and may include one or more processors 14A of any suitable type or suitable combination of types.
- the processing circuitry 14 may be a programmable processor that interprets computer program instructions and processes data.
- the processing circuitry 14 may include plural programmable processors.
- the processing circuitry 14 may be, for example, programmable hardware with embedded firmware.
- the processing circuitry 14 may be termed processing means.
- the processing circuitry 14 may alternatively or additionally include one or more
- processing circuitry 14 may be referred to as computing apparatus.
- the processing circuitry 14 described with reference to Figure 8 is coupled to the memory 11 (or one or more storage devices) and is operable to read/write data to/from the memory.
- the memory may comprise a single memory unit or a plurality of memory units 13 upon which the computer readable instructions 13A (or code) is stored.
- the memory 11 may comprise both volatile memory 12 and non-volatile memory 13.
- the computer readable instructions 13A may be stored in the non-volatile memory 13 and may be executed by the processing circuitry 14 using the volatile memory 12 for temporary storage of data or data and instructions.
- volatile memory include RAM, DRAM, and SDRAM etc.
- Examples of non-volatile memory include ROM, PROM, EEPROM, flash memory, optical storage, magnetic storage, etc.
- the memories 11 in general may be referred to as non-transitory computer readable memory media.
- the term 'memory' in addition to covering memory comprising both non-volatile memory and volatile memory, may also cover one or more volatile memories only, one or more non-volatile memories only, or one or more volatile memories and one or more nonvolatile memories.
- the computer readable instructions 13A described herein with reference to Figure 8 may be pre-programmed into the automatic video summariser. Alternatively, the computer readable instructions 13A may arrive at the automatic video summariser via an electromagnetic carrier signal or may be copied from a physical entity such as a computer program product, a memory device or a record medium such as a CD-ROM or DVD.
- the computer readable instructions 13A may provide the logic and routines that enable the automatic video summariser to perform the functionalities described above.
- the video-text module 20, AI attention module 30, the feature extraction module, and the word-embedding module may be implemented as computer readable instructions stored on one or more memories, which, when executed by the processor circuitry, cause processing input data according to embodiments of the invention.
- the combination of computer-readable instructions stored on memory (of any of the types described above) may be referred to as a computer program or a computer program product.
- Figure 9 illustrates an example of a computer-readable medium 16 with computer- readable instructions (code) stored thereon.
- the computer-readable instructions (code) when executed by a processor, may cause any one of or any combination of the operations described above to be performed.
- the automatic video summariser described herein may include various hardware components which have may not been shown in the Figures since they may not have direct interaction with the shown features.
- Embodiments may be implemented in software, hardware, application logic or a combination of software, hardware and application logic.
- the software, application logic and/or hardware may reside on memory, or any computer media.
- the application logic, software or an instruction set is maintained on any one of various conventional computer-readable media.
- a "memory" or “computer-readable medium” may be any non-transitory media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer.
- references to, where relevant, "computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc., or a “processor” or “processing circuitry” etc. should be understood to encompass not only computers having differing architectures such as single/multi-processor architectures and sequencers/parallel architectures, but also specialised circuits such as field programmable gate arrays FPGA, application specific circuits ASIC, signal processing devices and other devices.
- References to computer program, instructions, code etc. should be understood to express software for a programmable processor firmware such as the programmable content of a hardware device as instructions for a processor or configured or configuration settings for a fixed function device, gate array, programmable logic device, etc.
- circuitry refers to all of the following: (a) hardware- only circuit implementations (such as implementations in only analogue and/or digital circuitry) and (b) to combinations of circuits and software (and/or firmware), such as (as applicable): (i) to a combination of processor(s) or (ii) to portions of processor(s)/software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile device or server, to perform various functions) and (c) to circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present.
- circuitry would also cover an implementation of merely a processor (or multiple processors) or portion of a processor and its (or their) accompanying software and/or firmware.
- circuitry would also cover, for example and if applicable to the particular claim element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in server, a cellular network device, or other network device.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Theoretical Computer Science (AREA)
- Computer Security & Cryptography (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
This specification describes a method comprising: analysing (S2400), using a neural network, a text description of an input video and an input question; causing production of an attention map (S2500) based on the text description and the input question; and determining locations of the attention map having the highest attention value.
Description
Method and Apparatus for Automatic Video Summarisation Field
This specification generally relates to automatic video summarisation.
Background
Video summarisation includes producing a video which is smaller in size. Temporal video summarisation includes producing a shorter video. Spatial video summarisation includes producing a video which has less spatial extent that the original. Video summarisation may include detecting events in the video which are relatively more interesting than other events in the video.
Summary
According to a first aspect, the specification describes a method comprising: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
The attention map may be a temporal attention map, wherein the locations correspond to temporal locations of the attention map having the highest attention value.
The attention map may be a spatial attention map, wherein the locations correspond to spatial locations of the attention map having the highest attention value. The attention map may be a spatio-temporal attention map, wherein the locations correspond to spatial and temporal locations of the attention map having the highest attention value.
The method may further comprise outputting a video summary having video portions corresponding to the locations of the attention map having the highest attention value.
The method may further comprise selecting a video portion based on the temporal location of the attention map having the highest attention value and surrounding temporal locations having an attention value above a threshold attention value.
The method may further comprise converting the input video to the text description.
The method may further comprise converting the text description and input question respectively to a text description summary vector and a question summary vector.
The method may further comprise providing the text description summary vector and the question summary vector to the neural network.
According to a second aspect, the specification describes a computer program comprising machine readable instructions that, when executed by computing apparatus, causes it to perform any method as described with reference to the first aspect.
According to a third aspect, the specification describes an apparatus configured to perform any method as described with reference to the first aspect.
According to a fourth aspect, the specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to perform a method comprising: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
The attention map may be a temporal attention map, wherein the locations correspond to temporal locations of the attention map having the highest attention value. The attention map may be a spatial attention map, wherein the locations correspond to spatial locations of the attention map having the highest attention value.
The attention map may be a spatio-temporal attention map, wherein the locations correspond to spatial and temporal locations of the attention map having the highest attention value.
The computer program code, when executed, may cause the apparatus to perform:
outputting a video summary having video portions corresponding to the locations of the attention map having the highest attention value.
The computer program code, when executed, may cause the apparatus to perform:
selecting a video portion based on the temporal location of the attention map having the
highest attention value and surrounding temporal locations having an attention value above a threshold attention value.
The computer program code, when executed, may cause the apparatus to perform:
converting the input video to the text description.
The computer program code, when executed, may cause the apparatus to perform:
converting the text description and input question respectively to a text description summary vector and a question summary vector.
The computer program code, when executed, may cause the apparatus to perform:
providing the text description summary vector and the question summary vector to the neural network. According to a fifth aspect, the specification describes a computer-readable medium having computer-readable code stored thereon, the computer-readable code, when executed by at least one processor, causes performance of at least: analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
According to a sixth aspect, there is provided an apparatus comprising means for:
analysing, using a neural network, a text description of an input video and an input question; causing production of an attention map based on the text description and the input question; and determining locations of the attention map having the highest attention value.
Brief Description of the Figures
For a more complete understanding of the methods, apparatuses and computer-readable instructions described herein, reference is now made to the following descriptions taken in connection with the accompanying drawings in which:
Figure 1 is a schematic illustration of an automatic video summariser, according to embodiments of this specification;
Figure 2 is a schematic illustration of temporal video summarisation according to embodiments of this specification;
Figure 3 is a schematic illustration of spatial video summarisation according to embodiments of this specification;
Figure 4 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to convert video to a text description according to embodiments of this specification;
Figure 5 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to produce a video summary based on a user's question according to embodiments of this specification;
Figure 6 is a flow chart illustrating operations which may be performed by the automatic video summariser in order to produce a spatio-temporal attention map according to embodiments of this specification;
Figure 7 illustrates an example of a spatio-temporal attention map produced by the automatic video summariser according to embodiments of this specification;
Figure 8 is a schematic illustration of an example configuration of the automatic video summariser according to embodiments of this specification;
Figure 9 is a computer-readable memory medium upon which computer-readable code may be stored, according to embodiments of this specification.
Detailed Description
In the description and drawings, like reference numerals may refer to like elements throughout.
Figure 1 is a schematic illustration of an automatic video summariser 10. The automatic video summariser 10 described herein make use of neural networks in order to produce spatio-temporal summaries including visual information relevant to a user's question or request. In this way, the events in the video which are considered to be relevant to the user's question are determined and video portions showing these events can be output as a spatio-temporal summary for the user.
The automatic video summariser 10 comprises a video-to-text module 20, an artificial intelligence (AI) attention module 30, a user interface 40 for receiving a user input, and an output 50, which may be a display, for example. The AI attention module may use deep learning methods such as attention mechanisms, neural attention mechanisms, or one or more neural networks outputting attention weights.
Figure 2 is a schematic illustration of temporal video summarisation. In temporal summarisation, the size of an input video 100 made up of video frames looa-iooi is reduced in size in terms of content by producing a video summary with a shorter time duration. A number of frames may be extracted from a video 100 formed of frames 100a-
i. For example, frames iooa,b,e,f,g,i may be extracted and joined temporally one after the other, maintaining the temporal order intact. The output video summary would comprise video portion 101 made up of frames looa-b, video portion 102 made up of frames looe-f, and video portion 103 made up of frames looh-i. Accordingly, the summary will be a video having fewer frames than the input video. The portions may be made up of any number of frames. The portions may contain different frame numbers to the other portions. The temporal portions may be determined based on events occurring in the video. For example, a temporal portion may relate to one specific event occurring in the video. Selection of the temporal portions of the video may be performed as described in more detail with reference to Figures 4 to 7.
The video 100 may be a virtual reality video, for example a 360 degree video shot by a camera having a 360 degree field of view, such as the Nokia OZO camera. An example of a frame 110 from a virtual reality video can be seen in figure 3. The video may include multiple events in different spatial sectors of the video. The video may therefore be spatially summarised. A spatial video summary is a video comprising video crops, i.e. spatial video portions extracted from the original video by cropping spatially. Figure 3 illustrates spatial crops 111, 112, and 113. In spatial summarisation, the size of the video crops may be the same for all crops. In embodiments where the crops are not the same size, a resizing step may be applied to increase the resolution of at least one video crop. Increasing the resolution may be performed, for example, by upsampling with or without interpolation. Increasing the resolution may also be performed by using neural super- resolution methods. Alternatively, the resizing step may involve decreasing the resolution of at least one video crop. Decreasing the resolution may be performed, for example, by down-sampling of the video crop. Selection of the spatial portions of the video may be performed as described in more detail with reference to Figures 4 to 7.
By performing both temporal and spatial summarisation, a spatio-temporal video summary can be produced. For example, the video 100 may be a full length 360 degree movie. The movie may include multiple events temporally and multiple events spatially.
Figure 4 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to convert video to a text description. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 4.
In operation Siooo the automatic video summariser may receive an input video from a video source. The video may be a video extract, or it may be a full length movie. The video may be provided from any suitable video source. For example, the video may be stored on a storage medium such as a DVD, Blue-Ray, hard drive, or any other suitable storage medium. Alternatively, the video may be obtained via streaming or download from an external server.
In operation SHOO an input video is analysed by a feature extraction module. The feature extraction module may comprise a Convolutional Neural Network (CNN). A CNN is an artificial neural network which represents currently the state-of-the-art for performing feature extraction from images and videos. A CNN consists of a sequence of computation layers, where the input is the data (a video frame or an image) and the output is a feature vector, i.e., a vector describing the input image. There may be different types of computation layers in a CNN, but the most important is the convolutional layer. A convolutional layer performs a convolution operation on its input, but using a set of convolution kernels. Other types of computation layers present in a CNN may be pooling layers, non-linear activation function layers, batch-normalization layers, etc. However, the present invention is not limited to a CNN and other feature extraction methodologies may be utilized.
In operation S1200, the features extracted in operation Siooo may be input to a temporal neural network. The temporal neural network may comprise a Recurrent Neural Network (RNN). A suitable RNN may be, for example, a Long Short-Term Memory network (LSTM).
In operation S1300, the temporal network outputs a "frame-description" vector, for each input video-frame. The frame description vector corresponds to a description of the video-frame. The frame-description vector may be used for generating a sentence or phrase describing the video frame, represented by a vector of real numbers.
In operation S1400, the frame description vectors may be analysed by a second RNN. The second RNN may also be a LSTM network, or any other suitable temporal neural network.
The second RNN generates a set of characters, or words, describing the input video-frame. As such a vector comprising a set of sentences describing the whole video is output.
In operation S1500, a softmax function is applied to the vector output by the second RNN as a result of operation S1400. This indicates the distribution of the words corresponding to the extracted features throughout the video. The vector which is output may be referred to as a "text description vector".
In operation S1600, an index synchronisation is performed. In order to determine the temporal locations of the features within the video, the text description is synchronised with the video. This includes associating each word or character with a certain video frame. A word or character may be associated with several adjacent video frames.
The association of the words or characters with corresponding video frames can be achieved by outputting a video-frame index for each word or character, corresponding to the index of the frame which is described by those words or characters. For example, in one case, one word may be associated with multiple adjacent frames.
In operation S1700, the automatic video summariser outputs a text description of the video associated with corresponding time indexes.
However, it will be recognised that any suitable implementation of a video to text module 20 can be utilised.
Figure 5 is a flow chart illustrating various operations which may be performed by the automatic video summariser in order to produce a spatio-temporal summary of an input video. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 5.
In operation S2000 an input video is received. In operation S2100, the video is converted to text, for example as described with reference to Figure 4. However, it will be understood that any suitable video to text conversion may be used.
In operation S2200, the automatic video summariser outputs text descriptions of the video.
In operation S2300 the automatic video summariser receives a user question or request. The question or request is input, or converted into, a text format. The question or request may relate to information the user would like to know about the input video. For example, the user may wish to find out whether there are any car crashes in the video. Therefore, the user may input a question such as "was there any car crashes in this movie?", or a request such as "would you summarise all the romantic scenes from the movie". The interface may be configured such that the user can input the question or request through user interface 40, for example by typing on a keyboard or on a touchscreen device connected to the automatic video summariser. Alternatively, the question or request may be verbally output by a user and received by voice recognition software to convert the question into text.
In operation S2400, the text question (or request) and the text descriptions of the video are input into an artificial intelligence (AI) attention module 30, which may comprise one or more neural networks, for example attention neural networks, and/or other operations which produce an "attention vector" . The text question and text description are analysed by the AI attention module. The question may be analysed before being input into the AI attention module 30. An example of how the question may be analysed is described in more detail with reference to Figure 6.
In operation S2500, the AI attention module 30 produces a spatio-temporal attention map representing the attention-intensity that a neural network has put at that point in time and spatial region when trying to answer the user's questions. In step S2600, the automatic video summariser retrieves the spatial and temporal portions of the input video corresponding to the temporal and spatial locations of the spatio-temporal attention map having the highest attention-intensity values.
In step S2700, the automatic video summariser outputs the selected video portions as a spatio-temporal video summary.
Figure 6 is a flow chart illustrating in more detail the steps involved in producing the spatio-temporal attention map used in order to produce the spatio-temporal video summarisation. In some embodiments, not all of the illustrated operations need to be performed. Operations may also be performed in a different order compared to the order presented in Figure 6.
In operation S3000, the text descriptions output as a result of operation S1700 of Figure 4 are input to a word-embedding module.
In operation S3100, the word-embedding module converts the text descriptions to a set of dense vectors. Each of the dense vectors may represent a single word with a plurality of real numbers. The words in the text description are each converted from a vocabulary representation to a vector of real numbers. The vector of real numbers may be of lower dimensionality than the input vector of vocabulary entries, for example a vector with less dimensions or axes. The new representation is a point in an "embedding space", where words with similar semantics are nearby. The word-embedding module may be implemented by a multi-layer perceptron network or alternatively a single fully-connected layer. In general, the word-embedding module may transform an input into a more convenient output representation. For example, words may be transformed into a new representation for which similar words lie close to each other in the new representation space.
In operation S3200, the text description vectors are input to an RNN where the vectors are analysed. The RNN outputs a single output vector, which will be referred to herein as a text description summary vector. The RNN may be an LSTM.
In operation S3300, the question is input to a word-embedding module.
In operation S3400, the word-embedding module converts the question to a set of dense vectors. The words in the question are each converted from a vocabulary representation to a vector of real numbers with lower dimensionality, in a similar way to the text
descriptions in operation S3100.
In operation S3500, the question vectors are input into an RNN where the vectors are analysed. The RNN outputs a single output vector which summarises the question, which will be referred to herein as a question summary vector. The RNN may be an LSTM.
In operation S3600, the text description summary vector and question summary vector are combined. The combination operation may be a concatenation in one of the dimensions of the input vectors, or an element-wise addition (if the input vectors have same dimensionalities). However, any suitable combination operation may be used at this step.
In operation S3700, the concatenated summary vectors are provided to a multi-layer perceptron (MLP) neural network. The MLP neural network may be referred to as an "attention neural network". The MLP is a neural network comprising a set of dense (i.e. fully connected) layers, followed by a softmax layer.
The dense layers of the MLP learn how to map the concatenated word-embedded text descriptions and user questions to an attention vector. The mapping is learned from data via a training process which happens offline, and which happens end-to-end for the whole model proposed in this invention. The input data is videos and a set of questions for each video, and the ground-truth output is the video segments which form the target video summary. The attention vector is in practice a set of attention weights (i.e., real numbers), summing up to 1, where each attention weight is associated to a certain temporal location of the video. The softmax layer will output a probability distribution over "temporal attention weights" w.
The size of the output vector (i.e. the number of weights w) is the number of temporal locations, which is the number of words in the text describing the input video. In an alternative implementation, the size of the output vector is less than the number of words in the video description, and thus an attention weight can refer to more than one word. This would be a case where the attention is "quantized".
The weights represent a l-dimensional "temporal attention map" (TAM), having bins which each correspond to a temporal location, and having a value of the value of the attention weight associated to that temporal location. The TAM value at location f, TAM[f], represents the attention-intensity that the attention neural network has put at that point in time when trying to answer the user's question. The temporal locations associated with each bin correspond to temporal location of the input video. The attention weights output by the MLP are a vector of N bins, where N is the number of total temporal locations of the video. Therefore, the attention weights correspond to words of the text description and are arranged in the same temporal order as the temporal order of the words of the text description of the video. Accordingly temporal synchronisation is achieved based on the temporal location of the attention weights and the corresponding words of the text description. The dimensionality of the
vector output by the MLP is determined automatically based on the number of words of the text description created by the video-to-text module 20.
In operation S3800, the attention neural network outputs the probability distribution over attention weights which can be represented as the temporal attention map. A temporal location t* of the TAM corresponding to the highest attention value in the TAM indicates the temporal location of the video which answers the user's question.
The temporal extent of the video portion is determined based on the temporal extent of the attention values around f *. For example, a threshold value of the attention weight values may determine the temporal boundaries of the video portion to extract. That is, the video portion is selected based on temporal locations of attention weights above a given threshold. However, the temporal extent of the video portion may be selected in any other suitable way.
Figure 7 illustrates an example of a spatio-temporal attention map (STAM) produced by the automatic video summariser. The STAM represents the attention weights
corresponding to each temporal location and spatial region of the video.
The TAM is extended to the spatial domain by analysing the video separately in the spatial dimension. For example, the video may be divided into a given number of angular sectors. Each sector is analysed separately by several attention networks. The joint output of the attention network is a 2-dimensional attention map, or "spatio-temporal attention map" (STAM).
The STAM is output as a matrix indexed using two indices, one for the time (f), and one for the space (the angular sector s). In Figure 7, the time runs along the x-axis of the map, and the space runs along the y-axis. In order to answer the user's question, the video portion (i.e. the particular temporal location and extent, and the spatial crop) will be determined by the highest value of attention within the STAM matrix. The video portion is based on the temporal location t* and angular sector s* having the highest attention value in the STAM matrix.
The video may be divided into a number of predetermined sectors. Alternatively, the division of the video may be dynamic. For example, the automatic video summariser may divide the video by means of "spatial scene cut detection". Spatial scene cut detection may be achieved by analysing the video with deep learning or multimedia analysis techniques
to detect objects, actions and activities, and then virtually cutting the scene to include the object, action and activity spatially. Therefore, the amount of data needed to analyse and summarize a spatial virtual reality video may be reduced. Spatial summary may be applicable to 360 degree videos in order to convert a 360 degree video to a standard size video. Spatial summary may also be performed without any temporal summarisation, if this is desired.
In Figure 7, the temporal and spatial locations determined by the automatic video summariser as answering the user's question are indicated by the indices ti, si, and t2, s2.
Therefore, the indices may be used to extract the corresponding temporal and spatial portions of the video for output to a user as a spatio-temporal summary of the video, based on the user's question. In this case, the user is provided with two video portions. The first video portion corresponds to the temporal location of the video indicated by indices ti, si. The temporal extent of the video portion may be determined as described above, for example by setting a threshold attention value for the values temporally adjacent to ti. The second video portion corresponds to the temporal location of the video indicated by indices t2, s2.
Accordingly, by determining the highest attention values, the automatic video summariser is able to output a video summary which is determined to be the most relevant to the user's question or request. The video portions may be output through the output 50. The output may be a display which forms part of the automatic video summariser 10.
Alternatively, the automatic video summariser 10 may be configured to output the video portions to a display which does not form part of the automatic video summariser 10, such as a display of a TV or PC, etc. For example, the automatic video summariser may be located on a server which is separate to the display through which the video portions are output. The automatic video summariser may be configured to output indicators of temporal locations of a video to be played in a video summary.
Figure 8 is a schematic block diagram of an example configuration of an automatic video summariser such as that described with reference to Figures 1 to 7. The video summariser may comprise memory and processing circuitry. The memory 11 may comprise any combination of different types of memory. In the example of Figure 8, the memory comprises one or more read-only memory (ROM) media 13 and one or more random
access memory (RAM) memory media 12. The processing circuitry 14 may be configured to process an input video and user question as described with reference to Figures 1 to 7.
The memory described with reference to Figure 8 may have computer readable instructions stored thereon 13A, which when executed by the processing circuitry 14 causes the processing circuitry 14 to cause performance of various ones of the operations described above. The processing circuitry 14 described above with reference to Figure 8 may be of any suitable composition and may include one or more processors 14A of any suitable type or suitable combination of types. For example, the processing circuitry 14 may be a programmable processor that interprets computer program instructions and processes data. The processing circuitry 14 may include plural programmable processors. Alternatively, the processing circuitry 14 may be, for example, programmable hardware with embedded firmware. The processing circuitry 14 may be termed processing means. The processing circuitry 14 may alternatively or additionally include one or more
Application Specific Integrated Circuits (ASICs). In some instances, processing circuitry 14 may be referred to as computing apparatus.
The processing circuitry 14 described with reference to Figure 8 is coupled to the memory 11 (or one or more storage devices) and is operable to read/write data to/from the memory. The memory may comprise a single memory unit or a plurality of memory units 13 upon which the computer readable instructions 13A (or code) is stored. For example, the memory 11 may comprise both volatile memory 12 and non-volatile memory 13. For example, the computer readable instructions 13A may be stored in the non-volatile memory 13 and may be executed by the processing circuitry 14 using the volatile memory 12 for temporary storage of data or data and instructions. Examples of volatile memory include RAM, DRAM, and SDRAM etc. Examples of non-volatile memory include ROM, PROM, EEPROM, flash memory, optical storage, magnetic storage, etc. The memories 11 in general may be referred to as non-transitory computer readable memory media. The term 'memory', in addition to covering memory comprising both non-volatile memory and volatile memory, may also cover one or more volatile memories only, one or more non-volatile memories only, or one or more volatile memories and one or more nonvolatile memories. The computer readable instructions 13A described herein with reference to Figure 8 may be pre-programmed into the automatic video summariser. Alternatively, the computer readable instructions 13A may arrive at the automatic video summariser via an
electromagnetic carrier signal or may be copied from a physical entity such as a computer program product, a memory device or a record medium such as a CD-ROM or DVD. The computer readable instructions 13A may provide the logic and routines that enable the automatic video summariser to perform the functionalities described above. For example, the video-text module 20, AI attention module 30, the feature extraction module, and the word-embedding module may be implemented as computer readable instructions stored on one or more memories, which, when executed by the processor circuitry, cause processing input data according to embodiments of the invention. The combination of computer-readable instructions stored on memory (of any of the types described above) may be referred to as a computer program or a computer program product.
Figure 9 illustrates an example of a computer-readable medium 16 with computer- readable instructions (code) stored thereon. The computer-readable instructions (code), when executed by a processor, may cause any one of or any combination of the operations described above to be performed.
As will be appreciated, the automatic video summariser described herein may include various hardware components which have may not been shown in the Figures since they may not have direct interaction with the shown features.
Embodiments may be implemented in software, hardware, application logic or a combination of software, hardware and application logic. The software, application logic and/or hardware may reside on memory, or any computer media. In an example embodiment, the application logic, software or an instruction set is maintained on any one of various conventional computer-readable media. In the context of this document, a "memory" or "computer-readable medium" may be any non-transitory media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer.
Reference to, where relevant, "computer-readable storage medium", "computer program product", "tangibly embodied computer program" etc., or a "processor" or "processing circuitry" etc. should be understood to encompass not only computers having differing architectures such as single/multi-processor architectures and sequencers/parallel architectures, but also specialised circuits such as field programmable gate arrays FPGA, application specific circuits ASIC, signal processing devices and other devices. References to computer program, instructions, code etc. should be understood to express software for
a programmable processor firmware such as the programmable content of a hardware device as instructions for a processor or configured or configuration settings for a fixed function device, gate array, programmable logic device, etc. As used in this application, the term 'circuitry' refers to all of the following: (a) hardware- only circuit implementations (such as implementations in only analogue and/or digital circuitry) and (b) to combinations of circuits and software (and/or firmware), such as (as applicable): (i) to a combination of processor(s) or (ii) to portions of processor(s)/software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile device or server, to perform various functions) and (c) to circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present. This definition of 'circuitry' applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" would also cover an implementation of merely a processor (or multiple processors) or portion of a processor and its (or their) accompanying software and/or firmware. The term
"circuitry" would also cover, for example and if applicable to the particular claim element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in server, a cellular network device, or other network device.
If desired, the different functions discussed herein may be performed in a different order and/or concurrently with each other. Furthermore, if desired, one or more of the above- described functions may be optional or may be combined. Similarly, it will also be appreciated that the flow diagrams of Figures 4 to 6 are examples only and that various operations depicted therein may be omitted, reordered and/or combined. Although various aspects are set out in the independent claims, other aspects comprise other combinations of features from the described embodiments and/or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims. It is also noted herein that while the above describes various examples, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and
modifications which may be made without departing from the scope of the appended claims.
Claims
1. A method comprising:
analysing, using a neural network, a text description of an input video and an input question;
causing production of an attention map based on the text description and the input question; and
determining locations of the attention map having the highest attention value.
2. A method according to claim 1, wherein
the attention map is a temporal attention map, and wherein the locations correspond to temporal locations of the attention map having the highest attention value.
3. A method according to claim 1, wherein the attention map is a spatial attention map, and wherein the locations correspond to spatial locations of the attention map having the highest attention value.
4. A method according to claim 1, wherein the attention map is a spatio-temporal attention map, wherein the locations correspond to spatial and temporal locations of the attention map having the highest attention value.
5. A method according to any preceding claim, comprising outputting a video summary having video portions corresponding to the locations of the attention map having the highest attention value.
6. A method according to claim 5, comprising selecting a video portion based on the temporal location of the attention map having the highest attention value and surrounding temporal locations having an attention value above a threshold attention value.
7. A method according to any preceding claim, further comprising converting the input video to the text description.
8. A method according to any preceding claim, further comprising converting the text description and input question respectively to a text description summary vector and a question summary vector.
9. A method according to claim 8, further comprising providing the text description summary vector and the question summary vector to the neural network.
10. A computer program comprising machine readable instructions that, when executed by computing apparatus, causes it to perform the method of any preceding claim.
11. Apparatus configured to perform the method of any of claims 1 to 9.
12. Apparatus comprising:
at least one processor; and
at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to perform a method comprising:
analysing, using a neural network, a text description of an input video and an input question;
causing production of an attention map based on the text description and the input question; and
determining locations of the attention map having the highest attention value.
13. A computer-readable medium having computer-readable code stored thereon, the computer-readable code, when executed by the at least one processor, causes performance of at least:
analysing, using a neural network, a text description of an input video and an input question;
causing production of an attention map based on the text description and the input question; and
determining locations of the attention map having the highest attention value.
14. Apparatus comprising means for:
analysing, using a neural network, a text description of an input video and an input question;
causing production of an attention map based on the text description and the input question; and
determining locations of the attention map having the highest attention value.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB1700265.0 | 2017-01-06 | ||
| GB1700265.0A GB2558582A (en) | 2017-01-06 | 2017-01-06 | Method and apparatus for automatic video summarisation |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018127627A1 true WO2018127627A1 (en) | 2018-07-12 |
Family
ID=58463740
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/FI2018/050001 Ceased WO2018127627A1 (en) | 2017-01-06 | 2018-01-02 | Method and apparatus for automatic video summarisation |
Country Status (2)
| Country | Link |
|---|---|
| GB (1) | GB2558582A (en) |
| WO (1) | WO2018127627A1 (en) |
Cited By (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109413448A (en) * | 2018-11-05 | 2019-03-01 | 中山大学 | Mobile device panoramic video play system based on deeply study |
| CN109871124A (en) * | 2019-01-25 | 2019-06-11 | 华南理工大学 | Emotion virtual reality scenario appraisal procedure based on deep learning |
| CN109889923A (en) * | 2019-02-28 | 2019-06-14 | 杭州一知智能科技有限公司 | A method for summarizing videos using a hierarchical self-attention network combined with video descriptions |
| CN109919114A (en) * | 2019-03-14 | 2019-06-21 | 浙江大学 | One kind is based on the decoded video presentation method of complementary attention mechanism cyclic convolution |
| CN110267051A (en) * | 2019-05-16 | 2019-09-20 | 北京奇艺世纪科技有限公司 | A kind of method and device of data processing |
| CN110414377A (en) * | 2019-07-09 | 2019-11-05 | 武汉科技大学 | A Scene Classification Method for Remote Sensing Image Based on Scale Attention Network |
| CN110933518A (en) * | 2019-12-11 | 2020-03-27 | 浙江大学 | Method for generating query-oriented video abstract by using convolutional multi-layer attention network mechanism |
| CN111241410A (en) * | 2020-01-22 | 2020-06-05 | 深圳司南数据服务有限公司 | Industry news recommendation method and terminal |
| WO2020197853A1 (en) * | 2019-03-22 | 2020-10-01 | Nec Laboratories America, Inc. | Efficient and fine-grained video retrieval |
| CN112016493A (en) * | 2020-09-03 | 2020-12-01 | 科大讯飞股份有限公司 | Image description method and device, electronic equipment and storage medium |
| CN113343821A (en) * | 2021-05-31 | 2021-09-03 | 合肥工业大学 | Non-contact heart rate measurement method based on space-time attention network and input optimization |
| WO2022134634A1 (en) * | 2020-12-22 | 2022-06-30 | 北京达佳互联信息技术有限公司 | Video processing method and electronic device |
| CN115334367A (en) * | 2022-07-11 | 2022-11-11 | 北京达佳互联信息技术有限公司 | Video summary information generation method, device, server and storage medium |
| US11663268B2 (en) | 2018-03-22 | 2023-05-30 | Guangdong Oppo Mobile Telecommunications Corp., Ltd. | Method and system for retrieving video temporal segments |
| WO2025061291A1 (en) * | 2023-09-21 | 2025-03-27 | Telefonaktiebolaget Lm Ericsson (Publ) | Systems and methods for generating and presenting a summary of a stream-of-interest |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115033736B (en) * | 2022-06-07 | 2025-04-15 | 浙江大学 | A natural language guided video summarization method |
| CN116089654B (en) * | 2023-04-07 | 2023-07-07 | 杭州东上智能科技有限公司 | A transferable audiovisual text generation method and system based on audio supervision |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130282747A1 (en) * | 2012-04-23 | 2013-10-24 | Sri International | Classification, search, and retrieval of complex video events |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8051446B1 (en) * | 1999-12-06 | 2011-11-01 | Sharp Laboratories Of America, Inc. | Method of creating a semantic video summary using information from secondary sources |
| US8869198B2 (en) * | 2011-09-28 | 2014-10-21 | Vilynx, Inc. | Producing video bits for space time video summary |
| KR102025362B1 (en) * | 2013-11-07 | 2019-09-25 | 한화테크윈 주식회사 | Search System and Video Search method |
-
2017
- 2017-01-06 GB GB1700265.0A patent/GB2558582A/en not_active Withdrawn
-
2018
- 2018-01-02 WO PCT/FI2018/050001 patent/WO2018127627A1/en not_active Ceased
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130282747A1 (en) * | 2012-04-23 | 2013-10-24 | Sri International | Classification, search, and retrieval of complex video events |
Non-Patent Citations (14)
| Title |
|---|
| HU , WEIMING ET AL.: "A Survey on Visual Content-Based Video Indexing and Retrieval", IEEE TRANSACTIONS ON SYSTEMS, MAN, AND CYBERNETICS, PART C (APPLICATIONS AND REVIEWS, vol. 41, no. 6, 10 March 2011 (2011-03-10), pages 797 - 819, XP011363202, DOI: doi:10.1109/TSMCC.2011.2109710 * |
| HU , WEIMING ET AL.: "A Survey on Visual Content-Based Video Indexing and Retrieval", IEEE TRANSACTIONS ON SYSTEMS, MAN, AND CYBERNETICS, PART C (APPLICATIONS AND REVIEWS, vol. 41, no. 6, 10 March 2011 (2011-03-10), pages 797 - 819, XP011479468, ISSN: 1094-6977, [retrieved on 20180417] * |
| LIN, DAHUA ET AL.: "Visual Semantic Search: Retrieving Videos via Complex Textual Queries", PROCEEDINGS OF THE 2014 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2014, 23 June 2014 (2014-06-23), COLUMBUS, OH, USA . LOS ALAMITOS, CA , USA, pages 2657 - 2664, XP032649129, DOI: doi:10.1109/CVPR.2014.340 * |
| LIN, DAHUA ET AL.: "Visual Semantic Search: Retrieving Videos via Complex Textual Queries", PROCEEDINGS OF THE 2014 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2014, 23 June 2014 (2014-06-23), Columbus, OH, USA . Los Alamitos, CA , USA, pages 2657 - 2664, XP032649129, ISSN: 1063-6919, ISBN: 978-1-4799-5118-5 * |
| MUN, JONGHWAN ET AL.: "MarioQA: Answering Questions by Watching Gameplay Videos", ARXIV.ORG, 6 December 2016 (2016-12-06), Retrieved from the Internet <URL:https://arxiv.org/pdf/1612.01669v1> [retrieved on 20180411] * |
| MUN, JONGHWAN ET AL.: "MarioQA: Answering Questions by Watching Gameplay Videos", ARXIV.ORG, 6 December 2016 (2016-12-06), XP080737096, Retrieved from the Internet <URL:https://arxiv.org/pdf/1612.01669v1> [retrieved on 20180411] * |
| SHARGHI, AIDEAN ET AL.: "Query-Focused Extractive Video Summarization", COMPUTER VISION - ECCV 2016, LECTURE NOTES IN COMPUTER SCIENCE, vol. 9912, 17 September 2016 (2016-09-17), SWITZERLAND, pages 3 - 19, XP047362379, DOI: doi:10.1007/978-3-319-46484-8_1 * |
| SHARGHI, AIDEAN ET AL.: "Query-Focused Extractive Video Summarization", COMPUTER VISION - ECCV 2016, LECTURE NOTES IN COMPUTER SCIENCE, vol. 9912, 17 September 2016 (2016-09-17), Switzerland, pages 3 - 19, XP047362379, ISBN: 978-3-319-46484-8, [retrieved on 20180417] * |
| TAPASWI, MAKARAND ET AL.: "Aligning plot synopses to videos for story-based retrieval", INTERNATIONAL JOURNAL OF MULTIMEDIA INFORMATION RETRIEVAL (2015), vol. 4, no. 1, March 2014 (2014-03-01), LONDON, pages 3 - 16, XP055511405 * |
| TAPASWI, MAKARAND ET AL.: "Aligning plot synopses to videos for story-based retrieval", INTERNATIONAL JOURNAL OF MULTIMEDIA INFORMATION RETRIEVAL (2015), vol. 4, no. 1, March 2014 (2014-03-01), London, pages 3 - 16, XP055511405, ISSN: 2192-662X, [retrieved on 20180416] * |
| TAPASWI, MAKARAND ET AL.: "MovieQA: Understanding Stories in Movies through Question-Answering", PROCEEDINGS OF THE 29TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2016, 26 June 2016 (2016-06-26), LAS VEGAS, NV, USA . LOS ALAMITOS, CA , USA, pages 4631 - 4640, XP033021654, DOI: doi:10.1109/CVPR.2016.501 * |
| TAPASWI, MAKARAND ET AL.: "MovieQA: Understanding Stories in Movies through Question-Answering", PROCEEDINGS OF THE 29TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2016, 26 June 2016 (2016-06-26), Las Vegas, NV, USA . Los Alamitos, CA , USA, pages 4631 - 4640, XP033021654, ISSN: 1063-6919, ISBN: 978-1-4673-8851-1, [retrieved on 20180417] * |
| TU, KEWEI ET AL.: "Joint Video and Text Parsing for Understanding Events and Answering Queries", IEEE MULTIMEDIA, 21 February 2014 (2014-02-21), CORNELL UNIVERSITY LIBRARY, pages 42 - 70, XP055457749, Retrieved from the Internet <URL:https://arxiv.org/pdf/1308.6628v2> [retrieved on 20180416] * |
| YU , YOUNGJAE ET AL.: "End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering", IEEE COMPUTER SOCIETY CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION. PROCEEDINGS, 21 July 2017 (2017-07-21), pages 3261 - 3269, XP033249673, Retrieved from the Internet <URL:https://arxiv.org/pdf/1610.02947v2> [retrieved on 20180416] * |
Cited By (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11663268B2 (en) | 2018-03-22 | 2023-05-30 | Guangdong Oppo Mobile Telecommunications Corp., Ltd. | Method and system for retrieving video temporal segments |
| CN109413448A (en) * | 2018-11-05 | 2019-03-01 | 中山大学 | Mobile device panoramic video play system based on deeply study |
| CN109871124B (en) * | 2019-01-25 | 2020-10-27 | 华南理工大学 | Emotional virtual reality scene evaluation method based on deep learning |
| CN109871124A (en) * | 2019-01-25 | 2019-06-11 | 华南理工大学 | Emotion virtual reality scenario appraisal procedure based on deep learning |
| CN109889923A (en) * | 2019-02-28 | 2019-06-14 | 杭州一知智能科技有限公司 | A method for summarizing videos using a hierarchical self-attention network combined with video descriptions |
| CN109889923B (en) * | 2019-02-28 | 2021-03-26 | 杭州一知智能科技有限公司 | A method for summarizing videos using a hierarchical self-attention network combined with video descriptions |
| CN109919114A (en) * | 2019-03-14 | 2019-06-21 | 浙江大学 | One kind is based on the decoded video presentation method of complementary attention mechanism cyclic convolution |
| WO2020197853A1 (en) * | 2019-03-22 | 2020-10-01 | Nec Laboratories America, Inc. | Efficient and fine-grained video retrieval |
| US11568247B2 (en) | 2019-03-22 | 2023-01-31 | Nec Corporation | Efficient and fine-grained video retrieval |
| CN110267051A (en) * | 2019-05-16 | 2019-09-20 | 北京奇艺世纪科技有限公司 | A kind of method and device of data processing |
| CN110414377A (en) * | 2019-07-09 | 2019-11-05 | 武汉科技大学 | A Scene Classification Method for Remote Sensing Image Based on Scale Attention Network |
| CN110933518A (en) * | 2019-12-11 | 2020-03-27 | 浙江大学 | Method for generating query-oriented video abstract by using convolutional multi-layer attention network mechanism |
| CN111241410A (en) * | 2020-01-22 | 2020-06-05 | 深圳司南数据服务有限公司 | Industry news recommendation method and terminal |
| CN111241410B (en) * | 2020-01-22 | 2023-08-22 | 深圳司南数据服务有限公司 | Industry news recommendation method and terminal |
| CN112016493A (en) * | 2020-09-03 | 2020-12-01 | 科大讯飞股份有限公司 | Image description method and device, electronic equipment and storage medium |
| US11651591B2 (en) | 2020-12-22 | 2023-05-16 | Beijing Dajia Internet Information Technology Co., Ltd. | Video timing labeling method, electronic device and storage medium |
| WO2022134634A1 (en) * | 2020-12-22 | 2022-06-30 | 北京达佳互联信息技术有限公司 | Video processing method and electronic device |
| CN113343821B (en) * | 2021-05-31 | 2022-08-30 | 合肥工业大学 | Non-contact heart rate measurement method based on space-time attention network and input optimization |
| CN113343821A (en) * | 2021-05-31 | 2021-09-03 | 合肥工业大学 | Non-contact heart rate measurement method based on space-time attention network and input optimization |
| CN115334367A (en) * | 2022-07-11 | 2022-11-11 | 北京达佳互联信息技术有限公司 | Video summary information generation method, device, server and storage medium |
| CN115334367B (en) * | 2022-07-11 | 2023-10-17 | 北京达佳互联信息技术有限公司 | Method, device, server and storage medium for generating abstract information of video |
| WO2025061291A1 (en) * | 2023-09-21 | 2025-03-27 | Telefonaktiebolaget Lm Ericsson (Publ) | Systems and methods for generating and presenting a summary of a stream-of-interest |
Also Published As
| Publication number | Publication date |
|---|---|
| GB2558582A (en) | 2018-07-18 |
| GB201700265D0 (en) | 2017-02-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2018127627A1 (en) | Method and apparatus for automatic video summarisation | |
| US12382115B2 (en) | Machine learning based media content annotation | |
| CN108763325B (en) | A kind of network object processing method and processing device | |
| CN113689440B (en) | Video processing method, device, computer equipment and storage medium | |
| CN109618222B (en) | A kind of splicing video generation method, device, terminal device and storage medium | |
| CN109740670B (en) | Video classification method and device | |
| CN114390217B (en) | Video synthesis method, device, computer equipment and storage medium | |
| CN108307229B (en) | Video and audio data processing method and device | |
| US20170065889A1 (en) | Identifying And Extracting Video Game Highlights Based On Audio Analysis | |
| US11556302B2 (en) | Electronic apparatus, document displaying method thereof and non-transitory computer readable recording medium | |
| US20170300752A1 (en) | Method and system for summarizing multimedia content | |
| US20250148658A1 (en) | Content generation method and apparatus, electronic device, and storage medium | |
| KR101617649B1 (en) | Recommendation system and method for video interesting section | |
| CN117609550A (en) | Video title generation method and training method of video title generation model | |
| CN116389849A (en) | Video generation method, device, equipment and storage medium | |
| CN117593473B (en) | Method, apparatus and storage medium for generating motion image and video | |
| WO2020011001A1 (en) | Image processing method and device, storage medium and computer device | |
| CN115359409A (en) | Video splitting method and device, computer equipment and storage medium | |
| CN114363694B (en) | Video processing method, device, computer equipment and storage medium | |
| CN110418148A (en) | Video generation method, video generation device and readable storage medium | |
| CN116612060B (en) | Video information processing method, device and storage medium | |
| US20250014204A1 (en) | Video engagement determination based on statistical positional object tracking | |
| CN114821794B (en) | Image processing method, model generating method, image processing device, vehicle, and storage medium | |
| CN110275988A (en) | Obtain the method and device of picture | |
| CN118227747A (en) | A text-based interactive method, device, equipment and storage medium |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18736142 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18736142 Country of ref document: EP Kind code of ref document: A1 |