WO2021184852A1 - 动作区域提取方法、装置、设备及计算机可读存储介质 - Google Patents

动作区域提取方法、装置、设备及计算机可读存储介质 Download PDF

Info

Publication number
WO2021184852A1
WO2021184852A1 PCT/CN2020/136320 CN2020136320W WO2021184852A1 WO 2021184852 A1 WO2021184852 A1 WO 2021184852A1 CN 2020136320 W CN2020136320 W CN 2020136320W WO 2021184852 A1 WO2021184852 A1 WO 2021184852A1
Authority
WO
WIPO (PCT)
Prior art keywords
action
action area
video
time period
sequence
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/136320
Other languages
English (en)
French (fr)
Inventor
张国辉
朱文和
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2021184852A1 publication Critical patent/WO2021184852A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/41Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs

Definitions

  • This application relates to the field of image processing technology, and in particular to a method, device, device, and computer-readable storage medium for extracting an action region.
  • Video content analysis is currently a hot research topic in the field of AI (Artificial Intelligence).
  • AI Artificial Intelligence
  • action recognition is an important branch of video analysis in many fields such as intelligent video surveillance, human-computer interaction, motion analysis, and video retrieval. , Has broad application prospects and has received extensive attention from scholars at home and abroad.
  • the video needs to be edited first, and multiple video clips containing only one action instance are obtained.
  • the video recorded in the real scene is usually very long and contains a lot of content that has nothing to do with the action instance.
  • timing detection is usually used to detect the action instances in the untrimmed video.
  • the timing detection task can be divided into two stages: the extraction and classification of the action area.
  • the extraction stage of the action area aims to extract the video area containing the action instance
  • the classification stage is to classify the action area. Therefore, obtaining a high-quality action area is the key to ensuring the accuracy of the action instance detection result.
  • the present application provides a method for extracting an action area.
  • the method for extracting an action area includes:
  • the corresponding action area is extracted from the video to be modified according to the time period information.
  • the present application also provides a device for extracting an action area, the device for extracting an action area includes:
  • a feature extraction module configured to obtain a video to be trimmed, and perform feature extraction on the video to be trimmed to obtain a first feature sequence
  • An information acquisition module configured to input the first characteristic sequence into a preset time sequence evaluation model to obtain time sequence information
  • the information detection module is used to detect whether the time sequence information meets preset conditions, and obtain the time period information of the action area according to the detection result;
  • the region extraction module is configured to extract the corresponding action region from the video to be modified according to the time period information.
  • the present application also provides an action area extraction device.
  • the action area extraction device includes a memory, a processor, and an action area extraction program stored on the memory and executable by the processor, wherein the action area extraction program When executed by the processor, the following steps are implemented:
  • the corresponding action area is extracted from the video to be modified according to the time period information.
  • the present application also provides a computer-readable storage medium on which an action area extraction program is stored, wherein when the action area extraction program is executed by a processor, the following steps are implemented:
  • the corresponding action area is extracted from the video to be modified according to the time period information.
  • FIG. 1 is a schematic diagram of a device structure of a hardware operating environment involved in a solution of an embodiment of the application
  • FIG. 2 is a schematic flowchart of a first embodiment of a method for extracting an action area according to this application;
  • step S10 is a schematic diagram of the detailed flow of step S10 in the first embodiment of the application.
  • step S30 is a schematic diagram of the detailed flow of step S30 in the first embodiment of the application.
  • FIG. 5 is a schematic flowchart of a second embodiment of a method for extracting an action area according to this application.
  • Fig. 6 is a schematic diagram of the functional modules of the first embodiment of the action area extraction device of this application.
  • FIG. 1 is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the application.
  • the action area extraction device involved in the embodiment of the present application may be a terminal device such as a PC (personal computer, personal computer), a notebook computer, and a server.
  • a terminal device such as a PC (personal computer, personal computer), a notebook computer, and a server.
  • the action area extraction device may include: a processor 1001, such as a CPU (Central Processing Unit, central processing unit), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005.
  • the communication bus 1002 is used to realize the connection and communication between these components;
  • the user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. .
  • the network interface 1004 can optionally include a standard wired interface and a wireless interface (such as Wireless-Fidelity, Wi-Fi interface); the memory 1005 can be a high-speed random access memory (random access memory, RAM) or A non-volatile memory (non-volatile memory), such as a disk memory, and the memory 1005 may optionally also be a storage device independent of the aforementioned processor 1001.
  • a standard wired interface and a wireless interface such as Wireless-Fidelity, Wi-Fi interface
  • the memory 1005 can be a high-speed random access memory (random access memory, RAM) or A non-volatile memory (non-volatile memory), such as a disk memory, and the memory 1005 may optionally also be a storage device independent of the aforementioned processor 1001.
  • RAM random access memory
  • non-volatile memory non-volatile memory
  • the structure of the action area extraction device shown in FIG. 1 does not constitute a limitation on the action area extraction device, and may include more or
  • the memory 1005 as a computer storage medium in FIG. 1 may include an operating system, a network communication module, and an action area extraction program.
  • the network communication module can be used to connect to the server and perform data communication with the server; and the processor 1001 can be used to call the action area extraction program stored in the memory 1005 and execute the action area extraction method provided in the embodiment of the present application.
  • This application provides a method for extracting an action area.
  • FIG. 2 is a schematic flowchart of a first embodiment of a method for extracting an action area according to this application.
  • the action area extraction method includes:
  • Step S10 Obtain a video to be trimmed, and perform feature extraction on the video to be trimmed to obtain a first feature sequence
  • the action area extraction method is implemented by an action area extraction device.
  • the action area extraction device may be a PC, a notebook computer, a server, etc.
  • the action area extraction device is described by taking a server as an example.
  • the method for extracting an action area in the embodiments of the present application can be applied to scenes such as security, surveillance, and wonderful action screen editing, and is used to extract an area containing an action from a video.
  • the video to be trimmed is first acquired, and then the feature extraction is performed on the video to be trimmed to obtain the first feature sequence.
  • the video to be trimmed is divided into frames first to obtain a video image sequence; then, the target video image is obtained every preset frame number in the video image sequence, and the RGB (Red-Green-Blue, red-green-blue) of the target video image is extracted.
  • (Blue) feature and optical flow feature to obtain the RGB feature sequence and the optical flow feature sequence; then the RGB feature sequence and the optical flow feature sequence are spliced to obtain the first feature sequence.
  • Step S20 input the first characteristic sequence into a preset time sequence evaluation model to obtain time sequence information
  • the first characteristic sequence is input to a preset time sequence evaluation model to obtain time sequence information.
  • the timing information includes the start probability of the action segment corresponding to each target video image, the intermediate probability of the action segment, and the end probability of the action segment.
  • the start probability of the action segment corresponding to each target video image means that each target video image belongs to the beginning part of the action segment.
  • the probability of each target video image corresponding to the middle of the action segment that is, the probability that each target video image belongs to the middle part of the action segment;
  • the end probability of each target video image is the end probability of each target video image, that is, each target video image belongs to the end part of the action segment The probability.
  • the start probability of the action segment, the middle probability of the action segment, and the end probability of the action segment are denoted as Ps, Pm, and Pe, respectively.
  • the preset time series evaluation model is based on training samples and a pre-built convolutional neural network model.
  • the preset time series evaluation model consists of 3 layers of convolutional layers. The first 2 layers are exactly the same, and the number of filters is 512.
  • the convolution kernel is 3, the activation function is Relu, and the step size is 1.
  • the number of filters in the last layer is 3, the convolution kernel is 3, and the activation function is Sigmoid, as follows:
  • the three filters of the last convolutional layer respectively output three probability values of start, middle, and end.
  • the Loss function is composed of 3 Loss: start (s), middle (m), and end (e). Can be expressed as:
  • L is the binary logistic regression function. Can be expressed as:
  • bi is an indicator function, which is 1 when it is true and 0 when it is false.
  • Step S30 detecting whether the time sequence information meets a preset condition, and obtaining time period information of the action area according to the detection result;
  • the time period information is the video time period information belonging to the action area, which may include multiple groups of time periods, and each time period is composed of the start time and the end time of the action segment.
  • the detection methods it is possible to detect whether there is a probability value greater than the first preset threshold in the start probability of the action segment, obtain the first detection result, and obtain the start time array of the action segment according to the first detection result; at the same time, Detect whether there is a probability value greater than the second preset threshold in the end probability of the action segment, obtain the second detection result, and obtain the end time array of the action segment according to the second detection result; then, according to the start time array of the action segment and the end time of the action segment Array, combined to get the time period information of the action area.
  • the time segment information of the action area is obtained by combining.
  • the detection conditions of the above detection methods can also be combined in the specific implementation. Only when any one of the above two detection conditions is met, or the above two conditions need to be met at the same time, the time of the corresponding target video image is stored To the start/end time array of the action segment, and then get the time period information.
  • the specific detection process can refer to the following embodiments, which will not be repeated here.
  • Step S40 Extract a corresponding action area from the video to be modified according to the time period information.
  • the corresponding action area is extracted from the video to be modified according to the time period information.
  • the embodiment of the present application provides a method for extracting an action region.
  • feature extraction is performed on the video to be trimmed to obtain a first feature sequence; then, the first feature sequence is input to a preset timing evaluation model to obtain timing information; It detects whether the timing information meets the preset condition, and obtains the time period information of the action area according to the detection result; and then extracts the corresponding action area from the video to be modified according to the time period information.
  • the action area extraction method of the embodiment of the present application is more Flexible, the precision and accuracy of the extraction results are higher.
  • FIG. 3 is a detailed flowchart of step S10 in the first embodiment of this application.
  • step S10 includes:
  • Step S11 Obtain a video to be trimmed, and perform framing processing on the video to be trimmed to obtain a video image sequence
  • the video to be trimmed is acquired first, and then the to-be trimmed video is subjected to framing processing to obtain a video image sequence.
  • the video image sequence includes frames of video images of the video to be trimmed, and the frames of video images are arranged in chronological order.
  • Step S12 acquiring a target video image every preset number of frames in the video image sequence, extracting the red, green, and blue RGB characteristics and optical flow characteristics of the target video image to obtain the RGB characteristic sequence and the optical flow characteristic sequence;
  • the target video image is acquired every preset number of frames in the video image sequence, and then the RGB (Red-Green-Blue) feature and optical flow feature of the target video image are extracted to obtain the RGB feature sequence and the optical flow feature sequence.
  • the target video image includes multiple, corresponding, RGB features and optical flow features also include multiple groups, and an RGB feature sequence can be formed based on the time sequence of the target video image corresponding to each group of RGB features. Similarly, based on each The time sequence of the target video images corresponding to the group of optical flow features constitutes an optical flow feature sequence.
  • the commonly used TSN Temporal Segment Networks, Time Sensitive Networks
  • the TSN is constructed based on the two stream method, and the specific feature extraction method can refer to the prior art, which will not be repeated here.
  • other types of features can also be extracted for use in extracting the action area.
  • Step S13 splicing the RGB feature sequence and the optical flow feature sequence to obtain a first feature sequence.
  • the RGB feature sequence and the optical flow feature sequence are spliced to obtain the first feature sequence.
  • each RGB feature is a 200*100-dimensional matrix
  • each optical flow feature is a 200*100-dimensional matrix.
  • the first segment of RGB features and the first segment of optical flow features can be spliced together to obtain a 400*100-dimensional matrix, that is, the first feature vector of the first feature sequence F is a 400*100-dimensional matrix.
  • the RGB features and optical flow features of each subsequent segment are spliced to obtain the first feature sequence.
  • the first feature sequence is formed by extracting the features of the corresponding video image every preset frame, which can save calculations and increase the extraction speed of the action area. It can be applied to long videos in security, surveillance and other scenarios. Extraction of the action area.
  • FIG. 4 is a detailed flowchart of step S30 in the first embodiment of this application.
  • the timing information includes the start probability of the action segment and the end probability of the action segment corresponding to each target video image
  • step S30 includes:
  • Step S31 Detect whether there is a probability value greater than a first preset threshold in the start probability of the action segment, obtain a first detection result, and obtain an action segment start time array according to the first detection result;
  • the timing information includes the start probability of the action segment and the end probability of the action segment corresponding to each target video image.
  • the start probability of the action segment corresponding to each target video image is the start probability of each target video image belonging to the beginning of the action segment.
  • Probability; the end probability of the action segment corresponding to each target video image that is, the probability that each target video image belongs to the end part of the action segment.
  • the timing information also includes the intermediate probability of the action segment corresponding to each target video image.
  • the starting probability of the action segment is denoted as Ps
  • the ending probability of the action segment is denoted as Pe. It can be understood that, in the above example, since there are L target video images acquired, correspondingly, there are L Ps and Pe.
  • the first detection result is obtained, and the start time array of the action segment is obtained according to the first detection result.
  • the first preset threshold is preset, which is not specifically limited here. Obtain Ps greater than the first preset threshold, and then determine the time Ts of the target video image corresponding to the Ps greater than the first preset threshold, and then form an action segment start time array ⁇ Ts ⁇ based on the determined Ts.
  • the start time array of the action segments may also be used. For example, it is detected whether there is a peak Ps at a certain moment, that is, greater than the Ps at the previous moment and the next moment. If the condition is met, the time Ts of the target video image corresponding to the Ps that meets the condition is stored in the array, Get the start time array ⁇ Ts ⁇ of the action segment.
  • the above two conditions can also be combined for detection. Only if any one of the above two conditions is met, or the above two conditions must be met at the same time, then the time of the corresponding target video image is stored to the start time of the action segment In the array ⁇ Ts ⁇ .
  • Step S32 Detect whether there is a probability value greater than a second preset threshold in the end probability of the action segment, obtain a second detection result, and obtain an action segment end time array according to the second detection result;
  • the second preset threshold is preset, which is not specifically limited here.
  • the second preset threshold may be the same as or different from the first preset threshold.
  • other detection rules may also be used to obtain an array of end time of the action segments. For example, detecting whether there is a peak value of Pe at a certain moment, that is, greater than the Pe at the previous moment and the next moment, if the condition is met, the time Te of the target video image corresponding to the eligible Pe is stored in the array, Get the end time array ⁇ Te ⁇ of the action segment.
  • the above two conditions can also be combined for detection. Only if any one of the above two conditions is met, or the above two conditions need to be met at the same time, then the time of the corresponding target video image is stored to the end time of the action segment In the array ⁇ Te ⁇ .
  • step S31 and step S32 is in no particular order.
  • Step S33 Combine the action segment start time array and the action segment end time array to obtain time segment information of the action area.
  • the action segment start time array ⁇ Ts ⁇ and the action segment end time array ⁇ Te ⁇ are obtained, the action segment start time array ⁇ Ts ⁇ and the action segment end time array ⁇ Te ⁇ are combined to obtain the time period information of the action area. Specifically, one value from each of ⁇ Ts ⁇ and ⁇ Te ⁇ is selected in turn and combined into an action area time period, and finally time period information of multiple action areas can be obtained, which can be expressed in the form of an action area time period array. Obviously, Te needs to be greater than Ts each time it is selected to form an array of action area time periods.
  • a high-probability boundary position that is, the start time and end time of the action area, is obtained, so as to facilitate subsequent extraction of the precise action area .
  • FIG. 5 is a schematic flowchart of a second embodiment of an action area extraction method according to this application.
  • the action area extraction method further includes:
  • Step S50 sampling the features of each action area based on the first feature sequence and the time period information to obtain a second feature sequence
  • the candidate area-level features are evaluated in the global scope to obtain reliable
  • the action area confidence score is retrieved to obtain the action area, which can further improve the accuracy and accuracy of the action area extraction result.
  • the characteristics of each action area are sampled based on the first feature sequence and the time period information to obtain the second feature sequence.
  • step S50 includes:
  • Step a1 sampling the features of each action area by using a linear difference method based on the first feature sequence and the time period information to obtain a first preset number of feature data;
  • Step a2 splicing the first preset number of feature data to obtain a second feature sequence.
  • linear interpolation refers to an interpolation method in which the interpolation function is a polynomial of the first degree, and the interpolation error at the interpolation node is zero.
  • other interpolation methods such as parabolic interpolation, can also be used, but linear interpolation is simpler and more convenient than other interpolation methods.
  • the first preset number is optionally set to 32, of course, it can also be specifically set according to actual conditions, and it is not used as a limitation to the application here.
  • the first preset number of feature data is obtained, the first preset number of feature data is spliced to obtain a second feature sequence.
  • Step S60 input the second feature sequence into a preset action area evaluation model to obtain an action area evaluation score
  • the second feature sequence is input to the preset action area evaluation model to obtain the action area evaluation score.
  • the preset action area evaluation model is obtained by training based on the training samples and the preset candidate area evaluation model.
  • the network structure of the preset action area evaluation model is a two-layer fully connected layer.
  • the first hidden layer has 512
  • the activation function is Relu
  • the second layer activation function is Sigmoid
  • the output is the confidence that the action area contains the action segment, as follows:
  • the Loss function is a simple regression Loss, using the variance of the intersection of the candidate region and the real segment with confidence and the variance of g. It is defined as follows:
  • N is the number of action areas.
  • step S40 includes:
  • Step S41 Extract a corresponding action area from the video to be modified according to the action area evaluation score and the time period information.
  • the corresponding action area is extracted from the video to be modified according to the action area evaluation score and time period information.
  • step S41 includes:
  • Step b1 sort the evaluation scores of the action areas in descending order
  • Step b2 Obtain a second preset number of action area evaluation scores according to the sorting result as the target action area evaluation score
  • Step b3 Obtain target time period information corresponding to the target action area evaluation score from the time period information;
  • Step b4 Extract a corresponding action area from the video to be modified according to the target time period information.
  • the extraction process of the action area is as follows: first sort the action area evaluation scores in descending order, obtain the second preset number of action area evaluation scores according to the sorting result, and use them as the target action area evaluation scores. Then, from time The target time period information corresponding to the evaluation score of the target action area is obtained from the segment information; further, the corresponding action area is extracted from the video to be modified according to the target time period information.
  • the candidate area-level features are further evaluated in the global scope to obtain reliable action area confidence scores for retrieval, so as to obtain the action area, which can further improve the action area extraction The precision and accuracy of the results.
  • the application also provides an action area extraction device.
  • FIG. 6 is a schematic diagram of the functional modules of the first embodiment of the action area extraction device of this application.
  • the action area extraction device includes:
  • the feature extraction module 10 is configured to obtain a video to be trimmed, and perform feature extraction on the video to be trimmed to obtain a first feature sequence
  • the information acquisition module 20 is configured to input the first characteristic sequence into a preset time sequence evaluation model to obtain time sequence information;
  • the information detection module 30 is configured to detect whether the timing information meets a preset condition, and obtain the time period information of the action area according to the detection result;
  • the region extraction module 40 is configured to extract a corresponding action region from the to-be-modified video according to the time period information.
  • each virtual function module of the above-mentioned action area extraction device is stored in the memory 1005 of the action area extraction device shown in FIG.
  • the function of the accuracy of the extraction result of the action area is stored in the memory 1005 of the action area extraction device shown in FIG. The function of the accuracy of the extraction result of the action area.
  • the feature extraction module 10 includes:
  • a framing processing unit configured to obtain a video to be trimmed, and perform framing processing on the video to be trimmed to obtain a video image sequence
  • the feature extraction unit is configured to obtain a target video image every preset number of frames in the video image sequence, extract the red, green, and blue RGB features and optical flow features of the target video image to obtain the RGB feature sequence and the optical flow feature sequence;
  • the first splicing unit is used for splicing the RGB characteristic sequence and the optical flow characteristic sequence to obtain a first characteristic sequence.
  • timing information includes the start probability of the action segment and the end probability of the action segment corresponding to each target video image
  • information detection module 30 includes:
  • the first detection unit is configured to detect whether there is a probability value greater than a first preset threshold in the start probability of the action segment, obtain a first detection result, and obtain an action segment start time array according to the first detection result;
  • the second detection unit is configured to detect whether there is a probability value greater than a second preset threshold in the end probability of the action segment, obtain a second detection result, and obtain an action segment end time array according to the second detection result;
  • the information combination unit is used to combine to obtain the time period information of the action area according to the start time array of the action segments and the end time array of the action segments.
  • the action area extraction device further includes:
  • a feature sampling module configured to sample features of each action area based on the first feature sequence and the time period information to obtain a second feature sequence
  • a score evaluation module configured to input the second feature sequence into a preset action area evaluation model to obtain an action area evaluation score
  • the region extraction module 40 is specifically configured to extract a corresponding action region from the video to be modified according to the action region evaluation score and the time period information.
  • the feature sampling module includes:
  • the feature sampling unit is configured to sample the features of each action area by using a linear difference method based on the first feature sequence and the time period information to obtain a first preset number of feature data;
  • the second splicing unit is used to splice the first preset number of feature data to obtain a second feature sequence.
  • region extraction module 40 includes:
  • the score sorting unit is used to sort the evaluation scores of the action area in descending order
  • the first obtaining unit is configured to obtain a second preset number of previous action area evaluation scores according to the sorting result, as the target action area evaluation score;
  • the second acquiring unit is configured to acquire target time period information corresponding to the target action area evaluation score from the time period information;
  • the area extraction unit is configured to extract the corresponding action area from the video to be modified according to the target time period information.
  • each module in the above-mentioned action region extraction device corresponds to each step in the above-mentioned embodiment of the action region extraction method, and the functions and implementation processes thereof will not be repeated here.
  • the present application also provides a computer-readable storage medium.
  • the computer-readable storage medium may be volatile or non-volatile.
  • the computer-readable storage medium stores an action area extraction program, and the action area When the extraction program is executed by the processor, the steps of the action area extraction method as described in any of the above embodiments are implemented.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Signal Processing (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Psychiatry (AREA)
  • Social Psychology (AREA)
  • Human Computer Interaction (AREA)
  • Computational Linguistics (AREA)
  • Image Analysis (AREA)
  • Television Signal Processing For Recording (AREA)

Abstract

一种动作区域提取方法、装置、设备及计算机可读存储介质,涉及图像处理技术领域,该方法包括:获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列(S10);将所述第一特征序列输入至预设时序评估模型,得到时序信息(S20);检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息(S30);根据所述时间段信息从所述待修改视频中提取出对应的动作区域(S40)。该方法能够解决现有的动作区域提取方法精确性较差的问题。

Description

动作区域提取方法、装置、设备及计算机可读存储介质
本申请要求于2020年3月16日提交中国专利局、申请号为CN202010185060.0、名称为“动作区域提取方法、装置、设备及计算机可读存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及图像处理技术领域,尤其涉及一种动作区域提取方法、装置、设备及计算机可读存储介质。
背景技术
视频内容分析是当前AI(Artificial Intelligence,人工智能)领域比较热门的研究课题,其中,行动识别作为视频分析的其中一个重要分支,在智能视频监控、人机交互、运动分析、视频检索等诸多领域,具有广阔的应用前景,受到了国内外学者的广泛关注。
在行动识别的过程中,需先对视频进行剪辑,得到多个仅包含一个动作实例的视频剪辑。但是现实场景中录制得到的视频通常很长、并且包含很多与动作实例无关的内容。此时,通常会通过时序检测的手段来检测未修剪视频中的动作实例。具体的,时序检测任务可以分为两个阶段:动作区域的提取和分类。其中,动作区域的提取阶段旨在提取出包含动作实例的视频区域,分类阶段则是对动作区域进行分类。因此,获取高质量的动作区域,是保证动作实例检测结果准确性的关键。
目前,通常是使用多个持续时间的滑动时间窗口以固定的间隔进行滑动,以提取动作区域,发明人意识到使用预定义的持续时间和间隔提取的动作区域具有以下缺陷:1)时间通常不精确;2)真实场景下动作实例的持续时间是复杂多变的,不能够灵活覆盖,尤其是在时间范围很大的情况下。因此,现有的动作区域提取方法精确性较差的问题。
发明内容
本申请提供一种动作区域提取方法,所述动作区域提取方法包括:
获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
将所述第一特征序列输入至预设时序评估模型,得到时序信息;
检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
本申请还提供一种动作区域提取装置,所述动作区域提取装置包括:
特征提取模块,用于获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
信息获取模块,用于将所述第一特征序列输入至预设时序评估模型,得到时序信息;
信息检测模块,用于检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
区域提取模块,用于根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
本申请还提供一种动作区域提取设备,所述动作区域提取设备包括存储器、处理器以及存储在所述存储器上并可被所述处理器执行的动作区域提取程序,其中所述动作区域提取程序被所述处理器执行时,实现如下步骤:
获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
将所述第一特征序列输入至预设时序评估模型,得到时序信息;
检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
本申请还提供一种计算机可读存储介质,所述计算机可读存储介质上存储有动作区域 提取程序,其中所述动作区域提取程序被处理器执行时,实现如下步骤:
获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
将所述第一特征序列输入至预设时序评估模型,得到时序信息;
检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
附图说明
图1为本申请实施例方案涉及的硬件运行环境的设备结构示意图;
图2为本申请动作区域提取方法第一实施例的流程示意图;
图3为本申请第一实施例中步骤S10的细化流程示意图;
图4为本申请第一实施例中步骤S30的细化流程示意图;
图5为本申请动作区域提取方法第二实施例的流程示意图;
图6为本申请动作区域提取装置第一实施例的功能模块示意图。
本申请目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
具体实施方式
应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
参照图1,图1为本申请实施例方案涉及的硬件运行环境的设备结构示意图。
本申请实施例涉及的动作区域提取设备可以是PC(personal computer,个人计算机)、笔记本电脑、服务器等终端设备。
如图1所示,该动作区域提取设备可以包括:处理器1001,例如CPU(Central Processing Unit,中央处理器),通信总线1002,用户接口1003,网络接口1004,存储器1005。其中,通信总线1002用于实现这些组件之间的连接通信;用户接口1003可以包括显示屏(Display)、输入单元比如键盘(Keyboard),可选用户接口1003还可以包括标准的有线接口、无线接口。网络接口1004可选的可以包括标准的有线接口、无线接口(如无线保真Wireless-Fidelity,Wi-Fi接口);存储器1005可以是高速随机存取存储器(random access memory,RAM),也可以是稳定的存储器(non-volatile memory),例如磁盘存储器,存储器1005可选的还可以是独立于前述处理器1001的存储装置。本领域技术人员可以理解,图1中示出的动作区域提取设备结构并不构成对动作区域提取设备的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
继续参照图1,图1中作为一种计算机存储介质的存储器1005中可以包括操作系统、网络通信模块以及动作区域提取程序。在图1中,网络通信模块可用于连接服务器,与服务器进行数据通信;而处理器1001可以用于调用存储器1005中存储的动作区域提取程序,并执行本申请实施例提供的动作区域提取方法。
基于上述硬件结构,提出本申请动作区域提取方法的各个实施例。
本申请提供一种动作区域提取方法。
请参照图2,图2为本申请动作区域提取方法第一实施例的流程示意图。
在本实施例中,该动作区域提取方法包括:
步骤S10,获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
在本实施例中,该动作区域提取方法由动作区域提取设备实现,该动作区域提取设备可以是PC、笔记本电脑、服务器等设备,该动作区域提取设备以服务器为例进行说明。本申请实施例的动作区域提取方法可应用于安防、监控、精彩动作画面剪辑等场景,用于从视频中提取获得包含动作的区域。
在本实施例中,先获取待修剪视频,然后,对待修剪视频进行特征提取,得到第一特征序列。具体的,先对待修剪视频进行分帧处理,得到视频图像序列;然后,在视频图像序列中每隔预设帧数获取目标视频图像,提取目标视频图像的RGB(Red-Green-Blue,红 绿蓝)特征和光流特征,得到RGB特征序列和光流特征序列;进而对RGB特征序列和光流特征序列进行拼接,得到第一特征序列。第一特征序列的具体获取过程可参照下述实施例,此处不作赘述。
步骤S20,将所述第一特征序列输入至预设时序评估模型,得到时序信息;
在得到第一特征序列之后,将所述第一特征序列输入至预设时序评估模型,得到时序信息。其中,时序信息包括各目标视频图像对应的动作片段开始概率、动作片段中间概率和动作片段结束概率,其中,各目标视频图像对应的动作片段开始概率,即为各目标视频图像属于动作片段开始部分的概率;各目标视频图像对应的动作片段中间概率,即为各目标视频图像属于动作片段中间部分的概率;各目标视频图像对应的动作片段结束概率,即为各目标视频图像属于动作片段结束部分的概率。为便于说明,将动作片段开始概率、动作片段中间概率和动作片段结束概率,分别记为Ps、Pm、Pe。其中,该预预设时序评估模型是基于训练样本和预先构建的卷积神经网络模型训练得到的,该预设时序评估模型由3层卷积层组成,前2层完全一致,过滤器数目为512、卷积核为3,激活函数为Relu,步长为1。最后一层过滤器数目为3,卷积核为3,激活函数为Sigmoid,如下:
Conv(512,3,Relu)→Conv(512,3,Relu)→Conv(3,1,Sigmoid)
其中,最后一层卷积层的3个过滤器,分别输出为开始、中间、结束3种概率值。
Loss函数由开始(s)、中间(m)、结束(e)3个Loss组成。可表示为:
J=λL(s)+L(m)+L(e)
L就是二元逻辑回归函数。可表示为:
L=∑bi*log(Pi)+(1-bi)*log(1-Pi)
其中,bi为指示函数,当为真时为1,假时为0。
步骤S30,检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
然后,检测时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息。其中,时间段信息即为属于动作区域的视频时间段信息,可包括多组时间段,每一时间段由动作片段的开始时间和结束时间组成。
具体的,作为其中一种检测方式,可以检测动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据第一检测结果得到动作片段开始时间数组;同时,检测动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据第二检测结果得到动作片段结束时间数组;然后,根据动作片段开始时间数组、动作片段结束时间数组,组合得到动作区域的时间段信息。
作为另一种检测方式,可以检测动作片段开始概率中是否存在峰值,即同时大于前一时刻和后一时刻的动作片段开始概率,得到第一检测结果,并根据第一检测结果得到动作片段开始时间数组;同时,检测动作片段结束概率中是否存在峰值,即同时大于前一时刻和后一时刻的动作片段结束概率,得到第二检测结果,并根据第二检测结果得到动作片段结束时间数组;然后,根据动作片段开始时间数组、动作片段结束时间数组,组合得到动作区域的时间段信息。
当然,在具体实施时也可以结合上述检测方式的检测条件,只需符合上述2种检测条件中的任一种,或需同时符合上述2种条件,才将其对应的目标视频图像的时间存储至动作片段开始/结束时间数组中,进而得到时间段信息。具体的检测过程可以参照下述实施例,此处不作赘述。
步骤S40,根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
在获取到时间段信息之后,根据时间段信息从待修改视频中提取出对应的动作区域。
本申请实施例提供一种动作区域提取方法,通过获取待修剪视频,对待修剪视频进行特征提取,得到第一特征序列;然后,将第一特征序列输入至预设时序评估模型,得到时 序信息;检测时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;进而根据时间段信息从待修改视频中提取出对应的动作区域。本申请实施例中通过提取特征,然后获取对应的时序信息,即各特征对应的时间位置分别属于动作片段开始、中间和结束的概率,进而筛选出高概率的边界位置,即动作区域的开始时间和结束时间,以提取得到精确的动作区域,相比于现有技术中采用使用多个持续时间的滑动时间窗口以固定的间隔进行滑动来提取动作区域,本申请实施例的动作区域提取方法更加灵活,提取结果的精确性和准确性更高。
进一步的,参照图3,图3为本申请第一实施例中步骤S10的细化流程示意图;
在本实施例中,步骤S10包括:
步骤S11,获取待修剪视频,对所述待修剪视频进行分帧处理,得到视频图像序列;
在本实施例中,先获取待修剪视频,然后对待修剪视频进行分帧处理,得到视频图像序列。其中,视频图像序列包括待修剪视频的各帧视频图像、各帧视频图像之间按时间顺序排列。
步骤S12,在所述视频图像序列中每隔预设帧数获取目标视频图像,提取所述目标视频图像的红绿蓝RGB特征和光流特征,得到RGB特征序列和光流特征序列;
在视频图像序列中每隔预设帧数获取目标视频图像,然后,提取目标视频图像的RGB(Red-Green-Blue,红绿蓝)特征和光流特征,得到RGB特征序列和光流特征序列。可以理解,由于目标视频图像包括多个,对应的,RGB特征和光流特征也包括多组,可基于各组RGB特征对应的目标视频图像的时间顺序,构成一RGB特征序列,类似地,基于各组光流特征对应的目标视频图像的时间顺序,构成一光流特征序列。
在获取目标视频图像时,假设待修剪视频有N帧,为了节约计算量,可以设定每隔M帧提取一次,因此,可获取得到L=N/M张目标视频图像,对应的,提取得到的特征也有L段。
对于RGB特征和光流特征的提取,可以采用常用的TSN(Temporal Segment Networks,时间敏感网络)算法对目标视频图像进行RGB特征、光流特征两种特征的提取。其中,TSN是基于two stream(双流)方法构建的,具体的特征提取方法可参照现有技术,此处不作赘述。此外,需要说明的是,在具体实施例中,还可以提取其他类型的特征,以用于提取动作区域。
步骤S13,对所述RGB特征序列和所述光流特征序列进行拼接,得到第一特征序列。
在得到RGB特征序列和光流特征序列后,对RGB特征序列和光流特征序列进行拼接,得到第一特征序列。例如,上述例中,得到L段RGB特征序列和L段光流特征序列后,每一段RGB特征为200*100维度的矩阵,每一段光流特征为200*100维度的矩阵,在拼接时,可以先对第一段RGB特征和第一段光流特征进行拼接,得到一400*100维度的矩阵,即第一特征序列F的第一个特征向量为400*100维度的矩阵,同样地,对后续每一段RGB特征和光流特征进行拼接,得到第一特征序列。
通过上述方式,对待修剪视频进行特征提取,得到对应的第一特征序列,以便于后续对目标动作区域进行提取。同时,本实施例中,通过每隔预设帧提取对应视频图像的特征,形成第一特征序列,可节约计算量,提高动作区域的提取速度,可适用于安防、监控等场景下长视频中的动作区域提取。
进一步地,参照图4,图4为本申请第一实施例中步骤S30的细化流程示意图
在本实施例中,所述时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结束概率,步骤S30包括:
步骤S31,检测所述动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组;
在本实施例中,时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结 束概率,其中,各目标视频图像对应的动作片段开始概率,即为各目标视频图像属于动作片段开始部分的概率;各目标视频图像对应的动作片段结束概率,即为各目标视频图像属于动作片段结束部分的概率。此外,时序信息还包括各目标视频图像对应的动作片段中间概率。为便于说明,将动作片段开始概率记为Ps,动作片段结束概率记为Pe。可以理解,上述例中,由于获取到的目标视频图像有L张,对应的,Ps和Pe有L个。
检测动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组。其中,第一预设阈值是预先设定,此处不作具体限定。获取大于第一预设阈值的Ps,进而确定该大于第一预设阈值的Ps所对应的目标视频图像的时间Ts,进而基于确定得到的Ts组成一动作片段开始时间数组{Ts}。
当然,在具体实施例中,还可以采用其他检测规则,以获取动作片段开始时间数组。例如,检测是否存在某一时刻的Ps为峰值,即大于前一时刻和后一时刻的Ps,若符合该条件,则将符合条件的Ps所对应的目标视频图像的时间Ts存储到数组中,得到动作片段开始时间数组{Ts}。当然,也可以结合上述2种条件进行检测,只需符合上述2种条件中的任一种,或需同时符合上述2种条件,才将其对应的目标视频图像的时间存储至动作片段开始时间数组{Ts}中。
步骤S32,检测所述动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组;
检测动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组。其中,第二预设阈值是预先设定,此处不作具体限定,第二预设阈值可以与第一预设阈值相同,也可以不同。获取大于第二预设阈值的Pe,进而确定该大于第二预设阈值的Pe所对应的目标视频图像的时间Te,进而基于确定得到的Te组成一动作片段结束时间数组{Te}。
同样的,在具体实施例中,也可以采用其他检测规则,以获取动作片段结束时间数组。例如,检测是否存在某一时刻的Pe为峰值,即大于前一时刻和后一时刻的Pe,若符合该条件,则将符合条件的Pe所对应的目标视频图像的时间Te存储到数组中,得到动作片段结束时间数组{Te}。当然,也可以结合上述2种条件进行检测,只需符合上述2种条件中的任一种,或需同时符合上述2种条件,才将其对应的目标视频图像的时间存储至动作片段结束时间数组{Te}中。
需要说明的是,步骤S31和步骤S32的执行顺序不分先后。
步骤S33,根据所述动作片段开始时间数组、所述动作片段结束时间数组,组合得到动作区域的时间段信息。
在得到动作片段开始时间数组{Ts}和动作片段结束时间数组{Te}之后,根据动作片段开始时间数组{Ts}、动作片段结束时间数组{Te},组合得到动作区域的时间段信息。具体的,依次从{Ts}、{Te}中各选一个值组合成一个动作区域时间段,最终可得到多个动作区域的时间段信息,可以以动作区域时间段数组的形式进行表示。显然,每次选取时,Te需大于Ts,从而组成动作区域时间段数组。
本实施例中,通过对时序信息中的动作片段开始概率和动作片段结束概率进行检测、筛选,得到高概率的边界位置,即动作区域的开始时间和结束时间,以便于后续提取精确的动作区域。
进一步地,基于上述各实施方式,提出本申请动作区域提取方法的第二实施例。参照图5,图5为本申请动作区域提取方法第二实施例的流程示意图;
在本实施例中,在上述步骤S40之前,该动作区域提取方法还包括:
步骤S50,基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
由于上述获取到的动作区域的时间段信息是针对局部范围而言,为进一步提高动作区 域提取结果的精确性和准确性,本实施例中在全局范围内评估候选区域级特征,以获得可靠的动作区域置信度分数进行检索,从而得到动作区域,可进一步提高动作区域提取结果的精确性和准确性。
在本实施例中,在根据检测结果得到动作区域的时间段信息之后,基于第一特征序列和时间段信息对各动作区域的特征进行采样,得到第二特征序列。
具体的,步骤S50包括:
步骤a1,基于所述第一特征序列和所述时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据;
步骤a2,对所述第一预设数量的特征数据进行拼接,得到第二特征序列。
由于每一个时间段信息所对应的动作区域的时间长短并不相同,因此,可先基于第一特征序列和时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据。即,采用线性差值的方法从各动作区域的时间段中采样第一预设数量的特征数据。其中,线性插值是指插值函数为一次多项式的插值方式,其在插值节点上的插值误差为零。当然,在具体实施时,还可以采用其他插值方式,如抛物线插值,但是线性插值相比其他插值方式,具有简单、方便的特点。第一预设数量可选地设为32,当然也可以根据实际情况具体设定,此处不作为对本申请的限定。
在得到第一预设数量的特征数据之后,对第一预设数量的特征数据进行拼接,得到第二特征序列。
步骤S60,将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
在采样得到第二特征序列之后,将第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数。其中,预设动作区域评估模型是基于训练样本和预设候选区域评估模型训练得到的,该预设动作区域评估模型的网络结构为两层全连接层,其中,第1层隐含层有512个单元,激活函数为Relu;第2层激活函数为Sigmoid,输出为该动作区域含有动作片段的置信度,如下:
FC(512,Relu)→FC(1,Sigmoid)
其中,Loss函数是一个简单的回归Loss,用置信度与该候选区域与真实片段的交并比g的方差。定义如下:
Figure PCTCN2020136320-appb-000001
其中,N为动作区域的个数。
此时,步骤S40包括:
步骤S41,根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
在得到动作区域评估分数之后,根据动作区域评估分数和时间段信息,从待修改视频中提取出对应的动作区域。
具体的,步骤S41包括:
步骤b1,按照从大到小的顺序对所述动作区域评估分数进行排序;
步骤b2,根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数;
步骤b3,从所述时间段信息中获取所述目标动作区域评估分数对应的目标时间段信息;
步骤b4,根据所述目标时间段信息从所述待修改视频中提取出对应的动作区域。
动作区域的提取过程如下:先按照从大到小的顺序对动作区域评估分数进行排序,根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数,然后,从时间段信息中获取目标动作区域评估分数对应的目标时间段信息;进而,根据目标时间 段信息从待修改视频中提取出对应的动作区域。
本实施例中,通过提取特征,然后获取对应的时序信息,即各特征对应的时间位置分别属于动作片段开始、中间和结束的概率,进而筛选出高概率的边界位置,即动作区域的开始时间和结束时间,以得到局部的精确的动作区域边界,接着,进一步在全局范围内评估候选区域级特征,以获得可靠的动作区域置信度分数进行检索,从而得到动作区域,可进一步提高动作区域提取结果的精确性和准确性。
本申请还提供一种动作区域提取装置。
参照图6,图6为本申请动作区域提取装置第一实施例的功能模块示意图。
在本实施例中,所述动作区域提取装置包括:
特征提取模块10,用于获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
信息获取模块20,用于将所述第一特征序列输入至预设时序评估模型,得到时序信息;
信息检测模块30,用于检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
区域提取模块40,用于根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
其中,上述动作区域提取装置的各虚拟功能模块存储于图1所示动作区域提取设备的存储器1005中,用于实现动作区域提取程序的所有功能;各模块被处理器1001执行时,可实现提高动作区域提取结果精确性的功能。
进一步地,所述特征提取模块10包括:
分帧处理单元,用于获取待修剪视频,对所述待修剪视频进行分帧处理,得到视频图像序列;
特征提取单元,用于在所述视频图像序列中每隔预设帧数获取目标视频图像,提取所述目标视频图像的红绿蓝RGB特征和光流特征,得到RGB特征序列和光流特征序列;
第一拼接单元,用于对所述RGB特征序列和所述光流特征序列进行拼接,得到第一特征序列。
进一步地,所述时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结束概率,所述信息检测模块30包括:
第一检测单元,用于检测所述动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组;
第二检测单元,用于检测所述动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组;
信息组合单元,用于根据所述动作片段开始时间数组、所述动作片段结束时间数组,组合得到动作区域的时间段信息。
进一步地,所述动作区域提取装置还包括:
特征采样模块,用于基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
分数评估模块,用于将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
所述区域提取模块40,具体用于根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
进一步地,所述特征采样模块包括:
特征采样单元,用于基于所述第一特征序列和所述时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据;
第二拼接单元,用于对所述第一预设数量的特征数据进行拼接,得到第二特征序列。
进一步地,所述区域提取模块40包括:
分数排序单元,用于按照从大到小的顺序对所述动作区域评估分数进行排序;
第一获取单元,用于根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数;
第二获取单元,用于从所述时间段信息中获取所述目标动作区域评估分数对应的目标时间段信息;
区域提取单元,用于根据所述目标时间段信息从所述待修改视频中提取出对应的动作区域。
其中,上述动作区域提取装置中各个模块的功能实现与上述动作区域提取方法实施例中各步骤相对应,其功能和实现过程在此处不再一一赘述。
本申请还提供一种计算机可读存储介质,所述计算机可读存储介质可以是易失性,也可以是非易失性,该计算机可读存储介质上存储有动作区域提取程序,所述动作区域提取程序被处理器执行时实现如以上任一项实施例所述的动作区域提取方法的步骤。
本申请计算机可读存储介质的具体实施例与上述动作区域提取方法各实施例基本相同,在此不作赘述。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者系统不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者系统所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者系统中还存在另外的相同要素。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在如上所述的一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台设备(可以是手机,计算机,服务器,空调器,或者网络设备等)执行本申请各个实施例所述的方法。
以上仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。

Claims (20)

  1. 一种动作区域提取方法,其中,所述动作区域提取方法包括以下步骤:
    获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
    将所述第一特征序列输入至预设时序评估模型,得到时序信息;
    检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
    根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
  2. 如权利要求1所述的动作区域提取方法,其中,所述获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列的步骤包括:
    获取待修剪视频,对所述待修剪视频进行分帧处理,得到视频图像序列;
    在所述视频图像序列中每隔预设帧数获取目标视频图像,提取所述目标视频图像的红绿蓝RGB特征和光流特征,得到RGB特征序列和光流特征序列;
    对所述RGB特征序列和所述光流特征序列进行拼接,得到第一特征序列。
  3. 如权利要求2所述的动作区域提取方法,其中,所述时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结束概率,所述检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息的步骤包括:
    检测所述动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组;
    检测所述动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组;
    根据所述动作片段开始时间数组、所述动作片段结束时间数组,组合得到动作区域的时间段信息。
  4. 如权利要求1至3中任一项所述的动作区域提取方法,其中,所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤之前,还包括:
    基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
    将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
    所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤包括:
    根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
  5. 如权利要求4所述的动作区域提取方法,其中,所述基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列的步骤包括:
    基于所述第一特征序列和所述时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据;
    对所述第一预设数量的特征数据进行拼接,得到第二特征序列。
  6. 如权利要求4所述的动作区域提取方法,其中,所述根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域的步骤包括:
    按照从大到小的顺序对所述动作区域评估分数进行排序;
    根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数;
    从所述时间段信息中获取所述目标动作区域评估分数对应的目标时间段信息;
    根据所述目标时间段信息从所述待修改视频中提取出对应的动作区域。
  7. 一种动作区域提取装置,其中,所述动作区域提取装置包括:
    特征提取模块,用于获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
    信息获取模块,用于将所述第一特征序列输入至预设时序评估模型,得到时序信息;
    信息检测模块,用于检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
    区域提取模块,用于根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
  8. 如权利要求7所述的动作区域提取装置,其中,所述动作区域提取装置还包括:
    特征采样模块,用于基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
    分数评估模块,用于将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
    所述区域提取模块,具体用于根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
  9. 一种动作区域提取设备,其中,所述动作区域提取设备包括存储器、处理器以及存储在所述存储器上并可被所述处理器执行的动作区域提取程序,其中所述动作区域提取程序被所述处理器执行时,实现如下步骤:
    获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
    将所述第一特征序列输入至预设时序评估模型,得到时序信息;
    检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
    根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
  10. 如权利要求9所述的动作区域提取设备,其中,所述获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列的步骤包括:
    获取待修剪视频,对所述待修剪视频进行分帧处理,得到视频图像序列;
    在所述视频图像序列中每隔预设帧数获取目标视频图像,提取所述目标视频图像的红绿蓝RGB特征和光流特征,得到RGB特征序列和光流特征序列;
    对所述RGB特征序列和所述光流特征序列进行拼接,得到第一特征序列。
  11. 如权利要求10所述的动作区域提取设备,其中,所述时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结束概率,所述检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息的步骤包括:
    检测所述动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组;
    检测所述动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组;
    根据所述动作片段开始时间数组、所述动作片段结束时间数组,组合得到动作区域的时间段信息。
  12. 如权利要求9至11中任一项所述的动作区域提取设备,其中,所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤之前,所述动作区域提取程序被所述处理器执行时还实现如下步骤:
    基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
    将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
    所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤包括:
    根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
  13. 如权利要求12所述的动作区域提取设备,其中,所述基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列的步骤包括:
    基于所述第一特征序列和所述时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据;
    对所述第一预设数量的特征数据进行拼接,得到第二特征序列。
  14. 如权利要求12所述的动作区域提取设备,其中,所述根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域的步骤包括:
    按照从大到小的顺序对所述动作区域评估分数进行排序;
    根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数;
    从所述时间段信息中获取所述目标动作区域评估分数对应的目标时间段信息;
    根据所述目标时间段信息从所述待修改视频中提取出对应的动作区域。
  15. 一种计算机可读存储介质,其中,所述计算机可读存储介质上存储有动作区域提取程序,其中所述动作区域提取程序被处理器执行时,实现如下步骤:
    获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列;
    将所述第一特征序列输入至预设时序评估模型,得到时序信息;
    检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息;
    根据所述时间段信息从所述待修改视频中提取出对应的动作区域。
  16. 如权利要求15所述的计算机可读存储介质,其中,所述获取待修剪视频,对所述待修剪视频进行特征提取,得到第一特征序列的步骤包括:
    获取待修剪视频,对所述待修剪视频进行分帧处理,得到视频图像序列;
    在所述视频图像序列中每隔预设帧数获取目标视频图像,提取所述目标视频图像的红绿蓝RGB特征和光流特征,得到RGB特征序列和光流特征序列;
    对所述RGB特征序列和所述光流特征序列进行拼接,得到第一特征序列。
  17. 如权利要求16所述的计算机可读存储介质,其中,所述时序信息包括各目标视频图像对应的动作片段开始概率和动作片段结束概率,所述检测所述时序信息是否符合预设条件,并根据检测结果得到动作区域的时间段信息的步骤包括:
    检测所述动作片段开始概率中是否存在大于第一预设阈值的概率值,得到第一检测结果,并根据所述第一检测结果得到动作片段开始时间数组;
    检测所述动作片段结束概率中是否存在大于第二预设阈值的概率值,得到第二检测结果,并根据所述第二检测结果得到动作片段结束时间数组;
    根据所述动作片段开始时间数组、所述动作片段结束时间数组,组合得到动作区域的时间段信息。
  18. 如权利要求15至17中任一项所述的计算机可读存储介质,其中,所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤之前,所述动作区域提取程序被处理器执行时还实现如下步骤:
    基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列;
    将所述第二特征序列输入至预设动作区域评估模型,得到动作区域评估分数;
    所述根据所述时间段信息从所述待修改视频中提取出对应的动作区域的步骤包括:
    根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域。
  19. 如权利要求18所述的计算机可读存储介质,其中,所述基于所述第一特征序列和所述时间段信息对各动作区域的特征进行采样,得到第二特征序列的步骤包括:
    基于所述第一特征序列和所述时间段信息、采用线性差值方法对各动作区域的特征进行采样,得到第一预设数量的特征数据;
    对所述第一预设数量的特征数据进行拼接,得到第二特征序列。
  20. 如权利要求18所述的计算机可读存储介质,其中,所述根据所述动作区域评估分数和所述时间段信息,从所述待修改视频中提取出对应的动作区域的步骤包括:
    按照从大到小的顺序对所述动作区域评估分数进行排序;
    根据排序结果获取前第二预设数量的动作区域评估分数,作为目标动作区域评估分数;
    从所述时间段信息中获取所述目标动作区域评估分数对应的目标时间段信息;
    根据所述目标时间段信息从所述待修改视频中提取出对应的动作区域。
PCT/CN2020/136320 2020-03-16 2020-12-15 动作区域提取方法、装置、设备及计算机可读存储介质 Ceased WO2021184852A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202010185060.0 2020-03-16
CN202010185060.0A CN111368786B (zh) 2020-03-16 2020-03-16 动作区域提取方法、装置、设备及计算机可读存储介质

Publications (1)

Publication Number Publication Date
WO2021184852A1 true WO2021184852A1 (zh) 2021-09-23

Family

ID=71206848

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/136320 Ceased WO2021184852A1 (zh) 2020-03-16 2020-12-15 动作区域提取方法、装置、设备及计算机可读存储介质

Country Status (2)

Country Link
CN (1) CN111368786B (zh)
WO (1) WO2021184852A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114420075A (zh) * 2022-01-24 2022-04-29 腾讯科技(深圳)有限公司 音频处理方法及装置、设备、计算机可读存储介质
CN114445732A (zh) * 2021-12-22 2022-05-06 北京理工大学 一种面向视频的时间动作检测方法
CN114463679A (zh) * 2022-01-27 2022-05-10 中国建设银行股份有限公司 视频的特征构造方法、装置及设备
CN114758277A (zh) * 2022-04-13 2022-07-15 中国工商银行股份有限公司 异常行为分类模型的训练方法、异常行为分类方法
CN115412765A (zh) * 2022-08-31 2022-11-29 北京奇艺世纪科技有限公司 视频精彩片段确定方法、装置、电子设备及存储介质
US20240037940A1 (en) * 2022-07-28 2024-02-01 International Business Machines Corporation Temporal Action Localization with Mutual Task Guidance

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111368786B (zh) * 2020-03-16 2024-11-29 平安科技(深圳)有限公司 动作区域提取方法、装置、设备及计算机可读存储介质
CN111898461B (zh) * 2020-07-08 2022-08-30 贵州大学 一种时序行为片段生成方法
CN112364835B (zh) * 2020-12-09 2023-08-11 武汉轻工大学 视频信息取帧方法、装置、设备及存储介质
CN115909504A (zh) * 2022-12-26 2023-04-04 京东科技信息技术有限公司 视频检测、视频检测模型的训练方法、装置、设备及介质

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108234821A (zh) * 2017-03-07 2018-06-29 北京市商汤科技开发有限公司 检测视频中的动作的方法、装置和系统
US20180349704A1 (en) * 2017-06-02 2018-12-06 Canon Kabushiki Kaisha Interaction classification using the role of people interacting over time
CN109784269A (zh) * 2019-01-11 2019-05-21 中国石油大学(华东) 一种基于时空联合的人体动作检测和定位方法
CN110414367A (zh) * 2019-07-04 2019-11-05 华中科技大学 一种基于gan和ssn的时序行为检测方法
CN110796071A (zh) * 2019-10-28 2020-02-14 广州博衍智能科技有限公司 行为检测方法、系统、机器可读介质及设备
CN111368786A (zh) * 2020-03-16 2020-07-03 平安科技(深圳)有限公司 动作区域提取方法、装置、设备及计算机可读存储介质

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110263733B (zh) * 2019-06-24 2021-07-23 上海商汤智能科技有限公司 图像处理方法、提名评估方法及相关装置
CN110852256B (zh) * 2019-11-08 2023-04-18 腾讯科技(深圳)有限公司 时序动作提名的生成方法、装置、设备及存储介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108234821A (zh) * 2017-03-07 2018-06-29 北京市商汤科技开发有限公司 检测视频中的动作的方法、装置和系统
US20180349704A1 (en) * 2017-06-02 2018-12-06 Canon Kabushiki Kaisha Interaction classification using the role of people interacting over time
CN109784269A (zh) * 2019-01-11 2019-05-21 中国石油大学(华东) 一种基于时空联合的人体动作检测和定位方法
CN110414367A (zh) * 2019-07-04 2019-11-05 华中科技大学 一种基于gan和ssn的时序行为检测方法
CN110796071A (zh) * 2019-10-28 2020-02-14 广州博衍智能科技有限公司 行为检测方法、系统、机器可读介质及设备
CN111368786A (zh) * 2020-03-16 2020-07-03 平安科技(深圳)有限公司 动作区域提取方法、装置、设备及计算机可读存储介质

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114445732A (zh) * 2021-12-22 2022-05-06 北京理工大学 一种面向视频的时间动作检测方法
CN114420075A (zh) * 2022-01-24 2022-04-29 腾讯科技(深圳)有限公司 音频处理方法及装置、设备、计算机可读存储介质
CN114420075B (zh) * 2022-01-24 2024-12-03 腾讯科技(深圳)有限公司 音频处理方法及装置、设备、计算机可读存储介质
CN114463679A (zh) * 2022-01-27 2022-05-10 中国建设银行股份有限公司 视频的特征构造方法、装置及设备
CN114758277A (zh) * 2022-04-13 2022-07-15 中国工商银行股份有限公司 异常行为分类模型的训练方法、异常行为分类方法
US20240037940A1 (en) * 2022-07-28 2024-02-01 International Business Machines Corporation Temporal Action Localization with Mutual Task Guidance
US12555375B2 (en) * 2022-07-28 2026-02-17 International Business Machines Corporation Temporal action localization with mutual task guidance
CN115412765A (zh) * 2022-08-31 2022-11-29 北京奇艺世纪科技有限公司 视频精彩片段确定方法、装置、电子设备及存储介质
CN115412765B (zh) * 2022-08-31 2024-03-26 北京奇艺世纪科技有限公司 视频精彩片段确定方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
CN111368786B (zh) 2024-11-29
CN111368786A (zh) 2020-07-03

Similar Documents

Publication Publication Date Title
WO2021184852A1 (zh) 动作区域提取方法、装置、设备及计算机可读存储介质
CN109284729B (zh) 基于视频获取人脸识别模型训练数据的方法、装置和介质
US11768597B2 (en) Method and system for editing video on basis of context obtained using artificial intelligence
EP3989158B1 (en) Method, apparatus and device for video similarity detection
CN111061915B (zh) 视频人物关系识别方法
CN116166843B (zh) 基于细粒度感知的文本视频跨模态检索方法和装置
CN108446681B (zh) 行人分析方法、装置、终端及存储介质
CN111931859B (zh) 一种多标签图像识别方法和装置
US11756301B2 (en) System and method for automatically detecting and marking logical scenes in media content
WO2021104097A1 (zh) 表情包生成方法、装置及终端设备
CN109960988A (zh) 图像分析方法、装置、电子设备及可读存储介质
CN112052375A (zh) 舆情获取和词粘度模型训练方法及设备、服务器和介质
CN110414433A (zh) 图像处理方法、装置、存储介质和计算机设备
CN110941978A (zh) 一种未识别身份人员的人脸聚类方法、装置及存储介质
CN106372603A (zh) 遮挡人脸识别方法及装置
CN114625918B (zh) 视频推荐方法、装置、设备、存储介质及程序产品
EP2864906A2 (en) Searching for events by attendants
WO2021228148A1 (zh) 保护个人数据隐私的特征提取方法、模型训练方法及硬件
WO2022111688A1 (zh) 人脸活体检测方法、装置及存储介质
CN114519831A (zh) 电梯场景识别方法、装置、电子设备及存储介质
CN111507289A (zh) 视频匹配方法、计算机设备和存储介质
CN115719428A (zh) 基于分类模型的人脸图像聚类方法、装置、设备及介质
US20250245986A1 (en) Scene recognition
CN114697763A (zh) 一种视频处理方法、装置、电子设备及介质
CN119942651A (zh) 一种基于姿态估计与时序预测的不良行为智能审核方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20925316

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20925316

Country of ref document: EP

Kind code of ref document: A1