WO2019201008A1 - 直播视频审查的方法和装置 - Google Patents
直播视频审查的方法和装置 Download PDFInfo
- Publication number
- WO2019201008A1 WO2019201008A1 PCT/CN2019/075141 CN2019075141W WO2019201008A1 WO 2019201008 A1 WO2019201008 A1 WO 2019201008A1 CN 2019075141 W CN2019075141 W CN 2019075141W WO 2019201008 A1 WO2019201008 A1 WO 2019201008A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video
- terminal device
- image
- feature vectors
- server
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/21—Server components or server architectures
- H04N21/218—Source of audio or video content, e.g. local disk arrays
- H04N21/2187—Live feed
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/854—Content authoring
- H04N21/8547—Content authoring involving timestamps for synchronizing content
Definitions
- the present application relates to the field of computer technologies, and in particular, to a method and apparatus for live video review.
- Video content is increasingly used, and the review of video content is an important part of the processing of video content.
- a common method is direct interaction between servers, that is, the first server acquires the video being played from the second server, and performs time-framed recognition on the video to identify yellow, violent, evil, and disgusting. Restricted video such as political sensitivity; another commonly used method is that the server receives an image of a video being played periodically sent by the terminal device, and recognizes the received image to identify yellow, violent, evil, disgusting, Restricted video such as political sensitivity.
- the above methods all implement video review on the server side, and the massive live video stream for the Internet needs to be implemented by a high-performance server, which is costly and has low review efficiency; and, for the server receiving terminal device, the periodic transmission is being played.
- the video review method of the video image also has the problem of large video review time delay and large traffic consumption.
- the embodiment of the present application provides a method and device for live video review, which does not require a high-performance server, has high review efficiency, small traffic consumption, and small video review time delay.
- the embodiment of the present application provides a method for live video review, including:
- the video review information including respective identification and execution times of N feature vectors, N ⁇ 1; the N feature vectors are partial feature vectors among all feature vectors stored by the server;
- the server determines that the video is not After passing the test, the video is marked as unqualified.
- the method further includes: when the terminal device acquires the matching degree between the M images corresponding to the N feature vectors and the first video frame image, acquiring the second video frame image;
- the server performs the unqualified labeling of the video after determining that the video is unqualified, M ⁇ N.
- the obtaining the N feature vectors according to the respective identifiers of the N feature vectors includes:
- the first video frame image is a next frame image of a video frame image being displayed at the execution time, and the first video frame image is a video frame of a storage terminal device to be displayed.
- the image in the cache is a next frame image of a video frame image being displayed at the execution time, and the first video frame image is a video frame of a storage terminal device to be displayed. The image in the cache.
- the playing time of the second video frame image is later than the first video frame image, and the second video frame image is an image in a buffer of a video frame of a storage terminal device to be displayed.
- the embodiment of the present application provides a method for live video review, including:
- the video review information includes respective identifiers and execution times of the N feature vectors; the execution time is used to indicate that the terminal device first acquires the M images corresponding to the N feature vectors for matching.
- the time of the frame image, the N feature vectors are partial feature vectors among all feature vectors stored by the server;
- the target feature vector is a feature vector corresponding to the target video frame image in the M images determined by the terminal device a feature vector corresponding to the first image whose matching degree is greater than the first preset threshold;
- the video is unqualified, M ⁇ N.
- the method further includes:
- the target video frame image is structured to obtain the target video frame image. Structured data; the structured data is used to acquire a model for acquiring feature vectors.
- the method before the sending the video review information to the terminal device, the method further includes:
- the video is marked as unqualified, including:
- the identification of the video is stored in association with the failure type of the second image.
- an embodiment of the present application provides a device for performing live video review, including:
- a receiving module configured to receive video review information from a server, where the video review information includes respective identifiers and execution times of N feature vectors, N ⁇ 1; the N feature vectors are part of all feature vectors stored by the server Feature vector;
- An obtaining module configured to acquire N feature vectors according to respective identifiers of the N feature vectors
- the acquiring module is further configured to: acquire, at the execution time, a first video frame image of a video being broadcasted, at least part of the M images corresponding to the N feature vectors, and the first video frame image Matching degree;
- a sending module configured to: if the first image corresponding to the first video frame image is greater than or equal to the first preset threshold in the M images corresponding to the N feature vectors, the first image is corresponding to An identifier of the first feature vector, the first video frame image, the identifier of the video is sent to the server, so that the server performs the unqualified labeling of the video after determining that the video is unqualified, M ⁇ N.
- the method further includes: the acquiring module, configured to: when the terminal device acquires the matching degree between the M images corresponding to the N feature vectors and the first video frame image Obtaining a second video frame image;
- the server performs a failure labeling on the video after determining that the video is unqualified.
- the sending module is further configured to
- the server Before receiving video review information from the server: transmitting information of the video to the server, the information of the video including an identifier of the video and an identifier of the terminal device; an identifier of the video and an identifier of the terminal device are used for the
- the server determines a first number of terminal devices that play the video, the first number and a second number of all feature vectors stored by the server for determining a value of the N and an execution time.
- the obtaining module is specifically configured to:
- the embodiment of the present application provides a device for performing live video review, including:
- a sending module configured to send video review information to the terminal device, where the video review information includes an identifier and an execution time of each of the N feature vectors; the execution time is used to instruct the terminal device to acquire the M frames corresponding to the N feature vectors for the first time.
- the time at which the image is matched to the video frame image, and the N feature vectors are partial feature vectors in all feature vectors stored by the server;
- a receiving module configured to receive an identifier of a target feature vector sent by the terminal device, a target video frame image, and an identifier of the video;
- the target feature vector is an image of the target video frame in the M image determined by the terminal device a feature vector corresponding to the first image whose matching degree is greater than the first preset threshold;
- the labeling module is configured to perform the unqualified labeling of the video if the second image corresponding to the target video frame image is greater than the second preset threshold in the first image corresponding to the target feature vector, M ⁇ N .
- the device further includes:
- a processing module configured to perform structural processing on the target video frame image if a second image that has a matching degree with the target video frame image that is greater than a second preset threshold is present in the first image corresponding to the target feature vector, Obtaining structured data of the target video frame image; the structured data is used to acquire a model, and the model is used to acquire a feature vector.
- the receiving module is further configured to receive information about the video that is sent by the terminal device, where the information of the video includes an identifier of the video and an identifier of the terminal device;
- the device further includes: a determining module
- the determining module is configured to determine, according to the information of the video that includes the identifier of the video that is received at the current time, the first quantity of the terminal device that plays the video;
- the labeling module is specifically configured to store the identifier of the video in association with the type of the second image.
- an embodiment of the present application provides a terminal device, including a processor
- the processor is operative to couple with a memory to read and execute instructions in the memory to implement the first aspect and any of the methods described in the design.
- the memory is also included.
- an embodiment of the present application provides a server, including a processor
- the processor is operative to couple with a memory to read and execute instructions in the memory to implement the second aspect and any of the methods described in the design.
- the memory is also included.
- embodiments of the present application provide a computer storage medium comprising instructions that, when executed on a communication device, cause the communication device to perform the first aspect and any of the methods described in the design.
- an embodiment of the present application provides a computer storage medium, comprising instructions that, when executed on a communication device, cause the communication device to perform the second aspect and any of the methods described in the design.
- FIG. 1 is a system architecture diagram provided by an embodiment of the present application.
- FIG. 3 is a schematic structural diagram 1 of an apparatus for performing live video review according to an embodiment of the present application
- FIG. 4 is a schematic structural diagram 2 of an apparatus for performing live video review according to an embodiment of the present application
- FIG. 5 is a schematic structural diagram 3 of a device for performing live video review according to an embodiment of the present disclosure
- FIG. 6 is a schematic structural diagram 1 of a terminal device provided by the present application.
- FIG. 7 is a schematic structural diagram 2 of a terminal device provided by the present application.
- FIG. 8 is a schematic structural diagram 1 of a server provided by the present application.
- FIG. 9 is a schematic structural diagram 2 of a server provided by the present application.
- the machine learning algorithm can be a convolutional neural network algorithm
- the feature vector is the basic data of the image via a convolutional neural network algorithm
- the convolutional neural network outputs the vector of the fully connected layer
- the basic data of the image can be the gray of the pixel of the image.
- LBP Local Binary Pattern
- the convolutional neural network algorithm is an algorithm commonly used in the prior art, and is not described in this embodiment.
- Terminal equipment may also be called a user equipment (UE), an access terminal, a subscriber unit, a subscriber station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, and a wireless communication.
- UE user equipment
- An access terminal a subscriber unit, a subscriber station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, and a wireless communication.
- UE user equipment
- an access terminal a subscriber unit, a subscriber station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, and a wireless communication.
- Device user agent, or user device.
- the terminal device may be a station (station, ST) in a wireless local area network (WLAN), and may be a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, or a wireless local loop (wireless local Loop, WLL) station, personal digital assistant (PDA) device, handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, in-vehicle device, wearable device, and next-generation communication system, For example, a terminal device in a fifth-generation (5G) network or a terminal device in a future public land mobile network (PLMN) network, a new radio (NR) communication system Terminal equipment, etc.
- 5G fifth-generation
- PLMN public land mobile network
- NR new radio
- the terminal device may also be a wearable device.
- a wearable device which can also be called a wearable smart device, is a general term for applying wearable technology to intelligently design and wear wearable devices such as glasses, gloves, watches, clothing, and shoes.
- a wearable device is a portable device that is worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are more than just a hardware device, but they also implement powerful functions through software support, data interaction, and cloud interaction.
- Generalized wearable smart devices include full-featured, large-size, non-reliable smartphones for full or partial functions, such as smart watches or smart glasses, and focus on only one type of application, and need to work with other devices such as smartphones. Use, such as various smart bracelets for smart signs monitoring, smart jewelry, etc.
- the terminal device may also include a drone, such as an onboard communication device on the drone.
- FIG. 1 is a system architecture diagram of an embodiment of the present application.
- the system architecture includes a server 11 and a terminal device 12.
- the server stores a plurality of feature vectors, and the plurality of feature vectors are obtained by the server according to the model and the machine learning algorithm after the model is trained by the machine learning algorithm according to the basic data of the plurality of images that are prohibited from being played.
- the machine learning algorithm can be a deep convolutional neural network algorithm
- the model is a deep convolutional neural network model.
- An image that is prohibited from playing may correspond to at least one feature vector.
- the terminal device 12 receives the video review information from the server 11, and acquires the first video frame image of the video being broadcasted at the execution time indicated by the video review information, if the M image corresponding to the N feature vectors indicated by the video review information exists.
- the first image of the first feature vector corresponding to each first image, the first video frame image, and the video, the first image of the first video image that matches the first video frame image of the live video is greater than or equal to the first preset threshold.
- the identity is sent to the server 11.
- the N feature vectors are partial feature vectors in all feature vectors stored in the server.
- the server 11 acquires the matching degree of the first image corresponding to each first feature vector of the first video frame image, and when there is a first image whose matching degree with the first video frame image is greater than the second preset threshold, the sending stops.
- the control command of the video is played to the terminal device 12, and the video is marked as unqualified.
- each terminal device completes matching of the M image corresponding to the N feature vectors indicated by the review information by the video frame image to be identified, that is, the video frame image of the terminal device to be identified and completed
- the matching between all the images corresponding to all the feature vectors stored by the server the server only needs to match the video frame image and the feature vector after the terminal device reports the matching between the reported video frame image and the image indicated by the feature vector.
- the terminal device reports the video frame image and the feature vector identifier to the server only when the matching degree of the image corresponding to the feature vector indicated by the video frame image and the video review information is greater than or equal to the first preset threshold. The amount of calculation is very small, the ordinary server can be completed, and the cost of the server is low. Since the terminal device division completes the matching of the video frame image to be identified with the feature vector, the efficiency of video review is greatly improved.
- the terminal device reports the video frame image only when the matching degree of the image corresponding to the feature vector indicated by the video frame image and the video review information is greater than or equal to the first preset threshold.
- the identification of the feature vector to the server eliminates the need for the image of the currently playing video sent by the terminal device in the prior art to the server, and the consumed traffic is small, and the video review delay is small.
- the method for the live video review provided by the embodiment of the present application is described in detail below with reference to a specific embodiment.
- the execution subject in the following embodiments may be the server 11 in FIG.
- FIG. 2 is an interaction flowchart of a video review method according to an embodiment of the present disclosure. Referring to FIG. 2, the method in this embodiment of the present application includes:
- Step S201 The terminal device sends information of the video being broadcasted to the server, where the information of the video includes the identifier of the video and the identifier of the terminal device.
- Step S202 The server determines, according to the information of the video that includes the identifier of the video that is received at the current time, the first quantity of the terminal device that plays the video.
- Step S203 The server determines, according to the first quantity of the terminal device that plays the video and the second quantity of all the feature vectors stored in the server, the video review information sent to each terminal device that plays the video, where the video review information includes N features.
- the respective identification and execution time of the vector, N ⁇ 1, and the N feature vectors are partial feature vectors among all feature vectors stored by the server;
- Step S204 The server sends corresponding video review information to the terminal device.
- Step S205 The terminal device acquires N feature vectors according to respective identifiers of the N feature vectors in the video review information.
- Step S206 The terminal device acquires a first video frame image of the currently playing video at the execution time indicated by the video review information, and acquires a matching degree of the M image corresponding to the N feature vectors of the first video frame image;
- the video frame image is the next frame image of the video frame image being displayed at the above execution time, M ⁇ N;
- Step S207 If there is a first image in which the matching degree of the first video frame image is greater than or equal to the first preset threshold in the M images corresponding to the N feature vectors, the terminal device sets the first feature vector corresponding to the first image.
- the identifier, the first video frame image, and the identifier of the video are sent to the server;
- Step S208 The server acquires a matching degree of the first image corresponding to the first feature vector of the first video frame image.
- steps S209-S210 are performed:
- Step S209 The server performs the unqualified labeling on the video
- Step S210 the server sends a control instruction to stop playing the video to the terminal device
- step S211 to step S213 are performed.
- Step S211 when the terminal device acquires the matching degree between the M images corresponding to the N feature vectors and the first video frame image, the terminal device acquires the second video frame image, and acquires the second video frame image and the N features.
- Step S212 If there is a second image in which the matching degree of the second video frame image is greater than or equal to the first preset threshold in the M images corresponding to the N feature vectors, the terminal device sets the second feature vector corresponding to the second image.
- the identifier, the second video frame image, and the identifier of the video are sent to the server;
- Step S213 The server acquires a matching degree of the second image corresponding to the second video frame image and the second feature vector;
- step S214 to step S215 are performed.
- Step S214 The server performs the unqualified labeling on the video.
- Step S215 The server sends a control instruction to stop playing the video to the terminal device.
- the terminal device when the terminal device receives the target video of the live broadcast from the server capable of providing the live video, the terminal device sends the information of the target video of the live broadcast to the server; wherein the information of the target video may include: the target video The identifier, the identifier of the terminal device, and the information of the target video may further include: an IP address of the terminal device, a name of the target video, and an address of the target video.
- the server receives the information of the target video being broadcasted by the terminal device, it can be understood that the server receives the information of multiple videos of the plurality of terminal devices at the same time, and each video information has a corresponding video. Identification and identification of the terminal device. For the target video that the terminal device is currently broadcasting, the server counts the number of the identifiers of the target video in the information of the video received at the current moment, and the number is the first number of the terminal devices that are playing the target video.
- the number of the identifiers of the terminals included in the information of all the videos having the identifier of the target video is the first number of the terminal devices that are playing the target video, that is, the information included in the information of all the videos having the identifier of the target video.
- the terminal device indicated by the terminal identifier is the terminal device that plays the target video.
- the server stores a plurality of feature vectors, and the plurality of feature vectors are obtained by the server according to the model and the machine learning algorithm after the model is trained by the machine learning algorithm according to the basic data of the plurality of images that are prohibited from being played.
- the plurality of feature vectors are obtained by the server according to the model and the machine learning algorithm after the model is trained by the machine learning algorithm according to the basic data of the plurality of images that are prohibited from being played.
- the types of videos that are prohibited from playing can be: yellow, violent, suspected, disgusting, politically sensitive, and so on. Since the video is composed of images of multiple frames, the feature vectors stored in the server are based on images that are prohibited from being played.
- the server After acquiring the first number of terminal devices playing the target video, the terminal device playing the target video, and the second number of all feature vectors stored in the server, the server is configured according to the first quantity of the terminal device that plays the target video.
- the second quantity of all feature vectors stored in the server determines video review information sent to the terminal device, the video review information includes respective identification and execution time of the N feature vectors, N ⁇ 1, and the N feature vectors are all stored by the server. Part of the feature vector in the feature vector.
- the server determines, according to the first quantity and the second quantity, video review information sent to each terminal device that is playing the target video, as follows:
- the server divides the K terminal devices that play the target video into L groups, each Each terminal device in the group is assigned the same feature vector, and the number of feature vectors assigned to each terminal device in each group is 1, and the feature vectors assigned to each group are different.
- the server also configures a different execution time for each terminal device in each group, and the execution time is used to indicate the time when the terminal device first acquires the video frame image that matches the image corresponding to the feature vector.
- the first quantity is K
- the second quantity is L
- K ⁇ L indicating that the number of terminal devices is small and the number of feature vectors is large.
- the server also sets an execution time for each terminal device, and the execution time is used to indicate the time when the terminal device first acquires the video frame image that matches the image corresponding to the feature vector. At this time, the execution time of each terminal device can be the same.
- the video review information determined by the server and sent to each terminal device that plays the target video includes the identifier and execution time of the N feature vectors, N ⁇ 1, and each terminal device corresponds to The N may or may not be the same.
- step S204 the server transmits respective video review information configured for each terminal device that plays the target video to each terminal device.
- the terminal device After receiving the video review information sent by the server, the terminal device acquires N feature vectors according to the respective identifiers of the N feature vectors in the video review information.
- N feature vectors according to respective identifiers of the N feature vectors in the video review information, specifically: determining, according to respective identifiers of the N feature vectors, whether the terminal device has feature vectors in the N feature vectors;
- the identifier of the fourth feature vector is sent to the server, and the fourth feature vector sent by the server is received.
- the terminal device Since the terminal device performs multiple video reviews, the terminal device itself stores some feature vectors acquired during the previous video review. For each of the N feature vectors not stored in the terminal device, it can be downloaded from the server.
- the terminal device For the execution time indicated by the video review information, the terminal device acquires the first video frame image of the currently playing video, and acquires at least part of the M image corresponding to the N feature vectors of the first video frame image. suitability.
- the first video frame image is the next frame image of the video frame image being displayed at the execution time, that is, the first video frame image is the most in the video frame buffer of the storage terminal device to be displayed.
- the video display frame buffer is a video frame buffer to be displayed of the storage terminal device, where the stored video frames of several frames are to be played.
- the matching degree of at least part of the M images corresponding to the N feature vectors of the first video frame image may be acquired.
- the matching degree may also be referred to as similarity, at least partially in whole or in part.
- the N feature vectors indicated by the identifiers of the N feature vectors included in the video review information are N feature vectors corresponding to the M forbidden playback images, and one image forbidden to play corresponds to at least one feature vector.
- One commonly used method is: if the above N feature vectors are obtained by using a convolutional neural network algorithm, the same convolutional neural network algorithm is used to obtain the first video frame. a feature vector corresponding to the image, calculating a distance between the feature vector corresponding to the first video frame image and the feature vector in the N feature vectors, and then normalizing the distance value into a matching degree by using a normalization method, wherein, normalizing
- the methods are linear mapping, piecewise linear mapping, and other methods of monotonic functions.
- N feature vectors are obtained by using a convolutional neural network algorithm, the same convolutional neural network is used to obtain the feature vector corresponding to the first video frame image, directly according to the corresponding image of the first video frame image.
- the feature vector and the feature vector of the N feature vectors acquire the matching degree of the image corresponding to the feature vector of the first video frame image and the N feature vectors.
- the matching degree of the image corresponding to the feature vector of the N feature vectors is obtained by using the feature vector of the N feature vectors and the feature vector corresponding to the first video frame image.
- the degree of matching between images is obtained by using the feature vector of the N feature vectors and the feature vector corresponding to the first video frame image.
- the terminal device sends the identifier of the first feature vector corresponding to the first image, the first video frame image, the time stamp of the first video frame image, and the identifier of the video to the server.
- the feature vector x is the first feature vector, and the identifier of the feature vector x is It is sent to the server by the terminal device.
- the first preset threshold may be between 80 and 85%.
- the matching degree of the first image corresponding to the first feature vector and the first video frame image is greater than or equal to the first preset threshold, indicating that the first video frame image is similar to the first image corresponding to the first feature vector, that is, the target
- the probability that the video is a non-conforming video is very large. Therefore, the terminal device sends the identifier of the first feature vector, the first video frame image, the time stamp of the first video frame image, and the identifier of the video to the server.
- the number of the first feature vectors may be one or more.
- the terminal device may sequentially acquire the matching degree of the image corresponding to the first video frame image and the multiple feature vectors, and the terminal device only needs to obtain the first video image.
- the matching degree of the image corresponding to the feature vector is greater than or equal to the first preset threshold, the identifier of the corresponding feature vector, the first video frame image, the time stamp of the first video frame image, and the identifier of the target video are sent to the server.
- the identifier of the first feature vector of the first image with the matching degree of the first video image being greater than the first preset threshold
- the identity of the target video is sent to the server.
- the identifiers of the feature vectors carried by the video review information received by the terminal device 3 are two, one indicating the feature vector 1, and the other indicating the feature vector 2, and the unqualified type of the image corresponding to the feature vector 1 is a politically sensitive type and feature.
- the unqualified type of the image corresponding to the vector 2 is a horror type, and the first matching degree of the first video frame image acquired by the terminal device and the image 1 corresponding to the feature vector 1 is obtained, and the first matching degree is 87%.
- the preset threshold is 80%
- the feature vector 1 is the first feature vector
- the terminal device immediately identifies the feature vector 1, the first video frame image, the time stamp of the first video frame image, and the target video identifier. Will be sent to the server.
- the terminal device may then acquire the second matching degree of the first video frame image and the feature vector 2, that is, the terminal device will feature
- the identifier of the vector 1, the first video frame image, the time stamp of the first video frame image, and the identifier of the target video are sent to the server, and the terminal device acquires the second matching degree of the first video frame image and the feature vector 2 If the obtained second matching degree is 70%, the terminal device does not send any information to the server, and if the obtained second matching degree is 85%, indicating that the feature vector 2 is the first feature vector, the terminal device will use the feature vector.
- the identifier of 2 the first video frame image, the timestamp of the first video frame image, and the identifier of the target video are sent to the server.
- the server receives the identifier of the first feature vector sent by the terminal device, the first video frame image, the timestamp of the first video frame image, and the identifier of the target video, and reacquires the first video frame image and the first a matching degree of the first image corresponding to the feature vector. If there is a second image that matches the first video frame image by a second preset threshold, the first feature vector corresponding to the second image is called the first target feature vector. , the control command to stop playing the target video is sent to the terminal device and the video is marked as unqualified.
- the first preset threshold may be between 90 and 95%.
- the matching degree between the second image corresponding to the first target feature vector acquired by the server and the first video frame image is greater than or equal to a second preset threshold, indicating that the second image corresponding to the first target frame vector is very Similarly, if the target video is basically determined to be an unqualified video, the target video is marked as unqualified.
- the server determines that the target video is a non-conforming video
- the target video is an unqualified video after the first time that the matching between the video frame image of the target video and the first image corresponding to the first feature vector is greater than the second preset threshold.
- the control command to stop playing the target video is immediately sent to each terminal device.
- the unqualified labeling of the target video includes: storing the identifier of the target video in association with the unqualified type of the second image, that is, storing the unqualified type of the image corresponding to the first target feature vector.
- the terminal device After the terminal device receives the control instruction to stop playing the target video, if the terminal device has n feature vectors in the N feature vectors indicated by the video review information, the n feature vectors are not used to calculate the matching degree (corresponding to the step S206).
- the partial image or all the images in the M images corresponding to the N feature vectors may stop acquiring the matching degree of the image corresponding to the first video frame image of the target video and the n feature vectors, and may continue to acquire n features.
- the degree of matching between the respective images of the vectors and the first video frame image (corresponding to all of the M images corresponding to the N feature vectors in step S206).
- the matching degree of the image corresponding to each of the n feature vectors is continuously obtained, and the possibility of completely labeling the unqualified type of the target video can be increased. This is because the type of the image corresponding to the feature vector may not be the same. If all the feature vectors are used for the calculation of the matching degree, the more unqualified type of the obtained target video is.
- the identifiers of the feature vectors carried by the video review information received by the terminal device 3 are two, one indicating the feature vector 1, and the other indicating the feature vector 2, and the unqualified type of the image corresponding to the feature vector 1 is a politically sensitive type and feature.
- the unqualified type of the image corresponding to the vector 2 is a horror type, and the first matching degree of the first video frame image acquired by the terminal device and the image corresponding to the feature vector 1 is obtained, and the first matching degree is 87%, when the first When the preset threshold is 80%, the feature vector 1 is the first feature vector, and the terminal device immediately sends the identifier of the feature vector 1, the first video frame image, the time stamp of the first video frame image, and the identifier of the target video to server.
- the matching degree of the first video frame image re-acquired by the server and the image corresponding to the feature vector 1 is 93%, and the second preset threshold is 90%, indicating that the feature vector 1 is the first target feature vector, if the first video frame image
- the matching degree of the image corresponding to the feature vector 1 is 93%.
- the matching degree of the video frame image obtained by the server for the first time and the image corresponding to the feature vector is greater than the second preset threshold, then the target video is determined by the server.
- the server immediately controls each terminal device that is currently broadcasting the target to continue to play the live video, and marks the target video as: a politically sensitive type, that is, the identifier of the target video is associated with "politically sensitive".
- the server will continue to receive (if any) the identifier of the feature vector transmitted by the terminal device (including the terminal device 3) of each live target video, and the first video frame.
- the image, the timestamp of the first video frame image, and the identification of the target video will be received (if any) the identifier of the feature vector transmitted by the terminal device (including the terminal device 3) of each live target video, and the first video frame. The image, the timestamp of the first video frame image, and the identification of the target video.
- the terminal device 3 obtains the second matching degree of the image corresponding to the first video frame image and the feature vector 2, if The obtained second matching degree is 85%, indicating that the feature vector 2 is also the first feature vector, and the terminal device will identify the feature vector 2, the first video frame image, the time stamp of the first video frame image, and the target video.
- the identifier is sent to the server; the first video frame image reacquired by the server and the image corresponding to the feature vector 2 have a matching degree of 90%, and the second preset threshold is 90%, indicating that the feature vector 2 is also the first target feature vector.
- the identity of the target video is also stored in association with "terror".
- the identifier of the feature vector carried by the video review information received by the terminal device 3 is one, indicating the feature vector 1, and the unqualified type of the image corresponding to the feature vector 1 is a politically sensitive type, which is acquired by the terminal device.
- the matching degree of the image corresponding to the first video frame image and the feature vector 1 is 87%, then the feature vector 1 is the first feature vector, the identifier of the feature vector 1, the first video frame image, and the time stamp of the first video frame image.
- the identifier of the target video is sent to the server; the identifier of the feature vector carried by the video review information received by the terminal device 4 is one, indicating the feature vector 2, and the image unqualified type corresponding to the feature vector 2 is a politically sensitive type, and the terminal
- the first video frame image obtained by the device and the image corresponding to the feature vector 2 have a matching degree of 85%, and the feature vector 2 is the first feature vector, the identifier of the feature vector 2, the first video frame image, and the first video frame image.
- Timestamp the identifier of the target video is sent to the server; the first video frame image retrieved by the server and the map corresponding to the feature vector 1
- the matching degree is 93%
- the matching degree of the image corresponding to the first video frame image and the feature vector 2 is 90%
- the second preset threshold is 90%
- the feature vector 1 and the feature vector 2 are the first target feature vector.
- the target video is stored in association with "politically sensitive” and "terror”, that is, the target video is labeled as: “politically sensitive” and "terror” type target types.
- the server performs a structured process on the first video frame to obtain the image of the first video frame.
- the first structured data is used to acquire a model
- the model is used to acquire a feature vector. That is, the first structured data is used as a training sample to acquire a model for acquiring a feature vector, and the model is also used to acquire a feature vector corresponding to the first video frame image to enrich the feature vector stored in the server, thereby improving the accuracy of the video review.
- the time stamp of the first video frame image and the first video frame image may be entered into the unqualified filing library as evidence for reporting.
- the method for the structuring of the first video frame by the server is a method in the prior art, which is not described in this embodiment.
- the terminal device acquires at the terminal device. After completing the matching degree of the M video images corresponding to the N feature vectors, the terminal device acquires the second video frame image, and acquires at least part of the M images corresponding to the N feature vectors of the second video frame image. The matching degree of the image.
- the “before the terminal device acquires the matching degree of the M video images corresponding to the N feature vectors of the first video frame image” includes: acquiring, after the terminal device, the M frames corresponding to the N feature vectors of the first video frame image. When the image is matched.
- the second video frame image is an image in the video display frame buffer that is different from the first video frame image, and the second video frame image is played later than the first video frame image. It can be understood that, when acquiring the second video frame image, the first video frame image may have been played, or may be in the video display frame buffer. If the first video frame image is still in the video display frame buffer, then The second video frame image is the next video frame image adjacent to the first video frame image.
- the time relationship between the first video frame image and the second video frame image of each terminal device is not necessarily the same.
- the method for acquiring the matching degree of the second video frame image and the at least part of the M images corresponding to the N feature vectors is the same as acquiring the at least part of the M images corresponding to the first video frame image and the N feature vectors.
- the method of matching degree is not described in this embodiment.
- Steps S207 to S210 are referred to in steps S212 to S215, and details are not described herein again.
- the terminal device does not receive the control instruction of the server to stop the playback of the target video before the terminal device acquires the matching degree of the M image corresponding to the N feature vectors, the terminal device When the terminal device acquires the matching degree of the M video images corresponding to the N feature vectors, the terminal device acquires the third video frame image, and acquires the M image corresponding to the N feature vectors of the third video frame image.
- the third video frame image is an image that is different from the second video frame image in the video display frame buffer, and The playback time of the three video frame images is later than the second video frame image. It can be understood that when acquiring the third video frame image, the second video frame image may have been played, or may be in the video display frame buffer, if the second video frame image is still in the video display frame buffer, then The third video frame image is the next video frame image adjacent to the second video frame image.
- the terminal device obtains the matching degree of the M image corresponding to the N feature vectors after the terminal device acquires the matching degree of the M image corresponding to the N feature vectors.
- the corresponding process of S215; the m+1th video frame image is an image in the video display frame buffer that is different from the mth video frame image, and the play time of the m+1th video frame image is later than the mth video frame image.
- the mth video frame image when acquiring the m+1th video frame image, the mth video frame image may have been played, or may be in the video display frame buffer, if the mth video frame image is still in the video display frame buffer.
- the m+1th video frame image is the next video frame image adjacent to the mth video frame image.
- the terminal device after receiving the control instruction to stop playing the target video, the terminal device has n feature vectors in the N feature vectors indicated by the current video review information, and is not used to calculate the matching degree. Then, the matching degree of the image corresponding to the image frame image of the target video and the n feature vectors may be stopped, and the matching degree between the image corresponding to each of the n feature vectors and the corresponding video frame image may be continuously obtained. Wherein, the matching degree of the image corresponding to each of the n feature vectors and the corresponding video frame image is continuously obtained, which may increase the possibility of completely labeling the unqualified type of the target video. This is because the types of unqualified images corresponding to the feature vectors may be different. The more feature vectors are used for the calculation of the matching degree, the more comprehensive the unqualified type of the obtained target video.
- the server needs to match the video frame image with the image corresponding to the stored feature vector, until an image matching the video frame image with the preset threshold is obtained.
- a lot of calculations are performed, so a high-performance server is required, which is costly, and the efficiency of video review is low due to the large amount of calculation.
- the stored feature vector of the server is allocated to each terminal device that is broadcasting the same video, and each terminal device is assigned N, so that the terminal device divides the work to complete the matching of the image corresponding to the feature vector image and the feature vector.
- the server only needs to match the video frame image and the feature vector after the terminal device reports the matching between the reported video frame image and the feature vector indicated by the feature vector identifier, and the terminal device only has the video frame image.
- the matching degree of the image corresponding to the feature vector assigned thereto is greater than or equal to the first preset threshold, the image frame image and the feature vector are reported to the server, and therefore, the calculation amount of the service is very small, the ordinary server It can be done, and the cost of the server is low. Since the division of the terminal device completes the matching of the image of the video frame to be identified and the image corresponding to the feature vector, the efficiency of the video review is greatly improved compared with the prior art.
- the terminal device only reports the video frame image when the matching degree of the video frame image and the image corresponding to the feature vector assigned thereto is greater than or equal to the first preset threshold.
- the identification of the feature vector to the server does not require the image of the currently playing video sent by the terminal device in the prior art to the server, the consumed traffic is small, and the video review delay is small.
- the terminal devices are divided into 50 groups, each group including two terminal devices, each The execution time of the video review information corresponding to the two terminal devices in the group is different.
- the execution time of the terminal device 1 in the group A is 8:20 mS
- the execution time of the terminal device 2 in the group B is 8:70 mS.
- the time taken by the terminal device 1 to obtain the matching degree between the feature vector and the first video frame image is 100 ms, which indicates that the terminal device 1 can obtain the matching result between the first video frame images at 8:120 mS, and the terminal device 2
- the matching result between the second video frame images can be obtained at 8:220 mS;
- the time taken by the terminal device 2 to obtain the matching degree between the feature vector and the first video frame image is 80 ms, indicating that the terminal device 2 is at 8:150 mS.
- the result of the matching between the first video frame images is obtained. It can be understood that the first video frame image corresponding to the terminal device 1 is different from the first video frame image corresponding to the terminal device 2, if the terminal device 1 is at 8 o'clock.
- the terminal device determines that the target video is qualified at 8:120 mS, and does not report to the server, and the terminal device 1 obtains the second video frame image at 8:220 mS. If the matching result is greater than or equal to the first preset threshold, the terminal device determines that the target video is unqualified at 8:220 mS and reports to the server. If the terminal device 2 obtains the first video frame image at 8:150 mS, If the matching result is greater than or equal to the first preset threshold, the terminal device determines that the target video is unqualified at the current time, and reports the result to the server.
- the server does not need to wait until 8:20 mS, the server waits for the report result. Therefore, in this scenario, the execution time of the terminal devices in the group is different, and the speed at which the matching between the feature vector and the video frame image of the target video is greater than the first preset threshold is accelerated, that is, the speed is accelerated.
- the speed of reviewing the unqualified video can prohibit the playback of the target video in a short period of time and reduce the adverse effects on society.
- the method for live video review of the present embodiment includes receiving video review information from a server, where the video review information includes respective identification and execution times of N feature vectors, N ⁇ 1; N feature vectors are all feature vectors stored by the server. Partial feature vectors; acquiring N feature vectors according to respective identifiers of the N feature vectors; acquiring the first video frame image of the live video during execution time, if the M images corresponding to the N feature vectors exist and are first Sending, by the identifier of the first frame, the identifier of the first feature vector corresponding to the first image, the identifier of the first video frame, and the identifier of the video to the server, so that the matching degree of the video frame image is greater than or equal to the first image of the first preset threshold. After the server determines that the video is unqualified, the server fails to mark the video.
- the live video review method of the present implementation has low performance requirements on the server, fast video review, low traffic consumption, and small delay in video review.
- the solution provided by the embodiment of the present application is introduced for the functions implemented by the server and the terminal device.
- the server and the terminal device include corresponding hardware structures and/or software modules for performing the respective functions in order to implement the respective functions described above.
- the embodiments of the present application can be implemented in a combination of hardware or hardware and computer software in combination with the examples and steps described in the embodiments disclosed in the application. Whether a function is implemented in hardware or computer software to drive hardware depends on the specific application and design constraints of the solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the technical solutions of the embodiments of the present application.
- the embodiment of the present application may divide the function module into the server and the terminal device according to the foregoing method example.
- each function module may be divided according to each function, or two or more functions may be integrated into one processing unit.
- the above integrated unit can be implemented in the form of hardware or in the form of a software function module.
- FIG. 3 is a schematic structural diagram of a device for performing live video review according to an embodiment of the present disclosure.
- the device in this embodiment includes a receiving module 31, an obtaining module 32, and a sending module 33.
- the receiving module 31 is configured to receive video review information from the server, where the video review information includes an identifier and an execution time of each of the N feature vectors, N ⁇ 1; the N feature vectors are in all feature vectors stored by the server. Partial feature vector;
- the obtaining module 32 is configured to acquire N feature vectors according to respective identifiers of the N feature vectors;
- the obtaining module 32 is further configured to: acquire, in the execution time, a first video frame image of a video being broadcasted, at least part of the M images corresponding to the N feature vectors, and the first video frame Image matching degree;
- the sending module 33 is configured to: if the first image corresponding to the first video frame image is greater than or equal to the first preset threshold in the M images corresponding to the N feature vectors, the first image is corresponding to The identifier of the first feature vector, the first video frame image, the identifier of the video is sent to the server, so that the server performs the unqualified labeling of the video after determining that the video is unqualified.
- the device in this embodiment may be used to implement the technical solution of the foregoing method embodiment, and the implementation principle and the technical effect are similar, and details are not described herein again.
- the acquiring module is further configured to: when the terminal device acquires the matching degree between the M images corresponding to the N feature vectors and the first video frame image, acquire the second video frame. image;
- the server performs a failure labeling on the video after determining that the video is unqualified.
- the sending module is further configured to
- the server Before receiving video review information from the server: transmitting information of the video to the server, the information of the video including an identifier of the video and an identifier of the terminal device; an identifier of the video and an identifier of the terminal device are used for the
- the server determines a first number of terminal devices that play the video, the first number and a second number of all feature vectors stored by the server for determining a value of the N and an execution time.
- the obtaining module is specifically configured to:
- the device in this embodiment may be used to implement the technical solution of the foregoing method embodiment, and the implementation principle and the technical effect are similar, and details are not described herein again.
- FIG. 4 is a schematic structural diagram of a device for performing live video review according to an embodiment of the present disclosure.
- the device in this embodiment includes: a sending module 41, a receiving module 42, and an labeling module 43.
- the sending module 41 is configured to send video review information to the terminal device, where the video review information includes an identifier and an execution time of each of the N feature vectors, where the execution time is used to instruct the terminal device to acquire the M corresponding to the N feature vectors for the first time.
- the time at which the image is matched to the video frame image, and the N feature vectors are partial feature vectors in all feature vectors stored by the server;
- the receiving module 42 is configured to receive an identifier of a target feature vector, a target video frame image, and a video identifier sent by the terminal device, where the target feature vector is the target video frame in the M image determined by the terminal device And the matching degree of the feature vector corresponding to the image is greater than the feature vector corresponding to the first image of the first preset threshold;
- the labeling module 43 is configured to perform the unqualified labeling on the video if the second image corresponding to the target video frame image is greater than the second preset threshold in the first image corresponding to the target feature vector.
- the device in this embodiment may be used to implement the technical solution of the foregoing method embodiment, and the implementation principle and the technical effect are similar, and details are not described herein again.
- the labeling module is specifically configured to store the identifier of the video in association with the type of the second image.
- the device in this embodiment may be used to implement the technical solution of the foregoing method embodiment, and the implementation principle and the technical effect are similar, and details are not described herein again.
- FIG. 5 is a schematic structural diagram 3 of a device for performing live video review according to an embodiment of the present disclosure.
- the present embodiment further includes: a determining module 44 and a processing module 45, based on the device shown in FIG.
- the receiving module 42 is further configured to receive information about the video that is sent by the terminal device, where the information of the video includes an identifier of the video and an identifier of the terminal device;
- the determining module 44 is configured to determine, according to the information of the video that includes the identifier of the video that is received at the current time, the first quantity of the terminal device that plays the video;
- the processing module 45 is configured to perform structural processing on the target video frame image if there is a second image in the first image corresponding to the target feature vector that has a matching degree with the target video frame image that is greater than a second preset threshold. Obtaining structured data of the target video frame image; the structured data is used to acquire a model, and the model is used to acquire a feature vector.
- the device in this embodiment may be used to implement the technical solution of the foregoing method embodiment, and the implementation principle and the technical effect are similar, and details are not described herein again.
- FIG. 6 is a schematic structural diagram 1 of a terminal device provided by the present application, including a processor 51 and a communication bus 52.
- the processor 51 is configured to invoke a program instruction stored in a memory to implement a method performed by a terminal device in the foregoing method embodiment. It is a memory external to the communication device.
- FIG. 7 is a schematic structural diagram of a terminal device provided by the present application, including a processor 61, a memory 62, and a communication bus 63.
- the processor 61 is configured to invoke a program instruction stored in the memory 62 to implement the foregoing method embodiment. Methods.
- FIG. 8 is a schematic structural diagram 1 of a server provided by the present application, including a processor 71 and a communication bus 72.
- the processor 71 is configured to invoke a program instruction stored in a memory to implement a method performed by a server in the foregoing method embodiment, where the memory is a communication. Memory external to the device.
- FIG. 9 is a schematic structural diagram 2 of a server provided by the present application, including a processor 81, a memory 82, and a communication bus 83.
- the processor 81 is configured to invoke a program instruction stored in the memory 82 to implement a method executed by the server in the foregoing method embodiment. .
- Embodiments of the present application are also directed to a computer storage medium including instructions that, when executed on a communication device, cause the communication device to perform a method corresponding to the terminal device.
- Embodiments of the present application are also directed to a computer storage medium comprising instructions that, when executed on a communication device, cause the communication device to perform a method corresponding to the server.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Security & Cryptography (AREA)
- Databases & Information Systems (AREA)
- Image Analysis (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
Abstract
本申请实施例提供一种直播视频审查的方法和装置,该方法包括接收来自服务器的视频审查信息,视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;N个特征向量是服务器存储的所有特征向量中的部分特征向量;根据N个特征向量各自的标识,获取N个特征向量;在执行时间,获取正在直播的视频的第一视频帧图像,若N个特征向量对应的M张图像中存在与第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,第一视频帧图像,视频的标识发送至服务器,以使服务器在确定视频不合格后,对视频进行不合格标注。本申请提供的方法,对服务器性能要求低、视频审查速度快、流量消耗少且视频审查的时延小。
Description
本申请要求于2018年4月20日提交中国专利局、申请号为2018103612436、申请名称为“直播视频审查的方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及计算机技术领域,尤其涉及一种直播视频审查的方法和装置。
视频内容的应用日益广泛,对视频内容的审查为视频内容的处理的重要的一部分。
现有技术中,一种常用的方法为服务器之间直接交互,即第一服务器从第二服务器上获取正在播放的视频,对视频进行定时抽帧识别,以识别出黄色、暴力、恐怖、恶心、政治敏感等限制级视频;另一种常用的方法为服务器接收终端设备周期性发送的正在播放的视频的图像,并对接收到的图像进行识别,以识别出黄色、暴力、恐怖、恶心、政治敏感等限制级视频。
但是,上述方法均在服务器端实现视频的审查,而针对互联网的海量直播视频流,需要高性能服务器来实现,成本巨大且审查效率低;并且,对于服务器接收终端设备周期性发送的正在播放的视频的图像的视频审查方法,还存在视频审查时延大、流量消耗大的问题。
发明内容
本申请实施例提供一种直播视频审查的方法和装置,无需高性能服务器、审查效率高且流量消耗小、视频审查时延小。
第一方面,本申请实施例提供一种直播视频审查的方法,包括:
接收来自服务器的视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;所述N个特征向量是服务器存储的所有特征向量中的部分特征向量;
根据所述N个特征向量各自的标识,获取N个特征向量;
在所述执行时间,获取正在直播的视频的第一视频帧图像,若所述N个特征向量对应的M张图像中存在与所述第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,所述第一视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
在一种可能的设计中,所述方法还包括:在所述终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,获取第二视频帧图像;
若所述M张图像中存在与所述第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,将第二图像对应的第二特征向量的标识,所述第二视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注,M≤N。
在一种可能的设计中,在接收来自服务器的视频审查信息之前,还包括:
向服务器发送所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述视频的标识和终端设备的标识用于所述服务器确定播放所述视频的终端设备的第一数量,所述第一数量和所述服务器存储的所有特征向量的第二数量用于确定所述N的值和执行时间。
在一种可能的设计中,所述根据所述N个特征向量各自的标识,获取N个特征向量,包括:
根据所述N个特征向量各自的标识,判断终端设备中是否具有所述N个特征向量中的特征向量;
对于终端设备中存储的所述N个特征向量中的每个第三特征向量:
从所述终端设备的存储器中获取所述第三特征向量;
对于终端设备中未存储的所述N个特征向量中的每个第四特征向量:
将所述第四特征向量的标识发送至服务器;
接收所述服务器发送的第四特征向量。
在一种可能的设计中,所述第一视频帧图像为在所述执行时间正在显示的视频帧图像的下一帧图像,所述第一视频帧图像为存储终端设备的待显示的视频帧的缓存中的图像。
在一种可能的设计中,所述第二视频帧图像的播放时间晚于所述第一视频帧图像,所述第二视频帧图像为存储终端设备的待显示的视频帧的缓存中的图像。
第二方面,本申请实施例提供一种直播视频审查的方法,包括:
向终端设备发送视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间;所述执行时间用于指示终端设备首次获取与N个特征向量对应的M张图像进行匹配的视频帧图像的时间,N个特征向量是服务器存储的所有特征向量中的部分特征向量;
接收终端设备发送的目标特征向量的标识,目标视频帧图像,视频的标识;所述目标特征向量为所述终端设备确定的所述M张图像中与所述目标视频帧图像对应的特征向量的匹配度大于第一预设阈值的第一图像对应的特征向量;
若目标特征向量对应的图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述视频进行不合格标注,M≤N。
在一种可能的设计中,所述方法还包括:
若目标特征向量对应的图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述目标视频帧图像进行结构化处理,得到所述目标视频帧图像的结构化数据;所述结构化数据用于获取模型,所述模型用于获取特征向量。
在一种可能的设计中,所述向终端设备发送视频审查信息之前,所述方法还包括:
接收终端设备发送的所述视频的信息,所述视频的信息包括所述视频的标识和终 端设备的标识;
根据当前时刻接收到的包括所述视频的标识的视频的信息,确定播放所述视频的终端设备的第一数量;
根据所述第一数量和所述服务器存储的所有特征向量的第二数量确定发送至播放所述视频的终端设备各自的视频审查信息。
在一种可能的设计中,对所述视频进行不合格标注,包括:
将所述视频的标识与第二图像的不合格类型关联存储。
第三方面,本申请实施例提供一种直播视频审查的装置,包括:
接收模块,用于接收来自服务器的视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;所述N个特征向量是服务器存储的所有特征向量中的部分特征向量;
获取模块,用于根据所述N个特征向量各自的标识,获取N个特征向量;
所述获取模块,还用于在所述执行时间,获取正在直播的视频的第一视频帧图像,所述N个特征向量对应的M张图像中的至少部分图像与所述第一视频帧图像的匹配度;
发送模块,用于若所述N个特征向量对应的M张图像中存在与所述第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,所述第一视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注,M≤N。
在一种可能的设计中,所述方法还包括:所述获取模块,还用于在所述终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,获取第二视频帧图像;
若所述M张图像中存在与所述第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,将第二图像对应的第二特征向量的标识,所述第二视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
在一种可能的设计中,所述发送模块,还用于,
在接收来自服务器的视频审查信息之前:向服务器发送所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述视频的标识和终端设备的标识用于所述服务器确定播放所述视频的终端设备的第一数量,所述第一数量和所述服务器存储的所有特征向量的第二数量用于确定所述N的值和执行时间。
在一种可能的设计中,所述获取模块,具体用于:
根据所述N个特征向量各自的标识,判断终端设备中是否具有所述N个特征向量中的特征向量;
对于终端设备中存储的所述N个特征向量中的每个第三特征向量:
从所述终端设备的存储器中获取所述第三特征向量;
对于终端设备中未存储的所述N个特征向量中的每个第四特征向量:
将所述第四特征向量的标识发送至服务器;
接收所述服务器发送的第四特征向量。
第四方面,本申请实施例提供一种直播视频审查的装置,包括:
发送模块,用于向终端设备发送视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间;所述执行时间用于指示终端设备首次获取与N个特征向量对应的M张图像进行匹配的视频帧图像的时间,N个特征向量是服务器存储的所有特征向量中的部分特征向量;
接收模块,用于接收终端设备发送的目标特征向量的标识,目标视频帧图像,视频的标识;所述目标特征向量为所述终端设备确定的所述M张图像中与所述目标视频帧图像的匹配度大于第一预设阈值的第一图像对应的特征向量;
标注模块,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述视频进行不合格标注,M≤N。
在一种可能的设计中,所述装置还包括:
处理模块,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述目标视频帧图像进行结构化处理,得到所述目标视频帧图像的结构化数据;所述结构化数据用于获取模型,所述模型用于获取特征向量。
在一种可能的设计中,其特征在于,
所述接收模块,还用于,接收终端设备发送的所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;
所述装置还包括:确定模块;
所述确定模块,用于根据当前时刻接收到的包括所述视频的标识的视频的信息,确定播放所述视频的终端设备的第一数量;
根据所述第一数量和所述服务器存储的所有特征向量的第二数量确定发送至播放所述视频的终端设备各自的视频审查信息。
在一种可能的设计中,所述标注模块,具体用于将所述视频的标识与第二图像的不合格类型关联存储。
第五方面,本申请实施例提供一种终端设备,包括处理器;
所述处理器用于与存储器耦合,读取并执行所述存储器中的指令,以实现第一方面以及任一可能的设计所述的方法。
在一可能的设计中,还包括所述存储器。
第六方面,本申请实施例提供一种服务器,包括处理器;
所述处理器用于与存储器耦合,读取并执行所述存储器中的指令,以实现第二方面以及任一可能的设计所述的方法。
在一可能的设计中,还包括所述存储器。
第七方面,本申请实施例提供一种计算机存储介质,包括指令,当所述指令在通信装置上运行时,使得所述通信装置执行第一方面以及任一可能的设计所述的方法。
第八方面,本申请实施例提供一种计算机存储介质,包括指令,当所述指令在通信装置上运行时,使得所述通信装置执行第二方面以及任一可能的设计所述的方法。
图1为本申请实施例提供的系统架构图;
图2为本申请实施例提供的视频审查的方法的交互流程图;
图3为本申请实施例提供的直播视频审查的装置的结构示意图一;
图4为本申请实施例提供的直播视频审查的装置的结构示意图二;
图5为本申请实施例提供的直播视频审查的装置的结构示意图三;
图6为本申请提供的终端设备的结构示意图一;
图7为本申请提供的终端设备的结构示意图二;
图8为本申请提供的服务器的结构示意图一;
图9为本申请提供的服务器的结构示意图二。
首先对申请实施例涉及的技术名词进行解释。
特征向量:为对图像进行机器学习算法后,最终得到的向量。比如,机器学习算法可为卷积神经网络算法,特征向量为图像的基本数据经卷积神经网络算法,卷积神经网络的全连接层输出的向量;图像的基本数据可为图像的像素的灰度值或者图像的像素的局部二值模式(Local Binary Pattern,简称LBP)值。其中,卷积神经网络算法为现有技术中常用的算法,本实施例中不再赘述。
终端设备:终端设备也可以称为用户设备(user equipment,UE)、接入终端、用户单元、用户站、移动站、移动台、远方站、远程终端、移动设备、用户终端、终端、无线通信设备、用户代理或用户装置。终端设备可以是无线局域网(wireless local area networks,WLAN)中的站点(station,ST),可以是蜂窝电话、无绳电话、会话启动协议(session initiation protocol,SIP)电话、无线本地环路(wireless local loop,WLL)站、个人数字处理(personal digital assistant,PDA)设备、具有无线通信功能的手持设备、计算设备或连接到无线调制解调器的其它处理设备、车载设备、可穿戴设备以及下一代通信系统,例如,第五代通信(fifth-generation,5G)网络中的终端设备或者未来演进的公共陆地移动网络(public land mobile network,PLMN)网络中的终端设备,新空口(new radio,NR)通信系统中的终端设备等。
作为示例而非限定,在本申请实施例中,该终端设备还可以是可穿戴设备。可穿戴设备也可以称为穿戴式智能设备,是应用穿戴式技术对日常穿戴进行智能化设计、开发出可以穿戴的设备的总称,如眼镜、手套、手表、服饰及鞋等。可穿戴设备即直接穿在身上,或是整合到用户的衣服或配件的一种便携式设备。可穿戴设备不仅仅是一种硬件设备,更是通过软件支持以及数据交互、云端交互来实现强大的功能。广义穿戴式智能设备包括功能全、尺寸大、可不依赖智能手机实现完整或者部分的功能,例如:智能手表或智能眼镜等,以及只专注于某一类应用功能,需要和其它设备如智能手机配合使用,如各类进行体征监测的智能手环、智能首饰等。
另外,终端设备还可以包括无人机,如无人机上的机载通信设备等。
图1为本申请实施例提供的系统架构图,参见图1,该系统架构包括服务器11和终端设备12。其中,服务器中存储有多个特征向量,多个特征向量是服务器根据多张禁止播放的图像的基本数据采用机器学习算法训练得到模型后,根据模型以及该机器学习算法得到的。比如,该机器学习算法可为深度卷积神经网络算法,模型为深度卷 积神经网络模型。一张禁止播放的图像可对应至少一个特征向量。
终端设备12接收来自服务器11的视频审查信息,在视频审查信息指示的执行时间,获取正在直播的视频的第一视频帧图像,若视频审查信息指示的N个特征向量对应的M张图像中存在与正在直播的视频的第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将每个第一图像对应的第一特征向量的标识,第一视频帧图像,视频的标识发送至服务器11。其中,N个特征向量为服务器中存储的全部特征向量中的部分特征向量。
服务器11获取第一视频帧图像与每个第一特征向量对应的第一图像的匹配度,当存在与第一视频帧图像的匹配度大于第二预设阈值的第一图像时,则发送停止播放该视频的控制指令至终端设备12,并对视频进行不合格标注。
本申请实施的直播视频的审查方法,每个终端设备完成待识别的视频帧图像与审查信息指示的N个特征向量对应的M张图像的匹配,即终端设备分工完成待识别的视频帧图像与服务器存储的所有特征向量对应的所有图像的匹配,服务器只需在终端设备上报了视频帧图像与特征向量的标识后,进行上报的视频帧图像与特征向量的标识指示的图像之间的匹配,而终端设备只有在视频帧图像与视频审查信息中指示的特征向量对应的图像的匹配度大于或等于第一预设阈值时,才会上报视频帧图像与特征向量的标识至服务器,因此,服务的计算量非常小,普通的服务器即可完成,服务器的成本低。由于终端设备分工完成待识别的视频帧图像与特征向量的匹配,视频审查的效率大大提高。
此外,本实施例的直播视频的审查方法,终端设备只有在视频帧图像与视频审查信息中指示的特征向量对应的图像的匹配度大于或等于第一预设阈值时,才会上报视频帧图像与特征向量的标识至服务器,无需现有技术中终端设备周期性发送的正在播放的视频的图像至服务器,消耗的流量小,视频审查时延小。
下面采用具体的实施例对本申请实施例提供的直播视频审查的方法进行详细的阐述,以下实施例中的执行主体可为图1中的服务器11。
图2为本申请实施例提供的视频审查的方法的交互流程图,参见图2,本申请实施例的方法包括:
步骤S201、终端设备向服务器发送正在直播的视频的信息,视频的信息包括该视频的标识和终端设备的标识;
步骤S202、服务器根据当前时刻接收到的包括该视频的标识的视频的信息,确定播放该视频的终端设备的第一数量;
步骤S203、服务器根据播放该视频的终端设备的第一数量和服务器中存储的所有特征向量的第二数量确定发送至各播放该视频的终端设备各自的视频审查信息,视频审查信息包括N个特征向量各自的标识和执行时间,N≥1,N个特征向量是服务器存储的所有特征向量中的部分特征向量;
步骤S204、服务器发送相应的视频审查信息至终端设备;
步骤S205、终端设备根据视频审查信息中的N个特征向量各自的标识,获取N个特征向量;
步骤S206、终端设备在视频审查信息指示的执行时间,获取当前正在播放的视频 的第一视频帧图像,并获取第一视频帧图像与N个特征向量对应的M张图像的匹配度;第一视频帧图像为在上述执行时间正在显示的视频帧图像的下一帧图像,M≤N;
步骤S207、若N个特征向量对应的M张图像中存在与第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,终端设备将第一图像对应的第一特征向量的标识,第一视频帧图像,视频的标识发送至服务器;
步骤S208、服务器获取第一视频帧图像与第一特征向量对应的第一图像的匹配度;
若存在与第一视频帧图像的匹配度大于第二预设阈值的第一图像,执行步骤S209~S210:
步骤S209、服务器对该视频进行不合格标注;
步骤S210、服务器发送停止播放该视频的控制指令至终端设备;
若在终端设备获取完N个特征向量对应的M张图像与第一视频帧图像的匹配度之前,终端设备没有接收到服务器发送停止播放目标视频的控制指令,则执行步骤S211~步骤S213。
步骤S211、在终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,终端设备获取第二视频帧图像,并获取第二视频帧图像与N个特征向量对应的M张图像的匹配度;第二视频帧图像的播放时间晚于第一视频帧图像,第二视频帧图像为存储终端设备的待显示的视频帧的缓存中的图像;
步骤S212、若N个特征向量对应的M张图像中存在与第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,终端设备将第二图像对应的第二特征向量的标识,第二视频帧图像,视频的标识发送至服务器;
步骤S213、服务器获取第二视频帧图像与第二特征向量对应的第二图像的匹配度;
若存在与第二视频帧图像的匹配度大于第二预设阈值的第二图像,执行步骤S214~步骤S215
步骤S214:服务器对该视频进行不合格标注。
步骤S215、服务器发送停止播放该视频的控制指令至终端设备。
具体地,对于步骤S201、当终端设备从能够提供直播视频的服务器接收直播的目标视频时,终端设备将该直播的目标视频的信息发送至服务器;其中,目标视频的信息可包括:目标视频的标识、终端设备的标识,目标视频的信息还可包括:终端设备的IP地址、目标视频的名称、目标视频的地址。
对于步骤S202、服务器接收终端设备发送的正在直播的目标视频的信息,可以理解的是,服务器在同一时刻会接收到多个终端设备的多个视频的信息,每个视频信息中都具有相应视频的标识和终端设备的标识。对于该终端设备正在直播的目标视频,服务器统计当前时刻接收到的视频的信息中具有目标视频的标识的数量,该数量便为正在播放该目标视频的终端设备的第一数量。即具有目标视频的标识的所有视频的信息中包括的终端的标识的数量,便是正在播放该目标视频的终端设备的第一数量,即具有目标视频的标识的所有视频的信息中包括的各终端标识各自指示的终端设备便为在播放该目标视频的终端设备。
对于步骤S203、服务器中存储有多个特征向量,多个特征向量是服务器根据多张禁止播放的图像的基本数据采用机器学习算法训练得到模型后,根据模型以及该机器 学习算法得到的。比如,政治敏感人物A的人脸图像对应的至少一个特征向量、涉恐图像对应的至少一个特征向量等等。其中,禁止播放的视频的类型可为:黄色、暴力、恐怖、恶心、政治敏感等等。由于视频是由多帧的图像组成,因此,服务器中存储的特征向量基于的是禁止播放的图像获取的。
服务器在获取到播放该目标视频的终端设备的第一数量、播放该目标视频的终端设备以及服务器中存储的所有特征向量的第二数量后,根据播放该目标视频的终端设备的第一数量和服务器中存储的所有特征向量的第二数量确定发送至该终端设备的视频审查信息,视频审查信息包括N个特征向量各自的标识和执行时间,N≥1,N个特征向量是服务器存储的所有特征向量中的部分特征向量。
具体地,服务器会根据第一数量和第二数量,确定发送至正在播放该目标视频的各终端设备各自的视频审查信息,具体如下:
若第一数量为K,第二数量为L,且K>L,说明终端设备的数量多,特征向量的数量少,此时,服务器将播放该目标视频的K个终端设备分成L组,每组内的每个终端设备被分配相同的特征向量,且每组内的每个终端设备被分配的特征向量的数量为1,每个组被分配的特征向量不相同。服务器还为每组内的每个终端设备配置不同的执行时间,执行时间用于指示终端设备首次获取与特征向量对应的图像进行匹配的视频帧图像的时间。
此时,在K>L的场景下,服务器确定的发送至播放该目标视频的每个终端设备的视频审查信息中包括1张特征向量的标识和执行时间,即N=1。
第一数量为K,第二数量为L,且K≤L,说明终端设备的数量少,特征向量的数量多,此时,每个播放该目标视频的终端设备将至少被分配一个特征向量;比如,若K=50,L=100,则每个播放该目标视频的终端设备被分配两个特征向量,若K=50,L=80,则有的终端设备被分配一个特征向量,有的终端设备被分配两个特征向量,每个终端设备被分配的特征向量不相同。服务器还为每个终端设备设置执行时间,执行时间用于指示终端设备首次获取与特征向量对应的图像进行匹配的视频帧图像的时间。此时,每个终端设备的执行时间可相同。
此时,在K≤L的场景下,服务器确定的发送至播放该目标视频的每个终端设备的视频审查信息中包括N张特征向量的标识和执行时间,N≥1,每个终端设备对应的N可能相同也可能不相同。
对于步骤S204、服务器将为各播放该目标视频的各终端设备配置的各自视频审查信息发送至各终端设备。
对于步骤S205、终端设备接收到服务器发送的视频审查信息后,根据视频审查信息中的N个特征向量各自的标识,获取N个特征向量;
根据视频审查信息中的N个特征向量各自的标识,获取N个特征向量,具体包括:根据N个特征向量各自的标识,判断终端设备中是否具有N个特征向量中的特征向量;
对于终端设备中存储的N个特征向量中的每个第三特征向量:从终端设备的存储器中获取第三特征向量;
对于终端设备中未存储的N个特征向量中的每个第四特征向量:将第四特征向量的标识发送至服务器,接收服务器发送的第四特征向量。
由于终端设备会进行多次的视频审查,因此终端设备自身会存储一些之前视频审查时获取的特征向量。对于终端设备中未存储的N个特征向量中的每个第四特征向量,则从服务器下载即可。
对于步骤S206、终端设备在视频审查信息指示的执行时间,获取当前正在播放的视频的第一视频帧图像,并获取第一视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度。
具体地,可选地,第一视频帧图像为在执行时间正在显示的视频帧图像的下一帧图像,也就是第一视频帧图像为存储终端设备的待显示的视频帧缓存中将被最先进行播放的图像,或者说位于视频显示帧缓存中的将被最先进行播放的图像。其中,视频显示帧缓存即为存储终端设备的待显示的视频帧缓存,其中存储的是即将进行播放的几帧视频帧图像。
获取到第一视频帧图像和N个特征向量后,便可以获取第一视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度。其中,匹配度也可称为相似度,至少部分为全部或者部分。
如上所述,视频审查信息中包括的N个特征向量的标识指示的N个特征向量,是M个禁止播放的图像各自对应的N特征向量,一张禁止播放的图像对应至少一个特征向量。
获取两张图像的匹配度的方法具有很多,一种常用的方法为:若上述N个特征向量是采用卷积神经网络算法获取得到的,那么采用相同的卷积神经网络算法获取第一视频帧图像对应的特征向量,计算第一视频帧图像对应的特征向量与N个特征向量中的特征向量之间的距离,然后用归一化法将距离值归一化为匹配度,其中,归一化方法为线性映射、分段线性映射以及其他单调函数的方法。上述各归一化方法均为现有技术中的方法,本实施例中不再赘述。
还有一种常用的方法:若上述N个特征向量是采用卷积神经网络算法获取得到的,采用相同卷积神经网络获取第一视频帧图像对应的特征向量,直接根据第一视频帧图像对应的特征向量与N个特征向量中的特征向量获取第一视频帧图像与N个特征向量中的特征向量对应的图像的匹配度。直接根据第一视频帧图像对应的特征向量与N个特征向量中的特征向量获取第一视频帧图像与N个特征向量中的特征向量对应的图像的匹配度的计算公式为现有的公式,本实施例中不再赘述。
也就是说:获取第一视频帧图像与N个特征向量中的特征向量对应的图像的匹配度,就是采用N个特征向量中的特征向量与第一视频帧图像对应的特征向量计算相应两张图像之间的匹配度。
对于步骤S207、若N个特征向量对应的M张图像中存在与第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则终端设备将第一图像对应的第一特征向量的标识,第一视频帧图像,视频的标识发送至服务器;或者,若N个特征向量对应的M张图像中存在与第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则终端设备将第一图像对应的第一特征向量的标识,第一视频帧图像,第一视频帧图像的时间戳,视频的标识发送至服务器。
即若采用N个特征向量中的特征向量x与第一视频帧图像对应的特征向量计算得 到的匹配度大于第一预设阈值,则特征向量x为第一特征向量,特征向量x的标识会被终端设备发送至服务器。
其中,第一预设阈值可为80~85%之间。第一特征向量对应的第一图像与第一视频帧图像的匹配度大于或等于第一预设阈值,说明第一视频帧图像与第一特征向量对应的第一图像很相似,也就是说目标视频为不合格视频的概率很大,因此,终端设备将第一特征向量的标识,第一视频帧图像,第一视频帧图像的时间戳,视频的标识发送至服务器。
其中,第一特征向量的个数可为一个或多个。
若终端设备中接收到的视频审查信息中包括多张特征向量的标识,终端设备可依次获取第一视频帧图像和多个特征向量各自对应的图像的匹配度,终端设备只要得到第一视频图像和特征向量对应的图像的匹配度大于等于第一预设阈值的情况,就立即将相应特征向量的标识,第一视频帧图像,第一视频帧图像的时间戳和目标视频的标识发送至服务器,而不是在得到多张特征向量对应的图像与第一视频帧图像的匹配度后,才将与第一视频图像的匹配度大于第一预设阈值的第一图像的第一特征向量的标识,第一视频帧图像,目标视频的标识发送至服务器。
比如,终端设备3接收到的视频审查信息携带的特征向量的标识为两个,一个指示特征向量1,另一个指示特征向量2,特征向量1对应的图像的不合格类型为政治敏感类型、特征向量2对应的图像的不合格类型为恐怖类型,且终端设备先获取的第一视频帧图像和特征向量1对应的图像1的第一匹配度,得到的第一匹配度为87%,当第一预设阈值为80%时,说明特征向量1为第一特征向量,终端设备会立即将特征向量1的标识,第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识均会发送至服务器。在得到第一视频帧图像和特征向量1对应的图像1的第一匹配度后,终端设备会接着获取第一视频帧图像和特征向量2的第二匹配度,也就是在终端设备会将特征向量1的标识,第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识均会发送至服务器的同时,终端设备会获取第一视频帧图像和特征向量2的第二匹配度;若得到的第二匹配度为70%,则终端设备不发送任何信息至服务器,若得到的第二匹配度为85%,说明特征向量2为第一特征向量,则终端设备会将特征向量2的标识、第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识发送至服务器。
对于步骤S208~S210、服务器接收到终端设备发送的第一特征向量的标识,第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识后,重新获取第一视频帧图像与第一特征向量对应的第一图像的匹配度,若存在与第一视频帧图像的匹配度大于第二预设阈值的第二图像,第二图像对应的第一特征向量称为第一目标特征向量,则发送停止播放目标视频的控制指令至终端设备并对视频进行不合格标注。
其中,第一预设阈值可为90~95%之间。服务器获取的第一目标特征向量对应的第二图像与第一视频帧图像的匹配度大于或等于第二预设阈值,说明第一视频帧图像与该第一目标特征向量对应的第二图像十分相似,基本可认定该目标视频为不合格视频,则对目标视频进行不合格标注。
可以理解的是,在服务器在确定目标视频为不合格视频后,需要立即控制各正在直播目标的终端设备不能继续播放直播视频,即发送停止播放目标视频的控制指令至 各终端设备。为了降低不合格视频直播的时间,需要在服务器首次得到目标视频的视频帧图像和第一特征向量对应的第一图像的匹配度大于第二预设阈值后,即确定目标视频为不合格视频,立即发送停止播放目标视频的控制指令至各终端设备。
其中,对目标视频进行不合格标注,包括:将目标视频的标识与第二图像的不合格类型关联存储,也就是与第一目标特征向量对应的图像的不合格类型关联存储。
终端设备在接收到停止播放目标视频的控制指令后,对于某终端设备而言,若视频审查信息指示的N个特征向量中还具有n个特征向量未用于计算匹配度(对应步骤S206中的N个特征向量对应的M张图像中的部分图像或全部图像),则可停止获取目标视频的第一视频帧图像与n个特征向量各自对应的图像的匹配度,也可继续获取n个特征向量各自对应的图像与第一视频帧图像的匹配度(对应步骤S206中的N个特征向量对应的M张图像中的全部图像)。其中,继续获取n个特征向量各自对应的图像的匹配度,可增大对目标视频的不合格类型进行完整的标注的可能性。这是因为特征向量对应的图像的不合格的类型可能不相同,若所有的特征向量都用于了匹配度的计算,则得到的目标视频的不合格类型就越全面。
比如,终端设备3接收到的视频审查信息携带的特征向量的标识为两个,一个指示特征向量1,另一个指示特征向量2,特征向量1对应的图像的不合格类型为政治敏感类型、特征向量2对应的图像的不合格类型为恐怖类型,且终端设备先获取的第一视频帧图像和特征向量1对应的图像的第一匹配度,得到的第一匹配度为87%,当第一预设阈值为80%时,说明特征向量1为第一特征向量,终端设备会立即将特征向量1的标识、第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识发送至服务器。服务器重新获取的第一视频帧图像和特征向量1对应的图像的匹配度为93%,第二预设阈值为90%,则说明特征向量1为第一目标特征向量,若第一视频帧图像和特征向量1对应的图像的匹配度为93%是服务器首次得到目标视频的视频帧图像和特征向量对应的图像的匹配度大于第二预设阈值,则此时,即为服务器确定目标视频为不合格视频的时刻,服务器立即控制各正在直播目标的终端设备不能继续播放直播视频,并将目标视频标注为:政治敏感类型,即将目标视频的标识和“政治敏感”进行关联存储。
但是为了能够对目标视频的不合格类型进行全面的标注,服务器还会继续接收(如果有的话)各直播目标视频的终端设备(包括终端设备3)发送的特征向量的标识、第一视频帧图像、第一视频帧图像的时间戳,目标视频的标识。比如:终端设备3在获取完第一视频帧图像和特征向量1对应的图像的第一匹配度后,会接着获取的第一视频帧图像和特征向量2对应的图像的第二匹配度,若得到的第二匹配度为85%,说明特征向量2也为第一特征向量,则终端设备会将特征向量2的标识、第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识发送至服务器;服务器重新获取的第一视频帧图像和特征向量2对应的图像的匹配度为90%,第二预设阈值为90%,则说明特征向量2也为第一目标特征向量,将目标视频的标识也和“恐怖”进行关联存储。
此外,还存在如下情况:终端设备3接收到的视频审查信息携带的特征向量的标识为1个,指示特征向量1,特征向量1对应的图像的不合格类型为政治敏感类型,终端设备获取的第一视频帧图像和特征向量1对应的图像的匹配度为87%,则特征向 量1为第一特征向量,则特征向量1的标识,第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识均会发送至服务器;终端设备4接收到的视频审查信息携带的特征向量的标识为1个,指示特征向量2,特征向量2对应的图像不合格类型为政治敏感类型,终端设备获取的第一视频帧图像和特征向量2对应的图像的匹配度为85%,则特征向量2为第一特征向量,则特征向量2的标识,第一视频帧图像,第一视频帧图像的时间戳,目标视频的标识均会发送至服务器;服务器重新获取的第一视频帧图像和特征向量1对应的图像的匹配度为93%,第一视频帧图像和特征向量2对应的图像的匹配度为90%,第二预设阈值为90%,则特征向量1、特征向量2均为第一目标特征向量,则将目标视频的标识与“政治敏感”和“恐怖”关联存储,即将目标视频的标注为:“政治敏感”和“恐怖”类型的目标类型。
此外,若存在与第一视频帧图像的匹配度大于第二预设阈值的第一图像对应的第一特征向量,则服务器会对第一视频帧进行结构化处理,得到第一视频帧图像的第一结构化数据;第一结构化数据用于获取模型,模型用于获取特征向量。即第一结构化数据作为训练样本,获取用于获取特征向量的模型,且还会采用模型获取第一视频帧图像对应的特征向量,以丰富服务器中存储的特征向量,提高视频审查的准确率。第一视频帧图像和第一视频帧图像的时戳可录入不合格备案库作为举报的证据。其中,服务器会对第一视频帧进行结构化处理的方法为现有技术中的方法,本实施例中不再赘述。
对于步骤S211、若在终端设备获取完第一视频帧图像与N个特征向量对应的M张图像的匹配度之前,终端设备没有接收到服务器发送停止播放目标视频的控制指令,则在终端设备获取完第一视频帧图像与N个特征向量对应的M张图像的匹配度时,终端设备获取第二视频帧图像,并获取第二视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度。
此处的“在终端设备获取完第一视频帧图像与N个特征向量对应的M张图像的匹配度之前”包括:在终端设备获取完第一视频帧图像与N个特征向量对应的M张图像的匹配度时。
第二视频帧图像为视频显示帧缓存中与第一视频帧图像不相同的图像,且第二视频帧图像的播放时间晚于第一视频帧图像。可以理解的是,在获取第二视频帧图像时,第一视频帧图像可能已经被播放,也可能还处在视频显示帧缓存中,若第一视频帧图像还处在视频显示帧缓存,则第二视频帧图像为与第一视频帧图像相邻的下一视频帧图像。
对于不同的终端设备,由于终端设备的能力不相同,因此,每个终端设备的第一视频帧图像与第二视频帧图像的时间关系并不一定一样。
其中,获取第二视频帧图像与N个特征向量中对应的M张图像中至少部分图像的匹配度的方法同获取第一视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度的方法,本实施例中不再赘述。
对于步骤S212~步骤S215参照步骤S207~步骤S210,此处不再赘述。
本领域技术人员应当明白,若在终端设备获取完第二视频帧图像与N个特征向量对应的M张图像的匹配度之前,终端设备没有接收到服务器发送停止播放目标视频的 控制指令,则在终端设备获取完第二视频帧图像与N个特征向量对应的M张图像的匹配度时,终端设备获取第三视频帧图像,并获取第三视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度,并可进行步骤S207~步骤S210或者步骤S212~步骤S215相应的过程;第三视频帧图像为视频显示帧缓存中与第二视频帧图像不相同的图像,且第三视频帧图像的播放时间晚于第二视频帧图像。可以理解的是,在获取第三视频帧图像时,第二视频帧图像可能已经被播放,也可能还处在视频显示帧缓存中,若第二视频帧图像还处在视频显示帧缓存,则第三视频帧图像为与第二视频帧图像相邻的下一视频帧图像。
也就是说,只要终端设备没有接收到服务器发送停止播放目标视频的控制指令,终端设备就会在终端设备获取完第m视频帧图像与N个特征向量对应的M张图像的匹配度时,终端设备获取第m+1视频帧图像,并获取第m+1视频帧图像与N个特征向量对应的M张图像中至少部分图像的匹配度,并可进行步骤S207~步骤S210或者步骤S212~步骤S215相应的过程;第m+1视频帧图像为视频显示帧缓存中与第m视频帧图像不相同的图像,且第m+1视频帧图像的播放时间晚于第m视频帧图像。可以理解的是,在获取第m+1视频帧图像时,第m视频帧图像可能已经被播放,也可能还处在视频显示帧缓存中,若第m视频帧图像还处在视频显示帧缓存,则第m+1视频帧图像为与第m视频帧图像相邻的下一视频帧图像。
综上可得,终端设备在接收到停止播放目标视频的控制指令后,对于某终端设备而言,若当前视频审查信息指示的N个特征向量中还具有n个特征向量未用于计算匹配度,则可停止获取目标视频的视频帧图像与n个特征向量各自对应的图像的匹配度,也可继续获取n个特征向量各自对应的图像与相应的视频帧图像的匹配度。其中,继续获取n个特征向量各自对应的图像与相应的视频帧图像的匹配度,可增大对目标视频的不合格类型进行完整的标注的可能性。这是因为特征向量对应的图像的不合格的类型可能不相同,越多的特征向量都用于了匹配度的计算,则得到的目标视频的不合格类型就越全面。
现有技术中,对于每一帧待识别的视频帧图像,服务器要将视频帧图像依次与其存储的特征向量对应的图像进行匹配,直至得到与视频帧图像匹配度大于预设阈值的图像,需要进行大量的计算,因此,需要高性能的服务器,成本巨大,且由于计算量巨大,视频审查的效率也较低。
本实施例中,服务器其存储的特征向量分配给正在直播同一视频的各终端设备,每个终端设备被分配N个,使终端设备分工完成待识别的视频帧图像与特征向量对应的图像的匹配,服务器只需在终端设备上报了视频帧图像与特征向量的标识后,进行上报的视频帧图像与特征向量的标识指示的特征向量对应的图像之间的匹配,而终端设备只有在视频帧图像与分配给其的特征向量对应的图像的匹配度大于或等于第一预设阈值时,才会上报的视频帧图像与特征向量的标识至服务器,因此,服务的计算量非常小,普通的服务器即可完成,服务器的成本低。由于终端设备分工完成待识别的视频帧图像与特征向量对应的图像的匹配,视频审查的效率相对于现有技术大大提高。
此外,本实施例的直播视频审查的方法,终端设备只有在视频帧图像与分配给其的特征向量对应的图像的匹配度大于或等于第一预设阈值时,才会上报的视频帧图像 与特征向量的标识至服务器,无需现有技术中终端设备周期性发送的正在播放的视频的图像至服务器,消耗的流量小,视频审查时延小。
进一步地,本实施例的直播视频审查的方法中对于上述K>L的场景,比如,K=100,L=50,则100个终端设备被分成50组,每组包括两个终端设备,每组中两个终端设备对应的视频审查信息中的执行时间不相同,比如A组的中的终端设备1的执行时间为8点20mS,B组的中的终端设备2的执行时间为8点70mS,终端设备1得到特征向量与第一视频帧图像之间的匹配度所用的时间为100ms,则说明终端设备1在8点120mS可得到第一视频帧图像之间的匹配度结果,终端设备2在8点220mS可得到第二视频帧图像之间的匹配度结果;终端设备2得到特征向量与第一视频帧图像之间的匹配度所用的时间为80ms,则说明终端设备2在8点150mS可得到第一视频帧图像之间的匹配度结果,可以理解的是,终端设备1对应的第一视频帧图像与终端设备2对应的第一视频帧图像不相同,若终端设备1在8点120mS可得到第一视频帧图像之间的匹配度结果小于第一预设阈值,终端设备在8点120mS会判断该目标视频合格,不上报服务器,终端设备1在8点220mS可得到第二视频帧图像之间的匹配度结果大于或等于第一预设阈值,终端设备在8点220mS会判断该目标视频不合格,上报服务器,若终端设备2在8点150mS可得到第一视频帧图像之间的匹配度结果为大于或等于第一预设阈值,则终端设备在当前时刻会判断该目标视频不合格,上报服务器,无需等到8点220mS时,服务器才等到上报结果。因此,在该场景下,一组内的终端设备的执行时间不相同,可加快终端设备获取到特征向量与目标视频的视频帧图像的匹配度大于第一预设阈值的结果的速度,即加快了将不合格视频审查出来的速度,从而可在较短的时间内禁止该目标视频的播放,降低对社会造成的不良影响。
本实施的直播视频审查的方法,包括接收来自服务器的视频审查信息,视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;N个特征向量是服务器存储的所有特征向量中的部分特征向量;根据N个特征向量各自的标识,获取N个特征向量;在执行时间,获取正在直播的视频的第一视频帧图像,若N个特征向量对应的M张图像中存在与第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,第一视频帧图像,视频的标识发送至所述服务器,以使服务器在确定视频不合格后,对视频进行不合格标注。本实施的直播视频审查的方法,对服务器的性能要求低、视频审查速度快、流量消耗少且视频审查的时延小。
上述针对服务器和终端设备所实现的功能,对本申请实施例提供的方案进行了介绍。可以理解的是,服务器和终端设备为了实现上述各自的功能,其包含了执行各个功能相应的硬件结构和/或软件模块。结合本申请中所公开的实施例描述的各示例及步骤,本申请实施例能够以硬件或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。本领域技术人员可以对每个特定的应用来使用不同的方法来实现所描述的功能,但是这种实现不应认为超出本申请实施例的技术方案的范围。
本申请实施例可以根据上述方法示例对服务器和终端设备进行功能模块的划分,例如,可以对应各个功能划分各个功能模块,也可以将两个或两个以上的功能集成在一个处理单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能 模块的形式实现。
图3为本申请实施例提供的直播视频审查的装置的结构示意图一,参见图3,本实施例的装置包括接收模块31、获取模块32和发送模块33。
接收模块31,用于接收来自服务器的视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;所述N个特征向量是服务器存储的所有特征向量中的部分特征向量;
获取模块32,用于根据所述N个特征向量各自的标识,获取N个特征向量;
所述获取模块32,还用于在所述执行时间,获取正在直播的视频的第一视频帧图像,所述N个特征向量对应的M张图像中的至少部分图像与所述第一视频帧图像的匹配度;
发送模块33,用于若所述N个特征向量对应的M张图像中存在与所述第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,所述第一视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
本实施例的装置,可以用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
在一种可能的设计中,所述获取模块,还用于在所述终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,获取第二视频帧图像;
若所述M张图像中存在与所述第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,将第二图像对应的第二特征向量的标识,所述第二视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
在一种可能的设计中,所述发送模块,还用于,
在接收来自服务器的视频审查信息之前:向服务器发送所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述视频的标识和终端设备的标识用于所述服务器确定播放所述视频的终端设备的第一数量,所述第一数量和所述服务器存储的所有特征向量的第二数量用于确定所述N的值和执行时间。
在一种可能的设计中,所述获取模块,具体用于:
根据所述N个特征向量各自的标识,判断终端设备中是否具有所述N个特征向量中的特征向量;
对于终端设备中存储的所述N个特征向量中的每个第三特征向量:
从所述终端设备的存储器中获取所述第三特征向量;
对于终端设备中未存储的所述N个特征向量中的每个第四特征向量:
将所述第四特征向量的标识发送至服务器;
接收所述服务器发送的第四特征向量。
本实施例的装置,可以用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
图4为本申请实施例提供的直播视频审查的装置的结构示意图二,参见图4,本实施例的装置包括:发送模块41、接收模块42和标注模块43。
发送模块41,用于向终端设备发送视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间;所述执行时间用于指示终端设备首次获取与N个特征向量对应的M张图像进行匹配的视频帧图像的时间,N个特征向量是服务器存储的所有特征向量中的部分特征向量;
接收模块42,用于接收终端设备发送的目标特征向量的标识,目标视频帧图像,视频的标识;所述目标特征向量为所述终端设备确定的所述M张图像中与所述目标视频帧图像对应的特征向量的匹配度大于第一预设阈值的第一图像对应的特征向量;
标注模块43,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述视频进行不合格标注。
本实施例的装置,可以用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
在一种可能的设计中,所述标注模块,具体用于将所述视频的标识与第二图像的不合格类型关联存储。
本实施例的装置,可以用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
图5为本申请实施例提供的直播视频审查的装置的结构示意图三,参见图5,本实施例在图4所示的装置的基础上,还包括:确定模块44和处理模块45;
所述接收模块42,还用于,接收终端设备发送的所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;
所述确定模块44,用于根据当前时刻接收到的包括所述视频的标识的视频的信息,确定播放所述视频的终端设备的第一数量;
根据所述第一数量和所述服务器存储的所有特征向量的第二数量确定发送至播放所述视频的终端设备各自的视频审查信息。
处理模块45,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述目标视频帧图像进行结构化处理,得到所述目标视频帧图像的结构化数据;所述结构化数据用于获取模型,所述模型用于获取特征向量。
本实施例的装置,可以用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
图6为本申请提供的终端设备的结构示意图一,包括处理器51和通信总线52,处理器51用于调用存储器中存储的程序指令,以实现上述方法实施例中终端设备执行的方法,存储器为通信装置外部的存储器。
图7为本申请提供的终端设备的结构示意图二,包括处理器61、存储器62和通信总线63,处理器61用于调用存储器62中存储的程序指令,以实现上述方法实施例终端设备执行中的方法。
图8为本申请提供的服务器的结构示意图一,包括处理器71和通信总线72,处理器71用于调用存储器中存储的程序指令,以实现上述方法实施例中服务器执行的方法,存储器为通信装置外部的存储器。
图9为本申请提供的服务器的结构示意图二,包括处理器81、存储器82和通信 总线83,处理器81用于调用存储器82中存储的程序指令,以实现上述方法实施例中服务器执行的方法。
本申请实施例还涉及一种计算机存储介质,包括指令,当所述指令在通信装置上运行时,使得所述通信装置执行终端设备对应的方法。
本申请实施例还涉及一种计算机存储介质,包括指令,当所述指令在通信装置上运行时,使得所述通信装置执行服务器对应的方法。
以上所述,仅为本发明的具体实施方式,但本发明的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本发明揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本发明的保护范围之内。因此,本发明的保护范围应以所述权利要求的保护范围为准。
Claims (24)
- 一种直播视频审查的方法,其特征在于,包括:接收来自服务器的视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;所述N个特征向量是服务器存储的所有特征向量中的部分特征向量;根据所述N个特征向量各自的标识,获取N个特征向量;在所述执行时间,获取正在直播的视频的第一视频帧图像,若所述N个特征向量对应的M张图像中存在与所述第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,所述第一视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注;M≤N。
- 根据权利要求1所述的方法,其特征在于,所述方法还包括:在所述终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,获取第二视频帧图像;若所述M张图像中存在与所述第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,将第二图像对应的第二特征向量的标识,所述第二视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
- 根据权利要求1所述的方法,其特征在于,在接收来自服务器的视频审查信息之前,还包括:向服务器发送所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述视频的标识和终端设备的标识用于所述服务器确定播放所述视频的终端设备的第一数量,所述第一数量和所述服务器存储的所有特征向量的第二数量用于确定所述N的值和执行时间。
- 根据权利要求1~3任一项所述的方法,其特征在于,所述根据所述N个特征向量各自的标识,获取N个特征向量,包括:根据所述N个特征向量各自的标识,判断终端设备中是否具有所述N个特征向量中的特征向量;对于终端设备中存储的所述N个特征向量中的每个第三特征向量:从所述终端设备的存储器中获取所述第三特征向量;对于终端设备中未存储的所述N个特征向量中的每个第四特征向量:将所述第四特征向量的标识发送至服务器;接收所述服务器发送的第四特征向量。
- 根据权利要求1~3任一项所述的方法,其特征在于,所述第一视频帧图像为在所述执行时间正在显示的视频帧图像的下一帧图像,所述第一视频帧图像为存储终端设备的待显示的视频帧的缓存中的图像。
- 根据权利要求2所述的方法,其特征在于,所述第二视频帧图像的播放时间晚于所述第一视频帧图像,所述第二视频帧图像为存储终端设备的待显示的视频帧的缓 存中的图像。
- 一种直播视频审查的方法,其特征在于,包括:向终端设备发送视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间;所述执行时间用于指示终端设备首次获取与N个特征向量对应的M张图像进行匹配的视频帧图像的时间,N个特征向量是服务器存储的所有特征向量中的部分特征向量;接收终端设备发送的目标特征向量的标识,目标视频帧图像,视频的标识;所述目标特征向量为所述终端设备确定的所述M张图像中与所述目标视频帧图像对应的特征向量的匹配度大于第一预设阈值的第一图像对应的特征向量;M≤N;若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述视频进行不合格标注。
- 根据权利要求7所述的方法,其特征在于,所述方法还包括:若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述目标视频帧图像进行结构化处理,得到所述目标视频帧图像的结构化数据;所述结构化数据用于获取模型,所述模型用于获取特征向量。
- 根据权利要求7所述的方法,其特征在于,所述向终端设备发送视频审查信息之前,所述方法还包括:接收终端设备发送的所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;根据当前时刻接收到的包括所述视频的标识的视频的信息,确定播放所述视频的终端设备的第一数量;根据所述第一数量和所述服务器存储的所有特征向量的第二数量确定发送至播放所述视频的终端设备各自的视频审查信息。
- 根据权利要求7~9任一项所述的方法,其特征在于,对所述视频进行不合格标注,包括:将所述视频的标识与所述第二图像的不合格类型关联存储。
- 一种直播视频审查的装置,其特征在于,包括:接收模块,用于接收来自服务器的视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间,N≥1;所述N个特征向量是服务器存储的所有特征向量中的部分特征向量;获取模块,用于根据所述N个特征向量各自的标识,获取N个特征向量;所述获取模块,还用于在所述执行时间,获取正在直播的视频的第一视频帧图像,所述N个特征向量对应的M张图像中的至少部分图像与所述第一视频帧图像的匹配度;发送模块,用于若所述N个特征向量对应的M张图像中存在与所述第一视频帧图像的匹配度大于或等于第一预设阈值的第一图像,则将第一图像对应的第一特征向量的标识,所述第一视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注;M≤N。
- 根据权利要求11所述的装置,其特征在于,所述获取模块,还用于在所述终端设备获取完N个特征向量对应的M张图像与所述第一视频帧图像的匹配度时,获取 第二视频帧图像;若所述M张图像中存在与所述第二视频帧图像的匹配度大于或等于第一预设阈值的第二图像,将第二图像对应的第二特征向量的标识,所述第二视频帧图像,所述视频的标识发送至所述服务器,以使所述服务器在确定所述视频不合格后,对所述视频进行不合格标注。
- 根据权利要求11所述的装置,其特征在于,所述发送模块,还用于,在接收来自服务器的视频审查信息之前:向服务器发送所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述视频的标识和终端设备的标识用于所述服务器确定播放所述视频的终端设备的第一数量,所述第一数量和所述服务器存储的所有特征向量的第二数量用于确定所述N的值和执行时间。
- 根据权利要求11~13任一项所述的装置,其特征在于,所述获取模块,具体用于:根据所述N个特征向量各自的标识,判断终端设备中是否具有所述N个特征向量中的特征向量;对于终端设备中存储的所述N个特征向量中的每个第三特征向量:从所述终端设备的存储器中获取所述第三特征向量;对于终端设备中未存储的所述N个特征向量中的每个第四特征向量:将所述第四特征向量的标识发送至服务器;接收所述服务器发送的第四特征向量。
- 一种直播视频审查的装置,其特征在于,包括:发送模块,用于向终端设备发送视频审查信息,所述视频审查信息包括N个特征向量各自的标识和执行时间;所述执行时间用于指示终端设备首次获取与N个特征向量对应的M张图像进行匹配的视频帧图像的时间,N个特征向量是服务器存储的所有特征向量中的部分特征向量;接收模块,用于接收终端设备发送的目标特征向量的标识,目标视频帧图像,视频的标识;所述目标特征向量为所述终端设备确定的所述M张图像中与所述目标视频帧图像的匹配度大于第一预设阈值的第一图像对应的特征向量;标注模块,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述视频进行不合格标注;M≤N。
- 根据权利要求15所述的装置,其特征在于,还包括:处理模块,用于若目标特征向量对应的第一图像中存在与所述目标视频帧图像的匹配度大于第二预设阈值的第二图像,则对所述目标视频帧图像进行结构化处理,得到所述目标视频帧图像的结构化数据;所述结构化数据用于获取模型,所述模型用于获取特征向量。
- 根据权利要求15所述的装置,其特征在于,所述接收模块,还用于,接收终端设备发送的所述视频的信息,所述视频的信息包括所述视频的标识和终端设备的标识;所述装置还包括:确定模块;所述确定模块,用于根据当前时刻接收到的包括所述视频的标识的视频的信息, 确定播放所述视频的终端设备的第一数量;根据所述第一数量和所述服务器存储的所有特征向量的第二数量确定发送至播放所述视频的终端设备各自的视频审查信息。
- 根据权利要求15~17任一项所述的装置,其特征在于,所述标注模块,具体用于将所述视频的标识与第二图像的不合格类型关联存储。
- 一种终端设备,其特征在于,包括处理器;所述处理器用于与存储器耦合,读取并执行所述存储器中的指令,以实现如权1-6任一所述的方法。
- 根据权利要求19所述的终端设备,其特征在于,还包括所述存储器。
- 一种服务器,其特征在于,包括处理器;所述处理器用于与存储器耦合,读取并执行所述存储器中的指令,以实现如权7-10任一所述的方法。
- 根据权利要求21所述的服务器,其特征在于,还包括所述存储器。
- 一种计算机存储介质,包括指令,其特征在于,当所述指令在通信装置上运行时,使得所述通信装置执行如权1-6任一所述的方法。
- 一种计算机存储介质,包括指令,其特征在于,当所述指令在通信装置上运行时,使得所述通信装置执行如权7-10任一所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810361243.6 | 2018-04-20 | ||
| CN201810361243.6A CN110392271A (zh) | 2018-04-20 | 2018-04-20 | 直播视频审查的方法和装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019201008A1 true WO2019201008A1 (zh) | 2019-10-24 |
Family
ID=68240432
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/075141 Ceased WO2019201008A1 (zh) | 2018-04-20 | 2019-02-15 | 直播视频审查的方法和装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110392271A (zh) |
| WO (1) | WO2019201008A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111222397A (zh) * | 2019-10-25 | 2020-06-02 | 深圳市优必选科技股份有限公司 | 一种绘本识别方法、装置及机器人 |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111861636B (zh) * | 2020-06-19 | 2023-04-18 | 武汉理工大学 | 基于居家养老的服务数据处理方法、系统和存储介质 |
| CN114339292A (zh) * | 2021-12-31 | 2022-04-12 | 安徽听见科技有限公司 | 一种直播流的审查干预方法、装置、存储介质及设备 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103544498A (zh) * | 2013-09-25 | 2014-01-29 | 华中科技大学 | 基于自适应抽样的视频内容检测方法与系统 |
| CN103902954A (zh) * | 2012-12-26 | 2014-07-02 | 中国移动通信集团贵州有限公司 | 一种不良视频的鉴别方法和系统 |
| CN104268446A (zh) * | 2014-09-30 | 2015-01-07 | 小米科技有限责任公司 | 防止视频二次传播的方法及装置 |
| CN104540024A (zh) * | 2014-12-18 | 2015-04-22 | 网宿科技股份有限公司 | 视频终端及其限制视频播放的方法和系统 |
| CN105956550A (zh) * | 2016-04-29 | 2016-09-21 | 浪潮电子信息产业股份有限公司 | 一种视频鉴别的方法和装置 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107580259A (zh) * | 2016-07-04 | 2018-01-12 | 北京新岸线网络技术有限公司 | 一种视频内容审核方法和系统 |
| CN106412618A (zh) * | 2016-09-09 | 2017-02-15 | 上海斐讯数据通信技术有限公司 | 一种视频审核的方法及系统 |
| CN106791517A (zh) * | 2016-11-21 | 2017-05-31 | 广州爱九游信息技术有限公司 | 直播视频检测方法、装置及服务端 |
| CN106686395B (zh) * | 2016-12-29 | 2019-12-13 | 北京奇艺世纪科技有限公司 | 一种直播非法视频的检测方法及系统 |
| CN107622466A (zh) * | 2017-07-31 | 2018-01-23 | 上海与德科技有限公司 | 信息审核系统及方法 |
-
2018
- 2018-04-20 CN CN201810361243.6A patent/CN110392271A/zh active Pending
-
2019
- 2019-02-15 WO PCT/CN2019/075141 patent/WO2019201008A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103902954A (zh) * | 2012-12-26 | 2014-07-02 | 中国移动通信集团贵州有限公司 | 一种不良视频的鉴别方法和系统 |
| CN103544498A (zh) * | 2013-09-25 | 2014-01-29 | 华中科技大学 | 基于自适应抽样的视频内容检测方法与系统 |
| CN104268446A (zh) * | 2014-09-30 | 2015-01-07 | 小米科技有限责任公司 | 防止视频二次传播的方法及装置 |
| CN104540024A (zh) * | 2014-12-18 | 2015-04-22 | 网宿科技股份有限公司 | 视频终端及其限制视频播放的方法和系统 |
| CN105956550A (zh) * | 2016-04-29 | 2016-09-21 | 浪潮电子信息产业股份有限公司 | 一种视频鉴别的方法和装置 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111222397A (zh) * | 2019-10-25 | 2020-06-02 | 深圳市优必选科技股份有限公司 | 一种绘本识别方法、装置及机器人 |
| CN111222397B (zh) * | 2019-10-25 | 2023-10-13 | 深圳市优必选科技股份有限公司 | 一种绘本识别方法、装置及机器人 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110392271A (zh) | 2019-10-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10997788B2 (en) | Context-aware tagging for augmented reality environments | |
| WO2021159609A1 (zh) | 一种视频卡顿识别方法、装置及终端设备 | |
| US20220279223A1 (en) | Video object tagging based on machine learning | |
| US12380362B2 (en) | Continuous learning models across edge hierarchies | |
| WO2021120866A1 (zh) | 目标物体跟踪方法、装置及终端设备 | |
| CN103080951A (zh) | 用于识别媒体内容中的对象的方法和装置 | |
| CN110097004B (zh) | 面部表情识别方法和装置 | |
| WO2019201008A1 (zh) | 直播视频审查的方法和装置 | |
| US10432853B2 (en) | Image processing for automatic detection of focus area | |
| US20250104423A1 (en) | Video processing method and device | |
| CN113747245A (zh) | 多媒体资源上传方法、装置、电子设备以及可读存储介质 | |
| CN106940880A (zh) | 一种美颜处理方法、装置和终端设备 | |
| CN117978650A (zh) | 数据处理方法、装置、终端及网络侧设备 | |
| CN113194281B (zh) | 视频解析方法、装置、计算机设备和存储介质 | |
| EP2564334A1 (en) | Methods and apparatuses for facilitating remote data processing | |
| US11699463B1 (en) | Video processing method, electronic device, and non-transitory computer-readable storage medium | |
| CN114978585B (zh) | 基于流量特征的深度学习对称加密协议识别方法 | |
| CN112906551B (zh) | 视频处理方法、装置、存储介质及电子设备 | |
| US11080094B2 (en) | Method, apparatus, and electronic device for improving parallel performance of CPU | |
| CN111445499B (zh) | 用于识别目标信息的方法及装置 | |
| CN117114073A (zh) | 数据处理方法、装置、设备及介质 | |
| CN113129360B (zh) | 视频内对象的定位方法、装置、可读介质及电子设备 | |
| Hasper et al. | Remote execution vs. simplification for mobile real-time computer vision | |
| CN115883768B (zh) | 屏幕共享方法、装置、计算机设备和存储介质 | |
| CN116310966A (zh) | 视频动作定位模型训练方法、视频动作定位方法和系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19788703 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19788703 Country of ref document: EP Kind code of ref document: A1 |