WO2017161756A1 - 视频鉴别方法及系统 - Google Patents
视频鉴别方法及系统 Download PDFInfo
- Publication number
- WO2017161756A1 WO2017161756A1 PCT/CN2016/088889 CN2016088889W WO2017161756A1 WO 2017161756 A1 WO2017161756 A1 WO 2017161756A1 CN 2016088889 W CN2016088889 W CN 2016088889W WO 2017161756 A1 WO2017161756 A1 WO 2017161756A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- images
- video
- identified
- key
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/192—Recognition using electronic means using simultaneous comparisons or correlations of the image signals with a plurality of references
- G06V30/194—References adjustable by an adaptive method, e.g. learning
Definitions
- the embodiments of the present invention relate to the field of information security technologies, and in particular, to a video authentication method and system.
- people can use computers to replace humans to complete some visual recognition tasks.
- people can use the computer's monitoring system to complete intelligent monitoring, and also use the computer to complete the identification and review of video content.
- the use of computers instead of humans to complete video authentication and auditing requires the creation of complex computational models for large-volume data operations.
- the inventors found that if the calculation model created is poor and the operation error is accumulated, the computer may recognize the error or the recognition is slow, and the requirements for accuracy and timeliness cannot be met.
- the embodiment of the invention provides a video identification method and system, which are used to solve the problems of low recognition accuracy, fault tolerance and generalization ability in the prior art.
- An embodiment of the present invention provides a video authentication method, where the method includes:
- Pre-processing a plurality of images of a known type, the pre-processing including at least data augmentation;
- the plurality of to-be-identified images are identified using an optimized authentication model in the convolutional neural network.
- An embodiment of the present invention provides a video authentication system, where the system includes:
- An image preprocessing unit configured to preprocess a plurality of images of a known type, the preprocessing including at least data augmentation;
- An image discriminating training unit configured to input the preprocessed image into the convolutional neural network, perform type identification training using the discriminant model, and optimize the discriminant model according to the type discriminating result and the known type;
- An image acquisition unit to be identified configured to acquire a plurality of images to be identified
- an image discriminating unit configured to identify the plurality of to-be-identified images by using an optimized discriminant model in the convolutional neural network.
- the convolutional neural network Since the convolutional neural network has its own learning function, with the enhancement of its generalization ability, the accuracy of the target recognition and classification using the deep neural network will be continuously enhanced. Therefore, the present invention can convolve.
- neural network can enhance the generalization ability of the model of convolutional neural network through augmented image discrimination training. Convolutional neural networks and their models are simpler and more efficient to use than traditional complex computational recognition models. At the same time, by using the above-mentioned optimized convolutional neural network for video identification, the accuracy and speed of video authentication are improved.
- FIG. 1 is a flow chart of a video authentication method according to an embodiment of the present invention.
- FIG. 2 is a flow chart of acquiring a plurality of images to be identified according to an embodiment of the present invention
- 3(a) is a schematic diagram showing image rotation, cropping, and enlargement processing in a data augmentation process according to an embodiment of the present invention
- 3(b) is a schematic diagram of augmenting an image into eight images in accordance with one embodiment of the present invention.
- FIG. 4 is a flow chart of generating a low brightness image in accordance with one embodiment of the present invention.
- FIG. 5 is a flowchart of acquiring a plurality of images to be identified according to an embodiment of the present invention
- FIG. 6 is a schematic structural diagram of a video authentication system according to an embodiment of the present invention.
- FIG. 7 is a structural diagram of an image generation unit to be identified according to an embodiment of the present invention.
- FIG. 8 is a schematic structural diagram of a user equipment according to an embodiment of the present application.
- the video authentication method may include the following steps:
- Step 11 The video discriminating device performs preprocessing on a plurality of images of a known type, the preprocessing including at least data augmentation;
- Step 12 The video discriminating device inputs the pre-processed multiple images into the convolutional neural network to perform type identification training using the discriminant model, and the discriminating model may be optimized according to the type discriminating result and the known type;
- Step 13 The video discriminating device acquires an image to be identified, wherein the number of images to be identified may be determined as one or more frames according to actual conditions;
- Step 14 The video discriminating device utilizes the optimized discriminant model pair in the convolutional neural network
- the image to be identified is authenticated.
- This embodiment can be used to identify illegal video content such as redundancy, duplication, infringement of intellectual property rights, blood, violence, terror or obscenity.
- the convolutional neural network Since the convolutional neural network has its own learning function, with the enhancement of its generalization ability, the accuracy of the target recognition and classification using the deep neural network will be continuously enhanced. Therefore, the present invention can convolve.
- neural network can enhance the generalization ability of the model of convolutional neural network through augmented image discrimination training. Convolutional neural networks and their models are simpler and more efficient to use than traditional complex computational recognition models. At the same time, by using the optimized convolutional neural network for video authentication, the accuracy and speed of video authentication can be improved.
- acquiring the image to be identified may include:
- Step 131 The video discriminating device extracts a first number of key image frames in the to-be-identified video
- Step 132 The video discriminating device compares the first quantity (eg, X1) with a set threshold (eg, Y) to determine a second number (eg, X2) of key image frames;
- Step 133 The video discriminating device decodes the second number of key image frames to generate a series of images.
- Step 134 The video discriminating device performs normalization processing based on the series of images to generate a plurality of images to be identified.
- the embodiment of the present invention decodes the video image frame that satisfies the condition by extracting a certain number of key image frames of the video and setting a threshold value for the number of key image frames. Subsequent identification.
- the embodiment of the invention can reduce the number of image frames, reduce the amount of data calculation, reduce the data operation time, reduce the computing load of the processor, and reduce the hardware configuration cost under the premise of ensuring image frame quality (key frame). Low devices can also take on the task of video authentication.
- the video authentication method can include:
- Step 11' the video discriminating device acquires an image to be authenticated
- Step 12' The video discriminating device batch inputs the preprocessed image into the convolutional neural network
- the identification model is used for identification, and the authentication model is updated according to the identification result
- Step 13' The video discriminating device performs the identification of the next round of to-be-identified video by using the updated authentication model.
- obtaining the image to be authenticated may include:
- Step 111' the video discriminating device extracts a first number of key image frames in the video to be authenticated
- Step 112' the video discriminating device compares the first quantity with a set threshold to determine a second number of key image frames
- Step 113' the video discriminating device decodes the second number of key image frames to generate a series of images
- Step 114' The video discriminating device performs pre-processing on the series of images, and the pre-processing may include data augmentation and image-by-image mean reduction.
- the embodiment can make the authentication model continuously self-learning and self-updating, and can further improve the identification accuracy in the later stage.
- the discriminant model can be trained to increase the accuracy of picture recognition.
- This embodiment can perform effective data augmentation on pictures.
- Data augmentation includes, for example, rotation, random cropping, scaling, or color blurring.
- the applicant has found through a large number of experiments that the equi-angle rotation is more powerful and accurate than the horizontal and vertical flip.
- the image 1 of the vertically upward arrow is taken as an example in FIGS. 3(a) and 3(b), and an implementation of data augmentation of the image is specifically described.
- the image 1 is first obtained by rotating the image 1 (the size of which matches the size of the display frequency) by 45 degrees clockwise. Obviously, the size of image a no longer matches the display. In order to unify the size of the image a to match the size of the display screen and to maximize the integrity of the information, the present embodiment can crop the image b in the image a and then enlarge the image b into the image 2.
- the present embodiment can augment the image 1 into the image 2 by rotating, cropping, and enlarging processing, and can effectively save the effective information (usually the information in the middle of the screen is valid information, such as vertical Up arrow).
- FIG. 2 can be rotated clockwise by 45 degrees, and then subjected to the above-described cropping and enlargement processing, whereby the image 2 can be augmented into the image 3.
- the original key image (image 1) can be rotated 45 degrees counterclockwise or clockwise using the method of equal angle rotation, cropping and zooming. After completing the 360 degree rotation process for one week, the image 2 can be respectively obtained. Image 3, Image 4, Image 5, Image 6, Image 7, and Image 8. At this point, an original image can be turned into eight images, which greatly increases the amount of data in the image, enhances the generalization ability of the model, and thus improves the accuracy of the convolutional neural network training model.
- This embodiment can train the convolutional neural network model to enhance its generalization ability and robustness. Reusing the model after training to batch image recognition can improve the accuracy of video authentication and speed up video authentication.
- the training of the convolutional neural network model may adopt a data augmentation method (the data augmentation operation may also be completed before training).
- Data augmentation methods can include equiangular rotation, cropping, and scaling.
- the amount of data augmented can be increased by reducing the rotation angle. For example, you can adjust the angle from 45 degrees to 10 degrees, so that an original image can only be augmented to 8 images, but now it can be expanded to 36 images, which can increase the amount of data and increase the convolutional nerve.
- the generalization ability of the network training model can improve the accuracy of the image recognition in the later stage, but this will increase the amount of data calculation, resulting in an increase in training time.
- the amount of data that is augmented can be reduced by increasing the angle of rotation. For example, adjust the angle from 45 degrees to 90 degrees, so that an original image can be augmented to 8 images before, but now it can only be expanded to 4 images, although the training speed will be improved, but It will affect the generalization ability of the convolutional neural network training model, which will affect the accuracy of post-video identification.
- the training time is proved when the angle of rotation is 45 degrees.
- the accuracy of video identification can achieve a relatively balanced optimization effect.
- FIG. 4 is a flow diagram of generating a low brightness image in accordance with one embodiment of the present invention.
- the data augmentation may also include image brightness processing.
- the requirement for authenticating whether the video contains pornographic content may be in a training sample (ie, a known type of image, such as for pornographic content, the training sample is an erotic image).
- a training sample ie, a known type of image, such as for pornographic content, the training sample is an erotic image.
- a sample image with a lower brightness can be generated by reducing the brightness of the image by copying the existing sample.
- image brightness processing includes:
- 10 images can be rotated into a 45-degree angle to form 80 images.
- the gray values ga(1), ga(2), ... ga(80) of the first to eighth images are counted.
- Step 42 The video discriminating device determines the gray level mean ga of the plurality of images according to the gray value of each image pixel of the plurality of images.
- Step 43 The video discriminating device compares each of the grayscale values with the grayscale mean value, and when there is a certain grayscale value greater than the grayscale mean value, the image corresponding to the certain grayscale value , produces a lower brightness image copy.
- the calculation formula for determining the gray mean ga of all images can be as follows:
- Ri, Gi, and Bi are the values of the current sample images r, g, and b, respectively.
- Ri, Gi, and Bi are two-dimensional matrices, and their sizes correspond to the length and width of the image. Each element of the matrix needs to be processed separately, that is, each pixel of the image is processed.
- the image transformation formula is as follows:
- the sample with low brightness corresponding to the image sample with higher brightness can be added, which enriches the total number of samples on the one hand, and increases the generalization ability and robustness of the final model of the convolutional neural network on the other hand. Improve the accuracy of late video identification.
- the gray mean value can also be determined according to the gray value of all image pixel points, and then the gray mean value of all the images is counted, and then the gray mean value of each image is calculated, and the object of the present invention can also be achieved.
- the calculation time is longer than the above.
- the pre-processing further includes image-by-image mean reduction (eg, subtracting values of R, G, and B of the image) or further pre-processing the image using color jitter.
- image-by-image mean reduction eg, subtracting values of R, G, and B of the image
- color jitter e.g., subtracting values of R, G, and B of the image
- the step of extracting the first number of key image frames in the to-be-identified video (ie, step 131 in FIG. 2) of the embodiment may include the following steps:
- Step 1311 The video discriminating device extracts a plurality of image frames of the video to be authenticated.
- Step 1312 The video discriminating device filters out the first number of key image frames from the plurality of image frames.
- the video in this embodiment may be composed of a series of image frames. If the video frame rate is 25fps, then there are 25 pictures per second. If the video is very long, the video will contain a huge number of image frames.
- the first number of key image frames (including complete and clear image information) are filtered out from the plurality of image frames of the extracted video to be identified, so that the selected key image frames are not only suitable for detection. Tasks can also improve detection accuracy, reduce detection time, and facilitate subsequent image identification processing.
- the video of some full I frames is prevented from having excessive key frame effects.
- the detection speed this embodiment can limit the maximum number of key frames.
- embodiments of the present invention refer to a large amount of experimental data (e.g., discrimination speed and discrimination time), and the threshold Y is preferably 5000.
- X1 in the embodiment takes a value of 1000, and X1 ⁇ Y at this time, indicating that X1 does not exceed the threshold range, then the value of X2 is also taken as 1000.
- 1000 of the extracted video to be authenticated may be used.
- the key image frames of the picture are all decoded.
- X1 is 20000, when X1>Y, it means that X1 has exceeded the threshold range, which will affect the speed of video review. Therefore, it is determined that X2 is one-N of X1 such that the second number is less than or equal to the threshold, wherein N is an integer greater than or equal to two.
- N is specific, it can be customized according to the precision or time requirement of the operation. For example, when N is 10, only 2000 image frames in the 20,000 key image frames in the authentication video need to be decoded.
- the threshold value can be set to control the number of key image frames that need to be decoded, and the discrimination speed due to the increase in the number of samples is prevented under the premise of extracting as many samples as possible (key image frames). problem.
- the threshold can be set large enough to improve the video authentication accuracy.
- the normalization process can include performing image-by-image mean subtraction on the series of images.
- the speed of video detection can be improved by buffering the decoded images and then performing parallel detection on the images of the batch.
- batch_size a certain number of key frames of the video
- the key frames are sent to the convolutional neural network model for detection.
- multi-threading prepares the next batch of key frames in parallel, which can save a lot of time.
- the last batch of keyframes is insufficient (that is, when the last batch of keyframes is less than batch_size)
- the insufficient portion can be filled with a pure black image.
- the video authentication system may include: an image preprocessing unit, and image authentication training. Unit, image acquisition unit to be identified, and image authentication unit. among them:
- the image pre-processing unit is configured to pre-process a plurality of images of a known type, the pre-processing including at least data augmentation.
- the image discriminating training unit is configured to input the image preprocessed by the image preprocessing unit into the convolutional neural network for type identification training using the discriminant model, and optimize the discriminant model according to the type discriminating result and the known type.
- the image acquisition unit to be identified is used to acquire a plurality of images to be identified.
- the image discriminating unit is configured to identify the plurality of to-be-identified images acquired by the to-be-identified image acquiring unit by using the optimized discriminating model in the convolutional neural network.
- the to-be-identified image acquisition unit may include: a key image frame extraction module, a key image frame determination module, an image decoding module, and an image recognition module to be identified. among them:
- the key image frame extraction module is configured to extract a first number of key image frames in the video to be authenticated.
- the key image frame determination module is configured to compare the first quantity with a set threshold to determine a second number of key image frames.
- the image decoding module is configured to decode the second number of key image frames to generate a series of images.
- the image to be identified module is configured to perform normalization processing based on the series of images to generate an image to be identified.
- the data augmentation can include at least: an equiangular rotation, preferably, the equiangular angle is 45 degrees.
- the data augmentation further includes image brightness processing, the image brightness processing comprising:
- the pre-processing may also include image-by-image mean subtraction.
- the key image frame extraction unit is configured to extract a plurality of image frames of the video to be authenticated; and filter the first number of key image frames from the plurality of image frames.
- the key image frame determination unit is for:
- the determining module determines that the second quantity is the first quantity
- the determining module determines that the second quantity is one of N of the first quantity, such that the second quantity is less than or equal to the threshold, Where N is an integer greater than or equal to two.
- the normalization process can include image-by-image mean subtraction.
- the system or device described above may be a server or a server cluster, and the corresponding units may also be related processing units in one server or one or more servers in the server cluster.
- the related unit is one or more servers in the server cluster
- the interaction between the corresponding units is represented as an interaction between the servers, and the present invention is not limited in this respect.
- a related function module can be implemented by a hardware processor.
- the present invention also provides a non-transitory computer readable storage medium having stored therein one or more programs including execution instructions executable by an electronic device with a control interface, For performing the relevant steps in the above method embodiments, for example:
- Pre-processing a plurality of images of a known type, the pre-processing including at least data augmentation;
- the plurality of to-be-identified images are identified using an optimized authentication model in the convolutional neural network.
- FIG. 8 is a schematic structural diagram of a user equipment 800 according to an embodiment of the present disclosure.
- the specific implementation of the user equipment 800 is not limited in the specific embodiment of the present application.
- the user equipment 800 can include:
- a processor 810 a communications interface 820, a memory 830, and a communication bus 840. among them:
- the processor 810, the communication interface 820, and the memory 830 complete communication with each other via the communication bus 840.
- the communication interface 820 is configured to communicate with a network element such as a client.
- the processor 810 is configured to execute the program 832 in the memory 830 to perform the related steps in the foregoing method embodiments.
- program 832 can include program code, the program code including computer operating instructions.
- the processor 810 may be a central processing unit CPU, or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
- CPU central processing unit
- ASIC Application Specific Integrated Circuit
- the memory 830 is configured to store the program 832.
- the memory 830 may include a high speed RAM memory and may also include a non-volatile memory such as at least one disk memory.
- the program 832 can be specifically configured to cause the user device 800 to perform the following operations:
- the image discriminating training step inputs the pre-processed plurality of images into the convolutional neural network, performs type identification training using the discriminant model, and optimizes the discriminant model according to the type discriminating result and the known type;
- the image acquisition step to be identified, the image to be identified is obtained;
- the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, ie may be located A place, or it can be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of the embodiment. Those of ordinary skill in the art can understand and implement without deliberate labor.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Multimedia (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Evolutionary Computation (AREA)
- Evolutionary Biology (AREA)
- General Engineering & Computer Science (AREA)
- Bioinformatics & Computational Biology (AREA)
- Artificial Intelligence (AREA)
- Databases & Information Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Image Analysis (AREA)
Abstract
一种视频鉴别方法,包括:对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广(11);将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型(12);获取待鉴别图像(13);利用所述卷积神经网络中的优化后的鉴别模型对所述多幅待鉴别图像进行鉴别(14)。还提供一种视频鉴别系统。通过增广的鉴别训练,能增强卷积神经网络的模型的泛化能力;通过将视频处理成图像,利用卷积神经网络进行鉴别,提高了视频鉴别的准确率和速度。
Description
本发明实施例涉及信息安全技术领域,尤其涉及一种视频鉴别方法及系统。
随着计算机硬件及互联网大数据的快速发展,互联网中视频数量呈现出爆炸式增长的态势。这其中存在着大量冗余、重复的视频内容,除此之外还有一些涉及侵犯知识产权、血腥、暴力、恐怖或淫秽等的非法视频内容。
目前,人们可以利用计算机替代人类完成一些视觉识别任务。例如人们可以利用计算机的监控系统完成智能监视,还可以利用计算机完成视频内容的识别与审核等。通常,利用计算机替代人类完成视频鉴别和审核需要创建复杂的计算模型,进行大批量数据的运算。发明人在实现本发明的过程中发现如果创建的计算模型不佳和运算误差累积,可能会导致计算机识别错误或者识别缓慢,无法满足人们对精确度和及时性的要求。
发明内容
本发明实施例提供一种视频鉴别方法及系统,用以解决现有技术中识别准确度低,容错能力和泛化能力差等问题。
本发明实施例提供一种视频鉴别方法,该方法包括:
对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;
将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴
别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;
获取多幅待鉴别图像;
利用所述卷积神经网络中的优化后的鉴别模型对所述多幅待鉴别图像进行鉴别。
本发明实施例提供一种视频鉴别系统,该系统包括:
图像预处理单元,用于对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;
图像鉴别训练单元,用于将预处理后的图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;
待鉴别图像获取单元,用于获取多幅待鉴别图像;
图像鉴别单元,用于利用所述卷积神经网络中的优化后的鉴别模型对所述多幅待鉴别图像进行鉴别。
由于卷积神经网络拥有自己学习的功能,随着其泛化能力的增强,利用深层次的神经网络进行目标的识别与分类的精度也会随之不断的增强,因此,本发明可以将卷积神经网络作为识别的主要工具,通过增广的图像鉴别训练,可以增强卷积神经网络的模型的泛化能力。相比于传统的复杂的计算识别模型来说,卷积神经网络及其模型运用起来更加简单高效。同时通过利用上述优化后的卷积神经网络进行视频鉴别,提高了视频鉴别的准确率和速度。
为了更清楚地说明本发明实施例的技术方案,下面将对实施例描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为根据本发明一个实施例的视频鉴别方法流程图;
图2为根据本发明一个实施例的获取多幅待鉴别图像流程图;
图3(a)为根据本发明一个实施例的数据增广过程中图像旋转45度、裁剪和放大处理的示意图;
图3(b)为根据本发明一个实施例的将一幅图像增广为八幅图像的示意图;
图4为根据本发明一个实施例的生成低亮度图像的流程图;
图5为根据本发明一个实施例的获取多幅待鉴别图像的流程图;
图6为根据本发明一个实施例的视频鉴别系统结构示意图;
图7为根据本发明一个实施例的待鉴别图像生成单元的结构图;
图8为根据本申请实施例提供的一种用户设备的结构示意图。
为使本发明实施例的目的、技术方案和优点更加清楚,下面将结合本发明实施例中的附图,对本发明实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的范围。
如图1所示,视频鉴别方法可以包括如下步骤:
步骤11:视频鉴别装置对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;
步骤12:视频鉴别装置将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,可以根据类型鉴别结果和所述已知类型优化所述鉴别模型;
步骤13:视频鉴别装置获取待鉴别图像,其中,待鉴别图像的幅数可以按实际情况确定为一幅或者多幅;
步骤14:视频鉴别装置利用所述卷积神经网络中的优化后的鉴别模型对
所述待鉴别图像进行鉴别。
本实施例可以用于鉴别冗余、重复、侵犯知识产权、血腥、暴力、恐怖或淫秽等非法视频内容。
由于卷积神经网络拥有自己学习的功能,随着其泛化能力的增强,利用深层次的神经网络进行目标的识别与分类的精度也会随之不断的增强,因此,本发明可以将卷积神经网络作为识别的主要工具,通过增广的图像鉴别训练,可以增强了卷积神经网络的模型的泛化能力。相比于传统的复杂的计算识别模型来说,卷积神经网络及其模型运用起来更加简单高效。同时通过利用上述优化后的卷积神经网络进行视频鉴别,可以提高视频鉴别的准确率和速度。
如图2所示,获取待鉴别图像(即图1中步骤13)可以包括:
步骤131:视频鉴别装置提取待鉴别视频中的第一数量的关键图像帧;
步骤132:视频鉴别装置将所述第一数量(例如X1)与设定的阈值(例如Y)进行比较,确定第二数量(例如X2)的关键图像帧;
步骤133:视频鉴别装置对所述第二数量的关键图像帧进行解码,生成一系列图像;
步骤134:视频鉴别装置基于所述一系列图像进行归一化处理以生成多幅待鉴别图像。
为了使卷积神经网络能够应对视频的鉴别任务,本发明实施例通过提取视频的一定数量的关键图像帧,通过对关键图像帧的数量设定阈值,对满足条件的视频图像帧才进行解码和后续的鉴别。本发明实施例在保证了图像帧质量(关键帧)的前提下,可以减少图像帧的数量,降低的数据运算量,减少了数据运算时间,降低了处理器的运算负荷,使得硬件配置成本较低的设备也能承担视频鉴别的任务。
在一些实施例中,视频鉴别方法可以包括:
步骤11’:视频鉴别装置获取待鉴别图像;
步骤12’:视频鉴别装置将预处理后的图像批量输入到卷积神经网络中
利用鉴别模型进行鉴别,并根据鉴别结果更新鉴别模型;
步骤13’:视频鉴别装置利用更新后的鉴别模型进行下一轮待鉴别视频的鉴别。
在一些实施例中,获取待鉴别图像(步骤11’)可以包括:
步骤111’:视频鉴别装置提取待鉴别视频中的第一数量的关键图像帧;
步骤112’:视频鉴别装置将所述第一数量与设定的阈值进行比较,确定第二数量的关键图像帧;
步骤113’:视频鉴别装置对所述第二数量的关键图像帧进行解码,生成一系列图像;
步骤114’:视频鉴别装置对所述一系列图像进行预处理,所述预处理可以包括数据增广和逐图像均值消减。
由此,本实施例可以使鉴别模型不断地自我学习、自我更新,可以进一步提高后期的鉴别准确度。
为了提高卷积神经网络中鉴别模型的泛化能力,可以对鉴别模型进行训练,从而可以增加图片识别的精度。本实施例可以对图片进行有效的数据增广。数据增广例如包括旋转、随机裁剪、缩放或颜色抖动等。另外,申请人经过大量的实验发现,等角度旋转相比于水平及垂直方向上的翻转而言,泛化能力和准确性更强。
为了能够形象的体现图像的方向,下面在图3(a)和图3(b)中以竖直向上的箭头的图像1为例,具体说明对该图像进行数据增广的实现方式。
如图3(a)所示,首先将图像1(其尺寸与显示频的尺寸相匹配)顺时针旋转45度得到图像a。显然,图像a的尺寸与显示屏不再匹配。为了将图像a的尺寸统一为与显示屏相匹配的尺寸,并最大限度的保护信息的完整性,本实施例可以在图像a中裁剪出图像b,然后,再将图像b放大为图像2。
由此,本实施例可以通过旋转、裁剪和放大处理,将图像1增广为图像2,且可以有效的保存有效信息(通常屏幕中间的信息是有效信息,例如竖直
向上的箭头)。
同理,如图3(b)所示,可以将图2顺时针旋转45度,再通过上述的裁剪和放大处理后,从而可以将图像2增广为图像3。当然也可以直接将图像1顺时针旋转90度后,再通过裁剪和放大处理,将图像1增广为图像3。
本实施例可以使用等角度旋转、裁剪和缩放的方式,对原始关键图像(图像1)逆时针或者顺时针每次旋转45度,在完成一周360度的旋转处理后,可以分别得到图像2、图像3、图像4、图像5、图像6、图像7和图像8。此时,一张原始图像,就可以变成了八张图像,大幅度增加了图像的数据量,增强了模型的泛化能力,进而可以提高卷积神经网络训练模型的准确度。
本实施例可以对卷积神经网络模型进行训练,以增强其泛化能力和鲁棒性。再利用训练之后的模型对图像进行批量识别,可以提高视频鉴别的准确率,并能加快视频鉴别速度。
在本实施例中,对卷积神经网络模型进行训练可以采用数据增广的方式(该数据增广的操作也可以在训练之前完成)。数据增广方式可以包括等角度旋转、裁剪和缩放等。
其中,为了进一步提高卷积神经网络训练模型的泛化能力,可以通过减小旋转角度来增加增广的数据量。例如,可以将角度由45度调整为10度,这样一张原始图像,之前只能增广为8幅图像,而现在却能增广为36幅图像,这样可以增加数据量,提高卷积神经网络训练模型的泛化能力,进而可以提高后期图像识别的精度,但是这样做会增加数据运算量,导致训练的时间增长。
同理,可以通过增大旋转角度来减少增广的数据量。例如,将角度由45度调整为90度,这样一张原始图像,之前可以增广为8幅图像,而现在却只能增广为4幅图像,虽然训练的速度会有所提高,但这会影响卷积神经网络训练模型的泛化能力,进而会影响后期视频鉴别的准确度。
由此,经过大量的试验数据证明,当旋转的角度为45度时,训练的时间
和视频鉴别的精度可以达到相对平衡的优化效果。
图4为根据本发明一个实施例的生成低亮度图像的流程图。数据增广还可以包括图像亮度处理,本实施例中,针对需要鉴别视频中是否包含色情内容的要求,可以在训练样本(即已知类型的图像,例如针对色情内容,训练样本就是色情图片)中,人为增加一些亮度较低的样本图像(由于关于色情的视频内容通常在昏暗的环境下,所以图像的亮度较低)。亮度较低的样本图像可以通过现有样本的副本通过降低图像的亮度处理生成。如图4所示,图像亮度处理包括:
步骤41:视频鉴别装置获取多幅图像的每幅图像像素的灰度值ga(i),(i=1、2、3…n)。
例如,10幅图像通过45度等角度旋转后可以形成80幅图像。统计第1至80幅图像的灰度值ga(1)、ga(2)……ga(80)。
步骤42:视频鉴别装置根据多幅图像的每幅图像像素的灰度值,确定多幅图像的灰度均值ga。
步骤43:视频鉴别装置将所述各个灰度值分别与所述灰度均值进行比较,当存在某一灰度值大于所述灰度均值时,针对所述某一灰度值所对应的图像,生成亮度较低的图像副本。
具体的,确定所有图像(例如80幅)的灰度均值ga的计算公式可以如下:
其中,n是样本图像总数,Ri、Gi、Bi分别为当前样本图像r、g、b分量值。其中Ri、Gi、Bi为二维矩阵,其大小分别对应图像的长和宽。需要分别对矩阵的每个元素进行处理,即对图像的每个像素点进行处理。
在本实施例中,图像变换公式如下所示:
经过上述处理后,可以增加与亮度较高的图像样本对应的亮度低的样本,一方面丰富了样本总数,另一方面也增加了卷积神经网络最终模型的泛化能力和鲁棒性,可以提高后期的视频鉴别的准确度。
当然,上述方法中,还可以根据所有图像像素点的灰度值来确定灰度均值,之后统计所有图像的灰度均值,然后再计算各个图像的灰度均值,也可以达到本发明的目的,只是,运算时间比上述方式要长。
在一些实施例中,预处理还包括:逐图像均值消减(例如对图像的R、G和B的数值进行消减)或者利用颜色抖动(color jitter)的方法对图像做进一步预处理。这样做便于数据加工和处理(可以是归一化数据处理方式),加快视频鉴别的速度。
如图5所示,本实施例的提取待鉴别视频中的第一数量的关键图像帧的步骤(即图2中步骤131)可以包括如下步骤:
步骤1311:视频鉴别装置提取待鉴别视频的多幅图像帧。
步骤1312:视频鉴别装置从多幅图像帧中筛选出第一数量的关键图像帧。
本实施例中的视频可以是由一系列图像帧组成。如果视频帧率为25fps,那么每秒钟的视频就有25张图片。如果视频时长很长,那么该视频包含的图像帧的数量就会非常巨大。本实施例通过从提取的待鉴别视频的多幅图像帧中筛选出第一数量的关键图像帧(包含完整、清晰的图像信息),使得筛选出的关键图像帧不仅能够很好的适用于检测任务,还可以提高检测准确度,减少检测时间,而且便于后续的图像鉴别处理。
具体的,在一些实施例中,为了控制关键帧数量,防止一些全I帧(MPEG编码中的内部编码帧,代表一个完整的画面)的视频含有过多的关键帧影响
检测速度,本实施例可以限制最大关键帧数量。为了提高视频鉴别的精度和减少鉴别的时间,本发明实施例参考大量的实验数据(例如鉴别速度和鉴别时间),阈值Y优选为5000。
具体的,如果本实施例中的X1取值1000时,此时X1≤Y,说明X1没有超出阈值范围,那么,X2的值也取1000,此时,可以对提取的待鉴别视频中的1000张的关键图像帧进行全部解码。
如果X1取值为20000时,此时X1>Y时,说明X1已经超出阈值范围,这会影响视频审核的速度。因此,确定X2为X1的N分之一,以使所述第二数量小于或者等于所述阈值,其中,N为大于或者等于二的整数。N在具体取值时,可以按运算的精度或者时间要求,进行自定义。例如,N取10时,只需要对待鉴别视频中的20000张的关键图像帧中的2000张图像帧进行解码。
由此,本实施例可以通过设定阈值,对需要解码的关键图像帧进行数量控制,在尽量多的提取样本(关键图像帧)的前提下,防止因样本数量增多带来的鉴别速度下降的问题。当然,如果硬件配置较高,处理器运算速度较快时,可以将阈值设置得足够大,以提高视频鉴别准确度。
在一些实施例中,归一化处理可以包括对所述一系列图像进行逐图像均值消减。
在一些实施例中,可以通过将解码的图像进行缓存,然后可以对批量的图像进行并行检测,从而可以提高视频检测的速度。
具体的,在进行批量检测时,首先提取视频一定数量(batch_size)的关键帧、之后将这批关键帧送入卷积神经网络模型进行检测。在检测的同时,多线程并行地准备下一批关键帧,这样可以大幅度节省时间。此外,当最后一批关键帧数量不足时(即最后一批关键帧数量小于batch_size时),不足的部分可以用纯黑色图像补齐。
如图6所示,视频鉴别系统可以包括:图像预处理单元、图像鉴别训练
单元、待鉴别图像获取单元和图像鉴别单元。其中:
图像预处理单元用于对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广。
图像鉴别训练单元用于将图像预处理单元预处理后的图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型。
待鉴别图像获取单元用于获取多幅待鉴别图像。
图像鉴别单元用于利用所述卷积神经网络中的优化后的鉴别模型对所述待鉴别图像获取单元获取的多幅待鉴别图像进行鉴别。
在一些实施例中,所述待鉴别图像获取单元可以包括:关键图像帧提取模块、关键图像帧确定模块、图像解码模块和待鉴别图像生成模块。其中:
关键图像帧提取模块用于提取待鉴别视频中的第一数量的关键图像帧。
关键图像帧确定模块用于将所述第一数量与设定的阈值进行比较,确定第二数量的关键图像帧。
图像解码模块用于对所述第二数量的关键图像帧进行解码,生成一系列图像。
待鉴别图像生成模块用于基于所述一系列图像进行归一化处理以生成待鉴别图像。
在一些实施例中,所述数据增广至少可以包括:等角度旋转,较佳的,所述等角度为45度。
在一些实施例中,所述数据增广还包括图像亮度处理,所述图像亮度处理包括:
获取所述多幅图像的每幅图像像素的灰度值;
根据所述多幅图像的每幅图像像素的灰度值,确定所述多幅图像的灰度均值;
将所述各个灰度值分别与所述灰度均值进行比较,当存在某一灰度值大
于所述灰度均值时,可以针对所述某一灰度值所对应的图像,生成亮度较低的图像副本。
在一些实施例中,所述预处理还可以包括逐图像均值消减。
在一些实施例中,关键图像帧提取单元用于提取待鉴别视频的多幅图像帧;以及从所述多幅图像帧中筛选出第一数量的关键图像帧。
在一些实施例中,关键图像帧确定单元用于:
当所述比较模块判定所述第一数量小于或者等于设定的阈值时,所述确定模块确定第二数量为第一数量;以及
当所述比较模块判定所述第一数量大于设定的阈值时,所述确定模块确定第二数量为第一数量的N分之一,以使所述第二数量小于或者等于所述阈值,其中,N为大于或者等于二的整数。
在一些实施例中,归一化处理可以包括逐图像均值消减。
以上所述的系统或装置可以是一个服务器或者服务器集群,相应的各个单元也可以为一个服务器中的相关处理单元或者为服务器集群中的一个或多个服务器。当相关的单元为服务器集群中的一个或多个服务器时,相应的单元之间的交互则表现为服务器之间的交互,本发明在此方面没有限制。
由于上述实施例的视频鉴别系统与视频鉴别方法的功能相对应,在此,不再赘述视频鉴别系统与视频鉴别方法相关的内容。本发明实施例中可以通过硬件处理器(hardware processor)来实现相关功能模块。
进一步地,本发明还提供一种非瞬时的计算机可读存储介质,所述存储介质中存储有一个或多个包括执行指令的程序,所述执行指令能够被带有控制界面的电子设备执行,以用于执行上述方法实施例中的相关步骤,例如:
对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;
将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;
获取待鉴别图像;
利用所述卷积神经网络中的优化后的鉴别模型对所述多幅待鉴别图像进行鉴别。
图8为本申请实施例提供的一种用户设备800的结构示意图,本申请具体实施例并不对用户设备800的具体实现做限定。如图8所示,该用户设备800可以包括:
处理器(processor)810、通信接口(Communications Interface)820、存储器(memory)830、以及通信总线840。其中:
处理器810、通信接口820、以及存储器830通过通信总线840完成相互间的通信。
通信接口820,用于与比如客户端等的网元通信。
处理器810,用于执行存储器830中的程序832,以执行上述方法实施例中的相关步骤。
具体地,程序832可以包括程序代码,所述程序代码包括计算机操作指令。
处理器810可能是一个中央处理器CPU,或者是特定集成电路ASIC(Application Specific Integrated Circuit),或者是被配置成实施本申请实施例的一个或多个集成电路。
存储器830,用于存放程序832。存储器830可能包含高速RAM存储器,也可能还包括非易失性存储器(non-volatile memory),例如至少一个磁盘存储器。程序832具体可以用于使得所述用户设备800执行以下操作:
图像预处理步骤,对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;
图像鉴别训练步骤,将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;
待鉴别图像获取步骤,获取待鉴别图像;
图像鉴别步骤,利用所述卷积神经网络中的优化后的鉴别模型对所述待鉴别图像进行鉴别。
程序832中各步骤的具体实现可以参见上述实施例中的相应步骤和单元中对应的描述,在此不赘述。所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的设备和模块的具体工作过程,可以参考前述方法实施例中的对应过程描述,在此不再赘述。
以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。本领域普通技术人员在不付出创造性的劳动的情况下,即可以理解并实施。
通过以上的实施例的描述,本领域的技术人员可以清楚地了解到各实施例可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的理解,上述技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如ROM/RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行各个实施例或者实施例的某些部分所述的方法。
最后应说明的是:以上实施例仅用以说明本发明的技术方案,而非对其限制;尽管参照前述实施例对本发明进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的精神和范围。
Claims (18)
- 一种视频鉴别方法,包括:对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;将预处理后的多幅图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;获取待鉴别图像;利用所述卷积神经网络中的优化后的鉴别模型对所述待鉴别图像进行鉴别。
- 根据权利要求1所述的方法,其中,所述数据增广至少包括:等角度旋转。
- 根据权利要求2所述的方法,其中,所述等角度为45度。
- 根据权利要求2所述的方法,其中,所述数据增广还包括图像亮度处理,所述图像亮度处理包括:获取所述多幅图像的每幅图像像素的灰度值;根据所述多幅图像的每幅图像像素的灰度值,确定所述多幅图像的灰度均值;将所述各个灰度值分别与所述灰度均值进行比较,当存在某一灰度值大于所述灰度均值时,针对所述某一灰度值所对应的图像,生成亮度较低的图像副本。
- 根据权利要求1所述的方法,其中,所述预处理还包括:逐图像均值消减。
- 根据权利要求1-5中任一项所述的方法,其中,所述获取多幅待鉴别图像包括:提取待鉴别视频中的第一数量的关键图像帧;将所述第一数量与设定的阈值进行比较,确定第二数量的关键图像帧;对所述第二数量的关键图像帧进行解码,生成一系列图像;基于所述一系列图像进行归一化处理以生成多幅待鉴别图像。
- 根据权利要求6所述的方法,其中,所述提取待鉴别视频中的第一数量的关键图像帧包括:提取待鉴别视频的多幅图像帧;从所述多幅图像帧中筛选出第一数量的关键图像帧。
- 根据权利要求6所述的方法,其中,所述将所述第一数量与设定的阈值进行比较,确定第二数量的关键图像帧包括:当所述第一数量小于或者等于设定的阈值时,确定第二数量为第一数量;当所述第一数量大于设定的阈值时,确定第二数量为第一数量的N分之一,以使所述第二数量小于或者等于所述阈值,其中,N为大于或者等于二的整数。
- 根据权利要求6所述的方法,其中,所述归一化处理包括:逐图像均值消减。
- 一种视频鉴别系统,包括:图像预处理单元,用于对已知类型的多幅图像进行预处理,所述预处理至少包括数据增广;图像鉴别训练单元,用于将预处理后的图像输入到卷积神经网络中利用鉴别模型进行类型鉴别训练,根据类型鉴别结果和所述已知类型优化所述鉴别模型;待鉴别图像获取单元,用于获取待鉴别图像;图像鉴别单元,用于利用所述卷积神经网络中的优化后的鉴别模型对所述待鉴别图像进行鉴别。
- 根据权利要求10所述的系统,其中,所述数据增广至少包括:等角度旋转。
- 根据权利要求11所述的系统,其中,所述等角度为45度。
- 根据权利要求11所述的系统,其中,所述数据增广还包括图像亮度处理,所述图像亮度处理包括:获取所述多幅图像的每幅图像像素的灰度值;根据所述多幅图像的每幅图像像素的灰度值,确定所述多幅图像的灰度均值;将所述各个灰度值分别与所述灰度均值进行比较,当存在某一灰度值大于所述灰度均值时,针对所述某一灰度值所对应的图像,生成亮度较低的图像副本。
- 根据权利要求10所述的系统,其中,所述预处理还包括:逐图像均值消减。
- 根据权利要求10-14中任一项所述的系统,其中,所述待鉴别图像 获取单元包括:关键图像帧提取模块,用于提取待鉴别视频中的第一数量的关键图像帧;关键图像帧确定模块,用于将所述第一数量与设定的阈值进行比较,确定第二数量的关键图像帧;图像解码模块,用于对所述第二数量的关键图像帧进行解码,生成一系列图像;待鉴别图像生成模块,用于基于所述一系列图像进行归一化处理以生成待鉴别图像。
- 根据权利要求15所述的系统,其中,所述关键图像帧提取单元用于:提取待鉴别视频的多幅图像帧;以及从所述多幅图像帧中筛选出第一数量的关键图像帧。
- 根据权利要求15所述的系统,其中,所述关键图像帧确定单元用于:当所述比较模块判定所述第一数量小于或者等于设定的阈值时,所述确定模块确定第二数量为第一数量;以及当所述第一数量大于设定的阈值时,确定第二数量为第一数量的N分之一,以使所述第二数量小于或者等于所述阈值,其中,N为大于或者等于二的整数。
- 根据权利要求15所述的系统,其中,所述归一化处理包括:逐图像均值消减。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/246,166 US20170277955A1 (en) | 2016-03-23 | 2016-08-24 | Video identification method and system |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610168258.1 | 2016-03-23 | ||
| CN201610168258.1A CN105844238A (zh) | 2016-03-23 | 2016-03-23 | 视频鉴别方法及系统 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US15/246,166 Continuation US20170277955A1 (en) | 2016-03-23 | 2016-08-24 | Video identification method and system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017161756A1 true WO2017161756A1 (zh) | 2017-09-28 |
Family
ID=56583014
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/088889 Ceased WO2017161756A1 (zh) | 2016-03-23 | 2016-07-06 | 视频鉴别方法及系统 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN105844238A (zh) |
| WO (1) | WO2017161756A1 (zh) |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111091489A (zh) * | 2019-11-01 | 2020-05-01 | 平安科技(深圳)有限公司 | 图片优化方法、装置、电子设备及存储介质 |
| CN111314733A (zh) * | 2020-01-20 | 2020-06-19 | 北京百度网讯科技有限公司 | 用于评估视频清晰度的方法和装置 |
| CN112100075A (zh) * | 2020-09-24 | 2020-12-18 | 腾讯科技(深圳)有限公司 | 一种用户界面回放方法、装置、设备及存储介质 |
| CN112131907A (zh) * | 2019-06-24 | 2020-12-25 | 华为技术有限公司 | 一种训练分类模型的方法及装置 |
| EP3770851A4 (en) * | 2019-05-23 | 2021-07-21 | Wangsu Science & Technology Co., Ltd. | QUALITY INSPECTION PROCESS FOR TARGET OBJECT, AND PERIPHERAL COMPUTER DEVICE |
| CN113408589A (zh) * | 2021-05-26 | 2021-09-17 | 北京迈格威科技有限公司 | 训练图像识别模型的方法、图像识别方法和图像识别设备 |
Families Citing this family (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106682569A (zh) * | 2016-09-28 | 2017-05-17 | 天津工业大学 | 一种基于卷积神经网络的快速交通标识牌识别方法 |
| CN106504064A (zh) * | 2016-10-25 | 2017-03-15 | 清华大学 | 基于深度卷积神经网络的服装分类与搭配推荐方法及系统 |
| CN107016356A (zh) * | 2017-03-21 | 2017-08-04 | 乐蜜科技有限公司 | 特定内容识别方法、装置和电子设备 |
| CN109685756A (zh) * | 2017-10-16 | 2019-04-26 | 乐达创意科技有限公司 | 影像特征自动辨识装置、系统及方法 |
| CN107818302A (zh) * | 2017-10-20 | 2018-03-20 | 中国科学院光电技术研究所 | 基于卷积神经网络的非刚性多尺度物体检测方法 |
| CN108154134B (zh) * | 2018-01-11 | 2019-07-23 | 天格科技(杭州)有限公司 | 基于深度卷积神经网络的互联网直播色情图像检测方法 |
| CN108419091A (zh) * | 2018-03-02 | 2018-08-17 | 北京未来媒体科技股份有限公司 | 一种基于机器学习的视频内容审核方法及装置 |
| CN108810547A (zh) * | 2018-07-03 | 2018-11-13 | 电子科技大学 | 一种基于神经网络和pca-knn的高效vr视频压缩方法 |
| CN109242788A (zh) * | 2018-08-21 | 2019-01-18 | 福州大学 | 一种基于编码-解码卷积神经网络低照度图像优化方法 |
| CN109344883A (zh) * | 2018-09-13 | 2019-02-15 | 西京学院 | 一种基于空洞卷积的复杂背景下果树病虫害识别方法 |
| CN110334693A (zh) * | 2019-07-17 | 2019-10-15 | 中国电子科技集团公司第五十四研究所 | 一种面向深度学习的遥感图像目标样本生成方法 |
| CN110796098B (zh) | 2019-10-31 | 2021-07-27 | 广州市网星信息技术有限公司 | 内容审核模型的训练及审核方法、装置、设备和存储介质 |
| CN111125539B (zh) * | 2019-12-31 | 2024-02-02 | 武汉市烽视威科技有限公司 | 一种基于人工智能的cdn有害信息阻断方法及系统 |
| CN111241969A (zh) * | 2020-01-06 | 2020-06-05 | 北京三快在线科技有限公司 | 目标检测方法、装置及相应模型训练方法、装置 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103902954A (zh) * | 2012-12-26 | 2014-07-02 | 中国移动通信集团贵州有限公司 | 一种不良视频的鉴别方法和系统 |
| CN104281858A (zh) * | 2014-09-15 | 2015-01-14 | 中安消技术有限公司 | 三维卷积神经网络训练方法、视频异常事件检测方法及装置 |
| CN105303179A (zh) * | 2015-10-28 | 2016-02-03 | 小米科技有限责任公司 | 指纹识别方法、装置 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104751107B (zh) * | 2013-12-30 | 2018-08-03 | 中国移动通信集团公司 | 一种视频关键数据确定方法、装置及设备 |
-
2016
- 2016-03-23 CN CN201610168258.1A patent/CN105844238A/zh active Pending
- 2016-07-06 WO PCT/CN2016/088889 patent/WO2017161756A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103902954A (zh) * | 2012-12-26 | 2014-07-02 | 中国移动通信集团贵州有限公司 | 一种不良视频的鉴别方法和系统 |
| CN104281858A (zh) * | 2014-09-15 | 2015-01-14 | 中安消技术有限公司 | 三维卷积神经网络训练方法、视频异常事件检测方法及装置 |
| CN105303179A (zh) * | 2015-10-28 | 2016-02-03 | 小米科技有限责任公司 | 指纹识别方法、装置 |
Non-Patent Citations (1)
| Title |
|---|
| ZHAO, KAIXUAN ET AL.: "Recognition of Individual Dairy Cattle Based on Convolutional Neural Networks", TRANSACTIONS OF THE CHINESE SOCIETY OF AGRICULTURAL ENGINEERING, vol. 31, no. 5, 31 March 2015 (2015-03-31), pages 182 * |
Cited By (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3770851A4 (en) * | 2019-05-23 | 2021-07-21 | Wangsu Science & Technology Co., Ltd. | QUALITY INSPECTION PROCESS FOR TARGET OBJECT, AND PERIPHERAL COMPUTER DEVICE |
| CN112131907A (zh) * | 2019-06-24 | 2020-12-25 | 华为技术有限公司 | 一种训练分类模型的方法及装置 |
| CN111091489A (zh) * | 2019-11-01 | 2020-05-01 | 平安科技(深圳)有限公司 | 图片优化方法、装置、电子设备及存储介质 |
| CN111091489B (zh) * | 2019-11-01 | 2024-05-07 | 平安科技(深圳)有限公司 | 图片优化方法、装置、电子设备及存储介质 |
| CN111314733A (zh) * | 2020-01-20 | 2020-06-19 | 北京百度网讯科技有限公司 | 用于评估视频清晰度的方法和装置 |
| CN111314733B (zh) * | 2020-01-20 | 2022-06-10 | 北京百度网讯科技有限公司 | 用于评估视频清晰度的方法和装置 |
| CN112100075A (zh) * | 2020-09-24 | 2020-12-18 | 腾讯科技(深圳)有限公司 | 一种用户界面回放方法、装置、设备及存储介质 |
| CN112100075B (zh) * | 2020-09-24 | 2024-03-15 | 腾讯科技(深圳)有限公司 | 一种用户界面回放方法、装置、设备及存储介质 |
| CN113408589A (zh) * | 2021-05-26 | 2021-09-17 | 北京迈格威科技有限公司 | 训练图像识别模型的方法、图像识别方法和图像识别设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN105844238A (zh) | 2016-08-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017161756A1 (zh) | 视频鉴别方法及系统 | |
| JP6994588B2 (ja) | 顔特徴抽出モデル訓練方法、顔特徴抽出方法、装置、機器および記憶媒体 | |
| CN110222573B (zh) | 人脸识别方法、装置、计算机设备及存储介质 | |
| WO2021139324A1 (zh) | 图像识别方法、装置、计算机可读存储介质及电子设备 | |
| EP2806374B1 (en) | Method and system for automatic selection of one or more image processing algorithm | |
| WO2022161286A1 (zh) | 图像检测方法、模型训练方法、设备、介质及程序产品 | |
| WO2021137946A1 (en) | Forgery detection of face image | |
| US20170277955A1 (en) | Video identification method and system | |
| CN111079816A (zh) | 图像的审核方法、装置和服务器 | |
| CN114743121A (zh) | 图像处理方法、图像处理模型的训练方法及装置 | |
| CN116746155A (zh) | 端到端加水印系统 | |
| Balafrej et al. | Enhancing practicality and efficiency of deepfake detection | |
| TWI803243B (zh) | 圖像擴增方法、電腦設備及儲存介質 | |
| CN115082667A (zh) | 图像处理方法、装置、设备及存储介质 | |
| CN118506427A (zh) | 伪造人脸检测方法及装置、终端和存储介质 | |
| CN114529750A (zh) | 图像分类方法、装置、设备及存储介质 | |
| WO2024260302A1 (zh) | 活体检测模型训练方法、活体检测方法和装置 | |
| Gu et al. | Deepfake detection and localisation based on illumination inconsistency | |
| Mustafa et al. | Obscenity Detection Using Haar‐Like Features and Gentle Adaboost Classifier | |
| CN113989548B (zh) | 证件分类模型训练方法、装置、电子设备及存储介质 | |
| CN113762060B (zh) | 人脸图像检测方法、装置、可读介质及电子设备 | |
| WO2023137905A1 (zh) | 图像处理方法、装置、电子设备以及存储介质 | |
| CN114118412A (zh) | 证件识别模型训练及证件识别的方法、系统、设备及介质 | |
| CN114842478A (zh) | 文本区域的识别方法、装置、设备及存储介质 | |
| CN114463734A (zh) | 文字识别方法、装置、电子设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16895106 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16895106 Country of ref document: EP Kind code of ref document: A1 |
