WO2018233256A1 - 直播视频监控方法、存储介质、电子设备及系统 - Google Patents

直播视频监控方法、存储介质、电子设备及系统 Download PDF

Info

Publication number
WO2018233256A1
WO2018233256A1 PCT/CN2017/117973 CN2017117973W WO2018233256A1 WO 2018233256 A1 WO2018233256 A1 WO 2018233256A1 CN 2017117973 W CN2017117973 W CN 2017117973W WO 2018233256 A1 WO2018233256 A1 WO 2018233256A1
Authority
WO
WIPO (PCT)
Prior art keywords
signal
decomposition
imf
speech
live video
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2017/117973
Other languages
English (en)
French (fr)
Inventor
李振华
张文明
陈少杰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Wuhan Douyu Network Technology Co Ltd
Original Assignee
Wuhan Douyu Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Wuhan Douyu Network Technology Co Ltd filed Critical Wuhan Douyu Network Technology Co Ltd
Publication of WO2018233256A1 publication Critical patent/WO2018233256A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L63/00Network architectures or network communication protocols for network security
    • H04L63/02Network architectures or network communication protocols for network security for separating internal from external traffic, e.g. firewalls
    • H04L63/0227Filtering policies
    • H04L63/0245Filtering by information in the payload
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/18Speech classification or search using natural language modelling
    • G10L15/1822Parsing for meaning understanding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L63/00Network architectures or network communication protocols for network security
    • H04L63/02Network architectures or network communication protocols for network security for separating internal from external traffic, e.g. firewalls
    • H04L63/0227Filtering policies
    • H04L63/0236Filtering by address, protocol, port number or service, e.g. IP-address or URL
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • H04L65/61Network streaming of media packets for supporting one-way streaming services, e.g. Internet radio
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/24Monitoring of processes or resources, e.g. monitoring of server load, available bandwidth, upstream requests

Definitions

  • the present invention relates to the field of live video surveillance, and in particular, to a live video monitoring method, a storage medium, an electronic device, and a system.
  • the existing method for monitoring the live video is generally: manually performing live monitoring on the live video of all live broadcasts in the live broadcast platform.
  • the live broadcast is marked as Focus on monitoring the live broadcast room; then arrange manual monitoring of the key monitoring live broadcasts, thereby improving monitoring efficiency and purifying live content.
  • the technical problem solved by the present invention is how to automatically and accurately screen a live webcast video with bad information.
  • the invention can significantly reduce the noise of the voice data signal, thereby maximizing the perfect recognition of the live voice data, greatly improving the monitoring precision, and providing a powerful guarantee for preventing the spread of the bad information of the network.
  • the live video monitoring method provided by the present invention includes the following steps:
  • S1 The server obtains the live video address of the live broadcast, obtains the live video stream according to the live video address, parses the live video stream to obtain the voice data signal, and transfers to S2;
  • S2 The server modally decomposes the voice data through CEEMDAN to obtain an nth order IMF component, n ⁇ 3, and goes to S3;
  • S3 the server splits the nth order IMF component by FastICA to obtain m reconstructed speech signals, 1 ⁇ m ⁇ n-2, and goes to S4;
  • S4 The server identifies all reconstructed voice signals to determine whether there is a reconstructed voice signal containing bad information. If yes, the current live broadcast is abnormally marked; otherwise, the current live broadcast will be normally marked.
  • the invention provides a storage medium on which a computer program is stored, and the computer program is executed by a processor to implement the live video monitoring method.
  • the electronic device comprises a memory and a processor, wherein the memory stores a computer program running on the processor, and the live video monitoring method is implemented when the processor executes the computer program.
  • the live video monitoring system comprises a voice signal parsing module, a voice signal decomposition module, a voice signal recombining module and a voice signal identifying module, which are disposed on the server;
  • the voice signal parsing module is configured to: obtain a live video address between the live broadcasts, obtain a live video stream according to the live video address, parse the live video stream to obtain a voice data signal, and send a voice signal decomposition signal to the voice signal decomposition module;
  • the voice signal decomposition module is configured to: after receiving the signal decomposition signal, perform modal decomposition on the voice data by CEEMDAN to obtain an n-th order IMF component, n ⁇ 3, and send a voice signal recombination signal to the voice signal recombination module;
  • the speech signal recombining module is configured to: after receiving the recombination signal of the speech signal, separating the n-th order IMF components by the FastICA to obtain m reconstructed speech signals, 1 ⁇ m ⁇ n-2, and transmitting the speech signal identification signal to the speech signal recognition module;
  • the voice signal identification module is configured to: after receiving the voice signal identification signal, identify all reconstructed voice signals, determine whether there is a reconstructed voice signal containing bad information, and if so, perform an abnormal mark on the current live broadcast; otherwise, the current live broadcast Normal marking will be performed.
  • the algorithm developed by the present invention can reduce the noise of the live voice data signal without prior selection of the basis function, and completely relies on the characteristics of the detection signal to perform modal decomposition, and obtains several independent functions. IMF component. Therefore, the invention can effectively reduce the calculation cost and overcome the modal aliasing problem, thereby significantly reducing the noise of the voice data signal, thereby maximizing the perfect recognition of the live voice data, greatly improving the monitoring precision, and eliminating the bad information of the network. Communication provides a strong guarantee.
  • the invention realizes automatic monitoring of the network live broadcast room, which not only significantly reduces the labor cost, but also greatly improves the work efficiency.
  • FIG. 1 is a flowchart of a live video monitoring method according to an embodiment of the present invention.
  • FIG. 2 is a block diagram showing the connection of an electronic device in an embodiment of the present invention.
  • a live video monitoring method in an embodiment of the present invention includes the following steps:
  • the server obtains the live video address of the live broadcast, obtains the live video stream according to the live video address, parses the live video stream to obtain the voice data signal, and transfers to the S2.
  • the server uses CEEMDAN (CEEmpirical Mode DecompositionAN, a further improved algorithm of noise-assisted data analysis method.
  • the signal with poor performance can be decomposed by EMD to obtain a series of IMFs arranged from high frequency to low frequency, ie IMFIntrinsic Mode Function, intrinsic
  • the modulo function modally decomposes the speech data signal to obtain an n-th order IMF component, n ⁇ 3, thereby overcoming the modal aliasing problem while effectively reducing the computational cost, and shifting to S3.
  • S201 adding j-order Gaussian white noise to the speech data signal x(t), j is a positive integer, forming a Gaussian white noise sequence v j (t), using V j (t) as a decomposition signal, performing EMD decomposition, and obtaining an IMF component.
  • IMF jn , n represents the order of the IMF component, and goes to S202.
  • S202 Perform an average calculation on the IMF jn to obtain a first eigenmode function IMF of the CEEMDAN: IMF- ⁇ IMF jn , where J is a positive integer, and go to S203.
  • S205 The IMF re-executed as the IMF S203, to give R n and R n, determines whether the present step R n not decomposed (i.e., R n the extreme value is less than 2), if, according to the present step R n and IMF
  • S2a F0 is fitted to the maximal envelope e + (t) of the signal through the 3 times spline function; the minimum signal is fitted to f(t) through the 3 times spline function.
  • S2c Determine whether z(t) satisfies the preset criterion.
  • the preset criterion is: z(t) meets any of the following conditions: z(t) is a monotonic function, z(t) is a constant, and z(t) is less than a pre- Set the value; if so, use z(t) as the IMF component, otherwise use z(t) as f(t) and re-execute S2a.
  • S3 The server separates the n-th order IMF components by FastICA (Fast Independent Component Analysis) to obtain m reconstructed speech signals, 1 ⁇ m ⁇ n-2, to enhance the speech signal, and shift to S4.
  • FastICA Fast Independent Component Analysis
  • S4 The server recognizes all reconstructed speech signals, determines whether there is a reconstructed speech signal containing bad information, and if so, goes to S5; otherwise, goes to S6.
  • the server will mark the current live broadcast abnormally, that is, it is marked as the key monitoring live broadcast room, so that the staff can focus on the abnormally marked live broadcast room, that is, the similarity algorithm is used to determine whether the live broadcast room is a bad live broadcast room.
  • the purpose of purifying live content is used.
  • S6 The server will mark the current live broadcast normally, so that the staff randomly monitors the normal marked live broadcast room.
  • S5 and S6 can be completed in S4 during actual use, that is, S4 to S6 can be aggregated into one step.
  • the algorithm independently developed by the embodiment of the present invention can effectively reduce the calculation cost and overcome the modal aliasing problem when the live voice data signal is denoised, thereby significantly reducing the noise of the voice data signal.
  • the monitoring accuracy is greatly improved, and a strong guarantee for preventing the spread of bad information on the network is provided.
  • the embodiment of the invention further provides a storage medium, where the computer program is stored on the storage medium, and the live video monitoring method is implemented when the computer program is executed by the processor.
  • the storage medium includes a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a disk or an optical disk, and the like. The medium of the code.
  • an embodiment of the present invention further provides an electronic device, including a memory and a processor.
  • the memory stores a computer program running on the processor, and the live video monitoring method is implemented when the processor executes the computer program.
  • the live video monitoring system includes a voice signal parsing module, a voice signal decomposition module, a voice signal recombining module and a voice signal identification module.
  • the voice signal parsing module is configured to: obtain a live video address between live broadcasts, obtain a live video stream according to the live video address, parse the live video stream to obtain a voice data signal, and send a voice signal decomposition signal to the voice signal decomposition module.
  • the voice signal decomposition module is configured to: after receiving the voice signal decomposition signal, perform modal decomposition on the voice data by CEEMDAN to obtain an n-th order IMF component, n ⁇ 3, and send a voice signal recombination signal to the voice signal recombination module.
  • the workflow of the speech signal decomposition module includes:
  • Speech signal decomposition 01 Add j times Gaussian white noise to the speech data signal x(t), j is a positive integer, form a Gaussian white noise sequence v j (t), and use V j (t) as a decomposition signal to perform EMD decomposition. Obtaining the IMF component IMF jn , where n represents the order of the IMF component, and goes to the speech signal decomposition 02;
  • Speech signal decomposition 02 According to IMF ji calculated eigenmode function IMF: IMF- ⁇ IMF jn , where J is a positive integer, go to the speech signal decomposition 03;
  • Speech signal decomposition 03 Calculate the margin R n according to the IMF:
  • the workflow of the voice signal decomposition module to perform EMD decomposition on the decomposition signal f(t) includes:
  • EMD decomposition c Determine whether z(t) meets the preset criteria.
  • the preset criterion is: z(t) meets any of the following conditions: z(t) is a monotonic function, z(t) is a constant, z(t) Less than the preset value; if yes, z(t) is taken as the IMF component, otherwise z(t) is taken as f(t), and the EMD decomposition a is re-executed;
  • the voice signal recombination module is configured to: after receiving the voice signal recombination signal, the m-th order IMF component is separated by FastICA to obtain m reconstructed voice signals, 1 ⁇ m ⁇ n-2, and the voice signal identification signal is sent to the voice signal recognition module.
  • the workflow of the speech signal recombination module includes: after all the IMF components are divided into m groups, FastICA calculation is performed on each group of IMF components to obtain m reconstructed speech signals, and the order of each group of IMF components is m, m+1. m+2.
  • the voice signal identification module is configured to: after receiving the voice signal identification signal, identify all reconstructed voice signals, determine whether there is a reconstructed voice signal containing bad information, and if so, perform an abnormal mark on the current live broadcast; otherwise, the current live broadcast Normal marking will be performed.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Computer Hardware Design (AREA)
  • Computer Security & Cryptography (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)

Abstract

本发明公开了一种直播视频监控方法、存储介质、电子设备及系统,涉及直播视频监控领域。该方法的步骤为:服务端对直播视频流进行解析得到语音数据信号,服务端通过CEEMDAN对语音数据进行模态分解,得到n阶IMF分量;服务端将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2;服务端对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。本发明能够显著减少语音数据信号的噪声,以此最大化完美识别直播语音数据,大幅度提高监控精度,为杜绝网络不良信息的传播提供了有力保障。

Description

直播视频监控方法、存储介质、电子设备及系统 技术领域
本发明涉及直播视频监控领域,具体涉及一种直播视频监控方法、存储介质、电子设备及系统。
背景技术
随着直播行业的快速发展,越来越多的用户喜爱观看网络直播。为了实现避免网络不良信息的传播,直播平台需要对直播视频进行监控。现有的对直播视频进行监控的方法一般为:人工对直播平台中的所有直播间的直播视频,进行随机监控,当监控到某个直播间存在不良信息的语音时,将该直播间标记为重点监控直播间;随后安排人工对重点监控直播间进行全程监控,进而提高监控效率,净化直播内容。
但是,上述对直播视频进行监控的方法存在以下缺陷:
人工监控不仅所需的人力成本较大,而且当直播间数量较多时,每个直播间的直播视频的监控精度较低,进而使得不良信息的直播视频的被发现性较低,无法良好的实现杜绝网络不良信息的传播。
发明内容
针对现有技术中存在的缺陷,本发明解决的技术问题为:如何自动和精准的筛选出具有不良信息的网络直播视频。本发明能够显著减少语音数据信号的噪声,以此最大化完美识别直播语音数据,大幅度提高监控精度,为杜绝网络不良信息的传播提供了有力保障。
为达到以上目的,本发明提供的直播视频监控方法,包括以下步骤:
S1:服务端获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,转到S2;
S2:服务端通过CEEMDAN对语音数据进行模态分解,得到n阶IMF分量,n≥3,转到S3;
S3:服务端将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2,转到S4;
S4:服务端对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。
本发明提供的存储介质,该存储介质上存储有计算机程序,所述计算机程序被处理器执行时实现上述直播视频监控方法。
本发明提供的电子设备,包括存储器和处理器,存储器上储存有在处理器上运行的计算机程序,处理器执行计算机程序时实现上述直播视频监控方法。
本发明提供的直播视频监控系统,包括设置于服务端上的语音信号解析模块、语音信号分解模块、语音信号重组模块和语音信号识别模块;
语音信号解析模块用于:获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,向语音信号分解模块发送语音信号分解信号;
语音信号分解模块用于:收到语音信号分解信号后,通过CEEMDAN对语音数据进行模态分解,得到n阶IMF分量,n≥3,向语音信号重组模块发送语音信号重组信号;
语音信号重组模块用于:收到语音信号重组信号后,将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2,向语音 信号识别模块发送语音信号识别信号;
语音信号识别模块用于:收到语音信号识别信号后,对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。
与现有技术相比,本发明的优点在于:
通过本发明的S1至S6可知,本发明能够自主研制的算法,对直播语音数据信号进行降噪时,无需事先选定基函数,完全依靠检测信号本身的特征进行模态分解,得到若干独立的IMF分量。因此,本发明能够有效的减少计算成本并克服模态混叠问题,进而显著减少了语音数据信号的噪声,以此最大化完美识别直播语音数据,大幅度提高监控精度,为杜绝网络不良信息的传播提供了有力保障。
在此基础上,本发明实现了自动监控网络直播间,不仅显著降低了人力成本,而且大幅度提升了工作效率。
附图说明
图1为本发明实施例中直播视频监控方法的流程图;
图2为本发明实施例中电子设备的连接框图。
具体实施方式
以下结合附图及实施例对本发明作进一步详细说明。
参见图1所示,本发明实施例中的直播视频监控方法,包括以下步骤:
S1:服务端获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,转到S2。
S2:服务端通过CEEMDAN(CEEmpirical Mode DecompositionAN,噪声辅助数据分析方法的进一步改进算法,性能 不好的信号经过EMD分解后能够得到一系列由高频到低频排列的IMF,即IMFIntrinsic Mode Function,本征模函数)对语音数据信号进行模态分解,得到n阶IMF分量,n≥3,进而在有效减少计算成本的同时克服模态混叠问题,转到S3。
S2的具体流程为:
S201:对语音数据信号x(t)加入j次高斯白噪声,j为正整数,形成高斯白噪声序列v j(t),将v j(t)作为分解信号后进行EMD分解,得到IMF分量IMF jn,n代表IMF分量的阶数,转到S202。
通过S201可知,高斯白噪声序列v j(t)可以表示为:v j(t)=∑ n IMF jn+r j,其中r j为加入不同高斯白噪声的IMF分量的信号分解得出的趋势项。
S202:对IMF jn进行平均计算,得到CEEMDAN的第一个本征模态函数IMF:IMF-∑ IMF jn,其中J为正整数,转到S203。
S203:根据IMF计算余量R n
Figure PCTCN2017117973-appb-000001
转到S204。
S204:将R n-1+W作为分解信号进行EMD分解,获取EMD分解时的第一个模态E1,W=ε nE n(w j,ε n为常量(ε n的选择为本领域公知常识),w j为0至1之间的高斯白噪声序列,E n(w j为将w j作为分解信号进行EMD分解后,获取EMD分解时的第n个模态;根据E1计算模态IMF n
IMF-∑(R n-1+W),转到S205。
S205:将IMF作为IMF后重新执行S203,得到R n和R n,判断本步骤中的R n是否无法分解(即R n极点值是否小于2),若是,根据本步骤中的R n和IMF得到IMF分量x(t):x(t)=∑ n IMF n+R n,此时 S2结束,转到S3;否则根据本步骤中的R n-1重新执行S204。
通过S201至S205可知,通CEEMDAN加入自适应的高斯白噪声,紧接着计算其特定余量来获取相应的IMF分量,克服了原本EEMD加入的白噪声会使得分解失去完备性、产生重构误差的问题。
S201和S204中作为分解信号f(t)进行EMD分解的具体流程为:
S2a:对f(t)通过3次样条函数,拟合出信号的极大值包络线e +(t);对f(t)通过3次样条函数,拟合出信号的极小值包络线e -(t);根据e +(t)和e -(t)计算均值包络m i(t),m i(t)=(e +(t)+e -(t))/2;根据f(t)和m j(t)得到去掉低频的信号h,h=f(t)-m i(t),转到S2b。
S2b:判断h是否满足IMF的两个条件(即设定值范围):1、信号的极值点(极大值或极小值)数目和过零点数目相等或最多相差一个;2、由局部极大值构成的上包络线和由局部极小值构成的下包络线的平均值为零;若是,根据h和v j(t)得到去掉高频成分的信号z(t),z(t)=v j(t)-h,转到S2c;否则将h作为f(t)后,重新执行S2a。
S2b的实际操作流程为本领域技术人员的常规手段,在此不做赘述。
S2c:判断z(t)是否满足预设标准,预设标准为:z(t)符合以下任意一种情况:z(t)为单调函数、z(t)为常量、z(t)小于预设值;若是,将z(t)作为IMF分量,否则将z(t)作为f(t)后,重新执行S2a。
当f(t)为S201中的v j(t)时,IMF jn为z(t),当f(t)为S204中的R n-1+W时,E1为h,当f(t)为S204中的w j时,E n(w j为h n
S3:服务端将n阶IMF分量通过FastICA(FastIndependent Component Analysis,快速独立成分分析算法)分离得到m个重构语音信号,1≤m≤n-2,以增强语音信号,转到S4。
S3的原理和具体流程为:由于语音信号经过CEEMDAN自适应 分解后得到IMF分量的数目无法确定,因此单纯的选取某几阶IMF分量重构语音信号,可能会产生丢失语音信息的情况。而FastICA是基于高阶统计特性的分析方法,在消除噪声的同时,对其它信号的细节几乎没有破坏,去噪性能也往往要比传统的滤波方法好很多,对高阶统计特性的分析更符合实际。由于一定存在2阶IMF分量分别包含主要的噪声信号和语音信号。因此在无法确定噪声信号和语音信号对应的IMF分量的具体阶数n时,将所有IMF分量分为m组后,对每组IMF分量进行FastICA计算,得到m个重构语音信号,1≤m≤n-2,每组IMF分量的阶数为m,m+1,m+2,转到S4。
S4:服务端对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,转到S5;否则转到S6。
S5:服务端将当前直播间进行异常标记,即标记为重点监控直播间,进而使得工作人员对异常标记的直播间进行重点监控,即通过相似度算法判断直播间是否为不良直播间,以达到净化直播内容的目的。
S6:服务端将当前直播将进行正常标记,进而使得工作人员对正常标记的直播间进行随机监控。
在实际使用过程中S5和S6可以在S4中完成,即S4至S6可汇聚为1个步骤。
通过S1至S6可知,本发明实施例能够自主研制的算法,对直播语音数据信号进行降噪时,能够有效的减少计算成本并克服模态混叠问题,进而显著减少了语音数据信号的噪声,以此最大化完美识别直播语音数据,大幅度提高监控精度,为杜绝网络不良信息的传播提供了有力保障。
本发明实施例还提供一种存储介质,存储介质上存储有计算机程序,计算机程序被处理器执行时实现上述直播视频监控方法。需要说 明的是,所述存储介质包括U盘、移动硬盘、ROM(Read-Only Memory,只读存储器)、RAM(Random Access Memory,随机存取存储器)、磁碟或者光盘等各种可以存储程序代码的介质。
参见图2所示,本发明实施例还提供一种电子设备,包括存储器和处理器,存储器上储存有在处理器上运行的计算机程序,处理器执行计算机程序时实现上述直播视频监控方法。
本发明实施例提供的直播视频监控系统,包括设置于服务端上的语音信号解析模块、语音信号分解模块、语音信号重组模块和语音信号识别模块。
语音信号解析模块用于:获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,向语音信号分解模块发送语音信号分解信号。
语音信号分解模块用于:收到语音信号分解信号后,通过CEEMDAN对语音数据进行模态分解,得到n阶IMF分量,n≥3,向语音信号重组模块发送语音信号重组信号。
语音信号分解模块的工作流程包括:
语音信号分解01:对语音数据信号x(t)加入j次高斯白噪声,j为正整数,形成高斯白噪声序列v j(t),将v j(t)作为分解信号后进行EMD分解,得到IMF分量IMF jn,n代表IMF分量的阶数,转到语音信号分解02;
语音信号分解02:根据IMF ji计算得到本征模态函数IMF:IMF-∑ IMF jn,其中J为正整数,转到语音信号分解03;
语音信号分解03:根据IMF计算余量R n
Figure PCTCN2017117973-appb-000002
转到语音信号分解04;
语音信号分解04:将R n-1+W作为分解信号进行EMD分解,获取EMD分解时的第一个模态E1,W=ε nE n(w j,ε n为常量,w j为0至1之间的高斯白噪声序列,E n(w j为将w j作为分解信号进行EMD分解后,获取EMD分解时的第n个模态;根据E1计算模态IMF n
IMF -∑(R n-1+W),转到语音信号分解05;
语音信号分解05:将IMF作为IMF后重新执行语音信号分解03,得到R n和R n;判断本步骤中的R n是否无法分解,若是,根据本步骤中的R n和IMF得到IMF分量x(t):x(t)=∑ n IMF n+R n,向语音信号重组模块发送语音信号重组信号;否则根据本步骤中的R n-1重新执行语音信号分解04。
语音信号分解模块将分解信号f(t)进行EMD分解的工作流程包括:
EMD分解a:对f(t)通过3次样条函数,拟合出信号的极大值包络线e +(t);对f(t)通过3次样条函数,拟合出信号的极小值包络线e -(t);根据e +(t)和e -(t)计算均值包络m i(t),m i(t)=(e +(t)+e -(t))/2;根据f(t)和m j(t)得到去掉低频的信号h,h=f(t)-m i(t),转到EMD分解b;
EMD分解b:判断h是否满足设定值范围,若是,根据h和v j(t)得到去掉高频成分的信号z(t),z(t)=v j(t)-h,转到EMD分解c;否则将h作为f(t)后,重新执行EMD分解a;
EMD分解c:判断z(t)是否满足预设标准,预设标准为:z(t)符合以下任意一种情况:z(t)为单调函数、z(t)为常量、z(t)小于预设值;若是,将z(t)作为IMF分量,否则将z(t)作为f(t)后,重新执行EMD分解a;
当f(t)为所述语音信号分解01中的v j(t)时,IMF jn为z(t),当f(t)为所述语音信号分解04中的R n-1+W时,E1为h,当f(t)为所述语 音信号分解04中的w j时,E n(w j为h n
语音信号重组模块用于:收到语音信号重组信号后,将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2,向语音信号识别模块发送语音信号识别信号。语音信号重组模块的工作流程包括:将所有IMF分量分为m组后,对每组IMF分量进行FastICA计算,得到m个重构语音信号,每组IMF分量的阶数为m,m+1,m+2。
语音信号识别模块用于:收到语音信号识别信号后,对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。
需要说明的是:本发明实施例提供的系统在进行模块间通信时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将系统的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。
进一步,本发明不局限于上述实施方式,对于本技术领域的普通技术人员来说,在不脱离本发明原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也视为本发明的保护范围之内。本说明书中未作详细描述的内容属于本领域专业技术人员公知的现有技术。

Claims (10)

  1. 一种直播视频监控方法,其特征在于,该方法包括以下步骤:
    S1:服务端获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,转到S2;
    S2:服务端通过噪声辅助数据分析方法的进一步改进算法CEEMDAN对语音数据进行模态分解,得到n阶IMF分量,n≥3,转到S3;
    S3:服务端将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2,转到S4;
    S4:服务端对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。
  2. 如权利要求1所述的直播视频监控方法,其特征在于,S2的流程包括:
    S201:对语音数据信号x(t)加入j次高斯白噪声,j为正整数,形成高斯白噪声序列v j(t),将v j(t)作为分解信号后进行EMD分解,得到IMF分量IMF jn,n代表IMF分量的阶数,转到S202;
    S202:根据IMF jn计算得到本征模态函数IMF 1
    Figure PCTCN2017117973-appb-100001
    Figure PCTCN2017117973-appb-100002
    其中J为正整数,转到S203;
    S203:根据IMF 1计算余量R n
    Figure PCTCN2017117973-appb-100003
    转到S204;
    S204:将R n-1+W作为分解信号进行EMD分解,获取EMD分解时的第一个模态E1,W=ε nE n(w j,ε n为常量,w j为0至1之间的高斯白噪声序列,E n(w j为将w j作为分解信号进行EMD分解后,获 取EMD分解时的第n个模态;根据E1计算模态IMF n
    Figure PCTCN2017117973-appb-100004
    转到S205;
    S205:将IMF作为IMF 1后重新执行S203,得到R n和R n  1;判断本步骤中的R n是否无法分解,若是,根据本步骤中的R n和IMF得到IMF分量x(t):x(t)=∑ n  1IMF n+R n,转到S3;否则根据本步骤中的R n-1重新执行S204。
  3. 如权利要求2所述的直播视频监控方法,其特征在于,S201和S204中所述作为分解信号f(t)进行EMD分解的流程包括:
    S2a:对f(t)通过3次样条函数,拟合出信号的极大值包络线e +(t);对f(t)通过3次样条函数,拟合出信号的极小值包络线e -(t);根据e +(t)和e -(t)计算均值包络m i(t),m i(t)=(e +(t)+e -(t))/2;根据f(t)和m j(t)得到去掉低频的信号h 1,h 1=f(t)-m i(t),转到S2b;
    S2b:判断h 1是否满足设定值范围,若是,根据h 1和v j(t)得到去掉高频成分的信号z(t),z(t)=v j(t)-h 1,转到S2c;否则将h 1作为f(t)后,重新执行S2a;
    S2c:判断z(t)是否满足预设标准,预设标准为:z(t)符合以下任意一种情况:z(t)为单调函数、z(t)为常量、z(t)小于预设值;若是,将z(t)作为IMF分量,否则将z(t)作为f(t)后,重新执行S2a;
    当f(t)为S201中的v j(t)时,IMF jn为z(t),当f(t)为S204中的R n-1+W时,E1为h 1,当f(t)为S204中的w j时,E n(w j
    Figure PCTCN2017117973-appb-100005
  4. 如权利要求1至3任一项所述的直播视频监控方法,其特征在于,S3的流程包括:将所有IMF分量分为m组后,对每组IMF分量进行FastICA计算,得到m个重构语音信号,每组IMF分量的阶数为m,m+1,m+2。
  5. 一种存储介质,该存储介质上存储有计算机程序,其特征在 于:所述计算机程序被处理器执行时实现权利要求1至4任一项所述的方法。
  6. 一种电子设备,包括存储器和处理器,存储器上储存有在处理器上运行的计算机程序,其特征在于:处理器执行计算机程序时实现权利要求1至4任一项所述的方法。
  7. 一种直播视频监控系统,其特征在于,该系统包括设置于服务端上的语音信号解析模块、语音信号分解模块、语音信号重组模块和语音信号识别模块;
    语音信号解析模块用于:获取直播间的直播视频地址,根据直播视频地址得到直播视频流,对直播视频流进行解析得到语音数据信号,向语音信号分解模块发送语音信号分解信号;
    语音信号分解模块用于:收到语音信号分解信号后,通过CEEMDAN对语音数据进行模态分解,得到n阶IMF分量,n≥3,向语音信号重组模块发送语音信号重组信号;
    语音信号重组模块用于:收到语音信号重组信号后,将n阶IMF分量通过FastICA分离得到m个重构语音信号,1≤m≤n-2,向语音信号识别模块发送语音信号识别信号;
    语音信号识别模块用于:收到语音信号识别信号后,对所有重构语音信号进行识别,判断是否存在含有不良信息的重构语音信号,若是,将当前直播间进行异常标记;否则将当前直播将进行正常标记。
  8. 如权利要求7所述的直播视频监控系统,其特征在于,所述语音信号分解模块的工作流程包括:
    语音信号分解01:对语音数据信号x(t)加入j次高斯白噪声,j为正整数,形成高斯白噪声序列v j(t),将v j(t)作为分解信号后进行EMD分解,得到IMF分量IMF jn,n代表IMF分量的阶数,转到语 音信号分解02;
    语音信号分解02:根据IMF jn计算得到本征模态函数IMF 1
    Figure PCTCN2017117973-appb-100006
    其中J为正整数,转到语音信号分解03;
    语音信号分解03:根据IMF 1计算余量R n
    Figure PCTCN2017117973-appb-100007
    转到语音信号分解04;
    语音信号分解04:将R n-1+W作为分解信号进行EMD分解,获取EMD分解时的第一个模态E1,W=ε nE n(w j,ε n为常量,w j为0至1之间的高斯白噪声序列,E n(w j为将w j作为分解信号进行EMD分解后,获取EMD分解时的第n个模态;根据E1计算模态IMF n
    Figure PCTCN2017117973-appb-100008
    转到语音信号分解05;
    语音信号分解05:将IMF作为IMF 1后重新执行语音信号分解03,得到R n和R n  1;判断本步骤中的R n是否无法分解,若是,根据本步骤中的R n和IMF得到IMF分量x(t):x(t)=∑ n  1IMF n+R n,向语音信号重组模块发送语音信号重组信号;否则根据本步骤中的R n-1重新执行语音信号分解04。
  9. 如权利要求8所述的直播视频监控系统,其特征在于,所述语音信号分解模块将分解信号f(t)进行EMD分解的工作流程包括:
    EMD分解a:对f(t)通过3次样条函数,拟合出信号的极大值包络线e +(t);对f(t)通过3次样条函数,拟合出信号的极小值包络线e -(t);根据e +(t)和e -(t)计算均值包络m i(t),m i(t)=(e +(t)+e -(t))/2;根据f(t)和m j(t)得到去掉低频的信号h 1,h 1=f(t)-m i(t),转到EMD分解b;
    EMD分解b:判断h 1是否满足设定值范围,若是,根据h 1和v j(t)得到去掉高频成分的信号z(t),z(t)=v j(t)-h 1,转到EMD分解c;否则将h 1作为f(t)后,重新执行EMD分解a;
    EMD分解c:判断z(t)是否满足预设标准,预设标准为:z(t)符合以下任意一种情况:z(t)为单调函数、z(t)为常量、z(t)小于预设值;若是,将z(t)作为IMF分量,否则将z(t)作为f(t)后,重新执行EMD分解a;
    当f(t)为所述语音信号分解01中的v j(t)时,IMF jn为z(t),当f(t)为所述语音信号分解04中的R n-1+W时,E1为h 1,当f(t)为所述语音信号分解04中的w j时,E n(w j
    Figure PCTCN2017117973-appb-100009
  10. 如权利要求7至9任一项所述的直播视频监控系统,其特征在于:所述语音信号重组模块的工作流程包括:将所有IMF分量分为m组后,对每组IMF分量进行FastICA计算,得到m个重构语音信号,每组IMF分量的阶数为m,m+1,m+2。
PCT/CN2017/117973 2017-06-22 2017-12-22 直播视频监控方法、存储介质、电子设备及系统 Ceased WO2018233256A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201710483586.5 2017-06-22
CN201710483586.5A CN107465657A (zh) 2017-06-22 2017-06-22 直播视频监控方法、存储介质、电子设备及系统

Publications (1)

Publication Number Publication Date
WO2018233256A1 true WO2018233256A1 (zh) 2018-12-27

Family

ID=60544023

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2017/117973 Ceased WO2018233256A1 (zh) 2017-06-22 2017-12-22 直播视频监控方法、存储介质、电子设备及系统

Country Status (2)

Country Link
CN (1) CN107465657A (zh)
WO (1) WO2018233256A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110096956A (zh) * 2019-03-25 2019-08-06 中国地质大学(武汉) 基于eemd和排列熵二阶差分的信号去噪方法及装置
CN111272429A (zh) * 2020-03-04 2020-06-12 贵州大学 一种轴承故障诊断方法
CN114371326A (zh) * 2021-12-06 2022-04-19 国网江苏省电力有限公司电力科学研究院 一种光纤电流互感器的非线性误差识别方法、装置及存储介质
CN115144902A (zh) * 2021-03-30 2022-10-04 中国石油化工股份有限公司 储层含油气分析方法、装置、存储介质及电子设备
CN115988216A (zh) * 2022-12-27 2023-04-18 天翼云科技有限公司 无损关联编码的方法、系统、设备及介质
CN115988216B (zh) * 2022-12-27 2026-05-19 天翼云科技有限公司 无损关联编码的方法、系统、设备及介质

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107465657A (zh) * 2017-06-22 2017-12-12 武汉斗鱼网络科技有限公司 直播视频监控方法、存储介质、电子设备及系统
CN110493615A (zh) * 2018-05-15 2019-11-22 武汉斗鱼网络科技有限公司 一种直播视频监控方法及电子设备
CN111107380B (zh) * 2018-10-10 2023-08-15 北京默契破冰科技有限公司 一种用于管理音频数据的方法、设备和计算机存储介质
CN111031329B (zh) * 2018-10-10 2023-08-15 北京默契破冰科技有限公司 一种用于管理音频数据的方法、设备和计算机存储介质
CN111383659B (zh) * 2018-12-28 2021-03-23 广州市百果园网络科技有限公司 分布式语音监控方法、装置、系统、存储介质和设备
CN114187555A (zh) * 2021-12-14 2022-03-15 四川省人工智能研究院(宜宾) 一种面向社交视频直播的异常事件检测方法

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103106903A (zh) * 2013-01-11 2013-05-15 太原科技大学 一种单通道盲源分离法
CN104375973A (zh) * 2014-11-24 2015-02-25 沈阳建筑大学 一种基于集合经验模态分解的盲源信号去噪方法
CN104882140A (zh) * 2015-02-05 2015-09-02 宇龙计算机通信科技(深圳)有限公司 基于盲信号提取算法的语音识别方法及系统
CN105976077A (zh) * 2016-03-04 2016-09-28 国家电网公司 一种输变电工程造价动态控制目标计算系统及计算方法
CN106101819A (zh) * 2016-06-21 2016-11-09 武汉斗鱼网络科技有限公司 一种基于语音识别的直播视频敏感内容过滤方法及装置
CN106228529A (zh) * 2016-09-05 2016-12-14 上海理工大学 一种激光散斑图像处理分析方法
CN106250837A (zh) * 2016-07-27 2016-12-21 腾讯科技(深圳)有限公司 一种视频的识别方法、装置和系统
CN107465657A (zh) * 2017-06-22 2017-12-12 武汉斗鱼网络科技有限公司 直播视频监控方法、存储介质、电子设备及系统

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103106903A (zh) * 2013-01-11 2013-05-15 太原科技大学 一种单通道盲源分离法
CN104375973A (zh) * 2014-11-24 2015-02-25 沈阳建筑大学 一种基于集合经验模态分解的盲源信号去噪方法
CN104882140A (zh) * 2015-02-05 2015-09-02 宇龙计算机通信科技(深圳)有限公司 基于盲信号提取算法的语音识别方法及系统
CN105976077A (zh) * 2016-03-04 2016-09-28 国家电网公司 一种输变电工程造价动态控制目标计算系统及计算方法
CN106101819A (zh) * 2016-06-21 2016-11-09 武汉斗鱼网络科技有限公司 一种基于语音识别的直播视频敏感内容过滤方法及装置
CN106250837A (zh) * 2016-07-27 2016-12-21 腾讯科技(深圳)有限公司 一种视频的识别方法、装置和系统
CN106228529A (zh) * 2016-09-05 2016-12-14 上海理工大学 一种激光散斑图像处理分析方法
CN107465657A (zh) * 2017-06-22 2017-12-12 武汉斗鱼网络科技有限公司 直播视频监控方法、存储介质、电子设备及系统

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
XIANRONG DENG ET AL.: "Single-Channel Blind Singal Separation Method For Time-frequency Overlapped Signal Based on CEEMD-FastICA", 2016 IEEE 13TH INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING PROCEEDINGS (ICSP2016), 6 November 2016 (2016-11-06), pages 1440 - 1445, XP033076187 *

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110096956A (zh) * 2019-03-25 2019-08-06 中国地质大学(武汉) 基于eemd和排列熵二阶差分的信号去噪方法及装置
CN110096956B (zh) * 2019-03-25 2023-02-10 中国地质大学(武汉) 基于eemd和排列熵二阶差分的信号去噪方法及装置
CN111272429A (zh) * 2020-03-04 2020-06-12 贵州大学 一种轴承故障诊断方法
CN111272429B (zh) * 2020-03-04 2021-08-17 贵州大学 一种轴承故障诊断方法
CN115144902A (zh) * 2021-03-30 2022-10-04 中国石油化工股份有限公司 储层含油气分析方法、装置、存储介质及电子设备
CN114371326A (zh) * 2021-12-06 2022-04-19 国网江苏省电力有限公司电力科学研究院 一种光纤电流互感器的非线性误差识别方法、装置及存储介质
CN115988216A (zh) * 2022-12-27 2023-04-18 天翼云科技有限公司 无损关联编码的方法、系统、设备及介质
CN115988216B (zh) * 2022-12-27 2026-05-19 天翼云科技有限公司 无损关联编码的方法、系统、设备及介质

Also Published As

Publication number Publication date
CN107465657A (zh) 2017-12-12

Similar Documents

Publication Publication Date Title
CN111160469B (zh) 一种目标检测系统的主动学习方法
CN112492343B (zh) 一种视频直播监控方法及相关装置
US11748595B2 (en) Convolution acceleration operation method and apparatus, storage medium and terminal device
CN108172213B (zh) 娇喘音频识别方法、装置、设备及计算机可读介质
CN105095675B (zh) 一种开关柜故障特征选择方法及装置
CN111294819A (zh) 一种网络优化方法及装置
CN118968038B (zh) 目标检测方法、装置、设备、存储介质及产品
CN106547744A (zh) 一种图像检索方法及系统
CN104518905A (zh) 一种故障定位方法及装置
Doan et al. S-SOM v1. 0: A structural self-organizing map algorithm for weather typing
CN107465657A (zh) 直播视频监控方法、存储介质、电子设备及系统
CN111177193A (zh) 一种基于Flink的日志流式处理方法及系统
CN109146923B (zh) 一种目标跟踪丢断帧的处理方法及系统
CN114330542A (zh) 一种基于目标检测的样本挖掘方法、装置及存储介质
US20250292059A1 (en) Data processing method and apparatus, device, and medium
CN110245684B (zh) 数据处理方法、电子设备和介质
CN110096605B (zh) 图像处理方法及装置、电子设备、存储介质
CN111027316A (zh) 文本处理方法、装置、电子设备及计算机可读存储介质
CN117457017B (zh) 语音数据的清洗方法及电子设备
CN112863548A (zh) 训练音频检测模型的方法、音频检测方法及其装置
CN117496263A (zh) 一种困难样本预标注方法、系统、存储介质和电子设备
CN113449062B (zh) 轨迹处理方法、装置、电子设备和存储介质
CN113947195A (zh) 模型确定方法、装置、电子设备和存储器
CN113936693A (zh) 一种音频筛选方法及系统
CN111258788A (zh) 磁盘故障预测方法、装置及计算机可读存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17914748

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17914748

Country of ref document: EP

Kind code of ref document: A1