WO2018184306A1 - 话题预警的方法、装置、计算机设备及存储介质 - Google Patents
话题预警的方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2018184306A1 WO2018184306A1 PCT/CN2017/090579 CN2017090579W WO2018184306A1 WO 2018184306 A1 WO2018184306 A1 WO 2018184306A1 CN 2017090579 W CN2017090579 W CN 2017090579W WO 2018184306 A1 WO2018184306 A1 WO 2018184306A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- keyword
- extended
- keywords
- similarity
- target
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/284—Lexical analysis, e.g. tokenisation or collocates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/40—Business processes related to social networking or social networking services
- G06Q10/44—Identification of trends within social networks, e.g. identification of trending topics
Definitions
- the present invention relates to the field of computer processing, and in particular, to a method, device, computer device and storage medium for topic early warning.
- social media is also an important way to spread publicity.
- the traditional monitoring of social media topics is based on the analysis of the historical data obtained, and then the labeling of different topics. Since the topic update speed is very fast, the results obtained only by analyzing the historical data are obviously not accurate enough, and the traditional topic monitoring is to monitor all the topics without considering the individual needs of the user.
- a method, apparatus, computer device, and storage medium for topic warning are provided.
- a method of topic warning including:
- the topic is alerted.
- a device for early warning of topics including:
- a custom keyword acquisition module for obtaining a custom keyword
- An extended keyword obtaining module configured to calculate a similarity between the custom keyword and each word in the corpus, and obtain an extended keyword related to the customized keyword from the corpus according to the similarity;
- a target keyword screening module configured to filter a target keyword from the extended keyword according to a type of the extended keyword and a similarity between the extended keyword and the customized keyword, and join a target key List of words;
- a monitoring module configured to perform real-time monitoring according to the target keyword in the target keyword list
- the warning module is configured to perform topic warning when the topic amount corresponding to the target keyword is monitored to reach a preset threshold.
- a computer device comprising a memory and a processor, the memory storing computer readable instructions, the computer readable instructions being executed by the processor such that the processor performs the following steps:
- the topic is alerted.
- One or more non-transitory readable storage mediums storing computer readable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps:
- the topic is alerted.
- FIG. 1 is a block diagram showing the internal structure of a terminal in an embodiment
- FIG. 2 is a block diagram showing the internal structure of a server in an embodiment
- FIG. 3 is a flow chart of a method for alerting a topic in an embodiment
- FIG. 4 is a flowchart of a method for filtering a target keyword from an extended keyword according to the type of the extended keyword and the similarity between the extended keyword and the customized keyword in one embodiment
- FIG. 5 is a flow chart of a method for warning a topic in another embodiment
- FIG. 6 is a flow chart of a method for calculating a similarity between a custom keyword and each corpus in a corpus in an embodiment, and obtaining an extended keyword from the corpus according to the similarity;
- FIG. 7 is a structural block diagram of an apparatus for alerting a topic in an embodiment
- FIG. 8 is a structural block diagram of a target keyword screening module in an embodiment
- Figure 9 is a block diagram showing the structure of an apparatus for alerting a topic in another embodiment.
- the internal structure of the terminal 102 is as shown in FIG. 1, including a processor connected through a system bus, a non-volatile storage medium, an internal memory, a network interface, a display screen, and an input device. .
- the processor of the terminal 102 is used to provide computing and control capabilities to support the operation of the entire terminal 102.
- the non-volatile storage medium stores operating systems and computer readable instructions executable by the processor to implement a method for topic alerting of the terminal 102.
- the internal memory in terminal 102 provides an environment for the operation of operating systems and computer readable instructions in a non-volatile storage medium.
- the network interface is used to connect to the network for communication.
- the display screen of the terminal 102 may be a liquid crystal display or an electronic ink display screen.
- the input device may be a touch layer covered on the display screen, or may be a button, a trackball or a touchpad provided on the outer casing of the electronic device, or may be An external keyboard, trackpad, or mouse.
- the terminal 102 can be a tablet, a laptop, a desktop computer, or the like.
- FIG. 1 A person skilled in the art can understand that the structure shown in FIG. 1 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the terminal to which the solution of the present application is applied.
- the specific terminal may include a ratio. More or fewer components are shown in the figures, or some components are combined, or have different component arrangements.
- the internal structure of server 104 includes a processor coupled through a system bus, a non-volatile storage medium, an internal memory, and a network interface.
- the processor of the server 104 is used to provide computing and control capabilities to support the operation of the entire server.
- the non-volatile storage medium includes an operating system and computer readable instructions.
- the computer readable instructions are executable by a processor to implement a method for alerting a topic to the server 104.
- the internal memory of the server 104 is an operating system and computer readable instructions in a non-volatile storage medium. Providing an environment, the server's network interface is used to communicate with external servers and terminals over a network connection.
- FIG. 2 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the server to which the solution of the present application is applied.
- the specific server may include a ratio. More or fewer components are shown in the figures, or some components are combined, or have different component arrangements.
- a method for alerting a topic is proposed, which may be applied to a computer device, where the computer device may be a terminal or a server, and specifically includes the following steps:
- step 302 a custom keyword is obtained.
- the custom keyword refers to a keyword given by the user that meets the user's listening needs.
- the setting of the monitoring keyword is set according to the user-defined keyword. Due to the complexity of social media information in the era of big data, the main body is diverse, and the topics of different users are different. Among them, the topic refers to the subject of discussion, because different subjects pay attention to different subjects, so through the custom keywords Not only can bring friendly user interaction, but more can realize the personalization and diversification of user monitoring needs.
- Step 304 Calculate the similarity between the custom keyword and each word in the corpus, and obtain the extended keyword related to the customized keyword from the corpus according to the similarity.
- the customized keywords given by the user are often not complete and comprehensive, it is necessary to extend the custom keywords.
- Obtaining the extended keywords related to the customized keyword is beneficial to ensure that the user is more comprehensive and complete on the topics to be monitored, thereby ensuring the integrity and diversity of the monitoring results.
- the words with similarities with the custom keywords are selected from the corpus as the extended keywords. The greater the similarity, the closer the semantics of the word to the custom keyword.
- the similarity between words can be calculated by using the synonym word forest, or the Pearson correlation coefficient can be used to calculate the similarity between words. The calculation method of word similarity is not limited here.
- the calculation of similarity is obtained by calculating the similarity between word vectors.
- the word2vec model is used to calculate the word vector corresponding to the custom keyword.
- Word2vec is an efficient tool for characterizing words as real-value vectors. It can use the idea of deep learning to simplify the processing of text content through training. It is a vector operation in k-dimensional vector space, and the similarity in vector space can be used to represent the semantic similarity of text.
- the custom keyword is input as the word2vec model, and the word vector representation of the custom keyword is output.
- the extended keyword of the custom keyword is filtered from the corpus by calculating the similarity between the word vectors.
- the words in the corpus can be stored in the form of word vectors.
- n represents the nth word vector feature of the word vector and i represents the ith word vector feature in the word vector.
- the extended keywords related to the customized keywords are filtered by calculating the similarity between the custom keywords and each word in the corpus. Specifically, the similarities may be arranged in descending order, and the top k words with the highest similarity are selected as the extended keywords of the customized keywords.
- the expansion of custom keywords makes the keywords more diverse, ensuring that the topic monitoring results have contrast with similar keywords, which is convenient for providing decision makers with richer information.
- Step 306 Filter the target keyword from the extended keyword according to the type of the extended keyword and the similarity between the extended keyword and the customized keyword, and add the target keyword list.
- the information will be confusing. Therefore, in order to ensure the clarity of the information, it is necessary to further filter the obtained extended keywords.
- first, all the obtained extended keywords are classified, and then selected and customized from each class.
- the first h extended keywords with the highest similarity of the keyword are used as the target keywords, wherein h is a positive integer greater than 0, and the target keywords filtered by each class are aggregated to generate a target keyword list for monitoring.
- the types corresponding to all the extended words are acquired, and then the keywords of the same type are grouped.
- extension words corresponding to each type of extended keyword and use the type with the smallest number of extended words as the benchmark. If the number of types with the smallest number of extended words corresponds to X, then X is also selected from each of the other types.
- the extended keywords are used as the target keywords. Among them, the X extended keywords are filtered from the other types according to the similarity, and the highest similarity among the other extended keywords is selected. The X extended keywords are used as target keywords to join the target keyword list.
- Step 308 Perform real-time monitoring according to the target keyword in the target keyword list.
- a sliding window based timing management framework can be used.
- the main idea of the sliding window-based timing management framework is: for each target keyword in the target listening list, the topic data stream is managed in the form of a sliding window, and each target keyword maintains a certain size of the cache, each time A time slice (for real-time monitoring, the time slice setting is usually small, such as 5 minutes), the data window slides, and then the data in the cache is processed.
- Step 310 When the number of topics corresponding to the target keyword is monitored reaches a preset threshold, the topic is alerted.
- a good monitoring must require an early warning, and the topic is alerted by monitoring whether the topic amount corresponding to the target keyword reaches a preset threshold.
- the early warning can be considered from two aspects. First, the monitoring amount of the topic amount in the preset time slice is monitored. Since the time slice is a short time, it is possible to alert an unexpected event in a short time by monitoring the topic in a short time. Second, for the early warning of the topic of a period of time, many times the occurrence of events or the trend of public opinion is not necessarily sharp. Therefore, examining the hotspots of topics over a period of time can help decision-makers discover the rise of events or the gradual trend of public opinion. . Specifically, two evaluation strategies are used.
- the real-time warning of keywords is to use the topic heat to conduct early warning.
- the critical threshold of heat is determined empirically, when the target keyword of the monitor is in a sliding window.
- An early warning response is made when the frequency occurring within the time slice is greater than the critical temperature threshold.
- One is to use the emotional polarity ratio to make an early warning, and to analyze the emotional polarity of the social network text related to the target keyword list of the monitoring, mainly including the positive, neutral and negative emotional polarity, when the negative emotion is at all
- an early warning is performed.
- the method of early warning of this topic can be applied to many fields, especially in the financial field. Take the application of financial products as an example to illustrate the benefits of early warning on this topic.
- the Internet is closely related to the financial industry. According to the monitoring of Internet data, financial products can avoid many losses.
- the keywords related to finance are relatively regular, and relatively fixed. By monitoring and warning the topics related to financial products, it can achieve rapid response without losing accuracy.
- the target keyword that is finally used for monitoring is selected, and then real-time monitoring is performed on the social media according to the target keyword, and when the topic amount of the target keyword is monitored reaches a preset threshold, the topic is alerted.
- the method not only can monitor the topic in real time, but also can be monitored based on the user-defined keywords to meet the needs of the user's personalized monitoring and early warning.
- the target keyword is selected from the extended keyword according to the type of the extended keyword and the similarity between the extended keyword and the customized keyword, and the step of adding the target keyword list is performed.
- step 306A the extended keywords are classified according to a preset type.
- the extended keywords are classified into three categories according to "brand", “product”, and “competition”. This way, it is easy to follow each
- the class picks out the same number of target keywords for monitoring, which helps to ensure clear and comprehensive monitoring information.
- Step 306B Filter out the top h extended keywords with the highest similarity with the customized keywords from the extended keywords of each class as the target keywords, where h is a positive integer greater than 0.
- the crowdsourcing strategy is used to filter out the top h extended keywords with the highest similarity with the customized keywords from each type of extended keywords.
- Target keyword For example, the top 5 words with the highest similarity to the custom keywords are selected from each category, and finally the selected target keywords of each category are aggregated.
- Step 306C The target keywords selected by each class are aggregated to generate a target keyword list for monitoring.
- the target keywords of each type are filtered. Up, put in the same list, that is, generate a target keyword list, and then facilitate real-time monitoring according to the target keyword in the target keyword list. For example, if the extended keywords are classified into three categories according to "brand", "product", and "competition”. If 5 target keywords are selected for each category, then 15 target keywords will be selected for monitoring. By classifying the extended keywords and then screening for each category, the content of the monitoring is more clear and comprehensive, and there is no biased result.
- a method for topic early warning comprising:
- Step 502 Obtain a custom keyword.
- Step 504 Calculate a word vector corresponding to the customized keyword.
- Step 506 Calculate the similarity between the word vector of the customized keyword and the word vector of each word in the corpus, and obtain the extended keyword related to the customized keyword from the corpus according to the similarity between the word vectors.
- Step 508 Filter the target keyword from the extended keyword according to the type of the extended keyword and the similarity between the extended keyword and the customized keyword, and add the target keyword list.
- Step 510 Perform real-time monitoring according to the target keyword in the target keyword list.
- Step 512 When the number of topics corresponding to the target keyword is monitored reaches a preset threshold, the topic is alerted.
- the word vector corresponding to the custom keyword needs to be calculated first, and the custom keyword is used as the input of the word2vec model. , generating a word vector corresponding to the custom keyword and outputting it.
- the extended keyword related to the customized keyword is obtained by calculating the similarity between the custom keyword and each word in the corpus, wherein the higher the similarity, the closer the semantics of the custom keyword are.
- the Pearson Correlation Coefficient method can be used to calculate the similarity between the word vector of the custom keyword and the word vector of each word in the corpus, and the highest similarity with the custom keyword is selected.
- each category selects the top h words with the highest similarity with the customized keyword as the target. Key words, then the target keywords selected by each class are summarized and placed in the same list, that is, the target keyword list is added. Then, according to the target keyword list, the monitoring is performed, and corresponding warnings are performed.
- the method expands the user-defined keywords to ensure the diversity and comprehensiveness of the monitoring. The further selection of the extended keywords by the crowdsourcing technology ensures that the monitoring results are not biased.
- calculating the similarity between the custom keyword and each word in the corpus, and obtaining the extended keyword related to the customized keyword from the corpus according to the similarity includes:
- Step 304A using the Pearson correlation coefficient method to calculate each of the custom keywords and the corpus The similarity between words.
- Step 304B Acquire the top K words with the highest similarity with the customized keyword as the extended keyword of the customized keyword, where K is a positive integer greater than 0.
- the step of performing real-time monitoring according to the target keyword in the target keyword list comprises: performing real-time monitoring on each target keyword in the target keyword list in the form of a sliding window.
- real-time monitoring since the social media data is generated all the time, and is fast and large in scale, in order to achieve real-time monitoring of the topic, it is necessary to solve how to conduct the topic in the context of the data stream.
- Real-time monitoring In this embodiment, real-time monitoring of each target keyword in the target keyword column is performed by adopting a form based on a sliding window. That is, the topic data stream is managed in the form of a sliding window, each target keyword maintains a certain size of the cache, and the data window slides after each time slice, and then the data in the cache is processed, thereby realizing Target keywords for real-time monitoring.
- a device 700 for topic warning is proposed, the device comprising:
- the custom keyword acquisition module 702 is configured to obtain a custom keyword.
- the extended keyword obtaining module 704 is configured to calculate a similarity between the customized keyword and each word in the corpus, and obtain an extended keyword related to the customized keyword from the corpus according to the similarity.
- the target keyword screening module 706 is configured to filter the target keyword from the extended keyword according to the type of the extended keyword and the similarity between the extended keyword and the customized keyword, and add the target keyword list.
- the monitoring module 708 is configured to perform real-time monitoring according to the target keyword in the target keyword list.
- the warning module 710 is configured to perform a topic warning when the topic amount corresponding to the target keyword reaches a preset threshold.
- the target keyword screening module 706 includes:
- the classification module 706A is configured to classify the extended keywords according to a preset type.
- the screening module 706B is configured to filter, from each of the extended keywords, the top h extended keywords with the highest similarity with the customized keywords as the target keywords, where h is a positive integer greater than 0.
- the aggregation module 706C is configured to aggregate the target keywords filtered by each class to generate a target keyword list for monitoring.
- a device for alerting a topic is proposed.
- the method further includes:
- the calculation module 703 is configured to calculate a word vector corresponding to the customized keyword.
- the extended keyword acquisition module 704 is also used to calculate a word vector and a corpus of a custom keyword.
- the similarity between the word vectors of each word, and the extended keywords related to the customized keywords are obtained from the corpus according to the similarity between the word vectors.
- the extended word acquisition module is further configured to calculate a similarity between the customized keyword and each word in the corpus by using a Pearson correlation coefficient method, and obtain the top K words with the highest similarity with the customized keyword.
- K is a positive integer greater than zero.
- the early warning module is further configured to perform real-time monitoring on each target keyword in the target keyword list in the form of a sliding window.
- the various modules in the above-mentioned topic early warning device may be implemented in whole or in part by software, hardware, and combinations thereof.
- the network interface may be an Ethernet card or a wireless network card.
- the above modules may be embedded in the hardware in the processor or in the memory in the server, or may be stored in the memory in the server, so that the processor calls the corresponding operations of the above modules.
- the processor can be a central processing unit (CPU), a microprocessor, a microcontroller, or the like.
- the storage medium may be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Business, Economics & Management (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Artificial Intelligence (AREA)
- Entrepreneurship & Innovation (AREA)
- Economics (AREA)
- Human Resources & Organizations (AREA)
- Marketing (AREA)
- Operations Research (AREA)
- Quality & Reliability (AREA)
- Strategic Management (AREA)
- Tourism & Hospitality (AREA)
- General Business, Economics & Management (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种话题预警的方法,包括:获取自定义关键词,计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词,根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表,根据所述目标关键词列表中的目标关键词进行实时监听,当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
Description
本申请要求于2017年4月7日提交中国专利局、申请号为2017102256853、发明名称为“话题预警的方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本发明涉及计算机处理领域,特别是涉及一种话题预警的方法、装置、计算机设备及存储介质。
随着社交媒体的发展,社交网站、在线社区、微博等已逐渐成为人们生活中不可或缺的一部分,也是当今时代信息传播的主要渠道,与此同时,社交媒体也是舆情传播的重要途径。通过对社交媒体的话题监听预警,能够为决策者提供科学化的信息支持。传统的对社交媒体话题监听预警是通过对获取到的历史数据进行分析,然后针对不同的话题进行标签分级。由于话题更新速度非常快,仅仅针对历史数据进行分析得出的结果显然不够准确,且传统的话题监听是针对所有的话题进行监听,没有考虑到用户的个性化需求。
发明内容
根据本申请的各种实施例,提供一种话题预警的方法、装置、计算机设备及存储介质。
一种话题预警的方法,包括:
获取自定义关键词;
计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相
似度从语料库中获取与所述自定义关键词相关的扩展关键词;
根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;
根据所述目标关键词列表中的目标关键词进行实时监听;及
当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
一种话题预警的装置,包括:
自定义关键词获取模块,用于获取自定义关键词;
扩展关键词获取模块,用于计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;
目标关键词筛选模块,用于根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;
监听模块,用于根据所述目标关键词列表中的目标关键词进行实时监听;
及
预警模块,用于当监听所述目标关键词所对应的话题量达到预设阈值时,进行话题预警。
一种计算机设备,包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被所述处理器执行时,使得所述处理器执行以下步骤:
获取自定义关键词;
计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;
根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;
根据所述目标关键词列表中的目标关键词进行实时监听;及
当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
一个或多个存储有计算机可读指令的非易失性可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行以下步骤:
获取自定义关键词;
计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;
根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;
根据所述目标关键词列表中的目标关键词进行实时监听;及
当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
本发明的一个或多个实施例的细节在下面的附图和描述中提出。本发明的其它特征、目的和优点将从说明书、附图以及权利要求书变得明显。
为了更清楚地说明本发明实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为一个实施例中终端的内部结构框图;
图2为一个实施例中服务器的内部结构框图;
图3为一个实施例中话题预警的方法流程图;
图4为一个实施例中根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词的方法流程图;
图5为另一个实施例中话题预警的方法流程图;
图6为一个实施例中计算自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取扩展关键词的方法流程图;
图7为一个实施例中话题预警的装置结构框图;
图8为一个实施例中目标关键词筛选模块的结构框图;
图9为另一个实施例中话题预警的装置结构框图。
为了使本发明的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本发明进行进一步详细说明。应当理解,此处所描述的具体实施例仅仅用以解释本发明,并不用于限定本发明。
如图1所示,在一个实施例中,终端102的内部结构如图1所示,包括通过系统总线连接的处理器、非易失性存储介质、内存储器、网络接口、显示屏和输入装置。其中,终端102的处理器用于提供计算和控制能力,支撑整个终端102的运行。非易失性存储介质存储有操作系统和计算机可读指令,该计算机可读指令可被处理器执行以实现适用于终端102的一种话题预警的方法。终端102中的内存储器为非易失性存储介质中的操作系统和计算机可读指令的运行提供环境。网络接口用于连接到网络进行通信。终端102的显示屏可以是液晶显示屏或者电子墨水显示屏等,输入装置可以是显示屏上覆盖的触摸层,也可以是电子设备外壳上设置的按键、轨迹球或触控板,也可以是外接的键盘、触控板或鼠标等。该终端102可以是平板电脑、笔记本电脑、台式计算机等。本领域技术人员可以理解,图1中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的终端的限定,具体的终端可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
如图2所示,在一个实施例中,服务器104的内部结构如图2所示,包括通过系统总线连接的处理器、非易失性存储介质、内存储器和网络接口。其中,该服务器104的处理器用于提供计算和控制能力,支撑整个服务器的运行。该非易失存储介质包括操作系统和计算机可读指令。该计算机可读指令可被处理器执行以实现适用于服务器104的一种话题预警的方法该服务器104的内存储器为非易失性存储介质中的操作系统和计算机可读指令的运行
提供环境,该服务器的网络接口用于与外部的服务器和终端通过网络连接通信。本领域技术人员可以理解,图2中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的服务器的限定,具体的服务器可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
如图3所示,在一个实施例中,提出了一种话题预警的方法,该方法可应用于计算机设备中,其中,计算机设备可以是终端,也可以是服务器,具体包括以下步骤:
步骤302,获取自定义关键词。
在本实施例中,自定义关键词是指用户给出的符合用户监听需求的关键词。为了能够满足用户的个性化的监听需求,监听关键词的设定是根据用户自定义关键词来设定的。由于大数据时代的社交媒体信息错综复杂,主体多种多样,而不同的用户所关注的话题不尽相同,其中,话题是指谈论的主体,由于不同人关注的主体不同,所以通过自定义关键词不仅能带来友好的用户交互,更多的是能够实现用户监听需求的个性化以及多元化。
步骤304,计算自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与自定义关键词相关的扩展关键词。
在本实施例中,由于用户给定的自定义关键词往往不够完整和全面,因此有必要对该自定义关键词进行一定的扩展。获取与该自定义关键词相关的扩展关键词,有利于保证用户对所需要监听的话题更加全面和完整,从而保证监听结果的完整性和多样性。通过计算自定义关键词与语料库中每个词语之间的相似度,从语料库中选取与自定义关键词相似度比较大的词语作为扩展关键词。相似度越大,说明该词语与自定义关键词的语义越相近。词语相似度的计算方法有多种,比如,可以采用同义词词林的方式计算词语之间的相似度,也可以采用皮尔森相关系数来计算词语之间的相似度。这里并不对词语相似度的计算方法进行限定。
在一个实施例中,相似度的计算是通过计算词向量之间的相似度得到的。
首先,采用word2vec模型计算自定义关键词对应的词向量,其中,word2vec是一款将词表征为实数值向量的高效工具,其利用深度学习的思想,可以通过训练,把对文本内容的处理简化为k维向量空间中的向量运算,而向量空间上的相似度可以用来表示文本语义上的相似度。具体地,将自定义关键词作为word2vec模型的输入,输出该自定义关键词的词向量表示。获取到自定义关键词的词向量表示之后,通过计算词向量之间的相似度从语料库中筛选出自定义关键词的扩展关键词。为了能够更快的获取到与自定义关键词相关的扩展关键词,可以将语料库中的词语均以词向量的形式存储。在一个实施例中,采用皮尔森相关系数(Pearson Correlation Coefficient)来计算词向量之间的相似度。假设自定义关键词的向量表示为W=(w1,w2,…,wn),语料库中任一词语的向量表示为X=(x1,x2,…,xn),那么它们之间的相似度s(W,X)为:
其中,n表示词向量的第n个词向量特征,i表示词向量中的第i个词向量特征。通过计算自定义关键词与语料库中每个词语的相似度筛选出与自定义关键词相关的扩展关键词。具体地,可以将相似度按照从高到低的顺序进行排列,选出相似度最高的前k个词语作为自定义关键词的扩展关键词。将自定义关键词进行扩展,使得关键词更具多样性,保证了话题监听结果具有与相似关键词的对比性,便于为决策者提供更丰富的信息。
步骤306,根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词,加入目标关键词列表。
在本实施例中,如果对步骤204得到的扩展关键词全部监听,将会使得信息错杂冗乱。所以为了保证信息的清楚,需要对获取到的扩展关键词进行进一步的筛选。根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词的方法有多种。在一个实施例中,首先,将获取到的全部扩展关键词进行分类,然后从每一类中选取出与自定
义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数,将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。在另一个实施例中,首先,获取全部扩展词对应的类型,然后将相同类型的关键词分为一组。分别获取每一类扩展关键词对应的扩展词数目,以扩展词数目最少的类型为基准,假设扩展词数目最少的类型对应的数目为X个,那么分别从其他每一类型中也筛选出X个扩展关键词作为目标关键词,其中,从其他每一类型中筛选出X个扩展关键词是根据相似度的大小进行筛选的,分别筛选出其他每一类扩展关键词中相似度最高的前X个扩展关键词作为目标关键词,加入目标关键词列表。
步骤308,根据目标关键词列表中的目标关键词进行实时监听。
在本实施例中,当确定了目标关键词列表后,根据目标关键词列表中的目标关键词进行实时监听。由于社交媒体数据每时每刻都在产生,迅速而规模庞大,形成了庞大的网络数据流。为了更好的对话题进行监听,可以采用基于滑动窗口的时序管理框架。基于滑动窗口的时序管理框架的主要思想是:对于目标监听列表中的每一个目标关键词,以滑动窗口的形式对话题数据流进行管理,每个目标关键词维护一个一定大小的缓存,每过一个时间片(为了实时监听,时间片的设置通常很小,比如5分钟),数据窗口进行滑动,然后对缓存中的数据进行处理。
步骤310,当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
在本实施例中,良好的监听必定需要预警,通过监听目标关键词所对应的话题量是否达到预设阈值,对话题进行预警。预警可以从两个方面来进行考虑,第一,对预设的时间片内的话题量进行监听预警。由于时间片是一个较短的时间,所以通过对短时间内的话题监听,能够对短时间内的突发事件进行预警。第二,对于一段时间段的话题进行预警,很多时候事件的发生或舆情的走势并不一定是急剧的,因此,考察一段时间内话题的热点能够帮助决策者发现事件的兴起或舆情的逐渐走势。具体地,采用两种评价策略进行
关键词的实时预警,一种是采用话题热度进行预警,通过分析大量的关键词的热度变化趋势及其生命周期,以经验的方式确定热度临界阈值,当监听的目标关键词在一个滑动窗口的时间片内出现的频率大于该热度临界阈值时,进行预警响应。一种是采用情感极性比率进行预警,对监听的目标关键词列表相关的社会网络文本进行情感极性分析,主要包括正面、中性和负面三个方面的情感极性,当负面情感在所有该目标关键词对应的话题量中占的比率大于情感极性阈值时,进行预警。该话题预警的方法可以应用于很多领域,尤其是可以应用于金融领域。以应用于金融产品为例,说明一下该话题预警的益处。首先,互联网与金融产业息息相关,根据对互联网数据的监控可以为金融产品避免诸多损失。其次,与金融相关的关键词比较有规律,而且相对比较固定,通过对金融产品相关的话题进行监听预警,可以实现快速响应而不失准确率。
在本实施例中,通过获取用户自定义关键词,然后在语料库中根据相似度对该自定义关键词进行扩展,获取相关的扩展关键词,再根据扩展关键词的类型和相似度进行筛选,筛选出最终用于监听的目标关键词,之后在社交媒体上根据该目标关键词进行实时监听,当监听到目标关键词的话题量达到预设阈值时,进行话题预警。该方法不仅能够实时对话题进行监听,而且可以基于用户自定义的关键词有针对性的进行监控,满足了用户的个性化监听预警的需求。通过对用户所要监控的自定义关键词进行扩展和筛选,保证了监听的多样性和全面性。
如图4所示,在一个实施例中,根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词,加入目标关键词列表的步骤包括:
步骤306A,将扩展关键词按照预设的类型进行分类。
在本实施例中,为了对基于自定义关键词的监听能够监听的更加全面和平衡化。首先,需要对扩展关键词按照预设的类型进行分类,比如,将扩展关键词按照“品牌”、“产品”、“竞品”分为三类。这样,便于后续针对每一
类挑选出相同个数的目标关键词进行监听,有利于保证监听信息的清楚全面和平衡。
步骤306B,从每一类的扩展关键词中筛选出与自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数。
在本实施例中,将扩展关键词按照预设的类型进行分类后,采用众包策略从每一类的扩展关键词中筛选出与自定义关键词相似度最高的前h个扩展关键词作为目标关键词。例如,从每一类中挑选出与自定义关键词相似度最高的前5个词语,最后将挑选出的每一类的目标关键词进行聚合。
步骤306C,将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
在本实施例中,通过从每一类的扩展关键词中筛选出与自定义关键词相似度最高的前h个扩展关键词作为目标关键词后,将每一类筛选出来的目标关键词聚集起来,放在同一张列表中,即生成目标关键词列表,后续便于根据该目标关键词列表中的目标关键词进行实时监听。比如,若将扩展关键词按照“品牌”、“产品”、“竞品”分为三类。若每一类都挑选出5个目标关键词,那么将总共挑选出15个目标关键词进行监听。通过将扩展关键词进行分类,然后再针对每一类进行筛选有利于监听的内容更加清晰和全面,不会出现偏激化的结果。
如图5所示,在一个实施例中,提出了一种话题预警的方法,该方法包括:
步骤502,获取自定义关键词。
步骤504,计算自定义关键词对应的词向量。
步骤506,计算自定义关键词的词向量与语料库中每个词语的词向量之间的相似度,根据词向量之间的相似度从语料库中获取与自定义关键词相关的扩展关键词。
步骤508,根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词,加入目标关键词列表。
步骤510,根据目标关键词列表中的目标关键词进行实时监听。
步骤512,当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
在本实施例中,当获取到自定义关键词后,为了后续计算词向量之间的相似度,首先需要计算该自定义关键词对应的词向量,通过将自定义关键词作为word2vec模型的输入,生成与该自定义关键词对应的词向量并输出。为了监听的更加全面,需要对自定义关键词进行扩展,即找出相关的与该自定义关键词语义相近的词语表示。通过计算自定义关键词与语料库中的每个词语之间的相似度来获取与自定义关键词相关的扩展关键词,其中,相似度越高,说明与自定义关键词的语义越相近。具体地,可以采用皮尔森相关系数(Pearson Correlation Coefficient)方法计算自定义关键词的词向量与语料库中每个词语的词向量之间的相似度,从中挑选出与自定义关键词相似度最高的前K个(比如,设K=50)词语作为扩展关键词。如果对挑选出来的扩展关键词全部进行监听,将会使得信息显得冗杂,为了解决这一问题,还需要对挑选出来的扩展关键词进行进一步的筛选。基于众包策略对扩展关键词进行进一步的筛选,首先对挑选出来的扩展关键词进行分类,比如,按照“品牌”、“产品”、“竞品”分为三类。分类完成后,针对每一类,根据之前计算得到的每个扩展关键词与自定义关键词之间的相似度,每一类选出与自定义关键词相似度最高的前h个词语作为目标关键词,然后将每一类筛选出来的目标关键词进行汇总,放在同一个列表中,即都加入目标关键词列表。之后根据该目标关键词列表进行监听,并进行相应的预警。该方法通过对用户自定义关键词进行扩展,保证了监听的多样性和全面性,结合众包技术对扩展关键词进行进一步甄选保证了监听结果不具有偏激化。
如图6所示,在一个实施例中,计算自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与自定义关键词相关的扩展关键词的步骤包括:
步骤304A,采用皮尔森相关系数方法计算自定义关键词与语料库中每个
词语之间的相似度。
在本实施例中,为了对自定义关键词进行扩展,找出与自定义关键词语义相近的扩展关键词,通过采用皮尔森相关系数方法来计算自定义关键词与语料库中每个词语之间的相似度。相似度越大,语义越相近。具体地,首先,获取自定义关键词的词向量表示,可以通过word2vec方法计算得到。然后计算自定义关键词的词向量与语料库中词语的词向量之间的相似度。为了能够更加快捷的计算自定义关键词与语料库中词语之间的相似度,在语料库中,词语是以词向量的形式存在的。假设自定义关键词的词向量表示为W=(w1,w2,…,wn),语料库中任一词语的词向量表示为X=(x1,x2,…,xn),那么它们之间的相似度s(W,X)为:
步骤304B,获取与自定义关键词相似度最高的前K个词语作为自定义关键词的扩展关键词,其中,K为大于0的正整数。
在本实施例中,显然,对自定义关键词进行无限扩展是不切实际的,所以需要从语料库中筛选出相似度比较大的词语作为扩展关键词。具体地,采用贪心策略选择与自定义关键词相似度最高的前K个词语作为自定义关键词的扩展,设扩展关键词集合为ES(W),那么ES(W)={X|s(W,X)≥s(W,Xk)},其中,W表示自定义关键词,Xk表示与自定义关键词相似度第K大的词汇,比如,可以设置K=50,即选取与自定义关键词相似度最高的前50个词汇作为其扩展关键词集合。
在一个实施例中,根据目标关键词列表中的目标关键词进行实时监听的步骤包括:采用滑动窗口的形式对目标关键词列表中的每一个目标关键词进行实时监听。
在本实施例中由于社交媒体数据每时每刻都在产生,且迅速而规模庞大,为了达到对话题进行实时监听,需要解决如何在数据流的环境下进行话题的
实时监听。在该实施例中,通过采用基于滑动窗口的形式对目标关键词列中的每一个目标关键词进行实时监听。即以滑动窗口的形式对话题数据流进行管理,每个目标关键词维护一个一定大小的缓存,每过一个时间片,数据窗口进行滑动,然后对缓存中的数据进行处理,从而实现了对每个目标关键词进行实时监听。
如图7所示,在一个实施例中,提出了一种话题预警的装置700,该装置包括:
自定义关键词获取模块702,用于获取自定义关键词。
扩展关键词获取模块704,用于计算自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与自定义关键词相关的扩展关键词。
目标关键词筛选模块706,用于根据扩展关键词的类型和扩展关键词与自定义关键词之间的相似度从扩展关键词中筛选出目标关键词,加入目标关键词列表。
监听模块708,用于根据目标关键词列表中的目标关键词进行实时监听。
预警模块710,用于当监听目标关键词所对应的话题量达到预设阈值时,进行话题预警。
如图8所示,在一个实施例中,目标关键词筛选模块706包括:
分类模块706A,用于将扩展关键词按照预设的类型进行分类。
筛选模块706B,用于从每一类的扩展关键词中筛选出与自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数。
聚合模块706C,用于将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
如图9所示,在一个实施例中,提出了一种话题预警的装置900,除了包括上述模块702-710,还包括:
计算模块703,用于计算自定义关键词对应的词向量。
扩展关键词获取模块704还用于计算自定义关键词的词向量与语料库中
每个词语的词向量之间的相似度,根据词向量之间的相似度从语料库中获取与自定义关键词相关的扩展关键词。
在一个实施例中,扩展词获取模块还用于采用皮尔森相关系数方法计算自定义关键词与语料库中每个词语之间的相似度,获取与自定义关键词相似度最高的前K个词语作为自定义关键词的扩展关键词,其中,K为大于0的正整数。
在一个实施例中,预警模块还用于采用滑动窗口的形式对目标关键词列表中的每一个目标关键词进行实时监听。
上述话题预警的装置中的各个模块可全部或部分通过软件、硬件及其组合来实现。其中,网络接口可以是以太网卡或无线网卡等。上述各模块可以硬件形式内嵌于或独立于服务器中的处理器中,也可以以软件形式存储于服务器中的存储器中,以便于处理器调用执行以上各个模块对应的操作。该处理器可以为中央处理单元(CPU)、微处理器、单片机等。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,该计算机程序可存储于一计算机可读取存储介质中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,前述的存储介质可为磁碟、光盘、只读存储记忆体(Read-Only Memory,ROM)等非易失性存储介质,或随机存储记忆体(Random Access Memory,RAM)等。
以上所述实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上所述实施例仅表达了本发明的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对发明专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本发明构思的前提下,还可以做出若干变形和改进,这些都属于本发明的保护范围。因此,本发明专利的保护范围应以所附权利要求为准。
Claims (20)
- 一种话题预警的方法,包括:获取自定义关键词;计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;根据所述目标关键词列表中的目标关键词进行实时监听;及当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
- 根据权利要求1所述的方法,其特征在于,所述根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表包括:将所述扩展关键词按照预设的类型进行分类;从每一类的扩展关键词中筛选出与所述自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数;将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
- 根据权利要求1所述的方法,其特征在于,在获取自定义关键词之后还包括:计算所述自定义关键词对应的词向量;所述计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词包括:计算自定义关键词的词向量与所述语料库中每个词语的词向量之间的相似度;根据词向量之间的相似度从语料库中获取与所述自定义关键词相关的扩展关键词。
- 根据权利要求1所述的方法,其特征在于,所述计算所述自定义关键 词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与所述自定义关键词相关的扩展关键词包括:采用皮尔森相关系数方法计算所述自定义关键词与语料库中每个词语之间的相似度;获取与所述自定义关键词相似度最高的前K个词语作为所述自定义关键词的扩展关键词,其中,K为大于0的正整数。
- 根据权利要求1所述的方法,其特征在于,所述根据所述目标关键词列表中的目标关键词进行实时监听包括:采用滑动窗口的形式对所述目标关键词列表中的每一个目标关键词进行实时监听。
- 一种话题预警的装置,包括:自定义关键词获取模块,用于获取自定义关键词;扩展关键词获取模块,用于计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;目标关键词筛选模块,用于根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;监听模块,用于根据所述目标关键词列表中的目标关键词进行实时监听;及预警模块,用于当监听所述目标关键词所对应的话题量达到预设阈值时,进行话题预警。
- 根据权利要求6述的装置,其特征在于,所述目标关键词筛选模块包括:分类模块,用于将所述扩展关键词按照预设的类型进行分类;筛选模块,用于从每一类的扩展关键词中筛选出与所述自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数;聚合模块,用于将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
- 根据权利要求6述的装置,其特征在于,所述装置还包括:计算模块,用于计算所述自定义关键词对应的词向量;扩展关键词获取模块还用于计算自定义关键词的词向量与所述语料库中每个词语的词向量之间的相似度,根据词向量之间的相似度从语料库中获取与所述自定义关键词相关的扩展关键词。
- 根据权利要求6所述的装置,其特征在于,所述扩展词获取模块还用于采用皮尔森相关系数方法计算所述自定义关键词与语料库中每个词语之间的相似度,获取与所述自定义关键词相似度最高的前K个词语作为所述自定义关键词的扩展关键词,其中,K为大于0的正整数。
- 根据权利要求6所述的装置,其特征在于,所述预警模块还用于采用滑动窗口的形式对所述目标关键词列表中的每一个目标关键词进行实时监听。
- 一种计算机设备,包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被所述处理器执行时,使得所述处理器执行以下步骤:获取自定义关键词;计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;根据所述目标关键词列表中的目标关键词进行实时监听;及当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
- 根据权利要求11所述的计算机设备,其特征在于,所述处理器所执行的所述根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表的步骤包括:将所述扩展关键词按照预设的类型进行分类;从每一类的扩展关键词中筛选出与所述自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数;将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
- 根据权利要求11所述的计算机设备,其特征在于,在获取自定义关键词的步骤之后,所述处理器还用于执行以下步骤:计算所述自定义关键词对应的词向量;所述计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词包括:计算自定义关键词的词向量与所述语料库中每个词语的词向量之间的相似度;根据词向量之间的相似度从语料库中获取与所述自定义关键词相关的扩展关键词。
- 根据权利要求11所述的计算机设备,其特征在于,所述处理器所执行的所述计算所述自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与所述自定义关键词相关的扩展关键词的步骤包括:采用皮尔森相关系数方法计算所述自定义关键词与语料库中每个词语之间的相似度;获取与所述自定义关键词相似度最高的前K个词语作为所述自定义关键词的扩展关键词,其中,K为大于0的正整数。
- 根据权利要求11所述的计算机设备,其特征在于,所述处理器所执行的所述根据所述目标关键词列表中的目标关键词进行实时监听的步骤包括:采用滑动窗口的形式对所述目标关键词列表中的每一个目标关键词进行实时监听。
- 一个或多个存储有计算机可读指令的非易失性可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执 行以下步骤:获取自定义关键词;计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词;根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表;根据所述目标关键词列表中的目标关键词进行实时监听;及当监听到目标关键词所对应的话题量达到预设阈值时,进行话题预警。
- 根据权利要求16所述的非易失性可读存储介质,其特征在于,所述处理器所执行的所述根据所述扩展关键词的类型和所述扩展关键词与所述自定义关键词之间的相似度从所述扩展关键词中筛选出目标关键词,加入目标关键词列表的步骤包括:将所述扩展关键词按照预设的类型进行分类;从每一类的扩展关键词中筛选出与所述自定义关键词相似度最高的前h个扩展关键词作为目标关键词,其中,h为大于0的正整数;将每一类筛选出来的目标关键词进行聚合,生成用于监听的目标关键词列表。
- 根据权利要求16所述的非易失性可读存储介质,其特征在于,在获取自定义关键词的步骤之后,所述处理器还用于执行以下步骤:计算所述自定义关键词对应的词向量;所述计算所述自定义关键词与语料库中每个词语之间的相似度,根据所述相似度从语料库中获取与所述自定义关键词相关的扩展关键词包括:计算自定义关键词的词向量与所述语料库中每个词语的词向量之间的相似度;根据词向量之间的相似度从语料库中获取与所述自定义关键词相关的扩展关键词。
- 根据权利要求16所述的非易失性可读存储介质,其特征在于,所述 处理器所执行的所述计算所述自定义关键词与语料库中每个词语之间的相似度,根据相似度从语料库中获取与所述自定义关键词相关的扩展关键词的步骤包括:采用皮尔森相关系数方法计算所述自定义关键词与语料库中每个词语之间的相似度;获取与所述自定义关键词相似度最高的前K个词语作为所述自定义关键词的扩展关键词,其中,K为大于0的正整数。
- 根据权利要求16所述的非易失性可读存储介质,其特征在于,所述处理器所执行的所述根据所述目标关键词列表中的目标关键词进行实时监听的步骤包括:采用滑动窗口的形式对所述目标关键词列表中的每一个目标关键词进行实时监听。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/090,351 US11205046B2 (en) | 2017-04-07 | 2017-06-28 | Topic monitoring for early warning with extended keyword similarity |
| SG11201809697YA SG11201809697YA (en) | 2017-04-07 | 2017-06-28 | Topic alarm method, device, computer apparatus, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710225685.3 | 2017-04-07 | ||
| CN201710225685.3A CN107168943B (zh) | 2017-04-07 | 2017-04-07 | 话题预警的方法和装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018184306A1 true WO2018184306A1 (zh) | 2018-10-11 |
Family
ID=59849735
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/090579 Ceased WO2018184306A1 (zh) | 2017-04-07 | 2017-06-28 | 话题预警的方法、装置、计算机设备及存储介质 |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US11205046B2 (zh) |
| CN (1) | CN107168943B (zh) |
| SG (1) | SG11201809697YA (zh) |
| TW (1) | TWI663520B (zh) |
| WO (1) | WO2018184306A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109684483A (zh) * | 2018-12-11 | 2019-04-26 | 平安科技(深圳)有限公司 | 知识图谱的构建方法、装置、计算机设备及存储介质 |
| CN110427492A (zh) * | 2019-07-10 | 2019-11-08 | 阿里巴巴集团控股有限公司 | 生成关键词库的方法、装置和电子设备 |
| CN112650791A (zh) * | 2020-12-29 | 2021-04-13 | 招联消费金融有限公司 | 字段处理方法、装置、计算机设备和存储介质 |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107862015A (zh) * | 2017-10-30 | 2018-03-30 | 北京奇艺世纪科技有限公司 | 一种关键词关联扩展方法和装置 |
| TWI716761B (zh) * | 2018-11-08 | 2021-01-21 | 鯨動智能科技股份有限公司 | 智能會計帳務系統與會計憑證的辨識入帳方法 |
| CN109635286B (zh) * | 2018-11-26 | 2022-04-12 | 平安科技(深圳)有限公司 | 政策热点分析的方法、装置、计算机设备和存储介质 |
| CN110457672B (zh) * | 2019-06-25 | 2023-01-17 | 平安科技(深圳)有限公司 | 关键词确定方法、装置、电子设备及存储介质 |
| CN111859013B (zh) * | 2020-07-17 | 2024-11-19 | 腾讯音乐娱乐科技(深圳)有限公司 | 数据处理方法、装置、终端和存储介质 |
| CN114528406A (zh) * | 2022-02-23 | 2022-05-24 | 国泰新点软件股份有限公司 | 一种文本事件确定方法、装置、电子设备及存储介质 |
| CN115545022A (zh) * | 2022-09-30 | 2022-12-30 | 加和(北京)信息科技有限公司 | 词联想方法及装置、存储介质、计算设备 |
| CN116681086B (zh) * | 2023-07-31 | 2024-04-02 | 深圳市傲天科技股份有限公司 | 数据分级方法、系统、设备及存储介质 |
| CN118780280B (zh) * | 2024-09-06 | 2024-11-26 | 浙江寻常问道网络信息科技有限公司 | 一种多平台融合的智能选题灵感生成方法及选题灵感引擎 |
| CN119089025B (zh) * | 2024-11-06 | 2025-03-25 | 广东东软学院 | 一种舆情预警方法、装置、设备及存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102194001A (zh) * | 2011-05-17 | 2011-09-21 | 杭州电子科技大学 | 网络舆情危机预警方法 |
| CN103853720A (zh) * | 2012-11-28 | 2014-06-11 | 苏州信颐系统集成有限公司 | 基于用户关注度的网络敏感信息监控系统及方法 |
| CN104408157A (zh) * | 2014-12-05 | 2015-03-11 | 四川诚品电子商务有限公司 | 一种网络舆情漏斗式数据采集分析推送系统及方法 |
| CN104516903A (zh) * | 2013-09-29 | 2015-04-15 | 北大方正集团有限公司 | 关键词扩展方法及系统、及分类语料标注方法及系统 |
Family Cites Families (37)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7346492B2 (en) * | 2001-01-24 | 2008-03-18 | Shaw Stroz Llc | System and method for computerized psychological content analysis of computer and media generated communications to produce communications management support, indications, and warnings of dangerous behavior, assessment of media images, and personnel selection support |
| US20050102278A1 (en) * | 2003-11-12 | 2005-05-12 | Microsoft Corporation | Expanded search keywords |
| WO2006099621A2 (en) * | 2005-03-17 | 2006-09-21 | University Of Southern California | Topic specific language models built from large numbers of documents |
| US8898134B2 (en) * | 2005-06-27 | 2014-11-25 | Make Sence, Inc. | Method for ranking resources using node pool |
| IN2005MU00878A (zh) | 2005-07-22 | 2009-06-29 | ||
| US7627561B2 (en) * | 2005-09-12 | 2009-12-01 | Microsoft Corporation | Search and find using expanded search scope |
| TWI317488B (en) | 2005-11-04 | 2009-11-21 | Webgenie Information Ltd | Method for automatically detecting similar documents |
| US8280877B2 (en) * | 2007-02-22 | 2012-10-02 | Microsoft Corporation | Diverse topic phrase extraction |
| CN101295319B (zh) * | 2008-06-24 | 2010-06-02 | 北京搜狗科技发展有限公司 | 一种扩展查询的方法、装置及搜索引擎系统 |
| US9892103B2 (en) * | 2008-08-18 | 2018-02-13 | Microsoft Technology Licensing, Llc | Social media guided authoring |
| US7974983B2 (en) | 2008-11-13 | 2011-07-05 | Buzzient, Inc. | Website network and advertisement analysis using analytic measurement of online social media content |
| US8768960B2 (en) * | 2009-01-20 | 2014-07-01 | Microsoft Corporation | Enhancing keyword advertising using online encyclopedia semantics |
| US20110004465A1 (en) * | 2009-07-02 | 2011-01-06 | Battelle Memorial Institute | Computation and Analysis of Significant Themes |
| CN101751458A (zh) * | 2009-12-31 | 2010-06-23 | 暨南大学 | 一种网络舆情监控系统及方法 |
| CN102195899B (zh) * | 2011-05-30 | 2014-05-07 | 中国人民解放军总参谋部第五十四研究所 | 通信网络的信息挖掘方法与系统 |
| US8909643B2 (en) * | 2011-12-09 | 2014-12-09 | International Business Machines Corporation | Inferring emerging and evolving topics in streaming text |
| TW201324199A (zh) | 2011-12-13 | 2013-06-16 | Chunghwa Telecom Co Ltd | 一種基於相似度比對的內容分析方法 |
| CN103853722B (zh) * | 2012-11-29 | 2017-09-22 | 腾讯科技(深圳)有限公司 | 一种基于检索串的关键词扩展方法、装置和系统 |
| CN103268350B (zh) * | 2013-05-29 | 2017-02-08 | 安徽雷越网络科技有限公司 | 一种互联网舆情信息监测系统及监测方法 |
| CN104281607A (zh) * | 2013-07-08 | 2015-01-14 | 上海锐英软件技术有限公司 | 微博热点话题分析方法 |
| US20150213002A1 (en) * | 2014-01-24 | 2015-07-30 | International Business Machines Corporation | Personal emotion state monitoring from social media |
| US20150286627A1 (en) * | 2014-04-03 | 2015-10-08 | Adobe Systems Incorporated | Contextual sentiment text analysis |
| US10409912B2 (en) * | 2014-07-31 | 2019-09-10 | Oracle International Corporation | Method and system for implementing semantic technology |
| US20160062967A1 (en) * | 2014-08-27 | 2016-03-03 | Tll, Llc | System and method for measuring sentiment of text in context |
| US10417338B2 (en) * | 2014-09-02 | 2019-09-17 | Hewlett-Packard Development Company, L.P. | External resource identification |
| US10157225B2 (en) * | 2014-12-17 | 2018-12-18 | Bogazici Universitesi | Content sensitive document ranking method by analyzing the citation contexts |
| CN104573008B (zh) | 2015-01-08 | 2017-11-21 | 广东小天才科技有限公司 | 一种网络信息的监控方法及装置 |
| CN104915405B (zh) * | 2015-06-02 | 2018-10-23 | 华东师范大学 | 一种基于多层次的微博查询扩展方法 |
| CN104933183B (zh) * | 2015-07-03 | 2018-02-06 | 重庆邮电大学 | 一种融合词向量模型和朴素贝叶斯的查询词改写方法 |
| US9880999B2 (en) * | 2015-07-03 | 2018-01-30 | The University Of North Carolina At Charlotte | Natural language relatedness tool using mined semantic analysis |
| US10394953B2 (en) * | 2015-07-17 | 2019-08-27 | Facebook, Inc. | Meme detection in digital chatter analysis |
| CN105045875B (zh) * | 2015-07-17 | 2018-06-12 | 北京林业大学 | 个性化信息检索方法及装置 |
| US11068926B2 (en) * | 2016-09-26 | 2021-07-20 | Emm Patents Ltd. | System and method for analyzing and predicting emotion reaction |
| CN105631037B (zh) * | 2015-12-31 | 2019-02-22 | 北京恒冠网络数据处理有限公司 | 一种图像检索方法 |
| US20170213138A1 (en) * | 2016-01-27 | 2017-07-27 | Machine Zone, Inc. | Determining user sentiment in chat data |
| US9864743B2 (en) * | 2016-04-29 | 2018-01-09 | Fujitsu Limited | Textual emotion detection |
| US10558740B1 (en) * | 2017-03-13 | 2020-02-11 | Intuit Inc. | Serving different versions of a user interface in response to user emotional state |
-
2017
- 2017-04-07 CN CN201710225685.3A patent/CN107168943B/zh active Active
- 2017-06-28 US US16/090,351 patent/US11205046B2/en not_active Expired - Fee Related
- 2017-06-28 SG SG11201809697YA patent/SG11201809697YA/en unknown
- 2017-06-28 WO PCT/CN2017/090579 patent/WO2018184306A1/zh not_active Ceased
- 2017-11-28 TW TW106141314A patent/TWI663520B/zh not_active IP Right Cessation
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102194001A (zh) * | 2011-05-17 | 2011-09-21 | 杭州电子科技大学 | 网络舆情危机预警方法 |
| CN103853720A (zh) * | 2012-11-28 | 2014-06-11 | 苏州信颐系统集成有限公司 | 基于用户关注度的网络敏感信息监控系统及方法 |
| CN104516903A (zh) * | 2013-09-29 | 2015-04-15 | 北大方正集团有限公司 | 关键词扩展方法及系统、及分类语料标注方法及系统 |
| CN104408157A (zh) * | 2014-12-05 | 2015-03-11 | 四川诚品电子商务有限公司 | 一种网络舆情漏斗式数据采集分析推送系统及方法 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109684483A (zh) * | 2018-12-11 | 2019-04-26 | 平安科技(深圳)有限公司 | 知识图谱的构建方法、装置、计算机设备及存储介质 |
| CN110427492A (zh) * | 2019-07-10 | 2019-11-08 | 阿里巴巴集团控股有限公司 | 生成关键词库的方法、装置和电子设备 |
| CN110427492B (zh) * | 2019-07-10 | 2023-08-15 | 创新先进技术有限公司 | 生成关键词库的方法、装置和电子设备 |
| CN112650791A (zh) * | 2020-12-29 | 2021-04-13 | 招联消费金融有限公司 | 字段处理方法、装置、计算机设备和存储介质 |
| CN112650791B (zh) * | 2020-12-29 | 2023-12-26 | 招联消费金融有限公司 | 字段处理方法、装置、计算机设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| US11205046B2 (en) | 2021-12-21 |
| CN107168943A (zh) | 2017-09-15 |
| TWI663520B (zh) | 2019-06-21 |
| SG11201809697YA (en) | 2018-11-29 |
| US20210224481A1 (en) | 2021-07-22 |
| CN107168943B (zh) | 2018-07-03 |
| TW201837755A (zh) | 2018-10-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN107168943B (zh) | 话题预警的方法和装置 | |
| US10534635B2 (en) | Personal digital assistant | |
| US20220092446A1 (en) | Recommendation method, computing device and storage medium | |
| JP2021174516A (ja) | ナレッジグラフ構築方法、装置、電子機器、記憶媒体およびコンピュータプログラム | |
| US11609942B2 (en) | Expanding search engine capabilities using AI model recommendations | |
| US20180276553A1 (en) | System for querying models | |
| CN114996562A (zh) | 利用数据驱动分析确定数字角色 | |
| WO2023124029A1 (zh) | 深度学习模型的训练方法、内容推荐方法和装置 | |
| CN105283839A (zh) | 用以将命令显现在生产力应用用户界面内的个性化社区模型 | |
| CN106407425A (zh) | 基于人工智能的推送信息的方法和装置 | |
| CN116955817A (zh) | 内容推荐方法、装置、电子设备以及存储介质 | |
| CN111932308A (zh) | 数据推荐方法、装置和设备 | |
| CN113722593B (zh) | 事件数据处理方法、装置、电子设备和介质 | |
| CN110287313A (zh) | 一种风险主体的确定方法及服务器 | |
| US12056160B2 (en) | Contextualizing data to augment processes using semantic technologies and artificial intelligence | |
| CN109582967B (zh) | 舆情摘要提取方法、装置、设备及计算机可读存储介质 | |
| EP3304343A1 (en) | Systems and methods for providing a comment-centered news reader | |
| Kandanaarachchi et al. | Leave-one-out kernel density estimates for outlier detection | |
| CN113010769A (zh) | 基于知识图谱的物品推荐方法、装置、电子设备及介质 | |
| CN111414455B (zh) | 舆情分析方法、装置、电子设备及可读存储介质 | |
| CN116821493A (zh) | 消息推送方法、装置、计算机设备及存储介质 | |
| CN112818221B (zh) | 实体的热度确定方法、装置、电子设备及存储介质 | |
| CN105431841A (zh) | 跨模型过滤 | |
| CN119338554A (zh) | 基于大模型的商品对比方法、装置、电子设备及介质 | |
| CN116597443B (zh) | 素材标签处理方法、装置、电子设备及介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17904742 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 15/01/2020) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17904742 Country of ref document: EP Kind code of ref document: A1 |

