WO2019085355A1 - 互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质 - Google Patents
互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2019085355A1 WO2019085355A1 PCT/CN2018/077638 CN2018077638W WO2019085355A1 WO 2019085355 A1 WO2019085355 A1 WO 2019085355A1 CN 2018077638 W CN2018077638 W CN 2018077638W WO 2019085355 A1 WO2019085355 A1 WO 2019085355A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- news
- words
- clustering
- sentence
- reserved
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/35—Clustering; Classification
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/953—Querying, e.g. by the use of web search engines
- G06F16/9535—Search customisation based on user profiles and personalisation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/40—Business processes related to social networking or social networking services
- G06Q10/44—Identification of trends within social networks, e.g. identification of trending topics
Definitions
- the present application relates to the field of data analysis technologies, and in particular, to a public opinion clustering analysis method, an application server, and a computer readable storage medium for Internet news.
- the news hotspots of major Internet stations are constantly emerging. Grasping and identifying these news hotspots is of great significance to the release of Internet information and public opinion supervision.
- the present application proposes a public opinion clustering analysis method, an application server, and a computer readable storage medium for Internet news, to solve the problem of how to find a hot spot from the massive information of the Internet and present it to the user.
- the present application provides a public opinion clustering analysis method for Internet news, which method includes the following steps:
- the clustered news and the topic summary are output and displayed to the user.
- the present application further provides an application server, including a memory and a processor, where the memory stores a public opinion clustering analysis system for Internet news that can be run on the processor, the Internet news.
- the lyric cluster analysis system is implemented by the processor to implement the steps of the public opinion clustering analysis method of the Internet news as described above.
- the present application further provides a computer readable storage medium storing a public opinion clustering analysis system of Internet news, wherein the Internet news sensation clustering analysis system can be At least one processor executes the steps of causing the at least one processor to perform a public opinion clustering analysis method of Internet news as described above.
- the public opinion clustering analysis method, the application server and the computer readable storage medium of the Internet news proposed by the present application first obtain news information in an information source through a distributed crawler, and store it in a public opinion database; Secondly, denoising, segmenting, and clustering the data in the lyric database; then, summarizing the different types of news after clustering; and finally, outputting the clustered news and the topic summary, and Displayed to the user.
- the public opinion clustering analysis method, the application server and the computer readable storage medium of the Internet news proposed by the application can quickly obtain news on the network, cluster the acquired news to obtain hot news, and automatically obtain the news obtained. Keywords and abstracts are more convenient, faster and more accurate than the prior art.
- 1 is a schematic diagram of an optional hardware architecture of an application server of the present application
- FIG. 2 is a schematic diagram of a program module of an embodiment of a public opinion clustering analysis system of the Internet news of the present application;
- FIG. 3 is a flowchart of a first embodiment of a public opinion clustering analysis method for Internet news according to the present application
- FIG. 4 is a flowchart of a second embodiment of a public opinion clustering analysis method for Internet news according to the present application.
- FIG. 5 is a flowchart of a third embodiment of a public opinion clustering analysis method of the Internet news of the present application.
- FIG. 6 is a flowchart of a fourth embodiment of a public opinion clustering analysis method of the Internet news of the present application.
- FIG. 7 is a flowchart of a fifth embodiment of a public opinion clustering analysis method for Internet news according to the present application.
- FIG. 8 is a flowchart of a sixth embodiment of a public opinion clustering analysis method for Internet news according to the present application.
- FIG. 9 is a flow chart of a seventh embodiment of the public opinion clustering analysis method of the Internet news of the present application.
- FIG. 1 it is a schematic diagram of an optional hardware architecture of the application server 1 of the present application.
- the application server 1 may include, but is not limited to, the memory 11, the processor 12, and the network interface 13 being communicably connected to each other through a system bus. It is pointed out that Figure 1 only shows the application server 1 with components 11-13, but it should be understood that not all illustrated components may be implemented, and more or fewer components may be implemented instead.
- the application server 1 may be a computing device such as a rack server, a blade server, a tower server, or a rack server.
- the application server 1 may be an independent server or a server cluster composed of multiple servers. .
- the memory 11 includes at least one type of readable storage medium including a flash memory, a hard disk, a multimedia card, a card type memory (eg, SD or DX memory, etc.), a random access memory (RAM), a static Random access memory (SRAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), programmable read only memory (PROM), magnetic memory, magnetic disk, optical disk, and the like.
- the memory 11 may be an internal storage unit of the application server 1, such as a hard disk or memory of the application server 1.
- the memory 11 may also be an external storage device of the application server 1, such as a plug-in hard disk equipped on the application server 1, a smart memory card (SMC), and a secure digital number. (Secure Digital, SD) card, flash card, etc.
- SMC smart memory card
- SD Secure Digital
- the memory 11 can also include both the internal storage unit of the application server 1 and its external storage device.
- the memory 11 is generally used to store an operating system installed in the application server 1 and various types of application software, such as program code of the public opinion clustering analysis system 200 of Internet news. Further, the memory 11 can also be used to temporarily store various types of data that have been output or are to be output.
- the processor 12 may be a Central Processing Unit (CPU), controller, microcontroller, microprocessor, or other data processing chip in some embodiments.
- the processor 12 is typically used to control the overall operation of the application server 1.
- the processor 12 is configured to run program code or processing data stored in the memory 11, such as the public opinion clustering analysis system 200 that runs the Internet news.
- the network interface 13 may comprise a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the application server 1 and other electronic devices.
- the present application proposes a public opinion clustering analysis system 200 for Internet news.
- FIG. 2 it is a program module diagram of the first embodiment of the public opinion clustering analysis system 200 of the Internet news of the present application.
- the public opinion cluster analysis system 200 of the Internet news includes a series of computer program instructions stored on the memory 11, and when the computer program instructions are executed by the processor 12, embodiments of the present application may be implemented.
- the public opinion cluster analysis operation of Internet news can be divided into one or more modules based on the particular operations implemented by the various portions of the computer program instructions. For example, in FIG. 2, the public opinion cluster analysis system 200 of the Internet news may be divided into an acquisition module 21, a processing module 22, an induction module 23, and an output display module 24. among them:
- the obtaining module 21 is configured to obtain news class information in an information source through a distributed crawler, and store the information in the public opinion database.
- the distributed crawler can perform multi-task capture on webpages distributed on different servers, thereby improving information retrieval efficiency.
- the distributed crawler framework is mainly divided into two parts: a downloader and a parser.
- the downloader is responsible for crawling the web page
- the parser is responsible for parsing the web page into the library.
- the two rely on the message queue to communicate, the two can be distributed in different machines, but also distributed on the same machine.
- the number of the two is also flexible. For example, there may be five machines in the download and two machines in the analysis, which can be adjusted according to the status of the crawler system.
- the message queue has two pipes: an HTML/JS file and a seed to be crawled.
- the downloader takes a seed from the seed to be crawled, calls the corresponding crawl module to fetch the webpage according to the seed information, and then saves it into the HTML/JS file channel; the parser gets a webpage content from the HTML/JS file.
- the corresponding parsing module is called for parsing, the target field is put into the library, and a new seed to be crawled is parsed if necessary.
- the downloader includes a User-Agent pool, a Proxy pool, and a cookie pool, and can be adapted to the crawling of a complex website.
- the information source includes, but is not limited to, a news website, a microblog, a WeChat, a Post Bar, a forum, and the like.
- news websites include, but are not limited to, Sina, Netease, Phoenix News, Xinhuanet, People.com, Tencent.com and other major domestic news websites and local online newspapers.
- the processing module 22 is configured to perform denoising, word segmentation, and clustering on the data in the database.
- the inductive module 23 is configured to summarize the topic summary for different types of news after clustering.
- the news digest in the news event is the enrichment of the news content.
- the purpose is to further understand the important information related to the news after the user reads the news headline, in order to decide whether to further read the details of the news.
- Most of the users read the news using the mobile phone. Because the mobile phone screen is small, in order to maximize the information transmitted to the user with limited text, the duplicate information is minimized, and the summary of the different types of news after clustering is intelligently and automatically summarized. Therefore, it is presented to the user, which can save the user's event and enhance the user's experience.
- the output display module 24 is configured to output the clustered news and the topic summary to the user.
- the output display module 24 includes a display screen such as an LCD or an LED.
- the present application also proposes a public opinion clustering analysis method for Internet news.
- FIG. 3 it is a schematic flowchart of the first embodiment of the public opinion clustering analysis method of the Internet news of the present application.
- the order of execution of the steps in the flowchart shown in FIG. 3 may be changed according to different requirements, and some steps may be omitted.
- Step S110 Obtain news information in the information source through the distributed crawler, and store the information in the public opinion database.
- Step S120 performing denoising, word segmentation, and clustering on the data in the public opinion database.
- Step S130 summarizing the topic abstracts for different types of news after clustering.
- the induction method includes the steps of:
- Step 1 The clause of the body of the news is sentenced, and the sentence whose sentence length is within the preset length is reserved, and is recorded as a reserved sentence.
- the length of the sentence can be limited by this step, thereby defining the length of the title, and at the same time, selecting a sentence within the preset length range as a reserved sentence also has the advantage of being convenient to process.
- Step 2 Calculate the similarity S(s) of each reserved sentence with the title, and the weight Q(s) of each reserved sentence.
- the similarity between the reserved sentence and the title is introduced in order to make the similarity of the last selected abstract and the title low, and the weight of the sentence indicates the value of the sentence in the news, usually the more keywords the sentence contains, The greater its value.
- Step 1 synonym conversion of the reserved sentence and the title based on the synonym lexicon
- Step 2 calculating the reserved sentence and the title for the synonym conversion by Jaccard distance
- the similarity S(s) of the sentence and the title is reserved.
- the intersection of the phrase in the sentence and the title is divided by the union of the phrases to obtain the similarity S(s).
- Step 4 Select the reserved sentence with the highest sorting score as the summary of the same kind of news.
- the reserved sentence with the highest sorting score represents that the more the reserved sentence can represent the main content of the news, and a threshold may be set.
- the sentence corresponding to the sorting score is taken as a reserved sentence.
- step S140 the clustered news and the topic summary are output and displayed to the user.
- the step of denoising in the step of “denoising, segmenting, and clustering data in the database of the public opinion” includes the following steps:
- Step S210 filtering the picture, the copyright description, the advertisement, and obtaining the document information.
- the collected webpage information includes noise data such as advertisements, navigation information, pictures, copyright descriptions, etc. What is really needed for the analysis of public opinion information is meta information of the body part, clearing the irrelevant content, and retaining the document information in the collection webpage. .
- Step S220 filtering the stop words.
- the participle in the step of “denoising, word segmentation, and clustering data in the database of the public opinion” specifically includes the following steps:
- Step S310 using the Chinese word segmentation technology to segment the collected webpage text data.
- the Chinese word segmentation technique includes two types, and the first type of method applies dictionary matching, Chinese lexical or other Chinese language knowledge for word segmentation, such as: maximum matching method, minimum word segmentation method, and the like.
- the second type of statistical-based word segmentation method is based on the statistical information of words and words, such as the application of information between adjacent words, word frequency and corresponding co-occurrence information to the word segmentation.
- the Chinese word segmentation technique includes the following methods:
- the maximum forward matching method the basic idea is: assuming that the longest word in the word segmentation dictionary has i Chinese characters, the first i words in the current string of the processed document are used as matching fields to search for a dictionary. If such an i word exists in the dictionary, the match is successful and the matching field is segmented as a word. If such an i word is not found in the dictionary, the match fails, the last word in the matching field is removed, and the remaining strings are re-matched... so go on until the match is successful, ie, split The length of a word or remaining string is zero. This completes a round of matching and then takes an i-string to match until the document is scanned.
- the inverse maximum matching method starts the matching scan from the end of the processed document, and takes the last 2i characters (i word string) as the matching field each time. If the matching fails, the first one of the matching fields is removed. Word, continue to match.
- the word segmentation dictionary it uses is a reverse dictionary, in which each term is stored in reverse order. In the actual processing, the document is first inverted to generate a reverse-order document. Then, according to the reverse order dictionary, the reverse order document can be processed by the forward maximum matching method.
- FIG. 6 it is a schematic flowchart of a fourth embodiment of the public opinion clustering analysis method of the Internet news of the present application.
- the clustering in the step of “denoising, segmenting, and clustering the data in the database” includes the following steps:
- Step S410 setting a reference table including sensitive words and emotional words.
- the preset emotional words include adverbs, conjunctions, and opinion words with strong emotions.
- conjunctions include nothing, but, yes, and so on; adverbs include singular, perfect, almost, absolute, etc.; opinion words include perception, discovery, belief, assertion, conjecture, representation, thought, and so on.
- Sensitive words can be forbidden words, infringing words, indecent words, political and inflammatory words.
- step S420 keywords are obtained, and a keyword reference table is set according to the obtained keywords.
- the step may include the following steps:
- Step 1 Analyze the document after the word segmentation, count the frequency of occurrence of the word, the position of the occurrence (title, abstract, body, remarks), historical average frequency (if any);
- Step 2 Obtain the importance D of the word according to the formula
- a, b, c are the frequency, position, historical average frequency corresponding to the weight value at the time, Fn, Wi, Fh correspond to the frequency of occurrence of the word, the position of occurrence, the historical average frequency, and also the position where the word appears.
- Step 3 Sort the words according to the importance D, and use the words whose importance D is greater than the preset value as keywords and generate a keyword reference table.
- Step S430 analyzing keywords, sensitive words, and words with emotional inclinations in the webpage text against the sensitive words, emotional words, and keyword reference tables.
- the sensitive words, the emotional words and the keywords are analyzed, the keywords, their synonyms and synonyms are uniformly listed as keywords, and the synonym of the sensitive words and the emotional words are also compared and analyzed.
- the sensitive words, keywords and sentiment words located in the database server lexicon can also be added to the vocabulary by the administrator according to the needs of the development of the social lyrics, so as to realize the real-time update of the vocabulary.
- the filtering and early warning of the webpage information is realized (that is, the webpage text information is marked according to the keyword tag or according to the sentiment orientation word mark), and the processing result is submitted to the sensitive word processing module.
- Sensitive word processing; sensitive word processing module is used to filter and block sensitive words in text data processed by the keyword processing module according to relevant national laws and regulations, and the processing result is processed by the cluster analysis module, wherein the sensitive The word comes from the processing results provided by the word segmentation module.
- Step S440 automatically clustering the obtained webpage data according to the category of the webpage according to the keyword, the emotional word, and the sensitive word.
- the news webpages in the database are classified according to the obtained keywords, for example, political, economic, military, social, technology, games, fashion, sports, movies, etc.; if the keywords are geographical names, they may also be based on different geographical names. Classification, for example, Beijing, Shanghai, Guangzhou, Shenzhen, Hong Kong, Taiwan, etc.;
- the news webpages in the database can be classified according to the emotional words, for example, can be divided into love, affection, friendship, lovelorn, cure, and the like.
- news pages in the database can be classified according to sensitive words, such as classification of sensitive policies, sensitive areas, and sensitive people.
- the categories after clustering contain groups of multiple news.
- step S450 the clustered news is sorted according to the heat.
- the number of comments, the number of clicks, and the number of forwards of the news can also be obtained, and sorted according to the above data.
- the heat may be calculated according to the number of the same or similar news, and the different news of different clusters are separately calculated and ranked according to the heat.
- FIG. 7 it is a schematic flowchart of a fifth embodiment of the public opinion clustering analysis method of the Internet news of the present application.
- the step of the public opinion clustering analysis method of the Internet news "obtaining the keyword, and setting a keyword reference table according to the obtained keyword” specifically includes:
- Step S510 analyzing the document after the word segmentation, and counting the frequency of occurrence of the word, the position of occurrence, and the historical average frequency.
- Step S520 obtaining the importance D of the word according to the following formula:
- Step S530 sorting the words according to the importance D, using the words whose importance D is greater than the preset value as the keyword and generating the keyword reference table.
- a, b, c are the weights corresponding to the frequency, position and historical average frequency at the time of the word; Fn, Wi, Fh correspond to the frequency of occurrence of the word, the position of occurrence, and the historical average frequency.
- FIG. 8 it is a schematic flowchart of a sixth embodiment of the public opinion clustering analysis method of the Internet news of the present application.
- the step of the public opinion clustering analysis method of the Internet news “individing the summary of the different types of news after clustering” specifically includes:
- Step S610 the clause of the body of the news is sentenced, and the sentence whose sentence length is within the preset length is reserved, and is recorded as a reserved sentence.
- the length of the sentence can be defined by this step, thereby defining the length of the title.
- Step S620 respectively calculating the similarity S(s) of the reserved sentence to the title, and the weight Q(s) of the reserved sentence.
- the similarity between the reserved sentence and the title is introduced so that the similarity between the last selected abstract and the title is low, and the weight of the sentence indicates the value of the sentence in the news, usually the more keywords the sentence contains, Then the value is greater.
- Step S640 selecting the reserved sentence with the highest sorting score as a summary of the same kind of news.
- the step of calculating the similarity S(s) of the reserved sentence and the title, and the weight Q(s) of the reserved sentence respectively, in the step of analyzing the public opinion clustering analysis method of the Internet news specifically includes:
- Step S710 synchronizing the reserved sentence and the title based on the synonym vocabulary.
- Step S720 calculating the similarity S(s) of the reserved sentence and the title by using the Jaccard distance for the reserved sentence and the title after the synonym conversion.
- the similarity S(s) of the reserved sentence and the title is calculated by using the Jaccard distance for the reserved sentence and the title after the synonym conversion, that is, the intersection of the reserved sentence and the phrase in the title is divided by the union of the phrase to obtain the similarity S. (s).
- the technical solution of the present application which is essential or contributes to the prior art, may be embodied in the form of a software product stored in a storage medium (such as ROM/RAM, disk,
- a storage medium such as ROM/RAM, disk
- the optical disc includes a number of instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to perform the methods described in the various embodiments of the present application.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种互联网新闻的舆情聚类分析方法,该方法包括:通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中(S110);对所述舆情数据库中的数据进行去噪、分词、聚类(S120);对聚类后的不同类新闻分别归纳主题摘要(S1130);及将聚类后的新闻及所述主题摘要输出,并显示给用户(S140)。还提供一种应用服务器及计算机可读存储介质。提供的互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质能够快速获得网络上的新闻,对获取的新闻进行聚类获取热点新闻,并且可以对获取的新闻进行自动关键词及摘要,相较于现有技术,更加方便、快捷、准确。
Description
本申请要求于2017年11月01日提交中国专利局、申请号为201711060246.8、发明名称为“互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质”的中国专利申请的优先权,其全部内容通过引用结合在申请中。
本申请涉及数据分析技术领域,尤其涉及一种互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质。
随着Internet的迅猛发展,网络信息已经成为人们生活中必不可少的一部分,目前中国网民数量已经超过2亿,中国网页数量也超过了80亿。网络媒体已被公认为继报纸、广播和电视之后的"第四媒体",网络成为反应社会舆情的主要载体之一。网络舆情与社会舆情相互作用、相互影响,网络舆情与社会舆情在内容表现形态方面具有一致性,网络舆情一定程度上会影响社会舆情的发展趋势,因此网络舆情热点话题的发现具有十分重要的意义。
而各大互联网站的新闻热点层出不穷,抓取并找出这些新闻舆情热点,对互联网信息发布、舆情监督等都有重要意义。
因此,如何从互联网的海量信息中发现热点并呈现给用户,成为当下亟需解决的一大问题。
发明内容
有鉴于此,本申请提出一种互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质,以解决如何从互联网的海量信息中发现热点并呈 现给用户的问题。
首先,为实现上述目的,本申请提出一种互联网新闻的舆情聚类分析方法,该方法包括步骤:
通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中;
对所述舆情数据库中的数据进行去噪、分词、聚类;
对聚类后的不同类新闻分别归纳主题摘要;及
将聚类后的新闻及所述主题摘要输出,并显示给用户。
此外,为实现上述目的,本申请还提供一种应用服务器,包括存储器、处理器,所述存储器上存储有可在所述处理器上运行的互联网新闻的舆情聚类分析系统,所述互联网新闻的舆情聚类分析系统被所述处理器执行时实现如上述的互联网新闻的舆情聚类分析方法的步骤。
进一步地,为实现上述目的,本申请还提供一种计算机可读存储介质,所述计算机可读存储介质存储有互联网新闻的舆情聚类分析系统,所述互联网新闻的舆情聚类分析系统可被至少一个处理器执行,以使所述至少一个处理器执行如上述的互联网新闻的舆情聚类分析方法的步骤。
相较于现有技术,本申请所提出的互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质,首先通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中;其次,对所述舆情数据库中的数据进行去噪、分词、聚类;然后,对聚类后的不同类新闻分别归纳主题摘要;最后,将聚类后的新闻及所述主题摘要输出,并显示给用户。采用本申请所提出的互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质可以快速获得网络上的新闻,对获取的新闻进行聚类获取热点新闻,并且可以对获取的新闻进行自动关键词及摘要,相较于现有技术,更加方便、快捷、准确。
图1是本申请应用服务器一可选的硬件架构的示意图;
图2是本申请互联网新闻的舆情聚类分析系统实施方式的程序模块示意图;
图3是本申请互联网新闻的舆情聚类分析方法第一实施方式的流程图;
图4是本申请互联网新闻的舆情聚类分析方法第二实施方式的流程图;
图5是本申请互联网新闻的舆情聚类分析方法第三实施方式的流程图;
图6是本申请互联网新闻的舆情聚类分析方法第四实施方式的流程图;
图7是本申请互联网新闻的舆情聚类分析方法第五实施方式的流程图;
图8是本申请互联网新闻的舆情聚类分析方法第六实施方式的流程图;
图9是本申请互联网新闻的舆情聚类分析方法第七实施方式的流程图。
本申请目的的实现、功能特点及优点将结合实施方式,参照附图做进一步说明。
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施方式,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施方式仅用以解释本申请,并不用于限定本申请。基于本申请中的实施方式,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施方式,都属于本申请保护的范围。
需要说明的是,在本申请中涉及“第一”、“第二”等的描述仅用于描述目的,而不能理解为指示或暗示其相对重要性或者隐含指明所指示的技术特征的数量。由此,限定有“第一”、“第二”的特征可以明示或者隐含地包括至少一个该特征。另外,各个实施方式之间的技术方案可以相互结合,但是必须是以本领域普通技术人员能够实现为基础,当技术方案的结合出现相互矛盾或无法实现时应当认为这种技术方案的结合不存在,也不在本申请要求的保护范围之内。
参阅图1所示,是本申请应用服务器1一可选的硬件架构的示意图。
本实施方式中,所述应用服务器1可包括,但不仅限于,可通过系统总线相互通信连接存储器11、处理器12、网络接口13。需要指出的是,图1仅示出了具有组件11-13的应用服务器1,但是应理解的是,并不要求实施所有示出的组件,可以替代的实施更多或者更少的组件。
其中,所述应用服务器1可以是机架式服务器、刀片式服务器、塔式服务器或机柜式服务器等计算设备,该应用服务器1可以是独立的服务器,也可以是多个服务器所组成的服务器集群。
所述存储器11至少包括一种类型的可读存储介质,所述可读存储介质包括闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等。在一些实施方式中,所述存储器11可以是所述应用服务器1的内部存储单元,例如该应用服务器1的硬盘或内存。在另一些实施方式中,所述存储器11也可以是所述应用服务器1的外部存储设备,例如该应用服务器1上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。当然,所述存储器11还可以既包括所述应用服务器1的内部存储单元也包括其外部存储设备。本实施方式中,所述存储器11通常用于存储安装于所述应用服务器1的操作系统和各类应用软件,例如互联网新闻的舆情聚类分析系统200的程序代码等。此外,所述存储器11还可以用于暂时地存储已经输出或者将要输出的各类数据。
所述处理器12在一些实施方式中可以是中央处理器(Central Processing Unit,CPU)、控制器、微控制器、微处理器、或其他数据处理芯片。该处理器12通常用于控制所述应用服务器1的总体操作。本实施方式中,所述处理器12用于运行所述存储器11中存储的程序代码或者处理数据,例如运行所述的互联网新闻的舆情聚类分析系统200等。
所述网络接口13可包括无线网络接口或有线网络接口,该网络接口13通常用于在所述应用服务器1与其他电子设备之间建立通信连接。
至此,己经详细介绍了本申请相关设备的硬件结构和功能。下面,将基于上述介绍提出本申请的各个实施方式。
首先,本申请提出一种互联网新闻的舆情聚类分析系统200。
参阅图2所示,是本申请互联网新闻的舆情聚类分析系统200第一实施方式的程序模块图。
在一实施方式中,所述互联网新闻的舆情聚类分析系统200包括一系列的存储于存储器11上的计算机程序指令,当该计算机程序指令被处理器12执行时,可以实现本申请各实施方式的互联网新闻的舆情聚类分析操作。在一些实施方式中,基于该计算机程序指令各部分所实现的特定的操作,互联网新闻的舆情聚类分析系统200可以被划分为一个或多个模块。例如,在图2中,所述互联网新闻的舆情聚类分析系统200可以被分割成获取模块21、处理模块22、归纳模块23及输出显示模块24。其中:
所述获取模块21,用于通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中。
具体地,所述分布式爬虫可对分布于不同服务器上的网页进行多任务抓取,提高信息的抓取效率。所述分布式爬虫框架主要分成两部分:下载器和解析器。下载器负责抓取网页,解析器负责解析网页并入库。两者之间依靠消息队列进行通信,两者可以分布在不同机器,也可分布在同一台机器。两者的数量也是灵活可变的,例如可能有五台机在做下载、两台机在做解析,这都是可以根据爬虫系统的状态及时调整的。
具体地,所述消息队列有两个管道:HTML/JS文件和待爬种子。下载器从待爬种子里拿到一条种子,根据种子信息调用相应的抓取模块进行网页抓取,然后存入HTML/JS文件这个通道;解析器从HTML/JS文件里拿到一条网页内容,根据里面的信息调用相应的解析模块进行解析,将目标字段入库, 需要的话还会解析出新的待爬种子。
具体地,所述下载器包含User-Agent池、Proxy池、Cookie池的,可以适应复杂网站的抓取。
具体地,所述信息源包括但不限于新闻网站、微博、微信、贴吧、论坛等平台。其中,新闻网站包括但不限于新浪、网易、凤凰新闻、新华网、人民网、腾讯网等国内主要新闻网站及各地地方网络报纸等。
所述处理模块22,用于对所述舆情数据库中的数据进行去噪、分词、聚类。
所述归纳模块23,用于对聚类后的不同类新闻分别归纳主题摘要。
具体地,举例而言,对同一事件不同新闻网站、媒体会有不同角度、形式的报道,但各种报道的主题都是大同小异的。新闻事件中的新闻摘要为该新闻内容的浓缩,目的是在用户阅读了新闻标题后,进一步了解新闻相关的重要信息,以便决定是否进一步阅读新闻的详细内容。用户阅读新闻大多利用手机,由于手机屏幕小,为了使有限的文字传递给用户的信息最大化的同时,尽可能减少重复信息,对聚类后的不同类新闻的主题摘要进行智能、自动归纳,从而呈现给用户,可节省用户的事件及增强用户的体验度。
所述输出显示模块24,用于将聚类后的新闻及所述主题摘要输出,并显示给用户。
具体地,所述输出显示模块24包括LCD、LED等显示屏。
此外,本申请还提出一种互联网新闻的舆情聚类分析方法。
参阅图3所示,是本申请互联网新闻的舆情聚类分析方法第一实施方式的流程示意图。在本实施方式中,根据不同的需求,图3所示的流程图中的步骤的执行顺序可以改变,某些步骤可以省略。
步骤S110,通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中。
步骤S120,对所述舆情数据库中的数据进行去噪、分词、聚类。
步骤S130,对聚类后的不同类新闻分别归纳主题摘要。
具体地,所述归纳方法包括步骤:
步骤一:对该新闻的正文进行分句,并保留句子长度在预设长度范围内的句子,记为保留句子。
具体地,通过该步骤可以限定句子的长度,从而限定了标题的长度,同时,选择预设长度范围内的句子做保留句子也有方便处理的好处。
步骤二:分别计算每个保留句子与标题的相似度S(s),以及每个保留句子的权重Q(s)。其中,引入保留句子与标题的相似度是为了使最后选取的摘要与标题的相似度低,而句子的权重则表明该句子在该新闻中的价值,通常是句子包含的关键词越多,则其价值越大。
其中,计算保留句子与标题的相似度S(s)的步骤如下:步骤1,基于同义词词库对保留句子和标题进行同义词转换;步骤2,针对同义词转换后的保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s)。即将保留句子和标题中的词组的交集除以词组的并集得到相似度S(s)。
步骤三,根据公式R(s)=Q(s)/S(S)计算每个保留句子的排序分,其中,R(s)为保留句子的排序分。通过上述公式,排序分越高,则对应的句子越可能成为摘要。
步骤四,选取排序分最高的保留句子作为同类新闻的摘要。
具体地,排序分最高的保留句子代表这些保留句子越能代表新闻的主要内容,也可以设定一个阈值,当排序分大于所述阈值,就将该排序分所对应的的句子作为保留句子,最后经过分析处理后聚合形成摘要。
步骤S140,将聚类后的新闻及所述主题摘要输出,并显示给用户。
如图4所示,是本申请互联网新闻的舆情聚类分析方法的第二实施方式的流程示意图。本实施方式中,所述步骤“对所述舆情数据库中的数据进行去噪、分词、聚类”中去噪具体包括以下步骤:
步骤S210,过滤图片、版权说明、广告,获得文档信息。
具体地,采集的网页信息含有广告、导航信息、图片、版权说明等噪声数据,对舆情信息分析来说真正需要的是正文部分的元信息,清除掉这些无关内容,保留采集网页中的文档信息。
步骤S220,过滤停用词。
具体地,在文本信息中,还包括很多无意义的词语、符号等,这些词统称为停用词。设置一个无意词表,在无意词表中添加无意义的词语、符号,比如“了”,“的”,“和”,“或”,符号等,将无意词表中的词语、符号从上一步骤获得的文档信息中去除。
如图5所示,是本申请互联网新闻的舆情聚类分析方法的第三实施方式的流程示意图。本实施方式中,所述步骤“对所述舆情数据库中的数据进行去噪、分词、聚类”中分词具体包括以下步骤:
步骤S310,运用中文分词技术对采集到的网页文本数据进行分词。
具体地,所述的中文分词技术包括两类,第一类方法应用词典匹配、汉语词法或其它汉语语言知识进行分词,如:最大匹配法、最小分词方法等。第二类基于统计的分词方法则基于字和词的统计信息,如把相邻字间的信息、词频及相应的共现信息等应用于分词。
具体地,所述的中文分词技术包括以下方法:
最大正向匹配法,其基本思想为:假定分词词典中的最长词有i个汉字字符,则用被处理文档的当前字串中的前i个字作为匹配字段,查找字典。若字典中存在这样的一个i字词,则匹配成功,匹配字段被作为一个词切分出来。如果词典中找不到这样的一个i字词,则匹配失败,将匹配字段中的最后一个字去掉,对剩下的字串重新进行匹配处理……如此进行下去,直到匹配成功,即切分出一个词或剩余字串的长度为零为止。这样就完成了一轮匹配,然后取下一个i字字串进行匹配处理,直到文档被扫描完为止。
逆向最大匹配法,逆向最大匹配法从被处理文档的末端开始匹配扫描,每次取最末端的2i个字符(i字字串)作为匹配字段,若匹配失败,则去掉匹 配字段最前面的一个字,继续匹配。相应地,它使用的分词词典是逆序词典,其中的每个词条都将按逆序方式存放。在实际处理时,先将文档进行倒排处理,生成逆序文档。然后,根据逆序词典,对逆序文档用正向最大匹配法处理即可。
如图6所示,是本申请互联网新闻的舆情聚类分析方法的第四实施方式的流程示意图。本实施方式中,所述步骤“对所述舆情数据库中的数据进行去噪、分词、聚类”中聚类具体包括以下步骤:
步骤S410,设置包括敏感词、情感词的参照表。
具体地,预设的情感词包括具有强烈情感的副词、连词以及观点词。例如,连词包括不过、但是、于是、此外等等;副词包括相当、完美、几乎、绝对等等;观点词包括察觉、发现、认为、主张、猜想、表示、以为等等。而敏感词,可为禁止词,侵权词,不雅词,政治性、煽动性的词语。
步骤S420,获得关键词,并根据获得的关键词设置关键词参照表。
具体地,该步骤可包括如下步骤:
步骤一:对分词之后的文档进行分析,统计词语出现的频率,出现的位置(标题,摘要,正文,备注),历史平均频率(若存在);
步骤二:根据公式获得词语的重要度D;
D=a*Fn+∑bi*Wi+c*Fh,i=1,2,3…n,
其中,a,b,c为词语当时出现的频率,位置,历史平均频率对应权重值,Fn,Wi,Fh分别对应词语出现的频率,出现的位置,历史平均频率,对于词语出现的位置也设置不同的权重值;
步骤三:根据重要度D对各词语进行排序,将重要度D大于预设值的词语作为关键词并生成关键词参照表。
步骤S430,对照所述敏感词、情感词及关键词参照表分析出网页文本中的关键词、敏感词和带有情感倾向的词语。
进一步地,对比近义词、同义词数据库对所述敏感词、情感词及关键词 进行分析,将关键词及其同义词、近义词统一列为关键词,将敏感词、情感词的同义词也进行对比分析。
进一步地,位于数据库服务器词库中的敏感词、关键词和情感词,也可以根据社会舆情发展变化的需要由管理员添加新的词汇到此词库,以实现词库的实时更新。
进一步地,依据处理结果提供的关键词以及情感词实现对网页信息的过滤、预警(即对网页文本信息进行标记如依照关键词标记或依照情感倾向词标记),处理结果提交给敏感词处理模块进行敏感词处理;敏感词处理模块,用于根据国家相关法律法规,对经关键词处理模块处理的文本数据中的敏感词进行过滤、屏蔽,处理结果交由聚类分析模块进行处理,其中敏感词来自分词处理模块提供的处理结果。
步骤S440,根据关键词、情感词、敏感词将从获得的网页数据按照网页所属类别自动聚类。
具体地,根据获得的关键词对数据库中的新闻网页进行分类,例如可为政治、经济、军事、社会、科技、游戏、时尚、体育及电影等;若关键词为地名,也可根据不同地名分类,例如可为北京、上海、广州、深圳,香港、台湾等等;
具体地,根据情感词可对数据库中的新闻网页进行分类,例如可分为爱情、亲情、友情、失恋、治愈等等。
具体地,根据敏感词可将数据库中的新闻网页进行分类,例如对敏感政策、敏感地域、敏感人物的分类。
进一步地,聚类之后的各类包含多个新闻的分组。
步骤S450,对聚类后的新闻按照热度进行排序。
具体地,还可获取新闻的评论数、点击数、转发数,根据以上数据进行排序。
具体地,还可根据同一或者相似新闻的数目计算热度,对不同聚类出的 不同新闻分别计算热度后根据热度排序。
如图7所示,是本申请互联网新闻的舆情聚类分析方法的第五实施方式的流程示意图。本实施方式中,所述互联网新闻的舆情聚类分析方法的步骤“获得所述关键词,并根据获得的所述关键词设置关键词参照表”具体包括:
步骤S510,对分词之后的文档进行分析,统计词语出现的频率,出现的位置及历史平均频率。
步骤S520,根据如下公式获得词语的重要度D:
D=a*Fn+∑bi*Wi+c*Fh,i=1,2,3…n。
步骤S530,根据所述重要度D对各词语进行排序,将所述重要度D大于预设值的词语作为所述关键词并生成所述关键词参照表。
其中,a,b,c为词语当时出现的频率,位置及历史平均频率对应的权重值;Fn,Wi,Fh分别对应词语出现的频率,出现的位置,所述历史平均频率。
如图8所示,是本申请互联网新闻的舆情聚类分析方法的第六实施方式的流程示意图。本实施方式中,所述互联网新闻的舆情聚类分析方法的步骤“对聚类后的不同类新闻分别归纳主题摘要”具体包括:
步骤S610,对该新闻的正文进行分句,并保留句子长度在预设长度范围内的句子,记为保留句子。
具体地,通过该步骤可以限定句子的长度,从而限定了标题的长度。
步骤S620,分别计算所述保留句子与标题的相似度S(s),以及所述保留句子的权重Q(s)。
具体地,引入保留句子与标题的相似度是为了使最后选取的摘要与标题的相似度低,而句子的权重则表明该句子在该新闻中的价值,通常是句子包含的关键词越多,则其价值越大。
步骤S630,根据公式R(s)=Q(s)/S(S)计算所述保留句子的排序分。
步骤S640,选取排序分最高的所述保留句子作为同类新闻的摘要。
如图9所示,是本申请互联网新闻的舆情聚类分析方法的第七实施方式 的流程示意图。本实施方式中,所述互联网新闻的舆情聚类分析方法的步骤“分别计算所述保留句子与标题的相似度S(s),以及所述保留句子的权重Q(s)”具体包括:
步骤S710,基于同义词词库对所述保留句子和标题进行同义词转换。
步骤S720,针对同义词转换后的所述保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s)。
具体地,针对同义词转换后的所述保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s),即将保留句子和标题中的词组的交集除以词组的并集得到相似度S(s)。
上述本申请实施方式序号仅仅为了描述,不代表实施方式的优劣。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施方式方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,空调器,或者网络设备等)执行本申请各个实施方式所述的方法。
以上仅为本申请的优选实施方式,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。
Claims (19)
- 一种互联网新闻的舆情聚类分析方法,应用于应用服务器,其特征在于,所述方法包括步骤:通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中;对所述舆情数据库中的数据进行去噪、分词、聚类;对聚类后的不同类新闻分别归纳主题摘要;及将聚类后的新闻及所述主题摘要输出,并显示给用户。
- 如权利要求1所述的互联网新闻的舆情聚类分析方法,其特征在于,所述信息源包括新闻网站、微博、微信、贴吧及论坛。
- 如权利要求1所述的互联网新闻的舆情聚类分析方法,其特征在于,所述去噪步骤包括:过滤图片、版权说明、广告,获得文档信息;及过滤停用词。
- 如权利要求1所述的互联网新闻的舆情聚类分析方法,其特征在于,所述的分词步骤包括:运用中文分词技术对采集到的所述新闻类信息进行分词。
- 如权利要求1所述的互联网新闻的舆情聚类分析方法,其特征在于,所述聚类步骤包括:设置包括敏感词、情感词的参照表;获得关键词,并根据获得的关键词设置关键词参照表;对照所述敏感词、情感词及关键词参照表分析出所述新闻类信息中的关键词、敏感词和带有情感倾向的词语;根据关键词、情感词、敏感词将所述新闻类信息按照网页所属类别自动聚类;及对聚类后的新闻按照热度进行排序。
- 如权利要求5所述的互联网新闻的舆情聚类分析方法,其特征在于, 所述获得所述关键词,并根据获得的所述关键词设置关键词参照表的步骤还包括:对分词之后的所述新闻类信息进行分析,统计词语出现的频率,出现的位置及历史平均频率;根据如下公式获得词语的重要度D:D=a*Fn+∑bi*Wi+c*Fh,i=1,2,3…n;及根据所述重要度D对各词语进行排序,将所述重要度D大于预设值的词语作为所述关键词并生成所述关键词参照表;其中,a,b,c为词语当时出现的频率,位置及历史平均频率对应的权重值;Fn,Wi,Fh分别对应词语出现的频率,出现的位置,所述历史平均频率。
- 如权利要求1所述的互联网新闻的舆情聚类分析方法,其特征在于,所述对聚类后的不同类新闻分别归纳主题摘要的步骤还包括:对该新闻的正文进行分句,并保留句子长度在预设长度范围内的句子,记为保留句子;分别计算所述保留句子与标题的相似度S(s),以及所述保留句子的权重Q(s);根据公式R(s)=Q(s)/S(S)计算所述保留句子的排序分;及选取排序分最高的所述保留句子作为同类新闻的摘要;其中,R(s)为所述保留句子的排序分。
- 如权利要求7所述的互联网新闻的舆情聚类分析方法,其特征在于,计算所述相似度S(s)的步骤如下:基于同义词词库对所述保留句子和标题进行同义词转换;及针对同义词转换后的所述保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s)。
- 一种应用服务器,其特征在于,所述应用服务器包括存储器、处理器,所述存储器上存储有可在所述处理器上运行的互联网新闻的舆情聚类分析系 统,所述互联网新闻的舆情聚类分析系统被所述处理器执行时实现如下步骤:通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中;对所述舆情数据库中的数据进行去噪、分词、聚类;对聚类后的不同类新闻分别归纳主题摘要;及将聚类后的新闻及所述主题摘要输出,并显示给用户。
- 如权利要求9所述的应用服务器,其特征在于,所述去噪步骤包括:过滤图片、版权说明、广告,获得文档信息;及过滤停用词。
- 如权利要求9所述的应用服务器,其特征在于,所述的分词步骤包括:运用中文分词技术对采集到的所述新闻类信息进行分词。
- 如权利要求9所述的应用服务器,其特征在于,所述聚类步骤包括:设置包括敏感词、情感词的参照表;获得关键词,并根据获得的关键词设置关键词参照表;对照所述敏感词、情感词及关键词参照表分析出所述新闻类信息中的关键词、敏感词和带有情感倾向的词语;根据关键词、情感词、敏感词将所述新闻类信息按照网页所属类别自动聚类;及对聚类后的新闻按照热度进行排序。
- 如权利要求12所述的应用服务器,其特征在于,所述获得所述关键词,并根据获得的所述关键词设置关键词参照表的步骤还包括:对分词之后的所述新闻类信息进行分析,统计词语出现的频率,出现的位置及历史平均频率;根据如下公式获得词语的重要度D:D=a*Fn+∑bi*Wi+c*Fh,i=1,2,3…n;及根据所述重要度D对各词语进行排序,将所述重要度D大于预设值的词语作为所述关键词并生成所述关键词参照表;其中,a,b,c为词语当时出现的频率,位置及历史平均频率对应的权重值;Fn,Wi,Fh分别对应词语出现的频率,出现的位置,所述历史平均频率。
- 如权利要求9所述的应用服务器,其特征在于,所述对聚类后的不同类新闻分别归纳主题摘要的步骤还包括:对该新闻的正文进行分句,并保留句子长度在预设长度范围内的句子,记为保留句子;分别计算所述保留句子与标题的相似度S(s),以及所述保留句子的权重Q(s);根据公式R(s)=Q(s)/S(S)计算所述保留句子的排序分;及选取排序分最高的所述保留句子作为同类新闻的摘要;其中,R(s)为所述保留句子的排序分。15.如权利要求14所述的应用服务器,其特征在于,计算所述相似度S(s)的步骤如下:基于同义词词库对所述保留句子和标题进行同义词转换;及针对同义词转换后的所述保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s)。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有互联网新闻的舆情聚类分析系统,所述互联网新闻的舆情聚类分析系统可被至少一个处理器执行,以使所述至少一个处理器执行如下步骤:通过分布式爬虫在信息源获取新闻类信息,并存储到舆情数据库中;对所述舆情数据库中的数据进行去噪、分词、聚类;对聚类后的不同类新闻分别归纳主题摘要;及将聚类后的新闻及所述主题摘要输出,并显示给用户。
- 如权利要求16所述的计算机可读存储介质,其特征在于,所述聚类步骤包括:设置包括敏感词、情感词的参照表;获得关键词,并根据获得的关键词设置关键词参照表;对照所述敏感词、情感词及关键词参照表分析出所述新闻类信息中的关键词、敏感词和带有情感倾向的词语;根据关键词、情感词、敏感词将所述新闻类信息按照网页所属类别自动聚类;及对聚类后的新闻按照热度进行排序。
- 如权利要求17所述的计算机可读存储介质,其特征在于,所述获得所述关键词,并根据获得的所述关键词设置关键词参照表的步骤还包括:对分词之后的所述新闻类信息进行分析,统计词语出现的频率,出现的位置及历史平均频率;根据如下公式获得词语的重要度D:D=a*Fn+∑bi*Wi+c*Fh,i=1,2,3…n;及根据所述重要度D对各词语进行排序,将所述重要度D大于预设值的词语作为所述关键词并生成所述关键词参照表;其中,a,b,c为词语当时出现的频率,位置及历史平均频率对应的权重值;Fn,Wi,Fh分别对应词语出现的频率,出现的位置,所述历史平均频率。
- 如权利要求16所述的计算机可读存储介质,其特征在于,所述对聚类后的不同类新闻分别归纳主题摘要的步骤还包括:对该新闻的正文进行分句,并保留句子长度在预设长度范围内的句子,记为保留句子;分别计算所述保留句子与标题的相似度S(s),以及所述保留句子的权重Q(s);根据公式R(s)=Q(s)/S(S)计算所述保留句子的排序分;及选取排序分最高的所述保留句子作为同类新闻的摘要;其中,R(s)为所述保留句子的排序分。
- 如权利要求19所述的计算机可读存储介质,其特征在于,计算所述相似度S(s)的步骤如下:基于同义词词库对所述保留句子和标题进行同义词转换;及针对同义词转换后的所述保留句子和标题采用Jaccard距离计算保留句子和标题的相似度S(s)。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201711060246.8 | 2017-11-01 | ||
| CN201711060246.8A CN107908694A (zh) | 2017-11-01 | 2017-11-01 | 互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019085355A1 true WO2019085355A1 (zh) | 2019-05-09 |
Family
ID=61843090
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/077638 Ceased WO2019085355A1 (zh) | 2017-11-01 | 2018-02-28 | 互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN107908694A (zh) |
| WO (1) | WO2019085355A1 (zh) |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110533212A (zh) * | 2019-07-04 | 2019-12-03 | 西安理工大学 | 基于大数据的城市内涝舆情监测预警方法 |
| CN115271296A (zh) * | 2022-04-13 | 2022-11-01 | 浪潮软件科技有限公司 | 一种农产品新闻舆情分析方法及装置 |
| WO2022204435A3 (en) * | 2021-03-24 | 2022-11-24 | Trust & Safety Laboratory Inc. | Multi-platform detection and mitigation of contentious online content |
| CN116306604A (zh) * | 2023-03-03 | 2023-06-23 | 北京众辉科技有限公司 | 一种新闻事件发现方法、装置及电子设备 |
| GR1010585B (el) * | 2022-11-10 | 2023-12-12 | Παναγιωτης Τσαντιλας | Ανιχνευση ιστου και συνοψη περιεχομενου |
| CN120429488A (zh) * | 2025-04-21 | 2025-08-05 | 淮安一逗洣科技有限公司 | 企业舆情监测处理方法、装置、计算机设备和存储介质 |
Families Citing this family (27)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110399478A (zh) * | 2018-04-19 | 2019-11-01 | 清华大学 | 事件发现方法和装置 |
| CN108776671A (zh) * | 2018-05-12 | 2018-11-09 | 苏州华必讯信息科技有限公司 | 一种网络舆情监控系统及方法 |
| CN110019814B (zh) * | 2018-07-09 | 2021-07-27 | 暨南大学 | 一种基于数据挖掘与深度学习的新闻信息聚合方法 |
| CN110852078B (zh) * | 2018-07-27 | 2025-09-12 | 北京京东尚科信息技术有限公司 | 生成标题的方法和装置 |
| CN109446409A (zh) * | 2018-09-19 | 2019-03-08 | 杭州安恒信息技术股份有限公司 | 一种疑似传销行为的目标对象的识别方法 |
| CN109299277A (zh) * | 2018-11-20 | 2019-02-01 | 中山大学 | 舆情分析方法、服务器及计算机可读存储介质 |
| CN109492162A (zh) * | 2018-11-23 | 2019-03-19 | 四川工大创兴大数据有限公司 | 一种智能化粮情监测方法及其系统 |
| CN109657137B (zh) * | 2018-11-26 | 2024-05-31 | 平安科技(深圳)有限公司 | 舆情新闻分类模型构建方法、装置、计算机设备和存储介质 |
| CN109815391A (zh) * | 2018-12-14 | 2019-05-28 | 深圳壹账通智能科技有限公司 | 基于大数据的新闻数据分析方法及装置、电子终端 |
| CN109753596B (zh) * | 2018-12-29 | 2021-05-25 | 中国科学院计算技术研究所 | 用于大规模网络数据采集的信源管理与配置方法和系统 |
| CN110134876B (zh) * | 2019-01-29 | 2021-10-26 | 国家计算机网络与信息安全管理中心 | 一种基于群智传感器的网络空间群体性事件感知与检测方法 |
| CN110489541B (zh) * | 2019-07-26 | 2021-02-05 | 昆明理工大学 | 基于案件要素及BiGRU的涉案舆情新闻文本摘要方法 |
| CN110929145B (zh) * | 2019-10-17 | 2023-07-21 | 平安科技(深圳)有限公司 | 舆情分析方法、装置、计算机装置及存储介质 |
| CN110990676A (zh) * | 2019-11-28 | 2020-04-10 | 福建亿榕信息技术有限公司 | 一种社交媒体热点主题提取方法与系统 |
| CN111209390B (zh) * | 2020-01-06 | 2023-09-05 | 新方正控股发展有限责任公司 | 新闻展示方法和系统、计算机可读存储介质 |
| CN112100535A (zh) * | 2020-09-16 | 2020-12-18 | 南京智数云信息科技有限公司 | 一种基于dfa算法进行网络舆情分析系统及其方法 |
| CN112148936A (zh) * | 2020-10-10 | 2020-12-29 | 广州瀚信通信科技股份有限公司 | 一种基于scrapy爬虫架构及文本分析的商旅舆情分析方法 |
| CN112183093A (zh) * | 2020-11-02 | 2021-01-05 | 杭州安恒信息安全技术有限公司 | 一种企业舆情分析方法、装置、设备及可读存储介质 |
| CN112612867B (zh) * | 2020-11-24 | 2024-09-17 | 中国传媒大学 | 新闻稿件传播分析方法、计算机可读存储介质及电子设备 |
| CN112860971A (zh) * | 2021-02-05 | 2021-05-28 | 浙江华坤道威数据科技有限公司 | 一种基于分布式多任务的社会负面舆情实时分析方法 |
| CN112989161A (zh) * | 2021-03-10 | 2021-06-18 | 平安科技(深圳)有限公司 | 新闻舆情监控方法、装置、电子设备及存储介质 |
| CN113377956B (zh) * | 2021-06-11 | 2025-01-14 | 中国工商银行股份有限公司 | 用于预测黑产攻击趋势的方法、装置、电子设备及介质 |
| CN113703603A (zh) * | 2021-07-22 | 2021-11-26 | 广东食品药品职业学院 | 一种资讯信息处理的方法及其终端 |
| CN114139528A (zh) * | 2021-11-22 | 2022-03-04 | 深圳深度赋智科技有限公司 | 一种结合依存句法分析和规则的中英文评论观点挖掘方法 |
| CN114528460A (zh) * | 2022-02-23 | 2022-05-24 | 云基华海信息技术股份有限公司 | 一种面向失信主体行为的信用风险监测方法 |
| CN117112904A (zh) * | 2023-09-01 | 2023-11-24 | 上海捷晓信息技术有限公司 | 基于大语言模型的智能资讯推荐及资讯搜索系统 |
| CN119002763B (zh) * | 2024-07-26 | 2026-03-13 | 北京百度网讯科技有限公司 | 一种内容提供方法、装置、设备以及存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102708096A (zh) * | 2012-05-29 | 2012-10-03 | 代松 | 一种基于语义的网络智能舆情监测系统及其工作方法 |
| CN103544255A (zh) * | 2013-10-15 | 2014-01-29 | 常州大学 | 基于文本语义相关的网络舆情信息分析方法 |
| CN104933093A (zh) * | 2015-05-19 | 2015-09-23 | 武汉泰迪智慧科技有限公司 | 基于大数据的地区舆情监控及决策辅助系统和方法 |
| CN105824959A (zh) * | 2016-03-31 | 2016-08-03 | 首都信息发展股份有限公司 | 舆情监控方法及系统 |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101751458A (zh) * | 2009-12-31 | 2010-06-23 | 暨南大学 | 一种网络舆情监控系统及方法 |
| US8849812B1 (en) * | 2011-08-31 | 2014-09-30 | BloomReach Inc. | Generating content for topics based on user demand |
| CN103593418B (zh) * | 2013-10-30 | 2017-03-29 | 中国科学院计算技术研究所 | 一种面向大数据的分布式主题发现方法及系统 |
| CN103559310A (zh) * | 2013-11-18 | 2014-02-05 | 广东利为网络科技有限公司 | 一种从文章中提取关键词的方法 |
| CN105760546B (zh) * | 2016-03-16 | 2019-07-30 | 广州索答信息科技有限公司 | 互联网新闻摘要的自动生成方法和装置 |
| CN106055541B (zh) * | 2016-06-29 | 2018-12-28 | 清华大学 | 一种新闻内容敏感词过滤方法及系统 |
| CN107103043A (zh) * | 2017-03-29 | 2017-08-29 | 国信优易数据有限公司 | 一种文本聚类方法及系统 |
-
2017
- 2017-11-01 CN CN201711060246.8A patent/CN107908694A/zh active Pending
-
2018
- 2018-02-28 WO PCT/CN2018/077638 patent/WO2019085355A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102708096A (zh) * | 2012-05-29 | 2012-10-03 | 代松 | 一种基于语义的网络智能舆情监测系统及其工作方法 |
| CN103544255A (zh) * | 2013-10-15 | 2014-01-29 | 常州大学 | 基于文本语义相关的网络舆情信息分析方法 |
| CN104933093A (zh) * | 2015-05-19 | 2015-09-23 | 武汉泰迪智慧科技有限公司 | 基于大数据的地区舆情监控及决策辅助系统和方法 |
| CN105824959A (zh) * | 2016-03-31 | 2016-08-03 | 首都信息发展股份有限公司 | 舆情监控方法及系统 |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110533212A (zh) * | 2019-07-04 | 2019-12-03 | 西安理工大学 | 基于大数据的城市内涝舆情监测预警方法 |
| WO2022204435A3 (en) * | 2021-03-24 | 2022-11-24 | Trust & Safety Laboratory Inc. | Multi-platform detection and mitigation of contentious online content |
| CN115271296A (zh) * | 2022-04-13 | 2022-11-01 | 浪潮软件科技有限公司 | 一种农产品新闻舆情分析方法及装置 |
| GR1010585B (el) * | 2022-11-10 | 2023-12-12 | Παναγιωτης Τσαντιλας | Ανιχνευση ιστου και συνοψη περιεχομενου |
| CN116306604A (zh) * | 2023-03-03 | 2023-06-23 | 北京众辉科技有限公司 | 一种新闻事件发现方法、装置及电子设备 |
| CN120429488A (zh) * | 2025-04-21 | 2025-08-05 | 淮安一逗洣科技有限公司 | 企业舆情监测处理方法、装置、计算机设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN107908694A (zh) | 2018-04-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019085355A1 (zh) | 互联网新闻的舆情聚类分析方法、应用服务器及计算机可读存储介质 | |
| US10169449B2 (en) | Method, apparatus, and server for acquiring recommended topic | |
| US8656266B2 (en) | Identifying comments to show in connection with a document | |
| CN103493045B (zh) | 对在线问题的自动回答 | |
| US8972413B2 (en) | System and method for matching comment data to text data | |
| US8041713B2 (en) | Systems and methods for analyzing boilerplate | |
| KR100996311B1 (ko) | 스팸 ucc를 감지하기 위한 방법 및 시스템 | |
| JP6538277B2 (ja) | 検索クエリ間におけるクエリパターンおよび関連する総統計の特定 | |
| US20090287676A1 (en) | Search results with word or phrase index | |
| US20070276801A1 (en) | Systems and methods for constructing and using a user profile | |
| CN102915380A (zh) | 用于对数据进行搜索的方法和系统 | |
| CN102930054A (zh) | 数据搜索方法及系统 | |
| CN101305371A (zh) | 对博客文档进行排名 | |
| CN103853822A (zh) | 一种在浏览器中推送新闻信息的方法和装置 | |
| WO2007143914A1 (en) | Method, device and inputting system for creating word frequency database based on web information | |
| US10417334B2 (en) | Systems and methods for providing a microdocument framework for storage, retrieval, and aggregation | |
| CN112269906B (zh) | 网页正文的自动抽取方法及装置 | |
| CN105528416A (zh) | 一种网站更新内容的监测方法及系统 | |
| CN103530389B (zh) | 一种提高停用词搜索有效性的方法和装置 | |
| CN109933691B (zh) | 用于内容检索的方法、装置、设备和存储介质 | |
| CN108701133A (zh) | 提供推荐内容 | |
| CN111723201A (zh) | 一种用于文本数据聚类的方法和装置 | |
| CN105824951A (zh) | 检索方法和装置 | |
| JP2020042545A (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| CN111259259A (zh) | 大学生新闻推荐方法、装置、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 08.10.2020) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18874919 Country of ref document: EP Kind code of ref document: A1 |