WO2014040570A1 - 一种垃圾模板文章识别方法和设备 - Google Patents

一种垃圾模板文章识别方法和设备 Download PDF

Info

Publication number
WO2014040570A1
WO2014040570A1 PCT/CN2013/083613 CN2013083613W WO2014040570A1 WO 2014040570 A1 WO2014040570 A1 WO 2014040570A1 CN 2013083613 W CN2013083613 W CN 2013083613W WO 2014040570 A1 WO2014040570 A1 WO 2014040570A1
Authority
WO
WIPO (PCT)
Prior art keywords
segment
features
article
feature
template
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2013/083613
Other languages
English (en)
French (fr)
Inventor
郝志新
何建国
张国强
何小晨
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology Shenzhen Co Ltd
Original Assignee
Tencent Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology Shenzhen Co Ltd filed Critical Tencent Technology Shenzhen Co Ltd
Priority to US14/428,314 priority Critical patent/US9330075B2/en
Publication of WO2014040570A1 publication Critical patent/WO2014040570A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • G06F16/353Clustering; Classification into predefined classes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/951Indexing; Web crawling techniques
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/953Querying, e.g. by the use of web search engines
    • G06F16/9535Search customisation based on user profiles and personalisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/26Techniques for post-processing, e.g. correcting the recognition result
    • G06V30/262Techniques for post-processing, e.g. correcting the recognition result using context analysis, e.g. lexical, syntactic or semantic context
    • G06V30/268Lexical context
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/40Document-oriented image-based pattern recognition
    • G06V30/41Analysis of document content
    • G06V30/418Document matching, e.g. of document images
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L51/00User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
    • H04L51/21Monitoring or handling of messages
    • H04L51/212Monitoring or handling of messages using filtering or selective blocking
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L51/00User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
    • H04L51/52User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L63/00Network architectures or network communication protocols for network security
    • H04L63/14Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic
    • H04L63/1441Countermeasures against malicious traffic
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L67/00Network arrangements or protocols for supporting network services or applications
    • H04L67/01Protocols
    • H04L67/10Protocols in which an application is distributed across nodes in the network

Definitions

  • the invention relates to the field of network communication, in particular to a garbage template article identification method and device. Background technique
  • Weibo APP applications
  • Similar template articles which caused a large number of spam templates in the Weibo platform.
  • These junk template articles are generally duplicated, or some texts are randomly modified according to the personal information or a certain law of the forwarder.
  • the amount of information contained is very small, but the amount of data is very large.
  • the garbage template article accounts for the total amount. 10% of the blog post. If these garbage template articles are not recognized and filtered, the search engine resources will be greatly wasted, and a large number of duplicate templates will seriously affect the user experience.
  • the same type of junk template article has some common features. At present, it mainly analyzes the semantics of the article by manual to determine whether a microblog article is a spam template article.
  • the manual identification method is slow in speed, low in efficiency, and unable to cope with the huge amount of data on the Weibo platform, and it is impossible for each Weibo article to be Perform manual identification.
  • the embodiment of the present invention provides a garbage template article identification method and device.
  • the technical solution is as follows:
  • the embodiment of the present invention provides a method for identifying a spam template article, the method comprising: extracting a feature from an eligible microblog article, and generating an article feature; wherein the article feature includes at least a punctuation feature, Topic features, parenthesis features, link features, and account name characteristics;
  • Obtaining a garbage template list where the garbage template list includes a garbage template feature;
  • the garbage template feature is an article feature whose frequency reaches a preset threshold, and the garbage template feature is extracted in the same manner as the article feature is extracted;
  • the microblog article is determined to be a junk template article.
  • the qualified microblog article is an original form and includes a microblog article of a link and a picture
  • the extracting the feature to the qualified microblog article further includes:
  • the extracting the feature of the qualified microblog article includes:
  • the qualified microblog articles are segmented by punctuation, and the segment numbers are sequentially generated in order;
  • a topic with a segment of the topic and a corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a character string to generate the topic feature;
  • a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a character string to generate the Bracket characteristics;
  • a sequence is generated as a link feature according to whether there is a link in each segment;
  • a sequence is generated as the account name feature based on whether there is an account name identifier in each of the segments.
  • the article feature further includes a content feature
  • the extracting the feature of the qualified microblog article further includes:
  • each of the segments is removed from all punctuation, topics, parentheses, links, and contents remaining after the account name is identified, and assembled in order to generate the content features.
  • the article feature further includes a previous content feature
  • the matching the microblog article extraction feature includes:
  • each of the segments is removed from all the punctuation, the topic, the parentheses, the link, and the content remaining after the account name identifier, and only the previous portion is taken in a predetermined number of bytes to generate the previous segment.
  • Content characteristics
  • the feature of the article further includes a content feature of the subsequent segment
  • the extracting the feature of the microblog article that meets the condition further includes:
  • each of the segments is removed from all punctuation, topics, parentheses, links, and account name identifiers, and the remaining portions are taken only by the predetermined number of bytes, and the generated portions are generated. Segment content characteristics.
  • the embodiment of the present invention further provides a junk template article identification device, where the device includes: a feature extraction module, configured to extract features from the qualified microblog articles, and generate article features; wherein the article features include at least punctuation features , topic features, bracket features, link features, and account name characteristics;
  • An acquisition module configured to obtain a garbage template list, where the garbage template list includes a garbage template feature;
  • the garbage template feature is an article feature whose appearance frequency reaches a preset threshold, and the garbage template feature extraction manner and the article feature The same method of extraction;
  • the identification module is configured to determine that the microblog article is a junk template article when the article feature is the same as the junk template feature in the junk template list.
  • the device further includes:
  • a pre-processing module configured to remove numbers and letters in the Weibo article before extracting features from the qualified Weibo article, and remove the contents of the various brackets in the Weibo article to retain the brackets
  • the qualified Weibo article is an original form and contains Weibo articles of links and pictures.
  • the feature extraction module includes:
  • a segmentation unit configured to segment the qualified microblog articles by punctuation, and sequentially generate segment numbers in sequence
  • a punctuation feature unit configured to extract punctuation of the segment in each segment, and form the extracted punctuation into a character string to generate the punctuation feature
  • a topic feature unit configured to extract, in each of the segments, a topic of a segment with a topic and a feature of the corresponding topic
  • a parenthesis feature unit configured to extract a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment in each segment, and combine the extracted segment number and the parenthesis type a string, generating the bracket feature;
  • a link feature unit configured to generate, in each of the segments, a sequence according to whether there is a link in each segment, as the link feature;
  • the account name feature unit is configured to generate, in each of the segments, a sequence according to whether there is an account name identifier in each segment, as the account name feature.
  • the feature extraction module further includes:
  • a content feature unit configured to, in each of the segments, remove all topics, parentheses, links, and content remaining after the account name identification, and assemble the pieces in order to generate the content features.
  • the feature extraction module further includes:
  • a preceding content feature unit configured to remove all topics, parentheses, links, and account name identifiers in each of the segments, and take only the previous portion by a predetermined number of bytes , generating the previous piece of content features.
  • the feature extraction module further includes:
  • a subsequent content feature unit configured to remove all topics, parentheses, links, and account name identifiers in each of the segments by a predetermined number of bytes In part, generating the latter piece of content features.
  • An embodiment of the present invention further provides a garbage template article identification device, where the device includes: one or more processors; and
  • the memory stores one or more programs, the one or more programs being configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations:
  • the conditional microblog article extracts features, and generates article features; wherein, the article features include at least punctuation features, topic features, parenthesis features, link features, and account name features;
  • the garbage template list includes a garbage template feature
  • the garbage template feature is an article feature whose frequency reaches a preset threshold, and the garbage template feature is extracted in the same manner as the article feature is extracted;
  • the microblog article is determined to be a junk template article.
  • an instruction for performing the following operations is further included;
  • an instruction for performing the following operations is further included;
  • a topic with a segment of the topic and a corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a character string to generate the topic feature;
  • a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a character string to generate the Bracket characteristics;
  • a sequence is generated as a link feature according to whether there is a link in each segment;
  • a sequence is generated as the account name feature based on whether there is an account name identifier in each of the segments.
  • an instruction for performing the following operations is further included;
  • each of the segments is removed from all punctuation, topics, parentheses, links, and contents remaining after the account name is identified, and assembled in order to generate the content features.
  • an instruction for performing the following operations is further included;
  • each of the segments is removed from all the punctuation, the topic, the parentheses, the link, and the content remaining after the account name identifier, and only the previous portion is taken in a predetermined number of bytes to generate the previous segment.
  • Content characteristics
  • an instruction for performing the following operations is further included;
  • each of the segments is removed from all punctuation, topics, parentheses, links, and account name identifiers, and the remaining portions are taken only by the predetermined number of bytes, and the generated portions are generated. Segment content characteristics.
  • the garbage template article identification method and device can determine whether the microblog article is a spam template article by extracting multiple features of the microblog article, and solves the problem that a large number of spam template articles in the current microblog platform cannot be effectively identified.
  • the problem is that the effect of the spam template article can be automatically and efficiently recognized automatically and without the need of manual, only need to extract the article features in the microblog article to do the logic operation.
  • FIG. 1 is a flowchart of a garbage template article identification method according to an embodiment of the present invention
  • FIG. 2 is a flowchart of a garbage template article identification method according to another embodiment of the present invention
  • FIG. 3 is an embodiment of the present invention
  • a schematic diagram of a garbage template article identification device provided by the example
  • FIG. 4 is a schematic diagram of another structure of a junk template article identification device according to another embodiment of the present invention.
  • FIG. 5 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention. detailed description
  • FIG. 1 is a flowchart of a garbage template article identification method according to an embodiment of the present invention.
  • the embodiment includes:
  • the garbage template list includes a garbage template feature;
  • the garbage template feature is an article feature whose appearance frequency reaches a preset threshold, and the garbage template feature extraction manner is the same as the microblog article feature extraction manner;
  • the microblog article feature is the article feature extracted by the microblog article in step 101.
  • the microblog article is determined to be a junk template article.
  • the qualified microblog article in the embodiment of the present invention is an original form and includes a microblog article of a link and a picture. Before extracting the feature from the qualified microblog article, the method further includes:
  • extracting features for qualified Weibo articles including: The qualified microblog articles are segmented by punctuation, and the segment numbers are sequentially generated in order; in each segment, the segmentation punctuation is extracted, and the extracted punctuation marks are formed into strings to generate punctuation features;
  • each segment the topic of the segment with the topic and the corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a string to generate a topic feature;
  • each segment the segment number corresponding to the segment with the parentheses and the parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a string to generate a bracket feature;
  • each segment a sequence is generated as a link feature according to whether there is a link in each segment; in each segment, a sequence is generated according to whether there is an account name identifier in each segment, as an account name feature .
  • the feature of the article further includes content features, and extracting features of the qualified microblog articles, and the following:
  • each segment is removed from all punctuation, topics, parentheses, links, and the remaining content after the account name is identified, assembled in order, to generate content features.
  • the feature of the article further includes the feature of the previous paragraph, and extracts features for the qualified microblog articles, and further includes:
  • each segment is removed from all punctuation, topics, parentheses, links, and the remaining contents of the account name to take only the previous portion in a predetermined number of bytes to generate the previous segment content feature.
  • the article feature further includes a later content feature, and extracts characteristics for the qualified microblog article, and further includes:
  • the predetermined operation includes not displaying, not feeding back to the terminal, deleting, masking, and alerting as a search result.
  • the garbage template article identification method determines whether the article is a garbage template article by using multiple features of the microblog article, and solves the problem that a large number of garbage template articles in the current microblog platform cannot be effectively identified, and the problem is achieved. You don't need to be artificial, you only need to extract the article features in the Weibo article to do logical operations, and you can automatically and accurately identify the effect of the spam template article.
  • FIG. 2 is a flowchart of a garbage template article identification method according to another embodiment of the present invention. Referring to FIG. 2, the embodiment includes: 201. Obtain a garbage template article, and perform preprocessing and feature extraction on the garbage template article respectively, and generate a garbage template feature to be stored in the garbage template list;
  • the step may include two sub-steps of pre-processing and feature extraction:
  • Spam template articles are generally original and contain links and images, remove the numbers and letters from the Weibo article, and remove the brackets from the various brackets in the Weibo article.
  • the pre-processed junk template article is segmented by punctuation such as a comma, a period, an exclamation point, a question mark, and a semicolon, and the segment numbers are sequentially generated in order;
  • each segment extract the segmentation punctuation in each segment in order, and form the extracted punctuation into a string to generate punctuation features
  • each segment determines whether there is a topic. If there is a topic in the segment, extract the topic corresponding to the segment and the corresponding segment number, and form the extracted topic and the segment number into a string to generate Topic features; for example, there are #topic 1# in the second segment and #topic2# in the fourth segment, then "topic 1 , 2; topic 2, 4" is generated;
  • each segment extract the segment number corresponding to the segment with the parentheses and the parenthesis type corresponding to the segment, and form the extracted segment number and the bracket type into a string to generate a bracket feature; for example, the first In the segment ( ), ⁇ ⁇ in the third segment, then generate "1 ( ), 3 ⁇ ⁇ ";
  • each segment a sequence is generated according to whether there is a link in each segment as a link feature; for example, if there is a link in the first and second segments, it is 1, and if there is no segment in the third and fourth segments, The link is 0, generating "1100";
  • each segment a sequence is generated according to whether there is an account name identifier in each segment, as an account name feature; for example, if there is an account name identifier in the first and third segments, it is 1, 2, 4 If there is no account name identifier in the segment, it is 0, and "1010" is generated; f. In each segment, remove all topics, parentheses, links, and remaining contents after the account name identification in each segment, and assemble them in order to generate content features;
  • each segment remove all topics, parentheses, links, and account name identifiers for each segment, and take only the previous portion in a predetermined number of bytes to generate the previous content features; for example, The first 4 bytes of the content, generating the front-end content features;
  • each segment remove all topics, parentheses, links, and account name identifiers for each segment, and take only the following parts according to the predetermined number of bytes to generate the content of the latter segment; for example, Take the last 4 bytes of the content, and generate the back-end content features;
  • the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the content feature may be combined in sequence to generate a junk template feature including the content feature;
  • the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the previous content feature may also be combined in order to generate a junk template feature including the previous content feature;
  • Punctuation features, topic features, parenthesis features, link features, account name features, and post-content features can also be combined in order to generate junk template features containing post-content features.
  • the above-mentioned punctuation feature, the topic feature, the parenthesis feature, the link feature, the account name feature, and the content feature, the previous content feature, and the subsequent content feature may be exchanged.
  • it is necessary to generate a junk template feature corresponding to all content features, a junk template feature including the previous content feature, and a junk template feature including the back content feature according to the order of extracting features, and subsequent extraction of features from the microblog article The order of the steps is the same as the order in which the garbage template features are extracted.
  • the garbage template feature is an article feature whose frequency reaches a preset threshold and the garbage template feature is extracted in the same manner as the subsequent microblog article feature; for example, the pre-processing is performed on the microblog article satisfying the condition every 12 hours.
  • feature extraction respectively generating an article feature including content features, an article feature including the previous content feature, and an article feature including the latter content feature, and offline calculating the frequency of occurrence of each feature. When the frequency reaches the threshold, the article is identified as garbage.
  • the template article, and the extracted three article features including content features, the article features including the previous content features, and the article features including the latter content features are determined as garbage template features, and are saved in the garbage template list, thereby continuously updating the garbage.
  • Garbage template in template list special 202. Obtain a microblog article published by the user, and preprocess the microblog article;
  • preprocessing the Weibo article includes two sub-steps as follows:
  • the method for extracting features in this step is the same as step 201 above, and details are not described herein again.
  • the article features extracted in this step include at least: punctuation feature, topic feature, parenthesis feature, link feature, and account name feature, wherein the content feature, the previous content feature, and the subsequent content feature of the microblog article may also be extracted.
  • the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the content feature of the extracted microblog article may be combined to generate all article features in sequence;
  • the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the previous content feature of the extracted microblog article may be combined to generate the previous article feature in sequence;
  • the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the subsequent content feature of the extracted microblog article may also be combined in order to generate a poster feature.
  • the above-mentioned punctuation feature, the topic feature, the parenthesis feature, the link feature, the account name feature, and the content feature, the previous content feature, and the subsequent content feature may be exchanged.
  • the garbage template feature included in the garbage template list generated in step 201 Specifically, the garbage template feature that includes all the content features generated in step 201, the garbage template feature that includes the previous content feature, and the garbage that includes the back content feature are obtained. Template feature. 205. When the feature of the article is the same as the feature of the junk template in the junk template list, determine that the microblog article is a junk template article;
  • the microblog article is determined to be a spam template article; specifically,
  • the microblog article is determined to be a junk template article
  • the microblog article is determined to be a spam template article
  • the microblog article is determined to be a spam template article.
  • the microblog article is determined to be a junk template article; if only all the article features match the junk template feature including all the content features, then it may be caused by a different name.
  • the Weibo article which is originally the same template, cannot be identified. Therefore, the garbage template feature including the previous content feature and the garbage template feature including the back content feature can be added to identify this, which can increase the recall rate of the template recognition. It may also lead to misjudgment, but because of the joint judgment with punctuation features, topic features, bracket features, link features, and account name characteristics, the probability of misjudgment is still relatively low.
  • microblog article is determined to be a junk template article
  • the microblog article is not displayed when the microblog article is retrieved.
  • FIG. 3 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention.
  • the device includes:
  • the feature extraction module 301 is configured to extract features from the qualified microblog articles and generate article features; wherein the article features include at least punctuation features, topic features, parenthesis features, link features, and account name features;
  • the obtaining module 302 is configured to obtain a garbage template list, and the garbage template list includes a garbage template feature; the garbage template feature is an article feature whose frequency reaches a preset threshold and the garbage template feature is extracted in the same manner as the microblog article feature;
  • the identification module 303 is configured to determine that the microblog article is a spam template article when the article feature is the same as the junk template feature in the junk template list.
  • the device further includes: a pre-processing module 304, as shown in FIG. 4;
  • the pre-processing module 304 is configured to remove the numbers and letters in the microblog article before extracting the feature from the qualified microblog article, and remove the brackets in the various brackets in the microblog article;
  • the blog post is a microblog article that is original and contains links and images.
  • the feature extraction module 301 includes:
  • a segmentation unit configured to segment the qualified microblog articles by punctuation, and sequentially generate the segment numbers in order
  • Punctuation feature unit used to extract the punctuation of the segment in each segment, and group the extracted punctuation into a string to generate punctuation features
  • a topic feature unit configured to extract, in each segment, a topic with a segment of the topic and a corresponding segment number, and form the extracted topic and the segment number into a string to generate a topic feature
  • a parenthesis feature unit in each segment, extracting a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment, and forming the extracted segment number and the parenthesis type into a string to generate a parenthesis feature;
  • a link feature unit configured to generate a sequence in each segment according to whether there is a link in each segment, as a link feature
  • the account name feature unit is configured to generate a sequence in each segment according to whether there is an account name identifier in each segment, as an account name feature.
  • the feature extraction module 301 further includes:
  • the content feature unit is configured to, in each segment, remove all topics, parentheses, links, and remaining contents after the account name identification in each segment, and assemble them in order to generate content features.
  • the feature extraction module 301 further includes:
  • the previous piece of content feature unit used to remove all topics in each segment, including After the number, link, and account name are identified, only the previous part is taken in a predetermined number of bytes to generate the previous content feature.
  • the feature extraction module 301 further includes:
  • the following content feature unit is used to remove all the topics, parentheses, links, and account name identifiers in each segment, and the remaining content is only taken in the predetermined number of bytes, and then generated. Segment content characteristics.
  • the junk template article identification device determines whether the microblog article is a junk template article and does not display the microblog article determined as the junk template article by extracting multiple features of the microblog article, and solves the current micro A large number of junk template articles in the blog platform can not be effectively identified, and the effect of the spam template article can be automatically and efficiently recognized automatically without the need for manual, only need to extract the article features in the microblog article to perform logical operations.
  • FIG. 5 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention.
  • the junk template article identification device 500 can be a server, and the junk template article identification device 500 includes a central processing unit (CPU) 501, a system memory 504 including a random access memory (RAM) 502 and a read only memory (ROM) 503. And a system bus 505 that connects system memory 504 and central processing unit 501.
  • the junk template article identification device 500 also includes a basic input/output system (I/O system) 506 that facilitates transfer of information between various devices within the computer, and for storing the operating system 513, applications 514, and other program modules 515.
  • the basic input/output system 506 includes a display 508 for displaying information and an input device 509 such as a mouse, keyboard for inputting information by the user. Both the display 508 and the input device 509 are connected to the central processing unit 501 via an input and output controller 510 that is coupled to the system bus 505.
  • the basic input/output system 506 can also include an input and output controller 510 for receiving and processing input from a plurality of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, input-output controller 510 also provides output to a display screen, printer, or other type of output device.
  • the mass storage device 507 is connected to the central processing unit 501 by a mass storage controller (not shown) connected to the system bus 505.
  • the mass storage device 507 and its associated computer readable medium provide non-volatile storage for the client device 500. That is, the mass storage device 507 can include a computer readable medium (not shown) such as a hard disk or a CD-ROM drive.
  • the computer readable medium can include computer storage media and communication media.
  • Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.
  • Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid state storage technologies, CD-ROM, DVD or other optical storage, tape cartridges, magnetic tape, disk storage or other magnetic storage devices.
  • RAM random access memory
  • ROM read only memory
  • EPROM Erasable programmable read-only memory
  • EEPROM electrically erasable programmable read-only memory
  • the junk template article identification device 500 can also be operated by a remote computer connected to the network via a network such as the Internet. That is, the junk template article identification device 500 can be connected to the network 512 through a network interface unit 511 connected to the system bus 505, or can be connected to other types of networks or remote computer systems using the network interface unit 511 ( Not shown).
  • the memory also includes one or more programs, the one or more programs being stored in a memory, and configured to be executed by one or more central processing units 501, the one or more programs comprising The garbage template article identification method provided by the embodiment shown in FIG. 1 and the garbage template article identification method provided by the embodiment shown in FIG.
  • a person skilled in the art may understand that all or part of the steps of implementing the above embodiments may be completed by hardware, or may be instructed by a program to execute related hardware, and the program may be stored in a computer readable storage medium.
  • the storage medium mentioned may be a read only memory, a magnetic disk or an optical disk or the like.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Computer Security & Cryptography (AREA)
  • Multimedia (AREA)
  • Data Mining & Analysis (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computing Systems (AREA)
  • Computational Linguistics (AREA)
  • Artificial Intelligence (AREA)
  • Computer Hardware Design (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Processing Of Solid Wastes (AREA)

Description

说 明 书 一种垃圾模板文章识别方法和设备
本申请要求于 2012 年 09 月 17 日提交中国专利局、 申请号为 201210344209.0、 发明名称为 "一种垃圾模板文章识别方法和设备,, 的中国专 利申请的优先权, 其全部内容通过引用结合在本申请中。 技术领域
本发明涉及网络通讯领域, 特别涉及一种垃圾模板文章识别方法和设备。 背景技术
随着微博的飞速发展, 某些微博用户为了达到广告或活动宣传的目的制作 微博 APP ( application, 应用程序), 发文吸引其他用户点击并自动发表转播文 章, 在短时间内制造大量的格式相似的模板文章, 这就造成在微博平台中, 垃 圾模板文章大量存在。 这些垃圾模板文章一般都是重复的, 或者根据转发人的 个人信息或某种规律随机的修改某些文字, 包含的信息量非常少, 但是数据量 却很大, 据统计垃圾模板文章约占全量博文的 10%。 如果不对这些垃圾模板文 章进行识别以及过滤, 会极大的浪费搜索引擎资源, 大量的重复模板也会严重 影响用户体验。
同一类的垃圾模板文章具有某些共性特征, 目前, 主要通过人工对文章所 包含的语义进行分析, 从而判断某一篇微博文章是否为垃圾模板文章。
在实现本发明的过程中, 发明人发现现有技术至少存在以下问题: 人工识别的方式速度比较慢,效率较低,无法应对微博平台庞大的数据量, 不可能对每篇微博文章都进行人工识别。 发明内容
为了有效解决目前微博平台中大量的垃圾模板文章无法有效识别的问题, 本发明实施例提供了一种垃圾模板文章识别方法和设备。 所述技术方案如下: 本发明实施例提供了一种垃圾模板文章识别方法, 所述方法包括: 对符合条件的微博文章提取特征, 生成文章特征; 其中, 所述文章特征至 少包括标点特征、 话题特征、 括号特征、 链接特征以及账户名特征; 获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特征; 所述垃圾模 板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特征的提取方式 与所述文章特征的提取方式相同;
当所述文章特征与所述垃圾模板列表中的垃圾模板特征相同时, 判定所述 微博文章为垃圾模板文章。
具体地, 所述符合条件的微博文章为原创形式且包含链接和图片的微博文 章, 所述对符合条件的微博文章提取特征之前, 还包括:
将所述符合条件的微博文章中的数字以及字母去掉, 并将所述微博文章中 的各种括号中的内容去掉保留所述括号。
具体地, 所述对符合条件的微博文章提取特征, 包括:
将所述符合条件的微博文章以标点进行分段, 并按顺序依次生成分段编 号;
在所述每个分段中, 提取所述分段的标点, 并将提取的所述标点组成字符 串, 生成所述标点特征;
在所述每个分段中, 提取有话题的分段的话题和对应的分段编号, 并将提 取的所述话题以及所述分段编号组成字符串, 生成所述话题特征;
在所述每个分段中,提取有括号的分段对应的分段编号和所述分段对应的 括号类型, 将提取的所述分段编号以及所述括号类型组成字符串, 生成所述括 号特征;
在所述每个分段中, 根据所述每个分段中是否有链接而生成序列, 作为所 述链接特征;
在所述每个分段中, 根据所述每个分段中是否有账户名标识而生成序列, 作为所述账户名特征。
进一步地, 所述文章特征还包括内容特征, 所述对符合条件的微博文章提 取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容特征。
进一步地, 所述文章特征还包括前段内容特征, 所述对符合条件的微博文 章提取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取前面的部分, 生成所述前段 内容特征。
进一步地, 所述文章特征还包括后段内容特征, 所述对符合条件的微博文 章提取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取后面的部分, 生成所述后段 内容特征。
本发明实施例还提供了一种垃圾模板文章识别设备, 所述设备包括: 特征提取模块, 用于对符合条件的微博文章提取特征, 生成文章特征; 其 中, 所述文章特征至少包括标点特征、 话题特征、 括号特征、 链接特征以及账 户名特征;
获取模块, 用于获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特 征; 所述垃圾模板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特 征的提取方式与所述文章特征的提取方式相同;
识别模块, 用于当所述文章特征与所述垃圾模板列表中的垃圾模板特征相 同时, 判定所述微博文章为垃圾模板文章。
具体地, 所述设备还包括:
预处理模块, 用于对符合条件的微博文章提取特征之前, 将所述微博文章 中的数字以及字母去掉, 并将所述微博文章中的各种括号中的内容去掉保留所 述括号; 所述符合条件的微博文章为原创形式且包含链接和图片的微博文章。
具体地, 所述特征提取模块, 包括:
分段单元, 用于将所述符合条件的微博文章以标点进行分段, 并按顺序依 次生成分段编号;
标点特征单元, 用于在所述每个分段中, 提取所述分段的标点, 并将提取 的所述标点组成字符串, 生成所述标点特征;
话题特征单元, 用于在所述每个分段中, 提取有话题的分段的话题和对应 题特征;
括号特征单元, 用于在所述每个分段中, 提取有括号的分段对应的分段编 号和所述分段对应的括号类型,将提取的所述分段编号以及所述括号类型组成 字符串, 生成所述括号特征; 链接特征单元, 用于在所述每个分段中, 根据所 述每个分段中是否有链接而生成序列, 作为所述链接特征; 账户名特征单元, 用于在所述每个分段中, 根据所述每个分段中是否有账 户名标识而生成序列, 作为所述账户名特征。
进一步地, 所述特征提取模块, 还包括:
内容特征单元,用于在所述每个分段中,将所述每个分段去除所有的话题、 括号、 链接以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容 特征。
进一步地, 所述特征提取模块, 还包括:
前段内容特征单元, 用于在所述每个分段中, 将所述每个分段去除所有的 话题、 括号、 链接以及账户名标识后剩余的内容按预定的字节数只取前面的部 分, 生成所述前段内容特征。
进一步地, 所述特征提取模块, 还包括:
后段内容特征单元, 用于在所述每个分段中, 将所述每个分段去除所有的 话题、 括号、 链接以及账户名标识后剩余的内容按预定的字节数只取后面的部 分, 生成所述后段内容特征。
本发明实施例还提供了一种垃圾模板文章识别设备, 所述设备包括: 一个或多个处理器; 和
存储器;
所述存储器存储有一个或多个程序, 所述一个或多个程序被配置成由所述 一个或多个处理器执行, 所述一个或多个程序包含用于进行以下操作的指令: 对符合条件的微博文章提取特征, 生成文章特征; 其中, 所述文章特征至 少包括标点特征、 话题特征、 括号特征、 链接特征以及账户名特征;
获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特征; 所述垃圾模 板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特征的提取方式 与所述文章特征的提取方式相同;
当所述文章特征与所述垃圾模板列表中的垃圾模板特征相同时, 判定所述 微博文章为垃圾模板文章。
优选地, 还包含用于进行以下操作的指令;
将所述符合条件的微博文章中的数字以及字母去掉, 并将所述微博文章中 的各种括号中的内容去掉保留所述括号。
优选地, 还包含用于进行以下操作的指令;
将所述符合条件的微博文章以标点进行分段, 并按顺序依次生成分段编 号;
在所述每个分段中, 提取所述分段的标点, 并将提取的所述标点组成字符 串, 生成所述标点特征;
在所述每个分段中, 提取有话题的分段的话题和对应的分段编号, 并将提 取的所述话题以及所述分段编号组成字符串, 生成所述话题特征;
在所述每个分段中,提取有括号的分段对应的分段编号和所述分段对应的 括号类型, 将提取的所述分段编号以及所述括号类型组成字符串, 生成所述括 号特征;
在所述每个分段中, 根据所述每个分段中是否有链接而生成序列, 作为所 述链接特征;
在所述每个分段中, 根据所述每个分段中是否有账户名标识而生成序列, 作为所述账户名特征。
优选地, 还包含用于进行以下操作的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容特征。
优选地, 还包含用于进行以下操作的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取前面的部分, 生成所述前段 内容特征。
优选地, 还包含用于进行以下操作的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取后面的部分, 生成所述后段 内容特征。
本发明实施例提供的技术方案带来的有益效果是:
本发明实施例提供的垃圾模板文章识别方法和设备,通过提取微博文章的 多个特征来判断微博文章是否为垃圾模板文章, 解决了目前微博平台中大量的 垃圾模板文章无法有效识别的问题, 达到了不需要人工、 只需要提取微博文章 中的文章特征来做逻辑运算就可以高效、 准确地自动识别出垃圾模板文章的效 果。 附图说明 为了更清楚地说明本发明实施例中的技术方案, 下面将对实施例描述中所 需要使用的附图作筒单地介绍, 显而易见地, 下面描述中的附图仅仅是本发明 的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下, 还可以根据这些附图获得其他的附图。
图 1是本发明一个实施例提供的一种垃圾模板文章识别方法的流程图; 图 2是本发明另一个实施例提供的一种垃圾模板文章识别方法的流程图; 图 3 是本发明一个实施例提供的一种垃圾模板文章识别设备的结构示意 图;
图 4是本发明另一个实施例提供的一种垃圾模板文章识别设备的另一结构 示意图;
图 5是本发明一个实施例提供的垃圾模板文章识别设备的结构示意图。 具体实施方式
为使本发明的目的、 技术方案和优点更加清楚, 下面将结合附图对本发明 实施方式作进一步地详细描述。 图 1是本发明一个实施例提供的一种垃圾模板文章识别方法的流程图, 参 见图 1 , 该实施例包括:
101、 对符合条件的微博文章提取特征, 生成文章特征; 其中, 文章特征 至少包括标点特征、 话题特征、 括号特征、 链接特征以及账户名特征;
102、 获取垃圾模板列表, 垃圾模板列表中包含垃圾模板特征; 垃圾模板 特征为出现频率达到预设阈值的文章特征且垃圾模板特征的提取方式与微博 文章特征的提取方式相同;
该微博文章特征即为步骤 101中的微博文章提取的文章特征。
103、 当文章特征与垃圾模板列表中的垃圾模板特征相同时, 判定微博文 章为垃圾模板文章。
具体地, 本发明实施例中的符合条件的微博文章为原创形式且包含链接和 图片的微博文章, 对符合条件的微博文章提取特征之前, 还包括:
将符合条件的微博文章中的数字以及字母去掉, 并将微博文章中的各种括 号中的内容去掉保留括号。
具体地, 对符合条件的微博文章提取特征, 包括: 将符合条件的微博文章以标点进行分段, 并按顺序依次生成分段编号; 在每个分段中, 提取分段的标点, 并将提取的标点组成字符串, 生成标点 特征;
在每个分段中, 提取有话题的分段的话题和对应的分段编号, 并将提取的 话题以及分段编号组成字符串, 生成话题特征;
在每个分段中, 提取有括号的分段对应的分段编号和分段对应的括号类 型, 将提取的分段编号以及括号类型组成字符串, 生成括号特征;
在每个分段中, 根据每个分段中是否有链接而生成序列, 作为链接特征; 在每个分段中, 根据每个分段中是否有账户名标识而生成序列, 作为账户 名特征。
进一步地, 文章特征还包括内容特征, 对符合条件的微博文章提取特征, 还包括:
在每个分段中, 将每个分段去除所有的标点、 话题、 括号、 链接以及账户 名标识后剩余的内容, 按顺序拼装在一起, 生成内容特征。
进一步地, 文章特征还包括前段内容特征, 对符合条件的微博文章提取特 征, 还包括:
在每个分段中, 将每个分段去除所有的标点、 话题、 括号、 链接以及账户 名标识后剩余的内容按预定的字节数只取前面的部分, 生成前段内容特征。
进一步地, 文章特征还包括后段内容特征, 对符合条件的微博文章提取特 征, 还包括:
在每个分段中, 将每个分段去除所有的标点、 话题、 括号、 链接以及账户 名标识后剩余的内容按预定的字节数只取后面的部分, 生成后段内容特征。
进一步地, 判定微博文章为垃圾模板文章后, 执行预定操作, 预定操作包 括不予显示、 不作为搜索结果反馈给终端、 删除、 屏蔽和告警中的任意一种。
本发明实施例提供的垃圾模板文章识别方法,通过对微博文章的多个特征 来判断文章是否为垃圾模板文章, 解决了目前微博平台中大量的垃圾模板文章 无法有效识别的问题, 达到了不需要人工、 只需要提取微博文章中的文章特征 来做逻辑运算就可以高效、 准确地自动识别出垃圾模板文章的效果。 图 2是本发明另一实施例提供的一种垃圾模板文章识别方法的流程图。 参 见图 2, 该实施例包括: 201、 获取垃圾模板文章, 并对垃圾模板文章分别进行预处理和特征提取, 并生成垃圾模板特征存储在垃圾模板列表中;
具体地, 该步骤可以包括预处理和特征提取两个子步骤:
( 1 )获取垃圾模板文章, 并对垃圾模板文章分别进行预处理:
垃圾模板文章一般为原创形式且同时包含链接和图片, 将微博文章中的数 字以及字母去掉, 并将微博文章中的各种括号中的内容去掉保留括号。
比如 "QQ等级加速 443天" 和 "QQ等级加速 373天", 该类垃圾模板 文章除了里面的数字不一样, 其他的都一样, 因此去掉字母数字更能提高模板 的召回率; 由于有些类模板仅改变括号里面的内容, 所以将文章中各种括号如 (),[],<>, ( ), 【】, 《》, ",,等中的内容去掉, 括号本身要保留, 供后续特征提取 时使用,
( 2 )对预处理后的垃圾模板文章提取特征, 分别生成包含全部内容特征 的垃圾模板特征、 包含前段内容特征的垃圾模板特征和包含后段内容特征的垃 圾模板特征, 包括:
将预处理后的垃圾模板文章以标点比如逗号、 句号、 感叹号、 问号、 分号 进行分段, 按顺序依次生成分段编号;
a、 在每个分段中, 按顺序在每个分段中提取分段的标点, 将提取的标点 组成字符串, 生成标点特征;
b、 在每个分段中, 判断是否有话题, 如果分段中有话题, 则提取该分段 对应的话题和对应的分段编号, 并将提取的话题以及分段编号组成字符串, 生 成话题特征;比如第 2分段有 #话题 1#和第 4分段有 #话题 2#,则生成 "话题 1 , 2; 话题 2, 4" ;
c、 在每个分段中, 提取有括号的分段对应的分段编号和该分段对应的括 号类型, 将提取的分段编号以及该括号类型组成字符串, 生成括号特征; 比如 第 1分段中有 ( ), 第 3分段中 { } , 则生成 "1 ( ) , 3 { }";
d、 在每个分段中, 根据每个分段中是否有链接而生成序列, 作为链接特 征; 比如第 1、 2分段中如果有链接则为 1 , 第 3、 4分段中如果没有链接则为 0, 生成 "1100" ;
e、 在每个分段中, 根据每个分段中是否有账户名标识而生成序列, 作为 账户名特征; 比如第 1、 3分段中如果有账户名标识则为 1 , 第 2、 4分段中如 果没有账户名标识则为 0, 生成 "1010" ; f、 在每个分段中, 将每个分段去除所有的话题、 括号、 链接以及账户名 标识后剩余的内容, 按顺序拼装在一起, 生成内容特征;
g、 在每个分段中, 将每个分段去除所有的话题、 括号、 链接以及账户名 标识后剩余的内容按预定的字节数只取前面的部分, 生成前段内容特征; 比如 可以取内容的前 4个字节, 生成前段内容特征;
h、 在每个分段中, 将每个分段去除所有的话题、 括号、 链接以及账户名 标识后剩余的内容按预定的字节数只取后面的部分, 生成后段内容特征; 比如 可以取内容的后 4个字节, 生成后段内容特征;
可以将上述标点特征、 话题特征、 括号特征、 链接特征、 账户名特征以 及内容特征, 按顺序组合生成包含内容特征的垃圾模板特征;
也可以将标点特征、 话题特征、 括号特征、 链接特征、 账户名特征以及 前段内容特征, 按顺序组合生成包含前段内容特征的垃圾模板特征;
还可以将标点特征、 话题特征、 括号特征、 链接特征、 账户名特征以及 后段内容特征, 按顺序组合生成包含后段内容特征的垃圾模板特征。
需要说明的是, 上述标点特征、 话题特征、 括号特征、 链接特征、 账户 名特征以及内容特征、 前段内容特征、 后段内容特征的提取先后顺序可以相互 调换, 对此本发明实施例不做限制, 但是需要按照提取特征的先后顺序生成相 应地包含全部内容特征的垃圾模板特征、 包含前段内容特征的垃圾模板特征和 包含后段内容特征的垃圾模板特征, 并且后续的对微博文章提取特征时的先后 顺序与垃圾模板特征提取时的先后顺序要相同。
( 3 )将生成的包含全部内容特征的垃圾模板特征、 包含前段内容特征的 垃圾模板特征和包含后段内容特征的垃圾模板特征保存到垃圾模板列表中; 需要特别说明的是, 本发明实施例的垃圾模板特征为出现频率达到预设阈 值的文章特征且垃圾模板特征的提取方式与后续的微博文章特征的提取方式 相同; 比如按照每 12小时定时对满足条件的微博文章进行上述预处理和特征 提取, 分别生成包含内容特征的文章特征、 包含前段内容特征的文章特征和包 含后段内容特征的文章特征, 离线计算每个特征出现的频率, 当频率达到阈值 时,认定该文章为垃圾模板文章,并将提取到的 3个包含内容特征的文章特征、 包含前段内容特征的文章特征和包含后段内容特征的文章特征判定为垃圾模 板特征, 保存到垃圾模板列表中, 从而不断更新垃圾模板列表中的垃圾模板特 202、 获取用户发表的微博文章, 并对微博文章进行预处理;
具体地, 对微博文章进行预处理, 包括如两个子步骤:
( 1 ) 首先判定微博文章是否为原创形式以及是否包含链接和图片; 其中, 绝大多数垃圾模板文章都是以原创形式发表的, 为了达到病毒式的 宣传效果模板一般都会包含链接, 用户点击后自动发表, 而且为了达到吸引眼 球的目的, 绝大多数垃圾模板文章都包含图片。
( 2 ) 当微博文章为原创形式且同时包含链接和图片时, 将微博文章中的 数字以及字母去掉, 并将微博文章中的各种括号中的内容去掉且保留括号。
首先, 对满足为原创形式且同时包含链接和图片的微博文章, 将其中的数 字以及字母去掉, 比如将 "QQ等级加速 443天"里面的数字 443去掉; 然后, 由于有些类模板仅改变括号里面的内容,所以将微博文章中各种括号如 (),[],<>, ( ), 【】, 《》, ""等中的内容去掉,括号本身要保留,供后续特征提取时使用,
203、 对上述预处理后的微博文章提取特征, 生成文章特征;
具体地, 该步骤提取特征的方式与上述步骤 201相同, 在此不再赘述。 本 步骤所提取的文章特征至少包括: 标点特征、话题特征、括号特征、链接特征、 账户名特征, 其中还可以提取该微博文章的内容特征、 前段内容特征、 后段内 容特征。
其中, 可以将上述提取的该微博文章的标点特征、 话题特征、 括号特征、 链接特征、 账户名特征以及内容特征, 按顺序组合生成全部文章特征;
也可以将上述提取的该微博文章的标点特征、 话题特征、 括号特征、 链 接特征、 账户名特征以及前段内容特征, 按顺序组合生成前段文章特征;
还可以将上述提取的该微博文章的标点特征、 话题特征、 括号特征、 链 接特征、 账户名特征以及后段内容特征, 按顺序组合生成后段文章特征。
需要说明的是, 上述标点特征、 话题特征、 括号特征、 链接特征、 账户 名特征以及内容特征、 前段内容特征、 后段内容特征的提取先后顺序可以相互 调换, 对此本发明实施例不做限制, 但是需要按照提取特征的先后顺序生成相 应地文章特征、 前段文章特征、 后段文章特征, 并且与步骤 201生成的垃圾模 板特征的先后顺序相同。
204、 获取步骤 201生成的垃圾模板列表中包含的垃圾模板特征; 具体地, 获取步骤 201生成的包含全部内容特征的垃圾模板特征、 包含前 段内容特征的垃圾模板特征和包含后段内容特征的垃圾模板特征。 205、 当文章特征与垃圾模板列表中的垃圾模板特征相同时, 判定该微博 文章为垃圾模板文章;
具体地, 当步骤 203生成的全部文章特征、 前段文章特征和后段文章特征 中的任一特征与垃圾模板列表中的垃圾模板特征相同时, 判定微博文章为垃圾 模板文章; 具体地,
当全部文章特征与包含全部内容特征的垃圾模板特征相同时, 判定微博文 章为垃圾模板文章;
或,
当前段文章特征与包含前段内容特征的垃圾模板特征相同时, 判定微博文 章为垃圾模板文章;
或,
当后段文章特征与包含后段内容特征的垃圾模板特征相同时, 判定微博文 章为垃圾模板文章。
当满足上述条件中的任一条件时, 判定该微博文章为垃圾模板文章; 如果 仅用全部文章特征与包含全部内容特征的垃圾模板特征匹配, 那么可能由于某 个名字的不同, 就会导致本来是同一模板的微博文章识别不出来, 因此增加包 含前段内容特征的垃圾模板特征和包含后段内容特征的垃圾模板特征, 就可以 对此进行识别, 这样可以增加模板识别的召回率, 当然也有可能导致误判, 但 由于还要同标点特征、 话题特征、 括号特征、 链接特征、 账户名特征联合判 断, 误判的概率还是比较低的。
206、 当判定该微博文章为垃圾模板文章时, 在以后的微博文章检索时, 当检索到该微博文章时不予显示。
其中,当判定该微博文章为垃圾模板文章时,对于后续的微博文章检索中, 不管是什么形式的检索, 只要检索到该微博文章, 都不予显示
本发明实施例提供的垃圾模板文章识别方法,通过提取微博文章的多个特 征来判断微博文章是否为垃圾模板文章, 解决了目前微博平台中大量的垃圾模 板文章无法有效识别的问题, 达到了不需要人工、 只需要提取微博文章中的文 章特征来做逻辑运算就可以高效、 准确地自动识别出垃圾模板文章的效果。 图 3 是本发明一个实施例提供的一种垃圾模板文章识别设备的结构示意 图, 参见图 3 , 该设备包括: 特征提取模块 301 ,用于对符合条件的微博文章提取特征,生成文章特征; 其中, 文章特征至少包括标点特征、 话题特征、 括号特征、 链接特征以及账户 名特征;
获取模块 302, 用于获取垃圾模板列表, 垃圾模板列表中包含垃圾模板特 征; 垃圾模板特征为出现频率达到预设阈值的文章特征且垃圾模板特征的提取 方式与微博文章特征的提取方式相同;
识别模块 303 ,用于当文章特征与垃圾模板列表中的垃圾模板特征相同时, 判定微博文章为垃圾模板文章。
具体地, 设备还包括: 预处理模块 304, 如图 4所示;
预处理模块 304, 用于对符合条件的微博文章提取特征之前, 将微博文章 中的数字以及字母去掉, 并将微博文章中的各种括号中的内容去掉保留括号; 符合条件的微博文章为原创形式且包含链接和图片的微博文章。
具体地, 特征提取模块 301 , 包括:
分段单元, 用于将符合条件的微博文章以标点进行分段, 并按顺序依次生 成分段编号;
标点特征单元, 用于在每个分段中, 提取分段的标点, 并将提取的标点组 成字符串, 生成标点特征;
话题特征单元, 用于在每个分段中, 提取有话题的分段的话题和对应的分 段编号, 并将提取的话题以及分段编号组成字符串, 生成话题特征;
括号特征单元, 用于在每个分段中, 提取有括号的分段对应的分段编号和 分段对应的括号类型, 将提取的分段编号以及括号类型组成字符串, 生成括号 特征;
链接特征单元, 用于在每个分段中, 根据每个分段中是否有链接而生成序 列, 作为链接特征;
账户名特征单元, 用于在每个分段中, 根据每个分段中是否有账户名标识 而生成序列, 作为账户名特征。
进一步地, 特征提取模块 301 , 还包括:
内容特征单元, 用于在每个分段中, 将每个分段去除所有的话题、 括号、 链接以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成内容特征。
进一步地, 特征提取模块 301 , 还包括:
前段内容特征单元, 用于在每个分段中, 将每个分段去除所有的话题、 括 号、 链接以及账户名标识后剩余的内容按预定的字节数只取前面的部分, 生成 前段内容特征。
进一步地, 特征提取模块 301 , 还包括:
后段内容特征单元, 用于在每个分段中, 将每个分段去除所有的话题、 括 号、 链接以及账户名标识后剩余的内容按预定的字节数只取后面的部分, 生成 后段内容特征。
本发明实施例提供的垃圾模板文章识别设备,通过提取微博文章的多个特 征来判断微博文章是否为垃圾模板文章并对判定为垃圾模板文章的微博文章 不予显示, 解决了目前微博平台中大量的垃圾模板文章无法有效识别的问题, 达到了不需要人工、 只需要提取微博文章中的文章特征来做逻辑运算就可以高 效、 准确地自动识别出垃圾模板文章的效果。
需要说明的是: 上述实施例提供的垃圾模板文章识别设备在识别垃圾模板 文章时, 仅以上述各功能模块的划分进行举例说明, 实际应用中, 可以根据需 要而将上述功能分配由不同的功能模块完成, 即将垃圾模板文章识别设备的内 部结构划分成不同的功能模块, 以完成以上描述的全部或者部分功能。 另外, 上述实施例提供的垃圾模板文章识别设备与的垃圾模板文章识别方法实施例 属于同一构思, 其具体实现过程详见方法实施例, 这里不再赘述。 图 5是本发明一个实施例提供的垃圾模板文章识别设备的结构示意图。 所 述垃圾模板文章识别设备 500可以为服务器, 所述垃圾模板文章识别设备 500 包括中央处理单元(CPU ) 501、 包括随机存取存储器(RAM ) 502 和只读存 储器(ROM ) 503的系统存储器 504, 以及连接系统存储器 504和中央处理单 元 501的系统总线 505。 所述垃圾模板文章识别设备 500还包括帮助计算机内 的各个器件之间传输信息的基本输入 /输出系统(I/O 系统) 506, 和用于存储 操作系统 513、 应用程序 514和其他程序模块 515的大容量存储设备 507。
所述基本输入 /输出系统 506包括有用于显示信息的显示器 508和用于用户 输入信息的诸如鼠标、 键盘之类的输入设备 509。 其中所述显示器 508和输入 设备 509都通过连接到系统总线 505的输入输出控制器 510连接到中央处理单 元 501。所述基本输入 /输出系统 506还可以包括输入输出控制器 510以用于接 收和处理来自键盘、 鼠标、 或电子触控笔等多个其他设备的输入。 类似地, 输 入输出控制器 510还提供输出到显示屏、 打印机或其他类型的输出设备。 所述大容量存储设备 507通过连接到系统总线 505 的大容量存储控制器 (未示出)连接到中央处理单元 501。 所述大容量存储设备 507及其相关联的 计算机可读介质为客户端设备 500提供非易失性存储。 也就是说, 所述大容量 存储设备 507可以包括诸如硬盘或者 CD-ROM驱动器之类的计算机可读介质 (未示出)。
不失一般性, 所述计算机可读介质可以包括计算机存储介质和通信介质。 计算机存储介质包括以用于存储诸如计算机可读指令、 数据结构、 程序模块或 其他数据等信息的任何方法或技术实现的易失性和非易失性、可移动和不可移 动介质。 计算机存储介质包括 RAM、 ROM, EPROM、 EEPROM、 闪存或其他 固态存储其技术, CD-ROM、 DVD 或其他光学存储、 磁带盒、 磁带、 磁盘存 储或其他磁性存储设备。 当然, 本领域技术人员可知所述计算机存储介质不局 限于上述几种。上述的系统存储器 504和大容量存储设备 507可以统称为存储 器。
根据本发明的各种实施例, 所述垃圾模板文章识别设备 500还可以通过诸 如因特网等网络连接到网络上的远程计算机运行。也即垃圾模板文章识别设备 500可以通过连接在所述系统总线 505上的网络接口单元 511连接到网络 512, 或者说, 也可以使用网络接口单元 511来连接到其他类型的网络或远程计算机 系统(未示出)。
所述存储器还包括一个或者一个以上的程序, 所述一个或者一个以上程序 存储于存储器中, 且经配置以由一个或者一个以上中央处理单元 501执行所述 一个或者一个以上程序包含用于执行图 1所示实施例所提供的垃圾模板文章识 别方法和图 2所示实施例所提供的垃圾模板文章识别方法。
本领域普通技术人员可以理解实现上述实施例的全部或部分步骤可以通 过硬件来完成, 也可以通过程序来指令相关的硬件完成, 所述的程序可以存储 于一种计算机可读存储介质中, 上述提到的存储介质可以是只读存储器, 磁盘 或光盘等。
以上所述仅为本发明的较佳实施例, 并不用以限制本发明, 凡在本发明的 精神和原则之内, 所作的任何修改、 等同替换、 改进等, 均应包含在本发明的 保护范围之内。

Claims

权 利 要 求 书
1、 一种垃圾模板文章识别方法, 其特征在于, 所述方法包括:
对符合条件的微博文章提取特征, 生成文章特征; 其中, 所述文章特征至 少包括标点特征、 话题特征、 括号特征、 链接特征以及账户名特征;
获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特征; 所述垃圾模 板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特征的提取方式与 所述文章特征的提取方式相同;
当所述文章特征与所述垃圾模板列表中的垃圾模板特征相同时, 判定所述 微博文章为垃圾模板文章。
2、 根据权利要求 1所述的方法, 其特征在于, 所述符合条件的微博文章为 原创形式且包含链接和图片的微博文章, 所述对符合条件的微博文章提取特征 之前, 还包括:
将所述符合条件的微博文章中的数字以及字母去掉, 并将所述微博文章中 的各种括号中的内容去掉保留所述括号。
3、 根据权利要求 1所述的方法, 其特征在于, 所述对符合条件的微博文章 提取特征, 包括:
将所述符合条件的微博文章以标点进行分段, 并按顺序依次生成分段编号; 在所述每个分段中, 提取所述分段的标点, 并将提取的所述标点组成字符 串, 生成所述标点特征;
在所述每个分段中, 提取有话题的分段的话题和对应的分段编号, 并将提 取的所述话题以及所述分段编号组成字符串, 生成所述话题特征;
在所述每个分段中, 提取有括号的分段对应的分段编号和所述分段对应的 括号类型, 将提取的所述分段编号以及所述括号类型组成字符串, 生成所述括 号特征;
在所述每个分段中, 根据所述每个分段中是否有链接而生成序列, 作为所 述链接特征;
在所述每个分段中, 根据所述每个分段中是否有账户名标识而生成序列, 作为所述账户名特征。
4、 根据权利要求 3所述的方法, 其特征在于, 所述文章特征还包括内容特 征, 所述对符合条件的微博文章提取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容特征。
5、 根据权利要求 3所述的方法, 其特征在于, 所述文章特征还包括前段内 容特征, 所述对符合条件的微博文章提取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取前面的部分, 生成所述前段 内容特征。
6、 根据权利要求 3所述的方法, 其特征在于, 所述文章特征还包括后段内 容特征, 所述对符合条件的微博文章提取特征, 还包括:
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取后面的部分, 生成所述后段 内容特征。
7、 一种垃圾模板文章识别设备, 其特征在于, 所述设备包括:
特征提取模块, 用于对符合条件的微博文章提取特征, 生成文章特征; 其 中, 所述文章特征至少包括标点特征、 话题特征、 括号特征、 链接特征以及账 户名特征;
获取模块, 用于获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特 征; 所述垃圾模板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特 征的提取方式与所述文章特征的提取方式相同;
识别模块, 用于当所述文章特征与所述垃圾模板列表中的垃圾模板特征相 同时, 判定所述微博文章为垃圾模板文章。
8、 根据权利要求 7所述的设备, 其特征在于, 所述设备还包括: 预处理模块, 用于对符合条件的微博文章提取特征之前, 将所述微博文章 中的数字以及字母去掉, 并将所述微博文章中的各种括号中的内容去掉保留所 述括号; 所述符合条件的微博文章为原创形式且包含链接和图片的微博文章。
9、 根据权利要求 7所述的设备, 其特征在于, 所述特征提取模块, 包括: 分段单元, 用于将所述符合条件的微博文章以标点进行分段, 并按顺序依 次生成分段编号;
标点特征单元, 用于在所述每个分段中, 提取所述分段的标点, 并将提取 的所述标点组成字符串, 生成所述标点特征;
话题特征单元, 用于在所述每个分段中, 提取有话题的分段的话题和对应 题特征;
括号特征单元, 用于在在所述每个分段中, 提取有括号的分段对应的分段 编号和所述分段对应的括号类型, 将提取的所述分段编号以及所述括号类型组 成字符串, 生成所述括号特征;
链接特征单元, 用于在所述每个分段中, 根据所述每个分段中是否有链接 而生成序列, 作为所述链接特征;
账户名特征单元, 用于在所述每个分段中, 根据所述每个分段中是否有账 户名标识而生成序列, 作为所述账户名特征。
10、 根据权利要求 9所述的设备, 其特征在于, 所述特征提取模块, 还包 括:
内容特征单元, 用于在所述每个分段中, 将所述每个分段去除所有的话题、 括号、 链接以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容 特征。
11、 根据权利要求 9所述的设备, 其特征在于, 所述特征提取模块, 还包 括:
前段内容特征单元, 用于在所述每个分段中, 将所述每个分段去除所有的 话题、 括号、 链接以及账户名标识后剩余的内容按预定的字节数只取前面的部 分, 生成所述前段内容特征。
12、 根据权利要求 9所述的设备, 其特征在于, 所述特征提取模块, 还包 括:
后段内容特征单元, 用于在所述每个分段中, 将所述每个分段去除所有的 话题、 括号、 链接以及账户名标识后剩余的内容按预定的字节数只取后面的部 分, 生成所述后段内容特征。
13、 一种垃圾模板文章识别设备, 其特征在于, 所述设备包括:
一个或多个处理器; 和
存储器;
所述存储器存储有一个或多个程序, 所述一个或多个程序被配置成由所述 一个或多个处理器执行, 所述一个或多个程序包含用于进行以下操作的指令: 对符合条件的微博文章提取特征, 生成文章特征; 其中, 所述文章特征至 少包括标点特征、 话题特征、 括号特征、 链接特征以及账户名特征;
获取垃圾模板列表, 所述垃圾模板列表中包含垃圾模板特征; 所述垃圾模 板特征为出现频率达到预设阈值的文章特征且所述垃圾模板特征的提取方式与 所述文章特征的提取方式相同;
当所述文章特征与所述垃圾模板列表中的垃圾模板特征相同时, 判定所述 微博文章为垃圾模板文章。
14、 根据权利要求 13所述的设备, 其特征在于, 还包含用于进行以下操作 的指令;
将所述符合条件的微博文章中的数字以及字母去掉, 并将所述微博文章中 的各种括号中的内容去掉保留所述括号。
15、 根据权利要求 13所述的设备, 其特征在于, 还包含用于进行以下操作 的指令;
将所述符合条件的微博文章以标点进行分段, 并按顺序依次生成分段编号; 在所述每个分段中, 提取所述分段的标点, 并将提取的所述标点组成字符 串, 生成所述标点特征;
在所述每个分段中, 提取有话题的分段的话题和对应的分段编号, 并将提 取的所述话题以及所述分段编号组成字符串, 生成所述话题特征;
在所述每个分段中, 提取有括号的分段对应的分段编号和所述分段对应的 括号类型, 将提取的所述分段编号以及所述括号类型组成字符串, 生成所述括 号特征;
在所述每个分段中, 根据所述每个分段中是否有链接而生成序列, 作为所 述链接特征;
在所述每个分段中, 根据所述每个分段中是否有账户名标识而生成序列, 作为所述账户名特征。
16、 根据权利要求 15所述的设备, 其特征在于, 还包含用于进行以下操作 的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容, 按顺序拼装在一起, 生成所述内容特征。
17、 根据权利要求 15所述的设备, 其特征在于, 还包含用于进行以下操作 的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取前面的部分, 生成所述前段 内容特征。
18、 根据权利要求 15所述的设备, 其特征在于, 还包含用于进行以下操作 的指令;
在所述每个分段中, 将所述每个分段去除所有的标点、 话题、 括号、 链接 以及账户名标识后剩余的内容按预定的字节数只取后面的部分, 生成所述后段 内容特征。
PCT/CN2013/083613 2012-09-17 2013-09-17 一种垃圾模板文章识别方法和设备 Ceased WO2014040570A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US14/428,314 US9330075B2 (en) 2012-09-17 2013-09-17 Method and apparatus for identifying garbage template article

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210344209.0 2012-09-17
CN201210344209.0A CN103678373B (zh) 2012-09-17 2012-09-17 一种垃圾模板文章识别方法和设备

Publications (1)

Publication Number Publication Date
WO2014040570A1 true WO2014040570A1 (zh) 2014-03-20

Family

ID=50277651

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2013/083613 Ceased WO2014040570A1 (zh) 2012-09-17 2013-09-17 一种垃圾模板文章识别方法和设备

Country Status (3)

Country Link
US (1) US9330075B2 (zh)
CN (1) CN103678373B (zh)
WO (1) WO2014040570A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107239440A (zh) * 2017-04-21 2017-10-10 同盾科技有限公司 一种垃圾文本识别方法和装置
CN110209838A (zh) * 2019-06-10 2019-09-06 广东工业大学 一种文本模板获取方法及相关装置

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105120440B (zh) * 2015-08-26 2019-05-07 小米科技有限责任公司 信息处理方法及装置
CN107229638A (zh) * 2016-03-24 2017-10-03 北京搜狗科技发展有限公司 一种文本信息处理方法及装置
CN109033224B (zh) * 2018-06-29 2022-02-01 创新先进技术有限公司 一种风险文本识别方法和装置
CN111310465B (zh) * 2020-02-18 2021-07-23 北京字节跳动网络技术有限公司 平行语料获取方法、装置、电子设备、及存储介质
CN113535813B (zh) * 2021-06-30 2023-07-28 北京百度网讯科技有限公司 一种数据挖掘方法、装置、电子设备以及存储介质

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101350032A (zh) * 2008-09-23 2009-01-21 胡辉 判断网页内容是否相同的方法
CN101661468A (zh) * 2008-08-29 2010-03-03 中国科学院计算技术研究所 一种从论坛帖子列表页面中抽取帖子元数据的方法
CN101859309A (zh) * 2009-04-07 2010-10-13 慧科讯业有限公司 重复文本识别系统及方法
US20110185236A1 (en) * 2010-01-28 2011-07-28 Fujitsu Limited Common trouble case data generating method and non-transitory computer-readable medium storing common trouble case data generating program
CN102622365A (zh) * 2011-01-28 2012-08-01 北京百度网讯科技有限公司 一种网页重复的判断系统及其判断方法
CN102662965A (zh) * 2012-03-07 2012-09-12 上海引跑信息科技有限公司 一种自动发现互联网热点新闻主题的方法及系统

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050015626A1 (en) * 2003-07-15 2005-01-20 Chasin C. Scott System and method for identifying and filtering junk e-mail messages or spam based on URL content
US7912907B1 (en) * 2005-10-07 2011-03-22 Symantec Corporation Spam email detection based on n-grams with feature selection
CN101393555A (zh) * 2008-09-09 2009-03-25 浙江大学 一种垃圾博客检测方法
CN101404037B (zh) * 2008-11-18 2011-05-18 西安交通大学 一种检测及定位电子文本内容剽窃的方法
CN101706807B (zh) * 2009-11-27 2011-06-01 清华大学 一种中文网页新词自动获取方法
US20110271179A1 (en) * 2010-04-28 2011-11-03 Peter Jasko Methods and systems for graphically visualizing text documents

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101661468A (zh) * 2008-08-29 2010-03-03 中国科学院计算技术研究所 一种从论坛帖子列表页面中抽取帖子元数据的方法
CN101350032A (zh) * 2008-09-23 2009-01-21 胡辉 判断网页内容是否相同的方法
CN101859309A (zh) * 2009-04-07 2010-10-13 慧科讯业有限公司 重复文本识别系统及方法
US20110185236A1 (en) * 2010-01-28 2011-07-28 Fujitsu Limited Common trouble case data generating method and non-transitory computer-readable medium storing common trouble case data generating program
CN102622365A (zh) * 2011-01-28 2012-08-01 北京百度网讯科技有限公司 一种网页重复的判断系统及其判断方法
CN102662965A (zh) * 2012-03-07 2012-09-12 上海引跑信息科技有限公司 一种自动发现互联网热点新闻主题的方法及系统

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107239440A (zh) * 2017-04-21 2017-10-10 同盾科技有限公司 一种垃圾文本识别方法和装置
CN110209838A (zh) * 2019-06-10 2019-09-06 广东工业大学 一种文本模板获取方法及相关装置

Also Published As

Publication number Publication date
CN103678373B (zh) 2017-11-17
US9330075B2 (en) 2016-05-03
US20150227497A1 (en) 2015-08-13
CN103678373A (zh) 2014-03-26

Similar Documents

Publication Publication Date Title
WO2014040570A1 (zh) 一种垃圾模板文章识别方法和设备
CN112035669A (zh) 基于传播异质图建模的社交媒体多模态谣言检测方法
WO2018032937A1 (zh) 一种文本信息分类方法及其装置
JP5534280B2 (ja) テキストクラスタリング装置、テキストクラスタリング方法、およびプログラム
WO2022116435A1 (zh) 标题生成方法、装置、电子设备及存储介质
CN108920675B (zh) 一种信息处理的方法、装置、计算机存储介质及终端
CN104052714B (zh) 多媒体信息的推送方法及服务器
WO2018205389A1 (zh) 语音识别方法、系统、电子装置及介质
CN106776544A (zh) 人物关系识别方法及装置和分词方法
CN112199606B (zh) 一种基于层次用户表示的面向社交媒体的谣言检测系统
CN106407484A (zh) 一种基于弹幕语义关联的视频标签提取方法
CN106547875B (zh) 一种基于情感分析和标签的微博在线突发事件检测方法
CN106886567A (zh) 基于语义扩展的微博突发事件检测方法及装置
CN112287914A (zh) Ppt视频段提取方法、装置、设备及介质
CN103778200A (zh) 一种报文信息源抽取方法及其系统
CN114386392B (zh) 文案生成方法、装置、设备及存储介质
CN109446299B (zh) 基于事件识别的搜索电子邮件内容的方法及系统
CN105224604A (zh) 一种基于堆优化的微博突发事件检测方法及其检测装置
CN110457711A (zh) 一种基于主题词的社交媒体事件主题识别方法
WO2018205458A1 (zh) 获取目标用户的方法、装置、电子设备及介质
CN110688455A (zh) 基于人工智能过滤无效评论的方法、介质及计算机设备
CN114118937A (zh) 基于任务的信息推荐方法、装置、电子设备及存储介质
CN112328735A (zh) 热点话题确定方法、装置及终端设备
CN116244423A (zh) 实体关系挖掘方法、装置、电子设备及存储介质
JP2019520614A (ja) Sns情報に基づくリスクイベント認識システム、方法、電子装置及び記憶媒体

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13837007

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 14428314

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205N DATED 27/05/2015)

122 Ep: pct application non-entry in european phase

Ref document number: 13837007

Country of ref document: EP

Kind code of ref document: A1