WO2014040570A1 - 一种垃圾模板文章识别方法和设备 - Google Patents
一种垃圾模板文章识别方法和设备 Download PDFInfo
- Publication number
- WO2014040570A1 WO2014040570A1 PCT/CN2013/083613 CN2013083613W WO2014040570A1 WO 2014040570 A1 WO2014040570 A1 WO 2014040570A1 CN 2013083613 W CN2013083613 W CN 2013083613W WO 2014040570 A1 WO2014040570 A1 WO 2014040570A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- segment
- features
- article
- feature
- template
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/35—Clustering; Classification
- G06F16/353—Clustering; Classification into predefined classes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/951—Indexing; Web crawling techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/953—Querying, e.g. by the use of web search engines
- G06F16/9535—Search customisation based on user profiles and personalisation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/26—Techniques for post-processing, e.g. correcting the recognition result
- G06V30/262—Techniques for post-processing, e.g. correcting the recognition result using context analysis, e.g. lexical, syntactic or semantic context
- G06V30/268—Lexical context
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/40—Document-oriented image-based pattern recognition
- G06V30/41—Analysis of document content
- G06V30/418—Document matching, e.g. of document images
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/21—Monitoring or handling of messages
- H04L51/212—Monitoring or handling of messages using filtering or selective blocking
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/52—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L63/00—Network architectures or network communication protocols for network security
- H04L63/14—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic
- H04L63/1441—Countermeasures against malicious traffic
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
- H04L67/01—Protocols
- H04L67/10—Protocols in which an application is distributed across nodes in the network
Definitions
- the invention relates to the field of network communication, in particular to a garbage template article identification method and device. Background technique
- Weibo APP applications
- Similar template articles which caused a large number of spam templates in the Weibo platform.
- These junk template articles are generally duplicated, or some texts are randomly modified according to the personal information or a certain law of the forwarder.
- the amount of information contained is very small, but the amount of data is very large.
- the garbage template article accounts for the total amount. 10% of the blog post. If these garbage template articles are not recognized and filtered, the search engine resources will be greatly wasted, and a large number of duplicate templates will seriously affect the user experience.
- the same type of junk template article has some common features. At present, it mainly analyzes the semantics of the article by manual to determine whether a microblog article is a spam template article.
- the manual identification method is slow in speed, low in efficiency, and unable to cope with the huge amount of data on the Weibo platform, and it is impossible for each Weibo article to be Perform manual identification.
- the embodiment of the present invention provides a garbage template article identification method and device.
- the technical solution is as follows:
- the embodiment of the present invention provides a method for identifying a spam template article, the method comprising: extracting a feature from an eligible microblog article, and generating an article feature; wherein the article feature includes at least a punctuation feature, Topic features, parenthesis features, link features, and account name characteristics;
- Obtaining a garbage template list where the garbage template list includes a garbage template feature;
- the garbage template feature is an article feature whose frequency reaches a preset threshold, and the garbage template feature is extracted in the same manner as the article feature is extracted;
- the microblog article is determined to be a junk template article.
- the qualified microblog article is an original form and includes a microblog article of a link and a picture
- the extracting the feature to the qualified microblog article further includes:
- the extracting the feature of the qualified microblog article includes:
- the qualified microblog articles are segmented by punctuation, and the segment numbers are sequentially generated in order;
- a topic with a segment of the topic and a corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a character string to generate the topic feature;
- a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a character string to generate the Bracket characteristics;
- a sequence is generated as a link feature according to whether there is a link in each segment;
- a sequence is generated as the account name feature based on whether there is an account name identifier in each of the segments.
- the article feature further includes a content feature
- the extracting the feature of the qualified microblog article further includes:
- each of the segments is removed from all punctuation, topics, parentheses, links, and contents remaining after the account name is identified, and assembled in order to generate the content features.
- the article feature further includes a previous content feature
- the matching the microblog article extraction feature includes:
- each of the segments is removed from all the punctuation, the topic, the parentheses, the link, and the content remaining after the account name identifier, and only the previous portion is taken in a predetermined number of bytes to generate the previous segment.
- Content characteristics
- the feature of the article further includes a content feature of the subsequent segment
- the extracting the feature of the microblog article that meets the condition further includes:
- each of the segments is removed from all punctuation, topics, parentheses, links, and account name identifiers, and the remaining portions are taken only by the predetermined number of bytes, and the generated portions are generated. Segment content characteristics.
- the embodiment of the present invention further provides a junk template article identification device, where the device includes: a feature extraction module, configured to extract features from the qualified microblog articles, and generate article features; wherein the article features include at least punctuation features , topic features, bracket features, link features, and account name characteristics;
- An acquisition module configured to obtain a garbage template list, where the garbage template list includes a garbage template feature;
- the garbage template feature is an article feature whose appearance frequency reaches a preset threshold, and the garbage template feature extraction manner and the article feature The same method of extraction;
- the identification module is configured to determine that the microblog article is a junk template article when the article feature is the same as the junk template feature in the junk template list.
- the device further includes:
- a pre-processing module configured to remove numbers and letters in the Weibo article before extracting features from the qualified Weibo article, and remove the contents of the various brackets in the Weibo article to retain the brackets
- the qualified Weibo article is an original form and contains Weibo articles of links and pictures.
- the feature extraction module includes:
- a segmentation unit configured to segment the qualified microblog articles by punctuation, and sequentially generate segment numbers in sequence
- a punctuation feature unit configured to extract punctuation of the segment in each segment, and form the extracted punctuation into a character string to generate the punctuation feature
- a topic feature unit configured to extract, in each of the segments, a topic of a segment with a topic and a feature of the corresponding topic
- a parenthesis feature unit configured to extract a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment in each segment, and combine the extracted segment number and the parenthesis type a string, generating the bracket feature;
- a link feature unit configured to generate, in each of the segments, a sequence according to whether there is a link in each segment, as the link feature;
- the account name feature unit is configured to generate, in each of the segments, a sequence according to whether there is an account name identifier in each segment, as the account name feature.
- the feature extraction module further includes:
- a content feature unit configured to, in each of the segments, remove all topics, parentheses, links, and content remaining after the account name identification, and assemble the pieces in order to generate the content features.
- the feature extraction module further includes:
- a preceding content feature unit configured to remove all topics, parentheses, links, and account name identifiers in each of the segments, and take only the previous portion by a predetermined number of bytes , generating the previous piece of content features.
- the feature extraction module further includes:
- a subsequent content feature unit configured to remove all topics, parentheses, links, and account name identifiers in each of the segments by a predetermined number of bytes In part, generating the latter piece of content features.
- An embodiment of the present invention further provides a garbage template article identification device, where the device includes: one or more processors; and
- the memory stores one or more programs, the one or more programs being configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations:
- the conditional microblog article extracts features, and generates article features; wherein, the article features include at least punctuation features, topic features, parenthesis features, link features, and account name features;
- the garbage template list includes a garbage template feature
- the garbage template feature is an article feature whose frequency reaches a preset threshold, and the garbage template feature is extracted in the same manner as the article feature is extracted;
- the microblog article is determined to be a junk template article.
- an instruction for performing the following operations is further included;
- an instruction for performing the following operations is further included;
- a topic with a segment of the topic and a corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a character string to generate the topic feature;
- a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a character string to generate the Bracket characteristics;
- a sequence is generated as a link feature according to whether there is a link in each segment;
- a sequence is generated as the account name feature based on whether there is an account name identifier in each of the segments.
- an instruction for performing the following operations is further included;
- each of the segments is removed from all punctuation, topics, parentheses, links, and contents remaining after the account name is identified, and assembled in order to generate the content features.
- an instruction for performing the following operations is further included;
- each of the segments is removed from all the punctuation, the topic, the parentheses, the link, and the content remaining after the account name identifier, and only the previous portion is taken in a predetermined number of bytes to generate the previous segment.
- Content characteristics
- an instruction for performing the following operations is further included;
- each of the segments is removed from all punctuation, topics, parentheses, links, and account name identifiers, and the remaining portions are taken only by the predetermined number of bytes, and the generated portions are generated. Segment content characteristics.
- the garbage template article identification method and device can determine whether the microblog article is a spam template article by extracting multiple features of the microblog article, and solves the problem that a large number of spam template articles in the current microblog platform cannot be effectively identified.
- the problem is that the effect of the spam template article can be automatically and efficiently recognized automatically and without the need of manual, only need to extract the article features in the microblog article to do the logic operation.
- FIG. 1 is a flowchart of a garbage template article identification method according to an embodiment of the present invention
- FIG. 2 is a flowchart of a garbage template article identification method according to another embodiment of the present invention
- FIG. 3 is an embodiment of the present invention
- a schematic diagram of a garbage template article identification device provided by the example
- FIG. 4 is a schematic diagram of another structure of a junk template article identification device according to another embodiment of the present invention.
- FIG. 5 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention. detailed description
- FIG. 1 is a flowchart of a garbage template article identification method according to an embodiment of the present invention.
- the embodiment includes:
- the garbage template list includes a garbage template feature;
- the garbage template feature is an article feature whose appearance frequency reaches a preset threshold, and the garbage template feature extraction manner is the same as the microblog article feature extraction manner;
- the microblog article feature is the article feature extracted by the microblog article in step 101.
- the microblog article is determined to be a junk template article.
- the qualified microblog article in the embodiment of the present invention is an original form and includes a microblog article of a link and a picture. Before extracting the feature from the qualified microblog article, the method further includes:
- extracting features for qualified Weibo articles including: The qualified microblog articles are segmented by punctuation, and the segment numbers are sequentially generated in order; in each segment, the segmentation punctuation is extracted, and the extracted punctuation marks are formed into strings to generate punctuation features;
- each segment the topic of the segment with the topic and the corresponding segment number are extracted, and the extracted topic and the segment number are grouped into a string to generate a topic feature;
- each segment the segment number corresponding to the segment with the parentheses and the parenthesis type corresponding to the segment are extracted, and the extracted segment number and the parenthesis type are combined into a string to generate a bracket feature;
- each segment a sequence is generated as a link feature according to whether there is a link in each segment; in each segment, a sequence is generated according to whether there is an account name identifier in each segment, as an account name feature .
- the feature of the article further includes content features, and extracting features of the qualified microblog articles, and the following:
- each segment is removed from all punctuation, topics, parentheses, links, and the remaining content after the account name is identified, assembled in order, to generate content features.
- the feature of the article further includes the feature of the previous paragraph, and extracts features for the qualified microblog articles, and further includes:
- each segment is removed from all punctuation, topics, parentheses, links, and the remaining contents of the account name to take only the previous portion in a predetermined number of bytes to generate the previous segment content feature.
- the article feature further includes a later content feature, and extracts characteristics for the qualified microblog article, and further includes:
- the predetermined operation includes not displaying, not feeding back to the terminal, deleting, masking, and alerting as a search result.
- the garbage template article identification method determines whether the article is a garbage template article by using multiple features of the microblog article, and solves the problem that a large number of garbage template articles in the current microblog platform cannot be effectively identified, and the problem is achieved. You don't need to be artificial, you only need to extract the article features in the Weibo article to do logical operations, and you can automatically and accurately identify the effect of the spam template article.
- FIG. 2 is a flowchart of a garbage template article identification method according to another embodiment of the present invention. Referring to FIG. 2, the embodiment includes: 201. Obtain a garbage template article, and perform preprocessing and feature extraction on the garbage template article respectively, and generate a garbage template feature to be stored in the garbage template list;
- the step may include two sub-steps of pre-processing and feature extraction:
- Spam template articles are generally original and contain links and images, remove the numbers and letters from the Weibo article, and remove the brackets from the various brackets in the Weibo article.
- the pre-processed junk template article is segmented by punctuation such as a comma, a period, an exclamation point, a question mark, and a semicolon, and the segment numbers are sequentially generated in order;
- each segment extract the segmentation punctuation in each segment in order, and form the extracted punctuation into a string to generate punctuation features
- each segment determines whether there is a topic. If there is a topic in the segment, extract the topic corresponding to the segment and the corresponding segment number, and form the extracted topic and the segment number into a string to generate Topic features; for example, there are #topic 1# in the second segment and #topic2# in the fourth segment, then "topic 1 , 2; topic 2, 4" is generated;
- each segment extract the segment number corresponding to the segment with the parentheses and the parenthesis type corresponding to the segment, and form the extracted segment number and the bracket type into a string to generate a bracket feature; for example, the first In the segment ( ), ⁇ ⁇ in the third segment, then generate "1 ( ), 3 ⁇ ⁇ ";
- each segment a sequence is generated according to whether there is a link in each segment as a link feature; for example, if there is a link in the first and second segments, it is 1, and if there is no segment in the third and fourth segments, The link is 0, generating "1100";
- each segment a sequence is generated according to whether there is an account name identifier in each segment, as an account name feature; for example, if there is an account name identifier in the first and third segments, it is 1, 2, 4 If there is no account name identifier in the segment, it is 0, and "1010" is generated; f. In each segment, remove all topics, parentheses, links, and remaining contents after the account name identification in each segment, and assemble them in order to generate content features;
- each segment remove all topics, parentheses, links, and account name identifiers for each segment, and take only the previous portion in a predetermined number of bytes to generate the previous content features; for example, The first 4 bytes of the content, generating the front-end content features;
- each segment remove all topics, parentheses, links, and account name identifiers for each segment, and take only the following parts according to the predetermined number of bytes to generate the content of the latter segment; for example, Take the last 4 bytes of the content, and generate the back-end content features;
- the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the content feature may be combined in sequence to generate a junk template feature including the content feature;
- the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the previous content feature may also be combined in order to generate a junk template feature including the previous content feature;
- Punctuation features, topic features, parenthesis features, link features, account name features, and post-content features can also be combined in order to generate junk template features containing post-content features.
- the above-mentioned punctuation feature, the topic feature, the parenthesis feature, the link feature, the account name feature, and the content feature, the previous content feature, and the subsequent content feature may be exchanged.
- it is necessary to generate a junk template feature corresponding to all content features, a junk template feature including the previous content feature, and a junk template feature including the back content feature according to the order of extracting features, and subsequent extraction of features from the microblog article The order of the steps is the same as the order in which the garbage template features are extracted.
- the garbage template feature is an article feature whose frequency reaches a preset threshold and the garbage template feature is extracted in the same manner as the subsequent microblog article feature; for example, the pre-processing is performed on the microblog article satisfying the condition every 12 hours.
- feature extraction respectively generating an article feature including content features, an article feature including the previous content feature, and an article feature including the latter content feature, and offline calculating the frequency of occurrence of each feature. When the frequency reaches the threshold, the article is identified as garbage.
- the template article, and the extracted three article features including content features, the article features including the previous content features, and the article features including the latter content features are determined as garbage template features, and are saved in the garbage template list, thereby continuously updating the garbage.
- Garbage template in template list special 202. Obtain a microblog article published by the user, and preprocess the microblog article;
- preprocessing the Weibo article includes two sub-steps as follows:
- the method for extracting features in this step is the same as step 201 above, and details are not described herein again.
- the article features extracted in this step include at least: punctuation feature, topic feature, parenthesis feature, link feature, and account name feature, wherein the content feature, the previous content feature, and the subsequent content feature of the microblog article may also be extracted.
- the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the content feature of the extracted microblog article may be combined to generate all article features in sequence;
- the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the previous content feature of the extracted microblog article may be combined to generate the previous article feature in sequence;
- the punctuation feature, the topic feature, the bracket feature, the link feature, the account name feature, and the subsequent content feature of the extracted microblog article may also be combined in order to generate a poster feature.
- the above-mentioned punctuation feature, the topic feature, the parenthesis feature, the link feature, the account name feature, and the content feature, the previous content feature, and the subsequent content feature may be exchanged.
- the garbage template feature included in the garbage template list generated in step 201 Specifically, the garbage template feature that includes all the content features generated in step 201, the garbage template feature that includes the previous content feature, and the garbage that includes the back content feature are obtained. Template feature. 205. When the feature of the article is the same as the feature of the junk template in the junk template list, determine that the microblog article is a junk template article;
- the microblog article is determined to be a spam template article; specifically,
- the microblog article is determined to be a junk template article
- the microblog article is determined to be a spam template article
- the microblog article is determined to be a spam template article.
- the microblog article is determined to be a junk template article; if only all the article features match the junk template feature including all the content features, then it may be caused by a different name.
- the Weibo article which is originally the same template, cannot be identified. Therefore, the garbage template feature including the previous content feature and the garbage template feature including the back content feature can be added to identify this, which can increase the recall rate of the template recognition. It may also lead to misjudgment, but because of the joint judgment with punctuation features, topic features, bracket features, link features, and account name characteristics, the probability of misjudgment is still relatively low.
- microblog article is determined to be a junk template article
- the microblog article is not displayed when the microblog article is retrieved.
- FIG. 3 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention.
- the device includes:
- the feature extraction module 301 is configured to extract features from the qualified microblog articles and generate article features; wherein the article features include at least punctuation features, topic features, parenthesis features, link features, and account name features;
- the obtaining module 302 is configured to obtain a garbage template list, and the garbage template list includes a garbage template feature; the garbage template feature is an article feature whose frequency reaches a preset threshold and the garbage template feature is extracted in the same manner as the microblog article feature;
- the identification module 303 is configured to determine that the microblog article is a spam template article when the article feature is the same as the junk template feature in the junk template list.
- the device further includes: a pre-processing module 304, as shown in FIG. 4;
- the pre-processing module 304 is configured to remove the numbers and letters in the microblog article before extracting the feature from the qualified microblog article, and remove the brackets in the various brackets in the microblog article;
- the blog post is a microblog article that is original and contains links and images.
- the feature extraction module 301 includes:
- a segmentation unit configured to segment the qualified microblog articles by punctuation, and sequentially generate the segment numbers in order
- Punctuation feature unit used to extract the punctuation of the segment in each segment, and group the extracted punctuation into a string to generate punctuation features
- a topic feature unit configured to extract, in each segment, a topic with a segment of the topic and a corresponding segment number, and form the extracted topic and the segment number into a string to generate a topic feature
- a parenthesis feature unit in each segment, extracting a segment number corresponding to the parenthesized segment and a parenthesis type corresponding to the segment, and forming the extracted segment number and the parenthesis type into a string to generate a parenthesis feature;
- a link feature unit configured to generate a sequence in each segment according to whether there is a link in each segment, as a link feature
- the account name feature unit is configured to generate a sequence in each segment according to whether there is an account name identifier in each segment, as an account name feature.
- the feature extraction module 301 further includes:
- the content feature unit is configured to, in each segment, remove all topics, parentheses, links, and remaining contents after the account name identification in each segment, and assemble them in order to generate content features.
- the feature extraction module 301 further includes:
- the previous piece of content feature unit used to remove all topics in each segment, including After the number, link, and account name are identified, only the previous part is taken in a predetermined number of bytes to generate the previous content feature.
- the feature extraction module 301 further includes:
- the following content feature unit is used to remove all the topics, parentheses, links, and account name identifiers in each segment, and the remaining content is only taken in the predetermined number of bytes, and then generated. Segment content characteristics.
- the junk template article identification device determines whether the microblog article is a junk template article and does not display the microblog article determined as the junk template article by extracting multiple features of the microblog article, and solves the current micro A large number of junk template articles in the blog platform can not be effectively identified, and the effect of the spam template article can be automatically and efficiently recognized automatically without the need for manual, only need to extract the article features in the microblog article to perform logical operations.
- FIG. 5 is a schematic structural diagram of a garbage template article identification device according to an embodiment of the present invention.
- the junk template article identification device 500 can be a server, and the junk template article identification device 500 includes a central processing unit (CPU) 501, a system memory 504 including a random access memory (RAM) 502 and a read only memory (ROM) 503. And a system bus 505 that connects system memory 504 and central processing unit 501.
- the junk template article identification device 500 also includes a basic input/output system (I/O system) 506 that facilitates transfer of information between various devices within the computer, and for storing the operating system 513, applications 514, and other program modules 515.
- the basic input/output system 506 includes a display 508 for displaying information and an input device 509 such as a mouse, keyboard for inputting information by the user. Both the display 508 and the input device 509 are connected to the central processing unit 501 via an input and output controller 510 that is coupled to the system bus 505.
- the basic input/output system 506 can also include an input and output controller 510 for receiving and processing input from a plurality of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, input-output controller 510 also provides output to a display screen, printer, or other type of output device.
- the mass storage device 507 is connected to the central processing unit 501 by a mass storage controller (not shown) connected to the system bus 505.
- the mass storage device 507 and its associated computer readable medium provide non-volatile storage for the client device 500. That is, the mass storage device 507 can include a computer readable medium (not shown) such as a hard disk or a CD-ROM drive.
- the computer readable medium can include computer storage media and communication media.
- Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.
- Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid state storage technologies, CD-ROM, DVD or other optical storage, tape cartridges, magnetic tape, disk storage or other magnetic storage devices.
- RAM random access memory
- ROM read only memory
- EPROM Erasable programmable read-only memory
- EEPROM electrically erasable programmable read-only memory
- the junk template article identification device 500 can also be operated by a remote computer connected to the network via a network such as the Internet. That is, the junk template article identification device 500 can be connected to the network 512 through a network interface unit 511 connected to the system bus 505, or can be connected to other types of networks or remote computer systems using the network interface unit 511 ( Not shown).
- the memory also includes one or more programs, the one or more programs being stored in a memory, and configured to be executed by one or more central processing units 501, the one or more programs comprising The garbage template article identification method provided by the embodiment shown in FIG. 1 and the garbage template article identification method provided by the embodiment shown in FIG.
- a person skilled in the art may understand that all or part of the steps of implementing the above embodiments may be completed by hardware, or may be instructed by a program to execute related hardware, and the program may be stored in a computer readable storage medium.
- the storage medium mentioned may be a read only memory, a magnetic disk or an optical disk or the like.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Databases & Information Systems (AREA)
- General Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Computer Networks & Wireless Communication (AREA)
- Computer Security & Cryptography (AREA)
- Multimedia (AREA)
- Data Mining & Analysis (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computing Systems (AREA)
- Computational Linguistics (AREA)
- Artificial Intelligence (AREA)
- Computer Hardware Design (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Processing Of Solid Wastes (AREA)
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US14/428,314 US9330075B2 (en) | 2012-09-17 | 2013-09-17 | Method and apparatus for identifying garbage template article |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201210344209.0 | 2012-09-17 | ||
| CN201210344209.0A CN103678373B (zh) | 2012-09-17 | 2012-09-17 | 一种垃圾模板文章识别方法和设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2014040570A1 true WO2014040570A1 (zh) | 2014-03-20 |
Family
ID=50277651
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2013/083613 Ceased WO2014040570A1 (zh) | 2012-09-17 | 2013-09-17 | 一种垃圾模板文章识别方法和设备 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US9330075B2 (zh) |
| CN (1) | CN103678373B (zh) |
| WO (1) | WO2014040570A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107239440A (zh) * | 2017-04-21 | 2017-10-10 | 同盾科技有限公司 | 一种垃圾文本识别方法和装置 |
| CN110209838A (zh) * | 2019-06-10 | 2019-09-06 | 广东工业大学 | 一种文本模板获取方法及相关装置 |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105120440B (zh) * | 2015-08-26 | 2019-05-07 | 小米科技有限责任公司 | 信息处理方法及装置 |
| CN107229638A (zh) * | 2016-03-24 | 2017-10-03 | 北京搜狗科技发展有限公司 | 一种文本信息处理方法及装置 |
| CN109033224B (zh) * | 2018-06-29 | 2022-02-01 | 创新先进技术有限公司 | 一种风险文本识别方法和装置 |
| CN111310465B (zh) * | 2020-02-18 | 2021-07-23 | 北京字节跳动网络技术有限公司 | 平行语料获取方法、装置、电子设备、及存储介质 |
| CN113535813B (zh) * | 2021-06-30 | 2023-07-28 | 北京百度网讯科技有限公司 | 一种数据挖掘方法、装置、电子设备以及存储介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101350032A (zh) * | 2008-09-23 | 2009-01-21 | 胡辉 | 判断网页内容是否相同的方法 |
| CN101661468A (zh) * | 2008-08-29 | 2010-03-03 | 中国科学院计算技术研究所 | 一种从论坛帖子列表页面中抽取帖子元数据的方法 |
| CN101859309A (zh) * | 2009-04-07 | 2010-10-13 | 慧科讯业有限公司 | 重复文本识别系统及方法 |
| US20110185236A1 (en) * | 2010-01-28 | 2011-07-28 | Fujitsu Limited | Common trouble case data generating method and non-transitory computer-readable medium storing common trouble case data generating program |
| CN102622365A (zh) * | 2011-01-28 | 2012-08-01 | 北京百度网讯科技有限公司 | 一种网页重复的判断系统及其判断方法 |
| CN102662965A (zh) * | 2012-03-07 | 2012-09-12 | 上海引跑信息科技有限公司 | 一种自动发现互联网热点新闻主题的方法及系统 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050015626A1 (en) * | 2003-07-15 | 2005-01-20 | Chasin C. Scott | System and method for identifying and filtering junk e-mail messages or spam based on URL content |
| US7912907B1 (en) * | 2005-10-07 | 2011-03-22 | Symantec Corporation | Spam email detection based on n-grams with feature selection |
| CN101393555A (zh) * | 2008-09-09 | 2009-03-25 | 浙江大学 | 一种垃圾博客检测方法 |
| CN101404037B (zh) * | 2008-11-18 | 2011-05-18 | 西安交通大学 | 一种检测及定位电子文本内容剽窃的方法 |
| CN101706807B (zh) * | 2009-11-27 | 2011-06-01 | 清华大学 | 一种中文网页新词自动获取方法 |
| US20110271179A1 (en) * | 2010-04-28 | 2011-11-03 | Peter Jasko | Methods and systems for graphically visualizing text documents |
-
2012
- 2012-09-17 CN CN201210344209.0A patent/CN103678373B/zh active Active
-
2013
- 2013-09-17 WO PCT/CN2013/083613 patent/WO2014040570A1/zh not_active Ceased
- 2013-09-17 US US14/428,314 patent/US9330075B2/en active Active
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101661468A (zh) * | 2008-08-29 | 2010-03-03 | 中国科学院计算技术研究所 | 一种从论坛帖子列表页面中抽取帖子元数据的方法 |
| CN101350032A (zh) * | 2008-09-23 | 2009-01-21 | 胡辉 | 判断网页内容是否相同的方法 |
| CN101859309A (zh) * | 2009-04-07 | 2010-10-13 | 慧科讯业有限公司 | 重复文本识别系统及方法 |
| US20110185236A1 (en) * | 2010-01-28 | 2011-07-28 | Fujitsu Limited | Common trouble case data generating method and non-transitory computer-readable medium storing common trouble case data generating program |
| CN102622365A (zh) * | 2011-01-28 | 2012-08-01 | 北京百度网讯科技有限公司 | 一种网页重复的判断系统及其判断方法 |
| CN102662965A (zh) * | 2012-03-07 | 2012-09-12 | 上海引跑信息科技有限公司 | 一种自动发现互联网热点新闻主题的方法及系统 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107239440A (zh) * | 2017-04-21 | 2017-10-10 | 同盾科技有限公司 | 一种垃圾文本识别方法和装置 |
| CN110209838A (zh) * | 2019-06-10 | 2019-09-06 | 广东工业大学 | 一种文本模板获取方法及相关装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN103678373B (zh) | 2017-11-17 |
| US9330075B2 (en) | 2016-05-03 |
| US20150227497A1 (en) | 2015-08-13 |
| CN103678373A (zh) | 2014-03-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2014040570A1 (zh) | 一种垃圾模板文章识别方法和设备 | |
| CN112035669A (zh) | 基于传播异质图建模的社交媒体多模态谣言检测方法 | |
| WO2018032937A1 (zh) | 一种文本信息分类方法及其装置 | |
| JP5534280B2 (ja) | テキストクラスタリング装置、テキストクラスタリング方法、およびプログラム | |
| WO2022116435A1 (zh) | 标题生成方法、装置、电子设备及存储介质 | |
| CN108920675B (zh) | 一种信息处理的方法、装置、计算机存储介质及终端 | |
| CN104052714B (zh) | 多媒体信息的推送方法及服务器 | |
| WO2018205389A1 (zh) | 语音识别方法、系统、电子装置及介质 | |
| CN106776544A (zh) | 人物关系识别方法及装置和分词方法 | |
| CN112199606B (zh) | 一种基于层次用户表示的面向社交媒体的谣言检测系统 | |
| CN106407484A (zh) | 一种基于弹幕语义关联的视频标签提取方法 | |
| CN106547875B (zh) | 一种基于情感分析和标签的微博在线突发事件检测方法 | |
| CN106886567A (zh) | 基于语义扩展的微博突发事件检测方法及装置 | |
| CN112287914A (zh) | Ppt视频段提取方法、装置、设备及介质 | |
| CN103778200A (zh) | 一种报文信息源抽取方法及其系统 | |
| CN114386392B (zh) | 文案生成方法、装置、设备及存储介质 | |
| CN109446299B (zh) | 基于事件识别的搜索电子邮件内容的方法及系统 | |
| CN105224604A (zh) | 一种基于堆优化的微博突发事件检测方法及其检测装置 | |
| CN110457711A (zh) | 一种基于主题词的社交媒体事件主题识别方法 | |
| WO2018205458A1 (zh) | 获取目标用户的方法、装置、电子设备及介质 | |
| CN110688455A (zh) | 基于人工智能过滤无效评论的方法、介质及计算机设备 | |
| CN114118937A (zh) | 基于任务的信息推荐方法、装置、电子设备及存储介质 | |
| CN112328735A (zh) | 热点话题确定方法、装置及终端设备 | |
| CN116244423A (zh) | 实体关系挖掘方法、装置、电子设备及存储介质 | |
| JP2019520614A (ja) | Sns情報に基づくリスクイベント認識システム、方法、電子装置及び記憶媒体 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 13837007 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 14428314 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205N DATED 27/05/2015) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 13837007 Country of ref document: EP Kind code of ref document: A1 |