WO2012116587A1 - 相似邮件处理系统和方法 - Google Patents
相似邮件处理系统和方法 Download PDFInfo
- Publication number
- WO2012116587A1 WO2012116587A1 PCT/CN2012/070816 CN2012070816W WO2012116587A1 WO 2012116587 A1 WO2012116587 A1 WO 2012116587A1 CN 2012070816 W CN2012070816 W CN 2012070816W WO 2012116587 A1 WO2012116587 A1 WO 2012116587A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sample
- preset format
- similar
- preset
- original sample
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/42—Mailbox-related aspects, e.g. synchronisation of mailboxes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/21—Monitoring or handling of messages
- H04L51/216—Handling conversation history, e.g. grouping of messages in sessions or threads
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/06—Message adaptation to terminal or network requirements
- H04L51/066—Format adaptation, e.g. format conversion or compression
Definitions
- the present invention relates to the field of network technologies, and in particular, to a similar mail processing system and method. Background technique
- the spam system from statistics to interception, has a mature architecture.
- the system is based on a single-machine computing model. It can count a certain number of emails in a short period of time, and obtain similar relationships and similarities between emails. index. Because the system can identify spam that has been deformed by a certain amount and added interference elements, in practical applications, it has excellent indicators in terms of the size, quantity and accuracy of intercepting spam.
- the similar mail processing system in the prior art is based on a stand-alone computing mode, and has a large limitation on the size of input data and output data that can be processed.
- the operation data size of a single million or more input data has a slow operation speed and a high system load.
- the problem, unable to achieve real-time, can not be achieved in quasi-real-time statistics due to the long completion time. Summary of the invention
- Embodiments of the present invention provide a similar mail processing system and method.
- the technical solution is as follows:
- a similar mail processing system includes:
- control node configured to receive a sample in a preset format, and determine whether the sample in the preset format is a final result of the similar calculation, and if not, merge or split the sample in the preset format according to a preset criterion.
- the plurality of similar operation nodes are configured to perform a similarity calculation on the samples in the received subtask data packet to obtain a similar calculation intermediate result, where the similar calculation intermediate result is a preset format, and the similar calculation intermediate result is Feedback to the control node, the similar calculation intermediate result includes at least: a unique similar sample, a similarity relationship, and a similarity count of the unique similar sample.
- the system also includes:
- a data input node configured to collect the original sample and convert the original sample into a preset format, and send the converted original sample package to the control node as a sample of a preset format.
- the data input node includes:
- a data collection module configured to collect mails on a server or server cluster of a similar mail processing system, and use the mail as an original sample
- a conversion module configured to convert the original sample into a preset format that matches a similar calculation
- a sending module configured to allocate a task identifier to the converted original sample package, and send the converted original sample package as a sample of the preset format to the control node as a whole or in batches.
- the sending module includes:
- An optimized transmission unit configured to split the converted original sample packet into a plurality of data packets according to a network condition
- a sending unit configured to use the multiple data packets output by the optimized transmission unit as a preset format
- the samples are sent to the control node in batches.
- the control node includes:
- a receiving module configured to receive a sample in a preset format
- a determining module configured to determine whether the sample of the preset format meets a preset condition, and if yes, the sample of the preset format is a final result of the similarity calculation, and if not, the sample of the preset format is not a similar calculation The final result, and trigger the merge split module;
- the merge splitting module is configured to combine or split the samples of the preset format according to the heartbeat information of the similar computing node to obtain a plurality of subtask data packets; the heartbeat information is used for monitoring and describing The idle computing capability of the similar computing node;
- an allocating module configured to allocate the plurality of subtask data packets obtained by the merge splitting module to each similar operating node.
- the merge splitting module is specifically configured to collect data key indicators of the converted original sample package and the sample of the preset format, and perform the converted original according to the configuration file registration information and the data key indicator.
- the sample package and the sample of the preset format are sorted, and the converted original sample package or the sample of the preset format is combined or split according to a sorting order to obtain a plurality of subtask data packets.
- the control node further includes:
- the heartbeat information monitoring module is configured to acquire heartbeat information of the similar operation node every preset time period or when receiving samples of a preset format.
- the control node is further configured to save and record a sample of the preset format, record a mapping relationship between the plurality of subtask data packets and a similar operation node allocated by the subtask data packet, and record the similar operation node Heartbeat information.
- the heartbeat information monitoring module is further configured to: when the similar operation node does not return heartbeat information within a preset time period and continuously returns the heartbeat information for more than a preset number of times, mark the similar operation node to collapse, and mark the Similar operation section
- the subtask data packet running on the point fails, and the allocation module is triggered to allocate the subtask data packet with the failed label according to the heartbeat information of the similar operation node to the similar computing node that is not crashed and idle.
- a similar mail processing method including:
- the similarity calculation intermediate result is a sample of a preset format
- the sample of the preset format is fed back
- the result includes at least: a unique similar sample, a similarity relationship, and a similar count of the unique similar sample.
- Receive samples of the original samples and preset formats including:
- Determining whether the converted original sample package and the sample in the preset format are the similar result of the calculation specifically: determining whether the converted original sample package meets a preset condition, if the converted original sample package satisfies Presetting the condition, the converted original sample package is a similar calculation final result, and if the converted original sample does not satisfy the preset condition, the converted original sample package is not a similar calculation final result;
- the sample of the preset format is a similar result of the similar calculation, if the sample of the preset format is not If the preset condition is met, the sample of the preset format is not the final result of the similar calculation.
- the merged original sample package and the sample of the preset format are combined or split according to a preset standard, and multiple subtask data packets are obtained, which specifically includes:
- the sample of the preset format is a sample of at least one similarly calculated sample and there are at least two samples of a preset format returned by the task to which the sample of the preset format belongs to the local server, the at least two of the samples are Sample of preset format The samples of the preset format returned by the task are merged.
- the preset standard includes at least one of the following:
- the similar processing and calculation of mails of more than 10 million levels are realized by the control node merging or splitting the input samples, and allocating the obtained plurality of subtask data packets to a distributed system of multiple similar operation nodes. Thereby improving the computing speed and computing power, reducing the system load, and supporting real-time and quasi-real-time statistics and interception of anti-spam requirements.
- Figure la is a schematic diagram of a similar mail processing system according to an embodiment of the present invention.
- Figure lb is a schematic diagram of a similar mail processing system according to an embodiment of the present invention.
- FIG. 2 is a flowchart of a similar mail processing method according to an embodiment of the present invention.
- FIG. 3 is a flowchart of a similar mail processing method according to an embodiment of the present invention. detailed description
- the present invention is based on the following simple common sense: Spam must have a significant scale in quantity and scale, and must be identical in form. Phenomenon, it is not difficult to find that as long as we process and operate fast enough, we can identify spam (with a large number of scales) in the first time, thus implementing interception. It can be seen that the sooner a large-scale similar spam is discovered, the sooner intervention can be carried out, so that the earlier the spam is blocked outside the mailbox system (according to statistics, more than 60% of the mailbox system is spam). This is self-evident for the benefits of the user, and can also greatly reduce the operational The pressure of this (bandwidth, storage).
- the embodiment of the present invention provides a similar mail processing system.
- the system includes: a control node 101 and a plurality of similar computing nodes 102.
- the control node 101 is configured to receive a sample in a preset format, and determine whether the sample in the preset format is a final result of the similar calculation, and if not, merge the samples in the preset format according to a preset criterion or Splitting processing, obtaining a plurality of subtask data packets, and assigning the plurality of subtask data packets to the plurality of similar operation nodes;
- the plurality of similar operation nodes 102 are configured to perform a similarity calculation on the samples in the received subtask data packet to obtain a similar calculation intermediate result, where the similar calculation intermediate result is a sample of a preset format, and the pre The formatted sample is fed back to the control node, and the similarly calculated intermediate result includes at least: a unique similar sample, a similarity relationship, and a similarity count of the unique similar sample.
- the system further includes:
- the data input node 103 is configured to collect the original sample and convert the original sample into a preset format, and send the converted original sample packet to the control node as a sample of a preset format.
- the data input node 103 includes:
- the data collection module 1031 is configured to collect mails on a server or server cluster of a similar mail processing system, and use the mail as an original sample;
- the converting module 1032 is configured to convert the original sample into a preset format that matches a similar calculation
- the sending module 1033 is configured to allocate a task identifier to the converted original sample package, and send the converted original sample packet to the control node as a sample of a preset format as a whole or in batches.
- the sending module 1033 includes:
- the optimized transmission unit 1033a is configured to split the converted original sample packet into a plurality of data packets according to a network condition
- the sending unit 1033b is configured to send the plurality of data packets output by the optimized transmission unit to the control node in batches as samples of a preset format.
- the control node 101 includes:
- the receiving module 1011 is configured to receive a sample in a preset format.
- the determining module 1012 is configured to determine whether the sample of the preset format meets a preset condition, and if yes, the sample of the preset format is a final result of the similarity calculation, and if not, the sample of the preset format is not similar Calculate the final result and trigger the merge split module;
- the merge splitting module 1013 is configured to compare the heartbeat information of the similar computing node to the preset format. Performing a merge or split process to obtain a plurality of subtask data packets; the heartbeat information is used to describe an idle computing capability of the similar computing node;
- the merge splitting module 1013 is specifically configured to collect data key indicators of the converted original sample package and the sample of the preset format, and according to the configuration file registration information and the data key indicator pair Sorting the converted original sample package and the sample of the preset format, and merging or splitting the converted original sample package or the sample of the preset format according to a sorting order to obtain a plurality of Subtask packets.
- the allocating module 1014 is configured to allocate the plurality of subtask data packets obtained by the merge splitting module to the respective similar computing nodes 102.
- the control node 101 further includes:
- the heartbeat information monitoring module is configured to acquire heartbeat information of the similar operation node every preset time period or when receiving samples of a preset format.
- the control node 101 is further configured to save and record a sample of the preset format, record a mapping relationship between the plurality of subtask data packets and a similar operation node allocated by the subtask data packet, and record the similar operation node.
- Heartbeat information is further configured to save and record a sample of the preset format, record a mapping relationship between the plurality of subtask data packets and a similar operation node allocated by the subtask data packet, and record the similar operation node. Heartbeat information.
- the heartbeat information monitoring module is further configured to: when the similar operation node does not return heartbeat information within a preset time period and continuously returns the heartbeat information for more than a preset number of times, mark the similar operation node to collapse, and mark the The subtask data packet running on the similar computing node fails, and the allocation module is triggered to allocate the subtask data packet with the failed label according to the heartbeat information of the similar computing node to the similar computing node that is not crashed and idle.
- the similar processing and calculation of mails of more than 10 million levels are realized by the control node merging or splitting the input samples, and allocating the obtained plurality of subtask data packets to a distributed system of multiple similar operation nodes. Thereby improving the computing speed and computing power, reducing the system load, and supporting real-time and quasi-real-time statistics and interception of anti-spam requirements.
- the embodiment of the present invention provides a similar mail processing method, and the execution body of the method is the similar mail processing system provided by the above embodiment 1, see FIG. 2, the method Includes:
- the similar mail processing system receives the sample of the original sample and the preset format, and converts the received original sample into a preset format;
- the similar mail processing system determines whether the converted original sample package and the sample of the preset format are similar calculation final results
- the converted original sample package and the sample of the preset format are merged or split according to a preset criterion, and multiple subtask data packets are obtained; If yes, the sample of the preset format is a final result of the similarity calculation, and the sample of the preset format is output as the final result of the similarity calculation;
- the similar mail processing system performs a similarity relationship calculation on each sample in the subtask data packet, and obtains an intermediate result of the similar calculation, wherein the intermediate result of the similarity calculation is a sample of a preset format, and the sample of the preset format is fed back, the similarity
- the intermediate results of the calculation include a unique similar sample, a similarity relationship, and a similarity count for the unique similar sample.
- the sample that receives the original sample and the preset format includes:
- determining whether the converted original sample package and the sample of the preset format are the final result of the similar calculation specifically comprising:
- the sample of the preset format is a similar result of the similar calculation, if the sample of the preset format is not If the preset condition is met, the sample of the preset format is not the final result of the similar calculation.
- the original sample package and the sample of the preset format are combined or split according to a preset standard, and multiple subtask data packets are obtained, which specifically includes:
- the sample of the preset format is a sample that has undergone similar calculation at least once and there are at least two samples of a preset format returned by the task to which the sample of the preset format belongs on the local server
- the at least two presets are The samples of the preset format returned by the task to which the format belongs are merged.
- the preset standard includes at least one of the following:
- the converted original sample package is split
- the split original sample package is split
- the similar processing and calculation of mails of more than 10 million levels are realized by the control node merging or splitting the input samples, and allocating the obtained plurality of subtask data packets to a distributed system of multiple similar operation nodes. Thereby improving the computing speed and computing power, reducing the system load, and supporting real-time and quasi-real-time statistics and interception of anti-spam requirements.
- the embodiment of the present invention provides a similar mail processing method, and the execution body of the method is the different nodes of the similar mail processing system provided by the above embodiment 1, the similarity
- the mail processing system includes a data input node, a control node, and a similar computing node.
- a data input node, a control node, and four similar computing nodes are included in the similar mail processing system as an example.
- the control node can receive the original sample for conversion, and can also receive the sample from the data input node, and is converted by the data input node.
- the data input node performs conversion as an example, as shown in FIG. 3,
- An embodiment of the method specifically includes:
- the data collection module in the data input node collects the mail on the server or the server cluster of the similar mail processing system, and uses the mail as the original sample;
- the data input node is configured to collect the original sample and convert the original sample into a preset format, and send the converted original sample packet to the control node as a sample in a preset format.
- the data input node can be a server capable of communicating with the control node, or a server cluster composed of multiple servers.
- the conversion module in the data input node converts the original sample into a preset format that matches the similar calculation; it should be noted that, in the subsequent similar calculation, in order to improve the processing speed and conveniently record the processing result, the original sample is needed.
- the conversion is performed according to a similar calculation algorithm configured on a subsequent similar computing node, and is converted into a data format corresponding to the similar computing algorithm.
- the similarity calculation algorithm may be multiple, which is not limited by the present invention.
- the sending module in the data input node allocates a task identifier to the converted original sample package, and sends the converted original sample packet as a sample of the preset format to the control node as a whole or in batches;
- the task identifier is assigned to make the task that the system is running transparent, and the technician can identify the task. It is known which tasks are currently running on the system, and when it is necessary to terminate a task, the control node can send a termination instruction to the similar operation node of the subtask that is running the task according to the task identifier.
- the optimized transmission unit in the sending module splits the converted original sample packet into multiple data packets according to the network condition; and the optimization is performed by the sending unit.
- the plurality of data packets output by the transmission unit are sent to the control node in batches as samples of a preset format, occupying less memory and bandwidth resources.
- the data input node may be part of the control node, and the function of converting the format may also be performed by the control node.
- the control node includes the function, the data input node is responsible for collecting the mail, and packaging the mail as the original sample.
- the control node scans the original sample, converts the original sample into a sample of a preset format, and after performing the judgment of step 305, when the sample of the preset format is not the final result of the similar calculation, the statistical pre- Formatted key data indicators (including metrics such as packet size or record entries), sorted according to key data metrics based on sample configuration information (including the number of records included in each package or the size of each package)
- the subsequent alignment is split or merged into multiple subtask packets.
- the above steps are the processing of the original sample.
- the receiving module of the control node receives a sample in a preset format, where the sample of the preset format includes the converted original sample package and a similar calculated intermediate result fed back by the similar computing node;
- the control node is configured to receive a sample of the preset format, and determine whether the sample of the preset format is a final result of the similar calculation, and if not, merge or split the sample of the preset format according to a preset criterion, Obtaining a plurality of subtask data packets, and assigning the plurality of subtask data packets to the plurality of similar operation nodes;
- the samples of the preset format appearing in the subsequent steps may be divided into the converted original sample package converted by the data input node and the preset format sample not converted by the data input node according to the source and the processed processing steps.
- the data received by the control node is in a preset format. Therefore, the original sample packet after conversion and the sample in the preset format are not distinguished, and are collectively referred to as samples of the preset format.
- the sample is transmitted multiple times, the task life cycle is long or no termination time, and the similar relationship data that needs to be output should cover all input data, and can output similar results between the sample parts that have been transmitted without waiting for all samples. After all the transmission is completed, the similar calculation process is started;
- control node is a control part in the entire system, and the control node is further configured to process a request from a data input node.
- the request is used to request similar calculation processing on a sample of a preset format.
- the control node can verify the validity of the request. When the request verification is legal, the received sample in the preset format is processed.
- the control node is generally a server, and in the case of hot standby, it can be two or more.
- control node is further configured to save and record the sample of the preset format, record a mapping relationship between the plurality of subtask data packets and a similar operation node allocated by the subtask data packet, and record heartbeat information of the similar operation node. .
- the determining module of the control node determines whether the sample of the preset format meets a preset condition
- the sample of the preset format is a final result of the similarity calculation, and the sample of the preset format is output as the final result of the similar calculation;
- step 306 is performed
- the preset condition means that the similarity count of the sample reaches a preset threshold and the sample package has been filtered and the independent sample is excluded, and the independent sample means that it has no similar relationship with any other sample; or after the similar calculation, no new one is found.
- the similarity relationship for example, input 1000 samples, after calculation, there is no sample that can be merged, still 1000 samples.
- the preset condition is set by the technician according to the carrying capacity of the system or other elements, and is not specifically limited in the embodiment of the present invention.
- the difference between the record entries in the converted original sample package is large, and no similar calculation is needed, and at this time, after the conversion
- the original sample package can be used as the final result of the similar calculation.
- the merge splitting module of the control node combines or splits the samples of the preset format according to the heartbeat information of the similar computing node to obtain multiple subtask data packets.
- the heartbeat information is used to monitor and describe the idle computing capability of the similar computing node, including: a configuration of the CPU or memory and a computing capability and a list of currently running tasks.
- the heartbeat information monitoring module is configured to acquire heartbeat information of the similar operation node every preset time period or when receiving a sample of the preset format. Specifically, the heartbeat information monitoring module sends a heartbeat information request to the similar operation node every preset time period (for example, 1 minute) or triggers the heartbeat information monitoring module to send a heartbeat information request to the similar operation node when the control node receives the sample of the preset format.
- the similar computing node receives the heartbeat information request, it feeds back to the control node information such as the currently running subtask list.
- the heartbeat information monitoring module saves the feedback heartbeat information, periodically monitors the status of all similar computing nodes, and monitors the completion of running subtasks, including running, ending, or abnormal failures, etc., for dispatching subtask data packets and similar Compute the query processing when the node crashes.
- TCP long link is maintained between the control node and all similar computing modules.
- the number of the record entries in the sample of the preset format or the total size bytes after the data packet is exceeded exceeds a preset threshold, and the sample of the preset format is split.
- the sample of the preset format must satisfy any of the following aspects, the sample needs to be split: 1.
- the sample has been sorted according to key data indicators;
- the number of recorded entries exceeds a preset threshold, such as 100,000;
- the size of the data packet after the data packet exceeds a preset threshold, such as 1G;
- the sample when the sample must satisfy any of the following aspects, the sample needs to be merged:
- the similarity calculation is completed, and the unique sample step (that is, only one sample is retained, but the similarity index between all the samples merged and the unique sample is recorded) remains unchanged;
- the allocation module of the control node allocates the plurality of subtask data packets obtained by the merge splitting module to each similar computing node;
- the similar computing node receives one or more subtask data packets, and performs similarity calculation on the samples in the received subtask data packet to obtain a similar calculation intermediate result, where the intermediate result of the similar computing is a preset format sample, The sample of the preset format is fed back to the control node, and step 304 is performed until the task to which the sample belongs is completed.
- control node when the control node receives the sample in the preset format, it determines whether the subtask data packet in the task to which the sample belongs has been feedback according to the task identifier, and if yes, the task ends, and if not, the feedback The samples of the preset format and the subsequent input samples are then merged or split and again assigned to similar computing nodes for similar calculations.
- the similarity calculation intermediate result includes at least a unique similarity sample, a similarity relationship, and a similarity count of the unique similarity sample, and may include other information.
- the similar computing node is only responsible for the similarity calculation of the internal entries of each data packet, and feeds the similar computing intermediate results of each data packet to the control node without processing between the data packets.
- the arithmetic node unit is responsible for performing specific similar computing tasks, and does not make any changes to the original data except for the input and output of data.
- the similar computing node can be a server with different CPU computing power, and can use one or several similarly calculated core algorithms;
- the similar computing node does not actively report its own heartbeat information, and returns the necessary information to the control node only after receiving the heartbeat information request.
- each task has a maximum running time limit, that is, if the operation time exceeds the specified number of seconds, the task is invalidated, and only some similar samples complete the similar operation, and according to the configuration information of the subtask, whether to return is not required.
- the completed result is given to the control node.
- the receiving control node issues a termination instruction, the operation is immediately stopped and immediately discarded; when the subtask is completed, the similar computing node sends a request to the control node, and returns the result data, which has a timeout retry mechanism.
- the control node when the request sent by the similar computing node does not receive the feedback of the control node within the preset duration, it is resent, and when the number of retransmissions exceeds the preset number of times, the control node is considered to crash. If a similar computing node crash occurs, the data in the similar computing node and the unfinished subtask are not restored. After the similar computing node resumes the response, it waits for a new computing request;
- the original sample contains 9 samples of ABCDEFGHI, sorted according to the data key indicators, and then split into 3 packages, namely:
- a similar computing node crash occurs.
- the similar computing node does not return heartbeat information within a preset duration and continuously returns the heartbeat information for more than a preset number of times, the similar computing node is marked to crash. And marking the failure of the subtask data packet running on the similar operation node, and triggering the allocation module to assign the subtask data packet with the failed label to the uncombed and idle similar operation node according to the heartbeat information of the similar operation node .
- the similar mail processing system includes a control node and four similar computing nodes, wherein the four similar computing nodes are NodeK Node2, Node3, and Node4, and the running subtask data packets are P1, P2, P3, and P4, the subtask data packets running on similar computing nodes can be seen in Table 1 below.
- Node2 is running P3 when it crashes
- Table 2 can know that Node4 is idle, and Node3 has already finished running.
- Node3 has strong computing power, while P3 has a large amount of data, P3 is assigned to Node3 to perform similar calculations.
- the heartbeat service is used to collect the subtasks that are running at the moment, and the subtask list can be reconstructed in combination with the LOG data of the control node. It should be noted that in extreme cases, there is a possibility of losing part of the information. The missing information may be the part that has accepted the similar calculation request, but has not had time to split or has split but has not had time to distribute.
- the similar processing and calculation of mails of more than 10 million levels are realized by the control node merging or splitting the input samples, and allocating the obtained plurality of subtask data packets to a distributed system of multiple similar operation nodes.
- All or part of the above technical solutions provided by the embodiments of the present invention may be completed by hardware related to program instructions, and the program may be stored in a readable storage medium, including: wake up, RAM, disk or CD. And other media that can store program code.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Data Exchanges In Wide-Area Networks (AREA)
Abstract
Description
Claims
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| SG2013065685A SG193013A1 (en) | 2011-03-03 | 2012-02-01 | System and method for processing similar emails |
| KR1020137017886A KR101526344B1 (ko) | 2011-03-03 | 2012-02-01 | 유사 이메일을 처리하기 위한 시스템 및 방법 |
| US13/905,037 US20130282846A1 (en) | 2011-03-03 | 2013-05-29 | System and method for processing similar emails |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201110051222.2A CN102655480B (zh) | 2011-03-03 | 2011-03-03 | 相似邮件处理系统和方法 |
| CN201110051222.2 | 2011-03-03 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US13/905,037 Continuation US20130282846A1 (en) | 2011-03-03 | 2013-05-29 | System and method for processing similar emails |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2012116587A1 true WO2012116587A1 (zh) | 2012-09-07 |
Family
ID=46731006
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2012/070816 Ceased WO2012116587A1 (zh) | 2011-03-03 | 2012-02-01 | 相似邮件处理系统和方法 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US20130282846A1 (zh) |
| KR (1) | KR101526344B1 (zh) |
| CN (1) | CN102655480B (zh) |
| MY (1) | MY167496A (zh) |
| SG (1) | SG193013A1 (zh) |
| WO (1) | WO2012116587A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10347275B2 (en) | 2013-09-09 | 2019-07-09 | Huawei Technologies Co., Ltd. | Unvoiced/voiced decision for speech processing |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107087010B (zh) * | 2016-02-14 | 2020-10-27 | 阿里巴巴集团控股有限公司 | 中间数据传输方法及系统、分布式系统 |
| CN108259568B (zh) * | 2017-12-22 | 2021-05-04 | 东软集团股份有限公司 | 任务分配方法、装置、计算机可读存储介质及电子设备 |
| CN113094243B (zh) * | 2020-01-08 | 2024-08-20 | 北京小米移动软件有限公司 | 节点性能检测方法和装置 |
| CN114756366A (zh) * | 2022-04-08 | 2022-07-15 | 深圳英博达智能科技有限公司 | 一种边缘计算方法及边缘计算服务器 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1922837A (zh) * | 2004-05-14 | 2007-02-28 | 布赖特梅有限公司 | 基于相似性量度过滤垃圾邮件的方法和装置 |
| CN101159704A (zh) * | 2007-10-23 | 2008-04-09 | 浙江大学 | 基于微内容相似度的反垃圾方法 |
| US7590694B2 (en) * | 2004-01-16 | 2009-09-15 | Gozoom.Com, Inc. | System for determining degrees of similarity in email message information |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7543053B2 (en) * | 2003-03-03 | 2009-06-02 | Microsoft Corporation | Intelligent quarantining for spam prevention |
| US20050132197A1 (en) * | 2003-05-15 | 2005-06-16 | Art Medlar | Method and apparatus for a character-based comparison of documents |
| US7475118B2 (en) * | 2006-02-03 | 2009-01-06 | International Business Machines Corporation | Method for recognizing spam email |
-
2011
- 2011-03-03 CN CN201110051222.2A patent/CN102655480B/zh active Active
-
2012
- 2012-02-01 KR KR1020137017886A patent/KR101526344B1/ko active Active
- 2012-02-01 SG SG2013065685A patent/SG193013A1/en unknown
- 2012-02-01 MY MYPI2013002093A patent/MY167496A/en unknown
- 2012-02-01 WO PCT/CN2012/070816 patent/WO2012116587A1/zh not_active Ceased
-
2013
- 2013-05-29 US US13/905,037 patent/US20130282846A1/en not_active Abandoned
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7590694B2 (en) * | 2004-01-16 | 2009-09-15 | Gozoom.Com, Inc. | System for determining degrees of similarity in email message information |
| CN1922837A (zh) * | 2004-05-14 | 2007-02-28 | 布赖特梅有限公司 | 基于相似性量度过滤垃圾邮件的方法和装置 |
| CN101159704A (zh) * | 2007-10-23 | 2008-04-09 | 浙江大学 | 基于微内容相似度的反垃圾方法 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10347275B2 (en) | 2013-09-09 | 2019-07-09 | Huawei Technologies Co., Ltd. | Unvoiced/voiced decision for speech processing |
| US11328739B2 (en) | 2013-09-09 | 2022-05-10 | Huawei Technologies Co., Ltd. | Unvoiced voiced decision for speech processing cross reference to related applications |
Also Published As
| Publication number | Publication date |
|---|---|
| CN102655480B (zh) | 2015-12-02 |
| MY167496A (en) | 2018-08-30 |
| SG193013A1 (en) | 2013-10-30 |
| US20130282846A1 (en) | 2013-10-24 |
| KR20130109195A (ko) | 2013-10-07 |
| KR101526344B1 (ko) | 2015-06-05 |
| CN102655480A (zh) | 2012-09-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN105391742B (zh) | 一种基于Hadoop的分布式入侵检测系统 | |
| CN112118174B (zh) | 软件定义数据网关 | |
| CN110225074B (zh) | 一种基于设备地址域的通讯报文分发系统及分发方法 | |
| US20100229182A1 (en) | Log information issuing device, log information issuing method, and program | |
| CN103761309A (zh) | 一种运营数据处理方法及系统 | |
| CN110809060B (zh) | 一种应用服务器集群的监控系统及监控方法 | |
| WO2012116587A1 (zh) | 相似邮件处理系统和方法 | |
| CN107977167B (zh) | 一种基于纠删码的分布式存储系统的退化读优化方法 | |
| CN111726410B (zh) | 用于分散计算网络的可编程实时计算和网络负载感知方法 | |
| CN106815254A (zh) | 一种数据处理方法和装置 | |
| CN111241038A (zh) | 卫星数据处理方法及系统 | |
| CN108304293A (zh) | 一种基于大数据技术的软件系统监控方法 | |
| CN114710424B (zh) | 基于软件定义网络的主机侧数据包处理延时测量方法 | |
| CN120448213B (zh) | 一种基于消息队列的nginx日志监控方法 | |
| CN101350733B (zh) | 基于前置数据服务机的网元性能数据采集系统及实现方法 | |
| CN118626344B (zh) | Android应用ANR监控方法、装置、设备及介质 | |
| CN120614323A (zh) | 一种流数据处理的调度指令实时监控系统 | |
| CN115801562B (zh) | 一种高效可伸缩的cdn日志处理方法及系统 | |
| CN109992572A (zh) | 一种自适应均衡日志存储请求的方法 | |
| CN112003900A (zh) | 实现分布式系统中高负载场景下服务高可用的方法、系统 | |
| CN102256276B (zh) | 路测信息处理方法及装置 | |
| CN120336260B (zh) | 一种用于ftp文件的服务器智能监听方法 | |
| CN121239510B (zh) | 一种基于Flink的流量计费系统及计费方法 | |
| CN118381740B (zh) | 时序指标数据的历史监控报警回放系统 | |
| CN115865612B (zh) | 网络故障处理方法及装置、存储介质及电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 12752498 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 20137017886 Country of ref document: KR Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC OF 250214 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 12752498 Country of ref document: EP Kind code of ref document: A1 |



