WO2020062689A1 - 流量数据的聚类处理方法、装置及电子设备 - Google Patents
流量数据的聚类处理方法、装置及电子设备 Download PDFInfo
- Publication number
- WO2020062689A1 WO2020062689A1 PCT/CN2018/125246 CN2018125246W WO2020062689A1 WO 2020062689 A1 WO2020062689 A1 WO 2020062689A1 CN 2018125246 W CN2018125246 W CN 2018125246W WO 2020062689 A1 WO2020062689 A1 WO 2020062689A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- cluster
- traffic data
- data
- feature
- clusters
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/23—Clustering techniques
Definitions
- the present application relates to the field of big data technology, and in particular, to a method, an apparatus, and an electronic device for clustering processing of traffic data.
- the identification of abnormal traffic is generally determined by collecting user behavior buried points and sdk data to determine the path repeatability, the proportion of login buried points on the front and back of the device, the number of IP access accounts, the number of IP accesses, and the mobile phone in the period Features such as the mean and variance of the user's login in the number segment, and based on these characteristics of each piece of traffic data, determine the probability that the traffic data is abnormal.
- the inventor of the present application realizes that the shortcomings of the prior art are that the black industry often shows that the traffic data of the group is abnormal, and the identification of the abnormal traffic by the prior art is determined for each piece of traffic data in isolation, which cannot meet the needs of the group. Demand for overall analysis of traffic data.
- the present application provides a method, an apparatus, and an electronic device for clustering processing of traffic data.
- a clustering processing method of traffic data includes white data and black data, and the white data is a traffic extracted from data traffic of a user determined as a white user.
- Data the black data is traffic data extracted from data traffic of a user determined to be a black user, the white user is a user determined to not send abnormal traffic data, and the black user is determined to send abnormal traffic
- the method includes: selecting N features in a preset feature database, where N is a positive integer; obtaining a feature vector of the traffic data based on the feature value corresponding to the selected feature of the traffic data; the feature The vector includes feature values corresponding to the N features of the traffic data, where one of the features corresponds to one of the feature values; all the traffic data is clustered into M according to the feature vector of the traffic data Clusters, where M is a positive integer greater than or equal to 2; determine the total number of cluster errors of the cluster into which the traffic data is divided under various combinations of M and N values
- a clustering processing apparatus for traffic data includes white data and black data, and the white data is extracted from data traffic of a user determined as a white user.
- Traffic data the black data is traffic data extracted from data traffic of a user identified as a black user, the white user is a user determined to not issue abnormal traffic data, and the black user is determined to issue an exception
- the device includes: a selection unit for selecting N features from a preset feature database, where N is a positive integer; and an acquisition unit for obtaining feature values corresponding to the selected features based on the traffic data to obtain A feature vector of the traffic data; the feature vector includes a feature value corresponding to each of the N features of the traffic data; wherein one of the features corresponds to one of the feature values;
- a clustering unit is configured to The feature vector of the traffic data, clustering all the traffic data into M clusters, where M is a positive integer greater than or equal to 2; and a determining unit for determining
- the sum of the cluster error numbers is the result of adding the error numbers of each cluster divided.
- the error number of each cluster refers to the cluster. The smaller one of the number of white data and the number of black data; a setting unit for setting the number of features and the number of clusters corresponding to the sum of the smallest number of cluster errors as the target feature number selected when clustering the traffic data And the number of target clusters.
- a computer-readable storage medium which is characterized in that it stores a computer program, which when executed by a computer causes the computer to execute the method described above.
- an electronic device including: a memory, where the computer-readable instructions are stored on the memory; and a processor configured to execute the computer-readable instructions stored on the memory, When the computer-readable instructions are executed by the processor, the processor is caused to execute the method described above.
- the image control method provided in this application includes the following steps: selecting N features from a preset feature database, where N is a positive integer; and obtaining a feature vector of the traffic data based on feature values corresponding to the selected features of the traffic data;
- the feature vector includes feature values corresponding to the N features of the traffic data; one of the features corresponds to one of the feature values; and all the traffic data is aggregated according to the feature vector of the traffic data.
- the number of errors in each cluster refers to the smaller of the number of white data and the number of black data in the cluster; the number of features corresponding to the sum of the smallest cluster error numbers And the number of clusters are used as the number of target features and the number of target clusters selected when clustering the traffic data.
- the identification of the traffic anomaly is not determined in isolation, but the traffic data is divided into several clusters, and the combination of several clusters can reflect the traffic data in a group, or in an area, or in a group of people.
- the characteristics are conducive to analyzing the behavior of the black industry chain. In summary, the demand for the overall analysis of the traffic data of the group is satisfied.
- Fig. 1 is a schematic diagram showing a clustering processing apparatus for traffic data according to an exemplary embodiment
- Fig. 2 is a flow chart showing a method for processing clustering of traffic data according to an exemplary embodiment
- Fig. 3 is a flowchart showing details of step 230 according to the corresponding embodiment of Fig. 2;
- Fig. 4 is a flow chart showing a method for processing clustering of traffic data according to another exemplary embodiment
- Fig. 5 is a block diagram of a device for processing clustering of traffic data according to an exemplary embodiment.
- the implementation environment of this application may be a portable mobile device, such as a smart phone, a tablet computer, or a desktop computer.
- the clustering processing method of traffic data disclosed in this application can be applied to any application program running on a portable mobile device.
- Fig. 1 is a schematic diagram illustrating a clustering processing apparatus for traffic data according to an exemplary embodiment.
- the device 100 may be the aforementioned portable mobile device.
- the device 100 may include one or more of the following components: a processing component 102, a memory 104, a power component 106, a multimedia component 108, an audio component 110, a sensor component 114, and a communication component 116.
- the processing component 102 generally controls overall operations of the device 100, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations.
- the processing component 102 may include one or more processors 118 to execute instructions to complete all or part of the steps of the method described below.
- the processing component 102 may include one or more modules for facilitating interaction between the processing component 102 and other components.
- the processing component 102 may include a multimedia module to facilitate the interaction between the multimedia component 108 and the processing component 102.
- the memory 104 is configured to store various types of data to support operation at the device 100. Examples of such data include instructions for any application program or method for operating on the device 100.
- the memory 104 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (Static Random Access Memory) Access Memory (SRAM for short), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (Programmable Red-Only Memory (referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, magnetic disk or optical disk.
- the memory 104 also stores one or more modules for the one or more modules configured to be executed by the one or more processors 118 to complete all or part of the steps in the method shown below.
- the power supply assembly 106 provides power to various components of the device 100.
- the power component 106 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 100.
- the multimedia component 108 includes a screen that provides an output interface between the device 100 and a user.
- the screen may include a liquid crystal display (Liquid Crystal Display (abbreviated as LCD) and touch panel. If the screen includes a touch panel, the screen may be implemented as a touch screen to receive an input signal from a user.
- the touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel.
- the touch sensor may not only sense a boundary of a touch or slide action, but also detect duration and pressure related to the touch or slide operation.
- the screen may also include an organic electroluminescent display (Organic Light Emitting Display (OLED for short).
- the audio component 110 is configured to output and / or input audio signals.
- the audio component 110 includes a microphone (Microphone, MIC for short).
- the microphone is configured to receive an external audio signal.
- the received audio signals may be further stored in the memory 104 or transmitted via the communication component 116.
- the audio component 110 further includes a speaker for outputting audio signals.
- the sensor component 114 includes one or more sensors for providing status assessment of various aspects of the device 100.
- the sensor component 114 can detect the open / closed state of the device 100, the relative positioning of the components, and the sensor component 114 can also detect a change in the position of the device 100 or a component of the device 100 and a change in the temperature of the device 100.
- the sensor component 114 may further include a magnetic sensor, a pressure sensor, or a temperature sensor.
- the communication component 116 is configured to facilitate wired or wireless communication between the device 100 and other devices.
- the device 100 can access a wireless network based on a communication standard, such as WiFi (Wireless-Fidelity, wireless fidelity).
- the communication component 116 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
- the communication component 116 further includes a Near Field Communication (NFC) module for facilitating short-range communication.
- NFC Near Field Communication
- the NFC module can be based on radio frequency identification (Radio Frequency Identification (RFID for short) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth technology and other technologies.
- RFID Radio Frequency Identification
- IrDA Infrared Data Association
- UWB Ultra Wideband
- the device 100 may be implemented by one or more application-specific integrated circuits (Application Specific Integrated Circuit (ASIC for short), digital signal processors, digital signal processing equipment, programmable logic devices, field programmable gate arrays, controllers, microcontrollers, microprocessors, or other electronic components to implement the following method.
- ASIC Application Specific Integrated Circuit
- Fig. 2 is a flow chart showing a method for processing clustering of traffic data according to an exemplary embodiment. As shown in Figure 2, this method includes the following steps.
- Step 210 Select N features from a preset feature database, where N is a positive integer.
- each user's traffic data is specified with several characteristics in advance. For example, these characteristics may include path repetition, the proportion of front-end and back-end login buried points, the number of IP access accounts, the number of IP accesses, and the mobile phone in the period. Number of user logins such as mean and variance.
- the preset database includes but is not limited to the above-mentioned features. N features are selected from the features included in the preset feature database, where N may be less than or equal to the preset feature database. A positive integer for the number of all features. The features may be selected by the user, randomly selected, or other selection methods, which are not limited in the embodiments of the present application.
- selecting the N features in the preset feature library may include: selecting the first N features with a chi-square value from high to low in the preset feature database.
- the preset feature library contains 14 features. At this time, a total of features are selected. If the number of clusters is between 2 and 20, there are 19 values. There are several combinations of features and class numbers. If each combination is traversed, the amount of calculation is very large.
- the target feature can be selected according to the size of the chi-square value corresponding to each feature. For example, if N is 1, the feature with the highest chi-square value in the preset feature library is selected as the target feature.
- N 2
- the feature with the highest and second-highest chi-square value in the preset feature database is selected as the target feature.
- Step 220 Obtain a feature vector of the traffic data based on the feature values corresponding to the selected features of the traffic data.
- the feature vector includes feature values corresponding to the N features of the traffic data, and one feature corresponds to one feature value.
- a1, a2, ..., an are the eigenvalues of the first, second, ..., N features, respectively.
- the obtained feature vector of the traffic data is a set of (a1, a2, ..., an).
- the traffic data includes white data and black data.
- the white data is traffic data extracted from the data traffic of the user identified as the white user
- the black data is extracted from the data traffic of the user identified as the black user.
- white users are users who are determined not to send abnormal traffic data
- black users are users who are determined to send abnormal traffic data.
- the ratio of the white data to the black data in the traffic data is a preset ratio and the preset ratio is 1: 1.
- the preset ratio may also be another ratio, which is not limited in the embodiment of the present application.
- Step 230 Cluster all the traffic data into M clusters according to the feature vector of the traffic data.
- M is a positive integer greater than or equal to 2.
- Step 240 Determine the sum of the cluster error numbers of the clusters into which the traffic data is divided under various combinations of M and N values.
- the sum of the error numbers of the clusters is a result of adding the error numbers of each of the clusters, and the error number of each cluster refers to a smaller number of white data and black data in the cluster. . Specifically, if the cluster includes only white data or only black data, it is considered that the clustering effect is the best at this time.
- the clusters in which the number of white data is greater than the number of black data in the M clusters are determined as white clusters
- the clusters in which the number of black data is greater than the number of white data in the M clusters are determined as black clusters
- the number of cluster errors of the white cluster Is the number of black data in the white cluster
- the number of cluster errors in the black cluster is the number of white flow data in the black cluster
- the sum of the cluster error number of the M clusters is the sum of the cluster error number of all the white clusters and the error number of all the black clusters The sum of the total number of cluster errors.
- various combinations of M and N values are combinations of traversing all values of the N value range and all values of the M value range.
- Step 250 Use the number of features and the number of clusters corresponding to the sum of the smallest number of cluster errors as the target feature number and the number of target clusters selected when the traffic data is clustered.
- the formula of the cluster risk score is as follows: ,among them, Represent the number of white samples and black samples in the cluster, and score is the cluster risk score. Among them, the number of white samples is the number of white data in the cluster, and the number of black samples is the number of black data in the cluster.
- the value range of the cluster risk score is [0,1].
- the cluster risk score (the closer the cluster risk score is to 1), the larger the proportion of black samples in the cluster, and the greater the risk of abnormal flow in the cluster.
- the cluster number of the cluster is correspondingly stored with the cluster risk score corresponding to the cluster, and the manager can view the cluster risk score of each cluster, thereby making the presentation of the cluster risk situation more intuitive.
- the cluster risk score is greater than 0.5
- the cluster is determined to be a cluster with abnormal traffic.
- the cluster may also be determined to be a cluster with abnormal traffic when the cluster risk score is greater than 0.6 or 0.7.
- the specific cluster risk score is greater than A certain value is not limited in the embodiments of the present application.
- step 250 the following steps may be performed: determining whether the number of clusters aggregated is greater than a preset number; when it is determined that the number of clusters is greater than the preset number, determining each cluster The center point of each cluster; according to the center point of each cluster, all clusters are divided into preset clusters, where the preset clusters include black clusters, white clusters, and mixed clusters, and black clusters are black Data-dominated clusters, white clusters are clusters dominated by white data, and mixed clusters are clusters where neither black data nor white data dominate.
- the clusters clustered can be further divided into three clusters, which is beneficial to subsequent analysis of the behavior of the black industry chain based on the clusters obtained by the division.
- the identification of traffic anomalies is not determined in isolation, but the traffic data is divided into several clusters based on the number of target features and the number of target clusters. Combining several clusters can reflect the traffic data in a group or an area. Or, the characteristics of a group of people are conducive to analyzing the behavior of the black industry chain. In summary, the demand for the overall analysis of the traffic data of the group is satisfied.
- FIG. 3 is a flowchart of details of step 230 shown in the corresponding embodiment of FIG. 2. As shown in FIG. 3, step 230 includes:
- Step 231 Normalize each feature value included in the feature vector of the traffic data to obtain a normalized feature vector.
- the normalization process is a result of dividing a feature value of a feature included in a feature vector of traffic data by a maximum feature value of the feature included in a feature vector of all traffic data.
- Step 232 cluster the normalized feature vectors into M clusters.
- Fig. 4 is a flow chart showing a method for processing clustering of traffic data according to another exemplary embodiment. As shown in Figure 4, this method includes the following steps.
- Step 401 Select N features from a preset feature database, where N is a positive integer.
- selecting the N features in the preset feature library may include: selecting the first N features with a chi-square value from high to low in the preset feature database.
- Step 402 Obtain a feature vector of the traffic data based on the feature values corresponding to the selected features of the traffic data.
- the feature vector includes feature values corresponding to the N features of the traffic data, and one feature corresponds to one feature value.
- the traffic data includes white data and black data. The white data is extracted from the data traffic of the user determined as the white user, and the black data is extracted from the data traffic of the user determined as the black user. The user is a user determined to not send abnormal traffic data, and the black user is a user determined to send abnormal traffic data.
- the ratio of the white data to the black data in the traffic data is a preset ratio and the preset ratio is 1: 1.
- Step 403 Cluster a part of the traffic data into M clusters according to the feature vector of the traffic data, where M is a positive integer greater than or equal to 2.
- Step 404 Determine the sum of the cluster error numbers of the clusters into which the traffic data is divided under various combinations of M and N values, and the sum of the cluster error numbers is the result of adding the error numbers of each of the clusters.
- the number of errors in each cluster refers to a smaller one of the number of white data and the number of black data in the cluster.
- step 405 the combination of M and N corresponding to the sum of the number of cluster errors in a predetermined order from small to large is used as a combination of candidate feature numbers M and N.
- Step 406 Cluster all the traffic data into M clusters according to the feature vector of the traffic data.
- M is a positive integer greater than or equal to 2.
- Step 407 Determine the total number of cluster errors of the clusters into which the traffic data is divided under various combinations of candidate M and N values.
- the sum of the error numbers of the clusters is a result of adding the error numbers of each of the clusters, and the error number of each cluster refers to a smaller number of white data and black data in the cluster .
- Step 408 Use the number of features and the number of clusters corresponding to the sum of the minimum number of cluster errors as the target feature number and the number of target clusters selected when the traffic data is clustered.
- the formula of the cluster risk score is as follows: ,among them, Represent the number of white samples and black samples in the cluster, and score is the cluster risk score.
- the first clustering process of this process clusters part of the traffic data to obtain a better combination of M and N candidates, and the second clustering process selects the better and better candidate according to the first clustering.
- the combination of M and N values can be used to cluster all traffic data, which can take into account both processing efficiency and accuracy of clustering.
- Fig. 5 is a block diagram of a clustering processing apparatus 500 for traffic data according to an exemplary embodiment. As shown in Figure 5, the device includes:
- a selecting unit 501 is configured to select N features from a preset feature database, where N is a positive integer. As an optional implementation manner, the selection unit 501 selects N features from a preset feature database, where N is a positive integer, and the selection unit 501 may select the first N chi-square values from high to low in the preset feature database. feature.
- the obtaining unit 502 is configured to obtain the feature vector of the traffic data based on the feature values corresponding to the selected features of the traffic data; the feature vector includes the feature values corresponding to the N features of the traffic data; and one feature corresponds to one feature value.
- the ratio of the white data to the black data in the traffic data is a preset ratio, and the preset ratio may be 1: 1.
- the first clustering unit 503 is configured to cluster all the traffic data into M clusters according to a feature vector of the traffic data, where M is a positive integer greater than or equal to 2.
- the first determining unit 504 is configured to determine the total number of cluster errors of the cluster into which the traffic data is divided under various combinations of M and N values.
- the total number of cluster errors is a result of adding the error numbers of each cluster into which each The number of errors of a cluster refers to the smaller of the number of white data and the number of black data in the cluster.
- the first setting unit 505 is configured to use the number of features and the number of clusters corresponding to the sum of the minimum number of cluster errors as the target feature number and the number of target clusters selected when clustering the traffic data.
- the first clustering unit 503 may include a normalization unit 5031 configured to perform normalization processing on each feature value included in the feature vector of the traffic data to obtain a normalization.
- the feature vector where the normalization process is the result of dividing the feature value of a feature included in the feature vector of the traffic data by the maximum feature value of the feature included in the feature vectors of all traffic data; the clustering subunit 5032, It is configured to cluster the normalized feature vector into M clusters.
- the device may further include a second clustering unit 506 configured to cluster all the traffic data in the first clustering unit 503.
- a part of the traffic data is clustered into M clusters, where M is a positive integer greater than or equal to 2; and a second determining unit 507 is configured to: The sum of the cluster error numbers of the clusters into which the part of the traffic data is divided under various combinations of M and N values, and the sum of the cluster error numbers is the result of the sum of the error numbers of each cluster divided.
- the number of errors refers to the smaller one of the number of white data and the number of black data in the cluster; the second setting unit 508 is configured to: set the predetermined order from small to large determined by the second determining unit 507
- the combination of M and N corresponding to the sum of the number of cluster errors is the combination of the candidate feature numbers M and N, wherein the first determining unit 504 is configured to implement the determination in various M and
- the flow data is divided into a combination of N values.
- Sum of cluster error numbers of clusters Determines the sum of cluster error numbers of clusters into which the traffic data is divided under various combinations of candidate M and N values.
- the device may further include a risk determination unit 509, which is configured to: in the first setting unit 505, use the number of features and the number of clusters corresponding to the minimum sum of the number of cluster errors as traffic After the number of target features and the number of target clusters selected during data clustering, determine the cluster risk score of each cluster after clustering according to the selected target feature number and the number of target clusters, and the formula for the cluster risk score as follows: ,among them, Represent the number of white samples and black samples in the cluster, and score is the cluster risk score.
- a risk determination unit 509 is configured to: in the first setting unit 505, use the number of features and the number of clusters corresponding to the minimum sum of the number of cluster errors as traffic After the number of target features and the number of target clusters selected during data clustering, determine the cluster risk score of each cluster after clustering according to the selected target feature number and the number of target clusters, and the formula for the cluster risk score as follows: ,among them, Represent the number of white samples and black samples in the cluster
- the first setting unit 505 the number of features and the number of clusters corresponding to the sum of the minimum number of cluster errors are used as the target feature number and the target cluster number selected when the traffic data is clustered.
- the first setting unit 505 may be further configured to: determine whether the number of clusters clustered is greater than a preset number; when it is determined that the number of clusters is greater than the preset number, determine a center point of each cluster cluster; The center point of the clustered cluster divides all clusters into preset clusters.
- the preset clusters include black clusters, white clusters, and mixed clusters. Black clusters are clusters dominated by black data and white clusters are white clusters.
- clusters clustered can be further divided into three clusters, which is conducive to subsequent analysis of the behavior of the black industry chain based on the clusters obtained by the clustering.
- each unit / module in the above device and related details are described in detail in the implementation process of the corresponding steps in the above method embodiment, and are not repeated here.
- the device embodiments in the above embodiments may be implemented by means of hardware, software, firmware, or a combination thereof, and they may be implemented as a single device, or each component unit / module may be dispersed in one or more A logic integrated system in each computing device and each performing a corresponding function.
- the units / modules constituting the device in the above embodiments are divided according to logical functions, and they may be re-divided according to logical functions.
- the device may be implemented by more or fewer units / modules.
- constituent units / modules may be implemented by means of hardware, software, firmware, or a combination thereof. They may be separate independent components or integrated units / modules in which multiple components are combined to perform corresponding logic functions.
- the manner of the hardware, software, firmware, or a combination thereof may include: separated hardware components, a functional module implemented by a programming manner, a functional module implemented by a programmable logic device, or the like, or a combination of the foregoing manners.
- the present application also provides an electronic device, the electronic device comprising: a memory having computer-readable instructions stored thereon; and a processor configured to execute the computer-readable instructions stored on the memory, the computer-readable instructions When executed by the processor, the computer is caused to execute the clustering processing method of the traffic data as described above.
- the processor described in the above embodiments may refer to a single processing unit, such as a central processing unit CPU, or a distributed processor system including a plurality of distributed processing units.
- the memory described in the above embodiments may include one or more memories, which may be internal memory of the computing device, such as various kinds of memories, whether transient or non-transitory, or external to the computing device through a memory interface. Storage device.
- the electronic device may be the traffic data clustering processing apparatus 100 shown in FIG. 1, and its structure and composition may refer to FIG. 1 and the description of FIG. 1 above.
- the present application further provides a computer-readable storage medium on which a computer program is stored.
- the processor causes the processor to perform clustering of traffic data as described above. Approach.
- the storage medium may be any tangible device that can hold and store instructions that can be used by the instruction execution device.
- it may be-but not limited to-an electric storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
- storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory) , Static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding device, such as a punch card with instructions stored on it, or A raised structure in the groove, and any suitable combination of the above.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- SRAM Static random access memory
- CD-ROM compact disc read-only memory
- DVD digital versatile disc
- memory stick floppy disk
- mechanical encoding device such as a punch card with instructions stored on it, or A raised structure in the groove, and any suitable combination of the above.
- the computer programs / computer instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.
- the network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
- the network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device .
- the computer program instructions described in this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or any arbitrary one or more programming languages.
- the programming languages include object-oriented programming languages—such as Smalltalk, C ++, and the like—and conventional procedural programming languages—such as the "C" language or similar programming languages.
- Computer-readable program instructions may be executed entirely on a user's computer, partly on a user's computer, as a stand-alone software package, partly on a user's computer, partly on a remote computer, or entirely on a remote computer or server carried out.
- the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (such as through the Internet using an Internet service provider) connection).
- the electronic circuit is personalized by utilizing the state information of the computer-readable program instructions, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA).
- the electronic circuit may Computer-readable program instructions are executed to implement various aspects of the invention.
- These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing device, thereby producing a machine such that when executed by a processor of a computer or other programmable data processing device , Means for implementing the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams.
- These computer-readable program instructions may also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Therefore, a computer-readable medium storing instructions includes: An article of manufacture that includes instructions to implement various aspects of the functions / acts specified in one or more blocks in the flowcharts and / or block diagrams.
- Computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other device, so that a series of operating steps can be performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process , So that the instructions executed on the computer, other programmable data processing apparatus, or other equipment can implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
Landscapes
- Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Biology (AREA)
- Evolutionary Computation (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种流量数据的聚类处理方法、装置及电子设备,该方法包括:在预置特征库中选取N个特征(201);基于流量数据的所选取的特征对应的特征值,得到流量数据的特征向量(202);根据流量数据的特征向量,将所有流量数据聚类成M个簇(203);确定在各种M和N取值的组合下流量数据分成的簇的簇错误数总和(204),簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数(205),从而利用聚类算法对大量流量数据进行聚类处理,得以满足对于群体的流量数据整体分析的需求。
Description
相关申请的交叉引用:本申请基于并要求2018年9月27日递交、发明名称为“一种流量数据的聚类处理方法、装置及电子设备”的中国专利申请CN
201811128269.2的优先权,在此通过引用将其全部内容合并于此。
本申请涉及大数据技术领域,特别涉及流量数据的聚类处理方法、装置及电子设备。
目前,随着互联网用户的日益增多,互联网领域正面临着大流量数据的挑战。大流量数据中难免会出现异常流量,这些异常流量会给互联网带来巨大的冲击与损失,例如,黑色产业形成的木马播种、流量交易和虚拟财产套现等诸多黑色产业链都会产生大量的异常流量。在现有技术的实现中,流量异常的识别一般是通过采集用户行为埋点和sdk数据来确定路径重复度、设备前后端登录埋点占比、ip访问账号数、ip访问次数、周期内手机号段用户登录均值和方差等特征,根据每一条流量数据的这些特征,确定该流量数据异常的概率。
本申请的发明人意识到,现有技术的缺陷在于,黑色产业往往表现为群体的流量数据出现异常,而现有技术对于流量异常的识别是针对每一条流量数据孤立确定的,无法满足对于群体的流量数据整体分析的需求。
为了解决相关技术中存在的无法满足对于群体的流量数据整体分析的需求的问题,本申请提供了一种流量数据的聚类处理方法、装置及电子设备。
根据本申请实施例的一方面,提供一种流量数据的聚类处理方法,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述方法包括:在预置特征库中选取N个特征,N为正整数;基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量;所述特征向量包括所述流量数据的所述N个特征各自对应的特征值,其中,一个所述特征对应一个所述特征值;根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
根据本申请实施例的另一方面,提供一种流量数据的聚类处理装置,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述装置包括:选取单元,用于在预置特征库中选取N个特征,N为正整数;获取单元,用于基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量;所述特征向量包括所述流量数据的所述N个特征各自对应的特征值;其中,一个所述特征对应一个所述特征值;聚类单元,用于根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定单元,用于确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;设置单元,用于将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
根据本申请实施例的又一方面,提供一种计算机可读存储介质,其特征在于,其存储计算机程序,所述计算机程序在被计算机执行时使得计算机执行如前所述的方法。
根据本申请实施例的又一方面,提供一种电子设备,包括:存储器,所述存储器上存储有计算机可读指令;处理器,其被配置为执行所述存储器上存储的计算机可读指令,其中所述计算机可读指令被所述处理器执行时,使得所述处理器执行如前所述的方法。
本申请的实施例提供的技术方案可以包括以下有益效果:
本申请所提供的图像控制方法包括如下步骤,在预置特征库中选取N个特征,N为正整数;基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量;所述特征向量包括所述流量数据的所述N个特征各自对应的特征值;其中,一个所述特征对应一个所述特征值;根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
此方法下,对于流量异常的识别不是孤立确定的,而是将流量数据划分为若干个簇,结合若干个簇能够反映流量数据在一个群体、或在一个区域、或在一类人中呈现出的特点,有利于分析黑产业链的行为。综上,对于群体的流量数据整体分析的需求得以满足。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性的,并不能限制本申请。
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本申请的实施例,并与说明书一起用于解释本申请的原理。
图1是根据一示例性实施例示出的一种流量数据的聚类处理装置的示意图;
图2是根据一示例性实施例示出的一种流量数据的聚类处理方法的流程图;
图3是根据图2对应实施例示出的步骤230的细节的流程图;
图4是根据另一示例性实施例示出的一种流量数据的聚类处理方法的流程图;
图5是根据一示例性实施例示出的一种流量数据的聚类处理装置的框图。
这里将详细地对示例性实施例执行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。以下示例性实施例中所描述的实施方式并不代表与本申请相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本申请的一些方面相一致的装置和方法的例子。
本申请的实施环境可以是便携移动设备,如智能手机、平板电脑、台式电脑。本申请所公开的流量数据的聚类处理方法可以适用于运行于便携移动设备上的任意应用程序。
图1是根据一示例性实施例示出的一种流量数据的聚类处理装置的示意图。装置100可以是上述便携移动设备。如图1所示,装置100可以包括以下一个或多个组件:处理组件102,存储器104,电源组件106,多媒体组件108,音频组件110,传感器组件114以及通信组件116。处理组件102通常控制装置100的整体操作,诸如与显示,电话呼叫,数据通信,相机操作以及记录操作相关联的操作等。处理组件102可以包括一个或多个处理器118来执行指令,以完成下述的方法的全部或部分步骤。此外,处理组件102可以包括一个或多个模块,用于便于处理组件102和其他组件之间的交互。例如,处理组件102可以包括多媒体模块,用于以方便多媒体组件108和处理组件102之间的交互。存储器104被配置为存储各种类型的数据以支持在装置100的操作。这些数据的示例包括用于在装置100上操作的任何应用程序或方法的指令。存储器104可以由任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(Static Random
Access Memory,简称SRAM),电可擦除可编程只读存储器(Electrically
Erasable Programmable Read-Only Memory,简称EEPROM),可擦除可编程只读存储器(Erasable Programmable Read Only Memory,简称EPROM),可编程只读存储器(Programmable
Red-Only Memory,简称PROM),只读存储器(Read-Only Memory,简称ROM),磁存储器,快闪存储器,磁盘或光盘。存储器104中还存储有一个或多个模块,用于该一个或多个模块被配置成由该一个或多个处理器118执行,以完成如下所示方法中的全部或者部分步骤。电源组件106为装置100的各种组件提供电力。电源组件106可以包括电源管理系统,一个或多个电源,及其他与为装置100生成、管理和分配电力相关联的组件。多媒体组件108包括在所述装置100和用户之间的提供一个输出接口的屏幕。在一些实施例中,屏幕可以包括液晶显示器(Liquid Crystal
Display,简称LCD)和触摸面板。如果屏幕包括触摸面板,屏幕可以被实现为触摸屏,以接收来自用户的输入信号。触摸面板包括一个或多个触摸传感器以感测触摸、滑动和触摸面板上的手势。所述触摸传感器可以不仅感测触摸或滑动动作的边界,而且还检测与所述触摸或滑动操作相关的持续时间和压力。屏幕还可以包括有机电致发光显示器(Organic
Light Emitting Display,简称OLED)。音频组件110被配置为输出和/或输入音频信号。例如,音频组件110包括一个麦克风(Microphone,简称MIC),当装置100处于操作模式,如呼叫模式、记录模式和语音识别模式时,麦克风被配置为接收外部音频信号。所接收的音频信号可以被进一步存储在存储器104或经由通信组件116发送。在一些实施例中,音频组件110还包括一个扬声器,用于输出音频信号。传感器组件114包括一个或多个传感器,用于为装置100提供各个方面的状态评估。例如,传感器组件114可以检测到装置100的打开/关闭状态,组件的相对定位,传感器组件114还可以检测装置100或装置100一个组件的位置改变以及装置100的温度变化。在一些实施例中,该传感器组件114还可以包括磁传感器,压力传感器或温度传感器。通信组件116被配置为便于装置100和其他设备之间有线或无线方式的通信。装置100可以接入基于通信标准的无线网络,如WiFi(Wireless-Fidelity,无线保真)。在一个示例性实施例中,通信组件116经由广播信道接收来自外部广播管理系统的广播信号或广播相关信息。在一个示例性实施例中,所述通信组件116还包括近场通信(Near Field Communication,简称NFC)模块,用于以促进短程通信。例如,在NFC模块可基于射频识别(Radio
Frequency Identification,简称RFID)技术,红外数据协会(Infrared Data Association,简称IrDA)技术,超宽带(Ultra Wideband,简称UWB)技术,蓝牙技术和其他技术来实现。
在示例性实施例中,装置100可以被一个或多个应用专用集成电路(Application
Specific Integrated Circuit,简称ASIC)、数字信号处理器、数字信号处理设备、可编程逻辑器件、现场可编程门阵列、控制器、微控制器、微处理器或其他电子元件实现,用于执行下述方法。
图2是根据一示例性实施例示出的一种流量数据的聚类处理方法的流程图。如图2所示,此方法包括以下步骤。
步骤210,在预置特征库中选取N个特征,N为正整数。
本申请实施例中,每个用户的流量数据事先被规定若干个特征,例如,这些特征可以包括路径重复度、设备前后端登录埋点占比、ip访问账号数、ip访问次数、周期内手机号段用户登录均值和方差等,预置数据库中包括但不限于上述若干个特征,从预置特征库包括的若干个特征中选取N个特征,其中,N可以为小于等于预置特征库中所有特征的数量的正整数。其中,特征可以由用户指定选取,也可以随机选取,也可以采用其他选取方式,本申请实施例中不做限定。
作为一种可选实施方式,在预置特征库中选取N个特征可包括:在预置特征库中选取卡方值从高到低前N个特征。本申请实施例中,假设预置特征库中包含14个特征,这时选取特征共有 种情况,若聚成簇的类数在2-20之间取值,有19种取值,因此选取的特征和类数的组合有 种。如果每种组合都去遍历,计算量非常大。此时可以按照每一特征对应的卡方值大小来选取目标特征。例如,如果N为1,则选取预置特征库中卡方值最高的特征作为目标特征,如果N为2,则选取预置特征库中卡方值最高和次高的特征作为目标特征,由于卡方值越大,该卡方值对应的特征对于良好的聚类越重要,因此可以选取聚类效果最好的特征,提升聚类效果。
步骤220,基于流量数据的所选取的特征对应的特征值,得到流量数据的特征向量。
本申请实施例中,特征向量包括流量数据的N个特征各自对应的特征值;其中,一个特征对应一个特征值。例如,a1,a2,……,an分别是第1,2,……,N个特征的特征值,得到的流量数据的特征向量即为(a1,a2,……,an)构成的集合。本申请实施例中,流量数据包括白数据和黑数据,白数据是从确定为白用户的用户的数据流量中抽取的流量数据,黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,白用户是确定为不会发出异常流量数据的用户,黑用户是确定为会发出异常流量数据的用户。可选的,流量数据中白数据和黑数据的比为预设比例且预设比例为1:1,预设比例也可以为其他比例,本申请实施例中不做限定。通过实施这种可选的实施方式,减少了因白数据与黑数据选取比例失衡导致局部最优的情况出现的概率。
步骤230,根据流量数据的特征向量,将所有流量数据聚类成M个簇。
本申请实施例中,M为大于等于2的正整数。
步骤240,确定在各种M和N取值的组合下流量数据分成的簇的簇错误数总和。
本申请实施例中,簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量。具体的,簇中如果只包括白数据或者只包括黑数据,认为此时聚类效果最佳。因此,将M个簇中白数据的数量大于黑数据的数量的簇确定为白簇,将M个簇中黑数据的数量大于白数据的数量的簇确定为黑簇,白簇的簇错误数为白簇中黑数据的数量,黑簇的簇错误数为黑簇中白流量数据的数量,M个簇的簇错误数总和即为所有白簇的簇错误数总和与所有黑簇的错误数总和累加得到的簇错误数总和。并且,各种M和N取值的组合为遍历N值取值范围的所有值与M值取值范围的所有值的组合。
步骤250,将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
作为一种可选的实施方式,在将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数之后,还可以执行以下步骤:确定在按照选取的目标特征数和目标簇个数聚类之后,聚成的每个簇的簇风险评分,该簇风险评分的公式如下:
,其中,
分别表示该簇中白样本个数和黑样本个数,score为簇风险评分。其中,白样本个数即为该簇中白数据个数,黑样本个数即为该簇中黑数据个数。簇风险评分的取值范围为[0,1],簇风险评分越大(簇风险评分越靠近1),表示该簇黑样本比例越大,该簇存在流量异常的风险也就越大。且将该簇的簇编号与该簇对应的簇风险评分对应存储,管理人员可以查看每一簇的簇风险评分,从而使得簇风险情况的呈现更加直观。可选的,当簇风险评分大于0.5时,确定该簇为流量异常的簇,其中,也可以在簇风险评分大于0.6或者0.7时确定该簇为流量异常的簇,具体的簇风险评分大于的某一数值本申请实施例中不做限定。
作为另一种可选的实施方式,在执行完步骤250之后,还可以执行以下步骤:判断聚成的簇的数量是否大于预设数量;当判断出大于预设数量时,确定出每一聚成的簇的中心点;根据每一聚成的簇的中心点,将所有聚成的簇划分至预设簇中,其中,预设簇包括黑簇、白簇以及混合簇,黑簇为黑数据占主导的簇,白簇为白数据占主导的簇,混合簇为黑数据与白数据均不做主导的簇。通过实施这种可选实施方式,当聚成的簇的数量过多时,可进一步将聚成的簇划分得到三个簇,有利于后续根据划分得到的簇分析黑产业链的行为。
上述方法下,对于流量异常的识别不是孤立确定的,而是依据目标特征数和目标簇个数将流量数据划分为若干个簇,结合若干个簇能够反映流量数据在一个群体、或在一个区域、或在一类人中呈现出的特点,有利于分析黑产业链的行为。综上,对于群体的流量数据整体分析的需求得以满足。
图3是图2对应实施例示出的步骤230的细节的流程图。如图3所示,步骤230包括:
步骤231,对流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量。本申请实施例中,归一化处理是用流量数据的特征向量所包括的一个特征的特征值除以所有流量数据的特征向量所包括的该特征的最大特征值的结果。
步骤232,将归一化特征向量聚类成M个簇。
图4是根据另一示例性实施例示出的一种流量数据的聚类处理方法的流程图。如图4所示,此方法包括以下步骤。
步骤401,在预置特征库中选取N个特征,N为正整数。作为一种可选实施方式,在预置特征库中选取N个特征可包括:在预置特征库中选取卡方值从高到低前N个特征。
步骤402,基于流量数据的所选取的特征对应的特征值,得到流量数据的特征向量。本申请实施例中,特征向量包括流量数据的N个特征各自对应的特征值;其中一个特征对应一个特征值。本申请实施例中,流量数据包括白数据和黑数据,白数据是从确定为白用户的用户的数据流量中抽取的,黑数据是从确定为黑用户的用户的数据流量中抽取的,白用户是确定为不会发出异常流量数据的用户,黑用户是确定为会发出异常流量数据的用户。可选的,流量数据中白数据和黑数据的比为预设比例且预设比例为1:1。
步骤403,根据流量数据的特征向量,将一部分流量数据聚类成M个簇,M为大于等于2的正整数。
步骤404,确定在各种M和N取值的组合下流量数据分成的簇的簇错误数总和,簇错误数总和是分成的每个簇的错误数相加的结果。本申请实施例中,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量。
步骤405,将从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合。
步骤406,根据流量数据的特征向量,将所有流量数据聚类成M个簇。本申请实施例中,M为大于等于2的正整数。
步骤407,确定在各种候选M和N取值的组合下流量数据分成的簇的簇错误数总和。
本申请实施例中,簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量。
步骤408,将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
作为一种可选的实施方式,在将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数之后,还可以执行以下步骤:确定在按照选取的目标特征数和目标簇个数聚类之后,聚成的每个簇的簇风险评分,该簇风险评分的公式如下:
,其中,
分别表示该簇中白样本个数和黑样本个数,score为簇风险评分。上述方法下,能够在对预设数量的流量数据聚类成初始簇时,从中选取较优的候选M和N取值的组合,并在该选取较优的候选M和N取值的组合下针对所有流量数据进行聚类,从中选取簇错误数总和取值最小的簇错误数总和。这一过程的第一次聚类过程对部分流量数据聚类获取较优的候选M和N取值的组合,第二次聚类过程按照第一次聚类选取的较优的较优的候选M和N取值的组合,对全部流量数据聚类,可以同时兼顾处理效率与聚类的准确性。以下是本申请的装置实施例。
图5是根据一示例性实施例示出的一种流量数据的聚类处理装置500的框图。如图5所示,该装置包括:
选取单元501,用于在预置特征库中选取N个特征,N为正整数。作为一种可选的实施方式,选取单元501在预置特征库中选取N个特征,N为正整数可以包括:选取单元501在预置特征库中选取卡方值从高到低前N个特征。获取单元502,用于基于流量数据的所选取的特征对应的特征值,得到流量数据的特征向量;特征向量包括流量数据的N个特征各自对应的特征值;其中,一个特征对应一个特征值。本实施例中,流量数据中白数据和黑数据的比为预设比例,预设比例可以为1:1。第一聚类单元503,用于根据流量数据的特征向量,将所有流量数据聚类成M个簇,M为大于等于2的正整数。第一确定单元504,用于确定在各种M和N取值的组合下流量数据分成的簇的簇错误数总和,簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量。第一设置单元505,用于将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
作为一种可选的实施方式,第一聚类单元503可以包括:归一化单元5031,其被配置为:对流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量,其中归一化处理是用流量数据的特征向量所包括的一个特征的特征值除以所有流量数据的特征向量所包括的该特征的最大特征值的结果;聚类子单元5032,其被配置为:将归一化特征向量聚类成M个簇。
作为一种可选的实施方式,如图5所示,该装置还可以包括:第二聚类单元506,其被配置为:在所述第一聚类单元503将所有所述流量数据聚类成M个簇之前,根据所述流量数据的特征向量,将一部分所述流量数据聚类成M个簇,M为大于等于2的正整数;第二确定单元507,其被配置为:确定在各种M和N取值的组合下所述一部分所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;第二设置单元508,其被配置为:将所述第二确定单元507确定的从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合,其中,所述第一确定单元504被配置为通过执行如下步骤来实现所述确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和:确定在各种候选M和N取值的组合下所述流量数据分成的簇的簇错误数总和。作为一种可选的实施方式,该装置还可以包括风险确定单元509,其被配置为:在第一设置单元505将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数之后,确定在按照选取的目标特征数和目标簇个数聚类之后,聚成的每个簇的簇风险评分,该簇风险评分的公式如下:
,其中,
分别表示该簇中白样本个数和黑样本个数,score为簇风险评分。
作为另一种可选的实施方式,在第一设置单元505将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数之后,第一设置单元505还可以被配置为:判断聚成的簇的数量是否大于预设数量;当判断出大于预设数量时,确定出每一聚成的簇的中心点;根据每一聚成的簇的中心点,将所有聚成的簇划分至预设簇中,其中,预设簇包括黑簇、白簇以及混合簇,黑簇为黑数据占主导的簇,白簇为白数据占主导的簇,混合簇为黑数据与白数据均不做主导的簇。通过实施这种可选的实施方式,当聚成的簇的数量过多时,可以进一步将聚成的簇划分得到三个簇,有利于后续根据划分得到的簇分析黑产业链的行为。
上述装置中各个单元/模块的功能和作用的实现过程以及相关细节具体详见上述方法实施例中对应步骤的实现过程,在此不再赘述。以上各实施例中的装置实施例可以通过硬件、软件、固件或其组合的方式来实现,并且其可以被实现为一个单独的装置,也可以被实现为各组成单元/模块分散在一个或多个计算设备中并分别执行相应功能的逻辑集成系统。以上各实施例中组成装置的各单元/模块是根据逻辑功能而划分的,它们可以根据逻辑功能被重新划分,例如可以通过更多或更少的单元/模块来实现该装置。这些组成单元/模块分别可以通过硬件、软件、固件或其组合的方式来实现,它们可以是分别的独立部件,也可以是多个组件组合起来执行相应的逻辑功能的集成单元/模块。所述硬件、软件、固件或其组合的方式可以包括:分离的硬件组件,通过编程方式实现的功能模块、通过可编程逻辑器件实现的功能模块,等等,或者以上方式的组合。
本申请还提供一种电子设备,该电子设备包括:存储器,该存储器上存储有计算机可读指令;处理器,其被配置为执行所述存储器上存储的计算机可读指令,该计算机可读指令被处理器执行时,使得计算机执行如前所述的流量数据的聚类处理方法。上面的实施例中所述的处理器可以指单个的处理单元,如中央处理单元CPU,也可以是包括多个分散的处理单元的分布式处理器系统。上面的实施例中所述的存储器可以包括一个或多个存储器,其可以是计算设备的内部存储器,例如暂态或非暂态的各种存储器,也可以是通过存储器接口连接到计算设备的外部存储装置。
该电子设备可以是图1所示的流量数据聚类处理装置100,其结构组成可以参考图1及上面对图1的描述。
在一示例性实施例中,本申请还提供一种计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时,使得处理器执行如前所述的流量数据的聚类处理方法。
该存储介质可以是任何可以保持和存储可由指令执行设备使用的指令的有形设备。例如,其可以是――但不限于――电存储设备、磁存储设备、光存储设备、电磁存储设备、半导体存储设备或者上述的任意合适的组合。存储介质的更具体的例子(非穷举的列表)包括:便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、静态随机存取存储器(SRAM)、便携式压缩盘只读存储器(CD-ROM)、数字多功能盘(DVD)、记忆棒、软盘、机械编码设备、例如其上存储有指令的打孔卡或凹槽内凸起结构、以及上述的任意合适的组合。
这里所描述的计算机程序/计算机指令可以从计算机可读存储介质下载到各个计算/处理设备,或者通过网络、例如因特网、局域网、广域网和/或无线网下载到外部计算机或外部存储设备。网络可以包括铜传输电缆、光纤传输、无线传输、路由器、防火墙、交换机、网关计算机和/或边缘服务器。每个计算/处理设备中的网络适配卡或者网络接口从网络接收计算机可读程序指令,并转发该计算机可读程序指令,以供存储在各个计算/处理设备中的计算机可读存储介质中。
本公开中所述的计算机程序指令可以是汇编指令、指令集架构(ISA)指令、机器指令、机器相关指令、微代码、固件指令、状态设置数据、或者以一种或多种编程语言的任意组合编写的源代码或目标代码,所述编程语言包括面向对象的编程语言—诸如Smalltalk、C++等,以及常规的过程式编程语言—诸如“C”语言或类似的编程语言。计算机可读程序指令可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络—包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。在一些实施例中,通过利用计算机可读程序指令的状态信息来个性化定制电子电路,例如可编程逻辑电路、现场可编程门阵列(FPGA)或可编程逻辑阵列(PLA),该电子电路可以执行计算机可读程序指令,从而实现本发明的各个方面。
这里参照根据本发明实施例的方法、装置(系统)和计算机程序产品的流程图和/或框图描述了本发明的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其它可编程数据处理装置的处理器,从而生产出一种机器,使得这些指令在通过计算机或其它可编程数据处理装置的处理器执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
也可以把计算机可读程序指令加载到计算机、其它可编程数据处理装置、或其它设备上,使得在计算机、其它可编程数据处理装置或其它设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其它可编程数据处理装置、或其它设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
应当理解的是,本申请并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围执行各种修改和改变。本申请的范围仅由所附的权利要求来限制。
Claims (28)
- 一种流量数据的聚类处理方法,其特征在于,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述方法包括:在预置特征库中选取N个特征,N为正整数;基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量;所述特征向量包括所述流量数据的所述N个特征各自对应的特征值,其中,一个所述特征对应一个所述特征值;根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
- 根据权利要求1所述的方法,其特征在于,所述在预置特征库中选取N个特征包括:在预置特征库中选取卡方值从高到低前N个特征。
- 如权利要求1所述的方法,其特征在于,流量数据中白数据和黑数据的比为预设比例。
- 根据权利要求3所述的方法,其特征在于,所述预设比例为1:1。
- 根据权利要求1所述的方法,其特征在于,所述根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,包括:对所述流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量,其中归一化处理是用所述流量数据的特征向量所包括的一个特征的特征值除以所有所述流量数据的特征向量所包括的该特征的最大特征值的结果;将所述归一化特征向量聚类成M个簇。
- 根据权利要求1所述的方法,其特征在于,在根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇之前,所述方法还包括:根据流量数据的特征向量将一部分流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合,且所述确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和包括:确定在各种候选M和N取值的组合下所述流量数据分成的簇的簇错误数总和。
- 一种流量数据的聚类处理装置,其特征在于,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述装置包括:选取单元,其被配置为:在预置特征库中选取N个特征,N为正整数;获取单元,其被配置为:基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量,所述特征向量包括所述流量数据的所述N个特征各自对应的特征值;其中,一个所述特征对应一个所述特征值;第一聚类单元,其被配置为:根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;第一确定单元,其被配置为:确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;第一设置单元,其被配置为:将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
- 根据权利要求8所述的装置,其特征在于,所述选取单元被配置为:在预置特征库中选取卡方值从高到低的前N个特征。
- 如权利要求8所述的装置,其特征在于,流量数据中白数据和黑数据的比为预设比例。
- 根据权利要求10所述的装置,其特征在于,所述预设比例为1:1。
- 根据权利要求8所述的装置,其特征在于,所述第一聚类单元包括:归一化单元,其被配置为:对所述流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量,其中归一化处理是用所述流量数据的特征向量所包括的一个特征的特征值除以所有所述流量数据的特征向量所包括的该特征的最大特征值的结果;聚类子单元,其被配置为:将所述归一化特征向量聚类成M个簇。
- 根据权利要求8所述的装置,其特征在于,还包括:第二聚类单元,被配置为:在第一聚类单元将所有所述流量数据聚类成M个簇之前,根据流量数据的特征向量将一部分所述流量数据聚类成M个簇,M为大于等于2的正整数;第二确定单元,其被配置为:确定在各种M和N取值的组合下所述一部分所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;第二设置单元,其被配置为:将所述第二确定单元确定的从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合,其中,所述第一确定单元被配置为通过执行如下步骤来实现所述确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和:确定在各种候选M和N取值的组合下所述流量数据分成的簇的簇错误数总和。
- 一种电子设备,其特征在于,所述电子设备包括:存储器,所述存储器上存储有计算机可读指令;处理器,其被配置为执行所述存储器上存储的计算机可读指令,其中,所述处理器通过执行所述计算机可读指令被配置为执行一种流量数据的聚类处理方法,其中,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述流量数据的聚类处理方法包括:在预置特征库中选取N个特征,N为正整数;基于流量数据的所选取的特征对应的特征值得到所述流量数据的特征向量;所述特征向量包括流量数据的所述N个特征各自对应的特征值,其中一个所述特征对应一个所述特征值;根据流量数据的特征向量将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
- 根据权利要求15所述的电子设备,其特征在于,所述在预置特征库中选取N个特征包括:在预置特征库中选取卡方值从高到低前N个特征。
- 根据权利要求15所述的电子设备,其特征在于,流量数据中白数据和黑数据的比为预设比例。
- 根据权利要求17所述的电子设备,其特征在于,所述预设比例为1:1。
- 根据权利要求15所述的电子设备,其特征在于,所述处理器通过执行所述计算机可读指令被配置为通过执行如下步骤来实现所述根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇:对所述流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量,其中归一化处理是用所述流量数据的特征向量所包括的一个特征的特征值除以所有所述流量数据的特征向量所包括的该特征的最大特征值的结果;将所述归一化特征向量聚类成M个簇。
- 根据权利要求15所述的电子设备,其特征在于,所述处理器通过执行所述计算机可读指令被配置为:在根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇之前,执行如下步骤:根据流量数据的特征向量将一部分流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合,其中,所述处理器通过执行所述计算机可读指令被配置为通过执行如下步骤来实现所述确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和:确定在各种候选M和N取值的组合下所述流量数据分成的簇的簇错误数总和。
- 一种计算机可读存储介质,其特征在于,其存储计算机程序,所述计算机程序在被计算机执行时使得所述计算机被配置为执行一种流量数据的聚类处理方法,其中,所述流量数据包括白数据和黑数据,所述白数据是从确定为白用户的用户的数据流量中抽取的流量数据,所述黑数据是从确定为黑用户的用户的数据流量中抽取的流量数据,所述白用户是确定为不会发出异常流量数据的用户,所述黑用户是确定为会发出异常流量数据的用户,所述流量数据的聚类处理方法包括:在预置特征库中选取N个特征,N为正整数;基于流量数据的所选取的特征对应的特征值,得到所述流量数据的特征向量;所述特征向量包括所述流量数据的所述N个特征各自对应的特征值,其中,一个所述特征对应一个所述特征值;根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将最小的簇错误数总和所对应的特征个数和簇个数,作为流量数据聚类时选取的目标特征数和目标簇个数。
- 根据权利要求22所述的计算机可读存储介质,其特征在于,所述在预置特征库中选取N个特征包括:在预置特征库中选取卡方值从高到低前N个特征。
- 根据权利要求22所述的计算机可读存储介质,其特征在于,流量数据中白数据和黑数据的比为预设比例。
- 根据权利要求24所述的计算机可读存储介质,其特征在于,所述预设比例为1:1。
- 根据权利要求22所述的计算机可读存储介质,其特征在于,所述计算机程序在被计算机执行时使得所述计算机被配置为通过执行如下步骤来实现所述根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇:对所述流量数据的特征向量所包括的各特征值进行归一化处理,得到归一化特征向量,其中归一化处理是用所述流量数据的特征向量所包括的一个特征的特征值除以所有所述流量数据的特征向量所包括的该特征的最大特征值的结果;将所述归一化特征向量聚类成M个簇。
- 根据权利要求22所述的计算机可读存储介质,其特征在于,所述计算机程序在被计算机执行时使得所述计算机被配置为:在根据所述流量数据的特征向量,将所有所述流量数据聚类成M个簇之前,执行如下步骤:根据所述流量数据的特征向量,将一部分所述流量数据聚类成M个簇,M为大于等于2的正整数;确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和,所述簇错误数总和是分成的每个簇的错误数相加的结果,每个簇的错误数是指该簇中白数据的数量和黑数据的数量中较少的一个数量;将从小到大前预定名次的簇错误数总和所对应的M和N的组合,作为候选特征数M和N的组合,其中,所述处理器通过执行所述计算机可读指令被配置为通过执行如下步骤来实现所述确定在各种M和N取值的组合下所述流量数据分成的簇的簇错误数总和:确定在各种候选M和N取值的组合下所述流量数据分成的簇的簇错误数总和。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811128269.2A CN109284307B (zh) | 2018-09-27 | 2018-09-27 | 一种流量数据的聚类处理方法、装置及电子设备 |
| CN201811128269.2 | 2018-09-27 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020062689A1 true WO2020062689A1 (zh) | 2020-04-02 |
Family
ID=65181859
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/125246 Ceased WO2020062689A1 (zh) | 2018-09-27 | 2018-12-29 | 流量数据的聚类处理方法、装置及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109284307B (zh) |
| WO (1) | WO2020062689A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210035025A1 (en) * | 2019-07-29 | 2021-02-04 | Oracle International Corporation | Systems and methods for optimizing machine learning models by summarizing list characteristics based on multi-dimensional feature vectors |
| CN119596713A (zh) * | 2025-02-10 | 2025-03-11 | 佛山市南海英吉威铝建材有限公司 | 一种仿大理石铝单板冲压设备的运行状态监控方法 |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110209260B (zh) * | 2019-04-26 | 2024-02-23 | 平安科技(深圳)有限公司 | 耗电量异常检测方法、装置、设备及计算机可读存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103001825A (zh) * | 2012-11-15 | 2013-03-27 | 中国科学院计算机网络信息中心 | Dns流量异常的检测方法和系统 |
| CN105141604A (zh) * | 2015-08-19 | 2015-12-09 | 国家电网公司 | 一种基于可信业务流的网络安全威胁检测方法及系统 |
| US20170134401A1 (en) * | 2015-11-05 | 2017-05-11 | Radware, Ltd. | System and method for detecting abnormal traffic behavior using infinite decaying clusters |
| CN107592323A (zh) * | 2017-11-02 | 2018-01-16 | 江苏物联网研究发展中心 | 一种DDoS检测方法及检测装置 |
-
2018
- 2018-09-27 CN CN201811128269.2A patent/CN109284307B/zh active Active
- 2018-12-29 WO PCT/CN2018/125246 patent/WO2020062689A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103001825A (zh) * | 2012-11-15 | 2013-03-27 | 中国科学院计算机网络信息中心 | Dns流量异常的检测方法和系统 |
| CN105141604A (zh) * | 2015-08-19 | 2015-12-09 | 国家电网公司 | 一种基于可信业务流的网络安全威胁检测方法及系统 |
| US20170134401A1 (en) * | 2015-11-05 | 2017-05-11 | Radware, Ltd. | System and method for detecting abnormal traffic behavior using infinite decaying clusters |
| CN107592323A (zh) * | 2017-11-02 | 2018-01-16 | 江苏物联网研究发展中心 | 一种DDoS检测方法及检测装置 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210035025A1 (en) * | 2019-07-29 | 2021-02-04 | Oracle International Corporation | Systems and methods for optimizing machine learning models by summarizing list characteristics based on multi-dimensional feature vectors |
| US12393860B2 (en) * | 2019-07-29 | 2025-08-19 | Oracle International Corporation | Systems and methods for optimizing machine learning models by summarizing list characteristics based on multi-dimensional feature vectors |
| CN119596713A (zh) * | 2025-02-10 | 2025-03-11 | 佛山市南海英吉威铝建材有限公司 | 一种仿大理石铝单板冲压设备的运行状态监控方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109284307A (zh) | 2019-01-29 |
| CN109284307B (zh) | 2021-06-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10733538B2 (en) | Techniques for querying a hierarchical model to identify a class from multiple classes | |
| US11062698B2 (en) | Image-based approaches to identifying the source of audio data | |
| EP3635578B1 (en) | Systems and methods for crowdsourced actions and commands | |
| US20210081720A1 (en) | Techniques for the automated customization and deployment of a machine learning application | |
| US20180276553A1 (en) | System for querying models | |
| US11790375B2 (en) | Flexible capacity in an electronic environment | |
| US11165779B2 (en) | Generating a custom blacklist for a listening device based on usage | |
| US20200273453A1 (en) | Topic based summarizer for meetings and presentations using hierarchical agglomerative clustering | |
| US20150213127A1 (en) | Method for providing search result and electronic device using the same | |
| US20200272693A1 (en) | Topic based summarizer for meetings and presentations using hierarchical agglomerative clustering | |
| US9747175B2 (en) | System for aggregation and transformation of real-time data | |
| US20180366113A1 (en) | Robust replay of digital assistant operations | |
| CN110780955B (zh) | 一种用于处理表情消息的方法与设备 | |
| US12307389B2 (en) | Predicting events based on time series data | |
| US20190138511A1 (en) | Systems and methods for real-time data processing analytics engine with artificial intelligence for content characterization | |
| WO2020062689A1 (zh) | 流量数据的聚类处理方法、装置及电子设备 | |
| US11914667B2 (en) | Managing multi-dimensional array of data definitions | |
| US20240236091A1 (en) | Systems and methods for recognizing new devices | |
| US10909206B2 (en) | Rendering visualizations using parallel data retrieval | |
| US11068558B2 (en) | Managing data for rendering visualizations | |
| US20190220901A1 (en) | Mobile terminal and method of managing application thereof, and system for providing target advertisement using the same | |
| US10409916B2 (en) | Natural language processing system | |
| WO2025055714A1 (zh) | 行人重识别方法、装置、电子设备及存储介质 | |
| US20220276067A1 (en) | Method and apparatus for guiding voice-packet recording function, device and computer storage medium | |
| CN115910062A (zh) | 音频识别方法、装置、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18935895 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18935895 Country of ref document: EP Kind code of ref document: A1 |


