WO2018014717A1 - 聚类方法、装置及电子设备 - Google Patents

聚类方法、装置及电子设备 Download PDF

Info

Publication number
WO2018014717A1
WO2018014717A1 PCT/CN2017/091432 CN2017091432W WO2018014717A1 WO 2018014717 A1 WO2018014717 A1 WO 2018014717A1 CN 2017091432 W CN2017091432 W CN 2017091432W WO 2018014717 A1 WO2018014717 A1 WO 2018014717A1
Authority
WO
WIPO (PCT)
Prior art keywords
cluster
clusters
sample data
similarity
initial
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2017/091432
Other languages
English (en)
French (fr)
Inventor
潘薪宇
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to US15/859,345 priority Critical patent/US11080306B2/en
Publication of WO2018014717A1 publication Critical patent/WO2018014717A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/28Databases characterised by their database models, e.g. relational or object models
    • G06F16/284Relational databases
    • G06F16/285Clustering or classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/23Clustering techniques
    • G06F18/232Non-hierarchical techniques
    • G06F18/2321Non-hierarchical techniques using statistics or function optimisation, e.g. modelling of probability density functions
    • G06F18/23211Non-hierarchical techniques using statistics or function optimisation, e.g. modelling of probability density functions with adaptive number of clusters
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/22Matching criteria, e.g. proximity measures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F7/00Methods or arrangements for processing data by operating upon the order or content of the data handled
    • G06F7/06Arrangements for sorting, selecting, merging, or comparing data on individual record carriers
    • G06F7/08Sorting, i.e. grouping record carriers in numerical or other ordered sequence according to the classification of at least some of the information they carry
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F7/00Methods or arrangements for processing data by operating upon the order or content of the data handled
    • G06F7/06Arrangements for sorting, selecting, merging, or comparing data on individual record carriers
    • G06F7/14Merging, i.e. combining at least two sets of record carriers each arranged in the same ordered sequence to produce a single set having the same ordered sequence

Definitions

  • the present invention relates to the field of data processing technologies, and in particular, to a clustering method, device, and electronic device.
  • the present disclosure discloses a clustering technical solution.
  • the present disclosure discloses a clustering method, the method comprising: acquiring inter-sample similarity between every two sample data in M sample data, M is a positive integer; and the M is obtained according to the obtained inter-sample similarity
  • the sample data is merged into N initial clusters, and N is a positive integer smaller than M.
  • the N initial clusters are clustered and merged to obtain a plurality of clusters corresponding to the M sample data.
  • the combining the M sample data into the N initial cluster clusters according to the obtained inter-sample similarity includes: For the two sample data, if the similarity between the samples of the two sample data is greater than the first threshold, the two sample data are combined into one suspected initial cluster cluster; for the plurality of suspected initial cluster clusters obtained, At least two suspected initial cluster clusters containing the same number of sample data not less than the second threshold are merged into one initial cluster cluster.
  • the combining the M sample data into the N initial cluster clusters according to the obtained inter-sample similarity includes: For the two sample data, if the similarity between the samples of the two sample data is greater than the first threshold, the two sample data are combined into one suspected initial cluster cluster; the number of data containing the same sample is not less than the first a first suspected initial clustering cluster and a second suspected initial clustering cluster of two thresholds, calculating a modulus of the intersection of the first and second suspected initial clustering clusters, and calculating a modulus of the first initial clustering cluster and a second initializing cluster Merging the sum of the modules of the clusters with the first constant, and combining the first and second suspected initial clusters into a suspected initialization when the difference between the modulus of the intersection and the quotient is not less than the second constant Clustering clusters.
  • the combining the M sample data into the N initial clusters according to the acquired inter-sample similarity further includes: initializing multiple suspected pairs Cluster clusters, each suspected initial cluster cluster that cannot be merged with other suspected initial cluster clusters is used as an initial cluster cluster.
  • the second threshold is linear according to the number of sample data included in one or more suspected initial cluster clusters according to the at least two suspected initial cluster clusters. The function is determined.
  • the method further includes: performing outlier separation on at least one of the plurality of cluster clusters, and obtaining the outliers after the separation process All cluster clusters are used as a plurality of cluster clusters corresponding to the M sample data.
  • the outlier separation includes: obtaining, for any cluster cluster, a to-be-away cluster and a non-out-of-group corresponding to each sample data in the cluster cluster.
  • a cluster wherein the sample data is included in the cluster to be separated from each sample data, and the non-out-of-group cluster includes: other sample data in the cluster cluster except the sample data;
  • the performing clustering and merging the N initial cluster clusters includes: using the N initial cluster clusters as multiple clusters to be clustered; Whether the degree of similarity between the clusters to be clustered and the other clusters to be clustered is determined; whether the maximum inter-cluster similarity among all the cluster similarities corresponding to the plurality of clusters to be clustered is greater than a fourth threshold; The maximum inter-cluster similarity is greater than the fourth threshold, and the two clusters to be clustered corresponding to the maximum inter-cluster similarity are combined to obtain a new cluster to be clustered; and the new cluster to be clustered The new plurality of clusters to be clustered, which are composed of other clusters to be clustered, are not clustered until there is no cluster to be clustered.
  • the M sample data is M images
  • the inter-sample similarity between the two images includes: two feature vectors corresponding to the two images respectively Cosine distance between.
  • the present disclosure also provides a clustering device, wherein the device comprises:
  • an acquiring unit configured to acquire the inter-sample similarity between each of the M sample data, where M is a positive integer
  • a merging unit configured to merge the M sample data according to the obtained inter-sample similarity to N initial cluster clusters, N is a positive integer smaller than M
  • a clustering unit is configured to cluster and merge the N initial cluster clusters to obtain a plurality of cluster clusters corresponding to the M sample data.
  • the merging unit includes: a first merging subunit, in the M sample data, if any two sample data, if the two sample data If the similarity between samples is greater than the first threshold, the two sample data are merged into a suspected initial cluster; the second merged subunit is used to initialize the cluster cluster for the plurality of suspected clusters, and the same sample data is included. At least two suspected initial cluster clusters whose number is not less than the second threshold are merged into one initial cluster cluster.
  • the merging unit is specifically configured to: in the M sample data, for any two sample data, if the similarity between samples of the two sample data is greater than a first threshold, the two sample data are merged into a suspected initial cluster cluster; and the first suspected initial cluster cluster and the second suspected initial cluster cluster containing the same sample data are not less than the second threshold, Calculating a modulus of the intersection of the first and second suspected initial cluster clusters, and calculating a quotient of a sum of a modulus of the first initial cluster cluster and a second initial cluster cluster and a first constant, the modulus at the intersection When the difference from the quotient is not less than the second constant, the first and second suspected initial cluster clusters are merged into a suspected initial cluster cluster.
  • the merging unit further includes: first, as a sub-unit, for initializing the cluster cluster for the plurality of suspected ones, and not combining with other suspected initial clusters Each suspected initial cluster cluster is used as an initial cluster cluster.
  • the second threshold is linear according to the number of sample data included in one or more suspected initial cluster clusters according to the at least two suspected initial cluster clusters. The function is determined.
  • the device further includes:
  • a separating unit configured to perform outlier separation on at least one of the plurality of cluster clusters, and all cluster clusters obtained after the outlier separation process are corresponding to the M sample data Multiple clusters of clusters.
  • the separating unit is specifically configured to perform outlier separation on a cluster of clusters; the separating unit includes:
  • the sample data is included in the cluster, and the non-out-of-group cluster includes: other sample data in the cluster cluster except the sample data;
  • a first acquiring subunit configured to acquire an inter-cluster similarity between the to-be-away cluster and the non-out-of-group cluster corresponding to each sample data
  • a first determining subunit configured to determine whether a minimum inter-cluster similarity among the plurality of inter-cluster similarities corresponding to all the sample data in the cluster cluster is less than a third threshold
  • a first response subunit configured to respond to the minimum inter-cluster similarity to be smaller than the third threshold, and to treat the minimum inter-cluster similarity as the two groups to be separated and the non-outgoing cluster respectively New clustering clusters.
  • the clustering unit includes:
  • a second as a subunit configured to use the N initial cluster clusters as a plurality of clusters to be clustered
  • a second obtaining subunit configured to acquire the inter-cluster similarity between each cluster to be clustered and other clusters to be clustered
  • a second determining subunit configured to determine whether a maximum inter-cluster similarity among all inter-cluster similarities corresponding to the plurality of clusters to be clustered is greater than a fourth threshold
  • a second response subunit configured to combine the two clusters to be clustered corresponding to the maximum inter-cluster similarity to obtain a new cluster to be clustered, in response to the maximum inter-cluster similarity being greater than the fourth threshold And triggering the second acquiring sub-unit, and continuing to cluster and merge the new cluster to be clustered with the other clusters to be clustered that are not merged this time until there is no Merged clusters to be clustered.
  • the M sample data is M images
  • the inter-sample similarity between the two images includes: two feature vectors corresponding to the two images respectively Cosine distance between.
  • the present disclosure also discloses an electronic device comprising: a housing, a processor, a memory, a circuit board, and a power supply circuit, wherein the circuit board is disposed inside a space enclosed by the housing, the processor and the The memory is disposed on the circuit board; the power circuit is configured to supply power to each circuit or device of the terminal; the memory is configured to store executable program code; and the processor reads the stored in the memory by reading The executable program code runs a program corresponding to the executable program code for performing the operation corresponding to the clustering method described above.
  • the present disclosure also discloses a non-transitory computer storage medium that can store computer readable instructions that, when executed, cause a processor to perform operations corresponding to the clustering methods described above.
  • the sample data when clustering the sample data, the sample data is first combined according to the similarity between samples of each of the plurality of sample data to obtain an initial cluster, and the cluster is initialized when the cluster is reduced.
  • the number of clusters is clustered and merged according to the initial cluster clusters at this time, and multiple cluster clusters corresponding to the plurality of sample data are obtained, which is beneficial to improving the clustering speed.
  • FIG. 1 is a schematic diagram of an application scenario of the present disclosure
  • FIG. 2 is a schematic diagram of another application scenario of the present disclosure.
  • FIG. 3 shows a block diagram of an exemplary device that implements the present disclosure
  • FIG. 4 shows a block diagram of another exemplary device that implements the present disclosure
  • FIG. 5 is a schematic flowchart diagram of a clustering method provided by the present disclosure.
  • FIG. 6 is a schematic diagram of inter-cluster similarity between initial clusters provided by the present disclosure.
  • FIG. 7 is a schematic flowchart diagram of an initial clustering cluster clustering merging method provided by the present disclosure.
  • FIG. 8 is a schematic flowchart diagram of a method for determining the similarity between clusters provided by the present disclosure
  • FIG. 9 is a schematic structural diagram of a clustering apparatus provided by the present disclosure.
  • FIG. 10 is a schematic structural diagram of a clustering apparatus provided by the present disclosure.
  • FIG. 11 is a schematic structural diagram of an electronic device provided by the present disclosure.
  • FIG. 1 schematically illustrates an application scenario in which a clustering technical solution provided in accordance with the present disclosure may be implemented.
  • a plurality of pictures are stored in the user's smart mobile phone 1000.
  • the pictures stored in the user's smart mobile phone 1000 include: a first picture 1001, a second picture 1002, a third picture 1003, ... and the A total of M pictures such as M picture 1004, wherein the size of M can be determined by the number of pictures stored in smart mobile phone 1000, and in general, M is a positive integer not less than 2.
  • the picture stored in the smart mobile phone 1000 may be a picture taken by the user using the smart mobile phone 1000, or may be a picture in which the user transmits and stores the picture stored in the other terminal device in the smart mobile phone 1000, or may be used by the user.
  • the smart mobile phone 1000 downloads pictures and the like from the network.
  • the technical solution provided by the present disclosure may enable the first picture 1001, the second picture 1002, the third picture 1003, ..., and the Mth picture 1004 stored in the smart mobile phone 1000 to be automatically divided into multiple sets according to a predetermined classification manner.
  • the predetermined classification manner is based on the person in the picture
  • the technical solution provided by the present disclosure can automatically make the first picture 1001 and the M picture 1004 stored in the smart mobile phone 1000 automatically. Dividing into the picture set of the first user, and automatically dividing the second picture 1002 and the third picture 1003 stored in the smart mobile phone 1000 into the picture set of the Nth user,
  • the above N is a positive integer smaller than M.
  • the present disclosure is useful for improving the manageability of a picture by automatically dividing the first picture 1001, the second picture 1002, the third picture 1003, ..., and the M picture 1004 stored in the smart mobile phone 1000 into a plurality of sets. Thereby, it is possible to bring convenience to the user, for example, the user can conveniently browse all the pictures stored in the smart mobile phone 1000 containing the portrait of the first user.
  • FIG. 2 schematically illustrates another application scenario in which a clustering technical scheme provided in accordance with the present disclosure may be implemented.
  • the device 2000 (such as a computer or a server) of the e-commerce platform stores all the product information or part of the product information that can be provided by the e-commerce platform.
  • the product information stored in the device 2000 of the e-commerce platform includes: A total of M pieces of product information, such as the first product information 2001, the second product information 2002, the third product information 2003, ..., and the Mth product information 2004, wherein the size of M can be determined by the number of product information stored in the device 2000.
  • M is a positive integer not less than 2.
  • the item information stored in the device 2000 may include: a picture of the item, a text description of the item, and the like.
  • the technical solution provided by the present disclosure can automatically divide the first product information 2001, the second product information 2002, the third product information 2003, ..., and the Mth product information 2004 stored in the device 2000 into a plurality of predetermined classification manners.
  • the predetermined classification method is divided according to the product category
  • the product corresponding to the first product information 2001 stored in the device 2000 and the product corresponding to the M-th product information 2004 belong to the first category (
  • the products corresponding to the second product information 2002 and the products corresponding to the third product information 2003 stored in the device 2000 belong to the Nth category (such as the infant formula)
  • the present disclosure provides
  • the technical solution can automatically divide the first product information 2001 and the Mth product information 2004 stored in the device 2000 into a first category (such as a ladies travel shoe category) set, and enable the second product information 2002 and the first stored in the device 2000.
  • the three commodity information 2003 is automatically divided into a set of the Nth category (such as a baby milk powder category), and the above N
  • the present disclosure helps to improve the manageability of product information by dividing the first product information 2001, the second product information 2002, the third product information 2003, ..., and the Mth product information 2004 stored in the device 2000 into a plurality of sets.
  • sexuality so that it can bring operational convenience to the e-commerce platform and the consumer.
  • the e-commerce platform can locate the first item information 2001.
  • Other product information eg, Mth item information
  • Mth item information in the first category set is recommended for display to the user.
  • FIG. 3 illustrates a block diagram of an exemplary device 3000 (eg, a computer system/server) suitable for implementing the present disclosure.
  • the device 3000 shown in FIG. 3 is merely an example and should not impose any limitation on the function and scope of use of the present disclosure.
  • device 3000 can be embodied in the form of a general purpose computing device.
  • Components of device 3000 may include, but are not limited to, one or more processors or processing units 3001, system memory 3002, and bus 3003 that connect different system components, including system memory 3002 and processing unit 3001.
  • Device 3000 can include a variety of computer system readable media. These media can be any available media that can be accessed by device 3000, including volatile and non-volatile media, removable and non-removable media, and the like.
  • System memory 3002 can include computer system readable media in the form of volatile memory, such as random access storage (RAM) 3021 and/or cache memory 3022.
  • Device 3000 may further include other removable/non-removable, volatile/non-volatile computer system storage media.
  • ROM 3023 can be used to read and write non-removable, non-volatile magnetic media (not shown in Figure 3, commonly referred to as "hard disk drives").
  • a disk drive for reading and writing to a removable non-volatile disk (eg, a "floppy disk"), and a removable non-volatile disk (eg, CD-ROM, DVD-) may be provided.
  • ROM or other optical media read and write optical drive.
  • each drive can be coupled to bus 300 via one or more data medium interfaces.
  • At least one program product may be included in system memory 3002, the program product having a set (eg, at least one) of program modules configured to perform the functions of the present disclosure.
  • Program module 3024 typically performs the functions and/or methods described in this disclosure.
  • Device 3000 can also be in communication with one or more external devices 3004 (eg, a keyboard, pointing device, display, etc.). Such communication may be through an input/output (I/O) interface 3005, and the device 3000 may also be through a network adapter 3006 with one or more networks (eg, a local area network (LAN), a wide area network (WAN), and/or a public network, For example, the Internet) communication. As shown in FIG. 3, network adapter 3006 communicates with other modules of device 3000 (e.g., processing unit 3001, etc.) via bus 3003. It should be understood that although not shown in FIG. 3, other hardware and/or software modules may be utilized in conjunction with device 3000.
  • I/O input/output
  • network adapter 3006 communicates with other modules of device 3000 (e.g., processing unit 3001, etc.) via bus 3003. It should be understood that although not shown in FIG. 3, other hardware and/or software modules may be utilized in conjunction with device 3000.
  • the processing unit 3001 executes various functional applications and data processing by executing a computer program stored in the system memory 3002, for example, executing instructions for implementing the steps in the above methods; specifically, processing The unit 3001 can execute a computer program stored in the system memory 3002, and when the computer program is executed, the following steps are implemented: acquiring inter-sample similarity between every two sample data in the M sample data, M being a positive integer And merging the M sample data into N initial cluster clusters according to the obtained similarity between samples, N is a positive integer smaller than M; clustering and merging the N initial cluster clusters to obtain the M Multiple clusters of clusters corresponding to sample data.
  • device 4000 includes one or more processors, communication portions, etc., which may be: one or more central processing units (CPUs) 4001, and/or one or more images
  • processors central processing units
  • GPU graphics processing unit
  • the processor may perform various appropriate operations according to executable instructions stored in a read only memory (ROM) 4002 or executable instructions loaded from the storage portion 4008 into the random access memory (RAM) 4003.
  • the communication unit 4120 may include, but is not limited to, a network card, which may include, but is not limited to, an IB (Infiniband) network card.
  • the processor can communicate with the read only memory 4002 and/or the random access memory 4300 to execute executable instructions, connect to the communication portion 4120 via the bus 4004, and communicate with other target devices via the communication portion 4120, thereby completing the corresponding in the present disclosure.
  • the operation performed by the processor includes: acquiring inter-sample similarity between each of the M sample data, M is a positive integer; merging the M sample data into N according to the obtained inter-sample similarity Initializing the cluster cluster, N is a positive integer smaller than M; clustering and merging the N initial cluster clusters to obtain a plurality of cluster clusters corresponding to the M sample data.
  • RAM 4003 various programs and data required for the operation of the device can also be stored.
  • the CPU 4001, the ROM 4002, and the RAM 4003 are connected to each other through a bus 4004.
  • ROM 4002 is an optional module.
  • the RAM 4003 stores executable instructions, or writes executable instructions to the ROM 4002 at runtime, the executable instructions cause the central processing unit 4001 The operation corresponding to the above communication method is performed.
  • An input/output (I/O) interface 4005 is also coupled to bus 4004.
  • the communication unit 4120 may be integrated, or may be configured to have a plurality of sub-modules (for example, a plurality of IB network cards) and be respectively connected to the bus.
  • the following components are connected to the I/O interface 4005: an input portion 4006 including a keyboard, a mouse, etc.; an output portion 4007 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, etc.; a storage portion 4008 including a hard disk or the like And a communication portion 4009 including a network interface card such as a LAN card, a modem, or the like.
  • the communication section 4009 performs communication processing via a network such as the Internet.
  • the driver 4100 is also connected to the I/O interface 4005 as needed.
  • a removable medium 4110 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like, is mounted on the drive 4100 as needed so that a computer program read therefrom is installed into the storage portion 4008 as needed.
  • FIG. 4 is only an optional implementation manner.
  • the number and type of components in the foregoing FIG. 4 may be selected, deleted, added, or replaced according to actual needs; Different function components can be set up, such as separate settings or integrated settings.
  • the GPU and the CPU can be separated, and the GPU can be integrated on the CPU.
  • the communication unit can be separated or integrated. CPU or GPU, etc.
  • embodiments of the present disclosure include a computer program product comprising tangibly embodied on a machine readable medium
  • Computer program comprising program code for performing the steps shown in the flowchart, the program code comprising executable instructions corresponding to the steps performed by the present application, for example, for acquiring each of the M sample data
  • An executable instruction for the similarity between samples between sample data, M is a positive integer
  • an executable instruction for combining the M sample data into N initial cluster clusters according to the obtained similarity between samples, N is a positive integer smaller than M
  • the computer program can be downloaded and installed from the network via the communication portion 4009, and/or installed from the removable medium 4110.
  • the computer program is executed by the central processing unit (CPU) 4001
  • the above-described executable instructions described in the present disclosure are executed.
  • the clustering technical solution provided by the present disclosure will be described below with reference to FIGS. 5-11. Any of the clustering technical solutions provided by the present disclosure may be exemplified by software or hardware or a combination of software and hardware.
  • the clustering technical solution provided by the present disclosure may be implemented by a certain electronic device or by a certain processor, and the disclosure is not limited, and the electronic device may include, but is not limited to, a terminal or a server, and the processor may include but Not limited to CPU or GPU. The details are not described below.
  • S101 acquires the inter-sample similarity between every two sample data in the M sample data, and M is a positive integer. Usually, M is greater than 2, such as M can be 100 or 1001 or 5107, and the like.
  • step S101 may be performed by a processor calling an instruction stored in a memory, or may be performed by an obtaining unit 501 executed by a processor.
  • the sample data in the present disclosure may be an image (ie, a picture), a voice, a video, or a text, or the like.
  • the processor may extract the feature of the sample data by using a convolutional neural network or other conventional local feature descriptors, for example, when the sample data is an image,
  • the processor may perform feature extraction on the image by using Sift (Scale-invariant feature transform), HOG (Histogram of Oriented Gradien), etc., to obtain feature vectors of each sample data;
  • the feature vector of the extracted sample data constructs a feature matrix. Assume that there are currently n sample data and extract each sample data.
  • the m-class feature constructs an n*m-order feature matrix based on the extracted features.
  • the inter-sample similarity between the two sample data in the present disclosure may be represented by, but not limited to, a cosine distance, an Euclidean distance, or the like using the sample data feature vector.
  • the processor After the processor establishes the feature matrix, the cosine distance of the feature between the sample data can be calculated to generate a similarity matrix, where all features of one sample data are used as one feature vector.
  • the processor generates an n*m-order feature matrix, according to which the processor can generate an n*n-order similarity matrix as follows:
  • M ij is the inter-sample similarity between the i-th sample data and the j-th sample data
  • X represents the feature vector of the i-th sample data
  • Y represents the feature vector of the j-th sample data
  • the similarity matrix characterizes the similarity measure between the sample data.
  • the similarity between the samples of the two sample data is very high (such as exceeding a certain threshold, or the similarity satisfies a certain predetermined condition, etc.), the approximation can be approximated.
  • the two samples are considered to belong to the same class.
  • the similarity matrix is a symmetric matrix.
  • the processor may obtain a similarity matrix of M sample data by receiving the input information, however, the input information in the present disclosure is M sample data.
  • the processor can obtain a similarity matrix of M sample data by using the scheme described above.
  • the present disclosure does not limit the specific implementation of the similarity matrix in which the processor obtains M sample data.
  • S102 Combine the M sample data into N initial cluster clusters according to the obtained inter-sample similarity, where N is a positive integer smaller than M.
  • step S102 may be performed by the processor invoking a corresponding instruction stored in the memory, or may be performed by a merging unit 502 that is executed by the processor.
  • each initialization cluster family contains at least one sample data.
  • the present disclosure generally requires determining an initial cluster of sample data prior to clustering the sample data.
  • the usual practice is to determine each sample data separately as a cluster cluster. These cluster clusters are used as initial clusters. For example, there are currently 4 sample data, which are a, b, c, and d, respectively. 4 initial clusters. This method, based on each sample data, has good robustness, and the final clustering results obtained have good accuracy.
  • the method for determining each sample data separately as an initial cluster cluster, and obtaining a final clustering result according to the determined initial cluster cluster has a computational complexity of O(n 3 log n), where n is sample data. Quantity, this method is suitable for scenarios where high precision is required and speed performance is not high. In the scenario where the speed performance of the processed data is high, the clustering technical solution of the present disclosure reduces the number of initial clusters by merging the sample data, which is beneficial to improve the speed performance of the cluster.
  • the processor combines the M sample data into N initial clusters according to the obtained inter-sample similarity, which may include the following steps:
  • step S01 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a first merging sub-unit executed by the processor.
  • the processor may combine a and b into a suspected initial cluster.
  • C 1 combining a and d into a suspected initial cluster cluster C 2 , and b and c into a suspected initial cluster cluster C 3 , since the similarity between the samples of e and all other sample data is not greater than the first Threshold, therefore, the processor can initialize e as a suspect cluster cluster C 4 .
  • the first threshold may be a relatively high threshold. For example, if the range of similarity between samples is [0, 1], the first threshold may be 0.9, 0.95, etc. Therefore, the accuracy of the suspected initial cluster cluster classification can be guaranteed as much as possible.
  • the first threshold may be adjusted according to actual conditions. For example, in a case where each suspected initial cluster cluster includes one sample data, the processor may appropriately lower the first threshold; for example, include two sample data. In the case where the ratio of the number of suspected initial cluster clusters to the number of all suspected initial cluster clusters is lower than a predetermined ratio (for example, 1/100), the processor may appropriately lower the first threshold.
  • S02. Combine, for the obtained plurality of suspected initial cluster clusters, at least two suspected initial cluster clusters including the number of identical sample data not less than the second threshold into an initial cluster cluster.
  • step S02 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a second merging sub-unit executed by the processor.
  • the foregoing second threshold may be set to a fixed positive integer. For example, if the second threshold is 1, if two suspected initial clusters each contain the same sample data, The processor merges the two suspected initial clusters; in addition, a suspected initial cluster that cannot be merged with other suspected initial clusters is also determined to be an initial cluster, for example, the first run by the processor As a subunit, a suspected initial cluster cluster that cannot be merged with other suspected initial cluster clusters is determined as an initial cluster cluster.
  • the similarity between samples is transitive, so the processor can merge the suspected initial clusters containing the same sample data into one initial cluster.
  • initialization suspected suspected clades C. 1 and 2 in the initialization clades C contains sample data a
  • C suspected initialization clades. 1 and 3 suspected initialization clades C contains sample data b
  • the processor may merge the suspected initial cluster cluster C 1 , the suspect initial cluster cluster C 2 and the suspect initial cluster cluster C 3 into one initial cluster cluster C 5 ; in addition, the suspect cluster cluster C 4 and other initialization clades suspected sample not contain the same data, the processor does not initialize the clades C 4 suspected merging process will initialize the pseudo clades C 4 determination initialization clades.
  • the computational complexity of the clustering technical scheme of the present disclosure is O(n 2 log n), and each sample data is separately determined as an initial cluster cluster, and the final clustering result is obtained according to the determined initial cluster cluster.
  • O(n 3 log n) of the technical solution the calculation speed of clustering the sample data to obtain the final clustering result is improved to a large extent.
  • One way in which the present disclosure incorporates a suspected initial clustering cluster is that the two suspected initial cluster clusters are merged as long as the number of identical data contained in the two suspected initial cluster clusters is not less than the second threshold.
  • the two suspected initial cluster clusters may not belong to the same category, but the similarity between the sample data between the two suspected initial cluster clusters is greater than the first threshold. That is, since it is suspected that the noise point sample data exists in the cluster cluster, two suspected initial cluster clusters are merged, so that an error may occur in determining the stage of initializing the cluster cluster.
  • X i is sample data, and there is a solid line connection between the sample data indicating that the similarity between the samples between the two sample data is greater than the first threshold, and the suspected initial cluster cluster C 1 and the suspected initial cluster are suspected. Only X 19 and X 12 are connected between cluster C 2 . If the above-mentioned suspected initial cluster clustering method is adopted, the suspected initial cluster cluster C 1 and the suspect initial cluster cluster C 2 are merged, which makes it impossible to merge two One suspected initial clustering cluster is merged into one, and an error occurs in determining the stage of initializing the cluster cluster.
  • the second threshold in the present disclosure may be based on sample data included in one or more suspected initial cluster clusters in the at least two suspected initial cluster clusters.
  • the linear function of the number is determined. For example, if the slope k is 0.5, when a suspected initial cluster cluster C 1 contains 2 sample data, if it wants to merge with the suspected initial cluster cluster C 2 , the second threshold is 1; when it is suspected to initialize the cluster cluster When C 1 contains 4 sample data, if it wants to merge with the suspected initial cluster cluster C 2 , the second threshold is 2, and so on.
  • Another alternative implementation manner in which the present disclosure incorporates a suspected initial clustering cluster is that not only the number of identical sample data included in the two suspected initial cluster clusters is not less than a second threshold, but also two suspected initial cluster clusters It should also satisfy the following formula (1) in order to be merged:
  • S 1 is a sample data set corresponding to a suspected initial cluster cluster.
  • the similarity between samples of one sample data and other sample data is greater than a first threshold
  • S 2 is another suspect initialization.
  • the sample data set corresponding to the cluster cluster, in the S 2 the sample similarity between one sample data and the other sample data is greater than the first threshold
  • ⁇ and ⁇ are preset parameters.
  • the process of merging two suspected initial cluster clusters can be specifically as follows:
  • the first suspected initial cluster cluster and the second suspected initial cluster cluster are refused to be merged into a new suspected initial cluster cluster.
  • all suspected initial clusters are determined to be initialized clusters.
  • step S01 since both the suspected initial cluster cluster C 1 and the suspected initial cluster cluster C 2 contain sample data a, the processor will suspect the initial cluster cluster C 1 and the suspect initialize cluster cluster C 2 It is regarded as a sample data set respectively. If the modulus of the intersection of C 1 and C 2 , the mode of C 1 and the mode of C 2 satisfy the formula (1), that is, the following formula holds:
  • the processor merges the suspected initial cluster cluster C 1 and the suspected initial cluster cluster C 2 into a new suspected initial cluster cluster C 5 , at which time the suspected initial cluster cluster C 5 contains sample data a sample data b And the sample data d, and the suspected initial cluster cluster C 3 contains the sample data b and the sample data c, and the suspected initial cluster cluster C 5 and the suspect initial cluster cluster C 3 both contain the sample data b, and the processor judges again It is suspected that the initial cluster cluster C 5 and the suspected initial cluster cluster C 3 satisfy the formula (1).
  • the processor refuses to merge the suspected initial cluster cluster C 5 and the suspect initial cluster cluster C 3 Since the suspected initial cluster cluster C 4 contains the sample data e, the suspected initial cluster cluster C 4 and the suspected initial cluster cluster C 5 and the suspected initial cluster cluster C 3 do not contain the same sample data, therefore, the processor respectively Determining the initial cluster cluster C 5 , the suspect initial cluster cluster C 3 , and the suspect initial cluster cluster C 4 as initial cluster clusters;
  • the processor refuses to merge the suspected initial cluster cluster C 1 and the suspected initial cluster cluster C 2 due to the suspect
  • the initial cluster cluster C 1 and the suspected initial cluster cluster C 3 both contain sample data b, and the processor again determines whether the suspected initial cluster cluster C 1 and the suspected initial cluster cluster C 3 satisfy the formula (1), and the assumption is still not satisfied.
  • the suspected initial cluster cluster C 4 contains sample data e, suspected initial cluster cluster C 4 and suspected initial cluster cluster C 1 , suspected initial cluster cluster C 2 and suspected initial cluster cluster C 3 The same sample data is not included, so the processor does not combine them, the processor will suspect the initial cluster cluster C 1 , the suspect cluster cluster C 2 , the suspect cluster cluster C 3 and the suspect cluster cluster C 4 is determined to initialize cluster clusters, respectively.
  • the disclosure discloses that the number of the same data included in the two suspected initial cluster clusters is not less than the second threshold, and in the case that the formula (1) is satisfied, the two suspected initial cluster clusters are merged, in two In the case of suspected initial clustering clusters, the case of (
  • S103 Perform clustering and combining the N initial cluster clusters to obtain a plurality of cluster clusters corresponding to the M sample data.
  • step S103 may be performed by a processor calling a corresponding instruction stored in the memory, or may be performed by a clustering unit 503 executed by the processor.
  • clustering and combining the N initial cluster clusters may include the following steps:
  • S301 The N initial cluster clusters are used as multiple clusters to be clustered.
  • step S301 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a second sub-unit executed by the processor.
  • S302 Acquire an inter-cluster similarity between each cluster to be clustered and other clusters to be clustered.
  • step S302 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a second acquisition sub-unit operated by the processor.
  • FIG. 8 is a method for determining the similarity between clusters provided by the present disclosure. A schematic diagram of the process, the method comprising the following steps:
  • the sample data contained in the first cluster to be clustered is: a, b; the sample data contained in the second cluster to be clustered is: c, d; then the determined similarity between the first samples includes: a and The similarity between samples between c, the similarity between samples between a and d, the similarity between samples between b and c, and the similarity between samples between b and d.
  • the minimum value of the similarity range is greater than the smallest inter-sample similarity among the first samples, and the maximum value of the similarity range is smaller than the largest inter-sample similarity among the first samples.
  • the sample data corresponding to the larger value and the smaller value in the first sample similarity is likely to be the noise point data in general, and therefore, when determining the similarity range, the first sample is taken.
  • the middle range in the similarity is the similarity range. According to the similarity between the first samples in the similarity range, the inter-cluster similarity between the two clusters to be clustered is calculated, and the noise point data pair can be removed to the greatest extent. The influence of the accuracy of the similarity between clusters between the two clusters is improved, and the accuracy of the final clustering result is improved.
  • the similarity between the first samples obtained by the processor is: 0.2, 0.32, 0.4, 0.3, 0.7, 0.5, 0.75, 0.8, 0.9, 0.92.
  • the similarity range determined by the processor may be: 0.3-0.75. It may also be: 0.4-0.7; of course, other similarity ranges satisfying the above conditions may also be used, and the disclosure does not limit this.
  • the processor may determine the similarity range by:
  • the processor sorts the obtained similarity between the first samples, for example, e1 ⁇ e2 ⁇ e3 ⁇ ... ⁇ eE, where E is the number of similarities between the obtained first samples, and e is the first Similarity between samples;
  • the processor determines the similarity range according to the parameters l and k.
  • the parameter l may take 0.2E
  • the parameter k may take 0.8E.
  • the processor sorts the similarity between the first samples obtained by the processor, 0.2 ⁇ 0.3 ⁇ 0.32 ⁇ 0.4 ⁇ 0.5 ⁇ 0.7 ⁇ 0.75 ⁇ 0.8 ⁇ 0.9 ⁇ 0.92, and the processor is based on the similarity between the first samples.
  • the processor may calculate the inter-cluster similarity between the first cluster to be clustered and the second cluster to be clustered according to the following formula:
  • the i-th first sample-to-sample similarity, l and k are the above-determined parameters.
  • S303 Determine whether a maximum inter-cluster similarity among all inter-cluster similarities corresponding to the plurality of clusters to be clustered is greater than a fourth threshold.
  • step S303 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a second determining sub-unit operated by the processor.
  • the fourth threshold may take a value in a range of 0.75-0.95. If the similarity between the clusters is greater than the fourth threshold, it indicates that the two clusters to be clustered corresponding to the similarity between the clusters are very similar, and the two clusters to be clustered can be considered as one class, and the two clusters are combined. Clusters to be clustered.
  • step S304 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a second response sub-unit operated by the processor.
  • S305 continue to cluster and merge the new cluster to be clustered with the other clusters to be clustered that are not merged, until there is no cluster to be clustered, that is, repeat The above steps S302 to S304 until there are no clusters to be clustered that can be merged.
  • the fourth threshold is set to 0.75. Since 0.95>0.75, 0.95 corresponds.
  • the two clusters to be clustered are merged into a new cluster to be clustered, and the process returns to step S302 to continue to acquire the similarity between clusters of each cluster to be clustered and other clusters to be clustered until the largest cluster is similar.
  • the degree is not greater than the fourth threshold, that is, until there are no clusters to be clustered that can be merged.
  • the current plurality of clusters to be clustered may be used as the plurality of cluster clusters corresponding to the plurality of sample data.
  • the processor may further perform outliers on each of the plurality of cluster clusters. Point separation is performed to obtain clustering results optimized by the plurality of sample data.
  • the operation of performing an outlier separation process on each of the plurality of cluster clusters may be performed by a processor invoking a corresponding instruction stored in a memory, or may be The separation unit 601 in which the processor operates is executed.
  • the separation process includes the following steps:
  • S31, 1 is obtained for each sample data corresponding to clusters to be outliers and C clades be non-outlier clusters.
  • step S31 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by an obtaining sub-unit operated by the processor.
  • X corresponds to the same data to be present outlier clusters comprising: a sample data X
  • the sample data corresponding to X to be a non-outlier clusters comprises: clustering the cluster C 1 in addition to other samples of data samples X of data.
  • clades comprising sample data C 1 are: X 1, X 2, X 3 and X 4, then, X 1 be the corresponding outlier clusters ⁇ X 1 ⁇ , X 1 be the corresponding non-outlier clusters to ⁇ X 2, X 3, X 4 ⁇ ; X 2 corresponding to clusters to be outliers ⁇ X 2 ⁇ , X 2 be the corresponding non-outlier clusters ⁇ X 1, X 3, X 4 ⁇ , X 3 corresponding to It is outlier clusters ⁇ X 3 ⁇ , X 3 corresponding to the non-outlier clusters to be ⁇ X 1, X 2, X 4 ⁇ , X 4 to be outlier clusters corresponding to ⁇ X 4 ⁇ , X 4 corresponding The non-outlier cluster is ⁇ X 1 , X 2 , X 3 ⁇ .
  • step S32 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a first acquisition sub-unit operated by the processor.
  • step S33 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a first determining sub-unit operated by the processor.
  • the third threshold may take a value in the range of 0.3-0.5. If the smallest inter-cluster similarity is less than the third threshold, then say It is obvious that the clusters to be separated from the clusters are very similar to the non-out clusters. It can be understood that the clusters to be separated from the clusters are not one type, and the sample data of the clusters to be separated are included. The processor needs to separate the cluster to be separated from the cluster that is not to be separated from the cluster.
  • the minimum inter-cluster similarity is less than a third threshold, and the minimum inter-cluster similarity corresponding to the to-be-out cluster and the non-out-of-out cluster are respectively taken as two new clusters.
  • step S34 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a first response sub-unit operated by the processor.
  • the inter-cluster similarity calculated by the processor has 0.25, 0.2, 0.7, and 0.5, and the smallest inter-cluster similarity obtained therefrom is 0.2 (0.2 corresponds to the cluster to be ⁇ X 2 ⁇
  • the third threshold is 0.3. Since 0.2 ⁇ 0.3, the processor can determine that the cluster to be separated ⁇ X 2 ⁇ is an outlier, and the processor treats the corresponding clusters and non-out clusters corresponding to 0.2 as two.
  • the point separation operation is performed until the minimum inter-cluster similarity is not less than the third threshold, that is, until there is no separable sample data.
  • the processor clusters the sample data
  • the sample data is first combined according to the similarity between samples of each of the plurality of sample data to obtain initial clustering.
  • Clustering reducing the number of clustering clusters when clustering is performed, and the processor performs clustering and merging according to the initial clustering clusters at this time, and obtains multiple clustering clusters corresponding to the plurality of sample data, thereby effectively improving clustering. speed.
  • the apparatus shown in FIG. 9 includes: an obtaining unit 501, configured to acquire an inter-sample similarity between every two sample data in the M sample data, where M is a positive integer;
  • the merging unit 502 is configured to merge the M sample data into N initial cluster clusters according to the obtained inter-sample similarity, where N is a positive integer smaller than M;
  • the clustering unit 503 is configured to perform clustering and combining the N initial cluster clusters to obtain a plurality of cluster clusters corresponding to the M sample data.
  • the merging unit 502 can include:
  • a first merging subunit for the M sample data, for any two sample data, if the inter-sample similarity of the two sample data is greater than the first threshold, then Combining the two sample data into a suspected initial cluster cluster;
  • a second merging subunit (not shown in FIG. 9), configured to merge at least two suspected initial cluster clusters including the number of identical sample data not less than a second threshold for the obtained plurality of suspected initial cluster clusters Initialize the cluster cluster for one.
  • the merging unit 502 may further include:
  • the first is a sub-unit (not shown in FIG. 9) for initializing the cluster cluster for the plurality of suspected initial clusters, and the suspected initial cluster cluster that cannot be merged with other suspected initial cluster clusters is used as an initial cluster cluster.
  • the second threshold is determined according to a linear function of a sum of the number of sample data included in each of the at least two suspected initial cluster clusters.
  • the device may further include:
  • the separating unit 601 is configured to perform outlier separation on each of the plurality of cluster clusters to obtain clustering results optimized by the plurality of sample data.
  • the separating unit 601 is specifically configured to perform outlier separation on a cluster of clusters
  • the separation unit 601 can include:
  • Obtaining a subunit (not shown in FIG. 10), configured to obtain a to-be-off cluster and a non-out-of-out cluster corresponding to each sample data in the cluster cluster, wherein the sample data corresponds to the to-be-disengaged
  • the sample data is included in the cluster, and the non-out-of-group cluster includes: other sample data in the cluster cluster except the sample data;
  • a first acquiring sub-unit (not shown in FIG. 10), configured to acquire an inter-cluster similarity between the to-be-away cluster and the non-out-of-group cluster corresponding to each sample data;
  • a first determining sub-unit (not shown in FIG. 10), configured to determine whether a minimum inter-cluster similarity among a plurality of inter-cluster similarities corresponding to all sample data in the cluster cluster is less than a third threshold;
  • a first response subunit (not shown in FIG. 10), configured to: in response to the minimum inter-cluster similarity being less than the third threshold, the minimum inter-cluster similarity corresponding to the cluster to be separated and non- The clusters to be separated are respectively used as two new cluster clusters, and the obtained subunits are triggered, and the cluster clusters corresponding to the non-outgoing clusters are further subjected to the outlier separation operation until there is no separable cluster. cluster.
  • the clustering unit 503 may include:
  • a second as a sub-unit configured to use the N initial cluster clusters as a plurality of clusters to be clustered;
  • a second acquisition sub-unit (not shown in FIG. 9), configured to acquire the inter-cluster similarity between each cluster to be clustered and other clusters to be clustered;
  • a second determining sub-unit (not shown in FIG. 9), configured to determine whether a maximum inter-cluster similarity among all inter-cluster similarities corresponding to the plurality of clusters to be clustered is greater than a fourth threshold;
  • a second response subunit (not shown in FIG. 9), configured to merge the two clusters to be clustered corresponding to the maximum inter-cluster similarity in response to the maximum inter-cluster similarity being greater than the fourth threshold Obtaining a new cluster to be clustered, and triggering the second acquiring subunit, and continuing to construct a new plurality of clusters to be clustered by the cluster to be clustered and other clusters to be clustered Clustering is performed until there are no clusters to be clustered that can be merged.
  • the M sample data is M images.
  • the inter-sample similarity between two images includes: a cosine distance between two feature vectors corresponding to the two images respectively.
  • the sample data is first combined according to the similarity between samples of each of the plurality of sample data to obtain an initial cluster, and the clustering is reduced.
  • the number of clustering clusters is clustered and merged according to the initial clustering clusters at this time, and multiple clustering clusters corresponding to the plurality of sample data are obtained, which effectively improves the clustering speed.
  • FIG. 11 is a schematic structural diagram of an electronic device according to the present disclosure.
  • the electronic device includes a housing 701 , a processor 702 , a memory 703 , a circuit board 704 , and a power circuit 705 .
  • the circuit board 704 is disposed on the circuit board 704 .
  • the power supply circuit 705 is used to supply power to various circuits or devices of the electronic device;
  • the memory 703 is used to store executable program code; 702 runs a program corresponding to the executable program code by reading executable program code stored in the memory 703 for performing the following steps:
  • the sample data when clustering the sample data, the sample data is first combined according to the similarity between samples of each of the plurality of sample data to obtain an initial cluster, and the cluster is initialized when the cluster is reduced.
  • the number of clusters is clustered and merged according to the initial cluster clusters at this time, and multiple cluster clusters corresponding to the plurality of sample data are obtained, which is beneficial to improving the clustering speed.
  • the electronic device exists in a variety of forms including, but not limited to:
  • Mobile communication devices These devices are characterized by mobile communication functions and are mainly aimed at providing voice and data communication.
  • Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones.
  • Ultra-mobile personal computer equipment This type of equipment belongs to the category of personal computers, has computing and processing functions, and generally has mobile Internet access.
  • Such terminals include: PDAs, MIDs, and UMPC devices, such as the iPad.
  • Portable entertainment devices These devices can display and play multimedia content. Such devices include: audio, video players (such as iPod), handheld game consoles, e-books, and smart toys and portable car navigation devices.
  • the server consists of a processor, a hard disk, a memory, a system bus, etc.
  • the server is similar to a general-purpose computer architecture, but because of the need to provide highly reliable services, processing power and stability High reliability in terms of reliability, security, scalability, and manageability.
  • the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Evolutionary Computation (AREA)
  • Evolutionary Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Artificial Intelligence (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Probability & Statistics with Applications (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种聚类方法、装置及电子设备,该方法包括:获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数(S101);根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数(S102);对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇(S103)。

Description

聚类方法、装置及电子设备
本公开要求在2016年7月22日提交中国专利局、申请号为201610586139.8、发明名称为“聚类方法、装置及电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本发明涉及数据处理技术领域,特别涉及一种聚类方法、装置及电子设备。
背景技术
生活中常常存在大量错综复杂的数据,为了便于管理或者查找这些数据等因素,往往需要对这些数据进行聚类,将具有相同或相似特征的数据聚为一类,如:将包含花朵的图像聚为一类等,这样,用户就可以按照类别查找或管理需要的数据。如何快速准确的将具有相同或相似特征的数据聚为一类,是非常值得关注的。
发明内容
本公开公开一种聚类技术方案。
本公开公开了一种聚类方法,所述方法包括:获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
在本公开的一种可选的实现方式中,所述根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,包括:所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
在本公开的一种可选的实现方式中,所述根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,包括:所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;针对包含相同样本数据的个数不小于第二阈值的第一疑似初始化聚类簇和第二疑似初始化聚类簇,计算第一和第二疑似初始化聚类簇的交集的模,并计算第一初始化聚类簇的模与第二初始化聚类簇的模之和与第一常数的商,在所述交集的模与所述商之差不小于第二常数时,将所述第一和第二疑似初始化聚类簇合并为一疑似初始化聚类簇。
在本公开的一种可选的实现方式中,所述根据获取的所述样本间相似度将所述M个样本数据合并为N个初始化聚类簇,还包括:针对得到的多个疑似初始化聚类簇,将不能与其他疑似初始化聚类簇合并的各疑似初始化聚类簇分别作为一初始化聚类簇。
在本公开的一种可选的实现方式中,所述第二阈值根据所述至少二个疑似初始化聚类簇中的一个或者多个疑似初始化聚类簇所包括的样本数据的个数的线性函数确定。
在本公开的一种可选的实现方式中,所述方法还包括:对所述多个聚类簇中的至少一个聚类簇进行离群点分离,且所述离群点分离处理后获得的所有聚类簇被作为所述M个样本数据对应的多个聚类簇。
在本公开的一种可选的实现方式中,所述离群点分离包括:针对任一聚类簇,获得所述聚类簇内每一样本数据对应的待离群簇和非待离群簇,其中,每一样本数据对应的所述待离群簇中包括所述样本数据,所述非待离群簇中包括:所述聚类簇中除所述样本数据外的其它样本数据;获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度;确定所述聚类簇中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值;响应于所述最小的簇间相似度小于所述第三阈值,将所述最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇。
在本公开的一种可选的实现方式中,所述对所述N个初始化聚类簇进行聚类合并,包括:将所述N个初始化聚类簇作为多个待聚类簇;获取每个待聚类簇与其他待聚类簇之间的簇间相似度;确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值;响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇;对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇。
在本公开的一种可选的实现方式中,所述M个样本数据为M个图像,且两个图像之间的样本间相似度包括:所述两个图像分别对应的两个特征向量之间的余弦距离。
本公开还提供了一种聚类装置,其特征在于,所述装置包括:
获取单元,用于获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;合并单元,用于根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;聚类单元,用于对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
在本公开的一种可选的实现方式中,所述合并单元包括:第一合并子单元,用于所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;第二合并子单元,用于针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
在本公开的一种可选的实现方式中,所述合并单元具体用于:所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;针对包含相同样本数据的个数不小于第二阈值的第一疑似初始化聚类簇和第二疑似初始化聚类簇,计算第一和第二疑似初始化聚类簇的交集的模,并计算第一初始化聚类簇的模与第二初始化聚类簇的模之和与第一常数的商,在所述交集的模与所述商之差不小于第二常数时,将所述第一和第二疑似初始化聚类簇合并为一疑似初始化聚类簇。
在本公开的一种可选的实现方式中,所述合并单元还包括:第一作为子单元,用于针对得到的多个疑似初始化聚类簇,将不能与其他疑似初始化聚类簇合并的各疑似初始化聚类簇分别作为一初始化聚类簇。
在本公开的一种可选的实现方式中,所述第二阈值根据所述至少二个疑似初始化聚类簇中的一个或者多个疑似初始化聚类簇所包括的样本数据的个数的线性函数确定。
在本公开的一种可选的实现方式中,所述装置还包括:
分离单元,用于对所述多个聚类簇中的至少一个聚类簇进行离群点分离,且所述离群点分离处理后获得的所有聚类簇被作为所述M个样本数据对应的多个聚类簇。
在本公开的一种可选的实现方式中,所述分离单元具体用于对一所述聚类簇进行离群点分离;所述分离单元包括:
获得子单元,针对任一聚类簇,用于获得所述聚类簇内每一样本数据对应的待离群簇和非待离群簇,其中,每一样本数据对应的所述待离群簇中包括所述样本数据,所述非待离群簇中包括:所述聚类簇中除所述样本数据外的其它样本数据;
第一获取子单元,用于获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度;
第一确定子单元,用于确定所述聚类簇中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值;
第一响应子单元,用于响应于所述最小的簇间相似度小于所述第三阈值,将所述最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇。
在本公开的一种可选的实现方式中,所述聚类单元包括:
第二作为子单元,用于将所述N个初始化聚类簇作为多个待聚类簇;
第二获取子单元,用于获取每个待聚类簇与其他待聚类簇之间的簇间相似度;
第二确定子单元,用于确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值;
第二响应子单元,用于响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇,并触发所述第二获取子单元,对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇。
在本公开的一种可选的实现方式中,所述M个样本数据为M个图像,且两个图像之间的样本间相似度包括:所述两个图像分别对应的两个特征向量之间的余弦距离。
本公开还公开了一种电子设备,包括:壳体、处理器、存储器、电路板和电源电路,其中,所述电路板安置在所述壳体围成的空间内部,所述处理器和所述存储器设置在所述电路板上;所述电源电路,用于为终端的各个电路或器件供电;所述存储器用于存储可执行程序代码;所述处理器通过读取所述存储器中存储的可执行程序代码来运行与可执行程序代码对应的程序,以用于执行上述聚类方法对应的操作。
本公开还公开了一种非暂时性计算机存储介质,该介质可存储计算机可读指令,当这些指令被执行时可使处理器执行上述聚类方法对应的操作。
本公开中,对样本数据进行聚类时,先根据多个样本数据中每两个样本数据之间的样本间相似度对样本数据进行合并,获得初始化聚类簇,减少聚类时初始化聚类簇的数量,根据此时的初始化聚类簇进行聚类合并,得到所述多个样本数据对应的多个聚类簇,有利于提高聚类速度。
附图说明
为了更清楚地说明本公开或现有技术中的技术方案,下面将对本公开或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1示出了本公开的一应用场景示意图;
图2示出了本公开的另一应用场景示意图;
图3示出了实现本公开的一示例性设备的框图;
图4示出了实现本公开的另一示例性设备的框图;
图5为本公开提供的聚类方法的流程示意图;
图6本公开提供的初始聚类簇之间的簇间相似度示意图;
图7为本公开提供的初始化聚类簇聚类合并方法的流程示意图;
图8为本公开提供的簇间相似度的确定方法的流程示意图;
图9为本公开提供的聚类装置的结构示意图;
图10为本公开提供的聚类装置的结构示意图;
图11为本公开提供的电子设备的结构示意图。
具体实施例
下面将结合本公开中的附图,对本公开中的技术方案进行清楚、完整地描述,显然,所描述的技术方案仅仅是本公开一部分技术方案,而不是全部的技术方案。基于本公开中的技术方案,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
图1示意性地示出了根据本公开提供的聚类技术方案可以在其中实现的一应用场景。
图1中,用户的智能移动电话1000中存储有多张图片,例如,用户的智能移动电话1000中存储的图片包括:第一图片1001、第二图片1002、第三图片1003、……以及第M图片1004等共M张图片,其中,M的大小可以由智能移动电话1000中存储的图片数量决定,且在通常情况下,M为不小于2的正整数。存储于智能移动电话1000中的图片可以是用户利用智能移动电话1000拍摄的图片,也可以是用户将其他终端设备中存储的图片传输并存储在智能移动电话1000中的图片,还可以是用户利用智能移动电话1000从网络中下载的图片等。
本公开提供的技术方案可以使智能移动电话1000中存储的第一图片1001、第二图片1002、第三图片1003、……以及第M图片1004按照预定分类方式而被自动的划分在多个集合中,例如,在预定分类方式为根据图片中的人进行划分的情况下,由于智能移动电话1000中存储的第一图片1001和第M图片1004中均包含有第一用户的肖像,而智能移动电话1000中存储的第二图片1002和第三图片1003均包含有第N用户的肖像,因此,本公开提供的技术方案可以使智能移动电话1000中存储的第一图片1001和第M图片1004自动划分在第一用户的图片集合中,并使智能移动电话1000中存储的第二图片1002和第三图片1003自动划分在第N用户的图片集合中, 上述N为小于M的正整数。
本公开通过将智能移动电话1000中存储的第一图片1001、第二图片1002、第三图片1003、……以及第M图片1004自动划分在多个集合中,有利于提高图片的可管理性,从而可以为用户带来操作上的便利,例如,用户可以方便的浏览智能移动电话1000中存储的包含有第一用户的肖像的所有图片。
图2示意性地示出了根据本公开提供的聚类技术方案可以在其中实现的另一应用场景。
图2中,电商平台的设备2000(如计算机或者服务器等)中存储有电商平台所能提供的所有商品信息或者部分商品信息,例如,电商平台的设备2000中存储的商品信息包括:第一商品信息2001、第二商品信息2002、第三商品信息2003、……以及第M商品信息2004等共M个商品信息,其中,M的大小可以由设备2000中存储的商品信息的数量决定,且在通常情况下,M为不小于2的正整数。存储于设备2000中的商品信息可以包括:商品的图片以及商品的文字描述等。
本公开提供的技术方案可以使设备2000中存储的第一商品信息2001、第二商品信息2002、第三商品信息2003、……以及第M商品信息2004按照预定分类方式而被自动划分在多个集合中,例如,在预定分类方式为根据商品类别进行划分的情况下,由于设备2000中存储的第一商品信息2001所对应的商品和第M商品信息2004所对应的商品均属于第一类别(如女士旅游鞋类别),而设备2000中存储的第二商品信息2002所对应的商品和第三商品信息2003所对应的商品均属于第N类别(如婴儿奶粉类别),因此,本公开提供的技术方案可以使设备2000中存储的第一商品信息2001和第M商品信息2004自动划分在第一类别(如女士旅游鞋类别)集合中,并使设备2000中存储的第二商品信息2002和第三商品信息2003自动划分在第N类别(如婴儿奶粉类别)集合中,上述N为小于M的正整数。
本公开通过将设备2000中存储的第一商品信息2001、第二商品信息2002、第三商品信息2003、……以及第M商品信息2004划分在多个集合中,有利于提高商品信息的可管理性,从而可以为电商平台以及消费者带来操作上的便利,例如,在消费者查看设备2000中存储的第一商品信息2001的情况下,电商平台可以将第一商品信息2001所在的第一类别集合中的其他商品信息(例如,第M商品信息)推荐展示给用户。
然而,本领域技术人员可以理解,本公开还可以适用于其他应用场景中,即本公开所能够适用的应用场景并不会受上述例举的两个应用场景的限制。
下面结合附图通过具体的实施例对本公开的聚类技术方案进行详细介绍。
图3示出了适于实现本公开的示例性设备3000(例如,计算机系统/服务器)的框图。图3显示的设备3000仅仅是一个示例,不应对本公开的功能和使用范围带来任何限制。
如图3所示,设备3000可以以通用计算设备的形式表现。设备3000的组件可以包括但不限于:一个或者多个处理器或者处理单元3001,系统存储器3002,连接不同系统组件(包括系统存储器3002和处理单元3001)的总线3003。设备3000可以包括多种计算机系统可读介质。这些介质可以是任何能够被设备3000访问的可用介质,包括易失性和非易失性介质,可移动的和不可移动的介质等。
系统存储器3002可以包括易失性存储器形式的计算机系统可读介质,例如,随机存取存储 器(RAM)3021和/或高速缓存存储器3022。设备3000可以进一步包括其他可移动的/不可移动的、易失性/非易失性计算机系统存储介质。仅作为举例,ROM 3023可以用于读写不可移动的、非易失性磁介质(图3中未显示,通常称为“硬盘驱动器”)。尽管未在图3中示出,可以提供用于对可移动非易失性磁盘(例如“软盘”)读写的磁盘驱动器,以及对可移动非易失性光盘(例如CD-ROM,DVD-ROM或者其他光介质)读写的光盘驱动器。在这些情况下,每个驱动器可以通过一个或者多个数据介质接口与总线3003相连。系统存储器3002中可以包括至少一个程序产品,该程序产品具有一组(例如至少一个)程序模块,这些程序模块被配置以执行本公开的功能。
具有一组(至少一个)程序模块3024的程序/实用工具3025,可以存储在例如系统存储器3002中,这样的程序模块3024包括但不限于:操作系统、一个或者多个应用程序、其他程序模块以及程序数据,这些示例中的每一个或某种组合中可能包括网络环境的实现。程序模块3024通常执行本公开所描述的功能和/或方法。
设备3000也可以与一个或多个外部设备3004(如键盘、指向设备、显示器等)通信。这种通信可以通过输入/输出(I/O)接口3005进行,并且,设备3000还可以通过网络适配器3006与一个或者多个网络(例如局域网(LAN),广域网(WAN)和/或者公共网络,例如因特网)通信。如图3所示,网络适配器3006通过总线3003与设备3000的其他模块(如处理单元3001等)通信。应当明白,尽管图3中未示出,可以结合设备3000使用其他硬件和/或软件模块。
处理单元3001(即处理器)通过运行存储在系统存储器3002中的计算机程序,从而执行各种功能应用以及数据处理,例如,执行用于实现上述方法中的各步骤的指令;具体而言,处理单元3001可以执行系统存储器3002中存储的计算机程序,且该计算机程序被执行时,下述步骤被实现:获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
图4示出了适于实现本公开的示例性设备4000,设备4000可以是移动终端、个人计算机(PC)、平板电脑以及服务器等。图4中,设备4000包括一个或者多个处理器、通信部等,所述一个或者多个处理器可以为:一个或者多个中央处理单元(CPU)4001,和/或,一个或者多个图像处理器(GPU)4130等,处理器可以根据存储在只读存储器(ROM)4002中的可执行指令或者从存储部分4008加载到随机访问存储器(RAM)4003中的可执行指令而执行各种适当的动作和处理。通信部4120可以包括但不限于网卡,所述网卡可以包括但不限于IB(Infiniband)网卡。处理器可与只读存储器4002和/或随机访问存储器4300中通信以执行可执行指令,通过总线4004与通信部4120相连、并经通信部4120与其他目标设备通信,从而完成本公开中的相应操作。处理器所执行的操作包括:获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
此外,在RAM 4003中,还可以存储有装置操作所需的各种程序以及数据。CPU4001、ROM4002以及RAM4003通过总线4004彼此相连。在有RAM4003的情况下,ROM4002为可选模块。RAM4003存储可执行指令,或在运行时向ROM4002中写入可执行指令,可执行指令使中央处理单元4001 执行上述通信方法对应的操作。输入/输出(I/O)接口4005也连接至总线4004。通信部4120可以集成设置,也可以设置为具有多个子模块(例如,多个IB网卡),并分别与总线连接。
以下部件连接至I/O接口4005:包括键盘、鼠标等的输入部分4006;包括诸如阴极射线管(CRT)、液晶显示器(LCD)等以及扬声器等的输出部分4007;包括硬盘等的存储部分4008;以及包括诸如LAN卡、调制解调器等的网络接口卡的通信部分4009。通信部分4009经由诸如因特网的网络执行通信处理。驱动器4100也根据需要连接至I/O接口4005。可拆卸介质4110,诸如磁盘、光盘、磁光盘、半导体存储器等等,根据需要安装在驱动器4100上,以便于从其上读出的计算机程序根据需要被安装入存储部分4008。
需要说明的,如图4所示的架构仅为一种可选实现方式,在具体实践过程中,可根据实际需要对上述图4的部件数量和类型进行选择、删减、增加或替换;在不同功能部件设置上,也可采用分离设置或集成设置等实现方式,例如,GPU和CPU可分离设置,再如理,可将GPU集成在CPU上,通信部可分离设置,也可集成设置在CPU或GPU上等。这些可替换的实施方式均落入本公开的保护范围。
特别地,根据本公开的实施方式,下文参考流程图描述的过程可以被实现为计算机软件程序,例如,本公开的实施方式包括一种计算机程序产品,其包括有形地包含在机器可读介质上的计算机程序,计算机程序包含用于执行流程图所示的步骤的程序代码,程序代码可包括对应执行本申请提供的步骤对应的可执行指令,例如,用于获取M个样本数据中每两个样本数据之间的样本间相似度的可执行指令,M为正整数;用于根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇的可执行指令,N为小于M的正整数;用于对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇的可执行指令。
在这样的实施方式中,该计算机程序可以通过通信部分4009从网络上被下载及安装,和/或从可拆卸介质4110被安装。在该计算机程序被中央处理单元(CPU)4001执行时,执行本公开中记载的上述可执行指令。
下面结合图5-图11对本公开提供的聚类技术方案进行说明。本公开提供的任一种聚类技术方案可由软件或者硬件或者软硬结合的方式进行示例。例如,本公开提供的聚类技术方案可由某一电子设备实施或者由某一处理器实施,本公开并不限制,所述电子设备可包括但不限于终端或服务器,所述处理器可包括但不限于CPU或GPU。以下不再赘述。
图5中,S101:获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数。通常情况下,M大于2,如M可为100或者1001或者5107等。
一种可选的实现方式中,步骤S101可以由处理器调用存储器存储的指令执行,或者,可以由被处理器运行的获取单元501执行。
本公开中的样本数据可以为图像(即图片)、语音、视频或者文本等。
一种可选的实现方式中,在处理器获得样本数据后,处理器可以采用卷积神经网络或其他传统的局部特征描述子对样本数据的特征进行提取,如:当样本数据为图像时,处理器可以采用Sift(Scale-invariant feature transform,尺度不变特征转换)、HOG(Histogram of Oriented Gradien,方向梯度直方图)等对图像进行特征提取,获得每个样本数据的特征向量;处理器根据提取的样本数据的特征向量构建特征矩阵。假设,当前有n个样本数据,提取每一样本数据 的m类特征,根据提取到的特征构建n*m阶特征矩阵。
本公开中的两个样本数据之间的样本间相似度可通过但不限于采用样本数据特征向量的余弦距离、欧式距离等表示。
例如,处理器在建立特征矩阵后,可以计算样本数据间特征的余弦距离,生成相似度矩阵,这里,一个样本数据的所有特征作为一个特征向量。根据上述假设,处理器生成了n*m阶特征矩阵,根据该特征矩阵,处理器可以生成n*n阶相似度矩阵,如下所示:
Figure PCTCN2017091432-appb-000001
其中,Mij为第i个样本数据和第j个样本数据之间的样本间相似度,
Figure PCTCN2017091432-appb-000002
X表示第i个样本数据的特征向量,Y表示第j个样本数据的特征向量。
相似度矩阵表征的是样本数据间的相似度度量,当两个样本数据的样本间相似度非常高(如超过某一设定阈值,或者,相似度满足一定的预定条件等)时,可以近似地认为这两个样本数据属于同一个类。另外,该相似度矩阵为对称矩阵。在本公开的输入信息为M个样本数据的相似度矩阵的情况下,处理器可以通过接收输入信息而获得M个样本数据的相似度矩阵,然而,在本公开的输入信息为M个样本数据的情况下,处理器可以采用上述描述的方案获得M个样本数据的相似度矩阵。本公开不限制处理器获得M个样本数据的相似度矩阵的具体实现方式。
S102:根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数。
一种可选的实现方式中,步骤S102可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的合并单元502执行。其中,每个初始化聚类族包含至少一个样本数据。
本公开在对样本数据进行聚类前,一般需要确定样本数据的初始化聚类簇。通常的做法是将每一样本数据单独确定为一个聚类簇,这些聚类簇作为初始化聚类簇,如,当前有4个样本数据,分别为a、b、c和d,则可以确定出4个初始化聚类簇。这种方法,以每一个样本数据为基础,具有良好的鲁棒性,获得的最终的聚类结果具有良好的准确率。
上述将每一样本数据单独确定为一个初始化聚类簇,根据确定的初始化聚类簇获得最终的聚类结果的方法的计算复杂度为O(n3log n),其中,n为样本数据的数量,这种方法适用于对精度要求高、而对速度性能要求不高的场景。在对处理数据的速度性能要求较高的场景中,本公开的聚类技术方案通过对样本数据进行合并,减少了初始化聚类簇的个数,有利于提高聚类的速度性能。
一种可选的实现方式中,处理器根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,可以包括下述步骤:
S01、针对一个样本数据而言,查找与该样本数据之间的样本间相似度大于第一阈值的所有样本数据,并将查找到的所有样本数据分别与该样本数据合并为一个疑似初始化聚类簇,从而处理器可以获得多个疑似初始化聚类簇。
一种可选的实现方式中,步骤S01可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第一合并子单元执行。
假设,当前有5个样本数据,分别为a、b、c、d和e,且此时处理器已经获得了a分别与b、c、d、e的样本间相似度、b分别与a、c、d、e的样本间相似度、c分别与a、b、d、e的样本间相似度、d分别与a、b、c、e的样本间相似度、e分别与a、b、c、d的样本间相似度。
根据上述假设,若a与b、d的样本间相似度大于第一阈值,且b和c的样本间相似度大于第一阈值,则处理器可以将a与b合并为一个疑似初始化聚类簇C1,将a与d合并为一个疑似初始化聚类簇C2,将b与c合并为一个疑似初始化聚类簇C3,由于e和其他所有样本数据的样本间相似度都不大于第一阈值,因此,处理器可以将e作为一个疑似初始化聚类簇C4
值得一提的是,第一阈值可选的可以为比较高的阈值,如,在样本间相似度的取值范围为[0,1]的情况下,第一阈值可以为0.9、0.95等,由此可以尽可能保证疑似初始化聚类簇分类的准确性。另外,第一阈值可以根据实际情况进行调整,例如,在每一个疑似初始化聚类簇均包括一个样本数据的情况下,处理器可以适当降低第一阈值;再例如,在包含有两个样本数据的疑似初始化聚类簇的数量与所有疑似初始化聚类簇的数量的比值低于预定比值(例如,1/100)的情况下,处理器可以适当降低第一阈值。
S02、针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
一种可选的实现方式中,步骤S02可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第二合并子单元执行。
一种可选的实现方式中,上述第二阈值可以设定为固定的正整数,例如,假设第二阈值为1,如果两个疑似初始化聚类簇中均包含有一个相同的样本数据,则处理器将这两个疑似初始化聚类簇进行合并;另外,不能与其他疑似初始化聚类簇合并的一疑似初始化聚类簇也确定为一初始化聚类簇,例如,被处理器运行的第一作为子单元将不能与其他疑似初始化聚类簇合并的一疑似初始化聚类簇确定为一初始化聚类簇。
样本间相似度具有传递性,因此处理器可以将包含相同样本数据的疑似初始化聚类簇合并为一个初始化聚类簇。
如步骤S01中的假设,疑似初始化聚类簇C1和疑似初始化聚类簇C2中都包含样本数据a,疑似初始化聚类簇C1和疑似初始化聚类簇C3中都包含样本数据b,则处理器可以将疑似初始化聚类簇C1、疑似初始化聚类簇C2和疑似初始化聚类簇C3合并为一个初始化聚类簇C5;另外,疑似初始化聚类簇C4与其他疑似初始化聚类簇不包含相同样本数据,则处理器不对疑似初始化聚类簇C4进行合并处理,将疑似初始化聚类簇C4确定为初始化聚类簇。
本公开的聚类技术方案的计算复杂度为O(n2log n),与上述将每一样本数据单独确定为一 个初始化聚类簇,并根据确定的初始化聚类簇获得最终的聚类结果的技术方案的计算复杂度O(n3log n)相比,在较大程度上提高了对样本数据进行聚类获得最终的聚类结果的计算速度。
本公开合并疑似初始化聚类簇的一种方式为,只要两个疑似初始化聚类簇所包含的相同样数据的个数不小于第二阈值,就会将这两个疑似初始化聚类簇合并。
实际应用的某些情形中,两个疑似初始化聚类簇可能并不属于同一类,可这两个疑似初始化聚类簇中存在极少数的样本数据间的样本间相似度大于第一阈值,也就是,由于疑似初始化聚类簇中存在噪声点样本数据,导致两个疑似初始化聚类簇合并,这样在确定初始化聚类簇阶段可能出现误差。
如图6所示,图6中,Xi为样本数据,样本数据间存在实线连接表示两样本数据间的样本间相似度大于第一阈值,疑似初始化聚类簇C1和疑似初始化聚类簇C2之间只有X19和X12相连,若采用上述疑似初始化聚类簇合并方法,那么会将疑似初始化聚类簇C1和疑似初始化聚类簇C2合并,这使得不应合并两个疑似初始化聚类簇合并为了一个,在确定初始化聚类簇阶段出现了误差。
为了减小甚至避免在确定初始化聚类簇阶段出现误差,本公开中的第二阈值可以根据所述至少二个疑似初始化聚类簇中一个或者多个疑似初始化聚类簇所包括的样本数据的个数的线性函数确定。如:假设斜率k为0.5,当一个疑似初始化聚类簇C1中包含2个样本数据时,若其想与疑似初始化聚类簇C2合并,第二阈值为1;当疑似初始化聚类簇C1中包含4个样本数据时,若其想与疑似初始化聚类簇C2合并,第二阈值为2,依此类推。
本公开合并疑似初始化聚类簇的另一种可选的实现方式为:不仅两个疑似初始化聚类簇所包含的相同样本数据的个数不小于第二阈值,而且两个疑似初始化聚类簇还应满足下述公式(1)才能够被合并:
|S1∩S2|≥(|S1|+|S2|)/β+γ;        (1)
其中,S1为一个疑似初始化聚类簇对应的样本数据集合,在所述S1中,一个样本数据与其他各样本数据的样本间相似度均大于第一阈值,S2为另一个疑似初始化聚类簇对应的样本数据集合,在所述S2中,一个样本数据与其他各样本数据之间的样本相似度均大于第一阈值,β和γ为预设参数。其中,可选地,2≤β≤3,-3≤γ≤-1,例如,β=2.5,γ=-1。
合并两个疑似初始化聚类簇的过程可以具体为:
判断第一疑似初始化聚类簇和第二疑似初始化聚类簇是否满足公式(1),其中,第一疑似初始化聚类簇和第二疑似初始化聚类簇包含相同的样本数据;
若为是,将第一疑似初始化聚类簇和第二疑似初始化聚类簇合并为一个新的疑似初始化聚类簇,再次根据公式(1)将该新的疑似初始化聚类簇与其他疑似初始化聚类簇进行合并;
若为否,拒绝将第一疑似初始化聚类簇和第二疑似初始化聚类簇合并为一个新的疑似初始化聚类簇。
当所有疑似初始化聚类簇都不能再合并时,将所有疑似初始化聚类簇,都确定为初始化聚类簇。
如步骤S01中的假设,由于疑似初始化聚类簇C1和疑似初始化聚类簇C2中都包含样本数据 a,因此,处理器将疑似初始化聚类簇C1和疑似初始化聚类簇C2分别看作是一个样本数据集合,若C1和C2的交集的模、C1的模和C2的模满足公式(1),也就是,下述公式成立:
|C1∩C2|≥(|C1|+|C2|)/β+γ,
那么处理器就将疑似初始化聚类簇C1和疑似初始化聚类簇C2合并为一个新的疑似初始化聚类簇C5,此时疑似初始化聚类簇C5中包含样本数据a样本数据b和样本数据d、而疑似初始化聚类簇C3中包含样本数据b和样本数据c,疑似初始化聚类簇C5和疑似初始化聚类簇C3中都包含样本数据b,则处理器再次判断疑似初始化聚类簇C5和疑似初始化聚类簇C3是否满足公式(1),若不满足公式(1),则处理器拒绝合并疑似初始化聚类簇C5和疑似初始化聚类簇C3,由于疑似初始聚类簇C4包含样本数据e,疑似初始聚类簇C4与疑似初始聚类簇C5和疑似初始化聚类簇C3未包含有相同的样本数据,因此,处理器分别将疑似初始化聚类簇C5、疑似初始化聚类簇C3和疑似初始聚类簇C4确定为初始化聚类簇;
若C1和C2的交集的模、C1的模和C2的模不满足公式(1),则处理器拒绝合并疑似初始化聚类簇C1和疑似初始化聚类簇C2,由于疑似初始聚类簇C1和疑似初始化聚类簇C3都包含样本数据b,处理器再次判断疑似初始化聚类簇C1和疑似初始化聚类簇C3是否满足公式(1),假设还是不满足公式(1),另外,疑似初始聚类簇C4包含样本数据e,疑似初始聚类簇C4与疑似初始聚类簇C1、疑似初始聚类簇C2和疑似初始化聚类簇C3未包含有相同的样本数据,因此,处理器不对其进行合并处理,处理器将疑似初始化聚类簇C1、疑似初始化聚类簇C2、疑似初始化聚类簇C3和疑似初始化聚类簇C4分别确定为初始化聚类簇。
本公开在两个疑似初始化聚类簇所包含的相同样数据的个数不小于第二阈值,且在满足公式(1)的情况下,才将两个疑似初始化聚类簇进行合并,在两个疑似初始化聚类簇合并过程中考虑了(|S1∩S2|)的情况,因此有利于提高初始化聚类簇的准确性,有利于提高整个聚类技术的鲁棒性。
S103:对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
一种可选的实现方式中,步骤S103可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的聚类单元503执行。
一种可选的实现方式中,参考图7,对所述N个初始化聚类簇进行聚类合并,可以包括如下步骤:
S301:将所述N个初始化聚类簇作为多个待聚类簇。
一种可选的实现方式中,步骤S301可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第二作为子单元执行。
S302:获取每个待聚类簇与其他待聚类簇之间的簇间相似度。
一种可选的实现方式中,步骤S302可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第二获取子单元执行。
处理器(如第二获取子单元)在计算每两个待聚类簇之间的簇间相似度时,可参考图8,图8为本公开提供的一种簇间相似度的确定方法的流程示意图,该方法包括如下步骤:
S401、获得第一待聚类簇中的每一样本数据与第二待聚类簇中的每一样本数据之间的至少 一个第一样本间相似度。
假设,第一待聚类簇中包含的样本数据为:a、b;第二待聚类簇中包含的样本数据为:c、d;那么确定的第一样本间相似度包括:a和c间的样本间相似度、a和d间的样本间相似度、b和c间的样本间相似度、b和d间的样本间相似度。
S402、确定相似度范围。
其中,相似度范围的最小值大于第一样本间相似度中的最小的样本间相似度,相似度范围的最大值小于第一样本间相似度中的最大的样本间相似度。实际应用中,第一样本间相似度中的较大值和较小值对应的样本数据在一般情况下很可能为噪声点数据,因此,在确定相似度范围时,取第一样本间相似度中的中间范围为相似度范围,根据该相似度范围内的第一样本间相似度,计算两待聚类簇之间的簇间相似度,能够最大程度的去除噪声点数据对计算两待聚类簇之间的簇间相似度准确性的影响,进而提高最终聚类结果的准确率。
假设,处理器获得的第一样本间相似度为:0.2、0.32、0.4、0.3、0.7、0.5、0.75、0.8、0.9、0.92,此时处理器确定的相似度范围可以为:0.3-0.75;也可以为:0.4-0.7;当然也可以为其他满足上述条件的相似度范围,本公开对此不进行限定。
一种具体实现方式中,处理器可以通过以下方式确定相似度范围:
处理器对获得的第一样本间相似度进行排序,如,e1≥e2≥e3≥...≥eE,其中,E为获得的第一样本间相似度的个数,e为第一样本间相似度;
处理器根据参数l、k确定相似度范围,这里,可选的,参数l可以取0.2E,参数k可以取0.8E。
根据上述假设,处理器对其获得的第一样本间相似度进行排序,0.2<0.3<0.32<0.4<0.5<0.7<0.75<0.8<0.9<0.92,处理器根据第一样本间相似度的个数E=10,确定l=2,k=8,此时处理器确定的相似度范围可以为0.3-0.9。
S403、根据相似度范围内的第一样本间相似度,计算第一待聚类簇和第二待聚类簇之间的簇间相似度。
具体地,处理器可以根据下述公式计算第一待聚类簇和第二待聚类簇之间的簇间相似度:
Figure PCTCN2017091432-appb-000003
其中,α为预设参数,可选地,0.85<α<0.95;ei为对获得的第一样本间相似度按照从小到大的顺序排序后,从最小侧起(或从最大侧起)的第i个第一样本间相似度,l和k为上述确定出的参数。
S303:确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值。
一种可选的实现方式中,步骤S303可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第二确定子单元执行。
可选地,所述第四阈值可以在0.75-0.95的范围内取值。若簇间相似度大于第四阈值,则说明该簇间相似度对应的两个待聚类簇非常相似,可以认为这两个待聚类簇为一类,合并这两 个待聚类簇。
S304:响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇。
一种可选的实现方式中,步骤S304可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第二响应子单元执行。
S305:对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇,即重复上述步骤S302至S304,直到没有可合并的待聚类簇。
假设,当前计算得到的簇间相似度有0.85、0.92、0.7、0.6、0.95,从中获得的最大的簇间相似度为0.95,设定第四阈值为0.75,由于0.95>0.75,因此将0.95对应的两个待聚类簇合并为一个新的待聚类簇,并返回步骤S302,继续获取每个待聚类簇与其他待聚类簇之间的簇间相似度,直到最大的簇间相似度不大于第四阈值为止,也就是,直到没有可合并的待聚类簇为止。此时,可以将当前多个待聚类簇作为所述多个样本数据对应的多个聚类簇。
由于样本数据本身存在一定的噪声,并且聚类过程中存在一些可以影响最终的聚类结果鲁棒性的因素,因此,导致最终的聚类结果中可能会存在一些样本数据被错误的聚类。为了使最终呈现的聚类结果更加准确,在得到所述多个样本数据对应的多个聚类簇之后,处理器还可以对所述多个聚类簇中的每个聚类簇进行离群点分离,得到所述多个样本数据优化后的聚类结果。一种可选的实现方式中,对所述多个聚类簇中的每个聚类簇进行离群点分离处理的操作可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的分离单元601执行。
以一个聚类簇C1进行离群点分离为例,分离的过程包括如下步骤:
S31、获得聚类簇C1内每一样本数据对应的待离群簇和非待离群簇。
一种可选的实现方式中,步骤S31可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的获得子单元执行。
其中,一样本数据X对应的待离群簇中包括:样本数据X,所述样本数据X对应的非待离群簇中包括:聚类簇C1中除样本数据X外的其他样本数据。
假设,聚类簇C1包括的样本数据有:X1、X2、X3和X4,那么,X1对应的待离群簇为{X1},X1对应的非待离群簇为{X2,X3,X4};X2对应的待离群簇为{X2},X2对应的非待离群簇为{X1,X3,X4},X3对应的待离群簇为{X3},X3对应的非待离群簇为{X1,X2,X4},X4对应的待离群簇为{X4},X4对应的非待离群簇为{X1,X2,X3}。
S32、获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度。
一种可选的实现方式中,步骤S32可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第一获取子单元执行。
S33、确定聚类簇C1中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值。
一种可选的实现方式中,步骤S33可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第一确定子单元执行。
这里,第三阈值可以在0.3-0.5的范围内取值。若最小的簇间相似度小于第三阈值,则说 明该簇间相似度对应的待离群簇与非待离群簇非常不相似,可以理解为该待离群簇与该非待离群簇不是一类,该待离群簇包含的样本数据相对于该非待离群簇来说为离群点,处理器需要将该待离群簇与该非待离群簇分离。
S34、响应于最小的簇间相似度小于第三阈值,将最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇。
一种可选的实现方式中,步骤S34可以由处理器调用存储器存储的相应指令来执行,或者,可以由被处理器运行的第一响应子单元执行。
S35、对非待离群簇对应的聚类簇继续进行离群点分离操作,如重复上述步骤S32至S34,直到聚类簇中没有可分离的样本数据。
根据S31中的例子,假设,处理器计算得到的簇间相似度有0.25、0.2、0.7、0.5,从中获得的最小的簇间相似度为0.2(0.2对应的待离群簇为{X2}),第三阈值为0.3,由于0.2<0.3,因此处理器可以确定待离群簇{X2}为离群点,处理器将0.2对应的待离群簇和非待离群簇分别作为两个新的聚类簇,即聚类簇{X2}和聚类簇{X1,X3,X4};处理器对聚类簇{X1,X3,X4}继续进行离群点分离操作,直到最小的簇间相似度不小于第三阈值为止,也就是,直到没有可分离的样本数据为止。
应用图5-图8所示技术方案,处理器对样本数据进行聚类时,先根据多个样本数据中每两个样本数据之间的样本间相似度对样本数据进行合并,获得初始化聚类簇,减少了聚类时初始化聚类簇的数量,处理器根据此时的初始化聚类簇进行聚类合并,得到所述多个样本数据对应的多个聚类簇,有效地提高了聚类速度。
参考图9所示的装置包括:中,获取单元501,用于获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;
合并单元502,用于根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;
聚类单元503,用于对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
一种可选的实现方式中,所述合并单元502可以包括:
第一合并子单元(图9中未示出),用于所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;
第二合并子单元(图9中未示出),用于针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
一种可选的实现方式中,所述合并单元502还可以包括:
第一作为子单元(图9中未示出),用于针对得到的多个疑似初始化聚类簇,将不能与其他疑似初始化聚类簇合并的疑似初始化聚类簇作为一初始化聚类簇。
一种可选的实现方式中,所述第二阈值根据所述至少二个疑似初始化聚类簇中各疑似初始化聚类簇所包括的样本数据的个数总和的线性函数确定。
一种可选的实现方式中,可参考图10,所述装置在图9所示装置的基础上还可以包括:
分离单元601,用于对所述多个聚类簇中的每个聚类簇进行离群点分离,得到所述多个样本数据优化后的聚类结果。
一种可选的实现方式中,所述分离单元601具体用于对一所述聚类簇进行离群点分离;
所述分离单元601可以包括:
获得子单元(图10中未示出),用于获得所述聚类簇内每一样本数据对应的待离群簇和非待离群簇,其中,每一样本数据对应的所述待离群簇中包括所述样本数据,所述非待离群簇中包括:所述聚类簇中除所述样本数据外的其它样本数据;
第一获取子单元(图10中未示出),用于获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度;
第一确定子单元(图10中未示出),用于确定所述聚类簇中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值;
第一响应子单元(图10中未示出),用于响应于所述最小的簇间相似度小于所述第三阈值,将所述最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇,并触发所述获得子单元,对所述非待离群簇对应的聚类簇继续进行离群点分离操作,直到没有可分离的聚类簇。
一种可选的实现方式中,所述聚类单元503可以包括:
第二作为子单元(图9中未示出),用于将所述N个初始化聚类簇作为多个待聚类簇;
第二获取子单元(图9中未示出),用于获取每个待聚类簇与其他待聚类簇之间的簇间相似度;
第二确定子单元(图9中未示出),用于确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值;
第二响应子单元(图9中未示出),用于响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇,并触发所述第二获取子单元,对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇。
一种可选的实现方式中,所述M个样本数据为M个图像。
一种可选的实现方式中,两个图像之间的样本间相似度包括:所述两个图像分别对应的两个特征向量之间的余弦距离。
应用图9所示,对样本数据进行聚类时,先根据多个样本数据中每两个样本数据之间的样本间相似度对样本数据进行合并,获得初始化聚类簇,减少聚类时初始化聚类簇的数量,根据此时的初始化聚类簇进行聚类合并,得到所述多个样本数据对应的多个聚类簇,有效地提高了聚类速度。
参考图11,图11为本公开提供的一种电子设备的结构示意图,该电子设备包括:壳体701、处理器702、存储器703、电路板704和电源电路705,其中,电路板704安置在壳体701围成的空间内部,处理器702和存储器703设置在电路板704上;电源电路705,用于为电子设备的各个电路或器件供电;存储器703用于存储可执行程序代码;处理器702通过读取存储器703中存储的可执行程序代码来运行与可执行程序代码对应的程序,以用于执行以下步骤:
获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;
根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;
对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
处理器702对上述步骤的具体执行过程以及处理器702通过运行可执行程序代码来进一步执行的步骤,可以参见本公开针对图5-11所示的描述,在此不再赘述。
本公开中,对样本数据进行聚类时,先根据多个样本数据中每两个样本数据之间的样本间相似度对样本数据进行合并,获得初始化聚类簇,减少聚类时初始化聚类簇的数量,根据此时的初始化聚类簇进行聚类合并,得到所述多个样本数据对应的多个聚类簇,有利于提高聚类速度。
该电子设备以多种形式存在,包括但不限于:
(1)移动通信设备:这类设备的特点是具备移动通信功能,并且以提供话音、数据通信为主要目标。这类终端包括:智能手机(例如iPhone)、多媒体手机、功能性手机,以及低端手机等。
(2)超移动个人计算机设备:这类设备属于个人计算机的范畴,有计算和处理功能,一般也具备移动上网特性。这类终端包括:PDA、MID和UMPC设备等,例如iPad。
(3)便携式娱乐设备:这类设备可以显示和播放多媒体内容。该类设备包括:音频、视频播放器(例如iPod),掌上游戏机,电子书,以及智能玩具和便携式车载导航设备。
(4)服务器:提供计算服务的设备,服务器的构成包括处理器、硬盘、内存、系统总线等,服务器和通用的计算机架构类似,但是由于需要提供高可靠的服务,因此在处理能力、稳定性、可靠性、安全性、可扩展性、可管理性等方面要求较高。
(5)其他具有数据交互功能的电子装置。
对于装置、电子设备实施例而言,由于其基本相似于方法实施例,所以描述的比较简单,相关之处参见方法实施例的部分说明即可。
需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
本领域普通技术人员可以理解实现上述方法实施方式中的全部或部分步骤是可以通过程序来指令相关的硬件来完成,所述的程序可以存储于计算机可读取存储介质中,这里所称得的存储介质,如:ROM/RAM、磁碟、光盘等。
以上所述仅为本公开的较佳实施例而已,并非用于限定本公开的保护范围。凡在本公开的精神和原则之内所作的任何修改、等同替换、改进等,均包含在本公开的保护范围内。

Claims (21)

  1. 一种聚类方法,其特征在于,所述方法包括:
    获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;
    根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;
    对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
  2. 根据权利要求1所述的方法,其特征在于,所述根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,包括:
    所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;
    针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
  3. 根据权利要求1所述的方法,其特征在于,所述根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,包括:
    所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;
    针对包含相同样本数据的个数不小于第二阈值的第一疑似初始化聚类簇和第二疑似初始化聚类簇,计算第一和第二疑似初始化聚类簇的交集的模,并计算第一初始化聚类簇的模与第二初始化聚类簇的模之和与第一常数的商,在所述交集的模与所述商之差不小于第二常数时,将所述第一和第二疑似初始化聚类簇合并为一疑似初始化聚类簇。
  4. 根据权利要求2-3任一所述的方法,其特征在于,所述根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,还包括:
    针对得到的多个疑似初始化聚类簇,将不能与其他疑似初始化聚类簇合并的各疑似初始化聚类簇分别作为一初始化聚类簇。
  5. 根据权利要求2-4任一所述的方法,其特征在于,所述第二阈值根据所述至少二个疑似初始化聚类簇中的一个或者多个疑似初始化聚类簇所包括的样本数据的个数的线性函数确定。
  6. 根据权利要求1-5任一所述的方法,其特征在于,所述方法还包括:对所述多个聚类簇中的至少一个聚类簇进行离群点分离,且所述离群点分离处理后获得的所有聚类簇被作为所述M个样本数据对应的多个聚类簇。
  7. 根据权利要求6所述的方法,其特征在于,所述离群点分离包括:
    针对任一聚类簇,获得所述聚类簇内每一样本数据对应的待离群簇和非待离群簇,其中,每一样本数据对应的所述待离群簇中包括所述样本数据,所述非待离群簇中包括:所述聚类簇中除所述样本数据外的其它样本数据;
    获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度;
    确定所述聚类簇中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值;
    响应于所述最小的簇间相似度小于所述第三阈值,将所述最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇。
  8. 根据权利要求1-7任一所述的方法,其特征在于,所述对所述N个初始化聚类簇进行聚类合并,包括:
    将所述N个初始化聚类簇作为多个待聚类簇;
    获取每个待聚类簇与其他待聚类簇之间的簇间相似度;
    确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值;
    响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇;
    对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇。
  9. 根据权利要求2-8任一所述的方法,其特征在于,所述M个样本数据为M个图像,且两个图像之间的样本间相似度包括:所述两个图像分别对应的两个特征向量之间的余弦距离。
  10. 一种聚类装置,其特征在于,所述装置包括:
    获取单元,用于获取M个样本数据中每两个样本数据之间的样本间相似度,M为正整数;
    合并单元,用于根据获取的样本间相似度将所述M个样本数据合并为N个初始化聚类簇,N为小于M的正整数;
    聚类单元,用于对所述N个初始化聚类簇进行聚类合并,得到所述M个样本数据对应的多个聚类簇。
  11. 根据权利要求10所述的装置,其特征在于,所述合并单元包括:
    第一合并子单元,用于所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;
    第二合并子单元,用于针对得到的多个疑似初始化聚类簇,将包含相同样本数据的个数不小于第二阈值的至少二个疑似初始化聚类簇合并为一初始化聚类簇。
  12. 根据权利要求10所述的装置,其特征在于,所述合并单元具体用于:
    所述M个样本数据中,对于任意两个样本数据而言,如果这两个样本数据的样本间相似度大于第一阈值,则将这两个样本数据合并为一疑似初始化聚类簇;
    针对包含相同样本数据的个数不小于第二阈值的第一疑似初始化聚类簇和第二疑似初始化聚类簇,计算第一和第二疑似初始化聚类簇的交集的模,并计算第一初始化聚类簇的模与第二初始化聚类簇的模之和与第一常数的商,在所述交集的模与所述商之差不小于第二常数时,将所述第一和第二疑似初始化聚类簇合并为一疑似初始化聚类簇。
  13. 根据权利要求11-12任一所述的装置,其特征在于,所述合并单元还包括:
    第一作为子单元,用于针对得到的多个疑似初始化聚类簇,将不能与其他疑似初始化聚类簇合并的各疑似初始化聚类簇分别作为一初始化聚类簇。
  14. 根据权利要求11-13任一所述的装置,其特征在于,所述第二阈值根据所述至少二个疑似初始化聚类簇中的一个或者多个疑似初始化聚类簇所包括的样本数据的个数的线性函数确定。
  15. 根据权利要求10-13任一所述的装置,其特征在于,所述装置还包括:
    分离单元,用于对所述多个聚类簇中的至少一个聚类簇进行离群点分离,且所述离群点分离处理后获得的所有聚类簇被作为所述M个样本数据对应的多个聚类簇。
  16. 根据权利要求15所述的装置,其特征在于,所述分离单元具体用于对一所述聚类簇进行离群点分离;
    所述分离单元包括:
    获得子单元,针对任一聚类簇,用于获得所述聚类簇内每一样本数据对应的待离群簇和非待离群簇,其中,每一样本数据对应的所述待离群簇中包括所述样本数据,所述非待离群簇中包括:所述聚类簇中除所述样本数据外的其它样本数据;
    第一获取子单元,用于获取每一样本数据对应的待离群簇与非待离群簇之间的簇间相似度;
    第一确定子单元,用于确定所述聚类簇中所有样本数据分别对应的多个簇间相似度中最小的簇间相似度是否小于第三阈值;
    第一响应子单元,用于响应于所述最小的簇间相似度小于所述第三阈值,将所述最小的簇间相似度对应的待离群簇和非待离群簇分别作为两个新的聚类簇。
  17. 根据权利要求10-16任一所述的装置,其特征在于,所述聚类单元包括:
    第二作为子单元,用于将所述N个初始化聚类簇作为多个待聚类簇;
    第二获取子单元,用于获取每个待聚类簇与其他待聚类簇之间的簇间相似度;
    第二确定子单元,用于确定所述多个待聚类簇对应的所有簇间相似度中的最大簇间相似度是否大于第四阈值;
    第二响应子单元,用于响应于所述最大簇间相似度大于所述第四阈值,将所述最大簇间相似度对应的两个待聚类簇进行合并得到一新的待聚类簇,并触发所述第二获取子单元,对所述新的待聚类簇与本次未合并的其它待聚类簇构成的新的多个待聚类簇继续进行聚类合并,直到没有可合并的待聚类簇。
  18. 根据权利要求10-17任一所述的装置,其特征在于,所述M个样本数据为M个图像,且两个图像之间的样本间相似度包括:所述两个图像分别对应的两个特征向量之间的余弦距离。
  19. 一种电子设备,其特征在于,包括:壳体、处理器、存储器、电路板和电源电路,其中,所述电路板安置在所述壳体围成的空间内部,所述处理器和所述存储器设置在所述电路板上;所述电源电路,用于为终端的各个电路或器件供电;所述存储器用于存储可执行程序代码;所述处理器通过读取所述存储器中存储的可执行程序代码来运行与可执行程序代码对应的程序,以用于执行权利要求1-9任一项所述的聚类方法对应的操作。
  20. 一种计算机程序,包括计算机可读代码,当所述计算机可读代码在设备中运行时,所述设备中的处理器执行用于实现权利要求1-9中的任一权利要求所述的聚类方法中的步骤的可执行指令。
  21. 一种计算机可读介质,用于存储权利要求20所述的计算机程序。
PCT/CN2017/091432 2016-07-22 2017-07-03 聚类方法、装置及电子设备 Ceased WO2018014717A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US15/859,345 US11080306B2 (en) 2016-07-22 2017-12-30 Method and apparatus and electronic device for clustering

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201610586139.8 2016-07-22
CN201610586139.8A CN106228188B (zh) 2016-07-22 2016-07-22 聚类方法、装置及电子设备

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US15/859,345 Continuation US11080306B2 (en) 2016-07-22 2017-12-30 Method and apparatus and electronic device for clustering

Publications (1)

Publication Number Publication Date
WO2018014717A1 true WO2018014717A1 (zh) 2018-01-25

Family

ID=57531356

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2017/091432 Ceased WO2018014717A1 (zh) 2016-07-22 2017-07-03 聚类方法、装置及电子设备

Country Status (3)

Country Link
US (1) US11080306B2 (zh)
CN (1) CN106228188B (zh)
WO (1) WO2018014717A1 (zh)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110532379A (zh) * 2019-07-08 2019-12-03 广东工业大学 一种基于lstm的用户评论情感分析的电子资讯推荐方法
CN111201524A (zh) * 2018-08-30 2020-05-26 谷歌有限责任公司 百分位链接聚类
CN111915391A (zh) * 2020-06-16 2020-11-10 北京迈格威科技有限公司 商品数据的处理方法、装置及电子设备
CN112784893A (zh) * 2020-12-29 2021-05-11 杭州海康威视数字技术股份有限公司 图像数据的聚类方法、装置、电子设备及存储介质
CN118519666A (zh) * 2024-07-25 2024-08-20 武汉中原电子信息有限公司 一种用于电力采集终端的升级方法及系统

Families Citing this family (30)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106228188B (zh) * 2016-07-22 2020-09-08 北京市商汤科技开发有限公司 聚类方法、装置及电子设备
CN108322428B (zh) * 2017-01-18 2021-11-05 阿里巴巴集团控股有限公司 一种异常访问检测方法及设备
CN108461110B (zh) * 2017-02-21 2021-07-23 阿里巴巴集团控股有限公司 医疗信息处理方法、装置及设备
CN109784354B (zh) * 2017-11-14 2021-07-09 中移(杭州)信息技术有限公司 基于改进分类效用的无参数聚类方法及电子设备
CN107941812B (zh) * 2017-12-20 2021-07-16 联想(北京)有限公司 信息处理方法及电子设备
CN107995518B (zh) * 2017-12-22 2021-02-26 海信视像科技股份有限公司 图像显示方法、装置和计算机存储介质
CN108062576B (zh) * 2018-01-05 2019-05-03 百度在线网络技术(北京)有限公司 用于输出数据的方法和装置
CN108229419B (zh) * 2018-01-22 2022-03-04 百度在线网络技术(北京)有限公司 用于聚类图像的方法和装置
CN108833156B (zh) * 2018-06-08 2022-08-30 中国电力科学研究院有限公司 一种针对电力通信网的仿真性能指标的评估方法及系统
CN108921204B (zh) * 2018-06-14 2023-12-26 平安科技(深圳)有限公司 电子装置、图片样本集生成方法和计算机可读存储介质
CN109272040B (zh) * 2018-09-20 2020-08-14 中国科学院电子学研究所苏州研究院 一种雷达工作模式生成方法
CN109447186A (zh) * 2018-12-13 2019-03-08 深圳云天励飞技术有限公司 聚类方法及相关产品
CN109658572B (zh) * 2018-12-21 2020-09-15 上海商汤智能科技有限公司 图像处理方法及装置、电子设备和存储介质
CN109886300A (zh) * 2019-01-17 2019-06-14 北京奇艺世纪科技有限公司 一种用户聚类方法、装置及设备
CN110191085B (zh) * 2019-04-09 2021-09-10 中国科学院计算机网络信息中心 基于多分类的入侵检测方法、装置及存储介质
CN110046586B (zh) * 2019-04-19 2024-09-27 腾讯科技(深圳)有限公司 一种数据处理方法、设备及存储介质
CN110045371A (zh) * 2019-04-28 2019-07-23 软通智慧科技有限公司 一种鉴定方法、装置、设备及存储介质
CN110414429A (zh) * 2019-07-29 2019-11-05 佳都新太科技股份有限公司 人脸聚类方法、装置、设备和存储介质
CN110232373B (zh) * 2019-08-12 2020-01-03 佳都新太科技股份有限公司 人脸聚类方法、装置、设备和存储介质
CN110730270B (zh) * 2019-09-09 2021-09-14 上海斑马来拉物流科技有限公司 一种短信分组方法、装置及计算机存储介质、电子设备
CN110704708B (zh) * 2019-09-27 2023-04-07 深圳市商汤科技有限公司 数据处理方法、装置、设备和存储介质
CN110838123B (zh) * 2019-11-06 2022-02-11 南京止善智能科技研究院有限公司 一种室内设计效果图像光照高亮区域的分割方法
CN111310834B (zh) * 2020-02-19 2024-05-28 深圳市商汤科技有限公司 数据处理方法及装置、处理器、电子设备、存储介质
CN111340084B (zh) * 2020-02-20 2024-05-17 北京市商汤科技开发有限公司 数据处理方法及装置、处理器、电子设备、存储介质
CN115982634A (zh) * 2021-10-13 2023-04-18 中国移动通信集团江苏有限公司 应用程序分类方法、装置、电子设备及计算机程序产品
CN113936162B (zh) * 2021-10-26 2025-06-10 恒睿(重庆)人工智能技术研究院有限公司 增量聚类方法、装置、计算机设备和存储介质
CN113936161B (zh) * 2021-10-26 2025-04-25 恒睿(重庆)人工智能技术研究院有限公司 基于组属性的增量聚类方法、装置、设备和存储介质
CN114358110A (zh) * 2021-10-28 2022-04-15 腾讯科技(深圳)有限公司 对数据的聚类处理方法、装置、计算机设备和存储介质
CN114662578A (zh) * 2022-03-10 2022-06-24 中国工商银行股份有限公司 均值聚类方法及装置
CN116720094A (zh) * 2023-06-16 2023-09-08 中国工商银行股份有限公司 客户信息的聚类方法、装置、处理器及电子设备

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103679190A (zh) * 2012-09-20 2014-03-26 富士通株式会社 分类装置、分类方法以及电子设备
CN103699653A (zh) * 2013-12-26 2014-04-02 沈阳航空航天大学 数据聚类方法和装置
WO2014067296A1 (zh) * 2012-11-05 2014-05-08 深圳市恩普电子技术有限公司 一种血管内外膜识别、描记和测量的方法
CN104252627A (zh) * 2013-06-28 2014-12-31 广州华多网络科技有限公司 Svm分类器训练样本获取方法、训练方法及其系统
CN104281569A (zh) * 2013-07-01 2015-01-14 富士通株式会社 构建装置和方法、分类装置和方法以及电子设备
CN106228188A (zh) * 2016-07-22 2016-12-14 北京市商汤科技开发有限公司 聚类方法、装置及电子设备

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6941321B2 (en) * 1999-01-26 2005-09-06 Xerox Corporation System and method for identifying similarities among objects in a collection
EP2297703A1 (en) * 2008-06-03 2011-03-23 ETH Zurich Method and system for generating a pictorial reference database using geographical information
CN102915347B (zh) * 2012-09-26 2016-10-12 中国信息安全测评中心 一种分布式数据流聚类方法及系统
CN104123279B (zh) * 2013-04-24 2018-12-07 腾讯科技(深圳)有限公司 关键词的聚类方法和装置
CN104731789A (zh) * 2013-12-18 2015-06-24 北京慧眼智行科技有限公司 一种聚类簇获取方法及装置
US9589045B2 (en) * 2014-04-08 2017-03-07 International Business Machines Corporation Distributed clustering with outlier detection
CN104268149A (zh) * 2014-08-28 2015-01-07 小米科技有限责任公司 聚类方法及装置

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103679190A (zh) * 2012-09-20 2014-03-26 富士通株式会社 分类装置、分类方法以及电子设备
WO2014067296A1 (zh) * 2012-11-05 2014-05-08 深圳市恩普电子技术有限公司 一种血管内外膜识别、描记和测量的方法
CN104252627A (zh) * 2013-06-28 2014-12-31 广州华多网络科技有限公司 Svm分类器训练样本获取方法、训练方法及其系统
CN104281569A (zh) * 2013-07-01 2015-01-14 富士通株式会社 构建装置和方法、分类装置和方法以及电子设备
CN103699653A (zh) * 2013-12-26 2014-04-02 沈阳航空航天大学 数据聚类方法和装置
CN106228188A (zh) * 2016-07-22 2016-12-14 北京市商汤科技开发有限公司 聚类方法、装置及电子设备

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111201524A (zh) * 2018-08-30 2020-05-26 谷歌有限责任公司 百分位链接聚类
CN111201524B (zh) * 2018-08-30 2023-08-25 谷歌有限责任公司 百分位链接聚类
CN110532379A (zh) * 2019-07-08 2019-12-03 广东工业大学 一种基于lstm的用户评论情感分析的电子资讯推荐方法
CN110532379B (zh) * 2019-07-08 2023-01-20 广东工业大学 一种基于lstm的用户评论情感分析的电子资讯推荐方法
CN111915391A (zh) * 2020-06-16 2020-11-10 北京迈格威科技有限公司 商品数据的处理方法、装置及电子设备
CN112784893A (zh) * 2020-12-29 2021-05-11 杭州海康威视数字技术股份有限公司 图像数据的聚类方法、装置、电子设备及存储介质
CN112784893B (zh) * 2020-12-29 2024-03-01 杭州海康威视数字技术股份有限公司 图像数据的聚类方法、装置、电子设备及存储介质
CN118519666A (zh) * 2024-07-25 2024-08-20 武汉中原电子信息有限公司 一种用于电力采集终端的升级方法及系统
CN118519666B (zh) * 2024-07-25 2024-10-01 武汉中原电子信息有限公司 一种用于电力采集终端的升级方法及系统

Also Published As

Publication number Publication date
CN106228188B (zh) 2020-09-08
CN106228188A (zh) 2016-12-14
US11080306B2 (en) 2021-08-03
US20180129727A1 (en) 2018-05-10

Similar Documents

Publication Publication Date Title
US11080306B2 (en) Method and apparatus and electronic device for clustering
Doan et al. One loss for quantization: Deep hashing with discrete wasserstein distributional matching
US10607062B2 (en) Grouping and ranking images based on facial recognition data
WO2017045443A1 (zh) 一种图像检索方法及系统
US9148619B2 (en) Music soundtrack recommendation engine for videos
CN110334356B (zh) 文章质量的确定方法、文章筛选方法、以及相应的装置
US20150039583A1 (en) Method and system for searching images
CN110909222B (zh) 基于聚类的用户画像建立方法、装置、介质及电子设备
WO2017101506A1 (zh) 信息处理方法及装置
CN113553386B (zh) 嵌入表示模型训练方法、基于知识图谱的问答方法及装置
CN116310994B (zh) 一种视频片段提取方法、装置、电子设备及介质
US11868358B1 (en) Contextualized novelty for personalized discovery
CN112784102A (zh) 视频检索方法、装置和电子设备
US20190394156A1 (en) Methods, servers, and non-transitory computer readable record media for converting image to location data
US9734434B2 (en) Feature interpolation
CN116010694A (zh) 搜索筛选项展示方法、装置、电子设备及存储介质
KR101931859B1 (ko) 전자문서의 대표 단어 선정 방법, 전자 문서 제공 방법, 및 이를 수행하는 컴퓨팅 시스템
CN111858966A (zh) 知识图谱的更新方法、装置、终端设备及可读存储介质
WO2026026718A1 (zh) 内容搜索方法、装置、设备和存储介质
WO2019120024A1 (zh) 用户性别识别方法、装置、存储介质及电子设备
CN111738009A (zh) 实体词标签生成方法、装置、计算机设备和可读存储介质
CN114298123A (zh) 聚类处理方法、装置、电子设备及可读存储介质
CN115392389B (zh) 跨模态信息匹配、处理方法、装置、电子设备及存储介质
EP4738078A1 (en) Method and apparatus for searching, device, and storage medium
CN113849666B (zh) 一种数据处理、模型训练方法、装置、设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17830351

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17830351

Country of ref document: EP

Kind code of ref document: A1