WO2012109959A1 - 检索词的聚类方法和装置 - Google Patents
检索词的聚类方法和装置 Download PDFInfo
- Publication number
- WO2012109959A1 WO2012109959A1 PCT/CN2012/070824 CN2012070824W WO2012109959A1 WO 2012109959 A1 WO2012109959 A1 WO 2012109959A1 CN 2012070824 W CN2012070824 W CN 2012070824W WO 2012109959 A1 WO2012109959 A1 WO 2012109959A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- search term
- search
- clustering
- candidate
- term
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/35—Clustering; Classification
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/02—Marketing; Price estimation or determination; Fundraising
- G06Q30/0241—Advertisements
- G06Q30/0251—Targeted advertisements
- G06Q30/0255—Targeted advertisements based on user history
- G06Q30/0256—User search
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/3332—Query translation
- G06F16/3338—Query expansion
Definitions
- the invention relates to a network search technology, in particular to a clustering method and device for searching words. Background of the invention
- search term may be specifically identified as an advertisement advertisement provided by the advertiser, and may also be referred to as a purchase word, so as to facilitate the user to search for the corresponding advertisement through the search term.
- the process of clustering search terms can be abstracted into a process of clustering a set of short text strings.
- the most commonly used clustering method is: For a search term provided by an advertiser, only the term, the search term provided by the advertiser and the found search term are clustered together. Thus, when the search engine user retrieves the corresponding advertisement through a search term, the advertisement corresponding to the search term and the advertisement corresponding to the search term clustered by the search term are displayed to the user.
- search terms although the advertiser does not provide it, but it is substantially related to the advertisement corresponding to the search term provided by the advertiser, and the aforementioned clustering method is to perform only the word-related clustering of the search terms provided by the advertiser. , which does not take into account these other search terms related to the search term provided by the advertiser and which have not yet been provided by the advertiser, which reduces the accuracy of clustering of the search terms.
- the present invention provides a clustering method and apparatus for search terms to improve the accuracy and relevance of clustering of search terms.
- a clustering method for search terms including:
- a clustering operation is performed on the first search term in the candidate search term set and the second search term associated with the first search term based on text features and/or semantic features of the search term.
- a clustering device for searching words comprising:
- a establishing unit configured to establish a candidate search term set, where the candidate search term set includes a first search term provided by a user, and a second search term related to the first search term;
- a clustering unit configured to perform a clustering operation on the first search term in the set of candidate search words and the second search term related to the first search term according to text features and/or semantic features of the search term.
- the clustering method and apparatus for searching for a search term provided by the present invention does not cluster the search terms provided by the user only by the user, as in the prior art. Considering simultaneously the search term provided by the user, and other search terms related to the search term provided by the user, and the search term provided to the user according to the text feature and/or semantic feature of the search term, and the search term provided by the user The related other search terms are clustered, thereby increasing the accuracy and relevance of the search term clustering.
- FIG. 1 is a basic flowchart of an embodiment of the present invention
- FIG. 2a is a flowchart of step 102 according to an embodiment of the present invention
- FIG. 2b is a flowchart of potential clustering relationship mining according to an embodiment of the present invention
- FIG. 3 is a first schematic diagram of a topology structure between search terms according to an embodiment of the present invention
- FIG. 3b is a second schematic diagram of a topology structure between search terms according to an embodiment of the present invention
- FIG. 3 is a third schematic diagram of a topology structure when a search term is added according to an embodiment of the present invention
- FIG. 4 is a schematic diagram of a new search term provided by an embodiment of the present invention
- FIG. 5 is a basic structural diagram of an apparatus according to an embodiment of the present invention.
- FIG. 6 is a detailed structural diagram of an apparatus according to an embodiment of the present invention. Mode for carrying out the invention
- the present invention does not cluster only the search terms provided by the user, such as an advertiser, as in the prior art, but provides the user with the text features and/or semantic features of the search terms.
- the search term, and the search term cluster associated with the search term are added to increase the accuracy of clustering of the search terms.
- the method provided by the present invention is described below.
- FIG. 1 is a basic flowchart of an embodiment of the present invention. As shown in Figure 1, the process can include the following steps:
- Step 101 Establish a candidate search term set, where the candidate search term set includes a first search term provided by a user, and a second search term related to the first search term.
- the second search term related to the first search term provided by the user may specifically include the following two methods or any one of the following manners: mode 1, determining a search term that matches the first search term provided by the user, Determining the determined search term as a second search term related to the first search term provided by the user; mode 2, searching for the first search term provided by the user as a keyword, and determining the search term in the search result as The second term related to the first search term provided by the user Search terms.
- the search term obtained by the method 1 may be: a search term obtained by performing a string conversion process on the first search term provided by the user; or a first search term determined according to actual experience.
- the search term used together for example, if the first search term provided by the user is a coffee maker, it is known from experience that the coffee maker is usually used frequently with a coffee cup or the like, based on which it can be determined that the coffee maker provided by the user matches.
- the search term can be a coffee cup or the like.
- the search term obtained by the method 2 may specifically be: searching for the first search term provided by the user as a keyword, and the search term in the obtained search result.
- the search may be specifically implemented by a user search string and a search term mapping integration system (QBM: Query Bidterm Merge), wherein the QBM implementation may be: searching with the first search term provided by the user as input, from the search.
- QBM search term mapping integration system
- the QBM implementation may be: searching with the first search term provided by the user as input, from the search
- the search term is obtained from the search result, and the obtained search term is used as a search term related to the first search term provided by the user.
- the candidate search term set can be obtained through step 101. It should be noted that, in this embodiment, it is necessary to ensure that there are no duplicate search terms in the candidate search term set obtained in step 101.
- Step 102 Perform a clustering operation on the first search term in the candidate search term set and the second search term related to the first search term according to the text feature and/or the semantic feature of the search term.
- the first search term and the second search term related to the first search term in the candidate search term set may be calculated according to the text feature and/or the semantic feature of the first search term.
- the similarity value, the first search term and the second search term having a higher similarity value with the first search term are clustered together.
- the step 102 can be embodied by the flow shown in Figure 2a.
- FIG. 2a is a flowchart of step 102 according to an embodiment of the present invention.
- the flow shows the specific implementation principle of the basic clustering relationship, as shown in FIG. 2a, the process may include Next steps:
- Step 201a Calculate a similarity value between the first search term and each of the related second search terms according to the text feature and/or the semantic feature of the first search term.
- Step 202a If the similarity value between the first search term and the second search term is greater than or equal to the first preset threshold, the first search term and the second search term are clustered together.
- step 202a the first search term and the second search term associated with the first search term having a similarity value greater than or equal to the first preset threshold may be clustered together, that is, the present search word is implemented.
- the present search word is implemented.
- the embodiment further provides a mining process of the potential clustering relationship, which can be embodied by the process shown in FIG. 2b.
- FIG. 2b is a flowchart of potential clustering relationship mining according to an embodiment of the present invention. As shown in Figure 2b, the process can include the following steps:
- Step 201b Select, from each of the second search terms related to the first search term, a second search term whose similarity value with the first search term is greater than or equal to a second preset threshold.
- the step 201b may be further replaced by: selecting and selecting the second search term from the first search term.
- a second search term whose similarity value between the search terms is greater than or equal to the second predetermined threshold.
- the second preset threshold in the step 201b is independent of the first preset threshold in step 202a, and the two may be equal or unequal.
- Step 202b calculating a similarity value between the selected two second search terms, and if the calculated similarity value is greater than or equal to the first preset threshold, clustering the two second search terms Together.
- the first search term and the second clustered together in step 202a are used in the embodiment of the present invention.
- a search term ie, a cluster relationship between the first search term and the second search term
- the second search terms clustered together in 202b are combined to form a full-scale clustering result of the embodiment of the present invention.
- the clustering of step 202a and the clustering of step 202b may be implemented according to a similar machine learning model, which is not specifically limited herein.
- the first search terms provided by the user are bl, b3, b4 and b5 respectively, wherein, by step 101, it can be obtained that: the second search terms related to bl are b2, b3 and b4, and the second search term related to b3 For b5, b6 and b4, the second search terms associated with b4 are b7, b8 and b9, and the second search term associated with b5 is b3. All search terms are represented by the graph data structure shown in Figure 3a. Referring to FIG. 3a, FIG. 3a is a first schematic diagram of a topology structure between search terms according to an embodiment of the present invention. In Fig.
- each search term is taken as an arrow of node bi (i takes a value of 1 to 9), from node bi to node bj (j takes a value of 1 to 9), indicating that bi can be expanded to bj, that is, , the related term with bi is bj.
- the topology shown in Fig. 3a is a directed acyclic graph, that is to say, the correlation between the two search terms is not guaranteed to be bidirectional, specifically:
- the search term related to bi is the search term bj, but the search term bj does not necessarily extend the search term related to the search term bj as the search term bio
- step 201a it can be obtained that: for bl, the similarity value wl2 between bl and b2, the similarity values wl3, bl and b4 between bl and b3 are calculated according to the text feature and/or semantic feature of bl.
- the similarity value wl4; for b3, the similarity value wl4 between b3 and b4, the similarity value between b3 and b5, the similarity between b3 and b6 is calculated according to the text feature and/or semantic feature of b3 Degree value w36; for b4, calculate the similarity value w47 between b4 and b7 according to the text feature and/or semantic feature of b4, the similarity value w48 between b4 and b8, the similarity value w49 between b4 and b9 ; for b5, according to the text characteristics of b5 and / or semantic feature calculates the similarity value w53 between b5 and b3.
- FIG. 3b is a second schematic diagram of a topological structure between search terms provided by an embodiment of the present invention.
- Figure 3b shows the clustering relationship between the interconnected search terms, wherein the two search terms connected by the solid line indicate that the two search terms have a cluster relationship: the two are considered equivalent and can be clustered Together; the two search terms connected by the dotted line have the clustering relationship: The two are not equivalent, and cannot be clustered together, and the dotted line can be removed later.
- the second search terms with bl are: b2, b3, and b4, so, based on step 201b, when the similarity values between b2, b3, and b4 and bl are greater than or equal to
- the present invention can supplement three potential clustering relationships: clustering relationship between b2 and b3, clustering relationship between b2 and b4, and clustering between b3 and b4 relationship.
- the clustering relationship between b3 and b4 has been determined in the above step 202a.
- the present invention may omit the operation of determining the clustering relationship between b3 and b4, It is necessary to increase the clustering relationship between b2 and b3 and the clustering relationship between b2 and b4. Then calculate the similarity value between b2 and b3, and the similarity value between b2 and b4, determine whether the cluster relationship between b2 and b3 and the cluster relationship between b2 and b4 meet the criteria of clustering.
- the clustering relationship between the search terms is represented by a solid line (also referred to as an edge relationship) between the search terms. Therefore, the embodiment of the present invention can only traverse the edge relationship, so that the present invention can be made.
- the complexity of the embodiment is reduced to 0 (n + e), where n represents the number of search terms and e represents the number of edge relationships.
- the second search term related to the first search term provided by the user in FIG. 3a may be further mined, and the second search term is N (for example, N is 3)
- N for example, N is 3
- the potential clustering relationship between the "children" nodes within the hop For the specific implementation, refer to the process shown in Figure 2b, which will not be described in detail here.
- the set of candidate search terms is not fixed, and the search terms can be incremented over time.
- the candidate search term set newly adds the first search term provided by the user, and the newly added first search term is newly appearing relative to all previous search terms.
- the newly added first search term it is also necessary to perform a clustering operation similar to that shown in Fig. 2a and Fig. 2b, and at the same time, integrate the result obtained after performing the clustering operation with the previous clustering result. See the process shown in Figure 4.
- FIG. 4 is a flowchart of a process of adding a first search term (indicated as an incremental update process) according to an embodiment of the present invention.
- the process may include the following steps: Step 401: Determine a second search term related to the added first search term, and compare the added first search term with the determined second search term related to the added first search term. A second search term different from any one of the candidate search term sets is added to the candidate search term set.
- the search terms stored in the candidate search term set before the execution of step 401 are bl to b9 shown in Fig. 3a, and when executed to this step 401, if the following two first search words are newly added: nl and n2.
- the second search terms related to nl are b5 and b6, and the second search terms related to n2 are bl, b2, b3, b4, b8, and n3, as shown in FIG. 3d. Since b5 and b6 associated with nl, and bl, b2, b3, b4, b8 associated with n2 are already stored in the set of candidate search terms, this step 401 can only refer to nl, n2, and n2. N3 is added to the set of candidate terms.
- Step 402 Perform a clustering operation on the newly added first search term in the candidate search term set and the second search term related to the first search term according to the text feature and/or the semantic feature of the search term.
- step 402 is described by taking the newly added first search term as nl as an example, and the added other search terms are similar in principle.
- step 402 when performing this step 402, based on the flow shown in FIG. 2a, the similarity value between n1 and b5 is calculated according to the text feature and/or the semantic feature of n1, and the similarity value between n1 and b6 is calculated. Then, it is determined whether the similarity value between nl and b5 is greater than or equal to the first preset threshold, and if so, it is determined that nl and b5 are equivalent, and the two can be clustered together, otherwise, nl and b5 are not clustered. Together. The same operation is performed for the similarity value between nl and b6.
- Step 403 Perform mining of a potential clustering relationship on the second search term related to the added first search term in the candidate search term set.
- the process of the potential clustering relationship may be performed by using the process shown in FIG. 2b, and the bill is described as: each second search term related to the added first search term from the candidate search term set, or increased from And selecting, from each of the second search words, the first search term, the second search term having a similarity value with the first search term greater than or equal to a second preset threshold; calculating any two selected And a similarity value between the two search terms, if the calculated similarity value is greater than or equal to the first preset threshold, the two second search terms are clustered together.
- the search term nl As an example, since it is determined in step 401 that the second search term related to the n1 is b5 and b6, when performing this step 403, if b5 and b6 respectively If the similarity value between n1 and n1 is greater than the second preset threshold, the similarity value between b5 and b6 may be calculated, and if the calculated similarity value is greater than or equal to the first preset threshold, the two are The search terms b5 and b6 are clustered together, otherwise b5 and b6 are not clustered together.
- Incremental clustering results the cluster relationship between the newly added first search term (indicated as an incremental search term) and the original existing search term (recorded as the old search term) is realized (hereinafter referred to as Incremental clustering results).
- the incremental clustering result and the pre-existing full-quantity clustering result are collectively referred to as the final clustering result of the present invention.
- the second search term related to the first search term is not fixed, and the search term is also changed according to the user, and the method provided by the embodiment of the present invention should also be able to Reflect this change.
- the change is implemented by periodically updating the set of candidate search terms (referred to as full update), and is specifically implemented as: when the set full amount update time arrives, determining, for the first search term in the candidate search term set, the first Searching for a second search term related to the word, placing the first search term and the determined second search term related to the first search term into a new candidate search term set, and then according to FIG. 2a and FIG. 2
- the illustrated process clusters the search terms in the new candidate search term set to obtain a full-scale clustering result. This Can be described by the image of Table 1.
- the corresponding QBM extension result of the first search term is Q ⁇ i B, and the extended result is mainly a set of second search terms related to the first search term.
- the clustering result obtained by clustering the first search term and the second search term based on the flow shown in Fig. 2a and Fig. 2b is: CF CXB);
- the full update starts on the ith day, the end of the kth day, and on the k+1th (ie, L) day, the synchronous operation of the full amount of data and the incremental data is performed, that is, the kth All of the first search terms in the +1 (i.e., L) day candidate search term set perform the flow shown in FIG.
- the device provided by the embodiment of the present invention is described below.
- FIG. 5 is a basic structural diagram of an apparatus according to an embodiment of the present invention. As shown in Figure 5, the device can include:
- the establishing unit 501 is configured to establish a candidate search term set, where the candidate search term set includes a first search term provided by the user, and a second search term related to the first search term; and a clustering unit 502, configured to perform the search according to the search
- the textual feature and/or semantic feature of the word performs a clustering operation on the first search term in the set of candidate search terms and the second search term associated with the first search term.
- the device shown in FIG. 5 can be specifically seen in FIG. 6.
- FIG. 6 is a detailed structural diagram of an apparatus according to an embodiment of the present invention.
- the apparatus may include an establishing unit 601 and a clustering unit 602, wherein the establishing unit 601 and the clustering unit 602 have functions similar to the establishing unit 501 and the clustering unit 502 shown in FIG. 5, respectively. No longer.
- the apparatus may further include:
- An adding unit 603, configured to: when the user adds a new first search term, determine a second search term related to the added first search term, and add the added first search term and the determined a second search term different from any one of the candidate search term sets in the second search term related to the first search term is added to the candidate search term set;
- the clustering unit 602 is further configured to perform, according to the text feature and/or the semantic feature of the search term, the newly added first search term in the candidate search term set and the second search term related to the first search term. Clustering operation.
- the apparatus further includes:
- the updating unit 604 is configured to: when the set full amount update time arrives, determine, for the first search term in the candidate search term set, a second search term related to the first search term, the first search term And determining the second search term associated with the first search term into one A new set of candidate search terms.
- the clustering unit 602 is further configured to perform clustering on the first search term in the new candidate search term set and the second search term related to the first search term according to the text feature and/or the semantic feature of the search term. operating.
- the clustering unit 602 performs a clustering operation through the following subunits:
- a calculating subunit 6021 configured to calculate, according to a text feature and/or a semantic feature of the first search term, a similarity value between the first search term and each second search term related to the first search term;
- the clustering sub-unit 6022 is configured to cluster the first search term and the second search term when the similarity value between the first search term and the second search term is greater than or equal to the first preset threshold .
- the clustering sub-unit 6022 is further configured to select the first search term from each of the second search terms related to the first search term, or from each of the second search terms clustered with the first search term. a second search term having a similarity value greater than or equal to a second predetermined threshold; and calculating a similarity value between the selected two second search terms, if the calculated similarity value is greater than or equal to The first preset threshold is clustered, and the two second search terms are clustered together, and the first preset threshold is independent of the second preset threshold.
- the clustering method and apparatus for searching for a search term provided by the present invention does not cluster the search terms provided by the user only by the user, as in the prior art. Considering simultaneously the search term provided by the user, and other search terms related to the search term provided by the user, and the search term provided to the user according to the text feature and/or semantic feature of the search term, and the search term provided by the user Clustering related related terms, which obviously increases the accuracy of clustering of search terms;
- the present invention also excavates various first related to the first search term provided by the user.
- the clustering relationship between the two search terms compared with the prior art, the clustering relationship between the search terms can be deeply explored, and the clustering of the search words is more accurate.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Business, Economics & Management (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Accounting & Taxation (AREA)
- Development Economics (AREA)
- Finance (AREA)
- Strategic Management (AREA)
- General Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Entrepreneurship & Innovation (AREA)
- General Business, Economics & Management (AREA)
- Marketing (AREA)
- Economics (AREA)
- Game Theory and Decision Science (AREA)
- Computational Linguistics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US14/000,083 US20140019452A1 (en) | 2011-02-18 | 2012-02-01 | Method and apparatus for clustering search terms |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201110043030.7A CN102646103B (zh) | 2011-02-18 | 2011-02-18 | 检索词的聚类方法和装置 |
| CN201110043030.7 | 2011-02-18 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2012109959A1 true WO2012109959A1 (zh) | 2012-08-23 |
Family
ID=46658926
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2012/070824 Ceased WO2012109959A1 (zh) | 2011-02-18 | 2012-02-01 | 检索词的聚类方法和装置 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20140019452A1 (zh) |
| CN (1) | CN102646103B (zh) |
| WO (1) | WO2012109959A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103744889A (zh) * | 2013-12-23 | 2014-04-23 | 百度在线网络技术(北京)有限公司 | 一种用于对问题进行聚类处理的方法与装置 |
| WO2015016908A1 (en) | 2013-07-30 | 2015-02-05 | Intuit Inc. | Method and system for clustering similar items |
Families Citing this family (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103699550B (zh) * | 2012-09-27 | 2017-12-12 | 腾讯科技(深圳)有限公司 | 数据挖掘系统及数据挖掘方法 |
| CN103853722B (zh) * | 2012-11-29 | 2017-09-22 | 腾讯科技(深圳)有限公司 | 一种基于检索串的关键词扩展方法、装置和系统 |
| CN104123279B (zh) * | 2013-04-24 | 2018-12-07 | 腾讯科技(深圳)有限公司 | 关键词的聚类方法和装置 |
| CN104933081B (zh) | 2014-03-21 | 2018-06-29 | 阿里巴巴集团控股有限公司 | 一种搜索建议提供方法及装置 |
| TW201619853A (zh) * | 2014-11-21 | 2016-06-01 | 財團法人資訊工業策進會 | 檢索過濾方法及其處理裝置 |
| CN104462272B (zh) * | 2014-11-25 | 2018-05-04 | 百度在线网络技术(北京)有限公司 | 搜索需求分析方法和装置 |
| CN106326259A (zh) * | 2015-06-26 | 2017-01-11 | 苏宁云商集团股份有限公司 | 搜索引擎中商品标签的构建方法、系统及搜索方法和系统 |
| CN106610989B (zh) * | 2015-10-22 | 2021-06-01 | 北京国双科技有限公司 | 搜索关键词聚类方法及装置 |
| CN106951511A (zh) * | 2017-03-17 | 2017-07-14 | 福建中金在线信息科技有限公司 | 一种文本聚类方法及装置 |
| US11409799B2 (en) * | 2017-12-13 | 2022-08-09 | Roblox Corporation | Recommendation of search suggestions |
| CN111259058B (zh) * | 2020-01-16 | 2023-09-15 | 北京百度网讯科技有限公司 | 数据挖掘方法、数据挖掘装置和电子设备 |
| CN112650907B (zh) * | 2020-12-25 | 2023-07-14 | 百度在线网络技术(北京)有限公司 | 搜索词的推荐方法、目标模型的训练方法、装置及设备 |
| CN112905765B (zh) * | 2021-02-09 | 2024-06-18 | 联想(北京)有限公司 | 一种信息处理方法及装置 |
| CN115376054B (zh) * | 2022-10-26 | 2023-03-24 | 浪潮电子信息产业股份有限公司 | 一种目标检测方法、装置、设备及存储介质 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100131563A1 (en) * | 2008-11-25 | 2010-05-27 | Hongfeng Yin | System and methods for automatic clustering of ranked and categorized search objects |
| KR20100106718A (ko) * | 2009-03-24 | 2010-10-04 | 엔에이치엔(주) | 연관 키워드에 따른 클러스터를 이용하여 검색 키워드를 분류하는 시스템 및 방법 |
Family Cites Families (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5488725A (en) * | 1991-10-08 | 1996-01-30 | West Publishing Company | System of document representation retrieval by successive iterated probability sampling |
| US5931907A (en) * | 1996-01-23 | 1999-08-03 | British Telecommunications Public Limited Company | Software agent for comparing locally accessible keywords with meta-information and having pointers associated with distributed information |
| US6502091B1 (en) * | 2000-02-23 | 2002-12-31 | Hewlett-Packard Company | Apparatus and method for discovering context groups and document categories by mining usage logs |
| EP1182581B1 (en) * | 2000-08-18 | 2005-01-26 | Exalead | Searching tool and process for unified search using categories and keywords |
| KR20020049164A (ko) * | 2000-12-19 | 2002-06-26 | 오길록 | 유전자 알고리즘을 이용한 카테고리 학습과 단어클러스터에 의한 문서 자동 분류 시스템 및 그 방법 |
| US20030120630A1 (en) * | 2001-12-20 | 2003-06-26 | Daniel Tunkelang | Method and system for similarity search and clustering |
| US6947930B2 (en) * | 2003-03-21 | 2005-09-20 | Overture Services, Inc. | Systems and methods for interactive search query refinement |
| US7428529B2 (en) * | 2004-04-15 | 2008-09-23 | Microsoft Corporation | Term suggestion for multi-sense query |
| US7260568B2 (en) * | 2004-04-15 | 2007-08-21 | Microsoft Corporation | Verifying relevance between keywords and web site contents |
| US7689585B2 (en) * | 2004-04-15 | 2010-03-30 | Microsoft Corporation | Reinforced clustering of multi-type data objects for search term suggestion |
| US7756855B2 (en) * | 2006-10-11 | 2010-07-13 | Collarity, Inc. | Search phrase refinement by search term replacement |
| US7792858B2 (en) * | 2005-12-21 | 2010-09-07 | Ebay Inc. | Computer-implemented method and system for combining keywords into logical clusters that share similar behavior with respect to a considered dimension |
| US8799285B1 (en) * | 2007-08-02 | 2014-08-05 | Google Inc. | Automatic advertising campaign structure suggestion |
| US7962486B2 (en) * | 2008-01-10 | 2011-06-14 | International Business Machines Corporation | Method and system for discovery and modification of data cluster and synonyms |
| US20100094673A1 (en) * | 2008-10-14 | 2010-04-15 | Ebay Inc. | Computer-implemented method and system for keyword bidding |
| US8463783B1 (en) * | 2009-07-06 | 2013-06-11 | Google Inc. | Advertisement selection data clustering |
| US9002857B2 (en) * | 2009-08-13 | 2015-04-07 | Charite-Universitatsmedizin Berlin | Methods for searching with semantic similarity scores in one or more ontologies |
| US20110295678A1 (en) * | 2010-05-28 | 2011-12-01 | Google Inc. | Expanding Ad Group Themes Using Aggregated Sequential Search Queries |
| US9830379B2 (en) * | 2010-11-29 | 2017-11-28 | Google Inc. | Name disambiguation using context terms |
-
2011
- 2011-02-18 CN CN201110043030.7A patent/CN102646103B/zh active Active
-
2012
- 2012-02-01 US US14/000,083 patent/US20140019452A1/en not_active Abandoned
- 2012-02-01 WO PCT/CN2012/070824 patent/WO2012109959A1/zh not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100131563A1 (en) * | 2008-11-25 | 2010-05-27 | Hongfeng Yin | System and methods for automatic clustering of ranked and categorized search objects |
| KR20100106718A (ko) * | 2009-03-24 | 2010-10-04 | 엔에이치엔(주) | 연관 키워드에 따른 클러스터를 이용하여 검색 키워드를 분류하는 시스템 및 방법 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2015016908A1 (en) | 2013-07-30 | 2015-02-05 | Intuit Inc. | Method and system for clustering similar items |
| EP3031024A4 (en) * | 2013-07-30 | 2017-01-11 | Intuit Inc. | Method and system for clustering similar items |
| CN103744889A (zh) * | 2013-12-23 | 2014-04-23 | 百度在线网络技术(北京)有限公司 | 一种用于对问题进行聚类处理的方法与装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN102646103B (zh) | 2016-03-16 |
| CN102646103A (zh) | 2012-08-22 |
| US20140019452A1 (en) | 2014-01-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2012109959A1 (zh) | 检索词的聚类方法和装置 | |
| CN103187052B (zh) | 一种建立用于语音识别的语言模型的方法及装置 | |
| CN109523991B (zh) | 语音识别的方法及装置、设备 | |
| US10698654B2 (en) | Ranking and boosting relevant distributable digital assistant operations | |
| EP2518978A2 (en) | Context-Aware Mobile Search Based on User Activities | |
| US11264006B2 (en) | Voice synthesis method, device and apparatus, as well as non-volatile storage medium | |
| CN105224554A (zh) | 推荐搜索词进行搜索的方法、系统、服务器和智能终端 | |
| CN103699689A (zh) | 事件知识库的构建方法及装置 | |
| CN110019348A (zh) | 扩展搜索查询 | |
| CN105900081A (zh) | 基于自然语言处理的搜索 | |
| CN104216906A (zh) | 语音搜索方法和设备 | |
| CN103383699A (zh) | 字符串检索方法及系统 | |
| CN106777328B (zh) | 一种移动终端的题目推荐方法及装置 | |
| WO2017054332A1 (zh) | 路径查询方法、装置、设备及非易失性计算机存储介质 | |
| CN102955821A (zh) | 一种对查询序列进行扩展处理的方法与设备 | |
| CN104036018A (zh) | 视频获取方法和装置 | |
| CN105930376A (zh) | 一种搜索方法和装置 | |
| CN110232129A (zh) | 场景纠错方法、装置、设备和存储介质 | |
| CN102063194A (zh) | 用于供用户进行文字输入的方法、设备、服务器和系统 | |
| CN103970756A (zh) | 热点话题提取方法、装置和服务器 | |
| CN103092928A (zh) | 语音查询方法及系统 | |
| CN110246493A (zh) | 通讯录联系人查找方法、装置及存储介质 | |
| US20170228402A1 (en) | Inconsistency Detection And Correction System | |
| CN106682190A (zh) | 标签知识库的构建方法、装置、应用搜索方法和服务器 | |
| CN103064825B (zh) | 模糊音对建立、设置方法和输入法及其装置和系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 12747612 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 14000083 Country of ref document: US |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205N DATED 15/01/2014) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 12747612 Country of ref document: EP Kind code of ref document: A1 |

