WO2012116562A1 - 一种跨语言搜索的方法和装置 - Google Patents
一种跨语言搜索的方法和装置 Download PDFInfo
- Publication number
- WO2012116562A1 WO2012116562A1 PCT/CN2011/083420 CN2011083420W WO2012116562A1 WO 2012116562 A1 WO2012116562 A1 WO 2012116562A1 CN 2011083420 W CN2011083420 W CN 2011083420W WO 2012116562 A1 WO2012116562 A1 WO 2012116562A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- query
- translation
- source language
- search
- result
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/3332—Query translation
- G06F16/3337—Translation of the query language, e.g. Chinese to English
Definitions
- the feature vector of each language is trained in the excavation of existing resources of each language.
- the result integration unit is configured to integrate the search results obtained by the search processing unit to form a final search result set; wherein, in the final search result set, according to the row of each search result in the belonging category The ranking weight of the sub-category and the sub-category, and sorts each search result.
- the translation processing unit for each target language, the translation result of the target language corresponding to the source language query, and the translation result having the highest translation score as the target language query;
- E translation results translated at least one score is determined by the following factors: Translation translation corpus used in the translation result of the combination of the statistical probability of the number of e e and the translation of each word.
- the device further comprises: an optimization processing unit;
- the optimization processing unit is configured to perform optimization processing on the source language query input by the user, and then provide the translation processing unit, where the optimization processing includes any one or combination of a query error correction process and a query extension process;
- the optimization processing unit performs a query error correction process on the source language query input by the user to obtain nl queries.
- the source language query collection Ql, nl is a preset positive integer;
- the translation processing unit performs translation for each of the target languages by using each query in the Q1, and determines a translation result with the highest total translation score as the target language query; wherein the translation result sum of the translation results is ⁇ ( ⁇ ) , ⁇ ( ⁇ is translated as e in Ql
- the optimization processing unit performs a query extension process on the source language query input by the user to obtain a source language query set Q2 including n2 queries, where n2 is a preset positive integer. ;
- E translation result corresponding to the value of at least one translation determined by the following factors: the translation used in the corpus translation translation combined probability statistics and the number of e e translation of each word.
- the optimization processing unit performs a query error correction process and a query extension process on the source language query input by the user to obtain a source language including n queries.
- the query set Q, n is a preset positive integer;
- the translation processing unit performs translation for each of the target languages by using each query in the Q to determine a translation result with the highest total translation score as the target language query; wherein the translation result sum of the translation results is Q Translated as e Translation score
- the translation score of the translation result e is determined by at least one of the following factors: the number of statistics of the translation result e in the translation corpus used for translation and the combination probability of the words in the translation result e.
- the optimization processing unit may include: a first error correction module and a first expansion module;
- the first error correction module is configured to perform a query error correction process on the source language query input by the user to obtain a source language query set Q1, where nl is a preset positive integer;
- the first expansion module is configured to perform query expansion processing on each query in the Q1 to obtain a source language query set Q including n queries.
- the second error correction module is configured to perform query error correction processing on each query in the Q2 to obtain a source language query set Q including n queries.
- the optimization processing unit specifically includes: a third error correction module, a third expansion module, and a merge processing module;
- the merge processing module is configured to obtain the source language query set Q including n queries after the Q1 and Q2 are combined. Training the corpus, determining whether there is an error query in the error correction training corpus that is the same as the source language query input by the user, and if so, determining all correct queries corresponding to the same error query as the source language query input by the user, Selecting the correct query of the first nl corresponding error correction probability from the determined correct queries constitutes the source language query set Q1; otherwise, the Q1 includes only the source language query input by the user;
- the error correction training corpus includes: an error query collected in advance from the search log and a query pair corresponding to the correct query, and the error query is corrected to the error correction probability corresponding to the correct query.
- the rehearsal resources of the source language By finding the rehearsal resources of the source language, the synonym of each word obtained by the word segmentation is determined, and the words obtained by the word segmentation and the synonym of each word are combined, and the extended scores of the combined queries are ranked in the first n2
- the query constitutes the Q2;
- the device further includes: selecting a query in the source language query set;
- the search processing unit is further configured to acquire a search result corresponding to the query selected by the source request selection unit.
- the result integration unit includes: a merge processing module, a deduplication processing module, and a sort processing module;
- the feature vector of each language is trained by excavating the existing resources of each language in advance.
- Embodiment 1 is a flowchart of a method according to Embodiment 1 of the present invention.
- Embodiment 2 is a schematic diagram of Embodiment 2 of the present invention.
- FIG. 3 is a structural diagram of a device according to Embodiment 3 of the present invention.
- FIG. 4 are three structural diagrams of an optimization processing unit provided in the third embodiment of the present invention.
- FIG. 1 is a flowchart of a method according to Embodiment 1 of the present invention. As shown in FIG. 1, the method may include the following steps:
- Step 101 Receive the source language query entered by the user.
- the method can be implemented on the server side of the search engine. After the user inputs the source language query, the browser sends the source language query input by the user to the server side of the search engine.
- Step 102 Optimize the source language query.
- This step is an optional step, and the optimization process of the source language query may include: any one or combination of query error correction processing and query expansion processing.
- the purpose of the query error correction processing is to improve the possibility that the source language query is correctly translated.
- the purpose of the query extension processing is to expand the resource recall rate of the query.
- the query error correction process is designed to correct errors in the query, mainly spelling errors, which can be implemented based on the noise channel model.
- the method includes: collecting a large number of error queries from the search log and a query pair corresponding to the correct query, and calculating a probability that the error query is corrected to the correct query, that is, the error correction probability ⁇ ( ⁇
- p) is the error correction probability that the source language query p is corrected to p'
- ⁇ ') is the probability that ⁇ ' is miswritten into ⁇ in the search
- ⁇ ( ⁇ ') is The number of times ⁇ ' appears in the search log
- ⁇ ( ⁇ ) is the number of times ⁇ appears in the search log.
- the purpose of query extension processing is to extend the synonymous query to avoid translation difficulties caused by the inclusion of specific nouns or new words in the source language query input by the user. For example, if the source language query input by the user is "Nokia n8 Mito", and "Mei Tu” appears in the network language, it does not appear frequently in the general translation model, which will bring the subsequent translation source language query. Difficulty, but if you expand it synonymously to "picture", it will reduce the difficulty of translation.
- the received source language query can be extended to a query set Q2, assuming that the source language query is q, and the extended Q2 is ⁇ , 3 ⁇ 4 2 ,... , 3 ⁇ 4 2 ⁇ , where n2
- the expanded query set Q2 contains q and the synonymous query extended by q.
- the source language query may be first processed by word segmentation, and the metaphorical resources of the source language, such as a synonym dictionary, may be searched to determine synonymous words of each word obtained after the word segmentation process, and the words obtained by the word segmentation process and their synonyms are combined. , select the combination of the query scores obtained in the query, the first n2 queries constitute the set Q2.
- the extended score of the query may be determined by the number of statistics of the query generated in the retelling resource.
- the rehearsal resources are not limited to words, but may also be phrases, even sentences, such as lexicographic annotation based substitution, word order transformation, sentence structure transformation, sentence splitting and merging, or reasoning based retelling, as long as The things described are the same, the meanings are the same, and they can all be considered as retelling resources. For example, if n2 is 2, the source language query input by the user is "Chinese cuisine", and after the word segmentation, "China” and "Gourmet" are obtained.
- the query error correction processing may be first performed to obtain the source language query set Q1, and each query in the Q1 obtained after the query error correction processing is respectively subjected to query expansion processing, and finally Get the collection
- Q ⁇ qi, q2, ..., q n ⁇ , n is a preset positive integer; you can also first perform query expansion processing to obtain the source language query set Q2, and then query the query extension to obtain the query in Q2.
- the query error correction processing is performed separately, and finally the set Q ⁇ ,..., 3 ⁇ 4 ⁇ is obtained, or the query error correction processing and the query expansion processing can be performed simultaneously, and the set Q1 and Q2 obtained by the query error correction processing and the query expansion processing are performed simultaneously.
- the set 0 ⁇ , 3 ⁇ 42, ..., q n ⁇ is obtained.
- Step 103 Translate the optimized source language query into N target languages query, where N is an integer greater than 1.
- N common language types may be used as a target language in advance, for example, commonly used English, French, Japanese, German, and the like are set as target languages in advance.
- the specific target language can be flexibly set according to requirements, and the present invention does not limit this.
- step 102 If the optimization process in step 102 is not performed on the source language query input by the user, the source language query input by the user is directly translated into N target language queries.
- the translation score of the translation result may be determined by at least one of the following two factors: the number of statistics of the translation result in the translation corpus used in the translation and the combination probability of each word in the translation result.
- the optimized source language query is translated into N target language queries. If the expanded query set Q is obtained after the optimization process, for each target language, each query in Q is translated into the target language query, and the translation result with the highest sum of translation scores is determined as The translation result of the target language.
- the source language query input by the user is "Chinese cuisine”
- the collection Q obtained after optimization processing is ⁇ "Chinese cuisine”, “Chinese cuisine”, “Chinese cuisine”, “Chinese cuisine”, for the target language is English.
- the translation model may include a translation model trained in advance using the extracted specific nouns and new words.
- Step 104 Obtain search results corresponding to the N target language queries respectively.
- the search is performed in the resource library corresponding to the target language.
- searching in the resource library of each target language it may be a normal search, that is, searching in an unstructured resource library corresponding to the target language; or a vertical search, that is, a structured resource library corresponding to the target language. Search in.
- the search results of the source language query can also be obtained at the same time. Specifically, if the source language query is not optimized, the search result is directly obtained by using the source language query input by the user. If the source language query is optimized, a query is selected from the source language query set Q obtained after the optimization process, and the selected query is used to search in the source language resource library to obtain the corresponding search result.
- the selected selection criteria may include but are not limited to:
- the search effect may be reflected by the number of search results, or the number of search results in the search result that are related to the source language query satisfying the preset relevance requirement, or the number of search results published within the set time, or, search The source of the results meets the number of search results required by the default source, and so on.
- Step 105 Integrate and sort the obtained search results to form a final search result set, and provide the final search result set to the user, where the ranking of each search result in the belonging category and the sorting weight of the belonging category are in the final search result set. Sort each search result.
- the search results are obtained from the resource libraries of the N target languages, or the search results are further obtained from the source language resource library, the obtained search results need to be integrated, and the integration includes: Merger and de-emphasis.
- the ranking weights of the classifications to which the search results belong and the ranks in the classified categories can be used to score the search results, and then sorted according to the score results from high to low.
- the scoring result score (Rst) of the search result Rst is:
- the classification of the search results can be determined by its corresponding language, ie, the search for which language the search results are derived from. Since in some cases, search results obtained by different languages may overlap, equation (2) is actually a process of weighting the average. For example, if a search result is both a Chinese search result and an English search result, the ranking of the search result in the English search result is multiplied by the English sort weight, and the row of the search result in the Chinese search result is added. Multiply by Chinese sort weight, get and divide by search The total number of languages used to obtain the score of the search results.
- the existing resources of each language may be mined in advance to train the feature vector; when determining the sorting weight of each language, the feature of the source language query is extracted, and the feature will be extracted.
- the feature is compared with the feature vector of each language to determine the language whose similarity exceeds the preset similarity threshold is the mapping language of the source language query.
- the existing resources corresponding to the Japanese are mined, and the feature vectors are trained to include: Tokyo, Japan, Sho, Kimono, Koizumi, Sakura, Kimura Takuya user input source language query for "Kimura Takuya" What are the famous movies?, the character of the source language query is extracted as "Kimura Takuya Movie", and the similarity between the extracted features and the feature vectors of each language is calculated, and the similarity between the feature vectors and the Japanese classification is determined to exceed the pre-preparation. If the similarity value is set, the mapping language of the source language query is determined to be Japanese.
- the sorting weight W of the classification is the first set value a; if the search result belongs to the source language and the source language is not the mapping language of the classified category, the sorting weight of the classified category is 1 ⁇ is the second set value b; if the classification to which the search result belongs is neither the mapping language nor the source language, the sorting weight 1 ⁇ of the belonging classification is the third set value c.
- the rank of the search result in the category to which it belongs may be determined by one or any combination of the degree of correlation between the search result and the query used by the search, the weight of the site from which the search result is derived, the time of publication of the search result, and the like.
- the query is optimized, after the error correction process is determined, the query is correct without error correction, and then the synonym dictionary is searched.
- the word “Beckham” can be used to expand “Beckham”, and the word “beauty” can be expanded. " , after the final extension, get the set Q as: ⁇ "Beckham picture”, “Beckham picture”, “Little Bmite”, “Beckham Meitu” ⁇ .
- the search engine's server side has three target languages set in advance: English, Japanese, and French.
- “Beckham Picture” and “Small Pome” can be translated into “Beckham picture”, “Cockles picture”, “Beckham picture” and “Little Pome” translated into “Beckham picture” "The corresponding translation scores are 8 and 6, respectively.
- the translation scores for "Beckham Pictures” and “Little Bmite” translated into “Cockles picture” are 1 and 2 respectively; “Beckham Pictures” and “Baker Hammet pictures can be translated into “Beckham picture”, and the corresponding translation scores are 9 and 7.
- the target language query is searched into the resource library of the corresponding target language, and each search result is obtained.
- the source language For the source language, select the best search "Beckham Picture" from the collection Q and get the search result of the selected query. Integrate search results in Chinese, English, Japanese, and French. When sorting the integrated search results, it is performed in accordance with the formula (2). After the feature extraction and similarity matching of the source language query "Beiismeeitu", if the mapping language is determined to be English, the sorting weight of English is 2, the Chinese sorting weight is 1, and the Japanese and French sorting weight is 0.5. . For one of the search results Pagel, if the Pagel exists in the Chinese search result, the English search result, and the Japanese search result, and the ranking in the Chinese search result is 2, the rank in the English search result is 5, in The ranking in the Japanese search result is 8, then the Pagel
- FIG. 3 is a structural diagram of a device according to an embodiment of the present invention.
- the device may include The user side interaction unit 300, the translation processing unit 310, the search processing unit 320, and the result integration unit 330.
- the user side interaction unit 300 is configured to receive a source language search request query input by the user, and perform a search formed by integrating the result integration unit 330.
- the result set is provided to the user.
- the translation processing unit 310 is configured to translate the source language query into N target language queries, and N is an integer greater than 1.
- the search processing unit 320 is configured to respectively obtain search results corresponding to the N target language queries.
- the result integration unit 330 is configured to integrate the search results obtained by the search processing unit 320 to form a final search result set; wherein, in the final sort result, according to the ranking of each search result in the belonging category and the classification thereof Sort weights, sorting each search result.
- the translation processing unit 310 may use, as the target language query, a translation result having the highest translation score among the translation results of the target language corresponding to the source language query for each target language.
- the translation score of the translation result e is determined by at least one of the following factors: the number of statistics of the translation result e in the translation corpus used in the translation and the combination probability of the words in the translation result e.
- the apparatus may further comprise: an optimization processing unit 340.
- the optimization processing unit 340 is configured to perform optimization processing on the source language query input by the user, and then provide the translation processing unit 310, wherein the optimization processing may include any one or combination of a query error correction processing and a query expansion processing.
- the translation processing unit 310 translates the source language query optimized by the optimization processing unit 340 into the N standard language query.
- the device includes an optimization processing unit 340, there may be three cases depending on the form of the optimization process:
- the merge processing module 423 is configured to combine Q1 and Q2 to obtain a source language query set Q containing n queries.
- the result integration unit 330 includes: a merge processing module 331, a deduplication processing module 332, and a sort processing module 333.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Description
一种跨语言搜索的方法和装置 本申请要求了申请日为 2011年 02月 28日,申请号为 201110047892.7 发明名称为"一种跨语言搜索的方法和装置"的中国专利申请的优先权, 其全部内容通过引用结合在本申请中。
技术领域
本发明涉及互联网技术领域, 特别涉及一种跨语言搜索的方法和装 置。
背景技术
随着互联网信息的不断增长,人们对于信息搜索提出了更高的要求, 不再满足于在同一种语种文档集中搜索, 而要求获取多种语种文档。 例 如, 如果用户输入的搜索词 (query ) 为 "贝克汉姆 图片" , 则中文文 档集中的搜索可能并不能最大程度地满足用户需求, 欧美网站的英文文 档集中可能具有更优、 更多的搜索结果。 当从多语种文档集中进行搜索的需求越来越高时, 为了获得更多、 更全面、 更准确的信息, 同时为了跨越语言障碍, 人们希望能够以一种 自己熟悉的语言描述 query, 而搜索结果中能够包括多语言的文档, 即进 行两语种之间的跨语言搜索。 发明内容
有鉴于此, 本发明提供了一种跨语言搜索的方法和装置, 以便于实 现包含多语言文档的搜索结果, 为用户提供更优、 更多的搜索结果。
具体技术方案如下:
一种跨语言搜索的方法, 该方法包括:
A、 接收用户输入的源语言搜索请求 query;
B、 将所述源语言 query翻译为 N种目标语言 query, N为大于 1的 整数;
C、 分别获取所述 N种目标语言 query对应的搜索结果;
D、 将步骤 C获取的搜索结果进行整合后形成最终的搜索结果集合 提供给用户;
其中在所述最终的搜索结果集合中, 根据各搜索结果在所属分类中 的排次以及所属分类的排序权重, 对各搜索结果进行排序。
在步骤 B中, 针对每一种目标语言, 将所述源语言 query对应的该 种目标语言的翻译结果中, 翻译分值最高的一种翻译结果作为目标语言 query;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
较优地, 所述步骤 B具体包括:
Bl、 对所述源语言 query进行优化处理, 所述优化处理包括 query 纠错处理和 query扩展处理中的任一种或组合;
B2、 将优化处理后的源语言 query翻译为 N种目标语言 query。 其中, 如果所述优化处理仅包括 query纠错处理, 则对所述用户输 入的源语言 query进行 query纠错处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的正整数;
所述步骤 B2具体为: 针对每一种目标语言, 分别利用所述 Q1中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言
query; 其中, 翻译结果的翻译分值总和为 为 Q 1中 被
翻译为 e的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
如果所述优化处理仅包括 query扩展处理, 则对所述用户输入的源 语言 query进行 query扩展处理后得到包含 n2个 query的源语言 query 集合 Q2 , n2为预设的正整数;
所述步骤 B2具体为: 针对每一种目标语言, 分别利用所述 Q2中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 为 Q2中 被
翻译为 e的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
如果所述优化处理既包括 query纠错处理又包括 query扩展处理,则 对所述用户输入的源语言 query进行 query纠错处理和 query扩展处理后 得到包含 n个 query的源语言 query集合 Q , n为预设的正整数;
所述步骤 B2具体为: 针对每一种目标语言, 分别利用所述 Q中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , τ ψ^)为 Q中 被翻
=1 译为 e的翻译分值;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
其中,对所述用户输入的源语言 query进行 query纠错处理后和 query 扩展处理后得到包含 n个 query的源语言 query集合 Q具体包括:
对所述用户输入的源语言 query进行 query纠错处理后得到包含 nl 个 query的源语言 query集合 Ql , nl为预设的正整数, 将所述 Q1中的 各 query分别进行 query扩展处理, 得到包含 n个 query的源语言 query 集合 Q; 或者,
对所述用户输入的源语言 query进行 query扩展处理后得到包含 n2 个 query的源语言 query集合 Q2 , n2为预设的正整数, 将所述 Q2中的 各 query分另 进行 query糾错处理 , 得 j包含 n个 query的源、语言 query 集合 Q; 或者,
对所述用户输入的源语言 query同时进行 query纠错处理和 query扩 展处理后, 分别得到包含 nl个 query的源语言 query集合 Q1和包含 n2 个 query的源语言 query集合 Q2 , 将所述 Q1和 Q2取并集后, 得到包含 n个 query的源语言 query集合 Q。
对所述用户输入的源语言 query进行 query纠错处理具体包括: 利用所述用户输入的源语言 query查找纠错训练语料, 判断纠错训 练语料中是否存在与所述用户输入的源语言 query相同的错误 query,如 果是, 则确定与所述用户输入的源语言 query相同的错误 query所对应 的所有正确 query,从确定的所有正确 query中选择对应纠错概率排在前 nl个的正确 query构成源语言 query集合 Q1 ; 否则, 所述 Q1中仅包括 所述用户输入的源语言 query;
其中, 所述纠错训练语料包括: 预先从搜索日志中收集的错误 query 和对应正确 query构成的 query对, 以及错误 query被纠错为对应正确 query的纠错概率。
对所述用户输入的源语言 query进行 query扩展处理具体包括: 将所述用户输入的源语言 query进行分词处理, 通过查找源语言的 复述资源确定分词处理后得到的各词语的同义词, 利用分词处理后得到 的各词语及各词语的同义词进行组合, 取组合得到的 query中扩展分值 排在前 n2个的 query构成所述 Q2;
query的扩展分值由创建所述复述资源中该 query的统计次数确定。 另外, 所述步骤 C还包括:
获取所述源语言 query对应的搜索结果。
对于对源语言 query进行优化处理的情况, 所述步骤 C还包括: 从 优化处理后得到的源语言 query集合中选择一个 query,获取选择的 query 对应的搜索结果。
所述从优化处理后得到的源语言 query集合中选择一个 query时,使 用的选择策略包括:
对优化处理后得到的源语言 query集合中的各 query逐一进行搜索, 直至找到搜索效果满足预设要求的 query,选择该搜索效果满足预设要求 的 query; 或者,
对优化处理后得到的源语言 query集合中的各 query进行搜索,选择 搜索效果最优的 query。
具体地, 步骤 D中所述整合包括: 对步骤 C获取的搜索结果进行合 并和去重。
所述根据各搜索结果在所属分类中的排次以及所属分类的排序权重, 对各搜索结果进行排序具体包括:
利用各搜索结果在所属分类中的排次以及所属分类的排序权重, 对 各搜索结果进行打分, 按照打分结果从高到低对各搜索结果进行排序; 其中, 搜索结果 Rst的打分结果 ore(Rst)为: ore(Rst) = 丄!^ ' ^ 1 ^ , m为搜索结果分类的总个数, 为第 i种分类的排序权 m ~^ ranki (Rst) 重, ra/^(Rst)为 Rst在第 i种分类中的排次。
特别地, 搜索结果所属分类为: 搜索结果对应的语言;
第 i种分类的排序权重的确定方法具体为:
S1、 提取所述用户输入的源语言 query的特征;
52、 将步骤 S1提取的特征与各语言的特征向量进行相似度计算, 确定相似度超过预设的相似度阔值的语言为所述用户输入的源语言 query的映射语 ^,
53、 对于搜索结果 Rst, 如果 Rst所属分类为映射语言, 则该所属分 类的排序权重 W为第一设定值 a;如果 Rst所属分类为源语言且源语言不 是该所属分类的映射语言, 则该所属分类的排序权重 为第二设定值 b; 如果 Rst所属分类既不是映射语言也不是源语言, 则该所属分类的排序 权重 为第三设定值 c;
其中, a>b>c ,各语言的特征向量是预先对各语言的已有资源进行挖 掘所训练出来的。
一种跨语言搜索的装置, 该装置包括: 用户侧交互单元、 翻译处理 单元、 搜索处理单元和结果整合单元;
所述用户侧交互单元, 用于接收用户输入的源语言搜索请求 query,
将所述结果整合单元整合后形成的搜索结果集合提供给所述用户; 所述翻译处理单元, 用于将所述源语言 query翻译为 N种目标语言 query, N为大于 1的整数;
所述搜索处理单元, 用于分别获取所述 N种目标语言 query对应的 搜索结果;
所述结果整合单元, 用于将所述搜索处理单元获取的搜索结果进行 整合后形成最终的搜索结果集合;其中,在所述最终的搜索结果集合中, 根据各搜索结果在所属分类中的排次以及所属分类的排序权重, 对各搜 索结果进行排序。
所述翻译处理单元针对每一种目标语言, 将所述源语言 query对应 的该种目标语言的翻译结果中, 翻译分值最高的一种翻译结果作为目标 语言 query;
翻译结果 e的翻译分值由以下因素中的至少一种确定: 翻译所使用 的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概 率。
较优地, 该装置还包括: 优化处理单元;
所述优化处理单元, 用于对所述用户输入的源语言 query进行优化 处理后提供给所述翻译处理单元, 所述优化处理包括 query糾错处理和 query扩展处理中的任一种或组合;
所述翻译处理单元将所述优化处理单元进行优化处理后的源语言 query翻译为 N种 标语言 query。
如果所述优化处理仅包括 query糾错处理, 则所述优化处理单元对 所述用户输入的源语言 query进行 query纠错处理后得到包含 nl个 query
的源语言 query集合 Ql , nl为预设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q1中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , ^(^^为 Ql中 被翻译为 e
=1 的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定: 翻译所 使用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组 合概率。
如果所述优化处理仅包括 query扩展处理, 则所述优化处理单元对 所述用户输入的源语言 query进行 query扩展处理后得到包含 n2个 query 的源语言 query集合 Q2 , n2为预设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q2中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , ^(^^为 Q2中 被翻译为 e
=1 的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
如果所述优化处理既包括 query纠错处理又包括 query扩展处理,则 所述优化处理单元对所述用户输入的源语言 query进行 query纠错处理 和 query扩展处理后得到包含 n个 query的源语言 query集合 Q , n为预 设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中,翻译结果的翻译分值总和为 Q中 被翻译为 e的
翻译分值;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
具体地, 所述优化处理单元可以包括: 第一纠错模块和第一扩展模 块;
所述第一纠错模块,用于对所述用户输入的源语言 query进行 query 纠错处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的 正整数;
所述第一扩展模块, 用于将所述 Q1中的各 query分别进行 query扩 展处理, 得到包含 n个 query的源语言 query集合 Q。
或者,所述优化处理单元具体包括:第二扩展模块和第二糾错模块; 所述第二扩展模块,用于对所述用户输入的源语言 query进行 query 扩展处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的 正整数;
所述第二纠错模块, 用于将所述 Q2中的各 query分别进行 query纠 错处理, 得到包含 n个 query的源语言 query集合 Q。
或者, 所述优化处理单元具体包括: 第三纠错模块、 第三扩展模块 和合并处理模块;
所述第三纠错模块,用于对所述用户输入的源语言 query进行 query 纠错处理, 得到包含 nl个 query的源语言 query集合 Q1;
所述第三扩展模块,用于对所述用户输入的源语言 query进行 query 扩展处理, 得到包含 n2个 query的源语言 query集合 Q2;
所述合并处理模块, 用于将所述 Q1和 Q2取并集后, 得到包含 n个 query的源语言 query集合 Q。 训练语料,判断纠错训练语料中是否存在与所述用户输入的源语言 query 相同的错误 query, 如果是, 则确定与所述用户输入的源语言 query相同 的错误 query所对应的所有正确 query , 从确定的所有正确 query中选择 对应纠错概率排在前 nl个的正确 query构成源语言 query集合 Q1 ;否则, 所述 Q1中仅包括所述用户输入的源语言 query;
其中, 所述纠错训练语料包括: 预先从搜索日志中收集的错误 query 和对应正确 query构成的 query对, 以及错误 query被纠错为对应正确 query的纠错概率。 理,通过查找源语言的复述资源确定分词处理后得到的各词语的同义词, 利用分词处理后得到的各词语及各词语的同义词进行组合, 取组合得到 的 query中扩展分值排在前 n2个的 query构成所述 Q2;
query的扩展分值由创建所述复述资源中该 query的统计次数确定。 另外, 所述搜索处理单元, 还用于获取所述源语言 query对应的搜 索结果。
对应于包括优化处理单元的情况, 该装置还包括: 源语言 query集合中选择一个 query;
所述搜索处理单元, 还用于获取所述源请求选择单元选择的 query 对应的搜索结果。
其中, 所述源请求选择单元釆用的选择策略包括:
对所述优化处理单元进行优化处理后得到的源语言 query集合中的 各 query逐一进行搜索, 直至找到搜索效果满足预设要求的 query, 选择 该搜索效果满足预设要求的 query; 或者,
对所述优化处理单元进行优化处理后得到的源语言 query集合中的 各 query进行搜索, 选择搜索效果最优的 query。
具体地, 所述结果整合单元包括: 合并处理模块、 去重处理模块和 排序处理模块;
所述合并处理模块, 用于将所述搜索处理单元获取的搜索结果进行 合并处理;
所述去重处理单元, 用于将所述合并处理模块合并处理后的搜索结 果进行去重处理得到搜索结果集合;
所述排序处理模块, 用于在所述搜索结果集合中, 根据各搜索结果 在所属分类中的排次以及所属分类的排序权重,对各搜索结果进行排序。
所述排序处理模块具体利用所述搜索结果集合中各搜索结果在所属 分类中的排次以及所属分类的排序权重, 对各搜索结果进行打分, 按照 打分结果从高到低对各搜索结果进行排序;
其中, 搜索结果 Rst的打分结果 ore(Rst)为: ore(Rst) = 丄!^ ' ^ 1 ^ , m为搜索结果分类的总个数, 为第 i种分类的排序权 m ~^ ranki (Rst) 重, ra/^(Rst)为 Rst在第 i种分类中的排次。
搜索结果所属分类为: 搜索结果对应的语言;
所述排序处理模块在确定第 i种分类的排序权重时, 具体执行以下 操作: 提取所述用户输入的源语言 query的特征; 将提取的特征与各语 言的特征向量进行相似度计算, 确定相似度超过预设的相似度阈值的语 言为所述用户输入的源语言 query的映射语言; 对于搜索结果 Rst, 如果 Rst所属分类为映射语言, 则该所属分类的排序权重 为第一设定值 a; 如果 Rst所属分类为源语言且源语言不是该所属分类的映射语言, 则该 所属分类的排序权重 W为第二设定值 b;如果 Rst所属分类既不是映射语 言也不是源语言, 则该所属分类的排序权重 W为第三设定值 c;
其中, a>b>c, 各语言的特征向量是预先对各语言的已有资源进行挖 掘所训练出来的。
由以上技术方案可以看出, 本发明将源语言 query翻译为多种目标 语言 query后, 获取 N种目标语言 query对应的搜索结果, 对搜索结果 进行整合后形成最终的搜索结果集合提供给用户, 其中在最终的搜索结 果集合中, 根据各搜索结果在所属分类中的排次以及所属分类的排序权 重, 对各搜索结果进行排序。 通过本发明实现了包含多语言文档的搜索 结果, 为用户提供了更优、 更多的搜索结果。
附图说明
图 1为本发明实施例一提供的方法流程图;
图 2为本发明实施例二提供的示意图;
图 3为本发明实施例三提供的装置结构图;
图 4中的(a)(b)(c)为本发明实施例三提供的优化处理单元的三种结 构图。
具体实施方式
为了使本发明的目的、 技术方案和优点更加清楚, 下面结合附图和 具体实施例对本发明进行详细描述。
实施例一、
图 1为本发明实施例一提供的方法流程图, 如图 1所示, 该方法可 以包括以下步骤:
步骤 101 : 接收用户输入的源语言 query。
该方法可以在搜索引擎的服务器端实现, 在用户输入源语言 query 后, 浏览器将用户输入的源语言 query发送给搜索引擎的服务器端。
步骤 102: 对源语言 query进行优化处理。
本步骤为可选步骤,对源语言 query进行的优化处理可以包括: query 纠错处理和 query扩展处理中的任一种或组合。 其中 query纠错处理的 目的是为了提高源语言 query被正确翻译的可能性, query扩展处理的目 的是为了扩大 query的资源召回率。
其中, query纠错处理旨在纠正 query中的错误, 主要是拼写错误, 可以基于噪声信道模型实现。 具体包括: 预先从搜索日志中收集大量的 错误 query和对应正确 query构成的 query对, 并计算错误 query被纠错 为正确 query的概率, 即纠错概率 Ρ(ρ|ρ,), 组成纠错训练语料; 针对用 户输入的源语言 query查询纠错训练语料, 判断纠错训练语料中是否存 在与源语言 query相同的错误 query , ^口果是, 则确定与源语言 query 目 同的错误 query所对应的所有正确 query,从中选择对应纠错概率排在前 nl个的正确 query构成源语言 query进行纠错处理后的源语言 query集 合 Ql , nl为预设的正整数; 如果否, 则无需对源语言 query进行纠错处 理, 最终得到的 Q1中仅包括用户输入的源语言 query。
其中, (p'|p)= P( ')P( ')/P( )
P(p'|p)为源语言 query p被纠错为 p'的纠错概率, Ρ(ρ|ρ')为搜索曰志 中 ρ'被错写成 ρ的概率, Ρ(ρ')是 ρ'在搜索日志中出现的次数, Ρ(ρ)是 ρ 在搜索日志中出现的次数。
query扩展处理的目的在于扩展出同义 query, 以避免由于用户输入 的源语言 query中包含专用名词或新词所带来的翻译困难的问题。例如, 如果用户输入的源语言 query为 "诺基亚 n8 美图" , 其中 "美图" 多出 现在网络用语中, 在一般的翻译模型中不常出现, 这就会给后续翻译源 语言 query带来难度, 但倘若将其进行同义扩展为 "图片" , 就会降低 翻译难度。
在进行同义扩展时, 可以将接收到的源语言 query扩展为一个 query 集合 Q2 ,假设源语言 query为 q,扩展出的 Q2为 { , ¾2,... , ¾2 } ,其中, n2为预设的正整数, 扩展出的 query集合 Q2中包含 q以及由 q扩展出 的同义 query„
具体地, 可以首先将源语言 query进行分词处理, 通过查找源语言 的复述资源,例如同义词词典,确定分词处理后得到的各词语的同义词, 利用分词处理后得到的各词语及其同义词进行组合后, 选取组合得到的 query中扩展分值在前 n2个的 query构成集合 Q2。 其中, query的扩展 分值可以由创建复述资源中 query出现的统计次数确定。 需要特别说明 的是, 复述资源不限于词, 也可以为短语, 甚至为句子, 例如基于词典 注释的替换、 语序变换、 句子结构变换、 句子拆分与合并或基于推理的 复述得到的资源, 只要描述的事物相同, 表达的含义相同, 都可以认为 是复述资源。
例如, 如果 n2为 2, 用户输入的源语言 query为 "中国 美食" , 进 行分词处理后得到 "中国" 和 "美食" , "中国" 的同义词有 "中华" ; "美食 "的同义词有 "佳肴",进行组合后得到的 query为"中国 美食", "中国 佳肴" , "中华 美食" , "中华 佳肴" , 选取扩展分值排在 前 2个的 query构成 Q2{ "中国 美食" , "中华 美食" }。
如果上述优化处理包括: query纠错处理以及 query扩展处理, 可以 首先进行 query纠错处理, 得到源语言 query集合 Q1 , 对 query纠错处 理后得到的 Q1中的各 query分别进行 query扩展处理, 最终得到集合
Q{qi, q2, ..., qn}, n为预设的正整数; 也可以首先进行 query扩展处理, 得到源语言 query集合 Q2, 然后对 query扩展处理后得到的 Q2中的各 query分别进行 query纠错处理, 最终得到集合 Q{ ^,..., ¾}, 或者, query纠错处理以及 query扩展处理可以同时进行, 将 query纠错处理以 及 query扩展处理得到的集合 Q1和 Q2取并集后,得到集合 0{ , ¾2, ... , qn }。
步骤 103:将优化处理后的源语言 query,翻译为 N种目标语言 query, 其中 N为大于 1的整数。
本发明实施例中可以预先将 N种常用的语言种类作为目标语言, 例 如, 预先将常用的英文、 法文、 日文、 德文等设置为目标语言。 具体设 置哪几种目标语言可以根据需求灵活设置, 本发明对此并不加限制。
如果没有对用户输入的源语言 query执行步骤 102中的优化处理, 则直接将用户输入的源语言 query翻译为 N种目标语言 query。
在对源语言 query翻译成第 m种目标语言时, 可能会得到多种不同 的翻译结果, 针对这种情况, 可以选择多种不同的翻译结果中翻译分值
最高的一种翻译结果。 即如果源语言 query为 q, 则针对第 m种目标语 言的翻译结果 e为: e =argmax P(e|q) , 其中, P(e|q)表示 q被翻译为 e的翻 译分值。
其中,翻译结果的翻译分值可以由以下两种因素中的至少一种确定: 翻译所使用的翻译语料库中该翻译结果出现的统计次数以及翻译结果中 各词的组合概率。
如果对用户输入的源语言 query执行了步骤 102中的优化处理, 则 将优化处理后的源语言 query翻译为 N种目标语言 query。如果在进行优 化处理后得到了扩展出的 query集合 Q , 则针对每一种目标语言, 分别 将 Q中的各 query都翻译为目标语言 query,从中确定出一个翻译分值总 和最高的翻译结果作为该种目标语言的翻译结果。
即针对每一种目标语言,其翻译结果 e为: e =argmax ^JP(e|qi)。 ( 1 ) 其中 表示 q扩展出的集合 Q中 翻译为 e的翻译分值, 如果 不能被翻译为 e , 则对应翻译分值为零, !^^^表示集合 Q中能够被 翻译为 e的所有 的翻译分值总和。例如,用户输入的源语言 query为 "中 国 美食 ",优化处理后得到的集合 Q为 { "中国 美食", "中国 佳肴", "中华 美食" , "中华 佳肴" } , 针对目标语言为英文的情况, "中国 美食" 、 "中华 美食"和 "中华 佳肴"都被翻译为 "Chinese cuisine" , 则^口果这三个 query被翻译为 "Chinese cuisine" 的翻译分值总和最大, 则将 "Chinese cuisine" 作为英文的翻译结果。
需要说明的是, 如果在优化处理时仅进行了 query纠错处理, 则 Q 就是步骤 102中所述的 Q1 , 公式( 1 ) 中 n即为 nl ; 如果在优化处理时
仅进行了 query扩展处理, 则 Q就是步骤 102中所述的 Q2 , 公式 ( 1 ) 中 n即为 n2。
另外,在对用户输入的源语言 query或者由源语言 query扩展出的 Q 中各 query进行翻译时, 可以通过匹配翻译模型的方式实现。 更优地, 翻译模型中可以包含预先利用挖掘出的专用名词、 新词所训练出的翻译 模型。
步骤 104: 分别获取 N种目标语言 query对应的搜索结果。
在翻译得到 N种目标语言 query时, 分别在对应目标语言的资源库 中进行搜索。 在每一种目标语言的资源库中进行搜索时, 可以是普通搜 索, 即在对应目标语言的非结构化资源库中进行搜索; 也可以是垂直搜 索, 即在对应目标语言的结构化资源库中进行搜索。
另外, 除了将源语言 query翻译为 N种目标语言 query并获取目标 语言 query对应的搜索结果之外, 还可以同时获取源语言 query的搜索 结果。 具体地, 如果不对源语言 query进行优化处理, 则直接利用用户 输入的源语言 query获取搜索结果。 如果对源语言 query进行了优化处 理, 则从优化处理后得到的源语言 query集合 Q中选择一个 query, 利用 该选择的 query在源语言的资源库中进行搜索获取对应的搜索结果。
在从源语言 query集合 Q中选择 query时 , 采用的选择的策格可以 包括但不限于:
对集合 Q中的各 query逐一进行搜索, 直至找到搜索效果满足预设 要求的 query,选择该搜索效果满足预设要求的 query作为目标语言 query; 或者, 对集合 Q中的各 query进行搜索, 选择搜索效果最优的 query作 为 标语言 query。
其中, 搜索效果可以体现为搜索结果数量, 或者, 搜索结果中与源 语言 query的相关度满足预设相关度要求的搜索结果数量, 或者, 在设 定时间内发布的搜索结果数量, 或者, 搜索结果的来源满足预设来源要 求的搜索结果数量, 等等。
步骤 105 : 将获取的各搜索结果进行整合和排序后形成最终的搜索 结果集合提供给用户, 其中根据各搜索结果在所属分类中的排次以及所 属分类的排序权重, 在最终的搜索结果集合中对各搜索结果进行排序。
由于从 N种目标语言的资源库中都获得了搜索结果, 或者更进一步 从源语言的资源库中也获得了搜索结果, 因此, 需要对获得的各搜索结 果进行整合, 其中整合包括: 搜索结果的合并和去重。
在对整合后的搜索结果进行排序时, 可以利用各搜索结果所属分类 的排序权重以及在所属分类中的排次为各搜索结果打分, 然后按照打分 结果从高到低进行排序。
具体地 , 搜索结果 Rst的打分结果 score (Rst)为:
丄
通常, 搜索结果所属分类可以由其对应的语言确定, 即搜索结果来 源于哪种语言的搜索。 由于在有些情况下, 通过不同的语言得到的搜索 结果可能存在重合, 因此, 公式(2 ) 实际上是求加权平均值的处理。 例 如, 某一个搜索结果既是中文的搜索结果也是英文的搜索结果, 则将该 搜索结果在英文搜索结果中的排次乘以英文的排序权重, 再加上该搜索 结果在中文搜索结果中的排次乘以中文的排序权重, 获得的和再除以搜
索所使用语言的总个数, 得到该搜索结果的打分结果。
本发明实施例中, 可以预先对各语言 (包括各目标语言和源语言) 的已有资源进行挖掘, 训练出特征向量; 在确定各语言的排序权重时, 提取源语言 query的特征, 将提取的特征与各语言的特征向量进行相似 度计算, 确定相似度超过预设的相似度阔值的语言为该源语言 query的 映射语言。
例如, 对于按照语言进行分类, 对日文对应的已有资源进行挖掘, 训练出特征向量包括: 东京、 日本、 相朴、 和服、 小泉、 樱花、 木村拓 当用户输入的源语言 query为 "木村拓哉有哪些著名电影" , 则提 取源语言 query的特征为 "木村拓哉 电影" , 对提取的特征与各语言的 特征向量进行相似度计算, 确定与日文的分类的特征向量之间的相似度 超过预设的相似度阔值, 则确定该源语言 query的映射语言为日文。
如果搜索结果所属分类为映射语言, 则该分类的排序权重 W为第一 设定值 a; 如果搜索结果所属分类为源语言且源语言不是该所属分类的 映射语言, 则所属分类的排序权重1^为第二设定值 b; 如果搜索结果所 属分类既不是映射语言也不是源语言, 则所属分类的排序权重1^为第三 设定值 c。 其中, a>b>c , 例如可以取 b=l , a>l,c<l。
另外, 搜索结果在其所属分类中的排次可以由搜索结果与搜索所使 用 query之间的相关度、 搜索结果所来源的站点权重、 搜索结果的发布 时间等中的一种或任意组合确定。
至此实施例一所示流程结束。 下面通过实施例二中一个具体实例对 上述方法进行描述。
实施例二、
实施例二对应的示意图如图 2所示, 假设用户输入中文 query: "小 贝 美图" 。
对该 query进行优化处理,经过纠错处理后确定该 query正确无需纠 错, 然后查找同义词词典, 利用词语 "小贝"可以扩展出 "贝克汉姆" 、 词语 "美图" 可以扩展出 "图片" , 最终扩展后得到集合 Q为: { "小 贝 图片" , "贝克汉姆 图片" , "小贝 美图" , "贝克汉姆 美图" }。
假设搜索引擎的服务器端预先设置了 3种目标语言: 英文、 日文和 法文, 另外还需要针对源语言 (中文) 进行搜索, 即最终利用 4种语言 进行搜索。 其中, 在翻译成英文时, "小贝 图片" 和 "小贝 美图" 可 以翻译成 "Beckham picture" 、 "Cockles picture" , "小贝 图片" 和 "小贝 美图 "翻译成 "Beckham picture"对应的翻译分值分别为 8和 6 , "小贝 图片" 和 "小贝 美图" 翻译成 "Cockles picture" 对应的翻译分 值分别为 1和 2; "贝克汉姆 图片" 和 "贝克汉姆 美图" 均可以翻译 成 "Beckham picture" , 对应的翻译分值分别为 9和 7。
计算 "Beckham picture" 的翻译分值总和为 8+6+9+7=30 , 计算 "Cockles picture" 的翻译分值总和为 1+2=3。
最终选取英文的翻译结果为 "Beckham picture" 。
其他目标语言的翻译结果类似, 不再举例描述。
然后, 分别利用各目标语言的翻译结果, 作为目标语言 query分别 到对应目标语言的资源库中进行搜索, 得到各搜索结果。
针对源语言, 从集合 Q中选择搜索效果最好的 query "贝克汉姆 图 片" , 并获取该选择的 query的搜索结果。
将中文、 英文、 日文和法文对应的搜索结果进行整合。 在对整合后的搜索结果进行排序时, 按照公式(2 )的方式进行。 在 对源语言 query "小贝 美图" 进行特征提取以及相似度匹配后, 确定映 射语言为英文, 则取英文的排序权重为 2 , 中文的排序权重为 1 , 日文和 法文的排序权重为 0.5。 对于其中一个搜索结果 Pagel , 如果该 Pagel在中文检索结果、 英 文检索结果和日文检索结果中都存在, 且在中文检索结果中的排次为 2 , 在英文检索结果中的排次为 5 ,在日文检索结果中的排次为 8 ,则该 Pagel
— x (lx— + 2x— + 0 5 x— )
的打分结果 5Core(Pa§e 为: score(Page 1) = 4 2 5 8 =0.24。 实施例三、 图 3为本发明实施例提供的装置结构图, 如图 3所示, 该装置可以 包括: 用户侧交互单元 300、 翻译处理单元 310、 搜索处理单元 320和结 果整合单元 330。 用户侧交互单元 300 , 用于接收用户输入的源语言搜索请求 query, 将结果整合单元 330整合后形成的搜索结果集合提供给用户。 翻译处理单元 310 ,用于将源语言 query翻译为 N种目标语言 query, N为大于 1的整数。 搜索处理单元 320 , 用于分别获取 N种目标语言 query对应的搜索 结果。 结果整合单元 330 , 用于将搜索处理单元 320获取的搜索结果进行 整合后形成最终的搜索结果集合; 其中, 在最终的排序结果中, 根据各 搜索结果在所属分类中的排次以及所属分类的排序权重, 对各搜索结果 进行排序。
具体地, 翻译处理单元 310可以针对每一种目标语言, 将源语言 query对应的该种目标语言的翻译结果中 ,翻译分值最高的一种翻译结果 作为目标语言 query。 其中, 翻译结果 e的翻译分值由以下因素中的至少 一种确定: 翻译所使用的翻译语料库中翻译结果 e的统计次数以及翻译 结果 e中各词的组合概率。
较优地, 该装置还可以包括: 优化处理单元 340。
优化处理单元 340 ,用于对用户输入的源语言 query进行优化处理后 提供给翻译处理单元 310 , 其中优化处理可以包括 query纠错处理和 query扩展处理中的任一种或组合。
翻译处理单元 310将优化处理单元 340进行优化处理后的源语言 query翻译为 N种 标语言 query。
如果装置中包含优化处理单元 340 , 则根据优化处理的形式可以存 在以下三种情况:
第一种情况: 如果优化处理仅包括 query糾错处理, 则优化处理单 元 340对用户输入的源语言 query进行 query纠错处理后得到包含 nl个 query的源语言 query集合 Q 1 , nl为预设的正整数。
此时, 翻译处理单元 310针对每一种目标语言, 分别利用 Q1中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , ^(ψ^)为 Q1中 被
=1
翻译为 e的翻译分值。
翻译结果 e对应的翻译分值由以下因素中的至少一种确定: 翻译所 使用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组 合概率。
第二种情况: 如果优化处理仅包括 query扩展处理, 则优化处理单 元 340对用户输入的源语言 query进行 query扩展处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的正整数。
翻译处理单元 310针对每一种目标语言, 分别利用 Q2中的各 query 进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query;其中, 翻译结果的翻译分值总和为 !^(+1 ,
Q2中 被翻译为 e的翻译 分值。
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
第三种情况:如果优化处理既包括 query纠错处理又包括 query扩展 处理, 则优化处理单元 340对用户输入的源语言 query进行 query纠错 处理和 query扩展处理后得到包含 n个 query的源语言 query集合 Q, n 为预设的正整数。
翻译处理单元 310针对每一种目标语言, 分别利用 Q中的各 query 进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query;其中, 翻译结果的翻译分值总和为 ^^;!为 Q中 被翻译为 e的翻译
分值。
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
在第三种情况中, 优化处理单元 340可以存在三种结构:
第一种结构:如图 4中的(a)所示,优化处理单元 340可以具体包括:
第一纠错模块 401和第一扩展模块 402。
第一纠错模块 401 ,用于对用户输入的源语言 query进行 query纠错 处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的正整 数。
第一扩展模块 402 , 用于将 Q1中的各 query分别进行 query扩展处 理, 得到包含 n个 query的源语言 query集合 Q。
第二种结构:如图 4中的 (b)所示,优化处理单元 340可以具体包括: 第二扩展模块 411和第二纠错模块 412。
第二扩展模块 411 , 用于对用户输入的源语言 query进行 query扩展 处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的正整 数。
第二纠错模块 412 , 用于将 Q2中的各 query分别进行 query糾错处 理 , 得到包含 n个 query的源语言 query集合 Q。
第三种结构:如图 4中的(c)所示,优化处理单元 340可以具体包括: 第三纠错模块 421、 第三扩展模块 422和合并处理模块 42
第三纠错模块 421 , 用于对用户输入的源语言 query进行 query纠错 处理, 得到包含 nl个 query的源语言 query集合 Q1。
第三扩展模块 422 , 用于对用户输入的源语言 query进行 query扩展 处理, 得到包含 n2个 query的源语言 query集合 Q2。
合并处理模块 423 ,用于将 Q1和 Q2取并集后,得到包含 n个 query 的源语言 query集合 Q。
具体地, 图 3中所示的优化处理单元 340进行的 query纠错处理可 以具体为: 利用用户输入的源语言 query查找纠错训练语料, 判断纠错
训练语料中是否存在与用户输入的源语言 query相同的错误 query,如果 是, 则确定与用户输入的源语言 query相同的错误 query所对应的所有 正确 query, 从确定的所有正确 query中选择对应纠错概率排在前 nl个 的正确 query构成源语言 query集合 Q1 ; 否则, Q1中仅包括用户输入的 源语言 query„
其中, 纠错训练语料包括: 预先从搜索日志中收集的错误 query和 对应正确 query构成的 query对,以及错误 query被纠错为对应正确 query 的纠错概率。
优化处理单元 340进行的 query扩展处理可以具体为: 将用户输入 的源语言 query进行分词处理, 通过查找源语言的复述资源, 例如同义 词词典, 确定分词处理后得到的各词语的同义词, 利用分词处理后得到 的各词语及各词语的同义词进行组合, 取组合得到的 query中扩展分值 排在前 n2个的 query构成 Q2。
其中, query的扩展分值由创建复述资源中该 query的统计次数确定。 需要特别说明的是, 复述资源不限于词, 也可以为短语, 甚至为句子, 例如基于词典注释的替换、 语序变换、 句子结构变换、 句子拆分与合并 或基于推理的复述得到的资源 ,只要描述的事物相同 ,表达的含义相同, 都可以认为是复述资源。
另外, 为了同时获取到源语言的搜索结果, 上述搜索处理单元 320 还可以用于获取源语言 query对应的搜索结果。
如果对源语言 query进行了优化处理, 则该装置可以进一步包括: 源请求选择单元 350 , 用于从优化处理单元 340进行优化处理后得到的 源语言 query集合中选择一个 query。
此时, 搜索处理单元 320在针对源语言获取搜索结果时, 具体是获 取源请求选择单元 350选择的 query对应的搜索结果。
源请求选择单元 350釆用的选择策略可以包括但不限于: 对优化处 理单元 340进行优化处理后得到的源语言 query集合中的各 query逐一 进行搜索, 直至找到搜索效果满足预设要求的 query,选择该搜索效果满 足预设要求的 query; 或者,对优化处理单元 340进行优化处理后得到的 源语言 query集合中的各 query进行搜索, 选择搜索效果最优的 query。
另外, 结果整合单元 330包括: 合并处理模块 331、 去重处理模块 332和排序处理模块 333。
合并处理模块 331 , 用于将搜索处理单元 320获取的搜索结果进行 合并处理。
去重处理单元 332 , 用于将合并处理模块 331合并处理后的搜索结 果进行去重处理得到搜索结果集合。
排序处理模块 333 , 用于在上述搜索结果集合中, 根据各搜索结果 在所属分类中的排次以及所属分类的排序权重,对各搜索结果进行排序。
具体地, 排序处理模块 333利用搜索结果集合中各搜索结果在所属 分类中的排次以及所属分类的排序权重, 对各搜索结果进行打分, 按照 打分结果从高到低对各搜索结果进行排序。
上述搜索结果所属分类可以为: 搜索结果对应的语言, 该语言可以 包括目标语言, 还可以包括源语言。
排序处理模块 333在确定第 i种分类的排序权重时, 具体执行以下 操作: 提取用户输入的源语言 query的特征; 将提取的特征与各语言的 特征向量进行相似度计算, 确定相似度超过预设的相似度阈值的语言为 用户输入的源语言 query的映射语言; 对于搜索结果 Rst, 如果 Rst所属 分类为映射语言, 则该所属分类的排序权重 W为第一设定值 a; 如果 Rst 所属分类为源语言且源语言不是该所属分类的映射语言, 则该所属分类 的排序权重 W为第二设定值 b ;如果 Rst所属分类既不是映射语言也不是 源语言, 则该所属分类的排序权重 W为第三设定值 c。 其中, a>b>c , 各 语言的特征向量是预先对各语言的已有资源进行挖掘所训练出来的。
以上所述仅为本发明的较佳实施例而已, 并不用以限制本发明, 凡 在本发明的精神和原则之内, 所做的任何修改、 等同替换、 改进等, 均 应包含在本发明保护的范围之内。
Claims
1、 一种跨语言搜索的方法, 其特征在于, 该方法包括:
A、 接收用户输入的源语言搜索请求 query;
B、 将所述源语言 query翻译为 N种目标语言 query, N为大于 1的 整数;
C、 分别获取所述 N种目标语言 query对应的搜索结果;
D、 将步骤 C获取的搜索结果进行整合后形成最终的搜索结果集合 提供给用户;
其中在所述最终的搜索结果集合中, 根据各搜索结果在所属分类中 的排次以及所属分类的排序权重, 对各搜索结果进行排序。
2、 根据权利要求 1所述的方法, 其特征在于, 在步骤 B中, 针对 每一种目标语言, 将所述源语言 query对应的该种目标语言的翻译结果 中, 翻译分值最高的一种翻译结果作为目标语言 query;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
3、 根据权利要求 1所述的方法, 其特征在于, 所述步骤 B具体包 括:
Bl、 对所述源语言 query进行优化处理, 所述优化处理包括 query 纠错处理和 query扩展处理中的任一种或组合;
B2、 将优化处理后的源语言 query翻译为 N种目标语言 query。
4、 根据权利要求 3所述的方法, 其特征在于, 如果所述优化处理仅 包括 query纠错处理,则对所述用户输入的源语言 query进行 query纠错 处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的正整 数;
所述步骤 B2具体为: 针对每一种目标语言, 分别利用所述 Q1中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中,翻译结果的翻译分值总和为 ^^(ψ^) , ^(ψ^)为所述 Q1中
=1 被翻译为 e的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
5、 根据权利要求 3所述的方法, 其特征在于, 如果所述优化处理仅 包括 query扩展处理 ,则对所述用户输入的源语言 query进行 query扩展 处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的正整 数;
所述步骤 B2具体为: 针对每一种目标语言, 分别利用所述 Q2中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中,翻译结果的翻译分值总和为 ^^(ψ^) , ^(ψ^)为所述 Q2中
=1 被翻译为 e的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
6、 根据权利要求 3所述的方法, 其特征在于, 如果所述优化处理既 包括 query纠错处理又包括 query扩展处理, 则对所述用户输入的源语 言 query进行 query纠错处理和 query扩展处理后得到包含 n个 query的 源语言 query集合 Q , η为预设的正整数;
所述步骤 Β2具体为: 针对每一种目标语言, 分别利用所述 Q中的 各 query进行翻译, 确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , τ ψ^)为 Q中 被翻
=1 译为 e的翻译分值;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
7、 根据权利要求 6所述的方法, 其特征在于, 对所述用户输入的源 语言 query进行 query纠错处理后和 query扩展处理后得到包含 n个 query 的源语言 query集合 Q具体包括:
对所述用户输入的源语言 query进行 query纠错处理后得到包含 nl 个 query的源语言 query集合 Ql , nl为预设的正整数, 将所述 Q1中的 各 query分别进行 query扩展处理, 得到包含 n个 query的源语言 query 集合 Q; 或者,
对所述用户输入的源语言 query进行 query扩展处理后得到包含 n2 个 query的源语言 query集合 Q2 , n2为预设的正整数, 将所述 Q2中的 各 query分另 进行 query糾错处理 , 得 j包含 n个 query的源、语言 query 集合 Q; 或者,
对所述用户输入的源语言 query同时进行 query纠错处理和 query扩 展处理后, 分别得到包含 nl个 query的源语言 query集合 Q1和包含 n2 个 query的源语言 query集合 Q2 , 将所述 Q1和 Q2取并集后, 得到包含 n个 query的源语言 query集合 Q。
8、 根据权利要求 3、 4或 7所述的方法, 其特征在于, 对所述用户 输入的源语言 query进行 query纠错处理具体包括:
利用所述用户输入的源语言 query查找纠错训练语料, 判断纠错训 练语料中是否存在与所述用户输入的源语言 query相同的错误 query,如 果是, 则确定与所述用户输入的源语言 query相同的错误 query所对应 的所有正确 query,从确定的所有正确 query中选择对应纠错概率排在前 nl个的正确 query构成源语言 query集合 Ql , nl为预设的正整数;否则, 所述 Q1中仅包括所述用户输入的源语言 query;
其中, 所述纠错训练语料包括: 预先从搜索日志中收集的错误 query 和对应正确 query构成的 query对, 以及错误 query被纠错为对应正确 query的纠错概率。
9、 根据权利要求 3、 5或 7所述的方法, 其特征在于, 对所述用户 输入的源语言 query进行 query扩展处理具体包括:
将所述用户输入的源语言 query进行分词处理, 通过查找源语言的 复述资源确定分词处理后得到的各词语的同义词, 利用分词处理后得到 的各词语及各词语的同义词进行组合, 取组合得到的 query中扩展分值 排在前 n2个的 query构成所述 Q2 , n2为预设的正整数;
query的扩展分值由创建所述复述资源中该 query的统计次数确定。
10、 根据权利要求 1至 7任一权项所述的方法, 其特征在于, 所述 步骤 C还包括:
获取所述源语言 query对应的搜索结果。
11、 根据权利要求 4、 5、 6或 7所述的方法, 其特征在于, 所述步 骤 C还包括: 从优化处理后得到的源语言 query集合中选择一个 query, 获取所述选择的 query对应的搜索结果。
12、 根据权利要求 11所述的方法, 其特征在于, 所述从优化处理后 得到的源语言 query集合中选择一个 query时, 使用的选择策略包括: 对优化处理后得到的源语言 query集合中的各 query逐一进行搜索, 直至找到搜索效果满足预设要求的 query,选择该搜索效果满足预设要求 的 query; 或者,
对优化处理后得到的源语言 query集合中的各 query进行搜索,选择 搜索效果最优的 query。
13、 根据权利要求 1所述的方法, 其特征在于, 步骤 D中所述整合 包括: 对步骤 C获取的搜索结果进行合并和去重。
14、 根据权利要求 1所述的方法, 其特征在于, 所述根据各搜索结 果在所属分类中的排次以及所属分类的排序权重, 对各搜索结果进行排 序具体包括:
利用各搜索结果在所属分类中的排次以及所属分类的排序权重, 对 各搜索结果进行打分, 按照打分结果从高到低对各搜索结果进行排序; 其中, 搜索结果 Rst的打分结果 ore(Rst)为: ore(Rst) = 丄!^ ' ^ 1 ^ , m为搜索结果分类的总个数, 为第 i种分类的排序权 m ~^ ranki (Rst) 重, ra/^(Rst)为 Rst在第 i种分类中的排次。
15、 根据权利要求 14所述的方法, 其特征在于, 搜索结果所属分类 为: 搜索结果对应的语言;
第 i种分类的排序权重的确定方法具体为:
S1、 提取所述用户输入的源语言 query的特征;
S2、 将步骤 S1提取的特征与各语言的特征向量进行相似度计算, 确定相似度超过预设的相似度阔值的语言为所述用户输入的源语言 query的映射语
S3、 对于搜索结果 Rst, 如果 Rst所属分类为映射语言, 则该所属分 类的排序权重 W为第一设定值 a;如果 Rst所属分类为源语言且源语言不 是该所属分类的映射语言, 则该所属分类的排序权重 为第二设定值 b; 如果 Rst所属分类既不是映射语言也不是源语言, 则该所属分类的排序 权重 为第三设定值 c ;
其中, a>b>c ,各语言的特征向量是预先对各语言的已有资源进行挖 掘所训练出来的。
16、 一种跨语言搜索的装置, 其特征在于, 该装置包括: 用户侧交 互单元、 翻译处理单元、 搜索处理单元和结果整合单元;
所述用户侧交互单元, 用于接收用户输入的源语言搜索请求 query, 将所述结果整合单元整合后形成的搜索结果集合提供给所述用户;
所述翻译处理单元, 用于将所述源语言 query翻译为 N种目标语言 query, N为大于 1的整数;
所述搜索处理单元, 用于分别获取所述 N种目标语言 query对应的 搜索结果;
所述结果整合单元, 用于将所述搜索处理单元获取的搜索结果进行 整合后形成最终的搜索结果集合;其中,在所述最终的搜索结果集合中, 根据各搜索结果在所属分类中的排次以及所属分类的排序权重, 对各搜 索结果进行排序。
17、 根据权利要求 16所述的装置, 其特征在于, 所述翻译处理单元 针对每一种目标语言, 将所述源语言 query对应的该种目标语言的翻译 结果中, 翻译分值最高的一种翻译结果作为目标语言 query; 翻译结果 e的翻译分值由以下因素中的至少一种确定: 翻译所使用 的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概 率。
18、 根据权利要求 16所述的装置, 其特征在于, 该装置还包括: 优 化处理单元;
所述优化处理单元, 用于对所述用户输入的源语言 query进行优化 处理后提供给所述翻译处理单元, 所述优化处理包括 query糾错处理和 query扩展处理中的任一种或组合;
所述翻译处理单元将所述优化处理单元进行优化处理后的源语言 query翻译为 N种 标语言 query。
19、 根据权利要求 18所述的装置, 其特征在于, 如果所述优化处理 仅包括 query纠错处理, 则所述优化处理单元对所述用户输入的源语言 query进行 query纠错处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q1中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^) , ^(^^为 Ql中 被翻译为 e ι=1 的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定: 翻译所 使用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组 合概率。
20、 根据权利要求 18所述的装置, 其特征在于, 如果所述优化处理 仅包括 query扩展处理, 则所述优化处理单元对所述用户输入的源语言 query进行 query扩展处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q2中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中, 翻译结果的翻译分值总和为 ^^(ψ^), ^(^^为 Q2中 被翻译为 e
=1 的翻译分值;
翻译结果 e对应的翻译分值由以下因素中的至少一种确定:翻译所使 用的翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合 概率。
21、 根据权利要求 18所述的装置, 其特征在于, 如果所述优化处理 既包括 query纠错处理又包括 query扩展处理, 则所述优化处理单元对 所述用户输入的源语言 query进行 query纠错处理和 query扩展处理后得 到包含 n个 query的源语言 query集合 Q, n为预设的正整数;
所述翻译处理单元针对每一种目标语言, 分别利用所述 Q中的各 query进行翻译,确定翻译分值总和最高的翻译结果作为目标语言 query; 其中,翻译结果的翻译分值总和为 Q中 被翻译为 e的 翻译分值;
翻译结果 e的翻译分值由以下因素中的至少一种确定:翻译所使用的 翻译语料库中翻译结果 e的统计次数以及翻译结果 e中各词的组合概率。
22、 根据权利要求 21所述的装置, 其特征在于, 所述优化处理单元 具体包括: 第一纠错模块和第一扩展模块;
所述第一纠错模块,用于对所述用户输入的源语言 query进行 query 纠错处理后得到包含 nl个 query的源语言 query集合 Ql , nl为预设的 正整数;
所述第一扩展模块, 用于将所述 Q1中的各 query分别进行 query扩 展处理, 得到包含 n个 query的源语言 query集合 Q。
23、 根据权利要求 21所述的装置, 其特征在于, 所述优化处理单元 具体包括: 第二扩展模块和第二纠错模块;
所述第二扩展模块,用于对所述用户输入的源语言 query进行 query 扩展处理后得到包含 n2个 query的源语言 query集合 Q2 , n2为预设的 正整数;
所述第二纠错模块, 用于将所述 Q2中的各 query分别进行 query纠 错处理, 得到包含 n个 query的源语言 query集合 Q。
24、 根据权利要求 21所述的装置, 其特征在于, 所述优化处理单元 具体包括: 第三纠错模块、 第三扩展模块和合并处理模块;
所述第三纠错模块,用于对所述用户输入的源语言 query进行 query 纠错处理, 得到包含 nl个 query的源语言 query集合 Q1;
所述第三扩展模块,用于对所述用户输入的源语言 query进行 query 扩展处理, 得到包含 n2个 query的源语言 query集合 Q2;
所述合并处理模块, 用于将所述 Q1和 Q2取并集后, 得到包含 n个 query的源语言 query集合 Q。
25、 根据权利要求 19或 21所述的装置, 其特征在于, 所述优化处 理单元具体利用所述用户输入的源语言 query查找纠错训练语料, 判断 纠错训练语料中是否存在与所述用户输入的源语言 query相同的错误 query,如果是,则确定与所述用户输入的源语言 query相同的错误 query 所对应的所有正确 query,从确定的所有正确 query中选择对应纠错概率 排在前 nl个的正确 query构成源语言 query集合 Q1 ; 否则, 所述 Q1中 仅包括所述用户输入的源语言 query;
其中, 所述纠错训练语料包括: 预先从搜索日志中收集的错误 query 和对应正确 query构成的 query对, 以及错误 query被纠错为对应正确 query的纠错概率。
26、 根据权利要求 20或 21所述的装置, 其特征在于, 所述优化处 理单元具体将所述用户输入的源语言 query进行分词处理, 通过查找源 语言的复述资源确定分词处理后得到的各词语的同义词, 利用分词处理 后得到的各词语及各词语的同义词进行组合, 取组合得到的 query中扩 展分值排在前 n2个的 query构成所述 Q2;
query的扩展分值由创建所述复述资源中该 query的统计次数确定。
27、 根据权利要求 16至 24任一权项所述的装置, 其特征在于, 所 述搜索处理单元, 还用于获取所述源语言 query对应的搜索结果。
28、 根据权利要求 19、 20、 21或 22任一权项所述的装置, 其特征 在于, 该装置还包括: 源语言 query集合中选择一个 query;
所述搜索处理单元, 还用于获取所述源请求选择单元选择的 query 对应的搜索结果。
29、 根据权利要求 28所述的装置, 其特征在于, 所述源请求选择单 元釆用的选择策略包括:
对所述优化处理单元进行优化处理后得到的源语言 query集合中的 各 query逐一进行搜索, 直至找到搜索效果满足预设要求的 query, 选择 该搜索效果满足预设要求的 query; 或者,
对所述优化处理单元进行优化处理后得到的源语言 query集合中的 各 query进行搜索, 选择搜索效果最优的 query。
30、 根据权利要求 16所述的装置, 其特征在于, 所述结果整合单元 包括: 合并处理模块、 去重处理模块和排序处理模块;
所述合并处理模块, 用于将所述搜索处理单元获取的搜索结果进行 合并处理;
所述去重处理单元, 用于将所述合并处理模块合并处理后的搜索结 果进行去重处理得到搜索结果集合;
所述排序处理模块, 用于在所述搜索结果集合中, 根据各搜索结果 在所属分类中的排次以及所属分类的排序权重,对各搜索结果进行排序。
31、 根据权利要求 30所述的装置, 其特征在于, 所述排序处理模块 具体利用所述搜索结果集合中各搜索结果在所属分类中的排次以及所属 分类的排序权重, 对各搜索结果进行打分, 按照打分结果从高到低对各 搜索结果进行排序;
其中, 搜索结果 Rst的打分结果 ore(Rst)为: ore(Rst) = 丄!^ ' ^ 1 ^ , m为搜索结果分类的总个数, 为第 i种分类的排序权 m ~^ ranki (Rst) 重, ra/^(Rst)为 Rst在第 i种分类中的排次。
32、 根据权利要求 31所述的装置, 其特征在于, 搜索结果所属分类 为: 搜索结果对应的语言;
所述排序处理模块在确定第 i种分类的排序权重时, 具体执行以下 操作: 提取所述用户输入的源语言 query的特征; 将提取的特征与各语 言的特征向量进行相似度计算, 确定相似度超过预设的相似度阈值的语 言为所述用户输入的源语言 query的映射语言; 对于搜索结果 Rst, 如果 Rst所属分类为映射语言, 则该所属分类的排序权重 为第一设定值 a; 如果 Rst所属分类为源语言且源语言不是该所属分类的映射语言, 则该 所属分类的排序权重 W为第二设定值 b ;如果 Rst所属分类既不是映射语 言也不是源语言, 则该所属分类的排序权重 为第三设定值 c;
其中, a>b>c ,各语言的特征向量是预先对各语言的已有资源进行挖 掘所训练出来的。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201110047892.7A CN102651003B (zh) | 2011-02-28 | 2011-02-28 | 一种跨语言搜索的方法和装置 |
| CN201110047892.7 | 2011-02-28 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2012116562A1 true WO2012116562A1 (zh) | 2012-09-07 |
Family
ID=46693011
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2011/083420 Ceased WO2012116562A1 (zh) | 2011-02-28 | 2011-12-03 | 一种跨语言搜索的方法和装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN102651003B (zh) |
| WO (1) | WO2012116562A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014204658A1 (en) * | 2013-06-17 | 2014-12-24 | Ronin Ilya | Cross-lingual e-commerce |
Families Citing this family (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103699545A (zh) * | 2012-09-28 | 2014-04-02 | 摩根全球购物有限公司 | 网络搜寻系统及其网络搜寻方法 |
| CN103268326A (zh) * | 2013-05-02 | 2013-08-28 | 百度在线网络技术(北京)有限公司 | 一种个性化的跨语言检索方法及装置 |
| US9852129B2 (en) | 2013-11-26 | 2017-12-26 | International Business Machines Corporation | Language independent processing of logs in a log analytics system |
| CN104573019B (zh) * | 2015-01-12 | 2019-04-02 | 百度在线网络技术(北京)有限公司 | 信息检索方法和装置 |
| CN105404688A (zh) * | 2015-12-11 | 2016-03-16 | 北京奇虎科技有限公司 | 搜索方法和搜索设备 |
| CN107273372A (zh) * | 2016-04-06 | 2017-10-20 | 北京搜狗科技发展有限公司 | 一种搜索方法、装置和设备 |
| CN106777261A (zh) * | 2016-12-28 | 2017-05-31 | 深圳市华傲数据技术有限公司 | 基于多源异构数据集的数据查询方法及装置 |
| CN106919642B (zh) * | 2017-01-13 | 2021-04-16 | 北京搜狗科技发展有限公司 | 一种跨语言搜索方法和装置、一种用于跨语言搜索的装置 |
| CN108304412B (zh) * | 2017-01-13 | 2022-09-30 | 北京搜狗科技发展有限公司 | 一种跨语言搜索方法和装置、一种用于跨语言搜索的装置 |
| CN111400464B (zh) * | 2019-01-03 | 2023-05-26 | 百度在线网络技术(北京)有限公司 | 一种文本生成方法、装置、服务器及存储介质 |
| CN109933724B (zh) * | 2019-03-07 | 2022-01-14 | 上海智臻智能网络科技股份有限公司 | 知识搜索方法、系统、问答装置、电子设备及存储介质 |
| CN110110171A (zh) * | 2019-05-09 | 2019-08-09 | 上海泰豪迈能能源科技有限公司 | 企业信息搜索方法、装置及电子设备 |
| CN112446222B (zh) * | 2019-08-16 | 2024-12-27 | 阿里巴巴集团控股有限公司 | 翻译优化方法、装置及处理器 |
| CN111666417B (zh) * | 2020-04-13 | 2023-06-23 | 百度在线网络技术(北京)有限公司 | 生成同义词的方法、装置、电子设备以及可读存储介质 |
| CN114896990B (zh) * | 2022-04-22 | 2025-02-14 | 北京捷通华声科技股份有限公司 | 一种语言翻译方法、电子设备以及存储介质 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101271461A (zh) * | 2007-03-19 | 2008-09-24 | 株式会社东芝 | 跨语言检索请求的转换及跨语言信息检索方法和系统 |
| CN101743544A (zh) * | 2007-05-16 | 2010-06-16 | 谷歌公司 | 跨语言信息检索 |
| CN101868797A (zh) * | 2007-09-21 | 2010-10-20 | 谷歌公司 | 跨语言搜索 |
-
2011
- 2011-02-28 CN CN201110047892.7A patent/CN102651003B/zh active Active
- 2011-12-03 WO PCT/CN2011/083420 patent/WO2012116562A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101271461A (zh) * | 2007-03-19 | 2008-09-24 | 株式会社东芝 | 跨语言检索请求的转换及跨语言信息检索方法和系统 |
| CN101743544A (zh) * | 2007-05-16 | 2010-06-16 | 谷歌公司 | 跨语言信息检索 |
| CN101868797A (zh) * | 2007-09-21 | 2010-10-20 | 谷歌公司 | 跨语言搜索 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014204658A1 (en) * | 2013-06-17 | 2014-12-24 | Ronin Ilya | Cross-lingual e-commerce |
| US9678952B2 (en) | 2013-06-17 | 2017-06-13 | Ilya Ronin | Cross-lingual E-commerce |
Also Published As
| Publication number | Publication date |
|---|---|
| CN102651003A (zh) | 2012-08-29 |
| CN102651003B (zh) | 2014-08-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN102651003B (zh) | 一种跨语言搜索的方法和装置 | |
| JP5425820B2 (ja) | ターゲットページとは異なる文字セットおよび/または言語で書かれたクエリを使用する検索のためのシステムおよび方法 | |
| CN104011712B (zh) | 对跨语言查询建议的查询翻译进行评价 | |
| CN102831246B (zh) | 藏文网页分类方法和装置 | |
| JP5379696B2 (ja) | 概念ベースの検索とランク付けを伴う情報検索のシステム、方法およびソフトウェア | |
| KR101721338B1 (ko) | 검색 엔진 및 그의 구현 방법 | |
| CN103678576B (zh) | 基于动态语义分析的全文检索系统 | |
| TWI434187B (zh) | 文字轉換方法與系統 | |
| CN102760142A (zh) | 一种针对搜索请求抽取搜索结果主题标签的方法和装置 | |
| CN102654867B (zh) | 一种跨语言搜索中的网页排序方法和系统 | |
| CN102662936B (zh) | 融合Web挖掘、多特征与有监督学习的汉英未登录词翻译方法 | |
| CN102779135B (zh) | 跨语言获取搜索资源的方法和装置及对应搜索方法和装置 | |
| JP2015523659A (ja) | 多言語混合検索方法およびシステム | |
| CN104008126A (zh) | 一种基于网页内容分类进行分词处理的方法和装置 | |
| CN103838833A (zh) | 基于相关词语语义分析的全文检索系统 | |
| CN102855263A (zh) | 一种对双语语料库进行句子对齐的方法及装置 | |
| CN101271461A (zh) | 跨语言检索请求的转换及跨语言信息检索方法和系统 | |
| CN107690634A (zh) | 自动查询模式生成 | |
| CN105912662A (zh) | 基于Coreseek的垂直搜索引擎研究与优化的方法 | |
| CN102955853A (zh) | 一种跨语言文摘的生成方法及装置 | |
| CN110134970A (zh) | 标题纠错方法和装置 | |
| Priyatam et al. | Domain specific search in indian languages | |
| Hatakeyama et al. | Statistical analysis of automatic seed word acquisition to improve harmful expression extraction in cyberbullying detection | |
| CN106708808B (zh) | 一种信息挖掘方法及装置 | |
| CN102486770B (zh) | 文字转换方法与系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 11859730 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 11859730 Country of ref document: EP Kind code of ref document: A1 |