WO2016101765A1 - 问答页面相关问题推荐方法及装置 - Google Patents

问答页面相关问题推荐方法及装置 Download PDF

Info

Publication number
WO2016101765A1
WO2016101765A1 PCT/CN2015/095853 CN2015095853W WO2016101765A1 WO 2016101765 A1 WO2016101765 A1 WO 2016101765A1 CN 2015095853 W CN2015095853 W CN 2015095853W WO 2016101765 A1 WO2016101765 A1 WO 2016101765A1
Authority
WO
WIPO (PCT)
Prior art keywords
question
browsing
questions
user
group
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/095853
Other languages
English (en)
French (fr)
Inventor
沈亮
周伟
梁任鹏
项碧波
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Qihoo Technology Co Ltd
Qizhi Software Beijing Co Ltd
Original Assignee
Beijing Qihoo Technology Co Ltd
Qizhi Software Beijing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from CN201410827521.4A external-priority patent/CN104462552B/zh
Priority claimed from CN201410828977.2A external-priority patent/CN104462554B/zh
Priority claimed from CN201410830054.0A external-priority patent/CN104462556B/zh
Priority claimed from CN201410828866.1A external-priority patent/CN104462553B/zh
Application filed by Beijing Qihoo Technology Co Ltd, Qizhi Software Beijing Co Ltd filed Critical Beijing Qihoo Technology Co Ltd
Publication of WO2016101765A1 publication Critical patent/WO2016101765A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor

Definitions

  • the present invention relates to the field of search technology, and in particular, to a method and apparatus for recommending questions related to a question and answer page.
  • the search question on the current Q&A page is: “What should I do with a cold cough?”
  • the relevant questions recommended for the user on the current Q&A page may include: “What should I do with a cold?”, “What about a cold, cough, runny nose?”, “Children’s colds” What about coughing?”, wait.
  • the search word input by the user is generally used as the core word to acquire, which is relatively simple and straightforward, but the correlation between the obtained related problem and the user input question is not very good.
  • the matching between the related problems obtained by the user and the answers to the questions that the user really wants to obtain is relatively poor, resulting in poor accuracy of the question and answer page question retrieval, and
  • the fit of the user's needs is relatively poor, and it cannot solve the search matching requirement that the user wants to view the more consistent answer to the question that is closer to the retrieved question on the current question and answer page.
  • the present invention has been made in order to provide a question and answer page related problem recommendation method and a corresponding question and answer page related problem recommendation device in a search process that overcomes the above problems or at least partially solves the above problems.
  • a question and answer page related question recommendation method comprising: acquiring at least one related question related to the search term by a database according to a search term from a user; and acquiring the obtained according to at least one preset rule The related problem is filtered; according to the screening result of the related problem, the related question recommended by the question and answer page to the user is determined.
  • a core word extraction method for a question and answer page including: extracting a core word candidate string from a question and answer page; classifying the core word candidate string, and extracting a classification feature of each candidate string segmentation word; Whether each candidate string segmentation word is a core word is filtered according to the classification feature.
  • a question and answer page related question recommendation method comprising: acquiring at least one related question related to the search term in a database according to a search term from a first user; a browsing behavior log of the second user in the segment, determining a browsing weight of the obtained related question; sorting the related related questions according to the browsing weight; and determining a recommendation in the question and answer page according to the sorting result of the related question A related question for a user.
  • a question and answer page related question recommendation method including: acquiring at least one related question related to the search term in a database according to a search term from a first user; according to a selected time period a search activity log of the second user, determining a click weight of the related question obtained; sorting related questions obtained according to the click weight; and determining a question and answer page recommendation to the first user according to the sorting result of the related question Related issues.
  • a question and answer page related problem recommendation apparatus comprising: an acquirer adapted to acquire at least one related question related to the search term by a database according to a search term from a user; a filter, The method is adapted to filter the obtained related question according to the at least one preset rule; the recommender is adapted to determine a related question recommended by the question and answer page to the user according to the screening result of the related question.
  • a question and answer page core word extracting apparatus including: a candidate string extracting module, The core word candidate module is used for extracting core word candidate strings from the question and answer page; the feature extraction module is configured to perform word segmentation on the core word candidate string, and extract the classification features of each candidate string segmentation word; the core word determining module is configured to filter each candidate according to the classification feature Whether the string word is the core word.
  • a question and answer page related problem recommendation apparatus including: a problem obtaining module, configured to acquire at least one related question related to the search term in a database according to a search term from a first user
  • the weight determination module determines the browsing weight of the related question obtained according to the browsing behavior log of the second user in the selected time period; and the ranking recommendation module is configured to sort the related questions obtained according to the browsing weight;
  • the sorting result of the related question determines the related question recommended to the first user in the question and answer page.
  • a question and answer page related problem recommendation apparatus including: a problem obtaining module, configured to acquire at least one related question related to the search term in a database according to a search term from a first user a weight determining module, configured to determine a click weight of the related question obtained according to a search behavior log of the second user in the selected time period; and a ranking recommendation module, configured to sort the related questions obtained according to the click weight According to the sorting result of the related question, the related question recommended by the question and answer page to the first user is determined.
  • a computer program comprising computer readable code, when the computer readable code is run on a computing device, causing the computing device to perform the question and answer page related question recommendation method described above .
  • a computer readable medium storing a computer program as described above is provided.
  • the question and answer page related problem recommendation method of the embodiment of the present invention at least one related problem related to the search term of the database is obtained according to the search term from the user, and the related problem acquired is filtered according to at least one preset rule, according to the screening The result determines the relevant questions that are recommended to the user. It can be seen that, according to the method for recommending questions related to the question and answer page according to the embodiment of the present invention, after obtaining related questions related to the search term, the related questions are filtered by using a preset rule, and the correlation of the search words input by the user is better reflected. The problem is to get the answer to the question the user really wants to get.
  • the obtained related questions are filtered by using at least one preset rule, that is, in this example, multiple preset rules may be used to filter related related questions.
  • multiple preset rules may be used to filter related related questions.
  • FIG. 1 is a flow chart showing a process of recommending a question related to a question and answer page according to an embodiment of the present invention
  • FIG. 2 illustrates a process flow diagram for screening related questions and recommending based on core words, in accordance with one embodiment of the present invention
  • FIG. 3 is a flowchart showing a process of screening related questions according to core words and recommending them according to another embodiment of the present invention
  • FIG. 4 is a flowchart showing a process of screening related questions and recommending according to core words according to still another embodiment of the present invention.
  • FIG. 5 is a flowchart showing a process of filtering and recommending related questions according to a browsing history log of a user according to an embodiment of the present invention
  • FIG. 6 is a flowchart showing a process of filtering and recommending related questions according to a browsing history log of a user according to another embodiment of the present invention.
  • FIG. 7 is a flowchart showing a process of filtering and recommending related questions according to a search click activity log of a user according to an embodiment of the present invention
  • FIG. 8 is a flowchart showing a process of filtering and recommending related questions according to a search click activity log of a user according to another embodiment of the present invention.
  • FIG. 9 is a schematic diagram showing a system environment for implementing a question related to a question and answer page according to an embodiment of the present invention.
  • FIG. 10 is a schematic diagram showing a process flow for screening and recommending related questions according to the above three preset rules according to a preferred embodiment of the present invention.
  • FIG. 11 is a schematic structural diagram of a question and answer page related problem recommendation device according to an embodiment of the present invention.
  • FIG. 12 is a block diagram showing the structure of a question and answer page related problem recommendation device according to a preferred embodiment of the present invention.
  • FIG. 13 is a flowchart of a core word extraction method of a question and answer page according to Embodiment 8 of the present invention.
  • FIG. 14 is a flowchart of a core word extraction method of a question and answer page according to Embodiment 9 of the present invention.
  • FIG. 15 is a flowchart of a core word extraction method of a question and answer page according to Embodiment 10 of the present invention.
  • FIG. 16 is a schematic structural diagram of a core word extracting apparatus of a question and answer page in an embodiment of the present invention.
  • FIG. 17 is a flowchart of a method for recommending a question related to a question and answer page in the eleventh embodiment of the present invention.
  • FIG. 18 is a flowchart of a method for recommending a question related to a question and answer page in the twelfth embodiment of the present invention.
  • FIG. 19 is a schematic diagram of a system environment for implementing a question related to a question and answer page in an embodiment of the present invention.
  • FIG. 20 is a schematic structural diagram of a device for recommending a question related to a question and answer page in an embodiment of the present invention
  • 21 is a flowchart of a method for recommending a question related to a question and answer page in the thirteenth embodiment of the present invention.
  • FIG. 22 is a flowchart of a method for recommending a question related to a question and answer page in the fourteenth embodiment of the present invention.
  • FIG. 23 is a schematic diagram of a system environment for implementing a question related to a question and answer page in an embodiment of the present invention
  • FIG. 24 is a schematic structural diagram of a device for recommending a question related to a question and answer page in an embodiment of the present invention.
  • 25 is a block diagram schematically showing a computing device for performing a question and answer page related question recommendation method according to the present invention.
  • Fig. 26 schematically shows a storage unit for holding or carrying program code implementing the question and answer page related question recommendation method according to the present invention.
  • FIG. 1 is a flow chart showing the processing of a question and answer page related question recommendation method according to an embodiment of the present invention. Referring to FIG. 1, the flow includes at least steps S102 to S106.
  • Step S102 Acquire at least one related problem related to the search term in the database according to the search term from the user;
  • Step S104 Filter related related questions according to at least one preset rule.
  • Step S106 Determine, according to the screening result of the related question, a related question recommended to the user by the question and answer page.
  • the question and answer page related problem recommendation method of the embodiment of the present invention at least one related problem related to the search term of the database is obtained according to the search term from the user, and the related problem acquired is filtered according to at least one preset rule, according to the screening The result determines the relevant questions that are recommended to the user. It can be seen that, according to the method for recommending questions related to the question and answer page according to the embodiment of the present invention, after obtaining related questions related to the search term, the related questions are filtered by using a preset rule, and the correlation of the search words input by the user is better reflected. The problem is to get the answer to the question the user really wants to get.
  • the obtained related questions are filtered by using at least one preset rule, that is, in this example, multiple preset rules may be used to filter related related questions.
  • multiple preset rules may be used to filter related related questions.
  • the embodiment of the present invention filters related questions related to the search term according to at least one preset rule.
  • the preset rule based on the screening of related questions may be any rule that can further filter related issues.
  • the preset rule may be to filter related questions according to the user behavior log, or to filter related questions according to the degree of fit of the search words and related questions.
  • the related problem is preferably filtered according to the following preset rules:
  • the related questions may be filtered according to only one of the above preset rules, and related questions may be filtered according to several or all of the above preset rules.
  • the related questions recommended to the user are determined based on the screening results.
  • the related questions are separately screened according to the preset rules, and then the respective screening results are fitted to obtain related questions recommended to the user, which can be seen in the basis of When multiple preset rules filter related questions, it is still necessary to perform a single preset rule to filter related issues. Therefore, in this example, the process of separately filtering related questions according to each preset rule and determining related questions recommended to the user according to the screening result is introduced.
  • the search is performed only on the basis of the search term, and the core word extracted in the search term at the time of the search is not suitable, and the question of answering the question and answer question which is more suitable for the user and cannot meet the user's demand cannot be obtained. Therefore, in this example, the question and answer page corresponding to the search term is first obtained. Second, extract the core words in the Q&A page and filter related questions based on the extracted core words.
  • FIG. 2 illustrates a process flow diagram for screening related questions and recommendations based on core words, in accordance with one embodiment of the present invention.
  • the process includes the following steps:
  • Step S201 Acquire a corresponding question and answer page and related questions according to the search term input by the user.
  • Step S202 Extract a core word candidate string from the question and answer page.
  • the core word candidate string for determining the core word is extracted from the question and answer page, and the qualified core word is selected from the candidate string.
  • the core word candidate string can be extracted from the question and answer page.
  • the core word candidate string can be extracted from the title of the question and answer page, or can be extracted from the page content of the question and answer page, or extracted from the title of the question and answer page and the page content of the question and answer page.
  • Extracting the core word candidate string from the question and answer page includes: obtaining a question and answer page corresponding to the search word input by the user; and extracting the core word candidate string from the title of the obtained question and answer page. And/or extracting a character string related to the search term input by the user from the page content of the obtained question and answer page as a core word candidate string.
  • Step S203 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • each candidate string segmentation word is divided into several candidate string segmentation words, and the classification features of the candidate string segmentation words are extracted.
  • the classification feature of the candidate string segmentation word includes at least one of the following features: noun, heat vocabulary, hyperlink, co-occurrence rate of related questions, document word frequency, and the like.
  • Step S204 Filter whether each candidate string segmentation word is a core word according to the extracted classification feature.
  • the candidate string segmentation words are classified according to the classification features, and whether each candidate segmentation word segment is a core word is determined according to the classification result.
  • the classification feature of the candidate string segmentation includes at least one of a noun, a heat vocabulary, a hyperlink, a co-occurrence rate of a related question, and a word frequency of the document, and all the nouns in the candidate string segmentation can be classified into one class.
  • the participles in the candidate vocabulary in the candidate vocabulary are classified into one class, the sub-words in the candidate categorical vocabulary that are hyperlinks are classified into one class, or all the nouns in the vocabulary in the candidate categorical categorization can also be categorized. For a class, ..., and so on.
  • the core words may be filtered according to the classification result, for example, according to the matching degree of each candidate string word in each category and the search words input by the user, or according to each candidate string word in each category The frequency of use statistics and other factors for screening, or comprehensive consideration of the above various factors for screening.
  • the usage frequency statistics value of the candidate string segmentation word includes one of the following parameters: the number of times of being searched, the number of times of being clicked, the number of times that have been used as a core word, and the number of times that have been used as a search term.
  • a database can be established to count the number of times the candidate string segmentation is searched by the user, the number of times the user clicks has been determined as the number of core words, the number of times the user has been used as a search term, and the like.
  • Step S205 The related problem acquired in step S201 is filtered by using the core word determined in step S204.
  • FIG. 3 is a flowchart showing a process of screening related questions according to a core word and recommending it according to another embodiment of the present invention. As shown in FIG. 3, the method includes the following steps:
  • Step S301 Acquire a question and answer page corresponding to the search term input by the user and related questions.
  • the user enters the search term "What should I do if the child has a cold cough?”, and the corresponding question and answer page is obtained according to the search term.
  • the Q&A page on the Q&A page has a Q&A page title, at least one question answer, and at least one related question.
  • the related question can be "What should I do if my child has a cold and cough?", "What kind of medicine is better for children with cold and cough?"
  • Step S302 Extract a core word candidate string from the title of the obtained question and answer page.
  • the extracted core word candidate string may be “how to do a child's cold cough”.
  • the core word candidate string can also be extracted from the content of the question and answer content of the question and answer page, related questions and the like.
  • Step S303 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • the extracted core word candidate string "how to deal with children's cold cough" can be divided into: “child”, “cold”, “cough”, "how to do” and other candidate word segmentation.
  • the classification feature extraction is performed on the candidate segmentation words of the word segmentation.
  • the classification features of the candidate word segmentation term "child” include: noun, etc.
  • the classification features of the two candidate string segmentation words "cold” and “cough” include: noun, It is a word in a heat vocabulary, a hyperlink, etc.
  • the classification feature of the "how to do" candidate word segmentation includes a hyperlink.
  • Step S304 classify the candidate string segmentation words according to the extracted classification features.
  • the candidate words such as “child”, “cold”, “cough” and “how to do” are classified, for example, “child”, “cold”, “cough” are nouns. Classified into one category; “cold”, “cough” are words in the heat vocabulary, classified into one category; “cold”, “cough”, “how to do” are hyperlinks, classified into one category.
  • Step S305 For each category, match each candidate string segmentation word in the category with a search term input by the user.
  • the search words input by the user are matched for each category separately.
  • each candidate string segmentation in the noun classification, the hot vocabulary classification, and the hyperlink classification is matched with the search word input by the user.
  • Step S306 Filter out the set number of candidate string words with the highest matching degree as the core word.
  • the two candidate string parts with high matching degree are selected as: “cold” and “cough”, then “cold” and “cough” are determined as core words; or 3 matches with higher matching degree are selected.
  • the candidate word segmentation is: “cold”, “cough”, "child”, then determine “cold”, “cough”, "child” as the core word.
  • Step S307 Filter related questions according to the determined core words.
  • search words, Q&A page titles and the like listed in the above embodiments are simple examples.
  • the search terms input by the user may be simpler, and the number of candidate crosswords obtained according to the Q&A page may be more.
  • the matching process may be more complicated, so that the function of the method of the present invention can be better utilized, and will not be enumerated here.
  • Step S305 and step S306 in the above second embodiment may be replaced with the screening methods disclosed in the following steps S405 and S406.
  • FIG. 4 is a flowchart showing a process for screening related questions and recommending according to a core word according to still another embodiment of the present invention. As shown in FIG. 4, the process includes the following steps:
  • Step S401 Acquire a question and answer page corresponding to the search term input by the user and related questions.
  • the user inputs the search term "What should I do if the child has a cold cough?”, according to the search term, the corresponding question and answer page is obtained, and the obtained Q&A page has the title of the question and answer page, at least one question answer, and at least one related question.
  • the answer to the question may include “choose the right cold (cough) medicine” and “Chinese medicine for cold and cough”.
  • the related question may be “What should the child do with a cold cough?”, “What kind of medicine is better for children with cold and cough? ?”And other issues.
  • Step S402 Extract a character string related to the search term input by the user from the page content of the obtained question and answer page as a core word candidate string.
  • the word segmentation input by the user is segmented, and a character string including at least one search term segmentation word is extracted from the page content of the obtained question and answer page.
  • the user-entered search term "What should I do if my child has a cold and cough?"
  • the words are "children”, “cold”, “cough”, "how to do” and other search terms.
  • the page content of the question and answer page, related questions, and the like may be extracted from the page content including "child”, “cold”, “cough", "how A string of at least one search term segmentation in the book is used as a core word candidate string.
  • the extracted core word candidate can have: “What should the child do with a cold cough?”, “Choose the right cold (cough) medicine”, “Chinese medicine for cold and cough”, “What should I do if I have a cold and cough?", "Children have a cold” What kind of medicine is better for coughing?" Wait.
  • Step S403 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • the extracted core word candidate string "how to deal with children's cold cough” is divided into words, such as "children", “cold”, “cough”, “how to do” and other candidate crosswords.
  • the extracted core word candidate string “choose the right cold (cough) medicine” is divided into words, for example, the words can be divided into: “choice”, “correct”, “cold”, “cough”, “medicine” and other candidate crosswords .
  • the extracted core word candidate string “Chinese medicine for cold and cough” is segmented. For example, the words can be divided into: “cold”, “cough”, “Chinese medicine” and other candidate crosswords.
  • the extracted core word candidate strings are segmented in turn, and are not enumerated here.
  • the classification feature extraction is performed on the candidate segmentation words of the word segmentation.
  • the classification features of the candidate word segmentation term "child” include: noun, etc.
  • the classification features of the two candidate string segmentation words "cold” and “cough” include: noun, It is a word in a heat vocabulary, a hyperlink, etc.
  • the classification features of the two candidate string words "Chinese medicine” and “medicine” include: a noun, etc.
  • the classification feature of the candidate censor word “cough” includes: a heat word Words in the table, etc.
  • the classification features of the candidate stringifier include: hyperlinks, and the like.
  • the classification feature extraction is performed on all candidate crosswords out of the word segmentation, and the classification features are not listed one by one in each of the candidate strings in the above example.
  • Step S404 classify the candidate string segmentation words according to the extracted classification features.
  • Candidate strings such as “child”, “cold”, “cough”, “how to do”, “select”, “correct”, “medicine”, “cough cough”, “Chinese medicine”, etc. according to the extracted classification features
  • Classification of word segmentation for example: “child”, “cold”, “cough”, “Chinese medicine”, “medicine” are nouns, classified into one category; “cold”, “cough”, “cough” are hot words The words in the table are classified into one category; “cold”, “cough”, and “how to do” are hyperlinks and fall into one category.
  • all candidate segmentation words for the word segmentation are classified according to the classification features.
  • the candidate strings in the above example are not listed one by one.
  • Step S405 For each classification, determine a usage frequency statistics value of each candidate string segmentation word in the classification.
  • the statistical values of the use frequency of each candidate string segment are respectively determined.
  • the usage frequency statistics value of the candidate string segmentation word may be at least one of factors such as the number of times the candidate string segmentation word is searched by the user, the number of times the user is clicked, the number of times the user has been determined as the core word, the number of times the user has been used as the search term, and the like. Factors are counted.
  • Step S406 According to the statistical value of the use frequency of each candidate string segmentation, the candidate string segmentation with the highest number of used frequency statistics is selected as the core word.
  • the three candidate string words with the highest frequency of use are selected as: “cold”, “cough”, “cough cough”, and then “cold”, “cough” and “cough cough” are determined as core words;
  • the three candidate serial words with the highest usage frequency statistics are selected as: “cold”, “cough”, “child”, and then “cold”, “cough” and “child” are determined as core words.
  • Step S407 Filter related questions according to the determined core words.
  • FIG. 5 illustrates a process flow diagram for screening and recommending related questions based on a user's browsing behavior log, in accordance with one embodiment of the present invention.
  • the process includes the following steps:
  • Step S501 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term for a question and answer search, and when the question and answer page is generated, the generated question and answer page includes but is not limited to the title of the question and answer page, at least one question answer, and at least one related question.
  • the search term input by the first user several related questions are obtained from the database, and the related questions are related questions in the question and answer question or the question and answer page in the question and answer page browsed by the second user in the database.
  • the first user refers to the current user
  • the second user refers to the historical user
  • Step S502 Determine, according to the browsing behavior log of the second user in the selected time period, the browsing weight of the related question obtained.
  • the browsing behavior log of the second user corresponding to the related problem acquired in the above step S501 is obtained from the database. Analyze the browsing behavior log to determine the browsing weight of related questions.
  • the related browsing weights may be calculated for the related questions obtained, and the relevant browsing weights of the same related problem are weighted according to the calculated related browsing weights, and the browsing weights of the related questions are obtained. .
  • the acquired related questions may also be grouped according to the set grouping conditions, and in each related problem group, the relevant browsing weights of each related problem and other related questions in the group are respectively calculated, and then the calculation results of each group are integrated. Weighting the relevant browsing weights of the same related problem appearing in each group to obtain the browsing weight of each related question.
  • the process of determining the browsing weight of the related problem is illustrated by taking the grouping according to the browsing user as an example.
  • Step S503 Sort the acquired related questions according to the determined browsing weights.
  • the related issues are sorted. For example, you can sort by browsing weights from high to low.
  • Step S504 Filter the related questions according to the sorted result of the obtained related questions, and then determine related questions recommended to the first user according to the screening result.
  • the related questions are screened, and the related questions obtained by the screening are recommended to the user.
  • the related questions of the set number of the highest browsing weights of all related questions are filtered out and recommended as the screening result to the user; or the related questions of the set number are respectively selected in the related questions corresponding to each browsing user as The screening results are recommended to the first user.
  • FIG. 6 A flow of processing for screening related questions according to a browsing history log of a user according to another embodiment of the present invention is shown in FIG. 6. Referring to Figure 6, the process includes the following steps:
  • Step S601 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term "What should the child have to do with a cold?", and generates a corresponding question and answer page according to the search term.
  • the generated question and answer page has a title of the question and answer page, at least one question answer, and at least one related question.
  • the related question can be "What should I do if I have a cold and cough?”, "What should I do if I have a cold and a fever?”, "What kind of medicine is better for children with cold and cough?", "What about children with cold and stuffy nose?”, "Baby has a cold and cough What to do”, “How to do a baby with a cold cough and runny nose”, “What kind of medicine is better for a baby's cold and cough?", “How to do a baby with a cold and stuffy nose”, “How to do a child with a cold and cough”, “How to do a child with a cold and stuffy nose” "How do children have a fever?”
  • Step S602 Group the related questions obtained according to the browsing user who browses the related question.
  • each related question group includes some or all related questions corresponding to one browsing user.
  • the browsing feature vector ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn ⁇ of each browsing user is obtained, where Ti represents a correlation. problem.
  • the attribute of the element Ti in the browsing feature vector includes at least one of the following parameters:
  • the generation time of the Q&A page the number of answers, the number of favorable comments, the number of bad reviews, the length of questions and answers, the time of user browsing, and the time of user stay.
  • Step S603 In each related problem group, calculate the relevant browsing weights of each related question in the group and other related questions in the group.
  • Time(i) is the user browsing time of a question and answer question
  • Time(i+1) is the user browsing time for other question and answer questions in the group
  • A1, a2 are empirical value constants.
  • each related question and other related questions in the group separately. For example, for the first related problem group that is the same for the browsing user, calculate “What should I do with a cold cough?”, “ What kind of medicine is better for children with cold and cough?”, “What should I do with a baby’s cold and cough?” “What kind of medicine is better for a baby’s cold and cough?”, “How to deal with a child’s cold and cough” and related weights related to other related issues in the group . Other related problem groups are also calculated.
  • the related browsing weights of each related question in the computing group and other related questions in the group include: in each related problem group, according to the browsing time of the browsing user browsing each related question, all relevant in the relevant problem grouping The problem is sorted; according to the sorting result, the related question that the browsing time interval is less than the preset time interval threshold is divided into the same conversation group; in each conversation group, the related browsing weights of each related problem in the group and other related questions in the group are calculated. .
  • different session groups can be further divided according to the browsing time, and the browsing time difference of the related questions in the same conversation group is less than or equal to a set time threshold.
  • the session can be divided according to the browsing feature vector of the browsing user. Calculate the browsing weight of related questions in the same session.
  • step S604 the related browsing weights calculated in the related problem groups of the same related problem are obtained, and the obtained related browsing weights are weighted to obtain the browsing weights of the obtained related questions.
  • the same related questions in each related problem group are extracted. For example, what is the related problem of "What is the child's cold and stuffy nose?" The first related problem grouping and the related browsing weights calculated in the third related question are weighted.
  • the related browsing weights calculated by the same related problem in different related problem groupings may be directly added, or may be added after multiplying the corresponding weighting coefficients respectively, or may be weighted by other weighting rules. deal with.
  • Step S605 Sort the acquired related questions according to the determined browsing weights of the related questions.
  • Step S606 Filter the related questions according to the sorted result of the obtained related questions, and then determine related questions recommended to the first user according to the screening result.
  • the first few questions with the highest browsing weight are selected as the screening result and recommended to the first user, and added to the question and answer page generated according to the search term input by the user.
  • the search click behaviors of a plurality of historical users are analyzed, and related questions are filtered according to the analysis results, and related problems that are better matched with the answers of the questions that the user really wants to obtain are obtained.
  • FIG. 7 illustrates a process flow diagram for screening and recommending related questions based on a user's search click activity log, in accordance with one embodiment of the present invention.
  • the process includes the following steps:
  • Step S701 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term for a question and answer search, and when the question and answer page is generated, the generated question and answer page includes but is not limited to the title of the question and answer page, at least one question answer, and at least one related question.
  • the search term input by the first user several related questions are obtained from the database, and the related questions are related questions in the question and answer question or the question and answer page in the question and answer page of the second user search click in the database.
  • the first user refers to the current user
  • the second user refers to the historical user
  • Step S702 Determine, according to the search behavior log of the second user in the selected time period, the click weight of the acquired related question.
  • the search behavior log of the second user corresponding to the related problem acquired in the above step S701 is obtained from the database. Analyze the search behavior log to determine the click weight of the relevant question. In the process of determining the weight of the hit, the relevant click weights of each other may be calculated for the related questions obtained, and the relevant click weights of the same related question are weighted according to the calculated relevant click weights, and the click weights of the related questions are obtained. .
  • the acquired related questions may also be grouped according to the set grouping conditions, and in each related problem group, the relevant click weights of each related problem and other related questions in the group are respectively calculated, and then the calculation results of each group are integrated. The weights of the relevant click weights of the same related problem appearing in each group are weighted to obtain the click weights of each related question.
  • the process of determining the click weight of the related problem is illustrated by taking the grouping according to the query request string as an example.
  • Step S703 Sort the acquired related questions according to the determined click weight of the related question.
  • the related questions are sorted. For example, you can sort by the order of the click weights from high to low.
  • Step S704 Filter the related questions according to the sorted result of the obtained related questions, and then determine related questions recommended to the first user according to the screening result.
  • the related questions are screened, and the related questions obtained by the screening are recommended to the first user.
  • the related questions of the set weights with the highest click weight among all related questions are filtered out and recommended as the screening result to the first user; or the set number is separately selected in the related questions corresponding to each query request string. Relevant questions are recommended as a screening result to the first user.
  • FIG. 8 is a flowchart showing a process of filtering and recommending related questions according to a search click history log of a user according to another embodiment of the present invention. Referring to Figure 8, the process includes the following steps:
  • Step S801 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term "What should the child have to do with a cold?”, and generates a corresponding question and answer page according to the search term.
  • the generated question and answer page has a title of the question and answer page, at least one question answer, and at least one related question.
  • the related question can be "What should I do if I have a cold and cough?”, "What should I do if I have a cold and a fever?", "What kind of medicine is better for children with cold and cough?", "What about children with cold and stuffy nose?”, "Baby has a cold and cough What should I do?
  • Step S802 Group the related questions obtained according to the query request string corresponding to the obtained related question.
  • each related problem group includes some or all related questions corresponding to one query request string.
  • the click feature vector ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn ⁇ of each query request string is obtained, where Ti represents a related problem. .
  • the related problems of the acquisition are grouped.
  • the attribute of the element Ti in the click feature vector includes at least one of the following parameters: the generation time of the question and answer page, the number of answers, the number of praises, the number of bad reviews, the length of the question and answer, the number of impressions, the number of clicks, and the like.
  • the query request string is "baby cold”. As a group;
  • the query request string is "Children's cold” and is grouped into one group;
  • the query request string is "Fever and fever” and is grouped into one group;
  • the query request string is "slow nose and thiophene" and is grouped into one group;
  • Step S803 In each related problem group, calculate the relevant click weights of each related question in the group and other related questions in the group.
  • the relevant click weight of 1 is W(Ti, Ti+I):
  • Ti+I indicates other question and answer questions included in the click feature vector
  • Query Request String) indicates the probability of obtaining Ti when using the query request string
  • Query Request String) indicates the probability of getting Ti+I when using the query request string.
  • each related question and other related questions in the group separately. For example, group the related questions for the query request string as “pediatric cold” and calculate “What should I do with a cold cough?” "What should I do if my child has a cold?", "What kind of medicine is better for children with cold and cough?”, "What about children with cold and stuffy nose?” Relevant click weights for other related issues. Other related problem groups are also calculated.
  • step S804 the relevant click weights calculated in the related problem groups of the same related problem are obtained, and the obtained related click weights are weighted to obtain the click weights of the obtained related questions.
  • the same related questions in each related problem group are extracted. For example, for the related question "What should I do with a cold cough?" Grouping related questions for "pediatric colds" and weighting the relevant click weights calculated in the group of related questions in which the query request string is "cold cough”.
  • the relevant click weights calculated by the same related problem in different related problem groups may be directly added, or may be added after multiplying the corresponding weight coefficients, or may be weighted by other weighting rules. deal with.
  • Step S805 Sort the acquired related questions according to the determined click weights of the related questions.
  • Step S806 Filter the related questions according to the sorted result of the obtained related questions, and then determine related questions recommended to the first user according to the screening result.
  • the first few questions with the highest click weight are selected as the screening result and recommended to the first user, and added to the question and answer page generated according to the search term input by the user.
  • the method for filtering and/or recommending related questions is a log and/or a search click activity log, and a system environment for implementing the question related to the question and answer page is shown in FIG. 9 .
  • the system includes a database, and stores related questions of a plurality of second users (historical users).
  • the question and answer page problem recommendation device can acquire the search words input by the first user, and obtain a number of historical user browsing and/or search clicks from the database according to the search words. Relevant problems and historical data of related problems are recommended to the first user by analyzing and processing the historical data to achieve better related problems.
  • FIG. 10 is a schematic diagram showing a process flow for screening and recommending related questions according to the above three preset rules according to a preferred embodiment of the present invention. Referring to Figure 10, the process includes the following steps:
  • Step S1001 Acquire a related question corresponding to the search term input by the user.
  • the user inputs a search term "What to do with a flu in a child", and obtains a corresponding related question based on the search term.
  • related issues obtained include:
  • Step S1002 screening related questions according to the core words.
  • Step S1003 Filter related questions according to the browsing behavior log of the user.
  • the obtained screening result is:
  • Step S1004 Filter related questions according to the user's search click activity log.
  • step S1001 The calculation of the search click weight value is performed on each related question mentioned in step S1001, and each related problem is sorted according to the obtained search click weight value, and the sort result is obtained as follows:
  • the screening result is:
  • Step S1005 Determine related questions recommended to the user according to the respective screening results obtained in step S1002, step S1003, and step S1004.
  • each screening result obtained in step S1002, step S1003, and step S1004 may be sorted and sorted.
  • the three screening results obtained include the related question "What should I do if I have a cold and cough?"
  • two of the three screening results obtained include "what is the cause of a flu in a child" and "how to stop coughing.” If the questions recommended to the user in the Q&A page can be:
  • each screening result mentioned in the above example, and/or the related problem of determining the recommendation in step S1005 are examples, and cannot represent the screening result obtained in the actual application and/or the related problem of determining the recommendation.
  • the embodiment of the present invention further provides a question and answer page related problem recommendation device.
  • the structure of the device is as shown in FIG. 11 , and includes an acquirer 1110 , a filter 1120 , and a recommender 1130 .
  • the acquirer 1110 is adapted to acquire at least one related problem related to the search term in the database according to the search term from the user;
  • the filter 1120 is coupled to the acquirer 1110, and is adapted to filter related related questions according to at least one preset rule;
  • the recommender 1130 is coupled to the filter 1120 and is adapted to determine a related question recommended by the question and answer page to the user according to the screening result of the related question.
  • FIG. 12 is a block diagram showing the structure of a question and answer page related question recommending apparatus according to a preferred embodiment of the present invention.
  • the filter 1120 further includes:
  • the first screening module 1121 is respectively coupled to the acquirer 1110 and the recommender 1130, and is adapted to filter related issues according to the browsing behavior log of the user;
  • the second screening module 1122 is coupled to the acquirer 1110 and the recommender 1130 respectively, and is adapted to filter related issues according to the user's search click activity log;
  • the third screening module 1123 is coupled to the acquirer 1110 and the recommender 1130 respectively, and is adapted to filter related questions according to the core words.
  • the third screening module 1123 further includes:
  • the obtaining unit 11231 is adapted to obtain a question and answer page corresponding to the search term
  • the extracting unit 11232 is coupled to the extracting unit 11231, and is adapted to extract a core word in the question and answer page;
  • the determining unit 11233 is coupled to the extracting unit 11232 and is adapted to filter related questions according to the core words.
  • the extraction unit 11232 is further adapted to:
  • each candidate string segmentation word is a core word is filtered according to the classification feature.
  • the extraction unit 11232 is further adapted to:
  • a string related to the search term is extracted as a core word candidate string.
  • the extraction unit 11232 is further adapted to:
  • a string containing at least one search term segmentation word is extracted from the page content of the question and answer page.
  • the extraction unit 11232 is further adapted to:
  • the classification feature includes at least one of the following features: noun, heat vocabulary, hyperlink, co-occurrence rate of related questions, and word frequency of the document.
  • the extraction unit 11232 is further adapted to:
  • the candidate number segmentation with the highest number of used frequency statistics is selected as the core word; wherein the usage frequency statistics of the candidate string segmentation include One of the following parameters: number of searches, number of clicks, number of times as a core word, number of times as a search term.
  • the first screening module 1121 further includes:
  • the first weight determining unit 11211 is adapted to determine a browsing weight of the related question acquired according to the browsing behavior log of the user in the selected time period;
  • the first sorting unit 11212 is coupled to the weight determining unit 11211, and is adapted to sort the acquired related questions according to the browsing weight;
  • the first screening unit 11213 is coupled to the sorting unit 11212 and is adapted to filter related questions according to the sorting result.
  • the first screening unit 11213 is further adapted to: extract a first predetermined number of related questions according to the sorting result.
  • the first weight determining unit 11211 is further adapted to:
  • each related problem group includes a part or all related questions corresponding to one browsing user;
  • the first weight determining unit 11211 is further adapted to:
  • the browsing feature vectors ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn ⁇ of each browsing user are obtained, where Ti represents a related problem.
  • the first weight determining unit 11211 is further adapted to:
  • each related problem group all related questions in the related problem group are sorted according to the browsing time of the browsing user browsing each related question;
  • the related problem that the browsing time interval is less than the preset time interval threshold is divided into the same session group
  • the relevant browsing weights for each related question in the group and other related questions in the group are calculated.
  • the second screening module 1122 further includes:
  • the second weight determining unit 11221 is adapted to determine a click weight of the related question obtained according to the search click log of the user in the selected time period;
  • the second sorting unit 11222 is coupled to the second weight determining unit 11221, and is adapted to sort the acquired related questions according to the click weight;
  • the second screening unit 11223 is coupled to the second sorting unit 11222 and is adapted to filter related questions according to the sorting result.
  • the second weight determining unit 11221 is further adapted to:
  • each related problem group includes some or all related problems corresponding to one query request string;
  • the second weight determining unit 11221 is further adapted to:
  • the click feature vectors ⁇ T1, T2, ..., Tn ⁇ of each query request string are obtained, and the related problems acquired are grouped; wherein Ti represents a related problem.
  • the second weight determining unit 11221 is further adapted to:
  • the attribute of the element Ti in the obtained click feature vector includes at least one of the following parameters:
  • the generation time of the Q&A page the number of answers, the number of favorable comments, the number of bad reviews, the length of questions and answers, the number of impressions, and the number of clicks.
  • the embodiments of the present invention can achieve the following beneficial effects:
  • the question and answer page related problem recommendation method of the embodiment of the present invention at least one related problem related to the search term of the database is obtained according to the search term from the user, and the related problem acquired is filtered according to at least one preset rule, according to the screening The result determines the relevant questions that are recommended to the user. It can be seen that, according to the method for recommending questions related to the question and answer page according to the embodiment of the present invention, after obtaining related questions related to the search term, the related questions are filtered by using a preset rule, and the correlation of the search words input by the user is better reflected. The problem is to get the answer to the question the user really wants to get.
  • the obtained related questions are filtered by using at least one preset rule, that is, in this example, multiple preset rules may be used to filter related related questions.
  • multiple preset rules may be used to filter related related questions.
  • an embodiment of the present invention provides a core word extraction method for a question and answer page.
  • the core word extraction method of the question and answer page provided in the eighth embodiment of the present invention, the flow of which is shown in FIG. 13 and includes the following steps:
  • Step S1301 Extract a core word candidate string from the question and answer page.
  • the core word candidate string for determining the core word is extracted from the question and answer page, and the qualified core word is selected from the candidate string.
  • the core word candidate string can be extracted from the question and answer page.
  • the core word candidate string can be extracted from the title of the question and answer page, or can be extracted from the page content of the question and answer page, or extracted from the title of the question and answer page and the page content of the question and answer page.
  • Extracting the core word candidate string from the question and answer page includes: obtaining a question and answer page corresponding to the search word input by the user; and extracting the core word candidate string from the title of the obtained question and answer page. And/or extracting a character string related to the search term input by the user from the page content of the obtained question and answer page as a core word candidate string.
  • Step S1302 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • each candidate string segmentation word is divided into several candidate string segmentation words, and the classification features of the candidate string segmentation words are extracted.
  • the classification feature of the candidate string segmentation word includes at least one of the following features: noun, heat vocabulary, hyperlink, co-occurrence rate of related questions, document word frequency, and the like.
  • Step S1303 Filter whether each candidate string segmentation word is a core word according to the extracted classification feature.
  • the candidate string segmentation words are classified according to the classification features, and whether each candidate segmentation word segment is a core word is determined according to the classification result.
  • the classification feature of the candidate string segmentation includes at least one of a noun, a heat vocabulary, a hyperlink, a co-occurrence rate of a related question, and a word frequency of the document, and all the nouns in the candidate string segmentation can be classified into one class.
  • the participles in the candidate vocabulary in the candidate vocabulary are classified into one class, the sub-words in the candidate categorical vocabulary that are hyperlinks are classified into one class, or all the nouns in the vocabulary in the candidate categorical categorization can also be categorized. For a class, ..., and so on.
  • the core words may be filtered according to the classification result, for example, according to the matching degree of each candidate string word in each category and the search words input by the user, or according to each candidate string word in each category The frequency of use statistics and other factors for screening, or comprehensive consideration of the above various factors for screening.
  • the usage frequency statistics value of the candidate string segmentation word includes one of the following parameters: the number of times of being searched, the number of times of being clicked, the number of times that have been used as a core word, and the number of times that have been used as a search term.
  • a database can be established to count the number of times the candidate string segmentation is searched by the user, the number of times the user clicks has been determined as the number of core words, the number of times the user has been used as a search term, and the like.
  • the core word extraction method of the question and answer page provided in the ninth embodiment of the present invention describes a specific implementation manner of the core word extraction.
  • the flow is shown in FIG. 14 and includes the following steps:
  • Step S1401 Acquire a question and answer page corresponding to the search term input by the user.
  • the user inputs the search term "What should I do if the child has a cold cough?”, according to the search term, the corresponding question and answer page is obtained, and the obtained Q&A page has the title of the question and answer page, at least one question answer, and at least one related question.
  • the related question can be "What should I do if I have a cold and cough?”, "What kind of medicine is better for children with cold and cough?"
  • Step S1402 Extract the core word candidate string from the title of the obtained question and answer page.
  • the extracted core word candidate string may be “how to do a child's cold cough”.
  • the core word candidate string can also be extracted from the content of the question and answer content of the question and answer page, related questions and the like.
  • Step S1403 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • the extracted core word candidate string "how to deal with children's cold cough" can be divided into: “child”, “cold”, “cough”, "how to do” and other candidate word segmentation.
  • the classification feature extraction is performed on the candidate segmentation words of the word segmentation.
  • the classification features of the candidate word segmentation term "child” include: noun, etc.
  • the classification features of the two candidate string segmentation words "cold” and “cough” include: noun, It is a word in a heat vocabulary, a hyperlink, etc.
  • the classification feature of the "how to do" candidate word segmentation includes a hyperlink.
  • Step S1404 classify the candidate string segmentation words according to the extracted classification features.
  • Candidates such as “child”, “cold”, “cough”, “how to do”, etc. based on the extracted classification features
  • Classification of words such as: “child”, “cold”, “cough” are nouns, classified into one category; “cold”, “cough” are words in the heat vocabulary, classified as one; Colds, "coughs,” and “how to do” are hyperlinks that fall into one category.
  • Step S1405 For each category, each candidate string word in the category is matched with the search word input by the user.
  • the search words input by the user are matched for each category separately.
  • each candidate string segmentation in the noun classification, the hot vocabulary classification, and the hyperlink classification is matched with the search word input by the user.
  • Step S1406 Filter out the set number of candidate crosswords with the highest matching degree as the core word.
  • the two candidate string parts with high matching degree are selected as: “cold” and “cough”, then “cold” and “cough” are determined as core words; or 3 matches with higher matching degree are selected.
  • the candidate word segmentation is: “cold”, “cough”, "child”, then determine “cold”, “cough”, "child” as the core word.
  • search words, Q&A page titles and the like listed in the above embodiments are simple examples.
  • the search terms input by the user may be simpler, and the number of candidate crosswords obtained according to the Q&A page may be more.
  • the matching process may be more complicated, so that the function of the method of the present invention can be better utilized, and will not be enumerated here.
  • step S1405 and step S1406 realize whether or not each candidate string participle is a core word according to the classification result.
  • Steps S1405 and S1406 in the above-described Embodiment 9 may be replaced with the screening methods disclosed in the following steps S1505 and S1506.
  • the core word extraction method of the question and answer page provided by the tenth embodiment of the present invention describes another specific implementation manner of the core word extraction.
  • the flow is shown in FIG. 15 and includes the following steps:
  • Step S1501 Acquire a question and answer page corresponding to the search term input by the user.
  • the user inputs the search term "What should I do if the child has a cold cough?”
  • the corresponding question and answer page is obtained, and the obtained Q&A page has the title of the question and answer page, at least one question answer, and at least one related question.
  • the answer to the question may include “choose the right cold (cough) medicine”, “Chinese medicine for cold and cough”, and the related question may be “What should I do with a cold and cough in children?”, “What kind of medicine is better for children with cold and cough? ?”And other issues.
  • Step S1502 Extract a character string related to the search term input by the user from the page content of the obtained question and answer page as a core word candidate string.
  • the word segmentation input by the user is segmented, and a character string including at least one search term segmentation word is extracted from the page content of the obtained question and answer page.
  • the user-entered search term "What should I do if my child has a cold and cough?" can be divided into words such as "children”, “cold”, “cough”, "how to do” and other search terms.
  • the page content of the question and answer page, related questions, and the like may be extracted from the page content including "child”, “cold”, “cough", "how A string of at least one search term segmentation in the book is used as a core word candidate string.
  • the extracted core word candidate can have: “What should the child do with a cold cough?”, “Choose the right cold (cough) medicine”, “Chinese medicine for cold and cough”, “What should I do if I have a cold and cough?", "Children have a cold” What kind of medicine is better for coughing?" Wait.
  • Step S1503 Perform word segmentation on the extracted core word candidate strings, and extract classification features of each candidate string segmentation word.
  • the extracted core word candidate string "how to deal with children's cold cough” is divided into words, such as "children", “cold”, “cough”, “how to do” and other candidate crosswords.
  • the extracted core word candidate string “choose the right cold (cough) medicine” is divided into words, for example, the words can be divided into: “choice”, “correct”, “cold”, “cough”, “medicine” and other candidate crosswords .
  • the extracted core word candidate string “Chinese medicine for cold and cough” is segmented. For example, the words can be divided into: “cold”, “cough”, “Chinese medicine” and other candidate crosswords.
  • the extracted core word candidate strings are segmented in turn, and are not enumerated here.
  • the classification feature extraction is performed on the candidate segmentation words of the word segmentation.
  • the classification features of the candidate word segmentation term "child” include: noun, etc.
  • the classification features of the two candidate string segmentation words "cold” and “cough” include: noun, It is a word in a heat vocabulary, a hyperlink, etc.
  • the classification features of the two candidate string words "Chinese medicine” and “medicine” include: a noun, etc.
  • the classification feature of the candidate censor word “cough” includes: a heat word Words in the table, etc.
  • "What to do” is a classification of the candidate stringifiers.
  • the levy includes: hyperlinks, etc.
  • the classification feature extraction is performed on all candidate crosswords out of the word segmentation, and the classification features are not listed one by one in each of the candidate strings in the above example.
  • Step S1504 classify the candidate string segmentation words according to the extracted classification features.
  • Candidate strings such as “child”, “cold”, “cough”, “how to do”, “select”, “correct”, “medicine”, “cough cough”, “Chinese medicine”, etc. according to the extracted classification features
  • Classification of word segmentation for example: “child”, “cold”, “cough”, “Chinese medicine”, “medicine” are nouns, classified into one category; “cold”, “cough”, “cough” are hot words The words in the table are classified into one category; “cold”, “cough”, and “how to do” are hyperlinks and fall into one category.
  • all candidate segmentation words for the word segmentation are classified according to the classification features.
  • the candidate strings in the above example are not listed one by one.
  • Step S1505 For each classification, determine a statistical value of the usage frequency of each candidate string segmentation word in the classification.
  • the statistical values of the use frequency of each candidate string segment are respectively determined.
  • the usage frequency statistics value of the candidate string segmentation word may be at least one of factors such as the number of times the candidate string segmentation word is searched by the user, the number of times the user is clicked, the number of times the user has been determined as the core word, the number of times the user has been used as the search term, and the like. Factors are counted.
  • Step S1506 According to the statistical value of the use frequency of each candidate string segmentation, the candidate string segmentation with the highest number of used frequency statistics is selected as the core word.
  • the three candidate string words with the highest frequency of use are selected as: “cold”, “cough”, “cough cough”, and then “cold”, “cough” and “cough cough” are determined as core words;
  • the three candidate serial words with the highest usage frequency statistics are selected as: “cold”, “cough”, “child”, and then “cold”, “cough” and “child” are determined as core words.
  • step S1505 and step S1506 are implemented to determine whether each candidate string segmentation word is a core word according to the classification result.
  • Steps S1505 and S1506 in the above third embodiment may be replaced with the screening methods disclosed in step S1405 and step S1406.
  • the embodiment of the present invention further provides a question and answer page core word extracting device.
  • the device has the structure shown in FIG. 16 and includes a candidate string extracting module 1601, a feature extracting module 1602, and a core word determining module 1603.
  • the candidate string extraction module 1601 is configured to extract a core word candidate string from the question and answer page.
  • the feature extraction module 1602 is configured to perform segmentation on the core word candidate string, and extract classification features of each candidate string segmentation word.
  • the core word determining module 1603 is configured to filter, according to the extracted classification feature, whether each candidate string segmentation word is a core word.
  • the candidate string extraction module 1601 is configured to obtain a Q&A page corresponding to the search term input by the user, extract a core word candidate string from the title of the obtained Q&A page, and/or a page content from the obtained Q&A page. In the middle, a character string related to the search term input by the user is extracted as a core word candidate string.
  • the candidate string extraction module 1601 is specifically configured to perform word segmentation on the search term, and extract a character string including at least one search term segmentation from the page content of the obtained question and answer page.
  • the core word determining module 1603 is configured to classify the candidate string segmentation words according to the extracted classification features, and determine, according to the classification result, whether each candidate string segmentation word is a core word; wherein the classification feature includes at least one of the following features: : nouns, heat vocabulary, hyperlinks, co-occurrence rates of related questions, document frequency.
  • the core word determining module 1603 is configured to: for each category, match each candidate string segmentation word in the category with a search term input by the user, and filter out the set number of candidate string segmentation words with the highest matching degree as Core words; or for each classification, according to the statistical value of the use frequency of each candidate string segmentation in the classification, screening the candidate number segmentation words with the highest number of used frequency statistics as the core word; wherein, the use of the candidate string word segmentation
  • the frequency statistics include one of the following parameters: number of searches, number of clicks, number of times the core word was used, and number of times the search term was used.
  • the core word extraction method and apparatus for the above-mentioned question and answer page provided by the embodiment of the present invention can extract a core word that is more in line with the user's search requirement according to the question and answer page corresponding to the search word input by the user, so that the search word input by the user can be obtained according to the core word.
  • Relevant questions with higher relevance, in the current question and answer page provide users with relevant issues that are more conformable to user needs and more in line with user needs, and improve the accuracy of question and answer page retrieval.
  • the method and apparatus for extracting core words of the question and answer page provided by the above embodiment, extracting core word candidate strings from the question and answer page,
  • the extracted core word candidate string is segmented, the classification features of each candidate string segmentation word are extracted, and each candidate string segmentation word is selected as a core word according to the classification feature.
  • the program extracts the core word from the analysis of the question and answer page, so that the determined
  • the core words can better reflect the problems input by the users, and have higher relevance to the questions input by the users, so that they can obtain more questions and user needs according to the extracted core words, and more satisfy the user's needs of the question and answer questions, and obtain the users really want to obtain
  • the answer to the question improves the accuracy of the Q&A page search.
  • the core word can be extracted according to the title or page content of the question and answer page corresponding to the search word input by the user, so that the extraction of the core word can be more accurate and more suitable for the user.
  • the candidate word classification features can be comprehensively considered, and the core words are determined according to the comprehensive considerations of different categories, so that the appropriate core words can be determined more objectively and reasonably.
  • the embodiment of the present invention provides a question and answer page related question recommendation.
  • the method by analyzing the browsing behavior of a number of historical users, obtains a problem that is more relevant to the answer matching of the questions that the user really wants to obtain.
  • a first embodiment of the present invention provides a method for recommending a question related to a question and answer page.
  • the method is as shown in FIG. 17 and includes the following steps:
  • Step S1701 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term for a question and answer search, and when the question and answer page is generated, the generated question and answer page includes but is not limited to the title of the question and answer page, at least one question answer, and at least one related question.
  • the search term input by the first user several related questions are obtained from the database, and the related questions are related questions in the question and answer question or the question and answer page in the question and answer page browsed by the second user in the database.
  • the first user refers to the current user
  • the second user refers to the historical user
  • Step S1702 Determine the browsing weight of the related question obtained according to the browsing behavior log of the second user in the selected time period.
  • the browsing behavior log of the second user corresponding to the related problem obtained in the above step S1701 is obtained from the database. Analyze the browsing behavior log to determine the browsing weight of related questions.
  • the related browsing weights may be calculated for the related questions obtained, and the relevant browsing weights of the same related problem are weighted according to the calculated related browsing weights, and the browsing weights of the related questions are obtained. .
  • the acquired related questions may also be grouped according to the set grouping conditions, and in each related problem group, the relevant browsing weights of each related problem and other related questions in the group are respectively calculated, and then the calculation results of each group are integrated. Weighting the relevant browsing weights of the same related problem appearing in each group to obtain the browsing weight of each related question.
  • Step S1703 Sort the acquired related questions according to the determined browsing weights.
  • the related issues are sorted. For example, you can sort by browsing weights from high to low.
  • Step S1704 Determine a related question recommended to the first user in the question and answer page according to the sorted result of the obtained related question.
  • the sorting result of related questions select relevant questions to recommend to the user. For example, the related question of obtaining the set number of the highest browsing weight among all related questions is recommended to the user; or the related question of obtaining the set quantity in each related question of each browsing user is recommended to the first user.
  • a twelfth embodiment of the present invention provides a method for recommending a question related to a question and answer page.
  • the process of the method is as shown in FIG. 18, and includes the following steps:
  • Step S1801 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term "What should the child have to do with a cold?”, and generates a corresponding question and answer page according to the search term.
  • the generated question and answer page has a title of the question and answer page, at least one question answer, and at least one related question.
  • the related question can be "What should I do if I have a cold and cough?”, "What should I do if I have a cold and a fever?", "What kind of medicine is better for children with cold and cough?", "What about children with cold and stuffy nose?”, "Baby has a cold and cough What should I do?
  • Step S1802 group the related questions obtained according to the browsing user who browses the related question.
  • each related question group includes some or all related questions corresponding to one browsing user.
  • the browsing feature vector ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn ⁇ of each browsing user is obtained, where Ti represents a correlation. problem.
  • the attribute of the element Ti in the browsing feature vector includes at least one of the following parameters:
  • the generation time of the Q&A page the number of answers, the number of favorable comments, the number of bad reviews, the length of questions and answers, the time of user browsing, and the time of user stay.
  • Step S1803 In each related problem group, calculate the relevant browsing weights of each related question in the group and other related questions in the group.
  • Time(i) is the user browsing time of a question and answer question
  • Time(i+1) is the user browsing time for other question and answer questions in the group
  • A1, a2 are empirical value constants.
  • each related question and other related questions in the group separately. For example, for the first related problem group that is the same for the browsing user, calculate “What should I do with a cold cough?”, “ What kind of medicine is better for children with cold and cough?”, “What should I do with a baby’s cold and cough?” “What kind of medicine is better for a baby’s cold and cough?”, “How to deal with a child’s cold and cough” and related weights related to other related issues in the group . Other related problem groups are also calculated.
  • the related browsing weights of each related question in the computing group and other related questions in the group include: in each related problem group, according to the browsing time of the browsing user browsing each related question, all relevant in the relevant problem grouping Sort the problem; according to the sorting result, the related question of dividing the browsing time interval is smaller than the preset time interval threshold to the same meeting Groups; in each conversation group, calculate the relevant browsing weights of each related question in the group and other related questions in the group.
  • different session groups can be further divided according to the browsing time, and the browsing time difference of the related questions in the same conversation group is less than or equal to a set time threshold.
  • the session can be divided according to the browsing feature vector of the browsing user. Calculate the browsing weight of related questions in the same session.
  • Step S1804 Acquire related browsing weights calculated by the related related questions in each related problem group, and weight the obtained related browsing weights to obtain the browsing weights of the obtained related questions.
  • the same related questions in each related problem group are extracted. For example, what is the related problem of "What is the child's cold and stuffy nose?" The first related problem grouping and the related browsing weights calculated in the third related question are weighted.
  • the related browsing weights calculated by the same related problem in different related problem groupings may be directly added, or may be added after multiplying the corresponding weighting coefficients respectively, or may be weighted by other weighting rules. deal with.
  • Step S1805 Sort the acquired related questions according to the determined browsing weights of the related questions.
  • Step S1806 Determine a related question recommended to the first user in the question and answer page according to the sorted result of the obtained related question.
  • the first few questions with the highest browsing weight are recommended as related questions to the first user, and are added to the question and answer page generated according to the search words input by the user.
  • the above method analyzes the browsing behavior of the historical user browsing each related question according to the historical data in the database, determines the browsing weight parameter of the related question, and determines the recommendation priority of recommending the related problem to the user, thereby obtaining the search term input by the user.
  • Relevant questions with higher matching degree provide users with relevant issues that are more conformable to user needs and more in line with user needs in the current question and answer page, and improve the accuracy of question and answer page question retrieval.
  • the system includes a database, and stores related questions of a plurality of second users (historical users).
  • the question and answer page problem recommendation device can obtain the search words input by the first user, and obtain related questions and related questions that the historical users have browsed from the database according to the search words.
  • the historical data of the problem is recommended to the first user by analyzing and processing the historical data to achieve better related problems.
  • the embodiment of the present invention further provides a question and answer page related problem recommendation device.
  • the structure of the device is as shown in FIG. 20, and includes: a problem obtaining module 2001, a weight determining module 2002, and a sorting recommending module 2003.
  • the problem obtaining module 2001 is configured to acquire at least one related problem related to the search term in the database according to the search term from the first user.
  • the weight determination module 2002 determines the browsing weight of the related question obtained according to the browsing behavior log of the second user in the selected time period.
  • the ranking recommendation module 2003 is configured to sort related related questions according to the determined browsing weights, and determine related questions recommended to the first user in the question and answer page according to the sorted results of the obtained related questions.
  • the weight determination module 2002 includes the problem grouper 20021, the correlation weight calculator 20022, and the browsing weight calculator 20023.
  • the problem grouper 20021 is configured to group the related questions obtained according to the browsing user who browses the related question; wherein each related problem group includes a part or all related questions corresponding to one browsing user.
  • Correlation weight calculator 20022 used to calculate related problems in the group and other related groups in the group in each related problem group The relevant browsing weight of the question.
  • the browsing weight calculator 20023 is configured to obtain related browsing weights calculated in each related problem group by the same related problem, and weight the obtained related browsing weights to obtain the browsing weights of the obtained related questions.
  • the problem grouper 20021 is specifically configured to obtain the browsing feature vector ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn of each browsing user according to the browsing behavior log in the selected time period. ⁇ , to achieve the grouping of related problems acquired; where Ti represents a related problem.
  • the correlation weight calculator 20022 is specifically configured to sort all related questions in the related problem group according to the browsing time of the browsing user to browse each related question in each related problem group; according to the sorting result, the browsing is divided.
  • the related problem that the time interval is less than the preset time interval threshold is to the same session group; in each session group, the relevant browsing weights of each related question in the group and other related questions in the group are calculated.
  • the above problem grouper 20021 specifically uses the attribute of the element Ti in the obtained browsing feature vector to include at least one of the following parameters:
  • the generation time of the Q&A page the number of answers, the number of favorable comments, the number of bad reviews, the length of questions and answers, the time of user browsing, and the time of user stay.
  • the question and answer page related problem recommendation method and apparatus provided in the foregoing embodiment, when a question and answer page needs to be generated for the first user who inputs the search term to obtain the relevant question, the related problem is determined according to the browsing behavior logs of a plurality of second users in a period of time. Browsing weights, obtaining better related questions based on browsing weights, thereby obtaining a problem that is more relevant to the relevance of the questions input by the first user, and matching the related questions obtained with the answers to the questions that the user really wants to obtain. Better, better able to meet user needs, so that users can see more close to the questions retrieved on the Q&A page.
  • the related browsing weights of each related problem can be calculated according to the grouping, so that the browsing weights of each related problem are obtained, and the browsing behaviors of each related problem based on a plurality of second users are measured to measure the acquired
  • the matching of the relevant problems to the user's needs is high and low, so as to achieve the purpose of obtaining related problems with better matching.
  • the embodiment of the present invention provides a question and answer page related question recommendation.
  • the method by analyzing the search behavior of several historical users, obtains a problem that is more relevant to the answer matching of the questions that the user really wants to obtain.
  • the thirteenth embodiment of the present invention provides a method for recommending a question related to a question and answer page.
  • the method is as shown in FIG. 21 and includes the following steps:
  • Step S2101 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term for a question and answer search, and when the question and answer page is generated, the generated question and answer page includes but is not limited to the title of the question and answer page, at least one question answer, and at least one related question.
  • the search term input by the first user several related questions are obtained from the database, and the related questions are related questions in the question and answer question or the question and answer page in the question and answer page of the second user search click in the database.
  • the first user refers to the current user
  • the second user refers to the historical user
  • Step S2102 Determine the click weight of the obtained related question according to the search behavior log of the second user in the selected time period.
  • the search behavior log of the second user corresponding to the related problem obtained in the above step S2101 is obtained from the database. Analyze the search behavior log to determine the click weight of the relevant question. In the process of determining the weight of the hit, the relevant click weights of each other may be calculated for the related questions obtained, and the relevant click weights of the same related question are weighted according to the calculated relevant click weights, and the click weights of the related questions are obtained. .
  • the acquired related questions may also be grouped according to the set grouping conditions, and in each related problem group, the relevant click weights of each related problem and other related questions in the group are respectively calculated, and then the calculation results of each group are integrated. The weights of the relevant click weights of the same related problem appearing in each group are weighted to obtain the click weights of each related question.
  • the grouping according to the query request string is taken as an example to illustrate the determination of the click weight of the related question. process.
  • Step S2103 Sort the acquired related questions according to the determined click weights of the related questions.
  • the related questions are sorted. For example, you can sort by the order of the click weights from high to low.
  • Step S2104 Determine a related question recommended to the first user according to the sorted result of the obtained related question.
  • the relevant questions are selected and recommended to the first user. For example, the related problem of obtaining the set number of the highest click weight among all related questions is recommended to the first user; or the related question of obtaining the set quantity respectively in the related question corresponding to each query request string is recommended to the first user.
  • a fourteenth embodiment of the present invention provides a method for recommending a question related to a question and answer page.
  • the method is as shown in FIG. 22 and includes the following steps:
  • Step S2201 Acquire at least one related question in the database related to the search term from the first user according to the search term from the first user.
  • the first user inputs a search term "What should the child have to do with a cold?”, and generates a corresponding question and answer page according to the search term.
  • the generated question and answer page has a title of the question and answer page, at least one question answer, and at least one related question.
  • the related question can be "What should I do if I have a cold and cough?”, "What should I do if I have a cold and a fever?", "What kind of medicine is better for children with cold and cough?", "What about children with cold and stuffy nose?”, "Baby has a cold and cough What should I do?
  • Step S2202 Group the related questions obtained according to the query request string corresponding to the obtained related question.
  • each related problem group includes some or all related questions corresponding to one query request string.
  • the click feature vector ⁇ T1, T2, ..., Ti, Ti+1, ..., Tn ⁇ of each query request string is obtained, where Ti represents a related problem. .
  • the related problems of the acquisition are grouped.
  • the attribute of the element Ti in the click feature vector includes at least one of the following parameters: the generation time of the question and answer page, the number of answers, the number of praises, the number of bad reviews, the length of the question and answer, the number of impressions, the number of clicks, and the like.
  • the query request string is "baby cold”. As a group;
  • the query request string is "Children's cold” and is grouped into one group;
  • the query request string is "Fever and fever” and is grouped into one group;
  • the query request string is "slow nose and thiophene" and is grouped into one group;
  • Step S2203 In each related problem group, calculate related clicks of each related question in the group and other related questions in the group Weights.
  • the relevant click weight of 1 is W(Ti, Ti+I):
  • Ti+I indicates other question and answer questions included in the click feature vector
  • Query Request String) indicates the probability of obtaining Ti when using the query request string
  • Query Request String) indicates the probability of getting Ti+I when using the query request string.
  • each related question and other related questions in the group separately. For example, group the related questions for the query request string as “pediatric cold” and calculate “What should I do with a cold cough?” "What should I do if my child has a cold?", "What kind of medicine is better for children with cold and cough?”, "What about children with cold and stuffy nose?” The weight of the clicks related to other related issues in the group. Other related problem groups are also calculated.
  • step S2204 the relevant click weights calculated in the related problem groups of the same related problem are obtained, and the obtained related click weights are weighted to obtain the click weights of the obtained related questions.
  • the same related questions in each related problem group are extracted. For example, for the related question "What should I do with a cold cough?" Grouping related questions for "pediatric colds" and weighting the relevant click weights calculated in the group of related questions in which the query request string is "cold cough”.
  • the relevant click weights calculated by the same related problem in different related problem groups may be directly added, or may be added after multiplying the corresponding weight coefficients, or may be weighted by other weighting rules. deal with.
  • Step S2205 Sort the acquired related questions according to the determined click weights of the related questions.
  • Step S2206 Determine related questions recommended to the first user according to the sorted result of the obtained related questions.
  • the first few questions with the highest click weight are recommended as related questions to the first user, and are added to the question and answer page generated according to the search words input by the user.
  • the historical user analyzes the search click behavior of each related question by the historical user, determines the click weight parameter of the related question, and determines the recommendation priority of recommending the relevant question to the user, thereby obtaining the search input with the user.
  • Relevant questions with higher word matching degree provide users with relevant questions that are more conformable to user needs and more in line with user needs in the current question and answer page, and improve the accuracy of question and answer page question retrieval.
  • the system includes a database, and stores related questions of a plurality of second users (historical users).
  • the question and answer page problem recommendation device can obtain the search words input by the first user, and obtains a plurality of related questions from the database that are searched by the historical user. And the historical data of related problems, through the analysis and processing of historical data, to achieve better and relevant issues recommended to the first user.
  • the embodiment of the present invention further provides a question and answer page related problem recommendation device.
  • the structure of the device is as shown in FIG. 24, and includes: a problem obtaining module 2401, a weight determining module 2402, and a sorting recommending module 2403.
  • the problem obtaining module 2401 is configured to acquire at least one related problem related to the search term in the database according to the search term from the first user.
  • the weight determining module 2402 is configured to determine a click weight of the acquired related question according to the search behavior log of the second user in the selected time period.
  • the ranking recommendation module 2403 is configured to sort related related questions according to the determined click weights, and determine related questions recommended by the question and answer page to the first user according to the sorted results of the obtained related questions.
  • the weight determining module 2402 further includes a question grouper 24021, an associated weight calculator 24022, and a click weight calculator 24423.
  • the problem grouper 24021 is configured to group the obtained related questions according to the query request string corresponding to the obtained related question; wherein each related question group includes some or all related questions corresponding to one query request string.
  • the correlation weight calculator 24022 is configured to calculate related click weights of each related question in the group and other related questions in the group in each related problem group.
  • Clicking the weight calculator 24023 is configured to obtain the relevant click weights calculated in the related problem groups of the same related problem, and weight the obtained related click weights to obtain the click weights of the obtained related questions.
  • the problem grouper 24021 is specifically configured to obtain a click feature vector ⁇ T1, T2, ..., Tn ⁇ of each query request string according to the query request string corresponding to the obtained related question, and implement related problems of the acquisition. Grouping; where Ti represents a related problem.
  • the correlation weight calculator 24022 is specifically configured to calculate a related click weight W of each related question in the group and other related questions in the group by using the following formula:
  • Ti+I indicates other question and answer questions included in the click feature vector
  • Query Request String) indicates the probability of obtaining Ti when using the query request string
  • Query Request String) indicates the probability of getting Ti+1 when using the query request string.
  • the problem grouper 24021 specifically for the attribute of the element Ti in the obtained click feature vector, includes at least one of the following parameters: a generation time of the question and answer page, an answer number, a favorable number, a bad evaluation number, a question and answer length, Impressions, clicks, etc.
  • the related problem is determined according to the search behavior logs of a plurality of second users in a period of time. Click weights, get better related questions based on click weights, and thus get a better correlation with the relevance of the questions input by the first user, so that the matching related questions and the answers to the questions that the user really wants to obtain are matched. Better, better able to meet user needs, so that users can see more close to the questions retrieved on the Q&A page.
  • the relevant click weights of each related question can be calculated according to the group, thereby obtaining the click weights of each related problem, and the search click behavior based on a plurality of second users for each related problem is measured to measure the acquisition.
  • the related problems of the relevant users meet the matching degree of the user's needs, thus achieving the purpose of obtaining related problems with better matching.
  • modules in the devices of the embodiments can be adaptively changed and placed in one or more devices different from the embodiment.
  • the modules or units or components of the embodiments may be combined into one module or unit or component, and further they may be divided into a plurality of sub-modules or sub-units or sub-components.
  • any combination of the features disclosed in the specification, including the accompanying claims, the abstract and the drawings, and any methods so disclosed, or All processes or units of the device are combined.
  • Each feature disclosed in this specification (including the accompanying claims, the abstract and the drawings) may be replaced by alternative features that provide the same, equivalent or similar purpose.
  • the various component embodiments of the present invention may be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof.
  • a microprocessor or digital signal processor may be used in practice to implement some or all of the functionality of some or all of the components of the device in accordance with embodiments of the present invention.
  • the invention can also be implemented as a device or device program (e.g., a computer program and a computer program product) for performing some or all of the methods described herein.
  • a program implementing the invention may be stored on a computer readable medium or may be in the form of one or more signals.
  • Such signals may be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
  • FIG. 25 illustrates a computing device that can implement a question and answer page related question recommendation method in accordance with the present invention.
  • the computing device conventionally includes a processor 2510 and a computer program product or computer readable medium in the form of a memory 2520.
  • the memory 2520 may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read Only Memory), an EPROM, a hard disk, or a ROM.
  • the memory 2520 has a storage space 2530 for executing program code 2531 of any of the above method steps.
  • storage space 2530 for program code may include various program code 2531 for implementing various steps in the above methods, respectively.
  • the program code can be read from or written to one or more computer program products.
  • Such computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks.
  • Such a computer program product is typically a portable or fixed storage unit as described with reference to FIG.
  • the storage unit may have a storage segment, a storage space, and the like that are similarly arranged to the storage 2520 in the computing device of FIG.
  • the program code can be compressed, for example, in an appropriate form.
  • the storage unit includes computer readable code 2531', ie, code that can be read by, for example, a processor such as 2510, which when executed by a computing device causes the computing device to perform each of the methods described above step.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种问答页面相关问题推荐方法及装置,其中,该方法包括:根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题(S102);根据至少一个预设规则对获取的相关问题进行筛选(S104);根据相关问题的筛选结果,确定推荐给用户的相关问题(S106)。依据问答页面相关问题推荐方法,能够得到更准确、更贴合用户需要的相关问题,因此能够提高问答页面检索的准确性。

Description

问答页面相关问题推荐方法及装置 技术领域
本发明涉及搜索技术领域,特别是涉及一种问答页面相关问题推荐方法及装置。
背景技术
随着互联网技术的发展,互联网数据早已呈现爆炸性增长的趋势,人们对知识的需求越来越渴望,越来越多的人们开始使用搜索引擎搜索来满足对未知知识的查询与搜索。大型搜索引擎(比如谷歌google、360、百度等)可以很方便快捷的提供相关问答的搜索。其中相关问答搜索是指用户输入一个问题,搜索引擎检索与该问题相对应的答案。在不同的问答知识页面,不仅提供了针对用户输入的问题进行回答的相关答复内容,还提供了与当前问答页面的用户输入问题相关的问题链接,供用户参考使用,方便用户在进行问答搜索时从不同角度综合得到该问题的解决答案。
例如:当前问答页面的搜索问题为:“感冒咳嗽怎么办?”在当前问答页面为用户推荐的相关问题可以包括:“感冒怎么办?”,“感冒咳嗽流鼻涕怎么办?”,“小孩感冒咳嗽怎么办?”,等等。
现有技术中获取相关问题时,一般是根据用户输入的搜索词作为核心词来进行获取的,这种方式比较简单直接,但获取到的相关问题与用户输入的问题的相关度并不是很好,往往不能很好地满足用户的需求,也就是说,其所获取的相关问题与用户真正想要获得的问题答案之间的匹配度比较差,导致问答页面问题检索的准确性比较差,与用户需求的贴合性比较差,不能解决用户想在当前问答页面查看与所检索的问题更贴近的、更吻合的问题答案的检索匹配需求。
因此,如何获取更合适的相关问题推荐给用户,成为问答页面相关问题获取推荐过程中亟待解决的技术问题。
发明内容
鉴于上述问题,提出了本发明以便提供一种克服上述问题或者至少部分地解决上述问题的搜索过程中的问答页面相关问题推荐方法和相应的问答页面相关问题推荐装置。
依据本发明的一个方面,提供了一种问答页面相关问题推荐方法,包括:根据来自用户的搜索词,获取数据库与所述搜索词相关的至少一个相关问题;根据至少一个预设规则对获取的所述相关问题进行筛选;根据所述相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
依据本发明的另一个方面,提供了一种问答页面核心词提取方法,包括:从问答页面中提取核心词候选串;对所述核心词候选串进行分词,提取各个候选串分词的分类特征;根据所述分类特征筛选各个候选串分词是否是核心词。
依据本发明的再一个方面,还提供了一种问答页面相关问题推荐方法,包括:根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;根据选定时间段内第二用户的浏览行为日志,确定获取的所述相关问题的浏览权重;根据所述浏览权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
依据本发明的又一个方面,提供了一种问答页面相关问题推荐方法,包括:根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;根据选定时间段内第二用户的搜索行为日志,确定获取的所述相关问题的点击权重;根据所述点击权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面推荐给第一用户的相关问题。
依据本发明的又一方面,提供了一种问答页面相关问题推荐装置,包括:获取器,适于根据来自用户的搜索词,获取数据库与所述搜索词相关的至少一个相关问题;筛选器,适于根据至少一个预设规则对获取的所述相关问题进行筛选;推荐器,适于根据所述相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
依据本发明的又一方面,提供了一种问答页面核心词提取装置,包括:候选串提取模块, 用于从问答页面中提取核心词候选串;特征提取模块,用于对核心词候选串进行分词,提取各个候选串分词的分类特征;核心词确定模块,用于根据所述分类特征筛选各个候选串分词是否是核心词。
依据本发明的又一方面,提供了一种问答页面相关问题推荐装置,包括:问题获取模块,用于根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;权重确定模块,根据选定时间段内第二用户的浏览行为日志,确定获取的所述相关问题的浏览权重;排序推荐模块,用于根据所述浏览权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
依据本发明的又一方面,提供了一种问答页面相关问题推荐装置,包括:问题获取模块,用于根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;权重确定模块,用于根据选定时间段内第二用户的搜索行为日志,确定获取的所述相关问题的点击权重;排序推荐模块,用于根据所述点击权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面推荐给第一用户的相关问题。
依据本发明的又一个方面,提供了一种计算机程序,其包括计算机可读代码,当所述计算机可读代码在计算设备上运行时,导致所述计算设备执行上述的问答页面相关问题推荐方法。
依据本发明的又一个方面,提供了一种计算机可读介质,其中存储了如上述的计算机程序。
依据本发明实施例的问答页面相关问题推荐方法,能够根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题,并根据至少一个预设规则对获取的相关问题进行筛选,根据筛选结果确定推荐给用户的相关问题。可知,依据本发明实施例的问答页面相关问题推荐方法,在获取到与搜索词相关的相关问题后,利用预设规则对相关问题进行筛选,得到能够更好地反映用户输入的搜索词的相关问题,从而获取到用户真正想要获得的问题答案。另外,本例中利用至少一个预设规则对获取的相关问题进行筛选,即,本例中可以利用多个预设规则对获取的相关问题进行筛选。而利用多个预设规则对获取的相关问题进行多次筛选,能够得到更准确、更贴合用户需要的相关问题,因此能够提高问答页面检索的准确性。
上述说明仅是本发明技术方案的概述,为了能够更清楚了解本发明的技术手段,而可依照说明书的内容予以实施,并且为了让本发明的上述和其它目的、特征和优点能够更明显易懂,以下特举本发明的具体实施方式。
附图说明
通过阅读下文优选实施方式的详细描述,各种其他的优点和益处对于本领域普通技术人员将变得清楚明了。附图仅用于示出优选实施方式的目的,而并不认为是对本发明的限制。而且在整个附图中,用相同的参考符号表示相同的部件。在附图中:
图1示出了根据本发明一个实施例的问答页面相关问题推荐方法的处理流程图;
图2示出了根据本发明一个实施例的根据核心词筛选相关问题并推荐的处理流程图;
图3示出了根据本发明另一个实施例的根据核心词筛选相关问题并推荐的处理流程图;
图4示出了根据本发明又一个实施例的根据核心词筛选相关问题并推荐的处理流程图;
图5示出了根据本发明一个实施例的根据用户的浏览行为日志对相关问题进行筛选并推荐的处理流程图;
图6示出了根据本发明另一个实施例的根据用户的浏览行为日志对相关问题进行筛选并推荐的处理流程图;
图7示出了根据本发明一个实施例的根据用户的搜索点击行为日志对相关问题进行筛选并推荐的处理流程图;
图8示出了根据本发明另一个实施例的根据用户的搜索点击行为日志对相关问题进行筛选并推荐的处理流程图;
图9示出了根据本发明一个实施例的实现问答页面相关问题推荐的系统环境示意图;
图10示出了根据本发明一个优选实施例的根据以上三项预设规则对相关问题进行筛选并推荐的处理流程示意图;
图11示出了根据本发明一个实施例的问答页面相关问题推荐装置的结构示意图;
图12示出了根据本发明一个优选实施例的问答页面相关问题推荐装置的结构示意图;
图13是本发明实施例八的问答页面核心词提取方法的流程图;
图14是本发明实施例九的问答页面核心词提取方法的流程图;
图15是本发明实施例十的问答页面核心词提取方法的流程图;
图16是本发明实施例中问答页面核心词提取装置的结构示意图;
图17是本发明实施例十一中问答页面相关问题推荐方法的流程图;
图18是本发明实施例十二中问答页面相关问题推荐方法的流程图;
图19是本发明实施例中实现问答页面相关问题推荐的系统环境示意图;
图20是本发明实施例中问答页面相关问题推荐装置的结构示意图;
图21是本发明实施例十三中问答页面相关问题推荐方法的流程图;
图22是本发明实施例十四中问答页面相关问题推荐方法的流程图;
图23是本发明实施例中实现问答页面相关问题推荐的系统环境示意图;
图24是本发明实施例中问答页面相关问题推荐装置的结构示意图;
图25示意性地示出了用于执行根据本发明的问答页面相关问题推荐方法的计算设备的框图;以及
图26示意性地示出了用于保持或者携带实现根据本发明的问答页面相关问题推荐方法的程序代码的存储单元。
具体实施方式
下面结合附图和具体的实施方式对本发明作进一步的描述。
下面将参照附图更详细地描述本公开的示例性实施例。虽然附图中显示了本公开的示例性实施例,然而应当理解,可以以各种形式实现本公开而不应被这里阐述的实施例所限制。相反,提供这些实施例是为了能够更透彻地理解本公开,并且能够将本公开的范围完整的传达给本领域的技术人员。
为解决上述技术问题,本发明实施例提供了一种问答页面相关问题推荐方法。图1示出了根据本发明一个实施例的问答页面相关问题推荐方法的处理流程图。参见图1,该流程至少包括步骤S102至步骤S106。
步骤S102、根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题;
步骤S104、根据至少一个预设规则对获取的相关问题进行筛选;
步骤S106、根据相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
依据本发明实施例的问答页面相关问题推荐方法,能够根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题,并根据至少一个预设规则对获取的相关问题进行筛选,根据筛选结果确定推荐给用户的相关问题。可知,依据本发明实施例的问答页面相关问题推荐方法,在获取到与搜索词相关的相关问题后,利用预设规则对相关问题进行筛选,得到能够更好地反映用户输入的搜索词的相关问题,从而获取到用户真正想要获得的问题答案。另外,本例中利用至少一个预设规则对获取的相关问题进行筛选,即,本例中可以利用多个预设规则对获取的相关问题进行筛选。而利用多个预设规则对获取的相关问题进行多次筛选,能够得到更准确、更贴合用户需要的相关问题,因此能够提高问答页面检索的准确性。
上文提及,为保证能够为用户提供更贴合用户需求的检索结果,本发明实施例根据至少一个预设规则对与搜索词相关的相关问题进行筛选。本例中,对相关问题进行筛选所依据的预设规则可以是任意能够对相关问题进行进一步筛选的规则。例如,预设规则可以是根据用户行为日志对相关问题进行筛选,还可以是根据搜索词与相关问题的贴合程度对相关问题进行筛选。
本发明实施例中,优选根据以下预设规则对相关问题进行筛选:
(1)根据核心词对相关问题进行筛选;
(2)根据用户的浏览行为日志对相关问题进行筛选;
(3)根据用户的搜索点击行为日志对相关问题进行筛选。
另外,本例中可以仅根据以上预设规则中的一项对相关问题进行筛选,还可以根据以上预设规则中的几项或全部对相关问题进行筛选。之后,根据筛选结果确定推荐给用户的相关问题。在根据以上预设规则中的几项或全部对相关问题进行筛选时,先根据各预设规则分别对相关问题进行筛选,之后拟合各个筛选结果得到推荐给用户的相关问题,可见,在根据多个预设规则对相关问题进行筛选时,仍旧需要进行单个预设规则对相关问题进行筛选的过程。因此,本例中,对根据各个预设规则分别对相关问题进行筛选,并根据筛选结果确定推荐给用户的相关问题的过程进行介绍。
(1)根据核心词对相关问题进行筛选,并根据筛选结果确定推荐的相关问题。
现有技术中,仅根据搜索词进行检索,存在由于检索时在搜索词中提取的核心词不合适,而导致不能获取到匹配度较高的、更贴合用户需求的问答问题答案的问题,因此,本例中,首先获取与搜索词对应的问答页面。其次,提取问答页面中的核心词,并根据提取的核心词筛选相关问题。
实施例一
图2示出了根据本发明一个实施例的根据核心词筛选相关问题并推荐的处理流程图。参见图2,该流程包括如下步骤:
步骤S201:根据用户输入的搜索词获取对应的问答页面及相关问题。
步骤S202:从问答页面中提取核心词候选串。
提取核心词时,从问答页面中提取用于确定核心词的核心词候选串,从候选串中筛选出符合条件的核心词。
从问答页面中提取核心词候选串,可以从问答页面的标题中提取核心词候选串,也可以从问答页面的页面内容中提取,或者从问答页面的标题和问答页面的页面内容中提取。
从问答页面中提取核心词候选串,包括:获取与用户输入的搜索词对应的问答页面;从获取的问答页面的标题中提取核心词候选串。和/或从获取的问答页面的页面内容中,提取与用户输入的搜索词相关的字符串,作为核心词候选串。
步骤S203:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
提取到问答页面的核心词候选串后,进行分词处理,将每一个候选串分词划分为若干候选串分词,并提取出这些候选串分词的分类特征。其中,候选串分词的分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频等等。
步骤S204:根据提取出的分类特征筛选各个候选串分词是否是核心词。
提取出候选串分词的分类特征后,根据分类特征对候选串分词进行分类,并根据分类结果确定各个候选串分词是否是核心词。
如上所述,候选串分词的分类特征包括名词、热度词表、超链接、相关问题共现率、文档词频等特征中的至少一种,则可以候选串分词中所有的名词归为一类,将候选串分词中在热度词表中的分词归为一类,将候选串分词汇中是超级链接的分词归为一类,或者也可以将候选串分词中在热度词表中的所有名词归为一类,……,等等。
对候选串分词进行分类后,可以根据分类结果,进行核心词的筛选,比如,根据各个分类中各个候选串分词与用户输入的搜索词的匹配程度进行筛选,或者根据各个分类中各个候选串分词的使用频率统计值等因素进行筛选,或者综合考虑上述各种因素进行筛选。
其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。可以建立数据库,统计候选串分词被用户搜索的次数,被用户点击的次数曾经被确定为核心词的次数、曾经被用户用作搜索词的次数等。
步骤S205:利用步骤S204中确定的核心词筛选步骤S201中获取到的相关问题。
实施例二
图3示出了根据本发明另一个实施例的根据核心词筛选相关问题并推荐的处理流程图,如图3所示,包括如下步骤:
步骤S301:获取与用户输入的搜索词对应的问答页面及相关问题。
例如:用户输入搜索词“孩子感冒咳嗽怎么办?”,根据该搜索词获取到对应的问答页面, 获取到的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如相关问题可以是“孩子感冒咳嗽怎么办?”,“小儿感冒咳嗽用什么药比较好呢?”。
步骤S302:从获取的问答页面的标题中提取核心词候选串。
本实施例中以从问答页面的标题中提取核心词候选串为例,比如,提取到的核心词候选串可以是“孩子感冒咳嗽怎么办”。
实际操作中还可以从问答页面的问答内容、相关问题等页面内容中提取核心词候选串。
步骤S303:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
对提取的核心词候选串“孩子感冒咳嗽怎么办”进行分词,例如,可以分词为:“孩子”、“感冒”、“咳嗽”、“怎么办”等候选串分词。
对分词出的候选串分词进行分类特征提取,例如“孩子”这个候选串分词的分类特征包括:是名词等;“感冒”、“咳嗽”这两个候选串分词的分类特征包括:是名词、是热度词表中的词、是超链接等;“怎么办”这个候选串分词的分类特征包括是超链接等。
步骤S304:根据提取的分类特征对候选串分词进行分类。
根据提取的分类特征对上述分词出的“孩子”、“感冒”、“咳嗽”、“怎么办”等候选串分词进行分类,例如:“孩子”、“感冒”、“咳嗽”都是名词,归为一类;将“感冒”、“咳嗽”都是热度词表中的词,归为一类;“感冒”、“咳嗽”、“怎么办”都是超链接,归为一类。
步骤S305:针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配。
对候选串分词进行分类后,分别针对每个分类,与用户输入的搜索词进行匹配。
沿用上边的例子,根据上边的分类,将名词分类、热度词表分类和超链接分类中的各个候选串分词分别与用户输入的搜索词进行匹配。
步骤S306:筛选出匹配度最高的设定数量的候选串分词,作为核心词。
沿用上边的例子,筛选出匹配度较高的2个候选串分词为:“感冒”、“咳嗽”,则确定“感冒”、“咳嗽”为核心词;或筛选出匹配度较高的3个候选串分词为:“感冒”、“咳嗽”、“孩子”,则确定“感冒”、“咳嗽”、“孩子”为核心词。
步骤S307:根据确定的核心词筛选相关问题。
沿用上边的例子,根据核心词“感冒”、“咳嗽”、“孩子”筛选得到相关问题“孩子感冒咳嗽怎么办?”。
上述实施例中所列举的搜索词、问答页面标题等都属于简单的举例,实际应用中用户输入的检索词可能会更简单,而根据问答页面获取到的候选串分词的数量可能会更多,匹配过程可能会更复杂,从而能够更好地发挥本发明方法的作用,在此不再一一列举。
上述步骤S305和步骤S306实现了根据分类结果确定各个候选串分词是否是核心词。
上述实施例二中的步骤S305和步骤S306可替换为下面步骤S405和步骤S406所公开的筛选方式。
实施例三
图4示出了根据本发明又一个实施例的根据核心词筛选相关问题并推荐的处理流程图,如图4所示,该流程包括如下步骤:
步骤S401:获取与用户输入的搜索词对应的问答页面及相关问题。
例如:用户输入搜索词“孩子感冒咳嗽怎么办?”,根据该搜索词获取到对应的问答页面,获取到的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如,问答答案中可能包括“选择正确的感冒(咳嗽)药”、“感冒止咳的中药”等描述,相关问题可以是“孩子感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”等问题。
步骤S402:从获取的问答页面的页面内容中,提取与用户输入的搜索词相关的字符串,作为核心词候选串。
对用户输入的搜索词进行分词,从获取的问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
沿用上边的例子,对用户输入的搜索词“孩子感冒咳嗽怎么办?”进行分词,例如可以分 词为“孩子”、“感冒”、“咳嗽”、“怎么办”等搜索词分词。
本实施例中以从问答页面的页面内容中提取核心词候选串为例,可以从问答页面的问答内容、相关问题等页面内容中提取包括“孩子”、“感冒”、“咳嗽”、“怎么办”中至少一个搜索词分词的字符串作为核心词候选串。例如,提取到的核心词候选串可以有:“孩子感冒咳嗽怎么办”、“选择正确的感冒(咳嗽)药”、“感冒止咳的中药”、“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”等等。
步骤S403:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
沿用上边的例子,对提取的核心词候选串“孩子感冒咳嗽怎么办”进行分词,例如,可以分词为:“孩子”、“感冒”、“咳嗽”、“怎么办”等候选串分词。对提取的核心词候选串“选择正确的感冒(咳嗽)药”进行分词,例如,可以分词为:“选择”、“正确的”、“感冒”、“咳嗽”、“药”等候选串分词。对提取的核心词候选串“感冒止咳的中药”进行分词,例如,可以分词为:“感冒”、“止咳”、“中药”等候选串分词。依次对提取的核心词候选串进行分词,此处不再一一列举。
对分词出的候选串分词进行分类特征提取,例如“孩子”这个候选串分词的分类特征包括:是名词等;“感冒”、“咳嗽”这两个候选串分词的分类特征包括:是名词、是热度词表中的词、是超链接等;“中药”、“药”这两个候选串分词的分类特征包括:是名词等;“止咳”这个候选串分词的分类特征包括:是热度词表中的词等;“怎么办”这个候选串分词的分类特征包括:是超链接等。总之,对分词出的所有候选串分词都进行分类特征提取,此处不再对上边举例中的各候选串一一列举其分类特征。
步骤S404:根据提取的分类特征对候选串分词进行分类。
根据提取的分类特征对上述分词出的“孩子”、“感冒”、“咳嗽”、“怎么办”、“选择”、“正确的”、“药”、“止咳”、“中药”等候选串分词进行分类,例如:“孩子”、“感冒”、“咳嗽”、“中药”、“药”都是名词,归为一类;将“感冒”、“咳嗽”、“止咳”都是热度词表中的词,归为一类;“感冒”、“咳嗽”、“怎么办”都是超链接,归为一类。总之,对分词出的所有候选串分词都根据分类特征进行分类,此处不再对上边举例中的各候选串一一列举其分类。
步骤S405:针对每个分类,确定该分类中各个候选串分词的使用频率统计值。
沿用上边的例子,在名词分类中、热度词表中的词分类、超链接分类中,分别确定各候选串分词的使用频率统计值。
其中,候选串分词的使用频率统计值可以根据各候选串分词被用户搜索的次数、被用户点击的次数、曾经被确定为核心词的次数、曾经被作为搜索词的次数等因素中的至少一种因素进行统计。
步骤S406:根据各个候选串分词的使用频率统计值,筛选出使用频率统计值最高的设定数量的候选串分词,作为核心词。
沿用上边的例子,筛选出使用频率统计值最高的3个候选串分词为:“感冒”、“咳嗽”、“止咳”,则确定“感冒”、“咳嗽”、“止咳”为核心词;或筛选出使用频率统计值最高的3个候选串分词为:“感冒”、“咳嗽”、“孩子”,则确定“感冒”、“咳嗽”、“孩子”为核心词。
步骤S407:根据确定的核心词对相关问题进行筛选。
沿用上边的例子,根据确定的核心词“感冒”、“咳嗽”、“孩子”筛选得到相关问题“孩子感冒咳嗽怎么办?”。
上述步骤S405和步骤S406实现了根据分类结果确定各个候选串分词是否是核心词。
(2)根据用户的浏览行为日志对相关问题进行筛选,并根据筛选结果确定推荐给用户的相关问题。
本发明实施例中,通过对若干历史用户的浏览行为进行分析,并根据分析结果对相关问题进行筛选,获取到与用户真正想要获得的问题答案匹配度更好的相关问题。
实施例四
图5示出了根据本发明一个实施例的根据用户的浏览行为日志对相关问题进行筛选并推荐的处理流程图。参见图5,该流程包括如下步骤:
步骤S501:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
第一用户输入搜索词进行问答检索,生成问答页面时,生成的问答页面中包括但不限于问答页面的标题、至少一个问题答案,至少一个相关问题。在获取到第一用户输入的搜索词后,从数据库中获取若干相关问题,这些相关问题为数据库中第二用户浏览的问答页面中的问答问题或问答页面中的相关问题。
其中,第一用户是指当前用户,第二用户是指历史用户。
步骤S502:根据选定时间段内第二用户的浏览行为日志,确定获取的相关问题的浏览权重。
从数据库中获取上述步骤S501中获取到的相关问题对应的第二用户的浏览行为日志。对浏览行为日志进行分析,确定相关问题的浏览权重。确定浏览权重的过程中,可以对获取的相关问题,计算彼此之间的相关浏览权重,根据计算出来的相关浏览权重,对同一相关问题的相关浏览权重进行加权处理,得到各相关问题的浏览权重。
优选的,也可以根据设定的分组条件对获取的相关问题进行分组,在各个相关问题分组中,分别计算各相关问题与组中其他相关问题的相关浏览权重,然后综合各组的计算结果,对各组中出现的同一相关问题的相关浏览权重进行加权处理,得到各相关问题的浏览权重。
下面的实施例五中,以根据浏览用户进行分组为例,说明相关问题的浏览权重的确定过程。
步骤S503:根据确定的浏览权重对获取的相关问题进行排序。
根据确定出的各相关问题的浏览权重,对各相关问题进行排序。比如可以按照浏览权重从高到低的顺序进行排序。对相关问题进行排序时,可以对获取所有的相关问题一起进行排序,也可以按照不同的浏览用户在个浏览用户分组中分别排序,或者按照其他的规则排序。
步骤S504:根据获取的相关问题的排序结果,对相关问题进行筛选,进而根据筛选结果确定推荐给第一用户的相关问题。
根据对相关问题的排序结果,按照设定的推荐规则,筛选相关问题,并将筛选得到的相关问题推荐给用户。比如,根据排序结果将所有的相关问题中浏览权重最高的设定数量的相关问题筛选出作为筛选结果推荐给用户;或者在各浏览用户对应的相关问题中分别筛选出设定数量的相关问题作为筛选结果推荐给第一用户。
实施例五
本发明另一个实施例的根据用户的浏览行为日志对相关问题进行筛选的处理的流程如图6所示。参见图6,该流程包括如下步骤:
步骤S601:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
例如:第一用户输入搜索词“孩子感冒怎么办?”,根据该搜索词生成对应的问答页面,生成的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如:相关问题可以是“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”等等。
这些相关问题为数据库中存储的历史用户曾经浏览过的问答页面上的问答问题或问答页面上的相关问题。
步骤S602:根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组。
对获取的相关问题进行分组时,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题。
可选的,根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
其中,浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间等。
沿用上边的例子,对上边获取到的各相关问题进行分组如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”被同一个浏览用户浏览过,归为一组。
“小儿感冒发烧怎么办?”、“儿童感冒发烧怎么办”、“小儿感冒鼻塞怎么办?”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”被同一个浏览用户浏览过,归为一组。
“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”被同一个浏览用户浏览过,归为一组。
……
以此类推,对所有获取的相关问题进行分组,实现将被同一用户浏览过的相关问题归为一组。
步骤S603:在各相关问题分组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
根据上述各浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},利用如下公式计算每个相关问题与组中其它相关问题的相关浏览权重W(Ti,Ti+1):
log(a1/(|Time(i)–Time(i+1)|+a2))
其中,Time(i)一个问答问题的用户浏览时间;
Time(i+1)为组中其它问答问题的用户浏览时间;
a1,a2为经验值常数。
当然,也可以计算组中各相关问题Ti与组中其他相关问题Ti-1的相关浏览权重W。
沿用上边的例子,针对每个分组,分别计算每个相关问题与组中其他相关问题的,例如,针对浏览用户相同的第一个相关问题分组,分别计算“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”与组中其他相关问题的相关浏览权重。其他相关问题分组也同样进行计算。
进一步可选的,计算组中各相关问题与组中其它相关问题的相关浏览权重,包括:在各相关问题分组中,根据浏览用户浏览各相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;在各会话组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
也就是说,对于浏览用户相同的相关问题分组中的用户,可以进一步根据浏览时间划分出不同的会话组(session),同一会话组中的相关问题的浏览时间差小于等于某个设定的时间阈值。可以根据浏览用户的浏览特征向量进行session划分。在同一session内,计算相关问题的浏览权重。
步骤S604:获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的各相关问题的浏览权重。
上边计算出各相关问题分组中的各相关问题的相关浏览权重后,将各相关问题分组中相同的相关问题提取出来,例如,对于“小儿感冒鼻塞怎么办?”这个相关问题,在浏览用户相同的第一个相关问题分组和第三个相关问题中计算得到的相关浏览权重进行加权。
可选的,可以把同一相关问题在不同相关问题分组中计算得到的相关浏览权重直接进行相加,也可以分别乘上相应的权重系数后在进行相加,也可以通过其它的加权规则进行加权处理。
步骤S605:根据确定出的相关问题的浏览权重对获取的相关问题进行排序。
沿用上边的例子,以获取所有的相关问题一起进行排序为例,按照浏览权重从高到低的顺序进行排序,得到排序结果如下:
“小儿感冒发烧怎么办?”、“小儿感冒咳嗽怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、 “宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”。
步骤S606:根据获取的相关问题的排序结果,对相关问题进行筛选,进而根据筛选结果确定推荐给第一用户的相关问题。
根据排序结果,筛选出浏览权重最高的前几个问题作为筛选结果推荐给第一用户,加入到根据用户输入的搜索词生成的问答页面中。
例如:将“小儿感冒发烧怎么办?”、“小儿感冒咳嗽怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”作为相关问题加入到问答页面中。
(3)根据用户的搜索点击行为日志对相关问题进行筛选,并根据筛选结果确定推荐给用户的相关问题。
本发明实施例中,通过对若干历史用户的搜索点击行为进行分析,并根据分析结果对相关问题进行筛选,获取到与用户真正想要获得的问题答案匹配度更好的相关问题。
实施例六
图7示出了根据本发明一个实施例的根据用户的搜索点击行为日志对相关问题进行筛选并推荐的处理流程图。参见图7,该流程包括如下步骤:
步骤S701:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
第一用户输入搜索词进行问答检索,生成问答页面时,生成的问答页面中包括但不限于问答页面的标题、至少一个问题答案,至少一个相关问题。在获取到第一用户输入的搜索词后,从数据库中获取若干相关问题,这些相关问题为数据库中第二用户搜索点击的问答页面中的问答问题或问答页面中的相关问题。
其中,第一用户是指当前用户,第二用户是指历史用户。
步骤S702:根据选定时间段内第二用户的搜索行为日志,确定获取的相关问题的点击权重。
从数据库中获取上述步骤S701中获取到的相关问题对应的第二用户的搜索行为日志。对搜索行为日志进行分析,确定相关问题的点击权重。确定击权重的过程中,可以对获取的相关问题,计算彼此之间的相关点击权重,根据计算出来的相关点击权重,对同一相关问题的相关点击权重进行加权处理,得到各相关问题的点击权重。
优选的,也可以根据设定的分组条件对获取的相关问题进行分组,在各个相关问题分组中,分别计算各相关问题与组中其他相关问题的相关点击权重,然后综合各组的计算结果,对各组中出现的同一相关问题的相关点击权重进行加权处理,得到各相关问题的点击权重。
下面的实施例七中,以根据查询请求串进行分组为例,说明相关问题的点击权重的确定过程。
步骤S703:根据确定出的相关问题的点击权重对获取的相关问题进行排序。
根据确定出的各相关问题的点击权重,对各相关问题进行排序。比如可以按照点击权重从高到低的顺序进行排序。对相关问题进行排序时,可以对获取所有的相关问题一起进行排序,也可以按照不同的查询请求串在各查询串分组中分别排序,或者按照其他的规则排序。
步骤S704:根据获取的相关问题的排序结果,对相关问题进行筛选,进而根据筛选结果确定推荐给第一用户的相关问题。
根据对相关问题的排序结果,按照设定的推荐规则,筛选相关问题,并将筛选得到的相关问题推荐给第一用户。比如,根据排序结果将所有的相关问题中点击权重最高的设定数量的相关问题筛选出作为筛选结果推荐给第一用户;或者在各查询请求串对应的相关问题中分别筛选出设定数量的相关问题作为筛选结果推荐给第一用户。
实施例七
图8示出了根据本发明另一个实施例的根据用户的搜索点击行为日志对相关问题进行筛选并推荐的处理流程图。参见图8,该流程包括如下步骤:
步骤S801:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
例如:第一用户输入搜索词“孩子感冒怎么办?”,根据该搜索词生成对应的问答页面,生成的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如:相关问题可以是“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽怎么办”“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”等等。
这些相关问题为数据库中存储的历史用户曾经搜索过的问答页面上的问答问题或问答页面上的相关问题。
步骤S802:根据获取的相关问题对应的查询请求串,对获取的相关问题进行分组。
对获取的相关问题进行分组时,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题。
可选的,根据获取的相关问题对应的查询请求串,得到各查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中Ti表示一个相关问题。从而实现对获取的相关问题进行分组。
其中,点击特征向量中的元素Ti的属性包括下列参数中的至少一个:问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
沿用上边的例子,对上边获取到的各相关问题进行分组如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”对应的查询请求串为“小儿感冒”,归为一组。
“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”对应的查询请求串为“宝宝感冒”,归为一组;
“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”对应的查询请求串为“儿童感冒”,归为一组;
“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”,“宝宝感冒咳嗽怎么办”,“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”对应的查询请求串为“感冒咳嗽”,归为一组;
“小儿感冒发烧怎么办?”、“小儿感冒发烧怎么办?”、“儿童感冒发烧怎么办”对应的查询请求串为“感冒发烧”,归为一组;
“小儿感冒鼻塞怎么办?”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”对应的查询请求串为“感冒鼻噻”,归为一组;
……
以此类推,对所有获取的相关问题进行分组,实现将查询请求串相同的相关问题归为一组。
步骤S803:在各相关问题分组中,计算组中各相关问题与组中其他相关问题的相关点击权重。
根据上述生成的各查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},利用如下公式计算组中各相关问题Ti与组中其他相关问题Ti+1的相关点击权重W(Ti,Ti+I):
W=P((Ti)|查询请求串)*P((Ti+I)|查询请求串)
其中,Ti表示一个相关问题;
Ti+I表示点击特征向量中包括的其他问答问题;
P((Ti)|查询请求串)表示使用查询请求串时得到Ti的概率;
P((Ti+I)|查询请求串)表示使用查询请求串时得到Ti+I的概率。
当然,也可以计算组中各相关问题Ti与组中其他相关问题Ti-I的相关点击权重W。
沿用上边的例子,针对每个分组,分别计算每个相关问题与组中其他相关问题的,例如,针对查询请求串为“小儿感冒”的相关问题分组,分别计算“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”与组 中其他相关问题的相关点击权重。其他相关问题分组也同样进行计算。
步骤S804:获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的各相关问题的点击权重。
上边计算出各相关问题分组中的各相关问题的相关点击权重后,将各相关问题分组中相同的相关问题提取出来,例如,对于“小儿感冒咳嗽怎么办?”这个相关问题,在查询请求串为“小儿感冒”的相关问题分组和在查询请求串为“感冒咳嗽”的相关问题分组中计算得到的相关点击权重进行加权。
可选的,可以把同一相关问题在不同相关问题分组中计算得到的相关点击权重直接进行相加,也可以分别乘上相应的权重系数后在进行相加,也可以通过其它的加权规则进行加权处理。
步骤S805:根据确定出的相关问题的点击权重对获取的相关问题进行排序。
沿用上边的例子,以获取所有的相关问题一起进行排序为例,按照点击权重从高到低的顺序进行排序,得到排序结果如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“宝宝感冒咳嗽怎么办”、“儿童感冒发烧怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”。
步骤S806:根据获取的相关问题的排序结果,对相关问题进行筛选,进而根据筛选结果确定推荐给第一用户的相关问题。
根据排序结果,筛选出点击权重最高的前几个问题作为筛选结果推荐给第一用户,加入到根据用户输入的搜索词生成的问答页面中。
例如:将“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“宝宝感冒咳嗽怎么办”、“儿童感冒发烧怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”作为相关问题加入到问答页面中。
上述根据用户的浏览性为日志和/或搜索点击行为日志对相关问题进行筛选和/或推荐的流程中,根据数据库中的历史数据,分析历史用户浏览各个相关问题的浏览行为,和/或点击各个相关问题的搜索点击行为,确定相关问题的浏览权重参数和/或点击权重参数,从而确定向用户推荐相关问题的推荐优先级,从而获取到与用户输入的搜索词匹配度更高的相关问题,在当前问答页面为用户提供与用户需求的贴合性更好、更符合用户需求的相关问题,提高问答页面问题检索的准确性。
针对本发明实施例根据用户的浏览性为日志和/或搜索点击行为日志对相关问题进行筛选和/或推荐的方法,实现问答页面相关问题推荐的系统环境示意如图9所示。该系统包括数据库,存储若干第二用户(历史用户)的相关问题,问答页面问题推荐装置能够获取第一用户输入的搜索词,并根据搜索词从数据库获取若干历史用户浏览和/或搜索点击过的相关问题及相关问题的历史数据,通过对历史数据的分析处理,实现获取更优的相关问题推荐给第一用户。
上文对分别根据各预设规则对相关问题进行筛选,并根据筛选结果推荐相关问题的过程进行了介绍。本例中,当根据预设规则中的几项或全部对相关问题进行筛选时,首先根据各个预设规则分别对相关问题进行筛选,其次,拟合各个筛选结果,得到推荐给用户的相关问题。如图10示出了根据本发明一个优选实施例的根据以上三项预设规则对相关问题进行筛选并推荐的处理流程示意图。参见图10,该流程包括如下步骤:
步骤S1001:获取与用户输入的搜索词对应的相关问题。
例如,用户输入搜索词“小儿感冒怎么办”,根据该搜索词获取到对应的相关问题。例如,获取到的相关问题包括:
“小儿感冒咳嗽怎么办”;
“孩子感冒流鼻涕怎么办”;
“感冒的症状是什么”;
“宝宝感冒的常见问题有什么”;
“感冒发烧怎么办”;
“小儿感冒病因有什么”;
“儿童感冒有没有食疗”;
“怎样停止咳嗽”。
步骤S1002:根据核心词对相关问题进行筛选。
当提取到核心词为“小儿”、“感冒”,根据该核心词筛选到的相关问题为:
“小儿感冒咳嗽怎么办”;
“小儿感冒病因有什么”。
步骤S1003:根据用户的浏览行为日志对相关问题进行筛选。
对步骤S1001中提及的各个相关问题进行浏览权重值的计算,并根据得到的浏览权重值对各个相关问题进行排序,得到排序结果为:
“小儿感冒咳嗽怎么办”;
“怎样停止咳嗽”;
“小儿感冒病因有什么”;
“儿童感冒有没有食疗”;
“宝宝感冒的常见问题有什么”;
“感冒发烧怎么办”;
“孩子感冒流鼻涕怎么办”;
“感冒的症状是什么”。
根据排序结果提取3个相关问题,即得到的筛选结果为:
“小儿感冒咳嗽怎么办”;
“怎样停止咳嗽”;
“小儿感冒病因有什么”。
步骤S1004:根据用户的搜索点击行为日志对相关问题进行筛选。
对步骤S1001中提及的各个相关问题进行搜索点击权重值的计算,并根据得到的搜索点击权重值对各个相关问题进行排序,得到排序结果为:
“怎样停止咳嗽”;
“孩子感冒流鼻涕怎么办”;
“小儿感冒咳嗽怎么办”;
“小儿感冒病因有什么”;
“儿童感冒有没有食疗”;
“宝宝感冒的常见问题有什么”;
“感冒发烧怎么办”;
“感冒的症状是什么”。
根据排序结果提取3个相关问题,即筛选结果为:
“怎样停止咳嗽”;
“孩子感冒流鼻涕怎么办”;
“小儿感冒咳嗽怎么办”。
步骤S1005:根据步骤S1002、步骤S1003以及步骤S1004中得到的各个筛选结果,确定推荐给用户的相关问题。
优选地,本例中可以对步骤S1002、步骤S1003以及步骤S1004中得到的各个筛选结果进行整理排序。例如,得到的三个筛选结果中均包括相关问题“小儿感冒咳嗽怎么办”。再例如,得到的三个筛选结果中的两个筛选结果包括“小儿感冒病因有什么”及“怎样停止咳嗽”。若在问答页面中推荐给用户的相关问题可以是:
“小儿感冒咳嗽怎么办”;
“小儿感冒病因有什么”;
“怎样停止咳嗽”。
需要说明的是,上例中提及的各个筛选结果,和/或步骤S1005中确定推荐的相关问题均为示例,不能够代表实际应用中得到的筛选结果和/或确定推荐的相关问题。
基于同一发明构思,本发明实施例还提供了一种问答页面相关问题推荐装置,该装置的结构如图11所示,包括获取器1110、筛选器1120以及推荐器1130。
现介绍本发明实施例的问答页面相关问题推荐装置的各器件或组成的功能以及各部分间的连接关系:
获取器1110,适于根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题;
筛选器1120,与获取器1110相耦合,适于根据至少一个预设规则对获取的相关问题进行筛选;
推荐器1130,与筛选器1120相耦合,适于根据相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
图12示出了根据本发明一个优选实施例的问答页面相关问题推荐装置的结构示意图。参见图12,筛选器1120还包括:
第一筛选模块1121,与获取器1110以及推荐器1130分别耦合,适于根据用户的浏览行为日志对相关问题进行筛选;
第二筛选模块1122,与获取器1110以及推荐器1130分别耦合,适于根据用户的搜索点击行为日志对相关问题进行筛选;
第三筛选模块1123,与获取器1110以及推荐器1130分别耦合,适于根据核心词对相关问题进行筛选。
在一个优选的实施例中,第三筛选模块1123还包括:
获取单元11231,适于获取与搜索词对应的问答页面;
提取单元11232,与提取单元11231相耦合,适于提取问答页面中的核心词;
确定单元11233,与提取单元11232相耦合,适于根据核心词筛选相关问题。
在一个优选的实施例中,提取单元11232还适于:
从问答页面中提取核心词候选串;
对核心词候选串进行分词,提取各个候选串分词的分类特征;
根据分类特征筛选各个候选串分词是否是核心词。
在一个优选的实施例中,提取单元11232还适于:
从问答页面的标题中提取核心词候选串;和/或
从问答页面的页面内容中,提取与搜索词相关的字符串,作为核心词候选串。
在一个优选的实施例中,提取单元11232还适于:
对搜索词进行分词;
从问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
在一个优选的实施例中,提取单元11232还适于:
根据分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;
分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
在一个优选的实施例中,提取单元11232还适于:
针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为核心词;
针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出使用频率统计值最高的设定数量的候选串分词,作为核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
在一个优选的实施例中,第一筛选模块1121还包括:
第一权重确定单元11211,适于根据选定时间段内用户的浏览行为日志,确定获取的相关问题的浏览权重;
第一排序单元11212,与权重确定单元11211相耦合,适于根据浏览权重对获取的相关问题进行排序;
第一筛选单元11213,与排序单元11212相耦合,适于根据排序结果对相关问题进行筛选。
在一个优选的实施例中,第一筛选单元11213还适于:根据排序结果提取第一预定个数个相关问题。
在一个优选的实施例中,第一权重确定单元11211还适于:
根据浏览相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题;
在每个相关问题分组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重;
获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的每个相关问题的浏览权重。
在一个优选的实施例中,第一权重确定单元11211还适于:
根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
在一个优选的实施例中,第一权重确定单元11211还适于:
在每个相关问题分组中,根据浏览用户浏览每个相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;
根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;
在每个会话组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重。
在一个优选的实施例中,第二筛选模块1122还包括:
第二权重确定单元11221,适于根据选定时间段内用户的搜索点击日志,确定获取的相关问题的点击权重;
第二排序单元11222,与第二权重确定单元11221相耦合,适于根据点击权重对获取的相关问题进行排序;
第二筛选单元11223,与第二排序单元11222相耦合,适于根据排序结果对相关问题进行筛选。
在一个优选的实施例中,第二权重确定单元11221还适于:
根据相关问题对应的查询请求串,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题;
在每个相关问题分组中,计算组中每个相关问题与组中其他相关问题的相关点击权重;
获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的每个相关问题的点击权重。
在一个优选的实施例中,第二权重确定单元11221还适于:
根据相关问题对应的查询请求串,得到每个查询请求串的点击特征向量{T1、T2、……、Tn},实现对获取的相关问题进行分组;其中,Ti表示一个相关问题。
在一个优选的实施例中,第二权重确定单元11221还适于:
得到的点击特征向量中的元素Ti的属性包括下列参数中的至少一个:
问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
根据上述任意一个实施例或多个实施例的组合,本发明实施例能够达到如下有益效果:
依据本发明实施例的问答页面相关问题推荐方法,能够根据来自用户的搜索词,获取数据库与搜索词相关的至少一个相关问题,并根据至少一个预设规则对获取的相关问题进行筛选,根据筛选结果确定推荐给用户的相关问题。可知,依据本发明实施例的问答页面相关问题推荐方法,在获取到与搜索词相关的相关问题后,利用预设规则对相关问题进行筛选,得到能够更好地反映用户输入的搜索词的相关问题,从而获取到用户真正想要获得的问题答案。另外,本例中利用至少一个预设规则对获取的相关问题进行筛选,即,本例中可以利用多个预设规则对获取的相关问题进行筛选。而利用多个预设规则对获取的相关问题进行多次筛选,能够得到更准确、更贴合用户需要的相关问题,因此能够提高问答页面检索的准确性。
为了解决现有技术中存在的检索过程中,由于核心词确定的不是很合适,而导致不能获取到匹配度较高的、更贴合用户需求的问答问题答案的问题,为用户提供更贴合用户需求的检索 结果,本发明实施例提供一种问答页面核心词提取方法。
实施例八
本发明实施例八提供的问答页面核心词提取方法,其流程如图13所示,包括如下步骤:
步骤S1301:从问答页面中提取核心词候选串。
提取核心词时,从问答页面中提取用于确定核心词的核心词候选串,从候选串中筛选出符合条件的核心词。
从问答页面中提取核心词候选串,可以从问答页面的标题中提取核心词候选串,也可以从问答页面的页面内容中提取,或者从问答页面的标题和问答页面的页面内容中提取。
从问答页面中提取核心词候选串,包括:获取与用户输入的搜索词对应的问答页面;从获取的问答页面的标题中提取核心词候选串。和/或从获取的问答页面的页面内容中,提取与用户输入的搜索词相关的字符串,作为核心词候选串。
步骤S1302:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
提取到问答页面的核心词候选串后,进行分词处理,将每一个候选串分词划分为若干候选串分词,并提取出这些候选串分词的分类特征。其中,候选串分词的分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频等等。
步骤S1303:根据提取出的分类特征筛选各个候选串分词是否是核心词。
提取出候选串分词的分类特征后,根据分类特征对候选串分词进行分类,并根据分类结果确定各个候选串分词是否是核心词。
如上所述,候选串分词的分类特征包括名词、热度词表、超链接、相关问题共现率、文档词频等特征中的至少一种,则可以候选串分词中所有的名词归为一类,将候选串分词中在热度词表中的分词归为一类,将候选串分词汇中是超级链接的分词归为一类,或者也可以将候选串分词中在热度词表中的所有名词归为一类,……,等等。
对候选串分词进行分类后,可以根据分类结果,进行核心词的筛选,比如,根据各个分类中各个候选串分词与用户输入的搜索词的匹配程度进行筛选,或者根据各个分类中各个候选串分词的使用频率统计值等因素进行筛选,或者综合考虑上述各种因素进行筛选。
其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。可以建立数据库,统计候选串分词被用户搜索的次数,被用户点击的次数曾经被确定为核心词的次数、曾经被用户用作搜索词的次数等。
实施例九
本发明实施例九提供的问答页面核心词提取方法,描述核心词提取的一种具体实现方式,其流程如图14所示,包括如下步骤:
步骤S1401:获取与用户输入的搜索词对应的问答页面。
例如:用户输入搜索词“孩子感冒咳嗽怎么办?”,根据该搜索词获取到对应的问答页面,获取到的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如相关问题可以是“小儿感冒咳嗽怎么办?”,“小儿感冒咳嗽用什么药比较好呢?”。
步骤S1402:从获取的问答页面的标题中提取核心词候选串。
本实施例中以从问答页面的标题中提取核心词候选串为例,比如,提取到的核心词候选串可以是“孩子感冒咳嗽怎么办”。
实际操作中还可以从问答页面的问答内容、相关问题等页面内容中提取核心词候选串。
步骤S1403:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
对提取的核心词候选串“孩子感冒咳嗽怎么办”进行分词,例如,可以分词为:“孩子”、“感冒”、“咳嗽”、“怎么办”等候选串分词。
对分词出的候选串分词进行分类特征提取,例如“孩子”这个候选串分词的分类特征包括:是名词等;“感冒”、“咳嗽”这两个候选串分词的分类特征包括:是名词、是热度词表中的词、是超链接等;“怎么办”这个候选串分词的分类特征包括是超链接等。
步骤S1404:根据提取的分类特征对候选串分词进行分类。
根据提取的分类特征对上述分词出的“孩子”、“感冒”、“咳嗽”、“怎么办”等候选 串分词进行分类,例如:“孩子”、“感冒”、“咳嗽”都是名词,归为一类;将“感冒”、“咳嗽”都是热度词表中的词,归为一类;“感冒”、“咳嗽”、“怎么办”都是超链接,归为一类。
步骤S1405:针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配。
对候选串分词进行分类后,分别针对每个分类,与用户输入的搜索词进行匹配。
沿用上边的例子,根据上边的分类,将名词分类、热度词表分类和超链接分类中的各个候选串分词分别与用户输入的搜索词进行匹配。
步骤S1406:筛选出匹配度最高的设定数量的候选串分词,作为核心词。
沿用上边的例子,筛选出匹配度较高的2个候选串分词为:“感冒”、“咳嗽”,则确定“感冒”、“咳嗽”为核心词;或筛选出匹配度较高的3个候选串分词为:“感冒”、“咳嗽”、“孩子”,则确定“感冒”、“咳嗽”、“孩子”为核心词。
上述实施例中所列举的搜索词、问答页面标题等都属于简单的举例,实际应用中用户输入的检索词可能会更简单,而根据问答页面获取到的候选串分词的数量可能会更多,匹配过程可能会更复杂,从而能够更好地发挥本发明方法的作用,在此不再一一列举。
上述步骤S1405和步骤S1406实现了根据分类结果确定各个候选串分词是否是核心词。
上述实施例九中的步骤S1405和步骤S1406可替换为下面步骤S1505和步骤S1506所公开的筛选方式。
实施例十
本发明实施例十提供的问答页面核心词提取方法,描述核心词提取的另一种具体实现方式,其流程如图15所示,包括如下步骤:
步骤S1501:获取与用户输入的搜索词对应的问答页面。
例如:用户输入搜索词“孩子感冒咳嗽怎么办?”,根据该搜索词获取到对应的问答页面,获取到的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如,问答答案中可能包括“选择正确的感冒(咳嗽)药”、“感冒止咳的中药”等描述,相关问题可以是“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”等问题。
步骤S1502:从获取的问答页面的页面内容中,提取与用户输入的搜索词相关的字符串,作为核心词候选串。
对用户输入的搜索词进行分词,从获取的问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
沿用上边的例子,对用户输入的搜索词“孩子感冒咳嗽怎么办?”进行分词,例如可以分词为“孩子”、“感冒”、“咳嗽”、“怎么办”等搜索词分词。
本实施例中以从问答页面的页面内容中提取核心词候选串为例,可以从问答页面的问答内容、相关问题等页面内容中提取包括“孩子”、“感冒”、“咳嗽”、“怎么办”中至少一个搜索词分词的字符串作为核心词候选串。例如,提取到的核心词候选串可以有:“孩子感冒咳嗽怎么办”、“选择正确的感冒(咳嗽)药”、“感冒止咳的中药”、“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”等等。
步骤S1503:对提取的核心词候选串进行分词,提取各个候选串分词的分类特征。
沿用上边的例子,对提取的核心词候选串“孩子感冒咳嗽怎么办”进行分词,例如,可以分词为:“孩子”、“感冒”、“咳嗽”、“怎么办”等候选串分词。对提取的核心词候选串“选择正确的感冒(咳嗽)药”进行分词,例如,可以分词为:“选择”、“正确的”、“感冒”、“咳嗽”、“药”等候选串分词。对提取的核心词候选串“感冒止咳的中药”进行分词,例如,可以分词为:“感冒”、“止咳”、“中药”等候选串分词。依次对提取的核心词候选串进行分词,此处不再一一列举。
对分词出的候选串分词进行分类特征提取,例如“孩子”这个候选串分词的分类特征包括:是名词等;“感冒”、“咳嗽”这两个候选串分词的分类特征包括:是名词、是热度词表中的词、是超链接等;“中药”、“药”这两个候选串分词的分类特征包括:是名词等;“止咳”这个候选串分词的分类特征包括:是热度词表中的词等;“怎么办”这个候选串分词的分类特 征包括:是超链接等。总之,对分词出的所有候选串分词都进行分类特征提取,此处不再对上边举例中的各候选串一一列举其分类特征。
步骤S1504:根据提取的分类特征对候选串分词进行分类。
根据提取的分类特征对上述分词出的“孩子”、“感冒”、“咳嗽”、“怎么办”、“选择”、“正确的”、“药”、“止咳”、“中药”等候选串分词进行分类,例如:“孩子”、“感冒”、“咳嗽”、“中药”、“药”都是名词,归为一类;将“感冒”、“咳嗽”、“止咳”都是热度词表中的词,归为一类;“感冒”、“咳嗽”、“怎么办”都是超链接,归为一类。总之,对分词出的所有候选串分词都根据分类特征进行分类,此处不再对上边举例中的各候选串一一列举其分类。
步骤S1505:针对每个分类,确定该分类中各个候选串分词的使用频率统计值。
沿用上边的例子,在名词分类中、热度词表中的词分类、超链接分类中,分别确定各候选串分词的使用频率统计值。
其中,候选串分词的使用频率统计值可以根据各候选串分词被用户搜索的次数、被用户点击的次数、曾经被确定为核心词的次数、曾经被作为搜索词的次数等因素中的至少一种因素进行统计。
步骤S1506:根据各个候选串分词的使用频率统计值,筛选出使用频率统计值最高的设定数量的候选串分词,作为核心词。
沿用上边的例子,筛选出使用频率统计值最高的3个候选串分词为:“感冒”、“咳嗽”、“止咳”,则确定“感冒”、“咳嗽”、“止咳”为核心词;或筛选出使用频率统计值最高的3个候选串分词为:“感冒”、“咳嗽”、“孩子”,则确定“感冒”、“咳嗽”、“孩子”为核心词。
上述步骤S1505和步骤S1506实现了根据分类结果确定各个候选串分词是否是核心词。
上述实施例三中的步骤S1505和步骤S1506可替换为步骤S1405和步骤S1406所公开的筛选方式。
基于同一发明构思,本发明实施例还提供一种问答页面核心词提取装置,该装置的结构如图16所示,包括:候选串提取模块1601、特征提取模块1602和核心词确定模块1603。
候选串提取模块1601,用于从问答页面中提取核心词候选串。
特征提取模块1602,用于对核心词候选串进行分词,提取各个候选串分词的分类特征。
核心词确定模块1603,用于根据提取的分类特征筛选各个候选串分词是否是核心词。
优选的,上述候选串提取模块1601,具体用于获取与用户输入的搜索词对应的问答页面,从获取的问答页面的标题中提取核心词候选串;和/或从获取的问答页面的页面内容中,提取与用户输入的搜索词相关的字符串,作为核心词候选串。
优选的,上述候选串提取模块1601,具体用于对所述搜索词进行分词,从获取的问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
优选的,上述核心词确定模块1603,具体用于根据提取的分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;其中,分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
优选的,上述核心词确定模块1603,具体用于针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为核心词;或针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出使用频率统计值最高的设定数量的候选串分词,作为核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
本发明实施例提供的上述问答页面核心词提取方法和装置,能够根据用户输入的搜索词对应的问答页面提取更符合用户搜索需求的核心词,从而能够根据核心词获取到与用户输入的搜索词相关度更高的相关问题,在当前问答页面为用户提供与用户需求的贴合性更好、更符合用户需求的相关问题,提高问答页面问题检索的准确性。
上述实施例提供的问答页面核心词提取方法和装置,从问答页面中提取核心词候选串,对 提取的核心词候选串进行分词,提取各个候选串分词的分类特征,根据分类特征筛选各个候选串分词是否是核心词,该方案从对问答页面的分析中实现核心词的提取,使所确定的核心词能够更好地反映用户输入的问题,与用户输入的问题相关性更高,从而能够根据提取的核心词获得更贴和用户需求、更符合用户需要的问答问题,获得用户真正想要获得的问题答案,提高了问答页面检索的准确性。
进一步地,能够根据用户输入的搜索词所对应的问答页面的标题或页面内容中提取核心词,从而使核心词的提取能够更准确、更贴合用户需要。且能够综合考虑各个候选串分类特征,根据不同类别的综合考量确定核心词,从而能够更客观、合理的确定出合适的核心词。
为了解决现有技术中存在的获取到的相关问题与用户输入的问题的相关度并不是很好,往往不能很好地满足用户的需求的问题,本发明实施例提供一种问答页面相关问题推荐方法,通过对若干历史用户的浏览行为进行分析,获取到与用户真正想要获得的问题答案匹配度更好地相关问题。
实施例十一
本发明实施例一提供一种问答页面相关问题推荐方法,该方法流程如图17所示,包括如下步骤:
步骤S1701:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
第一用户输入搜索词进行问答检索,生成问答页面时,生成的问答页面中包括但不限于问答页面的标题、至少一个问题答案,至少一个相关问题。在获取到第一用户输入的搜索词后,从数据库中获取若干相关问题,这些相关问题为数据库中第二用户浏览的问答页面中的问答问题或问答页面中的相关问题。
其中,第一用户是指当前用户,第二用户是指历史用户。
步骤S1702:根据选定时间段内第二用户的浏览行为日志,确定获取的相关问题的浏览权重。
从数据库中获取上述步骤S1701中获取到的相关问题对应的第二用户的浏览行为日志。对浏览行为日志进行分析,确定相关问题的浏览权重。确定浏览权重的过程中,可以对获取的相关问题,计算彼此之间的相关浏览权重,根据计算出来的相关浏览权重,对同一相关问题的相关浏览权重进行加权处理,得到各相关问题的浏览权重。
优选的,也可以根据设定的分组条件对获取的相关问题进行分组,在各个相关问题分组中,分别计算各相关问题与组中其他相关问题的相关浏览权重,然后综合各组的计算结果,对各组中出现的同一相关问题的相关浏览权重进行加权处理,得到各相关问题的浏览权重。
下面的实施例十二中,以根据浏览用户进行分组为例,说明相关问题的浏览权重的确定过程。
步骤S1703:根据确定的浏览权重对获取的相关问题进行排序。
根据确定出的各相关问题的浏览权重,对各相关问题进行排序。比如可以按照浏览权重从高到低的顺序进行排序。对相关问题进行排序时,可以对获取所有的相关问题一起进行排序,也可以按照不同的浏览用户在个浏览用户分组中分别排序,或者按照其他的规则排序。
步骤S1704:根据获取的相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
根据对相关问题的排序结果,按照设定的推荐规则,选择相关问题推荐给用户。比如,将获取所有的相关问题中浏览权重最高的设定数量的相关问题推荐给用户;或者在各浏览用户对应的相关问题中分别获取设定数量的相关问题推荐给第一用户。
实施例十二
本发明实施例十二提供一种问答页面相关问题推荐方法,该方法流程如图18所示,包括如下步骤:
步骤S1801:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
例如:第一用户输入搜索词“孩子感冒怎么办?”,根据该搜索词生成对应的问答页面,生成的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如:相关问题可以是“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽怎么办”“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”等等。
这些相关问题为数据库中存储的历史用户曾经浏览过的问答页面上的问答问题或问答页面上的相关问题。
步骤S1802:根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组。
对获取的相关问题进行分组时,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题。
可选的,根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
其中,浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间等。
沿用上边的例子,对上边获取到的各相关问题进行分组如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”被同一个浏览用户浏览过,归为一组。
“小儿感冒发烧怎么办?”、“儿童感冒发烧怎么办”、“小儿感冒鼻塞怎么办?”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”被同一个浏览用户浏览过,归为一组。
“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”被同一个浏览用户浏览过,归为一组。
……
以此类推,对所有获取的相关问题进行分组,实现将被同一用户浏览过的相关问题归为一组。
步骤S1803:在各相关问题分组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
根据上述各浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},利用如下公式计算每个相关问题与组中其它相关问题的相关浏览权重W(Ti,Ti+1):
log(a1/(|Time(i)–Time(i+1)|+a2))
其中,Time(i)一个问答问题的用户浏览时间;
Time(i+1)为组中其它问答问题的用户浏览时间;
a1,a2为经验值常数。
当然,也可以计算组中各相关问题Ti与组中其他相关问题Ti-1的相关浏览权重W。
需要说明的是,上述各个公式并不是实现本发明的唯一公式,仅作为实施例的一种实现方式。技术人员可以根据业务需要对公式做适当变形,例如增加常量或变量或系数等方式,依然落在本发明的保护范围之内。
沿用上边的例子,针对每个分组,分别计算每个相关问题与组中其他相关问题的,例如,针对浏览用户相同的第一个相关问题分组,分别计算“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”与组中其他相关问题的相关浏览权重。其他相关问题分组也同样进行计算。
进一步可选的,计算组中各相关问题与组中其它相关问题的相关浏览权重,包括:在各相关问题分组中,根据浏览用户浏览各相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会 话组;在各会话组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
也就是说,对于浏览用户相同的相关问题分组中的用户,可以进一步根据浏览时间划分出不同的会话组(session),同一会话组中的相关问题的浏览时间差小于等于某个设定的时间阈值。可以根据浏览用户的浏览特征向量进行session划分。在同一session内,计算相关问题的浏览权重。
步骤S1804:获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的各相关问题的浏览权重。
上边计算出各相关问题分组中的各相关问题的相关浏览权重后,将各相关问题分组中相同的相关问题提取出来,例如,对于“小儿感冒鼻塞怎么办?”这个相关问题,在浏览用户相同的第一个相关问题分组和第三个相关问题中计算得到的相关浏览权重进行加权。
可选的,可以把同一相关问题在不同相关问题分组中计算得到的相关浏览权重直接进行相加,也可以分别乘上相应的权重系数后在进行相加,也可以通过其它的加权规则进行加权处理。
步骤S1805:根据确定出的相关问题的浏览权重对获取的相关问题进行排序。
沿用上边的例子,以获取所有的相关问题一起进行排序为例,按照浏览权重从高到低的顺序进行排序,得到排序结果如下:
“小儿感冒发烧怎么办?”、“小儿感冒咳嗽怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”。
步骤S1806:根据获取的相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
根据排序结果,将浏览权重最高的前几个问题作为相关问题推荐给第一用户,加入到根据用户输入的搜索词生成的问答页面中。
例如:将“小儿感冒发烧怎么办?”、“小儿感冒咳嗽怎么办?”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”作为相关问题加入到问答页面中。
上述方法,根据数据库中的历史数据,分析历史用户浏览各个相关问题的浏览行为,确定相关问题的浏览权重参数,从而确定向用户推荐相关问题的推荐优先级,从而获取到与用户输入的搜索词匹配度更高的相关问题,在当前问答页面为用户提供与用户需求的贴合性更好、更符合用户需求的相关问题,提高问答页面问题检索的准确性。
针对本发明实施例提供的问答页面相关问题推荐方法,实现问答页面相关问题推荐的系统环境示意如图19所示。该系统包括数据库,存储若干第二用户(历史用户)的相关问题,问答页面问题推荐装置能够获取第一用户输入的搜索词,并根据搜索词从数据库获取若干历史用户浏览过的相关问题及相关问题的历史数据,通过对历史数据的分析处理,实现获取更优的相关问题推荐给第一用户。
基于同一发明构思,本发明实施例还提供一种问答页面相关问题推荐装置,该装置的结构如图20所示,包括:问题获取模块2001、权重确定模块2002和排序推荐模块2003。
问题获取模块2001,用于根据来自第一用户的搜索词,获取数据库中与搜索词相关的至少一个相关问题.
权重确定模块2002,根据选定时间段内第二用户的浏览行为日志,确定获取的相关问题的浏览权重。
排序推荐模块2003,用于根据确定出的浏览权重对获取的相关问题进行排序;根据获取的相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
优选的,上述权重确定模块2002,具体包括:问题分组器20021、相关权重计算器20022和浏览权重计算器20023。
问题分组器20021,用于根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题。
相关权重计算器20022,用于在各相关问题分组中,计算组中各相关问题与组中其它相关 问题的相关浏览权重。
浏览权重计算器20023,用于获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的各相关问题的浏览权重。
优选的,上述问题分组器20021,具体用于根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},实现对获取的相关问题进行分组;其中,Ti表示一个相关问题。
优选的,上述相关权重计算器20022,具体用于在各相关问题分组中,根据浏览用户浏览各相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;在每个会话组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
优选的,上述问题分组器20021,具体用于得到的浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间等。
上述实施例提供的问答页面相关问题推荐方法和装置,当需要为输入搜索词的第一用户获取相关问题生成问答页面时,根据一段时间内的若干第二用户的浏览行为日志确定获取的相关问题的浏览权重,根据浏览权重获取较佳的相关问题,从而获取到与第一用户输入的问题相关度更好地相关问题,使获取的相关问题与用户真正想要获得的问题答案之间的匹配度更好,能够更好地满足用户需求,使用户在问答页面上查看到与所检索的问题更贴近的、更吻合的问题答案。
进一步地,能够根据不同的浏览用户,按照分组计算各相关问题的相关浏览权重,从而获取到各相关问题的浏览权重,实现基于若干第二用户对各相关问题的浏览行为,来衡量获取的各相关问题对用户需求的满足匹配度高低,从而达到了获取匹配度更好的相关问题的目的。
为了解决现有技术中存在的获取到的相关问题与用户输入的问题的相关度并不是很好,往往不能很好地满足用户的需求的问题,本发明实施例提供一种问答页面相关问题推荐方法,通过对若干历史用户的搜索行为进行分析,获取到与用户真正想要获得的问题答案匹配度更好地相关问题。
实施例十三
本发明实施例十三提供一种问答页面相关问题推荐方法,该方法流程如图21所示,包括如下步骤:
步骤S2101:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
第一用户输入搜索词进行问答检索,生成问答页面时,生成的问答页面中包括但不限于问答页面的标题、至少一个问题答案,至少一个相关问题。在获取到第一用户输入的搜索词后,从数据库中获取若干相关问题,这些相关问题为数据库中第二用户搜索点击的问答页面中的问答问题或问答页面中的相关问题。
其中,第一用户是指当前用户,第二用户是指历史用户。
步骤S2102:根据选定时间段内第二用户的搜索行为日志,确定获取的相关问题的点击权重。
从数据库中获取上述步骤S2101中获取到的相关问题对应的第二用户的搜索行为日志。对搜索行为日志进行分析,确定相关问题的点击权重。确定击权重的过程中,可以对获取的相关问题,计算彼此之间的相关点击权重,根据计算出来的相关点击权重,对同一相关问题的相关点击权重进行加权处理,得到各相关问题的点击权重。
优选的,也可以根据设定的分组条件对获取的相关问题进行分组,在各个相关问题分组中,分别计算各相关问题与组中其他相关问题的相关点击权重,然后综合各组的计算结果,对各组中出现的同一相关问题的相关点击权重进行加权处理,得到各相关问题的点击权重。
下面的实施例十四中,以根据查询请求串进行分组为例,说明相关问题的点击权重的确定 过程。
步骤S2103:根据确定出的相关问题的点击权重对获取的相关问题进行排序。
根据确定出的各相关问题的点击权重,对各相关问题进行排序。比如可以按照点击权重从高到低的顺序进行排序。对相关问题进行排序时,可以对获取所有的相关问题一起进行排序,也可以按照不同的查询请求串在各查询串分组中分别排序,或者按照其他的规则排序。
步骤S2104:根据获取的相关问题的排序结果,确定推荐给第一用户的相关问题。
根据对相关问题的排序结果,按照设定的推荐规则,选择相关问题推荐给第一用户。比如,将获取所有的相关问题中点击权重最高的设定数量的相关问题推荐给第一用户;或者在各查询请求串对应的相关问题中分别获取设定数量的相关问题推荐给第一用户。
实施例十四
本发明实施例十四提供一种问答页面相关问题推荐方法,该方法流程如图22所示,包括如下步骤:
步骤S2201:根据来自第一用户的搜索词,获取数据库中与来自第一用户的搜索词相关的至少一个相关问题。
例如:第一用户输入搜索词“孩子感冒怎么办?”,根据该搜索词生成对应的问答页面,生成的问答页面上有问答页面的标题,至少一个问题答案,至少一个相关问题。比如:相关问题可以是“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽怎么办”“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”等等。
这些相关问题为数据库中存储的历史用户曾经搜索过的问答页面上的问答问题或问答页面上的相关问题。
步骤S2202:根据获取的相关问题对应的查询请求串,对获取的相关问题进行分组。
对获取的相关问题进行分组时,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题。
可选的,根据获取的相关问题对应的查询请求串,得到各查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中Ti表示一个相关问题。从而实现对获取的相关问题进行分组。
其中,点击特征向量中的元素Ti的属性包括下列参数中的至少一个:问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
沿用上边的例子,对上边获取到的各相关问题进行分组如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”对应的查询请求串为“小儿感冒”,归为一组。
“宝宝感冒咳嗽怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”对应的查询请求串为“宝宝感冒”,归为一组;
“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”、“儿童感冒发烧怎么办”对应的查询请求串为“儿童感冒”,归为一组;
“小儿感冒咳嗽怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”,“宝宝感冒咳嗽怎么办”,“宝宝感冒咳嗽流鼻涕怎么办”、“宝宝感冒咳嗽用什么药比较好呢?”、“儿童感冒咳嗽怎么办”对应的查询请求串为“感冒咳嗽”,归为一组;
“小儿感冒发烧怎么办?”、“小儿感冒发烧怎么办?”、“儿童感冒发烧怎么办”对应的查询请求串为“感冒发烧”,归为一组;
“小儿感冒鼻塞怎么办?”、“宝宝感冒鼻塞怎么办”、“儿童感冒鼻塞怎么办”对应的查询请求串为“感冒鼻噻”,归为一组;
……
以此类推,对所有获取的相关问题进行分组,实现将查询请求串相同的相关问题归为一组。
步骤S2203:在各相关问题分组中,计算组中各相关问题与组中其他相关问题的相关点击 权重。
根据上述生成的各查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},利用如下公式计算组中各相关问题Ti与组中其他相关问题Ti+1的相关点击权重W(Ti,Ti+I):
W=P((Ti)|查询请求串)*P((Ti+I)|查询请求串)
其中,Ti表示一个相关问题;
Ti+I表示点击特征向量中包括的其他问答问题;
P((Ti)|查询请求串)表示使用查询请求串时得到Ti的概率;
P((Ti+I)|查询请求串)表示使用查询请求串时得到Ti+I的概率。
当然,也可以计算组中各相关问题Ti与组中其他相关问题Ti-I的相关点击权重W。
需要说明的是,上述各个公式并不是实现本发明的唯一公式,仅作为实施例的一种实现方式。技术人员可以根据业务需要对公式做适当变形,例如增加常量或变量或系数等方式,依然落在本发明的保护范围之内。
沿用上边的例子,针对每个分组,分别计算每个相关问题与组中其他相关问题的,例如,针对查询请求串为“小儿感冒”的相关问题分组,分别计算“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”与组中其他相关问题的相关点击权重。其他相关问题分组也同样进行计算。
步骤S2204:获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的各相关问题的点击权重。
上边计算出各相关问题分组中的各相关问题的相关点击权重后,将各相关问题分组中相同的相关问题提取出来,例如,对于“小儿感冒咳嗽怎么办?”这个相关问题,在查询请求串为“小儿感冒”的相关问题分组和在查询请求串为“感冒咳嗽”的相关问题分组中计算得到的相关点击权重进行加权。
可选的,可以把同一相关问题在不同相关问题分组中计算得到的相关点击权重直接进行相加,也可以分别乘上相应的权重系数后在进行相加,也可以通过其它的加权规则进行加权处理。
步骤S2205:根据确定出的相关问题的点击权重对获取的相关问题进行排序。
沿用上边的例子,以获取所有的相关问题一起进行排序为例,按照点击权重从高到低的顺序进行排序,得到排序结果如下:
“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”、“小儿感冒咳嗽用什么药比较好呢?”、“小儿感冒鼻塞怎么办?”、“宝宝感冒咳嗽用什么药比较好呢?”、“宝宝感冒鼻塞怎么办”、“儿童感冒咳嗽怎么办”、“儿童感冒鼻塞怎么办”。
步骤S2206:根据获取的相关问题的排序结果,确定推荐给第一用户的相关问题。
根据排序结果,将点击权重最高的前几个问题作为相关问题推荐给第一用户,加入到根据用户输入的搜索词生成的问答页面中。
例如:将“小儿感冒咳嗽怎么办?”、“小儿感冒发烧怎么办?”、“宝宝感冒咳嗽怎么办”“儿童感冒发烧怎么办”、“宝宝感冒咳嗽流鼻涕怎么办”作为相关问题加入到问答页面中。
上述方法,根据数据库中的历史数据,分析历史用户点击各个相关问题的搜索点击行为,确定相关问题的点击权重参数,从而确定向用户推荐相关问题的推荐优先级,从而获取到与用户输入的搜索词匹配度更高的相关问题,在当前问答页面为用户提供与用户需求的贴合性更好、更符合用户需求的相关问题,提高问答页面问题检索的准确性。
针对本发明实施例提供的问答页面相关问题推荐方法,实现问答页面相关问题推荐的系统环境示意如图23所示。该系统包括数据库,存储若干第二用户(历史用户)的相关问题,问答页面问题推荐装置能够获取第一用户输入的搜索词,并更具搜索词从数据库获取若干历史用户搜索点击过的相关问题及相关问题的历史数据,通过对历史数据的分析处理,实现获取更优的相关问题推荐给第一用户。
基于同一发明构思,本发明实施例还提供一种问答页面相关问题推荐装置,该装置的结构如图24所示,包括:问题获取模块2401、权重确定模块2402和排序推荐模块2403。
问题获取模块2401,用于根据来自第一用户的搜索词,获取数据库中与搜索词相关的至少一个相关问题.
权重确定模块2402,用于根据选定时间段内第二用户的搜索行为日志,确定获取的相关问题的点击权重。
排序推荐模块2403,用于根据确定出的点击权重对获取的相关问题进行排序;根据获取的相关问题的排序结果,确定问答页面推荐给第一用户的相关问题。
优选的,上述权重确定模块2402,具体包括:问题分组器24021、相关权重计算器24022和点击权重计算器24023。
问题分组器24021,用于根据获取的相关问题对应的查询请求串,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题。
相关权重计算器24022,用于在各相关问题分组中,计算组中各相关问题与组中其他相关问题的相关点击权重。
点击权重计算器24023,用于获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的各相关问题的点击权重。
优选的,上述问题分组器24021,具体用于根据获取的相关问题对应的查询请求串,得到每个查询请求串的点击特征向量{T1、T2、……、Tn},实现对获取的相关问题进行分组;其中Ti表示一个相关问题。
优选的,上述相关权重计算器24022,具体用于利用如下公式计算组中各相关问题与组中其他相关问题的相关点击权重W:
W=P((Ti)|查询请求串)*P((Ti+I)|查询请求串)
其中,Ti表示一个相关问题;
Ti+I表示点击特征向量中包括的其他问答问题;
P((Ti)|查询请求串)表示使用查询请求串时得到Ti的概率;
P((Ti+I)|查询请求串)表示使用查询请求串时得到Ti+1的概率。
优选的,上述问题分组器24021,具体用于得到的点击特征向量中的元素Ti的属性包括下列参数中的至少一个:问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
需要说明的是,上述各个公式并不是实现本发明的唯一公式,仅作为实施例的一种实现方式。技术人员可以根据业务需要对公式做适当变形,例如增加常量或变量或系数等方式,依然落在本发明的保护范围之内。
上述施例提供的问答页面相关问题推荐方法和装置,当需要为输入搜索词的第一用户获取相关问题生成问答页面时,根据一段时间内的若干第二用户的搜索行为日志确定获取的相关问题的点击权重,根据点击权重获取较佳的相关问题,从而获取到与第一用户输入的问题相关度更好地相关问题,使获取的相关问题与用户真正想要获得的问题答案之间的匹配度更好,能够更好地满足用户需求,使用户在问答页面上查看到与所检索的问题更贴近的、更吻合的问题答案。
进一步地,能够根据不同的查询请求串,按照分组计算各相关问题的相关点击权重,从而获取到各相关问题的点击权重,实现基于若干第二用户对各相关问题的搜索点击行为,来衡量获取的各相关问题对用户需求的满足匹配度高低,从而达到了获取匹配度更好的相关问题的目的。
在此处所提供的说明书中,说明了大量具体细节。然而,能够理解,本发明的实施例可以在没有这些具体细节的情况下实践。在一些实例中,并未详细示出公知的方法、结构和技术,以便不模糊对本说明书的理解。
类似地,应当理解,为了精简本公开并帮助理解各个发明方面中的一个或多个,在上面对本发明的示例性实施例的描述中,本发明的各个特征有时被一起分组到单个实施例、图、或者对其的描述中。然而,并不应将该公开的方法解释成反映如下意图:即所要求保护的本发明要求比在每个权利要求中所明确记载的特征更多的特征。更确切地说,如下面的权利要求书所反映的那样,发明方面在于少于 前面公开的单个实施例的所有特征。因此,遵循具体实施方式的权利要求书由此明确地并入该具体实施方式,其中每个权利要求本身都作为本发明的单独实施例。
本领域那些技术人员可以理解,可以对实施例中的设备中的模块进行自适应性地改变并且把它们设置在与该实施例不同的一个或多个设备中。可以把实施例中的模块或单元或组件组合成一个模块或单元或组件,以及此外可以把它们分成多个子模块或子单元或子组件。除了这样的特征和/或过程或者单元中的至少一些是相互排斥之外,可以采用任何组合对本说明书(包括伴随的权利要求、摘要和附图)中公开的所有特征以及如此公开的任何方法或者设备的所有过程或单元进行组合。除非另外明确陈述,本说明书(包括伴随的权利要求、摘要和附图)中公开的每个特征可以由提供相同、等同或相似目的的替代特征来代替。
此外,本领域的技术人员能够理解,尽管在此所述的一些实施例包括其它实施例中所包括的某些特征而不是其它特征,但是不同实施例的特征的组合意味着处于本发明的范围之内并且形成不同的实施例。例如,在下面的权利要求书中,所要求保护的实施例的任意之一都可以以任意的组合方式来使用。
本发明的各个部件实施例可以以硬件实现,或者以在一个或者多个处理器上运行的软件模块实现,或者以它们的组合实现。本领域的技术人员应当理解,可以在实践中使用微处理器或者数字信号处理器(DSP)来实现根据本发明实施例的***设备中的一些或者全部部件的一些或者全部功能。本发明还可以实现为用于执行这里所描述的方法的一部分或者全部的设备或者装置程序(例如,计算机程序和计算机程序产品)。这样的实现本发明的程序可以存储在计算机可读介质上,或者可以具有一个或者多个信号的形式。这样的信号可以从因特网网站上下载得到,或者在载体信号上提供,或者以任何其他形式提供。
例如,图25示出了可以实现根据本发明的问答页面相关问题推荐方法的计算设备。该计算设备传统上包括处理器2510和以存储器2520形式的计算机程序产品或者计算机可读介质。存储器2520可以是诸如闪存、EEPROM(电可擦除可编程只读存储器)、EPROM、硬盘或者ROM之类的电子存储器。存储器2520具有用于执行上述方法中的任何方法步骤的程序代码2531的存储空间2530。例如,用于程序代码的存储空间2530可以包括分别用于实现上面的方法中的各种步骤的各个程序代码2531。这些程序代码可以从一个或者多个计算机程序产品中读出或者写入到这一个或者多个计算机程序产品中。这些计算机程序产品包括诸如硬盘,紧致盘(CD)、存储卡或者软盘之类的程序代码载体。这样的计算机程序产品通常为如参考图26所述的便携式或者固定存储单元。该存储单元可以具有与图25的计算设备中的存储器2520类似布置的存储段、存储空间等。程序代码可以例如以适当形式进行压缩。通常,存储单元包括计算机可读代码2531’,即可以由例如诸如2510之类的处理器读取的代码,这些代码当由计算设备运行时,导致该计算设备执行上面所描述的方法中的各个步骤。
本文中所称的“一个实施例”、“实施例”或者“一个或者多个实施例”意味着,结合实施例描述的特定特征、结构或者特性包括在本发明的至少一个实施例中。此外,请注意,这里“在一个实施例中”的词语例子不一定全指同一个实施例。
应该注意的是上述实施例对本发明进行说明而不是对本发明进行限制,并且本领域技术人员在不脱离所附权利要求的范围的情况下可设计出替换实施例。在权利要求中,不应将位于括号之间的任何参考符号构造成对权利要求的限制。单词“包含”不排除存在未列在权利要求中的元件或步骤。位于元件之前的单词“一”或“一个”不排除存在多个这样的元件。本发明可以借助于包括有若干不同元件的硬件以及借助于适当编程的计算机来实现。在列举了若干装置的单元权利要求中,这些装置中的若干个可以是通过同一个硬件项来具体体现。单词第一、第二、以及第三等的使用不表示任何顺序。可将这些单词解释为名称。
此外,还应当注意,本说明书中使用的语言主要是为了可读性和教导的目的而选择的,而不是为了解释或者限定本发明的主题而选择的。因此,在不偏离所附权利要求书的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。对于本发明的范围,对本发明所做的公开是说明性的,而非限制性的,本发明的范围由所附权利要求书限定。

Claims (68)

  1. 一种问答页面相关问题推荐方法,包括:
    根据来自用户的搜索词,获取数据库与所述搜索词相关的至少一个相关问题;
    根据至少一个预设规则对获取的所述相关问题进行筛选;
    根据所述相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
  2. 根据权利要求1所述的方法,其中,所述至少一个预设规则包括下列至少之一:
    根据核心词对所述相关问题进行筛选;
    根据用户的浏览行为日志对所述相关问题进行筛选;
    根据用户的搜索点击行为日志对所述相关问题进行筛选。
  3. 根据权利要求1-2任一项所述的方法,其中,所述根据核心词对所述相关问题进行筛选,包括:
    获取与所述搜索词对应的问答页面;
    提取所述问答页面中的核心词,并根据所述核心词筛选所述相关问题。
  4. 根据权利要求1-3任一项所述的方法,其中,提取所述问答页面中的核心词,包括:
    从问答页面中提取核心词候选串;
    对所述核心词候选串进行分词,提取各个候选串分词的分类特征;
    根据所述分类特征筛选各个候选串分词是否是核心词。
  5. 根据权利要求1-4任一项所述的方法,其中,从问答页面中提取核心词候选串,包括:
    从所述问答页面的标题中提取核心词候选串;和/或
    从所述问答页面的页面内容中,提取与所述搜索词相关的字符串,作为核心词候选串。
  6. 根据权利要求1-5任一项所述的方法,其中,提取与所述搜索词相关的字符串,包括:
    对所述搜索词进行分词;
    从所述问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
  7. 根据权利要求1-6任一项所述的方法,其中,根据所述分类特征筛选各个候选串分词是否是核心词,包括:
    根据所述分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;
    所述分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
  8. 根据权利要求1-7任一项所述的方法,其中,根据分类结果确定各个候选串分词是否是核心词,具体包括:
    针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为所述核心词;
    针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出所述使用频率统计值最高的设定数量的候选串分词,作为所述核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
  9. 根据权利要求1-8任一项所述的方法,其中,所述根据用户的浏览行为日志对所述相关问题进行筛选,包括:
    根据选定时间段内用户的浏览行为日志,确定获取的所述相关问题的浏览权重;
    根据所述浏览权重对所述相关问题进行排序;
    根据排序结果对所述相关问题进行筛选。
  10. 根据权利要求1-9任一项所述的方法,其中,所述根据排序结果对所述相关问题进行筛选,包括:
    根据所述排序结果提取第一预定个数个所述相关问题。
  11. 根据权利要求1-10任一项所述的方法,其中,所述根据选定时间段内用户的浏览行为日志,确定获取的所述相关问题的浏览权重,包括:
    根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题;
    在每个相关问题分组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重;
    获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的每个相关问题的浏览权重。
  12. 根据权利要求1-11任一项所述的方法,其中,根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组,包括:
    根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
  13. 根据权利要求1-12任一项所述的方法,其中,计算组中每个相关问题与组中其它相关问题的相关浏览权重,包括:
    在每个相关问题分组中,根据浏览用户浏览每个相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;
    根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;
    在每个会话组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重。
  14. 根据权利要求1-13任一项所述的方法,其中,所述浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间等。
  15. 根据权利要求1-14任一项所述的方法,其中,所述根据用户的搜索点击行为日志对所述相关问题进行筛选,包括:
    根据选定时间段内用户的搜索点击日志,确定获取的所述相关问题的点击权重;
    根据所述点击权重对获取的相关问题进行排序;
    根据排序结果对所述相关问题进行筛选。
  16. 根据权利要求1-15任一项所述的方法,其中,所述根据排序结果对所述相关问题进行筛选,包括:
    根据所述排序结果提取第二预定个数个所述相关问题。
  17. 根据权利要求1-16任一项所述的方法,其中,根据设定时间段内用户的搜索点击日志,确定获取的所述相关问题的点击权重,包括:
    根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题;
    在每个相关问题分组中,计算组中每个相关问题与组中其他相关问题的相关点击权重;
    获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的每个相关问题的点击权重。
  18. 根据权利要求1-17任一项所述的方法,其中,根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组,包括:
    根据所述相关问题对应的查询请求串,得到每个查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
  19. 根据权利要求1-18任一项所述的方法,其中,点击特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
  20. 一种问答页面核心词提取方法,包括:
    从问答页面中提取核心词候选串;
    对所述核心词候选串进行分词,提取各个候选串分词的分类特征;
    根据所述分类特征筛选各个候选串分词是否是核心词。
  21. 根据权利要求20所述的方法,其中,从问答页面中提取核心词候选串,包括:
    获取与用户输入的搜索词对应的问答页面;
    从所述问答页面的标题中提取核心词候选串;和/或从所述问答页面的页面内容中,提取与所述搜索词相关的字符串,作为核心词候选串。
  22. 根据权利要求20-21任一项所述的方法,其中,提取与所述搜索词相关的字符串,包括:
    对所述搜索词进行分词;
    从所述问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
  23. 根据权利要求20-22任一项所述的方法,其中,根据所述分类特征筛选各个候选串分词是否是核心词,包括:
    根据所述分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;
    所述分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
  24. 根据权利要求20-23任一项所述的方法,其中,根据分类结果确定各个候选串分词是否是核心词,具体包括:
    针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为所述核心词;
    针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出所述使用频率统计值最高的设定数量的候选串分词,作为所述核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
  25. 一种问答页面相关问题推荐方法,包括:
    根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;
    根据选定时间段内第二用户的浏览行为日志,确定获取的所述相关问题的浏览权重;
    根据所述浏览权重对获取的相关问题进行排序;
    根据所述相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
  26. 根据权利要求25所述的方法,其中,根据选定时间段内第二用户的浏览行为日志,确定获取的所述相关问题的浏览权重,包括:
    根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题;
    在各相关问题分组中,计算组中各相关问题与组中其它相关问题的相关浏览权重;
    获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的各相关问题的浏览权重。
  27. 根据权利要求25-26任一项所述的方法,其中,根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组,包括:
    根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中,Ti表示一个相关问题。
  28. 根据权利要求25-27任一项所述的方法,其中,计算组中各相关问题与组中其它相关问题的相关浏览权重,包括:
    在各相关问题分组中,根据浏览用户浏览各相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;
    根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;
    在各会话组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
  29. 根据权利要求25-28任一项所述的方法,其中,浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间。
  30. 一种问答页面相关问题推荐方法,包括:
    根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;
    根据选定时间段内第二用户的搜索行为日志,确定获取的所述相关问题的点击权重;
    根据所述点击权重对获取的相关问题进行排序;
    根据所述相关问题的排序结果,确定问答页面推荐给第一用户的相关问题。
  31. 根据权利要求30所述的方法,其中,根据设定时间段内第二用户的搜索行为日志,确定获取的所述相关问题的点击权重,包括:
    根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题;
    在各相关问题分组中,计算组中各相关问题与组中其他相关问题的相关点击权重;
    获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的各相关问题的点击权重。
  32. 根据权利要求30-31任一项所述的方法,其中,根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组,包括:
    根据所述相关问题对应的查询请求串,得到各查询请求串的点击特征向量{T1、T2、……、Ti、Ti+1、……、Tn},其中Ti表示一个相关问题。
  33. 根据权利要求30-32任一项所述的方法,其中,计算组中各相关问题与组中其他相关问题的相关点击权重,包括:
    利用如下公式计算组中各相关问题与组中其他相关问题的相关点击权重W:
    W=P((Ti)|查询请求串)*P((Ti+I)|查询请求串)
    其中,Ti表示一个相关问题;
    Ti+I表示点击特征向量中包括的其他问答问题;
    P((Ti)|查询请求串)表示使用查询请求串时得到Ti的概率;
    P((Ti+I)|查询请求串)表示使用查询请求串时得到Ti+1的概率。
  34. 根据权利要求30-33任一项所述的方法,其中,点击特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数。
  35. 一种问答页面相关问题推荐装置,包括:
    获取器,适于根据来自用户的搜索词,获取数据库与所述搜索词相关的至少一个相关问题;
    筛选器,适于根据至少一个预设规则对获取的所述相关问题进行筛选;
    推荐器,适于根据所述相关问题的筛选结果,确定问答页面推荐给用户的相关问题。
  36. 根据权利要求35所述的装置,其中,所述筛选器还包括:
    第一筛选模块,适于根据用户的浏览行为日志对所述相关问题进行筛选;
    第二筛选模块,适于根据用户的搜索点击行为日志对所述相关问题进行筛选;
    第三筛选模块,适于根据核心词对所述相关问题进行筛选。
  37. 根据权利要求35-36任一项所述的装置,其中,所述第三筛选模块还包括:
    获取单元,适于获取与所述搜索词对应的问答页面;
    提取单元,适于提取所述问答页面中的核心词;
    确定单元,适于根据所述核心词筛选所述相关问题。
  38. 根据权利要求35-37-任一项所述的装置,其中,所述提取单元还适于:
    从问答页面中提取核心词候选串;
    对所述核心词候选串进行分词,提取各个候选串分词的分类特征;
    根据所述分类特征筛选各个候选串分词是否是核心词。
  39. 根据权利要求35-38任一项所述的装置,其中,所述提取单元还适于:
    从所述问答页面的标题中提取核心词候选串;和/或
    从所述问答页面的页面内容中,提取与所述搜索词相关的字符串,作为核心词候选串。
  40. 根据权利要求35-39任一项所述的装置,其中,所述提取单元还适于:
    对所述搜索词进行分词;
    从所述问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
  41. 根据权利要求35-40任一项所述的装置,其中,所述提取单元还适于:
    根据所述分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;
    所述分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
  42. 根据权利要求35-41任一项所述的装置,其中,所述提取单元还适于:
    针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为所述核心词;
    针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出所述使用频率统计值最高的设定数量的候选串分词,作为所述核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
  43. 根据权利要求35-42任一项所述的装置,其中,所述第一筛选模块还包括:
    第一权重确定单元,适于根据选定时间段内用户的浏览行为日志,确定获取的所述相关问题的浏览权重;
    第一排序单元,适于根据所述浏览权重对获取的相关问题进行排序;
    第一筛选单元,适于根据排序结果对所述相关问题进行筛选。
  44. 根据权利要求35-43任一项所述的装置,其中,所述第一筛选单元还适于:
    根据所述排序结果提取第一预定个数个所述相关问题。
  45. 根据权利要求35-44任一项所述的装置,其中,所述第一权重确定单元还适于:
    根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题;
    在每个相关问题分组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重;
    获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的每个相关问题的浏览权重。
  46. 根据权利要求35-45任一项所述的装置,其中,所述第一权重确定单元还适于:
    根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},实现对获取的相关问题进行分组;其中,Ti表示一个相关问题。
  47. 根据权利要求35-46任一项所述的装置,其中,所述第一权重确定单元还适于:
    在每个相关问题分组中,根据浏览用户浏览每个相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;
    根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;
    在每个会话组中,计算组中每个相关问题与组中其它相关问题的相关浏览权重。
  48. 根据权利要求35-47任一项所述的装置,其中,所述第二筛选模块还包括:
    第二权重确定单元,适于根据选定时间段内用户的搜索点击日志,确定获取的所述相关问题的点击权重;
    第二排序单元,适于根据所述点击权重对获取的相关问题进行排序;
    第二筛选单元,适于根据排序结果对所述相关问题进行筛选。
  49. 根据权利要求35-48任一项所述的装置,其中,所述第二权重确定单元还适于:
    根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题;
    在每个相关问题分组中,计算组中每个相关问题与组中其他相关问题的相关点击权重;
    获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的每个相关问题的点击权重。
  50. 根据权利要求35-49任一项所述的装置,其中,所述第二权重确定单元还适于:
    根据所述相关问题对应的查询请求串,得到每个查询请求串的点击特征向量{T1、T2、……、Tn},实现对获取的相关问题进行分组;其中,Ti表示一个相关问题。
  51. 根据权利要求35-50任一项所述的装置,其中,所述第二权重确定单元还适于:
    得到的点击特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数等。
  52. 一种问答页面核心词提取装置,包括:
    候选串提取模块,用于从问答页面中提取核心词候选串;
    特征提取模块,用于对核心词候选串进行分词,提取各个候选串分词的分类特征;
    核心词确定模块,用于根据所述分类特征筛选各个候选串分词是否是核心词。
  53. 根据权利要求52所述的装置,其中,所述候选串提取模块,具体用于:
    获取与用户输入的搜索词对应的问答页面;
    从所述问答页面的标题中提取核心词候选串;和/或从所述问答页面的页面内容中,提取与所述搜索词相关的字符串,作为核心词候选串。
  54. 根据权利要求52-53任一项所述的装置,其中,所述候选串提取模块,具体用于:
    对所述搜索词进行分词;
    从所述问答页面的页面内容中提取包括至少一个搜索词分词的字符串。
  55. 根据权利要求52-54任一项所述的装置,其中,所述核心词确定模块,具体用于:
    根据所述分类特征对候选串分词进行分类,根据分类结果确定各个候选串分词是否是核心词;
    所述分类特征包括下列特征中的至少一种:名词、热度词表、超链接、相关问题共现率、文档词频。
  56. 根据权利要求52-55任一项所述的装置,其中,所述核心词确定模块,具体用于:
    针对每个分类,将该分类中各个候选串分词与用户输入的搜索词进行匹配,筛选出匹配度最高的设定数量的候选串分词,作为所述核心词;或
    针对每个分类,根据该分类中各个候选串分词的使用频率统计值,筛选出所述使用频率统计值最高的设定数量的候选串分词,作为所述核心词;其中,候选串分词的使用频率统计值包括下列参数之一:被搜索次数、被点击次数、曾作为核心词的次数、曾作为搜索词的次数。
  57. 一种问答页面相关问题推荐装置,包括:
    问题获取模块,用于根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;
    权重确定模块,根据选定时间段内第二用户的浏览行为日志,确定获取的所述相关问题的浏览权重;
    排序推荐模块,用于根据所述浏览权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面中推荐给第一用户的相关问题。
  58. 根据权利要求57所述的装置,其中,所述权重确定模块,具体包括:
    问题分组器,用于根据浏览所述相关问题的浏览用户,对获取的相关问题进行分组;其中,每个相关问题分组中包括一个浏览用户对应的部分或者全部相关问题;
    相关权重计算器,用于在各相关问题分组中,计算组中各相关问题与组中其它相关问题的相关浏览权重;
    浏览权重计算器,用于获取同一相关问题在各相关问题分组中计算得到的相关浏览权重,将获取到的相关浏览权重进行加权,得到获取的各相关问题的浏览权重。
  59. 根据权利要求57-58任一项所述的装置,其中,所述问题分组器,具体用于:
    根据选定时间段内的浏览行为日志,得到每个浏览用户的浏览特征向量{T1、T2、……、Ti、Ti+1、……、Tn},实现对获取的相关问题进行分组;其中,Ti表示一个相关问题。
  60. 根据权利要求57-59任一项所述的装置,其中,所述相关权重计算器,具体用于:
    在各相关问题分组中,根据浏览用户浏览各相关问题的浏览时间对该相关问题分组中的所有相关问题进行排序;
    根据排序结果中,划分浏览时间间隔小于预设的时间间隔阈值的相关问题至同一会话组;
    在每个会话组中,计算组中各相关问题与组中其它相关问题的相关浏览权重。
  61. 根据权利要求57-60任一项所述的装置,其中,所述问题分组器,具体用于:
    得到的浏览特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、用户浏览时间、用户停留时间。
  62. 一种问答页面相关问题推荐装置,包括:
    问题获取模块,用于根据来自第一用户的搜索词,获取数据库中与所述搜索词相关的至少一个相关问题;
    权重确定模块,用于根据选定时间段内第二用户的搜索行为日志,确定获取的所述相关问题的点击权重;
    排序推荐模块,用于根据所述点击权重对获取的相关问题进行排序;根据所述相关问题的排序结果,确定问答页面推荐给第一用户的相关问题。
  63. 根据权利要求62所述的装置,其中,所述权重确定模块,具体包括:
    问题分组器,用于根据所述相关问题对应的查询请求串,对获取的所述相关问题进行分组;其中,每个相关问题分组中包括一个查询请求串对应的部分或全部相关问题;
    相关权重计算器,用于在各相关问题分组中,计算组中各相关问题与组中其他相关问题的相关点击权重;
    点击权重计算器,用于获取同一相关问题在各相关问题分组中计算得到的相关点击权重,将获取到的相关点击权重进行加权,得到获取的各相关问题的点击权重。
  64. 根据权利要求62-63任一项所述的装置,其中,所述问题分组器,具体用于:
    根据所述相关问题对应的查询请求串,得到每个查询请求串的点击特征向量{T1、T2、……、Tn},实现对获取的相关问题进行分组;其中Ti表示一个相关问题。
  65. 根据权利要求62-64任一项所述的装置,其中,所述相关权重计算器,具体用于:
    利用如下公式计算组中各相关问题与组中其他相关问题的相关点击权重W:
    W=P((Ti)|查询请求串)*P((Ti+I)|查询请求串)
    其中,Ti表示一个相关问题;
    Ti+I表示点击特征向量中包括的其他问答问题;
    P((Ti)|查询请求串)表示使用查询请求串时得到Ti的概率;
    P((Ti+I)|查询请求串)表示使用查询请求串时得到Ti+1的概率。
  66. 根据权利要求62-65任一项所述的装置,其中,所述问题分组器,具体用于:
    得到的点击特征向量中的元素Ti的属性包括下列参数中的至少一个:
    问答页面的生成时间、答案数、好评数、差评数、问答长度、展示次数、被点击次数。
  67. 一种计算机程序,包括计算机可读代码,当所述计算机可读代码在计算设备上运行时,导致所述计算设备执行根据权利要求1-34中的任一个所述的方法。
  68. 一种计算机可读介质,其中存储了如权利要求67所述的计算机程序。
PCT/CN2015/095853 2014-12-25 2015-11-27 问答页面相关问题推荐方法及装置 Ceased WO2016101765A1 (zh)

Applications Claiming Priority (8)

Application Number Priority Date Filing Date Title
CN201410827521.4A CN104462552B (zh) 2014-12-25 2014-12-25 问答页面核心词提取方法和装置
CN201410828866.1 2014-12-25
CN201410828977.2A CN104462554B (zh) 2014-12-25 2014-12-25 问答页面相关问题推荐方法和装置
CN201410828977.2 2014-12-25
CN201410830054.0A CN104462556B (zh) 2014-12-25 2014-12-25 问答页面相关问题推荐方法和装置
CN201410828866.1A CN104462553B (zh) 2014-12-25 2014-12-25 问答页面相关问题推荐方法及装置
CN201410830054.0 2014-12-25
CN201410827521.4 2014-12-25

Publications (1)

Publication Number Publication Date
WO2016101765A1 true WO2016101765A1 (zh) 2016-06-30

Family

ID=56149210

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/095853 Ceased WO2016101765A1 (zh) 2014-12-25 2015-11-27 问答页面相关问题推荐方法及装置

Country Status (1)

Country Link
WO (1) WO2016101765A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113505293A (zh) * 2021-06-15 2021-10-15 深圳追一科技有限公司 信息推送方法、装置、电子设备及存储介质
CN115080806A (zh) * 2022-06-16 2022-09-20 北京字跳网络技术有限公司 一种数据搜索方法、装置、设备及介质

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090138443A1 (en) * 2007-11-23 2009-05-28 Institute For Information Industry Method and system for searching for a knowledge owner in a network community
CN103123624A (zh) * 2011-11-18 2013-05-29 阿里巴巴集团控股有限公司 确定中心词的方法及装置、搜索方法及装置
CN103365899A (zh) * 2012-04-01 2013-10-23 腾讯科技(深圳)有限公司 一种问答社区中的问题推荐方法及系统
CN104133817A (zh) * 2013-05-02 2014-11-05 深圳市世纪光速信息技术有限公司 网络社区交互方法、装置及网络社区平台
CN104462552A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面核心词提取方法和装置
CN104462554A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法和装置
CN104462553A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法及装置
CN104462556A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法和装置

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090138443A1 (en) * 2007-11-23 2009-05-28 Institute For Information Industry Method and system for searching for a knowledge owner in a network community
CN103123624A (zh) * 2011-11-18 2013-05-29 阿里巴巴集团控股有限公司 确定中心词的方法及装置、搜索方法及装置
CN103365899A (zh) * 2012-04-01 2013-10-23 腾讯科技(深圳)有限公司 一种问答社区中的问题推荐方法及系统
CN104133817A (zh) * 2013-05-02 2014-11-05 深圳市世纪光速信息技术有限公司 网络社区交互方法、装置及网络社区平台
CN104462552A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面核心词提取方法和装置
CN104462554A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法和装置
CN104462553A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法及装置
CN104462556A (zh) * 2014-12-25 2015-03-25 北京奇虎科技有限公司 问答页面相关问题推荐方法和装置

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113505293A (zh) * 2021-06-15 2021-10-15 深圳追一科技有限公司 信息推送方法、装置、电子设备及存储介质
CN113505293B (zh) * 2021-06-15 2024-03-19 深圳追一科技有限公司 信息推送方法、装置、电子设备及存储介质
CN115080806A (zh) * 2022-06-16 2022-09-20 北京字跳网络技术有限公司 一种数据搜索方法、装置、设备及介质
CN115080806B (zh) * 2022-06-16 2024-12-24 北京字跳网络技术有限公司 一种数据搜索方法、装置、设备及介质

Similar Documents

Publication Publication Date Title
Yao et al. Dynamic word embeddings for evolving semantic discovery
Cao et al. Attsum: Joint learning of focusing and summarization with neural attention
He et al. Trirank: Review-aware explainable recommendation by modeling aspects
CN108280114B (zh) 一种基于深度学习的用户文献阅读兴趣分析方法
CN105224699B (zh) 一种新闻推荐方法及装置
CN103377232B (zh) 标题关键词推荐方法及系统
CN104462553B (zh) 问答页面相关问题推荐方法及装置
US9558267B2 (en) Real-time data mining
CN109189990B (zh) 一种搜索词的生成方法、装置及电子设备
CN111460251A (zh) 数据内容个性化推送冷启动方法、装置、设备和存储介质
CN111444304A (zh) 搜索排序的方法和装置
CN111324771A (zh) 视频标签的确定方法、装置、电子设备及存储介质
CN103714084A (zh) 推荐信息的方法和装置
CN104217030A (zh) 一种根据服务器搜索日志数据进行用户分类的方法和装置
CN107193883B (zh) 一种数据处理方法和系统
Huang et al. Kb-enabled query recommendation for long-tail queries
CN101826102B (zh) 一种图书关键字自动生成的方法
CN104462554A (zh) 问答页面相关问题推荐方法和装置
CN116431908A (zh) 推荐方法、装置、电子设备及存储介质
CN104462552B (zh) 问答页面核心词提取方法和装置
Coelho et al. Transformer-based Language Models for Semantic Search and Mobile Applications Retrieval.
JP4569380B2 (ja) ベクトル生成方法及び装置及びカテゴリ分類方法及び装置及びプログラム及びプログラムを格納したコンピュータ読み取り可能な記録媒体
Song et al. OpenFact: Factuality enhanced open knowledge extraction
CN114282536A (zh) 一种基于ai算法的智能推荐引擎系统
Xu et al. Measuring semantic relatedness between flickr images: from a social tag based view

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15871834

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15871834

Country of ref document: EP

Kind code of ref document: A1