WO2019214236A1 - 原创内容摘要确定和原创内容推荐 - Google Patents
原创内容摘要确定和原创内容推荐 Download PDFInfo
- Publication number
- WO2019214236A1 WO2019214236A1 PCT/CN2018/121321 CN2018121321W WO2019214236A1 WO 2019214236 A1 WO2019214236 A1 WO 2019214236A1 CN 2018121321 W CN2018121321 W CN 2018121321W WO 2019214236 A1 WO2019214236 A1 WO 2019214236A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- sentence
- determining
- content
- original content
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/258—Heading extraction; Automatic titling; Numbering
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/02—Marketing; Price estimation or determination; Fundraising
- G06Q30/0201—Market modelling; Market analysis; Collecting market data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/34—Browsing; Visualisation therefor
- G06F16/345—Summarisation for human users
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/12—Use of codes for handling textual entities
- G06F40/131—Fragmentation of text files, e.g. creating reusable text-blocks; Linking to fragments, e.g. using XInclude; Namespaces
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/12—Use of codes for handling textual entities
- G06F40/137—Hierarchical processing, e.g. outlines
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/211—Syntactic parsing, e.g. based on context-free grammar [CFG] or unification grammars
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/237—Lexical tools
- G06F40/242—Dictionaries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/253—Grammatical analysis; Style critique
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/284—Lexical analysis, e.g. tokenisation or collocates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
- G06F40/295—Named entity recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
Definitions
- the present application relates to an original content summary determination method and apparatus in the field of computer technology, and an original content recommendation method and apparatus.
- a summary is a brief description of an article or paragraph, usually expressing the core meaning of an article or text.
- the method of automatically generating a summary of an article can be regarded as an information compression process, which compresses the input article or text into a short summary, which inevitably has information loss.
- the present application provides an original content summary determination method and apparatus, and a user original content recommendation method and apparatus.
- an embodiment of the present application provides a method for determining an original content summary, including: determining a plurality of sentences arranged before and after a user original content; determining a quality score of each of the sentences; and constraining conditions according to a maximum character length of the abstract And the quality score of each of the sentences, the sentence group having the highest quality score is determined as a summary of the user-originated content, wherein the sentences included in the sentence group are continuous.
- an embodiment of the present application provides an original content summary determining apparatus, including: a sentence determining module, configured to determine a plurality of sentences arranged before and after a user original content; a sentence quality score determining module, configured to determine each a quality score of the sentence; a digest determination module, configured to determine a sentence group having the highest quality score according to a constraint condition of a maximum character length of the digest and a quality score of each of the sentences, as a summary of the original content of the user, where The sentences included in the sentence group are continuous.
- the embodiment of the present application further discloses a method for recommending user-originated content, including: determining a target merchant of a user; determining an original content of the candidate user according to an evaluation score of the user-originated content of the target merchant; determining the candidate a user-originated content that matches the user in the user-originated content; a user-originated content summary determining method according to an embodiment of the present application, determining a summary of the original content with the target user; and recommending the target user to the user A summary of the content.
- the embodiment of the present application further discloses a user-originated content recommendation device, including: a target merchant determining module, configured to determine a target merchant of the user; and a candidate user original content determining module, configured to be used according to the target merchant user An evaluation score of the original content, the candidate user original content is determined; the matching candidate original content determining module is configured to determine the target user original content in the candidate user original content that matches the user; the original content summary determining module is configured to The user original content summary determining method of the embodiment of the present application determines a summary of the target user original content; and a recommendation module is configured to recommend the abstract of the target user original content to the user.
- a target merchant determining module configured to determine a target merchant of the user
- a candidate user original content determining module configured to be used according to the target merchant user An evaluation score of the original content, the candidate user original content is determined
- the matching candidate original content determining module is configured to determine the target user original content in the candidate user original content that matches the user
- an embodiment of the present application further discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and operable on the processor, where the processor implements the computer program
- the user original content summary determining method and the user original content recommending method described in the embodiments of the present application are disclosed.
- the embodiment of the present application provides a computer readable storage medium, where a computer program is stored, and when the program is executed by the processor, the original content summary determining method and the user original content recommending method disclosed in the embodiment of the present application are step.
- the user-originated content summary determining method disclosed in the embodiment of the present application determines the quality scores of each of the sentences by determining the plurality of sentences arranged by the user original content; and finally, according to the constraint condition of the maximum character length of the abstract, A sentence group having the highest quality score is determined as a summary of the user-originated content, wherein sentences included in the sentence group are continuous. This method is capable of extracting user-generated content summaries efficiently and accurately.
- FIG. 1 is a flowchart of a method for determining a user original content summary according to Embodiment 1 of the present application;
- FIG. 2 is a flowchart of a method for determining a user original content summary according to Embodiment 2 of the present application
- FIG. 3 is a flowchart of a user original content recommendation method according to Embodiment 3 of the present application.
- FIG. 4 is a flowchart of a user original content recommendation method according to Embodiment 4 of the present application.
- FIG. 5 is a schematic structural diagram of a user original content summary determining apparatus according to Embodiment 5 of the present application.
- FIG. 6 is a schematic structural diagram of a user original content recommendation apparatus according to Embodiment 6 of the present application.
- FIG. 7 is a second schematic structural diagram of a user original content recommendation apparatus according to Embodiment 6 of the present application.
- This embodiment discloses an original content summary determining method. As shown in FIG. 1 , the method includes: Step 110 to Step 130.
- Step 110 Determine a plurality of sentences arranged in front of the user including the original content of the user.
- the user-originated content is first subjected to data processing, sentences in the user-originated content are extracted, and the extracted sentences are arranged in the order in which the sentences appear in the user-generated content.
- the user-originated content is divided into a plurality of sentences by using a preset punctuation mark as a separation mark between sentences.
- the preset punctuation marks include, but are not limited to, any one or more of the following: a period, an exclamation point, a question mark, a comma, a space, a semicolon, a comma, an ellipsis, an emoji, and a wavy symbol.
- Standard punctuation marks include at least a period, an exclamation point, a question mark, a comma, a semicolon, a comma, a colon, and an ellipsis.
- the user's original content is firstly sentenced by standard punctuation, and if the sentence is still too long, the other symbols are used again.
- the sentences are arranged in the order of the positions in which the sentences appear in the user-originated content, and the M sentences arranged before and after the user-originated content are obtained. Where M is a natural number greater than or equal to 1.
- step 120 a quality score for each of the sentences is determined.
- the quality score of the sentence may be determined from features of the information dimension such as text, viewpoints, and entities included in the sentence.
- the text may further include: information of dimensions such as location, length, keyword emotional attribute, and description of the keyword to the merchant feature.
- the information of the viewpoint dimension may be information such as an evaluation object and an evaluation word included in the viewpoint.
- the information of the entity dimension may be information such as the frequency of occurrence of the entity word, the type of the entity word, and the like.
- the quality of a sentence is used to indicate the contribution or performance of the sentence to the core idea of the user's original content.
- Step 130 Determine, according to a constraint condition of a maximum character length of the digest and a quality score of each of the sentences, a sentence group with the highest quality score as a summary of the original content of the user, wherein the sentence included in the sentence group is continuous of.
- the sentence group having the highest information content is selected as the abstract of the user original content.
- a plurality of sets of sentences are included that contain a character length that satisfies a preset character length condition.
- the score of the sentence group is determined according to the quality score of each sentence in each group of sentences. Finally, select the sentence group with the highest score as a summary of the user's original content.
- the user original content summary determining method disclosed in the embodiment of the present application determines the quality score of each sentence by determining at least one sentence arranged before and after the user original content; the constraint according to the maximum character length of the digest and each of the described
- the quality score of the sentence, the sentence group with the highest quality score is determined, and as a summary of the user-originated content, the abstract of the user-originated content can be extracted efficiently and accurately.
- An original content summary determining method disclosed in this embodiment is shown in FIG. 2, and the method includes: Step 210 to Step 240.
- Step 210 Construct an evaluation object library, an evaluation vocabulary, and an entity vocabulary.
- the evaluation object library, the evaluation vocabulary and the entity vocabulary are constructed, and then the sentence is determined based on the evaluation object library, the evaluation vocabulary and the entity vocabulary. Entities included, evaluation objects, and keywords of emotion classes included in sentences.
- a lexical analyzer is used to obtain keywords such as nouns and adjectives, and the contents of the preset POI knowledge base are combined.
- the N-Gram technology is used to obtain the keyword in the UGC comment and the part of speech of the query keyword (for example, an attraction, a movie theater, a business district, a shopping mall, etc.). Then, through the evaluation object mining, a library of evaluation objects with high coverage rate can be built to provide support for subsequent comment mining.
- An entity is a subset of the evaluation objects, and is selected from keywords in structured data such as merchants, users, etc., such as a business name, a dish category, a dish name, and the like.
- Keywords refer to meaningful words after UGC text has been segmented.
- Evaluation words refer to keywords such as adjectives, adverbs and idioms.
- the evaluation words of the high frequency in the UGC comment are obtained, and the distribution of the evaluation words in the 5-star comment and the 1-star comment is counted, and the polarity (positive, negative, and neutral) of the evaluation word is obtained. For example, the number of "very good” evaluation words appears in the praise comments is much larger than the number in the bad reviews, then the polarity of the "very good” evaluation words is positive.
- an evaluation vocabulary can be built to support the subsequent comment mining. The emotional information of the sentence can be determined by the evaluation word.
- Step 220 Determine a plurality of sentences arranged in front of the user including the original content of the user.
- the user-originated content is first subjected to data processing, sentences in the user-originated content are extracted, and the extracted sentences are arranged in the order in which the sentences appear in the user-generated content.
- the user-originated content is divided into a plurality of sentences according to a preset punctuation mark as a separation mark between sentences.
- the preset punctuation marks include, but are not limited to, any one or more of the following: a period, an exclamation point, a question mark, a comma, a space, a semicolon, a comma, a colon, an ellipsis, an emoji, and a wavy symbol.
- Standard punctuation marks include at least a period, an exclamation point, a question mark, a comma, a semicolon, a comma, a colon, and an ellipsis.
- the user's original content is firstly sentenced by standard punctuation, and if the sentence is still too long, the other symbols are used again.
- the sentences are arranged in the order of the positions in which the sentences appear in the user-originated content, and the M sentences arranged before and after the user-originated content are obtained. Where M is a natural number greater than or equal to 1.
- determining the at least one sentence of the user-originated content includes: pre-scoring the user-originated content based on the standard punctuation, obtaining the first sentence included in the user-originated content; and based on the extended punctuation
- the first sentence in the first sentence in which the character length is greater than the preset sentence character length threshold is re-sentenced to obtain the second sentence corresponding to the first sentence; the character length in the first sentence is not re-commented
- the first sentence and the second sentence are arranged in the order of the positions appearing in the user original content, and the M sentences arranged before and after the user original content are obtained.
- M is a natural number greater than or equal to 1.
- the standard punctuation marks include at least: a period, a comma, a question mark, an exclamation point, an ellipsis, a colon, a comma, and a semicolon
- the extended punctuation includes: a space, an emoji, a broken wave symbol, and the like.
- the sentence includes an emoji " ⁇ _ ⁇ "
- the sentence is divided based on the extended punctuation, and two second sentences are obtained, namely: "with the pollution-free Longli fish from Vietnam” and "The taste is fresh and tender.”
- the four sentences included in the original content of the user are: the first sentence “Original Bayan aged sauerkraut”, “three years of fermentation”, and the second sentence "cooperating with the pollution-free Longli fish from Vietnam” And "the taste is fresh and incomparable”.
- a quality score for each of the sentences is determined.
- determining a quality score of each of the sentences includes: determining a quality score of the sentence according to information of a preset dimension of the sentence, wherein the preset dimension includes one of the following dimensions Or multiple: text, entities, and opinions. Determining the quality score of the sentence according to the information of the preset dimension of the sentence, comprising: weighting the entity dimension score and the viewpoint dimension score of the sentence to obtain an initial quality score; and scoring the text dimension of the sentence The initial quality score is adjusted; and the adjusted initial quality score is determined as the quality score of the sentence.
- score(sentence i ) represents the quality score of sentence i
- score_sentence i word ⁇ entity
- score_sentence i word ⁇ evaluation object
- w′ represents sentence i Text dimension score.
- the evaluation object is an evaluation object included in the viewpoint included in the sentence i, ⁇ represents a first weight adjustment factor corresponding to the entity dimension score, and ⁇ represents a second weight adjustment factor corresponding to the viewpoint dimension score. That is, first, the initial value quality score is calculated by the following formula:
- the initial value quality score is adjusted by the text dimension score w' to obtain the quality score of the sentence i.
- determining the textual dimension score of the sentence according to the position of the sentence in the user's original content, the negative emotion information of the sentence, and the merchant characteristic information includes: improving the quality score of the sentence near the head of the user original content, and reducing the content
- the quality score of the sentence of negative sentiment information includes the quality score of the sentence including the characteristic information of the merchant. For example, for the first three sentences that appear in the user's original content, the quality scores of the first three sentences are increased, for example, by adding 10 points, thereby increasing the probability that the user's original content head position sentence appears in the digest.
- the sentence includes a negative word in the preset evaluation vocabulary
- it is determined that the sentence contains a negative emotion and by reducing the quality score of the sentence, for example, by 20 points, the probability that the sentence appears in the abstract is reduced.
- the probability of the sentence appearing in the abstract is reduced by lowering the quality score of the sentence, for example, by 10 points.
- the quality score of the sentence is increased, for example, by adding 10 points, thereby increasing the probability that the sentence appears in the abstract.
- the entity dimension score reflects the weight of the entity in the user-generated content.
- the entity dimension score of the sentence is determined based on the inverse text word frequency of the entity word included in the sentence.
- the entity dimension score is the sum of the inverse text word frequencies of the entities included in the sentence, and the entity dimension score of the sentence is determined by the following formula:
- idf(word j ) is the inverse text word frequency of the entity word word j included in the sentence.
- the inverse text word frequency of the entity can be determined by the following formula:
- the viewpoint dimension score of the sentence is determined based on the reverse text word frequency of the evaluation object involved in the viewpoint included in the sentence.
- the perspective dimension score reflects the weight of the evaluation object in the user's original content.
- the viewpoint dimension score of the sentence is determined based on the inverse text word frequency of the evaluation object words included in the sentence.
- the information of the viewpoint dimension is the sum of the reverse text word frequencies of the evaluation objects involved in the viewpoint included in the sentence, and the viewpoint dimension score of the sentence is determined by the following formula:
- idf(word l ) is the inverse text word frequency of the evaluation object word l included in the sentence.
- the reverse text word frequency of the evaluation object is determined by the following formula:
- the viewpoint dimension score of the sentence is determined based on the reverse text word frequency of the evaluation object involved in the viewpoint included in the sentence. For example, using a formula to determine the perspective dimension of a sentence,
- idf(word l ) is the inverse text word frequency of the evaluation object word l included in the sentence.
- the weight score of the sentence is obtained by weighting the entity dimension score and the viewpoint dimension score.
- the weights of the entity dimension score and the viewpoint dimension score are set by empirical statistics.
- Step 240 Determine, according to a constraint condition of a maximum character length of the digest and a quality score of each of the sentences, a sentence group having the highest quality score as a summary of the original content of the user, wherein the sentence included in the sentence group is continuous of.
- the sentence group having the highest information content is selected as the abstract of the user original content.
- the sentence group between begin and end is determined by the following formula as a summary of the user-generated content:
- begin and end are the sequence numbers of the sentences in the original content of the user
- max_length is the maximum character length of the preset digest
- length(sentence i ) is the length of the characters in the sentence i
- w is the total score adjustment factor
- w is based on the sentence sentence i
- begin ⁇ i ⁇ end contains entities and opinions, and determine.
- Determining the sentence group with the highest quality score according to the constraint condition of the maximum character length of the abstract and the quality score of each of the sentences, as a summary of the original content of the user including: determining, by using a sliding window technology, a constraint condition that satisfies the maximum character length of the digest At least one set of sentences; for each of the sentence groups, determining a weighted sum of quality scores of respective sentences included in the sentence group as a quality score of the sentence group; determining a sentence group having the highest quality score as a A summary of user-generated content.
- the weight of the quality score of each sentence in the quality score of the sentence group is based on whether each sentence in the sentence group contains an entity and a viewpoint, a character length of the sentence group, and whether the sentence group includes Any one or more of the first sentence or the last sentence of the user-originated content is determined.
- the maximum character length of the preset digest is 35, and the nine original sentences are arranged in a user original content, and the quality score and the character length of each sentence are as shown in the following table.
- the sentence numbers 1 to 9 are the sequence numbers of the sentences before and after the sentence, and the weights of the quality scores of the respective sentences are the same, for example, both are 1.
- a sentence group having a length of no more than 35 characters such as ⁇ sentence 1 ⁇ , ⁇ sentence 1, sentence 2 ⁇ , ⁇ sentence 1, sentence 2, is found. , sentence 3 ⁇ , ⁇ sentence 1, sentence 2, sentence 3, sentence 4 ⁇ .
- the quality scores of each group of sentences are respectively determined, and a group of sentences with the highest quality score is retained, such as a sentence group composed of ⁇ sent 1, sentence 2, sentence 3, sentence 4 ⁇ as a candidate summary, the candidate summary
- the quality is divided into 3.7 points.
- the quality score of the candidate summary consisting of ⁇ sentence 1, sentence 2, sentence 3, and sentence 4 ⁇ is greater than the quality score of the sentence group consisting of ⁇ sentence 2, sentence 3, and sentence 4 ⁇ (3.2 points), therefore, temporarily retain ⁇ sentence 1 , sentence 2, sentence 3, sentence 4 ⁇ candidate summary of sentence group composition.
- the sentence group with the highest quality score is determined according to the constraint condition of the maximum character length of the digest and the quality score of each of the sentences, as a summary of the original content of the user, including: determining the satisfaction summary by sliding window technology
- One or more sentence groups of the constraint of the maximum character length for each of the sentence groups, the weighted sum of the quality scores of the respective sentences included in the sentence group is determined as the quality score of the sentence group;
- a sentence group with the highest quality score as a summary of the user-generated content.
- the quality scores of the individual sentences in the sentence group may have the same weight or may have different weights.
- the quality scores of the sentences in the sentence group have different weights, and if the entity dimension score of the sentence is zero, for example, the entity is not included in the sentence, the weight of the quality score of the sentence is decreased; If the viewpoint dimension score of the sentence is zero, for example, the evaluation object is not included in the sentence, the weight of the quality score of the sentence is lowered; if the sentence is the first sentence or the last sentence of the user original content, then the promotion is performed.
- the weight of the quality of the sentence The integrity of the sentence in the determined digest can be improved by determining whether the sentence is the weight of the quality score of the first sentence or the last sentence of the user original content.
- the user original content summary determining method disclosed in the embodiment of the present application determines the quality scores of each of the sentences by determining the plurality of sentences arranged by the user original content; and finally, the root according to the constraint of the maximum character length of the abstract And the quality score of each of the sentences, the sentence group having the highest quality score is determined, and as the abstract of the user original content, the abstract question of the user original content can be extracted efficiently and accurately.
- the quality score of the sentence is obtained by weighting the three dimensions of the text, the entity and the viewpoint of the user original content, and by this method, the sentence group with the highest information value density in the user original content can be found.
- the user original content summary determining method disclosed in the embodiment of the present application supports the extraction of punctuation marks that are not standardized, and even the extraction of the user original content summary that is not fluent, and the robustness is stronger; A summary of user-generated content adapted to the characteristics of the merchant.
- An original content recommendation method disclosed in this embodiment is shown in FIG. 3, and the method includes: Step 310 to Step 350.
- step 310 the target merchant of the user is determined.
- the merchant that has performed the preset historical behavior by the user is determined as the first target merchant; and then, the merchant that is similar to the first target merchant is determined as the second target merchant. Finally, the first target merchant and the second target merchant are used as target merchants of the user.
- Step 320 Determine candidate user original content according to the evaluation score of the user original content of the target merchant.
- the evaluation score of the user's original content may be determined according to text information, entity information, viewpoint information, and the like of the user-originated content. In an embodiment, the higher the evaluation score indicates that the quality of the user-generated content is higher, that is, the information that the user-generated content presents to the user is more valuable. Then, the user-originated content of each of the target users is separately sorted in order of the evaluation score of the user-originated content from high to low. Thereafter, for each target user, a preset number of user-originated content having the highest evaluation score is selected as the candidate user original content.
- Step 330 Determine target user original content that matches the user in the candidate user original content.
- the feature vector of the user and the feature vector of each candidate user original content may be separately extracted, and then determined by calculating the similarity between the feature vector of the user and the feature vector of each candidate user original content.
- the target user original content in the candidate user original content that matches the user.
- the degree of matching between the user and the original content of a certain candidate user may be determined by calculating a similarity distance between the feature vector of the user and the feature vector of the candidate user original content; or, by pre-training
- the machine learning sorting model calculates a matching degree between the user and the original content of the certain candidate user according to the input feature vector of the user and a feature vector of the candidate user original content.
- one or a preset number of the candidate user original content having the highest degree of matching with the user is selected as the target user original content.
- Step 340 determining a summary of the original content of the target user.
- the user original content summary determining method determines a summary of the original content of the target user.
- Step 350 recommending a summary of the target user original content to the user.
- the user-originated content recommendation method disclosed in the embodiment of the present application determines the target user's original merchant according to the evaluation score of the user-originated content of the target merchant; and determines the candidate user's original content and the user. a matching target user original content; finally, recommending a summary of the target user original content to the user, wherein the target user original content is summarized according to the user original content summary determining method according to the first embodiment or the second embodiment It is determined that, in this way, it is possible to recommend more accurate user-originated content according to user requirements than a scheme of recommending user-originated content to the user based on the heat of the user's original content.
- the user-originated content recommendation method disclosed in the embodiment of the present application by recommending the user-originated content matched with the user to the user, achieves targeted information recommendation, and effectively improves the accuracy of the user-originated content recommendation.
- the original content for the user only the summary of the original content is displayed, and the key information recommended by the user is displayed in a clear and clear manner, so that the user can make an accurate and rapid decision, thereby further improving the user experience.
- a user-originated content recommendation method disclosed in this embodiment is shown in FIG. 4, and the method includes: Step 410 to Step 470.
- Step 410 Construct an evaluation object library, an evaluation vocabulary, and an entity vocabulary.
- the evaluation vocabulary and the entity vocabulary refer to the second embodiment, which is not described in this embodiment.
- step 420 the target merchant of the user is determined.
- determining a target merchant of the user includes: determining a merchant that the user has generated a preset behavior as the first target merchant; and determining a similarity to the first target merchant based on the similarity of the merchant vector Two target merchants; the first target merchant and the second target merchant are used as target merchants of the user.
- the merchant that the user has performed the preset historical behavior is determined as the first target merchant.
- the merchants that have generated the preset behavior by the user include, but are not limited to, the merchant that the user clicked, the merchant that the user browsed, the merchant that the user has collected, and the merchant that the user purchased the commodity.
- the merchant similar to the first target merchant is further determined as the second target merchant.
- the method before determining the second target merchant similar to the first target merchant based on the similarity of the merchant vector, the method further includes: inputting the merchant sequence clicked by the user as a word vector model, and training the merchant vector model; The merchant vector model determines the merchant's merchant vector.
- the user's behavior on the merchant is transformed into a time series event, and then the time series event is taken as input, and the deep learning algorithm is used to train the merchant vector model, that is, the merchant feature is mapped from the high dimensional discrete space to the low The continuous space of the dimension.
- the merchant vector model that is, the merchant feature is mapped from the high dimensional discrete space to the low The continuous space of the dimension.
- a second target merchant similar to the first target merchant may be determined.
- the first target merchant and the second target merchant are used as target merchants of the user. For example, according to the historical behavior of the user, it is determined that the user has clicked on the merchant 1, and the merchant 1 is the first target merchant of the user. Then, by calculating the similarity of the merchant vector, the merchant 2 similar to the merchant 1 is determined, and the merchant 2 is taken as the second target merchant of the user. Finally, merchant 1 and merchant 2 are used as target merchants of the user.
- Step 430 Determine, according to information of three dimensions of text, entity, and viewpoint of the user-originated content of the target merchant, an evaluation score of the user-originated content.
- the method further includes: determining the user original content according to the information of the three dimensions of the text, the entity, and the viewpoint of the user-originated content of the target merchant. Evaluation score. For example, determining the evaluation score of the user-originated content according to the three dimensions of the text, the entity, and the viewpoint of the user-originated content of the target merchant may include: a text score, an entity score, and a viewpoint score by using the original content of the user. A weighted summation is performed, the evaluation score of the user-originated content.
- the user-originated content of the platform such as a user review
- the user-originated content of the most recent preset time e.g, within six months
- the evaluation score of the user-originated content is determined. Because high-quality merchants or high-star users also have low-quality user-originated content, when rating user-originated content, regardless of the characteristics of the merchant and user, only the content quality of the user-originated content itself is analyzed, through the text.
- the evaluation dimensions of the user-generated content are calculated from the three dimensions of entity, viewpoint and viewpoint.
- the text score is proportional to the number of different words contained in the user's original content. That is, the more different texts included in the user-generated content, the higher the text score. Determining the text score based on the number of different texts contained in the user's original content can effectively filter out user-originated content in which the user reuses the same punctuation or text to serve as the number of words.
- the entity score may be represented by the inverse text word frequency of the entity contained in the user-originated content; the opinion score may be represented by the reverse text word frequency of the evaluation object involved in the viewpoint contained in the user-originated content.
- the user-generated content is divided into a plurality of sentences.
- a specific method for dividing the user-originated content into a plurality of sentences refer to the method for determining a sentence in the user-originated content in the second embodiment, which is not described in this embodiment.
- Entity refers to the comment objects involved in the user's original content, such as business name, address, category, shopping mall, star hotel, shopping mall, community, cinema, administrative district and city.
- An entity is important information in user-generated content. For example, information about recommended dishes, addresses, and categories mentioned in a user-generated content can be an important feature of the original content of the user.
- the information extraction in the O2O (online to offline, O2O) scenario is different from the traditional name, place name and company name identification. It is necessary to mine the weight information of different keywords in different dimensions, for example, in the business reviews under the US food category, “ The Dream of the Dragon has very few comments, and its reverse text word frequency is higher than that of “Cantonese cuisine”.
- the entity score of a user-generated content can be determined by the following formula:
- idf(word p ) is the inverse text word frequency of the entity word p included in the original content of the user.
- the reverse text word frequency of the entity word is determined by the following formula.
- the viewpoint indicates the subjective and objective judgment information for the specific evaluation target, and in the present application, the viewpoint is mainly extracted from the sentence.
- the specific method of extracting opinions from the sentence is as follows:
- the evaluation object included in the sentence can be determined. Yes: coffee beans; according to the pre-built evaluation vocabulary, it can be determined that the evaluation words included in the sentence are: “concentrated” and “classic”; the evaluation object included in the sentence is combined with the evaluation word to obtain the sentence included in the sentence.
- the views are: "coffee beans - classic" and "coffee beans - concentrated”.
- the confidence of each viewpoint is obtained. In an embodiment, the more frequently the viewpoint appears, the higher the confidence. Finally, all the opinions in the original content of the user are obtained, as well as the confidence of each viewpoint.
- a vector representation of the viewpoint is obtained by summing the evaluation object and the word vector of the evaluation word included in the viewpoint. After the viewpoint is represented by the vector, the cosine theorem can be used to calculate the distance between the vectors to determine the similarity between the viewpoints.
- the following view data structure table can be obtained:
- the training samples are obtained by word segmentation, and the word vector of each keyword in the training sample is obtained using a word vector technique well known to those skilled in the art.
- the keywords include entity words, evaluation words, and various meaningful general terms.
- a word vector is a vector representation of a keyword.
- the word vector of the keyword is a fixed-length floating-point one-dimensional vector, for example, a word vector model is trained using a negative sampling method of the skip-gram model.
- all keywords can be represented by a fixed-length vector, compressing the original sparsely large dimension into a smaller dimensional space. For example, the words "pizza" and "pizza" are not similar in text. Sex, but after the word vector is represented, its semantic distance is relatively close.
- the viewpoint scores of the viewpoints, and the text scores are weighted, and the weighted values of the scores are set according to specific needs. Generally, the viewpoint score has the highest weight and the text score has the lowest weight.
- Step 440 Determine candidate user original content according to the evaluation score of the user original content of the target merchant.
- the evaluation scores are respectively selected in the user-originated content of the merchant 1 and the merchant 2 according to the evaluation score of the user-originated content.
- a plurality of user-originated content is used as a candidate user-originated content of the user. For example, according to the order of the evaluation scores from high to low, the user original content of the merchant 1 and the merchant 2 are respectively sorted, and then the M original content of the merchant 1 and the M original content of the merchant 2 with the highest evaluation score are selected. , as a candidate user original content.
- Step 450 Determine target user original content that matches the user in the candidate user original content.
- determining the target user original content that matches the user in the candidate user original content includes: determining, according to each of the candidate user original content ranking features and the user user characteristics, And matching the candidate user original content with the user; determining the candidate user original content whose matching degree satisfies a preset condition as the target user original content that matches the user.
- the matching model may be trained by machine learning based on the ranking features of the user-originated content and the user characteristics of the user. For example, combining the sorting feature of the user-originated content and the user feature of the user who published the original content into a positive sample, combining the sorting feature of the user-originated content and the user feature of the user who stepped on the original content into a negative sample, and training matching degree Identify the model. Then, the matching degree recognition model identifies the user-originated content and the matching degree of the user based on the sorted feature of the input user-originated content and the user feature of the user.
- the sorting feature includes: a number of likes, a number of comments, a share number, a text quality score, a picture quality score, an entity word, a user-originated content publisher level, a relationship between the publisher and the user, or
- the user feature includes: any one or more of a user historical behavior feature, a business district preference feature, a category preference feature, and a similar user feature, and the user historical behavior feature includes: searching, browsing, purchasing, A feature of any one or more of the behaviors of the store.
- the preset number of the candidate user original content with the highest matching score may be determined as the target user original content that matches the user; or the candidate user corresponding to each merchant is determined.
- the candidate user original content having the highest matching score with the user is the target user original content that matches the user. Since the matching degree is combined, the user preference, the user social relationship and the like are combined, and thus the determined target user original content is the user-originated content preferred by the user.
- Step 460 determining a summary of the original content of the target user.
- the user original content summary determination method described in the first embodiment and the second embodiment is used to determine the abstract of the target user's original content. In this embodiment, the specific determination method of the abstract is not described again.
- Step 470 recommending a summary of the target user original content to the user.
- the user original content recommendation method disclosed in the embodiment of the present application determines the target merchant of the user by determining the target merchant of the target merchant, and determines the candidate according to the evaluation score of the user original content of the target merchant.
- the user-originated content recommendation method disclosed in the embodiment of the present application by recommending the user-originated content matched with the user to the user, achieves targeted information recommendation, and effectively improves the accuracy of the user-originated content recommendation.
- by displaying the user-originated content for the user only the summary of the user's original content is displayed, and the key information recommended by the user is displayed in a clear and clear manner, so that the user can make an accurate and rapid decision, thereby further improving the user experience.
- the accuracy of the user's original content quality evaluation can be improved, and the accuracy of the user's original content recommendation can be further improved.
- An original content summary determining apparatus disclosed in this embodiment as shown in FIG. 5, the apparatus includes:
- the sentence determination module 510 is configured to determine at least one sentence arranged before and after the user original content is included.
- the sentence quality score determining module 520 is configured to determine the quality score of each of the sentences.
- a summary determining module 530 configured to determine, according to a constraint condition of a maximum character length of the digest and a quality score of each of the sentences, a sentence group with the highest quality score as a summary of the user-originated content, where the sentence group includes The sentences are continuous.
- the sentence quality score determining module 520 is further configured to:
- a quality score for each of the sentences is determined according to information of a preset dimension of each of the sentences, wherein the preset dimensions include one or more of the following dimensions: text, entity, and viewpoint.
- the determining, according to the information of the preset dimension of each of the sentences, the quality score of each of the sentences comprising: weighting and summing the entity dimension score and the viewpoint dimension score of each of the sentences An initial quality score, and the initial quality score is adjusted by a textual dimension score of the sentence; and the adjusted initial quality score is determined as the quality score of the sentence.
- the entity dimension score and the viewpoint dimension score of each of the sentences are weighted and summed to obtain an initial quality score, and the initial quality score is performed by a text dimension score of the sentence.
- Adjusting; and determining the adjusted initial mass score as the quality score of the sentence further comprising:
- the quality score for each of the sentences is determined according to the following formula:
- Score(sentence i ) w' ⁇ ( ⁇ score_sentence i (word ⁇ entity)+ ⁇ score_sentence i (word ⁇ evaluation object))
- score(sentence i ) represents the quality score of sentence i
- score_sentence i word ⁇ entity
- score_sentence i word ⁇ evaluation object
- viewpoint dimension score of sentence i
- w′ represents sentence i Text dimension score.
- the evaluation object is an evaluation object for which the viewpoint included in the sentence is directed
- ⁇ represents a first weight adjustment factor corresponding to the entity dimension score
- ⁇ represents a second weight adjustment factor corresponding to the viewpoint dimension score.
- the digest determining module 530 is further configured to:
- the weighted sum of the quality scores of the respective sentences included in the sentence group is determined as the quality score of the sentence group
- a sentence group having the highest quality score is determined as a summary of the user-originated content.
- the weight of the quality score of each sentence in the quality score of the sentence group is based on whether each sentence in the sentence group contains an entity and a viewpoint, a character length of the sentence group, and whether the sentence group includes the Any one or more of the first sentence or the last sentence of the user-generated content is determined.
- This embodiment is an embodiment of the device corresponding to the first embodiment and the second embodiment.
- the modules in this embodiment refer to the description of related steps in the first embodiment and the second embodiment, and details are not described herein again.
- Determining the quality scores of each of the sentences by determining a plurality of sentences arranged before and after the user-originated content, and then determining the sentence with the highest quality score according to the constraint of the maximum character length of the digest and the quality score of each of the sentences A group, as a summary of the user-generated content, wherein the sentences included in the sentence group are consecutive.
- the user-originated content recommendation device in the example of the present disclosure solves the problem that the original content summary cannot be accurately extracted. After a large amount of user-originated content testing, the user-originated content summary determining apparatus disclosed in the present application can efficiently and accurately determine the abstract of the user-originated content.
- the method for calculating the sentence quality by weighting the three dimensions of the text, the entity and the viewpoint of the user-originated content can find the sentence group with the highest information value density in the user-originated content.
- the method for determining the remote content summary disclosed in the embodiment of the present application supports the extraction of the punctuation of the punctuation of the user, and even the extraction of the user-originated content summary of the statement is more robust; the robustness is stronger according to the different requirements for the length of the digest; A summary of user-generated content adapted to the characteristics of the merchant.
- An original content recommendation device disclosed in this embodiment includes:
- the target merchant determination module 610 is configured to determine a target merchant of the user.
- the candidate user original content determining module 620 is configured to determine the candidate user original content according to the evaluation score of the user original content of the target merchant.
- the matching candidate original content determining module 630 is configured to determine target user original content that matches the user in the candidate user original content.
- the original content summary determining module 640 is configured to determine a summary of the original content of the target user according to the user original content summary determining method according to the embodiment of the present application.
- the recommendation module 650 is configured to recommend a summary of the target user original content to the user, where the summary of the target user original content is determined according to the user original content summary determining method according to the first embodiment and the second embodiment.
- the apparatus further includes:
- the user-originated content evaluation score determining module 660 is configured to determine an evaluation score of the user-originated content according to information of three dimensions of text, entity, and viewpoint of the user-originated content.
- the target merchant determining module 610 is further configured to:
- the target merchant determining module 610 is further configured to:
- the merchant sequence clicked by the user is used as an input of the word vector model, and the merchant vector model is trained; the merchant vector of the first target merchant is determined by the merchant vector model.
- the matching candidate original content determining module 630 is further configured to:
- the candidate user original content is the target user original content that matches the user.
- the sorting feature includes: a number of likes, a number of comments, a share number, a text quality score, a picture quality score, an entity word, a user-originated content publisher level, a relationship between the publisher and the user, or
- the user feature includes: any one or more of a user historical behavior feature, a business district preference feature, a category preference feature, and a similar user feature, and the user historical behavior feature includes: searching, browsing, purchasing, A feature of any one or more of the behaviors of the store.
- This embodiment is an embodiment of the device corresponding to the third embodiment and the fourth embodiment.
- the modules in this embodiment refer to the descriptions of related steps in the third embodiment and the fourth embodiment, and details are not described herein again.
- the user-originated content recommendation device in the example of the present disclosure solves the problem that the recommended user-originated content is inaccurate and cannot satisfy the user's demand when the user-originated content is recommended for the user according to the heat of the user-originated content.
- the user-originated content recommendation device in the example of the present disclosure effectively improves the accuracy of the user-originated content recommendation by recommending the user-originated content that matches the user to the user.
- the user-originated content recommendation device by recommending the original content for the user, only the summary of the original content is displayed, and the key information recommended by the user is displayed in a clear and clear manner, so that the user can make an accurate and rapid decision, thereby further improving the user experience.
- the accuracy of the user's original content quality evaluation can be improved, and the accuracy of the user's original content recommendation can be further improved.
- the present application also discloses an electronic device including a memory, a processor, and a computer program stored on the memory and operable on the processor, the processor executing the computer program to implement the present application
- the original content summary determining method, the third embodiment and the fourth embodiment described in the first embodiment and the second embodiment are used for the original content recommendation method.
- the electronic device can be a PC, a mobile terminal, a personal digital assistant, a tablet, or the like.
- the present application also discloses a computer readable storage medium, on which a computer program is stored, and the program is executed by the processor to implement the original content summary determination method or the third embodiment as described in the first embodiment and the second embodiment of the present application. And the user-originated content recommendation method described in the fourth embodiment.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Artificial Intelligence (AREA)
- General Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Health & Medical Sciences (AREA)
- Business, Economics & Management (AREA)
- Strategic Management (AREA)
- Accounting & Taxation (AREA)
- Development Economics (AREA)
- Finance (AREA)
- Entrepreneurship & Innovation (AREA)
- Data Mining & Analysis (AREA)
- Software Systems (AREA)
- Game Theory and Decision Science (AREA)
- Economics (AREA)
- Marketing (AREA)
- General Business, Economics & Management (AREA)
- Medical Informatics (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Databases & Information Systems (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Machine Translation (AREA)
Abstract
一种原创内容摘要确定方法,所述方法包括:确定用户原创内容包括的前后排列的多个句子(110);然后,确定每个所述句子的质量分(120);最后,根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要(130),其中,所述句组中包括的句子是连续的。
Description
相关申请的交叉引用
本专利申请要求于2018年5月11日提交的、申请号为201810447372.7、发明名称为“原创内容摘要确定方法及装置,原创内容推荐方法及装置”的中国专利申请的优先权,该申请的全文以引用的方式并入本文中。
本申请涉及计算机技术领域的原创内容摘要确定方法及装置,原创内容推荐方法及装置。
摘要是一篇文章或一段文字的简要描述,通常表达了文章或文字的核心含义。可以把文章自动生成摘要的方法看作是一个信息压缩过程,将输入的文章或文字压缩为一篇简短的摘要,该过程不可避免有信息损失。
发明内容
本申请提供一种原创内容摘要确定方法和装置以及用户原创内容推荐方法和装置。
第一方面,本申请实施例提供了一种原创内容摘要确定方法包括:确定用户原创内容包括的前后排列的多个句子;确定每个所述句子的质量分;根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
第二方面,本申请实施例提供了一种原创内容摘要确定装置,包括:句子确定模块,用于确定用户原创内容包括的前后排列的多个句子;句子质量分确定模块,用于确定每个所述句子的质量分;摘要确定模块,用于根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
第三方面,本申请实施例还公开了一种用户原创内容推荐方法,包括:确定用户的目标商户;根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;确 定所述候选用户原创内容中与所述用户匹配的目标用户原创内容;根据本申请实施例所述用户原创内容摘要确定方法,确定与所述目标用户原创内容的摘要;向所述用户推荐所述目标用户原创内容的摘要。
第四方面,本申请实施例还公开了一种用户原创内容推荐装置,包括:目标商户确定模块,用于确定用户的目标商户;候选用户原创内容确定模块,用于根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;匹配候选用户原创内容确定模块,用于确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容;原创内容摘要确定模块,用于根据本申请实施例所述用户原创内容摘要确定方法,确定所述目标用户原创内容的摘要;推荐模块,用于向所述用户推荐所述目标用户原创内容的摘要。
第五方面,本申请实施例还公开了一种电子设备,包括存储器、处理器及存储在所述存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现本申请实施例所述的用户原创内容摘要确定方法和用户原创内容推荐方法。
第六方面,本申请实施例提供了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时本申请实施例公开的原创内容摘要确定方法和用户原创内容推荐方法的步骤。
本申请实施例公开的用户原创内容摘要确定方法,通过确定用户原创内容包括的前后排列的多个句子;然后,确定每个所述句子的质量分;最后,根据摘要最大字符长度的约束条件,确定所述质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。该方法能够高效并准确地提取用户原创内容摘要。
为了更清楚地说明本申请实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1是本申请实施例一的用户原创内容摘要确定方法流程图;
图2是本申请实施例二的用户原创内容摘要确定方法流程图;
图3是本申请实施例三的用户原创内容推荐方法流程图;
图4是本申请实施例四的用户原创内容推荐方法流程图;
图5是本申请实施例五的用户原创内容摘要确定装置的结构示意图之一;
图6是本申请实施例六的用户原创内容推荐装置的结构示意图之一;
图7是本申请实施例六的用户原创内容推荐装置的结构示意图之二。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
在确定摘要的过程中,为了保留尽可能多的重要信息,常用的方法包括信息抽取、文章分类和词法分析等,然后根据获取的信息生成摘要。与传统文章相比,用户原创内容UGC(User created Content)具有篇幅更短、段落不明显、句子结构不规范、用词相对随意的特点,传统的提取文章或文字摘要的方法可能无法准确提取出用户原创内容的摘要。
实施例一
本实施例公开一种原创内容摘要确定方法,如图1所示,该方法包括:步骤110至步骤130。
步骤110,确定用户原创内容包括的前后排列的多个句子。
在一实施例中,首先对用户原创内容进行数据处理,提取出所述用户原创内容中的句子,并将提取出的句子按照各句子在所述用户原创内容中出现的先后顺序进行排列。
由于用户原创内容,如用户点评,没有固定格式要求,所以内容和格式多样。在一实施例中,以预设标点符号作为句子之间的分隔标记,将所述用户原创内容划分为多个句子。其中,所述预设标点符号包括但不限于以下任意一种或多种:句号、感叹号、问号、逗号、空格、分号、顿号、省略号、表情符号、波浪符号。标准标点符号至少包括句号、感叹号、问号、逗号、分号、顿号、冒号、省略号。在一实施例中,先用标准标点符号对用户原创内容进行分句,如果分句后句子还是过长采用其他符号再次分句。按照各句子在所述用户原创内容中出现位置的前后顺序进行排列,得到所述用户原创内容包括的前后排列的M个句子。其中,M为大于等于1的自然数。
步骤120,确定每个所述句子的质量分。
在一实施例中,可以从句子包括的文本、观点和实体等信息维度的特征确定所述句子的质量分。其中,文本进一步可以包括:位置、长度、关键词情感属性、关键词对商户特征的描述等维度的信息。观点维度的信息可以为观点中包括的评价对象、评价词等信息。实体维度的信息可以为实体词的出现频次、实体词的类型等维度的信息。
句子的质量分用于表示该句子对所述用户原创内容的核心思想的贡献或表现能力。
步骤130,根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
在确定了用户原创内容中包括的前后排列的多个句子之后,选择信息含量最高的句组作为所述用户原创内容的摘要。在一实施例中,通过滑动窗口,找到包含的字符长度满足预设字符长度条件的多组句组。然后,根据每组句组中各句子的质量分,确定所述句组的评分。最后,选择评分最高的句组,作为所述用户原创内容的摘要。
本申请实施例公开的用户原创内容摘要确定方法,通过确定用户原创内容包括的前后排列的至少一个句子,确定每个所述句子的质量分;根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,能够高效并准确地提取用户原创内容的摘要。
实施例二
本实施例公开的一种原创内容摘要确定方法,如图2所示,该方法包括:步骤210至步骤240。
步骤210,构建评价对象库、评价词库和实体词库。
在一实施例中,为了确定用户原创内容中包括的句子的质量分,首先构建评价对象库、评价词库和实体词库,再基于评价对象库、评价词库和实体词库,确定句子中包括的实体、评价对象,以及句子中包括的情感类的关键词等。
在一实施例中,根据海量用户在平台上生成的数亿条UGC评论和每日千万级别的查询关键词,使用词法分析器得到名词和形容词等关键词,结合预设POI知识库的内容,使用N-Gram技术得到UGC评论中的所述关键词和所述查询关键词的词性类别(例如:景点、电影院、商区、商场等)。然后,通过评价对象挖掘,可以建成一个覆盖率比较高的评价对象库,为后续评论挖掘提供支持。
实体是评价对象中的一个子集,选自于商户、用户等的结构化数据中的关键词,例如:商家名称、菜品类别、菜品名称等。
关键词是指UGC文本经过分词后的有意义的词。评价词是指形容词、副词和成语等关键词。在一实施例中,获取UGC评论中高频的评价词,统计这些评价词在5星评论和1星评论中的分布情况,得到评价词的极性(正面、负面和中性)。比如“很好”这个评价词出现在好评的评论中的数量要远大于在差评中的数量,则“很好”这个评价词的极性为正面。通过评价词挖掘,可以建成一个评价词库,为后续评论挖掘提供支持。通过评价词可以确定句子的情感信息。
步骤220,确定用户原创内容包括的前后排列的多个句子。
在一实施例中,首先对用户原创内容进行数据处理,提取出所述用户原创内容中的句子,并将提取出的句子按照各句子在所述用户原创内容中出现的先后顺序进行排列。
由于用户原创内容,如用户点评,没有固定格式要求,所以内容和格式多样。在一实施例中,按照预设标点符号作为句子之间的分隔标记,将所述用户原创内容划分为多个句子。其中,所述预设标点符号包括但不限于以下任意一种或多种:句号、感叹号、问号、逗号、空格、分号、顿号、冒号、省略号、表情符号、波浪符号。标准标点符号至少包括句号、感叹号、问号、逗号、分号、顿号、冒号、省略号。在一实施例中,先用标准标点符号对用户原创内容进行分句,如果分句后句子还是过长采用其他符号再次分句。按照各句子在所述用户原创内容中出现位置的前后顺序进行排列,得到所述用户原创内容包括的前后排列的M个句子。其中,M为大于等于1的自然数。
在一实施例中,确定用户原创内容包括的前后排列的至少一个句子包括:基于标准标点符号对用户原创内容进行分句,得到所述用户原创内容包括的第一句子;基于扩展标点符号对所述第一句子中字符长度大于预设句子字符长度阈值的第一句子进行再次分句,得到所述第一句子对应的第二句子;将所述第一句子中字符长度未进行再次分句的所述第一句子和所述第二句子,按照在所述用户原创内容中出现位置的前后顺序进行排列,得到所述用户原创内容包括的前后排列的M个句子。其中,M为大于等于1的自然数。其中,标准标点符号至少包括:句号、逗号、问号、感叹号、省略号、冒号、顿号和分号,扩展标点符号包括:空格、表情符号、破浪符号等。
以一条用户原创内容为“地道巴蜀陈年酸菜,三年发酵而成,配合来自越南的无污染的龙利鱼^_^味道鲜嫩无比!”且预设句子字符长度阈值为10为例,说明如何确定用 户原创内容包括的前后排列的多个句子。首先,基于标准标点符号对用户原创内容进行分句,可以得到“地道巴蜀陈年酸菜”、“三年发酵而成”和“配合来自越南的无污染的龙利鱼^_^味道鲜嫩无比”共3个第一句子。对于第一句子“配合来自越南的无污染的龙利鱼^_^味道鲜嫩无比”,其字符长度为21,大于预设句子字符长度阈值,因此需要基于扩展标点符号进一步对其进行句子划分。由于该句子中包括一个表情符号“^_^”,因此,该句子基于该扩展标点符号进行划分后,得到2个第二句子,分别为:“配合来自越南的无污染的龙利鱼”和“味道鲜嫩无比”。最后,确定该用户原创内容中包括的4个句子为:第一句子“地道巴蜀陈年酸菜”、“三年发酵而成”,以及第二句子“配合来自越南的无污染的龙利鱼”和“味道鲜嫩无比”。之后,按照上述4个句子在所述用户原创内容中出现位置的前后顺序进行排列,得到所述用户原创内容包括的前后排列的4个句子,分别为:地道巴蜀陈年酸菜”、“三年发酵而成”、“配合来自越南的无污染的龙利鱼”和“味道鲜嫩无比”。
步骤230,确定每个所述句子的质量分。
句子的质量分用于表示该句子对所述用户原创内容的核心思想的贡献或表现能力。在一实施例中,确定每个所述句子的质量分,包括:根据所述句子的预设维度的信息,确定所述句子的质量分,其中,所述预设维度包括以下维度中的一个或多个:文本、实体和观点。根据所述句子的预设维度的信息,确定所述句子的质量分,包括:对所述句子的实体维度评分和观点维度评分进行加权求和得到初始质量分;通过所述句子的文本维度评分对所述初始质量分进行调整;并将调整后的初始质量分确定为所述句子的质量分。
在一实施例中,对所述句子的实体维度评分和观点维度评分进行加权求和得到初始质量分,并通过所述句子的文本维度评分对所述初始质量分进行调整,并将调整后的初始质量分确定为所述句子的质量分,包括根据以下公式确定所述句子的质量分:score(sentence
i)=w'×(α×score_sentence
i(word∈实体)+β×score_sentence
i(word∈评价对象))
其中,score(sentence
i)表示句子i的质量分,score_sentence
i(word∈实体)表示句子i的实体维度评分,score_sentence
i(word∈评价对象)表示句子i的观点维度评分,w'表示句子i的文本维度评分。
其中,评价对象为句子i中包括的观点包括的评价对象,α表示实体维度评分对应的第一权重调节因子,β表示观点维度评分对应的第二权重调节因子。即,首先,通过 以下公式计算初始值质量分:
α×score_sentence
i(word∈实体)+β×score_sentence
i(word∈评价对象)。
然后,通过文本维度评分w'对初始值质量分进行调整,得到句子i的质量分。
在一实施例中,根据句子在所述用户原创内容中的前后位置、句子的负面情感信息、商户特色信息确定句子的文本维度评分包括:提升靠近用户原创内容首部的句子的质量分、降低含有负面情感信息的句子的质量分、提升包括商户特色信息的句子的质量分。例如,对于出现在用户原创内容中的前三个句子,则提高这前三个句子的质量分,如加10分,以此提升用户原创内容首部位置句子出现在摘要中的概率。例如,如果句子中包括预设评价词库中的负面词语,则确定所述句子包含负面情感,通过降低句子的质量分,如减20分,降低这个句子出现在摘要中的概率。如果句子中包括预设评价词库中的广告词语,则通过降低句子的质量分,如减10分,降低该句子出现在摘要中的概率。再例如,如果句子中含有商户排名前三的推荐菜,或者含有商户类目下特色的评价对象,提高该句子的质量分,如加10分,从而提升该句子出现在摘要中的概率。
实体维度评分反映了实体在用户原创内容中的权重。在一实施例中,根据句子中包括的实体词的逆向文本词频确定句子的实体维度评分。例如,实体维度评分为所述句子中包括的实体的逆向文本词频之和,通过以下公式确定句子的实体维度评分:
该公式中,idf(word
j)为句子包括的实体词word
j的逆向文本词频。其中,所述实体的逆向文本词频可通过如下公式确定:
该公式中,|shop_num|为所有用户原创内容覆盖的商户总数,{k:word(j)∈shop
k}表示出现关键词word(j)的商户总数。
在一实施例中,根据句子中包括的观点涉及的评价对象的逆向文本词频确定句子的观点维度评分。
观点维度评分反映了观点中的评价对象在用户原创内容中的权重。在一实施例中,根据句子中包括的评价对象词的逆向文本词频确定句子的观点维度评分。例如,所述观点维度的信息为所述句子中包括的观点所涉及的评价对象的逆向文本词频之和,通过如 下公式确定句子的观点维度评分:
该公式中idf(word
l)为句子包括的评价对象word
l的逆向文本词频。其中,所述评价对象的逆向文本词频通过如下公式确定:
该公式中,|shop_num|为所有用户原创内容覆盖的商户总数,{k:word(l)∈shop
k}表示出现关键词word(l)的商户总数。
在一实施例中,根据句子中包括的观点涉及的评价对象的逆向文本词频确定句子的观点维度评分。例如,通过公式确定句子的观点维度评分,
该公式中,idf(word
l)为句子包括的评价对象word
l的逆向文本词频。
通过上述公式可以看出,如果实体或评价对象出现在用户原创内容(如商户评论)中的频率低,则相应的实体维度评分或观点维度评分权重高。进一步的,通过对实体维度评分和观点维度评分进行加权求和,得到句子的质量分。在一实施例中,实体维度评分和观点维度评分的权值通过经验统计设置。
步骤240,根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
在确定了用户原创内容中包括的前后排列的多个句子之后,选择信息含量最高的句组作为所述用户原创内容的摘要。
在一实施例中,通过如下公式确定begin和end之间的句组,作为所述用户原创内容的摘要:
其中,begin和end是所述用户原创内容中句子的顺序号,max_length为预设摘要 最大字符长度,length(sentence
i)为句子i中的字符长度,w是总分调节因子,w根据句子sentence
i,begin≤i≤end是否含有实体和观点、以及
确定。
根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,包括:通过滑窗技术确定满足摘要最大字符长度的约束条件的至少一组句组;针对各个所述句组,将所述句组中包括的各个句子的质量分的加权和确定为所述句组的质量分;;确定质量分最高的句组作为所述用户原创内容的摘要。在一实施例中,所述句组的质量分中各个句子的质量分的权重根据所述句组中各个句子是否含有实体和观点、所述句组的字符长度、所述句组中是否包含所述用户原创内容的首个句子或末尾句子中的任意一项或多项因素确定。
在一实施例中,假设预设摘要最大字符长度为35,以某条用户原创内容中包括前后排列的9个句子,每个句子的质量分和字符长度如下表所示为例,说明确定摘要的方法。其中,句子编号1至9为句子的前后排列序号,且各个句子的质量分的权重相同,例如,均为1。
| 句子1 | 句子2 | 句子3 | 句子4 | 句子5 | 句子6 | 句子7 | 句子8 | 句子9 | |
| 字符长度 | 10 | 9 | 6 | 8 | 16 | 7 | 8 | 9 | 10 |
| 质量分 | 0.5 | 0.2 | 1 | 2 | -10 | 2 | 3 | 3 | 2 |
在一实施例中,首先,从句子1开始,通过调整窗口的长度,找到长度不超过35个字符的句组,如{句子1},{句子1,句子2},{句子1,句子2,句子3},{句子1,句子2,句子3,句子4}。然后,分别确定每组句组的质量分,并保留质量分最高的一组句组,如{句子1,句子2,句子3,句子4}组成的句组作为候选摘要,所述候选摘要的质量分为3.7分。
接下来,滑动窗口,从句子2开始,通过调整窗口的长度,找到长度不超过35个字符的句组,如{句子2},{句子2,句子2},{句子2,句子3,句子4}。然后,分别确定每组句组的质量分,并保留质量分最高的一组句组,如{句子2,句子3,句子4}组成的句组,质量分为3.2分。
{句子1,句子2,句子3,句子4}组成的候选摘要的质量分大于{句子2,句子3,句子4}组成的句组的质量分(3.2分),因此,暂时保留{句子1,句子2,句子3,句子4}句组构成的候选摘要。
以此类推,通过滑窗技术,分别确定以每个句子开始的长度不超过35个字符的多组句组,并确定每组句组的质量分,以通过质量分更高的句组更新暂时保留的候选摘要,直至最后找到最高得分的句组,作为所述用户原创内容的摘要。以上表中的句子为例,最终将确定质量分为10分的句组{句子6,句子7,句子8,句子9}作为所述用户原创内容的摘要。
在一实施例中,根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,包括:通过滑窗技术确定满足摘要最大字符长度的约束条件的一个或多个句组;针对各个所述句组,将所述句组中包括的各个句子的质量分的加权和确定为所述句组的质量分;确定所述质量分最高的句组,作为所述用户原创内容的摘要。
在确定所述句组的质量分时,所述句组中各个句子的质量分可具有相同的权重,也可以具有不同的权重。
在一实施例中,假设所述句组中各个句子的质量分具有相同的权重,该权重与句组字符长度和预设摘要最大字符长度的比值为T,T为大于1的数,如T=1.5,这样可以避免确定的摘要的字符长度过短。在一实施例中,假设所述句组中各个句子的质量分具有不同的权重,如果句子的实体维度评分为零,例如该句子中不包括实体,则降低所述句子的质量分的权重;如果句子的观点维度评分为零,例如句子中不包括评价对象,则降低所述句子的质量分的权重;如果句子是所述用户原创内容的第一个句子或最后一个句子,则提升所述句子的质量分的权重。根据所述句子是否是所述用户原创内容的首个句子或末尾句子确定该句子的质量分的权重,可以提升确定的摘要中句子的完整性。
本申请实施例公开的用户原创内容摘要确定方法,通过确定用户原创内容包括的前后排列的多个句子;然后,确定每个所述句子的质量分;最后,根根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,能够高效并准确地提取用户原创内容的摘要题。本申请实施例中,通过所述用户原创内容的文本、实体和观点三个维度加权计算得到句子的质量分,通过这种方法,能找到用户原创内容中信息价值密度最高的句组。并且,本申请实施例公开的用户原创内容摘要确定方法,支持标点符号使用不规范,甚至语句不通顺的用户原创内容摘要的抽取,鲁棒性更强;可以根据对摘要长度的不同要求,自适应抽取商户特色的用户原创内容摘要。
实施例三
本实施例公开的一种原创内容推荐方法,如图3所示,该方法包括:步骤310至步骤350。
步骤310,确定用户的目标商户。
在一实施例中,首先根据用户的历史行为数据,确定用户发生过预设历史行为的商户,作为第一目标商户;然后,确定与所述第一目标商户相似的商户,作为第二目标商户;最后,将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。
步骤320,根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容。
获取所述目标商户的用户原创内容,并进一步确定每条用户原创内容的评价得分。在一实施例中,可以根据用户原创内容的文本信息、实体信息以及观点信息等确定用户原创内容的评价得分。在一实施例中,评价得分越高表示所述用户原创内容的质量越高,即所述用户原创内容展示给用户的信息更有价值。然后,按照用户原创内容的评价得分由高到低的顺序,对每个所述目标用户的用户原创内容分别进行排序。之后,对于每一个目标用户,分别选择评价得分最高的预设数量的用户原创内容,作为候选用户原创内容。
步骤330,确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容。
在一实施例中,可以分别提取用户的特征向量,以及每一条候选用户原创内容的特征向量,然后,通过计算用户的特征向量与每一条候选用户原创内容的特征向量之间的相似度,确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容。在一实施例中,可以通过计算用户的特征向量与候选用户原创内容的特征向量之间的相似度距离,确定用户与某一条候选用户原创内容的之间的匹配度;或者,通过预先训练的机器学习排序模型,根据输入的用户的特征向量和某一条所述候选用户原创内容的特征向量,计算用户与所述某一条候选用户原创内容之间的匹配度。
然后,选择与所述用户匹配度最高的一个或预设数量的所述候选用户原创内容,作为所述目标用户原创内容。
步骤340,确定所述目标用户原创内容的摘要。
根据实施例一和实施例二所述用户原创内容摘要确定方法,确定所述目标用户原创内容的摘要。
步骤350,向所述用户推荐所述目标用户原创内容的摘要。
在确定了与所述用户匹配的目标用户原创内容后,向所述用户推荐所述目标用户原创内容的摘要。
本申请实施例公开的用户原创内容推荐方法,通过确定用户的目标商户;根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容;最后,向所述用户推荐所述目标用户原创内容的摘要,其中,所述目标用户原创内容的摘要根据实施例一或实施例二所述的用户原创内容摘要确定方法确定,这样,与根据用户原创内容的热度为用户推荐用户原创内容的方案相比,实现了根据用户需求推荐更准确的用户原创内容。本申请实施例公开的用户原创内容推荐方法,通过把与用户匹配的用户原创内容推荐给用户,实现了有针对性的进行信息推荐,有效提升了用户原创内容推荐的准确性。同时,通过在为用户推荐原创内容时,仅展示原创内容的摘要,简洁清晰的为用户展示推荐的关键信息,便于用户准确快速的做出决策,进一步提升了用户体验。
实施例四
本实施例公开的一种用户原创内容推荐方法,如图4所示,该方法包括:步骤410至步骤470。
步骤410,构建评价对象库、评价词库和实体词库。
构建评价对象库、评价词库和实体词库的具体实施方式参见实施例二,本实施例不再赘述。
步骤420,确定用户的目标商户。
在一实施例中,确定用户的目标商户,包括:确定所述用户产生过预设行为的商户,作为第一目标商户;基于商户向量的相似度,确定与所述第一目标商户相似的第二目标商户;将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。
在一实施例中,首先根据用户的历史行为数据,确定用户发生过预设历史行为的商户,作为第一目标商户。其中,用户产生过预设行为的商户包括但不限于:用户点击过的商户、用户浏览过的商户、用户收藏过的商户、用户购买过商品的商户。
然后,进一步确定与所述第一目标商户相似的商户,作为第二目标商户。
在一实施例中,基于商户向量的相似度,确定与所述第一目标商户相似的第二目标商户之前,还包括:将用户点击的商户序列作为词向量模型输入,训练商户向量模型; 通过所述商户向量模型确定商户的商户向量。
在一实施例中,把用户在商户上的行为转变为时间序列事件,然后,把时间序列事件作为输入,采用深度学习算法训练商户向量模型,即把商户特征从高维的离散空间映射到低维的连续空间。例如,当用户先后点击了商户1、商户2和商户3,那么,可以把商户1、商户2和商户3的商户标识序列作为输入样本,用于训练商户向量模型。然后,通过预先训练的商户向量模型,可以获得某个商户标识对应的商户向量。
在确定了每个商户的商户向量之后,通过计算各个商户向量与所述第一目标商户之间的相似度,可以确定与所述第一目标商户相似的第二目标商户。
最后,将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。例如,根据用户的历史行为,确定用户曾经点击过商户1,则将商户1作为用户的第一目标商户。然后,通过计算商户向量的相似度,确定与商户1相似的商户2,则将商户2作为所述用户的第二目标商户。最后,将商户1和商户2作为所述用户的目标商户。
步骤430,根据所述目标商户的用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。
根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容之前,还包括:根据所述目标商户的用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。例如,根据所述目标商户的用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分,可以包括:通过对用户原创内容的文本得分、实体得分和观点得分进行加权求和,所述用户原创内容的评价得分。
在一实施例中,首先,对于平台的用户原创内容,如用户点评,选取最近预设时间(如半年内)的用户原创内容。然后,根据用户原创内容中的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。因为优质商户或者高星级用户下也存在低质的用户原创内容,所以在对用户原创内容进行评分时,不考虑商户和用户的特征,只从用户原创内容的内容质量本身来分析,通过文本、实体和观点三个维度计算得到用户原创内容的评价得分。
在一实施例中,文本得分与用户原创内容中包含的不同文字的数量成正比。即,用户原创内容中包含的不同文字越多,文本得分越高。根据用户原创内容中包含的不同文字的数量确定文本得分,可以有效过滤掉用户重复使用同一个标点符号或者文字来充当字数的用户原创内容。
在一实施例中,实体得分可以通过用户原创内容中包含的实体的逆向文本词频表示;观点得分可以通过用户原创内容中包含的观点所涉及的评价对象的逆向文本词频表示。
在确定实体得分和观点得分之前,首先,将用户原创内容划分为多个句子。将用户原创内容划分为多个句子的具体方法可以参考实施例二中确定用户原创内容中的句子的方法,本实施例不再赘述。
然后,通过预设的实体词库,确定由用户原创内容中划分得到的每个句子中包括的实体和观点。
实体是指用户原创内容中涉及的评论对象,例如,商户名、地址、类目、商场、星级酒店、商场、小区、电影院、行政区和城市等。实体是用户原创内容中的重要信息,例如,一条用户原创内容中提到的推荐菜、地址和类目等内容的信息,可以作为该条用户原创内容的重要特征。O2O(online to offline,O2O)场景下的信息抽取有别于传统的人名、地名和公司名识别,需要挖掘不同维度下不同关键词的权重信息,例如在美食品类下的商家评论中,“龙之梦”的评论数很少,其逆向文本词频要高于“粤菜”。在一实施例中,可以通过如下公式确定一条用户原创内容的实体得分:
该公式中,idf(word
p)为该条用户原创内容包括的实体词word
p的逆向文本词频。其中,所述实体词的逆向文本词频通过如下公式确定。
该公式中,|shop_num|为所有用户原创内容覆盖的商户总数,{k:word(p)∈shop
k}表示出现关键词word
p的商户总数。
观点表示对具体的评价对象的主客观判断信息,本申请中,主要从句子中抽取观点。例如,对于一条用户原创内容中的一个句子“浓缩咖啡豆是皮爷家的经典”,从该句子中抽取观点的具体方法如下:根据预先构建的评价对象库可以确定该句子中包括的评价对象是:咖啡豆;根据预先构建的评价词库可以确定该句子中包括的评价词是:“浓缩”、“经典”;将该句子中包括的评价对象和评价词组合,得到该句子中包括的观点,即:“咖啡豆-经典”和“咖啡豆-浓缩”。再后,根据上述两个观点在所述用户原创内容中出现的比率得到每个观点的置信度,在一实施例中,观点出现越频繁则置信度越高。最 后得到一条所述用户原创内容中所有的观点,以及每个观点的置信度。
对于一条用户原创内容中得到的每个观点,通过对所述观点包括的评价对象和评价词的词向量进行求和,得到所述观点的向量表示。通过向量对观点进行表示之后,就可以采用余弦定理计算向量之间的距离,来判断观点之间的相似关系。在一实施例中,通过对句子进行分析,可以得到如下观点数据结构表:
| 字段名称 | 字段说明 | 示例 |
| Opinion | 观点 | 咖啡豆-经典 |
| SemanticVector | 词向量 | [0,1,0.32,0.16,0.07…] |
| Aspect | 评价对象 | 咖啡豆 |
| Evaluate | 评价词 | 经典 |
| Confidence | 置信度 | 0.87 |
| Updatetime | 更新时间 | 2018-03-12 09:00:00 |
在一实施例中,基于用户产生的全部用户原创内容,通过分词处理后得到训练样本,使用本领域技术人员熟知的词向量技术,得到训练样本中每个关键词的词向量。在一实施例中,关键词包括实体词、评价词以及各种有意义的通用词汇。词向量是关键词的向量表示。在一实施例中,关键词的词向量为固定长度的浮点型一维向量,例如,采用skip-gram模型的负采样方法训练词向量模型。采用词向量技术后,所有关键词都可以用一个固定长度的向量表示,将原来稀疏的巨大维度压缩到一个更小维度空间,例如“披萨”和“pizza”这两个词在文本上没有相似性,但是通过词向量表示后,其语义距离比较接近。
最后,通过对一条用户原创内容中包括的实体的实体得分、观点的观点得分,以及文本得分进行加权求和,将得到的总得分和作为该条用户原创内容的评价得分。在一实施例中,对实体得分、观点得分以及文本得分进行加权,各项得分的加权值根据具体需求设置,通常,观点得分的权值最高,文本得分的权值最低。
步骤440,根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容。
如前所述,假设将商户1和商户2作为用户的目标商户,则进一步根据用户原创内容的评价得分,在所述商户1和商户2的用户原创内容中,分别选择评价得分满足预设条件的多条用户原创内容作为用户的候选用户原创内容。例如,按照评价得分由高到低的顺序,分别对商户1和商户2的用户原创内容排序,然后选择商户1评价得分最高 的M条用户原创内容和商户2评价得分最高的M条用户原创内容,作为候选用户原创内容。
步骤450,确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容。
在一实施例中,确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容,包括:根据每条所述候选用户原创内容的排序特征和所述用户的用户特征,分别确定每条所述候选用户原创内容与所述用户的匹配度;确定所述匹配度满足预设条件的所述候选用户原创内容,作为与所述用户匹配的目标用户原创内容。
在一实施例中,可以首先基于用户原创内容的排序特征和用户的用户特征通过机器学习训练匹配度识别模型。例如,将用户原创内容的排序特征和发布该原创内容的用户的用户特征组合为正样本,将用户原创内容的排序特征和踩了该原创内容的用户的用户特征组合为负样本,训练匹配度识别模型。然后,通过该匹配度识别模型基于输入的用户原创内容的排序特征和用户的用户特征识别所述用户原创内容和所述用户的匹配度。其中,所述排序特征包括:点赞数、评论数、分享数、文本质量分、图片质量分、实体词、用户原创内容发布者等级、发布者与所述用户的关系中的任意一项或多项;所述用户特征包括:用户历史行为特征、商区偏好特征、类目偏好特征、相似用户特征中的任意一项或多项,所述用户历史行为特征包括:搜索、浏览、购买、到店行为中的任意一项或多项的特征。
在一实施例中,可以确定所述匹配度得分最高的预设数量的所述候选用户原创内容,作为与所述用户匹配的目标用户原创内容;或者,确定每个商户对应的所述候选用户原创内容中,与所述用户的所述匹配度得分最高的一条所述候选用户原创内容,作为与所述用户匹配的目标用户原创内容。由于进行匹配度识别时,结合了用户偏好、用户社交关系等特征,因此,所确定的目标用户原创内容,是用户偏好的用户原创内容。
步骤460,确定所述目标用户原创内容的摘要。
在一实施例中,通过是实施例一和实施例二所述的用户原创内容摘要确定方法确定所述目标用户原创内容的摘要,本实施例中,对摘要的具体确定方法不再赘述。
步骤470,向所述用户推荐所述目标用户原创内容的摘要。
在确定了与所述用户匹配的目标用户原创内容后,向所述用户推荐所述目标用户原创内容的摘要。
本申请实施例公开的用户原创内容推荐方法,通过确定用户的目标商户;然后, 确定所述目标商户的用户原创内容的评价得分,并根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容及其摘要;最后,向所述用户推荐所述目标用户原创内容的摘要。这样,与根据用户原创内容的热度为用户推荐用户原创内容的方案相比,能够根据用户需求推荐更加准确的用户原创内容。本申请实施例公开的用户原创内容推荐方法,通过把与用户匹配的用户原创内容推荐给用户,实现了有针对性的进行信息推荐,有效提升了用户原创内容推荐的准确性。同时,通过在为用户推荐用户原创内容时,仅展示用户原创内容的摘要,简洁清晰的为用户展示推荐的关键信息,便于用户准确快速的做出决策,进一步提升了用户体验。
通过用户原创内容的文本、实体和观点的信息,确定用户原创内容的评价得分,能够提升用户原创内容质量评价的准确性,进一步提升用户原创内容推荐的准确性。
实施例五
本实施例公开的一种原创内容摘要确定装置,如图5所示,所述装置包括:
句子确定模块510,用于确定用户原创内容包括的前后排列的至少一个句子。
句子质量分确定模块520,用于确定每个所述句子的质量分。
摘要确定模块530,用于根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
可选的,所述句子质量分确定模块520进一步用于:
根据每个所述句子的预设维度的信息,确定每个所述句子的质量分,其中,所述预设维度包括以下维度中的一个或多个:文本、实体和观点。
可选的,所述根据每个所述句子的预设维度的信息,确定每个所述句子的质量分,包括:对每个所述句子的实体维度评分和观点维度评分进行加权求和得到初始质量分,并通过所述句子的文本维度评分对所述初始质量分进行调整;并将调整后的初始质量分确定为所述句子的质量分。在本申请的一个实施例中,所述对每个所述句子的实体维度评分和观点维度评分进行加权求和得到初始质量分,并通过所述句子的文本维度评分对所述初始质量分进行调整;并将调整后的初始质量分确定为所述句子的质量分,进一步包括:
根据如下公式确定每个所述句子的质量分:
score(sentence
i)=w'×(α×score_sentence
i(word∈实体)+β×score_sentence
i(word∈评价对象))
其中,score(sentence
i)表示句子i的质量分,score_sentence
i(word∈实体)表示句子i的实体维度评分,score_sentence
i(word∈评价对象)表示句子i的观点维度评分,w'表示句子i的文本维度评分。其中,评价对象为句子中包括的观点针对的评价对象,α表示实体维度评分对应的第一权重调节因子,β表示观点维度评分对应的第二权重调节因子。
可选的,所述摘要确定模块530进一步用于:
通过滑窗技术确定满足摘要最大字符长度的约束条件的一个或多个句组;
针对各个所述句组,将所述句组中包括的各个句子的质量分的加权和确定为所述句组的质量分;
确定所述质量分最高的句组,作为所述用户原创内容的摘要。
可选的,所述句组的质量分中各个句子的质量分的权重根据所述句组中各个句子是否含有实体和观点、所述句组的字符长度、所述句组中是否包含所述用户原创内容的首个句子或末尾句子中的任意一项或多项因素确定。
本实施例是与实施例一和实施例二对应的装置实施例,本实施例中各模块的具体实现方式参见实施例一和实施例二中的相关步骤的描述,此处不再赘述。
通过确定用户原创内容包括的前后排列的多个句子并确定每个所述句子的质量分,然后,根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。本公开实例中的用户原创内容推荐装置解决了无法准确提取用于原创内容摘要的问题。经过大量用户原创内容的测试,本申请公开的用户原创内容摘要确定装置,可以高效、准确的确定用户原创内容的摘要。通过用户原创内容的文本、实体和观点三个维度加权计算得到句子质量的方法,本公开实例能找到用户原创内容中信息价值密度最高的句组。并且,本申请实施例公开的远传内容摘要确定方法,支持标点符号使用不规范,甚至语句不通顺的用户原创内容摘要的抽取,鲁棒性更强;可以根据对摘要长度的不同要求,自适应抽取商户特色的用户原创内容摘要。
实施例六
本实施例公开的一种原创内容推荐装置,如图6所示,所述装置包括:
目标商户确定模块610,用于确定用户的目标商户。
候选用户原创内容确定模块620,用于根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容。
匹配候选用户原创内容确定模块630,用于确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容。
原创内容摘要确定模块640,用于根据本申请实施例所述用户原创内容摘要确定方法,确定所述目标用户原创内容的摘要。
推荐模块650,用于向所述用户推荐所述目标用户原创内容的摘要,其中,所述目标用户原创内容的摘要根据实施例一和实施例二所述的用户原创内容摘要确定方法确定。
可选的,如图7所示,所述装置还包括:
用户原创内容评价得分确定模块660,用于根据所述用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。
可选的,所述目标商户确定模块610进一步用于:
确定所述用户产生过预设行为的商户,作为第一目标商户;基于商户向量的相似度,确定与所述第一目标商户相似的第二目标商户;将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。
可选的,所述目标商户确定模块610还用于:
将用户点击的商户序列作为词向量模型的输入,训练商户向量模型;通过所述商户向量模型确定所述第一目标商户的所述商户向量。
可选的,所述匹配候选用户原创内容确定模块630进一步用于:
根据每条所述候选用户原创内容的排序特征和所述用户的用户特征,分别确定每条所述候选用户原创内容与所述用户的匹配度;确定所述匹配度满足预设条件的所述候选用户原创内容,作为与所述用户匹配的目标用户原创内容。
其中,所述排序特征包括:点赞数、评论数、分享数、文本质量分、图片质量分、实体词、用户原创内容发布者等级、发布者与所述用户的关系中的任意一项或多项;所 述用户特征包括:用户历史行为特征、商区偏好特征、类目偏好特征、相似用户特征中的任意一项或多项,所述用户历史行为特征包括:搜索、浏览、购买、到店行为中的任意一项或多项的特征。
本实施例是与实施例三和实施例四对应的装置实施例,本实施例中各模块的具体实现方式参见实施例三和实施例四中的相关步骤的描述,此处不再赘述。
通过确定用户的目标商户;然后,确定所述目标商户的用户原创内容的评价得分,并根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容及其摘要;最后,向所述用户推荐所述目标用户原创内容的摘要。本公开实例中的用户原创内容推荐装置解决了根据用户原创内容的热度为用户推荐用户原创内容时,推荐的用户原创内容不准确,无法满足用户需求的问题。通过把与用户匹配的用户原创内容推荐给用户,实现了有针对性的进行信息推荐,本公开实例中的用户原创内容推荐装置有效提升了用户原创内容推荐的准确性。同时,通过在为用户推荐原创内容时,仅展示原创内容的摘要,简洁清晰的为用户展示推荐的关键信息,便于用户准确快速的做出决策,进一步提升了用户体验。
通过用户原创内容的文本、实体和观点的信息,确定用户原创内容的评价得分,能够提升用户原创内容质量评价的准确性,进一步提升用户原创内容推荐的准确性。
相应的,本申请还公开了一种电子设备,包括存储器、处理器及存储在所述存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如本申请实施例一和实施例二所述的原创内容摘要确定方法、实施例三和实施例四所述的用于原创内容推荐方法。所述电子设备可以为PC机、移动终端、个人数字助理、平板电脑等。
本申请还公开了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如本申请实施例一和实施例二所述的原创内容摘要确定方法或实施例三和实施例四所述的用户原创内容推荐方法。
本说明书中的各个实施例均采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似的部分互相参见即可。对于装置实施例而言,由于其与方法实施例基本相似,所以描述的比较简单,相关之处参见方法实施例的部分说明即可。
以上对本申请提供的一种用户原创内容摘要确定方法及装置,用户原创内容推荐 方法及装置进行了详细介绍,本文中应用了具体个例对本申请的原理及实施方式进行了阐述,以上实施例的说明只是用于帮助理解本申请的方法及其核心思想;同时,对于本领域的一般技术人员,依据本申请的思想,在具体实施方式及应用范围上均会有改变之处,综上所述,本说明书内容不应理解为对本申请的限制。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到各实施方式可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件实现。基于这样的理解,上述技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如ROM/RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行各个实施例或者实施例的某些部分所述的方法。
Claims (22)
- 一种用户原创内容的摘要确定方法,包括:确定用户原创内容包括的前后排列的多个句子;确定每个所述句子的质量分;根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
- 根据权利要求1所述的方法,确定所述句子的质量分,包括:根据所述句子的预设维度的信息,确定所述句子的质量分,其中,所述预设维度包括以下维度中的一个或多个:文本、实体和观点。
- 根据权利要求2所述的方法,根据所述句子的预设维度的信息,确定所述句子的质量分,包括:对所述句子的实体维度评分和观点维度评分进行加权求和,得到初始质量分;通过所述句子的文本维度评分对所述初始质量分进行调整;将调整后的所述初始质量分确定为所述句子的质量分。
- 根据权利要求1所述的方法,根据所述摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的所述句组,作为所述用户原创内容的摘要,包括:通过滑窗技术确定满足所述摘要最大字符长度的约束条件的一个或多个句组;针对各个所述句组,将所述句组中包括的各个句子的质量分的加权和确定为所述句组的质量分;确定所述质量分最高的句组,作为所述用户原创内容的摘要。
- 根据权利要求4所述的方法,所述句组中包括的各个句子的质量分的权重由以下任意一项或多项因素确定:针对所述句组中包括的各个句子,所述句子是否含有实体和观点;所述句组的字符长度;和所述句组中是否包含所述用户原创内容的首个句子或末尾句子。
- 一种用户原创内容推荐方法,包括:确定用户的目标商户;根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容;根据权利要求1至5任一项所述用户原创内容摘要确定方法,确定所述目标用户原创内容的摘要;向所述用户推荐所述目标用户原创内容的摘要。
- 根据权利要求6所述的方法,所述方法还包括:根据所述用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。
- 根据权利要求6所述的方法,确定所述用户的目标商户,包括:确定所述用户产生过预设行为的商户,作为第一目标商户;基于商户向量的相似度,确定与所述第一目标商户相似的第二目标商户;将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。
- 根据权利要求8所述的方法,所述方法还包括:将用户点击的商户序列作为词向量模型的输入,训练商户向量模型;通过所述商户向量模型确定所述第一目标商户的所述商户向量。
- 根据权利要求6所述的方法,确定所述候选用户原创内容中与所述用户匹配的所述目标用户原创内容,包括:根据每条所述候选用户原创内容的排序特征和所述用户的用户特征,分别确定每条所述候选用户原创内容与所述用户的匹配度;确定所述匹配度满足预设条件的所述候选用户原创内容,作为与所述用户匹配的所述目标用户原创内容;其中,所述排序特征包括:点赞数、评论数、分享数、文本质量分、图片质量分、实体词、用户原创内容发布者等级、发布者与所述用户的关系中的任意一项或多项;所述用户特征包括:用户历史行为特征、商区偏好特征、类目偏好特征、相似用户特征中的任意一项或多项;所述用户历史行为特征包括:搜索、浏览、购买、到店行为中的任意一项或多项的特征。
- 一种原创内容摘要确定装置,包括:句子确定模块,用于确定用户原创内容包括的前后排列的多个句子;句子质量分确定模块,用于确定每个所述句子的质量分;摘要确定模块,用于根据摘要最大字符长度的约束条件和每个所述句子的质量分,确定质量分最高的句组,作为所述用户原创内容的摘要,其中,所述句组中包括的句子是连续的。
- 根据权利要求11所述的装置,所述句子质量分确定模块进一步用于:根据每个所述句子的预设维度的信息,确定每个所述句子的质量分,其中,所述预 设维度包括以下维度中的一个或多个:文本、实体和观点。
- 根据权利要求12所述的装置,根据每个所述句子的预设维度的信息,确定每个所述句子的质量分,包括:对所述句子的实体维度评分和观点维度评分进行加权求和,得到初始质量分;通过所述句子的文本维度评分对所述初始质量分进一步加权调整;将调整后的所述初始质量分确定为所述句子的质量分。
- 根据权利要求11所述的装置,所述摘要确定模块进一步用于:通过滑窗技术确定满足所述摘要最大字符长度的约束条件的一个或多个句组;针对各个所述句组,将所述句组中包括的各个句子的质量分的加权和确定为所述句组的质量分;确定所述质量分最高的句组。
- 根据权利要求14所述的装置,所述句组中包括的各个句子的质量分的权重由以下任意一项或多项因素确定:针对所述句组中包括的各个句子,所述句子是否含有实体和观点;所述句组的字符长度;和所述句组中是否包含所述用户原创内容的首个句子或末尾句子。
- 一种用户原创内容推荐装置,包括:目标商户确定模块,用于确定用户的目标商户;候选用户原创内容确定模块,用于根据所述目标商户的用户原创内容的评价得分,确定候选用户原创内容;匹配候选用户原创内容确定模块,用于确定所述候选用户原创内容中与所述用户匹配的目标用户原创内容;原创内容摘要确定模块,用于根据权利要求1至5任一项所述用户原创内容摘要确定方法,确定所述目标用户原创内容的摘要;推荐模块,用于向所述用户推荐所述目标用户原创内容的摘要。
- 根据权利要求16所述的装置,还包括:用户原创内容评价得分确定模块,用于根据所述用户原创内容的文本、实体和观点三个维度的信息,确定所述用户原创内容的评价得分。
- 根据权利要求16所述的装置,所述目标商户确定模块进一步用于:确定所述用户产生过预设行为的商户,作为第一目标商户;基于商户向量的相似度,确定与所述第一目标商户相似的第二目标商户;将所述第一目标商户和所述第二目标商户,作为所述用户的目标商户。
- 根据权利要求16所述的装置,所述目标商户确定模块还用于:将用户点击的商户序列作为词向量模型的输入,训练商户向量模型;通过所述商户向量模型确定所述第一目标商户的所述商户向量。
- 根据权利要求16所述的装置,所述匹配候选用户原创内容确定模块进一步用于:根据每条所述候选用户原创内容的排序特征和所述用户的用户特征,分别确定每条所述候选用户原创内容与所述用户的匹配度;确定所述匹配度满足预设条件的所述候选用户原创内容,作为与所述用户匹配的所述目标用户原创内容;其中,所述排序特征包括:点赞数、评论数、分享数、文本质量分、图片质量分、实体词、用户原创内容发布者等级、发布者与所述用户的关系中的任意一项或多项;所述用户特征包括:用户历史行为特征、商区偏好特征、类目偏好特征、相似用户特征中的任意一项或多项;所述用户历史行为特征包括:搜索、浏览、购买、到店行为中的任意一项或多项的特征。
- 一种电子设备,包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现权利要求1至5任意一项所述的用户原创内容摘要确定方法或实现权利要求6至10任意一项所述的用户原创内容推荐方法。
- 一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现权利要求1至5任意一项所述用户原创内容摘要确定方法或实现权利要求6至10任意一项所述用户原创内容推荐方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/093,969 US20210056571A1 (en) | 2018-05-11 | 2020-11-10 | Determining of summary of user-generated content and recommendation of user-generated content |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810447372.7 | 2018-05-11 | ||
| CN201810447372.7A CN108628833B (zh) | 2018-05-11 | 2018-05-11 | 原创内容摘要确定方法及装置,原创内容推荐方法及装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/093,969 Continuation US20210056571A1 (en) | 2018-05-11 | 2020-11-10 | Determining of summary of user-generated content and recommendation of user-generated content |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019214236A1 true WO2019214236A1 (zh) | 2019-11-14 |
Family
ID=63692812
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/121321 Ceased WO2019214236A1 (zh) | 2018-05-11 | 2018-12-14 | 原创内容摘要确定和原创内容推荐 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20210056571A1 (zh) |
| CN (1) | CN108628833B (zh) |
| WO (1) | WO2019214236A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111858873A (zh) * | 2020-04-21 | 2020-10-30 | 北京嘀嘀无限科技发展有限公司 | 一种推荐内容的确定方法、装置、电子设备及存储介质 |
Families Citing this family (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108628833B (zh) * | 2018-05-11 | 2021-01-22 | 北京三快在线科技有限公司 | 原创内容摘要确定方法及装置,原创内容推荐方法及装置 |
| CN109151521B (zh) * | 2018-10-15 | 2021-03-02 | 北京字节跳动网络技术有限公司 | 一种用户原创值获取方法、装置、服务器及存储介质 |
| CN110334192B (zh) * | 2019-07-15 | 2021-09-24 | 河北科技师范学院 | 文本摘要生成方法及系统、电子设备及存储介质 |
| CN110688845B (zh) * | 2019-10-10 | 2024-02-13 | 汉海信息技术(上海)有限公司 | 菜谱类内容的识别方法、装置、终端及可读存储介质 |
| CN111241242B (zh) * | 2020-01-09 | 2023-05-30 | 北京百度网讯科技有限公司 | 目标内容的确定方法、装置、设备及计算机可读存储介质 |
| CN111737382A (zh) * | 2020-05-15 | 2020-10-02 | 百度在线网络技术(北京)有限公司 | 地理位置点的排序方法、训练排序模型的方法及对应装置 |
| CN112579800A (zh) * | 2020-08-28 | 2021-03-30 | 太极计算机股份有限公司 | 一种融媒体新闻原创作品及首发媒体自动识别方法 |
| CN113535942B (zh) * | 2021-07-21 | 2022-08-19 | 北京海泰方圆科技股份有限公司 | 一种文本摘要生成方法、装置、设备及介质 |
| WO2023007270A1 (en) * | 2021-07-26 | 2023-02-02 | Carl Wimmer | Foci analysis tool |
| CN114281981B (zh) * | 2021-12-22 | 2023-05-02 | 北京百度网讯科技有限公司 | 新闻简报的生成方法、装置和电子设备 |
| CN114428837B (zh) * | 2021-12-31 | 2026-03-10 | 东软集团股份有限公司 | 内容质量评价方法、装置、介质及电子设备 |
| US12307188B2 (en) | 2022-04-13 | 2025-05-20 | Servicenow, Inc. | Labeled clustering preprocessing for natural language processing |
| US12271699B2 (en) * | 2022-04-13 | 2025-04-08 | Servicenow, Inc. | Multi-dimensional N-gram preprocessing for natural language processing |
| CN115221863B (zh) * | 2022-07-18 | 2023-08-04 | 桂林电子科技大学 | 一种文本摘要评价方法、装置以及存储介质 |
| US20240062020A1 (en) * | 2022-08-16 | 2024-02-22 | Microsoft Technology Licensing, Llc | Unified natural language model with segmented and aggregate attention |
| US12430511B2 (en) * | 2022-08-24 | 2025-09-30 | Maplebear Inc. | Generating suggested instructions through natural language processing of instruction examples |
| CN115795025A (zh) * | 2022-11-29 | 2023-03-14 | 华为技术有限公司 | 一种摘要生成方法及其相关设备 |
| CN116433800B (zh) * | 2023-06-14 | 2023-10-20 | 中国科学技术大学 | 基于社交场景用户偏好与文本联合指导的图像生成方法 |
| US12561525B2 (en) * | 2023-07-31 | 2026-02-24 | Paypal, Inc. | Systems and methods for establishing multilingual context-preserving chunk library |
| CN118694818B (zh) * | 2024-08-22 | 2024-11-08 | 青岛漫斯特数字科技有限公司 | 基于人工智能的旅游服务信息推送方法及系统 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101667194A (zh) * | 2009-09-29 | 2010-03-10 | 北京大学 | 基于用户评论文本特征的自动摘要方法及其自动摘要系统 |
| CN105868175A (zh) * | 2015-12-03 | 2016-08-17 | 乐视网信息技术(北京)股份有限公司 | 摘要生成方法及装置 |
| CN107609960A (zh) * | 2017-10-18 | 2018-01-19 | 口碑(上海)信息技术有限公司 | 推荐理由生成方法及装置 |
| CN108628833A (zh) * | 2018-05-11 | 2018-10-09 | 北京三快在线科技有限公司 | 原创内容摘要确定方法及装置,原创内容推荐方法及装置 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002132677A (ja) * | 2000-10-20 | 2002-05-10 | Oki Electric Ind Co Ltd | 電子メール転送装置及び電子メール装置 |
| US20040133560A1 (en) * | 2003-01-07 | 2004-07-08 | Simske Steven J. | Methods and systems for organizing electronic documents |
| CN100492366C (zh) * | 2007-06-28 | 2009-05-27 | 腾讯科技(深圳)有限公司 | 摘要提取方法以及摘要提取模块 |
| CN104615772B (zh) * | 2015-02-16 | 2017-11-03 | 重庆大学 | 一种用于电子商务的文本评价数据专业程度分析方法 |
| US20170186102A1 (en) * | 2015-12-29 | 2017-06-29 | Linkedin Corporation | Network-based publications using feature engineering |
| WO2018058096A1 (en) * | 2016-09-26 | 2018-03-29 | Contiq, Inc. | Systems and methods for constructing presentations |
| CN106600360B (zh) * | 2016-11-11 | 2020-05-12 | 北京星选科技有限公司 | 推荐对象的排序方法及装置 |
| CN108959312B (zh) * | 2017-05-23 | 2021-01-29 | 华为技术有限公司 | 一种多文档摘要生成的方法、装置和终端 |
-
2018
- 2018-05-11 CN CN201810447372.7A patent/CN108628833B/zh active Active
- 2018-12-14 WO PCT/CN2018/121321 patent/WO2019214236A1/zh not_active Ceased
-
2020
- 2020-11-10 US US17/093,969 patent/US20210056571A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101667194A (zh) * | 2009-09-29 | 2010-03-10 | 北京大学 | 基于用户评论文本特征的自动摘要方法及其自动摘要系统 |
| CN105868175A (zh) * | 2015-12-03 | 2016-08-17 | 乐视网信息技术(北京)股份有限公司 | 摘要生成方法及装置 |
| CN107609960A (zh) * | 2017-10-18 | 2018-01-19 | 口碑(上海)信息技术有限公司 | 推荐理由生成方法及装置 |
| CN108628833A (zh) * | 2018-05-11 | 2018-10-09 | 北京三快在线科技有限公司 | 原创内容摘要确定方法及装置,原创内容推荐方法及装置 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111858873A (zh) * | 2020-04-21 | 2020-10-30 | 北京嘀嘀无限科技发展有限公司 | 一种推荐内容的确定方法、装置、电子设备及存储介质 |
| CN111858873B (zh) * | 2020-04-21 | 2024-06-04 | 北京嘀嘀无限科技发展有限公司 | 一种推荐内容的确定方法、装置、电子设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN108628833B (zh) | 2021-01-22 |
| CN108628833A (zh) | 2018-10-09 |
| US20210056571A1 (en) | 2021-02-25 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019214236A1 (zh) | 原创内容摘要确定和原创内容推荐 | |
| CN111221962B (zh) | 一种基于新词扩展与复杂句式扩展的文本情感分析方法 | |
| US10748164B2 (en) | Analyzing sentiment in product reviews | |
| CN108491377B (zh) | 一种基于多维度信息融合的电商产品综合评分方法 | |
| CN108694647B (zh) | 一种商户推荐理由的挖掘方法及装置,电子设备 | |
| CN103425635B (zh) | 一种答案推荐方法和装置 | |
| CN108388660B (zh) | 一种改进的电商产品痛点分析方法 | |
| CN108269125B (zh) | 评论信息质量评估方法及系统、评论信息处理方法及系统 | |
| CN110134792B (zh) | 文本识别方法、装置、电子设备以及存储介质 | |
| CN111260437A (zh) | 一种基于商品方面级情感挖掘和模糊决策的产品推荐方法 | |
| CN105183833A (zh) | 一种基于用户模型的微博文本推荐方法及其推荐装置 | |
| KR101540683B1 (ko) | 감정어의 극성을 분류하는 방법 및 서버 | |
| CN110706028A (zh) | 基于属性特征的商品评价情感分析系统 | |
| CN110598219A (zh) | 一种面向豆瓣网电影评论的情感分析方法 | |
| CN109298796B (zh) | 一种词联想方法及装置 | |
| CN106294744A (zh) | 兴趣识别方法及系统 | |
| KR101652433B1 (ko) | Sns 문서에서 추출된 토픽을 기반으로 파악된 감정에 따른 개인화 광고 제공 방법 | |
| CN107944911A (zh) | 一种基于文本分析的推荐系统的推荐方法 | |
| CN103279504B (zh) | 一种基于歧义消解的搜索方法及装置 | |
| CN110110225A (zh) | 基于用户行为数据分析的在线教育推荐模型及构建方法 | |
| WO2023159766A1 (zh) | 餐饮数据分析方法、装置、电子设备及存储介质 | |
| Al-Sheikh et al. | Social media mining for assessing brand popularity | |
| CN105955957A (zh) | 一种商家总体评论中方面评分的确定方法及装置 | |
| CN108733652B (zh) | 基于机器学习的影评情感倾向性分析的测试方法 | |
| Zhang et al. | A novel approach to recommender system based on aspect-level sentiment analysis |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18917718 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18917718 Country of ref document: EP Kind code of ref document: A1 |
