WO2022068543A1 - 一种多媒体内容发布的方法、装置、电子设备及存储介质 - Google Patents

一种多媒体内容发布的方法、装置、电子设备及存储介质 Download PDF

Info

Publication number
WO2022068543A1
WO2022068543A1 PCT/CN2021/117199 CN2021117199W WO2022068543A1 WO 2022068543 A1 WO2022068543 A1 WO 2022068543A1 CN 2021117199 W CN2021117199 W CN 2021117199W WO 2022068543 A1 WO2022068543 A1 WO 2022068543A1
Authority
WO
WIPO (PCT)
Prior art keywords
multimedia content
label
aging
candidate
aging label
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/117199
Other languages
English (en)
French (fr)
Inventor
苏凯
王长虎
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing ByteDance Network Technology Co Ltd
Original Assignee
Beijing ByteDance Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing ByteDance Network Technology Co Ltd filed Critical Beijing ByteDance Network Technology Co Ltd
Priority to US18/029,074 priority Critical patent/US12182195B2/en
Publication of WO2022068543A1 publication Critical patent/WO2022068543A1/zh
Anticipated expiration legal-status Critical
Priority to US18/957,626 priority patent/US12566793B2/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/783Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • G06F16/7844Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using original textual content or text extracted from visual content or transcript of audio data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/48Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/489Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using time information
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/48Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/44Browsing; Visualisation therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/45Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/7867Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using information manually generated, e.g. tags, keywords, comments, title and artist information, manually generated time, location and usage information, user ratings

Definitions

  • the present disclosure relates to the technical field of information processing, and in particular, to a method, an apparatus, an electronic device and a storage medium for publishing multimedia content.
  • the embodiments of the present disclosure provide at least a solution for publishing multimedia content, providing a candidate aging label for the publisher of multimedia content, so that the publisher can select a target aging label, and the added target aging label can be used for the user who initiates the search request. It is used as a reference when providing relevant search results to improve the accuracy of the search results.
  • an embodiment of the present disclosure provides a method for publishing multimedia content, the method comprising:
  • the method further includes:
  • the generated multimedia content release information including the target aging tag is released to the outside.
  • the acquiring at least one candidate aging label matching the multimedia content includes:
  • At least one candidate time-limited tag matching the multimedia content is acquired.
  • the determining the multimedia content to be published includes:
  • an embodiment of the present disclosure further provides a method for publishing multimedia content, the method comprising:
  • the multimedia content is distributed based on the multimedia content distribution information.
  • the method further includes:
  • the selecting at least one candidate aging label matching the multimedia content from the candidate aging label set includes:
  • At least one candidate aging label is selected from the candidate aging label set.
  • the determining the correlation between the multimedia content and each time-limited label in the candidate time-limited label set includes:
  • a correlation degree between the multimedia content and each aging label in the candidate aging label set is determined.
  • the multimedia content to be published includes the multimedia content uploaded by the first user terminal and the title information added for the multimedia content; the multimedia feature vector is extracted from the multimedia content ,include:
  • the extracted content feature vector and the text feature vector are determined as the multimedia feature vector.
  • the determining the vector correlation between the multimedia feature vector and each of the text feature vectors includes:
  • the vector correlation between the multimedia feature vector and each of the text feature vectors is determined.
  • the relevance model is trained according to the following steps:
  • the historical search term corresponding to the multimedia content search result is used as the positive type aging label of the multimedia content search result, and other multimedia content search results other than the multimedia content search result corresponding to the The historical search term is used as the negative aging label of the multimedia content search result;
  • the correlation degree to be trained is The model is trained to obtain the trained correlation model.
  • the negative aging label of each multimedia content search result is determined according to the following steps:
  • For each multimedia content search result determine other multimedia content search results that are different from the identification information of the historical search term corresponding to the multimedia content search result, and use the determined historical search term corresponding to the other multimedia content search result as the multimedia content Negative class aging labels for search results.
  • the method further includes:
  • For each aging label group calculate the word overlap between each aging label in the aging label group and other aging labels in the aging label group except the aging label; Update the aging label group according to the word overlap, and obtain the updated aging label group;
  • the time-limited label group is updated according to the calculated multiple overlapping degrees of the words to obtain an updated time-limited label group, including:
  • the aging label with the largest number of words in the aging label group is assigned to the updated aging label group;
  • the plurality of word overlap degrees include a first word overlap degree greater than a preset threshold, and include a second word overlap degree less than or equal to a preset threshold, the first word overlap degree
  • the aging label with the largest number of characters among the pointed aging labels is assigned to the updated aging label group, and the multiple aging labels pointed to by the second word overlap degree are respectively attributed to the updated aging label Group;
  • each aging label in the aging label group is respectively assigned to the updated aging label group.
  • the word overlap is determined as follows:
  • the selecting at least one candidate aging label matching the multimedia content from the candidate aging label set includes:
  • At least one candidate aging label matching the multimedia content is determined.
  • an embodiment of the present disclosure further provides an apparatus for publishing multimedia content, the apparatus comprising:
  • a content determination module used to determine the multimedia content to be published
  • a label obtaining module configured to obtain at least one candidate aging label matching the multimedia content;
  • the candidate aging label is a media content label dynamically updated based on user real-time search data;
  • a label determination module configured to determine at least one target aging label selected from the at least one candidate aging label
  • An information generating module is used for generating multimedia content release information including the target aging label.
  • an embodiment of the present disclosure further provides an apparatus for publishing multimedia content, the apparatus comprising:
  • a content acquisition module used to acquire multimedia content to be published
  • the label selection module is configured to select at least one candidate aging label matching the multimedia content from the candidate aging label set, and return the selected at least one candidate aging label to the first user terminal;
  • an information receiving module configured to receive multimedia content release information including at least one target aging label, and the target aging label belongs to the candidate aging label;
  • a content publishing module configured to publish the multimedia content based on the multimedia content publishing information.
  • embodiments of the present disclosure further provide an electronic device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the Communication between the processor and the memory is through a bus, and the machine-readable instructions are executed as described in any of the first aspect and its various embodiments, the second aspect and its various embodiments when executed by the processor.
  • an electronic device including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the Communication between the processor and the memory is through a bus, and the machine-readable instructions are executed as described in any of the first aspect and its various embodiments, the second aspect and its various embodiments when executed by the processor.
  • an embodiment of the present disclosure further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is run by an electronic device, the electronic device executes the first Aspects and various embodiments thereof, the second aspect and the steps of a method for publishing multimedia content according to any one of the various embodiments thereof.
  • the target aging label selected is also the aging label. Therefore, to a certain extent, it can provide more time-sensitive multimedia content for subsequent searches; in addition, the target aging label is further confirmed and selected by the publisher on the basis of the candidate aging label, which further improves the target aging label.
  • the accuracy of the query index of multimedia content more accurate and effective search results can be provided for users who initiate search requests, and the service quality of the search platform can be improved.
  • FIG. 1 shows a flowchart of a method for publishing multimedia content provided by Embodiment 1 of the present disclosure
  • FIG. 2( a ) shows an application schematic diagram of a method for publishing multimedia content provided by Embodiment 1 of the present disclosure
  • FIG. 2(b) shows an application schematic diagram of another method for publishing multimedia content provided by Embodiment 1 of the present disclosure
  • Fig. 2(c) shows an application schematic diagram of another method for publishing multimedia content provided by Embodiment 1 of the present disclosure
  • Embodiment 3 shows a flowchart of a method for publishing multimedia content provided by Embodiment 2 of the present disclosure
  • FIG. 4 shows a flowchart of a specific method for determining similarity in the method for publishing multimedia content provided by Embodiment 2 of the present disclosure
  • FIG. 5 shows a flowchart of a specific method for updating an aging tag set in the method for publishing multimedia content provided by Embodiment 2 of the present disclosure
  • FIG. 6 shows a schematic diagram of an apparatus for publishing multimedia content according to Embodiment 3 of the present disclosure
  • FIG. 7 shows a schematic diagram of another apparatus for publishing multimedia content provided by Embodiment 3 of the present disclosure.
  • FIG. 8 shows a schematic diagram of an electronic device according to Embodiment 4 of the present disclosure.
  • FIG. 9 shows a schematic diagram of another electronic device provided by Embodiment 4 of the present disclosure.
  • the embodiments of the present disclosure provide at least one scheme for publishing multimedia content, providing users with a function of choosing and adding time-sensitive tags independently, so as to provide users with more time-sensitive and accurate search results.
  • the electronic equipment includes, for example: terminal equipment or server or other processing equipment, the terminal equipment can be user equipment (User Equipment, UE), mobile equipment, user terminal, terminal, cellular phone, cordless phone, personal digital processing (Personal digital processing equipment) Digital Assistant, PDA), handheld devices, computing devices, in-vehicle devices, wearable devices, etc.
  • the method for publishing multimedia content may be implemented by a processor invoking computer-readable instructions stored in a memory.
  • the method for publishing multimedia content provided by the embodiment of the present disclosure will be described below by taking the execution subject as the client as an example.
  • FIG. 1 is a flowchart of a method for publishing multimedia content provided by an embodiment of the present disclosure
  • the method includes steps S101-S104, wherein:
  • the above method for publishing multimedia content is mainly suitable for application scenarios with multimedia content publishing requirements.
  • relevant tags can be provided to provide search reference for subsequent multimedia content query. If you only provide some fixed types of tags such as location and category, it cannot reflect the time-sensitive hotspot information of some multimedia content, resulting in some high-value multimedia content being ignored by users on the search side, so that some highly time-sensitive multimedia content cannot be effectively utilized. content, resulting in a waste of resources and a decrease in the quality of search services.
  • the embodiments of the present disclosure provide a solution for publishing multimedia content based on time-sensitive tags.
  • the solution provides users with the function of adding time-sensitive labels independently based on dynamically updated time-sensitive labels, so as to provide time-sensitive comparisons for multimedia content searches. Strong search results.
  • the above-mentioned multimedia content to be published may include multimedia content uploaded by users, and the multimedia content here may be pictures, videos, or other forms of multimedia content. Considering the wide application of video search, video is mostly used below. Example for specific description, in addition, in a specific application, the above-mentioned multimedia content to be published may further include title information added for the uploaded multimedia content.
  • a corresponding upload button and an information input box may be set on the publishing page of the client.
  • the multimedia content uploaded by the user is acquired, and at the same time, corresponding title information can be input in the information input box, and the title information can be represented by relevant keywords of the multimedia content.
  • title information in this embodiment of the present disclosure may be manually input by the user, or may be title information automatically parsed based on the multimedia content uploaded by the user after the user end receives the multimedia content. This example does not make specific restrictions.
  • the embodiment of the present disclosure can acquire at least one candidate aging label matching the multimedia content, so as to facilitate the user to select the target aging label directly related to the intention.
  • the candidate timeliness tags in the embodiments of the present disclosure may be media content tags that are dynamically updated based on real-time search data of users.
  • the media content tags may change dynamically with real-time search operations of users and have high timeliness.
  • real-time search data of users can be obtained from various search platforms.
  • the search platform here can be an encyclopedia search platform, a multimedia search platform, or other search platforms.
  • the real-time search data here can be a distance from the current.
  • the information obtained from the search records of the above search platforms within a recent period of time of publication may be search terms, or may be search results returned by initiating a search based on the search terms.
  • the embodiment of the present disclosure can determine an updated candidate aging label based on the analysis result of the user's real-time search data, that is, once the access data is determined within a certain period of time
  • the corresponding candidate aging label will be updated.
  • the above-mentioned candidate aging labels may be displayed on the user end first.
  • the trigger operation of the button can be added based on the label set on the publishing page on the client side, and then the above-mentioned candidate aging labels can be displayed.
  • the method for publishing multimedia content may be, in the case of determining that the multimedia content to be published is time-sensitive content, and then obtaining a candidate time-limited label that matches the multimedia content, or may be a response time-limited label. Get operation to get candidate aging labels.
  • the above process of judging that the multimedia content to be published is time-sensitive content may be determined by the user terminal based on content attribute information and/or author attribute information corresponding to the multimedia content uploaded by the user.
  • the content attribute information here can be determined based on the analysis of the multimedia content.
  • various types of multimedia content with relatively strong timeliness can be preset, so that , when it is determined that the uploaded multimedia content belongs to the above-mentioned multimedia content type, it can be determined that the uploaded content is time-sensitive;
  • the author attribute information here can be the relevant information of the provider of the multimedia content, for example, for news authors , the possibility of publishing time-sensitive content will also be higher, so that it can be determined that time-sensitive content is uploaded based on the author's news identity.
  • a target aging label related to the user's intention may be selected based on the user's selection operation, and multimedia content publishing information including the target aging label may be generated.
  • the multimedia content publishing information including the target aging tag can be published to the outside.
  • the multimedia content publishing information can be sent to the end.
  • multimedia content release information may include multimedia content uploaded by the user, may also include title information input for the multimedia content, and in addition, may also include release time, release location and other information.
  • the information related to the user's privacy rights may be collected after obtaining the user's authorization.
  • the multimedia content corresponding to the search request can be pushed to the user based on the target aging label contained in the multimedia content. Since the target aging label here identifies the The multimedia content with high timeliness can improve the browsing volume of the multimedia content to a certain extent.
  • a corresponding publishing button may be set on the publishing page of the user terminal.
  • the target aging label may be included in response to the triggering operation for the publishing button.
  • the multimedia content release information is released.
  • the multimedia content release information in the embodiment of the present disclosure may include not only relevant multimedia content, but also other release information.
  • the tag addition information such as the source of the multimedia content may also be information related to publishing authority and publishing time, which is not specifically limited in this embodiment of the present disclosure.
  • the publishing page presented by the client includes an upload button and an information input box.
  • the AA video can be uploaded, and the title information of the AA event can also be input in the information input box.
  • the server can, based on the uploaded AA video and the input AA event, determine multiple candidate time-limited tags that match the AA video to be published, that is, AA medium risk, AA event escalation, AA secondary response, and AA related personnel.
  • the user terminal can obtain the above-mentioned candidate aging labels from the server based on the trigger operation of the label adding button included on the presented publishing page, and can display the obtained candidate aging labels in the display area corresponding to the label adding button, such as Figure 2(b).
  • a selection operation can be performed to select the target aging label that is closest to the user's intention, that is, risk in AA and AA event escalation, as shown in Figure 2(c).
  • a release button is provided on the release page, and after the release button is triggered, the above-mentioned multimedia content release information including risks in AA and AA event upgrade can be released to the server.
  • FIG. 3 is a flowchart of a method for publishing multimedia content according to Embodiment 2 of the present disclosure, the method includes steps S301-S304, wherein:
  • the embodiment of the present disclosure may rely on the correlation between the candidate time-limited label set and the multimedia content, that is, the candidate time-limited label set related to the multimedia content may be selected from the candidate time-limited label set
  • the aging label with a relatively high degree is used as the candidate aging label of the multimedia content.
  • the above-mentioned candidate aging tag set may be generated based on the media content tags dynamically updated by the user's real-time search data.
  • the update process of the media content tags please refer to the relevant description of the above-mentioned Embodiment 1, which will not be repeated here.
  • the updated media content label can be placed in the candidate aging label set, that is, as the user real-time search data is captured, the candidate aging label set is also updated accordingly. Strong sex.
  • the selected one or more candidate aging labels can be pushed to the client, so that the client can select the target aging labels that meet its own publishing intention, and can publish the target aging labels containing
  • the multimedia content publishing information of the target aging tag is sent to the server.
  • the server After receiving the multimedia content release information sent by the client, the server can release the multimedia content release information and release the multimedia content. Once the corresponding multimedia content is released, because the time-sensitive target time-sensitive tag carried in the multimedia content release information can be used as a search basis, compared with the commonly released multimedia content, the possibility of its subsequent real-time search will be greatly improved. Thereby, the exposure of the multimedia content can be enhanced.
  • the candidate aging label set in the embodiment of the present disclosure may be composed of several aging labels.
  • the multimedia content and each aging label in the candidate aging label set can be determined.
  • the correlation between the two, and one or more candidate aging labels are selected from the aging label set based on the correlation.
  • each relevancy degree may be ranked first, and then, from the ranking result, an aging label with a preset degree of relevancy may be selected as a candidate aging label.
  • the aging labels ranked in the top 10 may be selected as candidate aging labels. Label.
  • the process of calculating correlation includes the following steps:
  • the multimedia feature vector can be extracted from the multimedia content and the text feature vector can be extracted from each aging label in the aging label set, and then the multimedia feature vector and each text feature vector can be determined based on the calculation method of vector similarity. Since the vector correlation largely represents the correlation between the multimedia content and the aging label, the correlation between the multimedia content and the aging label can be determined based on the vector correlation between the vectors.
  • the above-mentioned extracted multimedia feature vector may be directly extracted from the multimedia content to be published, such as video scene information, video duration information and other features of a video, or may be extracted based on a pre-trained multimedia feature extraction model. of.
  • the multimedia content to be published in the embodiment of the present disclosure may be the multimedia content uploaded by the first user terminal, and may also be title information added for the multimedia content, therefore, the multimedia feature vector here may be extracted for the uploaded multimedia content.
  • the content feature vector of , or the text feature vector extracted for the added title information may be the multimedia feature vector of , or the text feature vector extracted for the added title information.
  • the multimedia feature extraction model here can be obtained by training a convolutional neural network (Convolutional Neural Networks, CNN), and the training of the network can be input
  • a convolutional neural network Convolutional Neural Networks, CNN
  • the relationship between multimedia content and its various dimensional attributes For example, for a video, a vector of 128-dimensional multimedia features can be obtained by training; when the extracted multimedia feature vector is the feature vector of the title information of the multimedia content,
  • the multimedia feature extraction model here can be obtained by one-hot encoding, and can also be trained by word vector encoding model Word2vec, which trains the relationship between title information and title vector, for example, for the title information of a video , a 128-dimensional feature vector can be extracted.
  • the text feature vector of the above-mentioned aging label may be obtained by encoding the aging label.
  • it may also be obtained by one-hot encoding of read heat, and may also be obtained by using Word2vec training, which is not specifically limited in this embodiment of the present disclosure.
  • Word2vec training which is not specifically limited in this embodiment of the present disclosure.
  • a 128-dimensional text feature vector can be extracted for each aging label in the aging label set.
  • the vector correlation between the vectors can be determined.
  • the vector correlation can be directly determined by the vector cosine formula in the embodiment of the present disclosure.
  • the vector correlation can also be determined based on the trained correlation model. Considering that the correlation model can excavate richer and deeper features to a certain extent, in the embodiment of the present disclosure, a trained correlation model can be used to determine the vector correlation.
  • the relevance model in the embodiment of the present disclosure may be a model related to multi-classification of labels, and in the process of model training, it is intended to select a label that is more relevant to the input multimedia content from each label.
  • tag identification IDs to represent tags
  • the tag set needs to be fixed. For example, No. 1 corresponds to the first tag in the tag set, and No. 2 corresponds to the second tag in the tag set.
  • the embodiment of the present disclosure adopts the method of text encoding to achieve the purpose of generalization. In this way, no matter how the aging labels are transformed, the aging labels can be aligned to the text space, and then the correlation learning and multimedia content can be carried out. match.
  • the text feature vector obtained by the text encoding adopted in the embodiments of the present disclosure can mine more abundant information, which will help to train the subsequent relevance model.
  • the training sample data for training the relevance model in the embodiment of the present disclosure may be determined based on the specific application scenario of the method for publishing multimedia content provided by the present disclosure, that is, the corresponding training sample data may be obtained based on the scenario application, and then the The training of the correlation model includes the following steps:
  • Step 1 Obtain each historical search term and the multimedia content search result returned by initiating a search based on each historical search term;
  • Step 2 for each multimedia content search result, use the historical search term corresponding to the multimedia content search result as the positive time-sensitive label of the multimedia content search result, and search for other multimedia contents except the multimedia content search result.
  • the historical search term corresponding to the result is used as the negative aging label of the multimedia content search result;
  • Step 3 Take each multimedia content search result, the positive time-sensitive label of the multimedia content search result and the negative time-limited label of the multimedia content search result as a set of training sample data, based on the correlations to be trained based on multiple sets of training sample data The model is trained to obtain a trained correlation model.
  • relevant historical search data such as each historical search term and the multimedia content search result returned by the search based on each historical search term can be obtained from each search platform.
  • the multimedia content search result here corresponds to its historical search term, that is, , the two are bound together based on the search relationship, so that the time-consuming and laborious problem caused by the need for manual labeling in related technologies can be avoided.
  • the search term is "BB new song”
  • the returned search results may be videos and titles related to "BB new song”.
  • the above-mentioned correlation model as a multi-classification model, for each multimedia content search result, can determine the positive class aging label and the negative class aging label of the search result, and the multimedia content search result has a higher correlation with the positive class aging label, The correlation with negative class aging labels is even lower.
  • the positive-type aging label of a multimedia content search result may be the initiating search word for obtaining the multimedia content search result, and the negative-type aging label may be the initiating search word of other multimedia content search results.
  • each multimedia content search result, the positive-type aging label of the multimedia content search result, and the negative-type aging label of the multimedia content search result can be used as a set of training sample data to perform the training of the correlation model, so that the correlation model can be trained. Get the model parameters of the correlation model. In this way, after acquiring the multimedia content to be published, the degree of correlation between the multimedia content and each time-limited label in the time-limited label set can be determined based on this model parameter.
  • each group can be The same identification information is added to the same historical search term in the training sample data.
  • the same search term has the same identification information. If the search term identification of the target multimedia content search result is the same as the search word identification of another multimedia content search result, the embodiment of the present disclosure will not identify the other multimedia content
  • the historical search words corresponding to the search results are used as negative classes of the target multimedia content search results.
  • Using the above-mentioned negative-class aging label determination scheme can avoid sampling data with the same identification information as a pseudo-negative class, improve the high discrimination ability of positive and negative classes, and further improve the accuracy of the correlation model.
  • time-limited label set in the embodiment of the present disclosure may be obtained from real-time search data analysis of users of various search platforms, it is difficult to avoid redundant time-sensitive labels.
  • This time-limited label, and another search platform corresponds to the time-limited label "B location parade", which is the redundancy of the time-limited label. If the redundant time-sensitive tags are directly pushed to users, it will not only occupy unnecessary display space, but also cause the user's experience of the publishing platform to decline to a certain extent.
  • an embodiment of the present disclosure provides a method for deduplication processing on an aging label set. As shown in FIG. 5 , the above deduplication processing is specifically implemented through the following steps:
  • the aging label set can be clustered based on the semantic similarity of each aging label in the candidate aging label set to obtain each aging label group.
  • any two aging labels in the aging label group can be calculated.
  • the word overlap can represent the possibility of redundancy between the two aging labels to a certain extent. The greater the word overlap, the greater the possibility of redundancy, and the word overlap The smaller the value, the smaller the possibility of redundancy.
  • the aging label group Based on the overlapping degree of multiple words corresponding to each aging label group, the aging label group can be updated. In this way, based on the update result of the aging label group, it can be determined The updated candidate aging label set.
  • the semantic vector may be extracted for each aging label in the candidate aging label set, and then the semantic similarity between the aging labels may be determined by calculating the similarity of the semantic vectors.
  • the time-limited labels with the same or similar semantics can be included in the same time-limited label group, so that the word overlap degree of any two time-limited labels in one time-limited label group can be determined, and the word overlap degree can be determined according to follows these steps to determine:
  • Step 1 Perform word segmentation processing on each of the two aging labels for the two aging labels whose word overlap degree is to be calculated to obtain multiple aging label words corresponding to each aging label;
  • Step 2 Perform intersection processing on multiple aging label words corresponding to the two aging labels to obtain a processed first aging label word group, and combine the multiple aging label words corresponding to the two aging labels respectively. Set processing to obtain the processed second time-limited label word group;
  • Step 3 Determine the proportion of the first time-limited label word group in the second time-limit label word group, and use the determined proportion as the word overlap between the two time-limited labels.
  • the two time-limited labels can be calculated and processed separately to obtain a plurality of time-labeled words corresponding to each time-limited label.
  • the word segmentation can be word-by-word segmentation, that is, as many words as a time-limited label include, it can be divided into several parts.
  • the embodiment of the present disclosure can also determine the words that can be segmented based on a dictionary.
  • the time-limited label “parade at location B in country A” can be divided into “country A", “location B”, " There are three time-limited label words such as “parade”, and “parade at location B” can be divided into two time-limited label words such as “location B” and “parade”.
  • the intersection processing and union processing of the aging label words can be performed, and the first aging label word group obtained by the intersection processing is stored in the The proportion of the word group of the second time-limited label obtained by the union process can be determined as the word overlap between the two time-limited labels.
  • the three time-limited label words corresponding to the time-limited label "parade at location B in country A” are "country A”, “Location B", “Parade”, and “Parade at Location B” when the two time-limited label words corresponding to the time-limited label are "Location B” and "Parade”, take the first time-limited label word obtained from the intersection result
  • the phrase is "parade at location B” (corresponding to 5 words), and the result of the union is "parade at location B in country A” (corresponding to 8 words).
  • 5/8 can be used as the words of the above two time-limited labels degree of overlap.
  • the update of the time-limited label group can be implemented by setting a preset threshold of the word overlap degree (for example, set to 0.5).
  • the aging label with the largest number of words in the aging label group is assigned to the updated aging label group. , that is, select the aging label with more abundant label information as the aging label in the updated aging label group, and the corresponding other aging labels can be deleted; If the overlapping degree of words is less than or equal to the preset threshold, then each aging label in the aging label group is assigned to the updated aging label group, that is, if the overlapping degree of each word is less than the preset threshold value To a certain extent, it indicates that the possibility of redundancy in the aging label group is small. In this case, the updated aging label group can be directly determined based on the original aging label of the aging label group.
  • the aging label corresponding to the first word overlap degree can be selected according to the above-mentioned first processing method to update the aging label group, and the aging label corresponding to the second word overlap degree can be selected according to the above-mentioned second type.
  • the processing method is to update the aging label group.
  • the method for publishing multimedia content provided by the embodiment of the present disclosure can update the aging label set correspondingly after the update of the aging label group is implemented.
  • the aging label corresponding to the current sampling moment can be obtained from the candidate aging label set according to the preset snapshot sampling frequency.
  • the tag set initiates a snapshot fetch.
  • all aging labels in the candidate aging label set captured by the current snapshot can be regarded as candidate aging labels matching the multimedia content, or all the captured aging labels can be filtered first, based on the part obtained by the filtering operation.
  • the aging label is used to determine the candidate aging label that matches the multimedia content. For example, the aging label whose label time is not more than a preset time interval from the current sampling time can be determined as the candidate aging label.
  • target aging tags corresponding to each multimedia content to be uploaded can also be stored, and searching for high-efficiency multimedia content can be implemented based on each stored target aging tag. Steps to achieve:
  • Step 1 storing the target aging label corresponding to the multimedia content
  • Step 2 in the case of receiving the search request initiated by the second user terminal, search for the target aging label that matches the search request from the stored target aging label corresponding to the multimedia content;
  • Step 3 Push the multimedia content corresponding to the found target aging tag to the second client.
  • the multimedia content corresponding to the search request may be found out based on the matching relationship between the search term carried in the search request and each stored target aging tag.
  • the matching relationship between the search word and the target aging label can be determined based on the similarity of the word vector.
  • the second user terminal here may be different from the first user terminal.
  • the second user terminal may be a multimedia content search terminal, because the target aging tag may be based on high-efficiency requirements.
  • the established time-limited label can therefore be pushed to the second client terminal with more time-sensitive multimedia content, which can meet the user's high-time search requirements and improve the service quality of the search platform.
  • the writing order of each step does not mean a strict execution order but constitutes any limitation on the implementation process, and the specific execution order of each step should be based on its function and possible Internal logic is determined.
  • an apparatus for publishing multimedia content corresponding to the method for publishing multimedia content is also provided in the embodiment of the present disclosure, because the principle of solving the problem of the apparatus in the embodiment of the present disclosure is the same as the method for publishing the multimedia content in the embodiment of the present disclosure. Similar, therefore, the implementation of the apparatus may refer to the implementation of the method, and repeated descriptions will not be repeated.
  • the apparatus includes: a content determination module 601, a tag acquisition module 602, a tag determination module 603, and an information generation module 604; wherein,
  • the label determination module 603 is configured to determine at least one selected target aging label in the at least one candidate aging label
  • the information generation module 604 is configured to generate multimedia content release information including the target aging tag.
  • the candidate aging label may be a dynamically updated aging label provided based on the real-time search data of the user.
  • the target aging label selected by the publisher is the label with higher timeliness. Therefore, to a certain extent, it can provide multimedia content with strong timeliness for subsequent searches; in addition, the target aging label is further confirmed and selected by the publisher on the basis of the candidate aging label, which further improves the query of the target aging label as multimedia content.
  • the accuracy of the index can provide more accurate and effective search results for users who initiate search requests, and improve the service quality of the search platform.
  • the above device further includes:
  • the content publishing module 605 is configured to, after generating the multimedia content publishing information including the target aging label, in response to the media content publishing request, publish the generated multimedia content publishing information including the target aging label to the outside.
  • the label obtaining module 602 is used to obtain at least one candidate aging label matching the multimedia content according to the following steps:
  • At least one candidate time-limiting tag matching the multimedia content is acquired.
  • the content determination module 601 is configured to determine the multimedia content to be published according to the following steps:
  • the multimedia content uploaded by the target user and the title information added to the multimedia content are acquired, and the multimedia content and the title information uploaded by the target user are used as the multimedia content to be published.
  • the apparatus includes: a content acquisition module 701, a label selection module 702, an information receiving module 703, and a content publishing module 704; wherein,
  • the label selection module 702 is configured to select at least one candidate aging label matching the multimedia content from the candidate aging label set, and return the selected at least one candidate aging label to the first user terminal;
  • an information receiving module 703, configured to receive multimedia content release information including at least one target aging label, and at least one target aging label belongs to a candidate aging label;
  • the content publishing module 704 is configured to publish multimedia content based on multimedia content publishing information.
  • the above device further includes:
  • the content push module 705 is used to store the target aging label corresponding to the multimedia content; in the case of receiving the search request initiated by the second user terminal, find the target aging label corresponding to the multimedia content from the stored target aging label that matches the search request. Target aging label; push the multimedia content corresponding to the found target aging label to the second client.
  • the label selection module 702 is configured to select at least one candidate aging label matching the multimedia content from the candidate aging label set according to the following steps:
  • At least one candidate aging label is selected from the candidate aging label set.
  • the label selection module 702 is configured to determine the correlation between the multimedia content and each aging label in the candidate aging label set according to the following steps:
  • the correlations between the multimedia content and each aging label in the candidate aging label set are determined.
  • the multimedia content to be published includes the multimedia content uploaded by the first user terminal and the title information added for the multimedia content; the label selection module 702 is configured to extract the multimedia feature vector from the multimedia content according to the following steps :
  • the extracted content feature vector and text feature vector are determined as multimedia feature vectors.
  • the label selection module 702 is used to determine the vector correlation between the multimedia feature vector and each text feature vector according to the following steps:
  • the vector correlation between the multimedia feature vector and each text feature vector is determined.
  • the label selection module 702 is used to train the relevance model according to the following steps:
  • the historical search term corresponding to the multimedia content search result is used as the positive type aging label of the multimedia content search result, and other multimedia content search results other than the multimedia content search result corresponding to the The historical search term is used as the negative aging label of the multimedia content search result;
  • the label selection module 702 is configured to determine the negative-type aging label of each multimedia content search result according to the following steps:
  • For each multimedia content search result determine other multimedia content search results that are different from the identification information of the historical search term corresponding to the multimedia content search result, and use the determined historical search term corresponding to the other multimedia content search result as the multimedia content search result
  • the negative class aging label For each multimedia content search result, determine other multimedia content search results that are different from the identification information of the historical search term corresponding to the multimedia content search result, and use the determined historical search term corresponding to the other multimedia content search result as the multimedia content search result The negative class aging label.
  • the above device further includes:
  • the label set updating module 706 is configured to use an aging label whose semantic similarity in the candidate aging label set is greater than a preset threshold as an aging label group; for each aging label group, calculate the relationship between each aging label in the aging label group and the aging label. The word overlap between other aging labels in the label group except the aging label; update the aging label group according to the calculated overlapping degrees of multiple words, and obtain the updated aging label group; Combine each aging label group of , to obtain the updated candidate aging label set.
  • the label set updating module 706 is used to update the aging label group according to the calculated overlapping degrees of multiple words according to the following steps to obtain the updated aging label group:
  • the aging label with the largest number of words in the aging label group is assigned to the updated aging label group;
  • the multiple word overlap degrees include a first word overlap degree greater than a preset threshold, and include a second word overlap degree less than or equal to a preset threshold
  • the aging label with the largest number of words in the aging label belongs to the updated aging label group, and the multiple aging labels pointed to by the second word overlap are respectively attributed to the updated aging label group;
  • each aging label in the aging label group is respectively assigned to the updated aging label group.
  • the label set update module 706 is configured to determine the degree of word overlap according to the following steps:
  • the label selection module 702 is configured to select at least one candidate aging label matching the multimedia content from the candidate aging label set according to the following steps:
  • At least one candidate aging label matching the multimedia content is determined.
  • An embodiment of the present disclosure also provides an electronic device, where the electronic device may be a server or a client.
  • the electronic device may be a server or a client.
  • a schematic structural diagram of the electronic device provided by the embodiment of the present disclosure includes: a processor 801 , a memory 802 , and a bus 803 .
  • the memory 802 stores machine-readable instructions executable by the processor 801 (in the apparatus for publishing multimedia content as shown in FIG. instruction), when the electronic device is running, the processor 801 communicates with the memory 802 through the bus 803, and the machine-readable instruction is executed by the processor 801 to perform the following processing:
  • the instructions executed by the processor 801 further include:
  • the generated multimedia content release information including the target time-limited label is released to the outside.
  • the instructions executed by the processor 801 to obtain at least one candidate aging tag matching the multimedia content include:
  • At least one candidate time-limiting tag matching the multimedia content is acquired.
  • the instructions executed by the processor 801 to determine the multimedia content to be published include:
  • the multimedia content uploaded by the target user and the title information added to the multimedia content are acquired, and the multimedia content and the title information uploaded by the target user are used as the multimedia content to be published.
  • a schematic structural diagram of the electronic device includes: a processor 901 , a memory 902 , and a bus 903 .
  • the memory 902 stores machine-readable instructions executable by the processor 901 (in the apparatus for publishing multimedia content as shown in FIG. instruction), when the electronic device is running, the communication between the processor 901 and the memory 902 is through the bus 903, and the machine-readable instruction is executed by the processor 901 to perform the following processing:
  • Receive multimedia content release information that includes at least one target aging label, and the at least one target aging label belongs to a candidate aging label;
  • the multimedia content is distributed based on the multimedia content distribution information.
  • the instructions executed by the processor 901 further include:
  • At least one candidate aging label matching the multimedia content is selected from the candidate aging label set, including:
  • At least one candidate aging label is selected from the candidate aging label set.
  • the instructions executed by the processor 901 above determine the correlation between the multimedia content and each time-limited label in the candidate time-limited label set, including:
  • the correlations between the multimedia content and each aging label in the candidate aging label set are determined.
  • the multimedia content to be published includes the multimedia content uploaded by the first user terminal and the title information added for the multimedia content; in the instructions executed by the processor 901, a multimedia feature vector is extracted from the multimedia content, include:
  • the extracted content feature vector and text feature vector are determined as multimedia feature vectors.
  • the instructions executed by the processor 901 to determine the vector correlation between the multimedia feature vector and each text feature vector include:
  • the vector correlation between the multimedia feature vector and each text feature vector is determined.
  • the correlation model is trained according to the following steps:
  • the historical search term corresponding to the multimedia content search result is used as the positive type aging label of the multimedia content search result, and other multimedia content search results other than the multimedia content search result corresponding to the The historical search term is used as the negative aging label of the multimedia content search result;
  • the negative-type aging label of each multimedia content search result is determined according to the following steps:
  • For each multimedia content search result determine other multimedia content search results that are different from the identification information of the historical search term corresponding to the multimedia content search result, and use the determined historical search term corresponding to the other multimedia content search result as the multimedia content search result
  • the negative class aging label For each multimedia content search result, determine other multimedia content search results that are different from the identification information of the historical search term corresponding to the multimedia content search result, and use the determined historical search term corresponding to the other multimedia content search result as the multimedia content search result The negative class aging label.
  • the instructions executed by the processor 901 further include:
  • the aging labels whose semantic similarity is greater than the preset threshold in the candidate aging label set are regarded as an aging label group;
  • For each aging label group calculate the word overlap between each aging label in the aging label group and other aging labels in the aging label group except the aging label; The overlapping degree updates the aging label group to obtain the updated aging label group;
  • the aging label group is updated according to the calculated overlapping degrees of multiple words, and an updated aging label group is obtained, including:
  • the aging label with the largest number of words in the aging label group is assigned to the updated aging label group;
  • the multiple word overlap degrees include a first word overlap degree greater than a preset threshold, and include a second word overlap degree less than or equal to a preset threshold
  • the aging label with the largest number of words in the aging label belongs to the updated aging label group, and the multiple aging labels pointed to by the second word overlap are respectively attributed to the updated aging label group;
  • each aging label in the aging label group is respectively assigned to the updated aging label group.
  • the degree of word overlap is determined according to the following steps:
  • At least one candidate aging label matching the multimedia content is selected from the candidate aging label set, including:
  • At least one candidate aging label matching the multimedia content is determined.
  • Embodiments of the present disclosure further provide a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the multimedia content described in the first and second embodiments of the foregoing method is executed The steps of the published method.
  • the storage medium may be a volatile or non-volatile computer-readable storage medium.
  • the computer program product of the method for publishing multimedia content includes a computer-readable storage medium storing program codes, and the instructions included in the program code can be used to execute the multimedia content publishing described in the above method embodiments.
  • the steps of the method reference may be made to the foregoing method embodiments, which will not be repeated here.
  • Embodiments of the present disclosure also provide a computer program, which implements any one of the methods in the foregoing embodiments when the computer program is executed by a processor.
  • the computer program product can be specifically implemented by hardware, software or a combination thereof.
  • the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (Software Development Kit, SDK), etc. Wait.
  • the units described as separate components may or may not be physically separated, and components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution in this embodiment.
  • each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
  • the functions, if implemented in the form of software functional units and sold or used as stand-alone products, may be stored in a processor-executable non-volatile computer-readable storage medium.
  • the computer software products are stored in a storage medium, including Several instructions are used to cause an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure.
  • the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disk or optical disk and other media that can store program codes .

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Multimedia (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Library & Information Science (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

本公开提供了一种多媒体内容发布的方法、装置、电子设备及存储介质,其中,该方法包括:确定待发布的多媒体内容;获取与多媒体内容匹配的至少一个候选时效标签;确定至少一个候选时效标签中被选中的至少一个目标时效标签;生成包含目标时效标签的多媒体内容发布信息。本公开为多媒体内容的发布者提供候选时效标签,以供发布者从中选择目标时效标签,添加的目标时效标签能够用于在为发起搜索请求的用户提供相关搜索结果时参考使用,以提高搜索结果的准确性。

Description

一种多媒体内容发布的方法、装置、电子设备及存储介质
相关申请的交叉引用
本申请基于申请号为202011061795.9、申请日为2020年9月30日、名称为“一种多媒体内容发布的方法、装置、电子设备及存储介质”的中国专利申请提出,并要求该中国专利申请的优先权,该中国专利申请的全部内容在此引入本申请作为参考。
技术领域
本公开涉及信息处理技术领域,具体而言,涉及一种多媒体内容发布的方法、装置、电子设备及存储介质。
背景技术
随着互联网技术的发展,出现了各种各样的自媒体社交应用程序(Application,APP),所有用户都可以在APP上上传自己的多媒体内容(如视频、图片等),并可以为上传的多媒体内容添加对应的标题后,进行发布。由于热门APP上发布的多媒体内容数量非常大,在基于用户的搜索请求为用户查找多媒体内容时很难定位到较为准确的搜索结果。
对于一些高热的实时新闻内容,很多有价值的多媒体内容在发布后就可能无法被用户及时搜索并阅读到,从而一方面导致了资源的浪费,另一方面也无法很好地满足用户的搜索需求。
发明内容
本公开实施例至少提供一种多媒体内容发布的方案,为多媒体内容的发布者提供候选时效标签,以供发布者从中选择目标时效标签,添加的目标时效标签能够用于在为发起搜索请求的用户提供相关搜索结果时参考使用,以提高搜索结果的准确性。
主要包括以下几个方面:
第一方面,本公开实施例提供了一种多媒体内容发布的方法,所述方法包括:
确定待发布的多媒体内容;
获取与所述多媒体内容匹配的至少一个候选时效标签;
确定所述至少一个候选时效标签中被选中的至少一个目标时效标签;
生成包含所述目标时效标签的多媒体内容发布信息。
在一种实施方式中,生成包含所述目标时效标签的多媒体内容发布信息之后,所述方法还包括:
响应媒体内容发布请求,将生成的包含所述目标时效标签的多媒体内容发布信息向外发布。
在一种实施方式中,所述获取与所述多媒体内容匹配的至少一个候选时效标签,包括:
响应于时效标签获取操作,获取与所述多媒体内容匹配的至少一个候选时效标签;或者,
在根据所述多媒体内容对应的内容属性信息和/或作者属性信息,确定所述多媒体内容为时效性内容后,获取与所述多媒体内容匹配的至少一个候选时效标签。
在一种实施方式中,所述确定待发布的多媒体内容,包括:
获取目标用户上传的多媒体内容,以及为所述多媒体内容添加的标题信息,将所述目标用户上传的多媒体内容以及所述标题信息作为所述待发布的多媒体内容。
第二方面,本公开实施例还提供了一种多媒体内容发布的方法,所述方法包括:
获取待发布的多媒体内容;
从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
接收包含至少一个目标时效标签的多媒体内容发布信息,所述至少一个目标时效标签属于所述候选时效标签;
基于所述多媒体内容发布信息,发布所述多媒体内容。
在一种实施方式中,所述方法还包括:
存储与所述多媒体内容对应的目标时效标签;
在接收到第二用户端发起的搜索请求的情况下,从存储的与所述多媒体内容对应的目标时效标签中查找与所述搜索请求匹配的目标时效标签;
将查找到的所述目标时效标签对应的多媒体内容推送至所述第二用户端。
在一种实施方式中,所述从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,包括:
确定所述多媒体内容与候选时效标签集中的每个时效标签之间的相关度;
基于所述相关度,从所述候选时效标签集中选取至少一个所述候选时效标签。
在一种实施方式中,所述确定所述多媒体内容与候选时效标签集中的每个时效标签之间的相关度,包括:
从所述多媒体内容中提取出多媒体特征向量,以及从所述候选时效标签集中的每个时效标签中提取出文本特征向量;
确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度;
基于确定出的每个所述向量相关度,确定所述多媒体内容与所述候选时效标签集中的每个时效标签之间的相关度。
在一种实施方式中,所述待发布的多媒体内容包括所述第一用户端上传的多媒体内容,以及为所述多媒体内容添加的标题信息;所述从所述多媒体内容中提取出多媒 体特征向量,包括:
从所述第一用户端上传的多媒体内容中提取出内容特征向量,以及,从为所述多媒体内容添加的标题信息中提取出文本特征向量;
将提取出的所述内容特征向量和所述文本特征向量,确定为所述多媒体特征向量。
在一种实施方式中,所述确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度,包括:
利用训练好的相关度模型,确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度。
在一种实施方式中,按照如下步骤训练所述相关度模型:
获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果;
针对每个多媒体内容搜索结果,将该多媒体内容搜索结果所对应的历史搜索词作为该多媒体内容搜索结果的正类时效标签,并将除该多媒体内容搜索结果之外的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签;
将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据,基于多组训练样本数据对所述待训练的相关度模型进行训练,得到所述训练好的相关度模型。
在一种实施方式中,按照如下步骤确定每个多媒体内容搜索结果的负类时效标签:
针对各组训练样本数据的同一个历史搜索词,为该历史搜索词添加同一标识信息;
针对每个多媒体内容搜索结果,确定与该多媒体内容搜索结果对应历史搜索词的标识信息不同的其它多媒体内容搜索结果,并将确定的所述其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签。
在一种实施方式中,所述方法还包括:
将所述候选时效标签集中语义相似度大于预设阈值的时效标签作为一个时效标签组;
针对每个时效标签组,计算该时效标签组中的每个时效标签与该时效标签组中除该时效标签之外的其它时效标签之间的字词重叠度;根据计算得到的多个所述字词重叠度对该时效标签组进行更新,得到更新后的时效标签组;
将更新后的各个时效标签组进行组合,得到更新后的候选时效标签集。
在一种实施方式中,所述根据计算得到的多个所述字词重叠度对该时效标签组进行更新,得到更新后的时效标签组,包括:
若多个所述字词重叠度均大于预设阈值,则将该时效标签组中字数最多的时效标签归属至所述更新后的时效标签组;
若多个所述字词重叠度中包括大于预设阈值的第一字词重叠度、且包括小于或等于预设阈值的第二字词重叠度,则将所述第一字词重叠度所指向的多个时效标签中字数最多的时效标签归属至所述更新后的时效标签组,并将所述第二字词重叠度所指向 的多个时效标签分别归属至所述更新后的时效标签组;
若多个所述字词重叠度均小于或等于预设阈值,则将该时效标签组中的各个时效标签分别归属至所述更新后的时效标签组。
在一种实施方式中,按照如下步骤确定所述字词重叠度:
针对待计算字词重叠度的两个时效标签,将所述两个时效标签中的每个所述时效标签进行字词切分处理,得到与每个时效标签对应的多个时效标签字词;
将所述两个时效标签分别对应的多个时效标签字词进行交集处理,得到处理后的第一时效标签字词组,以及将所述两个时效标签分别对应的多个时效标签字词进行并集处理,得到处理后的第二时效标签字词组;
确定所述第一时效标签字词组在所述第二时效标签字词组中的占比,将确定的所述占比作为所述两个时效标签之间的字词重叠度。
在一种实施方式中,所述从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,包括:
按照预设快照采样频率从候选时效标签集中获取与当前采样时刻对应的时效标签;
基于获取的所述时效标签,确定至少一个与所述多媒体内容匹配的候选时效标签。
第三方面,本公开实施例还提供了一种多媒体内容发布的装置,所述装置包括:
内容确定模块,用于确定待发布的多媒体内容;
标签获取模块,用于获取与所述多媒体内容匹配的至少一个候选时效标签;所述候选时效标签为基于用户实时搜索数据动态更新的媒体内容标签;
标签确定模块,用于确定所述至少一个候选时效标签中被选中的至少一个目标时效标签;
信息生成模块,用于生成包含所述目标时效标签的多媒体内容发布信息。
第四方面,本公开实施例还提供了一种多媒体内容发布的装置,所述装置包括:
内容获取模块,用于获取待发布的多媒体内容;
标签选取模块,用于从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
信息接收模块,用于接收包含至少一个目标时效标签的多媒体内容发布信息,所述目标时效标签属于所述候选时效标签;
内容发布模块,用于基于所述多媒体内容发布信息,发布所述多媒体内容。
第五方面,本公开实施例还提供了一种电子设备,包括:处理器、存储器和总线,所述存储器存储有所述处理器可执行的机器可读指令,当电子设备运行时,所述处理器与所述存储器之间通过总线通信,所述机器可读指令被所述处理器执行时执行如第一方面及其各种实施方式、第二方面及其各种实施方式任一项所述的多媒体内容发布的方法的步骤。
第六方面,本公开实施例还提供了一种计算机可读存储介质,所述计算机可读存 储介质上存储有计算机程序,所述计算机程序被电子设备运行时,所述电子设备执行如第一方面及其各种实施方式、第二方面及其各种实施方式任一项所述的多媒体内容发布的方法的步骤。
采用上述多媒体内容发布的方案,其在确定待发布的多媒体内容的情况下,可以获取与多媒体内容匹配的至少一个候选时效标签,这样,在确定至少一个候选时效标签中被选中的目标时效标签的情况下,可以生成包含所述目标时效标签的多媒体内容发布信息,这里的候选时效标签可以是基于用户实时搜索数据提供的动态更新的时效标签,这样,发布者从中选择的目标时效标签也就是时效性较高的标签,因而一定程度上可以为后续搜索提供时效性较强的多媒体内容;另外,目标时效标签是在候选时效标签的基础上由发布者进一步确认选择的,进一步提升了目标时效标签作为多媒体内容的查询索引的准确性,从而能够为发起搜索请求的用户提供更准确有效的搜索结果,提升搜索平台的服务质量。
为使本公开的上述目的、特征和优点能更明显易懂,下文特举较佳实施例,并配合所附附图,作详细说明如下。
附图说明
为了更清楚地说明本公开实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,此处的附图被并入说明书中并构成本说明书中的一部分,这些附图示出了符合本公开的实施例,并与说明书一起用于说明本公开的技术方案。应当理解,以下附图仅示出了本公开的某些实施例,因此不应被看作是对范围的限定,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他相关的附图。
图1示出了本公开实施例一所提供的一种多媒体内容发布的方法的流程图;
图2(a)示出了本公开实施例一所提供的一种多媒体内容发布的方法的应用示意图;
图2(b)示出了本公开实施例一所提供的另一种多媒体内容发布的方法的应用示意图;
图2(c)示出了本公开实施例一所提供的另一种多媒体内容发布的方法的应用示意图;
图3示出了本公开实施例二所提供的一种多媒体内容发布的方法的流程图;
图4示出了本公开实施例二所提供的多媒体内容发布的方法中,确定相似度具体方法的流程图;
图5示出了本公开实施例二所提供的多媒体内容发布的方法中,更新时效标签集具体方法的流程图;
图6示出了本公开实施例三所提供的一种多媒体内容发布的装置的示意图;
图7示出了本公开实施例三所提供的另一种多媒体内容发布的装置的示意图;
图8示出了本公开实施例四所提供的一种电子设备的示意图;
图9示出了本公开实施例四所提供的另一种电子设备的示意图。
具体实施方式
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分实施例,而不是全部的实施例。通常在此处描述和示出的本公开实施例的组件可以以各种不同的配置来布置和设计。因此,以下对本公开的实施例的详细描述并非旨在限制要求保护的本公开的范围,而是仅仅表示本公开的选定实施例。基于本公开的实施例,本领域技术人员在没有做出创造性劳动的前提下所获得的所有其他实施例,都属于本公开保护的范围。
经研究发现,由于一些热门的网站或APP上发布的多媒体内容数量很多,对于一些高热的实时新闻内容,很多有价值的多媒体内容在被发布后可能无法被用户及时搜索并阅读到,导致资源没有得到充分利用,也没有很好地满足用户的搜索需求。
基于上述研究,本公开实施例提供了至少一种多媒体内容发布的方案,为用户提供时效标签自主选择添加的功能,以便为用户提供时效性较强、较准确的搜索结果。
针对以上方案所存在的缺陷,均是发明人在经过实践并仔细研究后得出的结果,因此,上述问题的发现过程以及下文中本公开针对上述问题所提出的解决方案,都应该是发明人在本公开过程中对本公开做出的贡献。
应注意到:相似的标号和字母在下面的附图中表示类似项,因此,一旦某一项在一个附图中被定义,则在随后的附图中不需要对其进行进一步定义和解释。
为便于对本实施例进行理解,首先对本公开实施例所公开的一种多媒体内容发布的方法进行详细介绍,本公开实施例所提供的多媒体内容发布的方法的执行主体一般为具有一定计算能力的电子设备,该电子设备例如包括:终端设备或服务端或其它处理设备,终端设备可以为用户设备(User Equipment,UE)、移动设备、用户端、终端、蜂窝电话、无绳电话、个人数字处理(Personal Digital Assistant,PDA)、手持设备、计算设备、车载设备、可穿戴设备等。在一些可能的实现方式中,该多媒体内容发布的方法可以通过处理器调用存储器中存储的计算机可读指令的方式来实现。
下面以执行主体为用户端为例对本公开实施例提供的多媒体内容发布的方法加以说明。
参见图1所示,为本公开实施例提供的多媒体内容发布的方法的流程图,方法包括步骤S101~S104,其中:
S101、确定待发布的多媒体内容;
S102、获取与多媒体内容匹配的至少一个候选时效标签;
S103、确定至少一个候选时效标签中被选中的至少一个目标时效标签;
S104、生成包含目标时效标签的多媒体内容发布信息。
上述多媒体内容发布的方法主要适用于具有多媒体内容发布需求的应用场景中,在发布多媒体内容时,可以提供相关标签,为后续进行多媒体内容查询提供搜索参考。如果只是提供位置、类别等一些固定类型的标签,无法体现出一些多媒体内容的时效性热点信息,导致在搜索侧,一些高价值多媒体内容被用户忽略,从而无法有效利用好一些高时效性的多媒体内容,导致资源的浪费和搜索服务质量的降低。
为了解决上述问题,本公开实施例提供了一种基于时效标签进行多媒体内容发布的方案,该方案基于动态更新的时效标签为用户提供时效性标签自主添加功能,以便为多媒体内容搜索提供时效性较强的搜索结果。
其中,上述待发布的多媒体内容可以包括用户上传的多媒体内容,这里的多媒体内容可以是图片、还可以是视频、还可以是其它多媒体内容形式,考虑到视频搜索的广泛应用,以下多以视频为例进行具体说明,此外,在具体应用中,上述待发布的多媒体内容还可以包括为上传的多媒体内容添加的标题信息。
本公开实施例中,为了便于确定待发布的多媒体内容,可以在用户端的发布页面上设置相应的上传按钮和信息输入框,例如,可以在用户进入发布页面之后,响应针对上传按钮的触发操作,获取用户上传的多媒体内容,与此同时,还可以在信息输入框输入相应的标题信息,该标题信息可以采用多媒体内容的有关关键词来表征。
需要说明的是,本公开实施例中的标题信息可以是用户手动输入的,也可以是在用户端接收到用户上传的多媒体内容之后,基于这一多媒体内容自动解析得到的标题信息,本公开实施例对此不做具体的限制。
在确定待发布的多媒体内容的情况下,本公开实施例可以获取与多媒体内容匹配的至少一个候选时效标签,从而便于用户从中选取出与意图直接相关的目标时效标签。
本公开实施例中的候选时效标签可以是基于用户实时搜索数据动态更新的媒体内容标签,该媒体内容标签可以随着用户的实时搜索操作而动态变化,具有较高的时效性。
在具体应用中,可以从各种搜索平台获取用户实时搜索数据,这里的搜索平台可以是百科搜索平台,还可以是多媒体搜索平台,还可以是其它搜索平台,这里的实时搜索数据可以是距离当前发布时间最近的一段时间内从上述各个搜索平台的搜索记录中获取的,可以是搜索词,还可以是基于搜索词发起搜索所返回的搜索结果。
这样,在从搜索平台获取到用户实时搜索数据的情况下,本公开实施例可以基于对用户实时搜索数据的分析结果确定更新的候选时效标签,也即,一旦在一定时间内确定接入的各个搜索平台发生了搜索更新,则对应的候选时效标签将产生更新。
本公开实施例中,为了便于从候选时效标签中选取符合用户意图的目标时效标签,可以先在用户端展示上述候选时效标签。为了提升用户与发布平台的交互体验,这里可以基于用户端发布页面上设置的标签添加按钮的触发操作,再进行上述候选时效标签的展示。
需要说明的是,本公开实施例提供的多媒体内容发布的方法可以是在确定待发布 的多媒体内容为时效性内容的情况下,再获取与多媒体内容匹配的候选时效标签,还可以是响应时效标签获取操作来获取候选时效标签。
其中,上述判断待发布的多媒体内容为时效性内容的过程可以是用户端基于用户上传的多媒体内容对应的内容属性信息和/或作者属性信息确定的。这里的内容属性信息可以是基于对该多媒体内容进行解析之后所确定的,在具体应用中,可以预设多种时效性比较强的多媒体内容类型(例如军事类、生活类、娱乐类),这样,在确定上传的多媒体内容属于上述多媒体内容类型的情况下,即可以确定上传的是时效性内容;这里的作者属性信息可以是多媒体内容的提供者的相关信息,例如,针对新闻类作者而言,其发布时效性内容的可能性也会更高,这样,即可以基于该作者的新闻类身份确定上传的是时效性内容。
本公开实施例中,在确定各个候选时效标签之后,可以基于用户的选取操作选取与用户意图相关的目标时效标签,并生成包含目标时效标签的多媒体内容发布信息。这时,在响应媒体内容发布请求的前提下,可以将包含目标时效标签的多媒体内容发布信息向外发布,例如,为了便于服务端实现后续的多媒体搜索,可以将上述多媒体内容发布信息发送给服务端。
其中,上述多媒体内容发布信息可以包括用户上传的多媒体内容,还可以包括为该多媒体内容输入的标题信息,除此之外,还可以包括发布时间、发布位置等信息。
需要说明的是,有关发布位置等涉及用户隐私权限的信息,可以是在获得用户授权之后再采集的。
针对发布后的各个多媒体内容,可以在用户向服务端发起搜索请求之后,基于该多媒体内容所包含的目标时效标签向用户推送与搜索请求对应的多媒体内容,由于这里的目标时效标签所标识的是时效性较高的多媒体内容,因而一定程度上可以提升多媒体内容的浏览量。
其中,本公开实施例提供的多媒体内容发布的方法可以在用户端的发布页面上设置相应的发布按钮,例如,可以在用户选中目标时效标签之后,响应针对发布按钮的触发操作,将包含目标时效标签的多媒体内容发布信息发布出去。
本公开实施例中的多媒体内容发布信息除了可以包括有关多媒体内容,还可以包括其它发布信息,例如,可以是在用户端的发布页面上进行多媒体内容发布时的封面设置信息,还可以是添加位置、多媒体内容来源等标签添加信息,还可以是与发布权限和发布时间等相关的信息,本公开实施例对此不做具体的限制。
接下来可以下面结合图2(a)、图2(b)以及图2(c)所示的用户端界面呈现效果图对本公开实施例提供的上述多媒体内容发布的方法进行示例说明。
如图2(a)所示,用户端所呈现的发布页面上包括有上传按钮和信息输入框。用户触发上传按钮之后,可以上传AA视频,还可以在信息输入框输入AA事件这一标题信息。
这样,服务端即可以基于上传的AA视频以及输入的AA事件,确定与待发布的 AA视频匹配的多个候选时效标签,即AA中风险、AA事件升级、AA二级响应、AA相关人员。这时,用户端可以基于呈现的发布页面上包括的标签添加按钮的触发操作从服务端获取上述候选时效标签并可以将获取的候选时效标签对应显示在标签添加按钮所对应的显示区域内,如图2(b)所示。
针对用户端当前发布页面显示的各个候选时效标签,可以执行选取操作,以选取与用户意图最接近的目标时效标签,即AA中风险、AA事件升级,如图2(c)所示。
如图2(c)所示,在发布页面上设置有发布按钮,在该发布按钮被触发之后,即可以将上述包含有AA中风险、AA事件升级的多媒体内容发布信息发布给服务端。
除此之外,如图2(a)、图2(b)以及图2(c)所示,还可以设置其它多媒体内容发布信息,如地理标签添加信息、封面设置信息、发布设置等相关信息,在此不再赘述。
接下来从服务端侧,对本公开实施例提供的多媒体内容发布的方法作进一步说明。
实施例二
参见图3所示,为本公开实施例二提供的多媒体内容发布的方法的流程图,方法包括步骤S301~S304,其中:
S301、获取待发布的多媒体内容;
S302、从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
S303、接收包含至少一个目标时效标签的多媒体内容发布信息,至少一个目标时效标签属于候选时效标签;
S304、基于多媒体内容发布信息,发布多媒体内容。
上述步骤中,有关多媒体内容、多媒体内容发布信息的相关描述内容参照本公开实施例一的相关描述,在此不再赘述。
为了确定与获取的待发布的多媒体内容匹配的候选时效标签,本公开实施例可以依赖于候选时效标签集与多媒体内容之间的相关度,也即,可以从候选时效标签集中选取与多媒体内容相关度比较高的时效标签作为该多媒体内容的候选时效标签。
其中,上述候选时效标签集可以是基于用户实时搜索数据动态更新的媒体内容标签而生成的,有关媒体内容标签的更新过程参见上述实施例一的相关描述,在此不再赘述。这样,每当媒体内容标签产生更新,即可以将更新后的媒体内容标签置入候选时效标签集中,也即,随着用户实时搜索数据的捕获,候选时效标签集也随之产生更新,因而时效性较强。
在为多媒体内容选取了匹配的候选时效标签之后,即可以将选取的一个或多个候选时效标签推送给用户端,以便于用户端从中选取出符合自身发布意图的目标时效标签,并可以发布包含该目标时效标签的多媒体内容发布信息到服务端。其中,有关目标时效标签的选取以及多媒体内容发布信息的发布具体参见上述实施例一的相关描述内容,在此不再赘述。
服务端在接收到用户端发送的多媒体内容发布信息之后,即可以多媒体内容发布信息,发布多媒体内容。一旦对应的多媒体内容得以发布,由于多媒体内容发布信息中携带时效性比较强的目标时效标签可以作为搜索依据,相对一般发布的多媒体内容而言,其后续被实时搜索到的可能性将大大提升,从而可以提升多媒体内容的曝光度。
本公开实施例中的候选时效标签集可以是由若干个时效标签构成的,这样,在从用户端获取到待发布的多媒体内容之后,即可以确定该多媒体内容与候选时效标签集中的各个时效标签之间的相关度,基于相关度从时效标签集中选取一个或多个候选时效标签。
本公开实施例中,可以先将各个相关度进行排名,然后从排名结果中选取相关度在预设名次的时效标签作为候选时效标签,例如,可以选取排名在前10名的时效标签作为候选时效标签。
考虑到相关度计算对候选时效标签选取的关键作用,接下来可以对相关度的计算过程进行具体描述,如图4所示,计算相关度的过程具体包括如下步骤:
S401、从多媒体内容中提取出多媒体特征向量,以及从候选时效标签集中的每个时效标签中提取出文本特征向量;
S402、确定多媒体特征向量与每个文本特征向量之间的向量相关度;
S403、基于确定出的每个向量相关度,确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度。
这里,首先可以分别从多媒体内容中提取出多媒体特征向量以及从时效标签集中的每个时效标签中提取出文本特征向量,而后可以基于向量相似度的计算方法确定多媒体特征向量与每个文本特征向量之间的向量相关度,由于向量相关度很大程度上表征了多媒体内容与时效标签的相关度,因而可以基于向量之间的向量相关度来确定多媒体内容与时效标签之间的相关度。
其中,上述提取多媒体特征向量可以是直接从待发布的多媒体内容中提取出的,如,一个视频的视频场景信息、视频时长信息等特征,还可以是基于预先训练好的多媒体特征提取模型提取得到的。
考虑到本公开实施例中待发布的多媒体内容可以是第一用户端上传的多媒体内容,还可以是为多媒体内容添加的标题信息,因而,这里的多媒体特征向量可以是针对上传的多媒体内容提取出的内容特征向量,还可以是针对添加的标题信息提取出的文本特征向量。
其中,在提取的多媒体特征向量是有关多媒体内容的特征向量的情况下,这里的多媒体特征提取模型可以是卷积神经网络(Convolutional Neural Networks,CNN)训练得到的,该网络训练的可以是输入的多媒体内容与其各种维度属性之间的关联关系,例如,针对一个视频,可以训练得到128维的多媒体特征这个向量;在提取的多媒体特征向量是有关多媒体内容的标题信息的特征向量的情况下,这里的多媒体特征提取模型可以是采用独热one-hot编码得到,还可以采用词向量编码模型Word2vec训练 得到,训练的是标题信息与标题向量之间的关联关系,例如,针对一个视频的标题信息,可以提取出128维的特征向量。
另外,上述时效标签的文本特征向量可以是针对时效标签进行编码得到的,这里也可以采用读热one-hot编码得到,还可以采用Word2vec训练得到,本公开实施例对此不做具体的限制,例如,针对时效标签集中的每个时效标签可以提取出128维的文本特征向量。
本公开实施例中,在确定出多媒体内容的多媒体特征向量以及各个时效标签的文本特征向量之后,可以确定向量之间的向量相关度,本公开实施例一方面可以直接由向量余弦公式确定向量相关度,另一方面还可以基于训练好的相关度模型来确定向量相关度。考虑到相关度模型一定程度上可以挖掘出更为丰富、更为深层次的特征,因此,本公开实施例中可以采用训练好的相关度模型的方式来确定向量相关度。
本公开实施例中的相关度模型可以是有关标签多分类的模型,在模型训练的过程中,旨在从各个标签中选取出与输入的多媒体内容相关性更高的标签。这里,考虑到传统的分类模型一般都通过标签标识ID来表示标签,因此需要固定住标签集,例如,1号对应标签集中的第一个标签,2号对应标签集中的第二个标签。
然而,在本公开实施例的应用下,标签集(对应时效标签集)在实时变化,旧的标签标识会失效,因此标签标识将变得不再具有实时的泛化性。为了解决上述问题,本公开实施例才选用了文本编码的方式达到泛化的目的,这样,不管时效标签如何变换,均可以将时效标签对齐到文本空间,而后再和多媒体内容进行相关性学习和匹配。与此同时,相比标识编码方式,本公开实施例中所采用的文本编码所得到的文本特征向量能够挖掘出更为丰富的信息,这将有助于进行后续相关度模型的训练。
本公开实施例中训练相关度模型的训练样本数据可以是基于本公开提供的多媒体内容发布的方法的具体应用场景所确定的,也即,基于场景应用可以获取到相应的训练样本数据,进而进行相关度模型的训练,具体包括如下步骤:
步骤一、获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果;
步骤二、针对每个多媒体内容搜索结果,将该多媒体内容搜索结果所对应的历史搜索词作为该多媒体内容搜索结果的正类时效标签,并将除该多媒体内容搜索结果之外的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签;
步骤三、将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据,基于多组训练样本数据对待训练的相关度模型进行训练,得到训练好的相关度模型。
这里,首先可以从各搜索平台获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果等相关历史搜索数据,这里的多媒体内容搜索结果与其历史搜索词相对应,也即,两者基于搜索关系而绑定在一起,从而可以避免相关 技术中需要进行人工标签标注所带来的费时费力的问题。例如,搜索词是“BB新歌”,返回的搜索结果可以是“BB新歌”相关的视频和标题。
上述相关度模型作为一个多分类模型,针对每个多媒体内容搜索结果,可以确定该搜索结果的正类时效标签以及负类时效标签,该多媒体内容搜索结果与正类时效标签的相关度更高,与负类时效标签的相关度更低。一个多媒体内容搜索结果的正类时效标签可以是得到该多媒体内容搜索结果的发起搜索词,负类时效标签则可以是其它多媒体内容搜索结果的发起搜索词。
本公开实施例中,将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据可进行相关度模型的训练,从而得到相关度模型的模型参数。这样,在获取到待发布的多媒体内容之后,即可以基于这一模型参数确定该多媒体内容与时效标签集中的每个时效标签之间的相关度。
考虑到在进行多媒体内容搜索的过程中,一个历史搜索词,所搜索得到的多媒体内容搜索结果往往为多个,也即,多媒体内容搜索结果与历史搜索词存在多对一的关系,因而,在针对一个多媒体内容搜索结果进行负类时效标签的确定时,为了避免与该多媒体内容搜索结果同步被一个搜索词搜索出来的其它多媒体内容搜索结果对训练识别率的影响,这里,可以先为各组训练样本数据的同一个历史搜索词添加同一标识信息。
这样,针对每个多媒体内容搜索结果,可以确定与该多媒体内容搜索结果对应历史搜索词的标识信息不同的其它多媒体内容搜索结果,该其它多媒体内容搜索结果对应的历史搜索词即可作为该多媒体内容搜索结果的负类时效标签。
本公开实施例中,相同搜索词具有相同的标识信息,如果目标多媒体内容搜索结果的搜索词标识,和另一个多媒体内容搜索结果的搜索词标识一致,本公开实施例不会将另一个多媒体内容搜索结果所对应的历史搜索词作为目标多媒体内容搜索结果的负类。
采用上述负类时效标签确定方案,可以避免采样相同标识信息的数据作为伪负类,提升了正负类的高判别能力,进而提升相关度模型的准确率。
考虑到本公开实施例中的时效标签集可以是从各个搜索平台的用户实时搜索数据分析得到的,这难以避免产生冗余的时效标签,例如,在一个搜索平台对应有“A国家B地点游行”这一时效标签,而另一个搜索平台对应有“B地点游行”这一时效标签,这即是产生了时效标签的冗余。如果直接将存在冗余的时效标签推送给用户,不仅会占用不必要的展示位,还一定程度上会造成用户对发布平台体验度的下降。
为了解决上述问题,本公开实施例提供了一种对时效标签集进行消重处理的方法,如图5所示,上述消重处理具体通过如下步骤实现:
S501、将候选时效标签集中语义相似度大于预设阈值的时效标签作为一个时效标签组;
S502、针对每个时效标签组,计算该时效标签组中的每个时效标签与该时效标签组中除该时效标签之外的其它时效标签之间的字词重叠度;根据计算得到的多个字词重叠度对该时效标签组进行更新,得到更新后的时效标签组;
S503、将更新后的各个时效标签组进行组合,得到更新后的候选时效标签集。
这里,首先可以基于候选时效标签集中的各个时效标签的语义相似度对时效标签集进行聚类,得到各个时效标签组,这样,针对每个时效标签组,可以计算该时效标签组内的任意两个时效标签的字词重叠度,字词重叠度一定程度上可以表征两个时效标签存在冗余的可能性,字词重叠度越大,存在冗余的可能性也越大,字词重叠度越小,存在冗余的可能性也越小,基于每个时效标签组对应的多个字词重叠度即可以对这以时效标签组进行更新,这样,基于时效标签组的更新结果,可以确定更新后的候选时效标签集。
其中,本公开实施例可以先对候选时效标签集中的各个时效标签进行语义向量的提取,而后通过计算语义向量的相似度来确定时效标签之间的语义相似度。
本公开实施例中,针对语义相同或相似的时效标签可以纳入同一个时效标签组,这样,即可以确定一个时效标签组中任意两个时效标签的字词重叠度,上述字词重叠度可以按照如下步骤来确定:
步骤一、针对待计算字词重叠度的两个时效标签,将两个时效标签中的每个时效标签进行字词切分处理,得到与每个时效标签对应的多个时效标签字词;
步骤二、将两个时效标签分别对应的多个时效标签字词进行交集处理,得到处理后的第一时效标签字词组,以及将两个时效标签分别对应的多个时效标签字词进行并集处理,得到处理后的第二时效标签字词组;
步骤三、确定第一时效标签字词组在第二时效标签字词组中的占比,将确定的占比作为两个时效标签之间的字词重叠度。
这里,首先可以针对待计算字词重叠度的两个时效标签,分别计算对两个进行字词切分处理,得到每个时效标签对应的多个时效标签字词,本公开实施例中的字词切分可以是逐字切分,也即一个时效标签包括多少字,就可以对应切分几份,除此之外,本公开实施例还可以基于词典确定可切分的字词。例如,针对“A国家B地点游行”和“B地点游行”这两个时效标签而言,“A国家B地点游行”这一时效标签可以切分为“A国家”、“B地点”、“游行”等三个时效标签字词,“B地点游行”可以切分为“B地点”、“游行”等两个时效标签字词。
在确定两个时效标签中每个时效标签对应的多个时效标签字词之后,即可以进行时效标签字词的交集处理和并集处理,将交集处理所得到的第一时效标签字词组在并集处理所得到的第二时效标签字词组的占比,即可以确定为这两个时效标签之间的字词重叠度。
仍以“A国家B地点游行”和“B地点游行”这两个时效标签为例,在“A国家B地点游行”这一时效标签所对应的三个时效标签字词为“A国家”、“B地点”、 “游行”,“B地点游行”这一时效标签所对应的两个时效标签字词为“B地点”、“游行”的情况下,取交集结果得到的第一时效标签字词组为“B地点游行”(对应5个字),取并集结果得到“A国家B地点游行”(对应8个字),此时5/8即可作为上述两个时效标签的字词重叠度。
本公开实施例中,重复上述字词重叠度的计算过程,即可以确定每个时效标签组对应的多个字词重叠度。
本公开实施例可以通过设定字词重叠度的预设阈值(如设置为0.5)实现时效标签组的更新。
在具体应用中,其一、可以是在确定一个时效标签组对应的多个字词重叠度均大于预设阈值,则将该时效标签组中字数最多的时效标签归属至更新后的时效标签组,也即,选取标签信息更为丰富的时效标签作为更新后的时效标签组内的时效标签,对应其它时效标签可以进行删减操作;其二、还可以在确定一个时效标签组对应的多个字词重叠度均小于或等于预设阈值,则将该时效标签组中的各个时效标签分别归属至更新后的时效标签组,也即,在各个字词重叠度均不到预设阈值的前提下,一定程度上说明该时效标签组存在冗余的可能性较小,此时,可以直接基于该时效标签组的原始时效标签确定更新后的时效标签组。
除了以上两种情形,在一个时效标签组对应的多个字词重叠度中既存在大于预设阈值的第一字词重叠度又存在小于或等于预设阈值的第二字词重叠度的情况下,可以对第一字词重叠度对应的时效标签按照上述第一种处理方式选取字数最多的进行时效标签组的更新,还可以对第二字词重叠度对应的时效标签按照上述第二种处理方式进行时效标签组的更新。
本公开实施例提供的多媒体内容发布的方法在实现时效标签组的更新之后,可以对应更新时效标签集。
这里,为了便于为待发布的多媒体内容匹配时效性更强的候选时效标签,可以按照预设快照采样频率从候选时效标签集中获取与当前采样时刻对应的时效标签,例如,每秒即向候选时效标签集发起一次快照抓取。
这里,可以将当前快照抓取的候选时效标签集中的所有时效标签均作为与多媒体内容匹配的候选时效标签,还可以是先对抓取的所有时效标签进行筛选操作,基于筛选操作所得到的部分时效标签来确定与多媒体内容匹配的候选时效标签,例如,可以将标签时间与当前采样时间相隔不超过预设时长的时效标签确定为候选时效标签。
本公开实施例中,在发布多媒体内容的同时,还可以存储与各个待上传的多媒体内容对应的目标时效标签,基于存储的各个目标时效标签可实现有关高时效多媒体内容的搜索,具体可以通过如下步骤来实现:
步骤一、存储与多媒体内容对应的目标时效标签;
步骤二、在接收到第二用户端发起的搜索请求的情况下,从存储的与多媒体内容对应的目标时效标签中查找与搜索请求匹配的目标时效标签;
步骤三、将查找到的目标时效标签对应的多媒体内容推送至第二用户端。
这里,在接受到第二用户端发起的搜索请求的情况下,可以基于搜索请求中携带的搜索词与存储的各个目标时效标签之间的匹配关系,从中查找出与搜索请求对应的多媒体内容。
其中,有关搜索词与目标时效标签之间的匹配关系,这里可以基于词向量相似度来确定。
这里的第二用户端可以与第一用户端不同,例如,在第一用户端作为多媒体内容发布端的情况下,第二用户端可以是多媒体内容搜索端,由于目标时效标签可以是基于高时效需求所建立的时效标签,因而这里所推送给第二用户端的可以是时效性更高的多媒体内容,能够满足用户的高时效搜索需求,提升搜索平台的服务质量。
本领域技术人员可以理解,在具体实施方式的上述方法中,各步骤的撰写顺序并不意味着严格的执行顺序而对实施过程构成任何限定,各步骤的具体执行顺序应当以其功能和可能的内在逻辑确定。
基于同一发明构思,本公开实施例中还提供了与多媒体内容发布的方法对应的多媒体内容发布的装置,由于本公开实施例中的装置解决问题的原理与本公开实施例上述多媒体内容发布的方法相似,因此装置的实施可以参见方法的实施,重复之处不再赘述。
实施例三
参照图6所示,为本公开实施例提供的一种多媒体内容发布的装置示意图,装置包括:内容确定模块601、标签获取模块602、标签确定模块603和信息生成模块604;其中,
内容确定模块601,用于确定待发布的多媒体内容;
标签获取模块602,用于获取与多媒体内容匹配的至少一个候选时效标签;
标签确定模块603,用于确定至少一个候选时效标签中被选中的至少一个目标时效标签;
信息生成模块604,用于生成包含目标时效标签的多媒体内容发布信息。
本公开实施例提供的多媒体内容发布的装置中,候选时效标签可以是是基于用户实时搜索数据提供的动态更新的时效标签,这样,发布者从中选择的目标时效标签也就是时效性较高的标签,因而一定程度上可以为后续搜索提供时效性较强的多媒体内容;另外,目标时效标签是在候选时效标签的基础上由发布者进一步确认选择的,进一步提升了目标时效标签作为多媒体内容的查询索引的准确性,从而能够为发起搜索请求的用户提供更准确有效的搜索结果,提升搜索平台的服务质量。
在一种实施方式中,上述装置还包括:
内容发布模块605,用于生成包含目标时效标签的多媒体内容发布信息之后,响应媒体内容发布请求,将生成的包含目标时效标签的多媒体内容发布信息向外发布。
在一种实施方式中,标签获取模块602,用于按照以下步骤获取与多媒体内容匹 配的至少一个候选时效标签:
响应于时效标签获取操作,获取与多媒体内容匹配的至少一个候选时效标签;或者,
在根据多媒体内容对应的内容属性信息和/或作者属性信息,确定多媒体内容为时效性内容后,获取与多媒体内容匹配的至少一个候选时效标签。
在一种实施方式中,内容确定模块601,用于按照以下步骤确定待发布的多媒体内容:
获取目标用户上传的多媒体内容,以及为多媒体内容添加的标题信息,将目标用户上传的多媒体内容以及标题信息作为待发布的多媒体内容。
参照图7所示,为本公开实施例提供的另一种多媒体内容发布的装置示意图,装置包括:内容获取模块701、标签选取模块702、信息接收模块703和内容发布模块704;其中,
内容获取模块701,用于获取待发布的多媒体内容;
标签选取模块702,用于从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
信息接收模块703,用于接收包含至少一个目标时效标签的多媒体内容发布信息,至少一个目标时效标签属于候选时效标签;
内容发布模块704,用于基于多媒体内容发布信息,发布多媒体内容。
在一种实施方式中,上述装置还包括:
内容推送模块705,用于存储与多媒体内容对应的目标时效标签;在接收到第二用户端发起的搜索请求的情况下,从存储的与多媒体内容对应的目标时效标签中查找与搜索请求匹配的目标时效标签;将查找到的目标时效标签对应的多媒体内容推送至第二用户端。
在一种实施方式中,标签选取模块702,用于按照以下步骤从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签:
确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度;
基于相关度,从候选时效标签集中选取至少一个候选时效标签。
在一种实施方式中,标签选取模块702,用于按照以下步骤确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度:
从多媒体内容中提取出多媒体特征向量,以及从候选时效标签集中的每个时效标签中提取出文本特征向量;
确定多媒体特征向量与每个文本特征向量之间的向量相关度;
基于确定出的每个向量相关度,确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度。
在一种实施方式中,待发布的多媒体内容包括第一用户端上传的多媒体内容,以及为多媒体内容添加的标题信息;标签选取模块702,用于按照以下步骤从多媒体内 容中提取出多媒体特征向量:
从第一用户端上传的多媒体内容中提取出内容特征向量,以及,从为多媒体内容添加的标题信息中提取出文本特征向量;
将提取出的内容特征向量和文本特征向量,确定为多媒体特征向量。
在一种实施方式中,标签选取模块702,用于按照以下步骤确定多媒体特征向量与每个文本特征向量之间的向量相关度:
利用训练好的相关度模型,确定多媒体特征向量与每个文本特征向量之间的向量相关度。
在一种实施方式中,标签选取模块702,用于按照以下步骤训练相关度模型:
获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果;
针对每个多媒体内容搜索结果,将该多媒体内容搜索结果所对应的历史搜索词作为该多媒体内容搜索结果的正类时效标签,并将除该多媒体内容搜索结果之外的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签;
将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据,基于多组训练样本数据对待训练的相关度模型进行训练,得到训练好的相关度模型。
在一种实施方式中,标签选取模块702,用于按照如下步骤确定每个多媒体内容搜索结果的负类时效标签:
针对各组训练样本数据的同一个历史搜索词,为该历史搜索词添加同一标识信息;
针对每个多媒体内容搜索结果,确定与该多媒体内容搜索结果对应历史搜索词的标识信息不同的其它多媒体内容搜索结果,并将确定的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签。
在一种实施方式中,上述装置还包括:
标签集更新模块706,用于将候选时效标签集中语义相似度大于预设阈值的时效标签作为一个时效标签组;针对每个时效标签组,计算该时效标签组中的每个时效标签与该时效标签组中除该时效标签之外的其它时效标签之间的字词重叠度;根据计算得到的多个字词重叠度对该时效标签组进行更新,得到更新后的时效标签组;将更新后的各个时效标签组进行组合,得到更新后的候选时效标签集。
在一种实施方式中,标签集更新模块706,用于按照以下步骤根据计算得到的多个字词重叠度对该时效标签组进行更新,得到更新后的时效标签组:
若多个字词重叠度均大于预设阈值,则将该时效标签组中字数最多的时效标签归属至更新后的时效标签组;
若多个字词重叠度中包括大于预设阈值的第一字词重叠度、且包括小于或等于预设阈值的第二字词重叠度,则将第一字词重叠度所指向的多个时效标签中字数最多的时效标签归属至更新后的时效标签组,并将第二字词重叠度所指向的多个时效标签分 别归属至更新后的时效标签组;
若多个字词重叠度均小于或等于预设阈值,则将该时效标签组中的各个时效标签分别归属至更新后的时效标签组。
在一种实施方式中,标签集更新模块706,用于按照如下步骤确定字词重叠度:
针对待计算字词重叠度的两个时效标签,将两个时效标签中的每个时效标签进行字词切分处理,得到与每个时效标签对应的多个时效标签字词;
将两个时效标签分别对应的多个时效标签字词进行交集处理,得到处理后的第一时效标签字词组,以及将两个时效标签分别对应的多个时效标签字词进行并集处理,得到处理后的第二时效标签字词组;
确定第一时效标签字词组在第二时效标签字词组中的占比,将确定的占比作为两个时效标签之间的字词重叠度。
在一种实施方式中,标签选取模块702,用于按照以下步骤从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签:
按照预设快照采样频率从候选时效标签集中获取与当前采样时刻对应的时效标签;
基于获取的时效标签,确定至少一个与多媒体内容匹配的候选时效标签。
关于装置中的各模块的处理流程、以及各模块之间的交互流程的描述可以参照上述方法实施例中的相关说明,这里不再详述。
实施例四
本公开实施例还提供了一种电子设备,该电子设备可以是服务端,也可以是用户端。在以用户端作为电子设备时,如图8所示,为本公开实施例提供的电子设备的结构示意图,包括:处理器801、存储器802、和总线803。存储器802存储有处理器801可执行的机器可读指令(如图6所示多媒体内容发布的装置中,内容确定模块601、标签获取模块602、标签确定模块603和信息生成模块604所对应执行的指令),当电子设备运行时,处理器801与存储器802之间通过总线803通信,机器可读指令被处理器801执行时执行如下处理:
确定待发布的多媒体内容;
获取与多媒体内容匹配的至少一个候选时效标签;
确定至少一个候选时效标签中被选中的至少一个目标时效标签;
生成包含目标时效标签的多媒体内容发布信息。
在一种实施方式中,生成包含目标时效标签的多媒体内容发布信息之后,上述处理器801执行的指令还包括:
响应媒体内容发布请求,将生成的包含目标时效标签的多媒体内容发布信息向外发布。
在一种实施方式中,上述处理器801执行的指令中,获取与多媒体内容匹配的至少一个候选时效标签,包括:
响应于时效标签获取操作,获取与多媒体内容匹配的至少一个候选时效标签;或者,
在根据多媒体内容对应的内容属性信息和/或作者属性信息,确定多媒体内容为时效性内容后,获取与多媒体内容匹配的至少一个候选时效标签。
在一种实施方式中,上述处理器801执行的指令中,确定待发布的多媒体内容,包括:
获取目标用户上传的多媒体内容,以及为多媒体内容添加的标题信息,将目标用户上传的多媒体内容以及标题信息作为待发布的多媒体内容。
在以服务端作为电子设备时,如图9所示,为本公开实施例提供的电子设备的结构示意图,包括:处理器901、存储器902、和总线903。存储器902存储有处理器901可执行的机器可读指令(如图7所示多媒体内容发布的装置中,内容获取模块701、标签选取模块702、信息接收模块703和内容发布模块704所对应执行的指令),当电子设备运行时,处理器901与存储器902之间通过总线903通信,机器可读指令被处理器901执行时执行如下处理:
获取待发布的多媒体内容;
从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
接收包含至少一个目标时效标签的多媒体内容发布信息,至少一个目标时效标签属于候选时效标签;
基于多媒体内容发布信息,发布多媒体内容。
在一种实施方式中,上述处理器901执行的指令还包括:
存储与多媒体内容对应的目标时效标签;
在接收到第二用户端发起的搜索请求的情况下,从存储的与多媒体内容对应的目标时效标签中查找与搜索请求匹配的目标时效标签;
将查找到的目标时效标签对应的多媒体内容推送至第二用户端。
在一种实施方式中,上述处理器901执行的指令中,从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签,包括:
确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度;
基于相关度,从候选时效标签集中选取至少一个候选时效标签。
在一种实施方式中,上述处理器901执行的指令中,确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度,包括:
从多媒体内容中提取出多媒体特征向量,以及从候选时效标签集中的每个时效标签中提取出文本特征向量;
确定多媒体特征向量与每个文本特征向量之间的向量相关度;
基于确定出的每个向量相关度,确定多媒体内容与候选时效标签集中的每个时效标签之间的相关度。
在一种实施方式中,待发布的多媒体内容包括第一用户端上传的多媒体内容,以及为多媒体内容添加的标题信息;上述处理器901执行的指令中,从多媒体内容中提取出多媒体特征向量,包括:
从第一用户端上传的多媒体内容中提取出内容特征向量,以及,从为多媒体内容添加的标题信息中提取出文本特征向量;
将提取出的内容特征向量和文本特征向量,确定为多媒体特征向量。
在一种实施方式中,上述处理器901执行的指令中,确定多媒体特征向量与每个文本特征向量之间的向量相关度,包括:
利用训练好的相关度模型,确定多媒体特征向量与每个文本特征向量之间的向量相关度。
在一种实施方式中,上述处理器901执行的指令中,按照如下步骤训练相关度模型:
获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果;
针对每个多媒体内容搜索结果,将该多媒体内容搜索结果所对应的历史搜索词作为该多媒体内容搜索结果的正类时效标签,并将除该多媒体内容搜索结果之外的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签;
将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据,基于多组训练样本数据对待训练的相关度模型进行训练,得到训练好的相关度模型。
在一种实施方式中,上述处理器901执行的指令中,按照如下步骤确定每个多媒体内容搜索结果的负类时效标签:
针对各组训练样本数据的同一个历史搜索词,为该历史搜索词添加同一标识信息;
针对每个多媒体内容搜索结果,确定与该多媒体内容搜索结果对应历史搜索词的标识信息不同的其它多媒体内容搜索结果,并将确定的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签。
在一种实施方式中,上述处理器901执行的指令还包括:
将候选时效标签集中语义相似度大于预设阈值的时效标签作为一个时效标签组;
针对每个时效标签组,计算该时效标签组中的每个时效标签与该时效标签组中除该时效标签之外的其它时效标签之间的字词重叠度;根据计算得到的多个字词重叠度对该时效标签组进行更新,得到更新后的时效标签组;
将更新后的各个时效标签组进行组合,得到更新后的候选时效标签集。
在一种实施方式中,上述处理器901执行的指令中,根据计算得到的多个字词重叠度对该时效标签组进行更新,得到更新后的时效标签组,包括:
若多个字词重叠度均大于预设阈值,则将该时效标签组中字数最多的时效标签归属至更新后的时效标签组;
若多个字词重叠度中包括大于预设阈值的第一字词重叠度、且包括小于或等于预设阈值的第二字词重叠度,则将第一字词重叠度所指向的多个时效标签中字数最多的时效标签归属至更新后的时效标签组,并将第二字词重叠度所指向的多个时效标签分别归属至更新后的时效标签组;
若多个字词重叠度均小于或等于预设阈值,则将该时效标签组中的各个时效标签分别归属至更新后的时效标签组。
在一种实施方式中,上述处理器901执行的指令中,按照如下步骤确定字词重叠度:
针对待计算字词重叠度的两个时效标签,将两个时效标签中的每个时效标签进行字词切分处理,得到与每个时效标签对应的多个时效标签字词;
将两个时效标签分别对应的多个时效标签字词进行交集处理,得到处理后的第一时效标签字词组,以及将两个时效标签分别对应的多个时效标签字词进行并集处理,得到处理后的第二时效标签字词组;
确定第一时效标签字词组在第二时效标签字词组中的占比,将确定的占比作为两个时效标签之间的字词重叠度。
在一种实施方式中,上述处理器901执行的指令中,从候选时效标签集中,选取至少一个与多媒体内容匹配的候选时效标签,包括:
按照预设快照采样频率从候选时效标签集中获取与当前采样时刻对应的时效标签;
基于获取的时效标签,确定至少一个与多媒体内容匹配的候选时效标签。
上述指令的具体执行过程可以参考本公开实施例一和实施例二中的多媒体内容发布的方法的步骤,此处不再赘述。
本公开实施例还提供一种计算机可读存储介质,该计算机可读存储介质上存储有计算机程序,该计算机程序被处理器运行时执行上述方法实施例一和实施例二中所述的多媒体内容发布的方法的步骤。其中,该存储介质可以是易失性或非易失的计算机可读取存储介质。
本公开实施例所提供的多媒体内容发布的方法的计算机程序产品,包括存储了程序代码的计算机可读存储介质,所述程序代码包括的指令可用于执行上述方法实施例中所述的多媒体内容发布的方法的步骤,具体可参见上述方法实施例,在此不再赘述。
本公开实施例还提供一种计算机程序,该计算机程序被处理器执行时实现前述实施例的任意一种方法。该计算机程序产品可以具体通过硬件、软件或其结合的方式实现。在一个可选实施例中,所述计算机程序产品具体体现为计算机存储介质,在另一个可选实施例中,计算机程序产品具体体现为软件产品,例如软件开发包(Software Development Kit,SDK)等等。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统和装置的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。在 本公开所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,又例如,多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些通信接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个处理器可执行的非易失的计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台电子设备(可以是个人计算机,服务端,或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
最后应说明的是:以上所述实施例,仅为本公开的具体实施方式,用以说明本公开的技术方案,而非对其限制,本公开的保护范围并不局限于此,尽管参照前述实施例对本公开进行了详细的说明,本领域的普通技术人员应当理解:任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,其依然可以对前述实施例所记载的技术方案进行修改或可轻易想到变化,或者对其中部分技术特征进行等同替换;而这些修改、变化或者替换,并不使相应技术方案的本质脱离本公开实施例技术方案的精神和范围,都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应所述以权利要求的保护范围为准。

Claims (20)

  1. 一种多媒体内容发布的方法,其特征在于,所述方法包括:
    确定待发布的多媒体内容;
    获取与所述多媒体内容匹配的至少一个候选时效标签;
    确定所述至少一个候选时效标签中被选中的至少一个目标时效标签;
    生成包含所述目标时效标签的多媒体内容发布信息。
  2. 根据权利要求1所述的方法,其特征在于,生成包含所述目标时效标签的多媒体内容发布信息之后,所述方法还包括:
    响应媒体内容发布请求,将生成的包含所述目标时效标签的多媒体内容发布信息向外发布。
  3. 根据权利要求1或2所述的方法,其特征在于,所述获取与所述多媒体内容匹配的至少一个候选时效标签,包括:
    响应于时效标签获取操作,获取与所述多媒体内容匹配的至少一个候选时效标签;或者,
    在根据所述多媒体内容对应的内容属性信息和/或作者属性信息,确定所述多媒体内容为时效性内容后,获取与所述多媒体内容匹配的至少一个候选时效标签。
  4. 根据权利要求1所述的方法,其特征在于,所述确定待发布的多媒体内容,包括:
    获取目标用户上传的多媒体内容,以及为所述多媒体内容添加的标题信息,将所述目标用户上传的多媒体内容以及所述标题信息作为所述待发布的多媒体内容。
  5. 一种多媒体内容发布的方法,其特征在于,所述方法包括:
    获取待发布的多媒体内容;
    从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,并将选取的至少一个候选时效标签返回给第一用户端;
    接收包含至少一个目标时效标签的多媒体内容发布信息,所述至少一个目标时效标签属于所述候选时效标签;
    基于所述多媒体内容发布信息,发布所述多媒体内容。
  6. 根据权利要求5所述的方法,其特征在于,所述方法还包括:
    存储与所述多媒体内容对应的目标时效标签;
    在接收到第二用户端发起的搜索请求的情况下,从存储的与所述多媒体内容对应的目标时效标签中查找与所述搜索请求匹配的目标时效标签;
    将查找到的所述目标时效标签对应的多媒体内容推送至所述第二用户端。
  7. 根据权利要求5或6所述的方法,其特征在于,所述从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,包括:
    确定所述多媒体内容与候选时效标签集中的每个时效标签之间的相关度;
    基于所述相关度,从所述候选时效标签集中选取至少一个所述候选时效标签。
  8. 根据权利要求7所述的方法,其特征在于,所述确定所述多媒体内容与候选时效标签集中的每个时效标签之间的相关度,包括:
    从所述多媒体内容中提取出多媒体特征向量,以及从所述候选时效标签集中的每个时效标签中提取出文本特征向量;
    确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度;
    基于确定出的每个所述向量相关度,确定所述多媒体内容与所述候选时效标签集中的每个时效标签之间的相关度。
  9. 根据权利要求8所述的方法,其特征在于,所述待发布的多媒体内容包括所述第一用户端上传的多媒体内容,以及为所述多媒体内容添加的标题信息;所述从所述多媒体内容中提取出多媒体特征向量,包括:
    从所述第一用户端上传的多媒体内容中提取出内容特征向量,以及,从为所述多媒体内容添加的标题信息中提取出文本特征向量;
    将提取出的所述内容特征向量和所述文本特征向量,确定为所述多媒体特征向量。
  10. 根据权利要求8或9所述的方法,其特征在于,所述确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度,包括:
    利用训练好的相关度模型,确定所述多媒体特征向量与每个所述文本特征向量之间的向量相关度。
  11. 根据权利要求10所述的方法,其特征在于,按照如下步骤训练所述相关度模型:
    获取各个历史搜索词以及基于每个历史搜索词发起搜索所返回的多媒体内容搜索结果;
    针对每个多媒体内容搜索结果,将该多媒体内容搜索结果所对应的历史搜索词作为该多媒体内容搜索结果的正类时效标签,并将除该多媒体内容搜索结果之外的其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签;
    将每个多媒体内容搜索结果、该多媒体内容搜索结果的正类时效标签以及该多媒体内容搜索结果的负类时效标签作为一组训练样本数据,基于多组训练样本数据对所述待训练的相关度模型进行训练,得到所述训练好的相关度模型。
  12. 根据权利要求11所述的方法,其特征在于,按照如下步骤确定每个多媒体内容搜索结果的负类时效标签:
    针对各组训练样本数据的同一个历史搜索词,为该历史搜索词添加同一标识信息;
    针对每个多媒体内容搜索结果,确定与该多媒体内容搜索结果对应历史搜索词的标识信息不同的其它多媒体内容搜索结果,并将确定的所述其它多媒体内容搜索结果对应的历史搜索词作为该多媒体内容搜索结果的负类时效标签。
  13. 根据权利要求5所述的方法,其特征在于,所述方法还包括:
    将所述候选时效标签集中语义相似度大于预设阈值的时效标签作为一个时效标签组;
    针对每个时效标签组,计算该时效标签组中的每个时效标签与该时效标签组中除该时效标签之外的其它时效标签之间的字词重叠度;根据计算得到的多个所述字词重叠度对该时效标签组进行更新,得到更新后的时效标签组;
    将更新后的各个时效标签组进行组合,得到更新后的候选时效标签集。
  14. 根据权利要求13所述的方法,其特征在于,所述根据计算得到的多个所述字词重叠度对该时效标签组进行更新,得到更新后的时效标签组,包括:
    若多个所述字词重叠度均大于预设阈值,则将该时效标签组中字数最多的时效标签归属至所述更新后的时效标签组;
    若多个所述字词重叠度中包括大于预设阈值的第一字词重叠度、且包括小于或等于预设阈值的第二字词重叠度,则将所述第一字词重叠度所指向的多个时效标签中字数最多的时效标签归属至所述更新后的时效标签组,并将所述第二字词重叠度所指向的多个时效标签分别归属至所述更新后的时效标签组;
    若多个所述字词重叠度均小于或等于预设阈值,则将该时效标签组中的各个时效标签分别归属至所述更新后的时效标签组。
  15. 根据权利要求13所述的方法,其特征在于,按照如下步骤确定所述字词重叠度:
    针对待计算字词重叠度的两个时效标签,将所述两个时效标签中的每个所述时效标签进行字词切分处理,得到与每个时效标签对应的多个时效标签字词;
    将所述两个时效标签分别对应的多个时效标签字词进行交集处理,得到处理后的第一时效标签字词组,以及将所述两个时效标签分别对应的多个时效标签字词进行并集处理,得到处理后的第二时效标签字词组;
    确定所述第一时效标签字词组在所述第二时效标签字词组中的占比,将确定的所述占比作为所述两个时效标签之间的字词重叠度。
  16. 根据权利要求5所述的方法,其特征在于,所述从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,包括:
    按照预设快照采样频率从候选时效标签集中获取与当前采样时刻对应的时效标签;
    基于获取的所述时效标签,确定至少一个与所述多媒体内容匹配的候选时效标签。
  17. 一种多媒体内容发布的装置,其特征在于,所述装置包括:
    内容确定模块,用于确定待发布的多媒体内容;
    标签获取模块,用于获取与所述多媒体内容匹配的至少一个候选时效标签;
    标签确定模块,用于确定所述至少一个候选时效标签中被选中的至少一个目标时效标签;
    信息生成模块,用于生成包含所述目标时效标签的多媒体内容发布信息。
  18. 一种多媒体内容发布的装置,其特征在于,所述装置包括:
    内容获取模块,用于获取待发布的多媒体内容;
    标签选取模块,用于从候选时效标签集中,选取至少一个与所述多媒体内容匹配的候选时效标签,并将选取的所述至少一个候选时效标签返回给第一用户端;
    信息接收模块,用于接收包含至少一个目标时效标签的多媒体内容发布信息,所述至少一个目标时效标签属于所述候选时效标签;
    内容发布模块,用于基于所述多媒体内容发布信息,发布所述多媒体内容。
  19. 一种电子设备,其特征在于,包括:处理器、存储器和总线,所述存储器存储有所述处理器可执行的机器可读指令,当电子设备运行时,所述处理器与所述存储器之间通过总线通信,所述机器可读指令被所述处理器执行时执行如权利要求1至16任一项所述的多媒体内容发布的方法的步骤。
  20. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质上存储有计算机程序,所述计算机程序被电子设备运行时,所述电子设备执行如权利要求1至16任一项所述的多媒体内容发布的方法的步骤。
PCT/CN2021/117199 2020-09-30 2021-09-08 一种多媒体内容发布的方法、装置、电子设备及存储介质 Ceased WO2022068543A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US18/029,074 US12182195B2 (en) 2020-09-30 2021-09-08 Multimedia content publishing method and apparatus, and electronic device and storage medium
US18/957,626 US12566793B2 (en) 2020-09-30 2024-11-22 Multimedia content publishing method and apparatus, and electronic device and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202011061795.9A CN112199526B (zh) 2020-09-30 2020-09-30 一种多媒体内容发布的方法、装置、电子设备及存储介质
CN202011061795.9 2020-09-30

Related Child Applications (2)

Application Number Title Priority Date Filing Date
US18/029,074 A-371-Of-International US12182195B2 (en) 2020-09-30 2021-09-08 Multimedia content publishing method and apparatus, and electronic device and storage medium
US18/957,626 Continuation US12566793B2 (en) 2020-09-30 2024-11-22 Multimedia content publishing method and apparatus, and electronic device and storage medium

Publications (1)

Publication Number Publication Date
WO2022068543A1 true WO2022068543A1 (zh) 2022-04-07

Family

ID=74013555

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/117199 Ceased WO2022068543A1 (zh) 2020-09-30 2021-09-08 一种多媒体内容发布的方法、装置、电子设备及存储介质

Country Status (3)

Country Link
US (2) US12182195B2 (zh)
CN (1) CN112199526B (zh)
WO (1) WO2022068543A1 (zh)

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112199526B (zh) 2020-09-30 2023-03-14 抖音视界有限公司 一种多媒体内容发布的方法、装置、电子设备及存储介质
CN112530016B (zh) 2020-10-30 2022-11-11 北京字跳网络技术有限公司 一种道具吸附的方法、装置、设备及存储介质
CN112637517B (zh) 2020-11-16 2022-10-28 北京字节跳动网络技术有限公司 视频处理方法、装置、电子设备及存储介质
CN112396679B (zh) 2020-11-20 2022-09-13 北京字节跳动网络技术有限公司 虚拟对象显示方法及装置、电子设备、介质
CN112529997B (zh) 2020-12-28 2022-08-09 北京字跳网络技术有限公司 烟花视觉效果的生成方法、视频生成方法、电子设备
CN112906553B (zh) 2021-02-09 2022-05-17 北京字跳网络技术有限公司 图像处理方法、装置、设备及介质
CN113360657B (zh) * 2021-06-30 2023-10-24 安徽商信政通信息技术股份有限公司 一种公文智能分发办理方法、装置及计算机设备
CN113420224B (zh) * 2021-07-19 2025-03-25 抖音视界有限公司 一种信息处理的方法、装置以及计算机存储介质
CN113641247A (zh) 2021-08-31 2021-11-12 北京字跳网络技术有限公司 视线角度调整方法、装置、电子设备及存储介质
CN114168763B (zh) * 2021-11-12 2025-11-21 北京达佳互联信息技术有限公司 一种多媒体资源召回方法、装置、设备及存储介质
CN117009679A (zh) * 2023-06-25 2023-11-07 雪球(北京)技术开发有限公司 信息发布方法、装置、存储介质和电子设备
CN116913460B (zh) * 2023-09-13 2023-12-29 福州市迈凯威信息技术有限公司 一种药械及检验试剂的营销业务合规性判断分析方法
CN118331455A (zh) * 2024-04-19 2024-07-12 北京字跳网络技术有限公司 发布作品的方法、装置、设备和存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20130346588A1 (en) * 2012-06-20 2013-12-26 Google Inc. Status Aware Media Play
CN108829893A (zh) * 2018-06-29 2018-11-16 北京百度网讯科技有限公司 确定视频标签的方法、装置、存储介质和终端设备
CN111353071A (zh) * 2018-12-05 2020-06-30 阿里巴巴集团控股有限公司 标签生成方法及装置
CN112199526A (zh) * 2020-09-30 2021-01-08 北京字节跳动网络技术有限公司 一种多媒体内容发布的方法、装置、电子设备及存储介质

Family Cites Families (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5832499A (en) * 1996-07-10 1998-11-03 Survivors Of The Shoah Visual History Foundation Digital library system
US8090659B2 (en) * 2001-09-18 2012-01-03 Music Public Broadcasting, Inc. Method and system for providing location-obscured media delivery
US9092434B2 (en) * 2007-01-23 2015-07-28 Symantec Corporation Systems and methods for tagging emails by discussions
US8694533B2 (en) * 2010-05-19 2014-04-08 Google Inc. Presenting mobile content based on programming context
US9158775B1 (en) * 2010-12-18 2015-10-13 Google Inc. Scoring stream items in real time
US20130091421A1 (en) * 2011-10-11 2013-04-11 International Business Machines Corporation Time relevance within a soft copy document or media object
US10331661B2 (en) * 2013-10-23 2019-06-25 At&T Intellectual Property I, L.P. Video content search using captioning data
CN107577753A (zh) * 2017-08-31 2018-01-12 维沃移动通信有限公司 一种多媒体数据的播放方法及移动终端
CN108304761A (zh) * 2017-09-25 2018-07-20 腾讯科技(深圳)有限公司 文本检测方法、装置、存储介质和计算机设备
CN108416279B (zh) * 2018-02-26 2022-04-19 北京阿博茨科技有限公司 文档图像中的表格解析方法及装置
US20190340255A1 (en) * 2018-05-07 2019-11-07 Apple Inc. Digital asset search techniques
CN111611492A (zh) * 2020-05-26 2020-09-01 北京字节跳动网络技术有限公司 一种触发搜索的方法、装置、电子设备及存储介质
CN111949864B (zh) * 2020-08-10 2022-02-25 北京字节跳动网络技术有限公司 一种搜索方法、装置、电子设备及存储介质

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20130346588A1 (en) * 2012-06-20 2013-12-26 Google Inc. Status Aware Media Play
CN108829893A (zh) * 2018-06-29 2018-11-16 北京百度网讯科技有限公司 确定视频标签的方法、装置、存储介质和终端设备
CN111353071A (zh) * 2018-12-05 2020-06-30 阿里巴巴集团控股有限公司 标签生成方法及装置
CN112199526A (zh) * 2020-09-30 2021-01-08 北京字节跳动网络技术有限公司 一种多媒体内容发布的方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
US12182195B2 (en) 2024-12-31
CN112199526A (zh) 2021-01-08
US20230367804A1 (en) 2023-11-16
US12566793B2 (en) 2026-03-03
US20250086224A1 (en) 2025-03-13
CN112199526B (zh) 2023-03-14

Similar Documents

Publication Publication Date Title
CN112199526B (zh) 一种多媒体内容发布的方法、装置、电子设备及存储介质
CN108733766B (zh) 一种数据查询方法、装置和可读介质
US8577882B2 (en) Method and system for searching multilingual documents
CN101639857B (zh) 构建知识问答分享平台的方法、装置及系统
CN108701161B (zh) 为搜索查询提供图像
CN111708943B (zh) 一种搜索结果展示方法、装置和用于搜索结果展示的装置
WO2022033321A1 (zh) 一种搜索方法、装置、电子设备及存储介质
CN112035687B (zh) 一种多媒体内容发布的方法、装置、电子设备及存储介质
US9798776B2 (en) Systems and methods for parsing search queries
US20130013591A1 (en) Image re-rank based on image annotations
CN103678576A (zh) 基于动态语义分析的全文检索系统
CN102402593A (zh) 对于搜索查询输入的多模态方式
CN110543484A (zh) 提示词的推荐方法及装置、存储介质和处理器
CN105956053A (zh) 一种基于网络信息的搜索方法及装置
JP2017220204A (ja) 検索クエリに応答してホワイトリストとブラックリストを使用し画像とコンテンツをマッチングする方法及びシステム
CN116361428A (zh) 一种问答召回方法、装置和存储介质
CN104142955A (zh) 一种推荐学习课程的方法和终端
US20160154885A1 (en) Method for searching a database
CN107992563B (zh) 一种用户浏览内容的推荐方法及系统
CN110851560B (zh) 信息检索方法、装置及设备
CN110704654A (zh) 一种图片搜索方法和装置
WO2020191706A1 (zh) 主动学习自动图像标注系统及方法
CN111753861A (zh) 主动学习自动图像标注系统及方法
CN113609372A (zh) 搜索方法、装置、服务器、介质及产品
CN117033797A (zh) 一种基于联邦学习的信息检索方法、装置、设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21874204

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 060723)

122 Ep: pct application non-entry in european phase

Ref document number: 21874204

Country of ref document: EP

Kind code of ref document: A1