Detailed Description
In order that those skilled in the art will better understand the present application, a technical solution in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in which it is apparent that the described embodiments are only some embodiments of the present application, not all embodiments. All other embodiments, which can be made by those skilled in the art based on the embodiments of the present application without making any inventive effort, shall fall within the scope of the present application.
It should be noted that the terms "first," "second," and the like in the description and the claims of the present application and the above figures are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the data so used may be interchanged where appropriate such that the embodiments of the application described herein may be implemented in sequences other than those illustrated or otherwise described herein. Furthermore, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, system, article, or apparatus that comprises a list of steps or elements is not necessarily limited to those steps or elements expressly listed but may include other steps or elements not expressly listed or inherent to such process, method, article, or apparatus.
The technical scheme in the embodiment of the application can follow legal rules in the implementation process, and when the operation is executed according to the technical scheme in the embodiment, the used data can not relate to user privacy, and the safety of the data is ensured while the operation process is ensured to be a compliance method. In addition, when the above embodiments of the present application are applied to a specific product or technology, user approval or consent needs to be obtained, and the collection, use and processing of relevant data needs to comply with relevant regulations and standards of the relevant country or region.
According to an aspect of an embodiment of the present application, there is provided a method for distributing media resources. As an alternative embodiment, the above-mentioned method for distributing media resources may be applied, but not limited to, to the application scenario shown in fig. 1. In an application scenario as shown in fig. 1, the target terminal 102 may be, but is not limited to being, in communication with the server 106 via the network 104, and the server 106 may be, but is not limited to being, performing operations on the database 108, such as, for example, write data operations or read data operations. The target terminal 102 may include, but is not limited to, a man-machine interaction screen, a processor, and a memory. The man-machine interaction screen may be, but not limited to, a media resource screen distributed by the technical scheme of the present application, etc., for displaying on the target terminal 102. The processor may be, but is not limited to being, configured to perform a corresponding operation in response to the man-machine interaction operation, or generate a corresponding instruction and send the generated instruction to the server 106. The memory is used for storing related processing data, such as a data set to be processed, a weight value, filtered data and the like.
Optionally, in this embodiment, the target terminal may be a terminal configured with a target client, and may include, but is not limited to, at least one of a Mobile phone (such as an Android Mobile phone, an iOS Mobile phone, etc.), a notebook computer, a tablet computer, a palm computer, a MID (Mobile INTERNET DEVICES, a Mobile internet device), a PAD, a desktop computer, a smart tv, etc. The target client may be a video client, an instant messaging client, a browser client, an educational client, and the like. Such networks may include, but are not limited to, wired networks including local area networks, metropolitan area networks, and wide area networks, wireless networks including bluetooth, WIFI, and other networks that enable wireless communications. The server may be a single server, a server cluster composed of a plurality of servers, or a cloud server.
The technical scheme of the application can be widely applied to intelligent media resource self-adaptive delivery scenes based on deep learning, and specific examples of the intelligent media resource self-adaptive delivery scenes are given below:
(1) Under the condition of possessing a plurality of overseas platform accounts, the technical scheme of the application can intelligently match the most suitable video content by analyzing the historical data and user behaviors of each account, thereby realizing personalized and differentiated delivery of the multi-account content, ensuring the consistency of the content style, avoiding homogenization competition and effectively improving the market influence and user viscosity of the whole account matrix.
(2) The intelligent recommendation and optimization of the content can dynamically adjust the content recommendation strategy according to the user preference, the watching habit and the algorithm rule of the platform aiming at a specific media platform, such as Netflix International edition, hulu and the like, ensure that the pushed video content can meet the requirement of a target audience to the maximum extent, improve the watching rate and the retention rate, and simultaneously provide precious feedback for a content creator, and assist content creation and optimization.
(3) In the advertisement activity positioning, the advertisement materials with the most attractive and conversion potential can be intelligently selected by analyzing the video characteristics of the target market and the bid in the advertisement putting or brand activity, so that the accurate putting and the maximization of the effect are realized, the effect of half effort can be achieved no matter the brand awareness is improved or the sales of specific products is promoted, and the advantage is particularly obvious on various globalization advertisement platforms.
In conclusion, the technical scheme of the application not only can improve the operation efficiency and the intelligentization level of the media resource system on overseas platforms, but also provides powerful technical support for content individuation and cultural propagation, and has wide practical value.
As can be seen from the description in the foregoing embodiments, the technical problem of low accuracy easily occurs in the process of distributing the media resources by adopting the conventional method, and in order to solve the problem, in an embodiment of the present application, a method for distributing the media resources is provided, and fig. 2 is a flowchart of a method for distributing the media resources according to an embodiment of the present application, where the flowchart includes the following steps S202 to S208.
It should be noted that, the method for distributing the media resources shown in step S202 to step S208 may be performed by, but not limited to, an electronic device, which may be, but not limited to, a target terminal or a server as shown in fig. 1.
Step S202, an account feature vector of a target account is obtained, wherein the account feature vector is used for representing style features of media resources published by the target account on an external media resource platform;
Step S204, fusing the account feature vector and the media resource feature vector of the candidate media resource to obtain a fused feature vector;
step S206, obtaining a recommended resource list by inputting the fusion feature vector into a neural network, wherein the recommended resource list is obtained by sequencing the candidate media resources according to the scores of the recommended behaviors;
And step S208, distributing target media resources based on the recommended resource list, and reducing the difference between the accumulated rewards of the target media resources and the scores of the recommended behaviors in a given state by optimizing the neural network.
Before explaining the technical solution in this embodiment, first, the technical terms related to the present application will be briefly described.
The media asset management system can be understood as a media asset management system, and comprises a core database module, a network storage module and a digital processing module, wherein the core database module, the network storage module and the digital processing module are configured to perform integrated management of full life cycle on audio and video, graphics and other forms of digital resources generated by media institutions, enterprises and other entity institutions. The technical aim of the media asset system is to realize the persistent storage, the structured efficient retrieval, the effective utilization and the value increment of the media asset by utilizing a digital processing flow, a standardized metadata framework and an intelligent analysis algorithm, technically support a downstream content production flow, a copyright management and protection mechanism and face the automatic distribution requirement of a multi-target platform.
DQN (Deep Q-Network) is a Deep reinforcement learning algorithm, and the decision problem in a high-dimensional state space is solved by combining a Deep neural Network with a Q-learning framework. As an extension of Q-learning, DQN employs function approximation to replace the traditional Q table, overcoming the dimension disaster problem. Q-learning is a value-based method in reinforcement learning, and the core Q function Q (S, a) is defined as the expected cumulative return (expected cumulative reward) that an agent (agent) can obtain after performing action a e A in state S e S. The environment will feed back the instant prize r according to action a and iteratively update the Q value by bellman equation to approach the optimal strategy.
Account feature vector of the target account, wherein the account feature vector of the target account on an external media information platform (such as YouTube, tikTok and the like) is a series of values extracted after the comprehensive analysis of the released video by a deep learning method (such as a 3D-CNN network), and the values comprise data of the visual style, the content type, the audience preference and other dimensionalities of the video and are used for accurately representing the style features and the video content attributes of the account.
The candidate media resource refers to a video list to be recommended in the media resource system, and the feature vector of each video is extracted through a 3D-CNN network, so that the features of visual elements, subject content, duration, format and the like of the video are covered, and the comprehensive properties of the video can be reflected.
And the fusion feature vector is a new vector formed by combining the account feature vector of the target account and the feature vector of each candidate media resource in a linear connection mode and is used for representing the association degree and the suitability between the account and the video.
Neural network in the embodiment of the application, the structure of the neural network can be, but is not limited to, a DQN network, and the neural network is combined with a linear layer (LL layer) and an attention layer (AL layer) for predicting and optimizing the scores of recommended behaviors in a given state so as to realize the intelligent sequencing of recommended resources and the optimization of a delivery strategy.
And the recommended resource list is generated by the media resource system after sequencing candidate media resources based on the recommendation behavior scores predicted by the DQN, and the list ranks videos from high to low according to the scores so as to guide the optimal media resource recommendation and distribution.
Cumulative rewards-cumulative rewards are the sum of rewards obtained by an agent (a media asset system) in a series of states and actions, which reflect the operational effects of recommendation and delivery strategies in long-term operation, such as vermicelli growth, play volume improvement, user interaction increase, etc.
For easy understanding, the method for dividing the media resources will be described briefly by taking video content as media resources as an example, with reference to the overall architecture diagram shown in fig. 8.
Step 1, obtaining a target account A and a video list which can be issued from a media resource system (which can be understood as a media system platform);
The specific process is as follows:
step 1.1, obtaining a large number of video lists which are available for being released to an external media platform by the media system platform.
And step 1.2, the user i authorizes the account number of the external media resource platform to the media resource system, and the system acquires all videos p e of the account number on the external media resource platform.
And 1.3, simply screening out a video list p i which can be issued to the external media platform according to the internal limit of the media system and the related specification of the external media platform.
And 2, extracting the characteristics of the obtained video list (forming candidate videos) which can be released to the external media resource platform and the video p e of which the account number is released on the external media resource platform, and carrying out fusion processing on the extracted characteristic vectors.
In this process, the extraction of vector features for video p e into a normalized floating point array list [ f e ] of 1024 dimensions in length may be accomplished, but is not limited to, using a video feature vector extraction technique, convolutional neural network (3D-CNN). The following will explain the implementation process of extracting the video feature vector by using the 3D-CNN in combination with the specific embodiment.
And 3, uploading and distributing the video.
And inputting the feature vectors subjected to the fusion processing into the DQN network, and selecting the video with the highest adaptation degree from the candidate videos to upload according to the decision of the DQN algorithm by the system.
And 4, data such as play quantity, interactive feedback and the like generated on the external media resource platform by the uploaded video are collected by the media resource system and fed back to the DQN algorithm. Through the stepwise evaluation strategy, these data are structured for optimizing the decision model of the DQN, forming a closed loop for effect acquisition and learning.
And 5, dynamic training and strategy optimization.
The DQN system can continuously adjust a strategy model of the DQN system based on real-time feedback of the social media platform, so that dynamic optimization of video delivery decisions is realized. The Q value is updated in the process, so that the performance of the model in practice is closer to the requirements of users and the platform rules of an external media resource platform, the matching degree of the style characteristics of the video to be distributed and the style characteristics of the video published through the target account is improved, and meanwhile, the viscosity of the users is improved.
It should be noted that, step ① shown in fig. 8 indicates that the developer account corresponding to the social media platform corresponding to the application is provided to the distribution platform, step ② indicates that the development application is configured according to the social media platform requirement specification, and step ③ indicates that the user account a grants part of the rights (for example, the publishing rights and the viewing account basic information rights) of the user account a on the social media platform to the development application through a protocol.
Step ④ shows that account a can choose to upload its approved video onto the social media platform. After the account A is configured with an automatic uploading strategy, n videos with the recommended values being the front can be automatically uploaded to a social media platform. Step ⑤ shows that the development application can also collect the playing condition of the video on the social media platform, and the summary induction transfer process is omitted through the step-type evaluation strategy, and the result is provided for Q-Network learning.
The overall architecture diagram depicted in fig. 8 illustrates the complete flow from media asset management, external platform account authorization, to content matching decision making, video upload distribution, up to effect feedback and policy optimization. Through the processing flow, the intelligent release problem of the medium resource system on the overseas platform is solved, and the continuous evolution and optimization of the strategy are realized through the deep reinforcement learning technology. The system design also fully considers the technical advancement and the application practicability, ensures the tight combination of the media resource management and the intelligent decision, and provides technical support for the internal culture output.
It should be noted that, assuming that the target account a exists in the foregoing media resource system, the system first needs to extract style features from the video that has been published on the external media resource platform by the account, which includes analyzing the key frames of the video by using the 3D-CNN network, extracting multidimensional features such as color distribution, motion pattern, content category, and the like, and converting the multidimensional features into a feature vector with a fixed length. For example, keyframe analysis of a video may result in a 1024-dimensional feature vector, where each dimension corresponds to a different visual or content feature, respectively.
Then, candidate media resources which can be used for being released to an external media resource platform in a media resource system platform (an internal platform) are obtained, and feature extraction is carried out on the candidate media resources to obtain a list containing each video feature vector. Then, the system linearly connects the feature vector of the target account A with the feature vector of each candidate video in the list to form a fusion feature vector, and a fusion feature vector with the dimension of 2048 is obtained.
By inputting the fusion feature vector into the DQN network, a score of the recommended action executed for distributing each candidate media resource is obtained, wherein the score is used for representing the matching degree between the candidate media resource and the published historical media resource of the target account.
After the target media asset distribution is completed, iterative optimization of the DQN network can be performed, but is not limited to, to achieve continuous learning and adaptation improvement of the network structure or model, specific reasons include:
(1) Reinforcement learning is a dynamic process whose core idea is to learn the optimal strategy from instant rewards obtained and potential rewards in the future by interacting with the environment, with the agent constantly trying different actions. In this process, the agent (here the DQN model) optimizes its decision by continually accumulating experience. Thus, even videos that have been recommended and released, their performance in the user (e.g., number of views, comments, praise, etc.) can be used as feedback to evaluate the quality of decision before the model and adjust the policy accordingly.
(2) The external environment changes, such as user preference, platform rules, trending, etc., of the external media asset platform. By collecting and analyzing the effect feedback after each video delivery, the changes can be captured in time, so that the model can be quickly adapted to and continue to make high-quality decisions. The dynamic updating mechanism ensures that the DQN can keep up with the rapid change of the outside, and the timeliness and the accuracy of the decision are maintained.
(3) Fine-tuning of network weights when recommended videos are delivered and feedback is received, such feedback information (e.g., popularity of the videos, viewing behavior of the viewer, etc.) is converted into jackpots for updating weight parameters in the DQN network structure. Such updating is typically based on so-called time-series differential (Temporal Difference, TD) learning, where the TD error reflects the gap between the model predicted Q value and the actual rewards obtained. By minimizing the TD error, the model can gradually correct its assessment of different states and action combinations, thereby predicting future rewards more accurately and making better recommended decisions.
(4) The robustness of the model is improved, namely the robustness of the model can be enhanced through continuous training and optimization, so that the model can make a relatively stable decision when facing various uncertain factors and abnormal conditions.
By adopting the mode, the adaptation degree and the potential influence between the candidate media resources and the published historical media resources of the target account can be learned by acquiring the style characteristics of the target account and fusing the media resource characteristic vectors of the candidate media resources, and the recommendation resource list with higher adaptation degree is screened out by inputting the fused characteristic vectors of the candidate media resources and the published historical media resources of the target account into the neural network, so that the accuracy of media resource distribution is improved. Meanwhile, after the media resource distribution is completed, the neural network is continuously optimized by taking the score of the minimum accumulated rewards and the recommended behaviors of the media resource as a training mode of a training target, so that the accumulated rewards which can be acquired after the distribution is completed can be more accurately predicted, the decision-making capability of the neural network is improved, and the technical effect of improving the accuracy of the media resource distribution result is realized.
As an optional example, the fusing the account feature vector and the media resource feature vector of the candidate media resource to obtain a fused feature vector includes:
Acquiring a historical media resource list of media resources published by the target account on the external media resource platform;
extracting the characteristics of the media resources in the history media resource list to obtain the account characteristic vector;
And obtaining the fusion feature vector by linearly connecting the account feature vector and the media resource feature vector of the candidate media resource.
As can be seen from the description of the above embodiment, by obtaining the account feature vector corresponding to the target account and the media resource feature vector corresponding to the candidate video that can be released, the associated feature vector (which can be understood as the fusion feature vector) of the target account and the video is obtained through linear connection.
In this embodiment, the extraction and processing of the feature vector may be completed by using a convolutional neural network (3D-CNN) extraction technique, which is specifically implemented as follows:
s11, video preprocessing and frame sampling;
as shown in fig. 3, the video is first preprocessed and the key frames are extracted.
(1) Consecutive frame truncation-dividing the video into fixed length segments (10 frames), the step size can be set to 2 frames to avoid redundancy.
(2) Resolution unification-scaling to a uniform size (64 x 64 pixels) per frame, reducing the computational effort.
S12, generating multichannel information;
In this process, it is implemented mainly depending on different types of derived images processed according to specific properties of source video (published video), wherein the derived images include, but are not limited to, gray scale images, gradient images, and light flow images as shown in fig. 4. The following describes the embodiments in detail.
And S13, carrying out data enhancement and normalization processing on the derivative image to obtain the account feature vector.
And extracting the characteristics of the videos in the optional video list (namely the candidate media resources) by adopting the same video characteristic vector extraction algorithm to obtain the media resource characteristic vectors of the candidate media resources.
And obtaining a fusion feature vector by simply and linearly connecting the account feature vector and the media resource feature vector of the candidate media resource. In the case of a normalized floating point array with a dimension of 1024 for account feature vectors, the dimension of the fused feature vector after fusion processing is 2048, e.g., [0.34, -0.48, ], 0.79.
By means of deep analysis of media resources published on an external media platform and feature fusion of features of the media resources and features of candidate media resources published by a current internal platform through a target account, a referential data basis is provided for a decision making process of a media resource management system, and the system is ensured to be more accurate and efficient in video delivery decisions of the external media resource platform. Not only is invalid delivery reduced, but also establishment of user viscosity and deepening of cultural acceptance are accelerated, and accurate matching between media resources and target audiences is promoted.
As an optional example, the obtaining the account feature vector by extracting features of the media resources in the historical media resource list includes:
Preprocessing the media resources in the history media resource list to obtain preprocessed media resources;
Performing image processing on the preprocessed media resources based on specific attributes to obtain derivative images of different types, wherein the derivative images comprise visual features and description information which highlight the media resources in the history media resource list in different dimensions;
And carrying out data enhancement and normalization processing on the feature vector of the derivative image to obtain the account feature vector.
Wherein the preprocessing includes, but is not limited to, continuous frame truncation and resolution unification in the above-described embodiments, and the like.
The image processing is performed on the preprocessed media resources based on the specific attribute, and the obtaining of the derivative images of different types includes generating the gray scale map data stream, the gradient map data stream and the optical flow map data stream shown in fig. 4 through additional calculation besides the original RGB frames of the source video.
The gray level map mainly retains brightness information, and compresses RGB three-channel information into single-channel brightness values. The calculation formula is gray=0.299×r+0.587×g+0.114×b, R is image information of red channel, G is image information of green channel, and B is image information of blue channel.
The gradient map (x/y direction) captures mainly edge variations, describing intensity and direction of brightness variations of each pixel in the image for edge detection and feature extraction. The image is intended to be generated into a gray scale map, which is obtained by convolution with a gradient operator.
The light flow graph (x/y direction) mainly describes inter-frame motion vectors and describes motion vectors of pixels between successive frames, and can be obtained by solving a motion equation by adopting a Lucas-Kanade (sparse optical flow) formula, but is not limited to.
The gray map data stream is calculated by calculating the weighted sum of red, green and blue channels, so that the brightness information of the video is reserved, the data dimension is simplified, the gradient map data stream is generated mainly by calculating the brightness change of pixels in the x and y directions, the edge and texture characteristics in the image are captured, richer details are provided for subsequent analysis, and the optical flow map data stream is generated by tracking the motion vectors of pixels between continuous frames, so that the dynamic change of the video is facilitated to be understood, and the generation of the optical flow map data stream is very important for identifying the motion and scene conversion in the video.
The account feature vector is obtained by performing data enhancement and normalization processing on the feature vector of the derived image, where the data enhancement process may be shown in fig. 5, and is specifically as follows:
S21, geometric transformation enhancement;
The same transformation parameters as the original image, such as the rotation angle θ, the translation amounts Δx, Δy, are applied to the dataflow graph, in particular by rotation and translation.
S22, calculating disturbance of a motion mode;
Motion amplitude scaling, namely multiplying the whole optical flow field by a random coefficient alpha epsilon [0.8,1.2], and simulating a speed motion difference scaled_flow=flow alpha.
And (3) motion direction offset, namely adding random noise E-N (0,0.1) to a direction component for acting to enhance the adaptability of the model to motion speed and noise.
S23, enhancing the time sequence.
First, frame interpolation/deletion, namely randomly skipping or copying frames in a video sequence, generating discontinuous frame pairs, and forcing a model to learn large displacement motion.
Second, timing skew, applying nonlinear stretching/compression in the time dimension (interpolation function t' =t+βsin (t)) to successive multiframe optical streams, simulates variable motion.
The normalization process is as follows:
S31, amplitude normalization;
Global normalization mapping the optical flow magnitude to [ -1,1] or [0,1].
Local normalization-computing local mean and variance based on image blocks (Patches), to accommodate motion non-uniformities.
S32, direction normalization;
And the direction periodic processing, namely mapping the rotation angle to [ -pi, pi ], so as to avoid the problem of discontinuity between 0 degree and 360 degrees.
S33, sub-channel normalization.
Based on the preprocessed media resources, the system generates multi-dimensional derived images which can highlight visual features of the video in terms of brightness, edges, motion and the like, and more comprehensive description information is provided for the deep learning model. The system generates a final account feature vector by carrying out data enhancement and normalization processing on the feature vector of the derivative image. The account feature vector not only contains the static and dynamic visual information of the video, but also improves the diversity and adaptability of the features through data enhancement, so that the model can more accurately understand and match the style preference of the target account. Through the series of processing, high-quality input is provided for the DQN model, the accuracy and efficiency of intelligent delivery decision are ensured, and the intelligent operation level and content transmission effect of the media resource system are greatly improved.
As an optional example, the obtaining the recommended resource list by inputting the fused feature vector into a neural network includes:
Predicting a score of each recommended action when distributing the candidate media resources to the external media resource platform through the target account by inputting the fusion feature vector into the neural network;
And sequencing the candidate media resources based on the scores of the recommendation behaviors to obtain the recommendation resource list.
After the fusion feature vector obtained by fusing the account feature vector and the media resource feature vector of the candidate media resource is obtained, the fusion feature vector is input into a pre-trained deep reinforcement learning neural network, namely a DQN network. The DQN network performs complex computation and feature conversion through its internal LL layer (LINEAR LAYERS) and AL layer (intent Layers) according to the input feature vector, and finally outputs the score of each recommended behavior.
The score of the recommended behavior reflects the expected performance of the candidate media resource when the candidate media resource is distributed on the external media resource platform through the target account, namely the matching degree of the candidate media resource and the target audience, the potential watching amount, comments, praise and other interactive indexes. For example, for a segment of video of a dance show, if the audience population of the target account shows a high interest in dance-like content, the DQN network may assign a higher score to the segment of video, indicating that it is expected to achieve good distribution.
The scores output by the DQN network are used to rank candidate media assets, and the system will generate a list of recommended assets based on the scores. In the video distribution process, video or media resources with high scores are preferentially considered to be delivered so as to achieve the optimal distribution effect.
As shown in fig. 6, a DQN system architecture design is provided that models the matching problem of accounts and video content that the media asset system needs to solve as a markov decision process (Markov Decision Process) based primarily on DQN algorithms, optimizing the allocation strategy. Is implemented by the server through the python language + pytorch deep learning framework.
Step ①, obtain account A and its list of available videos from the Environment (Environment of the agent).
And ②, extracting a feature vector f e corresponding to the account A and a feature vector f pi corresponding to the video which can be released, and forming an associated feature vector f i of the account and the video through simple linear connection.
And ③, predicting the scores of the account A recommendation behaviors by inputting the associated feature vectors into the DQN network, and sorting the recommendation list according to the scores.
Step ④, according to the DQN algorithm, the Q value Q (s, a; θ) is mapped to all actions (which can be understood as recommended actions) for summarizing the effect of the recommendation list, thereby training the DQN.
And ⑤, putting the transfer process obtained in the step ④ into a training pool of the Q-Network.
And ⑥, training the data of the training pool for training of the Q-Network.
The system utilizes the DQN network to accurately predict the expected score of each recommended action based on multidimensional feature matching of the account and the video content, thereby quantifying the potential performance of the media resource on an external media resource platform. The intelligent decision process ensures the accuracy and the high efficiency of resource release, and the content highly matched with the target user group is preferentially considered through the recommended resource list generated by automatic sequencing, so that the user viscosity, the content touch rate and the cultural spreading effect are improved.
As an optional example, predicting the score of each recommended action when distributing the candidate media resource to the external media resource platform through the target account by inputting the fusion feature vector into the neural network includes:
Performing linear transformation on the fusion feature vector through a first linear layer of the neural network to obtain a transformed feature vector, wherein the neural network comprises the first linear layer, an attention layer and a second linear layer which are sequentially connected;
Capturing diversified features and dependency relations indicated by the transformed feature vectors by using the attention layer to obtain processed feature vectors, and obtaining the score of the recommended behavior by inputting the processed feature vectors into the second linear layer.
In this embodiment, a Q-Network hierarchical structure design diagram as shown in fig. 7 is provided, and in the Q-Network of this structure, the DQN system achieves good effects in calculating the Q value and learning training.
The method specifically comprises the following steps:
S41, in the first linear layer (LL layer), the input data is mapped to the new feature space by linear transformation.
Specifically, the method is realized by the following formula (1):
rFF(X)=relu(XW+b) (1)
Where X is the input fusion feature vector, W and b are learning parameters, relu is the activation function.
S42, after the data mapping is completed in the first linear layer and mapped to the new feature space, continuing to process the data in the attention layer;
The Attention layer (AL layer) adopts a Multi-Head Attention mechanism (Multi-Head Attention), and a plurality of independent Attention heads are calculated in parallel by combining Soft-Attention layer with Self-Attention layer, so that diversified features and dependency relations in an input sequence are captured from different angles.
The multi-layer combination of the LL layer and the AL layer in the Q-Network in S43, the Network hierarchy is shown in figure 7.
And S44, the second linear layer of the Q-Network outputs the Q value finally.
Wherein, the Q value represents the score corresponding to the fusion feature vector.
It should be noted that, the linear transformation of the first linear layer mainly performs weight multiplication and offset addition on the input fusion feature vector, so as to map the original feature to a new feature space, so as to more effectively capture the complex modes critical to the recommendation decision. For example, assuming the fused feature vector is a 2048-dimensional vector, the first linear layer might transform it into a 512-dimensional vector, where each new dimension represents a different combination of original features that might be more relevant to the popularity trends of the video, the user points of interest, or platform algorithm preferences.
The feature vector converted by the first linear layer enters the attention layer. The attention mechanism allows the model to focus on the part of the input vector that has a decisive influence on the score of the recommended behavior, thereby enabling the most critical information points to be extracted when processing fusion features that contain a large amount of irrelevant or redundant information. For example, a video of a traditional food product may be emphasized by the attention layer with respect to taste descriptions, cooking skills, or cultural backgrounds, ignoring secondary features such as video length or resolution, and thereby more accurately assessing the degree of matching between the video and a particular user population.
The feature vector after the attention layer processing is fed into the second linear layer of the neural network. The task of this layer is to further convert the refined features screened by the attention mechanism into a single scoring value, i.e. a score of the recommended action, for quantifying the expected performance of the candidate media resource when distributed through the target account on the external platform.
Through the improved Q-Network structure provided by the embodiment, feature space conversion, attention layer capturing and information refining are sequentially carried out on the input fusion features, and finally the score of the recommended behavior is output on the second linear layer, so that the accuracy and efficiency of media resource recommendation are improved, and a technical background is provided for efficient operation and content transmission of a media resource system on an external media resource platform.
As an optional example, the distributing the target media resources based on the recommended resource list comprises determining at least part of media resources which are ranked at the front in the recommended resource list as the target media resources, and sequentially distributing the media resources in the at least part of media resources to the external media resource platform.
By using the score of the recommended behavior obtained in the embodiment, the neural network accurately evaluates the media resource release effect, and ensures that the finally selected media resource has the highest expected value. For example, assuming that the recommended resources list contains 100 candidate videos, the top 10 videos are determined to be the most suitable target media resources for distribution through the target account because of their high score by the ordering of DQN, these top 10 videos may cover recent trending topics, types that highly match target user group preferences, or premium content that is of strong cultural appeal.
After determining the target media resources, the system will distribute these resources to the external media resource platform in turn according to the ranking in the recommended resource list. The process includes, but is not limited to, automatic uploading of the target media resource, metadata optimization, and tag addition operations to ensure efficient exposure and accurate positioning of the resource on the target platform.
The mode of distributing at least part of media resources which are ranked at the front is selected from the recommended resource list directly according to the score of the recommended behavior, so that the preferential release of the high-potential media resources is ensured. And meanwhile, the media resources are sequentially uploaded to an external media resource platform, so that the content access speed is accelerated, and the content access speed is more in accordance with the platform algorithm and user preference.
As an optional example, the foregoing distributing the target media resource based on the recommended resource list further includes:
Determining an evaluation index of each media resource in the recommended resource list, wherein the evaluation index represents instant rewards generated after each media resource is distributed to the external media resource platform;
Based on the evaluation index, adjusting the sequence of each media resource in the recommended resource list to obtain an updated recommended resource list;
determining at least part of media resources which are ranked at the front in the updated recommended resource list as the target media resources; and sequentially distributing the media resources in the at least part of media resources to the external media resource platform.
In addition to the manner of distributing the part of the media resources ranked forward in the recommended resource list obtained by ranking the candidate media resources directly according to the score of the recommended behavior in the above embodiment, the ranked media resource list may be adjusted by a step-like evaluation policy, so as to determine the target media resource.
The ladder type evaluation strategy is to map all action pairs Q values (s, a; theta) according to the DQN algorithm, and is used for summarizing the effect of the recommended resource list. s represents the state (a multi-dimensional vector containing the integrated features of account number, video content, environmental conditions and historical feedback), a represents the action (recommended behavior), and θ represents the expected value of the cumulative rewards that may be obtained after a dispensing action is performed under the given state s.
The specific implementation process is as follows:
s51, browsing the recommended video list (which can be understood as a recommended resource list) sequentially, and determining whether to recommend the issued video;
S52, after the current ordering and the score of the recommended behavior of the ith video are approved by the user, determining an evaluation index r i of the ith row of the list in the Q-Network according to the following formula (2):
The evaluation index may be used, but is not limited to, to determine the rationality of the current media asset sequence, and may also represent the acceptance of the media asset to be distributed. For example, video 1 ranked 1 has a score of 0.8, the rating index of 0.6, video 2 ranked 2 has a score of 0.6, the rating index of 0.4, video 3 ranked 3 has a score of 0.6, and the rating index of 0.2, and then the ranking of video 3 and video 2 may be considered to be reversed. That is, the earlier the video is ordered in the recommended resources list, the higher the rating it gets, and when the recommendation does not get approval for the video, the rating r i is 0.
S53, constructing a state transition process of the Q-Network based on the information, wherein the successful state transition process is (S i,ai,ri,Si+1), and the transition failure pair transition process is (S i,ai,0,Si+1).
Wherein S i is in matrix form, the 1 st row of the matrix represents the feature vector of the 1 st video in the recommended resource list, the 2 nd row represents the feature vector of the 2 nd video, and so on. S i+1 is a feature vector after a video is distributed and the state is changed, wherein the 1 st row in S i+1 represents the feature vector after the state of the 1 st video in the recommended resource list is changed, and the 2 nd row represents the feature vector after the state of the 2 nd video is changed.
And according to the evaluation index, adjusting the sequence of each media resource in the current recommended resource list, determining a target media resource according to the adjusted media resource list, and distributing the target media resource to an external media resource platform.
By introducing an instant rewarding mechanism and a dynamic ordering strategy, the high efficiency and the intellectualization of the distribution of the media resources on an external media resource platform are realized. The system can quickly respond to the change of the user and the platform, optimize the selection and the sequencing of recommended resources, and further improve the touch effect of the content. The dynamic decision process not only reduces manual intervention, but also enhances the flexibility and accuracy of media resource distribution.
As an alternative implementation manner, the reducing the difference between the cumulative rewards of the target media resources and the scores of the recommended actions in the given state by optimizing the neural network includes:
A state transition process after distributing the target media resource is constructed, wherein the state transition process is used for describing a state after executing a distribution action according to the given state and the score of the recommended action and obtaining rewards and updated states after completing the distribution action;
Storing the state transition process to an experience playback pool;
randomly extracting a batch of empirical data from the empirical playback pool and training the neural network based on the batch of empirical data, wherein the batch of empirical data includes at least a portion of the data during the state transition.
And combining the state transition process constructed in the step S53, putting the state transition process into a training pool of the Q-Network, and using the data of the training pool for training of the Q-Network.
This is because the system builds a state transition process after each distribution of a media asset, which details the start-stop state of the distribution behavior, the score of the recommended behavior, the immediate rewards obtained after the execution of the distribution action, and the new state. The instant rewards may be based on the index of the video watching times, praise numbers and comment numbers, and the new state includes the latest interaction data of the target account and the feedback of the audience.
Each time a media resource is distributed, the state changes, and a mode of a state transition process is constructed based on the change, which comprises the following steps:
sequentially acquiring each media resource from the target media resource as a current media resource;
And after the current media resource is distributed, updating the media resource feature vectors of the candidate media resources in the current state to obtain feature vectors after the state change, wherein one row of the media resource feature vectors of the candidate media resources represents the feature vector of one media resource.
The feature vector of each media resource in the recommended resource list in the current state is S i, and after the ith video is distributed, the feature vector is updated to the value of each element in S i, so as to obtain a feature vector S i+1 after the state change.
The state transition process is stored in an empirical playback pool for subsequent model training. The storage of the empirical playback pool is not limited to data of a single distribution action, but also includes a series of consecutive state transition sequences of distribution actions, providing a rich data set for learning and optimizing strategies for the DQN algorithm.
The system randomly extracts batches of empirical data from the empirical playback pool, which are used to train neural Network models, particularly Q-networks in DQN, by learning and updating weight parameters so that the models can better predict the score of future actions while optimizing recommendation strategies. For example, the system may draw a sample from the pool containing 1000 state transition processes, perform training of the neural network, each training aimed at reducing the gap between the predicted value and the actual instant prize, and improve the decision accuracy of the model.
The system periodically and randomly extracts batch experience data from the experience playback pool, and iteratively trains the neural network model through back propagation and weight updating. The training process is the key of model iterative learning, and enables the system to continuously adjust and optimize the recommendation strategy based on the feedback of the historical distribution behavior, so that the accuracy and effect of decision making are improved, the accurate throwing and efficient propagation of media resources on an external platform are finally realized, the intelligent decision making capability of the system is enhanced, the viscosity and cultural identity of a user are improved, meanwhile, the operation cost is reduced, and the distribution efficiency is improved.
By constructing an experience playback pool, the deep reinforcement learning training is implemented, and the intelligent optimization and dynamic decision of the media resource management system are realized. The system can continuously adjust the recommendation strategy based on the historical distribution data and the instant feedback, and improves the touch effect and the user interaction level of the content on an external media resource platform, so that the accuracy and the efficiency of content transmission are enhanced.
As shown in fig. 8, for the video approved by the target account, the video is uploaded through the distribution platform, so as to upload the video to the external social media platform, which specifically includes:
s61, applying for a developer account corresponding to the corresponding social media platform and providing the developer account to the distribution platform;
s62, configuring development application according to social media platform requirement specifications;
s63, the user account A grants the user account A to the development application through a protocol on part of rights (such as publishing rights, account viewing and basic information rights) of the social media platform;
s64, the account A can select to upload the approved video to the social media platform;
after the account A is configured with an automatic uploading strategy, n videos with the recommended values being the front can be automatically uploaded to an external media resource platform.
S65, developing application, and collecting the playing condition of the video on the social media platform, wherein the playing condition is provided for Q-Network learning through a step type evaluation strategy omitting summary and induction transfer process, and the evaluation index r' i is determined, and the state transfer process is that (S i,ai,r'i,Si+1) is provided for a training pool for training.
As can be seen from the description of the above embodiments, in the embodiments of the present application, feature vectors are extracted through 3D-CNN, and the candidate media resources in a recommended resource list with a higher degree of matching with the style features of media resources issued by a target account are determined by means of deep reinforcement learning decision, deep reinforcement learning DQN decision video matching, external media resource platform authorization delivery and post-delivery acquisition effect dynamic training DQN neural network, so as to ensure that the degree of matching between the media resources distributed to the external media resource platform and the style features of the account is higher, and simultaneously, the DQN policy model is repeatedly trained by combining with external media resource platform feedback, thereby continuously improving model decision capability.
It should be noted that, for simplicity of description, the foregoing method embodiments are all described as a series of acts, but it should be understood by those skilled in the art that the present application is not limited by the order of acts described, as some steps may be performed in other orders or concurrently in accordance with the present application. Further, those skilled in the art will also appreciate that the embodiments described in the specification are all preferred embodiments, and that the acts and modules referred to are not necessarily required for the present application.
According to still another aspect of the embodiment of the present application, there is also provided a media resource distribution apparatus as shown in fig. 9, the apparatus including:
a first obtaining unit 902, configured to obtain an account feature vector of a target account, where the account feature vector is used to characterize style features of media resources published by the target account on an external media resource platform;
a first processing unit 904, configured to perform fusion processing on the account feature vector and a media resource feature vector of a candidate media resource to obtain a fusion feature vector;
A second processing unit 906, configured to obtain a recommended resource list by inputting the fused feature vector into a neural network, where the recommended resource list is obtained by sorting the candidate media resources according to scores of recommended behaviors;
A third processing unit 908 for distributing a target media asset based on the list of recommended assets and reducing the difference between the cumulative rewards of the target media asset and the score of the recommended behavior for a given state by optimizing the neural network.
Optionally, the first processing unit 904 includes:
The first acquisition module is used for acquiring a historical media resource list of media resources published by the target account on the external media resource platform;
the feature extraction module is used for extracting features of the media resources in the history media resource list to obtain the account feature vector;
and the first processing module is used for obtaining the fusion feature vector by linearly connecting the account feature vector and the media resource feature vector of the candidate media resource.
Optionally, the feature extraction module includes:
the preprocessing sub-module is used for preprocessing the media resources in the history media resource list to obtain preprocessed media resources;
The first processing sub-module is used for carrying out image processing on the preprocessed media resources based on specific attributes to obtain derivative images of different types, wherein the derivative images comprise visual features and description information which highlight the media resources in the history media resource list in different dimensions;
and the second processing sub-module is used for carrying out data enhancement and normalization processing on the feature vector of the derivative image to obtain the account feature vector.
Optionally, the second processing unit 906 includes:
the prediction module is used for predicting the score of each recommended action when the candidate media resources are distributed to the external media resource platform through the target account by inputting the fusion feature vector into the neural network;
And the ranking module is used for ranking the candidate media resources based on the scores of the recommendation behaviors to obtain the recommendation resource list.
Optionally, the prediction module includes:
The third processing sub-module is used for carrying out linear transformation on the fusion feature vector through a first linear layer of the neural network to obtain a transformed feature vector, wherein the neural network comprises the first linear layer, an attention layer and a second linear layer which are sequentially connected;
The capturing submodule is used for capturing diversified features and dependency relations indicated by the transformed feature vectors by utilizing the attention layer to obtain processed feature vectors;
And a fourth processing sub-module, configured to obtain a score of the recommended behavior by inputting the processed feature vector into the second linear layer.
Optionally, the third processing unit 908 includes:
a second processing module, configured to determine at least a part of the media resources ranked first in the recommended resource list as the target media resources;
and the first distribution module is used for sequentially distributing the media resources in the at least part of media resources to the external media resource platform.
Optionally, the third processing unit 908 includes:
The third processing module is used for determining an evaluation index of each media resource in the recommended resource list, wherein the evaluation index represents instant rewards generated after each media resource is distributed to the external media resource platform;
the adjustment module is used for adjusting the ordering of each media resource in the recommended resource list based on the evaluation index to obtain an updated recommended resource list;
a fourth processing module, configured to determine at least a part of media resources ranked earlier in the updated recommended resource list as the target media resource;
and the second distribution module is used for sequentially distributing the media resources in the at least part of media resources to the external media resource platform.
Optionally, the third processing unit 908 includes:
A building module, configured to build a state transition process after the target media resource is distributed, where the state transition process is configured to describe executing a distribution action according to the given state and the score of the recommended behavior, and a reward obtained after the completion of the distribution action and an updated state;
The storage module is used for storing the state transition process to an experience playback pool;
And the extraction module is used for randomly extracting batch experience data from the experience playback pool and training the neural network based on the batch experience data, wherein the batch experience data comprises at least part of data in the state transition process.
Optionally, the apparatus further includes:
a second obtaining unit, configured to obtain each media resource in turn from the target media resource as a current media resource;
And the updating unit is used for updating the media resource feature vectors of the candidate media resources in the current state after the current media resources are distributed, so as to obtain feature vectors after the state change, wherein one row of the media resource feature vectors of the candidate media resources represents the feature vector of one media resource.
It should be noted that, the embodiments of the media resource distribution apparatus herein may refer to the embodiments of the media resource distribution method described above, and will not be described herein again.
According to still another aspect of the embodiment of the present application, there is also provided an electronic device for implementing the foregoing media resource distribution method, where the electronic device may be a target terminal or a server shown in fig. 1. The present embodiment is described taking the electronic device as a target terminal as an example. As shown in fig. 10, the electronic device comprises a memory 1002 and a processor 1004, the memory 1002 having stored therein a computer program, the processor 1004 being arranged to perform the steps of any of the method embodiments described above by means of the computer program.
Alternatively, the electronic device may be located in at least one of a plurality of network devices of a computer.
Alternatively, the above processor may be arranged to perform the following steps by a computer program:
s1, acquiring an account feature vector of a target account, wherein the account feature vector is used for representing style features of media resources published by the target account on an external media resource platform;
s2, fusing the account feature vector and the media resource feature vector of the candidate media resource to obtain a fused feature vector;
s3, inputting the fusion feature vector into a neural network to obtain a recommended resource list, wherein the recommended resource list is obtained by sequencing the candidate media resources according to scores of recommended behaviors;
And S4, distributing target media resources based on the recommended resource list, and reducing the difference between the accumulated rewards of the target media resources and the scores of the recommended behaviors in a given state by optimizing the neural network.
Alternatively, it will be appreciated by those of ordinary skill in the art that the configuration shown in fig. 10 is merely illustrative, and that fig. 10 is not intended to limit the configuration of the electronic device electronics described above. For example, the electronics may also include more or fewer components (e.g., network interfaces, etc.) than shown in FIG. 10, or have a different configuration than shown in FIG. 10.
The memory 1002 may be configured to store software programs and modules, such as program instructions/modules corresponding to the method and apparatus for distributing media resources in the embodiment of the present application, and the processor 1004 executes the software programs and modules stored in the memory 1002, thereby executing various functional applications and data processing, that is, implementing the method for distributing media resources described above. The memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid state memory. In some examples, the memory 1002 may further include memory located remotely from the processor 1004, which may be connected to the terminal via a network. Examples of such networks include, but are not limited to, the internet, intranets, local area networks, mobile communication networks, and combinations thereof. The memory 1002 may be used for storing, but is not limited to, a data set to be processed, a current screening radius, and screened data. As an example, as shown in fig. 10, the memory 1002 may include, but is not limited to, a first acquiring unit 902, a first processing unit 904, a second processing unit 906, a third processing unit 908, and the like in the distribution apparatus including the media resource. In addition, other module units in the media resource distribution device may be included, but are not limited to, and are not described in detail in this example.
Optionally, the transmission device 906 is used to receive or transmit data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission means 906 includes a network adapter (Network Interface Controller, NIC) that can connect to other network devices and routers via a network cable to communicate with the internet or a local area network. In one example, the transmission device 906 is a Radio Frequency (RF) module for communicating wirelessly with the internet.
The electronic device further includes a display 908 for displaying video pictures of the media assets, and a connection bus 1010 for connecting the various modular components of the electronic device.
In other embodiments, the target terminal or server may be a node in a distributed system. The distributed system may be a blockchain system, which may be a distributed system formed by the plurality of nodes connected by a network communication. The nodes may form a point-to-point network, and any type of computing device, such as a server, a target terminal, etc., may become a node in the blockchain system by joining the point-to-point network.
According to yet another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program comprising computer instructions stored in a computer readable storage medium. The computer instructions are read by a processor of a computer device from a computer readable storage medium, and executed by the processor, to cause the computer device to perform a method of distributing media resources provided in various alternative implementations of the server verification process described above, wherein the computer program is configured to perform the steps of any of the method embodiments described above when run.
Alternatively, in the present embodiment, the above-described computer-readable storage medium may be configured to store a computer program for executing the steps of:
s1, acquiring an account feature vector of a target account, wherein the account feature vector is used for representing style features of media resources published by the target account on an external media resource platform;
s2, fusing the account feature vector and the media resource feature vector of the candidate media resource to obtain a fused feature vector;
s3, inputting the fusion feature vector into a neural network to obtain a recommended resource list, wherein the recommended resource list is obtained by sequencing the candidate media resources according to scores of recommended behaviors;
And S4, distributing target media resources based on the recommended resource list, and reducing the difference between the accumulated rewards of the target media resources and the scores of the recommended behaviors in a given state by optimizing the neural network.
Alternatively, in embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program having a predetermined function and working together with other relevant parts to achieve a predetermined object, and may be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Also, a processor (or multiple processors or memories) may be used to implement one or more modules or units. Furthermore, each module or unit may be part of an overall module or unit that incorporates the functionality of the module or unit.
Alternatively, in this embodiment, it will be understood by those skilled in the art that all or part of the steps in the various methods of the above embodiments may be implemented by a program for instructing the target terminal related hardware, and the program may be stored in a computer readable storage medium, where the storage medium may include a flash disk, a Read-Only Memory (ROM), a random access Memory (Random Access Memory, RAM), a magnetic disk, or an optical disk.
The foregoing embodiment numbers of the present application are merely for the purpose of description, and do not represent the advantages or disadvantages of the embodiments. The integrated units in the above embodiments may be stored in the above-described computer-readable storage medium if implemented in the form of software functional units and sold or used as separate products. Based on such understanding, the technical solution of the present application may be embodied in essence or a part contributing to the prior art or all or part of the technical solution in the form of a software product stored in a storage medium, comprising several instructions for causing one or more computer devices (which may be personal computers, servers or network devices, etc.) to perform all or part of the steps of the method of the various embodiments of the present application.
In the foregoing embodiments of the present application, the descriptions of the embodiments are emphasized, and for a portion of this disclosure that is not described in detail in this embodiment, reference is made to the related descriptions of other embodiments. In several embodiments provided by the present application, it should be understood that the disclosed client may be implemented in other manners. The above-described embodiments of the apparatus are merely exemplary, and are merely a logical functional division, and there may be other manners of dividing the apparatus in actual implementation, for example, multiple units or components may be combined or integrated into another system, or some features may be omitted, or not performed. Alternatively, the coupling or direct coupling or communication connection shown or discussed with each other may be through some interfaces, units or modules, or may be in electrical or other forms.
The units described as separate units may or may not be physically separate, and units shown as units may or may not be physical units, may be located in one place, or may be distributed over a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, each functional unit in the embodiments of the present application may be integrated in one processing unit, or each unit may exist alone physically, or two or more units may be integrated in one unit. The integrated units may be implemented in hardware or in software functional units.
The foregoing is merely a preferred embodiment of the present application and it should be noted that modifications and adaptations to those skilled in the art may be made without departing from the principles of the present application, which are intended to be comprehended within the scope of the present application.