WO2021160141A1 - 视频处理方法、装置、可读介质和电子设备 - Google Patents
视频处理方法、装置、可读介质和电子设备 Download PDFInfo
- Publication number
- WO2021160141A1 WO2021160141A1 PCT/CN2021/076409 CN2021076409W WO2021160141A1 WO 2021160141 A1 WO2021160141 A1 WO 2021160141A1 CN 2021076409 W CN2021076409 W CN 2021076409W WO 2021160141 A1 WO2021160141 A1 WO 2021160141A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target video
- historical
- video segment
- quality score
- video
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/222—Studio circuitry; Studio devices; Studio equipment
- H04N5/262—Studio circuits, e.g. for mixing, switching-over, change of character of image, other special effects ; Cameras specially adapted for the electronic generation of special effects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/47205—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for manipulating displayed content, e.g. interacting with MPEG-4 objects, editing locally
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/10—Indexing; Addressing; Timing or synchronising; Measuring tape travel
- G11B27/19—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier
- G11B27/28—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier by using information signals recorded by the same method as the main recording
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/10—Indexing; Addressing; Timing or synchronising; Measuring tape travel
- G11B27/34—Indicating arrangements
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/36—Monitoring, i.e. supervising the progress of recording or reproducing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/45—Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
- H04N21/466—Learning process for intelligent management, e.g. learning user preferences for recommending movies
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
- H04N21/8456—Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/854—Content authoring
- H04N21/8549—Creating video summaries, e.g. movie trailer
Definitions
- the present disclosure relates to the field of computer technology, and in particular, to a video processing method, device, readable medium, and electronic equipment.
- the present disclosure provides a video processing method, the method including:
- the quality score corresponding to the target video segment is displayed at the time position corresponding to the target video segment, where the time position corresponding to the target video segment is the time position of the target video segment in the The time position that appears in the target video.
- the present disclosure provides a video processing device, the device including:
- the dividing module is used to divide the target video to obtain the target video segment
- the first determining module is configured to determine the quality score corresponding to the target video segment according to the video frame images contained in the target video segment;
- the first display module is configured to display the quality score corresponding to the target video segment at the time position corresponding to the target video segment on the quality score display time axis, wherein the time position corresponding to the target video segment is the time position corresponding to the target video segment. The time position where the target video segment appears in the target video.
- the present disclosure provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect of the present disclosure are implemented.
- an electronic device including:
- a storage device on which a computer program is stored
- the processing device is configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect of the present disclosure.
- the present disclosure provides a computer program product, the program product comprising: a computer program that, when executed by a processing device, implements the steps of the method described in the first aspect of the present disclosure.
- the present disclosure provides a computer program that, when executed by a processing device, implements the steps of the method described in the first aspect of the present disclosure.
- the target video is divided to obtain the target video segment, and according to the video frame images contained in the target video segment, the quality score corresponding to the target video segment is determined, and the quality score is displayed on the time axis in the target video
- the quality score corresponding to the target video segment is displayed at the time position corresponding to the segment. That is, after the quality score corresponding to the target video segment in the target video is obtained, the quality score corresponding to the target video segment can be displayed at the corresponding position of the quality score display time axis.
- Fig. 1 is a flowchart of a video processing method provided according to an embodiment of the present disclosure
- 2A-2C are several exemplary schematic diagrams showing quality scores in the video processing method provided according to the present disclosure.
- 3A-3C are several exemplary schematic diagrams showing quality scores in the video processing method provided according to the present disclosure.
- 4A-4E are exemplary schematic diagrams displayed by the client in the video processing method provided according to the present disclosure.
- Fig. 5 is a block diagram of a video processing device according to an embodiment of the present disclosure.
- Fig. 6 is a block diagram of a device according to an exemplary embodiment.
- the present disclosure provides a video processing method, device, readable medium, and electronic equipment.
- the target video segment is obtained by dividing the target video, and the quality score corresponding to the target video segment is determined according to the video frame images contained in the target video segment. And, on the quality score display timeline, the quality score corresponding to the target video clip is displayed at the time position corresponding to the target video clip, so as to provide users with a visual display result about the quality score of the target video clip for the user to view. It solves the problem that the user manual selection of video clips in the prior art is not efficient.
- audio/video editing includes a three-layer structure, namely the business layer (front desk), the SDK layer (middle stage), and the algorithm layer (backstage).
- SDK is the abbreviation of Software Development Kit , Chinese means "software development kit”.
- the business layer is responsible for receiving user operations, that is, the client; the SDK layer is responsible for data transmission, such as passing the data to be processed to the algorithm layer to obtain the processing results of the algorithm layer, and further process the data according to the obtained processing results.
- the layer can be responsible for audio/video frame extraction, encoding and decoding, transmission, etc.
- the SDK layer can set data processing strategies; the algorithm layer is responsible for processing the data passed in by the SDK layer, and output the processing results obtained to the SDK layer .
- the method provided in the present disclosure is mainly applied to video editing scenarios, and the related algorithms used in the present disclosure are integrated in the algorithm layer.
- the steps related to data processing in the method provided in the present disclosure can be executed by the SDK layer (middle stage), and the final The data processing result (for example, quality score) can be displayed on the client.
- Fig. 1 is a flowchart of a video processing method provided according to an embodiment of the present disclosure. As shown in Figure 1, the method may include the following steps.
- step 11 the target video is divided to obtain the target video segment.
- the target video is the video for which the user needs to perform video editing.
- the video can be used as the target video, and a series of steps in the video processing method provided in the present disclosure can be executed.
- the user can determine the target video through the visual interface of the device (for example, a terminal).
- the target video segment may include one frame of video frame image or multiple frames of video frame image.
- the “multiple” in this solution refers to two or more than two, and correspondingly, “multi-frame” refers to two or more frames.
- step 11 may include the following steps:
- the target video is divided to obtain multiple video clips including the target video clip.
- the slicing algorithm can be obtained by training the machine learning algorithm through the first training data.
- the first training data may include the first historical video and historical segmentation time points used to divide the first historical video.
- the historical segmentation time point used to divide the first historical video can be marked by the user. For example, the user can start from the perspectives of the main body of the screen, the color of the screen, the similarity of the screen, the position change of the screen content, etc., and focus on the first historical video. Video marks the time point of historical segmentation.
- the user can place the video frame image A1 in the first historical video A at the corresponding time
- the point is regarded as one of the historical division time points of the first historical video A.
- the algorithm can learn the picture change characteristics of the video frame images before and after the historical segmentation time point in the first historical video, that is, the image characteristics corresponding to the "picture transition", so that the final slice algorithm is obtained It has the ability to recognize this "picture transition", can determine the video frame image suitable for use as a division point based on the video frame image in the input video, and use the video frame image suitable for the division point in the input video
- the time point (that is, the time point in the input video that is suitable for cropping) is used as the output of the slicing algorithm.
- the target segmentation time point output by the slicing algorithm can be obtained.
- multiple target video frame images at different time positions in the target video may include all the video frame images in the target video to ensure the accuracy of the output result of the slicing algorithm.
- one or more video frame images can be extracted from the video frame images at various time positions in the target video as multiple target video frame images at different time positions in the target video, so as to reduce the data processing of the slicing algorithm pressure.
- each video segment may include one frame of video image or multiple frames of video frame image.
- the divided multiple video clips including the target video clip may have only one video clip containing one video frame image, or only a video clip containing multiple video frame images, or both There are video clips containing one frame of video frame images, and there are video clips containing multiple frames of video frame images.
- the target segmentation time point corresponding to the target video can be quickly obtained, so that the target video can be divided according to the target segmentation time point, and multiple segments including the target video segment can be obtained. Video clips to speed up the speed of obtaining the target video clips.
- each video frame image of the target video can also be directly used as a video segment, that is, each video segment corresponds to a video frame image of the target video, so that the target video segment It is a video frame image in the target video.
- step 12 the quality score corresponding to the target video segment is determined according to the video frame images contained in the target video segment.
- the quality score of the video clip can reflect the performance of the video clip in the preset effect measurement dimension.
- the preset effect measurement dimensions may include, but are not limited to, the following: richness of the picture, wonderfulness of the picture, prominence of the subject in the picture, degree of light change, degree of movement change, aesthetic quality, and composition quality. For example, if the character in the video clip is clearly distinguished from the background (that is, the protruding degree of the subject in the screen is high), the video clip corresponds to a higher quality score, and if the difference between the character and the background in the video clip is small (that is, the subject in the screen The degree of prominence is low), the video clip corresponds to a lower quality score.
- the video clip corresponds to a higher quality score
- the screen content in the video clip is single (that is, the screen richness is low)
- the video clip corresponds to Lower quality score
- the quality score corresponding to the target video segment may be determined in the following manner:
- the video frame images contained in the target video segment are input to the first quality evaluation model, and the quality score corresponding to the target video segment is obtained.
- the first quality evaluation model is obtained by training on the second training data, for example, obtained by training a machine learning algorithm on the second training data.
- the second training data includes the second historical video and historical quality scores corresponding to the second historical video.
- a second historical video is used as input data, and the historical quality score corresponding to the second historical video is used as output data, and training is performed based on a machine learning algorithm to obtain a first quality evaluation model.
- the historical quality score corresponding to the second historical video may be determined by at least one of the following:
- the first score can reflect the manual annotation made by the user for the second historical video, and the intuitive evaluation made by the user from the perspective of the second historical video can be obtained.
- the second score can reflect the clarity of the second historical video, and the higher the clarity of the second historical video, the higher the second score corresponding to the second historical video.
- the third score can reflect the degree of deviation of the first object (which can be preset, for example, a certain person, a certain building, etc.) in the second historical video screen from the center of the screen in the second historical video. The closer an object is to the center of the screen in the second historical video (that is, the smaller the deviation of the first object from the center of the screen in the second historical video is), the higher the third score corresponding to the second historical video.
- the fourth score can reflect the proportion of the second object (which can be preset, for example, a certain person, a certain building, etc.) in the second historical video frame in the second historical video, and the second object is in the second historical video.
- the weight corresponding to each score can be set. Combining the actual score and the weight corresponding to the score, the historical quality score corresponding to the second historical video is calculated.
- the first quality evaluation model is trained based on the second historical video and the historical quality scores corresponding to the second historical video.
- the evaluation result therefore, the video frame image contained in the target video segment is input to the first quality evaluation model, and the output result of the first quality evaluation model obtained is the quality score corresponding to the target video segment.
- all the video frame images included in the target video segment may be input to the first quality evaluation model to ensure the accuracy of the output result of the first quality evaluation model.
- a part of the video frame image contained in the target video segment may be input to the first quality evaluation model to reduce the data processing pressure of the first quality evaluation model.
- the quality score corresponding to the target video segment can be determined in the following manner:
- the quality score corresponding to the target video segment is determined.
- the second quality evaluation model is obtained through training on the third training data, for example, obtained by training a machine learning algorithm through the third training data.
- the third training data may include historical images and historical quality scores corresponding to the historical images.
- a historical image is used as input data
- the historical quality score corresponding to the historical image is used as output data, and training is performed based on a machine learning algorithm to obtain a second quality evaluation model.
- the historical quality score corresponding to the historical image may be determined by at least one of the following:
- the eighth score is determined according to the proportion of the fourth object in the historical image in the overall picture, wherein the eighth score of the historical image is in a positive correlation with the proportion of the fourth object in the historical image in the overall picture of the historical image.
- the fifth score can reflect the manual annotation made by the user for the historical image, and the intuitive evaluation of the historical image from the user's perspective can be obtained.
- the sixth score can reflect the sharpness of the historical image, and the higher the sharpness of the historical image, the higher the sixth score corresponding to the historical image.
- the seventh score can reflect the degree of deviation of the third object (which can be preset, for example, a certain person, a certain building, etc.) in the historical image frame from the center of the screen in the historical image, and the third object in the historical image The closer the image is to the center of the screen (that is, the smaller the deviation of the third object from the center of the screen in the historical image screen), the higher the seventh score corresponding to the historical image.
- the eighth score can reflect the proportion of the fourth object (which can be preset, for example, a certain person, a certain building, etc.) in the historical image frame in the historical image, and the fourth object occupies the historical image The larger the ratio of the screen, the higher the eighth score corresponding to the historical image.
- the weight corresponding to each point value can be set, combining the actual The score and the weight corresponding to the score are used to calculate the historical quality score corresponding to the historical image.
- the second quality evaluation model is trained based on historical images and historical quality scores corresponding to the historical images. That is to say, the second quality evaluation model can obtain the evaluation result for the image based on the input single image, so ,
- the video frame images contained in the target video segment are respectively input to the second quality evaluation model, and the output result of the second quality evaluation model obtained is the initial quality corresponding to the video frame image in the target video segment that is input to the second evaluation model Score, that is, the initial quality score corresponding to a single video frame image.
- all the video frame images included in the target video segment may be input to the second quality evaluation model to ensure the accuracy of the output result of the second quality evaluation model.
- a part of the video frame images contained in the target video segment may be input to the second quality assessment model respectively, so as to reduce the data processing pressure of the second quality assessment model.
- the overall quality score corresponding to the target video segment is not yet known. Therefore, it is necessary to determine the quality score corresponding to the target video segment based on these initial quality scores.
- determining the quality score corresponding to the target video segment according to the initial quality score may include any one of the following:
- the median of the initial quality score is determined as the quality score corresponding to the target video segment.
- step 13 on the quality score display time axis, the quality score corresponding to the target video segment is displayed at the time position corresponding to the target video segment.
- the time position corresponding to the target video segment is the time position where the target video segment appears in the target video.
- the quality score display time axis is a time axis generated according to the target video, and each time point on it corresponds to a corresponding time point in the target video.
- step 13 may include any one of the following:
- the quality score corresponding to the target video clip is displayed in digital form at the time position corresponding to the target video clip;
- the quality score corresponding to the target video segment is displayed with a dot mark
- the quality score corresponding to the target video segment is displayed with a bar mark, where the maximum score of the bar mark corresponding to the target video segment is the value corresponding to the target video segment Quality score.
- the display mode of the quality score can be as the following examples.
- FIG. 2A shows an exemplary display diagram showing the quality score corresponding to the target video segment in digital form at the time position corresponding to the target video segment on the quality score display time axis.
- Figure 2B shows that on the quality score display time axis, at the time position corresponding to the target video segment and at the score position corresponding to the quality score of the target video segment, the quality score corresponding to the target video segment is displayed with a dot mark
- An exemplary display diagram An exemplary display diagram.
- FIG. 2C shows an exemplary display diagram showing the quality score corresponding to the target video segment with bar marks at the time position corresponding to the target video segment on the quality score display time axis.
- the target video can be divided into multiple video segments including the target video segment. Therefore, when the quality score is displayed, the respective quality scores of the multiple video segments can be displayed.
- the quality score corresponding to each video segment can be determined with reference to the process of determining the quality score corresponding to the target video segment, which will not be repeated here. Therefore, on the basis of the above-mentioned embodiments, the quality scores corresponding to multiple video segments in the target video can be obtained, and the quality scores corresponding to these video segments can be displayed to visually display the respective quality of the multiple video segments in the target video. Fraction.
- the quality score corresponding to the video segment composed of 0-50s in the target video is 0.8
- the quality score corresponding to the video segment composed of 50s-70s in the target video is 0.6
- the video composed of 70s-90s in the target video The quality score corresponding to the fragment
- the quality score corresponding to the video fragment composed of 90s to 150s in the target video is 0.3
- the exemplary display of the quality score can be shown in Figures 3A, 3B, and 3C (corresponding to the above Figure 2A, 2B, 2C display mode).
- FIG. 3B after displaying the quality scores corresponding to each video segment in the target video, it can also be connected in sequence according to the displayed dot marks to show the change of the quality scores more intuitively.
- the target video is divided to obtain the target video segment, and according to the video frame images contained in the target video segment, the quality score corresponding to the target video segment is determined, and the quality score is displayed on the time axis in the target video
- the quality score corresponding to the target video segment is displayed at the time position corresponding to the segment. That is, after the quality score corresponding to the target video segment in the target video is obtained, the quality score corresponding to the target video segment can be displayed at the corresponding position of the quality score display time axis.
- the quality score display timeline may be a video editing timeline corresponding to the target video, that is, the quality score of the target video segment in the target video is displayed on the editing interface of the target video.
- the target video segment corresponds to the start segmentation time point and the end segmentation time point in the target video, that is, the start and end time corresponding to the target video segment in the target video.
- the present disclosure may also include the following steps:
- the cutting mark corresponding to the target video segment is displayed at the start segmentation time point and the end segmentation time point corresponding to the target video segment.
- the cutting mark is used to show the user the starting point and the ending point of the target video clip in the target video, so that the user can know the specific position of the target video clip in the target video, the corresponding picture, etc.
- the cutting mark can provide users with a one-click cutting function, that is, when the cutting mark is clicked, the target video clip is copied (or cut) from the target video to form the target video.
- the video file corresponding to the video clip is used to show the user the starting point and the ending point of the target video clip in the target video, so that the user can know the specific position of the target video clip in the target video, the corresponding picture, etc.
- the cutting mark can provide users with a one-click cutting function, that is, when the cutting mark is clicked, the target video clip is copied (or cut) from the target video to form the target video.
- the video file corresponding to the video clip is a one-click cutting function
- the target video is divided to obtain multiple video segments including the target video, and each video segment can display the cutting mark corresponding to the video segment in the above-mentioned manner. Therefore, when the cutting mark corresponding to a certain video clip is clicked, a file corresponding to the video clip corresponding to the clicked cutting mark can be formed.
- the corresponding position of the target video segment in the target video is displayed through the cutting mark, and the detailed content of the target video segment is displayed to the user, which is convenient for the user to save the video segment according to their own needs.
- the method of the present disclosure may further include the following steps:
- the target video segment is determined as a candidate material for video splicing.
- the target video segment can be determined as a candidate material for video splicing to provide usable material for subsequent video splicing.
- multiple candidate materials including the target video segment may also be synthesized into the target spliced video.
- other candidate materials other than the target video fragment may be other video fragments other than the target video fragment in the target video, or may be taken from other videos other than the target video
- this disclosure does not limit this.
- the quality score corresponding to each video segment in the target video can be obtained by referring to the method given above, and the first few with higher quality scores are used as candidate materials to synthesize the target splicing video.
- the quality score corresponding to the target video segment is higher than the quality scores corresponding to other video segments in the target video, you can also directly add other content in the target video except the target video segment. Delete, but only keep the target video clip.
- the target video segment is a single-frame video frame image
- the above steps are equivalent to retaining only the video frame image with the highest quality score in the target video, that is, the "high light moment" corresponding to the target video.
- the target video segment is a multi-frame video frame image
- the above steps are equivalent to retaining the highest score segment in the target video. In this way, the highest quality part of the video can be automatically reserved for the user, without the user having to view it frame by frame, and the user's video editing efficiency can be improved.
- the interactive page of the client can be displayed as shown in Figure 4A to Figure 4E.
- a video editing page is shown, which displays the video editing timeline corresponding to the target video.
- the button Bu1 at the bottom right corner of the page is used to trigger the "intelligent evaluation" function of the target video, for example, displaying the target video clip. Quality score, retain video clips with high quality scores in the target video, etc.
- the button Bu1 in FIG. 4A it indicates that the user has a need for intelligent evaluation of the current video.
- the video displayed on the page is the target video
- the relevant steps provided in the present disclosure can be performed to process the target video, for example, to determine the quality score corresponding to the target video segment, or if the quality score corresponding to the target video segment is high
- the client's display content can be shown in Figure 4B, where the percentage in the center of Figure 4B is used to indicate the progress of data processing.
- the current data processing is completed, that is, the middle station pair
- the intelligent evaluation of the target video has been completed, and the content displayed on the client at this time can be as shown in Figure 4C.
- the specific intelligent evaluation result can be displayed, as shown in FIG. 4D or FIG. 4E.
- the curve represents the quality score curve formed by the quality score of each video segment in the target video.
- the video clips retained after the target video is clipped are also displayed below the quality score curve.
- corresponding video frame images are displayed at several higher quality score positions in the quality score curve, that is, the "highlight moment" corresponding to the target video.
- Fig. 5 is a block diagram of a video processing device provided according to an embodiment of the present disclosure. As shown in Fig. 5, the device 40 may include:
- the dividing module 41 is used to divide the target video to obtain a target video segment
- the first determining module 42 is configured to determine the quality score corresponding to the target video segment according to the video frame images contained in the target video segment;
- the first display module 43 is configured to display the quality score corresponding to the target video segment at the time position corresponding to the target video segment on the quality score display time axis, where the time position corresponding to the target video segment is The time position where the target video segment appears in the target video.
- the dividing module 41 includes:
- the slicing sub-module is used to input the multiple frames of target video frame images at different time positions in the target video and the time when the multiple frames of target video frame images appear in the target video to the slicing algorithm to obtain the slices
- the target segmentation time point of the algorithm output, the slice algorithm is obtained by training the machine learning algorithm through the first training data, the first training data includes the first historical video and the history used to divide the first historical video Split point in time;
- the dividing sub-module is configured to divide the target video according to the target segmentation time point to obtain multiple video segments including the target video segment.
- the first determining module 42 is configured to determine the quality score corresponding to the target video segment in the following manner:
- the video frame images contained in the target video segment are input to a first quality evaluation model to obtain a quality score corresponding to the target video segment, wherein the first quality evaluation model is obtained by training through second training data, and
- the second training data includes a second historical video and historical quality scores corresponding to the second historical video.
- the historical quality score corresponding to the second historical video is determined by at least one of the following:
- a fourth score determined according to the proportion of the second object in the second historical video in the overall picture, where the fourth score of the second historical video and the proportion of the second object in the overall picture of the second historical video The relationship is positively correlated.
- the first determining module 42 includes:
- the processing sub-module is configured to input the video frame images contained in the target video segment into a second quality evaluation model, and obtain the second quality evaluation model for each video frame input to the second quality evaluation model
- the initial quality score of the image output wherein the second quality assessment model is obtained through training of third training data, and the third training data includes historical images and historical quality scores corresponding to the historical images;
- the determining sub-module is configured to determine the quality score corresponding to the target video segment according to the initial quality score.
- the determining submodule is configured to determine the quality score corresponding to the target video segment through any one of the following:
- the median of the initial quality score is determined as the quality score corresponding to the target video segment.
- the historical quality score corresponding to the historical image is determined by at least one of the following:
- a fifth score determined according to the user's mark score for the historical image is
- a sixth score determined according to the resolution of the historical image wherein the sixth score of the historical image has a positive correlation with the resolution of the historical image;
- the target video is divided into multiple video segments including the target video segment;
- the device 40 also includes:
- the second determining module is configured to, in response to receiving the editing instruction for the target video, if the quality score corresponding to the target video segment is higher than the quality scores corresponding to other video segments in the target video, the target video The segment is determined as the candidate material for video splicing.
- the device 40 further includes:
- the synthesis module is used to synthesize multiple candidate materials including the target video segment into a target spliced video.
- the first display module 43 includes any one of the following:
- the first display sub-module is configured to display the quality score corresponding to the target video segment in digital form at the time position corresponding to the target video segment on the quality score display time axis;
- the second display submodule is used to mark the time position corresponding to the target video segment and the score position corresponding to the quality score of the target video segment on the quality score display time axis Displaying the quality score corresponding to the target video segment;
- the third display sub-module is configured to display the quality score corresponding to the target video segment with a bar mark at the time position corresponding to the target video segment on the quality score display time axis, where the target video segment The maximum score of the corresponding bar mark is the quality score corresponding to the target video segment.
- the quality score display time axis is a video editing time axis corresponding to the target video.
- the target video segment corresponds to a start segmentation time point and an end segmentation time point in the target video
- the device 40 also includes:
- the second display module is configured to display the data corresponding to the target video segment at the start segmentation time point and the end segmentation time point corresponding to the target video segment on the video editing time axis corresponding to the target video Cutting mark.
- FIG. 6 shows a schematic structural diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure.
- the terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PAD (Tablet PC, Portable Android Device), PMP (Portable Multimedia Players, mobile terminals such as Personal Multimedia Player, in-vehicle terminals (for example, in-vehicle navigation terminals), and fixed terminals such as digital televisions (TV), desktop computers, etc.
- the electronic device shown in FIG. 6 is only an example, and should not bring any limitation to the function and scope of use of the embodiments of the present disclosure.
- the electronic device 600 may include a processing device (such as a central processing unit, a graphics processor, etc.) 601, which may be based on a program stored in a read-only memory (Read-Only Memory, ROM) 602 or from a storage device 608 is loaded into a random access memory (Random Access Memory, RAM) 603 program to execute various appropriate actions and processing.
- a processing device such as a central processing unit, a graphics processor, etc.
- ROM read-only memory
- RAM Random Access Memory
- various programs and data required for the operation of the electronic device 600 are also stored.
- the processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604.
- An input/output (Input/Output, I/O) interface 605 is also connected to the bus 604.
- the following devices can be connected to the I/O interface 605: including input devices 606 such as touch screen, touch panel, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; including, for example, a liquid crystal display (LCD) Output devices 607 such as speakers, vibrators, etc.; storage devices 608 such as magnetic tapes, hard disks, etc.; and communication devices 609.
- the communication device 609 may allow the electronic device 600 to perform wireless or wired communication with other devices to exchange data.
- FIG. 6 shows an electronic device 600 having various devices, it should be understood that it is not required to implement or have all of the illustrated devices. It may be implemented alternatively or provided with more or fewer devices.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program may be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602.
- the processing device 601 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
- the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or a combination of any of the above.
- Computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable removable Programmable read-only memory (Electrical Programmable ROM, EPROM or flash memory), optical fiber, portable compact disc read-only memory (Compact Disc ROM, CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as a part of a carrier wave, and a computer-readable program code is carried therein.
- This propagated data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.
- the computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium.
- the computer-readable signal medium may send, propagate, or transmit the program for use by or in combination with the instruction execution system, apparatus, or device .
- the program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wire, optical cable, RF (Radio Frequency), etc., or any suitable combination of the foregoing.
- the terminal and the server can communicate with any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and can communicate with any form or medium of digital data.
- HTTP HyperText Transfer Protocol
- communication networks include local area networks (Local Area Network, "LAN”), Wide Area Networks ("WAN"), Internet (for example, the Internet), and end-to-end networks (for example, ADaptive Heuristic for Opponent Classification). , Ad hoc) end-to-end network), and any networks currently known or developed in the future.
- LAN Local Area Network
- WAN Wide Area Networks
- Internet for example, the Internet
- end-to-end networks for example, ADaptive Heuristic for Opponent Classification
- Ad hoc end-to-end network
- the above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist alone without being assembled into the electronic device.
- the above-mentioned computer-readable medium carries one or more programs.
- the electronic device When the above-mentioned one or more programs are executed by the electronic device, the electronic device: divides the target video to obtain the target video segment; The included video frame images determine the quality score corresponding to the target video segment; on the quality score display time axis, the quality score corresponding to the target video segment is displayed at the time position corresponding to the target video segment, where The time position corresponding to the target video segment is the time position where the target video segment appears in the target video.
- the computer program code used to perform the operations of the present disclosure can be written in one or more programming languages or a combination thereof.
- the above-mentioned programming languages include but are not limited to object-oriented programming languages such as Java, Smalltalk, C++, and Including conventional procedural programming languages-such as "C" language or similar programming languages.
- the program code can be executed entirely on the user's computer, partly on the user's computer, executed as an independent software package, partly on the user's computer and partly executed on a remote computer, or entirely executed on the remote computer or server.
- the remote computer can be connected to the user’s computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to Connect via the Internet).
- LAN local area network
- WAN wide area network
- each block in the flowchart or block diagram may represent a module, program segment, or part of code, and the module, program segment, or part of code contains one or more for realizing the specified logical function Executable instructions.
- the functions marked in the block may also occur in a different order from the order marked in the drawings. For example, two blocks shown one after another can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved.
- each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or operations Or it can be realized by a combination of dedicated hardware and computer instructions.
- the modules involved in the embodiments described in the present disclosure can be implemented in software or hardware.
- the name of the module does not constitute a limitation on the module itself under certain circumstances.
- the division module can also be described as "a module used to divide a target video to obtain a target video segment".
- exemplary types of hardware logic components include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and application specific standard products (Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programming Logic Device (CPLD), etc.
- FPGA Field Programmable Gate Array
- ASIC Application Specific Integrated Circuit
- ASSP Application Specific Standard Parts
- SOC System on Chip
- CPLD Complex Programming Logic Device
- a machine-readable medium may be a tangible medium, which may contain or store a program for use by the instruction execution system, apparatus, or device or in combination with the instruction execution system, apparatus, or device.
- the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- the machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing.
- machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM compact disk read only memory
- magnetic storage device or any suitable combination of the foregoing.
- a video processing method including:
- the quality score corresponding to the target video segment is displayed at the time position corresponding to the target video segment, where the time position corresponding to the target video segment is the time position of the target video segment in the The time position that appears in the target video.
- dividing the target video to obtain the target video segment includes:
- the slicing algorithm is obtained by training a machine learning algorithm through first training data, where the first training data includes a first historical video and a historical segmentation time point used to divide the first historical video;
- the target video is divided to obtain multiple video segments including the target video segment.
- a video processing method wherein the quality score corresponding to the target video segment is determined in the following manner:
- the video frame images contained in the target video segment are input to a first quality evaluation model to obtain a quality score corresponding to the target video segment, wherein the first quality evaluation model is obtained by training through second training data, and
- the second training data includes a second historical video and historical quality scores corresponding to the second historical video.
- the historical quality score corresponding to the second historical video is determined by at least one of the following:
- a fourth score determined according to the proportion of the second object in the second historical video in the overall picture, where the fourth score of the second historical video and the proportion of the second object in the overall picture of the second historical video The relationship is positively correlated.
- a video processing method wherein the quality score corresponding to the target video segment is determined in the following manner:
- the video frame images included in the target video segment are respectively input to the second quality evaluation model, and the initial quality score of the second quality evaluation model for each video frame image input to the second quality evaluation model is obtained.
- the second quality evaluation model is obtained through training of third training data, and the third training data includes historical images and historical quality scores corresponding to the historical images;
- the quality score corresponding to the target video segment is determined.
- determining the quality score corresponding to the target video segment according to the initial quality score includes any one of the following :
- the median of the initial quality score is determined as the quality score corresponding to the target video segment.
- a video processing method wherein the historical quality score corresponding to the historical image is determined by at least one of the following:
- a fifth score determined according to the user's mark score for the historical image is
- a sixth score determined according to the resolution of the historical image wherein the sixth score of the historical image has a positive correlation with the resolution of the historical image;
- the target video is divided into multiple video segments including the target video segment
- the method also includes:
- the target video segment is determined to be used for video splicing Alternative material.
- a video processing method wherein the method further includes:
- the multiple candidate materials including the target video segment are synthesized into a target spliced video.
- a video processing method wherein, on the quality score display time axis, the target video segment corresponding to the target video segment is displayed at a time position corresponding to the target video segment
- the quality score of including any of the following:
- the quality score corresponding to the target video segment is displayed in a digital form at a time position corresponding to the target video segment;
- the dot marks corresponding to the target video segment are displayed.
- the quality score corresponding to the target video segment is displayed with a bar mark, wherein the maximum score of the bar mark corresponding to the target video segment Is the quality score corresponding to the target video segment.
- the quality score display time axis is a video editing time axis corresponding to the target video.
- the target video segment corresponds to a start segmentation time point and an end segmentation time point in the target video
- the method also includes:
- a cutting mark corresponding to the target video segment is displayed at the start segmentation time point and the end segmentation time point corresponding to the target video segment.
- a video processing device including:
- the dividing module is used to divide the target video to obtain the target video segment
- the first determining module is configured to determine the quality score corresponding to the target video segment according to the video frame images contained in the target video segment;
- the first display module is configured to display the quality score corresponding to the target video segment at the time position corresponding to the target video segment on the quality score display time axis, wherein the time position corresponding to the target video segment is the time position corresponding to the target video segment. The time position where the target video segment appears in the target video.
- dividing module includes:
- the slicing submodule is used to input the multiple frames of target video frame images at different time positions in the target video and the time when the multiple frames of target video frame images appear in the target video to the slicing algorithm to obtain the slices
- the target segmentation time point of the algorithm output, the slice algorithm is obtained by training the machine learning algorithm through the first training data, the first training data includes the first historical video and the history used to divide the first historical video Split point in time;
- the dividing sub-module is configured to divide the target video according to the target segmentation time point to obtain multiple video segments including the target video segment.
- the first determining module is configured to determine the quality score corresponding to the target video segment in the following manner:
- the video frame images contained in the target video segment are input to a first quality evaluation model to obtain a quality score corresponding to the target video segment, wherein the first quality evaluation model is obtained by training through second training data, and
- the second training data includes a second historical video and historical quality scores corresponding to the second historical video.
- a video processing device wherein the historical quality score corresponding to the second historical video is determined by at least one of the following:
- a fourth score determined according to the proportion of the second object in the second historical video in the overall picture, where the fourth score of the second historical video and the proportion of the second object in the overall picture of the second historical video The relationship is positively correlated.
- the first determining module includes:
- the processing sub-module is configured to input the video frame images contained in the target video segment into a second quality evaluation model, and obtain the second quality evaluation model for each video frame input to the second quality evaluation model
- the initial quality score of the image output wherein the second quality assessment model is obtained through training of third training data, and the third training data includes historical images and historical quality scores corresponding to the historical images;
- the determining sub-module is configured to determine the quality score corresponding to the target video segment according to the initial quality score.
- determining submodule is configured to determine the quality score corresponding to the target video segment through any one of the following:
- the median of the initial quality score is determined as the quality score corresponding to the target video segment.
- a video processing device wherein the historical quality score corresponding to the historical image is determined by at least one of the following:
- a fifth score determined according to the user's mark score for the historical image is
- a sixth score determined according to the resolution of the historical image wherein the sixth score of the historical image has a positive correlation with the resolution of the historical image;
- a video processing device wherein the target video is divided into a plurality of video segments including the target video segment;
- the device also includes:
- the second determining module is configured to, in response to receiving the editing instruction for the target video, if the quality score corresponding to the target video segment is higher than the quality scores corresponding to other video segments in the target video, the target video The segment is determined as the candidate material for video splicing.
- a video processing device wherein the device further includes:
- the synthesis module is used to synthesize multiple candidate materials including the target video segment into a target spliced video.
- the first display module includes any one of the following:
- the first display sub-module is configured to display the quality score corresponding to the target video segment in digital form at the time position corresponding to the target video segment on the quality score display time axis;
- the second display submodule is used to mark the time position corresponding to the target video segment and the score position corresponding to the quality score of the target video segment on the quality score display time axis Displaying the quality score corresponding to the target video segment;
- the third display sub-module is configured to display the quality score corresponding to the target video segment with a bar mark at the time position corresponding to the target video segment on the quality score display time axis, where the target video segment The maximum score of the corresponding bar mark is the quality score corresponding to the target video segment.
- the quality score display time axis is a video editing time axis corresponding to the target video.
- a video processing device wherein the target video segment corresponds to a start segmentation time point and an end segmentation time point in the target video;
- the device also includes:
- the second display module is configured to display the data corresponding to the target video segment at the start segmentation time point and the end segmentation time point corresponding to the target video segment on the video editing time axis corresponding to the target video Cutting mark.
- a computer program product includes: a computer program, which, when executed by a processing device, implements the steps of the method described in any embodiment of the present disclosure.
- a computer program is also provided, which, when executed by a processing device, implements the steps of the method described in any embodiment of the present disclosure.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Databases & Information Systems (AREA)
- Human Computer Interaction (AREA)
- Computer Security & Cryptography (AREA)
- Television Signal Processing For Recording (AREA)
Abstract
本公开涉及一种视频处理方法、装置、可读介质和电子设备。所述方法包括:对目标视频进行划分,得到目标视频片段;根据目标视频片段中包含的视频帧图像,确定目标视频片段对应的质量分数;在质量分数展示时间轴上,在目标视频片段对应的时间位置处展示目标视频片段对应的质量分数,其中,目标视频片段对应的时间位置为目标视频片段在目标视频中出现的时间位置。由此,能够为用户提供有关目标视频片段质量分数的可视化展示结果,为用户的视频片段挑选提供参考,节省用户查看目标视频片段所花费的时间。另外,还能形成针对多个视频片段的质量分数的直观比较,为用户从目标视频中挑选视频片段提供参考依据,方便用户快速挑选视频片段。
Description
本申请要求于2020年02月11日提交中国专利局、申请号为202010087010.9、申请名称为“视频处理方法、装置、可读介质和电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本公开涉及计算机技术领域,具体地,涉及一种视频处理方法、装置、可读介质和电子设备。
在视频编辑场景下,用户在进行视频剪辑时,需要从一段较长的视频中挑选出较短的视频片段。一般情况下,用户需要对原始视频逐帧查看并从中筛选出合适的部分,而在原始视频很长的情况下,在挑选视频片段的过程中,用户需要花费大量的时间进行手动挑选,既耗费人力及时间,效率也不够高。
发明内容
提供该发明内容部分以便以简要的形式介绍构思,这些构思将在后面的具体实施方式部分被详细描述。该发明内容部分并不旨在标识要求保护的技术方案的关键特征或必要特征,也不旨在用于限制所要求的保护的技术方案的范围。
第一方面,本公开提供一种视频处理方法,所述方法包括:
对目标视频进行划分,得到目标视频片段;
根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;
在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
第二方面,本公开提供一种视频处理装置,所述装置包括:
划分模块,用于对目标视频进行划分,得到目标视频片段;
第一确定模块,用于根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;
第一展示模块,用于在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
第三方面,本公开提供一种计算机可读介质,其上存储有计算机程序,该程序被处理装置执行时实现本公开第一方面所述方法的步骤。
第四方面,本公开提供一种电子设备,包括:
存储装置,其上存储有计算机程序;
处理装置,用于执行所述存储装置中的所述计算机程序,以实现本公开第一方面所述方法的步骤。
第五方面,本公开提供一种计算机程序产品,所述程序产品包括:计算机程序,所述计算机程序被处理装置执行时实现本公开第一方面所述方法的步骤。
第六方面,本公开提供一种计算机程序,所述计算机程序被处理装置执行时实现本公开第一方面所述方法的步骤。
通过上述技术方案,对目标视频进行划分,得到目标视频片段,以及,根据目标视频片段中包含的视频帧图像,确定目标视频片段对应的质量分数,并在质量分数展示时间轴上,在目标视频片段对应的时间位置处展示目标视频片段对应的质量分数。也就是说,在得到目标视频中的目标视频片段对应的质量分数后,能够在质量分数展示时间轴的相应位置展示目标视频片段对应的质量分数。由此,能够为用户提供有关于目标视频片段质量分数的可视化展示结果,供用户查看,方便用户快速获知目标视频片段对应的质量分数,从而为用户的视频片段挑选提供参考,节省用户查看目标视频片段所花费的时间。另外,基于上述方案,能够确定并展示目标视频中多个视频片段各自对应的质量分数,从而能够形成针对多个视频片段的质量分数的直观比较,为用户从目标视频中挑选视频片段提供参考依据,方便用户快速挑选视频片段。
本公开的其他特征和优点将在随后的具体实施方式部分予以详细说明。
结合附图并参考以下具体实施方式,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。贯穿附图中,相同或相似的附图标记表示相同或相似的元素。应当理解附图是示意性的,原件和元素不一定按照比例绘制。在附图中:
图1是根据本公开的一种实施方式提供的视频处理方法的流程图;
图2A-2C是根据本公开提供的视频处理方法中,展示质量分数的几种示例性的示意图;
图3A-3C是根据本公开提供的视频处理方法中,展示质量分数的几种示例性的示意图;
图4A-4E是根据本公开提供的视频处理方法中,客户端展示的种示例性的示意图;
图5是根据本公开的一种实施方式提供的视频处理装置的框图;
图6是根据一示例性实施例示出的设备的框图。
下面将参照附图更详细地描述本公开的实施例。虽然附图中显示了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
应当理解,本公开的方法实施方式中记载的各个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
需要注意,本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
本公开提供了一种视频处理方法、装置、可读介质和电子设备,通过对目标视频进行划分得到目标视频片段,根据目标视频片段中包含的视频帧图像,确定目标视频片段对应的质量分数,以及,在质量分数展示时间轴上,在目标视频片段对应的时间位置处展示目标视频片段对应的质量分数,从而,为用户提供有关于目标视频片段质量分数的可视化展示结果,供用户查看,以解决现有技术中用户手动挑选视频片段效率不高的问题。
在音/视频处理领域,一般情况下,音/视频编辑包括三层结构,分别是业务层(前台)、SDK层(中台)、算法层(后台),其中,SDK是Software Development Kit的缩写,中文意思是“软件开发工具包”。业务层负责接收用户操作,即客户端;SDK层负责数据传递,比如将待处理数据传递给算法层,以获得算法层的处理结果,并根据得到的处理结果进一步处理数据,举例来说,SDK层可以负责音/视频的抽帧、编解码、传递等,同时,在SDK层可以设置针对数据的处理策略;算法层负责处理SDK层传入的数据,并将得到的处理结果输出给SDK层。
本公开提供的方法主要应用于视频编辑场景,并且,本公开所使用的相关算法集成在算法层,本公开提供的方法中有关于数据处理的步骤可以由SDK层(中台)执行,最终的数据处理结果(例如,质量分数)可以在客户端进行展示。
图1是根据本公开的一种实施方式提供的视频处理方法的流程图。如图1所示,该方法可以包括以下步骤。
在步骤11中,对目标视频进行划分,得到目标视频片段。
目标视频就是用户需要进行视频剪辑的视频,换言之,用户期望从哪个视频中挑选视频片段,就可以将该视频作为目标视频,并执行本公开提供的视频处理方法中的一系列步骤。实际应用时,用户可以通过设备(例如,终端)的可视化界面确定目标视频。
目标视频片段可以包括一帧视频帧图像或多帧视频帧图像。在未做另外说明的情况下,本方案中的“多个”指代两个或两个以上,相应地,“多帧”指代两帧或两帧以上。
在一种可能的实施方式中,对目标视频的划分可以基于预先训练所得的切片算法,相应地,步骤11可以包括以下步骤:
将目标视频中处于不同时间位置的多帧目标视频帧图像和该多帧目标视频帧图像在目标视频中出现的时间输入至切片算法,获得切片算法输出的目标分割时间点;
根据目标分割时间点,对目标视频进行划分,得到包括目标视频片段在内的多个视频片段。
切片算法可以通过第一训练数据对机器学习算法进行训练得到。第一训练数据可以包括第一历史视频和用于对第一历史视频进行划分的历史分割时间点。其中,用于对第一历史视频进行划分的历史分割时间点可以由用户标注,示例地,用户可以从画面主体、画面颜色、画面相似度、画面内容的位置变化等角度出发,针对第一历史视频标注历史分割时间点。例如,若第一历史视频A中,视频帧图像A1后的视频帧图像A2与视频帧图像A1的画面颜色相差较大,则用户可以将视频帧图像A1在第一历史视频A中对应的时间点作为第一历史视频A的历史分割时间点之一。
在训练过程中,每次以一个第一历史视频为输入数据、并以该第一历史视频对应的历史分割时间点为输出数据,对机器学习算法进行训练。经过多次的训练,算法内部能够学习到第一历史视频中处于历史分割时间点前后的视频帧图像的画面变化特征,也就是“画面转折”所对应的图像特征,从而,最终得到的切片算法具有能够识别这种“画面转折”的能力,能够基于输入视频中的视频帧图像确定出其中适合用于作为划分点的视频帧图像,并将适合用于作为划分点的视频帧图像在输入视频中所处的时间点(即,输入视频中适合裁剪的时间点)作为切片算法的输出。
需要说明的是,对机器学习算法进行训练的过程为本领域的公知常识,此处不进行详细描述。
因此,将目标视频中处于不同时间位置的多帧目标视频帧图像和多帧目标视频帧图像在目标视频中出现的时间输入至切片算法后,就能够获得切片算法输出的目标分割时间点。其中,目标视频中处于不同时间位置的多帧目标视频帧图像可以包含目标视频中的全部视频帧图像,以保证切片算法输出结果的准确性。或者,可以分别从目标视频中处于各个时间位置的视频帧图像中抽取出一帧或多帧视频帧图像作为目标视频中处于不同时间位置的多帧目标视频帧图像,以减轻切片算法的数据处理压力。
之后,根据目标分割时间点,对目标视频进行划分后,能够得到包括目标视频片段在内的多个视频片段。其与目标视频片段相似,基于目标分割时间点对目标视频进行划分后所得的多个视频片段中,每一视频片段可以包括一帧视频帧图像或多帧视频帧图像。对于同一目标视频来说,划分所得的包括目标视频片段在内的多个视频片段可以仅有包含一帧视频帧图像的视频片段,或者仅有包含多帧视频帧图像的视频片段,或者既有包含一帧视频帧图像的视频片段、又有包含多帧视频帧图像的视频片段。
采用上述方式,利用基于机器学习算法训练所得的切片算法,能够快速得到目标视频对应的目标分割时间点,从而,能够根据目标分割时间点对目标视频进行划分,得到包括目标视频片段在内的多个视频片段,加快获得目标视频片段的速度。
在另一种可能的实施方式中,还可以直接将目标视频的每一帧视频帧图像直接作为一个视频片段,也就是每个视频片段对应目标视频的一帧视频帧图像,从而,目标视频片段就是目标视频中的一帧视频帧图像。
在步骤12中,根据目标视频片段中包含的视频帧图像,确定目标视频片段对应的质量分数。
视频片段的质量分数可以反映视频片段在预设的效果衡量维度下的表现。其中,预设的效果衡量维度可以包括但不限于以下几者:画面丰富程度、画面精彩程度、画面内主体的突出程度、光线变化程度、运动变化程度、美学质量、构图质量。例如,若视频片段内人物与背景区别明显(即,画面内主体的突出程度高),该视频片段对应较高的质量分数,而若视频片段内人物与背景区别较小(即,画面内主体的突出程度低),该视频片段对应较低的质量分数。再例如,若视频片段内画面内容丰富(即,画面丰富程度高),该视频片段对应较高的质量分数,而若视频片段内画面内容单一(即,画面丰富程度低),该视频片段对应较低的质量分数。
在一种可能的实施方式中,目标视频片段对应的质量分数可以通过如下方式确定:
将目标视频片段中包含的视频帧图像输入至第一质量评估模型,获得目标视频片段对应的质量分数。
其中,第一质量评估模型通过第二训练数据训练得到,例如,通过第二训练数据对机器学习算法进行训练得到。第二训练数据包括第二历史视频和第二历史视频对应的历史质量分数。在训练过程中,将一个第二历史视频作为输入数据、并将该第二历史视频对应的历史质量分数作为输出数据,基于机器学习算法进行训练,以得到第一质量评估模型。
示例地,第二历史视频对应的历史质量分数可以通过以下中的至少一者确定:
根据用户针对第二历史视频的标记分数确定的第一分值;
根据第二历史视频的分辨率确定的第二分值,其中,第二历史视频的第二分值与该第二历史视频的分辨率呈正相关变化关系;
根据第二历史视频内第一对象在画面中所处位置确定的第三分值,其中,第二历史视频的第三分值与第一对象距第一对象所在视频帧图像的画面中心之间的距离呈负相关变化关系;
根据第二历史视频内第二对象占画面总体比例确定的第四分值,其中,第二历史视频的第四分值与第二对象占该第二历史视频的画面总体的比例呈正相关变化关系。
其中,第一分值能够反映用户针对第二历史视频所做的人工标注,能够获知用户角度对该第二历史视频做出的直观评价。第二分值能够反映第二历史视频的清晰程度,且第二历史视频的清晰程度越高该第二历史视频对应的第二分值越高。第三分值能够反映第二历史视频内第一对象(可以是预先设置的,例如,某个人、某个建筑物等)在第二历史视频画面中相对于画面中心的偏离程度,并且,第一对象在第二历史视频中越靠近画面中心(即,第一对象在第二历史视频画面中相对于画面中心的偏离程度小),该第二历史视频对应的第三分值越高。第四分值能够反映第二历史视频内第二对象(可以是预先设置的,例如,某个人、某个建筑物等)在第二历史视频画面中所占的比例,并且,第二对象在第二历史视频中占画面的比例越大,该第二历史视频对应的第四分值越高。
以及,在使用第一分值、第二分值、第三分值和第四分值中的多者确定第二历史视频对应的历史质量分数时,可以设置每一种分值对应的权重,结合实际的分值和分值对应的权重,计算出第二历史视频对应的历史质量分数。
由上所述,第一质量评估模型是基于第二历史视频和第二历史视频对应的历史质量分数训练得到的,也就是说,第一质量评估模型能够基于输入的视频直接得到针对该视频的评估结果,因此,将目标视频片段中包含的视频帧图像输入至第一质量评估模型,得到的第一质量评估模型的输出结果就是目标视频片段对应的质量分数。其中,可以将目标视频片段中包含的视频帧图像的全部输入至第一质量评估模型,以保证第一质量评估模型输出结果的准确性。或者,可以将目标视频片段中包含的视频帧图像的一部分输入至第一质量评估模型,以减轻第一质量评估模型的数据处理压力。
在另一种可能的实施方式中,目标视频片段对应的质量分数可以通过如下方式确定:
将目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,获得第二质量评估模型针对输入至该第二质量评估模型中的每一视频帧图像输出的初始质量分数;
根据初始质量分数,确定与目标视频片段对应的质量分数。
其中,第二质量评估模型通过第三训练数据训练得到,例如,通过第三训练数据对机器学习算法进行训练得到。第三训练数据可以包括历史图像和历史图像对应的历史质量分数。在训练过程中,将一个历史图像作为输入数据、并将该历史图像对应的历史质量分数作为输出数据,基于机器学习算法进行训练,以得到第二质量评估模型。
示例地,历史图像对应的历史质量分数可以通过以下中的至少一者确定:
根据用户针对历史图像的标记分数确定的第五分值;
根据历史图像的分辨率确定的第六分值,其中,历史图像的第六分值与该历史图像的分辨率呈正相关变化关系;
根据历史图像内第三对象在画面中所处位置确定的第七分值,其中,历史图像的第七分值与该历史图像内第三对象距该历史图像的画面中心之间的距离呈负相关变化关系;
根据历史图像内第四对象占画面总体比例确定的第八分值,其中,历史图像的第八分值与该历史图像内第四对象占该历史图像的画面总体的比例呈正相关变化关系。
其中,第五分值能够反映用户针对历史图像所做的人工标注,能够获知用户角度对该历史图像做出的直观评价。第六分值能够反映历史图像的清晰程度,且历史图像的清晰程度越高该历史图像对应的第六分值越高。第七分值能够反映历史图像内第三对象(可以是预先设置的,例如,某个人、某个建筑物等)在历史图像画面中相对于画面中心的偏离程度,并且,第三对象在历史图像中越靠近画面中心(即,第三对象在历史图像画面中相对于画面中心的偏离程度小),该历史图像对应的第七分值越高。第八分值能够反映历史图像内第四对象(可以是预先设置的,例如,某个人、某个建筑物等)在历史图像画面中所占的比例,并且,第四对象在历史图像中占画面的比例越大,该历史图像对应的第八分值越高。
以及,在使用第五分值、第六分值、第七分值和第八分值中的多者确定历史图像对应的历史质量分数时,可以设置每一种分值对应的权重,结合实际的分值和分值对应的权重,计算出历史图像对应的历史质量分数。
由上所述,第二质量评估模型是基于历史图像和历史图像对应的历史质量分数训练得到的,也就是说,第二质量评估模型能够基于输入的单个图像得到针对该图像的评估结果,因此,将目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,得到的第二质量评估模型的输出结果是目标视频片段中输入至第二评估模型中的视频帧图像对应的初始质量分数,也就是单帧视频帧图像对应的初始质量分数。其中,可以将目标视频片段中包含的视频 帧图像的全部分别输入至第二质量评估模型,以保证第二质量评估模型输出结果的准确性。或者,可以将目标视频片段中包含的视频帧图像的一部分分别输入至第二质量评估模型,以减轻第二质量评估模型的数据处理压力。
在得到初始质量分数后,目标视频片段整体对应的质量分数还未知,因此,需要根据这些初始质量分数,确定与目标视频片段对应的质量分数。
示例地,根据初始质量分数,确定与目标视频片段对应的质量分数,可以包括以下中的任意一者:
将初始质量分数的平均值确定为与目标视频片段对应的质量分数;
将初始质量分数中的最大值确定为与目标视频片段对应的质量分数;
将初始质量分数的中位数确定为与目标视频片段对应的质量分数。
在步骤13中,在质量分数展示时间轴上,在目标视频片段对应的时间位置处展示目标视频片段对应的质量分数。
其中,目标视频片段对应的时间位置为目标视频片段在目标视频中出现的时间位置。质量分数展示时间轴是根据目标视频而生成的时间轴,其上的各个时间点分别对应于目标视频中的相应时间点。
在一种可能的实施方式中,步骤13可以包括以下中的任意一者:
在质量分数展示时间轴上,在目标视频片段对应的时间位置处、以数字形式展示目标视频片段对应的质量分数;
在质量分数展示时间轴上,在目标视频片段对应的时间位置处、且在目标视频片段的质量分数对应的分数位置处,以点状标记展示目标视频片段对应的质量分数;
在质量分数展示时间轴上,在目标视频片段对应的时间位置处,以条形标记展示目标视频片段对应的质量分数,其中,目标视频片段对应的条形标记的最大分数为目标视频片段对应的质量分数。
示例地,若目标视频片段为目标视频中50s~70s所构成的视频片段,且目标视频片段对应的质量分数为0.6,则质量分数的展示方式可以如以下几种示例。
图2A示出了在质量分数展示时间轴上,在目标视频片段对应的时间位置处、以数字形式展示目标视频片段对应的质量分数的一种示例性的展示示意图。
图2B示出了在质量分数展示时间轴上,在目标视频片段对应的时间位置处、且在目标视频片段的质量分数对应的分数位置处,以点状标记展示目标视频片段对应的质量分数的一种示例性的展示示意图。
图2C示出了在质量分数展示时间轴上,在目标视频片段对应的时间位置处,以条形标记展示目标视频片段对应的质量分数的一种示例性的展示示意图。
如上所述,目标视频可以被划分为包括目标视频片段在内的多个视频片段,因此,在进行质量分数展示时,可以展示多个视频片段各自对应的质量分数。其中,每一视频片段对应的质量分数的确定均可参照确定目标视频片段对应的质量分数的过程,此处不再赘述。从而,在上述实施方式的基础上,能够得到目标视频中多个视频片段对应的质量分数,并将这些视频片段对应的质量分数进行展示,以直观展示目标视频中多个视频片段各自对应的质量分数。
示例地,若目标视频中0~50s所构成的视频片段对应的质量分数为0.8,目标视频中50s~70s所构成的视频片段对应的质量分数为0.6,目标视频中70s~90s所构成的视频片段对应 的质量分数为0.4,目标视频中90s~150s所构成的视频片段对应的质量分数为0.3,则质量分数的示例性的展示方式可以如图3A、3B、3C所示(依次对应于上述图2A、2B、2C的展示方式)。并且,在图3B中,展示目标视频中各个视频片段对应的质量分数之后,还可以根据已展示的点状标记依次连成连线,更加直观地展示质量分数的变化情况。
通过上述技术方案,对目标视频进行划分,得到目标视频片段,以及,根据目标视频片段中包含的视频帧图像,确定目标视频片段对应的质量分数,并在质量分数展示时间轴上,在目标视频片段对应的时间位置处展示目标视频片段对应的质量分数。也就是说,在得到目标视频中的目标视频片段对应的质量分数后,能够在质量分数展示时间轴的相应位置展示目标视频片段对应的质量分数。由此,能够为用户提供有关于目标视频片段质量分数的可视化展示结果,供用户查看,方便用户快速获知目标视频片段对应的质量分数,从而为用户的视频片段挑选提供参考,节省用户查看目标视频片段所花费的时间。另外,基于上述方案,能够确定并展示目标视频中多个视频片段各自对应的质量分数,从而能够形成针对多个视频片段的质量分数的直观比较,为用户从目标视频中挑选视频片段提供参考依据,方便用户快速挑选视频片段。
在一种可能的实施方式中,质量分数展示时间轴可以为目标视频对应的视频编辑时间轴,也就是在目标视频的编辑界面展示目标视频中目标视频片段的质量分数。以及,目标视频片段对应有在目标视频中的起始分割时间点和结束分割时间点,也就是目标视频片段在目标视频中对应的起始、结束时间。在这一实施方式中,除图1所示各步骤的基础上,本公开还可以包括以下步骤:
在目标视频对应的视频编辑时间轴上,在目标视频片段对应的起始分割时间点和结束分割时间点展示目标视频片段对应的切割标记。
其中,切割标记用于向用户展示目标视频片段在目标视频中的起点和终点,便于用户知晓目标视频片段在目标视频中具体的位置、对应的画面等。并且,在实际应用中,切割标记可以为用户提供一键裁剪的功能,即,在切割标记被点击的情况下,从目标视频中将目标视频片段中复制(或剪切)出来,形成与目标视频片段对应的视频文件。
另外,如上所述,对目标视频进行划分得到包括目标视频在内的多个视频片段,每一视频片段均可按照上述方式展示与该视频片段对应的切割标记。从而,在某个视频片段对应的切割标记被点击时,就可以形成对应于这个被点击的切割标记对应视频片段的文件。
采用上述方式,通过切割标记展示目标视频片段在目标视频中对应的位置,为用户展示目标视频片段的详细内容,方便用户根据自身需求对视频片段进行保存。
在一种可能的实施方式中,在图1所示各步骤的基础上,本公开的方法还可以包括以下步骤:
响应于接收到针对目标视频的剪辑指令,若目标视频片段对应的质量分数高于目标视频中其他视频片段对应的质量分数,将目标视频片段确定为用于视频拼接的备选素材。
若目标视频片段对应的质量分数高于目标视频中其他视频片段对应的质量分数,说明目标视频片段是目标视频中质量分数最高的视频片段,具有较高的可利用性。因此,可以将目标视频片段确定为用于视频拼接的备选素材,以为后续的视频拼接提供可用素材。
在一种可能的实施方式中,在上述步骤的基础上,还可以将包括目标视频片段在内的多个备选素材合成为目标拼接视频。
其中,上述多个备选素材中除目标视频片段之外的其他备选素材,可以是目标视频中除目标视频片段之外的其他视频片段,或者,可以取自除目标视频之外的其他视频中,本公开对此不进行限定。示例地,可以参照上文中给出的方式得到目标视频中各个视频片段各自对应的质量分数,并将质量分数较高的前几者作为备选素材,以合成目标拼接视频。再例如,可以参照上文中给出的方式,对包括目标视频在内的多个视频进行处理,并从这多个视频的每个视频中,确定出质量分数最高的视频片段作为备选素材,以合成目标拼接视频。这样,能够自动为用户生成拼接视频,无需用户手动操作,提升用户的视频剪辑效率。
另外,响应于接收到针对目标视频的剪辑指令,若目标视频片段对应的质量分数高于目标视频中其他视频片段对应的质量分数,还可以直接将目标视频中除目标视频片段之外的其他内容删除,而只保留目标视频片段。示例地,若目标视频片段为单帧视频帧图像,上述步骤相当于仅保留目标视频中质量分数最高的视频帧图像,即目标视频对应的“高光时刻”。再例如,若目标视频片段为多帧视频帧图像,上述步骤相当于保留目标视频中的最高分片段。这样,能够自动为用户保留视频中质量最高的部分,无需用户逐帧查看,提升用户的视频剪辑效率。
在实际的应用场景中,客户端的交互页面可以展示如图4A至图4E所示。在图4A中,展示有视频编辑页面,其中展示有目标视频对应的视频编辑时间轴,处于页面右下角的按钮Bu1用于触发对目标视频的“智能评估”功能,例如,展示目标视频片段的质量分数、保留目标视频中质量分数高的视频片段等。当用户点击图4A中的按钮Bu1,表示用户对当前视频存在智能评估的需求。此时页面上所展示的视频就是目标视频,进而可以执行本公开提供的相关步骤,对目标视频进行处理,例如,确定目标视频片段对应的质量分数,或者,若目标视频片段对应的质量分数高于目标视频中其他视频片段对应的质量分数,直接将目标视频中除目标视频片段之外的其他内容删除,而只保留目标视频片段,等等。在执行上述步骤的过程中客户端显示内容可以如图4B所示,其中,图4B中心的百分数用于表示数据处理的进度,当百分数达到100%时,当前的数据处理完成,即中台对目标视频的智能评估已经完成,此时客户端显示内容可以如图4C所示。之后,可以展示具体的智能评估的结果,如图4D或图4E所示,在图4D和图4E中,曲线表示目标视频中各个视频片段的质量分数所构成的质量分数曲线。以及,在图4D中,除展示有质量分数曲线之外,还在质量分数曲线下方展示针对目标视频剪辑后所保留的视频片段。在图4E中,除展示有质量分数曲线之外,还在质量分数曲线中几个较高质量分数位置处展示有对应的视频帧图像,即目标视频对应的“高光时刻”。
图5是根据本公开的一种实施方式提供的视频处理装置的框图。如图5所示,该装置40可以包括:
划分模块41,用于对目标视频进行划分,得到目标视频片段;
第一确定模块42,用于根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;
第一展示模块43,用于在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
可选地,所述划分模块41包括:
切片子模块,用于将所述目标视频中处于不同时间位置的多帧目标视频帧图像和所述多帧目标视频帧图像在所述目标视频中出现的时间输入至切片算法,获得所述切片算法输出的目标分割时间点,所述切片算法通过第一训练数据对机器学习算法进行训练得到,所述第一训练数据包括第一历史视频和用于对所述第一历史视频进行划分的历史分割时间点;
划分子模块,用于根据所述目标分割时间点,对所述目标视频进行划分,得到包括所述目标视频片段在内的多个视频片段。
可选地,所述第一确定模块42用于通过如下方式确定目标视频片段对应的质量分数:
将所述目标视频片段中包含的视频帧图像输入至第一质量评估模型,获得所述目标视频片段对应的质量分数,其中,所述第一质量评估模型通过第二训练数据训练得到,所述第二训练数据包括第二历史视频和所述第二历史视频对应的历史质量分数。
可选地,所述第二历史视频对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述第二历史视频的标记分数确定的第一分值;
根据所述第二历史视频的分辨率确定的第二分值,其中,第二历史视频的第二分值与该第二历史视频的分辨率呈正相关变化关系;
根据所述第二历史视频内第一对象在画面中所处位置确定的第三分值,其中,第二历史视频的第三分值与所述第一对象距所述第一对象所在视频帧图像的画面中心之间的距离呈负相关变化关系;
根据所述第二历史视频内第二对象占画面总体比例确定的第四分值,其中,第二历史视频的第四分值与所述第二对象占该第二历史视频的画面总体的比例呈正相关变化关系。
可选地,所述第一确定模块42包括:
处理子模块,用于将所述目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,获得所述第二质量评估模型针对输入至该第二质量评估模型中的每一视频帧图像输出的初始质量分数,其中,所述第二质量评估模型通过第三训练数据训练得到,所述第三训练数据包括历史图像和所述历史图像对应的历史质量分数;
确定子模块,用于根据所述初始质量分数,确定与所述目标视频片段对应的质量分数。
可选地,所述确定子模块用于通过以下中的任意一者确定与所述目标视频片段对应的质量分数:
将所述初始质量分数的平均值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数中的最大值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数的中位数确定为与所述目标视频片段对应的质量分数。
可选地,所述历史图像对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述历史图像的标记分数确定的第五分值;
根据所述历史图像的分辨率确定的第六分值,其中,历史图像的第六分值与该历史图像的分辨率呈正相关变化关系;
根据所述历史图像内第三对象在画面中所处位置确定的第七分值,其中,历史图像的第七分值与该历史图像内第三对象距该历史图像的画面中心之间的距离呈负相关变化关系;
根据所述历史图像内第四对象占画面总体比例确定的第八分值,其中,历史图像的第八分值与该历史图像内第四对象占该历史图像的画面总体的比例呈正相关变化关系。
可选地,所述目标视频被划分为包括所述目标视频片段在内的多个视频片段;
所述装置40还包括:
第二确定模块,用于响应于接收到针对所述目标视频的剪辑指令,若所述目标视频片段对应的质量分数高于所述目标视频中其他视频片段对应的质量分数,将所述目标视频片段确定为用于视频拼接的备选素材。
可选地,所述装置40还包括:
合成模块,用于将包括所述目标视频片段在内的多个备选素材合成为目标拼接视频。
可选地,所述第一展示模块43包括以下中的任意一者:
第一展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、以数字形式展示所述目标视频片段对应的质量分数;
第二展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、且在所述目标视频片段的质量分数对应的分数位置处,以点状标记展示所述目标视频片段对应的质量分数;
第三展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处,以条形标记展示所述目标视频片段对应的质量分数,其中,目标视频片段对应的条形标记的最大分数为所述目标视频片段对应的质量分数。
可选地,所述质量分数展示时间轴为所述目标视频对应的视频编辑时间轴。
可选地,所述目标视频片段对应有在所述目标视频中的起始分割时间点和结束分割时间点;
所述装置40还包括:
第二展示模块,用于在所述目标视频对应的视频编辑时间轴上,在所述目标视频片段对应的所述起始分割时间点和所述结束分割时间点展示所述目标视频片段对应的切割标记。
关于上述实施例中的装置,其中各个模块执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
下面参考图6,其示出了适于用来实现本公开实施例的电子设备600的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理,Personal Digital Assistant)、PAD(平板电脑,Portable Android Device)、PMP(便携式多媒体播放器,Personal Multimedia Player)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字电视(TV)、台式计算机等等的固定终端。图6示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图6所示,电子设备600可以包括处理装置(例如中央处理器、图形处理器等)601,其可以根据存储在只读存储器(Read-Only Memory,ROM)602中的程序或者从存储装置608加载到随机访问存储器(Random Access Memory,RAM)603中的程序而执行各种适当的动作和处理。在RAM 603中,还存储有电子设备600操作所需的各种程序和数据。处理装置601、ROM 602以及RAM 603通过总线604彼此相连。输入/输出(Input/Output,I/O)接口605也连接至总线604。
通常,以下装置可以连接至I/O接口605:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置606;包括例如液晶显示器(Liquid Crystal Display,LCD)、扬声器、振动器等的输出装置607;包括例如磁带、硬盘等的存储装置608;以及通信装置609。通信装置609可以允许电子设备600与其他设备进行无线或有线通信以交换数据。 虽然图6示出了具有各种装置的电子设备600,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置609从网络上被下载和安装,或者从存储装置608被安装,或者从ROM 602被安装。在该计算机程序被处理装置601执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(Electrical Programmable ROM,EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(Compact Disc ROM,CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频,Radio Frequency)等等,或者上述的任意合适的组合。
在一些实施方式中,终端、服务器可以利用诸如HTTP(HyperText Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(Local Area Network,“LAN”),广域网(Wide Area Network,“WAN”),网际网(例如,互联网)以及端对端网络(例如,自组织(ADaptive Heuristic for Opponent Classification,ad hoc)端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:对目标视频进行划分,得到目标视频片段;根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言——诸如“C”语言或类似的程序设计语言。程序代码可以完 全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)——连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的模块可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,模块的名称在某种情况下并不构成对该模块本身的限定,例如,划分模块还可以被描述为“用于对目标视频进行划分,得到目标视频片段的模块”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(Field Programmable Gate Array,FPGA)、专用集成电路(Application Specific Integrated Circuit,ASIC)、专用标准产品(Application Specific Standard Parts,ASSP)、片上系统(System on Chip,SOC)、复杂可编程逻辑设备(Complex Programming Logic Device,CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
根据本公开的一个或多个实施例,提供了一种视频处理方法,所述方法包括:
对目标视频进行划分,得到目标视频片段;
根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;
在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述对目标视频进行划分,得到目标视频片段,包括:
将所述目标视频中处于不同时间位置的多帧目标视频帧图像和所述多帧目标视频帧图像在所述目标视频中出现的时间输入至切片算法,获得所述切片算法输出的目标分割时间点, 所述切片算法通过第一训练数据对机器学习算法进行训练得到,所述第一训练数据包括第一历史视频和用于对所述第一历史视频进行划分的历史分割时间点;
根据所述目标分割时间点,对所述目标视频进行划分,得到包括所述目标视频片段在内的多个视频片段。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述目标视频片段对应的质量分数通过如下方式确定:
将所述目标视频片段中包含的视频帧图像输入至第一质量评估模型,获得所述目标视频片段对应的质量分数,其中,所述第一质量评估模型通过第二训练数据训练得到,所述第二训练数据包括第二历史视频和所述第二历史视频对应的历史质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述第二历史视频对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述第二历史视频的标记分数确定的第一分值;
根据所述第二历史视频的分辨率确定的第二分值,其中,第二历史视频的第二分值与该第二历史视频的分辨率呈正相关变化关系;
根据所述第二历史视频内第一对象在画面中所处位置确定的第三分值,其中,第二历史视频的第三分值与所述第一对象距所述第一对象所在视频帧图像的画面中心之间的距离呈负相关变化关系;
根据所述第二历史视频内第二对象占画面总体比例确定的第四分值,其中,第二历史视频的第四分值与所述第二对象占该第二历史视频的画面总体的比例呈正相关变化关系。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述目标视频片段对应的质量分数通过如下方式确定:
将所述目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,获得所述第二质量评估模型针对输入至该第二质量评估模型中的每一视频帧图像输出的初始质量分数,其中,所述第二质量评估模型通过第三训练数据训练得到,所述第三训练数据包括历史图像和所述历史图像对应的历史质量分数;
根据所述初始质量分数,确定与所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述根据所述初始质量分数,确定与所述目标视频片段对应的质量分数,包括以下中的任意一者:
将所述初始质量分数的平均值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数中的最大值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数的中位数确定为与所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述历史图像对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述历史图像的标记分数确定的第五分值;
根据所述历史图像的分辨率确定的第六分值,其中,历史图像的第六分值与该历史图像的分辨率呈正相关变化关系;
根据所述历史图像内第三对象在画面中所处位置确定的第七分值,其中,历史图像的第七分值与该历史图像内第三对象距该历史图像的画面中心之间的距离呈负相关变化关系;
根据所述历史图像内第四对象占画面总体比例确定的第八分值,其中,历史图像的第八分值与该历史图像内第四对象占该历史图像的画面总体的比例呈正相关变化关系。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述目标视频被划分为包括所述目标视频片段在内的多个视频片段;
所述方法还包括:
响应于接收到针对所述目标视频的剪辑指令,若所述目标视频片段对应的质量分数高于所述目标视频中其他视频片段对应的质量分数,将所述目标视频片段确定为用于视频拼接的备选素材。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述方法还包括:
将包括所述目标视频片段在内的多个备选素材合成为目标拼接视频。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,包括以下中的任意一者:
在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、以数字形式展示所述目标视频片段对应的质量分数;
在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、且在所述目标视频片段的质量分数对应的分数位置处,以点状标记展示所述目标视频片段对应的质量分数;
在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处,以条形标记展示所述目标视频片段对应的质量分数,其中,目标视频片段对应的条形标记的最大分数为所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述质量分数展示时间轴为所述目标视频对应的视频编辑时间轴。
根据本公开的一个或多个实施例,还提供了一种视频处理方法,其中,所述目标视频片段对应有在所述目标视频中的起始分割时间点和结束分割时间点;
所述方法还包括:
在所述目标视频对应的视频编辑时间轴上,在所述目标视频片段对应的所述起始分割时间点和所述结束分割时间点展示所述目标视频片段对应的切割标记。
根据本公开的一个或多个实施例,提供了一种视频处理装置,所述装置包括:
划分模块,用于对目标视频进行划分,得到目标视频片段;
第一确定模块,用于根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;
第一展示模块,用于在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述划分模块包括:
切片子模块,用于将所述目标视频中处于不同时间位置的多帧目标视频帧图像和所述多帧目标视频帧图像在所述目标视频中出现的时间输入至切片算法,获得所述切片算法输出的 目标分割时间点,所述切片算法通过第一训练数据对机器学习算法进行训练得到,所述第一训练数据包括第一历史视频和用于对所述第一历史视频进行划分的历史分割时间点;
划分子模块,用于根据所述目标分割时间点,对所述目标视频进行划分,得到包括所述目标视频片段在内的多个视频片段。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述第一确定模块用于通过如下方式确定目标视频片段对应的质量分数:
将所述目标视频片段中包含的视频帧图像输入至第一质量评估模型,获得所述目标视频片段对应的质量分数,其中,所述第一质量评估模型通过第二训练数据训练得到,所述第二训练数据包括第二历史视频和所述第二历史视频对应的历史质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述第二历史视频对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述第二历史视频的标记分数确定的第一分值;
根据所述第二历史视频的分辨率确定的第二分值,其中,第二历史视频的第二分值与该第二历史视频的分辨率呈正相关变化关系;
根据所述第二历史视频内第一对象在画面中所处位置确定的第三分值,其中,第二历史视频的第三分值与所述第一对象距所述第一对象所在视频帧图像的画面中心之间的距离呈负相关变化关系;
根据所述第二历史视频内第二对象占画面总体比例确定的第四分值,其中,第二历史视频的第四分值与所述第二对象占该第二历史视频的画面总体的比例呈正相关变化关系。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述第一确定模块包括:
处理子模块,用于将所述目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,获得所述第二质量评估模型针对输入至该第二质量评估模型中的每一视频帧图像输出的初始质量分数,其中,所述第二质量评估模型通过第三训练数据训练得到,所述第三训练数据包括历史图像和所述历史图像对应的历史质量分数;
确定子模块,用于根据所述初始质量分数,确定与所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述确定子模块用于通过以下中的任意一者确定与所述目标视频片段对应的质量分数:
将所述初始质量分数的平均值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数中的最大值确定为与所述目标视频片段对应的质量分数;
将所述初始质量分数的中位数确定为与所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述历史图像对应的历史质量分数通过以下中的至少一者确定:
根据用户针对所述历史图像的标记分数确定的第五分值;
根据所述历史图像的分辨率确定的第六分值,其中,历史图像的第六分值与该历史图像的分辨率呈正相关变化关系;
根据所述历史图像内第三对象在画面中所处位置确定的第七分值,其中,历史图像的第七分值与该历史图像内第三对象距该历史图像的画面中心之间的距离呈负相关变化关系;
根据所述历史图像内第四对象占画面总体比例确定的第八分值,其中,历史图像的第八分值与该历史图像内第四对象占该历史图像的画面总体的比例呈正相关变化关系。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述目标视频被划分为包括所述目标视频片段在内的多个视频片段;
所述装置还包括:
第二确定模块,用于响应于接收到针对所述目标视频的剪辑指令,若所述目标视频片段对应的质量分数高于所述目标视频中其他视频片段对应的质量分数,将所述目标视频片段确定为用于视频拼接的备选素材。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述装置还包括:
合成模块,用于将包括所述目标视频片段在内的多个备选素材合成为目标拼接视频。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述第一展示模块包括以下中的任意一者:
第一展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、以数字形式展示所述目标视频片段对应的质量分数;
第二展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、且在所述目标视频片段的质量分数对应的分数位置处,以点状标记展示所述目标视频片段对应的质量分数;
第三展示子模块,用于在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处,以条形标记展示所述目标视频片段对应的质量分数,其中,目标视频片段对应的条形标记的最大分数为所述目标视频片段对应的质量分数。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述质量分数展示时间轴为所述目标视频对应的视频编辑时间轴。
根据本公开的一个或多个实施例,还提供了一种视频处理装置,其中,所述目标视频片段对应有在所述目标视频中的起始分割时间点和结束分割时间点;
所述装置还包括:
第二展示模块,用于在所述目标视频对应的视频编辑时间轴上,在所述目标视频片段对应的所述起始分割时间点和所述结束分割时间点展示所述目标视频片段对应的切割标记。
根据本公开的一个或多个实施例,还提供了一种计算机程序产品,程序产品包括:计算机程序,该计算机程序被处理装置执行时实现本公开任意实施例所述方法的步骤。
根据本公开的一个或多个实施例,还提供了一种计算机程序,该计算机程序被处理装置执行时实现本公开任意实施例所述方法的步骤。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围 的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。关于上述实施例中的装置,其中各个模块执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
Claims (17)
- 一种视频处理方法,其特征在于,所述方法包括:对目标视频进行划分,得到目标视频片段;根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
- 根据权利要求1所述的方法,其特征在于,所述对目标视频进行划分,得到目标视频片段,包括:将所述目标视频中处于不同时间位置的多帧目标视频帧图像和所述多帧目标视频帧图像在所述目标视频中出现的时间输入至切片算法,获得所述切片算法输出的目标分割时间点,所述切片算法通过第一训练数据对机器学习算法进行训练得到,所述第一训练数据包括第一历史视频和用于对所述第一历史视频进行划分的历史分割时间点;根据所述目标分割时间点,对所述目标视频进行划分,得到包括所述目标视频片段在内的多个视频片段。
- 根据权利要求1或2所述的方法,其特征在于,所述目标视频片段对应的质量分数通过如下方式确定:将所述目标视频片段中包含的视频帧图像输入至第一质量评估模型,获得所述目标视频片段对应的质量分数,其中,所述第一质量评估模型通过第二训练数据训练得到,所述第二训练数据包括第二历史视频和所述第二历史视频对应的历史质量分数。
- 根据权利要求3所述的方法,其特征在于,所述第二历史视频对应的历史质量分数通过以下中的至少一者确定:根据用户针对所述第二历史视频的标记分数确定的第一分值;根据所述第二历史视频的分辨率确定的第二分值,其中,第二历史视频的第二分值与该第二历史视频的分辨率呈正相关变化关系;根据所述第二历史视频内第一对象在画面中所处位置确定的第三分值,其中,第二历史视频的第三分值与所述第一对象距所述第一对象所在视频帧图像的画面中心之间的距离呈负相关变化关系;根据所述第二历史视频内第二对象占画面总体比例确定的第四分值,其中,第二历史视频的第四分值与所述第二对象占该第二历史视频的画面总体的比例呈正相关变化关系。
- 根据权利要求1或2所述的方法,其特征在于,所述目标视频片段对应的质量分数通过如下方式确定:将所述目标视频片段中包含的视频帧图像分别输入至第二质量评估模型,获得所述第二质量评估模型针对输入至该第二质量评估模型中的每一视频帧图像输出的初始质量分数,其中,所述第二质量评估模型通过第三训练数据训练得到,所述第三训练数据包括历史图像和所述历史图像对应的历史质量分数;根据所述初始质量分数,确定与所述目标视频片段对应的质量分数。
- 根据权利要求5所述的方法,其特征在于,所述根据所述初始质量分数,确定与所述目标视频片段对应的质量分数,包括以下中的任意一者:将所述初始质量分数的平均值确定为与所述目标视频片段对应的质量分数;将所述初始质量分数中的最大值确定为与所述目标视频片段对应的质量分数;将所述初始质量分数的中位数确定为与所述目标视频片段对应的质量分数。
- 根据权利要求5或6所述的方法,其特征在于,所述历史图像对应的历史质量分数通过以下中的至少一者确定:根据用户针对所述历史图像的标记分数确定的第五分值;根据所述历史图像的分辨率确定的第六分值,其中,历史图像的第六分值与该历史图像的分辨率呈正相关变化关系;根据所述历史图像内第三对象在画面中所处位置确定的第七分值,其中,历史图像的第七分值与该历史图像内第三对象距该历史图像的画面中心之间的距离呈负相关变化关系;根据所述历史图像内第四对象占画面总体比例确定的第八分值,其中,历史图像的第八分值与该历史图像内第四对象占该历史图像的画面总体的比例呈正相关变化关系。
- 根据权利要求1-7任一项所述的方法,其特征在于,所述目标视频被划分为包括所述目标视频片段在内的多个视频片段;所述方法还包括:响应于接收到针对所述目标视频的剪辑指令,若所述目标视频片段对应的质量分数高于所述目标视频中其他视频片段对应的质量分数,将所述目标视频片段确定为用于视频拼接的备选素材。
- 根据权利要求8所述的方法,其特征在于,所述方法还包括:将包括所述目标视频片段在内的多个备选素材合成为目标拼接视频。
- 根据权利要求1-9任一项所述的方法,其特征在于,所述在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,包括以下中的任意一者:在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、以数字形式展示所述目标视频片段对应的质量分数;在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处、且在所述目标视频片段的质量分数对应的分数位置处,以点状标记展示所述目标视频片段对应的质量分数;在所述质量分数展示时间轴上,在所述目标视频片段对应的时间位置处,以条形标记展示所述目标视频片段对应的质量分数,其中,目标视频片段对应的条形标记的最大分数为所述目标视频片段对应的质量分数。
- 根据权利要求1-10任一项所述的方法,其特征在于,所述质量分数展示时间轴为所述目标视频对应的视频编辑时间轴。
- 根据权利要求1-11任一项所述的方法,其特征在于,所述目标视频片段对应有在所述目标视频中的起始分割时间点和结束分割时间点;所述方法还包括:在所述目标视频对应的视频编辑时间轴上,在所述目标视频片段对应的所述起始分割时间点和所述结束分割时间点展示所述目标视频片段对应的切割标记。
- 一种视频处理装置,其特征在于,所述装置包括:划分模块,用于对目标视频进行划分,得到目标视频片段;第一确定模块,用于根据所述目标视频片段中包含的视频帧图像,确定所述目标视频片段对应的质量分数;第一展示模块,用于在质量分数展示时间轴上,在所述目标视频片段对应的时间位置处展示所述目标视频片段对应的质量分数,其中,所述目标视频片段对应的时间位置为所述目标视频片段在所述目标视频中出现的时间位置。
- 一种计算机可读介质,其上存储有计算机程序,其特征在于,该程序被处理装置执行时实现权利要求1-12中任一项所述方法的步骤。
- 一种电子设备,其特征在于,包括:存储装置,其上存储有计算机程序;处理装置,用于执行所述存储装置中的所述计算机程序,以实现权利要求1-12中任一项所述方法的步骤。
- 一种计算机程序产品,包括计算机程序,所述计算机程序被处理装置执行时实现权利要求1-12中任一项所述方法的步骤。
- 一种计算机程序,所述计算机程序被处理装置执行时实现权利要求1-12中任一项所述方法的步骤。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/819,238 US11996124B2 (en) | 2020-02-11 | 2022-08-11 | Video processing method, apparatus, readable medium and electronic device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010087010.9A CN113259601A (zh) | 2020-02-11 | 2020-02-11 | 视频处理方法、装置、可读介质和电子设备 |
| CN202010087010.9 | 2020-02-11 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/819,238 Continuation US11996124B2 (en) | 2020-02-11 | 2022-08-11 | Video processing method, apparatus, readable medium and electronic device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021160141A1 true WO2021160141A1 (zh) | 2021-08-19 |
Family
ID=77219559
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/076409 Ceased WO2021160141A1 (zh) | 2020-02-11 | 2021-02-09 | 视频处理方法、装置、可读介质和电子设备 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11996124B2 (zh) |
| CN (1) | CN113259601A (zh) |
| WO (1) | WO2021160141A1 (zh) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115904168A (zh) * | 2022-11-18 | 2023-04-04 | Oppo广东移动通信有限公司 | 基于多设备的影像素材处理方法及相关装置 |
| CN118283364A (zh) * | 2022-12-29 | 2024-07-02 | 北京字跳网络技术有限公司 | 在线视频编辑方法、装置、电子设备及存储介质 |
| CN117370602B (zh) * | 2023-04-24 | 2024-08-06 | 深圳云视智景科技有限公司 | 视频处理方法、装置、设备与计算机存储介质 |
| CN119342201B (zh) * | 2023-07-21 | 2025-11-04 | 北京字跳网络技术有限公司 | 一种全景视频的评估方法、装置和电子设备 |
| CN119172594B (zh) * | 2024-09-27 | 2025-11-25 | 维沃移动通信有限公司 | 视频处理方法及装置 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103945219A (zh) * | 2014-04-30 | 2014-07-23 | 北京邮电大学 | 一种网络侧视频质量监测系统 |
| CN104780454A (zh) * | 2015-03-17 | 2015-07-15 | 联想(北京)有限公司 | 一种多媒体数据播放控制方法及电子设备 |
| CN105493512A (zh) * | 2014-12-14 | 2016-04-13 | 深圳市大疆创新科技有限公司 | 一种视频处理方法、视频处理装置及显示装置 |
| CN106993227A (zh) * | 2016-01-20 | 2017-07-28 | 腾讯科技(北京)有限公司 | 一种进行信息展示的方法和装置 |
| CN108156528A (zh) * | 2017-12-18 | 2018-06-12 | 北京奇艺世纪科技有限公司 | 一种视频处理方法及装置 |
| CN109587578A (zh) * | 2018-12-21 | 2019-04-05 | 麒麟合盛网络技术股份有限公司 | 视频片段的处理方法及装置 |
| CN109963164A (zh) * | 2017-12-14 | 2019-07-02 | 北京搜狗科技发展有限公司 | 一种在视频中查询对象的方法、装置和设备 |
| CN110753246A (zh) * | 2018-07-23 | 2020-02-04 | 优视科技有限公司 | 视频播放方法、客户端、服务器及系统 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7483618B1 (en) * | 2003-12-04 | 2009-01-27 | Yesvideo, Inc. | Automatic editing of a visual recording to eliminate content of unacceptably low quality and/or very little or no interest |
| US20160037176A1 (en) * | 2014-07-30 | 2016-02-04 | Arris Enterprises, Inc. | Automatic and adaptive selection of profiles for adaptive bit rate streaming |
| CN104778230B (zh) * | 2015-03-31 | 2018-11-06 | 北京奇艺世纪科技有限公司 | 一种视频数据切分模型的训练、视频数据切分方法和装置 |
| CN106210902B (zh) * | 2016-07-06 | 2019-06-11 | 华东师范大学 | 一种基于弹幕评论数据的影视片段剪辑方法 |
| US20190095961A1 (en) * | 2017-09-22 | 2019-03-28 | Facebook, Inc. | Applying a trained model for predicting quality of a content item along a graduated scale |
| US10410060B2 (en) * | 2017-12-14 | 2019-09-10 | Google Llc | Generating synthesis videos |
| CN108198177A (zh) * | 2017-12-29 | 2018-06-22 | 广东欧珀移动通信有限公司 | 图像获取方法、装置、终端及存储介质 |
| CN109819338B (zh) * | 2019-02-22 | 2021-09-14 | 影石创新科技股份有限公司 | 一种视频自动剪辑方法、装置及便携式终端 |
| US11521386B2 (en) * | 2019-07-26 | 2022-12-06 | Meta Platforms, Inc. | Systems and methods for predicting video quality based on objectives of video producer |
| CN110650374B (zh) * | 2019-08-16 | 2022-03-25 | 咪咕文化科技有限公司 | 剪辑方法、电子设备和计算机可读存储介质 |
| CN110516749A (zh) * | 2019-08-29 | 2019-11-29 | 网易传媒科技(北京)有限公司 | 模型训练方法、视频处理方法、装置、介质和计算设备 |
| US11350169B2 (en) * | 2019-11-13 | 2022-05-31 | Netflix, Inc. | Automatic trailer detection in multimedia content |
-
2020
- 2020-02-11 CN CN202010087010.9A patent/CN113259601A/zh active Pending
-
2021
- 2021-02-09 WO PCT/CN2021/076409 patent/WO2021160141A1/zh not_active Ceased
-
2022
- 2022-08-11 US US17/819,238 patent/US11996124B2/en active Active
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103945219A (zh) * | 2014-04-30 | 2014-07-23 | 北京邮电大学 | 一种网络侧视频质量监测系统 |
| CN105493512A (zh) * | 2014-12-14 | 2016-04-13 | 深圳市大疆创新科技有限公司 | 一种视频处理方法、视频处理装置及显示装置 |
| CN104780454A (zh) * | 2015-03-17 | 2015-07-15 | 联想(北京)有限公司 | 一种多媒体数据播放控制方法及电子设备 |
| CN106993227A (zh) * | 2016-01-20 | 2017-07-28 | 腾讯科技(北京)有限公司 | 一种进行信息展示的方法和装置 |
| CN109963164A (zh) * | 2017-12-14 | 2019-07-02 | 北京搜狗科技发展有限公司 | 一种在视频中查询对象的方法、装置和设备 |
| CN108156528A (zh) * | 2017-12-18 | 2018-06-12 | 北京奇艺世纪科技有限公司 | 一种视频处理方法及装置 |
| CN110753246A (zh) * | 2018-07-23 | 2020-02-04 | 优视科技有限公司 | 视频播放方法、客户端、服务器及系统 |
| CN109587578A (zh) * | 2018-12-21 | 2019-04-05 | 麒麟合盛网络技术股份有限公司 | 视频片段的处理方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| US11996124B2 (en) | 2024-05-28 |
| US20220383910A1 (en) | 2022-12-01 |
| CN113259601A (zh) | 2021-08-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11996124B2 (en) | Video processing method, apparatus, readable medium and electronic device | |
| CN109168026B (zh) | 即时视频显示方法、装置、终端设备及存储介质 | |
| US20260038533A1 (en) | Video processing method and apparatus, electronic device, and computer readable storage medium | |
| US20190130185A1 (en) | Visualization of Tagging Relevance to Video | |
| CN116137662B (zh) | 页面展示方法及装置、电子设备、存储介质和程序产品 | |
| CN113886707B (zh) | 百科信息确定方法、显示方法、装置、设备和介质 | |
| WO2021057740A1 (zh) | 视频生成方法、装置、电子设备和计算机可读介质 | |
| WO2022042389A1 (zh) | 搜索结果的展示方法、装置、可读介质和电子设备 | |
| US11785195B2 (en) | Method and apparatus for processing three-dimensional video, readable storage medium and electronic device | |
| US12067053B2 (en) | Media file processing method, device, readable medium, and electronic apparatus | |
| KR20220103112A (ko) | 비디오 생성 방법 및 장치, 전자 장치, 및 컴퓨터 판독가능 매체 | |
| US10153003B2 (en) | Method, system, and apparatus for generating video content | |
| WO2021197023A1 (zh) | 多媒体资源筛选方法、装置、电子设备及计算机存储介质 | |
| US9961275B2 (en) | Method, system, and apparatus for operating a kinetic typography service | |
| CN114598815B (zh) | 一种拍摄方法、装置、电子设备和存储介质 | |
| WO2024165010A1 (zh) | 信息生成方法、信息显示方法、装置、设备和存储介质 | |
| EP4554236A1 (en) | Information display method and apparatus, electronic device, and computer readable medium | |
| JP2025521195A (ja) | テキスト素材取得方法、装置、機器、媒体及びプログラム製品 | |
| CN114117127B (zh) | 视频生成方法、装置、可读介质及电子设备 | |
| WO2023195914A2 (zh) | 处理方法、装置、终端设备及介质 | |
| US20180077460A1 (en) | Method, System, and Apparatus for Providing Video Content Recommendations | |
| WO2025036409A9 (zh) | 媒体内容处理方法、设备、存储介质及程序产品 | |
| WO2024169881A1 (zh) | 视频处理方法、装置、设备及介质 | |
| CN111314777A (zh) | 视频生成方法及装置、计算机存储介质、电子设备 | |
| WO2023174066A1 (zh) | 视频生成方法、装置、电子设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21753526 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21753526 Country of ref document: EP Kind code of ref document: A1 |