WO2025002073A1 - 视频数据处理方法、装置及电子设备 - Google Patents
视频数据处理方法、装置及电子设备 Download PDFInfo
- Publication number
- WO2025002073A1 WO2025002073A1 PCT/CN2024/101103 CN2024101103W WO2025002073A1 WO 2025002073 A1 WO2025002073 A1 WO 2025002073A1 CN 2024101103 W CN2024101103 W CN 2024101103W WO 2025002073 A1 WO2025002073 A1 WO 2025002073A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- camera movement
- video
- video frame
- camera
- frame image
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/60—Control of cameras or camera modules
- H04N23/64—Computer-aided capture of images, e.g. transfer from script file into camera, check of taken image quality, advice or proposal for image composition or decision on when to take image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/222—Studio circuitry; Studio devices; Studio equipment
- H04N5/262—Studio circuits, e.g. for mixing, switching-over, change of character of image, other special effects ; Cameras specially adapted for the electronic generation of special effects
- H04N5/2621—Cameras specially adapted for the electronic generation of special effects during image pickup, e.g. digital cameras, camcorders, video cameras having integrated special effects capability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/222—Studio circuitry; Studio devices; Studio equipment
- H04N5/262—Studio circuits, e.g. for mixing, switching-over, change of character of image, other special effects ; Cameras specially adapted for the electronic generation of special effects
- H04N5/2628—Alteration of picture size, shape, position or orientation, e.g. zooming, rotation, rolling, perspective, translation
Definitions
- the embodiments of the present disclosure relate to the field of image processing technology, and in particular to a video data processing method, device and electronic device.
- Camera movement usually refers to motion shots.
- camera movement as an important form of narrative expression, can reflect the creative agility and artistic value.
- the reasonable use of camera movement in the video can help to portray the character image, character characteristics, lay the scene atmosphere, and also promote the video narrative.
- Different motion shots can control different rhythms in the narrative, and can bring different visual experiences and psychological hints.
- the embodiments of the present disclosure provide a video data processing method, device and electronic device to provide a camera movement template to a user, thereby solving the problem that it is difficult for the user to apply the camera movement effect in a video.
- an embodiment of the present disclosure provides a method for processing video data, the method comprising: extracting camera movements corresponding to multiple frames of video images from a reference video, and composing a camera movement sequence from the multiple camera movements; based on the camera movement sequence, generating a camera movement template for adding camera movements to the video.
- an embodiment of the present disclosure provides a video data processing method, the method comprising: in response to receiving a trigger operation of applying camera movement to a target video performed by a user, calling a camera movement template; wherein the camera movement sequence in the camera movement template is obtained according to the camera movement sequence extracted from a reference video; and applying multiple camera movements in the camera movement sequence in the camera movement template to multiple frames of video images of the target video.
- the present disclosure provides a video data processing device, the device comprising: an acquisition unit, configured to extract camera movements corresponding to multiple frames of video images from a reference video, and to extract camera movements corresponding to multiple frames of video images from a reference video.
- a camera movement sequence is formed; and a generating unit is used to generate a camera movement template for adding camera movement to a video based on the camera movement sequence.
- an embodiment of the present disclosure provides a video data processing device, comprising a calling unit for calling a camera movement template in response to receiving a trigger operation of applying camera movement to a target video by a user; wherein the camera movement sequence in the camera movement template is obtained according to a camera movement sequence extracted from a reference video; and an application unit for applying multiple camera movements in the camera movement sequence in the camera movement template to multiple frames of video images of the target video.
- an embodiment of the present disclosure provides an electronic device, comprising: a processor and a memory; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the first aspect, the second aspect, and various possible video data processing methods of the first and second aspects above.
- an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored.
- a processor executes the computer execution instructions, various possible video data processing methods as described in the first aspect, the second aspect, and the first and second aspects are implemented.
- an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the first aspect, the second aspect, and various possible video data processing methods of the first aspect and the second aspect.
- FIG1 is a flow chart of a video data processing method according to an embodiment of the present disclosure
- FIG2 is a principle flow chart of a video data processing method provided by the embodiment shown in FIG1 ;
- FIG3 is a second flow chart of a video data processing method provided by an embodiment of the present disclosure.
- FIG4 is a schematic diagram showing a principle of adding enhancement information to a camera movement sequence
- FIG5 is a third flow chart of a video data processing method provided by an embodiment of the present disclosure.
- FIG7 is a structural block diagram of a video data processing device provided by an embodiment of the present disclosure.
- FIG8 is a structural block diagram of a video data processing device provided by an embodiment of the present disclosure.
- FIG. 9 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure.
- Camera movement usually refers to a moving lens.
- the images taken by the moving lens can also be simulated by geometric transformation of the images taken by the fixed lens.
- the "push” in the camera movement means gradually approaching the object being photographed. You can use the zooming in of the image taken by the fixed lens to simulate the "push” camera movement.
- the “pull” in the camera movement means gradually moving away from the object being photographed. You can use the zooming out of the image taken by the fixed lens to simulate the "pull” camera movement.
- the “shift” in the camera movement means moving the lens. Similarly, you can use the cropping of the image taken by the fixed lens and the appropriate expansion of the scene to simulate the "shift” camera movement, etc.
- the present disclosure provides a video data processing method, which extracts camera movement sequences from existing videos and generates camera movement templates based on the camera movement sequences. Users can use the camera movement templates to add simulated camera movements to the videos they shoot, which reduces the difficulty of adding simulated camera movements to the videos they shoot.
- FIG. 1 shows a flow chart 1 of the video data processing method provided by the present disclosure.
- S101 extracting camera movements corresponding to multiple frames of video images from a reference video, and composing a camera movement sequence from the multiple camera movements.
- the execution subject of the video processing method in the present disclosure may be a user terminal device or a server.
- the reference video here refers to a reference video with camera movement applied, which includes: actually using a motion lens in the process of shooting the reference video, or applying a simulated motion lens such as zooming and rotating to a video shot with a fixed lens.
- the reference video may be, for example, a dance video.
- a motion lens can be used to shoot dance videos.
- a user who is proficient in the application of simulated lens movement can also add simulated lens movement to the dance video shot by a fixed lens to obtain the above reference video.
- the camera movement corresponding to the plurality of video frames can be determined by using each video frame in the reference video.
- the camera movement corresponding to the plurality of video frames can be a simulated camera movement.
- the camera movement sequence includes multiple camera movements arranged in time sequence, that is, the camera movement information in the camera movement sequence may include multiple camera movements and time information corresponding to each camera movement.
- the above-mentioned camera movements include but are not limited to: translation, zooming in, zooming out, and rotation.
- the above-mentioned execution subject can analyze and process the reference video to obtain the camera movements corresponding to multiple frames of video images in the reference video.
- the simulated camera movement here can be regarded as a geometric transformation operation.
- P non-collinear pixel points can be determined on the video frame image first, and then the matching P pixel points can be found on the previous video frame image. Then, the geometric transformation operation corresponding to the video frame image is determined based on the geometric relationship between the above P pixel points corresponding to the two frames of video images. Assume that the position of the P pixel points in the previous video frame image is X1 (the position of the center point of the figure composed of P pixel points is X1), and the position of the P pixel points in the video frame image is x1+ ⁇ x. Due to the continuity between the video frame images, it can be determined that the camera movement corresponding to the video frame image is a translation on the x-axis. The size of the translation can be approximately equal to ⁇ x.
- the geometric relationship between the matching pixel points in the previous video frame image and the next video frame image in the two adjacent video frames can be determined in the above manner to determine the simulated camera movement corresponding to the next video frame image.
- the camera movement corresponding to each video frame image in the reference video can be obtained by the above method.
- the camera movements corresponding to each video frame image are arranged in sequence according to the time information corresponding to each video frame image to obtain a camera movement sequence.
- Each camera movement in the above camera movement sequence can carry the time information of the video frame image corresponding to the camera movement.
- S102 Based on the camera movement sequence, generate a camera movement template for adding camera movement to the video.
- the camera movement sequences extracted from the reference video may be packaged in their respective orders to generate a camera movement template for adding camera movement to the video.
- the camera movement here refers to the simulated camera movement of performing geometric transformations such as translation, rotation, and scaling on the video frame images in the video that has been shot.
- a camera movement template including enhancement information for enhancing the camera movement may be generated according to a preset enhancement rule.
- the camera movement in the camera movement sequence obtained in step S101 may also be smoothed to obtain a smoothed camera movement template.
- camera movements corresponding to multiple frames of video images are extracted from a reference video, and a camera movement sequence is formed from the multiple camera movements; based on the camera movement sequence, a camera movement template for adding camera movements to the video is generated.
- the user can use the camera movement template to add camera movements to the captured video, which reduces the difficulty for the user to use the camera movement effect, so that users with different picture editing and processing capabilities can achieve better application effects when applying camera movements to videos, thereby improving the user experience.
- step S101 includes the following steps:
- the camera movement corresponding to the video frame image is determined according to the same image information in a previous video frame image and the video frame image.
- the camera movements corresponding to the video frame images are arranged in sequence to obtain a camera movement sequence.
- the reference video includes a plurality of video frames arranged in chronological order.
- the camera movement of the video frame can be determined from the geometric correspondence between the previous video frame and the same image information in the video frame.
- the camera movement corresponding to the first video frame can be a preset value.
- the preset value can indicate that no camera movement is used.
- the camera movement corresponding to the latter video frame image of the two adjacent video frame images is determined, so that the obtained camera movement template has a smaller granularity.
- the camera movement template is used to apply the camera movement to the shot video, a camera movement application effect with better continuity can be obtained.
- determining the camera movement corresponding to the video frame image according to the same image information in the video frame image and the previous video frame image of the video frame image includes the following steps:
- background images corresponding to the same background are determined in the video frame image and the previous video frame image of the video frame image.
- the camera movement is determined according to the geometric transformation relationship between the previous video frame image of the video frame image and the above-mentioned background images respectively corresponding to the video frame image.
- the background here may be a stationary entity object.
- the image information corresponding to the background may be the imaging result of the stationary entity object in the video frame image.
- the background may be a wall in the captured space.
- an edge operator can be used to determine a background image with a regular edge shape (such as a rectangle) in the video frame image, and then an image matching method is used to search for a background image identical to the background image in a previous video frame image of the video frame image.
- a regular edge shape such as a rectangle
- a background image (deemed as the first background image) can be determined from the video frame image, and then the image features of the first background image in the video frame image can be determined. The image features are then matched in the previous video frame image of the video frame image, so as to match a background image (deemed as the second background image) that matches the first background image in the previous video frame image.
- the camera movement of the video frame image can be determined based on the geometric correspondence between the first background image and the second background image.
- the camera movement of the video frame may be determined according to a geometric transformation relationship between feature points in the first background image and corresponding feature points in the second background image.
- the moving object can be firstly identified in multiple video frames of the reference video according to an object detection method.
- the image of the moving object is removed from the video frame image data by covering or deleting the image data of the moving object.
- Each video frame image after the above processing only includes the entire background image.
- the N-1th frame of video image is the previous frame of video image of the Nth frame.
- a background image first background image
- the image features of the first background image are extracted.
- the image features are used to match in the N-1th frame of video image, so as to match the second background image in the N-1th frame of video image.
- the N-1th frame of video image and the Nth frame of video image can be background images of the image without the moving object.
- the regular shape may be, for example, a rectangle, a triangle, etc.
- the vertices in the outline of the first background image may be used as feature points.
- the corresponding vertices in the second background image may be used as corresponding feature points.
- the camera movement may be determined by the geometric transformation relationship between at least one feature point in the first background image and the corresponding feature points in the second background image.
- the above-mentioned geometric transformation relationship may include, for example, rotation, translation, scaling, etc.
- the step of determining the camera movement according to the geometric transformation relationship between a previous video frame image of the video frame image and a background image respectively corresponding to the video frame image includes:
- the camera movement is determined based on an affine matrix between feature points of a background image corresponding to the video frame image and corresponding feature points of a background image corresponding to a previous video frame image of the video frame image.
- an affine matrix including four degrees of freedom such as x-axis translation, y-axis translation, rotation, and scaling
- the affine matrix here may be, for example, an affine matrix.
- the values of the parameters corresponding to the above-mentioned degrees of freedom may be approximately determined based on the affine transformation relationship between the coordinates of the multiple corresponding feature points in the background image in the previous video frame image and the feature points in the background image in the video frame image. According to the values of the parameters of the various degrees of freedom, the quantized camera movement corresponding to the video frame image is determined.
- the geometric transformation relationship between the corresponding feature point X'(x', y') in the first background image and the feature point X(x, y) in the second background image can be represented by the following formula (1):
- It can be an affine matrix.
- a, b, c, d, e, f are the parameters to be determined.
- the above-mentioned multiple feature points and the coordinates of the above-mentioned multiple corresponding feature points can be used to solve the above-mentioned formula (2), so as to obtain a quantitative representation of the camera movement of the Nth frame of the video image.
- the above a, b, c, d, e, and f have different constraints.
- the above parameters can be solved using the constraints corresponding to the camera movement and the above feature points and the coordinates of the corresponding feature points. Then, a quantitative representation of the camera movement is obtained.
- the above formula (2) can be solved based on the gradient descent algorithm.
- initial values a0, b0, c0, d0, e0, f0 can be set for each parameter, and the result obtained by multiplying the affine matrix composed of the initial values of each parameter with the column vector [x, y, 1]' is then subtracted from the column vector [x', y', 1]', and the obtained difference is used as the initial difference (Loss value). That is, the above difference is calculated by the following formula (3):
- the change direction corresponding to the change is the opposite direction of the gradient.
- the change direction corresponding to the change is the opposite direction of the gradient.
- the updated initial values corresponding to each parameter are the updated initial values corresponding to each parameter.
- the values of a, b, c, d, e, and f can be obtained.
- the affine matrix that is, the quantized camera movement, is obtained.
- the simulated camera movement of each video frame image may be determined according to the above method.
- the reference video includes a video frame image sequence arranged in time order: video frame image 1, video frame image 2, ..., video frame image k, ... and video frame image n.
- the camera movement of video frame image 1 can be set to a parameter for indicating that the camera movement is not used.
- the target object in motion can be removed from the video frame image to determine the background image in the video frame image.
- an edge detection operator Using an edge detection operator, a first background image having a rectangular edge contour is determined in the background image.
- the first background image is registered with the previous video frame image using a preset registration method to obtain a second background image in the previous video frame image that matches the first background image.
- a human body recognition algorithm can be used to identify a human body from video frame image 2. Then, the image data of the human body is removed from the image data of video frame image 2 to obtain background image data. An image detection method or an edge detection operator is used to identify a first background image having a rectangular outline. The image features of the first background image are extracted. Then, a match is performed in video frame image 1 based on the image features, and a second background image is matched from video frame image 1. The outline of the second background image is determined using an edge detection algorithm.
- first feature points such as the four vertices of the rectangular outline
- second feature points such as the four vertices of the rectangular outline
- the parameters of the affine matrix corresponding to video frame image 2 are calculated.
- the value of each parameter can be gradually solved using the gradient descent method and the preset loss function, and finally the value of each parameter of the affine matrix is obtained.
- the affine matrix composed of the values of each parameter can be used as the simulated camera movement 2 of the video frame image 2. Then the simulated camera movement 2 of the video frame image 2 is obtained. And so on, until the simulated camera movement n of the video frame image n is determined. Time information is added to the above-mentioned simulated camera movement 1, simulated camera movement 2, ..., simulated camera movement k, ... and simulated camera movement n, respectively, and each simulated camera movement is arranged according to the corresponding time information to obtain a camera movement sequence. Among them, the simulated camera movement 1 of the video frame image 1 can correspond to a preset initial value.
- the affine matrix corresponding to the latter video frame image in two adjacent video frames is estimated by using the feature points of the same background image in the two adjacent video frames, and the camera movement corresponding to the latter video frame image is represented by the affine matrix.
- the specific representation of the camera movement can be determined by using the existing affine matrix solution method, and an accurate camera movement representation can be obtained.
- FIG3 shows a flow chart of the second video data processing method provided by the present disclosure.
- the video data processing method includes the following steps:
- S301 extracting camera movements corresponding to multiple frames of video images from a reference video, and composing a camera movement sequence from the multiple camera movements.
- step S301 can refer to the description of the embodiment shown in FIG1 , which will not be described in detail here.
- S302 Determine first enhancement information corresponding to at least one camera movement according to action information and/or aesthetic deconstruction features of a target object in video frames respectively involved in a plurality of camera movements.
- S303 Generate a camera movement template including first enhancement information based on the camera movement sequence.
- the execution subject of this embodiment may be a terminal device, or a server that provides a camera movement template.
- each camera movement in the camera movement sequence may correspond to an affine matrix quantizing the camera movement.
- first camera movement enhancement information may also be set in the camera movement template.
- corresponding enhancement information can be set for a preset action of an object.
- the preset action here can include, for example, a preset action performed by a person's limbs.
- the above-mentioned preset action can include, for example, one or more specified actions in a dance action. That is to say, when the action of the target object in the video frame image in the multiple frames of the reference video is the above-mentioned preset action, enhancement information can be set for the camera movement corresponding to the video frame image.
- the first enhancement information here includes a coefficient for enlarging or reducing the camera movement.
- magnification and reduction factors can be set according to specific application scenarios, which will not be discussed here.
- magnification factor and reduction factor can also be set based on experience.
- first enhancement information positively correlated with the movement amplitude of the target object can be set.
- the camera movement template can include first enhancement information corresponding to each camera movement.
- Each first enhancement information is positively correlated with the movement amplitude of the target object in the corresponding video frame.
- the aesthetic deconstruction features of the video frames corresponding to the multiple camera movements may be analyzed, and then one or more camera movements may be selected from the multiple camera movements according to the aesthetic deconstruction features, and the first enhancement information may be set for the selected one or more camera movements.
- Aesthetic deconstruction can be characterized by emphasizing contingency and particularity.
- one or more camera movements are selected for enhancement through aesthetic deconstruction features, which can increase the interest of the video using the above camera movement template.
- this embodiment adds the first enhancement information corresponding to the camera movement determined according to the action information and/or aesthetic structure characteristics of the target object in the reference video, and the generated camera movement template includes the first enhancement information, so that the video obtained using the camera movement template can have better dynamics and appeal, which can improve the user's visual experience.
- Figure 4 shows a schematic diagram of the principle of adding enhanced information to a camera movement sequence.
- a reference video is input to the execution body of the video data processing method.
- the above-mentioned execution body can use the steps of the embodiment shown in Figure 1 or Figure 3 to analyze the camera movement of the reference video, so as to obtain a camera movement sequence.
- the aesthetic deconstruction features in the video frame image can be analyzed.
- the first enhanced information of the camera movement corresponding to the video frame image can be determined based on the above-mentioned aesthetic deconstruction features and/or the user action information in the video frame image.
- the first enhanced information corresponding to each video frame is associated with the camera movement corresponding to the video frame and stored to obtain a camera movement template including the first enhanced information.
- FIG5 shows a flow chart of the video data processing method provided by the present disclosure.
- the video data processing method includes the following steps:
- S501 extracting camera movements corresponding to multiple frames of video images from a reference video, and composing a camera movement sequence from the multiple camera movements.
- step S501 can refer to the description of the embodiment shown in FIG1 , which will not be described in detail here.
- S502 Determine second enhancement information corresponding to at least one camera movement according to audio information corresponding to video frame images respectively involved in a plurality of camera movements.
- S503 Generate a camera movement template including second enhancement information based on the camera movement sequence.
- the execution subject of this embodiment may be a terminal device, or a server that provides a camera movement template.
- each camera movement in the camera movement sequence may correspond to an affine matrix for quantizing the camera movement.
- the at least one camera movement may be one or more of the multiple camera movements. In some application scenarios, the at least one camera movement may be all the camera movements in the camera movement sequence.
- the above audio information includes but is not limited to one or more of the following: music type, emotions conveyed by the music, music structure, beat, human voice in the music, and instrument type.
- the second enhancement information here includes a coefficient for enlarging or reducing the camera movement.
- corresponding second enhancement information may be set for specific audio information, such as drum beats, and second enhancement information corresponding to the drum beats may be set.
- different second enhancement information may be set for different music types. For example, for audio information with a soft melody, second enhancement information corresponding to a reduction factor may be set, and for music with a strong sense of rhythm, second enhancement information corresponding to an amplification factor may be set.
- the second enhanced information corresponding to the at least one camera movement in the camera movement sequence is respectively associated with the at least one camera movement, and then packaged to generate a camera movement template including the second enhanced information.
- one or more camera movements in the camera movement sequence in the camera movement template are respectively associated with the corresponding second enhancement information.
- the multiple audio frames in the audio information used, the camera movement sequence and the video frame sequence of the video to which the camera movement template is to be applied can be aligned on the time axis.
- the camera movement and audio frame corresponding to the video frame image can be searched on the time axis.
- the audio information of the audio frame is detected, and the second enhancement information matching the audio information is searched from the camera movement template, and the second enhancement information is applied to the camera movement, so as to apply music-aware camera movement to the video.
- the camera movement template in this embodiment adds transformation information for transforming the camera movement related to audio information, which is beneficial for users to apply music-perceived camera movement to the shot video, thereby enhancing the appeal and expressiveness of the video and further enhancing the user experience.
- FIG6 shows a flow chart of the video data processing method provided by the present disclosure.
- the video data processing method includes the following steps:
- S601 In response to receiving a trigger operation of applying a camera movement to a target video, calling A camera movement template; wherein the camera movement sequence in the camera movement template is obtained according to the camera movement sequence extracted from the reference video.
- the execution subject of this embodiment may be a terminal device used by a user.
- the target video here can be any video shot using a fixed lens.
- the target video can be a dance video.
- the user may apply camera movement to the target video, for example, by displaying a camera movement control for indicating the application of camera movement in the interface displaying the target video.
- the user may perform a trigger operation on the camera movement control.
- the execution subject may call a pre-stored camera movement template.
- the camera movement template may include a camera movement sequence.
- the camera movement template here may be generated based on a camera movement sequence extracted from a reference video.
- S602 Apply multiple camera movements in the camera movement sequence to multiple frames of video images of a target video.
- the camera movement sequence includes multiple camera movements arranged in chronological order.
- Each camera movement in the camera movement sequence can be applied to multiple frames of the target video in sequence according to their respective time information.
- the camera movement sequence can be aligned with the video frame image sequence of the target video on the time axis. For each frame of the target video, the camera movement aligned on the time axis is applied thereto.
- a camera movement template when applying camera movement to a target video, a camera movement template is called, and multiple camera movements in the camera movement template are applied to multiple frames of video images of the target video, thereby reducing the difficulty of adding simulated camera movements to the shot video.
- Users with different picture editing and processing capabilities can achieve better application effects when applying camera movements to videos, thereby improving the user experience.
- the above method further comprises the following steps:
- the action of the target object in the current video frame image is detected, and first enhancement information matching the action is determined.
- the camera movement corresponding to the current video frame image is enhanced using the first enhancement information.
- the target object's action is a body movement of the target object.
- first enhancement information matching the preset action can be determined from the camera movement template.
- the first enhancement information is also associated with the above-mentioned camera movement template.
- the first enhancement information can be associated with a specific action (preset action) of the target object in the target video frame. Taking the target video as a dance video as an example, when the user performs a preset dance action in the target video, the dance action can be enhanced. The camera movement of the video frame is enhanced.
- first enhancement information matching the movement can be determined from the camera movement template.
- a corresponding first enhancement information may be associated.
- the enhancement amplitude of the camera movement indicated by the first enhancement information may be positively correlated with the motion amplitude of the target object.
- the motion amplitude of the target object here may include, for example, the motion amplitude of a human limb.
- the size of the camera movement enhancement coefficient corresponding to a video frame image is set according to the size of the dance movement of the target object detected from a video frame image.
- the first enhancement information is determined according to the action of the target object, and the camera movement of the video frame image is enhanced according to the first enhancement information, so that the dynamics of the target video can be enhanced and the appeal of the target video can be enhanced.
- the above method further comprises the following steps:
- the second enhancement information is applied to enhance the camera movement corresponding to the current video frame image.
- the audio information here includes but is not limited to one or more of music style, structure, emotion, beat, instrument type, and vocals.
- the above-mentioned camera movement template may include second enhancement information corresponding to different audio information.
- the current audio content can be acquired in real time, and the audio information corresponding to the acquired audio content is detected.
- the camera movement corresponding to the audio information is determined from the camera movement template.
- the second enhancement information corresponding to the audio information is determined from the camera movement template, and the second enhancement information is applied to the camera movement.
- the camera movement to which the second enhancement information is applied is applied to the current video frame image.
- the intensity of camera movement is triggered, making the overall camera movement effect of the video richer and more varied.
- Music camera movement analyzes the type, structure, emotion, beat and other dimensions of the music in the video, and gives a corresponding trigger signal for the above audio information, thereby automatically adapting special effects such as camera movement based on the original video, enhancing the appeal and expressiveness of the video.
- FIG. 7 is a structural block diagram of a video data processing device provided by an embodiment of the present disclosure.
- the video data processing device 70 includes: an acquisition unit 701 and a generation unit 702.
- the unit 702 is formed. Among them,
- An acquisition unit 701 is used to extract camera movements corresponding to multiple frames of video images from a reference video, and to form a camera movement sequence from the multiple camera movements;
- a generating unit is used to generate a camera movement template for adding camera movement to a video based on the camera movement sequence.
- the acquiring unit 701 is further configured to:
- the acquiring unit 701 is further configured to:
- background images corresponding to the same background are determined in the previous video frame image of the video frame image and the video frame image;
- the camera movement is determined according to a geometric transformation relationship between a previous video frame image of the video frame image and background images respectively corresponding to the video frame image.
- the acquiring unit 701 is further configured to:
- the camera movement is determined based on an affine matrix between feature points of a background image corresponding to the video frame image and corresponding feature points of a background image corresponding to a previous video frame image of the video frame image.
- the generating unit 702 is further configured to:
- a camera movement template including the first enhancement information is generated.
- the generating unit 702 is further configured to:
- the audio information includes at least one of the following:
- the type of music the emotions it conveys, the structure of the music, the beat, the vocals in the music, the types of instruments.
- FIG. 8 is a structural block diagram of a video data processing device provided by an embodiment of the present disclosure.
- the video data processing device 80 includes: a calling unit 801 and an application unit 802.
- the calling unit 801 is used to call the camera movement template in response to receiving a trigger operation of applying camera movement performed by a user on a target video; wherein the camera movement sequence in the camera movement template is obtained according to the camera movement sequence extracted from the reference video;
- the application unit 802 is used to apply multiple camera movements in the camera movement sequence in the camera movement template to multiple frames of video images of the target video.
- the application unit 802 is further configured to:
- the first enhancement information is applied to enhance the camera movement corresponding to the current video frame image.
- the application unit 802 is further configured to:
- the second enhancement information is applied to enhance the camera movement corresponding to the current video frame image.
- the embodiment of the present disclosure also provides an electronic device.
- FIG9 it shows a schematic diagram of the structure of an electronic device 900 suitable for implementing the embodiment of the present disclosure
- the electronic device 900 may be a terminal device or a server.
- the terminal device may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- PDAs personal digital assistants
- PADs Portable Android Devices, PADs
- PMPs portable multimedia players
- vehicle terminals such as vehicle navigation terminals
- fixed terminals such as digital TVs, desktop computers, etc.
- the electronic device shown in FIG9 is only an example and should not bring any limitation to the functions and scope of use of the embodiment of the present disclosure.
- the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 to a random access memory (RAM) 903.
- a processing device e.g., a central processing unit, a graphics processing unit, etc.
- RAM random access memory
- Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903.
- the processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904.
- I/O Input/Output
- An interface 905 is also connected to the bus 904 .
- the following devices may be connected to the I/O interface 905: input devices 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909.
- the communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data.
- FIG. 9 shows an electronic device 900 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program can be downloaded and installed from a network through a communication device 909, or installed from a storage device 908, or installed from a ROM 902.
- the processing device 901 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
- the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
- This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
- the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium.
- the computer readable medium can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device.
- the program code contained on the computer readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
- the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
- the computer-readable medium carries one or more programs.
- the electronic device executes the method shown in the above embodiment.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
- the program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
- LAN Local Area Network
- WAN Wide Area Network
- each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments of the present disclosure may be implemented by software or hardware.
- the name of the unit does not constitute a limitation on the unit in some cases.
- the definition of the element itself, for example, the acquisition unit can also be described as "a unit that extracts the camera movements corresponding to multiple frames of video images from a reference video and composes a camera movement sequence from multiple camera movements.”
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chips
- CPLDs complex programmable logic devices
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing.
- a more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM portable compact disk read-only memory
- CD-ROM compact disk read-only memory
- magnetic storage device or any suitable combination of the foregoing.
- a video data processing method comprising:
- a camera movement template is generated for adding camera movement to the video.
- camera movements corresponding to multiple frames of video images are extracted from a reference video, and a camera movement sequence is formed from the multiple camera movements, including:
- determining a camera movement corresponding to the video frame image according to a previous video frame image of the video frame image and the same image information in the video frame image includes:
- background images corresponding to the same background are determined in the previous video frame image of the video frame image and the video frame image;
- the camera movement is determined according to a geometric transformation relationship between a previous video frame image of the video frame image and background images respectively corresponding to the video frame image.
- determining the camera movement according to the geometric transformation relationship between the previous video frame image of the video frame image and the background image respectively corresponding to the video frame image includes:
- the camera movement is determined based on an affine matrix between feature points of a background image corresponding to the video frame image and corresponding feature points of a background image corresponding to a previous video frame image of the video frame image.
- generating a camera movement template for adding camera movement to a video based on a camera movement sequence includes:
- a camera movement template including the first enhancement information is generated.
- generating a camera movement template for adding camera movement to a video based on a camera movement sequence includes:
- the audio information includes at least one of the following:
- the type of music the emotions it conveys, the structure of the music, the beat, the vocals in the music, the types of instruments.
- a video data processing method comprising:
- a camera movement template In response to receiving a trigger operation of applying camera movement performed by a user on a target video, calling a camera movement template; wherein a camera movement sequence in the camera movement template is obtained according to a camera movement sequence extracted from a reference video;
- the method further includes:
- the camera movement corresponding to the current video frame image is enhanced using the first enhancement information.
- the method further includes:
- the second enhancement information is applied to enhance the camera movement corresponding to the current video frame image.
- a video data processing device comprising: an acquisition unit and a generation unit.
- An acquisition unit used for extracting camera movements corresponding to multiple frames of video images from a reference video, and forming a camera movement sequence from the multiple camera movements;
- the generating unit is used to generate a camera movement template for adding camera movement to the video based on the camera movement sequence.
- the acquisition unit is further configured to:
- the acquisition unit is further configured to:
- background images corresponding to the same background are determined in the previous video frame image of the video frame image and the video frame image;
- the camera movement is determined according to a geometric transformation relationship between a previous video frame image of the video frame image and background images respectively corresponding to the video frame image.
- the acquisition unit is further configured to:
- the camera movement is determined based on an affine matrix between feature points of a background image corresponding to the video frame image and corresponding feature points of a background image corresponding to a previous video frame image of the video frame image.
- the generating unit is further configured to:
- a camera movement template including the first enhancement information is generated.
- the generating unit is further configured to:
- the audio information includes at least one of the following:
- the type of music the emotions it conveys, the structure of the music, the beat, the vocals in the music, the types of instruments.
- a video data processing system comprises: a calling unit and an application unit; wherein,
- a calling unit configured to call a camera movement template in response to receiving a trigger operation of applying camera movement performed by a user on a target video; wherein the camera movement sequence in the camera movement template is obtained according to the camera movement sequence extracted from the reference video;
- the application unit is used to apply multiple camera movements in the camera movement sequence to multiple frames of video images of the target video.
- the application unit is further configured to:
- the camera movement corresponding to the current video frame image is enhanced using the first enhancement information.
- the application unit is further configured to:
- the second enhancement information is applied to enhance the camera movement corresponding to the current video frame image.
- an electronic device comprising: at least one processor and a memory;
- Memory stores computer-executable instructions
- At least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the first aspect, the second aspect, and various video data processing methods that may be involved in the first aspect and the second aspect as described above.
- a computer-readable storage medium in which computer execution instructions are stored.
- a processor executes the computer execution instructions, the video data processing methods as described in the first aspect, the second aspect, and various possible aspects of the first and second aspects are implemented.
- a computer program product including a computer program, which, when executed by a processor, implements the first aspect, the second aspect, and various video data processing methods that may be involved in the first and second aspects.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Image Analysis (AREA)
Abstract
本公开实施例提供一种视频数据处理方法、装置及电子设备,该方法包括:从参考视频中提取对多帧视频图像分别对应的运镜,并由多个运镜组成运镜序列;基于所述运镜序列,生成用于对视频添加运镜的运镜模板。上述方案通过从应用了运镜的参考视频中提取运镜模板,用户可以利用运镜模板对拍摄的视频添加运镜,降低了用户使用运镜效果的难度,使得具有不同画面编辑处理能力的用户为视频应用运镜时均可以达到较好的应用效果,改善了用户体验。
Description
本申请要求2023年6月27日递交的、标题为“视频数据处理方法、装置及电子设备”、申请号为2023107744702的中国发明专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本公开实施例涉及图像处理技术领域,尤其涉及一种视频数据处理方法、装置及电子设备。
运镜通常是指运动镜头。在视频拍摄过程中,运镜作为一种重要的叙事表现形式,能够体现创作灵动性和艺术价值。在视频中合理运用运镜,有助于刻画人物形象、角色特点、铺垫场景氛围,还对视频叙事有推动作用。不同的运动镜头可以在叙事中把控不同的节奏,可以带来不同的视觉体验和心里暗示。
发明内容
本公开实施例提供一种视频数据处理方法、装置及电子设备,以向用户提供运镜模板,解决用户在视频中应用运镜效果难度大的问题。
第一方面,本公开实施例提供一种视频数据处理方法,该方法包括:从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;基于所述运镜序列,生成用于对视频添加运镜的运镜模板。
第二方面,本公开实施例提供一种视频数据处理方法,该方法包括:响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,所述运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;将运镜模板中的运镜序列中的多个运镜应用到目标视频的多帧视频图像。
第三方面,本公开实施例提供一种视频数据处理装置,该装置包括:获取单元,用于从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组
成运镜序列;生成单元,用于基于所述运镜序列,生成用于对视频添加运镜的运镜模板。
第四方面,本公开实施例提供一种视频数据处理装置,该装置包括调用单元,用于响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,所述运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;应用单元,用于将运镜模板中的运镜序列中的多个运镜应用到目标视频的多帧视频图像。
第五方面,本公开实施例提供一种电子设备,包括:处理器和存储器;所述存储器存储计算机执行指令;所述处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面、第二方面以及第一方面和第二方面各种可能的视频数据处理方法。
第六方面,本公开实施例提供一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面、第二方面以及第一方面和第二方面各种可能的视频数据处理方法。
第七方面,本公开实施例提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上第一方面、第二方面以及第一方面和第二方面各种可能的视频数据处理方法。
为了更清楚地说明本公开实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例提供的视频数据处理方法的流程示意图一;
图2为图1所示实施例提供的视频数据处理方法的一个原理性流程图;
图3为本公开实施例提供的视频数据处理方法的流程示意图二;
图4为对运镜序列添加增强信息的一个原理性示意图;
图5为本公开实施例提供的视频数据处理方法的流程示意图三;
图6为本公开实施例提供的视频数据处理方法的一个流程示意图;
图7为本公开实施例提供的视频数据处理装置的结构框图;
图8为本公开实施例提供的视频数据处理装置的结构框图;
图9为本公开实施例提供的电子设备的硬件结构示意图。
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
运镜通常是指运动镜头,对于固定场景中的图像而言,运动镜头拍摄的图像,也可以用使用固定镜头拍摄的图像进行几何变换来模拟。例如运镜中“推”是逐渐靠近被拍摄的对象,可以使用对使用固定镜头拍摄的图像进行放大来模拟“推”这个运镜。运镜中的“拉”是逐渐远离被拍摄的对象,可以使用对使用固定镜头拍摄的图像进行缩小来模拟“拉”这个运镜。对于运镜中的“移”,是移动镜头,同样的,可以使用对使用固定镜头拍摄的图像进行裁剪和适当扩展场景来模拟“移”这个运镜等。
为了解决大部分用户对视频使用运镜难度大的问题,本公开提供了的视频数据处理方法,通过从已有视频中提取运镜序列,根据运镜序列生成运镜模板。用户可以利用运镜模板为所拍摄的视频添加模拟运镜,降低了为所拍摄的视频添加模拟运镜的难度。
请参考图1,其示出了本公开提供的视频数据处理方法的流程示意图一。
S101:从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列。
本公开中的视频处理方法的执行主体可以是用户终端设备,也可以是服务端。
在本实施例中,这里的参考视频是指应用了运镜的参考视频。这里应用了运镜包括:在拍摄参考视频过程中实际使用了运动镜头,或者对使用固定镜头拍摄的视频应用了进行缩放、旋转等模拟运动镜头。
在一些应用场景中,上述参考视频例如可以是舞蹈类视频。作为一种实
现方式,为了达到舞蹈画面的动感,增强舞蹈的叙事效果,可以使用运动镜头来拍摄舞蹈类视频。作为另外一种实现方式,还可以由精通模拟运镜应用的用户对由固定镜头拍摄的舞蹈类视频添加模拟运镜,得到上述参考视频。
对于应用了运镜的参考视频,可以通过参考视频中的各个视频帧,确定出多个视频帧分别对应的运镜。多个视频帧分别对应的运镜可以是模拟运镜。
上述运镜序列包括由多个运镜按照时间顺序排列的运镜。也即上述运镜序列中的运镜信息可以包括多个运镜以及各个运镜分别对应的时间信息。
上述运镜包括但不限于:平移、放大、缩小、旋转。
上述执行主体可以对参考视频进行分析处理,得到参考视频中多帧视频图像对应的运镜。
这里的模拟运镜可以视为几何变换操作。
为了确定视频帧图像对应的几何变换操作,可以先在该视频帧图像上确定不共线的P个像素点,再在前一视频帧图像上查找匹配的P个像素点。然后根据该两帧视频图像分别对应的上述P个像素点之间的几何关系,来确定该视频帧图像对应的几何变换操作。假设前一视频帧图像中P个像素点的位置为X1(P个像素点构成的图形的中心点的位置为X1),该视频帧图像中P个像素点的位置为x1+Δx。由于视频帧图像之间的连续性,可以确定该视频帧图像对应的运镜方式为在x轴平移。平移大小可以近似等于Δx。
实践中,对于参考视频中的每相邻两帧视频图像,可以按照上述方式确定由该相邻两帧视频图像中前一视频帧图像与后一视频帧图像中匹配的像素点之间的几何关系,确定后一视频帧图像对应的模拟运镜。
这样一来,通过上述方式可以得到参考视频中每一视频帧图像对应的运镜。将各视频帧图像分别对应的运镜,按照各视频帧图像对应的时间信息依次排列,得到运镜序列。上述运镜序列中的每一运镜,可以携带该运镜对应的视频帧图像的时间信息。
S102:基于运镜序列,生成用于对视频添加运镜的运镜模板。
在一些应用场景中,可以将从参考视频中提取的运镜序列按照各自的顺序进行封装,生成用于对视频添加运镜的运镜模板。
这里的运镜是指对已经拍摄的视频中的视频帧图像进行平移、旋转、缩放等几何变换的模拟运镜。
在一些应用场景中,在由步骤S101得到了上述运镜序列之后,可以根据预设增强规则,生成包括用于对上述运镜进行增强处理的增强信息的运镜模板。
在一些应用场景中,还可以对由步骤S101得到的运镜序列中运镜进行平滑处理,得到平滑处理后的运镜模板。
本实施例中,从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;基于运镜序列,生成用于对视频添加运镜的运镜模板。从而用户可以利用运镜模板对拍摄的视频添加运镜,降低了用户使用运镜效果的难度,使得具有不同画面编辑处理能力的用户为视频应用运镜时均可以达到较好的应用效果,改善了用户体验。
在一些实施例中,上述步骤S101包括如下步骤:
首先,对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜。
其次,将视频帧图像对应的运镜依次排列,得到运镜序列。
在这些实施例中,上述参考视频包括按照时间顺序依次排列的多帧视频图像。对于从第二帧视频图像开始的任意帧视频图像,可以从该视频帧图像的前一视频帧图像和该视频帧图像中的相同图像信息之间的几何对应关系,来确定由该视频帧图像的运镜。首帧视频图像对应的运镜可以为预设值。该预设值可以表示没有使用运镜。
在这些实施例中,通过每相邻两个视频帧图像,来确定该相邻两个视频帧图像中后一个视频帧图像对应的运镜,从而得到的运镜模板的细粒度较小,在使用该运镜模板对所拍摄的视频应用运镜时,可以得到连续性较好的运镜应用效果。
在一些实施例中,上述对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜,包括如下步骤:
首先,通过特征信息匹配,在该视频帧图像的前一视频帧图像和该视频帧图像中确定出同一背景分别对应的背景图像。
其次,根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的上述背景图像之间的几何变换关系,确定运镜。
这里的背景可以是静止实体对象。背景对应的图像信息可以是静止实体对象在视频帧图像中的成像结果。优选地,上述背景可以为所拍摄的空间中的墙体。
实践中,可以利用边缘算子在该视频帧图像中确定出边缘为规则形状(例如矩形)的背景图像。然后利用图像匹配的方式在该视频帧图像的前一视频帧图像中查找与该背景图像相同的背景图像。
在这些实施例中,可以在从该视频帧图像中确定一个背景图像(视为第一背景图像),再确定该视频帧图像中的上述第一背景图像的图像特征。然后将该图像特征在该视频帧图像的前一视频帧图像中进行匹配,从而在上述前一视频帧图像中匹配出与上述第一背景图像匹配的背景图像(视为第二背景图像)。可以根据第一背景图像与第二背景图像之间的几何对应关系,确定该视频帧图像的运镜。
具体地,可以根据第一背景图像中的特征点与第二背景图像中对应特征点之间的几何变换关系,确定该视频帧的运镜。
作为示意性说明,对于包括运动对象的参考视频。可以首先根据对象检测方法在参考视频的多帧视频图像中识别出运动对象。并将运动对象的图像使用遮盖或者删除的方式从视频帧图像数据中移出运动对象的图像数据,经过上述处理的每一视频帧图像仅包括全部的背景图像。
第N-1帧视频图像为第N帧视频图像的前一视频帧图像。可以先在第N帧视频图像中查找轮廓为规则图形的一背景图像(第一背景图像)。然后提取第一背景图像的图像特征。利用该图像特征在第N-1帧视频图像中进行匹配,从而在第N-1帧视频图像中匹配出第二背景图像。这里的第N-1帧视频图像和第N帧视频图像可以为移出了运动对象的图像的背景图像。
上述规则图形,例如可以为矩形、三角形等。可以将第一背景图像的轮廓中的顶点作为特征点。将第二背景图像的对应顶点作为对应特征点。可以由第一背景图像的至少一个特征点与各自在第二背景图像中的对应特征点之间的几何变换关系,确定上述运镜。
上述几何变换关系例如可以包括旋转、平移、缩放等。
在这些实施例中,通过选择轮廓为规则图形的背景图像来确定模拟运镜,可以降低确定模拟运镜的复杂度。
在一些实施例中,上述根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的背景图像之间的几何变换关系,确定运镜,包括:
基于该视频帧图像对应的背景图像的特征点和该视频帧图像的前一视频帧图像对应的背景图像的对应特征点之间的仿射矩阵,确定运镜。
实践中,可以设置包括x轴平移、y轴平移、旋转、缩放等4个自由度的仿射矩阵。这里的仿射矩阵例如可以为affine矩阵。可以根据上述前一视频帧图像中背景图像中的多个对应特征点与该视频帧图像中背景图像中的特征点的坐标之间的仿射变换关系近似确定上述各个自由度对应的参数的值。根据各个自由度的参数的值,确定对应该视频帧图像的量化运镜。
仍以第N-1帧视频图像中的背景图像为第二背景图像,第N帧视频图像为第一背景图像为例。第一背景图像中的对应特征点X’(x’,y’),第二背景图像中的特征点X(x,y)之间的几何变换关系可以由如下公式(1)表征:
X’=AX+B(1);其中,
A为包括旋转、缩放两个自由度的矩阵;B为包括平移的矩阵。
上述公式(1)可以表示为如下公式(2):
其中,可以为仿射矩阵。a、b、c、d、e、f为待求参数。
可以利用上述多个特征点以及上述多个对应特征点的坐标,来求解上述公式(2),从而得到第N帧视频图像的运镜的量化表示。
此外,对于不同类型的运镜,上述a、b、c、d、e、f具有不同的约束。在求解上述公式(2)中可以对每一类型的运镜,利用该运镜对应的约束以及上述特征点以及对应特征点的坐标来求解上述各参数。进而得到该运镜的量化表示。
在一些实施例中,可以基于梯度下降算法,求解上述公式(2)。示意性地,可以为各参数设置初始值a0、b0、c0、d0、e0、f0,以各参数的初始值组成的仿射矩阵与列向量[x,y,1]’相乘得到的结果,再与列向量[x’,y’,1]’求差,将得到的差值作为初始差值(Loss值)。也即由如下公式(3)计算上述差值:
然后,设置各参数分别对应的变化量,将各参数的初始值与上述变化量的和作为更新后的初始值,重新利用上述公式(3)计算上述Loss。如果Loss增大,则调整上述变换量的变化方向,重新确定更新后的初始值,重新计算上述公式(3)。
如果Loss减小,则变化量对应的变化方向为梯度反方向,继续沿上述方向确定变化量,再确定各参数分别对应的更新后的初始值。然后根据更新后的初始值利用上述公式(3)计算Loss。逐步迭代,直至Loss的趋于平稳。从而可以得到上述a、b、c、d、e、f的值。从而得到仿射矩阵,也即量化的运镜。
可以根据上述方式确定各视频帧图像的模拟运镜。
请参考图2,其示出了本实施例提供的视频数据处理方法的一个原理性流程图。参考视频包括按照时间顺序排列的视频帧图像序列:视频帧图像1、视频帧图像2、...、视频帧图像k、...和视频帧图像n。
对于上述n帧视频图像,视频帧图像1的运镜可以设置为用于指示未使用运镜的参数。可以从视频帧图像2开始,在该视频帧图像中移出处于运动状态的目标对象,确定出该视频帧图像中的背景图像。使用边缘检测算子,在背景图像中确定出具有矩形边缘轮廓的第一背景图像。
将该第一背景图像与前一视频帧图像利用预设配准方式进行配准,得到前一视频帧图像中与第一背景图像匹配的第二背景图像。
若参考视频为舞蹈视频,以视频帧图像2为例,可以利用人体识别算法,从视频帧图像2中识别出人体。然后从视频帧图像2的图像数据中去除人体的图像数据,得到背景图像数据。利用图像检测方法或者边缘检测算子,来识别具有矩形轮廓的第一背景图像。提取第一背景图像的图像特征。然后根据该图像特征在视频帧图像1中进行匹配,从视频帧图像1中匹配出第二背景图像。利用边缘检测算法确定第二背景图像的轮廓。从第一背景图像中寻找若干第一特征点(例如矩形轮廓的四个顶点),从第二背景图像中寻找与第一特征点对应的若干第二特征点(例如矩形轮廓的四个顶点)。根据第一特征点和第二特征点,来计算视频帧图像2对应的仿射矩阵的各个参数。
具体地,在计算预设仿射矩阵的各个参数时,可以利用梯度下降法以及预设损失函数逐步求解各参数的值。最后得到仿射矩阵的各个参数的值。
可以将由各个参数的值组成的仿射矩阵作为视频帧图像2的模拟运镜2。进而得到视频帧图像2的模拟运镜2。依次类推,直至确定出视频帧图像n的模拟运镜n。对上述模拟运镜1、模拟运镜2、...、模拟运镜k、...和模拟运镜n分别添加时间信息,并按照各自对应的时间信息将各模拟运镜排列得到运镜序列。其中视频帧图像1的模拟运镜1可以对应预设初始值。
在这些实现方式中,通过相邻两帧视频图像中相同背景图像的特征点来估算该相邻两帧视频图像中后一视频帧图像对应的仿射矩阵,由该仿射矩阵表征该后一视频帧图像对应的运镜,从而可以利用现有仿射矩阵的求解方式来确定运镜的具体表征,可以得到精确的运镜表征。
请继续参考图3,其示出了本公开提供的视频数据处理方法的流程示意图二。如图3所示,视频数据处理方法包括如下步骤:
S301:从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列。
上述步骤S301的具体实施可以参考图1所示实施例的阐述,此处不赘述。
S302:根据多个运镜分别涉及的视频帧中目标对象的动作信息和/或美学解构特征,确定至少一个运镜对应的第一增强信息。
S303:基于运镜序列,生成包括第一增强信息的运镜模板。
本实施例的执行主体可以为终端设备,还可以为提供运镜模板的服务端。
在步骤S301中得到了上述运镜序列之后,运镜序列中的每一个运镜,都可以对应一个量化该运镜的仿射矩阵。为了增加应用运镜模板的视频的感染力,在运镜模板中还可以设置运镜的第一增强信息。
在一些应用场景中,可以为对象的预设动作设置对应的增强信息。这里的预设动作例如可以包括人的肢体执行的预设动作。上述预设动作例如可以包括舞蹈动作中的一个或多个指定动作。也就是说,当参考视频的多帧视频图像中,当视频帧图像中的目标对象的动作为上述预设动作时,可以为该视频帧图像对应的运镜设置增强信息。
这里的第一增强信息包括对运镜进行放大或者缩小的系数。
上述放大系数和缩小系数可以根据具体的应用场景进行设置,此处不进
行限制。此外,上述放大系数和缩小系数还可以根据经验进行设置。
在一些应用场景中,对于每一个运镜,可以设置与目标对象的动作幅度正相关的第一增强信息。这样一来,运镜模板中可以包括与各个运镜分别对应的第一增强信息。每一个第一增强信息与其所对应的视频帧中目标对像的动作幅度正相关。
在一些应用场景中,可以分析多个运镜分别对应的视频帧的美学解构特征,然后根据美学解构特征从多个运镜中选择一个或者多个,对所选择的这一个或多个运镜设置第一增强信息。
美学解构特征可以为强调偶然性和特殊性的特征。
在这些应用场景中,通过美学解构特征选择一个或多个运镜进行增强,可以为应用上述运镜模板的视频增加趣味性。
与图1实施例相比,本实施例增加了根据参考视频中目标对象的动作信息和/或美学结构特征,确定运镜对应的第一增强信息,所生成的运镜模板中包括第一增强信息,从而使得运用该运镜模板得到的视频可以具有较好的动感,富有感染力,可以改善用户的视觉体验。
请参考图4,其示出了对运镜序列添加增强信息的一个原理性示意图。如图4所示,向视频数据的处理方法的执行主体输入参考视频。上述执行主体可以利用上述图1或图3所示实施例步骤分析参考视频的运镜,从而得到运镜序列。对于参考视频中每一个视频帧图像,可以分析该视频帧图像中的美学解构特征。对于每一视频帧图像,可以根据上述美学解构特征和/或该视频帧图像中的用户动作信息确定与该视频帧图像对应的运镜的第一增强信息。将每一视频帧对应的第一增强信息与该视频帧对应的运镜进行关联存储,得到包括第一增强信息的运镜模板。
请继续参考图5,其示出了本公开提供的视频数据处理方法的流程示意图三。如图5所示,视频数据处理方法包括如下步骤:
S501:从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列。
上述步骤S501的具体实施可以参考图1所示实施例的阐述,此处不赘述。
S502:根据多个运镜分别涉及的视频帧图像所对应的音频信息,确定至少一个运镜对应的第二增强信息。
S503:基于运镜序列,生成包括第二增强信息的运镜模板。
本实施例的执行主体可以为终端设备,还可以为提供运镜模板的服务端。
在步骤S501中得到了上述运镜序列之后,运镜序列中的每一个运镜,可以对应一个量化该运镜的仿射矩阵。
在本实施例中,上述至少一个运镜可以是上述多个运镜中的一个或多个运镜。在一些应用场景中,上述至少一个运镜可以是上述运镜序列中的全部运镜。
上述音频信息包括但不限于以下中的一种或多种:音乐类型、音乐传达的情绪、音乐结构、节拍、音乐中的人声、乐器类型。
这里的第二增强信息包括对运镜进行放大或者缩小的系数。
在一些实施例中,对于特定的音频信息可以设置对应的第二增强信息。例如对于鼓点,可以设置与鼓点对应的第二增强信息。
在一些实施例中,对于不同音乐类型可以设置不同的第二增强信息。例如对于旋律柔和的音频信息,可以对应缩小系数的第二增强信息,对于节奏感较强的卡点音乐,可以设置对应放大系数的第二增强信息等。
将运镜序列中的上述至少一个运镜分别对应的第二增强信息分别与该至少一个运镜进行关联,然后进行封装,生成包括第二增强信息的运镜模板。
也就是说,运镜模板中上述运镜序列中的一个运镜或者多个运镜,分别关联了各自对应的第二增强信息。在应用运镜模板时,可以将所使用的音频信息中的多个音频帧、运镜序列和待应用运镜模板的视频的视频帧序列在时间轴上对齐。对于每一个视频帧图像,可以在时间轴上查找该视频帧图像对应的运镜和音频帧。检测音频帧的音频信息,并从运镜模板中查找与音频信息匹配的第二增强信息,对上述运镜应用上述第二增强信息,从而对上述视频应用音乐感知运镜。
与图1所示实施例相比,本实施例中的运镜模板中增加了与音频信息相关的对运镜进行变换的变换信息,有利于用户对所拍摄的视频应用音乐感知运镜,从而增强视频的感染力和表现力,进一步增强用户的体验。
请继续参考图6,其示出了本公开提供的视频数据处理方法的一个流程示意图。如图6所示,视频数据处理方法包括如下步骤:
S601:响应于接收到用户对目标视频执行的应用运镜的触发操作,调用
运镜模板;其中,运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到。
本实施例的执行主体可以是用户使用的终端设备。
这里的目标视频可以是使用固定镜头拍摄的任意视频。在一些应用场景中,上述目标视频可以是舞蹈类视频。
用户可以对上述目标视频应用运镜,例如在显示目标视频的界面中显示用于指示应用运镜的运镜控件。用户可以对上述运镜控件执行触发操作。上述执行主体在接收到对运镜控件的触发操作后,可以调用预先存储的运镜模板。运镜模板可以包括运镜序列。这里的运镜模板可以是根据由参考视频中提取的运镜序列生成。
S602:将上述运镜序列中的多个运镜应用到目标视频的多帧视频图像。
上述运镜序列包括由多个运镜按照发生时间先后排列的运镜。
可以依次将运镜序列中的各运镜按照各自的时间信息应用到目标视频的多帧视频图像。作为一种实现方式,可以在时间轴上将运镜序列与目标视频的视频帧图像序列对齐。对于目标视频中的每一帧视频图像,对其应用在时间轴上对齐的运镜。
本实施例中,通过对目标视频应用运镜时,调用运镜模板,将运镜模板中的多个运镜应用到目标视频的多帧视频图像,降低了为所拍摄的视频添加模拟运镜的难度,可以实现具有不同画面编辑处理能力的用户为视频应用运镜时均可以达到较好的应用效果,改善了用户体验。
在一些实施例中,上述方法还包括如下步骤:
首先,检测当前视频帧图像中目标对象的动作,确定与动作匹配的第一增强信息。
其次,对当前视频帧图像对应的运镜使用第一增强信息进行增强。
目标对象的动作为目标对象的肢体动作。在一些实施例中,当上述目标对象动作与预设动作匹配时,可以从运镜模板中确定与该预设动作匹配的第一增强信息。
也就是说,在上述运镜模板中还关联有第一增强信息。第一增强信息可以与目标视频帧目标对象的特定的动作(预设动作)相关联。以目标视频为舞蹈类视频为例,当目标视频中用户做出预设舞蹈动作时,可以对该舞蹈动
作的视频帧的运镜进行增强。
在一些实施例中,若目标对象的动作的幅度大于预设幅度阈值,可以从运镜模板中确定与该动作匹配的第一增强信息。
对于每一个运镜,可以关联一个对应的第一增强信息。上述第一增强信息指示的对运镜的增强幅度可以与目标对象的动作幅度正相关。这里的目标对象的动作幅度例如可以包括人的肢体的动作幅度。
以目标视频的场景为舞蹈场景为例,根据从一视频帧图像中检测到的目标对象的舞蹈动作的大小,设置对该视频帧图像对应的运镜增强系数的大小。舞蹈动作幅度越大,上述增强系数的大小越大。
在这些实施例中,通过根据目标对象的动作确定第一增强信息,按照第一增强信息对该视频帧图像的运镜进行增强,从而可以增强目标视频的动感,增强目标视频的感染力。
在一些实施例中,上述方法还包括如下步骤:
首先,检测当前音频信息,确定与当前音频信息匹配的第二增强信息。
其次,对当前视频帧图像对应的运镜应用第二增强信息进行增强。
这里的音频信息包括但不限于音乐风格、结构、情绪、节拍、乐器类型、人声中的一项或多项。
在上述运镜模板中,可以包括不同音频信息对应的第二增强信息。
在这些实施例中,可以实时获取当前音频内容,并对所获取的音频内容检测该音频内容对应的音频信息。从运镜模板中确定与该音频信息对应的运镜。同时从运镜模板中确定与该音频信息对应的第二增强信息,为该运镜应用上述第二增强信息。将将应用了上述第二增强信息的上述运镜应用到当前视频帧图像。
在这些实施例中,结合音频信息的分析,触发强度的运镜,使视频整体的运镜效果更加丰富和多变。音乐运镜通过对视频中音乐的类型、结构、情绪、节拍等维度进行分析,并为上述音频信息给出相应的触发信号,从而在原视频的基础上自动适配运镜等特效,增强了视频的感染力和表现力。
对应于上文图1所示实施例的视频数据处理方法,图7为本公开实施例提供的视频数据处理装置的结构框图。为了便于说明,仅示出了与本公开实施例相关的部分。参照图7,视频数据处理装置70包括:获取单元701和生
成单元702。其中,
获取单元701,用于从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;
702生成单元,用于基于运镜序列,生成用于对视频添加运镜的运镜模板。
在本公开的一个实施例中,获取单元701进一步用于:
对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜;
将视频帧图像对应的运镜依次排列,得到运镜序列。
在本公开的一个实施例中,获取单元701进一步用于:
通过特征信息匹配,在该视频帧图像的前一视频帧图像和该视频帧图像中确定出同一背景分别对应的背景图像;
根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的背景图像之间的几何变换关系,确定运镜。
在本公开的一个实施例中,获取单元701进一步用于:
基于该视频帧图像对应的背景图像的特征点和该视频帧图像的前一视频帧图像对应的背景图像的对应特征点之间的仿射矩阵,确定运镜。
在本公开的一个实施例中,生成单元702进一步用于:
根据多个运镜分别涉及的视频帧图像中目标对象的动作信息和/或美学解构特征,确定至少一个运镜对应的第一增强信息;
生成包括第一增强信息的运镜模板。
在本公开的一个实施例中,生成单元702进一步用于:
根据多个运镜分别涉及的视频帧图像所对应的音频信息,确定至少一个运镜对应的第二增强信息;
生成包括度第二增强信息的运镜模板。
在本公开的一个实施例中,音频信息包括以下至少之一:
音乐类型、音乐传达的情绪、音乐结构、节拍、音乐中的人声、乐器类型。
对应于上文图5所示实施例的视频数据处理方法,图8为本公开实施例提供的视频数据处理装置的结构框图。为了便于说明,仅示出了与本公开实
施例相关的部分。参照图8,视频数据处理装置80包括:调用单元801和应用单元802。其中,
调用单元801,用于响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;
应用单元802,用于将运镜模板中的运镜序列中的多个运镜应用到目标视频的多帧视频图像。
在一些实施例中,应用单元802进一步用于:
检测当前视频帧图像中目标对象的动作,确定与动作匹配的第一增强信息;
对当前视频帧图像对应的运镜应用第一增强信息进行增强。
在一些实施例中,应用单元802进一步用于:
检测当前音频信息,确定与当前音频信息匹配的第二增强信息;
对当前视频帧图像对应的运镜应用第二增强信息进行增强。
为了实现上述实施例,本公开实施例还提供了一种电子设备。
参考图9,其示出了适于用来实现本公开实施例的电子设备900的结构示意图,该电子设备900可以为终端设备或服务器。其中,终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,简称PDA)、平板电脑(Portable Android Device,简称PAD)、便携式多媒体播放器(Portable Media Player,简称PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图9示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图9所示,电子设备900可以包括处理装置(例如中央处理器、图形处理器等)901,其可以根据存储在只读存储器(Read Only Memory,简称ROM)902中的程序或者从存储装置908加载到随机访问存储器(Random Access Memory,简称RAM)903中的程序而执行各种适当的动作和处理。在RAM 903中,还存储有电子设备900操作所需的各种程序和数据。处理装置901、ROM902以及RAM 903通过总线904彼此相连。输入/输出(I/O)
接口905也连接至总线904。
通常,以下装置可以连接至I/O接口905:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置906;包括例如液晶显示器(Liquid Crystal Display,简称LCD)、扬声器、振动器等的输出装置907;包括例如磁带、硬盘等的存储装置908;以及通信装置909。通信装置909可以允许电子设备900与其他设备进行无线或有线通信以交换数据。虽然图9示出了具有各种装置的电子设备900,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置909从网络上被下载和安装,或者从存储装置908被安装,或者从ROM 902被安装。在该计算机程序被处理装置901执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信
号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备执行上述实施例所示的方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括面向对象的程序设计语言-诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言-诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(Local Area Network,简称LAN)或广域网(Wide Area Network,简称WAN)-连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元的名称在某种情况下并不构成对该单
元本身的限定,例如,获取单元还可以被描述为“从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
第一方面,根据本公开的一个或多个实施例,提供了一种视频数据处理方法,该方法包括:
从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;
基于运镜序列,生成用于对视频添加运镜的运镜模板。
根据本公开的一个或多个实施例,从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列,包括:
对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜;
将视频帧图像对应的运镜依次排列,得到运镜序列。
根据本公开的一个或多个实施例,对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜,包括:
通过特征信息匹配,在该视频帧图像的前一视频帧图像和该视频帧图像中确定出同一背景分别对应的背景图像;
根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的背景图像之间的几何变换关系,确定运镜。
根据本公开的一个或多个实施例,根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的背景图像之间的几何变换关系,确定运镜,包括:
基于该视频帧图像对应的背景图像的特征点和该视频帧图像的前一视频帧图像对应的背景图像的对应特征点之间的仿射矩阵,确定运镜。
根据本公开的一个或多个实施例,基于运镜序列,生成用于对视频添加运镜的运镜模板,包括:
根据多个运镜分别涉及的视频帧图像中目标对象的动作信息和/或美学解构特征,确定至少一个运镜对应的第一增强信息;
生成包括第一增强信息的运镜模板。
根据本公开的一个或多个实施例,基于运镜序列,生成用于对视频添加运镜的运镜模板,包括:
根据多个运镜分别涉及的视频帧图像所对应的音频信息,确定至少一个运镜对应的第二增强信息;
生成包括度第二增强信息的运镜模板。
根据本公开的一个或多个实施例,音频信息包括以下至少之一:
音乐类型、音乐传达的情绪、音乐结构、节拍、音乐中的人声、乐器类型。
第二方面,根据本公开的一个或多个实施例,提供了一种视频数据处理方法,该方法包括:
响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;
将运镜序列中的多个运镜应用到目标视频的多帧视频图像。
根据本公开的一个或多个实施例,该方法还包括:
检测当前视频帧图像中目标对象的动作,确定与动作匹配的第一增强信息;
对当前视频帧图像对应的运镜使用第一增强信息进行增强。
根据本公开的一个或多个实施例,该方法还包括:
检测当前音频信息,确定与当前音频信息匹配的第二增强信息;
对当前视频帧图像对应的运镜应用第二增强信息进行增强。
第三方面,根据本公开的一个或多个实施例,提供了一种视频数据处理装置,该装置包括:获取单元和生成单元。其中,
获取单元,用于从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;
生成单元,用于基于运镜序列,生成用于对视频添加运镜的运镜模板。
在本公开的一个实施例中,获取单元进一步用于:
对于参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜;
将视频帧图像对应的运镜依次排列,得到运镜序列。
在本公开的一个实施例中,获取单元进一步用于:
通过特征信息匹配,在该视频帧图像的前一视频帧图像和该视频帧图像中确定出同一背景分别对应的背景图像;
根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的背景图像之间的几何变换关系,确定运镜。
在本公开的一个实施例中,获取单元进一步用于:
基于该视频帧图像对应的背景图像的特征点和该视频帧图像的前一视频帧图像对应的背景图像的对应特征点之间的仿射矩阵,确定运镜。
在本公开的一个实施例中,生成单元进一步用于:
根据多个运镜分别涉及的视频帧图像中目标对象的动作信息和/或美学解构特征,确定至少一个运镜对应的第一增强信息;
生成包括第一增强信息的运镜模板。
在本公开的一个实施例中,生成单元进一步用于:
根据多个运镜分别涉及的视频帧图像所对应的音频信息,确定至少一个运镜对应的第二增强信息;
生成包括度第二增强信息的运镜模板。
在本公开的一个实施例中,音频信息包括以下至少之一:
音乐类型、音乐传达的情绪、音乐结构、节拍、音乐中的人声、乐器类型。
第四方面,根据本公开的一个或多个实施例,提供了一种视频数据处理
装置,该装置包括:调用单元和应用单元;其中,
调用单元,用于响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;
应用单元,用于将运镜序列中的多个运镜应用到目标视频的多帧视频图像。
在一些实施例中,应用单元进一步用于:
检测当前视频帧图像中目标对象的动作,确定与动作匹配的第一增强信息;
对当前视频帧图像对应的运镜使用第一增强信息进行增强。
在一些实施例中,应用单元进一步用于:
检测当前音频信息,确定与当前音频信息匹配的第二增强信息;
对当前视频帧图像对应的运镜应用第二增强信息进行增强。
第五方面,根据本公开的一个或多个实施例,提供了一种电子设备,包括:至少一个处理器和存储器;
存储器存储计算机执行指令;
至少一个处理器执行存储器存储的计算机执行指令,使得至少一个处理器执行如上第一方面、第二方面以及第一方面和第二方面各种可能涉及的视频数据处理方法。
第六方面,根据本公开的一个或多个实施例,提供了一种计算机可读存储介质,计算机可读存储介质中存储有计算机执行指令,当处理器执行计算机执行指令时,实现如上第一方面、第二方面以及第一方面和第二方面各种可能涉及的视频数据处理方法。
第七方面,根据本公开的一个或多个实施例,提供了一种计算机程序产品,包括计算机程序,计算机程序被处理器执行时实现如上第一方面、第二方面以及第一方面和第二方面各种可能涉及的视频数据处理方法。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上
述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。
Claims (15)
- 一种视频数据处理方法,包括:从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;基于所述运镜序列,生成用于对视频添加运镜的运镜模板。
- 根据权利要求1所述的方法,其中所述从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列,包括:对于所述参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜;将所述视频帧图像对应的运镜依次排列,得到所述运镜序列。
- 根据权利要求2所述的方法,其中所述对于所述参考视频中的视频帧图像,根据该视频帧图像的前一视频帧图像与该视频帧图像中的相同图像信息确定该视频帧图像对应的运镜,包括:通过特征信息匹配,在该视频帧图像的前一视频帧图像和该视频帧图像中确定出同一背景分别对应的背景图像;根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的所述背景图像之间的几何变换关系,确定所述运镜。
- 根据权利要求3所述的方法,其中所述根据该视频帧图像的前一视频帧图像和该视频帧图像分别对应的所述背景图像之间的几何变换关系,确定所述运镜,包括:基于该视频帧图像对应的所述背景图像的特征点和该视频帧图像的前一视频帧图像对应的所述背景图像的对应特征点之间的仿射矩阵,确定所述运镜。
- 根据权利要求1所述的方法,其中所述基于所述运镜序列,生成用于对视频添加运镜的运镜模板,包括:根据所述多个运镜分别涉及的视频帧图像中目标对象的动作信息和/或美学解构特征,确定至少一个运镜对应的第一增强信息;生成包括所述第一增强信息的运镜模板。
- 根据权利要求1所述的方法,其中所述基于所述运镜序列,生成用于对视频添加运镜的运镜模板,包括:根据所述多个运镜分别涉及的视频帧图像所对应的音频信息,确定至少一个运镜对应的第二增强信息;生成包括所述第二增强信息的运镜模板。
- 根据权利要求6所述的方法,其中所述音频信息包括以下至少之一:音乐类型、音乐传达的情绪、音乐结构、节拍、音乐中的人声、乐器类型。
- 一种视频数据处理方法,包括:响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,所述运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;将所述运镜序列中的多个运镜应用到目标视频的多帧视频图像。
- 根据权利要求8所述的方法,其中所述方法还包括:检测当前视频帧图像中目标对象的动作,确定与所述动作匹配的第一增强信息;对当前视频帧图像对应的运镜使用第一增强信息进行增强。
- 根据权利要求8所述的方法,其中所述方法还包括:检测当前音频信息,确定与当前音频信息匹配的第二增强信息;对当前视频帧图像对应的运镜应用第二增强信息进行增强。
- 一种视频数据处理装置,包括:获取单元,用于从参考视频中提取多帧视频图像对应的运镜,并由多个运镜组成运镜序列;生成单元,用于基于所述运镜序列,生成用于对视频添加运镜的运镜模板。
- 一种视频数据处理装置,包括:调用单元,用于响应于接收到用户对目标视频执行的应用运镜的触发操作,调用运镜模板;其中,所述运镜模板中的运镜序列根据从参考视频中提取的运镜序列得到;应用单元,用于将运镜模板中的运镜序列中的多个运镜应用到目标视频的多帧视频图像。
- 一种电子设备,包括:处理器和存储器;所述存储器存储计算机执行指令;所述处理器执行所述存储器存储的计算机执行指令,使得所述处理器执行如权利要求1至10中任一项所述的视频数据处理方法。
- 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如权利要求1至10中任一项所述的视频数据处理方法。
- 一种计算机程序产品,包括计算机程序,其特征在于,所述计算机程序被处理器执行时实现如权利要求1至10中任一项所述的视频数据处理方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310774470.2A CN116828282A (zh) | 2023-06-27 | 2023-06-27 | 视频数据处理方法、装置及电子设备 |
| CN202310774470.2 | 2023-06-27 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025002073A1 true WO2025002073A1 (zh) | 2025-01-02 |
Family
ID=88114151
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/101103 Ceased WO2025002073A1 (zh) | 2023-06-27 | 2024-06-24 | 视频数据处理方法、装置及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116828282A (zh) |
| WO (1) | WO2025002073A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116828282A (zh) * | 2023-06-27 | 2023-09-29 | 北京字跳网络技术有限公司 | 视频数据处理方法、装置及电子设备 |
| CN119893016A (zh) * | 2023-10-23 | 2025-04-25 | 北京小米移动软件有限公司 | 视频录制方法及装置、电子设备、存储介质 |
| CN119701350B (zh) * | 2024-12-31 | 2026-03-06 | 网易(杭州)网络有限公司 | 一种显示控制方法、装置、电子设备及可读存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114173060A (zh) * | 2021-12-10 | 2022-03-11 | 北京瞰瞰智能科技有限公司 | 智能移动拍摄方法、装置以及控制器 |
| CN114363689A (zh) * | 2022-01-11 | 2022-04-15 | 广州博冠信息科技有限公司 | 直播控制方法、装置、存储介质及电子设备 |
| CN114500851A (zh) * | 2022-02-23 | 2022-05-13 | 广州博冠信息科技有限公司 | 视频录制方法及装置、存储介质、电子设备 |
| WO2023065832A1 (zh) * | 2021-10-18 | 2023-04-27 | 华为技术有限公司 | 视频的制作方法和电子设备 |
| CN116828282A (zh) * | 2023-06-27 | 2023-09-29 | 北京字跳网络技术有限公司 | 视频数据处理方法、装置及电子设备 |
-
2023
- 2023-06-27 CN CN202310774470.2A patent/CN116828282A/zh active Pending
-
2024
- 2024-06-24 WO PCT/CN2024/101103 patent/WO2025002073A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023065832A1 (zh) * | 2021-10-18 | 2023-04-27 | 华为技术有限公司 | 视频的制作方法和电子设备 |
| CN114173060A (zh) * | 2021-12-10 | 2022-03-11 | 北京瞰瞰智能科技有限公司 | 智能移动拍摄方法、装置以及控制器 |
| CN114363689A (zh) * | 2022-01-11 | 2022-04-15 | 广州博冠信息科技有限公司 | 直播控制方法、装置、存储介质及电子设备 |
| CN114500851A (zh) * | 2022-02-23 | 2022-05-13 | 广州博冠信息科技有限公司 | 视频录制方法及装置、存储介质、电子设备 |
| CN116828282A (zh) * | 2023-06-27 | 2023-09-29 | 北京字跳网络技术有限公司 | 视频数据处理方法、装置及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116828282A (zh) | 2023-09-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2025002073A1 (zh) | 视频数据处理方法、装置及电子设备 | |
| US12155957B2 (en) | Video special effects processing method and apparatus | |
| WO2020186935A1 (zh) | 虚拟对象的显示方法、装置、电子设备和计算机可读存储介质 | |
| CN109348277B (zh) | 运动像素视频特效添加方法、装置、终端设备及存储介质 | |
| CN112258622B (zh) | 图像处理方法、装置、可读介质及电子设备 | |
| JP2022505118A (ja) | 画像処理方法、装置、ハードウェア装置 | |
| CN111210485A (zh) | 图像的处理方法、装置、可读介质和电子设备 | |
| CN111368668B (zh) | 三维手部识别方法、装置、电子设备及存储介质 | |
| WO2024198855A1 (zh) | 场景渲染方法、装置、设备、计算机可读存储介质及产品 | |
| CN109600559B (zh) | 一种视频特效添加方法、装置、终端设备及存储介质 | |
| CN112766215B (zh) | 人脸图像处理方法、装置、电子设备及存储介质 | |
| CN111833461A (zh) | 一种图像特效的实现方法、装置、电子设备及存储介质 | |
| WO2023193642A1 (zh) | 视频处理方法、装置、设备及介质 | |
| WO2024152723A1 (zh) | 表情信息识别方法、装置、设备、可读存储介质及产品 | |
| CN112132859A (zh) | 贴纸生成方法、装置、介质和电子设备 | |
| WO2025021172A1 (zh) | 视频生成方法、装置、电子设备及存储介质 | |
| CN116309983A (zh) | 虚拟人物模型的训练方法、生成方法、装置和电子设备 | |
| WO2025002075A1 (zh) | 视频生成方法、装置、电子设备及存储介质 | |
| WO2025011490A1 (zh) | 视频特效添加方法、装置、设备、存储介质和程序产品 | |
| WO2025113388A1 (zh) | 一种三维数据的生成方法、装置、电子设备及存储介质 | |
| WO2025161766A1 (zh) | 图像处理方法、装置、设备、计算机可读存储介质及产品 | |
| WO2025167333A1 (zh) | 一种图像生成方法、装置、设备、介质、产品 | |
| WO2025021169A1 (zh) | 图像处理方法、设备、存储介质及程序产品 | |
| WO2025011491A1 (zh) | 视频处理方法、装置、设备、存储介质和程序产品 | |
| WO2024198952A1 (zh) | 图像超分辨率方法、设备、存储介质及程序产品 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24830715 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |