WO2025001955A1 - 视频编码方法、装置、电子设备及可读存储介质 - Google Patents

视频编码方法、装置、电子设备及可读存储介质 Download PDF

Info

Publication number
WO2025001955A1
WO2025001955A1 PCT/CN2024/100314 CN2024100314W WO2025001955A1 WO 2025001955 A1 WO2025001955 A1 WO 2025001955A1 CN 2024100314 W CN2024100314 W CN 2024100314W WO 2025001955 A1 WO2025001955 A1 WO 2025001955A1
Authority
WO
WIPO (PCT)
Prior art keywords
video
video frames
macroblock
frame
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/100314
Other languages
English (en)
French (fr)
Inventor
曾达彬
付俊
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Vivo Mobile Communication Co Ltd
Original Assignee
Vivo Mobile Communication Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Vivo Mobile Communication Co Ltd filed Critical Vivo Mobile Communication Co Ltd
Publication of WO2025001955A1 publication Critical patent/WO2025001955A1/zh
Priority to US19/411,503 priority Critical patent/US20260095580A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • H04N19/137Motion inside a coding unit, e.g. average field, frame or block difference
    • H04N19/139Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/124Quantisation
    • H04N19/126Details of normalisation or weighting functions, e.g. normalisation matrices or variable uniform quantisers
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/156Availability of hardware or computational resources, e.g. encoding based on power-saving criteria
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/20Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video object coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/537Motion estimation other than block-based
    • H04N19/54Motion estimation other than block-based using feature points or meshes
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/625Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding using discrete cosine transform [DCT]

Definitions

  • the present application belongs to the field of video technology, and specifically relates to a video encoding method, device, electronic device and readable storage medium.
  • electronic devices when encoding multiple captured video frames, electronic devices need to perform at least one of the following: determine the encoding search range of each macroblock in the multiple video frames to identify redundant information between the multiple video frames, and then compress the redundant information; determine the internal coding frames in the multiple video frames according to the degree of change of the image content of each video frame to increase the video frame compression rate and reduce the size of the encoded video.
  • the electronic device can determine it by traversing multiple pixel points around each pixel point in each of the above-mentioned macroblocks through a search algorithm; and for the above-mentioned internal coding frame, the electronic device can determine it by comparing the pixel changes of the macroblocks in the above-mentioned multiple video frames.
  • the electronic device since a large amount of calculation is required to traverse the above-mentioned multiple pixel points and compare the above-mentioned pixel changes, the electronic device has a large computing power when determining the above-mentioned coding search range and at least one of the above-mentioned internal coding frames, which leads to high power consumption of the electronic device during the video encoding process.
  • the purpose of the embodiments of the present application is to provide a video encoding method, device, electronic device and readable storage medium, which can solve the problem of high power consumption of electronic devices during video encoding.
  • an embodiment of the present application provides a video encoding method, the method comprising: obtaining an affine transformation matrix based on image data of at least two video frames collected, and first data, the affine transformation matrix being used to indicate a mapping relationship between corresponding macroblocks between every two adjacent video frames in at least two video frames, the first data being: acceleration data and angular velocity data of an electronic device during the process of collecting at least two video frames; determining a first object based on the affine transformation matrix; and determining a first object based on the affine transformation matrix.
  • a first object encodes at least two video frames; wherein the first object includes at least one of the following: an encoding search range of each macroblock in the at least two video frames; and an intra-coded frame in the at least two video frames.
  • an embodiment of the present application provides a video encoding device, which includes an acquisition module, a determination module and an encoding module; the acquisition module is used to acquire an affine transformation matrix based on image data of at least two acquired video frames and first data, the affine transformation matrix being used to indicate a mapping relationship between corresponding macroblocks between every two adjacent video frames in at least two video frames, the first data being: acceleration data and angular velocity data of the electronic device during the process of acquiring at least two video frames; the determination module is used to determine a first object based on the affine transformation matrix acquired by the acquisition module; the encoding module is used to encode at least two video frames based on the first object determined by the determination module; wherein the first object includes at least one of the following: an encoding search range for each macroblock in at least two video frames; an internal encoding frame in at least two video frames.
  • an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
  • an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
  • an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method described in the first aspect.
  • an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
  • an affine transformation matrix can be obtained based on the image data of at least two video frames collected, and the first data, the affine transformation matrix is used to indicate the mapping relationship between the corresponding macroblocks between each two adjacent video frames in the at least two video frames, the first data is: the acceleration data and angular velocity data of the electronic device in the process of collecting at least two video frames; and based on the affine transformation matrix, the first object is determined; and based on the first object, at least two video frames are encoded; wherein the first object includes at least one of the following: the encoding search range of each macroblock in at least two video frames; the internal coding frame in at least two video frames.
  • the electronic device when the electronic device determines the encoding search range of each macroblock in the at least two video frames collected, it can be directly based on the affine transformation matrix used to indicate the mapping relationship between the corresponding macroblocks of the adjacent video frames, without traversing multiple pixel points around each pixel point in each macroblock; on the other hand, when the electronic device determines the internal coding frame in the at least two video frames, it can also be directly based on the affine transformation matrix, without comparing the pixel changes of the macroblocks in the at least two video frames. In this way, when encoding the at least two video frames based on the encoding search range and at least one of the internal encoding frames, the computing power of the electronic device can be greatly reduced, and the power consumption of the electronic device can be reduced.
  • FIG1 is a schematic diagram of a diamond search method in a conventional encoding process
  • FIG2 is a flowchart of a video encoding method according to an embodiment of the present application.
  • FIG3 is a schematic diagram of pixel affine transformation in a video encoding method provided in an embodiment of the present application.
  • FIG4 is a second flowchart of the video encoding method provided in an embodiment of the present application.
  • FIG5 is a third flowchart of the video encoding method provided in an embodiment of the present application.
  • FIG. 6 is a schematic diagram of determining a coding search range of a macroblock in a video coding method provided in an embodiment of the present application
  • FIG7 is a fourth flowchart of the video encoding method provided in an embodiment of the present application.
  • FIG8 is a fifth flowchart of the video encoding method provided in an embodiment of the present application.
  • FIG9 is a sixth flowchart of the video encoding method provided in an embodiment of the present application.
  • FIG10 is a schematic diagram of a video encoding device provided in an embodiment of the present application.
  • FIG11 is a schematic diagram of an electronic device provided in an embodiment of the present application.
  • FIG. 12 is a hardware schematic diagram of an electronic device provided in an embodiment of the present application.
  • first, second, etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first”, “second”, etc. are generally of one type, and the number of objects is not limited.
  • the first object can be one or more.
  • “and/or” in the specification and claims represents at least one of the connected objects, and the character “/" generally indicates that the objects associated with each other are in an "or” relationship.
  • At least one (item) refers to any one, any two or a combination of more than two of the objects included therein.
  • at least one (item) of a, b, and c can be represented by: “a”, “b”, “c”, “a and b”, “a and c", “b and c” and "a, b and c", where a, b, and c can be single or multiple.
  • at least two (items) refers to two or more, and its meaning is similar to that of "at least one (item)”.
  • Macroblock It is a basic concept in video coding technology. It divides the picture into blocks of different sizes to implement different compression strategies at different locations. In video coding, a video frame image is usually divided into several blocks of the same size, called macroblocks. The size of the macroblock can be 16 ⁇ 16, 16 ⁇ 8, 8 ⁇ 16 or 8 ⁇ 8, etc., but the minimum can be 4 ⁇ 4.
  • Intra-coded frame i.e. I-frame: also called intra-frame or key frame
  • I-frame Intra-coded frame
  • the picture of I-frame will be completely preserved.
  • decoding I-frame only the data of this frame is needed to reconstruct the complete image.
  • the frame method is an intra-frame compression method, also known as a key frame compression method.
  • the I frame method is a compression technology based on Discrete Cosine Transform (DCT). Using I frame compression can achieve a compression ratio of 1/6 without obvious compression marks.
  • DCT Discrete Cosine Transform
  • Unidirectional predictive coding frame (i.e. P frame): also known as difference frame, belongs to inter-frame compression.
  • the P frame represents the difference information between the current frame and the I frame, or the P frame before the current frame.
  • the P frame method compresses the data of the current frame according to the differences between the current frame and the adjacent previous frame (I frame or P frame). The method of jointly compressing P frames and I frames can achieve higher compression without obvious compression marks.
  • Bidirectional predictive coding frame also known as bidirectional difference frame
  • the encoded B frame records the difference information between the current frame (i.e. the current frame) and the previous and next frames; in other words, to decode a B frame, it is necessary not only to obtain the previous cached picture, but also to decode the next picture, and reconstruct the current frame image through the previous and next frames and the current frame coding data;
  • the B frame method is a bidirectional predictive inter-frame compression algorithm. When compressing a frame into a B frame, it compresses the current frame according to the differences between the adjacent previous frame, current frame and next frame data. Only by using B frame compression can a high compression of 200:1 be achieved.
  • the electronic anti-shake algorithm It is an algorithm that compensates for the displacement of the image in the video frames captured by the camera to achieve a smooth effect.
  • the electronic anti-shake algorithm usually includes the following steps: (1) Perform image stabilization processing on the video frames captured by the camera. Specifically, it can be achieved by calculating the displacement of each video frame relative to the previous frame. The stabilization processing will make the image transition smooth and avoid the jitter caused by camera shaking; (2) Estimate the movement of the camera. Specifically, it can be estimated by calculating the displacement change between two adjacent frames to estimate the acceleration, speed and direction of the camera during the movement; (3) According to the results of motion estimation, compensate the pixels in the video frame.
  • the method usually adopted is to perform transformations such as displacement, rotation and scaling on the video frame so that the pixel position in the video frame can be aligned with the previous frame.
  • the mainstream encoding methods mainly include Moving Pictures Experts Group (Mpeg) 1 encoding method, Mpeg2 encoding method, Mpeg4 encoding method, and H.26X encoding method.
  • Mpeg Moving Pictures Experts Group 1 encoding method
  • Mpeg2 encoding method Mpeg2 encoding method
  • Mpeg4 encoding method Mpeg4 encoding method
  • H.26X encoding method H.26X
  • Mpeg1 is the first video and audio lossy compression standard developed by the Mpeg organization. It mainly adopts block-based motion compensation, discrete cosine transform, quantization and other technologies, and is optimized for a 1.2Mbps transmission rate. Mpeg1 has the following main features: random access, flexible frame rate, variable image size, definition of I-frames, P-frames and B-frames, motion compensation that can span multiple frames, half-pixel precision motion vectors, quantization matrices, etc.
  • Mpeg2 is another video and audio lossy compression standard developed by the Mpeg organization after Mpeg1. Compared with the Mpeg1 encoding method, the Mpeg2 encoding method can make the encoded image have higher image quality, more image formats and transmission bit rates. Mpeg2 is a compression scheme for standard digital television and high-definition television in various applications, with a transmission rate between 3Mbit/s and 10Mbit/s.
  • Mpeg2 The principle of Mpeg2 is to utilize two characteristics of the image: spatial correlation and temporal correlation; any scene in a frame of image is composed of several pixels, so a pixel usually has a certain relationship with some pixels around it in brightness and chrominance, and this relationship is called spatial correlation; a plot in a program is often composed of an image sequence composed of several frames of continuous images, and an image There is also a certain relationship between the previous and next frame images in the sequence, which is called temporal correlation. These two correlations result in a large amount of redundant information in the image. Using the Mpeg2 encoding method to encode can remove this redundant information and only retain a small amount of irrelevant information for transmission, thereby greatly saving transmission bandwidth and improving encoding efficiency.
  • Mpeg4 is another video and audio lossy compression standard developed by the Mpeg organization after Mpeg2. Compared with Mpeg1 and Mpeg2, Mpeg4 is not only for video and audio encoding at a certain bit rate, but also pays more attention to the interactivity and flexibility of multimedia systems. Mpeg4 is mainly used in video calls, video emails, electronic news, etc., and its transmission rate requirements are relatively low, between 4800-64000 bits/s. Mpeg4 uses very narrow bandwidth, frame reconstruction technology, compression and transmission of data to obtain the best image quality with the least data. Mpeg4 proposes some new and innovative key technologies, including: video object extraction technology, video object plane video coding technology, video coding scalability technology, motion estimation and motion compensation technology, etc.
  • H.26X is a new generation of digital video coding standard jointly proposed by the International Organization for Standardization and the International Telecommunication Union.
  • the main parts of the H.264 standard include: access unit delimiter, additional enhancement information, basic image coding, redundant image coding, instant decoding refresh, hypothetical reference decoding, hypothetical bitstream scheduler, etc.
  • H.264 is built on the basis of Mpeg4 technology, and its encoding and decoding process mainly includes 5 parts: inter-frame and intra-frame prediction, transformation and inverse transformation, quantization and inverse quantization, loop filtering, and entropy coding.
  • the advantages of using the H.264 encoding method for encoding are as follows: 1.
  • High bit rate Under the same image quality, the amount of data compressed by H.264 technology is only 1/8 of Mpeg2 and 1/3 of Mpeg4; 2. High-quality images: H.264 can provide continuous, smooth, high-quality images; 3. Strong fault tolerance: H.264 provides the necessary tools to solve errors such as packet loss that are prone to occur in unstable network environments; 4. Strong network adaptability: H.264 provides a network abstraction layer, which allows H.264 files to be easily transmitted on different networks.
  • Step A The electronic device determines the coding search range of each macroblock in the above-mentioned multiple video frames to identify redundant information between the multiple video frames, and compresses the identified redundant information through data compression technology to reduce the amount of data during transmission and the burden of storing information.
  • the electronic device can determine it through a search algorithm by traversing multiple pixel points around each pixel point in each macroblock of the above-mentioned multiple video frames.
  • the search algorithm mainly includes diamond search method, hexagonal search method, full search method, etc.
  • pixel 11 is a vertex of a macroblock
  • the electronic device can traverse four pixel points in four directions around pixel 11, namely, pixel 12 to the left of pixel 11, pixel 13 below pixel 11, pixel 14 to the right of pixel 11, and pixel 15 above pixel 11.
  • the electronic device can determine the coding search range of the macroblock by comparing pixel values.
  • up, down, left, and right in the embodiments of the present application are all illustrated by taking the screen of the electronic device displaying an interface as an example where the screen is facing the user.
  • the electronic device determines the encoding search range of each macroblock by traversing multiple pixel points around each pixel point in each macroblock of the above-mentioned multiple video frames, the electronic device needs to perform a large amount of calculations when determining the encoding search range of each macroblock; especially when the motion amplitude of two adjacent frames is large, the large difference in image content will increase the number of pixel points traversed by the electronic device, which will further increase the computing power of the electronic device and consume additional encoding performance, thus resulting in higher power consumption of the electronic device during the video encoding process.
  • Step B The electronic device determines the internal coding frame among the plurality of video frames according to the degree of change of the image content of each video frame, so as to increase the video frame compression rate and reduce the volume of the encoded video.
  • the electronic device may first set a maximum value of the frame interval of an internal coding frame, and then select a video frame with a dramatic change in image content within the frame interval as an internal coding frame; wherein the method for determining a video frame with a dramatic change in image content is: comparing the pixel changes of macroblocks in the above-mentioned multiple video frames, and determining the video frame where the macroblock with a pixel change degree exceeding a certain threshold is located as the video frame with a dramatic change in image content.
  • the embodiments of the present application provide a video encoding method, device, electronic device and readable storage medium.
  • the video encoding method provided by the embodiments of the present application can be applied to the scenario where a mobile phone encodes collected video frames.
  • the mobile phone collects N (N is an integer greater than or equal to 2) video frames (for example, at least two video frames in the embodiment of the present application), and in the process of collecting the N video frames, obtains the acceleration data and angular velocity data of the mobile phone (for example, the first data in the embodiment of the present application); then the mobile phone can obtain a matrix (for example, the affine transformation matrix in the embodiment of the present application) for indicating the mapping relationship between the corresponding macroblocks between every two adjacent video frames in the N video frames based on the image data of the N video frames and the obtained acceleration data and angular velocity data of the mobile phone; then the mobile phone can determine at least one of the following based on the matrix: the encoding search range of each macroblock in the N video frames; the internal encoding frame in the N video frames.
  • the mobile phone can encode the N video frames based on at least one of the determined encoding search range and the internal encoding frame.
  • the mobile phone determines the coding search range of each macroblock in the above-mentioned N video frames collected, it can be directly based on the above-mentioned matrix used to indicate the mapping relationship between the corresponding macroblocks of adjacent video frames, without traversing multiple pixel points around each pixel point in each macroblock; on the other hand, when the mobile phone determines the internal coding frame in the N video frames, it can also be directly based on the matrix, without comparing the pixel changes of the macroblocks in the N video frames. In this way, when encoding the N video frames based on the coding search range and at least one of the internal coding frames, the computing power of the electronic device can be greatly reduced, and the power consumption of the electronic device can be reduced.
  • the video encoding method provided in the embodiments of the present application may be executed by a video encoding device, an electronic device, or a functional module in an electronic device, etc.
  • an electronic device executing the video encoding method is taken as an example to illustrate the video encoding method provided in the embodiments of the present application.
  • Fig. 2 shows a flow chart of a video encoding method provided by an embodiment of the present application.
  • the video encoding method provided by an embodiment of the present application may include the following steps 201 to 203.
  • Step 201 The electronic device obtains an affine transformation matrix based on the collected image data of at least two video frames and first data.
  • the affine transformation matrix is used to indicate the mapping relationship between corresponding macroblocks between every two adjacent video frames in the at least two video frames, and the first data is: acceleration data and angular velocity data of the electronic device during the process of collecting the at least two video frames.
  • each two adjacent video frames include: video frame 1 and video frame 2, video frame 2 and video frame 3, video frame 3 and video frame 4.
  • the at least two video frames are video frames continuously captured by the electronic device.
  • the at least two video frames may be captured by the electronic device through a camera in the electronic device.
  • the above-mentioned camera can be a traditional optical camera, an infrared camera, or a time-of-flight (TOF) camera in an electronic device.
  • TOF time-of-flight
  • the camera can be a telephoto camera, a short-focus camera or a zoom camera, etc.
  • the at least two video frames may be a sequence of video frames arranged in a fixed order, and the fixed order may be determined according to a time sequence in which each of the at least two video frames is captured.
  • the at least two video frames include: video frame A captured at time a, video frame B captured at time b, and video frame C captured at time c
  • the order of time a, time b, and time c is: time b, time a, time c
  • the video frame sequence arranged in a fixed order may be: video frame B, video frame A, video frame C. It can be seen that the fixed order is determined according to the order of time a, time b, and time c.
  • the above-mentioned image data may include: pixel coordinate data of pixel points or pixel values of pixel points, etc.
  • the first data may be acquired by an inertial measurement unit (IMU) sensor in the electronic device during the process of collecting the at least two video frames.
  • IMU inertial measurement unit
  • IMU sensor can be used to collect data such as acceleration, angular velocity, tilt, impact, vibration, rotation or multi-degree-of-freedom motion.
  • the IMU sensor may include a timer and a first-in-first-out (FIFO) stack; the IMU sensor may generate an interrupt at a fixed time interval through the timer to collect data, and after each data collection, save the collected sampled data to the FIFO stack, and send an interrupt signal to a processor in the electronic device to notify the processor to read the IMU sensor at a fixed time interval. In this way, the processor can read the sampled data after receiving the interrupt signal.
  • FIFO first-in-first-out
  • the above-mentioned acceleration data may be all acceleration data sampled by the above-mentioned IMU sensor in the process of acquiring the above-mentioned at least two video frames; or may be an acceleration data obtained after linear fitting of all the acceleration data; or may be an acceleration data obtained after averaging all the acceleration data.
  • the angular velocity data may be all angular velocity data sampled by the IMU sensor in the process of acquiring the at least two video frames; or may be an angular velocity data obtained after linear fitting of all angular velocity data; or may be an angular velocity data obtained after averaging all angular velocity data.
  • the mapping relationship between corresponding macroblocks between every two adjacent video frames in the at least two video frames is an affine transformation relationship between the corresponding macroblocks.
  • the corresponding macroblocks between two adjacent video frames include multiple groups of macroblocks, each group of macroblocks includes: a macroblock in one video frame of the two adjacent video frames (hereinafter referred to as macroblock 1), and a macroblock in the other video frame of the two adjacent video frames (hereinafter referred to as macroblock 2), and the degree of similarity between the pixel value of the pixel point in macroblock 1 and the pixel value of the pixel point in macroblock 2 is greater than or equal to a preset degree of similarity.
  • macroblock 1 a macroblock in one video frame of the two adjacent video frames
  • macroblock 2 a macroblock in the other video frame of the two adjacent video frames
  • affine transformation also known as affine mapping
  • affine mapping refers to a non-singular linear transformation of a vector space in geometry followed by a translation transformation to transform it into another vector space.
  • each affine transformation can be given by a matrix A and a vector b, which can be written as A and an additional column b.
  • An affine transformation corresponds to the multiplication of a matrix and a vector, while the composition of affine transformations corresponds to ordinary matrix multiplication, as long as an extra row is added to the bottom of the matrix. This row is all 0 except for a 1 on the rightmost side, and a 1 is added to the bottom of the column vector.
  • any two adjacent video frames are the i-th video frame and the i+1-th video frame in the at least two video frames, and i is a positive integer; then as shown in Figure 3, pixel 31 is the central pixel of macroblock 32 in the i-th video frame, and since the electronic device shakes during the process of acquiring the at least two video frames, the macroblock 34 corresponding to macroblock 32 in the i+1-th video frame is distorted, so that the geometric position of pixel 33 of macroblock 34 corresponding to pixel 31 is changed, and the transformation relationship of the geometric position is the affine transformation relationship between pixel 31 and pixel 33.
  • the above affine transformation relationship can be recorded by the above affine transformation matrix.
  • the affine transformation matrix can indicate the mapping relationship between corresponding macroblocks between any two adjacent video frames in the at least two video frames.
  • the above-mentioned affine transformation matrix may include at least two sub-matrices, each of the at least two sub-matrices corresponds to one of the at least two video frames, and each sub-matrix is used to indicate the geometric position of a macroblock in the corresponding video frame.
  • the geometric position of a macroblock in a video frame can be determined by the coordinate values of the pixel points in the macroblock.
  • the values in each of the above sub-matrices may include: coordinate values of a pixel point in a corresponding video frame.
  • macroblock 5 in video frame 5 includes pixel 1, pixel 2, pixel 3 and pixel 4, and the four pixel points are the four vertices of macroblock 5; if the coordinate value of pixel 1 is (2, 2), the coordinate value of pixel 2 is (4, 2), the coordinate value of pixel 3 is (2, 4), and the coordinate value of pixel 4 is (4, 4), then the geometric position of macroblock 5 in the video frame 5 can obviously be determined by the coordinate values of the four pixel points.
  • the at least two sub-matrices may also be arranged in the above fixed order.
  • any two adjacent sub-matrices in the at least two sub-matrices may indicate a mapping relationship between corresponding macroblocks between two adjacent video frames in the at least two video frames.
  • the above-mentioned affine transformation matrix may include: the above-mentioned at least two sub-matrices corresponding one-to-one to the above-mentioned at least two video frames, and each sub-matrix can be used to indicate the geometric position of a macroblock in a corresponding video frame, when the electronic device needs to obtain the mapping relationship between the corresponding macroblocks between any two adjacent video frames, it only needs to use the sub-matrices corresponding to the any two adjacent video frames without considering other sub-matrices, thereby reducing the computing power consumption of the electronic device.
  • step 201 may be specifically implemented through the following steps 201a to 201c.
  • Step 201a The electronic device inputs the image data and the first data into an electronic anti-shake algorithm.
  • the above-mentioned image data is the image data of the above-mentioned at least two video frames.
  • the above-mentioned electronic anti-shake algorithm is a commonly used algorithm for improving the shaking problem of electronic devices when shooting videos.
  • the electronic anti-shake algorithm can work in conjunction with the gyroscope in the electronic device (i.e., an angular motion detection device that uses the momentum of a high-speed rotating body to sense the angular motion of the shell relative to the inertial space around one or two axes orthogonal to the axis of rotation).
  • the electronic image stabilization algorithm can analyze and collect the image on the sensor, calculate the rotational changes in the posture of the electronic device, dynamically adjust the sensitivity, shutter, etc. to make blur corrections, and dynamically crop the video frames to reduce the impact of the shaking of the electronic device on the shooting, thereby effectively improving the stability of the video picture.
  • the electronic image stabilization algorithm may obtain a motion vector of a pixel based on the image content of a video frame or an IMU sensor to calculate a rotational change in the posture of the electronic device.
  • Step 201b The electronic device obtains pixel coordinate data of feature points in at least two video frames from the image data through an electronic anti-shake algorithm.
  • the feature points in the at least two video frames include: feature points in each video frame of the at least two video frames.
  • a feature point in a video frame may be: a vertex, a corner point, a center point in the video frame, or a point in the video frame where a grayscale value changes dramatically.
  • the feature points of an image can reflect the essential features of the image, can identify the target object in the image, and the image matching can be completed by matching the feature points.
  • pixel coordinate data of a pixel point is used to indicate: the geometric position of the pixel point in the video frame.
  • the geometric position of the pixel point in the video frame is: a pixels away from the origin (usually the upper left corner of the video frame) in the X-axis direction and b pixels away from the origin in the Y-axis direction.
  • Step 201c The electronic device calculates an affine transformation matrix based on the pixel coordinate data and the first data through an electronic anti-shake algorithm.
  • the above-mentioned pixel coordinate data is the pixel coordinate data of the feature points in the above-mentioned at least two video frames.
  • the electronic device can input the image data of the at least two video frames and the first data into the electronic anti-shake algorithm to obtain the affine transformation matrix, it can be ensured that the obtained affine transformation matrix can accurately indicate the mapping relationship between corresponding macroblocks between any two adjacent video frames in the at least two video frames based on the motion estimation function of the electronic anti-shake algorithm.
  • Step 202 The electronic device determines a first object based on an affine transformation matrix.
  • the first object includes at least one of the following:
  • the encoding search range of a macroblock is used to determine: the macroblock in the next video frame of the video frame where the macroblock is located that is most similar to the macroblock (ie, the degree of pixel value matching is greater than or equal to a preset threshold).
  • the above-mentioned coding search range may include: a coding search radius and a coding search direction.
  • the coding search radius of a macroblock is used to indicate: the coding search range of the macroblock The size of the circumference.
  • the encoding search range of macroblock 1 is: a circular range with a radius of a pixels and a geometric position of macroblock 1 in the video frame as the origin.
  • the encoding search direction of a macroblock is used to indicate: the encoding search range of the macroblock is compared to the direction where the macroblock is located.
  • the direction of the encoding search range of macroblock 2 is the lower right of macroblock 2.
  • the coding search range of the macroblock can be narrowed down to a smaller range, so as to determine the macroblock most similar to the macroblock in the next video frame of the video frame where the macroblock is located.
  • the first object includes a coding search range of each macroblock in the at least two video frames.
  • the step 202 can be implemented by the following steps 202a and 202b.
  • Step 202a The electronic device determines a motion vector corresponding to each macroblock according to at least two sub-matrices.
  • Each motion vector indicates a search range.
  • the electronic device may determine a motion vector corresponding to each macroblock of one of the at least two video frames based on any two adjacent sub-matrices of the at least two sub-matrices.
  • the process of the electronic device determining the motion vector corresponding to each macroblock of the at least two video frames may include the following steps 1 to N:
  • the electronic device determines the motion vector of each macroblock in the first video frame of the at least two video frames corresponding to the first submatrix by comparing the change between the value of the first submatrix of the at least two submatrices and the value of the second submatrix of the at least two submatrices;
  • the electronic device determines the motion vector of each macroblock in the second video frame of the at least two video frames corresponding to the second submatrix by comparing the change between the values of the second submatrix of the at least two submatrices and the values of the third submatrix of the at least two submatrices;
  • the electronic device determines the motion vector of each macroblock in the third video frame of the at least two video frames corresponding to the third submatrix by comparing the value of the third submatrix of the at least two submatrices with the change between the values of the fourth submatrix of the at least two submatrices;
  • the electronic device determines the value of the previous submatrix in the at least two video frames corresponding to the last submatrix of the at least two submatrices by comparing the change between the value of the previous submatrix and the value in the last submatrix.
  • the electronic device can determine the vectors corresponding to each macroblock of the last video frame to be 0;
  • the electronic device can determine the motion vector corresponding to each macroblock in the at least two video frames.
  • Step 202b For each macroblock, the electronic device determines the search range indicated by the motion vector corresponding to the macroblock as the encoding search range of the macroblock, and obtains the encoding search range of each macroblock.
  • the motion vector includes a size and a direction
  • a search range can be determined by the size and direction of a motion vector.
  • the xth video frame and the x+1th video frame are any two adjacent video frames of the at least two video frames, and x is a positive integer; the electronic device can obtain the geometric positions of the pixels in the xth video frame and the x+1th video frame according to the xth submatrix corresponding to the xth video frame and the x+1th submatrix corresponding to the x+1th video frame.
  • the coordinates of the pixel point 62 in the xth video frame are (50, 50), and the coordinates of the pixel point 63 in the x+1th video frame corresponding to the pixel point 62 are (100, 100), and then the motion vector corresponding to the macroblock 61 in the xth video frame can be determined; it can be seen that the search range indicated by the motion vector is: a range with a radius of 50 pixels and a direction to the lower right of the macroblock 61.
  • the electronic device can determine the range as the encoding search range of the macroblock 61.
  • the electronic device determines the search range indicated by the motion vector corresponding to each macroblock as the encoding search range of the corresponding macroblock, it obtains the encoding search range of each macroblock.
  • the electronic device can first determine the motion vector corresponding to each macroblock, and then determine the search range indicated by each motion vector as the encoding search range of the corresponding macroblock, the process of determining the encoding search range of each macroblock can be simplified, thereby greatly reducing the complexity of the calculation.
  • the first object includes an intra-coded frame in the at least two video frames.
  • the step 202 can be implemented by the following steps 202c and 202d.
  • Step 202c The electronic device determines a first change rate corresponding to each sub-matrix according to at least two sub-matrices.
  • the first change rate corresponding to a submatrix includes: a change rate of values in the submatrix compared to values in a previous submatrix of the submatrix.
  • a first change rate corresponding to a submatrix may be determined based on each value in the submatrix and a corresponding value in a previous submatrix of the submatrix.
  • submatrix 1 and submatrix 2 are two adjacent submatrices among the at least two submatrices mentioned above, and submatrix 1 is located before submatrix 2, submatrix 1 is [a], and submatrix 2 is [b]; then a and b are corresponding values, so that the electronic device can determine the first change rate corresponding to submatrix 2 as (b-a)/a based on a in submatrix 1 and b in submatrix 2.
  • submatrix 3 and submatrix 4 are two adjacent submatrices in the at least two submatrices, and the Submatrix 3 is located before submatrix 4, submatrix 3 is [c1, d1], and submatrix 4 is [c2, d2]; then c1 and c2 are corresponding values, d1 and d2 are corresponding values, so that the electronic device can first obtain the change rates of the two groups of corresponding values, namely (c2-c1)/c1 and (d2-d1)/d1, according to c1 in submatrix 1 and c2 in submatrix 2, as well as d1 in submatrix 1 and d2 in submatrix 2, and then determine the average value of (c2-c1)/c1 and (d2-d1)/d1 as the first change rate corresponding to submatrix 4.
  • the process of the electronic device determining the first change rate corresponding to each sub-matrix in the at least two sub-matrices may include the following steps a to n:
  • the electronic device determines the first change rate corresponding to the first submatrix to be 0;
  • the electronic device determines a first change rate corresponding to the second submatrix according to a value in a first submatrix of the at least two submatrices and a value in a second submatrix of the at least two submatrices;
  • the electronic device determines a first change rate corresponding to the third submatrix according to a value in a second submatrix of the at least two submatrices and a value in a third submatrix of the at least two submatrices;
  • the electronic device determines a first change rate corresponding to the fourth submatrix according to a value in a third submatrix of the at least two submatrices and a value in a fourth submatrix of the at least two submatrices;
  • the electronic device determines a first change rate corresponding to the last submatrix of the at least two submatrices according to a value in a submatrix before the last submatrix and a value in the last submatrix;
  • the electronic device can determine the first change rate corresponding to each of the above sub-matrices.
  • Step 202d The electronic device determines an internal coding frame based on the first change rate.
  • the above-mentioned first change rate is the first change rate corresponding to each of the above-mentioned sub-matrices.
  • the electronic device when the electronic device determines the internal coding frame in the at least two video frames mentioned above, it can directly determine the first change rate corresponding to each of the above sub-matrices without comparing the pixel changes of the macroblocks in the at least two video frames. Therefore, the process of determining the internal coding frame can be simplified, thereby reducing the complexity of the calculation.
  • the electronic device may determine the internal coding frame based on the first change rate corresponding to each of the sub-matrices through the following method 1 or method 2.
  • step 202d may be specifically implemented through the following step 202d1 .
  • Step 202d1 For at least one first video frame corresponding to a submatrix having a first change rate greater than or equal to a first threshold, the electronic device determines the at least one first video frame as an intra-coding frame.
  • the first threshold value may be a system default value, or may be set by a user according to actual usage requirements.
  • the first threshold may be any value such as 50%, 60% or 70%.
  • the electronic device determines that the first change rate corresponding to sub-matrix A is 0, the first change rate corresponding to sub-matrix B is 50%, and the first change rate corresponding to sub-matrix C is 68%, then the electronic device can determine the video frame corresponding to sub-matrix C as an internal coding frame.
  • the at least one first video frame is a video frame among the at least two video frames.
  • each submatrix in the at least two submatrices whose first change rate is greater than or equal to the first threshold corresponds to one first video frame in the at least one first video frame.
  • step 202d may be specifically implemented through the following steps 202d2 and 202d3 .
  • Step 202d2 For at least one second video frame corresponding to a submatrix having a first change rate less than a first threshold, the electronic device calculates a second change rate corresponding to each second video frame.
  • the second change rate corresponding to a second video frame includes: a change rate of a pixel value of a pixel in the second video frame compared to a pixel value of a pixel in a video frame preceding the second video frame.
  • Step 202d3 The electronic device determines all second video frames whose second change rate is greater than or equal to the second threshold as intra-coding frames.
  • the second threshold value may be a system default value, or may be set by a user according to actual usage requirements.
  • the second threshold may be any value such as 65%, 75% or 80%.
  • the electronic device determines that the first change rate corresponding to the sub-matrix a is 0, the first change rate corresponding to the sub-matrix b is 50%, and the first change rate corresponding to the sub-matrix c is 68%, then the electronic device can directly determine the video frame corresponding to the sub-matrix c as an internal coding frame, and then calculate the second change rate corresponding to the video frame corresponding to the sub-matrix a (hereinafter referred to as video frame a), and the second change rate corresponding to the video frame corresponding to the sub-matrix b (hereinafter referred to as video frame b); if the electronic device calculates that the second change rate corresponding to the video frame a is 0, and the second change rate corresponding to the video frame
  • the at least one second video frame is a video frame among the at least two video frames.
  • the electronic device may The second video frame above the second threshold is determined to be a P/B frame.
  • the electronic device can directly determine at least one first video frame corresponding to a submatrix whose first change rate is greater than or equal to a first threshold as the above-mentioned internal coding frame; or can first calculate the second change rate of each second video frame corresponding to a submatrix whose first change rate is less than the first threshold, and determine all second video frames whose second change rate is greater than or equal to the second threshold as the internal coding frame; therefore, the flexibility of determining the internal coding frame can be improved.
  • the electronic device may transmit the affine transformation matrix and the image data of the above-mentioned at least two video frames to an encoder in the electronic device, and then determine the above-mentioned first object based on the above-mentioned affine transformation matrix through the encoder.
  • the above-mentioned encoder is a device that compiles and converts signals (such as bit streams) or data into a signal form that can be used for communication, transmission and storage.
  • the encoder converts angular displacement or linear displacement into electrical signals.
  • the former is called a code disk and the latter is called a code scale.
  • the above-mentioned encoder can be an incremental encoder or an absolute encoder; wherein, the incremental encoder converts the displacement into a periodic electrical signal, and then converts the electrical signal into a counting pulse, and uses the number of pulses to represent the size of the displacement; each position of the absolute encoder corresponds to a certain digital code, so its indication is only related to the starting and ending positions of the measurement, but has nothing to do with the intermediate process of the measurement.
  • Step 203 The electronic device encodes at least two video frames based on the first object.
  • the electronic device encodes the at least two video frames based on the encoding search range, and can identify redundant information between the at least two video frames to compress the redundant information.
  • the electronic device encodes the at least two video frames based on the internal coding frames, which can increase the video frame compression rate and reduce the volume of the encoded video.
  • the electronic device when the electronic device determines the encoding search range of each macroblock in at least two captured video frames, it can be directly based on the affine transformation matrix used to indicate the mapping relationship between the corresponding macroblocks of adjacent video frames, without traversing multiple pixel points around each pixel point in each macroblock; on the other hand, when the electronic device determines the internal coding frame in the at least two video frames, it can also be directly based on the affine transformation matrix, without comparing the pixel changes of the macroblocks in the at least two video frames. In this way, when encoding the at least two video frames based on the encoding search range and at least one of the internal coding frames, the computing power of the electronic device can be greatly reduced, and the power consumption of the electronic device can be reduced.
  • the video encoding method provided in the embodiment of the present application can be executed by a video encoding device.
  • a video encoding device executing the video encoding method is taken as an example to illustrate the video encoding device provided in the embodiment of the present application.
  • an embodiment of the present application provides a video encoding device 10
  • the video encoding device 80 may include an acquisition module 11 , a determination module 12 and an encoding module 13 .
  • the acquisition module 11 can be used to acquire an affine transformation matrix based on the image data of at least two acquired video frames and the first data, wherein the affine transformation matrix is used to indicate the mapping relationship between the corresponding macroblocks between each two adjacent video frames in the at least two video frames, and the first data is: the acceleration data and angular velocity data of the electronic device during the acquisition of the at least two video frames.
  • the determination module 12 can be used to determine the first object based on the affine transformation matrix acquired by the acquisition module 11.
  • the encoding module 13 can be used to encode the at least two video frames based on the first object determined by the determination module 12.
  • the first object includes at least one of the following: the encoding search range of each macroblock in the at least two video frames; the internal encoding frame in the at least two video frames.
  • the affine transformation matrix may include at least two sub-matrices, each sub-matrix corresponds to one of the at least two video frames, and each sub-matrix is used to indicate the geometric position of a macroblock in the corresponding video frame.
  • the first object includes the coding search range of each macroblock in the at least two video frames.
  • the determination module 12 can be specifically used to determine the motion vector corresponding to each macroblock according to the at least two sub-matrices, each motion vector indicating a search range; and for each macroblock, determine the search range indicated by the motion vector corresponding to a macroblock as the coding search range of the macroblock, thereby obtaining the coding search range of each macroblock.
  • the first object includes an internal coding frame in the at least two video frames.
  • the determination module 12 can be specifically used to determine the first change rate corresponding to each submatrix in the at least two submatrices according to the at least two submatrices; and determine the internal coding frame based on the first change rate.
  • the first change rate corresponding to a submatrix includes: the change rate of the value in the submatrix compared to the value in the previous submatrix of the submatrix.
  • the determination module 12 can be specifically used to: for at least one first video frame corresponding to the submatrix whose first change rate is greater than or equal to the first threshold, determine the at least one first video frame as the internal coding frame; or, for at least one second video frame corresponding to the submatrix whose first change rate is less than the first threshold, calculate the second change rate corresponding to each second video frame, the second change rate corresponding to a second video frame including: the rate of change of the pixel value of the pixel in the second video frame compared to the pixel value of the pixel in the previous video frame of the second video frame; and determine all second video frames whose second change rate is greater than or equal to the second threshold as the internal coding frame.
  • the acquisition module 11 can be specifically used to input the above-mentioned image data and the above-mentioned first data into an electronic anti-shake algorithm; and through the electronic anti-shake algorithm, obtain the pixel coordinate data of the feature points in the above-mentioned at least two video frames from the image data; and through the electronic anti-shake algorithm, calculate the above-mentioned affine transformation matrix according to the pixel coordinate data and the first data.
  • the video encoding device determines at least two acquired When determining the coding search range of each macroblock in a video frame, the coding device can directly use the affine transformation matrix used to indicate the mapping relationship between the corresponding macroblocks of the adjacent video frames, without traversing multiple pixels around each pixel in each macroblock; on the other hand, when determining the internal coding frame in the at least two video frames, the video coding device can also directly use the affine transformation matrix, without comparing the pixel changes of the macroblocks in the at least two video frames. In this way, when encoding the at least two video frames based on at least one of the coding search range and the internal coding frame, the computing power and power consumption can be greatly reduced.
  • the video encoding device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip.
  • the electronic device can be a terminal or other devices other than a terminal.
  • the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR)/virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc.
  • It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., and the embodiment of the present application is not specifically limited.
  • Network Attached Storage NAS
  • PC personal computer
  • TV television
  • teller machine a self-service machine
  • the video encoding device in the embodiment of the present application may be a device having an operating system.
  • the operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
  • the video encoding device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be repeated here.
  • an embodiment of the present application also provides an electronic device 100, including a processor 101 and a memory 102, and the memory 102 stores a program or instruction that can be executed on the processor 101.
  • the program or instruction is executed by the processor 101, the various steps of the above-mentioned video encoding method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
  • the electronic devices in the embodiments of the present application include mobile electronic devices and non-mobile electronic devices.
  • FIG. 12 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
  • the electronic device 1000 includes but is not limited to components such as a radio frequency unit 1001 , a network module 1002 , an audio output unit 1003 , an input unit 1004 , a sensor 1005 , a display unit 1006 , a user input unit 1007 , an interface unit 1008 , a memory 1009 , and a processor 1010 .
  • the electronic device 1000 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1010 through a power management system, so that the power management system can manage charging, discharging, and power consumption management.
  • a power supply such as a battery
  • the electronic device structure shown in FIG12 does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than shown in the figure, or combine certain components, or arrange components differently, which will not be described in detail here.
  • the processor 1010 may be used to obtain an affine transformation matrix based on the image data of at least two video frames collected and the first data, wherein the affine transformation matrix is used to indicate that each two adjacent video frames in the at least two video frames.
  • the mapping relationship between the corresponding macroblocks between the at least two video frames is as follows: the acceleration data and the angular velocity data of the electronic device during the process of collecting the at least two video frames.
  • the processor 1010 can also be used to determine the first object based on the acquired affine transformation matrix.
  • the processor 1010 can also be used to encode the at least two video frames based on the determined first object.
  • the first object includes at least one of the following: the encoding search range of each macroblock in the at least two video frames; the internal encoding frame in the at least two video frames.
  • the affine transformation matrix may include at least two sub-matrices, each sub-matrix corresponds to one of the at least two video frames, and each sub-matrix is used to indicate the geometric position of a macroblock in the corresponding video frame.
  • the first object includes a coding search range of each macroblock in the at least two video frames.
  • the processor 1010 may be specifically configured to determine a motion vector corresponding to each macroblock according to the at least two sub-matrices, each motion vector indicating a search range; and for each macroblock, determine the search range indicated by the motion vector corresponding to a macroblock as the coding search range of the macroblock, thereby obtaining the coding search range of each macroblock.
  • the first object includes an internal coding frame in the at least two video frames.
  • the processor 1010 can be specifically configured to determine a first change rate corresponding to each of the at least two submatrices according to the at least two submatrices; and determine the internal coding frame based on the first change rate.
  • the first change rate corresponding to a submatrix includes: a change rate of a value in the submatrix compared to a value in a previous submatrix of the submatrix.
  • the processor 1010 may be specifically configured to: for at least one first video frame corresponding to a submatrix whose first change rate is greater than or equal to a first threshold, determine the at least one first video frame as the internal coding frame; or, for at least one second video frame corresponding to a submatrix whose first change rate is less than the first threshold, calculate a second change rate corresponding to each second video frame, wherein the second change rate corresponding to a second video frame includes: a rate of change of a pixel value of a pixel in the second video frame compared to a pixel value of a pixel in a video frame preceding the second video frame; and determine all second video frames whose second change rate is greater than or equal to the second threshold as the internal coding frame.
  • the processor 1010 can be specifically used to input the above-mentioned image data and the above-mentioned first data into an electronic anti-shake algorithm; and through the electronic anti-shake algorithm, obtain the pixel coordinate data of the feature points in the above-mentioned at least two video frames from the image data; and through the electronic anti-shake algorithm, calculate the above-mentioned affine transformation matrix according to the pixel coordinate data and the first data.
  • the electronic device when determining the coding search range of each macroblock in at least two captured video frames, the electronic device can directly be based on the affine transformation matrix used to indicate the mapping relationship between the macroblocks corresponding to the adjacent video frames, without traversing multiple pixels around each pixel in each macroblock; on the other hand, when determining the internal coding frame in the at least two video frames, the electronic device can also directly be based on the affine transformation matrix, without comparing the pixel changes of the macroblocks in the at least two video frames. In this way, based on the coding search range and the internal coding frame When at least one of the encodings encodes the at least two video frames, the computing power of the electronic device can be greatly reduced, and the power consumption of the electronic device can be reduced.
  • the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042, and the GPU 10041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode.
  • the display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc.
  • the user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072.
  • the touch panel 10071 is also called a touch screen.
  • the touch panel 10071 may include two parts: a touch detection device and a touch controller.
  • Other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
  • the memory 1009 can be used to store software programs and various data.
  • the memory 1009 can mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area can store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc.
  • the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both volatile and non-volatile memories.
  • the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
  • the volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM).
  • the memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
  • the processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 1010.
  • An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored.
  • a program or instruction is stored.
  • each process of the above-mentioned video encoding method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
  • the processor is the processor in the electronic device described in the above embodiment.
  • the readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
  • An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned video encoding method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
  • the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
  • An embodiment of the present application provides a computer program product, which is stored in a storage medium.
  • the program product is executed by at least one processor to implement the various processes of the above-mentioned video encoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
  • the technical solution of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM/RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
  • a storage medium such as ROM/RAM, a disk, or an optical disk
  • a terminal which can be a mobile phone, a computer, a server, or a network device, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computing Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Discrete Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

本申请公开了一种视频编码方法、装置、电子设备及可读存储介质,属于视频技术领域。该方法包括:基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,仿射变换矩阵用于指示至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,第一数据为:采集至少两个视频帧的过程中电子设备的加速度数据和角速度数据;基于仿射变换矩阵,确定第一对象;基于第一对象,编码至少两个视频帧;其中,第一对象包括以下至少之一:至少两个视频帧中每个宏块的编码搜索范围;至少两个视频帧中的内部编码帧。

Description

视频编码方法、装置、电子设备及可读存储介质
交叉引用
本申请要求在2023年06月26日提交中国专利局、申请号为202310763796.5、名称为“视频编码方法、装置、电子设备及可读存储介质”的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本申请属于视频技术领域,具体涉及一种视频编码方法、装置、电子设备及可读存储介质。
背景技术
目前,电子设备在编码采集的多个视频帧时,需执行以下至少之一:确定该多个视频帧中每个宏块的编码搜索范围,以识别该多个视频帧帧间的冗余信息,进而对该冗余信息进行压缩处理;根据每个视频帧的图像内容变化程度,确定该多个视频帧中的内部编码帧,以增加视频帧压缩率,减小编码后视频的体积。
具体地,对于上述编码搜索范围,电子设备可以通过搜索算法,遍历上述每个宏块中每个像素点周围的多个像素点确定;而对于上述内部编码帧,电子设备可以通过对比上述多个视频帧中宏块的像素变化确定。
然而,由于遍历上述多个像素点以及对比上述像素变化,均需进行大量的计算,因此使得电子设备在确定上述编码搜索范围和上述内部编码帧中至少之一时的算力较大,从而导致视频编码过程中电子设备的功耗较大。
发明内容
本申请实施例的目的是提供一种视频编码方法、装置、电子设备及可读存储介质,能够解决视频编码过程中电子设备的功耗较大的问题。
第一方面,本申请实施例提供了一种视频编码方法,该方法包括:基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,仿射变换矩阵用于指示至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,第一数据为:采集至少两个视频帧的过程中电子设备的加速度数据和角速度数据;基于仿射变换矩阵,确定第一对象;基于 第一对象,编码至少两个视频帧;其中,第一对象包括以下至少之一:至少两个视频帧中每个宏块的编码搜索范围;至少两个视频帧中的内部编码帧。
第二方面,本申请实施例提供了一种视频编码装置,该装置包括获取模块、确定模块和编码模块;获取模块,用于基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,仿射变换矩阵用于指示至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,第一数据为:采集至少两个视频帧的过程中电子设备的加速度数据和角速度数据;确定模块,用于基于获取模块获取的仿射变换矩阵,确定第一对象;编码模块,用于基于确定模块确定的第一对象,编码至少两个视频帧;其中,第一对象包括以下至少之一:至少两个视频帧中每个宏块的编码搜索范围;至少两个视频帧中的内部编码帧。
第三方面,本申请实施例提供了一种电子设备,该电子设备包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述的方法的步骤。
第四方面,本申请实施例提供了一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如第一方面所述的方法的步骤。
第五方面,本申请实施例提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述的方法。
第六方面,本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如第一方面所述的方法。
在本申请实施例中,可以基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,仿射变换矩阵用于指示至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,第一数据为:采集至少两个视频帧的过程中电子设备的加速度数据和角速度数据;且基于仿射变换矩阵,确定第一对象;并基于第一对象,编码至少两个视频帧;其中,第一对象包括以下至少之一:至少两个视频帧中每个宏块的编码搜索范围;至少两个视频帧中的内部编码帧。通过该方案,一方面电子设备在确定采集的至少两个视频帧中每个宏块的编码搜索范围时,可以直接基于用于指示相邻视频帧对应宏块的映射关系的仿射变换矩阵,而无需遍历该每个宏块中每个像素点周围的多个像素点;另一方面电子设备在确定该至少两个视频帧中的内部编码帧时,也可以直接基于该仿射变换矩阵,而无需对比该至少两个视频帧中宏块的像素变化。如此在基于该编码搜索范围和该内部编码帧中的至少之一编码该至少两个视频帧时,可以极大地减少电子设备的算力,降低电子设备的功耗。
附图说明
图1是传统编码过程中的钻石搜索法的示意图;
图2是本申请实施例提供的视频编码方法的流程图之一;
图3是本申请实施例提供的视频编码方法中像素点仿射变换的示意图;
图4是本申请实施例提供的视频编码方法的流程图之二;
图5是本申请实施例提供的视频编码方法的流程图之三;
图6是本申请实施例提供的视频编码方法中确定宏块的编码搜索范围的示意图;
图7是本申请实施例提供的视频编码方法的流程图之四;
图8是本申请实施例提供的视频编码方法的流程图之五;
图9是本申请实施例提供的视频编码方法的流程图之六;
图10是本申请实施例提供的视频编码装置的示意图;
图11是本申请实施例提供的电子设备的示意图;
图12是本申请实施例提供的电子设备的硬件示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员获得的所有其他实施例,都属于本申请保护的范围。
本申请的说明书和权利要求书中的术语“第一”、“第二”等是用于区别类似的对象,而不用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施,且“第一”、“第二”等所区分的对象通常为一类,并不限定对象的个数,例如第一对象可以是一个,也可以是多个。此外,说明书以及权利要求中“和/或”表示所连接对象的至少其中之一,字符“/”,一般表示前后关联对象是一种“或”的关系。
本申请的说明书和权利要求书中的术语“至少一个(项)”、“至少之一”等指其包含对象中的任意一个、任意两个或两个以上的组合。例如,a、b、c中的至少一个(项),可以表示:“a”、“b”、“c”、“a和b”、“a和c”、“b和c”以及“a、b和c”,其中a,b,c可以是单个,也可以是多个。同理,“至少两个(项)”是指两个或两个以上,其表达的含义与“至少一个(项)”类似。
下面首先对本申请的说明书和权利要求书中涉及的一些名词或者术语进行解释说明。
宏块:是视频编码技术中的一个基本概念,通过将画面分成一个个大小不同的块以在不同位置实行不同的压缩策略;在视频编码中,一个视频帧图像通常被划分为若干尺寸大小相同的块,称为宏块;宏块的尺寸可以为16×16、16×8、8×16或8×8等,但最小可以为4×4。
内部编码帧(即I帧):也称为内部帧或关键帧,它是帧间压缩编码里的重要帧,属于帧内压缩,I帧的画面会完整保留,解码I帧时只需要本帧数据就可以重构完整图像;I 帧法是帧内压缩法,也称为关键帧压缩法,I帧法是基于离散余弦变换(Discrete Cosine Transform,DCT)的压缩技术,采用I帧压缩可达到1/6的压缩比而无明显的压缩痕迹。
单向预测编码帧(即P帧):也称为差别帧,属于帧间压缩,P帧编码后表示的是当前帧与I帧,或与当前帧之前的P帧的差别信息;解码P帧时,需要用当前帧之前的P帧或I帧缓存的画面叠加上本帧定义的编码的差别信息,重建当前帧的画面;P帧法是根据本帧与相邻的前一帧(I帧或P帧)的不同点来压缩本帧数据,采取P帧和I帧联合压缩的方法可达到更高的压缩且无明显的压缩痕迹。
双向预测编码帧(即B帧):也称为双向差别帧,编码后的B帧记录的是本帧(即当前帧)与前后帧的差别信息;换言之,要解码B帧,不仅要取得之前的缓存画面,还要解码之后的画面,通过前后帧与本帧编码数据重建本帧图像;B帧法是双向预测的帧间压缩算法,当把一帧压缩成B帧时,它根据相邻的前一帧、本帧以及后一帧数据的不同点来压缩本帧。只有采用B帧压缩才能达到200:1的高压缩。
电子防抖算法:是在摄像头采集到的视频帧中,对图像的位移进行补偿,以达到平稳效果的算法。电子防抖算法通常包括以下步骤:(1)对摄像头采集到的视频帧进行图像稳定化处理,具体可以通过计算每个视频帧相对于前一帧的位移来实现,稳定化处理会使图像平滑过渡,避免摄像头晃动所带来的抖动;(2)对摄像头的运动进行估计,具体可以通过计算相邻两帧之间的位移变化,估算出摄像头在运动过程中的加速度、速度和方向等信息;(3)根据运动估计的结果,对视频帧中的像素进行补偿,通常采用的方法是对视频帧进行位移、旋转和缩放等变换,使得视频帧中的像素位置能够与前一帧对齐。
下面结合附图,通过具体的实施例及其应用场景对本申请实施例提供的视频编码方法、装置、电子设备及可读存储介质进行详细地说明。
主流的编码方法主要包括动态图像专家组(Moving Pictures Experts Group,Mpeg)1编码方法、Mpeg2编码方法、Mpeg4编码方法、H.26X编码方法。
其中,Mpeg1是Mpeg组织制定的第一个视频和音频有损压缩标准,主要采用了块方式的运动补偿、离散余弦变换、量化等技术,并为1.2Mbps传输速率进行了优化。Mpeg1主要具有以下特点:随机访问,灵活的帧率,可变的图像尺寸,定义了I-帧、P-帧和B-帧,运动补偿可跨越多个帧,半像素精度的运动向量,量化矩阵等。
Mpeg2是Mpeg组织继Mpeg1之后制定的又一视频和音频有损压缩标准,与Mpeg1编码方法相比,使用Mpeg2编码方法进行编码,可以使编码后的图像具有更高的图像质量、更多的图像格式和传输码率。Mpeg2是针对标准数字电视和高清晰电视在各种应用下的压缩方案,传输速率在3Mbit/s~10Mbit/s之间。Mpeg2的原理是利用了图像中的两种特性:空间相关性和时间相关性;一帧图像内的任何一个场景都是由若干像素点构成的,因此一个像素通常与它周围的某些像素在亮度和色度上存在一定的关系,这种关系叫作空间相关性;一个节目中的一个情节常常由若干帧连续图像组成的图像序列构成,一个图像 序列中前后帧图像间也存在一定的关系,这种关系叫作时间相关性。这两种相关性使得图像中存在大量的冗余信息,使用Mpeg2编码方法进行编码可以将这些冗余信息去除,只保留少量非相关信息进行传输,从而大大节省传输频带,提升编码效率。
Mpeg4是Mpeg组织继Mpeg2之后制定的又一视频和音频有损压缩标准,相较于Mpeg1和Mpeg2,Mpeg4不仅是针对一定比特率下的视频、音频编码,还更加注重多媒体系统的交互性和灵活性。Mpeg4主要应用于视像电话、视像电子邮件、电子新闻等,其传输速率要求较低,在4800-64000bits/s之间。Mpeg4利用很窄的带宽,通过帧重建技术、压缩和传输数据,以最少的数据获得最佳的图像质量。Mpeg4提出了一些新的有创见性的关键技术,包括:视频对象提取技术、视频对象平面视频编码技术、视频编码可分级性技术、运动估计与运动补偿技术等。
H.26X是国际标准化组织和国际电信联盟共同提出的新一代数字视频编码标准。以H.264为例,H.264标准的主要部分包括:访问单元分割符、附加增强信息、基本图像编码、冗余图像编码、即时解码刷新、假想参考解码、假想码流调度器等。H.264是在Mpeg4技术的基础之上建立起来的,其编解码流程主要包括5个部分:帧间和帧内预测、变换和反变换、量化和反量化、环路滤波、熵编码。使用H.264编码方法进行编码存在以下优势:1、低码率:在同等图像质量下,采用H.264技术压缩后的数据量只有Mpeg2的1/8,Mpeg4的1/3;2、高质量的图像:H.264能够提供连续、流畅的高质量图像;3、容错能力强:H.264提供了解决在不稳定网络环境下容易发生的丢包等错误的必要工具;4、网络适应性强:H.264提供了网络抽象层,使得H.264的文件可以容易地在不同网络上传输。
目前,电子设备采用上述任一种编码方法对摄像头采集的多个视频帧进行编码时,通常需执行下述的步骤A和步骤B中的至少之一:
步骤A、电子设备确定上述多个视频帧中每个宏块的编码搜索范围,以识别该多个视频帧帧间的冗余信息,并通过数据压缩技术将识别的冗余信息进行压缩,以减少传输过程中数据量以及存储信息的负担。
其中,对于上述编码搜索范围(包括编码搜索半径和编码搜索方向),电子设备可以通过搜索算法,遍历上述多个视频帧的每个宏块中的每个像素点周围的多个像素点确定,该搜索算法主要包括钻石搜索法、六边形搜索法、全搜索法等。
例如,以上述钻石搜索法为例,如图1所示,像素点11为一个宏块的一个顶点,电子设备可以遍历像素点11周围4个方向上的4个像素点,即位于像素点11左方的像素点12、位于像素点11下方的像素点13、位于像素点11右方的像素点14、位于像素点11上方的像素点15。从而电子设备在遍历该宏块的每个像素点周围的4个像素点之后,可以通过像素值比对确定出该宏块的编码搜索范围。
需要说明的是,本申请实施例中的上下左右,均是以电子设备的屏幕显示界面时,屏幕朝向用户为例进行示意的。
可以理解,由于电子设备是通过遍历上述多个视频帧的每个宏块中的每个像素点周围的多个像素点,确定该每个宏块的编码搜索范围的,因此电子设备在确定该每个宏块的编码搜索范围时,需要进行大量的计算;尤其是在相邻两帧运动幅度较大时,较大的图像内容差异会提升电子设备遍历的像素点的数量,从而会进一步加大电子设备的算力,额外耗费编码性能,如此导致视频编码过程中电子设备的功耗较大。
步骤B、电子设备根据每个视频帧的图像内容的变化程度,确定上述多个视频帧中的内部编码帧,以增加视频帧压缩率,减小编码后视频的体积。
具体地,电子设备可以先设置一个内部编码帧帧间隔的最大数值,然后在帧间隔范围内选择图像内容剧烈变化的视频帧作为内部编码帧;其中,图像内容剧烈变化的视频帧的确定方法为:对比上述多个视频帧中宏块的像素变化,并将像素变化程度超过一定阈值的宏块所在的视频帧确定为图像内容剧烈变化的视频帧。显然,由于对比该多个视频帧中宏块的像素变化同样需要电子设备进行大量的计算,因此电子设备也需耗费较大算力才能确定出该多个视频帧中合适的内容编码帧,从而也会导致视频编码过程中电子设备的功耗较大。
为了解决视频编码过程中电子设备的功耗较大的问题,本申请实施例提供一种视频编码方法、装置、电子设备及可读存储介质,本申请实施例提供的视频编码方法可以应用于手机对采集的视频帧进行编码的场景中。
示例性地,手机采集了N(N为大于或等于2的整数)个视频帧(例如本申请实施例中的至少两个视频帧),并在采集该N个视频帧的过程中,获取了手机的加速度数据和角速度数据(例如本申请实施例中的第一数据);然后手机可以基于该N个视频帧的图像数据,以及获取的手机的加速度数据和角速度数据,获取一个用于指示该N个视频帧中每两个相邻视频帧间对应宏块的映射关系的矩阵(例如本申请实施例中的仿射变换矩阵);之后手机可以基于该矩阵确定以下至少之一:该N个视频帧中每个宏块的编码搜索范围;该N个视频帧中的内部编码帧。从而手机可以基于确定的该编码搜索范围和该内部编码帧中的至少之一,编码该N个视频帧。
通过本申请的方案,一方面手机在确定采集的上述N个视频帧中每个宏块的编码搜索范围时,可以直接基于用于指示相邻视频帧对应宏块的映射关系的上述矩阵,而无需遍历该每个宏块中每个像素点周围的多个像素点;另一方面手机在确定该N个视频帧中的内部编码帧时,也可以直接基于该矩阵,而无需对比该N个视频帧中宏块的像素变化。如此在基于该编码搜索范围和该内部编码帧中的至少之一编码该N个视频帧时,可以极大地减少电子设备的算力,降低电子设备的功耗。
需要说明的是,本申请实施例提供的视频编码方法,执行主体可以为视频编码装置、电子设备或电子设备中的功能模块等。本申请的一些实施例中以电子设备执行视频编码方法为例,说明本申请实施例提供的视频编码方法。
图2示出了本申请实施例提供的视频编码方法的流程图。如图2所示,本申请实施例提供的视频编码方法可以包括下述的步骤201至步骤203。
步骤201、电子设备基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵。
其中,上述仿射变换矩阵用于指示上述至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,上述第一数据为:采集该至少两个视频帧的过程中电子设备的加速度数据和角速度数据。
例如,假设上述至少两个视频帧包括:视频帧1、视频帧2、视频帧3和视频帧4;那么上述每两个相邻视频帧包括:视频帧1和视频帧2,视频帧2和视频帧3,视频帧3和视频帧4。
可选地,本申请实施例中,上述至少两个视频帧是电子设备连续采集的视频帧。
可选地,本申请实施例中,上述至少两个视频帧可以是电子设备通过电子设备中的摄像头采集的。
可选地,本申请实施例中,上述摄像头可以为电子设备中的传统光学摄像头、红外摄像头或飞行时间(Time of Flight,TOF)摄像头等。
例如,以上述摄像头为上述传统光学摄像头为例,该摄像头可以为长焦摄像头、短焦摄像头或变焦摄像头等。
可选地,本申请实施例中,上述至少两个视频帧可以是按照固定顺序排列的视频帧序列,该固定顺序可以根据该至少两个视频帧中每个视频帧被采集的时间先后顺序确定。
例如,假设上述至少两个视频帧包括:在时间a采集的视频帧A,在时间b采集的视频帧B,以及在时间c采集的视频帧C,且时间a、时间b和时间c的先后顺序为:时间b、时间a、时间c;那么上述按照固定顺序排列的视频帧序列可以为:视频帧B、视频帧A、视频帧C。可以看出,该固定顺序是根据时间a、时间b和时间c的先后顺序确定的。
可选地,本申请实施例中,上述图像数据可以包括:像素点的像素坐标数据或像素点的像素值等。
可选地,本申请实施例中,上述第一数据可以是电子设备在采集上述至少两个视频帧的过程中,通过电子设备中的惯性测量单元(Inertial Measurement Unit,IMU)传感器获取的。
需要说明的是,上述IMU传感器可以用于采集加速度、角速度、倾斜、冲击、振动、旋转或多自由度运动等数据。
可选地,本申请实施例中,上述IMU传感器可以包括定时器和先进先出(First Input First Output,FIFO)堆栈;该IMU传感器可以通过该定时器以固定时间间隔产生中断,进行数据采集,并在每次采集数据之后将采集的采样数据记录保存至该FIFO堆栈,同时向电子设备中的处理器发送中断信号,以通知该处理器读取该IMU传感器以固定时间间 隔采集的采样数据。如此,该处理器可以在接收到该中断信号之后,对该采样数据进行读取。
可选地,本申请实施例中,上述加速度数据可以为在采集上述至少两个视频帧的过程中,上述IMU传感器采样的所有加速度数据;或可以为对该所有加速度数据进行线性拟合后得到的一个加速度数据;或可以为对该所有加速度数据进行平均后得到的一个加速度数据。
可选地,本申请实施例中,上述角速度数据可以为在采集上述至少两个视频帧的过程中,上述IMU传感器采样的所有角速度数据;或可以为对该所有角速度数据进行线性拟合后得到的一个角速度数据;或可以为对该所有角速度数据进行平均后得到的一个角速度数据。
可选地,本申请实施例中,上述至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,即该对应宏块的仿射变换关系。
可选地,本申请实施例中,两个相邻视频帧间的对应宏块包括多组宏块,每组宏块包括:该两个相邻视频帧的一个视频帧中的一个宏块(以下称为宏块1),以及该两个相邻视频帧的另一个视频帧中的一个宏块(以下称为宏块2),且宏块1中像素点的像素值与宏块2中像素点的像素值的相似程度大于或等于预设相似程度。
可以理解,上述每组宏块中的两个宏块在各自视频帧的图像中对应同一对象。
例如,假设视频帧a和视频帧b为两个相邻的视频帧,且该两个视频帧记录的是飞鸟的飞行轨迹,若该视频帧a中的宏块A对应该视频帧a的图像中的飞鸟,则与该宏块A对应的该视频帧b中的宏块B,对应该视频帧b的图像中的飞鸟。
需要说明的是,仿射变换又称仿射映射,是指几何中一个向量空间进行一次非奇异的线性变换并接上一个平移变换,变换为另一个向量空间。在有限维的情况,每个仿射变换可以由一个矩阵A和一个向量b给出,它可以写作A和一个附加的列b。一个仿射变换对应于一个矩阵和一个向量的乘法,而仿射变换的复合对应于普通的矩阵乘法,只要加入一个额外的行到矩阵的底下,这一行全部是0除了最右边是一个1,而列向量的底下要加上一个1。
下面以宏块中的一个像素点为例,对上述仿射变换关系进行示例性地说明。
示例性地,假设上述任意两个相邻视频帧为上述至少两个视频帧中的第i个视频帧和第i+1个视频帧,i为正整数;那么如图3所示,像素点31为第i个视频帧中宏块32的中心像素点,而由于电子设备在采集上述至少两个视频帧的过程中存在抖动,因此使得第i+1个视频帧中与宏块32对应的宏块34出现了扭曲,从而与像素点31对应的宏块34的像素点33的几何位置发生了变换,该几何位置的变换关系即像素点31与像素点33的仿射变换关系。
可选地,本申请实施例中,上述仿射变换关系可以通过上述仿射变换矩阵进行记录, 使得该仿射变换矩阵可以指示上述至少两个视频帧中任意两个相邻视频帧间对应宏块的映射关系。
可选地,本申请实施例中,上述仿射变换矩阵可以包括至少两个子矩阵,该至少两个子矩阵中的每个子矩阵与上述至少两个视频帧中的一个视频帧所对应,该每个子矩阵用于指示所对应的一个视频帧中宏块所处的几何位置。
可选地,本申请实施例中,一个宏块在一个视频帧中所处的几何位置,可以通过该宏块中像素点的坐标值确定。
可选地,本申请实施例中,上述每个子矩阵中的数值可以包括:所对应的一个视频帧中像素点的坐标值。
例如,假设视频帧5中的宏块5包括像素点1、像素点2、像素点3和像素点4,且该四个像素点为该宏块5的四个顶点;若该像素点1的坐标值为(2,2),该像素点2的坐标值为(4,2),该像素点3的坐标值为(2,4),该像素点4的坐标值为(4,4),则该宏块5在该视频帧5中的几何位置显然可以由该四个像素点的坐标值确定出。
可选地,本申请实施例中,上述至少两个子矩阵也可以是按照上述固定顺序排列的。
可选地,本申请实施例中,通过上述至少两个子矩阵中任意相邻的两个子矩阵,可以指示上述至少两个视频帧中两个相邻视频帧间对应宏块的映射关系。
本申请实施例中,由于上述仿射变换矩阵可以包括:与上述至少两个视频帧一一对应的上述至少两个子矩阵,且每个子矩阵可以用于指示所对应的一个视频帧中宏块的几何位置,因此当电子设备需获取上述任意两个相邻视频帧间对应宏块的映射关系时,只需通过该任意两个相邻视频帧对应的子矩阵,而无需考虑其它子矩阵,从而可以降低电子设备的算力功耗。
可选地,本申请实施例中,结合图1,如图4所示,上述步骤201具体可以通过下述的步骤201a至步骤201c实现。
步骤201a、电子设备将图像数据和第一数据输入电子防抖算法。
本申请实施例中,上述图像数据即上述至少两个视频帧的图像数据。
本申请实施例中,上述电子防抖算法是用于改善电子设备在拍摄视频时的抖动问题的常用算法,该电子防抖算法可以配合电子设备中的陀螺仪(即用高速回转体的动量矩敏感壳体相对惯性空间绕正交于自转轴的一个或二个轴的角运动检测装置)工作。
具体地,在拍摄视频的过程中,当上述陀螺仪检测到电子设备的震动时,上述电子防抖算法可以对传感器上的图像进行分析、采集,并通过计算电子设备姿态的旋转变化,动态调整感光度、快门等来做模糊修正,以及对视频帧进行动态的裁切,以减轻电子设备抖动对拍摄的影响,有效提高视频画面稳定性。
可选地,本申请实施例中,上述电子防抖算法可以基于视频帧的图像内容或IMU传感器,获取像素的运动向量,以计算电子设备姿态的旋转变化。
对上述电子防抖算法的具体描述,可以参照相关技术中的相关描述,为了避免重复,此处不再赘述。
步骤201b、电子设备通过电子防抖算法,从图像数据中获取至少两个视频帧中特征点的像素坐标数据。
可选地,本申请实施例中,上述至少两个视频帧中的特征点包括:该至少两个视频帧的每个视频帧中的特征点。
可选地,本申请实施例中,一个视频帧中的特征点可以为:该视频帧中的顶点、角点、中心点或该视频帧中灰度值发生剧烈变化的点等。
可选地,本申请实施例中,一个图像的特征点能够反映该图像的本质特征,能够标识该图像中的目标对象,通过特征点的匹配能够完成图像的匹配。
可选地,本申请实施例中,一个像素点的像素坐标数据用于指示:该像素点在视频帧中所处的几何位置。
例如,若一个像素点的坐标数据为(a,b),则该像素点在视频帧中所处的几何位置为:在X轴方向上距离原点(通常为视频帧的左上角点)a像素,在Y轴方向上距离原点b像素的位置。
步骤201c、电子设备通过电子防抖算法,根据像素坐标数据和第一数据,计算得到仿射变换矩阵。
本申请实施例中,上述像素坐标数据即上述至少两个视频帧中特征点的像素坐标数据。
对电子设备根据像素坐标数据,以及角速度数据和加速度数据,计算仿射变换矩阵的具体方法,可以参照相关技术中的相关描述,为了避免重复,此处不再赘述。
本申请实施例中,由于电子设备可以将上述至少两个视频帧的图像数据以及第一数据输入上述电子防抖算法,以得到上述仿射变换矩阵,因此可以基于该电子防抖算法的运动估计功能,确保得到的该仿射变换矩阵可以准确指示该至少两个视频帧中任意两个相邻视频帧间对应宏块的映射关系。
步骤202、电子设备基于仿射变换矩阵,确定第一对象。
其中,上述第一对象包括以下至少之一:
上述至少两个视频帧中每个宏块的编码搜索范围;
上述至少两个视频帧中的内部编码帧。
本申请实施例中,一个宏块的编码搜索范围用于确定:该宏块所在视频帧的下一视频帧中,与该宏块最相似(即像素值的匹配程度大于或等于预设阈值)的宏块。
可选地,本申请实施例中,上述编码搜索范围可以包括:编码搜索半径和编码搜索方向。
可选地,本申请实施例中,一个宏块的编码搜索半径用于指示:该宏块的编码搜索范 围的大小。
例如,假设宏块1的编码搜索半径为a像素,那么该宏块1的编码搜索范围为:以宏块1在视频帧中所处的几何位置为原点,半径为a像素的圆形范围。
可选地,本申请实施例中,一个宏块的编码搜索方向用于指示:该宏块的编码搜索范围相较于该宏块所在的方向。
例如,假设宏块2的编码搜索方向为右下方,那么该宏块2的编码搜索范围所在的方向为该宏块2的右下方。
可以看出,通过一个宏块的编码搜索半径和编码搜索方向,可以将该宏块的编码搜索范围缩小为某个较小的范围,以便于确定该宏块所在视频帧的下一视频帧中,与该宏块最相似的宏块。
对上述内部编码帧的具体描述,可以参照上述对本申请的说明书和权利要求书中涉及的一些名词或者术语的解释说明中的相关描述,为了避免重复,此处不再赘述。
下面对电子设备确定上述第一对象的具体方法进行详细地说明。
可选地,本申请实施例中,上述第一对象包括上述至少两个视频帧中每个宏块的编码搜索范围。示例性地,结合图1,如图5所示,上述步骤202具体可以通过下述的步骤202a和步骤202b实现。
步骤202a、电子设备根据至少两个子矩阵,确定每个宏块对应的运动向量。
其中,每个运动向量指示一个搜索范围。
可选地,本申请实施例中,电子设备可以根据上述至少两个子矩阵中任意相邻的两个子矩阵,确定上述至少两个视频帧中一个视频帧的每个宏块对应的运动向量。
具体地,电子设备确定上述至少两个视频帧的每个宏块对应的运动向量的过程可以包括下述的1~N:
1、电子设备通过对比上述至少两个子矩阵中的第一个子矩阵的数值,与该至少两个子矩阵中的第二个子矩阵的数值间的变化,确定该第一个子矩阵对应的上述至少两个视频帧中的第一个视频帧中每个宏块的运动向量;
2、电子设备通过对比上述至少两个子矩阵中的第二个子矩阵的数值,与该至少两个子矩阵中的第三个子矩阵的数值间的变化,确定该第二个子矩阵对应的上述至少两个视频帧中的第二个视频帧中每个宏块的运动向量;
3、电子设备通过对比上述至少两个子矩阵中的第三个子矩阵的数值,与该至少两个子矩阵中的第四个子矩阵的数值间的变化,确定该第三个子矩阵对应的上述至少两个视频帧中的第三个视频帧中每个宏块的运动向量;
……
N-1、电子设备通过对比上述至少两个子矩阵中最后一个子矩阵的前一子矩阵的数值,与该最后一个子矩阵中的数值间的变化,确定该前一子矩阵对应的上述至少两个视频帧中 的最后一个视频帧的前一视频帧中每个宏块的运动向量;
N、由于上述至少两个视频帧中的最后一个视频帧的宏块无需再进行运动估计,因此电子设备可以将该最后一个视频帧的每个宏块对应的向量均确定为0;
如此,电子设备可以确定上述至少两个视频帧中每个宏块对应的运动向量。
步骤202b、对于每个宏块,电子设备将一个宏块对应的运动向量指示的搜索范围,确定为一个宏块的编码搜索范围,得到每个宏块的编码搜索范围。
本申请实施例中,运动向量包括大小和方向,通过一个运动向量的大小和方向可以确定一个搜索范围。
示例性地,如图6所示,第x个视频帧和第x+1个视频帧为上述至少两个视频帧中的任意两个相邻视频帧,x为正整数;电子设备可以根据该第x个视频帧对应的第x个子矩阵,以及该第x+1个视频帧对应的第x+1个子矩阵,获取该第x个视频帧和该第x+1个视频帧中像素的几何位置,例如,该第x个视频帧中像素点62的坐标为(50,50),与该像素点62对应的该第x+1个视频帧中的像素点63的坐标为(100,100),进而可以确定该第x个视频帧中宏块61对应的运动向量;可以看出,该运动向量指示的搜索范围为:半径为50像素、方向为该宏块61的右下方的范围。从而电子设备可以将该范围确定为该宏块61的编码搜索范围。
可以理解,电子设备在将上述每个宏块对应的运动向量指示的搜索范围,确定为对应宏块的编码搜索范围之后,得到该每个宏块的编码搜索范围。
本申请实施例中,由于电子设备可以先确定每个宏块对应的运动向量,然后将每个运动向量指示的搜索范围确定为对应宏块的编码搜索范围,因此可以简化确定该每个宏块的编码搜索范围的过程,从而可以极大地减少计算的复杂度。
可选地,本申请实施例中,上述第一对象包括上述至少两个视频帧中的内部编码帧。示例性地,结合图1,如图7所示,上述步骤202具体可以通过下述的步骤202c和步骤202d实现。
步骤202c、电子设备根据至少两个子矩阵,确定每个子矩阵对应的第一变化率。
其中,一个子矩阵对应的第一变化率包括:该一个子矩阵中的数值,相较于该一个子矩阵的前一子矩阵中的数值的变化率。
可选地,本申请实施例中,一个子矩阵对应的第一变化率可以根据该一个子矩阵中的每个数值,与该一个子矩阵的前一子矩阵中的对应数值确定。
例如,假设子矩阵1和子矩阵2为上述至少两个子矩阵中相邻的两个子矩阵,且该子矩阵1位于该子矩阵2之前,该子矩阵1为[a],该子矩阵2为[b];那么a和b为相对应的数值,从而电子设备可以根据该子矩阵1中的a和该子矩阵2中的b,确定该子矩阵2对应的第一变化率为(b-a)/a。
又例如,假设子矩阵3和子矩阵4为上述至少两个子矩阵中相邻的两个子矩阵,且该 子矩阵3位于该子矩阵4之前,该子矩阵3为[c1,d1],该子矩阵4为[c2,d2];那么c1和c2为相对应的数值,d1和d2为相对应的数值,从而电子设备可以根据该子矩阵1中的c1和该子矩阵2中的c2,以及该子矩阵1中的d1和该子矩阵2中的d2,先分别得到两组对应数值各自的变化率,即(c2-c1)/c1和(d2-d1)/d1,然后将(c2-c1)/c1与(d2-d1)/d1的平均值确定为该子矩阵4对应的第一变化率。
可选地,本申请实施例中,电子设备确定上述至少两个子矩阵中每个子矩阵对应的第一变化率的过程可以包括下述的a~n:
a、由于上述至少两个子矩阵中第一个子矩阵之前不存在子矩阵,因此该第一个子矩阵中的数值,不存在相较于该第一个子矩阵的前一子矩阵中的数值的变化率,从而电子设备将该第一个子矩阵对应的第一变化率确定为0;
b、电子设备根据上述至少两个子矩阵中第一个子矩阵中的数值,与该至少两个子矩阵中第二个子矩阵中的数值,确定该第二个子矩阵对应的第一变化率;
c、电子设备根据上述至少两个子矩阵中第二个子矩阵中的数值,与该至少两个子矩阵中第三个子矩阵中的数值,确定该第三个子矩阵对应的第一变化率;
d、电子设备根据上述至少两个子矩阵中第三个子矩阵中的数值,与该至少两个子矩阵中第四个子矩阵中的数值,确定该第四个子矩阵对应的第一变化率;
……
n、电子设备根据上述至少两个子矩阵中最后一个子矩阵的前一子矩阵中的数值,与该最后一个子矩阵中的数值,确定该最后一个子矩阵对应的第一变化率;
如此,电子设备可以确定上述每个子矩阵对应的第一变化率。
步骤202d、电子设备基于第一变化率,确定内部编码帧。
本申请实施例中,上述第一变化率即上述每个子矩阵对应的第一变化率。
本申请实施例中,由于电子设备在确定上述至少两个视频帧中的内部编码帧时,可以直接基于确定的上述每个子矩阵对应的第一变化率,而无需再对比该至少两个视频帧中宏块的像素变化,因此可以简化确定内部编码帧的过程,从而可以减少计算的复杂度。
下面对电子设备基于上述每个子矩阵对应的第一变化率,确定上述内部编码帧的具体方法进行详细地说明。
可选地,本申请实施例中,电子设备可以基于上述每个子矩阵对应的第一变化率,通过下述的方式一或方式二,确定上述内部编码帧。
方式一
可选地,本申请实施例中,结合图7,如图8所示,上述步骤202d具体可以通过下述的步骤202d1实现。
步骤202d1、对于第一变化率大于或等于第一阈值的子矩阵所对应的至少一个第一视频帧,电子设备将至少一个第一视频帧确定为内部编码帧。
可选地,本申请实施例中,上述第一阈值可以为系统默认的,或可以为用户根据实际使用需求设置的。
可选地,本申请实施例中,上述第一阈值可以为50%、60%或70%等任意的数值。
示例性地,假设上述至少两个子矩阵依次为子矩阵A、子矩阵B和子矩阵C,且上述第一阈值为60%,那么若电子设备确定该子矩阵A对应的第一变化率为0,该子矩阵B对应的第一变化率为50%,该子矩阵C对应的第一变化率为68%,则电子设备可以将该子矩阵C所对应的视频帧确定为内部编码帧。
本申请实施例中,上述至少一个第一视频帧为上述至少两个视频帧中的视频帧。
可选地,本申请实施例中,上述至少两个子矩阵中第一变化率大于或等于第一阈值的每个子矩阵,对应上述至少一个第一视频帧中的一个第一视频帧。
方式二
可选地,本申请实施例中,结合图7,如图9所示,上述步骤202d具体可以通过下述的步骤202d2和步骤202d3实现。
步骤202d2、对于第一变化率小于第一阈值的子矩阵所对应的至少一个第二视频帧,电子设备计算每个第二视频帧对应的第二变化率。
其中,一个第二视频帧对应的第二变化率包括:该一个第二视频帧中像素的像素值,相较于该一个第二视频帧的前一视频帧中像素的像素值的变化率。
对电子设备计算上述每个第二视频帧对应的第二变化率的具体描述,可以参照相关技术中的相关描述,为了避免重复,此处不再赘述。
步骤202d3、电子设备将第二变化率大于或等于第二阈值的所有第二视频帧,确定为内部编码帧。
可选地,本申请实施例中,上述第二阈值可以为系统默认的,或可以为用户根据实际使用需求设置的。
可选地,本申请实施例中,上述第二阈值可以为65%、75%或80%等任意的数值。
示例性地,假设上述至少两个子矩阵依次为子矩阵a、子矩阵b和子矩阵c,且上述第一阈值为60%,上述第二阈值为75%,那么若电子设备确定该子矩阵a对应的第一变化率为0,该子矩阵b对应的第一变化率为50%,该子矩阵c对应的第一变化率为68%,则电子设备可以将该子矩阵c所对应的视频帧直接确定为内部编码帧,然后计算该子矩阵a所对应的视频帧(以下称为视频帧a)对应的第二变化率,以及该子矩阵b所对应的视频帧(以下称为视频帧b)对应的第二变化率;若电子设备计算得到的该视频帧a对应的第二变化率为0,该视频帧b对应的第二变化率为78%,则电子设备可以将该视频帧b也确定为内部编码帧。
本申请实施例中,上述至少一个第二视频帧为上述至少两个视频帧中的视频帧。
可选地,本申请实施例中,电子设备可以将上述至少一个第二视频帧中第二变化率小 于上述第二阈值的第二视频帧,确定为P/B帧。
本申请实施例中,由于电子设备可以直接将第一变化率大于或等于第一阈值的子矩阵所对应的至少一个第一视频帧,确定为上述内部编码帧;或者可以先计算每个第一变化率小于第一阈值的子矩阵所对应的第二视频帧的第二变化率,并将第二变化率大于或等于第二阈值的所有第二视频帧,确定为该内部编码帧;因此可以提高确定内部编码帧的灵活性。
可选地,本申请实施例中,电子设备在获取上述仿射变换矩阵之后,可以将该仿射变换矩阵以及上述至少两个视频帧的图像数据传输到电子设备中的编码器,然后通过该编码器基于上述仿射变换矩阵,确定上述第一对象。
可选地,本申请实施例中,上述编码器是将信号(例如比特流)或数据进行编制、转换为可用于通讯、传输和存储的信号形式的设备,编码器把角位移或直线位移转换成电信号,前者称为码盘,后者称为码尺。
可选地,本申请实施例中,上述编码器可以为增量式编码器或绝对式编码器;其中,增量式编码器是将位移转换成周期性的电信号,再把这个电信号转变成计数脉冲,用脉冲的个数表示位移的大小;绝对式编码器的每一个位置对应一个确定的数字码,因此它的示值只与测量的起始和终止位置有关,而与测量的中间过程无关。
步骤203、电子设备基于第一对象,编码至少两个视频帧。
可选地,本申请实施例中,电子设备基于上述编码搜索范围编码上述至少两个视频帧,可以识别出该至少两个视频帧帧间的冗余信息,以对该冗余信息进行压缩处理。
可选地,本申请实施例中,电子设备基于上述内部编码帧编码上述至少两个视频帧,可以增加视频帧压缩率,减小编码后视频的体积。
对电子设备编码视频帧的具体方法,可以参照相关技术中的相关描述,为了避免重复,此处不再赘述。
在本申请实施例提供的视频编码方法中,一方面电子设备在确定采集的至少两个视频帧中每个宏块的编码搜索范围时,可以直接基于用于指示相邻视频帧对应宏块的映射关系的仿射变换矩阵,而无需遍历该每个宏块中每个像素点周围的多个像素点;另一方面电子设备在确定该至少两个视频帧中的内部编码帧时,也可以直接基于该仿射变换矩阵,而无需对比该至少两个视频帧中宏块的像素变化。如此在基于该编码搜索范围和该内部编码帧中的至少之一编码该至少两个视频帧时,可以极大地减少电子设备的算力,降低电子设备的功耗。
上述各个方法实施例,或者各个方法实施例中的各种可能的实现方式可以单独执行,或者,在不存在矛盾的前提下,也可以相互结合执行,具体可以根据实际使用需求确定,本申请实施例对此不做限制。
本申请实施例提供的视频编码方法,执行主体可以为视频编码装置。本申请实施例中以视频编码装置执行视频编码方法为例,说明本申请实施例提供的视频编码装置。
如图10所示,本申请实施例提供一种视频编码装置10,该视频编码装置80可以包括获取模块11、确定模块12和编码模块13。
其中,获取模块11,可以用于基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,该仿射变换矩阵用于指示该至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,该第一数据为:采集该至少两个视频帧的过程中电子设备的加速度数据和角速度数据。确定模块12,可以用于基于获取模块11获取的该仿射变换矩阵,确定第一对象。编码模块13,可以用于基于确定模块12确定的该第一对象,编码该至少两个视频帧。其中,该第一对象包括以下至少之一:该至少两个视频帧中每个宏块的编码搜索范围;该至少两个视频帧中的内部编码帧。
一种可能的实现方式中,上述仿射变换矩阵可以包括至少两个子矩阵,每个子矩阵与上述至少两个视频帧中的一个视频帧所对应,每个子矩阵用于指示所对应的一个视频帧中宏块所处的几何位置。
一种可能的实现方式中,上述第一对象包括上述至少两个视频帧中每个宏块的编码搜索范围。确定模块12,具体可以用于根据上述至少两个子矩阵,确定该每个宏块对应的运动向量,每个运动向量指示一个搜索范围;并对于该每个宏块,将一个宏块对应的运动向量指示的搜索范围,确定为该一个宏块的编码搜索范围,得到该每个宏块的编码搜索范围。
一种可能的实现方式中,上述第一对象包括上述至少两个视频帧中的内部编码帧。确定模块12,具体可以用于根据上述至少两个子矩阵,确定该至少两个子矩阵中每个子矩阵对应的第一变化率;并基于该第一变化率,确定上述内部编码帧。其中,一个子矩阵对应的第一变化率包括:该一个子矩阵中的数值,相较于该一个子矩阵的前一子矩阵中数值的变化率。
一种可能的实现方式中,确定模块12,具体可以用于:对于上述第一变化率大于或等于第一阈值的子矩阵所对应的至少一个第一视频帧,将该至少一个第一视频帧确定为上述内部编码帧;或者,对于该第一变化率小于该第一阈值的子矩阵所对应的至少一个第二视频帧,计算每个第二视频帧对应的第二变化率,一个第二视频帧对应的第二变化率包括:该一个第二视频帧中像素的像素值,相较于该一个第二视频帧的前一视频帧中像素的像素值的变化率;并将该第二变化率大于或等于第二阈值的所有第二视频帧,确定为该内部编码帧。
一种可能的实现方式中,获取模块11,具体可以用于将上述图像数据和上述第一数据输入电子防抖算法;且通过该电子防抖算法,从该图像数据中获取上述至少两个视频帧中特征点的像素坐标数据;并通过该电子防抖算法,根据该像素坐标数据和该第一数据,计算得到上述仿射变换矩阵。
在本申请实施例提供的视频编码装置中,一方面该视频编码装置在确定采集的至少两 个视频帧中每个宏块的编码搜索范围时,可以直接基于用于指示相邻视频帧对应宏块的映射关系的仿射变换矩阵,而无需遍历该每个宏块中每个像素点周围的多个像素点;另一方面该视频编码装置在确定该至少两个视频帧中的内部编码帧时,也可以直接基于该仿射变换矩阵,而无需对比该至少两个视频帧中宏块的像素变化。如此在基于该编码搜索范围和该内部编码帧中的至少之一编码该至少两个视频帧时,可以极大地减少算力,降低功耗。
本申请实施例中的视频编码装置可以是电子设备,也可以是电子设备中的部件,例如集成电路或芯片。该电子设备可以是终端,也可以为除终端之外的其他设备。示例性的,电子设备可以为手机、平板电脑、笔记本电脑、掌上电脑、车载电子设备、移动上网装置(Mobile Internet Device,MID)、增强现实(augmented reality,AR)/虚拟现实(virtual reality,VR)设备、机器人、可穿戴设备、超级移动个人计算机(ultra-mobile personal computer,UMPC)、上网本或者个人数字助理(personal digital assistant,PDA)等,还可以为服务器、网络附属存储器(Network Attached Storage,NAS)、个人计算机(personal computer,PC)、电视机(television,TV)、柜员机或者自助机等,本申请实施例不作具体限定。
本申请实施例中的视频编码装置可以为具有操作系统的装置。该操作系统可以为安卓(Android)操作系统,可以为IOS操作系统,还可以为其他可能的操作系统,本申请实施例不作具体限定。
本申请实施例提供的视频编码装置能够实现上述方法实施例实现的各个过程,为避免重复,这里不再赘述。
如图11所示,本申请实施例还提供一种电子设备100,包括处理器101和存储器102,存储器102上存储有可在所述处理器101上运行的程序或指令,该程序或指令被处理器101执行时实现如上述视频编码方法实施例的各个步骤,且能达到相同的技术效果,为避免重复,这里不再赘述。
需要说明的是,本申请实施例中的电子设备包括移动电子设备和非移动电子设备。
图12为实现本申请实施例的一种电子设备的硬件结构示意图。
如图12所示,电子设备1000包括但不限于:射频单元1001、网络模块1002、音频输出单元1003、输入单元1004、传感器1005、显示单元1006、用户输入单元1007、接口单元1008、存储器1009、以及处理器1010等部件。
本领域技术人员可以理解,电子设备1000还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器1010逻辑相连,从而通过电源管理系统实现管理充电、放电、以及功耗管理等功能。图12中示出的电子设备结构并不构成对电子设备的限定,电子设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。
其中,处理器1010,可以用于基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,该仿射变换矩阵用于指示该至少两个视频帧中每两个相邻视频帧 间对应宏块的映射关系,该第一数据为:采集该至少两个视频帧的过程中电子设备的加速度数据和角速度数据。处理器1010,还可以用于基于获取的该仿射变换矩阵,确定第一对象。处理器1010,还可以用于基于确定的该第一对象,编码该至少两个视频帧。其中,该第一对象包括以下至少之一:该至少两个视频帧中每个宏块的编码搜索范围;该至少两个视频帧中的内部编码帧。
一种可能的实现方式中,上述仿射变换矩阵可以包括至少两个子矩阵,每个子矩阵与上述至少两个视频帧中的一个视频帧所对应,每个子矩阵用于指示所对应的一个视频帧中宏块所处的几何位置。
一种可能的实现方式中,上述第一对象包括上述至少两个视频帧中每个宏块的编码搜索范围。处理器1010,具体可以用于根据上述至少两个子矩阵,确定该每个宏块对应的运动向量,每个运动向量指示一个搜索范围;并对于该每个宏块,将一个宏块对应的运动向量指示的搜索范围,确定为该一个宏块的编码搜索范围,得到该每个宏块的编码搜索范围。
一种可能的实现方式中,上述第一对象包括上述至少两个视频帧中的内部编码帧。处理器1010,具体可以用于根据上述至少两个子矩阵,确定该至少两个子矩阵中每个子矩阵对应的第一变化率;并基于该第一变化率,确定上述内部编码帧。其中,一个子矩阵对应的第一变化率包括:该一个子矩阵中的数值,相较于该一个子矩阵的前一子矩阵中数值的变化率。
一种可能的实现方式中,处理器1010,具体可以用于:对于上述第一变化率大于或等于第一阈值的子矩阵所对应的至少一个第一视频帧,将该至少一个第一视频帧确定为上述内部编码帧;或者,对于该第一变化率小于该第一阈值的子矩阵所对应的至少一个第二视频帧,计算每个第二视频帧对应的第二变化率,一个第二视频帧对应的第二变化率包括:该一个第二视频帧中像素的像素值,相较于该一个第二视频帧的前一视频帧中像素的像素值的变化率;并将该第二变化率大于或等于第二阈值的所有第二视频帧,确定为该内部编码帧。
一种可能的实现方式中,处理器1010,具体可以用于将上述图像数据和上述第一数据输入电子防抖算法;且通过该电子防抖算法,从该图像数据中获取上述至少两个视频帧中特征点的像素坐标数据;并通过该电子防抖算法,根据该像素坐标数据和该第一数据,计算得到上述仿射变换矩阵。
在本申请实施例提供的电子设备中,一方面该电子设备在确定采集的至少两个视频帧中每个宏块的编码搜索范围时,可以直接基于用于指示相邻视频帧对应宏块的映射关系的仿射变换矩阵,而无需遍历该每个宏块中每个像素点周围的多个像素点;另一方面该电子设备在确定该至少两个视频帧中的内部编码帧时,也可以直接基于该仿射变换矩阵,而无需对比该至少两个视频帧中宏块的像素变化。如此在基于该编码搜索范围和该内部编码帧 中的至少之一编码该至少两个视频帧时,可以极大地减少电子设备的算力,降低电子设备的功耗。
应理解的是,本申请实施例中,输入单元1004可以包括图形处理器(Graphics Processing Unit,GPU)10041和麦克风10042,GPU10041对在视频捕获模式或图像捕获模式中由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元1006可包括显示面板10061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板10061。用户输入单元1007包括触控面板10071以及其他输入设备10072中的至少一种。触控面板10071,也称为触摸屏。触控面板10071可包括触摸检测装置和触摸控制器两个部分。其他输入设备10072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。
存储器1009可用于存储软件程序以及各种数据。存储器1009可主要包括存储程序或指令的第一存储区和存储数据的第二存储区,其中,第一存储区可存储操作系统、至少一个功能所需的应用程序或指令(比如声音播放功能、图像播放功能等)等。此外,存储器1009可以包括易失性存储器或非易失性存储器,或者,存储器1009可以包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请实施例中的存储器1009包括但不限于这些和任意其它适合类型的存储器。
处理器1010可包括一个或多个处理单元;可选的,处理器1010集成应用处理器和调制解调处理器,其中,应用处理器主要处理涉及操作系统、用户界面和应用程序等的操作,调制解调处理器主要处理无线通信信号,如基带处理器。可以理解的是,上述调制解调处理器也可以不集成到处理器1010中。
本申请实施例还提供一种可读存储介质,所述可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述视频编码方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
其中,所述处理器为上述实施例中所述的电子设备中的处理器。所述可读存储介质,包括计算机可读存储介质,如计算机只读存储器ROM、随机存取存储器RAM、磁碟或者光盘等。
本申请实施例另提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现上述视频编码方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
应理解,本申请实施例提到的芯片还可以称为系统级芯片、系统芯片、芯片系统或片上系统芯片等。
本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如上述视频编码方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去、或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以计算机软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例所述的方法。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式,均属于本申请的保护之内。

Claims (15)

  1. 一种视频编码方法,包括:
    基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,所述仿射变换矩阵用于指示所述至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,所述第一数据为:采集所述至少两个视频帧的过程中电子设备的加速度数据和角速度数据;
    基于所述仿射变换矩阵,确定第一对象;
    基于所述第一对象,编码所述至少两个视频帧;
    其中,所述第一对象包括以下至少之一:
    所述至少两个视频帧中每个宏块的编码搜索范围;
    所述至少两个视频帧中的内部编码帧。
  2. 根据权利要求1所述的方法,其中,所述仿射变换矩阵包括至少两个子矩阵,每个所述子矩阵与所述至少两个视频帧中的一个视频帧所对应,每个所述子矩阵用于指示所对应的一个视频帧中宏块所处的几何位置。
  3. 根据权利要求2所述的方法,其中,所述第一对象包括所述每个宏块的编码搜索范围;
    所述基于所述仿射变换矩阵,确定第一对象,包括:
    根据所述至少两个子矩阵,确定所述每个宏块对应的运动向量,每个运动向量指示一个搜索范围;
    对于所述每个宏块,将一个宏块对应的运动向量指示的搜索范围,确定为所述一个宏块的编码搜索范围,得到所述每个宏块的编码搜索范围。
  4. 根据权利要求2所述的方法,其中,所述第一对象包括所述内部编码帧;
    所述基于所述仿射变换矩阵,确定第一对象,包括:
    根据所述至少两个子矩阵,确定每个所述子矩阵对应的第一变化率;
    基于所述第一变化率,确定所述内部编码帧;
    其中,一个子矩阵对应的第一变化率包括:所述一个子矩阵中的数值,相较于所述一个子矩阵的前一子矩阵中数值的变化率。
  5. 根据权利要求4所述的方法,其中,所述基于所述第一变化率,确定所述内部编码帧,包括:
    对于所述第一变化率大于或等于第一阈值的子矩阵所对应的至少一个第一视频帧,将所述至少一个第一视频帧确定为所述内部编码帧;
    或者,
    对于所述第一变化率小于第一阈值的子矩阵所对应的至少一个第二视频帧,计算每个所述第二视频帧对应的第二变化率,一个第二视频帧对应的第二变化率包括:所述一个第 二视频帧中像素的像素值,相较于所述一个第二视频帧的前一视频帧中像素的像素值的变化率;
    将所述第二变化率大于或等于第二阈值的所有所述第二视频帧,确定为所述内部编码帧。
  6. 根据权利要求1或2所述的方法,其中,所述基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,包括:
    将所述图像数据和所述第一数据输入电子防抖算法;
    通过所述电子防抖算法,从所述图像数据中获取所述至少两个视频帧中特征点的像素坐标数据;
    通过所述电子防抖算法,根据所述像素坐标数据和所述第一数据,计算得到所述仿射变换矩阵。
  7. 一种视频编码装置,包括获取模块、确定模块和编码模块;
    所述获取模块,用于基于采集的至少两个视频帧的图像数据,以及第一数据,获取仿射变换矩阵,所述仿射变换矩阵用于指示所述至少两个视频帧中每两个相邻视频帧间对应宏块的映射关系,所述第一数据为:采集所述至少两个视频帧的过程中电子设备的加速度数据和角速度数据;
    所述确定模块,用于基于所述获取模块获取的所述仿射变换矩阵,确定第一对象;
    所述编码模块,用于基于所述确定模块确定的所述第一对象,编码所述至少两个视频帧;
    其中,所述第一对象包括以下至少之一:
    所述至少两个视频帧中每个宏块的编码搜索范围;
    所述至少两个视频帧中的内部编码帧。
  8. 根据权利要求7所述的装置,其中,所述仿射变换矩阵包括至少两个子矩阵,每个所述子矩阵与所述至少两个视频帧中的一个视频帧所对应,每个所述子矩阵用于指示所对应的一个视频帧中宏块所处的几何位置。
  9. 根据权利要求8所述的装置,其中,所述第一对象包括所述每个宏块的编码搜索范围;
    所述确定模块,具体用于根据所述至少两个子矩阵,确定所述每个宏块对应的运动向量,每个运动向量指示一个搜索范围;并对于所述每个宏块,将一个宏块对应的运动向量指示的搜索范围,确定为所述一个宏块的编码搜索范围,得到所述每个宏块的编码搜索范围。
  10. 根据权利要求8所述的装置,其中,所述第一对象包括所述内部编码帧;
    所述确定模块,具体用于根据所述至少两个子矩阵,确定每个所述子矩阵对应的第一变化率;并基于所述第一变化率,确定所述内部编码帧;
    其中,一个子矩阵对应的第一变化率包括:所述一个子矩阵中的数值,相较于所述一个子矩阵的前一子矩阵中数值的变化率。
  11. 根据权利要求7或8所述的装置,其中,
    所述获取模块,具体用于将所述图像数据和所述第一数据输入电子防抖算法;且通过所述电子防抖算法,从所述图像数据中获取所述至少两个视频帧中特征点的像素坐标数据;并通过所述电子防抖算法,根据所述像素坐标数据和所述第一数据,计算得到所述仿射变换矩阵。
  12. 一种电子设备,包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求1-6中任一项所述的视频编码方法的步骤。
  13. 一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如权利要求1-6中任一项所述的视频编码方法的步骤。
  14. 一种芯片,包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如权利要求1-6中任一项所述的视频编码方法。
  15. 一种计算机程序产品,所述程序产品被存储在存储介质中,所述程序产品被至少一个处理器执行以实现如权利要求1-6中任一项所述的视频编码方法。
PCT/CN2024/100314 2023-06-26 2024-06-20 视频编码方法、装置、电子设备及可读存储介质 Ceased WO2025001955A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US19/411,503 US20260095580A1 (en) 2023-06-26 2025-12-08 Video Encoding Method, Electronic Device and Non-Transitory Readable Storage Medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310763796.5A CN116723323A (zh) 2023-06-26 2023-06-26 视频编码方法、装置、电子设备及可读存储介质
CN202310763796.5 2023-06-26

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/411,503 Continuation US20260095580A1 (en) 2023-06-26 2025-12-08 Video Encoding Method, Electronic Device and Non-Transitory Readable Storage Medium

Publications (1)

Publication Number Publication Date
WO2025001955A1 true WO2025001955A1 (zh) 2025-01-02

Family

ID=87865939

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/100314 Ceased WO2025001955A1 (zh) 2023-06-26 2024-06-20 视频编码方法、装置、电子设备及可读存储介质

Country Status (3)

Country Link
US (1) US20260095580A1 (zh)
CN (1) CN116723323A (zh)
WO (1) WO2025001955A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116723323A (zh) * 2023-06-26 2023-09-08 维沃移动通信有限公司 视频编码方法、装置、电子设备及可读存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20060274159A1 (en) * 2005-06-01 2006-12-07 Canon Kabushiki Kaisha Image coding apparatus and image coding method
CN113228631A (zh) * 2019-01-12 2021-08-06 腾讯美国有限责任公司 视频编解码的方法和装置
CN113411585A (zh) * 2021-06-15 2021-09-17 广东工业大学 一种适用于高速飞行器的h.264运动视频编码方法及系统
CN115705651A (zh) * 2021-08-06 2023-02-17 武汉Tcl集团工业研究院有限公司 视频运动估计方法、装置、设备和计算机可读存储介质
CN116723323A (zh) * 2023-06-26 2023-09-08 维沃移动通信有限公司 视频编码方法、装置、电子设备及可读存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2001275103A (ja) * 2000-03-23 2001-10-05 Matsushita Electric Ind Co Ltd 監視システム及びその動き検出方法
CN101536036A (zh) * 2005-09-26 2009-09-16 皇家飞利浦电子股份有限公司 用于跟踪物体或人的运动的方法和设备
JP2012204914A (ja) * 2011-03-24 2012-10-22 Panasonic Corp 動きベクトル検出方法と装置、及びビデオカメラ
CN105141807B (zh) * 2015-09-23 2018-11-30 北京二郎神科技有限公司 视频信号图像处理方法和装置
CN107071421B (zh) * 2017-05-23 2019-11-22 北京理工大学 一种结合视频稳定的视频编码方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20060274159A1 (en) * 2005-06-01 2006-12-07 Canon Kabushiki Kaisha Image coding apparatus and image coding method
CN113228631A (zh) * 2019-01-12 2021-08-06 腾讯美国有限责任公司 视频编解码的方法和装置
CN113411585A (zh) * 2021-06-15 2021-09-17 广东工业大学 一种适用于高速飞行器的h.264运动视频编码方法及系统
CN115705651A (zh) * 2021-08-06 2023-02-17 武汉Tcl集团工业研究院有限公司 视频运动估计方法、装置、设备和计算机可读存储介质
CN116723323A (zh) * 2023-06-26 2023-09-08 维沃移动通信有限公司 视频编码方法、装置、电子设备及可读存储介质

Also Published As

Publication number Publication date
US20260095580A1 (en) 2026-04-02
CN116723323A (zh) 2023-09-08

Similar Documents

Publication Publication Date Title
US20190273929A1 (en) De-Blocking Filtering Method and Terminal
JP5031976B2 (ja) ディジタルビデオデータの処理
KR20190015093A (ko) 향상된 비디오 코딩을 위한 참조 프레임 재투영
US8019000B2 (en) Motion vector detecting device
US20090016623A1 (en) Image processing device, image processing method and program
US9083982B2 (en) Image combining and encoding method, image combining and encoding device, and imaging system
CN112449182B (zh) 视频编码方法、装置、设备及存储介质
US20260095580A1 (en) Video Encoding Method, Electronic Device and Non-Transitory Readable Storage Medium
US9319682B2 (en) Moving image encoding apparatus, control method therefor, and non-transitory computer readable storage medium
JP2014239428A (ja) デジタルビデオデータを符号化するための方法
CN101389032A (zh) 一种基于图像插值的帧内预测编码方法
JP4709155B2 (ja) 動き検出装置
JP4346573B2 (ja) 符号化装置と方法
CN117596392B (zh) 编码块的编码信息确定方法及相关产品
JPH11308617A (ja) ディジタル画像符号化装置とこれに用いる動きベクトル検出装置
US20240373076A1 (en) Methods and apparatus for generating a live stream for a connected device
US20080298769A1 (en) Method and system for generating compressed video to improve reverse playback
CN115776570A (zh) 视频流编解码方法、装置、处理系统及电子设备
CN100496126C (zh) 影像编码装置及其方法
JP2001045493A (ja) 動画像符号化装置、動画像出力装置、及び記憶媒体
CN115086665A (zh) 误码掩盖方法、装置、系统、存储介质和计算机设备
WO2007055013A1 (ja) 画像復号化装置および方法、画像符号化装置
JP2009081622A (ja) 動画像圧縮符号化装置
JP2001275103A (ja) 監視システム及びその動き検出方法
WO2013165624A1 (en) Mechanism for facilitating cost-efficient and low-latency encoding of video streams

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24830601

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE