WO2022147703A1 - 跟焦方法、装置、拍摄设备及计算机可读存储介质 - Google Patents
跟焦方法、装置、拍摄设备及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2022147703A1 WO2022147703A1 PCT/CN2021/070581 CN2021070581W WO2022147703A1 WO 2022147703 A1 WO2022147703 A1 WO 2022147703A1 CN 2021070581 W CN2021070581 W CN 2021070581W WO 2022147703 A1 WO2022147703 A1 WO 2022147703A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video frame
- region
- image
- interest
- sensor
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G03—PHOTOGRAPHY; CINEMATOGRAPHY; ANALOGOUS TECHNIQUES USING WAVES OTHER THAN OPTICAL WAVES; ELECTROGRAPHY; HOLOGRAPHY
- G03B—APPARATUS OR ARRANGEMENTS FOR TAKING PHOTOGRAPHS OR FOR PROJECTING OR VIEWING THEM; APPARATUS OR ARRANGEMENTS EMPLOYING ANALOGOUS TECHNIQUES USING WAVES OTHER THAN OPTICAL WAVES; ACCESSORIES THEREFOR
- G03B13/00—Viewfinders; Focusing aids for cameras; Means for focusing for cameras; Autofocus systems for cameras
- G03B13/32—Means for focusing
- G03B13/34—Power focusing
- G03B13/36—Autofocus systems
Definitions
- the present application relates to the field of photographing technologies, and in particular, to a follow focus method, apparatus, photographing device, and computer-readable storage medium.
- the purpose of focusing processing is to find an accurate focal position, so that the photographed target can form a clear image on the plane of the photosensitive device, so as to ensure the clearness of the main image on the photographed screen. Therefore, finding the focal position accurately is one of the key factors in determining the quality of shooting. Especially in the video shooting scene, the user needs to shoot continuously, and how to continuously find the focus position in real time is a technical problem that needs to be solved urgently.
- the present application provides a follow focus method, an apparatus, a photographing device, and a computer-readable storage medium, so as to solve the problem that automatic follow focus cannot be performed during shooting in the related art.
- a first aspect provides a follow focus method, the method comprising:
- the focus position is determined based on the depth information of the region of interest, and the focus position is used to control the shooting device to follow focus when shooting a video.
- a follow focus device in a second aspect, includes a processor, a memory, and a computer program stored on the memory and executable by the processor, the processor implements the following steps when executing the computer program :
- the focus position is determined based on the depth information of the region of interest, and the focus position is used to control the shooting device to follow focus when shooting a video.
- a photographing device including an image sensor and a depth sensor; and the follow focus device according to the second aspect.
- a computer-readable storage medium where several computer instructions are stored thereon, and when the computer instructions are executed, the follow focus method described in the first aspect is implemented.
- using an image sensor to acquire a first video frame can automatically acquire a region of interest in the first video frame; using a depth sensor to acquire a second video frame carrying depth information, because the first video frame and the second video frame contains the same picture content, so the depth information of the second video frame can be used to determine the depth information of the region of interest.
- the depth information of the region of interest represents the depth information of the photographed target. What is obtained is the depth information of the entire area instead of the single point depth information, so the focus position can be accurately found, and the shooting device can automatically follow the focus when shooting videos. Additional operation, and can accurately find the focus position; because the focus position can be found accurately, the real-time follow focus can be guaranteed during continuous video shooting, and it will not cause obvious focus overshoot or virtual focus.
- FIG. 1A is a schematic diagram of a follow focus method according to an embodiment of the present application.
- FIG. 1B is a schematic diagram of a follow focus method according to another embodiment of the present application.
- FIG. 1C is a schematic diagram of an image sensor and a depth sensor according to an embodiment of the present application.
- FIG. 1D is a schematic diagram of a weight of a face region according to an embodiment of the present application.
- FIG. 2 is a schematic diagram of a follow focus device according to an embodiment of the present application.
- FIG. 3 is a schematic diagram of a photographing device according to an embodiment of the present application.
- the ISP (Image Signal Processing) system of the shooting device is usually equipped with a focusing module, and some focusing processing schemes use phase detection autofocus, but this scheme has higher requirements on ambient light and is less effective in the night environment; Other focusing processing solutions use the laser method to obtain single-point depth information. Based on the single-point depth information, it may not be possible to obtain the depth information of the subject that the user wants, and accurate focusing cannot be achieved. Therefore, users often need to focus manually during actual shooting. This method is acceptable in a single image shooting scene, because in this shooting scene, the user does not have the need for continuous shooting. After the user operates the focus and then operates the shooting, the entire shooting process is completed, and manual focus will not. cause more disturbance to users.
- the continuous video shooting scene is actually quite different from the single image shooting scene; because the video shooting needs to ensure the continuity, allowing the user to interrupt the continuous shooting to select the focus position will bring the user a very poor shooting. experience.
- an additional follow focus engineer is required to manually adjust the focus, which brings a large cost for users to shoot high-quality videos.
- the image sensor is used to obtain the first video frame, and the region of interest in the first video frame can be automatically obtained; the depth sensor is used to obtain the second video frame carrying depth information.
- the depth information of the second video frame can be used to determine the depth information of the region of interest.
- the depth information of the region of interest characterizes the subject. Since the depth information obtained is the depth information of the entire area rather than the depth information of a single point, the focal position can be accurately found, and the shooting device can automatically follow the focus when shooting video.
- This solution can simultaneously ensure: video shooting Continuity, no additional user operation is required, and the focus position can be accurately found; because the focus position can be found accurately, the real-time follow focus can be guaranteed during continuous video shooting, without causing obvious focus overshoot or blur. focus phenomenon.
- the solution of this embodiment can be applied to a shooting device, and the follow-focus solution of this embodiment is executed by a built-in processor of the shooting device to realize automatic follow-focus during video shooting.
- the photographing device may also be mounted in an external device, where the external device may include devices such as a movable platform, and the movable platform may include a vehicle, an unmanned aerial vehicle, a mobile robot, or the like.
- FIG. 1A and FIG. 1B are respectively schematic diagrams of the follow focus method shown in this embodiment, which may include the following steps:
- step 102 after obtaining the first video frame by using the image sensor, obtain the region of interest in the first video frame;
- a depth sensor is used to obtain a second video frame carrying depth information; wherein, the first video frame and the second video frame contain the same picture content;
- step 106 use the second video frame to determine the depth information of the region of interest
- a focus position is determined based on the depth information of the region of interest, and the focus position is used to control the shooting device to follow focus when shooting a video.
- the follow focus solution in this embodiment is suitable for a video shooting scene.
- the shooting device is configured with an image sensor, and the image sensor is used to continuously capture images.
- This embodiment obtains the first video frame through the image sensor, and the first video frame may be an image sensor.
- the collected original image may also be an image obtained by performing some processing on the original image collected by the image sensor.
- the shooting device is further configured with a depth sensor to continuously collect images carrying depth information in the video shooting scene.
- the depth sensor is used to obtain a second video frame, and the second video frame may be a depth
- the original image collected by the sensor may also be an image obtained by performing some processing on the original image collected by the depth sensor.
- the first video frame and the second video frame contain the same picture content
- the region of interest is automatically identified by the first video frame
- the depth of the region of interest is determined by using the depth information carried by the second video frame information.
- the first video frame and the second video frame may contain the same picture content in various ways.
- the depth sensor and the image sensor can be made parallel to each other, so that the images captured by the depth sensor and the image sensor can be aligned, as shown in FIG. A schematic diagram of the depth sensor and the image sensor.
- the depth sensor is a TOF (Time of flight) sensor as an example.
- the depth sensor TOF is set up by a supporting device TOF holder and the image sensor Image sensor.
- being in a parallel state means that the angle X deg between the depth sensor and the horizontal plane is equal to the angle between the image sensor Image sensor and the horizontal plane, so the video frame collected by the depth sensor and the video collected by the image sensor can be made
- the frames correspond to the same picture, that is, the first video frame and the second video frame contain the same picture content.
- the depth sensor and the image sensor may not have the limitation of the above-mentioned parallel state; in other examples, the depth sensor and the image sensor may be in a non-parallel state, that is, the depth sensor and the image sensor can be configured to be in other positional relationships as required.
- the embodiment solution may further perform image alignment processing on the image collected by the image sensor and the image collected by the depth sensor to obtain the first video frame and the second video frame.
- the installation positions of the depth sensor and the image sensor may not be in a parallel state, and the image sensor and the depth sensor have an angular deviation, so the image collected by the image sensor has an angular deviation from the image collected by the depth sensor.
- this embodiment can also Performing image alignment processing, as an example, may include: performing contour alignment on the image acquired by the image sensor and the image acquired by the depth sensor, for example, according to the angle difference between the depth sensor and the image sensor, aligning the image sensor The acquired image is aligned with the contour of the image acquired by the depth sensor.
- the corresponding viewing angle ranges of the depth sensor and the image sensor may not be exactly the same, that is, the images contained in the images collected by the two are not exactly the same.
- the images included in the images captured by the depth sensor may contain images that are not included in the images captured by the image sensor, and images included in the images captured by the image sensor may also include images that are not included in the images captured by the depth sensor.
- the different pictures in the images collected by the two are called offset areas.
- the offset areas of both the image collected by the image sensor and the image collected by the depth sensor are cropped, Both the cropped first video frame and the second video frame contain the same pictures, so that depth information of the region of interest can be easily obtained.
- the depth sensor may include a 3D ToF (TOF, Time-of-Flight) sensor, and the 3D ToF sensor captures images without being affected by poor lighting such as at night, so that the follow focus solution of this embodiment is It can be applied to the night environment; on the other hand, the 3D ToF sensor can collect depth information in a large range, so that the follow focus solution of this embodiment can accurately find the focus position.
- TOF Time-of-Flight
- the subject is automatically determined by automatically identifying the region of interest from the first video frame, and the depth information of the region of interest is determined by using the second video frame.
- the processing process may include: using the The position of the region of interest in the first video frame, and the depth information of the region of interest is determined from the second video frame. For example, if the first video frame and the second video frame contain the same picture, the identified region of interest can be determined, and the position of the region of interest in the first video frame can be determined, then correspondingly, it can be determined that the region of interest is in the second video frame The position in the frame, the depth information of the region of interest can be determined according to the depth information carried by the second video frame.
- the focus position can be further determined.
- the region of interest represents the region where the shooting subject in the first video frame is located, and the focal position can be determined from the depth information of the region of interest.
- the specific determination method can be flexibly configured as required in practical applications.
- the region of interest includes multiple pixels, and the average depth information of the region of interest can be determined based on the depth information of each target pixel in the region of interest, and then the average depth information of the region of interest can be used to determine the average depth information of the region of interest. Determine the focus position.
- the average depth information of the region of interest is determined based on the depth information of each target pixel in the region of interest and the preset weight of each target pixel.
- the region of interest can be divided into different sub-regions in advance according to the corresponding shooting subjects in the region of interest, and weights can be set for different sub-regions.
- the pixel points are multiplied by the weight of the target pixel points and averaged, and then the average depth information of the region of interest is obtained, and the focus position can be further determined.
- the target pixels may be all pixels in the region of interest, or may be part of the pixels, for example, pixels obtained by denoising the depth information of the region of interest.
- the depth information of some pixels in the region of interest may be unreliable, and some unreliable pixels can be eliminated.
- the sampling point may not belong to the subject, or the data collected by the depth sensor is wrong. Based on this, in order to prevent the interference of these possible wrong pixels, these pixels can be eliminated.
- a confidence threshold can be pre-configured, that is, the target pixel is a pixel in the region of interest whose depth information is greater than the set confidence threshold, Pixels whose depth information meets the confidence threshold are retained, and those that do not meet the confidence threshold are eliminated.
- the confidence threshold is set based on the echo strength of the depth sensor.
- different depth sensors may have different echo sensing capabilities. Under normal circumstances, the depth information collected by the depth sensor will also be within a reasonable range. This embodiment can be determined based on the echo strength of the depth sensor in advance. Confidence threshold, so that the accuracy of the follow focus processing scheme is more accurate.
- regions of interest in different shooting scenarios can be determined as needed.
- the region of interest may be the face region identified from the first video frame; in other examples, when the face picture is not included, the region of interest may also be obtained from all face pictures.
- multiple regions of interest may also be automatically identified from the first video frame. For the identified multiple regions of interest, one with the highest priority may also be automatically reserved according to the set priority. For other regions of interest to be eliminated, for example, the human face region has the highest priority, the animal face region takes the second place, and other items of different categories come second, which can be configured as needed in practical applications, which is not limited in this embodiment. .
- multiple faces may also be recognized, and the region where one of the faces is located can be automatically selected as the region of interest as needed, for example, it can be determined based on the size of each face region, or based on the depth information of each face region For example, the area where the face is located is larger as the area of interest, or the depth information of the face is smaller as the area of interest, etc.
- the region of interest in the first video frame may also be determined according to a face region designated by the user that needs to be in focus.
- the user can specify the face to be in focus before or during the shooting, and according to the face specified by the user, in the subsequent continuous shooting process, the person to be focused by the user can be identified in the first video frame
- the region where the face is located is the region of interest.
- the acquiring the region of interest in the first video frame may include: highlighting the image acquired from the image sensor in the preview screen of the photographing device After at least two face regions identified in the user-specified face region that needs to follow focus; use the face region that needs to follow focus to obtain the face region in the first video frame; based on this, By highlighting, users can quickly and conveniently specify the face they want to focus on. After the user specifies the face that needs to be focused, when the video is captured next, the video frame when the user specifies the face is different from the next one.
- the region of interest can be determined from the first video frame by using the face feature of the face that needs to follow focus specified by the user.
- this embodiment can automatically identify the region of interest from the captured video frame through the face feature of the face that needs to be in focus specified by the user. Based on this, when shooting a video, the user only needs to perform a specified operation once, and then the whole process of automatic continuous focus can be achieved.
- weights can be set for different regions in the face region as needed.
- FIG. A schematic diagram of the weight of the face area, in the face area, the weight of the pixel point in the eye area, the weight of the pixel point in the mouth area, the weight of the pixel point in the nose area, and the weight of the pixel point in the edge area decrease sequentially.
- other implementation manners of the weight may be set as required, for example, it is divided into multiple other areas as required, and different weights are configured for different areas, which is not limited in this embodiment.
- the follow focus solution in this embodiment can be continuously executed to find the focus position.
- the follow focus solution of this embodiment may be continuously performed according to a set time interval; in other examples, this implementation may be performed based on the first video frame of each frame based on an image captured by an image sensor as a reference In other examples, the follow focus solution in this embodiment may also be performed based on the second video frame of each frame based on the image collected by the depth sensor.
- the frame rate (FPS, Frames Per Second) of the depth sensor is the same as the frame rate of the image sensor.
- the depth sensor and the image sensor are activated at the same time.
- the images acquired by the image sensor can be automatically aligned in time, that is, the time stamps of the images acquired by the image sensor and the depth sensor are the same.
- the frame rate of the depth sensor is different from the frame rate of the image sensor
- the second video frame may be the video frame closest to the acquisition time of the first video frame, so that the first video frame The time difference between the video frame and the second video frame is the smallest, so as to ensure that the depth information obtained by using the second video frame matches the shooting subject in the first video frame, and to ensure the accuracy of the follow-focus scheme.
- the follow focus solution in this embodiment may be based on the image acquired by the image sensor as the benchmark, and the first video frame with the closest acquisition time based on the first video frame. Two video frames; it is also possible to obtain the first video frame with the closest time based on the second video frame based on the image captured by the depth sensor.
- a buffer area may be set to buffer the image, so as to obtain the first video frame and the second video frame with the smallest time difference, thereby improving the Follow focus accuracy.
- the image buffered in the buffer area includes a second video frame, that is, at least one frame of image with a time stamp collected by the depth sensor is buffered in the buffer area;
- the second video frame with the closest time may be obtained from the buffer according to the time stamp of the image.
- the image buffered in the buffer area includes the first video frame, that is, at least one frame of image with a time stamp collected by the image sensor is buffered in the buffer area;
- the first video frame with the closest time may be obtained from the buffer area according to the timestamp of the image.
- this embodiment is also based on the buffering of the image buffered in the buffer area.
- the duration or the number of frames of the image cached in the buffer provides an innovative solution.
- the number of frames of the image buffered in the buffer area is determined based on the difference between the frame rate of the depth sensor and the frame rate of the image sensor.
- the buffering duration of the image buffered in the buffer area is greater than or equal to the difference between the inverse of the frame rate of the depth sensor and the inverse of the frame rate of the image sensor.
- the buffer area is used to store the image collected by the depth sensor
- the frame rate of the depth sensor is greater than the frame rate of the image sensor
- at least one frame of image collected by the depth sensor needs to be stored in the buffer area, so that when the second video frame is acquired based on the first video frame, it can be The second video frame with the closest frame time. Because the depth sensor has been able to collect at least one frame of image within the time interval when the image sensor collects one frame of image, it is necessary to buffer at least one frame of image collected by the depth sensor.
- the number of frames of the image buffered in the buffer area is greater than or equal to the inverse of the frame rate of the depth sensor and the image sensor.
- the frame rate of the depth sensor is 40FPS, and the frame rate of the image sensor is 30FPS;
- the number of frames of images collected by the sensor must be at least greater than the quotient of 1/30 and 1/40, that is, at least greater than 2 frames, that is, one frame of the image collected by the image sensor will be generated within the time when the depth sensor collects 2 frames of images. Therefore, At least 2 frames should be buffered, so as to ensure that when acquiring the second video frame based on the first video frame, the second video frame whose time is closest to the first video frame can be accurately found.
- the frame rate of the depth sensor is smaller than the frame rate of the image sensor
- at least one frame of the image collected by the depth sensor needs to be buffered in the buffer area for a certain period of time, so that when the second video frame is obtained based on the first video frame, it is possible to find the The second video frame with the closest time of the first video frame. Because the depth sensor has not been able to collect an image within the time interval when the image sensor collects a frame of image, the image collected by the depth sensor needs to be buffered for a certain period of time.
- the buffering duration of the image buffered in the buffer area is greater than or equal to the reciprocal frame rate of the depth sensor and the image sensor. The difference between the inverse of the frame rate.
- the frame rate of the depth sensor is 30FPS
- the frame rate of the image sensor is 40FPS
- the duration of the collected images must be at least greater than the difference between 1/30 and 1/40, that is, at least greater than 1/120 seconds, to ensure that when the second video frame is acquired based on the first video frame, it is possible to find the same value as the first video frame.
- the second video frame with the closest video frame time is 30FPS
- the frame rate of the image sensor is 40FPS
- the duration of the collected images must be at least greater than the difference between 1/30 and 1/40, that is, at least greater than 1/120 seconds, to ensure that when the second video frame is acquired based on the first video frame, it is possible to find the same value as the first video frame.
- the second video frame with the closest video frame time is the closest video frame time.
- the buffer area is used to store the image collected by the image sensor, which is the same principle as the previous embodiment:
- the buffer area needs to store multiple frames of images collected by the image sensor, so that when the first video frame is acquired based on the second video frame, it can be found and the second video frame.
- the first video frame with the closest time. Because the image sensor has not been able to capture an image within the time interval when the depth sensor captures a frame of image, the image captured by the image sensor needs to be buffered for a certain period of time.
- the buffering duration of the image buffered in the buffer area is greater than or equal to the inverse of the frame rate of the depth sensor and all The difference between the inverse of the frame rate of the image sensor.
- the frame rate of the depth sensor is 40FPS
- the frame rate of the image sensor is 30FPS
- the cache duration of the collected image must be at least greater than the difference between 1/30 and 1/40, that is, at least greater than 1/120 second, to ensure that when the first video frame is acquired based on the second video frame, it is possible to find the same value as the first video frame.
- the first video frame whose time is closest to the time when the two video frames are collected.
- the buffer area needs to store multiple frames of images collected by the image sensor, so that when the first video frame is acquired based on the second video frame, it is possible to find the same frame as the second video frame.
- the first video frame with the closest time. Because the image sensor has been able to collect at least one frame of image within the time interval when the depth sensor collects one frame of image, it is necessary to buffer at least one frame of image collected by the image sensor.
- the number of frames of the image buffered in the buffer area is greater than or equal to the inverse of the frame rate of the depth sensor and the image sensor.
- the quotient of the inverse of the frame rate is greater than or equal to the inverse of the frame rate of the depth sensor and the image sensor.
- the frame rate of the depth sensor is 30FPS, and that of the image sensor is 40FPS;
- the number of frames of the image collected by the sensor must be at least greater than the quotient of 1/30 and 1/40, that is, at least greater than 2 frames, that is, the image sensor will generate one frame of the image collected by the depth sensor within the time of collecting two frames of images. Therefore, at least 2 frames must be cached to ensure that when the first video frame is acquired based on the second video frame, the first video frame whose time is closest to the second video frame can be found.
- the images cached in the buffer area may also include other images. It is configured according to factors, for example, further buffering images with more frames, or further extending the buffering time of the images, which is not limited in this embodiment.
- the resolution of the image sensor is the same as the resolution of the depth sensor, and the depth information of the second video frame can be conveniently used to determine the depth information of the region of interest.
- the resolution of the image sensor is different from the resolution of the depth sensor
- the lower resolution image can be interpolated to make the resolution of the two the same, and then the depth information of the second video frame is used to determine Depth information of the region of interest.
- the first resolution of the image collected by the image sensor is higher than the second resolution of the image collected by the depth sensor
- the second video frame is to interpolate the image collected by the depth sensor.
- the video frame obtained by processing is the same as the first resolution.
- the implementation process of the interpolation processing may be flexibly configured as required, which is not limited in this embodiment.
- the solution of this embodiment may also simultaneously perform image alignment processing on the image collected by the depth sensor and the image collected by the image sensor, and may also perform interpolation processing on the image collected by the depth sensor at the same time.
- image alignment processing can be found in the foregoing description, and are not repeated here.
- the calculated focus position can be based on This controls the focus position of the shooting device when shooting the next video frame, so as to control the shooting device to follow focus when shooting video.
- the calculated focus position can be directly used as the focus position of the shooting device when shooting the next video frame.
- the follow focus scheme of this embodiment is fused with the results of other follow focus modes. Fusion can be handled in a variety of ways:
- an optional way is to use the focus position to control the shooting device to follow focus when shooting a video, and based on the overall confidence of the depth information of the region of interest, compare the focus position with those based on other The focus position determined by the follow focus mode is fused, and the result of fusion is used to control the shooting device to follow focus when shooting video.
- the overall confidence level of the depth information of the region of interest may be determined based on the average value of the confidence levels of the depth information of the target pixel points in the region of interest.
- the depth information of the region of interest may be compared with the depth information of the region of interest based on the overall confidence of the depth information of the region of interest.
- the depth information determined by other follow focus modes is fused, and the focus position determined by the fusion result is used to control the shooting device to follow focus when shooting video.
- the process can also be terminated, and the focus position can be obtained in other ways.
- the foregoing method embodiments may be implemented by software, and may also be implemented by hardware or a combination of software and hardware.
- Taking software implementation as an example as a device in a logical sense, it is formed by reading the corresponding computer program instructions in the memory into the memory through the processor where it is located.
- FIG. 2 which is a hardware structure diagram of a follow focus device 200 for implementing the image processing method of this embodiment, except for the processor 201 and the memory 202 shown in FIG. 2 , the embodiment
- the device for implementing the follow focus method in FIG. 2 may also include other hardware according to the actual function of the follow focus device, which will not be repeated here.
- the processor 201 implements the following steps when executing the computer program:
- the focus position is determined based on the depth information of the region of interest, and the focus position is used to control the shooting device to follow focus when shooting a video.
- the depth sensor is parallel to the image sensor.
- the depth sensor includes: a 3D ToF sensor.
- the determining the depth information of the region of interest using the second video frame includes:
- the depth information of the region of interest is determined from the second video frame.
- the focus position is determined using average depth information of the region of interest, and the average depth information of the region of interest is determined based on depth information of each target pixel of the region of interest.
- the average depth information of the region of interest is determined based on depth information of each target pixel in the region of interest and a preset weight of each target pixel.
- the target pixel is a pixel in the region of interest whose depth information is greater than a set confidence threshold.
- the confidence threshold is set based on the echo strength of the depth sensor.
- the region of interest includes a face region identified from the first video frame.
- the weight of the pixel points in the eye area, the weight of the pixel points in the mouth area, the weight of the pixel points in the nose area, and the weight of the pixel points in the edge area decrease sequentially.
- the region of interest is determined based on a user-specified face region that needs to be in focus.
- the obtaining a region of interest in the first video frame :
- the face area in the first video frame is acquired by using the face area to be focused.
- the frame rate of the depth sensor is different from the frame rate of the image sensor
- the second video frame is the video frame closest to the acquisition time of the first video frame.
- the second video frame is obtained from a buffer area, where at least one frame of image with a time stamp collected by the depth sensor is buffered in the buffer area.
- the first video frame is acquired from a buffer area, where at least one frame of image with a time stamp collected by the image sensor is buffered in the buffer area.
- the number of frames of the image buffered in the buffer area is determined based on the difference between the frame rate of the depth sensor and the frame rate of the image sensor.
- the buffering duration of the image buffered in the buffer area is determined based on the difference between the frame rate of the depth sensor and the frame rate of the image sensor.
- the number of frames of the image buffered in the buffer area is greater than or equal to the quotient of the inverse of the frame rate of the depth sensor and the inverse of the frame rate of the image sensor.
- the buffering duration of the image buffered in the buffer area is greater than or equal to the difference between the inverse of the frame rate of the depth sensor and the inverse of the frame rate of the image sensor.
- a first resolution of an image captured by the image sensor is higher than a second resolution of an image captured by the depth sensor
- the second video frame is a video frame with the same resolution as the first video frame obtained by performing interpolation processing on the image collected by the depth sensor.
- the first video frame and the second video frame are obtained after image alignment processing.
- the image alignment process includes any one of the following: performing contour alignment on the image acquired by the image sensor and the image acquired by the depth sensor, or, aligning the image acquired by the image sensor with the depth The offset areas of the two images collected by the sensor are cropped.
- the offset area is determined based on a positional difference between the installation position of the image sensor and the installation position of the depth sensor.
- using the focus position to control the shooting device to follow focus when shooting a video includes:
- the focus position is fused with focus positions determined based on other follow focus modes, and the fusion result is used to control the shooting device to follow focus when shooting a video.
- using the focus position to control the shooting device to follow focus when shooting a video includes:
- the depth information of the region of interest is fused with the depth information determined based on other follow focus modes, and the focus position determined by the fusion result is used to control the shooting device to shoot video. follow focus.
- an embodiment of the application further provides a photographing device 300 , including: an image sensor 301 , a depth sensor 302 , and the follow focus device 200 described in any embodiment.
- the embodiments of this specification further provide a computer-readable storage medium, where several computer instructions are stored on the readable storage medium, and when the computer instructions are executed, the steps of the focus method described in any one of the embodiments are performed.
- Embodiments of the present specification may take the form of a computer program product embodied on one or more storage media having program code embodied therein, including but not limited to disk storage, CD-ROM, optical storage, and the like.
- Computer-usable storage media includes permanent and non-permanent, removable and non-removable media, and storage of information can be accomplished by any method or technology.
- Information may be computer readable instructions, data structures, modules of programs, or other data.
- Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Flash Memory or other memory technology, Compact Disc Read Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, Magnetic tape cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
- PRAM phase-change memory
- SRAM static random access memory
- DRAM dynamic random access memory
- RAM random access memory
- ROM read only memory
- EEPROM Electrically Erasable Programmable Read Only Memory
- Flash Memory or other memory technology
- CD-ROM Compact Disc Read Only Memory
- CD-ROM Compact Disc Read Only Memory
- DVD Digital Versatile Disc
- Magnetic tape cassettes magnetic tape magnetic disk storage or other magnetic storage devices or any other non-
Landscapes
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Studio Devices (AREA)
Abstract
本申请提供一种跟焦方法、装置、拍摄设备及计算机可读存储介质,本申请实施例方案中,利用图像传感器获取第一视频帧,能够自动获取所述第一视频帧中的感兴趣区域;利用深度传感器获取携带有深度信息的第二视频帧,由于第一视频帧和第二视频帧包含相同画面内容,因此能够利用第二视频帧的深度信息来确定所述感兴趣区域的深度信息,如此,感兴趣区域的深度信息表征了被摄目标的深度信息,因此能准确地找到焦点位置,能够使拍摄设备在拍摄视频时实现自动跟焦,该方案能够同时保证:视频拍摄的连贯性、无需用户额外操作,并且还能准确地找到焦点位置。
Description
本申请涉及拍摄技术领域,具体而言,涉及一种跟焦方法、装置、拍摄设备及计算机可读存储介质。
拍摄设备在拍摄时,需要进行对焦处理,对焦处理的目的是找到准确的焦点位置,使被摄目标在感光器件平面上形成清晰的影像,保证所摄画面上主体影像的清晰。因此,准确地找到焦点位置是决定拍摄质量的关键因素之一。特别是在视频拍摄场景下,用户需要连续地拍摄,如何实时连续地找到焦点位置是亟待解决的技术问题。
发明内容
有鉴于此,本申请提供一种跟焦方法、装置、拍摄设备及计算机可读存储介质,以解决相关技术中拍摄时无法自动跟焦的问题。
第一方面,提供一种跟焦方法,所述方法包括:
利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;
利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;
利用所述第二视频帧确定所述感兴趣区域的深度信息;
基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
第二方面,提供一种跟焦装置,所述装置包括处理器、存储器、存储在所述存储器上可被所述处理器执行的计算机程序,所述处理器执行所述计算机程序时实现以下步骤:
利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;
利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;
利用所述第二视频帧确定所述感兴趣区域的深度信息;
基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
第三方面,提供一种拍摄设备,包括图像传感器和深度传感器;以及,如第二方面所述的跟焦装置。
第四方面,提供一种计算机可读存储介质,所述可读存储介质上存储有若干计算机指令,所述计算机指令被执行时实现第一方面所述的跟焦方法。
应用本申请提供的方案,利用图像传感器获取第一视频帧,能够自动获取所述第一视频帧中的感兴趣区域;利用深度传感器获取携带有深度信息的第二视频帧,由于第一视频帧和第二视频帧包含相同画面内容,因此能够利用第二视频帧的深度信息来 确定所述感兴趣区域的深度信息,如此,感兴趣区域的深度信息表征了被摄目标的深度信息,由于获取到的是整个区域的深度信息而并非单点深度信息,因此能准确地找到焦点位置,能够使拍摄设备在拍摄视频时实现自动跟焦,该方案能够同时保证:视频拍摄的连贯性、无需用户额外操作,并且还能准确地找到焦点位置;由于焦点位置能够准确地找到,因此在视频连续拍摄时可保证跟焦的实时性,不会导致明显的焦点过冲或虚焦的现象。
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1A是本申请一个实施例的跟焦方法的示意图。
图1B是本申请另一个实施例的跟焦方法的示意图。
图1C是本申请一个实施例的图像传感器和深度传感器的示意图。
图1D是本申请一个实施例的人脸区域的权重示意图。
图2是本申请一个实施例的跟焦装置的示意图。
图3是本申请一个实施例的拍摄设备的示意图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。
拍摄设备的ISP(Image Signal Processing,图像信号处理)系统中通常配置有对焦模块,一些对焦处理方案采用相位检测自动对焦,但此种方案对环境光线要求较高,在夜晚环境下效果较差;另一些对焦处理方案是采用激光方式获取到单点深度信息,基于单点深度信息可能无法获取到用户想要的拍摄主体的深度信息,无法实现准确对焦。因此,实际拍摄时很多时候用户都需要手动对焦。此种方式,在单次的图像拍摄场景下是可接受的,因为在此拍摄场景中,用户不具有连续拍摄的需求,用户操作对焦后再操作拍摄就完成了整个拍摄流程,手动对焦不会对用户造成较多干扰。
然而连续的视频拍摄场景,与单次的图像拍摄场景实际上是有诸多不同的;由于视频拍摄要保证连贯性,让用户中断连贯的拍摄去选取对焦位置,会给用户带来非常差的拍摄体验。特别是在一些专业视频拍摄领域,为了解决保证拍摄连贯性的同时准确对焦,需要有额外的跟焦师负责手动调节焦点,这为用户拍摄高质量的视频带来了较大的成本。
由此可知,在视频拍摄场景下,面临的问题是:一方面需要解放用户双手,无需用户手动操作对焦以保证视频拍摄的连贯性,一方面希望对焦处理只需要一个用户就能够完成以减少拍摄成本,再者,还要同时保证自动跟焦处理的准确性。
基于此,本申请实施例提供的跟焦方案,利用图像传感器获取第一视频帧,能够自动获取所述第一视频帧中的感兴趣区域;利用深度传感器获取携带有深度信息的第 二视频帧,由于第一视频帧和第二视频帧包含相同画面内容,因此能够利用第二视频帧的深度信息来确定所述感兴趣区域的深度信息,如此,感兴趣区域的深度信息表征了被摄目标的深度信息,由于获取到的是整个区域的深度信息而并非单点深度信息,因此能准确地找到焦点位置,能够使拍摄设备在拍摄视频时实现自动跟焦,该方案能够同时保证:视频拍摄的连贯性、无需用户额外操作,并且还能准确地找到焦点位置;由于焦点位置能够准确地找到,因此在视频连续拍摄时可保证跟焦的实时性,不会导致明显的焦点过冲或虚焦的现象。
本实施例方案可应用于拍摄设备中,由拍摄设备内置的处理器运行本实施例的跟焦方案以实现在视频拍摄时的自动跟焦。在一些例子中,拍摄设备还可以搭载在外部设备中,此处的外部设备可以包括可移动平台等设备,该可移动平台可以包括车辆、无人机或可移动机器人等。
结合图1A和图1B进行说明,图1A和图1B分别为本实施例示出的跟焦方法的示意图,可包括如下步骤:
在步骤102中,利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;
在步骤104中,利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;
在步骤106中,利用所述第二视频帧确定所述感兴趣区域的深度信息;
在步骤108中,基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
本实施例的跟焦方案中适用于视频拍摄场景,拍摄设备配置有图像传感器,图像传感器用于连续采集图像,本实施例通过图像传感器获取第一视频帧,该第一视频帧可以是图像传感器采集的原始图像,也可以是对图像传感器采集的原始图像经过一些处理后得到的图像。
本实施例方案中,拍摄设备还配置有深度传感器,以在视频拍摄场景中连续地采集携带有深度信息的图像,本实施例通过深度传感器获取第二视频帧,该第二视频帧可以是深度传感器采集的原始图像,也可以是对深度传感器采集的原始图像经过一些处理后得到的图像。
本实施例中,第一视频帧和第二视频帧包含相同画面内容,利用第一视频帧自动识别出感兴趣区域,并利用所述第二视频帧携带的深度信息来确定感兴趣区域的深度信息。
实际应用中可以通过多种方式令第一视频帧和第二视频帧包含相同画面内容。作为例子,可以在所述拍摄设备中,令所述深度传感器与所述图像传感器平行,使得深度传感器与图像传感器分别采集的图像可以对齐,如图1C所示,是一个实施例中拍摄设备内所述深度传感器与所述图像传感器的示意图,图1C中,深度传感器以TOF(Time of flight,飞行时间)传感器为例,该深度传感器TOF通过一支撑器件TOF holder架设与图像传感器Image sensor两者处于平行状态,处于平行状态是指深度传感器与水平面的夹角X deg,与图像传感器Image sensor与水平面的夹角相等,因此能够使得所述深度传感器采集的视频帧与所述图像传感器采集的视频帧对应同一画面,即第一视频帧和第二视频帧包含相同画面内容。
当然,深度传感器与图像传感器也可以不具有上述平行状态的限制;在另一些例子中,深度传感器与图像传感器可以处于非平行状态,即可以根据需要配置深度传感器与图像传感器处于其他位置关系,本实施例方案还可以对图像传感器采集的图像与深度传感器采集的图像进行图像对齐处理,以获得所述第一视频帧和第二视频帧。
作为例子,深度传感器与图像传感器的安装位置可能没有处于平行状态,图像传感器与深度传感器具有角度偏差,从而图像传感器采集的图像与深度传感器采集的图像具有角度偏差,基于此,本实施例还可以进行图像对齐处理,作为例子,可以包括:将所述图像传感器采集的图像与所述深度传感器采集的图像进行轮廓对齐,例如,根据深度传感器与图像传感器的角度差异,将所述所述图像传感器采集的图像与所述深度传感器采集的图像的轮廓对齐。
在另一些例子中,考虑到深度传感器与图像传感器两者无法完全重合,深度传感器与图像传感器分别对应的视角范围可能无法完全相同,即两者采集的图像中所包含的画面并未完全相同,深度传感器采集的图像所包含的画面,可能存在图像传感器采集的图像所未能包含的画面,图像传感器采集的图像所包含的画面,也可能存在深度传感器采集的图像所未能包含的画面,因此,本实施例将两者采集的图像中不同的画面称之为偏移区域,本实施例通过将所述图像传感器采集的图像与所述深度传感器采集的图像两者的偏移区域进行裁剪,使得裁剪后的第一视频帧和第二视频帧两者包含的画面相同,进而可以便捷地获取到感兴趣区域的深度信息。
在一些例子中,所述深度传感器可以包括3D ToF(TOF,Time-of-Flight)传感器,通过3D ToF传感器采集图像,不会受夜晚等光照较差的影响,使得本实施例的跟焦方案能够适用于夜晚环境;另一方面3D ToF传感器能够采集到大范围内的深度信息,从而使本实施例的跟焦方案能够准确地找到焦点位置。
本实施例通过对第一视频帧自动识别出感兴趣区域来自动确定拍摄主体,利用所述第二视频帧确定所述感兴趣区域的深度信息,作为例子,该处理过程可以包括:利用所述感兴趣区域在所述第一视频帧中的位置,从所述第二视频帧中确定所述感兴趣区域的深度信息。例如,第一视频帧和第二视频帧包含同一画面,识别出的感兴趣区域,能够确定该感兴趣区域在第一视频帧中的位置,则对应地可以确定该感兴趣区域在第二视频帧中的位置,根据第二视频帧携带的深度信息可以确定出所述感兴趣区域的深度信息。
在确定了感兴趣区域的深度信息后,可进一步确定焦点位置。感兴趣区域表征了第一视频帧中的拍摄主体所在的区域,从感兴趣区域的深度信息可以确定焦点位置,具体的确定方式实际应用中可以根据需要灵活配置。
作为例子,感兴趣区域包含了多个像素点,可以基于所述感兴趣区域的各个目标像素点的深度信息来确定感兴趣区域的平均深度信息,进而利用所述感兴趣区域的平均深度信息来确定焦点位置。
可选的,所述感兴趣区域的平均深度信息是基于所述感兴趣区域中各个目标像素点的深度信息以及各个所述目标像素点的预设权重确定的。例如,可以根据感兴趣区域中对应的拍摄主体,预先对感兴趣区域划分为不同的子区域,并对不同的子区域设置权重,在确定焦点位置时,将实际得到的感兴趣区域中各个目标像素点乘以该目标像素点的权重并求平均值,进而得到感兴趣区域的平均深度信息,进一步可确定焦点 位置。
其中,目标像素点可以是感兴趣区域中的所有像素点,也可以是部分像素点,例如,可以通过对感兴趣区域的深度信息进行去噪处理后得到的像素点。实际应用中,感兴趣区域中可能有部分像素点的深度信息不可靠,可以剔除掉部分不可靠的像素点。
作为例子,考虑到拍摄主体与拍摄设备的距离应该是在一个合理的范围内,如果第二视频帧中携带的深度信息表征拍摄主体非常近或者非常远,不在合理的范围内,可以认为该采样点的深度信息不可靠,该采样点可能不属于拍摄主体,或者是深度传感器的采集数据有误,基于此,为了防止这些可能错误的像素点的干扰,可以将这些像素点剔除。如何确定这些深度信息可能错误的像素点,在一些例子中,可以预先配置置信度阈值,也即是所述目标像素点是所述感兴趣区域中深度信息大于设定置信度阈值的像素点,深度信息满足置信度阈值的像素点则保留,不满足置信度阈值的像素点则剔除。
在一些例子中,所述置信度阈值是基于所述深度传感器的回波强度设定的。实际应用中,不同的深度传感器可能有不同的回波感应能力,正常情况下深度传感器采集到的深度信息也会处于一个合理的范围内,本实施例可以预先基于深度传感器的回波强度来确定置信度阈值,从而使得跟焦处理方案的准确度更为精准。
实际应用中,可以根据需要确定不同拍摄场景下的感兴趣区域。作为例子,若包含有人脸画面,感兴趣区域可以是从所述第一视频帧中识别得到的人脸区域;在另一些例子,在未包含有人脸画面时,感兴趣区域还可以是从所述第一视频帧中识别得到的其他动物的脸部区域;在其他拍摄场景中还可以是其他物体的区域;实际实现时,可以利用物体识别算法等方式自动识别出感兴趣区域。在一些例子中,从第一视频帧中可能还会自动识别出多个感兴趣区域,对于识别出的多个感兴趣区域,还可以根据设定的优先级自动保留其中一个优先级最高的,而其他的感兴趣区域剔除,例如,人脸区域的优先级最高,动物脸部区域次之,其他不同类别的物品次之等,实际应用中可以根据需要进行配置,本实施例对此不作限定。
在一些例子中,还可能识别到多个人脸,可以根据需要自动选取其中一个人脸所在区域作为感兴趣区域,例如可以基于各个人脸区域的大小确定,或者是基于各个人脸区域的深度信息的高低来确定,例如将人脸所在区域面积较大的作为感兴趣区域,或者是将人脸深度信息较小的作为感兴趣区域等。
在另一些例子中,还可以根据用户指定的需要跟焦的人脸区域来确定所述第一视频帧中的感兴趣区域。作为例子,用户可以在拍摄前或拍摄过程中指定要跟焦的人脸,根据用户所指定的人脸,在后续的连续拍摄过程中可以在第一视频帧中识别出用户要跟焦的人脸所在的区域为感兴趣区域。
为了便于用户指定要跟焦的人脸区域,在一些例子中,所述获取所述第一视频帧中的感兴趣区域,可以包括:在拍摄设备的预览画面中突出显示从图像传感器采集的图像中识别出的至少两个人脸区域后,获取用户指定的需跟焦的人脸区域;利用所述需跟焦的人脸区域,获取所述第一视频帧中的人脸区域;基于此,通过突出显示可以供用户快速便捷地指定其想要跟焦的人脸,用户指定的需跟焦的人脸后,接下来进行视频拍摄时,由于用户指定人脸时的视频帧与接下来的视频拍摄时需要进行跟焦的视频帧已经不同,本实施例可以利用用户指定的需跟焦的人脸的人脸特征,从第一视频 帧中确定感兴趣区域。如此,拍摄时在有多个人脸区域的情况下,本实施例通过用户指定的需跟焦的人脸的人脸特征,能够从拍摄的视频帧中自动识别出来作为感兴趣区域。基于此,在视频拍摄时,用户只需要执行一次指定的操作,后续就可以实现全程的自动连续跟焦。
以人脸作为感兴趣区域,在根据感兴趣区域的深度信息确定焦点位置时,可以根据需要为人脸区域中的不同区域设定权重,作为例子,如图1D所示,是本实施例中一种人脸区域的权重的示意图,所述人脸区域中,眼部区域像素点的权重、嘴部区域像素点的权重、鼻子区域像素点的权重和边缘区域像素点的权重依次递减。实际应用中可以根据需要设置该权重的其他实现方式,例如根据需要划分为多个其他区域,并对不同区域配置不同的权重,本实施例对此不作限定。
实际应用中,在连续拍摄时,图像传感器会持续采集图像,深度传感器也会持续地采集图像,本实施例的跟焦方案可以持续地执行以找出焦点位置。在一些例子中,可以是按照设定时间间隔持续地执行本实施例的跟焦方案;在另一些例子中,可以以图像传感器采集的图像为基准,基于每一帧第一视频帧执行本实施例的跟焦方案;在另一些例子中,还可以是以深度传感器采集的图像为基准,基于每一帧第二视频帧执行本实施例的跟焦方案。
其中,在一些例子中,深度传感器的帧率(FPS,Frames Per Second,每秒传输帧数)与所述图像传感器的帧率相同,在拍摄时深度传感器与图像传感器同时启动,则深度传感器与图像传感器分别采集的图像在时间上可以自动对齐,即图像传感器和深度传感器分别采集的图像的时间戳相同。
在另一些例子中,所述深度传感器的帧率与所述图像传感器的帧率不同,则所述第二视频帧可以是与第一视频帧的采集时刻时间最接近的视频帧,使得第一视频帧和第二视频帧在时间差异最小,以保证利用第二视频帧获取的深度信息与第一视频帧中拍摄主体相匹配,保证跟焦方案的精确度。
其中,在深度传感器的帧率与所述图像传感器的帧率不同的情况下,本实施例的跟焦方案,可以是以图像传感器采集图像为基准,基于第一视频帧获取时间最接近的第二视频帧;也可以是以深度传感器采集图像为基准,基于第二视频帧获取时间最接近的第一视频帧。
在深度传感器的帧率与所述图像传感器的帧率不同的情况下,本实施例中可以设置缓存区以缓存图像,以获取到时间差异最小的第一视频帧和第二视频帧,进而提高跟焦准确度。
在一些例子中,以图像传感器采集图像为基准,缓存区中缓存的图像包括第二视频帧,即缓存区中缓存有所述深度传感器采集的携带有时间戳的至少一帧图像;在基于第一视频帧获取时间最接近的第二视频帧时,可以是从缓存区中根据图像的时间戳获取时间最接近的第二视频帧。
在另一些例子中,以深度传感器采集图像为基准,缓存区中缓存的图像包括第一视频帧,即缓存区中缓存有所述图像传感器采集的携带有时间戳的至少一帧图像;在基于第二视频帧获取时间最接近的第一视频帧时,可以是从缓存区中根据图像的时间戳获取时间最接近的第一视频帧。
在深度传感器的帧率与所述图像传感器的帧率不同的情况下,为了保证能够精准 地从缓存区中获取到所需的图像,本实施例还基于所述缓存区中缓存的图像的缓存时长或缓存区中缓存的图像的帧数提供了创新的解决方案。
在一些例子中,所述缓存区中缓存的图像的帧数,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
在一些例子中,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
深度传感器的帧率与所述图像传感器的帧率不同的情况下,作为例子:
(1)在一些例子中,以图像传感器采集图像为基准,缓存区用于存储深度传感器采集的图像;
①假设深度传感器的帧率大于所述图像传感器的帧率,缓存区中需要存储深度传感器采集的至少一帧图像,以在基于第一视频帧获取第二视频帧时,能够找到与第一视频帧时间最接近的第二视频帧。因为在图像传感器采集到一帧图像的时间间隔内,深度传感器已经能够采集到至少一帧图像,因此需要缓存深度传感器采集到的至少一帧图像。
为了保证能够获取到与第一视频帧时间最接近的第二视频帧,作为例子,所述缓存区中缓存的图像的帧数,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的商值。
例如,深度传感器的帧率为40FPS,图像传感器的帧率为30FPS;深度传感器每隔1/40秒采集一帧图像,图像传感器每隔1/30秒采集一帧图像,缓存区中缓存的深度传感器采集的图像的帧数,至少要大于1/30与1/40的商值,即至少要大于2帧,即深度传感器采集2帧图像的时间内会产生一帧图像传感器采集的图像,因此至少要缓存2帧,这样可保证在基于第一视频帧获取第二视频帧时,能够准确找到与第一视频帧时间最接近的第二视频帧。
②假设深度传感器的帧率小于所述图像传感器的帧率,缓存区中需要将深度传感器采集的至少一帧图像缓存一定时间,以在基于第一视频帧获取第二视频帧时,能够找到与第一视频帧时间最接近的第二视频帧。因为在图像传感器采集到一帧图像的时间间隔内,深度传感器还未能够采集到图像,因此需要将深度传感器采集到的图像缓存一定时间。
为了保证能够获取到与第一视频帧时间最接近的第二视频帧,作为例子,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
例如,深度传感器的帧率为30FPS,图像传感器的帧率为40FPS;深度传感器每隔1/30秒采集一帧图像,图像传感器每隔1/40采集一帧图像,缓存区中缓存的深度传感器采集的图像的时长,至少要大于1/30与1/40的差值,即至少要大于1/120秒,才可保证在基于第一视频帧获取第二视频帧时,能够找到与第一视频帧时间最接近的第二视频帧。
(2)在另一些例子中,以深度传感器采集图像为基准,缓存区用于存储图像传感器采集的图像,与前述实施例相同的原理:
①假设深度传感器的帧率大于所述图像传感器的帧率,缓存区中需要存储图像传感器采集的多帧图像,以在基于第二视频帧获取第一视频帧时,能够找到与第二视频 帧时间最接近的第一视频帧。因为在深度传感器采集到一帧图像的时间间隔内,图像传感器还未能够采集到图像,因此需要将图像传感器采集到的图像缓存一定时间。
为了保证能够获取到与第二视频帧的采集时刻时间最接近的第一视频帧,作为例子,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
例如,深度传感器的帧率为40FPS,图像传感器的帧率为30FPS;深度传感器每隔1/40秒采集一帧图像,图像传感器每隔1/30采集一帧图像,缓存区中缓存的图像传感器采集的图像的缓存时长,至少要大于1/30与1/40的差值,即至少要大于1/120秒,才可保证在基于第二视频帧获取第一视频帧时,能够找到与第二视频帧采集时刻的时间最接近的第一视频帧。
②假设深度传感器的帧率小于所述图像传感器的帧率,缓存区中需要存储图像传感器采集的多帧图像,以在基于第二视频帧获取第一视频帧时,能够找到与第二视频帧时间最接近的第一视频帧。因为在深度传感器采集到一帧图像的时间间隔内,图像传感器已经能够采集到至少一帧图像,因此需要缓存图像传感器采集到的至少一帧图像。
为了保证能够获取到与第二视频帧时间最接近的第一视频帧,作为例子,所述缓存区中缓存的图像的帧数,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的商值。
例如,深度传感器的帧率为30FPS,图像传感器的帧率为40FPS;深度传感器每隔1/30秒采集一帧图像,图像传感器每隔1/40秒采集一帧图像,缓存区中缓存的图像传感器采集的图像的帧数,至少要大于1/30与1/40的商值,即至少要大于2帧,即图像传感器在采集两帧图像的时间内会产生一帧深度传感器采集的图像,因此至少要缓存2帧,才可保证在基于第二视频帧获取第一视频帧时,能够找到与第二视频帧时间最接近的第一视频帧。
实际应用中,考虑到跟焦方案的执行需要一定时间,其中可能还涉及其他的图像处理,缓存区中缓存的图像还可以包括其他图像,缓存的图像的帧数及时长还可以根据需要结合其他因素进行配置,例如进一步缓存更多帧数的图像,或者是将图像的缓存时间进一步延长,本实施例对此不作限定。
实际应用中,在一例子中,图像传感器的分辨率与深度传感器的分辨率相同,可以便捷地利用第二视频帧的深度信息确定感兴趣区域的深度信息。
在另一些例子中,图像传感器的分辨率与深度传感器的分辨率不同,可以对分辨率较低的图像进行插值处理以使两者的分辨率相同,之后再利用第二视频帧的深度信息确定感兴趣区域的深度信息。
实际应用中,通常是所述图像传感器采集的图像的第一分辨率高于所述深度传感器采集的图像的第二分辨率,所述第二视频帧是对所述深度传感器采集的图像进行插值处理得到的与所述第一分辨率相同的视频帧。其中,插值处理的实现过程可以根据需要灵活配置,本实施例对此不作限定。
在一些例子中,本实施例方案还可以同时对所述深度传感器采集的图像进行与图像传感器采集的图像的图像对齐处理时,还可以同时对所述深度传感器采集的图像进行插值处理。图像对齐处理的实施例可见于前述描述,在此不再赘述。
本实施例中,可以基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦;作为例子,可以将深度信息作为物距代入至高斯成像公式1/f=1/u+1/v,其中,u为物距,v为像距,f为焦距,进而可以得出像距,根据像距得到焦点位置,计算得的焦点位置,可以基于此控制拍摄设备在拍摄下一视频帧时的焦点位置,从而实现控制拍摄设备在拍摄视频时进行跟焦。
在一些例子中,可以将计算得到的焦点位置直接作为拍摄设备在拍摄下一视频帧时的焦点位置。
在另一些例子中,考虑到利用本实施例的跟焦方法得到的焦点位置的置信度也有可能无法满足拍摄要求,而拍摄设备中配置有其他跟焦模式,基于此,可以综合考虑各种不同跟焦模式获得的结果,将本实施例的跟焦方案与其他跟焦模式的结果进行融合。融合的处理方式可以有多种:
作为例子,一种可选的方式是,在利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,基于所述感兴趣区域的深度信息的整体置信度,将所述焦点位置与基于其他跟焦模式确定的焦点位置进行融合,利用融合结果控制拍摄设备在拍摄视频时进行跟焦。作为例子,感兴趣区域的深度信息的整体置信度,可以是基于感兴趣区域中目标像素点的深度信息的置信度的平均值而确定的。
在另一些例子中,在利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦时,可以基于所述感兴趣区域的深度信息的整体置信度,将所述感兴趣区域的深度信息与基于其他跟焦模式确定的深度信息进行融合,利用融合结果确定出的焦点位置控制拍摄设备在拍摄视频时进行跟焦。
当然,实际应用中,若利用本实施例的跟焦方案获得的感兴趣区域的深度信息的整体置信度较低,也可以终止流程,采用其他方式获得焦点位置。
上述方法实施例可以通过软件实现,也可以通过硬件或者软硬件结合的方式实现。以软件实现为例,作为一个逻辑意义上的装置,是通过其所在处理器将存储器中对应的计算机程序指令读取到内存中运行形成的。从硬件层面而言,如图2所示,为实施本实施例图像处理方法的跟焦装置200的一种硬件结构图,除了图2所示的处理器201、以及存储器202之外,实施例中用于实施本跟焦方法的装置,通常根据该跟焦装置的实际功能,还可以包括其他硬件,对此不再赘述。
本实施例中,所述处理器201执行所述计算机程序时实现以下步骤:
利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;
利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;
利用所述第二视频帧确定所述感兴趣区域的深度信息;
基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
在一些例子中,在所述拍摄设备中,所述深度传感器与所述图像传感器平行。
在一些例子中,所述深度传感器包括:3D ToF传感器。
在一些例子中,所述利用所述第二视频帧确定所述感兴趣区域的深度信息,包括:
利用所述感兴趣区域在所述第一视频帧中的位置,从所述第二视频帧中确定所述感兴趣区域的深度信息。
在一些例子中,所述焦点位置是利用所述感兴趣区域的平均深度信息确定的,所述感兴趣区域的平均深度信息是基于所述感兴趣区域的各个目标像素点的深度信息确定的。
在一些例子中,所述感兴趣区域的平均深度信息是基于所述感兴趣区域中各个目标像素点的深度信息以及各个所述目标像素点的预设权重确定的。
在一些例子中,所述目标像素点是所述感兴趣区域中深度信息大于设定置信度阈值的像素点。
在一些例子中,所述置信度阈值是基于所述深度传感器的回波强度设定的。
在一些例子中,所述感兴趣区域包括:从所述第一视频帧中识别得到的人脸区域。
在一些例子中,所述人脸区域中,眼部区域像素点的权重、嘴部区域像素点的权重、鼻子区域像素点的权重和边缘区域像素点的权重依次递减。
在一些例子中,所述感兴趣区域是基于用户指定的需跟焦的人脸区域确定的。
在一些例子中,所述获取所述第一视频帧中的感兴趣区域:
在拍摄设备的预览画面中突出显示从图像传感器采集的图像中识别出的至少两个人脸区域后,获取用户指定的需跟焦的人脸区域;
利用所述需跟焦的人脸区域,获取所述第一视频帧中的人脸区域。
在一些例子中,所述深度传感器的帧率与所述图像传感器的帧率不同;
所述第二视频帧是与第一视频帧的采集时刻时间最接近的视频帧。
在一些例子中,所述第二视频帧是从缓存区中获取的,所述缓存区中缓存有所述深度传感器采集的携带有时间戳的至少一帧图像。
在一些例子中,所述第一视频帧是从缓存区中获取的,所述缓存区中缓存有所述图像传感器采集的携带有时间戳的至少一帧图像。
在一些例子中,所述缓存区中缓存的图像的帧数,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
在一些例子中,所述缓存区中缓存的图像的缓存时长,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
在一些例子中,所述缓存区中缓存的图像的帧数,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的商值。
在一些例子中,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
在一些例子中,所述图像传感器采集的图像的第一分辨率高于所述深度传感器采集的图像的第二分辨率;
所述第二视频帧是对所述深度传感器采集的图像进行插值处理得到的与所述第一分辨率相同的视频帧。
在一些例子中,所述第一视频帧与第二视频帧是经过图像对齐处理后得到的。
在一些例子中,所述图像对齐处理,包括如下任一:将所述图像传感器采集的图像与所述深度传感器采集的图像进行轮廓对齐,或,将所述图像传感器采集的图像与所述深度传感器采集的图像两者的偏移区域进行裁剪处理。
在一些例子中,所述偏移区域,是基于所述图像传感器的安装位置和所述深度传感器的安装位置的位置差异确定的。
在一些例子中,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:
基于所述感兴趣区域的深度信息的整体置信度,将所述焦点位置与基于其他跟焦模式确定的焦点位置进行融合,利用融合结果控制拍摄设备在拍摄视频时进行跟焦。
在一些例子中,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:
基于所述感兴趣区域的深度信息的整体置信度,将所述感兴趣区域的深度信息与基于其他跟焦模式确定的深度信息进行融合,利用融合结果确定出的焦点位置控制拍摄设备在拍摄视频时进行跟焦。
如图3所示,是申请实施例还提供一种拍摄设备300,包括:图像传感器301和深度传感器302,以及任一实施例所述的跟焦装置200。
本说明书实施例还提供一种计算机可读存储介质,所述可读存储介质上存储有若干计算机指令,所述计算机指令被执行时实任一实施例所述跟焦方法的步骤。
本说明书实施例可采用在一个或多个其中包含有程序代码的存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。计算机可用存储介质包括永久性和非永久性、可移动和非可移动媒体,可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括但不限于:相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。
对于装置实施例而言,由于其基本对应于方法实施例,所以相关之处参见方法实施例的部分说明即可。以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。本领域普通技术人员在不付出创造性劳动的情况下,即可以理解并实施。
需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
以上对本发明实施例所提供的方法和装置进行了详细介绍,本文中应用了具体个例对本发明的原理及实施方式进行了阐述,以上实施例的说明只是用于帮助理解本发明的方法及其核心思想;同时,对于本领域的一般技术人员,依据本发明的思想,在具体实施方式及应用范围上均会有改变之处,综上所述,本说明书内容不应理解为对 本发明的限制。
Claims (52)
- 一种跟焦方法,其特征在于,所述方法包括:利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;利用所述第二视频帧确定所述感兴趣区域的深度信息;基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
- 根据权利要求1所述的方法,其特征在于,在所述拍摄设备中,所述深度传感器与所述图像传感器平行。
- 根据权利要求1所述的方法,其特征在于,所述深度传感器包括:3D ToF传感器。
- 根据权利要求2所述的方法,其特征在于,所述利用所述第二视频帧确定所述感兴趣区域的深度信息,包括:利用所述感兴趣区域在所述第一视频帧中的位置,从所述第二视频帧中确定所述感兴趣区域的深度信息。
- 根据权利要求1所述的方法,其特征在于,所述焦点位置是利用所述感兴趣区域的平均深度信息确定的,所述感兴趣区域的平均深度信息是基于所述感兴趣区域的各个目标像素点的深度信息确定的。
- 根据权利要求5所述的方法,其特征在于,所述感兴趣区域的平均深度信息是基于所述感兴趣区域中各个目标像素点的深度信息以及各个所述目标像素点的预设权重确定的。
- 根据权利要求5或6所述的方法,其特征在于,所述目标像素点是所述感兴趣区域中深度信息大于设定置信度阈值的像素点。
- 根据权利要求7所述的方法,其特征在于,所述置信度阈值是基于所述深度传感器的回波强度设定的。
- 根据权利要求6所述的方法,其特征在于,所述感兴趣区域包括:从所述第一视频帧中识别得到的人脸区域。
- 根据权利要求9所述的方法,其特征在于,所述人脸区域中,眼部区域像素点的权重、嘴部区域像素点的权重、鼻子区域像素点的权重和边缘区域像素点的权重依次递减。
- 根据权利要求1所述的方法,其特征在于,所述感兴趣区域是基于用户指定的需跟焦的人脸区域确定的。
- 根据权利要求11所述的方法,其特征在于,所述获取所述第一视频帧中的感兴趣区域:在拍摄设备的预览画面中突出显示从图像传感器采集的图像中识别出的至少两个人脸区域后,获取用户指定的需跟焦的人脸区域;利用所述需跟焦的人脸区域,获取所述第一视频帧中的人脸区域。
- 根据权利要求1所述的方法,其特征在于,所述深度传感器的帧率与所述图像传感器的帧率不同;所述第二视频帧是与第一视频帧的采集时刻时间最接近的视频帧。
- 根据权利要求13所述的方法,其特征在于,所述第二视频帧是从缓存区中获取的,所述缓存区中缓存有所述深度传感器采集的携带有时间戳的至少一帧图像。
- 根据权利要求13所述的方法,其特征在于,所述第一视频帧是从缓存区中获取的,所述缓存区中缓存有所述图像传感器采集的携带有时间戳的至少一帧图像。
- 根据权利要求14或15所述的方法,其特征在于,所述缓存区中缓存的图像的帧数,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
- 根据权利要求14或15所述的方法,其特征在于,所述缓存区中缓存的图像的缓存时长,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
- 根据权利要求16所述的方法,其特征在于,所述缓存区中缓存的图像的帧数,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的商值。
- 根据权利要求17所述的方法,其特征在于,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
- 根据权利要求1所述的方法,其特征在于,所述图像传感器采集的图像的第一分辨率高于所述深度传感器采集的图像的第二分辨率;所述第二视频帧是对所述深度传感器采集的图像进行插值处理得到的与所述第一分辨率相同的视频帧。
- 根据权利要求1所述的方法,其特征在于,所述第一视频帧与第二视频帧是经过图像对齐处理后得到的。
- 根据权利要求21所述的方法,其特征在于,所述图像对齐处理,包括如下任一:将所述图像传感器采集的图像与所述深度传感器采集的图像进行轮廓对齐,或,将所述图像传感器采集的图像与所述深度传感器采集的图像两者的偏移区域进行裁剪 处理。
- 根据权利要求22所述的方法,其特征在于,所述偏移区域,是基于所述图像传感器的安装位置和所述深度传感器的安装位置的位置差异确定的。
- 根据权利要求1所述的方法,其特征在于,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:基于所述感兴趣区域的深度信息的整体置信度,将所述焦点位置与基于其他跟焦模式确定的焦点位置进行融合,利用融合结果控制拍摄设备在拍摄视频时进行跟焦。
- 根据权利要求1所述的方法,其特征在于,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:基于所述感兴趣区域的深度信息的整体置信度,将所述感兴趣区域的深度信息与基于其他跟焦模式确定的深度信息进行融合,利用融合结果确定出的焦点位置控制拍摄设备在拍摄视频时进行跟焦。
- 一种跟焦装置,其特征在于,所述装置包括处理器、存储器、存储在所述存储器上可被所述处理器执行的计算机程序,所述处理器执行所述计算机程序时实现以下步骤:利用图像传感器获取第一视频帧后,获取所述第一视频帧中的感兴趣区域;利用深度传感器获取携带有深度信息的第二视频帧;其中,所述第一视频帧和第二视频帧包含相同画面内容;利用所述第二视频帧确定所述感兴趣区域的深度信息;基于所述感兴趣区域的深度信息确定焦点位置,利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦。
- 根据权利要求26所述的装置,其特征在于,在所述拍摄设备中,所述深度传感器与所述图像传感器平行。
- 根据权利要求27所述的装置,其特征在于,所述深度传感器包括:3D ToF传感器。
- 根据权利要求27所述的装置,其特征在于,所述利用所述第二视频帧确定所述感兴趣区域的深度信息,包括:利用所述感兴趣区域在所述第一视频帧中的位置,从所述第二视频帧中确定所述感兴趣区域的深度信息。
- 根据权利要求26所述的装置,其特征在于,所述焦点位置是利用所述感兴趣区域的平均深度信息确定的,所述感兴趣区域的平均深度信息是基于所述感兴趣区域 的各个目标像素点的深度信息确定的。
- 根据权利要求30所述的装置,其特征在于,所述感兴趣区域的平均深度信息是基于所述感兴趣区域中各个目标像素点的深度信息以及各个所述目标像素点的预设权重确定的。
- 根据权利要求30或31所述的装置,其特征在于,所述目标像素点是所述感兴趣区域中深度信息大于设定置信度阈值的像素点。
- 根据权利要求32所述的装置,其特征在于,所述置信度阈值是基于所述深度传感器的回波强度设定的。
- 根据权利要求31所述的装置,其特征在于,所述感兴趣区域包括:从所述第一视频帧中识别得到的人脸区域。
- 根据权利要求34所述的装置,其特征在于,所述人脸区域中,眼部区域像素点的权重、嘴部区域像素点的权重、鼻子区域像素点的权重和边缘区域像素点的权重依次递减。
- 根据权利要求26所述的装置,其特征在于,所述感兴趣区域是基于用户指定的需跟焦的人脸区域确定的。
- 根据权利要求36所述的装置,其特征在于,所述获取所述第一视频帧中的感兴趣区域:在拍摄设备的预览画面中突出显示从图像传感器采集的图像中识别出的至少两个人脸区域后,获取用户指定的需跟焦的人脸区域;利用所述需跟焦的人脸区域,获取所述第一视频帧中的人脸区域。
- 根据权利要求26所述的装置,其特征在于,所述深度传感器的帧率与所述图像传感器的帧率不同;所述第二视频帧是与第一视频帧的采集时刻时间最接近的视频帧。
- 根据权利要求38所述的装置,其特征在于,所述第二视频帧是从缓存区中获取的,所述缓存区中缓存有所述深度传感器采集的携带有时间戳的至少一帧图像。
- 根据权利要求38所述的装置,其特征在于,所述第一视频帧是从缓存区中获取的,所述缓存区中缓存有所述图像传感器采集的携带有时间戳的至少一帧图像。
- 根据权利要求39或40所述的装置,其特征在于,所述缓存区中缓存的图像的帧数,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
- 根据权利要求39或40所述的装置,其特征在于,所述缓存区中缓存的图像的缓存时长,是基于所述深度传感器的帧率与所述图像传感器的帧率的差异确定的。
- 根据权利要求41所述的装置,其特征在于,所述缓存区中缓存的图像的帧数,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的商值。
- 根据权利要求42所述的装置,其特征在于,所述缓存区中缓存的图像的缓存时长,大于或等于所述深度传感器的帧率倒数与所述图像传感器的帧率倒数的差值。
- 根据权利要求26所述的装置,其特征在于,所述图像传感器采集的图像的第一分辨率高于所述深度传感器采集的图像的第二分辨率;所述第二视频帧是对所述深度传感器采集的图像进行插值处理得到的与所述第一分辨率相同的视频帧。
- 根据权利要求45所述的装置,其特征在于,在对所述深度传感器采集的图像进行插值处理时,还对所述深度传感器采集的图像进行与所述第一视频帧的图像对齐处理。
- 根据权利要求46所述的装置,其特征在于,所述图像对齐处理,包括如下任一:将所述深度传感器采集的图像与所述第一视频帧的轮廓对齐,或,将所述深度传感器采集的图像与所述第一视频帧两者的偏移区域进行裁剪。
- 根据权利要求47所述的装置,其特征在于,所述偏移区域,是基于所述图像传感器的安装位置和所述深度传感器的安装位置的位置差异确定的。
- 根据权利要求26所述的装置,其特征在于,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:基于所述感兴趣区域的深度信息的整体置信度,将所述焦点位置与基于其他跟焦模式确定的焦点位置进行融合,利用融合结果控制拍摄设备在拍摄视频时进行跟焦。
- 根据权利要求26所述的装置,其特征在于,所述利用所述焦点位置控制拍摄设备在拍摄视频时进行跟焦,包括:基于所述感兴趣区域的深度信息的整体置信度,将所述感兴趣区域的深度信息与基于其他跟焦模式确定的深度信息进行融合,利用融合结果确定出的焦点位置控制拍摄设备在拍摄视频时进行跟焦。
- 一种拍摄设备,其特征在于,包括图像传感器和深度传感器;以及,如权利要求26至50任意一项所述跟焦装置。
- 一种计算机可读存储介质,其特征在于,所述可读存储介质上存储有若干计算机指令,所述计算机指令被执行时实现权利要求1至25任一项所述方法的步骤。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2021/070581 WO2022147703A1 (zh) | 2021-01-07 | 2021-01-07 | 跟焦方法、装置、拍摄设备及计算机可读存储介质 |
| CN202180079145.3A CN116507970A (zh) | 2021-01-07 | 2021-01-07 | 跟焦方法、装置、拍摄设备及计算机可读存储介质 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2021/070581 WO2022147703A1 (zh) | 2021-01-07 | 2021-01-07 | 跟焦方法、装置、拍摄设备及计算机可读存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022147703A1 true WO2022147703A1 (zh) | 2022-07-14 |
Family
ID=82357040
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/070581 Ceased WO2022147703A1 (zh) | 2021-01-07 | 2021-01-07 | 跟焦方法、装置、拍摄设备及计算机可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116507970A (zh) |
| WO (1) | WO2022147703A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115623318A (zh) * | 2022-12-20 | 2023-01-17 | 荣耀终端有限公司 | 对焦方法及相关装置 |
| WO2025180224A1 (zh) * | 2024-02-29 | 2025-09-04 | 荣耀终端股份有限公司 | 自动对焦方法、电子设备、存储介质及程序产品 |
| WO2026040469A1 (zh) * | 2024-08-21 | 2026-02-26 | 荣耀终端股份有限公司 | 图像处理方法、终端设备、芯片系统及可读存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105264436A (zh) * | 2013-04-05 | 2016-01-20 | 安德拉运动技术股份有限公司 | 用于控制与图像捕捉有关的设备的系统和方法 |
| US20170374354A1 (en) * | 2015-03-16 | 2017-12-28 | SZ DJI Technology Co., Ltd. | Apparatus and method for focal length adjustment and depth map determination |
| CN109696667A (zh) * | 2018-12-19 | 2019-04-30 | 哈工大机器人(合肥)国际创新研究院 | 一种基于3DToF摄像头的门禁道闸检测方法 |
| CN110381261A (zh) * | 2019-08-29 | 2019-10-25 | 重庆紫光华山智安科技有限公司 | 聚焦方法、装置、计算机可读存储介质及电子设备 |
| CN111226154A (zh) * | 2018-09-26 | 2020-06-02 | 深圳市大疆创新科技有限公司 | 自动对焦相机和系统 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| TWI471630B (zh) * | 2012-06-01 | 2015-02-01 | Hon Hai Prec Ind Co Ltd | 主動式距離對焦系統及方法 |
| CN106454090B (zh) * | 2016-10-09 | 2019-04-09 | 深圳奥比中光科技有限公司 | 基于深度相机的自动对焦方法及系统 |
| DE102019130963B3 (de) * | 2019-11-15 | 2020-09-17 | Sick Ag | Fokusmodul |
| WO2024138648A1 (zh) * | 2022-12-30 | 2024-07-04 | 深圳市大疆创新科技有限公司 | 拍摄设备的辅助对焦方法、装置、云台和跟焦装置 |
-
2021
- 2021-01-07 WO PCT/CN2021/070581 patent/WO2022147703A1/zh not_active Ceased
- 2021-01-07 CN CN202180079145.3A patent/CN116507970A/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105264436A (zh) * | 2013-04-05 | 2016-01-20 | 安德拉运动技术股份有限公司 | 用于控制与图像捕捉有关的设备的系统和方法 |
| US20170374354A1 (en) * | 2015-03-16 | 2017-12-28 | SZ DJI Technology Co., Ltd. | Apparatus and method for focal length adjustment and depth map determination |
| CN111226154A (zh) * | 2018-09-26 | 2020-06-02 | 深圳市大疆创新科技有限公司 | 自动对焦相机和系统 |
| CN109696667A (zh) * | 2018-12-19 | 2019-04-30 | 哈工大机器人(合肥)国际创新研究院 | 一种基于3DToF摄像头的门禁道闸检测方法 |
| CN110381261A (zh) * | 2019-08-29 | 2019-10-25 | 重庆紫光华山智安科技有限公司 | 聚焦方法、装置、计算机可读存储介质及电子设备 |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115623318A (zh) * | 2022-12-20 | 2023-01-17 | 荣耀终端有限公司 | 对焦方法及相关装置 |
| CN115623318B (zh) * | 2022-12-20 | 2024-04-19 | 荣耀终端有限公司 | 对焦方法及相关装置 |
| WO2025180224A1 (zh) * | 2024-02-29 | 2025-09-04 | 荣耀终端股份有限公司 | 自动对焦方法、电子设备、存储介质及程序产品 |
| WO2026040469A1 (zh) * | 2024-08-21 | 2026-02-26 | 荣耀终端股份有限公司 | 图像处理方法、终端设备、芯片系统及可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116507970A (zh) | 2023-07-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9521311B2 (en) | Quick automatic focusing method and image acquisition apparatus | |
| CN103988227B (zh) | 用于图像捕获目标锁定的方法和装置 | |
| CN107223330B (zh) | 一种深度信息获取方法、装置及图像采集设备 | |
| CN108076278B (zh) | 一种自动对焦方法、装置及电子设备 | |
| CN111147741A (zh) | 基于对焦处理的防抖方法和装置、电子设备、存储介质 | |
| US20050046706A1 (en) | Image data capture method and apparatus | |
| US20160191788A1 (en) | Image processing apparatus and image pickup apparatus | |
| WO2018201809A1 (zh) | 基于双摄像头的图像处理装置及方法 | |
| US20030002870A1 (en) | System for and method of auto focus indications | |
| WO2019037038A1 (zh) | 图像处理方法、装置及服务器 | |
| CN101410743A (zh) | 自动聚焦 | |
| US11190670B2 (en) | Method and a system for processing images based a tracked subject | |
| CN105611158A (zh) | 一种自动跟焦方法和装置、用户设备 | |
| CN116507970A (zh) | 跟焦方法、装置、拍摄设备及计算机可读存储介质 | |
| CN105847660A (zh) | 一种动态变焦方法、装置及智能终端 | |
| JP2015012482A (ja) | 画像処理装置及び画像処理方法 | |
| WO2018076529A1 (zh) | 场景深度计算方法、装置及终端 | |
| CN111246100B (zh) | 防抖参数的标定方法、装置和电子设备 | |
| JP2019062340A (ja) | 像振れ補正装置および制御方法 | |
| CN110602376A (zh) | 抓拍方法及装置、摄像机 | |
| CN111355891A (zh) | 基于ToF的微距对焦方法、微距拍摄方法及其拍摄装置 | |
| CN105745915B (zh) | 图像拍摄装置、方法和程序 | |
| JP6613149B2 (ja) | 像ブレ補正装置及びその制御方法、撮像装置、プログラム、記憶媒体 | |
| JP2024102803A (ja) | 撮像装置 | |
| CN114222059B (zh) | 拍照、拍照处理方法、系统、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21916761 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202180079145.3 Country of ref document: CN |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21916761 Country of ref document: EP Kind code of ref document: A1 |