WO2022040941A1 - 深度计算方法、装置、可移动平台及存储介质 - Google Patents

深度计算方法、装置、可移动平台及存储介质 Download PDF

Info

Publication number
WO2022040941A1
WO2022040941A1 PCT/CN2020/111156 CN2020111156W WO2022040941A1 WO 2022040941 A1 WO2022040941 A1 WO 2022040941A1 CN 2020111156 W CN2020111156 W CN 2020111156W WO 2022040941 A1 WO2022040941 A1 WO 2022040941A1
Authority
WO
WIPO (PCT)
Prior art keywords
depth
movable platform
depth map
captured image
tof ranging
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/111156
Other languages
English (en)
French (fr)
Inventor
丁晓飞
张鹏
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SZ DJI Technology Co Ltd
Original Assignee
SZ DJI Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by SZ DJI Technology Co Ltd filed Critical SZ DJI Technology Co Ltd
Priority to PCT/CN2020/111156 priority Critical patent/WO2022040941A1/zh
Publication of WO2022040941A1 publication Critical patent/WO2022040941A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis

Definitions

  • the present invention relates to the technical field of data processing, and in particular, to a depth computing method, a device, a removable platform and a storage medium.
  • the TOF ranging device may include a transmitting device and a receiving device, the transmitting device can transmit an optical signal, the optical signal is reflected by a target object in the environment, and the receiving device can receive the transmitted optical signal, and generate a depth image according to the received optical signal.
  • the traditional depth calculation method calculating the depth value corresponding to each pixel in the depth map according to the optical signal received by the receiving device will occupy a lot of computing resources and storage resources, resulting in high computational complexity, and the need for spatial memory. Greater demand.
  • the embodiments of the present application provide a depth calculation method, device, removable platform and storage medium, which can reduce the computing resources and storage resources occupied by the depth calculation, reduce the computational complexity, and reduce the demand for space memory smaller.
  • a first aspect of the embodiments of the present application provides a depth calculation method, which is applied to a movable platform, and the movable platform is configured with a photographing device and a Time of Flight (TOF) ranging device.
  • the distance device includes a transmitting device for transmitting an optical signal and a receiving device for receiving the optical signal reflected by the object, and the method includes:
  • the depth values of the pixels in the regional depth map are calculated according to the reflected optical signal received by the receiving device, and the depth map of regions other than the regional depth map in the depth map is not calculated based on the optical signal.
  • the depth value of the pixel is not calculated based on the optical signal.
  • a second aspect of an embodiment of the present application provides a depth computing device, including a memory and a processor, the depth computing device is applied to a movable platform, and the movable platform is configured with a photographing device and a TOF ranging device, the TOF
  • the ranging device includes transmitting means for transmitting an optical signal and receiving means for receiving said optical signal reflected by an object, wherein,
  • the memory for storing program codes
  • the processor calls the program code in the memory, and when the program code is executed, is used to perform the following operations:
  • the depth values of the pixels in the regional depth map are calculated according to the reflected optical signal received by the receiving device, and the depth map of regions other than the regional depth map in the depth map is not calculated based on the optical signal.
  • the depth value of the pixel is not calculated based on the optical signal.
  • a third aspect of the embodiments of the present application provides a movable platform, and the movable platform includes:
  • a photographing device installed on the body, for photographing the surrounding environment of the movable platform to obtain a photographed image
  • a TOF ranging device installed on the fuselage for acquiring a depth map, the TOF ranging device comprising a transmitting device for transmitting an optical signal and a receiving device for receiving the optical signal reflected by an object;
  • a fourth aspect of the embodiments of the present application provides a computer storage medium, where computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, are used to perform the depth calculation described in the first aspect method.
  • the movable platform may acquire a photographed image obtained by photographing the surrounding environment of the movable platform by the photographing device, and determine the area image of the target object in the photographed image in the surrounding environment. Then, the movable platform can obtain the relative pose parameters between the TOF ranging device and the photographing device, and determine the area of the target object on the depth map of the TOF ranging device according to the position of the area image in the captured image and the relative pose parameters Depth map, and then calculate the depth value of the pixel in the regional depth map according to the reflected optical signal received by the receiving device, and do not calculate the depth value of the pixel in the regional depth map other than the regional depth map in the depth map according to the optical signal.
  • the depth value of the pixels in the entire depth map is directly calculated according to the reflected light signal received by the receiving device.
  • the light signal calculation area depth value of the pixel in the depth map that is to say, the embodiment of the present application only calculates the depth value of the pixel in a part of the depth map, which can reduce the computing resources and storage resources occupied by the depth calculation and reduce the calculation Complexity, and the demand for space memory is small.
  • FIG. 1 is a schematic structural diagram of a TOF ranging device and a photographing device according to an embodiment of the present application
  • FIG. 2 is a schematic flowchart of a depth calculation method according to an embodiment of the present application.
  • FIG. 3 is a schematic flowchart of a depth calculation method according to another embodiment of the present invention.
  • FIG. 5 is a schematic structural diagram of a movable platform according to an embodiment of the present application.
  • the TOF ranging device may include a transmitting device and a receiving device.
  • the transmitting device can transmit an optical signal, and the optical signal is reflected by the target object in the environment.
  • the transmitting device may be a light emitting diode (Light Emitting Diode, LED for short) or a laser diode (Laser Diode, LD for short), wherein the transmitting device is driven by the driving module of the TOF ranging device, and the driving module is processed by the TOF ranging device.
  • the processing module controls the driving module to output the driving signal to drive the transmitting device, in which the frequency and duty cycle of the driving signal output by the driving module can be controlled by the processing module, the driving signal is used to drive the transmitting device, and the transmitting device sends out after modulation
  • the light signal hits the target object.
  • Target objects can be users, buildings, cars, and so on.
  • the receiving device may include a photodiode, an avalanche photodiode, and a charge-coupled element.
  • the receiving device may be a photodiode, avalanche photodiode, or a receiving array of charge-coupled elements.
  • the receiving device converts the optical signal into an electrical signal
  • the signal processing module of the TOF ranging device processes the electrical signal output by the receiving device, such as amplifying, filtering, etc.
  • the signal processed by the signal processing module is input into the processing module, and the processing module can The electrical signals are converted into depth images and grayscale images.
  • the captured image obtained by capturing the surrounding environment of the movable platform by the camera can be obtained, and the regional image of the target object in the captured image in the surrounding environment (that is, the sensory image) can be obtained.
  • the image corresponding to the region of interest (ROI)), and then the relative pose parameters between the TOF ranging device and the shooting device can be obtained, and the target object is determined according to the position and relative pose parameters of the regional image in the captured image.
  • ROI region of interest
  • the depth map can be divided into regions based on the regional depth map, and based on the corresponding importance of different image regions determined by the region division, the depth values of the pixels in the regional depth map can be calculated according to the reflected light signals received by the receiving device. , the depth values of pixels in the area depth map other than the area depth map in the depth map are not calculated according to the light signal. That is to say, after the depth map is divided into regions, only the depth values of the pixels in the regional depth map are calculated, which can reduce the computational resources and storage resources occupied by the depth calculation, reduce the computational complexity, and require more space and memory. little.
  • the movable platform may include a handheld stabilization system or an unmanned aerial vehicle for stabilization of the photographing device.
  • the relative pose parameters between the TOF ranging device and the photographing device are acquired in the storage device.
  • the relative pose of the TOF ranging device and the photographing device can be moved by external force, or the movement system of the movable platform itself changes, such as temperature changes or collision of the movable platform, etc., in order to improve the TOF ranging device and
  • the movable platform can regularly update the relative pose parameters of the TOF ranging device and the shooting device, or when the movable platform has depth calculation requirements, real-time calculation of the TOF ranging device and the relative pose parameters.
  • the relative pose parameters of the camera can regularly update the relative pose parameters of the TOF ranging device and the shooting device, or when the movable platform has depth calculation requirements, real-time calculation of the TOF ranging device and the relative pose parameters.
  • the TOF ranging device and the photographing device can be installed relatively movably.
  • the shooting angle of the shooting device can be changed according to the user's needs, and then the relative pose of the TOF ranging device and the shooting device will change.
  • the signal transmitting angle and/or the signal receiving angle of the TOF ranging device can be changed correspondingly according to the user's needs, so the relative pose of the TOF ranging device and the photographing device will change.
  • the movable platform can calculate the relative pose parameters of the TOF ranging device and the photographing device in real time, or the movable platform can detect the relative position of the TOF ranging device and the photographing device in real time.
  • the relative pose parameters of the TOF ranging device and the photographing device are calculated.
  • FIG. 2 is a schematic flowchart of a depth calculation method proposed by an embodiment of the present application. As shown in FIG. 2 , the method may include:
  • the movable platform can acquire a photographed image obtained by photographing the surrounding environment of the movable platform by the photographing device.
  • the surrounding environment of the movable platform refers to the environment in which the movable platform is located, that is, a three-dimensional environment in which a camera configured on the movable platform can capture images based on a field of view (FOV).
  • FOV field of view
  • the photographing device arranged at the bottom of the movable platform can take pictures below, in front of or behind the movable platform, and the photographing device arranged on the top of the movable platform can take pictures above, in front of or behind the movable platform.
  • the photographing device at the front end of the movable platform can photograph the front, left or right side of the movable platform, and so on.
  • a shooting instruction is generated, and the movable platform sends the shooting instruction to the shooting device, so that the shooting device can respond to the shooting instruction to shoot the surrounding environment of the movable platform to obtain a shot image; for another example, if the user sends voice information to the movable platform, the The voice information is used to instruct the surrounding environment of the movable platform to be photographed, the movable platform can generate a photographing instruction in response to the voice information, and the movable platform sends the photographing instruction to the photographing device, so that the photographing device can respond to the photographing instruction for the surrounding environment of the movable platform.
  • the environment is photographed to obtain a photographed image.
  • S202 Determine the area image of the target object in the captured image in the surrounding environment.
  • the target object may be selected by a user operating an interaction device displaying an image captured by the photographing device.
  • the interaction device may be configured in the movable platform, such as a display screen of the movable platform.
  • the captured image can be displayed in the interactive device, and the user can click or select a frame on the captured image displayed in the interactive device, and the movable platform detects the user's click or frame.
  • the target object may be determined based on the click or box selection operation, and then the region image of the target object in the captured image may be determined.
  • the interactive device can be configured in the control terminal, which can be a remote controller, a smart phone, or a ground station, and the interactive device is, for example, a display screen of the control terminal.
  • the control terminal displays the captured image in the interaction device, and the user can click or select a box on the captured image displayed in the interaction device.
  • the control terminal detects the user's click or frame selection operation, it can determine the target object based on the click or frame selection operation, and generate object indication information based on the target object, and the control terminal sends the object indication information to the movable platform.
  • the area image of the target object in the captured image may be determined based on the object indication information.
  • the object indication information may include information used to describe the target object, such as the type of the target object (such as a person, tree, car or bridge, etc.) and/or identity characteristics (such as color, size, clothing information or hairstyle, etc.), or the object
  • the indication information may be used to indicate the position of the target object in the captured image, for example, the target object is in the upper left corner of the captured image, or the target object is the second car at the top and from left to right in the captured image, and so on.
  • the control terminal may determine an area image of the target object in the captured image, and the control terminal may send the area image to the movable platform.
  • the target object may be a tracking photographing subject of the photographing device.
  • the user can operate the interactive device that displays the image captured by the camera.
  • the movable platform detects the user's operation, it can determine that the target object is the tracking shooting of the camera. object.
  • the movable platform can determine that the target object is the tracking object of the shooting device, and then the movable platform can obtain the position information of the tracking object, and then determine that the tracking object is in the tracking object according to the position information.
  • the location information can be the coordinate information of the tracking object in the surrounding environment.
  • the target object can be determined without user operation, which can improve the efficiency of determining the target object. sex.
  • the manner in which the movable platform determines the area image of the target object in the captured image in the surrounding environment may include one or more of the following:
  • the movable platform can determine the area image of the target object in the captured image in the surrounding environment according to the image tracking algorithm.
  • the movable platform can run an image tracking algorithm to determine the area of the target object in the captured image.
  • the image tracking algorithm may be a KLT tracking algorithm.
  • an image area most similar to the area image of the target object in the historical captured image may be determined in the acquired captured image, and the most similar area image may be determined as the area in the surrounding environment.
  • the historical photographing picture may be a photographing picture of a previous frame of the acquired photographing image.
  • the target object can be selected by the user.
  • the movable platform can input the captured image into the preset neural network model to determine the area image of the target object in the captured image in the surrounding environment.
  • the preset neural network model can be used to identify a specific object, exemplarily, the specific object can be a person, a building, a car, or the like.
  • the movable platform can input the captured image into the preset neural network model, and the movable platform can use the neural network model to identify whether there is a human body (such as a human face) in the captured image. , hands, feet, back, etc.), if there is a human body in the captured image, the movable platform can determine the area image of the human body in the captured image.
  • the movable platform can acquire attention indication information, wherein the attention indication information is used to indicate the user's attention area to the captured image. Then, the movable platform can determine the area image of the target object in the captured image in the surrounding environment according to the attention indication information.
  • the movable platform may acquire motion data for indicating the movement of the control terminal, and generate attention indication information according to the motion data; or, the movable platform may acquire motion data for instructing the rotation of the eyeball of the user wearing the control terminal information, and generate attention indication information according to the information; or, obtain a composition instruction determined according to the user's instruction, and generate attention indication information according to the composition instruction.
  • the motion data indicating the motion of the control terminal may be the head rotation data of the user wearing the control terminal, etc.
  • the control terminal may be video glasses. By wearing the video glasses, the video glasses can pass the built-in
  • the motion sensor detects the head rotation of the user to generate the head rotation data
  • the video glasses can send the head rotation data to the movable platform, and based on the motion data, attention indication information can be generated, so that the The image within the region of interest is taken as the region image.
  • the control terminal can detect the information of the eyeball rotation of the user wearing the control terminal, and send the eyeball rotation information to the movable platform, and the movable platform can obtain the information of the user wearing the control terminal.
  • the wearing control terminal may be, for example, video glasses or the like.
  • the movable platform may also acquire a composition instruction of the user, so that the target area image may be determined from the current image to be encoded according to the composition instruction, and the composition instruction of the user may be, for example, the user
  • the selection instruction issued by the control terminal the image area acted by the selection instruction can be used as an area image.
  • the manner of determining the area image of the target object in the captured image in this embodiment of the present application includes, but is not limited to, the above manner.
  • the movable platform may determine the target object in the surrounding environment according to preset composition rules. Take an image of the area in the image.
  • the preset composition rule may be, for example, determining the target area image from the current image according to the preset composition rule; wherein, the preset composition rule may be, for example, based on the composition, the central area of the image is used as the target. area image.
  • the relative pose parameter between the TOF ranging device and the photographing device is used to indicate the relative position and relative attitude between the TOF ranging device and the photographing device, for example, the relative pose parameter between the TOF ranging device and the photographing device
  • a rotational relationship and/or translational relationship between the TOF ranging device and the camera may be included.
  • the relative pose parameters may be pre-stored on a local storage device of the movable platform, and the relative pose parameters may be acquired from the local storage device.
  • This embodiment of the present application does not limit the sequence of execution of steps S201 to S203.
  • the movable platform detects that the photographing device has photographing behavior, the relative pose parameters between the TOF ranging device and the photographing device can be obtained.
  • the movable platform can determine the area image of the target object in the photographed image in the surrounding environment after acquiring the photographed image obtained by the photographing device of the surrounding environment of the movable platform, and obtain the distance between the TOF ranging device and the photographing device. relative pose parameters.
  • the movable platform can obtain the photographed image obtained by photographing the surrounding environment of the movable platform by the photographing device, and after determining the area image of the target object in the photographed image in the surrounding environment, obtain the TOF ranging device and the photographing device.
  • Relative pose parameters In another example, the movable platform can obtain the relative pose parameters between the TOF ranging device and the photographing device after acquiring the photographed image obtained by the photographing device of the surrounding environment of the movable platform, and then determine that the target object in the surrounding environment is in the photographed image. image of the area in .
  • the regional image can be projected and transformed into the TOF ranging device according to the position of the regional image in the captured image and the relative pose parameters on the depth map of the device to obtain the region depth map corresponding to the region image.
  • the image area and the area depth image are used to indicate the same target object in space.
  • S205 Calculate the depth value of the pixel in the regional depth map according to the reflected optical signal received by the receiving device, and do not calculate the depth value of the pixel in the regional depth map other than the regional depth map in the depth map according to the optical signal.
  • the receiving device receives the reflected optical signal, and calculates the depth value corresponding to each pixel in the depth map according to the received optical signal.
  • calculating the depth values of all pixels on the depth map of the TOF ranging device according to the received optical signals requires a lot of computing resources and storage resources.
  • the movable platform only needs to focus on the depth values corresponding to the pixels in the area depth image. Therefore, in the embodiments of the present application, the depth values of the pixels in the regional depth map are calculated according to the reflected optical signals received by the receiving device, and the depth values of the pixels in the depth map other than the regional depth map are not calculated according to the optical signals.
  • the depth value that is, the depth value of the pixel in the regional depth map is only calculated according to the reflected light signal received by the receiving device, and the depth value of the pixel in the regional depth map other than the regional depth map in the depth map is not calculated.
  • the movable platform only calculates the depth value of the pixels in the regional depth map according to the reflected optical signal received by the receiving device, and does not calculate the depth value of the pixels in the depth map other than the regional depth map according to the optical signal.
  • the depth value of the pixel can reduce the computing resources and storage resources occupied by the depth calculation, reduce the computing complexity, and reduce the demand for space memory.
  • FIG. 3 is a schematic flowchart of a depth calculation method proposed by another embodiment of the present invention. As shown in FIG. 3 , the method may include:
  • S302 Determine the area image of the target object in the captured image in the surrounding environment.
  • step S302 in this embodiment of the present application reference may be made to the specific description of step S202 in the foregoing embodiment, which is not repeated in this embodiment of the present application.
  • S304 Determine a regional depth map of the target object on the depth map of the TOF ranging device according to the position of the regional image in the captured image and the relative pose parameters.
  • the working state of the movable platform may include a first working state and a second working state.
  • the first working state includes at least one of the following: the photographing device is focusing or following the target object, the movable platform is in a face recognition state, a visual positioning state or a background blur state, the The movable platform is determining the position of the target object, or the photographing device is tracking and photographing the target object.
  • the second working state includes at least one of the following: the movable platform is in an obstacle avoidance mode, the movable platform is in a motion state, or the movable platform is manually controlled by a user through a control terminal.
  • the first working state includes, but is not limited to, the above-mentioned scene. As long as it is a scene where only depth values of some pixels of the depth map need to be acquired, the first working state can be set.
  • the second working state includes but is not limited to the above scenarios. As long as it is a scene where the depth value of each pixel of the depth map needs to be acquired, it can be set to the second working state, which is not specifically limited by the embodiments of the present application.
  • the movable platform controls the camera to focus, or in scenes such as face recognition, it is not necessary to obtain the depth information of the entire depth map, but only the depth information of the regional depth map. Based on this, in order to avoid obtaining the depth information of the entire depth map This leads to waste of computing resources and storage resources.
  • the depth value of the pixels in the regional depth map can be calculated only according to the reflected light signal received by the receiving device, Not calculating the depth value of the pixels in the depth map other than the regional depth map based on the optical signal can reduce the computing resources and storage resources occupied by the depth solution, reduce the computational complexity, and require more space and memory. little.
  • the movable platform can directly calculate the depth value of each pixel in the depth map according to the reflected light signal received by the receiving device. For example, when the movable platform is in the obstacle avoidance mode, the movable platform is in motion, or the movable platform is manually controlled by the user through the control terminal, the movable platform can determine that the current working state is the second state, and then receive according to the receiving device. The obtained reflected light signal calculates the depth value of each pixel in the depth map.
  • the movable platform can perform depth calculation on the pixels in the depth map by using different processing methods according to different working states. For example, when the current working state of the movable platform is the first working state, only the area depth is calculated. The depth value of the pixel in the figure, when the current working state of the movable platform is the second working state, the depth value of each pixel in the depth map is calculated, and different depth calculation methods can be selected based on different working states to improve the depth calculation. flexibility.
  • a relative pose parameter acquisition method includes:
  • S401 Acquire a depth map output by the TOF ranging device at a historical moment, wherein each pixel in the depth map output at the historical moment is calculated according to the received reflected light signal.
  • S402 Determine multiple depth transition points in the depth map output at the historical moment.
  • the movable platform may traverse the point clouds corresponding to the depth map output at historical moments, determine multiple target point clouds, and determine the multiple target point clouds as multiple depth transition points. Wherein, the depth value of the target point cloud and the depth values of multiple adjacent point clouds are greater than or equal to a preset depth threshold.
  • the positional relationship between the target point cloud and at least one adjacent point cloud adjacent to the target point cloud can be the target point cloud and the points in its eight fields, namely the nine-square grid, and the eight points other than the target point cloud are called eight points field.
  • the positional relationship between the target point cloud and at least one adjacent point cloud adjacent to the target point cloud may be the target point cloud and the points in its three fields, that is, the four-square grid, and the three points except the target point cloud.
  • the dots are called three domains.
  • the present invention does not specifically limit the "adjacent" relationship between the target point cloud and the adjacent point cloud, it can be set in advance, or in different application scenarios, different target point clouds and adjacent point clouds can be set. "adjacent" relationship.
  • the depth value of a point P1 in the point cloud is 3m
  • the point to the left of point P1 is P2
  • the depth value of P2 is 2.9m
  • the point to the right of point P1 is P3
  • the depth of point P3 If the value is 1.0m, then it can be determined that the P1 point is the point where the depth jump occurs, that is, the depth jump point.
  • the rotation and translation between the TOF ranging device and the camera can be written as:
  • R refers to the rotation between the TOF ranging device and the photographing device
  • t refers to the translation between the TOF ranging device and the photographing device.
  • the rotation R from the coordinate system of the TOF ranging device to the coordinate system of the photographing device is respectively represented; and the position t of the origin of the coordinates of the TOF ranging device in the coordinate system of the photographing device.
  • the original translation and original rotation between the TOF ranging device and the photographing device can be the translation and rotation between the TOF ranging device and the photographing device calibrated by the movable platform before leaving the factory, or it can be the original translation and rotation between the TOF ranging device and the photographing device that the movable platform is calibrated before leaving the factory. Panning and rotation between the TOF rangefinder and the camera after preset. Wherein, the preset translation and rotation between the TOF ranging device and the photographing device may be set according to experience.
  • the movable platform may run an edge detection algorithm on the captured image to obtain edge response values of pixels in the captured image.
  • the edge detection algorithm may include: Sobel edge detection algorithm, Isotropic Sobel edge detection algorithm, Roberts edge detection algorithm, Prewitt edge detection algorithm, Laplacian edge detection algorithm, Canny edge detection algorithm, and the like.
  • the edge detection algorithm as the Canny edge detection algorithm as an example, the movable platform detects the edge information in the photographed image collected by the photographing device, and can obtain the edge information of the pixel points corresponding to each of the depth transition points in the photographed image.
  • the edge response values of the pixels of the plurality of depth transition points in the captured image may be acquired for the movable platform.
  • the movable platform can determine the pixel points of the multiple depth jump points in the captured image according to the multiple depth jump points, the original translation and the original rotation, and obtain the pixel points of the multiple depth jump points in the captured image.
  • the edge response value of the pixel points taking the minimum sum of the edge response values of multiple depth transition points in the captured image as the optimization goal, taking the original translation and original rotation as the optimization objects to perform optimization operations, and optimize the original translation obtained by optimization. and the original rotation is determined as the rotation and translation between the TOF ranging device and the photographing device.
  • the movable platform projects the set P of multiple depth jump points on the photographed image, and the The pixel coordinates of the pixel points corresponding to each of the depth jump points are marked as p, where p is a set of pixel points, and p may include p1, p2, p3 . . . Among them, this projection process is recorded as ⁇ , then there are:
  • the movable platform sequentially finds the pixel point p1 corresponding to P1, the pixel point p2 corresponding to P2, the pixel point p3 corresponding to P3..., and after obtaining multiple groups (P, p), the pixels at each point p are
  • the sum of the edge response values E(p) is input as the Cost Function (cost function), and the optimization solution obtains the precise rotation R and displacement relationship t.
  • arg represents the optimized parameter is rotation R, displacement t.
  • the movable platform obtains the depth map output by the TOF ranging device at the historical moment, determines a plurality of depth transition points in the depth map output at the historical moment, and determines the edge response value of the pixel point in the captured image
  • the original translation and original rotation are optimized according to the position, original translation, original rotation and edge response values of multiple depth jump points to determine the rotation and translation between the TOF ranging device and the photographing device, which can improve the TOF ranging
  • FIG. 5 is a structural diagram of the depth computing device applied to a mobile platform provided by an embodiment of the present application.
  • the depth computing device 500 on the movable platform includes a memory 501 and a processor 502, the movable platform is configured with a photographing device and a time-of-flight TOF ranging device, and the TOF ranging device includes a transmitting device for transmitting optical signals and a device for receiving The device for receiving the optical signal reflected by the object, wherein the memory 502 stores program codes, the processor 502 calls the program codes in the memory, and when the program codes are executed, the processor 502 performs the following operations:
  • the depth values of the pixels in the regional depth map are calculated according to the reflected optical signal received by the receiving device, and the depth map of regions other than the regional depth map in the depth map is not calculated based on the optical signal.
  • the depth value of the pixel is not calculated based on the optical signal.
  • the relative pose parameters between the TOF ranging device and the photographing device are stored in a local storage device of the movable platform.
  • the TOF ranging device and the photographing device are installed relatively movably.
  • processor 502 is further configured to perform the following operations:
  • the processor 502 calculates the depth value of the pixel in the regional depth map according to the reflected optical signal received by the receiving device, and does not calculate the difference between the regional depth map in the depth map according to the optical signal.
  • the following steps are specifically performed:
  • the depth value of the pixel in the depth map is calculated according to the reflected light signal received by the receiving device, and the depth map is not calculated according to the light signal. Depth values of pixels in the region depth map other than the region depth map.
  • the first working state includes at least one of the following: the photographing device is focusing or following the target object, the movable platform is determining the position of the target object, or The photographing device is tracking and photographing the target object.
  • processor 502 is further configured to perform the following operations:
  • the depth value of each pixel in the depth map is calculated according to the reflected light signal received by the receiving device.
  • the second working state includes at least one of the following: the movable platform is in an obstacle avoidance mode, the movable platform is in a motion state, or the movable platform is manually controlled by a user through a control terminal .
  • the target object is selected by a user operating an interaction device displaying an image captured by the capturing device.
  • the target object is a tracking shooting object of the shooting device.
  • the processor 502 when determining the area image of the target object in the captured image in the surrounding environment, the processor 502 specifically performs the following operations:
  • the captured image is input into a preset neural network model to determine a region image of the target object in the captured image in the surrounding environment.
  • the processor 502 when determining the area image of the target object in the captured image in the surrounding environment, the processor 502 specifically performs the following operations:
  • the attention indication information is used to indicate the user's attention area to the captured image
  • the area image of the target object in the captured image in the surrounding environment is determined.
  • the processor 502 when acquiring the relative pose parameters between the TOF ranging device and the photographing device, the processor 502 specifically performs the following operations:
  • an optimization operation is performed on the original translation and the original rotation, so as to determine the TOF ranging device and the photographing device between rotation and translation.
  • the processor 502 performs an optimization operation on the original translation and the original rotation according to the positions of the plurality of depth transition points, the original translation, the original rotation and the edge response value , when determining the rotation and translation between the TOF ranging device and the photographing device, specifically perform the following operations:
  • the processor 502 when determining the edge response value of a pixel in the captured image, the processor 502 specifically performs the following operations:
  • An edge detection algorithm is run on the captured image to obtain edge response values of pixels in the captured image.
  • the processor 502 specifically performs the following operations when determining multiple depth jump points in the depth map output at the historical moment:
  • the plurality of target point clouds are determined as the plurality of depth jump points.
  • the relative pose parameter between the TOF ranging device and the photographing device includes a rotational relationship and/or a translational relationship between the TOF ranging device and the photographing device.
  • the TOF ranging device comprises a 3D-TOF camera or lidar or millimeter wave radar.
  • the depth calculation apparatus applied to the movable platform provided by this embodiment can execute the depth calculation methods shown in FIG. 2 to FIG. 4 provided by the foregoing embodiments, and the execution manner and beneficial effects are similar, and details are not repeated here.
  • the movable platform When the movable platform is an unmanned aerial vehicle, its power system may include a rotor, a motor that drives the rotor to rotate, and its electric regulator.
  • the UAV can be a quad-rotor, hexa-rotor, octa-rotor or other multi-rotor UAV, and the UAV takes off and lands vertically at this time. It is understood that the UAV may also be a fixed-wing movable platform or a hybrid-wing movable platform.
  • the TOF ranging device and the photographing device are installed relatively movably.
  • the movable platform further includes a first communication device, the first communication device is installed on the fuselage, and the first communication device is used for data interaction with the control terminal.
  • the movable platform further includes a second communication device, the second communication device is installed on the body, and the second communication device is used to communicate with the device displaying the image captured by the camera.
  • the interaction device performs data interaction.
  • the movable platform includes at least one of the following: a handheld stabilization system or an unmanned aerial vehicle for stabilization of the photographing device.
  • Embodiments of the present application further provide a computer storage medium, where computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by the processor, are used to execute the depth calculation method shown in FIG. 2 to FIG. 4 . .

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Studio Devices (AREA)

Abstract

一种深度计算方法、装置、可移动平台及存储介质,该方法应用于可移动平台,该方法包括:获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像(S201);确定周围环境中目标对象在拍摄图像中的区域图像(S202);获取TOF测距装置和拍摄装置之间的相对位姿参数(S203);根据区域图像在拍摄图像中的位置、相对位姿参数确定目标对象在TOF测距装置的深度图上的区域深度图(S204);根据TOF测距装置所包含的接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值(S205),该方法可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。

Description

深度计算方法、装置、可移动平台及存储介质 技术领域
本发明涉及数据处理技术领域,尤其涉及深度计算方法、装置、可移动平台及存储介质。
背景技术
TOF测距装置可以包括发射装置和接收装置,发射装置可以发射光信号,光信号经过环境中的目标对象反射,接收装置可以接收发射的光信号,根据接收到的光信号生成深度图像。然而,在传统的深度计算方法中,根据接收装置接收到的光信号计算深度图中每一个像素对应的深度值会占用大量的计算资源和存储资源,导致计算复杂度高,且对空间内存的需求较大。
发明内容
有鉴于此,本申请实施例提供了一种深度计算方法、装置、可移动平台及存储介质,可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。
本申请实施例第一方面提供了一种深度计算方法,该方法应用于可移动平台,所述可移动平台配置有拍摄装置和飞行时间(Time of flight,TOF)测距装置,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置,所述方法包括:
获取所述拍摄装置对所述可移动平台周围环境进行拍摄得到的拍摄图像;
确定所述周围环境中目标对象在所述拍摄图像中的区域图像;
获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数;
根据所述区域图像在所述拍摄图像中的位置、所述相对位姿参数确定所述目标对象在所述TOF测距装置的深度图上的区域深度图;
根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域 深度图中的像素的深度值。
本申请实施例第二方面提供了一种深度计算装置,包括存储器和处理器,所述深度计算装置应用于可移动平台,所述可移动平台配置有拍摄装置和TOF测距装置,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置,其中,
所述存储器,用于存储有程序代码;
所述处理器,调用存储器中的程序代码,当程序代码被执行时,用于执行如下操作:
获取所述拍摄装置对所述可移动平台周围环境进行拍摄得到的拍摄图像;
确定所述周围环境中目标对象在所述拍摄图像中的区域图像;
获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数;
根据所述区域图像在所述拍摄图像中的位置、所述相对位姿参数确定所述目标对象在所述TOF测距装置的深度图上的区域深度图;
根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
本申请实施例第三方面提供了一种可移动平台,该可移动平台包括:
机身;
拍摄装置,安装在所述机身,用于对所述可移动平台周围环境进行拍摄得到拍摄图像;
TOF测距装置,安装在所述机身,用于获取深度图,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置;
以及如第二方面中任一项所述的深度计算装置。
本申请实施例第四方面提供了一种计算机存储介质,所述计算机存储介质中存储有计算机程序指令,所述计算机程序指令被处理器执行时,用于执行如第一方面所述的深度计算方法。
在本申请实施例中,可移动平台可获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像,并确定周围环境中目标对象在拍摄图像中的区域图像。 然后,可移动平台可以获取TOF测距装置和拍摄装置之间的相对位姿参数,根据区域图像在拍摄图像中的位置、相对位姿参数确定目标对象在TOF测距装置的深度图上的区域深度图,进而根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值。相对传统的深度计算方法在获取TOF测距装置的深度图之后,直接根据接收装置接收到的反射的光信号计算整个深度图中像素的深度值,本申请实施例仅根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,也就是说,本申请实施例仅计算深度图中部分区域的像素的深度值,可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。
附图说明
为了更清楚地说明本申请实施例或现有技术中的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是本申请实施例的一种TOF测距装置和拍摄装置的结构示意图;
图2是本申请实施例的一种深度计算方法的流程示意图;
图3是本发明另一实施例的一种深度计算方法的流程示意图;
图4是本申请实施例的一种相对位姿参数获取方法的流程示意图;
图5是本申请实施例的一种可移动平台的结构示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
除非另有定义,本文所使用的所有的技术和科学术语与属于本发明的技术领域的技术人员通常理解的含义相同。本文中在本发明的说明书中所使用的术 语只是为了描述具体的实施例的目的,不是旨在于限制本发明。本文所使用的术语“及/或”包括一个或多个相关的所列项目的任意的和所有的组合。
下面结合附图,对本发明的一些实施方式作详细说明。在不冲突的情况下,下述的实施例及实施例中的特征可以相互组合。
TOF测距装置可以包括发射装置和接收装置,发射装置可以发射光信号,光信号经过环境中的目标对象反射,接收装置可以接收发射的光信号,根据接收到的光信号生成深度图像和与深度图像对应的灰度图像。其中发射装置可以为发光二极管(Light Emitting Diode,简称LED)或激光二极管(Laser Diode,简称LD),其中,发射装置由TOF测距装置的驱动模块来驱动,驱动模块由TOF测距装置的处理模块控制,处理模块控制驱动模块输出驱动信号来驱动发射装置,其中驱动模块输出的驱动信号的频率、占空比等都可以由处理模块控制,利用驱动信号驱动发射装置,发射装置发出经过调制以后的光信号,光信号打到目标对象上。目标对象可以是用户、建筑、汽车等等。其中,接收装置可以包括光电二极管、雪崩光电二极管、电荷耦合元件。所述接收装置可以是一种光电二极管、雪崩光电二极管、电荷耦合元件的接收阵列。接收装置将光信号转化成电信号,TOF测距装置信号处理模块对接收装置输出的电信号进行处理,例如放大、滤波等,经过信号处理模块处理的信号输入到处理模块中,处理模块可以将电信号转换成深度图像和灰度图像。
计算TOF测距装置的深度图上的所有像素的深度值,需要占用大量的计算资源和存储资源,导致计算复杂度高,且对空间内存的需求较大。所以,为了满足深度图上区域深度图的像素的深度计算需求,可获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像,确定周围环境中目标对象在拍摄图像中的区域图像(即感兴趣区域(region of interest,ROI)对应的图像),然后可以获取TOF测距装置和拍摄装置之间的相对位姿参数,根据区域图像在拍摄图像中的位置、相对位姿参数确定目标对象在TOF测距装置的深度图上的区域深度图。进一步的,可以基于区域深度图对该深度图进行区域划分,基于该区域划分确定的不同图像区域对应的重要性,可根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值。也就是说,在对深度图进行区域划 分之后,仅计算区域深度图中像素的深度值,可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。
其中,可移动平台可以包括用于对拍摄装置增稳的手持增稳系统或者无人机等。
TOF测距装置可以包括3D-TOF相机。拍摄装置可以为可见光相机或者红外相机等,拍摄装置用于对可移动平台周围环境进行拍摄得到拍摄图像。
其中,TOF测距装置和拍摄装置配置在可移动平台中,图1为本申请实施例提供的一种示例性的TOF测距装置和拍摄装置的结构示意图。在图1中,TOF测距装置和拍摄装置相对固定安装,例如,TOF测距装置和拍摄装置可以固定在同一结构件上,确保TOF测距装置和拍摄装置的相对位姿参数保持不变。在该实施例中,TOF测距装置和拍摄装置之间的相对位姿参数可以存储在可移动平台的本地存储装置中,那么可移动平台存在深度计算需求时,可直接从可移动平台的本地存储装置中获取TOF测距装置和拍摄装置之间的相对位姿参数。可选的,TOF测距装置和拍摄装置的相对位姿可通过外力移动,或者可移动平台自身的运动系统移动发生变化,例如温度变化或者可移动平台发生碰撞等,为了提高TOF测距装置和拍摄装置的相对位姿参数的准确度,可移动平台可定期对TOF测距装置和拍摄装置的相对位姿参数进行更新,或者在可移动平台存在深度计算需求时,实时计算TOF测距装置和拍摄装置的相对位姿参数。
在另一种实现方式中,TOF测距装置和拍摄装置可以相对活动安装。例如,拍摄装置的拍摄角度可以根据用户需求发生相应变化,那么TOF测距装置和拍摄装置的相对位姿就会发生变化。又如,TOF测距装置的信号发射角度和/或信号接收角度可以根据用户需求发生相应变化,那么TOF测距装置和拍摄装置的相对位姿就会发生变化。基于此,在可移动平台存在深度计算需求时,可移动平台可以实时计算TOF测距装置和拍摄装置的相对位姿参数,或者可移动平台可以在检测到TOF测距装置和拍摄装置的相对位姿发生变化时,计算TOF测距装置和拍摄装置的相对位姿参数。
请参见图2,是本申请实施例提出的一种深度计算方法的流程示意图,如 图2所示,该方法可包括:
S201,获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像。
可移动平台可以获取拍摄装置对该可移动平台周围环境进行拍摄得到的拍摄图像。可移动平台周围环境指的是可移动平台所处环境,即配置在可移动平台上的拍摄装置基于视场角(Field of view,FOV)可拍摄得到图像的三维环境。例如,配置在可移动平台底部的拍摄装置可以对可移动平台的下方、前方或者后方进行拍摄,配置在可移动平台顶部的拍摄装置可以对可移动平台的上方、前方或者后方进行拍摄,配置在可移动平台前端的拍摄装置可以对可移动平台的前方、左方或者右方进行拍摄,等等。
在一示例性场景中,用户想要对可移动平台周围环境进行拍摄时,可以通过对可移动平台进行操作的方式或者对可移动平台的控制终端(例如遥控器、智能手机、平板电脑、膝上型电脑、台式电脑中的一种或多种)向可移动平台提交拍摄指令,可移动平台将该拍摄指令发送给拍摄装置,以便拍摄装置响应该拍摄指令对可移动平台周围环境进行拍摄得到拍摄图像。例如用户点击可移动平台中具有拍摄功能的虚拟按键或者物理按键或者控制终端中具有拍摄功能的虚拟按键或者物理按键,可移动平台或者控制终端检测到用户对具有拍摄功能的虚拟按键或者物理按键的点击操作时生成拍摄指令,可移动平台将该拍摄指令发送给拍摄装置,以便拍摄装置响应该拍摄指令对可移动平台周围环境进行拍摄得到拍摄图像;又如用户向可移动平台发送语音信息,该语音信息用于指示对可移动平台周围环境进行拍摄,可移动平台可以响应该语音信息生成拍摄指令,可移动平台将该拍摄指令发送给拍摄装置,以便拍摄装置响应该拍摄指令对可移动平台周围环境进行拍摄得到拍摄图像。
在另一示例性场景中,可移动平台确定满足拍摄条件时,可以生成拍摄指令,可移动平台将该拍摄指令发送给拍摄装置,以便拍摄装置响应该拍摄指令对可移动平台周围环境进行拍摄得到拍摄图像。例如,在可移动平台进行目标跟踪(例如拍摄装置在对目标对象进行追踪拍摄)时,可移动平台可以确定满足拍摄条件。又如,在可移动平台确定目标对象的位置(例如视觉定位)时,可移动平台可以确定满足拍摄条件。
S202,确定周围环境中目标对象在拍摄图像中的区域图像。
可移动平台获取到拍摄图像之后,可以确定周围环境中目标对象在拍摄图 像中的区域图像。
在一个实施例中,目标对象可以是由用户对显示有拍摄装置拍摄的图像的交互装置进行操作而选中的。
例如,假设可移动平台为手持相机或者用于对拍摄装置增稳的手持增稳系统,那么交互装置可以配置在可移动平台中,交互装置例如为可移动平台的显示屏幕。基于此,可移动平台获取到拍摄图像之后,可以将该拍摄图像显示在交互装置中,用户可以对交互装置中显示的拍摄图像进行点击或者框选操作,可移动平台检测到用户的点击或者框选操作之后,可以基于该点击或者框选操作确定目标对象,进而确定该目标对象在拍摄图像中的区域图像。
又如,假设可移动平台为无人机,那么交互装置可以配置在控制终端中,控制终端可以为遥控器、智能手机或者地面站等,交互装置例如为控制终端的显示屏幕。基于此,可移动平台获取到拍摄图像之后,可以将该拍摄图像发送给控制终端,控制终端将该拍摄图像显示在交互装置中,用户可以对交互装置中显示的拍摄图像进行点击或者框选操作,控制终端检测到用户的点击或者框选操作之后,可以基于该点击或者框选操作确定目标对象,并基于目标对象生成对象指示信息,控制终端将对象指示信息发送给可移动平台,可移动平台可以基于对象指示信息确定目标对象在拍摄图像中的区域图像。其中对象指示信息可以包括用于描述目标对象的信息,例如目标对象的类型(例如人、树、汽车或者桥等)和/或身份特征(例如颜色、尺寸、服装信息或者发型等),或者对象指示信息可以用于指示目标对象在拍摄图像中的位置,例如目标对象在拍摄图像中的左上角,或者目标对象为拍摄图像中最上方且从左往右的第二辆汽车,等等。示例性的,控制终端基于该点击或者框选操作确定目标对象之后,可以确定该目标对象在拍摄图像中的区域图像,控制终端可以将该区域图像发送给可移动平台。
在一个实施例中,目标对象可以为拍摄装置的追踪拍摄对象。例如,用户想要获取跟踪拍摄对象的深度信息,那么用户可以对显示有拍摄装置拍摄的图像的交互装置进行操作,可移动平台检测到用户的操作之后,可以确定目标对象为拍摄装置的跟踪拍摄对象。又如,在可移动平台进行目标跟踪时,可移动平台可以确定目标对象为拍摄装置的跟踪拍摄对象,然后可移动平台可以获取跟踪拍摄对象的位置信息,进而根据该位置信息确定跟踪拍摄对象在拍摄图像 中的区域图像,举例来说,位置信息可以为跟踪拍摄对象在周围环境中的坐标信息,本实施例在无需用户操作的情况下,就可以确定目标对象,可提升确定目标对象的有效性。
在一个实施例中,可移动平台确定周围环境中目标对象在拍摄图像中的区域图像的方式可以包括如下一种或多种:
一、可移动平台可以根据图像跟踪算法确定周围环境中目标对象在拍摄图像中的区域图像。
可移动平台可以运行图像跟踪算法来确定,目标对象在拍摄图像中的区域图像。其中,所述图像跟踪算法可以为KLT跟踪算法。具体地,在可移动平台获取到拍摄图像之后,可在获取的拍摄图像中确定与历史拍摄图像中目标对象的区域图像最相似的图像区域,可以将该最相似的区域图像确定为周围环境中目标对象在拍摄图像中的区域图像。所述历史拍摄画面可以是所述获取的拍摄图像的上一帧拍摄画面。目标对象可以是由用户选中的。
二、可移动平台可以将拍摄图像输入到预设的神经网络模型以确定周围环境中目标对象在拍摄图像中的区域图像。
其中,预设的神经网络模型可以用于识别特定对象,示例性的,特定对象可以为人、建筑或者汽车等。假设预设的神经网络模型用于识别人,那么可移动平台可以将拍摄图像输入到预设的神经网络模型,可移动平台可以通过该神经网络模型识别该拍摄图像中是否存在人体(例如人脸、手足或者背部等),若该拍摄图像中存在人体,那么可移动平台可以确定人体在拍摄图像中的区域图像。
三、可移动平台可以获取关注指示信息,其中关注指示信息用于指示用户对拍摄图像的关注区域。然后,可移动平台可以根据该关注指示信息,确定周围环境中目标对象在拍摄图像中的区域图像。
其中,可移动平台可以获取用于指示所述控制终端的运动的运动数据,根据所述运动数据生成关注指示信息;或者,可移动平台可以获取用于指示佩戴控制终端的用户的眼球的转动的信息,根据所述信息生成关注指示信息;或者,获取根据用户的指令确定的构图指令,根据所述构图指令生成关注指示信息。
所述指示所述控制终端的运动的运动数据可以佩戴控制终端的用户的头 部转动数据等,例如,所述控制终端可以为视频眼镜,用户通过佩戴视频眼镜,使得所述视频眼镜可以通过内置的运动传感器检测用户的头部转动以生成所述头部转动数据,视频眼镜可以向所述可移动平台发送所述头部转动数据,基于所述运动数据可生成关注指示信息,从而可将该关注区域内的图像作为区域图像。或者,控制终端可以检测佩戴所述控制终端的用户的眼球的转动的信息,并将所述眼球的转动的信息发送给可移动平台,所述可移动平台可获取佩戴所述控制终端的用户的眼球的转动的信息,并根据所述信息生成关注指示信息,所述佩戴的控制终端例如可以是视频眼镜等。又或者,所述可移动平台还可获取用户的构图指令,从而可根据所述构图指令从所述当前待编码图像中确定所述目标区域图像,所述用户的构图指令例如可以是所述用户通过所述控制终端发出的选择指令,从而可将所述选择指令作用的图像区域作为区域图像。
在一个实施例中,本申请实施例确定周围环境中目标对象在拍摄图像中的区域图像的方式包含但不仅限于上述方式,例如可移动平台可以按照预设的构图规则确定周围环境中目标对象在拍摄图像中的区域图像。所述预设的构图规则例如可以是按照预设的构图规则从所述当前图像中确定所述目标区域图像;其中,所述预设的构图规则例如可以是基于构图将图像的中心区域作为目标区域图像。
S203,获取TOF测距装置和拍摄装置之间的相对位姿参数。
其中,TOF测距装置和拍摄装置之间的相对位姿参数用于指示TOF测距装置和拍摄装置之间的相对位置和相对姿态,例如TOF测距装置和拍摄装置之间的相对位姿参数可以包括TOF测距装置和拍摄装置之间的旋转关系和/或平移关系。所述相对位姿参数可以预存在可移动平台的本地存储装置上,从所述本地存储装置中获取所述相对位姿参数。
本申请实施例并不限定步骤S201至步骤S203的先后执行顺序,例如可移动平台只要检测到拍摄装置存在拍摄行为,那么就可以获取TOF测距装置和拍摄装置之间的相对位姿参数。又如可移动平台可以在获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像之后,同时确定周围环境中目标对象在拍摄图像中的区域图像,并获取TOF测距装置和拍摄装置之间的相对位姿参数。又如可移动平台可以在获取拍摄装置对可移动平台周围环境进行拍摄得到 的拍摄图像,并确定周围环境中目标对象在拍摄图像中的区域图像之后,获取TOF测距装置和拍摄装置之间的相对位姿参数。又如可移动平台可以在获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像之后,获取TOF测距装置和拍摄装置之间的相对位姿参数,然后确定周围环境中目标对象在拍摄图像中的区域图像。
S204,根据区域图像在拍摄图像中的位置、相对位姿参数确定目标对象在TOF测距装置的深度图上的区域深度图。
具体地,由于相对位姿关系表征了TOF测距装置与拍摄装置之间的安装位置关系,可以根据区域图像在拍摄图像中的位置、相对位姿参数将所述区域图像投影变换到TOF测距装置的深度图上以获取与所述区域图像对应的区域深度图。所述图像区域和所述区域深度图像用于指示空间中的同一个目标对象。
S205,根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值。
具体地,在现有技术中,接收装置接收反射的光信号,会根据接收到的光信号计算深度图中每一个像素对应的深度值。如前所述,根据接收到光信号解算深度值计算TOF测距装置的深度图上的所有像素的深度值,需要占用大量的计算资源和存储资源。在一些情况中,可移动平台只需要关注区域深度图像中像素对应的深度值。因此,在本申请实施例中,根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值,即仅根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不计算计算深度图中区域深度图之外的区域深度图中的像素的深度值。
在本申请实施例中,可移动平台仅根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值,可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。
请参见图3,是本发明另一实施例提出的一种深度计算方法的流程示意图,如图3所示,该方法可包括:
S301,获取拍摄装置对可移动平台周围环境进行拍摄得到的拍摄图像。
本申请实施例中的步骤S301可以参见上述实施例中步骤S201的具体描述,本申请实施例不再赘述。
S302,确定周围环境中目标对象在拍摄图像中的区域图像。
本申请实施例中的步骤S302可以参见上述实施例中步骤S202的具体描述,本申请实施例不再赘述。
S303,获取TOF测距装置和拍摄装置之间的相对位姿参数。
S304,根据区域图像在拍摄图像中的位置、相对位姿参数确定目标对象在TOF测距装置的深度图上的区域深度图。
S305,确定可移动平台的当前的工作状态。
可移动平台的工作状态可以包括第一工作状态和第二工作状态。其中,所述第一工作状态包括以下至少一种:所述拍摄装置在对所述目标对象进行对焦或跟焦,可移动平台处于人脸识别状态、视觉定位状态或背景虚化状态,所述可移动平台在确定所述目标对象的位置,或所述拍摄装置在对所述目标对象进行追踪拍摄。所述第二工作状态包括以下至少一种:所述可移动平台处于避障模式,所述可移动平台处于运动状态,或所述可移动平台由用户通过控制终端手动控制。
在一个实施例中,第一工作状态包含但不仅限于上述场景,只要是处于仅需获取深度图的部分像素的深度值的场景,都可以设置为第一工作状态。同理,第二工作状态包含但不仅限于上述场景,只要是处于需要获取深度图的每一个像素的深度值的场景,都可以设置为第二工作状态,具体不受本申请实施例的限制。
S306,在当前的工作状态为第一工作状态时,根据接收装置接收到的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值。
例如,在可移动平台控制拍摄装置对焦,或者人脸识别等场景时,无需获取整个深度图的深度信息,仅需获取区域深度图的深度信息,基于此,为了避免获取整个深度图的深度信息导致计算资源和存储资源的浪费,本申请实施例在可移动平台当前的工作状态为第一工作状态时,可以仅根据接收装置接收到 的反射的光信号计算区域深度图中像素的深度值,不根据光信号计算深度图中区域深度图之外的区域深度图中的像素的深度值,可减少深度解算所占用的计算资源和存储资源,降低计算复杂度,且对空间内存的需求较小。
S307,在当前的工作状态为第二工作状态时,根据接收装置接收到的反射的光信号计算深度图中每一个像素的深度值。
举例来说,在可移动平台处于需要获取整个深度图的深度信息的场景中,可移动平台可以直接根据接收装置接收到的反射的光信号计算深度图中每一个像素的深度值。例如,在可移动平台处于避障模式,可移动平台处于运动状态,或可移动平台由用户通过控制终端手动控制时,可移动平台可以确定当前的工作状态为第二状态,进而根据接收装置接收到的反射的光信号计算深度图中每一个像素的深度值。
在本申请实施例中,可移动平台可以根据不同工作状态,采用不同处理方式对深度图中的像素进行深度计算,例如在可移动平台当前的工作状态为第一工作状态时,仅计算区域深度图中像素的深度值,在可移动平台当前的工作状态为第二工作状态时,计算深度图中每一个像素的深度值,可基于不同工作状态选择不同的深度计算方式,以提高深度计算的灵活性。
在一个实施例中,基于上述实施例的描述,为了对可移动平台获取TOF测距装置和拍摄装置之间的相对位姿参数的方法进行具体描述,请参见图4,是本申请实施例提出的一种相对位姿参数获取方法,该方法包括:
S401,获取TOF测距装置在历史时刻输出的深度图,其中,历史时刻输出的深度图中的每一个像素是根据接收到的反射的光信号计算得到的。
S402,在历史时刻输出的深度图中确定多个深度跳变点。
在一个实施例中,可移动平台可以遍历历史时刻输出的深度图对应的点云,确定多个目标点云,将多个目标点云确定为多个深度跳变点。其中,目标点云的深度值与其相邻的多个点云的深度值大于或等于预设的深度阈值。
其中,目标点云与目标点云相邻的至少一个相邻点云之间的位置关系可以是目标点云与其八领域内的点,即九宫格,除目标点云以外的八个点称为八领域。可选的,目标点云与目标点云相邻的至少一个相邻点云之间的位置关系可 以是目标点云与其三领域内的点,即四宫格,除目标点云以外的三个点称为三领域。本发明对目标点云与相邻点云之间的“相邻”关系不作具体限定,可以预先提前设定,也可以在不同应用场景下,设置不同的目标点云与相邻点云之间的“相邻”关系。
举例来说,可移动平台遍历点云传感器采集到的点云,当某个点云的深度值与其八领域内任一点云的深度值之间的差异大于一定预设阈值时,即可判定此该点云为深度跳变点,其中深度跳变点可能是物体边缘点。可移动平台将确定得到的多个深度跳变点记作P,P为点的集合,例如P包含P1,P2,P3…。假设相邻关系为九宫格,点云中某一个点P1的深度值是3m,位于P1点左边的点为P2,P2的深度值是2.9m,位于P1点右边的点为P3,P3点的深度值是1.0m,那么可以确定P1点就是深度发生跳变处的点,即深度跳变点。
S403,获取TOF测距装置和拍摄装置之间的原始平移和原始旋转。
举例来说,TOF测距装置和拍摄装置之间的旋转和平移,可以写成:
Figure PCTCN2020111156-appb-000001
其中,R指的是TOF测距装置和拍摄装置之间的旋转,t指的是TOF测距装置和拍摄装置之间的平移。具体可以为,分别表示TOF测距装置坐标系到拍摄装置坐标系的旋转R;以及在拍摄装置坐标系下TOF测距装置坐标原点的位置t。
其中,TOF测距装置和拍摄装置之间的原始平移和原始旋转可以是可移动平台在出厂前标定的TOF测距装置和拍摄装置之间的的平移和旋转,也可以是在可移动平台出厂后预设的TOF测距装置和拍摄装置之间的平移和旋转。其中,预设的TOF测距装置和拍摄装置之间的平移和旋转可以是根据经验设定的。
S404,确定拍摄图像中像素点的边缘响应值。
在一个实施例中,可移动平台可以对拍摄图像运行边缘检测算法,以获取拍摄图像中像素点的边缘响应值。
举例来说,边缘检测算法可以包括:Sobel边缘检测算法、Isotropic Sobel边缘检测算法、Roberts边缘检测算法、Prewitt边缘检测算法、Laplacian边缘检测算法以及Canny边缘检测算法等。以边缘检测算法为Canny边缘检测算法为例,可移动平台检测拍摄装置采集到的拍摄图像中的边缘信息,可以得到 拍摄图像中的各个所述深度跳变点对应的像素点的边缘信息,具体可以为可移动平台获取所述多个深度跳变点在拍摄图像中的像素点的边缘响应值。
S405,根据多个深度跳变点的位置、原始平移、原始旋转和边缘响应值对原始平移和原始旋转进行优化运算,以确定TOF测距装置与拍摄装置之间的旋转和平移。
在一个实施例中,可移动平台可以根据多个深度跳变点、原始平移和原始旋转确定多个深度跳变点在拍摄图像中的像素点,获取多个深度跳变点在拍摄图像中的像素点的边缘响应值,将多个深度跳变点在拍摄图像中的像素点的边缘响应值总和最小作为优化目标,以原始平移和原始旋转为优化对象进行优化运算,将优化得到的原始平移和原始旋转确定为TOF测距装置与拍摄装置之间的旋转和平移。
举例来说,可移动平台按照TOF测距装置与拍摄装置之间的原始旋转R与原始位移关系t,可移动平台将多个深度跳变点的集合P投影到拍摄图像上,拍摄图像中的各个所述深度跳变点对应的像素点的像素坐标记为p,p为像素点的集合,p可以包括p1,p2,p3…。其中,这个投影过程记为π,则有:
p i=π(RP i+t)
其中,P,p,t是向量,R是矩阵。
在一种实现方式中,可移动平台依次找到P1对应的像素点p1,P2对应的像素点p2,P3对应的像素点p3…,得到多组(P,p)后,将各个p点处的边缘响应值E(p)总和作为Cost Function(代价函数)输入,优化求解得到精准的旋转R以及位移关系t。
Figure PCTCN2020111156-appb-000002
其中,arg代表优化的参数是旋转R,位移t。
在本申请实施例中,可移动平台获取TOF测距装置在历史时刻输出的深度图,在历史时刻输出的深度图中确定多个深度跳变点,确定拍摄图像中像素点的边缘响应值,根据多个深度跳变点的位置、原始平移、原始旋转和边缘响应值对原始平移和原始旋转进行优化运算,以确定TOF测距装置与拍摄装置之间的旋转和平移,可提高TOF测距装置与拍摄装置的相对位姿参数的准确 性。
本申请实施例提供了一种深度计算装置,应用于可移动平台中,图5是本申请实施例提供的应用于可移动平台的深度计算装置的结构图,如图5所示,所述应用于可移动平台的深度计算装置500包括存储器501和处理器502,可移动平台配置有拍摄装置和飞行时间TOF测距装置,TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置,其中,存储器502中存储有程序代码,处理器502调用存储器中的程序代码,当程序代码被执行时,处理器502执行如下操作:
获取所述拍摄装置对所述可移动平台周围环境进行拍摄得到的拍摄图像;
确定所述周围环境中目标对象在所述拍摄图像中的区域图像;
获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数;
根据所述区域图像在所述拍摄图像中的位置、所述相对位姿参数确定所述目标对象在所述TOF测距装置的深度图上的区域深度图;
根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
在一个实施例中,所述TOF测距装置和所述拍摄装置相对固定安装。
在一个实施例中,所述TOF测距装置和所述拍摄装置之间的相对位姿参数存储在所述可移动平台的本地存储装置中。
在一个实施例中,所述TOF测距装置和所述拍摄装置相对活动安装。
在一个实施例中,所述处理器502还用于执行如下操作:
确定所述可移动平台的当前的工作状态;
所述处理器502在根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值时,具体执行如下步骤:
若当前的工作状态为第一工作状态时,根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
在一个实施例中,所述第一工作状态包括以下至少一种:所述拍摄装置在对所述目标对象进行对焦或跟焦,所述可移动平台在确定所述目标对象的位置,或所述拍摄装置在对所述目标对象进行追踪拍摄。
在一个实施例中,所述处理器502还用于执行如下操作:
若当前的工作状态为第二工作状态时,根据所述接收装置接收到的反射的光信号计算所述深度图中每一个像素的深度值。
在一个实施例中,所述第二工作状态包括以下至少一种:所述可移动平台处于避障模式,所述可移动平台处于运动状态,或所述可移动平台由用户通过控制终端手动控制。
在一个实施例中,所述目标对象是由用户对显示有所述拍摄装置拍摄的图像的交互装置进行操作而选中的。
在一个实施例中,所述目标对象为所述拍摄装置的追踪拍摄对象。
在一个实施例中,所述处理器502在确定所述周围环境中目标对象在所述拍摄图像中的区域图像时,具体执行如下操作:
根据图像跟踪算法确定所述周围环境中目标对象在所述拍摄图像中的区域图像;或者,
将所述拍摄图像输入到预设的神经网络模型以确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
在一个实施例中,所述处理器502在确定所述周围环境中目标对象在所述拍摄图像中的区域图像时,具体执行如下操作:
获取关注指示信息,所述关注指示信息用于指示用户对所述拍摄图像的关注区域;
根据所述关注指示信息,确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
在一个实施例中,所述处理器502在获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数时,具体执行如下操作:
获取所述TOF测距装置在历史时刻输出的深度图,其中,所述历史时刻输出的深度图中的每一个像素是根据接收到的反射的所述光信号计算得到的;
在所述历史时刻输出的深度图中确定多个深度跳变点;
获取所述TOF测距装置和所述拍摄装置之间的原始平移和原始旋转;
确定所述拍摄图像中像素点的边缘响应值;
根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移。
在一个实施例中,所述处理器502在根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移时,具体执行如下操作:
根据所述多个深度跳变点和所述原始平移和原始旋转确定所述多个深度跳变点在所述拍摄图像中的像素点;
获取所述多个深度跳变点在所述拍摄图像中的像素点的边缘响应值;
将多个深度跳变点在拍摄图像中的像素点的边缘响应值总和最小作为优化目标,以所述原始平移和原始旋转为优化对象进行优化运算;
将优化得到的原始平移和原始旋转确定为所述TOF测距装置与所述拍摄装置之间的旋转和平移。
在一个实施例中,所述处理器502在确定所述拍摄图像中像素点的边缘响应值时,具体执行如下操作:
对所述拍摄图像运行边缘检测算法,以获取所述拍摄图像中像素点的边缘响应值。
在一个实施例中,所述处理器502在所述历史时刻输出的深度图中确定多个深度跳变点时,具体执行如下操作:
遍历所述历史时刻输出的深度图对应的点云,确定多个目标点云,其中,所述目标点云的深度值与其相邻的多个点云的深度值大于或等于预设的深度阈值;
将所述多个目标点云确定为所述多个深度跳变点。
在一个实施例中,所述TOF测距装置和所述拍摄装置之间的相对位姿参数包括所述TOF测距装置和所述拍摄装置之间的旋转关系和/或平移关系。
在一个实施例中,所述TOF测距装置包括3D-TOF相机或激光雷达或毫 米波雷达。
在一个实施例中,所述可移动平台包括用于对所述拍摄装置增稳的手持增稳系统或者无人机。
本实施例提供的应用于可移动平台的深度计算装置能执行前述实施例提供的如图2至图4所示的深度计算方法,且执行方式和有益效果类似,在这里不再赘述。
本申请实施例提供一种可移动平台,包括机身,动力系统,拍摄装置,TOF测距装置以及如前所述的深度计算装置。可移动平台的深度计算装置工作与前述相同或类似,此处不再赘述。动力系统,安装在所述机身,用于为所述可移动平台提供动力。拍摄装置,安装在所述机身,用于对所述可移动平台周围环境进行拍摄得到拍摄图像。TOF测距装置,安装在所述机身,用于获取深度图,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置。当所述可移动平台为无人机时,其动力系统可以包括旋翼、驱动旋翼旋转的电机及其电调。所述无人机可以是四旋翼、六旋翼、八旋翼或其他多旋翼无人机,此时无人机垂直起降进行工作。可以理解的是,无人机还可以是固定翼可移动平台或混合翼可移动平台。
在一个实施例中,所述TOF测距装置和所述拍摄装置相对固定安装。
在一个实施例中,所述TOF测距装置和所述拍摄装置相对活动安装。
在一个实施例中,所述可移动平台还包括第一通信设备,所述第一通信设备安装在所述机身,所述第一通信设备用于与控制终端进行数据交互。
在一个实施例中,所述可移动平台还包括第二通信设备,所述第二通信设备安装在所述机身,所述第二通信设备用于与显示有所述拍摄装置拍摄的图像的交互装置进行数据交互。
在一个实施例中,所述可移动平台至少包括如下的一种:用于对所述拍摄装置增稳的手持增稳系统或者无人机。
本申请实施例还提供一种计算机存储介质,所述计算机存储介质中存储有计算机程序指令,所述计算机程序指令被处理器执行时,用于执行如图2至图4所示的深度计算方法。
可以理解,以上所揭露的仅为本申请实施例的部分实施例而已,当然不能 以此来限定本发明之权利范围,本领域普通技术人员可以理解实现上述实施例的全部或部分流程,并依本发明权利要求所作的等同变化,仍属于发明所涵盖的范围。

Claims (45)

  1. 一种深度计算方法,其特征在于,所述方法应用于可移动平台,所述可移动平台配置有拍摄装置和TOF测距装置,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置,所述方法包括:
    获取所述拍摄装置对所述可移动平台周围环境进行拍摄得到的拍摄图像;
    确定所述周围环境中目标对象在所述拍摄图像中的区域图像;
    获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数;
    根据所述区域图像在所述拍摄图像中的位置、所述相对位姿参数确定所述目标对象在所述TOF测距装置的深度图上的区域深度图;
    根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
  2. 根据权利要求1所述的方法,其特征在于,所述TOF测距装置和所述拍摄装置相对固定安装。
  3. 根据权利要求2所述的方法,其特征在于,所述TOF测距装置和所述拍摄装置之间的相对位姿参数存储在所述可移动平台的本地存储装置中。
  4. 根据权利要求1所述的方法,其特征在于,所述TOF测距装置和所述拍摄装置相对活动安装。
  5. 根据权利要求1-4任一项所述的方法,其特征在于,所述方法还包括:
    确定所述可移动平台的当前的工作状态;
    所述根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值,包括:
    若当前的工作状态为第一工作状态时,根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
  6. 根据权利要求5所述的方法,其特征在于,所述第一工作状态包括以下至少一种:所述拍摄装置在对所述目标对象进行对焦或跟焦,所述可移动平台在确定所述目标对象的位置,或所述拍摄装置在对所述目标对象进行追踪拍摄。
  7. 根据权利要求5或6所述的方法,其特征在于,所述方法还包括:
    若当前的工作状态为第二工作状态时,根据所述接收装置接收到的反射的光信号计算所述深度图中每一个像素的深度值。
  8. 根据权利要求7所述的方法,其特征在于,所述第二工作状态包括以下至少一种:所述可移动平台处于避障模式,所述可移动平台处于运动状态,或所述可移动平台由用户通过控制终端手动控制。
  9. 根据权利要求1-8任一项所述的方法,其特征在于,所述目标对象是由用户对显示有所述拍摄装置拍摄的图像的交互装置进行操作而选中的。
  10. 根据权利要求1-9任一项所述的方法,其特征在于,所述目标对象为所述拍摄装置的追踪拍摄对象。
  11. 根据权利要求1-10任一项所述的方法,其特征在于,所述确定所述周围环境中目标对象在所述拍摄图像中的区域图像,包括:
    根据图像跟踪算法确定所述周围环境中目标对象在所述拍摄图像中的区域图像;或者,
    将所述拍摄图像输入到预设的神经网络模型以确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
  12. 根据权利要求1-10任一项所述的方法,其特征在于,所述确定所述周围环境中目标对象在所述拍摄图像中的区域图像,包括:
    获取关注指示信息,所述关注指示信息用于指示用户对所述拍摄图像的关注区域;
    根据所述关注指示信息,确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
  13. 根据权利要求1-12任一项所述的方法,其特征在于,所述获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数,包括:
    获取所述TOF测距装置在历史时刻输出的深度图,其中,所述历史时刻输出的深度图中的每一个像素是根据接收到的反射的所述光信号计算得到的;
    在所述历史时刻输出的深度图中确定多个深度跳变点;
    获取所述TOF测距装置和所述拍摄装置之间的原始平移和原始旋转;
    确定所述拍摄图像中像素点的边缘响应值;
    根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移。
  14. 根据权利要求13所述的方法,其特征在于,所述根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移,包括:
    根据所述多个深度跳变点和所述原始平移和原始旋转确定所述多个深度跳变点在所述拍摄图像中的像素点;
    获取所述多个深度跳变点在所述拍摄图像中的像素点的边缘响应值;
    将多个深度跳变点在拍摄图像中的像素点的边缘响应值总和最小作为优化目标,以所述原始平移和原始旋转为优化对象进行优化运算;
    将优化得到的原始平移和原始旋转确定为所述TOF测距装置与所述拍摄 装置之间的旋转和平移。
  15. 根据权利要求13或14所述的方法,其特征在于,所述确定所述拍摄图像中像素点的边缘响应值,包括:
    对所述拍摄图像运行边缘检测算法,以获取所述拍摄图像中像素点的边缘响应值。
  16. 根据权利要求13-15任一项所述的方法,其特征在于,所述在所述历史时刻输出的深度图中确定多个深度跳变点,包括:
    遍历所述历史时刻输出的深度图对应的点云,确定多个目标点云,其中,所述目标点云的深度值与其相邻的多个点云的深度值大于或等于预设的深度阈值;
    将所述多个目标点云确定为所述多个深度跳变点。
  17. 根据权利要求1-16任一项所述的方法,其特征在于,所述TOF测距装置和所述拍摄装置之间的相对位姿参数包括所述TOF测距装置和所述拍摄装置之间的旋转关系和/或平移关系。
  18. 根据权利要求1-17任一项所述的方法,其特征在于,所述TOF测距装置包括3D-TOF相机或激光雷达或毫米波雷达。
  19. 根据权利要求1-18任一项所述的方法,其特征在于,所述可移动平台包括用于对所述拍摄装置增稳的手持增稳系统或者无人机。
  20. 一种深度计算装置,其特征在于,包括存储器和处理器,所述深度计算装置应用于可移动平台,所述可移动平台配置有拍摄装置和TOF测距装置,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置,其中,
    所述存储器,用于存储有程序代码;
    所述处理器,调用存储器中的程序代码,当程序代码被执行时,用于执行如下操作:
    获取所述拍摄装置对所述可移动平台周围环境进行拍摄得到的拍摄图像;
    确定所述周围环境中目标对象在所述拍摄图像中的区域图像;
    获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数;
    根据所述区域图像在所述拍摄图像中的位置、所述相对位姿参数确定所述目标对象在所述TOF测距装置的深度图上的区域深度图;
    根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
  21. 根据权利要求20所述的装置,其特征在于,所述TOF测距装置和所述拍摄装置相对固定安装。
  22. 根据权利要求21所述的装置,其特征在于,所述TOF测距装置和所述拍摄装置之间的相对位姿参数存储在所述可移动平台的本地存储装置中。
  23. 根据权利要求20所述的装置,其特征在于,所述TOF测距装置和所述拍摄装置相对活动安装。
  24. 根据权利要求20-23任一项所述的装置,其特征在于,所述处理器还用于执行如下操作:
    确定所述可移动平台的当前的工作状态;
    所述处理器在根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值时,具体执行如下步骤:
    若当前的工作状态为第一工作状态时,根据所述接收装置接收到的反射的所述光信号计算所述区域深度图中像素的深度值,不根据所述光信号计算所述深度图中所述区域深度图之外的区域深度图中的像素的深度值。
  25. 根据权利要求24所述的装置,其特征在于,所述第一工作状态包括以下至少一种:所述拍摄装置在对所述目标对象进行对焦或跟焦,所述可移动平台在确定所述目标对象的位置,或所述拍摄装置在对所述目标对象进行追踪拍摄。
  26. 根据权利要求24或25所述的装置,其特征在于,所述处理器还用于执行如下操作:
    若当前的工作状态为第二工作状态时,根据所述接收装置接收到的反射的光信号计算所述深度图中每一个像素的深度值。
  27. 根据权利要求26所述的装置,其特征在于,所述第二工作状态包括以下至少一种:所述可移动平台处于避障模式,所述可移动平台处于运动状态,或所述可移动平台由用户通过控制终端手动控制。
  28. 根据权利要求20-27任一项所述的装置,其特征在于,所述目标对象是由用户对显示有所述拍摄装置拍摄的图像的交互装置进行操作而选中的。
  29. 根据权利要求20-28任一项所述的装置,其特征在于,所述目标对象为所述拍摄装置的追踪拍摄对象。
  30. 根据权利要求20-29任一项所述的装置,其特征在于,所述处理器在确定所述周围环境中目标对象在所述拍摄图像中的区域图像时,具体执行如下操作:
    根据图像跟踪算法确定所述周围环境中目标对象在所述拍摄图像中的区域图像;或者,
    将所述拍摄图像输入到预设的神经网络模型以确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
  31. 根据权利要求20-29任一项所述的装置,其特征在于,所述处理器在确定所述周围环境中目标对象在所述拍摄图像中的区域图像时,具体执行如下操作:
    获取关注指示信息,所述关注指示信息用于指示用户对所述拍摄图像的关注区域;
    根据所述关注指示信息,确定所述周围环境中目标对象在所述拍摄图像中的区域图像。
  32. 根据权利要求20-31任一项所述的装置,其特征在于,所述处理器在获取所述TOF测距装置和所述拍摄装置之间的相对位姿参数时,具体执行如下操作:
    获取所述TOF测距装置在历史时刻输出的深度图,其中,所述历史时刻输出的深度图中的每一个像素是根据接收到的反射的所述光信号计算得到的;
    在所述历史时刻输出的深度图中确定多个深度跳变点;
    获取所述TOF测距装置和所述拍摄装置之间的原始平移和原始旋转;
    确定所述拍摄图像中像素点的边缘响应值;
    根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移。
  33. 根据权利要求32所述的装置,其特征在于,所述处理器在根据所述多个深度跳变点的位置、所述原始平移、原始旋转和所述边缘响应值对所述原始平移和所述原始旋转进行优化运算,以确定所述TOF测距装置与拍摄装置之间的旋转和平移时,具体执行如下操作:
    根据所述多个深度跳变点和所述原始平移和原始旋转确定所述多个深度跳变点在所述拍摄图像中的像素点;
    获取所述多个深度跳变点在所述拍摄图像中的像素点的边缘响应值;
    将多个深度跳变点在拍摄图像中的像素点的边缘响应值总和最小作为优化目标,以所述原始平移和原始旋转为优化对象进行优化运算;
    将优化得到的原始平移和原始旋转确定为所述TOF测距装置与所述拍摄装置之间的旋转和平移。
  34. 根据权利要求32或33所述的装置,其特征在于,所述处理器在确定所述拍摄图像中像素点的边缘响应值时,具体执行如下操作:
    对所述拍摄图像运行边缘检测算法,以获取所述拍摄图像中像素点的边缘响应值。
  35. 根据权利要求32-34任一项所述的装置,其特征在于,所述处理器在所述历史时刻输出的深度图中确定多个深度跳变点时,具体执行如下操作:
    遍历所述历史时刻输出的深度图对应的点云,确定多个目标点云,其中,所述目标点云的深度值与其相邻的多个点云的深度值大于或等于预设的深度阈值;
    将所述多个目标点云确定为所述多个深度跳变点。
  36. 根据权利要求20-35任一项所述的装置,其特征在于,所述TOF测距装置和所述拍摄装置之间的相对位姿参数包括所述TOF测距装置和所述拍摄装置之间的旋转关系和/或平移关系。
  37. 根据权利要求20-36任一项所述的装置,其特征在于,所述TOF测距装置包括3D-TOF相机或激光雷达或毫米波雷达。
  38. 根据权利要求20-37任一项所述的装置,其特征在于,所述可移动平台包括用于对所述拍摄装置增稳的手持增稳系统或者无人机。
  39. 一种可移动平台,其特征在于,包括:
    机身;
    拍摄装置,安装在所述机身,用于对所述可移动平台周围环境进行拍摄得到拍摄图像;
    飞行时间TOF测距装置,安装在所述机身,用于获取深度图,所述TOF测距装置包括用于发射光信号的发射装置和用于接收由对象反射的所述光信号的接收装置;
    以及如权利要求20-38中任一项所述的深度计算装置。
  40. 根据权利要求39所述的可移动平台,其特征在于,所述TOF测距装置和所述拍摄装置相对固定安装。
  41. 根据权利要求39所述的可移动平台,其特征在于,所述TOF测距装置和所述拍摄装置相对活动安装。
  42. 根据权利要求39-41任一项所述的可移动平台,其特征在于,所述可移动平台还包括第一通信设备,所述第一通信设备安装在所述机身,所述第一通信设备用于与控制终端进行数据交互。
  43. 根据权利要求39-42任一项所述的可移动平台,其特征在于,所述可移动平台还包括第二通信设备,所述第二通信设备安装在所述机身,所述第二通信设备用于与显示有所述拍摄装置拍摄的图像的交互装置进行数据交互。
  44. 根据权利要求39-43任一项所述的可移动平台,其特征在于,所述可移动平台至少包括如下的一种:用于对所述拍摄装置增稳的手持增稳系统或者无人机。
  45. 一种计算机存储介质,其特征在于,所述计算机存储介质中存储有计算机程序指令,所述计算机程序指令被处理器执行时,用于执行如权利要求1-19任一项所述的深度计算方法。
PCT/CN2020/111156 2020-08-25 2020-08-25 深度计算方法、装置、可移动平台及存储介质 Ceased WO2022040941A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2020/111156 WO2022040941A1 (zh) 2020-08-25 2020-08-25 深度计算方法、装置、可移动平台及存储介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2020/111156 WO2022040941A1 (zh) 2020-08-25 2020-08-25 深度计算方法、装置、可移动平台及存储介质

Publications (1)

Publication Number Publication Date
WO2022040941A1 true WO2022040941A1 (zh) 2022-03-03

Family

ID=80352378

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/111156 Ceased WO2022040941A1 (zh) 2020-08-25 2020-08-25 深度计算方法、装置、可移动平台及存储介质

Country Status (1)

Country Link
WO (1) WO2022040941A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2024188219A1 (zh) * 2023-03-13 2024-09-19 鹏城实验室 目标定位与识别方法、设备及可读存储介质
CN119887945A (zh) * 2025-01-02 2025-04-25 杭州电子科技大学 一种大场景高速的深度计算方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103198473A (zh) * 2013-03-05 2013-07-10 腾讯科技(深圳)有限公司 一种深度图生成方法及装置
US20170337826A1 (en) * 2016-05-23 2017-11-23 Intel Corporation Flight Management and Control for Unmanned Aerial Vehicles
CN109583304A (zh) * 2018-10-23 2019-04-05 宁波盈芯信息科技有限公司 一种基于结构光模组的快速3d人脸点云生成方法及装置
WO2019144300A1 (zh) * 2018-01-23 2019-08-01 深圳市大疆创新科技有限公司 目标检测方法、装置和可移动平台
CN110291771A (zh) * 2018-07-23 2019-09-27 深圳市大疆创新科技有限公司 一种目标对象的深度信息获取方法及可移动平台

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103198473A (zh) * 2013-03-05 2013-07-10 腾讯科技(深圳)有限公司 一种深度图生成方法及装置
US20170337826A1 (en) * 2016-05-23 2017-11-23 Intel Corporation Flight Management and Control for Unmanned Aerial Vehicles
WO2019144300A1 (zh) * 2018-01-23 2019-08-01 深圳市大疆创新科技有限公司 目标检测方法、装置和可移动平台
CN110291771A (zh) * 2018-07-23 2019-09-27 深圳市大疆创新科技有限公司 一种目标对象的深度信息获取方法及可移动平台
CN109583304A (zh) * 2018-10-23 2019-04-05 宁波盈芯信息科技有限公司 一种基于结构光模组的快速3d人脸点云生成方法及装置

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2024188219A1 (zh) * 2023-03-13 2024-09-19 鹏城实验室 目标定位与识别方法、设备及可读存储介质
CN119887945A (zh) * 2025-01-02 2025-04-25 杭州电子科技大学 一种大场景高速的深度计算方法

Similar Documents

Publication Publication Date Title
US12416918B2 (en) Unmanned aerial image capture platform
US12586204B2 (en) Detecting optical discrepancies in captured images
CN112740269B (zh) 一种目标检测方法及装置
CN111344644B (zh) 用于基于运动的自动图像捕获的技术
JP6571274B2 (ja) レーザ深度マップサンプリングのためのシステム及び方法
KR20220028042A (ko) 포즈 결정 방법, 장치, 전자 기기, 저장 매체 및 프로그램
CN121411450A (zh) 用于对机器人进行初始化以沿着训练路线自主行进的系统和方法
CN108476288A (zh) 拍摄控制方法及装置
CN113378605A (zh) 多源信息融合方法及装置、电子设备和存储介质
WO2019051832A1 (zh) 可移动物体控制方法、设备及系统
CN110187720B (zh) 无人机导引方法、装置、系统、介质及电子设备
WO2021078003A1 (zh) 无人载具的避障方法、避障装置及无人载具
CN110337806A (zh) 集体照拍摄方法和装置
CN110880161A (zh) 一种多主机多深度摄像头的深度图像拼接融合方法及系统
EP4163675B1 (en) Illuminating an environment for localisation
CN116124119B (zh) 一种定位方法、定位设备及系统
WO2022040940A1 (zh) 标定方法、装置、可移动平台及存储介质
TW202539260A (zh) 自移動設備的環境感知方法、自移動設備、電腦儲存媒體及電子設備
WO2021217403A1 (zh) 可移动平台的控制方法、装置、设备及存储介质
CN118409589A (zh) 限制区域设置方法、自移动设备、终端设备及存储介质
CN116777965A (zh) 虚拟视窗配置装置、方法及系统
CN117057086A (zh) 基于目标识别与模型匹配的三维重建方法、装置及设备
CN118587346A (zh) 渲染方法及装置
US20260133592A1 (en) Unmanned Aerial Image Capture Platform
CN111950420A (zh) 一种避障方法、装置、设备和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20950614

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20950614

Country of ref document: EP

Kind code of ref document: A1