EP4690109A1 - Estimating a distance to one or more objects using an event based vision sensor - Google Patents
Estimating a distance to one or more objects using an event based vision sensorInfo
- Publication number
- EP4690109A1 EP4690109A1 EP24712840.8A EP24712840A EP4690109A1 EP 4690109 A1 EP4690109 A1 EP 4690109A1 EP 24712840 A EP24712840 A EP 24712840A EP 4690109 A1 EP4690109 A1 EP 4690109A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- evs
- pixels
- adjacent
- distance
- array
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/50—Depth or shape recovery
- G06T7/55—Depth or shape recovery from multiple images
- G06T7/557—Depth or shape recovery from multiple images from light fields, e.g. from plenoptic cameras
Definitions
- the present disclosure relates to methods and apparatuses for estimating a distance to one or more objects, and, more particularly, to methods and apparatuses for estimating distances using event-based vision sensors (EVS).
- EVS event-based vision sensors
- Depth estimation is a process of determining the distance of objects in a scene from an observer. Depth estimation is a fundamental problem in computer vision and may be essential for many applications, including robotics, augmented reality, and autonomous vehicles. There are several known concepts and techniques for depth estimation, some of which are: Stereo vision: This technique involves capturing images of a scene from two or more cameras with overlapping fields of view. By analyzing the differences in the images captured by the two cameras, it is possible to triangulate the position of objects in the scene and estimate their depth.
- Time-of-flight (ToF) ToF sensors use a modulated light source and directly or indirectly measure the time it takes for the light to travel to the object and back to the sensor. This time measurement can be used to estimate the distance of objects in the scene.
- Structured light sensors project a pattern of light onto the scene and measure the deformation of the pattern as it interacts with the objects in the scene. This deformation can be used to estimate the depth of objects.
- Monocular depth estimation This technique uses a single camera to estimate depth by analyzing features such as texture, edges, and gradients in the image. This approach typically requires the use of deep learning algorithms to train a model to estimate depth from a single image.
- LiDAR sensors use laser beams to scan the environment and measure the time it takes for the light to bounce back to the sensor. By analyzing the time-of-flight data, it is possible to create a 3D map of the environment and estimate depth.
- Conventional depth estimation concepts may involve relatively high latency, high complexity and/or high power-consumption, making them suboptimum for fast-paced and demanding environments, where real-time information and rapid response times are critical.
- the present disclosure provides an apparatus for estimating a distance to at least one object or portions thereof.
- the apparatus comprises a main lens, an array of micro-lenses downstream to the main lens, and an event-based vision sensor (EVS) comprising an array of EVS pixels downstream to the array of micro-lenses.
- EVS event-based vision sensor
- the array of EVS pixels comprising one or more groups of adjacent EVS pixels.
- a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels.
- a group of adjacent EVS pixels may comprise at least two adjacent EVS pixels.
- the apparatus further comprises processing circuitry configured to estimate the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
- each group may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group.
- the processing circuitry may then be configured to estimate the distance to the object based on a comparison of the respective first events and the respective second events of the plurality of groups of adjacent EVS pixels.
- Embodiments of the present disclosure invention may thus relate to an event camera system that may enhance object tracking capabilities by integrating an EVS with a micro-lens alignment.
- the EVS may capture image data in a unique way, providing high dynamic range, low latency, and low power information about the distance and movement of objects.
- an event camera or EVS only captures information when there is a change in the scene, resulting in a significant reduction in power consumption and latency. This makes EVS ideal for fast-paced and demanding environments, where real-time information and rapid response times are critical.
- the processing circuitry is configured to estimate the distance to the object based on a parity or disparity between events generated by EVS pixels of the one or more groups of adjacent EVS pixels.
- each group of EVS pixels may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group.
- the processing circuitry may be configured to estimate the distance to the object based on a parity or disparity between the respective first events and a parity or disparity between the respective second events of the one or more groups of adjacent EVS pixels.
- an “event” refers to a change in brightness of a EVS pixel in the camera's sensor. Unlike traditional cameras that capture a series of frames at fixed intervals, EVS capture changes in the scene at a very high temporal resolution, with each event being timestamped and associated with a specific pixel location. When an EVS pixel in the sensor detects a change in brightness, it generates an event that is transmitted to the camera's output in realtime. The event may contain information about a position, a polarity (+ or -) as well as a precise time at which it occurred.
- Parity/disparity between respective single or accumulated events generated by the group of adjacent EVS pixels may thus refer to parity/disparity with regards at least to the polarity or magnitude of the change in brightness.
- adjacent pixels of a pixel group under a common micro-lens detect events at the same time, i.e., in parallel or concurrently.
- a plurality of subsequent events of one EVS pixel may be converted into an image or accumulated event whose intensity or magnitude is a function of the motion history at the location of the EVS pixel, for example according to where tiast(x) ⁇ t is a timestamp of the last event at x and 5 is a constant decay rate parameter (i.e., 30ms).
- This equation may convert (accumulated) events into an image whose intensity is a function of the motion history at that location.
- Disparity may refer to a lack of equality or similarity between single or accumulated events, while parity may refer to a state of equality or similarity between single or accumulated events.
- the processing circuitry may be configured to estimate the distance to correspond to the focal length of the main lens in case of parity between the events (or event magnitudes) generated by the one or more groups of adjacent EVS pixels.
- the processing circuitry may be configured to estimate the distance to be smaller or larger than the focal length of the main lens in case of disparity between the events (or event magnitudes) generated by the one or more groups of adjacent EVS pixels.
- the processing circuitry may be configured to estimate the distance to be smaller than the focal length in case the disparity between the events is negative and to estimate the distance to be larger than the focal length in case the disparity between the events is positive, or vice versa.
- a slope or gradient of events (or accumulated event magnitudes) generated by the one or more groups of adjacent EVS pixels may be determined.
- Each group may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group.
- the processing circuitry may then be configured to estimate the distance to the object based on a comparison of the gradient of the first events (or accumulated event magnitudes) and the gradient of the second events (or accumulated event magnitudes) of the plurality of groups of adjacent EVS pixels.
- the processing circuitry may be configured to estimate the distance to be smaller than the focal length in case the gradient of the first events (or accumulated event magnitudes) is negative and to estimate the distance to be larger than the focal length in case the gradient of the first events (or accumulated event magnitudes) is positive, or vice versa.
- the processing circuitry may also be configured to estimate the distance to be smaller than the focal length in case the gradient of the second events (or accumulated event magnitudes) is positive and to estimate the distance to be larger than the focal length in case the gradient of the second events (or accumulated event magnitudes) is negative, or vice versa.
- the processing circuitry is configured to estimate the distance to the object further based on a distance between the main lens and the EVS. That is, the distance between the main lens and the EVS (or array of EVS pixels) may be necessary to draw conclusions on the distance to the object.
- the distance between the main lens and the EVS may be fixed and may correspond to the focal length of the main lens.
- the distance between the main lens and the EVS may be variable (adjustable). This may be beneficial for estimating distances to non-moving (still) objects, for example. In this case, events may be generated by moving the main lens instead of moving the object(s).
- the processing circuitry may be further configured to estimate motion of the object based on a sequence of subsequent events generated by one or more groups of adjacent EVS pixels.
- Motion estimation may, for example, be based on an optical flow algorithm calculating the motion of pixels, points, or objects in an image. This may involve computing a velocity vector of each pixel, point, or object which describes how much it has moved between two consecutive time instants.
- the processing circuitry may be further configured to reduce noise based on filtering a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels.
- the filtering may be a lowpass filtering, for example.
- the apparatus may further comprise an RGB (Red, Green, Blue) pixel array.
- An RGB pixel may comprise at least one R-subpixel, at least one G-subpixel, and at least one B-subpixel.
- the RGB pixel may be combined with a group of at least two adjacent EVS pixels.
- a single micro-lens of the micro-lens array may cover the group of adjacent EVS pixels and at least one RGB pixel. In this way, more information on a scene or objects thereof may be obtained and combined.
- the information may include depth, motion, speed, color, etc.
- the array of EVS pixels will typically comprise a plurality of groups of adjacent EVS pixels, wherein each group of adjacent EVS pixels is covered by a respective single micro-lens of the array of micro-lenses.
- the processing circuitry is configured to estimate a respective distance value for each group of adjacent EVS pixels based on a comparison of the events generated by a respective group of adjacent EVS pixels. For example, the processing circuitry is configured to estimate a first distance value for a first group of adjacent EVS pixels based on a comparison of first events generated by the first group of adjacent EVS pixels, estimate a second distance value for a second group of adjacent EVS pixels based on a comparison of second events generated by the second group of adjacent EVS pixels, etc.
- Each group of EVS pixels “sees” a different portion of an object or a scene.
- the processing circuitry is configured to estimate different distance values for different groups of adjacent EVS pixels asynchronously.
- An EVS is a type of camera that is designed to capture changes in a visual scene in a highly efficient and low- latency manner. Unlike traditional cameras that capture images at fixed intervals of time, EVS detect changes in the intensity of light on a per-pixel (or per pixel-group-basis) basis, and only report changes as they occur. EVS are also sometimes referred to as neurom orphic cameras, as they are designed to emulate the way that neurons in the brain process visual information. EVS operate by detecting changes in the intensity of light on individual pixels (or pixel groups), and then generating an event or signal that indicates the direction and magnitude of the change.
- the processing circuitry may be configured as an asynchronous neuromorphic processing circuit.
- the processing circuitry comprises a machine learning processor configured to map the events generated by the group of adjacent EVS pixels to a distance value associated with the group of adjacent EVS pixels.
- a machine learning processor configured to map the events generated by the group of adjacent EVS pixels to a distance value associated with the group of adjacent EVS pixels.
- CNNs Convolutional Neural Networks
- CNNs are a type of deep learning algorithm that have been highly successful in image and video recognition tasks. They work by automatically learning hierarchical features from the input data, which can be used to classify or detect objects in the scene. CNNs have been shown to be effective in processing event-based data, and can be used for tasks such as object recognition, motion detection, and optical flow estimation.
- RNNs Recurrent Neural Networks
- RNNs are a type of neural network that can process sequences of input data, such as the events generated by an event-based vision sensor. They can be used for tasks such as gesture recognition, action recognition, and activity recognition.
- SVMs Support Vector Machines
- SVMs are a type of supervised learning algorithm that can be used for classification tasks. They work by finding the optimal hyperplane that separates the data into different classes. SVMs have been shown to be effective in processing event-based data, and can be used for tasks such as object recognition, gesture recognition, and action recognition.
- Clustering Algorithms Clustering algorithms are unsupervised learning algorithms that can be used to group similar events together. They can be used for tasks such as anomaly detection, event segmentation, and feature extraction.
- DRL Deep Reinforcement Learning
- the present disclosure also provides a manned or unmanned vehicle comprising an apparatus for estimating a distance to an object according to any one of the previous claims.
- a manned or unmanned vehicle comprising an apparatus for estimating a distance to an object according to any one of the previous claims. Examples of such vehicles are robots, drones, or autonomous cars.
- the present disclosure also provides method for estimating a distance to a scene or an object.
- the method includes providing a main lens, providing an array of micro-lenses downstream to the main lens, and providing an EVS comprising an array of EVS pixels downstream to the array of micro-lenses.
- the array of EVS pixels comprises one or more groups of adjacent EVS pixels.
- a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels.
- the method further includes estimating the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
- FIG. 1 schematically illustrates an embodiment of an apparatus for estimating a distance from the apparatus to a scene/object
- Fig. 2A schematically shows a first example of an array of EVS pixels covered with micro-lenses
- Fig. 2B schematically shows a second example of an array of EVS pixels covered with micro-lenses
- Fig. 3A illustrates a scenario where an object is located at a distance from the apparatus which is larger than the focal length of main lens
- Fig. 3B illustrates a scenario where an object is located at a distance from the apparatus which is smaller than the focal length of main lens
- Fig. 4 illustrates a schematic flowchart how depth may be calculated based on event streams
- Fig. 5 schematically illustrates a combination of EVS pixels and RGB pixels into a hybrid pixel array.
- Fig. 1 schematically illustrates an apparatus 100 for estimating a distance c from the apparatus 100 to a scene 101 or object 102 in the scene 101.
- the apparatus 100 Seen in the direction of light traveling from the scene 101/object 102 toward the apparatus 100 (e.g., event camera), the apparatus 100 comprises a main lens 103, downstream to the main lens 103, an array of micro-lenses 104, and downstream to the array of micro-lenses 104, an EVS comprising an array of EVS pixels 105.
- downstream may be understood as further along the travel direction of light from the scene 101/object 102 toward the apparatus 100.
- the array of EVS pixels 105 may be a two-dimensional (2D) array comprising Ni xMi EVS pixels.
- the EVS 105 may span a plane. In the illustrated embodiment, the plane corresponds to a first y-z-plane in an x-y-z coordinate system.
- the array of micro-lenses 104 may be a 2D array comprising N2 x micro-lenses.
- the array of micro-lenses 104 may span a plane. In the illustrated embodiment, the spanned plane corresponds to a second y-z-plane in the x-y-z coordinate system.
- the first and the second y-z-planes may be essentially identical and the array of micro-lenses 104 and the EVS 105 may be regarded as one entity.
- groups of at least two adjacent EVS pixels are covered by a respective single micro-lens of the array of micro-lenses 104 to focus light onto the respective group of adjacent EVS pixels. That is, the EVS 105 may comprise a plurality of groups of adjacent EVS pixels. Each group comprises at least two adjacent EVS pixels covered by a single common micro-lens of the array of micro-lenses 104. This means that in general Ni > Nz and/or Mi > Mz. Different example configurations of micro-lenses 104 attached on top of groups of adjacent EVS pixels are illustrated in Fig. 2A, 2B.
- EVS 104, 105 An individual EVS pixel may be implemented using a photodiode, a comparator, and a reset circuit.
- a photodiode When light from scene 101/object 102 hits the photodiode, it generates a current that is proportional to the intensity of the light. This current may be compared to a reference current in the comparator circuit. If the current from the photodiode exceeds the reference current, the comparator generates an output pulse, or event. The event indicates that a change in illumination has been detected by the pixel.
- the reset circuit may then recharge the photodiode to its original state, ready for the next event.
- the EVS pixels 105 are designed to operate independently and asynchronously, detecting changes in the scene 101 and generating events in real-time. Because EVS pixels only generate events when a change in illumination is detected, they consume much less power than traditional cameras, and they are well-suited for high-speed, low-latency applications.
- the multiple subsequent events of one EVS pixel may be converted into an image whose intensity or magnitude is a function of the motion history at the location of the EVS pixel, for example according to where tiast(x) ⁇ t is a timestamp of the last event at x and 5 is a constant decay rate parameter (i.e., 30ms).
- This equation may convert (accumulated) events into an image whose intensity is a function of the motion history at that location.
- the micro-lenses 104 may improve the overall optical performance of the apparatus 100, particularly in terms of light gathering ability and event quality. These small lenses 104 are placed on top of the individual groups of EVS pixels 105.
- the micro-lenses 104 may help increase the amount of light that reaches the EVS pixels 105, by focusing incoming light onto sensitive areas of the EVS. This may be beneficial in low-light situations, where the amount of light available is limited.
- the micro-lenses 104 may help improve the signal -to-noise ratio and reduce noise in the resulting signal.
- the micro-lenses 104 may also help reduce optical aberrations, such as distortion and vignetting, which can cause image quality issues. By focusing the light more precisely onto the EVS, the micro-lenses 104 may help produce EVS images that are sharper and more detailed.
- the micro-lenses 104 may also enhance captured light/motion information by creating additional disparity based on the distance tfe between the main lens 103 and the EVS 104, 105, allowing the apparatus 100 (or event camera) to estimate a position and depth of moving objects 102 in real-time.
- the micro-lens alignment works by positioning one micro-lens 104-A over two or more EVS pixels 105-1, 105-2.
- the apparatus 100 (or event camera) can receive different events based on the distance d of moving objects. This additional information may enable to estimate the depth and motion of objects 102 in the scene 101, a capability that conventional event cameras do not possess.
- the output events from the EVS pixels 105-1, 105-2 or EVS 105 in general may be processed by specialized algorithms and/or neural networks to extract information about the scene 101 or object 102, such as the distance d from the apparatus 100 (or its main lens 103) to the scene 101/object 102.
- apparatus 100 further comprises processing circuitry 106 configured to estimate the distance d to the scene 101/object 102 based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels 105-1, 105-2 (or 105-1, 105-2, 105-3, 105-4).
- processing circuitry 106 is coupled to the EVS 105 to receive events generated by the one or more groups of adjacent EVS pixels.
- the principle of distance or depth estimation will be described in greater detail referring to Figs. 3A, 3B.
- the distance between the main lens 103 and EVS 104, 105 may be denoted by tfe.
- the focal length f of main lens 103 is the distance between the main lens 103 and its focal point when the main lens 103 is focused at infinity. When light passes through main lens 103, it is refracted or bent, causing the light rays to converge at a point known as the focal point.
- the distance between the center C of the main lens 103 and the focal point is the focal length f of the main lens 103.
- the distance tfe may also be variable or adjustable by adequate actuators. This may be beneficial for estimating distances to non-moving (still) objects. In this case, events may be generated by moving the main lens 103 back and forth instead of moving the object(s).
- three different regions related to the focal length f of the main lens 103 may be distinguished.
- Fig. 3A illustrates a scenario where (moving) object 102 is located at a distance d from the main lens 103 which is larger than the focal length f of main lens 103.
- object 102 is depicted as a point source at distance di > f from the center C of the main lens 103.
- FIG. 3 A In the upper left portion of Fig. 3 A, three adjacent pairs of EVS pixels at the center of EVS sensor 104, 105 are highlighted for illustrative purposes.
- a left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) is covered by micro-lens 104-A.
- Adjacent to the left pair a center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) is covered by micro-lens 104-B.
- Adjacent to the center pair, a right pair of adjacent EVS pixels 105-5 (L), 105-5 (R) is covered by microlens 104-C.
- point object 102 is located at a distance d from the main lens 103 which is larger than the focal length f of main lens 103, point object 102 is projected not only to the center pair of EVS pixels 105-3 (L), 105-4 (R) but also to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) as well as to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R).
- the (accumulated) events generated by the EVS pixels 105-3 (L), 105-4 (R) (center pair) are essentially equal in magnitude (parity)
- the events generated by the EVS pixels 105-1 (L), 105-2 (R) (left pair) and the events generated by the EVS pixels 105-5 (L), 105-6 (R) (right pair) are unequal in magnitude (disparity).
- the magni- tude of the events generated by left EVS pixel 105-1 (L) is smaller than that of the events generated by adjacent right EVS pixel 105-2 (R) of the left pair.
- the magnitude of the events generated by left EVS pixel 105-5 (L) is larger than that of the events generated by adjacent right EVS pixel 105-6 (R) of the right pair.
- processing circuitry 106 may be configured to estimate the distance d to the scene 101 or object 102 based on a disparity (or parity) between events generated by EVS pixels of the one or more pairs/groups of adjacent EVS pixels.
- a disparity or parity between events generated by EVS pixels of the one or more pairs/groups of adjacent EVS pixels.
- Further details on the distance d to object 102 may be obtained when taking the distribution or gradient (slope) of the (accumulated) events into account.
- the gradient of the events (magnitudes) of the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs/groups is positive, while the gradient of the events (magnitudes) of the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs/groups is negative. This may indicate that the distance d is larger than the focal length of main lens 103.
- the processing circuitry 106 may be configured to estimate the distance d to object 102 based on a comparison of the gradient of events generated by left EVS pixels of the adjacent EVS pixel groups with the gradient of events generated by right EVS pixels of the adjacent EVS pixel groups. For example, if the gradient of events generated by left EVS pixels 105-1, 105-3, 105-5 is positive while the gradient of events generated by the right EVS pixels 105-2, 105-4, 105-6 is negative, the distance d ⁇ to object 102 may be estimated to be larger than the focal length of main lens 103. Referring now to Fig.
- EVS pixels at the center of EVS sensor 104, 105 are highlighted for illustrative purposes.
- a left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) is covered by micro-lens 104-A.
- Adjacent to the left pair a center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) is covered by micro-lens 104-B.
- Adjacent to the center pair a right pair of adjacent EVS pixels 105-5 (L), 105-5 (R) is covered by micro-lens 104-C.
- point object 102 is located at distance d from the main lens 103 which is smaller than the focal length f of main lens 103, point object 102 is projected not only to the center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) but also to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) as well as to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R).
- the (accumulated) events generated by the two EVS pixels 105-3 (L), 105-4 (R) (center pair) are essentially equal in magnitude (parity)
- the (accumulated) events generated by the two EVS pixels 105-1 (L), 105-2 (R) (left pair) and the events generated by the two EVS pixels 105-5 (L), 105-6 (R) (right pair) are unequal in magnitude (disparity), respectively.
- the magnitude of the event generated by left EVS pixel 105-1 (L) is larger than that of the event generated by adjacent right EVS pixel 105-2 (R) of the left pair.
- the magnitude of the event generated by left EVS pixel 105-5 (L) is smaller than that of the event generated by adjacent right EVS pixel 105-6 (R) of the right pair.
- processing circuitry 106 may be configured to estimate the distance d to the scene 101 or object 102 based on a disparity (or parity) between events generated by EVS pixels of the one or more groups of adjacent EVS pixels.
- Further information on the distance d to object 102 may be obtained when taking the distribution or gradient of the events into account.
- the gradient of the events of the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs is negative, while the gradient of the events of the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs is positive.
- the processing circuitry 106 may be configured to estimate the distance d to object 102 based on a comparison of the gradient of events generated by left EVS pixels of the adjacent EVS pixel pairs with the gradient of events generated by right EVS pixels of the adjacent EVS pixel pairs. For example, if the gradient of events generated by left EVS pixels is negative, while the gradient of events generated by the right EVS pixels is positive, the distance d to object 102 may be estimated to be smaller than the focal length of main lens 103.
- point object 102 would be projected only to the center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) but not to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) and also not to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R). In this case, the (accumulated) events generated by the two EVS pixels 105-3 (L), 105-4 (R) (center pair) would still be essentially equal (parity).
- the processing circuitry 106 may be configured to estimate the distance d to correspond to the focal length of the main lens 103 in case of parity between events generated by EVS pixels of a group/pair of adjacent EVS pixels (covered by a common micro-lens).
- processing circuitry 106 may be configured to estimate the distance d ⁇ to be smaller or larger than the focal length of the main lens 103 in case of disparity between events generated by EVS pixels of a group or pair of adjacent EVS pixels (covered by a common micro-lens). In some implementations, the processing circuitry 106 may configured to estimate the distance d to be smaller than the focal length in case the disparity (or gradient of events) between a first group of (left) EVS pixels is negative while the disparity (or gradient of events) between a second group of (right) EVS pixels is positive.
- the processing circuitry 106 may be configured to estimate the distance d to be larger than the focal length in case the disparity (or gradient of events) between the first group of (left) EVS pixels is positive while the disparity (or gradient of events) between a second group of (right) EVS pixels is negative.
- noise of the events may be reduced based on filtering a sequence of subsequent events generated by EVS pixel of the array 105.
- filtering may correspond to lowpass filtering, which may be implemented by averaging or integrating subsequent events of EVS pixel, for example.
- the disparity that can be derived between adjacent EVS pixels of a group of EVS pixels is either positive, zero or negative.
- the depth d of the objects 102 may be calibrated and computed in a traditional way, such as with dynamic programming or semi global matching
- An additional interesting feature of case 2) (where we expect to have a disparity of zero) may be a low computational option to define a perimeter / safety distance, which can be defined just by the focal distance of the main lens 103.
- respective event intensities or magnitudes EDR(left) and EDR(right) may be obtained for each of the left and right EVS pixels.
- the scene 101 may thus be represented by the accumulated (e.g., summed) events EDR(left) and EDR (right) of the left and right EVS pixels, making up a left and a right half-image.
- EDR(left) and EDR (right) may be input into a spatial correspondence search processor 430, for example.
- SIFT Scale-Invariant Feature Transform
- the output of the spatial correspondence search block 430 may be a depth or distance of scene 101 or objects 102 relative to the focal length of main lens 103. With a previous calibration 440 of the focal length, absolute depth values may be obtained.
- processing circuitry 106 may estimate the relative or absolute distance d to object 102 based on disparity/parity between (accumulated) events of adjacent EVS pixels or groups of pixels, processing circuitry 106 may be further configured to also estimate motion of object 102 based on a sequence of subsequent events generated by one or more groups of adjacent EVS pixels of array 104 to which object 102 is projected to. For example, processing circuitry 106 may be configured to apply an optical flow algorithm for motion estimation.
- An optical flow algorithm works by analyzing the motion of events across EVS pixels in a sequence of subsequent events generated by one or more groups of adjacent EVS pixels.
- the algorithm may compute a vector for each event that describes its motion between two occurrences.
- a basic idea behind the algorithm is that the apparent motion of an object can be represented as a vector field, where each vector corresponds to the displacement of a EVS pixel or event in the image.
- the optical flow algorithm attempts to estimate this vector field by solving an optimization problem that minimizes the difference between the observed image and a predicted image based on the estimated motion.
- There are several techniques used to estimate the optical flow such as the Lucas-Kanade method, Horn-Schunck method, and the Farneback method.
- the Lucas-Kanade method is one of the most popular techniques used to estimate the optical flow. It works by assuming that the motion between two events is small and then estimating the flow vector for each EVS pixel by computing the image gradient at that EVS pixel and solving a set of linear equations.
- the Horn- Schunck method assumes that the motion between two frames is smooth and tries to estimate a smooth vector field that minimizes the difference between the observed image and a predicted image.
- the Farneback method is a dense optical flow algorithm that computes the optical flow for all EVS pixels in the image. It works by approximating the image patch around each EVS pixel with a polynomial and then estimating the motion using the polynomial coefficients.
- the proposed event camera system may use either conventional computer vision, machine learning, such as deep learning or self-supervised learning methods, to obtain the disparity and herewith depth, optical flow and herewith motion, and speed of an object.
- This information can be processed using an edge device or a neuromorphic processing unit, which may enable fully asynchronous depth and motion perception.
- the use of an edge device or neuromorphic processing unit would allow for extremely high temporally resolved information captured by the EVS, without the need for a separate computer or other processing device enabling potentially also a close loop system in the robotics sense.
- EVS are also sometimes referred to as neuromorphic cameras, as they are designed to emulate the way that neurons in the brain process visual information.
- EVS operate by detecting changes in the intensity of light on individual pixels (or pixel groups), and then generating an event or signal that indicates the direction and magnitude of the change. These events are reported to processing circuitry 106 asynchronously and in real-time, which means that the latency between the time of an event and its detection may be extremely low. Due to this asynchro- nism, the processing circuitry 106 may be configured as an asynchronous neuromorphic processing circuit. It is also noted that the proposed array of EVS pixels 105 may also be combined with an array of RGB pixels in some implementations. For example, a micro-lens of the micro-lens array 104 may cover groups of pixels, wherein each group comprises at least one RGB pixel and a pair of adjacent EVS pixels. Fig.
- Fig. 5 illustrates an embodiment where a group of pixels covered by a micro-lens consists of three RGB subpixels 505-R (Red), 505-G (Green), 505- B (Blue) and two adjacent EVS pixels 105-1, 105-2.
- Fig. 5 illustrates a 3 x 2 array of such groups of pixels. Each group of pixels is covered by an associated micro-lens (not shown).
- Fig. 5 illustrates only one of many possible layouts of combined EVS pixels and RGB pixels. With such a combination of EVS pixels and RGB pixels a more complete picture of scene 101 or object 102 may be gained, including color, depth, motion, and speed of an object 102.
- the camera system 100 may enable unmanned vehicles, such as robots and drones, to quickly adapt to their environment and navigate efficiently through areas with many obstacles.
- unmanned vehicles such as robots and drones
- the ability to estimate the depth of objects would allow for improved obstacle avoidance and more precise navigation, even in cluttered environments.
- the invention could also have a significant impact on the automotive industry, providing a highly targeted and effective camera system 100 for use in driver assistance and autonomous driving systems.
- the ability to estimate the depth of objects in real-time would allow for improved decision making and more precise control in these applications.
- Benefits of the proposed camera system 100 may extend beyond robotics and automotive applications.
- the improved accuracy and precision in object tracking and depth estimation could be of immense value.
- the high dynamic range and low latency of the EVS would allow for real-time monitoring of large areas, even in challenging lighting conditions, while the low power consumption would make it possible to use the system for extended periods without needing to replace batteries or connect to an external power source.
- the combination of the EVS and micro-lens alignment in the proposed camera system may offer improved accuracy and precision in object tracking and depth estimation, as well as exciting new opportunities for innovation in a variety of fields.
- the camera system With its low power consumption, high dynamic range, and low latency, the camera system is poised to have a significant impact on a wide range of applications, from robotics and driver assistance systems to surveillance and tracking systems.
- An example (e.g., example 1) relates to an apparatus for estimating a distance to an object, the apparatus comprising a main lens, an array of micro-lenses downstream to the main lens, and an EVS comprising an array of EVS pixels downstream to the array of micro-lenses.
- the array of EVS pixels comprises one or more groups of adjacent EVS pixels.
- a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels, and processing circuitry configured to estimate the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
- Another example relates to a previous example (e.g., example 1) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to the object based on a parity or disparity between events generated by EVS pixels of the one or more groups of adjacent EVS pixels.
- Another example e.g., example 3) relates to a previous example (e.g., one of the examples 1 or 2) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to correspond to the focal length of the main lens in case of parity between the events generated by the one or more groups of adjacent EVS pixels.
- Another example (e.g., example 4) relates to a previous example (e.g., one of the examples 1 to 3) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to be smaller or larger than the focal length of the main lens in case of disparity between the events generated by the one or more groups of adjacent EVS pixels.
- Another example (e.g., example 5) relates to a previous example (e.g., example 4) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to be smaller than the focal length in case the disparity between the events is negative and to estimate the distance to be larger than the focal length in case the disparity between the events is positive, or vice versa.
- Another example (e.g., example 6) relates to a previous example (e.g., one of the examples 1 to 5) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to the object further based on a distance between the main lens and the EVS.
- Another example (e.g., example 7) relates to a previous example (e.g., example 6) or to any other example, further comprising that the distance between the main lens and the EVS corresponds to the focal length of the main lens.
- Another example (e.g., example 8) relates to a previous example (e.g., one of the examples 1 to 7) or to any other example, further comprising that a distance between the main lens and the EVS is variable/adjustable.
- Another example (e.g., example 9) relates to a previous example (e.g., one of the examples 1 to 8) or to any other example, further comprising that the processing circuitry is further configured to estimate motion of the object based on a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels.
- Another example (e.g., example 10) relates to a previous example (e.g., one of the examples 1 to 9) or to any other example, further comprising that the processing circuitry is further configured to reduce noise based on filtering a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels.
- Another example (e.g., example 11) relates to a previous example (e.g., one of the examples 1 to 10) or to any other example, further comprising an RGB pixel array.
- Another example (e.g., example 12) relates to a previous example (e.g., example 11) or to any other example, further comprising that a single micro-lens covers a group of adjacent EVS pixels and at least one RGB pixel.
- Another example (e.g., example 15) relates to a vehicle comprising an apparatus for estimating a distance to an object according to any one of the previous examples.
- An example (e.g., example 16) relates to a method for estimating a distance to an object, the method comprising providing a main lens, providing an array of micro-lenses downstream to the main lens, and providing an EVS comprising an array of EVS pixels downstream to the array of micro-lenses.
- the array of EVS pixels comprises one or more groups of adja- cent EVS pixels.
- a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels.
- the method further includes estimating the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
- Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component.
- steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components.
- Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machine-executable, processorexecutable or computer-executable programs and instructions.
- Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example.
- Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
- FPLAs field programmable logic arrays
- F field) programmable gate arrays
- GPU graphics processor units
- ASICs application-specific integrated circuits
- ICs integrated circuits
- SoCs system-on-a-chip
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Measurement Of Optical Distance (AREA)
Abstract
The present disclosure relates to an apparatus (100) for estimating a distance to an object (102). The apparatus comprises a main lens (102), an array of micro-lenses (104) downstream to the main lens (102), and an event-based vision sensor (EVS) comprising an array of EVS pixels (105) downstream to the array of micro-lenses (104). The array of EVS pixels (105) comprises one or more groups of adjacent EVS pixels (105-1; 105-2), wherein a group of adjacent EVS pixels (105-1; 105-2) is covered by a respective micro-lens (104-A) of the array of micro-lenses to focus light onto the group of adjacent EVS pixels. The apparatus (100) also comprises processing circuitry (106) configured to estimate the distance to the object (102) based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels (105-1; 105-2).
Description
ESTIMATING A DISTANCE TO ONE OR MORE OBJECTS USING AN EVENT BASED VISION SENSOR
Field
The present disclosure relates to methods and apparatuses for estimating a distance to one or more objects, and, more particularly, to methods and apparatuses for estimating distances using event-based vision sensors (EVS).
Background
Depth estimation is a process of determining the distance of objects in a scene from an observer. Depth estimation is a fundamental problem in computer vision and may be essential for many applications, including robotics, augmented reality, and autonomous vehicles. There are several known concepts and techniques for depth estimation, some of which are: Stereo vision: This technique involves capturing images of a scene from two or more cameras with overlapping fields of view. By analyzing the differences in the images captured by the two cameras, it is possible to triangulate the position of objects in the scene and estimate their depth.
Time-of-flight (ToF): ToF sensors use a modulated light source and directly or indirectly measure the time it takes for the light to travel to the object and back to the sensor. This time measurement can be used to estimate the distance of objects in the scene.
Structured light: Structured light sensors project a pattern of light onto the scene and measure the deformation of the pattern as it interacts with the objects in the scene. This deformation can be used to estimate the depth of objects.
Monocular depth estimation: This technique uses a single camera to estimate depth by analyzing features such as texture, edges, and gradients in the image. This approach typically requires the use of deep learning algorithms to train a model to estimate depth from a single image.
LiDAR: LiDAR sensors use laser beams to scan the environment and measure the time it takes for the light to bounce back to the sensor. By analyzing the time-of-flight data, it is possible to create a 3D map of the environment and estimate depth.
Conventional depth estimation concepts may involve relatively high latency, high complexity and/or high power-consumption, making them suboptimum for fast-paced and demanding environments, where real-time information and rapid response times are critical.
Thus, there may be a need for improved concepts of depth estimation.
Summary
This need is addressed by methods and apparatuses in accordance with the appended independent claims. Possibly advantageous embodiments are addressed by the dependent claims.
According to a first aspect, the present disclosure provides an apparatus for estimating a distance to at least one object or portions thereof. The apparatus comprises a main lens, an array of micro-lenses downstream to the main lens, and an event-based vision sensor (EVS) comprising an array of EVS pixels downstream to the array of micro-lenses. The array of EVS pixels comprising one or more groups of adjacent EVS pixels. A group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels. A group of adjacent EVS pixels may comprise at least two adjacent EVS pixels. The apparatus further comprises processing circuitry configured to estimate the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
The skilled person having benefit from the present disclosure will appreciate that there will not only be one group of adjacent EVS pixels covered by a micro-lens of the array of microlenses. In practical applications, there may be a plurality of groups of adjacent EVS pixels, each group being covered by a respective micro-lens of the array of micro-lenses. Each group may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group. The processing circuitry may then be configured to estimate the distance to the object based on a comparison of the respective first events and the respective second events of the plurality of groups of adjacent EVS pixels.
Embodiments of the present disclosure invention may thus relate to an event camera system that may enhance object tracking capabilities by integrating an EVS with a micro-lens alignment. The EVS may capture image data in a unique way, providing high dynamic range, low latency, and low power information about the distance and movement of objects. Unlike traditional cameras, which continuously capture images and transmit them to a processor, an event camera or EVS only captures information when there is a change in the scene, resulting in a significant reduction in power consumption and latency. This makes EVS ideal for fast-paced and demanding environments, where real-time information and rapid response times are critical.
In some embodiments, the processing circuitry is configured to estimate the distance to the object based on a parity or disparity between events generated by EVS pixels of the one or more groups of adjacent EVS pixels. For example, each group of EVS pixels may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group. The processing circuitry may be configured to estimate the distance to the object based on a parity or disparity between the respective first events and a parity or disparity between the respective second events of the one or more groups of adjacent EVS pixels.
In an EVS, an "event" refers to a change in brightness of a EVS pixel in the camera's sensor. Unlike traditional cameras that capture a series of frames at fixed intervals, EVS capture changes in the scene at a very high temporal resolution, with each event being timestamped and associated with a specific pixel location. When an EVS pixel in the sensor detects a change in brightness, it generates an event that is transmitted to the camera's output in realtime. The event may contain information about a position, a polarity (+ or -) as well as a precise time at which it occurred. Parity/disparity between respective single or accumulated events generated by the group of adjacent EVS pixels may thus refer to parity/disparity with regards at least to the polarity or magnitude of the change in brightness. Here, it may be assumed that adjacent pixels of a pixel group under a common micro-lens detect events at the same time, i.e., in parallel or concurrently. For example, a plurality of subsequent events of one EVS pixel may be converted into an image or accumulated event whose intensity or magnitude is a function of the motion history at the location of the EVS pixel, for example according to
where tiast(x) < t is a timestamp of the last event at x and 5 is a constant decay rate parameter (i.e., 30ms). This equation may convert (accumulated) events into an image whose intensity is a function of the motion history at that location.
Disparity may refer to a lack of equality or similarity between single or accumulated events, while parity may refer to a state of equality or similarity between single or accumulated events. For example, the processing circuitry may be configured to estimate the distance to correspond to the focal length of the main lens in case of parity between the events (or event magnitudes) generated by the one or more groups of adjacent EVS pixels. On the other hand, the processing circuitry may be configured to estimate the distance to be smaller or larger than the focal length of the main lens in case of disparity between the events (or event magnitudes) generated by the one or more groups of adjacent EVS pixels. For example, the processing circuitry may be configured to estimate the distance to be smaller than the focal length in case the disparity between the events is negative and to estimate the distance to be larger than the focal length in case the disparity between the events is positive, or vice versa.
For this purpose, a slope or gradient of events (or accumulated event magnitudes) generated by the one or more groups of adjacent EVS pixels may be determined. Each group may be configured to generate a respective first event of a respective first EVS pixel and at least a respective second event of at least a respective second EVS pixel of said group. The processing circuitry may then be configured to estimate the distance to the object based on a comparison of the gradient of the first events (or accumulated event magnitudes) and the gradient of the second events (or accumulated event magnitudes) of the plurality of groups of adjacent EVS pixels. For example, the processing circuitry may be configured to estimate the distance to be smaller than the focal length in case the gradient of the first events (or accumulated event magnitudes) is negative and to estimate the distance to be larger than the focal length in case the gradient of the first events (or accumulated event magnitudes) is positive, or vice versa. The processing circuitry may also be configured to estimate the distance to be smaller than the focal length in case the gradient of the second events (or accumulated event magnitudes) is positive and to estimate the distance to be larger than the focal length in case the gradient of the second events (or accumulated event magnitudes) is negative, or vice versa.
In some embodiments, the processing circuitry is configured to estimate the distance to the object further based on a distance between the main lens and the EVS. That is, the distance between the main lens and the EVS (or array of EVS pixels) may be necessary to draw conclusions on the distance to the object. For example, the distance between the main lens and the EVS may be fixed and may correspond to the focal length of the main lens.
In some embodiments, the distance between the main lens and the EVS may be variable (adjustable). This may be beneficial for estimating distances to non-moving (still) objects, for example. In this case, events may be generated by moving the main lens instead of moving the object(s).
In some embodiments, the processing circuitry may be further configured to estimate motion of the object based on a sequence of subsequent events generated by one or more groups of adjacent EVS pixels. Motion estimation may, for example, be based on an optical flow algorithm calculating the motion of pixels, points, or objects in an image. This may involve computing a velocity vector of each pixel, point, or object which describes how much it has moved between two consecutive time instants.
In some embodiments, the processing circuitry may be further configured to reduce noise based on filtering a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels. The filtering may be a lowpass filtering, for example.
In some embodiments, the apparatus may further comprise an RGB (Red, Green, Blue) pixel array. An RGB pixel may comprise at least one R-subpixel, at least one G-subpixel, and at least one B-subpixel. The RGB pixel may be combined with a group of at least two adjacent EVS pixels. Thus, a single micro-lens of the micro-lens array may cover the group of adjacent EVS pixels and at least one RGB pixel. In this way, more information on a scene or objects thereof may be obtained and combined. The information may include depth, motion, speed, color, etc.
As mentioned before, the array of EVS pixels will typically comprise a plurality of groups of adjacent EVS pixels, wherein each group of adjacent EVS pixels is covered by a respective single micro-lens of the array of micro-lenses. The processing circuitry is configured to
estimate a respective distance value for each group of adjacent EVS pixels based on a comparison of the events generated by a respective group of adjacent EVS pixels. For example, the processing circuitry is configured to estimate a first distance value for a first group of adjacent EVS pixels based on a comparison of first events generated by the first group of adjacent EVS pixels, estimate a second distance value for a second group of adjacent EVS pixels based on a comparison of second events generated by the second group of adjacent EVS pixels, etc. Each group of EVS pixels “sees” a different portion of an object or a scene.
In some embodiments, the processing circuitry is configured to estimate different distance values for different groups of adjacent EVS pixels asynchronously. An EVS is a type of camera that is designed to capture changes in a visual scene in a highly efficient and low- latency manner. Unlike traditional cameras that capture images at fixed intervals of time, EVS detect changes in the intensity of light on a per-pixel (or per pixel-group-basis) basis, and only report changes as they occur. EVS are also sometimes referred to as neurom orphic cameras, as they are designed to emulate the way that neurons in the brain process visual information. EVS operate by detecting changes in the intensity of light on individual pixels (or pixel groups), and then generating an event or signal that indicates the direction and magnitude of the change. These events are reported asynchronously and in real-time, which means that the latency between the time of an event and its detection is extremely low, typically in the range of microseconds. Due to this asynchronism, the processing circuitry may be configured as an asynchronous neuromorphic processing circuit.
In some embodiments, the processing circuitry comprises a machine learning processor configured to map the events generated by the group of adjacent EVS pixels to a distance value associated with the group of adjacent EVS pixels. There are several machine learning algorithms that can be suitable for use in the context of EVS, depending on the specific application and task at hand.
Convolutional Neural Networks (CNNs): CNNs are a type of deep learning algorithm that have been highly successful in image and video recognition tasks. They work by automatically learning hierarchical features from the input data, which can be used to classify or detect objects in the scene. CNNs have been shown to be effective in processing event-based data, and can be used for tasks such as object recognition, motion detection, and optical flow estimation.
Recurrent Neural Networks (RNNs): RNNs are a type of neural network that can process sequences of input data, such as the events generated by an event-based vision sensor. They can be used for tasks such as gesture recognition, action recognition, and activity recognition.
Support Vector Machines (SVMs): SVMs are a type of supervised learning algorithm that can be used for classification tasks. They work by finding the optimal hyperplane that separates the data into different classes. SVMs have been shown to be effective in processing event-based data, and can be used for tasks such as object recognition, gesture recognition, and action recognition.
Clustering Algorithms: Clustering algorithms are unsupervised learning algorithms that can be used to group similar events together. They can be used for tasks such as anomaly detection, event segmentation, and feature extraction.
Deep Reinforcement Learning (DRL): DRL is a type of machine learning that combines deep learning with reinforcement learning. It has been used in event-based robotics for tasks such as navigation, grasping, and manipulation.
According to a further aspect, the present disclosure also provides a manned or unmanned vehicle comprising an apparatus for estimating a distance to an object according to any one of the previous claims. Examples of such vehicles are robots, drones, or autonomous cars.
According to yet a further aspect, the present disclosure also provides method for estimating a distance to a scene or an object. The method includes providing a main lens, providing an array of micro-lenses downstream to the main lens, and providing an EVS comprising an array of EVS pixels downstream to the array of micro-lenses. The array of EVS pixels comprises one or more groups of adjacent EVS pixels. A group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels. The method further includes estimating the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
Brief description of the Figures
Some examples of apparatuses and/or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
Fig. 1 schematically illustrates an embodiment of an apparatus for estimating a distance from the apparatus to a scene/object;
Fig. 2A schematically shows a first example of an array of EVS pixels covered with micro-lenses;
Fig. 2B schematically shows a second example of an array of EVS pixels covered with micro-lenses;
Fig. 3A illustrates a scenario where an object is located at a distance from the apparatus which is larger than the focal length of main lens;
Fig. 3B illustrates a scenario where an object is located at a distance from the apparatus which is smaller than the focal length of main lens;
Fig. 4 illustrates a schematic flowchart how depth may be calculated based on event streams; and
Fig. 5 schematically illustrates a combination of EVS pixels and RGB pixels into a hybrid pixel array.
Detailed Description
Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
Throughout the description of the figures same or similar reference numerals refer to same or similar elements and/or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and/or areas in the figures may also be exaggerated for clarification.
When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and/or B" may be used. This applies equivalently to combinations of more than two elements.
If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and/or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and/or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and/or a group thereof.
Fig. 1 schematically illustrates an apparatus 100 for estimating a distance c from the apparatus 100 to a scene 101 or object 102 in the scene 101.
Seen in the direction of light traveling from the scene 101/object 102 toward the apparatus 100 (e.g., event camera), the apparatus 100 comprises a main lens 103, downstream to the main lens 103, an array of micro-lenses 104, and downstream to the array of micro-lenses 104, an EVS comprising an array of EVS pixels 105. Here, “downstream” may be understood as further along the travel direction of light from the scene 101/object 102 toward the apparatus 100.
The array of EVS pixels 105 (or in short EVS) may be a two-dimensional (2D) array comprising Ni xMi EVS pixels. The EVS 105 may span a plane. In the illustrated embodiment, the plane corresponds to a first y-z-plane in an x-y-z coordinate system. Likewise, the array of micro-lenses 104 may be a 2D array comprising N2 x micro-lenses. The array of micro-lenses 104 may span a plane. In the illustrated embodiment, the spanned plane corresponds to a second y-z-plane in the x-y-z coordinate system. As the array of micro-lenses 104 may be directly attached to the EVS 105, the first and the second y-z-planes may be
essentially identical and the array of micro-lenses 104 and the EVS 105 may be regarded as one entity.
According to embodiments of the present disclosure, groups of at least two adjacent EVS pixels are covered by a respective single micro-lens of the array of micro-lenses 104 to focus light onto the respective group of adjacent EVS pixels. That is, the EVS 105 may comprise a plurality of groups of adjacent EVS pixels. Each group comprises at least two adjacent EVS pixels covered by a single common micro-lens of the array of micro-lenses 104. This means that in general Ni > Nz and/or Mi > Mz. Different example configurations of micro-lenses 104 attached on top of groups of adjacent EVS pixels are illustrated in Fig. 2A, 2B.
Fig. 2A shows an example of the array of EVS pixels 105 comprising a plurality of groups of adjacent EVS pixels. Here, a respective group of adjacent EVS pixels consists of two adjacent EVS pixels 105-1, 105-2 which are covered by a common micro-lens 104-A of the array of micro-lenses 104. EVS pixels 105-1, 105-2 may be directly adjacent. Adjacent to the group of the two adjacent EVS pixels 105-1, 105-2 are further groups of respective two adjacent EVS pixels. Each group of two adjacent EVS pixels is covered by an associated micro-lens 104-A, 104-B, 104-C.
Fig. 2B shows another example of the array of EVS pixels 105 comprising a plurality of groups of adjacent EVS pixels. Here, an individual group of adjacent EVS pixels consists of four adjacent EVS pixels 105-1, 105-2, 105-3, 105-4 which are arranged in a 2 x 2 matrix and covered by a common micro-lens 104-A of the array of micro-lenses 104. Adjacent to the group of four adjacent EVS pixels 105-1, 105-2, 105-3, 105-4 are further groups of respective four adjacent EVS pixels 105-1, 105-2, 105-3, 105-4. Each group of four adjacent EVS pixels is covered by an associated micro-lens. The four adjacent EVS pixels are arranged in a 2 x 2 matrix but a 1 x 4 configuration would also be conceivable. The skilled person having benefit from the present disclosure will appreciate that various other configurations are conceivable.
In the following, the array of micro-lenses 104 and the array of EVS pixels 105 together may be referred to as EVS 104, 105.
An individual EVS pixel may be implemented using a photodiode, a comparator, and a reset circuit. When light from scene 101/object 102 hits the photodiode, it generates a current that is proportional to the intensity of the light. This current may be compared to a reference current in the comparator circuit. If the current from the photodiode exceeds the reference current, the comparator generates an output pulse, or event. The event indicates that a change in illumination has been detected by the pixel. The reset circuit may then recharge the photodiode to its original state, ready for the next event. In contrast to traditional pixels used in cameras, the EVS pixels 105 are designed to operate independently and asynchronously, detecting changes in the scene 101 and generating events in real-time. Because EVS pixels only generate events when a change in illumination is detected, they consume much less power than traditional cameras, and they are well-suited for high-speed, low-latency applications.
Events in event-based vision sensors are generated by individual EVS pixels, and each event may comprise at least three pieces of information: the x and y coordinates of the EVS pixel generating the event, and the polarity of the change (either an increase or a decrease in luminance). Further, the EVS 105 may be designed to provide more detailed information about the magnitude of the change at each pixel. For example, EVS 105 may be capable of generating multiple events for a single EVS pixel if the change in luminance exceeds a certain threshold, with the number or timing of these events providing information about the magnitude of the change. The multiple subsequent events of one EVS pixel may be converted into an image whose intensity or magnitude is a function of the motion history at the location of the EVS pixel, for example according to
where tiast(x) < t is a timestamp of the last event at x and 5 is a constant decay rate parameter (i.e., 30ms). This equation may convert (accumulated) events into an image whose intensity is a function of the motion history at that location.
The micro-lenses 104 may improve the overall optical performance of the apparatus 100, particularly in terms of light gathering ability and event quality. These small lenses 104 are placed on top of the individual groups of EVS pixels 105. The micro-lenses 104 may help increase the amount of light that reaches the EVS pixels 105, by focusing incoming light onto sensitive areas of the EVS. This may be beneficial in low-light situations, where the
amount of light available is limited. By focusing the light onto the EVS, the micro-lenses 104 may help improve the signal -to-noise ratio and reduce noise in the resulting signal. In addition to improving light gathering ability, the micro-lenses 104 may also help reduce optical aberrations, such as distortion and vignetting, which can cause image quality issues. By focusing the light more precisely onto the EVS, the micro-lenses 104 may help produce EVS images that are sharper and more detailed.
The micro-lenses 104 may also enhance captured light/motion information by creating additional disparity based on the distance tfe between the main lens 103 and the EVS 104, 105, allowing the apparatus 100 (or event camera) to estimate a position and depth of moving objects 102 in real-time. As explained above, the micro-lens alignment works by positioning one micro-lens 104-A over two or more EVS pixels 105-1, 105-2. Depending on the focal distance of the main lens 103 in front of the EVS 104, 105, the apparatus 100 (or event camera) can receive different events based on the distance d of moving objects. This additional information may enable to estimate the depth and motion of objects 102 in the scene 101, a capability that conventional event cameras do not possess.
The output events from the EVS pixels 105-1, 105-2 or EVS 105 in general may be processed by specialized algorithms and/or neural networks to extract information about the scene 101 or object 102, such as the distance d from the apparatus 100 (or its main lens 103) to the scene 101/object 102. For this purpose, apparatus 100 further comprises processing circuitry 106 configured to estimate the distance d to the scene 101/object 102 based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels 105-1, 105-2 (or 105-1, 105-2, 105-3, 105-4). For this purpose, processing circuitry 106 is coupled to the EVS 105 to receive events generated by the one or more groups of adjacent EVS pixels. The principle of distance or depth estimation will be described in greater detail referring to Figs. 3A, 3B.
The distance between the main lens 103 and EVS 104, 105 (made up of micro-lens array 104 and EVS pixel array 105) may be denoted by tfe. In some embodiments, this distance tfe may be fixed and may correspond to the focal length f of main lens 103, i.e., tfe =f The focal length f of main lens 103 is the distance between the main lens 103 and its focal point when the main lens 103 is focused at infinity. When light passes through main lens 103, it is refracted or bent, causing the light rays to converge at a point known as the focal point. The
distance between the center C of the main lens 103 and the focal point is the focal length f of the main lens 103. In some embodiments, the distance tfe may also be variable or adjustable by adequate actuators. This may be beneficial for estimating distances to non-moving (still) objects. In this case, events may be generated by moving the main lens 103 back and forth instead of moving the object(s).
For example, with regards to the distance d to the scene 101/object 102, three different regions related to the focal length f of the main lens 103 may be distinguished.
1) Objects 102 in larger distance than the focal length (d > f)
2) Objects in focus (d ==f)
3) Objects closer than the focal length (di < f)
Fig. 3A illustrates a scenario where (moving) object 102 is located at a distance d from the main lens 103 which is larger than the focal length f of main lens 103. The distance tfe between main lens 103 and EVS sensor 104, 105 corresponds to the focal length f of main lens 103 (tfe =f). For illustrative purposes, object 102 is depicted as a point source at distance di > f from the center C of the main lens 103.
In the upper left portion of Fig. 3 A, three adjacent pairs of EVS pixels at the center of EVS sensor 104, 105 are highlighted for illustrative purposes. A left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) is covered by micro-lens 104-A. Adjacent to the left pair, a center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) is covered by micro-lens 104-B. Adjacent to the center pair, a right pair of adjacent EVS pixels 105-5 (L), 105-5 (R) is covered by microlens 104-C.
As point object 102 is located at a distance d from the main lens 103 which is larger than the focal length f of main lens 103, point object 102 is projected not only to the center pair of EVS pixels 105-3 (L), 105-4 (R) but also to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) as well as to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R). While the (accumulated) events generated by the EVS pixels 105-3 (L), 105-4 (R) (center pair) are essentially equal in magnitude (parity), the events generated by the EVS pixels 105-1 (L), 105-2 (R) (left pair) and the events generated by the EVS pixels 105-5 (L), 105-6 (R) (right pair) are unequal in magnitude (disparity). In the illustrated example, the magni-
tude of the events generated by left EVS pixel 105-1 (L) is smaller than that of the events generated by adjacent right EVS pixel 105-2 (R) of the left pair. The magnitude of the events generated by left EVS pixel 105-5 (L) is larger than that of the events generated by adjacent right EVS pixel 105-6 (R) of the right pair.
When looking at the magnitudes of the events generated by the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs, it may be observed that the respective magnitude is increasing from left to right, i.e., events of EVS pixel 105-1 < events of EVS pixel 105-3 < events of EVS pixel 105-5. When looking at the magnitudes of the events generated by the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs, it may be observed that the respective magnitude is decreasing from left to right, i.e., events of EVS pixel 105-2 > events of EVS pixel 105-4 > events of EVS pixel 105-6. This may be used by processing circuitry 106 and processing circuitry 106 may be configured to estimate the distance d to the scene 101 or object 102 based on a disparity (or parity) between events generated by EVS pixels of the one or more pairs/groups of adjacent EVS pixels. In the illustrated example, there is a disparity between the events generated by the adjacent EVS pixels 105-1 to 105-6. Therefore, it may be concluded that point object 102 is located at a distance d from the main lens 103 which is different from the the focal length of main lens 103.
Further details on the distance d to object 102 may be obtained when taking the distribution or gradient (slope) of the (accumulated) events into account. In the illustrated example, the gradient of the events (magnitudes) of the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs/groups is positive, while the gradient of the events (magnitudes) of the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs/groups is negative. This may indicate that the distance d is larger than the focal length of main lens 103. Thus, the processing circuitry 106 may be configured to estimate the distance d to object 102 based on a comparison of the gradient of events generated by left EVS pixels of the adjacent EVS pixel groups with the gradient of events generated by right EVS pixels of the adjacent EVS pixel groups. For example, if the gradient of events generated by left EVS pixels 105-1, 105-3, 105-5 is positive while the gradient of events generated by the right EVS pixels 105-2, 105-4, 105-6 is negative, the distance d\ to object 102 may be estimated to be larger than the focal length of main lens 103.
Referring now to Fig. 3B, it is illustrated another scenario where (moving) object 102 is located at a distance d from the main lens 103 which is smaller than the focal length f of main lens 103. Again, the distance
between main lens 103 and EVS sensor 104, 105 corresponds to the focal length f of main lens 103 (tfe =f). For illustrative purposes, object 102 is again depicted as a point source at distance d <f from the center of the main lens 103. The skilled person will appreciate, however, that the concept presented herein works for any object shapes.
Again, three adjacent pairs of EVS pixels at the center of EVS sensor 104, 105 are highlighted for illustrative purposes. A left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) is covered by micro-lens 104-A. Adjacent to the left pair, a center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) is covered by micro-lens 104-B. Adjacent to the center pair, a right pair of adjacent EVS pixels 105-5 (L), 105-5 (R) is covered by micro-lens 104-C.
As point object 102 is located at distance d from the main lens 103 which is smaller than the focal length f of main lens 103, point object 102 is projected not only to the center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) but also to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) as well as to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R). While the (accumulated) events generated by the two EVS pixels 105-3 (L), 105-4 (R) (center pair) are essentially equal in magnitude (parity), the (accumulated) events generated by the two EVS pixels 105-1 (L), 105-2 (R) (left pair) and the events generated by the two EVS pixels 105-5 (L), 105-6 (R) (right pair) are unequal in magnitude (disparity), respectively. In the illustrated example, the magnitude of the event generated by left EVS pixel 105-1 (L) is larger than that of the event generated by adjacent right EVS pixel 105-2 (R) of the left pair. The magnitude of the event generated by left EVS pixel 105-5 (L) is smaller than that of the event generated by adjacent right EVS pixel 105-6 (R) of the right pair.
When looking at the magnitudes of the events generated by the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs, it may be observed that the respective magnitude is decreasing from left to right, i.e., events of EVS pixel 105-1 > events of EVS pixel 105-3 > events of EVS pixel 105-5. When looking at the magnitudes of the events generated by the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs, it may be observed that the respective magnitude is increasing from left to right, i.e., events of right EVS pixel 105-2 < events of right EVS pixel 105-4 < events of right EVS
pixel 105-6. This confirms that processing circuitry 106 may be configured to estimate the distance d to the scene 101 or object 102 based on a disparity (or parity) between events generated by EVS pixels of the one or more groups of adjacent EVS pixels. In the illustrated example, there is a disparity between the events generated by the adjacent EVS pixels 105-1 to 105-6. Therefore, it may be concluded that point object 102 is located at a distance d from the main lens 103 which is different from the focal length of main lens 103.
Further information on the distance d to object 102 may be obtained when taking the distribution or gradient of the events into account. In the illustrated example of Fig. 3B, the gradient of the events of the respective left EVS pixels 105-1, 105-3, 105-5 of the adjacent EVS pixel pairs is negative, while the gradient of the events of the respective right EVS pixels 105-2, 105-4, 105-6 of the adjacent EVS pixel pairs is positive. This may indicate that the distance d is smaller than the focal length of main lens 103. Thus, the processing circuitry 106 may be configured to estimate the distance d to object 102 based on a comparison of the gradient of events generated by left EVS pixels of the adjacent EVS pixel pairs with the gradient of events generated by right EVS pixels of the adjacent EVS pixel pairs. For example, if the gradient of events generated by left EVS pixels is negative, while the gradient of events generated by the right EVS pixels is positive, the distance d to object 102 may be estimated to be smaller than the focal length of main lens 103.
If object 102 illustrated in Figs. 3 A, 3B was located at distance d equal to the focal length of main lens 103, point object 102 would be projected only to the center pair of adjacent EVS pixels 105-3 (L), 105-4 (R) but not to the left pair of adjacent EVS pixels 105-1 (L), 105-2 (R) and also not to the right pair of adjacent EVS pixels 105-5 (L), 105-5 (R). In this case, the (accumulated) events generated by the two EVS pixels 105-3 (L), 105-4 (R) (center pair) would still be essentially equal (parity). Also, the (accumulated) events generated by the two EVS pixels 105-1 (L), 105-2 (R) (left pair) and the events generated by the two EVS pixels 105-5 (L), 105-6 (R) (right pair) would still be essentially equal (e.g., zero), respectively. In this case, the processing circuitry 106 may be configured to estimate the distance d to correspond to the focal length of the main lens 103 in case of parity between events generated by EVS pixels of a group/pair of adjacent EVS pixels (covered by a common micro-lens).
On the other hand, processing circuitry 106 may configured to estimate the distance d\ to be smaller or larger than the focal length of the main lens 103 in case of disparity between events generated by EVS pixels of a group or pair of adjacent EVS pixels (covered by a common micro-lens). In some implementations, the processing circuitry 106 may configured to estimate the distance d to be smaller than the focal length in case the disparity (or gradient of events) between a first group of (left) EVS pixels is negative while the disparity (or gradient of events) between a second group of (right) EVS pixels is positive. The processing circuitry 106 may configured to estimate the distance d to be larger than the focal length in case the disparity (or gradient of events) between the first group of (left) EVS pixels is positive while the disparity (or gradient of events) between a second group of (right) EVS pixels is negative.
The skilled person having benefit from the present disclosure will appreciate that noise of the events may be reduced based on filtering a sequence of subsequent events generated by EVS pixel of the array 105. Such filtering may correspond to lowpass filtering, which may be implemented by averaging or integrating subsequent events of EVS pixel, for example.
Depending on the cases 1), 2), or 3), the disparity that can be derived between adjacent EVS pixels of a group of EVS pixels is either positive, zero or negative. Based on the focal length the depth d of the objects 102 may be calibrated and computed in a traditional way, such as with dynamic programming or semi global matching An additional interesting feature of case 2) (where we expect to have a disparity of zero) may be a low computational option to define a perimeter / safety distance, which can be defined just by the focal distance of the main lens 103.
Fig- 4 illustrates a schematic flowchart 400 how depth may be calculated based on event streams.
Block 410 represents EVS sensor 104, 105 comprising TV micro-lenses pLl, ..., pLN. In the illustrated example, each micro-lens pLl, ..., pLN of EVS sensor 104, 105 covers a respective left and right EVS pixel. More EVS pixels covered by a micro-lens are also possible. Thus, events generated at the output of EVS sensor 104, 105 may be grouped into events E(left) from respective left EVS pixels and events E(right) from respective right EVS pixels. At block 420, the events E(left) from the left EVS pixels and the events E(right) from the
right EVS pixels may be accumulated (e.g., summed) over a predefined time interval of t = 10 ms, for example. In this way, respective event intensities or magnitudes EDR(left) and EDR(right) may be obtained for each of the left and right EVS pixels. The scene 101 may thus be represented by the accumulated (e.g., summed) events EDR(left) and EDR (right) of the left and right EVS pixels, making up a left and a right half-image. EDR(left) and EDR (right) may be input into a spatial correspondence search processor 430, for example. Spatial correspondence search is a process used in computer vision and image processing to find correspondences between features in different images or frames. The goal of spatial correspondence search is to identify which features in one image (e.g., EDR(left)) correspond to features in another image (e.g., EDR (right)). For example, if we have two images of the same scene taken from different viewpoints, we might want to find the correspondence between specific points in one image and their corresponding points in the other image. This can be useful for tasks such as 3D reconstruction or motion tracking. One common approach to spatial correspondence search is to use feature descriptors, which are compact representations of local image features that can be compared across images. For example, SIFT (Scale-Invariant Feature Transform) is a popular feature descriptor that is invariant to changes in scale, rotation, and illumination. Once feature descriptors have been computed for each image, the next step is to match them across the images. This is typically done by comparing the descriptors and finding the closest match between them.
The output of the spatial correspondence search block 430 may be a depth or distance of scene 101 or objects 102 relative to the focal length of main lens 103. With a previous calibration 440 of the focal length, absolute depth values may be obtained.
While it has previously been shown that processing circuitry 106 may estimate the relative or absolute distance d to object 102 based on disparity/parity between (accumulated) events of adjacent EVS pixels or groups of pixels, processing circuitry 106 may be further configured to also estimate motion of object 102 based on a sequence of subsequent events generated by one or more groups of adjacent EVS pixels of array 104 to which object 102 is projected to. For example, processing circuitry 106 may be configured to apply an optical flow algorithm for motion estimation.
An optical flow algorithm works by analyzing the motion of events across EVS pixels in a sequence of subsequent events generated by one or more groups of adjacent EVS pixels.
The algorithm may compute a vector for each event that describes its motion between two occurrences. A basic idea behind the algorithm is that the apparent motion of an object can be represented as a vector field, where each vector corresponds to the displacement of a EVS pixel or event in the image. The optical flow algorithm attempts to estimate this vector field by solving an optimization problem that minimizes the difference between the observed image and a predicted image based on the estimated motion. There are several techniques used to estimate the optical flow, such as the Lucas-Kanade method, Horn-Schunck method, and the Farneback method. The Lucas-Kanade method is one of the most popular techniques used to estimate the optical flow. It works by assuming that the motion between two events is small and then estimating the flow vector for each EVS pixel by computing the image gradient at that EVS pixel and solving a set of linear equations. The Horn- Schunck method, on the other hand, assumes that the motion between two frames is smooth and tries to estimate a smooth vector field that minimizes the difference between the observed image and a predicted image. The Farneback method is a dense optical flow algorithm that computes the optical flow for all EVS pixels in the image. It works by approximating the image patch around each EVS pixel with a polynomial and then estimating the motion using the polynomial coefficients.
The proposed event camera system may use either conventional computer vision, machine learning, such as deep learning or self-supervised learning methods, to obtain the disparity and herewith depth, optical flow and herewith motion, and speed of an object. This information can be processed using an edge device or a neuromorphic processing unit, which may enable fully asynchronous depth and motion perception. The use of an edge device or neuromorphic processing unit would allow for extremely high temporally resolved information captured by the EVS, without the need for a separate computer or other processing device enabling potentially also a close loop system in the robotics sense. EVS are also sometimes referred to as neuromorphic cameras, as they are designed to emulate the way that neurons in the brain process visual information. EVS operate by detecting changes in the intensity of light on individual pixels (or pixel groups), and then generating an event or signal that indicates the direction and magnitude of the change. These events are reported to processing circuitry 106 asynchronously and in real-time, which means that the latency between the time of an event and its detection may be extremely low. Due to this asynchro- nism, the processing circuitry 106 may be configured as an asynchronous neuromorphic processing circuit.
It is also noted that the proposed array of EVS pixels 105 may also be combined with an array of RGB pixels in some implementations. For example, a micro-lens of the micro-lens array 104 may cover groups of pixels, wherein each group comprises at least one RGB pixel and a pair of adjacent EVS pixels. Fig. 5 illustrates an embodiment where a group of pixels covered by a micro-lens consists of three RGB subpixels 505-R (Red), 505-G (Green), 505- B (Blue) and two adjacent EVS pixels 105-1, 105-2. Fig. 5 illustrates a 3 x 2 array of such groups of pixels. Each group of pixels is covered by an associated micro-lens (not shown). The skilled person having benefit from the present disclosure will appreciate that Fig. 5 illustrates only one of many possible layouts of combined EVS pixels and RGB pixels. With such a combination of EVS pixels and RGB pixels a more complete picture of scene 101 or object 102 may be gained, including color, depth, motion, and speed of an object 102.
Advantages of the proposed monocular depth and motion estimation system 100 may be numerous and far-reaching. In the field of robotics, the camera system 100 may enable unmanned vehicles, such as robots and drones, to quickly adapt to their environment and navigate efficiently through areas with many obstacles. The ability to estimate the depth of objects would allow for improved obstacle avoidance and more precise navigation, even in cluttered environments. The invention could also have a significant impact on the automotive industry, providing a highly targeted and effective camera system 100 for use in driver assistance and autonomous driving systems. The ability to estimate the depth of objects in real-time would allow for improved decision making and more precise control in these applications.
Benefits of the proposed camera system 100 may extend beyond robotics and automotive applications. In surveillance and tracking systems, the improved accuracy and precision in object tracking and depth estimation could be of immense value. The high dynamic range and low latency of the EVS would allow for real-time monitoring of large areas, even in challenging lighting conditions, while the low power consumption would make it possible to use the system for extended periods without needing to replace batteries or connect to an external power source.
Further embodiments of the present disclosure may be considered:
1) Addition of structured active light (e.g. similar to a LIDAR system) to enable also the depth acquisition without apparent motion. This could also enable night vision like applications, which could rely on active Infra-red light.
2) High speed and object locked tracking since we could capture motion and depth at a very high speed. This could be realized effectively with a closed loop control mechanism of the camera focus based on the EVS pixel.
3) Mixed /hybrid pixel solution with a RGB sensor to enable the above-mentioned features (e.g. high speed tracking of objects motion and depth) and transfer them to the RGB vision based domain, for instance for blur reduction by signal processing or controlling the RGB pixel parameters (exposure time etc.)
In conclusion, the combination of the EVS and micro-lens alignment in the proposed camera system may offer improved accuracy and precision in object tracking and depth estimation, as well as exciting new opportunities for innovation in a variety of fields. With its low power consumption, high dynamic range, and low latency, the camera system is poised to have a significant impact on a wide range of applications, from robotics and driver assistance systems to surveillance and tracking systems.
In the following, some examples of the proposed concept are presented:
An example (e.g., example 1) relates to an apparatus for estimating a distance to an object, the apparatus comprising a main lens, an array of micro-lenses downstream to the main lens, and an EVS comprising an array of EVS pixels downstream to the array of micro-lenses. The array of EVS pixels comprises one or more groups of adjacent EVS pixels. A group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels, and processing circuitry configured to estimate the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
Another example (e.g., example 2) relates to a previous example (e.g., example 1) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to the object based on a parity or disparity between events generated by EVS pixels of the one or more groups of adjacent EVS pixels.
Another example (e.g., example 3) relates to a previous example (e.g., one of the examples 1 or 2) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to correspond to the focal length of the main lens in case of parity between the events generated by the one or more groups of adjacent EVS pixels.
Another example (e.g., example 4) relates to a previous example (e.g., one of the examples 1 to 3) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to be smaller or larger than the focal length of the main lens in case of disparity between the events generated by the one or more groups of adjacent EVS pixels.
Another example (e.g., example 5) relates to a previous example (e.g., example 4) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to be smaller than the focal length in case the disparity between the events is negative and to estimate the distance to be larger than the focal length in case the disparity between the events is positive, or vice versa.
Another example (e.g., example 6) relates to a previous example (e.g., one of the examples 1 to 5) or to any other example, further comprising that the processing circuitry is configured to estimate the distance to the object further based on a distance between the main lens and the EVS.
Another example (e.g., example 7) relates to a previous example (e.g., example 6) or to any other example, further comprising that the distance between the main lens and the EVS corresponds to the focal length of the main lens.
Another example (e.g., example 8) relates to a previous example (e.g., one of the examples 1 to 7) or to any other example, further comprising that a distance between the main lens and the EVS is variable/adjustable.
Another example (e.g., example 9) relates to a previous example (e.g., one of the examples 1 to 8) or to any other example, further comprising that the processing circuitry is further configured to estimate motion of the object based on a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels.
Another example (e.g., example 10) relates to a previous example (e.g., one of the examples 1 to 9) or to any other example, further comprising that the processing circuitry is further configured to reduce noise based on filtering a sequence of subsequent events generated by the one or more groups of adjacent EVS pixels.
Another example (e.g., example 11) relates to a previous example (e.g., one of the examples 1 to 10) or to any other example, further comprising an RGB pixel array.
Another example (e.g., example 12) relates to a previous example (e.g., example 11) or to any other example, further comprising that a single micro-lens covers a group of adjacent EVS pixels and at least one RGB pixel.
Another example (e.g., example 13) relates to a previous example (e.g., one of the examples 1 to 11) or to any other example, further comprising that the array of EVS pixels comprises a plurality of groups of at least two adjacent EVS pixels, each group being covered by a respective single micro-lens of the array of micro-lenses. Each group is configured to generate a respective first event of a first EVS pixel and a respective second event of a second EVS pixel. The processing circuitry is configured to estimate the distance to the object based on a comparison of the respective first events and the respective second events.
Another example (e.g., example 14) relates to a previous example (e.g., one of the examples 1 to 13) or to any other example, further comprising that the processing circuitry comprises a machine learning processor configured to map the events generated by the one or more groups of adjacent EVS pixels to a distance value associated with the one or more groups of adjacent EVS pixels.
Another example (e.g., example 15) relates to a vehicle comprising an apparatus for estimating a distance to an object according to any one of the previous examples.
An example (e.g., example 16) relates to a method for estimating a distance to an object, the method comprising providing a main lens, providing an array of micro-lenses downstream to the main lens, and providing an EVS comprising an array of EVS pixels downstream to the array of micro-lenses. The array of EVS pixels comprises one or more groups of adja-
cent EVS pixels. A group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels. The method further includes estimating the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and/or be broken up into several sub-steps, -functions, -processes or -operations.
If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
1. An apparatus for estimating a distance to an object, the apparatus comprising a main lens; an array of micro-lenses downstream to the main lens; an event-based vision sensor, EVS, comprising an array of EVS pixels downstream to the array of micro-lenses, the array of EVS pixels comprising one or more groups of adjacent EVS pixels, wherein a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels; and processing circuitry configured to estimate the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
2. The apparatus of claim 1, wherein the processing circuitry is configured to estimate the distance to the object based on a parity or disparity between events generated by EVS pixels of the one or more groups of adjacent EVS pixels.
3. The apparatus of claim 1, wherein the processing circuitry is configured to estimate the distance to correspond to the focal length of the main lens in case of parity between the events generated by the one or more groups of adjacent EVS pixels.
4. The apparatus of claim 1, wherein the processing circuitry is configured to estimate the distance to be smaller or larger than the focal length of the main lens in case of disparity between the events generated by the one or more groups of adjacent EVS pixels.
5. The apparatus of claim 4, wherein the processing circuitry is configured to estimate the distance to be smaller than the focal length in case the disparity between the events is negative and to estimate the distance to be larger than the focal length in case the disparity between the events is positive, or vice versa.
6. The apparatus of claim 1, wherein the processing circuitry is configured to estimate the distance to the object further based on a distance between the main lens and the EVS.
7. The apparatus of claim 6, wherein the distance between the main lens and the EVS corresponds to the focal length of the main lens.
8. The apparatus of claim 1, wherein a distance between the main lens and the EVS is variable.
9. The apparatus of claim 1, wherein the processing circuitry is further configured to estimate motion of the object based on a sequence of subsequent events generated by one or more groups of adjacent EVS pixels.
10. The apparatus of claim 1, wherein the processing circuitry is further configured to reduce noise based on filtering a sequence of subsequent events generated by the group of adjacent EVS pixels.
11. The apparatus of claim 1, further comprising an RGB pixel array.
12. The apparatus of claim 11, wherein a single micro-lens covers the group of adjacent
EVS pixels and at least one RGB pixel.
13. The apparatus of claim 1, wherein the array of EVS pixels comprises a plurality of groups of at least two adjacent EVS pixels, each group being covered by a respective single micro-lens of the array of microlenses, wherein each group is configured to generate a respective first event of a first EVS pixel and a respective second event of a second EVS pixel; the processing circuitry is configured to estimate the distance to the object based on a comparison of the respective first events and the respective second events.
14. The apparatus of claim 1, wherein the processing circuitry comprises a machine learning processor configured to map the events generated by the one or more groups of adjacent EVS pixels to a distance value associated with the one or more groups of adjacent EVS pixels.
15. A method for estimating a distance to an object, the method comprising
providing a main lens; providing an array of micro-lenses downstream to the main lens; providing an EVS comprising an array of EVS pixels downstream to the array of micro-lenses, the array of EVS pixels comprising one or more groups of adjacent EVS pix- els, wherein a group of adjacent EVS pixels is covered by a respective micro-lens of the array of micro-lenses to focus light onto the group of adjacent EVS pixels; and estimating the distance to the object based on a comparison of respective events generated by the one or more groups of adjacent EVS pixels.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23164894 | 2023-03-29 | ||
| PCT/EP2024/057782 WO2024200273A1 (en) | 2023-03-29 | 2024-03-22 | Estimating a distance to one or more objects using an event based vision sensor |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4690109A1 true EP4690109A1 (en) | 2026-02-11 |
Family
ID=85781612
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24712840.8A Pending EP4690109A1 (en) | 2023-03-29 | 2024-03-22 | Estimating a distance to one or more objects using an event based vision sensor |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4690109A1 (en) |
| CN (1) | CN120937045A (en) |
| WO (1) | WO2024200273A1 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022184557A1 (en) * | 2021-03-03 | 2022-09-09 | Sony Semiconductor Solutions Corporation | Time-of-flight data generation circuitry and time-of-flight data generation method |
| CN114401358A (en) * | 2022-01-28 | 2022-04-26 | 苏州华兴源创科技股份有限公司 | Event camera imaging device and method and event camera |
-
2024
- 2024-03-22 CN CN202480021018.1A patent/CN120937045A/en active Pending
- 2024-03-22 EP EP24712840.8A patent/EP4690109A1/en active Pending
- 2024-03-22 WO PCT/EP2024/057782 patent/WO2024200273A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| CN120937045A (en) | 2025-11-11 |
| WO2024200273A1 (en) | 2024-10-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP6509027B2 (en) | Object tracking device, optical apparatus, imaging device, control method of object tracking device, program | |
| Meyer et al. | Sensor fusion for joint 3d object detection and semantic segmentation | |
| US20220383530A1 (en) | Method and system for generating a depth map | |
| US20170124693A1 (en) | Pose Estimation using Sensors | |
| US20180173983A1 (en) | Digital neuromorphic (nm) sensor array, detector, engine and methodologies | |
| US8928737B2 (en) | System and method for three dimensional imaging | |
| Wang et al. | Stereo hybrid event-frame (shef) cameras for 3d perception | |
| CN116258734B (en) | Monocular three-dimensional instance segmentation method based on depth information guidance | |
| Ophoff et al. | Improving real-time pedestrian detectors with RGB+ depth fusion | |
| JP6675510B2 (en) | Subject tracking device and its control method, image processing device and its control method, imaging device and its control method, and program | |
| Nguyen et al. | CalibBD: Extrinsic calibration of the LiDAR and camera using a bidirectional neural network | |
| EP4690109A1 (en) | Estimating a distance to one or more objects using an event based vision sensor | |
| Xu et al. | A survey on event-driven 3d reconstruction: Development under different categories | |
| CN119555105A (en) | Visual odometer method, system, device and medium based on point and line features | |
| Shi et al. | A Review of Event-Based Indoor Positioning and Navigation. | |
| Kameyama et al. | Generation of Multi-Level Disparity Map from Stereo Wide Angle Fovea Vision System | |
| KR20160120533A (en) | Image sementation method in light field image | |
| Sun et al. | EV-FuseMODNet: moving object detection using the fusion of event camera and frame camera | |
| Shalma et al. | Structure tensor-based Gaussian kernel edge-adaptive depth map refinement with triangular point view in images | |
| CN117934577B (en) | A method, system, and device for achieving microsecond-level 3D detection based on binocular DVS | |
| US20260072167A1 (en) | Distance measuring apparatus, movable-unit control apparatus, distance measuring method, and storage medium | |
| Rego et al. | Event Camera Depth Estimation from Epipolar Plane Images | |
| Thaler et al. | Stereo Vision: Camera Agnostic Asymmetric Stereo | |
| Ding | Research on imaging system of artificial compound-eye and moving object detection | |
| Chen et al. | CalibCfC: LiDAR-Camera Self-Calibration with Convolutional Closed-Form Continuous-Time Layer |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251029 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |