EP4634612A1 - Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtung - Google Patents
Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtungInfo
- Publication number
- EP4634612A1 EP4634612A1 EP23798977.7A EP23798977A EP4634612A1 EP 4634612 A1 EP4634612 A1 EP 4634612A1 EP 23798977 A EP23798977 A EP 23798977A EP 4634612 A1 EP4634612 A1 EP 4634612A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- information
- light
- illumination
- sensor device
- generate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01B—MEASURING LENGTH, THICKNESS OR SIMILAR LINEAR DIMENSIONS; MEASURING ANGLES; MEASURING AREAS; MEASURING IRREGULARITIES OF SURFACES OR CONTOURS
- G01B11/00—Measuring arrangements characterised by the use of optical techniques
- G01B11/24—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures
- G01B11/25—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures by projecting a pattern, e.g. one or more lines, moiré fringes on the object
- G01B11/2513—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures by projecting a pattern, e.g. one or more lines, moiré fringes on the object with several lines being projected in more than one direction, e.g. grids, patterns
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01B—MEASURING LENGTH, THICKNESS OR SIMILAR LINEAR DIMENSIONS; MEASURING ANGLES; MEASURING AREAS; MEASURING IRREGULARITIES OF SURFACES OR CONTOURS
- G01B11/00—Measuring arrangements characterised by the use of optical techniques
- G01B11/24—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures
- G01B11/25—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures by projecting a pattern, e.g. one or more lines, moiré fringes on the object
- G01B11/2545—Measuring arrangements characterised by the use of optical techniques for measuring contours or curvatures by projecting a pattern, e.g. one or more lines, moiré fringes on the object with one projection direction and several detection directions, e.g. stereo
Definitions
- the present disclosure relates to a sensor device and a method for operating a sensor device.
- the present disclosure is related to the generation of depth information, preferably for hand tracking.
- Such techniques comprise the usage of structured light, i.e. the illumination of an object with static or time varying sparse light patterns in various solid angles such as to generate e.g. line, bar or checkerboard patterns, or active stereo depth sensing.
- structured light i.e. the illumination of an object with static or time varying sparse light patterns in various solid angles such as to generate e.g. line, bar or checkerboard patterns, or active stereo depth sensing.
- For a known orientation of light source and camera it is possible to determine the shape and the distance of an object from triangulation based on the known positions of the light source, the camera, the orientation of the emitted light in space, and the position of the according light signal on the camera.
- a set of illumination patterns providing high intensities at predetermined solid angles is sent out to an object and the distribution of light reflected from the object is measured by a receiver such as a camera (or two receivers for active stereo depth sensing).
- the task is then to find for the known solid angles of light emission, the solid angles of maximum light reception on the receiver. Due to the limited density of intensity changes in the illumination pattern and the limited pixel resolution, for the determination of the solid angle of maximum light reception a fit of the expected intensity distribution to the measured intensity values has to be made. In conventional systems this requires storage of all intensity values obtained at all pixels of the camera for all different illuminations. Only after all intensity values have been stored, a depth map can be generated. Thus, in conventional systems memory space must be large. Further, complete storage of intensity values leads to an enhanced latency in the system. Also, the available pixel resolution is limited by the readout speed, if applications with real-time behavior are envisaged, since too many pixels will lead to too long processing times.
- a sensor device for generating depth information for an object comprises a projector unit that is configured to project in a temporally consecutive manner a plurality of different illumination patterns in a projection solid angle to the object, where the projection solid angle consists of a predefined number of predetermined solid angles and each illumination pattern is generated by deciding for each of the predetermined solid angles whether or not to illuminate the respective predetermined solid angle by projecting light into it, a receiver unit that comprises a plurality of pixels, the receiver unit being configured to detect on each pixel intensities of light reflected from the object stemming from the illumination with the illumination patterns and/or from illumination with ambient light, and to generate an event at one of the pixels if the intensity detected at the pixel changes by more than a predetermined threshold, and a control unit that is configured to generate the depth information based on all events generated during a predetermined time period.
- a method for operating a sensor device for generating depth information for an object comprising: projecting in a temporally consecutive manner a plurality of different illumination patterns in a projection solid angle to the object, where the projection solid angle consists of a predefined number of predetermined solid angles and each illumination pattern is generated by deciding for each of the predetermined solid angles whether or not to illuminate the respective predetermined solid angle by projecting light into it; detecting on each pixel intensities of light reflected from the object stemming from the illumination with the illumination patterns and/or from illumination with ambient light, and to generate an event at one of the pixels if the intensity detected at the pixel changes by more than a predetermined threshold; and generating the depth information based on all events generated during a predetermined time period.
- Fig. 1 A is a simplified block diagram of the event detection circuitry of a solid-state imaging device including a pixel array.
- Fig. IB is a simplified block diagram of the pixel array illustrated in Fig. 1 A.
- Fig. 1C is a simplified block diagram of the imaging signal read-out circuitry of the solid-state imaging device of Fig. 1A.
- Fig. 2 shows schematically a sensor device.
- Fig. 3 shows schematically another sensor device.
- Fig. 4 shows another schematic representation of a sensor device.
- Fig. 5 shows an exemplary series of code words that encode illumination patterns.
- Figs. 6 shows schematically generation of depth information in a sensor device.
- Figs. 7 shows schematically a pixel array comprising pixels adapted to different wavelengths.
- Fig. 8 shows schematically generation of depth information in a sensor device.
- Figs. 9 shows schematically another sensor device.
- Fig. 10 shows schematically generation of depth information in a sensor device.
- Fig. 11 shows schematically spatially tiled illumination patterns.
- Figs. 12A and 12B show schematically different exemplary applications of a camera comprising a depth sensor device.
- Fig. 13 shows schematically a head mounted display comprising a depth sensor device.
- Fig. 14 shows schematically an industrial production device comprising a depth sensor device.
- Fig. 15 shows a schematic process flow of a method of operating a sensor device.
- Fig. 16 is a simplified perspective view of a solid-state imaging device with laminated structure according to an embodiment of the present disclosure.
- Fig. 17 illustrates simplified diagrams of configuration examples of a multi-layer solid-state imaging device to which a technology according to the present disclosure may be applied.
- Fig. 18 is a block diagram depicting an example of a schematic configuration of a vehicle control system.
- Fig. 19 is a diagram of assistance in explaining an example of installation positions of an outside-vehicle information detecting section and an imaging section of the vehicle control system of Fig. 18.
- the present disclosure relies on event detection by event visions sensor/dynamic vision sensors. Although these sensors are in principle known to a skilled person a brief overview will be given with respect to Figs. 1 A to 1C.
- Fig. 1A is a block diagram of a solid-state imaging device 100 employing event-based change detection.
- the solid-state imaging device 100 includes a pixel array 110 with one or more imaging pixels 111, wherein each pixel 111 includes a photoelectric conversion element PD.
- the pixel array 110 may be a one-dimensional pixel array with the photoelectric conversion elements PD of all pixels arranged along a straight or meandering line (line sensor).
- the pixel array 110 may be a two-dimensional array, wherein the photoelectric conversion elements PDs of the pixels 111 may be arranged along straight or meandering rows and along straight or meandering lines.
- the illustrations show a two dimensional array of pixels 111, wherein the pixels 111 are arranged along straight rows and along straight columns running orthogonal the rows.
- Each pixel 111 converts incoming light into an imaging signal representing the incoming light intensity and an event signal indicating a change of the light intensity, e.g. an increase by at least an upper threshold amount (positive polarity) and/or a decrease by at least a lower threshold amount (negative polarity).
- the function of each pixel 111 regarding intensity and event detection may be divided and different pixels observing the same solid angle can implement the respective functions.
- These different pixels may be subpixels and can be implemented such that they share part of the circuitry.
- the different pixels may also be part of different image sensors.
- a pixel capable of generating an imaging signal and an event signal this should be understood to include also a combination of pixels separately carrying out these functions as described above.
- a controller 120 performs a flow control of the processes in the pixel array 110.
- the controller 120 may control a threshold generation circuit 130 that determines and supplies thresholds to individual pixels 111 in the pixel array 110.
- a readout circuit 140 provides control signals for addressing individual pixels 111 and outputs information about the position of such pixels 111 that indicate an event. Since the solid-state imaging device 100 employs event-based change detection, the readout circuit 140 may output a variable amount of data per time unit.
- Fig. IB shows exemplarily details of the imaging pixels 111 in Fig. 1 A as far as their event detection capabilities are concerned. Of course, any other implementation that allows detection of events can be employed.
- Each pixel 111 includes a photoreceptor module PR and is assigned to a pixel back-end 300, wherein each complete pixel back-end 300 may be assigned to one single photoreceptor module PR.
- a pixel back-end 300 or parts thereof may be assigned to two or more photoreceptor modules PR, wherein the shared portion of the pixel back-end 300 may be sequentially connected to the assigned photoreceptor modules PR in a multiplexed manner.
- the photoreceptor module PR includes a photoelectric conversion element PD, e.g. a photodiode or another type of photosensor.
- the photoelectric conversion element PD converts impinging light 9 into a photocurrent Iphoto through the photoelectric conversion element PD, wherein the amount of the photocurrent Iphoto is a function of the light intensity of the impinging light 9.
- a photoreceptor circuit PRC converts the photocurrent Iphoto into a photoreceptor signal Vpr.
- the voltage of the photoreceptor signal Vpr is a function of the photocurrent Iphoto.
- a memory capacitor 310 stores electric charge and holds a memory voltage whose amount depends on a past photoreceptor signal Vpr.
- the memory capacitor 310 receives the photoreceptor signal Vpr such that a first electrode of the memory capacitor 310 carries a charge that is responsive to the photoreceptor signal Vpr and thus the light received by the photoelectric conversion element PD.
- a second electrode of the memory capacitor Cl is connected to the comparator node (inverting input) of a comparator circuit 340.
- the voltage of the comparator node, Vdiff varies with changes in the photoreceptor signal Vpr.
- the comparator circuit 340 compares the difference between the current photoreceptor signal Vpr and the past photoreceptor signal to a threshold.
- the comparator circuit 340 can be in each pixel back-end 300, or shared between a subset (for example a column) of pixels.
- each pixel 111 includes a pixel back-end 300 including a comparator circuit 340, such that the comparator circuit 340 is integral to the imaging pixel 111 and each imaging pixel 111 has a dedicated comparator circuit 340.
- a memory element 350 stores the comparator output in response to a sample signal from the controller 120.
- the memory element 350 may include a sampling circuit (for example a switch and a parasitic or explicit capacitor) and/or a digital memory circuit such as a latch or a flip-flop).
- the memory element 350 may be a sampling circuit.
- the memory element 350 may be configured to store one, two or more binary bits.
- An output signal of a reset circuit 380 may set the inverting input of the comparator circuit 340 to a predefined potential.
- the output signal of the reset circuit 380 may be controlled in response to the content of the memory element 350 and/or in response to a global reset signal received from the controller 120.
- the solid-state imaging device 100 is operated as follows: A change in light intensity of incident radiation 9 translates into a change of the photoreceptor signal Vpr. At times designated by the controller 120, the comparator circuit 340 compares Vdiff at the inverting input (comparator node) to a threshold Vb applied on its non-inverting input. At the same time, the controller 120 operates the memory element 350 to store the comparator output signal Vcomp.
- the memory element 350 may be located in either the pixel circuit 111 or in the readout circuit 140 shown in Fig. 1 A.
- conditional reset circuit 380 If the state of the stored comparator output signal indicates a change in light intensity AND the global reset signal GlobalReset (controlled by the controller 120) is active, the conditional reset circuit 380 outputs a reset output signal that resets Vdiff to a known level.
- the memory element 350 may include information indicating a change of the light intensity detected by the pixel 111 by more than a threshold value.
- the solid-state imaging device 120 may output the addresses (where the address of a pixel 111 corresponds to its row and column number) of those pixels 111 where a light intensity change has been detected.
- a detected light intensity change at a given pixel is called an event.
- the term ‘event’ means that the photoreceptor signal representing and being a function of light intensity of a pixel has changed by an amount greater than or equal to a threshold applied by the controller through the threshold generation circuit 130.
- the address of the corresponding pixel 111 is transmitted along with data indicating whether the light intensity change was positive or negative.
- the data indicating whether the light intensity change was positive or negative may include one single bit.
- each pixel 111 stores a representation of the light intensity at the previous instance in time.
- each pixel 111 stores a voltage Vdiff representing the difference between the photoreceptor signal at the time of the last event registered at the concerned pixel 111 and the current photoreceptor signal at this pixel 111.
- Vdiff at the comparator node may be first compared to a first threshold to detect an increase in light intensity (ON-event), and the comparator output is sampled on a (explicit or parasitic) capacitor or stored in a flip-flop. Then Vdiff at the comparator node is compared to a second threshold to detect a decrease in light intensity (OFF-event) and the comparator output is sampled on a (explicit or parasitic) capacitor or stored in a flip-flop.
- the global reset signal is sent to all pixels 111, and in each pixel 111 this global reset signal is logically ANDed with the sampled comparator outputs to reset only those pixels where an event has been detected. Then the sampled comparator output voltages are read out, and the corresponding pixel addresses sent to a data receiving device.
- Fig. 1C illustrates a configuration example of the solid-state imaging device 100 including an image sensor assembly 10 that is used for readout of intensity imaging signals in form of an active pixel sensor, APS.
- Fig. 1C is purely exemplary. Readout of imaging signals can also be implemented in any other known manner.
- the image sensor assembly 10 may use the same pixels 111 or may supplement these pixels 111 with additional pixels observing the respective same solid angles. In the following description the exemplary case of usage of the same pixel array 110 is chosen.
- the image sensor assembly 10 includes the pixel array 110, an address decoder 12, a pixel timing driving unit 13, an ADC (analog-to-digital converter) 14, and a sensor controller 15.
- the pixel array 110 includes a plurality of pixel circuits I IP arranged matrix-like in rows and columns.
- Each pixel circuit I IP includes a photosensitive element and FETs (field effect transistors) for controlling the signal output by the photosensitive element.
- the address decoder 12 and the pixel timing driving unit 13 control driving of each pixel circuit 1 IP disposed in the pixel array 110. That is, the address decoder 12 supplies a control signal for designating the pixel circuit 1 IP to be driven or the like to the pixel timing driving unit 13 according to an address, a latch signal, and the like supplied from the sensor controller 15.
- the pixel timing driving unit 13 drives the FETs of the pixel circuit I IP according to driving timing signals supplied from the sensor controller 15 and the control signal supplied from the address decoder 12.
- each ADC 14 performs an analog-to-digital conversion on the pixel output signals successively output from the column of the pixel array unit 11 and outputs the digital pixel data DPXS to a signal processing unit.
- each ADC 14 includes a comparator 23, a digital-to-analog converter (DAC) 22 and a counter 24.
- DAC digital-to-analog converter
- the sensor controller 15 controls the image sensor assembly 10. That is, for example, the sensor controller 15 supplies the address and the latch signal to the address decoder 12, and supplies the driving timing signal to the pixel timing driving unit 13. In addition, the sensor controller 15 may supply a control signal for controlling the ADC 14.
- the pixel circuit IIP includes the photoelectric conversion element PD as the photosensitive element.
- the photoelectric conversion element PD may include or may be composed of, for example, a photodiode. With respect to one photoelectric conversion element PD, the pixel circuit IIP may have four FETs serving as active elements, i.e., a transfer transistor TG, a reset transistor RST, an amplification transistor AMP, and a selection transistor SEL.
- the photoelectric conversion element PD photoelectrically converts incident light into electric charges (here, electrons).
- the amount of electric charge generated in the photoelectric conversion element PD corresponds to the amount of the incident light.
- the transfer transistor TG is connected between the photoelectric conversion element PD and a floating diffusion region FD.
- the transfer transistor TG serves as a transfer element for transferring charge from the photoelectric conversion element PD to the floating diffusion region FD.
- the floating diffusion region FD serves as temporary local charge storage.
- a transfer signal serving as a control signal is supplied to the gate (transfer gate) of the transfer transistor TG through a transfer control line.
- the transfer transistor TG may transfer electrons photoelectrically converted by the photoelectric conversion element PD to the floating diffusion FD.
- the reset transistor RST is connected between the floating diffusion FD and a power supply line to which a positive supply voltage VDD is supplied.
- a reset signal serving as a control signal is supplied to the gate of the reset transistor RST through a reset control line.
- the reset transistor RST serving as a reset element resets a potential of the floating diffusion FD to that of the power supply line.
- the floating diffusion FD is connected to the gate of the amplification transistor AMP serving as an amplification element. That is, the floating diffusion FD functions as the input node of the amplification transistor AMP serving as an amplification element.
- the amplification transistor AMP and the selection transistor SEL are connected in series between the power supply line VDD and a vertical signal line VSL.
- the amplification transistor AMP is connected to the signal line VSL through the selection transistor SEL and constitutes a source-follower circuit with a constant current source 21 illustrated as part of the ADC 14.
- a selection signal serving as a control signal corresponding to an address signal is supplied to the gate of the selection transistor SEL through a selection control line, and the selection transistor SEL is turned on.
- the amplification transistor AMP amplifies the potential of the floating diffusion FD and outputs a voltage corresponding to the potential of the floating diffusion FD to the signal line VSL.
- the signal line VSL transfers the pixel output signal from the pixel circuit IIP to the ADC 14.
- the ADC 14 may include a DAC 22, the constant current source 21 connected to the vertical signal line VSL, a comparator 23, and a counter 24.
- the vertical signal line VSL, the constant current source 21 and the amplifier transistor AMP of the pixel circuit 1 IP combine to a source follower circuit.
- the DAC 22 generates and outputs a reference signal.
- the DAC 22 may generate a reference signal including a reference voltage ramp. Within the voltage ramp, the reference signal steadily increases per time unit. The increase may be linear or not linear.
- the comparator 23 has two input terminals.
- the reference signal output from the DAC 22 is supplied to a first input terminal of the comparator 23 through a first capacitor CL
- the pixel output signal transmitted through the vertical signal line VSL is supplied to the second input terminal of the comparator 23 through a second capacitor C2.
- the comparator 23 compares the pixel output signal and the reference signal that are supplied to the two input terminals with each other, and outputs a comparator output signal representing the comparison result. That is, the comparator 23 outputs the comparator output signal representing the magnitude relationship between the pixel output signal and the reference signal. For example, the comparator output signal may have high level when the pixel output signal is higher than the reference signal and may have low level otherwise, or vice versa.
- the comparator output signal VCO is supplied to the counter 24.
- the counter 24 counts a count value in synchronization with a predetermined clock. That is, the counter 24 starts the count of the count value from the start of a P phase or a D phase when the DAC 22 starts to decrease the reference signal, and counts the count value until the magnitude relationship between the pixel output signal and the reference signal changes and the comparator output signal is inverted. When the comparator output signal is inverted, the counter 24 stops the count of the count value and outputs the count value at that time as the AD conversion result (digital pixel data DPXS) of the pixel output signal.
- event sensor as described above might be used in the following, when it is referred to event detection. However, any other manner of implementation of event detection might be applicable. In particular, event detection may also be carried out in sensors directed to external influences other than light, like e.g. sound, pressure, temperature or the like. In principle, the below description could be applied to any sensor that provides a binary output in response to the detection of intensities.
- Fig. 2 shows schematically a sensor device 1000 for generating depth information for an object O, i.e. a device that allows deduction of distances of surface elements of the object O or the posture of the object O in three- dimensional space to the sensor device 1000.
- the sensor device 1000 may be capable to generate the depth information itself or may only generate data based on which the depth information can be established in further processing steps.
- the sensor device 1000 comprises a projector unit 1010 configured to illuminate different locations of the object O during different time periods with an illumination pattern.
- the projector unit 1010 is configured to project in a temporally consecutive manner a plurality of different illumination patterns in a projection solid angle PS to the object O, where the projection solid angle PS consists of a predefined number of predetermined solid angles and each illumination pattern is generated by deciding for each of the predetermined solid angles whether or not to illuminate the respective predetermined solid angle by projecting light into it.
- the predetermined solid angles have linear or rectangular cross sections and are parallel to each other in a cross-sectional plane.
- the predetermined solid angles are adjacent to each other such that they completely fill the projection solid angle PS.
- the predetermined solid angles may also be separated from each other by a certain distance such that predetermined solid angles are separated by non-illuminated regions.
- Fig. 2 shows an example, in which only one line is projected to the object O
- several lines may be projected at the same time, as schematically shown in Fig. 3.
- the equidistant arrangement of lines in Fig. 3 is only chosen for simplicity.
- the lines may have arbitrary positions.
- the number of lines, i.e. the number of illuminated predetermined solid angles may change with time.
- the change of the illumination may be effected e.g. by using a fixed light source, the light of which is deflected at different times at different angles.
- a mirror tilted by a micro-electro-mechanical system (MEMS) might be used to deflect the illumination pattern and/or a refractive grating may be used to produce a plurality of lines.
- MEMS micro-electro-mechanical system
- an array of vertical-cavity surface -emitting lasers (VCSELs) or any other laser LEDs might be used that illuminate different parts of the object O at different times.
- shielding optics like slit plates or LCD-panels to produce time varying illumination patterns.
- the projector unit 1010 may be arranged such that some points in the field of view of the projector unit 1010 are never illuminated with the illumination patterns.
- the illumination pattern sent out from the projector unit 1010 may be fixed, while the object O moves across the illumination pattern.
- the precise manner of the generation of the illumination pattern and its movement across the object is arbitrary, as long as different positions of the object O are illuminated during different time periods.
- the sensor device 1000 comprises a receiver unit 1020 comprising a plurality of pixels 1025. Due to the surface structure of the object O, the illumination patterns are reflected from the object O in distorted form and forms an image I of the illumination pattern on the receiver unit 1020.
- the pixels 1025 of the receiver unit 1020 may in principle be capable to generate a full intensity image of the received reflection. More importantly, the receiver unit 1020 is configured to detect on each pixel 1025 intensities of light reflected from the object O while it is illuminated with the illumination pattern, and to generate an event at one of the pixels 1025 if the intensity detected at the pixel 1025 changes by more than a predetermined threshold.
- the receiver unit 1020 can act as an event sensor as described above with respect to Figs.
- 1A to 1C that can detect changes in the received intensity that exceed a given threshold.
- positive and negative changes might be detectable, leading to events of so-called positive or negative polarity.
- the event detection thresholds might be dynamically adaptable and might differ for positive and negative polarities.
- the receiver unit 1020 is also capable to detect intensities that stem from illumination with ambient light, and to generate an event at one of the pixels 1025 if the intensity detected at the pixel 1025 changes by more than a predetermined threshold for both kinds of light sources. This is schematically illustrated in Fig. 4, where the receiver unit 120 not only receives and detects the reflected illumination patterns, but also light from external light sources, such as the sun or lamps, that are reflected on the object O.
- the reflected ambient light will be mainly visible light, while the light of the illumination patterns may have any wavelengths.
- both the ambient light and the light of the illumination patterns may be visible light.
- the light of the illumination patterns may also be infrared light, if it is intended that the illumination patterns are not to be seen on the object O.
- the pixels 1025 of the receiver unit 1020 are then capable to detect infrared light as well as visible light, i.e. they have sensitivity for light having wavelengths between 1,000 pm to 380 nm.
- Fig. 5 shows a symbolization of the change of illumination patterns over time as used in the projector unit 1010.
- the illumination patterns are formed by illuminating 8 different predetermined solid angles, e.g. by projecting lines at 8 different locations onto an object, or by illuminating 8 different (preferably rectangular) areas on the object, which might even have a resolution comparably to those of the pixels 1025 of the reception unit 1020.
- the number of different predetermined solid angles might be different.
- Projecting light into one of the 8 predetermined solid angles of Fig. 5 is indicated by a white square, while missing illumination is illustrated by a black square.
- illumination/white may be represented by a “1” and missing illumination/black by a “0”.
- An according representation of changes of illuminations as code words projected to a given predetermined solid angle is particularly adapted to the usage of an event-based vision sensor. In fact, each transition from “0” to “1” in a code word, will trigger a positive polarity event, while transitions from “ 1” to “0” trigger a negative polarity event.
- the control unit can compare event sequences generated at certain pixels with illumination patterns projected into specific predetermined solid angles. Matching event sequences and code words allows then to identify the optical path of the light of the illumination pattern via the object, i.e. to determine the distance via triangulation.
- the control unit 1030 receives all the events generated during a predetermined time period, i.e. the events generated due to illumination by the projector unit 1010 and the events generated due to the illumination with ambient light. The control unit 1030 generates depth information based on all these events.
- control unit 1030 may be any arrangement of circuitry that is capable to carry out the functions described herein.
- the control unit 1030 may be constituted by a processor.
- the control unit 1030 may be part of the pixel section of the sensor device 1000 and may be placed on the same die(s) as the other components of the sensor device 1000. But the control unit 1030 may also be arranged separately, e.g. on a separate die.
- the functions of the control unit 1030 may be fully implemented in hardware, in software or may be implemented as a mixture of hardware and software functions.
- the dataset on which the control unit 1030 operates to determine the depth information contains therefore a part that is related to overall shape and texture of the objects (ambient light) and a part dedicated to determining the distance between object (O) and sensor device 1000 (illumination patterns). This increase in information increases the accuracy with which the depth information can be generated. However, since only a single receiver unit 1020 is used, this improvement comes without a raise in costs and/or power consumption.
- events from ambient light will be generated with a higher frequency than the frequency of changes between differing illumination patterns. This allows refining at high temporal rate and by using the events generated due to the ambient light core estimates made via the events caused by the illumination patterns. For example, in monitoring an object O a distance between sensor device 1000 and the object O may be established with a first frequency by using the events caused by the illumination patterns. From such a measurement it can be established how big the object O looks for a given distance. The events generated by ambient light, which represent basically a two-dimensional image of the observed scene can then be used to determine changes in the apparent size of the object O, which allow to deduce changes of the distance between object O and sensor device 1000.
- This adaption of the distance can be carried out with a considerably larger frequency than the original distance estimation.
- the temporal resolution of the generation of depth information is increased. Just the same the frequency of the changes of illumination patterns can be lowered to save energy. By using the events generated due to the incident ambient light the temporal resolution of the generation of depth information can still be kept in an acceptable range.
- Fig. 6 shows schematically how the control unit 1030 operates on the event data such as to generate the depth information. All the blocks shown in Fig. 6 may be constituted by hardware, i.e. processors or circuitry, or software and/or a mixture thereof.
- the control device 1030 receives all the events E generated during the predetermined time period. Further, the control device 1030 is provided from the projector unit 1010 with information P that indicates whether during a given time period within the predetermined time period the projector unit 1010 did not project light on the object O. Alternatively, the control unit 1030 generates the information P itself and controls the projector unit 1010 accordingly, i.e. the control unit 1030 decides when to project the illumination patterns and when not. In this manner, the control unit 1030 is configured to determine whether or not the projector unit 1010 projects light during a given time period into the projection solid angle PS. This might be done in a discriminator block 1032 as shown in Fig. 6.
- the control unit 1030 determines that no light has been projected into the projection solid angle PS during the given time period (“N” in Fig. 6), the control unit 1030 is configured to generate first information by processing the events on the assumption that all events generated during the given time period are caused by ambient light. Of course this assumption is adequate since if no illumination patterns are projected onto the object, events can only be caused by changes in the ambient light. Event processing is therefore executed as if no projector unit 1010 were present.
- the first information may then be the mere event data, i.e. the position of the event on the receiver unit 1020, its time stamp, and its polarity.
- metadata M may be added to the event data to generate the first information, e.g. by concatenating event data and metadata M.
- the metadata may e.g. indicate the bearing vector of the event generating pixel, i.e. the vector pointing from the camera center to the pixel.
- control unit 1030 determines that light has been projected into the projection solid angle PS during the given time period
- the control unit 1030 is configured to generate second information by processing the events on the assumption that all events generated during the given time period are caused by light projected by the projector unit 1010.
- the control unit 1030 will operate as if there was no ambient light. The error introduced by this assumption is small enough to be negligible or compensated during further processing the second information.
- the control unit 1030 operates with the knowledge that information on the distance between the object O and the sensor device 1000 is encoded in the event data and will extract this information on the distance. As shown by the dashed boxes in Fig. 6 this may be done by generating a depth map of the object O based on the obtained events.
- a code word extractor module 1033 and a triangulation module 1034 may be provided.
- the code word extractor module 1033 operates on the known distribution L of the illumination patterns and on the event data E.
- the control unit 1030 tries to establish a correspondence between the sequences of positive and negative polarity events received in each pixel 1025 and the known illumination patterns that can be expressed as code words as explained above with respect to Fig.
- the events can be accumulated in a temporal histogram per pixel, where time bins of the histogram match the change frequency of illumination patterns. It is then possible to cross correlate the histogram of one pixel (and optionally of its neighboring pixels) with the projected code words. The code word with the maximum correlation will be assumed to have generated the corresponding events.
- a neural network that takes as input the histogram of one pixel (and optionally of its neighboring pixel) and directly outputs correlation scores of the different code words or the responsible code word.
- any other method that allows to identify the part of the illumination that caused an event sequence on a certain pixel, may be used.
- the correlation of pixel 1025 and code word is then forwarded to the triangulation module 1034, which establishes the depth map based on this correlation and the known geometry G of the setup, i.e. the relative positions of projector unit 1010 and receiver unit 1020, by triangulation in an in principle known manner.
- the triangulation module 1034 which establishes the depth map based on this correlation and the known geometry G of the setup, i.e. the relative positions of projector unit 1010 and receiver unit 1020, by triangulation in an in principle known manner.
- the distance between sensor device 1000 and object O can be determined.
- control unit 1030 identifies in the second case, i.e. projector on, based on the temporal and spatial distribution of the events, which events were caused by which illumination pattern, and generates based on this identification and the known geometric relation of projector unit 1010 and receiver unit 1020 a depth map of the object O as the second information.
- the second information may contain in addition the event data E and the metadata M.
- any other method for obtaining second information can be used, as long as the second information represents somehow the fact that due to the usage of the illumination patterns knowledge about the distance between sensor device 1000 and object O has been introduced.
- the histograms showing event numbers over time for each pixel 1025 can be directly compared with histograms pre-derived for specific illumination conditions and object distances. From this comparison, the distance can be directly deduced, if the measured histogram matches one of the pre-derived histograms.
- the comparison can be done by a neural network that has been trained based on simulated results of illuminating object at varying distances with varying illumination patterns.
- the first and the second information are then provided to a predictor module 1031 of the control unit 1030.
- the predictor module 1031 generates the depth information based on both the first information and the second information.
- the predetermined time interval that is used to determine the depth information will most often contain both, given time periods without projection of illumination patterns, and given time periods with projection of illumination patterns.
- the predictor module 1031 gathers the information generated for each of these given time periods and provides depth information upon input of either first or second information. In this manner, it is e.g. possible to provide depth maps also for time instances at which the projector unit 1010 is turned off, by updating the depth maps contained in the second information based on the first information. Moreover, it is possible to derive more than the pure distance information from the depth maps.
- the depth information may also include such additional information.
- the predictor unit 1031 may recognize a specific gesture or sign made with the hand. In this manner, refined information can be obtained without increase of the production cost or the energy consumption.
- a receiving unit 1020 in which all pixels 1025 are in principle capable to receive illumination light and ambient light.
- a receiver unit 1020 that comprises first pixels 1025a that detect only intensities of light reflected from the object O that stem from illumination with ambient light and second pixels 1025b that detect only intensities of light reflected from the object O that stem from the illumination with the illumination patterns.
- the first pixels 1025a may be provided with color filters that only transmit ambient light
- the second pixels 1025b are provided with color filters that transmit only the illumination light.
- first pixels 1025a will always contribute to the generation of the first information
- second pixels 1025b will always contribute to the generation of second information.
- first and second information may also be provided in parallel to the predictor module 1031. If pixels 1025 capable to operate based on illumination light and based on ambient light are present, these pixels 1025 will alternatively contribute to the generation of first and second information as explained above. In this manner, errors occurring due to interpreting ambient light generated events as illumination light generated events can be avoided or at least suppressed.
- the predictor module 1031 may comprise different predictor submodules 1035, 1036, 1037 that are configured to derive state variables of a given time step based on state variables of a previous time step, the first information, and the second information, where the state variables describe the state of the object (O). Further, the predictor module 1031 comprises at least one task submodule 1038 that is configured to generate for each time step the depth information from the state variables of the respective time step.
- the different predictor submodules 1035, 1036, 1037 receive information regarding the, or derived from the detected events as described above. Based on this information the predictor submodules 1035, 1036, 1037 set state variables Zk for each instance of time k. This is done in an iterating manner, i.e. previous state variables Zk-i of the previous time k-1 are updated based on the most recent information obtained by the predictor submodules 1035, 1036, 1037.
- One predictor module may operate consecutively on the information for a given time, and only sparsely another predictor submodule will take over.
- the state variables Zk are used as input for the at least one task submodule 1038 which derives the desired depth information from the state variable Zk for time k.
- different task submodules 1038 may be used at different times if the output of different depth information is desired.
- predictor submodule 1035, 1036, 1037 for generating and updating state variables and task submodules for generating depth information based on the state variables allows for a most flexible implementation.
- the predictor submodules 1035, 1036, 1037 can be optimized for the function of translating information on the/derived from the event data into state variables that describe the object O and its position in space.
- the task module(s) 1038 can in turn be optimized to generate the desired depth information from the state variables.
- each of the predictor submodules 1035, 1036, 1037 and the task submodule(s) 1038 may be configured by in principle known neural networks. These neural networks can be trained to optimize their respective functions, which allows to use highly specific networks for different duties. Moreover, it is even possible to let the neural networks decide during the training process which kind of state variables to use. This means that the question which predictor submodules 1035, 1036, 1037 are used, will influence the form of the state variables. Just the same, also the desired depth information generated by the task submodule(s) 1038 will influence the form of the state variables if a modification is allowed in order to optimize the input of the task submodule(s).
- CNN neural network convolutional neural networks
- transformer architectures may be used as neural network convolutional neural networks. These operate on fixed size tensors or tokens as state variables, respectively. They can be trained by end-to-end supervised learning with data generated by a simulator environment that can generate sensor data for a variety of geometric arrangements of projector unit 1010 and receiver unit 1020, of illumination patterns, and of observed scenes. Since such neural networks are in principle known, a detailed description of their functioning can be omitted here. What is crucial in the present context is that different types of information on/derived from event data can be processed by the neural networks such as to provide state variables, which can be updated also by neural networks operating on different information, and which can further be used by the neural network(s) constituting the task submodule(s) to generate the depth information.
- the predictor submodules may comprise a first predictor submodule 1035 that is used to derive the state variables, if first information is provided to the predictor module 1031, and a second predictor submodule 1036 that is used to derive the state variables, if second information is provided to the predictor module 1031.
- First information will basically consist of event data (plus optionally metadata).
- the first information contains basically two-dimensional information and can be most efficiently used for two-dimensional shape recognition or object classification.
- the first information may be sufficient to identify human hands or faces in the observed scene.
- the size of the identified two-dimensional features can then already be an indicator for a distance of these features from sensor device 1000, in particular if the feature has a standard size.
- the second information contains basically a depth map of the object O, i.e. three-dimensional information. This can on the one hand be used as a basic information regarding a distance of an object O that is then updated based on the two-dimensional first information.
- the depth map contained in the second information can also be used to correct and thus refine state variables derived based on the first information, in particular, if for a certain time period only first information is available.
- First and second information therefore support each other such that omission of the second information, i.e. switch off of the projector for power saving, does not become critical for the accuracy of the state variables.
- the state variables might be constituted by a mere depth map
- the state variables may also take a form that is not as easy to understand for a human as a depth map.
- the state variables are set such as to optimize the processing and will most often have the form of mere datasets that do not allow a direct deduction of the meaning encoded therein.
- the state variables may constitute such a depth map. But if more information is requested, as e.g. the orientation of an object (such as a hand) in space, the recognition of a specific gesture, the classification of a facial expression, or the like, the state variables will take a form that makes processing of this request most reliable and fast.
- the receiver unit 1020 may not only be configured to generate events but may also be configured to generate for each pixel 1025 intensity information indicating the intensity of the light reflected from the object O, i.e. to generate a normal RGB or grayscale frame image of the observed scene.
- the predictor submodules 1035, 1036, 1037 may comprise a third predictor submodule 1037 that is used to derive state variables of a given time step based on state variables of a previous time step and the intensity information.
- image data will only be available from time to time.
- the image data can e.g. help to identify fine textures that may be helpful in the generation of depth maps or the determination of an orientation of an object O.
- the image data contain two-dimensional information that can be used to support the generation of depth information just as the event data stemming from ambient light.
- image data can make the process more accurate and reliable since a further source of information is included.
- the control unit 1030 may be configured to turn the projector unit 1010 on and off and/or to control the projection solid angle PS of the projector unit 1010 based on the depth information.
- the depth information may be generated based on the two- dimensional information obtainable via ambient light. This makes the projection of further illumination patterns superfluous for a certain time period. Only, if the control unit 1030 determines that the accuracy and/or reliability of the depth information (depth map, three-dimensional orientation, gesture classification or the like) is no longer good enough, the projector unit 1010 is turned on again.
- a coarse depth map may be sufficient.
- the control unit 1030 may decide that the quality of the depth information needs to be improved/may worsen and may control the projector unit 1010 accordingly.
- the control unit 1030 will see to project illumination patterns only to this region.
- the sensor device 1000 comprises only a single receiver unit 1020, which is used to generate depth information basically due to the distortion of the form of a known illumination pattern.
- the above concept can be extended to the use of a stereo camera for depth estimation by adding a further receiver unit 1040 that is arranged a predetermined distance away from the receiver unit 1020 and that has the same functions as the receiver unit 1020.
- the presence of the projector unit 1010 makes the sensor device 1000 to an active stereo depth sensing device.
- This active stereo depth sensing device is in principle capable to determine depth maps on the one hand in a passive manner by using ambient light and determining the parallax of ambient light images captured with both receiver units 1020, 1040.
- the projector unit 1010 is used for active stereo sensing, i.e. the changes of distortion of the known illumination patterns captured by the two receiver units 1020, 1040 can be used to determine the distance to the object O based on triangulation.
- the active stereo depth sensing device is capable to perform both modes of depth sensing simultaneously, even if illumination light and ambient light have different wavelengths.
- the active stereo depth sensing device is capable to detect events as described above.
- the control unit 1030 is then configured to determine a disparity between the events generated by the receiver unit 1020 and the events generated by the further receiver unit 1040 and to generate based on the disparity and the known geometric relation G of the projector unit 1010, the receiver unit 1020, and the further receiver unit 1040 a depth map of the object O.
- the control unit 1030 comprises further a predictor module 1031 that is configmed to generate the depth information based on the depth map and the events used to generate the depth map.
- Fig. 10 This is schematically illustrated in Fig. 10.
- all events E from both receiver units 1020, 1040 are input into the control unit 1030.
- the control unit 1030 determines with a disparity estimator module 1039 the disparity between the events generated by the different receiver units 1020, 1040 using in principle known methods. For example, event sequences are identified and the spatial shift in the frame coordinates of the respective pixels 1025 is determined as parallax information.
- the disparity information is then input into the triangulation module 1034 that uses the known geometric arrangement G to determine from the disparity, e.g. parallax information, the distance between the object O and the sensor device 1000.
- This depth map is then concatenated with the event data and (optionally) metadata and provided to the estimator module 1031 that generates depth information from it in the maimer described above.
- Fig. 10 could be modified in the manner of Fig. 6 such that while the projector unit 1010 does not operate, depth map generation is performed based on ambient light, while depth map generation is performed based on the illumination patterns if the projector unit 1010 operates. This might be used to switch on the pattern projection only for scenes showing little texture and/or no motion, i.e. in situations where passive stereo depth sensing is not possible.
- the filed of view of the receiver units 1020, 1040 may not overlap in all regions of the observed object O. This means that no depth map can be established in these regions. However, the depth of the neighboring regions with overlapping field of view can be used to propagate the known depth to these regions.
- illumination patterns may be repeated after a given number of solid angles.
- this ambiguity can be resolved by the control unit 1030 by recurring to the fact that only for one match a depth map showing a meaningful result will be generated. For example, if the desired depth information relates to the position of a hand in three-dimensional space, it can be checked which of the generated depth maps will show a hand, thus eliminating wrong matches.
- ambiguities can be resolved, which allows usage of repeating illumination patterns. This makes the design of the projector unit 1010 simpler.
- Figs. 12A and 12B show schematically camera devices 2000 that comprise the sensor device 1000 described above.
- the camera device 2000 is configured to generate depth information on a captured scene containing the object O in the manner described above.
- Fig. 12A shows a smart phone that is used to obtain depth information such as a depth map of an object O. This might be used to improve augmented reality functions of the smart phone or to enhance game experiences available on the smart phone.
- Fig. 12B shows a face capture sensor that might be used e.g. for face recognition at airports or boarder control, for viewpoint correction or artificial makeup in web meetings, or to animate chat avatars for web meeting or gaming. Further, movie/animation creators might use such an EVS-enhanced face capture sensor to adapt animated figures to real live persons.
- Fig. 13 shows as further example a head mounted display 3000 that comprises a sensor device 1000 as described above, wherein the head mounted display 3000 is configured to generate depth information of an object O viewed through the head mounted display 3000 as described above.
- This example might be used for accurate hand tracking or gesture recognition in augmented reality or virtual reality applications, e.g. in aiding complicated medical tasks.
- Fig. 14 shows schematically an industrial production device 4000 that comprises a sensor device 1000 as described above, wherein the industrial production device 4000 comprises means 4010 to move objects O in front of the projector unit 1010 in order to (partly) achieve the projection of the illumination pattern onto different locations of the objects O, and the industrial production device 4000 is configured to generate depth information for the objects O based on the positions of the images of the illumination patterns.
- This application is particularly adapted to EVS-enhanced depth sensors, since conveyor belts constituting e.g. the means 4010 to move objects O have a high movement speed that allows generation of depth information only if the receiver unit 1020 has a sufficiently high time resolution.
- the depth information may contain a depth map and/or a classification of object position on the conveyor belt, information on deviations from desired production standards, error classification and the like.
- Fig. 15 summarizes the steps of the method for generating depth information for an object O with a sensor device 1000 described above.
- the method for operating a sensor device 1000 for generating depth information for the object O comprises:
- Fig. 16 is a perspective view showing an example of a laminated structure of a solid-state imaging device 23020 with a plurality of pixels arranged matrix-like in array form in which the functions described above may be implemented.
- Each pixel includes at least one photoelectric conversion element.
- the solid-state imaging device 23020 has the laminated structure of a first chip (upper chip) 910 and a second chip (lower chip) 920.
- the laminated first and second chips 910, 920 may be electrically connected to each other through TC(S)Vs (Through Contact (Silicon) Vias) formed in the first chip 910.
- the solid-state imaging device 23020 may be formed to have the laminated structure in such a manner that the first and second chips 910 and 920 are bonded together at wafer level and cut out by dicing.
- the first chip 910 may be an analog chip (sensor chip) including at least one analog component of each pixel, e.g., the photoelectric conversion elements arranged in array form.
- the first chip 910 may include only the photoelectric conversion elements.
- the first chip 910 may include further elements of each photoreceptor module.
- the first chip 910 may include, in addition to the photoelectric conversion elements, at least some or all of the n- channel MOSFETs of the photoreceptor modules.
- the first chip 910 may include each element of the photoreceptor modules.
- the first chip 910 may also include parts of the pixel back-ends 300.
- the first chip 910 may include the memory capacitors, or, in addition to the memory capacitors sample/hold circuits and/or buffer circuits electrically connected between the memory capacitors and the event-detecting comparator circuits.
- the first chip 910 may include the complete pixel back-ends.
- the first chip 910 may also include at least portions of the readout circuit 140, the threshold generation circuit 130 and/or the controller 120 or the entire control unit.
- the second chip 920 may be mainly a logic chip (digital chip) that includes the elements complementing the circuits on the first chip 910 to the solid-state imaging device 23020.
- the second chip 920 may also include analog circuits, for example circuits that quantize analog signals transferred from the first chip 910 through the TCVs.
- the second chip 920 may have one or more bonding pads BPD and the first chip 910 may have openings OPN for use in wire-bonding to the second chip 920.
- the solid-state imaging device 23020 with the laminated structure of the two chips 910, 920 may have the following characteristic configuration:
- the electrical connection between the first chip 910 and the second chip 920 is performed through, for example, the TCVs.
- the TCVs may be arranged at chip ends or between a pad region and a circuit region.
- the TCVs for transmitting control signals and supplying power may be mainly concentrated at, for example, the four comers of the solid-state imaging device 23020, by which a signal wiring area of the first chip 910 can be reduced.
- the first chip 910 includes a p-type substrate and formation of p-channel MOSFETs typically implies the formation of n-doped wells separating the p-type source and drain regions of the p-channel MOSFETs from each other and from further p-type regions. Avoiding the formation of p-channel MOSFETs may therefore simplify the manufacturing process of the first chip 910.
- Fig. 17 illustrates schematic configuration examples of solid- state imaging devices 23010, 23020.
- the single-layer solid-state imaging device 23010 illustrated in part A of Fig. 17 includes a single die (semiconductor substrate) 23011. Mounted and/or formed on the single die 23011 are a pixel region 23012 (photoelectric conversion elements), a control circuit 23013 (readout circuit, threshold generation circuit, controller, control unit), and a logic circuit 23014 (pixel back-end). In the pixel region 23012, pixels are disposed in an array form.
- the control circuit 23013 performs various kinds of control including control of driving the pixels.
- the logic circuit 23014 performs signal processing.
- Parts B and C of Fig. 17 illustrate schematic configuration examples of multi-layer solid-state imaging devices
- first chip and a logic die 23024 (second chip), are stacked in a solid-state imaging device 23020. These dies are electrically connected to form a single semiconductor chip.
- the pixel region 23012 and the control circuit 23013 are formed or mounted on the sensor die 23021, and the logic circuit 23014 is formed or mounted on the logic die 23024.
- the logic circuit 23014 may include at least parts of the pixel back-ends.
- the pixel region 23012 includes at least the photoelectric conversion elements.
- the pixel region 23012 is formed or mounted on the sensor die 23021, whereas the control circuit 23013 and the logic circuit 23014 are formed or mounted on the logic die 23024.
- the pixel region 23012 and the logic circuit 23014, or the pixel region 23012 and parts of the logic circuit 23014 may be formed or mounted on the sensor die 23021, and the control circuit 23013 is formed or mounted on the logic die 23024.
- all photoreceptor modules PR may operate in the same mode.
- a first subset of the photoreceptor modules PR may operate in a mode with low SNR and high temporal resolution and a second, complementary subset of the photoreceptor module may operate in a mode with high SNR and low temporal resolution.
- the control signal may also not be a function of illumination conditions but, e.g., of user settings.
- the technology according to the present disclosure may be realized, e.g., as a device mounted in a mobile body of any type such as automobile, electric vehicle, hybrid electric vehicle, motorcycle, bicycle, personal mobility, airplane, drone, ship, or robot.
- Fig. 18 is a block diagram depicting an example of schematic configuration of a vehicle control system as an example of a mobile body control system to which the technology according to an embodiment of the present disclosure can be applied.
- the vehicle control system 12000 includes a plurality of electronic control units connected to each other via a communication network 12001.
- the vehicle control system 12000 includes a driving system control unit 12010, a body system control unit 12020, an outside-vehicle information detecting unit 12030, an in-vehicle information detecting unit 12040, and an integrated control unit 12050.
- a microcomputer 12051, a sound/image output section 12052, and a vehicle-mounted network interface (I/F) 12053 are illustrated as a functional configuration of the integrated control unit 12050.
- the driving system control unit 12010 controls the operation of devices related to the driving system of the vehicle in accordance with various kinds of programs.
- the driving system control unit 12010 functions as a control device for a driving force generating device for generating the driving force of the vehicle, such as an internal combustion engine, a driving motor, or the like, a driving force transmitting mechanism for transmitting the driving force to wheels, a steering mechanism for adjusting the steering angle of the vehicle, a braking device for generating the braking force of the vehicle, and the like.
- the body system control unit 12020 controls the operation of various kinds of devices provided to a vehicle body in accordance with various kinds of programs.
- the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various kinds of lamps such as a headlamp, a backup lamp, a brake lamp, a turn signal, a fog lamp, or the like.
- radio waves transmitted from a mobile device as an alternative to a key or signals of various kinds of switches can be input to the body system control unit 12020.
- the body system control unit 12020 receives these input radio waves or signals, and controls a door lock device, the power window device, the lamps, or the like of the vehicle.
- the outside-vehicle information detecting unit 12030 detects information about the outside of the vehicle including the vehicle control system 12000.
- the outside-vehicle information detecting unit 12030 is connected with an imaging section 12031.
- the outside-vehicle information detecting unit 12030 makes the imaging section 12031 imaging an image of the outside of the vehicle, and receives the imaged image.
- the outside-vehicle information detecting unit 12030 may perform processing of detecting an object such as a human, a vehicle, an obstacle, a sign, a character on a road surface, or the like, or processing of detecting a distance thereto.
- the imaging section 12031 may be or may include a solid-state imaging sensor with event detection and photoreceptor modules according to the present disclosure.
- the imaging section 12031 may output the electric signal as position information identifying pixels having detected an event.
- the light received by the imaging section 12031 may be visible light, or may be invisible light such as infrared rays or the like.
- the in-vehicle information detecting unit 12040 detects information about the inside of the vehicle and may be or may include a solid-state imaging sensor with event detection and photoreceptor modules according to the present disclosure.
- the in-vehicle information detecting unit 12040 is, for example, connected with a driver state detecting section 12041 that detects the state of a driver.
- the driver state detecting section 12041 for example, includes a camera focused on the driver.
- the in-vehicle information detecting unit 12040 may calculate a degree of fatigue of the driver or a degree of concentration of the driver, or may determine whether the driver is dozing.
- the microcomputer 12051 can calculate a control target value for the driving force generating device, the steering mechanism, or the braking device on the basis of the information about the inside or outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040, and output a control command to the driving system control unit 12010.
- the microcomputer 12051 can perform cooperative control intended to implement functions of an advanced driver assistance system (ADAS) which functions include collision avoidance or shock mitigation for the vehicle, following driving based on a following distance, vehicle speed maintaining driving, a warning of collision of the vehicle, a warning of deviation of the vehicle from a lane, or the like.
- ADAS advanced driver assistance system
- the microcomputer 12051 can perform cooperative control intended for automatic driving, which makes the vehicle to travel autonomously without depending on the operation of the driver, or the like, by controlling the driving force generating device, the steering mechanism, the braking device, or the like on the basis of the information about the outside or inside of the vehicle which information is obtained by the outsidevehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040.
- the microcomputer 12051 can output a control command to the body system control unit 12020 on the basis of the information about the outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030.
- the microcomputer 12051 can perform cooperative control intended to prevent a glare by controlling the headlamp so as to change from a high beam to a low beam, for example, in accordance with the position of a preceding vehicle or an oncoming vehicle detected by the outsidevehicle information detecting unit 12030.
- the sound/image output section 12052 transmits an output signal of at least one of a sound or an image to an output device capable of visually or audible notifying information to an occupant of the vehicle or the outside of the vehicle.
- an audio speaker 12061, a display section 12062, and an instrument panel 12063 are illustrated as the output device.
- the display section 12062 may, for example, include at least one of an on-board display or a head-up display.
- Fig. 19 is a diagram depicting an example of the installation position of the imaging section 12031, wherein the imaging section 12031 may include imaging sections 12101, 12102, 12103, 12104, and 12105.
- the imaging sections 12101, 12102, 12103, 12104, and 12105 are, for example, disposed at positions on a front nose, side-view mirrors, a rear bumper, and a back door of the vehicle 12100 as well as a position on an upper portion of a windshield within the interior of the vehicle.
- the imaging section 12101 provided to the front nose and the imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle obtain mainly an image of the front of the vehicle 12100.
- the imaging sections 12102 and 12103 provided to the side view mirrors obtain mainly an image of the sides of the vehicle 12100.
- the imaging section 12104 provided to the rear bumper or the back door obtains mainly an image of the rear of the vehicle 12100.
- the imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle is used mainly to detect a preceding vehicle, a pedestrian, an obstacle, a signal, a traffic sign, a lane, or the like.
- Fig. 19 depicts an example of photographing ranges of the imaging sections 12101 to 12104.
- An imaging range 12111 represents the imaging range of the imaging section 12101 provided to the front nose.
- Imaging ranges 12112 and 12113 respectively represent the imaging ranges of the imaging sections 12102 and 12103 provided to the side view mirrors.
- An imaging range 12114 represents the imaging range of the imaging section 12104 provided to the rear bumper or the back door.
- a bird's-eye image of the vehicle 12100 as viewed from above is obtained by superimposing image data imaged by the imaging sections 12101 to 12104, for example.
- At least one of the imaging sections 12101 to 12104 may have a function of obtaining distance information.
- at least one of the imaging sections 12101 to 12104 may be a stereo camera constituted of a plurality of imaging elements, or may be an imaging element having pixels for phase difference detection.
- the microcomputer 12051 can determine a distance to each three-dimensional object within the imaging ranges 12111 to 12114 and a temporal change in the distance (relative speed with respect to the vehicle 12100) on the basis of the distance information obtained from the imaging sections 12101 to 12104, and thereby extract, as a preceding vehicle, a nearest three-dimensional object in particular that is present on a traveling path of the vehicle 12100 and which travels in substantially the same direction as the vehicle 12100 at a predetermined speed (for example, equal to or more than 0 km/hour). Further, the microcomputer 12051 can set a following distance to be maintained in front of a preceding vehicle in advance, and perform automatic brake control (including following stop control), automatic acceleration control (including following start control), or the like. It is thus possible to perform cooperative control intended for automatic driving that makes the vehicle travel autonomously without depending on the operation of the driver or the like.
- automatic brake control including following stop control
- automatic acceleration control including following start control
- the microcomputer 12051 can classify three-dimensional object data on three-dimensional objects into three-dimensional object data of a two-wheeled vehicle, a standard-sized vehicle, a large-sized vehicle, a pedestrian, a utility pole, and other three-dimensional objects on the basis of the distance information obtained from the imaging sections 12101 to 12104, extract the classified three-dimensional object data, and use the extracted three-dimensional object data for automatic avoidance of an obstacle.
- the microcomputer 12051 identifies obstacles around the vehicle 12100 as obstacles that the driver of the vehicle 12100 can recognize visually and obstacles that are difficult for the driver of the vehicle 12100 to recognize visually. Then, the microcomputer 12051 determines a collision risk indicating a risk of collision with each obstacle.
- the microcomputer 12051 In a situation in which the collision risk is equal to or higher than a set value and there is thus a possibility of collision, the microcomputer 12051 outputs a warning to the driver via the audio speaker 12061 or the display section 12062, and performs forced deceleration or avoidance steering via the driving system control unit 12010. The microcomputer 12051 can thereby assist in driving to avoid collision.
- At least one of the imaging sections 12101 to 12104 may be an infrared camera that detects infrared rays.
- the microcomputer 12051 can, for example, recognize a pedestrian by determining whether or not there is a pedestrian in imaged images of the imaging sections 12101 to 12104. Such recognition of a pedestrian is, for example, performed by a procedure of extracting characteristic points in the imaged images of the imaging sections 12101 to 12104 as infrared cameras and a procedure of determining whether or not it is the pedestrian by performing pattern matching processing on a series of characteristic points representing the contour of the object.
- the sound/image output section 12052 controls the display section 12062 so that a square contour line for emphasis is displayed so as to be superimposed on the recognized pedestrian.
- the sound/image output section 12052 may also control the display section 12062 so that an icon or the like representing the pedestrian is displayed at a desired position.
- the image data transmitted through the communication network may be reduced and it may be possible to reduce power consumption without adversely affecting driving support.
- embodiments of the present technology are not limited to the above-described embodiments, but various changes can be made within the scope of the present technology without departing from the gist of the present technology.
- the solid-state imaging device may be any device used for analyzing and/or processing radiation such as visible light, infrared light, ultraviolet light, and X-rays.
- the solid-state imaging device may be any electronic device in the field of traffic, the field of home appliances, the field of medical and healthcare, the field of security, the field of beauty, the field of sports, the field of agriculture, the field of image reproduction or the like.
- the solid-state imaging device may be a device for capturing an image to be provided for appreciation, such as a digital camera, a smart phone, or a mobile phone device having a camera function.
- the solid-state imaging device may be integrated in an in- vehicle sensor that captures the front, rear, peripheries, an interior of the vehicle, etc. for safe driving such as automatic stop, recognition of a state of a driver, or the like, in a monitoring camera that monitors traveling vehicles and roads, or in a distance measuring sensor that measures a distance between vehicles or the like.
- the solid-state imaging device may be integrated in any type of sensor that can be used in devices provided for home appliances such as TV receivers, refrigerators, and air conditioners to capture gestures of users and perform device operations according to the gestures. Accordingly the solid-state imaging device may be integrated in home appliances such as TV receivers, refrigerators, and air conditioners and/or in devices controlling the home appliances. Furthermore, in the field of medical and healthcare, the solid- state imaging device may be integrated in any type of sensor, e.g. a solid-state image device, provided for use in medical and healthcare, such as an endoscope or a device that performs angiography by receiving infrared light.
- a solid-state image device provided for use in medical and healthcare, such as an endoscope or a device that performs angiography by receiving infrared light.
- the solid-state imaging device can be integrated in a device provided for use in security, such as a monitoring camera for crime prevention or a camera for person authentication use.
- the solid-state imaging device can be used in a device provided for use in beauty, such as a skin measuring instrument that captures skin or a microscope that captures a probe.
- the solid- state imaging device can be integrated in a device provided for use in sports, such as an action camera or a wearable camera for sport use or the like.
- the solid-state imaging device can be used in a device provided for use in agriculture, such as a camera for monitoring the condition of fields and crops.
- the present technology can also be configured as described below:
- a sensor device for generating depth information for an object comprising: a projector unit that is configured to project in a temporally consecutive manner a plurality of different illumination patterns in a projection solid angle to the object, where the projection solid angle consists of a predefined number of predetermined solid angles and each illumination pattern is generated by deciding for each of the predetermined solid angles whether or not to illuminate the respective predetermined solid angle by projecting light into it; a receiver unit that comprises a plurality of pixels, the receiver unit being configured to detect on each pixel intensities of light reflected from the object stemming from the illumination with the illumination patterns and/or from illumination with ambient light, and to generate an event at one of the pixels if the intensity detected at the pixel changes by more than a predetermined threshold; and a control unit that is configured to generate the depth information based on all events generated during a predetermined time period.
- control unit is configured to determine whether or not the projector unit projected light during a given time period into the projection solid angle; if, as a first case, the control unit determines that no light has been projected into the projection solid angle during the given time period, the control unit is configured to generate first information by processing the events on the assumption that all events generated during the given time period are caused by ambient light; if, as a second case, the control unit determines that light has been projected into the projection solid angle during the given time period, the control unit is configured to generate second information by processing the events on the assumption that all events generated during the given time period are caused by light projected by the projector unit; and the control unit comprises a predictor module that is configured to generate the depth information based on both the first information and the second information.
- the receiver unit comprises first pixels that detect only intensities of light reflected from the object (O) that stem from illumination with ambient light and second pixels that detect only intensities of light reflected from the object (O) that stem from the illumination with the illumination patterns; and the control unit is configured to generate the first information based on events generated by the first pixels and to generate the second information based on events generated by the second pixels.
- control unit is configured to identify based on the temporal and spatial distribution of the events, which events were caused by which illumination pattern, and to generate based on this identification and the known geometric relation of projector unit and receiver unit a depth map of the object as the second information.
- the predictor module comprises predictor submodules that are configured to derive state variables of a given time step based on state variables of a previous time step, the first information, and the second information, the state variables describing the state of the object; and the predictor module comprises at least one task submodule that is configured to generate for each time step the depth information from the state variables of the respective time step.
- the predictor submodules comprise a first predictor submodule that is used to derive the state variables, if first information is provided to the predictor module, and a second predictor submodule that is used to derive the state variables, if second information is provided to the predictor module.
- the receiver unit is configured to generate for each pixel intensity information indicating the intensity of the light reflected from the object; and the predictor submodules comprise a third predictor submodule that is used to derive state variables of a given time step based on state variables of a previous time step and the intensity information.
- control unit is configured to turn the projector unit on and off and/or to control the projection solid angle of the projector unit based on the depth information.
- the sensor device according to any one of (1) to (10), further comprising a further receiver unit that is arranged a predetermined distance away from the receiver unit and that has the same functions as the receiver unit; wherein the control unit is configured to determine a disparity between the events generated by the receiver unit and the events generated by the further receiver unit and to generate based on the disparity and the known geometric relation of the projector unit, the receiver unit, and the further receiver unit a depth map of the object; and the control unit comprises a predictor module that is configured to generate the depth information based on the depth map and the events used to generate the depth map.
- a method for operating a sensor device for generating depth information for an object comprising: projecting in a temporally consecutive manner a plurality of different illumination patterns in a projection solid angle to the object, where the projection solid angle consists of a predefined number of predetermined solid angles and each illumination pattern is generated by deciding for each of the predetermined solid angles whether or not to illuminate the respective predetermined solid angle by projecting light into it; detecting on each pixel intensities of light reflected from the object stemming from the illumination with the illumination patterns and/or from illumination with ambient light, and to generate an event at one of the pixels if the intensity detected at the pixel changes by more than a predetermined threshold; and generating the depth information based on all events generated during a predetermined time period.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Length Measuring Devices By Optical Means (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP22214219 | 2022-12-16 | ||
| PCT/EP2023/080927 WO2024125892A1 (en) | 2022-12-16 | 2023-11-07 | Depth sensor device and method for operating a depth sensor device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4634612A1 true EP4634612A1 (de) | 2025-10-22 |
Family
ID=84537676
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23798977.7A Pending EP4634612A1 (de) | 2022-12-16 | 2023-11-07 | Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtung |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4634612A1 (de) |
| WO (1) | WO2024125892A1 (de) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10516876B2 (en) * | 2017-12-19 | 2019-12-24 | Intel Corporation | Dynamic vision sensor and projector for depth imaging |
| JP7549025B2 (ja) * | 2020-09-07 | 2024-09-10 | ファナック株式会社 | 三次元計測装置 |
| EP4314704A1 (de) * | 2021-03-29 | 2024-02-07 | Sony Semiconductor Solutions Corporation | Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtung |
-
2023
- 2023-11-07 WO PCT/EP2023/080927 patent/WO2024125892A1/en not_active Ceased
- 2023-11-07 EP EP23798977.7A patent/EP4634612A1/de active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024125892A1 (en) | 2024-06-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112640428B (zh) | 固态成像装置、信号处理芯片和电子设备 | |
| US11336860B2 (en) | Solid-state image capturing device, method of driving solid-state image capturing device, and electronic apparatus | |
| US12598403B2 (en) | Hybrid image and event sensing with rolling shutter compensation | |
| CN114424022A (zh) | 测距设备,测距方法,程序,电子装置,学习模型生成方法,制造方法和深度图生成方法 | |
| US20240259703A1 (en) | Sensor device and method for operating a sensor device | |
| US20250020455A1 (en) | Depth sensor device and method for operating a depth sensor device | |
| US20240323552A1 (en) | Solid-state imaging device and method for operating a solid-state imaging device | |
| EP4431869B1 (de) | Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtung | |
| EP4689549A1 (de) | Tiefensensorvorrichtung und verfahren zum betreiben einer tiefensensorvorrichtung | |
| WO2024125892A1 (en) | Depth sensor device and method for operating a depth sensor device | |
| US20260016289A1 (en) | Depth sensor device and method for operating a depth sensor device | |
| WO2022254792A1 (ja) | 受光素子およびその駆動方法、並びに、測距システム | |
| WO2025062888A1 (ja) | 光検出素子及びシステム | |
| US20250175716A1 (en) | Sensor device and method for operating a sensor device | |
| WO2024199931A1 (en) | Sensor device and method for operating a sensor device | |
| WO2025172318A1 (en) | Sensor device and method for operating a sensor device | |
| CN121942211A (zh) | 传感器装置以及用于操作传感器装置的方法 | |
| CN120380773A (zh) | 光检测装置及光检测装置的控制方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250603 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |