EP4673912A1 - Method and system for depth completion - Google Patents

Method and system for depth completion

Info

Publication number
EP4673912A1
EP4673912A1 EP24718887.3A EP24718887A EP4673912A1 EP 4673912 A1 EP4673912 A1 EP 4673912A1 EP 24718887 A EP24718887 A EP 24718887A EP 4673912 A1 EP4673912 A1 EP 4673912A1
Authority
EP
European Patent Office
Prior art keywords
downscaling
upscaling
data
depth
processing level
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24718887.3A
Other languages
German (de)
French (fr)
Inventor
Csaba BENEDEK
Tamás SZIRÁNYI
Örkény H. ZOVÁTHI
Zsolt JANKÓ
Marcell KÉGL
Lóránt KOVÁCS
Balázs PÁLFFY
Zoltán RÓZSA
József KÖVENDI
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hun Ren Szamitastechnikai Es Automatizalasi Kutatointezet
Original Assignee
Hun Ren Szamitastechnikai Es Automatizalasi Kutatointezet
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hun Ren Szamitastechnikai Es Automatizalasi Kutatointezet filed Critical Hun Ren Szamitastechnikai Es Automatizalasi Kutatointezet
Publication of EP4673912A1 publication Critical patent/EP4673912A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/77Retouching; Inpainting; Scratch removal
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10016Video; Image sequence
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10028Range image; Depth image; 3D point clouds
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20016Hierarchical, coarse-to-fine, multiscale or multiresolution image processing; Pyramid transform
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20021Dividing image into blocks, subimages or windows
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30248Vehicle exterior or interior
    • G06T2207/30252Vehicle exterior; Vicinity of vehicle
    • G06T2207/30261Obstacle

Definitions

  • the invention relates to a method and system for depth completion using neural network.
  • the invention also relates to a training method therefor.
  • BACKGROUND ART Accurate and dense depth map prediction is an essential problem in 3D scene understanding.
  • depth information is often obtained from commercially available lidar (light detection and ranging, may also be written as Lidar or LIDAR) sensors as they are able to map their environment in real-time by emitting multiple laser beams and receiving their returns.
  • lidar light detection and ranging
  • LIDAR Lidar or LIDAR
  • the data captured by lidars is often very sparse, while its characteristics may vary depending on the sensors' scanning technology.
  • depth completion algorithms need to estimate dense depth images from the lidar-acquired sparse range measurements.
  • RMB typically rotating multi-beam
  • lidar sensors e.g., Ouster OS1 or Velodyne Puck models; see ⁇ . Zováthi, B. Nagy, and C. Benedek, “Point cloud registration and change detection in urban environment using an onboard lidar sensor and MLS reference data,” Int. J. Appl. Earth Obs. Geoinf., vol.110, p.102767, 2022.
  • RMB multi-beam
  • ConvLSTM operation i.e. convolution Long-Short Term Memory operation, see also below
  • ConvLSTM operation is mentioned in a depth completion method in US 2023/0230269 A1 using also depth data and RGB images on its input.
  • optical images may not provide eligible information in cases of sudden illumination changes (D. Nazir, A. Pagani, M. Liwicki, D. Stricker, and M. Z. Afzal, “SemAttNet: Towards attention-based semantic aware guided depth completion,” IEEE Access, pp. 1 1, 2022.) or in low-light environments (M. F. F.
  • edges and other finely textured structures on the generated depth images are often missing, blurred or distorted (R. Li, D. Xue, Y. Zhu, H. Wu, J. Sun, and Y. Zhang, “Selfsupervised monocular depth estimation with frequency-based recurrent refinement,” IEEE Trans. Multimedia, pp.1 12, 2022. and A. Savkin, Y. Wang, S. Wirkert, N. Navab, and F. Tombari, “lidar upsampling with sliced wasserstein distance,” IEEE Rob. Autom. Lett., pp.1 8, 2022.).
  • R. Li D. Xue, Y.
  • the primary object of the invention is to provide a method and system for depth completion, which are free of the disadvantages of prior art approaches to the greatest possible extent.
  • the object of the invention is to produce a high-quality, dense and spatially precise point cloud stream from measurements of a single lidar device with non-repetitive scanning pattern.
  • the objects of the invention can be achieved by providing the method according to claim 1 and the system according to claim 6, as well as the training method according to claim 7.
  • Preferred embodiments of the invention are defined in the dependent claims.
  • Light detection and ranging (lidar) is a powerful technology to quickly obtain a large amount of accurate spatial data from the environment. However, by using currently available lidar sensors, we still face a trade-off between the temporal and the spatial resolution of the measurements.
  • Real time lidar sensors such as non-repetitive circular scanning (NRCS) lidars and rotating multi-beam (RMB) lidars enable us to observe and analyse dynamic scenes (which is needed e.g. by the perception of self driving cars, or in surveillance applications), however the recorded point cloud measurement sequences have a limited resolution either in space or in time, which facts induce hard challenges for the environment analysis capabilities of the available machine perception systems.
  • NRCS non-repetitive circular scanning
  • RMB rotating multi-beam
  • lidar device with non-repetitive scanning pattern are also capable of providing measurements for real-time scene analysis in robotics and autonomous driving, at a significantly lower cost compared to the RMB technology (C. Glennie and P. Hartzell, ”Accuracy assessment and calibration of low-cost autonomous lidar sensors,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol.
  • the laser beams cover a higher proportion (around 90%) of the FOV yielding high spatial measurement resolution.
  • the potential ego-motion of the lidar's platform (e.g., vehicle or robot) and the dynamic objects in the surrounding area induce various artifacts, such as blurred shapes of the observed vehicles, pedestrians or buildings, which phenomena complicate dynamic event analysis.
  • a proposed spatio-temporal (ST) deep network may also be called ST-DepthNet, see Fig.
  • the network operates in the range domain and expects as input multiple sparse depth maps captured by a NRCS lidar consecutively in time, with using a narrow (i.e., for example, 200 ms) integration window for each frame.
  • the network provides a dense, high-quality range image of the same FOV, which does not reflect the sensor's original scanning artifacts (i.e., visible trails of the circular scanning pattern).
  • the architecture of the proposed spatio-temporal deep network applied preferably in the invention was directly designed to exploit both spatial and temporal patterns in the input NRCS data for depth completion, by extending a U-Net-like architecture (O. Ronneberger, P. Fischer, and T.
  • our solution connects the last sparse input image to the output by a direct skip connection to force our proposed model to complete the original sparse, but precise range map instead of overwriting it with completely new values.
  • FIG.3 shows the flow chart of an embodiment of the method according to the invention
  • Figs.4A-4D show reconstructions with different prior art methods, as well as an embodiment of the invention, and the ground truth
  • Figs.5A-5F show results on real measurements from the LivoxBudapest test set
  • Figs.6A and 6B show a sparse measurement and a depth image completed by an embodiment of the invention
  • Figs.7A-7C show a depth image exported from the simulator, a pattern of the applied sensor, and a created realistic depth image
  • Figs.3 shows the flow chart of an embodiment of the method according to the invention
  • Figs.4A-4D show reconstructions with different prior art methods, as well as an embodiment of the invention, and the ground truth
  • Figs.5A-5F show results on real measurements from the LivoxBudapest test set
  • Figs.6A and 6B show a sparse measurement and a depth image completed by an embodiment of the invention
  • FIGS. 8A-8C show a depth image exported from the simulator, a mask corresponding to the sensor and a ground truth data generated with the mask
  • Figs.9A-9C show an actual input, a ground truth and a prediction
  • Figs.10A-10D show an actual input, a ground truth and two predictions
  • Figs.11A and 11B show a greyscaled depth image and a corresponding 3D point cloud
  • Fig.12A shows lidar and camera measurements at timestamp t-1
  • Fig.12B shows camera measurements at timestamp t and generated virtual lidar measurement to the given time moment
  • Fig.13 shows details of the virtual point cloud generated based on the earlier illustrated inputs
  • Fig.14 shows a flowchart of an embodiment for upsampling in time
  • Figs.14 shows a flowchart of an embodiment for upsampling in time
  • Figs.14 shows a flowchart of an embodiment for upsampling in time
  • Figs.14 shows
  • 15A-15C show lidar measurement at timestamp t-1, and camera measurements at timestamps t-1 and t
  • Fig.16 shows optical flow field from inputs of Figs.15A-15C
  • Fig.17 shows motion-in-depth results from inputs of Figs.15A-15C and Fig.
  • Fig.18 show a segmented ground from input of Figs.15A-15C
  • Fig.19 shows estimated displacements based on Figs.15A-15C
  • Fig.20 shows estimated virtual measurement at timestamp t, based on the inputs of Figs.15A-15C
  • Figs.21A-21I show illustration of generated point clouds for a first exemplary scene with different methods including an embodiment of the invention, together with the ground truth
  • Figs. 22A-22I show illustration of generated point clouds for a second exemplary scene with different methods including an embodiment of the invention, together with the ground truth
  • Figs.23A-23I show illustration of generated point clouds for a third exemplary scene with different methods including an embodiment of the invention, together with the ground truth.
  • Depth completion preferably comprises the followings: First, the consecutive measurements of the NRCS lidar are grouped to form discrete time frames, using a narrow, 200 ms integration window (up to 40% FOV coverage in each frame). Thereafter, within a frame the distances of the measured 3D field points from the sensor are assigned to corresponding pixels in a high-resolution range image. By each actual time frame, the last five collected depth images (covering together around 95% of the FOV) are fed to the proposed spatio-temporal depth completion network, which composes a high-quality range image as output, with almost 100% FOV coverage, also eliminating the motion blurring artifacts.
  • Range image generation Range images are widely used, compact representations of lidar-based depth measurements (see ⁇ . Zováthi, B. Nagy, and C. Benedek, “Point cloud registration and change detection in urban environment using an onboard lidar sensor and MLS reference data,” Int. J. Appl. Earth Obs. Geoinf., vol. 110, p. 102767, 2022.; L. Kovács, M. Kégl, and C. Benedek, ”Real-time foreground segmentation for surveillance applications in NRCS lidar sequences,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol.
  • the horizontal and vertical pixel coordinates represent the polar azimuth and elevation angles, while the pixel’s depth value encodes the distance of the corresponding point.
  • the parameters of the exemplary Livox AVIA state-of-the-art NRCS lidar sensor see also for the sensor, L. Kovács, M. Kégl, and C. Benedek, “Real-time foreground segmentation for surveillance applications in NRCS lidar sequences,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol. XLIII-B1-2022, pp.45 51, 052022.).
  • LSTM Long-Short Term Memory
  • a Conv2DLSTM cell operates similarly to a regular LSTM cell, with an extension that the input Xt, memory state Ct and final state Ht with their respective gates (it, ft, ot) are 3D tensors – with one temporal and two spatial dimensions – and both the spatial and recurrent transformations are convolutional (marked by * in Equation system (1)) and not element-wise (marked by ⁇ ), making it able to propagate spatio-temporal features: (1) Accordingly, these are available conv2DLSTM equations performed on the data blocks of a level in the U-net hierarchy from the first time instant until the last time instant.
  • ft, ot gates the value is calculated by the help of a ⁇ activation function (ft forget gate is applied on Ct-1 memory state).
  • W ⁇ and b ⁇ are parameters with ⁇ and ⁇ indices for input and output. These are calculated based on subequations 2-4 of equation (1) for subequation 1, as well as subequation 5.
  • subequation 5 the result of subequation 1 is needed.
  • the upscaling branch of our proposed network is purely two-dimensional, in order to accurately restore the single output image of our interest.
  • Preferably applied skip connections at each level are performed by recurrent pooling utilizing the last output of a Conv2DLSTM layer which represents features of the last 200 ms measurement (but through this, thanks to the application of Conv2DLSTM, the whole sequence, see also the code part in Section II.G. below).
  • operational step S24b the dimension changes to 5x32x200x200 since operational step S24b is a strided convolution (downscaling convolution; strided ConvLSTM, spatial and recurrent, as also operational step S24a, see below).
  • the channel number increases by the help of strided convolutions (here a stride of 2 is applied) and the dimensions of the data blocks 23a-23e is decreased compared to data blocks 25a-25e.
  • data blocks 21a-21e and data blocks 23a-23e three dots are taken for illustrative purposes, showing that a large number of downscaling processing levels can be applied (see also the upscaling branch 15 where the three dots have similar illustrative purposes; the numbers of downscaling and upscaling processing levels are always the same).
  • data blocks 21a-21e have the dimensions of 5x64x100x100 (as a result of a further strided convolution having stride of 2), so the data blocks 21a-21e and the data blocks 23a-23e could have been illustrated also as a single level, but in this case too few levels would have been illustrated for the U-net like structure, a higher number of levels is typical.
  • an upscaling processing level 30 is illustrated in upscaling branch 15.
  • the output range image 40 is generated by a convolution in operational step S13.
  • the last range image 10e is added (this is also a skip connection applied in an addition layer) to the output obtained from the convolution applied in operational step S13 in order to obtain the dense output range image 40 (in this respect, see also below).
  • the invention is described herebelow with reference to Fig. 3 illustrating an embodiment.
  • the invention is a method for depth completion, comprising the steps of - downscaling is performed (see operational steps S24a, S24b in Fig.
  • a downscaling processing level 20 is illustrated in Fig.3 of a downscaling branch of a processing unit (see the above description of Fig.3 for the branches 5, 15 of the processing unit 50) based on neural network with starting, at the highest downscaling processing level of the downscaling branch, on a plurality of initial first data blocks (the downscaling is started on data blocks 25a-25e) corresponding to respective input depth information images (preferably, input range images, see Fig.
  • the downscaling is continued on all of the data blocks kept for a level) to obtain respective one or more downscaled first data block of the respective downscaling processing level (this is illustrated in Fig. 3 for downscaling in operational step S24b between data blocks 25a-25e and data blocks 23a-23e), wherein, in case of having one or more downscaling processing level below the lowest downscaling processing level on which a convolution operation is combined with performing convolution Long- Short Term Memory operation (i.e.
  • a processed last data block outputted by an auxiliary convolution Long-Short Term Memory operation performed on the plurality of downscaled data blocks of the lowest downscaling processing level (the processed last data block – or by other name the output first data block, since it is a ‘first’ data block of the downscaling branch – is the last data block processed by the auxiliary convLSTM operation, preferably the only output of it; accordingly, on the contrary to the above-mentioned (base) convLSTM operations, which can output one (the last) or more data blocks, the auxiliary convLSTM operation outputs only a single data block) is forwarded (one of the above three possibilities: when there is only one downscaled first data block on the lowest downscaling processing level, then it is forwarded, and when the plurality of downscaled first data blocks is kept, there is two possibility) on the lowest downscaling processing level (i.e.
  • the method is started on inputs at the highest level of the downscaling branch with inputting the data blocks corresponding to input depth information images (may be called simply depth images) generated from respective point sets recorded by means of a lidar device with non-repetitive scanning pattern (it is noted that according to the above there are multiple data blocks on the input, i.e. the group of them comprises at least two data blocks).
  • depth information images may be called simply depth images
  • the group of them comprises at least two data blocks.
  • Fig.1 an exemplary total 1 s is illustrated, during which time a comparatively high FOV coverage can be reached.
  • Fig.1 also shows that for the exemplary selected first integration window, an approx.40% FOV coverage corresponds.
  • the 1 s mentioned also elsewhere is a typical length for which a good FOV coverage can be reached by this device independently from the number of integration windows included in this period.
  • a time period with a good FOV coverage is divided between a group of consecutive integration windows (the number of which has lower significance), i.e. a single depth completed output can be obtained from a plurality of inputs from this time period with good FOV coverage (in the invention, a single output is generated from more input).
  • the downscaling branch and upscaling branch is preferably symmetrical like in Fig.3 and in the code part of Section II.G. below (it is also noted that from the point of view of the changes of channel numbers, the code part is not symmetrical, see below).
  • stride factors may be chosen for the different levels.
  • branches are symmetrical, levels with same stride factor can easily be identified in this. It is noted that it is technically not excluded level pairs in the branches with the same stride, but the overall upscaling is obtained by different stride values than the overall downscaling. It is furthermore noted that ‘initial’ in the name of some data blocks is called with other words: input (of the level) or level-input. Moreover, ‘first’ and ‘second’ in the name of the respective data blocks shows that these correspond to the downscaling branch and upscaling branch, respectively. Thus, ‘first’ and ‘second’ could be ‘first- branch’ and ‘second-branch’, respectively.
  • a convolution operation applied on a respective downscaling processing level is preferably combined with performing convolution Long-Short Term Memory (LSTM) operation for all downscaling processing levels (for this embodiment, see Late fusion below). Accordingly, in this embodiment, applying of convolution LSTM operation is required for all downscaling processing levels, not at least on the highest downscaling processing level.
  • the illustrated embodiment is an embodiment in which (for the illustration of channel operations see the code part in Section II.G. below and also its interpretation herebelow) - on a downscaling processing level, a channel increasing (multiplying) convolution operation is performed (see e.g. operational step S11 in Fig.
  • channel is modified also on other levels) for increasing a channel number of a data block by a channel increasing factor (may be different for the levels) preferably on first data blocks of the initial group of the respective downscaling processing level before performing the downscaling convolution operation on the respective downscaling processing level, and - in the upscaling branch a respective – i.e. the same, but reverse direction change (i.e. demultiplication instead of multiplication with a factor) in channel number as in the above channel increasing convolution operation – channel decreasing (demultiplying) convolution operation is performed (see e.g. operational step S13 in Fig.
  • skip connections are performed in the embodiment illustrated in Fig. 3.
  • a skip connection is preferably performed after a respective upscaling processing level, adding to the upscaled second data block of the respective upscaling processing level (the possibilities are the same as in the base skip connection, i.e. the forwarding between the downscaling and upscaling branches) o in case of having a single initial first data block on the corresponding downscaling processing level (i.e.
  • a downscaling level with the same size values for the data blocks, and, if any, with the same channel number), the single initial first data block of the corresponding downscaling processing level, or o in case of having a plurality of initial first data blocks on the corresponding downscaling processing level, ⁇ the last of the plurality of initial first data blocks of the corresponding downscaling processing level, or ⁇ a processed last data block outputted by an auxiliary convolution Long- Short Term Memory operation performed on the plurality of initial first data blocks of the corresponding downscaling processing level.
  • a channel increasing (convolution) 1 to 8 is performed (accordingly, with a first exemplary channel increasing factor of 8), the output of which is x1, after that – on a further layer of the same level – a downscaling convolution of the size (this word can also be used for the data block, e.g.400x400, just as resolution, it can be also said that the block has dimensions) from 400 to 200 is done.
  • a channel increasing 8 to 32 is performed (a second exemplary channel increasing factor is 4), the output of which is x2.
  • a downscaling of the size from 200 to 100 is done.
  • a channel increasing 32 to 64 is performed (a third exemplary channel increasing factor is 2) with an output of x3 and a downscaling of the size from 100 to 50.
  • the application of an auxiliary convLSTM operation outputting a processed last data block is chosen for the exemplary size of 50x50 and channel number 64.
  • first an upscaling of the size from 50 to 100 is done, after that channel number of 64 is maintained with a respective convolution (channel decreasing is shifted for facilitating the skip connection, i.e.
  • a skip connection is performed by adding x1.
  • this skip connection there is a further channel decreasing convolution is done in the example of the code part to decrease the channel number from 8 to 1. This is the pair of the channel increasing 1 to 8.
  • the channel increasing/decreasing convolution operation and the downscaling convolution operation/upscaling operation are performed on different (consecutive) layers of a respective downscaling/upscaling processing level (see the respective lines of the code part in Section II.G. part below).
  • a convolution operation can be downscaling convolution or channel increasing convolution since this this is relevant only for the downscaling branch, in the invention convolution LSTM operation is applied only in this branch and it is not applied in upscaling branch
  • convolution LSTM operation can be combined with performing convolution LSTM operation in the following way: ⁇ with downscaling convolution: see Equation system (1) in which the operations applicable for certain pixels are applied only for every second pixels (in case for a stride of two); ⁇ with channel increasing convolution: in this case – as a result of applying different filters – the number of the channels are increased, thus, also the number of pixels is multiplied to the number of the channels (see Fig.3) on which the operations of Equation system (1) are to be performed.
  • the last of the input depth information images is added (cf. operational step S32) to an upscaled depth information image corresponding to the upscaled data block of the highest upscaling processing level to obtain the output depth information image.
  • depth information image preferably range image is applied. Accordingly, an upscaled depth information image is obtained from the highest upscaled data block (e.g. by decreasing the channel number of it, so these correspond to each other), and the output depth information image is obtained from the addition of the two other depth information images. Accordingly, in general, the output depth information image is obtained based on the upscaled second data block of the highest upscaling processing level.
  • the proposed spatio-temporal network (having the above detailed architecture) is responsible for learning and predicting a high-density range image using a sparse input range image sequence.
  • a proposed loss function L is preferably composed of three main components to address handling of the artifacts in a targeted way.
  • LL1 Loss L1 Loss
  • Our exemplary LivoxCARLA dataset consists of 11726 randomly sampled input- output range image pairs, from which 10000 were used for training, 500 as validation and 1226 for testing.
  • Validation is the process of monitoring the evolution of a model with unknown data during the learning process. Typically, we look at how the model improves its accuracy as learns on data unknown to the model/validation data. Testing is the process of examining the effectiveness of a ready-made, learnt model.
  • Each pair consists of 400 x 400 images: the input range images were generated with NRCS-characteristics by a Livox AVIA sensor model, in 200 ms integration windows (with ca.40% FOV coverage), while a high-resolution ground truth range image was sampled by each fifth input frame, which is matched to the recording of the input range images, the last of which corresponds to the ground truth. Accordingly, the data corresponding to the respective integration time windows is captured at the end of each integration window. Thus, the GT data is recorded at the same time as the data of the fifth integration window.
  • LivoxBudapest does not include GT data, it enables us to validate the effectiveness of the proposed algorithm in real environment, despite the fact that the network is purely trained on synthetic data. According to our tests, the trained method for depth completion performs well on the LivoxBudapest dataset.
  • the system according to the invention comprises a processing unit based on neural network, the processing unit having a downscaling branch and an upscaling branch (see as an embodiment processing unit 50 with downscaling branch 5 and upscaling 15 in Fig.3; we refer to Fig.3 also in connection with other details of the system according to the invention).
  • the downscaling branch is adapted for performing downscaling in a plurality of downscaling processing levels with starting, at the highest downscaling processing level of the downscaling branch, on a plurality of initial first data blocks corresponding to respective input depth information images each generated from a respective point set, wherein the respective point sets have been recorded in consecutive integration windows of a scene by means of a lidar device with non-repetitive scanning pattern, wherein o on the highest downscaling processing level or on more consecutive downscaling processing levels from the highest downscaling processing level a convolution operation applied on the respective downscaling processing level is combined with performing convolution Long-Short Term Memory (convLSTM) operation, o in a respective downscaling processing level, a downscaling convolution operation is performed by a downscaling factor (stride) on all of the one or more initial first data block of the respective downscaling processing level to obtain respective one or more downscaled first data block of
  • the processing unit may be realized by a computer or with a dedicated hardware thereof.
  • the branches of the processing unit may be implemented by software code running on the computer implementing the processing unit or also by the help of dedicated hardware units.
  • the computer implementing the processing unit handles the input depth information data and performing the relevant operations to get the depth completed output depth information data.
  • some other embodiments of the invention relate to a training method (it has been touched upon above).
  • the training method according to the invention comprises training of the processing unit based on neural network utilized in the method according to the invention and comprised in the system according to the invention.
  • a loss function is utilized for training which comprises - a pixel-level term, - a Structure Similarity Index Measure term, and - an edge loss term.
  • the training is performed by the help of a training dataset generated by the help of CARLA simulator (see in particular Section II.F.
  • the training dataset comprises pairs of - input depth information images generated with characteristics of the lidar device with non-repetitive scanning pattern, and - output depth information image extracted to be dense using a mask corresponding to the field of view of the lidar device with non-repetitive scanning pattern.
  • the ground truth data can be obtained by the help of the CARLA simulator, which gives a kind of ideal data.
  • IPi denotes the ith pixel of the image generated by the actual method
  • IGTi is the ith pixel of the corresponding GT image
  • the inverse depth-based errors focus on relative error improvements and may be more relevant in scenes with objects at varying distances.
  • 3D errors Besides range image based evaluation, we also compared the generated point clouds to the reference model in the 3D space. Let us denote the GT and a predicted point cloud by PGT and PP, and the number of points in PGT and PP by #PGT and #PP, respectively. We evaluate the quality of the predicted point cloud with respect to the GT data using the symmetric Normalized Chamfer Distance (NCD) and Normalized Median Distance (NMD) ( ⁇ . Zováthi, B. Nagy, and C.
  • NCD symmetric Normalized Chamfer Distance
  • NMD Normalized Median Distance
  • Fig.4A-4D show that fine structures (marked by green ellipse in Fig.4C) recognized by the proposed spatio-temporal deep network approach (Fig. 4C), but remained partly or fully unrecognized (merged to wall or background) by the reference IP-Basic++ (Fig.4A) and Sparse-to-Dense (Fig.4B) approaches with respect to the Ground Truth data (Fig.4D: in this case the ground truth data is the ideal data which corresponds to the time instant at the end of the time window covered by the proposed spatio-temporal deep network approach, i.e. to the status of the scene being frozen at the end of the time window, cf. with Fig.3).
  • Fig.4C – illustrating an embodiment shows much better resemblance to the ground truth shown in Fig.4D.
  • the LivoxBudapest test set contains three different scenarios: two pathway recordings from the city center (a boulevard and a narrow street), both around 1 km long, and a speedway section near the city, recorded on a path of around 3.5 km.
  • Fig.5A-5F displays selected relevant sample frames from the three scenarios.
  • Figs.5A-5F show results on real measurements from the LivoxBudapest test set.
  • Fig.5A corresponds to sparse measurements (sparse input data captured in a 200 ms time window), and the row of Fig. 5B corresponds to RGB images for visual reference only.
  • Fig.5C As for the Large integration method (Fig.5C), it performs significantly worse on real data than on the simulated samples: its generated range images are extremely noisy, and if the platform is moving, structures are barely recognizable. Note that by large integration time the regions of moving street objects become blurred even if the lidar platform is static.
  • IP-Basic++ approach images of Fig.5D
  • the Sparse-to-Dense method cannot eliminate the rosetta patterns of the input lidar measurements, which are typically visible on ground and wall areas (e.g., second column in Fig.5E).
  • the tendency of merging fine structures into larger surfaces is also notable: In the first and third columns of Fig.5E, vehicles and pedestrians are falsely merged to the wall behind them. Such artefacts can mean critical problems for urban scene understanding tasks, while as shown, they are handled better by the proposed spatio-temporal deep network approach (see regions marked by green ellipses in Fig.5F).
  • the Sparse-to-Dense method heavily blurs other fine structures (columns, traffic signs and lights, etc.) as well, as displayed in Fig.5E. Moreover, while objects close to the sensor are usually well recognizable for the human eye, they are often predicted at inaccurate distances with this method (e.g., cyclist in the second column of Fig.5E).
  • we conducted a survey for visual verification of the generated depth image streams by asking 20 computer vision related experts to rate the quality of the input and the output of each method in all three videos with scores between 1 and 9, where 9 is the best possible score.
  • the results provided in Table V confirm that the test experts found the proposed method significantly better than the reference techniques.
  • Source code is attached in Section G below.
  • F Datasets and results
  • the purpose of project connected to the invention is effective deep learning based depth completion for sparse measurements captured by the Livox AVIA sensor (displayed in Figs 6A-6B, showing a sparse measurement by the Livox AVIA sensor (Fig. 6A) and the completed depth image by the proposed spatio-temporal deep network approach (Fig.6B)).
  • the sensor specification sheet can be found in Section II.H. below.
  • We use training, validation and test data with ground truth information from the Carla simulator see above, Y. Kwon, M. Sung, and S. Yoon, ”Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob.
  • the dataset was generated from the Carla simulator that gives the opportunity to export perfect depth images without any distortion or blurring.
  • the dense depth images were sampled with rosetta scanning pattern of the Livox AVIA sensor.
  • the capturing platform (a simulated) vehicle was dynamically moving, and to augment on the extractable information (e.g., vary the ground level), the capturing sensor’s position was randomly rotated along the up axis by [ ⁇ 22.5°, 22.5°], and its height was randomly adjusted between [1.5m, 2.5m], see above.
  • the final dataset consists of 11726 randomly sampled input-output data pairs.
  • the data generation process is illustrated in Figs 7A-7C and 8A-8C, respectively.
  • the input data is a sparse depth image sequence, that consists of five consecutive sparse depth images, each sampled after 200 ms.
  • the patterns of the Livox AVIA sensor (displayed in Fig.7B) are used to filter the depth image exported from the simulator (Fig.7A) resulting in realistic, Livox-like depth images (Fig.7C).
  • Figs. 8A-8C it is illustrated that the output (ground truth) data in Fig. 8C is generated from the depth image exported from the simulator (Fig.8A) at the end of each input sequence using the mask (obtained from real measurements, see Fig.
  • a grayscale depth images can be converted to corresponding 3D point clouds, an example is shown from these in Figs.11A-11B, respectively.
  • G. Code part Herebelow, a code part is illustrated (with comments written with capitalized letters) for network architecture of the neural network utilized in the processing unit in an embodiment. This is a Python source code part.
  • the processing ReLU rectified linear unit
  • ReLU is an activation function defined as the positive part of its argument (if input is negative, then the output is zero, and, if the input is positive, then the output is equal to the input).
  • batch normalization is a method used to make training of artificial neural networks faster and more stable through normalization of the layers' inputs by re-centering and re-scaling. ReLU introduces nonlinearity to the processing in a usual way.
  • H Datasheet of the device applied in the experiments Some data (main parameters) from the datasheet of the Livox AVIA device applied in the experiments are given in Table VI.
  • the FOV value corresponds to a single sensor which “sees” in front of it within a cone which is a not regular cone according to the above data (see the different horizontal and vertical data). In the description, it is discussed how the coverage of this FOV builds up (cf. Fig.1).
  • the datasheet shows that the Livox AVIA sensor as an exemplary NRCS lidar has an approximately circular FOV.
  • range and angular precision This specifies how accurate the detected point is in distance (range) and laterally (angular). So how accurate it is that what we see is at this angle and at this distance.
  • the depth map is generated for a single (optical) image.
  • a depth data is generated for each (optical) image input.
  • the neural network based approach preferably utilizes sparse range image inputs.
  • the consecutive sparse range images have ca.40% FOV coverage each and the overlap in their content is relatively low; this is a result of the non-repetitiveness of the pattern in the FOV which is an inherent feature of the NRCS lidar device. Such a good results with convLSTM are thus not expected.
  • convLSTM works together very well with consecutive frames recorded by the NRCS lidar device giving a good coverage of the selected FOV altogether thanks to the non-repetitiveness.
  • the non-repetitiveness of the applied device results in highly advantageous performance.
  • the application of a rotating multi-beam lidar would be based on hugely different principles than the invention, since it is repetitive in its nature.
  • a point cloud of a rotating multi-beam lidar would not help the depth completion since – because of the repetitiveness – it would not yield information for depth completion: it scans the same field of view, the consecutive rotations bring the same information in this respect.
  • We hereby summarize the main preferred contributions of invention as follows: ⁇ We proposed a novel deep learning based solution (proposed spatio- temporal deep network approach), which in an embodiment extends the classical U-Net architecture with a spatio-temporal downscaling branch for utilizing consecutive sparse measurements captured by NRCS lidars. Our model produces spatially precise high-density depth data using a spatial upscaling branch following effective temporal pooling steps. ⁇ We provided a new synthetic urban dataset called LivoxCarla, which contains simulated NRCS lidar data with corresponding dense depth Ground Truth (GT) information.
  • GT Ground Truth
  • LivoxCarla a real-life dataset called LivoxBudapest, which contains real NRCS lidar measurement data collected in Budapest, Hungary, both in downtown and speedway areas, by a sensor mounted on a moving vehicle. Testing with the real measurements allows us to clearly demonstrate the usability of the method according to the invention trained on synthetic data in real-life urban scenarios.
  • proposed spatio-temporal deep network is composed of a spatio-temporally (ST) extended U-Net architecture, which accepts a very sparse range data sequence as input and produces a dense depth image stream of the same field-of-view ensuring a high level of spatial details and accuracy.
  • ST spatio-temporally
  • the pipeline first utilizes optical flow estimations from the available camera frames. Next, optical expansion (or, equivalently, motion in depth, see below) is used to upgrade it to 3D scene flow. Following that, ground plane fitting is made on the previous lidar point cloud. Finally, the estimated scene flow is applied to the previously measured object points (i.e. points corresponding to the objects, see also below) to generate the new point cloud.
  • the efficiency of the framework of the present embodiment is proven in performance compared to prior art achieved on the KITTI dataset using measurements of a Rotating Multi-Beam (RMB) lidar sensor.
  • RMB Rotating Multi-Beam
  • the steps of the present embodiment can be performed independently from the above detailed depth completion method according to the invention.
  • the initial depth data would not be connected to the output depth information image (formally, we would not start from initial depth data corresponding the output depth information image, but simply from a general initial depth data).
  • depth completed data originating from data recorded by a lidar device with non-repetitive scanning pattern will be utilized for these steps as an input.
  • initial depth data of any origin could be used as an input. It is in line with these statements, that herebelow the details are shown simply for a lidar and for initial depth data originating from any source.
  • any kind of depth data may be the input.
  • ADAS advanced driver-assistance systems
  • passive cameras and lidars are usually basic components of the sensor systems equipped for the given vehicle. In this way, sensor fusion (Nagy, B.; Benedek, C. On-the-Fly Camera and lidar Calibration.
  • Remote Sensing 2020, 12. is frequently applied by utilizing the benefits of both sensors to solve different problems like 3D object detection (Wu, X.; Peng, L.; Yang, H.; Xie, L.; Huang, C.; Deng, C.; Liu, H.; Cai, D. Sparse Fuse Dense: Towards High Quality 3D Detection With Depth Completion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp.5418 5427.) or road detection (Chen, Z.; Zhang, J.; Tao, D. Progressive LiDAR adaptation for road detection. IEEE/CAA Journal of Automatica Sinica 2019, 6, 693 702.).
  • 3D object detection Wang, X.; Peng, L.; Yang, H.; Xie, L.; Huang, C.; Deng, C.; Liu, H.; Cai, D. Sparse Fuse Dense: Towards High Quality 3D Detect
  • the advantages compared to each other include but are not limited to colour imaging, high resolution (millions of pixels) and high framerate (generally, at least 30 FPS) in case of cameras, and working in the dark or the possibility of direct depth measurement in case of lidars.
  • Lidars not only have much lower resolution (a few hundreds of thousands of measurement points) and measurement frequency (usually 5-20 Hz) than the passive visual sensors, but in most cases, it suffers the problem of spatial and temporal resolution also competing with each other. Meaning, increasing the spatial (horizontal) frequency decrease the temporal one and vice versa. This fact is unfortunate, as both of them being high can be essential in decision- making systems of intelligent vehicles (e.g., we should recognize an object accurately and as soon as possible).
  • the output of which is utilized in the present embodiment the input spatial resolution is advantageously high, and due to the present embodiment also the temporal resolution can be enhanced. Accordingly, we propose a solution in the present embodiment to increase the measurement frequency of lidar sensors virtually; this enables the maximization of angular resolution. Also, enhancing temporal resolution beyond the limit is already important in itself to avoid hazardous scenarios, as dynamic traffic participants can be present with high acceleration, deceleration, or angular acceleration.
  • the present embodiment can be applied in the presence of a lidar sensor and a calibrated camera (intrinsic and extrinsic) with a higher frame rate.
  • Figs.12A-12B show lidar (Pt ⁇ 1) and camera (It ⁇ 1) measurements at timestamp t - 1. Fig.
  • FIG. 12B shows camera measurements (It) at timestamp t and the generated virtual lidar measurement (Pt,v) to the given time moment.
  • the Point cloud and the optical images show that these data have been recorded in a road-crossing (i.e. where four roads encounter). Because of the views – there are also intermediate views shown by optical images not just those which see along the roads – there are many common details in the optical images. Accordingly, based on the seven optical images we can form a very good vision idea about the road-crossing (see e.g. the left and left-down image, in both of those images two vans can be observed one after the other, see also below). With the comparison of optical images around the point cloud in Fig.12A and Fig.
  • Figs.12A-12B show illustration of the up-sampling problem and our solution on the Argoverse dataset (Chang, M.F.; Lambert, J.; Sangkloy, P.; Singh, J.; Bak, S.; Hartnett, A.; Wang, D.; Carr, P.; Lucey, S.; Ramanan, D.; et al. Argoverse: 3D Tracking and Forecasting With Rich Maps. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.).
  • camera images of Fig. 12B at timestamp t
  • the corresponding lidar measurement does not.
  • We generated this (coloured) lidar point cloud utilizing Pt-1, It-1 and It.
  • deep network-based optical flow estimation is made between the previous and the current camera frames.
  • the estimated displacements are applied to the last measurement points to generate the virtual point cloud to the current timestamp. While it is widespread to apply cameras to spatially up-sample lidar point clouds, and there are various solutions available (Zhao, S.; Gong, M.; Fu, H.; Tao, D. Adaptive Context-Aware Multi-Modal Network for Depth Completion. IEEE Transactions on Image Processing 2021, 30, 5264 5276.
  • the pipeline in the present embodiment adapts to the point cloud characteristics and generates virtual clouds with similar characteristics to the real measurements.
  • the point-level transformation of the system explains this fact. It can be observed by seeing our result on different lidar sensors (e.g., Fig.13 and Figs.21A-23I).
  • lidar sensors e.g., Fig.13 and Figs.21A-23I.
  • ego-motion of the vehicle does not need to be known or estimated as in previous works (Rozsa, Z.; Sziranyi, T. Temporal Up-Sampling of lidar Measurements Based on a Mono Camera.
  • Fig.13 shows details of the virtual point cloud generated with inputs visible in Figs.12A-12B (it is noted that Fig.12A shows Pt-1, as well as Fig.12B shows Pt).
  • Fig.13 shows point clouds for three different time instants (the last for t+1 only for illustration purposes), with the help of slightly different colouring. Colormap in Fig. 13: Green (lightest) – Pt-1 (last available measurement), Blue (next in the details, where displacement is shown) – Pt,v (generated point cloud), and Red (last in the details, where displacement is shown) – Pt+1 (future point cloud, only serves illustration purposes).
  • Section D. shows performance measures from our tests
  • Section E. provides an ablation study and further discussion.
  • Section F. B. Related Works This section is divided into five subsections. First, spatial up-sampling of point clouds is investigated, as it has mature methods and is similar to our approach in terms of sensor fusion (which is also done according to the invention). Second, future frame prediction literature is introduced, which is a recent research interest and similar to our approach as it aims for the virtual lidar frame generation but ignores actual information.
  • lidar frame interpolation methods are examined, which have the same purpose as our approach, but they can be applied offline, as they require a 'future' frame for the generation of in-between frames.
  • some earlier approaches for temporal up-sampling are discussed.
  • some prior art patent documents are also described.
  • B.1. Spatial Up-sampling The spatial up-sampling of lidar point clouds is usually driven by camera images. These methods aim to estimate depth for every pixel of a depth image having the same resolution as a corresponding RGB image with an input very sparse depth image initialized with projected lidar data points.
  • Point Cloud Prediction Point cloud prediction or sequential point cloud forecasting is a hot research topic.
  • the methods of solving this problem are important to our research. They also generate virtual point clouds, as our proposal uses previous frame information. In this way, they can be alternatives to the pipeline we propose. However, they aim differently; they try to predict the future. That is why they have a different approach with relevant differences, too. In their alternative approach, they do not use current (camera) information, which even theoretically limits their accuracy.
  • Point Cloud Interpolation Point Cloud Interpolation methods like Liu, H.; Liao, K.; Zhao, Y.; Liu, M. PLIN: A Network for Pseudo-LiDAR Point Cloud Interpolation. Sensors 2020, 20, 1573.; Liu, H.; Liao, K.; Lin, C.; Zhao, Y.; Guo, Y.
  • Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. refer to the temporal up-sampling problem as predicting future Pseudo-lidar frames, which could be directly used as an alternative to our proposal. However, they require three camera frames as input and two previous lidar frames. The only method which only needs two camera frames and one previous lidar frame (as the one proposed here) is published in Rozsa, Z.; Sziranyi, T. Temporal Up- Sampling of lidar Measurements Based on a Mono Camera.
  • CN 111612728 A a point cloud densification method based on binocular RGB images in which ground segmentation is applied is disclosed. Depth image completion is applied in US 2023/0245282 A1. In US 2023/0136235 A1 the sparse depth data along with image input are processed together in connection with point cloud densification. A vehicle driving safety early warning method using FastFlowNet is disclosed in CN 114983328 A.
  • Fig. 14 shows the proposed pipeline of generating virtual point clouds to the intermediate time stamp t (or 'future pseudo- lidar' frame prediction) for up-sampling.
  • some of the illustrative images of Fig.14 are shown in larger size, i.e. in the further figures the exemplary scene illustrated in Fig.14 is further illustrated.
  • Figs.15A-15C show example inputs, such as in Fig.15A lidar (Pt ⁇ 1) measurement (as disparity image), in Fig.15B camera measurement (It ⁇ 1), as well as in Fig.15C camera measurement (It).
  • Figs.15B and 15C show the same images as the first optical image 102 and second optical image 104 of Fig.14 (these are optical images on which optical flow and optical expansion/motion in depth can be calculated, i.e. images recorded by a camera), but in larger size, the content of them can be better observed.
  • Fig.15A corresponds to the initial depth data 100, but shows a different representation, see below for details.
  • the camera intrinsic Zhang, Z. A flexible new technique for camera calibration.
  • the depth image containing Zt-1 values are determined using the lidar- camera TL,C transformation (e.g. transformation matrix from the lidar coordinate system to the camera coordinate system) and K intrinsic matrix (3x3 matrix of camera parameters which projects a 3D point from 3D camera coordinate system onto the image plane).
  • the intermediate results of the pipeline of the present embodiment with exemplary parameters and inputs will be shown enlarged in the detailed explanation part.
  • the results of certain steps performed (optical flow estimation, ground segmentation, etc.) have a physical meaning, which has been visualized in each case. It is noted that on the contrary to some offline prior art approaches, the present embodiment is performed online, i.e.
  • optical flow estimation In the framework of the present embodiment, the optical flow is obtained by means of a neural network (see Fig. 14 and also below). However, also theoretical background is given (the expression can also be utilized in another equation).
  • the flow field describes the velocity of image pixels as: where u and v are the flow components, between pt and pt-1 image points of different time stamps with x and y are the pixel coordinates.
  • the two images (pt-1 and pt) are utilized, which are determined for every x, y coordinates, having a value at every x, y pixel coordinates, just like ⁇ (in Eq. (11) the respective quantities are given as column vectors according to the transposed representation).
  • the optical flow can be calculated for every x, y pixel coordinates.
  • FastFlowNet Kong, L.; Shen, C.; Yang, J. FastFlowNet: A Lightweight Network for Fast Optical Flow Estimation. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021.
  • FastFlowNet is a lightweight model for fast and accurate prediction with only 1.37M parameters enabling us real-time run. It uses a coarse-to-fine paradigm with a head- enhanced pooling pyramid (HEPP) feature extractor to intensify high-resolution pyramid features while reducing parameters. Besides, the center dense dilated correlation (CDDC) layer is applied to compact cost volume that can keep a large search radius with reduced computation time, and a shuffle block decoder (SBD) is implemented into pyramid levels to accelerate the estimation. Further details can be found in the original paper cited at the beginning of this paragraph.
  • HEPP head- enhanced pooling pyramid
  • CDDC center dense dilated correlation
  • SBD shuffle block decoder
  • Fig.16 shows that FastFlowNet resulted optical flow field from inputs of Figs. 15A-15C.
  • Fig.16 shows the same image as optical flow data 108 in Fig.14 (about Fig.16, see at Fig.17) with grayscaled Middlebury colour code (see the previously referred article).
  • Motion-in-depth estimation Similarly to the optical flow, in the framework of the present embodiment, the motion- in-depth estimation is obtained by means of a further neural network (see Fig.14 and also below). However, also theoretical background is given herebelow (the expression can also be utilized in another equation).
  • Motion-in-depth by definition, is: where Zt and Zt-1 are the depth values at time moments t and t-1 respectively.
  • Z(x,y) is a depth map which stores a depth value at each x,y pixel coordinate. Accordingly, ⁇ is the ratio of depths of the current and previous time instants. In this step, we estimate an 'image' of depth ratios which will relate our 2D flow estimations to a 3D motion estimation (as we see later).
  • is the ratio of depth values, where it is given in the arguments of Z where it is evaluated: Zt-1 is evaluated for pt-1 (pt-1 points to a point with x, y coordinates and Z is the depth there) and Zt is evaluated for pt-1 + ⁇ (pt-1) (the latter is ⁇ at t-1, i.e. ⁇ starting from t-1 same as in Eq. (11), which gives this ⁇ ), i.e. pt based Eq. (11). Zt and Zt-1 are both functions of x, y coordinates. Thus, ⁇ is a function of x, y coordinates.
  • Optical expansion is the reciprocal of the motion-in-depth (see further details in Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.). Thus, if we can estimate the scale change of the scene elements, we will get how much they moved closer or farther away. This principle is utilized in the work Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  • Fig.17 shows motion-in- depth results with the network of (Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion.
  • Figs.15A-C and 16 shows the same image as motion-in-depth related data 112 of Fig.14.
  • the motion-in-depth related data 112 shown in Fig.17 is the motion in depth itself.
  • the ⁇ values which can vary between about 0.5-2 (how the new depth is proportional to the old depth, the approach, the distance determines the ratio of the Z's) are rescaled to a grayscale image.
  • the car is very dark, it has a minimum ratio (if we stick to 0.5 as an example: for a value of 0.5, the car is half as far away at time t as it was at time t-1).
  • Fig.16 where the magnitude of the lateral displacements is represented by the colours according to the optical flow (again, we refer to the colouring of the car).
  • Ground model estimation is applied in the present embodiment for the following reason: Finding the corresponding locations of points of Pt-1 at time t (estimating 3D scene flow) differs from the problem of estimating lidar measurement at time t. Due to the movements and sensor characteristics, the objects will be hit by the sensor in different parts. This will mean a big difference between the scene flow estimated points and real measurements in the case of the ground and far points. The phenomena can be observed, e.g., in the last column of Figs.21A-21I, i.e. in Figs. 21C, 21F and 21I, and see also the importance of our compensation in Table XII. So, in our solution, in the present embodiment, we estimate a ground model, and for the ground points, no displacement is applied.
  • MLESAC Maximum Likelihood Estimation Sample Consensus, see Torr, P.; Zisserman, A. MLESAC: A New Robust Estimator with Application to Estimating Image Geometry.
  • a variant of the RANSAC (Random Sample Consensus) robust model-fitting method is preferably applied to fit a ground plane to the lidar points of Pt-1 with a reference normal vector of [010] T in camera coordinate system (column vector according to the widely applied conventions in the field).
  • Estimated ground points are illustrated in Fig.18 showing segmented ground from input of Figs.15A-C. Ground points are coloured grey, and object points with black.
  • Fig.18 shows the same point set as ground segmented depth data 106 in Fig.14. 3.4.
  • 3D scene flow is defined as the three-dimensional motion of 3D points, just as optical flow is the 2D motion of points in an image.
  • optical flow is the 2D motion of points in an image.
  • the motion-in-depth values, and the depths, Zt-1 projected from the lidar
  • the corresponding depth values, Zt can be determined by rearranging the following equation (Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion.
  • can be calculated (for x, y, z coordinates), and if it has been calculated, also Zt can be expressed from Eq. (15).
  • Eqs. (15) and (16) ⁇ and ⁇ (both depending on x, y coordinates) obtained by respective neural networks according to the approach of the present embodiment are utilized. In this step, calculations are made using the results of the neural networks ( ⁇ and ⁇ ), so here the expressions of Eqs. (15) and (16) are not inserted only for illustration of the theoretical background, but these also show the calculation process.
  • ) estimated from the 3D scene flow is illustrated in Fig.19. In Fig.
  • FIG. 19 estimated displacements (for ground points assumed to be 0) are shown corresponding to lidar data points (Pt-1) from the input data of Figs.15A-15C.
  • Fig. 19 shows the same point set as 3D flow data 114 of Fig.14.
  • the displacements are illustrating with different colours (see the colour bar of Fig.19), with darker points corresponding to smaller displacements (cf. with the ground points having 0 displacement) and with lighter points corresponding to larger displacements.
  • Generating virtual point cloud The virtual point cloud to time t is generated from the points of Pt by the following rule: where the upper index i indicates the i th point of the point cloud and G represents ground points in the point cloud P.
  • Fig. 20 shows estimated virtual measurement (Pt,v) from the inputs of Figs.15A-15C.
  • Fig.20 shows the same point set as subsequent depth data 116 of Fig.14.
  • the appropriate estimation of static and moving objects can be observed in Fig.20, where approaching vehicle points have about 2 m estimated displacement.
  • static environment points including parking cars
  • a temporally up-sampled point cloud sequence can be produced.
  • - initial depth data corresponding to the output depth information image of the scene corresponding to an initial point of time is generally collected from multiple integration windows, but a characteristic time instance (e.g. starting or final time instance) of collecting can be assigned to the data, which will be here the initial point of time, point of time called time stamp elsewhere), and - a first optical image of the scene corresponding to the initial point of time and a second image of the scene corresponding to a subsequent point of time being subsequent to the initial point of time (i.e.
  • the depth data may be e.g. point set data (point cloud data) or depth information image (the latter is preferably range image).
  • the initial depth data may be the output depth information image (e.g. output range image) of the previous part of the depth completion method itself, or, when the depth data is point set data, it can be obtained (generated) by a transformation from that.
  • the above quantities are inputs of the present embodiment, thus the above introduction of these inputs shows that in the present embodiment the depth completion method according to the invention is continued from its output.
  • the present embodiment of the method further – i.e. after obtaining output depth information image – comprises (most of the steps and substeps are illustrated in Fig.14 corresponding to an embodiment for point sets) - an input generating step (this is meant that a step for generating inputs for generating 3D flow data) comprising the substeps of (the substeps below can be performed in any order – cf.
  • optical flow data between the first optical image and the second image is generated by means of an optical flow generation module based on a first auxiliary neural network (see operational step S115 in Fig.14 illustrated by a stylized neural network labelled by ‘FastFlow network’, its output is optical flow data 108), and motion-in-depth related data is generated based on the first optical image, the second image and the optical flow data (see the three inputs in Fig.
  • a motion in depth module by means of a motion in depth module based on a second auxiliary neural network (see operational step S120 in Fig.14 illustrated also by a stylized neural network labelled by ‘Motion-in-depth-network’, its output is motion-in-depth related data 112, which can be either the motion in depth itself, i.e. ⁇ , or the optical expansion, which is chosen for further procession), and o ground depth data and non-ground object depth data are determined by segmenting the initial depth data (see operational step S110 of ground segmentation in Fig.
  • - subsequent depth data is generated based on (see operational step S130 with the output of subsequent depth data 116) o the ground depth data, and o the non-ground object depth data modified based on 3D flow data.
  • 3D flow data correspond to changes compared to the initial depth data, i.e. it contains information how the initial depth data is to be modified. Based on the ground segmentation, the changes for the ground depth data can be considered to be zero by definition.
  • 3D flow data is applied only to non-ground object depth data.
  • 3D flow data 114 in Fig.14, as well as in Fig.19 to the point other than the ground shifting values (for shifting in space, i.e. in x, y, z coordinates) are assigned.
  • the transformed subsequent depth data see subsequent depth data 116 in Fig.14, as well as Fig.20
  • Performing the steps of the present embodiment can also be an option for the depth completion system according to the invention just like other optional features of the depth completion method according to the invention.
  • the steps of the present embodiment are performed by means of a temporal upsampling processing subunit integrated into the – main or depth completion – processing unit.
  • a temporal upsampling processing subunit integrated into the – main or depth completion – processing unit.
  • ground segmentation ground points are segmented, i.e. a respective ground point set and object point set can be determined.
  • the above cited ⁇ is determined.
  • the 3D flow data is determined based on Eq. (16): it is zero for the ground points (by definition) and is utilized for other points.
  • the subsequent point set can be generated by unifying the following two point sets: the ground point set and the object point set modified based on 3D flow data (the latter is an addition in this case, the value of 3D flow data for each coordinate (x, y, z) is added the respective coordinate values of the point set, see Eq. (16)).
  • the above can also be proceeded in the range image representation.
  • the ground segmentation can be performed within the range image itself (there are available approaches) or by transforming the range image into a point set, performing the ground segmentation and transforming the result back to range image representation (then the range image parts of the ground range image and the object range image can be determined).
  • Zt corresponding to the depths has to be calculated (determined) after the determination of ⁇ based on Eq. (15).
  • the ground segmentation is available for the range image for t-1 (i.e. Zt-1).
  • Zt is applied for the object range image part of Zt-1 and the other part (ground range image part thereof) remains unchanged.
  • a range image representation result can also be obtained in such a way that the result of Eq. (16) is transformed from point set representation to range image representation.
  • the result of Eq. (16) is transformed from point set representation to range image representation.
  • Fig.14 gives good illustration for the steps and substeps by the help of a flowchart. It is illustrated in Fig.14 at its left side that the inputs in the illustrated embodiment are It-1 (the first optical image 102) and It (the second image 104), as well as Pt-1 (initial depth data 100).
  • Subsection D.2 describes the evaluation procedure and introduces our results in the case of the Odometry dataset, subsection D.3 in the case of the Depth Completion dataset, while subsection D.4 shows qualitative examples.
  • the most commonly applied error metrics for point cloud generation are involved, namely Chamfer Distance (CD) and Earth Movers Distance (EMD).
  • CD Chamfer Distance
  • EMD Earth Movers Distance
  • Pt,v N_t,v x 3 and Pt N_t x 3 are the data points of the predicted (virtual) Pt,v and ground truth Pt point clouds respectively.
  • the CD metric measures the goodness of the virtually generated point cloud. Namely, by generating point clouds with the method described that have been measured in reality. Our virtual point cloud generated for time instant t will always consist of as many points as we had at time instant t-1, because we generate offsets for these points one by one.
  • SPINet self-supervised point cloud frame interpolation network. Neural Computing and Applications 2022.
  • PointINet Lu, F.; Chen, G.; Qu, S.; Li, Z.; Liu, Y.; Knoll, A.
  • PointINet Point Cloud Frame Interpolation Network. In Proceedings of the Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021.
  • Rigid body based up-sampling Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera.
  • ⁇ Prediction Deng, D.; Zakhor, A. Temporal LiDAR Frame Prediction for Autonomous Driving. In Proceedings of the 2020 International Conference on 3D Vision (3DV); IEEE Computer Society: Los Alamitos, CA, USA, 2020; pp. 829–837.
  • ⁇ PLIN Liu, H.; Liao, K.; Zhao, Y.; Liu, M.
  • PLIN A Network for Pseudo-LiDAR Point Cloud Interpolation. Sensors 2020, 20, 1573.
  • ⁇ PLIN+ Liu, H.; Liao, K.; Lin, C.; Zhao, Y.; Guo, Y.
  • FIGs.21A-I, 22A-I, and 23A-I show for different exemplary scenes illustration of generated clouds with different methods together with the ground truth (black - virtual measurement, grey – ground truth measurement).
  • the first columns of the example figures Figs.21A, 21D, 21G; Figs.22A, 22D, 22G; Figs.
  • FIGS.21B, 21E, 21H and Figs.21C, 21F, 21I; Figs.22B, 22E, 22H and Figs.22C, 22F, 22I; Figs.23B, 23E, 23H and Figs.23C, 23F, 23I) contain zoomed parts, i.e. enlarged details (left-bottom part and right-up part, respectively).
  • FIG.22I Figs.23A-23I illustrate a further example where the limitation of the proposal can be seen (ground model fitting limits the precision of generated ground points, see Fig. 23H). However, dynamic object parts are still the most accurate with our proposal (see the comparison of Figs.23B, 23E and 23H).
  • This Section contains further discussion beyond our test results provided in the previous Section. First, the computation efficiency discussion; next, a separate evaluation for dynamic objects is presented, and finally, an ablation study is presented. E.1. Computation efficiency Real-time running is essential as our goal is to temporally up-sample point clouds.
  • the run time should be less than the time elapsed between measurements with a given frequency we aim to up-sample.
  • the running time of the pipeline components is listed in Table IX with the following configuration: Intel Core i7-4790K @ 4.00GHz processor, 32 GB RAM, Nvidia GTX 1080 GPU, Windows 1064 bit.
  • KITTI-sized images were used (about 1200x350 resolution).
  • the ego-motion can be determined with different localization sensors (like in He, L.; Jin, Z.; Gao, Z. De-Skewing LiDAR Scan for Refinement of Local Mapping. Sensors 2020, 20, 1846.) or by methods (like Campos, C.; Elvira, R.; Gomez, J.J.; Montiel, J.M.M.; Tardos, J.D.
  • ORB-SLAM3 An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM. IEEE Transactions on Robotics 2021, 37, 1874 1890.), and the motion of static scene elements can be calculated based on that.
  • Dynamic objects generally pose a greater threat as they change their position, and they also can change their state variables (angular and linear velocity, acceleration).
  • state variables angular and linear velocity, acceleration.
  • Semantic KITTI A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proceedings of the Proc. of the IEEE/CVF International Conf. on Computer Vision (ICCV), 2019.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Image Processing (AREA)

Abstract

The invention is a method for depth completion, in which a neural network-based processing unit (50) with downscaling and upscaling branches (5, 15), the method comprises - downscaling (S24a, S24b) in downscaling branch levels (20) starting on data blocks (25a-25e) of an initial group corresponding to input depth information images (10a-10e) generated from point sets recorded in consecutive integration windows by a lidar device with non-repetitive scanning pattern, wherein o downscaling convolution is performed in the levels (20), and o at least on highest level (20) convolution is combined with convolution LSTM, - forwarding (S42a) last data block (21e) of the lowest downscaling level as an initial data block (31) to the lowest upscaling level, then - upscaling (S26a, S26b) in upscaling branch levels (30). The invention is further a system for depth completion, and a training method for training the processing unit (50).

Description

METHOD AND SYSTEM FOR DEPTH COMPLETION TECHNICAL FIELD The invention relates to a method and system for depth completion using neural network. The invention also relates to a training method therefor. BACKGROUND ART Accurate and dense depth map prediction is an essential problem in 3D scene understanding. In applications such as dynamic environment analysis, 3D mapping, and virtual city generation, depth information is often obtained from commercially available lidar (light detection and ranging, may also be written as Lidar or LIDAR) sensors as they are able to map their environment in real-time by emitting multiple laser beams and receiving their returns. However, the data captured by lidars is often very sparse, while its characteristics may vary depending on the sensors' scanning technology. In this context, depth completion algorithms need to estimate dense depth images from the lidar-acquired sparse range measurements. For dynamic environment perception and recognition tasks such as advanced scene analysis and understanding, repetitive, typically rotating multi-beam (RMB) lidar sensors (e.g., Ouster OS1 or Velodyne Puck models; see Ö. Zováthi, B. Nagy, and C. Benedek, “Point cloud registration and change detection in urban environment using an onboard lidar sensor and MLS reference data,” Int. J. Appl. Earth Obs. Geoinf., vol.110, p.102767, 2022.) are commonly utilized devices. RMB lidars can produce real-time point cloud streams (300k-2M points/s), however, their measurements have low spatial density, and their field of view (FOV) coverage is constant through the whole scanning process: Their vertical resolution is fixed by the number of the laser beams (16-128), while their horizontal resolution depends on the sensor's rotation frequency (5-20Hz). A study about the state-of-the-art of depth completion techniques and challenges are hereby presented. In the past few years, research on lidar-based approaches emerged as a hot topic in the literature, due to the availability of popular public datasets like the KITTI Depth Completion Benchmark (J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger, “Sparsity invariant CNNs,” in Int. Conf. 3D Vis., 2017.). The majority of the recent methods focus on completing depth maps obtained from RMB lidars fused with optical images as guidance to recover the pixels with missing depth measurements (S. Zhao, M. Gong, H. Fu, and D. Tao, “Adaptive context- aware multimodal network for depth completion,” IEEE Trans. Image Process., vol. 30, pp.5264 5276, 2021. and D. Nazir, A. Pagani, M. Liwicki, D. Stricker, and M. Z. Afzal, ”SemAttNet: Towards attention-based semantic aware guided depth completion,” IEEE Access, pp.1 1, 2022.). Similar approaches are disclosed in US 11,222,217 B2 and US 2022/0262023 A1, which utilize ConvLSTM operation (i.e. convolution Long-Short Term Memory operation, see also below) in the processing neural network in a specific processing location. ConvLSTM operation is mentioned in a depth completion method in US 2023/0230269 A1 using also depth data and RGB images on its input. However, optical images may not provide eligible information in cases of sudden illumination changes (D. Nazir, A. Pagani, M. Liwicki, D. Stricker, and M. Z. Afzal, “SemAttNet: Towards attention-based semantic aware guided depth completion,” IEEE Access, pp. 1 1, 2022.) or in low-light environments (M. F. F. Khan, N. D. Troncoso Aldas, A. Kumar, S. Advani, and V. Narayanan, ”Sparse to dense depth completion using a generative adversarial network with intelligent sampling strategies” in Proc. ACM Int. Conf. Multimedia, New York, NY, USA, 2021, p.5528 5536.). In these cases, depth completion must be performed solely based on sparse lidar range measurement samples, which includes significantly harder challenges (R. Li, D. Xue, Y. Zhu, H. Wu, J. Sun, and Y. Zhang, ”Selfsupervised monocular depth estimation with frequency-based recurrent refinement,” IEEE Trans. Multimedia, pp.1 12, 2022.). First, without relying on external sources (e.g., high-resolution RGB images), edges and other finely textured structures on the generated depth images are often missing, blurred or distorted (R. Li, D. Xue, Y. Zhu, H. Wu, J. Sun, and Y. Zhang, “Selfsupervised monocular depth estimation with frequency-based recurrent refinement,” IEEE Trans. Multimedia, pp.1 12, 2022. and A. Savkin, Y. Wang, S. Wirkert, N. Navab, and F. Tombari, “lidar upsampling with sliced wasserstein distance,” IEEE Rob. Autom. Lett., pp.1 8, 2022.). In R. Li, D. Xue, Y. Zhu, H. Wu, J. Sun, and Y. Zhang, “Selfsupervised monocular depth estimation with frequency- based recurrent refinement,” IEEE Trans. Multimedia, pp.1 12, 2022., global and local depth variations are separated based on the fact that in the wavelet representation of the images, the fine structures mainly appear in the high-frequency domain while the global regions are defined by the low-frequency coefficients. In order to exploit this phenomenon, they introduce a frequency-based recurrent depth coefficient refinement scheme. Similarly to the previous article, the difficulty of data upsampling near the edges also appears in A. Savkin, Y. Wang, S. Wirkert, N. Navab, and F. Tombari, “lidar upsampling with sliced wasserstein distance,” IEEE Rob. Autom. Lett., pp.1 8, 2022., where feature extraction by an edge convolution layer is used to strengthen the precision at fine 3D structures. Second, a limitation of many existing depth completion networks is that they generate new range values for all image pixels, instead of filling only the missing information (Y. Kwon, M. Sung, and S. Yoon, “Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob. Autom., 2022, pp.8424 8430.). Therefore, the Implicit lidar Network (Y. Kwon, M. Sung, and S. Yoon, “Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob. Autom., 2022, pp. 8424 8430.) learns the weights of an interpolation function for 3D point cloud completion, thus the original measurements are not modified and only the missing points are estimated. In a field other than depth completion, in CN 115100090 A and US 10,929,996 B2 ConvLSTM operation is utilized in connection with images sequences applied as input. A scene depth completion system and method based on deep learning is disclosed in CN 114004754 A. A further depth completion method is disclosed in US 10,929,995 B2. In view of the known approaches, there is a need for overcoming the technical limitations of the prior art. DESCRIPTION OF THE INVENTION The primary object of the invention is to provide a method and system for depth completion, which are free of the disadvantages of prior art approaches to the greatest possible extent. The object of the invention is to produce a high-quality, dense and spatially precise point cloud stream from measurements of a single lidar device with non-repetitive scanning pattern. The objects of the invention can be achieved by providing the method according to claim 1 and the system according to claim 6, as well as the training method according to claim 7. Preferred embodiments of the invention are defined in the dependent claims. Light detection and ranging (lidar) is a powerful technology to quickly obtain a large amount of accurate spatial data from the environment. However, by using currently available lidar sensors, we still face a trade-off between the temporal and the spatial resolution of the measurements. Real time lidar sensors, such as non-repetitive circular scanning (NRCS) lidars and rotating multi-beam (RMB) lidars enable us to observe and analyse dynamic scenes (which is needed e.g. by the perception of self driving cars, or in surveillance applications), however the recorded point cloud measurement sequences have a limited resolution either in space or in time, which facts induce hard challenges for the environment analysis capabilities of the available machine perception systems. To overcome the above technological limitations, the present invention covers two main contributions. First, lidar measurement up-sampling in space (see the details of this from here below) and, secondly, lidar measurement up-sampling in time in an embodiment (see the details of this in Section IV below). Alternatively to rotating multi-beam (RMB) lidars, recent non-repetitive circular scanning (NRCS) lidar sensors – in general, lidar device with non-repetitive scanning pattern (see the next paragraph for the scanning patterns) – are also capable of providing measurements for real-time scene analysis in robotics and autonomous driving, at a significantly lower cost compared to the RMB technology (C. Glennie and P. Hartzell, ”Accuracy assessment and calibration of low-cost autonomous lidar sensors,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol. XLIII-B1-2020, pp.371 376, 082020.), by using single- or multi-line lasers combined with high-speed scanning on a circular path (for the scanning properties of an exemplary NRCS lidar, see also Section I.A. and Table VI below). Unlike repetitive RMB lidars (see Background art above), NRCS lidars (e.g., the Livox AVIA sensor) are able to densely map large areas from a given scanning position due to their special scanning technology which follows non-repetitive e.g., rosetta patterns (Fig. 1, the pattern under the graphs, showing the gradually increasing FOV (or FoV) coverage; non-repetitive sampling strategy of the Livox AVIA NRCS lidar is illustrated, the circular scanning produces typical rosetta patterns which are varying across different time frames). The main challenge is here to efficiently balance between the spatial and the temporal resolution of the recorded range data using a suitable integration window (L. Kovács, M. Kégl, and C. Benedek, ”Real-time foreground segmentation for surveillance applications in NRCS lidar sequences,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol. XLIII-B1-2022, pp.45 51, 052022.). On one hand, as shown in Fig.2A, allowing larger integration time (tΔ > 1 s), the laser beams cover a higher proportion (around 90%) of the FOV yielding high spatial measurement resolution. However, the potential ego-motion of the lidar's platform (e.g., vehicle or robot) and the dynamic objects in the surrounding area induce various artifacts, such as blurred shapes of the observed vehicles, pedestrians or buildings, which phenomena complicate dynamic event analysis. On the other hand, if the measurements are collected within a narrow time window (e.g., in 200 ms) they are spatially more precise, however, the resulting point clouds are notably sparse (around 48k points, up to 40% FOV coverage), which fact yields a significant loss of details across the spatial dimension of the FOV (see Fig.2B). By the help of the invention, we aim to overcome the above-mentioned challenges caused by the spatio-temporal trade-off of the NRCS lidar based perception, and propose a novel deep learning based approach for densifying sparse NRCS lidar data while keeping its spatial accuracy high. An embodiment, called a proposed spatio-temporal (ST) deep network (may also be called ST-DepthNet, see Fig. 3 showing its architecture), operates in the range domain and expects as input multiple sparse depth maps captured by a NRCS lidar consecutively in time, with using a narrow (i.e., for example, 200 ms) integration window for each frame. As output, the network provides a dense, high-quality range image of the same FOV, which does not reflect the sensor's original scanning artifacts (i.e., visible trails of the circular scanning pattern). The architecture of the proposed spatio-temporal deep network applied preferably in the invention was directly designed to exploit both spatial and temporal patterns in the input NRCS data for depth completion, by extending a U-Net-like architecture (O. Ronneberger, P. Fischer, and T. Brox, ”U-Net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comp.- Ass. Interv., 2015, pp.234 241. and T. Shan, J. Wang, F. Chen, P. Szenher, and B. Englot, “Simulation-based lidar super-resolution for ground vehicles, ”Rob. Autom. Sys., vol.134, p.103647, 2020.) with Conv2DLSTM layers (mentioned also simply as ConvLSTM throughout the description, see: X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo, ”Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” in Proc. Int. Conf. NIPS, 2015, p. 802 810.). From the aspect that focus on lidar-only depth completion, the most closely related methods to our approach are J. Ku, A. Harakeh, and S. L. Waslander, “In defense of classical image processing: Fast depth completion on the CPU,” Conf. Comput. Rob. Vis., pp. 16 22, 2018. which effectively combines morphological operations and bilateral filtering, and M. F. F. Khan, N. D. Troncoso Aldas, A. Kumar, S. Advani, and V. Narayanan, “Sparse to dense depth completion using a generative adversarial network with intelligent sampling strategies,” in Proc. ACM Int. Conf. Multimedia, New York, NY, USA, 2021, p. 5528 5536. that investigates different sampling strategies for training a generative adversarial network. However, as our experiments show in Section II, both approaches are highly sensitive to the measurement characteristics of the applied lidar sensor and fail to accurately compensate for the irregular, non-repetitive sampling pattern of NRCS lidars. As NRCS lidars are relatively new to the market, to the best of our knowledge, it is connected to this description to provide first a dataset and method utilizing information for depth completion propagated from their measurements. Referring to the respective documents in the introduction, in an embodiment of approach according to the invention, we recover the fine structures by adding an appropriate edge-loss term (C. Godard, O. Aodha, M. Firman, and G. Brostow, “Digging into selfsupervised monocular depth estimation,” in Int. Conf. Comput. Vis., 112019.) to the loss function applied in our training method, instead of performing edge enhancement by a dedicated sub-network. The invention is a spatio-temporal deep network for depth completion using a single non-repetitive circular scanning lidar in an embodiment. In other words, it is a depth image completion method in an embodiment. In a preferred embodiment, our solution connects the last sparse input image to the output by a direct skip connection to force our proposed model to complete the original sparse, but precise range map instead of overwriting it with completely new values. BRIEF DESCRIPTION OF THE DRAWINGS Preferred embodiments of the invention are described below by way of example with reference to the following drawings, where Fig. 1 illustrates non-repetitive sampling strategy together with a graph showing the FOV coverage, Figs. 2A and 2B show dynamic scene captured with different integration windows, Fig.3 shows the flow chart of an embodiment of the method according to the invention, Figs.4A-4D show reconstructions with different prior art methods, as well as an embodiment of the invention, and the ground truth, Figs.5A-5F show results on real measurements from the LivoxBudapest test set, Figs.6A and 6B show a sparse measurement and a depth image completed by an embodiment of the invention, Figs.7A-7C show a depth image exported from the simulator, a pattern of the applied sensor, and a created realistic depth image, Figs. 8A-8C show a depth image exported from the simulator, a mask corresponding to the sensor and a ground truth data generated with the mask, Figs.9A-9C show an actual input, a ground truth and a prediction, Figs.10A-10D show an actual input, a ground truth and two predictions, Figs.11A and 11B show a greyscaled depth image and a corresponding 3D point cloud, Fig.12A shows lidar and camera measurements at timestamp t-1, Fig.12B shows camera measurements at timestamp t and generated virtual lidar measurement to the given time moment, Fig.13 shows details of the virtual point cloud generated based on the earlier illustrated inputs, Fig.14 shows a flowchart of an embodiment for upsampling in time, Figs. 15A-15C show lidar measurement at timestamp t-1, and camera measurements at timestamps t-1 and t, Fig.16 shows optical flow field from inputs of Figs.15A-15C, Fig.17 shows motion-in-depth results from inputs of Figs.15A-15C and Fig. 16, Fig.18 show a segmented ground from input of Figs.15A-15C, Fig.19 shows estimated displacements based on Figs.15A-15C Fig.20 shows estimated virtual measurement at timestamp t, based on the inputs of Figs.15A-15C, Figs.21A-21I show illustration of generated point clouds for a first exemplary scene with different methods including an embodiment of the invention, together with the ground truth, Figs. 22A-22I show illustration of generated point clouds for a second exemplary scene with different methods including an embodiment of the invention, together with the ground truth, and Figs.23A-23I show illustration of generated point clouds for a third exemplary scene with different methods including an embodiment of the invention, together with the ground truth. MODES FOR CARRYING OUT THE INVENTION I. Depth completion Our approach preferably comprises the followings: First, the consecutive measurements of the NRCS lidar are grouped to form discrete time frames, using a narrow, 200 ms integration window (up to 40% FOV coverage in each frame). Thereafter, within a frame the distances of the measured 3D field points from the sensor are assigned to corresponding pixels in a high-resolution range image. By each actual time frame, the last five collected depth images (covering together around 95% of the FOV) are fed to the proposed spatio-temporal depth completion network, which composes a high-quality range image as output, with almost 100% FOV coverage, also eliminating the motion blurring artifacts. The output high-quality range image can be backprojected to the 3D space as well. A. Range image generation Range images are widely used, compact representations of lidar-based depth measurements (see Ö. Zováthi, B. Nagy, and C. Benedek, “Point cloud registration and change detection in urban environment using an onboard lidar sensor and MLS reference data,” Int. J. Appl. Earth Obs. Geoinf., vol. 110, p. 102767, 2022.; L. Kovács, M. Kégl, and C. Benedek, ”Real-time foreground segmentation for surveillance applications in NRCS lidar sequences,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol. XLIII-B1-2022, pp.45 51, 052022. and B. Nagy, L. Kovács, and C. Benedek, “ChangeGAN: A deep network for change detection in coarsely registered point clouds,” IEEE Rob. Autom. Lett., vol.6, no.4, pp.8277 8284, 2021.), which enable to adopt 2D convolution operations and effective image- based neural network architectures (O. Ronneberger, P. Fischer, and T. Brox, “U- Net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comp.-Ass. Interv., 2015, pp.234 241. and X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo, “Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” in Proc. Int. Conf. NIPS, 2015, p.802 810.) during processing. In our approach, preferably, the captured sparse point clouds (collected within tΔ = 200 ms) are converted from the Cartesian (x,y,z) to the spherical (distance, azimuth, elevation) polar coordinate system. Then, a 2D pixel lattice is generated by quantizing the horizontal (azimuth) and vertical (elevation) FOVs. In the resulting range images, the horizontal and vertical pixel coordinates represent the polar azimuth and elevation angles, while the pixel’s depth value encodes the distance of the corresponding point. In our experiments, we have exploited the parameters of the exemplary Livox AVIA state-of-the-art NRCS lidar sensor (see also for the sensor, L. Kovács, M. Kégl, and C. Benedek, “Real-time foreground segmentation for surveillance applications in NRCS lidar sequences,” ISPRS Arch. Photogramm. Remote Sens. Spatial Inf. Sci., vol. XLIII-B1-2022, pp.45 51, 052022.). The sensor's FOV has been exemplary mapped onto a 400x400 pixel lattice, which resolution (5.6 px/°) yields both high spatial accuracy and reasonable computational requirements. As experienced, the density of the recorded valid range values is decreasing towards the peripheral regions of the range image due to the nature of the circular scanning technique: the scanning pattern crosses the optical center of the sensor significantly more frequently, than the FOV's perimeter, making the central regions of the range images densely filled, and leaving peripheral areas notably sparse (see Fig.1 and Figs.2A- 2B, and, furthermore, Table VI below for the datasheet of Livox AVIA sensor). In Figs. 2A and 2B a dynamic scene captured by a NRCS lidar with different tΔ integration windows is shown. A large integration time (tΔ=1 s) induce several blurring artifacts in Fig. 2A, while a narrow integration window (tΔ=200 ms) yields the loss of details. Blurred pedestrians are marked by red ellipses in Fig.2A. As a result of using an exemplary integration time window of 200 ms for collecting the consecutive time frames, around 60% of the range image pixels receive undefined range values. Such a level of sparseness of the range image makes it difficult to efficiently visualize the data or to perform scene analysis, emerging the need for the proposed depth estimation approach. B. Architecture of the proposed spatio-temporal deep network Next, we have used a range image sequence acquired by the NRCS lidar as input to the proposed spatio-temporal deep network (see Fig. 3). As discussed earlier, sparse measurement frames collected in 200 ms time windows cover only a low proportion of the defined 400x400 range image lattice. On the other hand, using a 1 s time frame, the collected point set covers almost fully the sensor's FOV (cf. Fig. 1), but it is affected by motion blur as discussed above in connection with Figs.2A and 2B. In the graph of Fig.1, the FOV coverage (in %) as a function of integration window for a repetitive lidar (Velodyne HDL 64E) as well as for a non-repetitive lidar sensor (Livox AVIA) preferably utilized in the invention. It is observable in the graph of Fig. 1 that the FOV coverage saturates early for Velodyne HDL 64E and the saturation value is not too high. However, the FOV coverage for Livox AVIA increases throughout the investigated integration window high FOV coverage is reached at the end of the window (1 s). Nevertheless, based on Fig.1, we can expect that the measurements from the last 1 s time interval always contain dense range information from the scene. Thus to also prevent blurring, we take five consecutive ”sparse” range images (each one recorded in 200 ms) as our network's input in the embodiment illustrated in Fig.3. Since the main goal is to generate a high-quality output image from the sparse range image inputs, we have adopted an image-to-image U-Net like architecture (for U- net see O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comp.- Ass. Interv., 2015, pp.234 241.): More specifically, we extended the downscaling part of a U-Net network enabling to exploit temporal connections where the input is an image sequence, by utilizing Conv2DLSTM layers presented first in X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo, “Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” in Proc. Int. Conf. NIPS, 2015, p.802 810. Based on this article, let us introduce a regular Long-Short Term Memory (LSTM) cell, which has a memory state Ct and a final state Ht. At each timestep t, the memory is updated as a function of its current input Xt, previous final state Ht-1 based on an input gate it, while the propagation of its previous value Ct-1 depends on a forget gate ft. The propagation of the memory state Ct to the final state Ht depends on the output gate ot. In each dependency, from state α to β, there is a weight term Wαβ and a bias term bα. A Conv2DLSTM cell operates similarly to a regular LSTM cell, with an extension that the input Xt, memory state Ct and final state Ht with their respective gates (it, ft, ot) are 3D tensors – with one temporal and two spatial dimensions – and both the spatial and recurrent transformations are convolutional (marked by * in Equation system (1)) and not element-wise (marked by ), making it able to propagate spatio-temporal features: (1) Accordingly, these are available conv2DLSTM equations performed on the data blocks of a level in the U-net hierarchy from the first time instant until the last time instant. For the it, ft, ot gates, the value is calculated by the help of a σ activation function (ft forget gate is applied on Ct-1 memory state). Wαβ and bα are parameters with α and β indices for input and output. These are calculated based on subequations 2-4 of equation (1) for subequation 1, as well as subequation 5. For subequation 5, the result of subequation 1 is needed. These equations go through the time instants of e.g. a group of five time instants, from the first time instant until the last one. For the first time instant (t=0), Ct-1 and Ht-1 are set to zero (another initial value is also possible), since there is no earlier time instance which is taken into account. This simplifies the equations, but all of the required contributions can be calculated (at the end Ct and Ht for that time instance). Thanks to this, for the second time instant (and for further time instants), all parameters will be available (beside Xt input – the only one which was available at the first time instant, former Ct and Ht as Ct-1 and Ht-1). The equations always point back only to the just previous time instance, earlier time instances are taken into account through the processed values having an effect on the previous time instance. Hence, in our proposed approach, we keep a spatio-temporal three-dimensional (two spatial and one temporal) downscaling branch at the whole left side of the U- Net structure. On the other hand, the upscaling branch of our proposed network is purely two-dimensional, in order to accurately restore the single output image of our interest. Preferably applied skip connections at each level are performed by recurrent pooling utilizing the last output of a Conv2DLSTM layer which represents features of the last 200 ms measurement (but through this, thanks to the application of Conv2DLSTM, the whole sequence, see also the code part in Section II.G. below). Last but not least, we preferably directly connect the last input image to our output (see operational step S32 in Fig.3). With this modification, which is also supported by our ablation experiments (see Section II.), we can exploit that the last and most up-to-date input image contains spatially precise points, and therefore, our network only has to learn the missing regions of the range image (Y. Kwon, M. Sung, and S. Yoon, “Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob. Autom., 2022, pp.8424 8430.). According to the above short references to it, the description of Fig. 3 is hereby given. In Fig. 3 an embodiment is illustrated which is elaborated in high details. Accordingly, for Fig.3 e.g. even resolutions for images (within the U-net these blocks have also channels) are given, see these. These exemplary values expedite the illustration, but it is to be understood that naturally, other values can be selected for the various parameters (resolutions, number of input images, etc.). In Fig.3 a group of sparse input range images 10a-10e are observable on the left side (these are consecutive images as illustrated by the below arrow showing the time), as well as dense (i.e. depth completed) output range image 40 is observable on the right side of the figure. Between the input and output, a processing unit 50 is shown with its two branches, namely a downscaling branch 5 (spatio-temporal branch – 3D) and an upscaling branch 15 (spatial branch – 2D). In the illustrated embodiment, these branches 5, 15 are branches of a modified U-net (U-net like structure) having asymmetric branches. It Is also shown in Fig.3 that skip connections are applied, in respective operation steps S42a, S42b and S42c from the downscaling branch 5 to the upscaling branch 15. These are skip connections performing recurrent poolings by outputting the last element of an applied CONVLSTM operation (illustrated by two-ended arrows under the row of data blocks in the downscaling branch 5; these are recurrent pooling with convLSTM). It is naturally noted that the skip connection corresponding to operational step S42a is a compulsory connection between the branches 5, 15 since the branches are needed to be connected (for the skip connections, see also below). In Fig.3, the downscaling processing levels (see a downscaling processing level 20) and the upscaling processing levels (see an upscaling processing level 30) are illustrated. From the levels the ones provided with reference numbers are illustrated in details. Thus, both input data blocks 25a-25e and output data blocks 23a-23e of the downscaling processing level 20 are shown in the figure, between them performing downscaling in operational step S24b is illustrated. Data blocks 25a-25e are generated from the sparse input range images 10a-10e by a (spatial and/or recurrent) convolution in operational step S11. In the illustrative example of Fig. 3, in operational step S11 the dimensions are changed as the following: from 5x1x400x400 (where 5 is the number of range images 10a-10e, 1 is the channel number and 400x400 is the spatial resolution of the range images 10a- 10e) 5x8x400x400 is obtained, i.e. each of the five data blocks obtained 8 channels in operational step S11. Moreover, in operational step S24b the dimension changes to 5x32x200x200 since operational step S24b is a strided convolution (downscaling convolution; strided ConvLSTM, spatial and recurrent, as also operational step S24a, see below). Thus, according to the U-net like structure, the channel number increases by the help of strided convolutions (here a stride of 2 is applied) and the dimensions of the data blocks 23a-23e is decreased compared to data blocks 25a-25e. Between data blocks 21a-21e and data blocks 23a-23e three dots are taken for illustrative purposes, showing that a large number of downscaling processing levels can be applied (see also the upscaling branch 15 where the three dots have similar illustrative purposes; the numbers of downscaling and upscaling processing levels are always the same). However, in the exemplary illustration, data blocks 21a-21e have the dimensions of 5x64x100x100 (as a result of a further strided convolution having stride of 2), so the data blocks 21a-21e and the data blocks 23a-23e could have been illustrated also as a single level, but in this case too few levels would have been illustrated for the U-net like structure, a higher number of levels is typical. Beside this, the connection between the dimensions of data blocks 21a-21e and data blocks 23a-23e illustrates that an output of a level (data blocks 23a-23e) become the input of the next level for generating new output (data blocks 21a-21e). Conventional convolution uses a step size (or stride) of 1 meaning that the sliding filter moves 1 sample (e.g. pixel in the case of images) at a time. On the contrary, strided convolution introduces a stride variable that controls the step of the filter as it moves over the input. Similar considerations can be applied for the upscaling branch 15, where – according to the only spatial characteristics – only single data blocks are upscaled. In the illustrative example – according to the downscaling branch 5 – exemplary dimensions are the followings: 64x100x100 for data block 31, 32x200x200 for data block 33 and 8x400x400 for data block 35. Similarly to the downscaling branch, upscaling is performed in operational steps S26a and S26b (spatial, Conv2dUpsampling, i.e. two-dimensional upsampling (upscaling) convolution, comprises in this embodiment an upscaling operation and a channel decreasing convolution operation, see also the code part in Section II.G. below, that is why the term upscaling convolution can used in the upscaling branch). Also, similarly to the downscaling branch 5, an upscaling processing level 30 is illustrated in upscaling branch 15. From the data block 35 in the upmost level, the output range image 40 is generated by a convolution in operational step S13. Additionally, in the illustrated embodiment in operational step S32, the last range image 10e is added (this is also a skip connection applied in an addition layer) to the output obtained from the convolution applied in operational step S13 in order to obtain the dense output range image 40 (in this respect, see also below). According to the above, the invention is described herebelow with reference to Fig. 3 illustrating an embodiment. Thus, the invention is a method for depth completion, comprising the steps of - downscaling is performed (see operational steps S24a, S24b in Fig. 3) on a plurality of downscaling processing levels (or: in a plurality of downscaling processing levels; a downscaling processing level 20 is illustrated in Fig.3) of a downscaling branch of a processing unit (see the above description of Fig.3 for the branches 5, 15 of the processing unit 50) based on neural network with starting, at the highest downscaling processing level of the downscaling branch, on a plurality of initial first data blocks (the downscaling is started on data blocks 25a-25e) corresponding to respective input depth information images (preferably, input range images, see Fig. 3 showing the correspondence between input range images 10a-10e and data blocks 25a-25e, see also below) each generated from a respective point set, wherein the respective point sets (or simply, the point sets, i.e. those which correspond to the input range images) have been recorded in consecutive integration windows (i.e. these data are collected in an integration window) of a scene by means of a lidar device with non-repetitive scanning pattern, wherein o on the highest downscaling processing level (see downscaling processing level 20 in Fig.3) or on more consecutive downscaling processing levels from the highest downscaling processing level a convolution operation applied on the respective downscaling processing level is combined with performing convolution Long-Short Term Memory (convLSTM) operation (accordingly, performing convolution LSTM is required on the highest downscaling processing level, see also below in connection with Early fusion: convolution LSTM is applied only on the highest level, or on more consecutive downscaling processing levels from the highest downscaling processing level, see also below in connection with Late Fusion: convolution LSTM is applied on all levels; sometimes convLSTM is called conv2DLSTM showing that it is applied on 2D images), o in a respective downscaling processing level, a downscaling convolution operation (a convolution operation for downscaling wherein a stride is applied) is performed by a downscaling factor on all of the one or more initial first data block of the respective downscaling processing level (i.e. the downscaling is continued on all of the data blocks kept for a level) to obtain respective one or more downscaled first data block of the respective downscaling processing level (this is illustrated in Fig. 3 for downscaling in operational step S24b between data blocks 25a-25e and data blocks 23a-23e), wherein, in case of having one or more downscaling processing level below the lowest downscaling processing level on which a convolution operation is combined with performing convolution Long- Short Term Memory operation (i.e. if there is one or more level under the levels where the convolution is combined with convLSTM), for the downscaling processing levels below the lowest of the one or more downscaling processing level on which a convolution operation is combined with performing convolution Long-Short Term Memory operation (i.e. for the levels below the one or more levels on which convLSTM is applied), downscaling is performed further on at least a part of the plurality of downscaled first data blocks including the last of the plurality of downscaled first data blocks (according to the above definition, in these levels – under the levels in which convLSTM is applied – the last of the data blocks is always kept – it is included in the ‘at least a part of’, see above – and, on the top of it others can be kept; preferably, the last or all of the data blocks are kept for further downscaling), and - after performing downscaling in the plurality of downscaling processing levels (i.e. in all of the downscaling processing levels), - in case of having a single downscaled first data block on the lowest downscaling processing level, the single downscaled first data block, or - in case of having a plurality of downscaled first data blocks on the lowest downscaling processing level, ^ the last of the plurality of downscaled first data blocks (i.e. simply the last is chosen), or ^ a processed last data block outputted by an auxiliary convolution Long-Short Term Memory operation performed on the plurality of downscaled data blocks of the lowest downscaling processing level (the processed last data block – or by other name the output first data block, since it is a ‘first’ data block of the downscaling branch – is the last data block processed by the auxiliary convLSTM operation, preferably the only output of it; accordingly, on the contrary to the above-mentioned (base) convLSTM operations, which can output one (the last) or more data blocks, the auxiliary convLSTM operation outputs only a single data block) is forwarded (one of the above three possibilities: when there is only one downscaled first data block on the lowest downscaling processing level, then it is forwarded, and when the plurality of downscaled first data blocks is kept, there is two possibility) on the lowest downscaling processing level (i.e. the forwarding is performed, see Fig. 3 for an example, where forwarding is performed in operational step S42a) as an initial second data block (in the example of Fig.3 this is data block 31) to the lowest upscaling processing level of a plurality of upscaling processing levels of an upscaling branch of the processing unit, then - upscaling is performed (see operational steps S26a, S26b) in the plurality of upscaling processing levels (an upscaling processing level 30 is illustrated in Fig. 3) of the upscaling branch of the processing unit with starting on the initial second data block of the lowest upscaling processing level of the upscaling branch, wherein in a respective upscaling processing level an upscaling operation (generally, this is an upscaling operation, in which an upscaling with an appropriate factor is done) is performed by an upscaling factor (preferably corresponding to – i.e. having the same value as – the downscaling factor of a corresponding downscaling processing level, i.e. preferably a downscaling level with the same stride can be found) on the initial second data block of the respective upscaling processing level to obtain upscaled second data block of the respective upscaling processing level (this is illustrated in Fig.3 for upscaling in operational step S26b between data block 33 and data block 35), and an output depth information image is obtained based on the upscaled second data block of the highest upscaling processing level (see the output range image 40 – as a preferred output depth information image – obtained based on data block 35). It is noted that, of course, a trained neural network is utilized in the above method according to the invention, in other words, the processing unit is based on a trained neural network. This also holds for the system according to the invention (see below), as well as for the auxiliary neural networks of an embodiment (see also below in Section IV.). According to the above, the method is started on inputs at the highest level of the downscaling branch with inputting the data blocks corresponding to input depth information images (may be called simply depth images) generated from respective point sets recorded by means of a lidar device with non-repetitive scanning pattern (it is noted that according to the above there are multiple data blocks on the input, i.e. the group of them comprises at least two data blocks). As it will be seen below, also channel increasing will be optionally done, this is why the correspondence is defined this way between the data blocks and depth information images, which are thus generated from respective point sets. According to this definition, there is only depth information on the input, i.e. no optical image information is inputted. This is clear from the definition, since the input is exactly defined. In other words, the above mentioned data blocks constitute the input of the downscaling branch (and nothing else), i.e. these are the inputs exclusively (i.e. being exclusive inputs). This ‘only depth input’ characteristics of the invention leads to the advantage that the depth completion method according to the invention – since depending only on depth information – is not influenced or misled by images coming from optical cameras. In the above definition, it has been also given that the respective point sets have been recorded by means of a lidar device with non-repetitive scanning pattern. Accordingly, these have been already recorded (previously recorded, recorded in the past, cf. “last collected” above) when taken to the input. It is emphasized by this wording that the method according to the invention is not such a method for which future time frames are to be collected to have the input (see also below the analysation of certain prior art documents). Thus, the depth completion according to the invention can be performed real-time, i.e. online very advantageously. Furthermore, a lidar device with (or: having) non-repetitive scanning pattern scans the FOV continuously which leads to larger and larger coverage (cf. with the incremental patterns in the bottom row of Fig.1 below the graph). However, in the consecutive integration windows point clouds recorded by this device are typically sparse (the sparseness depends of the length of an integration window). In Fig.1 an exemplary total 1 s is illustrated, during which time a comparatively high FOV coverage can be reached. Fig.1 also shows that for the exemplary selected first integration window, an approx.40% FOV coverage corresponds. The 1 s mentioned also elsewhere is a typical length for which a good FOV coverage can be reached by this device independently from the number of integration windows included in this period. Thus, advantageously, a time period with a good FOV coverage is divided between a group of consecutive integration windows (the number of which has lower significance), i.e. a single depth completed output can be obtained from a plurality of inputs from this time period with good FOV coverage (in the invention, a single output is generated from more input). It is noted in connection with the processing levels that in the exemplary Fig.3 the data blocks 23a-23e are output of a strided convolution and input of a further level, and data blocks 25a-25e are output of an earlier operation – see operational step S11 – beside that these are input of a strided convolution. Thus, based on the detailed description a level can be handled in other way than it is illustrated by the downscaling processing level 20. It is furthermore noted that in connection with the point sets (or depth data) it is given that these are recorded of a scene. This is meant that the consecutively recorded data pieces are considered to be recorded from the same scene (even if the recording device is moving), since these are recorded close in time to each other. These are also in connection in connection with that the depth completion is to be done in approximately the same place, i.e. of the same scene. Furthermore, accordingly, these can be considered to be implicitly included in the other features, i.e. the term „of a/the scene” can even be omitted. It is noted also that from the view of the value of the stride factors (i.e. downscaling and upscaling factors) the downscaling branch and upscaling branch is preferably symmetrical like in Fig.3 and in the code part of Section II.G. below (it is also noted that from the point of view of the changes of channel numbers, the code part is not symmetrical, see below). However, also different stride factors may be chosen for the different levels. If the branches are symmetrical, levels with same stride factor can easily be identified in this. It is noted that it is technically not excluded level pairs in the branches with the same stride, but the overall upscaling is obtained by different stride values than the overall downscaling. It is furthermore noted that ‘initial’ in the name of some data blocks is called with other words: input (of the level) or level-input. Moreover, ‘first’ and ‘second’ in the name of the respective data blocks shows that these correspond to the downscaling branch and upscaling branch, respectively. Thus, ‘first’ and ‘second’ could be ‘first- branch’ and ‘second-branch’, respectively. As mentioned above, a convolution operation applied on a respective downscaling processing level is preferably combined with performing convolution Long-Short Term Memory (LSTM) operation for all downscaling processing levels (for this embodiment, see Late fusion below). Accordingly, in this embodiment, applying of convolution LSTM operation is required for all downscaling processing levels, not at least on the highest downscaling processing level. As shown in Fig.3, the illustrated embodiment is an embodiment in which (for the illustration of channel operations see the code part in Section II.G. below and also its interpretation herebelow) - on a downscaling processing level, a channel increasing (multiplying) convolution operation is performed (see e.g. operational step S11 in Fig. 3, channel is modified also on other levels) for increasing a channel number of a data block by a channel increasing factor (may be different for the levels) preferably on first data blocks of the initial group of the respective downscaling processing level before performing the downscaling convolution operation on the respective downscaling processing level, and - in the upscaling branch a respective – i.e. the same, but reverse direction change (i.e. demultiplication instead of multiplication with a factor) in channel number as in the above channel increasing convolution operation – channel decreasing (demultiplying) convolution operation is performed (see e.g. operational step S13 in Fig. 3) for decreasing a channel number of a data block by a channel decreasing factor corresponding to the respective channel increasing factor (a pair of channel increasing and decreasing can be found in the downscaling branch and the upscaling branch) preferably on the upscaled data block of a corresponding (based on the channel numbers) upscaling processing level after performing the upscaling operation (see the example detailed in the code part in Section II.G. below, where the pairs of channel increasing and decreasing can be found for the example; it is noted that the channel increasing/decreasing convolution does not alter that a data block is initial or downscaled/upscaled; it is also noted that the order of the channel changing operation and the scaling operation can be freely chosen, not just as illustrated). Preferably, also skip connections are performed in the embodiment illustrated in Fig. 3. As shown in the example of the code part in Section II.G. below, a skip connection is preferably performed after a respective upscaling processing level, adding to the upscaled second data block of the respective upscaling processing level (the possibilities are the same as in the base skip connection, i.e. the forwarding between the downscaling and upscaling branches) o in case of having a single initial first data block on the corresponding downscaling processing level (i.e. a downscaling level with the same size values for the data blocks, and, if any, with the same channel number), the single initial first data block of the corresponding downscaling processing level, or o in case of having a plurality of initial first data blocks on the corresponding downscaling processing level, ^ the last of the plurality of initial first data blocks of the corresponding downscaling processing level, or ^ a processed last data block outputted by an auxiliary convolution Long- Short Term Memory operation performed on the plurality of initial first data blocks of the corresponding downscaling processing level. The above description of the channel increasing/decreasing, as well as the skip connections are illustrated by the code part in Section II.G. below, the interpretation (a kind of summary) of which is given hereby. In the code part In the example of code part, first in the downscaling branch, on the highest downscaling processing level a channel increasing (convolution) 1 to 8 is performed (accordingly, with a first exemplary channel increasing factor of 8), the output of which is x1, after that – on a further layer of the same level – a downscaling convolution of the size (this word can also be used for the data block, e.g.400x400, just as resolution, it can be also said that the block has dimensions) from 400 to 200 is done. On the next downscaling processing level, a channel increasing 8 to 32 is performed (a second exemplary channel increasing factor is 4), the output of which is x2. Furthermore, in the same level a downscaling of the size from 200 to 100 is done. On the lowest downscaling processing level, a channel increasing 32 to 64 is performed (a third exemplary channel increasing factor is 2) with an output of x3 and a downscaling of the size from 100 to 50. In the forwarding between the two branches, the application of an auxiliary convLSTM operation outputting a processed last data block is chosen for the exemplary size of 50x50 and channel number 64. On the lowest upscaling processing level, first an upscaling of the size from 50 to 100 is done, after that channel number of 64 is maintained with a respective convolution (channel decreasing is shifted for facilitating the skip connection, i.e. in the lowest level no channel decreasing convolution, but a channel maintaining convolution is done). A skip connection (by performing the auxiliary convLSTM) adding x3 to the output of the lowest upscaling processing level is done for size of 100x100 and channel number 64 (in the above detailed way, there is a fitting in the parameters for the skip connection). On the next upscaling processing level, first an upscaling of the size from 100 to 200 is done, after that channel number is decreased from 64 to 32. At this point, the next skip connection is performed by adding x2 to the output of the level (parameter fitting: size is 200, channel number is 32). On the further next upscaling processing level, an upscaling of the size from 200 to 400 is done, after that channel number is decreased from 32 to 8. For this output, a skip connection is performed by adding x1. After this skip connection, there is a further channel decreasing convolution is done in the example of the code part to decrease the channel number from 8 to 1. This is the pair of the channel increasing 1 to 8. According to the above, the channel increasing/decreasing convolution operation and the downscaling convolution operation/upscaling operation are performed on different (consecutive) layers of a respective downscaling/upscaling processing level (see the respective lines of the code part in Section II.G. part below). On a certain level, a convolution operation (can be downscaling convolution or channel increasing convolution since this this is relevant only for the downscaling branch, in the invention convolution LSTM operation is applied only in this branch and it is not applied in upscaling branch) can be combined with performing convolution LSTM operation in the following way: ^ with downscaling convolution: see Equation system (1) in which the operations applicable for certain pixels are applied only for every second pixels (in case for a stride of two); ^ with channel increasing convolution: in this case – as a result of applying different filters – the number of the channels are increased, thus, also the number of pixels is multiplied to the number of the channels (see Fig.3) on which the operations of Equation system (1) are to be performed. Furthermore, as touched above, in the embodiment illustrated in Fig. 3, after performing upscaling in the plurality of upscaling processing levels, the last of the input depth information images is added (cf. operational step S32) to an upscaled depth information image corresponding to the upscaled data block of the highest upscaling processing level to obtain the output depth information image. As depth information image, preferably range image is applied. Accordingly, an upscaled depth information image is obtained from the highest upscaled data block (e.g. by decreasing the channel number of it, so these correspond to each other), and the output depth information image is obtained from the addition of the two other depth information images. Accordingly, in general, the output depth information image is obtained based on the upscaled second data block of the highest upscaling processing level. Above, last of the plurality of downscaled first data blocks has been written, and here last of the input depth information images is mentioned. According to the fact, that these correspond to point sets have been recorded in consecutive integration windows (i.e. in integration windows being consecutive in time), it is clear that the ‘last’ of them is to be understood ‘last in time’. It is also noted that in the method according to the invention, always a single depth completed output is generated from multiple inputs, so this within this group of multiple inputs the time order can be determined (cf. with Fig.3 where the direction of time is indicated under the input depth information images 10a-10e). C. Training process The proposed spatio-temporal network (having the above detailed architecture) is responsible for learning and predicting a high-density range image using a sparse input range image sequence. To deal with the challenging artifacts presented in the introduction (e.g. typically edges are blurred, or a trade-off is introduced, or detailedness disappears, or noise is observed in larger regions), a proposed loss function L is preferably composed of three main components to address handling of the artifacts in a targeted way. First, we calculate the L1 Loss (LL1) as the mean absolute error between the generated and the GT depth images to force detailed, pixel-level accurate predictions. Second, we adopt the Structure Similarity Index Measure LSSIM proposed by Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600 612, 2004., which quantifies the perceived difference in luminance, contrast and structural information between the predicted and GT depth images using a variety of known properties of the human visual system. Third, we also utilize a smoothness or edge loss term LEDGE specifically proposed for depth images by (C. Godard, O. Aodha, M. Firman, and G. Brostow, “Digging into selfsupervised monocular depth estimation,” in Int. Conf. Comput. Vis., 11 2019.), which induces sharp contours on the generated images, thus spatially precise boundaries are enforced between objects in the 3D space. The preferably applied final loss function can be expressed therefore as follows: Following a parameter optimization step (see Section II. about this), α1= 0.7, α2 = 1.4 and α3 = 1.5 were used in the final model. The loss function was minimized by the Adam optimizer. The learning rate was set to 0.0002 and the decay rate of the first moment to 0.5. We have trained our model on 10 epochs which took around 27 hours on a NVIDIA GTX 1080 Ti graphical processing unit (GPU). D. Datasets While the main goal of the invention is to propose an algorithm which can accurately deal with real NRCS lidar measurement sequences (i.e. to achieve depth completion on the data), it is challenging to provide dense, spatially precise GT depth information for real data due to the independent movements of dynamic objects of the scene including the ego robot or vehicle (see Fig. 2A and 2B). Instead, we constructed a synthetic range image dataset called LivoxCARLA from a realistic virtual world using the CARLA simulator (Y. Kwon, M. Sung, and S. Yoon, ”Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob. Autom., 2022, pp.8424 8430. and A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Proc. Ann. Conf. Rob. Learn., 2017, pp.1 16.), where the behaviour of the Livox AVIA NRCS lidar sensor (see Fig.1) was simulated. The virtual world allows us to extract dense, spatially precise depth information, used as GT for the lidar’s sparse, rosetta patterned samples. During data extraction, the synthetic NRCS sensor was placed by default on the front-top of the capturing vehicle and was pointing forwards. The vehicle was dynamically moving during the whole data recording. To augment on the extractable information (e.g., due to varying ground level), the sensor’s position was randomly rotated along the up axis by [-22.5°, 22.5°], and its height was randomly adjusted between [1.5m, 2.5m]. We performed it this way because if the sensor was always in the same position, it would be too attached to the ground level, etc. We also want it to work on any kind of car (e.g. for cars with different heights), as many cars as the sensor is placed in different positions. We have prepared for this by varying the training data within meaningful configurations. Our exemplary LivoxCARLA dataset consists of 11726 randomly sampled input- output range image pairs, from which 10000 were used for training, 500 as validation and 1226 for testing. Validation is the process of monitoring the evolution of a model with unknown data during the learning process. Typically, we look at how the model improves its accuracy as learns on data unknown to the model/validation data. Testing is the process of examining the effectiveness of a ready-made, learnt model. We look at how the model that is judged to perform best during teaching behaves on unknown data not seen during teaching. These 3 sets are a baseline for any teaching process. Each pair consists of 400 x 400 images: the input range images were generated with NRCS-characteristics by a Livox AVIA sensor model, in 200 ms integration windows (with ca.40% FOV coverage), while a high-resolution ground truth range image was sampled by each fifth input frame, which is matched to the recording of the input range images, the last of which corresponds to the ground truth. Accordingly, the data corresponding to the respective integration time windows is captured at the end of each integration window. Thus, the GT data is recorded at the same time as the data of the fifth integration window. Besides the LivoxCARLA dataset, we also collected real measurement sequences from Budapest. In these experiments, we used the Livox AVIA sensor mounted on the front-top of our test vehicle on a driving path of total 5.5 kilometres in both speedways and in the city center. Similarly to synthetic data generation, the real test vehicle was continuously moving during the measurements, while many different traffic participants were captured. Although this real dataset, referred as LivoxBudapest, does not include GT data, it enables us to validate the effectiveness of the proposed algorithm in real environment, despite the fact that the network is purely trained on synthetic data. According to our tests, the trained method for depth completion performs well on the LivoxBudapest dataset. Some embodiments of the invention relate to a system for depth completion. The system according to the invention comprises a processing unit based on neural network, the processing unit having a downscaling branch and an upscaling branch (see as an embodiment processing unit 50 with downscaling branch 5 and upscaling 15 in Fig.3; we refer to Fig.3 also in connection with other details of the system according to the invention). In the system according to the invention - the downscaling branch is adapted for performing downscaling in a plurality of downscaling processing levels with starting, at the highest downscaling processing level of the downscaling branch, on a plurality of initial first data blocks corresponding to respective input depth information images each generated from a respective point set, wherein the respective point sets have been recorded in consecutive integration windows of a scene by means of a lidar device with non-repetitive scanning pattern, wherein o on the highest downscaling processing level or on more consecutive downscaling processing levels from the highest downscaling processing level a convolution operation applied on the respective downscaling processing level is combined with performing convolution Long-Short Term Memory (convLSTM) operation, o in a respective downscaling processing level, a downscaling convolution operation is performed by a downscaling factor (stride) on all of the one or more initial first data block of the respective downscaling processing level to obtain respective one or more downscaled first data block of the respective downscaling processing level, wherein, in case of having one or more downscaling processing level below the lowest downscaling processing level on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, for the downscaling processing levels below the lowest of the one or more downscaling processing level on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, downscaling is performed further on at least a part of the plurality of downscaled first data blocks including the last of the plurality of downscaled first data blocks, and - the processing unit is adapted for forwarding, after performing downscaling in the plurality of downscaling processing levels, o in case of having a single downscaled first data block on the lowest downscaling processing level, the single downscaled first data block, or o in case of having a plurality of downscaled first data blocks on the lowest downscaling processing level, ^ the last of the plurality of downscaled first data blocks, or ^ a processed last data block outputted by an auxiliary convolution Long-Short Term Memory operation performed on the plurality of downscaled data blocks of the lowest downscaling processing level on the lowest downscaling processing level as an initial second data block to the lowest upscaling processing level of a plurality of upscaling processing levels of the upscaling branch, - the upscaling branch is adapted for performing upscaling in the plurality of upscaling processing levels with starting on the initial data block of the lowest upscaling processing level of the upscaling branch, wherein in a respective upscaling processing level an upscaling operation is performed by an upscaling factor on the initial data block of the respective upscaling processing level to obtain upscaled initial data block of the respective upscaling processing level, and the upscaling branch is further adapted for obtaining an output depth information image based on the initial upscaled data block of the highest upscaling processing level (the output depth information image can be simply obtained at the highest upscaling processing level). It is noted that the optional features of the method according to the invention are compatible also with the system according to the invention (until there is some explicit incompatibility). Thus, all the embodiments of the method are considered to be embodiments of the system (see also in this respect the relevant notes given in Section IV. below). It is furthermore noted that the processing unit may be realized by a computer or with a dedicated hardware thereof. The branches of the processing unit may be implemented by software code running on the computer implementing the processing unit or also by the help of dedicated hardware units. The computer implementing the processing unit handles the input depth information data and performing the relevant operations to get the depth completed output depth information data. Furthermore, some other embodiments of the invention relate to a training method (it has been touched upon above). The training method according to the invention comprises training of the processing unit based on neural network utilized in the method according to the invention and comprised in the system according to the invention. As detailed above (see in particular Eq. (2) above), in an embodiment of the training method a loss function is utilized for training which comprises - a pixel-level term, - a Structure Similarity Index Measure term, and - an edge loss term. In a further embodiment of the training method, the training is performed by the help of a training dataset generated by the help of CARLA simulator (see in particular Section II.F. above), wherein the training dataset comprises pairs of - input depth information images generated with characteristics of the lidar device with non-repetitive scanning pattern, and - output depth information image extracted to be dense using a mask corresponding to the field of view of the lidar device with non-repetitive scanning pattern. Accordingly, the ground truth data can be obtained by the help of the CARLA simulator, which gives a kind of ideal data. After training with the help of this training dataset, there is the possibility advantageously for checking with real data by the help of the LivoxBudapest dataset (see also below). II. Experiments We have trained and quantitatively evaluated the proposed method using the LivoxCarla dataset, exploiting its sparse input-dense output range image pairs generated by a simulated car-mounted NRCS lidar sensor during virtual drives in dense city environments with several dynamic traffic participants (humans, vehicles, bikes). Besides quantitative validation, we also evaluated qualitatively the performance of the proposed method on real data using the LivoxBudapest dataset. Both datasets are introduced earlier in Section I.D. A. Evaluation metrics During quantitative analysis, we performed evaluation in both 2D and 3D, analysing the generated range images, and the backprojected 3D point clouds, respectively. 1) 2D errors For measuring the similarity between the generated range images to the GT, we adopted the following metrics from the KITTI Depth Completion Benchmark (J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger, “Sparsity invariant CNNs,” in Int. Conf.3D Vis., 2017.). In the above equations (3)-(6), IPi denotes the ith pixel of the image generated by the actual method, while IGTi is the ith pixel of the corresponding GT image. N denotes the number of pixels, in our exemplary case N = 400 x 400. Using the two direct depth-based errors (RMSE, MAE), we can compare the absolute accuracy of depth estimates in meters. However, in some cases, for example the large errors for distant objects can disproportionately affect the overall error value using these metrics. On the other hand, the inverse depth-based errors (iRMSE, iMAE) focus on relative error improvements and may be more relevant in scenes with objects at varying distances. Motivated by this, we used both depth and inverse depth-based errors. 2) 3D errors Besides range image based evaluation, we also compared the generated point clouds to the reference model in the 3D space. Let us denote the GT and a predicted point cloud by PGT and PP, and the number of points in PGT and PP by #PGT and #PP, respectively. We evaluate the quality of the predicted point cloud with respect to the GT data using the symmetric Normalized Chamfer Distance (NCD) and Normalized Median Distance (NMD) (Ö. Zováthi, B. Nagy, and C. Benedek, “Point cloud registration and change detection in urban environment using an onboard lidar sensor and MLS reference data,” Int. J. Appl. Earth Obs. Geoinf., vol. 110, p. 102767, 2022.), while these evaluation measures are also used to compare the performance of different baseline algorithms in the 3D space: B. Ablation study and hyperparameters For optimizing the network structure, we investigated the effect of how deeply we integrate temporal information in the network architecture. In a first setup (No fusion), we trained the network without utilizing temporal data and considering only the measurements from the last 200 ms (i.e. the last range image; with this single range image no ConvLSTM can be applied). In a second setup (called Early fusion, a first embodiment), we fused the multitemporal information only in the first Conv2DLSTM (this name can be used instead ConvLSTM) layer and the remaining layers remained pure spatial convolutions. Finally, we propagated the temporal information through the whole feature downscaling branch (called Late fusion; a second embodiment, such an embodiment is illustrated in Fig.3). Furthermore, in each setup, we examined the effect of including/excluding U-Net- like skip connections in the network (Inner levels; this corresponds to the case that we apply the skip connections, thus, in the illustrated example, operational steps S42b and S42c, besides S42a which is the basic connection between the two branches) and to directly bind the last input and the output depth image (Output; see the also optional operational step S32 in Fig.3). The comparative results are displayed in Table I. Temporal fusion Skip connections RMSE ↓ MAE ↓ X X 3512.95 1549.36 X Inner levels 3367.75 1353.49 X Inner levels+Output 2869.84 1383.10 Early X 2969.22 883.92 Early Inner levels 2767.87 832.87 Early Inner levels+Output 2129.39 686.34 Late X 2672.20 800.16 Late Inner levels 1897.99 523.57 Late Inner levels+Output 1799.16 440.42 Table I – An ablation study of the proposed spatio-temporal deep network architecture According to the comparative results in Table I, the proposed late fusion approach produced the less RMSE and MAE rates (and also, the Early fusion yields better results than the respective No fusion results), while allowing direct skip connections between the latest sparse input frame and the predicted output significantly improved on the results at each temporal setup. Using these connections, it is noted that the network learns to complete the missing regions of the sensor's sparse range map, while keeping high fidelity to the accurate range measurements from the last 200 ms time frame. Comparing the results of Table I, it is clear that the best results are obtained when there are Late fusion, Skip connections, as well as adding the last range image to the output (shortly, ‘output connection’), but it can be observed in Table I that in also with Late fusion and Skip connections, but without output connection, a comparatively good results can be obtained. Leaving the other optional feature, the Skip connections, the results remain better than most of the other “lower level” (i.e. with less connections) results. It is also observable in Table I that the third best results by Early fusion, Skip connection and output connection (also substantiating the advantages of the embodiment based on Early fusion). Next, we also performed hyperparameter optimization steps in the final late fusion based model where we compared different weight combinations of the L loss function's subterms. The most relevant configurations are summarized in Table II. Output skip layer α1 α2 α3 RMSE↓ MAE↓ X 0.85 1.00 0.90 4267.88 1494.07 X 0.85 1.00 1.00 3871.45 1366.16 X 0.60 2.50 2.50 2938.23 1512.53 X 0.70 1.00 1.20 2000.33 541.06 X 0.70 1.40 1.50 1897.99 523.57 ^ 0.70 1.40 1.50 1799.16 440.42 Table II – Different hyperparameter setups for the final model First, using a relatively higher weight for the LSSIM loss term (α1) results in smoothed edges and blurred fine structures and therefore it produces higher RMSE and MAE errors. On the other hand, if the weight of LSSIM is significantly smaller than the weight of LL1 (α2), image regions with uniform depth remain noisy, resulting again in higher RMSE and MAE rates. As a good balance, we experienced that an optimal ratio between the weights of LSSIM and LL1 is around 1:2. Second, based on experiments, the point level LL1 and edge based LEDGE loss terms (α3) are in the best balance with a weight ratio of around 1:1. The values show on what ‘around’ is meant. It is noted in connection with Table II (see Table I for comparison) that the hyperparameter setups were investigated for a structure with Late fusion and Skip connections (thus, the hyperparameter setup ends to the results of the penultimate row of Table II which is the same as the penultimate row of Table I), and then, output connection (operational step S32) has been “switched on” with the same hyperparameters achieving even better results in the ultimate row of Table II (similarly to Table I). C. Reference methods We have compared the results of the proposed spatio-temporal deep network model to related approaches published in the recent years. Note that the majority of existing methods (J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger, “Sparsity invariant CNNs,” in Int. Conf.3D Vis., 2017.; S. Zhao, M. Gong, H. Fu, and D. Tao, “Adaptive context-aware multimodal network for depth completion,” IEEE Trans. Image Process., vol.30, pp.5264 5276, 2021. and D. Nazir, A. Pagani, M. Liwicki, D. Stricker, and M. Z. Afzal, “SemAttNet: Towards attention-based semantic aware guided depth completion,” IEEE Access, pp.1 1, 2022.) relies on fused lidar based sparse depth maps and dense RGB images, therefore we cannot directly compare the proposed method to them, as we address lidar-only scenarios (M. F. F. Khan, N. D. Troncoso Aldas, A. Kumar, S. Advani, and V. Narayanan, “Sparse to dense depth completion using a generative adversarial network with intelligent sampling strategies,” in Proc. ACM Int. Conf. Multimedia, New York, NY, USA, 2021, p.5528 5536.). As the first baseline for comparison, we investigated how the sensor itself can produce high density images, by allowing a large integration window (tΔ = 1s) to cover a high proportion (>95%) of the FOV. We refer to this method from now on as “Large integration”; accordingly, in the Large integration method the data is not processed but the data is collected for longer times than in case of the spatio- temporal deep network approach. As the second reference, we adopted an improved version of the method presented in J. Ku, A. Harakeh, and S. L. Waslander, “In defense of classical image processing: Fast depth completion on the CPU,” Conf. Comput. Rob. Vis., pp. 16 22, 2018., called hereafter as “IP-Basic++”, by optimizing its morphological operations to our irregular NRCS data and extending it with bilateral blurring. We have chosen as the third reference the approach of M. F. F. Khan, N. D. Troncoso Aldas, A. Kumar, S. Advani, and V. Narayanan, “Sparse to dense depth completion using a generative adversarial network with intelligent sampling strategies,” in Proc. ACM Int. Conf. Multimedia, New York, NY, USA, 2021, p.5528 5536., called henceforward “Sparse-to-Dense”, which is proposed directly for lidar- only perception, and we trained it on our LivoxCarla dataset, with the parameters described in this M. F. F. Khan et al. article. To adopt the latter method to our dataset, we changed the size of the input layer from 480 x 480 to our range image lattice 400 x 400. D. Comparative results Next, we compare the proposed spatio-temporal deep network approach to the above three reference methods on the LivoxCarla test set, in both 2D range image based and 3D point cloud based representations. The overall mean values of the calculated 2D error rates – see the definitions in Eqs. (3)-(6) – are displayed in Table III. Method iRMSE↓ iMAE↓ RMSE↓ MAE↓ Large integration 74.745 21.12 4259.83 1119.41 IP-Basic++ 170.02 24.31 2918.33 574.99 Sparse-to-Dense 493.65 151.55 4583.75 1672.79 Proposed spa^o- 59.46 15.94 1799.16 440.42 temporal deep network approach Table III – Comparative results between 2D range images Regarding all numerical quality measures, the performances of the Large integration and the Sparse-to-Dense approaches are quite similar, the IP-Basic++ works better in average, while the proposed spatio-temporal deep network approach significantly outperforms all of them, reducing their RMSE errors by more than 1 m (cf. with the definitions of errors in Eqs. (3)-(6)). First, we can observe that the main sources for large errors of the Large integration method are the movement of the capturing platform and the presence of dynamic objects of the scene. Second, the IP-Basic++ approach predicts missing depth values more robustly on large, homogeneous surfaces, but fails estimating the fine details. Third, the depth image estimation by the Sparse-to-Dense method keeps the trails of the circular scanning pattern of the NRCS sensor still visible, while scene objects and finely textured regions are often merged with their background. Figs. 4A-4D demonstrate these limitations of the Sparse-to-Dense method and the IP- Basic++ approach on a range image sample, where the proposed method performs significantly better. Fig.4A-4D show that fine structures (marked by green ellipse in Fig.4C) recognized by the proposed spatio-temporal deep network approach (Fig. 4C), but remained partly or fully unrecognized (merged to wall or background) by the reference IP-Basic++ (Fig.4A) and Sparse-to-Dense (Fig.4B) approaches with respect to the Ground Truth data (Fig.4D: in this case the ground truth data is the ideal data which corresponds to the time instant at the end of the time window covered by the proposed spatio-temporal deep network approach, i.e. to the status of the scene being frozen at the end of the time window, cf. with Fig.3). Compared to the prior art approaches of Figs.4A and 4B, Fig.4C – illustrating an embodiment shows much better resemblance to the ground truth shown in Fig.4D. We conducted further analysis in the 3D domain, by comparing the 3D Ground Truth scene models to point clouds backprojected from the range images generated by the proposed and reference methods. Errors obtained by calculating the symmetric Normalized Chamfer and Median Distance metrics – introduced above in Eqs. (7)- (10) – are displayed in Table IV. Method QNMD[mm]↓ QNCD[mm]↓ Large integration 1754.12 3830.62 IP-Basic++ 1072.60 2241.69 Sparse-to-Dense 4466.14 6065.44 Proposed spa^o-temporal 687.36 1718.53 deep network approach Table IV – Comparative results in the 3D space As shown, the error of the proposed method is the smallest, by a margin of around a half meter regarding both metrics. Note that, while Large integration seems to work better than Sparse-to-Dense in 3D, this observation is mainly the consequence of the fact that object regions affected by motion blur can still have points close to GT in the 3D space, and vice versa. E. Analysis on real measurements Beyond a comprehensive numerical evaluation on our synthetic LivoxCarla dataset, we also validated the proposed method on real lidar measurement sequences of the LivoxBudapest set, supporting its future real-life application. The LivoxBudapest test set contains three different scenarios: two pathway recordings from the city center (a boulevard and a narrow street), both around 1 km long, and a speedway section near the city, recorded on a path of around 3.5 km. Fig.5A-5F displays selected relevant sample frames from the three scenarios. Thus, Figs.5A-5F show results on real measurements from the LivoxBudapest test set. Namely, the row of Fig.5A corresponds to sparse measurements (sparse input data captured in a 200 ms time window), and the row of Fig. 5B corresponds to RGB images for visual reference only. Rows of Figs. 5C-5F corresponds to predicted depth maps by different methods (Fig.5C: Large integration with integration time tΔ = 1s; Fig.5D: IP-Basic++, Fig.5E Sparse-to-Dense and Fig.5F: proposed spatio- temporal deep network approach). Accurately predicted fine object structures by the proposed spatio-temporal deep network approach are highlighted with green ellipses in Fig.5F. As for the Large integration method (Fig.5C), it performs significantly worse on real data than on the simulated samples: its generated range images are extremely noisy, and if the platform is moving, structures are barely recognizable. Note that by large integration time the regions of moving street objects become blurred even if the lidar platform is static. We can observe that similarly to the experiments with synthetic data, the IP-Basic++ approach (images of Fig.5D) robustly completes missing values in case of larger surfaces (walls, ground areas and even vehicles), but fails to accurately estimate fine structures. For example, pedestrians in the first column of Fig.5D are blurred into one object, while traffic lights and signs are partly merged to the background in the second and fourth columns of Fig.5D. The Sparse-to-Dense method (images of Fig. 5E) cannot eliminate the rosetta patterns of the input lidar measurements, which are typically visible on ground and wall areas (e.g., second column in Fig.5E). The tendency of merging fine structures into larger surfaces is also notable: In the first and third columns of Fig.5E, vehicles and pedestrians are falsely merged to the wall behind them. Such artefacts can mean critical problems for urban scene understanding tasks, while as shown, they are handled better by the proposed spatio-temporal deep network approach (see regions marked by green ellipses in Fig.5F). The Sparse-to-Dense method heavily blurs other fine structures (columns, traffic signs and lights, etc.) as well, as displayed in Fig.5E. Moreover, while objects close to the sensor are usually well recognizable for the human eye, they are often predicted at inaccurate distances with this method (e.g., cyclist in the second column of Fig.5E). Besides the above qualitative analysis, we conducted a survey for visual verification of the generated depth image streams, by asking 20 computer vision related experts to rate the quality of the input and the output of each method in all three videos with scores between 1 and 9, where 9 is the best possible score. The results provided in Table V confirm that the test experts found the proposed method significantly better than the reference techniques. Method /score↑ Boulevard Narrow st. Speedway Mean Sparse input data 2.95 2.75 2.60 2.76 Large integration 4.30 4.45 3.75 4.17 IP-Basic++ 5.75 5.65 5.60 5.67 Sparse-to-Dense 6.10 6.35 6.40 6.28 Proposed spa^o- 7.15 7.20 6.65 7.00 temporal deep network approach Table V – Comparative survey results on real measurements In summary, the proposed spatio-temporal deep network method can better compensate for both the noisiness and the sampling pattern of the sensor data, while the predicted distance values of dynamic objects or background scene structures are more accurate than in the depth maps of the reference approaches. Regarding the computation time, for the prediction of a single depth frame, the proposed spatio-temporal deep network method needs 100 ms. The datasets and the data generation process is detailed in Section F below. Source code is attached in Section G below. F. Datasets and results The purpose of project connected to the invention is effective deep learning based depth completion for sparse measurements captured by the Livox AVIA sensor (displayed in Figs 6A-6B, showing a sparse measurement by the Livox AVIA sensor (Fig. 6A) and the completed depth image by the proposed spatio-temporal deep network approach (Fig.6B)). The sensor specification sheet can be found in Section II.H. below. We use training, validation and test data with ground truth information from the Carla simulator (see above, Y. Kwon, M. Sung, and S. Yoon, ”Implicit lidar network: lidar superresolution via interpolation weight prediction,” in Proc. Int. Conf. Rob. Autom., 2022, pp. 8424 8430. and A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Proc. Ann. Conf. Rob. Learn., 2017, pp. 1 16.). Our exemplary dataset consists of 11726 randomly sampled sparse input-dense output range image pairs. Each sample consists of a sequence of five consecutive range images with height and width of 400 pixels. The dataset is arranged in three folders named: ^ Train: 10000 sparse samples with ground truth data, ^ Validation: 500 sparse samples with ground truth data, ^ Test: 1226 sparse samples with ground truth data. We provide the data in raw image format. The dataset was generated from the Carla simulator that gives the opportunity to export perfect depth images without any distortion or blurring. To simulate realistic Livox AVIA measurements, the dense depth images were sampled with rosetta scanning pattern of the Livox AVIA sensor. As mentioned above, during the whole data recording, the capturing platform (a simulated) vehicle was dynamically moving, and to augment on the extractable information (e.g., vary the ground level), the capturing sensor’s position was randomly rotated along the up axis by [−22.5°, 22.5°], and its height was randomly adjusted between [1.5m, 2.5m], see above. The final dataset consists of 11726 randomly sampled input-output data pairs. The data generation process is illustrated in Figs 7A-7C and 8A-8C, respectively. In Figs. 7A-7C it is shown that in each sample, the input data is a sparse depth image sequence, that consists of five consecutive sparse depth images, each sampled after 200 ms. In the figure, the patterns of the Livox AVIA sensor (displayed in Fig.7B) are used to filter the depth image exported from the simulator (Fig.7A) resulting in realistic, Livox-like depth images (Fig.7C). In Figs. 8A-8C it is illustrated that the output (ground truth) data in Fig. 8C is generated from the depth image exported from the simulator (Fig.8A) at the end of each input sequence using the mask (obtained from real measurements, see Fig. 8B) with the full field of view of the Livox AVIA sensor. We also provide three real-life measurement sequences in video format (LivoxBudapest Dataset) and code for testing the trained models in real-world scenarios. There are different ways to plot the model's predictions, that are displayed in Figs 9A-9C showing the model input, ground truth and prediction in one row, respectively. For plotting the prediction of two model variants for comparison (i.e. comparison of two different models/learning, cf. with Tables I and II, where different models for different embodiments are parametrized), see Figs 10A-10D showing the model input, the ground truth and two predictions, respectively. It is it is observable in Figs. 10C and 10D that, within certain limits, the result can be good or even better. A grayscale depth images can be converted to corresponding 3D point clouds, an example is shown from these in Figs.11A-11B, respectively. G. Code part Herebelow, a code part is illustrated (with comments written with capitalized letters) for network architecture of the neural network utilized in the processing unit in an embodiment. This is a Python source code part. INPUT SEQUENCE inputs = tf.keras.Input(shape=(5, 400, 400, 1)) FIFTH INPUT RANGE IMAGE FOR DIRECT FEEDBACK (SEE OPERATIONAL STEP S32) input5 = layers.Lambda(lambda x: x[:,4,:,:,:])(inputs) DOWNSCALING BRANCH FIRST CONVLSTM LAYER WITH EXTENDING CHANNEL NUMBER 1->8 x = layers.ConvLSTM2D(8, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=True)(inputs) x = layers.Activation('relu')(x) x1 = layers.BatchNormalization()(x) FIRST CONVLSTM LAYER WITH DOWNSCALING 400->200 x = layers.ConvLSTM2D(8, kernel_size=(3, 3), padding='same', strides = (2, 2), return_sequences=True)(x1) x = layers.Activation('relu')(x) x = layers.BatchNormalization()(x) SECOND CONVLSTM LAYER WITH EXTENDING CHANNEL NUMBER 8->32 x = layers.ConvLSTM2D(32, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=True)(x) x = layers.Activation('relu')(x) x2 = layers.BatchNormalization()(x) SECOND CONVLSTM LAYER WITH DOWNSCALING 200 -> 100 x = layers.ConvLSTM2D(32, kernel_size=(3, 3), padding='same', strides = (2, 2), return_sequences=True)(x2) x = layers.Activation('relu')(x) x = layers.BatchNormalization()(x) THIRD CONVLSTM LAYER WITH EXTENDING CHANNEL NUMBER 32->64 x = layers.ConvLSTM2D(64, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=True)(x) x = layers.Activation('relu')(x) x3 = layers.BatchNormalization()(x) THIRD CONVLSTM LAYER WITH DOWNSCALING 100 -> 50 x = layers.ConvLSTM2D(64, kernel_size=(3, 3), padding='same', strides = (2, 2), return_sequences=True)(x3) x = layers.Activation('relu')(x) x = layers.BatchNormalization()(x) FIRST SKIP CONNECTION AT THE BOTTOM OF THE TWO BRANCHES, WITH 64 CHANNELS AND WITH KEEPING THE LAST ELEMENT OF THE SEQUENCE x = layers.ConvLSTM2D(64, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=False)(x) x = layers.Activation("relu")(x) x = layers.BatchNormalization()(x) UPSCALING BRANCH UPSCALING 50->100 x = layers.UpSampling2D(size=2)(x) CONVOLUTION x = Conv2D(64, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu')(x) xx3 = layers.BatchNormalization()(x) SECOND SKIP CONNECTION inp3 = layers.ConvLSTM2D(64, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=False)(x3) x = layers.Add()([xx3, inp3]) UPSCALING 100->200 x = layers.UpSampling2D(size=2)(x) CONVOLUTION x = Conv2D(32, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu')(x) xx2 = layers.BatchNormalization()(x) THIRD SKIP CONNECTION inp2 = layers.ConvLSTM2D(32, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=False)(x2) x = layers.Add()([xx2, inp2]) UPSCALING 200->400 x = layers.UpSampling2D(size=2)(x) CONVOLUTION x = Conv2D(8, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu')(x) xx1 = layers.BatchNormalization()(x) FOURTH SKIP CONNECTION inp1 = layers.ConvLSTM2D(8, kernel_size=(3, 3), padding='same', strides = (1, 1), return_sequences=False)(x1) x = layers.Add()([xx1, inp1]) CONVOLUTION FOR OUTPUT outputs = layers.Conv2D(1, 3, activation="sigmoid", padding="same")(x) ADDING OPTIONALLY THE FIFTH INPUT RANGE IMAGE (SEE OPERATIONAL STEP S32) x = layers.Add()([outputs, input5]) Hereabove, a code part corresponding to a network being a little bit different than the one introduced above with exemplary parameters, since in this case the downscaling goes down from 400x400 to 50x50 (there are four skip connections instead of the three illustrated in Fig. 3), which is forwarded between the two branches. In the code part it can be followed how the inputs and outputs for the layers and within the layers are handled (sometimes both input and output called ‘x’, but sometimes auxiliary variables are introduced to show how the values are handled: x1, x2, x3 and xx1, xx2, xx3). Further details are given in the above code part, e.g. performing the skip connections is illustrated. Exemplary stride value (two) is given in the above code part, as well as exemplary kernel size for the convolution. It is also specified in the code part that there are five of the 400x400 input images. In the arguments of layers.ConvLSTM2D function, the first is the output channel number of it, and the ‘return_sequences’ argument decides that all of the input data blocks are outputted after performing the convLSTM operation on them (in case of value ‘True’), or only the last data block is outputted (in case of value ‘False’). Accordingly, in case ‘return_sequences=False’ is the argument, then make one out of five data blocks (this is called recurrent pooling). Furthermore, same padding adds additional rows and columns of pixels around the edges of the input data so that the size of the output data is the same as the size of the input data. In the layers of the processing ReLU (rectified linear unit) activation function and batch normalization is utilized. These are functions normally used in neural networks. ReLU is an activation function defined as the positive part of its argument (if input is negative, then the output is zero, and, if the input is positive, then the output is equal to the input). Furthermore, batch normalization is a method used to make training of artificial neural networks faster and more stable through normalization of the layers' inputs by re-centering and re-scaling. ReLU introduces nonlinearity to the processing in a usual way. H. Datasheet of the device applied in the experiments Some data (main parameters) from the datasheet of the Livox AVIA device applied in the experiments are given in Table VI. Livox AVIA datasheet Laser Wavelength 905 nm Detection Range (@ 100 klx) 190 m @ 10% reflectivity 230 m @ 20% reflectivity 320 m @ 80% reflectivity Detection Range (@ 0 klx) 190 m @ 10% reflectivity 260 m @ 20% reflectivity 450 m @ 80% reflectivity FOV Non-repetitive scanning pattern: 70.4° (Horizontal) x77.2° (Vertical) Repetitive line scanning: 70.4° (Horizontal) *4.5° (Vertical) Range Precision (1σ @ 20m) 2 cm Angular Precision (1σ) < 0.05° Beam Divergence 0.28° (Vertical) x 0.03° (Horizontal) Point Rate 240,000 points/s (first or strongest return) 480,000 points/s (dual return) 720,000 points/s (triple return) Data Latency < 2 ms Noise 40cm omnidirectional <45 dBA Table VI – Livox AVIA datasheet In connection with the data in Table VI, the followings are given. In the detection range klx corresponds to the solar irradiance. In strong sunlight, the sensor sensitivity is slightly worse. The principle of operation is not affected by this. The FOV value corresponds to a single sensor which “sees” in front of it within a cone which is a not regular cone according to the above data (see the different horizontal and vertical data). In the description, it is discussed how the coverage of this FOV builds up (cf. Fig.1). The datasheet shows that the Livox AVIA sensor as an exemplary NRCS lidar has an approximately circular FOV. In connection with range and angular precision: This specifies how accurate the detected point is in distance (range) and laterally (angular). So how accurate it is that what we see is at this angle and at this distance. The value of data latency is relevant in that e.g.200 ms is the integration window, compared to this less than 1% is spent for data aggregation, that is the latency. III. Comments on prior art and conclusion A. Comments on prior art Herebelow, some of the prior art documents mentioned in the introduction are commented. It is noted that the approach of the invention shows huge differences from CN 115100090 A and US 10,929,996 B2 – among others – based on the fact that this latter prior art approaches are applied for images which is in connection with the fact that these documents are not related to depth completion, i.e. have different objectives. In the prior art, ConvLSTM is applied for (optical) images since there are a lot of similarity between the images of the consecutive frames based on which the convLSTM module can work. It is also noted in connection with CN 115100090 A that in the central part of the processing, the data is processed in a complicated way using also data from future time frames (t for t-1, t+1 for t, t+2 for t+1, etc.). The utilization of future data blocks the possibility of real-time processing in the case of this prior art approach. On the contrary as touched upon above, in the invention only previously recorded data is utilized and thus, very advantageously, the depth completion according to the invention can be performed real-time. It is furthermore noted that in US 10,929,996 B2, according to Fig. 2 thereof, the depth map is generated for a single (optical) image. Also, in CN 115100090 A depth data is generated for each (optical) image input. Accordingly, differently from the invention which do that, none of these approaches intends to perform depth completion based on a group of inputs, but these approaches generate output on a one-to-one basis for each of the inputs. According to the invention, the neural network based approach preferably utilizes sparse range image inputs. The consecutive sparse range images have ca.40% FOV coverage each and the overlap in their content is relatively low; this is a result of the non-repetitiveness of the pattern in the FOV which is an inherent feature of the NRCS lidar device. Such a good results with convLSTM are thus not expected. According to our experiences, convLSTM works together very well with consecutive frames recorded by the NRCS lidar device giving a good coverage of the selected FOV altogether thanks to the non-repetitiveness. Thus, somehow synergically the non-repetitiveness of the applied device results in highly advantageous performance. It is noted that the application of a rotating multi-beam lidar would be based on hugely different principles than the invention, since it is repetitive in its nature. A point cloud of a rotating multi-beam lidar would not help the depth completion since – because of the repetitiveness – it would not yield information for depth completion: it scans the same field of view, the consecutive rotations bring the same information in this respect. It is also noted that difference of the consecutive scans may be in the dynamic changes, which are typically also the part of the input of the invention, but – because of non-repetitiveness – the invention also brings in new information (in other words, new coverage contributions) with each consecutive frames for helping the depth completion (cf. how the exemplary rosetta pattern builds up from the coverage contributions coming time after time). These are lidar related aspects, do not have analogy with (optical) image inputs applied in the prior art approaches, i.e. the inputs cannot be considered interchangeable. In connection with US 11,222,217 B2 and US 2022/0262023 A1, it is noted that in these documents the ConvLSTM block is utilized in a central portion of the processing. B. Conclusion To conclude, hereabove we proposed a novel depth completion method called the proposed spatio-temporal deep network approach above, which is capable of creating high-density depth images from sparse consecutive depth maps preferably acquired by a NRCS lidar. For training and quantitative evaluation, we constructed a new synthetic benchmark set called LivoxCarla, and we showed that the invention outperforms two state-of-the-art reference methods. The usability of the proposed method on real NRCS measurement data has also been demonstrated using our recorded LivoxBudapest real-life dataset. In the future, we aim to use the proposed method in intelligent robot and vehicle platforms, for improving the limited spatial resolution of NRCS lidars. We hereby summarize the main preferred contributions of invention as follows: ^ We proposed a novel deep learning based solution (proposed spatio- temporal deep network approach), which in an embodiment extends the classical U-Net architecture with a spatio-temporal downscaling branch for utilizing consecutive sparse measurements captured by NRCS lidars. Our model produces spatially precise high-density depth data using a spatial upscaling branch following effective temporal pooling steps. ^ We provided a new synthetic urban dataset called LivoxCarla, which contains simulated NRCS lidar data with corresponding dense depth Ground Truth (GT) information. We demonstrate that with the LivoxCarla dataset, we are able to simulate realistic urban NRCS lidar measurements, which can provide a basis for training and evaluation of methods developed for processing real NRCS sensor data. ^ We provided a real-life dataset called LivoxBudapest, which contains real NRCS lidar measurement data collected in Budapest, Hungary, both in downtown and speedway areas, by a sensor mounted on a moving vehicle. Testing with the real measurements allows us to clearly demonstrate the usability of the method according to the invention trained on synthetic data in real-life urban scenarios. ^ We qualitatively and quantitatively evaluated the proposed method (algorithm), and experimentally demonstrate its advantages against state-of- the-art methods. We also share the datasets and the source code of the proposed method with the community. Furthermore, in summary, in the above description we proposed a novel depth image completion technique based on sparse consecutive measurements of preferably a non-repetitive circular scanning (NRCS) lidar, demonstrating the capabilities of a new, compact, and accessible sensor technology for dense range mapping of highly dynamic scenes. In an embodiment our deep network called proposed spatio-temporal deep network is composed of a spatio-temporally (ST) extended U-Net architecture, which accepts a very sparse range data sequence as input and produces a dense depth image stream of the same field-of-view ensuring a high level of spatial details and accuracy. For evaluation, we have constructed a new urban dataset, that – to our best knowledge as the first open benchmark in this field – comprises various simulated and real-world NRCS lidar data samples, allowing us to simultaneously train our model on synthetic data with Ground Truth (GT), and to validate the result via real NRCS lidar measurements. Using this new dataset, we have shown the superiority of our method against a densified depth map obtained from the raw sensor stream, and against two independent state-of-the-art deep-learning based lidar-only depth completion methods. IV. Temporal upsampling A. Introduction In summary, in the next sections, we propose a framework that enables the online generation of virtual point clouds relying only on a previous camera measurement and point cloud and a current camera measurement (the point cloud originates from the depth completion method according to the invention detailed above, beside it, in this embodiment, camera measurements are also captured). The continuous usage of the proposed pipeline – see the details thereof below, especially in Section IV.C – generating virtual lidar measurements makes the temporal up-sampling of point clouds possible. The only requirement of this embodiment is a camera with a higher frame rate than the lidar equipped to the same vehicle, which is usually provided. The data is obtained preferably from the depth information image generated based on the input depth information images originating from lidar device with non- repetitive scanning pattern utilized in the above depth completion method. If this input is utilized then – considering the FOV of this lidar device – only a single optical image is sufficient to have correspondence between the depth data and the corresponding data coming from optical image(s). The pipeline first utilizes optical flow estimations from the available camera frames. Next, optical expansion (or, equivalently, motion in depth, see below) is used to upgrade it to 3D scene flow. Following that, ground plane fitting is made on the previous lidar point cloud. Finally, the estimated scene flow is applied to the previously measured object points (i.e. points corresponding to the objects, see also below) to generate the new point cloud. The efficiency of the framework of the present embodiment is proven in performance compared to prior art achieved on the KITTI dataset using measurements of a Rotating Multi-Beam (RMB) lidar sensor. It is noted that theoretically, the steps of the present embodiment can be performed independently from the above detailed depth completion method according to the invention. In this case the initial depth data would not be connected to the output depth information image (formally, we would not start from initial depth data corresponding the output depth information image, but simply from a general initial depth data). In summary, if the above steps are steps of an embodiment, depth completed data originating from data recorded by a lidar device with non-repetitive scanning pattern will be utilized for these steps as an input. But, if these steps would be considered independently from the depth completion data, initial depth data of any origin could be used as an input. It is in line with these statements, that herebelow the details are shown simply for a lidar and for initial depth data originating from any source. In other words, the tests below were performed on RMB lidars, but framework of the present embodiment does not depend on the type of the lidar (generally, of the depth data generating device), any kind of depth data may be the input. In researches related to advanced driver-assistance systems (ADAS), intelligent vehicles or autonomous driving, passive cameras and lidars are usually basic components of the sensor systems equipped for the given vehicle. In this way, sensor fusion (Nagy, B.; Benedek, C. On-the-Fly Camera and lidar Calibration. Remote Sensing 2020, 12.) is frequently applied by utilizing the benefits of both sensors to solve different problems like 3D object detection (Wu, X.; Peng, L.; Yang, H.; Xie, L.; Huang, C.; Deng, C.; Liu, H.; Cai, D. Sparse Fuse Dense: Towards High Quality 3D Detection With Depth Completion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp.5418 5427.) or road detection (Chen, Z.; Zhang, J.; Tao, D. Progressive LiDAR adaptation for road detection. IEEE/CAA Journal of Automatica Sinica 2019, 6, 693 702.). The advantages compared to each other include but are not limited to colour imaging, high resolution (millions of pixels) and high framerate (generally, at least 30 FPS) in case of cameras, and working in the dark or the possibility of direct depth measurement in case of lidars. Lidars not only have much lower resolution (a few hundreds of thousands of measurement points) and measurement frequency (usually 5-20 Hz) than the passive visual sensors, but in most cases, it suffers the problem of spatial and temporal resolution also competing with each other. Meaning, increasing the spatial (horizontal) frequency decrease the temporal one and vice versa. This fact is unfortunate, as both of them being high can be essential in decision- making systems of intelligent vehicles (e.g., we should recognize an object accurately and as soon as possible). Thanks to the depth completion method according to the invention, the output of which is utilized in the present embodiment, the input spatial resolution is advantageously high, and due to the present embodiment also the temporal resolution can be enhanced. Accordingly, we propose a solution in the present embodiment to increase the measurement frequency of lidar sensors virtually; this enables the maximization of angular resolution. Also, enhancing temporal resolution beyond the limit is already important in itself to avoid hazardous scenarios, as dynamic traffic participants can be present with high acceleration, deceleration, or angular acceleration. The present embodiment can be applied in the presence of a lidar sensor and a calibrated camera (intrinsic and extrinsic) with a higher frame rate. If the whole 360° field of view of the lidar is covered by cameras, we can generate circular lidar frames (see Figs.12A-12B). With a single camera, we can only generate virtual lidar frames in the field of view of the camera (because it is actually the camera that tells you where the moving objects are). Since RMB lidars are of 360 degrees, we need more cameras to cover the whole field of view. The colourings in Fig.12B give an example of how different lidar parts (different colours) were generated with 7 cameras. Thus, circular lidar frames can be generated. Fig.12A shows lidar (Pt−1) and camera (It−1) measurements at timestamp t - 1. Fig. 12B shows camera measurements (It) at timestamp t and the generated virtual lidar measurement (Pt,v) to the given time moment. The Point cloud and the optical images show that these data have been recorded in a road-crossing (i.e. where four roads encounter). Because of the views – there are also intermediate views shown by optical images not just those which see along the roads – there are many common details in the optical images. Accordingly, based on the seven optical images we can form a very good vision idea about the road-crossing (see e.g. the left and left-down image, in both of those images two vans can be observed one after the other, see also below). With the comparison of optical images around the point cloud in Fig.12A and Fig. 12B it can be observed what has been changed between timestamps t-1 and t. The most remarkable differences can be observed at the left and left-down images, since the vans proceeded a little between the two images. Furthermore, in Fig. 12B colormap around the camera images and in the point cloud indicates which image (with a common field of view) was used to generate the given point cloud part. Thus, Figs.12A-12B show illustration of the up-sampling problem and our solution on the Argoverse dataset (Chang, M.F.; Lambert, J.; Sangkloy, P.; Singh, J.; Bak, S.; Hartnett, A.; Wang, D.; Carr, P.; Lucey, S.; Ramanan, D.; et al. Argoverse: 3D Tracking and Forecasting With Rich Maps. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.). In the dataset, camera images of Fig. 12B (at timestamp t) exist, but the corresponding lidar measurement does not. We generated this (coloured) lidar point cloud utilizing Pt-1, It-1 and It. In summary, in the present embodiment deep network-based optical flow estimation is made between the previous and the current camera frames. After that, we use another deep network to upgrade the optical flow to scene flow. Next, we estimate the ground points in the last available lidar frame in order to establish a ground model and determine measurements on the ground model in the next timestamp. Finally, the estimated displacements are applied to the last measurement points to generate the virtual point cloud to the current timestamp. While it is widespread to apply cameras to spatially up-sample lidar point clouds, and there are various solutions available (Zhao, S.; Gong, M.; Fu, H.; Tao, D. Adaptive Context-Aware Multi-Modal Network for Depth Completion. IEEE Transactions on Image Processing 2021, 30, 5264 5276. or Hu, M.; Wang, S.; Li, B.; Ning, S.; Fan, L.; Gong, X. PENet: Towards Precise and Efficient Image Guided Depth Completion. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021, pp. 13656 13662.). Thus, these apply spatial upsampling. Only a few research considered relying on cameras for the temporal up-sampling but with totally different approach (Rozsa, Z.; Sziranyi, T. Temporal Up-Sampling of lidar Measurements Based on a Mono Camera. In Proceedings of the Image Analysis and Processing ICIAP 2022; Sclaroff, S.; Distante, C.; Leo, M.; Farinella, G.M.; Tombari, F., Eds.; Springer International Publishing: Cham, 2022; pp.51 64.; Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.05192, 2023 or Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst.2022, 28, 1611 1620.). The approach of the last article needs multiple camera frames. Some research relies even on the generating between lidar frames at all (Lu, F.; Chen, G.; Qu, S.; Li, Z.; Liu, Y.; Knoll, A. PointINet: Point Cloud Frame Interpolation Network. In Proceedings of the Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021., and Xu, J.; Le, X.; Chen, C. SPINet: self-supervised point cloud frame interpolation network. Neural Computing and Applications 2022.). This approach needs future frames. See Section IV.B below for details. The advantages of our proposed pipeline compared to previous approaches: ^ Point cloud prediction methods – like (Weng, X.; Wang, J.; Levine, S.; Kitani, K.; Rhinehart, N. Inverting the Pose Forecasting Pipeline with SPF2: Sequential Pointcloud Forecasting for Sequential Pose Forecasting. In Proceedings of the Proceedings of (CoRL) Conference on Robot Learning, 2020. or Lu, F.; Chen, G.; Li, Z.; Zhang, L.; Liu, Y.; Qu, S.; Knoll, A. MoNet: Motion-Based Point Cloud Prediction Network. IEEE Transactions on Intelligent Transportation Systems 2021, pp. 1 11.) – need five or more preceding frames to generate a virtual measurement (we need only one). ^ The pipeline in the present embodiment adapts to the point cloud characteristics and generates virtual clouds with similar characteristics to the real measurements. The point-level transformation of the system explains this fact. It can be observed by seeing our result on different lidar sensors (e.g., Fig.13 and Figs.21A-23I). ^ For our solution of this embodiment, advantageously, ego-motion of the vehicle does not need to be known or estimated as in previous works (Rozsa, Z.; Sziranyi, T. Temporal Up-Sampling of lidar Measurements Based on a Mono Camera. In Proceedings of the Image Analysis and Processing ICIAP 2022; Sclaroff, S.; Distante, C.; Leo, M.; Farinella, G.M.; Tombari, F., Eds.; Springer International Publishing: Cham, 2022; pp.51 64.; Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023.). Camera measurements in time moments when complete lidar frames are not available allow our method to generate virtual point clouds and, in this way, enable the temporally up-sampling. The problem and result are illustrated in Figs.12A-12B and 13 (see also above about Figs.12A-12B). Fig.13 shows details of the virtual point cloud generated with inputs visible in Figs.12A-12B (it is noted that Fig.12A shows Pt-1, as well as Fig.12B shows Pt). Fig.13 shows point clouds for three different time instants (the last for t+1 only for illustration purposes), with the help of slightly different colouring. Colormap in Fig. 13: Green (lightest) – Pt-1 (last available measurement), Blue (next in the details, where displacement is shown) – Pt,v (generated point cloud), and Red (last in the details, where displacement is shown) – Pt+1 (future point cloud, only serves illustration purposes). Looking into the encircled enlarged part of the point clouds, it is visible that both dynamic and static objects occupy intermediate positions (in terms of position and orientation) in t (the timestamp of the generation) relative to t- 1 and t+1, as it was expected. This phenomenon can be observed best in the left- down circle which shows a displacement in the middle of the circle (starting a little bit above the middle), that the different colours are incrementally beside each other from the top of the feature (starting with the lightest). ‘P’ and ‘I’ means lidar and camera measurements respectively, ‘v’ index indicates the virtually generated point clouds and ‘t’ is the timestamp. In the remaining part of the description, the following notation is used: The data points (with arbitrary dimension) and an array of data points are indicated with the same letter, but in the latter case, it is bolded (e.g., point clouds – P P, images – I I). A.1 Contribution The main contributions of the present embodiment are the following: ^ We propose a framework, applying optical principles (flow and expansion) to solve one of the critical problems of autonomous driving researches, namely the balancing between the spatial and temporal resolution of 3D lidar measurements. ^ We extend the state-of-the-art with a new optical flow calculation method, enabling real-time run of our system and temporal up-sampling of lidar measurements. ^ The baseline was enhanced among others (see below) by ground estimation, which ensured virtual measurement generation with higher accuracy. ^ Our proposal includes motion vector estimation (point-wise) of surrounding objects (without the requirement of solving the challenging dynamic object segmentation problem, see Chen, X.; Li, S.; Mersch, B.;Wiesmann, L.; Gall, J.; Behley, J.; Stachniss, C. Moving Object Segmentation in 3D LiDAR Data: A Learning-based Approach Exploiting Sequential Data. IEEE Robot. Autom. Lett. (RA-L) 2021, 6, 6529 6536.). This is a significant advantage compared to prior art approaches. A.2. Outline of the sections below The below sections are organized as follows: Section B. studies the related literature. Section C. introduces the proposed pipeline, i.e. the solution according to the present embodiment. Section D. shows performance measures from our tests, while Section E. provides an ablation study and further discussion. Finally, the conclusions are drawn in Section F. B. Related Works This section is divided into five subsections. First, spatial up-sampling of point clouds is investigated, as it has mature methods and is similar to our approach in terms of sensor fusion (which is also done according to the invention). Second, future frame prediction literature is introduced, which is a recent research interest and similar to our approach as it aims for the virtual lidar frame generation but ignores actual information. Third, lidar frame interpolation methods are examined, which have the same purpose as our approach, but they can be applied offline, as they require a 'future' frame for the generation of in-between frames. Fourth, some earlier approaches for temporal up-sampling are discussed. Finally, some prior art patent documents are also described. B.1. Spatial Up-sampling The spatial up-sampling of lidar point clouds is usually driven by camera images. These methods aim to estimate depth for every pixel of a depth image having the same resolution as a corresponding RGB image with an input very sparse depth image initialized with projected lidar data points. That is why this problem is commonly referred to as depth completion (Uhrig, J.; Schneider, N.; Schneider, L.; Franke, U.; Brox, T.; Geiger, A. Sparsity Invariant CNNs. In Proceedings of the International Conference on 3D Vision (3DV), 2017.). As the targeted resolution comes from an image, images help (with a few exceptions, e.g., Premebida, C.; Garrote, L.; Asvadi, A.; Ribeiro, A.; Nunes, U. High-resolution lidar-based depth mapping using bilateral filter.2016, pp.2469 2474 or Ku, J.; Harakeh, A.;Waslander, S.L. In Defense of Classical Image Processing: Fast Depth Completion on the CPU. In Proceedings of the 201815th Conference on Computer and Robot Vision (CRV), 2018, pp.16 22.) in the process of pixel-wise depth estimation. There are different approaches to do this, like semantic-based up-sampling Schneider, N.; Schneider, L.; Pinggera, P.; Franke, U.; Pollefeys, M.; Stiller, C. Semantically Guided Depth Upsampling. In Proceedings of the Pattern Recognition. Springer International Publishing, 2016, Vol.9796, pp. 37 48.), but deep learning-based methods (e.g., Zhao, S.; Gong, M.; Fu, H.; Tao, D. Adaptive Context-Aware Multi-Modal Network for Depth Completion. IEEE Transactions on Image Processing 2021, 30, 5264 5276. or Hu, M.; Wang, S.; Li, B.; Ning, S.; Fan, L.; Gong, X. PENet: Towards Precise and Efficient Image Guided Depth Completion. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021, pp. 13656 13662.) are currently the most successful ones. These methods are similar to our solution for the temporal up-sampling problem in the sense that our approach is also camera-driven, but we utilize richer camera information to do the up-sampling with another processing tools, and we expect the output to be the same (temporal) resolution as the corresponding camera has. B.2. Point Cloud Prediction Point cloud prediction or sequential point cloud forecasting is a hot research topic. The methods of solving this problem are important to our research. They also generate virtual point clouds, as our proposal uses previous frame information. In this way, they can be alternatives to the pipeline we propose. However, they aim differently; they try to predict the future. That is why they have a different approach with relevant differences, too. In their alternative approach, they do not use current (camera) information, which even theoretically limits their accuracy. Also, these typically neural-network-based methods (like Weng, X.; Wang, J.; Levine, S.; Kitani, K.; Rhinehart, N. Inverting the Pose Forecasting Pipeline with SPF2: Sequential Pointcloud Forecasting for Sequential Pose Forecasting. In Proceedings of the Proceedings of (CoRL) Conference on Robot Learning, 2020.; Wencan, C.; Ko, J.H. Segmentation of Points in the Future: Joint Segmentation and Prediction of a Point Cloud. IEEE Access 2021, 9, 52977 52986.; Lu, F.; Chen, G.; Li, Z.; Zhang, L.; Liu, Y.; Qu, S.; Knoll, A. MoNet: Motion-Based Point Cloud Prediction Network. IEEE Transactions on Intelligent Transportation Systems 2021, pp. 1 11.; Deng, D.; Zakhor, A. Temporal LiDAR Frame Prediction for Autonomous Driving. In Proceedings of the 2020 International Conference on 3D Vision (3DV); IEEE Computer Society: Los Alamitos, CA, USA, 2020; pp.829 837. or He, P.; Emami, P.; Ranka, S.; Rangarajan, A. Learning Scene Dynamics from Point Cloud Sequences. International Journal of Computer Vision 2022, 130, 1 27.) have other drawbacks too: ^ As several previous frames are necessary for the prediction (usually 5), it implies that the motion model is embedded in the system resulting in a loss of generality. ^ End-to-end training of point cloud prediction could result in weak robustness against different datasets and point cloud characteristics. ^ Most of these methods operate only near real-time and in close range. We present a comparison in Section D below to prove the superiority of our solution to these types of methods in the temporal up-sampling problem. B.3. Point Cloud Interpolation Point Cloud Interpolation methods like Liu, H.; Liao, K.; Zhao, Y.; Liu, M. PLIN: A Network for Pseudo-LiDAR Point Cloud Interpolation. Sensors 2020, 20, 1573.; Liu, H.; Liao, K.; Lin, C.; Zhao, Y.; Guo, Y. Pseudo-LiDAR Point Cloud Interpolation Based on 3D Motion Representation and Spatial Supervision. IEEE Transactions on Intelligent Transportation Systems 2021, pp.1 11.; Lu, F.; Chen, G.; Qu, S.; Li, Z.; Liu, Y.; Knoll, A. PointINet: Point Cloud Frame Interpolation Network. In Proceedings of the Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021.and Xu, J.; Le, X.; Chen, C. SPINet: self-supervised point cloud frame interpolation network. Neural Computing and Applications 2022. have similar intermediate aims as our one, namely generating virtual lidar frames between two real lidar measurements (but we do not generate between two lidar measurements, but based only on an earlier one). However, they have a final goal, offering a solution to the frequency mismatching problem of lidar and cameras, which is very distinct from ours. In this way, these methods cannot be applied to our problem. They work offline, as they utilize frames from time t+1, naturally not available at time moment t to generate measurements to timestamp t. Still, in our preferably online method, we outperform these offline methods in certain performance measures (see Section D below). These methods may be used as preprocessing to our one if the synchronization of camera and lidar is not solved; in practice, in most cases, using camera frames closest in time to the lidar frame is considered accurate enough. Thus, offline interpolation is optional. Offline interpolation cannot be an alternative to our proposal in autonomous driving, while we provide an online substitute for these methods. B.4. Temporal Up-sampling Temporal up-sampling of measurements is not completely unknown to the literature (Beck, H.; Kuhn, M. Temporal Up-Sampling of Planar Long-Range Doppler LiDAR Wind Speed Measurements Using Space-Time Conversion. Remote Sensing 2019, 11.), yet, there are only very few researches available trying to solve the temporal up-sampling problem of point clouds. Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. refer to the temporal up-sampling problem as predicting future Pseudo-lidar frames, which could be directly used as an alternative to our proposal. However, they require three camera frames as input and two previous lidar frames. The only method which only needs two camera frames and one previous lidar frame (as the one proposed here) is published in Rozsa, Z.; Sziranyi, T. Temporal Up- Sampling of lidar Measurements Based on a Mono Camera. In Proceedings of the Image Analysis and Processing ICIAP 2022; Sclaroff, S.; Distante, C.; Leo, M.; Farinella, G.M.; Tombari, F., Eds.; Springer International Publishing: Cham, 2022; pp. 51 64. and Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023. In Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. motion-in-depth estimation network is proposed using two camera frames. This method cannot be directly applied to the temporal up-sampling problem. However, it should be mentioned here as we extended this method to be applicable to the given issue. Comparisons to these methods are presented in Section D below. On the top of the comparisons, many of the above prior art references are mentioned elsewhere, where characterised differences between them and the proposed method are investigated. B.5. Prior art patent documents In US 9,872,011 B2 such densifying of depth information is disclosed, which is based on a video camera having higher frame rate. Optical flow, as well as perspective depth correction is applied in the document. An approach similar to the above is disclosed in WO 2020/104423 A1, in which optical flow is utilized and ground plane is handled. In CN 111612728 A a point cloud densification method based on binocular RGB images in which ground segmentation is applied is disclosed. Depth image completion is applied in US 2023/0245282 A1. In US 2023/0136235 A1 the sparse depth data along with image input are processed together in connection with point cloud densification. A vehicle driving safety early warning method using FastFlowNet is disclosed in CN 114983328 A. C. The Proposed Method From here, we focus on describing the generation of only one virtual point cloud measurement. As we can generate virtual measurements to any t time moment (with camera measurement), the following steps can be repeated as many times as you want, resulting in a temporally upscaled point cloud stream. In a variation of the present embodiment, it is an optical flow and expansion based deep temporal up-sampling of lidar point clouds. The pipeline corresponding to the proposed method of this embodiment has five important steps, and these are the following: 1. Estimate optical flow (ũ) between images acquired at t-1 and t. 2. Estimate optical expansion (s) and motion in depth ( ^ =1/s) (both of these may be the motion-in-depth related data, i.e. preferably one of them is estimated) from the previously estimated flow. 3. Estimate ground model and points on Pt-1. 4. Calculate scene flow, utilizing the estimations and lidar measurements from t-1. 5. Transform the object points with the estimated scene flow to generate the virtual measurement (Pt,v) at t. The differences from an exemplary baseline (Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020., for the differences, see also below) are in steps 1 and 3-5; in 2 we utilize the network introduced by them for motion-in-depth estimation. In step 1, utilizing FastFlowNet enables real-time application. Step 3 and utilizing it in step 4 (also applying the exact formula in step 5) increases the similarity to real measurements (see for further details in Table XII below). Maybe the biggest advantage of our pipeline compared to prior art approaches (e.g. Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.05192, 2023) is that differentiating static and moving object is unnecessary; movement calculation of surrounding objects is included. The pipeline of the proposed method in the present embodiment is illustrated in Fig. 14 for the inputs of Figs. 15A-15C. Fig. 14 shows the proposed pipeline of generating virtual point clouds to the intermediate time stamp t (or 'future pseudo- lidar' frame prediction) for up-sampling. In the further figures, some of the illustrative images of Fig.14 are shown in larger size, i.e. in the further figures the exemplary scene illustrated in Fig.14 is further illustrated. Figs.15A-15C show example inputs, such as in Fig.15A lidar (Pt−1) measurement (as disparity image), in Fig.15B camera measurement (It−1), as well as in Fig.15C camera measurement (It). Figs.15B and 15C show the same images as the first optical image 102 and second optical image 104 of Fig.14 (these are optical images on which optical flow and optical expansion/motion in depth can be calculated, i.e. images recorded by a camera), but in larger size, the content of them can be better observed. Fig.15A corresponds to the initial depth data 100, but shows a different representation, see below for details. Beforehand using it, the camera intrinsic (Zhang, Z. A flexible new technique for camera calibration. IEEE Transactions on Pattern Analysis and Machine Intelligence 2000, 22, 1330 1334.) and camera-lidar (extrinsic) (Zhou, L.; Li, Z.; Kaess, M. Automatic Extrinsic Calibration of a Camera and a 3D LiDAR Using Line and Plane Correspondences. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2018, pp.5562 5569.) calibration should be executed. The depth image containing Zt-1 values are determined using the lidar- camera TL,C transformation (e.g. transformation matrix from the lidar coordinate system to the camera coordinate system) and K intrinsic matrix (3x3 matrix of camera parameters which projects a 3D point from 3D camera coordinate system onto the image plane). For better visualization, the disparity image is illustrated here (disparity is d = b * f/Z with b baseline and f focal length, and it depends on 1/Z, i.e. inverted distance; this is a virtual disparity image, it gives disparity with a virtual stereo pair with a shift of the baseline to convert depth values to disparity values for illustration purposes). The intermediate results of the pipeline of the present embodiment with exemplary parameters and inputs (point cloud and camera image) will be shown enlarged in the detailed explanation part. The results of certain steps performed (optical flow estimation, ground segmentation, etc.) have a physical meaning, which has been visualized in each case. It is noted that on the contrary to some offline prior art approaches, the present embodiment is performed online, i.e. not needing future data for performing the steps for temporal upsampling (point cloud with time stamp t is generated based on image – camera measurement – of time stamp t). C.1. Optical flow estimation In the framework of the present embodiment, the optical flow is obtained by means of a neural network (see Fig. 14 and also below). However, also theoretical background is given (the expression can also be utilized in another equation). The flow field describes the velocity of image pixels as: where u and v are the flow components, between pt and pt-1 image points of different time stamps with x and y are the pixel coordinates. Accordingly, here, the two images (pt-1 and pt) are utilized, which are determined for every x, y coordinates, having a value at every x, y pixel coordinates, just like ũ (in Eq. (11) the respective quantities are given as column vectors according to the transposed representation). Thus, the optical flow can be calculated for every x, y pixel coordinates. We preferably adapt FastFlowNet (Kong, L.; Shen, C.; Yang, J. FastFlowNet: A Lightweight Network for Fast Optical Flow Estimation. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021.) to estimate the optical flow field (for the first auxiliary neural network, see also below). FastFlowNet is a lightweight model for fast and accurate prediction with only 1.37M parameters enabling us real-time run. It uses a coarse-to-fine paradigm with a head- enhanced pooling pyramid (HEPP) feature extractor to intensify high-resolution pyramid features while reducing parameters. Besides, the center dense dilated correlation (CDDC) layer is applied to compact cost volume that can keep a large search radius with reduced computation time, and a shuffle block decoder (SBD) is implemented into pyramid levels to accelerate the estimation. Further details can be found in the original paper cited at the beginning of this paragraph. Illustration of a resulted flow field coloured according to the Middlebury colour code (Baker, S.; Scharstein, D.; Lewis, J.; Roth, S.; Black, M.; Szeliski, R. A Database and Evaluation Methodology for Optical Flow. International Journal of Computer Vision 2007, 92, 1 31.) with input images of Figs.15A-C can be seen in Fig.16, which latter shows that FastFlowNet resulted optical flow field from inputs of Figs. 15A-15C. Fig.16 shows the same image as optical flow data 108 in Fig.14 (about Fig.16, see at Fig.17) with grayscaled Middlebury colour code (see the previously referred article). C.2. Motion-in-depth estimation Similarly to the optical flow, in the framework of the present embodiment, the motion- in-depth estimation is obtained by means of a further neural network (see Fig.14 and also below). However, also theoretical background is given herebelow (the expression can also be utilized in another equation). Motion-in-depth, by definition, is: where Zt and Zt-1 are the depth values at time moments t and t-1 respectively. In the expression Z(x,y) is a depth map which stores a depth value at each x,y pixel coordinate. Accordingly, ^ is the ratio of depths of the current and previous time instants. In this step, we estimate an 'image' of depth ratios which will relate our 2D flow estimations to a 3D motion estimation (as we see later). In this theoretical background, ^ is the ratio of depth values, where it is given in the arguments of Z where it is evaluated: Zt-1 is evaluated for pt-1 (pt-1 points to a point with x, y coordinates and Z is the depth there) and Zt is evaluated for pt-1 + ũ(pt-1) (the latter is ũ at t-1, i.e. ũ starting from t-1 same as in Eq. (11), which gives this ũ), i.e. pt based Eq. (11). Zt and Zt-1 are both functions of x, y coordinates. Thus, ^ is a function of x, y coordinates. Optical expansion (in case of not rotating scene elements and orthographic camera model) is the reciprocal of the motion-in-depth (see further details in Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.). Thus, if we can estimate the scale change of the scene elements, we will get how much they moved closer or farther away. This principle is utilized in the work Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. where local affine transformation was used to estimate scale changes: where A is 2x2 matrix describing the local affine transformation, c index indicates a center pixel of a given coordinate; pt and pt-1 corresponding pixels of images acquired at different time stamps are related by equation (11). Accordingly, Eq. (12) is the definition of ^, however, Eq. (14) shows how to get ^ based on the local changes of two images (pt and pt-1). Later on, the estimated scale ratios were used to train a network (namely, the second auxiliary neural network, see also below) for the estimation of depth change. We applied this network to get motion-in-depth estimations. The image generated from the motion of depth estimation is illustrated in Fig.17, which shows motion-in- depth results with the network of (Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.) from inputs of Figs.15A-C and 16. Fig.17 shows the same image as motion-in-depth related data 112 of Fig.14. The motion-in-depth related data 112 shown in Fig.17 is the motion in depth itself. The ^ values, which can vary between about 0.5-2 (how the new depth is proportional to the old depth, the approach, the distance determines the ratio of the Z's) are rescaled to a grayscale image. For example, the car is very dark, it has a minimum ratio (if we stick to 0.5 as an example: for a value of 0.5, the car is half as far away at time t as it was at time t-1). A similar interpretation can be made for the Fig.16, where the magnitude of the lateral displacements is represented by the colours according to the optical flow (again, we refer to the colouring of the car). C.3. Ground model estimation Ground model estimation is applied in the present embodiment for the following reason: Finding the corresponding locations of points of Pt-1 at time t (estimating 3D scene flow) differs from the problem of estimating lidar measurement at time t. Due to the movements and sensor characteristics, the objects will be hit by the sensor in different parts. This will mean a big difference between the scene flow estimated points and real measurements in the case of the ground and far points. The phenomena can be observed, e.g., in the last column of Figs.21A-21I, i.e. in Figs. 21C, 21F and 21I, and see also the importance of our compensation in Table XII. So, in our solution, in the present embodiment, we estimate a ground model, and for the ground points, no displacement is applied. This is based on the assumption that as we move on the plane and measure, the sensor characteristic (laser emitting angles) do not change. Thus, we will find points approximately at the same distances on this plane. (A more complex model could be applied; however, this simple assumption proved to be useful, accurate, and efficient.) MLESAC (Maximum Likelihood Estimation Sample Consensus, see Torr, P.; Zisserman, A. MLESAC: A New Robust Estimator with Application to Estimating Image Geometry. Computer Vision and Image Understanding 2000, 78, 138 156.), a variant of the RANSAC (Random Sample Consensus) robust model-fitting method, is preferably applied to fit a ground plane to the lidar points of Pt-1 with a reference normal vector of [010]T in camera coordinate system (column vector according to the widely applied conventions in the field). Estimated ground points are illustrated in Fig.18 showing segmented ground from input of Figs.15A-C. Ground points are coloured grey, and object points with black. Fig.18 shows the same point set as ground segmented depth data 106 in Fig.14. 3.4. Calculate 3D scene flow 3D scene flow is defined as the three-dimensional motion of 3D points, just as optical flow is the 2D motion of points in an image. For further explanation see Vedula, S.; Rander, P.; Collins, R.; Kanade, T. Three-dimensional scene flow. IEEE Trans. Pattern Anal. Mach. Intell.2005, 27, 475 480. Utilizing the estimated optical flow, the motion-in-depth values, and the depths, Zt-1 (projected from the lidar), the corresponding depth values, Zt can be determined by rearranging the following equation (Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.): In Eq. (15), Ũ (3D scene flow) is given as the difference of two point clouds, Pt and Pt-1. This difference is expressed by the help (of inverse) of K intrinsic matrix, which is multiplied in the third expression of Eq. (15) with multiplications of x, y coordinate dependent functions (Z, p). Moreover, in the fourth expression of Eq. (15), inverse of K is multiplied both from left and right sides: from the right side, with ‘ ^*pt+pt’, since, based on Eq. (11) above pt= ũ+pt-1. Based on the fourth expression, Ũ can be calculated (for x, y, z coordinates), and if it has been calculated, also Zt can be expressed from Eq. (15). In Eqs. (15) and (16), ũ and ^ (both depending on x, y coordinates) obtained by respective neural networks according to the approach of the present embodiment are utilized. In this step, calculations are made using the results of the neural networks (ũ and ^), so here the expressions of Eqs. (15) and (16) are not inserted only for illustration of the theoretical background, but these also show the calculation process. The displacement (|Ũ|) estimated from the 3D scene flow is illustrated in Fig.19. In Fig. 19 estimated displacements (for ground points assumed to be 0) are shown corresponding to lidar data points (Pt-1) from the input data of Figs.15A-15C. Fig. 19 shows the same point set as 3D flow data 114 of Fig.14. In Fig.19 the displacements are illustrating with different colours (see the colour bar of Fig.19), with darker points corresponding to smaller displacements (cf. with the ground points having 0 displacement) and with lighter points corresponding to larger displacements. 3.5. Generating virtual point cloud The virtual point cloud to time t is generated from the points of Pt by the following rule: where the upper index i indicates the ith point of the point cloud and G represents ground points in the point cloud P. The above equation means that scene flow is only applied for non-ground points (as it is visible in the case of Fig.19). Addition of Ũ to Pt-1 can be understood based on Eq. (15). As 3D scene flow is estimated for each point, movement estimation of moving objects is included in the pipeline and we do not have to consider them separately. All the points of static objects (e.g. parking cars in the exemplary point set) should have the same scene flow value (the ego-motion vector in opposite direction), and points of a dynamic object should have some other value (same for the same objects). This is because of the displacement of the car holding the recording devices (camera and lidar: the changes of the lidar data will be obtained from the changes can be observed on the subsequent images from the camera). By our proposed method these can be handled automatically. The resulted point cloud is illustrated in Fig. 20, which shows estimated virtual measurement (Pt,v) from the inputs of Figs.15A-15C. Fig.20 shows the same point set as subsequent depth data 116 of Fig.14. The appropriate estimation of static and moving objects can be observed in Fig.20, where approaching vehicle points have about 2 m estimated displacement. In comparison, static environment points (including parking cars) have a roughly 1 m estimated displacement. By repeating all the steps listed here for every t without lidar measurement (but with camera measurement), a temporally up-sampled point cloud sequence can be produced. In the present embodiment of the depth completion method according to the invention, - initial depth data corresponding to the output depth information image of the scene corresponding to an initial point of time (as given above the data incorporated into the output depth information image is generally collected from multiple integration windows, but a characteristic time instance (e.g. starting or final time instance) of collecting can be assigned to the data, which will be here the initial point of time, point of time called time stamp elsewhere), and - a first optical image of the scene corresponding to the initial point of time and a second image of the scene corresponding to a subsequent point of time being subsequent to the initial point of time (i.e. subsequent point of time is after the initial point of time) are utilized (see Fig.14 illustrating this embodiment, where, in the left side of the figure, initial depth data 100, a first optical image 102 and a second image 104 are shown). The depth data may be e.g. point set data (point cloud data) or depth information image (the latter is preferably range image). Thus, the initial depth data may be the output depth information image (e.g. output range image) of the previous part of the depth completion method itself, or, when the depth data is point set data, it can be obtained (generated) by a transformation from that. The above quantities are inputs of the present embodiment, thus the above introduction of these inputs shows that in the present embodiment the depth completion method according to the invention is continued from its output. The present embodiment of the method further – i.e. after obtaining output depth information image – comprises (most of the steps and substeps are illustrated in Fig.14 corresponding to an embodiment for point sets) - an input generating step (this is meant that a step for generating inputs for generating 3D flow data) comprising the substeps of (the substeps below can be performed in any order – cf. with that these are two branches in Fig.14 – the outputs of them are to be utilized afterwards) o optical flow data between the first optical image and the second image is generated by means of an optical flow generation module based on a first auxiliary neural network (see operational step S115 in Fig.14 illustrated by a stylized neural network labelled by ‘FastFlow network’, its output is optical flow data 108), and motion-in-depth related data is generated based on the first optical image, the second image and the optical flow data (see the three inputs in Fig. 14) by means of a motion in depth module based on a second auxiliary neural network (see operational step S120 in Fig.14 illustrated also by a stylized neural network labelled by ‘Motion-in-depth-network’, its output is motion-in-depth related data 112, which can be either the motion in depth itself, i.e. ^, or the optical expansion, which is chosen for further procession), and o ground depth data and non-ground object depth data are determined by segmenting the initial depth data (see operational step S110 of ground segmentation in Fig. 14, where – after the homogeneous point set corresponds to initial depth data 100 – in ground segmented depth data 106, ground depth data 106a and non-ground object depth data 106b are shown by grey and black, respectively), - after the input generating step has been performed, 3D flow data corresponding to changes compared to the initial depth data is generated based on the ground depth data, the non-ground object depth data and the motion-in-depth related data (see operational step S125 labelled by ‘scene flow estimation’ with output 3D flow data 114; as touched upon above; scene flow is calculated according to Eq. (15), it is an estimation because its exact determination is theoretically not possible, since there are no one to one point matches between lidar point clouds and in some cases since the optically obtained data – optical flow data and motion-in-depth related data – have to be combined with the point cloud type data), then - subsequent depth data is generated based on (see operational step S130 with the output of subsequent depth data 116) o the ground depth data, and o the non-ground object depth data modified based on 3D flow data. According to the above, 3D flow data correspond to changes compared to the initial depth data, i.e. it contains information how the initial depth data is to be modified. Based on the ground segmentation, the changes for the ground depth data can be considered to be zero by definition. This is in line with the fact that 3D flow data is applied only to non-ground object depth data. As can observed in 3D flow data 114 in Fig.14, as well as in Fig.19, to the point other than the ground shifting values (for shifting in space, i.e. in x, y, z coordinates) are assigned. Based on this information the transformed subsequent depth data (see subsequent depth data 116 in Fig.14, as well as Fig.20) can be generated. Performing the steps of the present embodiment can also be an option for the depth completion system according to the invention just like other optional features of the depth completion method according to the invention. In a respective system the steps of the present embodiment are performed by means of a temporal upsampling processing subunit integrated into the – main or depth completion – processing unit. It is noted to the above that when 3D flow data is generated the optical flow data can be used directly or indirectly, in the latter case pt = ũ + pt-1 expression is utilized, and then only ^ is utilized, i.e. the motion-in-depth itself from the motion-in-depth related data (see Eq. (15) where motion-in-depth is used directly, but the utilization of this expression can be understood based on this). There are several ways for generating the subsequent depth data. Hereby, two ways are given in detail for the above mentioned two representations of depth data (point set, range image). In the above figures (particularly in Fig.14 and subsequent figures) and description part, the generation of the subsequent depth data is detailed for point sets. In this case in the ground segmentation ground points are segmented, i.e. a respective ground point set and object point set can be determined. In connection with the optical flow and motion-in-depth, the above cited Ũ is determined. The 3D flow data is determined based on Eq. (16): it is zero for the ground points (by definition) and is utilized for other points. Accordingly, in the point set representation the subsequent point set can be generated by unifying the following two point sets: the ground point set and the object point set modified based on 3D flow data (the latter is an addition in this case, the value of 3D flow data for each coordinate (x, y, z) is added the respective coordinate values of the point set, see Eq. (16)). The above can also be proceeded in the range image representation. In this case, the ground segmentation can be performed within the range image itself (there are available approaches) or by transforming the range image into a point set, performing the ground segmentation and transforming the result back to range image representation (then the range image parts of the ground range image and the object range image can be determined). In the determination of the 3D flow data, in this case Zt corresponding to the depths has to be calculated (determined) after the determination of Ũ based on Eq. (15). In this case the ground segmentation is available for the range image for t-1 (i.e. Zt-1). Based on the result of the ground segmentation, Zt is applied for the object range image part of Zt-1 and the other part (ground range image part thereof) remains unchanged. A 3D point cloud is basically a sequence of x,y,z points, and the scene flow U=dx,dy,dz offsets are written in a similar form. If we want to do this in a range image, we need to apply ũ (optical flow) on the pixel coordinates and ^ depth ratio on the depth value. Moreover, a range image representation result can also be obtained in such a way that the result of Eq. (16) is transformed from point set representation to range image representation. In this case it is e.g. possible that we have started from range image representation for the initial depth data and afterwards it has been transformed to point set representation for the ground segmentation. Referring to the above description, a summary about Fig.14 is given. Fig.14 gives good illustration for the steps and substeps by the help of a flowchart. It is illustrated in Fig.14 at its left side that the inputs in the illustrated embodiment are It-1 (the first optical image 102) and It (the second image 104), as well as Pt-1 (initial depth data 100). At all of the steps and substeps, including the inputs corresponding (related) example data (images and point sets) are shown as an illustration. In accordance with this, for the point sets, we can see a homogeneous point set as the initial depth data 100, the ground depth data 106a and non-ground object depth data 106b (the latter may be called simply object depth data or object depth data different from the ground) are separated by the help of grey and black colours after ground segmentation. Furthermore, after the scene flow estimation it is shown for 3D flow data 114 by the help of colours how much the individual points of the point set are to be shifted according to the 3D scene flow. Finally, a transformed (homogeneous) point set is illustrated for the subsequent depth data 116. D. Results Here, we provide a quantitative and qualitative evaluation of the proposed pipeline. D.1. Data for Comparison Ground truth is unavailable for generating point clouds to time stamps without measurement. That is why we generated point clouds to time stamps with measurements (as others in similar problems). This means that for a sequence with x (number of frames) consecutive lidar frames, we can generate x-1 succeeding point clouds (using the previous one as input) and this way the x-1 point clouds can be compared to the x-1 real measurements. As different competitors used different datasets with different conditions to evaluate their method, we also used more of them for a fair comparison. Subsection D.2 describes the evaluation procedure and introduces our results in the case of the Odometry dataset, subsection D.3 in the case of the Depth Completion dataset, while subsection D.4 shows qualitative examples. In our tests, the most commonly applied error metrics for point cloud generation are involved, namely Chamfer Distance (CD) and Earth Movers Distance (EMD). As the literature used different EMD-s, we define them in the subsections. The formula of CD: where Pt,v N_t,v x 3 and Pt N_t x 3 are the data points of the predicted (virtual) Pt,v and ground truth Pt point clouds respectively. The number of generated points (Nt,v denoted previously I the upper index as N_t,v) can differ from the number of data points of the ground truth Nt (denoted previously as N_t). That is why the denser point cloud is randomly downsampled to N for these scenarios, where N=min(Nt,v, Nt). Here the CD metric measures the goodness of the virtually generated point cloud. Namely, by generating point clouds with the method described that have been measured in reality. Our virtual point cloud generated for time instant t will always consist of as many points as we had at time instant t-1, because we generate offsets for these points one by one. However, the ground truth (measured) point cloud at time instant t will most likely not consist of as many points as the measured one at time instant t-1, so it will not be the same, but it must have the same number of points to evaluate the metric. Therefore, a point cloud with more points is undersampled (in practice, this means dropping a few points). D.2. Odometry dataset The papers used the KITTI Odometry dataset (Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2012.) for evaluation and downsampled the lidar frames to 16384 data points. We also did that, and after that, we only considered data points for which the proposed up-sampling is applied (points seen by camera no.2). Sequences 08-10 (altogether 6863 frames) were used for testing. The definition of Earth Movers Distance: where Φ is bijection, which calculates the point-to-point mapping between two point cloud Pt and Pt,v. EMD measures the similarity between point clouds by calculating the cost of the global matching problem. For EMD, the approximation of (Bertsekas, D.P. A distributed asynchronous relaxation algorithm for the assignment problem. In Proceedings of the 198524th IEEE Conference on Decision and Control, 1985, pp. 1703 1704.) is applied, which is used by (Fan, H.; Su, H.; Guibas, L. A Point Set Generation Network for 3D Object Reconstruction from a Single Image. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp.2463 2471.) and other literature. The results of our proposed pipeline in the dataset compared to other methods are visible in Table VII. The applied approaches (these are mentioned also above) the results of which is involved in Table VII are the followings: ^ MoNet (LSTM and GRU): Lu, F.; Chen, G.; Li, Z.; Zhang, L.; Liu, Y.; Qu, S.; Knoll, A. MoNet: Motion-Based Point Cloud Prediction Network. IEEE Transactions on Intelligent Transportation Systems 2021, pp.1 11. ^ SPINet: Xu, J.; Le, X.; Chen, C. SPINet: self-supervised point cloud frame interpolation network. Neural Computing and Applications 2022. ^ PointINet: Lu, F.; Chen, G.; Qu, S.; Li, Z.; Liu, Y.; Knoll, A. PointINet: Point Cloud Frame Interpolation Network. In Proceedings of the Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021. ^ Rigid body based up-sampling: Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023. Methods Application CD [m2] EMD [m2] MoNet (LSTM) Forecasting 0.573 91.79 MoNet (GRU) Forecasting 0.554 91.97 SPINet Offline interpolation 0.465 40.69 PointINet Offline interpolation 0.457 39.46 Rigid body based up- sampling Online interpolation 0.471 33.98 Proposed pipeline Online interpolation 0.486 28.51 Table VII – Quantitative evaluation and comparison of the proposed pipeline of the present embodiment on KITTI odometry dataset. In the case of CD, PointINet performed best; however, it cannot be applied online as it needs a 'future frame' for operation. From the alternatives, Rigid body-based up-sampling performed slightly better. We achieved the best performance in the case of EMD, which describes the similarity of point cloud distributions the best. D.3. Depth completion dataset As other competitors used the Depth Completion Dataset (Uhrig, J.; Schneider, N.; Schneider, L.; Franke, U.; Brox, T.; Geiger, A. Sparsity Invariant CNNs. In Proceedings of the International Conference on 3D Vision (3DV), 2017.) for evaluation, we tested our proposal of the present embodiment on this, too, on the validation subset containing 1000 frames. The performance measuring procedure (used by Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. and others) does not include downsampling. (Downsampling is usually used for faster calculation of the measures.) However, in this case, only points considered in the field of view of the camera image are cropped to 1216×256 pixel resolution (leaving out distant points). The Earth Movers definition applied by Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. is the following (without squared norm): Our results compared to other methods are in Table VIII. The applied approaches the results of which is involved in Table VIII are the followings: ^ Prediction: Deng, D.; Zakhor, A. Temporal LiDAR Frame Prediction for Autonomous Driving. In Proceedings of the 2020 International Conference on 3D Vision (3DV); IEEE Computer Society: Los Alamitos, CA, USA, 2020; pp. 829–837. ^ PLIN: Liu, H.; Liao, K.; Zhao, Y.; Liu, M. PLIN: A Network for Pseudo-LiDAR Point Cloud Interpolation. Sensors 2020, 20, 1573. ^ PLIN+: Liu, H.; Liao, K.; Lin, C.; Zhao, Y.; Guo, Y. Pseudo-LiDAR Point Cloud Interpolation Based on 3D Motion Representation and Spatial Supervision. IEEE Transactions on Intelligent Transportation Systems 2021, pp.1–11. ^ Future pseudo-lidar: Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo-LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. Methods Application CD [m2] EMD [m] Prediction Forecasting 0.202 11.498 PLIN Offline interpolation 0.21 - PLIN+ Offline interpolation 0.12 - Future pseudo-lidar Online interpolation 0.157 3.303 Proposed pipeline Online interpolation 0.141 0.806 Table VIII – Quantitative evaluation and comparison of the proposed pipeline of the present embodiment on KITTI depth completion dataset. One can see that in the case of CD, PLIN+ performed the best; however, this method cannot be applied to our problem in practice, as it needs a future point cloud. Apart from that, our proposed pipeline has the best performance in both measures. 4.4. Qualitative examples This subsection illustrates our method's performance compared to the prior art approaches through qualitative examples. These can be seen in Figs.21A-I, 22A-I, and 23A-I. These figures show for different exemplary scenes illustration of generated clouds with different methods together with the ground truth (black - virtual measurement, grey – ground truth measurement). The first columns of the example figures (Figs.21A, 21D, 21G; Figs.22A, 22D, 22G; Figs. 23A, 23D, 23G) show the generated point clouds from a distance, and the columns afterward (Figs.21B, 21E, 21H and Figs.21C, 21F, 21I; Figs.22B, 22E, 22H and Figs.22C, 22F, 22I; Figs.23B, 23E, 23H and Figs.23C, 23F, 23I) contain zoomed parts, i.e. enlarged details (left-bottom part and right-up part, respectively). Furthermore, for the comparison, - the results of the first row (Figs.21A, 21B, 21C; Figs.22A, 22B, 22C; Figs.23A, 23B, 23C) have been obtained by the approach of Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.; - the results of the second row (Figs.21D, 21E, 21F; Figs.22D, 22E, 22F; Figs. 23D, 23E, 23F) have been obtained by the approach of Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023. - the results of the third row (Figs.21G, 21H, 21I; Figs.22G, 22H, 22I; Figs.23G, 23H, 23I) have been obtained by the proposed solution of the present embodiment. The following can be observed in Figs. 21A-21I: The point clouds of the vehicles (second column, i.e. Figs. 21B, 21E, 21H) are less accurate in the case of the second row and about the same accuracy in the case of the first raw and our proposal (last, third row). However, regarding the accuracy of ground points, our proposal (recommendation) is the best (last column, i.e. Figs. 21C, 21F, 21I, it is observable by the close fitting of the lines that our proposal, Fig.21I shows the best results). In Figs. 22A-I, the static scene (second column, i.e. Figs. 22B, 22E, 22H) is estimated with the most considerable precision with the approach applied in the second row (using ego-motion estimation), but our ones are almost as good, first row is the worst of the three. The dynamic points of far vehicles (third column, i.e. Figs.22C, 22F, 22I) are the most precise in the case of our proposal (i.e. in Fig.22I). Figs.23A-23I illustrate a further example where the limitation of the proposal can be seen (ground model fitting limits the precision of generated ground points, see Fig. 23H). However, dynamic object parts are still the most accurate with our proposal (see the comparison of Figs.23B, 23E and 23H). E. Discussion This Section contains further discussion beyond our test results provided in the previous Section. First, the computation efficiency discussion; next, a separate evaluation for dynamic objects is presented, and finally, an ablation study is presented. E.1. Computation efficiency Real-time running is essential as our goal is to temporally up-sample point clouds. (As a matter of fact, the run time should be less than the time elapsed between measurements with a given frequency we aim to up-sample.) The running time of the pipeline components is listed in Table IX with the following configuration: Intel Core i7-4790K @ 4.00GHz processor, 32 GB RAM, Nvidia GTX 1080 GPU, Windows 1064 bit. For the tests, KITTI-sized images were used (about 1200x350 resolution). Component Average running time (ms) Optical flow estimation 14 Motion-in-depth estimation 29 Ground estimation 8 Frame generation (scene-flow + transformation) ≈0 Total 51 Table IX – Average running time of the different pipeline components As the Total running time of the system in the above example is 51 ms, we can generate virtual point clouds with almost 20 Hz, meaning we can up-sample 5 and 10 Hz lidar measurements with the current limited research configuration. In Table X below, our method of the present embodiment is compared to alternatives that can run in real-time and have similar inputs and the same goal as us, namely up-sampling point cloud measurements. The approaches have been used in Table X for comparison with the proposed method of the present embodiment: ^ Future pseudo-lidar: Huang, X.; Lin, C.; Liu, H.; Nie, L.; Zhao, Y. Future Pseudo- LiDAR Frame Prediction for Autonomous Driving. Multimedia Syst. 2022, 28, 1611 1620. ^ Rigid body based up-sampling: Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023. Methods GPU Average running time [ms] Future pseudo-lidar Nvidia RTX 2080Ti 52 Rigid body based up- Nvidia GTX 1080 62 sampling Proposed method Nvidia GTX 1080 51 Table X – Running time of different methods with the type of the GPU on which it is measured As shown in Table X, we outperform the alternatives in terms of running time (even without considering our GPU disadvantage). E.2. Dynamic objects We made an evaluation separately only for dynamic objects as (Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023.). There are two main reasons why this is important to investigate: ^ The ego-motion can be determined with different localization sensors (like in He, L.; Jin, Z.; Gao, Z. De-Skewing LiDAR Scan for Refinement of Local Mapping. Sensors 2020, 20, 1846.) or by methods (like Campos, C.; Elvira, R.; Gomez, J.J.; Montiel, J.M.M.; Tardos, J.D. ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM. IEEE Transactions on Robotics 2021, 37, 1874 1890.), and the motion of static scene elements can be calculated based on that. However, for dynamic objects, we cannot do that. ^ Dynamic objects generally pose a greater threat as they change their position, and they also can change their state variables (angular and linear velocity, acceleration). For the evaluation, we used the annotations of the Semantic KITTI dataset (Behley, J.; Garbade, M.; Milioto, A.; Quenzel, J.; Behnke, S.; Stachniss, C.; Gall, J. SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proceedings of the Proc. of the IEEE/CVF International Conf. on Computer Vision (ICCV), 2019.) to select vehicles as dynamic object candidates on the sequences of 08-10. The evaluation results can be observed in Table XI. The approaches have been used in Table XI for comparison with the proposed pipeline of the present embodiment: ^ Point based and Range map based: Weng, X.; Wang, J.; Levine, S.; Kitani, K.; Rhinehart, N. Inverting the Pose Forecasting Pipeline with SPF2: Sequential Pointcloud Forecasting for Sequential Pose Forecasting. In Proceedings of the Proceedings of (CoRL) Conference on Robot Learning, 2020. ^ Rigid body based up-sampling: Rozsa, Z.; Sziranyi, T. Virtually increasing the measurement frequency of lidar sensor utilizing a single RGB camera. arXiv preprint arXiv:2302.051922023. Methods CD [m2] EMD [m2] Point based 2.37 211.47 Range map based 0.92 128.81 Rigid body based up- 0.63 31.03 sampling Proposed pipeline 0.26 17.50 Table XI – Quantitative evaluation and comparison of the proposed pipeline on KITTI odometry dataset (just vehicles). As Table XI shows, we provide the best performance in the case of generating point clouds of dynamic objects. E.3. Ablation Study Here, we provide an ablation study to prove the significance of our contributions of the present embodiment. It can be seen in table form in Table XII. It is observable from the table that the all the examined proposed pipeline components have enhanced error metrics or running time of the algorithm, or both. Components CD [m2] EMD [m2] Runtime [ms] Proposed pipeline 0.486 28.51 51 without ground 0.508 34.50 43 estimation without FastFlowNet 0.508 34.40 95 Baseline (appr.) 0.547 34.39 270 Table XII – Ablation study on KITTI odometry dataset. In Table XII, Baseline (appr.) – for which temporal upsampling was never considered as an option – stands for the demonstration code provided by the authors of Yang, G.; Ramanan, D. Upgrading Optical Flow to 3D Scene Flow Through Optical Expansion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020 (briefly, Baseline document). Abbreviation appr. is for approximation as the authors in their demo code used the approximation of Zt(pt-1)= ^Zt-1(pt-1) ignoring ũ in position (see Eq. (12)). Table XII above shows the followings, from bottom to top: 1. In the Baseline row, we evaluate the Baseline method one by one for our problem (i.e. we add LIDAR depth data, so we get 3D point clouds as the result, not normalized 3D coordinates, which would give the result without a specific depth). 2. The next line, 'without FastFlowNet', contains the addition of "minor" modifications to Baseline (even without the FastFlowNet modification). This basically consists of two things: A) code optimization, a grid generation step was taken out of the inside of the mesh, saving runtime. B) Baseline uses ^=Zt(pt-1)/Zt-1(pt-1) for the definition of ^ (see above, this approximation, which effectively says ũ=0, is a deviation from our exact evaluation), and this is reflected in the accuracy. 3. The next line up from the bottom, 'without ground estimation', refers to when FastFlowNet is already being used (ground estimation is not yet). Using a network that can operate in real time (and the emphasis here is on real time, not just FastFlowNet) is an important step in determining optical flow, because it allows the method to actually be used for this problem, to make sense of it all. (Since these LIDARs have a measurement frequency of 10 Hz, if we want to generate only one new one between two measurements, we need to generate frames at about 50 ms, otherwise we will get the "t+1" time instant measurement before we generate what the "t" would look like). 4. Topmost row: This is the complete system, which therefore already includes the ground estimation step. Moreover, also the below details are to be considered. The results for the proposed pipeline of the present embodiment are also given without a building block of it. The 'without FastFlowNet' means our development where the exact equation is used (reducing errors), and also, the runtime is decreased as we optimized the code handed out by taking a grid generation step out of the network. This run time is closer reported by Baseline document. 'Without ground estimation' indicates our extension using FastFlowNet instead of the Volumetric Correspondence Networks (VCN) (Yang, G.; Ramanan, D. Volumetric Correspondence Networks for Optical Flow. In Proceedings of the NeurIPS, 2019.), it is a very significant reduction in the runtime, making the up- sampling possible. Our proposed pipeline of the present invention is the one described in this description, using both the optical flow estimation and the ground model estimation. Adding the ground estimation step somewhat increased the runtime; however, there was a significant improvement in the error metrics. F. Conclusions In Section IV. we introduced a methodology to solve the problem of temporally up- sampling lidar point clouds utilizing a mono camera. We extended the state-of-the- art to be able to do this in real time. The proposed pipeline of the present embodiment has several advantages compared to alternatives; it is not designed for specific point cloud characteristics, and it does not require ego-motion estimation. Our evaluation of the KITTI dataset shows that our framework has good accuracy in terms of similarity to real measurements. In the future, we intend to apply a more sophisticated sensor model to calculate more accurately the point locations of the generated virtual point clouds. The invention is, of course, not limited to the preferred embodiments described in detail above, but further variants, modifications and developments are possible within the scope of protection determined by the claims.

Claims

CLAIMS 1. A method for depth completion, comprising the steps of - performing downscaling (S24a, S24b) on a plurality of downscaling processing levels (20) of a downscaling branch (5) of a processing unit (50) based on neural network with starting, at the highest downscaling processing level (20) of the downscaling branch (5), on a plurality of initial first data blocks (25a-25e) corresponding to respective input depth information images (10a-10e) each generated from a respective point set, wherein the respective point sets have been recorded in consecutive integration windows of a scene by means of a lidar device with non- repetitive scanning pattern, wherein o on the highest downscaling processing level (20) or on more consecutive downscaling processing levels (20) from the highest downscaling processing level (20) a convolution operation applied on the respective downscaling processing level (20) is combined with performing convolution Long-Short Term Memory operation, o in a respective downscaling processing level (20), a downscaling convolution operation is performed by a downscaling factor on all of the one or more initial first data block (25a-25e) of the respective downscaling processing level (20) to obtain respective one or more downscaled first data block (21a-21e, 23a-23e) of the respective downscaling processing level (20), wherein, in case of having one or more downscaling processing level (20) below the lowest downscaling processing level on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, for the downscaling processing levels (20) below the lowest of the one or more downscaling processing level (20) on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, downscaling is performed further on at least a part of the plurality of downscaled first data blocks (21a-21e, 23a-23e) including the last of the plurality of downscaled first data blocks (21a-21e, 23a-23e), and - after performing downscaling in the plurality of downscaling processing levels (20), forwarding (S42a) o in case of having a single downscaled first data block on the lowest downscaling processing level, the single downscaled first data block, or o in case of having a plurality of downscaled first data blocks (21a- 21e) on the lowest downscaling processing level, ^ the last of the plurality of downscaled first data blocks (21a- 21e), or ^ a processed last data block outputted by an auxiliary convolution Long-Short Term Memory operation performed on the plurality of downscaled data blocks of the lowest downscaling processing level on the lowest downscaling processing level as an initial second data block (31) to the lowest upscaling processing level of a plurality of upscaling processing levels (30) of an upscaling branch (15) of the processing unit (50), then - performing upscaling (S26a, S26b) in the plurality of upscaling processing levels (30) of the upscaling branch (15) of the processing unit (50) with starting on the initial second data block (31) of the lowest upscaling processing level of the upscaling branch (15), wherein in a respective upscaling processing level (30) an upscaling operation is performed by an upscaling factor on the initial second data block (31) of the respective upscaling processing level (30) to obtain upscaled second data block (33, 35) of the respective upscaling processing level (30), and an output depth information image (40) is obtained based on the upscaled second data block (35) of the highest upscaling processing level (30).
2. The method according to claim 1, characterised by adding (S32), after performing upscaling (S26a, S26b) in the plurality of upscaling processing levels (30), the last of the input depth information images (10e) to an upscaled depth information image corresponding to the upscaled initial data block of the highest upscaling processing level (30) to obtain the output depth information image (40).
3. The method according to claim 1 or claim 2, characterized in that a convolution operation applied on a respective downscaling processing level (20) is combined with performing convolution Long-Short Term Memory operation for all downscaling processing levels (20).
4. The method according to any of claims 1-3, characterized by utilizing - initial depth data (100) corresponding to the output depth information image (40) of the scene corresponding to an initial point of time, and - a first optical image (102) of the scene corresponding to the initial point of time and a second image (104) of the scene corresponding to a subsequent point of time being subsequent to the initial point of time, in the method, which method further comprising - an input generating step comprising the substeps of o generating (S115) optical flow data (108) between the first optical image (102) and the second image (104) by means of an optical flow generation module based on a first auxiliary neural network, and generating (S120) motion-in-depth related data (112) based on the first optical image (102), the second image (104) and the optical flow data (108) by means of a motion-in-depth module based on a second auxiliary neural network, and o determining (S110) ground depth data (106a) and non-ground object depth data (106b) by segmenting the initial depth data (100), - after the input generating step has been performed, generating (S125) 3D flow data (114) corresponding to changes compared to the initial depth data (100) based on the ground depth data (106a), the non-ground object depth data (106b) and the motion-in-depth related data (112), then - generating (S130) subsequent depth data (116) based on o the ground depth data (106a), and o the non-ground object depth data (106b) modified based on 3D flow data (114).
5. The method according to claim 4, characterized in that the depth data is point set data or depth information image.
6. A system for depth completion, the system comprising a processing unit (50) based on neural network, the processing unit (50) having a downscaling branch (5) and an upscaling branch (15), wherein - the downscaling branch (5) is adapted for performing downscaling on a plurality of downscaling processing levels (20) of a downscaling branch (5) of a processing unit (50) based on neural network with starting, at the highest downscaling processing level (20) of the downscaling branch (5), on a plurality of initial first data blocks (25a-25e) corresponding to respective input depth information images (10a-10e) each generated from a respective point set, wherein the respective point sets have been recorded in consecutive integration windows of a scene by means of a lidar device with non-repetitive scanning pattern, wherein o on the highest downscaling processing level (20) or on more consecutive downscaling processing levels (20) from the highest downscaling processing level (20) a convolution operation applied on the respective downscaling processing level (20) is combined with performing convolution Long-Short Term Memory operation, o in a respective downscaling processing level (20), a downscaling convolution operation is performed by a downscaling factor on all of the one or more initial first data block (25a-25e) of the respective downscaling processing level (20) to obtain respective one or more downscaled first data block (21a-21e, 23a-23e) of the respective downscaling processing level (20), wherein, in case of having one or more downscaling processing level (20) below the lowest downscaling processing level on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, for the downscaling processing levels (20) below the lowest of the one or more downscaling processing level (20) on which a convolution operation is combined with performing convolution Long-Short Term Memory operation, downscaling is performed further on at least a part of the plurality of downscaled first data blocks (21a-21e, 23a-23e) including the last of the plurality of downscaled first data blocks (21a-21e, 23a-23e), and - the processing unit (50) is adapted for forwarding, after performing downscaling in the plurality of downscaling processing levels (20), o in case of having a single downscaled first data block on the lowest downscaling processing level, the single downscaled first data block, or o in case of having a plurality of downscaled first data blocks (21a- 21e) on the lowest downscaling processing level, ^ the last of the plurality of downscaled first data blocks (21a- 21e) or ^ a processed last data block outputted by an auxiliary convolution Long-Short Term Memory operation performed on the plurality of downscaled data blocks of the lowest downscaling processing level on the lowest downscaling processing level as an initial second data block (31) to the lowest upscaling processing level of a plurality of upscaling processing levels (30) of the upscaling branch (15), - the upscaling branch (15) is adapted for, after forwarding as an initial second data block (31) to the lowest upscaling processing level, performing upscaling in the plurality of upscaling processing levels (30) with starting on the initial data block (31) of the lowest upscaling processing level of the upscaling branch (15), wherein in a respective upscaling processing level (30) an upscaling operation is performed by an upscaling factor on the initial data block (31) of the respective upscaling processing level (30) to obtain upscaled initial data block (33, 35) of the respective upscaling processing level (30), and the upscaling branch (15) is further adapted for obtaining an output depth information image (40) based on the initial upscaled data block (35) of the highest upscaling processing level (30).
7. A training method characterized by comprising training of the processing unit (50) based on neural network utilized in the method according to any of claims 1-5 and comprised in the system according to claim 6.
8. The training method according to claim 7, characterized by utilizing for training a loss function comprising - a pixel-level term, - a Structure Similarity Index Measure term, and - an edge loss term.
9. The training method according to claim 7 or claim 8, characterized in that the training is performed by the help of a training dataset generated by the help of CARLA simulator, wherein the training dataset comprises pairs of - input depth information images generated with characteristics of the lidar device with non-repetitive scanning pattern, and - output depth information image extracted to be dense using a mask corresponding to the field of view of the lidar device with non-repetitive scanning pattern.
EP24718887.3A 2023-03-01 2024-02-29 Method and system for depth completion Pending EP4673912A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
HUP2300075 2023-03-01
PCT/HU2024/050016 WO2024180357A1 (en) 2023-03-01 2024-02-29 Method and system for depth completion

Publications (1)

Publication Number Publication Date
EP4673912A1 true EP4673912A1 (en) 2026-01-07

Family

ID=92590115

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24718887.3A Pending EP4673912A1 (en) 2023-03-01 2024-02-29 Method and system for depth completion

Country Status (2)

Country Link
EP (1) EP4673912A1 (en)
WO (1) WO2024180357A1 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119649052B (en) * 2024-12-12 2025-08-22 中科领航智能科技(苏州)有限公司 A highly scalable feature extraction network detection method for efficient detection of road objects

Family Cites Families (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9872011B2 (en) 2015-11-24 2018-01-16 Nokia Technologies Oy High-speed depth sensing with a hybrid camera setup
CN109964237B (en) 2016-09-15 2020-07-17 谷歌有限责任公司 Image depth prediction neural network
WO2020104423A1 (en) 2018-11-20 2020-05-28 Volkswagen Aktiengesellschaft Method and apparatus for data fusion of lidar data and image data
US10929995B2 (en) 2019-06-24 2021-02-23 Great Wall Motor Company Limited Method and apparatus for predicting depth completion error-map for high-confidence dense point-cloud
US12026899B2 (en) 2019-07-22 2024-07-02 Toyota Motor Europe Depth maps prediction system and training method for such a system
CN111612728B (en) 2020-05-25 2023-07-25 北京交通大学 A 3D point cloud densification method and device based on binocular RGB images
CN111950467B (en) 2020-08-14 2021-06-25 清华大学 Fusion network lane line detection method and terminal device based on attention mechanism
CN114004754B (en) 2021-09-13 2022-07-26 北京航空航天大学 A system and method for scene depth completion based on deep learning
US12145617B2 (en) 2021-10-28 2024-11-19 Nvidia Corporation 3D surface reconstruction with point cloud densification using artificial intelligence for autonomous systems and applications
KR102694068B1 (en) 2022-01-20 2024-08-09 숭실대학교산학협력단 Depth completion method and apparatus using a spatial-temporal
US12444029B2 (en) 2022-01-29 2025-10-14 Samsung Electronics Co., Ltd. Method and device for depth image completion
CN114983328A (en) 2022-05-09 2022-09-02 安阳市人民医院 Gastroenterology esophagus detection device
CN115100090B (en) 2022-06-09 2025-04-25 北京邮电大学 A monocular image depth estimation system based on spatiotemporal attention

Also Published As

Publication number Publication date
WO2024180357A1 (en) 2024-09-06

Similar Documents

Publication Publication Date Title
Mostafavi et al. Learning to reconstruct hdr images from events, with applications to depth and flow prediction
Shah et al. Traditional and modern strategies for optical flow: an investigation
CN111325794B (en) A Visual Simultaneous Localization and Mapping Method Based on Depth Convolutional Autoencoder
Guo et al. Context-enhanced stereo transformer
Lv et al. Learning rigidity in dynamic scenes with a moving camera for 3d motion field estimation
CN110651310B (en) Deep learning methods and related methods and software for estimating object density and/or traffic
Yin et al. Scale recovery for monocular visual odometry using depth estimated with deep convolutional neural fields
Kumar et al. Monocular fisheye camera depth estimation using sparse lidar supervision
Sun et al. Mm3dgs slam: Multi-modal 3d gaussian splatting for slam using vision, depth, and inertial measurements
Bešić et al. Dynamic object removal and spatio-temporal RGB-D inpainting via geometry-aware adversarial learning
CN118552596B (en) Depth estimation method based on multi-view self-supervision learning
Tong et al. Large-scale aerial scene perception based on self-supervised multi-view stereo via cycled generative adversarial network
Haji-Esmaeili et al. Large-scale monocular depth estimation in the wild
Li et al. Dgnr: Density-guided neural point rendering of large driving scenes
Zováthi et al. ST-DepthNet: A spatio-temporal deep network for depth completion using a single non-repetitive circular scanning Lidar
Liu et al. NeRF dynamic scene reconstruction based on motion, semantic information and inpainting
EP4673912A1 (en) Method and system for depth completion
Vaquero et al. Hallucinating dense optical flow from sparse lidar for autonomous vehicles
Da Silveira et al. Indoor depth estimation from single spherical images
CN120689585A (en) An anti-interference target detection method and system based on multimodal fusion of autonomous driving scenarios
Reza et al. FarSight: Long-range depth estimation from outdoor images
Chen et al. Neural radiance fields for dynamic view synthesis using local temporal priors
Saleh et al. Real-time 3D Perception of Scene with Monocular Camera
Dong et al. Framework of degraded image restoration and simultaneous localization and mapping for multiple bad weather conditions
Xu et al. Video-object segmentation and 3D-trajectory estimation for monocular video sequences

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250926

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR