WO2024258697A1 - Method and system for calibrating a camera - Google Patents
Method and system for calibrating a camera Download PDFInfo
- Publication number
- WO2024258697A1 WO2024258697A1 PCT/US2024/032500 US2024032500W WO2024258697A1 WO 2024258697 A1 WO2024258697 A1 WO 2024258697A1 US 2024032500 W US2024032500 W US 2024032500W WO 2024258697 A1 WO2024258697 A1 WO 2024258697A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- machine
- intrinsic
- incidence
- determining
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/80—Analysis of captured images to determine intrinsic or extrinsic camera parameters, i.e. camera calibration
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30248—Vehicle exterior or interior
- G06T2207/30252—Vehicle exterior; Vicinity of vehicle
Definitions
- the present disclosure relates to camera calibration and, more specifically, to a system and method for in-the-wild monocular camera calibration.
- Camera calibration is typically the first step in numerous vision and robotics applications that involve three-dimensional (3D) sensing.
- Classic methods enable accurate calibration by imaging a specific 3D structure such as a checkerboard.
- 3D sensing from in- the-wild images, such as monocular depth estimation, 3D object detection, and 3D reconstruction.
- 3D sensing techniques over in-the-wild images are developed, camera calibration for such in-the-wild images continues to pose significant challenges.
- the two-view camera pose is determined up to a projective ambiguity if images are uncalibrated.
- Alternative solutions exist by employing deep networks to regress the pose.
- regression hinders the usage of geometric constraints, which proves crucial in calibrated two-view pose estimation.
- Other works use more than two uncalibrated images for pose estimation.
- the present method makes no assumption except for undistorted images. This enables us to calibrate four degrees of freedom (DoF) intrinsic and train with any calibrated images. Further as compared to monocular camera calibration with object the present method can uses any image.
- DoF degrees of freedom
- calibration also detects geometric manipulation.
- the present method does not need to encrypt images, complementing photometric manipulation detection.
- the present system uses an incidence field is also a learnable monocular 3D prior.
- a common characteristic of monocular 3D priors is their exclusive relationship with 3D structures, making them invariant to 2D manipulations, e.g., cropping, and resizing.
- computing the groundtruth incidence field is straightforward, relying solely on intrinsic and image coordinates.
- a method for calibrating a camera includes communicating an image to a neural network comprising a plurality of pixels, determining an incidence field corresponding to the plurality of pixels using a neural network, removing outliers from the incidence field, determining intrinsic parameters based on the incidence, providing a two-dimensional image to a machine, providing the intrinsic parameters to the machine, determining a three-dimensional representation of the two- dimensional image based on the intrinsic parameters and controlling the machine based on the three-dimensional representation.
- a method includes mapping an input image having a plurality of pixels to an incidence field, removing outliers from the incidence field to form a modified incidence field, based on the modified incidence field, determining an intrinsic matrix comprising intrinsic parameters, providing a two- dimensional image to a machine, providing the intrinsic matrix to the machine, determining a three-dimensional representation of the two-dimensional image based on the intrinsic parameters and controlling the machine based on the three dimensional representation.
- a system includes a processor and a non-transitory computer readable medium that includes machine-readable instructions that are executable by the processor.
- the machine-readable instructions when executed perform the steps of determining an incidence field corresponding to a plurality of pixels of an image using a neural network, removing outliers from the incidence field, determining intrinsic parameters based on the incidence, providing a two-dimensional image to a machine and providing the intrinsic parameters to the machine, a machine comprising an image processor determining a three-dimensional representation of the two dimensional image based on the intrinsic parameters and a controller controlling the machine based on the three dimensional representation.
- Fig. 1 A is a block diagrammatic view of a calibration system relative to a machine.
- Fig. 1 B is a diagram illustrating an incidence ray relative to a three dimensional view.
- Fig. 1 C is a diagrammatic representation of an estimated depth map relative to a surface normal and a noisy incidence field.
- Fig. 2A is a flowchart of a method for controlling a machine based upon intrinsic parameters for calibration.
- Fig. 2B is a plot showing estimated verses Random Sampling Consensus (RANSAC) vectors for an image.
- Fig. 2C is a flowchart of the RANSAC sampling process without assumptions.
- Fig. 2D is a flowchart of a method for determining the focal length for a simple camera.
- Fig. 3B is an image using estimated restoration according to the present disclosure.
- Fig. 3C is a restoration using the prior known ground truth.
- Fig. 3D is an original image for Fig 3A.
- Fig. 4A is an input image having vectors for a reduced side image situation.
- Fig. 4B is an image using estimated restoration according to the present disclosure.
- Fig. 5A is a plot of pose versus focal length error using mega depth data.
- Fig. 5B is a plot of pose versus focal length error for a synthetic noisy intrinsic or the estimated intrinsic according to the present disclosure.
- the control system includes a calibration system 12 that is used to generate intrinsic parameters (intrinsics of the camera) that are used for calibrating a machine controller 14.
- the calibration system receives a two-dimensional initial image 16 that is communicated to a neural network 18.
- the neural network 18 generates an incidence field V that is provided to an outlier removal system 20.
- the outlier removal system 20 and the neural network 18 are described in greater detail below.
- the outlier removal system 20 removes outliers in the data.
- intrinsic parameters are formed in the intrinsic parameter system 24.
- the intrinsic parameter system 24 may generate a matrix with the intrinsic parameters therein.
- the machine controller 14 includes an image processor 40 that receives two dimensional images 42 from an image device 44.
- the image device 44 may be a camera such as a digital camera.
- the image processor 40 and the machine controller 14 may include a microprocessor or processor 46 that is coupled to a memory 48.
- the memory 48 may also be a non-transitory computer readable medium that includes machine-readable instructions that are executable by the processor 46 to control a machine and the image processor 40.
- more than one microprocessor and memory may be provided within the machine controller.
- the machine controller 14 may be used to control an actuator 50.
- the machine controller and actuator 50 are provided in a machine 52. Examples of a machine include an autonomous vehicle.
- the actuator 50 may, for example, be brakes, a throttle control and/or a steering control.
- the machine 52 may also include other safety systems within a vehicle including accident avoidance, pedestrian avoidance, active cruise control and the like.
- the machine 52 may also be a robot and the actuator 50 may be a robotic arm.
- the machine 52 may also be an augmented reality or virtual reality device in which the actuator 50 is a display.
- the machine 52 may include the calibration system 12.
- the calibration system 12 may also be coupled to the machine 52 by a direct wire or network 56.
- an incidence ray 60 has a point of origin 62 that is located in the 3D image.
- a camera plane 64 has a point 66 that corresponds to the point of origin 62 in the 3D image.
- an estimated depthmap 66A is converted to surface normals 72A and 72B in images 68B and 70B and using a groundtruth (GT) and a noisy intrinsic, respectively.
- GT groundtruth
- a noisy intrinsic distorts the point cloud 66B, consequently leading to inaccurate surface normal 72A and 72B.
- the surface normal 72A presents differently than the surface normal 72B in 70B.
- a solver that utilizes the consistency between the two to recover the intrinsic is set forth. However, this solution exhibits numerical instability.
- an alternative 3D monocular prior, the incidence field is used. As shown in Fig.
- the neural network 18 is used to learn in-the-wild incidence field and develop the outlier removal system 20, which in this example is a random sample consensus (RANSAC algorithm) to recover the intrinsic from the estimated incidence field generated by the neural network 18.
- RANSAC algorithm random sample consensus
- the present disclosure is motivated by the consistency between the monocular depthmap and surface normal map.
- an incorrect intrinsic distorts the back-projected 3D point cloud from the depthmap, resulting in distorted surface normals.
- intrinsic is optimal when the estimated monocular depthmap aligns consistently with the surface normal.
- a previously known solution to recover a four DoF intrinsic by leveraging the consistency between the surface normal and depthmap has been proposed.
- the system is numerically ill-conditioned as its computation depends on the accurate gradient of depthmap. This requires depthmap estimation with low bias and variance.
- the incidence field depicts the incidence ray between the observed 3D point and the projected 2D pixel on the imaging plane, as shown in Fig. 1 B.
- the combination of the incidence field and the monocular depthmap describes a 3D point cloud.
- the incidence field is a direct pixel-wise parameterization of the camera intrinsic. This implies that a minimal solver based on the incidence field only needs to have low bias.
- a deep neural network is used to perform the incidence field estimation.
- a non-learning RANSAC algorithm is developed to recover the intrinsic parameters from the estimated incidence field.
- the incidence field is also a monocular 3D prior. Similar to depthmap and surface normal, the incidence field is invariant to the image cropping or resizing. This encourages its generalization over in-the-wild images. Theoretically, the incidence field is pixel-wisely determined by depthmap and normal map. Empirically, multiple public datasets were combined into a comprehensive dataset with diverse indoor, outdoor, and object-centric images captured by different imaging devices. The variety of intrinsic is boosted by resizing and cropping the images in a similar manner. Zero-shot testing samples were included to benchmark real-world monocular camera calibration performance.
- Downstream applications benefit from monocular camera calibration.
- two additional applications detecting and restoring image resizing and cropping.
- an image is cropped or resized, it disrupts the assumption of a central focal point and identical focal length.
- the edited image is restored by adjusting its intrinsic to a regularized form.
- Another application involves two-view uncalibrated camera pose estimation. With established image correspondence, a fundamental matrix is determined. However, there does not exist an injective mapping between the fundamental matrix and camera pose. This raises a counter-intuitive fact: inferring the pose from two uncalibrated images is infeasible. But the present method enables uncalibrated two-view pose estimation by applying monocular camera calibration.
- the present approach tackles monocular camera calibration from a novel perspective by relying on monocular 3D priors.
- the present method makes no assumption for the to-be-calibrated undistorted image.
- the present method provides robust monocular intrinsic estimation for in-the-wild images, accompanied by extensive benchmarking and comparisons against other baselines. The benefits on diverse and novel downstream applications are set forth.
- step 210 the input image I is obtained.
- the input image I is mapped to the incidence field V in step 212.
- outliers are removed from the incidence field.
- the outliers may be removed using the random sampling consensus as described above.
- FIG. 2B a single iteration of RANSAC is used to obtain the intrinsic.
- the intrinsic is computed with two incidence vectors randomly sampled at pixel locations 216A, 216B. From Eq. (2), an intrinsic determines the incidence vector at a given location. The optimal intrinsic maximizes the consistency with the network prediction between the estimated incidence field and the data with the outliers removed by RANSAC.
- Figs. 2C and 2D details the RANSAC algorithm and an enumerated system for use with a simple camera are illustrated respectively. Different strategies are applied depending on if a simple camera is assumed. If not assumed (f x , bx) and (f y , by ) are independently computed. If assumed, there is only 1 DoF of intrinsic. The focal length within a predefined range was enumerated to determine the optimal value.
- the RANSAC sampling process has a first stage 218A that determines a focal point Fx in the x direction of the focal point bx.
- the first stage also determines the focal length F y and the focal point b y .
- Stages 218A and 218B use a minimal solver.
- a scoring function is used to generate a score for the X values at 222A and for the Y values at 222B.
- a maximum is obtained from the score valves 220A and 220B at 222A and 222B, respectively.
- matrix valves 224 are obtained for the focal length coordinates in the X and Y directions (X-axis focal length and Y-axis focal length) fx, f y and the center point coordinates in the X and Y directions bx, b y .
- the matrix values are formed for the intrinsic in step 226.
- Eq. 18 determines the focal length at 230.
- the score using Eq. (19) is generated at 232.
- a maximizing function 234 is used to obtain the maximum value so that the focal length is obtained in 236.
- step 2308 the intrinsic parameters are communicated to the image processor or the camera to provide values for calibration.
- step 240 a two-dimensional image is communicated to the image processor.
- step 242 the three-dimensional representation of the two-dimensional image calibrated based on the intrinsic parameters is determined.
- a machine is controlled based on the 3D representation in step 244.
- the vector v is an incidence ray, originating from the 3D point P, directed towards the 2D pixel p, and passing through the camera’s origin.
- Eq. (9) relies on the consistency between the surface normal and depthmap gradient, which may require a low-variance depthmap estimate. From Fig. 3, even groundtruth depthmap leads to spurious normal due to its inherent high variance. Thus, minimal solver in Eq. (9) can lead to a poor solution.
- the incidence field V is chosen and is a monocular 3D prior pixel wise determined by monocular depth D and surface normal N. Expanding Eq. (4) using the incidence vector v in Eq. (2):
- Neural Window Conditional Random Field (NeWCRFs) a neural network used in monocular depth estimation, is adopted for incidence field estimation.
- the last output head is changed to output a three-dimensional normalized incidence field V with the same resolution as the input image I.
- a cosine similarity loss is defined as: [0074]
- the last dimension of the incidence field is normalized to one before feeding to the RANSAC Algorithm (step 214). That is to say,
- the processor 28 executes the calibration system, a GPU-based RANSAC algorithm is used to recover the intrinsic K from the incidence field V. Unlike a CPU-based RANSAC, fixed Kr iterations of RANSAC are performed without termination. In RANSAC, the minimal solver is used to generate Kc candidates and select the optimal one that maximizes a scoring function (see Fig. 2).
- scoring function is defined in x-axis and y-axis, respectively:
- Nk and kx/ k y are the number of sampled pixels and the threshold for axis-x/y directions.
- Figs. 3A-3D and Figs 4A-4D manipulated images such as cropped images and resize detection may be restored.
- Eq. (10) defines a crop and resize function:
- the original intrinsic K is unknown. It is presumed that the genuine image possess an identical focal length and central focal point. Any resizing and cropping are detected when matrix K' breaks this assumption. Note, the rule cannot detect aspect ratio preserving resize or centered crop.
- the original image is restored by defining an inverse operation AK restore K' to an intrinsic fits the assumption.
- Intrinsic estimation can enable multiple applications, e.g., depthmap to point cloud, uncalibrated two-view pose estimation, etc.
- FIGs. 3A-3D one example of image editing including cropping is set forth.
- the monocular calibration is applicable to detect and restore images that are both cropped and resized.
- Fig. 3A an image having the estimated and RANSAC vectors imposed thereon is set forth.
- Fig. 3B the estimated restoration according to the present disclosure is provided.
- Fig. 3C a ground truth restoration according to the prior art is performed.
- Fig. 3D the original image with a box illustrating the position is set forth. As can be seen, the restoration locates the image in the proper original location relative to the frame.
- Figs. 6A and 6B plot the uncalibrated two-view camera pose estimation performance with respect to an increasingly noisy intrinsic.
- a naive way to synthesize noise into intrinsic is set forth as :
- c is a random sign whose value is either 1 or -1.
- Variable XX are noisy and groundtruth focal length in axis X and Y respectively.
- Figs. 6A and 6B show the uncalibrated pose performance is highly sensitive to intrinsic noise, suggesting itself a challenging problem.
- uncalibrated two-view pose estimation can also be utilized to evaluate the quality of intrinsic parameter estimation.
- first, second, third, etc. may be used herein to describe various elements, components, regions, layers and/or sections, these elements, components, regions, layers and/or sections should not be limited by these terms. These terms may be only used to distinguish one element, component, region, layer or section from another region, layer or section. Terms such as “first,” “second,” and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of the example embodiments.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Image Processing (AREA)
Abstract
A method and system for calibrating a camera comprises communicating an image to a neural network comprising a plurality of pixels, determining an incidence field corresponding to the plurality of pixels using a neural network, removing outliers from the incidence field, determining intrinsic parameters based on the incidence, providing a two-dimensional image to a machine, providing the intrinsic parameters to the machine, determining a three-dimensional representation of the two-dimensional image based on the intrinsic parameters and controlling the machine based on the three-dimensional representation.
Description
METHOD AND SYSTEM FOR CALIBRATING A CAMERA
CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63/521 ,367, filed on June 16, 2023. The entire disclosure of the above application is incorporated herein by reference.
FIELD
[0002] The present disclosure relates to camera calibration and, more specifically, to a system and method for in-the-wild monocular camera calibration.
BACKGROUND
[0003] Camera calibration is typically the first step in numerous vision and robotics applications that involve three-dimensional (3D) sensing. Classic methods enable accurate calibration by imaging a specific 3D structure such as a checkerboard. With the rapid growth of monocular 3D vision, there is an increasing focus on 3D sensing from in- the-wild images, such as monocular depth estimation, 3D object detection, and 3D reconstruction. While 3D sensing techniques over in-the-wild images are developed, camera calibration for such in-the-wild images continues to pose significant challenges.
[0004] Classic methods for monocular calibration use strong geometry prior, such as using a checkerboard. However, such 3D structures are not always available for in- the-wild images. As a solution, alternative methods relax the assumptions. For example, calibration using common objects such as human faces and 3D bounding boxes. Another significant line of research is based on the Manhattan World assumption, which posits that all planes within a scene are either parallel or perpendicular to each other. This assumption is further relaxed to estimate the lines that are either parallel or perpendicular to the direction of gravity. The intrinsic parameters are recovered by determining the intersected vanishing points of detected lines, assuming a central focal point and an identical focal length.
[0005] While the assumptions are relaxed, they may still not hold true for in-the- wild images. This creates a contradiction: although robust models to estimate in-the-wild monocular depthmap are developed, generating the corresponding 3D point cloud remains infeasible due to the missing intrinsic. A similar challenge arises in monocular 3D object detection, as limitations in projecting the detected 3D bounding boxes onto the
2D image are faced. In augmented reality and virtual reality applications, the absence of intrinsic precludes placing multiple reconstructed 3D objects within a canonical 3D space. The absence of a reliable, assumption-free monocular intrinsic calibrator has become a bottleneck in deploying these 3D sensing applications.
[0006] Monocular camera calibration with geometry has been used in one line of work and assumes the Manhattan World assumption, where all planes in 3D space are either parallel or perpendicular. Under the assumption, line segments in the image converge at the vanishing points, from which the intrinsic is recovered. Line Segment Detector (LSD) and Edge Drawing Line (EDLine) develop robust line estimators. Others jointly estimate the horizon line and the vanishing points. Recent learning-based methods relax the assumption by using the gravity-aligned panorama images during training. Such data provides the groundtruth vanishing point and horizon lines which are originally computed with the Manhattan World assumption. Still, the assumption constrains in modeling intrinsic as 1 degree of freedom (DoF) camera. Recently, another study relaxed the assumption to 3 DoF via regressing the focal point.
[0007] In another system monocular camera calibration with object Is used. The method based on a checkerboard pattern is widely regarded as the standard for camera calibration. Several works generalize this method to other geometric patterns such as 1 D objects, line segments, and spheres. Recent works extend camera calibration to real- world objects such as human faces. Optimizers, including Backpropagation Perspective- n-Points (BPnP) and Perspective-n-Points (PnP) have also been developed. However, the usage of specific objects restricts their applications.
[0008] With respect to image cropping and resizing, detecting photometric image manipulation has been extensively researched. But few studies detect geometric manipulation, e.g., resizing and cropping. On resizing, one study regresses the image aspect ratio with a deep model.
[0009] With the fundamental matrix estimated, the two-view camera pose is determined up to a projective ambiguity if images are uncalibrated. Alternative solutions exist by employing deep networks to regress the pose. However, regression hinders the usage of geometric constraints, which proves crucial in calibrated two-view pose estimation. Other works use more than two uncalibrated images for pose estimation.
[0010] Other systems use learnable monocular 3D priors. Monocular depth and surface normals are two established 3D priors. Numerous studies have shown a learnable mapping between monocular images and corresponding 3D priors. Recently,
one study introduces the perspective field as a monocular 3D prior for inferring gravity direction.
SUMMARY
[0011] This section provides a general summary of the disclosure and is not a comprehensive disclosure of its full scope or all of its features.
[0012] In comparison to monocular calibration with geometry, the present method makes no assumption except for undistorted images. This enables us to calibrate four degrees of freedom (DoF) intrinsic and train with any calibrated images. Further as compared to monocular camera calibration with object the present method can uses any image.
[0013] With respect to image cropping and resizing, calibration also detects geometric manipulation. The present method does not need to encrypt images, complementing photometric manipulation detection.
[0014] The present system and method also enables an uncalibrated two-view solution.
[0015] The present system uses an incidence field is also a learnable monocular 3D prior. A common characteristic of monocular 3D priors is their exclusive relationship with 3D structures, making them invariant to 2D manipulations, e.g., cropping, and resizing. In contrast, a camera intrinsic changes on spatial manipulation. Unlike other monocular priors, computing the groundtruth incidence field is straightforward, relying solely on intrinsic and image coordinates.
[0016] In one aspect of the disclosure, a method for calibrating a camera includes communicating an image to a neural network comprising a plurality of pixels, determining an incidence field corresponding to the plurality of pixels using a neural network, removing outliers from the incidence field, determining intrinsic parameters based on the incidence, providing a two-dimensional image to a machine, providing the intrinsic parameters to the machine, determining a three-dimensional representation of the two- dimensional image based on the intrinsic parameters and controlling the machine based on the three-dimensional representation.
[0017] In another aspect of the disclosure, a method includes mapping an input image having a plurality of pixels to an incidence field, removing outliers from the incidence field to form a modified incidence field, based on the modified incidence field, determining an intrinsic matrix comprising intrinsic parameters, providing a two-
dimensional image to a machine, providing the intrinsic matrix to the machine, determining a three-dimensional representation of the two-dimensional image based on the intrinsic parameters and controlling the machine based on the three dimensional representation.
[0018] In yet another aspect of the disclosure, a system includes a processor and a non-transitory computer readable medium that includes machine-readable instructions that are executable by the processor. The machine-readable instructions when executed perform the steps of determining an incidence field corresponding to a plurality of pixels of an image using a neural network, removing outliers from the incidence field, determining intrinsic parameters based on the incidence, providing a two-dimensional image to a machine and providing the intrinsic parameters to the machine, a machine comprising an image processor determining a three-dimensional representation of the two dimensional image based on the intrinsic parameters and a controller controlling the machine based on the three dimensional representation.
[0019] Further areas of applicability will become apparent from the description provided herein. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure.
DRAWINGS
[0020] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations and are not intended to limit the scope of the present disclosure.
[0021] Fig. 1 A is a block diagrammatic view of a calibration system relative to a machine.
[0022] Fig. 1 B is a diagram illustrating an incidence ray relative to a three dimensional view.
[0023] Fig. 1 C is a diagrammatic representation of an estimated depth map relative to a surface normal and a noisy incidence field.
[0024] Fig. 2A is a flowchart of a method for controlling a machine based upon intrinsic parameters for calibration.
[0025] Fig. 2B is a plot showing estimated verses Random Sampling Consensus (RANSAC) vectors for an image.
[0026] Fig. 2C is a flowchart of the RANSAC sampling process without assumptions.
[0027] Fig. 2D is a flowchart of a method for determining the focal length for a simple camera.
[0028] Fig. 3A is an input image having vectors for a cropping situation.
[0029] Fig. 3B is an image using estimated restoration according to the present disclosure.
[0030] Fig. 3C is a restoration using the prior known ground truth.
[0031] Fig. 3D is an original image for Fig 3A.
[0032] Fig. 4A is an input image having vectors for a reduced side image situation.
[0033] Fig. 4B is an image using estimated restoration according to the present disclosure.
[0034] Fig. 4C is a restoration using the prior known ground truth.
[0035] Fig. 4D is an original image for Fig 4A.
[0036] Fig. 5A is a plot of pose versus focal length error using mega depth data.
[0037] Fig. 5B is a plot of pose versus focal length error for a synthetic noisy intrinsic or the estimated intrinsic according to the present disclosure.
[0038] Corresponding reference numerals indicate corresponding parts throughout the several views of the drawings.
DETAILED DESCRIPTION
[0039] Example embodiments will now be described more fully with reference to the accompanying drawings.
[0040] Referring now to Figure 1 A, a control system 10 that uses a monocular 3D system is set forth. The control system includes a calibration system 12 that is used to generate intrinsic parameters (intrinsics of the camera) that are used for calibrating a machine controller 14. The calibration system receives a two-dimensional initial image 16 that is communicated to a neural network 18. The neural network 18 generates an incidence field V that is provided to an outlier removal system 20. The outlier removal system 20 and the neural network 18 are described in greater detail below. The outlier removal system 20 removes outliers in the data. Ultimately, intrinsic parameters are formed in the intrinsic parameter system 24. The intrinsic parameter system 24 may generate a matrix with the intrinsic parameters therein.
[0041] As will be described in greater detail, an image restoration system 26 may be used to restore modified images such as cropped images and resized images. The image restoration system 26, in general, may be used to position a restored image such as a cropped image or a resized image in the relative position of the received image. Details of components 18-26 are described in greater detail below. The calibration system may include one or more microprocessors or processors 28 that are in communication with a memory 30. The memory 30 is a non-transitory computer-readable medium that includes machine readable instructions that are executable by the processor 28. The machine readable instructions include instructions for forming the intrinsic parameters. Ultimately, the intrinsic parameters are provided to a machine controller 14. The intrinsic parameters are part of a calibration. The machine controller 14 includes an image processor 40 that receives two dimensional images 42 from an image device 44. The image device 44 may be a camera such as a digital camera. The image processor 40 and the machine controller 14 may include a microprocessor or processor 46 that is coupled to a memory 48. The memory 48 may also be a non-transitory computer readable medium that includes machine-readable instructions that are executable by the processor 46 to control a machine and the image processor 40. Of course, more than one microprocessor and memory may be provided within the machine controller. The machine controller 14 may be used to control an actuator 50. In one example, the machine controller and actuator 50 are provided in a machine 52. Examples of a machine include an autonomous vehicle. The actuator 50 may, for example, be brakes, a throttle control and/or a steering control. The machine 52 may also include other safety systems within a vehicle including accident avoidance, pedestrian avoidance, active cruise control and the like. The machine 52 may also be a robot and the actuator 50 may be a robotic arm. The machine 52 may also be an augmented reality or virtual reality device in which the actuator 50 is a display. The machine 52 may include the calibration system 12. The calibration system 12 may also be coupled to the machine 52 by a direct wire or network 56.
[0042] Referring now to Fig. 1 B, an incidence ray 60 has a point of origin 62 that is located in the 3D image. A camera plane 64 has a point 66 that corresponds to the point of origin 62 in the 3D image.
[0043] In the Fig. 1 C, an estimated depthmap 66A is converted to surface normals 72A and 72B in images 68B and 70B and using a groundtruth (GT) and a noisy intrinsic, respectively. A noisy intrinsic distorts the point cloud 66B, consequently leading to
inaccurate surface normal 72A and 72B. In 68B, the surface normal 72A presents differently than the surface normal 72B in 70B. Motivated by the observation, a solver that utilizes the consistency between the two to recover the intrinsic is set forth. However, this solution exhibits numerical instability. In the present system an alternative 3D monocular prior, the incidence field, is used. As shown in Fig. 1 B, the collection of the pixel-wise incidence rays 60, which originate from a 3D point of origin 62, targets at a 2D pixel 66, and crosses the camera origin. Similar to depthmap and normal, a noisy intrinsic leads to a noisy incidence field.
[0044] The neural network 18 is used to learn in-the-wild incidence field and develop the outlier removal system 20, which in this example is a random sample consensus (RANSAC algorithm) to recover the intrinsic from the estimated incidence field generated by the neural network 18. The present disclosure is motivated by the consistency between the monocular depthmap and surface normal map. In the images 66A, 68B and 70A an incorrect intrinsic distorts the back-projected 3D point cloud from the depthmap, resulting in distorted surface normals. Based on this, intrinsic is optimal when the estimated monocular depthmap aligns consistently with the surface normal. A previously known solution to recover a four DoF intrinsic by leveraging the consistency between the surface normal and depthmap has been proposed. However, the system is numerically ill-conditioned as its computation depends on the accurate gradient of depthmap. This requires depthmap estimation with low bias and variance.
[0045] To resolve it, an alternative approach is set forth herein by introducing a novel 3D monocular prior in complementation to the depthmap and surface normal map. This is referred to as the incidence field, which depicts the incidence ray between the observed 3D point and the projected 2D pixel on the imaging plane, as shown in Fig. 1 B. The combination of the incidence field and the monocular depthmap describes a 3D point cloud. Compared to the previous solution, the incidence field is a direct pixel-wise parameterization of the camera intrinsic. This implies that a minimal solver based on the incidence field only needs to have low bias. A deep neural network is used to perform the incidence field estimation. A non-learning RANSAC algorithm is developed to recover the intrinsic parameters from the estimated incidence field.
[0046] The incidence field is also a monocular 3D prior. Similar to depthmap and surface normal, the incidence field is invariant to the image cropping or resizing. This encourages its generalization over in-the-wild images. Theoretically, the incidence field is pixel-wisely determined by depthmap and normal map. Empirically, multiple public
datasets were combined into a comprehensive dataset with diverse indoor, outdoor, and object-centric images captured by different imaging devices. The variety of intrinsic is boosted by resizing and cropping the images in a similar manner. Zero-shot testing samples were included to benchmark real-world monocular camera calibration performance.
[0047] Downstream applications benefit from monocular camera calibration. In addition to the aforementioned 3D sensing tasks, two additional applications, detecting and restoring image resizing and cropping, are provided. When an image is cropped or resized, it disrupts the assumption of a central focal point and identical focal length. Using the estimated intrinsic parameters, the edited image is restored by adjusting its intrinsic to a regularized form. Another application involves two-view uncalibrated camera pose estimation. With established image correspondence, a fundamental matrix is determined. However, there does not exist an injective mapping between the fundamental matrix and camera pose. This raises a counter-intuitive fact: inferring the pose from two uncalibrated images is infeasible. But the present method enables uncalibrated two-view pose estimation by applying monocular camera calibration.
[0048] The present approach tackles monocular camera calibration from a novel perspective by relying on monocular 3D priors. The present method makes no assumption for the to-be-calibrated undistorted image. The present method provides robust monocular intrinsic estimation for in-the-wild images, accompanied by extensive benchmarking and comparisons against other baselines. The benefits on diverse and novel downstream applications are set forth.
[0049] Various non-learning methods rely on strong assumptions. Learning-based methods relax the assumptions to using gravity-aligned panorama images during training. The present method makes no assumptions except for undistorted images. This enables training with any calibrated images and calibrates 4 DoF intrinsic.
[0050] Referring now to Figs 2A, and 2B, in step 210, the input image I is obtained. The input image I is mapped to the incidence field V in step 212. In step 214, outliers are removed from the incidence field. The outliers may be removed using the random sampling consensus as described above.
[0051] In FIG. 2B, a single iteration of RANSAC is used to obtain the intrinsic. The intrinsic is computed with two incidence vectors randomly sampled at pixel locations 216A, 216B. From Eq. (2), an intrinsic determines the incidence vector at a given location.
The optimal intrinsic maximizes the consistency with the network prediction between the estimated incidence field and the data with the outliers removed by RANSAC.
[0052] Referring now also to Figs. 2C and 2D details the RANSAC algorithm and an enumerated system for use with a simple camera are illustrated respectively. Different strategies are applied depending on if a simple camera is assumed. If not assumed (fx, bx) and (fy, by ) are independently computed. If assumed, there is only 1 DoF of intrinsic. The focal length within a predefined range was enumerated to determine the optimal value.
[0053] As will be described in greater detail below, the RANSAC sampling process has a first stage 218A that determines a focal point Fx in the x direction of the focal point bx. The first stage also determines the focal length Fy and the focal point by. Stages 218A and 218B, as will be described below, use a minimal solver. At stage 222A and 222B, Eq. (17), a scoring function is used to generate a score for the X values at 222A and for the Y values at 222B. A maximum is obtained from the score valves 220A and 220B at 222A and 222B, respectively. Ultimately, matrix valves 224 are obtained for the focal length coordinates in the X and Y directions (X-axis focal length and Y-axis focal length) fx, fy and the center point coordinates in the X and Y directions bx, by. In Fig. 2A, the matrix values are formed for the intrinsic in step 226.
[0054] Referring now to Fig. 2D, when a simple camera is presumed, Eq. 18 determines the focal length at 230. The score using Eq. (19) is generated at 232. A maximizing function 234 is used to obtain the maximum value so that the focal length is obtained in 236.
[0055] In step 238, the intrinsic parameters are communicated to the image processor or the camera to provide values for calibration. In step 240, a two-dimensional image is communicated to the image processor. In step 242, the three-dimensional representation of the two-dimensional image calibrated based on the intrinsic parameters is determined. Ultimately, a machine is controlled based on the 3D representation in step 244.
Monocular Intrinsic Calibration
[0056] Details of the method for calibrating a camera using intrinsics are provided in greater detail below. The present method uses generalizable monocular 3D priors without assuming the 3D scene geometry. Hence, start with monocular depthmap D and surface normal map N. It is assumed there exists a learnable mapping from the input
image I to depthmap D and normal map N: D, N = D0(l), where D0 is a learned network.
[0057] The notation Ksimpie suggests a simple camera model with the identical focal length and central focal point assumption. Given a 2D homogeneous pixel location p = [x y 1 ]T and its depth value d = D(p), the corresponding 3D point P = [X Y Z]T is defined as:
[0058] where the vector v is an incidence ray, originating from the 3D point P, directed towards the 2D pixel p, and passing through the camera’s origin. The incidence field V is the collection of incidence rays v associated with all pixels at location p, where v = V(p).
Monocular Intrinsic Calibration with Monocular Depthmap and Surface Normal
[0059] In this section, the intrinsic matrix K using the estimated surface normal map N and depthmap D is explained. Given the estimated depth d = D(p) and normal n = N(p) at 2D pixel location p, a local 3D plane is described as: nT * ’ v 4- = 0. (3)
[0061] Note the bias c of the 3D local plane is independent of the camera projection process. Without loss of generality, the case of our method for x-direction. Expanding Eq. (4), is then:
[0062] where Vx(d) represents the gradient of the depthmap D in the x-axis and can be computed, for example, using a Sobel filter. Next, the unknowns in Eq. (5) are reparametrize to get:
[0064] where r = fy/fx. By stacking Eq. (7) with N > 4 randomly sampled pixels, a linear system is acquired:
A VX Xix l — B4x r. (8) where the intrinsic parameter to be solved is stored in a vectorX4xl = [fy by rbx r] .
[0065] The known constants are stored in matrix ANX4 and B4xi. If N = 4 is chosen in Eq. (8), a minimal solver is obtained where the solution X is computed by performing Gauss-Jordan Elimination. Conversely, when N > 4, the linear system is overdetermined, and X is obtained using a least squares solver. The above suggests the intrinsic is recoverable from the monocular 3D prior.
Monocular Incidence Field as Monocular 3D Prior
[0066] Eq. (9) relies on the consistency between the surface normal and depthmap gradient, which may require a low-variance depthmap estimate. From Fig. 3, even groundtruth depthmap leads to spurious normal due to its inherent high variance. Thus, minimal solver in Eq. (9) can lead to a poor solution.
[0067] As a solution to this, directly learning the incidence field V as a monocular 3D prior is set forth. In Eq. (2) and Figs. 1 A-1 C, the combination of the incidence field V and the monocular depthmap D creates a 3D point cloud. In Eq. (3), the incidence field V can measure the observation angle between a 3D plane and the camera. Similar to depthmap D and surface normal map N, the incidence field V is invariant to the image cropping and resizing. Consider an image cropping and resizing described as:
[0068] where K' is the intrinsic after transformation. The surface normal map N and depthmap D after transformation is defined as:,
[0070] Eq. (12) shows that the incidence field V is a parameterization of the intrinsic matrix that is invariant to image resizing and image cropping. Other invariant parameterizations of the intrinsic matrix, such as the camera field of view (FoV), rely on the central focal point assumption and only cover a 2 DoF intrinsic matrix. An illustration is shown in Fig. 3.
[0071] Next, the incidence field V is chosen and is a monocular 3D prior pixel wise determined by monocular depth D and surface normal N. Expanding Eq. (4) using the incidence vector v in Eq. (2):
[0072] Combined with the axis-y constraint, the 2 DoF incidence vector v = [vx, vy, 1]T is uniquely solved. Eq. (13) gives identical solutions when depthmap D is adjusted by a scalar. This implies that the incidence field V is pixel-wise determined by the up-to- scale depth map and surface normal map, indicating a learnable mapping from the monocular image I to the incidence field V exists.
[0073] Given the strong connection between the monocular depthmap D and camera incidence field V, Neural Window Conditional Random Field (NeWCRFs), a neural network used in monocular depth estimation, is adopted for incidence field estimation. The last output head is changed to output a three-dimensional normalized incidence field V with the same resolution as the input image I. A cosine similarity loss is defined as:
[0074] The last dimension of the incidence field is normalized to one before feeding to the RANSAC Algorithm (step 214). That is to say,
Monocular Intrinsic Calibration with Incidence Field
[0075] The processor 28 executes the calibration system, a GPU-based RANSAC algorithm is used to recover the intrinsic K from the incidence field V. Unlike a CPU-based RANSAC, fixed Kr iterations of RANSAC are performed without termination. In RANSAC, the minimal solver is used to generate Kc candidates and select the optimal one that maximizes a scoring function (see Fig. 2).
[0076] Details of RANSAC without assumption from Fig. 2C are set forth. From Eq. (2), the incidence vector v relates to the intrinsic K as:
[0077] From Eq. (15), a minimal solver for intrinsic is straightforward. In the incidence field, randomly sample two incidence vectors
[0079] where Nk and kx/ ky are the number of sampled pixels and the threshold for axis-x/y directions.
[0080] RANSAC w/ Assumption. If a simple camera model is assumed, i.e., intrinsic only has an unknown focal length, it only needs to estimate 1 -DoF intrinsic. The focal length candidates are:
[0081] The scoring function under the scenario is defined as the summation over x-axis and y-axis:
[0082] Referring now to Figs. 3A-3D and Figs 4A-4D manipulated images such as cropped images and resize detection may be restored.
[0084] When a modified image I' is presented, the present system calibrates its intrinsic K'.
[0085] In a first case, the original intrinsic K is known. That is, K is obtained from the image-associated file. Image manipulation is computed as AK = (K')(K“1). A manipulation is detected if AK deviates from an identity matrix. The original image restores as l(x) = l'(AK x). Interestingly, the four corners of image I' are mapped to a bounding box in original image I under manipulation AK. Thus, the restoration is quantified by measuring the bounding box.
[0086] In the second case, the original intrinsic K is unknown. It is presumed that the genuine image possess an identical focal length and central focal point. Any resizing and cropping are detected when matrix K' breaks this assumption. Note, the rule cannot detect aspect ratio preserving resize or centered crop. The original image is restored by defining an inverse operation AK restore K' to an intrinsic fits the assumption.
[0087] Intrinsic estimation can enable multiple applications, e.g., depthmap to point cloud, uncalibrated two-view pose estimation, etc.
[0088] Referring now to Figs. 3A-3D, one example of image editing including cropping is set forth. In this example, the monocular calibration is applicable to detect and restore images that are both cropped and resized. In Fig. 3A, an image having the estimated and RANSAC vectors imposed thereon is set forth. In Fig. 3B, the estimated restoration according to the present disclosure is provided. In Fig. 3C, a ground truth restoration according to the prior art is performed. In Fig. 3D, the original image with a box illustrating the position is set forth. As can be seen, the restoration locates the image in the proper original location relative to the frame.
[0089] Referring now to Figs. 4A-4D, an original image that has been resized is set forth along with the vectors according to the estimated system or a grounds truth
system. The ground truth and estimated vectors are set forth adjacent to each other. In Fig. 4B, the estimated restoration according to the present disclosure is set forth. In Fig. 4C, the ground truth restoration is set forth. As can be seen by the image boxes in Fig. 4D, the estimated restoration provides the location in the proper location relative to the original image.
[0090] Referring now to Fig. 6A and 6B, the performance of uncalibrated two-view pose estimation with respect to the monocular camera calibration error is set forth. The steep curve indicates that uncalibrated two-view pose estimation is a challenging problem. The dot in each view marks the posed performance using the estimated intrinsic as set forth herein.
[0091] Figs. 6A and 6B plot the uncalibrated two-view camera pose estimation performance with respect to an increasingly noisy intrinsic. A naive way to synthesize noise into intrinsic is set forth as :
[0092] where c is a random sign whose value is either 1 or -1. Variable
XX are noisy and groundtruth focal length in axis X and Y respectively. Figs. 6A and 6B show the uncalibrated pose performance is highly sensitive to intrinsic noise, suggesting itself a challenging problem. Just as the geometric matching community employs two-view pose estimation to assess the quality of correspondences, uncalibrated two-view pose estimation can also be utilized to evaluate the quality of intrinsic parameter estimation.
[0093] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a,” "an," and "the" may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises," "comprising," “including,” and “having,” are inclusive and therefore specify the presence of stated features, integers, steps, operations, elements, and/or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and/or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order discussed or illustrated, unless specifically identified as an order of
performance. It is also to be understood that additional or alternative steps may be employed.
[0094] When an element or layer is referred to as being "on," “engaged to,” "connected to," or "coupled to" another element or layer, it may be directly on, engaged, connected or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being "directly on," “directly engaged to,” "directly connected to," or "directly coupled to" another element or layer, there may be no intervening elements or layers present. Other words used to describe the relationship between elements should be interpreted in a like fashion (e.g., “between” versus “directly between,” “adjacent” versus “directly adjacent,” etc.). As used herein, the term "and/or" includes any and all combinations of one or more of the associated listed items.
[0095] Although the terms first, second, third, etc. may be used herein to describe various elements, components, regions, layers and/or sections, these elements, components, regions, layers and/or sections should not be limited by these terms. These terms may be only used to distinguish one element, component, region, layer or section from another region, layer or section. Terms such as “first,” “second,” and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of the example embodiments.
[0096] The foregoing description of the embodiments has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure. Individual elements or features of a particular embodiment are generally not limited to that particular embodiment, but, where applicable, are interchangeable and can be used in a selected embodiment, even if not specifically shown or described. The same may also be varied in many ways. Such variations are not to be regarded as a departure from the disclosure, and all such modifications are intended to be included within the scope of the disclosure.
Claims
1 . A method comprising: communicating an image to a neural network comprising a plurality of pixels; determining an incidence field corresponding to the plurality of pixels using a neural network; removing outliers from the incidence field; determining intrinsic parameters based on the incidence; providing a two-dimensional image to a machine; providing the intrinsic parameters to the machine determining a three-dimensional representation of the two-dimensional image based on the intrinsic parameters; and controlling the machine based on the three-dimensional representation.
2. The method of claim 1 wherein the image comprises a cropped image.
3. The method of claim 1 wherein the image comprises a resized image.
4. The method of claim 1 wherein the neural network comprises a neural window conditional random field.
5. The method of claim 1 wherein the incidence field comprises a three-dimensional normalized incidence field.
6. The method of claim 1 wherein the incidence field comprises a group comprising three vectors for each pixel of the plurality of pixels.
7. The method of claim 1 wherein the incidence field corresponds to rays from a 3D point and a projected 2D pixel passing through a camera origin on an imaging plane.
8. The method of claim 1 wherein removing outliers comprises removing outliers using a random sample consensus.
9. The method of claim 1 wherein the intrinsic parameters comprise a focal length.
10. The method of claim 1 wherein the intrinsic parameters comprise a focal length and a center point.
11 . The method of claim 1 wherein the intrinsic parameters comprise an X direction a Y direction focal length, an X location of a center point and a Y location of the center point.
12. The method of claim 1 further comprising determining an image manipulation and restoring the image based on the image manipulation.
13. The method of claim 12 wherein determining the image manipulation comprises restoring a resizing of the image.
14. The method of claim 12 wherein determining the image manipulation comprises restoring a cropping of the image.
15. The method of claim 1 wherein controlling a machine comprises controlling an autonomous vehicle.
16. The method of claim 1 wherein controlling a machine comprises controlling a robot.
17. A method comprising: mapping an input image having a plurality of pixels to an incidence field; removing outliers from the incidence field to form a modified incidence field; based on the modified incidence field, determining an intrinsic matrix comprising intrinsic parameters; providing a two-dimensional image to a machine; providing the intrinsic matrix to the machine; determining a three-dimensional representation of the two-dimensional image based on the intrinsic parameters; and controlling the machine based on the three dimensional representation.
18. The method of claim 17 wherein mapping is performed using a neural network.
19. The method of claim 17 wherein determining the intrinsic matrix is performed by using a random sample consensus to determine the intrinsic.
20. A system comprising: a processor; and a non-transitory computer readable medium that includes machine-readable instructions that are executable by the processor, the machine-readable instructions when executed perform the steps of, determining an incidence field corresponding to a plurality of pixels of an image using a neural network; removing outliers from the incidence field; determining intrinsic parameters based on the incidence; providing a two-dimensional image to a machine; providing the intrinsic parameters to the machine; a machine comprising an image processor determining a three-dimensional representation of the two dimensional image based on the intrinsic parameters and a controller controlling the machine based on the three dimensional representation.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363521367P | 2023-06-16 | 2023-06-16 | |
| US63/521,367 | 2023-06-16 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024258697A1 true WO2024258697A1 (en) | 2024-12-19 |
Family
ID=91585626
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/032500 Ceased WO2024258697A1 (en) | 2023-06-16 | 2024-06-05 | Method and system for calibrating a camera |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024258697A1 (en) |
-
2024
- 2024-06-05 WO PCT/US2024/032500 patent/WO2024258697A1/en not_active Ceased
Non-Patent Citations (7)
| Title |
|---|
| ANONYMOUS: "How will the Camera Intrinsics change if an image is cropped/resized?", STACK OVERFLOW, 10 December 2022 (2022-12-10), XP093201439, Retrieved from the Internet <URL:https://stackoverflow.com/questions/74749690/how-will-the-camera-intrinsics-change-if-an-image-is-cropped-resized> [retrieved on 20240904] * |
| FANFANI MARCO ET AL: "A vision-based fully automated approach to robust image cropping detection", SIGNAL PROCESSING. IMAGE COMMUNICATION, ELSEVIER SCIENCE PUBLISHERS, AMSTERDAM, NL, vol. 80, 23 September 2019 (2019-09-23), XP085915419, ISSN: 0923-5965, [retrieved on 20190923], DOI: 10.1016/J.IMAGE.2019.115629 * |
| GROSSBERG M D ET AL: "A general imaging model and a method for finding its parameters", PROCEEDINGS OF THE EIGHT IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION. (ICCV). VANCOUVER, BRITISH COLUMBIA, CANADA, JULY 7 - 14, 2001; [INTERNATIONAL CONFERENCE ON COMPUTER VISION], LOS ALAMITOS, CA : IEEE COMP. SOC, US, vol. 2, 7 July 2001 (2001-07-07), pages 108 - 115, XP010554075, ISBN: 978-0-7695-1143-6 * |
| LINYI JIN ET AL: "Perspective Fields for Single Image Camera Calibration", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 16 March 2023 (2023-03-16), XP091456009 * |
| SHENGJIE ZHU ET AL: "Tame a Wild Camera: In-the-Wild Monocular Camera Calibration", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 19 June 2023 (2023-06-19), XP091542805 * |
| VASILJEVIC IGOR ET AL: "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion", 2020 INTERNATIONAL CONFERENCE ON 3D VISION (3DV), IEEE, 25 November 2020 (2020-11-25), pages 1 - 11, XP033880198, DOI: 10.1109/3DV50981.2020.00010 * |
| WEIHAO YUAN ET AL: "NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 6 June 2022 (2022-06-06), XP091239707 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11455746B2 (en) | System and methods for extrinsic calibration of cameras and diffractive optical elements | |
| US9183631B2 (en) | Method for registering points and planes of 3D data in multiple coordinate systems | |
| Urban et al. | Mlpnp-a real-time maximum likelihood solution to the perspective-n-point problem | |
| US5706419A (en) | Image capturing and processing apparatus and image capturing and processing method | |
| US10469828B2 (en) | Three-dimensional dense structure from motion with stereo vision | |
| US5818959A (en) | Method of producing a three-dimensional image from two-dimensional images | |
| Becker et al. | Semiautomatic 3D-model extraction from uncalibrated 2D-camera views | |
| US10762654B2 (en) | Method and system for three-dimensional model reconstruction | |
| KR20180087947A (en) | Modeling method and modeling apparatus using 3d point cloud | |
| JP6172432B2 (en) | Subject identification device, subject identification method, and subject identification program | |
| US10346949B1 (en) | Image registration | |
| EP2840550A1 (en) | Camera pose estimation | |
| Baráth et al. | Optimal multi-view surface normal estimation using affine correspondences | |
| Deschenes et al. | An unified approach for a simultaneous and cooperative estimation of defocus blur and spatial shifts | |
| EP1580684B1 (en) | Face recognition from video images | |
| WO2024258697A1 (en) | Method and system for calibrating a camera | |
| Georgiev et al. | A fast and accurate re-calibration technique for misaligned stereo cameras | |
| Flusser et al. | Invariants to convolution in arbitrary dimensions | |
| JP7813503B1 (en) | Method for aligning 2D images and plan views based on 3D geometry and image alignment device using the same | |
| Brooks et al. | Data registration | |
| CN120027782B (en) | Multi-sensor coupling method, device and equipment | |
| Tekla et al. | A Minimal Solution for Image-Based Sphere Estimation | |
| JP2001338279A (en) | 3D shape measuring device | |
| Kovalenko et al. | Correction of the interpolation effect in modeling the process of estimating image spatial deformations | |
| Garg et al. | High-Precision Positioning of Known Objects Using a Static Monocular Camera |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24734774 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |







