WO2015085779A1 - Method and system for calibrating surveillance cameras - Google Patents
Method and system for calibrating surveillance cameras Download PDFInfo
- Publication number
- WO2015085779A1 WO2015085779A1 PCT/CN2014/083329 CN2014083329W WO2015085779A1 WO 2015085779 A1 WO2015085779 A1 WO 2015085779A1 CN 2014083329 W CN2014083329 W CN 2014083329W WO 2015085779 A1 WO2015085779 A1 WO 2015085779A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- images
- point set
- dimensional point
- surveillance cameras
- monitoring scene
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/80—Analysis of captured images to determine intrinsic or extrinsic camera parameters, i.e. camera calibration
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/52—Surveillance or monitoring of activities, e.g. for recognising suspicious objects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/18—Closed-circuit television [CCTV] systems, i.e. systems in which the video signal is not broadcast
- H04N7/181—Closed-circuit television [CCTV] systems, i.e. systems in which the video signal is not broadcast for receiving images from a plurality of remote sources
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10016—Video; Image sequence
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10028—Range image; Depth image; 3D point clouds
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30232—Surveillance
Definitions
- Embodiments of the present disclosure generally relate to an image processing technology, and more particularly, to a method and a system for calibrating a plurality of surveillance cameras.
- Video monitoring is an important part of a safeguard system, in which one key issue is how to calibrate surveillance cameras in a monitoring scene to obtain their intrinsic parameters (such as a coordinate of a principle point, a focal length and a distortion factor) and external parameters (such as a rotation matrix and a translation vector of a coordinate system of the surveillance camera with respect to a universal coordinate system) thereof, so as to obtain a relative erecting position, a facing direction and a viewing filed of each surveillance camera in the safeguard system.
- intrinsic parameters such as a coordinate of a principle point, a focal length and a distortion factor
- external parameters such as a rotation matrix and a translation vector of a coordinate system of the surveillance camera with respect to a universal coordinate system
- a conventional calibrating method a calibrating method based on active vision
- a self -calibrating method a calibrating method based on active vision
- an precisely machined calibration object such as a calibration block, a calibration plate or a calibration bar
- the intrinsic and external parameters of the surveillance camera are obtained by establishing a relationship between a given coordinate of a point on the calibration object and a coordinate of the point on an image.
- This method wastes time and energy because above operation is required to be performed on each surveillance camera, and only the intrinsic parameters can be calibrated by this method.
- the outdoor monitoring scene is usually large and a size of the calibration object is relatively small and occupies a little proportion of the image, a calibration error of this method is large.
- the surveillance camera is controlled actively to move in a special mode (such as a pure translation motion or a rotation around an optical center), and the intrinsic parameters of the surveillance camera are calculated according to the specificity of the motion.
- a special mode such as a pure translation motion or a rotation around an optical center
- this method has a poor applicability.
- it is also required to calibrate each surveillance camera respectively, thus wasting time and energy, and moreover, the external parameters cannot be calibrated.
- the self-calibrating method no calibration object is needed and the surveillance camera is not required to move in the special mode.
- the intrinsic and external parameters of the surveillance camera is directly calibrated according to the relationship between pixels on a plurality of images sampled by the surveillance camera (maybe a plurality of surveillance cameras) and constraints of the intrinsic and external parameters.
- Some self-calibrating methods can calibrate a plurality of surveillance cameras simultaneously by means of a multiple view geometry method. However, the viewing fields of the plurality of surveillance cameras are required to have a large area overlapped with each other, otherwise the plurality of surveillance cameras will not be calibrated due to an image matching failure. However, in practice, the overlapped area between viewing fields of the plurality of surveillance cameras is small. Thus, the currently known self-calibrating methods have difficulties in calibrating the plurality of surveillance cameras simultaneously.
- Embodiments of the present disclosure seek to solve at least one of the problems existing in the related art to at least some extent.
- one object of the present disclosure is to provide a method for calibrating a plurality of surveillance cameras which can calibrate intrinsic and external parameters of the surveillance cameras having no overlapped viewing field.
- Another object of the present disclosure is to provide a system for calibrating a plurality of surveillance cameras which can calibrate intrinsic and external parameters of the surveillance cameras having no overlapped viewing field.
- a method for calibrating a plurality of surveillance cameras includes: sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device and sampling a second image in the monitoring scene by each of the plurality of surveillance cameras respectively; performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images; reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result between the plurality of first images; and calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
- sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device includes: calibrating intrinsic parameters of the sampling device; sampling the plurality of first images in the monitoring scene of the plurality of surveillance cameras by the sampling device with calibrated intrinsic parameters, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
- performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images includes: when the sampling device has a GPS sensor, obtaining first images adjacent to each other according to GPS location information obtained during sampling the plurality of first images, and matching the adjacent first images with each other; and when the sampling device has no GPS sensor, matching the plurality of first images with each other.
- reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result includes: selecting two first images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation degree between the initial images is smaller than a predetermined degree; calculating camera matrices corresponding to the initial images according to the matching point pairs on the initial images; reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching relationship between the initial images ; expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images.
- expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images includes: determining whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene; if no, finding out a first image at most matched with the reconstructed three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculating a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set; expanding the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices; optimizing the calculated camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
- reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality of first images includes: obtaining GPS location information and gesture information corresponding to each of the plurality of first images and calculating the camera matrix corresponding to each of the plurality of first images according to the GPS location information, the gesture information and the intrinsic parameters of the sampling device; for each two adjacent first images, calculating three-dimensional coordinates of matching points on the two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm; optimizing the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
- calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras includes: for each surveillance camera to be calibrated, matching the sampled second image with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculating an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set includes points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set; determining whether a number of feature points in the intersection set is larger than a predetermined number; if no, sampling at least one third image in the monitoring scene of the plurality of surveillance cameras by the sampling device and expanding the three-dimensional point set according to the at least one third image, so as to obtain an updated intersection set in which the number of feature points is larger than the predetermined number; calculating the parameters of the surveillance camera to be calibrated
- a system for calibrating a plurality of surveillance cameras includes: at least one sampling device, configured to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras; and a computing device, configured to perform a feature matching on the plurality of first images to obtain a matching result between the plurality of first images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result between the plurality of first images, and to calculate parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
- intrinsic parameters of the at least one sampling device are calibrated, and the at least one sampling device with calibrated intrinsic parameters actively samples the plurality of first images in the monitoring scene of plurality of the surveillance cameras, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
- the computing device is configured to: select two images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation between the initial images is smaller than a predetermined degree, to calculate camera matrices corresponding to the initial images according to the matching point pairs on the initial images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching relationship between the initial images and the camera matrices corresponding to the initial images, and to expand the three-dimensional point set of the monitoring scene according to the three-dimensional point set and other first images except for the initial images.
- the computing device is further configured to determine whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene, if no, find out a first image at most matched with the reconstructed three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculate a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set, expand the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices, optimize the camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
- the computing device is configured to obtain GPS location information and gesture information corresponding to each of the plurality of first images and calculate the camera matrix corresponding to each of the plurality of first images according to the GPS location information, the gesture information and the intrinsic parameters of the at least one sampling device, calculate three-dimensional coordinates of matching points on each two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm, optimize the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
- the computing device is further configured to match the second image sampled by each of the plurality of surveillance cameras with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculate an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set includes points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set, determine whether a number of feature points in the intersection set is larger than a predetermined number; if no, receive at least one third image sampled by the at least one sampling device and expand the three-dimensional point set according to the at least one third image so as to obtain an updated intersection set in which the number of feature points is larger than the predetermined number, calculate the parameters of each of the plurality of surveillance cameras according to the intersection set and a three-dimensional point set corresponding to the intersection set.
- the three-dimensional point set of the monitoring scene can be accurately reconstructed according to the relatively complete image information of the monitoring scene obtained by the sampling device, and the camera matrices of the surveillance cameras can be calculated according to the relationship between the three-dimensional point set and the feature point set on the second images sampled by the surveillance cameras, thus achieving the calibration of the surveillance cameras.
- the calibrating process is dynamic, i.e., when the number of the reconstructed three-dimensional points is not large enough to calibrate the surveillance cameras, more images can be sampled to expand the reconstructed three-dimensional point set, thus ensuring that all the surveillance cameras can be calibrated.
- the calibrating method in the present disclosure does not require a calibration object, a special motion of the surveillance camera or overlapped viewing fields of the surveillance cameras. Rather, the calibrating method according to embodiments of the present disclosure can calibrate a plurality of surveillance cameras simultaneously as long as the sampling device samples the images of the monitoring scene to associate the viewing fields of the surveillance cameras with each other. The method according to embodiments of the present disclosure does not require the plurality of surveillance cameras to be synchronous, and can deal with conditions in which the viewing fields the surveillance cameras do not overlap with each other. When a street view car is used to sample the images, an intelligent analysis of a surveillance camera network in a city level can be implemented.
- Fig. 1 is a flow chart of a for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure
- Fig. 2 is a flow chart showing a method for reconstructing a three-dimensional point set in a monitoring scene of a plurality of surveillance cameras according to a matching result between a plurality of images according to an embodiment of the present disclosure
- Fig. 3 is a flow chart showing calculating parameters of a surveillance camera according to a reconstructed three-dimensional point set and a second image sampled by a surveillance camera according to an embodiment of the present disclosure
- Fig. 4 is a block diagram of a system for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure.
- first and second are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.
- features limited by “first” and “second” are intended to indicate or imply including one or more than one these features.
- a plurality of relates to two or more than two.
- Fig. 1 is a flow chart of a method for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure. As shown in Fig. 1, the calibrating method according to an embodiment of the present disclosure includes the following steps.
- a sampling device is used to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras, and each of the plurality of surveillance cameras is used to sample a second image.
- intrinsic parameters such as a focal distance, a length-width ratio, a principal point and a distortion factor
- the sampling device may be a camera, a cell phone, a PTZ lens or a panoramic sampling device.
- the sampling device with calibrated intrinsic parameters samples the plurality of first images or videos in the monitoring scene actively to obtain relatively complete image information of the monitoring scene, thus facilitating a subsequent three-dimensional reconstruction.
- the plurality of videos are sampled, the plurality of first images are captured from them. Furthermore, it should be ensured that at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold, such that the plurality of first images can associate the viewing fields of the surveillance cameras with each other.
- the plurality of first images or videos may be obtained at an equal interval (for example, google street view, baidu street view and soso street view), or may be obtained at an unequal interval (for example, images of the monitoring scene which have been on the internet).
- the plurality of first images or videos sampled by the sampling device and the second images sampled by the surveillance cameras are not required to be synchronous.
- a plurality of sampling devices having different positions and gestures can be used to sample the plurality of first images, or one sampling device can be used to sample the plurality of first images at different positions and with different gestures.
- the calibrating method can work as long as the positions and gestures corresponding to the plurality of first images are different from each other.
- a feature matching is performed on the plurality of first images to obtain a matching result between the plurality of first images.
- feature points such as Sift feature points, Surf feature points and Harris feature points
- feature points on each first image are firstly extracted to obtain a position and a descriptor of each feature point on the first image.
- the feature matching is performed on each two first images to obtain the matching relationship between the each two first images.
- This can be implemented automatically such as by a computer or can be implemented manually (i.e., the feature points matching with each other on the two first images are designated by person).
- adjacent first images can be obtained firstly and then the adjacent first images are matched with each other, thus reducing the working amount of matching.
- the plurality of first images can be clustered according to Gist global features so as to obtain adjacent images, or a lexicographic tree can be established according to Sift features of the first images to quantize each first image so as to obtain the adjacent images, or GPS location information obtained when the sampling device samples the first images can be used to obtain the adjacent images, or adjacent frames captured from the videos can be used as the adjacent images.
- step S3 a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras is reconstructed according to the matching result between the plurality of first images.
- This step can be implemented by many methods, such as a SFM (Structure From Motion) method based on an IBA (Incremental Bundle Adjustment) frame (refer to S. Agarwal, N. Snavely, I. Simon, S. Seitz, R. Szeliski. Building Rome in a Day, ICCV, 2009), a SFM method based on discrete belief propagation and Levenberg-Marquardt method (refer to Crandall D, Owens A, Snavely N, et al. Discrete-continuous optimization for large-scale structure from motion[C], CVPR, 2011) or a method based on the GPS information and gesture information obtained by a sensor on the sampling device.
- a SFM Structure From Motion
- IBA Intelligent Bundle Adjustment
- the SFM method based on the IBA frame is used to reconstruct the three-dimensional point set of the monitoring scene. Specifically, the three-dimensional point set of the monitoring scene is reconstructed by following steps.
- two first images satisfying a predetermined condition are selected from the plurality of first images as initial images.
- a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation degree between the initial images is smaller than a predetermined degree.
- an initial three-dimensional point set is reconstructed according to the initial images.
- a three-dimensional coordinate X . of the feature point pair (x jV x j2 ) on the initial images can be obtained according to the camera matrices P l and P 2 by means of a triangulation algorithm.
- the three-dimensional point set of the monitoring scene is expanded according to the initial three-dimensional point set and other first images except for the initial images, thus obtaining a denser representation of the monitoring scene. Specifically, following steps are executed to expand the three-dimensional point set.
- step S3131 it is determined whether each of the plurality of first images has been used to reconstruct the three-dimensional point set; if yes, the process is stopped and the reconstructed three-dimensional point set is output; if no, execute step S3132.
- a first image at most matched with the reconstructed three-dimensional point is found out from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set is calculated.
- the reconstructed three-dimensional point set is denoted as X _ set
- a corresponding set of feature points on images with calculated camera matrices (hereinafter referred to as a feature point set corresponding the three-dimensional point set) is denoted as x_ set
- the first image at most matched with the reconstructed three-dimensional point set is the first image having most feature points matched with the feature points in the set x_ set
- a set of these feature points matched with the feature points in the set x_ set is denoted as x_ set l
- a three-dimensional point set corresponding to the set x_ set is a subset of X _ set and is denoted as X _ set l .
- the camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point can be calculated according to sets X _ set l and x_ set l by means of direct linear transformation.
- the three-dimensional point set of the monitoring scene is expanded according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices, and the calculated camera matrices and the three-dimensional point set are optimized.
- all the matching points between the image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices are read (for example, directly read from the matching relationship obtained at step S2) and added into the set x_ set . Then, three-dimensional coordinates of these added matching points are calculated by means of the triangulation algorithm and added into the set X _ set .
- a bundle adjustment method can be used to optimize the camera matrices and the three-dimensional points.
- the calculated camera matrix P i and the coordinate of the three-dimensional point X j can be used as variables to make an iterative optimization to minimize a total back-projection error, i.e., an equation min arg P:X .
- n is a number of the calculated camera matrices
- m is a number of the reconstructed three-dimensional points
- x tj is a homogeneous coordinate of the j m feature point in the i m image, w..
- the three-dimensional point set X _set of the monitoring scene can be reconstructed by means of the SFM method based on the discrete belief propagation and the Levenberg-Marquardt algorithm.
- Specific implementations can be referred to Crandall, D., Owens, A., Snavely, N., Huttenlocher, D.P: Discrete-continuous optimization for large-scale structure from motion. CVPR, 2011, the entire content of which is included herein by reference.
- a set of images used to reconstruct the three-dimensional point set is denoted as 7_set
- a set comprising feature points in the set 7_set corresponding to the three-dimensional points in the set I_set is denoted as x_set .
- the three-dimensional point set of the monitoring scene can be directly reconstructed according to the GPS location information and gesture information obtained by the sensor on the sampling device. Specifically, the GPS location information and gesture information corresponding to each of the plurality of first images are obtained and the camera matrix , corresponding to each of the plurality of first images is calculated. Subsequently, for each two adjacent first images, the three-dimensional coordinate X j of the matching point pair (x ⁇ , x/2) on the two adjacent first images is calculated according to the camera matrices corresponding to the two adjacent first images by means of the triangulation algorithm. Finally, the camera matrices and the three-dimensional coordinates of the matching point pairs are optimized by means of the bundle adjustment method described at step S3133.
- the set of the images used to reconstruct the three-dimensional point set is denoted as 7_set (since some images do not have matching points with other images, they are not used to reconstruct the three-dimensional point set), the obtained three-dimensional point set is denoted as I _set and the set comprising feature points in / set corresponding to three-dimensional points in I _set is denoted as x_set .
- the SFM method based on the IBA frame can be used to reconstruct the three-dimensional point set of the monitoring scene, although a calculating speed of the method is relatively low.
- the SFM method based on the discrete belief propagation and the Levenberg-Marquardt algorithm is adapted for reconstructing a large-scale scene, during which various prior information of the image (such as lines on the image) can be used to make the reconstruction more accurate, thus providing a high robustness against noises (such as noises of the GPS sensor or angle sensor).
- step S4 parameters of each of the plurality of surveillance camera are calculated according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras. Specifically, following steps are carried out to calculate the parameters of each surveillance camera.
- the second image sampled by the surveillance camera is matched with the first images used to reconstruct the three-dimensional point set respectively to obtain a first feature point set.
- the first feature point set is matched with the feature point set corresponding to the reconstructed three-dimensional point set.
- the second image sampled by the surveillance camera is matched with each image in / set , and a set comprising matched feature points (i.e., the first feature point set) denoted as x_set is obtained. Then, an intersection set between x_set and x_set is calculated and denoted as x_result . Three-dimensional coordinates of points in x_result have been reconstructed (recorded in X set ). When a number of points in x result is large enough, the camera matrix of the surveillance camera can be calculated according to the points in x result .
- step S413 it is determined whether the number of feature points in the intersection set x_result is larger than a predetermined number.
- the predetermined number is 10.
- step S414 when the number of feature points in the intersection set x result is smaller than the predetermined number, at least one third image is sampled by the sampling device in the monitoring scene of the plurality of surveillance cameras.
- the number of feature points in the intersection set x_result is smaller than the predetermined number, some images close to the viewing filed of the surveillance camera and having a relatively large area overlapped with the viewing field of the previously sampled first images are sampled. For example, when a mounting height of the surveillance camera is 10m and an erecting height of the sampling device at step S I is 2m, some images can be sampled at the height of 4m, 6m or 8m. By sampling more images, the three-dimensional point set of the monitoring scene is expanded, such that more feature points on the images sampled by the surveillance camera can be matched with the reconstructed three-dimensional point set of the monitoring scene, thus improving the calibration precision.
- step S415 the reconstructed three-dimensional point set X _set and the intersection set x_result are expanded by the at least one third image.
- features points on the third image are extracted and matched with the each image in 7_set (by the same method as described at step S2), and a set comprising matched feature points is denoted as x_new .
- a set comprising matched feature points is denoted as x_new .
- an intersection set between x_set and x_new is obtained and denoted as x_set 2
- a three-dimensional point set corresponding to x_set 2 which is a subset of I _set is obtained and denoted as Z_set 2 .
- a direct linear transformation is performed on the sets x_set 2 and Z_set 2 to calculate a camera matrix corresponding to the third image.
- the at least one third image is added into 7_set , all the points in x_new are added into x_set , and three-dimensional coordinates of these newly added points in x_set (i.e., points except for points in x_set 2 ) are calculated by the triangulation algorithm and added into X _set .
- step S416 when the number of points in x result is larger than the predetermined number, the parameters of the surveillance camera are calculated.
- CVPR 2008 Anchorage, Alaska, USA, June 2008
- CVPR 2008 Anchorage, Alaska, USA, June 2008
- X result a result of an optical center of the surveillance camera
- the p4p method is stable, but just adapted for a condition in which the intrinsic matrix has only one unknown parameter (i.e., supposing that the principal point of the surveillance camera is a center of the image and the image is not distorted).
- All the intrinsic parameters (the focal distance, a coordinate of the principle point and the distortion factor) of the surveillance camera can be calculated by the direct linear transformation, but this method is not stable enough. For the sake of clarity, these two methods will not be described in detail herein.
- a system for calibrating a plurality of surveillance cameras is provided.
- Fig. 4 is a block diagram of a system for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure.
- the calibrating system includes at least one sampling device 1000 and a computing device 2000.
- the at least one sampling device 1000 is configured to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras.
- the computing device 2000 is configured to perform a feature matching on the plurality of first images to obtain a matching result between the plurality of first images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result, and to calculate parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and a second image sampled by each of the plurality of surveillance cameras.
- the at least one sampling device 1000 may include a GPS sensor.
- the at least one sampling device 1000 includes the GPS sensor, adjacent images of each of the plurality of first images can be obtained according to the GPS location information sampled by the GPS sensor when the sampling device 1000 samples the plurality of first images. Then, the adjacent images can be matched with each other, thus improving the matching efficiency.
- the at least one sampling device 1000 may also sample images actively.
- a working process of the computing device 2000 can be referred to above description, which will be omitted herein.
- the three-dimensional point set of the monitoring scene can be accurately reconstructed according to the relatively complete image information of the monitoring scene obtained by the sampling device, and the camera matrices of the surveillance cameras can be calculated according to the relationship between the three-dimensional point set and the feature point set on the second image sampled by the surveillance camera, thus achieving the calibration of the surveillance camera.
- the calibrating process is dynamic, i.e., when the number of the reconstructed three-dimensional points is not large enough to calibrate the surveillance cameras, more images can be sampled to expand the reconstructed three-dimensional point set, thus ensuring that all the surveillance cameras can be calibrated.
- the calibrating method in the present disclosure does not require a calibration object, a special motion of the surveillance camera or overlapped viewing fields of the surveillance cameras. Rather, the calibrating method according to embodiments of the present disclosure can calibrate a plurality of surveillance cameras simultaneously as long as the sampling device samples the images of the monitoring scene to associate the viewing fields of the surveillance cameras with each other. The method according to embodiments of the present disclosure does not require the plurality of surveillance cameras to be synchronous, and can deal with conditions in which the viewing fields the surveillance cameras do not overlap with each other. When a street view car is used to sample the images, an intelligent analysis of a surveillance camera network in a city level can be implemented.
- Any procedure or method described in the flow charts or described in any other way herein may be understood to comprise one or more modules, portions or parts for storing executable codes that realize particular logic functions or procedures.
- advantageous embodiments of the present disclosure comprises other implementations in which the order of execution is different from that which is depicted or discussed, including executing functions in a substantially simultaneous manner or in an opposite order according to the related functions. This should be understood by those skilled in the art which embodiments of the present disclosure belong to.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Signal Processing (AREA)
- Closed-Circuit Television Systems (AREA)
- Studio Devices (AREA)
- Image Analysis (AREA)
Abstract
A method and a system for calibrating a plurality of surveillance cameras are provided. The method includes: sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device and sampling a second image by each of the plurality of surveillance cameras; performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images; reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality first of images; and calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
Description
METHOD AND SYSTEM FOR CALIBRATING SURVEILLANCE CAMERAS
CROSS REFERENCE TO RELATED APPLICATION
This application claims priority and benefits of Chinese Patent Application No. 201310670982.0, filed with State Intellectual Property Office on December 10, 2013, the entire content of which is incorporated herein by reference.
FIELD
Embodiments of the present disclosure generally relate to an image processing technology, and more particularly, to a method and a system for calibrating a plurality of surveillance cameras.
BACKGROUND
Video monitoring is an important part of a safeguard system, in which one key issue is how to calibrate surveillance cameras in a monitoring scene to obtain their intrinsic parameters (such as a coordinate of a principle point, a focal length and a distortion factor) and external parameters (such as a rotation matrix and a translation vector of a coordinate system of the surveillance camera with respect to a universal coordinate system) thereof, so as to obtain a relative erecting position, a facing direction and a viewing filed of each surveillance camera in the safeguard system.
Currently, there are three calibrating methods for the surveillance camera: a conventional calibrating method, a calibrating method based on active vision and a self -calibrating method. In the conventional calibrating method, an precisely machined calibration object (such a calibration block, a calibration plate or a calibration bar) is placed in the viewing field of the surveillance camera, and then the intrinsic and external parameters of the surveillance camera are obtained by establishing a relationship between a given coordinate of a point on the calibration object and a coordinate of the point on an image. This method wastes time and energy because above operation is required to be performed on each surveillance camera, and only the intrinsic parameters can be calibrated by this method. Furthermore, since the outdoor monitoring scene is usually large and a size of the calibration object is relatively small and occupies a little proportion of the image, a calibration error of this method is large. In the calibrating method based on active vision, the surveillance camera is controlled actively to move in a special mode (such as a pure translation motion or a rotation around an optical center), and the intrinsic parameters of the surveillance
camera are calculated according to the specificity of the motion. However, since most of the surveillance cameras are installed on the fixed location and the motion thereof is hard to control, this method has a poor applicability. Furthermore, with this method, it is also required to calibrate each surveillance camera respectively, thus wasting time and energy, and moreover, the external parameters cannot be calibrated. In the self-calibrating method, no calibration object is needed and the surveillance camera is not required to move in the special mode. The intrinsic and external parameters of the surveillance camera is directly calibrated according to the relationship between pixels on a plurality of images sampled by the surveillance camera (maybe a plurality of surveillance cameras) and constraints of the intrinsic and external parameters. Some self-calibrating methods can calibrate a plurality of surveillance cameras simultaneously by means of a multiple view geometry method. However, the viewing fields of the plurality of surveillance cameras are required to have a large area overlapped with each other, otherwise the plurality of surveillance cameras will not be calibrated due to an image matching failure. However, in practice, the overlapped area between viewing fields of the plurality of surveillance cameras is small. Thus, the currently known self-calibrating methods have difficulties in calibrating the plurality of surveillance cameras simultaneously.
SUMMARY
Embodiments of the present disclosure seek to solve at least one of the problems existing in the related art to at least some extent.
For this, one object of the present disclosure is to provide a method for calibrating a plurality of surveillance cameras which can calibrate intrinsic and external parameters of the surveillance cameras having no overlapped viewing field.
Another object of the present disclosure is to provide a system for calibrating a plurality of surveillance cameras which can calibrate intrinsic and external parameters of the surveillance cameras having no overlapped viewing field.
According to embodiments of a first broad aspect of the present disclosure, a method for calibrating a plurality of surveillance cameras is provided. The method includes: sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device and sampling a second image in the monitoring scene by each of the plurality of surveillance cameras respectively; performing a feature matching on the plurality of first images to
obtain a matching result between the plurality of first images; reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result between the plurality of first images; and calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
In some embodiments, sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device includes: calibrating intrinsic parameters of the sampling device; sampling the plurality of first images in the monitoring scene of the plurality of surveillance cameras by the sampling device with calibrated intrinsic parameters, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
In some embodiments, performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images includes: when the sampling device has a GPS sensor, obtaining first images adjacent to each other according to GPS location information obtained during sampling the plurality of first images, and matching the adjacent first images with each other; and when the sampling device has no GPS sensor, matching the plurality of first images with each other.
In some embodiments, reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result includes: selecting two first images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation degree between the initial images is smaller than a predetermined degree; calculating camera matrices corresponding to the initial images according to the matching point pairs on the initial images; reconstructing a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching relationship between the initial images ; expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images.
In some embodiments, expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images includes: determining whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene; if no, finding out a first image
at most matched with the reconstructed three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculating a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set; expanding the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices; optimizing the calculated camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
In other embodiments, reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality of first images includes: obtaining GPS location information and gesture information corresponding to each of the plurality of first images and calculating the camera matrix corresponding to each of the plurality of first images according to the GPS location information, the gesture information and the intrinsic parameters of the sampling device; for each two adjacent first images, calculating three-dimensional coordinates of matching points on the two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm; optimizing the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
In some embodiments, calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras includes: for each surveillance camera to be calibrated, matching the sampled second image with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculating an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set includes points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set; determining whether a number of feature points in the intersection set is larger than a predetermined number; if no, sampling at least one third image in the monitoring scene of the plurality of surveillance cameras by the sampling device and expanding the three-dimensional point set according to the at least one third image, so as to obtain an updated intersection set in which the number of feature points is larger
than the predetermined number; calculating the parameters of the surveillance camera to be calibrated according to the updated intersection set and a three-dimensional point set corresponding to the updated intersection set.
According to embodiments of a second broad aspect of the present disclosure, a system for calibrating a plurality of surveillance cameras is provided, and the system includes: at least one sampling device, configured to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras; and a computing device, configured to perform a feature matching on the plurality of first images to obtain a matching result between the plurality of first images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result between the plurality of first images, and to calculate parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
In some embodiments, intrinsic parameters of the at least one sampling device are calibrated, and the at least one sampling device with calibrated intrinsic parameters actively samples the plurality of first images in the monitoring scene of plurality of the surveillance cameras, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
In some embodiments, the computing device is configured to: select two images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation between the initial images is smaller than a predetermined degree, to calculate camera matrices corresponding to the initial images according to the matching point pairs on the initial images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching relationship between the initial images and the camera matrices corresponding to the initial images, and to expand the three-dimensional point set of the monitoring scene according to the three-dimensional point set and other first images except for the initial images.
In some embodiments, the computing device is further configured to determine whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene, if no, find out a first image at most matched with the reconstructed
three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculate a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set, expand the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices, optimize the camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
In other embodiments, the computing device is configured to obtain GPS location information and gesture information corresponding to each of the plurality of first images and calculate the camera matrix corresponding to each of the plurality of first images according to the GPS location information, the gesture information and the intrinsic parameters of the at least one sampling device, calculate three-dimensional coordinates of matching points on each two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm, optimize the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
In some embodiments, the computing device is further configured to match the second image sampled by each of the plurality of surveillance cameras with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculate an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set includes points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set, determine whether a number of feature points in the intersection set is larger than a predetermined number; if no, receive at least one third image sampled by the at least one sampling device and expand the three-dimensional point set according to the at least one third image so as to obtain an updated intersection set in which the number of feature points is larger than the predetermined number, calculate the parameters of each of the plurality of surveillance cameras according to the intersection set and a three-dimensional point set corresponding to the intersection set.
With the calibrating method and system according to embodiments of the present disclosure, the three-dimensional point set of the monitoring scene can be accurately reconstructed according to the relatively complete image information of the monitoring scene obtained by the sampling
device, and the camera matrices of the surveillance cameras can be calculated according to the relationship between the three-dimensional point set and the feature point set on the second images sampled by the surveillance cameras, thus achieving the calibration of the surveillance cameras. Moreover, the calibrating process is dynamic, i.e., when the number of the reconstructed three-dimensional points is not large enough to calibrate the surveillance cameras, more images can be sampled to expand the reconstructed three-dimensional point set, thus ensuring that all the surveillance cameras can be calibrated. Compared with conventional calibrating methods, the calibrating method in the present disclosure does not require a calibration object, a special motion of the surveillance camera or overlapped viewing fields of the surveillance cameras. Rather, the calibrating method according to embodiments of the present disclosure can calibrate a plurality of surveillance cameras simultaneously as long as the sampling device samples the images of the monitoring scene to associate the viewing fields of the surveillance cameras with each other. The method according to embodiments of the present disclosure does not require the plurality of surveillance cameras to be synchronous, and can deal with conditions in which the viewing fields the surveillance cameras do not overlap with each other. When a street view car is used to sample the images, an intelligent analysis of a surveillance camera network in a city level can be implemented.
Additional aspects and advantages of embodiments of present disclosure will be given in part in the following descriptions, become apparent in part from the following descriptions, or be learned from the practice of the embodiments of the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
These and other aspects and advantages of embodiments of the present disclosure will become apparent and more readily appreciated from the following descriptions made with reference to the accompanying drawings, in which:
Fig. 1 is a flow chart of a for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure;
Fig. 2 is a flow chart showing a method for reconstructing a three-dimensional point set in a monitoring scene of a plurality of surveillance cameras according to a matching result between a plurality of images according to an embodiment of the present disclosure;
Fig. 3 is a flow chart showing calculating parameters of a surveillance camera according to a
reconstructed three-dimensional point set and a second image sampled by a surveillance camera according to an embodiment of the present disclosure;
Fig. 4 is a block diagram of a system for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure.
DETAILED DESCRIPTION
Reference will be made in detail to embodiments of the present disclosure. The same or similar elements and the elements having same or similar functions are denoted by like reference numerals throughout the descriptions. The embodiments described herein with reference to drawings are explanatory, illustrative, and used to generally understand the present disclosure. The embodiments shall not be construed to limit the present disclosure.
In addition, terms such as "first" and "second" are used herein for purposes of description and are not intended to indicate or imply relative importance or significance. Thus, features limited by "first" and "second" are intended to indicate or imply including one or more than one these features. In the description of the present disclosure, "a plurality of relates to two or more than two.
Fig. 1 is a flow chart of a method for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure. As shown in Fig. 1, the calibrating method according to an embodiment of the present disclosure includes the following steps.
At step SI, a sampling device is used to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras, and each of the plurality of surveillance cameras is used to sample a second image.
In an embodiment of the present disclosure, intrinsic parameters (such as a focal distance, a length-width ratio, a principal point and a distortion factor) of the sampling device may be calibrated by conventional calibrating methods, such that a precision of a scene reconstruction can be improved, and thus the surveillance camera can be calibrated more precisely. In an embodiment of the present disclosure, the sampling device may be a camera, a cell phone, a PTZ lens or a panoramic sampling device.
Then, the sampling device with calibrated intrinsic parameters samples the plurality of first images or videos in the monitoring scene actively to obtain relatively complete image information of the monitoring scene, thus facilitating a subsequent three-dimensional reconstruction. When the
plurality of videos are sampled, the plurality of first images are captured from them. Furthermore, it should be ensured that at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold, such that the plurality of first images can associate the viewing fields of the surveillance cameras with each other. The plurality of first images or videos may be obtained at an equal interval (for example, google street view, baidu street view and soso street view), or may be obtained at an unequal interval (for example, images of the monitoring scene which have been on the internet). In the embodiment of the present disclosure, the plurality of first images or videos sampled by the sampling device and the second images sampled by the surveillance cameras are not required to be synchronous. In the present disclosure, a plurality of sampling devices having different positions and gestures can be used to sample the plurality of first images, or one sampling device can be used to sample the plurality of first images at different positions and with different gestures. In other words, the calibrating method can work as long as the positions and gestures corresponding to the plurality of first images are different from each other.
At step S2, a feature matching is performed on the plurality of first images to obtain a matching result between the plurality of first images.
Specifically, feature points (such as Sift feature points, Surf feature points and Harris feature points) on each first image are firstly extracted to obtain a position and a descriptor of each feature point on the first image. Then, the feature matching is performed on each two first images to obtain the matching relationship between the each two first images. This can be implemented automatically such as by a computer or can be implemented manually (i.e., the feature points matching with each other on the two first images are designated by person). Furthermore, many methods for implementing the feature matching automatically have been presented, such as a feature matching method based on the nearest neighbor ratio constraint of the feature point descriptor (refer to Lowe, D.G., "Distinctive Image Feature from Scale-Invariant Keypoints " , International Journal of Computer Vision, 60, 2, pages 91-110, 2004), a feature matching method based on the epipolar geometry constraint of the feature point position (refer to M.A. Fischer and R.C.Bolles, "Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography" , Commun.Acm, volume 24, pages 381-395, June. 1981) and a feature matching method based on a WGTM algorithm (referto Saeedi P, Izadi M. "A Robust Weighted Graph Transformation Matching for Rigid and Non-rigid Image Registration [J] " .
2012), the entire content of which are included herein by reference.
Since a working amount of matching each two first images is huge, in this embodiment, adjacent first images can be obtained firstly and then the adjacent first images are matched with each other, thus reducing the working amount of matching. Specifically, the plurality of first images can be clustered according to Gist global features so as to obtain adjacent images, or a lexicographic tree can be established according to Sift features of the first images to quantize each first image so as to obtain the adjacent images, or GPS location information obtained when the sampling device samples the first images can be used to obtain the adjacent images, or adjacent frames captured from the videos can be used as the adjacent images.
At step S3, a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras is reconstructed according to the matching result between the plurality of first images.
This step can be implemented by many methods, such as a SFM (Structure From Motion) method based on an IBA (Incremental Bundle Adjustment) frame (refer to S. Agarwal, N. Snavely, I. Simon, S. Seitz, R. Szeliski. Building Rome in a Day, ICCV, 2009), a SFM method based on discrete belief propagation and Levenberg-Marquardt method (refer to Crandall D, Owens A, Snavely N, et al. Discrete-continuous optimization for large-scale structure from motion[C], CVPR, 2011) or a method based on the GPS information and gesture information obtained by a sensor on the sampling device.
In an embodiment of the present disclosure, the SFM method based on the IBA frame is used to reconstruct the three-dimensional point set of the monitoring scene. Specifically, the three-dimensional point set of the monitoring scene is reconstructed by following steps.
At step S311, two first images satisfying a predetermined condition are selected from the plurality of first images as initial images. A number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation degree between the initial images is smaller than a predetermined degree.
At step S312, an initial three-dimensional point set is reconstructed according to the initial images.
Firstly, an essential matrix E which satisfies xj T lExj2 = O is calculated, such as by a five-spot method proposed by David Nister (refer to Nister D. An efficient solution to the five-point relative pose problem [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2004), in
which ( ji, Xj2) is a feature point pair on the initial images. Then, the essential matrix E is decomposed to obtain two camera matrices Px and P2 corresponding to the initial images respectively (referring to Richard Hartley, "Multiple view geometry in computer vision", Chapter 8, Section 6, a coordinate system corresponding to one of the initial images is used as a world coordinate system, i.e., the camera matrix corresponding to one of the initial images is represented as Pl = ^Tj[II0] , where Kx is a known intrinsic matrix of the sampling device, / is a 3 x 3 unit matrix, and 0 is a 3 x 1 zero vector). Finally, a three-dimensional coordinate X . of the feature point pair (xjV xj2) on the initial images can be obtained according to the camera matrices Pl and P2 by means of a triangulation algorithm.
At step S313, the three-dimensional point set of the monitoring scene is expanded according to the initial three-dimensional point set and other first images except for the initial images, thus obtaining a denser representation of the monitoring scene. Specifically, following steps are executed to expand the three-dimensional point set.
At step S3131, it is determined whether each of the plurality of first images has been used to reconstruct the three-dimensional point set; if yes, the process is stopped and the reconstructed three-dimensional point set is output; if no, execute step S3132.
At step S3132, a first image at most matched with the reconstructed three-dimensional point is found out from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set is calculated.
The reconstructed three-dimensional point set is denoted as X _ set , and a corresponding set of feature points on images with calculated camera matrices (hereinafter referred to as a feature point set corresponding the three-dimensional point set) is denoted as x_ set . The first image at most matched with the reconstructed three-dimensional point set is the first image having most feature points matched with the feature points in the set x_ set , and a set of these feature points matched with the feature points in the set x_ set is denoted as x_ setl . A three-dimensional point set corresponding to the set x_ set is a subset of X _ set and is denoted as X _ setl . Thus, the camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point can be calculated according to sets X _ setl and x_ setl by means of direct linear transformation.
At step S3133, the three-dimensional point set of the monitoring scene is expanded according
to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices, and the calculated camera matrices and the three-dimensional point set are optimized.
Specifically, all the matching points between the image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices are read (for example, directly read from the matching relationship obtained at step S2) and added into the set x_ set . Then, three-dimensional coordinates of these added matching points are calculated by means of the triangulation algorithm and added into the set X _ set .
In addition, a bundle adjustment method can be used to optimize the camera matrices and the three-dimensional points. Specifically, the calculated camera matrix Pi and the coordinate of the three-dimensional point Xj can be used as variables to make an iterative optimization to minimize a total back-projection error, i.e., an equation min arg P:X . (where n is a number of the calculated camera matrices, m is a number of the reconstructed three-dimensional points, xtj is a homogeneous coordinate of the jm feature point in the im image, w.. = 1, the im feature point occurring in the im imag &e Λ ) i ·s so ,lved ,. S co ,lvi ·ng the equat .i·on is a
^ [0, other conditions
nonlinear least square problem and can be dealt with a LM algorithm (refer to Kenneth Levenberg (1944). "A Method for the Solution of Certain Non-Linear Problems in Least Squares". Quarterly of Applied Mathematics 2: 164-168, the calculated camera matrix Pi and the three-dimensional coordinate X j can be used as initial values of the iterative optimization).
The steps S3131 to S3133 are repeated until no more three-dimensional point can be found.
In another embodiment of the present disclosure, with the matching result obtained at step S2, the three-dimensional point set X _set of the monitoring scene can be reconstructed by means of the SFM method based on the discrete belief propagation and the Levenberg-Marquardt algorithm. Specific implementations can be referred to Crandall, D., Owens, A., Snavely, N., Huttenlocher, D.P: Discrete-continuous optimization for large-scale structure from motion. CVPR, 2011, the entire content of which is included herein by reference. A set of images used to reconstruct the three-dimensional point set is denoted as 7_set , and a set comprising feature points in the set 7_set corresponding to the three-dimensional points in the set I_set is denoted as x_set .
In yet another embodiment of the present disclosure, the three-dimensional point set of the
monitoring scene can be directly reconstructed according to the GPS location information and gesture information obtained by the sensor on the sampling device. Specifically, the GPS location information and gesture information corresponding to each of the plurality of first images are obtained and the camera matrix , corresponding to each of the plurality of first images is calculated. Subsequently, for each two adjacent first images, the three-dimensional coordinate Xj of the matching point pair (x^, x/2) on the two adjacent first images is calculated according to the camera matrices corresponding to the two adjacent first images by means of the triangulation algorithm. Finally, the camera matrices and the three-dimensional coordinates of the matching point pairs are optimized by means of the bundle adjustment method described at step S3133. The set of the images used to reconstruct the three-dimensional point set is denoted as 7_set (since some images do not have matching points with other images, they are not used to reconstruct the three-dimensional point set), the obtained three-dimensional point set is denoted as I _set and the set comprising feature points in / set corresponding to three-dimensional points in I _set is denoted as x_set .
It should be understood that, those having ordinary skills in the related art can select one of the above three methods according to different requirements and conditions. For example, when a sampling device with a GPS sensor and an angle sensor is used to sample images in the monitoring scene, it is easy to obtain sensor information (location information and angle information obtained by the GPS sensor and the angle sensor). Thus, the third method can be used, and the three-dimensional point set of the monitoring scene can be reconstructed simply and efficiently according to the GPS information and gesture information sampled by the sensors on the sampling device. However, when the angle information, location information or intrinsic parameters of the image (such as internet images and indoor sampled images) is missed, the SFM method based on the IBA frame can be used to reconstruct the three-dimensional point set of the monitoring scene, although a calculating speed of the method is relatively low. The SFM method based on the discrete belief propagation and the Levenberg-Marquardt algorithm is adapted for reconstructing a large-scale scene, during which various prior information of the image (such as lines on the image) can be used to make the reconstruction more accurate, thus providing a high robustness against noises (such as noises of the GPS sensor or angle sensor).
At step S4, parameters of each of the plurality of surveillance camera are calculated according to the three-dimensional point set and the second image sampled by each of the plurality of
surveillance cameras. Specifically, following steps are carried out to calculate the parameters of each surveillance camera.
At step S411, the second image sampled by the surveillance camera is matched with the first images used to reconstruct the three-dimensional point set respectively to obtain a first feature point set.
At step S412, the first feature point set is matched with the feature point set corresponding to the reconstructed three-dimensional point set.
The second image sampled by the surveillance camera is matched with each image in / set , and a set comprising matched feature points (i.e., the first feature point set) denoted as x_set is obtained. Then, an intersection set between x_set and x_set is calculated and denoted as x_result . Three-dimensional coordinates of points in x_result have been reconstructed (recorded in X set ). When a number of points in x result is large enough, the camera matrix of the surveillance camera can be calculated according to the points in x result .
At step S413, it is determined whether the number of feature points in the intersection set x_result is larger than a predetermined number. In an embodiment of the present disclosure, the predetermined number is 10.
At step S414, when the number of feature points in the intersection set x result is smaller than the predetermined number, at least one third image is sampled by the sampling device in the monitoring scene of the plurality of surveillance cameras.
When the number of feature points in the intersection set x_result is smaller than the predetermined number, some images close to the viewing filed of the surveillance camera and having a relatively large area overlapped with the viewing field of the previously sampled first images are sampled. For example, when a mounting height of the surveillance camera is 10m and an erecting height of the sampling device at step S I is 2m, some images can be sampled at the height of 4m, 6m or 8m. By sampling more images, the three-dimensional point set of the monitoring scene is expanded, such that more feature points on the images sampled by the surveillance camera can be matched with the reconstructed three-dimensional point set of the monitoring scene, thus improving the calibration precision.
At step S415, the reconstructed three-dimensional point set X _set and the intersection set x_result are expanded by the at least one third image.
Specifically, features points on the third image are extracted and matched with the each image
in 7_set (by the same method as described at step S2), and a set comprising matched feature points is denoted as x_new . Then, an intersection set between x_set and x_new is obtained and denoted as x_set2 , and a three-dimensional point set corresponding to x_set2 which is a subset of I _set is obtained and denoted as Z_set2 . Subsequently, a direct linear transformation is performed on the sets x_set2 and Z_set2 to calculate a camera matrix corresponding to the third image. Finally, the at least one third image is added into 7_set , all the points in x_new are added into x_set , and three-dimensional coordinates of these newly added points in x_set (i.e., points except for points in x_set2 ) are calculated by the triangulation algorithm and added into X _set .
The above steps are repeated until the number of points in x_result is larger than the predetermined number.
At step S416, when the number of points in x result is larger than the predetermined number, the parameters of the surveillance camera are calculated.
Firstly, a three-dimensional point set X _result (a subset of X _set ) corresponding to x_result is obtained. Then, the direct linear transformation (refer to "Multiple view geometry in computer vision", Chapter 2 Section 5) or p4p method (refer to M. Bujnak, Z. Kukelova, T. Pajdla. A general solution to the P4P problem for camera with unknown focal length. CVPR 2008, Anchorage, Alaska, USA, June 2008) is performed on the sets x result and X result to calculate an intrinsic matrix (including a focal distance of the surveillance camera) and external parameters (a rotation matrix of the coordinate system of the surveillance camera with respect to the world coordinate system and a location of an optical center of the surveillance camera) of surveillance camera. The p4p method is stable, but just adapted for a condition in which the intrinsic matrix has only one unknown parameter (i.e., supposing that the principal point of the surveillance camera is a center of the image and the image is not distorted). All the intrinsic parameters (the focal distance, a coordinate of the principle point and the distortion factor) of the surveillance camera can be calculated by the direct linear transformation, but this method is not stable enough. For the sake of clarity, these two methods will not be described in detail herein.
According to embodiments of the present disclosure, a system for calibrating a plurality of surveillance cameras is provided. Fig. 4 is a block diagram of a system for calibrating a plurality of surveillance cameras according to an embodiment of the present disclosure. As shown in Fig. 4, the calibrating system includes at least one sampling device 1000 and a computing device 2000.
The at least one sampling device 1000 is configured to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras. The computing device 2000 is configured to perform a feature matching on the plurality of first images to obtain a matching result between the plurality of first images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result, and to calculate parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and a second image sampled by each of the plurality of surveillance cameras.
In an embodiment of the present disclosure, the at least one sampling device 1000 may include a GPS sensor. When the at least one sampling device 1000 includes the GPS sensor, adjacent images of each of the plurality of first images can be obtained according to the GPS location information sampled by the GPS sensor when the sampling device 1000 samples the plurality of first images. Then, the adjacent images can be matched with each other, thus improving the matching efficiency. The at least one sampling device 1000 may also sample images actively.
A working process of the computing device 2000 can be referred to above description, which will be omitted herein.
With the calibrating method and system according to embodiments of the present disclosure, the three-dimensional point set of the monitoring scene can be accurately reconstructed according to the relatively complete image information of the monitoring scene obtained by the sampling device, and the camera matrices of the surveillance cameras can be calculated according to the relationship between the three-dimensional point set and the feature point set on the second image sampled by the surveillance camera, thus achieving the calibration of the surveillance camera. Moreover, the calibrating process is dynamic, i.e., when the number of the reconstructed three-dimensional points is not large enough to calibrate the surveillance cameras, more images can be sampled to expand the reconstructed three-dimensional point set, thus ensuring that all the surveillance cameras can be calibrated. Compared with conventional calibrating methods, the calibrating method in the present disclosure does not require a calibration object, a special motion of the surveillance camera or overlapped viewing fields of the surveillance cameras. Rather, the calibrating method according to embodiments of the present disclosure can calibrate a plurality of surveillance cameras simultaneously as long as the sampling device samples the images of the
monitoring scene to associate the viewing fields of the surveillance cameras with each other. The method according to embodiments of the present disclosure does not require the plurality of surveillance cameras to be synchronous, and can deal with conditions in which the viewing fields the surveillance cameras do not overlap with each other. When a street view car is used to sample the images, an intelligent analysis of a surveillance camera network in a city level can be implemented.
Any procedure or method described in the flow charts or described in any other way herein may be understood to comprise one or more modules, portions or parts for storing executable codes that realize particular logic functions or procedures. Moreover, advantageous embodiments of the present disclosure comprises other implementations in which the order of execution is different from that which is depicted or discussed, including executing functions in a substantially simultaneous manner or in an opposite order according to the related functions. This should be understood by those skilled in the art which embodiments of the present disclosure belong to.
Reference throughout this specification to "an embodiment," "some embodiments," "an example," "a specific example," or "some examples," means that a particular feature, structure, material, or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present disclosure. The appearances of the phrases throughout this specification are not necessarily referring to the same embodiment or example of the present disclosure. Furthermore, the particular features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
Although explanatory embodiments have been shown and described, it would be appreciated by those skilled in the art that the above embodiments cannot be construed to limit the present disclosure, and changes, alternatives, and modifications can be made in the embodiments without departing from spirit, principles and scope of the present disclosure.
Claims
1. A method for calibrating a plurality of surveillance cameras, comprising:
sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device and sampling a second image in the monitoring scene by each of the plurality of surveillance cameras;
performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images;
reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality of first images; and
calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras.
2. The method according to claim 1, wherein sampling a plurality of first images in a monitoring scene of the plurality of surveillance cameras by a sampling device comprises:
calibrating intrinsic parameters of the sampling device;
sampling the plurality of first images in the monitoring scene of the plurality of surveillance cameras by the sampling device with calibrated intrinsic parameters, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
3. The method according to claim 1 or 2, wherein performing a feature matching on the plurality of first images to obtain a matching result between the plurality of first images comprises: when the sampling device has a GPS sensor, obtaining first images adjacent to each other according to GPS location information obtained during sampling the plurality of first images, and matching the adjacent first images with each other; and
when the sampling device has no GPS sensor, matching the plurality of first images with each other.
4. The method according to claim 1, wherein reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality of first images comprises:
selecting two first images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a
degradation degree between the initial images is smaller than a predetermined degree; calculating camera matrices corresponding to the initial images according to the matching point pairs on the initial images;
reconstructing an initial three-dimensional point set of the monitoring scene according to the matching relationship between the initial images;
expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images.
5. The method according to claim 4, wherein expanding the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images comprises:
determining whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene;
if no, finding out a first image at most matched with the reconstructed three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculating a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set;
expanding the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices;
optimizing the calculated camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
6. The method according to claim 1 or 2, wherein reconstructing a three-dimensional point set of the monitoring scene according to the matching result between the plurality of first images comprises:
obtaining GPS location information and gesture information corresponding to each of the plurality of first images and calculating the camera matrix corresponding to each of the plurality of first images according to the GPS location information, the gesture information and the intrinsic parameters of the sampling device;
for each two adjacent first images, calculating three-dimensional coordinates of matching points on the two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm;
optimizing the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
7. The method according to claim 1, wherein calculating parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and the second image sampled by each of the plurality of surveillance cameras comprises:
for each surveillance camera to be calculated, matching the sampled second image with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculating an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set comprises points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set;
determining whether a number of feature points in the intersection set is larger than a predetermined number;
if no, sampling at least one third image in the monitoring scene of the plurality of surveillance cameras by the sampling device and expanding the three-dimensional point set according to the at least one third image, so as to obtain an updated intersection set in which the number of feature points is larger than the predetermined number;
calculating the parameters of the surveillance camera to be calibrated according to the updated intersection set and a three-dimensional point set corresponding to the updated intersection set.
8. A system for calibrating a plurality of surveillance cameras, comprising:
at least one sampling device, configured to sample a plurality of first images in a monitoring scene of the plurality of surveillance cameras; and
a computing device, configured to perform a feature matching on the plurality of first images so as to obtain a matching result between the plurality of first images, to reconstruct a three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching result between the plurality of first images, and to calculate parameters of each of the plurality of surveillance cameras according to the three-dimensional point set and a second image sampled by each of the plurality of surveillance cameras.
9. The system according to claim 8, wherein intrinsic parameters of the at least one sampling
device are calibrated, and the at least one sampling device with calibrated intrinsic parameters actively samples the plurality of first images in the monitoring scene of the plurality of surveillance cameras, in which at least some of the plurality of first images have an area overlapped with a total viewing field of the plurality of surveillance cameras and larger than a predetermined area threshold.
10. The system according to claim 8, wherein the computing device is configured to:
select two images from the plurality of first images as initial images, in which a number of matching point pairs on the initial images is larger than a predetermined threshold and a degradation degree between the initial images is smaller than a predetermined degree;
calculate camera matrices corresponding to the initial images according to the matching point pairs on the initial images;
reconstruct an initial three-dimensional point set of the monitoring scene of the plurality of surveillance cameras according to the matching relationship between the initial images and the camera matrices corresponding to the initial images; and
expand the three-dimensional point set of the monitoring scene according to the initial three-dimensional point set and other first images except for the initial images.
11. The system according to claim 10, wherein the computing device is further configured to: determine whether each of the plurality of first images has been used to reconstruct the three-dimensional point set of the monitoring scene;
if no, find out a first image at most matched with the reconstructed three-dimensional point set from images having not been used to reconstruct the three-dimensional point set of the monitoring scene, and calculate a camera matrix corresponding to the first image at most matched with the reconstructed three-dimensional point set;
expand the three-dimensional point set of the monitoring scene according to matching point pairs between the first image at most matched with the reconstructed three-dimensional point set and other first images with calculated camera matrices; and
optimize the calculated camera matrices and three-dimensional coordinates of points in the three-dimensional point set.
12. The system according to claim 8 or 9, wherein the computing device is configured to: obtain GPS location information and gesture information corresponding to each of the plurality of first images and calculate the camera matrix corresponding to each of the plurality of
first images according to the GPS location information, the gesture information and the intrinsic parameters of the at least one sampling device;
calculate three-dimensional coordinates of matching points on each two adjacent first images according to the camera matrices corresponding to the two adjacent first images by means of a triangulation algorithm; and
optimize the camera matrices and the three-dimensional coordinates of the matching points by means of a bundle adjustment method.
13. The system according claim 8, wherein the computing device is further configured to: match the second image sampled by each of the plurality of surveillance cameras with the first images used to reconstruct the three-dimensional point set of the monitoring scene respectively to obtain a first feature point set, and calculate an intersection set between the first feature point set and a feature point set corresponding to the three-dimensional point set, in which the feature point set corresponding to the three-dimensional point set comprises points located on the first images used to reconstruct the three-dimensional point set and corresponding to points in the three-dimensional point set;
determine whether a number of feature points in the intersection set is bigger than a predetermined number;
if no, receive at least one third image sampled by the at least one sampling device and expand the three-dimensional point set according to the at least one third image so as to obtain an updated intersection set in which the number of feature points is bigger than the predetermined number; and
calculate the parameters of the surveillance camera according to the updated intersection set and a three-dimensional point set corresponding to the updated intersection set.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US14/765,748 US9684962B2 (en) | 2013-12-10 | 2014-07-30 | Method and system for calibrating surveillance cameras |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201310670982.0 | 2013-12-10 | ||
| CN201310670982.0A CN103824278B (en) | 2013-12-10 | 2013-12-10 | The scaling method of CCTV camera and system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2015085779A1 true WO2015085779A1 (en) | 2015-06-18 |
Family
ID=50759320
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2014/083329 Ceased WO2015085779A1 (en) | 2013-12-10 | 2014-07-30 | Method and system for calibrating surveillance cameras |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US9684962B2 (en) |
| CN (1) | CN103824278B (en) |
| WO (1) | WO2015085779A1 (en) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112639883A (en) * | 2020-03-17 | 2021-04-09 | 华为技术有限公司 | Relative attitude calibration method and related device |
| CN113298885A (en) * | 2021-06-23 | 2021-08-24 | Oppo广东移动通信有限公司 | Binocular calibration method and device, equipment and storage medium |
| CN115733949A (en) * | 2021-08-25 | 2023-03-03 | 浙江宇视科技有限公司 | A method and device for creating a three-dimensional field of view model of a shooting device |
| KR102943303B1 (en) | 2025-10-13 | 2026-03-24 | 오픈인 주식회사 | Multi-camera based 3D shape measurement device with PTZ control |
Families Citing this family (38)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103824278B (en) | 2013-12-10 | 2016-09-21 | 清华大学 | The scaling method of CCTV camera and system |
| CN104766292B (en) * | 2014-01-02 | 2017-10-20 | 株式会社理光 | Many stereo camera calibration method and systems |
| CN105513056B (en) * | 2015-11-30 | 2018-06-19 | 天津津航技术物理研究所 | Join automatic calibration method outside vehicle-mounted monocular infrared camera |
| CN105678748B (en) * | 2015-12-30 | 2019-01-15 | 清华大学 | Interactive calibration method and device in three-dimension monitoring system based on three-dimensionalreconstruction |
| CN105844696B (en) * | 2015-12-31 | 2019-02-05 | 清华大学 | Image positioning method and device based on 3D reconstruction of ray model |
| CN105744223B (en) * | 2016-02-04 | 2019-01-29 | 北京旷视科技有限公司 | Video data processing method and device |
| CN105931229B (en) * | 2016-04-18 | 2019-02-05 | 东北大学 | Wireless camera sensor pose calibration method for wireless camera sensor network |
| CN105976391B (en) * | 2016-05-27 | 2018-12-14 | 西北工业大学 | Multiple cameras calibration method based on ORB-SLAM |
| CN106127115B (en) * | 2016-06-16 | 2020-01-31 | 哈尔滨工程大学 | A hybrid vision target localization method based on panoramic and conventional vision |
| CN106652026A (en) * | 2016-12-23 | 2017-05-10 | 安徽工程大学机电学院 | Three-dimensional space automatic calibration method based on multi-sensor fusion |
| CN107358633A (en) * | 2017-07-12 | 2017-11-17 | 北京轻威科技有限责任公司 | A Calibration Method of Internal and External Parameters of Multiple Cameras Based on Three-point Calibration Objects |
| CN107437264B (en) * | 2017-08-29 | 2020-06-19 | 重庆邮电大学 | Automatic detection and correction method of external parameters of vehicle camera |
| CN107657644B (en) * | 2017-09-28 | 2019-11-15 | 浙江大华技术股份有限公司 | Sparse scene flows detection method and device under a kind of mobile environment |
| CN107944455B (en) * | 2017-11-15 | 2020-06-02 | 天津大学 | An Image Matching Method Based on SURF |
| WO2019119328A1 (en) * | 2017-12-20 | 2019-06-27 | 深圳市大疆创新科技有限公司 | Vision-based positioning method and aerial vehicle |
| CN108346164B (en) * | 2018-01-09 | 2022-05-20 | 云南大学 | Method for calibrating conical mirror catadioptric camera by using nature of intrinsic matrix |
| CN108534797A (en) * | 2018-04-13 | 2018-09-14 | 北京航空航天大学 | A kind of real-time high-precision visual odometry method |
| CN109461189A (en) * | 2018-09-04 | 2019-03-12 | 顺丰科技有限公司 | Pose calculation method, device, equipment and the storage medium of polyphaser |
| CN109360250A (en) * | 2018-12-27 | 2019-02-19 | 爱笔(北京)智能科技有限公司 | Scaling method, equipment and the system of a kind of pair of photographic device |
| CN109508702A (en) * | 2018-12-29 | 2019-03-22 | 安徽云森物联网科技有限公司 | A kind of three-dimensional face biopsy method based on single image acquisition equipment |
| WO2020181509A1 (en) * | 2019-03-12 | 2020-09-17 | 深圳市大疆创新科技有限公司 | Image processing method, apparatus and system |
| CN111698455B (en) * | 2019-03-13 | 2022-03-11 | 华为技术有限公司 | Method, device and medium for controlling linkage between ball machine and bolt |
| CN111862225A (en) * | 2019-04-30 | 2020-10-30 | 罗伯特·博世有限公司 | Image calibration method, calibration system and vehicle with the same |
| CN110081841B (en) * | 2019-05-08 | 2021-07-02 | 上海鼎盛汽车检测设备有限公司 | Method and system for determining three-dimensional coordinates of target disc of 3D four-wheel aligner |
| US10999075B2 (en) | 2019-06-17 | 2021-05-04 | Advanced New Technologies Co., Ltd. | Blockchain-based patrol inspection proof storage method, apparatus, and electronic device |
| CN110427517B (en) * | 2019-07-18 | 2023-04-25 | 华戎信息产业有限公司 | Picture searching video method and device based on scene dictionary tree and computer readable storage medium |
| CN112348899A (en) * | 2019-08-07 | 2021-02-09 | 虹软科技股份有限公司 | Calibration parameter obtaining method and device, processor and electronic equipment |
| CN110827361B (en) * | 2019-11-01 | 2023-06-23 | 清华大学 | Camera group calibration method and device based on global calibration frame |
| CN110969097B (en) * | 2019-11-18 | 2023-05-12 | 浙江大华技术股份有限公司 | Monitoring target linkage tracking control method, equipment and storage device |
| CN111311693B (en) * | 2020-03-16 | 2023-11-14 | 威海经济技术开发区天智创新技术研究院 | An online calibration method and system for multi-camera |
| CN114170302B (en) * | 2020-08-20 | 2024-10-29 | 北京达佳互联信息技术有限公司 | Camera external parameter calibration method and device, electronic equipment and storage medium |
| CN111932627B (en) * | 2020-09-15 | 2021-01-05 | 蘑菇车联信息科技有限公司 | Marker drawing method and system |
| CN112365722B (en) * | 2020-09-22 | 2022-09-06 | 浙江大华系统工程有限公司 | Road monitoring area identification method and device, computer equipment and storage medium |
| CN114445583B (en) * | 2020-10-30 | 2025-12-12 | 阿里巴巴集团控股有限公司 | Data processing methods, apparatus, electronic devices and storage media |
| CN113923406B (en) * | 2021-09-29 | 2023-05-12 | 四川警察学院 | Method, device, equipment and storage medium for adjusting video monitoring coverage area |
| CN113794841B (en) * | 2021-11-16 | 2022-02-08 | 浙江原数科技有限公司 | Monitoring equipment adjusting method and equipment based on contact ratio |
| WO2024156022A1 (en) * | 2023-01-24 | 2024-08-02 | Visionary Machines Pty Ltd | Systems and methods for calibrating cameras and camera arrays |
| EP4655939A1 (en) * | 2023-01-24 | 2025-12-03 | Visionary Machines Pty Ltd | Systems and methods for automatically calibrating cameras within camera arrays |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101630406A (en) * | 2008-07-14 | 2010-01-20 | 深圳华为通信技术有限公司 | Camera calibration method and camera calibration device |
| CN101763632A (en) * | 2008-12-26 | 2010-06-30 | 华为技术有限公司 | Method for demarcating camera and device thereof |
| CN101894369A (en) * | 2010-06-30 | 2010-11-24 | 清华大学 | A Real-time Method for Computing Camera Focal Length from Image Sequence |
| CN102110292A (en) * | 2009-12-25 | 2011-06-29 | 新奥特(北京)视频技术有限公司 | Zoom lens calibration method and device in virtual sports |
| CN103024350A (en) * | 2012-11-13 | 2013-04-03 | 清华大学 | Master-slave tracking method for binocular PTZ (Pan-Tilt-Zoom) visual system and system applying same |
| CN103824278A (en) * | 2013-12-10 | 2014-05-28 | 清华大学 | Monitoring camera calibration method and system |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2006084385A1 (en) * | 2005-02-11 | 2006-08-17 | Macdonald Dettwiler & Associates Inc. | 3d imaging system |
| US9191650B2 (en) * | 2011-06-20 | 2015-11-17 | National Chiao Tung University | Video object localization method using multiple cameras |
| US8803943B2 (en) * | 2011-09-21 | 2014-08-12 | National Applied Research Laboratories | Formation apparatus using digital image correlation |
-
2013
- 2013-12-10 CN CN201310670982.0A patent/CN103824278B/en not_active Expired - Fee Related
-
2014
- 2014-07-30 US US14/765,748 patent/US9684962B2/en not_active Expired - Fee Related
- 2014-07-30 WO PCT/CN2014/083329 patent/WO2015085779A1/en not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101630406A (en) * | 2008-07-14 | 2010-01-20 | 深圳华为通信技术有限公司 | Camera calibration method and camera calibration device |
| CN101763632A (en) * | 2008-12-26 | 2010-06-30 | 华为技术有限公司 | Method for demarcating camera and device thereof |
| CN102110292A (en) * | 2009-12-25 | 2011-06-29 | 新奥特(北京)视频技术有限公司 | Zoom lens calibration method and device in virtual sports |
| CN101894369A (en) * | 2010-06-30 | 2010-11-24 | 清华大学 | A Real-time Method for Computing Camera Focal Length from Image Sequence |
| CN103024350A (en) * | 2012-11-13 | 2013-04-03 | 清华大学 | Master-slave tracking method for binocular PTZ (Pan-Tilt-Zoom) visual system and system applying same |
| CN103824278A (en) * | 2013-12-10 | 2014-05-28 | 清华大学 | Monitoring camera calibration method and system |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112639883A (en) * | 2020-03-17 | 2021-04-09 | 华为技术有限公司 | Relative attitude calibration method and related device |
| CN112639883B (en) * | 2020-03-17 | 2021-11-19 | 华为技术有限公司 | Relative attitude calibration method and related device |
| CN113298885A (en) * | 2021-06-23 | 2021-08-24 | Oppo广东移动通信有限公司 | Binocular calibration method and device, equipment and storage medium |
| CN115733949A (en) * | 2021-08-25 | 2023-03-03 | 浙江宇视科技有限公司 | A method and device for creating a three-dimensional field of view model of a shooting device |
| KR102943303B1 (en) | 2025-10-13 | 2026-03-24 | 오픈인 주식회사 | Multi-camera based 3D shape measurement device with PTZ control |
Also Published As
| Publication number | Publication date |
|---|---|
| CN103824278B (en) | 2016-09-21 |
| CN103824278A (en) | 2014-05-28 |
| US20150371385A1 (en) | 2015-12-24 |
| US9684962B2 (en) | 2017-06-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9684962B2 (en) | Method and system for calibrating surveillance cameras | |
| Chen et al. | Indoor camera pose estimation via style‐transfer 3D models | |
| CN103026385B (en) | Method, apparatus and computer program product for providing object tracking using template switching and feature adaptation | |
| CN104778675B (en) | A kind of coal mining fully mechanized workface dynamic video image fusion method | |
| Zhang et al. | Improved feature point extraction method of ORB-SLAM2 dense map | |
| TW202244680A (en) | Pose acquisition method, electronic equipment and storage medium | |
| CN118840427B (en) | BIM-based model labeling and live-action point three-dimensional positioning method and device | |
| CN111161334B (en) | A Semantic Map Construction Method Based on Deep Learning | |
| Ding et al. | Stereo vision SLAM-based 3D reconstruction on UAV development platforms | |
| CN112652020A (en) | Visual SLAM method based on AdaLAM algorithm | |
| CN111951158B (en) | Unmanned aerial vehicle aerial image splicing interruption recovery method, device and storage medium | |
| Gao et al. | Multi-source data-based 3D digital preservation of largescale ancient chinese architecture: A case report | |
| CN111829522B (en) | Instant positioning and map construction method, computer equipment and device | |
| Cai et al. | Visual Expansion and Real‐Time Calibration for Pan‐Tilt‐Zoom Cameras Assisted by Panoramic Models | |
| CN117522963A (en) | Checkerboard corner point positioning method, device, storage medium and electronic equipment | |
| CN116524382A (en) | Bridge swivel closure accuracy inspection method system and equipment | |
| CN116884065A (en) | Face recognition method and computer-readable storage medium | |
| Yuan et al. | Object-Based Semantic Fusion Algorithm of Lidar and Camera via Inverse Projection | |
| Geng et al. | Gaze control system for tracking Quasi-1D high-speed moving object in complex background | |
| CN116071396B (en) | A synchronous positioning method | |
| CN116597001B (en) | Indoor ceiling boundary position detection method, device, robot and storage medium | |
| CN114608558B (en) | SLAM method, system, equipment and storage medium based on feature matching network | |
| CN114155147B (en) | A wide-area ocean panoramic fusion method and system based on multi-source data | |
| Li et al. | Efficient and precise visual location estimation by effective priority matching-based pose verification in edge-cloud collaborative IoT | |
| Huang | Research on binocular vision ranging based on YOLO algorithm and stereo matching algorithm |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 14868855 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 14765748 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 14868855 Country of ref document: EP Kind code of ref document: A1 |