WO2024239207A1 - Three-dimensional object distance estimation - Google Patents
Three-dimensional object distance estimation Download PDFInfo
- Publication number
- WO2024239207A1 WO2024239207A1 PCT/CN2023/095578 CN2023095578W WO2024239207A1 WO 2024239207 A1 WO2024239207 A1 WO 2024239207A1 CN 2023095578 W CN2023095578 W CN 2023095578W WO 2024239207 A1 WO2024239207 A1 WO 2024239207A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- distance
- perspective transformation
- location
- target
- transformed
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/70—Determining position or orientation of objects or cameras
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30232—Surveillance
Definitions
- Various example embodiments of the present disclosure generally relate to the field of telecommunication and in particular, to methods, devices, apparatuses and computer readable storage medium for three-dimensional (3D) object distance estimation.
- 3D distances between objects in an environment For example, to avoid a hazard event (e.g., a collision) between two objects, it needs to estimate a distance between the two objects and determine whether the distance is too low (lower than a threshold) .
- a hazard event e.g., a collision
- artificial intelligence technology and imaging control technology are developing rapidly, and a large number of models have been implemented with the model architecture based on various neural networks, such as target detection models, image generation models, and the like.
- the main challenge is how to accurately restore the 3D scene.
- an apparatus comprising at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: detect a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; determine a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; transform, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and determine a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- an apparatus comprising means for detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; means for determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; means for transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and means for determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- a computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the second aspect.
- FIG. 1 illustrates an example environment in which example embodiments of the present disclosure can be implemented
- FIG. 3 illustrates a flowchart of a process for model training and inference for model-based object distance estimation in accordance with some example embodiments of the present disclosure
- FIG. 4 illustrates a flowchart of a process for training the object detection model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure
- FIG. 5 illustrates a flowchart of a process for training the perspective transformation model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure
- FIG. 6A and FIG. 6B illustrate schematic diagrams of example image transformation and normalized object distances in a 3D top plane view in accordance with some example embodiments of the present disclosure
- FIG. 7 illustrates a processing flow for training the perspective transformation model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure
- FIG. 8 illustrates a schematic diagram of floating key points in the model training in accordance with some example embodiments of the present disclosure
- FIG. 9 illustrates a flowchart of a process for object distance estimation and abnormal distance detection in an inference stage in accordance with some example embodiments of the present disclosure
- FIG. 10 illustrates a simplified block diagram of a device that is suitable for implementing example embodiments of the present disclosure.
- FIG. 11 illustrates a block diagram of an example computer readable medium in accordance with some example embodiments of the present disclosure.
- references in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
- first, ” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments.
- the term “and/or” includes any and all combinations of one or more of the listed terms.
- performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.
- circuitry may refer to one or more or all of the following:
- circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
- Communication in the present disclosure can follow any suitable type of communication standards, including cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof.
- cellular communication Wi-Fi communication
- LPWAN communication data communication
- voice communication multimedia communication
- short-range communications such as Bluetooth
- near-field communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof.
- GPS global positioning system
- Communications may be implemented in according to one or more communication protocols and/or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA) , Wideband Code Division Multiple Access (WCDMA) , GSM, LTE, New Radio (NR) , UMTS, WiMax, Ethernet, transmission control protocol/internet protocol (TCP/IP) , synchronous optical networking (SONET) , Asynchronous Transfer Mode (ATM) , QUIC, Hypertext Transfer Protocol (HTTP) , and so forth.
- CDMA Code Division Multiplexing Access
- WCDMA Wideband Code Division Multiple Access
- WCDMA Wideband Code Division Multiple Access
- GSM Global System for Mobile communications
- LTE Long Term Evolution
- NR New Radio
- UMTS Universal Mobile communications
- WiMax Ethernet
- TCP/IP transmission control protocol/internet protocol
- SONET synchronous optical networking
- ATM Asynchronous Transfer Mode
- QUIC Hypertext Transfer Protocol
- HTTP Hypertext Transfer Protocol
- model is referred to as an association between an input and an output learned from training data, and thus a corresponding output may be generated for a given input after the training.
- the generation of the model may be based on a machine learning technique.
- the machine learning techniques may also be referred to as artificial intelligence (AI) techniques.
- AI artificial intelligence
- a machine learning model can be built, which receives input information and makes predictions based on the input information.
- a classification model may predict a class of the input information among a predetermined set of classes.
- model may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network, ” which are used interchangeably herein.
- machine learning may usually involve three stages, i.e., a training stage, a validation stage, and an inference stage (also referred to as an inference stage) .
- a given machine learning model may be trained (or optimized) iteratively using a great amount of training data until the model can obtain, from the training data, consistent inference similar to those that human intelligence can make.
- a set of parameter values of the model is iteratively updated until a training objective is reached.
- the machine learning model may be regarded as being capable of learning the association between the input and the output (also referred to an input-output mapping) from the training data.
- a validation input is applied to the trained machine learning model to test whether the model can provide a correct output, so as to determine the performance of the model.
- the validation stage may be considered as a step in a training process, or sometimes may be omitted.
- the resulting machine learning model may be used to process a real-world model input based on the set of parameter values obtained from the training process and to determine the corresponding model output.
- 3D distance estimation there is a disagreement in whether the scene can be accurately restored to a 3D space by single-view two-dimensional (2D) images.
- a majority of works aim at accurate 3D scene reduction by adding more information, including, for example, Light Detection And Ranging () sensing data, images captured by multi-perspective camera, and the like.
- some solutions propose to obtain the multimodal data of 3D scene through the LIDAR-based sensor devices or imaging devices, combine the multimodal data to perform data fusion, and then generate a complete 3D scene for object distance estimation.
- Some other solution propose to directly estimate uncertain 3D objects by two-dimensional image data through a trained model, where the estimate is made by combining multi-perspective cameras with conversion of the 3D scene.
- those solutions often significantly increase the design and inference cost of the model as well as high costs of hardware device for data collection.
- a solution for 3D object distance estimation from 2D images is based on ML models.
- a perspective transformation model is trained for a deployment scene of a camera and is used to determine a perspective transformation and a distance scale for a specific image captured by the camera in the deployment scene.
- the perspective transformation is used to transform a first location of a first object and a second location of a second object within a 2D space corresponding to a target image into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a 3D space, respectively.
- a target distance between the first object and the second object in a three-dimensional space is then determined by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- the present solution can effectively utilize the 2D image and greatly reduce the modeling and usage costs for achieving 2D-to-3D scene restoration by increasing the information perception dimension (multi-vision camera, LIDAR, etc. ) in the application field. Further, the present solution does not require distance metric calculation on the basis of 3D model, and ca estimates the plane distance of 2D object in the 3D space directly by model reduction, which greatly reduces the technical development threshold and usage cost in the application field, and enhances the convenience in different application scenarios. As such, the present solution can be easily adopted in light application scenarios.
- FIG. 1 illustrates an example environment 100 in which example embodiments of the present disclosure can be implemented.
- a camera is mounted in a physical environment to capture images.
- a camera 102-1 is mounted with a specific deployment in a scene 101-1 (represented as “Scene 1” )
- a camera 102-N is mounted with a specific deployment in a scene 101-N (represented as “Scene N” )
- N may be an integer larger than or equal to one.
- the cameras 102-1, ..., 102-N may be the same or different types of cameras, e.g., with different imaging parameters.
- the deployments of the cameras 102-1, ..., 102-N may be the same or different although the scenes may be different.
- the cameras 102-1, ..., 102-N may be monocular images, capturing monochrome images or colored images (such as RGB images) .
- the cameras 102-1, ..., 102-N may be collectively or individually referred to as cameras 102, and the scenes 101-1, ..., 101-N may be e collectively or individually referred to as scenes 101.
- the images captured by the cameras 102 may be transmitted to a device (s) /system (s) for object distance estimation and in some cases, for abnormal distance detection.
- the images may be transmitted via various networking technology.
- the camera 102-1 may connect to a radio access network (RAN) and thus can transmit its images to a network device 103 in the RAN.
- the network device 103 may convey the images to a core transmission network 110, e.g., via a switch 104 and an access router 105.
- the images may be routed to an object distance estimation device in an information management monitoring room 120 which can access to the core transmission network 110 via an access router 122.
- the camera 102-N may have a network connection in a local access network (LAN) and thus can transmit its images to the core transmission network 110 via a switch 106 in the LAN.
- the images captured by the camera 102-N may also be conveyed to an object distance estimation device in the information management monitoring room 120.
- the devices in the information management monitoring room 120 may subscribe to remote computing resources from a computing center 130 for processing the images and/or other processing operations required in the object distance estimation and the abnormal distance detection.
- the computing center 130 may be accessed through the core transmission network 110 via an access router 132.
- the computing center 130 may contain one or more computing resource clusters to provide the processing, memory, storage, and networking resources for the information management monitoring room 120.
- the computing center 130 may be implemented as a cloud platform.
- images captured by a camera 102 may be processed remotely (e.g., in the information management monitoring room 120 and/or in the computing center 130) for the object distance estimation and abnormal distance detection, and in some other embodiments, the images captured by the camera 102 may be processed locally, e.g., by a device deployed locally in the scene and communicatively coupled with the camera 102.
- the scope of the present disclosure is not limited in this regard.
- the example embodiments of the present disclosure propose a solution to derive 3D distance information from an 2D image (e.g., an image captured by a monocular camera) using a perspective transformation model.
- the perspective transformation model is trained for a target deployment of a target camera, and is used to generate a perspective transformation and a distance scale for a target image captured by the target camera.
- an object detection model may be trained to detect the objects from the target image. As the models are introduced, there may be a training stage for the models and then an inference stage to apply the trained models for the distance estimation.
- FIG. 2 illustrates a framework 200 for model-based object distance estimation in accordance with some example embodiments of the present disclosure.
- the framework 200 comprises a model training system 210 in a training stage and a model inference system 220 in an inference stage.
- the model training system 210 is configured to train one or more perspective transformation models 204-1, ..., 204-N for one or more scenes (Scene 1, ..., Scene N) respectively.
- the perspective transformation models 204-1, ..., 204-N may be collectively or individually referred to as perspective transformation models 204.
- the model training system 210 may be further configured to train an object detection model 202 to detect an object (s) in an image.
- the model inference system 220 is configured to determine a distance between objects detected from a target image 232 using the trained object detection model 202 and the trained perspective transformation model 204.
- the model training system 210 and the model inference system 220 may be the same or different systems.
- the model training system 210 may be implemented at the computing center 130, and the model inference system 220 may be implemented at one or more devices in the information management monitoring room 120.
- both the model training system 210 and the model inference system 220 may be implemented as a same system/device in the computing center 130 and/or in the information management monitoring room 120.
- FIG. 3 illustrates a flowchart of a process 300 for model training and inference for model-based object distance estimation in accordance with some example embodiments of the present disclosure.
- the process 300 will be described with reference to FIG. 2.
- Some training-related operations in the process 300 may be performed by the model training system 210, and some inference-related operations in the process 300 may be performed by the model inference system 220.
- the model training system 210 collects first sample images 212 and first label information 213 indicating objects detected in the first sample images 212.
- the model training system 210 trains the object detection model 202 with the first sample images 212 and first label information 213.
- the input to the object detection model 202 may be an image, and the output from the object detection model 202 may be a location (s) of an object (s) in the input image.
- the object detection model 202 may provide a bounding box to locate an object in the image.
- the training data for the object detection model 202 may comprise sample images 212 (sometimes referred to as first sample images 212) and first label information 213 indicating objects detected from the first sample images 212.
- the object detection model 202 is used to detect one or more objects of interest from an image (the distance (s) of which are to be estimated) .
- the object detection 202 may not be associated with a specific scene or a deployment of a camera in the scene.
- the first sample images 212 may be collected from various sources, for example, by a plurality of cameras with a plurality of deployments in Scene 1, ..., Scene N.
- the training of the object detection model 202 will be discussed in more detail below with reference to FIG. 4.
- the model training system 210 collects second sample images 214-1, ..., 214-N and second label information 215-1, ..., 215-N for the second sample images 214-1, ..., 214-N.
- the model training system 210 trains the perspective transformation model 204-1, ..., 204-N with the second sample images 214-1, ..., 214-N and second label information 215-1, ..., 215-N, respectively.
- the second sample images 214-1, ..., 214-N may be collectively or individually referred to as second sample images 214; and the second label information 215-1, ..., 215-N may be collectively or individually referred to as second label information 215.
- the input to a perspective transformation model 204 may be an image, with one or more objects detected in the image, and the output from the perspective transformation model 204 may be a perspective transformation and a distance scale for the input image.
- the perspective transformation is used to transform the input image (which is in the 2D pixel space) into a top plane view of a 3D space, to obtain a transformed image.
- the distance scale is used to scale distances between pixels (or object locations) in the transformed image into the physical distance in the scene.
- the perspective transformation model 204 may also be referred to as a “perspective transformation and filed of view size” model.
- Each perspective transformation model 204 may be associated with a scene, and it is assumed that a camera is deployed in the scene with a specific deployment.
- the perspective transformation model 204 is trained to adopt to the perspective transformation and field of view scaling for the specific deployment of the camera, and thus can be used for inference of images captured by a camera with the corresponding deployment.
- training data for a perspective transformation model 204-1 may include second sample images 214-1 captured by one or more cameras with a deployment (Deployment 1) in Scene 1
- training data for a perspective transformation model 204-N may include second sample images 214-N captured by one or more cameras with a different deployment (Deployment N) in Scene N.
- the second sample images 214-1, ..., 214-N used for training the perspective transformation models may be divided from the first sample images 212 according to the scenes or deployments.
- a deployment of a camera may be identified by a mounting height, a shooting orientation, and a type of the camera in the scene. For example, for the same type of image, if it is deployed at different heights and/or with different orientations, then the modelling for 2D-to-3D scene restoration may be different. For different types of cameras, it may have different focal lengths, and then the modelling for 2D-to-3D scene restoration may be different. Then different perspective transformation models 204 may need to be trained for the different deployments of cameras.
- the training data for a perspective transformation model 204 may further include second label information 215 for the second sample images 214.
- the second label information 215 may indicate a labeled indication of whether a distance between a pair of objects detected in the second sample image is abnormal.
- the second label information may further indicate labeled transformed locations of the pair of objects in the top plane view.
- the model inference system 220 determines a target distance between a first object and a second object in a target image 232 using the object detection model 202 and a perspective transformation model 204.
- the model inference system 220 may comprise a distance estimator 230 to determine the target distance between the first and second objects using the outputs from the object detection model 202 and the perspective transformation model 204.
- the perspective transformation model 204 applied for the target image 232 may depend on a target deployment of a target camera which captures the target image 232. It is assumed that the target image 232 is captured by a target camera with a target deployment in Scene X. Among all the perspective transformation models 204, the perspective transformation model 204-X associated with the target deployment in Scene X may be apply by the model inference system 220 for determining the distance.
- the object detection model 202 is used to detect two or more objects from the target image 232, to determine locations of the two or more objects in the target image 232.
- the locations may represent location information of the objects within a 2D space (also referred to as a 2D pixel space) corresponding to the target image 232.
- the perspective transformation model 204-X is used to transform the location of a first object and the location of a second object within the 2D space into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a 2D space, respectively.
- the transformed locations are considered as 3D locations for the first and second objects.
- the model inference system 220 may perform abnormal distance detection based on the target distance between the first object and the second object.
- the model inference system 220 may comprise an distance anomaly detector 235 to perform the abnormal distance detection.
- the distance anomaly detector 235 may determine an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
- the present solution of the example embodiments disclosure herein can effectively improve the efficiency of 2D imaging to 3D distance detection, greatly reduce the hardware cost and model training and inference costs for scenarios such as in-site hazard distance detection, while achieving the expected 3D distance detection in multiple crossover application areas, including but not limited to pedestrian distance detection, construction vehicle safety distance detection, hazardous object safety distance detection, and the like.
- FIG. 4 illustrates a flowchart of a process 400 for training the object detection model 202 in the framework 200 of FIG. 2 in accordance with some example embodiments of the present disclosure.
- the process 400 will be described from the perspective of the model training system 210 in FIG. 2.
- the model training system 210 obtains first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images, for example, the first sample images 212 and the first label information 213 as shown in FIG. 2.
- the first sample images 212 may be collected from different dataset sources (for example, captured by different cameras with different deployments in the scenes) .
- the cameras used for capturing the first sample images 212 may be monocular cameras.
- a first sample image 212 may be stored as (scene ID, img_x) , where the scene ID is used to identify a deployment of a camera in the scene, and one scene or deployment may correspond to one camera.
- the differential data collection for multiple environmental states of the scenes may be considered, including but not limited to, different weather conditions (e.g., sunny, cloudy, rainy, etc. ) , different time conditions (e.g., early morning, midday, evening, etc. ) , and the like. If a specific deployment of a camera or the corresponding scene is a special deployment or a special scene, the impact of the environmental changes in that scene on the collection of the sample images may need to be taken into consideration.
- different weather conditions e.g., sunny, cloudy, rainy, etc.
- time conditions e.g., early morning, midday, evening, etc.
- the first sample images 212 are used to perform the training of the object detection model 202 which is primarily configured to localize objects in 2D pixel space.
- the first sample images 212 may be labeled with the objects occurred therein.
- the first sample images 212 may be labeled with bounding boxes of the objects to be detected, respectively.
- the labelling may be performed through manual annotations or through any suitable automation tools, to generate the first label information 213 for the first sample images 212.
- the first label information 213 may indicate a bounding box of the area of the object in the first sample images 212.
- the first label information 213 may include coordinates of the detected object in the first sample images 212, such as the four vertexes of the bounding box, , including the left, top, right, bottom vertexes (in this context, denotes l as left, t as top, r as right, b as bottom) for the pixel coordinate points.
- the first sample images 212 containing the specific types of objects of interest may be sampled from all the available sample images. For example, if it is expected to determine an abnormal distance between a forklift and a pedestrian in a scene monitored by a camera, then the first sample images 212 containing a forklift and one or more pedestrians may be sampled from all the images captured by the camera. It would be appreciated that the object detection model 202 may be trained to detect various types of objects at least including some types of objects whose distances are to be estimated, and in this case, more objects may be labeled in the first sample images 212.
- the model training system 210 constructs the object detection model 202.
- the object detection model 202 may be constructed using a generic detection model framework in computer vision. The scope of the present disclosure is not limited in the specific model structure for the object detection model.
- the model training system 210 trains the object detection model 202 with the first sample images 212 and the first label information 213.
- the model training system 210 may iteratively optimize the object detection model 202 for completing the object detection task.
- the model training system 210 may apply any suitable model training algorithms and strategies to complete the training of the object detection model 202.
- some hyperparameters may be set for the training of the object detection model 202.
- a hyperparameter confidence_threshold may be set to a predetermined value (in a range between 0 and 1) for thresholding the selection of detected pre-selected bounding boxes of objects.
- Another hyperparameter NMS_IOU_threshold may be typically set to a predetermined value between 0 and 1 for controlling the screening of the pre-selected bounding boxes with high overlapping.
- the iterative optimization may be completed if the F1-score of the object detection model 202 meets a training objective, e.g., the F1-score being close to 1 on a training dataset and the generalization of the object detection model 202 can prevent from overfitting on a validate dataset.
- a training objective e.g., the F1-score being close to 1 on a training dataset and the generalization of the object detection model 202 can prevent from overfitting on a validate dataset.
- the first sample images 212 used for training the object detection model 202 may be further inspected, to guarantee the accuracy of the next model construction, and removing the missed and over detected RGB images from the sample data.
- the first sample images 212 may be evaluated to determine which objects are missed or mis-checked by the trained object detection model 202. If it is obvious that the object detection model 202 has not yet trained to converge, re-training of the object detection model 202 may be performed with the first sample images 212 whose objects are missed. During the re-training, the gradient weights for those first sample images 212 may be enhanced.
- the model training may be iterated and the object detection model whose F1-score is at or near 1 on the training dataset and which has a strong generalization in the validation dataset, may be selected as the trained model for inference.
- FIG. 5 illustrates a flowchart of a process 500 for training the perspective transformation model (s) 204 in the framework 200 of FIG. 2 in accordance with some example embodiments of the present disclosure.
- the process 600 will be described from the perspective of the model training system 210 in FIG. 2.
- the training data for a perspective transformation model 204 associated with a specific deployment of a scene may include second sample images 214 and second label information 215 for the second sample images 214.
- the training of one perspective transformation model 204 is discussed below and all the other perspective transformation models for other scenes may be trained in similar ways.
- the model training system 210 obtains the second sample images 214 with objects detected.
- the second sample images 214 used for training a perspective transformation model 204 for a specific scene may include a subset of the first sample images 212 used for training the object detection model 202 which is collected from the specific scene.
- the first image samples 212 may be grouped according to the scene IDs, and the second sample images 214 may be selected from each scene (or deployment of the camera) .
- a part or all of the second sample images 214 may be different from those sample images used for training the object detection model 202. Then the objects in the second sample images 214 may be separately labeled from the labeling of the first sample images 212. In some example embodiments, if the object detection model 202 has been trained, the objects in the second sample images 214 may be detected using the trained object detection model 202.
- the objects labeled in the second sample images 214 may be the types of objects whose distance (s) are to be estimated or the abnormal distance (s) are to be detected in the corresponding scene. For example, in one scene, it is expected to estimate a distance between a forklift and a pedestrian, and then these two types of objects are labeled in the second sample images 214 captured from this scene. For another scene, it is expected to estimate a distance between two cars, and then a type of car object is labeled in the second sample images 214 captured from the other scene.
- the model training system 210 determines a fitted perspective transformation based on the second sample images 214 with the objects detected.
- the model training system 210 determines second label information 215 for the second sample images 214 based at least on the fitted perspective transformation.
- the second label information 215 may indicate a labeled indication of whether a distance between a pair of objects detected in the second sample image is abnormal.
- the second label information may further indicate labeled transformed locations of the pair of objects in the top plane view.
- the distances between one or more pairs of objects are also labeled for anomalies.
- the annotators may evaluate whether the distance between a pair of objects in a second sample image 214 is abnormal, e.g., according to some reference objects or reference distance in the second sample image 214.
- the second sample image 214 and its second label information 215 may be stored as (scene ID, img_x, an indication of whether the distance is abnormal or not) .
- a perspective transformation may be fitted from the labeled locations of the objects in the second sample images 214, for training of the perspective transformation model 204.
- the perspective transformation model 204 is configured to output a perspective transformation and a distance scale for a specific deployment in a scene.
- the perspective transformation is used to transform a 2D location of an object in the 2D pixel space into a 3D location in a 3D space, or more specifically, a planar location in a top plane view of the 3D space.
- the distance scale is used to scale a distance between transformed location in the top plane view into the actual physical distance in the 3D space.
- a perspective transformation (also referred to as a perspective transformation matrix) may be represented as the distance scale may be represented as Scaler.
- the M perspective may be flattened, for example, as a (1, 9) matrix, and the distance scale Scaler is a (1, 1) matrix.
- pts co-pix represents the location in the image (e.g., the pixel coordinates in the image)
- pts co-bev represents a transformed location (e.g., the 2D top plane coordinates) in the 3D top plane view.
- the transformation may allow the 2D top plane coordinates obeying three degrees of freedom.
- the distance scaler Scaler is introduced, and then the physical distance between two objects (object i and j) in an image may be calculated as follows:
- ⁇ l2 represents the L2-distance calculated from the 2D top plane coordinates of the two objects in the tope plane view.
- FIG. 6A and FIG. 6B illustrate schematic diagrams of example image transformation and normalized object distances in a 3D top plane view.
- a captured image 610 or 640 which may be a 2D image captured by a monocular camera, may be transformed through a corresponding perspective transformation into a transformed image 620 or 650 which is considered as a view of the 3D scene from the tope plane view.
- the locations of objects 611, 612, and 613 in the image 610 may be mapped to respective transformed locations 621, 622, and 623 of the corresponding objects within the transformed image 620.
- the transformed locations 621, 622, and 623 of the objects may be represented in a normalized top plane view 630.
- the locations of objects 641, 642, and 643 in the image 640 may be mapped to respective transformed locations 651, 652, and 653 of the corresponding objects within the transformed image 650.
- the transformed locations 651, 652, and 653 of the objects may be represented in a normalized top plane view 660.
- the training of the perspective transformation model 204 is to learn an accurate perspective transformation as well as an accurate distance scale to estimate the 3D distance between two objects.
- a fitted perspective transformation may be determined for a specific deployment in a scene.
- the locations of the objects labeled in the second sample images 214 e.g., coordinates of four vertexes of a bounding box of an object in a second sample image 214) may be selected and used to calculate the fitted perspective transformation of the scene.
- the fitted perspective transformation is used for transforming the locations of the objects in the second sample images 214 into transformed locations in the top plane view of the 3D space.
- the fitting process may be considered as there are errors in the manual evaluation, so a multi-labeling cross-validation approach is considered.
- One of the second sample images 214 collected for the scene i may be randomly selected for construction of the fitted perspective transformation.
- Another second sample image 214 is randomly selected for verifying whether the perspective transformation meets the expectations of the transformation.
- coordinates of four foot points in this image may be recorded as [pts 1 , pts 2 , pts 3 , pts 4 ]
- the coordinates of the 4 foot points after the perspective transformation are defined as [pts′ 1 , pts′ 2 , pts′ 3 , pts′ 4 ]
- the transformed location (s) of one or more object (s) in the second sample image 214 e.g., the coordinates of the four vertexes of the bounding box
- the coordinates of the four foot points after the perspective transformation may be adjusted until they meet the expectations of the perspective transformation. That is, the fitting is repeated until the obtained fitted perspective transformation meets the expectation.
- the fitted perspective transformation of each scene is recorded as (scene ID, img_x, ) .
- the fitted perspective transformation may be combined with the location of the objects detected in the second sample images 214 (e.g., the bounding boxes) .
- the model training system 210 may transform, based on the fitted perspective transformation, the respective locations of the objects into respective transformed locations within the top plane view of the 3D space.
- a location of an object detected from an image may be defined as a plurality of coordinates of a first plurality of key points of the object detected from the image.
- the key points for an object may include the four vertexes of a bounding box of the object.
- the key points may further include a centroid of the object.
- a location of an object in a second sample image 214 may be represented as coordinates of the four vertexes (left, top, right, and bottom) of a bounding box of the object within the second sample image 214, which may be represented as (Co left , Co top , Co right , Co bottom ) .
- a centroid of the object may also be taken into account to represent the location of the object.
- a key point transformation mapping may be determined for a second sample image 214 based on the fitted perspective transformation.
- the location of an object detected in a second sample image 214 may thus be represented as (Co left , Co top , Co right , Co bottom , (Co c_x , Co c_y ) ) , i.e., five coordinates in the 2D space corresponding to the second sample image 214.
- the 2D coordinates of the location of the object may be filled with a predetermined value (e.g., 1.0) for the third dimensional, to transform to Then Kpt may be transformed based on the fitted perspective transformation, to obtain a transformed location Kpt′in the top plane view.
- the transformed location may be
- a second sample image 214 and its second label information 215 may be constructed as a training sample for the perspective transformation model 204 as ⁇ image_x, Kpt, Kpt′, i, j, p/n ⁇ , where i and j represent a pair of objects in the second sample image 214 whose distance is to be estimated, the indicator “p” or “n” indicates whether the distance between the objects i and j is abnormal or not, Kpt indicates locations of the objects within the 2D space corresponding to the second sample image 214, and Kpt′ indicates transformed locations of the objects in the 3D top plane view (which may be used to calculate the labeled distance between a pair of objects in the 3D top plane view) .
- the objects i, j may be selected as the pair of object with the lowest transformed centroid distance, e.g., if the abnormal distance is determined as a distance lower than a distance threshold. For example, if a forklift and two pedestrians are detected in a second sample image 214 (as in the examples of FIG. 6A and FIG, 6B) , then a pair of forklift and a pedestrian with the lowest distance may be selected for labeling the second sample image 214. In the case where an abnormal distance is determined as a distance exceeding a distance threshold, the pair of object with the highest transformed centroid distance may be selected.
- the second sample image 214 may be labeled with the second label information associated with more than one pair of objects whose distance (s) is to be estimated.
- the real object size in the images corresponding to the aspect ratio of the sample images may have a large difference. Therefore, in the model training phase, it is considered to train the perspective transformation model 204 by varying the scale of the second sample image 214.
- the resolution of a second sample image is represented as (h, w) .
- the model training system 210 constructs the perspective transformation model 204.
- the perspective transformation model 204 may be constructed with any suitable model structure in the computer vision (e.g., with MobileNet_v2 as the backbone model) .
- the modeling of the perspective transformation model 204 may be represented as with an image (represented as “img” ) as its input, and output a perspective transformation and a distance scale Scaler for the input image.
- the model training system 210 trains the perspective transformation model 204 with the second sample images 214 and the second label information 215.
- a training sample may be constructed as follows: ⁇ image_x, Kpt, Kpt′, i, j, p/n ⁇ , image_x represents a second sample image 214, and (Kpt, Kpt′, i, j, p/n) represents the second label information 215.
- the training of the perspective transformation model 204 may be based on a first loss function (also referred to as an objective function) , which is used to evaluate a first error between a predicted indication of distance anomaly and a labeled indication of distance anomaly for a second sample image (s) 214.
- the first loss function may be represented as follows:
- the training objective is to decrease or minimize a loss value of the first loss function by iteratively updating or optimizing the perspective transformation model 204 (e.g., updating the model parameter values) .
- the calculation of a loss value of the first loss function is as follows. For each second sample image 214, the model training system 210 may determine a predicted perspective transformation output by the perspective transformation model 204 for the second sample image 214 (based on the current model parameter values) , and then determine respective predicted transformed locations of a pair of objects in the top plane view of the three-dimensional space using the predicted perspective transformation. A predicted transformed location of an object i is determined as is determined based on the location (coordinates) of the object i within the second sample image 214.
- the model training system 210 may further determine a predicted distance between the pair of objects (e.g., objects i and j) by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image.
- the predicted distance may be determined as as in Equation (3) .
- the model training system 210 may determine a predicted indication of whether a distance between the pair of objects is abnormal based on a comparison result between the predicted distance and a distance threshold. For example, if a distance lower than the distance threshold is determined as an abnormal distance, the predicted distance lower than the distance threshold may be indicated as an abnormal distance. Otherwise, if a distance exceeding a distance threshold is determined as an abnormal distance, the predicted distance exceeding the distance threshold may be indicated as an abnormal distance.
- the parameter a margin in Equation (3) represents a distance threshold, and in this example, a distance lower than the distance threshold is determined as an abnormal distance.
- the loss value of the first loss function is lower or decreased as the predicted indication approximates to the labeled indication, which means that the perspective transformation model 204 has learned to predict relatively accurate perspective transformation and distance scale to make the correct anormal distance detection. Otherwise, the loss value of the first loss function is high or increased if the predicted indication is different from the labeled indication, which means that the predicted perspective transformation and the distance scale do not meet the expectation.
- the perspective transformation model 204 may need to be further optimized.
- the training of the perspective transformation model 204 may be further based on a second loss function, which is used to evaluate a second error between a predicted transformed location (s) of an object (s) in a second sample image 214 and a labeled transformed location (s) of the object (s) in the second sample image 214.
- the second loss function may be represented as follows:
- the training objective is to decrease or minimize a loss value of the second loss function by iteratively updating or optimizing the perspective transformation model 204 (e.g., updating the model parameter values) .
- the calculation of a loss value of the first loss function is as follows.
- Equation (3) for each second sample image 214, a predicted transformed location of an object i is determined as as described above.
- Kpt′ i indicates a labeled transformed location of the object i . represents the difference (or error) between the predicted transformed location and the labeled transformed location.
- the predicted transformed locations and the labeled transformed locations for more objects detected in the second sample image 214 may be determined, to calculate the loss value of the second loss function L kpt .
- the weight a may be initialized (e.g., to 0.5, 0.8, or the like) and may be adjusted or gradually reduced during the training (e.g., gradually reduced by 0.01 after a number of epochs of iterations) .
- one or more additional or alternative loss functions may be designed for training the perspective transformation model 204 based on the second label information 215 for the second sample images 214.
- the training of the perspective transformation model 204 may reference to the architecture 700 in FIG. 7.
- the labeled transformed locations of the objects e.g., Kpt′ i and Kpt′ j may be determined by mapping Kpt i and Kpt j with the fitted perspective transformation.
- the perspective transformation model 204 is trained, and the predicted transformed locations of the objects may be determined and used to calculate the loss values of the first and/or second loss function.
- the model training stage involves segmented training process (the object detection model 202 and the perspective transformation model 204) , considering the order of model optimization, the optimization accuracy of the bounding box of the object directly affects the accuracy of the subsequent calculation. Therefore, to ensure the accuracy of the perspective transformation models 204, in the process of model training, the generalization and robustness performance of the perspective model 204 may be further enhanced by randomly changing the locations for objects in the sample images.
- FIG. 8 illustrates an example 800 of floating key points in the model training.
- the locations for objects 811 and 812, Kpt m and Kpt n , in a sample image may be transformed into transformed locations Kpt′ m and Kpt′ n 812 may be adjusted to and and the transformed locations may also be changed as and
- FIG. 9 illustrates a flowchart of a process 900 for object distance estimation and abnormal distance detection in an inference stage in accordance with some example embodiments of the present disclosure.
- the process 900 will be described from the perspective of the model inference system 220 in FIG. 2.
- the model inference system 220 detects a first location of a first object and a second location of a second object within a 2D space corresponding to a target image 232 captured by a camera with a target deployment in a scene.
- the model inference system 220 may detect the first location and the second location using the trained object detection model 202.
- the first object and the second object may be objects whose distance therebetween is to be estimated, e.g., for the anomaly detection or for other purposes.
- the first object or the second object may be detected in the target image 232 by respective bounding box.
- the first location of the first object may include a first plurality of coordinates of a first plurality of key points of the first object detected from the target image 232
- the second location may include a second plurality of coordinates of a second plurality of key points of the second object detected from the target image.
- the first plurality of key points may at least include a first centroid of the first object
- the second plurality of key points may at least include a second centroid of the second object.
- the key points for an object may include coordinates of the four vertexes of a bounding box of the object and the centroid of the object, which may be represented as
- the model inference system 220 determines a perspective transformation and a distance scale Scaler for the target image 232 by providing the target image 232 into a trained perspective transformation model 204 associated with the target deployment.
- the perspective transformation model 204 used for the target image 232 may be selected from a plurality of perspective transformation models 204 based on a target deployment identity of the target deployment.
- the plurality of perspective transformation models 204 are each identified with respective deployment identities.
- the deployment identity may sometimes correspond to a scene ID for the scene where the camera is deployed.
- the model inference system 220 transforms, based on the determined perspective transformation the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively.
- a transformed location for an object may be determined as
- the model inference system 220 determines a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale Scaler.
- the first transformed location and the second transformed location may be normalized, though
- the coordinates of the centroids of the first and second objects may be used to determine the distance between two objects.
- the model inference system 220 may further determine an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
- the distance threshold may be configured according to the actual applications of distance anomaly detection. In some example embodiments, if a distance lower than the distance threshold is considered as abnormal, the model inference system 220 may determine whether the target distance is abnormal or not based on a determination of whether the target distance is lower than the distance threshold. In some example embodiments, if a distance exceeding the distance threshold is considered as abnormal, the model inference system 220 may determine whether the target distance is abnormal or not based on a determination of whether the target distance exceeds the distance threshold.
- actions may be determined and performed according to the distance anomaly detection.
- alarms may be generated and transmitted to one or more related devices, and/or operation commands may be sent to the related devices.
- an operation command of applying a force stop may be sent to the vehicle or forklift, and/or an alarming device may be activated to warn the driver and the pedestrian (s) of the dangerous situation in the scene.
- the present disclosure provides solutions for object distance estimation and distance anomaly detection based on monocular image data, which provides a lightweight detection method for video distance anomaly monitoring application scenarios, and effectively reduces the implementation cost and usage cost.
- the models instead of performing a complete 2D-3D scene restoration to achieve the estimation of the distance between objects, the models are trained and applied to transform pixel location in an image into locations in the 3D top plane view, to calculate the distance between the objects is a dangerous distance. This can avoid a lot of tedious information collection and feature construction work in the early stage, and reduce the complexity of the model and its absolute requirements for data accuracy.
- the example embodiments of the present disclosure may not need to strictly define the absolute coordinate positions fitted in the 3D scenes. Due to the full consideration of the fitting of perspective transformation and the fitting of abnormal distance, the distance estimation and the anomaly detection based on the model output are still reliable.
- the present application may be described in the general context of computer-executable instructions executed by a computer, such as a program module or unit.
- a program module or unit may include routines, programs, objects, components, data structures, and the like that perform a particular task to implement a particular abstract data type.
- the program module or unit may be implemented by software, hardware, or a combination of both.
- the present application may also be practiced in distributed computing environments where the program modules or units may be located in local and process computer storage media, including storage devices, or by storage in portable computing devices for abnormally distant data transfer.
- an apparatus capable of performing any of the processes 300, 400, 500, and 900 may comprise means for performing the respective operations of those processes.
- the means may be implemented in any suitable form.
- the means may be implemented in a circuitry or software module.
- the first apparatus may be implemented as or included in the model inference system 220 and/or the model training system 210 in FIG. 2.
- the apparatus comprises means for means for detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; means for determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; means for transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and means for determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- the first apparatus further comprises: means for determining an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
- the means for determining the perspective transformation and the distance scale for the target image comprises: means for selecting the perspective transformation model associated with the target deployment from a plurality of perspective transformation models based on a target deployment identity of the target deployment, the plurality of perspective transformation models being identified with respective deployment identities.
- the target deployment is identified based on a mounting height, a shooting orientation, and a type of the target camera.
- the means for detecting the first location and the second location comprises: means for detecting the first location and the second location using an object detection model, wherein the object detection model is trained with first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images.
- the first location comprises a first plurality of coordinates of a first plurality of key points of the first object detected from the target image
- the second location comprises a second plurality of coordinates of a second plurality of key points of the second object detected from the target image
- the first plurality of key points at least comprises a first centroid of the first object
- the second plurality of key points at least comprises a second centroid of the second object.
- the perspective transformation model is trained using second sample images captured by a camera with the target deployment and second label information indicating at least one of the following for each second sample image: a labeled indication of whether a distance between the pair of objects is abnormal, the pair of objects being detected in the second sample image, or respective labeled transformed locations of the pair of objects in the top plane view of the three-dimensional space.
- the perspective transformation model is trained by the following: determining respective predicted transformed locations of the pair of objects in the top plane view of the three-dimensional space using a predicted perspective transformation output by the perspective transformation model for the second sample image; determining a predicted distance between the pair of objects by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image; determine a predicted indication of whether a distance between the pair of objects is abnormal based on a comparison result between the predicted distance and a distance threshold; determining a first loss value of a first loss function based on a first error between the predicted indication and the labeled indication for the second sample image; and updating the perspective transformation model by decreasing the first loss value of the first loss function.
- the perspective transformation model is further trained by the following: determining a second loss value of a second loss function based on second errors between the respective predicted transformed locations and the respective labeled transformed locations; and updating the perspective transformation model by decreasing the second loss value of the second loss function.
- the respective labeled transformed locations of the pair of objects are determined by the following: determining a fitted perspective transformation associated with the target deployment based on respective locations of the pair of objects within a two-dimensional space corresponding to the second sample image; transforming, based on the fitted perspective transformation, the respective locations of the pair of objects into the respective transformed locations within the top plane view of the three-dimensional space.
- the target camera is a monocular camera.
- the apparatus further comprises means for performing other operations in some example embodiments of the present disclosure.
- the means comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the performance of the apparatus.
- FIG. 10 is a simplified block diagram of a device 1000 that is suitable for implementing example embodiments of the present disclosure.
- the device 1000 may be provided to implement a communication device, for example, the model inference system 220 and/or the model training system 210 as shown in FIG. 2.
- the device 1000 includes one or more processors 1010, one or more memories 1020 coupled to the processor 1010, and one or more communication modules 1040 coupled to the processor 1010.
- the communication module 1040 is for bidirectional communications.
- the communication module 1040 has one or more communication interfaces to facilitate communication with one or more other modules or devices.
- the communication interfaces may represent any interface that is necessary for communication with other network elements.
- the communication module 1040 may include at least one antenna.
- the processor 1010 may be of any type suitable to the local technical network and may include one or more of the following: general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multicore processor architecture, as non-limiting examples.
- the device 1000 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.
- the memory 1020 may include one or more non-volatile memories and one or more volatile memories.
- the non-volatile memories include, but are not limited to, a Read Only Memory (ROM) 1024, an electrically programmable read only memory (EPROM) , a flash memory, a hard disk, a compact disc (CD) , a digital video disk (DVD) , an optical disk, a laser disk, and other magnetic storage and/or optical storage.
- ROM Read Only Memory
- EPROM electrically programmable read only memory
- flash memory a hard disk
- CD compact disc
- DVD digital video disk
- optical disk a laser disk
- RAM random access memory
- a computer program 1030 includes computer executable instructions that are executed by the associated processor 1010.
- the instructions of the program 1030 may include instructions for performing operations/acts of some example embodiments of the present disclosure.
- the program 1030 may be stored in the memory, e.g., the ROM 1024.
- the processor 1010 may perform any suitable actions and processing by loading the program 1030 into the RAM 1022.
- the example embodiments of the present disclosure may be implemented by means of the program 1030 so that the device 1000 may perform any process of the disclosure as discussed above.
- the example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.
- the program 1030 may be tangibly contained in a computer readable medium which may be included in the device 1000 (such as in the memory 1020) or other storage devices that are accessible by the device 1000.
- the device 1000 may load the program 1030 from the computer readable medium to the RAM 1022 for execution.
- the computer readable medium may include any types of non-transitory storage medium, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like.
- the term “non-transitory, ” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM) .
- FIG. 11 shows an example of the computer readable medium 1100 which may be in form of CD, DVD or other optical storage disk.
- the computer readable medium 1100 has the program 1030 stored thereon.
- various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device. Although various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
- Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer readable medium, such as a non-transitory computer readable medium.
- the computer program product includes computer-executable instructions, such as those included in program modules, being executed in a device on a target physical or virtual processor, to carry out any of the methods as described above.
- program modules include routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types.
- the functionality of the program modules may be combined or split between program modules as desired in various embodiments.
- Machine-executable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.
- Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages.
- the program code may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, cause the functions/operations specified in the flowcharts and/or block diagrams to be implemented.
- the program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
- the computer program code or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above.
- Examples of the carrier include a signal, computer readable medium, and the like.
- the computer readable medium may be a computer readable signal medium or a computer readable storage medium.
- a computer readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Image Analysis (AREA)
Abstract
Example embodiments of the present disclosure relate to three-dimensional (3D) object distance estimation. According to a method, a first location of a first object and a second location of a second object are detected within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene. A perspective transformation and a distance scale are determined for the target image by providing the target image into a perspective transformation model associated with the target deployment. Based on the determined perspective transformation, the first and the second locations are transformed into a first transformed location and a second transformed location within a top plane view of a three-dimensional space, respectively. A target distance between the first and the second objects in a three-dimensional space is determined by scaling a distance between the first and the second transformed location with the distance scale.
Description
FIELDS
Various example embodiments of the present disclosure generally relate to the field of telecommunication and in particular, to methods, devices, apparatuses and computer readable storage medium for three-dimensional (3D) object distance estimation.
In various applications, there is a need to estimate 3D distances between objects in an environment. For example, to avoid a hazard event (e.g., a collision) between two objects, it needs to estimate a distance between the two objects and determine whether the distance is too low (lower than a threshold) . At present, artificial intelligence technology and imaging control technology are developing rapidly, and a large number of models have been implemented with the model architecture based on various neural networks, such as target detection models, image generation models, and the like. To estimate 3D object distances, the main challenge is how to accurately restore the 3D scene.
In a first aspect of the present disclosure, there is provided an apparatus. The apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: detect a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; determine a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; transform, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and determine a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
In a second aspect of the present disclosure, there is provided a method. The method comprises: detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
In a third aspect of the present disclosure, there is provided an apparatus. The apparatus comprises means for detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; means for determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; means for transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and means for determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
In a fourth aspect of the present disclosure, there is provided a computer readable medium. The computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the second aspect.
It is to be understood that the Summary section is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become easily comprehensible through the following description.
Some example embodiments will now be described with reference to the accompanying drawings, where:
FIG. 1 illustrates an example environment in which example embodiments of the present disclosure can be implemented;
FIG. 2 illustrates a schematic diagram of a framework for model-based object distance estimation in accordance with some example embodiments of the present disclosure;
FIG. 3 illustrates a flowchart of a process for model training and inference for model-based object distance estimation in accordance with some example embodiments of the present disclosure;
FIG. 4 illustrates a flowchart of a process for training the object detection model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure;
FIG. 5 illustrates a flowchart of a process for training the perspective transformation model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure;
FIG. 6A and FIG. 6B illustrate schematic diagrams of example image transformation and normalized object distances in a 3D top plane view in accordance with some example embodiments of the present disclosure;
FIG. 7 illustrates a processing flow for training the perspective transformation model in the framework of FIG. 2 in accordance with some example embodiments of the present disclosure;
FIG. 8 illustrates a schematic diagram of floating key points in the model training in accordance with some example embodiments of the present disclosure;
FIG. 9 illustrates a flowchart of a process for object distance estimation and abnormal distance detection in an inference stage in accordance with some example embodiments of the present disclosure;
FIG. 10 illustrates a simplified block diagram of a device that is suitable for implementing example embodiments of the present disclosure; and
FIG. 11 illustrates a block diagram of an example computer readable medium in accordance with some example embodiments of the present disclosure.
Throughout the drawings, the same or similar reference numerals represent the same or similar element.
Principle of the present disclosure will now be described with reference to some example embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. Embodiments described herein can be implemented in various manners other than the ones described below.
In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
It shall be understood that although the terms “first, ” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and/or” includes any and all combinations of one or more of the listed terms.
As used herein, “at least one of the following: <a list of two or more elements>”
and “at least one of <a list of two or more elements>” and similar wording, where the list of two or more elements are joined by “and” or “or” , mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
As used herein, unless stated explicitly, performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.
The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and/or “including” , when used herein, specify the presence of stated features, elements, and/or components etc., but do not preclude the presence or addition of one or more other features, elements, components and/or combinations thereof.
As used in this application, the term “circuitry” may refer to one or more or all of the following:
(a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry) and
(b) combinations of hardware circuits and software, such as (as applicable) :
(i) a combination of analog and/or digital hardware circuit (s) with software/firmware and
(ii) any portions of hardware processor (s) with software (including digital signal processor (s) ) , software, and memory (ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and
(c) hardware circuit (s) and or processor (s) , such as a microprocessor (s) or a portion of a microprocessor (s) , that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
This definition of circuitry applies to all uses of this term in this application,
including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Communication in the present disclosure can follow any suitable type of communication standards, including cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof. Communications may be implemented in according to one or more communication protocols and/or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA) , Wideband Code Division Multiple Access (WCDMA) , GSM, LTE, New Radio (NR) , UMTS, WiMax, Ethernet, transmission control protocol/internet protocol (TCP/IP) , synchronous optical networking (SONET) , Asynchronous Transfer Mode (ATM) , QUIC, Hypertext Transfer Protocol (HTTP) , and so forth.
As used herein, the term “model” is referred to as an association between an input and an output learned from training data, and thus a corresponding output may be generated for a given input after the training. The generation of the model may be based on a machine learning technique. The machine learning techniques may also be referred to as artificial intelligence (AI) techniques. In general, a machine learning model can be built, which receives input information and makes predictions based on the input information. For example, a classification model may predict a class of the input information among a predetermined set of classes. As used herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network, ” which are used interchangeably herein.
Generally, machine learning may usually involve three stages, i.e., a training stage, a validation stage, and an inference stage (also referred to as an inference stage) .
At the training stage, a given machine learning model may be trained (or optimized) iteratively using a great amount of training data until the model can obtain, from the training data, consistent inference similar to those that human intelligence can make. During the training, a set of parameter values of the model is iteratively updated until a training objective is reached. Through the training process, the machine learning model may be regarded as being capable of learning the association between the input and the output (also referred to an input-output mapping) from the training data. At the validation stage, a validation input is applied to the trained machine learning model to test whether the model can provide a correct output, so as to determine the performance of the model. Generally, the validation stage may be considered as a step in a training process, or sometimes may be omitted. At the inference stage, the resulting machine learning model may be used to process a real-world model input based on the set of parameter values obtained from the training process and to determine the corresponding model output.
For 3D distance estimation, there is a disagreement in whether the scene can be accurately restored to a 3D space by single-view two-dimensional (2D) images. In this context, a majority of works aim at accurate 3D scene reduction by adding more information, including, for example, Light Detection And Ranging () sensing data, images captured by multi-perspective camera, and the like. For example, some solutions propose to obtain the multimodal data of 3D scene through the LIDAR-based sensor devices or imaging devices, combine the multimodal data to perform data fusion, and then generate a complete 3D scene for object distance estimation. Some other solution propose to directly estimate uncertain 3D objects by two-dimensional image data through a trained model, where the estimate is made by combining multi-perspective cameras with conversion of the 3D scene. However, those solutions often significantly increase the design and inference cost of the model as well as high costs of hardware device for data collection.
In the case where images are collected by monocular cameras, the following data are missing for 2D-to-3D object scene conversion: depth detection data in the images, and the intrinsic and extrinsic information corresponding to the camera. In order to collect the above data, the cost is large or even impossible in the real application scenes. Therefore, it is desired to effectively estimate 3D object distance through low-cost manners.
According to example embodiments of the present disclosure, there is proposed a solution for 3D object distance estimation from 2D images. This solution is based on
ML models. Specifically, a perspective transformation model is trained for a deployment scene of a camera and is used to determine a perspective transformation and a distance scale for a specific image captured by the camera in the deployment scene. For a first object and a second object detected in a target image, the perspective transformation is used to transform a first location of a first object and a second location of a second object within a 2D space corresponding to a target image into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a 3D space, respectively. A target distance between the first object and the second object in a three-dimensional space is then determined by scaling a distance between the first transformed location and the second transformed location with the distance scale.
The present solution can effectively utilize the 2D image and greatly reduce the modeling and usage costs for achieving 2D-to-3D scene restoration by increasing the information perception dimension (multi-vision camera, LIDAR, etc. ) in the application field. Further, the present solution does not require distance metric calculation on the basis of 3D model, and ca estimates the plane distance of 2D object in the 3D space directly by model reduction, which greatly reduces the technical development threshold and usage cost in the application field, and enhances the convenience in different application scenarios. As such, the present solution can be easily adopted in light application scenarios.
Example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It is noted that any section/subsection headings provided herein are not intended to be limiting. Embodiments are described throughout this document, and any type of embodiment may be included under any section/subsection. Furthermore, embodiments disclosed in any section/subsection may be combined with any other embodiments described in the same section/subsection and/or a different section/subsection in any manner.
Example Environment
FIG. 1 illustrates an example environment 100 in which example embodiments of the present disclosure can be implemented. In the environment 100, a camera is mounted in a physical environment to capture images. As shown, a camera 102-1 is mounted with a specific deployment in a scene 101-1 (represented as “Scene 1” ) , ..., a camera 102-N is mounted with a specific deployment in a scene 101-N (represented as “Scene N” ) , where N may be an integer larger than or equal to one. The cameras 102-1, ...,
102-N may be the same or different types of cameras, e.g., with different imaging parameters. The deployments of the cameras 102-1, ..., 102-N may be the same or different although the scenes may be different. In some example embodiments of the present disclosures, the cameras 102-1, ..., 102-N may be monocular images, capturing monochrome images or colored images (such as RGB images) .
For the purpose of discussion, the cameras 102-1, ..., 102-N may be collectively or individually referred to as cameras 102, and the scenes 101-1, …, 101-N may be e collectively or individually referred to as scenes 101.
In embodiments of the present disclosure, the images captured by the cameras 102 may be transmitted to a device (s) /system (s) for object distance estimation and in some cases, for abnormal distance detection. The images may be transmitted via various networking technology. For example, the camera 102-1 may connect to a radio access network (RAN) and thus can transmit its images to a network device 103 in the RAN. The network device 103 may convey the images to a core transmission network 110, e.g., via a switch 104 and an access router 105. The images may be routed to an object distance estimation device in an information management monitoring room 120 which can access to the core transmission network 110 via an access router 122. The camera 102-N may have a network connection in a local access network (LAN) and thus can transmit its images to the core transmission network 110 via a switch 106 in the LAN. The images captured by the camera 102-N may also be conveyed to an object distance estimation device in the information management monitoring room 120.
In some example embodiments, the devices in the information management monitoring room 120 may subscribe to remote computing resources from a computing center 130 for processing the images and/or other processing operations required in the object distance estimation and the abnormal distance detection. The computing center 130 may be accessed through the core transmission network 110 via an access router 132. The computing center 130 may contain one or more computing resource clusters to provide the processing, memory, storage, and networking resources for the information management monitoring room 120. In some example embodiments, the computing center 130 may be implemented as a cloud platform.
It would be appreciated that in some example embodiments, images captured by a camera 102 may be processed remotely (e.g., in the information management monitoring
room 120 and/or in the computing center 130) for the object distance estimation and abnormal distance detection, and in some other embodiments, the images captured by the camera 102 may be processed locally, e.g., by a device deployed locally in the scene and communicatively coupled with the camera 102. The scope of the present disclosure is not limited in this regard.
Work Principle and Example Framework
The example embodiments of the present disclosure propose a solution to derive 3D distance information from an 2D image (e.g., an image captured by a monocular camera) using a perspective transformation model. The perspective transformation model is trained for a target deployment of a target camera, and is used to generate a perspective transformation and a distance scale for a target image captured by the target camera. In some example embodiments, as will described in detail below, an object detection model may be trained to detect the objects from the target image. As the models are introduced, there may be a training stage for the models and then an inference stage to apply the trained models for the distance estimation.
FIG. 2 illustrates a framework 200 for model-based object distance estimation in accordance with some example embodiments of the present disclosure. The framework 200 comprises a model training system 210 in a training stage and a model inference system 220 in an inference stage. The model training system 210 is configured to train one or more perspective transformation models 204-1, …, 204-N for one or more scenes (Scene 1, …, Scene N) respectively. For the purpose of discussion, the perspective transformation models 204-1, …, 204-N may be collectively or individually referred to as perspective transformation models 204. In some example embodiments, the model training system 210 may be further configured to train an object detection model 202 to detect an object (s) in an image. The model inference system 220 is configured to determine a distance between objects detected from a target image 232 using the trained object detection model 202 and the trained perspective transformation model 204.
The model training system 210 and the model inference system 220 may be the same or different systems. For example, in the example environment 100 as shown in FIG. 1, the model training system 210 may be implemented at the computing center 130, and the model inference system 220 may be implemented at one or more devices in the information management monitoring room 120. In another example, both the model
training system 210 and the model inference system 220 may be implemented as a same system/device in the computing center 130 and/or in the information management monitoring room 120.
The object detection model 202 and the perspective transformation models 204 are first trained and then applied for inference. FIG. 3 illustrates a flowchart of a process 300 for model training and inference for model-based object distance estimation in accordance with some example embodiments of the present disclosure. The process 300 will be described with reference to FIG. 2. Some training-related operations in the process 300 may be performed by the model training system 210, and some inference-related operations in the process 300 may be performed by the model inference system 220.
At block 310, the model training system 210 collects first sample images 212 and first label information 213 indicating objects detected in the first sample images 212. At block 320, the model training system 210 trains the object detection model 202 with the first sample images 212 and first label information 213.
The input to the object detection model 202 may be an image, and the output from the object detection model 202 may be a location (s) of an object (s) in the input image. In some example embodiments, the object detection model 202 may provide a bounding box to locate an object in the image.
As shown in FIG. 3, the training data for the object detection model 202 may comprise sample images 212 (sometimes referred to as first sample images 212) and first label information 213 indicating objects detected from the first sample images 212. The object detection model 202 is used to detect one or more objects of interest from an image (the distance (s) of which are to be estimated) . Thus it is expected that the object detection 202 may not be associated with a specific scene or a deployment of a camera in the scene. In some example embodiments, the first sample images 212 may be collected from various sources, for example, by a plurality of cameras with a plurality of deployments in Scene 1, …, Scene N.
The training of the object detection model 202 will be discussed in more detail below with reference to FIG. 4.
At block 330, the model training system 210 collects second sample images 214-1, …, 214-N and second label information 215-1, …, 215-N for the second sample images
214-1, …, 214-N. At block 340, the model training system 210 trains the perspective transformation model 204-1, …, 204-N with the second sample images 214-1, …, 214-N and second label information 215-1, …, 215-N, respectively.
For the purpose of discussion, the second sample images 214-1, …, 214-N may be collectively or individually referred to as second sample images 214; and the second label information 215-1, …, 215-N may be collectively or individually referred to as second label information 215.
The input to a perspective transformation model 204 may be an image, with one or more objects detected in the image, and the output from the perspective transformation model 204 may be a perspective transformation and a distance scale for the input image. The perspective transformation is used to transform the input image (which is in the 2D pixel space) into a top plane view of a 3D space, to obtain a transformed image. The distance scale is used to scale distances between pixels (or object locations) in the transformed image into the physical distance in the scene. The perspective transformation model 204 may also be referred to as a “perspective transformation and filed of view size” model.
Each perspective transformation model 204 may be associated with a scene, and it is assumed that a camera is deployed in the scene with a specific deployment. The perspective transformation model 204 is trained to adopt to the perspective transformation and field of view scaling for the specific deployment of the camera, and thus can be used for inference of images captured by a camera with the corresponding deployment. Specifically, training data for a perspective transformation model 204-1 may include second sample images 214-1 captured by one or more cameras with a deployment (Deployment 1) in Scene 1, …, training data for a perspective transformation model 204-N may include second sample images 214-N captured by one or more cameras with a different deployment (Deployment N) in Scene N. In some example embodiments, the second sample images 214-1, …, 214-N used for training the perspective transformation models may be divided from the first sample images 212 according to the scenes or deployments.
In some example embodiments, a deployment of a camera may be identified by a mounting height, a shooting orientation, and a type of the camera in the scene. For example, for the same type of image, if it is deployed at different heights and/or with
different orientations, then the modelling for 2D-to-3D scene restoration may be different. For different types of cameras, it may have different focal lengths, and then the modelling for 2D-to-3D scene restoration may be different. Then different perspective transformation models 204 may need to be trained for the different deployments of cameras.
The training data for a perspective transformation model 204 may further include second label information 215 for the second sample images 214. For each second sample image 214, the second label information 215 may indicate a labeled indication of whether a distance between a pair of objects detected in the second sample image is abnormal. In some example embodiments, for a second sample image 214, the second label information may further indicate labeled transformed locations of the pair of objects in the top plane view.
The training of the perspective transformation models 204 will be described in detail below with reference to FIG. 5.
With the object detection model 202 and the perspective transformation models 204 have been trained, at block 350, the model inference system 220 determines a target distance between a first object and a second object in a target image 232 using the object detection model 202 and a perspective transformation model 204. As shown in FIG. 2, the model inference system 220 may comprise a distance estimator 230 to determine the target distance between the first and second objects using the outputs from the object detection model 202 and the perspective transformation model 204.
The perspective transformation model 204 applied for the target image 232 may depend on a target deployment of a target camera which captures the target image 232. It is assumed that the target image 232 is captured by a target camera with a target deployment in Scene X. Among all the perspective transformation models 204, the perspective transformation model 204-X associated with the target deployment in Scene X may be apply by the model inference system 220 for determining the distance.
The object detection model 202 is used to detect two or more objects from the target image 232, to determine locations of the two or more objects in the target image 232. The locations may represent location information of the objects within a 2D space (also referred to as a 2D pixel space) corresponding to the target image 232. The perspective transformation model 204-X is used to transform the location of a first object
and the location of a second object within the 2D space into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a 2D space, respectively. The transformed locations are considered as 3D locations for the first and second objects.
In some example embodiments, at block 360, the model inference system 220 may perform abnormal distance detection based on the target distance between the first object and the second object. As shown in FIG. 2, the model inference system 220 may comprise an distance anomaly detector 235 to perform the abnormal distance detection. The distance anomaly detector 235 may determine an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
The object distance estimation and abnormal distance detection will be further discussed in more detail below with reference to FIG. 9.
The present solution of the example embodiments disclosure herein can effectively improve the efficiency of 2D imaging to 3D distance detection, greatly reduce the hardware cost and model training and inference costs for scenarios such as in-site hazard distance detection, while achieving the expected 3D distance detection in multiple crossover application areas, including but not limited to pedestrian distance detection, construction vehicle safety distance detection, hazardous object safety distance detection, and the like.
Training of Object Detection Model
FIG. 4 illustrates a flowchart of a process 400 for training the object detection model 202 in the framework 200 of FIG. 2 in accordance with some example embodiments of the present disclosure. For the purpose of discussion, the process 400 will be described from the perspective of the model training system 210 in FIG. 2.
At block 410, the model training system 210 obtains first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images, for example, the first sample images 212 and the first label information 213 as shown in FIG. 2.
In some example embodiments, in order to train a generalized object detection model, the first sample images 212 may be collected from different dataset sources (for
example, captured by different cameras with different deployments in the scenes) . In some example embodiments, the cameras used for capturing the first sample images 212 may be monocular cameras. In some example embodiments, a first sample image 212 may be stored as (scene ID, img_x) , where the scene ID is used to identify a deployment of a camera in the scene, and one scene or deployment may correspond to one camera.
To ensure the generality of the first sample images 212 for training the object detection model 202, the differential data collection for multiple environmental states of the scenes may be considered, including but not limited to, different weather conditions (e.g., sunny, cloudy, rainy, etc. ) , different time conditions (e.g., early morning, midday, evening, etc. ) , and the like. If a specific deployment of a camera or the corresponding scene is a special deployment or a special scene, the impact of the environmental changes in that scene on the collection of the sample images may need to be taken into consideration.
The first sample images 212 are used to perform the training of the object detection model 202 which is primarily configured to localize objects in 2D pixel space. The first sample images 212 may be labeled with the objects occurred therein. In some example embodiments, the first sample images 212 may be labeled with bounding boxes of the objects to be detected, respectively. The labelling may be performed through manual annotations or through any suitable automation tools, to generate the first label information 213 for the first sample images 212. For an object labeled in a first sample images 212, the first label information 213 may indicate a bounding box of the area of the object in the first sample images 212. In some examples, the first label information 213 may include coordinates of the detected object in the first sample images 212, such as the four vertexes of the bounding box, , including the left, top, right, bottom vertexes (in this context, denotes l as left, t as top, r as right, b as bottom) for the pixel coordinate points.
In some example embodiments, for some use cases, it is expected to estimate a distance (s) between some specific types of objects, e.g., to evaluate whether those objects are too close to each other or too far away from each other. In those cases, the first sample images 212 containing the specific types of objects of interest may be sampled from all the available sample images. For example, if it is expected to determine an abnormal distance between a forklift and a pedestrian in a scene monitored by a camera, then the first sample images 212 containing a forklift and one or more pedestrians may be sampled from all the images captured by the camera. It would be appreciated that the object
detection model 202 may be trained to detect various types of objects at least including some types of objects whose distances are to be estimated, and in this case, more objects may be labeled in the first sample images 212.
At block 420, the model training system 210 constructs the object detection model 202. In some example embodiments, the object detection model 202 may be constructed using a generic detection model framework in computer vision. The scope of the present disclosure is not limited in the specific model structure for the object detection model.
At block 430, the model training system 210 trains the object detection model 202 with the first sample images 212 and the first label information 213.
During the training, the model training system 210 may iteratively optimize the object detection model 202 for completing the object detection task. The model training system 210 may apply any suitable model training algorithms and strategies to complete the training of the object detection model 202.
In some example embodiments, some hyperparameters may be set for the training of the object detection model 202. For example, a hyperparameter confidence_threshold may be set to a predetermined value (in a range between 0 and 1) for thresholding the selection of detected pre-selected bounding boxes of objects. Another hyperparameter NMS_IOU_threshold may be typically set to a predetermined value between 0 and 1 for controlling the screening of the pre-selected bounding boxes with high overlapping.
In some example embodiments, the iterative optimization may be completed if the F1-score of the object detection model 202 meets a training objective, e.g., the F1-score being close to 1 on a training dataset and the generalization of the object detection model 202 can prevent from overfitting on a validate dataset.
In some example embodiments, the first sample images 212 used for training the object detection model 202 may be further inspected, to guarantee the accuracy of the next model construction, and removing the missed and over detected RGB images from the sample data. After the training of the object detection model 202 is completed, the first sample images 212 may be evaluated to determine which objects are missed or mis-checked by the trained object detection model 202. If it is obvious that the object detection
model 202 has not yet trained to converge, re-training of the object detection model 202 may be performed with the first sample images 212 whose objects are missed. During the re-training, the gradient weights for those first sample images 212 may be enhanced. The model training may be iterated and the object detection model whose F1-score is at or near 1 on the training dataset and which has a strong generalization in the validation dataset, may be selected as the trained model for inference.
Training of Perspective Transformation Model
FIG. 5 illustrates a flowchart of a process 500 for training the perspective transformation model (s) 204 in the framework 200 of FIG. 2 in accordance with some example embodiments of the present disclosure. For the purpose of discussion, the process 600 will be described from the perspective of the model training system 210 in FIG. 2.
As mentioned above, the training data for a perspective transformation model 204 associated with a specific deployment of a scene may include second sample images 214 and second label information 215 for the second sample images 214. The training of one perspective transformation model 204 is discussed below and all the other perspective transformation models for other scenes may be trained in similar ways.
At block 510, the model training system 210 obtains the second sample images 214 with objects detected.
In some example embodiments, the second sample images 214 used for training a perspective transformation model 204 for a specific scene may include a subset of the first sample images 212 used for training the object detection model 202 which is collected from the specific scene. The first image samples 212 may be grouped according to the scene IDs, and the second sample images 214 may be selected from each scene (or deployment of the camera) .
In some example embodiments, a part or all of the second sample images 214 may be different from those sample images used for training the object detection model 202. Then the objects in the second sample images 214 may be separately labeled from the labeling of the first sample images 212. In some example embodiments, if the object detection model 202 has been trained, the objects in the second sample images 214 may be detected using the trained object detection model 202.
In some example embodiments, the objects labeled in the second sample images
214 may be the types of objects whose distance (s) are to be estimated or the abnormal distance (s) are to be detected in the corresponding scene. For example, in one scene, it is expected to estimate a distance between a forklift and a pedestrian, and then these two types of objects are labeled in the second sample images 214 captured from this scene. For another scene, it is expected to estimate a distance between two cars, and then a type of car object is labeled in the second sample images 214 captured from the other scene.
At block 520, the model training system 210 determines a fitted perspective transformation based on the second sample images 214 with the objects detected. At block 530, the model training system 210 determines second label information 215 for the second sample images 214 based at least on the fitted perspective transformation. For each second sample image 214, the second label information 215 may indicate a labeled indication of whether a distance between a pair of objects detected in the second sample image is abnormal. In some example embodiments, for a second sample image 214, the second label information may further indicate labeled transformed locations of the pair of objects in the top plane view.
For the anomaly indication, in some example embodiments, in conjunction with labeling of the bounding boxes of the objects, the distances between one or more pairs of objects are also labeled for anomalies. For example, the annotators may evaluate whether the distance between a pair of objects in a second sample image 214 is abnormal, e.g., according to some reference objects or reference distance in the second sample image 214. The second sample image 214 and its second label information 215 may be stored as (scene ID, img_x, an indication of whether the distance is abnormal or not) .
For the labeled transformed locations, a perspective transformation may be fitted from the labeled locations of the objects in the second sample images 214, for training of the perspective transformation model 204. As mentioned above, the perspective transformation model 204 is configured to output a perspective transformation and a distance scale for a specific deployment in a scene. The perspective transformation is used to transform a 2D location of an object in the 2D pixel space into a 3D location in a 3D space, or more specifically, a planar location in a top plane view of the 3D space. The distance scale is used to scale a distance between transformed location in the top plane view into the actual physical distance in the 3D space.
A perspective transformation (also referred to as a perspective transformation matrix)
may be represented asthe distance scale may be represented as Scaler. The Mperspective may be flattened, for example, as a (1, 9) matrix, and the distance scale Scaler is a (1, 1) matrix. A position transformation for a location in the 2D space corresponding to an image (e.g., a 2D pixel coordinate point) is defined as follows:
ptsco-bev=ptsco-pix×Mperspective (1)
ptsco-bev=ptsco-pix×Mperspective (1)
where ptsco-pix represents the location in the image (e.g., the pixel coordinates in the image) , and ptsco-bev represents a transformed location (e.g., the 2D top plane coordinates) in the 3D top plane view. The transformation may allow the 2D top plane coordinates obeying three degrees of freedom. Further, in order to keep the scale consistency, the distance scaler Scaler is introduced, and then the physical distance between two objects (object i and j) in an image may be calculated as follows:
whererepresents a transformed location of object i in the top plane view, represents a transformed location of object j in the top plane view, ‖‖l2 represents the L2-distance calculated from the 2D top plane coordinates of the two objects in the tope plane view.
FIG. 6A and FIG. 6B illustrate schematic diagrams of example image transformation and normalized object distances in a 3D top plane view. As shown, a captured image 610 or 640, which may be a 2D image captured by a monocular camera, may be transformed through a corresponding perspective transformation into a transformed image 620 or 650 which is considered as a view of the 3D scene from the tope plane view. The locations of objects 611, 612, and 613 in the image 610 may be mapped to respective transformed locations 621, 622, and 623 of the corresponding objects within the transformed image 620. After scale normalization, the transformed locations 621, 622, and 623 of the objects may be represented in a normalized top plane view 630. Similarly, the locations of objects 641, 642, and 643 in the image 640 may be mapped to respective transformed locations 651, 652, and 653 of the corresponding objects within the transformed image 650. After normalization, the transformed locations 651, 652, and 653 of the objects may be represented in a normalized top plane view 660.
The training of the perspective transformation model 204 is to learn an accurate perspective transformation as well as an accurate distance scale to estimate the 3D distance between two objects. For this purpose, in the training stage, a fitted perspective
transformation may be determined for a specific deployment in a scene. In some example embodiments, the locations of the objects labeled in the second sample images 214 (e.g., coordinates of four vertexes of a bounding box of an object in a second sample image 214) may be selected and used to calculate the fitted perspective transformation of the scene.
The fitted perspective transformation is used for transforming the locations of the objects in the second sample images 214 into transformed locations in the top plane view of the 3D space. The fitting process may be considered as there are errors in the manual evaluation, so a multi-labeling cross-validation approach is considered. Take a scene i as an example. One of the second sample images 214 collected for the scene i may be randomly selected for construction of the fitted perspective transformation. Another second sample image 214 is randomly selected for verifying whether the perspective transformation meets the expectations of the transformation. More specifically, for a selected second sample image 214, coordinates of four foot points in this image may be recorded as [pts1, pts2, pts3, pts4] , and the coordinates of the 4 foot points after the perspective transformation are defined as [pts′1, pts′2, pts′3, pts′4] . The transformed location (s) of one or more object (s) in the second sample image 214 (e.g., the coordinates of the four vertexes of the bounding box) may be evaluated to determine whether the locations (e.g., the coordinates) match the expected coordinates in the top view plane. The coordinates of the four foot points after the perspective transformation may be adjusted until they meet the expectations of the perspective transformation. That is, the fitting is repeated until the obtained fitted perspective transformation meets the expectation.
The fitted perspective transformation of each scene (Scene n) is recorded as (scene ID, img_x, ) . The fitted perspective transformation may be combined with the location of the objects detected in the second sample images 214 (e.g., the bounding boxes) . For each second sample image 214, the model training system 210 may transform, based on the fitted perspective transformation, the respective locations of the objects into respective transformed locations within the top plane view of the 3D space.
In some example embodiment, a location of an object detected from an image may be defined as a plurality of coordinates of a first plurality of key points of the object detected from the image. The key points for an object may include the four vertexes of a bounding box of the object. In some examples, the key points may further include a centroid of the object. Specifically, a location of an object in a second sample image 214
may be represented as coordinates of the four vertexes (left, top, right, and bottom) of a bounding box of the object within the second sample image 214, which may be represented as (Coleft, Cotop, Coright, Cobottom) . In some example embodiments, a centroid of the object may also be taken into account to represent the location of the object. In some example embodiments, since the centroid point is not easy to locate and due to the absence of object information with 3D, the coordinates of the centroid may be represented as
and Coc_y=Cobottom.
In some example embodiments, a key point transformation mapping may be determined for a second sample image 214 based on the fitted perspective transformation. The location of an object detected in a second sample image 214 may thus be represented as (Coleft, Cotop, Coright, Cobottom, (Coc_x, Coc_y) ) , i.e., five coordinates in the 2D space corresponding to the second sample image 214. The five coordinates may be split into five coordinate points, Kpt= [ (Coleft, Cotop) , (Coright, Cotop) , (Coc_x, Coc_y) , (Coleft, Cobottom) , (Coright, Cobottom) ] , corresponding to the coordinate points of top-left, top-right, center, bottom-left, bottom-right, where the center point is denoted as the center of the rigid object.
Further, the 2D coordinates of the location of the object may be filled with a predetermined value (e.g., 1.0) for the third dimensional, to transformto
Then Kpt may be transformed based on the fitted perspective transformation, to obtain a transformed location Kpt′in the top plane view. The mapping may be represented as Kpt′=Kpt×Mperspective. After scale normalization, the transformed location may be
As a result, a second sample image 214 and its second label information 215 may be constructed as a training sample for the perspective transformation model 204 as {image_x, Kpt, Kpt′, i, j, p/n} , where i and j represent a pair of objects in the second sample image 214 whose distance is to be estimated, the indicator “p” or “n” indicates whether the distance between the objects i and j is abnormal or not, Kpt indicates locations of the objects within the 2D space corresponding to the second sample image 214, and Kpt′ indicates transformed locations of the objects in the 3D top plane view (which may be used to calculate the labeled distance between a pair of objects in the 3D top plane view) .
In some example embodiments, if more than two pairs of objects of interest are detected from the second sample image 214, the objects i, j may be selected as the pair of object with the lowest transformed centroid distance, e.g.,
if the abnormal distance is determined as a distance lower than a distance threshold. For example, if a forklift and two pedestrians are detected in a second sample image 214 (as in the examples of FIG. 6A and FIG, 6B) , then a pair of forklift and a pedestrian with the lowest distance may be selected for labeling the second sample image 214. In the case where an abnormal distance is determined as a distance exceeding a distance threshold, the pair of object with the highest transformed centroid distance may be selected. In some example embodiments, the second sample image 214 may be labeled with the second label information associated with more than one pair of objects whose distance (s) is to be estimated.
In some example embodiments, due to the large difference in the resolution of the scene camera, the real object size in the images corresponding to the aspect ratio of the sample images may have a large difference. Therefore, in the model training phase, it is considered to train the perspective transformation model 204 by varying the scale of the second sample image 214.
The resolution of a second sample image is represented as (h, w) . To keep the spatial dimensions in the image from being further stretched in a destructive manner, the size of the second sample image may therefore be changed to Size= (h /max (w, h) *IMAGE_SIZE, w/max (w, h) *IMAGE_SIZE) , where IMAGE_SIZE is set to the image scaling size.
At block 540, the model training system 210 constructs the perspective transformation model 204. The perspective transformation model 204 may be constructed with any suitable model structure in the computer vision (e.g., with MobileNet_v2 as the backbone model) . The modeling of the perspective transformation model 204 may be represented aswith an image (represented as “img” ) as its input, and output a perspective transformationand a distance scale Scaler for the input image.
At block 550, the model training system 210 trains the perspective transformation model 204 with the second sample images 214 and the second label
information 215. As mentioned above, for a specific perspective transformation model 204, a training sample may be constructed as follows: {image_x, Kpt, Kpt′, i, j, p/n} , image_x represents a second sample image 214, and (Kpt, Kpt′, i, j, p/n) represents the second label information 215.
In some example embodiments, the training of the perspective transformation model 204 may be based on a first loss function (also referred to as an objective function) , which is used to evaluate a first error between a predicted indication of distance anomaly and a labeled indication of distance anomaly for a second sample image (s) 214. The first loss function may be represented as follows:
The training objective is to decrease or minimize a loss value of the first loss function by iteratively updating or optimizing the perspective transformation model 204 (e.g., updating the model parameter values) . The calculation of a loss value of the first loss function is as follows. For each second sample image 214, the model training system 210 may determine a predicted perspective transformationoutput by the perspective transformation model 204 for the second sample image 214 (based on the current model parameter values) , and then determine respective predicted transformed locations of a pair of objects in the top plane view of the three-dimensional space using the predicted perspective transformation. A predicted transformed location of an object i is determined asis determined based on the location (coordinates) of the object i within the second sample image 214.
The model training system 210 may further determine a predicted distance between the pair of objects (e.g., objects i and j) by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image. The predicted distance may be determined asas in Equation (3) .
The model training system 210 may determine a predicted indication of whether a distance between the pair of objects is abnormal based on a comparison result between the predicted distance and a distance threshold. For example, if a distance lower than the distance threshold is determined as an abnormal distance, the predicted distance lower
than the distance threshold may be indicated as an abnormal distance. Otherwise, if a distance exceeding a distance threshold is determined as an abnormal distance, the predicted distance exceeding the distance threshold may be indicated as an abnormal distance. The parameter amargin in Equation (3) represents a distance threshold, and in this example, a distance lower than the distance threshold is determined as an abnormal distance.
In Equation (3) , yi is the labeled indication of whether a distance between objects i and j is abnormal or not, and yi =1 indicates an abnormal distance, yi = -1 indicates a normal distance. According to Equation (3) , the loss value of the first loss function is lower or decreased as the predicted indication approximates to the labeled indication, which means that the perspective transformation model 204 has learned to predict relatively accurate perspective transformation and distance scale to make the correct anormal distance detection. Otherwise, the loss value of the first loss function is high or increased if the predicted indication is different from the labeled indication, which means that the predicted perspective transformation and the distance scale do not meet the expectation. The perspective transformation model 204 may need to be further optimized.
In some example embodiments, the training of the perspective transformation model 204 may be further based on a second loss function, which is used to evaluate a second error between a predicted transformed location (s) of an object (s) in a second sample image 214 and a labeled transformed location (s) of the object (s) in the second sample image 214. The second loss function may be represented as follows:
The training objective is to decrease or minimize a loss value of the second loss function by iteratively updating or optimizing the perspective transformation model 204 (e.g., updating the model parameter values) . The calculation of a loss value of the first loss function is as follows. In Equation (3) , for each second sample image 214, a predicted transformed location of an object i is determined asas described above. Kpt′i indicates a labeled transformed location of the object i .
represents the difference (or error) between the predicted transformed location and the labeled transformed location. In some example embodiments, the predicted transformed locations and the labeled transformed locations for more objects detected in
the second sample image 214 (not only the objects whose distance is labeled for anomaly) may be determined, to calculate the loss value of the second loss function Lkpt.
In some example embodiments, the overall objective function for training the perspective transformation model 204 may be set to Lopt=a·Lkpt+ (1.0-a) ·Ldis, where a∈ [0, 1] is used to control the weights of the two loss functions during the training process. The weight a may be initialized (e.g., to 0.5, 0.8, or the like) and may be adjusted or gradually reduced during the training (e.g., gradually reduced by 0.01 after a number of epochs of iterations) . It would be appreciated that one or more additional or alternative loss functions may be designed for training the perspective transformation model 204 based on the second label information 215 for the second sample images 214.
In some example embodiments, the training of the perspective transformation model 204 may reference to the architecture 700 in FIG. 7. As illustrated, at Phase 1, the trained object detection model 202 is applied to detect objects from the second sample images 214, and thus the five coordinate points, Kpt= [ (Coleft, Cotop) , (Coright, Cotop) , (Coc_x, Coc_y) , (Coleft, Cobottom) , (Coright, Cobottom) ] for each object may be derived based on the location of the detected object in the second sample images 214, for example, Kpti and Kptj. Through offline fitting of perspective transformation, the labeled transformed locations of the objects, e.g., Kpt′i and Kpt′j may be determined by mapping Kpti and Kptj with the fitted perspective transformation. At phase 2, the perspective transformation model 204 is trained, and the predicted transformed locations of the objectsmay be determined and used to calculate the loss values of the first and/or second loss function.
In some example embodiments, because the model training stage involves segmented training process (the object detection model 202 and the perspective transformation model 204) , considering the order of model optimization, the optimization accuracy of the bounding box of the object directly affects the accuracy of the subsequent calculation. Therefore, to ensure the accuracy of the perspective transformation models 204, in the process of model training, the generalization and robustness performance of the perspective model 204 may be further enhanced by randomly changing the locations for objects in the sample images. FIG. 8 illustrates an example 800 of floating key points in the model training. As illustrated, the locations for objects 811 and 812, Kptm and Kptn, in a sample image may be transformed into transformed locations Kpt′m and Kpt′n
812 may be adjusted toandand the transformed locations may also be changed asand
Object Distance Estimation and Abnormal Distance Detection in Inference Stage
FIG. 9 illustrates a flowchart of a process 900 for object distance estimation and abnormal distance detection in an inference stage in accordance with some example embodiments of the present disclosure. For the purpose of discussion, the process 900 will be described from the perspective of the model inference system 220 in FIG. 2.
At block 910, the model inference system 220 detects a first location of a first object and a second location of a second object within a 2D space corresponding to a target image 232 captured by a camera with a target deployment in a scene.
In some example embodiments, the model inference system 220 may detect the first location and the second location using the trained object detection model 202. The first object and the second object may be objects whose distance therebetween is to be estimated, e.g., for the anomaly detection or for other purposes.
In some example embodiments, through the detection, the first object or the second object may be detected in the target image 232 by respective bounding box. The first location of the first object may include a first plurality of coordinates of a first plurality of key points of the first object detected from the target image 232, and the second location may include a second plurality of coordinates of a second plurality of key points of the second object detected from the target image. The first plurality of key points may at least include a first centroid of the first object, and the second plurality of key points may at least include a second centroid of the second object.
The key points for an object may include coordinates of the four vertexes of a bounding box of the object and the centroid of the object, which may be represented as In some examples, the five coordinates may be split into five coordinate points to represent, Kpt= [ (Coleft, Cotop) , (Coright, Cotop) , (Coc_x, Coc_y) , (Coleft, Cobottom) , (Coright, Cobottom) ] , to represent a location of the corresponding location.
At block 920, the model inference system 220 determines a perspective
transformationand a distance scale Scaler for the target image 232 by providing the target image 232 into a trained perspective transformation model 204 associated with the target deployment. The perspective transformation model 204 used for the target image 232 may be selected from a plurality of perspective transformation models 204 based on a target deployment identity of the target deployment. The plurality of perspective transformation models 204 are each identified with respective deployment identities. The deployment identity may sometimes correspond to a scene ID for the scene where the camera is deployed.
At block 930, the model inference system 220 transforms, based on the determined perspective transformationthe first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively. For example, a transformed location for an object may be determined as
At block 940, the model inference system 220 determines a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale Scaler.
In some example, the first transformed location and the second transformed location may be normalized, thoughIn some example embodiments, the target distance between the first transformed location and the second transformed location may be calculated as distance=‖Kpt′ (i) -Kpt′ (j) ‖l2Scaler . In some example embodiments, the coordinates of the centroids of the first and second objects may be used to determine the distance between two objects.
In some example embodiments, the model inference system 220 may further determine an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold. The distance threshold may be configured according to the actual applications of distance anomaly detection. In some example embodiments, if a distance lower than the distance threshold is considered as abnormal, the model inference system 220 may determine whether the target distance is abnormal or not based on a
determination of whether the target distance is lower than the distance threshold. In some example embodiments, if a distance exceeding the distance threshold is considered as abnormal, the model inference system 220 may determine whether the target distance is abnormal or not based on a determination of whether the target distance exceeds the distance threshold.
In some example embodiments, actions may be determined and performed according to the distance anomaly detection. As some examples, if it is determined that the distance (s) between objects is abnormal, alarms may be generated and transmitted to one or more related devices, and/or operation commands may be sent to the related devices. In the example use case of detecting whether a vehicle or forklift is too close to a pedestrian (s) , in response to a detection of the abnormal distance between the vehicle or forklift, an operation command of applying a force stop may be sent to the vehicle or forklift, and/or an alarming device may be activated to warn the driver and the pedestrian (s) of the dangerous situation in the scene.
The present disclosure provides solutions for object distance estimation and distance anomaly detection based on monocular image data, which provides a lightweight detection method for video distance anomaly monitoring application scenarios, and effectively reduces the implementation cost and usage cost. According to the example embodiments of the present disclosure, instead of performing a complete 2D-3D scene restoration to achieve the estimation of the distance between objects, the models are trained and applied to transform pixel location in an image into locations in the 3D top plane view, to calculate the distance between the objects is a dangerous distance. This can avoid a lot of tedious information collection and feature construction work in the early stage, and reduce the complexity of the model and its absolute requirements for data accuracy.
Further, in the training of the models, there is no need for the precise distance labeling for the 3D scenes in absolute sense. Thus, from the perspective of modelling, the example embodiments of the present disclosure may not need to strictly define the absolute coordinate positions fitted in the 3D scenes. Due to the full consideration of the fitting of perspective transformation and the fitting of abnormal distance, the distance estimation and the anomaly detection based on the model output are still reliable.
Example Apparatus, Device and Medium
The present application may be described in the general context of computer-executable instructions executed by a computer, such as a program module or unit. Generally, a program module or unit may include routines, programs, objects, components, data structures, and the like that perform a particular task to implement a particular abstract data type. In general, the program module or unit may be implemented by software, hardware, or a combination of both. The present application may also be practiced in distributed computing environments where the program modules or units may be located in local and process computer storage media, including storage devices, or by storage in portable computing devices for abnormally distant data transfer.
In some example embodiments, an apparatus capable of performing any of the processes 300, 400, 500, and 900 may comprise means for performing the respective operations of those processes. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or software module. The first apparatus may be implemented as or included in the model inference system 220 and/or the model training system 210 in FIG. 2.
In some example embodiments, the apparatus comprises means for means for detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene; means for determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment; means for transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; and means for determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
In some example embodiments, the first apparatus further comprises: means for determining an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
In some example embodiments, the means for determining the perspective
transformation and the distance scale for the target image comprises: means for selecting the perspective transformation model associated with the target deployment from a plurality of perspective transformation models based on a target deployment identity of the target deployment, the plurality of perspective transformation models being identified with respective deployment identities.
In some example embodiments, the target deployment is identified based on a mounting height, a shooting orientation, and a type of the target camera.
In some example embodiments, the means for detecting the first location and the second location comprises: means for detecting the first location and the second location using an object detection model, wherein the object detection model is trained with first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images.
In some example embodiments, the first location comprises a first plurality of coordinates of a first plurality of key points of the first object detected from the target image, and the second location comprises a second plurality of coordinates of a second plurality of key points of the second object detected from the target image, wherein the first plurality of key points at least comprises a first centroid of the first object, and the second plurality of key points at least comprises a second centroid of the second object.
In some example embodiments, the perspective transformation model is trained using second sample images captured by a camera with the target deployment and second label information indicating at least one of the following for each second sample image: a labeled indication of whether a distance between the pair of objects is abnormal, the pair of objects being detected in the second sample image, or respective labeled transformed locations of the pair of objects in the top plane view of the three-dimensional space.
In some example embodiments, the perspective transformation model is trained by the following: determining respective predicted transformed locations of the pair of objects in the top plane view of the three-dimensional space using a predicted perspective transformation output by the perspective transformation model for the second sample image; determining a predicted distance between the pair of objects by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image; determine a predicted indication of whether a distance between the pair of objects is abnormal based on a
comparison result between the predicted distance and a distance threshold; determining a first loss value of a first loss function based on a first error between the predicted indication and the labeled indication for the second sample image; and updating the perspective transformation model by decreasing the first loss value of the first loss function.
In some example embodiments, the perspective transformation model is further trained by the following: determining a second loss value of a second loss function based on second errors between the respective predicted transformed locations and the respective labeled transformed locations; and updating the perspective transformation model by decreasing the second loss value of the second loss function.
In some example embodiments, the respective labeled transformed locations of the pair of objects are determined by the following: determining a fitted perspective transformation associated with the target deployment based on respective locations of the pair of objects within a two-dimensional space corresponding to the second sample image; transforming, based on the fitted perspective transformation, the respective locations of the pair of objects into the respective transformed locations within the top plane view of the three-dimensional space.
In some example embodiments, the target camera is a monocular camera.
In some example embodiments, the apparatus further comprises means for performing other operations in some example embodiments of the present disclosure. In some example embodiments, the means comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the performance of the apparatus.
FIG. 10 is a simplified block diagram of a device 1000 that is suitable for implementing example embodiments of the present disclosure. The device 1000 may be provided to implement a communication device, for example, the model inference system 220 and/or the model training system 210 as shown in FIG. 2. As shown, the device 1000 includes one or more processors 1010, one or more memories 1020 coupled to the processor 1010, and one or more communication modules 1040 coupled to the processor 1010.
The communication module 1040 is for bidirectional communications. The
communication module 1040 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interfaces may represent any interface that is necessary for communication with other network elements. In some example embodiments, the communication module 1040 may include at least one antenna.
The processor 1010 may be of any type suitable to the local technical network and may include one or more of the following: general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multicore processor architecture, as non-limiting examples. The device 1000 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.
The memory 1020 may include one or more non-volatile memories and one or more volatile memories. Examples of the non-volatile memories include, but are not limited to, a Read Only Memory (ROM) 1024, an electrically programmable read only memory (EPROM) , a flash memory, a hard disk, a compact disc (CD) , a digital video disk (DVD) , an optical disk, a laser disk, and other magnetic storage and/or optical storage. Examples of the volatile memories include, but are not limited to, a random access memory (RAM) 1022 and other volatile memories that will not last in the power-down duration.
A computer program 1030 includes computer executable instructions that are executed by the associated processor 1010. The instructions of the program 1030 may include instructions for performing operations/acts of some example embodiments of the present disclosure. The program 1030 may be stored in the memory, e.g., the ROM 1024. The processor 1010 may perform any suitable actions and processing by loading the program 1030 into the RAM 1022.
The example embodiments of the present disclosure may be implemented by means of the program 1030 so that the device 1000 may perform any process of the disclosure as discussed above. The example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.
In some example embodiments, the program 1030 may be tangibly contained in a computer readable medium which may be included in the device 1000 (such as in the memory 1020) or other storage devices that are accessible by the device 1000. The device
1000 may load the program 1030 from the computer readable medium to the RAM 1022 for execution. In some example embodiments, the computer readable medium may include any types of non-transitory storage medium, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like. The term “non-transitory, ” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM) .
FIG. 11 shows an example of the computer readable medium 1100 which may be in form of CD, DVD or other optical storage disk. The computer readable medium 1100 has the program 1030 stored thereon.
Generally, various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device. Although various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer readable medium, such as a non-transitory computer readable medium. The computer program product includes computer-executable instructions, such as those included in program modules, being executed in a device on a target physical or virtual processor, to carry out any of the methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various embodiments. Machine-executable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.
Program code for carrying out methods of the present disclosure may be written
in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, cause the functions/operations specified in the flowcharts and/or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
In the context of the present disclosure, the computer program code or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above. Examples of the carrier include a signal, computer readable medium, and the like.
The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
Further, although operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Unless explicitly stated, certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, unless explicitly stated, various features that are described in the context of a single embodiment may also be implemented in a plurality of embodiments separately or in any suitable sub-combination.
Although the present disclosure has been described in languages specific to structural features and/or methodological acts, it is to be understood that the present disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims (23)
- An apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:detect a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene;determine a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment;transform, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; anddetermine a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- The apparatus of claim 1, wherein the at least one memory and the at least one processor further cause the apparatus to:determine an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
- The apparatus of claim 1 or 2, wherein the at least one memory and the at least one processor cause the apparatus to determine the perspective transformation and the distance scale for the target image by:selecting the perspective transformation model associated with the target deployment from a plurality of perspective transformation models based on a target deployment identity of the target deployment, the plurality of perspective transformation models being identified with respective deployment identities.
- The apparatus of claim 3, wherein the target deployment is identified based on a mounting height, a shooting orientation, and a type of the target camera.
- The apparatus of any of claims 1 to 4, wherein the at least one memory and the at least one processor cause the apparatus to detect the first location and the second location by:detecting the first location and the second location using an object detection model,wherein the object detection model is trained with first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images.
- The apparatus of any of claims 1 to 5, wherein the first location comprises a first plurality of coordinates of a first plurality of key points of the first object detected from the target image, and the second location comprises a second plurality of coordinates of a second plurality of key points of the second object detected from the target image, wherein the first plurality of key points at least comprises a first centroid of the first object, and the second plurality of key points at least comprises a second centroid of the second object.
- The apparatus of any of claims 1 to 6, wherein the perspective transformation model is trained using second sample images captured by a camera with the target deployment and second label information indicating at least one of the following for each second sample image:a labeled indication of whether a distance between the pair of objects is abnormal, the pair of objects being detected in the second sample image, orrespective labeled transformed locations of the pair of objects in the top plane view of the three-dimensional space.
- The apparatus of claim 7, wherein the perspective transformation model is trained by the following:determining respective predicted transformed locations of the pair of objects in the top plane view of the three-dimensional space using a predicted perspective transformation output by the perspective transformation model for the second sample image;determining a predicted distance between the pair of objects by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image;determine a predicted indication of whether a distance between the pair of objects is abnormal based on a comparison result between the predicted distance and a distance threshold;determining a first loss value of a first loss function based on a first error between the predicted indication and the labeled indication for the second sample image; andupdating the perspective transformation model by decreasing the first loss value of the first loss function.
- The apparatus of claim 7, wherein the perspective transformation model is further trained by the following:determining a second loss value of a second loss function based on second errors between the respective predicted transformed locations and the respective labeled transformed locations; andupdating the perspective transformation model by decreasing the second loss value of the second loss function.
- The apparatus of claim 7, wherein the respective labeled transformed locations of the pair of objects are determined by the following:determining a fitted perspective transformation associated with the target deployment based on respective locations of the pair of objects within a two-dimensional space corresponding to the second sample image;transforming, based on the fitted perspective transformation, the respective locations of the pair of objects into the respective transformed locations within the top plane view of the three-dimensional space.
- The apparatus of any of claims 1 to 10, wherein the target camera is a monocular camera.
- A method comprising:detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene;determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment;transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; anddetermining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- The method of claim 12, further comprising:determining an indication of whether the target distance between the first object and the second object is abnormal based on a comparison result between the target distance and a distance threshold.
- The method of claim 12 or 13, wherein determining the perspective transformation and the distance scale for the target image comprises:selecting the perspective transformation model associated with the target deployment from a plurality of perspective transformation models based on a target deployment identity of the target deployment, the plurality of perspective transformation models being identified with respective deployment identities.
- The method of claim 14, wherein the target deployment is identified based on a mounting height, a shooting orientation, and a type of the target camera.
- The method of any of claims 12 to 15, wherein detecting the first location and the second location comprises:detecting the first location and the second location using an object detection model,wherein the object detection model is trained with first sample images captured by a plurality of cameras with a plurality of deployments and first label information indicating objects in the first sample images.
- The method of any of claims 12 to 16, wherein the first location comprises a first plurality of coordinates of a first plurality of key points of the first object detected from the target image, and the second location comprises a second plurality of coordinates of a second plurality of key points of the second object detected from the target image, wherein the first plurality of key points at least comprises a first centroid of the first object, and the second plurality of key points at least comprises a second centroid of the second object.
- The method of any of claims 12 to 17, wherein the perspective transformation model is trained using second sample images captured by a camera with the target deployment and second label information indicating at least one of the following for each second sample image:a labeled indication of whether a distance between the pair of objects is abnormal, the pair of objects being detected in the second sample image, orrespective labeled transformed locations of the pair of objects in the top plane view of the three-dimensional space.
- The method of claim 18, wherein the perspective transformation model is trained by the following:determining respective predicted transformed locations of the pair of objects in the top plane view of the three-dimensional space using a predicted perspective transformation output by the perspective transformation model for the second sample image;determining a predicted distance between the pair of objects by scaling a distance between the pair of transformed locations with a predicted distance scale output by the perspective transformation model for the second sample image;determine a predicted indication of whether a distance between the pair of objects is abnormal based on a comparison result between the predicted distance and a distance threshold;determining a first loss value of a first loss function based on a first error between the predicted indication and the labeled indication for the second sample image; andupdating the perspective transformation model by decreasing the first loss value of the first loss function.
- The method of claim 18, wherein the perspective transformation model is further trained by the following:determining a second loss value of a second loss function based on second errors between the respective predicted transformed locations and the respective labeled transformed locations; andupdating the perspective transformation model by decreasing the second loss value of the second loss function.
- The method of claim 18, wherein the respective labeled transformed locations of the pair of objects are determined by the following:determining a fitted perspective transformation associated with the target deployment based on respective locations of the pair of objects within a two-dimensional space corresponding to the second sample image;transforming, based on the fitted perspective transformation, the respective locations of the pair of objects into the respective transformed locations within the top plane view of the three-dimensional space.
- An apparatus comprising:means for detecting a first location of a first object and a second location of a second object within a two-dimensional space corresponding to a target image captured by a camera with a target deployment in a scene;means for determining a perspective transformation and a distance scale for the target image by providing the target image into a perspective transformation model associated with the target deployment;means for transforming, based on the determined perspective transformation, the first location and the second location into a first transformed location of the first object and a second transformed location of the second object within a top plane view of a three-dimensional space, respectively; andmeans for determining a target distance between the first object and the second object in a three-dimensional space by scaling a distance between the first transformed location and the second transformed location with the distance scale.
- A computer readable medium comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least the method of any of claims 12 to 21.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/095578 WO2024239207A1 (en) | 2023-05-22 | 2023-05-22 | Three-dimensional object distance estimation |
| CN202380099759.7A CN121420321A (en) | 2023-05-22 | 2023-05-22 | 3D object distance estimation |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/095578 WO2024239207A1 (en) | 2023-05-22 | 2023-05-22 | Three-dimensional object distance estimation |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024239207A1 true WO2024239207A1 (en) | 2024-11-28 |
Family
ID=93588667
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/095578 Ceased WO2024239207A1 (en) | 2023-05-22 | 2023-05-22 | Three-dimensional object distance estimation |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121420321A (en) |
| WO (1) | WO2024239207A1 (en) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102254318A (en) * | 2011-04-08 | 2011-11-23 | 上海交通大学 | Method for measuring speed through vehicle road traffic videos based on image perspective projection transformation |
| WO2021110497A1 (en) * | 2019-12-04 | 2021-06-10 | Valeo Schalter Und Sensoren Gmbh | Estimating a three-dimensional position of an object |
| CN113836964A (en) * | 2020-06-08 | 2021-12-24 | 北京图森未来科技有限公司 | Method and device for detecting lane line corner |
| CN114663507A (en) * | 2022-03-16 | 2022-06-24 | 深圳市商汤科技有限公司 | Distance detection method and device, electronic equipment and computer storage medium |
-
2023
- 2023-05-22 WO PCT/CN2023/095578 patent/WO2024239207A1/en not_active Ceased
- 2023-05-22 CN CN202380099759.7A patent/CN121420321A/en active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102254318A (en) * | 2011-04-08 | 2011-11-23 | 上海交通大学 | Method for measuring speed through vehicle road traffic videos based on image perspective projection transformation |
| WO2021110497A1 (en) * | 2019-12-04 | 2021-06-10 | Valeo Schalter Und Sensoren Gmbh | Estimating a three-dimensional position of an object |
| CN113836964A (en) * | 2020-06-08 | 2021-12-24 | 北京图森未来科技有限公司 | Method and device for detecting lane line corner |
| CN114663507A (en) * | 2022-03-16 | 2022-06-24 | 深圳市商汤科技有限公司 | Distance detection method and device, electronic equipment and computer storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121420321A (en) | 2026-01-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Yuan et al. | Keypoints-based deep feature fusion for cooperative vehicle detection of autonomous driving | |
| US10817752B2 (en) | Virtually boosted training | |
| KR102217549B1 (en) | Method and system for soar photovoltaic power station monitoring | |
| US20190291723A1 (en) | Three-dimensional object localization for obstacle avoidance using one-shot convolutional neural network | |
| US20220335258A1 (en) | Systems and methods for dataset and model management for multi-modal auto-labeling and active learning | |
| US11882262B2 (en) | System and method for stereoscopic image analysis | |
| CN113887400B (en) | Obstacle detection method, model training method and device and automatic driving vehicle | |
| CN112528974B (en) | Distance measuring method and device, electronic equipment and readable storage medium | |
| US20220327676A1 (en) | Method and system for detecting change to structure by using drone | |
| Zhang et al. | Improved feature point extraction method of ORB-SLAM2 dense map | |
| WO2024087962A1 (en) | Truck bed orientation recognition system and method, and electronic device and storage medium | |
| WO2019111932A1 (en) | Model learning device, model learning method, and recording medium | |
| Guo et al. | Hawkdrive: A transformer-driven visual perception system for autonomous driving in night scene | |
| US20230401748A1 (en) | Apparatus and methods to calibrate a stereo camera pair | |
| Wang et al. | 3D-LIDAR based branch estimation and intersection location for autonomous vehicles | |
| CN117876901A (en) | A fixed-point landing method based on UAV visual perception | |
| CN120411223B (en) | An efficient method for positioning photovoltaic inspection results using drones | |
| EP4564293A1 (en) | Method of training a neural network for vision-based tracking and assiociated apparatus and system | |
| Lezki et al. | Joint exploitation of features and optical flow for real-time moving object detection on drones | |
| Mahmudah et al. | Digital twin: challenge road damage detection on edge device | |
| CN121420321A (en) | 3D object distance estimation | |
| Xie et al. | Accurate localization of moving objects in dynamic environment for small unmanned aerial vehicle platform using global averaging | |
| CN117475397B (en) | Target annotation data acquisition method, medium and device based on multi-mode sensor | |
| US20260011113A1 (en) | Information processing apparatus, information processing method, and information generation method | |
| US20260120309A1 (en) | Extrinsic parameter prediction for image sensor(s) |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23937884 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |