WO2019101720A1 - Methods for scene classification of an image in a driving support system - Google Patents
Methods for scene classification of an image in a driving support system Download PDFInfo
- Publication number
- WO2019101720A1 WO2019101720A1 PCT/EP2018/081874 EP2018081874W WO2019101720A1 WO 2019101720 A1 WO2019101720 A1 WO 2019101720A1 EP 2018081874 W EP2018081874 W EP 2018081874W WO 2019101720 A1 WO2019101720 A1 WO 2019101720A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- features
- scene
- dbm
- image
- layer
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/088—Non-supervised learning, e.g. competitive learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
- G06F18/2411—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on the proximity to a decision surface, e.g. support vector machines
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/047—Probabilistic or stochastic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/0895—Weakly supervised learning, e.g. semi-supervised or self-supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/46—Descriptors for shape, contour or point-related descriptors, e.g. scale invariant feature transform [SIFT] or bags of words [BoW]; Salient regional features
- G06V10/462—Salient features, e.g. scale invariant feature transforms [SIFT]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/35—Categorising the entire scene, e.g. birthday party or wedding scene
- G06V20/38—Outdoor scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S13/00—Systems using the reflection or reradiation of radio waves, e.g. radar systems; Analogous systems using reflection or reradiation of waves whose nature or wavelength is irrelevant or unspecified
- G01S13/88—Radar or analogous systems specially adapted for specific applications
- G01S13/89—Radar or analogous systems specially adapted for specific applications for mapping or imaging
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S17/00—Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
- G01S17/88—Lidar systems specially adapted for specific applications
- G01S17/93—Lidar systems specially adapted for specific applications for anti-collision purposes
- G01S17/931—Lidar systems specially adapted for specific applications for anti-collision purposes of land vehicles
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S13/00—Systems using the reflection or reradiation of radio waves, e.g. radar systems; Analogous systems using reflection or reradiation of waves whose nature or wavelength is irrelevant or unspecified
- G01S13/88—Radar or analogous systems specially adapted for specific applications
- G01S13/93—Radar or analogous systems specially adapted for specific applications for anti-collision purposes
- G01S13/931—Radar or analogous systems specially adapted for specific applications for anti-collision purposes of land vehicles
- G01S2013/9323—Alternative operation using light waves
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S13/00—Systems using the reflection or reradiation of radio waves, e.g. radar systems; Analogous systems using reflection or reradiation of waves whose nature or wavelength is irrelevant or unspecified
- G01S13/88—Radar or analogous systems specially adapted for specific applications
- G01S13/93—Radar or analogous systems specially adapted for specific applications for anti-collision purposes
- G01S13/931—Radar or analogous systems specially adapted for specific applications for anti-collision purposes of land vehicles
- G01S2013/9324—Alternative operation using ultrasonic waves
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S7/00—Details of systems according to groups G01S13/00, G01S15/00, G01S17/00
- G01S7/02—Details of systems according to groups G01S13/00, G01S15/00, G01S17/00 of systems according to group G01S13/00
- G01S7/41—Details of systems according to groups G01S13/00, G01S15/00, G01S17/00 of systems according to group G01S13/00 using analysis of echo signal for target characterisation; Target signature; Target cross-section
- G01S7/417—Details of systems according to groups G01S13/00, G01S15/00, G01S17/00 of systems according to group G01S13/00 using analysis of echo signal for target characterisation; Target signature; Target cross-section involving the use of neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
- G06N20/10—Machine learning using kernel methods, e.g. support vector machines [SVM]
Definitions
- the invention relates to a method for scene classification of an image in image processing in a driving support system of an automotive vehicle.
- the invention further relates to a method for scene classification of an image in image processing in a driving support system of an automotive vehicle.
- Driving support systems such as driver assistance systems are systems developed to automate, adapt and enhance vehicle systems for safety and better driving.
- Safety features are designed to avoid collisions and accidents by offering technologies that alert the driver to potential problems, or to avoid collisions by implementing safeguards and taking over control of the vehicle.
- the driving support systems provide input to perform a control of the vehicle.
- Adaptive features may automate lighting, provide adaptive cruise control, automate braking, incorporate traffic warnings, connect to smartphones, alert e.g. the driver in respect to other cars or different types of dangers, keep the vehicle in the correct lane, or show what is located in blind spots.
- Driving support systems including the aforementioned driver assistance systems, frequently rely on inputs from multiple data sources such as automotive imaging, image processing, radar sensors, LiDAR, ultrasonic sensors and other sources.
- Neural networks have recently been implicated in processing such inputs of data within driver assistance systems, or in general in driving support systems.
- DBMs Deep Boltzmann Machines
- CNNs convolution neural networks
- a Deep Boltzmann Machine is a stochastic Hopfield net with hidden layers.
- Hopfield net is an energy-based model. Whereas the Hopfield net is used as a content- addressable memory system, the Boltzmann Machine learns to represent its inputs. It is a generative model, which means that it learns the joint probability distribution of all its inputs. Once the Boltzmann Machine has learned its input (i.e. when it has reached thermal equilibrium), the configuration of weights at the (multiple) hidden layers constitute a representation of the inputs presented at the visible layer.
- RBMs are Restricted Boltzmann Machines, wherein the restriction is that the neurons form a bipartite graph with no intra layer connections. This restriction allows the use of the highly efficient Contrastive
- a Deep Boltzmann Machine is a stack of RBMs.
- a DBN Deep Belief Net also contains RBMs, but they have RBMs only in the top two layers and the layers below are sigmoid belief nets which are directed graphical models. In contrast, the DBM is an undirected graphical model through and through.
- CNNs Convolutional neuronal networks
- a convolutional neural network is a class of deep, feed-forward artificial neural networks that has successfully been applied to analyzing visual imagery.
- CNNs use a variation of multilayer perceptrons designed to require minimal preprocessing.
- Convolutional networks were inspired by biological processes, in which the connectivity pattern between neurons is inspired by the organization of the animal visual cortex. Individual cortical neurons respond to stimuli only in a restricted region of the visual field known as the receptive field. The receptive fields of different neurons partially overlap such that they cover the entire visual field.
- CNNs use relatively little pre-processing compared to other image classification algorithms. This means that the network learns the filters that in traditional algorithms were hand- engineered. This independence from prior knowledge and human effort in feature design is a major advantage. CNNs have applications in image and video recognition, recommender systems and natural language processing.
- scene classification may be performed e.g. based on distinguishing between one or all of the following three categories:
- the above classification can be used by a layer that runs across all algorithms in a computer vision product.
- the classification can thus be used:
- 3DOD Dimensional Object Detection
- CPU, Memory can have an algorithm for sparsely populated scenes, which runs most times and consumes fewer resources (CPU, Memory), and can further have an intensive variant for densely populated scenes. So, if the“master / supervisor algorithm” knows that the scene is densely or sparsely populated, it can activate the corresponding variant of the 3DOD algorithm.
- snow and rain are well-known difficult conditions for Computer Vision algorithms. It is made all the more difficult for the algorithm, because it has to have the same configuration and learning parameters for sunny weather as for snowy conditions.
- the supervisor algorithm knows that the scene is snowy, raining or sunny, it can activate different variants of 3DOD, Pedestrian Detection (PD), Parking Slot Marker Detection (PSMD) and so on, while each of those variants only learns to handle one weather condition.
- PD Pedestrian Detection
- PSMD Parking Slot Marker Detection
- US 2007/0282506 A1 discloses a method for image processing for vehicular applications that determines edges of objects in images and feeds these data to a neural network that can provide a classification, identification and/or location of an object.
- the method comprises the steps of obtaining information about objects in an environment in or around a vehicle comprising training pattern recognition algorithm, such as a neural network, to provide information about objects in the environment upon receiving as input information about edges of unknown objects, installing the pattern recognition algorithm in a processor on the vehicle, operatively obtaining images of the environment deriving data about edges of objects in the obtained images and providing the data to the pattern recognition algorithm in the processor to receive as output, information about the object, such as a classification, identification and/or location of an object.
- training pattern recognition algorithm such as a neural network
- US 2008/0144944 A1 discloses a method for obtaining information about an occupying item in a vehicle, such as a human being, comprising obtaining images of an area above a seat in the vehicle in which the occupying item is situated and classifying the occupying item by inputting signals derived from the images into a trained neural network form which is trained to output an indication of the class of occupying item from one of a predetermined number of possible classes.
- the images may be pre-processed to remove background portions of the images and then converted into signals for input into the neural network form.
- Driving support systems like driver assistance systems are one of the fastest-growing segments in automotive electronics and there is a need for improved methods and systems for image processing in driving support assistance systems.
- the invention provides a method for scene classification of an image in image processing in a driving support system of an automotive vehicle comprising the steps of:
- DBM Deep Boltzmann Machine
- RBM Restricted Boltzmann Machines
- RBF-SVM Radial Basis Filter Support Vector Machine
- DBM Deep Boltzmann Machine
- this embodiment of the invention to combine the following three main steps in a unique manner: spatially ordering regions of the image, modelling the underlying joint probability distribution with a generative model, i.e. the Deep Boltzmann Machine (DBM), and then classifying scenes based on that generative model.
- a generative model i.e. the Deep Boltzmann Machine (DBM)
- DBM Deep Boltzmann Machine
- One advantage of the invention is no loss of region order (i.e. discarding improbable explanations for a scene), which contains a lot of information useful for classifying scenes, e.g.“sky above ground”,“road below sky” and“tree top above road”.
- the human brain uses such information to understand a scene and to discard the improbable explanations for what is being seen.
- the invention uses an unsupervised, generative model, i.e. the Deep Boltzmann Machine (DBM), which provides the advantage of reducing the amount of labeled data necessary. Labeled data are only used to fine-tune the DBM. Thus, very little annotated data is required, which will significantly reduce cost and reduce effort of annotation.
- DBM Deep Boltzmann Machine
- a further advantage of the invention making use of a DBM is that the method comprises more applicability to a broader range of tasks, which means that one does not have to go through the expensive step of annotating images for applying this to another task, e.g. segmentation.
- transformations e.g. lighting, perspective, and occlusion
- a scene can be a motorway at night, which can be composed of a lot of combinations of regions and of features, ii) there is also a lot of overlap among various kinds of scenes and these overlaps are not as well captured through a human-made annotation as they are through learning the underlying probability distributions of various scenes, iii) the annotated data at hand for scene classification is not likely to have enough representation.
- the spatially ordering regions of the image comprises using one or more region descriptors to capture a semantically uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, and iii) Co-occurrence features that are Haralick features derived from the Gray Level Co-occurrence Matrix (GLCM).
- region descriptors to capture a semantically uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, and iii) Co-occurrence features that are Haralick features derived from the Gray Level Co-occurrence Matrix (GLCM).
- the captured low-level representation of each region is semantically uncorrelated in that the above used features i), ii) and iii) are mutually uncorrelated in a semantic sense.
- Gabor filter is a linear filter used for texture analysis, which analyses whether there is any specific frequency content in the image in specific directions in a localized region around the point or region of analysis. Frequency and orientation representations of Gabor filters are proven to be a feature used by the human vision.
- the ii) HSV color space is a color space that defines the localization of a color by the features Hue, Saturation and Value.
- the iii) Co-occurrence features can be used for intra-region co-occurrence statistics (mean and range), wherein said features are selected from the group consisting of Angular Second Moment, Contrast, Sum Average, Sum Variance and Difference Variance.
- the driving support systems include driver assistance systems are systems, which are already known and used in state of the Art vehicles.
- the developed driving support systems are provided to automate, adapt and enhance vehicle systems for safety and better driving.
- Safety features are designed to avoid collisions and accidents by offering technologies that alert the driver to potential problems, or to avoid collisions by implementing safeguards and taking over control of the vehicle.
- the driving support systems provide input to perform a control of the vehicle.
- Adaptive features may automate lighting, provide adaptive cruise control, automate braking, incorporate traffic warnings, connect to smartphones, alert e.g. the driver in respect to other cars or different types of dangers, keep the vehicle in the correct lane, or show what is located in blind spots.
- Driving support systems including the aforementioned driver assistance systems frequently rely on inputs from multiple data sources such as automotive imaging, image processing, radar sensors, LiDAR, ultrasonic sensors and other sources.
- spatially ordering regions of the image further comprises adding spatial relationships among neighboring regions to create a Spatially Ordered Region Descriptor (SORD), wherein further Haralick features are used that are selected from the group consisting of the mean and range of iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
- SORD Spatially Ordered Region Descriptor
- the Spatially Ordered Region Descriptor comprises the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, iii) Co-occurrence features, iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation. That means that the aforementioned eight features are added to the Region Descriptor to convert it into the SORD.
- SORD Spatially Ordered Region Descriptor
- DBM Deep Boltzmann Machine
- RBMs Restricted Boltzmann Machines
- classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs). That means that the DBM preferably is topped off with a softmax layer in order to perform the classification of the scene.
- RBMs Restricted Boltzmann Machines
- the softmax layer added on top does the actual classification.
- a RBF-SVM or a CNN can also be used for the classification. This final classification layer requires a far smaller training set because its weights are initialized by the output of the DBM and hence require only fine-tuning for the specific scene categories in the intended application.
- the invention further provides a method for scene classification of an image in image processing in a driving support system of an automotive vehicle comprising the steps of:
- CNN convolutional neuronal network
- DBM Deep Boltzmann Machine
- RBM Restricted Boltzmann Machine
- RBF- SVM Radial Basis Filter Support Vector Machine
- CNN convolutional neuronal network
- DBM Deep Boltzmann Machine
- a CNN with a DBM in a method for scene classification of an image in image processing in a driving support system of an automotive vehicle is unique and provides several advantages. It is a particular advantage of this embodiment of the invention that features used as input to the DBM are not hand-crafted but discovered by the CNN as the most useful features, which significantly reduces cost and efforts.
- this embodiment of the invention uses a single CNN at multiple resolutions. The features at higher resolution capture detail, whereas the features learned at lower resolution capture the “big picture”, i.e. the region-level and scene-level information in the image. The single CNN provides these features at multiple resolutions or scales. Since the lower resolution images capture the“big picture”, a CNN trained at that resolution and performing inference at that low resolution will provide the region-level and scene-level features to the DBM, which then performs the scene classification in its softmax layer that is added on top.
- the DBM has advantages over a DBN (Deep Belief Network) in that the DBN is a DAG (Directed Acyclic Graphical model), whereas the DBM is an undirected graphical model.
- DBN Deep Belief Network
- the approximate inference procedure in DBMs in addition to an initial bottom-up pass, can incorporate top-down feedback, allowing DBMs to better propagate uncertainty about and hence deal more robustly with ambiguous inputs.
- this embodiment of the invention can achieve fast
- each layer of hidden units can be activated in a single bottom-up pass by doubling the bottom-up input to compensate for the lack of top-down feedback (except for the very top layer, which does not have a top-down input).
- This fast approximate inference is used to initialize the mean- field method, which converges much faster than with random initialization.
- scene classification is realized by using a CNN to learn what features are most useful for the classification of scenes of an image, wherein a CNN of multiple resolutions is used for capturing detail and“big picture” as separate inputs followed by modeling the joint probability distribution of each scene category using a DBM, which is then followed by classifying the scenes based on the learning result of the DBM that has been generated based on the inputs of the separate layers of the CNN.
- this embodiment of the invention can be described as a Hybrid CNN-DBM model, in which a DBM learns the inter-relationships among regions and multi-resolution features of the same region with an unsupervised, generative method based on the separate inputs of the various layers of the CNN. That means that the DBM learns an internal representation of the scene category using features at multiple resolutions extracted from the image that are fed by each layer of the CNN to the visible layer of the DBM as separate inputs.
- each layer’s output of the CNN is fed separately to the visible layer of the DBM, which is then followed by the stack of Restricted Boltzmann Machines (RBMs) learning the underlying joint probability distribution of each scene category followed by classifying the scene based on the learning result of the DBM.
- RBMs Restricted Boltzmann Machines
- This architecture of this Hybrid CNN-DBM model of the invention allows the DBM to not only learn the inter-relationships among regions in a scene, but also to learn the inter-relationships among multi-resolution features of the same regions. This is a major advantage of this embodiment of the invention over purely discriminative networks.
- the primary modeling of scene categories is done by an unsupervised, generative method, i.e. the DBM, thus
- the DBM learns the joint probability distribution of all its inputs.
- the DBM inputs are composed of the discriminative features learned at various abstraction levels. So, the DBM not only learns the inter-relationships among regions in a scene, but also the inter-relationships among the multi-resolution features of the same region;
- the scene can be a motorway at night, which can be composed of a lot of combinations of regions and of features;
- the convolutional neuronal network is pre-trained using supervised training and labeled data, wherein classifying the scene is done by a temporary softmax layer as the final layer of the Deep Boltzmann Machine (DBM), wherein the temporary softmax layer is stripped off once the convolutional neuronal network (CNN) has learned the features, followed by feeding the output of each layer of the convolutional neuronal network (CNN) into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input.
- DBM Deep Boltzmann Machine
- the Deep Boltzmann Machine is pre-trained using greedy layer-wise pre-training to learn the internal representation of the combination of multiple features in a scene and multi-resolution features of the same region, followed by adding on and pre-training the softmax layer using labelled data.
- classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs).
- the DBM preferably is topped off with a softmax layer, where classifying the scene based on the learning result of the DBM is performed, once the stack of RBMs has learned the underlying joint probability distributions of each scene category.
- a RBF-SVM or a CNN can also be used for the classification.
- This final classification layer requires a far smaller training set because its weights are initialized by the output of the DBM and hence require only fine- tuning for the specific scene categories in the intended application.
- the invention also provides the use of the methods described herein in a driving support system of an automotive vehicle. Specifically, the invention provides the use of the methods for scene classification of an image in image processing in a driving support system of an automotive vehicle described above.
- the invention further provides a driving support system for an automotive vehicle comprising a camera for providing images for classification, wherein the driving support system is configured for performing the methods described herein.
- the invention further provides a non-transitory computer-readable medium, comprising instructions stored thereon, that when executed on a processor, induce a driving support system to perform the methods described herein.
- the invention also provides an automotive vehicle comprising:
- non-transitory computer-readable medium comprising instructions stored thereon, that when executed on a processor, induce a driving support system to perform the methods described herein, and
- a driving support system for an automotive vehicle comprising a camera for providing images for classification, wherein the driving support system is configured for performing the methods described herein.
- Fig. 1 schematically depicts an automotive vehicle with a driving support system and a camera according to a first, preferred embodiment of the invention
- Fig. 2 schematically depicts the classification of an image in an image processing in a driving support system of the automotive vehicle according to the first embodiment of the invention
- Fig. 3 schematically depicts a second embodiment of the classification of an image in
- Fig. 1 schematically depicts an automotive vehicle 1 with a driving support system 2 and a camera 3 according to a first, preferred embodiment of the invention.
- the camera 3 provides images of a scene, e.g. a motorway at night, for classification by the driving support system 2, which is configured for performing the methods for scene classification of an image in image processing, as described herein.
- Fig. 2 schematically depicts the classification of an image in an image processing in a driving support system 2 of an automotive vehicle 1 according to a preferred embodiment of the invention.
- the image of a scene, e.g. a motorway at night, which is taken by the camera 3 is spatially ordered into regions by clustering the image pixels into regions with high inter-class variance and low intra-class variance.
- Spatially ordering regions of the image preferably is done by using region descriptors to capture a semantically uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features and iii) Co occurrence features that are Haralick features derived from the Gray Level Co-occurrence Matrix (GLCM, intra-region and inter-region).
- the captured low-level representation of each region is semantically uncorrelated in that the above used features i), ii) and iii) are mutually uncorrelated in a semantic sense.
- the i) Gabor filter is a linear filter used for texture analysis, which analyses whether there is any specific frequency content in the image in specific directions in a localized region around the point or region of analysis. Frequency and orientation representations of Gabor filters are proven to be a feature used by the human vision.
- the ii) HSV color space is a color space that defines the localization of a color by the features Hue, Saturation and Value.
- Co-occurrence features are used for intra-region co-occurrence statistics (mean and range), wherein said features are selected from the group consisting of Angular Second Moment, Contrast, Sum Average, Sum Variance and Difference Variance.
- Spatially ordering the regions of the image is further done by adding spatial relationships among neighboring regions to create the Spatially Ordered Region Descriptor (SORD), wherein further Haralick features are used that are selected from the group consisting of the mean and range of iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
- SORD Spatially Ordered Region Descriptor
- the SORD that is created in this way comprises the features i) Gabor filter, ii) Hue,
- HSV Saturation and Value
- the SORD is then fed into the neurons of the visible layer (having e.g. 1024 units) of the DBM followed by processing through the hidden layers (a Hidden Layer 1 having e.g. 512 units and a Hidden Layer 2 having e.g. 256 units) of the DBM for modelling the underlying joint probability distribution of each scene category.
- the hidden layers of the DBM are a stack of RBMs.
- the final step of classifying the scene is then performed by the final softmax layer having e.g. 1000 units on top of the hidden layers (stack of RBMs) of the DBM. That means that the DBM is topped off with a softmax layer in order to perform the classification of the scene.
- the softmax layer added on top does the actual classification.
- Fig. 3 schematically depicts a second embodiment of the classification of an image in image processing in a driving support system 2 of an automotive vehicle 1 that is based on a Hybrid-CNN-DBM model.
- the image of a scene e.g. a motorway at night, that is taken by the camera 3 is fed into a CNN comprising a plurality of layers (Layer 1 , Layer 2, ..., Layer n) to learn what features of the image are most useful for the classification of scenes.
- this embodiment of the invention uses a single CNN at multiple resolutions.
- the features at higher resolution capture detail, whereas the features learned at lower resolution capture the“big picture”, i.e. the region-level and scene-level information in the image.
- the single CNN provides these features at multiple resolutions or scales. Since the lower resolution images capture the“big picture”, the CNN trained at that resolution and performing inference at that low resolution provides the region-level and scene-level features for the DBM.
- the hidden layers of the DBM are a stack of RBMs, which learn the underlying joint probability distribution of each scene category.
- the CNN is pre-trained using supervised training and labeled data, wherein classifying the scene is done by a temporary softmax layer as the final layer of the DBM, wherein the temporary softmax layer is stripped off once the CNN has learned the features, followed by feeding the output of each layer of the CNN into the visible layer of the DBM as a separate input.
- the DBM is pre-trained using greedy layer-wise pre-training to learn the internal
- the final step of classifying the scene is then performed by the final softmax layer on top of the hidden layers (stack of RBMs) of the DBM.
- the DBM preferably is topped off with a softmax layer, where classifying the scene based on the learning result of the DBM is performed, once the stack of RBMs has learned the underlying joint probability
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Life Sciences & Earth Sciences (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Computing Systems (AREA)
- Molecular Biology (AREA)
- Computational Linguistics (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Probability & Statistics with Applications (AREA)
- Image Analysis (AREA)
- Traffic Control Systems (AREA)
Abstract
The invention relates to a method for scene classification of an image in image processing in a driving support system (2) of an automotive vehicle (1) comprising the steps of spatially ordering regions of the image by clustering the image pixels into regions with high inter-class variance and low intra-class variance, modelling the underlying joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) as a generative model comprising a stack of Restricted Boltzmann Machines (RBMs) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF-SVM), a second CNN or simply a softmax layer as the final layer, and classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).The invention further relates to a method for scene classification of an image in image processing in a driving support system (2) of an automotive vehicle (1) comprising the steps of providing a convolutional neuronal network (CNN) comprising a plurality of layers to learn what features of the image are most useful for the classification of scenes, wherein multiple resolutions of features are generated for capturing detail of features at higher resolution and "big picture" at lower resolution, modelling the joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) comprising a stack of Restricted Boltzmann Machines (RBMs) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF-SVM), a second CNN or simply a softmax layer as the final layer, wherein the output of each layer of the convolutional neuronal network (CNN) is fed into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input, followed by the stack of Restricted Boltzmann Machines (RBM) learning the underlying joint probability distribution of each scene category, and classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).
Description
Methods for scene classification of an image in a driving support system
The invention relates to a method for scene classification of an image in image processing in a driving support system of an automotive vehicle.
The invention further relates to a method for scene classification of an image in image processing in a driving support system of an automotive vehicle.
Driving support systems such as driver assistance systems are systems developed to automate, adapt and enhance vehicle systems for safety and better driving. Safety features are designed to avoid collisions and accidents by offering technologies that alert the driver to potential problems, or to avoid collisions by implementing safeguards and taking over control of the vehicle. In autonomous vehicles, the driving support systems provide input to perform a control of the vehicle. Adaptive features may automate lighting, provide adaptive cruise control, automate braking, incorporate traffic warnings, connect to smartphones, alert e.g. the driver in respect to other cars or different types of dangers, keep the vehicle in the correct lane, or show what is located in blind spots. Driving support systems including the aforementioned driver assistance systems, frequently rely on inputs from multiple data sources such as automotive imaging, image processing, radar sensors, LiDAR, ultrasonic sensors and other sources. Neural networks have recently been implicated in processing such inputs of data within driver assistance systems, or in general in driving support systems.
In recent times, there has been a surge of research on Deep Boltzmann Machines (DBMs) and convolution neural networks (CNNs). Their design has been aided by increase in computational power in computer architectures and the availability of large annotated datasets.
A Deep Boltzmann Machine (DBM) is a stochastic Hopfield net with hidden layers. A
Hopfield net is an energy-based model. Whereas the Hopfield net is used as a content- addressable memory system, the Boltzmann Machine learns to represent its inputs. It is a generative model, which means that it learns the joint probability distribution of all its inputs. Once the Boltzmann Machine has learned its input (i.e. when it has reached thermal equilibrium), the configuration of weights at the (multiple) hidden layers constitute a representation of the inputs presented at the visible layer. RBMs are Restricted Boltzmann
Machines, wherein the restriction is that the neurons form a bipartite graph with no intra layer connections. This restriction allows the use of the highly efficient Contrastive
Divergence algorithm for training. A Deep Boltzmann Machine (DBM) is a stack of RBMs. A DBN (Deep Belief Net) also contains RBMs, but they have RBMs only in the top two layers and the layers below are sigmoid belief nets which are directed graphical models. In contrast, the DBM is an undirected graphical model through and through.
Convolutional neuronal networks (CNNs) are highly successful at classification and categorization tasks, but much of the research is on standard photometric RGB images and is not focused on embedded automotive devices. Automotive hardware devices need to have low power consumption requirements and thus low computational power.
In machine learning, a convolutional neural network is a class of deep, feed-forward artificial neural networks that has successfully been applied to analyzing visual imagery. CNNs use a variation of multilayer perceptrons designed to require minimal preprocessing. Convolutional networks were inspired by biological processes, in which the connectivity pattern between neurons is inspired by the organization of the animal visual cortex. Individual cortical neurons respond to stimuli only in a restricted region of the visual field known as the receptive field. The receptive fields of different neurons partially overlap such that they cover the entire visual field.
CNNs use relatively little pre-processing compared to other image classification algorithms. This means that the network learns the filters that in traditional algorithms were hand- engineered. This independence from prior knowledge and human effort in feature design is a major advantage. CNNs have applications in image and video recognition, recommender systems and natural language processing.
For the methods described herein, scene classification may be performed e.g. based on distinguishing between one or all of the following three categories:
a. Types of scenes
i. Country-side
ii. City
iii. Parking lot in the open
iv. Parking lot in the basement of a mall
b. Weather conditions
i. Snow
ii. Sunny
c. Density of scenes
i. Sparse
ii. Dense / Busy Scene.
The above classification can be used by a layer that runs across all algorithms in a computer vision product. The classification can thus be used:
a) to determine the activation logic of an algorithm variant. For instance3
Dimensional Object Detection (3DOD) can have an algorithm for sparsely populated scenes, which runs most times and consumes fewer resources (CPU, Memory), and can further have an intensive variant for densely populated scenes. So, if the“master / supervisor algorithm” knows that the scene is densely or sparsely populated, it can activate the corresponding variant of the 3DOD algorithm.
b) In addition, snow and rain are well-known difficult conditions for Computer Vision algorithms. It is made all the more difficult for the algorithm, because it has to have the same configuration and learning parameters for sunny weather as for snowy conditions. However, if the supervisor algorithm knows that the scene is snowy, raining or sunny, it can activate different variants of 3DOD, Pedestrian Detection (PD), Parking Slot Marker Detection (PSMD) and so on, while each of those variants only learns to handle one weather condition.
c) Similarly, for a Parking Slot Marker Detection algorithm, parking slots in the open are very different from parking slots in an underground basement with artificial lights all around. With a supervisor algorithm and a Scene Classification algorithm to guide it, the PSMD algorithm then only needs to learn one scenario and learn it well.
In this respect, US 2007/0282506 A1 discloses a method for image processing for vehicular applications that determines edges of objects in images and feeds these data to a neural network that can provide a classification, identification and/or location of an object. The method comprises the steps of obtaining information about objects in an environment in or around a vehicle comprising training pattern recognition algorithm, such as a neural network, to provide information about objects in the environment upon receiving as input information about edges of unknown objects, installing the pattern recognition algorithm in a processor on the vehicle, operatively obtaining images of the environment deriving data about edges of objects in the obtained images and providing the data to the pattern recognition algorithm in the processor to receive as output, information about the object, such as a classification, identification and/or location of an object.
US 2008/0144944 A1 discloses a method for obtaining information about an occupying item in a vehicle, such as a human being, comprising obtaining images of an area above a seat in the vehicle in which the occupying item is situated and classifying the occupying item by inputting signals derived from the images into a trained neural network form which is trained to output an indication of the class of occupying item from one of a predetermined number of possible classes. The images may be pre-processed to remove background portions of the images and then converted into signals for input into the neural network form.
Driving support systems like driver assistance systems are one of the fastest-growing segments in automotive electronics and there is a need for improved methods and systems for image processing in driving support assistance systems.
It is an objective of the present invention to provide methods for classifying scenes in driving support systems more accurately than current methods and classifying scenes in a better way in order to eliminate hand-crafted features being used as inputs.
This objective is addressed by the subject-matter of the independent claims. Preferred embodiments are described in the sub claims.
The invention provides a method for scene classification of an image in image processing in a driving support system of an automotive vehicle comprising the steps of:
- spatially ordering regions of the image by clustering the image pixels into regions with high inter-class variance and low intra-class variance,
- modelling the underlying joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) as a generative model comprising a stack of Restricted Boltzmann Machines (RBM) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF-SVM), a second CNN or simply a softmax layer as the final layer, and
- classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).
Thus, it is an essential idea of this embodiment of the invention to combine the following three main steps in a unique manner: spatially ordering regions of the image, modelling the underlying joint probability distribution with a generative model, i.e. the Deep Boltzmann Machine (DBM), and then classifying scenes based on that generative model. One advantage of the invention is no loss of region order (i.e. discarding improbable explanations for a scene), which contains a lot of information useful for classifying scenes, e.g.“sky above
ground”,“road below sky” and“tree top above road”. The human brain uses such information to understand a scene and to discard the improbable explanations for what is being seen.
In addition, the invention uses an unsupervised, generative model, i.e. the Deep Boltzmann Machine (DBM), which provides the advantage of reducing the amount of labeled data necessary. Labeled data are only used to fine-tune the DBM. Thus, very little annotated data is required, which will significantly reduce cost and reduce effort of annotation. A further advantage of the invention making use of a DBM is that the method comprises more applicability to a broader range of tasks, which means that one does not have to go through the expensive step of annotating images for applying this to another task, e.g. segmentation. In addition, there are a lot more transformations (e.g. lighting, perspective, and occlusion) of scenes available than there are annotated data at hand. Given that fact, a generative model, i.e. the DBM, is more likely to give a better classification. Furthermore, using an
unsupervised, generative method provides the further advantage of providing a much better and more expansive representation, as specifically with respect to scene classification, there are numerous combinations that can make up the same scene, e.g. i) a scene can be a motorway at night, which can be composed of a lot of combinations of regions and of features, ii) there is also a lot of overlap among various kinds of scenes and these overlaps are not as well captured through a human-made annotation as they are through learning the underlying probability distributions of various scenes, iii) the annotated data at hand for scene classification is not likely to have enough representation.
Preferably, the spatially ordering regions of the image comprises using one or more region descriptors to capture a semantically uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, and iii) Co-occurrence features that are Haralick features derived from the Gray Level Co-occurrence Matrix (GLCM).
The captured low-level representation of each region is semantically uncorrelated in that the above used features i), ii) and iii) are mutually uncorrelated in a semantic sense. The i)
Gabor filter is a linear filter used for texture analysis, which analyses whether there is any specific frequency content in the image in specific directions in a localized region around the point or region of analysis. Frequency and orientation representations of Gabor filters are proven to be a feature used by the human vision. The ii) HSV color space is a color space that defines the localization of a color by the features Hue, Saturation and Value.
The iii) Co-occurrence features can be used for intra-region co-occurrence statistics (mean and range), wherein said features are selected from the group consisting of Angular Second Moment, Contrast, Sum Average, Sum Variance and Difference Variance.
The driving support systems include driver assistance systems are systems, which are already known and used in state of the Art vehicles. The developed driving support systems are provided to automate, adapt and enhance vehicle systems for safety and better driving. Safety features are designed to avoid collisions and accidents by offering technologies that alert the driver to potential problems, or to avoid collisions by implementing safeguards and taking over control of the vehicle. In autonomous vehicles, the driving support systems provide input to perform a control of the vehicle. Adaptive features may automate lighting, provide adaptive cruise control, automate braking, incorporate traffic warnings, connect to smartphones, alert e.g. the driver in respect to other cars or different types of dangers, keep the vehicle in the correct lane, or show what is located in blind spots. Driving support systems including the aforementioned driver assistance systems, frequently rely on inputs from multiple data sources such as automotive imaging, image processing, radar sensors, LiDAR, ultrasonic sensors and other sources.
Further, according to a preferred embodiment of the invention, spatially ordering regions of the image further comprises adding spatial relationships among neighboring regions to create a Spatially Ordered Region Descriptor (SORD), wherein further Haralick features are used that are selected from the group consisting of the mean and range of iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
Preferably, the Spatially Ordered Region Descriptor (SORD) comprises the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, iii) Co-occurrence features, iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation. That means that the aforementioned eight features are added to the Region Descriptor to convert it into the SORD.
The Spatially Ordered Region Descriptor (SORD) can be the input to the neurons of the visible layer of the Deep Boltzmann Machine (DBM) when modelling the underlying joint probability distribution of each scene category followed by processing through a stack of Restricted Boltzmann Machines (RBMs) that learn the underlying joint probability
distributions of each scene category.
Preferably, classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs). That means that the DBM preferably is topped off with a softmax layer in order to perform the classification of the scene. Once the stack of RBMs has learned the underlying joint probability distributions of each scene category, the softmax layer added on top does the actual classification. In place of the softmax layer, a RBF-SVM or a CNN can also be used for the classification. This final classification layer requires a far smaller training set because its weights are initialized by the output of the DBM and hence require only fine-tuning for the specific scene categories in the intended application.
The invention further provides a method for scene classification of an image in image processing in a driving support system of an automotive vehicle comprising the steps of:
- providing a convolutional neuronal network (CNN) comprising a plurality of layers to learn what features of the image are most useful for the classification of scenes, wherein multiple resolutions of features are generated for capturing detail of features at higher resolution and“big picture” at lower resolution,
- modelling the joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) comprising a stack of Restricted Boltzmann Machines (RBM) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF- SVM), a second CNN or simply a softmax layer as the final layer, wherein the output of each layer of the convolutional neuronal network (CNN) is fed into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input, followed by the stack of Restricted Boltzmann Machines (RBM) learning the underlying joint probability distribution of each scene category, and
- classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).
The combination of a CNN with a DBM in a method for scene classification of an image in image processing in a driving support system of an automotive vehicle is unique and provides several advantages. It is a particular advantage of this embodiment of the invention that features used as input to the DBM are not hand-crafted but discovered by the CNN as the most useful features, which significantly reduces cost and efforts. In addition, this embodiment of the invention uses a single CNN at multiple resolutions. The features at higher resolution capture detail, whereas the features learned at lower resolution capture the “big picture”, i.e. the region-level and scene-level information in the image. The single CNN provides these features at multiple resolutions or scales. Since the lower resolution images capture the“big picture”, a CNN trained at that resolution and performing inference at that
low resolution will provide the region-level and scene-level features to the DBM, which then performs the scene classification in its softmax layer that is added on top.
In addition, the DBM has advantages over a DBN (Deep Belief Network) in that the DBN is a DAG (Directed Acyclic Graphical model), whereas the DBM is an undirected graphical model. Unlike DBNs, the approximate inference procedure in DBMs, in addition to an initial bottom-up pass, can incorporate top-down feedback, allowing DBMs to better propagate uncertainty about and hence deal more robustly with ambiguous inputs. Also, through greedy layer-wise pre-training, this embodiment of the invention can achieve fast
approximate inference in DBMs. That is, given a data vector on the visible units, each layer of hidden units can be activated in a single bottom-up pass by doubling the bottom-up input to compensate for the lack of top-down feedback (except for the very top layer, which does not have a top-down input). This fast approximate inference is used to initialize the mean- field method, which converges much faster than with random initialization.
In this embodiment of the invention scene classification is realized by using a CNN to learn what features are most useful for the classification of scenes of an image, wherein a CNN of multiple resolutions is used for capturing detail and“big picture” as separate inputs followed by modeling the joint probability distribution of each scene category using a DBM, which is then followed by classifying the scenes based on the learning result of the DBM that has been generated based on the inputs of the separate layers of the CNN.
Thus, this embodiment of the invention can be described as a Hybrid CNN-DBM model, in which a DBM learns the inter-relationships among regions and multi-resolution features of the same region with an unsupervised, generative method based on the separate inputs of the various layers of the CNN. That means that the DBM learns an internal representation of the scene category using features at multiple resolutions extracted from the image that are fed by each layer of the CNN to the visible layer of the DBM as separate inputs. In other words, each layer’s output of the CNN is fed separately to the visible layer of the DBM, which is then followed by the stack of Restricted Boltzmann Machines (RBMs) learning the underlying joint probability distribution of each scene category followed by classifying the scene based on the learning result of the DBM. This architecture of this Hybrid CNN-DBM model of the invention allows the DBM to not only learn the inter-relationships among regions in a scene, but also to learn the inter-relationships among multi-resolution features of the same regions. This is a major advantage of this embodiment of the invention over purely discriminative networks.
Thus, in this Hybrid CNN-DBM model of the invention, the primary modeling of scene categories is done by an unsupervised, generative method, i.e. the DBM, thus
a. naturally providing a representation that incorporates inter-feature relationship. The DBM learns the joint probability distribution of all its inputs. The DBM inputs are composed of the discriminative features learned at various abstraction levels. So, the DBM not only learns the inter-relationships among regions in a scene, but also the inter-relationships among the multi-resolution features of the same region;
b. providing a much better, more expansive representation as specifically with respect to scene classification, there are numerous combinations that can make up the same scene; i. for example, the scene can be a motorway at night, which can be composed of a lot of combinations of regions and of features;
ii. there is also a lot of overlap among various kinds of scenes and these overlaps are not as well captured through a human annotation as they are through learning the underlying probability distribution of various scenes.
Preferably, the convolutional neuronal network (CNN) is pre-trained using supervised training and labeled data, wherein classifying the scene is done by a temporary softmax layer as the final layer of the Deep Boltzmann Machine (DBM), wherein the temporary softmax layer is stripped off once the convolutional neuronal network (CNN) has learned the features, followed by feeding the output of each layer of the convolutional neuronal network (CNN) into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input.
Further, according to a preferred embodiment of the invention, the Deep Boltzmann Machine (DBM) is pre-trained using greedy layer-wise pre-training to learn the internal representation of the combination of multiple features in a scene and multi-resolution features of the same region, followed by adding on and pre-training the softmax layer using labelled data.
Preferably, classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs). In other words, the DBM preferably is topped off with a softmax layer, where classifying the scene based on the learning result of the DBM is performed, once the stack of RBMs has learned the underlying joint probability distributions of each scene category. In place of the softmax layer, a RBF-SVM or a CNN can also be used for the classification. This final classification layer requires a far smaller training set because its weights are initialized by the output of the DBM and hence require only fine- tuning for the specific scene categories in the intended application.
The invention also provides the use of the methods described herein in a driving support system of an automotive vehicle. Specifically, the invention provides the use of the methods for scene classification of an image in image processing in a driving support system of an automotive vehicle described above.
The invention further provides a driving support system for an automotive vehicle comprising a camera for providing images for classification, wherein the driving support system is configured for performing the methods described herein.
The invention further provides a non-transitory computer-readable medium, comprising instructions stored thereon, that when executed on a processor, induce a driving support system to perform the methods described herein.
The invention also provides an automotive vehicle comprising:
a data processing apparatus,
a non-transitory computer-readable medium comprising instructions stored thereon, that when executed on a processor, induce a driving support system to perform the methods described herein, and
a driving support system for an automotive vehicle comprising a camera for providing images for classification, wherein the driving support system is configured for performing the methods described herein.
These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments and examples described hereinafter. Individual features disclosed in the embodiments constitute alone or in combination an aspect of the present invention. Features of the different embodiments can be carried over from one embodiment to another embodiment. Embodiments of the present disclosure are further described in the following examples, which are offered by way of illustration, and are not intended to limit the invention in any manner.
In the drawings:
Fig. 1 schematically depicts an automotive vehicle with a driving support system and a camera according to a first, preferred embodiment of the invention,
Fig. 2 schematically depicts the classification of an image in an image processing in a driving support system of the automotive vehicle according to the first embodiment of the invention, and
Fig. 3 schematically depicts a second embodiment of the classification of an image in
image processing in a driving support system of an automotive vehicle that is based on a Hybrid-CNN-DBM model.
Example 1
Fig. 1 schematically depicts an automotive vehicle 1 with a driving support system 2 and a camera 3 according to a first, preferred embodiment of the invention. The camera 3 provides images of a scene, e.g. a motorway at night, for classification by the driving support system 2, which is configured for performing the methods for scene classification of an image in image processing, as described herein.
Fig. 2 schematically depicts the classification of an image in an image processing in a driving support system 2 of an automotive vehicle 1 according to a preferred embodiment of the invention. The image of a scene, e.g. a motorway at night, which is taken by the camera 3 is spatially ordered into regions by clustering the image pixels into regions with high inter-class variance and low intra-class variance.
Spatially ordering regions of the image preferably is done by using region descriptors to capture a semantically uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features and iii) Co occurrence features that are Haralick features derived from the Gray Level Co-occurrence Matrix (GLCM, intra-region and inter-region). The captured low-level representation of each region is semantically uncorrelated in that the above used features i), ii) and iii) are mutually uncorrelated in a semantic sense. The i) Gabor filter is a linear filter used for texture analysis, which analyses whether there is any specific frequency content in the image in specific directions in a localized region around the point or region of analysis. Frequency and orientation representations of Gabor filters are proven to be a feature used by the human vision. The ii) HSV color space is a color space that defines the localization of a color by the features Hue, Saturation and Value.
The iii) Co-occurrence features are used for intra-region co-occurrence statistics (mean and range), wherein said features are selected from the group consisting of Angular Second Moment, Contrast, Sum Average, Sum Variance and Difference Variance. Spatially ordering
the regions of the image is further done by adding spatial relationships among neighboring regions to create the Spatially Ordered Region Descriptor (SORD), wherein further Haralick features are used that are selected from the group consisting of the mean and range of iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
The SORD that is created in this way comprises the features i) Gabor filter, ii) Hue,
Saturation and Value (HSV) color space features, iii) Co-occurrence features, iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation. That means that the aforementioned eight features are added to the Region Descriptor to convert it into the SORD.
The SORD is then fed into the neurons of the visible layer (having e.g. 1024 units) of the DBM followed by processing through the hidden layers (a Hidden Layer 1 having e.g. 512 units and a Hidden Layer 2 having e.g. 256 units) of the DBM for modelling the underlying joint probability distribution of each scene category. The hidden layers of the DBM are a stack of RBMs.
The final step of classifying the scene is then performed by the final softmax layer having e.g. 1000 units on top of the hidden layers (stack of RBMs) of the DBM. That means that the DBM is topped off with a softmax layer in order to perform the classification of the scene. Once the stack of RBMs has learned the underlying joint probability distributions of each scene category, the softmax layer added on top does the actual classification.
Example 2
Fig. 3 schematically depicts a second embodiment of the classification of an image in image processing in a driving support system 2 of an automotive vehicle 1 that is based on a Hybrid-CNN-DBM model. The image of a scene, e.g. a motorway at night, that is taken by the camera 3 is fed into a CNN comprising a plurality of layers (Layer 1 , Layer 2, ..., Layer n) to learn what features of the image are most useful for the classification of scenes.
Multiple resolutions of features are generated for capturing detail of features at higher resolution as well as“big picture” at lower resolution. The output of each individual layer of the CNN is then fed into the visible layer of the DBM as a separate input for modelling the joint probability distribution of each scene category. In other words, this embodiment of the invention uses a single CNN at multiple resolutions. The features at higher resolution capture detail, whereas the features learned at lower resolution capture the“big picture”, i.e.
the region-level and scene-level information in the image. The single CNN provides these features at multiple resolutions or scales. Since the lower resolution images capture the“big picture”, the CNN trained at that resolution and performing inference at that low resolution provides the region-level and scene-level features for the DBM.
This is followed by processing through the hidden layers (Hidden Layer 1 , Hidden Layer 2) of the DBM. The hidden layers of the DBM are a stack of RBMs, which learn the underlying joint probability distribution of each scene category. The CNN is pre-trained using supervised training and labeled data, wherein classifying the scene is done by a temporary softmax layer as the final layer of the DBM, wherein the temporary softmax layer is stripped off once the CNN has learned the features, followed by feeding the output of each layer of the CNN into the visible layer of the DBM as a separate input.
The DBM is pre-trained using greedy layer-wise pre-training to learn the internal
representation of the combination of multiple features in a scene and multi-resolution features of the same region, followed by adding on and pre-training the softmax layer using labelled data.
The final step of classifying the scene is then performed by the final softmax layer on top of the hidden layers (stack of RBMs) of the DBM. In other words, the DBM preferably is topped off with a softmax layer, where classifying the scene based on the learning result of the DBM is performed, once the stack of RBMs has learned the underlying joint probability
distributions of each scene category.
Reference signs list
1 automotive vehicle
2 driving support system
3 camera
Claims
1 . A method for scene classification of an image in image processing in a driving
support system (2) of an automotive vehicle (1 ) comprising the steps of:
- spatially ordering regions of the image by clustering the image pixels into regions with high inter-class variance and low intra-class variance,
- modelling the underlying joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) as a generative model comprising a stack of Restricted Boltzmann Machines (RBMs) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF-SVM), a second CNN or a softmax layer as the final layer, and
- classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).
2. The method according to claim 1 , wherein spatially ordering regions of the image comprises using one or more region descriptors to capture a semantically
uncorrelated low-level representation of each region based on the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, and iii) Co occurrence features that are Haralick features derived from the Gray Level Co occurrence Matrix (GLCM).
3. The method according to claim 2, wherein the iii) Co-occurrence features are used for intra-region co-occurrence statistics (mean and range), wherein said features are selected from the group consisting of Angular Second Moment, Contrast, Sum Average, Sum Variance and Difference Variance.
4. The method according to claims 2 or 3, wherein spatially ordering regions of the image further comprises adding spatial relationships among neighboring regions to create a Spatially Ordered Region Descriptor (SORD), wherein further Haralick features are used that are selected from the group consisting of the mean and range of iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
5. The method according to claim 4, wherein the Spatially Ordered Region Descriptor (SORD) comprises the features i) Gabor filter, ii) Hue, Saturation and Value (HSV) color space features, iii) Co-occurrence features, iv) Correlation, v) Entropy, vi) Sum Variance, vii) Difference Variance and viii) Information Measures of Correlation.
6. The method according to claims 4 or 5, wherein the Spatially Ordered Region
Descriptor (SORD) is the input to the neurons of the visible layer of the Deep Boltzmann Machine (DBM) when modelling the underlying joint probability distribution of each scene category followed by processing through a stack of Restricted Boltzmann Machines (RBMs) that learn the underlying joint probability distributions of each scene category.
7. The method according to any of the previous claims, wherein classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs).
8. A method for scene classification of an image in image processing in a driving
support system (2) of an automotive vehicle (1 ) comprising the steps of:
- providing a convolutional neuronal network (CNN) comprising a plurality of layers to learn what features of the image are most useful for the classification of scenes, wherein multiple resolutions of features are generated for capturing detail of features at higher resolution and“big picture” at lower resolution,
- modelling the joint probability distribution of each scene category using a Deep Boltzmann Machine (DBM) comprising a stack of Restricted Boltzmann Machines (RBMs) and a classifier such as a Radial Basis Filter Support Vector Machine (RBF- SVM), a second CNN or a softmax layer as the final layer, wherein the output of each layer of the convolutional neuronal network (CNN) is fed into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input, followed by the stack of Restricted Boltzmann Machines (RBM) learning the underlying joint probability distribution of each scene category, and
- classifying the scene based on the learning result of the Deep Boltzmann Machine (DBM).
9. The method according to claim 8, wherein the convolutional neuronal network (CNN) is pre-trained using supervised training and labeled data, and wherein classifying the scene is done by a temporary softmax layer as the final layer of the Deep Boltzmann Machine (DBM), wherein the temporary softmax layer is stripped off once the
convolutional neuronal network (CNN) has learned the features, followed by feeding the output of each layer of the convolutional neuronal network (CNN) into the visible layer of the Deep Boltzmann Machine (DBM) as a separate input.
10. The method according to claim 8 or 9, wherein the Deep Boltzmann Machine (DBM) is pre-trained using greedy layer-wise pre-training to learn the internal representation of the combination of multiple features in a scene and multi-resolution features of the same region, followed by adding on and pre-training the softmax layer using labelled data.
1 1 . The method according to any of claims 8 to 10, wherein classifying the scene is done by a softmax layer on top of the stack of Restricted Boltzmann Machines (RBMs).
12. Use of the method according to any of claims 1 to 7 or any of claims 8 to 1 1 in a driving support system (2) of an automotive vehicle (1 ).
13. A driving support system (2) for an automotive vehicle (1 ) comprising a camera (3) for providing images for classification, wherein the driving support system (2) is configured for performing the method according to any of claims 1 to 7 or any of claims 8 to 12.
14. A non-transitory computer-readable medium (4), comprising instructions stored
thereon, that when executed on a processor, induce a driving support system (2) to perform the method of any of claims 1 to 7 or any of claims 8 to 12.
15. An automotive vehicle (1 ) comprising:
a data processing apparatus (5),
the non-transitory computer-readable medium (4) according to claim 14, and the driving support system (2) according to claim 13.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE102017127592.4A DE102017127592A1 (en) | 2017-11-22 | 2017-11-22 | A method of classifying image scenes in a driving support system |
| DE102017127592.4 | 2017-11-22 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019101720A1 true WO2019101720A1 (en) | 2019-05-31 |
Family
ID=64604600
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2018/081874 Ceased WO2019101720A1 (en) | 2017-11-22 | 2018-11-20 | Methods for scene classification of an image in a driving support system |
Country Status (2)
| Country | Link |
|---|---|
| DE (1) | DE102017127592A1 (en) |
| WO (1) | WO2019101720A1 (en) |
Cited By (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200005154A1 (en) * | 2018-02-01 | 2020-01-02 | Siemens Healthcare Limited | Data encoding and classification |
| CN110956146A (en) * | 2019-12-04 | 2020-04-03 | 新奇点企业管理集团有限公司 | Road background modeling method and device, electronic equipment and storage medium |
| CN111008664A (en) * | 2019-12-05 | 2020-04-14 | 上海海洋大学 | A hyperspectral sea ice detection method based on combined spatial and spectral features |
| CN111339834A (en) * | 2020-02-04 | 2020-06-26 | 浙江大华技术股份有限公司 | Method for recognizing vehicle traveling direction, computer device, and storage medium |
| CN111382685A (en) * | 2020-03-04 | 2020-07-07 | 电子科技大学 | Scene recognition method and system based on deep learning |
| CN111694973A (en) * | 2020-06-09 | 2020-09-22 | 北京百度网讯科技有限公司 | Model training method and device for automatic driving scene and electronic equipment |
| CN112270397A (en) * | 2020-10-26 | 2021-01-26 | 西安工程大学 | Color space conversion method based on deep neural network |
| CN112581498A (en) * | 2020-11-17 | 2021-03-30 | 东南大学 | Roadside sheltered scene vehicle robust tracking method for intelligent vehicle road system |
| CN112637487A (en) * | 2020-12-17 | 2021-04-09 | 四川长虹电器股份有限公司 | Television intelligent photographing method based on time stack expression recognition |
| CN113254468A (en) * | 2021-04-20 | 2021-08-13 | 西安交通大学 | Fault query and reasoning method for certain type of equipment |
| CN113378973A (en) * | 2021-06-29 | 2021-09-10 | 沈阳雅译网络技术有限公司 | Image classification method based on self-attention mechanism |
| CN114255268A (en) * | 2020-09-24 | 2022-03-29 | 武汉Tcl集团工业研究院有限公司 | Disparity map processing and deep learning model training method and related equipment |
| CN115187474A (en) * | 2022-06-23 | 2022-10-14 | 电子科技大学 | A two-stage dehazing method for dense fog images based on inference |
| CN115438686A (en) * | 2022-07-29 | 2022-12-06 | 西北工业大学 | Underwater sound target identification method based on data enhancement and residual CNN |
| EP4170378A1 (en) * | 2021-10-20 | 2023-04-26 | Aptiv Technologies Limited | Methods and systems for processing radar sensor data |
| CN116310970A (en) * | 2023-03-03 | 2023-06-23 | 中南大学 | Autonomous Driving Scene Classification Algorithm Based on Deep Learning |
| CN116385707A (en) * | 2023-04-04 | 2023-07-04 | 河海大学 | Deep learning scene recognition method based on multi-scale features and feature enhancement |
| CN117076969A (en) * | 2023-08-25 | 2023-11-17 | 重庆大学 | Automatic driving test scene extraction method based on one-dimensional residual convolutional autoencoder |
| US12412363B2 (en) | 2020-09-01 | 2025-09-09 | Volkswagen Aktiengesellschaft | Method and device for sensing the environment of a vehicle driving with at least semi-automation |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110126846B (en) * | 2019-05-24 | 2021-07-23 | 北京百度网讯科技有限公司 | Representation method, device, system and storage medium of driving scene |
| DE102019216628A1 (en) * | 2019-10-29 | 2021-04-29 | Zf Friedrichshafen Ag | Device and method for recognizing and classifying a closed state of a vehicle door |
| CN110954933B (en) * | 2019-12-09 | 2023-05-23 | 王相龙 | Mobile platform positioning device and method based on scene DNA |
| US11847831B2 (en) * | 2020-12-30 | 2023-12-19 | Zoox, Inc. | Multi-resolution top-down prediction |
| CN114220439A (en) * | 2021-12-24 | 2022-03-22 | 北京金山云网络技术有限公司 | Method, device, system, equipment and medium for acquiring voiceprint recognition model |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20070282506A1 (en) | 2002-09-03 | 2007-12-06 | Automotive Technologies International, Inc. | Image Processing for Vehicular Applications Applying Edge Detection Technique |
| US20080144944A1 (en) | 1992-05-05 | 2008-06-19 | Automotive Technologies International, Inc. | Neural Network Systems for Vehicles |
| US20170206440A1 (en) * | 2016-01-15 | 2017-07-20 | Ford Global Technologies, Llc | Fixation generation for machine learning |
-
2017
- 2017-11-22 DE DE102017127592.4A patent/DE102017127592A1/en not_active Withdrawn
-
2018
- 2018-11-20 WO PCT/EP2018/081874 patent/WO2019101720A1/en not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20080144944A1 (en) | 1992-05-05 | 2008-06-19 | Automotive Technologies International, Inc. | Neural Network Systems for Vehicles |
| US20070282506A1 (en) | 2002-09-03 | 2007-12-06 | Automotive Technologies International, Inc. | Image Processing for Vehicular Applications Applying Edge Detection Technique |
| US20170206440A1 (en) * | 2016-01-15 | 2017-07-20 | Ford Global Technologies, Llc | Fixation generation for machine learning |
Non-Patent Citations (5)
| Title |
|---|
| ANONYMOUS: "Otsu's method", 23 October 2017 (2017-10-23), XP002788905, Retrieved from the Internet <URL:https://en.wikipedia.org/w/index.php?title=Otsu%27s_method&oldid=806653495> [retrieved on 20190213] * |
| GAO JINGYU ET AL: "Natural scene recognition based on Convolutional Neural Networks and Deep Boltzmannn Machines", 2015 IEEE INTERNATIONAL CONFERENCE ON MECHATRONICS AND AUTOMATION (ICMA), IEEE, 2 August 2015 (2015-08-02), pages 2369 - 2374, XP033216758, ISBN: 978-1-4799-7097-1, [retrieved on 20150902], DOI: 10.1109/ICMA.2015.7237857 * |
| NIXON, MARK AND AGUADO, ALBERTO: "Feature Extraction and Image Processing", 1 January 2002, ELSEVIER, Newnes, ISBN: 0750650788, article "Chapter 8.3.3 Statistical approaches", pages: 297 - 299, XP002788906 * |
| PATTERSON, JOSH AND GIBSON, ADAM: "Deep Learning: A Practitioner's Approach", 28 July 2017, O'REILLY MEDIA INC., ISBN: 1491914238, article "Output layer for classification", pages: 180, XP002788908 * |
| RUSLAN SALAKHUTDINOV ET AL: "Deep Boltzmann Machines", PROCEEDINGS OF AISTATS 2009, 18 April 2009 (2009-04-18), United States, pages 448 - 455, XP055556013, Retrieved from the Internet <URL:http://proceedings.mlr.press/v5/salakhutdinov09a/salakhutdinov09a.pdf> [retrieved on 20190213] * |
Cited By (29)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10977558B2 (en) * | 2018-02-01 | 2021-04-13 | Siemens Healthcare Gmbh | Data encoding and classification |
| US20200005154A1 (en) * | 2018-02-01 | 2020-01-02 | Siemens Healthcare Limited | Data encoding and classification |
| CN110956146A (en) * | 2019-12-04 | 2020-04-03 | 新奇点企业管理集团有限公司 | Road background modeling method and device, electronic equipment and storage medium |
| CN110956146B (en) * | 2019-12-04 | 2024-04-12 | 新奇点企业管理集团有限公司 | Road background modeling method and device, electronic equipment and storage medium |
| CN111008664A (en) * | 2019-12-05 | 2020-04-14 | 上海海洋大学 | A hyperspectral sea ice detection method based on combined spatial and spectral features |
| CN111008664B (en) * | 2019-12-05 | 2023-04-07 | 上海海洋大学 | Hyperspectral sea ice detection method based on space-spectrum combined characteristics |
| CN111339834B (en) * | 2020-02-04 | 2023-06-02 | 浙江大华技术股份有限公司 | Method for identifying vehicle driving direction, computer device and storage medium |
| CN111339834A (en) * | 2020-02-04 | 2020-06-26 | 浙江大华技术股份有限公司 | Method for recognizing vehicle traveling direction, computer device, and storage medium |
| CN111382685A (en) * | 2020-03-04 | 2020-07-07 | 电子科技大学 | Scene recognition method and system based on deep learning |
| CN111694973A (en) * | 2020-06-09 | 2020-09-22 | 北京百度网讯科技有限公司 | Model training method and device for automatic driving scene and electronic equipment |
| CN111694973B (en) * | 2020-06-09 | 2023-10-13 | 阿波罗智能技术(北京)有限公司 | Model training methods, devices, and electronic equipment for autonomous driving scenarios |
| US12412363B2 (en) | 2020-09-01 | 2025-09-09 | Volkswagen Aktiengesellschaft | Method and device for sensing the environment of a vehicle driving with at least semi-automation |
| CN114255268A (en) * | 2020-09-24 | 2022-03-29 | 武汉Tcl集团工业研究院有限公司 | Disparity map processing and deep learning model training method and related equipment |
| CN112270397A (en) * | 2020-10-26 | 2021-01-26 | 西安工程大学 | Color space conversion method based on deep neural network |
| CN112270397B (en) * | 2020-10-26 | 2024-02-20 | 西安工程大学 | A color space conversion method based on deep neural network |
| CN112581498B (en) * | 2020-11-17 | 2024-03-29 | 东南大学 | Road side shielding scene vehicle robust tracking method for intelligent vehicle road system |
| CN112581498A (en) * | 2020-11-17 | 2021-03-30 | 东南大学 | Roadside sheltered scene vehicle robust tracking method for intelligent vehicle road system |
| CN112637487A (en) * | 2020-12-17 | 2021-04-09 | 四川长虹电器股份有限公司 | Television intelligent photographing method based on time stack expression recognition |
| CN113254468A (en) * | 2021-04-20 | 2021-08-13 | 西安交通大学 | Fault query and reasoning method for certain type of equipment |
| CN113254468B (en) * | 2021-04-20 | 2023-03-31 | 西安交通大学 | Equipment fault query and reasoning method |
| CN113378973B (en) * | 2021-06-29 | 2023-08-08 | 沈阳雅译网络技术有限公司 | An image classification method based on self-attention mechanism |
| CN113378973A (en) * | 2021-06-29 | 2021-09-10 | 沈阳雅译网络技术有限公司 | Image classification method based on self-attention mechanism |
| EP4170378A1 (en) * | 2021-10-20 | 2023-04-26 | Aptiv Technologies Limited | Methods and systems for processing radar sensor data |
| CN115187474A (en) * | 2022-06-23 | 2022-10-14 | 电子科技大学 | A two-stage dehazing method for dense fog images based on inference |
| CN115438686A (en) * | 2022-07-29 | 2022-12-06 | 西北工业大学 | Underwater sound target identification method based on data enhancement and residual CNN |
| CN116310970A (en) * | 2023-03-03 | 2023-06-23 | 中南大学 | Autonomous Driving Scene Classification Algorithm Based on Deep Learning |
| CN116310970B (en) * | 2023-03-03 | 2025-09-23 | 中南大学 | Autonomous driving scene classification algorithm based on deep learning |
| CN116385707A (en) * | 2023-04-04 | 2023-07-04 | 河海大学 | Deep learning scene recognition method based on multi-scale features and feature enhancement |
| CN117076969A (en) * | 2023-08-25 | 2023-11-17 | 重庆大学 | Automatic driving test scene extraction method based on one-dimensional residual convolutional autoencoder |
Also Published As
| Publication number | Publication date |
|---|---|
| DE102017127592A1 (en) | 2019-05-23 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019101720A1 (en) | Methods for scene classification of an image in a driving support system | |
| CN111344739B (en) | Space-time action and character positioning | |
| US10776628B2 (en) | Video action localization from proposal-attention | |
| US12205041B2 (en) | Training a generative adversarial network for performing semantic segmentation of images | |
| EP4238067B1 (en) | Neural network models for semantic image segmentation | |
| Zeng et al. | Traffic sign recognition using kernel extreme learning machines with deep perceptual features | |
| US20200012865A1 (en) | Adapting to appearance variations when tracking a target object in video sequence | |
| Kuang et al. | Nighttime vehicle detection based on bio-inspired image enhancement and weighted score-level feature fusion | |
| CN107851213B (en) | Transfer learning in neural networks | |
| CN107533669B (en) | Filter specificity as a training criterion for neural networks | |
| US20200302185A1 (en) | Recognizing minutes-long activities in videos | |
| Lee et al. | Recognizing pedestrian’s unsafe behaviors in far-infrared imagery at night | |
| WO2018089158A1 (en) | Natural language object tracking | |
| Munian et al. | Intelligent system for detection of wild animals using HOG and CNN in automobile applications | |
| Uçar et al. | Moving towards in object recognition with deep learning for autonomous driving applications | |
| Nguyen et al. | Hybrid deep learning-Gaussian process network for pedestrian lane detection in unstructured scenes | |
| US12380689B2 (en) | Managing occlusion in Siamese tracking using structured dropouts | |
| WO2023029704A1 (en) | Data processing method, apparatus and system | |
| bin Che Mansor et al. | Emergency vehicle type classification using convolutional neural network | |
| Jensen et al. | Parking space occupancy verification-improving robustness using a convolutional neural network | |
| KR20250065594A (en) | Meta-pre-training with augmentations to generalize neural network processing for domain adaptation | |
| JP2023110875A (en) | Kernel transfer | |
| Mahima et al. | Highway Collision Avoidance by Detection of Animal’s Images | |
| Oviedo | Detection and tracking of motorcycles in urban environments by using video sequences with high level of oclussion | |
| Dong | Machine Learning and Deep Learning Algorithms in Thermal Imaging Vehicle Perception Systems |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18814493 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18814493 Country of ref document: EP Kind code of ref document: A1 |