WO2020177217A1 - 基于变尺度多特征融合卷积网络的路侧图像行人分割方法 - Google Patents
基于变尺度多特征融合卷积网络的路侧图像行人分割方法 Download PDFInfo
- Publication number
- WO2020177217A1 WO2020177217A1 PCT/CN2019/087164 CN2019087164W WO2020177217A1 WO 2020177217 A1 WO2020177217 A1 WO 2020177217A1 CN 2019087164 W CN2019087164 W CN 2019087164W WO 2020177217 A1 WO2020177217 A1 WO 2020177217A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- layer
- convolution
- standard
- step size
- dimension
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
- G06V20/58—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
- G06F18/253—Fusion techniques of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/10—Segmentation; Edge detection
- G06T7/11—Region-based segmentation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
- G06V10/267—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion by performing operations on regions, e.g. growing, shrinking or watersheds
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
- G06V10/443—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
- G06V10/449—Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
- G06V10/451—Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
- G06V10/454—Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T10/00—Road transport of goods or passengers
- Y02T10/10—Internal combustion engine [ICE] based vehicles
- Y02T10/40—Engine management systems
Definitions
- the invention belongs to the technical field of computer vision and intelligent roadside perception, and relates to an image pedestrian segmentation method of an intelligent drive test terminal, in particular to a roadside image pedestrian segmentation method based on a variable-scale multi-feature fusion convolution network.
- the continuous development of deep learning technology has provided a new solution for the task of segmenting pedestrians in intelligent roadside terminal images.
- the outstanding advantage of deep learning is its powerful feature expression ability.
- the pedestrian segmentation method based on deep neural network has good adaptability to complex traffic scenes, and can obtain more accurate segmentation performance.
- the current methods of using deep neural networks for pedestrian segmentation mainly use a single network structure. It is difficult to accurately extract the boundary local features of large-scale pedestrians and the global features of small-scale pedestrians in the intelligent roadside terminal image based on the network depth. Fuzzy boundaries and even missing segmentation limit the further improvement of pedestrian segmentation accuracy and fail to achieve satisfactory results.
- the present invention discloses a roadside image pedestrian segmentation method based on a variable-scale multi-feature fusion convolutional network, which effectively solves the problem that most current pedestrian segmentation methods based on a single network structure are difficult to apply to variable-scale pedestrians.
- the problem further improves the accuracy and robustness of pedestrian segmentation.
- the present invention provides the following technical solutions:
- a roadside image pedestrian segmentation method based on variable-scale multi-feature fusion convolutional network includes the following steps:
- Sub-step 1 Design the first convolutional neural network for small-scale pedestrians, including:
- the number of standard convolutional layers is 18, of which the size of the 8-layer convolution kernel is 3 ⁇ 3, and the number of convolution kernels are 64, 64, 128, 128, 256, 256, 256, 2.
- the steps are all 1, and the remaining 10 layers of convolution kernels are all 1 ⁇ 1, and the number of convolution kernels are 32, 32, 64, 64, 128, 128, 128, 128, 128, and the steps are all 1;
- the optimal network structure is as follows:
- Standard convolution layer 1_1 Convolve with 64 3 ⁇ 3 convolution kernels and A ⁇ A pixel input samples, with a step size of 1, and then activate ReLU to obtain a feature map with a dimension of A ⁇ A ⁇ 64;
- Standard convolution layer 1_1_1 Convolve with 32 1 ⁇ 1 convolution kernels and the feature map output by the standard convolution layer 1_1, with a step size of 1, and then activate ReLU to obtain features with dimensions A ⁇ A ⁇ 32 Figure;
- Standard convolution layer 1_1_2 Use 32 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1_1, with a step size of 1, and then activate ReLU to obtain features with dimensions A ⁇ A ⁇ 32 Figure;
- Standard convolution layer 1_2 Use 64 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1_2, with a step size of 1, and then activate ReLU to obtain features with dimensions A ⁇ A ⁇ 64 Figure;
- Pooling layer 1 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 1_2 for maximum pooling, the step size is 2, and the dimension is Characteristic map;
- Standard convolutional layer 2_1 Use 128 3 ⁇ 3 convolution kernels and the feature map output by pooling layer 1 to do convolution, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolution layer 2_1_1 Use 64 1 ⁇ 1 convolution kernels and the feature map output by the standard convolution layer 2_1 to do convolution, with a step size of 1, and then activate ReLU to obtain the dimension Characteristic map;
- Standard convolution layer 2_1_2 Use 64 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 2_1_1, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolutional layer 2_2 Use 128 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolutional layer 2_1_2, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Pooling layer 2 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 2_2 for maximum pooling, the step size is 2, and the dimension is Characteristic map;
- Standard convolution layer 3_1 Use 256 3 ⁇ 3 convolution kernels and the feature map output by pooling layer 2 to do convolution, with a step size of 1, and then through ReLU activation to obtain a dimension of Characteristic map;
- Standard convolution layer 3_1_1 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolution layer 3_1_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1_1, with a step size of 1, and then activate ReLU to obtain the dimension Characteristic map;
- Standard convolution layer 3_2 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1_2, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolution layer 3_2_1 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2, with a step size of 1, and then through ReLU activation to get the dimension as Characteristic map;
- Standard convolution layer 3_2_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2_1, with a step size of 1, and then through ReLU activation to get the dimension as Characteristic map;
- Standard convolution layer 3_3 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2_2, with a step size of 1, and then through ReLU activation to get the dimension as Characteristic map;
- Standard convolution layer 3_3_1 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_3, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolution layer 3_3_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_3_1, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Standard convolution layer 3_4 Use two 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 3_3_2, with a step size of 1, and then through ReLU activation to obtain a dimension of Characteristic map;
- Deconvolution layer 4 Use 2 3 ⁇ 3 convolution kernels and standard convolution layer 3_4 output feature map to do deconvolution, the step size is 2, and the dimension is Characteristic map;
- Deconvolution layer 5 Use two 3 ⁇ 3 convolution kernels and the feature map output by the deconvolution layer 4 to do deconvolution, with a step size of 2, and obtain a feature map with a dimension of A ⁇ A ⁇ 2;
- Sub-step 2 Design a second convolutional neural network for large-scale pedestrians, including:
- the number of expanded convolutional layers is 7, the expansion rate is 2, 4, 8, 2, 4, 2, 4, the size of the convolution kernel is 3 ⁇ 3, the step length is 1, the volume
- the number of product cores are 128, 128, 256, 256, 256, 512, 512 respectively;
- the optimal network structure is as follows:
- Standard convolution layer 1_1 Convolve with 64 3 ⁇ 3 convolution kernels and A ⁇ A pixel input samples, with a step size of 1, and then activate ReLU to obtain a feature map with a dimension of A ⁇ A ⁇ 64;
- Standard convolution layer 1_2 Use 64 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1, with a step size of 1, and then activate ReLU to obtain features with dimensions of A ⁇ A ⁇ 64 Figure;
- Pooling layer 1 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 1_2 for maximum pooling, the step size is 2, and the dimension is Characteristic map;
- Expanded convolutional layer 2_1 Use 128 3 ⁇ 3 convolution kernels and the feature map output by pooling layer 1 to do convolution, with a step size of 1, an expansion rate of 2, and then ReLU activation to obtain a dimension of Characteristic map;
- Expanded convolutional layer 2_2 Use 128 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolutional layer 2_1 to do convolution, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of Characteristic map;
- Pooling layer 2 Use a 2 ⁇ 2 check to do the maximum pooling of the feature map output by the expanded convolutional layer 2_2, the step size is 2, and the dimension is Characteristic map;
- Expanded convolutional layer 3_1 Convolution with 256 3 ⁇ 3 convolution kernels and the feature map output by pooling layer 2, step size is 1, expansion rate is 8, and then activated by ReLU, the dimension is Characteristic map;
- Expanded convolutional layer 3_2 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 3_1.
- the step size is 1, the expansion rate is 2, and the ReLU activation is performed to get the dimension as Characteristic map;
- Expanded convolutional layer 3_3 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 3_2, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of Characteristic map;
- Standard convolutional layer 3_4 Convolution with 512 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolutional layer 3_3, with a step size of 1, and then activation by ReLU to obtain a dimension of Characteristic map;
- Expanded convolutional layer 3_5 Use 512 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolutional layer 3_4.
- the step size is 1, the expansion rate is 2, and the ReLU activation is performed to obtain the dimension Characteristic map;
- Expanded convolutional layer 3_6 Use 512 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 3_5, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of Characteristic map;
- Standard convolutional layer 3_7 Use two 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolutional layer 3_6 to do convolution, with a step size of 1, and then activate ReLU to obtain a dimension of Characteristic map;
- Deconvolution layer 4 Use two 3 ⁇ 3 convolution kernels and the feature map output by the standard convolution layer 3_7 to do deconvolution.
- the step size is 2, and the dimension is Characteristic map;
- Deconvolution layer 5 Use two 3 ⁇ 3 convolution kernels and the feature map output by the deconvolution layer 4 to do deconvolution, with a step size of 2, and obtain a feature map with a dimension of A ⁇ A ⁇ 2;
- Substep 3 Propose a two-level fusion strategy to fuse the features extracted by the two-way network, which specifically includes:
- the local features are located in the 9th convolutional layer from left to right, and the global features are located in the 18th convolutional layer from left to right;
- the local features are located in the fifth convolutional layer from left to right, and the global features are located in the 11th convolutional layer from left to right;
- 3Fuse the variable-scale features of the two networks at the same level merge the local features extracted by the 9th convolutional layer of the first network with the local features extracted by the 5th convolutional layer of the second network, and then merge the first network
- the global features extracted by the 18th convolutional layer are merged with the global features extracted by the 11th convolutional layer of the second network;
- the present invention has the following advantages and beneficial effects:
- the present invention proposes a pedestrian segmentation method suitable for intelligent roadside terminal images.
- the car sensor detects pedestrians and is prone to insufficient sight distance blind spots, effectively reducing the missed pedestrian detection rate;
- the present invention designs two parallel convolutional neural networks for pedestrians of different scales to extract the pedestrian features in the smart roadside terminal image, and then proposes a two-level fusion strategy to fuse the extracted features, first through the same level Feature fusion obtains the local features and global features of the variable-scale pedestrian, and then performs a secondary fusion of the fused local features and global features to obtain a variable-scale multi-feature fusion convolutional neural network.
- the network not only greatly reduces the impact of pedestrian scale differentiation on segmentation accuracy, but also takes into account the local detailed information and global information of pedestrians at different scales. Compared with most current pedestrian segmentation methods based on a single network structure, it effectively solves the segmentation.
- the boundary blur and missing segmentation issues improve the accuracy and robustness of pedestrian segmentation.
- Figure 1 is a design flow chart of the variable-scale multi-feature fusion convolutional neural network of the present invention
- FIG. 2 is a schematic diagram of the structure of the variable-scale multi-feature fusion convolutional neural network designed by the present invention
- Fig. 3 is a training flowchart of a variable-scale multi-feature fusion convolutional neural network designed by the present invention.
- the invention discloses a roadside image pedestrian segmentation method based on a variable-scale multi-feature fusion convolution network. This method designs two parallel convolutional neural networks to extract the local and global features of pedestrians at different scales in the image, and then proposes a two-level fusion strategy to fuse the extracted features. First, fuse features of the same level at different scales.
- the network effectively solves the problem that most current pedestrian segmentation methods based on a single network structure are difficult to apply to variable-scale pedestrians, and further improves the accuracy and robustness of pedestrian segmentation.
- the roadside image pedestrian segmentation method based on a variable-scale multi-feature fusion convolutional network includes the following steps:
- Sub-step 1 Design the first convolutional neural network for small-scale pedestrians, including:
- the pooling layer can reduce the size of the feature map on the one hand to reduce the amount of calculation, and on the other hand, it can expand the receptive field to capture more complete pedestrian information. Frequent pooling operations are likely to cause the loss of pedestrian location information, which hinders the improvement of segmentation accuracy. On the contrary, although the poolless operation retains as much spatial location information as possible, it increases the computational burden. Therefore, in the design, the influence of these two aspects is considered comprehensively, and the number of pooling layers is set to n p1 , the value range is 2 to 3, and the maximum pooling operation is adopted.
- the sampling size is 2 ⁇ 2, and the step size is 2. ;
- the number of standard convolutional layers with a convolution kernel of 1 ⁇ 1 is n f
- the value range is 2-12
- N b is generally taken as an integer power of 2
- the step size is 1
- the number of standard convolutional layers with a convolution kernel of 3 ⁇ 3 is n s1
- the value range is 5-10
- step (2) substep 1 Design the deconvolution layer. Since n p1 pooling operations are performed in step (2) substep 1, the feature map is reduced by 1/n p1 times. In order to restore the feature map to the original image size, and avoid Introduce a lot of noise, and use n p1 parameter learnable deconvolution layer to decouple the pedestrian features included in the feature map. Since the pedestrian segmentation task is to classify each pixel twice, the convolution kernel of the deconvolution layer The number is 2, the size of the convolution kernel is 3 ⁇ 3, and the step size is 2.
- the specific structure of the first convolutional neural network is expressed as follows:
- Standard convolution layer 1_1 Convolution with 64 3 ⁇ 3 convolution kernels and 227 ⁇ 227 pixel input samples, with a step size of 1, and then ReLU activation to obtain a feature map with a dimension of 227 ⁇ 227 ⁇ 64;
- Standard convolution layer 1_1_1 Use 32 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1, with a step size of 1, and then activate ReLU to obtain features with a dimension of 227 ⁇ 227 ⁇ 32 Figure;
- Standard convolution layer 1_1_2 Use 32 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1_1, with a step size of 1, and then activate ReLU to obtain features with a dimension of 227 ⁇ 227 ⁇ 32 Figure;
- Standard convolution layer 1_2 Use 64 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 1_1_2, with a step size of 1, and then activate ReLU to obtain features with a dimension of 227 ⁇ 227 ⁇ 64 Figure;
- Pooling layer 1 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 1_2 for maximum pooling, with a step size of 2, and get a feature map with a dimension of 113 ⁇ 113 ⁇ 64;
- Standard convolutional layer 2_1 Convolution with 128 3 ⁇ 3 convolution kernels and the feature map output by the pooling layer 1, with a step size of 1, and then through ReLU activation to obtain a feature map with a dimension of 113 ⁇ 113 ⁇ 128 ;
- Standard convolution layer 2_1_1 Convolution with 64 1 ⁇ 1 convolution kernels and the feature map output by the standard convolution layer 2_1, with a step size of 1, and then ReLU activation to obtain a feature with a dimension of 113 ⁇ 113 ⁇ 64 Figure;
- Standard convolution layer 2_1_2 Use 64 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 2_1_1, with a step size of 1, and then activate ReLU to obtain a feature with a dimension of 113 ⁇ 113 ⁇ 64 Figure;
- Standard convolution layer 2_2 Use 128 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 2_1_2, with a step size of 1, and then activate ReLU to obtain features with a dimension of 113 ⁇ 113 ⁇ 128 Figure;
- Pooling layer 2 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 2_2 for maximum pooling, with a step size of 2, and obtain a feature map with a dimension of 56 ⁇ 56 ⁇ 128;
- Standard convolution layer 3_1 Convolution with 256 3 ⁇ 3 convolution kernels and the feature map output by the pooling layer 2, step size is 1, and then activated by ReLU to obtain a feature map with a dimension of 56 ⁇ 56 ⁇ 256 ;
- Standard convolution layer 3_1_1 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolution layer 3_1_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1_1, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolution layer 3_2 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 3_1_2, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 256 Figure;
- Standard convolution layer 3_2_1 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2, with a step size of 1, and then activate ReLU to obtain a feature with a dimension of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolution layer 3_2_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2_1, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolution layer 3_3 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolution layer 3_2_2, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 256 Figure;
- Standard convolution layer 3_3_1 Convolution with 128 1 ⁇ 1 convolution kernels and the feature map output by the standard convolution layer 3_3, with a step size of 1, and then ReLU activation to obtain a feature with a dimension of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolution layer 3_3_2 Use 128 1 ⁇ 1 convolution kernels to convolve with the feature map output by the standard convolution layer 3_3_1, with a step size of 1, and then activate ReLU to obtain features with dimensions of 56 ⁇ 56 ⁇ 128 Figure;
- Standard convolutional layer 3_4 Convolution with two 3 ⁇ 3 convolution kernels and the feature map output by the standard convolutional layer 3_3_2, with a step size of 1, and then ReLU activation to obtain a feature with a dimension of 56 ⁇ 56 ⁇ 2 Figure;
- Deconvolution layer 4 Use two 3 ⁇ 3 convolution kernels and standard convolution layer 3_4 to perform deconvolution with the feature map output by the standard convolution layer 3_4, with a step size of 2, and obtain a feature map with a dimension of 113 ⁇ 113 ⁇ 2;
- Deconvolution layer 5 Use two 3 ⁇ 3 convolution kernels and the feature map output by the deconvolution layer 4 for deconvolution, with a step size of 2, and a feature map with a dimension of 227 ⁇ 227 ⁇ 2 is obtained.
- Sub-step 2 Design a second convolutional neural network for large-scale pedestrians, including:
- step (2) design the pooling layer.
- substep 1 the frequent use of the pooling layer causes a great loss of pedestrian spatial position information, which can easily cause a decrease in segmentation accuracy, although there is no pooling operation It can retain more spatial location information, but it increases the consumption of computing resources. Therefore, the two aspects are considered at the same time in the design, and the number of pooling layers is set to n p2 , the value range is 2 to 3, and the maximum pooling operation is adopted.
- the sampling size is 2 ⁇ 2, and the step size is 2. ;
- the number of expanded convolutional layers is n d
- the value range is 6-10
- dr is an even number
- the value range is Is 2 ⁇ 10
- n e is generally an integer power of 2
- the size of the convolution kernel is 3 ⁇ 3.
- the length is 1;
- step (2) design standard convolutional layers.
- the feature expression ability of the network increases as the number of convolutional layers increases, but stacking more convolutional layers increases the computational burden, while a small number of convolutional layers makes it difficult to extract To express pedestrian characteristics.
- step (2) substep 2 1 Design the deconvolution layer. Since n p2 pooling operations are performed in step (2) substep 2 1, the feature map is reduced by 1/n p2 times. In order to restore it to the original image size, while avoiding the introduction A lot of noise, the deconvolution layer with n p2 parameters that can be learned is used to decouple the pedestrian features included in the feature map.
- the number of convolution kernels in the deconvolution layer is 2, and the size of the convolution kernel is 3 ⁇ 3.
- the step length is 2.
- 5Determine the network architecture establish different network models according to the value range of each variable in step (2) and sub-step 2, and then use the data set established in step (1) to verify these models, and filter out the accuracy And real-time optimal network architecture.
- the specific structure of the second convolutional neural network is expressed as follows:
- Standard convolution layer 1_1 Convolution with 64 3 ⁇ 3 convolution kernels and 227 ⁇ 227 pixel input samples, with a step size of 1, and then ReLU activation to obtain a feature map with a dimension of 227 ⁇ 227 ⁇ 64;
- Standard convolution layer 1_2 Convolve with 64 3 ⁇ 3 convolution kernels and the feature map output by the standard convolution layer 1_1, with a step size of 1, and then activate ReLU to obtain features with a dimension of 227 ⁇ 227 ⁇ 64 Figure;
- Pooling layer 1 Use 2 ⁇ 2 to check the feature map output by the standard convolutional layer 1_2 for maximum pooling, with a step size of 2, and get a feature map with a dimension of 113 ⁇ 113 ⁇ 64;
- Expanded convolutional layer 2_1 Use 128 3 ⁇ 3 convolution kernels to convolve with the feature map output by pooling layer 1, with a step size of 1, an expansion rate of 2, and ReLU activation to obtain a dimension of 113 ⁇ 113 ⁇ 128 feature map;
- Expanded convolutional layer 2_2 Use 128 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 2_1, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of 113 ⁇ 113 ⁇ 128 feature map;
- Pooling layer 2 Use a 2 ⁇ 2 check to do the maximum pooling of the feature map output by the expanded convolutional layer 2_2, with a step size of 2, and get a feature map with a dimension of 56 ⁇ 56 ⁇ 128;
- Expanded convolutional layer 3_1 Convolution with 256 3 ⁇ 3 convolution kernels and the feature map output by pooling layer 2, step size is 1, expansion rate is 8, and then activated by ReLU, the dimension is 56 ⁇ 56 ⁇ 256 feature map;
- Expanded convolutional layer 3_2 Use 256 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 3_1.
- the step size is 1, the expansion rate is 2, and the ReLU activation is performed to obtain a dimension of 56 ⁇ 56 ⁇ 256 feature map;
- Expanded convolutional layer 3_3 Convolve with 256 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolutional layer 3_2, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of 56 ⁇ 56 ⁇ 256 feature map;
- Standard convolution layer 3_4 Convolution with 512 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolution layer 3_3, with a step size of 1, and then through ReLU activation to obtain a feature with a dimension of 56 ⁇ 56 ⁇ 512 Figure;
- Expanded convolutional layer 3_5 Use 512 3 ⁇ 3 convolution kernels to convolve with the feature map output by the standard convolutional layer 3_4.
- the step size is 1, the expansion rate is 2, and the ReLU activation is performed to obtain a dimension of 56 ⁇ 56 ⁇ 512 feature map;
- Expanded convolutional layer 3_6 Use 512 3 ⁇ 3 convolution kernels to convolve with the feature map output by the expanded convolutional layer 3_5, with a step size of 1, an expansion rate of 4, and ReLU activation to obtain a dimension of 56 ⁇ 56 ⁇ 512 feature map;
- Standard convolutional layer 3_7 Convolution with two 3 ⁇ 3 convolution kernels and the feature map output by the expanded convolution layer 3_6, with a step size of 1, and then ReLU activation to obtain a feature with a dimension of 56 ⁇ 56 ⁇ 2 Figure;
- Deconvolution layer 4 Use two 3 ⁇ 3 convolution kernels and the feature map output by the standard convolution layer 3_7 for deconvolution, with a step size of 2, and obtain a feature map with a dimension of 113 ⁇ 113 ⁇ 2;
- Deconvolution layer 5 Use two 3 ⁇ 3 convolution kernels and the feature map output by the deconvolution layer 4 for deconvolution, with a step size of 2, and a feature map with a dimension of 227 ⁇ 227 ⁇ 2 is obtained.
- Substep 3 Propose a two-level fusion strategy to fuse the features extracted by the two-way network, which specifically includes:
- step (2) 2Determine the location of the local feature and global feature of the second convolutional neural network, and determine the location of the local feature and global feature according to the method described in step (2), substep 3, where the location of the local feature is recorded as s l2 , the value range is 3-6, the global feature is located in the 11th convolutional layer from left to right;
- the value of s l1 is 9 and the value of s l2 is 5 through the feature visualization method, and the first network is The local features extracted by the 9 convolutional layers are merged with the local features extracted by the fifth convolutional layer of the second network, and then the global features extracted by the 18th convolutional layer of the first network are combined with the 11th network of the second network Global feature fusion extracted by convolutional layer;
- a convolution kernel size of 1 ⁇ 1 convolution is used to change the scale of the shallow layer of the second network.
- the dimensionality of the local features of pedestrians is reduced to have the same dimensions as the deep global features, and then a jump connection structure is constructed to fuse the local features with the global features, and the resulting variable-scale multi-feature fusion convolutional neural network architecture is shown in Figure 2. Show.
- the training process includes two stages of forward propagation and back propagation.
- the sample set (x, y) is input to the network, where x is the input image and y is the corresponding label.
- the actual output f(x) is obtained through the network layer by layer operation, and the cross-entropy cost function with L2 regularization term is used to measure the error between the ideal output y and the actual output f(x):
- the first term is the cross-entropy cost function
- the second term is the L2 regularization term to prevent overfitting.
- ⁇ represents the parameters to be learned by the convolutional neural network model
- M represents the number of training samples
- N represents the number of pixels in each image
- Q represents the number of semantic categories in the sample, for road segmentation
- Q 2
- ⁇ is the regularization coefficient
- Means Corresponding label Means
- the probability of belonging to the qth category is defined as:
- the network parameters are updated layer by layer from back to front through the stochastic gradient descent algorithm to minimize the error between the actual output and the ideal output.
- the parameter update formula is as follows:
- ⁇ is the learning rate
- J 0 ( ⁇ ) is the cross-entropy cost function
- Is the calculated gradient
- Sub-step 1 Select data sets related to autonomous driving, such as ApolloScape, Cityscapes, CamVid, process them to include only pedestrian categories, and then adjust the sample size to 227 ⁇ 227 pixels and record it as D c , then use D c Pre-train the two designed convolutional neural networks and set the pre-training hyperparameters respectively, where the maximum number of iterations are respectively I c1 and I c2 , the learning rates are respectively ⁇ c1 , ⁇ c2 , and the weight attenuation is ⁇ c1 respectively , ⁇ c2 , and finally save the network parameters obtained by pre-training;
- autonomous driving such as ApolloScape, Cityscapes, CamVid
- Sub-step 2 Use the data set D k established in step (1) to fine-tune the parameters of the two networks pre-trained in sub-step 1 of step (3), and set the maximum number of iterations to I k1 and I k2 respectively ,
- the learning rate is ⁇ k1 , ⁇ k2
- the weight attenuation is ⁇ k1 , ⁇ k2 respectively .
- Sub-step 3 Use the data set D k established in step (1) to train the variable-scale multi-feature fusion convolutional neural network obtained in sub-step 3 of step (2), and reset the maximum number of iterations to I k3 , The learning rate is ⁇ k3 and the weight attenuation is ⁇ k3 respectively. Then, according to the changes of the training loss curve and the verification loss curve, that is, when the training loss curve slowly decreases and tends to converge, and the verification loss curve is at the critical point of rising, the most parameters are obtained. Optimal variable scale multi-feature fusion convolutional neural network model.
- variable-scale multi-feature fusion convolutional neural network for pedestrian segmentation, adjust the size of the pedestrian sample obtained by the intelligent roadside terminal to 227 ⁇ 227 pixels and input it into the trained variable-scale multi-feature fusion convolutional neural network , Get the result of pedestrian segmentation.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- Life Sciences & Earth Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Mathematical Physics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biodiversity & Conservation Biology (AREA)
- Human Computer Interaction (AREA)
- Image Analysis (AREA)
Abstract
Description
Claims (1)
- 基于变尺度多特征融合卷积网络的路侧图像行人分割方法,其特征在于,包括以下步骤:(1)建立行人分割数据集;(2)构建变尺度多特征融合卷积神经网络,首先设计并行的两个卷积神经网络来提取图像中不同尺度行人的局部特征和全局特征,第一个网络针对小尺度行人设计了精细的特征提取结构,第二个网络针对大尺度行人扩大了网络在浅层处的感受野;进而提出两级融合策略对所提取的特征进行融合,首先对不同尺度的同级特征进行融合,得到适用于变尺度行人的局部特征和全局特征,然后构建跳跃连接结构将融合后的局部特征与全局特征进行二次融合,获取完备的变尺度行人局部细节信息和全局信息,最终得到变尺度多特征融合卷积神经网络,包括以下子步骤:子步骤1:设计第一个针对小尺度行人的卷积神经网络,具体包括:①设计池化层,池化层数量为2,均采用最大池化操作,采样尺寸均为2×2,步长均为2;②设计标准卷积层,标准卷积层数量为18,其中8层卷积核大小均为3×3,卷积核数量分别为64、64、128、128、256、256、256、2,步长均为1,剩下10层卷积核大小均为1×1,卷积核数量分别为32、32、64、64、128、128、128、128、128、128,步长均为1;③设计反卷积层,反卷积层数量为2,卷积核大小均为3×3,步 长均为2,卷积核数量分别为2、2;④确定网络架构,根据步骤(2)的子步骤1中①~③涉及的网络层参数建立不同的网络模型,然后利用步骤(1)所建立的数据集对这些模型进行验证,从中筛选出兼顾准确性和实时性的网络结构,得到最优网络架构如下:标准卷积层1_1:用64个3×3的卷积核与A×A像素的输入样本做卷积,步长为1,再经过ReLU激活,得到维度为A×A×64的特征图;标准卷积层1_1_1:用32个1×1的卷积核与标准卷积层1_1输出的特征图做卷积,步长为1,再经过ReLU激活,得到维度为A×A×32的特征图;标准卷积层1_1_2:用32个1×1的卷积核与标准卷积层1_1_1输出的特征图做卷积,步长为1,再经过ReLU激活,得到维度为A×A×32的特征图;标准卷积层1_2:用64个3×3的卷积核与标准卷积层1_1_2输出的特征图做卷积,步长为1,再经过ReLU激活,得到维度为A×A×64的特征图;反卷积层5:用2个3×3的卷积核与反卷积层4输出的特征图做反卷积,步长为2,得到维度为A×A×2的特征图;子步骤2:设计第二个针对大尺度行人的卷积神经网络,具体包括:①设计池化层,池化层数量为2,均采用最大池化操作,采样尺寸均为2×2,步长均为2;②设计扩张卷积层,扩张卷积层数量为7,扩张率分别为2、4、8、2、4、2、4,卷积核大小均为3×3,步长均为1,卷积核数量分别为128、128、256、256、256、512、512;③设计标准卷积层,标准卷积层数量为4,卷积核大小均为3×3,步长均为1,卷积核数量分别为64、64、512、2;④设计反卷积层,反卷积层数量为2,卷积核大小均为3×3,步长均为2,卷积核数量分别为2、2;⑤确定网络架构,根据步骤(2)的子步骤2中①~④涉及的网络层参数建立不同的网络模型,然后利用步骤(1)所建立的数据集对这些模型进行验证,从中筛选出兼顾准确性和实时性的网络结构,得到最优网络架构如下:标准卷积层1_1:用64个3×3的卷积核与A×A像素的输入样本做卷积,步长为1,再经过ReLU激活,得到维度为A×A×64的特征图;标准卷积层1_2:用64个3×3的卷积核与标准卷积层1_1输出的特征图做卷积,步长为1,再经过ReLU激活,得到维度为A×A×64的特征图;反卷积层5:用2个3×3的卷积核与反卷积层4输出的特征图做反卷积,步长为2,得到维度为A×A×2的特征图;子步骤3:提出两级融合策略对两路网络提取的特征进行融合,具体包括:①确定第一个卷积神经网络的局部特征和全局特征所在位置,局部特征位于从左至右第9个卷积层,全局特征位于从左至右第18个卷积层;②确定第二个卷积神经网络的局部特征和全局特征所在位置,局部特征位于从左至右第5个卷积层,全局特征位于从左至右第11个卷积层;③融合两个网络的变尺度同级特征,将第一个网络第9个卷积层提取的局部特征与第二个网络第5个卷积层提取的局部特征融合,再将第一个网络第18个卷积层提取的全局特征与第二个网络第11个卷 积层提取的全局特征融合;④融合第二个网络的局部特征与全局特征,使用1×1卷积对第二个网络浅层所包含的变尺度行人局部特征进行降维,使其具有与深层全局特征相同的维度,然后构建跳跃连接结构将局部特征与全局特征融合,得到变尺度多特征融合卷积神经网络架构;(3)训练设计的变尺度多特征融合卷积神经网络,获得网络参数;(4)使用变尺度多特征融合卷积神经网络进行行人分割。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/267,493 US11783594B2 (en) | 2019-03-04 | 2019-05-16 | Method of segmenting pedestrians in roadside image by using convolutional network fusing features at different scales |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910161808.0A CN109977793B (zh) | 2019-03-04 | 2019-03-04 | 基于变尺度多特征融合卷积网络的路侧图像行人分割方法 |
| CN201910161808.0 | 2019-03-04 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020177217A1 true WO2020177217A1 (zh) | 2020-09-10 |
Family
ID=67077920
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/087164 Ceased WO2020177217A1 (zh) | 2019-03-04 | 2019-05-16 | 基于变尺度多特征融合卷积网络的路侧图像行人分割方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11783594B2 (zh) |
| CN (1) | CN109977793B (zh) |
| WO (1) | WO2020177217A1 (zh) |
Cited By (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112232300A (zh) * | 2020-11-11 | 2021-01-15 | 汇纳科技股份有限公司 | 全局遮挡自适应的行人训练/识别方法、系统、设备及介质 |
| CN112242193A (zh) * | 2020-11-16 | 2021-01-19 | 同济大学 | 一种基于深度学习的自动血管穿刺方法 |
| CN112307982A (zh) * | 2020-11-02 | 2021-02-02 | 西安电子科技大学 | 基于交错增强注意力网络的人体行为识别方法 |
| CN112434744A (zh) * | 2020-11-27 | 2021-03-02 | 北京奇艺世纪科技有限公司 | 一种多模态特征融合模型的训练方法及装置 |
| CN112836673A (zh) * | 2021-02-27 | 2021-05-25 | 西北工业大学 | 一种基于实例感知和匹配感知的重识别方法 |
| CN113378792A (zh) * | 2021-07-09 | 2021-09-10 | 合肥工业大学 | 融合全局和局部信息的弱监督宫颈细胞图像分析方法 |
| CN113674844A (zh) * | 2021-08-19 | 2021-11-19 | 浙江远图互联科技股份有限公司 | 基于多头cnn网络的医院门诊人流量预测及分诊系统 |
| CN113762483A (zh) * | 2021-09-16 | 2021-12-07 | 华中科技大学 | 一种用于心电信号分割的1D U-net神经网络处理器 |
| CN114494879A (zh) * | 2022-01-28 | 2022-05-13 | 福建农林大学 | 一种基于对位叠加操作的鸟类识别方法、系统及存储介质 |
| CN114821258A (zh) * | 2022-04-26 | 2022-07-29 | 湖北工业大学 | 一种基于特征图融合的类激活映射方法及装置 |
| CN115273141A (zh) * | 2022-07-04 | 2022-11-01 | 电子科技大学 | 一种自选择感受野块、图像处理方法及应用 |
| CN115829101A (zh) * | 2022-11-21 | 2023-03-21 | 国网甘肃省电力公司酒泉供电公司 | 一种基于多尺度时间卷积网络的用电成本预测及优化方法 |
| CN115909171A (zh) * | 2022-12-19 | 2023-04-04 | 浙江金汇华特种耐火材料有限公司 | 钢包透气砖生产方法及其系统 |
| CN116503956A (zh) * | 2023-05-24 | 2023-07-28 | 电子科技大学 | 一种基于多扩展率压缩卷积的太极拳动作识别方法 |
| CN116563526A (zh) * | 2022-01-26 | 2023-08-08 | 北京沃东天骏信息技术有限公司 | 图像语义分割方法和装置 |
| CN118169752A (zh) * | 2024-03-13 | 2024-06-11 | 北京石油化工学院 | 一种基于多特征融合的地震相位拾取方法及系统 |
| CN118506185A (zh) * | 2024-05-24 | 2024-08-16 | 南京航空航天大学 | 一种基于遥感影像的损毁评估方法及系统 |
| CN119229539A (zh) * | 2024-11-28 | 2024-12-31 | 临沂大学 | 基于多传感器数据融合的盲人步行方向识别方法 |
| CN119784593A (zh) * | 2024-12-27 | 2025-04-08 | 中山大学 | 一种图像超分辨率方法、系统、计算机设备和存储介质 |
Families Citing this family (45)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102708717B1 (ko) * | 2019-05-17 | 2024-09-23 | 삼성전자주식회사 | 특정 순간에 관한 사진 또는 동영상을 자동으로 촬영하는 디바이스 및 그 동작 방법 |
| CN110378305B (zh) * | 2019-07-24 | 2021-10-12 | 中南民族大学 | 茶叶病害识别方法、设备、存储介质及装置 |
| CN110674685B (zh) * | 2019-08-19 | 2022-05-31 | 电子科技大学 | 一种基于边缘信息增强的人体解析分割模型及方法 |
| CN111079761B (zh) * | 2019-11-05 | 2023-07-18 | 北京航空航天大学青岛研究院 | 图像处理方法、装置及计算机存储介质 |
| CN111222465B (zh) * | 2019-11-07 | 2023-06-13 | 深圳云天励飞技术股份有限公司 | 基于卷积神经网络的图像分析方法及相关设备 |
| CN110929622B (zh) * | 2019-11-15 | 2024-01-05 | 腾讯科技(深圳)有限公司 | 视频分类方法、模型训练方法、装置、设备及存储介质 |
| CN111126561B (zh) * | 2019-11-20 | 2022-07-08 | 江苏艾佳家居用品有限公司 | 一种基于多路并行卷积神经网络的图像处理方法 |
| CN111178211B (zh) * | 2019-12-20 | 2024-01-12 | 天津极豪科技有限公司 | 图像分割方法、装置、电子设备及可读存储介质 |
| CN111695447B (zh) * | 2020-05-26 | 2022-08-12 | 东南大学 | 一种基于孪生特征增强网络的道路可行驶区域检测方法 |
| CN112001301B (zh) * | 2020-08-21 | 2021-07-20 | 江苏三意楼宇科技股份有限公司 | 基于全局交叉熵加权的楼宇监控方法、装置和电子设备 |
| US11886983B2 (en) * | 2020-08-25 | 2024-01-30 | Microsoft Technology Licensing, Llc | Reducing hardware resource utilization for residual neural networks |
| CN112256823B (zh) * | 2020-10-29 | 2023-06-20 | 众阳健康科技集团有限公司 | 一种基于邻接密度的语料数据抽样方法及系统 |
| CN112464743B (zh) * | 2020-11-09 | 2023-06-02 | 西北工业大学 | 一种基于多尺度特征加权的小样本目标检测方法 |
| US12243158B2 (en) * | 2020-12-29 | 2025-03-04 | Volvo Car Corporation | Ensemble learning for cross-range 3D object detection in driver assist and autonomous driving systems |
| CN112509190B (zh) * | 2021-02-08 | 2021-05-11 | 南京信息工程大学 | 基于屏蔽门客流计数的地铁车辆断面客流统计方法 |
| KR102860336B1 (ko) * | 2021-02-08 | 2025-09-16 | 삼성전자주식회사 | 연사 영상 기반의 영상 복원 방법 및 장치 |
| CN113988163B (zh) * | 2021-10-20 | 2024-12-10 | 向前 | 基于多尺度分组融合卷积的雷达高分辨距离像识别方法 |
| CN114092760B (zh) * | 2021-11-05 | 2025-06-10 | 通号通信信息集团有限公司 | 卷积神经网络中自适应特征融合方法及系统 |
| CN114049339B (zh) * | 2021-11-22 | 2023-05-12 | 江苏科技大学 | 一种基于卷积神经网络的胎儿小脑超声图像分割方法 |
| CN113902765B (zh) * | 2021-12-10 | 2022-04-12 | 聚时科技(江苏)有限公司 | 基于全景分割的半导体自动分区方法 |
| CN114295368B (zh) * | 2021-12-24 | 2024-10-15 | 江苏国科智能电气有限公司 | 一种多通道融合的风电行星齿轮箱故障诊断方法 |
| CN114462491B (zh) * | 2021-12-29 | 2025-11-04 | 浙江大华技术股份有限公司 | 一种行为分析模型训练方法、行为分析方法及其设备 |
| CN116452602A (zh) * | 2022-01-07 | 2023-07-18 | 中移(成都)信息通信科技有限公司 | 一种图像分割方法、装置、设备及存储介质 |
| CN114549583A (zh) * | 2022-01-18 | 2022-05-27 | 西南石油大学 | 一种用于无人机跟踪的特征信息增强的孪生网络模型 |
| CN114652326B (zh) * | 2022-01-30 | 2024-06-14 | 天津大学 | 基于深度学习的实时脑疲劳监测装置及数据处理方法 |
| CN114565792B (zh) * | 2022-02-28 | 2024-07-19 | 华中科技大学 | 一种基于轻量化卷积神经网络的图像分类方法和装置 |
| CN114612759B (zh) | 2022-03-22 | 2023-04-07 | 北京百度网讯科技有限公司 | 视频处理方法、查询视频的方法和模型训练方法、装置 |
| CN114863252B (zh) * | 2022-04-12 | 2025-02-14 | 中国飞机强度研究所 | 一种基于双方向门控机制的飞机内仓裂纹分割方法 |
| CN114862920B (zh) * | 2022-04-29 | 2025-03-21 | 西安电子科技大学 | 基于多尺度图像恢复的跨摄像机行人重识别方法及装置 |
| CN114882530B (zh) * | 2022-05-09 | 2024-07-12 | 东南大学 | 一种构建面向行人检测的轻量级卷积神经网络模型的方法 |
| CN114821064B (zh) * | 2022-05-13 | 2025-01-03 | 东南大学 | 基于多光谱图像融合的森林非结构化场景分割方法 |
| CN115187844A (zh) * | 2022-06-30 | 2022-10-14 | 深圳云天励飞技术股份有限公司 | 基于神经网络模型的图像识别方法、装置及终端设备 |
| US11676399B1 (en) * | 2022-07-18 | 2023-06-13 | Motional Ad Llc. | Object tracking |
| CN115500807B (zh) * | 2022-09-20 | 2024-10-15 | 山东大学 | 基于小型卷积神经网络的心律失常分类检测方法及系统 |
| CN115527168A (zh) * | 2022-10-08 | 2022-12-27 | 通号通信信息集团有限公司 | 行人重识别方法、存储介质、数据库编辑方法、存储介质 |
| CN115880312A (zh) * | 2022-11-18 | 2023-03-31 | 重庆邮电大学 | 一种三维图像自动分割方法、系统、设备和介质 |
| CN116091789B (zh) * | 2023-01-29 | 2025-11-14 | 国网江苏省电力有限公司南京供电分公司 | 一种电动汽车充电场异常行为识别方法 |
| CN116523933A (zh) * | 2023-04-21 | 2023-08-01 | 国网山西省电力公司信息通信分公司 | 变电站电力设备图像分割模型、方法及训练方法 |
| CN116883466B (zh) * | 2023-07-11 | 2025-08-15 | 中国人民解放军国防科技大学 | 基于位置感知的光学与sar图像配准方法、装置及设备 |
| CN117173633B (zh) * | 2023-09-18 | 2025-08-12 | 华东师范大学 | 一种基于旋转等变卷积神经网络的行人轨迹预测方法 |
| CN117274591A (zh) * | 2023-09-22 | 2023-12-22 | 中科超精(南京)科技有限公司 | 一种医学图像智能分割网络模型及方法 |
| CN117876835B (zh) * | 2024-02-29 | 2024-07-26 | 重庆师范大学 | 基于残差Transformer的医学图像融合方法 |
| CN118484555A (zh) * | 2024-05-24 | 2024-08-13 | 山东省计算中心(国家超级计算济南中心) | 基于特征融合的医学图像检索方法 |
| CN120823532B (zh) * | 2025-09-17 | 2025-12-23 | 杭州师范大学 | 一种基于多尺度动态特征融合的无人机水面漂浮垃圾检测方法及系统 |
| CN121353693A (zh) * | 2025-10-23 | 2026-01-16 | 北京人形机器人创新中心有限公司 | 一种图像处理方法及系统 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2018003212A1 (ja) * | 2016-06-30 | 2018-01-04 | クラリオン株式会社 | 物体検出装置及び物体検出方法 |
| CN107688786A (zh) * | 2017-08-30 | 2018-02-13 | 南京理工大学 | 一种基于级联卷积神经网络的人脸检测方法 |
| CN108520219A (zh) * | 2018-03-30 | 2018-09-11 | 台州智必安科技有限责任公司 | 一种卷积神经网络特征融合的多尺度快速人脸检测方法 |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10068024B2 (en) * | 2012-02-01 | 2018-09-04 | Sri International | Method and apparatus for correlating and viewing disparate data |
| JP5991332B2 (ja) * | 2014-02-05 | 2016-09-14 | トヨタ自動車株式会社 | 衝突回避制御装置 |
| CN105631413A (zh) * | 2015-12-23 | 2016-06-01 | 中通服公众信息产业股份有限公司 | 一种基于深度学习的跨场景行人搜索方法 |
| US10486707B2 (en) * | 2016-01-06 | 2019-11-26 | GM Global Technology Operations LLC | Prediction of driver intent at intersection |
| WO2017158983A1 (ja) * | 2016-03-18 | 2017-09-21 | 株式会社Jvcケンウッド | 物体認識装置、物体認識方法及び物体認識プログラム |
| CN106570564B (zh) * | 2016-11-03 | 2019-05-28 | 天津大学 | 基于深度网络的多尺度行人检测方法 |
| KR20180067909A (ko) * | 2016-12-13 | 2018-06-21 | 한국전자통신연구원 | 영상 분할 장치 및 방법 |
| JP6514736B2 (ja) * | 2017-05-17 | 2019-05-15 | 株式会社Subaru | 車外環境認識装置 |
| US11537868B2 (en) * | 2017-11-13 | 2022-12-27 | Lyft, Inc. | Generation and update of HD maps using data from heterogeneous sources |
| US10586132B2 (en) * | 2018-01-08 | 2020-03-10 | Visteon Global Technologies, Inc. | Map and environment based activation of neural networks for highly automated driving |
| CN108062756B (zh) * | 2018-01-29 | 2020-04-14 | 重庆理工大学 | 基于深度全卷积网络和条件随机场的图像语义分割方法 |
| US10468062B1 (en) * | 2018-04-03 | 2019-11-05 | Zoox, Inc. | Detecting errors in sensor data |
| US10627823B1 (en) * | 2019-01-30 | 2020-04-21 | StradVision, Inc. | Method and device for performing multiple agent sensor fusion in cooperative driving based on reinforcement learning |
-
2019
- 2019-03-04 CN CN201910161808.0A patent/CN109977793B/zh not_active Expired - Fee Related
- 2019-05-16 WO PCT/CN2019/087164 patent/WO2020177217A1/zh not_active Ceased
- 2019-05-16 US US17/267,493 patent/US11783594B2/en active Active
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2018003212A1 (ja) * | 2016-06-30 | 2018-01-04 | クラリオン株式会社 | 物体検出装置及び物体検出方法 |
| CN107688786A (zh) * | 2017-08-30 | 2018-02-13 | 南京理工大学 | 一种基于级联卷积神经网络的人脸检测方法 |
| CN108520219A (zh) * | 2018-03-30 | 2018-09-11 | 台州智必安科技有限责任公司 | 一种卷积神经网络特征融合的多尺度快速人脸检测方法 |
Non-Patent Citations (1)
| Title |
|---|
| JIANAN LI, LIANG XIAODAN, SHEN SHENGMEI, XU TINGFA, FENG JIASHI, YAN SHUICHENG: "Scale-Aware Fast R-CNN for Pedestrian Detection", IEEE TRANSACTIONS ON MULTIMEDIA, vol. 20, no. 4, 30 April 2018 (2018-04-30), pages 985 - 996, XP055736367 * |
Cited By (28)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112307982A (zh) * | 2020-11-02 | 2021-02-02 | 西安电子科技大学 | 基于交错增强注意力网络的人体行为识别方法 |
| CN112307982B (zh) * | 2020-11-02 | 2023-07-28 | 西安电子科技大学 | 基于交错增强注意力网络的人体行为识别方法 |
| CN112232300A (zh) * | 2020-11-11 | 2021-01-15 | 汇纳科技股份有限公司 | 全局遮挡自适应的行人训练/识别方法、系统、设备及介质 |
| CN112232300B (zh) * | 2020-11-11 | 2024-01-19 | 汇纳科技股份有限公司 | 全局遮挡自适应的行人训练/识别方法、系统、设备及介质 |
| CN112242193A (zh) * | 2020-11-16 | 2021-01-19 | 同济大学 | 一种基于深度学习的自动血管穿刺方法 |
| CN112242193B (zh) * | 2020-11-16 | 2023-03-31 | 同济大学 | 一种基于深度学习的自动血管穿刺方法 |
| CN112434744A (zh) * | 2020-11-27 | 2021-03-02 | 北京奇艺世纪科技有限公司 | 一种多模态特征融合模型的训练方法及装置 |
| CN112434744B (zh) * | 2020-11-27 | 2023-05-26 | 北京奇艺世纪科技有限公司 | 一种多模态特征融合模型的训练方法及装置 |
| CN112836673B (zh) * | 2021-02-27 | 2024-06-04 | 西北工业大学 | 一种基于实例感知和匹配感知的重识别方法 |
| CN112836673A (zh) * | 2021-02-27 | 2021-05-25 | 西北工业大学 | 一种基于实例感知和匹配感知的重识别方法 |
| CN113378792A (zh) * | 2021-07-09 | 2021-09-10 | 合肥工业大学 | 融合全局和局部信息的弱监督宫颈细胞图像分析方法 |
| CN113378792B (zh) * | 2021-07-09 | 2022-08-02 | 合肥工业大学 | 融合全局和局部信息的弱监督宫颈细胞图像分析方法 |
| CN113674844A (zh) * | 2021-08-19 | 2021-11-19 | 浙江远图互联科技股份有限公司 | 基于多头cnn网络的医院门诊人流量预测及分诊系统 |
| CN113762483A (zh) * | 2021-09-16 | 2021-12-07 | 华中科技大学 | 一种用于心电信号分割的1D U-net神经网络处理器 |
| CN113762483B (zh) * | 2021-09-16 | 2024-02-09 | 华中科技大学 | 一种用于心电信号分割的1D U-net神经网络处理器 |
| CN116563526A (zh) * | 2022-01-26 | 2023-08-08 | 北京沃东天骏信息技术有限公司 | 图像语义分割方法和装置 |
| CN114494879A (zh) * | 2022-01-28 | 2022-05-13 | 福建农林大学 | 一种基于对位叠加操作的鸟类识别方法、系统及存储介质 |
| CN114821258A (zh) * | 2022-04-26 | 2022-07-29 | 湖北工业大学 | 一种基于特征图融合的类激活映射方法及装置 |
| CN115273141A (zh) * | 2022-07-04 | 2022-11-01 | 电子科技大学 | 一种自选择感受野块、图像处理方法及应用 |
| CN115829101B (zh) * | 2022-11-21 | 2023-06-06 | 国网甘肃省电力公司酒泉供电公司 | 一种基于多尺度时间卷积网络的用电成本预测及优化方法 |
| CN115829101A (zh) * | 2022-11-21 | 2023-03-21 | 国网甘肃省电力公司酒泉供电公司 | 一种基于多尺度时间卷积网络的用电成本预测及优化方法 |
| CN115909171B (zh) * | 2022-12-19 | 2023-12-15 | 浙江金汇华特种耐火材料有限公司 | 钢包透气砖生产方法及其系统 |
| CN115909171A (zh) * | 2022-12-19 | 2023-04-04 | 浙江金汇华特种耐火材料有限公司 | 钢包透气砖生产方法及其系统 |
| CN116503956A (zh) * | 2023-05-24 | 2023-07-28 | 电子科技大学 | 一种基于多扩展率压缩卷积的太极拳动作识别方法 |
| CN118169752A (zh) * | 2024-03-13 | 2024-06-11 | 北京石油化工学院 | 一种基于多特征融合的地震相位拾取方法及系统 |
| CN118506185A (zh) * | 2024-05-24 | 2024-08-16 | 南京航空航天大学 | 一种基于遥感影像的损毁评估方法及系统 |
| CN119229539A (zh) * | 2024-11-28 | 2024-12-31 | 临沂大学 | 基于多传感器数据融合的盲人步行方向识别方法 |
| CN119784593A (zh) * | 2024-12-27 | 2025-04-08 | 中山大学 | 一种图像超分辨率方法、系统、计算机设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109977793B (zh) | 2022-03-04 |
| US20210303911A1 (en) | 2021-09-30 |
| CN109977793A (zh) | 2019-07-05 |
| US11783594B2 (en) | 2023-10-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020177217A1 (zh) | 基于变尺度多特征融合卷积网络的路侧图像行人分割方法 | |
| CN110009648B (zh) | 基于深浅特征融合卷积神经网络的路侧图像车辆分割方法 | |
| US12154352B2 (en) | Lane line detection method and related device | |
| CN110188705B (zh) | 一种适用于车载系统的远距离交通标志检测识别方法 | |
| CN109829400B (zh) | 一种快速车辆检测方法 | |
| CN108846328B (zh) | 基于几何正则化约束的车道检测方法 | |
| CN110766098A (zh) | 基于改进YOLOv3的交通场景小目标检测方法 | |
| CN111460919A (zh) | 一种基于改进YOLOv3的单目视觉道路目标检测及距离估计方法 | |
| CN111553201A (zh) | 一种基于YOLOv3优化算法的交通灯检测方法 | |
| CN113052071B (zh) | 危化品运输车驾驶员分心行为快速检测方法及系统 | |
| CN107563372A (zh) | 一种基于深度学习ssd框架的车牌定位方法 | |
| CN111882620A (zh) | 一种基于多尺度信息道路可行驶区域分割方法 | |
| CN106446914A (zh) | 基于超像素和卷积神经网络的道路检测 | |
| CN110414421B (zh) | 一种基于连续帧图像的行为识别方法 | |
| CN107301369A (zh) | 基于航拍图像的道路交通拥堵分析方法 | |
| CN110009095A (zh) | 基于深度特征压缩卷积网络的道路行驶区域高效分割方法 | |
| CN111695447B (zh) | 一种基于孪生特征增强网络的道路可行驶区域检测方法 | |
| CN114091598A (zh) | 一种基于语义级信息融合的多车协同环境感知方法 | |
| CN111652129A (zh) | 一种基于语义分割和多特征融合的车辆前障碍物检测方法 | |
| CN115512325A (zh) | 一种端到端的基于实例分割的车道检测方法 | |
| CN117765507A (zh) | 一种基于深度学习的雾天交通标志检测方法 | |
| CN114627381B (zh) | 一种基于改进半监督学习的电瓶车头盔识别方法 | |
| CN114863122A (zh) | 一种基于人工智能的智能化高精度路面病害识别方法 | |
| CN114639084A (zh) | 一种基于ssd改进算法的路侧端车辆感知方法 | |
| CN112052829B (zh) | 一种基于深度学习的飞行员行为监控方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19918275 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19918275 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19918275 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 24.03.2022) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19918275 Country of ref document: EP Kind code of ref document: A1 |
