WO2025001898A1 - 一种图像处理方法、装置、电子设备、计算机可读介质 - Google Patents
一种图像处理方法、装置、电子设备、计算机可读介质 Download PDFInfo
- Publication number
- WO2025001898A1 WO2025001898A1 PCT/CN2024/099597 CN2024099597W WO2025001898A1 WO 2025001898 A1 WO2025001898 A1 WO 2025001898A1 CN 2024099597 W CN2024099597 W CN 2024099597W WO 2025001898 A1 WO2025001898 A1 WO 2025001898A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target
- image
- data
- tracking
- query
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/41—Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V2201/00—Indexing scheme relating to image or video recognition or understanding
- G06V2201/07—Target detection
Definitions
- Embodiments of the present disclosure relate to an image processing method, an apparatus, an electronic device, and a computer-readable medium.
- Target tracking is used to analyze the motion state information of some targets from video data (for example, their position in each frame of video image, etc.).
- the present disclosure provides an image processing method, an apparatus, an electronic device, and a computer-readable medium.
- the present disclosure provides an image processing method, the method comprising:
- Decoding is performed based on the detection query characterization data and the tracking query characterization data corresponding to the image to be processed to obtain target description information of the image to be processed; the tracking query characterization data is determined based on target description information of multiple historical images corresponding to the image to be processed; the target description information of the image to be processed is used to describe at least one target in the image to be processed.
- encoding the feature extraction result to obtain detection query representation data includes:
- the tracking query table corresponding to the detection query characterization data and the image to be processed The target description information of the image to be processed is obtained by decoding the target data, including:
- Decoding is performed based on the image feature representation data, the detection query representation data, and the tracking query representation data corresponding to the image to be processed to obtain target description information of the image to be processed.
- the plurality of historical images include first-category images and second-category images
- the tracking query representation data is determined according to the first tracking target description information and the target description information of the second type of image.
- determining the tracking query representation data according to the first tracking target description information and the target description information of the second type of image includes:
- the tracking query representation data is determined according to the second tracking target description information and the first tracking target description information.
- the image to be processed and a plurality of historical images corresponding to the image to be processed belong to the same video data
- the timing of the to-be-processed images is later than the timing of the first type of images
- the timing of the first type of images is later than the timing of the second type of images.
- the decoding process is performed based on the detection query representation data and the tracking query representation data corresponding to the image to be processed to obtain the target description information of the image to be processed, including:
- Decoding is performed according to the query representation data to be processed to obtain target description information of the image to be processed.
- the decoding process is performed based on the image feature representation data, the detection query representation data, and the tracking query representation data corresponding to the image to be processed, to obtain
- the target description information of the image to be processed includes:
- Decoding is performed based on the image feature representation data and the query representation data to be processed to obtain target description information of the image to be processed.
- the target description information of the image to be processed is determined using a pre-trained decoder
- the decoder includes at least one decoding layer, and the self-attention module in the decoding layer is implemented using a first attention mask, which is used to shield the interaction between corresponding query representation data belonging to the same target in different historical images.
- the at least one target includes one or more targets to be tracked
- the at least one decoding layer comprises a first decoding layer and a second decoding layer
- the decoder also includes a feature fusion module
- the feature fusion module is used to perform fusion processing on the tracking query processing data output by the first decoding layer to obtain tracking query fusion data;
- the tracking query processing data is determined according to the first decoding layer and the tracking query representation data;
- the tracking query processing data includes the processed query representation data output by the first decoding layer for each of the targets to be tracked;
- the tracking query fusion data includes the fused query representation data output by the feature fusion module for each of the processed query representation data;
- the second decoding layer is used to process the tracking query fusion data and the detection query processing data output by the first decoding layer; the detection query processing data is determined according to the first decoding layer and the detection query representation data.
- the feature fusion module includes an information removal branch network and an information addition branch network;
- the process of determining the tracking query fusion data includes:
- the tracking query fusion data is determined according to the retained information features and the to-be-added information features.
- the process of determining the retained information feature includes:
- information removal processing is performed on the tracking query processing data to obtain the retained information characteristics.
- the process of determining the information feature to be added includes:
- the second time series information and the tracking query processing data are subjected to self-attention processing to obtain the information features to be added.
- the self-attention layer in the feature fusion module is implemented using a second attention mask, where the second attention mask is used to shield the interaction between query representation data belonging to different targets.
- the target description information of the image to be processed is determined using a pre-trained target tracking model
- the process of determining the training loss of the target tracking model includes:
- Target description information of the first image data according to the first image data, target description information of a plurality of historical images corresponding to the first image data, and the target tracking model
- a training loss of the target tracking model is determined based on the first loss and the second loss.
- the timing corresponding to the first group of information is later than the timing corresponding to the second group of information.
- the second set of information and the target tag number According to the data, the second loss includes:
- the second loss is determined according to the label matching result, the second group of information, and the target label data; the label matching result is determined according to the matching result between the first group of information and the target label data.
- the process of determining the second loss includes:
- the second loss is determined according to an average value of the losses corresponding to the at least one target prediction result.
- the target description information of the image to be processed is determined using a pre-trained target tracking model
- the training process of the target tracking model includes:
- the target tracking model is obtained by training a decoding module in the model to be optimized using at least one image sequence and a target tracking label of each image sequence.
- the at least one target includes one or more targets to be tracked; and the tracking query representation data includes query representation data corresponding to each of the targets to be tracked under multiple historical images.
- the method further includes:
- the image sequence is video data.
- the present disclosure provides an image processing device, comprising:
- the extraction unit is configured to perform feature extraction processing on the image to be processed to obtain a feature extraction result.
- An encoding unit configured to encode the feature extraction result to obtain detection query representation data
- a decoding unit is configured to perform decoding processing based on the detection query representation data and the tracking query representation data corresponding to the image to be processed to obtain target description information of the image to be processed; the tracking query representation data is determined based on target description information of multiple historical images corresponding to the image to be processed; the target description information of the image to be processed is used to describe at least one target in the image to be processed.
- the present disclosure provides an electronic device, the device comprising: a processor and a memory;
- the memory is configured to store instructions or computer programs
- the processor is configured to execute the instructions or computer programs in the memory so that the electronic device performs the image processing method provided by the present disclosure.
- the present disclosure provides a computer-readable medium, in which instructions or computer programs are stored.
- the instructions or computer programs are executed on a device, the device executes the image processing method provided by the present disclosure.
- the present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program codes for executing the image processing method provided by the present disclosure.
- FIG1 is a flow chart of an image processing method provided by an embodiment of the present disclosure.
- FIG2 is a schematic diagram of a target tracking model provided by an embodiment of the present disclosure.
- FIG3 is a schematic diagram of a feature fusion module provided in an embodiment of the present disclosure.
- FIG4 is a schematic diagram of the structure of an image processing device provided by an embodiment of the present disclosure.
- FIG5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure.
- an image processing method which includes: for an image to be processed (for example, a frame of video image in a video data), first performing feature extraction processing on the image to be processed to obtain a feature extraction result, so that the feature extraction result is used to characterize the image information carried by the image to be processed; then encoding the feature extraction result to obtain detection query representation data (for example, multiple detection queries); finally, decoding processing is performed based on the detection query representation data and the tracking query representation data corresponding to the image to be processed (for example, a tracking query determined based on multiple historical images) to obtain target description information of the image to be processed (for example, target content feature representation data, target position representation data, and category confidence, etc.), so that the target description information can represent some targets in the image to be processed.
- an image to be processed for example, a frame of video image in a video data
- the tracking query characterization data is determined based on the target description information of multiple historical images corresponding to the image to be processed, the tracking query characterization data can better represent the historical status of some targets in the image to be processed, so that the target description information determined based on the tracking query characterization data can better represent some targets in the image to be processed, which is conducive to improving the target tracking effect.
- the image processing method provided by the present invention when used to implement target tracking processing for video data, the image processing method needs to refer to the target description information of multiple historical images (for example, Query) to determine the target in a certain video image, so that in the process of determining the target for the video image, more reference content (for example, the state of the target in different historical images, etc.) can be obtained from the target description information of these historical images.
- This can effectively overcome the problems caused by the low frame rate of the video data, thereby effectively improving the target tracking effect for low frame rate videos, and thus making the image processing method provided by the present invention suitable for target tracking processing for video data with any frame rate, which is conducive to improving the universality of the image processing method in the field of target tracking.
- the present disclosure does not limit the execution subject of the image processing method.
- the present disclosure provides Any implementation of the image processing method can be applied to a device with data processing function such as a terminal device or a server.
- a device with data processing function such as a terminal device or a server.
- any implementation of the image processing method provided in the embodiment of the present disclosure can also be implemented by means of a data communication process between different devices (for example, a terminal device and a server, two terminal devices, or two servers).
- the terminal device can be a smart phone, a computer, a personal digital assistant (PDA) or a tablet computer, etc.
- PDA personal digital assistant
- the server can be an independent server, a cluster server or a cloud server.
- the image processing method provided by the present disclosure includes the following S1-S3.
- Figure 1 is a flow chart of an image processing method provided by the embodiment of the present disclosure.
- S1 Perform feature extraction processing on the image to be processed to obtain feature extraction results.
- the image to be processed refers to image data that needs to be processed for target determination (for example, the t-th frame video image I t shown in FIG. 2 ).
- t is a positive integer.
- the present disclosure does not limit the implementation method of the target.
- the target can refer to any detectable object (for example, a person, an animal, a plant, an object, etc.).
- the feature extraction result is used to represent the image information carried by the above image to be processed.
- the present disclosure does not limit the determination process of the above feature extraction results (that is, the implementation method of S1).
- it can be implemented by any existing or future method that can perform feature extraction processing on an image data (for example, with the help of a feature extractor, etc.).
- S1 above can specifically be: input the image to be processed into a pre-built feature extractor to obtain a feature extraction result output by the feature extractor, so that the feature extraction result can represent the image information carried by the image to be processed.
- the feature extractor is used to perform feature extraction processing on the input data of the feature extractor; and the present disclosure does not limit the implementation method of the feature extractor.
- the feature extractor can be implemented using any existing or future device with a feature extraction function.
- the feature extractor can be implemented using a convolutional neural network (CNN).
- CNN convolutional neural network
- the feature extractor can be implemented using the feature extraction module in the target tracking model below.
- CNN does not limit the implementation method of CNN in the above paragraph.
- it can be implemented using ResNet50 pre-trained based on the CoCo dataset.
- the above S1 can specifically be: utilizing the feature extraction module in the target tracking model (for example, the CNN shown in FIG2 ) to perform feature extraction processing on the above image to be processed, and obtain a feature extraction result corresponding to the image to be processed.
- the image to be processed for example, the t-th frame video image I t shown in FIG2
- feature extraction processing is performed on the image to be processed to obtain a feature extraction result (for example, the data output by the CNN shown in FIG2 ), so that the feature extraction result can represent the image information carried by the image to be processed, so that some targets in the image to be processed can be determined based on the feature extraction result later.
- the detection query representation data is used to represent some targets in the above-mentioned image to be processed; and the present disclosure does not limit the detection query representation data, for example, the detection query representation data may include query representation data (e.g., Query) corresponding to multiple targets in the image to be processed. Wherein, the query representation data corresponding to the target is used to describe a target in the image to be processed. It can be seen that under a possible implementation, the detection query representation data may include query representation data corresponding to one or more targets under the image to be processed. Wherein, the query representation data corresponding to the target under the image to be processed is used to represent the state of the target under the image to be processed.
- query representation data e.g., Query
- the present disclosure does not limit the implementation of the query representation data corresponding to the target in the above paragraph.
- it may include the target content feature representation data corresponding to the target and the target location corresponding to the target.
- Characterization data is used to characterize the features of the target in the image to be processed (for example, features in terms of posture, clothing color, etc.); and the present disclosure does not limit the implementation method of the target content feature characterization data.
- the target content feature characterization data can be implemented using a feature vector.
- the target position characterization data is used to characterize the position of the target in the image to be processed; and the present disclosure does not limit the implementation method of the target content feature characterization data. For example, it can be implemented using a bounding box.
- the present disclosure does not limit the implementation method of the encoding process in S2 above.
- it can be implemented by using any existing or future method that can perform encoding processing (for example, with the help of an encoder, etc.).
- S2 above can specifically be: inputting the feature extraction result above into a pre-built encoder to obtain the detection query representation data output by the encoder.
- the encoder is used to encode the input data of the encoder; and the present disclosure does not limit the implementation of the encoder.
- the encoder can be implemented using any existing or future device with encoding function.
- the encoder can be implemented using a Transformer Encoder.
- the encoder can be implemented using the encoding module in the target tracking model below.
- the present disclosure does not limit the implementation method of the Transformer Encoder in the previous paragraph.
- it can be implemented using the Transformer Encoder module in the DINO (DETR with Improved DeNoising Anchor Boxes) detection algorithm.
- the above S2 can specifically be: after determining the feature extraction result from the above image to be processed by using the feature extraction module in the target tracking model, the encoding module in the target tracking model performs encoding processing on the feature extraction result to obtain the detection query representation data corresponding to the image to be processed (for example, D t shown in FIG. 2 ).
- D t is used to represent some targets in the image data I t shown in FIG. 2 ; and D t can be represented by the following formula (1).
- N det represents a preset number of queries, so that N det is used to represent the number of queries involved in D t .
- S2 can specifically be: encoding the feature extraction result to obtain image feature representation data and detection query representation data.
- the image feature representation data is used to represent the information carried by the image to be processed above; and the present disclosure does not limit the implementation of the image feature representation data.
- it can be implemented using any existing or future feature encoding vector determined for image data (for example, the image feature map shown in Figure 2).
- the encoding process may specifically be: after determining the feature extraction result from the above image to be processed by the feature extraction module in the target tracking model, the encoding module in the target tracking model performs encoding processing on the feature extraction result to obtain image feature representation data corresponding to the image to be processed (for example, the image feature map shown in FIG2 ) and detection query representation data corresponding to the image to be processed (for example, D t shown in FIG2 ).
- an image to be processed for example, the t-th frame video image I t shown in FIG. 2
- a feature extraction result corresponding to the image to be processed for example, data output by the CNN shown in FIG. 2
- the feature extraction result is encoded to obtain image feature representation data (for example, the image feature map shown in FIG. 2 ) and detection query representation data (for example, D t shown in FIG. 2 ), so that some targets in the image to be processed can be determined based on the image feature representation data and the detection query representation data.
- the image to be processed for example, the t-th image shown in FIG. 2
- the feature extraction result is encoded to at least obtain detection query representation data (for example, D t shown in FIG. 2 ), so that some targets in the image to be processed can be determined based on the detection query representation data.
- S3 Decode the detection query representation data and the tracking query representation data corresponding to the image to be processed to obtain target description information of the image to be processed; the tracking query representation data is determined based on target description information of multiple historical images corresponding to the image to be processed; the target description information of the image to be processed is used to describe at least one target in the image to be processed.
- the tracking query representation data corresponding to the image to be processed refers to the tracking query required to be used as the basis for target determination processing on the image to be processed, so that the tracking query representation data can represent the status of some tracked targets in the image to be processed (for example, the "target to be tracked” below) in the historical stage.
- the tracking query characterization data corresponding to the image to be processed is determined based on the target description information of multiple historical images corresponding to the image to be processed.
- the time sequence of the historical images is earlier than the time sequence of the image to be processed; and the present disclosure does not limit the multiple historical images corresponding to the image to be processed.
- the multiple historical images corresponding to the image to be processed may include the t-1th frame video image I t-1 , the t-2th frame video image I t-2 , the t-3th frame video image I t-3 , etc. shown in FIG2 .
- the target description information of the historical image refers to the result obtained when the target determination processing is performed on the historical image, so that the target description information of the historical image can describe some targets in the historical image.
- the target description information may include query representation data corresponding to some targets (for example, the feature vector corresponding to the target + the bounding box corresponding to the target).
- the target description information may also include the category confidence corresponding to these targets. The category confidence is used to indicate the probability that the target belongs to a certain category.
- the present disclosure does not limit the implementation method of the “multiple” in the above “multiple historical images corresponding to the image to be processed”. For example, it can be 2, 3 or other positive integers greater than 1.
- the present disclosure does not limit the implementation method of the process of determining the tracking query representation data corresponding to the image to be processed above.
- the process of determining the tracking query representation data may specifically include: determining the query representation data (Query) corresponding to part or all of the targets in the target description information of the rth image data. is the tracking Query corresponding to the r-th image data, r is a positive integer, r ⁇ R, and R is a positive integer; then the tracking Queries corresponding to the R image data are aggregated to obtain the tracking query representation data, so that the tracking query representation data includes the tracking Queries corresponding to some tracked targets in the multiple historical images corresponding to the image to be processed.
- the present disclosure also provides a possible implementation method of the process of determining the tracking query representation data corresponding to the above-mentioned image to be processed.
- the process of determining the tracking query representation data corresponding to the image to be processed may include the following steps 11-12.
- Step 11 extracting first tracking target description information from the target description information of the first category of images.
- the above-mentioned first category of images refers to the historical image existing in the multiple historical images and having a timing closest to the timing of the image to be processed (for example, the t-1th frame video image I t-1 shown in FIG. 2 ).
- the first tracking target description information is used to describe the most recent historical information extracted from the first type of image above, which is needed to be referenced when determining some targets in the above image to be processed, so that the first tracking target description information can represent some tracked targets that appear in the historical image whose timing is closest to that of the image to be processed.
- the present disclosure does not limit the implementation method of the first tracking target description information above.
- it may include target content feature representation data and target position representation data of at least one tracked target presented in the first type of image above, so that the target content feature representation data can represent the state of the tracked target in the historical image whose time sequence is closest to the time sequence of the image to be processed, and the target position representation data can represent the position of the tracked target in the historical image whose time sequence is closest to the time sequence of the image to be processed.
- the present disclosure does not limit the determination process of the first tracking target description information above.
- it can be specifically as follows: after obtaining the target description information of the first type of image above, first extract at least one query that meets the preset filtering condition from the target description information; then determine the first tracking target description information based on these queries, so that the first tracking target description information includes these queries.
- the preset filtering condition can be pre-set, for example, the preset filtering condition can be: the query is assigned a tracking object identifier (for example, ID), and the IOU (Intersection over Union) of the query is greater than a preset threshold.
- the tracking object identifier is used to uniquely identify the tracking process for a tracked target.
- step 11 above Based on the relevant content of step 11 above, it can be known that for the above image to be processed (for example, the t-th frame video image I t shown in FIG. 2 ), after obtaining the target description information of a historical image (for example, the t-1-th frame video image I t-1 shown in FIG. 2 ) corresponding to the image to be processed and closest to the image to be processed, some target content feature representation data and target position representation data of tracked targets can be extracted from the target description information to be determined as the first tracking target description information (for example, the tracking Query corresponding to t-1 in X t shown in FIG. 2 ), so that the first tracking target description information can represent the status of these tracked targets in the historical image.
- the first tracking target description information for example, the tracking Query corresponding to t-1 in X t shown in FIG. 2
- Step 12 According to the first tracking target description information and the target description information of the second type of image, the tracking query representation data corresponding to the image to be processed is determined.
- the second category of images is used to represent other images in the multiple historical images corresponding to the above-mentioned image to be processed, except for the above-mentioned first category of images; and the present disclosure does not limit the second category of images.
- the timing of the image to be processed is later than the timing of the first category of images
- the timing of the first category of images is later than the timing of the second category of images.
- the above-mentioned second category of images refers to historical images whose timing existing in the multiple historical images is not the closest to the timing of the image to be processed (for example, the t-2th frame video image I t-2 , the t-3th frame video image I t-3 shown in Figure 2, etc.).
- the present disclosure does not limit the implementation of the above step 12.
- it can be specifically as follows: firstly extract some target content feature representation data of the tracked targets from the target description information of the second type of image above; then perform a preset processing (for example, integration processing, fusion processing) on the target content feature representation data and the target content feature representation data of the corresponding target in the above first tracking target description information; Or splicing processing, etc.), to obtain the tracking query representation data corresponding to the above-mentioned image to be processed, so that the tracking query representation data can not only represent the position of these tracked targets in the latest frame of historical images, but also can more comprehensively represent the characteristics of these tracked targets in the latest multiple frames of historical images.
- a preset processing for example, integration processing, fusion processing
- the tracking query representation data can not only represent the position of these tracked targets in the latest frame of historical images, but also can more comprehensively represent the characteristics of these tracked targets in the latest multiple frames of historical images.
- the present disclosure also provides another possible implementation of the above step 12, which may specifically include the following steps 121-122.
- Step 121 Determine second tracking target description information according to the target position representation data in the first tracking target description information and the target content feature representation data in the target description information of the second type of image.
- the second tracking target description information is used to describe the tracking query determined based on the second type of image mentioned above and required to be referenced when determining some targets in the image to be processed mentioned above.
- the present disclosure does not limit the implementation method of the above step 121.
- it can be specifically: combining the target position representation data in the above first tracking target description information and the target content feature representation data of the corresponding target in the above second type of image target description information to obtain at least one combined query of the tracked target, so that the combined query can represent the features presented by the tracked target in the second type of image and the position of the tracked target in the above first type of image; and then determining the second tracking target description information (for example, the tracking query corresponding to t-2 or the tracking query corresponding to t-3 in Xt shown in Figure 2) based on these combined queries, so that the second tracking target description information includes these combined queries.
- the second tracking target description information for example, the tracking query corresponding to t-2 or the tracking query corresponding to t-3 in Xt shown in Figure 2
- step 121 Based on the relevant content of step 121 above, it can be known that for the above image to be processed (for example, the t-th frame video image I t shown in FIG. 2 ), after obtaining the target description information of a historical image corresponding to the image to be processed and not the closest to the image to be processed (for example, the t-2-th frame video image I t-2 or the t-3-th frame video image I t-3 shown in FIG.
- some target content feature representation data of the tracked targets can be extracted from the target description information first, so that the target content feature representation data can represent the features presented by the tracked targets in the historical image; then, the target content feature representation data is combined with the target position representation data of the corresponding target in the above first tracking target description information to obtain a combined Query of the tracked target; finally, based on the combined Query, the above second tracking target description information is determined, so that the second tracking target description information can represent the tracking Query determined based on the historical image and required to be referenced when determining some targets in the above image to be processed.
- Step 122 Determine the tracking query representation data corresponding to the above image to be processed according to the above second tracking target description information and the above first tracking target description information.
- i, i+1, and i+2 shown in FIG. 2 respectively represent tracking object identifiers assigned to a tracked target, and i is a positive integer.
- the target position representation data in the target description information of the first type of image and the target content feature representation data in the target description information of the second type of image can be used to construct a tracking Query corresponding to the second type of image (that is, the second tracking target description information above); and then, the tracking query corresponding to the second type of image and the tracking query corresponding to the first type of image (that is, the first tracking target description information above) are combined to determine the tracking query representation data corresponding to the above image to be processed, so that the tracking query representation data can better represent the status of some tracked targets in the image to be processed in the historical stage.
- the query characterization data (eg, D t shown in FIG. 2 ) is used to detect a new target in the image to be processed, so that the target description information of the image to be processed can be determined based on the tracking query characterization data and the detection query characterization data.
- the tracking query representation data includes the query representation data corresponding to the first target to be tracked under multiple historical images, the query representation data corresponding to the second target to be tracked under multiple historical images, ... (and so on), and the query representation data corresponding to the Yth target to be tracked under multiple historical images.
- the yth target to be tracked refers to the target that exists in the image to be processed and has been tracked in the historical stage; and the query representation data corresponding to the yth target to be tracked in a historical image is used to represent the state of the yth target to be tracked in the historical image, y is a positive integer, y ⁇ Y, and Y is a positive integer.
- the tracking query representation data includes at least two tracking queries for each tracked target, so that in the subsequent processing process (for example, the decoding process, etc.), for each tracked target, each tracking query corresponding to the target is respectively used as an independent query to complete the tracking prediction processing for the target, and each query will make full use of the information carried in other queries corresponding to the target (for example, target content feature representation data + target position representation data, etc.) to complete its own iterative update processing in this process to achieve collaborative tracking, so that the collaborative tracking process executed for each query is used to jointly track the same target, which can better improve the reliability of the features predicted for the target, so that the features predicted for the target can better express the state of the target in the image to be processed, which is conducive to improving the target tracking effect.
- the “target description information of the image to be processed” is used to describe at least one target in the image to be processed; and the present disclosure does not limit the implementation method of the “target description information of the image to be processed”.
- the “target description information of the image to be processed” may include query representation data (Query) of some targets and the category confidence of these targets.
- the query representation data of the target may include target content feature representation data of the target and target position representation data of the target (such as a bounding box, etc.).
- the category confidence of the target is used to indicate how likely the target belongs to certain categories (such as people, animals, cars, etc.).
- the target description information of the image to be processed may include the target prediction results corresponding to each detection query and the target prediction results corresponding to each tracking query.
- the target prediction result corresponding to the j-th detection query refers to the state of a target in the image to be processed derived based on the j-th detection query; j is a positive integer, j ⁇ J, J is a positive integer, and J represents the number of the detection query.
- the target prediction result corresponding to the m-th tracking query refers to the state of a target in the image to be processed derived based on the m-th tracking query; m is a positive integer, m ⁇ M, M is a positive integer, and M represents the number of the tracking query.
- the present disclosure does not limit the determination process of the above "target description information of the image to be processed" (that is, the implementation of S3 above), for example, it can be implemented with the help of a pre-built decoder.
- the decoder is used to decode the input data of the decoder; and the present disclosure does not limit the decoder, for example, the decoder can be implemented using the decoding module in the multi-target tracking model shown in Figure 2.
- the above S3 can be specifically: decoding the image feature representation data, the detection query representation data and the tracking query representation data corresponding to the above image to be processed to obtain the target description information of the image to be processed. It should be noted that the implementation of S3 shown in this paragraph can also be implemented with the help of a pre-built decoder.
- the present disclosure also provides a possible implementation of the above S3, which may specifically include the following steps 21 and 22.
- Step 21 Determine the query representation data to be processed according to the above detection query representation data and the above tracking query representation data.
- the query representation data to be processed is used to describe all the queries (for example, multiple detection queries + multiple tracking queries) that need to be referenced when performing target determination processing on the above image to be processed, so that the query representation data to be processed can better represent the characteristics of some targets in the image to be processed.
- the present disclosure does not limit the determination process of the above query representation data to be processed.
- it can be specifically: cascade processing of the above detection query representation data and the above tracking query representation data (also referred to as That is, splicing processing) is performed to obtain the query representation data to be processed, so that the query representation data to be processed includes the detection query representation data and the tracking query representation data, so that the query representation data to be processed can better represent the characteristics of some targets in the image to be processed.
- Step 22 Decode the query representation data to be processed to obtain target description information of the image to be processed.
- the present disclosure does not limit the implementation method of the above step 22.
- it can specifically be: inputting the above query representation data to be processed into a pre-constructed decoder (for example, the decoding module in the multi-target tracking model shown in Figure 2), so that the decoder can perform decoding processing based on the query representation data to be processed, and obtain and output the target description information of the above image to be processed (for example, the data output by the last decoding layer shown in Figure 2), so that the target description information can describe some targets in the image to be processed.
- a pre-constructed decoder for example, the decoding module in the multi-target tracking model shown in Figure 2
- the target description information of the above image to be processed for example, the data output by the last decoding layer shown in Figure 2
- the step 22 can be specifically: decoding the image feature representation data and the above query representation data to be processed to obtain the target description information of the image to be processed.
- the present disclosure does not limit the implementation method of the step in the previous paragraph "decoding the image feature representation data and the above query representation data to be processed to obtain target description information of the image to be processed".
- it can specifically be: inputting the image feature representation data and the query representation data to be processed into a pre-constructed decoder (for example, the decoding module in the multi-target tracking model shown in Figure 2), so that the decoder can decode the image feature representation data and the query representation data to be processed, obtain and output the target description information of the above image to be processed (for example, the data output by the last decoding layer shown in Figure 2), so that the target description information can describe some targets in the image to be processed.
- a pre-constructed decoder for example, the decoding module in the multi-target tracking model shown in Figure 2
- the two queries can be cascaded first; and then the cascade results can be directly sent to a pre-built decoder so that the decoder can output the target description information of the image to be processed.
- S3 above can be specifically as follows: for the above image to be processed (for example, the t-th frame video image I t shown in FIG. 2 ), after obtaining the tracking query representation data corresponding to the image to be processed (for example, X t shown in FIG. 2 ), the detection query representation data corresponding to the image to be processed output by the encoding module in the target tracking model (for example, D t shown in FIG. 2 ) and the image feature representation data corresponding to the image to be processed (for example, the image feature map shown in FIG.
- the tracking query representation data and the detection query representation data can be cascaded first; and then the decoding module in the target tracking model performs decoding processing based on the cascade result and the image feature representation data, so that the decoding module can obtain and output the target description information of the image to be processed (for example, the data output by the last decoding layer shown in FIG. 2 ) by an autoregressive manner, so that the target description information includes some target queries and the category confidences of these targets.
- the target description information of the above image to be processed is determined by using a pre-trained decoder.
- the decoder can be used to perform decoding processing based on the above detection query representation data and the tracking query representation data corresponding to the above image to be processed to obtain the target description information of the image to be processed; or, the decoder can also be used to perform decoding processing based on the above image feature representation data, the detection query representation data and the tracking query representation data corresponding to the image to be processed to obtain the target description information of the image to be processed.
- the present disclosure also provides a possible implementation of the above decoder, under which the decoder may include at least one decoding layer, and the self-attention module in the decoding layer is implemented using a first attention mask, which is used to shield the interaction between the query representation data corresponding to the same target in different historical images, so that the decoding layer can achieve the effect of temporal blocking.
- the present disclosure does not limit the number of the decoding layers, for example, it can be 6.
- the present disclosure does not limit the implementation method of the "query representation data corresponding to the same target in different historical images" in the previous paragraph.
- the decoder shown in the previous paragraph refers to the h-th decoder
- the "query representation data corresponding to the same target in different historical images” refers to different tracking queries corresponding to the same target existing in the input data of the self-attention module in the h-th decoder; h is a positive integer, h ⁇ the number of decoders (for example, 6).
- the above decoder can be stacked by multiple (for example, 6) temporal blocking decoder layers, and the self-attention module in each temporal blocking decoder layer is implemented using the first attention mask. It is now possible to enable each self-attention module to use the first attention mask to shield the interaction between the query representation data corresponding to the same target in different historical images, so as to avoid mutual inhibition between different tracking queries of the same target.
- the first attention mask mentioned above refers to the attention mask configured for the self-attention module in the decoding layer above, so that the self-attention module can use the attention mask to shield the interaction between the query representation data corresponding to the same target under different historical images; and the first attention mask is configured to shield the interaction between the query representation data corresponding to the same target under different historical images.
- the present disclosure does not limit the implementation method of the first attention mask.
- the first attention mask in the self-attention module in the decoding layer shown in FIG2 can be implemented using the attention mask shown in Table 1 below.
- Table 1 A possible implementation of the attention mask in the self-attention module in the decoding layer It should be noted that for the above Table 1, represents the tracking Query corresponding to the target identified by i in the image data I t-1 ; It represents the tracking query corresponding to the target identified by i+1 in the image data I t-1 ; It represents the tracking query corresponding to the target identified by i+1 in the image data I t-2 ; It represents the tracking Query corresponding to the target identified by i+1 as the tracking object in the image data I t-3 ; ... (and so on).
- the above decoder can be stacked by multiple time-sequentially blocked decoding layers, and the attention mask in the self-attention module in each time-sequentially blocked decoding layer can be configured to shield the interaction between different tracking queries belonging to the same target, so as to avoid the interaction between different tracking queries of the same target.
- Mutual inhibition For example, for the self-attention module in the decoding layer, the self-attention module can be used not only for interactive processing between queries of different targets, but also for interactive processing between the same query; however, the self-attention module will shield the interactive processing between different tracking queries under the same target.
- a mask for example, the first attention mask above or the mask shown in Table 1 above
- the self-attention module in the decoding layer is the self-attention module with the mask added.
- the association relationship between different elements in the triplet can be specifically: the feature fusion module is used to perform fusion processing on the tracking query processing data output by the first decoding layer to obtain tracking query fusion data; and the second decoding layer is used to perform fusion processing on the tracking query processing data output by the first decoding layer to obtain tracking query fusion data.
- the tracking query fusion data output by the feature fusion module and the detection query processing data output by the first decoding layer are processed.
- the tracking query processing data is determined based on the first decoding layer and the above tracking query representation data; the tracking query processing data includes the processed query representation data output by the first decoding layer for each target to be tracked; the tracking query fusion data includes the fused query representation data output by the feature fusion module for each processed query representation data; the detection query processing data is determined based on the first decoding layer and the above detection query representation data.
- the tracking query processing data mentioned above refers to the Query output by the first decoding layer mentioned above for the tracked target; and the tracking query processing data is determined based on the first decoding layer and the tracking query representation data corresponding to the image to be processed mentioned above. For example, if the first decoding layer is the first decoding layer in the above decoder, the tracking query processing data output by the first decoding layer refers to the Query obtained by processing the above tracking query representation data by the first decoding layer, so that the tracking query processing data output by the first decoding layer includes the processed query representation data output by the first decoding layer for each tracked target; if the first decoding layer is the second decoding layer in the above decoder, the tracking query processing data output by the first decoding layer refers to the Query obtained by processing the Query used to describe the tracked target in the input data of the second decoding layer by the second decoding layer.
- the above detection query processing data refers to the query output by the above first decoding layer for the target to be detected; and the detection query processing data is determined according to the first decoding layer and the above detection query representation data. For example, if the first decoding layer is the first decoding layer in the above decoder, the detection query processing data output by the first decoding layer refers to the query obtained by the first decoding layer processing the above detection query representation data; if the first decoding layer is the above decoder If the first decoding layer is the second decoding layer in the decoder, the detection query processing data output by the first decoding layer refers to the Query obtained by processing the Query for describing the target to be detected existing in the input data of the second decoding layer by the second decoding layer; if the first decoding layer is the third decoding layer in the above decoder, the detection query processing data output by the first decoding layer refers to the Query obtained by processing the Query for describing the target to be detected existing in the input data of the third decoding layer by the third decoding layer
- the target A corresponds to three tracking query processing data Query A-1 , Query A-2 , and Query A-3 , then in the corresponding fusion process of Query A-1 , it is necessary to use the fusion module to interact the information carried by Query A-1 with Query A-2 and Query A-3 to obtain the fused query representation data Query A -1 ' corresponding to Query A-1; in the corresponding fusion process of Query A-2 , it is necessary to use the fusion module to interact the information carried by Query A-2 with Query A-1 and Query A- 3 to obtain the fused query representation data Query A-2 ' corresponding to Query A -2 ; in the corresponding fusion process of Query A-3 , it is necessary to use the fusion module to interact the information carried by Query A-3 with Query A-1 and Query A-2 to obtain the fused query representation data Query A-3 ' corresponding to Query A-3 .
- the present disclosure does not limit the implementation method of the above feature fusion module.
- the feature fusion module may include multiple fusion sub-modules, and different fusion sub-modules are used to perform fusion processing on the tracking queries associated with different tracked targets. It can be seen that the number of the fusion sub-modules can be determined according to the number of tracked targets.
- the feature fusion module can be used to perform the following steps: fuse the tracking queries associated with the first tracked target, obtain and output the fused query representation data corresponding to the first tracked target, so that the fused query representation data corresponding to the first tracked target includes the fusion results corresponding to each tracking query associated with the first tracked target, so that the fused query corresponding to the first tracked target
- the number of queries in the representation data is consistent with the number of tracking queries associated with the first tracked target; then the tracking queries associated with the second tracked target are fused to obtain and output fused query representation data corresponding to the second tracked target, so that the fused query representation data corresponding to the second tracked target includes the fusion results corresponding to each tracking query associated with the second tracked target, thereby making the number of queries in the fused query representation data corresponding to the second tracked target consistent with the number of tracking queries associated with the second tracked target; ... (and so on).
- the present disclosure also provides a possible implementation of the above-mentioned feature fusion module.
- the feature fusion module includes an information removal branch network and an information addition branch network; and the process of using the feature fusion module to determine the above-mentioned tracking query fusion data can specifically include the following steps 31-33.
- Step 31 Use the above information removal branch network to process the above tracking query processing data to obtain the retained information features.
- the information removal branch network is used to perform information removal processing on the input data of the above feature fusion module; and the present disclosure does not limit the information removal branch network.
- the information removal branch network may include two self-attention layers, a fully connected layer and a sigmoid layer.
- the retained information feature is used to represent the information retained after information removal processing is performed on the input data of the above feature fusion module.
- the present disclosure does not limit the determination process of the above-mentioned retained information features.
- it may specifically include the following steps 311 to 315.
- Step 311 Perform timing information extraction processing on the above tracking query processing data to obtain first timing information, so that the first timing information can represent the timing information carried in the tracking query processing data.
- step 311 can be implemented with the help of a self-attention layer.
- the information removal branch network receives the tracking query processing data output by the above first decoding layer (for example, the information removal branch network shown in FIG. 3 ).
- the first self-attention layer in the information removal branch network can perform time series information extraction processing on the tracking query processing data to obtain the first time series information (for example, the time series information 1 shown in FIG. 3 ), So that the first timing information can represent the timing information carried in the tracking query processing data.
- the It can represent the result obtained by the first decoding layer in FIG. 2 processing X t shown in FIG. 2; and Ft is used to describe the feature information predicted by the first decoding layer for some tracked targets; Represents the position information predicted by the first decoding layer for some tracked targets.
- step 312 can be implemented with the help of a self-attention layer.
- step 313 can be implemented with the help of a fully connected layer.
- step 313 Based on the relevant content of step 313 above, it can be known that for the above information removal branch network (for example, the information removal branch network shown in Figure 3), after the second self-attention layer in the information removal branch network outputs the above self-attention processing result, the fully connected layer in the information removal branch network can perform fully connected processing on the self-attention processing result to obtain a fully connected processing result.
- the above information removal branch network for example, the information removal branch network shown in Figure 3
- the fully connected layer in the information removal branch network can perform fully connected processing on the self-attention processing result to obtain a fully connected processing result.
- Step 314 Determine the information representation data to be removed according to the above full-connection processing result, so that the information representation data to be removed can indicate which information in the above tracking query processing data needs to be removed.
- step 314 can be implemented with the aid of a sigmoid layer.
- Step 315 Based on the above information characterization data to be removed, information removal processing is performed on the above tracking query processing data to obtain the retained information features.
- step 315 may specifically be: first determine the information representation data to be retained (for example, 1-Z shown in FIG3 ) based on the above information representation data to be removed; then multiply the target content feature representation data in the above tracking query processing data by the information representation data to be retained to obtain the retained information feature (for example, F t ⁇ (1-Z) shown in FIG3 ).
- the above information removal branch network may include two self-attention layers, a fully connected layer and a sigmoid layer; wherein, the first self-attention layer is used to collect time series information, the second self-attention layer is used to fuse the features of multiple queries, and the fully connected layer is used to perform fully connected processing on the output data of the second self-attention layer; the sigmoid layer is used to process the output data of the fully connected layer to obtain a gating value (that is, the data representing the information to be removed above), so that information removal processing can be performed on the input data of the information removal branch network based on the gating value (for example, Z shown in Figure 3) in the future.
- a gating value that is, the data representing the information to be removed above
- Step 32 Use the above information to add a branch network to process the above tracking query processing data to obtain the information features to be added.
- the feature of the information to be added refers to the information required as a basis for the information addition process.
- the present disclosure does not limit the above determination process of the information features to be added, for example, it may specifically include the following steps 321 and 322.
- step 321 can be implemented with the help of a self-attention layer.
- step 322 can be implemented with the help of a self-attention layer.
- the second self-attention layer in the information removal branch network can perform self-attention processing based on the second time series information and the above tracking query processing data to obtain the information feature to be added (for example, the time series information 2 shown in FIG3). ) so that the feature of the information to be added can represent the information required for the information adding process.
- the above information adding branch network may include two self-attention layers; wherein the first self-attention layer is used to collect time series information, and the second self-attention layer is used to fuse the features of multiple queries.
- the feature fusion module includes an information adding branch network, then when the feature fusion module receives the information added by the above step
- the tracking query processing data output by a decoding layer (for example, as shown in FIG. 3 )
- the information adding branch network performs information extraction processing on the tracking query processing data to obtain the information features to be added (for example, the information features shown in FIG. 3 ), so as to subsequently determine the query fusion result for the tracking query processing data based on the information feature to be added.
- Step 33 Determine the above tracking query fusion data according to the above retained information features and the above to-be-added information features.
- the feature fusion module can include two branch networks, namely, an information removal branch network and an information addition branch network.
- Each branch network includes two superimposed self-attention layers, the first self-attention layer is used to collect time series information, and the second self-attention layer is used to fuse the features of multiple queries according to the time series information.
- the output data of the second self-attention layer in the information removal branch network will further pass through a fully connected layer and a sigmoid layer to obtain a gated value, and the gated value and the output data of the information addition branch network jointly determine the fused query information (as shown in formula (2) above).
- the feature fusion module simulates the information interaction process through two branches, so that for a certain tracking query of any target, the feature fusion module can more efficiently integrate the valid information carried by other tracking queries of the target except the certain tracking query into the certain tracking query, so as to complete the self-iterative update for the certain tracking query, which is conducive to improving the interaction efficiency and effect between different tracking queries of the same target, thereby improving the target tracking effect.
- the present disclosure also provides a possible implementation of the feature fusion module.
- the self-attention layer in the feature fusion module is implemented using a second attention mask, and the second attention mask is used to shield the interaction between query representation data belonging to different targets, thereby ensuring that only interactions between all tracking queries of the same target are performed in the feature fusion module.
- the second attention mask refers to the attention mask configured for the self-attention layer in the feature fusion module above; and the second attention mask is configured to shield the interaction between the query representation data belonging to different targets.
- the present disclosure does not limit the implementation of the second attention mask.
- the second attention mask in the self-attention layer in the feature fusion module shown in FIG. 2 can be implemented using the attention mask shown in Table 2 below.
- each branch network in the above feature fusion module for example, an information removal branch network or an information addition branch network
- the branch network includes two self-attention layers
- the second attention mask used in each self-attention layer in the branch network can be implemented using an attention mask similar to that shown in Table 2 below.
- Table 2 A possible implementation of the attention mask in the self-attention layer in the feature fusion module
- the feature fusion module can be used to simultaneously perform query fusion processing for all tracked targets; and the attention mask in the self-attention layer in the feature fusion module can be configured to avoid interaction between queries of different targets, so as to ensure that only interactions between queries of the same target are performed within the feature fusion module.
- the second attention mask used by the g-th self-attention layer in the branch network is used to shield the interaction between the query representation data of different targets appearing in the input data of the g-th self-attention layer, where g is a positive integer, g ⁇ 2.
- the decoder can be composed of a stack of multiple time-blocked decoding layers, and a feature fusion module will be inserted between any two adjacent decoding layers, so as to ensure that the decoder has a better decoding effect.
- the image to be processed is first subjected to feature extraction processing to obtain a feature extraction result, so that the feature extraction result is used to characterize the image information carried by the image to be processed; the feature extraction result is then encoded to obtain at least detection query characterization data (e.g., multiple detection queries); finally, decoding processing is performed at least based on the detection query characterization data and the tracking query characterization data corresponding to the image to be processed (e.g., the tracking query determined based on multiple historical images) to obtain the target description information of the image to be processed (e.g., target content feature characterization data, target location characterization data, and category confidence, etc.), so that the target description information can represent some targets in the image to be processed.
- detection query characterization data e.g., multiple detection queries
- decoding processing is performed at least based on the detection query characterization data and the tracking query characterization data corresponding to the image to be processed (e.g., the tracking query determined based on multiple historical images) to obtain the target description information of the image to be processed
- the tracking query characterization data is determined based on the target description information of multiple historical images corresponding to the image to be processed, so that the tracking query characterization data can better represent the historical status of some targets in the image to be processed, so that the target description information determined based on the tracking query characterization data can better represent some targets in the image to be processed, which is conducive to improving the target tracking effect.
- the present disclosure also provides a possible implementation of the above image processing method.
- the image processing method may include step 41 below in addition to S1-S3 above. The execution time of step 41 is later than the execution time of S3 above.
- Step 41 After obtaining the target description information of the image to be processed, extract the target prediction result that meets the preset reference condition from the target description information to obtain the target determination result of the image to be processed.
- the preset reference condition refers to the condition required to be used when selecting high-quality target prediction results from the target description information of the image to be processed above; and the present disclosure does not limit the preset reference condition.
- it may be: a target prediction result corresponding to a detection query or a target prediction result corresponding to a tracking query of a recent historical image, and the category confidence reaches a preset threshold.
- the target determination result of the image to be processed is used to describe some targets in the image to be processed that are predicted more accurately.
- the present disclosure does not limit the determination process of the target determination result of the image to be processed above.
- it can be specifically as follows: after obtaining the target description information of the image to be processed, the target prediction result corresponding to the detection query and the target prediction result corresponding to the tracking query of the most recent historical image are extracted from the target description information, and both are used as the prediction results to be processed; then, it is determined whether the category confidence in each prediction result to be processed reaches a preset threshold. If so, it can be determined that the target described by the prediction result to be processed is relatively accurate. Therefore, the target determination result of the image to be processed can be determined based on the prediction result to be processed, so that the target determination result includes the prediction result to be processed.
- step 41 Based on the relevant content of step 41 above, it can be known that in some application scenarios, after determining the target description information of an image data, some high-quality target prediction results are extracted from the target description information to obtain the target determination result of the image data, so that the target determination result can better represent which targets exist in the image data, which is conducive to improving the target determination effect.
- the present disclosure also provides a possible implementation of the above image processing method.
- the image processing method may at least include the above S1-S3 and the following steps 51-52. Among them, the execution time of step 51 is later than the execution time of S3.
- Step 51 After obtaining the target description information of the image to be processed, extracting the target prediction result that meets the preset reference condition from the target description information.
- step 51 is similar to part of the content of step 41 above, and for the sake of brevity, it will not be repeated here.
- Step 52 Use the next frame image in the above image sequence to update the image to be processed, and use the query representation data in the target prediction result that meets the preset reference condition to update the image to be processed.
- the corresponding tracking query characterization data is returned and the above step S1 and subsequent steps are continued to be executed until all the image data in the image sequence are traversed.
- the image sequence in step 52 above refers to a sequence that needs to be processed for target tracking; and the present disclosure does not limit the implementation method of the image sequence, for example, it can be implemented using video data.
- the next frame image in the above step 52 refers to image data existing in the above image sequence, adjacent to the position of the image to be processed involved in the above step 51, and later than the position of the image to be processed involved in the step 51.
- the present disclosure does not limit the implementation method of the above step 52.
- the step 52 can be specifically as follows: after extracting the target prediction result that meets the preset reference conditions from the target description information of the image data It (for example, the data output by the multi-target tracking model shown in Figure 2), the image data It +1 can be used as a new image to be processed, and the query representation data (Query) in the target prediction result is used to update the tracking query representation data corresponding to the image to be processed, so that the target determination process for the image data It +1 can be performed based on the updated image to be processed and its corresponding tracking query representation data.
- the target prediction result for example, the data output by the multi-target tracking model shown in Figure 2
- the image data It +1 can be used as a new image to be processed
- the query representation data (Query) in the target prediction result is used to update the tracking query representation data corresponding to the image to be processed, so that the target determination process for the image data It +1 can be performed based on the updated image to be processed and its
- the target prediction result that meets the preset reference conditions can be extracted from the target description information; and then the Query in the target prediction result is used to determine the tracking query representation data corresponding to the next frame video image, so that the tracking query representation data includes the Query, which is conducive to improving the target tracking effect.
- the query in the target prediction result can be determined to be invalid.
- the invalidity lasts for more than a predetermined number of frames, the target is considered to have disappeared and will no longer be tracked, and the corresponding tracking query will be deleted.
- the image to be processed refers to an image extracted from an image sequence (for example, video data)
- the image to be processed is the first frame image in the image sequence (that is, the first frame image)
- the tracking query representation data corresponding to the image to be processed also does not exist.
- the image processing process for the image to be processed can be specifically as follows: first, feature extraction processing is performed on the image to be processed to obtain feature extraction results; then, encoding processing is performed on the feature extraction results to obtain detection query representation data; and then, according to The detection query representation data is decoded and processed to obtain the target description information of the image to be processed; however, if the image to be processed is not the first frame image in the image sequence (for example, the second frame image, the third frame image, etc.), the above S1-S3 can be used to implement the image processing process for the image to be processed.
- the image processing method can be implemented with the aid of a pre-trained target tracking model. That is, in a possible implementation, when the target tracking model includes a feature extraction module, an encoding module and a decoding module, the image processing method can be specifically as follows: the feature extraction module first performs feature extraction processing on the above image to be processed (for example, the image data I t shown in FIG. 2 ) to obtain a feature extraction result; then the encoding module performs encoding processing on the feature extraction result to obtain image feature representation data (for example, the image feature map shown in FIG. 2 ) and detection query representation data (for example, D t shown in FIG.
- the decoding module performs decoding processing based on the image feature representation data, the detection query representation data and the tracking query representation data corresponding to the image to be processed (for example, X t shown in FIG. 2 ) to obtain the target description information of the image to be processed (for example, the data output by the last decoding layer shown in FIG. 2 ).
- the present disclosure also provides a possible implementation of the training process of the target tracking model, which may specifically include the following steps 61-62.
- Step 61 Train the initial model using at least one second image data and the target detection label of each second image data to obtain a model to be optimized.
- the target detection label of the second image data is used to describe the actual location of each target in the second image data; and the present disclosure does not limit the method for obtaining the target detection label of the second image data, for example, it can be implemented by means of manual annotation.
- the target detection label of the second image data can be label information existing in the target detection training data set and having a corresponding relationship with the second image data.
- the initial model is used to represent the target tracking model to be trained in the first training stage; and the present disclosure does not limit the initial model, for example, the initial model may include a feature extraction module, an encoding module and a decoding model; and the feature extraction module refers to ResNet50 pre-trained based on the CoCo dataset, The encoding module refers to the Transformer Encoder module in the DINO detection algorithm.
- the model to be optimized refers to the model obtained by training the initial model above for target detection, so that the model to be optimized has a better target detection function. It can be seen that the model to be optimized can represent the target tracking model obtained after the first training stage.
- Step 62 Using at least one image sequence and the target tracking label of each image sequence, the decoder in the above model to be optimized is trained to obtain a target tracking model.
- the image sequence in step 62 above refers to the image data sequence required to be used in the second training phase for the target tracking model; and the present disclosure does not limit the image sequence, for example, it can refer to video data. In addition, the present disclosure does not limit the acquisition method of the image sequence, for example, the image sequence can refer to video data extracted from any target tracking training data set.
- the target tracking label of the image sequence is used to describe the actual state of each target in each image data in the image sequence; and the present disclosure does not limit the method of obtaining the target tracking label of the image sequence, for example, it can be implemented by means of manual annotation.
- the target tracking label of the image sequence can be label information existing in the target tracking training data set and having a corresponding relationship with the video data.
- a two-stage training method can be used to complete the training process for the above target tracking model, and the two-stage training method can be specifically as follows: first, the image-based target detection data is regarded as multi-target tracking training data including video data with a frame length of 1, and the multi-target tracking training data is used to perform the first training stage for the target tracking model; then, the feature extraction module and the encoding module in the trained model are fixed, and the multi-target tracking training data including multiple frames of video data is used to perform the second stage of model training for the decoding module in the model to obtain the final trained target tracking model, so that the target tracking model has better performance.
- some data enhancement methods can be used in the training process of the above target tracking model.
- the training process of the target tracking model can be specifically as follows: after obtaining at least one second image data, first perform data enhancement processing on these second image data to obtain enhanced image data; then use the enhanced image data and the target detection labels of these enhanced image data to train the initial model to obtain the model to be optimized, so as to obtain After performing data enhancement on at least one of the above image sequences to obtain an enhanced image sequence, the decoder in the above model to be optimized is trained using the enhanced image sequence and the target tracking label of the enhanced image sequence to obtain a target tracking model, which is conducive to improving the training effect of the target tracking model.
- a multi-scale training method can be used for model training. Based on this, it can be seen that for the above target tracking model, the scale parameters in the trained model can be randomly adjusted during the training process of the target tracking model to ensure that the finally trained target tracking model can be applied to process image data with any scale, which is conducive to improving the universality of the target tracking model.
- the present disclosure also provides a process for determining the training loss of the above target tracking model. For ease of understanding, it is explained below with examples.
- the process of determining the training loss of the target tracking model may specifically include the following steps 71-75.
- Step 71 Determine target description information of the first image data based on the first image data, target description information of a plurality of historical images corresponding to the first image data, and the above target tracking model.
- the first image data is used to represent the image data used in a round of training.
- the first image data may be the image data I t shown in FIG. 2 .
- the multiple historical images corresponding to the first image data refer to image data required for reference when the first image data is used for model training.
- the multiple historical images corresponding to the first image data may include image data It-1 , image data It-2 , image data It-3 , etc. shown in FIG2 .
- the target description information of the first image data is used to describe some targets of the first image data; and the determination process of the "target description information of the first image data" is similar to the determination process of the "target description information of the image to be processed” above. For the sake of brevity, it will not be repeated here.
- step 71 Based on the relevant content of step 71 above, it can be known that for a round of training process, after obtaining the first image data and the target description information of multiple historical images corresponding to the first image data, the target description information of the first image data can be determined based on this information and the target tracking model that needs to be trained, so that the performance of the target tracking model can be analyzed based on the target description information.
- Step 72 Extract the first group of information and the second group of information from the target description information of the first image data.
- the first group of information refers to the target prediction results that exist in the target description information of the first image data above and meet certain conditions; and the conditions can be set in advance, for example, the conditions can specifically be: the target prediction results corresponding to the detection query or the target prediction results corresponding to the tracking query of the most recent historical image (for example, the Query framed by the rectangular dotted box shown in Figure 2).
- the second group of information refers to other target prediction results in the target description information of the first image data except the first group of information (for example, the Query framed by the square dotted box shown in FIG. 2 ).
- the present disclosure does not limit the association relationship between the first group of information and the second group of information above, for example, the timing corresponding to the first group of information is later than the timing corresponding to the second group of information.
- the timing corresponding to the first group of information refers to the timing of the image data required to be used when predicting the first group of information. For example, if the first group of information includes the Query framed by the rectangular dashed frame shown in FIG2, the timing corresponding to the first group of information includes the timing of the image data It shown in FIG2 and the timing of the image data It -1 shown in FIG2.
- the timing corresponding to the second group of information refers to the timing of the image data required to be used when predicting the second group of information.
- the target description information can be split into two groups of information for loss calculation based on the time sequence corresponding to each target prediction result in the target description information.
- the first group of information includes the target prediction results corresponding to all detection queries and the target prediction results corresponding to the tracking query of the most recent historical image, and there is no situation in which multiple queries track the same target in the first group of information.
- the second group of information refers to the target prediction results of other tracking queries after removing the tracking query of the most recent historical image, and there is a situation in which multiple queries track the same target in the second group of information.
- Step 73 Obtain a first loss based on the first set of information and the target label data corresponding to the first image data.
- the target label data corresponding to the first image data refers to the label information required for model training using the first image data (for example, the tracking object identification label corresponding to the target, the
- the target label data may include a bounding box label corresponding to the target, a category label corresponding to the target, etc.), so that the target label data can indicate which targets actually exist in the first image data.
- the present disclosure does not limit the implementation method of the above step 73.
- it can be implemented using any existing or future bipartite matching loss.
- step 73 Based on the relevant content of step 73 above, it can be known that for a round of training process, after obtaining the first group of information (for example, the target prediction results corresponding to all detection queries and the target prediction results corresponding to the tracking query of the most recent historical image), there is no situation in the group where multiple queries track the same target, so the binary matching loss can be directly used to perform loss calculation for the first group of information; and the calculation process can be roughly as follows: for the target prediction result corresponding to the tracking query of the tracked target, directly associate the target prediction result corresponding to the query with the target with the tracking object identifier in the target label data corresponding to the first image data above based on the tracking object identifier corresponding to the tracking query, so that the tracking of the tracked target can be achieved.
- the first group of information for example, the target prediction results corresponding to all detection queries and the target prediction results corresponding to the tracking query of the most recent historical image
- the binary matching loss can be directly used to perform loss calculation for the first group of information
- the calculation process can be
- the target prediction results corresponding to the query are matched; then, the remaining targets in the target label data that are not associated with any target prediction results are regarded as newly appeared targets, and based on the matching process of bounding boxes and categories, the target prediction results associated with the newly appeared target are determined from the target prediction results corresponding to all detection queries, so that the matching process for the target prediction results corresponding to some detection queries can be realized; in addition, a new tracking object identifier can be configured for the target prediction results corresponding to the detection query that have been associated with the newly appeared target, so that the target prediction results associated with a target in the target label data can be used to perform binary matching loss calculations, and the loss values corresponding to the target prediction results associated with a target in the target label data can be obtained.
- the target prediction results that are not associated with any target in the target label data need to be regarded as the background category, so the category prediction loss corresponding to these target prediction results that are not associated with any target in the target label data can be calculated.
- the loss value corresponding to the target prediction result can be determined based on the similarity between the target position representation data in the target prediction result and the bounding box label of the associated target, and the similarity between the category confidence in the target prediction result and the category label of the associated target.
- the class confidence in the target prediction result and the class corresponding to the background can be used to determine the target prediction result.
- the similarity between the different labels is used to determine the loss value corresponding to the target prediction result.
- Step 74 Obtain a second loss based on the second set of information and the target label data corresponding to the first image data.
- the above step 74 may specifically include: determining the second loss based on the label matching result, the above second set of information, and the target label data corresponding to the above first image data.
- the label matching result is used to describe which target in the target label data corresponding to the first image data each target prediction result in the second set of information is associated with.
- the present disclosure does not limit the process of determining the above label matching results.
- it can be implemented by using any existing or future method of associating the target prediction result corresponding to the tracking query of a tracked target with the target in a certain label data (for example, the association method involved in the bipartite matching loss).
- the tracked target described in the second set of information above is roughly the same as the tracked target described in the first set of information above, so in order to better save computing resources, the label matching result above can be determined based on the matching result between the first set of information above and the target label data corresponding to the first image data.
- the following is an example.
- the second set of information above includes the target prediction result corresponding to tracking Query1
- the first set of information above includes the target prediction result corresponding to tracking Query2
- the tracking object identifier corresponding to the tracking Query1 is the same as the tracking object identifier corresponding to the tracking Query2.
- the tracking Query1 and the tracking Query2 are used to describe the same tracked target, so it is possible to determine which target in the target label data corresponding to the first image data the target prediction result corresponding to the tracking Query2 is associated with based on the matching result between the first set of information above and the target label data corresponding to the first image data; and directly determine the target associated with the target prediction result corresponding to the tracking Query2 as the target associated with the target prediction result corresponding to the tracking Query1, so that the purpose of determining the above label matching result based on the matching result between the first set of information and the target label data corresponding to the first image data can be achieved.
- the present disclosure does not limit the calculation process of the second loss above.
- the second group of information above includes at least one target prediction result (for example, the Query framed by the square dotted box shown in Figure 2)
- the calculation process of the second loss may specifically include the following steps 741-742.
- Step 741 Determine the loss corresponding to each target prediction result based on the above label matching result, the above second group of information, and the target label data corresponding to the above first image data.
- the present disclosure does not limit the implementation method of the above step 741.
- it can be specifically as follows: first, based on the above label matching result, the above second group of information, and the target label data corresponding to the above first image data, determine the label information corresponding to each target prediction result in the second group of information (for example, a bounding box label + a category label); then, based on the nth target prediction result in the second group of information and the label information corresponding to the nth target prediction result, determine the loss corresponding to the nth target prediction result (for example, based on the similarity between the target position representation data in the nth target prediction result and the bounding box label in the label information corresponding to the nth target prediction result, and based on the similarity between the category confidence in the nth target prediction result and the category label in the label information corresponding to the nth target prediction result, determine the loss corresponding to the nth target prediction result), where n is a positive integer, n ⁇ N, and N
- Step 742 Determine a second loss based on the average value of the losses corresponding to at least one target prediction result in the second set of information above.
- step 74 Based on the relevant content of step 74 above, it can be known that for the second group of information above, there are multiple queries in the group tracking the same target, so the present disclosure can use the tracking consistency loss (tracking object consistency loss) to determine the second loss corresponding to the second group of information.
- the tracking consistency loss maintains the same loss calculation method as the above binary matching loss, but in the calculation process of the tracking consistency loss, the label information associated with each target prediction result in the second group of information can be directly determined based on the matching result between the first group of information above and the target label data corresponding to the first image data.
- the second loss calculated based on the second group of information is averaged within this group, rather than being averaged after the losses corresponding to all target prediction results in the first group are merged, so that the adverse effects caused by too many target prediction results in the second group of information can be effectively avoided (for example, the loss determined based on the target prediction result corresponding to the detection query in the first group is diluted, etc.), which is conducive to improving the model training effect.
- Step 75 Determine the training loss of the target tracking model based on the first loss and the second loss.
- step 75 may specifically be: directly determining the sum of the first loss and the second loss as the training loss of the target tracking model.
- step 75 may specifically be: performing a weighted sum of the first loss and the second loss to obtain the training loss of the target tracking model.
- step 75 may specifically be: performing a weighted sum of the first loss and the second loss to obtain the training loss of the target tracking model. Combined with the second loss, the training loss of the target tracking model is obtained.
- the training loss of the target tracking model can be determined according to the sum of the bipartite matching loss and the tracking consistency loss, so that the training loss can better represent the performance of the model, thereby helping to improve the model training effect.
- the present disclosure does not limit the application scenarios of the training loss shown in steps 71 to 75 above.
- the training loss shown in steps 71 to 75 is applied to step 62 above (that is, the second training stage)
- the "target tracking model" appearing in steps 71 to 75 can be replaced by the model to be optimized above.
- the target prediction results corresponding to the tracking queries involved in the current frame (for example, the first image data described above) and the target prediction results corresponding to the detection queries associated with the newly appeared targets are all assigned corresponding tracking object identifiers. Therefore, for any target prediction result assigned with a tracking object identifier, if it is determined that the IOU between the target prediction result assigned with the tracking object identifier and the target associated with it is greater than a predetermined threshold, it can be determined that the target has been well tracked. Therefore, the query in the target prediction result associated with the target can be sent to the subsequent frame as a historical tracking query (that is, a positive sample) to participate in tracking according to a preset probability.
- a historical tracking query that is, a positive sample
- the query in the target prediction result corresponding to the detection query that is not assigned with a tracking object identifier and whose target position is close to the tracking query will be used as a noise tracking query (that is, a negative sample) with a certain probability, associated with a non-existent target (for example, a virtual target) and sent to the subsequent frame as a historical tracking query to participate in training, thereby increasing the robustness of the model to noise.
- a noise tracking query that is, a negative sample
- a non-existent target for example, a virtual target
- this target tracking model is an end-to-end multi-target tracking algorithm based on multiple queries. It uses multiple historical queries of each target to construct a collaborative tracking query for the target in the current frame, and jointly tracks the target to improve the reliability of the features, so that the target tracking model can achieve better performance in both high frame rate videos and low frame rate videos, and the processing speed is relatively fast.
- the embodiment of the present disclosure also provides an image processing device, which is explained and illustrated in conjunction with Figure 4.
- Figure 4 is a schematic diagram of the structure of an image processing device provided by the embodiment of the present disclosure. It should be noted that for the technical details of the image processing device provided by the embodiment of the present disclosure, please refer to the relevant content of the image processing method above.
- the image processing device 400 provided by the embodiment of the present disclosure includes:
- An extraction unit 401 is used to perform feature extraction processing on the image to be processed to obtain a feature extraction result
- An encoding unit 402 is used to encode the feature extraction result to obtain detection query representation data
- the decoding unit 403 is used to perform decoding processing based on the detection query representation data and the tracking query representation data corresponding to the image to be processed to obtain target description information of the image to be processed; the tracking query representation data is determined based on the target description information of multiple historical images corresponding to the image to be processed; the target description information of the image to be processed is used to describe at least one target in the image to be processed.
- the encoding unit 402 is specifically used to: encode the feature extraction result to obtain image feature representation data and detection query representation data;
- the decoding unit 403 is specifically used to: perform decoding processing according to the image feature representation data, the detection query representation data and the tracking query representation data corresponding to the image to be processed, so as to obtain the target description information of the image to be processed.
- the multiple historical images include a first category of images and a second category of images; the process of determining the tracking query representation data includes: extracting first tracking target description information from the target description information of the first category of images; and determining the tracking query representation data based on the first tracking target description information and the target description information of the second category of images.
- the process of determining the tracking query characterization data includes: determining the second tracking target description information based on the target position characterization data in the first tracking target description information and the target content feature characterization data in the target description information of the second type of image; and determining the tracking query characterization data based on the second tracking target description information and the first tracking target description information.
- the image to be processed and multiple historical images corresponding to the image to be processed belong to the same video data; the timing of the image to be processed is later than the timing of the first category of images; the timing of the first category of images is later than the timing of the second category of images.
- the decoding unit 403 is specifically used to: determine the query representation data to be processed based on the detection query representation data and the tracking query representation data; and perform decoding processing based on the query representation data to be processed to obtain target description information of the image to be processed.
- the decoding unit 403 is specifically configured to: The query characterization data and the tracking query characterization data are used to determine the query characterization data to be processed; decoding is performed according to the image feature characterization data and the query characterization data to be processed to obtain the target description information of the image to be processed.
- the target description information of the image to be processed is determined using a pre-trained decoder; the decoder includes at least one decoding layer, and the self-attention module in the decoding layer is implemented using a first attention mask, which is used to shield the interaction between the query representation data corresponding to the same target in different historical images.
- the at least one target includes one or more targets to be tracked;
- the at least one decoding layer includes a first decoding layer and a second decoding layer;
- the decoder also includes a feature fusion module;
- the feature fusion module is used to perform fusion processing on the tracking query processing data output by the first decoding layer to obtain tracking query fusion data;
- the tracking query processing data is determined based on the first decoding layer and the tracking query representation data;
- the tracking query processing data includes the processed query representation data output by the first decoding layer for each of the targets to be tracked;
- the tracking query fusion data includes the fused query representation data output by the feature fusion module for each of the processed query representation data;
- the second decoding layer is used to process the tracking query fusion data and the detection query processing data output by the first decoding layer;
- the detection query processing data is determined based on the first decoding layer and the detection query representation data.
- the feature fusion module includes an information removal branch network and an information addition branch network;
- the process of determining the tracking query fusion data includes: using the information removal branch network to process the tracking query processing data to obtain the retained information characteristics; using the information addition branch network to process the tracking query processing data to obtain the information characteristics to be added; determining the tracking query fusion data based on the retained information characteristics and the information characteristics to be added.
- the process of determining the retained information features includes: performing timing information extraction processing on the tracking query processing data to obtain first timing information; performing self-attention processing on the first timing information and the tracking query processing data to obtain a self-attention processing result; performing full connection processing on the self-attention processing result to obtain a full connection processing result; determining information representation data to be removed based on the full connection processing result; and performing information removal processing on the tracking query processing data based on the information representation data to be removed to obtain the retained information features.
- the process of determining the information features to be added includes: performing timing information extraction processing on the tracking query processing data to obtain second timing information; performing self-attention processing on the second timing information and the tracking query processing data to obtain the information features to be added.
- the self-attention layer in the feature fusion module is implemented using a second attention mask, where the second attention mask is used to shield the interaction between query representation data belonging to different targets.
- the target description information of the image to be processed is determined using a pre-trained target tracking model
- the process of determining the training loss of the target tracking model includes: determining the target description information of the first image data based on the first image data, the target description information of multiple historical images corresponding to the first image data, and the target tracking model; extracting a first group of information and a second group of information from the target description information of the first image data; obtaining a first loss based on the first group of information and the target label data corresponding to the first image data; obtaining a second loss based on the second group of information and the target label data; and determining the training loss of the target tracking model based on the first loss and the second loss.
- the timing corresponding to the first group of information is later than the timing corresponding to the second group of information.
- the process of determining the second loss includes: determining the second loss based on a label matching result, the second set of information, and the target label data; the label matching result is determined based on a matching result between the first set of information and the target label data.
- the second set of information includes at least one target prediction result
- the process of determining the second loss includes: determining the loss corresponding to each of the target prediction results based on the label matching result, the second set of information, and the target label data; and determining the second loss based on the average value of the loss corresponding to at least one target prediction result.
- the target description information of the image to be processed is determined using a pre-trained target tracking model
- the training process of the target tracking model includes: using at least one second image data and the target detection label of each second image data to train the initial model to obtain the model to be optimized;
- the target tracking model is obtained by training a decoding module in the model to be optimized using at least one image sequence and a target tracking label of each image sequence.
- the at least one target includes one or more targets to be tracked; and the tracking query representation data includes query representation data corresponding to each of the targets to be tracked under multiple historical images.
- the image to be processed refers to an image extracted from an image sequence;
- the target description information includes at least one target prediction result;
- the image processing device 400 further includes:
- a screening unit used to extract target prediction results that meet preset reference conditions from the target description information
- An updating unit is used to update the image to be processed using the next frame image in the image sequence, update the tracking query characterization data using the query characterization data in the target prediction result that meets the preset reference condition, and continue to execute the step of performing feature extraction processing on the image to be processed.
- the image sequence is video data.
- the image to be processed is first subjected to feature extraction processing to obtain a feature extraction result, so that the feature extraction result is used to characterize the image information carried by the image to be processed; the feature extraction result is then encoded to obtain detection query characterization data (e.g., multiple detection queries); finally, the detection query characterization data and the tracking query characterization data corresponding to the image to be processed (e.g., the tracking query determined based on multiple historical images) are decoded to obtain the target description information of the image to be processed (e.g., target content feature characterization data, target location characterization data, and category confidence, etc.), so that the target description information can represent some targets in the image to be processed.
- detection query characterization data e.g., multiple detection queries
- the detection query characterization data and the tracking query characterization data corresponding to the image to be processed e.g., the tracking query determined based on multiple historical images
- the target description information of the image to be processed e.g., target content feature characterization data, target location characterization data,
- the tracking query characterization data is determined based on the target description information of multiple historical images corresponding to the image to be processed, the tracking query characterization data can better represent the historical status of some targets in the image to be processed, so that the target description information determined based on the tracking query characterization data can better represent some targets in the image to be processed, which is conducive to improving the target tracking effect.
- the present disclosure also provides an electronic device, the device comprising a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes the present disclosure. Any embodiment of the image processing method provided.
- the terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- PDAs personal digital assistants
- PADs tablet computers
- PMPs portable multimedia players
- vehicle-mounted terminals such as vehicle-mounted navigation terminals
- fixed terminals such as digital TVs, desktop computers, etc.
- the electronic device shown in FIG5 is only an example and should not bring any limitation to the functions and scope of use of
- the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 to a random access memory (RAM) 503.
- a processing device 501 e.g., a central processing unit, a graphics processing unit, etc.
- RAM random access memory
- Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503.
- the processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504.
- An input/output (I/O) interface 505 is also connected to the bus 504.
- the following devices may be connected to the I/O interface 505: input devices 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 509.
- the communication devices 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data.
- FIG. 5 shows an electronic device 500 with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM502.
- the processing device 501 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
- the present disclosure also provides a computer-readable medium, in which instructions or computer programs are stored.
- the instructions or computer programs are executed on a device, the device executes any implementation of the image processing method provided by the present disclosure.
- the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
- This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
- the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
- the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
- the client and server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network).
- HTTP Hyper Text Transfer Protocol
- Examples of communication networks include a local area network ("LAN”), a wide area network ("WAN”), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
- the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
- the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages.
- the program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
- LAN local area network
- WAN wide area network
- Internet service provider e.g., AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
- each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit/module does not, in some cases, constitute a limitation on the unit itself.
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chips
- CPLDs complex programmable logic devices
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or Semiconductor system, device or apparatus, or any suitable combination of the above.
- machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM compact disk read-only memory
- magnetic storage devices or any suitable combination of the above.
- At least one (item) means one or more, and “plurality” means two or more.
- “And/or” is used to describe the association relationship of associated objects, indicating that three relationships may exist.
- a and/or B can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural.
- the character “/” generally indicates that the objects associated before and after are in an “or” relationship.
- At least one of the following” or similar expressions refers to any combination of these items, including any combination of single or plural items.
- At least one of a, b or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c", where a, b, c can be single or multiple.
- the steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, a software module executed by a processor, or a combination of the two.
- the software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a mobile disk, a CD-ROM, or any other form of storage medium known in the technical field. It should be noted that the present disclosure is not limited to the implementation of step 33 above.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Image Analysis (AREA)
Abstract
本公开的实施例提供一种图像处理方法、装置、电子设备、计算机可读介质,该方法包括:先对该待处理图像进行特征提取处理,得到特征提取结果;再对该特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据;最后,依据该图像特征表征数据、该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据进行解码处理,得到该待处理图像的目标描述信息,以使该目标描述信息能够表示出该待处理图像中一些目标。
Description
本申请要求于2023年6月29日递交的中国专利申请第202310787598.2号的优先权,在此全文引用上述中国专利申请公开的内容以作为本申请的一部分。
本公开的实施例涉及一种图像处理方法、装置、电子设备、计算机可读介质。
目标跟踪用于从视频数据中分析出一些目标的运动状态信息(比如,在每帧视频图像中所处位置等)。
发明内容
本公开提供了一种图像处理方法、装置、电子设备、计算机可读介质。
为了实现上述目的,本公开提供的技术方案如下:
本公开提供一种图像处理方法,所述方法包括:
对待处理图像进行特征提取处理,得到特征提取结果;
对所述特征提取结果进行编码处理,得到检测查询表征数据;
依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息;所述跟踪查询表征数据是依据所述待处理图像对应的多个历史图像的目标描述信息所确定的;所述待处理图像的目标描述信息用于描述所述待处理图像中的至少一个目标。
在一种可能的实施方式下,所述对所述特征提取结果进行编码处理,得到检测查询表征数据,包括:
对所述特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据;
所述依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表
征数据进行解码处理,得到所述待处理图像的目标描述信息,包括:
依据所述图像特征表征数据、所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述多个历史图像包括第一类图像和第二类图像;
所述跟踪查询表征数据的确定过程,包括:
从所述第一类图像的目标描述信息中提取第一跟踪目标描述信息;
依据所述第一跟踪目标描述信息以及所述第二类图像的目标描述信息,确定所述跟踪查询表征数据。
在一种可能的实施方式下,所述依据所述第一跟踪目标描述信息以及所述第二类图像的目标描述信息,确定所述跟踪查询表征数据,包括:
依据所述第一跟踪目标描述信息中的目标位置表征数据以及所述第二类图像的目标描述信息中的目标内容特征表征数据,确定第二跟踪目标描述信息;
根据所述第二跟踪目标描述信息和所述第一跟踪目标描述信息,确定所述跟踪查询表征数据。
在一种可能的实施方式下,所述待处理图像与所述待处理图像对应的多个历史图像属于同一个视频数据;
所述待处理图像的时序晚于所述第一类图像的时序;
所述第一类图像的时序晚于所述第二类图像的时序。
在一种可能的实施方式下,所述依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息,包括:
依据所述检测查询表征数据和所述跟踪查询表征数据,确定待处理查询表征数据;
依据所述待处理查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述依据所述图像特征表征数据、所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得
到所述待处理图像的目标描述信息,包括:
依据所述检测查询表征数据和所述跟踪查询表征数据,确定待处理查询表征数据;
依据所述图像特征表征数据和所述待处理查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的解码器所确定的;
所述解码器包括至少一个解码层,所述解码层中的自注意力模块是利用第一注意力掩模进行实现的,所述第一注意力掩模用于屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互。
在一种可能的实施方式下,所述至少一个目标包括一个或者多个待跟踪目标;
所述至少一个解码层包括第一解码层和第二解码层;
所述解码器还包括特征融合模块;
所述特征融合模块用于针对由所述第一解码层输出的跟踪查询处理数据进行融合处理,得到跟踪查询融合数据;所述跟踪查询处理数据是依据所述第一解码层以及所述跟踪查询表征数据所确定的;所述跟踪查询处理数据包括由所述第一解码层针对各所述待跟踪目标分别输出的处理后的查询表征数据;所述跟踪查询融合数据包括由所述特征融合模块针对各所述处理后的查询表征数据分别输出的融合后的查询表征数据;
所述第二解码层用于针对所述跟踪查询融合数据以及由所述第一解码层输出的检测查询处理数据进行处理;所述检测查询处理数据是依据所述第一解码层以及所述检测查询表征数据所确定的。
在一种可能的实施方式下,所述特征融合模块包括信息移除分支网络和信息添加分支网络;
所述跟踪查询融合数据的确定过程,包括:
利用所述信息移除分支网络对所述跟踪查询处理数据进行处理,得到保留信息特征;
利用所述信息添加分支网络对所述跟踪查询处理数据进行处理,得到待添加信息特征;
根据所述保留信息特征和所述待添加信息特征,确定所述跟踪查询融合数据。
在一种可能的实施方式下,所述保留信息特征的确定过程,包括:
对所述跟踪查询处理数据进行时序信息提取处理,得到第一时序信息;
对所述第一时序信息与所述跟踪查询处理数据进行自注意力处理,得到自注意力处理结果;
对所述自注意力处理结果进行全连接处理,得到全连接处理结果;
依据所述全连接处理结果,确定待移除信息表征数据;
依据所述待移除信息表征数据,对所述跟踪查询处理数据进行信息移除处理,得到所述保留信息特征。
在一种可能的实施方式下,所述待添加信息特征的确定过程,包括:
对所述跟踪查询处理数据进行时序信息提取处理,得到第二时序信息;
对所述第二时序信息与所述跟踪查询处理数据进行自注意力处理,得到所述待添加信息特征。
在一种可能的实施方式下,所述特征融合模块中的自注意力层是利用第二注意力掩模进行实现的,所述第二注意力掩模用于屏蔽属于不同目标的查询表征数据之间的交互。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;
所述目标跟踪模型的训练损失的确定过程,包括:
依据第一图像数据、所述第一图像数据对应的多个历史图像的目标描述信息、以及所述目标跟踪模型,确定所述第一图像数据的目标描述信息;
从所述第一图像数据的目标描述信息中提取第一组信息和第二组信息;
依据所述第一组信息和所述第一图像数据对应的目标标签数据,得到第一损失;
依据所述第二组信息和所述目标标签数据,得到第二损失;
依据所述第一损失和所述第二损失,确定所述目标跟踪模型的训练损失。
在一种可能的实施方式下,所述第一组信息对应的时序晚于所述第二组信息对应的时序。
在一种可能的实施方式下,所述依据所述第二组信息和所述目标标签数
据,得到第二损失,包括:
依据标签匹配结果、所述第二组信息、以及所述目标标签数据,确定所述第二损失;所述标签匹配结果是根据所述第一组信息与所述目标标签数据之间的匹配结果所确定的。
在一种可能的实施方式下,所述第二组信息包括至少一个目标预测结果;
所述第二损失的确定过程,包括:
依据所述标签匹配结果、所述第二组信息、以及所述目标标签数据,确定各所述目标预测结果对应的损失;
依据所述至少一个目标预测结果对应的损失的平均值,确定所述第二损失。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;
所述目标跟踪模型的训练过程,包括:
利用至少一个第二图像数据以及各所述第二图像数据的目标检测标签,对初始模型进行训练,得到待优化模型;
利用至少一个图像序列以及各所述图像序列的目标跟踪标签,对所述待优化模型中的解码模块进行训练,得到所述目标跟踪模型。
在一种可能的实施方式下,所述至少一个目标包括一个或者多个待跟踪目标;所述跟踪查询表征数据包括各所述待跟踪目标在多个历史图像下对应的查询表征数据。
在一种可能的实施方式下,所述待处理图像是指从图像序列中抽取的一个图像;所述目标描述信息包括至少一个目标预测结果;
所述得到所述待处理图像的目标描述信息之后,所述方法还包括:
从所述目标描述信息中提取满足预设参考条件的目标预测结果;
利用所述图像序列中的下一帧图像更新所述待处理图像,利用所述满足预设参考条件的目标预测结果中的查询表征数据,更新所述跟踪查询表征数据,并继续执行所述对待处理图像进行特征提取处理的步骤。
在一种可能的实施方式下,所述图像序列为视频数据。
本公开提供了一种图像处理装置,包括:
提取单元,被配置为对待处理图像进行特征提取处理,得到特征提取结
果;
编码单元,被配置为对所述特征提取结果进行编码处理,得到检测查询表征数据;
解码单元,被配置为依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息;所述跟踪查询表征数据是依据所述待处理图像对应的多个历史图像的目标描述信息所确定的;所述待处理图像的目标描述信息用于描述所述待处理图像中的至少一个目标。
本公开提供了一种电子设备,所述设备包括:处理器和存储器;
所述存储器,被配置为存储指令或计算机程序;
所述处理器,被配置为执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行本公开提供的图像处理方法。
本公开提供了一种计算机可读介质,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行本公开提供的图像处理方法。
本公开提供了一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行本公开提供的图像处理方法的程序代码。
为了更清楚地说明本公开实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开中记载的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。
图1为本公开实施例提供的一种图像处理方法的流程图;
图2为本公开实施例提供的一种目标跟踪模型的示意图;
图3为本公开实施例提供的一种特征融合模块的示意图;
图4为本公开实施例提供的一种图像处理装置的结构示意图;
图5为本公开实施例提供的一种电子设备的结构示意图。
经研究发现,对于上文基于卡尔曼滤波的两阶段跟踪方案来说,因该方案依赖于卡尔曼滤波器对于运动的建模,其要求目标的运动速度和方向不能变化太快。但是,因低帧率视频中的目标在帧间变化更加剧烈(比如,会出现快速的速度、方向、外观的变化等),如此导致该方案在低帧率视频下的目标跟踪效果严重下降,从而导致该方案不适于处理低帧率视频。
基于上述研究,为了更好地提高目标跟踪效果,本公开提供了一种图像处理方法,该方法包括:对于待处理图像(比如,一个视频数据中的某一帧视频图像)来说,先对该待处理图像进行特征提取处理,得到特征提取结果,以使该特征提取结果用于表征该待处理图像所携带的图像信息;再对该特征提取结果进行编码处理,得到检测查询表征数据(比如,多个检测Query);最后,依据该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据(比如,基于多个历史图像所确定的跟踪Query)进行解码处理,得到该待处理图像的目标描述信息(比如,目标内容特征表征数据、目标位置表征数据、以及类别置信度等),以使该目标描述信息能够表示出该待处理图像中一些目标。其中,因该跟踪查询表征数据是依据该待处理图像对应的多个历史图像的目标描述信息所确定的,以使该跟踪查询表征数据能够更好地表示出该待处理图像中一些目标的历史状态,从而使得基于该跟踪查询表征数据所确定的目标描述信息能够更好地表示出该待处理图像中一些目标,如此有利于提高目标跟踪效果。
另外,当本公开提供的图像处理方法用于实现针对视频数据的目标跟踪处理时,因该图像处理方法需要参考多个历史图像的目标描述信息(比如,Query),以确定某一个视频图像中的目标,以使在针对该视频图像的目标确定过程中能够从这些历史图像的目标描述信息中获取到比较多的可参考内容(比如,目标在不同历史图像中所处状态等),如此能够有效地克服因视频数据的帧率较低而导致的问题,从而能够有效地提高针对低帧率视频的目标跟踪效果,进而使得本公开提供的图像处理方法适用于针对具有任何帧率的视频数据进行目标跟踪处理,如此有利于提高该图像处理方法在目标跟踪领域中的普适性。
此外,本公开不限定图像处理方法的执行主体,例如,本公开实施例提供
的图像处理方法的任一实施方式可以应用于终端设备或服务器等具有数据处理功能的设备。又如,本公开实施例提供的图像处理方法的任一实施方式也可以借助不同设备(例如,终端设备与服务器、两个终端设备、或者两个服务器)之间的数据通信过程进行实现。其中,终端设备可以为智能手机、计算机、个人数字助理(Personal Digital Assitant,PDA)或平板电脑等。服务器可以为独立服务器、集群服务器或云服务器。
为了使本技术领域的人员更好地理解本公开方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
为了更好地理解本公开所提供的技术方案,下面先结合一些附图对本公开提供的图像处理方法进行说明。如图1所示,本公开实施例提供的图像处理方法,包括下文S1-S3。其中,该图1为本公开实施例提供的一种图像处理方法的流程图。
S1:对待处理图像进行特征提取处理,得到特征提取结果。
其中,待处理图像是指需要进行目标确定处理的图像数据(比如,图2所示的第t帧视频图像It)。t为正整数。需要说明的是,本公开不限定目标的实施方式,例如,该目标可以是指任意一种可检测的对象(比如,人、动物、植物、物体等)。
另外,本公开不限定上文待处理图像,比如,在一些应用场景(比如,针对某个图像序列进行目标跟踪处理等场景)下,该待处理图像是指从该图像序列中抽取的一个图像(比如,位于第t个排列位置的图像数据)。其中,该图像序列用于记录按照某种顺序(比如,时序)进行排列的多个图像数据;而且本公开不限定该图像序列的实施方式,例如,其可以采用现有的或者未来出现的任意一种图像序列(比如,视频数据)进行实施。
特征提取结果用于表示上文待处理图像所携带的图像信息。
另外,本公开不限定上文特征提取结果的确定过程(也就是,S1的实施方式),比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像数据进行特征提取处理的方法(比如,借助特征提取器等方式)进行实施。
可见,在一些应用场景下,上文S1具体可以为:将待处理图像输入预先构建的特征提取器,得到该特征提取器输出的特征提取结果,以使该特征提取结果能够表示出该待处理图像所携带的图像信息。其中,该特征提取器用于针对该特征提取器的输入数据进行特征提取处理;而且本公开不限定该特征提取器的实施方式,例如,该特征提取器可以采用现有的或者未来出现的任意一种具有特征提取功能的设备进行实施。又如,该特征提取器可以采用卷积神经网络(Convolutional Neural Networks,CNN)进行实施。还如,该特征提取器可以采用下文目标跟踪模型中的特征提取模块进行实施。
需要说明的是,本公开不限定上段中CNN的实施方式,例如,其可以采用基于CoCo数据集预训练的ResNet50进行实施。
基于上述两段内容可知,在一种可能的实施方式下,当本公开提供的图像处理方法借助预先训练好的目标跟踪模型(比如,图2所示的多目标跟踪模块)进行实现时,上文S1具体可以为:利用该目标跟踪模型中的特征提取模块(比如,图2所示的CNN),对上文待处理图像进行特征提取处理,得到该待处理图像对应的特征提取结果。
基于上文S1的相关内容可知,对于待处理图像(比如,图2所示的第t帧视频图像It)来说,在获取到该待处理图像之后,对该待处理图像进行特征提取处理,得到特征提取结果(比如,由图2所示的CNN所输出的数据),以使该特征提取结果能够表示出该待处理图像所携带的图像信息,以便后续能够基于该特征提取结果,确定该待处理图像中的一些目标。
S2:对特征提取结果进行编码处理,得到检测查询表征数据。
其中,检测查询表征数据用于表征上文待处理图像中一些目标;而且本公开不限定该检测查询表征数据,比如,该检测查询表征数据可以包括该待处理图像中多个目标对应的查询表征数据(例如,Query)。其中,该目标对应的查询表征数据用于描述该待处理图像中一个目标。可见,在一种可能的实施方式下,该检测查询表征数据可以包括一个或者多个目标在该待处理图像下对应的查询表征数据。其中,该目标在该待处理图像下对应的查询表征数据用于表征该目标在该待处理图像下所处状态。
另外,本公开不限定上段中目标对应的查询表征数据的实施方式,例如,其可以包括该目标对应的目标内容特征表征数据以及该目标对应的目标位置
表征数据。其中,该目标内容特征表征数据用于表征该目标在该待处理图像中所呈现的特征(比如,在姿态、衣服颜色等方面所呈现的特征);而且本公开不限定该目标内容特征表征数据的实施方式,例如,该目标内容特征表征数据可以采用特征向量进行实施。该目标位置表征数据用于表征该目标在该待处理图像中所处位置;而且本公开不限定该目标内容特征表征数据的实施方式,比如,其可以采用边界框进行实施。
此外,本公开不限定上文S2中编码处理的实施方式,例如,其可以采用现有的或者未来出现的任意一种能够进行编码处理的方法(比如,借助编码器等方式)进行实施。
可见,在一种可能的实施方式下,上文S2具体可以为:将上文特征提取结果输入预先构建的编码器,得到该编码器输出的检测查询表征数据。其中,该编码器用于针对该编码器的输入数据进行编码处理;而且本公开不限定该编码器的实施方式,例如,该编码器可以采用现有的或者未来出现的任意一种具有编码功能的设备进行实施。又如,该编码器可以采用Transformer Encoder进行实施。还如,该编码器可以采用下文目标跟踪模型中的编码模块进行实施。
需要说明的是,本公开不限定上段中Transformer Encoder的实施方式,例如,其可以采用DINO(DETR with Improved DeNoising Anchor Boxes)检测算法中的Transformer Encoder模块进行实施。
基于上述两段内容可知,在一种可能的实施方式下,当本公开提供的图像处理方法借助预先训练好的目标跟踪模型(比如,图2所示的多目标跟踪模块)进行实现时,上文S2具体可以为:在利用该目标跟踪模型中的特征提取模块从上文待处理图像中确定出特征提取结果之后,由该目标跟踪模型中的编码模块针对该特征提取结果进行编码处理,得到该待处理图像对应的检测查询表征数据(比如,图2所示的Dt)。
需要说明的是,对于图2所示的Dt来说,Dt用于表示图2所示的图像数据It中一些目标;而且该Dt可以采用下文公式(1)进行表示。
式中,Dt表示通过针对上文图像数据It进行处理(比如,特征提取处理+编码处理)所确定的Query;表示该图像数据It中第1个目标对应的Query,
而且该包括该第1个目标对应的目标内容特征表征数据以及该第1个目标对应的目标位置表征数据;表示该图像数据It中第2个目标对应的Query,而且该包括该第2个目标对应的目标内容特征表征数据以及该第2个目标对应的目标位置表征数据;……(以此类推);表示该图像数据It中第Ndet个目标对应的Query,而且该包括该第Ndet个目标对应的目标内容特征表征数据以及该第Ndet个目标对应的目标位置表征数据;Ndet表示预先设定的Query数目,以使该Ndet用于表示该Dt所涉及的Query个数。
实际上,在一些应用场景下,为了更好地提高目标跟踪效果,本公开还提供了上文S2中编码处理的一种可能的实施方式,在该实施方式中,该S2具体可以为:对特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据。其中,该图像特征表征数据用于表征上文待处理图像所携带的信息;而且本公开不限定该图像特征表征数据的实施方式,例如,其可以采用现有的或者未来出现的任意一种针对图像数据所确定的特征编码向量(比如,图2所示的图像特征图)进行实施。
另外,本公开不限定上段中编码处理的实施方式,例如,该编码处理具体可以为:将上文特征提取结果输入预先构建的编码器,得到该编码器输出的图像特征表征数据以及检测查询表征数据。又如,当本公开提供的图像处理方法借助预先训练好的目标跟踪模型(比如,图2所示的多目标跟踪模块)进行实现时,该编码处理具体可以为:在利用该目标跟踪模型中的特征提取模块从上文待处理图像中确定出特征提取结果之后,由该目标跟踪模型中的编码模块针对该特征提取结果进行编码处理,得到该待处理图像对应的图像特征表征数据(比如,图2所示的图像特征图)以及该待处理图像对应的检测查询表征数据(比如,图2所示的Dt)。
基于上段内容可知,在一种可能的实施方式下,对于待处理图像(比如,图2所示的第t帧视频图像It)来说,在获取到该待处理图像对应的特征提取结果(比如,由图2所示的CNN所输出的数据)之后,对该特征提取结果进行编码处理,得到图像特征表征数据(比如,图2所示的图像特征图)以及检测查询表征数据(比如,图2所示的Dt),以便后续能够基于该图像特征表征数据以及该检测查询表征数据,确定该待处理图像中的一些目标。
基于上文S2的相关内容可知,对于待处理图像(比如,图2所示的第t
帧视频图像It)来说,在获取到该待处理图像对应的特征提取结果(比如,由图2所示的CNN所输出的数据)之后,对该特征提取结果进行编码处理,至少得到检测查询表征数据(比如,图2所示的Dt),以便后续能够基于该检测查询表征数据,确定该待处理图像中的一些目标。
S3:依据检测查询表征数据以及待处理图像对应的跟踪查询表征数据进行解码处理,得到该待处理图像的目标描述信息;该跟踪查询表征数据是依据该待处理图像对应的多个历史图像的目标描述信息所确定的;该待处理图像的目标描述信息用于描述该待处理图像中的至少一个目标。
其中,待处理图像对应的跟踪查询表征数据是指在针对该待处理图像进行目标确定处理时所需依据的跟踪Query,以使该跟踪查询表征数据能够表示出该待处理图像中一些已跟踪到的目标(比如,下文“待跟踪目标”)在历史阶段所呈现的状态。
另外,待处理图像对应的跟踪查询表征数据是依据该待处理图像对应的多个历史图像的目标描述信息所确定的。其中,该历史图像的时序早于该待处理图像的时序;而且本公开不限定该待处理图像对应的多个历史图像,比如,当该待处理图像为图2所示的第t帧视频图像It时,该待处理图像对应的多个历史图像可以包括图2所示的第t-1帧视频图像It-1、第t-2帧视频图像It-2、第t-3帧视频图像It-3等。该历史图像的目标描述信息是指在针对该历史图像进行目标确定处理时所得到的结果,以使该历史图像的目标描述信息能够描述出该历史图像中一些目标。
需要说明的是,本公开不限定目标描述信息的实施方式,例如,该目标描述信息可以包括一些目标对应的查询表征数据(比如,该目标对应的特征向量+该目标对应的边界框)。又如,该目标描述信息可以还包括这些目标对应的类别置信度。该类别置信度用于表示该目标以多大概率属于某一类别。
还需要说明的是,本公开不限定上文“待处理图像对应的多个历史图像”中“多个”的实施方式,比如,其可以为2个、3个或者其他大于1的正整数。
此外,本公开不限定上文待处理图像对应的跟踪查询表征数据的确定过程的实施方式,例如,当该待处理图像对应的多个历史图像包括R个图像数据时,该跟踪查询表征数据的确定过程具体可以包括:将第r个图像数据的目标描述信息中存在的部分或者全部目标对应的查询表征数据(Query),确定
为该第r个图像数据对应的跟踪Query,r为正整数,r≤R,R为正整数;再将R个图像数据对应的跟踪Query进行集合,得到该跟踪查询表征数据,以使该跟踪查询表征数据包括该待处理图像对应的多个历史图像中一些已跟踪到的目标对应的跟踪Query。
实际上,为了更好地提高目标确定效果,本公开还提供了上文待处理图像对应的跟踪查询表征数据的确定过程的一种可能的实施方式,在该实施方式中,当该待处理图像对应的多个历史图像包括第一类图像(比如,图2所示的第t-1帧视频图像It-1)和第二类图像(比如,图2所示的第t-2帧视频图像It-2、第t-3帧视频图像It-3等)时,该待处理图像对应的跟踪查询表征数据的确定过程可以包括下文步骤11-步骤12。
步骤11:从第一类图像的目标描述信息中提取第一跟踪目标描述信息。
其中,第一类图像用于提供上文待处理图像中已跟踪到的目标的最近一次历史所处位置;而且本公开不限定该第一类图像,比如,当该待处理图像与该待处理图像对应的多个历史图像属于同一个视频数据时,该待处理图像的时序晚于该第一类图像的时序,而且该第一类图像的时序晚于该多个历史图像中除了该第一类图像以外其他图像(也就是,上文第二类图像)的时序。
可见,在一种可能的实施方式下,当上文待处理图像与该待处理图像对应的多个历史图像属于同一个视频数据时,上文第一类图像是指该多个历史图像中存在的、时序最接近于该待处理图像的时序的历史图像(比如,图2所示的第t-1帧视频图像It-1)。
第一跟踪目标描述信息用于描述从上文第一类图像提取到的、在确定上文待处理图像中一些目标时所需参考的最近一次历史信息,以使该第一跟踪目标描述信息能够表示出在时序最接近于该待处理图像的时序的历史图像中出现的一些已跟踪到的目标。
另外,本公开不限定上文第一跟踪目标描述信息的实施方式,例如,其可以包括至少一个已跟踪到的目标的在上文第一类图像中所呈现的目标内容特征表征数据以及目标位置表征数据,以使该目标内容特征表征数据能够表示出该已跟踪到的目标在时序最接近于该待处理图像的时序的历史图像中所处状态,并使得该目标位置表征数据能够表示出该已跟踪到的目标在时序最接近于该待处理图像的时序的历史图像中所处位置。
此外,本公开不限定上文第一跟踪目标描述信息的确定过程,例如,其具体可以为:在获取到上文第一类图像的目标描述信息之后,先从该目标描述信息中抽取至少一个满足预设筛选条件的Query;再依据这些Query,确定该第一跟踪目标描述信息,以使该第一跟踪目标描述信息包括这些Query。其中,预设筛选条件可以预先设定,例如,该预设筛选条件可以为:该Query被分配了跟踪对象标识(比如,ID),而且该Query的IOU(Intersection over Union)大于预设阈值。该跟踪对象标识用于唯一标识针对某个已跟踪到的目标的跟踪过程。
基于上文步骤11的相关内容可知,对于上文待处理图像(比如,图2所示的第t帧视频图像It)来说,在获取到该待处理图像对应的、距离该待处理图像最近的一个历史图像(比如,图2所示的第t-1帧视频图像It-1)的目标描述信息之后,可以从该目标描述信息抽取一些已跟踪到的目标的目标内容特征表征数据及其目标位置表征数据,确定为第一跟踪目标描述信息(比如,图2所示的Xt中t-1对应的跟踪Query),以使该第一跟踪目标描述信息能够表示出在这些已跟踪到的目标在该历史图像中所处状态。
步骤12:依据上文第一跟踪目标描述信息以及上文第二类图像的目标描述信息,确定上文待处理图像对应的跟踪查询表征数据。
其中,第二类图像用于代表上文待处理图像对应的多个历史图像中除了上文第一类图像以外的其他图像;而且本公开不限定该第二类图像,比如,当该待处理图像与该待处理图像对应的多个历史图像属于同一个视频数据时,该待处理图像的时序晚于该第一类图像的时序,而且该第一类图像的时序晚于该第二类图像的时序。
可见,在一种可能的实施方式下,当上文待处理图像与该待处理图像对应的多个历史图像属于同一个视频数据时,上文第二类图像是指该多个历史图像中存在的时序不是最接近于该待处理图像的时序的历史图像(比如,图2所示的第t-2帧视频图像It-2、第t-3帧视频图像It-3等)。
另外,本公开不限定上文步骤12的实施方式,例如,其具体可以为:先从上文第二类图像的目标描述信息中提取一些已跟踪到的目标的目标内容特征表征数据;再将该目标内容特征表征数据与上文第一跟踪目标描述信息中相应目标的目标内容特征表征数据进行预设处理(比如,整合处理、融合处理
或者拼接处理等),得到上文待处理图像对应的跟踪查询表征数据,以使该跟踪查询表征数据不仅能够表示出这些已跟踪到的目标的在最近一帧历史图像中所处位置,还能够比较全面的表示出这些已跟踪到的目标的在最近多帧历史图像中所呈现的特征。
实际上,为了更好地提高目标确定效果,本公开还提供了上文步骤12的另一种可能的实施方式,其具体可以包括下文步骤121-步骤122。
步骤121:依据上文第一跟踪目标描述信息中的目标位置表征数据以及上文第二类图像的目标描述信息中的目标内容特征表征数据,确定第二跟踪目标描述信息。
其中,第二跟踪目标描述信息用于描述依据上文第二类图像所确定的、在确定上文待处理图像中一些目标时所需参考的跟踪Query。
另外,本公开不限定上文步骤121的实施方式,例如,其具体可以为:将上文第一跟踪目标描述信息中的目标位置表征数据以及上文第二类图像的目标描述信息中相应目标的目标内容特征表征数据进行组合,得到至少一个已跟踪到的目标的组合Query,以使该组合Query能够表示出该已跟踪到的目标在该第二类图像中所呈现的特征、以及该已跟踪到的目标在上文第一类图像中所处位置;再依据这些组合Query,确定为第二跟踪目标描述信息(比如,图2所示的Xt中t-2对应的跟踪Query或者t-3对应的跟踪Query),以使该第二跟踪目标描述信息包括这些组合Query。
基于上文步骤121的相关内容可知,对于上文待处理图像(比如,图2所示的第t帧视频图像It)来说,在获取到该待处理图像对应的、距离该待处理图像不是最近的一个历史图像(比如,图2所示的第t-2帧视频图像It-2、或者第t-3帧视频图像It-3等)的目标描述信息之后,可以先从该目标描述信息中抽取一些已跟踪到的目标的目标内容特征表征数据,以使该目标内容特征表征数据能够表示出该已跟踪到的目标在该历史图像中所呈现的特征;再将该目标内容特征表征数据与上文第一跟踪目标描述信息中相应目标的目标位置表征数据进行组合,得到该已跟踪到的目标的组合Query;最后,依据该组合Query,确定上文第二跟踪目标描述信息,以使该第二跟踪目标描述信息能够表示出依据该历史图像所确定的、在确定上文待处理图像中一些目标时所需参考的跟踪Query。
步骤122:根据上文第二跟踪目标描述信息和上文第一跟踪目标描述信息,确定上文待处理图像对应的跟踪查询表征数据。
需要说明的是,本公开不限定上文步骤122的实施方式,例如,其具体可以为:针对上文第二跟踪目标描述信息与上文第一跟踪目标描述信息进行信息整合处理,得到上文待处理图像对应的跟踪查询表征数据(比如,图2所示的Xt),以使该跟踪查询表征数据包括一些已跟踪到的目标对应的跟踪Query序列(比如,图2所示的序列序列以及序列等)。
还需要说明的是,对于图2中所示的i、i+1、以及i+2分别表示针对一个已跟踪到的目标所分配的跟踪对象标识,i为正整数。
基于上文步骤121至步骤122的相关内容可知,在一些应用场景下,对于上文待处理图像(比如,图2所示的第t帧视频图像It)来说,如果该待处理图像对应的多个历史图像包括第一类图像(比如,图2所示的第t-1帧视频图像It-1)和第二类图像(比如,图2所示的第t-2帧视频图像It-2、第t-3帧视频图像It-3等),则可以利用该第一类图像的目标描述信息中的目标位置表征数据以及该第二类图像的目标描述信息中的目标内容特征表征数据,构建该第二类图像对应的跟踪Query(也就是,上文第二跟踪目标描述信息);再结合该第二类图像对应的跟踪Query以及该第一类图像对应的跟踪Query(也就是,上文第一跟踪目标描述信息),确定上文待处理图像对应的跟踪查询表征数据,以使该跟踪查询表征数据能够更好地表示出该待处理图像中一些已跟踪到的目标在历史阶段所呈现的状态。
基于上文步骤11至步骤12的相关内容可知,在一些应用场景下,对于上文待处理图像(比如,图2所示的第t帧视频图像It)来说,可以利用在历史时间段内针对距离该待处理图像最近的Nmax帧历史图像(比如,图2所示的第t-1帧视频图像It-1、第t-2帧视频图像It-2、第t-3帧视频图像It-3等)所确定的目标内容特征表征数据、以及针对距离该待处理图像最近的一帧历史图像(比如,图2所示的第t-1帧视频图像It-1等)所确定的目标位置表征数据(比如,边界框),构建该待处理图像对应的跟踪查询表征数据(比如,图2所示的Xt),以便后续能够利用该跟踪查询表征数据中所记录的所有跟踪Query协同跟踪该待处理图像中的一些目标,以使该跟踪查询表征数据可以被用于确定该待处理图像中出现的一些已经跟踪到的目标,并使得上文检测查
询表征数据(比如,由图2所示的Dt)被用于检测该待处理图像中新出现的目标,以便后续能够基于该跟踪查询表征数据以及该检测查询表征数据,确定该待处理图像的目标描述信息。
基于上文待处理图像对应的跟踪查询表征数据的相关内容可知,在一种可能的实施方式下,当上文待处理图像中存在至少一个目标时,如果该至少一个目标包括Y个待跟踪目标,则该跟踪查询表征数据包括第1个待跟踪目标在多个历史图像下分别对应的查询表征数据、第2个待跟踪目标在多个历史图像下分别对应的查询表征数据、……(以此类推)、以及第Y个待跟踪目标在多个历史图像下分别对应的查询表征数据。其中,第y个待跟踪目标是指在该待处理图像中存在的、而且在历史阶段已经跟踪到的目标;而且该第y个待跟踪目标在一个历史图像下对应的查询表征数据用于表征该第y个待跟踪目标在该历史图像中所呈现的状态,y为正整数,y≤Y,Y为正整数。
基于上段内容可知,对于上文待处理图像对应的跟踪查询表征数据来说,如果该跟踪查询表征数据用于描述一个或者多个已跟踪到的目标的历史状态,则该跟踪查询表征数据包括每个已跟踪到的目标的至少两个跟踪Query,以便在后续处理过程(比如,解码过程等)中,对于每个已跟踪到的目标来说,该目标对应的各个跟踪query分别作为一个独立的Query完成针对该目标的跟踪预测处理,且各Query在此过程中会充分利用该目标对应的其他Query中所携带的信息(比如,目标内容特征表征数据+目标位置表征数据等)完成针对其自身的迭代更新处理,以实现协同跟踪,以使针对各Query所执行的协同跟踪过程均用于共同跟踪同一个目标,如此能够更好地提高针对该目标预测所得的特征的可靠性,从而使得针对该目标预测所得的特征能够更好地表达该目标在该待处理图像中所处状态,如此有利于提高目标跟踪效果。
上文“待处理图像的目标描述信息”用于描述该待处理图像中的至少一个目标;而且本公开不限定该“待处理图像的目标描述信息”的实施方式,比如,该“待处理图像的目标描述信息”可以包括一些目标的查询表征数据(Query)以及这些目标的类别置信度。其中,该目标的查询表征数据可以包括该目标的目标内容特征表征数据以及该目标的目标位置表征数据(比如,边界框等)。该目标的类别置信度用于表示该目标分别以多大概率属于某些类别(比如,人、动物、汽车等)。
又如,当上文待处理图像对应的检测查询表征数据包括多个检测Query,而且上文待处理图像对应的跟踪查询表征数据包括多个跟踪Query时,该待处理图像的目标描述信息可以包括各个检测Query对应的目标预测结果以及各个跟踪Query对应的目标预测结果。其中,第j个检测Query对应的目标预测结果是指依据该第j个检测Query推导出一个目标在该待处理图像中所处状态;j为正整数,j≤J,J为正整数,J表示该检测Query的个数。第m个跟踪Query对应的目标预测结果是指依据该第m个跟踪Query推导出一个目标在该待处理图像中所处状态;m为正整数,m≤M,M为正整数,M表示该跟踪Query的个数。
另外,本公开不限定上文“待处理图像的目标描述信息”的确定过程(也就是,上文S3的实施方式),例如,其可以借助预先构建的解码器进行实施。其中,该解码器用于针对该解码器的输入数据进行解码处理;而且本公开不限定该解码器,比如,该解码器可以采用图2所示的多目标跟踪模型中的解码模块进行实施。
此外,在一些应用场景下,当上文S2采用“对特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据”这一步骤进行实施时,上文S3具体可以为:依据该图像特征表征数据、该检测查询表征数据以及上文待处理图像对应的跟踪查询表征数据进行解码处理,得到该待处理图像的目标描述信息。需要说明的是,该段中所示的S3的实施方式也可以借助预先构建的解码器进行实施。
还有,为了更好地提高目标确定效果,本公开还提供了上文S3的一种可能的实施方式,其具体可以包括下文步骤21-步骤22。
步骤21:依据上文检测查询表征数据和上文跟踪查询表征数据,确定待处理查询表征数据。
其中,待处理查询表征数据用于描述在针对上文待处理图像进行目标确定处理时所需参考的所有Query(比如,多个检测Query+多个跟踪Query),以使该待处理查询表征数据能够更好地表示出该待处理图像中一些目标所具有的特点。
另外,本公开不限定上文待处理查询表征数据的确定过程,例如,其具体可以为:将上文检测查询表征数据和上文跟踪查询表征数据进行级联处理(也
就是,拼接处理),得到待处理查询表征数据,以使该待处理查询表征数据包括该检测查询表征数据和该跟踪查询表征数据,从而使得该待处理查询表征数据能够更好地表示出该待处理图像中一些目标所具有的特点。
步骤22:依据上文待处理查询表征数据进行解码处理,得到待处理图像的目标描述信息。
需要说明的是,本公开不限定上文步骤22的实施方式,例如,其具体可以为:将上文待处理查询表征数据输入预先构建的解码器(比如,图2所示的多目标跟踪模型中的解码模块),以使该解码器能够依据该待处理查询表征数据进行解码处理,得到并输出上文待处理图像的目标描述信息(比如,由图2所示的最后一个解码层所输出的数据),以使该目标描述信息能够描述出该待处理图像中的一些目标。
又如,在一种可能的实施方式下,当上文S2采用“对特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据”这一步骤进行实施时,该步骤22具体可以为:依据该图像特征表征数据和上文待处理查询表征数据进行解码处理,得到待处理图像的目标描述信息。
需要说明的是,本公开不限定上段中步骤“依据该图像特征表征数据和上文待处理查询表征数据进行解码处理,得到待处理图像的目标描述信息”的实施方式,例如,其具体可以为:将该图像特征表征数据和该待处理查询表征数据输入预先构建的解码器(比如,图2所示的多目标跟踪模型中的解码模块),以使该解码器能够依据该图像特征表征数据和该待处理查询表征数据进行解码处理,得到并输出上文待处理图像的目标描述信息(比如,由图2所示的最后一个解码层所输出的数据),以使该目标描述信息能够描述出该待处理图像中的一些目标。
基于上文步骤21至步骤22的相关内容可知,在一些应用场景下,对于待处理图像来说,在获取到该待处理图像对应的检测查询表征数据以及该待处理图像对应的跟踪查询表征数据之后,可以先将这两种Query进行级联处理;再将级联结果直接送入预先构建的解码器,以使该解码器能够输出该待处理图像的目标描述信息。
基于上段内容可知,在一种可能的实施方式下,当本公开提供的图像处理方法借助预先训练好的目标跟踪模型(比如,图2所示的多目标跟踪模块)
进行实现时,上文S3具体可以为:对于上文待处理图像(比如,图2所示的第t帧视频图像It)来说,在获取到该待处理图像对应的跟踪查询表征数据(比如,图2所示的Xt)、由该目标跟踪模型中的编码模块输出的该待处理图像对应的检测查询表征数据(比如,图2所示的Dt)以及该待处理图像对应的图像特征表征数据(比如,图2所示的图像特征图)之后,可以先将该跟踪查询表征数据与该检测查询表征数据进行级联处理;再由该目标跟踪模型中的解码模块依据该级联结果以及该图像特征表征数据进行解码处理,以使该解码模块能够通过自回归的方式得到并输出该待处理图像的目标描述信息(比如,由图2所示的最后一个解码层所输出的数据),以使该目标描述信息包括一些目标的Query以及这些目标的类别置信度。
基于上文内容可知,在一种可能的实施方式下,上文待处理图像的目标描述信息是利用预先训练好的解码器所确定的。其中,该解码器可以用于依据上文检测查询表征数据以及上文待处理图像对应的跟踪查询表征数据进行解码处理,得到该待处理图像的目标描述信息;或者,该解码器也可以用于依据上文图像特征表征数据、该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据进行解码处理,得到该待处理图像的目标描述信息。
另外,为了更好地提高目标确定效果,本公开还提供了上文解码器的一种可能的实施方式,在该实施方式下,该解码器可以包括至少一个解码层,而且该解码层中的自注意力模块是利用第一注意力掩模进行实现的,该第一注意力掩模用于屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互,以使该解码层达到时序阻隔的效果。需要说明的是,本公开不限定该解码层的个数,比如,其可以为6个。
需要说明的是,本公开不限定上段中“同一个目标的在不同历史图像下对应的查询表征数据”的实施方式,例如,当上段所示的解码器是指第h个解码器时,该“同一个目标的在不同历史图像下对应的查询表征数据”是指在该第h个解码器中的自注意力模块的输入数据中存在的对应于同一个目标的不同跟踪Query;h为正整数,h≤解码器的个数(比如,6)。
基于上段内容可知,在一种可能的实施方式下,上文解码器可以是由多个(比如,6个)时序阻隔的解码层(temporal blocking decoder)堆叠而成的,而且每个时序阻隔的解码层中的自注意力模块均是利用第一注意力掩模进行实
现的,以使各个自注意力模块均能够借助该第一注意力掩模实现屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互,以避免同一个目标的不同跟踪Query之间相互抑制。
上文第一注意力掩模是指针对上文解码层中的自注意力模块所配置的注意力掩模,以使该自注意力模块能够借助该注意力掩模实现屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互;而且该第一注意力掩模被配置为屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互。另外,本公开不限定该第一注意力掩模的实施方式,例如,图2所示的解码层中的自注意力模块中的第一注意力掩模可以采用下文表1所示的注意力掩模进行实施。
表1解码层中的自注意力模块中的注意力掩模的一种可能的实施方式需要说明的是,对于上文表1来说,表示以i作为跟踪对象标识的目标在图像数据It-1下所对应的跟踪Query;表示以i+1作为跟踪对象标识的目标在图像数据It-1下所对应的跟踪Query;表示以i+1作为跟踪对象标识的目标在图像数据It-2下所对应的跟踪Query;表示以i+1作为跟踪对象标识的目标在图像数据It-3下所对应的跟踪Query;……(以此类推)。
基于上述解码器的相关内容可知,在一种可能的实施方式下,上文解码器可以是由多个时序阻隔的解码层堆叠而成的,而且可以通过将每个时序阻隔的解码层内自注意力模块中的注意力掩模被配置为屏蔽属于同一个目标的不同跟踪Query之间做交互的方式,以避免同一个目标的不同跟踪Query之间
相互抑制。例如,对于该解码层中的自注意力模块来说,该自注意力模块可以不仅用于进行不同目标的Query之间的交互处理,还用于进行同一个Query之间的交互处理;但是该自注意力模块会屏蔽同一个目标下的不同跟踪Query之间的交互处理。
可见,由于在传统的自注意力模块中可以让该自注意力模块各项输入信息进行平等的交互,故为了有效地避免同一个目标的多个跟踪query彼此抑制,可以向该自注意力模块中添加掩模(比如,上文第一注意力掩模或者上文表1所示的掩模),以得到上文解码层中的自注意力模块,以使该解码层中的自注意力模块就是添加了掩模的自注意力模块。其中,因在添加了掩模的自注意力模块中能够屏蔽属于同一个目标的不同跟踪Query之间做交互,如此能够有效地避免在添加了掩模的自注意力模块中出现同一个目标的多个跟踪query彼此抑制这一现象,从而有利于提高目标跟踪效果。
实际上,为了更好地提高解码效果,本公开还提供了上文解码器的另一种可能的实施方式,在该实施方式中,当该解码器包括至少一个解码层,而且该至少一个解码层包括第一解码层和第二解码层时,该解码器可以还包括特征融合模块(比如,图2所示的特征融合模块)。其中,该第一解码层用于代表该至少一个解码层中存在的任意两个相邻解码层中位置靠前的一个解码层。该第二解码层用于代表该至少一个解码层中存在的任意两个相邻解码层中位置靠后的一个解码层。该特征融合模块位于第一解码层与第二解码层之间。为了便于理解,下面结合示例进行说明。
作为示例,当上文解码器包括依次排列的6个解码层时,如果第一解码层为该解码器中的第1个解码层,则第二解码层为该解码器中的第2个解码层,而且在该第1个解码层与该第2个解码层之间存在一个特征融合模块;如果第一解码层为该解码器中的第2个解码层,则第二解码层为该解码器中的第3个解码层,而且在该第2个解码层与该第3个解码层之间存在一个特征融合模块;……(以此类推)。
基于上述两段内容可知,对于上文解码器中存在的任意一个三元组(第一解码层,特征融合模块,第二解码层)来说,该三元组中不同元素之间的关联关系具体可以为:该特征融合模块用于针对由该第一解码层输出的跟踪查询处理数据进行融合处理,得到跟踪查询融合数据;而且该第二解码层用于针对
由该特征融合模块输出的跟踪查询融合数据以及由该第一解码层输出的检测查询处理数据进行处理。其中,该跟踪查询处理数据是依据所述第一解码层以及上文跟踪查询表征数据所确定的;该跟踪查询处理数据包括由该第一解码层针对各待跟踪目标分别输出的处理后的查询表征数据;该跟踪查询融合数据包括由该特征融合模块针对各处理后的查询表征数据分别输出的融合后的查询表征数据;该检测查询处理数据是依据该第一解码层以及上文检测查询表征数据所确定的。
上文跟踪查询处理数据是指由上文第一解码层针对已跟踪到的目标所输出的Query;而且该跟踪查询处理数据是依据该第一解码层以及上文待处理图像对应的跟踪查询表征数据所确定的。例如,如果该第一解码层为上文解码器中的第1个解码层,则由该第一解码层所输出的跟踪查询处理数据是指由该第1个解码层针对上文跟踪查询表征数据进行处理所得到的Query,以使由该第一解码层所输出的跟踪查询处理数据包括由该第1个解码层针对各个已跟踪到的目标分别输出的处理后的查询表征数据;如果该第一解码层为上文解码器中的第2个解码层,则由该第一解码层所输出的跟踪查询处理数据是指由该第2个解码层针对该第2个解码层的输入数据中存在的用于描述已跟踪到的目标的Query进行处理所得到的Query,以使由该第一解码层所输出的跟踪查询处理数据包括由该第2个解码层针对各个已跟踪到的目标分别输出的处理后的查询表征数据;如果该第一解码层为上文解码器中的第3个解码层,则由该第一解码层所输出的跟踪查询处理数据是指由该第3个解码层针对该第3个解码层的输入数据中存在的用于描述已跟踪到的目标的Query进行处理所得到的Query,以使由该第一解码层所输出的跟踪查询处理数据包括由该第3个解码层针对各个已跟踪到的目标分别输出的处理后的查询表征数据;……(以此类推)。其中,该处理后的查询表征数据是指由该第一解码层针对一个已跟踪到的目标迭代所得的Query。
上文检测查询处理数据是指由上文第一解码层针对待检测的目标所输出的Query;而且该检测查询处理数据是依据该第一解码层以及上文检测查询表征数据所确定的。例如,如果该第一解码层为上文解码器中的第1个解码层,则由该第一解码层所输出的检测查询处理数据是指由该第1个解码层针对上文检测查询表征数据进行处理所得到的Query;如果该第一解码层为上文解码
器中的第2个解码层,则由该第一解码层所输出的检测查询处理数据是指由该第2个解码层针对该第2个解码层的输入数据中存在的用于描述待检测到的目标的Query进行处理所得到的Query;如果该第一解码层为上文解码器中的第3个解码层,则由该第一解码层所输出的检测查询处理数据是指由该第3个解码层针对该第3个解码层的输入数据中存在的用于描述待检测到的目标的Query进行处理所得到的Query;……(以此类推)。
上文跟踪查询融合数据是指由上文特征融合模块针对由上文第一解码层所输出的与各个已跟踪到的目标相关联的Query进行融合所得到的融合结果,以使该跟踪查询融合数据包括由该特征融合模块针对各已跟踪到的目标分别输出的融合后的查询表征数据,从而使得该跟踪查询融合数据能够更好地表示出各个已跟踪到的目标所呈现的特点。其中,该融合后的查询表征数据是指将该已跟踪到的目标以及与该已跟踪到的目标相关联的Query进行融合所得到的融合结果。比如,对于已跟踪到的目标A来说,如果该目标A对应于3个跟踪查询处理数据QueryA-1,QueryA-2,QueryA-3,则在该QueryA-1对应融合过程中,需要借助融合模块将该QueryA-1与QueryA-2,QueryA-3所携带的信息进行交互,得到该QueryA-1对应的融合后的查询表征数据QueryA-1’;在该QueryA-2对应融合过程中,需要借助融合模块将该QueryA-2与QueryA-1,QueryA-
3所携带的信息进行交互,得到该QueryA-2对应的融合后的查询表征数据QueryA-2’;在该QueryA-3对应融合过程中,需要借助融合模块将该QueryA-3与QueryA-1,QueryA-2所携带的信息进行交互,得到该QueryA-3对应的融合后的查询表征数据QueryA-3’。
另外,本公开不限定上文特征融合模块的实施方式,例如,在一些应用场景下,该特征融合模块可以包括多个融合子模块,不同融合子模块分别用于针对不同已跟踪到的目标所关联的跟踪Query进行融合处理。可见,该融合子模块的个数可以依据已跟踪到的目标的个数进行确定。又如,在另一种可能的实施方式下,该特征融合模块可以用于执行以下步骤:将第1个已跟踪到的目标所关联的跟踪Query进行融合处理,得到并输出该第1个已跟踪到的目标对应的融合后的查询表征数据,以使该第1个已跟踪到的目标对应的融合后的查询表征数据包括该第1个已跟踪到的目标所关联的各个跟踪Query分别对应的融合结果,从而使得该第1个已跟踪到的目标对应的融合后的查询
表征数据中的Query个数与该第1个已跟踪到的目标所关联的跟踪Query个数保持一致;再将第2个已跟踪到的目标所关联的跟踪Query进行融合处理,得到并输出该第2个已跟踪到的目标对应的融合后的查询表征数据,以使该第2个已跟踪到的目标对应的融合后的查询表征数据包括该第2个已跟踪到的目标所关联的各个跟踪Query分别对应的融合结果,从而使得该第2个已跟踪到的目标对应的融合后的查询表征数据中的Query个数与该第2个已跟踪到的目标所关联的跟踪Query个数保持一致;……(以此类推)。
实际上,为了更好地提高融合效率,本公开还提供了上文特征融合模块的一种可能的实施方式,在该实施方式下,该特征融合模块包括信息移除分支网络和信息添加分支网络;而且利用该特征融合模块确定上文跟踪查询融合数据的过程,具体可以包括下文步骤31-步骤33。
步骤31:利用上文信息移除分支网络对上文跟踪查询处理数据进行处理,得到保留信息特征。
其中,信息移除分支网络用于针对上文特征融合模块的输入数据进行信息移除处理;而且本公开不限定该信息移除分支网络,比如,该信息移除分支网络可以包括两个自注意力层、一个全连接层以及一个sigmoid层。
保留信息特征用于表示通过针对上文特征融合模块的输入数据进行信息移除处理之后所保留下来信息。
另外,本公开不限定上文保留信息特征的确定过程,例如,其具体可以包括下文步骤311-步骤315。
步骤311:对上文跟踪查询处理数据进行时序信息提取处理,得到第一时序信息,以使该第一时序信息能够表示出该跟踪查询处理数据中所携带的时序信息。
需要说明的是,本公开不限定步骤311的实施方式,例如,该步骤311可以借助一个自注意力层进行实施。
基于上文步骤311的相关内容可知,对于上文信息移除分支网络来(比如,图3所示的信息移除分支网络)说,在该信息移除分支网络接收到由上文第一解码层所输出的跟踪查询处理数据(比如,图3所示的)之后,可以由该信息移除分支网络中的第一个自注意力层针对该跟踪查询处理数据进行时序信息提取处理,得到第一时序信息(比如,图3所示的时序信息1),
以使该第一时序信息能够表示出该跟踪查询处理数据中所携带的时序信息。
需要说明的是,对于图3所示的来说,该可以表示由图2中第1个解码层针对图2所示的Xt进行处理所得到的结果;而且Ft用于描述由该第1个解码层针对一些已跟踪到的目标预测所得的特征信息;表示由该第1个解码层针对一些已跟踪到的目标预测所得的位置信息。
步骤312:对上文第一时序信息与上文跟踪查询处理数据进行自注意力处理,得到自注意力处理结果。
需要说明的是,本公开不限定步骤312的实施方式,例如,该步骤312可以借助一个自注意力层进行实施。
基于上文步骤312的相关内容可知,对于上文信息移除分支网络(比如,图3所示的信息移除分支网络)来说,在由该信息移除分支网络中的第一个自注意力层输出上文第一时序信息(比如,图3所示的时序信息1)之后,可以由该信息移除分支网络中的第二个自注意力层依据该第一时序信息与上文跟踪查询处理数据进行自注意力处理,得到自注意力处理结果,以使该自注意力处理结果能够表示出需要将哪些信息进行剔除处理。
步骤313:对上文自注意力处理结果进行全连接处理,得到全连接处理结果。
需要说明的是,本公开不限定步骤313的实施方式,例如,该步骤313可以借助一个全连接层进行实施。
基于上文步骤313的相关内容可知,对于上文信息移除分支网络(比如,图3所示的信息移除分支网络)来说,在由该信息移除分支网络中的第二个自注意力层输出上文自注意力处理结果之后,可以由该信息移除分支网络中的全连接层针对该自注意力处理结果进行全连接处理,得到全连接处理结果。
步骤314:依据上文全连接处理结果,确定待移除信息表征数据,以使该待移除信息表征数据能够表示出上文跟踪查询处理数据中哪些信息需要被移除。
需要说明的是,本公开不限定步骤314的实施方式,例如,该步骤314可以借助一个sigmoid层进行实施。
基于上文步骤314的相关内容可知,对于上文信息移除分支网络(比如,图3所示的信息移除分支网络)来说,在由该信息移除分支网络中的全连接
层输出上文全连接处理结果之后,可以由该信息移除分支网络中的sigmoid层针对该全连接处理结果进行处理,得到待移除信息表征数据(比如,图3所示的Z),以使该待移除信息表征数据能够表示出上文跟踪查询处理数据中哪些信息需要被移除。
步骤315:依据上文待移除信息表征数据,对上文跟踪查询处理数据进行信息移除处理,得到保留信息特征。
需要说明的是,本公开不限定步骤315的实施方式,例如,该步骤315具体可以为:先依据上文待移除信息表征数据,确定待保留信息表征数据(比如,图3所示的1-Z);再将上文跟踪查询处理数据中的目标内容特征表征数据与该待保留信息表征数据相乘,得到保留信息特征(例如,图3所示的Ft×(1-Z))。
基于上文步骤311至步骤315的相关内容可知,在一种可能的实施方式下,上文信息移除分支网络可以包括两个自注意力层、一个全连接层以及一个sigmoid层;其中,第1个自注意力层用于搜集时序信息,第2个自注意力层用于融合多个Query的特征,该全连接层用于针对该第2个自注意力层的输出数据进行全连接处理;该sigmoid层用于针对该全连接层的输出数据进行处理,以得到门控值(也就是,上文待移除信息表征数据),以便后续能够基于该门控值(比如,图3所示的Z),针对该信息移除分支网络的输入数据进行信息移除处理。
基于上文步骤31的相关内容可知,对于上文特征融合模块来说,如果该特征融合模块包括信息移除分支网络,则在该特征融合模块接收到由上文第一解码层所输出的跟踪查询处理数据(比如,图3所示的)之后,由该信息移除分支网络针对该跟踪查询处理数据进行信息移除处理,得到保留信息特征(例如,图3所示的Ft×(1-Z)),以便后续基于该保留信息特征,确定出针对该跟踪查询处理数据的Query融合结果。
步骤32:利用上文信息添加分支网络对上文跟踪查询处理数据进行处理,得到待添加信息特征。
其中,信息添加分支网络用于从上文特征融合模块的输入数据中提取出所需添加的信息;而且本公开不限定该信息添加分支网络,比如,该信息添加分支网络可以包括两个自注意力层。
待添加信息特征是指在进行信息添加处理时所需依据的信息。
另外,本公开不限定上文待添加信息特征的确定过程,例如,其具体可以包括下文步骤321-步骤322。
步骤321:对上文跟踪查询处理数据进行时序信息提取处理,得到第二时序信息,以使该第二时序信息能够表示出该跟踪查询处理数据中所携带的时序信息。
需要说明的是,本公开不限定步骤321的实施方式,例如,该步骤321可以借助一个自注意力层进行实施。
基于上文步骤321的相关内容可知,对于上文信息添加分支网络(比如,图3所示的信息添加分支网络)来说,在该信息添加分支网络接收到由上文第一解码层所输出的跟踪查询处理数据(比如,图3所示的)之后,可以由该信息添加分支网络中的第一个自注意力层针对该跟踪查询处理数据进行时序信息提取处理,得到第二时序信息(比如,图3所示的时序信息2),以使该第二时序信息能够表示出该跟踪查询处理数据中所携带的时序信息。
步骤322:对上文第二时序信息与上文跟踪查询处理数据进行自注意力处理,得到待添加信息特征。
需要说明的是,本公开不限定步骤322的实施方式,例如,该步骤322可以借助一个自注意力层进行实施。
基于上文步骤322的相关内容可知,对于上文信息移除分支网络(比如,图3所示的信息添加分支网络)来说,在由该信息添加分支网络中的第一个自注意力层输出上文第二时序信息(比如,图3所示的时序信息2)之后,可以由该信息移除分支网络中的第二个自注意力层依据该第二时序信息与上文跟踪查询处理数据进行自注意力处理,得到待添加信息特征(比如,图3所示的),以使该待添加信息特征能够表示出在进行信息添加处理时所需依据的信息。
基于上文步骤321至步骤322的相关内容可知,在一种可能的实施方式下,上文信息添加分支网络可以包括两个自注意力层;其中,第1个自注意力层用于搜集时序信息,第2个自注意力层用于融合多个Query的特征。
基于上文步骤32的相关内容可知,对于上文特征融合模块来说,如果该特征融合模块包括信息添加分支网络,则在该特征融合模块接收到由上文第
一解码层所输出的跟踪查询处理数据(比如,图3所示的)之后,由该信息添加分支网络针对该跟踪查询处理数据进行信息提取处理,得到待添加信息特征(比如,图3所示的),以便后续基于该待添加信息特征,确定出针对该跟踪查询处理数据的Query融合结果。
步骤33:根据上文保留信息特征和上文待添加信息特征,确定上文跟踪查询融合数据。
需要说明的是,本公开不限定上文步骤33的实施方式,例如,其可以采用下文公式(2)进行实施。
式中,表示上文特征融合模块的输出结果(比如,上文跟踪查询融合数据);LN()表示层级标准化处理;Ft表示上文特征融合模块的输入数据中的目标内容特征表征数据;Z表示由上文特征融合模块中信息移除分支网络内sigmoid层所输出的门控值;Ft×(1-Z)表示由上文特征融合模块中信息移除分支网络所输出的数据(比如,上文保留信息特征);表示由上文特征融合模块中信息添加分支网络内第二个自注意力层所输出的数据(比如,上文待添加信息特征)。
基于上文步骤31至步骤33的相关内容可知,对于上文特征融合模块来说,该特征融合模块可以包括两个分支网络,即信息移除分支网络和信息添加分支网络。每个分支网络都包含两个叠加的自注意力层,第一个自注意力层用于搜集时序信息,第二个自注意力层用于根据该时序信息融合多个Query的特征。另外,对于信息移除分支网络来说,该信息移除分支网络中的第二个自注意力层的输出数据会进一步经过一个全连接层和一个sigmoid层得到门控值,该门控值与该信息添加分支网络的输出数据共同决定了融合后的Query信息(如上文公式(2)所示)。可见,因该特征融合模块是通过两个分支来模拟信息交互过程的,以使对于任一目标的某个跟踪Query来说,该特征融合模块能够更高效地将该目标除了该某个跟踪Query以外的其他跟踪Query所携带的有效信息整合至该某个跟踪Query内,以完成针对该某个跟踪Query的自我迭代更新,如此有利于提高同一个目标的不同跟踪Query之间的交互效率及效果,从而有利于提高目标跟踪效果。
另外,对于上文特征融合模块来说,为了确保该特征融合模块可以同时执
行针对所有已跟踪到的目标的Query融合处理,本公开还提供了该特征融合模块的一种可能的实施方式,在该实施方式中,该特征融合模块中的自注意力层是利用第二注意力掩模进行实现的,该第二注意力掩模用于屏蔽属于不同目标的查询表征数据之间的交互,如此能够确保在该特征融合模块中只进行同一个目标的所有跟踪Query之间的交互。
其中,第二注意力掩模是指针对上文特征融合模块中的自注意力层所配置的注意力掩模;而且该第二注意力掩模被配置为屏蔽属于不同目标的查询表征数据之间的交互。另外,本公开不限定该第二注意力掩模的实施方式,例如,图2所示的特征融合模块中的自注意力层中的第二注意力掩模可以采用下文表2所示的注意力掩模进行实施。
需要说明的是,对于上文特征融合模块中的每个分支网络(比如,信息移除分支网络或者信息添加分支网络)来说,因该分支网络包括两个自注意力层,故在一种可能的实施方式下,该分支网络中每个自注意力层所使用的第二注意力掩模分别可以采用类似于下文表2所示的注意力掩模进行实施。
还需要说明的是,对于下文表2来说,该表2所涉及的
以及的相关内容请参见上文。
表2特征融合模块中的自注意力层中的注意力掩模的一种可能的实施方式
基于上文特征融合模块的相关内容可知,该特征融合模块可以用于同时执行针对所有已跟踪到的目标的Query融合处理;而且可以通过将该特征融合模块内自注意力层中的注意力掩模配置为避免不同目标的Query之间进行交互的方式,来确保在该特征融合模块内只进行同一个目标的Query之间交互。
另外,对于上文特征融合模块中的每个分支网络(比如,信息移除分支网络或者信息添加分支网络)来说,因该分支网络包括两个自注意力层,故该分
支网络中第g个自注意力层所使用的第二注意力掩模用于屏蔽该第g个自注意力层的输入数据中出现的不同目标的查询表征数据之间的交互,g为正整数,g≤2。
基于上文解码器的相关内容可知,在一种可能的实施方式下,该解码器可以由多个时序阻隔的解码层堆叠而成,而且在任意两个相邻解码层之间会插入一个特征融合模块,如此能够确保该解码器具有更好地解码效果。
基于上文S1至S3的相关内容可知,对于本公开实施例提供的图像处理方法来说,先对待处理图像进行特征提取处理,得到特征提取结果,以使该特征提取结果用于表征该待处理图像所携带的图像信息;再对该特征提取结果进行编码处理,至少得到检测查询表征数据(比如,多个检测Query);最后,至少依据该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据(比如,基于多个历史图像所确定的跟踪Query)进行解码处理,得到该待处理图像的目标描述信息(比如,目标内容特征表征数据、目标位置表征数据、以及类别置信度等),以使该目标描述信息能够表示出该待处理图像中一些目标。其中,因该跟踪查询表征数据是依据该待处理图像对应的多个历史图像的目标描述信息所确定的,以使该跟踪查询表征数据能够更好地表示出该待处理图像中一些目标的历史状态,从而使得基于该跟踪查询表征数据所确定的目标描述信息能够更好地表示出该待处理图像中一些目标,如此有利于提高目标跟踪效果。
实际上,在一些应用场景下,在获取到一个图像数据的目标描述信息之后,可以从该目标描述信息中挑选出一些高质量的目标预测结果作为该图像数据的目标确定结果。基于此,本公开还提供了上文图像处理方法的一种可能的实施方式,在该实施方式下,当上文待处理图像的目标描述信息包括至少一个目标预测结果(比如,多个检测Query对应的目标预测结果+多个跟踪Query对应的目标预测结果)时,该图像处理方法除了包括上文S1-S3以外,可以还包括下文步骤41。其中,该步骤41的执行时间晚于上文S3的执行时间。
步骤41:在得到待处理图像的目标描述信息之后,从该目标描述信息中提取满足预设参考条件的目标预测结果,得到该待处理图像的目标确定结果。
其中,预设参考条件是指在从上文待处理图像的目标描述信息中筛选高质量的目标预测结果时所需依据的条件;而且本公开不限定该预设参考条件,
例如,其可以为:属于检测Query对应的目标预测结果或者属于最近一个历史图像的跟踪Query对应的目标预测结果,而且类别置信度达到预设阈值。
待处理图像的目标确定结果用于描述该待处理图像中一些预测比较准确的目标。
另外,本公开不限定上文待处理图像的目标确定结果的确定过程,例如,其具体可以为:在得到待处理图像的目标描述信息之后,从该目标描述信息中提取出检测Query对应的目标预测结果和最近一个历史图像的跟踪Query对应的目标预测结果,均作为待处理预测结果;再判断每个待处理预测结果中的类别置信度是否达到预设阈值,若达到,则可以确定该待处理预测结果所描述的目标比较准确,故可以依据该待处理预测结果,确定该待处理图像的目标确定结果,以使该目标确定结果包括该待处理预测结果。
基于上文步骤41的相关内容可知,在一些应用场景下,在确定出一个图像数据的目标描述信息之后,从该目标描述信息中提取一些高质量的目标预测结果,得到该图像数据的目标确定结果,以使该目标确定结果能够更好地表示出该图像数据中存在哪些目标,如此有利于提高目标确定效果。
实际上,当本公开实施例提供的图像处理方法应用于目标跟踪领域时,在获取到某一帧图像数据的目标描述信息之后,可以进一步利用该目标描述信息,确定下一帧图像对应的跟踪查询表征数据,如此有利于提高针对下一帧图像的目标确定效果。基于此,本公开还提供了上文图像处理方法的一种可能的实施方式,在该实施方式下,当上文待处理图像是指从图像序列(比如,视频数据)中抽取的一个图像,而且该待处理图像的目标描述信息包括至少一个目标预测结果(比如,多个检测Query对应的目标预测结果+多个跟踪Query对应的目标预测结果)时,该图像处理方法可以至少包括上文S1-S3以及下文步骤51-步骤52。其中,该步骤51的执行时间晚于该S3的执行时间。
步骤51:在得到待处理图像的目标描述信息之后,从该目标描述信息中提取满足预设参考条件的目标预测结果。
需要说明的是,步骤51的相关内容类似于上文步骤41中部分内容,为了简要起见,在此不再赘述。
步骤52:利用上文图像序列中的下一帧图像更新待处理图像,利用上文满足预设参考条件的目标预测结果中的查询表征数据,更新该待处理图像对
应的跟踪查询表征数据,并返回继续执行上文步骤S1及其后续步骤,直至该图像序列中所有图像数据均被遍历。
上文步骤52中的图像序列是指需要进行目标跟踪处理的序列;而且本公开不限定该图像序列的实施方式,例如,其可以采用视频数据进行实施。
上文步骤52中的下一帧图像是指在上文图像序列中存在的、与上文步骤51所涉及的待处理图像位置相邻、而且位置比该步骤51所涉及的待处理图像位置靠后的图像数据。
另外,本公开不限定上文步骤52的实施方式,例如,当上文图像序列为图2所涉及的视频数据,而且上文S1-S3以及步骤51均用于针对图2所示的图像数据It进行处理时,该步骤52具体可以为:在从该图像数据It的目标描述信息中提取满足预设参考条件的目标预测结果(比如,由图2所示的多目标跟踪模型所输出的数据)之后,可以将图像数据It+1作为新的待处理图像,并利用该目标预测结果中的查询表征数据(Query)更新该待处理图像对应的跟踪查询表征数据,以便后续能够基于更新后的待处理图像及其对应的跟踪查询表征数据,执行针对该图像数据It+1的目标确定过程。
基于上文步骤51至步骤52的相关内容可知,在一些应用场景下,在获取到当前帧视频图像的目标描述信息之后,可以从该目标描述信息中提取满足预设参考条件的目标预测结果;再利用该目标预测结果中的Query,确定下一帧视频图像对应的跟踪查询表征数据,以使该跟踪查询表征数据包括该Query,如此有利于提高目标跟踪效果。
需要说明的是,在一些应用场景下,对于一个跟踪Query对应的目标预测结果来说,如果该目标预测结果中的类别置信度较低,则可以判定该目标预测结果中的Query失效。另外,如果失效持续帧数超过预定的数量时视为此目标消失并不再继续跟踪,相应的跟踪Query会被删除。
还需要说明的是,当上文待处理图像是指从图像序列(比如,视频数据)中抽取的一个图像时,如果该待处理图像为该图像序列中的首帧图像(也就是,第1帧图像),则因不存在该待处理图像对应的多个历史图像,以使该待处理图像对应的跟踪查询表征数据也是不存在的,故针对该待处理图像的图像处理过程具体可以为:先对该待处理图像进行特征提取处理,得到特征提取结果;再对该特征提取结果进行编码处理,得到检测查询表征数据;然后,依
据该检测查询表征数据进行解码处理,得到该待处理图像的目标描述信息;但是,如果该待处理图像为该图像序列中的非首帧图像(比如,第2帧图像、第3帧图像等),则可以利用上文S1-S3实现针对该待处理图像的图像处理过程。
基于上文图像处理方法的相关内容可知,该图像处理方法可以借助预先训练好的目标跟踪模型进行实施。也就是,在一种可能的实施方式下,当该目标跟踪模型包括特征提取模块、编码模块以及解码模块时,该图像处理方法具体可以为:先由该特征提取模块针对上文待处理图像(比如,图2所示的图像数据It)进行特征提取处理,得到特征提取结果;再由该编码模块针对该特征提取结果进行编码处理,得到图像特征表征数据(比如,图2所示的图像特征图)以及检测查询表征数据(比如,图2所示的Dt);最后,由该解码模块依据该图像特征表征数据、该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据(例如,图2所示的Xt)进行解码处理,得到该待处理图像的目标描述信息(比如,由图2所示的最后一个解码层所输出的数据)。
实际上,为了更好地提高上文目标跟踪模型的性能,本公开还提供了该目标跟踪模型的训练过程的一种可能的实施方式,其具体可以包括下文步骤61-步骤62。
步骤61:利用至少一个第二图像数据以及各第二图像数据的目标检测标签,对初始模型进行训练,得到待优化模型。
其中,第二图像数据是指在针对目标跟踪模型的第一训练阶段中所需使用的图像数据;而且本公开不限定该第二图像数据的获取方式,例如,该第二图像数据可以是指从任意一个目标检测训练数据集中所抽取的图像数据。
第二图像数据的目标检测标签用于描述该第二图像数据中每个目标实际所处位置;而且本公开不限定该第二图像数据的目标检测标签的获取方式,比如,可以借助人工标注的方式进行实施。又如,在一些应用场景下,如果该第二图像数据是指从某个目标检测训练数据集中所抽取的图像数据,则该第二图像数据的目标检测标签可以为该目标检测训练数据集中存在的、与该第二图像数据具有对应关系的标签信息。
初始模型用于代表在第一训练阶段内所需训练的目标跟踪模型;而且本公开不限定该初始模型,比如,该初始模型可以包括特征提取模块、编码模块以及解码模型;而且该特征提取模块是指基于CoCo数据集预训练的ResNet50,
该编码模块是指DINO检测算法中的Transformer Encoder模块。
待优化模型是指在针对上文初始模型进行目标检测方面训练所得到的模型,以使该待优化模型具有较好的目标检测功能。可见,该待优化模型可以代表经过第一训练阶段所得到的目标跟踪模型。
步骤62:利用至少一个图像序列以及各图像序列的目标跟踪标签,对上文待优化模型中的解码器进行训练,得到目标跟踪模型。
上文步骤62中的图像序列是指在针对目标跟踪模型的第二训练阶段中所需使用的图像数据序列;而且本公开不限定该图像序列,比如,其可以是指视频数据。另外,本公开不限定该图像序列的获取方式,例如,该图像序列可以是指从任意一个目标跟踪训练数据集中所抽取的视频数据。
图像序列的目标跟踪标签用于描述该图像序列中每个图像数据内各个目标实际所处状态;而且本公开不限定该图像序列的目标跟踪标签的获取方式,比如,可以借助人工标注的方式进行实施。又如,在一些应用场景下,如果该图像序列是指从某个目标跟踪训练数据集中所抽取的视频数据,则该图像序列的目标跟踪标签可以为该目标跟踪训练数据集中存在的、与该视频数据具有对应关系的标签信息。
基于上文步骤61至步骤62的相关内容可知,在一些应用场景下,可以采用两阶段训练方式完成针对上文目标跟踪模型的训练过程,而且该两阶段训练方式具体可以为:先将基于图片的目标检测数据视为包括帧长为1的视频数据的多目标跟踪训练数据,并利用该多目标跟踪训练数据,针对该目标跟踪模型进行第一训练阶段;再将训练好的模型中的特征提取模块以及编码模块进行固定,并利用包括多帧视频数据的多目标跟踪训练数据,针对该模型中的解码模块进行第二阶段的模型训练,以得到最终训练好的目标跟踪模型,以使该目标跟踪模型具有较好的性能。
另外,为了更好地提高上文目标跟踪模型的训练效果,可以在针对上文目标跟踪模型的训练过程中采用一些数据增强手段(比如,Mosaic和mixup等数据增强手段)。基于此可知,该目标跟踪模型的训练过程具体可以为:在获取到至少一个第二图像数据之后,先针对这些第二图像数据进行数据增强处理,得到增强后的图像数据;再利用增强后的图像数据以及这些增强后的图像数据的目标检测标签,对初始模型进行训练,得到待优化模型,以便在获取到
针对上文至少一个图像序列进行数据增强所得到的增强后的图像序列之后,利用该增强后的图像序列以及该增强后的图像序列的目标跟踪标签,对上文待优化模型中的解码器进行训练,得到目标跟踪模型,如此有利于提高该目标跟踪模型的训练效果。
此外,为了更好地提高上文目标跟踪模型的训练效果,可以采用多尺度训练方式进行模型训练。基于此可知,对于上文目标跟踪模型来说,在针对该目标跟踪模型的训练过程中可以随机地调整被训练模型中的尺度参数,以确保最终训练好的目标跟踪模型可以适用于处理具有任一种尺度的图像数据,如此有利于提高该目标跟踪模型的普适性。
还有,为了更好地提高模型训练效果,本公开还提供了一种上文目标跟踪模型的训练损失的确定过程,为了便于理解,下面结合示例进行说明。
作为示例,当上文目标跟踪模型的训练过程中使用了某个图像序列,而且该图像序列包括第一图像数据以及该第一图像数据对应的多个历史图像时,该目标跟踪模型的训练损失的确定过程,具体可以包括下文步骤71-步骤75。
步骤71:依据第一图像数据、该第一图像数据对应的多个历史图像的目标描述信息、以及上文目标跟踪模型,确定该第一图像数据的目标描述信息。
其中,第一图像数据用于代表在一轮训练过程中所使用的图像数据。例如,该第一图像数据可以是图2所示的图像数据It。
第一图像数据对应的多个历史图像是指在利用该第一图像数据进行模型训练时所需参考的图像数据。例如,当该第一图像数据是图2所示的图像数据It时,该第一图像数据对应的多个历史图像可以包括图2所示的图像数据It-1、图像数据It-2、图像数据It-3等。
第一图像数据的目标描述信息用于描述该第一图像数据的一些目标;而且该“第一图像数据的目标描述信息”的确定过程类似于上文“待处理图像的目标描述信息”的确定过程,为了简要起见,在此不再赘述。
基于上文步骤71的相关内容可知,对于一轮训练过程来说,在获取到第一图像数据以及该第一图像数据对应的多个历史图像的目标描述信息之后,可以基于这些信息以及需要训练的目标跟踪模型,确定出该第一图像数据的目标描述信息,以便后续能够基于该目标描述信息,分析该目标跟踪模型所具有的性能。
步骤72:从第一图像数据的目标描述信息中提取第一组信息和第二组信息。
其中,第一组信息是指上文第一图像数据的目标描述信息中存在的、满足一定条件的目标预测结果;而且该条件可以预先设定,例如,该条件具体可以为:属于检测Query对应的目标预测结果或者属于最近一个历史图像的跟踪Query对应的目标预测结果(比如,图2所示的由长方形虚线框所框定的Query)。
第二组信息是指上文第一图像数据的目标描述信息中除了第一组信息以外的其他目标预测结果(比如,图2所示的由正方形虚线框所框定的Query)。
另外,本公开不限定上文第一组信息和第二组信息之间的关联关系,例如,该第一组信息对应的时序晚于该第二组信息对应的时序。其中,该第一组信息对应的时序是指在预测该第一组信息时所需依据的图像数据的时序。比如,如果该第一组信息包括图2所示的由长方形虚线框所框定的Query,则该第一组信息对应的时序包括图2所示的图像数据It的时序以及图2所示的图像数据It-1的时序。该第二组信息对应的时序是指在预测该第二组信息时所需依据的图像数据的时序。比如,如果该第二组信息包括图2所示的由正方形虚线框所框定的Query,则该第二组信息对应的时序包括图2所示的图像数据It-2的时序、图2所示的图像数据It-3的时序等。
基于上文步骤72的线管内容可知,对于一轮训练过程来说,在获取到第一图像数据的目标描述信息之后,可以依据该目标描述信息中各个目标预测结果对应的时序,将该目标描述信息拆分成两组信息进行损失计算。其中,第一组信息包括所有检测Query对应的目标预测结果和最近一个历史图像的跟踪Query对应的目标预测结果组成,而且该第一组信息中不存在多个Query跟踪同一个目标的情况。第二组信息是指去掉了最近一个历史图像的跟踪Query之后的其他跟踪Query的目标预测结果,而且该第二组信息中存在多个Query跟踪同一个目标的情况。
步骤73:依据上文第一组信息和上文第一图像数据对应的目标标签数据,得到第一损失。
其中,第一图像数据对应的目标标签数据是指在利用该第一图像数据进行模型训练时所需依据的标签信息(比如,目标对应的跟踪对象标识标签、该
目标对应的边界框标签、该目标对应的类别标签等),以使该目标标签数据能够表示出该第一图像数据中实际存在哪些目标。
另外,本公开不限定上文步骤73的实施方式,例如,其可以采用现有的或者未来出现的任意一种二分匹配损失(bipartite matching loss)进行实施。
基于上文步骤73的相关内容可知,对于一轮训练过程来说,在获取到第一组信息(比如,所有检测Query对应的目标预测结果以及最近一个历史图像的跟踪Query对应的目标预测结果)之后,因此组中不存在多个Query跟踪同一个目标的情况,故可以直接利用二分匹配损失,针对该第一组信息进行损失计算;而且该计算过程大致可以为:对于已跟踪到的目标的跟踪Query对应的目标预测结果来说,直接依据该跟踪Query对应的跟踪对象标识,将该Query对应的目标预测结果与上文第一图像数据对应的目标标签数据中具有该跟踪对象标识的目标进行关联,如此能够实现针对已跟踪到的目标的跟踪Query对应的目标预测结果的匹配过程;然后,将该目标标签数据中剩余的未关联任何目标预测结果的目标视为新出现的目标,并基于边界框以及类别等方面的匹配过程,从所有检测Query对应的目标预测结果中确定出与该新出现的目标相关联的目标预测结果,如此能够实现针对一些检测Query对应的目标预测结果的匹配过程;还有,可以针对那些已经关联到新出现的目标的检测Query对应的目标预测结果配置一个新的跟踪对象标识,以便后续能够基于被关联至该目标标签数据中某个目标的目标预测结果进行二分匹配损失计算,得到这些被关联至该目标标签数据中某个目标的目标预测结果对应的损失值。另外,对于未关联到该目标标签数据中任何目标的目标预测结果需要视为背景这一类别,故可以计算这些未关联到该目标标签数据中任何目标的目标预测结果对应的类别预测损失。
需要说明的是,对于上段中“被关联至该目标标签数据中某个目标的目标预测结果”来说,可以依据该目标预测结果中的目标位置表征数据与该被关联的目标的边界框标签之间的相似度、以及该目标预测结果中的类别置信度与该被关联的目标的类别标签之间的相似度,确定该目标预测结果对应的损失值。
还需要说明的是,对于上文“未关联到该目标标签数据中任何目标的目标预测结果”来说,可以依据该目标预测结果中的类别置信度与该背景对应的类
别标签之间的相似度,确定该目标预测结果对应的损失值。
步骤74:依据上文第二组信息和上文第一图像数据对应的目标标签数据,得到第二损失。
需要说明的是,本公开不限定上文步骤74的实施方式,例如,在一些可能的实施方式下,上文步骤74具体可以包括:依据标签匹配结果、上文第二组信息、以及上文第一图像数据对应的目标标签数据,确定第二损失。其中,该标签匹配结果用于描述该第二组信息中每个目标预测结果与该第一图像数据对应的目标标签数据中哪个目标关联。
另外,本公开不限定上文标签匹配结果的确定过程,例如,其可以采用现有的或者未来出现的任意一种将一个已跟踪到的目标的跟踪Query对应的目标预测结果与某个标签数据中的目标进行关联的方法(比如,二分匹配损失所涉及的关联方式)进行实施。
实际上,在一些应用场景下,上文第二组信息所描述的已跟踪到的目标与上文第一组信息所描述的已跟踪到的目标大致相同,故为了更好地节省计算资源,上文标签匹配结果可以根据上文第一组信息与该第一图像数据对应的目标标签数据之间的匹配结果进行确定。为了便于理解,下面结合示例进行说明。
作为示例,假设上文第二组信息包括跟踪Query1对应的目标预测结果,上文第一组信息包括跟踪Query2对应的目标预测结果,而且该跟踪Query1对应的跟踪对象标识与该跟踪Query2对应的跟踪对象标识相同。基于此该假设可知,该跟踪Query1与该跟踪Query2用于描述同一个已跟踪到的目标,故可以依据上文第一组信息与该第一图像数据对应的目标标签数据之间的匹配结果,确定该跟踪Query2对应的目标预测结果与该第一图像数据对应的目标标签数据中哪个目标相关联;并将确定出的与该跟踪Query2对应的目标预测结果相关联的目标,直接确定为与该跟踪Query1对应的目标预测结果相关联的目标,如此能够实现基于该第一组信息与该第一图像数据对应的目标标签数据之间的匹配结果,确定上文标签匹配结果的目的。
另外,本公开不限定上文第二损失的计算过程,例如,当上文第二组信息包括至少一个目标预测结果(比如,图2所示的由正方形虚线框所框定的Query)时,该第二损失的计算过程具体可以包括下文步骤741-步骤742。
步骤741:依据上文标签匹配结果、上文第二组信息、以及上文第一图像数据对应的目标标签数据,确定各目标预测结果对应的损失。
需要说明的是,本公开不限定上文步骤741的实施方式,例如,其具体可以为:先依据上文标签匹配结果、上文第二组信息、以及上文第一图像数据对应的目标标签数据,确定该第二组信息中各目标预测结果对应的标签信息(比如,边界框标签+类别标签);再依据该第二组信息中的第n个目标预测结果与该第n个目标预测结果对应的标签信息,确定该第n个目标预测结果对应的损失(比如,依据该第n个目标预测结果中的目标位置表征数据与该第n个目标预测结果对应的标签信息中的边界框标签之间的相似度、以及依据第n个目标预测结果中的类别置信度与该第n个目标预测结果对应的标签信息中的类别标签之间的相似度,确定该第n个目标预测结果对应的损失),n为正整数,n≤N,N表示上文第二组信息中目标预测结果的个数。
步骤742:依据上文第二组信息中至少一个目标预测结果对应的损失的平均值,确定第二损失。
基于上文步骤74的相关内容可知,对于上文第二组信息来说,因此组中存在多个Query跟踪同一个目标的情况,故本公开可以借助跟踪一致性损失(tracking object consistency loss),确定该第二组信息对应的第二损失。其中,该跟踪一致性损失保持了与上文二分匹配损失一样的损失计算方式,但是在该跟踪一致性损失的计算过程中可以直接根据上文第一组信息与该第一图像数据对应的目标标签数据之间的匹配结果,确定与该第二组信息中各个目标预测结果相关联的标签信息。另外,基于第二组信息计算所得的第二损失是在此组内取平均,而不是与第一组中所有目标预测结果对应的损失合并之后才取平均的,如此能够有效地避免因第二组信息中的目标预测结果个数过多而导致的不良影响(比如,导致基于该第一组中检测Query对应的目标预测结果所确定的损失被稀释等问题),如此有利于提高模型训练效果。
步骤75:依据第一损失和第二损失,确定目标跟踪模型的训练损失。
需要说明的是,本公开不限定步骤75的实施方式,例如,该步骤75具体可以为:直接将第一损失和第二损失之间的和值,确定为目标跟踪模型的训练损失。又如,该步骤75具体可以为:将第一损失和第二损失进行加权求和,得到该目标跟踪模型的训练损失。还如,该步骤75具体可以为:将第一损失
和第二损失进行集合,得到该目标跟踪模型的训练损失。
基于上文步骤71至步骤75的相关内容可知,如图2所示,在针对上文目标跟踪模型的训练过程中,该目标跟踪模型的训练损失可以根据二分匹配损失与跟踪一致性损失之间的和值进行确定,如此使得该训练损失能够更好地表示出该模型的性能,从而有利于提高模型训练效果。
需要说明的是,本公开不限定上文步骤71-步骤75所示的训练损失的应用场景,例如,当该步骤71-步骤75所示的训练损失应用于上文步骤62(也就是,第二训练阶段)时,可以将该步骤71-步骤75中出现的“目标跟踪模型”替换为上文待优化模型即可。
另外,在针对上文目标跟踪模型的训练过程中,当前帧(比如,上文第一图像数据)所涉及的各个跟踪Query对应的目标预测结果和关联到新出现的目标的检测Query对应的目标预测结果都被分配了相应的跟踪对象标识,故对于任意一个被分配了跟踪对象标识的目标预测结果来说,如果确定被分配了跟踪对象标识的目标预测结果与其关联的目标之间的IOU大于一个预定阈值,则可以确定已较好地跟踪到该目标,故可以将与该目标相关联的目标预测结果中的Query,按照一个预设的概率送入到后续帧中作为历史跟踪Query(也就是,正样本)参与跟踪。同时与跟踪Query相关联的目标位置较近的未分配到跟踪对象标识的检测Query对应的目标预测结果中的Query,会以一定概率作为噪声跟踪Query(也就是,负样本),关联一个不存在的目标(比如,虚拟目标)并送入后续帧中作为历史跟踪Query参与训练,增加模型对噪声的鲁棒性。
基于上文目标跟踪模型的相关内容可知,该目标跟踪模型属于一种基于多Query的端到端多目标跟踪算法,其使用每个目标的多个历史Query构建当前帧中该目标的协同跟踪Query,共同跟踪目标以提升特征的可靠性,以使该目标跟踪模型在高帧率视频中和低帧率视频中都获得更好的性能,且处理速度比较快。
基于本公开实施例提供的图像处理方法,本公开实施例还提供了一种图像处理装置,下面结合图4进行解释和说明。其中,图4为本公开实施例提供的一种图像处理装置的结构示意图。需要说明的是,本公开实施例提供的图像处理装置的技术详情,请参照上文图像处理方法的相关内容。
如图4所示,本公开实施例提供的图像处理装置400,包括:
提取单元401,用于对待处理图像进行特征提取处理,得到特征提取结果;
编码单元402,用于对所述特征提取结果进行编码处理,得到检测查询表征数据;
解码单元403,用于依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息;所述跟踪查询表征数据是依据所述待处理图像对应的多个历史图像的目标描述信息所确定的;所述待处理图像的目标描述信息用于描述所述待处理图像中的至少一个目标。
在一种可能的实施方式下,所述编码单元402,具体用于:对所述特征提取结果进行编码处理,得到图像特征表征数据以及检测查询表征数据;
所述解码单元403,具体用于:依据所述图像特征表征数据、所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述多个历史图像包括第一类图像和第二类图像;所述跟踪查询表征数据的确定过程,包括:从所述第一类图像的目标描述信息中提取第一跟踪目标描述信息;依据所述第一跟踪目标描述信息以及所述第二类图像的目标描述信息,确定所述跟踪查询表征数据。
在一种可能的实施方式下,所述跟踪查询表征数据的确定过程,包括:依据所述第一跟踪目标描述信息中的目标位置表征数据以及所述第二类图像的目标描述信息中的目标内容特征表征数据,确定第二跟踪目标描述信息;根据所述第二跟踪目标描述信息和所述第一跟踪目标描述信息,确定所述跟踪查询表征数据。
在一种可能的实施方式下,所述待处理图像与所述待处理图像对应的多个历史图像属于同一个视频数据;所述待处理图像的时序晚于所述第一类图像的时序;所述第一类图像的时序晚于所述第二类图像的时序。
在一种可能的实施方式下,所述解码单元403,具体用于:依据所述检测查询表征数据和所述跟踪查询表征数据,确定待处理查询表征数据;依据所述待处理查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述解码单元403,具体用于:依据所述检测
查询表征数据和所述跟踪查询表征数据,确定待处理查询表征数据;依据所述图像特征表征数据和所述待处理查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的解码器所确定的;所述解码器包括至少一个解码层,所述解码层中的自注意力模块是利用第一注意力掩模进行实现的,所述第一注意力掩模用于屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互。
在一种可能的实施方式下,所述至少一个目标包括一个或者多个待跟踪目标;所述至少一个解码层包括第一解码层和第二解码层;所述解码器还包括特征融合模块;所述特征融合模块用于针对由所述第一解码层输出的跟踪查询处理数据进行融合处理,得到跟踪查询融合数据;所述跟踪查询处理数据是依据所述第一解码层以及所述跟踪查询表征数据所确定的;所述跟踪查询处理数据包括由所述第一解码层针对各所述待跟踪目标分别输出的处理后的查询表征数据;所述跟踪查询融合数据包括由所述特征融合模块针对各所述处理后的查询表征数据分别输出的融合后的查询表征数据;所述第二解码层用于针对所述跟踪查询融合数据以及由所述第一解码层输出的检测查询处理数据进行处理;所述检测查询处理数据是依据所述第一解码层以及所述检测查询表征数据所确定的。
在一种可能的实施方式下,所述特征融合模块包括信息移除分支网络和信息添加分支网络;
所述跟踪查询融合数据的确定过程,包括:利用所述信息移除分支网络对所述跟踪查询处理数据进行处理,得到保留信息特征;利用所述信息添加分支网络对所述跟踪查询处理数据进行处理,得到待添加信息特征;根据所述保留信息特征和所述待添加信息特征,确定所述跟踪查询融合数据。
在一种可能的实施方式下,所述保留信息特征的确定过程,包括:对所述跟踪查询处理数据进行时序信息提取处理,得到第一时序信息;对所述第一时序信息与所述跟踪查询处理数据进行自注意力处理,得到自注意力处理结果;对所述自注意力处理结果进行全连接处理,得到全连接处理结果;依据所述全连接处理结果,确定待移除信息表征数据;依据所述待移除信息表征数据,对所述跟踪查询处理数据进行信息移除处理,得到所述保留信息特征。
在一种可能的实施方式下,所述待添加信息特征的确定过程,包括:对所述跟踪查询处理数据进行时序信息提取处理,得到第二时序信息;对所述第二时序信息与所述跟踪查询处理数据进行自注意力处理,得到所述待添加信息特征。
在一种可能的实施方式下,所述特征融合模块中的自注意力层是利用第二注意力掩模进行实现的,所述第二注意力掩模用于屏蔽属于不同目标的查询表征数据之间的交互。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;
所述目标跟踪模型的训练损失的确定过程,包括:依据第一图像数据、所述第一图像数据对应的多个历史图像的目标描述信息、以及所述目标跟踪模型,确定所述第一图像数据的目标描述信息;从所述第一图像数据的目标描述信息中提取第一组信息和第二组信息;依据所述第一组信息和所述第一图像数据对应的目标标签数据,得到第一损失;依据所述第二组信息和所述目标标签数据,得到第二损失;依据所述第一损失和所述第二损失,确定所述目标跟踪模型的训练损失。
在一种可能的实施方式下,所述第一组信息对应的时序晚于所述第二组信息对应的时序。
在一种可能的实施方式下,所述第二损失的确定过程,包括:依据标签匹配结果、所述第二组信息、以及所述目标标签数据,确定所述第二损失;所述标签匹配结果是根据所述第一组信息与所述目标标签数据之间的匹配结果所确定的。
在一种可能的实施方式下,所述第二组信息包括至少一个目标预测结果;
所述第二损失的确定过程,包括:依据所述标签匹配结果、所述第二组信息、以及所述目标标签数据,确定各所述目标预测结果对应的损失;依据所述至少一个目标预测结果对应的损失的平均值,确定所述第二损失。
在一种可能的实施方式下,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;
所述目标跟踪模型的训练过程,包括:利用至少一个第二图像数据以及各所述第二图像数据的目标检测标签,对初始模型进行训练,得到待优化模型;
利用至少一个图像序列以及各所述图像序列的目标跟踪标签,对所述待优化模型中的解码模块进行训练,得到所述目标跟踪模型。
在一种可能的实施方式下,所述至少一个目标包括一个或者多个待跟踪目标;所述跟踪查询表征数据包括各所述待跟踪目标在多个历史图像下对应的查询表征数据。
在一种可能的实施方式下,所述待处理图像是指从图像序列中抽取的一个图像;所述目标描述信息包括至少一个目标预测结果;
所述图像处理装置400还包括:
筛选单元,用于从所述目标描述信息中提取满足预设参考条件的目标预测结果;
更新单元,用于利用所述图像序列中的下一帧图像更新所述待处理图像,利用所述满足预设参考条件的目标预测结果中的查询表征数据,更新所述跟踪查询表征数据,并继续执行所述对待处理图像进行特征提取处理的步骤。
在一种可能的实施方式下,所述图像序列为视频数据。
基于上述图像处理装置400的相关内容可知,对于本公开实施例提供的图像处理装置400来说,先对待处理图像进行特征提取处理,得到特征提取结果,以使该特征提取结果用于表征该待处理图像所携带的图像信息;再对该特征提取结果进行编码处理,得到检测查询表征数据(比如,多个检测Query);最后,依据该检测查询表征数据以及该待处理图像对应的跟踪查询表征数据(比如,基于多个历史图像所确定的跟踪Query)进行解码处理,得到该待处理图像的目标描述信息(比如,目标内容特征表征数据、目标位置表征数据、以及类别置信度等),以使该目标描述信息能够表示出该待处理图像中一些目标。其中,因该跟踪查询表征数据是依据该待处理图像对应的多个历史图像的目标描述信息所确定的,以使该跟踪查询表征数据能够更好地表示出该待处理图像中一些目标的历史状态,从而使得基于该跟踪查询表征数据所确定的目标描述信息能够更好地表示出该待处理图像中一些目标,如此有利于提高目标跟踪效果。
另外,本公开实施例还提供了一种电子设备,所述设备包括处理器以及存储器:所述存储器,用于存储指令或计算机程序;所述处理器,用于执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行本公开实施例
提供的图像处理方法的任一实施方式。
参见图5,其示出了适于用来实现本公开实施例的电子设备500的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图5示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图5所示,电子设备500可以包括处理装置(例如中央处理器、图形处理器等)501,其可以根据存储在只读存储器(ROM)502中的程序或者从存储装置508加载到随机访问存储器(RAM)503中的程序而执行各种适当的动作和处理。在RAM503中,还存储有电子设备500操作所需的各种程序和数据。处理装置501、ROM 502以及RAM 503通过总线504彼此相连。输入/输出(I/O)接口505也连接至总线504。
通常,以下装置可以连接至I/O接口505:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置506;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置507;包括例如磁带、硬盘等的存储装置508;以及通信装置509。通信装置509可以允许电子设备500与其他设备进行无线或有线通信以交换数据。虽然图5示出了具有各种装置的电子设备500,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置509从网络上被下载和安装,或者从存储装置508被安装,或者从ROM502被安装。在该计算机程序被处理装置501执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施例提供的电子设备与上述实施例提供的方法属于同一发明构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的有益效果。
本公开实施例还提供了一种计算机可读介质,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行本公开实施例提供的图像处理方法的任一实施方式。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(Hyper Text Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备可以执行上述方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元/模块的名称在某种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或
半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
需要说明的是,本公开中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似部分互相参见即可。对于实施例公开的系统或装置而言,由于其与实施例公开的方法相对应,所以描述的比较简单,相关之处参见方法部分说明即可。
应当理解,在本公开中,“至少一个(项)”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,用于描述关联对象的关联关系,表示可以存在三种关系,例如,“A和/或B”可以表示:只存在A,只存在B以及同时存在A和B三种情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b或c中的至少一项(个),可以表示:a,b,c,“a和b”,“a和c”,“b和c”,或“a和b和c”,其中a,b,c可以是单个,也可以是多个。
还需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
结合本文中所公开的实施例描述的方法或算法的步骤可以直接用硬件、处理器执行的软件模块,或者二者的结合来实施。软件模块可以置于随机存储器(RAM)、内存、只读存储器(ROM)、电可编程ROM、电可擦除可编程ROM、寄存器、硬盘、可需要说明的是,本公开不限定上文步骤33的实施方式,移动磁盘、CD-ROM、或技术领域内所公知的任意其它形式的存储介质中。
对所公开的实施例的上述说明,使本领域专业技术人员能够实现或使用本公开。对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开的精神或范围的情况下,在其它实施例中实现。因此,本公开将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。
Claims (22)
- 一种图像处理方法,包括:对待处理图像进行特征提取处理,得到特征提取结果;对所述特征提取结果进行编码处理,得到检测查询表征数据;依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息;所述跟踪查询表征数据是依据所述待处理图像对应的多个历史图像的目标描述信息所确定的;所述待处理图像的目标描述信息用于描述所述待处理图像中的至少一个目标。
- 根据权利要求1所述的方法,其中,所述多个历史图像包括第一类图像和第二类图像;所述跟踪查询表征数据的确定过程,包括:从所述第一类图像的目标描述信息中提取第一跟踪目标描述信息;依据所述第一跟踪目标描述信息以及所述第二类图像的目标描述信息,确定所述跟踪查询表征数据。
- 根据权利要求2所述的方法,其中,所述依据所述第一跟踪目标描述信息以及所述第二类图像的目标描述信息,确定所述跟踪查询表征数据,包括:依据所述第一跟踪目标描述信息中的目标位置表征数据以及所述第二类图像的目标描述信息中的目标内容特征表征数据,确定第二跟踪目标描述信息;根据所述第二跟踪目标描述信息和所述第一跟踪目标描述信息,确定所述跟踪查询表征数据。
- 根据权利要求2所述的方法,其中,所述待处理图像与所述待处理图像对应的多个历史图像属于同一个视频数据;所述待处理图像的时序晚于所述第一类图像的时序;所述第一类图像的时序晚于所述第二类图像的时序。
- 根据权利要求1所述的方法,其中,所述依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息,包括:依据所述检测查询表征数据和所述跟踪查询表征数据,确定待处理查询表征数据;依据所述待处理查询表征数据进行解码处理,得到所述待处理图像的目标描述信息。
- 根据权利要求1所述的方法,其中,所述待处理图像的目标描述信息是利用预先训练好的解码器所确定的;所述解码器包括至少一个解码层,所述解码层中的自注意力模块是利用第一注意力掩模进行实现的,所述第一注意力掩模用于屏蔽属于同一个目标的在不同历史图像下对应的查询表征数据之间的交互。
- 根据权利要求6所述的方法,其中,所述至少一个目标包括一个或者多个待跟踪目标;所述至少一个解码层包括第一解码层和第二解码层;所述解码器还包括特征融合模块;所述特征融合模块用于针对由所述第一解码层输出的跟踪查询处理数据进行融合处理,得到跟踪查询融合数据;所述跟踪查询处理数据是依据所述第一解码层以及所述跟踪查询表征数据所确定的;所述跟踪查询处理数据包括由所述第一解码层针对各所述待跟踪目标分别输出的处理后的查询表征数据;所述跟踪查询融合数据包括由所述特征融合模块针对各所述处理后的查询表征数据分别输出的融合后的查询表征数据;所述第二解码层用于针对所述跟踪查询融合数据以及由所述第一解码层输出的检测查询处理数据进行处理;所述检测查询处理数据是依据所述第一解码层以及所述检测查询表征数据所确定的。
- 根据权利要求7所述的方法,其中,所述特征融合模块包括信息移除分支网络和信息添加分支网络;所述跟踪查询融合数据的确定过程,包括:利用所述信息移除分支网络对所述跟踪查询处理数据进行处理,得到保留信息特征;利用所述信息添加分支网络对所述跟踪查询处理数据进行处理,得到待添加信息特征;根据所述保留信息特征和所述待添加信息特征,确定所述跟踪查询融合 数据。
- 根据权利要求8所述的方法,其中,所述保留信息特征的确定过程,包括:对所述跟踪查询处理数据进行时序信息提取处理,得到第一时序信息;对所述第一时序信息与所述跟踪查询处理数据进行自注意力处理,得到自注意力处理结果;对所述自注意力处理结果进行全连接处理,得到全连接处理结果;依据所述全连接处理结果,确定待移除信息表征数据;依据所述待移除信息表征数据,对所述跟踪查询处理数据进行信息移除处理,得到所述保留信息特征。
- 根据权利要求8所述的方法,其中,所述待添加信息特征的确定过程,包括:对所述跟踪查询处理数据进行时序信息提取处理,得到第二时序信息;对所述第二时序信息与所述跟踪查询处理数据进行自注意力处理,得到所述待添加信息特征。
- 根据权利要求7所述的方法,其中,所述特征融合模块中的自注意力层是利用第二注意力掩模进行实现的,所述第二注意力掩模用于屏蔽属于不同目标的查询表征数据之间的交互。
- 根据权利要求1所述的方法,其中,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;所述目标跟踪模型的训练损失的确定过程,包括:依据第一图像数据、所述第一图像数据对应的多个历史图像的目标描述信息、以及所述目标跟踪模型,确定所述第一图像数据的目标描述信息;从所述第一图像数据的目标描述信息中提取第一组信息和第二组信息;依据所述第一组信息和所述第一图像数据对应的目标标签数据,得到第一损失;依据所述第二组信息和所述目标标签数据,得到第二损失;依据所述第一损失和所述第二损失,确定所述目标跟踪模型的训练损失。
- 根据权利要求12所述的方法,其中,所述第一组信息对应的时序晚于所述第二组信息对应的时序。
- 根据权利要求12所述的方法,其中,所述依据所述第二组信息和所述目标标签数据,得到第二损失,包括:依据标签匹配结果、所述第二组信息、以及所述目标标签数据,确定所述第二损失;所述标签匹配结果是根据所述第一组信息与所述目标标签数据之间的匹配结果所确定的。
- 根据权利要求14所述的方法,其中,所述第二组信息包括至少一个目标预测结果;所述第二损失的确定过程,包括:依据所述标签匹配结果、所述第二组信息、以及所述目标标签数据,确定各所述目标预测结果对应的损失;依据所述至少一个目标预测结果对应的损失的平均值,确定所述第二损失。
- 根据权利要求1所述的方法,其中,所述待处理图像的目标描述信息是利用预先训练好的目标跟踪模型所确定的;所述目标跟踪模型的训练过程,包括:利用至少一个第二图像数据以及各所述第二图像数据的目标检测标签,对初始模型进行训练,得到待优化模型;利用至少一个图像序列以及各所述图像序列的目标跟踪标签,对所述待优化模型中的解码模块进行训练,得到所述目标跟踪模型。
- 根据权利要求1所述的方法,其中,所述至少一个目标包括一个或者多个待跟踪目标;所述跟踪查询表征数据包括各所述待跟踪目标在多个历史图像下对应的查询表征数据。
- 根据权利要求1-17任一项所述的方法,其中,所述待处理图像是指从图像序列中抽取的一个图像;所述目标描述信息包括至少一个目标预测结果;所述得到所述待处理图像的目标描述信息之后,所述方法还包括:从所述目标描述信息中提取满足预设参考条件的目标预测结果;利用所述图像序列中的下一帧图像更新所述待处理图像,利用所述满足预设参考条件的目标预测结果中的查询表征数据,更新所述跟踪查询表征数据,并继续执行所述对待处理图像进行特征提取处理的步骤。
- 根据权利要求18所述的方法,其中,所述图像序列为视频数据。
- 一种图像处理装置,包括:提取单元,被配置为对待处理图像进行特征提取处理,得到特征提取结果;编码单元,被配置为对所述特征提取结果进行编码处理,得到检测查询表征数据;解码单元,被配置为依据所述检测查询表征数据以及所述待处理图像对应的跟踪查询表征数据进行解码处理,得到所述待处理图像的目标描述信息;所述跟踪查询表征数据是依据所述待处理图像对应的多个历史图像的目标描述信息所确定的;所述待处理图像的目标描述信息用于描述所述待处理图像中的至少一个目标。
- 一种电子设备,包括:处理器和存储器;其中,所述存储器,被配置为存储指令或计算机程序;所述处理器,被配置为执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行权利要求1-19任一项所述的方法。
- 一种计算机可读介质,其中,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行权利要求1-19任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310787598.2A CN119229133A (zh) | 2023-06-29 | 2023-06-29 | 一种图像处理方法、装置、电子设备、计算机可读介质 |
| CN202310787598.2 | 2023-06-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025001898A1 true WO2025001898A1 (zh) | 2025-01-02 |
Family
ID=93937344
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/099597 Ceased WO2025001898A1 (zh) | 2023-06-29 | 2024-06-17 | 一种图像处理方法、装置、电子设备、计算机可读介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN119229133A (zh) |
| WO (1) | WO2025001898A1 (zh) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112347852A (zh) * | 2020-10-10 | 2021-02-09 | 上海交通大学 | 体育运动视频的目标追踪与语义分割方法及装置、插件 |
| CN113963032A (zh) * | 2021-12-01 | 2022-01-21 | 浙江工业大学 | 一种融合目标重识别的孪生网络结构目标跟踪方法 |
| US20220301183A1 (en) * | 2021-08-24 | 2022-09-22 | Beijing Baidu Netcom Science Technology Co., Ltd. | Method and apparatus for tracking object, electronic device, and readable storage medium |
| CN115861893A (zh) * | 2022-12-19 | 2023-03-28 | 哈尔滨工业大学 | 基于可见光图像的多飞行器跟踪方法及系统 |
-
2023
- 2023-06-29 CN CN202310787598.2A patent/CN119229133A/zh active Pending
-
2024
- 2024-06-17 WO PCT/CN2024/099597 patent/WO2025001898A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112347852A (zh) * | 2020-10-10 | 2021-02-09 | 上海交通大学 | 体育运动视频的目标追踪与语义分割方法及装置、插件 |
| US20220301183A1 (en) * | 2021-08-24 | 2022-09-22 | Beijing Baidu Netcom Science Technology Co., Ltd. | Method and apparatus for tracking object, electronic device, and readable storage medium |
| CN113963032A (zh) * | 2021-12-01 | 2022-01-21 | 浙江工业大学 | 一种融合目标重识别的孪生网络结构目标跟踪方法 |
| CN115861893A (zh) * | 2022-12-19 | 2023-03-28 | 哈尔滨工业大学 | 基于可见光图像的多飞行器跟踪方法及系统 |
Non-Patent Citations (1)
| Title |
|---|
| FANGAO ZENG; BIN DONG; YUANG ZHANG; TIANCAI WANG; XIANGYU ZHANG; YICHEN WEI: "MOTR: End-to-End Multiple-Object Tracking with Transformer", ARXIV.ORG, 19 July 2022 (2022-07-19), US, XP091274265 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN119229133A (zh) | 2024-12-31 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112559800B (zh) | 用于处理视频的方法、装置、电子设备、介质和产品 | |
| CN114612759A (zh) | 视频处理方法、查询视频的方法和模型训练方法、装置 | |
| JP7680633B2 (ja) | ビデオクリップの識別方法、装置、機器及び記憶媒体 | |
| WO2024131406A1 (zh) | 模型构建方法、图像分割方法、装置、设备、介质 | |
| CN113033682B (zh) | 视频分类方法、装置、可读介质、电子设备 | |
| CN112907628A (zh) | 视频目标追踪方法、装置、存储介质及电子设备 | |
| US12456171B2 (en) | Image quality adjustment method and apparatus, device, and medium | |
| CN112766284A (zh) | 图像识别方法和装置、存储介质和电子设备 | |
| CN114140613A (zh) | 图像检测方法、装置、电子设备及存储介质 | |
| CN113610034A (zh) | 识别视频中人物实体的方法、装置、存储介质及电子设备 | |
| CN111563398A (zh) | 用于确定目标物的信息的方法和装置 | |
| CN110633716A (zh) | 一种目标对象的检测方法和装置 | |
| CN111160410A (zh) | 一种物体检测方法和装置 | |
| CN118351393A (zh) | 数据生成方法、训练方法、时空定位方法、装置、电子设备、存储介质及程序产品 | |
| CN114419070A (zh) | 一种图像场景分割方法、装置、设备及存储介质 | |
| CN113887615A (zh) | 图像处理方法、装置、设备和介质 | |
| CN120653793A (zh) | 跨模态检索方法、装置、设备及存储介质 | |
| CN117743370A (zh) | 多模态数据检索方法、装置、设备及可读存储介质 | |
| CN115797833B (zh) | 镜头分割、视觉任务处理方法、装置、电子设备以及介质 | |
| CN113140012A (zh) | 图像处理方法、装置、介质及电子设备 | |
| CN119317912A (zh) | 用于数据聚类的采样技术 | |
| CN115019037A (zh) | 对象分割方法及对应模型的训练方法、装置及存储介质 | |
| CN111666449B (zh) | 视频检索方法、装置、电子设备和计算机可读介质 | |
| CN113239215B (zh) | 多媒体资源的分类方法、装置、电子设备及存储介质 | |
| CN111783889B (zh) | 图像识别方法、装置、电子设备和计算机可读介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24830544 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |