WO2019080668A1 - 人物跟踪方法、装置、电子装置及计算机可读介质 - Google Patents

人物跟踪方法、装置、电子装置及计算机可读介质

Info

Publication number
WO2019080668A1
WO2019080668A1 PCT/CN2018/106140 CN2018106140W WO2019080668A1 WO 2019080668 A1 WO2019080668 A1 WO 2019080668A1 CN 2018106140 W CN2018106140 W CN 2018106140W WO 2019080668 A1 WO2019080668 A1 WO 2019080668A1
Authority
WO
WIPO (PCT)
Prior art keywords
tracking
person
character
time window
time
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/106140
Other languages
English (en)
French (fr)
Inventor
周佩明
叶韵
张爱喜
武军晖
陈宇
翁志
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Jingdong Century Trading Co Ltd
Beijing Jingdong Shangke Information Technology Co Ltd
Original Assignee
Beijing Jingdong Century Trading Co Ltd
Beijing Jingdong Shangke Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Jingdong Century Trading Co Ltd, Beijing Jingdong Shangke Information Technology Co Ltd filed Critical Beijing Jingdong Century Trading Co Ltd
Priority to US16/757,412 priority Critical patent/US11270126B2/en
Publication of WO2019080668A1 publication Critical patent/WO2019080668A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques
    • G06F18/243Classification techniques relating to the number of classes
    • G06F18/24323Tree-organised classifiers
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/01Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/207Analysis of motion for motion estimation over a hierarchy of resolutions
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/215Motion-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/223Analysis of motion using block-matching
    • G06T7/231Analysis of motion using block-matching using full search
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/246Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/70Determining position or orientation of objects or cameras
    • G06T7/73Determining position or orientation of objects or cameras using feature-based methods
    • G06T7/74Determining position or orientation of objects or cameras using feature-based methods involving reference images or patches
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/52Surveillance or monitoring of activities, e.g. for recognising suspicious objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10016Video; Image sequence
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20072Graph-based image processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30232Surveillance
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30241Trajectory
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/07Target detection

Definitions

  • the present disclosure generally relates to the field of video processing technologies, and in particular, to a person tracking method, apparatus, electronic device, and computer readable medium.
  • the character tracking technology is widely used in video surveillance, and is mostly implemented by the "detection + association" scheme, that is, the characters are detected from each frame of the video, and the character detection frames are associated with the identity of the characters.
  • detection + association the characters are detected from each frame of the video, and the character detection frames are associated with the identity of the characters.
  • the character tracking technology can be divided into real-time processing and global processing. These two schemes have their own advantages and disadvantages.
  • the real-time processing scheme can get the tracking results in real time, it has poor processing ability in solving the occlusion, crossover and long-term disappearance of characters, when the judgment of a certain frame image is wrong.
  • the image of the subsequent frame of this frame image cannot be corrected.
  • the global processing scheme can continuously correct the error of the previous frame during the processing, the fault tolerance is high, but the processing speed is very slow, and the tracking result cannot be obtained in real time.
  • the present disclosure provides a person tracking method, apparatus, electronic device, and computer readable medium, which solve the above technical problems.
  • a person tracking method including:
  • a continuous tracking path is constructed through successive time windows to obtain a tracking result of the target person.
  • N the value of N is related to the response delay and the number of frames of the image per second
  • N Ts*Ns, where Ts is the response delay, and Ns is the each The number of frames of the image in seconds.
  • acquiring the tracking path of the target person in the time window according to the N frame image includes:
  • the global tracking algorithm is adopted according to the N frame image, and the character association matching is performed by:
  • the virtual child node is used to simulate a frame image that is not displayed by the target person in the time window;
  • Character association matching is performed based on the person association score.
  • calculating a person association score in the multi-fork decision tree includes:
  • the person association score is calculated based on the Euclidean distance between the features.
  • the continuous time window includes n time windows, and the constructing the continuous tracking path by the continuous time window includes:
  • a continuous tracking path is constructed by character association matching between successive time windows using a real-time tracking algorithm.
  • a person tracking device including:
  • An image acquisition module configured to acquire an N frame image in units of time windows
  • the time window tracking module is configured to acquire a tracking path of the target person in the time window according to the N frame image
  • the continuous tracking path module is configured to construct a continuous tracking path through consecutive time windows to obtain a tracking result of the target person.
  • an electronic device including a processor; a memory storing instructions for the processor to control operations as described above.
  • a computer readable medium having stored thereon computer executable instructions that, when executed by a processor, implement the information push method as described above.
  • a video with N frames of images is processed by a global tracking scheme in units of time windows.
  • the tracking path of the target person in the image is obtained, and the time window is processed by a similar real-time tracking scheme.
  • the method optimizes the tracking method of global tracking and real-time tracking in the prior art, and can realize real-time character tracking. And it also avoids the problem of exponential expansion of computation and memory as the depth of the decision tree increases.
  • FIG. 1 is a flow chart of a person tracking method provided in an embodiment of the present disclosure.
  • FIG. 2 shows a flow chart of step S120 of FIG. 1 according to an embodiment of the present disclosure.
  • FIG. 3 shows a schematic diagram of a multi-fork decision tree constructed in an embodiment of the present disclosure.
  • FIG. 4 shows a flow chart of step S24 of FIG. 2 in an embodiment of the present disclosure.
  • FIG. 5 is a diagram showing a comparison between the person tracking using the method provided by the embodiment and the method provided by the prior art for character tracking according to an embodiment of the present disclosure.
  • FIG. 6 shows a schematic diagram of a person tracking device provided in another embodiment of the present disclosure.
  • FIG. 7 shows a schematic diagram of the in-time window tracking module 620 of FIG. 6 in accordance with an embodiment of the present disclosure.
  • FIG. 8 shows a schematic diagram of the calculated molecular module 624 of FIG. 7 in accordance with another embodiment of the present disclosure.
  • FIG. 9 is a schematic structural diagram of a computer system suitable for implementing an electronic device according to an embodiment of the present application.
  • the character tracking technology can be divided into real-time processing and global processing.
  • the real-time processing refers to determining the information of the next frame image according to the information of the previous frame image frame by frame, that is, determining the position of the image of the next frame according to the historical tracking path;
  • Processing refers to the association of information such as the location of a person in all frame images after collecting the entire video, and finally obtaining the path of the identity of the person in the entire video.
  • Character tracking technology can be divided into location-based tracking, feature-based tracking, and location-based tracking based on the information that is tracked.
  • a multi-hypothesis tracking (MHT) algorithm can be used based on location and feature to construct a multi-fork decision tree structure of a character to achieve association matching between multiple frames, and multi-target character tracking.
  • MHT multi-hypothesis tracking
  • the algorithm can not achieve real-time processing, limiting its application range, especially in the field of video-aware processing.
  • the algorithm cannot accurately maintain the identity recognition function of the original character, that is, the person before the occlusion cannot be accurately associated with the person who reappears after occlusion.
  • the algorithm adopts a more complex scoring model in the case of character association matching, and associates it by online training classifier, resulting in high computational complexity.
  • the existing MHT algorithm has the above drawbacks, and the present disclosure improves the person tracking method.
  • the present disclosure improves the person tracking method.
  • FIG. 1 is a flowchart of a person tracking method according to an embodiment of the present disclosure, including the following steps:
  • step S110 an N-frame image is acquired in units of time windows.
  • step S120 the tracking path of the target person is acquired in the time window according to the N frame image.
  • step S130 a continuous tracking path is constructed through successive time windows to obtain a tracking result of the target person.
  • the character tracking method provided by the present disclosure processes a video with N frames of images in a time window unit by using a global tracking scheme to obtain a tracking path of a target person in the image, and time windows between the time windows. It is processed by a scheme similar to real-time tracking.
  • the method can not only realize quasi-real-time character tracking, but also avoid calculation and memory with increasing tree depth. The problem of expansion.
  • the length of the time window is N frames
  • 0 to (N-1), N to (2N-1), 2N to (3N-1)...(n-1)N ⁇ (nN-1) is divided into n time windows.
  • a multi-frame global tracking algorithm is adopted, and the tracking path in the N-frame image is given after the time window ends.
  • the tracking path here may be a tracking path or multiple targets of a target person.
  • the tracking path of the character needs to be determined according to the needs in the specific application scenario.
  • the time window and the time window can be used to perform character association matching in a manner similar to real-time tracking.
  • the tracking result of the (n-1)th time window can be given during the nth time window.
  • the method delays the N frame to obtain the tracking result, the real-time processing effect can be achieved.
  • the tracking calculation of the character path can be performed after acquiring all the frame images of the video.
  • the global tracking algorithm is performed every time the N-frame image is acquired, and the tracking path of the target person is obtained, which is obviously time-sensitive. It has been greatly improved, so it is possible to achieve a tracking path of the target person in near real time.
  • the N frame image is divided into a time window in step S110, wherein the value of N is related to the response delay.
  • Ts is the response delay
  • Ns is the number of frames of the image per second.
  • the value of N in this embodiment may be determined according to requirements such as algorithm performance, computing resources, and the like, in addition to the response delay and the number of frames of the image per second.
  • requirements such as algorithm performance, computing resources, and the like
  • step S120 of the embodiment a global tracking algorithm is used according to the N-frame image to perform character association matching, and a tracking path of one or more target persons is obtained.
  • FIG. 2 is a flowchart showing a tracking path of one or more target persons by using a global tracking algorithm according to an N-frame image in the step S120 of the embodiment to obtain a tracking path of one or more target persons, including the following steps:
  • step S21 one root node and a plurality of child nodes are acquired based on the N frame image.
  • step S22 a multi-fork decision tree of a certain target person is constructed in time order according to the root node and the child nodes.
  • step S23 a virtual child node is added to the multi-fork decision tree, and the virtual child node is used to simulate a frame of image that the target person does not display in the time window.
  • step S24 the person association score is calculated in the multi-fork decision tree.
  • step S25 person-related matching is performed based on the person-related score.
  • FIG. 3 is a schematic diagram of a multi-fork decision tree constructed in this embodiment, wherein the image of the (t-1 frame) is the root node of the multi-fork decision tree, and the t-th frame includes 3 sub-nodes and a t+1th frame. Contains 6 child nodes. As shown in FIG. 3, two images of the t-th frame are used as two child nodes, and one virtual child node is added on the basis of the above; the four images of the t+1th frame are used as four child nodes, and based on this, The virtual child node of the t-th frame further simulates two virtual child nodes, that is, the image in the dotted line frame is a virtual child node.
  • the method provided in this embodiment introduces a virtual "missing" node in the multi-fork decision tree structure of the traditional MHT algorithm, and is used to simulate that the character is missed, occluded, missing, etc. in the t-th frame.
  • the virtual child nodes in the multi-fork decision tree participate in the decision tree optimization algorithm like ordinary root nodes or child nodes.
  • the decision tree is optimized according to the character association score.
  • the difference from the ordinary child node is that the value of the character association score of the virtual child node should be set to a threshold value that can distinguish between the character features to represent the current t-frame.
  • the detection frame matching the root node (the t-1th frame) is not found in the picture, and the current frame (the t-th frame) is missing, and the speculation based on the picture of the t-1th frame may appear in the t-th frame.
  • the image of the picture is not found in the picture, and the current frame (the t-th frame) is missing, and the speculation based on the picture of the t-1th frame may appear in the t-th frame.
  • FIG. 4 shows a flowchart for calculating a person association score in the multi-fork decision tree in step S24 in the present embodiment, including the following steps:
  • the feature is extracted by the person convolutional neural network.
  • the feature here is a high-dimensional vector, for example, information including the posture, clothing, and position of the character belongs to the feature.
  • One-dimensional is a high-dimensional vector, for example, information including the posture, clothing, and position of the character belongs to the feature.
  • the person association score is calculated based on the Euclidean distance between the features.
  • the decision is based on an online training classifier, so that the computational complexity of the calculation process is high, resulting in a long program execution time and a relatively high memory consumption. Big.
  • the score is calculated by simplifying the calculation process of the character association score and directly calculating the Euclidean distance between the features extracted by the Convolutional Neural Network (CNN).
  • the Euclidean distance between the features of a target character in the two frames before and after is detected, the lower the similarity between the two, the lower the character correlation score; otherwise, if two frames are detected before and after
  • step S25 the character association matching is performed according to step S25, that is, based on the character association score when the decision is made according to the multi-fork decision tree shown in FIG. 3, the person whose association score is higher is a better match, and thus can be more The optimal tracking path for a certain target person is obtained in the fork decision tree.
  • the solution provided in this embodiment can greatly save computation time, thereby ensuring real-time processing can be realized.
  • the continuous time window in step S130 may include n time windows, that is, the time window may be continuously acquired during the video production process, and the N frames in the time window are completed after the end of a time window.
  • the image is globally processed to obtain the tracking path in the time window, so the tracking path of the target person in the n-1th time window is calculated in the nth time window, and then the real-time tracking algorithm is used between successive time windows. Construct a continuous tracking path through character association matching.
  • the method and principle of matching the character association match with the character association score in the above step S25 are the same, and details are not described herein again.
  • FIG. 5 is a view showing a comparison between the person tracking and the method provided by the prior art by using the method provided by the embodiment, and the time window of the length of 16 frames is taken as an example, and the first behavior is provided by the embodiment.
  • the method performs the result of the person tracking
  • the second behavior uses the method provided by the prior art to perform the result of the person tracking.
  • the characters in the image of the 0th to 15th frames in the second line have a phenomenon of "missing" in multiple frames due to problems such as missed detection and occlusion, such as frames 2, 3, 4, 6, and 7 are missing.
  • "Phenomenon; and the tracking path of the characters in the 0 to 15 frame images of the first line is continuous, and the tracking result is relatively complete. It can be seen that the method provided by the embodiment can effectively solve the phenomenon of "missing" of the person tracking and construct a more complete and continuous tracking path.
  • the character tracking method provided by the embodiment improves the current multi-hypothesis tracking (MHT) algorithm, introduces the concept of time window, performs global tracking in the time window, and obtains the target person tracking path, and retains the MHT multi-frame.
  • MHT multi-hypothesis tracking
  • real-time character tracking can be realized, and the application range of the algorithm can be expanded, which can be widely applied in real-time video monitoring, video analysis, security and the like in public places.
  • the character tracking method combines the advantages of multiple tracking algorithms (global basis and real-time tracking), and strives to improve the tracking accuracy under complex scenes (severe occlusion or long-term disappearance between people), and proposes to use A low-cost solution for real-time processing.
  • FIG. 6 shows a schematic diagram of a person tracking device provided in another embodiment of the present disclosure.
  • the person tracking device 600 includes an image acquisition module 610, a time window tracking module 620, and a continuous tracking path module 630.
  • the image acquisition module 610 is configured to acquire an N frame image in units of time windows; the time window tracking module 620 is configured to acquire a tracking path of the target person in the time window according to the N frame image; the continuous tracking path module 630 is configured to pass A continuous time window constructs a continuous tracking path to obtain the tracking result of the target person.
  • the image acquisition module 610 assumes that the length of the time window is N frames, then 0 to (N-1), N to (2N-1), 2N to (3N-1), (n-1)N to (nN) -1) are divided into n time windows, respectively.
  • a multi-frame global tracking algorithm is adopted, and the tracking path in the N-frame image is given after the time window ends.
  • the tracking path here may be a tracking path or multiple targets of a target person.
  • the tracking path of the character needs to be determined according to the needs in the specific application scenario.
  • the time window and the time window can be used to perform character association matching in a manner similar to real-time tracking.
  • the tracking result of the (n-1)th time window can be given during the nth time window.
  • the method delays the N frame to obtain the tracking result, the effect of real-time processing can be achieved.
  • the global tracking algorithm is performed every time the N-frame image is acquired, and the target person tracking path is obtained, which is obviously aging. The sex is greatly improved, so the tracking path of the character can be obtained in real time.
  • the image acquisition module 610 divides the N frame image into a time window, where the value of N is related to the response delay.
  • Ts is the response delay
  • Ns is the number of frames of the image per second.
  • the value of N in this embodiment may be determined according to requirements such as algorithm performance, computing resources, and the like, in addition to the response delay and the number of frames of the image per second.
  • requirements such as algorithm performance, computing resources, and the like
  • the time window tracking module 620 performs a person association matching according to the N frame image by using a global tracking algorithm to obtain a tracking path of one or more target characters.
  • FIG. 7 shows a schematic diagram of the tracking module 620 in the time window of FIG. 6 in this embodiment.
  • the time window tracking module 620 includes a node acquisition sub-module 621, a sequence construction sub-module 622, a virtual sub-module 623, a calculated molecular module 624, and an associated sub-module 625.
  • the node acquisition sub-module 621 is configured to acquire one root node and a plurality of child nodes according to the N-frame image.
  • the sequence construction sub-module 622 is configured to construct a multi-fork decision tree of a certain target person in chronological order according to the root node and the child nodes.
  • the virtual sub-module 623 is configured to add a virtual child node in the multi-fork decision tree, and the virtual child node is used to simulate a frame of image that the target person does not display in the time window.
  • the calculated numerator module 624 is configured to calculate a person association score in the multi-fork decision tree.
  • the association sub-module 625 is configured to perform character association matching based on the person association score.
  • the calculated molecular module 624 includes a feature extraction unit 6241 and an Euclidean distance calculation unit 6242, and the feature extraction unit 6241 is configured to utilize The character convolutional neural network extracts features; the Euclidean distance calculation unit 6242 is configured to calculate a person association score based on the Euclidean distance between the features.
  • the decision is based on an online training classifier, so that the computational complexity of the calculation process is high, resulting in a long program execution time and a relatively high memory consumption. Big.
  • This embodiment calculates the score by simplifying the calculation process of the character association score and directly calculating the Euclidean distance between the features extracted by the character convolutional neural network (CNN). If the Euclidean distance between the features of a target character in the two frames before and after is detected, the lower the similarity between the two, the lower the character correlation score; otherwise, if two frames are detected before and after The smaller the Euclidean distance between the features of a certain target character in the image, the higher the similarity between the two, and the higher the score of the character association score.
  • CNN character convolutional neural network
  • the character tracking device improves the current multi-hypothesis tracking (MHT) algorithm, introduces the concept of time window, performs global tracking in the time window, and obtains a target person tracking path, while retaining the MHT multi-frame.
  • MHT multi-hypothesis tracking
  • real-time character tracking can be realized, and the application range of the algorithm can be expanded, which can be widely applied in real-time video monitoring, video analysis, security and the like in public places.
  • a virtual "missing" node is introduced into the decision tree structure to simulate the frame when a target character is occluded, missed, and the latter leaves the screen to solve the occlusion, missing, etc. problem.
  • the character tracking device combines the advantages of multiple tracking algorithms (global basis and real-time tracking) to focus on improving the tracking accuracy under complex scenes (severe occlusion or long-term disappearance between people) and proposes to use A low-cost solution for real-time processing.
  • the present disclosure also provides an electronic device comprising a processor and a memory, the memory storing instructions for the processor to control the following operations:
  • the N-frame image is obtained in units of time windows; the tracking path of the target person is acquired in the time window according to the N-frame image; the continuous tracking path is constructed through the continuous time window, and the tracking result of the target person is obtained.
  • FIG. 9 a block diagram of a computer system 900 suitable for use in implementing the electronic device of the embodiments of the present application is shown.
  • the electronic device shown in FIG. 9 is merely an example, and should not impose any limitation on the function and scope of use of the embodiments of the present application.
  • computer system 900 includes a central processing unit (CPU) 901 that can be loaded into a program in random access memory (RAM) 903 according to a program stored in read only memory (ROM) 902 or from storage portion 907. And perform various appropriate actions and processes.
  • RAM random access memory
  • ROM read only memory
  • various programs and data required for the operation of the system 900 are also stored.
  • the CPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904.
  • An input/output (I/O) interface 905 is also coupled to bus 904.
  • the following components are connected to the I/O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a storage portion 908 including a hard disk or the like. And a communication portion 909 including a network interface card such as a LAN card, a modem, or the like. The communication section 909 performs communication processing via a network such as the Internet.
  • Driver 910 is also connected to I/O interface 905 as needed.
  • a removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 910 as needed so that a computer program read therefrom is installed into the storage portion 908 as needed.
  • an embodiment of the present disclosure includes a computer program product comprising a computer program embodied on a computer readable medium, the computer program comprising program code for executing the method illustrated in the flowchart.
  • the computer program can be downloaded and installed from the network via the communication portion 909, and/or installed from the removable medium 911.
  • the central processing unit (CPU) 901 the above-described functions defined in the system of the present application are executed.
  • the computer readable medium illustrated in the present application may be a computer readable signal medium or a computer readable medium or any combination of the two.
  • the computer readable medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of computer readable media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable Read only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
  • a computer readable medium can be any tangible medium that can contain or store a program, which can be used by or in connection with an instruction execution system, apparatus or device.
  • a computer readable signal medium may include a data signal that is propagated in the baseband or as part of a carrier, carrying computer readable program code. Such propagated data signals can take a variety of forms including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the foregoing.
  • the computer readable signal medium can also be any computer readable medium other than a computer readable medium that can transmit, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device.
  • Program code embodied on a computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
  • each block of the flowchart or block diagrams can represent a module, a program segment, or a portion of code that includes one or more Executable instructions.
  • the functions noted in the blocks may also occur in a different order than that illustrated in the drawings. For example, two successively represented blocks may in fact be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending upon the functionality involved.
  • each block of the block diagrams or flowcharts, and combinations of blocks in the block diagrams or flowcharts can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be used A combination of dedicated hardware and computer instructions is implemented.
  • the units involved in the embodiments of the present application may be implemented by software or by hardware.
  • the described unit may also be provided in the processor, for example, as a processor comprising a transmitting unit, an obtaining unit, a determining unit and a first processing unit.
  • the name of these units does not constitute a limitation on the unit itself in some cases.
  • the sending unit may also be described as “a unit that sends a picture acquisition request to the connected server”.
  • the present disclosure also provides a computer readable medium, which may be included in the apparatus described in the above embodiments, or may be separately present and not incorporated into the apparatus.
  • the computer readable medium carries one or more programs, and when the one or more programs are executed by the device, the device includes:
  • the N-frame image is obtained in units of time windows; the tracking path of the target person is acquired in the time window according to the N-frame image; the continuous tracking path is constructed through the continuous time window, and the tracking result of the target person is obtained.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Mathematical Physics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Computational Linguistics (AREA)
  • Medical Informatics (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Databases & Information Systems (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Human Computer Interaction (AREA)
  • Image Analysis (AREA)

Abstract

一种人物跟踪方法,以时间窗为单位获取N帧图像;根据N帧图像在时间窗内获取目标人物的跟踪路径;通过连续的时间窗构建连续的跟踪路径,得到目标人物的跟踪结果。该方法能够实现准实时的人物追踪,避免计算和内存随决策树的深度增加而指数膨胀。

Description

人物跟踪方法、装置、电子装置及计算机可读介质
交叉引用
本公开要求于2017年10月23日递交的、申请号为:201710996063.0,发明创造名称为“人物跟踪方法、装置、电子装置及计算机可读介质”的中国专利申请的优先权,在此将上述中国专利申请公开的内容以引入的方式并入本申请。
技术领域
本公开总体涉及视频处理技术领域,具体而言,涉及一种人物跟踪方法、装置、电子装置及计算机可读介质。
背景技术
人物跟踪技术在视频监控中使用广泛,多采用“检测+关联”的方案来实现,即先从视频的每帧图像中将人物检测出来,再把这些人物检测框与人物的身份关联起来,达到对人物轨迹进行跟踪的目的。
按照跟踪处理时延的不同,人物跟踪技术可以分为实时处理与全局处理。这两种方案各有优缺点,实时处理方案虽然能够实时得到跟踪结果,但是在解决人物之间的遮挡、交叉以及人物长时失踪方面,处理能力较差,当某一帧图像的判断出现错误时,无法对这一帧图像后续帧的图像进行纠正。而全局处理方案虽然可以在处理过程中不断纠正前帧的错误,容错性较高,但是处理速度很慢,导致无法实时得到跟踪结果。
因此,现有技术中的技术方案还存在有待改进之处。
在所述背景技术部分公开的上述信息仅用于加强对本公开的背景的理解,因此它可以包括不构成对本领域普通技术人员已知的现有技术的信息。
发明内容
本公开提供一种人物跟踪方法、装置、电子装置及计算机可读介质,解决上述技术问题。
本公开的其他特性和优点将通过下面的详细描述变得显然,或部分地通过本公开的实践而习得。
根据本公开的一方面,提供一种人物跟踪方法,包括:
以时间窗为单位获取N帧图像;
根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径;
通过连续的时间窗构建连续的跟踪路径,得到所述目标人物的跟踪结果。
在本公开的一个实施例中,其中N的取值与响应时延有关。
在本公开的一个实施例中,其中N的取值与所述响应时延以及每秒内图像的帧数有关,N=Ts*Ns,其中Ts为所述响应时延,Ns为所述每秒内图像的帧数。
在本公开的一个实施例中,根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径包括:
根据所述N帧图像采用全局跟踪算法,进行人物关联匹配,得到一个或多个所述目标人物的跟踪路径。
在本公开的一个实施例中,根据所述N帧图像采用全局跟踪算法,进行人物关联匹配包括:
根据所述N帧图像获取一个根节点和多个子节点;
根据所述一个根节点和所述多个子节点按照时间顺序构建某一目标人物的多叉决策树;
在所述多叉决策树中增加虚拟的子节点,所述虚拟的子节点用于模拟所述目标人物在所述时间窗内未显示的一帧图像;
在所述多叉决策树中计算人物关联得分;
根据所述人物关联得分进行人物关联匹配。
在本公开的一个实施例中,在所述多叉决策树中计算人物关联得分包括:
利用人物卷积神经网络提取特征;
根据所述特征之间的欧氏距离计算得到所述人物关联得分。
在本公开的一个实施例中,所述连续的时间窗包括n个时间窗,所述通过连续的时间窗构建连续的跟踪路径包括:
在第n个时间窗计算得到第n-1个时间窗的目标人物的跟踪路径;
采用实时跟踪算法在连续的时间窗之间通过人物关联匹配构建连续的跟踪路径。
根据本公开的再一方面,提供一种人物跟踪装置,包括:
图像获取模块,被配置于以时间窗为单位获取N帧图像;
时间窗内跟踪模块,被配置于根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径;
连续跟踪路径模块,被配置于通过连续的时间窗构建连续的跟踪路径,得到所述目标人物的跟踪结果。
根据本公开的又一方面,提供一种电子装置,包括处理器;存储器,存储用于所述处理器控制如上所述的操作的指令。
根据本公开的另一方面,提供一种计算机可读介质,其上存储有计算机可执行指令,所述可执行指令被处理器执行时实现如上所述的信息推送方法。
根据本公开实施例提供的人物跟踪方法、装置、电子装置及计算机可读介质,在现有MHT算法的基础上,以时间窗为单位对含有N帧图像的视频采用全局跟踪的方案进行处 理,得到图像中目标人物的追踪路径,而时间窗之间采用类似实时跟踪的方案进行处理,该方法通过对现有技术中全局跟踪与实时跟踪的追踪方法进行优化,不仅能够实现准实时的人物追踪,而且还能避免计算和内存随决策树的深度增加而指数膨胀的问题。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性的,并不能限制本公开。
附图说明
通过参照附图详细描述其示例实施例,本公开的上述和其它目标、特征及优点将变得更加显而易见。
图1示出本公开一实施例中提供的一种人物跟踪方法的流程图。
图2示出本公开一实施例图1中步骤S120的流程图。
图3示出本公开一实施例中构建的多叉决策树的示意图。
图4示出本公开一实施例中图2步骤S24的流程图。
图5示出本公开一实施例中采用本实施例提供的方法进行人物跟踪与采用现有技术提供的方法进行人物跟踪的对比图。
图6示出本公开另一实施例中提供的一种人物跟踪装置的示意图。
图7示出本公开一实施例图6中时间窗内跟踪模块620的示意图。
图8示出本公开另一实施例图7中计算得分子模块624的示意图。
图9示出本公开一实施例提供的适于用来实现本申请实施例的电子装置的计算机系统的结构示意图。
具体实施方式
现在将参考附图更全面地描述示例实施方式。然而,示例实施方式能够以多种形式实施,且不应被理解为限于在此阐述的范例;相反,提供这些实施方式使得本公开将更加全面和完整,并将示例实施方式的构思全面地传达给本领域的技术人员。附图仅为本公开的示意性图解,并非一定是按比例绘制。图中相同的附图标记表示相同或类似的部分,因而将省略对它们的重复描述。
此外,所描述的特征、结构或特性可以以任何合适的方式结合在一个或更多实施方式中。在下面的描述中,提供许多具体细节从而给出对本公开的实施方式的充分理解。然而,本领域技术人员将意识到,可以实践本公开的技术方案而省略所述特定细节中的一个或更多,或者可以采用其它的方法、组元、装置、步骤等。在其它情况下,不详细示出或描述公知结构、方法、装置、实现、材料或者操作以避免喧宾夺主而使得本公开的各方面变得模糊。
附图中所示的一些方框图是功能实体,不一定必须与物理或逻辑上独立的实体相对应。可以采用软件形式来实现这些功能实体,或在一个或多个硬件模块或集成电路中实现这些功能实体,或在不同网络和/或处理器装置和/或微控制器装置中实现这些功能实体。
为使本发明的目的、技术方案和优点更加清楚明白,以下结合具体实施例,并参照附图,对本发明进一步详细说明。
人物跟踪技术可以分为实时处理和全局处理,其中实时处理是指逐帧地依据前帧图像的信息去判断下一帧图像的信息,即根据历史跟踪路径判断下一帧图像人物的位置;全局处理是指在收集整段视频之后,依据所有帧图像中人的位置等信息进行关联,最终得到人物身份在整段视频中的路径。
人物跟踪技术按照跟踪所依据信息的不同,可以分为基于位置的跟踪、基于特征的跟踪以及基于位置和特征的跟踪。例如,可以基于位置和特征采用“多假设跟踪”(multi-hypothesis tracking,简称MHT)算法,构建人物的多叉决策树结构实现多帧之间的关联匹配,进行多目标的人物跟踪。但是现有技术中的MHT算法有以下主要问题:
a.该算法不能实现实时处理,限制其应用范围,尤其是在视频感知处理领域的应用。
b.尽管该算法已经对多叉决策树进行很多消除操作以减少计算量,但是当视频长度增加时,仍面临计算量和内存消耗的指数式膨胀。
c.该算法在人物之间发生遮挡后,不能准确地保持原人物的身份识别功能,即不能准确地将被遮挡前的人与遮挡后重新出现的人关联在一起。
d.该算法在人物关联匹配时,采用较为复杂的得分模型,并通过在线训练分类器的方式进行关联,导致其计算复杂度较高。
基于上述,现有MHT算法存在上述缺陷,本公开对人物跟踪方法进行改进,具体请参照下述实施例。
图1示出本公开一实施例中提供的一种人物跟踪方法的流程图,包括以下步骤:
如图1所示,在步骤S110中,以时间窗为单位获取N帧图像。
如图1所示,在步骤S120中,根据N帧图像在时间窗内获取目标人物的跟踪路径。
如图1所示,在步骤S130中,通过连续的时间窗构建连续的跟踪路径,得到目标人物的跟踪结果。
本公开提供的人物追踪方法在现有MHT算法的基础上,以时间窗为单位对含有N帧图像的视频采用全局跟踪的方案进行处理,得到图像中目标人物的追踪路径,而时间窗之间采用类似实时跟踪的方案进行处理,该方法通过对现有技术中全局跟踪与实时跟踪的追踪方法进行优化,不仅能够实现准实时的人物追踪,而且还能避免计算和内存随树深度增加而指数膨胀的问题。
在本公开的一些实施例中,假设时间窗的长度为N帧,则第0~(N-1),N~(2N-1),2N~(3N-1)…(n-1)N~(nN-1)分别划分为n个时间窗。首先,在每个时间窗内,均采用多帧全局跟踪算法,并在时间窗结束后给出此N帧图像中的跟踪路径,这里的跟踪路径可以为一个目标人物的跟踪路径或多个目标人物的跟踪路径,需要根据具体应用场景中的需求而决定。其次,在本实施例中时间窗与时间窗之间,可以采用类似实时跟踪的方式进行人物关联匹配。以此推算,可以在第n个时间窗期间给出第(n-1)个时间窗的跟踪结果,虽 然该方法会延迟N帧得出跟踪结果,但能够实现实时处理的效果。相比较于现有技术中的全局跟踪方案需要获取视频的所有帧图像之后才能进行人物路径的追踪计算,本实施例每获取N帧图像就进行全局跟踪算法,得到目标人物跟踪路径,显然时效性得到很大提升,因此可以实现准实时得出目标人物的跟踪路径。
本实施例中步骤S110中将N帧图像划分为一个时间窗,其中N的取值与响应时延有关。在具体应用场景中,N的取值可以与响应时延以及每秒内图像的帧数有关,即N=Ts*Ns,其中Ts为响应时延,Ns为每秒内图像的帧数。例如,如果应用场景允许有10秒钟的响应时延,则可取N的取值可以为10秒内的帧数,假设视频每秒播放24帧,则N=10*24=240。
还需要说明的是,本实施例中N的取值大小,除了与响应时延以及每秒内图像的帧数有关以外,还可以根据对算法性能、计算资源等要求来决定。一般来说,N的取值越大,算法的容错性越好,即解决长时遮挡等问题的效果越好;反之,N的取值越小,响应时延越小,而计算资源消耗也越小。因此,在具体应用中需要根据具体对响应时延、算法性能、计算资源等需求决定。
在本实施例的步骤S120中,根据N帧图像采用全局跟踪算法,进行人物关联匹配,得到一个或多个目标人物的跟踪路径。
图2示出本实施例步骤S120中根据N帧图像采用全局跟踪算法,进行人物关联匹配,得到一个或多个目标人物的跟踪路径的流程图,包括以下步骤:
如图2所示,在步骤S21中,根据N帧图像获取一个根节点和多个子节点。
如图2所示,在步骤S22中,根据根节点和子节点按照时间顺序构建某一目标人物的多叉决策树。
如图2所示,在步骤S23中,在多叉决策树中增加虚拟的子节点,虚拟的子节点用于模拟目标人物在时间窗内未显示的一帧图像。
如图2所示,在步骤S24中,在多叉决策树中计算人物关联得分。
如图2所示,在步骤S25中,根据人物关联得分进行人物关联匹配。
图3示出本实施例中构建的多叉决策树的示意图,其中第(t-1帧)的图像为该多叉决策树的根节点,第t帧包含3个子节点以及第t+1帧包含6个子节点。如图3所示,第t帧的两幅图像作为2个子节点,并在此基础上增加1个虚拟的子节点;第t+1帧的四幅图像作为4个子节点,并在此基础上根据第t帧虚拟的子节点进一步模拟得到2个虚拟的子节点,即虚线框中的图像为虚拟的子节点。
本实施例提供的方法通过在传统MHT算法的多叉决策树结构中引入虚拟“失踪”节点,用于模拟该人物在第t帧时被漏检、或者被遮挡、失踪等情形。该多叉决策树中虚拟的子节点与普通的根节点或子节点一样,参与决策树的优化算法。根据人物关联得分进行决策树的优化,与普通子节点不同之处在于,虚拟的子节点的人物关联得分的数值应设置为可以进行人物特征之间区分的阈值,用以表征在第t帧当前画面内没有找到与根节点(第 t-1帧)相匹配的检测框,认为当前帧(第t帧)已失踪,而且根据第t-1帧的画面进行推测得到的可能出现在第t帧画面的图像。
在本实施例中,图4示出本本实施例中步骤S24中在多叉决策树中计算人物关联得分的流程图,包括以下步骤:
如图4所示,在步骤S41中,利用人物卷积神经网络提取特征,需要说明的是,这里的特征为高维的向量,例如包括人物的体态、衣着、位置等信息均属于特征中的一维。
如图4所示,在步骤S42中,根据特征之间的欧氏距离计算得到人物关联得分。通常计算人物关联得分的过程中除了根据计算公式进行计算以外,还要基于一在线训练的分类器来进行决策,这样计算过程的计算复杂度较高,导致程序执行时间很长,内存消耗也较大。本实施例通过简化人物关联得分的计算过程,直接由人物卷积神经网络(Convolutional Neural Network,简称CNN)提取的特征之间的欧氏距离来计算得分。如果检测到前后两帧图像中某一目标人物的特征之间的欧氏距离越大,表示二者之间的相似度越低,人物关联得分也就越低;反之,如果检测到前后两帧图像中某一目标人物的特征之间的欧氏距离越小,表示二者之间的相似度越高,人物关联得分也就得分越高。
得到人物关联得分之后,根据步骤S25进行人物关联匹配,即根据图3所示的多叉决策树进行决策时以人物关联得分为依据,人物关联得分较高者为较优匹配,从而可以在多叉决策树中得到针对某一目标人物的最优的跟踪路径。本实施例提供的方案可以大大节省计算时间,从而保证可以实现实时处理。
在本实施例中,步骤S130中连续的时间窗可以包括n个时间窗,即在视频生产过程中就可以连续地获取时间窗,并在一个时间窗结束后就对该时间窗内的N帧图像进行全局处理,得到该时间窗内的跟踪路径,因此在第n个时间窗计算得到第n-1个时间窗的目标人物的跟踪路径,之后再采用实时跟踪算法在连续的时间窗之间通过人物关联匹配构建连续的跟踪路径。这里的人物关联匹配与上述步骤S25根据人物关联得分进行匹配的方法和原理相同,此处不再赘述。
图5示出采用本实施例提供的方法进行人物跟踪与采用现有技术提供的方法进行人物跟踪的对比图,均以长度为16帧的时间窗为例,第一行为采用本实施例提供的方法进行人物跟踪的结果,第二行为采用现有技术提供的方法进行人物跟踪的结果。如图5所示,第二行第0~15帧图像中的人物由于漏检、遮挡等问题出现多帧“失踪”的现象,如第2、3、4、6、7帧均存在“失踪”现象;而第一行第0~15帧图像中的人物的跟踪路径比较连续,跟踪结果较为完整。由此可见,本实施例提供的方法可以有效解决人物跟踪“失踪”现象,并构建出更加完整和连续的跟踪路径。
综上所述,本实施例提供的人物跟踪方法通过对当前多假设跟踪(MHT)算法进行改进,引入时间窗的概念,在时间窗进行全局跟踪,得到目标人物跟踪路径,在保留MHT多帧全局匹配的优势前提下,可实现准实时的人物跟踪,扩大了算法的应用范围,可以广泛应用在公共场所的实时的视频监控、视频分析、安防等方面。
其次,通过对MHT的树形结构进行改进,在决策树结构中引入虚拟的“失踪”节点来模拟该帧图像时某个目标人物被遮挡、漏检、后者离开画面,以解决遮挡、失踪等问题。
最后,还对人物关联匹配算法中人物关联得分的计算进行简化,直接依据人物之间的卷积神经网络(CNN)特征的欧氏距离作为人物关联得分,简化得分计算方法以减小计算复杂度,从而使计算效率得到较为明显的提升。
总之,该人物跟踪方法综合多种跟踪算法(全局根据和实时跟踪)的优点,着力提高复杂场景(人与人之间较严重的遮挡或长时失踪)下的跟踪准确率,并提出能够用于实时处理的、计算代价低的解决方案。
图6示出本公开另一实施例中提供的一种人物跟踪装置的示意图。如图6所示,该人物跟踪装置600包括:图像获取模块610、时间窗内跟踪模块620和连续跟踪路径模块630。
图像获取模块610被配置于以时间窗为单位获取N帧图像;时间窗内跟踪模块620被配置于根据N帧图像在时间窗内获取目标人物的跟踪路径;连续跟踪路径模块630被配置于通过连续的时间窗构建连续的跟踪路径,得到目标人物的跟踪结果。
其中图像获取模块610中假设时间窗的长度为N帧,则第0~(N-1),N~(2N-1),2N~(3N-1)…(n-1)N~(nN-1)分别划分为n个时间窗。首先,在每个时间窗内,均采用多帧全局跟踪算法,并在时间窗结束后给出此N帧图像中的跟踪路径,这里的跟踪路径可以为一个目标人物的跟踪路径或多个目标人物的跟踪路径,需要根据具体应用场景中的需求而决定。其次,在本实施例中时间窗与时间窗之间,可以采用类似实时跟踪的方式进行人物关联匹配。以此推算,可以在第n个时间窗期间给出第(n-1)个时间窗的跟踪结果,虽然该方法会延迟N帧得出跟踪结果,但能够实现实时处理的效果。相比较于现有技术中的全局跟踪方案需要获取视频的所有帧图像之后才能进行目标人物路径的追踪计算,本实施例每获取N帧图像就进行全局跟踪算法,得到目标人物跟踪路径,显然时效性得到很大提升,因此可以实现准实时得出人物的跟踪路径。
图像获取模块610将N帧图像划分为一个时间窗,其中N的取值与响应时延有关。在具体应用场景中,N的取值可以与响应时延以及每秒内图像的帧数有关,即N=Ts*Ns,其中Ts为响应时延,Ns为每秒内图像的帧数。例如,如果应用场景允许有10秒钟的响应时延,则可取N的取值可以为10秒内的帧数,假设视频每秒播放24帧,则N=10*24=240。
还需要说明的是,本实施例中N的取值大小,除了与响应时延以及每秒内图像的帧数有关以外,还可以根据对算法性能、计算资源等要求来决定。一般来说,N的取值越大,算法的容错性越好,即解决长时遮挡等问题的效果越好;反之,N的取值越小,响应时延越小,而计算资源消耗也越小。因此,在具体应用中需要根据具体对响应时延、算法性能、计算资源等需求决定。
时间窗内跟踪模块620根据N帧图像采用全局跟踪算法,进行人物关联匹配,得到一个或多个目标人物的跟踪路径。
图7示出本实施例图6中时间窗内跟踪模块620的示意图。如图7所示,时间窗内跟 踪模块620包括:节点获取子模块621、顺序构建子模块622、虚拟子模块623、计算得分子模块624和关联子模块625。
节点获取子模块621被配置于根据N帧图像获取一个根节点和多个子节点。顺序构建子模块622被配置于根据根节点和子节点按照时间顺序构建某一目标人物的多叉决策树。虚拟子模块623被配置于在多叉决策树中增加虚拟的子节点,虚拟的子节点用于模拟目标人物在时间窗内未显示的一帧图像。计算得分子模块624被配置于在多叉决策树中计算人物关联得分。关联子模块625被配置于根据人物关联得分进行人物关联匹配。
图8示出本实施例图7中计算得分子模块624的示意图,如图8所示,计算得分子模块624包括特征提取单元6241和欧氏距离计算单元6242,特征提取单元6241被配置于利用人物卷积神经网络提取特征;欧氏距离计算单元6242被配置于根据特征之间的欧氏距离计算得到人物关联得分。通常计算人物关联得分的过程中除了根据计算公式进行计算以外,还要基于一在线训练的分类器来进行决策,这样计算过程的计算复杂度较高,导致程序执行时间很长,内存消耗也较大。本实施例通过简化人物关联得分的计算过程,直接由人物卷积神经网络(CNN)提取的特征之间的欧氏距离来计算得分。如果检测到前后两帧图像中某一目标人物的特征之间的欧氏距离越大,表示二者之间的相似度越低,人物关联得分也就越低;反之,如果检测到前后两帧图像中某一目标人物的特征之间的欧氏距离越小,表示二者之间的相似度越高,人物关联得分也就得分越高。
该装置中各个模块的功能参见上述方法实施例中的相关描述,此处不再赘述。
综上所述,本实施例提供的人物跟踪装置通过对当前多假设跟踪(MHT)算法进行改进,引入时间窗的概念,在时间窗进行全局跟踪,得到目标人物跟踪路径,在保留MHT多帧全局匹配的优势前提下,可实现准实时的人物跟踪,扩大了算法的应用范围,可以广泛应用在公共场所的实时的视频监控、视频分析、安防等方面。
其次,通过对MHT的树形结构进行改进,在决策树结构中引入虚拟的“失踪”节点来模拟该帧时某个目标人物被遮挡、漏检、后者离开画面,以解决遮挡、失踪等问题。
最后,还对人物关联匹配算法中人物关联得分的计算进行简化,直接依据人物之间的卷积神经网络(CNN)特征的欧氏距离作为人物关联得分,简化得分计算方法以减小计算复杂度,从而使计算效率得到较为明显的提升。
总之,该人物跟踪装置综合多种跟踪算法(全局根据和实时跟踪)的优点,着力提高复杂场景(人与人之间较严重的遮挡或长时失踪)下的跟踪准确率,并提出能够用于实时处理的、计算代价低的解决方案。
另一方面,本公开还提供了一种电子装置,包括处理器和存储器,存储器存储用于上述处理器控制以下的操作的指令:
以时间窗为单位获取N帧图像;根据N帧图像在时间窗内获取目标人物的跟踪路径;通过连续的时间窗构建连续的跟踪路径,得到目标人物的跟踪结果。
下面参考图9,其示出了适于用来实现本申请实施例的电子装置的计算机系统900的 结构示意图。图9示出的电子装置仅仅是一个示例,不应对本申请实施例的功能和使用范围带来任何限制。
如图9所示,计算机系统900包括中央处理单元(CPU)901,其可以根据存储在只读存储器(ROM)902中的程序或者从存储部分907加载到随机访问存储器(RAM)903中的程序而执行各种适当的动作和处理。在RAM 903中,还存储有系统900操作所需的各种程序和数据。CPU 901、ROM 902以及RAM 903通过总线904彼此相连。输入/输出(I/O)接口905也连接至总线904。
以下部件连接至I/O接口905:包括键盘、鼠标等的输入部分906;包括诸如阴极射线管(CRT)、液晶显示器(LCD)等以及扬声器等的输出部分907;包括硬盘等的存储部分908;以及包括诸如LAN卡、调制解调器等的网络接口卡的通信部分909。通信部分909经由诸如因特网的网络执行通信处理。驱动器910也根据需要连接至I/O接口905。可拆卸介质911,诸如磁盘、光盘、磁光盘、半导体存储器等等,根据需要安装在驱动器910上,以便于从其上读出的计算机程序根据需要被安装入存储部分908。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信部分909从网络上被下载和安装,和/或从可拆卸介质911被安装。在该计算机程序被中央处理单元(CPU)901执行时,执行本申请的系统中限定的上述功能。
需要说明的是,本申请所示的计算机可读介质可以是计算机可读信号介质或者计算机可读介质或者是上述两者的任意组合。计算机可读介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本申请中,计算机可读介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本申请中,计算机可读的信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读的信号介质还可以是计算机可读介质以外的任何计算机可读介质,该计算机可读介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:无线、电线、光缆、RF等等,或者上述的任意合适的组合。
附图中的流程图和框图,图示了按照本申请各种实施例的系统、方法和计算机程序产 品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,上述模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图或流程图中的每个方框、以及框图或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本申请实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。所描述的单元也可以设置在处理器中,例如,可以描述为:一种处理器包括发送单元、获取单元、确定单元和第一处理单元。其中,这些单元的名称在某种情况下并不构成对该单元本身的限定,例如,发送单元还可以被描述为“向所连接的服务端发送图片获取请求的单元”。
另一方面,本公开还提供了一种计算机可读介质,该计算机可读介质可以是上述实施例中描述的设备中所包含的;也可以是单独存在,而未装配入该设备中。上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被一个该设备执行时,使得该设备包括:
以时间窗为单位获取N帧图像;根据N帧图像在时间窗内获取目标人物的跟踪路径;通过连续的时间窗构建连续的跟踪路径,得到目标人物的跟踪结果。
应清楚地理解,本公开描述了如何形成和使用特定示例,但本公开的原理不限于这些示例的任何细节。相反,基于本公开公开的内容的教导,这些原理能够应用于许多其它实施方式。
以上具体地示出和描述了本公开的示例性实施方式。应可理解的是,本公开不限于这里描述的详细结构、设置方式或实现方法;相反,本公开意图涵盖包含在所附权利要求的精神和范围内的各种修改和等效设置。

Claims (10)

  1. 一种人物跟踪方法,其中,包括:
    以时间窗为单位获取N帧图像;
    根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径;
    通过连续的时间窗构建连续的跟踪路径,得到所述目标人物的跟踪结果。
  2. 根据权利要求1所述的人物跟踪方法,其中,其中N的取值与响应时延有关。
  3. 根据权利要求2所述的人物跟踪方法,其中,其中N的取值与所述响应时延以及每秒内图像的帧数有关,N=Ts*Ns,其中Ts为所述响应时延,Ns为所述每秒内图像的帧数。
  4. 根据权利要求1所述的人物跟踪方法,其中,根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径包括:
    根据所述N帧图像采用全局跟踪算法,进行人物关联匹配,得到一个或多个所述目标人物的跟踪路径。
  5. 根据权利要求4所述的人物跟踪方法,其中,根据所述N帧图像采用全局跟踪算法,进行人物关联匹配包括:
    根据所述N帧图像获取一个根节点和多个子节点;
    根据所述一个根节点和所述多个子节点按照时间顺序构建某一目标人物的多叉决策树;
    在所述多叉决策树中增加虚拟的子节点,所述虚拟的子节点用于模拟所述目标人物在所述时间窗内未显示的一帧图像;
    在所述多叉决策树中计算人物关联得分;
    根据所述人物关联得分进行人物关联匹配。
  6. 根据权利要求5所述的人物跟踪方法,其中,在所述多叉决策树中计算人物关联得分包括:
    利用人物卷积神经网络提取特征;
    根据所述特征之间的欧氏距离计算得到所述人物关联得分。
  7. 根据权利要求1所述的人物跟踪方法,其中,所述连续的时间窗包括n个时间窗,所述通过连续的时间窗构建连续的跟踪路径包括:
    在第n个时间窗计算得到第n-1个时间窗的目标人物的跟踪路径;
    采用实时跟踪算法在连续的时间窗之间通过人物关联匹配构建连续的跟踪路径。
  8. 一种人物跟踪装置,其中,包括:
    图像获取模块,被配置于以时间窗为单位获取N帧图像;
    时间窗内跟踪模块,被配置于根据所述N帧图像在所述时间窗内获取目标人物的跟踪路径;
    连续跟踪路径模块,被配置于通过连续的时间窗构建连续的跟踪路径,得到所述目标人物的跟踪结果。
  9. 一种电子装置,其中,包括:
    处理器;
    存储器,存储用于所述处理器控制如权利要求1-7任一项所述的操作的指令。
  10. 一种计算机可读介质,其上存储有计算机可执行指令,其中,所述可执行指令被处理器执行时实现如权利要求1-7任一项所述的信息推送方法。
PCT/CN2018/106140 2017-10-23 2018-09-18 人物跟踪方法、装置、电子装置及计算机可读介质 Ceased WO2019080668A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US16/757,412 US11270126B2 (en) 2017-10-23 2018-09-18 Person tracking method, device, electronic device, and computer readable medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201710996063.0 2017-10-23
CN201710996063.0A CN109697393B (zh) 2017-10-23 2017-10-23 人物跟踪方法、装置、电子装置及计算机可读介质

Publications (1)

Publication Number Publication Date
WO2019080668A1 true WO2019080668A1 (zh) 2019-05-02

Family

ID=66226075

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/106140 Ceased WO2019080668A1 (zh) 2017-10-23 2018-09-18 人物跟踪方法、装置、电子装置及计算机可读介质

Country Status (3)

Country Link
US (1) US11270126B2 (zh)
CN (1) CN109697393B (zh)
WO (1) WO2019080668A1 (zh)

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110443833B (zh) * 2018-05-04 2023-09-26 佳能株式会社 对象跟踪方法和设备
CN110533685B (zh) 2019-08-30 2023-10-24 腾讯科技(深圳)有限公司 对象跟踪方法和装置、存储介质及电子装置
CN110533700B (zh) * 2019-08-30 2023-08-29 腾讯科技(深圳)有限公司 对象跟踪方法和装置、存储介质及电子装置
CN112181667B (zh) * 2020-10-30 2023-08-08 中国科学院计算技术研究所 一种用于目标跟踪的多假设树虚拟化管理方法
CN115236672A (zh) * 2021-05-12 2022-10-25 上海仙途智能科技有限公司 障碍物信息生成方法、装置、设备及计算机可读存储介质
US12143756B2 (en) * 2021-10-08 2024-11-12 Target Brands, Inc. Multi-camera person re-identification

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104200488A (zh) * 2014-08-04 2014-12-10 合肥工业大学 一种基于图表示和匹配的多目标跟踪方法
US20150081258A1 (en) * 2012-08-28 2015-03-19 Numerica Corp. Tracking multiple particles in biological systems
CN106527496A (zh) * 2017-01-13 2017-03-22 平顶山学院 面向无人机航拍图像序列的空中目标快速跟踪方法

Family Cites Families (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5813863A (en) * 1996-05-01 1998-09-29 Sloane; Sharon R. Interactive behavior modification system
US6954678B1 (en) * 2002-09-30 2005-10-11 Advanced Micro Devices, Inc. Artificial intelligence system for track defect problem solving
US8208067B1 (en) * 2007-07-11 2012-06-26 Adobe Systems Incorporated Avoiding jitter in motion estimated video
CN102436662B (zh) * 2011-11-29 2013-07-03 南京信息工程大学 一种非重叠视域多摄像机网络中的人体目标跟踪方法
CN103324937B (zh) * 2012-03-21 2016-08-03 日电(中国)有限公司 标注目标的方法和装置
US9570113B2 (en) * 2014-07-03 2017-02-14 Gopro, Inc. Automatic generation of video and directional audio from spherical content
CN104484571A (zh) * 2014-12-22 2015-04-01 深圳先进技术研究院 一种基于边缘距离排序的集成学习机修剪方法及系统
CN104881882A (zh) * 2015-04-17 2015-09-02 广西科技大学 一种运动目标跟踪与检测方法
CN105354588A (zh) * 2015-09-28 2016-02-24 北京邮电大学 一种构造决策树的方法
CN105205500A (zh) * 2015-09-29 2015-12-30 北京邮电大学 基于多目标跟踪与级联分类器融合的车辆检测
CN105577773B (zh) * 2015-12-17 2019-01-04 清华大学 基于分布式节点及虚拟总线模型的智能车数据平台架构
GB2545658A (en) * 2015-12-18 2017-06-28 Canon Kk Methods, devices and computer programs for tracking targets using independent tracking modules associated with cameras
CN107203756B (zh) * 2016-06-06 2020-08-28 亮风台(上海)信息科技有限公司 一种识别手势的方法与设备
CN106846361B (zh) * 2016-12-16 2019-12-20 深圳大学 基于直觉模糊随机森林的目标跟踪方法及装置
CN106846355B (zh) * 2016-12-16 2019-12-20 深圳大学 基于提升直觉模糊树的目标跟踪方法及装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150081258A1 (en) * 2012-08-28 2015-03-19 Numerica Corp. Tracking multiple particles in biological systems
CN104200488A (zh) * 2014-08-04 2014-12-10 合肥工业大学 一种基于图表示和匹配的多目标跟踪方法
CN106527496A (zh) * 2017-01-13 2017-03-22 平顶山学院 面向无人机航拍图像序列的空中目标快速跟踪方法

Also Published As

Publication number Publication date
US11270126B2 (en) 2022-03-08
US20200334466A1 (en) 2020-10-22
CN109697393B (zh) 2021-11-30
CN109697393A (zh) 2019-04-30

Similar Documents

Publication Publication Date Title
WO2019080668A1 (zh) 人物跟踪方法、装置、电子装置及计算机可读介质
CN110378264B (zh) 目标跟踪方法及装置
Zhou et al. Deep continuous conditional random fields with asymmetric inter-object constraints for online multi-object tracking
JP7270617B2 (ja) 歩行者流量ファネル生成方法及び装置、プログラム、記憶媒体、電子機器
US11468680B2 (en) Shuffle, attend, and adapt: video domain adaptation by clip order prediction and clip attention alignment
CN114913200A (zh) 基于时空轨迹关联的多目标跟踪方法及系统
CN110807410B (zh) 关键点定位方法、装置、电子设备和存储介质
CN112528786B (zh) 车辆跟踪方法、装置及电子设备
CN107274433A (zh) 基于深度学习的目标跟踪方法、装置及存储介质
TWI734375B (zh) 圖像處理方法、提名評估方法及相關裝置
Tan et al. A multiple object tracking algorithm based on YOLO detection
CN110390294B (zh) 一种基于双向长短期记忆神经网络的目标跟踪方法
CN119131265B (zh) 基于多视角一致性的三维全景场景理解方法及装置
JP7408898B2 (ja) 音声エンドポイント検出方法、装置、電子機器、及び記憶媒体
CN111402303A (zh) 一种基于kfstrcf的目标跟踪架构
JP2023530796A (ja) 認識モデルトレーニング方法、認識方法、装置、電子デバイス、記憶媒体及びコンピュータプログラム
CN110008789A (zh) 多类物体检测与识别的方法、设备及计算机可读存储介质
CN111815670A (zh) 多视图目标跟踪方法、装置、系统、电子终端、及存储介质
Han et al. ORT: Occlusion-robust for multi-object tracking
CN114627556A (zh) 动作检测方法、动作检测装置、电子设备以及存储介质
Shi et al. Lane detection by variational auto-encoder with normalizing flow for autonomous driving
CN114820723B (zh) 一种基于联合检测和关联的在线多目标跟踪方法
CN110458867B (zh) 一种基于注意力循环网络的目标跟踪方法
CN115880776B (zh) 关键点信息的确定方法和离线动作库的生成方法、装置
Quan et al. People tracking accuracy improvement in video by matching relevant trackers and YOLO family detectors

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18871562

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 19.08.2020)

122 Ep: pct application non-entry in european phase

Ref document number: 18871562

Country of ref document: EP

Kind code of ref document: A1