EP1751716A1 - A hybrid graphical model for on-line multicamera tracking - Google Patents

A hybrid graphical model for on-line multicamera tracking

Info

Publication number
EP1751716A1
EP1751716A1 EP05750259A EP05750259A EP1751716A1 EP 1751716 A1 EP1751716 A1 EP 1751716A1 EP 05750259 A EP05750259 A EP 05750259A EP 05750259 A EP05750259 A EP 05750259A EP 1751716 A1 EP1751716 A1 EP 1751716A1
Authority
EP
European Patent Office
Prior art keywords
model
distribution
variables
tracking
observations
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP05750259A
Other languages
German (de)
French (fr)
Inventor
Bernardus Johannes Anthonius KRÖSE
Wojciech Piotr Zajdel
Ali Taylan Cemgil
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Stichting voor de Technische Wetenschappen STW
Original Assignee
Stichting voor de Technische Wetenschappen STW
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Stichting voor de Technische Wetenschappen STW filed Critical Stichting voor de Technische Wetenschappen STW
Priority to EP05750259A priority Critical patent/EP1751716A1/en
Publication of EP1751716A1 publication Critical patent/EP1751716A1/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/277Analysis of motion involving stochastic approaches, e.g. using Kalman filters
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/23Clustering techniques
    • G06F18/232Non-hierarchical techniques
    • G06F18/2321Non-hierarchical techniques using statistics or function optimisation, e.g. modelling of probability density functions
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/292Multi-camera tracking
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/24Aligning, centring, orientation detection or correction of the image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/762Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks
    • G06V10/763Non-hierarchical techniques, e.g. based on statistics of modelling distributions
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30232Surveillance
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30241Trajectory

Definitions

  • Tracking objects e.g. people, cars
  • CCTV closed-circuit TV
  • These systems comprise a network of sparsely distributed cameras, where a camera provides only information about local appearances of objects within its field of view (FOV) .
  • FOV field of view
  • Wide-area tracking aims to reconstruct global trajectories of objects given their local appearances.
  • the number of objects is large and viewing fields of cameras do not cover the entire region of interest: Therefore, methods developed for visual tracking within a single-scene are not immediately applicable to problems with sparsely distributed cameras.
  • tracking takes place at a frame-to-frame level.
  • smooth motion provides es- sential cues for resolving association ambiguities.
  • multi-camera tracking takes place at a camera-to-camera level. Observations (e.g. local appearance features) arrive one at a time, within irregular time intervals as objects move from one camera location to another. The time gap between successive observations of the same object may be very long. This hinders the use of smooth motion cues that facilitate tracking in a single-scene setting.
  • MHT multiple hypotheses tracking
  • MHT is an on-line, deterministic algorithm that maintains a list of high probability associations and prunes trajectories with low probability. It is usually a fast and easy to implement method. However, it is a heuristic approximation method with limited accuracy. Moreover, pruning does not allow to preserve any information about the discarded trajectories. In noisy environments, such information could be necessary for proper association of future observations. Alternatively, data association can be solved using approaches with a sound theoretical basis, like stochastic ap- proximation methods: Markov Chain Monte Carlo (MCMC) [11] or sequential Monte Carlo (SMC, also known as particle filters) [8] . An MCMC method was applied for tracking vehicles with multiple cameras on a motorway network [20] .
  • MCMC Markov Chain Monte Carlo
  • SMC sequential Monte Carlo
  • Stochastic methods have the attractive theoretical property that with in- creased computational resources, on average, they are guaranteed to give better solutions.
  • the promised accuracy is obtained at a very high computational cost.
  • foreground/background classification of pixels is an important preprocessing step in many computer vision problems, such as object tracking and identification where one wishes to suppress background pixels to collect reliable statistics about object features of interest.
  • a practical method has to be robust since all the subsequent inference depends on it, but it needs to be also efficient to meet eventual real-time operation requirements.
  • object (s) should be assigned to the foreground.
  • the key assumption here is that the visual sensors and the background image are static and objects of interest are in motion (e.g.
  • the present invention provides a method for tracking of multiple objects with non- overlapping or partly overlapping cameras. Furthermore, the present invention provides a system comprising one or more processing units, two or more non-overlapping cameras operatively connected to the processing units and software stored in memory to execute the above mentioned methods.
  • the preferred embodiments are dedicated to non- overlapping cameras, although applications for partly overlapping cameras are very well conceivable, although not yet realised in experience.
  • the above method provides for tracking of multiple objects with no -overlapping cameras, wherein each object is identified with a label and wherein a probabilistic dependency between the labels and observations of objects is defined, wherein the posterior probability distributions in the model take the form of mixtures, wherein the mixture is approximated.
  • the above method preferably assumes that an observation of an object includes appearance features defined as three colour vectors, wherein each colour vector is the average colour of pixels within one of three horizontal regions selected from the object's image.
  • the method preferably assumes that the appearance features from all observations of a single object are generated by a single Gaussian distribution.
  • the method preferably estimates the parameters of object's specific Gaussian model (mean, covariance matrix) with a Bayesian inference scheme, where the mean and the covariance are given a joint "Normal-Inverse Wishart" prior. Such an approach avoids iterative search for maximum-likelihood estimates of these parameters .
  • the method preferably assumes that appearance features of all objects are generated by a Dirichlet process mixture model, which is used for estimation of the number of distinct objects from the data.
  • the Dirichlet process mixture model is preferably extended with auxiliary variables that allow to model Markovian dependencies between spatio-temporal observa- tions of the same object.
  • the method preferably estimates trajectories of objects by computing posterior distribution of the labels conditioned on the available observations, wherein an assumed- density filtering (ADF) algorithm is used for online computa- tion of filtered distributions of the model's latent variables .
  • ADF assumed- density filtering
  • the present invention provides a method for background/foreground classification of pixels in a visual scene, wherein the class of each pixel is identified with a binary mask variable, and wherein a probabilistic dependency between the observed pixels and their masks is defined, wherein long-range correlations between masks are a-priori imposed.
  • the method according to this further aspect imposes long-range correlations between pixel binary masks by assuming that each binary mask is obtained by taking the sign of a corresponding continuous variable.
  • the correlations are a-priori imposed by defining a joint prior Gaussian distribution on all continuous variables.
  • the method preferably maintains the mask's long-range correlations across consecutive frames of a video sequence by assuming the continuous variables evolve between frames according to linear Gaussian process.
  • the method preferably estimates the masks by computing posterior distributions of binary variables conditioned on the observed pixel values, wherein the posterior distribution are approximated with an Expectation-Propagation algorithm.
  • Fig. 1 shows an example of the considered tracking problem.
  • Figs. 2 - 7 depict a complete pass (event) of a person through a camera viewing field. Six observations (events) are being shown corresponding to three different persons. The goal of tracking is to reconstruct global trajectories between the cameras .
  • Fig. 8 shows a graphical representation of the generative model .
  • the continuous random variables are shown as ovals, the discrete - as rectangles.
  • the arcs be- tween observations that mark the dependencies induced by the dynamics Dk were not drawn.
  • Fig. 9 shows a compact graphical representation of the ⁇ generative model .
  • This representation captures the essential structure of the model .
  • a single node H k represents now all association-related variables ⁇ Sk, Ck, Z k ⁇ 1) , ... , Z k ⁇ k) ⁇ at slice k.
  • This representation shows the relationship with Dirichlet process mixture models and mixed-memory Markov models.
  • Figs. 10A and B show a building plan were observations for experiments were taken. The grey areas show camera viewing fields.
  • Fig. IOC shows the movement constraints assumed by the distributions P ⁇ and P ⁇ o - Arrows indicate possible transitions. All other transitions have zero probability.
  • Figs. 11A and B show results of tracking with the hybrid model.
  • Each column in Fig. 11A gives a filtered distribution of label; Qk (Sk) • The true labels are indicated as dots.
  • Fig. 11B shows a filtered distribution of counter; ' Q k ( Ck) in each column.
  • Fig. 12 shows appearance features of a person. The appearance is described with three colour vectors. Each colour vector is the average colour of pixels within one of the three horizontal regions R 1 , R 2 , R3 selected in the image. The height of a region is 25% of the persons' height. The "skip" areas are meant to remove parts that have little discriminative power. Their height is fixed at 12.5% of the total height.
  • Fig. 13 shows an example of a recovered trajectory.
  • the sequence of images represents local appearances at the indicated cameras.
  • the trajectory includes such observations Y that have labels S ⁇ - 5, as shown in Fig. 11A.
  • Fig. 14A shows results of tracking using only the ap- pearance features .
  • Fig. 15A shows the expected kernels (expected mean and expected covariance) shown as a 2D PCA projection of the original 9D feature space, obtained when tracking was based on dynamics and appearance features .
  • Fig. 15B shows the expected kernels (expected mean and expected covariance) shown as a 2D PCA projection of the original 9D feature space, the kernels obtained when tracking was based only on appearance features .
  • Figs. 16A-16C show a few video frames from an office scene.
  • Fig. 16A shows a vertical slice through time, taken along the vertical line shown in Fig. 16B on the video frames and variations in RGB values of a single pixel (at row 200) in Fig. 16C. Occlusions and shadows can be clearly seen as immediate changes in intensity. Fluctuations in a single color channel illustrate the difficulty of detection when pixels are processed independently.
  • Figs. 17A-17D show random draws from correlated Gaus- sians on 256 dimensions arranged as a 16 x 16 grey scale image at the top row.
  • the correlation matrix of these Gaussians is chosen such that the correlation between two pixels values is a function of spatial distance between them.
  • the left column corresponds to strong correlation and right column to weak correlation.
  • the bottom row shows the signs of greyscale images as a binary image.
  • Fig. 18A shows a graphical model of a single time frame.
  • the rectangles are plates that denote K repetitions of nodes inside. Square and oval nodes correspond to discrete and continuous variables, respectively.
  • Variables f, b and y are vectors denoting the RGB components. Dotted arcs correspond to the regime switch mechanism that is triggered when an object leaves the scene .
  • Fig. 18B shows a loopy factor graph of a single slice used in EP iterations. Forward messages to the next time slice (not shown) are passed only once.
  • Fig. 20A shows typical data generated from the model.
  • Fig. 2 OB shows estimated means of colour features ⁇ ⁇ t) .
  • Fig. 20C shows filtered estimates of object masks.
  • Fig. 20D shows standard deviations of colour features ⁇ (t) .
  • Fig. 21 shows colour features estimated from the video sequence in Fig. 16B with superimposed filtered estimate of asks at the top and colour histograms in HSV space at the bottom.
  • a camera can track a person through its FOV. As soon as the person leaves the viewing field, the camera reports a summarized description of the person. Such a description is our primarily input, so it is referred to as an observation. Observations from all cameras are processed centrally.
  • Y k , k 1,2,..., denote the th in time order observa- tion from any camera.
  • the symbol O k represents appearance, e.g. color features computed while a person was visible within an FOV.
  • the symbol D k represents spatio-temporal features of the observation, e.g. the location and the time of observation.
  • An important feature of the approach is not to assume a fixed number of tracked objects beforehand. Given a sequence of observations ⁇ Y ⁇ ,Y 2/ — ⁇ the task is on-line association of a new observation Y k and estimation of the number of distinct individuals found in the data.
  • (Labels) is the prior information about the trajecto- ries.
  • the posterior is used to select the most likely label for every observation.
  • Labels) is more intuitive to specify than direct finding of trajectories.
  • the proposed generative model allows to apply any in- ference method for computing the posterior probabilities given the observations. However, due to the intractability of the data association, only the approximate methods are practical [19] . In this paper, an on-line, assumed-density filtering algorithm [4] is applied. Note that, the generative model can also be a basis for a multiple hypothesis tracker; in this case the model evaluates the likelihood of trajectories.
  • Generative model for appearance The appearance is described with a -dimensional vector of colour features. It is assumed that the underlying color (or generally, any other intrinsic property) of a person remains constant. However, because of the varying illumination and viewing angle (pose) , the reported features O k will differ every time the person is observed. The effects of such variations are treated as observation noise.
  • a person is modelled as a colour process and O k is assumed to be a sample from a Normal distribution specific to the observed person. For every observed O k , a latent kernel is introduced from which O k is sampled.
  • the dxl vector m describes the expected, person's specific features.
  • the dxd covariance matrix V models the sensitivity of the person's appearance to the changing illumination conditions or pose. For instance, the appearance of somebody dressed uniformly in black is relatively independent of illumination or pose, so his/her covariance V has small eigenvalues. This is in opposition to a person dressed in white or non-uniform colours.
  • the observed colour features differ significantly under different colour of the illuminating light or different pose, so such a person is modelled by a "broad" kernel, i.e., V with large eigenvalues.
  • the model implies that the noise structure depends exclusively on the person and not on the local camera environment .
  • the observed features O k are independent of the association with other observations.
  • ⁇ m, V ⁇
  • ⁇ m, V ⁇
  • N Normal
  • IW Inverse Wishart
  • the IW model is a multivariate generalization of the Inverse Gamma distribution.
  • ⁇ 0 is a vector of hyperparameters defining the prior density. Appendix II-A provides details of this model.
  • D k of a detected person The spatio-temporal features D k of a detected person are referred to as dynamics . Only such dynamics that can be assumed noise- free are selected: the camera location, time of observation and the borders of entering and leaving an FOV.
  • a sequence of dynamics ⁇ D (n> , D 2 (n> , ... ⁇ assigned to nth person (denoted by the superscript) defines a path in the building.
  • the sequence [D ⁇ (n> , D 2 ⁇ n) , ... ⁇ , is modelled as a random, first order Markov chain.
  • the path is started by sampling from an initial distribution P ⁇ 0 , and extended by sampling from a transition distribution P ⁇ (D k+ ⁇ ⁇ > ⁇ D k (n) ) .
  • the assumption of first order Markov transitions implies that to generate dynamics D , only knowledge about the dynamics of the last observation assigned to a trajectory which D k extends, is needed.
  • the distributions P ⁇ and P ⁇ o follow from the topology of FOVs . For example, consider Fig. 1 and assume that D k includes only the camera location.
  • P ⁇ 0 (c) gives probabilities of starting a trajectory at a location c ⁇ ⁇ "A” , "B” , “ C” ⁇ ;
  • P ⁇ 0 ( "A” ) P ⁇ 0 ⁇ " C” ) » P ⁇ 0 ( "B” ) for it is very unlikely to appear first at “B” .
  • the quantity P ⁇ ( c ⁇ s) gives probabilities of appearing at a location c when starting from a location s. It can be assumed that P ⁇ ("A"
  • "A”) P ⁇ ("B”
  • Generative model for a sequence of observations The generative models for features O k , D k become components of the generative model for a sequence of observations ⁇ Yi, ⁇ 2f - ⁇ - Another component are the association variables that define trajectories.
  • the model is organized into slices that are indexed by corresponding observations.
  • the label S k is accompanied by a set of auxiliary variables: a counter C k , and pointers Z k ) , ..., Zk k) .
  • the counter, C ⁇ e ⁇ l , ..., k ⁇ indicates the number of trajectories present in the data Y ⁇ -. .
  • the nth pointer variable, Z ⁇ (n) e ⁇ 0 , ..., k - 1 ⁇ denotes the slice when the rath person was last observed before slice k.
  • the first step is to decide upon which person will be observed at the slice k. Recall that C k - ⁇ gives the number of different people observed so far. People enter FOVs irregu- larly, so a uniform choice is made between observing one of the known C k - ⁇ persons or introducing a new one;
  • n l , ..., k- l .
  • the pointers summarize associations before slice k, so they are updated using Sk-i, as in (4) -(5). A person labelled as k cannot be observed before slice k, so the pointer to his/her previous observation Z k k) is set to zero.
  • the pointer to the past observation of the nth, n ⁇ k, person either does not change or is set to the index of the preceding observation, k - 1, only if the label of this observation was
  • a random variable is represented as a node in an acyclic directed graph, and a causal relation as an edge.
  • Table I Example of a generated sequence with five observations of three persons is shown. In this toy example, the appearance is modelled with one dimensional features O k .
  • the features O are generated by sampling from one dimensional kernels.
  • the features D k include time and location, that correspond to Figs. 1-7. s k 1 2 2 1 3 c k 1 2 2 2 3
  • Fig. 9 shows that the discrete variables H k evolve as a Hidden Markov Model with a transition distribution following from (2)- (5) .
  • the auxiliary variables evolve deterministically when conditioned on the labels.
  • the evolutions of parameters X are a Dirichlet process (DP) mixture model.
  • DP models are applied for Bayesian inference of mixture distributions. Under such a model, the prior for parameters of a new component density can be represented as a mixture of a global prior and Delta densities centred around parameters of other components.
  • equation (6) when we integrate over variable Z k k) e [0 , ...
  • the evolutions of parameters ⁇ can be also interpreted as a mixed-memory Markov process [19] , [23] .
  • the transition distribution follows from (6)-(7).
  • the variables •Xi s j t -i become a "memory", since they have to be stored for future extensions of the graph.
  • the node H k becomes a discrete "switch" that selects one X from the past X 1:k - ⁇ .
  • the parameters X k are generated using only this selected variable (in our case simply copied) or sampled from the prior.
  • On-line tracking In the previous section a probabilistic model is defined as a procedure that in a single iteration generates a new observation on the basis of past observations. This model is applied to develop a procedure that in a single iteration finds the posterior distribution on the label of an incoming observation given the current and the past observations. When an observation Y arrives, the necessary hidden variables, including the label are estimated. First a predic- tive distribution is computed on the variables given the past data Yit k -i - Next, the predictive distribution is updated with Y , yielding a posterior distribution. From the posterior the most likely label S k is taken, which resolves the association of Y k . Such prediction-update steps are typical to filtering in probabilistic models (e.g. a Kalman filter).
  • ⁇ k, Hk is ⁇ i-.k-i , Hk-i ⁇ . It includes every past memory node ⁇ i.-k-i, thus to compute pre- dictive distribution at slice K, the term P ⁇ -.k- ⁇ , Hk- ⁇ ⁇ Yi k- ⁇ ) is needed.
  • P Y lxk the joint distribution after processing the kth observation Y k the joint distribution P Y lxk ) has to be found, because the variables ⁇ i:k, H k ⁇ will be the parent set for the next slice hidden nodes .
  • the main tracking objective is finding the distribution after receiving every new datum Y k (i.e., when k increases) . This distribution resolves the association of the current Y k and provides information necessary for associating future observations.
  • the joint predictive distribution is computed:
  • the predictive distribution is updated with the new observation (Bayes' rule) :
  • D 1 :k - ⁇ is not written explicitly, however always assumed.
  • the procedure of (10) -(11) is referred to as filtering, and the quantity as the filtered distribution.
  • a node A is a parent of node B if there is an arc going directly from A to B.
  • the parents of ⁇ k , X k ⁇ are found from the graphical model of Figs . 8 and 9. resentation will be a mixture with k (k + 1) pdfs that were used at slice k . Clearly, iterating this procedure yields a distribution that has a form of a mixture with 0(k!) components, making the exact computation intractable.
  • the intractability of the filtered density is typical to other switching state-space models, e.g. Switching Kalman Filters. There exist a family of approximate inference methods applicable to such models (see [19]).
  • ADF assumed-density filtering
  • Approximate on-line filtering 1 Representation: The joint filtered distribution is approximated by the following factorial family P (XwH k ⁇ Y ⁇ ) " Q k ⁇ S k ' C k )flQ k (X i ) Q k (Z k ⁇ i) ) r (12)
  • Q k denotes appropriate marginal distribution.
  • Q k denotes appropriate marginal distribution.
  • Factorizing the joint distribution of discrete nodes H ⁇ S k , C k , Z k x) ,..., Z k) ⁇ with a product of univariate and bivariate models sidesteps maintaining a large table with probabilities for every combi- nation of their states.
  • the nodes ⁇ S , C k ⁇ are directly coupled, so a bivariate model is chosen for a closer approximation.
  • every continuous node is represented separately, Q k ( ⁇ i) , although these are coupled with other variables.
  • a variable X jointly represents the mean and covariance of a Gaussian kernel.
  • the "Normal-Inverse Wishart" family is used to approximate the marginal
  • the expression (14) does not admit the assumed factorial representation (12) .
  • the nearest in the KL-sense factorial distribution is the product of marginals [7] , so ADF recovers the assumed family by computing the marginals of (14) .
  • ADF recovers the assumed family by computing the marginals of (14) .
  • only approximate marginals can be found because the previous filtering steps were already approximate.
  • the marginals are computed efficiently for two reasons.
  • weights Wj are probabilities of assigning Y k to Yj regardless of the counter and the label.
  • Mixture component ⁇ (X k ⁇ j ) is a Bayesian posterior obtained when X was a copy of Xj , and ⁇ (X k ⁇ °) - in case when X k was sampled from the prior.
  • the MMatch operation finds such unimo- dal pdf that preserves the moments of the mixture.
  • the first mixture component is the posterior, if the variable X k was a copy of ⁇ j .
  • the posterior is the same as in (19) .
  • the second mixture component corresponds to the case when ⁇ k was not a copy of X j .
  • variables Xj and Q k are independent, thus the posterior ⁇ (X j ⁇ j ⁇ k ⁇ does not change .
  • Limiting memory usage The generative model assumes up to k different individuals within k observations. Under our factorial approximation, representing pdf on association variables requires memory of size Oik 2 ) .
  • the prediction P S k ,C k ,Z k X j) weights the contribution (denoted ⁇ j ) of the jth memory element to the posterior on indicator S and counter C k (15) .
  • the prediction when summed over possible counters and indicators, measures the practical value of the jth memory element for resolving associations. This criterion selects an observation that is the least likely to be an end- point of any trajectory. After pruning, the contribution ⁇ j of the removed element cannot be calculated, so it has to be as-od that ⁇ j is weighted by zero in future filtering steps. A consistent way is by setting the posterior
  • the person's image was split into three fixed regions as in Fig. 12.
  • the regions are our heuristic for describing people.
  • For each region the 3D average colour was computed, obtaining in total a 9D feature.
  • Unlike colour histograms, such features provide a low dimensional summary of colour content and its geometrical layout.
  • the method requires a prior distribution ⁇ (X) on means and covariances of kernels ⁇ that are assumed to generate the appearance features.
  • This distribution ⁇ ⁇ ) is a "Normal- Inverse Wishart" density.
  • E, Q took the values: "left", "right", “other”.
  • the model P ⁇ of Markovian paths (transitions) is defined as p(L n ,E n
  • the first quantity is the probability of moving directly from location Lp when quitting the FOV via border Q p to location L n and entering the FOV via border E n .
  • the possible transitions in the environment are sketched in Fig. 10C.
  • a single filtering step deals with a single observation 50c, thus each column in the table presents filtered marginal distribution, where the probability is indicated in grey-scale (darker - higher) .
  • the true labels of objects are given as dots in the appropriate cells.
  • the MCMC approach [20] finds trajectories by sampling partitions of the complete set of observations. A single partition defines trajectories of all hypothetical objects. This method was tried for sampling 10 3 and 10 4 partitions. In both settings, after sampling the partition with the highest posterior probability was taken as the solution.
  • Table II The tracking accuracy of the compared methods is summarized in Table II.
  • the results for the MCMC method are shown as means from ten independent runs. It is observed that the described hybrid model found the correct number of persons and returns nearly exact trajectories. The accuracy of the trajectories estimated with MCMC or MHT method does not match with the accuracy of trajectories found by our method. Moreover, sampling always indicated 6 or 7, and MHT 8, distinct persons in the data set .
  • FIG. 14A presents the posterior distributions on labels obtained when tracking used exclusively the appearance features.
  • the average error was 40%.
  • This experiment shows that the low-dimensional appearance features have little dis- criminative power.
  • Fig. 14B shows the posterior distributions on labels obtained when tracking was based exclusively the dynamic features. In this case the average error was 38%.
  • the experiment shows, that the dynamics and the assumed model for Markovian transitions P ⁇ do not uniquely determine the correct trajectories. However, as seen from the previous experiments, the dynamics jointly with appearance features lead to accurate estimation of trajectories.
  • M 10 Gaussian kernels representing the intrinsic col- our properties ( is the assumed practical limit; in principle there should be 70 kernels) .
  • Every kernel Xj is represented by a posterior distribution over its parameters ⁇ (X. ⁇ ⁇ j 70 ) conditioned on all 70 observations, because filtering updates distributions on all kernels in memory.
  • ⁇ ,7o tne expected mean E [ ⁇ XJ] and the expected covariance ⁇ [Vj] of the jth kernel are found;
  • E [m- / ] a- //70 and
  • E [V- / ] C j , 70 / (n- /70 - d-1) .
  • a 2 dimensional PCA projection of the original data input is found and applied it to the ex- pected kernels. Figs.
  • FIG. 15A and 15B present the projections of the observed appearance features (as points) and the expected kernels (as ellipses) . Note, that both plots present 10 kernels; some kernels overlap. Figs . 15A and 15B confirm that the appearance features alone do not distinguish between the persons. The ten kernels inferred using only appearance similarities (Fig. 14B) form only two clusters. In Fig. 15A it is seen that when the dynamic features supported tracking, then the estimated kernels closely correspond to latent appearance features of the tracked persons. For the considered tracking application, the low-dimensional appearance features jointly with the spatio- temporal features, allow for accurate data association.
  • a technique for tracking multiple objects with non- overlapping cameras is described.
  • the technique is based on the idea to identify each object with an unique label and to define a probabilistic dependency between the labels and ob- servations of objects.
  • trajectories can be recovered by finding posterior probability distributions on the labels conditioned on the observations.
  • the posterior distribution cannot be computed exactly.
  • the essential feature of the approach is that by casting the tracking task into a probabilistic inference problem, a variety of approximate inference methods can be applied.
  • An assumed-density filtering (ADF) approximation for computing the posterior probability distributions has been applied in the present model. These distributions take the form of mixtures with an intractable number of component densities.
  • Each component arises from a combination of labels, i.e., it corresponds to a hypothetical trajectory.
  • ADF assumes a tractable family, and approximates the mixture with such a function from the family that is the closest in the Kullback- Leibler sense [4] .
  • the ADF approximation does not discard selected components, but replaces the entire mixture with a new density, in such way that the moments of the original density are preserved.
  • the experiments showed superior performance of the ADF approxima- tion compared with the MHT method.
  • the deterministic ADF approximation performed superior to a stochastic method (Markov chain Monte Carlo, MCMC) .
  • the experiments also confirmed that the accuracy of the MCMC approximation improves with increasing computational resources. However, the increased accuracy of the MCMC approximation comes at a computational cost that might be prohibitive for on-line applications.
  • the proposed model assumes Gaussian noise in the ap- pearance features of objects. A Bayesian estimation of mean and covariance of a Gaussian density under data association uncertainty is used. These parameters are considered as latent variables jointly distributed with a ⁇ Normal-Inverse Wishart' density. It has been shown how such a density can be used within an ADF algorithm. Our formulation avoids iterative search for maximum-likelihood estimates of the mean and co- variance (e.g. using Expectation-Maximization, as in [20]).
  • the EP [18] approximates an intractable posterior density with a tractable family, and iteratively improves the approximation using all available data (as opposed to the current ADF method, which performs only one iteration) .
  • the EP has been already shown successful for a "fixed-lag" improvement of posterior distributions in time-series models [21] .
  • Another line of future research involves application of sequential Monte Carlo inference (particle filter) to the proposed model. Particle filters immediately lead to on-line algorithms.
  • approximating an intractable posterior distributions with samples rather than functions allows to apply more complicated noise models for appearance features.
  • the presented Gaussian model already performed well, more complicated models might be necessary in environments with more difficult noise conditions.
  • Background estimation According to another aspect of the present invention, a probabilistic model for robust estimation of features of foreground objects is described, and it is demonstrated how spatial and temporal correlations can be taken into account in a flexible way, without much additional computational cost.
  • the model is generative in nature, and reflects our a-priori assumptions about the visual scene (such as smooth object/shadow regions and a quasi-static background process) . Exact computation of required quantities is intractable and we derive an approxi- mate Expectation Propagation (EP) algorithm that works essentially in linear time in the number of pixels, regardless of the correlation structure among them.
  • EP Expectation Propagation
  • An attractive feature of the model is that it allows for learning important image re- gions or object shapes easily.
  • background estimation can be readily coupled in a transparent and consistent way to subsequent processing.
  • o(t) is a binary variable that indicates a regime switch and [text] denotes an indicator, that evaluates to 1 (or 0) whenever the proposition "text" is true (or false) .
  • o(t) onset
  • we switch to a new regime by "reinitializing" the state vector from the prior.
  • the pixels of the background image are assumed to be quasi-static. However, due to camera jitter and long term deviation in the illumination conditions, the pixel values are not exactly constant (See Figs. 16A-16C) .
  • the background pixel values denoted as b (k, t) , are assumed to be generated from the following linear dynamical system:
  • ⁇ (t) is a parameter vector.
  • 0 ⁇ p ⁇ 1 is a fixed parameter modelling the amount of illumination drop due to shadows.
  • the graphical model is shown in Figure 18.
  • EP is an iterative message-passing algorithm and generalizes Loopy Belief Propagation (LBP) [28] , in that it is directly applica- ble to arbitrary hybrid graphical models, including the model we have introduced.
  • LBP Loopy Belief Propagation
  • factors potentials representing local dependencies
  • beliefs marginal potentials
  • EP replaces a belief with a potential from an approximating exponential family, usually in KL sense, i.e. by match- ing moments.
  • the quality of the approximation depends, how well these "summary statistics" represent the exact beliefs [32] .
  • the message update equations are derived as particular cases of the general scheme in [32] .
  • the approximation evaluates to four possible families: Gaussian, multinomial, mixture of Gaussians or mixture of clipped Gaussians. For the first two, the approximation is exact. For mixture of Gaussians and clipped Gaussians, moment matching equations can be derived easily.
  • Inference in this submodel is equivalent to calculating the integral of the Gaussian distribution Jds r p ix r ⁇ s r ) p i s r ) in one of the 2 K orthants specified by the mask configuration.
  • the sparse structure of the model is exploited to show that the multiple summations and multiple integrals are reduced to a sum w.r.t. a single Z ⁇ and an integral w.r.t. single X k .
  • This distribution can be expressed as Q k _ 1 ⁇ i )P x (S k ,C k ,Z k t - ) , because the filtered distribution at the slice k - l was approximated with a factorial family (see (10)
  • the integration variable x ⁇ m,V ⁇ denotes jointly mean and covariance.
  • the observation model is NiO k ; ,V) .
  • j > 0: ⁇ j P ⁇ iD k ⁇ Dj) TiO k ;aj, k - ⁇ , (I+XJ. ⁇ -I) Cj, ⁇ _ ⁇ , l+r/j.jt-i) , (25)
  • T denotes a multivariate T-distribution (see Appendix II-C) .
  • P ⁇ is replaced with P ⁇ 0 and ⁇ j lk - ⁇ with ⁇ 0 .
  • the normalization term L k is recovered by first finding an unnormalized table ⁇ , and then applying the normalization constraint in the standard way.
  • the mixing weights are the same.
  • This marginal is a mixture of two N-IW pdfs.
  • the second is the previous density on this node, represented by ⁇ j ;k - ⁇ .
  • the mixing weights are W and 1 - Wj as defined in (31) .
  • ⁇ j,k)- The closest density is found in the family to the mixture; ⁇ j ⁇ k MMatch i ⁇ j , ⁇ j ⁇ k - ⁇ ,Wj,l - Wj) .
  • Joint density can conveniently be represented on these variables [10] via a product P(m
  • the vector a is the marginal expected value of m.
  • the parameter K is a scaling term.
  • the scalar parameter ⁇ and the matrix C generalize the 'scale' and shape' parameters of the Gamma densities; C together with ⁇ gives the marginal expected for V: C/( ⁇ - d - 1 ) .
  • 1) Prior: Theoretically, the model of (34) does not allow to represent a noninformative distribution. Practically, such parameters can be set ⁇ 0 so that the resulting prior density IT (X) ⁇ ( ⁇ ⁇ 0 ) is sufficiently vague to express the practical ignorance about and V.
  • the IW model is a multivariate generalization of the Inverse Gamma distribution.
  • d matrix X is Inverse Wishart distributed with parameters: scalar ⁇ and matrix C, if the density of X is
  • a random d-dimensional vector x is T distributed with parameters: scalar ⁇ , vector and matrix V, if the density of x is

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Software Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Probability & Statistics with Applications (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Image Analysis (AREA)

Abstract

The present invention provides a method for tracking of multiple objects or individuals with non-overlapping or partly overlapping cameras, wherein each object is identified with a label and wherein a probabilistic dependency between the labels and the observations of objects is defined wherein the posterior probability distributions in the model are in the form of mixtures, wherein the mixture is approximated.

Description

A HYBRID GRAPHICAL MODEL FOR ON-LINE MULTICAMERA TRACKING
Tracking objects (e.g. people, cars) in wide areas such as airports or shopping centres becomes increasingly important for automated surveillance and security applications. Such places are usually equipped with closed-circuit TV (CCTV) systems. These systems comprise a network of sparsely distributed cameras, where a camera provides only information about local appearances of objects within its field of view (FOV) . Wide-area tracking aims to reconstruct global trajectories of objects given their local appearances. Typically, the number of objects is large and viewing fields of cameras do not cover the entire region of interest: Therefore, methods developed for visual tracking within a single-scene are not immediately applicable to problems with sparsely distributed cameras. In single-scene-single-camera [12] , [14] , [17] or single-scene-multi- camera [5] , [16] , [24] scenarios, tracking takes place at a frame-to-frame level. In this case, smooth motion provides es- sential cues for resolving association ambiguities. Distributed, multi-camera tracking takes place at a camera-to-camera level. Observations (e.g. local appearance features) arrive one at a time, within irregular time intervals as objects move from one camera location to another. The time gap between successive observations of the same object may be very long. This hinders the use of smooth motion cues that facilitate tracking in a single-scene setting. Therefore, tracking with sparse cameras relies on appearance similarities and movement constraints imposed by the topology of camera lo- cations. A main problem in multi-object tracking is the association of the incoming sensory observations with trajectories. Typical sensors, such as cameras, cannot directly observe the identity of an object. Due to this inherent observa- tion ambiguity, the number of possible associations grows rapidly with the number of objects and with the number of observations. This explosion renders exact calculation of interesting quantities such as the most likely association intractable [2] and one has to reside to approximate techniques. A popular approximate technique for finding likely associations is multiple hypotheses tracking (MHT) [2] . In the context of tracking with non-overlapping cameras, an MHT algorithm was applied in a campus monitoring system [6] . MHT is an on-line, deterministic algorithm that maintains a list of high probability associations and prunes trajectories with low probability. It is usually a fast and easy to implement method. However, it is a heuristic approximation method with limited accuracy. Moreover, pruning does not allow to preserve any information about the discarded trajectories. In noisy environments, such information could be necessary for proper association of future observations. Alternatively, data association can be solved using approaches with a sound theoretical basis, like stochastic ap- proximation methods: Markov Chain Monte Carlo (MCMC) [11] or sequential Monte Carlo (SMC, also known as particle filters) [8] . An MCMC method was applied for tracking vehicles with multiple cameras on a motorway network [20] . Stochastic methods have the attractive theoretical property that with in- creased computational resources, on average, they are guaranteed to give better solutions. However, in practice, the promised accuracy is obtained at a very high computational cost. Furthermore, foreground/background classification of pixels is an important preprocessing step in many computer vision problems, such as object tracking and identification where one wishes to suppress background pixels to collect reliable statistics about object features of interest. Clearly, a practical method has to be robust since all the subsequent inference depends on it, but it needs to be also efficient to meet eventual real-time operation requirements. In typical visual scenes populated with several distinct objects, it may not be clear which object (s) should be assigned to the foreground. The key assumption here is that the visual sensors and the background image are static and objects of interest are in motion (e.g. people walking in a room full of furniture) . Transient deviations from "steady-state" pixels values are detected and associated with foreground ob- jects. Unfortunately, in reality, a background scene is almost never entirely static due to illumination changes, shadows, reflections, fog, wind or camera jitter [35] . All these factors render simple heuristics such as thresholding pixel differences useless. From a statistical viewpoint, background estimation can be viewed as a novelty detection problem, e.g. [27]. Either implicitly or explicitly, probabilistic approaches attempt to estimate a process for the background pixels and try to detect "surprises". However, mainly due to computational considera- tions, many methods (e.g. [27, 34]) ignore spatial dependencies among neighbouring pixels and only take temporal correlations into account. More recently, spatial correlations among pixels are implicitly taken into account by using a fixed basis transform [35] or are explicitly modelled using Markov random fields [37, 33] .
Summary of invention According to a first aspect, the present invention provides a method for tracking of multiple objects with non- overlapping or partly overlapping cameras. Furthermore, the present invention provides a system comprising one or more processing units, two or more non-overlapping cameras operatively connected to the processing units and software stored in memory to execute the above mentioned methods. The preferred embodiments are dedicated to non- overlapping cameras, although applications for partly overlapping cameras are very well conceivable, although not yet realised in experience. The above method provides for tracking of multiple objects with no -overlapping cameras, wherein each object is identified with a label and wherein a probabilistic dependency between the labels and observations of objects is defined, wherein the posterior probability distributions in the model take the form of mixtures, wherein the mixture is approximated. The above method preferably assumes that an observation of an object includes appearance features defined as three colour vectors, wherein each colour vector is the average colour of pixels within one of three horizontal regions selected from the object's image. The method preferably assumes that the appearance features from all observations of a single object are generated by a single Gaussian distribution. The method preferably estimates the parameters of object's specific Gaussian model (mean, covariance matrix) with a Bayesian inference scheme, where the mean and the covariance are given a joint "Normal-Inverse Wishart" prior. Such an approach avoids iterative search for maximum-likelihood estimates of these parameters . The method preferably assumes that appearance features of all objects are generated by a Dirichlet process mixture model, which is used for estimation of the number of distinct objects from the data. The Dirichlet process mixture model is preferably extended with auxiliary variables that allow to model Markovian dependencies between spatio-temporal observa- tions of the same object. The method preferably estimates trajectories of objects by computing posterior distribution of the labels conditioned on the available observations, wherein an assumed- density filtering (ADF) algorithm is used for online computa- tion of filtered distributions of the model's latent variables . According to a further aspect, the present invention provides a method for background/foreground classification of pixels in a visual scene, wherein the class of each pixel is identified with a binary mask variable, and wherein a probabilistic dependency between the observed pixels and their masks is defined, wherein long-range correlations between masks are a-priori imposed. The method according to this further aspect imposes long-range correlations between pixel binary masks by assuming that each binary mask is obtained by taking the sign of a corresponding continuous variable. The correlations are a-priori imposed by defining a joint prior Gaussian distribution on all continuous variables. The method preferably maintains the mask's long-range correlations across consecutive frames of a video sequence by assuming the continuous variables evolve between frames according to linear Gaussian process. The method preferably estimates the masks by computing posterior distributions of binary variables conditioned on the observed pixel values, wherein the posterior distribution are approximated with an Expectation-Propagation algorithm.
Brief description of the drawings
Fig. 1 shows an example of the considered tracking problem. In this example, there are three non-overlapping cameras: "A", "B" and " C" that observe local scenes in a building. Figs. 2 - 7 depict a complete pass (event) of a person through a camera viewing field. Six observations (events) are being shown corresponding to three different persons. The goal of tracking is to reconstruct global trajectories between the cameras .
Fig. 8 shows a graphical representation of the generative model . The continuous random variables are shown as ovals, the discrete - as rectangles. For clarity the arcs be- tween observations that mark the dependencies induced by the dynamics Dk were not drawn.
Fig. 9 shows a compact graphical representation of the ■generative model . This representation captures the essential structure of the model . A single node Hk represents now all association-related variables { Sk, Ck, Zk {1) , ... , Zk {k) } at slice k. This representation shows the relationship with Dirichlet process mixture models and mixed-memory Markov models. Figs. 10A and B show a building plan were observations for experiments were taken. The grey areas show camera viewing fields. Fig. IOC shows the movement constraints assumed by the distributions Pδ and Pδo - Arrows indicate possible transitions. All other transitions have zero probability.
Figs. 11A and B show results of tracking with the hybrid model. The columns correspond to filtering steps; k = 1,...,70. Each column in Fig. 11A gives a filtered distribution of label; Qk (Sk) • The true labels are indicated as dots. Fig. 11B shows a filtered distribution of counter; ' Qk ( Ck) in each column. Fig. 12 shows appearance features of a person. The appearance is described with three colour vectors. Each colour vector is the average colour of pixels within one of the three horizontal regions R1 , R2 , R3 selected in the image. The height of a region is 25% of the persons' height. The "skip" areas are meant to remove parts that have little discriminative power. Their height is fixed at 12.5% of the total height.
Fig. 13 shows an example of a recovered trajectory. The sequence of images represents local appearances at the indicated cameras. The trajectory includes such observations Y that have labels S± - 5, as shown in Fig. 11A.
Fig. 14A shows results of tracking using only the ap- pearance features . Fig. 14B shows results of tracking using only the dynamic features. In both cases, columns show the filtered distribution of labels Q ( Sk) ι £or k = 1,...70. The true labels are indi- cated as dots.
Fig. 15A shows the expected kernels (expected mean and expected covariance) shown as a 2D PCA projection of the original 9D feature space, obtained when tracking was based on dynamics and appearance features . Fig. 15B shows the expected kernels (expected mean and expected covariance) shown as a 2D PCA projection of the original 9D feature space, the kernels obtained when tracking was based only on appearance features .
Figs. 16A-16C show a few video frames from an office scene. Fig. 16A shows a vertical slice through time, taken along the vertical line shown in Fig. 16B on the video frames and variations in RGB values of a single pixel (at row 200) in Fig. 16C. Occlusions and shadows can be clearly seen as immediate changes in intensity. Fluctuations in a single color channel illustrate the difficulty of detection when pixels are processed independently.
Figs. 17A-17D show random draws from correlated Gaus- sians on 256 dimensions arranged as a 16 x 16 grey scale image at the top row. The correlation matrix of these Gaussians is chosen such that the correlation between two pixels values is a function of spatial distance between them. The left column corresponds to strong correlation and right column to weak correlation. The bottom row shows the signs of greyscale images as a binary image. By adjusting the correlation matrix of the latent continuous variables x, we can generate different distributions over possible masks r.
Fig. 18A shows a graphical model of a single time frame. The rectangles are plates that denote K repetitions of nodes inside. Square and oval nodes correspond to discrete and continuous variables, respectively. Variables f, b and y are vectors denoting the RGB components. Dotted arcs correspond to the regime switch mechanism that is triggered when an object leaves the scene . Fig. 18B shows a loopy factor graph of a single slice used in EP iterations. Forward messages to the next time slice (not shown) are passed only once.
Fig. 19 shows comparison of EP with importance sa - pling, for K = 8. We use a truncated Fourier basis Wr (k) = [1, sin(ωT), cos ( ωk) ] with ω = 2πK with Rr = 0.12T, mr = 0. Note that many configurations have the same probability since mr=0 . For importance sampling we have drawn 107 samples independently from sr and integrate over xr to compute p(r) . Fig. 20A shows typical data generated from the model. Fig. 2 OB shows estimated means of colour features μ { t) . Fig. 20C shows filtered estimates of object masks. Fig. 20D shows standard deviations of colour features μ(t) .
Fig. 21 shows colour features estimated from the video sequence in Fig. 16B with superimposed filtered estimate of asks at the top and colour histograms in HSV space at the bottom.
Overview Tracked people are observed in a building that is equipped with a network of non-overlapping (or partly overlapping) cameras A, B, C, e.g. as in Fig. 1. A camera can track a person through its FOV. As soon as the person leaves the viewing field, the camera reports a summarized description of the person. Such a description is our primarily input, so it is referred to as an observation. Observations from all cameras are processed centrally.
Input Assumptions Let Yk, k = 1,2,..., denote the th in time order observa- tion from any camera. An observation includes two components; Yk = { k, Dk} . The symbol Ok represents appearance, e.g. color features computed while a person was visible within an FOV. The symbol Dk represents spatio-temporal features of the observation, e.g. the location and the time of observation. An important feature of the approach is not to assume a fixed number of tracked objects beforehand. Given a sequence of observations {Yι,Y2/—} the task is on-line association of a new observation Yk and estimation of the number of distinct individuals found in the data. Approach The idea underlying the approach is to identify each person with an unique label, and view the sequence of local appearances as noisy observations from the hidden labels. The observations and their labels are treated as random variables with a probabilistic dependency P(Events | Labels) in the form of a generative model. Given the model, tracking is simply computing the posterior probabilities of the labels condi- tioned on the observations:
P(Labels I Events) oc P(Events | Labels) P(Labels) ,
where (Labels) is the prior information about the trajecto- ries. The posterior is used to select the most likely label for every observation. Such a probabilistic strategy is attractive since the causal relationship P(Events | Labels) is more intuitive to specify than direct finding of trajectories. The proposed generative model allows to apply any in- ference method for computing the posterior probabilities given the observations. However, due to the intractability of the data association, only the approximate methods are practical [19] . In this paper, an on-line, assumed-density filtering algorithm [4] is applied. Note that, the generative model can also be a basis for a multiple hypothesis tracker; in this case the model evaluates the likelihood of trajectories.
Generative model
Generative model for appearance The appearance is described with a -dimensional vector of colour features. It is assumed that the underlying color (or generally, any other intrinsic property) of a person remains constant. However, because of the varying illumination and viewing angle (pose) , the reported features Ok will differ every time the person is observed. The effects of such variations are treated as observation noise. A person is modelled as a colour process and Ok is assumed to be a sample from a Normal distribution specific to the observed person. For every observed Ok, a latent kernel is introduced from which Ok is sampled. A kernel is a Normal density, represented with parameters ∑k = {m , Vk} . The dxl vector m describes the expected, person's specific features. The dxd covariance matrix V models the sensitivity of the person's appearance to the changing illumination conditions or pose. For instance, the appearance of somebody dressed uniformly in black is relatively independent of illumination or pose, so his/her covariance V has small eigenvalues. This is in opposition to a person dressed in white or non-uniform colours. The observed colour features differ significantly under different colour of the illuminating light or different pose, so such a person is modelled by a "broad" kernel, i.e., V with large eigenvalues. The model implies that the noise structure depends exclusively on the person and not on the local camera environment . Moreover, given ∑k, the observed features Ok are independent of the association with other observations. The parameters Σ = {m, V} are treated as hidden random variables that are given a joint prior distribution π (∑) . In Bayesian statistics [10] , a convenient joint model for the mean JΠ and covariance V of a Gaussian kernel is a product of a Normal (N) density w.r.t. m, and an Inverse Wishart ( IW) density w.r.t. V. The IW model is a multivariate generalization of the Inverse Gamma distribution. The joint N- IW density is denoted as : π (∑) = φ (x\ θ0) , ( 1 )
where θ0 is a vector of hyperparameters defining the prior density. Appendix II-A provides details of this model.
Generative model for dynamics The spatio-temporal features Dk of a detected person are referred to as dynamics . Only such dynamics that can be assumed noise- free are selected: the camera location, time of observation and the borders of entering and leaving an FOV. A sequence of dynamics {D (n> , D2 (n> , ... } assigned to nth person (denoted by the superscript) defines a path in the building. The sequence [Dι (n> , D2 {n) , ... }, is modelled as a random, first order Markov chain. The path is started by sampling from an initial distribution Pδ0, and extended by sampling from a transition distribution Pδ (Dk+ι < > \ Dk (n) ) .The assumption of first order Markov transitions implies that to generate dynamics D , only knowledge about the dynamics of the last observation assigned to a trajectory which Dk extends, is needed. The distributions Pδ and Pδo follow from the topology of FOVs . For example, consider Fig. 1 and assume that Dk includes only the camera location. In this case, Pδ0(c) gives probabilities of starting a trajectory at a location c <≡ { "A" , "B" , " C" } ; One can assume Pδ0 ( "A" ) = Pδ0 { " C" ) » Pδ0 ( "B" ) for it is very unlikely to appear first at "B" . The quantity Pδ ( c \ s) gives probabilities of appearing at a location c when starting from a location s. It can be assumed that Pδ("A"|"A") = Pδ("B"|"A") » Pδ("C" |"A") , because it is unlikely to move from "A" to "C" without being observed at "B" . Generative model for a sequence of observations The generative models for features Ok, Dk become components of the generative model for a sequence of observations {Yi,Ϊ2f-}- Another component are the association variables that define trajectories. The model is organized into slices that are indexed by corresponding observations. 1) Association variables: For every observation Yk there is a corresponding variable Sk that denotes the label of the person to which Yk is assigned. Within the first k data, Yi:κ { Yi, Y, —} i there may be at most k different people, so Sk has k different states; Sk e {l,...,J}. Given the labels, the trajectory of the nth person can be recovered by taking all observations { Yi} that satisfy Si = n . At every slice k, the label Sk is accompanied by a set of auxiliary variables: a counter Ck, and pointers Zk ) , ..., Zk k) . The counter, C^ e { l , ..., k} , indicates the number of trajectories present in the data Yχ-. . The nth pointer variable, ZΛ (n) e { 0 , ..., k - 1} , denotes the slice when the rath person was last observed before slice k. Value Zk {n) = 0 indicates that person n has not yet been observed. At the slice k there can be up to k persons, so variables Z^(π) for n = l , ..., k are needed. Auxiliary variables only summarize the past associations. For instance, the counter is the maximum label used so far; Ck = max{Sι, ..., Sk) . The auxiliary variables provide an instant look-up reference to the necessary information that is encoded by a sequence S1:k. Therefore, the transitions of the labels are first-order Markovian. 2) One-slice generation: Our generative model builds observations one-by-one. Initially, the counter indicates that there were no objects observed; C0 = 0. The first step is to decide upon which person will be observed at the slice k. Recall that Ck-ι gives the number of different people observed so far. People enter FOVs irregu- larly, so a uniform choice is made between observing one of the known Ck-χ persons or introducing a new one;
Sk ~ Uniform(l/...,Cjc.α,Cjc-ι + 1). (2)
The uniform distribution is one choice; other distributions are also applicable. Given the label, the counter and auxiliary pointers are updated deterministically Ck = Ck-i + [Sk > Cit-i] , (3) Zk m = 0 , (4) Zk inX Z^ iSk ≠ n] + (k- 1 ) [Sk = n] (5)
where n = l , ..., k- l . The symbol [f] is an indicator; [f] ≡ 1 if the binary proposition f is true, and [f] = 0 otherwise. If the label indicates a new person, Sk = Ck-ι+l , then the counter increases, as in (3) . The pointers summarize associations before slice k, so they are updated using Sk-i, as in (4) -(5). A person labelled as k cannot be observed before slice k, so the pointer to his/her previous observation Zk k) is set to zero. The pointer to the past observation of the nth, n < k, person either does not change or is set to the index of the preceding observation, k - 1, only if the label of this observation was The second step is generating the latent parameters ∑k of the person indicated by Sjt. If this person has been already observed, then the index of his last observation Zk (Sk> is nonzero. By the assumption that the latent parameters do not change, so X is simply copied from the slice Zk (Sk> . If the cur- rent person is observed for the first time, Zk (Skl = 0, then the parameters are sampled from prior π (X) . Let i = Zk k>
Xk = ∑i [i > 0] + J^ew [ i = 0] , (6) Xnew ~ π (∑) . ( 7 )
The final step is rendering the observation Yk - { θk, Dk} given the parameters of a Gaussian kernel Xk = [mk l Vk} and the pointer to the past dynamics of the current object Zk k> . The generative models of the colour features and dynamics are used. Let i = Z (sk) Ok ~ N(mk, Vk) , (8) Dk ~ Pδ (Dk \ Dι) [i>0] + Pδ0 (Dk) [i=0] (9)
Graphical representation It is useful to think about the interaction between labels, kernel parameters, and observations using the language of Dynamic Bayesian Networks (DBNs) [19] . In this formalism, a random variable is represented as a node in an acyclic directed graph, and a causal relation as an edge. Fig. 8 shows a graphical representation of the first five slices of our model. Table I shows an example of generated data. Every column denotes the variables corresponding to a single slice of the generative process. In the rest of the paper, a variable Hk = { Sk, k, k X) , —, Zk k) } is used to denote the discrete-valued association variables at the th slice. Fig. 9 shows the corresponding compact graph. Table I Example of a generated sequence with five observations of three persons is shown. In this toy example, the appearance is modelled with one dimensional features Ok . The features O are generated by sampling from one dimensional kernels. The features Dk include time and location, that correspond to Figs. 1-7. sk 1 2 2 1 3 ck 1 2 2 2 3
■ (1) ,~, W >k *k {0} { ,0} {1,2,0} {1,3,0,0} {4,3,0,0,0}
Xk={mk,Vk) {5,2} {1,0.5} {1,0.5} {5,2} {3,1} Ok 6.68 0.91 1.08 4.94 4.06 Dk 10:15am/A 10:25am/B 10:26am/A 10:42am/B 10:42am/C
Generating a new observation extends the graph with a new column of nodes that depend on the past ones. Fig. 9 shows that the discrete variables Hk evolve as a Hidden Markov Model with a transition distribution following from (2)- (5) . The auxiliary variables evolve deterministically when conditioned on the labels. It is also noted, that the evolutions of parameters X are a Dirichlet process (DP) mixture model. DP models are applied for Bayesian inference of mixture distributions. Under such a model, the prior for parameters of a new component density can be represented as a mixture of a global prior and Delta densities centred around parameters of other components. To see the relationship with the model, consider equation (6) : when we integrate over variable Zk k) e [0 , ... , k-l) then the variable ∑k will be distributed according to a density which is a mixture of the prior π{Xk) and k - 1 Delta distributions δ(∑k - ∑j) , j = 1,..., k - l. The evolutions of parameters Σ can be also interpreted as a mixed-memory Markov process [19] , [23] . The transition distribution follows from (6)-(7). The variables •Xisjt-i become a "memory", since they have to be stored for future extensions of the graph. The node Hk becomes a discrete "switch" that selects one X from the past X1:k-ι. The parameters Xk are generated using only this selected variable (in our case simply copied) or sampled from the prior. On-line tracking In the previous section a probabilistic model is defined as a procedure that in a single iteration generates a new observation on the basis of past observations. This model is applied to develop a procedure that in a single iteration finds the posterior distribution on the label of an incoming observation given the current and the past observations. When an observation Y arrives, the necessary hidden variables, including the label are estimated. First a predic- tive distribution is computed on the variables given the past data Yitk-i - Next, the predictive distribution is updated with Y , yielding a posterior distribution. From the posterior the most likely label Sk is taken, which resolves the association of Yk. Such prediction-update steps are typical to filtering in probabilistic models (e.g. a Kalman filter).
Background In Markovian models filtering exploits the fact the current hidden nodes Σk, Hk are conditionally independent from Yi-.k-i given the set of parents1 Pa. (Xk, Hk) . This allows to find predictive distribution by combining two terms: (i) the transition distribution from the parents to the current nodes; and (ii) the posterior distribution of the parents conditioned on past data Yχ.-k-ι . Typical DBNs are first-order Markovian, so the parents of the current hidden nodes are just the hidden nodes from the previous slice, and their posterior is automatically available (cf. forwards pass [19]). However, the model of Fig. 9 is non-Markovian. The set of parents l?a. {∑k, Hk) is {∑i-.k-i , Hk-i} . It includes every past memory node ∑i.-k-i, thus to compute pre- dictive distribution at slice K, the term P {∑ι-.k-ι , Hk-ι \ Yi k-ι) is needed. As a result, when processing the kth observation Yk the joint distribution P Ylxk) has to be found, because the variables {∑i:k, Hk} will be the parent set for the next slice hidden nodes . The main tracking objective is finding the distribution after receiving every new datum Yk (i.e., when k increases) . This distribution resolves the association of the current Yk and provides information necessary for associating future observations. First, the joint predictive distribution is computed:
The predictive distribution is updated with the new observation (Bayes' rule) :
where L is normalization constant. The distribution follows from (8) -(9); the dependency of 5c on past dynamics
D1 :k-ι is not written explicitly, however always assumed. The procedure of (10) -(11) is referred to as filtering, and the quantity as the filtered distribution.
Intractability of exact filtering The filtered distribution is difficult to maintain, since it involves continuous [∑i-.k) and discrete (H*)variables that are dependent on each other. For every state of Sk a separate probability density function (pdf) is needed for representing belief of Xχ.k conditioned on Sk. The predictive distribution for the next slice requires summing out Sk (see (10) , what yields a representation in the form of a mixture of k pdf's. At slice k + 2 (after summing over S +x) , the exact rep-
1 A node A is a parent of node B if there is an arc going directly from A to B. The parents of { k, Xk} are found from the graphical model of Figs . 8 and 9. resentation will be a mixture with k (k + 1) pdfs that were used at slice k . Clearly, iterating this procedure yields a distribution that has a form of a mixture with 0(k!) components, making the exact computation intractable. The intractability of the filtered density is typical to other switching state-space models, e.g. Switching Kalman Filters. There exist a family of approximate inference methods applicable to such models (see [19]). The assumed-density filtering (ADF) approach is followed for it is deterministic and suited for on-line implementations. ADF [4], [18] approximates the filtered distribution with a tractable family. After one- slice filtering (10) -(11), the filtered distribution will no longer be a member of the assumed family. ADF projects it back to such member of the family that offers the closest approxi- mation in the Kullback-Leibler (KL) [7] sense.
Approximate on-line filtering 1) Representation: The joint filtered distribution is approximated by the following factorial family P (XwHk \ Y^) " Qk<Sk' Ck)flQk(Xi ) Qk(Zk<i)) r (12)
where Qk denotes appropriate marginal distribution. Factorizing the joint distribution of discrete nodes H = { Sk, Ck, Zk x) ,..., Z k) } with a product of univariate and bivariate models sidesteps maintaining a large table with probabilities for every combi- nation of their states. The nodes { S , Ck} are directly coupled, so a bivariate model is chosen for a closer approximation. For simplicity, every continuous node is represented separately, Qk (∑i) , although these are coupled with other variables. A variable X jointly represents the mean and covariance of a Gaussian kernel. The "Normal-Inverse Wishart" family is used to approximate the marginal
Qjc(Xi) = <p(XiiJc), (13) because this family is conjugate to the Gaussian generative kernel (Appendix 11-A) . The hyperparameters θχι are specific to the posterior on ∑ι, after observing the th datum. 2) One-slice filtering: After processing observation -t-1 the ADF algorithm maintains an approximation to the filtered distribution P (∑ι:k-ι, Hk-ι \ Yi;k-ι) in the form (12). When Yk arrives, this distribution is updated. Executing (10) yields an approximate predictive density Pr (∑i:k, Hk) (∑\ :k, H \ Yi-.k-i) • Executing (11) yields
P (Σ1:k,Hk I Y1:k) - —P (Yk j ∑k,Hk) P ∑1:k k) (14)
The expression (14) does not admit the assumed factorial representation (12) . The nearest in the KL-sense factorial distribution is the product of marginals [7] , so ADF recovers the assumed family by computing the marginals of (14) . As with any ADF method, only approximate marginals can be found because the previous filtering steps were already approximate. The marginals are computed efficiently for two reasons. First, the model is sparse: as can be seen from (6)-(9), when a label Sk = m is set, then out of all auxiliary variables, only Zk(m) is referenced. If Zk {m) = j is set, only variable ∑j and dynamics Dj contribute to the likelihood of observation 50c. Thus, when the marginals of (14) are computed, the variables that do not affect the new observation are present only in the predictive distribution Pr (Hfc, L:.fc) and they marginalize to unity. Therefore, the joint distribution does not have to be found, only its necessary marginals. Second, finding the marginals of the predictive distribution is simple, because most of the hidden variables admit deterministic transitions. Algorithm 1) Computing marginals: Below the marginaliza ion results are discussed; details are provided in the Appendix I.
The marginal on the counter and switch is given by
λ 7 = p δ (D I DJ) [p (°k \xk =x) Q kJXj =*) de)
0 The term Pr and an analytical solution to the integral are given in Appendix I. For j >_ 1, the quantity λj is the likelihood of 5 when the past observation of the current person was Yj . The quantity λ0 is the likelihood of 0 when it starts a 5 new trajectory; in this case one has to replace Pδ with Pu and Qjc-ι(Xj) with n (∑) in (16). Note that, due to integration of model parameters Σ, the term λ acts as a Bayesian model selector. The "broad" prior density π leads to an integral that takes almost uniform, but small, values for a broad range of
»0 Ok. The integrals involving Q (∑j) are peaked around values similar to the past features Oj . Thus, when the observed appearance Qk is far from any of the past appearances Oj , then likelihood of introducing a new trajectory is high. When Qk is close to one of the past Oj , then introduction of a new trajec-
>5 tory is penalized by small λ0. The motion constraints, represented by Pδo [Dk) and Pδ Dj) , act similarly. Introduction of a new trajectory depends on a balance between the prediction of c by the initial distribution Pδ0, and the prediction of Dk by one of the past dynamics Dj .
30 The marginal on an ith, i = l , ..., k, auxiliary variable is given by
The summations in this expression can be computed efficiently because the pointers Zk {x) and Z {m) are deterministic when conditioned on Sk-x and the past pointers. The marginal Qk(∑k) s a mixture of k densities. A procedure MMatch (Appendix II) replaces the mixture with a single "Normal-Inverse Wishart" density as required by (13) (18) φ (Xk I &) ∞ P (Ok\ ∑k) φ <Xk. \ θjrk , (19)
where the weights Wj are probabilities of assigning Yk to Yj regardless of the counter and the label. Recall from the pre- vious discussion, that our model assumes a Dirichlet prior for variable ∑k. The predictive density for ∑k is a mixture of the global prior π{∑k) and k - l Delta distributions δ(∑k - ∑j) , j = l,...,k-l. Accordingly, the filtered density is a mixture obtained by updating each component for the predictive density. Mixture component φ(Xkj) is a Bayesian posterior obtained when X was a copy of Xj , and φ(Xk\θ°) - in case when Xk was sampled from the prior. The MMatch operation finds such unimo- dal pdf that preserves the moments of the mixture. The marginal Q (∑j),j = l,...,k - 1, is a mixture of two pdfs
Qk(∑j) = MMatch WJ φ [Xj | θj) + (l - wf) φ (Xj \ θjιk-ι) ) . (20)
The first mixture component, φi∑^Q-3), is the posterior, if the variable Xk was a copy of ∑j . In this case the posterior is the same as in (19) . The second mixture component corresponds to the case when ∑k was not a copy of Xj. In this case, variables Xj and Qk are independent, thus the posterior ψ(Xjjιkχ does not change . 2) Limiting memory usage: The generative model assumes up to k different individuals within k observations. Under our factorial approximation, representing pdf on association variables requires memory of size Oik2) . Also 0(k) dynamical components D1 k and 0(k) parameters θ1:kιk have to be stored. It is rare in the real-world that every observation corresponds to a different person. The number of simultaneously tracked persons is limited to some maximum K. At any slice k the number of variables Zk >, and the domains of the counter and indicator variables are kmax = m±n(k,K) . To further restrict the memory demand, a limit M of the number of dynam- ics D and hyperparameters Σ that are preserved in memory is set. After processing 50c, the corresponding {Dkk,k} is saved into the memory, and if k > M, such memory element {Djjιk} is pruned, 1 < j < k, that minimizes the criterion
v(j)=YYPr(Sk=n,Ck=m,Zk Xj)l (21) m=l n-1
The prediction P Sk,Ck,ZkX j) weights the contribution (denoted λj) of the jth memory element to the posterior on indicator S and counter Ck (15) . The prediction when summed over possible counters and indicators, measures the practical value of the jth memory element for resolving associations. This criterion selects an observation that is the least likely to be an end- point of any trajectory. After pruning, the contribution λj of the removed element cannot be calculated, so it has to be as- certained that λj is weighted by zero in future filtering steps. A consistent way is by setting the posterior
Qk(Zk nXj) = 0 and renor alizing, for all n = l,...,Jmax. The deter- ministic transitions of auxiliary variables in the model guarantee that the prediction for Zτ lnl = j will be zero for τ > k. 3) Estimating label and counter: After one-slice filtering, the distribution Qk { Sk, Ck) provides the solution to the 5 tracking task. The observation 10 is given the most likely label according to the marginal on indicator; 2k=argmax mQ (Sk=m) . The marginal on the counter, Q ( Ck) , estimates the number of distinct individuals observed within data Yi-.k-
L0 Experiments The model of the invention was tested by tracking people observed in a building. In the first experiment the tracking accuracy of the approach was measured and compared with the Markov chain Monte Carlo (MCMC) and MHT methods. In the
15 second experiment the accuracy was compared in two cases: (1) when tracking was based exclusively on the appearance, and (2) when tracking was based exclusively on dynamics. This experiment evaluates the impact of appearance features on the overall tracking accuracy. 0 Setup 70 observations were recorded, Yi:7o, of 5 persons who were observed in an office-like environment with 7 non- overlapping cameras. Figs. 10A and 10B show the topology of 5 camera locations and their viewing fields. 1) Colour features: From a video sequence with a pass of a person through a FOV a single frame was selected. In this frame the pixels representing the person were manually found. The original RGB space was transformed into a colour-channel 0 normalized space [9] to suppress the effects introduced by the colour of the illuminating light. The person's image was split into three fixed regions as in Fig. 12. The regions are our heuristic for describing people. For each region the 3D average colour was computed, obtaining in total a 9D feature. Unlike colour histograms, such features provide a low dimensional summary of colour content and its geometrical layout. The method requires a prior distribution π (X) on means and covariances of kernels Σ that are assumed to generate the appearance features. This distribution π {∑) is a "Normal- Inverse Wishart" density. Its parameters θ0 = {a0, κ0, ^o,C0} were set as follows: the expected features a0 = 09xι (9 dimensional zero vector); the scale κ0 = 100; the degrees of freedom η0 = 9 (dimensionality of observations) , the matrix C0 = 10"3I9x9, where I9x9 is a 9 dimensional identity matrix. Parameter C0 with small eigenvalues defines relatively "sharp" kernels. Since the means m are not known, the scale κ0 is set to a large value . 2) Dynamic features: A pass of a person through a FOV is described with dynamic features D = {L, E, Q, T) . The term L is the camera location; L e {1,2,3,4,5,6,7}; T is the time of observation; variable B {Q) denotes the frame border through which the object enters (leaves) the camera FOV. Variables E, Q took the values: "left", "right", "other". The model Pδ of Markovian paths (transitions) is defined as p(Ln,En| Lp, p)p(Tn| Tp) . The first quantity is the probability of moving directly from location Lp when quitting the FOV via border Qp to location Ln and entering the FOV via border En . The possible transitions in the environment are sketched in Fig. 10C. All transitions marked with an arrow have equal nonzero probability; all other have zero probability. The term p(Tn|Tp) is the probability of being observed at time Tn if the previous observation was at time Tp regardless of the locations. A binary model p(Tn|ϊp) is used that prevents a person appearing at different locations at the same time. The distribution Pδ0 {L, E) gives the likelihood of starting a path at FOV L via border E. In this case it was only possible via left border of FOV 1. Experiment 1 Given the videos, the true trajectories were manually recovered and marked with true labels. The method was evaluated by comparing the estimated trajectories with the true ones. The tracking accuracy is defined as follows. Given an estimated trajectory the true labels of observations are counted, and it is assumed that this trajectory describes the person with the most frequent label . The observations with other labels within this trajectory are considered as wrongly classified. The ratio of wrongly classified to all observations within the trajectory makes the classification error. Tracking accuracy is measured by the average of such errors over all estimated trajectories. An additional criterion is the number of estimated trajectories. The tracker is run with the maximum number of objects K - 10 and memory limit M = 10. Fig. 11A shows the marginal distributions on labels Qk (Sk) and Fig. 11B shows counter Qk ( Ck) after filtering subsequent observations, k = 1,...,70. A single filtering step deals with a single observation 50c, thus each column in the table presents filtered marginal distribution, where the probability is indicated in grey-scale (darker - higher) . The true labels of objects are given as dots in the appropriate cells. An example of a recovered trajectory is shown in Fig. 13 as a sequence of local appearances Yi with a common label Si = 5. The MCMC approach [20] finds trajectories by sampling partitions of the complete set of observations. A single partition defines trajectories of all hypothetical objects. This method was tried for sampling 103 and 104 partitions. In both settings, after sampling the partition with the highest posterior probability was taken as the solution. The tracking accuracy of the compared methods is summarized in Table II. The results for the MCMC method are shown as means from ten independent runs. It is observed that the described hybrid model found the correct number of persons and returns nearly exact trajectories. The accuracy of the trajectories estimated with MCMC or MHT method does not match with the accuracy of trajectories found by our method. Moreover, sampling always indicated 6 or 7, and MHT 8, distinct persons in the data set .
Table II Classification results for the proposed model versus the MHT and MCMC methods.
Hybrid model MHT MCMC MCMC with 103 samples With 104 samples average error 5% 21% 20% + 0.6 12% ± 0.5 number of 6.5 + 0.3 6.5 ± 0.3 ob ects
Experiment 2 An important feature of the approach is that the ob- servation noise is described with a Gaussian model that uses a full covariance matrix. Representing a distribution on a full covariance matrix is expensive, therefore, the approach is most effective with low dimensional appearance features. However, one expects that more accurate trackers can be build with the increasing dimensionality of the appearance features. In this experiment we evaluate the influence of the appearance features on the overall tracking accuracy. First, the accuracy of tracking based exclusively on the appearance is compared with the accuracy of tracking based exclusively on dynamics. Second, the kernels obtained in the following two cases are also compared: (i) when the association was based only on the colour features, and (ii) when the association was based on dynamics and colour features . Fig. 14A presents the posterior distributions on labels obtained when tracking used exclusively the appearance features. The average error was 40%. This experiment shows that the low-dimensional appearance features have little dis- criminative power. Fig. 14B shows the posterior distributions on labels obtained when tracking was based exclusively the dynamic features. In this case the average error was 38%. The experiment shows, that the dynamics and the assumed model for Markovian transitions Pδ do not uniquely determine the correct trajectories. However, as seen from the previous experiments, the dynamics jointly with appearance features lead to accurate estimation of trajectories. After processing 70 observations the algorithm maintains M = 10 Gaussian kernels representing the intrinsic col- our properties ( is the assumed practical limit; in principle there should be 70 kernels) . Every kernel Xj is represented by a posterior distribution over its parameters φ (X. \ θj 70) conditioned on all 70 observations, because filtering updates distributions on all kernels in memory. Given the hyperparameters θ,7o tne expected mean E [ΏXJ] and the expected covariance Ξ[Vj] of the jth kernel are found; E [m-/] = a-//70 and E [V-/] = Cj,70/ (n-/70 - d-1) . The kernels and the appearance features have d = 9 dimensions. For visualization, a 2 dimensional PCA projection of the original data input is found and applied it to the ex- pected kernels. Figs. 15A and 15B present the projections of the observed appearance features (as points) and the expected kernels (as ellipses) . Note, that both plots present 10 kernels; some kernels overlap. Figs . 15A and 15B confirm that the appearance features alone do not distinguish between the persons. The ten kernels inferred using only appearance similarities (Fig. 14B) form only two clusters. In Fig. 15A it is seen that when the dynamic features supported tracking, then the estimated kernels closely correspond to latent appearance features of the tracked persons. For the considered tracking application, the low-dimensional appearance features jointly with the spatio- temporal features, allow for accurate data association.
Conclusions A technique for tracking multiple objects with non- overlapping cameras is described. The technique is based on the idea to identify each object with an unique label and to define a probabilistic dependency between the labels and ob- servations of objects. In such approach, trajectories can be recovered by finding posterior probability distributions on the labels conditioned on the observations. However, the posterior distribution cannot be computed exactly. The essential feature of the approach is that by casting the tracking task into a probabilistic inference problem, a variety of approximate inference methods can be applied. An assumed-density filtering (ADF) approximation for computing the posterior probability distributions has been applied in the present model. These distributions take the form of mixtures with an intractable number of component densities. Each component arises from a combination of labels, i.e., it corresponds to a hypothetical trajectory. ADF assumes a tractable family, and approximates the mixture with such a function from the family that is the closest in the Kullback- Leibler sense [4] . In contrast to an MHT approximation, the ADF approximation does not discard selected components, but replaces the entire mixture with a new density, in such way that the moments of the original density are preserved. The experiments showed superior performance of the ADF approxima- tion compared with the MHT method. In the experiments, the deterministic ADF approximation performed superior to a stochastic method (Markov chain Monte Carlo, MCMC) . The experiments also confirmed that the accuracy of the MCMC approximation improves with increasing computational resources. However, the increased accuracy of the MCMC approximation comes at a computational cost that might be prohibitive for on-line applications. The proposed model assumes Gaussian noise in the ap- pearance features of objects. A Bayesian estimation of mean and covariance of a Gaussian density under data association uncertainty is used. These parameters are considered as latent variables jointly distributed with a λNormal-Inverse Wishart' density. It has been shown how such a density can be used within an ADF algorithm. Our formulation avoids iterative search for maximum-likelihood estimates of the mean and co- variance (e.g. using Expectation-Maximization, as in [20]). Representing a full covariance matrix can be expensive, therefore the proposed method is most effective with low dimen- sional appearance features. The experiments indicated, that such features, jointly with spatio-temporal information about objects, provide enough discriminative capability to distinguish between objects. A difficult aspect of tracking in wide areas is esti- mation of the number of individual trajectories, that is, objects represented by the observations. This task resembles statistical model selection problems, where one has to decide about the number of model parameters (e.g. hidden units, mixture components) that are necessary to explain the data. In principle, our method does not constrain the number of objects, since for every observation a new latent descriptor (a Gaussian kernel) is introduced, corresponding to a new object. Our model is closely related to a Dirichlet model, where over- fitting is avoided due to Bayesian treatment of model parame- ters, which are integrated from the likelihood of the data. Additionally to this property, the tracking model of the present invention takes advantage of (prior) motion constraints that penalize unnecessary introduction of trajectories. In the above on-line tracking has been considered. The focus was on inference with the assumed-density filtering algorithm. One of the extensions of this work is to allow improving association of an observation given some range of the future observations. A convenient inference method for this purpose is the Expectation-Propagation (EP) algorithm. The EP [18] approximates an intractable posterior density with a tractable family, and iteratively improves the approximation using all available data (as opposed to the current ADF method, which performs only one iteration) . The EP has been already shown successful for a "fixed-lag" improvement of posterior distributions in time-series models [21] . Another line of future research involves application of sequential Monte Carlo inference (particle filter) to the proposed model. Particle filters immediately lead to on-line algorithms. Moreover, approximating an intractable posterior distributions with samples rather than functions, allows to apply more complicated noise models for appearance features. Although the presented Gaussian model already performed well, more complicated models might be necessary in environments with more difficult noise conditions.
Background estimation According to another aspect of the present invention, a probabilistic model for robust estimation of features of foreground objects is described, and it is demonstrated how spatial and temporal correlations can be taken into account in a flexible way, without much additional computational cost. We describe our model as a hybrid dynamic Bayesian network (with latent discrete child and continuous parent nodes) . The model is generative in nature, and reflects our a-priori assumptions about the visual scene (such as smooth object/shadow regions and a quasi-static background process) . Exact computation of required quantities is intractable and we derive an approxi- mate Expectation Propagation (EP) algorithm that works essentially in linear time in the number of pixels, regardless of the correlation structure among them. An attractive feature of the model is that it allows for learning important image re- gions or object shapes easily. Moreover, background estimation can be readily coupled in a transparent and consistent way to subsequent processing.
Model We denote each RGB pixel of a video stream as y{k, t) where J = (i, j) with i and j corresponding to the vertical and horizontal spatial indices and t denoting the time frame. To simplify notation we treat k as a linear index and let k = 1 , ..., K with K being the number of pixels per frame. We also use the boldface notation y ( ) to denote all the pixels at t'th frame . Each pixel has a binary indicator r (k, t) = {back = -1, fore = 1} , that indicates whether y(k, t) is associated with background or foreground object. For modelling shadows, we in- troduce a similar and independent binary indicator, c (k, t) with domain {no-shadow = -1, shadow = l} . This enumeration leaves a possibility that a foreground object can be under a shadow as well . We will denote the collection of indicators in the t"th frame as r ( t) and c ( ) and refer' to them as masks. Visual scenes that we are interested in contain smooth shadow and object regions. Therefore, the corresponding masks, when viewed as a binary images, will exhibit long range correlations. A popular approach for modelling such correlations is by a Markov Random Field (MRF) , where a positive coupling is assumed between adjacent pixels. Whilst a powerful and compact model, inference and especially learning in MRF ' s tend to be computationally expensive. We present an alternative approach for introducing long range correlations on the binary masks. Our approach is based on the simple observation that signs of a collection of correlated Gaussian variables are also correlated. We define a linear dynamical system (a Kalman filter model) on a collection of latent variables x(k, t) , from which we obtain binary indi- cators r (k, t) by thresholding. Analogous models were proposed for unsupervised learning of static binary patterns [30] , visualisation [36] and for classification [38] . To our knowledge, however, these ideas were not employed in a dynamical visual scene analysis con- text .
Prior on masks Linear dynamical systems are widely used state space models for continuous time series. In this model, data is as- sumed to be generated independently from a latent Markov process that obeys a single linear regime
s(0) ~N{m, P) s { t) ~N{As { t- l) , Q) x ( t) ~N{ W s ( t) , R) (35) Here, m and P are prior mean and covariance, Q and R are diagonal covariance matrices and A and W are transition and observation matrices that describe the linear mappings between s(t), s(t - 1) and x(t) . By integrating over s(l : T) , it is easy to see that this model induces on x(l : T) a Gaussian dis- tribution with a constrained (but in general full) covariance matrix. A tractable extension to this model is useful for modelling piecewise linear regimes with occasional regime switches :
o(t) ~ p (o) s(t) ~ [o(t) ≠ onset ] N(s ( t - 1),Q) + [o(t) = onset] N(m, P)
Here, o(t) is a binary variable that indicates a regime switch and [text] denotes an indicator, that evaluates to 1 (or 0) whenever the proposition "text" is true (or false) . When o(t) = onset, we switch to a new regime by "reinitializing" the state vector from the prior. To convert this model for real valued data into one for binary data, we use a clipping mechanism analogous as described in [30]. See Figs. 17A-17D as an example. Under this model, quantization of the corresponding hidden variables yields binary masks r and c : onset iff r(k, t - 1) = back for all k - 1...M or t=l o(t)={ no-onset otherwise
sr(t) ~ [o(t) φ onset] JNT(Arsr( t-1) , Qr) + [o( t) =onset] N(m Pr) r(k, t) = sgn(x (k, t) )
Appropriately chosen Wr and Ar yield typical x(k, t) that are smooth functions of k and t, hence their signs also alter- nate slowly. We use this property to model smooth object masks. When an object leaves the scene in the previous time frame, i.e. when all indicators switch to r(k, t - 1) = back for k = 1,...,M, we trigger the onset indicator by setting o(t) = onset. For shadow masks, we use a simpler model without regime switching:
Sc(0) ~N(mc, Pc) sc(t) ~N(Aα sc(t - 1) , Qc) xe(t) ~N(WC Bα{t) ,Ra) c(k, t) =sgn{xc(k, t) )
Generative model for pixel values The pixels of the background image are assumed to be quasi-static. However, due to camera jitter and long term deviation in the illumination conditions, the pixel values are not exactly constant (See Figs. 16A-16C) . The background pixel values, denoted as b (k, t) , are assumed to be generated from the following linear dynamical system:
sb ( 0 ) ~N(mb, b) sb ( t) ~N(Ab sb(t - l),Qb) b ( t) ~N( Wb sb ( t) , Rb)
The rationale behind this model is simple: Suppose that Wb is equal to a "snap-shot" of the background image at time 0 and Ab = 1. In this case sώ(t) is just a scalar and we can interpret it as a global intensity variable that undergoes a random walk with transition noise variance Qb . In general, we can represent the background image with multiple basis vectors arranged as columns of Wb. The expansion coefficients s*, will correspond to a low dimensional representation and the associated variables [A, Q, W, R) h can be learned offine to reflect the expected variation in the background image. The foreground objects may have a variety of colours and textures depending upon the observed scene (e.g. highway, office) . Hence, the prior distribution should be selected on a case by case basis, according to the features which one be- lieves are important. We let
where μ(t) is a parameter vector. When an object leaves the scene, as indicated by o(t), the parameter vector is reinitialized
Mnew ~ p iμ) μ ( t) = [o ( t) ≠ onset] μ ( t - 1) + [σ ( t) = onset] μnew
where p iμ) is a suitable prior distribution. An unrealistic, but simple example is to model individual foreground pixels as independent Gaussian variables, given their mean and covariance p (f \ μ ( t) ) ~N(μ ( t) , Rf ) p iμ i t) ) ~Nimμ , Pμ) (36 )
where we assume the variance Rf and hyperparameters mμ, Pμ to be known. Pixel values and the masks render the observed image as
p (k, t) = [c( , t) = no-shadow] 1 + [c(k, t) = shadow] p yik, t) =p ik, t) i lrik, t) =fore] f(k, t) + [r(k, t) =back] b (k, t) )
Here, 0 < p < 1 is a fixed parameter modelling the amount of illumination drop due to shadows. The graphical model is shown in Figure 18.
Inference For object identification, we are interested in various marginals of the smoothed density p(μ(l : T) |y(l : T) ) or the filtering density p(μ(t) |y(l : t) ) (for online operation) . In either case, the latent binary variables c, r render the model an intractable hybrid graphical model [31] , where desired marginals can be computed only approximately. To select a suitable inference method, consider the model structure: given the masks r, c, we could integrate over the background pixel process analytically (since it is a KFM) . In principle, we could sample from the masks, however, unless the continuous latent parents sc and sr are low dimensional, the prior probabilities of masks can not be computed easily. One needs to sample from sc and sr and this renders Rao- Blackwellized particle filtering or Gibbs sampling computation- ally expensive [25] . Alternatively, one could approximate the prior by a mean-field approximation. However, due to the hard clipping mechanism, mean field with factorized Gaussians as the approximating family would break down (since Kullback-Leibler divergence between any Gaussian and a clipped Gaussian becomes oo) and we need to relax clipping with a sigmoidal soft threshold or use a more exotic approximating family.
Expectation Propagation Here, we investigate an alternative deterministic approximation method, based on Expectation propagation [32] . EP is an iterative message-passing algorithm and generalizes Loopy Belief Propagation (LBP) [28] , in that it is directly applica- ble to arbitrary hybrid graphical models, including the model we have introduced. One can view EP as a local message passing algorithm on a factor graph [29] where messages are passed between factors (potentials representing local dependencies) , and beliefs (marginal potentials) . Unlike multinomial or Gaussian models, where all marginals stay within a closed family, in hybrid models, beliefs computed from mixed messages may have complicated forms (e.g. mixtures, clipped Gaussians) . In such cases, EP replaces a belief with a potential from an approximating exponential family, usually in KL sense, i.e. by match- ing moments. The quality of the approximation depends, how well these "summary statistics" represent the exact beliefs [32] . We implemented each time-slice of our model with a simple factor graph as shown in Fig. 18B. This structure corresponds to a fully factorized approximation to the joint poste- rior. We use Gaussians and multinomials as natural approximating families for the continuous and discrete variables. The message update equations are derived as particular cases of the general scheme in [32] . In our case, the approximation evaluates to four possible families: Gaussian, multinomial, mixture of Gaussians or mixture of clipped Gaussians. For the first two, the approximation is exact. For mixture of Gaussians and clipped Gaussians, moment matching equations can be derived easily. Experiments We demonstrate first the accuracy of EP approximation to the prior p(r) . Inference in this submodel is equivalent to calculating the integral of the Gaussian distribution Jdsrp ixr\ sr) p i sr) in one of the 2K orthants specified by the mask configuration. In one dimension, this integral can be evaluated easily as shown in the previous section; however in higher dimensions we need to reside to approximations. Fortunately, we can ensure by construction that the Gaussian is positive definite, hence the clipped Gaussian will be unimodal and we expect a factorized EP approximation to converge. In the right panel of Fig. 19 we show results of an experiment where we have computed the likelihood of all configurations of r for K = 8, ranked them and compared to importance sampling. We observe in this and similar experiments that EP converges in 2-3 iterations and is indeed very accurate. In the following experiment we illustrate the quality of feature estimation using artificial data generated from the model where objects are assumed to have a single colour as de- fined in Eq. 36. Mask parameters are set to Ar = γl, Qr = (1 - γ) I with the fudge factor y = 0.95. Wr is taken as a truncated Fourier basis, the same as in the previous experiment. Other parameters are im Pr) = ([1.5, 0.5, 0.5], 0.2T). Parameters for shadow masks are identical. For the background process (A, Q, R, m, P, p) b = (1, 0.12, 501, 1, 0.12, 0.6), elements of Wb are drawn uniformly from [0, 255] . Figs. 20A-20D show a typical run where the colour features are estimated very accurately in this parameter regime. To illustrate the performance on real data, we use the video sequence shown in Figs. 16A-16C. Our goal is to identify persons and associate a unique identity based on colour features . We have trained the parameters of the background process (A, 0, W, R, m, P) b using an EM algorithm from frames of the empty scene. Mask parameters are the same as in the synthetic data experiment; in principle these can also be learned from presegmented data. A single Gaussian will not be powerful enough to represent the multimodal foreground colour distribution; a richer model such as a mixture would be more appropriate. For example, we might assume that the foreground pixels f ik, t) are drawn independently from a flat distribution and display estimated features as colour histograms in HSV space. The colour histograms reflect the similarities between objects, see Fig. 21.
Conclusion We have described a model for robust estimation of appearance features where background estimation need not be viewed as an independent preprocessing step. Automatic identity estimation based on these features requires additional machinery; which can be readily coupled in a transparent and consistent way to the presented model . This allows us properly characterize any uncertainty in the foreground detection, which arguably is important for robust object identification. An attractive feature of the model is that any image subregion with an arbitrary shape (e.g. full frames, scanlines or randomly scattered patches) may be processed, since given s variables, the spatial ordering of pixels is irrelevant. Additionally, the dimensionality of s variables can be tuned to trade off correlations and computation time (which scales as
O i S2TK) , where S is the maximum dimension of variables sr, sc,Si,. The model, as it now stands, assumes that at most one object is observed at a given time frame. This assumption holds for small image regions, but for larger images the model has to be extended, e.g. by introducing additional subgraphs for each new object. Whilst the local message passing mechanism offers a computationally tractable solution, it remains to be seen, whether such an extension to the framework described here can be implemented in practice. APPENDIX I
Filtering algorithm Probabilistic filtering in the model involves computing the marginals of the filtered distribution (14) . Below, the details of marginalization are reported. A symbol θ0Ji≡θ0 is used (the parameters of the fixed prior π(X)) , and
A. Computing marginals of the filtered distribution a) Indicator and Counter: The filtered distribution on these variables is represented as a probability table rkim,n) = Qk(Sk = m, Ck = n) . It is found by marginalizing the filtered distribution (14)
Qk(Sk,Ck) = ∑ \ P(Yk\Hkk)Pr1:k/Hk) . (22) Lk z«,w ***
The sparse structure of the model is exploited to show that the multiple summations and multiple integrals are reduced to a sum w.r.t. a single Z^ and an integral w.r.t. single Xk.
Next, a closed-from solution is provided to the integral. From now on, a fixed Sk = m is assumed. The likelihood can be simplified after (8)-(9);
The predictive distribution is expressed in a factorial form; = {πi∑k) [Zk {m) = 0] + δi∑k - Σj) [Zk {m) > 0] ) Pri∑l k-lrHk) (24) where (24) follows from (6) -(7) . Recall that [/] is a binary indicator symbol. These results are put back to the sought marginal (22) ;
Qk(Sk,Ck)
where j = Zk l !. Marginalization w.r.t. Zk ! , i≠m, removes this variable from the predictive distribution Pr (∑ι-.k- ,Hk) , which is reduced to Pr(∑1:k-ifSkfCk,Zk ) ') . Now, consider the case when Zk l =0. It can be seen that variables Xι:jc-ι appear only in the prediction, which after marginalization is reduced to P3.(Sk,Ck,Zk i ) . Now, consider the case, when Zk ( ! = j , such that j > 0. In this case, the variables ι:jc-ι other than ∑j marginalize out from the predictive distribution which is reduced to Prj/Sk,Ck,Zk > ) .
This distribution can be expressed as Qk_1{∑i)Px(Sk,Ck,Zk t- ) , because the filtered distribution at the slice k - l was approximated with a factorial family (see (10)
Qk(SkrCk)cc +
Integration w.r.t. Xj is reduced using the definition of a Delta function. This function is equal one if ∑j = ∑k, and zero otherwise. The marginal compactly is denoted
where λj denotes the integral w.r.t. ∑k. The likelihood is expanded using (23) : P (Yk \zk'm)k) = P (Dk \zk (m> ) P (Ok \∑k) , and the integral becomes =Pδ Dk)lP(Ok\∑k=x)Qk_1j=x)f j > 0,
The integration variable x = {m,V} denotes jointly mean and covariance. The observation model is NiOk; ,V) . The densities Qk-ι (Xj = x) and nix) are "Normal-Inverse Wishart" models, parameterized respectively as φ This integral has a standard solution [10] . For j > 0: λj=PδiDk\Dj) TiOk;aj,k-ι, (I+XJ.Λ-I) Cj,Λ_ι, l+r/j.jt-i) , (25)
where T denotes a multivariate T-distribution (see Appendix II-C) . In case when j=0, then Pδ is replaced with Pδ0 and θjlk-ι with θ0. The normalization term Lk is recovered by first finding an unnormalized table ^, and then applying the normalization constraint in the standard way. η k-l rk(m,n) = —∑λjPr(Sk=m,Ck=n,Zk m>=j). (26) "k j=0
The predictive distribution Pr(SkrCkZk (m) is discussed in Section I-B of this appendix. b) Auxiliary pointers: There are k variables Zk >,---/Zk <k> .
Each marginal distribution is represented as sl k(n) = Qk(Zk 1> =n) . The variable Zk k> is a priori deterministically set, so marginal distributions Qk(Zk')for i = l,...,k-l are found;
It can be shown that all auxiliary pointers other than Zk { and Zk k> marginalize to unity (similarly as in the previous marginal) . For a fixed Zk' =n or Zk k> = j , all memory variables different than Xπ or Xj integrate out to one. For clarity, the summation over Sk is being split into two cases: (1) Sk = i; (2) Sk≠i . The index n, 0 < n < k - 1, enumerates the domain of Zk a>
si,k( ) = ~λn jPr(Sk =±,0^ =n) + βi(n) , Lk ck
where βd(n) is a helper term, denoting the sum over Sk≠i : First a special case β±(k-l) is calculated. From deterministic transitions (5) follows that the variables Zk { and
Zk k> cannot be simultaneously equal to k - 1, so this state is removed from the integration. The predictive distribution in this case is very simple (see I-B)
β±(k-l) = - ∑ ∑λiPr(Sk,Ck,2fk i)=k-l,2fk a*>=j) . (28) fc Sk≠i,Ckj=0
Note, that the normalization term Lk is already available from (26) . Now, the normalization constraint is exploited to efficiently find βi in) , 0 < n < k - 2. The following relation is obtained β.(n) = siιk_1(n) (l- jrk(i,j)-βi(k-l)), 3=1 where and rk represents the normalized posterior Qk (Sk, Ck) . Except for Si_ι k.x in) , this term does not depend on n, hence it is only computed once. c) Current memory node: To find the assumed density Qk i∑k) first a marginal is computed
This marginal is a mixture of k pdfs, where the jth mixture component φ (Xk \ θj) corresponds to posterior on Xk if the last observation of the current object was at time j. Given this condition each component is found by Bayesian updating of φ (∑\ θj k_1) with the observed Ok. In case of Gaussian models, this is a standard procedure [10] ; it is reported in Appendix II-A.2 as "NiWUpdate" θj = NiWUpdate i Ok, θj ι k-λ) . (30)
The mixing weights are
Note that the summed terms are already computed once in (26) . The assumed density has to be a single N-IW function;
Qk i∑k) - rather than a mixture. A single pdf is found in this family that is the closest in the KL-sense to the mixture with the "moment matching" procedure (Appendix II-A.3) θk,k = MMatch (θ°,..., θk_1, Wo, ..., wjc-i) . d) Past memory nodes: To find the assumed density Qki∑j) , 1 < j < k first a marginal is computed
pA) = φ(∑ΛθiιkX
This marginal is a mixture of two N-IW pdfs. The first corresponds to update of the jth memory node if Zk k>=j; it is defined by parameters ΘJ from (30) . The second is the previous density on this node, represented by θj;k-χ. The mixing weights are W and 1 - Wj as defined in (31) . The factorial approximation requires that this marginal is expressed as a single pdf from the N-IW family; Qki∑j) = <p(Xj|θj,k)- The closest density is found in the family to the mixture; θjιk = MMatch iθj , θjιk-χ,Wj,l - Wj) .
B. Predictive distribution At the kth slice, the predictive distribution describes the dependency of the latent variables on the past data , or shortly Pr (Hjc, ∑1:k) • It has been argued that the ADF approximation does not require the joint distribution, but only two specific marginals. These marginals are derived below as case (A) and case (B) . Case A) : In (26) a distribution Pr(Sk,Ck,Zk l ) is used, where m = Sk. The predictive distribution is computed with the
HMM forward algorithm, because the considered variables evolve as a first-order Markov process as defined by our generative model (2) -(5) . Case B) : In (28) a predictive distribution is used P SkXk'Zk'* =k-l,Zk n) = j) where m≠n and j < k - 1. This distribution can also be computed with HMM forward algorithm. APPENDIX II
A. Joint density on mean and covariance of a Gaussian kernel A variable X = {m, V} jointly represents a mean m and covariance V of a Normal distribution. Joint density can conveniently be represented on these variables [10] via a product P(m|v)P(V), where the conditional model P(m|V) is a Normal (JNT) w.r.t. m; and the marginal P(V) is an Inverse Wishart (IW) w.r.t. V (see II-B) . The joint N-IW density is denoted as: φ (X \ θ) = N(m; a,κV)IW(V; η , C) , (34)
where θ = { a., K, η, C} are the parameters. The vector a is the marginal expected value of m. The parameter K is a scaling term. The scalar parameter η and the matrix C generalize the 'scale' and shape' parameters of the Gamma densities; C together with η gives the marginal expected for V: C/( η - d - 1 ) . 1) Prior: Theoretically, the model of (34) does not allow to represent a noninformative distribution. Practically, such parameters can be set θ0 so that the resulting prior density IT (X) = φ (∑\ θ0) is sufficiently vague to express the practical ignorance about and V. 2) Bayesian update: The N- IW model is convenient mathematically, because it is conjugate to N( 0;m, V) - the distribution that generates observed O. If the prior φ (X\ θ) on X = {m,V} is in the N- IW family, then the posterior φ (X\ θn) ∞NiO; , V) φ (∑\ θ)
is also expressed with the same family, only with the updated parameters θn [10] . The update equations are reported as a ΛNiWUpdate ' procedure : θn = NiWUpdate i O, θ) ,
Kn = 1/ ( 1 + 1/κ) , (0 + -&) ηn = 1 + η , K
Cπ C + ( O - a) ( O - a) ' 1 + κ
3) Moment matching: The N- IW family allows for convenient minimization
argminθKL φ
This operation finds the best approximation in the family to the mixture of N- IW densities θ = MMatch ( θ1 , ... , ΘN, wi, ... wN) .
The parameters of Normal are found exactly
K = Σ W3K3 3=1
and the parameters of Inverse Wishart approximately: = , argmaxjWj ' N C = η∑w^C-1 . 3=1
Inverse Wishart distribution The IW model is a multivariate generalization of the Inverse Gamma distribution. A random d x. d matrix X is Inverse Wishart distributed with parameters: scalar η and matrix C, if the density of X is
where
T distribution A random d-dimensional vector x is T distributed with parameters: scalar η , vector and matrix V, if the density of x is
REFERENCES
[1] CE. Antoniak. Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems. Annals of Sta- tlstics, 2:1152-1174, 1974. [2] Yaakov Bar-Shalom and Xiao-Rong Li. Estimation and Tracking: Principles, Techniques and Software . Artech House, 1993. [3] Matthew J. Beal, Zoubin Ghahramani, and Carl E. Ras- mussen. The infinite hidden Markov model. In Advances in Neural Information Processing Systems (NIPS) 14 , 2002. [4] Xavier Boyen and Daphne Koller. Tractable inference for complex stochastic processes. In Proc . of Conf. on Uncertainty in Artificial Intelligence, pages 33-42. Morgan Kaufman, 1998. [5] Q. Cai and J.K. Aggarwal . Tracking human motion in structured environments using a distributed-camera system. IEEE Transactions on Pattern Analysis and Machine Intelligence, 21(11) .1241-1247, November 1999. [6] R. Collins, A. Lipton, H. Fujiyoshi, and T. Kanade .
Algorithms for cooperative multisensor surveillance. Proceedings of the IEEE, 89 (10) : 1456-1477, October 2001. [7] Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. John Wiley & Sons, New York, 1991. [8] A. Doucet, N. de Freitas, and N.J. Gordon, editors.
Sequential Monte Carlo Methods in Practice . Springer Verlag, 2001. [9] M.S. Drew, J. Wei, and Z.N. Li. Illumination- invariant color object recognition via compressed chromaticity histograms of color-channel-normalized images. In Proc . of Int . Conf . on Computer Vision, pages 533-540, 1998. [10] Andrew Gelman, John B. Carlin, Hal S. Stern, and Donald D. Rubin. Bayesian Data Analysis . Chapman & Hall, 1995. [11] W.R. Gilks, S. Richardson, and D.J. Spiegelhalter, editors. Markov Chain Monte Carlo in Practice . CRC Press, London, 1996. [12] Gregory D. Hager and Peter N. Belhumeur. Efficient region tracking with parametric models of geometry and illumination. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20 (10) :1025-1039, October 1998. [13] Timothy Huang and Stuart Russell. Object identification: A Bayesian analysis with application to traffic surveil- lance. Artificial Intelligence, 103 (1-2) : 1-17, 1998. [14] Michael Isard and Andrew Blake. Condensation- conditional density propagation for visual tracking. International Journal of Computer Vision, 29(1) :5-28, 1998. [15] Finn V. Jensen. Bayesian Networks and Decision- Graphs . Springer-Verlag, New York, 2001. [16] Y. Li, A. Hilton, and J. Illingworth. A relaxation algorithm for real-time multiple view 3d-tracking. Image and Vision Computing, 20 (12) : 841-859, October 2002. [17] David G. Lowe. Robust model-based motion tracking- through the integration of search and estimation. International Journal of Computer Vision, 8 (2) : 113-122 , August 1992. [18] Thomas Minka. Expectation Propagation for approxi mate Bayesian inference . PhD thesis, MIT, 2001. [19] Kevin P. Murphy. Dynamic Bayesian Networks: Repre- sentation, Inference and Learning. PhD thesis, UC, Berkley, 2002. [20] H. Pasula, S. Russell, M. Ostland, and Y. Ritov. Tracking many objects with many sensors. In Proc . of Int . Joint Conf . on Artificial Intelligence, pages 1160-1171, 1999. Stock- holm. [21] Yuan Qi and Thomas Minka. Expectation propagation for signal detection in flat-fading channels. In Proc . of IEEE Int . Symposium on Information Theory, 2003. Yokohama, Japan. [22] Carl E. Rasmussen. The infinite Gaussian mixture model. In Advances in Neural Information Processing Systems (NIPS) 12 , pages 554-560, 2000. [23] Lawrence K. Saul and Michael I. Jordan. Mixed mem- ory Markov models: Decomposing complex stochastic processes as mixtures of simpler ones. Machine Learning, 37(l):75-87, 1999. [24] N. Ukita and T. Matsuyama. Real-time cooperative multi-target tracking by communicating active vision agents. In Int . Conf . on Pattern Recogni tion, pages II: 14-19, 2002. [25] C. Andrieu, N. de Freitas, A. Doucet, and M. I. Jordan. An introduction to MCMC for machine learning. Machine Learning, to appear, 2002. [26] P. Fearnhead. Exact and efficient Bayesian inference for multiple changepoint problems. Technical report, Dept . of Math, and Stat . , Lancaster University, 2003. [27] N. Friedman and S. Russell. Image segmentation in- video sequences. A probabilistic approach. In Proc . UAI, 1997 . [28] W. T. Freeman J. Yedidia and Y. Weiss. Understanding belief propagation and its generalizations. In IJCAI, 2001. Distinguished Papers Track. [29] F. R. Kschischang, B. J. Frey, and H.-A. Loeliger. Factor graphs and the sum-product algorithm. IEEE Transactions on Information Theory, 47 (2) -.498-519 , February 2001. [30] D. D. Lee and H. Sompolinsky. Learning a continuous hidden variable model for binary data. In NIPS*98, 1998. [31] U. Lerner and R. Parr. Inference in hybrid networks: Theoretical limits and practical algorithms. In Proc . UAI- 01 , pages 310-318, Seattle, Washington, August 2001. [32] T. Minka. Expectation Propagation for approximate Bayesian inference . PhD thesis, MIT, 2001. [33] J. Rittscher, J. Kato, S. Joga, and A. Blake. A probabilistic background model for tracking. In Proc. European Conf. Computer Vision, 2000. [34] C. Stauffer and W.E.L. Grimson. Adaptive background mixture modelling for realtime tracking. In CVPR ' 99, pages 246- 252, 1999. [35] J. Sullivan, A. Blake, and J. Rittscher. Statistical foreground modelling for object localisation. In Proc . of ECCV, pages 307-323, 2000. [36] M. E. Tipping. Probabilistic visualisation of high- dimensional binary data. In NIPS*98, 1998. [37] D. Wang, T. Feng, H.-Y. Shum, and S. Ma. A novel probability model for background maintenance and subtraction. In 15th Int. Conf . on Vision Interface, pages 109-117. [38] C.K.I. Williams and D. Barber. Bayesian classification with Gaussian processes. IEEE Transactions on PAMI, 20:1342.1351, 1998.

Claims

1. A method for tracking of multiple objects or individuals with non-overlapping or partly overlapping cameras, wherein each object is identified with a label and wherein a probabilistic dependency between the labels and the observations of objects is defined wherein the posterior probability distributions in the model are in the form of mixtures, wherein the mixture is approximated.
2. A model according to claim 1, wherein a Dirichlet process mixture model is used for estimation of the number of distinct objects from the data and wherein the Dirichlet process mixture model is extended with auxiliary variables that allow to model Markovian dependencies between observations of the same object.
3. A method according to claim 1 or 2, wherein the appearance features of a single object are assumed to be generated by a single Gaussian density, wherein the mean and covariance are estimated with Bayesian inference, wherein a "Normal -Inverse Wishart" is used as a prior joint density for mean and covariance .
4. A method according to any of claims 1-4, wherein an observation of an object includes appearance features defined as three color vectors, wherein each color vector is the average (R, G, B) color of pixels within one of three horizontal regions
RI, R2, R3 selected from the object's image.
5. A method according to any of claims 1-5, wherein a joint predictive distribution of the model latent variables, preferably according to formula (10) , is updated preferably according to formula (11) .
6. A method according to any of the claims 1-6 wherein the filtered distribution is approximated, preferably by formulas (11) and (14) .
7. A method for classification of pixels in a video sequence, wherein the class of each pixel is identified with a binary mask variable, and wherein a dependency between observed pixels and their masks is defined.
8. A method according to claim 8, wherein the binary masks are a priori correlated and each binary mask is associated with a continuous variable, wherein a mask is obtained by taking the sign of the continuous variable, and wherein all continuous variables are a priori jointly distributed according to a multivariate Gaussian density.
9. A method according to claim 8 or 9, wherein the masks" correlations are maintained across consecutive frames of a video stream, wherein the continuous variables are assumed to evolve between frames according to a linear Gaussian process.
10. A method according to any of claims 8-10, wherein the posterior distribution on the masks conditioned on the ob- served pixel colours is computed with the Expectation- Propagation algorithm, wherein a Gaussian densities are used as the approximating family for the continuous variables.
11. A method according to any of claims 1-7 and ac- cording to any of claims 8-11.
12. A system comprising one or more processing units, two or more non-overlapping cameras operatively connected to the processing units and software stored in memory to execute the method according to any of claims 1-10.
EP05750259A 2004-05-13 2005-05-10 A hybrid graphical model for on-line multicamera tracking Withdrawn EP1751716A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP05750259A EP1751716A1 (en) 2004-05-13 2005-05-10 A hybrid graphical model for on-line multicamera tracking

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
EP04076423 2004-05-13
EP04077917A EP1596334A1 (en) 2004-05-13 2004-10-21 A hybrid graphical model for on-line multicamera tracking
EP05750259A EP1751716A1 (en) 2004-05-13 2005-05-10 A hybrid graphical model for on-line multicamera tracking
PCT/EP2005/005310 WO2005114579A1 (en) 2004-05-13 2005-05-10 A hybrid graphical model for on-line multicamera tracking

Publications (1)

Publication Number Publication Date
EP1751716A1 true EP1751716A1 (en) 2007-02-14

Family

ID=34969978

Family Applications (2)

Application Number Title Priority Date Filing Date
EP04077917A Withdrawn EP1596334A1 (en) 2004-05-13 2004-10-21 A hybrid graphical model for on-line multicamera tracking
EP05750259A Withdrawn EP1751716A1 (en) 2004-05-13 2005-05-10 A hybrid graphical model for on-line multicamera tracking

Family Applications Before (1)

Application Number Title Priority Date Filing Date
EP04077917A Withdrawn EP1596334A1 (en) 2004-05-13 2004-10-21 A hybrid graphical model for on-line multicamera tracking

Country Status (2)

Country Link
EP (2) EP1596334A1 (en)
WO (1) WO2005114579A1 (en)

Families Citing this family (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8688614B2 (en) 2007-01-26 2014-04-01 Raytheon Company Information processing system
US8010658B2 (en) * 2007-02-09 2011-08-30 Raytheon Company Information processing system for classifying and/or tracking an object
JP4991876B2 (en) * 2007-11-12 2012-08-01 ビ−エイイ− システムズ パブリック リミテッド カンパニ− Sensor control
US7917332B2 (en) 2007-11-12 2011-03-29 Bae Systems Plc Sensor control
JP6573361B2 (en) * 2015-03-16 2019-09-11 キヤノン株式会社 Image processing apparatus, image processing system, image processing method, and computer program
US10102635B2 (en) 2016-03-10 2018-10-16 Sony Corporation Method for moving object detection by a Kalman filter-based approach
US10313422B2 (en) 2016-10-17 2019-06-04 Hitachi, Ltd. Controlling a device based on log and sensor data
CN115937265B (en) * 2022-10-31 2026-02-27 中国航空工业集团公司雷华电子技术研究所 A target tracking method based on inverse gamma-Gaussian inverse Wissaud distribution
CN116125451B (en) * 2022-11-16 2025-10-14 西北工业大学 A background suppression method for port active sonar echo images based on non-parametric Bayesian model
CN115578694A (en) * 2022-11-18 2023-01-06 合肥英特灵达信息技术有限公司 Video analysis computing power scheduling method, system, electronic equipment and storage medium

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See references of WO2005114579A1 *

Also Published As

Publication number Publication date
EP1596334A1 (en) 2005-11-16
WO2005114579A1 (en) 2005-12-01

Similar Documents

Publication Publication Date Title
Bouwmans Subspace learning for background modeling: A survey
Ross et al. Adaptive probabilistic visual tracking with incremental subspace update
Oliver et al. A Bayesian computer vision system for modeling human interactions
Chan et al. Modeling, clustering, and segmenting video with mixtures of dynamic textures
Bazzani et al. Decentralized particle filter for joint individual-group tracking
Chan et al. Mixtures of dynamic textures
Sehairi et al. Comparative study of motion detection methods for video surveillance systems
Reddy et al. Improved foreground detection via block-based classifier cascade with probabilistic decision integration
Chan et al. Classification and retrieval of traffic video using auto-regressive stochastic processes
Sjarif et al. Detection of abnormal behaviors in crowd scene: a review
Tavakkoli et al. Non-parametric statistical background modeling for efficient foreground region detection
Caillette et al. Real-time 3-D human body tracking using learnt models of behaviour
Santhosh et al. Temporal unknown incremental clustering model for analysis of traffic surveillance videos
Oliver et al. Statistical modeling of human interactions
EP1751716A1 (en) A hybrid graphical model for on-line multicamera tracking
Qian et al. Intelligent surveillance systems
Zajdel et al. A sequential Bayesian algorithm for surveillance with nonoverlapping cameras
Mei et al. Integrated Detection, Tracking and Recognition for IR Video-Based Vehicle Classification.
Doulamis Dynamic tracking re-adjustment: a method for automatic tracking recovery in complex visual environments
Lu et al. Detecting unattended packages through human activity recognition and object association
Aloysius et al. Human posture recognition in video sequence using pseudo 2-d hidden markov models
Zajdel et al. Online multicamera tracking with a switching state-space model
Zajdel Bayesian visual surveillance: from object detection to distributed cameras
Zajdel et al. IAS
Hu et al. Region covariance based probabilistic tracking

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20061213

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU MC NL PL PT RO SE SI SK TR

DAX Request for extension of the european patent (deleted)
17Q First examination report despatched

Effective date: 20090702

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20101201