WO2010109644A1 - 被写体識別方法、被写体識別プログラムおよび被写体識別装置 - Google Patents

被写体識別方法、被写体識別プログラムおよび被写体識別装置 Download PDF

Info

Publication number
WO2010109644A1
WO2010109644A1 PCT/JP2009/056229 JP2009056229W WO2010109644A1 WO 2010109644 A1 WO2010109644 A1 WO 2010109644A1 JP 2009056229 W JP2009056229 W JP 2009056229W WO 2010109644 A1 WO2010109644 A1 WO 2010109644A1
Authority
WO
WIPO (PCT)
Prior art keywords
discriminator
aggregate
weight
discriminators
aggregation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2009/056229
Other languages
English (en)
French (fr)
Inventor
亨 米澤
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Glory Ltd
Original Assignee
Glory Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Glory Ltd filed Critical Glory Ltd
Priority to JP2011505767A priority Critical patent/JP5290401B2/ja
Priority to PCT/JP2009/056229 priority patent/WO2010109644A1/ja
Publication of WO2010109644A1 publication Critical patent/WO2010109644A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting
    • G06F18/2148Generating training patterns; Bootstrap methods, e.g. bagging or boosting characterised by the process organisation or structure, e.g. boosting cascade

Definitions

  • the present invention relates to a subject identification method, a subject identification program, and a subject identification device for identifying a predetermined subject by performing learning for separating a predetermined subject image and a non-subject image using a boosting technique.
  • the present invention relates to a subject identification method, a subject identification program, and a subject identification device capable of reducing the time required for detection processing while improving detection accuracy.
  • a face image identification method for automatically identifying whether or not a human face is included in an image captured by a monitoring camera or an authentication camera.
  • a technique such as a subspace method is generally used for such a face image identification method.
  • a face image identification method using the Integral Image method a plurality of rectangular areas are set in an image, and the total value obtained by adding the feature values of all the pixels included in each rectangular area is used.
  • a technique for detecting a face image based on this see Patent Document 1, Patent Document 2, and Non-Patent Document 1).
  • the subspace method when a face image is detected using the subspace method, the subspace method requires a large amount of computation, and therefore the processing time required for the face image detection process is increased.
  • the area of the rectangular area that is the target of the feature value summation value comparison is compared. It is necessary to set large. However, if the area of the rectangular region is increased, the feature value summation value fluctuates greatly due to the influence of direct sunlight in an image where the direct sunlight hits the face, and the detection accuracy of the face image decreases.
  • Non-Patent Document 1 since the technique of Non-Patent Document 1 needs to have a threshold value for each rectangular feature, there is a problem that the ability to exclude non-face images is particularly poor at the initial stage of discrimination.
  • the present invention has been made to solve the above-described problems of the prior art, and can improve the detection accuracy of a subject and reduce the time required for detection processing, a subject identification method, a subject identification program, and a subject
  • An object is to provide an identification device.
  • the present invention provides a subject identification method for identifying a predetermined subject by performing learning for separating a predetermined subject image and a non-subject image using a boosting technique. And selecting the best discriminator having the lowest error rate from among the binarization discriminators respectively corresponding to predetermined feature amounts used for separation of the subject image sample and the non-subject image sample, and the best discriminator A repetition step of repeating the updating of the sample weight based on the weighting factor so that the best discriminator is a discriminator having an error rate of 0.5 in the next learning, If a predetermined number of the best discriminators are selected in the return step, the discriminator group consisting of the best discriminators already selected in the repetition step is assigned to each of the best discriminators.
  • An aggregate discriminator deriving step for deriving an aggregate discriminator corresponding to the discriminator group by performing linear discriminant analysis using the unbinarized data held therein, and the aggregation derived by the aggregate discriminator deriving step Determined by an aggregate weight coefficient determination step for determining an aggregate weight coefficient corresponding to the aggregate discriminator so that the discriminator becomes a discriminator having an error rate of 0.5 in the next learning, and the aggregate weight coefficient determination step
  • a sample weight updating step for updating the sample weight used by the repetition step based on the aggregated weighting factor, and the aggregation discriminator derived by the aggregation discriminating step and the aggregation weighting factor determination step.
  • a final discriminator determining step for determining a final discriminator for separating the subject image and the non-subject image based on the determined aggregate weight coefficient; Characterized in that it contains it.
  • the aggregation discriminator derivation step derives the aggregation discriminator candidates for each predetermined number that is equal to or greater than a predetermined minimum number and equal to or less than a predetermined maximum number, One aggregate discriminator is selected from the derived candidates.
  • the present invention is the above invention, wherein the aggregation discriminator derivation step performs a full scan on the non-subject image in the range from the minimum number to the maximum number up to the predetermined number.
  • the candidate that minimizes the sum of the scan area by the full scan and the partial scan is used as the aggregate discriminator. It is characterized by selecting.
  • the present invention is the above-described invention, wherein the aggregation discriminator deriving step is newly derived so that the aggregation discriminator derivation step is different from the combination of the binarization discriminators included in the already derived aggregation discriminator. A combination of the binarization discriminators included in the unit is determined.
  • the present invention provides the aggregate discriminator newly derived in the above invention so that the aggregation discriminator derivation step does not include the binarization discriminator included in the already derived aggregate discriminator.
  • the binarization discriminator included is determined.
  • the present invention also provides a subject identification program for identifying a predetermined subject by performing learning for separating a predetermined subject image and a non-subject image using a boosting technique, the subject image sample and the non-subject image sample
  • the best discriminator with the lowest error rate is selected from the binarized discriminators respectively corresponding to the predetermined feature amounts used for separation from the above, and the weighting coefficient corresponding to the best discriminator is determined, and the next learning is performed.
  • a repeat procedure for repeatedly updating the sample weight based on the weight coefficient so that the best discriminator is a discriminator having an error rate of 0.5, and a predetermined number of the best discriminators are selected by the repeat procedure.
  • An aggregate discriminator derivation procedure for deriving an aggregate discriminator corresponding to the classifier group by performing shape discriminant analysis, and the error rate in the next learning of the aggregate discriminator derived by the aggregate discriminator derivation procedure is An aggregation weighting factor determination procedure for determining an aggregation weighting factor corresponding to the aggregation discriminator so as to be a discriminator of 0.5, and the iteration based on the aggregation weighting factor determined by the aggregation weighting factor determination procedure.
  • the aggregate discriminator derived by the aggregate discriminator derivation procedure Based on the sample weight update procedure for updating the sample weight used by the return procedure, the aggregate discriminator derived by the aggregate discriminator derivation procedure, and the aggregate weight factor determined by the aggregate weight factor determination procedure Causing the computer to execute a final discriminator determination procedure for determining a final discriminator for separating the subject image and the non-subject image. And butterflies.
  • the present invention also provides a subject identification device for identifying a predetermined subject by performing learning for separating a predetermined subject image sample and a non-subject image sample using a boosting technique, the subject image sample and a non-subject
  • the best discriminator having the lowest error rate is selected from the binarization discriminators respectively corresponding to the predetermined feature amounts used for separation from the image sample, and the weight coefficient corresponding to the best discriminator is determined
  • a repeating means for repeatedly updating the sample weight based on the weighting coefficient so that the best discriminator is a discriminator having an error rate of 0.5, and a predetermined number of the best discriminators by the repeating means Is selected, the non-binarized data held in each of the best classifiers for the classifier group consisting of the best classifiers already selected by the repeating means.
  • An aggregate discriminator deriving unit for deriving an aggregate discriminator corresponding to the discriminator group by performing linear discriminant analysis using and the aggregate discriminator derived by the aggregate discriminator deriving unit in the next learning Based on the aggregate weight coefficient determined by the aggregate weight coefficient determining means, the aggregate weight coefficient determining means for determining the aggregate weight coefficient corresponding to the aggregate discriminator so as to be a discriminator having an error rate of 0.5 Sample weight updating means for updating the sample weight used by the repetition means, the aggregate discriminator derived by the aggregate discriminator deriving means, and the aggregate weight coefficient determined by the aggregate weight coefficient determining means. And a final discriminator determining means for determining a final discriminator for separating the subject image and the non-subject image on the basis thereof. .
  • the best discriminator having the lowest error rate is selected and selected from the binarization discriminators respectively corresponding to predetermined feature amounts used for separation of the subject image sample and the non-subject image sample.
  • the weighting factor corresponding to the best discriminator is determined, and in the next learning, updating of the sample weight based on the determined weighting factor is repeated so that the best discriminator is a discriminator having an error rate of 0.5.
  • a linear discriminant analysis is performed using the unbinarized data held in each of the best discriminators for the discriminator group including the best discriminators that have already been selected.
  • an aggregate classifier corresponding to this classifier so that the derived classifier becomes a classifier with an error rate of 0.5 in the next learning.
  • the sample weight is updated based on the determined aggregate weight coefficient, and the subject image and the non-subject image are separated based on the derived aggregate discriminator and the aggregate weight coefficient.
  • an aggregate discriminator candidate is derived for each predetermined number that is greater than or equal to a predetermined minimum number and less than or equal to a predetermined maximum number, and one aggregate discriminator is selected from the derived candidates.
  • a portion for the area that was not excluded by the full scan in the range larger than the predetermined number after performing the full scan on the non-subject image in the range up to the predetermined number is selected as the aggregate discriminator, so that the exclusion target can be efficiently excluded.
  • the combination of the binarization discriminators included in the newly derived aggregation discriminator is determined so as to be different from the combination of the binarization discriminators included in the already derived aggregation discriminator. As a result, it is possible to improve the discrimination accuracy by avoiding duplication of aggregate discriminators.
  • the binarization discriminator included in the newly derived aggregation discriminator is determined so as not to include the binarization discriminator included in the already derived aggregation discriminator.
  • FIG. 1 is a diagram showing an outline of a subject identification method according to the present invention.
  • FIG. 2 is a block diagram illustrating the configuration of the face image identification apparatus according to the present embodiment.
  • FIG. 3 is a diagram illustrating processing for acquiring a feature amount from a sample image.
  • FIG. 4 is a diagram illustrating a process of calculating an aggregate discriminator candidate.
  • FIG. 5 is a diagram illustrating a process of calculating the offset of the aggregation discriminator candidate.
  • FIG. 6 is a diagram illustrating an example of the aggregate discriminator selection.
  • FIG. 7 is a diagram illustrating a process for deriving an aggregation classifier.
  • FIG. 8 is a flowchart illustrating a processing procedure executed by the face image identification device.
  • FIG. 9 is a flowchart showing the processing procedure of the aggregate discriminator determination process.
  • FIG. 10 is a diagram showing an outline of the Adaboost method.
  • AdaBoost method widely used as a boosting learning method will be described with reference to FIG. 10 and the outline of the subject identification method according to the present invention will be described with reference to FIG.
  • An embodiment of a face image identification device to which the subject identification method according to the invention is applied will be described.
  • a case where the subject to be identified is a face image will be described.
  • FIG. 10 is a diagram showing an outline of the AdaBoost method.
  • the AdaBoost method is a learning method for deriving a final discriminator having a high correct answer rate by combining a large number of binarized discriminators that output binarized discrimination results such as YES / NO and positive / negative based on the learning results. It is.
  • the classifiers to be combined are weak classifiers (hereinafter referred to as “weak classifiers”) whose correct answer rate slightly exceeds 50%. That is, in the AdaBoost method, a final discriminator with a high correct answer rate is derived by combining a number of weak discriminators with a low correct answer rate.
  • the function sign () is a binarization function that is +1 if the value in the parentheses is 0 or more and -1 if the value is less than 0.
  • the discriminator h s (x) is a binarization discriminator that takes a value of ⁇ 1 or +1. If it is determined as class B, it takes a value of -1.
  • the discriminators h s (x) shown in the equation (1-1) are selected one by one in one learning, and the weighting coefficient ⁇ s corresponding to the selected discriminator h s (x) is selected.
  • the final discriminator H (x) is derived by repeating the sequential determination process.
  • the Adaboost method will be described in more detail.
  • the learning sample is ⁇ (x 1 , y 1 ), ( x 2 , y 2 ),..., (x N , y N ) ⁇ .
  • N is the total number of feature quantities to be discriminated.
  • D s (i) is a sample weight when the s-th learning is performed on the i-th learning sample
  • the discriminator corresponding to each feature quantity x i is h s (x i ) and the weighting coefficient of each discriminator is ⁇ s
  • each formula used in the Adaboost method is It becomes.
  • the error rate for each discriminator h s (for example, the probability of misclassifying a sample of class A as class B) ⁇ s is calculated using equation (2-1).
  • the learning sample distribution for each discriminator h s is shown in (1) of the figure, as shown in (4) of the figure. It will be different from the distribution. Then, the number of learning times s is counted up, the distribution shown in (1) in the figure is updated with the distribution calculated in (4) in the figure, and then the processes after (2) in the figure are repeated.
  • the equation (2-3) represents the next learning sample weight so that the best discriminator selected in (2) in the figure becomes a discriminator having an error rate of 0.5 in the next learning. It shows that D s + 1 is determined. In other words, the process of selecting the next best classifier is performed using the learning sample weight that the best classifier is not good at.
  • the AdaBoost method repeats learning to perform selection of the discriminator and optimization of the weight coefficient of each discriminator, and finally, a final discriminator having a high correct answer rate can be derived.
  • the discriminator h s (x) selected by the Adaboost method is a binarization discriminator, and finally the value held in the discriminator is 2 Output after converting to a value. That is, there is a problem in that a decision branch accompanying binary conversion is required, and the amount of calculation is increased.
  • the RealBoost method uses a multi-value discriminator, so it is possible to avoid the problem of increasing the amount of computation due to the decision branch that occurs in the Adaboost method, but for each of the multi-values held by the multi-value discriminator. Since it is necessary to hold the corresponding weighting coefficient, there is a problem that the memory usage increases.
  • FIG. 1 is a diagram showing an outline of a subject identification method according to the present invention.
  • (A) in the figure shows an outline of the Adaboost method described using FIG. 10, and (B) in the figure shows an outline of the subject identification technique according to the present invention.
  • the h i the binary discriminator shown in the figure (A), f i shown in the same figure (B) is a function before the h i is binarized by a predetermined threshold value Each unbinarized discriminator is shown.
  • the discriminator with the smallest error rate is determined as h 1 in the first learning (see (A-1) in FIG. 1). Then, the weighting factor of h 1 is determined (see (A-2) in the figure). In the next learning, the sample for each sample is set so that h 1 becomes a discriminator having an error rate of 0.5. The weight is updated (see (A-3) in the figure).
  • the final discriminator is derived by repeating selection of the discriminator, determination of the weight coefficient for the selected discriminator, and update of the sample weight.
  • a predetermined number of unbinarized discriminators fi are aggregated by using an LDA (Linear Discriminant Analysis) method.
  • LDA Linear Discriminant Analysis
  • the unbinarized discriminators are aggregated according to a predetermined procedure (see (B-1) in the figure), and the aggregate discriminators are derived using LDA (see (B-2) in the figure). ). Further, the weight coefficient of the derived aggregation discriminator is determined (see (B-3) in the figure), and the sample weight for each sample is updated (see (B-4) in the figure).
  • the selection of the aggregate classifier, the determination of the weighting coefficient for the selected aggregate classifier, and the update of the sample weight are repeated to derive one final classifier.
  • the subject identification method according to the present invention since a predetermined number of unbinarized discriminators are linearly combined, the amount of calculation involved in the discrimination processing can be reduced.
  • wasteful judgment branch (h i shown in FIGS. 1 (A) is always (Decision branch accompanying binary conversion to be performed) can be reduced.
  • the discrimination accuracy can be improved.
  • the method shown in FIG. 1B is referred to as “LDAArray method”.
  • LDAArray method the method shown in FIG. 1B is referred to as “LDAArray method”.
  • the LDAArray method is applied to a face image identification device that identifies a face image and a non-face image (for example, a background image).
  • the LDAArray method is not limited to the field of image identification, and can be widely applied to the field targeted by the Adaboost method.
  • FIG. 2 is a block diagram illustrating the configuration of the face image identification device 10 according to the present embodiment.
  • the face image identification device 10 includes a control unit 11 and a storage unit 12.
  • the control unit 11 further includes an Adaboost processing unit 11a, an aggregate discriminator deriving unit 11b, an aggregate weight coefficient determining unit 11c, a sample weight updating unit 11d, and a final discriminator determining unit 11e.
  • the storage unit 12 stores a face image sample 12a, a non-face image sample 12b, an aggregate discriminator candidate 12c, an aggregate discriminator 12d, and an aggregate weight coefficient 12e.
  • the control unit 11 is a processing unit that performs processing for deriving a final discriminator by learning using the above-described LDAArray method.
  • FIG. 2 only the processing unit used to determine the final discriminator is shown, but the processing unit that performs the facial image identification process using the final discriminator determined by the final discriminator determining unit 11e.
  • the face image identification device 10 may be configured to include the above.
  • the AdaBoost processing unit 11a is a processing unit that performs a process of executing the Adaboost method already described with reference to FIG. Further, the AdaBoost processing unit 11a repeats learning using the face image sample 12a and the non-face image sample 12b read from the storage unit 12 as samples, and collectively discriminates a set of the selected binarization discriminator and the determined weight coefficient. The processing to be passed to the container derivation unit 11b is also performed.
  • the AdaBoost processing unit 11a updates the sample weight D s (see FIG. 10) with the received sample weight. Subsequently, the Adaboost processing unit 11a starts over the selection of the binarization discriminator from the beginning. That is, after the learning frequency s shown in FIG. 10 is set to 1, the binarization discriminator selection process and the like are repeated.
  • FIG. 3 is a diagram illustrating processing for acquiring a feature amount from a sample image.
  • (A) in the figure shows a flow of processing for acquiring a feature amount from a face image
  • (B) in the same drawing shows a flow of processing for acquiring a feature amount from a non-face image such as a background image. Respectively. Also, it is assumed that the size of each face image and each non-face image shown in the figure has been adjusted by a prior enlargement / reduction process.
  • the face image is divided into blocks of a predetermined size (see (A-1) of the figure), and for each block, the edge direction, its strength (thickness), and overall strength (See (A-2) in the figure).
  • feature quantities such as an upward edge strength 32a, an upper right edge strength 32a, a right edge strength 32b, a right lower edge strength 32c, and an overall strength 32d of the block 31 are extracted.
  • the thicknesses of the arrows shown in 32a to 32e indicate the strength.
  • 32a to 32e shown in the figure are examples of feature amounts, and the types of feature amounts are not limited.
  • the face image sample 12a can be obtained by performing the same process for other face images.
  • the non-face image is divided into blocks similar to the face image (see (B-1) of the figure), and the same procedure as the face image is performed for each block.
  • feature quantities such as an upward edge strength 34a, an upper right edge strength 34a, a rightward edge strength 34b, a rightward downward edge strength 34c, and an overall strength 34d of the block 33 are extracted. Is done.
  • the feature values for one non-face image are aligned.
  • the non-face image sample 12b is obtained by performing the same process on other non-face images.
  • the aggregation discriminator derivation unit 11b is a processing unit that performs processing for deriving the aggregation discriminator 12d in the LDAArray method described above. Specifically, the aggregate discriminator deriving unit 11b, when a predetermined number of binarization discriminators are selected by the Adaboost processing unit 11a, sets a combination of the selected binarization discriminator and the determined weight coefficient. Is a processing unit that performs processing for deriving an aggregate discriminator by combining these binarization discriminators by LDA.
  • the aggregation discriminator derivation unit 11b derives an aggregation discriminator candidate 12c that is an aggregation discriminator candidate in accordance with the number of binarization discriminators, and one aggregation discriminator 12c is derived from the derived aggregation discriminator candidates 12c. A process for determining the discriminator 12d is also performed.
  • the LDAArray method will be described using each mathematical expression. Assuming that the aggregation counter representing the number of times of deriving the aggregation discriminator is t (1 ⁇ t ⁇ T), the feature quantity is x, the aggregation discriminator corresponding to the feature quantity x is K t (x), and the predetermined offset value is th,
  • the final discriminator F (x) It is expressed as equation (3-1).
  • the function sign () is a binarization function that is +1 if the value in the parentheses is 0 or more and -1 if the value is less than 0.
  • the offset value th can be calculated by a procedure similar to the offset t calculation procedure described later with reference to FIG.
  • the aggregate discriminator K t (x) Is expressed as in equation (3-2).
  • the offset value offset t in the equation (3-2) is not essential, and the final adjustment may be performed with the offset value th in the equation (3-1) after omitting the offset value offset t .
  • the relationship between the unbinarized discriminator f s (i) and the binarized discriminator h s (i) is: It is expressed by equation (4). That is, the binarized discriminator h s (i) is obtained by binarizing the unbinarized discriminator f s (i) with the function sign ().
  • each aggregation counter t for each aggregation counter t, one aggregation classifier Kt (x) is selected from among a plurality of aggregation classifier candidates, and the weight coefficient ⁇ corresponding to the selected aggregation classifier K t (x) is selected.
  • the final discriminator F (x) is derived by repeating the process of sequentially determining t .
  • the learning sample is ⁇ (x 1 , y 1 ), ( x 2 , y 2 ),..., (x N , y N ) ⁇ .
  • N is the total number of feature quantities to be discriminated.
  • L t (i) is a sample weight when the t-th discriminator aggregation is performed on the i-th learning sample
  • Expression (5-3) indicates that the aggregation discriminator K t determines the next learning sample weight L t + 1 so that the next discriminator becomes a discriminator having an error rate of 0.5. ing.
  • the learning sample weight L t + 1 in the next aggregation is updated, the learning sample weight L t is copied to the learning sample weight D s in the Adaboost process in the LDAarray method. Then, in AdaBoost process will be repeated classifier selection processing learning samples weights D s updated by LDAarray method as an initial value.
  • the aggregate discriminator deriving unit 11b has two dimension numbers, that is, the minimum LDA dimension number (min_lda_dim) and the maximum LDA dimension number (max_lda_dim).
  • the “dimension number” represents, for example, the number of feature quantities.
  • values empirical values derived from the balance between processing time and accuracy can be used.
  • an aggregate discriminator candidate 12c is derived by LDA. Then, the derivation process of the aggregate discriminator candidate 12c is repeated until the number of discriminators (s) becomes equal to the maximum number of LDA dimensions (max_lda_dim).
  • an aggregate classifier candidate 12c in which two classifiers are aggregated three classifiers are aggregated.
  • the aggregate discriminator candidate 12c, the aggregate discriminator candidate 12c in which the four discriminators are aggregated, and the aggregate discriminator candidate 12c in which the five discriminators are aggregated are derived, respectively.
  • One aggregation discriminator 12d is selected.
  • FIG. 4 is a diagram illustrating a process of calculating an aggregate discriminator candidate.
  • the minimum LDA dimension number (min_lda_dim) is 4 and the maximum LDA dimension number (max_lda_dim) is 20.
  • the aggregate discriminator deriving unit 11b is equal to the minimum number of LDA dimensions (min_lda_dim), class A (face image sample 12a) and class B Discriminant analysis by LDA is performed using (non-face image sample 12b).
  • the aggregate discriminator candidate k t4 (x) when s is 4 is calculated.
  • the same processing is repeated until s is equal to 20, that is, the maximum number of LDA dimensions (max_lda_dim).
  • FIG. 5 is a diagram illustrating a process for calculating the offset of the aggregation discriminator candidate 12c.
  • 51a, 52a and 53a shown in the figure are graphs representing the probability density distribution of class A (face image sample 12a), and 51b, 52b and 53b shown in the figure are class B (non-face image sample 12b).
  • Graphs representing the probability density distributions of are respectively shown.
  • the horizontal axis shown in the figure the values of the aggregate classifier candidate (k s), the vertical axis represents probability density shown in the figure, represents respectively.
  • offset t4 is calculated as a horizontal axis value corresponding to a point where the class A graph 51a and the class B graph 51b intersect. That is, offset t4 is adjusted so that the probability that a face image is mistakenly recognized as a non-face image is equal to the probability that a non-face image is mistakenly recognized as a face image. Further, the error rate ⁇ t4 is calculated as the area of the hatched portion shown in FIG.
  • the intensive discriminator deriving unit 11b calculates offset tn for each LDA dimension number (s).
  • the aggregate discriminator deriving unit 11b calculates the candidate k tn (x) of each aggregate discriminator by performing the processing shown in FIG. 4 and FIG. Subsequently, the aggregate discriminator deriving unit 11b performs a process of selecting one aggregate discriminator 12d from the calculated aggregate discriminator candidates 12c.
  • selection processing will be described with reference to FIG.
  • FIG. 6 is a diagram illustrating an example of the aggregate discriminator selection.
  • the total scan area for a sample image such as class B
  • the LDA function is executed only once between the minimum number of LDA dimensions (min_lda_dim) and the maximum number of LDA dimensions (max_lda_dim).
  • a graph 61 showing a change in the total scan area) is shown. Further, in the drawing, the graph 61 illustrates the case where the minimum value 62 is taken when the LDA dimension number (s) is 6.
  • the total scan area is n ⁇ image area + (max_lda_dim ⁇ n) ⁇ (area of area that could not be eliminated by n full scans). Become.
  • the relationship between the total scan area calculated in this way and n is as shown in a graph 61, for example.
  • the aggregation discriminator deriving unit 11b performs the determination process shown in FIG. 6 using the aggregation discriminator candidate 12c corresponding to the aggregation counter t, and the LDA dimension number (s) candidate that minimizes the total scan area.
  • k tn is selected as the aggregate discriminator K t .
  • FIG. 6 shows the case where the candidate k tn having the LDA dimension number (s) that minimizes the total scan area is selected as the aggregate discriminator K t , but the LDA dimension number (s) is fixed. It is good. By doing so, since the processing load of the LDA processing does not change depending on the aggregation counter t, parallel processing becomes possible. Therefore, the processing time can be shortened.
  • aggregate weight coefficient determination unit 11 c when the aggregate classifier outlet portion 11b has derived the aggregate classifier K t, and determines the weighting factor for the aggregate classifier K t (aggregate weight coefficient alpha t), as an aggregate weight factor 12e It is a processing unit that performs processing to be stored in the storage unit 12.
  • the aggregation weighting coefficient ⁇ t is calculated using the above equation (5-2).
  • the sample weight updating unit 11d uses each of the learning sample weights L in the next aggregation based on the aggregation discriminator K t derived by the aggregation discriminator deriving unit 11b and the aggregation weight coefficient ⁇ t determined by the aggregation weight coefficient determining unit 11c. This is a processing unit that performs a process of updating t + 1 (see Expression (5-3)). Further, the sample weight updating unit 11d, a learning sample weight L t, is also a processing unit performs a process of copying the learning samples weights D s of AdaBoost processing unit 11a is used.
  • the aggregation discriminator 12d and the aggregation weighting coefficient 12e corresponding to the aggregation counter t are stored in the storage unit 12 while counting up the aggregation counter t.
  • the final discriminator determination unit 11e sets the aggregation counter on the condition that the correct answer rate of the final discriminator F using the aggregation discriminator 12d (K t ) and the aggregation weighting factor 12e ( ⁇ t ) is equal to or greater than a predetermined value. End the loop using t. Note that the final discriminator determining unit 11e also ends this loop even when there is no binarization discriminator (h s ) to be aggregated.
  • Figure 7 is a diagram illustrating the process of deriving the aggregate classifier K t.
  • the control unit 11 performs LDA candidate (aggregate classifier candidate) extraction (see Fig (A)), to determine the aggregate classifiers K 1 learning first (in the figure (See (B)).
  • the storage unit 12 is a storage unit configured by a storage device such as a nonvolatile memory or a hard disk drive, and includes a face image sample 12a, a non-face image sample 12b, an aggregation discriminator candidate 12c, an aggregation discriminator 12d, and an aggregation discriminator.
  • the weight coefficient 12e is stored.
  • description here is abbreviate
  • FIG. 8 is a flowchart showing a processing procedure executed by the face image identification device 10.
  • the minimum LDA dimension (min_lda_dim) and the maximum LDA dimension (max_lda_dim) are set (step S101)
  • the aggregation counter (t) is set to 1 (step S102)
  • the Adaboost counter (s) is set. 1 (step S103). Note that if the discriminator f in FIG. 7 is represented using the aggregation counter (t) and the Adaboost counter (s), it becomes ft ⁇ s .
  • the Adaboost processing unit 11a selects the best discriminator (h s ) (step S104), calculates the weight coefficient ( ⁇ s ) of the best discriminator (h s ) selected in step S104 (step S105). ), The sample weight (D s ) for each sample is updated (step S106).
  • the aggregate discriminator deriving unit 11b determines whether or not the Adaboost counter (s) is equal to or greater than the minimum LDA dimension number (min_lda_dim) (Step S107), and the Adaboost counter (s) is the minimum LDA dimension number. If it is less than (min_lda_dim) (No at Step S107), the Adaboost counter (s) is counted up (Step S110), and the processes after Step S104 are repeated.
  • step S107 when the Adaboost counter (s) is equal to or greater than the minimum LDA dimension number (min_lda_dim) (Yes in step S107), LDA is performed on the unbinarized discriminators (f 1 to f s ), and the aggregate discriminator. A candidate (k s ) is calculated (step S108).
  • step S109 it is determined whether or not the Adaboost counter (s) is equal to the maximum LDA dimension number (max_lda_dim) (step S109). If the Adaboost counter (s) is not equal to the maximum LDA dimension number (max_lda_dim), (No at Step S109), the Adaboost counter (s) is counted up (Step S110), and the processing after Step S104 is repeated.
  • step S109 when the Adaboost counter (s) is equal to the maximum number of LDA dimensions (max_lda_dim) (step S109, Yes), a process for determining the aggregation discriminator (K t ) is performed (step S111).
  • the detailed processing procedure of step S111 will be described later with reference to FIG.
  • the aggregation weight coefficient determination unit 11c determines the weight coefficient ( ⁇ t ) of the aggregation discriminator (K t ) (step S112), and the sample weight update unit 11d updates the sample weight (L t ) ( Step S113). Then, the final discriminator determining unit 11e either determines whether the class A and the class B are sufficiently separated based on the discrimination result by the final discriminator (F) or there is no unaggregated discriminator. It is determined whether or not the condition is satisfied (step S114).
  • step S114 If the determination condition in step S114 is satisfied (step S114, Yes), the final discriminator (F) is determined and the process is terminated. On the other hand, when the determination condition of step S114 is not satisfied (step S114, No), the sample weight (L t ) used by the aggregate discriminator derivation unit 11b is copied to the sample weight (D s ) used by the Adaboost processing unit 11a. (Step S115). Then, the aggregation counter (t) is counted up (step S116), and the processes after step S103 are repeated.
  • FIG. 9 is a flowchart showing the processing procedure of the aggregate discriminator determination process.
  • the aggregate discriminator deriving unit 11b sets the initial value of the LDA dimension number (s) as the minimum LDA dimension number (min_lda_dim) (step S201), and calculates the total area of the entire scan (s ⁇ total area). (Step S202).
  • a partial scan total area ((max_lda_dim-s) ⁇ residual area) is calculated (step S204). . Then, a total scan area (total scan total area + partial scan total area) is calculated (step S205).
  • step S206 it is determined whether or not s is equal to the maximum number of LDA dimensions (max_lda_dim) (step S206). If s is not equal to the maximum number of LDA dimensions (max_lda_dim) (step S206, No), s is counted. (Step S207), and repeat the process after Step S202. On the other hand, when s is equal to the maximum number of LDA dimensions (max_lda_dim) (Yes in step S206), the aggregate discriminator candidate (k s ) corresponding to the LDA dimension number (s) having the smallest total scan area is the aggregate discriminator. (K t ) (step S208), and the process is terminated.
  • the Adaboost processing unit has the lowest error rate among the binarization discriminators respectively corresponding to the predetermined feature amounts used for separating the subject image sample and the non-subject image sample.
  • the best discriminator is selected and a weighting factor corresponding to the selected best discriminator is determined.
  • the best discriminator is determined to be a discriminator having an error rate of 0.5.
  • the update of the sample weight based on the weighting coefficient is repeated, and the aggregate discriminator deriving unit selects the best discriminator group consisting of the best discriminators already selected if a predetermined number of best discriminators are selected.
  • An aggregate discriminator corresponding to the classifier group is derived by performing linear discriminant analysis using the retained unbinarized data, and the aggregate weight coefficient determination unit determines that the derived aggregate discriminator is in the next learning.
  • the aggregate weighting factor corresponding to this aggregate discriminator is determined so that the discriminator has a rate of 0.5, and the sample weight update unit uses the sample weight used by the Adaboost processing unit based on the determined aggregate weight factor
  • the face image identification device is configured to separate the subject image and the non-subject image using the final discriminator determined by the final discriminator determination unit based on the derived aggregate discriminator and the aggregate weight coefficient .
  • the subject identification method, the subject identification program, and the subject identification device according to the present invention are useful when it is desired to perform processing for detecting a specific subject from a predetermined image with high speed and high accuracy. It is suitable for processing that detects a face image from an image.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Data Mining & Analysis (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Biology (AREA)
  • Evolutionary Computation (AREA)
  • General Engineering & Computer Science (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

 アダブースト処理部が、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択し、選択された最良判別器に対応する重み係数を決定し、次の学習ではこの最良判別器を誤り率が0.5である判別器とするように、決定済の重み係数に基づくサンプル重みの更新を繰り返し、集約判別器導出部が、既に選択された最良判別器からなる判別器群について最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって判別器群に対応する集約判別器を導出し、集約重み係数決定部が、導出された集約判別器が次の学習では誤り率が0.5である判別器となるようにこの集約判別器に対応する集約重み係数を決定し、サンプル重み更新部が、決定された集約重み係数に基づいてアダブースト処理部が用いるサンプル重みを更新するように顔画像識別装置を構成する。

Description

被写体識別方法、被写体識別プログラムおよび被写体識別装置
 本発明は、ブースティング手法を用いて所定の被写体画像と非被写体画像とを分離する学習を行うことで所定の被写体を識別する被写体識別方法、被写体識別プログラムおよび被写体識別装置に関し、特に、被写体の検出精度を向上しつつ、検出処理に要する時間を短縮することができる被写体識別方法、被写体識別プログラムおよび被写体識別装置に関するものである。
 従来から、監視カメラや認証用カメラによって撮像された画像に人の顔が含まれているか否かを自動的に識別する顔画像識別手法が知られている。そして、かかる顔画像識別手法には、部分空間法などの技術が一般的に用いられている。
 たとえば、Integral Image法を用いた顔画像識別手法としては、画像中に複数の矩形領域を設定したうえで、各矩形領域に含まれるすべての画素の特徴量を合算することで得られる合算値に基づいて顔画像を検出する技術がある(特許文献1、特許文献2および非特許文献1参照)。
特開2004-362468号公報 特開2007-34723号公報 Paul Viola, Michael Jones, "Rapid Object Detection using a Boosted Cascade of Simple Features", In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Volume 1, pp.511-518, December 2001
 しかしながら、上述した従来技術には、顔画像の検出処理に要する時間をさらに短縮しつつ、検出精度を向上させることが難しいという問題があった。
 具体的には、部分空間法を用いて顔画像を検出する場合、部分空間法は演算量が多いので、顔画像検出処理に要する処理時間がかさんでしまう。
 また、Integral Image法を用いた顔画像識別手法によって顔画像を検出する場合、顔画像検出処理に要する処理時間を短縮するためには、特徴量合算値の算出対象となる矩形領域の面積を比較的大きく設定する必要がある。しかし、矩形領域の面積を大きくすると、直射日光が顔に当たっている画像などでは、直射日光の影響で特徴量合算値が大きく変動し、顔画像の検出精度が低下してしまう。
 また、非特許文献1の技術は、矩形特徴ごとに閾値をもつ必要があるため、特に、判別初期段階で、非顔画像を排除する能力に乏しいという問題もあった。
 これらのことから、顔画像の検出精度を向上しつつ、検出処理に要する時間を短縮することができる顔画像識別方法、顔画像識別プログラムあるいは顔画像識別装置をいかにして実現するかが大きな課題となっている。なお、かかる課題は、顔画像を識別対象とする場合にのみ発生する課題ではなく、特定の被写体を識別対象とする場合についても同様に発生する課題である。
 本発明は、上述した従来技術の課題を解決するためになされたものであり、被写体の検出精度を向上しつつ、検出処理に要する時間を短縮することができる被写体識別方法、被写体識別プログラムおよび被写体識別装置を提供することを目的とする。
 上述した課題を解決し、目的を達成するために、本発明は、ブースティング手法を用いて所定の被写体画像と非被写体画像とを分離する学習を行うことで所定の被写体を識別する被写体識別方法であって、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返工程と、前記繰返工程によって所定個数の前記最良判別器が選択されたならば、前記繰返工程によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出工程と、前記集約判別器導出工程によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定工程と、前記集約重み係数決定工程によって決定された前記集約重み係数に基づいて前記繰返工程によって用いられる前記サンプル重みを更新するサンプル重み更新工程と、前記集約判別器導出工程によって導出された前記集約判別器および前記集約重み係数決定工程によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定工程とを含んだことを特徴とする。
 また、本発明は、上記の発明において、前記集約判別器導出工程は、所定の最小個数以上であって所定の最大個数以下となる前記所定個数ごとに前記集約判別器の候補をそれぞれ導出し、導出した前記候補の中から1つの前記集約判別器を選択することを特徴とする。
 また、本発明は、上記の発明において、前記集約判別器導出工程は、前記最小個数から前記最大個数までの範囲において前記所定個数までの範囲では前記非被写体画像に対する全面スキャンを行ったうえで前記所定個数より大きい範囲では前記全面スキャンで排除できなかったエリアに対する部分スキャンを行うと仮定した場合に、前記全面スキャンおよび前記部分スキャンによるスキャン面積の総和が最小となる前記候補を前記集約判別器として選択することを特徴とする。
 また、本発明は、上記の発明において、前記集約判別器導出工程は、既に導出した前記集約判別器に含まれる前記2値化判別器の組合せとは異なるように、あらたに導出する前記集約判別器に含まれる前記2値化判別器の組合せを決定することを特徴とする。
 また、本発明は、上記の発明において、前記集約判別器導出工程は、既に導出した前記集約判別器に含まれる前記2値化判別器を含まないように、あらたに導出する前記集約判別器に含まれる前記2値化判別器を決定することを特徴とする。
 また、本発明は、ブースティング手法を用いて所定の被写体画像と非被写体画像とを分離する学習を行うことで所定の被写体を識別する被写体識別プログラムであって、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返手順と、前記繰返手順によって所定個数の前記最良判別器が選択されたならば、前記繰返手順によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出手順と、前記集約判別器導出手順によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定手順と、前記集約重み係数決定手順によって決定された前記集約重み係数に基づいて前記繰返手順によって用いられる前記サンプル重みを更新するサンプル重み更新手順と、前記集約判別器導出手順によって導出された前記集約判別器および前記集約重み係数決定手順によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定手順とをコンピュータに実行させることを特徴とする。
 また、本発明は、ブースティング手法を用いて所定の被写体画像サンプルと非被写体画像サンプルとを分離する学習を行うことで所定の被写体を識別する被写体識別装置であって、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返手段と、前記繰返手段によって所定個数の前記最良判別器が選択されたならば、前記繰返手段によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出手段と、前記集約判別器導出手段によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定手段と、前記集約重み係数決定手段によって決定された前記集約重み係数に基づいて前記繰返手段によって用いられる前記サンプル重みを更新するサンプル重み更新手段と、前記集約判別器導出手段によって導出された前記集約判別器および前記集約重み係数決定手段によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定手段とを備えたことを特徴とする。
 本発明によれば、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、選択された最良判別器に対応する重み係数を決定し、次の学習ではこの最良判別器を誤り率が0.5である判別器とするように、決定された重み係数に基づくサンプル重みの更新を繰り返し、所定個数の最良判別器が選択されたならば、既に選択された最良判別器からなる判別器群について最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって判別器群に対応する集約判別器を導出し、導出された集約判別器が次の学習では誤り率が0.5である判別器となるようにこの集約判別器に対応する集約重み係数を決定し、決定された集約重み係数に基づいてサンプル重みを更新し、導出された集約判別器および集約重み係数に基づいて被写体画像と非被写体画像とを分離することとしたので、複数の未2値化判別器を線形判別分析で集約することによって集約判別器を導出し、導出した集約判別器を用いて最終判別器を決定することで、被写体の検出精度を向上しつつ、検出処理に要する時間を短縮することができるという効果を奏する。
 また、本発明によれば、所定の最小個数以上であって所定の最大個数以下となる所定個数ごとに集約判別器の候補をそれぞれ導出し、導出した候補の中から1つの集約判別器を選択することとしたので、集約判別器の選択を柔軟に行うことができるという効果を奏する。また、複数の集約判別器候補を比較することで最適な集約判別器を選択することができるという効果を奏する。
 また、本発明によれば、最小個数から最大個数までの範囲において所定個数までの範囲では非被写体画像に対する全面スキャンを行ったうえで所定個数より大きい範囲では全面スキャンで排除できなかったエリアに対する部分スキャンを行うと仮定した場合に、全面スキャンおよび部分スキャンによるスキャン面積の総和が最小となる候補を集約判別器として選択することとしたので、排除対象を効率的に排除することができるという効果を奏する。
 また、本発明によれば、既に導出した集約判別器に含まれる2値化判別器の組合せとは異なるように、あらたに導出する集約判別器に含まれる2値化判別器の組合せを決定することとしたので、集約判別器の重複を回避することで、判別精度を向上させることができるという効果を奏する。
 また、本発明によれば、既に導出した集約判別器に含まれる2値化判別器を含まないように、あらたに導出する集約判別器に含まれる2値化判別器を決定することとしたので、集約対象とならない2値化判別器をなくすことで、各2値化判別器を有効活用することができるという効果を奏する。
図1は、本発明に係る被写体識別手法の概要を示す図である。 図2は、本実施例に係る顔画像識別装置の構成を示すブロック図である。 図3は、サンプル画像から特徴量を取得する処理を示す図である。 図4は、集約判別器候補を算出する処理を示す図である。 図5は、集約判別器候補のオフセットを算出する処理を示す図である。 図6は、集約判別器選択の一例を示す図である。 図7は、集約判別器を導出する処理を示す図である。 図8は、顔画像識別装置が実行する処理手順を示すフローチャートである。 図9は、集約判別器決定処理の処理手順を示すフローチャートである。 図10は、アダブースト手法の概要を示す図である。
符号の説明
  10  顔画像識別装置
  11  制御部
  11a アダブースト処理部
  11b 集約判別器導出部
  11c 集約重み係数決定部
  11d サンプル重み更新部
  11e 最終判別器決定部
  12  記憶部
  12a 顔画像サンプル
  12b 非顔画像サンプル
  12c 集約判別器候補
  12d 集約判別器
  12e 集約重み係数
 以下に、添付図面を参照して、本発明に係る被写体識別手法の好適な実施例を詳細に説明する。なお、以下では、ブースティング学習手法として広く用いられているアダブースト(AdaBoost)手法について図10を用いて、本発明に係る被写体識別手法の概要について図1を用いて、それぞれ説明した後に、本発明に係る被写体識別手法を適用した顔画像識別装置についての実施例を説明する。また、以下では、識別対象とする被写体を、顔画像とした場合について説明することとする。
 図10は、アダブースト手法の概要を示す図である。アダブースト手法は、YES/NO、正/負といった2値化された判別結果を出力する2値化判別器を学習結果に基づいて多数組み合わせることによって、正答率が高い最終判別器を導出する学習手法である。
 ここで、組合せ対象となる判別器は、正答率が50%を若干超える程度の弱い判別器(以下、「弱判別器」と記載する)である。すなわち、アダブースト手法では、正答率が低い弱判別器を多数組み合わせることで、正答率が高い最終判別器を導出する。
 まず、アダブースト手法に用いられる数式について説明する。なお、以下では、顔画像のサンプル群をクラスA、非顔画像のサンプル群をクラスBとし、クラスAとクラスBとを判別する場合について説明することとする。
 アダブースト手法において、学習回数をs(1≦s≦S)、各特徴量をx、特徴量xに対応する判別器をh-(x)、判別器h(x)の重み係数をαとすると、最終判別器H(x)は、
Figure JPOXMLDOC01-appb-M000001
式(1-1)のようにあらわされる。
 ここで、関数sign()は、かっこ内の値が0以上であれば+1、0未満であれば-1とする2値化関数である。また、式(1-2)に示したように、判別器h(x)は、-1または+1の値をとる2値化判別器であり、クラスAと判別した場合には+1の値をとり、クラスBと判別した場合には-1の値をとる。
 アダブースト手法では、式(1-1)に示した判別器h(x)を1回の学習で1つずつ選択するとともに、選択した判別器h(x)に対応する重み係数αを逐次決定していく処理を繰り返すことで、最終判別器H(x)を導出する。以下では、アダブースト手法についてさらに詳細に説明する。
 xを各特徴量とし、yを{-1,+1}(上記したクラスAは+1、上記したクラスBは-1)とすると、学習サンプルは、{(x,y),(x,y),…,(x,y)}とあらわされる。ここで、Nは、判別対象とする特徴量の総数である。
 また、D(i)を、i番目の学習サンプルに対してs回目の学習を行った場合のサンプル重みとすると、D(i)の初期値は、式「D(i)=1/N」であらわされる。そして、各特徴量xに対応する判別器をh(x)、各判別器の重み係数をαとすると、アダブースト手法に用いられる各数式は、
Figure JPOXMLDOC01-appb-M000002
となる。
 以下では、図10を用いながら、上記した式(2-1)~式(2-4)についてそれぞれ説明する。同図の(1)に示したように、1回目の学習では、サンプル重みD(i)を1/Nとしたうえで、判別器hごとの学習サンプル分布を算出する。このようにすることで、同図に示したように、クラスAの分布とクラスBの分布とが得られる。
 そして、同図の(2)に示したように、式(2-1)を用いて判別器hごとの誤り率(たとえば、クラスAのサンプルをクラスBと誤判別した確率)εを算出し、最も誤り率εが低い、すなわち、最も良好な判別を行った判別器hを最良判別器として選択する。
 つづいて、同図の(3-1)に示したように、式(2-2)を用いて判別器h(同図の(2)で選択された最良判別器)の重み係数αを決定する。そして、式(2-3)を用いて次回の学習における各学習サンプル重みDs+1を更新する。なお、式(2-3)の分母であるZは、式(2-4)であらわされる。
 このようにして、次回の学習サンプル重みDs+1が更新されると、同図の(4)に示したように、判別器hごとの学習サンプル分布は、同図の(1)に示した分布とは異なるものとなる。そして、学習回数sをカウントアップし、同図の(4)で算出された分布で同図の(1)に示した分布を更新したうえで、同図の(2)以降の処理を繰り返す。
 ここで、式(2-3)は、同図の(2)で選択された最良判別器が、次回の学習では、誤り率が0.5である判別器となるように次回の学習サンプル重みDs+1を決定することを示している。すなわち、最良判別器が最も苦手とする学習サンプル重みを用いて次の最良判別器を選択する処理を行うことになる。
 このように、アダブースト手法は、学習を繰り返すことで、判別器の選択と各判別器の重み係数の最適化とを行い、最終的には、正答率が高い最終判別器を導出することができる。しかし、式(1-2)に示したように、アダブースト手法によって選択される判別器h(x)は、2値化判別器であり、判別器内部で保持する値を最終的には2値に変換したうえで出力する。すなわち、2値変換に伴う判断分岐が必要となり、演算量がかさむという問題がある。
 なお、リアルブースト(RealBoost)手法では、多値判別器を用いるので、アダブースト手法で発生する判断分岐による演算量増大の問題を回避することができるが、多値判別器が保持する多値それぞれに対応した重み係数を保持する必要があるため、メモリ使用量が増大するという問題がある。
 そこで、本発明に係る被写体識別手法では、アダブースト手法を改良することで、判断分岐による演算量増大という問題を回避するとともに、リアルブースト手法のように大きなメモリを必要とすることなく識別精度を向上させることとした。以下では、本発明に係る被写体識別手法の概要について図1を用いて説明する。
 図1は、本発明に係る被写体識別手法の概要を示す図である。なお、同図の(A)には、図10を用いて説明したアダブースト手法の概要について、同図の(B)には、本発明に係る被写体識別手法の概要についてそれぞれ示している。また、同図の(A)に示したhは2値化判別器を、同図の(B)に示したfは、hが所定の閾値で2値化する前の関数である未2値化判別器を、それぞれあらわしている。
 図1の(A)に示したように、アダブースト手法では、1回目の学習で、誤り率が最小の判別器をhとして決定する(同図の(A-1)参照)。そして、hの重み係数を決定し(同図の(A-2)参照)、次回の学習では、hが、誤り率が0.5である判別器となるように、各サンプルに対するサンプル重みを更新する(同図の(A-3)参照)。
 そして、判別器の選択、選択した判別器に対する重み係数の決定およびサンプル重みの更新を繰り返すことで、最終判別器を導出する。
 一方、図1の(B)に示したように、本発明に係る被写体識別手法では、所定個数の未2値化判別器fiをLDA(Linear Discriminant Analysis)法を用いて集約することで集約判別器を導出し、導出した1個または複数個の集約判別器に基づいて1個の最終判別器を導出する点に主たる特徴がある。
 具体的には、所定の手順に従って未2値化判別器を集約し(同図の(B-1)参照)、LDAを用いて集約判別器を導出する(同図の(B-2)参照)。また、導出した集約判別器の重み係数を決定するとともに(同図の(B-3)参照)、各サンプルに対するサンプル重みを更新する(同図の(B-4)参照)。
 そして、集約判別器の選択、選択した集約判別器に対する重み係数の決定およびサンプル重みの更新を繰り返すことで、1個の最終判別器を導出する。このように、本発明に係る被写体識別手法では、所定数の未2値化判別器を線形結合するので、判別処理に伴う演算量を削減することができる。
 すなわち、排除対象(上記したクラスB)をある程度分離することができるようになるまで未2値化判別器を集約するので、無駄な判断分岐(図1の(A)に示したhが必ず行う2値変換に伴う判断分岐)を削減することができる。また、図1の(A)に示したアダブースト手法では考慮されていなかった特徴量間の関係を、あらたな特徴として捉えることができるので、判別精度を向上させることができる。
 なお、以下では、図1の(B)に示した手法を、「LDAArray法」と呼ぶこととする。また、以下では、かかるLDAArray法を、顔画像と非顔画像(たとえば、背景画像)との識別を行う顔画像識別装置に適用した場合について説明する。なお、LDAArray法は、画像識別の分野には限らず、アダブースト手法が対象とする分野についても広く適用することができる。
 図2は、本実施例に係る顔画像識別装置10の構成を示すブロック図である。同図に示すように、顔画像識別装置10は、制御部11と、記憶部12とを備えている。また、制御部11は、アダブースト処理部11aと、集約判別器導出部11bと、集約重み係数決定部11cと、サンプル重み更新部11dと、最終判別器決定部11eとをさらに備えている。そして、記憶部12は、顔画像サンプル12aと、非顔画像サンプル12bと、集約判別器候補12cと、集約判別器12dと、集約重み係数12eとを記憶する。
 制御部11は、上記したLDAArray法を用いた学習によって最終判別器を導出する処理を行う処理部である。なお、図2では、最終判別器を決定するために用いられる処理部のみを示しているが、最終判別器決定部11eによって決定された最終判別器を用いて顔画像の識別処理を行う処理部等を含むように顔画像識別装置10を構成することとしてもよい。
 アダブースト処理部11aは、図10を用いて既に説明したアダブースト手法を実行する処理を行う処理部である。また、アダブースト処理部11aは、記憶部12から読み出した顔画像サンプル12aおよび非顔画像サンプル12bをサンプルとする学習を繰り返し、選択した2値化判別器と決定した重み係数との組を集約判別器導出部11bに渡す処理を併せて行う。
 そして、アダブースト処理部11aは、サンプル重み更新部11dから更新後のサンプル重みを受け取った場合には、受け取ったサンプル重みでサンプル重みD(図10参照)を更新する。つづいて、アダブースト処理部11aは、2値化判別器の選択を最初からやり直す。すなわち、図10に示した学習回数sを1としたうえで、2値化判別器の選択処理等を繰り返す。
 ここで、アダブースト処理部11aの学習に用いられる顔画像サンプル12aおよび非顔画像サンプル12bについて図3を用いて説明しておく。図3は、サンプル画像から特徴量を取得する処理を示す図である。
 なお、同図の(A)には、顔画像から特徴量を取得する処理の流れを、同図の(B)には、背景画像のような非顔画像から特徴量を取得する処理の流れを、それぞれ示している。また、同図に示した各顔画像および各非顔画像は、事前の拡大/縮小処理によってサイズ合わせがなされているものとする。
 同図の(A)に示したように、顔画像を所定サイズのブロックに分割し(同図の(A-1)参照)、各ブロックについて、エッジ方向とその強度(太さ)、全体強度といった特徴量を抽出する(同図の(A-2)参照)。
 たとえば、顔画像の左目に相当するブロック31については、上向きエッジ強度32a、右上向きエッジ強度32a、右向きエッジ強度32b、右下向きエッジ強度32c、ブロック31の全体強度32dといった特徴量が抽出される。なお、32a~32eに示した矢印の太さは強度をあらわしている。また、同図に示した32a~32eは、特徴量の一例であり、特徴量の種類は問わない。
 このように、各ブロックについて特徴量を抽出する処理を顔画像全体について繰り返すことで、1枚の顔画像についての特徴量が揃うことになる。そして、同様の処理を他の複数枚の顔画像に対しても行うことで、顔画像サンプル12aが得られる。
 また、同図の(B)に示したように、非顔画像についても顔画像と同様のブロック分割を行い(同図の(B-1)参照)、各ブロックについて、顔画像と同様の手順で特徴量を抽出する(同図の(B-2)参照)。たとえば、顔画像のブロック31に対応する位置のブロック33についても、上向きエッジ強度34a、右上向きエッジ強度34a、右向きエッジ強度34b、右下向きエッジ強度34c、ブロック33の全体強度34dといった特徴量が抽出される。
 このように、各ブロックについて特徴量を抽出する処理を非顔画像全体について繰り返すことで、1枚の非顔画像についての特徴量が揃うことになる。そして、同様の処理を他の複数枚の非顔画像に対しても行うことで、非顔画像サンプル12bが得られる。
 集約判別器導出部11bは、上記したLDAArray法における集約判別器12dを導出する処理を行う処理部である。具体的には、この集約判別器導出部11bは、アダブースト処理部11aによって所定個数の2値化判別器が選択されると、選択された2値化判別器と決定された重み係数との組を受け取り、これらの2値化判別器をLDAによって結合することで、集約判別器を導出する処理を行う処理部である。
 また、集約判別器導出部11bは、集約判別器の候補となる集約判別器候補12cを2値化判別器の個数に応じてそれぞれ導出し、導出した集約判別器候補12cの中から1つの集約判別器12dを決定する処理を併せて行う。
 ここで、LDAArray法について各数式を用いて説明しておく。集約判別器の導出回数をあらわす集約カウンタをt(1≦t≦T)、特徴量をx、特徴量xに対応する集約判別器をK(x)、所定のオフセット値をthとすると、最終判別器F(x)は、
Figure JPOXMLDOC01-appb-M000003
式(3-1)のようにあらわされる。ここで、関数sign()は、かっこ内の値が0以上であれば+1、0未満であれば-1とする2値化関数である。なお、オフセット値thは、図5を用いて後述するoffsetの算出手順と同様の手順で算出することができる。
 また、未2値化判別器をfts(x)、LDAによって算出されるfts(x)の重みをβts、所定のオフセット値をoffsetとすると、集約判別器K(x)は、式(3-2)のようにあらわされる。
 なお、オフセット値offsetの算出手順については、図5を用いて後述する。また、式(3-2)のオフセット値offsetは必須ではなく、オフセット値offsetを省略したうえで、式(3-1)のオフセット値thで最終的な調整を行うこととしてもよい。
 ここで、未2値化判別器f(i)と、2値化判別器h(i)との関係は、
Figure JPOXMLDOC01-appb-M000004
式(4)であらわされる。すなわち、未2値化判別器f(i)を関数sign()で2値化したものが2値化判別器h(i)となる。
 LDAarray法では、集約カウンタtごとに、複数の集約判別器候補の中から集約判別器Kt(x)を1つずつ選択するとともに、選択した集約判別器K(x)に対応する重み係数αを逐次決定していく処理を繰り返すことで、最終判別器F(x)を導出する。以下では、LDAarray法についてさらに詳細に説明する。
 xを各特徴量とし、yを{-1,+1}(上記したクラスAは+1、上記したクラスBは-1)とすると、学習サンプルは、{(x,y),(x,y),…,(x,y)}とあらわされる。ここで、Nは、判別対象とする特徴量の総数である。
 また、L(i)を、i番目の学習サンプルについて、t回目の判別器集約を行った場合のサンプル重みとすると、Lt(i)の初期値は、式「L(i)=1/N」であらわされる。そして、特徴量xに対応する集約判別器をK(x)とすると、LDAarray法に用いられる各数式は、
Figure JPOXMLDOC01-appb-M000005
となる。
 LDAarray法では、式(5-1)を用いて集約判別器Kごとの誤り率(たとえば、クラスAのサンプルをクラスBと誤判別した確率)εを算出する。そして、式(5-1)で算出された誤り率εおよび式(5-2)を用いて集約判別器Kの重み係数αを決定する。さらに、式(5-3)を用いて次回の集約における各学習サンプル重みLt+1を更新する。なお、式(5-3)の分母であるZは、Lt+1を「ΣLt+1(i)=1」とするための規格化因子であり、式(5-4)であらわされる。
 ここで、式(5-3)は、集約判別器Kが、次回の集約では、誤り率が0.5である判別器となるように次回の学習サンプル重みLt+1を決定することを示している。
 このようにして、次回の集約における学習サンプル重みLt+1が更新されると、LDAarray法では、学習サンプル重みLを、アダブースト処理における学習サンプル重みDへコピーする。そして、アダブースト処理では、LDAarray法によって更新された学習サンプル重みDを初期値として判別器選択処理を繰り返すことになる。
 図2の説明に戻り、集約判別器導出部11bについての説明をつづける。集約判別器導出部11bは、最小LDA次元数(min_lda_dim)および最大LDA次元数(max_lda_dim)という2つの次元数を有している。ここで、「次元数」とは、たとえば、特徴量の数をあらわすものとする。また、上記した2つの次元数(最小LDA次元数および最大LDA次元数)としては、処理時間と精度との兼ね合いから導出した値(経験値)を用いることができる。
 そして、アダブースト処理部11aによって選択された判別器の個数(s)が最小LDA次元数(min_lda_dim)以上となると、LDAによって集約判別器候補12cを導出する。そして、集約判別器候補12cの導出処理を、判別器の個数(s)が最大LDA次元数(max_lda_dim)と等しくなるまで繰り返す。
 たとえば、最小LDA次元数(min_lda_dim)が2であり、最大LDA次元数(max_lda_dim)が5である場合には、2個の判別器を集約した集約判別器候補12c、3個の判別器を集約した集約判別器候補12c、4個の判別器を集約した集約判別器候補12c、5個の判別器を集約した集約判別器候補12cをそれぞれ導出し、導出した集約判別器候補12cの中から1つの集約判別器12dを選択する。
 ここで、集約判別器導出部11bが行う集約判別器候補算出処理の概要について図4を用いて説明しておく。図4は、集約判別器候補を算出する処理を示す図である。なお、同図では、最小LDA次元数(min_lda_dim)が4であり、最大LDA次元数(max_lda_dim)が20である場合について示している。
 集約判別器導出部11bは、アダブースト処理部11aによって選択された判別器の個数(s)が4、すなわち、最小LDA次元数(min_lda_dim)と等しくなると、クラスA(顔画像サンプル12a)およびクラスB(非顔画像サンプル12b)を用いてLDAによる判別分析を行う。このようにして、sが4である場合の集約判別器の候補kt4(x)を算出する。そして、同様の処理をsが20、すなわち、最大LDA次元数(max_lda_dim)と等しくなるまで繰り返す。
 ここで、図4に示した各オフセット値(offsettn)の算出手順について図5を用いて説明しておく。図5は、集約判別器候補12cのオフセットを算出する処理を示す図である。なお、同図に示す51a、52aおよび53aは、クラスA(顔画像サンプル12a)の確率密度分布をあらわすグラフを、同図に示す51b、52bおよび53bは、クラスB(非顔画像サンプル12b)の確率密度分布をあらわすグラフを、それぞれ示している。また、同図に示した横軸は各集約判別器候補(k)の値を、同図に示した縦軸は確率密度を、それぞれあらわしている。
 図5に示したように、offsett4は、クラスAのグラフ51aとクラスBのグラフ51bとが、交差する点に対応する横軸値として算出される。すなわち、offsett4は、顔画像を非顔画像と誤認識した確率と非顔画像を顔画像と誤認識した確率とが等しいように調整される。また、誤り率εt4は、同図に示した斜線部の面積として算出される。
 なお、図5に示したように、LDA次元数(s)の変化にともなって、offsettnの値も変化する。このため、集約判別器導出部11bは、LDA次元数(s)ごとにoffsettnをそれぞれ算出する。
 集約判別器導出部11bは、図4および図5に示した処理を行うことで、各集約判別器の候補ktn(x)を、それぞれ算出する。つづいて、集約判別器導出部11bは、算出した集約判別器候補12cの中から1つの集約判別器12dを選択する処理を行う。ここで、かかる選択処理の一例について図6を用いて説明しておく。
 図6は、集約判別器選択の一例を示す図である。なお、同図には、最小LDA次元数(min_lda_dim)から最大LDA次元数(max_lda_dim)までの間で1回だけLDA関数を実行させると仮定した場合におけるスキャン総面積(クラスBなどのサンプル画像に対するスキャン総面積)の変化をあらわすグラフ61を示している。また、同図では、グラフ61が、LDA次元数(s)が6のときに最小値62をとる場合について例示している。
 たとえば、LDA関数を実行させるLDA次元数(s)をnとすると、スキャン総面積は、n×画像面積+(max_lda_dim-n)×(n回の全面スキャンで排除できなかったエリアの面積)となる。このようにして算出されたスキャン総面積とnとの関係は、たとえば、グラフ61のようになる。
 ここで、同図では、LDA次元数(s)が6の場合に最小値62をとる場合について示したが、集約カウンタをtが変化すると、スキャン総面積が最小となる次元数も変化する。このため、集約判別器導出部11bは、集約カウンタtに対応する集約判別器候補12cを用いて図6に示した判定処理を行い、スキャン総面積が最小となるLDA次元数(s)の候補ktnを、集約判別器Kとして選択する。
 なお、図6では、スキャン総面積が最小となるLDA次元数(s)を有する候補ktnを、集約判別器Kとして選択する場合について示したが、LDA次元数(s)を固定することとしてもよい。このようにすることで、LDA処理の処理負荷が集約カウンタtによって変化しないので、並列処理が可能となる。したがって、処理時間の短縮を図ることができる。
 図2の説明に戻り、集約重み係数決定部11cについて説明する。集約重み係数決定部11cは、集約判別器導出部11bが集約判別器Kを導出した場合に、集約判別器Kに対する重み係数(集約重み係数α)を決定し、集約重み係数12eとして記憶部12へ記憶させる処理を行う処理部である。なお、集約重み係数αは、上記した式(5-2)を用いて算出される。
 サンプル重み更新部11dは、集約判別器導出部11bによって導出された集約判別器Kおよび集約重み係数決定部11cによって決定された集約重み係数αに基づいて次回の集約における各学習サンプル重みLt+1を更新する処理(式(5-3)参照)を行う処理部である。また、サンプル重み更新部11dは、学習サンプル重みLを、アダブースト処理部11aが用いる学習サンプル重みDへコピーする処理を行う処理部でもある。
 このようにして、集約カウンタtをカウントアップしながら、集約カウンタtに対応する集約判別器12dおよび集約重み係数12eが記憶部12へ記憶されていく。そして、最終判別器決定部11eは、集約判別器12d(K)および集約重み係数12e(α)を用いた最終判別器Fの正答率が所定値以上となったことを条件として集約カウンタtを用いたループを終了する。なお、最終判別器決定部11eは、集約対象とする2値化判別器(h)がない場合にもかかるループを終了する。
 ここで、制御部11によって行われる集約判別器導出処理についてまとめておく。図7は、集約判別器Kを導出する処理を示す図である。同図に示したように、制御部11は、LDA候補(集約判別器候補)抽出を行い(同図の(A)参照)、学習1回目の集約判別器Kを決定する(同図の(B)参照)。
 そして、Kを決定したならば、つづいて、Kの決定処理を開始し(同図の(C)参照)、Kを決定する(同図の(D)参照)。さらに、Kの決定処理を開始し(同図の(E)参照)、K、Kを順次決定していく。なお、同図では、KのLDA次元数が4で、KのLDA次元数が5である場合について示しているが、このように、後続のKになるほどLDA次元数が増加するとは限らない。
 図2の説明に戻り、記憶部12について説明する。記憶部12は、不揮発性メモリやハードディスクドライブといった記憶デバイスで構成される記憶部であり、顔画像サンプル12aと、非顔画像サンプル12bと、集約判別器候補12cと、集約判別器12dと、集約重み係数12eとを記憶する。なお、記憶部12に記憶される各情報については、制御部11の説明において既に説明したので、ここでの説明は省略する。
 次に、顔画像識別装置10が実行する処理手順について図8を用いて説明する。図8は、顔画像識別装置10が実行する処理手順を示すフローチャートである。同図に示すように、最小LDA次元(min_lda_dim)および最大LDA次元(max_lda_dim)を設定し(ステップS101)、集約カウンタ(t)を1とするとともに(ステップS102)、アダブーストカウンタ(s)を1とする(ステップS103)。なお、集約カウンタ(t)およびアダブーストカウンタ(s)を用いて図7における判別器fをあらわすと、ft-sとなる。
 そして、アダブースト処理部11aは、最良判別器(h)を選択し(ステップS104)、ステップS104で選択された最良判別器(h)の重み係数(α)を算出するとともに(ステップS105)、各サンプルに対するサンプル重み(D)を更新する(ステップS106)。
 つづいて、集約判別器導出部11bは、アダブーストカウンタ(s)が最小LDA次元数(min_lda_dim)以上であるか否かを判定し(ステップS107)、アダブーストカウンタ(s)が最小LDA次元数(min_lda_dim)未満である場合には(ステップS107,No)、アダブーストカウンタ(s)をカウントアップし(ステップS110)、ステップS104以降の処理を繰り返す。
 一方、アダブーストカウンタ(s)が最小LDA次元数(min_lda_dim)以上である場合には(ステップS107,Yes)、未2値化判別器(f~f)についてLDAを行い、集約判別器候補(k)を算出する(ステップS108)。
 つづいて、アダブーストカウンタ(s)が最大LDA次元数(max_lda_dim)と等しいか否かを判定し(ステップS109)、アダブーストカウンタ(s)が最大LDA次元数(max_lda_dim)と等しくない場合には(ステップS109,No)、アダブーストカウンタ(s)をカウントアップし(ステップS110)、ステップS104以降の処理を繰り返す。
 一方、アダブーストカウンタ(s)が最大LDA次元数(max_lda_dim)と等しい場合には(ステップS109,Yes)、集約判別器(K)を決定する処理を行う(ステップS111)。なお、ステップS111の詳細な処理手順については、図9を用いて後述することとする。
 つづいて、集約重み係数決定部11cは、集約判別器(K)の重み係数(α)を決定し(ステップS112)、サンプル重み更新部11dは、サンプル重み(L)を更新する(ステップS113)。そして、最終判別器決定部11eは、最終判別器(F)による判別結果に基づいてクラスAとクラスBとの分離が十分であるか、または、未集約判別器がないか、のいずれかの条件を満たすか否かを判定する(ステップS114)。
 そして、ステップS114の判定条件を満たした場合には(ステップS114,Yes)、最終判別器(F)を決定して処理を終了する。一方、ステップS114の判定条件を満たさなかった場合には(ステップS114,No)、集約判別器導出部11bが用いるサンプル重み(L)をアダブースト処理部11aが用いるサンプル重み(D)へコピーする(ステップS115)。そして、集約カウンタ(t)をカウントアップし(ステップS116)、ステップS103以降の処理を繰り返す。
 次に、図8のステップS111に示した集約判別器決定処理の詳細な処理手順について図9を用いて説明する。図9は、集約判別器決定処理の処理手順を示すフローチャートである。同図に示すように、集約判別器導出部11bは、LDA次元数(s)の初期値を最小LDA次元数(min_lda_dim)とし(ステップS201)、全面スキャン総面積(s×全面積)を算出する(ステップS202)。
 つづいて、s回の全面スキャンで排除できなかったエリアの面積を残存面積としたうええで(ステップS203)、部分スキャン総面積((max_lda_dim-s)×残存面積)を算出する(ステップS204)。そして、総スキャン面積(全面スキャン総面積+部分スキャン総面積)を算出する(ステップS205)。
 つづいて、sが最大LDA次元数(max_lda_dim)と等しいか否かを判定し(ステップS206)、sが最大LDA次元数(max_lda_dim)と等しくない場合には(ステップS206,No)、sをカウントアップしたうえで(ステップS207)、ステップS202以降の処理を繰り返す。一方、sが最大LDA次元数(max_lda_dim)と等しい場合には(ステップS206,Yes)、総スキャン面積が最も小さいLDA次元数(s)に対応する集約判別器候補(k)を集約判別器(K)とし(ステップS208)、処理を終了する。
 上述してきたように、本実施例では、アダブースト処理部が、被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、選択された最良判別器に対応する重み係数を決定し、次の学習ではこの最良判別器を誤り率が0.5である判別器とするように、決定済の重み係数に基づくサンプル重みの更新を繰り返し、集約判別器導出部が、所定個数の最良判別器が選択されたならば、既に選択された最良判別器からなる判別器群について最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって判別器群に対応する集約判別器を導出し、集約重み係数決定部が、導出された集約判別器が次の学習では誤り率が0.5である判別器となるようにこの集約判別器に対応する集約重み係数を決定し、サンプル重み更新部が、決定された集約重み係数に基づいてアダブースト処理部が用いるサンプル重みを更新し、導出された集約判別器および集約重み係数に基づいて最終判別器決定部が決定した最終判別器を用いて被写体画像と非被写体画像とを分離するように顔画像識別装置を構成した。
 したがって、アダブースト手法における判断分岐による演算量増大という問題を回避するとともに、リアルブースト手法のように大きなメモリを必要とすることなく識別精度を向上させることができる。すわなち、被写体の検出精度を向上しつつ、検出処理に要する時間を短縮することが可能となる。
 以上のように、本発明に係る被写体識別方法、被写体識別プログラムおよび被写体識別装置は、所定の画像から特定の被写体を検出する処理を高速かつ高精度に行いたい場合に有用であり、特に、背景画像から顔画像を検出する処理に適している。

Claims (7)

  1.  ブースティング手法を用いて所定の被写体画像と非被写体画像とを分離する学習を行うことで所定の被写体を識別する被写体識別方法であって、
     被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返工程と、
     前記繰返工程によって所定個数の前記最良判別器が選択されたならば、前記繰返工程によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出工程と、
     前記集約判別器導出工程によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定工程と、
     前記集約重み係数決定工程によって決定された前記集約重み係数に基づいて前記繰返工程によって用いられる前記サンプル重みを更新するサンプル重み更新工程と、
     前記集約判別器導出工程によって導出された前記集約判別器および前記集約重み係数決定工程によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定工程と
     を含んだことを特徴とする被写体識別方法。
  2.  前記集約判別器導出工程は、
     所定の最小個数以上であって所定の最大個数以下となる前記所定個数ごとに前記集約判別器の候補をそれぞれ導出し、導出した前記候補の中から1つの前記集約判別器を選択することを特徴とする請求項1に記載の被写体識別方法。
  3.  前記集約判別器導出工程は、
     前記最小個数から前記最大個数までの範囲において前記所定個数までの範囲では前記非被写体画像に対する全面スキャンを行ったうえで前記所定個数より大きい範囲では前記全面スキャンで排除できなかったエリアに対する部分スキャンを行うと仮定した場合に、前記全面スキャンおよび前記部分スキャンによるスキャン面積の総和が最小となる前記候補を前記集約判別器として選択することを特徴とする請求項2に記載の被写体識別方法。
  4.  前記集約判別器導出工程は、
     既に導出した前記集約判別器に含まれる前記2値化判別器の組合せとは異なるように、あらたに導出する前記集約判別器に含まれる前記2値化判別器の組合せを決定することを特徴とする請求項1、2または3に記載の被写体識別方法。
  5.  前記集約判別器導出工程は、
     既に導出した前記集約判別器に含まれる前記2値化判別器を含まないように、あらたに導出する前記集約判別器に含まれる前記2値化判別器を決定することを特徴とする請求項1、2または3に記載の被写体識別方法。
  6.  ブースティング手法を用いて所定の被写体画像と非被写体画像とを分離する学習を行うことで所定の被写体を識別する被写体識別プログラムであって、
     被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返手順と、
     前記繰返手順によって所定個数の前記最良判別器が選択されたならば、前記繰返手順によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出手順と、
     前記集約判別器導出手順によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定手順と、
     前記集約重み係数決定手順によって決定された前記集約重み係数に基づいて前記繰返手順によって用いられる前記サンプル重みを更新するサンプル重み更新手順と、
     前記集約判別器導出手順によって導出された前記集約判別器および前記集約重み係数決定手順によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定手順と
     をコンピュータに実行させることを特徴とする被写体識別プログラム。
  7.  ブースティング手法を用いて所定の被写体画像サンプルと非被写体画像サンプルとを分離する学習を行うことで所定の被写体を識別する被写体識別装置であって、
     被写体画像サンプルと非被写体画像サンプルとの分離に用いる所定の特徴量にそれぞれ対応する2値化判別器の中から最も誤り率が低い最良判別器を選択するとともに、当該最良判別器に対応する重み係数を決定し、次の学習では当該最良判別器を誤り率が0.5である判別器とするように当該重み係数に基づくサンプル重みの更新を繰り返す繰返手段と、
     前記繰返手段によって所定個数の前記最良判別器が選択されたならば、前記繰返手段によって既に選択された前記最良判別器からなる判別器群について該最良判別器のそれぞれに保持される未2値化データを用いて線形判別分析を行うことによって当該判別器群に対応する集約判別器を導出する集約判別器導出手段と、
     前記集約判別器導出手段によって導出された前記集約判別器が次の学習では前記誤り率が0.5である判別器となるように当該集約判別器に対応する集約重み係数を決定する集約重み係数決定手段と、
     前記集約重み係数決定手段によって決定された前記集約重み係数に基づいて前記繰返手段によって用いられる前記サンプル重みを更新するサンプル重み更新手段と、
     前記集約判別器導出手段によって導出された前記集約判別器および前記集約重み係数決定手段によって決定された前記集約重み係数に基づいて前記被写体画像と前記非被写体画像とを分離する最終判別器を決定する最終判別器決定手段と
     を備えたことを特徴とする被写体識別装置。
PCT/JP2009/056229 2009-03-27 2009-03-27 被写体識別方法、被写体識別プログラムおよび被写体識別装置 Ceased WO2010109644A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2011505767A JP5290401B2 (ja) 2009-03-27 2009-03-27 被写体識別方法、被写体識別プログラムおよび被写体識別装置
PCT/JP2009/056229 WO2010109644A1 (ja) 2009-03-27 2009-03-27 被写体識別方法、被写体識別プログラムおよび被写体識別装置

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2009/056229 WO2010109644A1 (ja) 2009-03-27 2009-03-27 被写体識別方法、被写体識別プログラムおよび被写体識別装置

Publications (1)

Publication Number Publication Date
WO2010109644A1 true WO2010109644A1 (ja) 2010-09-30

Family

ID=42780353

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2009/056229 Ceased WO2010109644A1 (ja) 2009-03-27 2009-03-27 被写体識別方法、被写体識別プログラムおよび被写体識別装置

Country Status (2)

Country Link
JP (1) JP5290401B2 (ja)
WO (1) WO2010109644A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2012060463A1 (ja) * 2010-11-05 2012-05-10 グローリー株式会社 被写体検出方法および被写体検出装置

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003044853A (ja) * 2001-05-22 2003-02-14 Matsushita Electric Ind Co Ltd 顔検出装置、顔向き検出装置、部分画像抽出装置及びそれらの方法
JP2008217589A (ja) * 2007-03-06 2008-09-18 Toshiba Corp 学習装置及びパターン認識装置

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101325691B (zh) * 2007-06-14 2010-08-18 清华大学 融合不同生存期的多个观测模型的跟踪方法和跟踪装置

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003044853A (ja) * 2001-05-22 2003-02-14 Matsushita Electric Ind Co Ltd 顔検出装置、顔向き検出装置、部分画像抽出装置及びそれらの方法
JP2008217589A (ja) * 2007-03-06 2008-09-18 Toshiba Corp 学習装置及びパターン認識装置

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
KANNO ET AL.: "Tadan Ryushi Filter o Mochiita Buttai Ninshiki no Heiretsu Jisso", IEICE TECHNICAL REPORT SIS, vol. 108, no. 85, 5 June 2008 (2008-06-05), pages 11 - 16 *
MURATA: "Boosting no Kikagakuteki Kosatsu", IEICE TECHNICAL REPORT NC, vol. 102, no. 381, 10 October 2002 (2002-10-10), pages 37 - 42 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2012060463A1 (ja) * 2010-11-05 2012-05-10 グローリー株式会社 被写体検出方法および被写体検出装置
JP2012099070A (ja) * 2010-11-05 2012-05-24 Glory Ltd 被写体検出方法および被写体検出装置

Also Published As

Publication number Publication date
JPWO2010109644A1 (ja) 2012-09-27
JP5290401B2 (ja) 2013-09-18

Similar Documents

Publication Publication Date Title
Li et al. Background data resampling for outlier-aware classification
US10002290B2 (en) Learning device and learning method for object detection
US8331655B2 (en) Learning apparatus for pattern detector, learning method and computer-readable storage medium
WO2022001137A1 (zh) 一种行人重识别方法、装置、设备及介质
US20100290700A1 (en) Information processing device and method, learning device and method, programs, and information processing system
WO2011044058A2 (en) Detecting near duplicate images
CN101937513A (zh) 信息处理设备、信息处理方法和程序
CN106485260A (zh) 对图像的对象进行分类的方法和设备及计算机程序产品
JPWO2010004958A1 (ja) 個人認証システム、個人認証方法
JP6897749B2 (ja) 学習方法、学習システム、および学習プログラム
JP5706131B2 (ja) 被写体検出方法および被写体検出装置
JP2008262331A (ja) オブジェクト追跡装置およびオブジェクト追跡方法
US20190303714A1 (en) Learning apparatus and method therefor
TW201621754A (zh) 多類別物件分類方法及系統
WO2019188054A1 (en) Method, system and computer readable medium for crowd level estimation
JP5214679B2 (ja) 学習装置、方法及びプログラム
Shang et al. Improving training and inference of face recognition models via random temperature scaling
CN112560787A (zh) 一种行人重识别匹配边界阈值设置方法、装置及相关组件
WO2010109645A1 (ja) 被写体識別方法、被写体識別プログラムおよび被写体識別装置
JP5290401B2 (ja) 被写体識別方法、被写体識別プログラムおよび被写体識別装置
Rabinowitz et al. Ghost: Gaussian hypothesis open-set technique
JP5769488B2 (ja) 認識装置、認識方法及びプログラム
JP2016062249A (ja) 識別辞書学習システム、認識辞書学習方法および認識辞書学習プログラム
JPWO2014118976A1 (ja) 学習方法、情報変換装置および学習プログラム
JP2014153837A (ja) 識別装置、データ判別装置、ソフトカスケード識別器を構成する方法、データの識別方法、および、プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 09842258

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2011505767

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 09842258

Country of ref document: EP

Kind code of ref document: A1