WO2017150211A1 - 行動認識装置及び行動学習装置並びに行動認識プログラム及び行動学習プログラム - Google Patents
行動認識装置及び行動学習装置並びに行動認識プログラム及び行動学習プログラム Download PDFInfo
- Publication number
- WO2017150211A1 WO2017150211A1 PCT/JP2017/005850 JP2017005850W WO2017150211A1 WO 2017150211 A1 WO2017150211 A1 WO 2017150211A1 JP 2017005850 W JP2017005850 W JP 2017005850W WO 2017150211 A1 WO2017150211 A1 WO 2017150211A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- behavior
- recognition
- action
- time
- time series
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
Definitions
- the present invention relates to machine learning, and relates to the field of learning and recognizing a target action.
- the supervised learning in which the target value to be predicted is included in the training data includes an identification (classification) problem for predicting a class. Improvements in reliability, speeding up of processing, and the like are issues.
- a monitoring video of a person or the like is used as input data to recognize the action of the person or the like. In this case, continuous image frames are analyzed. When an action is recognized from a certain frame sequence, an action before the action at the current recognition time can be taken into consideration when recognizing the action in the subsequent frame sequence (the action at the current recognition time).
- Non-Patent Document 1 is a learning technique that can be used with Trancated BPTT: LSTM, etc., and features before a predetermined frame are not referenced during learning. Basically, the amount of data used for action recognition is determined in a certain time (number of frames). In the invention described in Patent Document 1, instead of explicitly giving the start point of the gesture in gesture recognition, an observation signal for a fixed length is generated with the current frame as the end point, and input to the HMM model database to determine the likelihood of each gesture. Ask. The invention also basically determines the amount of data used for action recognition in a certain time (number of frames).
- the present invention has been made in view of the above problems in the prior art, and stabilizes the action before the action at the current recognition time, not the length of time (number of frames), during learning and recognition of action recognition. It is an object to determine the amount of data used for action recognition taking into account appropriate and appropriate considerations, and to improve the accuracy and efficiency of action recognition.
- the behavior recognition apparatus of the present invention for solving the above problems recognizes a behavior based on time-series data of feature quantities of the behavior of the target extracted from data in which the behavior of the target is recorded in chronological order.
- the recognition unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and a series of identical behaviors that are not distinguished is regarded as one behavior, and is arranged in a time series. After the recognition of the number of actions is over, based on the time series data of the feature amount corresponding to a plurality of actions arranged in a time series from the time point before the predetermined number of actions to the current recognition time point, It is characterized by recognizing actions.
- the behavior learning device of the present invention includes a recognition unit that recognizes and learns the behavior based on the time-series data of the feature amount of the target behavior extracted from the training data in which the target behavior is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and recognizes a predetermined number of actions arranged in the time series as one continuous behavior of the same behavior that is not distinguished.
- the action at the current recognition time is recognized based on the time series data of the feature amount corresponding to a plurality of actions arranged in a time series from the time point before the predetermined number of actions to the current recognition time point. It is characterized by that.
- the behavior recognition program of the present invention causes a computer to function as a recognition unit that recognizes a behavior based on time-series data of feature quantities of the behavior of the target extracted from data in which the behavior of the target is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and a predetermined number of behaviors arranged in time series is regarded as one action that is not distinguished.
- the behavior at the current recognition time point is determined based on the time series data of the feature amount corresponding to a plurality of behaviors arranged in time series from the time point before the predetermined number of actions to the current recognition time point. It is characterized by recognition.
- the behavior learning program of the present invention uses a computer as a recognition unit for recognizing and learning the behavior based on the time series data of the feature amount of the target behavior extracted from the training data in which the target behavior is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and a sequence of the same behavior that is not distinguished is defined as one behavior in a time series. After the recognition of the number of actions is finished, based on the time series data of the feature amount corresponding to a plurality of actions arranged in time series from the time point before the predetermined number of actions to the current recognition time point, the current recognition time point It is characterized by recognizing the behavior.
- the present invention at the time of learning and recognition of behavior recognition, not a length of time (number of frames) but a predetermined number of behaviors before the behavior at the current recognition time is included and behaviors before that are not included. Since the amount of data used for action recognition is determined, it is possible to improve the accuracy and efficiency of action recognition regardless of changes in conditions such as early and late action by humans.
- the object to be recognized is the actions of the elderly and their caregivers.
- Specific actions for elderly people to recognize include “sleeping”, “getting up”, “getting up”, “sitting”, “squatting”, “walking”, “meal”, “toilet”, “going out”, “things”
- Basic actions in daily life such as “Take”, and actions that occur at the time of accidents such as falls and falls.
- assistance actions such as “supporting”, “holding”, and “feeding” are also included as the actions of the assistant.
- “conversation” which is an action by a plurality of people is also conceivable.
- FIG. 1 shows a conceptual diagram of a system including the action recognition (learning) device of the present embodiment.
- the action recognition (learning) apparatus is configured by installing an action recognition (learning) program for causing the computer to function as the following units.
- the target is a human
- “data in which the behavior of the target is recorded in time series” is moving image data.
- the moving image data 12 is input to the preprocessing unit 11.
- moving image data 12 as training data is input to the preprocessing unit 11.
- the feature amount 13 of the action is extracted from each frame of the moving image data 12, and time-series data (hereinafter referred to as “feature amount sequence”) 14 of the feature amount is generated.
- the feature quantity sequence 14 is input to the recognition unit 15.
- the behavior recognition (learning) device receives the feature quantity sequence 14 as an input, and recognizes the target behavior (recognition result 16) and the likelihood 17 based on the input feature quantity sequence 14, the likelihood 17,
- the action boundary determining unit 18 is configured to determine a boundary point of an action that switches to a different action based on a feature amount or the like.
- the recognizing unit 15 recognizes the behavior at each time point by following the special information amount sequence 14 in time series.
- the time point that is the recognition target is the current recognition time point.
- the recognizing unit 15 obtains an action reflected in the moving image data corresponding to the current recognition time point and its likelihood.
- an action and its likelihood are output in units of frames.
- the feature amount the case where the image itself of each frame of the moving image is used most simply can be considered.
- optical flow extracted from an image, person position / posture, time information, or the like may be used.
- a human posture joint point coordinates
- the feature amount is usually given as a fixed length value such as giving the feature amount for the past 10 frames starting from the frame at the current recognition time point, or giving all frames seamlessly from past information
- a predetermined number of actions N corresponding to the action of the learning / recognition target frame is set as the starting point.
- the action A is counted as one action and the action B is counted as one action, while these two actions are consecutive with the actions A and B, for example. For example, it counts as 2 actions. In the case of the same action that does not distinguish between “walking” and heel “sitting”, if “walking” “sitting” “walking” continues, it is counted as three actions.
- FIG. 2 is a conceptual diagram showing a feature string of length used by the recognition unit for action recognition.
- FIG. 2A shows a comparative example in which all frames are used
- FIG. 3 is a conceptual diagram showing a feature string of lengths used by the recognition unit for action recognition in a frame
- FIG. 3A shows a comparative example in which the length of the frame 301 is fixed at a fixed number of frames.
- the recognition unit 15 uses the boundary point 19 output by the behavior boundary determination unit 18 as a reference, and the time series from the time point that goes back a predetermined number of actions (two times in the examples of FIGS. 2B and 3B) to the current recognition time point.
- the behavior at the current recognition time is recognized based on the feature amount sequence 20 corresponding to a plurality of behaviors arranged in a row.
- the speed of action differs depending on the person. It is conceivable that the previous action information effective for specifying the action number is not included, but the number of frames of the feature amount sequence used by the recognition unit 15 for action recognition is set based on the action number as in the example of the present invention in FIG. 3B. By making it variable, it becomes possible to sufficiently obtain information on past actions that lead to the action at the time of the current recognition.
- the feature quantity sequence for three actions is obtained by recognizing a predetermined number of actions arranged in time series (two in the example of FIGS. 2B and 3B). After it is over.
- the recognition unit 15 uses a predetermined number of actions arranged in time series (2 in the examples of FIGS. 2B and 3B). Before the recognition of (1) is completed, the behavior at the current recognition time is recognized based on all the feature quantity sequences up to the current recognition time.
- RNN Recurrent Neural Network
- LSTM Short Term Memory
- LSTM is a technology that can hold past information for a longer period of time, and by combining the two, it becomes possible to utilize long-term past data for learning and recognition of the current input.
- RNN + LSTM can reset the internal state with a flag. When not reset, the information of all the frames up to that point is retained internally, but when reset, the internal state is initialized, so it is handled that there is no past input. Therefore, in this embodiment, based on the determination of the action boundary determination unit 18, the process of resetting the internal state and inputting the feature amount again is used as the process of resetting the action used for learning recognition.
- the recognition unit 15 needs to learn before recognition. In learning, moving image data whose correct behavior is known is input, and what is an effective feature amount for distinguishing each behavior is learned. At the time of recognition, it recognizes based on the process made by learning.
- the boundary of actions is known at the time of learning, so it may be reset according to the number of actions.However, since the action is unknown in advance and the same cannot be done at the time of recognition, The determination unit 18 is required. Note that the method used for recognition is not limited to LSTM.
- the recognition unit 15 outputs the likelihood of each target action as the recognition result 16. For example, when 10 kinds of actions are recognized, the likelihood is calculated for each of the 10 actions, and the action with the highest likelihood is output as the recognition result 16.
- the action boundary determination unit 18 determines a boundary point that becomes a break between actions being recognized, and inputs the boundary point to the recognition unit 15.
- the recognition result 16 changes to a different action (when there is a first place change)
- it is considered to be a boundary point, but in that case, the recognition result 16 changes to a different action Since the boundary point is determined for the first time later, the determination is delayed.
- the delay in determination is expected to be greater.
- a method using likelihood information of each action of action recognition can be considered.
- the behavior boundary determination unit 18 determines a time point when the difference 601 between the first and second rankings having the highest likelihood is equal to or less than a predetermined threshold value as a boundary point. That is, in FIG. 6, action 0 is first in frame 1-6, but the first place is not determined in the seventh frame when action 0 switches to action 2 or after, and the first and second positions in the sixth frame.
- the determination is made early by making a determination when the difference 601 is equal to or less than a predetermined threshold.
- the recognition unit 15 updates the feature amount sequence used for the action recognition to the range of the number of actions traced from the new boundary point, thereby improving the accuracy of the action recognition.
- the action boundary determination unit 18 starts from the frame at the current recognition time point, based on a statistic such as the average or median of each action likelihood value in a predetermined range, A method of determining that the behavior has been switched at the stage where the statistics are switched is conceivable. Also, a method of determining the end of the action when the action indicating the maximum likelihood does not change within a predetermined time (the number of frames) after the action indicating the maximum likelihood changes can be considered. In this case, the “mode” can be used as the statistic.
- a boundary point of an action that switches to a different action is determined based on the position information, such as at the moment of leaving the bed.
- a boundary point of an action for switching to a different action is determined based on position information indicating that the person has entered / exited a specific range such as in a bathroom.
- This position information may be target position information obtained by analyzing the moving image data 12 shown in FIG. 1, or may be input from the position detection unit 21 separately.
- the position detection unit 21 is not based on the moving image data 12 but cooperates with a sensing system that detects a target position. Thereby, when the place which performs action, such as bathing, is limited, recognition accuracy can be improved.
- a method of determining that there is a boundary point when the same action continues for a predetermined number of frames or more can be considered. This is because if the same action continues for a too long period of time, the relationship between the previous action and the next action is considered weakened.
- feature length sequences of length used by the recognition unit 15 for action recognition are indicated by frames 801 and 803, and current recognition points are indicated by pointers 802 and 804.
- the recognition unit 15 performs the following action recognition (recognition at the current recognition time point 804 illustrated in FIG. 8B). As shown by a frame 803, the past action used for recognition is shifted by one action for recognition.
- the behavior recognition apparatus of the present invention has a recognition unit that recognizes a behavior based on time-series data of feature quantities of the behavior of the target extracted from data in which the behavior of the target is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and recognizes a predetermined number of actions arranged in the time series as one continuous action of the same action that is not distinguished.
- the action at the current recognition time is recognized based on the time series data of the feature amount corresponding to a plurality of actions arranged in a time series from the time point before the predetermined number of actions to the current recognition time point. It is characterized by that.
- the recognition unit is configured to determine the current recognition time point based on the time series data of all the feature quantities up to the current recognition time point before the recognition of the predetermined number of actions arranged in time series is finished. It is characterized by recognizing actions.
- the behavior recognition apparatus further includes a behavior boundary determination unit that determines a boundary point of a behavior that switches to a different behavior, and the recognition unit uses the boundary point output by the behavior boundary determination unit as a reference. It is characterized in that the action at the current recognition time point is recognized based on the time series data of the feature amount corresponding to a plurality of actions arranged in a time series from the time point before the action number to the current recognition time point.
- the behavior boundary determination unit determines a boundary point of the behavior that switches to a different behavior based on the likelihood information of the behavior output by the recognition unit.
- the behavior boundary determination unit determines the time point when the difference between the first ranking and the second ranking having a high likelihood is equal to or less than a predetermined threshold as the boundary point. .
- the behavior boundary determination unit determines a boundary point of the behavior that switches to a different behavior based on a statistical amount of likelihood information output from the recognition unit a plurality of times within a predetermined length of time. It is characterized by determining.
- the action boundary determination unit determines a boundary point of an action that switches to a different action based on the position information of the target.
- the behavior learning device of the present invention includes a recognition unit that recognizes and learns the behavior based on the time-series data of the feature amount of the target behavior extracted from the training data in which the target behavior is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and recognizes a predetermined number of actions arranged in the time series as one continuous behavior of the same behavior that is not distinguished.
- the action at the current recognition time is recognized based on the time series data of the feature amount corresponding to a plurality of actions arranged in a time series from the time point before the predetermined number of actions to the current recognition time point. It is characterized by that.
- the behavior recognition program of the present invention causes a computer to function as a recognition unit that recognizes a behavior based on time-series data of feature quantities of the behavior of the target extracted from data in which the behavior of the target is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and a predetermined number of behaviors arranged in time series is regarded as one action that is not distinguished.
- the behavior at the current recognition time point is determined based on the time series data of the feature amount corresponding to a plurality of behaviors arranged in time series from the time point before the predetermined number of actions to the current recognition time point. It is characterized by recognition.
- the recognition unit is configured to determine the current recognition time point based on the time series data of all the feature quantities up to the current recognition time point before the recognition of the predetermined number of actions arranged in time series is completed. It is characterized by recognizing actions.
- the computer functions as an action boundary determination unit that determines a boundary point of an action to be switched to a different action, and the recognition unit is based on the boundary point output by the action boundary determination unit. It is characterized in that the action at the current recognition time point is recognized based on the time series data of the feature amount corresponding to a plurality of actions lined up in time series from the time point preceding the predetermined number of actions to the current recognition time point.
- the behavior boundary determination unit determines a boundary point of the behavior that switches to a different behavior based on the likelihood information of the behavior output by the recognition unit.
- the behavior boundary determination unit determines, as the boundary point, a time point when a difference between a first ranking and a second ranking having a high likelihood is equal to or less than a predetermined threshold value.
- the behavior boundary determination unit determines a boundary point of the behavior that switches to a different behavior based on a statistical amount of likelihood information output from the recognition unit a plurality of times within a predetermined length of time. It is characterized by determining.
- the action boundary determination unit determines a boundary point of an action that switches to a different action based on the target position information.
- the behavior learning program of the present invention uses a computer as a recognition unit for recognizing and learning the behavior based on the time series data of the feature amount of the target behavior extracted from the training data in which the target behavior is recorded in time series.
- the recognizing unit recognizes the behavior at each time point by following the time series data of the feature amount in time series, and a sequence of the same behavior that is not distinguished is defined as one behavior in a time series. After the recognition of the number of actions is finished, based on the time series data of the feature amount corresponding to a plurality of actions arranged in time series from the time point before the predetermined number of actions to the current recognition time point, the current recognition time point It is characterized by recognizing the behavior.
- the present invention can be used for action recognition of a person or the like by a computer or the like.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Image Analysis (AREA)
Abstract
行動認識の学習及び認識時において、時間の長さ(フレーム数)ではなく、現認識時点の行動の前の行動を安定的かつ適度に考慮に入れて行動認識に用いるデータ量を決定し、行動認識の高精度化及び効率化を図る。人などの対象の行動が時系列に記録されたデータ(動画像データ)から抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部(12)を有する行動認識装置において、認識部は、特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点(802,804)までの時系列に並ぶ複数(図8において3)の行動に相当する特徴量の時系列データに基づき、現認識時点の行動を認識する。
Description
本発明は、機械学習に係り、対象の行動を学習し認識する分野に関する。
従来、コンピューターに明示的にプログラミングすることなく行動させるようにする機械学習が研究されている。予測する目標の値が訓練データに含まれている教師あり学習には、クラスを予測する識別(分類)問題などがある。信頼性の向上、処理の高速化等が課題となっている。また、人等の監視動画を入力データとし、人等の行動を認識する分野がある。この場合、連続する画像フレームを解析することとなる。あるフレーム列から行動が認識されると、その後のフレーム列における行動(現認識時点の行動)を認識するにあたり、現認識時点の行動の前の行動を考慮することができる。
非特許文献1に記載の発明は、Trancated BPTT:LSTM等でもちいられる学習テクニックであり、学習時に、所定のフレームよりも前の特徴は参照しないようにする。基本的に一定の時間(フレーム数)で行動認識に用いるデータ量を決める。
特許文献1に記載の発明は、ジェスチャ認識においてジェスチャの始点を明示的に与える代わりに、現フレームを終点として固定長分の観測信号を生成し、HMMモデルデータベースに入力し各ジェスチャの尤度を求める。同発明も、基本的に一定の時間(フレーム数)で行動認識に用いるデータ量を決める。
非特許文献1に記載の発明は、Trancated BPTT:LSTM等でもちいられる学習テクニックであり、学習時に、所定のフレームよりも前の特徴は参照しないようにする。基本的に一定の時間(フレーム数)で行動認識に用いるデータ量を決める。
特許文献1に記載の発明は、ジェスチャ認識においてジェスチャの始点を明示的に与える代わりに、現フレームを終点として固定長分の観測信号を生成し、HMMモデルデータベースに入力し各ジェスチャの尤度を求める。同発明も、基本的に一定の時間(フレーム数)で行動認識に用いるデータ量を決める。
David Zipser(Department of Cognitive Science,University of California, San Diego,La Jolla, CA 92093) Subgrouping reduces complexity and speeds up learning in recurrent networks
Graves, Alan, Abdel-rahman Mohamed, and Geoffrey Hinton. "Speech recognition with deep recurrent neural networks." Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013.
Hochreiter, Sepp, and Jurgen Schmidhuber. "Long short-term memory." Neural computation 9.8 (1997): 1735-1780.
しかしながら、一定の時間(フレーム数)で行動認識に用いるデータ量を決める手法では、その一定の時間(フレーム数)内に、認識したい行動の一単位が入らない場合が生じたり、現認識時点の行動の前の行動が入らない場合が生じたり、逆に無駄に多く前の行動が入ったりするなど、人による動作の早い遅い、状況によるばらつきなどが吸収できず、十分な認識精度が得られなかった。
本発明は以上の従来技術における問題に鑑みてなされたものであって、行動認識の学習及び認識時において、時間の長さ(フレーム数)ではなく、現認識時点の行動の前の行動を安定的かつ適度に考慮に入れて行動認識に用いるデータ量を決定し、行動認識の高精度化及び効率化を図ることを課題とする。
以上の課題を解決するための本発明の行動認識装置は、対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部を有する行動認識装置において、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また本発明の行動学習装置は、対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部を有する行動学習装置において、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また本発明の行動認識プログラムは、対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部としてコンピューターを機能させるための行動認識プログラムにおいて、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また本発明の行動学習プログラムは、対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部としてコンピューターを機能させるための行動学習プログラムにおいて、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
本発明によれば、行動認識の学習及び認識時において、時間の長さ(フレーム数)ではなく、現認識時点の行動の前の所定数の行動を含めるとともにそれ以前の行動を含めないように行動認識に用いるデータ量を決定するので、人による行動の早い遅い等、条件の変化に依らず、行動認識の高精度化及び効率化を図ることができる。
以下に本発明の一実施形態につき図面を参照して説明する。以下は本発明の一実施形態であって本発明を限定するものではない。
例えば、高齢者介護見守りの現場において、高齢者の生活状況や事故を認識する仕組みを考える。
この場合、認識する対象は高齢者やその介助者の行動である。具体的に認識する高齢者の行動としては、”就寝”、”起床”、”離床”、”座る”、”しゃがむ”, ”歩行”、”食事”、”トイレ”、”外出”,”モノを取る”の様な日常生活における基本的な行動や転倒、転落などの事故時に起きる行動が対象となる。介助者の行動としては”歩行”などの基本的な行動の他、”支える”、”抱える”,”食べさせる”などの介助動作も対象となる。また、複数人による行動である”会話”等も考えられる。
これらの行動の内、多くの行動はその前に強く関係がある。例えば“就寝”という行動はベッドに歩いて近づき、横たわった後に発生することが考えられるが、歩行中突然就寝状態になる事は考えにくい。このことは、前の行動は現在の行動を決定する上で非常に重要な情報であることを示している。そのため、行動認識において過去の情報を使うことは精度の向上のために非常に重要である。
従来は、過去10フレーム分の特徴量を認識に用いる、という様に、固定時間分の情報を認識に用いる場合が多かったが、人によって行動の速度は異なり、また同一人物でも繰り返しバラつきがあり、固定時間を設定するのは難しかった。本発明は、そうした問題に対応するための提案である。
この場合、認識する対象は高齢者やその介助者の行動である。具体的に認識する高齢者の行動としては、”就寝”、”起床”、”離床”、”座る”、”しゃがむ”, ”歩行”、”食事”、”トイレ”、”外出”,”モノを取る”の様な日常生活における基本的な行動や転倒、転落などの事故時に起きる行動が対象となる。介助者の行動としては”歩行”などの基本的な行動の他、”支える”、”抱える”,”食べさせる”などの介助動作も対象となる。また、複数人による行動である”会話”等も考えられる。
これらの行動の内、多くの行動はその前に強く関係がある。例えば“就寝”という行動はベッドに歩いて近づき、横たわった後に発生することが考えられるが、歩行中突然就寝状態になる事は考えにくい。このことは、前の行動は現在の行動を決定する上で非常に重要な情報であることを示している。そのため、行動認識において過去の情報を使うことは精度の向上のために非常に重要である。
従来は、過去10フレーム分の特徴量を認識に用いる、という様に、固定時間分の情報を認識に用いる場合が多かったが、人によって行動の速度は異なり、また同一人物でも繰り返しバラつきがあり、固定時間を設定するのは難しかった。本発明は、そうした問題に対応するための提案である。
図1に本実施形態の行動認識(学習)装置が含まれるシステム概念図を示す。コンピューターを以下の各部として機能させるための行動認識(学習)プログラムをコンピューターにインストールすることで本行動認識(学習)装置が構成される。本実施形態において、対象は人間であり、「対象の行動が時系列に記録されたデータ」は動画像データである。前処理部11に動画像データ12が入力される。学習時は前処理部11に訓練データである動画像データ12が入力される。人間が観察するなどして得られた正解行動を与える教師あり学習などを行う。
この動画像データ12の各フレームから行動の特徴量13を抽出して特徴量の時系列データ(以下「特徴量列」)14を生成する。特徴量列14が認識部15に入力される。
行動認識(学習)装置は、特徴量列14を入力とし、入力された特徴量列14に基づき対象の行動(認識結果16)とその尤度17を出力する認識部15と、尤度17や特徴量などに基づき異なる行動に切り替わる行動の境界点を判定する行動境界判定部18により構成される。
この動画像データ12の各フレームから行動の特徴量13を抽出して特徴量の時系列データ(以下「特徴量列」)14を生成する。特徴量列14が認識部15に入力される。
行動認識(学習)装置は、特徴量列14を入力とし、入力された特徴量列14に基づき対象の行動(認識結果16)とその尤度17を出力する認識部15と、尤度17や特徴量などに基づき異なる行動に切り替わる行動の境界点を判定する行動境界判定部18により構成される。
認識部15は、特報量列14を時系列に追って各時点の行動を認識する。いま、認識対象としている時点を現認識時点とする。認識部15は、現認識時点に相当する動画像データに写っている行動とその尤度を求めることとなる。本実施形態では、フレーム単位で行動とその尤度を出力する。
特徴量はもっとも簡単には動画像の各フレームの画像そのものを用いる場合が考えられる。それ以外の特徴量としては、画像から抜き出したオプティカルフローや人物位置・姿勢、また時間情報などを用いる場合も考えられる。本実施形態では、1例として人物姿勢(関節点座標)を用いることとする。
特徴量はもっとも簡単には動画像の各フレームの画像そのものを用いる場合が考えられる。それ以外の特徴量としては、画像から抜き出したオプティカルフローや人物位置・姿勢、また時間情報などを用いる場合も考えられる。本実施形態では、1例として人物姿勢(関節点座標)を用いることとする。
従来の手法としては、通常、特徴量は現認識時点のフレームを起点として過去10フレーム分の特徴量をまとめて与えるなど、固定長の値を与えるか、過去の情報から切れ目なく全フレーム与える様な形が多かったが、本発明の手法では同じ行動が連続したフレームは常に同じ行動をしている=1行動として、学習・認識対象のフレームの行動を起点とした所定の行動数N分のフレームの特徴量を与える形とする。行動数Nには、現認識時点の行動も含まれるので、遡る過去の行動数としては(N-1)である。
行動数のカウント方法を、上述した”座る” ”歩行” ”食事”を例にして説明する。”歩行”を区別しない同一行動とする場合は、”歩行”が数フレームに亘って連続しても、1行動としてカウントする。また、”食事”を区別しない同一行動とする場合は、”食事” が数フレームに亘って連続しても、1行動としてカウントする。しかし、”食事”を例えば”手に持った食器から食物を取り上げて口に運ぶ行動(行動A)”と ”テーブル上の食器から食物を取り上げて口に運ぶ行動(行動B)”とに細分化して行動ラベルを定義する場合には、行動Aの連続は1行動としてカウントし、行動Bの連続は1行動としてカウントする一方で、これら2つの行動が、例えば行動A、行動Bと連続すれば、2行動とカウントする。”歩行”及び ”座る”のそれぞれを区別しない同一行動とする場合は、”歩行” ”座る” ”歩行”と連続すれば3行動とカウントする。
図2は、認識部が行動認識のために用いる長さの特徴量列を示した概念図であり、図2Aは全フレームとする場合の比較例を示し、図2BはN=3とする場合の本発明例を示す。図中の数字は各行動ラベルを示し、数値を囲む矩形はその行動が連続する長さを示す。
図3は、認識部が行動認識のために用いる長さの特徴量列を枠で示した概念図であり、図3Aは枠301の長さを一定のフレーム数で固定とした比較例を示し、図3BはN=3とする場合の本発明例におけるフレーム数可変の枠302,303を示す。
認識部15は、行動境界判定部18が出力した境界点19を基準に、所定の行動数前に遡った時点(図2B,図3Bの例で2つ遡る)から現認識時点までの時系列に並ぶ複数の行動に相当する特徴量列20に基づき、現認識時点の行動を認識する。
図3は、認識部が行動認識のために用いる長さの特徴量列を枠で示した概念図であり、図3Aは枠301の長さを一定のフレーム数で固定とした比較例を示し、図3BはN=3とする場合の本発明例におけるフレーム数可変の枠302,303を示す。
認識部15は、行動境界判定部18が出力した境界点19を基準に、所定の行動数前に遡った時点(図2B,図3Bの例で2つ遡る)から現認識時点までの時系列に並ぶ複数の行動に相当する特徴量列20に基づき、現認識時点の行動を認識する。
図3Aの比較例の場合のように認識部が行動認識に用いる特徴量列が時間(フレーム数)で固定長の場合、人によって行動の速度は異なるため、人によって固定長の範囲内に行動を特定するのに有効な前行動情報が含まれない場合が考えられるが、図3Bの本発明例のように行動数を基準にし、認識部15が行動認識に用いる特徴量列のフレーム数を可変にすることで、現認識時点の行動につながる過去の行動の情報を十分に得ることが可能となる。
現認識時点の行動は過去の行動と強く関連付けられているといっても時間的に離れた情報は相対的に関係性が薄いと考えられ、図2Aの比較例のように全フレームを使った場合多くのノイズが含まれてしまい、ノイズ比の大きい過大なデータ量による負荷、認識精度の低下が懸念される。見るべき行動数をある程度限定することで、現認識時点の行動を推定するのに重要な情報のみを選択的に扱うことが可能になり、行動認識の高精度化及び効率化を図ることができる。
図2B,図3Bのように、N=3として、3行動分の特徴量列が得られるのは、時系列に並ぶ所定の行動数(図2B,図3Bの例で2つ)の認識が終わった後である。
動画の最初のフレームが入力されている時など行動数Nに入力フレーム数が満たない時のために、認識部15は、時系列に並ぶ所定の行動数(図2B,図3Bの例で2つ)の認識が終わる前は、現認識時点までの全ての特徴量列に基づき、現認識時点の行動を認識する。
動画の最初のフレームが入力されている時など行動数Nに入力フレーム数が満たない時のために、認識部15は、時系列に並ぶ所定の行動数(図2B,図3Bの例で2つ)の認識が終わる前は、現認識時点までの全ての特徴量列に基づき、現認識時点の行動を認識する。
本発明においてどの様に特徴量を用いて行動を認識するかは、機械学習手法の一種である、図4に概要図を示すRecurrent Neural Network(以下RNN)に図5に概要図を示すLong-Short Term Memory (以下LSTM)を組み合わせた思想に基づく。RNNはDeep Learningで用いられるニューラルネットワークベースの1手法であり、過去の入力による行動認識の結果を内部状態として保持することが可能であり、そのため前後の入力で関連がある言語音声分野や動画像解析で多く使われている手法である。ただしRNNではニューラルネットワークにおける勾配消失問題から直近の情報しか保持できないため、LSTMを組み合わせる形を採用する。LSTMは過去の情報をより長期間保持することが可能な技術であり、両者を組み合わせることで長期間の過去のデータを現在の入力の学習・認識に生かすことが可能となる。(RNNの詳細は非特許文献2を、LSTMの詳細は非特許文献3を参照。)
また、RNN+LSTMは内部状態をフラグによりリセットすることが可能である。リセットしない場合、それまでの全フレームの情報が内部的に保持される形となるが、リセットすると内部状態は初期化されるため、過去の入力はないものと扱われる。そのため、本実施形態では、行動境界判定部18の判定に基づき、この内部状態をリセットし再度特徴量を入力する処理が学習認識に用いる行動をリセットする処理として用いられる。
機械学習手法を用いているため、認識部15は認識の前に学習を行う必要がある。学習は正解行動が既知の動画像データを入力として、各行動を区別するために有効な特徴量が何かを学習していく。認識時は学習によって作られた処理に基づいて認識を行う。
また、RNN+LSTMは内部状態をフラグによりリセットすることが可能である。リセットしない場合、それまでの全フレームの情報が内部的に保持される形となるが、リセットすると内部状態は初期化されるため、過去の入力はないものと扱われる。そのため、本実施形態では、行動境界判定部18の判定に基づき、この内部状態をリセットし再度特徴量を入力する処理が学習認識に用いる行動をリセットする処理として用いられる。
機械学習手法を用いているため、認識部15は認識の前に学習を行う必要がある。学習は正解行動が既知の動画像データを入力として、各行動を区別するために有効な特徴量が何かを学習していく。認識時は学習によって作られた処理に基づいて認識を行う。
行動数に応じた入力について、学習時は行動の境界が既知であるため行動数に応じてリセットを行えばよいが、認識時は事前に行動が未知であり同様のことができないため、行動境界判定部18が必要となる。
なお、認識に用いる手法はLSTMに限定されない。
なお、認識に用いる手法はLSTMに限定されない。
認識部15は認識結果16として、対象の各行動の尤度を出力する。たとえば10種の行動を認識する場合、10個の行動それぞれについて、尤度が算出され、最も尤度が高い行動を認識結果16として出力する。
一方、行動境界判定部18は、認識中の行動の切れ目となる境界点を判定し、認識部15へ入力する。一般的には認識結果16が異なる行動に変わった場合(1位の入れ替わりがあった場合)、そこを境界点とすれば良いと考えられるが、その場合、認識結果16が異なる行動に変わった後に初めて境界点の判定が行われるため判定が遅れてしまう。特に行動間の境界がわかりにくい場合、判定の遅れはより大きくなることが予想される。これらの事象を押さえる手として、行動認識の各行動の尤度情報を用いる方法が考えられる。
一方、行動境界判定部18は、認識中の行動の切れ目となる境界点を判定し、認識部15へ入力する。一般的には認識結果16が異なる行動に変わった場合(1位の入れ替わりがあった場合)、そこを境界点とすれば良いと考えられるが、その場合、認識結果16が異なる行動に変わった後に初めて境界点の判定が行われるため判定が遅れてしまう。特に行動間の境界がわかりにくい場合、判定の遅れはより大きくなることが予想される。これらの事象を押さえる手として、行動認識の各行動の尤度情報を用いる方法が考えられる。
ひとつには、行動の認識結果の最大尤度とそれ以外の尤度との差が所定の値よりも小さい場合に行動終了と判定する方法が考えられる。最大尤度と他の尤度の差が縮んだり最大尤度が低下したりしているということは行動の移り変わりが発生している可能性が高いため、こうした判定は有効である。
例えば、行動境界判定部18は、図6に示すように尤度が高い順位が1位と2位の差601が所定の閾値以下となった時点を境界点と判定する。すなわち、図6において、1-6フレームで行動0が1位であるが、1位が行動0から行動2に切り替わる7フレーム目やそれ以降で判定せず、6フレーム目の1位と2位の差601が所定の閾値以下となった時点で判定を下すことで早期に判定する。これにより、7フレーム目では、認識部15は行動認識に用いる特徴量列を新たな境界点から遡った行動数の範囲に更新し、行動認識の精度を向上する。
例えば、行動境界判定部18は、図6に示すように尤度が高い順位が1位と2位の差601が所定の閾値以下となった時点を境界点と判定する。すなわち、図6において、1-6フレームで行動0が1位であるが、1位が行動0から行動2に切り替わる7フレーム目やそれ以降で判定せず、6フレーム目の1位と2位の差601が所定の閾値以下となった時点で判定を下すことで早期に判定する。これにより、7フレーム目では、認識部15は行動認識に用いる特徴量列を新たな境界点から遡った行動数の範囲に更新し、行動認識の精度を向上する。
また、行動の認識結果には、ある程度の誤判定やノイズが混じることが予想され、認識結果の1つの瞬間値に基づき境界判定を行った場合、図7に示すように連続行動中に1フレームだけ別の行動が誤認識されただけで観測された行動数が1から3に大きく変化してしまう。すなわち、図7において、1-10フレームの行動0と、12-30フレームの行動0との間に11フレーム目で行動1が1位になっただけで、3行動がカウントされてしまう。この場合、11フレーム目の行動1はノイズとしてカットし、行動0が続いていると判定すべきである。
こうした場合に対応するため、行動境界判定部18は、現認識時点のフレームを起点に所定の範囲の各行動尤度値の平均や中央値などの統計量に基づき、この平均や中央値などの統計量が入れ替わった段階で行動が切り替わったと判定する方法が考えられる。また、最大尤度を示す行動が変化した後所定の時間(フレーム数)内で最大尤度を示す行動が変化しなかった場合行動終了と判定する方法も考えられる。この場合は統計量として「最頻値」を用いれば実施できる。
こうした場合に対応するため、行動境界判定部18は、現認識時点のフレームを起点に所定の範囲の各行動尤度値の平均や中央値などの統計量に基づき、この平均や中央値などの統計量が入れ替わった段階で行動が切り替わったと判定する方法が考えられる。また、最大尤度を示す行動が変化した後所定の時間(フレーム数)内で最大尤度を示す行動が変化しなかった場合行動終了と判定する方法も考えられる。この場合は統計量として「最頻値」を用いれば実施できる。
また、尤度を使わない場合も考えられる。例えば寝るという行為は一般にベッド上で行われる。そのため、ベッドから離れた瞬間のように、位置情報に基づき異なる行動に切り替わる行動の境界点を判定する。例えば、浴室になどの特定の範囲に入った/出たという位置情報に基づき異なる行動に切り替わる行動の境界点を判定する。
この位置情報は、図1に示した動画像データ12を解析して得られる対象の位置情報としてもよいし、別途、位置検出部21から入力されるものとしてもよい。位置検出部21は、動画像データ12に基づくものではなく、対象の位置を検出するセンシングシステムと連携するものである。これにより、入浴などの行動を行う場所が限定されている場合に認識精度を向上することができる。
この位置情報は、図1に示した動画像データ12を解析して得られる対象の位置情報としてもよいし、別途、位置検出部21から入力されるものとしてもよい。位置検出部21は、動画像データ12に基づくものではなく、対象の位置を検出するセンシングシステムと連携するものである。これにより、入浴などの行動を行う場所が限定されている場合に認識精度を向上することができる。
尤度を使わない別の例としては、所定のフレーム数以上同じ行動が続いた場合に境界点があったと判定する方法も考えられる。これは、あまりに長い期間同じ行動が続いている場合、その前の行動と次の行動の関連性は弱まっていると考えられるためである。
図8に、認識部15が行動認識のために用いる長さの特徴量列を枠801、803で示し、現認識時点を指針802,804で示す。
図8Aに示す現認識時点802で行動境界判定部18が行動の境界を判定した場合、次の行動の認識(図8Bに示す現認識時点804における認識)では、認識部15は図8Bに示す枠803のように認識に用いる過去の行動を1行動分ずらして認識を行う。
図8に、認識部15が行動認識のために用いる長さの特徴量列を枠801、803で示し、現認識時点を指針802,804で示す。
図8Aに示す現認識時点802で行動境界判定部18が行動の境界を判定した場合、次の行動の認識(図8Bに示す現認識時点804における認識)では、認識部15は図8Bに示す枠803のように認識に用いる過去の行動を1行動分ずらして認識を行う。
以上のように本発明の行動認識装置は、対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部を有する行動認識装置において、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
前記行動認識装置において好ましくは、前記認識部は、時系列に並ぶ前記所定の行動数の認識が終わる前は、現認識時点までの全ての前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また前記行動認識装置において好ましくは、異なる行動に切り替わる行動の境界点を判定する行動境界判定部を有し、前記認識部は、前記行動境界判定部が出力した境界点を基準に、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また前記行動認識装置において好ましくは、前記行動境界判定部は、前記認識部が出力する行動の尤度情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また前記行動認識装置において好ましくは、前記行動境界判定部は、尤度が高い順位が1位と2位の差が所定の閾値以下となった時点を前記境界点と判定することを特徴とする。
また前記行動認識装置において好ましくは、前記行動境界判定部は、所定長さの時間内に前記認識部から複数回出力される尤度情報の統計量に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また前記行動認識装置において好ましくは、前記行動境界判定部は、前記対象の位置情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また本発明の行動学習装置は、対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部を有する行動学習装置において、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また本発明の行動認識プログラムは、対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部としてコンピューターを機能させるための行動認識プログラムにおいて、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
前記行動認識プログラムにおいて好ましくは、前記認識部は、時系列に並ぶ前記所定の行動数の認識が終わる前は、現認識時点までの全ての前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また前記行動認識プログラムにおいて好ましくは、異なる行動に切り替わる行動の境界点を判定する行動境界判定部として前記コンピューターを機能させ、前記認識部は、前記行動境界判定部が出力した境界点を基準に、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
また前記行動認識プログラムにおいて好ましくは、前記行動境界判定部は、前記認識部が出力する行動の尤度情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また前記行動認識プログラムにおいて好ましくは、前記行動境界判定部は、尤度が高い順位が1位と2位の差が所定の閾値以下となった時点を前記境界点と判定することを特徴とする。
また前記行動認識プログラムにおいて好ましくは、前記行動境界判定部は、所定長さの時間内に前記認識部から複数回出力される尤度情報の統計量に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また前記行動認識プログラムにおいて好ましくは、前記行動境界判定部は、前記対象の位置情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする。
また本発明の行動学習プログラムは、対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部としてコンピューターを機能させるための行動学習プログラムにおいて、前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする。
本発明は、コンピューター等による人等の対象の行動認識に利用することができる。
11 前処理部
12 動画像データ
13 特徴量
14 特徴量列
15 認識部
16 認識結果
17 尤度
18 行動境界判定部
19 境界点
20 特徴量列
21 位置検出部
12 動画像データ
13 特徴量
14 特徴量列
15 認識部
16 認識結果
17 尤度
18 行動境界判定部
19 境界点
20 特徴量列
21 位置検出部
Claims (16)
- 対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部を有する行動認識装置において、
前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする行動認識装置。 - 前記認識部は、時系列に並ぶ前記所定の行動数の認識が終わる前は、現認識時点までの全ての前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする請求項1に記載の行動認識装置。
- 異なる行動に切り替わる行動の境界点を判定する行動境界判定部を有し、
前記認識部は、前記行動境界判定部が出力した境界点を基準に、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする請求項1又は請求項2に記載の行動認識装置。 - 前記行動境界判定部は、前記認識部が出力する行動の尤度情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項3に記載の行動認識装置。
- 前記行動境界判定部は、尤度が高い順位が1位と2位の差が所定の閾値以下となった時点を前記境界点と判定することを特徴とする請求項4に記載の行動認識装置。
- 前記行動境界判定部は、所定長さの時間内に前記認識部から複数回出力される尤度情報の統計量に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項4又は請求項5に記載の行動認識装置。
- 前記行動境界判定部は、前記対象の位置情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項3に記載の行動認識装置。
- 対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部を有する行動学習装置において、
前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする行動学習装置。 - 対象の行動が時系列に記録されたデータから抽出された前記対象の行動の特徴量の時系列データに基づき当該行動を認識する認識部としてコンピューターを機能させるための行動認識プログラムにおいて、
前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする行動認識プログラム。 - 前記認識部は、時系列に並ぶ前記所定の行動数の認識が終わる前は、現認識時点までの全ての前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする請求項9に記載の行動認識プログラム。
- 異なる行動に切り替わる行動の境界点を判定する行動境界判定部として前記コンピューターを機能させ、
前記認識部は、前記行動境界判定部が出力した境界点を基準に、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする請求項9又は請求項10に記載の行動認識プログラム。 - 前記行動境界判定部は、前記認識部が出力する行動の尤度情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項11に記載の行動認識プログラム。
- 前記行動境界判定部は、尤度が高い順位が1位と2位の差が所定の閾値以下となった時点を前記境界点と判定することを特徴とする請求項12に記載の行動認識プログラム。
- 前記行動境界判定部は、所定長さの時間内に前記認識部から複数回出力される尤度情報の統計量に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項12又は請求項13に記載の行動認識プログラム。
- 前記行動境界判定部は、前記対象の位置情報に基づき、異なる行動に切り替わる行動の境界点を判定することを特徴とする請求項11に記載の行動認識プログラム。
- 対象の行動が時系列に記録された訓練データから抽出された対象の行動の特徴量の時系列データに基づき当該行動を認識するとともに学習する認識部としてコンピューターを機能させるための行動学習プログラムにおいて、
前記認識部は、前記特徴量の時系列データを時系列に追って各時点の行動を認識し、区別しない同一行動の連続は1行動として、時系列に並ぶ所定の行動数の認識が終わった後は、前記所定の行動数前に遡った時点から現認識時点までの時系列に並ぶ複数の行動に相当する前記特徴量の時系列データに基づき、現認識時点の行動を認識することを特徴とする行動学習プログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018503027A JPWO2017150211A1 (ja) | 2016-03-03 | 2017-02-17 | 行動認識装置及び行動学習装置並びに行動認識プログラム及び行動学習プログラム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2016040656 | 2016-03-03 | ||
| JP2016-040656 | 2016-03-03 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017150211A1 true WO2017150211A1 (ja) | 2017-09-08 |
Family
ID=59742827
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2017/005850 Ceased WO2017150211A1 (ja) | 2016-03-03 | 2017-02-17 | 行動認識装置及び行動学習装置並びに行動認識プログラム及び行動学習プログラム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JPWO2017150211A1 (ja) |
| WO (1) | WO2017150211A1 (ja) |
Cited By (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108804995A (zh) * | 2018-03-23 | 2018-11-13 | 李春莲 | 设备分布图像识别平台 |
| KR20190061538A (ko) * | 2017-11-28 | 2019-06-05 | 영남대학교 산학협력단 | 멀티 인식모델의 결합에 의한 행동패턴 인식방법 및 장치 |
| JP2019096252A (ja) * | 2017-11-28 | 2019-06-20 | Kddi株式会社 | 撮影映像から人の行動を表すコンテキストを推定するプログラム、装置及び方法 |
| JP2020009141A (ja) * | 2018-07-06 | 2020-01-16 | 株式会社 日立産業制御ソリューションズ | 機械学習装置及び方法 |
| JP2020087437A (ja) * | 2018-11-27 | 2020-06-04 | 富士ゼロックス株式会社 | カメラシステムを使用した、ユーザの身体部分によって実行されるタスクの完了の評価のための方法、プログラム、及びシステム |
| JP2021022323A (ja) * | 2019-07-30 | 2021-02-18 | Necソリューションイノベータ株式会社 | 行動推定装置、行動推定方法およびプログラム |
| JP2021071773A (ja) * | 2019-10-29 | 2021-05-06 | 株式会社エクサウィザーズ | 動作評価装置、動作評価方法、動作評価システム |
| WO2021125521A1 (ko) * | 2019-12-16 | 2021-06-24 | 연세대학교 산학협력단 | 순차적 특징 데이터 이용한 행동 인식 방법 및 그를 위한 장치 |
| CN114241354A (zh) * | 2021-11-19 | 2022-03-25 | 上海浦东发展银行股份有限公司 | 仓库人员行为识别方法、装置、计算机设备、存储介质 |
| JP2023094083A (ja) * | 2021-12-23 | 2023-07-05 | 富士通株式会社 | 情報処理プログラム、情報処理方法、および情報処理装置 |
| JPWO2024062882A1 (ja) * | 2022-09-20 | 2024-03-28 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005215927A (ja) * | 2004-01-29 | 2005-08-11 | Mitsubishi Heavy Ind Ltd | 行動認識システム |
| JP2005258830A (ja) * | 2004-03-11 | 2005-09-22 | Yamaguchi Univ | 人物行動理解システム |
| JP2011215951A (ja) * | 2010-03-31 | 2011-10-27 | Toshiba Corp | 行動判定装置、方法及びプログラム |
-
2017
- 2017-02-17 WO PCT/JP2017/005850 patent/WO2017150211A1/ja not_active Ceased
- 2017-02-17 JP JP2018503027A patent/JPWO2017150211A1/ja active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005215927A (ja) * | 2004-01-29 | 2005-08-11 | Mitsubishi Heavy Ind Ltd | 行動認識システム |
| JP2005258830A (ja) * | 2004-03-11 | 2005-09-22 | Yamaguchi Univ | 人物行動理解システム |
| JP2011215951A (ja) * | 2010-03-31 | 2011-10-27 | Toshiba Corp | 行動判定装置、方法及びプログラム |
Cited By (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102440385B1 (ko) * | 2017-11-28 | 2022-09-05 | 영남대학교 산학협력단 | 멀티 인식모델의 결합에 의한 행동패턴 인식방법 및 장치 |
| KR20190061538A (ko) * | 2017-11-28 | 2019-06-05 | 영남대학교 산학협력단 | 멀티 인식모델의 결합에 의한 행동패턴 인식방법 및 장치 |
| JP2019096252A (ja) * | 2017-11-28 | 2019-06-20 | Kddi株式会社 | 撮影映像から人の行動を表すコンテキストを推定するプログラム、装置及び方法 |
| CN108804995B (zh) * | 2018-03-23 | 2019-04-16 | 新昌县夙凡软件科技有限公司 | 盥洗设备分布图像识别平台 |
| CN108804995A (zh) * | 2018-03-23 | 2018-11-13 | 李春莲 | 设备分布图像识别平台 |
| JP2020009141A (ja) * | 2018-07-06 | 2020-01-16 | 株式会社 日立産業制御ソリューションズ | 機械学習装置及び方法 |
| JP7392348B2 (ja) | 2018-11-27 | 2023-12-06 | 富士フイルムビジネスイノベーション株式会社 | カメラシステムを使用した、ユーザの身体部分によって実行されるタスクの完了の評価のための方法、プログラム、及びシステム |
| JP2020087437A (ja) * | 2018-11-27 | 2020-06-04 | 富士ゼロックス株式会社 | カメラシステムを使用した、ユーザの身体部分によって実行されるタスクの完了の評価のための方法、プログラム、及びシステム |
| JP2021022323A (ja) * | 2019-07-30 | 2021-02-18 | Necソリューションイノベータ株式会社 | 行動推定装置、行動推定方法およびプログラム |
| JP7368045B2 (ja) | 2019-07-30 | 2023-10-24 | Necソリューションイノベータ株式会社 | 行動推定装置、行動推定方法およびプログラム |
| JP2021071773A (ja) * | 2019-10-29 | 2021-05-06 | 株式会社エクサウィザーズ | 動作評価装置、動作評価方法、動作評価システム |
| WO2021125521A1 (ko) * | 2019-12-16 | 2021-06-24 | 연세대학교 산학협력단 | 순차적 특징 데이터 이용한 행동 인식 방법 및 그를 위한 장치 |
| KR20210076659A (ko) * | 2019-12-16 | 2021-06-24 | 연세대학교 산학협력단 | 순차적 특징 데이터 이용한 행동 인식 방법 및 그를 위한 장치 |
| KR102334388B1 (ko) * | 2019-12-16 | 2021-12-01 | 연세대학교 산학협력단 | 순차적 특징 데이터 이용한 행동 인식 방법 및 그를 위한 장치 |
| CN114241354A (zh) * | 2021-11-19 | 2022-03-25 | 上海浦东发展银行股份有限公司 | 仓库人员行为识别方法、装置、计算机设备、存储介质 |
| JP2023094083A (ja) * | 2021-12-23 | 2023-07-05 | 富士通株式会社 | 情報処理プログラム、情報処理方法、および情報処理装置 |
| JP7769203B2 (ja) | 2021-12-23 | 2025-11-13 | 富士通株式会社 | 情報処理プログラム、情報処理方法、および情報処理装置 |
| JPWO2024062882A1 (ja) * | 2022-09-20 | 2024-03-28 | ||
| WO2024062882A1 (ja) * | 2022-09-20 | 2024-03-28 | 株式会社Ollo | プログラム、情報処理方法、及び情報処理装置 |
| JP7570151B2 (ja) | 2022-09-20 | 2024-10-21 | 株式会社Ollo | プログラム、情報処理方法、及び情報処理装置 |
| US12387530B1 (en) | 2022-09-20 | 2025-08-12 | Ollo, Inc. | Program, information processing method, and information processing apparatus |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2017150211A1 (ja) | 2018-12-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2017150211A1 (ja) | 行動認識装置及び行動学習装置並びに行動認識プログラム及び行動学習プログラム | |
| JP6658331B2 (ja) | 行動認識装置及び行動認識プログラム | |
| Aminikhanghahi et al. | Enhancing activity recognition using CPD-based activity segmentation | |
| US11551103B2 (en) | Data-driven activity prediction | |
| Junker et al. | Gesture spotting with body-worn inertial sensors to detect user activities | |
| EP4098182B1 (en) | Machine-learning based gesture recognition with framework for adding user-customized gestures | |
| Pauwels et al. | Sensor networks for ambient intelligence | |
| WO2009090584A2 (en) | Method and system for activity recognition and its application in fall detection | |
| JP2010213782A (ja) | 行動認識方法、装置及びプログラム | |
| Kheratkar et al. | Gesture controlled home automation using CNN | |
| Devanne et al. | Recognition of activities of daily living via hierarchical long-short term memory networks | |
| CN112801000A (zh) | 一种基于多特征融合的居家老人摔倒检测方法及系统 | |
| Parate et al. | Detecting eating and smoking behaviors using smartwatches | |
| Kang et al. | Beyond superficial emotion recognition: Modality-adaptive emotion recognition system | |
| JP6274114B2 (ja) | 制御方法、制御プログラム、および制御装置 | |
| Diete et al. | Vision and acceleration modalities: Partners for recognizing complex activities | |
| Gharasuie et al. | Real-time dynamic hand gesture recognition using hidden Markov models | |
| Mahmood et al. | Contextual anomaly detection based video surveillance system | |
| KR101836742B1 (ko) | 제스쳐를 판단하는 장치 및 방법 | |
| Golda Jeyasheeli et al. | Deep learning based indian sign language words identification system | |
| Esther et al. | Predicting human activity-state of the art | |
| Georgakopoulos et al. | On-line fall detection via mobile accelerometer data | |
| Li et al. | Human Activity Recognition in Free Living Using a Wearable Sensor: A Two-Stage Approach | |
| Dharwarkar et al. | Enhancing temporal classification of aar parameters in eeg single-trial analysis for brain-computer interfacing | |
| CN113887696A (zh) | 用于创建机器学习系统的方法和设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| ENP | Entry into the national phase |
Ref document number: 2018503027 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17759680 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17759680 Country of ref document: EP Kind code of ref document: A1 |