WO2024257321A1 - ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 - Google Patents
ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 Download PDFInfo
- Publication number
- WO2024257321A1 WO2024257321A1 PCT/JP2023/022338 JP2023022338W WO2024257321A1 WO 2024257321 A1 WO2024257321 A1 WO 2024257321A1 JP 2023022338 W JP2023022338 W JP 2023022338W WO 2024257321 A1 WO2024257321 A1 WO 2024257321A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- person
- gesture
- detection unit
- mask
- determination unit
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
Definitions
- Patent Document 1 describes an information processing system including an authentication means for performing biometric authentication using biometric information of the first user, an acquisition means for acquiring information on the behavioral history of the first user when the biometric authentication is successful, a storage means for storing information on the behavioral history, a generation means for generating risk information indicating a risk related to the first user based on the stored information on the behavioral history, and an output means for outputting the risk information of the first user in response to a request from a second user different from the first user, in which the authentication means (biometric authentication unit) particularly includes a motion information acquisition unit, which is configured to detect gestures around a shield (e.g., a mask) that partially covers the face of the first user when performing biometric authentication of the first user, and acquire motion information corresponding to the gesture.
- the motion information acquisition unit includes, for example, a camera, and may detect
- Patent Document 1 describes gesture detection when a person's (first user's) face is covered with a mask, i.e., when the person is wearing a mask, but does not describe gesture detection when the person is not wearing a mask.
- a person eats or drinks, they remove their mask and bring their hand holding food or drink to their mouth.
- the action of the person bringing their hand to their mouth is not intended to be any kind of gesture, so it is not desirable for this action to be detected as a gesture.
- Patent Document 1 does not describe gesture detection when the person is not wearing a mask, so there is a risk that the system described in Patent Document 1 may erroneously detect actions made when a person is eating or drinking as gestures.
- This disclosure has been made to solve the problems described above, and aims to provide a gesture detection device that can prevent erroneous detection of actions taken when a person is eating or drinking as a gesture.
- the gesture detection device includes a hand candidate detection unit that detects hand candidates that are candidates for the hand of a person based on an image captured of the person, a gesture detection unit that detects a gesture of the person based on the hand candidates detected by the hand candidate detection unit, a mask determination unit that determines whether or not the person is wearing a mask based on the captured image, an occlusion detection unit that determines whether or not the person's mouth is covered based on facial information of the person obtained from the captured image when the mask determination unit determines that the person is not wearing a mask, and a determination unit that rejects the gesture detected by the gesture detection unit when the occlusion detection unit determines that the person's mouth is covered.
- FIG. 1 is a diagram showing an example of the configuration of an occupant monitoring system including a gesture detection device according to a first embodiment
- 4 is a flowchart for explaining an example of an operation of the gesture detection device according to the first embodiment
- 3A and 3B are diagrams illustrating an example of a hardware configuration of the gesture detection device according to the first embodiment.
- FIG. 1 is a diagram showing an example of the configuration of an occupant monitoring system 100 including a gesture detection device 1 according to embodiment 1.
- the gesture detection device 1 according to embodiment 1 is installed in the occupant monitoring system 100.
- the imaging device 110 is composed of, for example, a camera installed in a vehicle (not shown) and captures images of the interior of the vehicle, including the vehicle occupants, in a time-series manner.
- the imaging device 110 sequentially outputs images obtained by capturing images (hereinafter also referred to as “captured images") to the gesture detection device 1.
- the control device 120 executes a predetermined control based on the gesture indicated by the information output from the gesture detection device 1. For example, if the gesture indicated by the information output from the gesture detection device 1 is a gesture for operating in-vehicle equipment such as an air conditioner or audio, the control device 120 executes adjustment of the air conditioner temperature, adjustment of the audio volume, etc. based on the gesture.
- the in-vehicle equipment is not limited to an air conditioner and audio.
- the gesture detection device 1 includes an image acquisition unit 10, a face information acquisition unit 11, a hand candidate detection unit 12, a gesture detection unit 13, a mask determination unit 14, an occlusion detection unit 15, and a determination unit 16.
- the image acquisition unit 10 acquires the captured image output from the imaging device 110.
- the image acquisition unit 10 outputs the acquired captured image to the face information acquisition unit 11, the hand candidate detection unit 12, and the mask determination unit 14.
- the face information acquisition unit 11 acquires information about the face of the occupant (hereinafter also referred to as "face information") based on the captured image output from the image acquisition unit 10.
- face information information about the face of the occupant
- the face information acquisition unit 11 outputs the acquired face information to the occlusion detection unit 15.
- the facial information of the occupant is, for example, information indicating an image obtained by cutting out the area in which the occupant's face is captured from the captured image output from the image acquisition unit 10.
- the facial information of the occupant may include information indicating parts of the face, such as the eyes, eyebrows, nose, mouth, forehead, cheeks, or chin, and information indicating the positions of those parts in the captured image.
- the hand candidate detection unit 12 detects hand candidates that are candidates for the occupant's hand based on the captured image output from the image acquisition unit 10.
- the detected hand candidates also include the position and shape of the hand candidate.
- the hand candidate detection unit 12 outputs information indicating the detected hand candidates to the gesture detection unit 13.
- the hand candidate detection unit 12 may also detect hand candidates using, for example, a machine learning model.
- the machine learning model may be a trained model that has been trained to output a result of inferring hand candidates of an occupant in response to an input of an image of the occupant.
- the gesture detection unit 13 detects the gesture of the occupant based on the position and shape of the hand candidate indicated by the information output from the hand candidate detection unit 12.
- the gesture detection unit 13 outputs information indicating the detected gesture of the occupant to the determination unit 16.
- the gesture detection unit 13 detects the gestures of the occupant using, for example, a machine learning model.
- the machine learning model is a trained model that has been trained to, in response to input of information indicating potential hand positions of the occupant, output a result of inferring a gesture corresponding to the potential hand positions indicated by the information.
- the mask determination unit 14 determines whether or not an occupant is wearing a mask, for example, by using a machine learning model.
- the machine learning model is a trained model that has been trained to output an inference result as to whether or not an occupant is wearing a mask in response to an input of an image of the occupant.
- the mask determination unit 14 may determine whether or not an occupant is wearing a mask based on the facial information of the occupant acquired by the facial information acquisition unit 11. In this case, the mask determination unit 14 may use, as the trained model, a trained model that has been trained to output a result of inferring whether or not an occupant is wearing a mask in response to input of the occupant's facial information.
- the occlusion detection unit 15 When the information output from the mask determination unit 14 indicates that the occupant is not wearing a mask, the occupant's mouth is determined to be covered or not based on the occupant's facial information output from the facial information acquisition unit 11.
- the occlusion detection unit 15 outputs information indicating the result of the determination to the determination unit 16. Note that the occlusion detection unit 15 is only required to make the above determination when the information output from the mask determination unit 14 indicates that the occupant is not wearing a mask, and is not required to make the above determination when the information output from the mask determination unit 14 indicates that the occupant is wearing a mask.
- the occlusion detection unit 15 determines whether or not the occupant's mouth is occluded by using, for example, a machine learning model.
- the machine learning model is a trained model that has been trained to output a result of inferring whether or not the occupant's mouth is occluded in response to input of, for example, facial information of the occupant.
- the occlusion detection unit 15 has been described as determining whether or not the mouth of an occupant is occluded based on the facial information of the occupant output from the facial information acquisition unit 11.
- the occlusion detection unit 15 is not limited to this, and may determine whether or not part or all of the face of the occupant is occluded based on, for example, the facial information of the occupant output from the facial information acquisition unit 11. In this case, it is sufficient for the occupancy detection unit 15 to be able to determine whether or not at least the mouth of the occupant is occluded based on, for example, the facial information of the occupant output from the facial information acquisition unit 11.
- the determination unit 16 recognizes the gesture indicated by the information output from the gesture detection unit 13 as the gesture of the occupant.
- the determination unit 16 recognizes the gesture indicated by the information output from the gesture detection unit 13 as the gesture of the occupant.
- the determination unit 16 rejects the gesture indicated by the information output from the gesture detection unit 13 without recognizing it as a gesture by the occupant.
- the determination unit 16 may calculate the distance between the hand candidate and the occupant's mouth based on the information indicating the hand candidate detected by the hand candidate detection unit 12 and the occupant's face information acquired by the face information acquisition unit 11 before determining whether to reject or recognize the gesture based on the information output from the occlusion detection unit 15.
- the determination unit 16 acquires a detection box of hand candidates based on the information indicating hand candidates detected by the hand candidate detection unit 12, and acquires a detection box of the occupant's mouth based on the occupant's face information acquired by the face information acquisition unit 11.
- the determination unit 16 can acquire the detection box of the occupant's hand candidates and the detection box of the occupant's mouth using a known method. Then, the determination unit 16 calculates the minimum distance between the acquired detection box of the occupant's hand candidate and the detection box of the occupant's mouth as the distance between the hand candidate and the occupant's mouth.
- the determination unit 16 may recognize the gesture detected by the gesture detection unit 13 as a gesture of the occupant, regardless of the information output from the occlusion detection unit 15. Furthermore, the determination unit 16 may determine to reject or recognize the gesture based on the information output from the occlusion detection unit 15 only when the calculated distance is equal to or less than the predetermined threshold. In this way, the determination unit 16 can increase the opportunities to determine whether to reject or recognize the gesture by making a determination based on the distance between the hand candidate and the occupant's mouth prior to determining to reject or recognize the gesture based on the information output from the occlusion detection unit 15, thereby improving the accuracy of the determination regarding the rejection or recognition of the gesture.
- the determination unit 16 makes a determination based on the distance between a candidate hand and the occupant's mouth before determining whether to reject or recognize a gesture based on the information output from the occlusion detection unit 15.
- the image acquisition unit 10 acquires the captured image output from the imaging device 110 (step ST1).
- the image acquisition unit 10 outputs the acquired captured image to the face information acquisition unit 11, the hand candidate detection unit 12, and the mask determination unit 14.
- the face information acquisition unit 11 acquires information about the face of the occupant (face information) based on the captured image output from the image acquisition unit 10 (step ST2).
- the face information acquisition unit 11 outputs the acquired face information to the occlusion detection unit 15 and the determination unit 16.
- the hand candidate detection unit 12 detects hand candidates that are candidates for the occupant's hand based on the captured image output from the image acquisition unit 10 (step ST3).
- the hand candidate detection unit 12 outputs information indicating the detected hand candidates to the gesture detection unit 13 and the determination unit 16.
- the gesture detection unit 13 detects the occupant's gesture based on the position and shape of the hand candidate indicated by the information output from the hand candidate detection unit 12 (step ST4).
- the gesture detection unit 13 outputs information indicating the detected occupant's gesture to the determination unit 16.
- the mask determination unit 14 determines whether or not the occupant is wearing a mask based on the captured image output from the image acquisition unit 10 (step ST5).
- the mask determination unit 14 outputs information indicating the result of the determination to the occlusion detection unit 15 and the determination unit 16.
- the determination unit 16 checks whether the information output from the mask determination unit 14 indicates that the occupant is wearing a mask, i.e., whether the mask determination unit 14 has determined that the occupant is wearing a mask (step ST6). As a result, if it is confirmed that the mask determination unit 14 has determined that the occupant is wearing a mask (step ST6; YES), the determination unit 16 does not reject the gesture detected by the gesture detection unit 13 in step ST4, but recognizes it as a gesture by the occupant (step ST7).
- the determination unit 16 determines that the mask determination unit 14 has determined that the occupant is not wearing a mask (step ST6; NO), it calculates the distance between the hand candidate and the occupant's mouth based on the information indicating the hand candidate detected by the hand candidate detection unit 12 and the occupant's face information acquired by the face information acquisition unit 11 (step ST8).
- step ST9 the determination unit 16 checks whether the calculated distance is equal to or less than a predetermined threshold. As a result, if the calculated distance is not equal to or less than the threshold (step ST9; NO), the process proceeds to step ST7, and the determination unit 16 does not reject the gesture detected by the gesture detection unit 13 in step ST4, but recognizes it as a gesture by an occupant. On the other hand, if the calculated distance is equal to or less than the threshold (step ST9; YES), the process proceeds to step ST10.
- step ST10 the occlusion detection unit 15 determines whether or not the occupant's mouth is occluded based on the facial information of the occupant output from the facial information acquisition unit 11, and outputs information indicating the result of the determination to the determination unit 16.
- the determination unit 16 then checks whether or not the information output from the occlusion detection unit 15 indicates that the occupant's mouth is occluded (step ST10).
- the determination unit 16 rejects the gesture indicated by the information output from the gesture detection unit 13 in step ST4 without recognizing it as a gesture by the occupant (step ST11).
- the determination unit 16 confirms that the information output from the occlusion detection unit 15 indicates that the occupant's mouth is not occluded (step ST10; NO), it does not reject the gesture detected by the gesture detection unit 13 in step ST4, but recognizes it as a gesture by the occupant (step ST7).
- the determination unit 16 calculates the distance between the hand candidate and the occupant's mouth in steps ST8 to ST9 before determining in step ST10 whether to reject or recognize the gesture based on the information output from the occlusion detection unit 15, and performs a determination based on the calculated distance.
- the processing in steps ST8 to ST9 is not essential and may be omitted. In that case, if it is confirmed in step ST6 that the occupant is determined not to be wearing a mask, the processing may transition to step ST10. However, if the determination unit 16 performs the processing in steps ST8 to ST9, the opportunities to determine whether to reject or recognize a gesture can be increased, and the accuracy of the determination regarding the rejection or recognition of a gesture can be improved.
- the occlusion detection unit 15 determines whether the occupant's mouth is covered or not. Then, in the gesture detection device 1, if the occlusion detection unit 15 determines that the occupant's mouth is covered, the determination unit 16 rejects the gesture detected by the gesture detection unit 13. This allows the gesture detection device 1 to suppress erroneous detection of the actions of the occupant when eating or drinking as a gesture.
- the processing circuit When the processing circuit is a CPU 52, the functions of the image acquisition unit 10, face information acquisition unit 11, hand candidate detection unit 12, gesture detection unit 13, mask determination unit 14, occlusion detection unit 15, and determination unit 16 are realized by software, firmware, or a combination of software and firmware.
- the software and firmware are written as programs and stored in memory 53.
- the processing circuit realizes the functions of each unit by reading and executing the programs recorded in memory 53.
- the gesture detection device 1 has a memory for storing a program that, when executed by the processing circuit, results in the execution of each step shown in FIG. 2, for example.
- Examples of memory 53 include non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable ROM), EEPROM (Electrically EPROM), magnetic disk, flexible disk, optical disk, compact disk, mini disk, or DVD (Digital Versatile Disc), etc.
- the functions of the image acquisition unit 10, face information acquisition unit 11, hand candidate detection unit 12, gesture detection unit 13, mask determination unit 14, occlusion detection unit 15, and determination unit 16 may be partially realized by dedicated hardware and partially realized by software or firmware.
- the image acquisition unit 10 may be realized by a processing circuit as dedicated hardware
- the face information acquisition unit 11, hand candidate detection unit 12, gesture detection unit 13, mask determination unit 14, occlusion detection unit 15, and determination unit 16 may be realized by the processing circuit reading and executing a program stored in memory 53.
- the processing circuitry can realize each of the above-mentioned functions through hardware, software, firmware, or a combination of these.
- the gesture detection device 1 includes a hand candidate detection unit 12 that detects hand candidates that are candidates for the hand of a person based on an image of the person, a gesture detection unit 13 that detects a gesture of the person based on the hand candidates detected by the hand candidate detection unit 12, a mask determination unit 14 that determines whether or not the person is wearing a mask based on the captured image, an occlusion detection unit 15 that determines whether or not the mouth of the person is covered based on information about the person's face obtained based on the captured image when the mask determination unit 14 determines that the person is not wearing a mask, and a determination unit 16 that rejects the gesture detected by the gesture detection unit 13 when the occlusion detection unit 15 determines that the person's mouth is covered.
- the determination unit 16 recognizes the gesture detected by the gesture detection unit 13 as a human gesture. In this way, the gesture detection device 1 according to the first embodiment can detect gestures made when a person is wearing a mask.
- the determination unit 16 recognizes the gesture detected by the gesture detection unit 13 as a human gesture. In this way, the gesture detection device 1 according to the first embodiment can detect actions other than eating and drinking actions as gestures when the person is not wearing a mask.
- the determination unit 16 calculates the distance between the hand candidate and the person's mouth based on the position of the hand candidate detected by the hand candidate detection unit 12 and the person's face information, and when the calculated distance is equal to or less than a threshold and the occlusion detection unit 15 determines that the person's mouth is occluded, the determination unit 16 rejects the gesture detected by the gesture detection unit 13, and when the calculated distance exceeds the threshold and when the occlusion detection unit 15 determines that the person's mouth is not occluded even if the calculated distance is equal to or less than the threshold, the determination unit 16 recognizes the gesture detected by the gesture detection unit 13 as a human gesture. This allows the gesture detection device 1 according to embodiment 1 to increase the opportunities to determine whether to reject or recognize a gesture, improving the determination accuracy.
- the mask determination unit 14 also uses a trained model that outputs whether or not a person is wearing a mask in response to an input of an image of a person, to determine whether or not the person is wearing a mask. This allows the gesture detection device 1 according to embodiment 1 to accurately determine whether or not a person is wearing a mask, improving the accuracy of the determination regarding rejection or recognition of gestures.
- the occlusion detection unit 15 also uses a trained model that outputs whether or not a person's mouth is occluded in response to input of information about a person's face to determine whether or not the person's mouth is occluded. This allows the gesture detection device 1 according to embodiment 1 to accurately determine whether or not a person's mouth is occluded, improving the accuracy of the determination regarding rejection or recognition of a gesture.
- the gesture detection unit 13 also detects human gestures using a trained model that receives information indicating the position and shape of hand candidates and outputs gestures corresponding to the position and shape of the hand candidates indicated by the information. This allows the gesture detection device 1 according to embodiment 1 to accurately detect human gestures.
- the occupant monitoring system 100 is configured to include the gesture detection device 1 described above, and is equipped with an imaging device 110 that images the vehicle occupant as a person, and a control device 120 that executes predetermined control based on the gesture recognized by the determination unit 16.
- the occupant monitoring system 100 according to the first embodiment can suppress erroneous control caused by erroneously detecting the actions of the vehicle occupant when eating or drinking as a gesture.
- the present disclosure makes it possible to prevent the erroneous detection of actions taken when a person is eating or drinking as a gesture, and is suitable for use in gesture detection devices, passenger monitoring systems, and gesture detection methods.
- gesture detection device 10 image acquisition unit, 11 face information acquisition unit, 12 hand candidate detection unit, 13 gesture detection unit, 14 mask determination unit, 15 occlusion detection unit, 16 determination unit, 51 processing circuit, 52 CPU, 53 memory, 100 occupant monitoring system, 110 imaging device, 120 control device.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Image Analysis (AREA)
Abstract
ジェスチャ検出装置(1)は、人を撮像した撮像画像に基づいて、当該人の手の候補である手候補を検出する手候補検出部(12)と、手候補検出部により検出された手候補に基づいて、人のジェスチャを検出するジェスチャ検出部(13)と、撮像画像に基づいて、人がマスクをしているか否かを判定するマスク判定部(14)と、マスク判定部により、人がマスクをしていないと判定された場合に、撮像画像に基づいて得られた人の顔の情報に基づいて、当該人の口が遮蔽されているか否かを判定する遮蔽検知部(15)と、遮蔽検知部により、人の口が遮蔽されていると判定された場合、ジェスチャ検出部により検出されたジェスチャを棄却する判定部(16)とを備えた。
Description
本開示は、ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法に関するものである。
従来、カメラなどの撮像装置で人を撮像した画像から、当該人のジェスチャを検出する技術が知られている。例えば特許文献1には、第1ユーザの生体情報を用いて生体認証を実行する認証手段と、生体認証が成功した場合に、第1ユーザの行動履歴に関する情報を取得する取得手段と、行動履歴に関する情報を蓄積する蓄積手段と、蓄積された行動履歴に関する情報に基づいて、第1ユーザに関するリスクを示すリスク情報を生成する生成手段と、第1ユーザとは異なる第2ユーザの要求に応じて、第1ユーザのリスク情報を出力する出力手段とを備えた情報処理システムにおいて、認証手段(生体認証部)は特に、動作情報取得部を備え、当該動作情報取得部は、第1ユーザの生体認証を実行する際に、第1ユーザの顔を部分的に覆う遮蔽物(例えばマスク等)の周辺におけるジェスチャを検知して、ジェスチャに対応する動作情報を取得可能に構成されることが記載されている。また、動作情報取得部は例えばカメラを含んでおり、マスクと重なる位置に指を持ってくるような動作をジェスチャとして検知してよいことも特許文献1に記載されている。
特許文献1には上記のように、人(第1ユーザ)の顔がマスクで覆われている、すなわち人がマスクをしている際のジェスチャ検出については記載されているが、人がマスクをしていない際のジェスチャ検出については記載されていない。一方、人は飲食をする場合、マスクを外した状態で、飲食物を持った手を口元に運ぶ動作を行う。このとき、人が口元に手を運ぶ動作は、何らかのジェスチャを意図したものではないため、この動作がジェスチャとして検出されるのは本来的には望ましくない。しかしながら、特許文献1には上記のように、人がマスクをしていない際のジェスチャ検出については記載されていないため、特許文献1記載のシステムでは、人が飲食をする際の動作をジェスチャとして誤検出するおそれがあった。
本開示は上記のような課題を解決するためになされたもので、人が飲食をする際の動作をジェスチャとして誤検出することを抑制可能なジェスチャ検出装置を得ることを目的とする。
本開示に係るジェスチャ検出装置は、人を撮像した撮像画像に基づいて、当該人の手の候補である手候補を検出する手候補検出部と、手候補検出部により検出された手候補に基づいて、人のジェスチャを検出するジェスチャ検出部と、撮像画像に基づいて、人がマスクをしているか否かを判定するマスク判定部と、マスク判定部により、人がマスクをしていないと判定された場合に、撮像画像に基づいて得られた人の顔の情報に基づいて、当該人の口が遮蔽されているか否かを判定する遮蔽検知部と、遮蔽検知部により、人の口が遮蔽されていると判定された場合、ジェスチャ検出部により検出されたジェスチャを棄却する判定部とを備えたものである。
本開示によれば、人が飲食をする際の動作をジェスチャとして誤検出することを抑制可能となる。
以下、本開示の実施の形態について、図面を参照しながら詳細に説明する。
実施の形態1.
実施の形態1.
図1は、実施の形態1に係るジェスチャ検出装置1を含む乗員監視システム100の構成例を示す図である。以下の説明では、実施の形態1に係るジェスチャ検出装置1が乗員監視システム100に搭載された場合を例に説明する。
乗員監視システム100は、例えば図1に示すように、撮像装置110と、ジェスチャ検出装置1と、制御装置120とを含んで構成される。
撮像装置110は、例えば車両(不図示)に設置されたカメラにより構成され、車両の乗員を含む車両の内部を時系列的に撮像する。撮像装置110は、撮像により得られた画像(以下、「撮像画像」ともいう。)をジェスチャ検出装置1に順次出力する。
ジェスチャ検出装置1は、撮像装置110から出力された撮像画像に基づき、人(ここでは車両の乗員)によるジェスチャを検出(認識)し、当該検出したジェスチャを示す情報を制御装置120に出力する。
制御装置120は、ジェスチャ検出装置1から出力された情報が示すジェスチャに基づいて、所定の制御を実行する。例えば、制御装置120は、ジェスチャ検出装置1から出力された情報が示すジェスチャが、エアコン及びオーディオ等の車載機器を操作するためのジェスチャであれば、当該ジェスチャに基づいてエアコンの温度調節、オーディオの音量調節等を実行する。ただし、車載機器は、エアコンおよびオーディオに限定されるものではない。
<ジェスチャ検出装置1>
ジェスチャ検出装置1は、例えば図1に示すように、画像取得部10と、顔情報取得部11と、手候補検出部12と、ジェスチャ検出部13と、マスク判定部14と、遮蔽検知部15と、判定部16とを含んで構成される。
ジェスチャ検出装置1は、例えば図1に示すように、画像取得部10と、顔情報取得部11と、手候補検出部12と、ジェスチャ検出部13と、マスク判定部14と、遮蔽検知部15と、判定部16とを含んで構成される。
画像取得部10は、撮像装置110から出力された撮像画像を取得する。画像取得部10は、取得した撮像画像を、顔情報取得部11、手候補検出部12、及びマスク判定部14に出力する。
顔情報取得部11は、画像取得部10から出力された撮像画像に基づいて、乗員の顔の情報(以下、「顔情報」ともいう。)を取得する。顔情報取得部11は、取得した顔情報を遮蔽検知部15に出力する。
なお、乗員の顔情報とは、例えば、画像取得部10から出力された撮像画像から、乗員の顔が撮影された範囲を切り出して得られた画像を示す情報である。乗員の顔情報には、目、眉、鼻、口、おでこ、頬又は顎等のような、顔に属する部位を示す情報、及び、当該部位の撮像画像における位置を示す情報等が含まれていてもよい。
手候補検出部12は、画像取得部10から出力された撮像画像に基づいて、乗員の手の候補である手候補を検出する。なお、検出される手候補には、当該手候補の位置および形状も含まれる。手候補検出部12は、検出した手候補を示す情報をジェスチャ検出部13に出力する。
なお、手候補検出部12は、例えば画像取得部10から出力された撮像画像における物体の形状のパターン(輝度分布の情報)と、予め定められた手の形状のパターンとをマッチングすることにより、つまりパターンマッチング処理により、乗員の手候補を検出する。検出対象の手の形状は、開いた状態の手の形状および閉じた状態の手の形状のうちいずれであってもよい。また、検出対象の手の形状は、例えば、数を示す手の形状、方向を示す手の形状、乗員の意思(OKまたはGoodなど)を示す手の形状等であってもよい。
また、手候補検出部12は、例えば機械学習モデルを用いて手候補の検出を行ってもよい。この場合、機械学習モデルは、例えば乗員を撮像した撮像画像の入力に対し、当該乗員の手候補を推論した結果を出力するよう学習された学習済みモデルであればよい。
ジェスチャ検出部13は、手候補検出部12から出力された情報が示す手候補の位置および形状に基づいて、乗員のジェスチャを検出する。ジェスチャ検出部13は、検出した乗員のジェスチャを示す情報を判定部16に出力する。
なお、ジェスチャ検出部13は、例えば機械学習モデルを用いて乗員のジェスチャを検出する。この場合、機械学習モデルは、例えば乗員の手候補を示す情報の入力に対し、当該情報が示す手候補に対応するジェスチャを推論した結果を出力するよう学習された学習済みモデルである。
マスク判定部14は、画像取得部10から出力された撮像画像に基づいて、乗員がマスクをしているか否かを判定する。マスク判定部14は、判定した結果を示す情報を遮蔽検知部15及び判定部16に出力する。
なお、マスク判定部14は、例えば機械学習モデルを用いて、乗員がマスクをしているか否かを判定する。この場合、機械学習モデルは、例えば乗員を撮像した撮像画像の入力に対し、当該乗員がマスクをしているか否かを推論した結果を出力するよう学習された学習済みモデルである。
なお、マスク判定部14は、顔情報取得部11により取得された乗員の顔情報に基づいて、乗員がマスクをしているか否かを判定してもよい。この場合、マスク判定部14は、学習済みモデルとして、例えば乗員の顔情報の入力に対し、当該乗員がマスクをしているか否かを推論した結果を出力するよう学習された学習済みモデルを用いればよい。
遮蔽検知部15は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示している場合、顔情報取得部11から出力された乗員の顔情報に基づいて、当該乗員の口が遮蔽されているか否かを判定する。遮蔽検知部15は、判定した結果を示す情報を判定部16に出力する。なお、遮蔽検知部15は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示している場合にのみ上記判定を行えばよく、マスク判定部14から出力された情報が、乗員がマスクをしている旨を示している場合は、上記判定を行うことを要しない。
遮蔽検知部15は、例えば機械学習モデルを用いて、乗員の口が遮蔽されているか否かを判定する。この場合、機械学習モデルは、例えば乗員の顔情報の入力に対し、当該乗員の口が遮蔽されているか否かを推論した結果を出力するよう学習された学習済みモデルである。
なお、上記の説明では、遮蔽検知部15は、顔情報取得部11から出力された乗員の顔情報に基づいて、当該乗員の口が遮蔽されているか否かを判定するものとして説明した。しかしながら、遮蔽検知部15はこれに限らず、例えば顔情報取得部11から出力された乗員の顔情報に基づいて、当該乗員の顔の一部または全部が遮蔽されているか否かを判定するものであってもよい。また、この場合、遮蔽検知部15は、例えば顔情報取得部11から出力された乗員の顔情報に基づいて、少なくとも当該乗員の口が遮蔽されているか否かを判定できればよい。
判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしている旨を示している場合、ジェスチャ検出部13から出力された情報が示すジェスチャを、当該乗員のジェスチャとして認識する。
また、判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示しており、かつ遮蔽検知部15から出力された情報が、乗員の口が遮蔽されていない旨を示している場合、ジェスチャ検出部13から出力された情報が示すジェスチャを、当該乗員のジェスチャとして認識する。
一方、判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示しており、かつ遮蔽検知部15から出力された情報が、乗員の口が遮蔽されている旨を示している場合、ジェスチャ検出部13から出力された情報が示すジェスチャを、当該乗員のジェスチャとして認識せずに棄却する。
なお、判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示している場合、遮蔽検知部15から出力された情報に基づくジェスチャの棄却又は認識を判定する前に、手候補検出部12により検出された手候補を示す情報と、顔情報取得部11により取得された乗員の顔情報とに基づいて、手候補と乗員の口との間の距離を算出してもよい。
例えば、判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしていない旨を示している場合、手候補検出部12により検出された手候補を示す情報に基づいて、手候補の検出ボックスを取得するとともに、顔情報取得部11により取得された乗員の顔情報に基づいて、乗員の口の検出ボックスを取得する。なお、判定部16は、既知の手法を用いて、乗員の手候補の検出ボックス及び乗員の口の検出ボックスを取得することができる。そして、判定部16は、当該取得した乗員の手候補の検出ボックスと、乗員の口の検出ボックスとの最小距離を、手候補と乗員の口との間の距離として算出する。
そして、判定部16は、当該算出した距離が、予め定められた閾値を上回っている場合、遮蔽検知部15から出力された情報の如何にかかわらず、ジェスチャ検出部13により検出されたジェスチャを乗員のジェスチャとして認識してもよい。また、判定部16は、当該算出した距離が、予め定められた閾値以下である場合に限り、遮蔽検知部15から出力された情報に基づくジェスチャの棄却又は認識を判定してもよい。このように、判定部16は、遮蔽検知部15から出力された情報に基づくジェスチャの棄却又は認識の判定に先立ち、手候補と乗員の口との間の距離に基づく判定を行うようにすることで、ジェスチャを棄却するか又は認識するかを判定する機会を増やすことができ、ジェスチャの棄却又は認識に関する判定精度が向上する。
次に、実施の形態1に係るジェスチャ検出装置1の動作例について、図2に示すフローチャートを参照しながら説明する。なお、以下の説明では、判定部16は、遮蔽検知部15から出力された情報に基づくジェスチャの棄却又は認識の判定に先立ち、手候補と乗員の口との間の距離に基づく判定を行う場合を例に説明する。
まず、画像取得部10は、撮像装置110から出力された撮像画像を取得する(ステップST1)。画像取得部10は、取得した撮像画像を、顔情報取得部11、手候補検出部12、及びマスク判定部14に出力する。
次に、顔情報取得部11は、画像取得部10から出力された撮像画像に基づいて、乗員の顔の情報(顔情報)を取得する(ステップST2)。顔情報取得部11は、取得した顔情報を遮蔽検知部15及び判定部16に出力する。
次に、手候補検出部12は、画像取得部10から出力された撮像画像に基づいて、乗員の手の候補である手候補を検出する(ステップST3)。手候補検出部12は、検出した手候補を示す情報をジェスチャ検出部13及び判定部16に出力する。
次に、ジェスチャ検出部13は、手候補検出部12から出力された情報が示す手候補の位置および形状に基づいて、乗員のジェスチャを検出する(ステップST4)。ジェスチャ検出部13は、検出した乗員のジェスチャを示す情報を判定部16に出力する。
次に、マスク判定部14は、画像取得部10から出力された撮像画像に基づいて、乗員がマスクをしているか否かを判定する(ステップST5)。マスク判定部14は、判定した結果を示す情報を遮蔽検知部15及び判定部16に出力する。
次に、判定部16は、マスク判定部14から出力された情報が、乗員がマスクをしている旨を示しているか否か、すなわち、マスク判定部14により乗員がマスクをしていると判定されたか否かを確認する(ステップST6)。その結果、マスク判定部14により乗員がマスクをしていると判定されたことを確認した場合(ステップST6;YES)、判定部16は、ステップST4でジェスチャ検出部13により検出されたジェスチャを棄却せず、乗員のジェスチャとして認識する(ステップST7)。
一方、判定部16は、マスク判定部14により乗員がマスクをしていないと判定されたことを確認した場合(ステップST6;NO)、手候補検出部12により検出された手候補を示す情報と、顔情報取得部11により取得された乗員の顔情報とに基づいて、手候補と乗員の口との間の距離を算出する(ステップST8)。
そして、判定部16は、当該算出した距離が予め定められた閾値以下であるか否かを確認する(ステップST9)。その結果、算出した距離が閾値以下でない場合(ステップST9;NO)、処理はステップST7へ移り、判定部16は、ステップST4でジェスチャ検出部13により検出されたジェスチャを棄却せず、乗員のジェスチャとして認識する。一方、算出した距離が閾値以下である場合(ステップST9;YES)、処理はステップST10へ移る。
ステップST10において、遮蔽検知部15は、顔情報取得部11から出力された乗員の顔情報に基づいて、当該乗員の口が遮蔽されているか否かを判定し、判定した結果を示す情報を判定部16に出力する。そして、判定部16は、遮蔽検知部15から出力された情報が、乗員の口が遮蔽されている旨を示しているか否かを確認する(ステップST10)。
その結果、遮蔽検知部15から出力された情報が、乗員の口が遮蔽されている旨を示していることを確認した場合(ステップST10;YES)、判定部16は、ステップST4でジェスチャ検出部13から出力された情報が示すジェスチャを、当該乗員のジェスチャとして認識せずに棄却する(ステップST11)。
一方、判定部16は、遮蔽検知部15から出力された情報が、乗員の口が遮蔽されていない旨を示していることを確認した場合(ステップST10;NO)、ステップST4でジェスチャ検出部13により検出されたジェスチャを棄却せず、乗員のジェスチャとして認識する(ステップST7)。
なお、上記の説明では、判定部16が、ステップST10において遮蔽検知部15から出力された情報に基づくジェスチャの棄却又は認識を判定する前に、ステップST8~ST9において、手候補と乗員の口との間の距離を算出し、当該算出した距離に基づく判定を行う場合を説明した。しかしながら、このステップST8~ST9の処理は必須ではなく、省略されていてもよい。その場合、ステップST6において、乗員がマスクをしていないと判定されたことが確認されたら、処理はステップST10に遷移すればよい。ただし、判定部16は、ステップST8~ST9の処理を行うようにすれば、ジェスチャを棄却するか又は認識するかを判定する機会を増やすことができ、ジェスチャの棄却又は認識に関する判定精度が向上する。
このように、実施の形態1に係るジェスチャ検出装置1では、マスク判定部14により、乗員がマスクをしていないと判定された場合、遮蔽検知部15により、乗員の口が遮蔽されているか否かが判定される。そして、ジェスチャ検出装置1では、遮蔽検知部15により、乗員の口が遮蔽されていると判定された場合、判定部16により、ジェスチャ検出部13にて検出されたジェスチャが棄却される。これにより、ジェスチャ検出装置1は、乗員が飲食をする際の動作をジェスチャとして誤検出することを抑制することができる。
なお、上記の説明では、ジェスチャ検出装置1によるジェスチャの検出対象が車両の乗員である場合の例を説明したが、ジェスチャの検出対象は車両の乗員に限らず、撮像装置により撮像可能な人であればよい。
次に、図3を参照して、実施の形態1に係るジェスチャ検出装置1のハードウェア構成例を説明する。ジェスチャ検出装置1における画像取得部10、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16の各機能は、処理回路により実現される。処理回路は、図3Aに示すように、専用のハードウェアであってもよいし、図3Bに示すように、メモリ53に格納されるプログラムを実行するCPU(Central Processing Unit、中央処理装置、処理装置、演算装置、マイクロプロセッサ、マイクロコンピュータ、プロセッサ、又はDSP(Digital Signal Processor)ともいう)52であってもよい。
処理回路が専用のハードウェアである場合、処理回路51は、例えば、単一回路、複合回路、プログラム化したプロセッサ、並列プログラム化したプロセッサ、ASIC(Application Specific Integrated Circuit)、FPGA(Field Programmable Gate Array)、又はこれらを組み合わせたものが該当する。画像取得部10、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16の各部の機能それぞれを処理回路51で実現してもよいし、各部の機能をまとめて処理回路51で実現してもよい。
処理回路がCPU52の場合、画像取得部10、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16の機能は、ソフトウェア、ファームウェア、又はソフトウェアとファームウェアとの組み合わせにより実現される。ソフトウェア及びファームウェアはプログラムとして記述され、メモリ53に格納される。処理回路は、メモリ53に記録されたプログラムを読み出して実行することにより、各部の機能を実現する。すなわち、ジェスチャ検出装置1は、処理回路により実行されるときに、例えば図2に示した各ステップが結果的に実行されることになるプログラムを格納するためのメモリを備える。また、これらのプログラムは、画像取得部10、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16の手順及び方法をコンピュータに実行させるものであるともいえる。ここで、メモリ53としては、例えば、RAM(Random Access Memory)、ROM(Read Only Memory)、フラッシュメモリ、EPROM(Erasable Programmable ROM)、EEPROM(Electrically EPROM)等の不揮発性又は揮発性の半導体メモリ、磁気ディスク、フレキシブルディスク、光ディスク、コンパクトディスク、ミニディスク、又はDVD(Digital Versatile Disc)等が該当する。
なお、画像取得部10、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16の各機能について、一部を専用のハードウェアで実現し、一部をソフトウェア又はファームウェアで実現するようにしてもよい。例えば、画像取得部10については専用のハードウェアとしての処理回路でその機能を実現し、顔情報取得部11、手候補検出部12、ジェスチャ検出部13、マスク判定部14、遮蔽検知部15、及び判定部16については処理回路がメモリ53に格納されたプログラムを読み出して実行することによってその機能を実現することが可能である。
このように、処理回路は、ハードウェア、ソフトウェア、ファームウェア、又はこれらの組み合わせによって、上述の各機能を実現することができる。
以上のように、実施の形態1によれば、ジェスチャ検出装置1は、人を撮像した撮像画像に基づいて、当該人の手の候補である手候補を検出する手候補検出部12と、手候補検出部12により検出された手候補に基づいて、人のジェスチャを検出するジェスチャ検出部13と、撮像画像に基づいて、人がマスクをしているか否かを判定するマスク判定部14と、マスク判定部14により、人がマスクをしていないと判定された場合に、撮像画像に基づいて得られた人の顔の情報に基づいて、当該人の口が遮蔽されているか否かを判定する遮蔽検知部15と、遮蔽検知部15により、人の口が遮蔽されていると判定された場合、ジェスチャ検出部13により検出されたジェスチャを棄却する判定部16とを備えた。これにより、実施の形態1に係るジェスチャ検出装置1は、人が飲食をする際の動作をジェスチャとして誤検出することを抑制可能となる。
また、判定部16は、マスク判定部14により、人がマスクをしていると判定された場合、ジェスチャ検出部13により検出されたジェスチャを人のジェスチャとして認識する。これにより、実施の形態1に係るジェスチャ検出装置1は、人がマスクをしている際のジェスチャを検出することができる。
また、判定部16は、マスク判定部14により、人がマスクをしていないと判定された場合において、遮蔽検知部15により、人の口が遮蔽されていないと判定された場合、ジェスチャ検出部13により検出されたジェスチャを人のジェスチャとして認識する。これにより、実施の形態1に係るジェスチャ検出装置1は、人がマスクをしていない際の飲食動作以外の動作をジェスチャとして検出することができる。
また、判定部16は、マスク判定部14により、人がマスクをしていないと判定された場合、手候補検出部12により検出された手候補の位置と、人の顔の情報とに基づいて、手候補と人の口との間の距離を算出し、当該算出した距離が閾値以下であり、かつ、遮蔽検知部15により、人の口が遮蔽されていると判定された場合、ジェスチャ検出部13により検出されたジェスチャを棄却し、当該算出した距離が閾値を上回る場合、及び、当該算出した距離が閾値以下であっても、遮蔽検知部15により、人の口が遮蔽されていないと判定された場合、ジェスチャ検出部13により検出されたジェスチャを人のジェスチャとして認識する。これにより、実施の形態1に係るジェスチャ検出装置1は、ジェスチャを棄却するか又は認識するかを判定する機会を増やすことができ、判定精度が向上する。
また、マスク判定部14は、人を撮像した撮像画像の入力に対し、当該人がマスクをしているか否かを出力する学習済みモデルを用いて、当該人がマスクをしているか否かを判定する。これにより、実施の形態1に係るジェスチャ検出装置1は、人がマスクをしているか否かを精度よく判定することができ、ジェスチャの棄却又は認識に関する判定精度が向上する。
また、遮蔽検知部15は、人の顔の情報の入力に対し、当該人の口が遮蔽されているか否かを出力する学習済みモデルを用いて、当該人の口が遮蔽されているか否かを判定する。これにより、実施の形態1に係るジェスチャ検出装置1は、人の口が遮蔽されているか否かを精度よく判定することができ、ジェスチャの棄却又は認識に関する判定精度が向上する。
また、ジェスチャ検出部13は、手候補の位置及び形状を示す情報を入力とし、当該情報が示す手候補の位置及び形状に対応するジェスチャを出力する学習済みモデルを用いて、人のジェスチャを検出する。これにより、実施の形態1に係るジェスチャ検出装置1は、人のジェスチャを精度よく検出することができる。
また、実施の形態1に係る乗員監視システム100は、上記ジェスチャ検出装置1を含んで構成され、人として車両の乗員を撮像する撮像装置110と、判定部16により認識されたジェスチャに基づいて所定の制御を実行する制御装置120とを備えた。これにより、実施の形態1に係る乗員監視システム100は、車両の乗員が飲食をする際の動作をジェスチャとして誤検出することに起因する誤制御を抑制することができる。
なお、本開示は、実施の形態の任意の構成要素の変形、若しくは実施の形態において任意の構成要素の省略が可能である。
本開示は、人が飲食をする際の動作をジェスチャとして誤検出することを抑制可能となり、ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法に用いるのに適している。
1 ジェスチャ検出装置、10 画像取得部、11 顔情報取得部、12 手候補検出部、13 ジェスチャ検出部、14 マスク判定部、15 遮蔽検知部、16 判定部、51 処理回路、52 CPU、53 メモリ、100 乗員監視システム、110 撮像装置、120 制御装置。
Claims (9)
- 人を撮像した撮像画像に基づいて、当該人の手の候補である手候補を検出する手候補検出部と、
前記手候補検出部により検出された前記手候補に基づいて、前記人のジェスチャを検出するジェスチャ検出部と、
前記撮像画像に基づいて、前記人がマスクをしているか否かを判定するマスク判定部と、
前記マスク判定部により、前記人がマスクをしていないと判定された場合に、前記撮像画像に基づいて得られた前記人の顔の情報に基づいて、当該人の口が遮蔽されているか否かを判定する遮蔽検知部と、
前記遮蔽検知部により、前記人の口が遮蔽されていると判定された場合、前記ジェスチャ検出部により検出されたジェスチャを棄却する判定部と、
を備えたジェスチャ検出装置。 - 前記判定部は、
前記マスク判定部により、前記人がマスクをしていると判定された場合、前記ジェスチャ検出部により検出されたジェスチャを前記人のジェスチャとして認識する
ことを特徴とする請求項1記載のジェスチャ検出装置。 - 前記判定部は、
前記マスク判定部により、前記人がマスクをしていないと判定された場合において、前記遮蔽検知部により、前記人の口が遮蔽されていないと判定された場合、前記ジェスチャ検出部により検出されたジェスチャを前記人のジェスチャとして認識する
ことを特徴とする請求項1又は請求項2に記載のジェスチャ検出装置。 - 前記判定部は、
前記マスク判定部により、前記人がマスクをしていないと判定された場合、前記手候補検出部により検出された前記手候補の位置と、前記人の顔の情報とに基づいて、前記手候補と前記人の口との間の距離を算出し、
当該算出した距離が閾値以下であり、かつ、前記遮蔽検知部により、前記人の口が遮蔽されていると判定された場合、前記ジェスチャ検出部により検出されたジェスチャを棄却し、
当該算出した距離が閾値を上回る場合、及び、当該算出した距離が閾値以下であっても、前記遮蔽検知部により、前記人の口が遮蔽されていないと判定された場合、前記ジェスチャ検出部により検出されたジェスチャを前記人のジェスチャとして認識する
ことを特徴とする請求項3記載のジェスチャ検出装置。 - 前記マスク判定部は、
前記人を撮像した撮像画像の入力に対し、当該人がマスクをしているか否かを出力する学習済みモデルを用いて、当該人がマスクをしているか否かを判定する
ことを特徴とする請求項1から請求項4のうちのいずれか1項に記載のジェスチャ検出装置。 - 前記遮蔽検知部は、
前記人の顔の情報の入力に対し、当該人の口が遮蔽されているか否かを出力する学習済みモデルを用いて、当該人の口が遮蔽されているか否かを判定する
ことを特徴とする請求項1から請求項5のうちのいずれか1項に記載のジェスチャ検出装置。 - 前記ジェスチャ検出部は、
前記手候補の位置及び形状を示す情報を入力とし、当該情報が示す手候補の位置及び形状に対応するジェスチャを出力する学習済みモデルを用いて、前記人のジェスチャを検出する
ことを特徴とする請求項1から請求項6のうちのいずれか1項に記載のジェスチャ検出装置。 - 請求項1から請求項7のうちのいずれか1項に記載のジェスチャ検出装置を含んで構成される乗員監視システムであって、
前記人として車両の乗員を撮像する撮像装置と、
前記判定部により認識されたジェスチャに基づいて所定の制御を実行する制御装置と、
を備えた乗員監視システム。 - ジェスチャ検出装置によるジェスチャ検出方法であって、
手候補検出部が、人を撮像した撮像画像に基づいて、当該人の手の候補である手候補を検出するステップと、
ジェスチャ検出部が、前記手候補検出部により検出された前記手候補に基づいて、前記人のジェスチャを検出するステップと、
マスク判定部が、前記撮像画像に基づいて、前記人がマスクをしているか否かを判定するステップと、
前記マスク判定部により、前記人がマスクをしていないと判定された場合に、遮蔽検知部が、前記撮像画像に基づいて得られた前記人の顔の情報に基づいて、当該人の口が遮蔽されているか否かを判定するステップと、
前記遮蔽検知部により、前記人の口が遮蔽されていると判定された場合、判定部が、前記ジェスチャ検出部により検出されたジェスチャを棄却するステップと、
を有するジェスチャ検出方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2023/022338 WO2024257321A1 (ja) | 2023-06-16 | 2023-06-16 | ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 |
| JP2025527173A JP7721044B2 (ja) | 2023-06-16 | 2023-06-16 | ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2023/022338 WO2024257321A1 (ja) | 2023-06-16 | 2023-06-16 | ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024257321A1 true WO2024257321A1 (ja) | 2024-12-19 |
Family
ID=93851707
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2023/022338 Ceased WO2024257321A1 (ja) | 2023-06-16 | 2023-06-16 | ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP7721044B2 (ja) |
| WO (1) | WO2024257321A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014095766A (ja) * | 2012-11-08 | 2014-05-22 | Sony Corp | 情報処理装置、情報処理方法及びプログラム |
| WO2019229938A1 (ja) * | 2018-05-31 | 2019-12-05 | 三菱電機株式会社 | 画像処理装置、画像処理方法及び画像処理システム |
| JP2023053670A (ja) * | 2021-10-01 | 2023-04-13 | ソニーグループ株式会社 | 情報処理装置、情報処理方法、及びプログラム |
-
2023
- 2023-06-16 JP JP2025527173A patent/JP7721044B2/ja active Active
- 2023-06-16 WO PCT/JP2023/022338 patent/WO2024257321A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014095766A (ja) * | 2012-11-08 | 2014-05-22 | Sony Corp | 情報処理装置、情報処理方法及びプログラム |
| WO2019229938A1 (ja) * | 2018-05-31 | 2019-12-05 | 三菱電機株式会社 | 画像処理装置、画像処理方法及び画像処理システム |
| JP2023053670A (ja) * | 2021-10-01 | 2023-04-13 | ソニーグループ株式会社 | 情報処理装置、情報処理方法、及びプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2024257321A1 (ja) | 2024-12-19 |
| JP7721044B2 (ja) | 2025-08-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11535280B2 (en) | Method and device for determining an estimate of the capability of a vehicle driver to take over control of a vehicle | |
| CN107405121B (zh) | 用于对车辆的驾驶员的疲劳状态和/或睡觉状态进行识别的方法和装置 | |
| JPWO2018154709A1 (ja) | 動作学習装置、技能判別装置および技能判別システム | |
| JPWO2019229938A1 (ja) | 画像処理装置、画像処理方法及び画像処理システム | |
| WO2022113275A1 (ja) | 睡眠検出装置及び睡眠検出システム | |
| US11195108B2 (en) | Abnormality detection device and abnormality detection method for a user | |
| US12295731B2 (en) | Driver availability detection device and driver availability detection method | |
| CN111696312B (zh) | 乘员观察装置 | |
| WO2011058837A1 (ja) | 偽指判定装置、偽指判定方法および偽指判定プログラム | |
| JP7721044B2 (ja) | ジェスチャ検出装置、乗員監視システム、及びジェスチャ検出方法 | |
| CN119018052B (zh) | 车辆的后视镜调节方法、装置、车辆及存储介质 | |
| Constantin et al. | Driver monitoring using face detection and facial landmarks | |
| JP2009080706A (ja) | 個人認証装置 | |
| KR101561817B1 (ko) | 얼굴과 손 인식을 이용한 생체 인증 장치 및 방법 | |
| WO2021171538A1 (ja) | 表情認識装置及び表情認識方法 | |
| JP7374386B2 (ja) | 状態判定装置および状態判定方法 | |
| WO2019030855A1 (ja) | 運転不能状態判定装置および運転不能状態判定方法 | |
| JP7812001B2 (ja) | 開瞼度検出装置、開瞼度検出方法、および眠気判定システム | |
| JP7843933B2 (ja) | 眠気推定装置および眠気推定方法 | |
| JP6698966B2 (ja) | 誤検出判定装置及び誤検出判定方法 | |
| WO2021186710A1 (ja) | ジェスチャ検出装置及びジェスチャ検出方法 | |
| JP7721045B2 (ja) | 判定装置および判定方法 | |
| JP7483060B2 (ja) | 手検出装置、ジェスチャー認識装置および手検出方法 | |
| JP7325687B2 (ja) | 眠気推定装置及び眠気推定システム | |
| US20250232598A1 (en) | Driving assistance apparatus and driving assistance method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23941623 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2025527173 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |