WO2026014240A1 - 情報処理方法、プログラム、及び情報処理装置 - Google Patents
情報処理方法、プログラム、及び情報処理装置Info
- Publication number
- WO2026014240A1 WO2026014240A1 PCT/JP2025/022839 JP2025022839W WO2026014240A1 WO 2026014240 A1 WO2026014240 A1 WO 2026014240A1 JP 2025022839 W JP2025022839 W JP 2025022839W WO 2026014240 A1 WO2026014240 A1 WO 2026014240A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- operator
- information processing
- vehicle
- reference data
- processing device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
-
- G—PHYSICS
- G08—SIGNALLING
- G08G—TRAFFIC CONTROL SYSTEMS
- G08G1/00—Traffic control systems for road vehicles
- G08G1/16—Anti-collision systems
Definitions
- the present invention relates to an information processing method, a program, and an information processing device.
- a safe driving assessment device that captures an image of a vehicle driver using an imaging device, obtains an angle value indicating the angle of the driver's face relative to the vehicle's traveling direction, and determines whether the driver is not paying attention to the road ahead based on the obtained angle value (see, for example, Patent Document 1).
- One aspect of this is to provide an information processing method that can efficiently generate reference data when performing monitoring processing on an operator.
- An information processing method involves using a camera mounted on a vehicle to acquire an image including an operator operating the vehicle, determining whether the vehicle is in a predetermined driving state, and, if it is determined that the vehicle is in the predetermined driving state, causing a computer to execute processing to generate reference data for performing monitoring processing on the operator based on the acquired image.
- One aspect of the present disclosure provides an information processing method that efficiently generates reference data when performing monitoring processing on an operator.
- FIG. 1 is a schematic diagram illustrating a configuration of a camera unit including an information processing device (camera ECU) according to a first embodiment
- FIG. 2 is a block diagram illustrating a physical configuration of an information processing apparatus.
- FIG. 10 is an explanatory diagram relating to a generation process of a face detection model.
- FIG. 10 is an explanatory diagram relating to a generation process of a facial landmark model.
- FIG. 10 is an explanatory diagram regarding a generation process of a posture detection model.
- 10 is a flowchart illustrating a process (calibration process) of a control unit of the information processing apparatus.
- FIG. 10 is an explanatory diagram showing the detection result of the position of a face (bounding box) in an image including an operator.
- 10A and 10B are explanatory diagrams showing the detection results of the face direction (distance between the eyes) in an image including an operator; 10 is an explanatory diagram showing the detection result of the body orientation (distance between both shoulders) in an image including an operator; FIG. 10A and 10B are explanatory diagrams showing the detection results of the eye opening degree (aspect ratio) in an image including an operator.
- 10 is a flowchart illustrating a process (monitoring process) of a control unit of the information processing apparatus.
- 10 is a flowchart illustrating processing (calibration processing when an operator makes a change) by a control unit of an information processing apparatus according to a second embodiment.
- FIG. 1 is a schematic diagram illustrating the configuration of a camera unit 1 including an information processing device 2 (camera ECU) according to the first embodiment.
- FIG. 2 is a block diagram illustrating the physical configuration of the information processing device 2.
- the camera unit 1 includes a camera 11 mounted on, for example, an OHC (overhead console) in a vehicle C, and the information processing device 2.
- the information processing device 2 performs image processing on images captured by the camera 11 and uses the image processing results to perform monitoring of an operator operating the vehicle C.
- Such a camera unit 1 may function as, for example, a cabin monitor system.
- the camera 11 is, for example, a CMOS camera, and is placed, for example, on an OHC (overhead console). There may be any number of cameras 11 in this embodiment, provided that there is one or more.
- the camera 11 is communicably connected to the information processing device 2, for example, via a serial cable, and outputs captured images such as video to the information processing device 2 periodically or in real time. Note that "real time" in the various processes in this specification does not mean strictly immediate or simultaneous processing, but rather means that processing is performed as quickly as possible.
- the information processing device 2 includes a control unit 20, a memory unit 23, an input/output I/F 21, and a communication unit 22, and functions as a camera ECU that performs image processing on images captured by the camera 11 and performs processing according to the results of this image processing.
- the control unit 20 is composed of a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and is configured to perform various control processes and arithmetic processes by reading and executing a control program P (program product) and data pre-stored in the memory unit 23.
- the storage unit 23 is composed of a volatile memory element such as RAM (Random Access Memory), or a non-volatile memory element such as ROM (Read Only Memory), EEPROM (Electrically Erasable Programmable ROM), or flash memory.
- the storage unit 23 pre-stores the control program P and data that the control unit 20 references when processing various calculations.
- the storage unit 23 may also store actual files of learning models such as a face detection model 201, a face landmark model 202, and a posture detection model 203. These learning models will be described later.
- learning models such as the face detection model 201, face landmark model 202, and posture detection model 203 may be implemented on an AI chip, and the control unit 20 may input images to these learning models and obtain recognition results for objects from the learning models by performing input/output processing on the AI chip via an internal bus or the like.
- the input/output I/F 21 is, for example, a communications interface for serial communication.
- the information processing device 2 is connected to the camera 11 and the like via the input/output I/F 21 so that they can communicate with each other.
- a speaker or a display device with audio output functionality may be connected to the input/output I/F 21. A warning sound or the like is emitted from the speaker or the like based on the results of the monitoring process executed by the information processing device 2.
- the communication unit 22 is an input/output interface that uses a communication protocol such as CAN or Ethernet (registered trademark).
- the control unit 20 may communicate with other ECUs 3 connected to the in-vehicle network 4 via the communication unit 22.
- the information processing device 2 communicates with ECUs 3 via the in-vehicle network 4 and acquires various communication data (CAN messages, Ethernet frames, etc.) from these ECUs 3.
- FIG. 3 is an explanatory diagram regarding the generation process of the face detection model 201.
- the face detection model 201 is implemented, for example, on an AI chip or the like.
- the face detection model 201 may be configured as a software function unit controlled by the control unit 20.
- the information processing device 2 is described as generating the face detection model 201, but this is not limited to this; the face detection model 201 may also be generated by a model generation device such as a server device.
- the control unit 20 of the information processing device 2 may train a pre-learning machine learning model (for example, a neural network such as YOLO or R-CNN) using existing training data or training data prepared for the face detection model 201, thereby generating a face detection model 201 that recognizes (detects) an object (the operator's face) included in an input image (an image including the operator) when the input image is input.
- a pre-learning machine learning model for example, a neural network such as YOLO or R-CNN
- the object recognition results output by the face detection model 201 include at least whether or not an object (face) was recognized. Furthermore, if an object is recognized, the recognition result includes the area (bounding box) and type (class) of the object in the image. Furthermore, the recognition result may also include a confidence score (predicted probability of each class) of the recognized object.
- the training data includes question data and answer data, with the image acquired from camera 11 (including the operator) corresponding to the question data and the recognition results for the object (face) corresponding to the answer data.
- the area and type of the object (face) that constitutes the answer data may be set by annotating (adding) it to the image.
- the data set of question data and answer data included in the training data for learning face detection model 201 and the data set of input data and output data when face detection model 201 is used are synonymous, and anything defined in one data set will naturally apply to the other data set as well.
- the neural network (face detection model 201) trained using training data is expected to be used as a program module that is part of artificial intelligence software.
- the face detection model 201 is used in the information processing device 2 that includes the control unit 20 (CPU, etc.) and memory unit 23, and when executed by the information processing device 2 that has such processing capabilities, a neural network system is formed.
- the control unit 20 of the information processing device 2 performs calculations to extract feature quantities from the image input to the input layer in accordance with instructions from the face detection model 201 stored in the memory unit 23, and outputs the recognition result for the object.
- the face detection model 201 is configured using, for example, YOLO or R-CNN (Region Convolutional Neural Network), and has an input layer that accepts image input, an intermediate layer that extracts features from the image, and an output layer that outputs the recognition results for the object.
- the input layer has multiple neurons that accept image input and passes the input values to the intermediate layer.
- the intermediate layer is defined using an activation function such as a ReLU function or a sigmoid function, has multiple neurons that extract features for each input value, and passes the extracted features to the output layer. Parameters such as the weighting coefficients and bias values of the activation function are optimized using the backpropagation method.
- the output layer is configured, for example, with a fully connected layer, and outputs the recognition results for the object based on the features output from the intermediate layer.
- the output layer may also include, for example, a softmax layer, and output the probability (confidence score) for the object.
- the face detection model 201 is R-CNN or the like, but is not limited to this.
- the face detection model 201 may be constructed using other machine learning algorithms, such as neural networks other than R-CNN, transformers, BERTs, GPTs, recurrent neural networks (RNNs), long-short term models (LSTMs), support vector machines (SVMs), Bayesian networks, linear regression, regression trees, multiple regressions, random forests, and ensembles.
- the face detection model 201 may be constructed using an artificial intelligence chatbot based on an LLM, such as ChatGPT. In this case, ChatGPT may be fine-tuned to efficiently output recognition results for the target object.
- a question (prompt) generated using an external database, such as a WebDB may be input together with an image to a prompt input I/F, which is the input interface of ChatGPT.
- FIG. 4 is an explanatory diagram of the generation process of the facial landmark model 202.
- the facial landmark model 202 is implemented on an AI chip or the like, or configured as a software function unit controlled by the control unit 20.
- the facial landmark model 202 may also be configured using YOLO or R-CNN or the like.
- the training data includes question data and answer data, with images acquired from the camera 11 (images including the operator) corresponding to the question data, and the areas or positions of landmarks indicating characteristic facial features such as the eyes and nose, as well as the attributes of which facial features the landmarks indicate, corresponding to the answer data.
- the control unit 20 of the information processing device 2 may train a pre-training machine learning model (e.g., a neural network such as YOLO or R-CNN) using existing training data or training data prepared for the facial landmark model 202, thereby generating a facial landmark model 202 that recognizes (detects) objects (facial landmarks) included in an input image (an image including the operator) by inputting the image.
- a pre-training machine learning model e.g., a neural network such as YOLO or R-CNN
- FIG. 5 is an explanatory diagram regarding the generation process of the posture detection model 203.
- the posture detection model 203 is implemented on an AI chip or the like, or is configured as a software function unit by the control unit 20.
- the posture detection model 203 may also be configured using YOLO or R-CNN or the like.
- the training data includes question data and answer data, with the image acquired from the camera 11 (the image including the operator) corresponding to the question data, and the areas or positions of landmarks indicating characteristic parts of the upper body such as the shoulders, elbows, neck, and waist, and the attributes indicating which part of the human body the landmark indicates corresponding to the answer data.
- the control unit 20 of the information processing device 2 may use existing training data or training data prepared for the posture detection model 203 to train a pre-learning machine learning model (for example, a neural network such as YOLO or R-CNN), thereby generating a posture detection model 203 that recognizes (detects) objects (landmarks in the upper body) included in an input image (an image including an operator) when the image is input.
- a pre-learning machine learning model for example, a neural network such as YOLO or R-CNN
- FIG. 6 is a flowchart illustrating the processing (calibration processing) of the control unit 20 of the information processing device 2.
- the control unit 20 of the information processing device 2 performs the following processing, for example, when the vehicle C is in an activated state. While performing the following processing, the control unit 20 of the information processing device 2 steadily or periodically acquires video or still images of the operator from the camera 11, and stores the acquired images in the memory unit 23 with a timestamp indicating the time at which the image was captured. By storing each acquired image (video frame) in this way in association with the time at which the image was captured, images captured at multiple consecutive times can be stored (saved) in chronological order in the memory unit 23.
- the control unit 20 of the information processing device 2 determines whether vehicle C has started to move from a stopped state (S101). For example, the control unit 20 of the information processing device 2 periodically acquires messages including vehicle speed and the like from the ECU 3, which periodically outputs such messages via the in-vehicle network 4. Based on the acquired messages including vehicle speed and the like, the control unit 20 of the information processing device 2 determines whether vehicle C has started to move from a stopped state based on whether vehicle C has started to move after coming to a stop, i.e., after the vehicle speed has essentially become 0 km/h, i.e., whether the vehicle speed has become greater than 0 km/h.
- control unit 20 of the information processing device 2 may determine that vehicle C has started to move from a stopped state when it acquires a message (a driving start message) from the ECU 3, which outputs a message including information indicating that vehicle C has started to move when vehicle C starts to move from a stopped state.
- a message a driving start message
- the control unit 20 of the information processing device 2 acquires the vehicle speed and steering angle (S102).
- the control unit 20 of the information processing device 2 periodically acquires messages including the vehicle speed, etc., from the ECU 3 that manages the vehicle speed, via the in-vehicle network 4.
- the control unit 20 of the information processing device 2 periodically acquires messages including the steering angle, etc., from the ECU 3 that manages the steering angle, via the in-vehicle network 4.
- the control unit 20 of the information processing device 2 associates the acquired vehicle speed and steering angle with the acquisition time point at which the message including the vehicle speed or steering angle was acquired, and stores them in the memory unit 23. By storing the vehicle speed and steering angle in this manner and associating them with the acquisition time point, the vehicle speed and steering angle values acquired at multiple consecutive time points can be stored (saved) in the memory unit 23 as history data arranged in chronological order.
- the control unit 20 of the information processing device 2 determines whether the vehicle C is traveling straight based on the vehicle speed and steering angle (S103). Based on periodically acquired messages including the vehicle speed, etc. and messages including the steering angle, etc., the control unit 20 of the information processing device 2 determines that the vehicle C is traveling straight based on, for example, if the vehicle speed is equal to or greater than a predetermined speed and the steering angle is within a predetermined angle.
- the predetermined speed is, for example, 5 km/h
- the predetermined angle is, for example, ⁇ 5° when straight traveling is set as the reference angle.
- the control unit 20 of the information processing device 2 determines whether the straight-ahead state has continued for a predetermined period of time or longer (S104). When the control unit 20 of the information processing device 2 determines that vehicle C is in a predetermined straight-ahead state (vehicle speed is greater than or equal to a predetermined speed and steering angle is within a predetermined angle), it determines whether the straight-ahead state has continued (maintained) for a predetermined period of time, such as 5 seconds.
- the control unit 20 of the information processing device 2 If the straight-line driving state continues for a predetermined period or longer (S104: YES), the control unit 20 of the information processing device 2 generates reference data (S105). If the straight-line driving state continues for a predetermined period or longer, the control unit 20 of the information processing device 2 determines the start and end points of the predetermined period. The control unit 20 of the information processing device 2 determines the start and end points based on the acquisition points associated with the vehicle speed and steering angle used when determining whether the straight-line driving state has continued for a predetermined period or longer.
- control unit 20 of the information processing device 2 may determine the start point as the point at which it is determined that the vehicle C is in a straight-line driving state based on the vehicle speed and steering angle in the processing of S103, and the end point as the point at which it is determined that the straight-line driving state has continued for a predetermined period or longer.
- the control unit 20 of the information processing device 2 acquires images captured during a specified period according to the start and end points of the identified period.
- the storage unit 23 stores images acquired from the camera 11 in association with the image capture time, and the control unit 20 of the information processing device 2 acquires images captured during the specified period during which the straight-line state continued, i.e., multiple images captured from the start to the end of the specified period. If the images are a video, the multiple images correspond to multiple frames (frame groups) that make up the video.
- the control unit 20 of the information processing device 2 extracts the position, face direction, body direction, or eye opening of the operator contained in each of a plurality of images (frame group) captured over a predetermined period of time.
- the control unit 20 of the information processing device 2 calculates the average, median, or mode of each of the face positions extracted from each of these images, and generates this average as reference data indicating the face position (face position reference data).
- the control unit 20 of the information processing device 2 calculates the average, median, or mode of each face orientation extracted from each of these images and generates the average as reference data indicating the face orientation (reference data for face orientation).
- the control unit 20 of the information processing device 2 calculates the average, median, or mode of each body orientation extracted from each of these images and generates the average as reference data indicating the body orientation (reference data for body orientation).
- the control unit 20 of the information processing device 2 calculates the maximum value of the eye opening extracted from each of these images and generates the maximum value as reference data indicating the eye opening (reference data for eye opening).
- the control unit 20 of the information processing device 2 may generate reference data for all of the operator's face position, face orientation, body orientation, and eye opening, or may generate reference data for one or more of them. In other words, the control unit 20 of the information processing device 2 may generate reference data for at least one of the operator's face position, face orientation, body orientation, and eye opening, depending on the type of monitoring process to be executed. Below, we will explain the reference data for the operator's face position, face direction, body direction, and eye opening.
- Figure 7 is an explanatory diagram showing the detection results of the position of a face (bounding box) in an image including an operator.
- the control unit 20 of the information processing device 2 inputs each of a plurality of images (frame groups) into a face detection model 201, thereby detecting the position of the face in each of the images.
- the face detection model 201 is, for example, a FaceDetect model, and when an image including an operator is input, it detects the face of the operator and outputs a rectangular frame (bounding box) indicating the position of the detected face, i.e., the area of the face in the image, superimposed on the image.
- the bounding box is defined by the coordinates of, for example, the upper left vertex and the lower right vertex in the image's coordinate system.
- the control unit 20 of the information processing device 2 extracts a bounding box from each of multiple images (frame group) captured within a predetermined period of time while the vehicle C is traveling straight, and derives a bounding box that is the average value of these extracted bounding boxes.
- the control unit 20 of the information processing device 2 may derive the coordinate values of the upper left and lower right vertices of the bounding box that is the average value by calculating the average coordinate values of the upper left and lower right vertices of each bounding box.
- control unit 20 of the information processing device 2 may calculate the center of gravity or center point of each bounding box, and derive the coordinates of the center of gravity of the bounding box that is the average value by calculating the average X and Y coordinate values of each calculated center of gravity, etc.
- the control unit 20 of the information processing device 2 stores the calculated average value of the bounding boxes as reference data (reference data for face position) in the storage unit 23.
- the control unit 20 of the information processing device 2 may set an area expanded by, for example, 25% of the calculated average value of the bounding boxes as the reference data (detection frame). In this embodiment, this expanded area is used as the reference data (reference data for face position).
- the control unit 20 of the information processing device 2 may store it with a timestamp or the like indicating the time of generation.
- the reference data (face position reference data) set in this expanded area by 25% or more is used in the monitoring process for the operator.
- the control unit 20 of the information processing device 2 uses the face detection model 201 to detect the position of the operator's face in an image containing the operator acquired during the monitoring process, and if the detected face position deviates from the area indicated by the reference data by, for example, 25% or more, it may determine that the operator is not facing forward, and if the state of deviation is, for example, 70% or more for three seconds, it may be determined that the driver is drowsy and issue a warning sound or the like.
- Figure 8 is an explanatory diagram showing the detection results of the face orientation (distance between the eyes) in an image including an operator.
- the control unit 20 of the information processing device 2 inputs each of a plurality of images (frame groups) into the face landmark model 202, thereby detecting the face orientation in each of the images.
- the control unit 20 of the information processing device 2 may extract an image (face area image) of the face area (bounding box) detected by the face detection model 201 and input the extracted face area image to the face landmark model 202.
- the face landmark model 202 is, for example, a FaceLandmark model or a MediaPipe model.
- an image including a face When an image including a face is input, it detects multiple landmarks that are characteristic parts of the face, such as the eyes, nose, and mouth, and outputs the points indicated by these landmarks by superimposing them on the image. These detected landmarks are assigned coordinate values in the image coordinate system and attributes indicating each part of the face (eyes, nose, mouth, etc.).
- the control unit 20 of the information processing device 2 identifies two landmarks representing the left and right eyes, and calculates the distance between the two identified landmarks (landmarks representing the eyes) in the image coordinate system, thereby calculating the distance between the left and right eyes.
- the control unit 20 of the information processing device 2 calculates the average value of the distance between the eyes calculated for each of multiple images (frame groups), and derives this average value as reference data (reference data for face direction).
- the reference data (face direction reference data) set using the average value of the distance between the eyes derived in this way is used in the monitoring process for the operator.
- the control unit 20 of the information processing device 2 uses the facial landmark model 202 to detect the distance between the eyes of the operator in an image including the operator acquired during the monitoring process.
- the control unit 20 of the information processing device 2 may determine that the operator is looking to the left if the detected distance between the eyes is longer than the reference data, for example, or that the operator is looking to the right if the operator continues to look to the left or right for a predetermined period of time, such as three seconds, and may issue a warning sound or the like, indicating that the operator is in a state of inattentive driving.
- Figure 9 is an explanatory diagram showing the detection results of the body orientation (distance between both shoulders) in an image including an operator.
- the control unit 20 of the information processing device 2 inputs each of a plurality of images (frame groups) into the posture detection model 203, thereby detecting the posture of the operator in each of the images.
- the posture detection model 203 is, for example, a posedetect model, and when an image including an operator is input, it detects characteristic parts of the operator's upper body, such as joints such as the shoulders, and outputs the points indicating each of these parts by superimposing them on the image. These detected parts are assigned coordinate values in the image coordinate system and attributes indicating each part of the body (shoulders, elbows, wrists, neck, waist, etc.).
- the control unit 20 of the information processing device 2 for example, identifies two parts representing the left and right shoulders, and calculates the distance between the two identified parts (shoulders) in the image coordinate system, thereby calculating the distance between the left and right shoulders.
- the control unit 20 of the information processing device 2 calculates the average value of the distance between the shoulders calculated for each of multiple images (frame groups), and derives this average value as reference data (reference data for body orientation).
- the reference data (body orientation reference data) set using the average value of the distance between the shoulders derived in this way is used in the monitoring process for the operator.
- the control unit 20 of the information processing device 2 uses the posture detection model 203 to detect the distance between the operator's shoulders in an image including the operator acquired during the monitoring process.
- the control unit 20 of the information processing device 2 may determine that the operator is looking to the left if the detected distance between the shoulders is longer than the reference data, for example, or that the operator is looking to the right if the distance is shorter than the reference data, and may issue a warning sound or the like if the operator continues to look to the left or right for a predetermined period of time, such as three seconds, indicating that the operator is in a state of inattentive driving.
- the control unit 20 of the information processing device 2 inputs each of multiple images (frame groups) into the face landmark model 202 to detect the eye opening in each of the images.
- the control unit 20 of the information processing device 2 may extract an image (face area image) of the face area (bounding box) detected by the face detection model 201 and input the extracted face area image to the face landmark model 202.
- the face landmark model 202 is, for example, a Face Landmark model or a Mediapipe model.
- the face landmark model 202 detects multiple landmarks (P1, P2, P3, P4, P5, P6) around or along the outline of the eyes. These detected landmarks are assigned coordinate values in the image coordinate system and attributes indicating the name of the part of the eye (upper eyelid, lower eyelid, etc.).
- the control unit 20 of the information processing device 2 calculates the eye length (horizontal length) and the distance from the upper eyelid to the lower eyelid (vertical length: eye height) from the detected multiple landmarks (P1, P2, P3, P4, P5, P6), and calculates the eye aspect ratio (vertical eye length / horizontal eye length) by dividing the vertical eye length by the horizontal eye length, as data related to the eye opening.
- the distance from the upper eyelid to the lower eyelid is calculated by adding up the distances at two locations (P2-P6, P3-P5), so the eye length (horizontal length) is divided by twice the value: " ⁇ eye height (distance P2-P6) + (distance P3-P5) ⁇ / eye length (distance P1-P4) ⁇ 2.”
- the control unit 20 of the information processing device 2 derives the maximum value of the eye opening degree (aspect ratio) calculated for each of the multiple images (frame group) as reference data (eye opening degree reference data).
- the reference data (eye opening reference data) set using the maximum value of the eye opening (aspect ratio) derived in this way is used in the monitoring process for the operator.
- the control unit 20 of the information processing device 2 uses the facial landmark model 202 to detect the eye opening (aspect ratio) of the operator in an image including the operator acquired during the monitoring process.
- the control unit 20 of the information processing device 2 compares the detected eye opening with the reference data, and may issue a warning sound, for example, if the detected eye opening is less than 90% of the reference data (maximum eye opening).
- the control unit 20 of the information processing device 2 may perform the monitoring process for either the left or right eye that is recognized.
- control unit 20 of the information processing device 2 may vary the reference data (reference data for eye opening) depending on the orientation of the face (up and down).
- the control unit 20 of the information processing device 2 may determine whether the orientation of the operator's face is upward or downward based on the landmarks detected by the facial landmark model 202.
- the control unit 20 of the information processing device 2 may reduce the reference data (reference data for eye opening) by, for example, 10% when the operator's face is facing downward. When the operator looks down, their eyes appear to be less open, which could lead to a greater likelihood of their eyes being determined to be closed. By reducing the reference data (reference data for eye opening) (changing the threshold to a stricter value), it is possible to prevent erroneous determinations from occurring.
- the control unit 20 of the information processing device 2 may increase the reference data (reference data for eye opening) by, for example, 10% when the operator's face is facing upward. When the operator looks up, their eyes appear to be more open, which could lead to a greater likelihood of their eyes being determined to be closed. By increasing the reference data (reference data for eye opening) (changing the threshold to a looser value), it is possible to prevent erroneous determinations from occurring.
- control unit 20 of information processing device 2 applies the newly generated reference data in place of the currently applied reference data, i.e., overwrites the current reference data stored in memory unit 23 with the newly generated reference data, thereby updating the reference data used in the monitoring process.
- the control unit 20 of the information processing device 2 generates the reference data, but this is not limited to this.
- the reference data may also be generated by an external server such as a cloud server that is wirelessly connected to the information processing device 2.
- the information processing device 2 may output an image including the operator to the external server via an external network such as a TCU (Telematics Control Unit) mounted on the vehicle C and the Internet, and issue an instruction to generate the reference data.
- the external server generates the reference data based on the image from the information processing device 2 and outputs the generated reference data to the information processing device 2.
- the control unit 20 of the information processing device 2 may generate the reference data by obtaining the reference data from the external server.
- the control unit 20 of the information processing device 2 does not generate reference data and continues to use the currently applied reference data (S1031). If vehicle C is not traveling straight or if the straight-line state has not continued for a predetermined period or longer, it is difficult for the operator to imagine that the vehicle is facing forward, so the control unit 20 of the information processing device 2 does not generate reference data and continues to use the currently applied reference data.
- the control unit 20 of the information processing device 2 executes a monitoring process for the operator using the reference data (S106). If vehicle C has not started moving from a stopped state, that is, while driving after generating reference data by maintaining the initial straight-ahead state for a predetermined period after stopping, the control unit 20 of the information processing device 2 executes the monitoring process by continuously applying the reference data. If reference data has not been created when vehicle C has started moving from a stopped state, the control unit 20 of the information processing device 2 may execute the monitoring process using, for example, a preset initial value (default value).
- control unit 20 of the information processing device 2 executes monitoring processing corresponding to the type of reference data generated, depending on the type of reference data (face position reference data, face orientation reference data, body orientation reference data, eye opening reference data). Furthermore, the control unit 20 of the information processing device 2 may execute monitoring processing by combining two or more types of reference data.
- FIG. 11 is a flowchart illustrating the processing (monitoring processing) of the control unit 20 of the information processing device 2.
- the control unit 20 of the information processing device 2 may perform the processing shown in the following flowchart as a subroutine of processing S106, as an example of the monitoring processing.
- the control unit 20 of the information processing device 2 determines whether the operator's eyes are closed based on the acquired image (T101).
- the control unit 20 of the information processing device 2 uses the facial landmark model 202 to detect whether the operator's eyes are closed (Eye Closed) and derives this detection result as a provisional determination. If it is determined that the operator's eyes are not closed (T101: NO), the control unit 20 of the information processing device 2 detects whether the operator is blinking or the degree of eye opening (T1011).
- the control unit 20 of the information processing device 2 determines whether the operator is facing forward (T102). If it is determined that the operator's eyes are closed, the control unit 20 of the information processing device 2 checks the direction of the face based on the detection results of the facial landmark model 202, and determines whether the face is facing forward, i.e., whether the face is not looking down excessively or excessively.
- the control unit 20 of the information processing device 2 determines that the operator's facial orientation is inappropriate (T1022). If it is determined that the operator is not facing forward, that is, if it is determined that the operator is not facing either left or right, the control unit 20 of the information processing device 2 determines that the operator's facial orientation is inappropriate and that the operator is distracted driving.
- the control unit 20 of the information processing device 2 determines whether the operator's eye opening degree is small for a predetermined period of time (T103). If it is determined that the operator is facing forward, the control unit 20 of the information processing device 2 determines whether the operator's eye opening degree (EAR) is small for a predetermined period of time, for example, 2 seconds. If the operator's eye opening degree is not small for a predetermined period of time (T103: NO), the control unit 20 of the information processing device 2 determines that the operator is not sleeping (T104). If the operator's eye opening degree is small for a predetermined period of time (T103: YES), the control unit 20 of the information processing device 2 generates a warning sound (T105).
- T103 the operator's eye opening degree is small for a predetermined period of time
- the control unit 20 of the information processing device 2 determines whether the operator's eye opening degree is small for a predetermined period of time (T106). If the operator's eye opening degree is not small for a predetermined period of time (T106: NO), the control unit 20 of the information processing device 2 determines that the operator is not sleeping (T107). If the operator's eye opening degree is small for a predetermined period of time (T106: YES), the control unit 20 of the information processing device 2 determines that the operator is sleeping (T1061). After the sound warning, if the operator's eye opening degree (EAR) continues to be small, the control unit 20 of the information processing device 2 determines that the operator is sleeping, and if it is large, the control unit 20 determines that the operator is not sleeping.
- T106 predetermined period of time
- the control unit 20 of the information processing device 2 can combine the shaking of the face and posture with steering angle information from the vehicle C to efficiently determine whether the operator is actually asleep. In other words, even if the face is facing forward but the eyes are looking down, it can be determined that the operator is not asleep but has their eyes closed, and after a predetermined period of time has passed, it can be prevented from issuing a warning as a drowsy driver.
- the control unit 20 of the information processing device 2 determines whether vehicle C has been stopped (S107). For example, the control unit 20 of the information processing device 2 periodically acquires messages containing data indicating whether the ignition switch is on (started) or off (stopped) from the ECU 3 that manages the ignition switch, and determines whether vehicle C has been stopped (ignition switch: off) by referencing the acquired messages. If it is determined that vehicle C has not been stopped (S107: NO), the control unit 20 of the information processing device 2 performs loop processing to execute the processing from S101 again. By performing loop processing in this manner, the control unit 20 of the information processing device 2 can repeat the process of generating and updating reference data when vehicle C starts running (ignition switch: on) and travels straight for the first time after starting driving.
- control unit 20 of the information processing device 2 erases the reference data stored in the memory unit 23 (S108).
- control unit 20 of the information processing device 2 may erase the reference data stored in the memory unit 23 or initialize the reference data.
- the information processing device 2 continuously or periodically acquires images including the operator operating the vehicle C using the camera 11 mounted on the vehicle C.
- the images may be still images or videos captured at a predetermined frame rate.
- the information processing device 2 further acquires information regarding the driving state of the vehicle C (vehicle C state information) based on, for example, CAN messages acquired from various ECUs 3 communicably connected via the in-vehicle network 4.
- vehicle C state information includes, for example, vehicle speed and steering angle (angle of the steering shaft: steering angle).
- the information processing device 2 can recognize the driving state determined by the vehicle speed and steering angle based on the acquired vehicle C state information.
- the information processing device 2 determines that the driving state of the vehicle C is a straight-line state at a predetermined speed (e.g., 5 km/h) or higher and that this straight-line state has continued for a predetermined period (e.g., 5 seconds) or longer, it generates reference data for performing monitoring processing on the operator based on images captured during the predetermined period, i.e., the period during which the straight-line state continued.
- the vehicle may be determined to be in a straight-ahead state if the steering angle is within a predetermined angle range (e.g., ⁇ 5° when straight-ahead driving is set as 0°).
- Parameters such as a predetermined speed (vehicle speed), a predetermined angle range (steering angle), and a predetermined period (state duration) for determining whether the vehicle C is in a straight-ahead state
- a predetermined speed vehicle speed
- a predetermined angle range steering angle
- a predetermined period state duration
- the information processing device 2 generates reference data for performing monitoring processing based on images of the operator captured during a period in which the vehicle C maintains a straight-ahead state, which is a predetermined driving state based on the vehicle speed and steering angle, after the vehicle C starts traveling. This eliminates the need for a process to pre-register the reference data.
- the frames (frame group) included in a video capturing the operator will mostly contain frames in which the operator faces forward.
- the reference data is generated using multiple images (frame group constituting the video) in which the operator faces forward, the reference data corresponds to data when the operator faces forward.
- an alert is issued if the direction of the operator's face or upper body deviates to the left or right from the forward direction. Because the reference data corresponds to data when the operator is facing forward, the accuracy of the monitoring process can be guaranteed or improved.
- the information processing device 2 determines whether vehicle C is in a predetermined driving state the first time it travels after stopping; in other words, it determines that it is in a predetermined driving state the first time it travels straight after starting to travel. Therefore, after determining that vehicle C is in a predetermined driving state (straight driving state) and generating reference data the first time it travels after stopping, the information processing device 2 continues to apply the generated reference data without making this determination until vehicle C next stops.
- a driving state in which it is traveling at a predetermined speed or above and the steering angle is within a predetermined angle range, i.e., a state in which the straight driving state continues for a predetermined period or more, can occur multiple times on a regular basis.
- information processing device 2 determines whether vehicle C is in the predetermined driving state (straight driving state) the first time it travels after stopping, thereby preventing this determination from being made excessively.
- the reference data generated by the information processing device 2 is data obtained based on at least one of the operator's facial position, facial orientation, and body orientation contained in images (frame group) acquired during a predetermined period, i.e., a period during which the vehicle C maintains a straight-ahead state (a predetermined driving state).
- the reference data may be data obtained based on the operator's facial position, facial orientation, or body orientation contained in the images, or a combination of these data.
- the information processing device 2 may extract areas of body parts such as the face or upper body from the images captured during the predetermined period using an object detection model such as R-CNN or YOLO, or using a face detection model 201 (Facedetect), posture detection model 203 (Poseddetect), or posture estimation model (OpenPose), which are widely used as general-purpose learning models, and use these areas as reference data.
- an object detection model such as R-CNN or YOLO
- a face detection model 201 Face detection model 201 (Facedetect), posture detection model 203 (Poseddetect), or posture estimation model (OpenPose), which are widely used as general-purpose learning models, and use these areas as reference data.
- the facial area is indicated by a bounding box output by the object detection model, but an area obtained by expanding this area (bounding box) by, for example, 25% may be set as reference data (detection frame).
- the information processing device 2 may calculate the distance between the eyes on the operator's face and set this distance between the eyes as reference data indicating the orientation of the face.
- the information processing device 2 may calculate the distance between the shoulders of the operator's upper body and set this distance between the shoulders as reference data indicating the orientation of the body. In this way, by generating reference data using the position, orientation, or orientation of the operator's face, it is possible to ensure or improve the accuracy of the monitoring process.
- the images captured over a predetermined period of time are composed of a frame group including multiple frames (still images) arranged in chronological order according to the frame rate; that is, the images captured over the predetermined period of time include multiple images.
- the information processing device 2 may calculate data indicating the position of the face, the direction of the face, or the direction of the body for each of the multiple images captured over the predetermined period of time, and derive the average value of this calculated data as reference data.
- the information processing device 2 may average the coordinates of a bounding box obtained by object detection targeting the face for each of multiple frames (still images) included in a video captured over the predetermined period of time, to calculate an area indicating the position of the face when the operator is facing forward.
- the information processing device 2 may average the distance between the eyes on the face for each of multiple frames (still images) included in a video captured over the predetermined period of time, to obtain the distance between the eyes when the operator is facing forward.
- the information processing device 2 may average the distance between the shoulders of the upper body in each of multiple frames (still images) included in a video captured within a predetermined period to determine the distance between the shoulders when the operator is facing forward. While an averaging process is used to generate the reference data, this is not limited to this.
- the information processing device 2 may also generate the reference data using the median or mode of data (bounding box coordinates, distance between the eyes, distance between the shoulders) derived from each of multiple frames (still images) included in a video captured within a predetermined period.
- the reference data generated by the information processing device 2 is data obtained based on the degree of eye opening of the operator, i.e., the size of the eyes when they are wide open, contained in images (frame group) acquired within a predetermined period, i.e., a period during which the vehicle C maintains a straight-ahead state (a predetermined driving state).
- the information processing device 2 extracts a face area image showing the area of the operator's face contained in the acquired image, and detects the upper eyelid, lower eyelid, and points on the left and right of the eye, which form the periphery of the eye, for example, by inputting the face area image into a face landmark model 202 (e.g., MediaPope), which detects each part of the face.
- a face landmark model 202 e.g., MediaPope
- the information processing device 2 calculates the horizontal length of the eye using the distance between these detected points on the left and right of the eye, and calculates the vertical length of the eye using the distance between the upper eyelid and lower eyelid.
- the information processing device 2 may then derive the eye aspect ratio (vertical eye length/horizontal eye length) calculated by dividing the vertical length of the eye by the horizontal length of the eye, as data related to the degree of eye opening.
- data regarding the degree of eye opening (size) (eye aspect ratio) is calculated for each image (frame group) acquired while vehicle C maintains a straight-ahead state (predetermined driving state), and information processing device 2 derives the maximum value of the eye aspect ratio calculated for each of the images (frames) as reference data.
- information processing device 2 may generate reference data using the average, median, or mode of the eye aspect ratios calculated for each of the images (frames). In this way, when using the eye aspect ratio as data regarding the degree of eye opening (size) of the operator, applying the maximum value of the aspect ratio as reference data when performing monitoring processing on the operator makes it possible to efficiently detect, for example, drowsy driving based on the degree of eye opening (size) of the operator.
- (Embodiment 2) 12 is a flowchart illustrating the processing (calibration processing when an operator makes a change) of the control unit 20 of the information processing device 2 according to the second embodiment.
- the control unit 20 of the information processing device 2 performs the following processing, for example, when the vehicle C is in a started state.
- the control unit 20 of the information processing device 2 performs the processing from S201 to S204 and S2031, similar to the processing from S101 to S104 and S1031 in the first embodiment.
- the control unit 20 of the information processing device 2 determines whether the operator has changed (S205). Based on the image including the operator used last time, i.e., when generating the currently applied reference data, and the image including the operator acquired in the current process, the control unit 20 of the information processing device 2 determines whether the operator used when generating the previous reference data is different from the operator targeted in the current process, i.e., whether the operator has changed (been replaced).
- the control unit 20 of the information processing device 2 stores images acquired periodically or steadily from the camera 11 in the memory unit 23 in association with the time at which the image was captured, and identifies the image used when the previous reference data was generated based on the time at which the image was captured.
- the time at which the reference data was generated is also stored in the memory unit 23, and the reference data is an image captured during a period in which a straight-line driving state continued for a predetermined period or more.
- the control unit 20 of the information processing device 2 identifies an image captured during a period in which a straight-line driving state continued for a predetermined period or more.
- the control unit 20 of the information processing device 2 may derive the similarity between an image containing the operator used in the previous processing and an image containing the operator to be used in the current processing, for example, using an image similarity model that outputs the similarity between the two input images. Then, if the derived similarity is equal to or greater than a threshold indicating the same person, the control unit 20 of the information processing device 2 may determine that the operator has not changed and is the same person. Alternatively, the control unit 20 of the information processing device 2 may calculate image embedding vectors by vectorizing each of these two images, and if the cosine similarity of these image embedding vectors is equal to or greater than a threshold, determine that the operator has not changed and is the same person.
- control unit 20 of the information processing device 2 If it is determined that the operator has changed (S205: YES), the control unit 20 of the information processing device 2 generates reference data (S206). If it is determined that the operator has changed, that is, if the operator changed when vehicle C stopped this time, the control unit 20 of the information processing device 2 generates reference data in the same manner as S105 in embodiment 1.
- control unit 20 of the information processing device 2 does not generate reference data and continues to use the currently applied reference data (S2031). If it is determined that the operator has not changed, that is, if the operator was the same person when vehicle C stopped this time, the control unit 20 of the information processing device 2 does not generate reference data and continues to use the currently applied reference data, as in S1031 in embodiment 1.
- the control unit 20 of the information processing device 2 performs the processes from S208 to S209 in the same manner as the processes from S107 to S108 in embodiment 1.
- the information processing device 2 After the operator gets in and vehicle C is started (ignition switch is turned on) until it is stopped (ignition switch is turned off), the information processing device 2 generates reference data for the operator without performing image-based facial recognition each time it determines that the vehicle is in a predetermined driving state (after stopping, the initial straight-ahead state is maintained for a predetermined period of time). In this case, it is assumed that the operator may change (the driver may be changed) while vehicle C is running. In other words, if the operator (driver) is changed while vehicle C is stopped and idling, the information processing device 2 generates and updates reference data based on an image including the new operator and continues the monitoring process.
- the information processing device 2 determines that the vehicle is in a predetermined driving state, if the operator has not changed between the previous time the vehicle was determined to be in a driving state and the time the vehicle is currently determined to be in a driving state, i.e., if the operator is the same person, it continues to use the currently applied reference data without generating reference data based on an image including the operator acquired in the current determination. That is, the information processing device 2 generates reference data based on the image including the operator acquired in the current determination only if the operator has changed between the previous determination that the vehicle was in a predetermined driving state and the current determination that the vehicle was in a predetermined driving state.
- the information processing device 2 compares the image used when the previous determination that the vehicle was in a predetermined driving state with the image used when the current determination that the vehicle was in a predetermined driving state, and if the degree of match or similarity of the operator included in these images is equal to or greater than a predetermined value, determines that the operator has not changed, i.e., that the operator is the same.
- the information processing device 2 does not generate reference data and continues to use the currently applied reference data. This reduces the number of times the reference data is generated and reduces the computational load.
- the claims may be combined with one another regardless of the form of reference.
- the claims may contain multiple dependent claims that depend on multiple claims. Multiple dependent claims that depend on multiple dependent claims may also be contained. Even if a multiple dependent claim that depends on a multiple dependent claim is not contained, this does not limit the number of multiple dependent claims that depend on a multiple dependent claim.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Human Computer Interaction (AREA)
- Image Analysis (AREA)
Abstract
情報処理方法は、車両に搭載されたカメラによって前記車両を操作する操作者を含む画像を取得し、前記車両が予め定められた所定の走行状態にあるか否かを判定し、前記所定の走行状態にあると判定した場合、取得した前記画像に基づき、前記操作者に対する監視処理を行う際の基準データを生成する処理をコンピュータに実行させる。
Description
本発明は、情報処理方法、プログラム、及び情報処理装置に関する。
車両の運転者を撮像装置によって撮像して車両の進行方向を基準にしたときの運転者の顔向き角度を示す角度値を取得し、取得した角度値に基づいて運転者が前方不注視状態であるか否かを判定する安全運転判定装置が知られている(例えば特許文献1)。
一つの側面では、操作者に対する監視処理を行う際の基準データを効率的に生成することができる情報処理方法等を提供することを目的とする。
本開示の一態様に係る情報処理方法は、車両に搭載されたカメラによって前記車両を操作する操作者を含む画像を取得し、前記車両が予め定められた所定の走行状態にあるか否かを判定し、前記所定の走行状態にあると判定した場合、取得した前記画像に基づき、前記操作者に対する監視処理を行う際の基準データを生成する処理をコンピュータに実行させる。
本開示の一態様によれば、操作者に対する監視処理を行う際の基準データを効率的に生成する情報処理方法等を提供することができる。
(実施形態1)
以下、実施の形態について図面に基づいて説明する。図1は、実施形態1に係る情報処理装置2(カメラECU)を含むカメラユニット1の構成を例示する模式図である。図2は、情報処理装置2の物理構成を例示するブロック図である。カメラユニット1は、車両Cにおいて、例えばOHC(オーバヘッドコンソール)に搭載されたカメラ11、及び情報処理装置2を含み、情報処理装置2は、カメラ11が撮像した画像に対する画像処理を行い、当該画像処理結果を用いて、車両Cを操作する操作者に対する監視処理を行う。このようなカメラユニット1は、例えば、キャビンモニタシステムとして機能するものであってもよい。
以下、実施の形態について図面に基づいて説明する。図1は、実施形態1に係る情報処理装置2(カメラECU)を含むカメラユニット1の構成を例示する模式図である。図2は、情報処理装置2の物理構成を例示するブロック図である。カメラユニット1は、車両Cにおいて、例えばOHC(オーバヘッドコンソール)に搭載されたカメラ11、及び情報処理装置2を含み、情報処理装置2は、カメラ11が撮像した画像に対する画像処理を行い、当該画像処理結果を用いて、車両Cを操作する操作者に対する監視処理を行う。このようなカメラユニット1は、例えば、キャビンモニタシステムとして機能するものであってもよい。
カメラ11は、例えばCMOSカメラにて構成され、例えば、OHC(オーバヘッドコンソール)に配置される。本実施形態に係るカメラ11は1台以上であれば何台であってもよい。カメラ11は、例えばシリアルケーブルにて情報処理装置2と通信可能に接続されており、撮像した動画等の画像を情報処理装置2に周期的又はリアルタイムで出力する。なお、本明細書の各種処理における「リアルタイム」とは厳密な即時処理又は同時処理を意味するのではなく、可能な限り即応的に処理を実行することを意味する。
情報処理装置2は、制御部20、記憶部23、入出力I/F21、及び通信部22を含み、カメラ11が撮像した画像に対する画像処理を行い、当該画像処理結果に応じた処理を行うカメラECUとして機能する。制御部20は、CPU(Central Processing Unit)又はMPU(Micro Processing Unit)等により構成してあり、記憶部23に予め記憶された制御プログラムP(プログラム製品)及びデータを読み出して実行することにより、種々の制御処理及び演算処理等を行うようにしてある。
記憶部23は、RAM(Random Access Memory)等の揮発性のメモリ素子又は、ROM(Read Only Memory)、EEPROM(Electrically Erasable Programmable ROM)若しくはフラッシュメモリ等の不揮発性のメモリ素子により構成される。記憶部23には、制御プログラムP及び、制御部20が各種演算の処理時に参照するデータが予め記憶してある。更に、記憶部23には、顔検出モデル201、顔ランドマークモデル202、及び姿勢検出モデル203等、学習モデルの実体ファイルが記憶されているものであってもよい。これら学習モデルについては、後述する。
又は、顔検出モデル201、顔ランドマークモデル202、及び姿勢検出モデル203等の学習モデルはAIチップに実装されており、制御部20は、内部バス等を介して当該AIチップに対する入出力処理を行うことにより、これら学習モデルへの画像の入力、及び学習モデルから対象物に対する認識結果を取得するものであってもよい。
入出力I/F21は、例えばシリアル通信するための通信インターフェイスである。入出力I/F21を介して、情報処理装置2は、カメラ11等と通信可能に接続される。更に、入出力I/F21には、スピーカ又は音声出力機能を有するディスプレイ装置が接続されているものであってもよい。当該スピーカ等からは、情報処理装置2によって実行される監視処理の結果に基づき、警告音等が発せられる。
通信部22は、例えばCAN又はイーサネット(Ethernet/登録商標)等の通信プロトコルを用いた入出力インターフェイスである。制御部20は、通信部22を介して車載ネットワーク4に接続されている他のECU3と相互に通信してもよい。例えば、本実施形態に係る情報処理装置2は、図1及び図2に示すように、車載ネットワーク4を介してECU3と相互に通信し、これらECU3から、各種の通信データ(CANメッセージ、イーサネットフレーム等)を取得する。
図3は、顔検出モデル201の生成処理に関する説明図である。顔検出モデル201は、例えばAIチップ等に実装されて構成される。又は、顔検出モデル201は、制御部20によるソフトウェア機能部として構成されるものであってもよい。本実施形態では、情報処理装置2が顔検出モデル201を生成するものとして説明するが、これに限定されず、顔検出モデル201は、各種サーバ装置等のモデル生成装置によって生成されるものであってもよい。情報処理装置2の制御部20は、既存の訓練データ、又は顔検出モデル201用に用意した訓練データを用いて、学習前の機械学習モデル(例えば、YOLO又はR-CNN等のニューラルネットワーク等)を学習させることにより、画像(操作者を含む画像)を入力すると入力画像に含まれる対象物(操作者の顔)を認識(検出)する顔検出モデル201を生成してもよい。
顔検出モデル201が出力する、対象物に対する認識結果は、少なくとも認識できた対象物(顔)の有無を含む。また、認識できた対象物が有る場合は、前記認識結果は、当該対象物の画像における領域(バウンディングボックス)及び種類(クラス)を含む。また、前記認識結果は、認識できた対象物の信頼度スコア(各クラスの予測確率)を含んでいてもよい。
訓練データは、問題データ及び回答データを含み、カメラ11から取得した画像(操作者を含む画像)は、問題データに相当し、対象物(顔)に対する認識結果は、回答データに相当する。回答データである対象物(顔)の領域及び種類は、画像に対しアノテーション(付加)することにより、設定されるものであってもよい。顔検出モデル201を学習するための訓練データに含まれる問題データ及び回答データのデータセットと、顔検出モデル201を用いた際の入力データ及び出力データのデータセットとは同義であり、いずれかのデータセットにて定義されていれば、他方のデータセットにおいても、当然に適用される。
訓練データを用いて学習されたニューラルネットワーク(顔検出モデル201)は、人工知能ソフトウェアの一部であるプログラムモジュールとして利用が想定される。顔検出モデル201は、上述のごとく制御部20(CPU等)及び記憶部23を備える情報処理装置2にて用いられるものであり、このように演算処理能力を有する情報処理装置2にて実行されることにより、ニューラルネットワークシステムが構成される。すなわち、情報処理装置2の制御部20が、記憶部23に記憶された顔検出モデル201からの指令に従って、入力層に入力された画像の特徴量を抽出する演算を行い、対象物に対する認識結果を出力する。
顔検出モデル201は、例えばYOLO、又はR-CNN(Region Convolutional Neural Network)等にて構成され、画像の入力を受付ける入力層と、当該画像の特徴量を抽出する中間層と、対象物に対する認識結果を出力とする出力層とを有する。入力層は、画像の入力を受付ける複数のニューロンを有し、入力された値を中間層に受け渡す。中間層は、ReLU関数又はシグモイド関数等の活性化関数を用いて定義され、入力されたそれぞれの値の特徴量を抽出する複数のニューロンを有し、抽出した特徴量を出力層に受け渡す。当該活性化関数の重みづけ係数及びバイアス値等のパラメータは、誤差逆伝播法を用いて最適化される。出力層は、例えば全結合層により構成され、中間層から出力された特徴量に基づいて、対象物に対する認識結果を出力する。更に出力層は、例えばソフトマックス層を含み、対象物における確率(信頼度スコア)についても、出力するものであってもよい。
本実施形態では、顔検出モデル201は、R-CNN等であるとしたがこれに限定されず、R-CNN以外のニューラルネットワーク、トランスフォーマ、BERT、GPT、RNN(Recurrent Neural Network)、LSTM(Long-short term model)、SVM(Support Vector Machine)、ベイジアンネットワーク、線形回帰、回帰木、重回帰、ランダムフォレスト、アンサンブルなど、他の機械学習アルゴリズムで構築された顔検出モデル201であってよい。又は、顔検出モデル201は、ChatGPT等のLLMによる人工知能チャットボットを用いて構成されるものであってもよい。この場合、ChatGPTは、対象物に対する認識結果を効率的に出力するようにファインチューニング等がされているものであってもよい。又は、ChatGPTの入力インターフェイスであるプロンプト入力I/Fに対し、WebDB等の外部データベースを用いて生成した質問文(プロンプト)を、画像と共に入力するものであってもよい。
図4は、顔ランドマークモデル202の生成処理に関する説明図である。顔ランドマークモデル202は、顔検出モデル201と同様にAIチップ等に実装、又は制御部20によるソフトウェア機能部として構成される。顔ランドマークモデル202についても、顔検出モデル201と同様にYOLO又はR-CNN等にて構成されるものであってもよい。訓練データは、問題データ及び回答データを含み、カメラ11から取得した画像(操作者を含む画像)は問題データに相当し、顔における目、鼻等の特徴的部位を示すランドマークの領域又は位置、及び当該ランドマークが顔のどの部位を示すかの属性は、回答データに相当する。情報処理装置2の制御部20は、既存の訓練データ、又は顔ランドマークモデル202用に用意した訓練データを用いて、学習前の機械学習モデル(例えば、YOLO又はR-CNN等のニューラルネットワーク等)を学習させることにより、画像(操作者を含む画像)を入力すると入力画像に含まれる対象物(顔におけるランドマーク)を認識(検出)する顔ランドマークモデル202を生成するものであってもよい。
図5は、姿勢検出モデル203の生成処理に関する説明図である。姿勢検出モデル203は、顔検出モデル201と同様にAIチップ等に実装、又は制御部20によるソフトウェア機能部として構成される。姿勢検出モデル203についても、顔検出モデル201と同様にYOLO又はR-CNN等にて構成されるものであってもよい。訓練データは、問題データ及び回答データを含み、カメラ11から取得した画像(操作者を含む画像)は問題データに相当し、上半身における肩、肘、首、腰等の特徴的部位を示すランドマークの領域又は位置、及び当該ランドマークが人体のどの部位を示すかの属性は、回答データに相当する。情報処理装置2の制御部20は、既存の訓練データ、又は姿勢検出モデル203用に用意した訓練データを用いて、学習前の機械学習モデル(例えば、YOLO又はR-CNN等のニューラルネットワーク等)を学習させることにより、画像(操作者を含む画像)を入力すると入力画像に含まれる対象物(上半身におけるランドマーク)を認識(検出)する姿勢検出モデル203を生成するものであってもよい。
図6は、情報処理装置2の制御部20の処理(キャリブレーション処理)を例示するフローチャートである。情報処理装置2の制御部20は、例えば車両Cが起動状態において以下の処理を行う。情報処理装置2の制御部20は、以下の処理の実行中において、操作者を撮像した動画又は静止画の画像をカメラ11から定常的又は周期的に取得し、取得した画像を、当該画像の撮像時刻等を示すタイムスタンプを付与して、記憶部23に記憶している。このように取得した画像(動画のフレーム)それぞれに対し、当該画像の撮像時点を関連付けて記憶することにより、連続する複数時点にて撮像された画像を、時系列に並べて記憶部23に記憶(保存)することができる。
情報処理装置2の制御部20は、車両Cが停車状態から、走行を開始したか否かを判定する(S101)。情報処理装置2の制御部20は、例えば、車載ネットワーク4を介して、車速等を含むメッセージを周期的に出力するECU3にから、当該メッセージを周期的に取得する。情報処理装置2の制御部20は、取得した車速等を含むメッセージに基づき、車両Cが、停車した以降、すなわち車速が実質的に0km/hとなったの後、走行を開始したか否か、すなわち車速が0km/hよりも大きくなったか否かに応じて、車両Cが停車状態から、走行を開始したか否かを判定する。又は、情報処理装置2の制御部20は、車両Cが停車状態から走行を開始した際、当該走行開始の旨を示す事項を含むメッセージ(走行開始メッセージ)を出力するECU3にから、当該メッセージを取得した場合、車両Cが停車状態から走行を開始したと判定するものであってもよい。
車両Cが停車状態から、走行を開始した場合(S101:YES)、情報処理装置2の制御部20は、車速及び舵角を取得する(S102)。情報処理装置2の制御部20は、車載ネットワーク4を介して、車速を管理するECU3から、車速等を含むメッセージを周期的に取得する。更に、情報処理装置2の制御部20は、車載ネットワーク4を介して、舵角を管理するECU3から、舵角等を含むメッセージを周期的に取得する。情報処理装置2の制御部20は、取得した車速及び舵角を、当該車速又は舵角を含むメッセージを取得した取得時点と関連付けて、記憶部23に記憶する。このように車速及び舵角と取得時点とを関連付けて記憶することにより、連続する複数時点にて取得された車速及び舵角の値を、時系列に並べた履歴データとして、記憶部23に記憶(保存)することができる。
情報処理装置2の制御部20は、車速及び舵角に基づき、車両Cが直進状態であるか否かを判定する(S103)。情報処理装置2の制御部20は、周期的に取得した車速等を含むメッセージ及び舵角等を含むメッセージに基づき、例えば、車速が予め定められた所定速度以上であり、かつ、舵角が予め定められた所定角度以内である場合、車両Cが直進状態であると判定する。当該所定速度は、例えば5km/hであり、当該所定角度は、例えば直進時を0°基準とした際の±5°であり、これらは所定速度、及び所定角度等、制御部20が各種演算を行う際の所定値又は閾値は、予め記憶部23に記憶されている。
車両Cが直進状態である場合(S103:YES)、情報処理装置2の制御部20は、直進状態が所定期間以上、継続したか否かを判定する(S104)。情報処理装置2の制御部20は、当該車両Cが、予め定められた直進状態(車速が所定速度以上かつ舵角が所定角度以内で)に該当すると判定した際、当該直進状態が、例えば5秒等、予め定められた期間にて継続(維持)されたか否かを判定する。
直進状態が所定期間以上継続した場合(S104:YES)、情報処理装置2の制御部20は、基準データを生成する(S105)。直進状態が所定期間以上、継続した場合、情報処理装置2の制御部20は、当該所定期間の開始時点及び終了時点を特定する。情報処理装置2の制御部20は、直進状態が所定期間以上、継続したか否かを判定する際に用いた車速及び舵角それぞれに付随する取得時点それぞれに応じて、開始時点及び終了時点を特定する。この際、情報処理装置2の制御部20は、S103の処理にて、車速及び舵角に基づき車両Cが直進状態であると判定した時点を開始時点し、当該直進状態が所定期間以上継続したと判定した時点を終了時点として、特定するものであってもよい。
情報処理装置2の制御部20は、特定した所定期間の開始時点及び終了時点に応じて、当該所定期間にて撮像された画像を取得する。記憶部23には、カメラ11から取得した画像が、撮像時間と関連付けられて記憶されており、情報処理装置2の制御部20は、直進状態が継続した所定期間において撮像された画像、すなわち所定期間の開始時点から終了時点までの間に撮像された複数の画像を取得する。画像が動画の場合、当該複数の画像は、動画を構成する複数のフレーム(フレーム群)に相当する。
情報処理装置2の制御部20は、所定期間にて撮像された複数の画像(フレーム群)において、これら画像それぞれに含まれている操作者の顔の位置、顔の向き、身体の向き、又は目の開度を抽出する。情報処理装置2の制御部20は、これら画像それぞれにて抽出した顔の位置それぞれの平均値、中央値、又は最頻値を算出し、当該平均値を顔の位置を示す基準データ(顔の位置用基準データ)として生成する。
情報処理装置2の制御部20は、これら画像それぞれにて抽出した顔の向きそれぞれの平均値、中央値、又は最頻値を算出し、当該平均値を顔の向きを示す基準データ(顔の向き用基準データ)として生成する。情報処理装置2の制御部20は、これら画像それぞれにて抽出した身体の向きそれぞれの平均値、中央値、又は最頻値を算出し、当該平均値を身体の向きを示す基準データ(身体の向き用基準データ)として生成する。情報処理装置2の制御部20は、これら画像それぞれにて抽出した目の開度の最大値を算出し、当該最大値を目の開度を示す基準データ(目の開度用基準データ)として生成する。情報処理装置2の制御部20は、操作者の顔の位置、顔の向き、身体の向き、及び目の開度の全てにおいて、基準データを生成するものであってもよく、いずれか1以上の基準データを生成するものであってもよい。すなわち、情報処理装置2の制御部20は、実行する監視処理の種類に応じて、操作者の顔の位置、顔の向き、身体の向き、及び目の開度のうち、少なくとも1つ以上の基準データを生成するものであってもよい。以下、これら操作者の顔の位置、顔の向き、身体の向き、及び目の開度の基準データそれぞれについて、説明する。
図7は、操作者を含む画像における顔の位置(バウンディングボックス)の検出結果を示す説明図である。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれを、顔検出モデル201に入力することにより、当該画像それぞれにおける顔の位置を検出する。当該顔検出モデル201は、例えばFacedetectモデルであり、操作者を含む画像を入力した場合、当該操作者の顔を検出し、検出した顔の位置、すなわち画像における顔の領域を示す矩形の枠体(バウンディングボックス)を、当該画像に重畳させて出力する。当該バウンディングボックスは、画像の座標系における、例えば、左上の頂点と、右下の頂点との座標にて定義される。
情報処理装置2の制御部20は、車両Cが直進状態にある所定期間内において撮像された複数の画像(フレーム群)それぞれにおけるバウンディングボックスを抽出し、これら抽出した複数のバウンディングボックスにおける平均値となるバウンディングボックスを導出する。情報処理装置2の制御部20は、当該バウンディングボックスの平均値を導出するにあたり、個々のバウンディングボックスの左上及び右下の頂点の座標値の平均値を算出することにより、平均値となるバウンディングボックスの左上及び右下の頂点の座標値を導出するものであってもよい。又は、情報処理装置2の制御部20は、個々のバウンディングボックスの重心又は中心点それぞれを算出し、算出した重心等それぞれにおけるX及びY座標値の平均値を算出することにより、平均値となるバウンディングボックスの重心の座標を導出するものであってもよい。
情報処理装置2の制御部20は、算出したバウンディングボックスの平均値を、基準データ(顔の位置用基準データ)として記憶部23に記憶する。又は、情報処理装置2の制御部20は、算出したバウンディングボックスの平均値に対し、例えば、25%等にて拡張した領域を基準データ(検出用の枠)として設定するものであってもよい。本実施形態においては、当該拡張した領域を基準データ(顔の位置用基準データ)として用いるものとする。情報処理装置2の制御部20は、生成した基準データを記憶部23に記憶する際、当該生成時点を示すタイムスタンプ等を付与して記憶するものであってもよい。
このように25%等拡張した領域にて設定された基準データ(顔の位置用基準データ)は、操作者に対する監視処理にて用いられるものとなる。情報処理装置2の制御部20は、監視処理の実行中に取得した操作者を含む画像に対し、顔検出モデル201を用いて、当該画像における操作者の顔の位置を検出し、検出した顔の位置が、基準データが示す領域から、例えば25%以上外れると前を向いていないと判断し、例えば70%以上外れた状態が3秒継続すると、居眠り運転状態であるとして警告音等を発するものであってもよい。
図8は、操作者を含む画像における顔の向き(両目間の距離)の検出結果を示す説明図である。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれを、顔ランドマークモデル202に入力することにより、当該画像それぞれにおける顔の向きを検出する。又は、情報処理装置2の制御部20は、操作者を含む画像を顔検出モデル201に出力した際、当該顔検出モデル201が検出した顔の領域(バウンディングボックス)の画像(顔領域画像)を抽出し、抽出した顔領域画像を顔ランドマークモデル202に入力するものであってもよい。当該顔ランドマークモデル202は、例えば、Facelandmarkモデル又はMediapipeモデルであり、顔を含む画像を入力した場合、当該顔における目、鼻、口等の特徴的部位となる複数のランドマークを検出し、これらランドマークそれぞれの示す点を画像に重畳させて出力する。これら検出したランドマークには、画像座標系の座標値及び、顔の各部位(目、鼻、口等)を示す属性が付与されている。
情報処理装置2の制御部20は、左右の目を示す2つのランドマークを特定し、特定した2つのランドマーク(目を示すランドマーク)間の画像座標系における距離を算出することにより、当該左右の目の距離を算出する。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれにおいて算出した両目間の距離の平均値を算出することにより、当該平均値を基準データ(顔の向き用基準データ)として導出する。
このように導出された両目間の距離の平均値にて設定された基準データ(顔の向き用基準データ)は、操作者に対する監視処理にて用いられるものとなる。情報処理装置2の制御部20は、監視処理の実行中に取得した操作者を含む画像に対し、顔ランドマークモデル202を用いて、当該画像における操作者の両目の距離を検出する。情報処理装置2の制御部20は、検出した両目の距離が、例えば、基準データよりも長い場合、操作者は左に向けていると判断し、基準データよりも短い場合、操作者は右に向けていると判断し、左右いずれかを向いている時間が、例えば3秒等の所定期間以上継続した場合、わき見運転状態であるとして警告音等を発するものであってもよい。
図9は、操作者を含む画像における身体の向き(両肩間の距離)の検出結果を示す説明図である。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれを、姿勢検出モデル203に入力することにより、当該画像それぞれにおける操作者の姿勢を検出する。当該姿勢検出モデル203は、例えば、Posedetectモデルであり、操作者を含む画像を入力した場合、当該操作者の上半身における肩等の関節部等の特徴的部位となる部位検出し、これら部位それぞれの示す点を画像に重畳させて出力する。これら検出した部位には、画像座標系の座標値及び、身体の各部位(肩、肘、手首、首、腰等)を示す属性が付与されている。
情報処理装置2の制御部20は、例えば、左右の肩を示す2つの部位を特定し、特定した2つ部位(両肩)間の画像座標系における距離を算出することにより、左右の肩間の距離を算出する。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれにおいて算出した両肩間の距離の平均値を算出することにより、当該平均値を基準データ(身体の向き用基準データ)として導出する。
このように導出された両肩間の距離の平均値にて設定された基準データ(身体の向き用基準データ)は、操作者に対する監視処理にて用いられるものとなる。情報処理装置2の制御部20は、監視処理の実行中に取得した操作者を含む画像に対し、姿勢検出モデル203を用いて、当該画像における操作者の両肩間の距離を検出する。情報処理装置2の制御部20は、検出した両肩間の距離が、例えば、基準データよりも長い場合、操作者は左に向けていると判断し、基準データよりも短い場合、操作者は右に向けていると判断し、左右いずれかを向いている時間が、例えば3秒等の所定期間以上継続した場合、わき見運転状態であるとして警告音等を発するものであってもよい。
図10は、操作者を含む画像における目の開度(アスペクト比)の検出結果を示す説明図である。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれを、顔ランドマークモデル202に入力することにより、当該画像それぞれにおける目の開度を検出する。又は、情報処理装置2の制御部20は、操作者を含む画像を顔検出モデル201に出力した際、当該顔検出モデル201が検出した顔の領域(バウンディングボックス)の画像(顔領域画像)を抽出し、抽出した顔領域画像を顔ランドマークモデル202に入力するものであってもよい。当該顔ランドマークモデル202は、例えば、Facelandmarkモデル又はMediapipeモデルであり、顔を含む画像を入力した場合、当該顔における目、鼻、口等の特徴的部位となる複数のランドマークを検出し、これらランドマークそれぞれの示す点を画像に重畳させて出力する。本実施形態においては、顔ランドマークモデル202は、目の周囲又は輪郭に沿った複数のランドマーク(P1,P2,P3,P4,P5,P6)を検出する。これら検出したランドマークには、画像座標系の座標値及び、目における部位の名称(上まぶた、下まぶた等)を示す属性が付与されている。
情報処理装置2の制御部20は、検出された複数のランドマーク(P1,P2,P3,P4,P5,P6)から、目の長さ(横方向長さ)、上瞼から下瞼までの距離(縦方向長さ:目の高さ)を算出し、目の縦方向長さを、目の横方向長さで除算することにより算出した目のアスペクト比(目の縦方向長さ/目の横方向長さ)を、目の開度に関するデータとして算出する。本実施形態においては、上瞼から下瞼までの距離(縦方向長さ)を2つの箇所(P2-P6,P3-P5)における距離を合算しているため、目の長さ(横方向長さ)を2倍した値にて除算「{目の高さ(P2-P6の距離)+(P3-P5の距離)}/目の長さ(P1-P4の距離)×2」している。これにより、”目の高さ“と”目の長さ“の割合を示す目の開度計算(EAR:Eye Aspect Ratio(目のアスペクト比))を行うことができる。情報処理装置2の制御部20は、複数の画像(フレーム群)それぞれにおいて算出した目の開度(アスペクト比)において、最大値を基準データ(目の開度用基準データ)として導出する。
このように導出された目の開度(アスペクト比)の最大値にて設定された基準データ(目の開度用基準データ)は、操作者に対する監視処理にて用いられるものとなる。情報処理装置2の制御部20は、監視処理の実行中に取得した操作者を含む画像に対し、顔ランドマークモデル202を用いて、当該画像における操作者の目の開度(アスペクト比)を検出する。情報処理装置2の制御部20は、検出した目の開度と、基準データとを対比し、例えば、検出した目の開度が基準データ(目の開度の最大値)の90%未満である場合、警告音等を発するものであってもよい。情報処理装置2の制御部20は、左右いずれかの認識できた目において、当該監視処理を行うものであってもよい。
更に、情報処理装置2の制御部20は、顔の向き(上下方向)により、基準データ(目の開度用基準データ)を可変させるものであってもよい。情報処理装置2の制御部20は、顔ランドマークモデル202にて検出したランドマークにより、操作者の顔の向きが、上方向又は下方向であるかを判定するものであってもよい。
情報処理装置2の制御部20は、操作者の顔の向きが下方向である場合、基準データ(目の開度用基準データ)を例えば、10%減少させるものであってもよい。操作者が下を向いた場合、目の開き具合が小さく見えるため、目を閉じていると判定され易いものとなることが懸念されるところ、基準データ(目の開度用基準データ)を減少(閾値を厳しめに変更)することにより、誤判定が発生することを抑制することができる。情報処理装置2の制御部20は、操作者の顔の向きが上方向である場合、基準データ(目の開度用基準データ)を例えば、10%増加させるものであってもよい。操作者が上を向いた場合、目の開き具合が大きく見えるため、目を閉じていると判定されにくいものとなることが懸念されるところ、基準データ(目の開度用基準データ)を増加(閾値を緩めに変更)することにより、誤判定が発生することを抑制することができる。
今回の基準データの生成が、車両Cが起動した以降、最初の生成である場合、当該生成された基準データは、最初の基準データとして、記憶部23に記憶される。既に基準データが生成され、基準データが記憶部23に既に記憶されており、現在実行中の監視処理に適用されている場合、情報処理装置2の制御部20は、現在適用されている基準データに代えて、今回生成した基準データを適用、すなわち記憶部23に記憶されている現状の基準データを、今回生成した基準データにて上書きすることにより、監視処理に用いる基準データを更新する。
本実施形態において、情報処理装置2の制御部20が基準データを生成するとしたが、これに限定されない。基準データは、情報処理装置2と無線にて通信可能に接続されるクラウドサーバ等の外部サーバによって生成されるものであってもよい。この際、情報処理装置2は、例えば、車両Cに搭載されるTCU(Telematics Control Unit)及びインターネット等に外部ネットワークを介して、外部サーバに対し、操作者を含む画像を出力し、基準データを生成する旨の指示を行うものであってよい。外部サーバは、情報処理装置2からの画像に基づき基準データを生成し、生成した基準データを情報処理装置2に出力する。情報処理装置2の制御部20は、外部サーバからの基準データを取得することにより、当該基準データを生成するものであってもよい。
車両Cが直進状態でない場合(S103:NO)、直進状態が所定期間以上継続しなかった場合又は(S104:NO)、情報処理装置2の制御部20は、基準データを生成せず、現在適用されている基準データを継続して使用する(S1031)。車両Cが直進状態でない場合、又は直進状態が所定期間以上継続しなかった場合、操作者は、前方を向いていることを想定することが困難となるため、情報処理装置2の制御部20は、基準データを生成せず、現在適用されている基準データを継続して使用する。
車両Cが停車状態から、走行を開始した状態ではない場合(S101:NO)、S105又はS1031の実行後、情報処理装置2の制御部20は、基準データを用いて、操作者に対する監視処理を実行する(S106)。車両Cが停車状態から走行を開始した状態ではない場合、すなわち停車後、最初の直進状態が所定期間維持されることにより基準データを生成した後の走行中においては、情報処理装置2の制御部20は、当該基準データを継続的に適用することにより、監視処理を実行する。情報処理装置2の制御部20は、車両Cが停車状態から走行を開始した際、基準データが未作成である場合、例えば、予め設定されている初期値(デフォルト値)を用いて、監視処理を実行するものであってもよい。
情報処理装置2の制御部20は、上述のとおり、生成した基準データの種類(顔の位置用基準データ、顔の向き用基準データ、身体の向き用基準データ、目の開度用基準データ)に応じて、当該基準データの種類に対応した監視処理を実行する。更に、情報処理装置2の制御部20は、2つ以上の種類の基準データを組み合わせて、監視処理を実行するものであってもよい。
図11は、情報処理装置2の制御部20の処理(監視処理)を例示するフローチャートである。情報処理装置2の制御部20は、処理S106のサブルーチンとして以下のフローチャートにて示される処理を、監視処理の一例として行うものであってよい。
情報処理装置2の制御部20は、取得した画像に基づき、操作者の目が閉じているか否かを判定する(T101)。情報処理装置2の制御部20は、顔ランドマークモデル202を用いて、操作者の目が閉じている(Eye Closed)かを検出し、当該検出結果と仮判断として導出する。操作者の目が閉じていないと判定した場合(T101:NO)、情報処理装置2の制御部20は、操作者のまばたき、又は目の開度検知を実行する(T1011)。
操作者の目が閉じていると判定した場合(T101:YES)、情報処理装置2の制御部20は、操作者が正面を向いているか否かを判定する(T102)。操作者の目が閉じていると判定した場合、情報処理装置2の制御部20は、顔ランドマークモデル202の検出結果に基づき、顔の向きを確認し、顔の向きが正面であるか、すなわち極端又は過度に下を向いていないか否かを判定する。
操作者が正面を向いていないと判定した場合(T102:NO)、情報処理装置2の制御部20は、操作者の顔の向きが不適切であると判断する(T1022)。操作者が正面を向いていないと判定した場合、すなわち操作者が左右いずれかを向いていないと判定した場合、情報処理装置2の制御部20は、操作者の顔の向きが不適切であり、わき見運転をしていると判断する。
操作者が正面を向いていると判定した場合(T102:YES)、情報処理装置2の制御部20は、所定期間継続して、操作者の目の開度が小さいか否かを判定する(T103)。操作者が正面を向いていると判定した場合、情報処理装置2の制御部20は、例えば2秒間等、所定期間継続して、操作者の目の開度(EAR)が小さいか否かを判定する。所定期間継続して、操作者の目の開度が小さいものではない場合(T103:NO)、情報処理装置2の制御部20は、操作者は寝ていないと判定する(T104)。所定期間継続して、操作者の目の開度が小さい場合(T103:YES)、情報処理装置2の制御部20は、警告音を発生させる(T105)。
情報処理装置2の制御部20は、所定期間継続して、操作者の目の開度が小さいか否かを判定する(T106)。所定期間継続して、操作者の目の開度が小さいものではない場合(T106:NO)、情報処理装置2の制御部20は、操作者は寝ていないと判定する(T107)。所定期間継続して、操作者の目の開度が小さい場合(T106:YES)、情報処理装置2の制御部20は、操作者は寝ていると判定する(T1061)。情報処理装置2の制御部20は、音による警告後、継続して操作者の目の開度(EAR)が小さいと寝ていると判断し、大きければ寝ていないと判断する。
このような処理を行うことにより、情報処理装置2の制御部20は、顔や姿勢の揺れ、車両Cからのハンドルの舵角情報と組合せ、操作者が本当に寝ているか否かを効率的に判断することができる。すなわち、顔は正面を向いているが、目線が下を向いている場合であっても、寝てはいないが、目を閉じていると判断し、所定期間経過した後、居眠り検知として警告してしまうことを抑制することができる。
情報処理装置2の制御部20は、車両Cが停止されたか否かを判定する(S107)。情報処理装置2の制御部20は、例えばイグニッションスイッチを管理するECU3にから、当該イグニッションスイッチのオン(起動)又はオフ(停止)を示すデータを含むメッセージを周期的に取得しており、取得したメッセージを参照することにより、車両Cが停止された(イグニッションスイッチ:オフ)か否かを判定する。車両Cが停止されていないと判定した場合(S107:NO)、情報処理装置2の制御部20は、再度S101からの処理を実行すべくループ処理を行う。このようにループ処理を行うことにより、情報処理装置2の制御部20は、車両Cが起動中(イグニッションスイッチ:オン)において、走り初めて最初の直進が行われた場合において、基準データを生成し更新する処理を繰り返すことができる。
車両Cが停止された場合(S107:YES)、情報処理装置2の制御部20は、記憶部23に記憶されている基準データを消去する(S108)。車両Cが停止(イグニッションスイッチ:オフ)された場合、情報処理装置2の制御部20は、記憶部23に記憶されている基準データを消去する、又は基準データを初期化するものであってもよい。
本実施形態によれば、情報処理装置2は、車両Cに搭載されたカメラ11によって車両Cを操作する操作者を含む画像を継続的又は周期的に取得する。当該画像は、静止画又は、所定のフレームレートで撮像される動画であってもよい。情報処理装置2は、更に、例えば車載ネットワーク4を介して通信可能に接続される各種のECU3から取得したCANメッセージ等に基づき、車両Cの走行状態に関する情報(車両C状態情報)を取得する。当該車両C状態情報は、例えば、車速及び舵角(ステアリング軸の角度:操舵角)を含む。情報処理装置2は、取得した車両C状態情報に基づき、車速及び舵角にて定められる走行状態を認識することができる。情報処理装置2は、車両Cの走行状態が、所定速度(例えば5km/h)以上での直進状態であり、当該直進状態が所定期間(例えば5秒)以上継続したと判定した場合、当該所定期間、すなわち直進状態が継続した期間にて撮像した画像に基づき、操作者に対する監視処理を行う際の基準データを生成する。直進状態であるか否かについては、舵角が所定角度範囲(例えば、直進時を0°基準とした際の±5°)以内である場合、直進状態と判定するものであってもよい。車両Cが直進状態であることを判定するための所定速度(車速)、所定角度範囲(舵角)、及び所定期間(状態継続時間)等のパラメータ(判定パラメータ)は、例えば操作者の入力操作等により設定可変に構成されるものであってもよい。このように情報処理装置2は、車両Cが走行を開始した後、車速及び舵角にて予め定められた走行状態である直進状態が維持された期間において、撮像された操作者の画像(当該期間における動画を構成する複数のフレーム)に基づき、監視処理を行う際の基準データを生成するため、当該基準データを予め登録する処理を不要とすることができる。すなわち、車両Cに搭載されたカメラ11によって車両Cの操作者を撮像した画像を用いて、例えば、わき見運転又は居眠り運転等を検知しアラートを出力する監視処理を行うにあたり、当該操作者の顔又は上半身等を撮像した画像によるユーザ登録が要求される場合も想定されるところ、本実施形態においては、事前のユーザ登録を不要とすることができる。従って、操作者の顔を含む画像を用いた顔認証が、個人特定等の観点から禁止される規約が存在する場合であっても、本実施形態による処理を行うことにより、顔認証を要するユーザ登録を行うことなく、基準データを生成し、当該基準データを用いて監視処理を実行することができる。車速及び舵角にて予め定められた走行状態は、実質的に車両Cの直進状態に相当するものであり、当該直進状態の継続中は、操作者は、前方を向いている傾向にあることが想定される。従って、当該所定の走行状態(直進状態)が継続した期間にて、操作者を撮像した動画に含まれる各フレーム(フレーム群)においては、操作者が前方を向いているフレームが多数を占めることが想定される。このように操作者が前方を向いている複数の画像(動画を構成するフレーム群)を用いて基準データを生成するため、当該基準データは、操作者が前方を向いている際でのデータに相当する。わき見運転又は居眠り運転等を対象とする監視処理において、操作者の顔の向き、又は上半身等の身体の向きが、前方から左右に角に外れた場合、アラートが発生されるところ、基準データは、操作者が前方を向いている際でのデータに相当するため、当該監視処理の精度を担保又は向上させることができる。
本実施形態によれば、情報処理装置2は、車両Cが停車した以降、最初に走行した際に所定の走行状態にあるか否かを判定するものであり、すなわち、走り初めて最初の直進が行われた場合において、所定の走行状態にあると判定する。従って、情報処理装置2は、車両Cが停止した以降、最初に走行した際に、所定の走行状態(直進状態)であると判定して基準データを生成した後は、車両Cが次に停止するまで当該判定を行うことなく、生成した基準データの適用を継続する。車両Cの走行時において、所定速度以上であり、かつ舵角が所定角度範囲内となる走行状態、すなわち直進状態が所定期間以上継続する状態は、定常的に複数回、発生し得る。これに対し、情報処理装置2は、車両Cの停止後、最初に走行した際に、当該所定の走行状態(直進状態)であるか判定するため、当該判定が過度に行われることを防止することができる。
本実施形態によれば、情報処理装置2が生成する基準データは、所定期間内、すなわち車両Cが直進状態(所定の走行状態)を維持している期間において取得した画像(フレーム群)において、当該画像に含まれる操作者の顔の位置、顔の向き、及び身体の向きの内、少なくとも1つに基づき得られるデータである。すなわち、基準データは、画像に含まれる操作者の顔の位置、顔の向き、又は身体の向きに基づき得られるデータ、又は、これらデータの組合せによるものであってもよい。情報処理装置2は基準データを生成するにあたり、当該所定期間にて撮像した画像に対し、例えばR-CNN又はYOLO等にて構成した物体検出モデルを用いて、又は汎用的な学習モデルとして流布されている顔検出モデル201(Facedetect)、姿勢検出モデル203(Poseddetect)又が姿勢推定モデル(OpenPose)を用いて、顔又は上半身等の身体の部位の領域を抽出し、当該領域を基準データとするものであってもよい。この際、顔の領域は、物体検出モデルが出力するバウンディングボックスにて示されるものとなるが、当該領域(バウンディングボックス)に対し、例えば25%等にて拡張した領域を基準データ(検出用の枠)として設定するものであってもよい。又は、情報処理装置2は基準データを生成するにあたり、操作者の顔における両目の距離を演算し、当該両目間距離を、顔の向きを示す基準データとして設定するものであってよい。又は、情報処理装置2は基準データを生成するにあたり、操作者の上半身における両肩の距離を演算し、当該両肩間距離を、身体の向きを示す基準データとして設定するものであってよい。このように操作者の顔の位置、顔の向き、又は身体の向きを用いて基準データを生成することにより、監視処理の精度を担保又は向上させることができる。
本実施形態によれば、所定期間にて撮像された画像は、フレームレートに応じて、時系列に並ぶ複数のフレーム(静止画)を含むフレーム群にて構成されるものであり、すなわち当該所定期間における画像は、複数枚の画像を含む。情報処理装置2は、これら所定期間における複数枚の画像それぞれにおいて、顔の位置、顔の向き、又は身体の向きを示すデータを算出し、これら算出したデータの平均値を、基準データとして導出するものであってもよい。情報処理装置2は、顔の位置に基づく基準データを生成するにあたり、所定期間内での動画に含まれる複数のフレーム(静止画)それぞれにおける、顔を対象とした物体検出によるバウンディングボックスの座標を平均化して、操作者が前方を向いている際の顔の位置を示す領域を算出するものであってよい。又は、情報処理装置2は、顔の向きに基づく基準データを生成するにあたり、所定期間内での動画に含まれる複数のフレーム(静止画)それぞれにおける、顔における両目の距離を平均化して、操作者が前方を向いている際の両目の距離を出するものであってよい。又は、情報処理装置2は、身体の向きに基づく基準データを生成するにあたり、所定期間内での動画に含まれる複数のフレーム(静止画)それぞれにおける、上半身における両肩の距離を平均化して、操作者が前方を向いている際の両肩の距離を出するものであってよい。なお、基準データの生成化において、平均化処理を用いるとしたが、これに限定されず、情報処理装置2は、所定期間内での動画に含まれる複数のフレーム(静止画)それぞれに基づき導出したデータ(バウンディングボックスの座標、両目の距離、両肩の距離)における中央値、又は最頻値を用いて、基準データを生成するものであってもよい。このように所定期間内での動画に含まれる複数のフレーム(静止画)それぞれにおいて、操作者の顔の位置、向き又は身体の向きに多少の偏倚がある場合であっても、平均化等の処理を行うことにより当該偏倚による影響を緩和することができ、操作者が前方を向いている際の基準データを好適に生成することができる。
本実施形態によれば、情報処理装置2が生成する基準データは、所定期間内、すなわち車両Cが直進状態(所定の走行状態)を維持している期間において取得した画像(フレーム群)において、当該画像に含まれる操作者の目の開度、すなわち目が見開かれた際の大きさに基づき得られるデータである。情報処理装置2は、取得した画像に含まれる操作者の顔の領域を示す顔領域画像を抽出し、例えば顔における各部位を検出する顔ランドマークモデル202(例えばMediaPope)に、当該顔領域画像を入力することにより、目の周辺部となる上瞼、下瞼、及び目の左右の地点を検出する。情報処理装置2は、これら検出した目の左右の地点間の距離にて目の横方向長さを算出し、上瞼と下瞼との距離にて目の縦方向長さを算出する。その上で、情報処理装置2は、目の縦方向長さを、目の横方向長さで除算することにより算出した目のアスペクト比(目の縦方向長さ/目の横方向長さ)を、目の開度に関するデータとして導出するものであってもよい。この際、目の開度(大きさ)に関するデータ(目のアスペクト比)は、車両Cが直進状態(所定の走行状態)を維持している期間において取得した画像(フレーム群)それぞれにおいて算出されるものとなり、情報処理装置2は、当該画像(フレーム)それぞれにて算出した目のアスペクト比の最大値を、基準データとして導出する。又は、情報処理装置2は、当該画像(フレーム)それぞれにて算出した目のアスペクト比の平均値、中央値、又は最頻値を用いて、基準データを生成するものであってもよい。このように操作者の目の開度(大きさ)に関するデータとして目のアスペクト比を用いるあたり、当該アスペクト比の最大値を、操作者に対する監視処理を行う際の基準データとして適用することにより、操作者の目の開度(大きさ)に基づき、例えば居眠り運転等の検知を効率的に行うことができる。
(実施形態2)
図12は、実施形態2に係る情報処理装置2の制御部20の処理(操作者が変更時、キャリブレーション処理)を例示するフローチャートである。情報処理装置2の制御部20は、例えば車両Cが起動状態において以下の処理を行う。情報処理装置2の制御部20は、実施形態1の処理S101からS104、S1031と同様にS201からS204、S2031までの処理を行う。
図12は、実施形態2に係る情報処理装置2の制御部20の処理(操作者が変更時、キャリブレーション処理)を例示するフローチャートである。情報処理装置2の制御部20は、例えば車両Cが起動状態において以下の処理を行う。情報処理装置2の制御部20は、実施形態1の処理S101からS104、S1031と同様にS201からS204、S2031までの処理を行う。
情報処理装置2の制御部20は、操作者が変更したか否かを判定する(S205)。情報処理装置2の制御部20は、前回、すなわち現時点にて適用している基準データを生成する際に用いた操作者を含む画像と、今回の処理において取得した操作者を含む画像とに基づき、前回の基準データ生成時の操作者と、今回の処理にて対象となる操作者とが、異なるか、すなわち操作者が変更(交替)したか否かを判定する。
情報処理装置2の制御部20は、カメラ11から周期的又は定常的に取得した画像を、当該画像の撮像時点と関連付けて記憶部23に記憶しており、前回の基準データ生成時にて用いた画像を撮像時点により特定する。基準データの生成時点についても記憶部23に記憶されており、当該基準データは、直進状態が所定期間以上継続した期間にて撮像された画像である。情報処理装置2の制御部20は、今回の処理において、直進状態が所定期間以上継続した期間にて撮像された画像を特定する。
情報処理装置2の制御部20は、前回の処理にて用いた操作者を含む画像と、今回の処理にて用いる操作者を含む画像とを、例えば、入力された2つの画像の類似を出力する画像類似モデルを用いて、これら画像の類似度を導出するものであってもよい。その上で、情報処理装置2の制御部20は、導出した類似度が、同一人を示す閾値以上となる場合、操作者は変更されておらず、同一人であると判定するものであってもよい。又は、情報処理装置2の制御部20は、これら2つの画像それぞれをベクトル化した画像埋め込みベクトルそれぞれを算出し、これら画像埋め込みベクトルのコサイン類似度が閾値以上である場合、操作者は変更されておらず、同一人であると判定するものであってもよい。
操作者が変更したと判定した場合(S205:YES)、情報処理装置2の制御部20は、基準データを生成する(S206)。操作者が変更したと判定した場合、すなわち今回、車両Cが停車した際、操作者が交替した場合、情報処理装置2の制御部20は、実施形態1のS105と同様に基準データを生成する。
操作者が変更していない判定した場合(S205:NO)、情報処理装置2の制御部20は、基準データを生成せず、現在適用されている基準データを継続して使用する(S2031)。操作者が変更していない判定した場合、すなわち今回、車両Cが停車した際、操作者が交替しておらず同一人である場合、情報処理装置2の制御部20は、実施形態1のS1031と同様に、基準データを生成せず、現在適用されている基準データを継続して使用する。
情報処理装置2の制御部20は、実施形態1の処理S107からS108と同様にS208からS209までの処理を行う。
本実施形態によれば、操作者が乗車し車両Cが起動(イグニッションスイッチがオン)した後、停止(イグニッションスイッチがオフ)するまでの間、情報処理装置2は、所定の走行状態(停車後、最初の直進状態が所定期間維持)にあると判定する度に、操作者を画像に基づく顔認証を実行することなく、操作者の基準データを生成することを前提としている。この際、車両Cが起動中において、操作者が変更(運転手が交代)される場合が想定され、すなわち車両Cが停車しておりアイドリング状態において、操作者(運転手)が交代された場合、情報処理装置2は、当該交代した操作者を含む画像に基づき、基準データを生成及び更新して、監視処理を継続する。これに対し、情報処理装置2は、所定の走行状態にあると判定した場合であっても、走行状態にあると前回判定した時点と、今回判定した時点とにおいて、操作者が変更されていない場合、すなわち操作者が同一人である場合、今回の判定において取得した操作者を含む画像に基づく基準データを生成することなく、現在適用されている基準データを継続して用いる。すなわち、情報処理装置2は、所定の走行状態にあると前回判定した時点と、今回判定した時点とにおいて、操作者が変更された場合のみ、今回の判定において取得した操作者を含む画像に基づき、基準データを生成する。情報処理装置2は、操作者の変更の有無を判定するにあたり、前回の走行状態にあると判定した際に用いた画像と、今回の走行状態にあると判定した際に用いた画像とを対比し、これら画像に含まれる操作者の一致度又は類似度が所定値以上であれば、操作者の変更は無い、すなわち同一人であると判定するものであってもよい。このように情報処理装置2は、車両Cの起動(イグニッションスイッチ:オン)から停止(イグニッションスイッチ:オフ)までの間にて、所定の走行状態(停車後、最初の直進状態が所定期間維持)となる場合が複数回発生したとしても、操作者が変更されていない場合は、基準データを生成せず、現在適用されている基準データを継続して用いるため、当該基準データを生成する処理回数を低減でき、演算負荷を軽減することができる。
今回開示された実施形態は全ての点で例示であって、制限的なものではないと考えられるべきである。本開示の範囲は、上記した意味ではなく、請求の範囲によって示され、請求の範囲と均等の意味及び範囲内での全ての変更が含まれることが意図される。
請求の範囲に記載されている複数の請求項に関して、引用形式に関わらず、相互に組み合わせることが可能である。請求の範囲では、複数の請求項に従属する多項従属請求項を記載してもよい。多項従属請求項に従属する多項従属請求項を記載してもよい。多項従属請求項に従属する多項従属請求項が記載されていない場合であっても、これは、多項従属請求項に従属する多項従属請求項の記載を制限するものではない。
C 車両
1 カメラユニット(キャビンモニタ)
11 カメラ
2 情報処理装置(カメラECU)
20 制御部
21 入出力I/F
22 通信部
23 記憶部
P 制御プログラム(プログラム製品)
201 顔検出モデル
202 顔ランドマークモデル
203 姿勢検出モデル
3 ECU
4 車載ネットワーク
1 カメラユニット(キャビンモニタ)
11 カメラ
2 情報処理装置(カメラECU)
20 制御部
21 入出力I/F
22 通信部
23 記憶部
P 制御プログラム(プログラム製品)
201 顔検出モデル
202 顔ランドマークモデル
203 姿勢検出モデル
3 ECU
4 車載ネットワーク
Claims (12)
- 車両に搭載されたカメラによって前記車両を操作する操作者を含む画像を取得し、
前記車両が予め定められた所定の走行状態にあるか否かを判定し、
前記所定の走行状態にあると判定した場合、取得した前記画像に基づき、前記操作者に対する監視処理を行う際の基準データを生成する
処理をコンピュータに実行させる情報処理方法。 - 前記車両の車速が所定速度以上であり、前記車両の舵角が所定角度範囲以内となる状態が、所定期間以上継続した場合、前記所定の走行状態にあると判定する
請求項1に記載の情報処理方法。 - 前記基準データは、前記所定期間内の前記操作者の顔の位置、顔の向き、及び身体の向きの内、少なくとも1つに基づき得られるデータである
請求項2に記載の情報処理方法。 - 前記基準データは、前記所定期間内の前記操作者の顔の位置、顔の向き、及び身体の向きの内、少なくとも1つの平均値である
請求項3に記載の情報処理方法。 - 前記車両が停車した以降、最初に走行した際、前記所定の走行状態にあるか否かを判定する
請求項1から4のいずれか一項に記載の情報処理方法。 - 前記基準データは、前記所定期間内の前記操作者の目の開度に関するデータに基づき得られるデータである
請求項2に記載の情報処理方法。 - 前記操作者の目の開度に関するデータは、前記操作者の目における縦方向の最大長さを、横方向の最大長さで除算した際の最大値である
請求項6に記載の情報処理方法。 - 前記基準データを生成した以降、前記操作者を含む画像を取得する処理を継続し、
前記基準データの生成後に取得した画像と、前記基準データとに基づき、前記操作者に対する前記監視処理を実行する
請求項1から7のいずれか一項に記載の情報処理方法。 - 前記操作者の乗車後に前記所定の走行状態にあると判定する度に、前記操作者を前記画像に基づく顔認証を実行することなく、前記操作者の前記基準データを生成する
請求項1から8のいずれか一項に記載の情報処理方法。 - 前記所定の走行状態にあると判定した場合、現在適用されている前記基準データの生成時に用いた操作者を含む画像と、今回取得した操作者を含む画像とを対比することにより、前記操作者が変更されたか否かを判定し、
前記操作者が変更されたと判定した場合、今回取得した操作者を含む画像に基づき、前記操作者に対する前記監視処理を行う際の基準データを生成し、
前記操作者が変更されていないと判定した場合、今回取得した操作者を含む画像に基づく基準データを生成することなく、現在適用されている基準データを継続して用いる
請求項1から9のいずれか一項に記載の情報処理方法。 - 車両に搭載されたカメラによって前記車両を操作する操作者を含む画像を取得し、
前記車両が予め定められた所定の走行状態にあるか否かを判定し、
前記所定の走行状態にあると判定した場合、取得した前記画像に基づき、前記操作者に対する監視処理を行う際の基準データを生成する
処理をコンピュータに実行させるプログラム。 - 車両に搭載され、制御部を備える情報処理装置であって、
前記制御部は、
前記車両に搭載されたカメラによって前記車両を操作する操作者を含む画像を取得し、
前記車両が予め定められた所定の走行状態にあるか否かを判定し、
前記所定の走行状態にあると判定した場合、取得した前記画像に基づき、前記操作者に対する監視処理を行う際の基準データを生成する
情報処理装置。
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463668305P | 2024-07-08 | 2024-07-08 | |
| US63/668,305 | 2024-07-08 | ||
| JP2024210723A JP2026009802A (ja) | 2024-07-08 | 2024-12-03 | 情報処理方法、プログラム、及び情報処理装置 |
| JP2024-210723 | 2024-12-03 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026014240A1 true WO2026014240A1 (ja) | 2026-01-15 |
Family
ID=98386583
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2025/022839 Pending WO2026014240A1 (ja) | 2024-07-08 | 2025-06-25 | 情報処理方法、プログラム、及び情報処理装置 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026014240A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010029537A (ja) * | 2008-07-30 | 2010-02-12 | Toyota Motor Corp | 覚醒度判定装置 |
| JP2016115117A (ja) * | 2014-12-15 | 2016-06-23 | アイシン精機株式会社 | 判定装置および判定方法 |
| JP2017126151A (ja) * | 2016-01-13 | 2017-07-20 | アルプス電気株式会社 | 視線検出装置および視線検出方法 |
| JP2019088522A (ja) * | 2017-11-15 | 2019-06-13 | オムロン株式会社 | 情報処理装置、運転者モニタリングシステム、情報処理方法、及び情報処理プログラム |
| JP2019091255A (ja) * | 2017-11-15 | 2019-06-13 | オムロン株式会社 | 情報処理装置、運転者モニタリングシステム、情報処理方法、及び情報処理プログラム |
-
2025
- 2025-06-25 WO PCT/JP2025/022839 patent/WO2026014240A1/ja active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010029537A (ja) * | 2008-07-30 | 2010-02-12 | Toyota Motor Corp | 覚醒度判定装置 |
| JP2016115117A (ja) * | 2014-12-15 | 2016-06-23 | アイシン精機株式会社 | 判定装置および判定方法 |
| JP2017126151A (ja) * | 2016-01-13 | 2017-07-20 | アルプス電気株式会社 | 視線検出装置および視線検出方法 |
| JP2019088522A (ja) * | 2017-11-15 | 2019-06-13 | オムロン株式会社 | 情報処理装置、運転者モニタリングシステム、情報処理方法、及び情報処理プログラム |
| JP2019091255A (ja) * | 2017-11-15 | 2019-06-13 | オムロン株式会社 | 情報処理装置、運転者モニタリングシステム、情報処理方法、及び情報処理プログラム |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Chan et al. | A comprehensive review of driver behavior analysis utilizing smartphones | |
| CN109902562B (zh) | 一种基于强化学习的驾驶员异常姿态监测方法 | |
| CN110826370B (zh) | 车内人员的身份识别方法、装置、车辆及存储介质 | |
| CN107704805B (zh) | 疲劳驾驶检测方法、行车记录仪及存储装置 | |
| Rezaei et al. | Look at the driver, look at the road: No distraction! no accident! | |
| US11963759B2 (en) | State determination device, state determination method, and recording medium | |
| JP4420081B2 (ja) | 行動推定装置 | |
| EP3113073A1 (en) | Determination device, determination method, and non-transitory storage medium | |
| CN109977771A (zh) | 司机身份的验证方法、装置、设备及计算机可读存储介质 | |
| WO2017000218A1 (zh) | 活体检测方法及设备、计算机程序产品 | |
| JP2016115117A (ja) | 判定装置および判定方法 | |
| CN116189153A (zh) | 驾驶员的视线识别方法及装置、车辆和存储介质 | |
| CN118201834A (zh) | 设备端的驾驶风险行为干预方法、装置、设备及存储介质 | |
| TW202125441A (zh) | 安全警示語音提示方法 | |
| Khan et al. | Real time eyes tracking and classification for driver fatigue detection | |
| JP2016115120A (ja) | 開閉眼判定装置および開閉眼判定方法 | |
| Bisogni et al. | Iot-enabled biometric security: enhancing smart car safety with depth-based head pose estimation | |
| WO2026014240A1 (ja) | 情報処理方法、プログラム、及び情報処理装置 | |
| CN113420656A (zh) | 一种疲劳驾驶检测方法、装置、电子设备及存储介质 | |
| JP2026009802A (ja) | 情報処理方法、プログラム、及び情報処理装置 | |
| CN119502943B (zh) | 基于复杂情景的驾驶行为规范方法、系统、介质及设备 | |
| CN120942356A (zh) | 一种应用于商用车的疲劳驾驶预警处理系统及方法 | |
| WO2021024905A1 (ja) | 画像処理装置、モニタリング装置、制御システム、画像処理方法、コンピュータプログラム、及び記憶媒体 | |
| CN120708198A (zh) | 一种基于生物特征识别的分级式驾驶疲劳干预方法 | |
| CN116704482B (zh) | 疲劳驾驶的检测方法、检测系统、检测装置及计算机可读存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25836802 Country of ref document: EP Kind code of ref document: A1 |