EP4602568A1 - System and method for estimation of eye gaze direction of a user with or without eyeglasses - Google Patents

System and method for estimation of eye gaze direction of a user with or without eyeglasses

Info

Publication number
EP4602568A1
EP4602568A1 EP23838037.2A EP23838037A EP4602568A1 EP 4602568 A1 EP4602568 A1 EP 4602568A1 EP 23838037 A EP23838037 A EP 23838037A EP 4602568 A1 EP4602568 A1 EP 4602568A1
Authority
EP
European Patent Office
Prior art keywords
user
eye
key points
processing device
gaze direction
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23838037.2A
Other languages
German (de)
French (fr)
Inventor
Sourav Lakhotia
Shuaib AHMED
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Mercedes Benz Group AG
Original Assignee
Mercedes Benz Group AG
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Mercedes Benz Group AG filed Critical Mercedes Benz Group AG
Publication of EP4602568A1 publication Critical patent/EP4602568A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/22Image preprocessing by selection of a specific region containing or referencing a pattern; Locating or processing of specific regions to guide the detection or recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/762Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks
    • G06V10/7625Hierarchical techniques, i.e. dividing or merging patterns to obtain a tree-like representation; Dendograms
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/766Arrangements for image or video recognition or understanding using pattern recognition or machine learning using regression, e.g. by projecting features on hyperplanes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/59Context or environment of the image inside of a vehicle, e.g. relating to seat occupancy, driver state or inner lighting conditions
    • G06V20/597Recognising the driver's state or behaviour, e.g. attention or drowsiness
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B60VEHICLES IN GENERAL
    • B60YINDEXING SCHEME RELATING TO ASPECTS CROSS-CUTTING VEHICLE TECHNOLOGY
    • B60Y2302/00Responses or measures related to driver conditions
    • B60Y2302/03Actuating a signal or alarm device

Definitions

  • a driver of vehicles generally gets distracted while driving by engaging in other tasks such as looking at mobiles, taking care of a child or pet inside the vehicle, or manually handling the infotainment system of the vehicle, which may result in road accidents. It would, therefore, be advantageous from the point of view of safety if a simple, automated, and efficient operable solution to monitor where the driver of the vehicle is looking, could be provided, which should also remind or alert the user or switch the vehicle to an auto-mode and/or stop the vehicle for safety, if the vision or gaze direction of the driver gets off the road beyond a defined time period.
  • Patent document number EP3789848A1 discloses a method for estimating a gaze direction of a user, the method comprising the steps of: acquiring an image of a face of the user, determining an approximate gaze direction based on a current head pose and a relationship between head pose and gaze direction, determining an estimated gaze direction based on detected eye features, determining a precise gaze direction based on glint/reflction position and eye features, and combining the approximate gaze direction and at least one of the estimated gaze direction and the precise gaze direction to provide a corrected gaze direction.
  • the above-cited reference estimates the current gaze direction of the user based on reflected landmarks in the cornea of an eye, along with the head pose of the driver, which may require more computational power.
  • the driver may be wearing eyeglasses or sunglasses while driving the vehicle, so their eye may be occluded.
  • the eye gaze detection technique of the above-cited reference may not be able to properly estimate the eye gaze direction of the drier, as the key landmarks of the eye can not be easily inferred due to the eyeglasses or occlusion.
  • the eyeglasses on the eye may also reflect the surrounding environment and glares, which may again degrade the eye gaze detection process of the above-cited reference.
  • a general object of the present disclosure is to monitor the eye gaze direction of drivers with or without wearing eyeglasses.
  • Another object of the present disclosure is to accurately identify the direction in which the driver wearing eyeglasses is looking while driving a vehicle to avoid distractions and help prevent road accidents.
  • An object of the present disclosure is to provide an efficient, and reliable system and a method for estimation of the eye gaze direction of drivers with or without wearing eyeglasses, which reminds or alerts the driver or switch the vehicle to an auto-mode and stop the vehicle if the gaze direction gets off the road for few seconds.
  • aspects of the present disclosure relate to the field of eye-gaze estimation.
  • the present disclosure provides a system and a method for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle.
  • An aspect of the present disclosure pertains to a system for the estimation of an eye gaze direction of a user with or without wearing eyeglasses.
  • the system comprises a processing device incorporating a model with a first branch configured with a convolutional neural network (CNN) architecture.
  • the CNN architecture may comprise case-based reasoning (CBR) blocks, and wherein heat maps may be fed to the CNN architecture to detect the one or more key points.
  • the model is pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module.
  • the processing device comprises a processor coupled with a memory, wherein the memory stores one or more instructions executable by the processor to: detect, using the CNN architecture, one or more key points from one or more images associated with the one or more eye of the user of a vehicle in real-time; and determine, using the learning module, the eye gaze direction of the user based on the detected key points.
  • the system may comprise an image acquisition unit comprising a camera to capture one or more images of a face of the user.
  • the processing device may be configured to crop one or more eye region from the captured one or more images of the face of the user, detect, using the CNN architecture, the one or more key points from the cropped one or more images, and regress, using the learning module, the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
  • the pre-trained model of the processing device may comprise a second branch including a classification network.
  • the second branch may be fed with an output of at least one of the CBR blocks, wherein the second branch may e trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data.
  • the processing device may be configured to train the learning module, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which may enable the learning module to determine the gaze direction of the user.
  • the one or more key points may be associated with a sclera, an iris, and an eyeball center.
  • the learning module may be a random forest model.
  • the random forest model may be fed with four vectors derived from any one or a combination of the iris center and extreme joints of the sclera, and angular distance between the extreme sclera point and the iris center.
  • the random forest model may be trained with one or more noise, keypoints removal, and one or more additional features to identify the yaw and gaze.
  • the processing device may be configured to generate a set of alert signals when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, and stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI for the predefined time, wherein the ROI is a road where the vehicle is traveling.
  • ROI region of interest
  • the processing device may be configured to determine the gaze direction of the user using the one or more key points associated with another non-occluded eye of the user.
  • the method may comprise the steps of: capturing, by a camera, one or more images of a face of the user; cropping, by the processing device, one or more eye region from the captured one or more images of the face of the user; detecting, using the CNN architecture, the one or more key points from the cropped one or more images; regressing, using the learning module, the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
  • FIG. 1 illustrates an exemplary block diagram of the proposed system for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with embodiments of the present invention.
  • FIG. 2 illustrates an exemplary block diagram representing functional units of a processing device associated with the proposed system, in accordance with embodiments of the present invention.
  • FIG. 3A illustrates a flow diagram representing steps of the proposed method for for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with one or more embodiments of the present disclosure.
  • FIG. 3B illustrates an exemplary flow chart depicting the operation of the proposed system for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with one or more embodiments of the present disclosure.
  • FIG. 5 illustrates an exemplary architecture of the random forest model implemented in the proposed system, in accordance with one or more embodiments of the present disclosure.
  • FIG. 6 illustrates an exemplary representation of the key point architecture implemented in the proposed system, in accordance with one or more embodiments of the present disclosure.
  • Embodiments explained herein relate to the field of eye-gaze estimation.
  • the present disclosure provides a system and a method for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle.
  • the proposed system 100 for estimation of an eye gaze direction of a user (also referred to as driver, passenger, or occupant, herein) wearing eyeglasses is disclosed.
  • the disclosed system 100 accurately identifies the gaze direction (the direction in which the user/driver wearing eyeglasses is looking) while driving a vehicle to avoid distractions and help prevent road accidents, which alerts the user/driver or switch the vehicle to an auto-mode and stop the vehicle for the safety of the driver if the gaze direction of the user/driver gets off the road for a few seconds.
  • the system 100 can include a processing device 102, which can be in communication with a vehicle control unit (VCU) 106 of a vehicle.
  • VCU vehicle control unit
  • the processing device 102 can be a server that can remain in one or more vehicle unit associated with one or more vehicles.
  • the system 100 can further include an image acquisition unit 104 comprising one or more image sensors or cameras (collectively referred to as cameras or image sensors, herein) installed within the vehicle to monitor one or more images or videos of users (also referred to as driver or passenger or occupant, herein) sitting/traveling in the vehicle.
  • the image acquisition unit 104 can also capture images or videos of a road or outside view of the vehicle.
  • the image acquisition unit 104 can be in communication with the processing device 102 and/or the VCU 106 of the vehicle.
  • the image acquisition unit 104 comprising the camera(s) or image sensor(s) 104 can be installed in the vehicle interior to capture or monitor images or video of the face of the user/driver in a real-time.
  • One of the cameras 104 can be facing toward the front seats of the vehicle where the user/driver may be sitting.
  • Other cameras 104 can be facing toward the road or outside of the vehicle.
  • the camera or image sensors 104 may be positioned on the dashboard or ceiling of the vehicle for providing full coverage of the interiors of the vehicle to capture the images of the face of the user/driver and also capture the outside view or road.
  • the cameras 104 may be an Infra-Red (IR) camera such as a nearinfrared camera, a mid-wave infrared camera, and long-wave infrared camera, and so forth.
  • IR Infra-Red
  • the processing device 102 can be in communication with or operatively coupled to the image acquisition unit 104, and the VCU 106 of the vehicle, through a network.
  • the network can be a wireless network, a wired network or a combination thereof that can be implemented as one of the different types of networks, such as Intranet, Local Area Network (LAN), Wide Area Network (WAN), Internet, and the like.
  • the network can either be a dedicated network or a shared network.
  • the shared network can represent an association of different types of networks that can use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol/lnternet Protocol (TCP/IP), Wireless Application Protocol (WAP), and the like.
  • HTTP Hypertext Transfer Protocol
  • TCP/IP Transmission Control Protocol/lnternet Protocol
  • WAP Wireless Application Protocol
  • the processing device 102 can be implemented using any or a combination of hardware components and software components such as a cloud, a server, a computing system, a computing device, a network device, and the like. Further, the processing device 102 can interact with the image acquisition unit 104, and the VCU 106 through the wired or wireless network.
  • the image acquisition unit 104 can be configured to capture the images/videos of the face of the user/driver.
  • the processing device 102 can be configured with a model with a first branch configured with a convolutional neural network (CNN) architecture, where the model can be pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module.
  • CNN convolutional neural network
  • the one or more key points can be associated with a sclera, an iris, and an eyeball center of the user/driver.
  • the processing device 102 can be configured to detect, using the pre-trained CNN architecture, one or more key points associated with one or more eye of the user in real-time from the images captured by the image acquisition unit 104. Further, the processing device 102 can be further configured with a learning module such as a random forest model, but not limited to the like, that can enable the processing device 102 to determine the eye gaze direction of the user based on the detected key points. The learning module enables the processing device 102 to regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user/driver with or without eyeglasses. Accordingly, the processing device 102 can accurately identify the direction in which the user/driver wearing eyeglasses is looking while driving a vehicle, which can help avoid distractions and help prevent road accidents.
  • a learning module such as a random forest model, but not limited to the like
  • the CNN architecture can further include case-based reasoning (CBR) blocks. Further, heat maps can be fed to the CNN architecture to detect the one or more key points.
  • the pre-trained model of the processing device 102 can also include a second branch including a classification network, where the second branch can be fed with an output of at least one of the CBR blocks. The second branch can be trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data.
  • the random forest model or learning module can be fed with four vectors derived from the iris center and extreme joints of the sclera, and/or angular distance between the extreme sclera point and the iris center. Further, the random forest model can be trained with one or more noise, key-points removal, and one or more additional features to identify the yaw and gaze for better accuracy and reliability.
  • the processing device 102 can be configured to pre-train or train the learning module (random forest model) in real-time, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which enables the learning module to determine the gaze direction of the user.
  • the processing device 102 can receive the first set of images associated with the eye regions and the corresponding one or more key points present in the first set of images. Further, the processing device 102 can render the eyeglasses at a predefined transparency over at least one of the eye regions in a predefined percentage of the first set of images to provide the second set of images.
  • the processing device 102 can train and test the CNN and the random forest model with the synthetic data comprising the first set of images, the second set of images, and the corresponding one or more key points, which enables the learning module to accurately detect the one or more key points in the captured images of the user in the real-time.
  • the processing device 102 can be configured to determine the gaze direction of the user using the key points associated with another non-occluded eye of the user.
  • the processing device 102 can be configured to generate a set of alert signals using the VCU 106 of the vehicle when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, wherein the ROI is a road where the vehicle is traveling.
  • the processing device 102 can be configured to actuate the VCU 106 to stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI (road) for the predefined time.
  • the memory 204 can store one or more computer- readable instructions or routines, which may be fetched and executed to create or share the data units over a network service.
  • the memory 204 can include any non-transitory storage device including, for example, volatile memory such as RAM, or non-volatile memory such as EPROM, flash memory, and the like.
  • the processing device 102 can also include an interface(s) 206.
  • the interface(s) 206 may include a variety of interfaces, for example, interfaces for data input and output devices, referred to as I/O devices, storage devices, and the like.
  • the interface(s) 206 may facilitate communication of the processing device 102 with various devices coupled to the server 102, such as the infotainment system, the image acquisition unit 104, the VCU 106, and the power source of the vehicle.
  • the interface(s) 206 may also provide a communication pathway for one or more components of the processing device 102. Examples of such components include, but are not limited to, processing engine(s) 208 and database 210.
  • the processing engine(s) 208 can be implemented as a combination of hardware and programming (for example, programmable instructions) to implement one or more functionalities of the processing engine(s) 208.
  • programming for the processing engine(s) 208 may be processor executable instructions stored on a non-transitory machine-readable storage medium and the hardware for the processing engine(s) 208 may include a processing resource (for example, one or more processors), to execute such instructions.
  • the machine-readable storage medium may store instructions that, when executed by the processing resource, implement the processing engine(s) 208.
  • the processing device 102 can include the machine-readable storage medium storing the instructions and the processing resource to execute the instructions, or the machine-readable storage medium may be separate but accessible to the processing device 102 and the processing resource.
  • the processing engine(s) 208 may be implemented by electronic circuitry.
  • the database 210 can include data that is either stored or generated as a result of functionalities implemented by any of the components of the processing engine(s) 208.
  • the processing engine(s) 208 can include a pre-trained model 212 with a CNN architecture and a classification network, a learning module 214, an actuation and control unit 216, an alert unit 218, and other unit(s) 220.
  • the other unit(s) 220 can implement functionalities that supplement applications or functions performed by the processing device 102 or the processing engine(s) 208.
  • the processing device 102 can enable the image acquisition unit 104 to capture the images/videos of a face of the user/driver wearing eyeglasses as well as the images of the road or outside of the vehicle.
  • the pre-trained model 212 as shown in detail in FIG. 5, can cause the processing device 102 to crop one or more eye regions from the captured images of the face of the user/driver occluded with eyeglass and correspondingly detect one or more key points from the captured images.
  • the one or more key points can be associated with a sclera, an iris, and an eyeball center of the user/driver as shown in detail in FIG. 4.
  • the processing device 102 can be further configured with a learning module 214, such as a random forest model as shown in FIG. 6, which can cause the processing device 102 to to determine the eye gaze direction of the user based on the detected key points.
  • the learning module 214 enables the processing device 102 to regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user/driver with or without eyeglasses.
  • the pre-trained model 212 and learning module 214 can enable the processing device 102 to accurately identify the direction in which the user/driver wearing eyeglasses is looking while driving a vehicle, which can help avoid distractions and help prevent road accidents.
  • the pre-trained model 212 can include a first branch configured with a convolutional neural network (CNN) architecture and a second branch including a classification network.
  • the model is pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass.
  • the CNN architecture can include case-based reasoning (CBR) blocks.
  • heat maps can be fed to the CNN architecture to detect the one or more key points.
  • the second branch can be fed with an output of at least one of the CBR blocks, where the second branch can be trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data.
  • the random forest model 214 can be fed with four vectors derived from the iris center and extreme joints of the sclera, and/or angular distance between the extreme sclera point and the iris center. Further, the random forest model 214 can be trained with one or more noise, key-points removal, and one or more additional features to identify the yaw and gaze for better accuracy and reliability.
  • the processor 202 can cause the processing device 102 to pre-train or train the learning module 214 (random forest model) in real-time, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which can enable the learning module 214 to determine the gaze direction of the user.
  • the processing device 102 can receive the first set of images associated with the eye regions and the corresponding one or more key points present in the first set of images.
  • the processing device 102 can render the eyeglasses at a predefined transparency over at least one of the eye regions in a predefined percentage of the first set of images to provide the second set of images as shown in FIG. 3B. Accordingly, the processing device 102 can train and test the model 212 and the learning module (random forest model) 214 with the synthetic data comprising the first set of images, the second set of images, and the corresponding one or more key points, which can enable the learning module 214 to accurately detect the one or more key points in the captured images of the user in realtime.
  • the model 212 and the learning module (random forest model) 214 with the synthetic data comprising the first set of images, the second set of images, and the corresponding one or more key points, which can enable the learning module 214 to accurately detect the one or more key points in the captured images of the user in realtime.
  • the actuation and control unit 216 can cause the processing device 102 to generate and transmit a set of alert signals to the VCU 106 of the vehicle to alert the user/driver when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, wherein the ROI is a road where the vehicle is traveling.
  • the actuation and control unit 216 can cause the processing device 102 to actuate the VCU 106 to stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI (road) for the predefined time.
  • the proposed method 300 for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle involves the image acquisition unit, and the processing device connected with the VCU of the vehicle.
  • the method 300 includes step 302 of capturing, by the image acquisition unit, one or more images of a face of the user, followed by step 304 of cropping, by the processing device, one or more eye region from the captured one or more images of the face of the user.
  • the method 300 further includes step 306 of detecting, by a processing device configured with a CNN architecture, one or more key points from one or more images associated with one or more eye of a user of a vehicle in real-time.
  • the CNN architecture can be pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass.
  • the method 300 includes step 308 of determining, using a learning module associated with the processing device, the eye gaze direction of the user based on the detected key points.
  • the processing device can regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
  • the image acquisition unit can capture the images of the user/driver sitting in the vehicle and a face detector module can detect face in the captured images.
  • the processing device can then crop the face and further crop eye regions from the captured images.
  • the CNN model can then detect key-points in the cropped eye images. In case, no key points are detected, no gaze condition is identified by the processing device. Further, the random forest model can then select a non-occluded eye and estimate the gaze direction of the user/driver.
  • the present invention (system and method) accurately estimates the eye gaze direction of a user/driver with or without wearing eyeglasses while driving the vehicle.
  • the present invention alerts the driver or switches the vehicle to an auto-mode and stops the vehicle for the safety of the driver if the gaze direction of the driver gets off the road for a few seconds, thereby avoiding distractions and preventing road accidents.
  • the present disclosure monitors the eye gaze direction of drivers with or without wearing eyeglasses.
  • the present disclosure provides an efficient, reliable, and faster system and method for the estimation of the eye gaze direction of drivers wearing eyeglasses or when the eye of the driver is occluded.
  • the present disclosure accurately identifies the direction in which the driver wearing eyeglasses is looking while driving a vehicle to avoid distractions and help prevent road accidents.
  • the present disclosure alerts the driver or switches the vehicle to an auto-mode and stops the vehicle for the safety of the driver if the gaze direction of the driver gets off the road for a few seconds.
  • the present disclosure provides an efficient, and reliable system and a method for estimation of the eye gaze direction of drivers with or without wearing eyeglasses, which reminds or alerts the driver or switches the vehicle to an auto-mode and stops the vehicle if the gaze direction of the driver gets off the road for few seconds.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computing Systems (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Human Computer Interaction (AREA)
  • General Engineering & Computer Science (AREA)
  • Ophthalmology & Optometry (AREA)
  • Image Analysis (AREA)
  • Traffic Control Systems (AREA)

Abstract

A system (100) for the estimation of an eye gaze direction of a user with or without wearing eyeglasses comprises a processing device (102) incorporating a model (212) with a first branch configured with a convolutional neural network (CNN) architecture, the model (212) being pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module (214). The processing device (102) is configured to detect, using the CNN architecture, one or more key points from one or more images associated with the one or more eye of the user of a vehicle in real-time, and determine, using the learning module, the eye gaze direction of the user based on the detected key points.

Description

SYSTEM AND METHOD FOR ESTIMATION OF EYE GAZE DIRECTION OF A USER WITH OR WITHOUT EYEGLASSES
TECHNICAL FIELD
[0001] The present disclosure relates to the field of eye-gaze estimation. In particular, the present disclosure provides a system and a method for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle.
BACKGROUND
[0002] A driver of vehicles generally gets distracted while driving by engaging in other tasks such as looking at mobiles, taking care of a child or pet inside the vehicle, or manually handling the infotainment system of the vehicle, which may result in road accidents. It would, therefore, be advantageous from the point of view of safety if a simple, automated, and efficient operable solution to monitor where the driver of the vehicle is looking, could be provided, which should also remind or alert the user or switch the vehicle to an auto-mode and/or stop the vehicle for safety, if the vision or gaze direction of the driver gets off the road beyond a defined time period.
[0003] Patent document number EP3789848A1 discloses a method for estimating a gaze direction of a user, the method comprising the steps of: acquiring an image of a face of the user, determining an approximate gaze direction based on a current head pose and a relationship between head pose and gaze direction, determining an estimated gaze direction based on detected eye features, determining a precise gaze direction based on glint/reflction position and eye features, and combining the approximate gaze direction and at least one of the estimated gaze direction and the precise gaze direction to provide a corrected gaze direction.
[0004] The above-cited reference estimates the current gaze direction of the user based on reflected landmarks in the cornea of an eye, along with the head pose of the driver, which may require more computational power. Moreover, the driver may be wearing eyeglasses or sunglasses while driving the vehicle, so their eye may be occluded. Thus, the eye gaze detection technique of the above-cited reference may not be able to properly estimate the eye gaze direction of the drier, as the key landmarks of the eye can not be easily inferred due to the eyeglasses or occlusion. Moreover, the eyeglasses on the eye may also reflect the surrounding environment and glares, which may again degrade the eye gaze detection process of the above-cited reference. [0005] There is, therefore, a need in the art to overcome the above-mentioned drawbacks, limitations, and shortcomings associated with the conventional eye gaze detection techniques and the above-cited reference, by accurately estimating the eye gaze direction of a user/driver with or without wearing eyeglasses while driving the vehicle.
OBJECTS OF THE PRESENT DISCLOSURE
[0006] A general object of the present disclosure is to monitor the eye gaze direction of drivers with or without wearing eyeglasses.
[0007] An object of the present disclosure is to provide an efficient, reliable, and faster system and method for estimation of the eye gaze direction of drivers wearing eyeglasses or when the eye of the driver is occluded.
[0008] Another object of the present disclosure is to accurately identify the direction in which the driver wearing eyeglasses is looking while driving a vehicle to avoid distractions and help prevent road accidents.
[0009] Yet another object of the present disclosure is to alert the driver or switch the vehicle to an auto-mode and stop the vehicle for the safety of the driver if the gaze direction of the driver gets off the road for a few seconds.
[0010] An object of the present disclosure is to provide an efficient, and reliable system and a method for estimation of the eye gaze direction of drivers with or without wearing eyeglasses, which reminds or alerts the driver or switch the vehicle to an auto-mode and stop the vehicle if the gaze direction gets off the road for few seconds.
SUMMARY
[0011] Aspects of the present disclosure relate to the field of eye-gaze estimation. In particular, the present disclosure provides a system and a method for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle.
[0012] An aspect of the present disclosure pertains to a system for the estimation of an eye gaze direction of a user with or without wearing eyeglasses. The system comprises a processing device incorporating a model with a first branch configured with a convolutional neural network (CNN) architecture. The CNN architecture may comprise case-based reasoning (CBR) blocks, and wherein heat maps may be fed to the CNN architecture to detect the one or more key points. The model is pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module. The processing device comprises a processor coupled with a memory, wherein the memory stores one or more instructions executable by the processor to: detect, using the CNN architecture, one or more key points from one or more images associated with the one or more eye of the user of a vehicle in real-time; and determine, using the learning module, the eye gaze direction of the user based on the detected key points.
[0013] In an aspect, the system may comprise an image acquisition unit comprising a camera to capture one or more images of a face of the user. The processing device may be configured to crop one or more eye region from the captured one or more images of the face of the user, detect, using the CNN architecture, the one or more key points from the cropped one or more images, and regress, using the learning module, the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
[0014] In an aspect, the pre-trained model of the processing device may comprise a second branch including a classification network. The second branch may be fed with an output of at least one of the CBR blocks, wherein the second branch may e trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data.
[0015] In an aspect, the processing device may be configured to train the learning module, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which may enable the learning module to determine the gaze direction of the user. The one or more key points may be associated with a sclera, an iris, and an eyeball center.
[0016] In an aspect, the learning module may be a random forest model. The random forest model may be fed with four vectors derived from any one or a combination of the iris center and extreme joints of the sclera, and angular distance between the extreme sclera point and the iris center. The random forest model may be trained with one or more noise, keypoints removal, and one or more additional features to identify the yaw and gaze.
[0017] In an aspect, the processing device may be configured to generate a set of alert signals when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, and stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI for the predefined time, wherein the ROI is a road where the vehicle is traveling.
[0018] In an aspect, when one of the eyes of the user is at least partially occluded, the processing device may be configured to determine the gaze direction of the user using the one or more key points associated with another non-occluded eye of the user.
[0019] Another aspect of the present disclosure pertains to a method for estimation of an eye gaze direction of a user. The method comprises the steps of: detecting, by a processing device configured with a CNN architecture, one or more key points from one or more images associated with one or more eye of a user of a vehicle in real-time, wherein the CNN architecture has been pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass; and determine, using a learning module associated with the processing device, the eye gaze direction of the user based on the detected key points.
[0020] In an aspect, the method may comprise the steps of: capturing, by a camera, one or more images of a face of the user; cropping, by the processing device, one or more eye region from the captured one or more images of the face of the user; detecting, using the CNN architecture, the one or more key points from the cropped one or more images; regressing, using the learning module, the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
[0021] Various objects, features, aspects and advantages of the inventive subject matter will become more apparent from the following detailed description of preferred embodiments, along with the accompanying drawing figures in which like numerals represent like components.
BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0023] FIG. 1 illustrates an exemplary block diagram of the proposed system for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with embodiments of the present invention.
[0024] FIG. 2 illustrates an exemplary block diagram representing functional units of a processing device associated with the proposed system, in accordance with embodiments of the present invention.
[0025] FIG. 3A illustrates a flow diagram representing steps of the proposed method for for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with one or more embodiments of the present disclosure.
[0026] FIG. 3B illustrates an exemplary flow chart depicting the operation of the proposed system for estimation of the eye gaze direction of a user wearing eyeglasses in a vehicle, in accordance with one or more embodiments of the present disclosure.
[0027] FIG. 3C illustrates an exemplary logic diagram of the proposed system to elaborate upon the working of the system, in accordance with one or more embodiments of the present disclosure. [0028] FIG. 4 illustrates exemplary key points associated with an eye of users, which help detect eye gaze direction of the users, in accordance with one or more embodiments of the present disclosure.
[0029] FIG. 5 illustrates an exemplary architecture of the random forest model implemented in the proposed system, in accordance with one or more embodiments of the present disclosure.
[0030] FIG. 6 illustrates an exemplary representation of the key point architecture implemented in the proposed system, in accordance with one or more embodiments of the present disclosure.
DETAILED DESCRIPTION
[0031] The following is a detailed description of embodiments of the disclosure depicted in the accompanying drawings. The embodiments are in such details as to clearly communicate the disclosure. However, the amount of detail offered is not intended to limit the anticipated variations of embodiments; on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosures as defined by the appended claims.
[0032] Embodiments explained herein relate to the field of eye-gaze estimation. In particular, the present disclosure provides a system and a method for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle.
[0033] Referring to FIG. 1 , the proposed system 100 for estimation of an eye gaze direction of a user (also referred to as driver, passenger, or occupant, herein) wearing eyeglasses is disclosed. In particular, the disclosed system 100 accurately identifies the gaze direction (the direction in which the user/driver wearing eyeglasses is looking) while driving a vehicle to avoid distractions and help prevent road accidents, which alerts the user/driver or switch the vehicle to an auto-mode and stop the vehicle for the safety of the driver if the gaze direction of the user/driver gets off the road for a few seconds.
[0034] In an embodiment, the system 100 can include a processing device 102, which can be in communication with a vehicle control unit (VCU) 106 of a vehicle. In another embodiment, the processing device 102 can be a server that can remain in one or more vehicle unit associated with one or more vehicles.
[0035] In another embodiment, the system 100 can further include an image acquisition unit 104 comprising one or more image sensors or cameras (collectively referred to as cameras or image sensors, herein) installed within the vehicle to monitor one or more images or videos of users (also referred to as driver or passenger or occupant, herein) sitting/traveling in the vehicle. The image acquisition unit 104 can also capture images or videos of a road or outside view of the vehicle. The image acquisition unit 104 can be in communication with the processing device 102 and/or the VCU 106 of the vehicle.
[0036] In an exemplary embodiment, the image acquisition unit 104 comprising the camera(s) or image sensor(s) 104 can be installed in the vehicle interior to capture or monitor images or video of the face of the user/driver in a real-time. One of the cameras 104 can be facing toward the front seats of the vehicle where the user/driver may be sitting. Other cameras 104 can be facing toward the road or outside of the vehicle. The camera or image sensors 104 may be positioned on the dashboard or ceiling of the vehicle for providing full coverage of the interiors of the vehicle to capture the images of the face of the user/driver and also capture the outside view or road. The cameras 104 may be an Infra-Red (IR) camera such as a nearinfrared camera, a mid-wave infrared camera, and long-wave infrared camera, and so forth.
[0037] In an embodiment, the processing device 102 can be in communication with or operatively coupled to the image acquisition unit 104, and the VCU 106 of the vehicle, through a network. Further, the network can be a wireless network, a wired network or a combination thereof that can be implemented as one of the different types of networks, such as Intranet, Local Area Network (LAN), Wide Area Network (WAN), Internet, and the like. Further, the network can either be a dedicated network or a shared network. The shared network can represent an association of different types of networks that can use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol/lnternet Protocol (TCP/IP), Wireless Application Protocol (WAP), and the like. Further, the
[0038] In an embodiment, the processing device 102 can be implemented using any or a combination of hardware components and software components such as a cloud, a server, a computing system, a computing device, a network device, and the like. Further, the processing device 102 can interact with the image acquisition unit 104, and the VCU 106 through the wired or wireless network.
[0039] The image acquisition unit 104 can be configured to capture the images/videos of the face of the user/driver. The processing device 102 can be configured with a model with a first branch configured with a convolutional neural network (CNN) architecture, where the model can be pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module. In an embodiment, the one or more key points can be associated with a sclera, an iris, and an eyeball center of the user/driver. The processing device 102 can be configured to detect, using the pre-trained CNN architecture, one or more key points associated with one or more eye of the user in real-time from the images captured by the image acquisition unit 104. Further, the processing device 102 can be further configured with a learning module such as a random forest model, but not limited to the like, that can enable the processing device 102 to determine the eye gaze direction of the user based on the detected key points. The learning module enables the processing device 102 to regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user/driver with or without eyeglasses. Accordingly, the processing device 102 can accurately identify the direction in which the user/driver wearing eyeglasses is looking while driving a vehicle, which can help avoid distractions and help prevent road accidents.
[0040] In an embodiment, the CNN architecture can further include case-based reasoning (CBR) blocks. Further, heat maps can be fed to the CNN architecture to detect the one or more key points. Furthermore, the pre-trained model of the processing device 102 can also include a second branch including a classification network, where the second branch can be fed with an output of at least one of the CBR blocks. The second branch can be trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data. In an embodiment, the random forest model or learning module can be fed with four vectors derived from the iris center and extreme joints of the sclera, and/or angular distance between the extreme sclera point and the iris center. Further, the random forest model can be trained with one or more noise, key-points removal, and one or more additional features to identify the yaw and gaze for better accuracy and reliability.
[0041] In another embodiment, the processing device 102 can be configured to pre-train or train the learning module (random forest model) in real-time, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which enables the learning module to determine the gaze direction of the user. During the training, the processing device 102 can receive the first set of images associated with the eye regions and the corresponding one or more key points present in the first set of images. Further, the processing device 102 can render the eyeglasses at a predefined transparency over at least one of the eye regions in a predefined percentage of the first set of images to provide the second set of images. Accordingly, the processing device 102 can train and test the CNN and the random forest model with the synthetic data comprising the first set of images, the second set of images, and the corresponding one or more key points, which enables the learning module to accurately detect the one or more key points in the captured images of the user in the real-time.
[0042] In an example, when one of the eyes of the user is detected to be at least partially occluded, the processing device 102 can be configured to determine the gaze direction of the user using the key points associated with another non-occluded eye of the user. [0043] In an embodiment, the processing device 102 can be configured to generate a set of alert signals using the VCU 106 of the vehicle when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, wherein the ROI is a road where the vehicle is traveling. In another embodiment, the processing device 102 can be configured to actuate the VCU 106 to stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI (road) for the predefined time.
[0044] Referring to FIG. 2, block diagram shown therein depicts exemplary functional units of the processing device 102 that can include one or more processor(s) 202, memory 204, interface(s) 206, processing engine(s) 208, and database 210. The one or more processor(s) 202 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing device 102s, logic circuitries, and/or any devices that manipulate data based on operational instructions. Among other capabilities, the one or more processor(s) 202 are configured to fetch and execute computer-readable instructions stored in a memory 204 of the server 102. The memory 204 can store one or more computer- readable instructions or routines, which may be fetched and executed to create or share the data units over a network service. The memory 204 can include any non-transitory storage device including, for example, volatile memory such as RAM, or non-volatile memory such as EPROM, flash memory, and the like.
[0045] In an embodiment, the processing device 102 can also include an interface(s) 206. The interface(s) 206 may include a variety of interfaces, for example, interfaces for data input and output devices, referred to as I/O devices, storage devices, and the like. The interface(s) 206 may facilitate communication of the processing device 102 with various devices coupled to the server 102, such as the infotainment system, the image acquisition unit 104, the VCU 106, and the power source of the vehicle. The interface(s) 206 may also provide a communication pathway for one or more components of the processing device 102. Examples of such components include, but are not limited to, processing engine(s) 208 and database 210.
[0046] In an embodiment, the processing engine(s) 208 can be implemented as a combination of hardware and programming (for example, programmable instructions) to implement one or more functionalities of the processing engine(s) 208. In examples described herein, such combinations of hardware and programming may be implemented in several different ways. For example, the programming for the processing engine(s) 208 may be processor executable instructions stored on a non-transitory machine-readable storage medium and the hardware for the processing engine(s) 208 may include a processing resource (for example, one or more processors), to execute such instructions. In the present examples, the machine-readable storage medium may store instructions that, when executed by the processing resource, implement the processing engine(s) 208. In such examples, the processing device 102 can include the machine-readable storage medium storing the instructions and the processing resource to execute the instructions, or the machine-readable storage medium may be separate but accessible to the processing device 102 and the processing resource. In other examples, the processing engine(s) 208 may be implemented by electronic circuitry. The database 210 can include data that is either stored or generated as a result of functionalities implemented by any of the components of the processing engine(s) 208.
[0047] In an embodiment, the processing engine(s) 208 can include a pre-trained model 212 with a CNN architecture and a classification network, a learning module 214, an actuation and control unit 216, an alert unit 218, and other unit(s) 220. The other unit(s) 220 can implement functionalities that supplement applications or functions performed by the processing device 102 or the processing engine(s) 208.
[0048] According to an embodiment, the processing device 102 can enable the image acquisition unit 104 to capture the images/videos of a face of the user/driver wearing eyeglasses as well as the images of the road or outside of the vehicle. The pre-trained model 212 as shown in detail in FIG. 5, can cause the processing device 102 to crop one or more eye regions from the captured images of the face of the user/driver occluded with eyeglass and correspondingly detect one or more key points from the captured images. In an embodiment, the one or more key points can be associated with a sclera, an iris, and an eyeball center of the user/driver as shown in detail in FIG. 4.
[0049] In an embodiment, the processing device 102 can be further configured with a learning module 214, such as a random forest model as shown in FIG. 6, which can cause the processing device 102 to to determine the eye gaze direction of the user based on the detected key points. The learning module 214 enables the processing device 102 to regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user/driver with or without eyeglasses. Accordingly, the pre-trained model 212 and learning module 214 can enable the processing device 102 to accurately identify the direction in which the user/driver wearing eyeglasses is looking while driving a vehicle, which can help avoid distractions and help prevent road accidents.
[0050] In an exemplary embodiment, as shown in FIG. 5, the pre-trained model 212 can include a first branch configured with a convolutional neural network (CNN) architecture and a second branch including a classification network. The model is pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass. The CNN architecture can include case-based reasoning (CBR) blocks. In addition, heat maps can be fed to the CNN architecture to detect the one or more key points. The second branch can be fed with an output of at least one of the CBR blocks, where the second branch can be trained to reduce entropy loss for enabling the model to determine the one or more real key points of the eye based on training with synthetic data.
[0051] In an exemplary embodiment, as shown in FIG. 6, the random forest model 214 can be fed with four vectors derived from the iris center and extreme joints of the sclera, and/or angular distance between the extreme sclera point and the iris center. Further, the random forest model 214 can be trained with one or more noise, key-points removal, and one or more additional features to identify the yaw and gaze for better accuracy and reliability.
[0052] In another exemplary embodiment, the processor 202 can cause the processing device 102 to pre-train or train the learning module 214 (random forest model) in real-time, with synthetic training data and/or real-time data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which can enable the learning module 214 to determine the gaze direction of the user. During the training, the processing device 102 can receive the first set of images associated with the eye regions and the corresponding one or more key points present in the first set of images. Further, the processing device 102 can render the eyeglasses at a predefined transparency over at least one of the eye regions in a predefined percentage of the first set of images to provide the second set of images as shown in FIG. 3B. Accordingly, the processing device 102 can train and test the model 212 and the learning module (random forest model) 214 with the synthetic data comprising the first set of images, the second set of images, and the corresponding one or more key points, which can enable the learning module 214 to accurately detect the one or more key points in the captured images of the user in realtime.
[0053] In an embodiment, the actuation and control unit 216 can cause the processing device 102 to generate and transmit a set of alert signals to the VCU 106 of the vehicle to alert the user/driver when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time, wherein the ROI is a road where the vehicle is traveling. [0054] In another embodiment, the actuation and control unit 216 can cause the processing device 102 to actuate the VCU 106 to stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI (road) for the predefined time.
[0055] In another embodiment, the ROI can be a road where the vehicle is running. The alert unit 218 can cause the processing device 102 to generate an alert when the estimated gaze direction of the user/driver is estimated to be off the road (ROI) for a first predefined time indicating the user/driver is distracted. Further, alert unit 218 can also cause the processing device 102 to generate an alert when the eye of the user/driver is found to be closed for a second predefined time indicating the user/driver to be sleeping or unconscious.
[0056] Referring to FIG. 3A, the proposed method 300 for estimation of the eye gaze direction of a user with or without wearing eyeglasses in a vehicle, involves the image acquisition unit, and the processing device connected with the VCU of the vehicle. The method 300 includes step 302 of capturing, by the image acquisition unit, one or more images of a face of the user, followed by step 304 of cropping, by the processing device, one or more eye region from the captured one or more images of the face of the user.
[0057] The method 300 further includes step 306 of detecting, by a processing device configured with a CNN architecture, one or more key points from one or more images associated with one or more eye of a user of a vehicle in real-time. The CNN architecture can be pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass.
[0058] Further, the method 300 includes step 308 of determining, using a learning module associated with the processing device, the eye gaze direction of the user based on the detected key points. At step 308, the processing device can regress the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
[0059] Referring to FIG. 3B and 3C, the image acquisition unit can capture the images of the user/driver sitting in the vehicle and a face detector module can detect face in the captured images. The processing device can then crop the face and further crop eye regions from the captured images. The CNN model can then detect key-points in the cropped eye images. In case, no key points are detected, no gaze condition is identified by the processing device. Further, the random forest model can then select a non-occluded eye and estimate the gaze direction of the user/driver.
[0060] Thus, the present invention (system and method) accurately estimates the eye gaze direction of a user/driver with or without wearing eyeglasses while driving the vehicle. In addition, the present invention alerts the driver or switches the vehicle to an auto-mode and stops the vehicle for the safety of the driver if the gaze direction of the driver gets off the road for a few seconds, thereby avoiding distractions and preventing road accidents.
[0061] While the foregoing describes various embodiments of the invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof. The scope of the invention is determined by the claims that follow. The invention is not limited to the described embodiments, versions or examples, which are included to enable a person having ordinary skill in the art to make and use the invention when combined with information and knowledge available to the person having ordinary skill in the art. ADVANTAGES OF THE PRESENT DISCLOSURE
[0062] The present disclosure monitors the eye gaze direction of drivers with or without wearing eyeglasses.
[0063] The present disclosure provides an efficient, reliable, and faster system and method for the estimation of the eye gaze direction of drivers wearing eyeglasses or when the eye of the driver is occluded.
[0064] The present disclosure accurately identifies the direction in which the driver wearing eyeglasses is looking while driving a vehicle to avoid distractions and help prevent road accidents.
[0065] The present disclosure alerts the driver or switches the vehicle to an auto-mode and stops the vehicle for the safety of the driver if the gaze direction of the driver gets off the road for a few seconds.
[0066] The present disclosure provides an efficient, and reliable system and a method for estimation of the eye gaze direction of drivers with or without wearing eyeglasses, which reminds or alerts the driver or switches the vehicle to an auto-mode and stops the vehicle if the gaze direction of the driver gets off the road for few seconds.

Claims

Claims:
1. A system (100) for estimation of an eye gaze direction of a user with or without wearing eyeglasses, the system (100) comprising: a processing device (102) incorporating a model (212) with a first branch configured with a convolutional neural network (CNN) architecture, wherein the CNN architecture comprises case-based reasoning (CBR) blocks, and wherein heat maps are fed to the CNN architecture to detect the one or more key points, the model (212) being pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass, and a learning module (214); the processing device (102) comprising a processor (202) coupled with a memory (204), wherein the memory (204) stores one or more instructions executable by the processor (202) to: detect, using the CNN architecture of the model (212), one or more key points from one or more images associated with the one or more eye of the user of a vehicle in real-time; and determine, using the learning module (214), the eye gaze direction of the user based on the detected key points.
2. The system (100) as claimed in claim 1 , wherein the system (100) comprises an image acquisition unit (104) comprising a camera to capture one or more images of a face of the user, wherein the processing device (102) is configured to: crop one or more eye region from the captured one or more images of the face of the user; detect, using the CNN architecture of the model (212), the one or more key points from the cropped one or more images; and regress, using the learning module (214), the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
3. The system (100) as claimed in claim 1 , wherein the pre-trained model (212) of the processing device (102) also comprises a second branch including a classification network, the second branch is fed with an output of at least one of the CBR blocks, wherein the second branch is trained to reduce entropy loss for enabling the model (212) to determine the one or more real key points of the eye based on training with synthetic data.
4. The system (100) as claimed in claim 1 , wherein the processing device (102) is configured to train the learning module (214), with synthetic training data and/or realtime data comprising a first set of images associated with eye regions, and a second set of images associated with the eye regions being occluded with eyeglasses, which enables the learning module (214) to determine the gaze direction of the user; and wherein the one or more key points are associated with a sclera, an iris, and an eyeball center.
5. The system (100) as claimed in claim 4, wherein the learning module (214) is a random forest model; wherein the random forest model (214) is fed with four vectors derived from any one or a combination of the iris center and extreme joints of the sclera, and angular distance between the extreme sclera point and the iris center; and wherein, the random forest model (214) is trained with one or more noise, keypoints removal, and one or more additional features to identify the yaw and gaze.
6. The system (100) as claimed in claim 1 , wherein the processing device (102) is configured to: generate a set of alert signals when the gaze direction of the user is determined to be out of a region of interest (ROI) for a predefined time; and stop the vehicle or switch the vehicle into an auto-driving mode when the gaze direction of the user is determined to be out of the ROI for the predefined time, wherein the ROI is a road where the vehicle is traveling.
7. The system (100) as claimed in claim 1 , wherein when one of the eyes of the user is at least partially occluded, the processing device (102) is configured to determine the gaze direction of the user using the one or more key points associated with another non-occluded eye of the user.
8. A method (300) for estimation of an eye gaze direction of a user, the method (300) comprises the steps of: detecting (306), by a processing device (102) incorporating a model (212) with a first branch configured with a convolutional neural network (CNN) architecture, one or more key points from one or more images associated with one or more eye of a user of a vehicle in real-time, wherein the CNN architecture has been pre-trained to detect one or more key points from one or more images associated with one or more eye of a user occluded with an eyeglass; and determining (308), using a learning module (214) associated with the processing device (102), the eye gaze direction of the user based on the detected key points.
9. The method (300) as claimed in claim 8, wherein the method (300) comprises the steps of: capturing (302), by an image acquisition unit (104), one or more images of a face of the user; cropping (304), by the processing device (102), one or more eye region from the captured one or more images of the face of the user; detecting (306), using the CNN architecture, the one or more key points from the cropped one or more images; regressing (308), using the learning module (214), the detected key points to calculate a yaw and pitch indicative of the eye gaze direction of the user.
EP23838037.2A 2023-01-10 2023-12-21 System and method for estimation of eye gaze direction of a user with or without eyeglasses Pending EP4602568A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GB2300344.5A GB2626136A (en) 2023-01-10 2023-01-10 System and method for estimation of eye gaze direction of a user with or without eyeglasses
PCT/EP2023/087406 WO2024149603A1 (en) 2023-01-10 2023-12-21 System and method for estimation of eye gaze direction of a user with or without eyeglasses

Publications (1)

Publication Number Publication Date
EP4602568A1 true EP4602568A1 (en) 2025-08-20

Family

ID=89541969

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23838037.2A Pending EP4602568A1 (en) 2023-01-10 2023-12-21 System and method for estimation of eye gaze direction of a user with or without eyeglasses

Country Status (6)

Country Link
EP (1) EP4602568A1 (en)
JP (1) JP2026501804A (en)
KR (1) KR20250116768A (en)
CN (1) CN120500709A (en)
GB (1) GB2626136A (en)
WO (1) WO2024149603A1 (en)

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3789848B1 (en) 2019-09-05 2024-07-17 Smart Eye AB Determination of gaze direction
WO2021146177A1 (en) * 2020-01-13 2021-07-22 Eye Tech Digital Systems, Inc. Systems and methods for eye tracking using machine learning techniques
CN111966219B (en) * 2020-07-20 2024-04-16 中国人民解放军军事科学院国防科技创新研究院 Eye tracking method, device, equipment and storage medium
US11704814B2 (en) * 2021-05-13 2023-07-18 Nvidia Corporation Adaptive eye tracking machine learning model engine
CN113642393B (en) * 2021-07-07 2024-03-22 重庆邮电大学 Attention mechanism-based multi-feature fusion sight estimation method

Also Published As

Publication number Publication date
JP2026501804A (en) 2026-01-16
CN120500709A (en) 2025-08-15
WO2024149603A1 (en) 2024-07-18
GB2626136A (en) 2024-07-17
KR20250116768A (en) 2025-08-01

Similar Documents

Publication Publication Date Title
JP7369184B2 (en) Driver attention state estimation
US11783600B2 (en) Adaptive monitoring of a vehicle using a camera
US20160272217A1 (en) Two-step sleepy driving prevention apparatus through recognizing operation, front face, eye, and mouth shape
US20200334477A1 (en) State estimation apparatus, state estimation method, and state estimation program
EP3871204A1 (en) A drowsiness detection system
CN110448316A (en) Data processing apparatus and method, monitoring system, wake-up system, and recording medium
EP2060993B1 (en) An awareness detection system and method
JP6090129B2 (en) Viewing area estimation device
JP6926636B2 (en) State estimator
EP3912149B1 (en) Method and system for monitoring a person using infrared and visible light
US20200104617A1 (en) System and method for remote monitoring of a human
JP7046748B2 (en) Driver status determination device and driver status determination method
US20250276645A1 (en) Adaptive monitoring of a vehicle using a camera
WO2024149603A1 (en) System and method for estimation of eye gaze direction of a user with or without eyeglasses
JP2019087018A (en) Driver monitor system
JP7704008B2 (en) Driver state determination method and device
JP7697913B2 (en) Gaze estimation device, gaze estimation computer program, and gaze estimation method
AU2021105935A4 (en) System for determining physiological condition of driver in autonomous driving and alarming the driver using machine learning model
JP2024162514A (en) Open/closed eye estimation device, open/closed eye estimation method, and program
JP7745829B2 (en) Driver state determination method and device
JP7745830B2 (en) Driver state determination method and device
JP7745831B2 (en) Driver state determination method and device
KR20250116150A (en) Brightness control for in-vehicle infotainment systems using gaze estimation
JP7433155B2 (en) Electronic equipment, information processing device, estimation method, and estimation program
WO2025115202A1 (en) Occupant state detection device, occupant state detection program, and occupant state detection method

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250516

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)