EP4670132A1 - FACE RECOGNITION SYSTEM TO REDUCE AMBIENT LIGHT DISTURBANCE DURING IMAGE CAPTURE - Google Patents

FACE RECOGNITION SYSTEM TO REDUCE AMBIENT LIGHT DISTURBANCE DURING IMAGE CAPTURE

Info

Publication number
EP4670132A1
EP4670132A1 EP24705650.0A EP24705650A EP4670132A1 EP 4670132 A1 EP4670132 A1 EP 4670132A1 EP 24705650 A EP24705650 A EP 24705650A EP 4670132 A1 EP4670132 A1 EP 4670132A1
Authority
EP
European Patent Office
Prior art keywords
facial
face
exposure
image
module
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24705650.0A
Other languages
German (de)
French (fr)
Inventor
Hui Yang
Hao Wang
Yingxin SONG
Renning LIU
Xin Chen
Dongyu ZHOU
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Thales DIS France SAS
Original Assignee
Thales DIS France SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Thales DIS France SAS filed Critical Thales DIS France SAS
Publication of EP4670132A1 publication Critical patent/EP4670132A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/20Image signal generators
    • H04N13/204Image signal generators using stereoscopic image cameras
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/172Classification, e.g. identification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/10Image acquisition
    • G06V10/12Details of acquisition arrangements; Constructional details thereof
    • G06V10/14Optical characteristics of the device performing the acquisition or on the illumination arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/60Extraction of image or video features relating to illumination properties, e.g. using a reflectance or lighting model
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/94Hardware or software architectures specially adapted for image or video understanding
    • G06V10/95Hardware or software architectures specially adapted for image or video understanding structured as a network, e.g. client-server architectures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • G06V40/166Detection; Localisation; Normalisation using acquisition arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/50Maintenance of biometric data or enrolment thereof

Definitions

  • the present invention relates to the technical field of facial recognition and, namely, to the field of facial recognition techniques for alleviating ambient light illumination drawbacks.
  • Facial recognition is the process of identifying or verifying the identity of a person using his/her face.
  • an image of the person’s face is first captured, and then analyzed to identify patterns based on the facial details which are finally compared with records in a dataset.
  • detecting and locating human faces in images and videos i.e. , face detection plays an essential role to acquire good-quality pictures.
  • Facial recognition improves customer/passenger experience, increases operational efficiency, and enhances security.
  • One extended application is access control, allowing a fast, transparent and secure access control by capturing and processing the live image of an individual, identifying whether he or she is an authorized person and granting access to a building, a place, or other restricted places, accordingly.
  • facial recognition software can run on multiple platforms such as on-premise, on a PC, in the cloud, in a mobile phone, a pod or tablet, and other different types of embedded environments.
  • the pod, kiosk or tablet is a specifically designed device for face capture and identification, used at physical access control points.
  • these “off-the-shelf’ devices should deal with and adapt to changing ambient conditions (e.g., lighting) and human operation.
  • ambient conditions e.g., lighting
  • human operation When controlled lighting is not possible, repetitively failing in capturing good-quality images drastically degrades user experience and device performance.
  • - a strong ambient light directly towards the facial recognition device’s screen can bring extreme back-lighting and darkening of the face;
  • - a strong ambient light directly towards the person’s face can bring extreme lighting and over exposure on the face;
  • the device will have difficulty to accurately detect the face when a passenger passes by.
  • facial recognition devices are typically equipped with quality check modules which verify in real-time that a minimum set of visibly distinguishing pertinent features are present within the captured faces (either regular photos or frames from a video stream). Only if the image quality is acceptable, it is proceeded further with the actual face matching bringing forth a time consuming experience.
  • the present invention provides a solution for the aforementioned problems by a facial recognition system according to claim 1 , and a method for alleviating ambient light during image capturing using the facial recognition according to claim 9.
  • a facial recognition system according to claim 1
  • the invention provides a facial recognition system for alleviating ambient light during image capture, the system comprising: a stereo camera configured to capture an image and send it to a face detection module, wherein the image comprises depth information of objects within the depth of field, the face detection module configured to identify a face based on the depth information of the received image from the stereo camera, a processing module configured to o split the image into a facial area and a non-facial area, wherein the facial area corresponds to the identified face region by the face detection module, and o outcome exposure adjustment factors for each area, these exposure adjustment factors being calculated based at least on a brightness values comparison of both facial and non-facial areas, a dual-regional adjustment module configured to separately adjust the exposure of both the facial and non-facial areas according to the adjustment factors, and a face matching module configure to verify or enroll person’s face identity based on the exposure-adjusted facial area.
  • the system according to the present invention allows for a two-step face detection: the stereo camera together with the face detection module perform a pre-detection of a face via depth perception, while the processing module, the dual-regional adjustment and the face matching module perform a further-step facial recognition based on the enhanced face image, in that to have a higher accuracy.
  • a stereo camera is a type of camera with two or more lenses with a separate image sensor for each lens. This allows the camera to simulate human binocular vision, and therefore gives it the ability to capture three-dimensional images (i.e. , 3D images or “depth map”) in addition to or on top of color images (such as Red-Green-Blue, “RGB”, images).
  • the captured image may be a color image comprising underlying depth information.
  • the stereo camera can perceive the distance and position of objects within its depth of field thus producing a depth map or depth information.
  • the stereo camera makes it possible to detect a neighboring face with high accuracy regardless of the actual ambient light, i.e., it can easily detect faces under dark or strong ambient light.
  • the image can be individually captured by the stereo camera or being extracted as a frame from an on-live or already recorded video stream.
  • the captured image arrives at the processing module either directly from the stereo camera or indirectly through the face detection module.
  • the processing module knows on which region of the image the face is present and thus associates it to a “facial area”.
  • the remaining part of the image is tagged as “non-facial area”. Since face recognition devices are typically designed to be operated pointing towards target’s face (e.g., to minimize errors), the facial area may cover the central part of the captured image (e.g., a rectangle / squared region).
  • the processing module then compares at least the brightness values of both facial and nonfacial areas in order to outcome the needed exposure adjustment factors to compensate the original exposure of each area.
  • the respective compensating exposure adjustment required in each of the areas are passed to the dual-regional adjustment module which separately (or dedicatedly) adjusts the exposure of each facial and non-facial area according to their own adjustment factors.
  • the dual-regional adjustment module may also re-construct an output image by bringing back together the separately exposure-adjusted facial and non-facial areas which is, in overall, more homorganic than the originally un-adjusted captured image.
  • the dual-regional adjustment module sends to the face matching module at least the part of the image with the adjusted facial area.
  • the borders of the image with the exposure-adjusted facial area can adapt dynamically to better suit a default format or orientation of the portrait to be properly digested by the facial matching software.
  • This facial matching software can be used to extract the particulars of the face. If the user is enrolling to a service, the particulars of his captured face image will be linked to his declared identity to be able to be authenticated later on during an eventual verification process. Otherwise, if already enrolled, the particulars of his captured face image will be compared (e.g., 1 :1 or 1 :N matching) with already existing records in a database while assessing the likelihood they match.
  • this facial matching software applies a (convolutional) neural network (CNN) to the exposure-adjusted facial area image.
  • CNN Convolutional Neural Network
  • a Convolutional Neural Network (CNN) is a Deep Learning algorithm used to analyze imagery which takes in an input image, assign importance (e.g., learnable weights and biases) to various aspects and objects in the image and is able to differentiate one from the other.
  • CNNs are a specialized type of neural networks that use convolution in place of general matrix multiplication in at least one of their layers, mostly used in image recognition and processing specifically designed to process pixel data. It consists of an input layer, hidden layers and an output layer. The CNN is especially configured to reduce the images into a form which is simpler to process without sacrificing feature reduction while maintaining its ability to discriminate and classify images.
  • the exposure- adjusted facial area image can be further pre-processed for cropping and de-skewing the face area image as part of a normalization process (i.e., some CNN models require meeting a particular input format).
  • De-skewing is the process of straightening an image that has been scanned or captured crookedly; that is, an image that is slanting too far in one direction, or one that is misaligned.
  • Croping is the removal of unwanted outer areas, namely, those portions irrelevant, unnecessary, or serving of little value with respect to the security patterns or other image features relevant to categorizing or classifying the face in the image.
  • having the exposure compensation achieved by the present invention during enrolment and verification stages dispenses with the expensive and heavy-to-carry lighting equipment that are required today in kiosks, I Dentity verification spots, etc.
  • this dual-regional adjustment also works as a liveness solution since realistic lightening interaction with the face may ensure that a real person is physically present during the image capturing process and helps in detecting any spoof attack. For instance, if a just previously recognized face required certain dual-region exposure adjustment factors and, suddenly, a new captured face does not needed it, it may be an indicia of a spoof attack. Especially if the stereo camera didn’t move, e.g., known by non-facial background area analysis, geo-positioning data analysis, etc. Moreover, spoofing depth maps is more complex than a simple colour image.
  • these exposure adjustment factors are calculated based on brightness and exposure values comparison of both facial and non-facial areas.
  • the processing module is further configured to calculate the brightness and exposure values of each pixel in the facial and non-facial areas, calculate representative brightness and exposure values for each area, and compare them to outcome compensating exposure adjustment factors to be applied separately to each area by the dual-regional adjustment module.
  • the system further comprises a screen.
  • the dual-regional adjustment module is further configured to re-construct an image with the separately exposure-adjusted facial and non-facial areas to be displayed on such screen.
  • the dual-regional adjustment module displays on the screen the more homorganic reconstructed output image to be presented to the user as part of the user interface.
  • system further comprises supplementary lighting to be activated in case the face detection module identifies a face in any of the images coming from the stereo camera and the ambient light is below a pre-defined threshold.
  • the supplementary lighting is only activated when a face is detected via depth perception. So it is less costly and more efficient than any prior art apparatus which keeps the lights turning on and off in pre-defined times.
  • This embodiment is especially applicable to outdoor night-time or any dark environment.
  • the system further comprises a light sensor configured to detect if the ambient light around the stereo camera is below a pre-defined threshold. If that happens, the light sensor notifies the face detection module which will activate the supplementary lighting in case of detecting a face nearby.
  • the stereo camera itself is configured to detect if the neighboring ambient light is below a pre-defined threshold in order to activate the supplementary light in case the face detection module detects a face.
  • the dual-regional adjustment module is a built-in feature of the stereo camera so that the exposure adjustment factors calculated by the processing module are sent back to the stereo camera to apply the required exposure adjustment accordingly.
  • the stereo camera is capturing only once.
  • the dedicated exposure-adjustment is done by the stereo camera itself by feeding back the exposure adjustment factors. Therefore, the stereo camera maturely supports smart exposure feature.
  • the system further comprises a facial recognition device, such as a tablet or pod.
  • the stereo camera is operatively connected to the face recognition device, either embedded as an integrative solution, or integrated dedicatedly as a plug-in solution.
  • the stereo camera can be either embedded into the device body, or integrated dedicatedly via a wired or wireless connection such as cables, Bluetooth, Wi-Fi Direct, etc.
  • the device can be any type of portable computing device e.g., tablet, smartphone, laptop computer, navigation device, game console, desktop computer system, workstation, Internet appliance and the like, or a stationary device such as a pod or kiosk provided further with a scanner and other cameras for capturing other images.
  • portable computing device e.g., tablet, smartphone, laptop computer, navigation device, game console, desktop computer system, workstation, Internet appliance and the like
  • stationary device such as a pod or kiosk provided further with a scanner and other cameras for capturing other images.
  • the face recognition device e.g., tablet, pod
  • the face recognition device can be a standalone face recognition apparatus embedding any of the face detection module, processing module, dual-regional adjustment module, and face matching module. Hence, it can be used off-line.
  • the device When the device integrates all (or most of) the modules, it may comprise an installable package in the form of a Software Development Kit (SDK) configured to carry out by itself all or part of the steps that can be allocated to and performed by a server side.
  • SDK Software Development Kit
  • This SDK is designed to quickly build a standalone application that needs facial recognition using either photos or videos, including liveness detection features. It may run on many environments including embedded environment.
  • one or more of the modules are feature(s) of the operating system (OS) of the device, e.g., implemented as part of iOS or Android.
  • OS operating system
  • the system is a distributed system where at least one of the face detection module, the processing module, the dual-regional adjustment module, or the face matching module is running at server side as a backend service.
  • the captured image through the stereo camera can be partially or entirely processed in-situ thanks to the face recognition device or can be requested via a communication network to a remote server-based face recognition and analysis system comprising Artificial Intelligence Convolutional Neural Network models coupled to different databases (e.g., storing reference face images).
  • a remote server-based face recognition and analysis system comprising Artificial Intelligence Convolutional Neural Network models coupled to different databases (e.g., storing reference face images).
  • the electronic device will only capture or get capture the image and send it together with a verification request to backend services.
  • the invention provides a method for alleviating ambient light during image capturing using the facial recognition system according to any of the embodiments of the first inventive aspect.
  • the method comprises the following steps: capturing, by the stereo camera, an image comprising depth information of objects within the depth of field, sending, by the stereo camera, the captured image to a face detection module, identifying, by the face detection module, a face based on the depth information of the received image from the stereo camera, receiving, by the processing module, the image and the region of the image where a face is identified, splitting, by the processing module, the image into a facial area and a non-facial area, wherein the facial area corresponds to the identified face region by the face detection module, outcome, by the processing module, exposure adjustment factors for each area, wherein these exposure adjustment factors are calculated based at least on a brightness values comparison of both facial and non-facial areas, adjusting separately, by a dual-regional adjustment module, the exposure of both the facial and non-facial areas according to the adjustment factors, and verify
  • the processing module calculates the brightness and exposure values of each pixel in the facial and non-facial areas. Then, the processing module calculates representative brightness and exposure values for each area, and compares them to outcome compensating exposure adjustment factors to be applied separately to each of the areas by the dual-regional adjustment module. In a particular embodiment, the method further comprises the steps of: re-constructing, by the dual-regional adjustment module, a re-constructed image with the separately exposure-adjusted facial and non-facial areas, and sending the re-constructing image to the screen to be displayed thereon.
  • the method further comprises the steps of: detecting, preferably by a light sensor, if the ambient light around the stereo camera is below a pre-defined threshold, and activating, by the face detection module, the supplementary lighting.
  • FIG. 1 This figure shows an embodiment of a face recognition device according to some embodiments of the present invention.
  • FIG. 2 This figure shows a schematic workflow of a facial recognition system when the ambient light around the stereo camera is above a pre-defined threshold.
  • Figures 3a to 3c show examples of dual-regional adjustments according to embodiments of the invention, when (a) there is a strong ambient light directly towards the facial recognition device’s screen; (b) there is a strong ambient light directly towards the person’s face; and (c) there is a strong ambient light towards a side of the face.
  • FIG. 4 This figure shows a schematic workflow of a facial recognition system when the ambient light around the stereo camera is below a pre-defined threshold.
  • aspects of the present invention may be embodied either as a facial recognition system or a method for alleviating ambient light during image capture.
  • the present invention relates to products and solutions where the user needs to be authenticated through his or her face.
  • This step could be part of an authorization request where he or she should prove his or her identity when onboarding, enrolling, authenticating to, or requesting a service.
  • This “service” can be access control to a physical or virtual restricted area. If the user is pre-authorized (e.g., already enrolled and whitelisted in the access control list) and his/her identity proven through his/her face, access to the restricted area will be granted. Otherwise, access is rejected.
  • the identity verification may occur either through a dedicated Software application installed in the face recognition device itself, or in cooperation with another service provider’s mobile application to which the user has logged-in, or through a web browser in the electronic device.
  • the present invention may be used in remote identity verification (e.g., Know-Your-Customer, “KYC”, solutions) where two images are captured: an image containing an alleged user’s I Dentity document (e.g., national ID, passport, driving license) as well as the user’s face.
  • I Dentity document e.g., national ID, passport, driving license
  • prior art liveness detection techniques can be applied to the latter image. Facial biometrics allows recognizing and measuring facial features from a captured image of a face (e.g., from a selfie). Liveness detection further ensures that a real person was physically present during the capture and detects any spoof attack.
  • the authenticity of the captured document is checked to prevent fraud by verifying data integrity, consistency and security elements; while face matching assesses the likelihood that the captured person is the same person as the one on the ID document. If user’s identity is verified, he/she is enrolled or his/her access granted to the requested service.
  • figure 1 it is schematically represented a face recognition system (1) configured to alleviate ambient light during the capture of a face image according to the invention.
  • the face recognition system (1) is configured to verify identities based on photo(s) of the face (3) of a user.
  • the face recognition system (1) of figure 1 is exemplified and represented throughout the figures as a face recognition device (10) in the form of a tablet.
  • the tablet (10) includes one or several (micro)processors (and/or a (micro)controller(s)), as data processing means (not shown), comprising and/or being connected to one or several memories, as data storing means, comprising or being connected to means for interfacing with the user, such as a Man Machine Interface (or MMI), and comprising or being connected to an Input/Output (or I/O) interface(s) that are internally all connected, through an internal bidirectional data bus.
  • MMI Man Machine Interface
  • I/O Input/Output
  • the I/O interface(s) may include a wired and/or a wireless interface, to exchange, over a contact and/or Contactless (or CTL) link(s), with a user.
  • the MMI may include a display screen(s), a keyboard(s), a loudspeaker(s) and/or a camera(s) and allows the user to interact with the tablet.
  • the MMI may be used for getting data entered and/or provided by the user.
  • the MMI comprises a stereo camera (2) configured to capture images of the user’s face (3). These images (3) comprise depth information (3.1) of objects within the depth of field.
  • the tablet memory(ies) may include one or several volatile memories and/or one or several non-volatile memories (not shown).
  • the memory(ies) may store data, such as an ID(s) relating to the tablet or hosted chip, that allows identifying uniquely and addressing the tablet.
  • the memory(ies) also stores the Operating System (OS) and applications which are adapted to provide services to the user (not shown).
  • OS Operating System
  • the tablet has installed an orchestrator able to receive ID verification requests from service provider apps or web browsers and e.g., routing them towards backend services hosted in server(s)
  • the tablet (10) is configured to send information over a communication network
  • the orchestrator and the backend infrastructure may communicate via one or more Application Program Interfaces, APIs, using HTTP over TLS.
  • the tablet (10) can verify the identity of the user by itself or in cooperation with remote server(s) (11). Therefore, in a basic configuration, the tablet (10) only comprises the stereo camera (2) configured to capture images of the user’s face (3) and then send them to server(s) end (11) for further processing: face detection, image processing, dual- regional exposure adjustment, face or biometrics matching, and ID verification.
  • the tablet typically comprises an application or orchestrator configured, on one side, to receive requests from other service provider applications installed in the device, or from web browsers and, on the other side, communicate with backend services running in the server(s) (11) side.
  • These one or more servers may comprise (jointly or separately) others modules which allows performing the steps which are being described hereinafter in relation with the tablet.
  • the tablet then comprises, a face detection module (4) configured to identify a face based on the depth information (3.1) of the received image (3) from the stereo camera (2), a processing module (5) configured to split the image (3) into a facial and non-facial areas and to calculate exposure adjustment factors for each area based at least on a brightness values comparison, a dual-regional adjustment module (6) configured to separately adjust the exposure of both the facial and non-facial areas according to the adjustment factors, and a face matching module (7) configure to verify or enroll person’s face identity based on the exposure-adjusted facial area.
  • a face detection module (4) configured to identify a face based on the depth information (3.1) of the received image (3) from the stereo camera (2)
  • a processing module (5) configured to split the image (3) into a facial and non-facial areas and to calculate exposure adjustment factors for each area based at least on a brightness values comparison
  • a dual-regional adjustment module (6) configured to separately adjust the exposure of both the facial and non-facial areas according
  • the dual-regional adjustment module (6) is a built-in feature of the stereo camera (2) so that the exposure adjustment factors calculated by the processing module (5) are sent back to the stereo camera to apply the exposure adjustment accordingly.
  • the tablet (9) further comprises a screen (8) to display a re-constructed image (3’) with the separately exposure-adjusted facial and non-facial areas made by the dual-regional adjustment module.
  • This screen prompts hereby the user with a more homorganic representation than the originally un-adjusted captured image. Additional information can be conveyed to the user together with the homorganic image such as information about the success of the face matching and ID verification, or enrolment, and the % certitude. For instance, it can additionally prompt the message “Welcome to enterprise X, Mr. name”, “Have a nice trip to destiny”, etc.
  • the tablet (9) further comprises a built-in supplementary lighting (9) to be activated in case the face detection module identifies a face in any of the images coming from the stereo camera and the ambient light is below a pre-defined threshold.
  • a built-in supplementary lighting 9 to be activated in case the face detection module identifies a face in any of the images coming from the stereo camera and the ambient light is below a pre-defined threshold.
  • light sensors (not shown) may be used.
  • Figure 2 depicts a schematic workflow (20) of a facial recognition system when the ambient light around the stereo camera is above a pre-defined threshold. That is, for instance, image is captured under day-light or in a properly illuminated room.
  • the steps can be taken e.g., by the facial recognition system (1) described in relation to figure 1.
  • the stereo camera (2) captures (21) an image (3) comprising depth information of objects within the depth of field, and send it (3) to a face detection module (4).
  • the face detection module identifies (22) whether there is a face in the image thanks to the depth information embedded in the received image (3). If no face is detected, the process is aborted. This could be part of a loop analyzing frames from a video stream.
  • the face detection module (4) is configured to detect objects (so called, Regions of Interest or “ROI”) on the image, and to infer whether any of the found objects resembles a face thanks to an Machine Leaning, ML, -driven software.
  • ROIs are typically represented with bounding boxes for further class classification (e.g., faces). If more than ones classes are possible, e.g., faces, persons, pets, etc. the detection module can access a database or keyword table (not shown) of pre-determined classes to be possibly found in the captured images. Then, the detection module applies a nearest neighbor search within the reference database using the ROIs, namely an embedding vector for each ROI. From results of the search, it selects the most likely match candidate (if any).
  • the processing module (5) splits (24) the image (3) into a facial area (3.2) and a nonfacial area (3.3).
  • the facial area (3.2) corresponds to the identified face region (3.4) by the face detection module.
  • the processing module (5) then outcomes (25) exposure adjustment factors for each of the areas (3.2, 3.3), wherein these exposure adjustment factors are calculated based at least on a brightness values comparison of both facial (3.2) and nonfacial (3.3) areas.
  • the dual-regional adjustment module (6) then adjusts separately (26) the exposure of both the facial (3.2) and non-facial (3.3) areas according to the received adjustment factors.
  • the dual-regional adjustment module (6) can be a built-in feature of the stereo camera (2), so the smart exposure adjustment can be performed by the camera itself after receiving instruction about how to adjust the areas.
  • the un-adjusted image (3) and its facial (3.2) and non-facial (3.3) areas are distinguished from the dual-region exposure-adjusted image (or re-constructed image) in that the latter has a “ ‘ ” in its numeric references such as 3’, 3.2’, 3.3’.
  • the bounding box (3.4) bordering the facial region will be in principle the same both pre-adjusted (3, 3.2, 3.3) and post-adjusted (3’, 3.2’, 3.3’) images.
  • a face matching module (7) verifies a person’s face identity based on the exposure- adjusted facial area (3.2’).
  • the identification result will be accompanied by a % of certainty. Normally, the identification result not only states the name of the identified preenrolled user but also adjoins his/her name or other identifier.
  • the dual-regional adjustment module (6) can re-construct (28) the image (3’) with the separately exposure- adjusted facial (3.2’) and non-facial (3.3’) areas. Then, this re-constructed image can be sent to the screen (8) to be displayed thereon. Along with this re-constructed image (3’), the screen III can prompt the user with any of the following messages: the identification result, % certainty, or any customized welcome, informational or exception message.
  • Figures 3a to 3c depict examples of dual-regional adjustments according to embodiments of the invention. For each one, it is shown on the left-hand side the original image (3) captured and, on the right-hand side, the re-constructed (and more homorganic) image (3’).
  • Figure 3a depicts the situation when there is a strong ambient light directly towards the facial recognition device’s screen. As it can be seen on the left-hand side, it produces a dark facial area (3.2) and an over-exposed (i.e., too bright) non-facial area (3.3).
  • the light interference is innocuous for the depth perception feature of stereo camera and the face detection module can perform a first-step face detection.
  • the face detection module can perform a first-step face detection.
  • the box size can be adjusted.
  • the brightness values comparison of both facial (3.2) and non-facial areas (3.3) outcomes exposure adjustment factors that requires to dedicatedly compensate the exposure of the facial area (3.2) and decrease the exposure of the non-facial area (3.3), accordingly.
  • the resulting homorganic image (3’) can be seen on the right-hand side.
  • Figure 3b depicts the situation when there is there is a strong ambient light directly towards the person’s face. As it can be seen on the left-hand side, it produces an over-exposed facial area (3.2) and a relatively lesser-exposed / bright non-facial area (3.3).
  • Figure 3c depicts the situation when there is there is a strong ambient light towards a side of the face. As it can be seen on the left-hand side, it produces a half bright I half dark facial area (3.2) and a half bright I half dark non-facial area (3.3).
  • the brightness values comparison of both facial (3.2) and non-facial areas (3.3) outcomes exposure adjustment factors that requires to dedicatedly compensate the dark part of the facial area (3.2) with a higher exposure and decrease the over-exposure I bright part of the facial area (3.2) with lower exposure, while compensate the dark part of the non-facial area (3.3) and decrease the over-exposure I bright part of the non-facial area (3.3), accordingly.
  • Figure 4 depicts a schematic workflow (20) of a facial recognition system when the ambient light around the stereo camera is below a pre-defined threshold. That is, it is night-time, dark or there is insufficient light so as to properly capture visible face features.
  • the method of figure 4 is similar to the one of figure 2 but adds the turning-on of supplementary lights (9) if the ambient light around the stereo camera (2) is insufficient (i.e., below a pre-defined threshold). The lights will turn on in case the face detection module detects a face within the depth of field of the stereo camera.
  • the face detection module can be continuously analyzing a video stream to find faces and, if found, hence triggering the dual-regional adjustment.
  • the system can similarly analyse images captured by a stereo camera which are stored either in a database of photos or a collection of videos.
  • the supplementary lights can be switch off according to different use cases. For instance, for single captures it can simply flash, while for in-live video stream analysis it can be active for more time.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Software Systems (AREA)
  • Signal Processing (AREA)
  • Collating Specific Patterns (AREA)

Abstract

The present invention provides a facial recognition system for alleviating ambient light during image capture, the system comprising: - a stereo camera configured to capture an image and send it to a face detection module, wherein the image comprises depth information of objects within the depth of field, - the face detection module configured to identify a face based on the depth information of the received image from the stereo camera, - an processing module configured to - split the image into a facial area and a non-facial area, wherein the facial area corresponds to the identified face region by the face detection module, and - outcome exposure adjustment factors for each area, these exposure adjustment factors being calculated based at least on a brightness values comparison of both facial and non-facial areas, - a dual-regional adjustment module configured to separately adjust the exposure of both the facial and non-facial areas according to the adjustment factors, and - a face matching module configure to verify or enroll person's face identity based on the exposure-adjusted facial area.

Description

FACIAL RECOGNITION SYSTEM FOR ALLEVIATING AMBIENT LIGHT DURING
IMAGE CAPTURE
TECHNICAL FIELD
The present invention relates to the technical field of facial recognition and, namely, to the field of facial recognition techniques for alleviating ambient light illumination drawbacks.
BACKGROUND OF THE INVENTION
Facial recognition is the process of identifying or verifying the identity of a person using his/her face. In this process, an image of the person’s face is first captured, and then analyzed to identify patterns based on the facial details which are finally compared with records in a dataset. During this process, detecting and locating human faces in images and videos (i.e. , face detection) plays an essential role to acquire good-quality pictures.
Facial recognition improves customer/passenger experience, increases operational efficiency, and enhances security. One extended application is access control, allowing a fast, transparent and secure access control by capturing and processing the live image of an individual, identifying whether he or she is an authorized person and granting access to a building, a place, or other restricted places, accordingly.
To tackle this diverse use cases, facial recognition software can run on multiple platforms such as on-premise, on a PC, in the cloud, in a mobile phone, a pod or tablet, and other different types of embedded environments. Among them, the pod, kiosk or tablet is a specifically designed device for face capture and identification, used at physical access control points. When deployed either indoor or outdoor, these “off-the-shelf’ devices should deal with and adapt to changing ambient conditions (e.g., lighting) and human operation. When controlled lighting is not possible, repetitively failing in capturing good-quality images drastically degrades user experience and device performance.
Examples of typical ambient light interference challenges are:
- a strong ambient light directly towards the facial recognition device’s screen can bring extreme back-lighting and darkening of the face; - a strong ambient light directly towards the person’s face can bring extreme lighting and over exposure on the face;
- a strong ambient light towards the face from a side will cause different exposures (e.g., half dark and half bright) resulting in “light spot” or “hot spot” effects;
- within a dark environment without enough light, the device will have difficulty to accurately detect the face when a passenger passes by.
All above cases will interfere the face detection capability, efficiency and accuracy of the facial recognition device. To mitigate the above-listed and other situations (such as blurriness, exposure, luminosity, contrast), facial recognition devices are typically equipped with quality check modules which verify in real-time that a minimum set of visibly distinguishing pertinent features are present within the captured faces (either regular photos or frames from a video stream). Only if the image quality is acceptable, it is proceeded further with the actual face matching bringing forth a time consuming experience.
Thus, there is a need in the industry for a continuous improvement on the user experience for the face recognition solution.
SUMMARY OF THE INVENTION
The present invention provides a solution for the aforementioned problems by a facial recognition system according to claim 1 , and a method for alleviating ambient light during image capturing using the facial recognition according to claim 9. In dependent claims, preferred embodiments of the invention are defined.
In a first inventive aspect, the invention provides a facial recognition system for alleviating ambient light during image capture, the system comprising: a stereo camera configured to capture an image and send it to a face detection module, wherein the image comprises depth information of objects within the depth of field, the face detection module configured to identify a face based on the depth information of the received image from the stereo camera, a processing module configured to o split the image into a facial area and a non-facial area, wherein the facial area corresponds to the identified face region by the face detection module, and o outcome exposure adjustment factors for each area, these exposure adjustment factors being calculated based at least on a brightness values comparison of both facial and non-facial areas, a dual-regional adjustment module configured to separately adjust the exposure of both the facial and non-facial areas according to the adjustment factors, and a face matching module configure to verify or enroll person’s face identity based on the exposure-adjusted facial area.
Briefly, the system according to the present invention allows for a two-step face detection: the stereo camera together with the face detection module perform a pre-detection of a face via depth perception, while the processing module, the dual-regional adjustment and the face matching module perform a further-step facial recognition based on the enhanced face image, in that to have a higher accuracy.
A stereo camera (so-called depth camera) is a type of camera with two or more lenses with a separate image sensor for each lens. This allows the camera to simulate human binocular vision, and therefore gives it the ability to capture three-dimensional images (i.e. , 3D images or “depth map”) in addition to or on top of color images (such as Red-Green-Blue, “RGB”, images). The captured image may be a color image comprising underlying depth information.
The stereo camera can perceive the distance and position of objects within its depth of field thus producing a depth map or depth information. Thus, unlike infrared (“IR”) camera + RGB camera which are ineffective under extreme light conditions, the stereo camera makes it possible to detect a neighboring face with high accuracy regardless of the actual ambient light, i.e., it can easily detect faces under dark or strong ambient light.
The image can be individually captured by the stereo camera or being extracted as a frame from an on-live or already recorded video stream.
The captured image arrives at the processing module either directly from the stereo camera or indirectly through the face detection module. When it arrives, thanks to the face identified by the face detection module, the processing module knows on which region of the image the face is present and thus associates it to a “facial area”. The remaining part of the image is tagged as “non-facial area”. Since face recognition devices are typically designed to be operated pointing towards target’s face (e.g., to minimize errors), the facial area may cover the central part of the captured image (e.g., a rectangle / squared region).
The processing module then compares at least the brightness values of both facial and nonfacial areas in order to outcome the needed exposure adjustment factors to compensate the original exposure of each area.
While, in general, both the exposure and brightness corrections on an image act to brighten or darken it, they both do it in a different manner. In short, exposure has a heavier bias to highlight tones, while brightness has no bias and affects all tones equally. This means that adjusting exposure will affect highlights more in brightening or darkening an image than brightness.
Advantageously, using a dual exposure adjustment rather than brightness adjustment exposes a higher-quality face image. It was found that if the captured image is not clear and has some visual deficiencies, only adjusting the brightness will not improve the quality overall.
Therefore, the respective compensating exposure adjustment required in each of the areas are passed to the dual-regional adjustment module which separately (or dedicatedly) adjusts the exposure of each facial and non-facial area according to their own adjustment factors. The dual-regional adjustment module may also re-construct an output image by bringing back together the separately exposure-adjusted facial and non-facial areas which is, in overall, more homorganic than the originally un-adjusted captured image.
Finally, the dual-regional adjustment module sends to the face matching module at least the part of the image with the adjusted facial area. At this point, the borders of the image with the exposure-adjusted facial area can adapt dynamically to better suit a default format or orientation of the portrait to be properly digested by the facial matching software. This facial matching software can be used to extract the particulars of the face. If the user is enrolling to a service, the particulars of his captured face image will be linked to his declared identity to be able to be authenticated later on during an eventual verification process. Otherwise, if already enrolled, the particulars of his captured face image will be compared (e.g., 1 :1 or 1 :N matching) with already existing records in a database while assessing the likelihood they match.
Preferably, this facial matching software applies a (convolutional) neural network (CNN) to the exposure-adjusted facial area image. A Convolutional Neural Network (CNN) is a Deep Learning algorithm used to analyze imagery which takes in an input image, assign importance (e.g., learnable weights and biases) to various aspects and objects in the image and is able to differentiate one from the other. CNNs are a specialized type of neural networks that use convolution in place of general matrix multiplication in at least one of their layers, mostly used in image recognition and processing specifically designed to process pixel data. It consists of an input layer, hidden layers and an output layer. The CNN is especially configured to reduce the images into a form which is simpler to process without sacrificing feature reduction while maintaining its ability to discriminate and classify images.
Depending on the capture conditions of the face in the original image, the exposure- adjusted facial area image can be further pre-processed for cropping and de-skewing the face area image as part of a normalization process (i.e., some CNN models require meeting a particular input format). “De-skewing” is the process of straightening an image that has been scanned or captured crookedly; that is, an image that is slanting too far in one direction, or one that is misaligned. “Cropping” is the removal of unwanted outer areas, namely, those portions irrelevant, unnecessary, or serving of little value with respect to the security patterns or other image features relevant to categorizing or classifying the face in the image.
Advantageously, having the exposure compensation achieved by the present invention during enrolment and verification stages dispenses with the expensive and heavy-to-carry lighting equipment that are required today in kiosks, I Dentity verification spots, etc.
Collaterally, this dual-regional adjustment also works as a liveness solution since realistic lightening interaction with the face may ensure that a real person is physically present during the image capturing process and helps in detecting any spoof attack. For instance, if a just previously recognized face required certain dual-region exposure adjustment factors and, suddenly, a new captured face does not needed it, it may be an indicia of a spoof attack. Especially if the stereo camera didn’t move, e.g., known by non-facial background area analysis, geo-positioning data analysis, etc. Moreover, spoofing depth maps is more complex than a simple colour image.
In a particular embodiment, these exposure adjustment factors are calculated based on brightness and exposure values comparison of both facial and non-facial areas. In a preferred embodiment, the processing module is further configured to calculate the brightness and exposure values of each pixel in the facial and non-facial areas, calculate representative brightness and exposure values for each area, and compare them to outcome compensating exposure adjustment factors to be applied separately to each area by the dual-regional adjustment module.
In a particular embodiment, the system further comprises a screen. The dual-regional adjustment module is further configured to re-construct an image with the separately exposure-adjusted facial and non-facial areas to be displayed on such screen.
The dual-regional adjustment module displays on the screen the more homorganic reconstructed output image to be presented to the user as part of the user interface.
In a particular embodiment, the system further comprises supplementary lighting to be activated in case the face detection module identifies a face in any of the images coming from the stereo camera and the ambient light is below a pre-defined threshold.
Advantageously, the supplementary lighting is only activated when a face is detected via depth perception. So it is less costly and more efficient than any prior art apparatus which keeps the lights turning on and off in pre-defined times. This embodiment is especially applicable to outdoor night-time or any dark environment.
In a preferred embodiment, the system further comprises a light sensor configured to detect if the ambient light around the stereo camera is below a pre-defined threshold. If that happens, the light sensor notifies the face detection module which will activate the supplementary lighting in case of detecting a face nearby.
In another embodiment, the stereo camera itself is configured to detect if the neighboring ambient light is below a pre-defined threshold in order to activate the supplementary light in case the face detection module detects a face.
In a particular embodiment, the dual-regional adjustment module is a built-in feature of the stereo camera so that the exposure adjustment factors calculated by the processing module are sent back to the stereo camera to apply the required exposure adjustment accordingly.
As noted, the stereo camera is capturing only once. In this embodiment, the dedicated exposure-adjustment is done by the stereo camera itself by feeding back the exposure adjustment factors. Therefore, the stereo camera maturely supports smart exposure feature.
In a particular embodiment, the system further comprises a facial recognition device, such as a tablet or pod. The stereo camera is operatively connected to the face recognition device, either embedded as an integrative solution, or integrated dedicatedly as a plug-in solution.
That is, the stereo camera can be either embedded into the device body, or integrated dedicatedly via a wired or wireless connection such as cables, Bluetooth, Wi-Fi Direct, etc.
The device can be any type of portable computing device e.g., tablet, smartphone, laptop computer, navigation device, game console, desktop computer system, workstation, Internet appliance and the like, or a stationary device such as a pod or kiosk provided further with a scanner and other cameras for capturing other images.
The face recognition device (e.g., tablet, pod) can be a standalone face recognition apparatus embedding any of the face detection module, processing module, dual-regional adjustment module, and face matching module. Hence, it can be used off-line.
When the device integrates all (or most of) the modules, it may comprise an installable package in the form of a Software Development Kit (SDK) configured to carry out by itself all or part of the steps that can be allocated to and performed by a server side. This SDK is designed to quickly build a standalone application that needs facial recognition using either photos or videos, including liveness detection features. It may run on many environments including embedded environment.
In another example, one or more of the modules are feature(s) of the operating system (OS) of the device, e.g., implemented as part of iOS or Android.
In a particular embodiment, the system is a distributed system where at least one of the face detection module, the processing module, the dual-regional adjustment module, or the face matching module is running at server side as a backend service.
Therefore, the captured image through the stereo camera can be partially or entirely processed in-situ thanks to the face recognition device or can be requested via a communication network to a remote server-based face recognition and analysis system comprising Artificial Intelligence Convolutional Neural Network models coupled to different databases (e.g., storing reference face images). Thus, the electronic device will only capture or get capture the image and send it together with a verification request to backend services.
The skilled person in the art knows how to distribute functionalities among client and server side to enhance the operation and ease deployment.
In a second inventive aspect, the invention provides a method for alleviating ambient light during image capturing using the facial recognition system according to any of the embodiments of the first inventive aspect. The method comprises the following steps: capturing, by the stereo camera, an image comprising depth information of objects within the depth of field, sending, by the stereo camera, the captured image to a face detection module, identifying, by the face detection module, a face based on the depth information of the received image from the stereo camera, receiving, by the processing module, the image and the region of the image where a face is identified, splitting, by the processing module, the image into a facial area and a non-facial area, wherein the facial area corresponds to the identified face region by the face detection module, outcome, by the processing module, exposure adjustment factors for each area, wherein these exposure adjustment factors are calculated based at least on a brightness values comparison of both facial and non-facial areas, adjusting separately, by a dual-regional adjustment module, the exposure of both the facial and non-facial areas according to the adjustment factors, and verifying or enrolling, by a face matching module, a person’s face identity based on the exposure-adjusted facial area.
In a particular embodiment, the processing module calculates the brightness and exposure values of each pixel in the facial and non-facial areas. Then, the processing module calculates representative brightness and exposure values for each area, and compares them to outcome compensating exposure adjustment factors to be applied separately to each of the areas by the dual-regional adjustment module. In a particular embodiment, the method further comprises the steps of: re-constructing, by the dual-regional adjustment module, a re-constructed image with the separately exposure-adjusted facial and non-facial areas, and sending the re-constructing image to the screen to be displayed thereon.
In a particular embodiment, the method further comprises the steps of: detecting, preferably by a light sensor, if the ambient light around the stereo camera is below a pre-defined threshold, and activating, by the face detection module, the supplementary lighting.
All the features described in this specification (including the claims, description and drawings) and/or all the steps of the described method can be combined in any combination, with the exception of combinations of such mutually exclusive features and/or steps.
DESCRIPTION OF THE DRAWINGS
These and other characteristics and advantages of the invention will become clearly understood in view of the detailed description of the invention which becomes apparent from a preferred embodiment of the invention, given just as an example and not being limited thereto, with reference to the drawings.
Figure 1 This figure shows an embodiment of a face recognition device according to some embodiments of the present invention.
Figure 2 This figure shows a schematic workflow of a facial recognition system when the ambient light around the stereo camera is above a pre-defined threshold.
Figures 3a to 3c These figures show examples of dual-regional adjustments according to embodiments of the invention, when (a) there is a strong ambient light directly towards the facial recognition device’s screen; (b) there is a strong ambient light directly towards the person’s face; and (c) there is a strong ambient light towards a side of the face.
Figure 4 This figure shows a schematic workflow of a facial recognition system when the ambient light around the stereo camera is below a pre-defined threshold. DETAILED DESCRIPTION OF THE INVENTION
As it will be appreciated by one skilled in the art, aspects of the present invention may be embodied either as a facial recognition system or a method for alleviating ambient light during image capture.
The present invention relates to products and solutions where the user needs to be authenticated through his or her face. This step could be part of an authorization request where he or she should prove his or her identity when onboarding, enrolling, authenticating to, or requesting a service. This “service” can be access control to a physical or virtual restricted area. If the user is pre-authorized (e.g., already enrolled and whitelisted in the access control list) and his/her identity proven through his/her face, access to the restricted area will be granted. Otherwise, access is rejected.
The identity verification may occur either through a dedicated Software application installed in the face recognition device itself, or in cooperation with another service provider’s mobile application to which the user has logged-in, or through a web browser in the electronic device.
As another implementation example, the present invention may be used in remote identity verification (e.g., Know-Your-Customer, “KYC”, solutions) where two images are captured: an image containing an alleged user’s I Dentity document (e.g., national ID, passport, driving license) as well as the user’s face. Optionally, prior art liveness detection techniques can be applied to the latter image. Facial biometrics allows recognizing and measuring facial features from a captured image of a face (e.g., from a selfie). Liveness detection further ensures that a real person was physically present during the capture and detects any spoof attack. Then, the authenticity of the captured document is checked to prevent fraud by verifying data integrity, consistency and security elements; while face matching assesses the likelihood that the captured person is the same person as the one on the ID document. If user’s identity is verified, he/she is enrolled or his/her access granted to the requested service.
In figure 1 , it is schematically represented a face recognition system (1) configured to alleviate ambient light during the capture of a face image according to the invention. In other words, the face recognition system (1) is configured to verify identities based on photo(s) of the face (3) of a user. The face recognition system (1) of figure 1 is exemplified and represented throughout the figures as a face recognition device (10) in the form of a tablet. The tablet (10) includes one or several (micro)processors (and/or a (micro)controller(s)), as data processing means (not shown), comprising and/or being connected to one or several memories, as data storing means, comprising or being connected to means for interfacing with the user, such as a Man Machine Interface (or MMI), and comprising or being connected to an Input/Output (or I/O) interface(s) that are internally all connected, through an internal bidirectional data bus.
The I/O interface(s) may include a wired and/or a wireless interface, to exchange, over a contact and/or Contactless (or CTL) link(s), with a user. The MMI may include a display screen(s), a keyboard(s), a loudspeaker(s) and/or a camera(s) and allows the user to interact with the tablet. The MMI may be used for getting data entered and/or provided by the user. In particular, the MMI comprises a stereo camera (2) configured to capture images of the user’s face (3). These images (3) comprise depth information (3.1) of objects within the depth of field.
The tablet memory(ies) may include one or several volatile memories and/or one or several non-volatile memories (not shown). The memory(ies) may store data, such as an ID(s) relating to the tablet or hosted chip, that allows identifying uniquely and addressing the tablet. The memory(ies) also stores the Operating System (OS) and applications which are adapted to provide services to the user (not shown). Among these applications, the tablet has installed an orchestrator able to receive ID verification requests from service provider apps or web browsers and e.g., routing them towards backend services hosted in server(s)
(11). That is, the tablet (10) is configured to send information over a communication network
(12) to the server end (11). In particular, the orchestrator and the backend infrastructure may communicate via one or more Application Program Interfaces, APIs, using HTTP over TLS.
As mentioned, the tablet (10) can verify the identity of the user by itself or in cooperation with remote server(s) (11). Therefore, in a basic configuration, the tablet (10) only comprises the stereo camera (2) configured to capture images of the user’s face (3) and then send them to server(s) end (11) for further processing: face detection, image processing, dual- regional exposure adjustment, face or biometrics matching, and ID verification. In this basic embodiment, the tablet typically comprises an application or orchestrator configured, on one side, to receive requests from other service provider applications installed in the device, or from web browsers and, on the other side, communicate with backend services running in the server(s) (11) side. These one or more servers may comprise (jointly or separately) others modules which allows performing the steps which are being described hereinafter in relation with the tablet.
The tablet (or the server) then comprises, a face detection module (4) configured to identify a face based on the depth information (3.1) of the received image (3) from the stereo camera (2), a processing module (5) configured to split the image (3) into a facial and non-facial areas and to calculate exposure adjustment factors for each area based at least on a brightness values comparison, a dual-regional adjustment module (6) configured to separately adjust the exposure of both the facial and non-facial areas according to the adjustment factors, and a face matching module (7) configure to verify or enroll person’s face identity based on the exposure-adjusted facial area.
In the tablet (9) of figure 1 , the dual-regional adjustment module (6) is a built-in feature of the stereo camera (2) so that the exposure adjustment factors calculated by the processing module (5) are sent back to the stereo camera to apply the exposure adjustment accordingly.
The tablet (9) further comprises a screen (8) to display a re-constructed image (3’) with the separately exposure-adjusted facial and non-facial areas made by the dual-regional adjustment module. This screen prompts hereby the user with a more homorganic representation than the originally un-adjusted captured image. Additional information can be conveyed to the user together with the homorganic image such as information about the success of the face matching and ID verification, or enrolment, and the % certitude. For instance, it can additionally prompt the message “Welcome to enterprise X, Mr. name”, “Have a nice trip to destiny”, etc.
As depicted in figure 1 , the tablet (9) further comprises a built-in supplementary lighting (9) to be activated in case the face detection module identifies a face in any of the images coming from the stereo camera and the ambient light is below a pre-defined threshold. To recognize whether the ambient light is above or below the threshold, light sensors (not shown) may be used.
Figure 2 depicts a schematic workflow (20) of a facial recognition system when the ambient light around the stereo camera is above a pre-defined threshold. That is, for instance, image is captured under day-light or in a properly illuminated room.
The steps can be taken e.g., by the facial recognition system (1) described in relation to figure 1. First, the stereo camera (2) captures (21) an image (3) comprising depth information of objects within the depth of field, and send it (3) to a face detection module (4). The face detection module identifies (22) whether there is a face in the image thanks to the depth information embedded in the received image (3). If no face is detected, the process is aborted. This could be part of a loop analyzing frames from a video stream.
Otherwise, i.e. , if a face is found in the captured image, the image (3) and the region (3.4) of the image where a face is identified are further processed. Namely, the face detection module (4) is configured to detect objects (so called, Regions of Interest or “ROI”) on the image, and to infer whether any of the found objects resembles a face thanks to an Machine Leaning, ML, -driven software. ROIs are typically represented with bounding boxes for further class classification (e.g., faces). If more than ones classes are possible, e.g., faces, persons, pets, etc. the detection module can access a database or keyword table (not shown) of pre-determined classes to be possibly found in the captured images. Then, the detection module applies a nearest neighbor search within the reference database using the ROIs, namely an embedding vector for each ROI. From results of the search, it selects the most likely match candidate (if any).
Once the image (3) and the region (3.4) of the image where a face is identified are received by the processing module (5), it splits (24) the image (3) into a facial area (3.2) and a nonfacial area (3.3). The facial area (3.2) corresponds to the identified face region (3.4) by the face detection module. The processing module (5) then outcomes (25) exposure adjustment factors for each of the areas (3.2, 3.3), wherein these exposure adjustment factors are calculated based at least on a brightness values comparison of both facial (3.2) and nonfacial (3.3) areas.
The dual-regional adjustment module (6) then adjusts separately (26) the exposure of both the facial (3.2) and non-facial (3.3) areas according to the received adjustment factors. As noted, the dual-regional adjustment module (6) can be a built-in feature of the stereo camera (2), so the smart exposure adjustment can be performed by the camera itself after receiving instruction about how to adjust the areas.
Through this document, the un-adjusted image (3) and its facial (3.2) and non-facial (3.3) areas are distinguished from the dual-region exposure-adjusted image (or re-constructed image) in that the latter has a “ ‘ ” in its numeric references such as 3’, 3.2’, 3.3’. The bounding box (3.4) bordering the facial region will be in principle the same both pre-adjusted (3, 3.2, 3.3) and post-adjusted (3’, 3.2’, 3.3’) images.
Finally, a face matching module (7) verifies a person’s face identity based on the exposure- adjusted facial area (3.2’). As known, the identification result will be accompanied by a % of certainty. Normally, the identification result not only states the name of the identified preenrolled user but also adjoins his/her name or other identifier.
As mentioned, after separately adjusting the exposure of the two areas, the dual-regional adjustment module (6) can re-construct (28) the image (3’) with the separately exposure- adjusted facial (3.2’) and non-facial (3.3’) areas. Then, this re-constructed image can be sent to the screen (8) to be displayed thereon. Along with this re-constructed image (3’), the screen III can prompt the user with any of the following messages: the identification result, % certainty, or any customized welcome, informational or exception message.
Figures 3a to 3c depict examples of dual-regional adjustments according to embodiments of the invention. For each one, it is shown on the left-hand side the original image (3) captured and, on the right-hand side, the re-constructed (and more homorganic) image (3’).
Figure 3a depicts the situation when there is a strong ambient light directly towards the facial recognition device’s screen. As it can be seen on the left-hand side, it produces a dark facial area (3.2) and an over-exposed (i.e., too bright) non-facial area (3.3).
In this scenario, the light interference is innocuous for the depth perception feature of stereo camera and the face detection module can perform a first-step face detection. In each picture it is depicted an example of a face detection box resulting from the face detection module (4). As known, the box size can be adjusted.
The brightness values comparison of both facial (3.2) and non-facial areas (3.3) outcomes exposure adjustment factors that requires to dedicatedly compensate the exposure of the facial area (3.2) and decrease the exposure of the non-facial area (3.3), accordingly. The resulting homorganic image (3’) can be seen on the right-hand side.
Figure 3b depicts the situation when there is there is a strong ambient light directly towards the person’s face. As it can be seen on the left-hand side, it produces an over-exposed facial area (3.2) and a relatively lesser-exposed / bright non-facial area (3.3).
The brightness values comparison of both facial (3.2) and non-facial areas (3.3) outcomes exposure adjustment factors that requires to dedicatedly lower the exposure of the facial area (3.2) while compensating the exposure of the non-facial area (3.3), accordingly.
Figure 3c depicts the situation when there is there is a strong ambient light towards a side of the face. As it can be seen on the left-hand side, it produces a half bright I half dark facial area (3.2) and a half bright I half dark non-facial area (3.3).
The brightness values comparison of both facial (3.2) and non-facial areas (3.3) outcomes exposure adjustment factors that requires to dedicatedly compensate the dark part of the facial area (3.2) with a higher exposure and decrease the over-exposure I bright part of the facial area (3.2) with lower exposure, while compensate the dark part of the non-facial area (3.3) and decrease the over-exposure I bright part of the non-facial area (3.3), accordingly.
Figure 4 depicts a schematic workflow (20) of a facial recognition system when the ambient light around the stereo camera is below a pre-defined threshold. That is, it is night-time, dark or there is insufficient light so as to properly capture visible face features.
As it can be easily derived, the method of figure 4 is similar to the one of figure 2 but adds the turning-on of supplementary lights (9) if the ambient light around the stereo camera (2) is insufficient (i.e., below a pre-defined threshold). The lights will turn on in case the face detection module detects a face within the depth of field of the stereo camera.
To this end, there could be an on-demand capture and image analysis made by a user e.g., when switching on or awakening a sleeping face recognition system. Alternatively, the face detection module can be continuously analyzing a video stream to find faces and, if found, hence triggering the dual-regional adjustment. The system can similarly analyse images captured by a stereo camera which are stored either in a database of photos or a collection of videos.
The supplementary lights can be switch off according to different use cases. For instance, for single captures it can simply flash, while for in-live video stream analysis it can be active for more time.

Claims

1.- Facial recognition system (1) for alleviating ambient light during image (3) capture, the system comprising: a stereo camera (2) configured to capture an image (3) and send it to a face detection module, wherein the image comprises depth information (3.1) of objects within the depth of field, the face detection module (4) configured to identify a face based on the depth information (3.1) of the received image (3) from the stereo camera (2), an processing module (5) configured to o split the image (3) into a facial area (3.2) and a non-facial area (3.3), wherein the facial area (3.2) corresponds to the identified face region by the face detection module (4), and o outcome exposure adjustment factors for each area (3.2, 3.3), these exposure adjustment factors being calculated based at least on a brightness values comparison of both facial (3.2) and non-facial areas (3.3), a dual-regional adjustment module (6) configured to separately adjust the exposure of both the facial (3.2) and non-facial areas (3.3) according to the adjustment factors, and a face matching module (7) configure to verify or enroll person’s face identity based on the exposure-adjusted facial area (3.2’).
2.- The system according to claim 1 , wherein the these exposure adjustment factors are calculated based on brightness and exposure values comparison of both facial (3.2) and non-facial (3.3) areas.
3.- The system according to claim 2, wherein the processing module (5) is further configured to calculate the brightness and exposure values of each pixel in the facial (3.2) and non- facial (3.3) areas, calculate representative brightness and exposure values for each area (3.2, 3.3), and compare them to outcome compensating exposure adjustment factors to be applied separately to each area (3.2, 3.3) by the dual-regional adjustment module.
4.- The system according to any of claims 1 to 3, further comprising a screen (8), wherein the dual-regional adjustment module (6) is configured to re-construct an image (3’) with the separately exposure-adjusted facial (3.2’) and non-facial (3.3’) areas to be displayed on the screen (8).
5.- The system according to any of claims 1 to 4, further comprising supplementary lighting (9) to be activated in case the face detection module (4) identifies a face in any of the images (3) coming from the stereo camera (2) and the ambient light is below a pre-defined threshold.
6.- The system according to any of claims 1 to 5, wherein the dual-regional adjustment module (6) is a built-in feature of the stereo camera (2) so that the exposure adjustment factors calculated by the processing module (5) are sent back to the stereo camera (2) to apply the exposure adjustment accordingly.
7.- The system according to any of claims 1 to 6, further comprising a facial recognition device (10), such as a tablet or pod, wherein the stereo camera (2) is operatively connected to the face recognition device, either embedded as an integrative solution, or integrated dedicatedly as a plug-in solution.
8.- The system according to any of claims 1 to 7, being a distributed system where at least one of the face detection module (4), the processing module (5), the dual-regional adjustment module (6), or the face matching module (7) is running at server side (11) as a backend service.
9.- A method (20) for alleviating ambient light during image capturing using the facial recognition system (1) according to any of claims 1 to 8, the method comprising the following steps: capturing (21), by the stereo camera (2), an image (3) comprising depth information of objects within the depth of field, sending, by the stereo camera, the captured image (3) to a face detection module (4), identifying (22), by the face detection module, a face based on the depth information of the received image from the stereo camera, receiving, by the processing module (5), the image (3) and the region of the image where a face is identified, splitting (24), by the processing module (5), the image (3) into a facial area (3.2) and a non-facial area (3.3), wherein the facial area (3.2) corresponds to the identified face region by the face detection module, outcome (25), by the processing module (5), exposure adjustment factors for each area (3.2, 3.3), wherein these exposure adjustment factors are calculated based at least on a brightness values comparison of both facial (3.2) and non-facial (3.3) areas, adjusting separately (26), by a dual-regional adjustment module (6), the exposure of both the facial (3.2) and non-facial (3.3) areas according to the adjustment factors, and verifying or enrolling (27), by a face matching module (7), a person’s face identity based on the exposure-adjusted facial area (3.2’).
10.- The method according to claim 9, wherein the processing module (5) calculates the brightness and exposure values of each pixel in the facial (3.2) and non-facial (3.3) areas, calculate representative brightness and exposure values for each area, and compare them to outcome compensating exposure adjustment factors to be applied separately to each area by the dual-regional adjustment module (6).
11.- The method according to any of claims 9 or 10, further comprising the steps: re-constructing (28), by the dual-regional adjustment module (6), a re-constructed image (3’) with the separately exposure-adjusted facial (3.2’) and non-facial (3.3’) areas, and sending the re-constructed image (3’) to the screen (8) to be displayed thereon.
12.- The system according to any of claims 9 to 11 , further comprising the steps: detecting, preferably by a light sensor, if the ambient light around the stereo camera (2) is below a pre-defined threshold, and activating (29), by the face detection module, the supplementary lighting (9).
EP24705650.0A 2023-02-23 2024-02-15 FACE RECOGNITION SYSTEM TO REDUCE AMBIENT LIGHT DISTURBANCE DURING IMAGE CAPTURE Pending EP4670132A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310161642.9A CN118540591A (en) 2023-02-23 2023-02-23 Facial recognition system for mitigating ambient light during image capture
PCT/EP2024/053810 WO2024175453A1 (en) 2023-02-23 2024-02-15 Facial recognition system for alleviating ambient light during image capture

Publications (1)

Publication Number Publication Date
EP4670132A1 true EP4670132A1 (en) 2025-12-31

Family

ID=89977926

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24705650.0A Pending EP4670132A1 (en) 2023-02-23 2024-02-15 FACE RECOGNITION SYSTEM TO REDUCE AMBIENT LIGHT DISTURBANCE DURING IMAGE CAPTURE

Country Status (3)

Country Link
EP (1) EP4670132A1 (en)
CN (1) CN118540591A (en)
WO (1) WO2024175453A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111242086A (en) * 2020-01-21 2020-06-05 成都国翼电子技术有限公司 Image exposure adjusting method based on face recognition
CN114627524B (en) * 2021-11-16 2025-06-24 浙江光珀智能科技有限公司 A method for automatic face exposure based on depth camera

Also Published As

Publication number Publication date
WO2024175453A1 (en) 2024-08-29
CN118540591A (en) 2024-08-23

Similar Documents

Publication Publication Date Title
US10762334B2 (en) System and method for entity recognition
US12272173B2 (en) Method and apparatus with liveness test and/or biometric authentication
US11227170B2 (en) Collation device and collation method
US11989975B2 (en) Iris authentication device, iris authentication method, and recording medium
JP2020523665A (en) Biological detection method and device, electronic device, and storage medium
US11709919B2 (en) Method and apparatus of active identity verification based on gaze path analysis
US12361594B2 (en) Image processing method and apparatus, computer device and storage medium
US20170308763A1 (en) Multi-modality biometric identification
US10970953B2 (en) Face authentication based smart access control system
JP2020518879A (en) Detection system, detection device and method thereof
US9349071B2 (en) Device for detecting pupil taking account of illuminance and method thereof
US11721132B1 (en) System and method for generating region of interests for palm liveness detection
WO2022222957A1 (en) Method and system for identifying target
WO2024175453A1 (en) Facial recognition system for alleviating ambient light during image capture
EP4336469A1 (en) Method for determining the quality of a captured image
US11688204B1 (en) System and method for robust palm liveness detection using variations of images
JP2026505887A (en) Shooting control method and related device
US11941911B2 (en) System and method for detecting liveness of biometric information
RU2798179C1 (en) Method, terminal and system for biometric identification
JP7272418B2 (en) Spoofing detection device, spoofing detection method, and program
KR20250001301A (en) Method to authorize user, computer deviec performing the same, and method of determining spoofing
HK40101730A (en) Moiré pattern detection in digital images and a liveness detection system thereof
JP2022028850A (en) Spoofing detection device, spoofing detection method, and program

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250923

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR