WO2021235641A1 - 나이 추정 장치 및 나이를 추정하는 방법 - Google Patents

나이 추정 장치 및 나이를 추정하는 방법 Download PDF

Info

Publication number
WO2021235641A1
WO2021235641A1 PCT/KR2020/019328 KR2020019328W WO2021235641A1 WO 2021235641 A1 WO2021235641 A1 WO 2021235641A1 KR 2020019328 W KR2020019328 W KR 2020019328W WO 2021235641 A1 WO2021235641 A1 WO 2021235641A1
Authority
WO
WIPO (PCT)
Prior art keywords
age
face
facial feature
user
facial
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2020/019328
Other languages
English (en)
French (fr)
Inventor
임상섭
강명주
서현
조현수
안성권
최명제
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SNU R&DB Foundation
LG H&H Co Ltd
Original Assignee
LG Household and Health Care Ltd
Seoul National University R&DB Foundation
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by LG Household and Health Care Ltd, Seoul National University R&DB Foundation filed Critical LG Household and Health Care Ltd
Publication of WO2021235641A1 publication Critical patent/WO2021235641A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/178Human faces, e.g. facial parts, sketches or expressions estimating age from face image; using age information for improving recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F21/00Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
    • G06F21/30Authentication, i.e. establishing the identity or authorisation of security principals
    • G06F21/31User authentication
    • G06F21/32User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q30/00Commerce
    • G06Q30/02Marketing; Price estimation or determination; Fundraising
    • G06Q30/0282Rating or review of business operators or products
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q30/00Commerce
    • G06Q30/06Buying, selling or leasing transactions
    • G06Q30/0601Electronic shopping [e-shopping]
    • G06Q30/0631Recommending goods or services
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q50/00Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
    • G06Q50/10Services
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation

Definitions

  • the present disclosure relates to an age estimation apparatus and method for estimating a user's age based on a face image.
  • image processing technology or image recognition technology based on artificial intelligence includes a technology for recognizing the type of an object included in an image, a technology for specifically identifying an object included in an image, and the like.
  • This image recognition technology is also used to provide a face recognition security function that identifies a user's face.
  • the technology for recognizing a user's face can only distinguish several users from each other, and does not provide information about the identified user.
  • the age of the user can be estimated based on the face image, the age of the user can be easily determined, and various functions can be provided based on the estimated age.
  • An object of the present disclosure is to provide an age estimation apparatus and method for estimating a user's age based on a face image.
  • An apparatus for estimating age includes: a memory for storing a facial feature point extraction model and an age estimation model; and receiving at least one facial image, extracting a plurality of facial feature points from the facial image using a facial feature point extraction model, extracting a plurality of predetermined facial indicators from the plurality of facial feature points, and an age estimation model and a plurality of faces
  • a processor for estimating the age of the user using the index may be included.
  • the processor may normalize a plurality of facial feature points and extract a plurality of facial indicators from the plurality of standardized facial feature points.
  • the processor may determine a transformation rule of the face image such that a predetermined part of the face is arranged at predetermined coordinates based on at least a part of the plurality of facial feature points, and may normalize the plurality of facial feature points based on the transformation rule.
  • the facial feature point extraction model is a deep learning model composed of an artificial neural network, and when a facial image is input, a plurality of facial feature points corresponding to the input facial image may be extracted by a predetermined number.
  • the processor When receiving the plurality of face images, the processor extracts a plurality of facial feature points from each of the plurality of face images, calculates an average of facial feature points at positions corresponding to each other, and calculates a plurality of average facial feature points, and the plurality of average facial features
  • a plurality of face indicators may be extracted from the feature points.
  • the processor may determine an outlier face image from among the plurality of face images, and extract a plurality of facial feature points from the face image except for the outlier face image.
  • the plurality of face images may be face images captured within a predetermined period.
  • the age estimation model may be a linear regression model, a decision tree, or a deep learning model, and may be a model that outputs an estimated age when at least a plurality of face indicators are input.
  • the processor estimates the user's age by using at least one of a plurality of facial feature points, a plurality of facial indicators, or the user's body information and an age estimation model, and the body information is the user's height, the user's weight, or the user's body fat index ( BMI) may include at least one or more.
  • the plurality of face indicators may include at least one of a total face area, a lip area, a face-related indicator, a nose-related indicator, or a lip-related indicator, and the face-related indicator may include a slope of a face outline.
  • the age estimation apparatus may further include a camera, and the processor may receive at least one face image through the camera.
  • the age estimation apparatus may further include a communication unit communicating with the user terminal, and the processor may receive at least one face image from the user terminal through the communication unit.
  • the processor may determine that the estimated age is suitable for the age authentication if the estimated age is greater than or equal to the age required for the authentication by a predetermined margin.
  • a method of estimating an age includes receiving at least one or more face images; extracting a plurality of facial feature points from a facial image using a facial feature point extraction model; extracting a plurality of predetermined facial indicators from a plurality of facial feature points; and estimating the age of the user by using the age estimation model and a plurality of face indicators.
  • the age of the user may be estimated simply by photographing the user's face image.
  • an age-based service may be provided simply by photographing a user's face image.
  • FIG. 1 is a block diagram illustrating an apparatus for estimating an age according to an embodiment of the present disclosure.
  • FIG. 2 is a block diagram illustrating an artificial intelligence server according to an embodiment of the present disclosure.
  • FIG. 3 is a diagram illustrating an artificial intelligence system according to an embodiment of the present disclosure.
  • FIG. 4 is an operation flowchart illustrating a method of estimating a user's age based on a face image according to an embodiment of the present disclosure.
  • FIG. 5 is a diagram illustrating a method of photographing a user's face according to an embodiment of the present disclosure.
  • FIG. 6 is a diagram illustrating examples of a plurality of face images.
  • FIG. 7 is a diagram illustrating an example of a plurality of facial features extracted from a plurality of facial images shown in FIG. 6 .
  • FIG. 8 is a diagram illustrating an example of face indicators according to an embodiment of the present disclosure.
  • 9 to 13 are diagrams illustrating examples of facial indicators extracted from a plurality of facial feature points illustrated in FIG. 7 .
  • FIG. 14 is a diagram illustrating an example of facial feature points extracted from a facial image according to an embodiment of the present disclosure.
  • 15 and 16 are diagrams illustrating examples of facial indicators extracted from a plurality of facial feature points shown in FIG. 14 .
  • 17 is a diagram illustrating upper face indexes having high age estimation ability for an age estimation model based on linear regression according to an embodiment of the present disclosure.
  • FIG. 18 is a view showing the distribution of the jaw edge angle (LD4) for each age group.
  • LD4 chin edge angle
  • 20 is a diagram illustrating an age estimation model according to an embodiment of the present disclosure.
  • 21 is a diagram illustrating an age estimation model based on deep learning according to an embodiment of the present disclosure.
  • 22 to 25 are diagrams illustrating structures of learning data used for learning an age estimation model according to an embodiment of the present disclosure.
  • 26 is a diagram illustrating an embodiment of estimating a user's age.
  • 27 is a diagram illustrating an embodiment of performing age authentication by estimating a user's age.
  • 28 is a diagram illustrating an embodiment of performing age authentication by estimating a user's age.
  • 29 is a diagram illustrating an embodiment of estimating the age of a user included in an image.
  • 30 and 31 are diagrams illustrating embodiments in which an age-customized operation is performed by estimating a user's age.
  • 32 is a diagram illustrating an embodiment of estimating a user's age using a virtual face image.
  • a component When it is said that a component is 'connected' or 'connected' to another component, it is understood that it may be directly connected or connected to the other component, but other components may exist in between. It should be On the other hand, when it is mentioned that a certain element is 'directly connected' or 'directly connected' to another element, it should be understood that the other element does not exist in the middle.
  • FIG. 1 is a block diagram illustrating an age estimation apparatus 100 according to an embodiment of the present disclosure.
  • the age estimation device 100 includes a TV, a projector, a mobile phone, a smart phone, a desktop computer, a notebook computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a tablet PC, a wearable device, a set-top box ( STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, a digital signage, a robot, a vehicle, etc., may be implemented as a stationary device or a mobile device.
  • the age estimation apparatus 100 includes a communication unit 110 , an input unit 120 , a learning processor 130 , a sensing unit 140 , an output unit 150 , a memory 170 and a processor 180 . and the like.
  • the communication unit 110 may be referred to as a communication modem or a communication circuit.
  • the communication unit 110 may transmit/receive data to and from external devices such as the artificial intelligence server 200 using wired/wireless communication technology.
  • the communication unit 110 may transmit/receive sensor information, a user input, a learning model, a control signal, and the like with external devices.
  • the communication technology used by the communication unit 110 includes GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), There are Bluetooth (Bluetooth), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), and the like.
  • GSM Global System for Mobile communication
  • CDMA Code Division Multi Access
  • LTE Long Term Evolution
  • 5G Fifth Generation
  • WLAN Wireless LAN
  • Wi-Fi Wireless-Fidelity
  • Bluetooth Bluetooth
  • RFID Radio Frequency Identification
  • IrDA Infrared Data Association
  • ZigBee ZigBee
  • NFC Near Field Communication
  • the input unit 120 may be referred to as an input interface.
  • the input unit 120 may acquire various types of data.
  • the voice data or image data collected by the input unit 120 may be analyzed and processed as a user's control command.
  • the input unit 120 includes a camera 121 for receiving an image signal, a microphone 122 for receiving an audio signal, or a user input unit 123 for receiving information from a user. can do.
  • a signal obtained from the camera 121 or the microphone 122 may be referred to as sensor data or sensor information.
  • the camera 121 may process an image frame such as a still image or a moving image obtained by an image sensor in a video call mode or a shooting mode.
  • the camera 121 may be composed of one or a plurality of cameras.
  • the processed image frame may be displayed on the display unit 151 or stored in the memory 170 .
  • the microphone 122 may receive the sound wave and convert it into electrical voice data.
  • the converted voice data may be variously utilized according to a function (or a running application program) being performed by the age estimation apparatus 100 . Meanwhile, various noise removal algorithms for removing noise in the process of receiving an external sound wave may be applied to the microphone 122 .
  • the user input unit 123 is for receiving information from the user. When information is input through the user input unit 123 , the processor 180 may control the operation of the age estimation apparatus 100 to correspond to the input information. have.
  • the user input unit 123 may include a mechanical input means or a touch input means.
  • the input unit 120 may acquire training data for model training and input data to be used when acquiring an output using the training model.
  • the input unit 120 may acquire raw input data, and in this case, the processor 180 or the learning processor 130 may extract an input feature as a preprocessing for the input data.
  • the learning processor 130 may train a model composed of an artificial neural network by using the training data.
  • the learned artificial neural network may be referred to as a learning model.
  • the learning model may be used to infer a result value with respect to new input data other than the training data, and the inferred value may be used as a basis for a decision to perform a certain operation.
  • the learning processor 130 may perform artificial intelligence processing together with the learning processor 240 of the artificial intelligence server 200 .
  • the sensing unit 140 may be referred to as a sensor unit.
  • the sensing unit 140 may acquire at least one of internal information of the age estimating apparatus 100 , surrounding environment information of the age estimating apparatus 100 , and user information by using various sensors.
  • Sensors included in the sensing unit 140 include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, an optical sensor, a microphone, a lidar, and a radar. etc.
  • the output unit 150 may be referred to as an output interface.
  • the output unit 150 includes a display unit 151 for outputting visual information, a sound output unit 152 for outputting auditory information, and a haptic module 153 or optical output unit 152 for outputting tactile information. It may include an output unit (Optical Output Unit, 154) and the like.
  • the display unit 151 may display (output) information processed by the age estimation apparatus 100 .
  • the display unit 151 may display information on an execution screen of an application program driven by the age estimation apparatus 100 , or user interface (UI) and graphic user interface (GUI) information according to the information on the execution screen.
  • UI user interface
  • GUI graphic user interface
  • the display unit 151 may be implemented as a touch screen by forming a layer structure with the touch sensor or being formed integrally with the touch sensor. Such a touch screen may simultaneously provide an input interface and an output interface between the age estimation apparatus 100 and the user.
  • the sound output unit 152 may include at least one of a receiver, a speaker, and a buzzer to output audio data as sound waves.
  • the haptic module 153 may generate various tactile effects that the user can feel.
  • a representative example of the tactile effect generated by the haptic module 153 is vibration.
  • the light output unit 154 outputs a signal for notifying the occurrence of an event by using the light of the light source of the age estimation apparatus 100 .
  • Examples of the event generated by the age estimation apparatus 100 may be message reception, call signal reception, missed call, alarm, schedule notification, email reception, information reception through an application, and the like.
  • the memory 170 may store data supporting various functions of the age estimation apparatus 100 .
  • the memory 170 may store input data obtained from the input unit 120 , learning data, a learning model, a learning history, an application, and the like.
  • the memory 170 may refer to various memories that permanently, semi-permanently, or temporarily store data.
  • the memory 170 may include a hard disk drive (HDD), a solid state drive (SSD), a compact disk (CD), a random access memory (RAM), a read only memory (ROM), and the like.
  • HDD hard disk drive
  • SSD solid state drive
  • CD compact disk
  • RAM random access memory
  • ROM read only memory
  • the processor 180 may determine at least one executable operation of the age estimation apparatus 100 based on information determined or generated using a data analysis algorithm or a machine learning algorithm. In addition, the processor 180 may control the components of the age estimation apparatus 100 to perform the determined operation. To this end, the processor 180 may request, search, receive, or utilize the data of the learning processor 130 or the memory 170, and may perform a predicted operation or an operation determined to be desirable among the at least one executable operation. The components of the age estimating apparatus 100 may be controlled to be executed.
  • the processor 180 may generate a control signal for controlling the external device when the connection of the external device is required to perform the determined operation, and transmit the generated control signal to the corresponding external device through the communication unit 110 . .
  • the processor 180 may obtain intention information with respect to a user input and determine a user's requirement based on the obtained intention information.
  • the processor 180 uses at least one of a speech to text (STT) engine for converting a voice input into a character string or a natural language processing (NLP) engine for obtaining intention information of a natural language. Intention information corresponding to the input may be obtained.
  • STT speech to text
  • NLP natural language processing
  • At least one of the STT engine and the NLP engine may be configured as an artificial neural network, at least a part of which is learned according to a machine learning algorithm.
  • at least one or more of the STT engine or the NLP engine is learned by the learning processor 130 , or learned by the learning processor 240 of the artificial intelligence server 200 , or learned by their distributed processing. it may have been
  • the processor 180 collects history information including the user's feedback on the operation contents or the operation of the age estimation apparatus 100 and stores it in the memory 170 or the learning processor 130, or the artificial intelligence server 200 It can be transmitted to an external device such as The collected historical information may be used to update the learning model.
  • the processor 180 may control at least some of the components of the age estimation apparatus 100 to drive an application program stored in the memory 170 . Furthermore, in order to drive the application program, the processor 180 may operate two or more of the components included in the age estimation apparatus 100 in combination with each other.
  • FIG. 2 is a block diagram illustrating an artificial intelligence server 200 according to an embodiment of the present disclosure.
  • the artificial intelligence server 200 may refer to a device that trains an artificial intelligence model including an artificial neural network using a machine learning algorithm or uses the learned artificial intelligence model.
  • the artificial intelligence server 200 may be configured with a plurality of servers to perform distributed processing. Also, the artificial intelligence server 200 may distribute and process artificial intelligence processing together with the age estimation apparatus 100 .
  • the artificial intelligence server 200 may include a communication unit 210 , a memory 230 , a learning processor 240 , and a processor 260 .
  • the communication unit 210 may transmit/receive data to and from an external device such as the age estimation apparatus 100 .
  • the memory 230 may include a model storage unit 231 .
  • the model storage unit 231 may store a model (or artificial neural network, 231a) being trained or learned through the learning processor 240 .
  • the learning processor 240 may train the artificial neural network 231a using the training data.
  • the learning model may be used while being mounted on the artificial intelligence server 200 of the artificial neural network, or may be used while being mounted on an external device such as the age estimation apparatus 100 .
  • the learning model may be implemented in hardware, software, or a combination of hardware and software.
  • one or more instructions constituting the learning model may be stored in the memory 230 .
  • the processor 260 may infer a result value with respect to new input data using the learning model, and may generate a response or a control command based on the inferred result value.
  • FIG. 3 is a diagram illustrating an artificial intelligence system 1 according to an embodiment of the present disclosure.
  • the artificial intelligence system 1 may include an age estimation apparatus 100 and an artificial intelligence server 200 or a user terminal 300 .
  • the age estimation apparatus 100 may be connected to the artificial intelligence server 200 or the user terminal 300 through a wired network or a wireless network.
  • the artificial intelligence server 200 may be connected to a plurality of age estimating devices 100 , and may be performed instead of artificial intelligence processing of each age estimating device 100 , or distributedly.
  • the artificial intelligence server 200 may train an artificial neural network using a machine learning algorithm or a deep learning algorithm instead of the age estimation apparatus 100, and may directly store the learning model or transmit it to the age estimation apparatus 100 .
  • the artificial intelligence server 200 receives input data from the age estimation device 100, infers a result value corresponding to the received input data using a learning model, and generates a response or control command based on the inferred result value Thus, it may be transmitted to the age estimation apparatus 100 .
  • the age estimation apparatus 100 may infer a result value with respect to input data using a direct learning model, and may generate a response or a control command based on the inferred result value.
  • the user terminal 300 may be an image photographing device capable of directly photographing an image.
  • the user terminal 300 is connected to the age estimating apparatus 100 , and may receive input data or a user input and transmit it to the age estimating apparatus 100 .
  • the age estimation apparatus 100 infers a result value corresponding to the input data received from the user terminal 300 , generates a response or a control command based on the inferred result value, and transmits it to the user terminal 300 . have.
  • FIG. 4 is an operation flowchart illustrating a method of estimating a user's age based on a face image according to an embodiment of the present disclosure.
  • the processor 180 of the age estimation apparatus 100 receives a face image of a user ( S401 ).
  • the user's face image may mean image data including the user's face.
  • the processor 180 may receive the user's face image by photographing an image including the user's face through the camera 121 , or may receive the user's face image from the user terminal 300 through the communication unit 110 .
  • the user terminal 300 may be a professional device for photographing a face image, and may photograph a visible light image, a polarized image, a UV image, etc. in a state in which external light is blocked with a blackout screen.
  • the user terminal 300 is a terminal including a camera, such as a smartphone, and may take a visible light image.
  • the user's face image may be image data in various formats, such as an RGB image, a polarized image, a black-and-white image, a UV image, an IR image, or an RGB-IR image. Also, the user's face image may mean an image frame constituting a still image or a moving picture.
  • the processor 180 may receive a plurality of face images captured within a predetermined period.
  • the processor 180 may receive a plurality of face images of the same kind, a plurality of different kinds of face images, or a plurality of face images in which at least some are different kinds, photographed within a predetermined period.
  • the same type of face image may mean face images having the same type of light used for photographing the face image
  • the heterogeneous face image may mean face images with different types of light used for photographing the face image. have.
  • the user's face image may mean not only an image including an unmakeup face after washing face, but also an image including a makeup face.
  • the processor 180 may additionally receive the user's body information. In this case, the processor 180 may receive the user's body information through the user input unit 123 or the communication unit 110 .
  • the processor 180 of the age estimation apparatus 100 extracts a plurality of facial features from the face image (S403).
  • a facial feature may mean a key point or a facial landmark extracted from a face.
  • a key point or landmark may mean a point that can be easily identified even if the shape, size, or position of an object changes, and may be extracted according to a predetermined pattern or rule.
  • the processor 180 may extract a plurality of facial feature points from the face image by using the facial feature point extraction model.
  • the processor 180 uses the facial landmark extraction model of the Dlib library, the facial landmark extraction model consisting of HR-Net (High-Resolution Network), or the facial landmark extraction model consisting of SAN (Style Aggregated Network).
  • a plurality of facial feature points (eg, 68 facial feature points) may be extracted from the image.
  • the processor 180 uses a deep learning-based facial feature point extraction model composed of an artificial neural network including a pooling layer, a convolution layer, a fully-connected layer, and the like. It is also possible to extract a plurality of facial feature points from the image.
  • the processor 180 may extract a plurality of facial feature points from the face image using a machine learning model-based facial feature point extraction model such as a support vector machine or an ensemble regression tree. .
  • the facial feature point extraction model may be learned by at least one of the running processor 160 of the age estimation apparatus 100 or the running processor 240 of the artificial intelligence server 200 .
  • the facial feature point extraction model may be stored in the memory 170 or the memory 230 of the artificial intelligence server 200 .
  • the processor 180 transmits the face image to the artificial intelligence server 200 through the communication unit 110, and the artificial intelligence server 200
  • the processor 260 of the processor 260 extracts a plurality of facial feature points from the face image using the facial feature point extraction model stored in the memory 230 , and the processor 180 of the age estimation apparatus 100 performs artificial intelligence through the communication unit 110 .
  • a plurality of facial feature points extracted from the server 200 may be received.
  • the extracted facial feature points may be expressed as coordinates of each feature point.
  • the facial feature point extraction model may be trained using first learning data including a facial image and a plurality of facial feature points (or facial landmarks) corresponding thereto.
  • first learning data including a facial image and a plurality of facial feature points (or facial landmarks) corresponding thereto.
  • the iBUG 300-W data set, the Annotated Facial Landmarks in the Wild (AFLW) data set, or the Wide Facial Landmarks in the Wild (WFLW) data set may be used as the first training data used for training the facial feature point extraction model.
  • the processor 180 may detect a rectangular face region including the face from the face image and extract a plurality of facial feature points from the face region. In this process, the processor 180 may crop only the face region from the face image, and extract a plurality of facial feature points from the cropped face region.
  • the processor 180 of the age estimation apparatus 100 normalizes a plurality of facial feature points (S405).
  • the processor 180 may standardize or correct the plurality of facial feature points according to a predetermined criterion.
  • the facial feature point may mean a standardized facial feature point.
  • the processor 180 may determine a transformation rule of the face image so that predetermined parts of the face are arranged at predetermined coordinates based on at least some of the plurality of facial feature points, and standardize or standardize each facial feature point based on the determined transformation rule. can be corrected. For example, the processor 180 determines the positions of the eyes, nose, and mouth based on a plurality of facial feature points, and determines an enlargement/reduction magnification of the facial image, a translation distance, It is possible to determine transformation rules such as rotation angles. In addition, the processor 180 may standardize or correct each facial feature point based on a conversion rule such as an enlargement/reduction magnification, a parallel movement distance, and a rotation angle.
  • a conversion rule such as an enlargement/reduction magnification, a parallel movement distance, and a rotation angle.
  • the processor 180 may determine a transformation rule of the face image so that the forehead rest is disposed at predetermined coordinates, and standardize or correct the face image based on the determined transformation rule. have. If the face image includes a plurality of forehead supports, the processor 180 may determine a conversion rule including an enlargement/reduction magnification of the face image in consideration of distances between the plurality of forehead supports.
  • facial feature points corresponding to each other are arranged at the same or adjacent positions, which is suitable for comparison.
  • the processor 180 of the age estimation apparatus 100 extracts a plurality of predetermined facial indicators from the plurality of standardized facial feature points (S407).
  • Each facial feature point has a small meaning in absolute coordinates, and a relationship (relative coordinate relationship) between two or more facial feature points has a large meaning. That is, it can be seen that the distance or direction (or angle) between two or more facial feature points provides meaningful information on a person's face. Accordingly, the processor 180 may extract a plurality of facial indicators from the plurality of facial feature points based on a predetermined method.
  • the facial index may include a distance between predetermined facial feature points, a direction (or an angle) between predetermined facial feature points, a width of a face, a height of a face, a width of a face, an angle of an outline of the face, and the like.
  • the first facial indicator may be a distance between the first facial feature point and the second facial feature point
  • the second facial indicator may be a direction (or an angle) between the first facial feature point and the second facial feature point.
  • the processor 180 determines an outlier face image with a large error among a plurality of face images by comparing standardized facial feature points extracted from a plurality of face images, and is extracted from face images except for the outlier face image.
  • a plurality of facial indicators may be extracted using standardized facial feature points.
  • the outlier face image may mean a face image in which a distribution of standardized facial feature points deviates from a standard deviation of a predetermined multiple from an average.
  • the outlier face image may mean a predetermined number or ratio of face images in which the distribution of standardized facial feature points among all face images has a large deviation from the average.
  • the processor 180 may determine, as the outlier face image, one face image in which the distribution of standardized facial feature points has the greatest deviation from the average among the three face images.
  • the processor 180 may calculate a plurality of face indicators based on an average of standardized facial feature points corresponding to face images except for the outlier face image. Specifically, the processor 180 may calculate an average between facial feature points at positions corresponding to each other for face images other than the outlier face image, and calculate a plurality of facial indicators based on the average facial feature points. For example, when both the first face image and the second face image are not outlier face images, the processor 180 calculates the average of the facial feature points located at the nose tip of the first face image and the facial feature points located at the nose tip of the second face image. calculated, and a facial index may be extracted based on the average of facial feature points located at the tip of the nose.
  • the processor 180 may extract at least some of the facial indicators to be described later based on the standardized facial feature points.
  • the processor 180 of the age estimation apparatus 100 estimates the age of the user by using the age estimation model and a plurality of face indices ( S409 ).
  • the age estimation model may estimate the age of the user from the plurality of input facial indices. Since the age is estimated using the face index extracted from the user's face, the estimated age may mean the face age.
  • the processor 180 may estimate the age of the user by additionally using body information or extracted facial feature points.
  • the age estimation model may output the age of the user inferred from the input information.
  • the age estimation model may be implemented as a linear regression model, a decision tree, or a deep learning model composed of an artificial neural network.
  • the age estimation model may be configured as a fully connected network.
  • the processor 180 may use an age estimation model corresponding to data (eg, facial indicators, body information, facial feature points, etc.) to be used for estimating the age of the user, and vice versa, corresponding to the age estimation model to be used for estimating the age of the user. data can be used.
  • the age estimation model may be learned by at least one of the learning processor 160 of the age estimation apparatus 100 or the learning processor 240 of the artificial intelligence server 200 .
  • the age estimation model may be stored in the memory 170 or the memory 230 of the artificial intelligence server 200 .
  • the processor 180 transmits a plurality of face indicators to the artificial intelligence server 200 through the communication unit 110, and the artificial intelligence server 200 ) of the processor 260 estimates the user's age from a plurality of face indicators using the age estimation model stored in the memory 230 , and the processor 180 of the age estimation apparatus 100 performs artificial intelligence through the communication unit 110 .
  • the age of the user inferred from the intelligent server 200 may be received.
  • the age estimation model may be trained by using second learning data including a plurality of face indicators and a user's age corresponding thereto.
  • the age estimation model may be a model for estimating the age of a user corresponding to at least one of a face index, a facial feature point, or body information is input.
  • the age estimation model may be learned using second learning data that includes at least one of a plurality of face indicators, a plurality of facial feature points, and body information, and further includes a corresponding user's age.
  • the body information may include the user's height and the user's weight or the user's body mass index (BMI).
  • the order of the steps shown in FIG. 4 is merely an example, and the present disclosure is not limited thereto. That is, in an embodiment, the order of some of the steps shown in FIG. 4 may be reversed to be performed. Also, in an embodiment, some of the steps shown in FIG. 4 may be performed in parallel. Also, only some of the steps shown in FIG. 4 may be performed.
  • FIG. 5 is a diagram illustrating a method of photographing a user's face according to an embodiment of the present disclosure.
  • the user terminal 510 may be a device for generating a face image by photographing the face of the user 520 .
  • the user terminal 510 may include a light source (not shown) and a camera (not shown) for photographing a face.
  • the user terminal 510 may include a chinrest 511 or a headrest 512 that can be fixed by leaning against the user's 520 face.
  • the face of the user 520 is photographed using the user terminal 510 as shown in FIG. 5 , the face of the user 520 is irradiated with controlled light from a fixed position with respect to the camera (not shown). Because it is shooting, it is possible to obtain a stable image with little noise. In addition, there is an advantage that face images can be generated for a plurality of users under the same conditions.
  • FIG. 6 is a diagram illustrating examples of a plurality of face images.
  • the age estimation apparatus 100 may photograph or receive a plurality of face images 610 , 620 , and 630 to estimate the age of one user.
  • the plurality of face images 610 , 620 , and 630 may be captured within a predetermined period.
  • Each face image 610 , 620 , or 630 may be an image taken at a predetermined time interval or may be frame images at a predetermined interval in a video.
  • Each face image 610 , 620 , or 630 may be photographed using the same type of light or may be photographed using different types of light.
  • the first face image 610 may be a visible light image (or RGB image)
  • the second face image 620 may be a polarized image
  • the third face image 630 may be a UV image.
  • Each face image 610 , 620 , or 630 may include a chin rest 511 or a head rest 512 for fixing the user's face when capturing or generating the face image.
  • FIG. 7 is a diagram illustrating an example of a plurality of facial features extracted from a plurality of facial images shown in FIG. 6 .
  • facial feature points 721 , 722 , and 723 corresponding to predetermined positions may be extracted for each of a plurality of facial images 610 , 620 , and 630 .
  • 7 is an image showing facial feature points 721, 722, and 723 extracted from one facial image 710 or an arbitrary facial image among a plurality of facial images 610, 620, and 630 for visualization. In the process of estimating , such a visualized face image may not be generated.
  • a feature point 711 corresponding to the headrest 512 may be extracted, and the age estimating device 100 is the headrest (
  • the extracted facial feature points 721 , 722 , and 723 may be normalized based on the feature point 711 corresponding to 512 .
  • the facial feature points 721 , 722 , and 723 illustrated in FIG. 7 may be standardized facial feature points.
  • Each facial feature point 721 , 722 , and 723 may be extracted corresponding to a predetermined point.
  • the first facial feature point 721 is a facial feature point extracted from the first facial image 610
  • the second facial feature point 722 is a facial feature point extracted from the second facial image 620
  • the third facial feature point 723 may be facial feature points extracted from the third face image 630 .
  • the fourth facial feature point 731 or the average facial feature point 731 is the average position of the first facial feature point 721, the second facial feature point 722 and the third facial feature point 723 at each predetermined point (or center of gravity).
  • the age estimating apparatus 100 may extract facial indices using the average facial feature point 731 , and estimate the age of the user using the extracted facial indices.
  • FIG. 7 only illustrates facial feature points 721 , 722 , and 723 extracted from three facial images 610 , 620 , and 630 as an example, and the present disclosure is not limited thereto. That is, according to an embodiment, facial feature points may be extracted from more or fewer facial images, and an average position (or center of gravity) of the extracted facial feature points may be determined as the average facial feature point.
  • FIG. 8 is a diagram illustrating an example of face indicators according to an embodiment of the present disclosure.
  • a facial indicator that can be utilized in an embodiment of the present disclosure may include at least some of the 49 shown facial indicators.
  • the facial indicators according to an embodiment of the present disclosure may include a total face area, a lip area, a lip-related indicator, a nose-related indicator, a face-related indicator, and the like.
  • 9 to 13 are diagrams illustrating examples of facial indicators extracted from a plurality of facial feature points illustrated in FIG. 7 .
  • the total face area 911 may mean the face area under the eyebrows.
  • the total face area 911 may be extracted by calculating the area of a polygon connecting facial feature points extracted from the eyebrows and facial feature points extracted from the face outline.
  • the lip area 1011 may mean the area of the upper lip and the lower lip.
  • the lip area 1011 may be extracted by calculating the area of a polygon connecting the outer facial feature points among the facial feature points extracted from the upper lip and the lower lip.
  • the face-related indicators include the distance between the eyebrows (length of G1), the distance between the eyes (length of G2), the width of the face at the top of the cheek (length of W1), the width of the face in the middle of the cheek (length of W2), Chin width (length of W3), eye width (length of left/right EW), height from chin to top of eyebrow (length of H1), height from mouth to top of eyebrow (length of H2), height from pharynx to top of eyebrow ( The length of H3) or the inclination of the facial contour (the inverse of the inclination of left/right LD1 to LD4 or the inclination of left/right LD5 to LD8) may be included.
  • the slope of the left contour and the slope of the right contour can be calculated separately. Since it is steep compared to the inclination of the contour lines LD1 to LD4 at the top of the face, a reciprocal value can be used as an index.
  • the lip-related indicators include the height of the central upper lip (length of UH), the height of the upper lip on the side (average of lengths of left/right UH2), and diagonal length of upper lip corners (average of lengths of left/right UH3) , central lower lip height (length of LH), lateral lower lip height (average length of left/right LH2), lower lip lip diagonal length (average of length of left/right LH3), lip width (length of W), Upper lip V-line angle, upper lip crest width or upper lip v-line width (UDH length), upper lip crest-to-trough height or upper lip v-line height (UDV length), lip proportions, etc. can be included. .
  • the lip ratio may mean a ratio of the lip height to the lip width.
  • the lip height used for calculating the lip ratio may be the sum of the height of the upper lip (average of lengths of left/right UH2) and the height of the lower lip (average of lengths of left/right LH2).
  • the nose-related index includes a first nose height (length of L1), a second nose height (length of L2), a third nose height (length of L3), a fourth nose height (length of L4),
  • the first nose width (the length of W1), the second nose width (the length of W2), the height of the V-line at the tip of the nose (the length of the H), the angle of the V-line at the tip of the nose, the nose ratio, and the like may be included.
  • the nose ratio may mean a ratio of a nose height to a nose width.
  • the nose width used to calculate the nose ratio may be the first nose width (the length of W1).
  • the nose height used in the calculation of the nose ratio is the sum of the first nose height (length of L1), the second nose height (length of L2), the third nose height (length of L3), and the fourth nose height (length of L4) can
  • FIG. 14 is a diagram illustrating an example of facial feature points extracted from a facial image according to an embodiment of the present disclosure.
  • the apparatus 100 for estimating age may extract 68 facial feature points P1 to P68 from a facial image.
  • the positions of the extracted facial feature points P1 to P68 may correspond to positions of the facial outline, lips, nose, eyes, and eyebrows.
  • 15 and 16 are diagrams illustrating examples of facial indicators extracted from a plurality of facial feature points shown in FIG. 14 .
  • the face index includes the total face area (total_area), the face width at the top of the cheek (or the first face width, total_width1), the face width in the middle of the cheek (or the second face width, total_width2), and the chin width (or third face width, total_width3), eyebrow gap (H_gap1), eye gap (H_gap2), chin to top of eyebrow height (or first face height, total_height1), mouth to top of eyebrow height (or second The face height, total_height2), the height from the chin to the top of the eyebrow (or the third face height, total_height3), the face ratio (total_ratio), and the like may be included.
  • facial indicators include lip area (lip_area), central upper lip height (or first upper lip height, lip_upper_height1), side upper lip height (or second upper lip height, lip_upper_height2), upper lip lip diagonal length (or third upper lip height, lip_upper_height3), middle lower lip height (or first lower lip height, lip_lower_height1), side lower lip height (or second lower lip height, lip_lower_height2), lower lip lip diagonal length (or third lower lip height, lip_lower_height3), lip width (lip_width), upper lip v-line angle (lip_v_degree), upper lip crest width (or upper lip v-line width, lip_upper_dist_H), upper lip crest and trough height (or upper lip v-line height, lip_upper_dist_V), lip ratio (lip_ratio), and the like may be included.
  • the face index includes the first nose height (nose_L1), the second nose height (nose_L2), the third nose height (nose_L3), the fourth nose height (nose_L4), the first nose width (nose_W1), the second nose width ( nose_W2), the height of the V-line at the tip of the nose (nose_V), the angle of the V-line at the tip of the nose (nose_degree), and the ratio of the nose (nose_ratio) may be included.
  • the face index may include a left eye width (eye_width_L), a right eye width (eye_width_R), and the like.
  • the face index includes a first left contour slope (Line_degree1_L), a first right contour slope (Line_degree1_R), a second left contour slope (Line_degree2_L), a second right contour slope (Line_degree2_R), a third left contour slope (Line_degree3_L), 3rd right contour slope (Line_degree3_R), 4th left contour slope (Line_degree4_L), 4th right contour slope (Line_degree4_R), 5th left contour slope (Line_degree5_L), 5th right contour slope (Line_degree5_R), 6th left contour slope (Line_degree6_L), 6th right contour slope (Line_degree6_R), 7th left contour slope (Line_degree7_L), 7th right contour slope (Line_degree7_R), 8th left contour slope (Line_degree8_L), 8th right contour slope (Line_degree8_R), etc. may be included.
  • the first left contour slope (Line_degree1_L) to the fourth left contour slope (Line_degree4_L) and the first right contour slope (Line_degree1_R) to the fourth right contour slope (Line_degree4_R) are the slopes of the corresponding contour lines may be the reciprocal of
  • 17 is a diagram illustrating upper face indexes having high age estimation ability for an age estimation model based on linear regression according to an embodiment of the present disclosure.
  • the age estimation model is a linear regression model, and the age is estimated from the face indicators shown in FIGS. 14 and 15 .
  • the age estimation model may be expressed as in Equation 1 below.
  • y is the user's age
  • n may be the number of face indicators. For example, if the increase in the specific face surface (a i) by one, and the user's age (y), estimated by the age presumption model is increased by ⁇ i ( ⁇ i is a negative number, the user's age is estimated (y) decreases).
  • the age estimation model estimates the user's age by additionally using body information or facial feature points as well as facial indicators, a i is the face index, body information, or facial feature points, and n is the face used to estimate the user's age. It may be the number of indicators, body information, and facial feature points.
  • each face index (a i ) is different, and the contribution may mean the ability to estimate the age.
  • the age estimation ability for each face index a i may be determined based on r 2 and p-value.
  • r 2 means a variance that can be explained by each face index (a i ) among the total age variance, and is also called a coefficient of determinatnion.
  • the p-value means whether the effect of each face index (a i ) on the user's age (y) is significant. Therefore, a face index (a i ) with a large r 2 and a small p-value has a large age estimation ability and can be viewed as significant.
  • each of the fourth right contour length line_degree_4_R and the fourth left contour length line_degree_4_L may estimate the age y of the user by about 10% or more.
  • the age can be estimated by considering the difference in Furthermore, in various embodiments, the age may be estimated by using the combined index of the facial indices shown in FIGS. 14 and 15 as the facial index.
  • the age estimation apparatus 100 calculates a small amount using an age estimation model for estimating the age of the user using only facial indicators whose age estimation ability (or r 2 ) exceeds a predetermined reference value. It is also possible to estimate the user's age. For example, the age estimation apparatus 100 may estimate the age of the user using an age estimation model that estimates the age using only the top 20 face indicators shown in FIG. 17 .
  • the 49 face indicators can estimate the age (y) of the user by about 46% or more. Confirmed.
  • the age of the user can be estimated from the face image, and there is an advantage in that it is possible to determine which face index has a high correlation with the age of the face.
  • a face index highly correlated with age may be used as an aging index.
  • a makeup method for obtaining a face index to look younger from the face image of a specific user may be proposed.
  • FIG. 18 is a view showing the distribution of the jaw edge angle (LD4) for each age group.
  • Each sample may contain each user's actual age and chin angle (LD4).
  • Figure 18 (b) is a box plot showing the distribution of the chin edge angle (LD4) by age group, (c) is a violin plot (violin) showing the distribution of the chin edge angle (LD4) by age group plot). Referring to (b) and (c) of FIG. 18 , it can be seen that the chin edge angle LD4 tends to increase as the age group increases.
  • LD4 chin edge angle
  • FIG. 19 (b) is a box plot showing the distribution of the chin corner angle (LD4) by age. Referring to (b) of FIG. 19 , it can be seen that the chin edge angle LD4 tends to increase as the age increases.
  • 20 is a diagram illustrating an age estimation model according to an embodiment of the present disclosure.
  • the age estimation model 2020 calculates an age 2031 of a user estimated in response thereto. can be printed out.
  • the facial index 2011 may be extracted from the facial feature points 2012 .
  • the face indicator 2011 may include at least some of the face indicators shown in FIGS. 14 and 15 .
  • the facial feature points 2012 may be extracted from the face image.
  • the facial feature point 2012 may mean a standardized facial feature point.
  • the facial feature point 2012 may have a value of (x-coordinate, y-coordinate), and the x-coordinate and y-coordinate of each facial feature point 2012 may be input as individual items.
  • the body information 2013 may be obtained by a user's input.
  • the body information 2013 may include a user's height, a user's weight, a user's body mass index (BMI), and the like.
  • the age estimation model 2020 may be a linear regression model, a decision tree, or a deep learning model composed of an artificial neural network.
  • 21 is a diagram illustrating an age estimation model based on deep learning according to an embodiment of the present disclosure.
  • the deep learning-based age estimation model 2120 is composed of an artificial neural network, and the artificial neural network may include an input layer 2121, one or more hidden layers 2122, and an output layer 2123. .
  • the age estimation model 2120 may include a fully connected network.
  • An input feature vector is input to the input layer 2121 , and the input feature vector may include at least one of a facial index 2111 , a facial feature point 2112 , and body information 2113 .
  • the output layer 2123 may be configured to include only a single node that outputs the estimated user's age 2031 , but the present disclosure is not limited thereto.
  • 22 to 25 are diagrams illustrating structures of learning data used for learning an age estimation model according to an embodiment of the present disclosure.
  • the training data 2210 used for learning the age estimation model may include the face index 2211 and the age 2215 as label information corresponding thereto.
  • the age estimation model may be trained to output a value estimated by following the corresponding age 2215 .
  • the training data 2210 used for learning the age estimation model may include a face index 2211 , a facial feature point 2212 , and an age 2215 as label information corresponding thereto.
  • the facial feature point 2212 may include an x-coordinate and a y-coordinate as individual items.
  • the facial feature point 2212 may mean a standardized facial feature point, and the facial indicator 2211 may be extracted from the facial feature point 2212 .
  • the age estimation model may be trained to output a value estimated by following the corresponding age 2215. .
  • the training data 2210 used for learning the age estimation model may include a face index 2211 , body information 2213 , and age 2215 as label information corresponding thereto.
  • the body information 2213 may include at least one of height, weight, and BMI.
  • the age estimation model may be trained to output a value estimated by following the corresponding age 2215. .
  • the training data 2210 used for learning the age estimation model includes a face index 2211 , a facial feature point 2212 , body information 2213 , and an age 2215 as label information corresponding thereto. can do.
  • the age estimation model follows the corresponding age 2215 and estimates a value. It can be learned to output.
  • the face indicator 2211 included in the training data 2210 shown in FIGS. 22 to 25 may include at least a portion of the face indicators shown in FIGS. 14 and 15 .
  • the face index 2211 included in the learning data 2210 may include only a predetermined number of face indexes in the order of the highest age estimation ability, like the upper face indexes shown in FIG. 17 .
  • 26 is a diagram illustrating an embodiment of estimating a user's age.
  • a user 2610 may photograph a face image including his or her own face by using an age estimation apparatus 2620 having a built-in camera.
  • the age estimation apparatus 2620 may be implemented as a portable user terminal such as a smart phone.
  • the age estimation apparatus 2620 extracts facial feature points from the photographed face image, extracts facial indicators based on the extracted facial feature points, and estimates the age of the user 2610 based on at least one or more of the facial feature points and the facial indicators. can do.
  • the age estimation apparatus 2620 may output the estimated age through a display unit or a speaker.
  • the age estimator 2620 may output the age estimated by voice, such as "Your face is 25 years old" (2631) through the speaker.
  • 27 is a diagram illustrating an embodiment of performing age authentication by estimating a user's age.
  • the age estimation apparatus 2720 may request age information of the user 2710 to perform a specific task. For example, when ordering alcohol or cigarettes in a shopping application, the age estimating device 2720 needs to determine whether the user 2710 is an adult. In this case, the age estimating device 2720 may request identification information for authenticating that the user 2710 is an adult, but if the face age is prioritized to perform adult authentication (or age authentication) and it is unclear whether the user is an adult You may also be asked for additional identification information.
  • the age estimating device 2720 may request the user 2710 to take a face image by outputting "Adult authentication is required. Please enter ID information or take a face picture.” 2731 through the display unit or speaker. .
  • the user 2710 may output the authentication button 2721 displayed on the display unit of the age estimation apparatus 2721 to perform adult authentication through face image-based age estimation.
  • the user 2710 photographs a face image using the age estimating device 2720 , and the age estimating device 2720 is the age of the user 2710 based on the captured face image. (or face age) can be estimated.
  • the age estimating apparatus 2720 may determine whether the age estimated for the user 2710 is suitable for adult authentication.
  • the age estimating device 2720 performs adult authentication based on the estimated age, if the estimated age is greater than or equal to an age required for adult authentication by a predetermined margin, the estimated age is suitable for adult authentication and adult authentication can be considered successful.
  • the age required for successful adult authentication is 20 years old
  • the age estimating device 2720 is suitable for adult authentication and succeeds in adult authentication when the estimated age is equal to or greater than the age (eg, 25 years) by adding a predetermined margin to 20 years old. It can be judged that
  • the age estimating device 2720 If the estimated age is suitable for adult authentication, the age estimating device 2720 outputs "adult authentication succeeded" 2732 through the display unit or the speaker, and the authentication button 2721 is displayed on the order button ( 2722) to provide a function that can be provided as adult authentication is performed.
  • 28 is a diagram illustrating an embodiment in which age authentication is performed by estimating an age.
  • the age estimating apparatus 2720 may determine whether the estimated age of the user 2710 is suitable for adult authentication. When the age estimating device 2720 performs adult authentication based on the estimated age, if the estimated age is not equal to or greater than the age required for adult authentication by a predetermined margin, the estimated age is not suitable for adult authentication and It may be determined that additional adult authentication is required. For example, if the age required for successful adult authentication is 20 years old, the age estimation device 2720 is not suitable for adult authentication if the estimated age is less than 20 years plus a predetermined margin (eg, 25 years old), and additional adult authentication is not performed. may be deemed necessary.
  • a predetermined margin eg, 25 years old
  • the age estimating device 2720 may display “additional adult authentication” through a display unit or a speaker. is required.” (2733), and the authentication button 2721 can be changed to an ID authentication button 2723 to provide an additional adult authentication function.
  • 27 and 28 disclose embodiments in which adult authentication is performed by estimating age, the present disclosure includes embodiments in which various age authentication is performed as well as simple adult authentication.
  • 29 is a diagram illustrating an embodiment of estimating the age of a user included in an image.
  • various objects and people may be included in an image 2910 captured through CCTV or various cameras.
  • the age estimation apparatus 100 may recognize the objects 2910 , 2920 , 2930 , 2940 , and 2950 included in the image 2910 , and identify the recognized objects 2910 , 2920 , 2930 , 2940 , and 2950 . can do.
  • the identification information for the recognized objects 2910 , 2920 , 2930 , 2940 , and 2950 may include type information and attribute information of the recognized object, and the type information and attribute information of the object are obtained using various object identification models. can be obtained
  • the attribute information of the object may include age information, and this age information may be obtained using the age estimation apparatus 100 according to an embodiment of the present disclosure. That is, the age estimation apparatus 100 according to an embodiment of the present disclosure may estimate the age of a person included in the image data, and generate the estimated age as identification information of the person.
  • the identification information 2911 for the first object 2910 may be “vehicle, light vehicle, silver”
  • the identification information 2921 for the second object 2920 may be “vehicle, silver”.
  • identification information 2931 for the third object 2930 is “person, white long-sleeved shirt”
  • identification information 2941 for the fourth object 2940 is “person, male, white short-sleeved shirt, black color” Bottom, the estimated age is 40”
  • the identification information 2951 for the fifth object 2950 may be “person, male, black/gray short-sleeved top, black bottom”. Since the fourth object 2940 among the recognized objects is a person and includes a face, the age estimation apparatus 100 estimates the age of the fourth object 2940 using image data corresponding to the fourth object 2940 . and, accordingly, the user's identity can be more clearly specified.
  • An embodiment of estimating the age of the user included in the image may be usefully used to specify the user included in the image captured by the surveillance camera. Accordingly, it is possible to specifically and effectively track a specific target, such as a criminal or a missing person, using the surveillance camera images.
  • 30 and 31 are diagrams illustrating embodiments in which an age-customized operation is performed by estimating a user's age.
  • the age estimation apparatus 3010 acquires image data including the face of the user 3020 using the camera 3011 , estimates the age of the user 3020 , and the estimated age Based on , an age-customized operation based on the estimated age of the user 3020 may be performed.
  • the age estimation device 3010 estimates the age of the user 3020 and provides (or suggests) a recommendation advertisement corresponding to the estimated age through the display unit 3012, or provides recommended content corresponding to the estimated age. (or suggest)
  • the age estimating device 3010 may provide a recommendation advertisement for promoting a product popular at the age estimated through the display unit 3012 , or provide a TV program or movie popular at the estimated age as recommended content You may.
  • the recommended content may mean media content to be played, such as images and videos, or content as an action, such as shopping, games, reading, or media appreciation.
  • the age estimation device 3010 may be a personal terminal of the user 3020 or a public terminal installed in a public facility used by an unspecified number of people.
  • the age estimation device 3010 may be a terminal installed on a train, a bus, or an airplane, or a digital signage or digital advertisement terminal installed indoors or outdoors. If the age estimating device 3010 is a public terminal, the image data taken for estimating the age is stored only temporarily and temporarily, and the age estimating device 3010 stores the image data after estimating the age of the user 3020 . You can delete it to protect your privacy.
  • the age estimation device 3010 obtains image data including the face of the user 3020 for age estimation, and asks the user 3020 for consent to take an image for age estimation in advance in that there is a risk of invasion of privacy.
  • the age estimating apparatus 3010 may determine the estimated age of the user 3020 in consideration of the additional personal information and the estimated age. It is possible to provide (or suggest) the recommended content corresponding to the . For example, when the age estimating device 3010 is a terminal installed in an individual seat of an airplane, the age estimating device 3010 provides additional personal information (eg, gender, nationality, race, etc.) of the user 3020 based on the airplane seat assignment information. ), and may provide (or suggest) recommended content based on additional personal information and the estimated age. To this end, the age estimation apparatus 3010 may receive recommended content information corresponding to at least one of age and additional personal information from a recommended content database (not shown).
  • a recommended content database not shown.
  • the age estimating device 3010 may be a public terminal installed in a seat of a means of transportation, and the age estimating device 3010 may estimate (3031) the age of the user 3020 as 30.
  • a “popular movie in 30s” corresponding to the age of 30 estimated through the display unit 3012 may be provided as recommended content (3031).
  • the age estimating device 3010 may provide a popular movie in their thirties while outputting a phrase such as “popular movie in their thirties” together with age information of the user 3020 on the display unit 3012 , or the user 3020 ), it is also possible to provide popular movies in their 30s while outputting only phrases such as “popular movies” excluding age information.
  • the age estimation device 3010 may be a digital signage installed outdoors or indoors, and the age estimation device 3010 may estimate the age of the user 3020 as 30 years old (3031).
  • a “popular restaurant in 30s” corresponding to the age of 30 estimated through the display unit 3012 may be provided as recommended content (3131).
  • the age estimating device 3010 may provide a popular restaurant in their 30s while outputting a phrase such as “a popular restaurant in their 30s” along with age information of the user 3020 on the display unit 3012 , and the user 3020 It is also possible to provide popular restaurants in their 30s by outputting only phrases such as “popular restaurants”, excluding age information of .
  • 32 is a diagram illustrating an embodiment of estimating the age of a user's face using a virtual face image.
  • the age estimating apparatus 100 may estimate the user's age when original image data 3210 including the user's face is input.
  • the age estimating apparatus 100 directly generates virtual image data 3220 or 3230 including a virtual face when a specific makeup, procedure, or surgery is performed on the user's face from the original image data 3210, or
  • the device may receive the virtual image data 3220 or 3230 generated by the device, and estimate an age when a specific makeup, operation, or surgery is performed based on the virtual image data 3220 or 3230 .
  • the age estimating apparatus 100 may estimate the user's age to be 25.1 years old when the original image data 3210 is input.
  • the age estimation apparatus 100 may generate or receive a first virtual face image 3220 including a virtual face after the chin filler is performed on the user's face included in the original image data 3210 .
  • the age of the user after the chin filler operation may be estimated to be 23.3 years old.
  • the age estimating apparatus 100 may generate or receive a second virtual face image 3230 including a virtual face after shading of the outer part with respect to the user's face included in the original image data 3210, Based on the second virtual face image 3230 , the age of the user after shading the outer part may be estimated to be 22.9 years old.
  • the age estimation apparatus 100 may provide the user with an expected age after any makeup, surgery, or surgery, and may further suggest a recommended makeup, a recommended procedure, or a recommended surgery.
  • the recommended makeup, the recommended procedure, or the recommended surgery provided to the user may be the makeup, procedure, or surgery in which the estimated age appears the youngest.
  • the age estimating apparatus 100 estimates an age (eg, 25.1 years old) from the original image data 3210 , and adjusts the estimated age value based on an input of a user (not shown).
  • Virtual image data including a face may be generated and provided.
  • the age estimating apparatus 100 may recommend a makeup method, a procedure, an operation, a health care method, etc. for reaching the virtual face included in the generated virtual image data.
  • the recommended makeup method may include the type of cosmetic to be used, a recommended cosmetic brand, and the like.
  • the recommended health management method may include a recommended diet, recommended food information, a recommended exercise routine, and the like.
  • the above-described method may be implemented as computer-readable code on a medium in which a program is recorded.
  • the computer-readable medium includes all kinds of recording devices in which data readable by a computer system is stored. Examples of computer-readable media include Hard Disk Drive (HDD), Solid State Disk (SSD), Silicon Disk Drive (SDD), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. There is this.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Business, Economics & Management (AREA)
  • General Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Strategic Management (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Tourism & Hospitality (AREA)
  • Economics (AREA)
  • Marketing (AREA)
  • General Business, Economics & Management (AREA)
  • Finance (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Data Mining & Analysis (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Accounting & Taxation (AREA)
  • Mathematical Physics (AREA)
  • Multimedia (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Development Economics (AREA)
  • Human Computer Interaction (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computer Security & Cryptography (AREA)
  • Primary Health Care (AREA)
  • Human Resources & Organizations (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Computer Hardware Design (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Game Theory and Decision Science (AREA)
  • Image Analysis (AREA)

Abstract

본 개시의 일 실시 예는 나이 추정 장치에 있어서, 얼굴 특징점 추출 모델 및 나이 추정 모델을 저장하는 메모리; 및 적어도 하나 이상의 얼굴 이미지를 수신하고, 상기 얼굴 특징점 추출 모델을 이용하여 상기 얼굴 이미지에서 복수의 얼굴 특징점을 추출하고, 상기 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하고, 상기 나이 추정 모델과 상기 복수의 얼굴 지표를 이용하여 상기 사용자의 나이를 추정하는 프로세서를 포함하는, 나이 추정 장치를 제공한다.

Description

나이 추정 장치 및 나이를 추정하는 방법
본 개시(disclosure)는 얼굴 이미지에 기초하여 사용자의 나이를 추정하는 나이 추정 장치 그 방법에 관한 것이다.
최근에 인공 지능에 기반한 이미지 처리 기술 또는 이미지 인식 기술이 많이 개발 및 이용되고 있다. 이러한 이미지 처리 기술 또는 이미지 인식 기술에는 이미지에 포함된 객체의 종류를 인식하는 기술부터, 이미지에 포함된 객체를 구체적으로 식별하는 기술 등이 포함된다.
이러한 이미지 인식 기술은 사용자의 얼굴을 식별하는 얼굴 인식 보안 기능을 제공하는데에도 사용되기도 한다. 그러나, 사용자의 얼굴을 식별하는 기술은 여러 사용자들을 서로 구분할 수 있을 뿐으로, 식별된 사용자에 대한 정보를 제공해주지 못한다.
만약 얼굴 이미지에 기초하여 사용자의 나이를 추정할 수 있다면, 사용자의 나이를 간편하게 파악할 수 있으며, 추정한 나이에 기초하여 다양한 기능을 제공할 수 있을 것이다.
본 개시는 얼굴 이미지에 기초하여 사용자의 나이를 추정하는 나이 추정 장치 및 그 방법을 제공하고자 한다.
본 개시의 일 실시 예에 따른 나이 추정 장치는 얼굴 특징점 추출 모델 및 나이 추정 모델을 저장하는 메모리; 및 적어도 하나 이상의 얼굴 이미지를 수신하고, 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점을 추출하고, 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하고, 나이 추정 모델과 복수의 얼굴 지표를 이용하여 사용자의 나이를 추정하는 프로세서를 포함할 수 있다.
프로세서는 복수의 얼굴 특징점을 표준화하고, 복수의 표준화된 얼굴 특징점으로부터 복수의 얼굴 지표를 추출할 수 있다.
프로세서는 복수의 얼굴 특징점들 중에서 적어도 일부에 기초하여 얼굴의 미리 정해진 부위가 미리 정해진 좌표에 배치되도록 얼굴 이미지의 변환 규칙을 결정하고, 변환 규칙에 기초하여 복수의 얼굴 특징점들을 표준화할 수 있다.
얼굴 특징점 추출 모델은 인공 신경망(neural network)으로 구성된 딥 러닝 모델이고, 얼굴 이미지가 입력되면 입력된 얼굴 이미지에 대응하는 미리 정해진 개수만큼의 복수의 얼굴 특징점을 추출할 수 있다.
프로세서는 복수의 얼굴 이미지를 수신한 경우, 복수의 얼굴 이미지 각각으로부터 복수의 얼굴 특징점을 추출하고, 서로 대응하는 위치의 얼굴 특징점의 평균을 산출하여 복수의 평균 얼굴 특징점들을 산출하고, 복수의 평균 얼굴 특징점들로부터 복수의 얼굴 지표를 추출할 수 있다.
프로세서는 복수의 얼굴 이미지 중에서 아웃라이어(outlier) 얼굴 이미지를 결정하고, 아웃라이어 얼굴 이미지를 제외한 얼굴 이미지로부터 복수의 얼굴 특징점을 추출할 수 있다.
복수의 얼굴 이미지는 미리 정해진 기간 이내에 촬영된 얼굴 이미지들일 수 있다.
나이 추정 모델은 선형 회귀 모델, 의사 결정 나무 또는 딥 러닝 모델이고, 적어도 복수의 얼굴 지표가 입력되면 추정된 나이를 출력하는 모델일 수 있다.
프로세서는 복수의 얼굴 특징점, 복수의 얼굴 지표 또는 사용자의 신체 정보 중에서 적어도 하나 이상 및 나이 추정 모델을 이용하여 사용자의 나이를 추정하고, 신체 정보는 사용자의 신장, 사용자의 체중 또는 사용자의 체지방지수(BMI) 중에서 적어도 하나 이상을 포함할 수 있다.
복수의 얼굴 지표는 얼굴 전체 면적, 입술 면적, 얼굴 관련 지표, 코 관련 지표 또는 입술 관련 지표 중에서 적어도 하나 이상을 포함하고, 얼굴 관련 지표는 얼굴 윤곽선의 기울기를 포함할 수 있다.
나이 추정 장치는 카메라를 더 포함하고, 프로세서는 카메라를 통해 적어도 하나 이상의 얼굴 이미지를 수신할 수 있다.
나이 추정 장치는 사용자 단말기와 통신하는 통신부를 더 포함하고, 프로세서는 통신부를 통해 사용자 단말기로부터 적어도 하나 이상의 얼굴 이미지를 수신할 수 있다.
프로세서는 추정된 나이를 이용하여 나이 인증시, 추정된 나이가 인증에서 요구하는 나이보다 미리 정해진 마진만큼 크거나 같은 경우, 추정된 나이가 나이 인증에 적합하다고 판단할 수 있다.
본 개시의 일 실시 예에 따른 나이를 추정하는 방법은 적어도 하나 이상의 얼굴 이미지를 수신하는 단계; 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점을 추출하는 단계; 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하는 단계; 및 나이 추정 모델과 복수의 얼굴 지표를 이용하여 사용자의 나이를 추정하는 단계를 포함할 수 있다.
본 개시의 다양한 실시 예에 따르면, 사용자의 얼굴 이미지를 촬영하는 것만으로도 간단하게 사용자의 나이를 추정할 수 있다.
또한, 본 개시의 다양한 실시 예에 따르면, 사용자의 얼굴 이미지를 촬영하는 것만으로도 간단하게 나이 기반 서비스를 제공할 수 있다.
도 1은 본 개시의 일 실시 예에 따른 나이 추정 장치를 나타낸 블록도이다.
도 2는 본 개시의 일 실시 예에 따른 인공 지능 서버를 나타낸 블록도이다.
도 3은 본 개시의 일 실시 예에 따른 인공 지능 시스템을 나타낸 도면이다.
도 4는 본 개시의 일 실시 예에 따른 얼굴 이미지에 기초하여 사용자의 나이를 추정하는 방법을 나타낸 동작 흐름도이다.
도 5는 본 개시의 일 실시 예에 따른 사용자의 얼굴을 촬영하는 방법을 나타낸 도면이다.
도 6은 복수의 얼굴 이미지의 예시들을 나타낸 도면이다.
도 7은 도 6에 도시된 복수의 얼굴 이미지에서 추출한 복수의 얼굴 특징점(facial feature)의 예시를 나타낸 도면이다.
도 8은 본 개시의 일 실시 예 따른 얼굴 지표들의 예시를 나타낸 도면이다.
도 9 내지 13은 도 7에 도시된 복수의 얼굴 특징점으로부터 추출한 얼굴 지표들의 예시를 나타낸 도면이다.
도 14는 본 개시의 일 실시 예에 따른 얼굴 이미지에서 추출된 얼굴 특징점들의 예시를 나타낸 도면이다.
도 15 및 16은 도 14에 도시된 복수의 얼굴 특징점으로부터 추출한 얼굴 지표들의 예시를 나타낸 도면이다.
도 17은 본 개시의 일 실시 예에 따른 선형 회귀 기반의 나이 추정 모델에 대한 나이 추정 능력이 높은 상위 얼굴 지표들을 나타낸 도면이다.
도 18은 연령대별 턱 모서리 각도(LD4)의 분포를 나타낸 도면이다.
도 19는 연령별 턱 모서리 각도(LD4)의 분포를 나타낸 도면이다.
도 20은 본 개시의 일 실시 예에 따른 나이 추정 모델을 나타낸 도면이다.
도 21은 본 개시의 일 실시 예에 따른 딥 러닝 기반의 나이 추정 모델을 나타낸 도면이다.
도 22 내지 25는 본 개시의 일 실시 예에 따른 나이 추정 모델의 학습에 이용되는 학습 데이터의 구조들을 나타낸 도면이다.
도 26은 사용자의 나이를 추정하는 일 실시 예를 나타낸 도면이다.
도 27은 사용자의 나이를 추정하여 나이 인증을 수행하는 일 실시 예를 나타낸 도면이다.
도 28은 사용자의 나이를 추정하여 나이 인증을 수행하는 일 실시 예를 나타낸 도면이다.
도 29는 이미지에 포함된 사용자의 나이를 추정하는 일 실시 예를 나타낸 도면이다.
도 30 및 31은 사용자의 나이를 추정하여 나이 맞춤형 동작을 수행하는 실시 예들을 나타낸 도면이다.
도 32는 가상의 얼굴 이미지를 이용하여 사용자의 나이를 추정하는 실시 예를 나타낸 도면이다.
이하, 첨부된 도면을 참조하여 본 명세서에 개시된 실시 예를 상세히 설명하되, 도면 부호에 관계없이 동일하거나 유사한 구성요소는 동일한 참조 번호를 부여하고 이에 대한 중복되는 설명은 생략하기로 한다. 이하의 설명에서 사용되는 구성요소에 대한 접미사 '모듈' 및 '부'는 명세서 작성의 용이함만이 고려되어 부여되거나 혼용되는 것으로서, 그 자체로 서로 구별되는 의미 또는 역할을 갖는 것은 아니다. 또한, 본 명세서에 개시된 실시 예를 설명함에 있어서 관련된 공지 기술에 대한 구체적인 설명이 본 명세서에 개시된 실시 예의 요지를 흐릴 수 있다고 판단되는 경우 그 상세한 설명을 생략한다. 또한, 첨부된 도면은 본 명세서에 개시된 실시 예를 쉽게 이해할 수 있도록 하기 위한 것일 뿐, 첨부된 도면에 의해 본 명세서에 개시된 기술적 사상이 제한되지 않으며, 본 개시의 사상 및 기술 범위에 포함되는 모든 변경, 균등물 내지 대체물을 포함하는 것으로 이해되어야 한다.
제1, 제2 등과 같이 서수를 포함하는 용어는 다양한 구성요소들을 설명하는데 사용될 수 있지만, 상기 구성요소들은 상기 용어들에 의해 한정되지는 않는다. 상기 용어들은 하나의 구성요소를 다른 구성요소로부터 구별하는 목적으로만 사용된다.
어떤 구성요소가 다른 구성요소에 '연결되어' 있다거나 '접속되어' 있다고 언급된 때에는, 그 다른 구성요소에 직접적으로 연결되어 있거나 또는 접속되어 있을 수도 있지만, 중간에 다른 구성요소가 존재할 수도 있다고 이해되어야 할 것이다. 반면에, 어떤 구성요소가 다른 구성요소에 '직접 연결되어' 있다거나 '직접 접속되어' 있다고 언급된 때에는, 중간에 다른 구성요소가 존재하지 않는 것으로 이해되어야 할 것이다.
도 1은 본 개시의 일 실시 예에 따른 나이 추정 장치(100)를 나타낸 블록도이다.
나이 추정 장치(100)는 TV, 프로젝터, 휴대폰, 스마트폰, 데스크탑 컴퓨터, 노트북, 디지털방송용 단말기, PDA(personal digital assistants), PMP(portable multimedia player), 네비게이션, 태블릿 PC, 웨어러블 장치, 셋톱박스(STB), DMB 수신기, 라디오, 세탁기, 냉장고, 데스크탑 컴퓨터, 디지털 사이니지, 로봇, 차량 등과 같은, 고정형 기기 또는 이동 가능한 기기 등으로 구현될 수 있다.
도 1을 참조하면, 나이 추정 장치(100)는 통신부(110), 입력부(120), 러닝 프로세서(130), 센싱부(140), 출력부(150), 메모리(170) 및 프로세서(180) 등을 포함할 수 있다.
통신부(110)는 통신 모뎀(communication modem) 또는 통신 회로(communication circuit)라고 칭할 수 있다.
통신부(110)는 유무선 통신 기술을 이용하여 인공 지능 서버(200) 등의 외부 장치들과 데이터를 송수신할 수 있다. 예컨대, 통신부(110)는 외부 장치들과 센서 정보, 사용자 입력, 학습 모델, 제어 신호 등을 송수신할 수 있다.
통신부(110)가 이용하는 통신 기술에는 GSM(Global System for Mobile communication), CDMA(Code Division Multi Access), LTE(Long Term Evolution), 5G, WLAN(Wireless LAN), Wi-Fi(Wireless-Fidelity), 블루투스(Bluetooth쪠), RFID(Radio Frequency Identification), 적외선 통신(Infrared Data Association; IrDA), ZigBee, NFC(Near Field Communication) 등이 있다.
입력부(120)는 입력 인터페이스(input interface)라고 칭할 수 있다.
입력부(120)는 다양한 종류의 데이터를 획득할 수 있다. 입력부(120)에서 수집한 음성 데이터나 이미지 데이터는 분석되어 사용자의 제어 명령으로 처리될 수 있다.
입력부(120)는 영상 신호를 수신하기 위한 카메라(Camera, 121), 오디오 신호를 수신하기 위한 마이크로폰(Microphone, 122) 또는 사용자로부터 정보를 입력 받기 위한 사용자 입력부(User Input Unit, 123) 등을 포함할 수 있다. 여기서, 카메라(121)나 마이크로폰(122)을 센서로 취급하여, 카메라(121)나 마이크로폰(122)으로부터 획득한 신호를 센서 데이터 또는 센서 정보라고 칭할 수 있다.
카메라(121)는 화상 통화 모드 또는 촬영 모드에서 이미지 센서에 의해 얻어지는 정지 영상 또는 동영상 등의 화상 프레임을 처리할 수 있다. 카메라(121)는 하나 또는 복수 개로 구성될 수 있다. 처리된 화상 프레임은 디스플레이부(Display Unit, 151)에 표시되거나 메모리(170)에 저장될 수 있다.
마이크로폰(122)은 음파를 수신하여 전기적인 음성 데이터로 변환할 수 있다. 변환된 음성 데이터는 나이 추정 장치(100)에서 수행 중인 기능(또는 실행 중인 응용 프로그램)에 따라 다양하게 활용될 수 있다. 한편, 마이크로폰(122)에는 외부의 음파를 수신하는 과정에서 잡음(noise)을 제거하기 위한 다양한 잡음 제거 알고리즘이 적용될 수 있다.
사용자 입력부(123)는 사용자로부터 정보를 입력 받기 위한 것으로서, 사용자 입력부(123)를 통해 정보가 입력되면, 프로세서(180)는 입력된 정보에 대응되도록 나이 추정 장치(100)의 동작을 제어할 수 있다. 사용자 입력부(123)는 기계식 (mechanical) 입력 수단 또는 터치식 입력 수단을 포함할 수 있다.
입력부(120)는 모델 학습을 위한 학습 데이터 및 학습 모델을 이용하여 출력을 획득할 때 사용될 입력 데이터 등을 획득할 수 있다. 입력부(120)는 가공되지 않은 입력 데이터를 획득할 수도 있으며, 이 경우 프로세서(180) 또는 러닝 프로세서(130)는 입력 데이터에 대하여 전처리로써 입력 특징점(input feature)을 추출할 수 있다.
러닝 프로세서(130)는 학습 데이터를 이용하여 인공 신경망으로 구성된 모델을 학습시킬 수 있다. 여기서, 학습된 인공 신경망을 학습 모델이라 칭할 수 있다. 학습 모델은 학습 데이터가 아닌 새로운 입력 데이터에 대하여 결과 값을 추론해 내는데 사용될 수 있고, 추론된 값은 어떠한 동작을 수행하기 위한 판단의 기초로 이용될 수 있다.
러닝 프로세서(130)는 인공 지능 서버(200)의 러닝 프로세서(240)과 함께 인공 지능 프로세싱을 수행할 수 있다.
센싱부(140)는 센서부라고 칭할 수 있다.
센싱부(140)는 다양한 센서들을 이용하여 나이 추정 장치(100) 내부 정보, 나이 추정 장치(100)의 주변 환경 정보 및 사용자 정보 중 적어도 하나를 획득할 수 있다.
센싱부(140)에 포함되는 센서에는 근접 센서, 조도 센서, 가속도 센서, 자기 센서, 자이로 센서, 관성 센서, RGB 센서, IR 센서, 지문 인식 센서, 초음파 센서, 광 센서, 마이크로폰, 라이다, 레이더 등이 있다.
출력부(150)는 출력 인터페이스(output interface)라고 칭할 수 있다.
출력부(150)는 시각 정보를 출력하는 디스플레이부(Display Unit, 151), 청각 정보를 출력하는 음향 출력부(Sound Output Unit, 152), 촉각 정보를 출력햅틱 모듈(Haptic Module, 153) 또는 광 출력부(Optical Output Unit, 154) 등을 포함할 수 있다.
디스플레이부(151)는 나이 추정 장치(100)에서 처리되는 정보를 표시(출력)할 수 있다. 예컨대, 디스플레이부(151)는 나이 추정 장치(100)에서 구동되는 응용 프로그램의 실행화면 정보, 또는 이러한 실행화면 정보에 따른 UI(User Interface), GUI(Graphic User Interface) 정보를 표시할 수 있다.
디스플레이부(151)는 터치 센서와 상호 레이어 구조를 이루거나 일체형으로 형성됨으로써, 터치 스크린으로 구현될 수 있다. 이러한 터치 스크린은 나이 추정 장치(100)와 사용자 사이의 입력 인터페이스와 출력 인터페이스를 동시에 제공할 수 있다.
음향 출력부(152)는 리시버(receiver), 스피커(speaker), 버저(buzzer) 중 적어도 하나 이상을 포함하여, 오디오 데이터를 음파로 출력할 수 있다.
햅틱 모듈(haptic module)(153)은 사용자가 느낄 수 있는 다양한 촉각 효과를 발생시킬 수 있다. 햅틱 모듈(153)이 발생시키는 촉각 효과의 대표적인 예로는 진동이 있다.
광출력부(154)는 나이 추정 장치(100)의 광원의 빛을 이용하여 이벤트 발생을 알리기 위한 신호를 출력한다. 나이 추정 장치(100)에서 발생 되는 이벤트의 예로는 메시지 수신, 호 신호 수신, 부재중 전화, 알람, 일정 알림, 이메일 수신, 애플리케이션을 통한 정보 수신 등이 될 수 있다.
메모리(170)는 나이 추정 장치(100)의 다양한 기능을 지원하는 데이터를 저장할 수 있다. 예컨대, 메모리(170)는 입력부(120)에서 획득한 입력 데이터, 학습 데이터, 학습 모델, 학습 히스토리, 어플리케이션 등을 저장할 수 있다.
메모리(170)는 데이터를 영구적, 반영구적 또는 임시적으로 저장하는 다양한 메모리를 지칭할 수 있다. 예컨대, 메모리(170)에는 HDD(Hard Disk Drive), SSD(Solid State Drive), 컴팩트 디스크(CD), RAM(Random Access Memory), ROM(Read Only Memory) 등이 포함될 수 있다.
프로세서(180)는 데이터 분석 알고리즘 또는 머신 러닝 알고리즘을 사용하여 결정되거나 생성된 정보에 기초하여, 나이 추정 장치(100)의 적어도 하나의 실행 가능한 동작을 결정할 수 있다. 그리고, 프로세서(180)는 나이 추정 장치(100)의 구성 요소들을 제어하여, 결정된 동작을 수행할 수 있다. 이를 위해, 프로세서(180)는 러닝 프로세서(130) 또는 메모리(170)의 데이터를 요청, 검색, 수신 또는 활용할 수 있고, 상기 적어도 하나의 실행 가능한 동작 중 예측되는 동작이나, 바람직한 것으로 판단되는 동작을 실행하도록 나이 추정 장치(100)의 구성 요소들을 제어할 수 있다.
프로세서(180)는 결정된 동작을 수행하기 위하여 외부 장치의 연계가 필요한 경우, 해당 외부 장치를 제어하기 위한 제어 신호를 생성하고, 통신부(110)를 통해 해당 외부 장치에 생성된 제어 신호를 전송할 수 있다.
프로세서(180)는 사용자 입력에 대하여 의도 정보를 획득하고, 획득한 의도 정보에 기초하여 사용자의 요구 사항을 결정할 수 있다.
예컨대, 프로세서(180)는 음성 입력을 문자열로 변환하기 위한 STT(Speech To Text) 엔진 또는 자연어의 의도 정보를 획득하기 위한 자연어 처리(NLP: Natural Language Processing) 엔진 중에서 적어도 하나 이상을 이용하여, 사용자 입력에 상응하는 의도 정보를 획득할 수 있다. STT 엔진 또는 NLP 엔진 중에서 적어도 하나 이상은 적어도 일부가 머신 러닝 알고리즘에 따라 학습된 인공 신경망으로 구성될 수 있다. 그리고, STT 엔진 또는 NLP 엔진 중에서 적어도 하나 이상은 러닝 프로세서(130)에 의해 학습된 것이나, 인공 지능 서버(200)의 러닝 프로세서(240)에 의해 학습된 것이거나, 또는 이들의 분산 처리에 의해 학습된 것일 수 있다.
프로세서(180)는 나이 추정 장치(100)의 동작 내용이나 동작에 대한 사용자의 피드백 등을 포함하는 이력 정보를 수집하여 메모리(170) 또는 러닝 프로세서(130)에 저장하거나, 인공 지능 서버(200) 등의 외부 장치에 전송할 수 있다. 수집된 이력 정보는 학습 모델을 갱신하는데 이용될 수 있다.
프로세서(180)는 메모리(170)에 저장된 응용 프로그램을 구동하기 위하여, 나이 추정 장치(100)의 구성 요소들 중 적어도 일부를 제어할 수 있다. 나아가, 프로세서(180)는 상기 응용 프로그램의 구동을 위하여, 나이 추정 장치(100)에 포함된 구성 요소들 중 둘 이상을 서로 조합하여 동작시킬 수 있다.
도 2는 본 개시의 일 실시 예에 따른 인공 지능 서버(200)를 나타낸 블록도이다.
도 2를 참조하면, 인공 지능 서버(200)는 머신 러닝 알고리즘을 이용하여 인공 신경망을 포함하는 인공 지능 모델을 학습시키거나, 학습된 인공 지능 모델을 이용하는 장치를 의미할 수 있다. 여기서, 인공 지능 서버(200)는 복수의 서버들로 구성되어 분산 처리를 수행할 수 있다. 또한, 인공 지능 서버(200)는 나이 추정 장치(100)와 함께 인공 지능 프로세싱을 분산하여 처리할 수도 있다.
인공 지능 서버(200)는 통신부(210), 메모리(230), 러닝 프로세서(240) 및 프로세서(260) 등을 포함할 수 있다.
통신부(210)는 나이 추정 장치(100) 등의 외부 장치와 데이터를 송수신할 수 있다.
메모리(230)는 모델 저장부(231)를 포함할 수 있다. 모델 저장부(231)는 러닝 프로세서(240)을 통하여 학습 중인 또는 학습된 모델(또는 인공 신경망, 231a)을 저장할 수 있다.
러닝 프로세서(240)는 학습 데이터를 이용하여 인공 신경망(231a)을 학습시킬 수 있다. 학습 모델은 인공 신경망의 인공 지능 서버(200)에 탑재된 상태에서 이용되거나, 나이 추정 장치(100) 등의 외부 장치에 탑재되어 이용될 수 있다.
학습 모델은 하드웨어, 소프트웨어 또는 하드웨어와 소프트웨어의 조합으로 구현될 수 있다. 학습 모델의 일부 또는 전부가 소프트웨어로 구현되는 경우 학습 모델을 구성하는 하나 이상의 명령어(instruction)는 메모리(230)에 저장될 수 있다.
프로세서(260)는 학습 모델을 이용하여 새로운 입력 데이터에 대하여 결과 값을 추론하고, 추론한 결과 값에 기초한 응답이나 제어 명령을 생성할 수 있다.
도 3은 본 개시의 일 실시 예에 따른 인공 지능 시스템(1)을 나타낸 도면이다.
도 3을 참조하면, 인공 지능 시스템(1)은 나이 추정 장치(100) 및 인공 지능 서버(200) 또는 사용자 단말기(300) 등을 포함할 수 있다. 그리고, 나이 추정 장치(100)는 각 장치는 유선 네트워크 또는 무선 네트워크를 통해 인공 지능 서버(200) 또는 사용자 단말기(300)와 연결될 수 있다.
인공 지능 서버(200)는 복수의 나이 추정 장치(100)들과 연결될 수 있고, 각 나이 추정 장치(100)의 인공 지능 프로세싱을 대신하여 수행하거나, 분산하여 수행할 수 있다.
인공 지능 서버(200)는 나이 추정 장치(100)를 대신하여 머신 러닝 알고리즘 또는 딥 러닝 알고리즘을 이용하여 인공 신경망을 학습시킬 수 있고, 학습 모델을 직접 저장하거나 나이 추정 장치(100)에 전송할 수 있다.
인공 지능 서버(200)는 나이 추정 장치(100)로부터 입력 데이터를 수신하고, 학습 모델을 이용하여 수신한 입력 데이터에 대응하는 결과 값을 추론하고, 추론한 결과 값에 기초한 응답이나 제어 명령을 생성하여 나이 추정 장치(100)로 전송할 수 있다. 또는, 나이 추정 장치(100)는 직접 학습 모델을 이용하여 입력 데이터에 대하여 결과 값을 추론하고, 추론한 결과 값에 기초한 응답이나 제어 명령을 생성할 수 있다.
사용자 단말기(300)는 직접 이미지를 촬영할 수 있는 이미지 촬영 장치일 수 있다.
사용자 단말기(300)는 나이 추정 장치(100)와 연결되며, 입력 데이터 또는 사용자 입력을 수신하여 나이 추정 장치(100)에 전달할 수 있다. 이 경우, 나이 추정 장치(100)는 사용자 단말기(300)로부터 수신한 입력 데이터에 대응하는 결과 값을 추론하고, 추론한 결과 값에 기초한 응답이나 제어 명령을 생성하여 사용자 단말기(300)에 전달할 수 있다.
도 4는 본 개시의 일 실시 예에 따른 얼굴 이미지에 기초하여 사용자의 나이를 추정하는 방법을 나타낸 동작 흐름도이다.
도 4를 참조하면, 나이 추정 장치(100)의 프로세서(180)는 사용자의 얼굴 이미지를 수신한다(S401).
사용자의 얼굴 이미지는 사용자의 얼굴을 포함하는 이미지 데이터를 의미할 수 있다.
프로세서(180)는 카메라(121)를 통해 사용자의 얼굴을 포함하는 이미지를 촬영함으로써 사용자의 얼굴 이미지를 수신하거나, 통신부(110)를 통해 사용자 단말기(300)로부터 사용자의 얼굴 이미지를 수신할 수 있다. 예컨대, 사용자 단말기(300)는 얼굴 이미지를 촬영하기 위한 전문적인 장치일 수 있고, 암막으로 외부 광을 차단한 상태에서 가시광 이미지, 편광 이미지, UV 이미지 등을 촬영할 수 있다. 또는, 사용자 단말기(300)는 스마트폰과 같이 카메라를 포함하는 단말기로, 가시광 이미지를 촬영할 수 있다.
사용자의 얼굴 이미지는 RGB 이미지, 편광 이미지, 흑백 이미지, UV 이미지, IR 이미지 또는 RGB-IR 이미지 등의 다양한 형식의 이미지 데이터일 수 있다. 또한, 사용자의 얼굴 이미지는 정지 화상 이미지 또는 동영상을 구성하는 이미지 프레임을 의미할 수 있다.
프로세서(180)는 일정한 기간 이내에 촬영된 복수의 얼굴 이미지를 수신할 수 있다. 예컨대, 프로세서(180)는 일정한 기간 이내에 촬영된 동종의 복수의 얼굴 이미지, 이종의 복수의 얼굴 이미지 또는 적어도 일부가 이종인 복수의 얼굴 이미지를 수신할 수 있다. 여기서, 동종의 얼굴 이미지는 얼굴 이미지의 촬영에 이용된 빛의 종류가 서로 동일한 얼굴 이미지들을 의미하고, 이종의 얼굴 이미지는 얼굴 이미지의 촬영에 이용된 빛의 종류가 서로 다른 얼굴 이미지들을 의미할 수 있다.
사용자의 얼굴 이미지는 세안 이후의 메이크업 되지 않은 얼굴을 포함하는 이미지뿐만 아니라, 메이크업 된 얼굴을 포함하는 이미지도 의미할 수 있다.
만약, 나이 추정 장치(100)가 사용자의 나이를 추정하기 위하여 신체 정보를 추가적으로 이용하는 경우, 프로세서(180)는 사용자의 신체 정보를 추가적으로 수신할 수도 있다. 이 경우, 프로세서(180)는 사용자 입력부(123) 또는 통신부(110)를 통해 사용자의 신체 정보를 수신할 수 있다.
그리고, 나이 추정 장치(100)의 프로세서(180)는 얼굴 이미지에서 복수의 얼굴 특징점(facial feature)을 추출한다(S403).
얼굴 특징점(facial feature)은 얼굴에서 추출한 키 포인트(key point) 또는 얼굴 랜드마크(facial landmark)를 의미할 수 있다. 키 포인트 또는 랜드마크는 물체의 형태나 크기, 위치가 변하더라도 쉽게 식별할 수 있는 지점을 의미할 수 있고, 미리 정해진 패턴 또는 규칙에 의하여 추출될 수 있다.
프로세서(180)는 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점을 추출할 수 있다. 예컨대, 프로세서(180)는 Dlib 라이브러리의 얼굴 랜드마크 추출 모델, HR-Net(High-Resolution Network)으로 구성된 얼굴 랜드마크 추출 모델 또는 SAN(Style Aggregated Network)으로 구성된 얼굴 랜드마크 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점(예컨대, 68개의 얼굴 특징점)을 추출할 수 있다. 또는, 프로세서(180)는 풀링 레이어(pooling layer), 컨벌루션 레이어(convolution layer), 전연결 레이어(fully-connected layer) 등을 포함하는 인공 신경망으로 구성된 딥 러닝 기반의 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점을 추출할 수도 있다. 또는, 프로세서(180)는 서포트 벡터 머신(Support vector machine)이나 앙상블 회귀 나무(ensemble regression tree)과 같은 기계 학습 모델 기반의 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지에서 복수의 얼굴 특징점을 추출할 수도 있다.
얼굴 특징점 추출 모델은 나이 추정 장치(100)의 러닝 프로세서(160) 또는 인공 지능 서버(200)의 러닝 프로세서(240) 중에서 적어도 하나 이상에 의하여 학습될 수 있다.
얼굴 특징점 추출 모델은 메모리(170) 또는 인공 지능 서버(200)의 메모리(230)에 저장될 수 있다. 얼굴 특징점 추출 모델이 인공 지능 서버(200)의 메모리(230)에 저장된 경우, 프로세서(180)는 통신부(110)를 통해 얼굴 이미지를 인공 지능 서버(200)로 전송하고, 인공 지능 서버(200)의 프로세서(260)는 메모리(230)에 저장된 얼굴 특징점 추출 모델을 이용하여 얼굴 이미지로부터 복수의 얼굴 특징점을 추출하고, 나이 추정 장치(100)의 프로세서(180)는 통신부(110)를 통해 인공 지능 서버(200)로부터 추출된 복수의 얼굴 특징점을 수신할 수 있다.
추출된 얼굴 특징점들은 각 특징점들의 좌표로 표현될 수 있다.
얼굴 특징점 추출 모델은 얼굴 이미지와 그에 대응하는 복수의 얼굴 특징점 (또는 얼굴 랜드마크)을 포함하는 제1 학습 데이터를 이용하여 학습될 수 있다. 예컨대, iBUG 300-W 데이터 세트, AFLW (Annotated Facial Landmarks in the Wild) 데이터 세트 또는 WFLW (Wider Facial Landmarks in the Wild) 데이터 세트 등이 얼굴 특징점 추출 모델의 학습에 이용되는 제1 학습 데이터로서 이용될 수 있다.
프로세서(180)는 얼굴 이미지에서 얼굴을 포함하는 직사각형의 얼굴 영역을 검출하고, 얼굴 영역에서 복수의 얼굴 특징점을 추출할 수 있다. 이 과정에서, 프로세서(180)는 얼굴 이미지에서 얼굴 영역만을 크롭(crop)하고, 크롭된 얼굴 영역에서 복수의 얼굴 특징점을 추출할 수 있다.
그리고, 나이 추정 장치(100)의 프로세서(180)는 복수의 얼굴 특징점을 표준화한다(S405).
얼굴을 촬영하는 상황에 따라 얼굴 이미지에 포함되는 얼굴의 크기나 방향이 달라질 수 있기 때문에, 서로 다른 사용자(또는 사람)의 얼굴을 촬영한 얼굴 이미지들뿐만 아니라 동일한 사용자(또는 사람)의 얼굴을 촬영한 얼굴 이미지들에서도 포함된 얼굴의 크기나 모양이 달라질 수 있다. 서로 크기나 모양이 다른 얼굴 이미지에서 얼굴 특징점을 추출하게 될 경우, 여러 얼굴 이미지들 사이의 비교가 어려워진다. 이에, 프로세서(180)는 복수의 얼굴 특징점들을 미리 정해진 기준에 따라 표준화 또는 보정할 수 있다. 이하에서, 얼굴 특징점은 표준화된 얼굴 특징점을 의미할 수 있다.
프로세서(180)는 복수의 얼굴 특징점들 중에서 적어도 일부에 기초하여 얼굴의 미리 정해진 부위들이 미리 정해진 좌표에 배치되도록 얼굴 이미지의 변환 규칙을 결정할 수 있고, 결정한 변환 규칙에 기초하여 각 얼굴 특징점들을 표준화 또는 보정할 수 있다. 예컨대, 프로세서(180)는 복수의 얼굴 특징점들에 기초하여 눈, 코 및 입의 위치를 결정하고, 눈, 코 및 입이 미리 정해진 좌표에 배치되도록 얼굴 이미지의 확대/축소 배율, 평행 이동 거리, 회전 각도 등의 변환 규칙을 결정할 수 있다. 그리고, 프로세서(180)는 확대/축소 배율, 평행 이동 거리, 회전 각도 등의 변환 규칙에 기초하여 각 얼굴 특징점들을 표준화 또는 보정할 수 있다.
프로세서(180)는 얼굴 이미지에 이마 받침대(headrest)가 포함된 경우, 이마 받침대가 미리 정해진 좌표에 배치되도록 얼굴 이미지의 변환 규칙을 결정하고, 결정된 변환 규칙에 기초하여 얼굴 이미지를 표준화 또는 보정할 수 있다. 만약, 얼굴 이미지에 복수의 이마 받침대가 포함된 경우, 프로세서(180)는 복수의 이마 받침대 사이의 거리를 고려하여 얼굴 이미지의 확대/축소 배율을 포함하는 변환 규칙을 결정할 수 있다.
복수의 얼굴 특징점들이 표준화됨에 따라, 서로 다른 얼굴 이미지에서 추출된 복수의 얼굴 특징점들이라 하더라도 서로 대응되는 위치의 얼굴 특징점들은 동일하거나 인접한 위치에 배치되게 되며, 서로 비교하기에 적합하다.
그리고, 나이 추정 장치(100)의 프로세서(180)는 복수의 표준화된 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표(facial indicator)를 추출한다(S407).
각 얼굴 특징점은 절대적인 좌표가 갖는 의미가 작고, 둘 이상의 얼굴 특징점들 사이의 관계(상대적인 좌표 관계)가 갖는 의미가 크다. 즉, 둘 이상의 얼굴 특징점들 사이의 거리나 방향(또는 각도)이 사람의 얼굴에 있어서 의미있는 정보를 제공한다고 볼 수 있다. 이에, 프로세서(180)는 복수의 얼굴 특징점으로부터 미리 정해진 방법에 기초하여 복수의 얼굴 지표들을 추출할 수 있다.
얼굴 지표에는 미리 정해진 얼굴 특징점들 사이의 거리, 미리 정해진 얼굴 특징점 사이의 방향(또는 각도), 얼굴의 너비, 얼굴의 높이, 얼굴의 넓이, 얼굴의 윤곽선의 각도 등이 포함될 수 있다. 예컨대, 제1 얼굴 지표는 제1 얼굴 특징점과 제2 얼굴 특징점 사이의 거리, 제2 얼굴 지표는 제1 얼굴 특징점과 제2 얼굴 특징점 사이의 방향(또는 각도)일 수 있다. 얼굴 지표의 종류에 대한 구체적인 설명은 후술한다.
프로세서(180)는 복수의 얼굴 이미지로부터 추출되어 표준화된 얼굴 특징점들을 비교함으로써 복수의 얼굴 이미지 중에서 오차가 큰 아웃라이어(outlier) 얼굴 이미지를 결정하고, 아웃라이어 얼굴 이미지를 제외한 얼굴 이미지들에서 추출되어 표준화된 얼굴 특징점들을 이용하여 복수의 얼굴 지표를 추출할 수 있다.
아웃라이어 얼굴 이미지는 표준화된 얼굴 특징점의 분포가 평균으로부터 미리 정해진 배수의 표준 편차를 벗어나는 얼굴 이미지를 의미할 수 있다. 또는, 아웃라이어 얼굴 이미지는 전체 얼굴 이미지 중에서 표준화된 얼굴 특징점의 분포가 평균으로부터 편차가 큰 미리 정해진 개수 또는 비율만큼의 얼굴 이미지를 의미할 수 있다. 예컨대, 프로세서(180)는 3개의 얼굴 이미지 중에서 표준화된 얼굴 특징점의 분포가 평균으로부터 편차가 가장 큰 1개의 얼굴 이미지를 아웃라이어 얼굴 이미지로 결정할 수 있다.
프로세서(180)는 아웃라이어 얼굴 이미지를 제외한 얼굴 이미지들에 대응하는 표준화된 얼굴 특징점들의 평균에 기초하여 복수의 얼굴 지표를 계산할 수 있다. 구체적으로, 프로세서(180)는 아웃라이어 얼굴 이미지를 제외한 얼굴 이미지들에 대하여 서로 대응하는 위치의 얼굴 특징점끼리 평균을 계산하고, 평균 얼굴 특징점들에 기초하여 복수의 얼굴 지표를 계산할 수 있다. 예컨대, 제1 얼굴 이미지와 제2 얼굴 이미지가 모두 아웃라이어 얼굴 이미지가 아닌 경우, 프로세서(180)는 제1 얼굴 이미지의 코 끝에 위치한 얼굴 특징점과 제2 얼굴 이미지의 코 끝에 위치한 얼굴 특징점의 평균을 산출하고, 코 끝에 위치한 얼굴 특징점들의 평균에 기초하여 얼굴 지표를 추출할 수 있다.
후술하는 도 8, 도 15 및 16은 본 개시에서 사용할 수 있는 얼굴 지표들의 예시를 나타낸다. 프로세서(180)는 표준화된 얼굴 특징점에 기초하여 후술하는 얼굴 지표들 중에서 적어도 일부를 추출할 수 있다.
그리고, 나이 추정 장치(100)의 프로세서(180)는 나이 추정 모델과 복수의 얼굴 지표를 이용하여 사용자의 나이를 추정한다(S409).
나이 추정 모델은 복수의 얼굴 지표가 입력되면, 입력된 복수의 얼굴 지표로부터 사용자의 나이를 추정할 수 있다. 사용자의 얼굴에서 추출된 얼굴 지표를 이용하여 나이를 추정한다는 점에서, 추정되는 나이는 얼굴 나이를 의미할 수 있다.
나아가, 프로세서(180)는 신체 정보 또는 추출된 얼굴 특징점들을 추가로 이용하여 사용자의 나이를 추정할 수 있다. 이 경우, 나이 추정 모델은 복수의 얼굴 지표뿐만 아니라 신체 정보 또는 추출된 얼굴 특징점들이 입력되면, 입력된 정보로부터 추론한 사용자의 나이를 출력할 수 있다.
나이 추정 모델은 선형 회귀 모델, 의사 결정 나무(decision tree) 또는 인공 신경망으로 구성된 딥 러닝 모델 등으로 구현될 수 있다. 예컨대, 나이 추정 모델은 전연결 네트워크(Fully Connected Network)으로 구성될 수 있다. 프로세서(180)는 사용자의 나이를 추정하는데 이용할 데이터 (예컨대, 얼굴 지표, 신체 정보, 얼굴 특징점 등)에 대응하는 나이 추정 모델을 이용할 수도 있고, 반대로 사용자의 나이를 추정하는데 이용할 나이 추정 모델에 대응하는 데이터를 이용할 수 있다.
나이 추정 모델은 나이 추정 장치(100)의 러닝 프로세서(160) 또는 인공 지능 서버(200)의 러닝 프로세서(240) 중에서 적어도 하나 이상에 의하여 학습될 수 있다.
나이 추정 모델은 메모리(170) 또는 인공 지능 서버(200)의 메모리(230)에 저장될 수 있다. 나이 추정 모델이 인공 지능 서버(200)의 메모리(230)에 저장된 경우, 프로세서(180)는 통신부(110)를 통해 복수의 얼굴 지표를 인공 지능 서버(200)로 전송하고, 인공 지능 서버(200)의 프로세서(260)는 메모리(230)에 저장된 나이 추정 모델을 이용하여 복수의 얼굴 지표로부터 사용자의 나이를 추정하고, 나이 추정 장치(100)의 프로세서(180)는 통신부(110)를 통해 인공 지능 서버(200)로부터 추론된 사용자의 나이를 수신할 수 있다.
나이 추정 모델은 복수의 얼굴 지표와 그에 대응하는 사용자의 나이를 포함하는 제2 학습 데이터를 이용하여 학습될 수 있다.
일 실시 예에서, 나이 추정 모델은 얼굴 지표, 얼굴 특징점 또는 신체 정보 중에서 적어도 하나 이상이 입력되면, 그에 대응하는 사용자의 나이를 추정하는 모델일 수 있다. 이 경우, 나이 추정 모델은 복수의 얼굴 지표, 복수의 얼굴 특징점 또는 신체 정보 중에서 적어도 하나 이상을 포함하고, 그에 대응하는 사용자의 나이를 더 포함하는 제2 학습 데이터를 이용하여 학습될 수도 있다. 신체 정보에는 사용자의 신장과 사용자의 몸무게 또는 사용자의 체질량지수(BMI: Body Mass Index) 등이 포함될 수 있다.
도 4에 도시된 단계들(steps)의 순서는 하나의 예시에 불과하며, 본 개시가 이에 한정되지는 않는다. 즉, 일 실시 예에서, 도 4에 도시된 단계들 중 일부 단계의 순서가 서로 바뀌어 수행될 수도 있다. 또한, 일 실시 예에서, 도 4에 도시된 단계들 중 일부 단계는 병렬적으로 수행될 수도 있다. 또한, 도 4에 도시된 단계들 중 일부만 수행될 수도 있다.
도 5는 본 개시의 일 실시 예에 따른 사용자의 얼굴을 촬영하는 방법을 나타낸 도면이다.
도 5를 참조하면, 사용자 단말기(510)는 사용자(520)의 얼굴을 촬영하여 얼굴 이미지를 생성하는 장치일 수 있다. 사용자 단말기(510)은 얼굴을 촬영하기 위한 광원(미도시)와 카메라(미도시)를 포함할 수 있다. 또한, 사용자 단말기(510)는 사용자(520)의 얼굴을 기대어 고정할 수 있는 턱 받침대(chinrest, 511) 또는 머리 받침대(headrest, 512) 등을 포함할 수 있다.
도 5에 도시된 것과 같은 사용자 단말기(510)를 이용하여 사용자(520)의 얼굴을 촬영할 경우, 사용자(520)의 얼굴이 카메라(미도시)에 대하여 고정된 위치에서 통제된 빛의 조사를 통해 촬영므로, 노이즈가 적고 안정된 이미지를 획득할 수 있다. 또한, 복수의 사용자들에 대하여도 동일한 조건으로 얼굴 이미지를 생성할 수 있는 장점이 있다.
도 6은 복수의 얼굴 이미지의 예시들을 나타낸 도면이다.
도 6를 참조하면, 나이 추정 장치(100)는 한 명의 사용자의 나이를 추정하기 위하여 복수의 얼굴 이미지(610, 620 및 630)를 촬영 또는 수신할 수 있다.
복수의 얼굴 이미지(610, 620 및 630)는 일정한 기간 이내에 촬영될 수 있다. 각 얼굴 이미지(610, 620 또는 630)는 미리 정해진 시간 간격으로 촬영된 이미지일 수도 있고, 동영상에서 미리 정해진 간격의 프레임 이미지들일 수도 있다.
각 얼굴 이미지(610, 620 또는 630)는 동일한 빛의 종류를 이용하여 촬영될 수도 있고, 서로 다른 빛의 종류를 이용하여 촬영될 수도 있다. 예컨대, 제1 얼굴 이미지(610)는 가시광 이미지 (또는 RGB 이미지), 제2 얼굴 이미지(620)는 편광 이미지, 제3 얼굴 이미지(630)는 UV 이미지일 수 있다.
각 얼굴 이미지(610, 620 또는 630)에는 얼굴 이미지의 촬영 또는 생성시 사용자의 얼굴을 고정하기 위한 턱 받침대(511) 또는 머리 받침대(512) 등이 포함될 수 있다.
도 7은 도 6에 도시된 복수의 얼굴 이미지에서 추출한 복수의 얼굴 특징점(facial feature)의 예시를 나타낸 도면이다.
도 7을 참조하면, 복수의 얼굴 이미지(610, 620 및 630) 각각에 대하여 미리 정해진 위치에 대응하는 얼굴 특징점들(721, 722 및 723)이 추출될 수 있다. 도 7은 시각화를 위하여 복수의 얼굴 이미지(610, 620 및 630) 중에서 하나의 얼굴 이미지(710) 또는 임의의 얼굴 이미지에 추출된 얼굴 특징점들(721, 722 및 723)을 나타낸 이미지로, 실제 나이를 추정하는 과정에서는 이러한 시각화된 얼굴 이미지가 생성되지 않을 수 있다.
또한, 얼굴 이미지(610, 620 및 630)에 머리 받침대(512)가 포함되어 있다면, 머리 받침대(512)에 대응하는 특징점(711)이 추출될 수 있고, 나이 추정 장치(100)는 머리 받침대(512)에 대응하는 특징점(711)에 기초하여 추출된 얼굴 특징점들(721, 722 및 723)을 표준화할 수 있다. 도 7에 도시된 얼굴 특징점들(721, 722 및 723)은 표준화된 얼굴 특징점들일 수 있다.
각 얼굴 특징점(721, 722 및 723)은 미리 정해진 지점에 대응하여 추출될 수 있다. 제1 얼굴 특징점(721)은 제1 얼굴 이미지(610)에서 추출된 얼굴 특징점이고, 제2 얼굴 특징점(722)은 제2 얼굴 이미지(620)에서 추출된 얼굴 특징점이고, 제3 얼굴 특징점(723)은 제3 얼굴 이미지(630)에서 추출된 얼굴 특징점들일 수 있다. 그리고, 제4 얼굴 특징점(731) 또는 평균 얼굴 특징점(731)은 각 정해진 지점에서의 제1 얼굴 특징점(721), 제2 얼굴 특징점(722) 및 제3 얼굴 특징점(723)의 평균 위치 (또는 무게 중심)을 의미할 수 있다. 나이 추정 장치(100)는 평균 얼굴 특징점(731)을 이용하여 얼굴 지표들을 추출하고, 추출한 얼굴 지표들을 이용하여 사용자의 나이를 추정할 수 있다.
도 7은 하나의 예시로써 3개의 얼굴 이미지(610, 620 및 630)에서 추출한 얼굴 특징점들(721, 722, 723)을 도시한 것에 불과하며, 본 개시가 이에 한정되지 않는다. 즉, 실시 예에 따라 더 많거나 더 적은 얼굴 이미지들에서 얼굴 특징점들을 추출하고, 추출된 얼굴 특징점들의 평균 위치 (또는 무게 중심)을 평균 얼굴 특징점으로 결정할 수 있다.
도 8은 본 개시의 일 실시 예 따른 얼굴 지표들의 예시를 나타낸 도면이다.
도 8을 참조하면, 본 개시의 일 실시 예에서 활용할 수 있는 얼굴 지표(Facial Indicator)에는 도시된 49개의 얼굴 지표 중에서 적어도 일부가 포함될 수 있다.
본 개시의 일 실시 예에 따른 얼굴 지표들에는 얼굴 전체 면적, 입술 면적, 입술 관련 지표, 코 관련 지표, 얼굴 관련 지표 등이 포함될 수 있다.
도 9 내지 13은 도 7에 도시된 복수의 얼굴 특징점으로부터 추출한 얼굴 지표들의 예시를 나타낸 도면이다.
도 9를 참조하면, 얼굴 전체 면적(911)은 눈썹 아래의 얼굴 면적을 의미할 수 있다. 얼굴 전체 면적(911)은 눈썹에서 추출된 얼굴 특징점들과 얼굴 윤곽에서 추출된 얼굴 특징점들을 연결한 다각형의 면적을 계산함으로써 추출할 수 있다.
도 10을 참조하면, 입술 면적(1011)은 윗 입술과 아랫 입술의 면적을 의미할 수 있다. 입술 면적(1011)은 윗 입술에서 추출된 얼굴 특징점들과 아랫 입술에서 추출된 얼굴 특징점들 중에서 외각 얼굴 특징점들을 연결한 다각형의 면적을 계산함으로써 추출할 수 있다.
도 11을 참조하면, 얼굴 관련 지표에는 눈썹 사이 간격(G1의 길이), 눈 사이 간격(G2의 길이), 뺨 상단의 얼굴 너비(W1의 길이), 뺨 중간의 얼굴 너비(W2의 길이), 턱 너비(W3의 길이), 눈 너비(좌/우 EW의 길이), 턱에서 눈썹 상단까지 높이(H1의 길이), 입에서 눈썹 상단까지 높이(H2의 길이), 인중에서 눈썹 상단까지 높이(H3의 길이) 또는 얼굴 윤곽선 기울기(좌/우 LD1~LD4의 기울기의 역수 또는 좌/우 LD5~LD8의 기울기) 등이 포함될 수 있다.
얼굴 윤곽선의 기울기는 좌측 윤곽선의 기울기와 우측 윤곽선의 기울기가 개별적으로 산출될 수 있으며, 상대적으로 얼굴 상단에서의 윤곽선(LD1~LD4)의 기울기가 얼굴 하단에서의 윤곽선(LD5~LD8)의 기울기에 비하여 가파르다는 점에서, 얼굴 상단에서의 윤곽선(LD1~LD4)의 기울기는 역수의 값을 지표로 이용할 수 있다.
도 12를 참조하면, 입술 관련 지표에는 중앙 윗 입술 높이(UH의 길이), 옆 윗 입술 높이(좌/우 UH2의 길이의 평균), 윗 입술 입꼬리 대각 길이(좌/우 UH3의 길이의 평균), 중앙 아랫 입술 높이(LH의 길이), 옆 아랫 입술 높이(좌/우 LH2의 길이의 평균), 아랫 입술 입꼬리 대각 길이(좌/우 LH3의 길이의 평균), 입술 너비(W의 길이), 윗 입술 V라인 각도, 윗 입술 마루 너비 또는 윗 입술 V라인 너비(UDH의 길이), 윗 입술의 마루와 골 사이의 높이 또는 윗 입술 V라인 높이(UDV의 길이), 입술 비율 등이 포함될 수 있다.
입술 비율은 입술의 너비에 대한 입술의 높이의 비율을 의미할 수 있다. 입술 비율의 계산에 이용되는 입술 높이는 옆 윗 입술 높이(좌/우 UH2의 길이의 평균)와 옆 아랫 입술의 높이(좌/우 LH2의 길이의 평균)의 합일 수 있다.
도 13을 참조하면, 코 관련 지표에는 제1 코 높이(L1의 길이), 제2 코 높이(L2의 길이), 제3 코 높이(L3의 길이), 제4 코 높이(L4의 길이), 제1 코 너비(W1의 길이), 제2 코 너비(W2의 길이), 코 끝 V라인 높이(H의 길이), 코 끝 V라인 각도, 코 비율 등이 포함될 수 있다.
코 비율은 코 너비에 대한 코 높이의 비율을 의미할 수 있다. 코 비율의 계산에 이용되는 코 너비는 제1 코 너비(W1의 길이)일 수 있다. 코 비율의 계산에 이용되는 코 높이는 제1 코 높이(L1의 길이), 제2 코 높이(L2의 길이), 제3 코 높이(L3의 길이) 및 제4 코 높이(L4의 길이)의 합일 수 있다.
도 14는 본 개시의 일 실시 예에 따른 얼굴 이미지에서 추출된 얼굴 특징점들의 예시를 나타낸 도면이다.
도 14를 참조하면, 본 개시의 일 실시 예에서 나이 추정 장치(100)는 얼굴 이미지로부터 68개의 얼굴 특징점들(P1 내지 P68)을 추출할 수 있다. 추출된 얼굴 특징점들(P1 내지 P68)의 위치는 얼굴 윤곽선, 입술, 코, 눈, 눈썹의 위치에 대응할 수 있다.
도 15 및 16은 도 14에 도시된 복수의 얼굴 특징점으로부터 추출한 얼굴 지표들의 예시를 나타낸 도면이다.
도 15 및 16을 참조하면, 얼굴 지표에는 얼굴 전체 면적(total_area), 뺨 상단의 얼굴 너비(또는 제1 얼굴 너비, total_width1), 뺨 중간의 얼굴 너비(또는 제2 얼굴 너비, total_width2), 턱 너비(또는 제3 얼굴 너비, total_width3), 눈썹 사이 간격(H_gap1), 눈 사이 간격(H_gap2), 턱에서 눈썹 상단까지 높이(또는 제1 얼굴 높이, total_height1), 입에서 눈썹 상단까지 높이(또는 제2 얼굴 높이, total_height2), 인중에서 눈썹 상단까지 높이(또는 제3 얼굴 높이, total_height3), 얼굴 비율(total_ratio) 등이 포함될 수 있다.
또한, 얼굴 지표에는 입술 면적(lip_area), 중앙 윗 입술 높이(또는 제1 윗 입술 높이, lip_upper_height1), 옆 윗 입술 높이(또는 제2 윗 입술 높이, lip_upper_height2), 윗 입술 입꼬리 대각 길이(또는 제3 윗 입술 높이, lip_upper_height3), 중앙 아랫 입술 높이(또는 제1 아랫 입술 높이, lip_lower_height1), 옆 아랫 입술 높이(또는 제2 아랫 입술 높이, lip_lower_height2), 아랫 입술 입꼬리 대각 길이(또는 제3 아랫 입술 높이, lip_lower_height3), 입술 너비(lip_width), 윗 입술 V라인 각도(lip_v_degree), 윗 입술 마루 너비(또는 윗 입술 V라인 너비, lip_upper_dist_H), 윗 입술의 마루와 골 사이의 높이(또는 윗 입술 V라인 높이, lip_upper_dist_V), 입술 비율(lip_ratio) 등이 포함될 수 있다.
또한, 얼굴 지표에는 제1 코 높이(nose_L1), 제2 코 높이(nose_L2), 제3 코 높이(nose_L3), 제4 코 높이(nose_L4), 제1 코 너비(nose_W1), 제2 코 너비(nose_W2), 코 끝 V라인 높이(nose_V), 코 끝 V라인 각도(nose_degree), 코 비율(nose_ratio) 등이 포함될 수 있다.
또한, 얼굴 지표에는 왼쪽 눈 너비(eye_width_L), 오른쪽 눈 너비(eye_width_R) 등이 포함될 수 있다.
또한, 얼굴 지표에는 제1 왼쪽 윤곽선 기울기(Line_degree1_L), 제1 오른쪽 윤곽선 기울기(Line_degree1_R), 제2 왼쪽 윤곽선 기울기(Line_degree2_L), 제2 오른쪽 윤곽선 기울기(Line_degree2_R), 제3 왼쪽 윤곽선 기울기(Line_degree3_L), 제3 오른쪽 윤곽선 기울기(Line_degree3_R), 제4 왼쪽 윤곽선 기울기(Line_degree4_L), 제4 오른쪽 윤곽선 기울기(Line_degree4_R), 제5 왼쪽 윤곽선 기울기(Line_degree5_L), 제5 오른쪽 윤곽선 기울기(Line_degree5_R), 제6 왼쪽 윤곽선 기울기(Line_degree6_L), 제6 오른쪽 윤곽선 기울기(Line_degree6_R), 제7 왼쪽 윤곽선 기울기(Line_degree7_L), 제7 오른쪽 윤곽선 기울기(Line_degree7_R), 제8 왼쪽 윤곽선 기울기(Line_degree8_L), 제8 오른쪽 윤곽선 기울기(Line_degree8_R) 등이 포함될 수 있다.
상술하였듯, 일 실시 예에서, 제1 왼쪽 윤곽선 기울기(Line_degree1_L) 내지 제4 왼쪽 윤곽선 기울기(Line_degree4_L) 및 제1 오른쪽 윤곽선 기울기(Line_degree1_R) 내지 제4 오른쪽 윤곽선 기울기(Line_degree4_R)는 대응하는 윤곽선의 기울기의 역수일 수 있다.
도 17은 본 개시의 일 실시 예에 따른 선형 회귀 기반의 나이 추정 모델에 대한 나이 추정 능력이 높은 상위 얼굴 지표들을 나타낸 도면이다.
도 17은 나이 추정 모델이 선형 회귀 모델이고, 도 14 및 15에 도시된 얼굴 지표들로부터 나이를 추정하는 상황을 가정한다. 이 경우, 나이 추정 모델은 하기 [수학식 1]과 같이 표현될 수 있다.
Figure PCTKR2020019328-appb-img-000001
상기 [수학식 1]에서 y는 사용자의 나이이고, β i (i=0, ... , n)는 i번째 모델 계수이고, a i (i=1, ..., n)는 i번째 얼굴 지표이고, n은 얼굴 지표의 개수일 수 있다. 예컨대, 특정 얼굴 지표(a i)가 1만큼 증가한다면, 나이 추정 모델에 의하여 추정되는 사용자의 나이(y)는 β i만큼 증가한다 (β i가 음수일 경우, 추정되는 사용자의 나이(y)는 감소한다).
만약, 나이 추정 모델이 얼굴 지표뿐만 아니라 신체 정보 또는 얼굴 특징점을 추가적으로 이용하여 사용자의 나이를 추정하는 경우, a i는 얼굴 지표, 신체 정보 또는 얼굴 특징점이고, n은 사용자의 나이를 추정하는데 이용하는 얼굴 지표, 신체 정보 및 얼굴 특징점의 개수일 수 있다.
나이 추정 모델에서 각 얼굴 지표(a i)가 기여하는 정도가 상이하며, 그 기여도는 나이 추정 능력을 의미할 수 있다. 각 얼굴 지표(a i)에 대한 나이 추정 능력은 r 2와 p-value에 기초하여 판단할 수 있다. r 2는 전체 나이 분산 중에서 각 얼굴 지표(a i)로 설명 가능한 분산을 의미하며, 결정 계수(coefficient of determinatnion)이라 칭하기도 한다. p-value는 각 얼굴 지표(a i)가 사용자의 나이(y)에 미치는 영향이 유의한지(significant)를 의미한다. 따라서, r 2가 크고 p-value가 작은 얼굴 지표(a i)는 나이 추정 능력이 크며 유의하고 볼 수 있다.
도 17을 참조하면, 나이 추정 모델을 학습하는데 충분히 많은 학습 데이터가 사용되었고, 그에 따라 얼굴 지표들의 p-value가 충분히 작고, 얼굴 지표들이 유의함을 확인할 수 있다. 따라서, 각 얼굴 지표는 r 2이 클수록 나이 추정 능력이 좋다고 볼 수 있다.
실제 나이 추정 모델을 학습해본 결과, 턱 모서리 각도(LD4)인 제4 오른쪽 윤곽선 길이(line_degree_4_R)와 제4 왼쪽 윤곽선 길이(line_degree_4_L)가 다른 얼굴 지표들과 비교하여 나이 추정 능력이 크다는 것을 확인하였다. 즉, 제4 오른쪽 윤곽선 길이(line_degree_4_R)와 제4 왼쪽 윤곽선 길이(line_degree_4_L) 각각이 사용자의 나이(y)를 약 10% 이상 추정할 수 있다.
일 실시 예에서, 상술한 나이 추정 모델과 달리 제4 오른쪽 윤곽선 길이(line_degree_4_R)와 제4 왼쪽 윤곽선 길이(line_degree_4_L)뿐만 아니라, 제4 오른쪽 윤곽선 길이(line_degree_4_R)와 제4 왼쪽 윤곽선 길이(line_degree_4_L) 사이의 차이를 고려하여 나이를 추정할 수 있다. 나아가, 다양한 실시 예들에서, 도 14 및 15에 도시된 얼굴 지표들의 조합 지표도 얼굴 지표로써 이용하여 나이를 추정할 수 있다.
도 17에 도시된 것과 같이, 얼굴 지표마다 나이 추정 능력이 상이하다. 많은 얼굴 지표들을 이용할 경우 사용자의 나이를 보다 정확히 추정할 수 있지만 요구되는 연산이 크게 늘어날 수 있다. 따라서, 일 실시 예에서, 나이 추정 장치(100)는 나이 추정 능력이 (또는 r 2)가 미리 정해진 기준 값을 넘는 얼굴 지표들만을 이용하여 사용자의 나이를 추정하는 나이 추정 모델을 이용하여 적은 연산으로도 사용자의 나이를 추정할 수 있다. 예컨대, 나이 추정 장치(100)는 도 17에 도시된 상위 20개의 얼굴 지표들만을 이용하여 나이를 추정하는 나이 추정 모델을 이용하여 사용자의 나이를 추정할 수 있다.
특히, 도 14 및 15에 도시된 49개의 얼굴 지표들을 모두 이용하여 선형 회귀 기반의 얼굴 인식 모델을 학습시킨 경우, 49개의 얼굴 지표들이 사용자의 나이(y)를 약 46% 이상 추정할 수 있음을 확인하였다.
도 17에 도시된 것과 같이, 얼굴 인식 모델을 학습시키게 될 경우 얼굴 이미지로부터 사용자의 나이를 추정할 수도 있으며, 얼굴에서 어떠한 얼굴 지표가 나이와 상관관계가 높은지 파악할 수 있다는 장점이 있다. 이 경우, 나이와 상관관계가 높은 얼굴 지표는 노화 지표로 사용할 수 있다. 또한, 이를 활용할 경우, 특정 사용자의 얼굴 이미지로부터 더 젊게 보이기 위한 얼굴 지표를 얻을 수 있는 화장 방법을 제안할 수도 있다.
도 18은 연령대별 턱 모서리 각도(LD4)의 분포를 나타낸 도면이다.
도 18의 (a)는 연령대별 샘플 수를 나타낸다. 각 샘플은 각 사용자의 실제 나이와 턱 모서리 각도(LD4)를 포함할 수 있다.
도 18의 (b)는 연령대별 턱 모서리 각도(LD4)의 분포를 나타낸 박스 플롯(box plot)이고, 도 18의 (c)는 연령대별 턱 모서리 각도(LD4)의 분포를 나타낸 바이올린 플롯(violin plot)이다. 도 18의 (b) 및 (c)를 참조하면, 연령대가 증가함에 따라 턱 모서리 각도(LD4)가 증가하는 추세에 있음을 확인할 수 있다.
도 19는 연령별 턱 모서리 각도(LD4)의 분포를 나타낸 도면이다.
도 19의 (a)는 연령별 샘플 수를 나타낸다. 각 샘플은 각 사용자의 실제 나이와 턱 모서리 각도(LD4)를 포함할 수 있다.
도 19의 (b)는 연령별 턱 모서리 각도(LD4)의 분포를 나타낸 박스 플롯이다. 도 19의 (b)를 참조하면, 연령이 증가함에 따라 턱 모서리 각도(LD4)가 증가하는 추세에 있음을 확인할 수 있다.
도 20은 본 개시의 일 실시 예에 따른 나이 추정 모델을 나타낸 도면이다.
도 20을 참조하면, 나이 추정 모델(2020)는 얼굴 지표(2011), 얼굴 특징점(2012) 또는 신체 정보(2013) 중에서 적어도 하나 이상이 입력되면, 그에 대응하여 추정된 사용자의 나이(2031)를 출력할 수 있다.
얼굴 지표(2011)는 얼굴 특징점(2012)로부터 추출될 수 있다. 얼굴 지표(2011)에는 도 14 및 15에 도시된 얼굴 지표들 중에서 적어도 일부가 포함될 수 있다.
얼굴 특징점(2012)은 얼굴 이미지로부터 추출될 수 있다. 얼굴 특징점(2012)은 표준화된 얼굴 특징점을 의미할 수 있다. 그리고, 얼굴 특징점(2012)은 (x좌표, y좌표)의 값을 가질 수 있으며, 각 얼굴 특징점(2012)의 x좌표와 y좌표가 개별적인 항목으로 입력될 수 있다.
신체 정보(2013)는 사용자의 입력에 의해 획득될 수 있다. 신체 정보(2013)에는 사용자의 신장, 사용자의 체중, 사용자의 체질량지수(BMI) 등이 포함될 수 있다.
나이 추정 모델(2020)은 선형 회귀 모델, 의사 결정 나무 또는 인공 신경망으로 구성된 딥 러닝 모델일 수 있다.
도 21은 본 개시의 일 실시 예에 따른 딥 러닝 기반의 나이 추정 모델을 나타낸 도면이다.
도 21을 참조하면, 딥 러닝 기반의 나이 추정 모델(2120)은 인공 신경망으로 구성되고, 인공 신경망은 입력 레이어(2121), 하나 이상의 히든 레이어(2122) 및 출력 레이어(2123)을 포함할 수 있다. 그리고, 나이 추정 모델(2120)은 전연결 네트워크(Fully Connected Network)를 포함할 수 있다.
입력 레이어(2121)에는 입력 특징 벡터(input feature vector)가 입력되며, 입력 특징 벡터에는 얼굴 지표(2111), 얼굴 특징점(2112) 또는 신체 정보(2113) 중에서 적어도 하나 이상이 포함될 수 있다. 출력 레이어(2123)는 추정한 사용자의 나이(2031)를 출력하는 단일한 노드만을 포함하도록 구성될 수 있으나, 본 개시가 이에 한정되지는 않는다.
도 22 내지 25는 본 개시의 일 실시 예에 따른 나이 추정 모델의 학습에 이용되는 학습 데이터의 구조들을 나타낸 도면이다.
도 22를 참조하면, 나이 추정 모델의 학습에 이용되는 학습 데이터(2210)는 얼굴 지표(2211)와 그에 대응하는 라벨 정보로써 나이(2215)를 포함할 수 있다. 이 경우, 나이 추정 모델은 학습 데이터(2210)에 포함된 얼굴 지표(2211)가 입력되면, 그에 대응하는 나이(2215)를 추종하여 추정하는 값을 출력하도록 학습될 수 있다.
도 23을 참조하면, 나이 추정 모델의 학습에 이용되는 학습 데이터(2210)는 얼굴 지표(2211), 얼굴 특징점(2212)과 그에 대응하는 라벨 정보로써 나이(2215)를 포함할 수 있다. 얼굴 특징점(2212)은 x좌표와 y좌표가 개별 항목으로 포함될 수 있다. 이미 설명하였듯이, 얼굴 특징점(2212)는 표준화된 얼굴 특징점을 의미할 수 있고, 얼굴 지표(2211)은 얼굴 특징점(2212)로부터 추출될 수 있다. 이 경우, 나이 추정 모델은 학습 데이터(2210)에 포함된 얼굴 지표(2211) 및 얼굴 특징점(2212)이 입력되면, 그에 대응하는 나이(2215)를 추종하여 추정하는 값을 출력하도록 학습될 수 있다.
도 24를 참조하면, 나이 추정 모델의 학습에 이용되는 학습 데이터(2210)는 얼굴 지표(2211), 신체 정보(2213)와 그에 대응하는 라벨 정보로써 나이(2215)를 포함할 수 있다. 신체 정보(2213)에는 신장, 체중, BMI 중에서 적어도 하나 이상이 포함될 수 있다. 이 경우, 나이 추정 모델은 학습 데이터(2210)에 포함된 얼굴 지표(2211) 및 신체 정보(2213)가 입력되면, 그에 대응하는 나이(2215)를 추종하여 추정하는 값을 출력하도록 학습될 수 있다.
도 25를 참조하면, 나이 추정 모델의 학습에 이용되는 학습 데이터(2210)는 얼굴 지표(2211), 얼굴 특징점(2212), 신체 정보(2213)와 그에 대응하는 라벨 정보로써 나이(2215)를 포함할 수 있다. 이 경우, 나이 추정 모델은 학습 데이터(2210)에 포함된 얼굴 지표(2211), 얼굴 특징점(2212) 및 신체 정보(2213)가 입력되면, 그에 대응하는 나이(2215)를 추종하여 추정하는 값을 출력하도록 학습될 수 있다.
도 22 내지 25에 도시된 학습 데이터(2210)에 포함되는 얼굴 지표(2211)는 도 14 및 15에 도시된 얼굴 지표 중에서 적어도 일부를 포함할 수 있다. 일 실시 예에서, 학습 데이터(2210)에 포함되는 얼굴 지표(2211)는 도 17에 도시된 상위 얼굴 지표들과 같이, 나이 추정 능력이 높은 순서대로 미리 정해진 개수의 얼굴 지표만을 포함할 수 있다.
도 26은 사용자의 나이를 추정하는 일 실시 예를 나타낸 도면이다.
도 26을 참조하면, 사용자(2610)는 카메라를 내장한 나이 추정 장치(2620)를 이용하여 자신의 얼굴이 포함된 얼굴 이미지를 촬영할 수 있다. 도 26에 도시된 것과 같이, 나이 추정 장치(2620)는 스마트폰과 같은 휴대 가능한 사용자 단말기로 구현될 수 있다.
나이 추정 장치(2620)는 촬영한 얼굴 이미지에서 얼굴 특징점을 추출하고, 추출한 얼굴 특징점에 기초하여 얼굴 지표들을 추출하고, 얼굴 특징점 또는 얼굴 지표 중에서 적어도 하나 이상에 기초하여 사용자(2610)의 나이를 추정할 수 있다.
나이 추정 장치(2620)는 디스플레이부 또는 스피커를 통해 추정한 나이를 출력할 수 있다. 예컨대, 나이 추정 장치(2620)는 스피커를 통해 "얼굴 나이는 25세 입니다."(2631)와 같이 음성으로 추정한 나이를 출력할 수 있다.
도 27은 사용자의 나이를 추정하여 나이 인증을 수행하는 일 실시 예를 나타낸 도면이다.
도 27의 (a)를 참조하면, 나이 추정 장치(2720)는 특정 작업을 수행하기 위해 사용자(2710)의 나이 정보를 요구할 수 있다. 예컨대, 쇼핑 어플리케이션에서 주류나 담배를 주문할 경우, 나이 추정 장치(2720)는 사용자(2710)이 성인인지 판단할 필요가 있다. 이 경우, 나이 추정 장치(2720)는 사용자(2710)가 성인임을 인증할 수 있는 신분증 정보를 요구할 수도 있지만, 얼굴 나이를 우선적으로 판단하여 성인 인증 (또는 나이 인증)을 수행하고 성인인지 불분명할 경우에 추가적으로 신분증 정보를 요구할 수도 있다.
나이 추정 장치(2720)는 사용자(2710)에게 디스플레이부 또는 스피커를 통해 "성인 인증이 필요합니다. 신분증 정보를 입력하거나 얼굴을 촬영해주세요."(2731)와 같이 출력하여 얼굴 이미지 촬영을 요청할 수 있다. 사용자(2710)는 나이 추정 장치(2721)의 디스플레이부에 표시된 인증 버튼(2721)을 출력하여 얼굴 이미지 기반 나이 추정을 통한 성인 인증을 수행할 수 있다.
도 27의 (b)를 참조하면, 사용자(2710)는 나이 추정 장치(2720)를 이용하여 얼굴 이미지를 촬영하고, 나이 추정 장치(2720)는 촬영한 얼굴 이미지에 기초하여 사용자(2710)의 나이 (또는 얼굴 나이)를 추정할 수 있다.
도 27의 (c)를 참조하면, 나이 추정 장치(2720)는 사용자(2710)에 대하여 추정된 나이가 성인 인증에 적합한지 판단할 수 있다. 나이 추정 장치(2720)는 추정된 나이에 기초한 성인 인증을 수행할 때, 추정된 나이가 성인 인증에서 요구하는 나이보다 미리 정해진 마진만큼 크거나 같은 경우, 추정된 나이가 성인 인증에 적합하며 성인 인증에 성공하였다고 판단할 수 있다. 예컨대, 성공적인 성인 인증에 필요한 나이가 20세인 경우, 나이 추정 장치(2720)는 추정된 나이가 20세보다 미리 정해진 마진을 더한 나이(예컨대 25세) 이상일 경우에 성인 인증에 적합하며 성인 인증에 성공하였다고 판단할 수 있다.
만약, 추정된 나이가 성인 인증에 적합한 경우, 나이 추정 장치(2720)는 디스플레이부 또는 스피커를 통해 "성인 인증에 성공하였습니다."(2732)와 같이 출력하고, 인증 버튼(2721)을 주문 버튼(2722)으로 변경하여 성인 인증이 수행됨에 따라 제공할 수 있는 기능을 제공할 수 있다.
도 28은 나이를 추정하여 나이 인증을 수행하는 일 실시 예를 나타낸 도면이다.
도 28의 (a) 및 (b)는 도 27의 (a) 및 (b)와 동일하며, 중복되는 설명은 생략한다.
도 28의 (c)를 참조하면, 나이 추정 장치(2720)는 사용자(2710)에 대하여 추정된 나이가 성인 인증에 적합한지 판단할 수 있다. 나이 추정 장치(2720)는 추정된 나이에 기초한 성인 인증을 수행할 때, 추정된 나이가 성인 인증에서 요구하는 나이보다 미리 정해진 마진만큼 크거나 같지 않은 경우, 추정된 나이가 성인 인증에 적합하지 않으며 추가 성인 인증이 필요하다고 판단할 수 있다. 예컨대, 성공적인 성인 인증에 필요한 나이가 20세인 경우, 나이 추정 장치(2720)는 추정된 나이가 20세보다 미리 정해진 마진을 더한 나이(예컨대 25세) 미만일 경우에 성인 인증에 부적합하며 추가 성인 인증이 필요하다고 판단할 수 있다.
만약, 추정된 나이가 성인 인증에서 요구하는 나이보다 적거나, 성인 인증에서 요구하는 나이로부터 미리 정해진 마진 이내의 차이를 갖는 경우, 나이 추정 장치(2720)는 디스플레이부 또는 스피커를 통해 "추가 성인 인증이 필요합니다."(2733)와 같이 출력하고, 인증 버튼(2721)을 신분증 인증 버튼(2723)으로 변경하여 추가 성인 인증 기능을 제공할 수 있다.
상기 도 27 및 도 28는 나이를 추정하여 성인 인증을 수행하는 실시 예들을 개시하고 있으나, 본 개시는 단순히 성인 인증뿐만 아니라 다양한 나이 인증을 수행하는 실시 예도 포함한다.
도 29는 이미지에 포함된 사용자의 나이를 추정하는 일 실시 예를 나타낸 도면이다.
도 29를 참조하면, CCTV나 다양한 카메라를 통하여 촬영된 이미지(2910)에는 여러 사물과 사람들이 포함될 수 있다. 나이 추정 장치(100)는 이미지(2910)에 포함된 객체들(2910, 2920, 2930, 2940, 2950)을 인식할 수 있고, 인식된 객체들(2910, 2920, 2930, 2940, 2950)을 식별할 수 있다.
인식된 객체들(2910, 2920, 2930, 2940, 2950)에 대한 식별 정보는 인식된 객체의 종류 정보와 속성 정보를 포함할 수 있고, 객체의 종류 정보와 속성 정보는 다양한 객체 식별 모델을 이용하여 획득할 수 있다. 객체의 속성 정보에는 나이 정보가 포함될 수 있고, 이러한 나이 정보는 본 개시의 일 실시 예에 따른 나이 추정 장치(100)를 이용하여 획득할 수 있다. 즉, 본 개시의 일 실시 예에 따른 나이 추정 장치(100)는 이미지 데이터에 포함된 사람의 나이를 추정하고, 추정된 나이를 해당 사람의 식별 정보로 생성할 수 있다.
예컨대, 제1 객체(2910)에 대한 식별 정보(2911)는 "차량, 경차, 은색"이고, 제2 객체(2920)에 대한 식별 정보(2921)는 "차량, 은색"일 수 있다. 마찬가지로, 제3 객체(2930)에 대한 식별 정보(2931)는 "사람, 흰색 긴팔 상의"이고, 제4 객체(2940)에 대한 식별 정보(2941)는 "사람, 남성, 흰색 반팔 상의, 검은색 하의, 추정 나이 40"이고, 제5 객체(2950)에 대한 식별 정보(2951)는 "사람, 남성, 검은색/회색 반팔 상의, 검은색 하의"일 수 있다. 인식된 객체들 중에서 제4 객체(2940)가 사람이면서 얼굴을 포함하므로, 나이 추정 장치(100)는 제4 객체(2940)에 대응하는 이미지 데이터를 이용하여 제4 객체(2940)의 나이를 추정할 수 있고, 그에 따라 사용자의 신분을 보다 명확하게 특정할 수 있다.
이미지에 포함된 사용자의 나이를 추정하는 실시 예는 감시 카메라를 통해 촬영된 이미지에 포함된 사용자를 특정하는데 유용하게 이용될 수 있다. 이에 따라, 감시 카메라 이미지들을 이용하여 범죄자나 실종자 등의 특정한 대상을 구체적으로 특정하고 효과적으로 추적할 수 있다.
도 30 및 31은 사용자의 나이를 추정하여 나이 맞춤형 동작을 수행하는 실시 예들을 나타낸 도면이다.
도 30 및 31을 참조하면, 나이 추정 장치(3010)는 카메라(3011)를 이용하여 사용자(3020)의 얼굴을 포함하는 이미지 데이터를 획득하고, 사용자(3020)의 나이를 추정하고, 추정된 나이에 기초하여 사용자(3020)의 추정된 나이에 기초한 나이 맞춤형 동작을 수행할 수 있다.
나이 추정 장치(3010)는 사용자(3020)의 나이를 추정하고, 디스플레이부(3012)를 통해 추정한 나이에 대응하는 추천 광고를 제공 (또는 제안)하거나, 추정한 나이에 대응하는 추천 컨텐츠를 제공 (또는 제안)할 수 있다. 나이 추정 장치(3010)는 디스플레이부(3012)를 통해 추정한 나이에서 인기가 많은 제품을 홍보하기 위한 추천 광고를 제공할 수도 있고, 추정한 나이에서 인기가 많은 TV 프로그램이나 영화를 추천 컨텐츠로 제공할 수도 있다. 추천 컨텐츠는 이미지, 동영상과 같은 재생될 미디어 컨텐츠를 의미할 수도 있고, 쇼핑, 게임, 독서, 미디어 감상과 같은 행위로써의 컨텐츠를 의미할 수도 있다.
나이 추정 장치(3010)는 사용자(3020)의 개인 단말기일 수도 있지만, 불특정 다수가 이용하는 공용 시설에 설치된 공용 단말기일 수도 있다. 예컨대, 나이 추정 장치(3010)는 기차, 버스 또는 비행기 등에 설치된 단말기, 실내나 야외에 설치된 디지털 사이니지나 디지털 광고 단말기 등일 수 있다. 만약, 나이 추정 장치(3010)가 공용 단말기인 경우에는 나이 추정을 위하여 촬영한 이미지 데이터는 임시적 및 일시적으로만 저장되며, 나이 추정 장치(3010)는 사용자(3020)의 나이 추정 이후에 이미지 데이터를 삭제하여 프라이버시를 보호할 수 있다.
나이 추정 장치(3010)는 나이 추정을 위해 사용자(3020)의 얼굴을 포함하는 이미지 데이터를 획득하며, 이는 프라이버시 침해의 우려가 있다는 점에서 사용자(3020)에게 미리 나이 추정을 위한 이미지 촬영 동의를 구할 수 있다.
나이 추정 장치(3010)가 사용자(3020)에 대한 추가 개인 정보를 획득할 수 있는 경우라면, 나이 추정 장치(3010)는 추가 개인 정보와 추정된 나이를 고려하여 사용자(3020)에 대하여 추정된 나이에 대응하는 추천 컨텐츠를 제공 (또는 제안)할 수 있다. 예컨대, 나이 추정 장치(3010)가 비행기의 개별 좌석에 설치된 단말기인 경우, 나이 추정 장치(3010)는 비행기 좌석 배정 정보에 기초하여 사용자(3020)의 추가 개인 정보(예컨대, 성별, 국적, 인종 등)를 획득할 수 있고, 추가 개인 정보와 추정된 나이에 기초하여 추천 컨텐츠를 제공 (또는 제안)할 수 있다. 이를 위해, 나이 추정 장치(3010)는 추천 컨텐츠 데이터베이스(미도시)로부터 나이와 추가 개인 정보 중에서 적어도 하나 이상에 대응하는 추천 컨텐츠 정보를 수신할 수 있다.
도 30에 도시된 실시 예에서, 나이 추정 장치(3010)는 교통 수단의 좌석에 설치된 공용 단말기일 수 있고, 나이 추정 장치(3010)는 사용자(3020)의 나이를 30세로 추정(3031)할 수 있고, 디스플레이부(3012)를 통해 추정된 나이 30세에 대응하는 "30대 인기 영화"를 추천 컨텐츠로 제공(3031)할 수 있다. 이 경우, 나이 추정 장치(3010)는 디스플레이부(3012)에 사용자(3020)의 나이 정보와 함께 "30대 인기 영화"와 같은 문구를 출력하면서 30대의 인기 영화를 제공할 수도 있고, 사용자(3020)의 나이 정보를 제외하고 "인기 영화"와 같은 문구만을 출력하면서 30대의 인기 영화를 제공할 수도 있다.
도 31에서 도시된 실시 예에서, 나이 추정 장치(3010)는 실외나 실내에 설치된 디지털 사이니지일 수 있고, 나이 추정 장치(3010)는 사용자(3020)의 나이를 30세로 추정(3031)할 수 있고, 디스플레이부(3012)를 통해 추정된 나이 30세에 대응하는 "30대 인기 식당"을 추천 컨텐츠로 제공(3131)할 수 있다. 마찬가지로, 나이 추정 장치(3010)는 디스플레이부(3012)에 사용자(3020)의 나이 정보와 함께 "30대 인기 식당"과 같은 문구를 출력하면서 30대의 인기 식당을 제공할 수도 있고, 사용자(3020)의 나이 정보를 제외하고 "인기 식당"과 같은 문구만을 출력하면서 30대의 인기 식당을 제공할 수도 있다.
도 32는 가상의 얼굴 이미지를 이용하여 사용자의 얼굴의 나이를 추정하는 실시 예를 나타낸 도면이다.
도 32를 참조하면, 본 개시의 일 실시 예는 나이 추정 장치(100)는 사용자의 얼굴을 포함하는 원본 이미지 데이터(3210)가 입력되면, 사용자의 나이를 추정할 수 있다. 그리고, 나이 추정 장치(100)는 직접 원본 이미지 데이터(3210)로부터 사용자의 얼굴에 특정한 화장, 시술 또는 수술을 하였을 때의 가상의 얼굴을 포함하는 가상 이미지 데이터(3220 또는 3230)를 생성하거나, 다른 장치에서 생성된 가상 이미지 데이터(3220 또는 3230)를 수신하고, 가상 이미지 데이터(3220 또는 3230)에 기초하여 특정한 화장, 시술 또는 수술을 하였을 때의 나이를 추정할 수 있다.
도 32의 실시 예에서, 나이 추정 장치(100)는 원본 이미지 데이터(3210)가 입력되었을 때 사용자의 나이를 25.1세로 추정할 수 있다. 그리고, 나이 추정 장치(100)는 원본 이미지 데이터(3210)에 포함된 사용자의 얼굴에 대하여 턱 주위 필러 시술된 이후의 가상의 얼굴을 포함하는 제1 가상 얼굴 이미지(3220)를 생성하거나 수신할 수 있고, 제1 가상 얼굴 이미지(3220)에 기초하여 턱 주위 필러 시술 후의 사용자의 나이를 23.3세로 추정할 수 있다. 또한, 나이 추정 장치(100)는 원본 이미지 데이터(3210)에 포함된 사용자의 얼굴에 대하여 외곽 부위 쉐이딩 이후의 가상의 얼굴을 포함하는 제2 가상 얼굴 이미지(3230)를 생성하거나 수신할 수 있고, 제2 가상 얼굴 이미지(3230)에 기초하여 외곽 부위 쉐이딩 후의 사용자의 나이를 22.9세로 추정할 수 있다.
나이 추정 장치(100)는 사용자에게 어떠한 화장, 시술 또는 수술 이후에 예상되는 나이를 제공할 수 있고, 나아가 추천 화장, 추천 시술 또는 추천 수술을 제안할 수 있다. 사용자에게 제공하는 추천 화장, 추천 시술 또는 추천 수술은 추정되는 나이가 가장 어리게 나타나는 화장, 시술 또는 수술일 수 있다.
나이 추정 장치(100)는 원본 이미지 데이터(3210)로부터 나이(예컨대, 25.1세)를 추정하고, 사용자(미도시)의 입력에 기초하여 추정된 나이 값을 조정하면 조정된 나이에 대응하는 가상의 얼굴을 포함하는 가상 이미지 데이터를 생성하여 제공할 수 있다. 또한, 나이 추정 장치(100)는 생성된 가상 이미지 데이터에 포함된 가상 얼굴에 이르기 위한 화장법, 시술, 수술, 건강 관리법 등을 추천할 수 있다. 추천하는 화장법에는 사용할 화장품의 종류, 추천 화장품 브랜ㄷ크 등이 포함될 수 있다. 추천하는 건강 관리법에는 추천 식단, 추천 식료품 정보, 추천 운동 루틴 등이 포함될 수 있다.
본 개시의 일 실시 예에 따르면, 전술한 방법은 프로그램이 기록된 매체에 컴퓨터가 읽을 수 있는 코드로서 구현하는 것이 가능하다. 컴퓨터가 읽을 수 있는 매체는, 컴퓨터 시스템에 의하여 읽혀질 수 있는 데이터가 저장되는 모든 종류의 기록장치를 포함한다. 컴퓨터가 읽을 수 있는 매체의 예로는, HDD(Hard Disk Drive), SSD(Solid State Disk), SDD(Silicon Disk Drive), ROM, RAM, CD-ROM, 자기 테이프, 플로피 디스크, 광 데이터 저장 장치 등이 있다.

Claims (18)

  1. 나이 추정 장치에 있어서,
    얼굴 특징점 추출 모델 및 나이 추정 모델을 저장하는 메모리; 및
    사용자에 대한 적어도 하나 이상의 얼굴 이미지를 수신하고, 상기 얼굴 특징점 추출 모델을 이용하여 상기 얼굴 이미지에서 복수의 얼굴 특징점을 추출하고, 상기 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하고, 상기 나이 추정 모델과 상기 복수의 얼굴 지표를 이용하여 상기 사용자의 나이를 추정하는 프로세서
    를 포함하는, 나이 추정 장치.
  2. 청구항 1에 있어서,
    상기 프로세서는
    상기 복수의 얼굴 특징점을 표준화하고, 상기 복수의 표준화된 얼굴 특징점으로부터 상기 복수의 얼굴 지표를 추출하는, 나이 추정 장치.
  3. 청구항 2에 있어서,
    상기 프로세서는
    상기 복수의 얼굴 특징점들 중에서 적어도 일부에 기초하여 얼굴의 미리 정해진 부위가 미리 정해진 좌표에 배치되도록 상기 얼굴 이미지의 변환 규칙을 결정하고, 상기 변환 규칙에 기초하여 상기 복수의 얼굴 특징점들을 표준화하는, 나이 추정 장치.
  4. 청구항 1에 있어서,
    상기 얼굴 특징점 추출 모델은
    인공 신경망(neural network)으로 구성된 딥 러닝 모델이고, 상기 얼굴 이미지가 입력되면 상기 입력된 얼굴 이미지에 대응하는 미리 정해진 개수만큼의 상기 복수의 얼굴 특징점을 추출하는, 나이 추정 장치.
  5. 청구항 1에 있어서,
    상기 프로세서는
    복수의 얼굴 이미지를 수신한 경우, 상기 복수의 얼굴 이미지 각각으로부터 상기 복수의 얼굴 특징점을 추출하고, 서로 대응하는 위치의 얼굴 특징점의 평균을 산출하여 복수의 평균 얼굴 특징점들을 산출하고, 상기 복수의 평균 얼굴 특징점들로부터 상기 복수의 얼굴 지표를 추출하는, 나이 추정 장치.
  6. 청구항 5에 있어서,
    상기 프로세서는
    상기 복수의 얼굴 이미지 중에서 아웃라이어(outlier) 얼굴 이미지를 결정하고, 상기 아웃라이어 얼굴 이미지를 제외한 얼굴 이미지로부터 상기 복수의 얼굴 특징점을 추출하는, 나이 추정 장치.
  7. 청구항 5에 있어서,
    상기 복수의 얼굴 이미지는
    미리 정해진 기간 이내에 촬영된 얼굴 이미지들인, 나이 추정 장치.
  8. 청구항 1에 있어서,
    상기 나이 추정 모델은
    선형 회귀 모델, 의사 결정 나무 또는 딥 러닝 모델이고, 적어도 상기 복수의 얼굴 지표가 입력되면 상기 추정된 나이를 출력하는 모델인, 나이 추정 장치.
  9. 청구항 8에 있어서,
    상기 프로세서는
    상기 복수의 얼굴 특징점, 상기 복수의 얼굴 지표 또는 상기 사용자의 신체 정보 중에서 적어도 하나 이상 및 상기 나이 추정 모델을 이용하여 상기 사용자의 나이를 추정하고,
    상기 신체 정보는
    상기 사용자의 신장, 상기 사용자의 체중 또는 상기 사용자의 체지방지수(BMI) 중에서 적어도 하나 이상을 포함하는, 나이 추정 장치.
  10. 청구항 9에 있어서,
    상기 복수의 얼굴 지표는
    얼굴 전체 면적, 입술 면적, 얼굴 관련 지표, 코 관련 지표 또는 입술 관련 지표 중에서 적어도 하나 이상을 포함하고,
    상기 얼굴 관련 지표는
    얼굴 윤곽선의 기울기를 포함하는, 나이 추정 장치.
  11. 청구항 1에 있어서,
    카메라
    를 더 포함하고,
    상기 프로세서는
    상기 카메라를 통해 상기 적어도 하나 이상의 얼굴 이미지를 수신하는, 나이 추정 장치.
  12. 청구항 1에 있어서,
    사용자 단말기와 통신하는 통신부
    를 더 포함하고,
    상기 프로세서는
    상기 통신부를 통해 상기 사용자 단말기로부터 상기 적어도 하나 이상의 얼굴 이미지를 수신하는, 나이 추정 장치.
  13. 청구항 1에 있어서,
    상기 프로세서는
    상기 추정된 나이를 이용하여 상기 사용자에 대한 나이 인증을 수행하는, 나이 추정 장치.
  14. 청구항 13에 있어서,
    상기 프로세서는
    상기 추정된 나이를 이용하여 나이 인증시, 상기 추정된 나이가 상기 인증에서 요구하는 나이보다 미리 정해진 마진만큼 크거나 같은 경우, 상기 추정된 나이가 상기 나이 인증에 적합하다고 판단하는, 나이 추정 장치.
  15. 청구항 1에 있어서,
    상기 프로세서는
    상기 추정된 나이를 상기 사용자에 대한 식별 정보로 생성하는, 나이 추정 장치.
  16. 청구항 1에 있어서,
    상기 프로세서는
    상기 추정된 나이에 대응하는 추천 컨텐츠를 결정하고, 상기 결정된 추천 컨텐츠를 제안하는, 나이 추정 장치.
  17. 나이를 추정하는 방법에 있어서,
    적어도 하나 이상의 얼굴 이미지를 수신하는 단계;
    얼굴 특징점 추출 모델을 이용하여 상기 얼굴 이미지에서 복수의 얼굴 특징점을 추출하는 단계;
    상기 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하는 단계; 및
    나이 추정 모델과 상기 복수의 얼굴 지표를 이용하여 상기 사용자의 나이를 추정하는 단계
    를 포함하는, 방법.
  18. 나이를 추정하는 방법을 기록한 기록 매체에 있어서, 상기 방법은
    적어도 하나 이상의 얼굴 이미지를 수신하는 단계;
    얼굴 특징점 추출 모델을 이용하여 상기 얼굴 이미지에서 복수의 얼굴 특징점을 추출하는 단계;
    상기 복수의 얼굴 특징점으로부터 미리 정해진 복수의 얼굴 지표를 추출하는 단계; 및
    나이 추정 모델과 상기 복수의 얼굴 지표를 이용하여 상기 사용자의 나이를 추정하는 단계
    를 포함하는, 기록 매체.
PCT/KR2020/019328 2020-05-18 2020-12-29 나이 추정 장치 및 나이를 추정하는 방법 Ceased WO2021235641A1 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020200059050A KR20210142342A (ko) 2020-05-18 2020-05-18 나이 추정 장치 및 나이를 추정하는 방법
KR10-2020-0059050 2020-05-18

Publications (1)

Publication Number Publication Date
WO2021235641A1 true WO2021235641A1 (ko) 2021-11-25

Family

ID=78708703

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2020/019328 Ceased WO2021235641A1 (ko) 2020-05-18 2020-12-29 나이 추정 장치 및 나이를 추정하는 방법

Country Status (2)

Country Link
KR (2) KR20210142342A (ko)
WO (1) WO2021235641A1 (ko)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120877355A (zh) * 2025-09-25 2025-10-31 浙江爱我科技有限公司 一种基于多视角图像的年龄及性别预测方法与应用

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102889841B1 (ko) * 2022-07-26 2025-11-25 주식회사 용진 메타버스를 이용한 쉼터 융합 시스템
KR20260038628A (ko) * 2024-09-12 2026-03-19 주식회사 인바디 사용자의 체성분을 결정하는 전자 장치 및 그 동작 방법

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06333023A (ja) * 1993-05-26 1994-12-02 Casio Comput Co Ltd 年齢推定装置
JP2008003749A (ja) * 2006-06-21 2008-01-10 Fujifilm Corp 特徴点検出装置および方法並びにプログラム
KR20130029723A (ko) * 2011-09-15 2013-03-25 가부시끼가이샤 도시바 얼굴 인식 장치 및 얼굴 인식 방법
KR20150089370A (ko) * 2014-01-27 2015-08-05 주식회사 에스원 얼굴 포즈 변화에 강한 연령 인식방법 및 시스템
JP2016004149A (ja) * 2014-06-17 2016-01-12 カシオ計算機株式会社 情報処理装置、情報処理システム、コンテンツ出力方法、及びプログラム
KR101895001B1 (ko) * 2017-07-13 2018-09-05 (주)블루오션소프트 스마트폰기반 쌍방향 라이브방송 커머스 서비스 플랫폼

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06333023A (ja) * 1993-05-26 1994-12-02 Casio Comput Co Ltd 年齢推定装置
JP2008003749A (ja) * 2006-06-21 2008-01-10 Fujifilm Corp 特徴点検出装置および方法並びにプログラム
KR20130029723A (ko) * 2011-09-15 2013-03-25 가부시끼가이샤 도시바 얼굴 인식 장치 및 얼굴 인식 방법
KR20150089370A (ko) * 2014-01-27 2015-08-05 주식회사 에스원 얼굴 포즈 변화에 강한 연령 인식방법 및 시스템
JP2016004149A (ja) * 2014-06-17 2016-01-12 カシオ計算機株式会社 情報処理装置、情報処理システム、コンテンツ出力方法、及びプログラム
KR101895001B1 (ko) * 2017-07-13 2018-09-05 (주)블루오션소프트 스마트폰기반 쌍방향 라이브방송 커머스 서비스 플랫폼

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120877355A (zh) * 2025-09-25 2025-10-31 浙江爱我科技有限公司 一种基于多视角图像的年龄及性别预测方法与应用

Also Published As

Publication number Publication date
KR20210142342A (ko) 2021-11-25
KR20220151596A (ko) 2022-11-15

Similar Documents

Publication Publication Date Title
WO2020235852A1 (ko) 특정 순간에 관한 사진 또는 동영상을 자동으로 촬영하는 디바이스 및 그 동작 방법
WO2018117704A1 (en) Electronic apparatus and operation method thereof
WO2021230680A1 (en) Method and device for detecting object in image
WO2019132518A1 (en) Image acquisition device and method of controlling the same
WO2018128362A1 (en) Electronic apparatus and method of operating the same
WO2020050595A1 (ko) 음성 인식 서비스를 제공하는 서버
WO2014025185A1 (en) Method and system for tagging information about image, apparatus and computer-readable recording medium thereof
WO2015102361A1 (ko) 얼굴 구성요소 거리를 이용한 홍채인식용 이미지 획득 장치 및 방법
WO2021025509A1 (en) Apparatus and method for displaying graphic elements according to object
WO2022071695A1 (ko) 영상을 처리하는 디바이스 및 그 동작 방법
WO2021235641A1 (ko) 나이 추정 장치 및 나이를 추정하는 방법
EP3545436A1 (en) Electronic apparatus and method of operating the same
WO2016117836A1 (en) Apparatus and method for editing content
WO2019182378A1 (en) Artificial intelligence server
WO2019124963A1 (ko) 음성 인식 장치 및 방법
EP3539056A1 (en) Electronic apparatus and operation method thereof
WO2017043857A1 (ko) 어플리케이션 제공 방법 및 이를 위한 전자 기기
WO2021006482A1 (en) Apparatus and method for generating image
WO2021256781A1 (ko) 영상을 처리하는 디바이스 및 그 동작 방법
WO2022114731A1 (ko) 딥러닝 기반 비정상 행동을 탐지하여 인식하는 비정상 행동 탐지 시스템 및 탐지 방법
WO2021132798A1 (en) Method and apparatus for data anonymization
WO2017099314A1 (ko) 사용자 정보를 제공하는 전자 장치 및 방법
WO2022250388A1 (ko) 비디오 품질을 평가하는 전자 장치 및 그 동작 방법
WO2016021907A1 (ko) 웨어러블 디바이스를 이용한 정보처리 시스템 및 방법
WO2022045613A1 (ko) 비디오 품질 향상 방법 및 장치

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20936631

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 13.04.2023)

122 Ep: pct application non-entry in european phase

Ref document number: 20936631

Country of ref document: EP

Kind code of ref document: A1