EP4710306A1 - Training data for image morphing detection - Google Patents

Training data for image morphing detection

Info

Publication number
EP4710306A1
EP4710306A1 EP24728518.2A EP24728518A EP4710306A1 EP 4710306 A1 EP4710306 A1 EP 4710306A1 EP 24728518 A EP24728518 A EP 24728518A EP 4710306 A1 EP4710306 A1 EP 4710306A1
Authority
EP
European Patent Office
Prior art keywords
image
locations
landmarks
images
morphed
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24728518.2A
Other languages
German (de)
French (fr)
Inventor
Andreas Dr. WILKE
Oliver DR. MUTH
Lars DR. SIMON
Holger DR. EBLE
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Bundesdruckerei GmbH
Original Assignee
Bundesdruckerei GmbH
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Bundesdruckerei GmbH filed Critical Bundesdruckerei GmbH
Publication of EP4710306A1 publication Critical patent/EP4710306A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/172Classification, e.g. identification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N10/00Quantum computing, i.e. information processing based on quantum-mechanical phenomena
    • G06N10/60Quantum algorithms, e.g. based on quantum optimisation, quantum Fourier or Hadamard transforms
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/10Machine learning using kernel methods, e.g. support vector machines [SVM]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/098Distributed learning, e.g. federated learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/774Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/40Spoof detection, e.g. liveness detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F21/00Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
    • G06F21/30Authentication, i.e. establishing the identity or authorisation of security principals
    • G06F21/31User authentication
    • G06F21/32User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Evolutionary Computation (AREA)
  • Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Medical Informatics (AREA)
  • Computational Linguistics (AREA)
  • Databases & Information Systems (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Human Computer Interaction (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Molecular Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Pure & Applied Mathematics (AREA)
  • Computational Mathematics (AREA)
  • Condensed Matter Physics & Semiconductors (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Image Analysis (AREA)

Abstract

Disclosed is a method of generating a training dataset (500) for training a machine learning model for morphed image detection. The method comprises: repeatedly performing the following: receiving (101) an image (130, 150) of an object (134, 154); identifying (103) a set of landmarks of the object (134, 154) in accordance with a landmark pattern; determining (105) locations of the set of landmarks with respect to a coordinate system defined relative to the object (134, 154), and adding (107) an entry to the training dataset (500), the entry indicating the set of locations and a label, wherein the label indicates whether the received image (130, 150) is a real image or morphed image.

Description

TRAINING DATA FOR IMAGE MORPHING DETECTION
FIELD OF THE INVENTION
[0001] The disclosure relates to image morphing, and particularly to a method for generating training data for image morphing detection.
BACKGROUND
[0002] Morphing may be a special effect in motion pictures and animations that morphs one image into another through a seamless transition. Computer software may be used to create morphed images. However, there is a need for an improved processing of the morphed images.
[0003] Document “Towards Detection of Morphed Face Images in Electronic Travel Documents” published in 2018 13th IAPR International Workshop on Document Analysis Systems (DAS) discloses automated morph detection algorithms based on general purpose pattern recognition algorithms.
SUMMARY
[0004] Example embodiments provide a method of generating a training dataset for training a machine learning model for morphed image detection. The method comprises: repeatedly performing the following: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; and adding an entry to the training dataset, the entry indicating the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image.
[0005] Example embodiments provide a computer system of generating a training dataset for training a machine learning model for morphed image detection. The computer system is configured for: repeatedly performing the following: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; and adding an entry to the training dataset, the entry indicating the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image.
[0006] Example embodiments provide a computer program comprising instructions for causing a computer system for performing at least the following: repeatedly performing the following: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; and adding an entry to the training dataset, the entry indicating the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image.
[0007] Example embodiments provide a computer implemented data structure comprising training data for a morphed image detection model, the data structure comprising entries, wherein the entry comprises location information indicating a set of landmark locations of an imaged object, and a label indicating a real image or morphed image, the landmark locations being relative locations.
[0008] Example embodiments provide a method for morphed image detection. The method comprises: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting location information to a trained machine learning model, the location information indicating the set of locations; receiving an output of the trained machine learning model indicating whether the received image is a morphed image or non-morphed image.
[0009] Example embodiments provide a computer system for morphed image detection. The computer system is configured for: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting location information to a trained machine learning model, the location information indicating the set of locations; receiving an output of the trained machine learning model indicating whether the received image is a morphed image or non-morphed image.
[0010] Example embodiments provide a computer program comprising instructions for causing a computer system for performing at least the following:: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting location information to a trained machine learning model, the location information indicating the set of locations; receiving an output of the trained machine learning model indicating whether the received image is a morphed image or non-morphed image.
BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the following, examples are described in greater detail making reference to the drawings in which:
[0012] Fig. 1 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter.
[0013] Fig. 2A is a diagram illustrating a method for determining locations of landmarks in an image of a human face in accordance with an example of the present subject matter.
[0014] Fig. 2B is a diagram illustrating a method for determining locations of landmarks in an image of a human eyes in accordance with an example of the present subject matter.
[0015] Fig. 3 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter. [0016] Fig. 4 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter.
[0017] Fig. 5 is a flowchart of a method of determining landmarks of an imaged object in accordance with an example of the present subject matter.
[0018] Fig. 6 is a diagram of a data structure representing the training dataset generated in accordance with an example of the present subject matter.
[0019] Fig. 7 is a block diagram of an exemplary computer system for implementing at least part of the present method in accordance with an example of the present subject matter.
[0020] Fig. 8 is a flowchart of a method for training a classical machine learning model in accordance with an example of the present subject matter.
[0021] Fig. 9 is a flowchart of a method for training a quantum machine learning model in accordance with an example of the present subject matter.
[0022] Fig. 10 is a flowchart of a method for training a hybrid classical-quantum machine learning model in accordance with an example of the present subject matter.
[0023] Fig. 11 is a flowchart of a method for morphed image detection in accordance with an example of the present subject matter.
DETAILED DESCRIPTION
[0024] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, interfaces, techniques, etc., in order to provide a thorough understanding of the examples. However, it will be apparent to those skilled in the art that the disclosed subject matter may be practiced in other illustrative examples that depart from these specific details. In some instances, detailed descriptions of well-known devices and/or methods are omitted so as not to obscure the description with unnecessary detail. [0025] Authentication may be a process for verifying the identity of persons. This may, for example, enable to keep unauthorized persons from accessing sensitive information or services. For example, based on the authentication result of a user, access control signals may be generated to enable the user to access controlled services or controlled areas. The authentication may be performed using an image of the user. For example, an automated border control system or eGate may use the image to verify the user's identity. After the user's identity is verified, a physical barrier such as a gate opens to permit passage. The image may provide a visual representation of the user. The image of the user may be obtained by imaging the user’s face e.g., using a camera. Alternatively, the image of the user may be a reproduced version that is captured from an identity token of the user. The image may be stored in a memory as a digital image. The identity token may, for example, be an identification (ID) card, visa, driver's license, vehicle registration document, health card, company ID card, bank card, or another ID document which carries an identity-linked field comprising a photo of the user. In another example, the identity token may be provided in the form of an ID application installed on a user's mobile portable terminal. In this case, the user's image may be provided, for example, on a display of the user's mobile portable terminal.
[0026] However, the images may be morphed. Morphing refers to the changing of one image to obtain a morphed image. The morphing may, for example, modify an image of a given object such that the resulting object resembles the given object. The morphing may be misused to, for example, create a double-identity face image, a double-identity fingerprint or a double-identity iris, or to perform cyberattacks by creating hoax images. For example, in case of face morphing, a misuse of the morphing can enable two individuals to use one identity document. In another example, a fake fingerprint may be created using morphing so that it can be used to identify two different fingers. Thus, the detection of such morphed images may prevent cyberattacks and fake identities. Different techniques may be used to detect morphed images. For example, machine learning models may be used to detect morphed images. The present subject matter may provide optimal training data for training these models. The resulting trained models may have a higher detection efficiency of morphed image. The efficiency may be defined as the ration of the number of detected morphed images and the total number of morphed images.
[0027] For that, a training dataset may be generated in order to train a machine learning model to detect morphed images. The training dataset may be created using a set of images. The image may, for example, be a visual representation of an object or subject. The image of the object may be obtained by imaging the object e.g., using a camera. Alternatively, the image of the object may be a reproduced version that is captured from a document comprising the image of the object. For each image of the set of images, a set of landmarks of the object that is represented in the image may be identified. For example, the set of landmarks may be identified in the image in accordance with a landmark pattern. The landmark pattern may be a predefined landmark pattern or a dynamically created landmark pattern. The landmark may refer to a specific point of the object. The locations of the set of landmarks may, for example, be determined with respect to a coordinate system defined relative the object. The locations of the set of landmarks may, thus, be relative locations. For each image of the set of images, an entry may be included in the training dataset. The entry may comprise a location information (which may be referred to as Lb where the subscript i refers to the image) indicating the set of locations and a label indicating whether the image is a real image (i.e., nonmorphed image) or morphed image.
[0028] For example, each entry of the training dataset may comprise a tuple (Lb is the location information determined for the i-th image of the set of images and the labelt is the label of the i-th image. In one example, the training dataset may be stored in a storage system of the computer system. The computer system may control access to the training dataset. For example, the computer system may define permissions, such as read permission and execute permissions, for access to the training dataset. One or more users may use the training dataset based on permissions which are assigned to the users.
[0029] Hence, instead of using a whole image as a training entry, only a selected set of landmarks is used to represent the image. In addition, the landmarks are located relative to the object using a local coordinate system. The training data may, thus, enable a faster training of machine learning models. The resulting trained model may efficiently detect morphed images.
[0030] The object that is represented in each image of the set of images may refer to a tangible, physical object capable of being rendered in an image. The object may be any object whose visual representation in an image may be morphed. The object may be of a specific object type. In one example, the object type may be an individual, a body part, a human face, human eyes, fingerprints, animal face etc. Covering different object types may enable a wider application of the present subject matter e.g., the resulting trained model may be used in face recognition systems, fingerprint recognition systems, and iris recognition systems to detect morphed images.
[0031] The present subject matter may advantageously control the number and types of the objects in the set of images in order to find a desired balance between efficiency of the resulting trained model and the extent of application of the trained model. For example, the set of images used to generate the training dataset may be a homogeneous set of images or a heterogeneous set of images. The homogenous set of images may represent objects of the same object type, while the heterogeneous set of images may represent objects of different object types.
[0032] The homogeneous set of images may be advantageous for the following reasons. The homogeneous set of images may enable a systematic and faster generation of the training dataset compared to images of different object types e.g., with the homogenous set of images a smaller number e.g., one, of landmark patterns may be sufficient to find the landmarks in all set of images. The homogeneous set of images may provide homogeneous training samples, wherein the homogeneous training samples may have higher affinity among them enabling a fast convergence of the training process when applied on the created training dataset. The homogeneous set of images may enable a trained model that is more efficient in detecting any other morphed image of the object type represented by the set of images. In one example, the homogenous set of images may be images of one object type e.g., the set of images may be images of human faces, wherein the human faces may be of a same or different individuals. [0033] The heterogeneous set of images may be advantageous for the following reasons. The heterogeneous set of images may provide a larger training dataset due to the availability of a higher number of images. The heterogeneity may allow to build a robust learning system by leveraging the intrinsic knowledge among data. Indeed, with a set of images of different object types, the model may learn different morphing techniques, and may thus enable a morphing detection that works well not only for one specific type of morphing (e.g., face morphing) but also for other types of morphing. The heterogeneous set of images may thus enable a wider application of the present subject matter while still providing reliable detection results. In one example, the heterogeneous set of images may be images of different object types. The different object types may be human face, human eyes, animal face, fingerprints etc.
[0034] The present subject matter may advantageously control the set of images in order to balance the training dataset with respect to a set of image attributes. The set of image attributes may, for example, comprise at least one of lighting conditions, background color, skin tone, facial expression and any other attribute descriptive of an image or of the object represented in the image. Each image of the set of images may have a specific set of values of the set of image attributes respectively. For example, the set of images may comprise multiple subsets of images, each subset having a distinct set of values of the set of image attributes. For example, the number of subsets of the images may be higher than a threshold. This may enable to control the level of diversity of the set of images. For example, the first subset images may have a first set of values of the set of image attributes, the second subset images may have a second set of values of the set of image attributes, and so forth. Each pair of sets of values (e.g., the first set and second set of values) of the image attributes may differ in at least one image attribute. Using different attribute values in the set of images may enable landmarks which are independent from background, lighting conditions, skin tone etc. This may significantly reduce the amount of data necessary to train the machine learning model. According to one example, the set of images may be provided as the homogeneous set of images comprising the multiple subsets. Alternatively, the set of images may be provided as the heterogeneous set of images comprising the multiple subsets.
[0035] The set of images may, for example, comprise images of different resolutions. In one example, the set of images may comprise scanned images of the objects and/or images of the objects which are captured directly from the objects e.g., by a digital camera. For example, the set of images may comprise a minimum fraction of images which are scanned images. A scanned image may, for example, be obtained using a photo scanner or a camera for capturing an image in an identity token. The scanned images may provide accurate detection in spite of containing only smaller part of data. The resulting trained machine learning model may, for example, be useful for passport image verification during the passport application process or during check of the passports.
[0036] The present subject matter may control the process of generation of the training dataset by using different access methods to access the set of images. In one example, at least part of the set of images (named retrieved images) may be retrieved or received from one or more existing database systems. For example, the retrieved images may be the whole set of images or a subset of the set of images. The computer system may be configured to connect to the database systems and request or retrieve the at least part of the set of images. This may speed up the generation of the training data compared to a local generation of the set of images. Additionally, or alternatively, at least part of the set of images (named produced images) may be produced locally by the computer system. This may save processing resources such as the network resources that would otherwise be required to retrieve images. For example, the produced images may be the whole set of images or a subset of the set of images. Thus, the set of images may comprise retrieved images and/or produced images. In one example, the produced images may comprise morphed images and/or non-morphed images. In one example, the retrieved images may comprise morphed images and/or non-morphed images. This example may provide a flexible and controllable access to the images e.g., if one access method is not available, the present subject matter can still use alternative access methods to produce the training data. This may improve the process of generation of the training dataset.
[0037] The produced images of the set of images may, for example, be obtained as follows. The computer system may receive non-morphed images. A subset of the received non-morphed images may be used to generate morphed images by the computer system in order to obtain said produced images. The morphed images may, for example, be generated using one or more morphing algorithms. The morphing algorithm may, for example, be a Generative Adversary Networks (GANs) or landmark-based morphing algorithm. The morphing algorithm may, for example, be a face morphing algorithm that extracts feature points on the face, and based on these feature points images are partitioned and face morphing is performed.
[0038] In one example, a pre-processing of the set of images may be performed before using the pre-processed images to generate the training dataset. The preprocessing may, for example, comprise the reduction of the size of the set of images. This may save processing resources required for determining the set of landmarks. The set of images may be pre-processed in order to have one size. In addition, pixels of each image of the set of images may be represented by a vector (e.g., tensor). The vector may have fields, wherein the fields may include the pixel width, the pixel height, and pixel value(s) such as red, green, blue (RGB) values. This may enable a uniform representation of the set of images.
[0039] In one example, at least part of the set of images may be processed in parallel in order to generate the training dataset. This may speed up the process of generating the training dataset. Additionally, or alternatively, at least part of the set of images may be processed sequentially. This may enable a simplified implementation of the image processing. In one example, the set of images may be processed in batches, wherein each batch may have a size smaller than a maximum size. For example, the training dataset may be produced in a distributed computing system e.g., each system component of the distributed computing system may process its respective batches of images, and the resulting training data entries may be combined in one training dataset. [0040] Hence, as described above, the present subject matter may use different techniques to provide and process the set of images for identification of landmarks. In addition, the present subject matter may provide different techniques to improve the determination of the landmarks.
[0041] For example, a set of landmarks may be used to represent each image of the set of images. The landmark may refer to a specific point of the object. The set of landmarks may be defined as points bearing key information on the geometry of the object. In case the object is a human face, for example, the landmark may be a right eyebrow lateral point, right eyebrow medial point, left eyebrow lateral point, left eyebrow medial point, right eye lateral canthus etc. The set of landmarks may be identified using a landmark pattern. The landmark pattern may indicate features of the object which enable to identify the object and its structure. The landmark pattern may indicate a maximum number of landmarks and/or the landmarks per feature and/or the densities of landmarks and/or distances between landmarks. The landmark pattern may be a user defined pattern or may be defined by a computer- implemented tool. For example, the landmark pattern may be determined using distinctive features of the object in the image. In case the object is a human face, the distinctive features may include eye spacing, nose length, mouth width, head eccentricity etc.
[0042] In one example, a landmark pattern may be provided per object type of the set of images. For example, if the set of images is a homogeneous set of images, one landmark pattern may be used to generate the set of landmarks of each image of the set images. For example, one landmark pattern may be predefined for the whole set of images so that during the generation of the training dataset the landmark pattern may automatically be used to identify the set of landmarks in each image of the set of images. For example, if the set of images is a heterogeneous set of images, a distinct landmark pattern may be used per object type to generate the set of landmarks of each image of that object type. For example, multiple landmark patterns may be predefined for the set of images. For each received image of the set of images, the identification of the set of landmarks in the image may be performed by: determining the object type represented by the image, selecting from the predefined landmark patterns the landmark pattern associated with the determined object type and using the selected landmark pattern to identify the set of landmarks in the image.
[0043] In one example, the object represented in each image of the set of images may be detected in the image. The object may, for example, be detected using a computer vision technique that identifies and locates objects within an image. An area of the image defined by the detected object may be cropped. This may result in a cropped image. The set of landmarks are then identified in the cropped image and the locations are determined using the cropped image. This may enable to determine the landmarks after the object is detected and cropped. Thus, the landmarks may not depend on the location of the object within the image or on the distance of the object to the camera. This may significantly reduce the amount of data necessary to train the machine learning model.
[0044] The present subject matter may use different advantageous techniques to provide the landmark pattern(s). In one example, for each object type in the set of images a reference image of the set of images that represents the object type may be selected. The selected reference image may be a randomly selected image or an image whose image attribute values fulfil a predefined selection criterion. The selection criterion may require that the value of each image attribute of the set of image attributes has a specific value or is within a specific range of values. The landmark pattern(s) may be determined using the respective reference image(s). In one example, a feature extraction tool may be used to extract the features that identify the object structure from the reference image. The landmark pattern may be defined by assigning to each feature of the extracted features zero or more landmarks that would represent the feature. The landmark pattern may indicate the position of the landmarks with respect to the respective feature e.g., it may indicate 10 landmarks surrounding the eye of a human face etc. This example may enable to define a priori landmark patterns which may be used during the generation of the training dataset. [0045] In another example, a trained machine learning model may be used to create the landmark pattern that identifies the structure of the object in each image of the set of images. This may enable to create the landmark pattern dynamically or automatically during the generation of the training dataset. In this example, the creation of the landmark pattern may enable an automatic identification of the landmarks in the image e.g., the creation of the landmark pattern implicitly includes the step of identification of the landmarks.
[0046] The set of locations (or positions) of each identified set of landmarks may be determined. In one example, the steps of identifying the set of landmarks of an image and the determination of their locations may be performed, e.g., in one step, concurrently or in parallel. This may speedup the generation of the training dataset. Alternatively, the set of locations of each identified set of landmarks may be determined after the set of landmarks is identified.
[0047] The set of locations of the set of landmarks of each image of the set of images may be provided as relative locations e.g., which do not depend on the image size. For that, a coordinate system which is defined relative to the object in the image may be used to determine the locations of the set of landmarks. The coordinate system may be a local coordinate system associated with the object. The coordinate system may refer to a frame of reference defined by orthogonal directions and an origin. The origin may be a reference point that is part of the object. The origin may be used as a fixed point of reference for the geometry of the surrounding space. In case more than one object type is covered by the set of images, the local coordinate system may be defined per object type. Using a fixed point per object type as origin may provide locations which accurately reflect the shape of the object. The location of each landmark may be provided as coordinates relative to the point of origin. The location may be a direction vector in the local coordinate system. The location of the landmark may, for example, refer to a three- dimensional measurement from a position of the landmark to the origin.
[0048] In another example, initial locations of the set of landmarks of an image may be determined in a first coordinate system (e.g., an absolute coordinate system). In addition, a transformation of the first coordinate system to the local coordinate system may be applied to the initial locations. This may result in the locations of the set of landmarks of the image which are determined in the local coordinate system.
[0049] In one example, the local coordinate system may be a three-dimensional, 3D, coordinate system. The set of landmark locations may be provided as 3D coordinates. The set of locations of the set of landmarks may thus describe the three-dimensional geometry of the object. This may provide an accurate position of the landmarks and accurate distinction between the set of landmarks in three dimensions. This may enable more efficient morphed image detection, because the morphing algorithms may generate specific patterns in the three-dimensional geometry of the thusly created objects.
[0050] In one example, the local coordinate system may be a two-dimensional, 2D, coordinate system. The set of landmark locations may thus be provided as 2D coordinates. Compared to 3D coordinates, this may save processing resources such as storage resources while still providing reliable results. For example, each entry of the training dataset may comprise the tuple (3D coordinates, label) or (2D coordinates, label).
[0051] Hence, for each i-th image of the set images, a set landmarks may be identified and a set of locations LOC , LOC"1' may be determined respectively, where > 2. Each entry of the resulting training dataset may comprise a tuple (Lh label ), where Lt = {LOC , LOC"1') is the location information determined for the i-th image and which contains the set of locations determined for the i-th image.
[0052] The present subject matter may further reduce the size of the input data while still providing accurate representation of the input data. For that, for each image of the set of images, the set of locations of the set of landmarks of the image may be represented by a feature vector in a predefined fc-dimensional feature space having the dimension k which is smaller than the number of landmarks identified in the image. In this case, the location information of each entry of the training dataset may be a feature vector. The feature space may refer to a fc-dimensional space spanned up by k different features used to characterize the set of landmark locations e.g., a feature may be the number of locations, location density etc. In one example, the k different features may be determined, and for each image of the set of images, the set of locations of the set of landmarks of the image may be used to evaluate the k different features, and the resulting evaluations may be provided as the feature vector. The k different features may be user determined features or automatically determined features e.g., using machine learning techniques.
[0053] In one example, the feature vectors may be determined using transfer learning. The transfer learning may be used to apply knowledge gained while solving another task which is related to the morphing detection task. For example, knowledge gained while learning to recognize the objects may be applied when trying to detect morphing of the objects.
[0054] In one example, the set of locations of the set of landmarks of the image may be represented with the feature vector using a linear transformation of a vector (y) representing the set of locations into the feature vector (x), using a weight matrix ( W), where y = x x W. The linear transformation may, for example, be user defined e.g., the weight matrix may be a user defined matrix. Alternatively, the weight matrix may comprise learnable weights which may be provided using the transfer learning. For example, another machine learning model having the weight matrix 1/1/ as a trainable weight matrix may have been trained to generate the feature vector from a set of landmark locations (e.g., for object recognition). The weight matrix l/l/may thus be transferred from this other machine learning model in order to be used with the present example.
[0055] In one example, the set of landmark locations of the set of landmarks of the image may be represented with the feature vector using a trained neural network (referred to herein as first neural network) comprising a fully connected layer having nodes representing the set of locations and an output layer representing the feature vector. The first neural network is configured to receive the set of locations of the set of landmarks of the image and to output the feature vector. [0056] In one example, the size k of the feature vector that represents the object of each image of the set of images may be defined based on the number of qubits in a quantum processing unit (QPU) that enables to train a quantum machine learning model to detect morphed images. The size of the feature vector may, for example, be provided based on the encoding scheme used by the quantum machine learning model for encoding classical data into quantum states. For example, the size k of the feature vector may be equal to the number of qubits which are available in the quantum processing unit. This may particularly be advantageous in case the angle encoding scheme is used by the quantum machine learning model for encoding classical data into quantum states. Alternatively, the size k of the feature vector may be smaller than the number of qubits which are available in the quantum processing unit. This may particularly be advantageous in case the amplitude encoding scheme is used by the quantum machine learning model for encoding classical data into quantum states.
[0057] Hence, for each i-th image of the set of images a feature vector that represents the object in the image may be provided. The feature vector comprises k feature values F*, .... F-F Each entry of the resulting training dataset may comprise a tuple (Lh label ), where = {F^, F } is the location information determined for the i-th image and which contains the feature vector determined for the i-th image. The size of the present entry may be smaller than the entry of the previous example as the size of the feature vector {F^, F } may be smaller than the size of the set of locations {LO 1, ... ., LOC™1}.
[0058] In one example, the training dataset may be provided as a data structure. The data structure may, for example, comprise a table comprising a column representing the label and one or more columns representing the location information. Alternatively, the data structure may be an Extensible Markup Language (XML) file or JavaScript Object Notation (JSON) file.
[0059] The entries of the data structure may, for example, be represented as follows (e.g., in a table): (Li, labels),
(L2, label2),
(L3, label3), or Lt = {LOC , LOG™1}, that is the location information may be provided as feature vector or as a set of locations.
[0060] After being produced, the training dataset may advantageously be used to train a machine learning model for image morphing detection. In one example, the machine learning model may comprise a quantum machine learning model or a classical machine learning model or a hybrid classical-quantum machine learning model. The quantum machine learning model may, for example, be a quantum neural network (QNN), a quantum convolutional neural network (QCNN), or a quantum support vector machine (QSVM). The classical machine learning model may, for example, a deep neural network such as a CNN or a support vector machine. The document arXiv:2009.09423 provides an example implementation of CNNs in a quantum environment for image classification.
[0061] The resulting trained machine learning model may for example be used as follows. The trained machine learning model may receive an image, generate the location information of the image (e.g., the location information may be a feature vector), process the location information, and output a value indicating whether the received image is a morphed image or not.
[0062] The resulting trained machine learning model may be stored in the computer system and/or in one or more other systems. If stored in the computer system only, the computer system may be configured to receive a request from a remote system to classify an image as being morphed or not, the computer system may generate a feature vector representing the object in the received image, and the feature vector may be input to the trained machine learning model in order to obtain a value indicating whether the image is morphed or not. The resulting value (prediction) may be sent to the remote system.
[0063] The quantum machine learning model may comprise an encoding layer. The encoding layer may be configured to encode the feature vector into a quantum state using a set of qubits. The quantum state may refer to a mathematical entity that provides a probability distribution for the outcomes of each possible measurement on the system of the set of qubits. This may, for example, be performed using an encoding scheme, wherein the encoding scheme may be an angle encoding or amplitude encoding. The encoding of the feature vector into a quantum state may, for example, be done using single qubit Pauli rotation gates (R_X, R_Y, R_Z) which are applied to individual qubits. The encoding layer may or may not entangle the resulting quantum state. That is, the encoding layer may provide for a given feature vector a quantum state which may be an entangled quantum state or unentangled quantum state. The entangling may not involve any input data or trainable parameters as it may be implemented using multi-qubit entangling gates. The entangling may enable to access a higher dimensional state space.
[0064] The quantum machine learning model may further comprise a learning layer having one or more trainable or free parameters. The learning layer may be configured to change the quantum state by applying one or more unitary transformations. The learnable parameters may, for example, be the rotation angles of single qubit Pauli rotation gates e.g., the Pauli rotation angle may be applied on each qubit of the set of qubits after the quantum state has been created. The learnable parameters may, for example, comprise a number of rotation angles which is applied to the set of qubits respectively.
[0065] The quantum machine learning model may further comprise a measurement layer for measuring the set of qubits after the change is applied. This set of measurements may provide an indication whether the image represented by the feature vector is morphed or not morphed. A loss function may be evaluated using the set of measurements and the label of the image. The quantum machine learning model may be trained by backpropagation using the loss function and an optimization technique that is performed by a classical computer. The backpropagation may enable to update of the learnable parameters using gradient descent. The convergence criterion may, for example, require that the loss function exceeds a threshold. The trained quantum machine learning model may be used to determine whether an image is a morphed image or not.
[0066] In case the feature vector is provided by the first neural network as described herein, the quantum machine learning model and the first neural network may be jointly trained. For example, in each iteration of the training, a set of locations may be input to the first neural network to generate the feature vector, the feature vector is provided as input to the quantum machine learning model, and the resulting set of measurements may provide an indication whether the image represented by the feature vector is morphed or not morphed. A loss function may be evaluated using the set of measurements and the label of the image. In case the loss function does not fulfill a convergence criterion, the backpropagation is performed in order to update both the learnable parameters of the quantum machine learning model as well as the weights of the first neural network. The update of the first neural network weights and the learnable parameters may be performed using gradient descent. In case the loss function fulfills the convergence criterion, the resulting trained first neural network and quantum machine learning model may be provided. The convergence criterion may, for example, require that the loss function exceeds a threshold. The trained first neural network and quantum machine learning model may be used to determine whether an image is a morphed image or not.
[0067] In another example, the training of the quantum machine learning model may be performed jointly with the first neural network and a second neural network, wherein the second neural network is configured to receive as input the set of measurements and to provide as output a value indicating whether the image is morphed or not. For example, in each iteration of the training, a set of locations may be input to the first neural network to generate the feature vector which is provided as input to the quantum machine learning model which provide in turn the set of measurements as input to the second neural network, the second neural network provides a value indicating whether the image is morphed or not. A loss function may be evaluated using the set of measurements and the label of the image. In case the loss function does not fulfill a convergence criterion, the backpropagation is performed in order to update the learnable parameters of the quantum machine learning model the weights of the first neural network and the weights of the second neural network. In case the loss function fulfills the convergence criterion, the resulting trained first neural network, second neural network and quantum machine learning model may be provided. The convergence criterion may, for example, require that the loss function exceeds a threshold. The trained first neural network and quantum machine learning model and second neural network may be used to determine whether an image is a morphed image or not.
[0068] The present subject matter may, for example, use the trained machine learning model for determining whether an image is a morphed image or not morphed image. For that, an image of an object may be received. The image may, for example, be captured from an identity token or may be read from a storage device where the image is stored. A set of landmarks of the object may be identified in accordance with a landmark pattern. Locations of the set of landmarks may be determined in step with respect to a coordinate system defined relative to the object. Location information may be input to the trained machine learning model. The location information indicates the set of locations. An output of the trained machine learning model may be received from the trained machine learning model. The output indicates whether the received inference image is a morphed image or nonmorphed image.
[0069] In one example, the output of the trained model may be used to perform authentication of a user. If the image of the user is not morphed this may indicate that the user’s identity is authentic. The authenticated user may be allowed to get access to services. The computer system may, for example, be configured to generate a control signal for enabling access to the services by the authenticated user.
[0070] In one example application, an electric gate may be provided. The electric gate comprises an electric gate motor that enables it to automatically open and close. If the output of the trained model indicates that the user is authenticated, the user may be enabled access to an area by sending a control signal to the electric gate motor to open the electric gate. The electric gate may, for example, be a sliding door.
[0071] It is understood that one or more of the aforementioned examples may be combined as long as the combined embodiments are not mutually exclusive.
[0072] Fig. 1 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter. The training dataset may be generated using a set of images.
[0073] An image of an object may be received in step 101 . For example, the image may be received from a local storage of the computer system. Alternatively, the image may be received from a remote database system in which the image is stored e.g., step 101 may be performed in response to sending a request to the database system.
[0074] A set of landmarks of the object may be identified in step 103 in accordance with a landmark pattern.
[0075] Locations of the set of landmarks may be determined in step 105 with respect to a coordinate system defined relative to the object. Figs. 2A and 2B provide an example implementation of step 105.
[0076] An entry may be added in step 107 to the training dataset. The entry indicates the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image. The real image refers to a nonmorphed image.
[0077] As indicated in Fig. 1 , steps 101 to 107 may be repeated for each further image of the set of images. The repetition may be performed until a stopping criterion is fulfilled. The stopping criterion may require that all the set of images are processed or that a maximum number of repetitions is reached. [0078] In one example implementation of Fig. 1 , steps 103 and 105 may be performed concurrently, e.g., the location of the landmark may be determined in response to identifying the landmark.
[0079] In one example implementation of Fig. 1 , the method may be repeated for each further set of images, wherein the training dataset is updated/increased with new entries obtained for each further set of images.
[0080] Fig. 2A is a diagram illustrating a method for determining locations of landmarks in an image 130 of a human face 134 in accordance with an example of the present subject matter. For that, a local coordinate system 131 may be used. Fig. 2A further shows an absolute coordinate system 135. The absolute coordinate system 135 may be a world coordinate system.
[0081] The local coordinate system 131 may be defined by an origin 132 and three orthogonal directions. The origin may be a specific point of the human face 134. In one example, the origin may be a user defined point. Alternatively, the origin may be a randomly selected point of the human face. Alternatively, the origin may be a center point of the human face or center of mass point of the human face.
[0082] In one example, the locations of the landmarks of the human face 134 may be determined with respect to the local coordinate system 131 e.g., the location may be provided by a direction vector between the landmark and the origin 132. The locations may be provided as 3D coordinates representing the three directions.
[0083] In one example, initial locations of the landmarks of the human face 134 may first be determined with respect to the absolute coordinate system 135 and then a transformation from the absolute coordinate system 135 to the local coordinate system 131 may be applied to the initial locations in order to obtain locations in the local coordinate system 131.
[0084] Fig. 2B is a diagram illustrating a method for determining locations of landmarks in an image 150 of human eyes 154 in accordance with an example of the present subject matter. For that, a local coordinate system 151 may be used. Fig. 2B further shows an absolute coordinate system 155. The absolute coordinate system 155 may be a world coordinate system.
[0085] The local coordinate system 151 may be defined by an origin 152 and two orthogonal directions. The origin may be a specific point of the eyes 154. In one example, the origin may be a user defined point. Alternatively, the origin may be a randomly selected point of the eyes. Alternatively, the origin may be a center point of the eyes.
[0086] In one example, the locations of the landmarks of the eyes 154 may be determined with respect to the local coordinate system 151 e.g., the location may be provided as a direction vector between the landmark and the origin 152. The locations may be provided as 2D coordinates representing the two directions.
[0087] In one example, initial locations of the landmarks of the eyes 154 may first be determined with respect to the absolute coordinate system 155 and then a transformation from the absolute coordinate system 155 to the local coordinate system 151 may be applied to the initial locations in order to obtain locations in the local coordinate system 151.
[0088] Fig. 3 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter. The training dataset may be generated using a set of images of one specific object type such as the human face. The set of images may represent faces of different individuals (objects) or of the same individual.
[0089] An image of the specific object type of the set of images may be received in step 201 . For example, the image may be received from a local storage of the computer system. Alternatively, the image may be received from a remote database system in which the image is stored e.g., step 201 may be performed in response to sending a request to the database system.
[0090] A set of landmarks of the object may be identified in step 203 in accordance with a landmark pattern. The landmark pattern may be provided in advance for the specific object type of the set of images. Alternatively, the landmark pattern may be determined in step 203 for the object type of the image.
[0091] Locations of the set of landmarks may be determined in step 205 with respect to a coordinate system defined relative to the object.
[0092] An entry may be added in step 207 to the training dataset. The entry indicates the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image.
[0093] As indicated in Fig. 3, steps 201 to 207 may be repeated for each further image of the set of images. The repetition may be performed until a stopping criterion is fulfilled. The stopping criterion may require that all the set of images are processed or that a maximum number of repetitions is reached.
[0094] In one example implementation of Fig. 3, steps 203 and 205 may be performed concurrently, e.g., the location of the landmark may be determined as the landmark is identified.
[0095] In one example implementation of Fig. 3, the method may be repeated for each further set of images, wherein the training dataset is updated/increased with new entries obtained for each further set of images.
[0096] Fig. 4 is a flowchart of a method of generating a training dataset for training a machine learning model for morphed image detection in accordance with an example of the present subject matter. The training dataset may be generated using a set of images of different object types such as the human face, human hand, eyes etc.
[0097] An image of the set of images may be received in step 301. For example, the image may be received from a local storage of the computer system. Alternatively, the image may be received from a remote database system in which the image is stored e.g., step 301 may be performed in response to sending a request to the database system. [0098] The object type represented by the image may be determined in step 303. The determination of the object type may, for example, be performed using a computer vision technique.
[0099] In step 305, a landmark pattern of the determined object type may be created or selected among provided landmark patterns.
[0100] A set of landmarks of the object may be identified in step 307 in accordance with the landmark pattern.
[0101] Locations of the set of landmarks may be determined in step 309 with respect to a coordinate system defined relative to the object in the image.
[0102] An entry may be added in step 311 to the training dataset. The entry indicates the set of locations and a label, wherein the label indicates whether the received image is a real image or morphed image.
[0103] As indicated in Fig. 4, steps 301 to 311 may be repeated for each further image of the set of images. The repetition may be performed until a stopping criterion is fulfilled. The stopping criterion may require that all the set of images are processed or that a maximum number of repetitions is reached.
[0104] In one example implementation of Fig. 4, steps 307 and 309 may be performed concurrently, e.g., the location of the landmark may be determined as the landmark is identified.
[0105] In one example implementation of Fig. 4, the method may be repeated for each further set of images, wherein the training dataset is updated/increased with new entries obtained for each further set of images.
[0106] Fig. 5 is a flowchart of a method of determining landmarks of an imaged object in accordance with an example of the present subject matter.
[0107] An image of an object may be received in step 401. The object may be detected in the image in step 403. This detection may, for example, be performed using a computer vision technique. The received image may be cropped in step 405 in an area of the image defined by the detected object. This may result in a cropped image. A set of landmarks of the object may be identified in the cropped image in step 407 in accordance with a landmark pattern. Locations of the set of landmarks may be determined in step 409 with respect to a coordinate system defined relative to the object in the cropped image.
[0108] Fig. 6 is a diagram of a data structure representing the training dataset generated in accordance with an example of the present subject matter. The data structure 500 comprises n entries 501 .1 through 501 .n. Each i-th entry 501 .i of the data structure comprises a tuple (Lh labelt), where is the location information determined for the i-th image and the labelt is the label of the i-th image.
[0109] Fig. 7 is a block diagram of an exemplary computer system for implementing at least part of the present method in accordance with an example of the present subject matter.
[0110] The components of the computer system 602 may include, but are not limited to, one or more processors or processing units 603, a storage system 611 , a memory unit 605, and a bus 607 that couples various system components including memory unit 605 to processor 603. The storage system 611 may include for example a hard disk drive (HDD). The memory unit 605 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and/or cache memory.
[0111] The computer system 602 may also communicate with one or more external devices such as a keyboard, a pointing device, a display 613, etc.; one or more devices that enable a user to interact with computer system 602; and/or any devices (e.g., network card, modem, etc.) that enable the computer system 602 to communicate with one or more other computing devices. Such communication can occur via I/O interface(s) 619. Still yet, the computer system 602 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and/or a public network (e.g., the Internet) via a network adapter 609. As depicted, the network adapter 609 communicates with the other components of the client system 602 via bus 607. [0112] The memory unit 605 is configured to store applications that are executable on the processor 603. For example, the memory unit 605 may comprise an operating system as well as one or more application programs. The application programs comprise instructions that when executed enable to perform the method described with reference to Fig. 1 , 3, 4, 5, 8, 9, 10 or 11 .
[0113] Fig. 8 is a flowchart of a method for training a classical machine learning model in accordance with an example of the present subject matter. The training may be performed using the training dataset as generated e.g., by the method of Fig. 1. For simplification, Fig. 8 may be described with reference to the data structure of Fig. 6. The entries 501.1 -n may be processed one by one using the method. In this example, the location information stored in each entry of the training dataset may be the set of locations of the set of landmarks of the image, e.g., L( = {LO 1, ...., LOC™1}. The classical machine learning model may, for example, be a deep neural network such as a CNN or a support vector machine.
[0114] In step 701 , the location information L( of the current i-th entry 501 .i may be input to the classical machine learning model. In response, the machine learning model may ouput in step 703 a value indicating whether the image represented by the location information L( is morphed or not. A loss function may be evaluated in step 705 using the value and the label labelt of the current entry 501 .i. In case (707) the loss function does not fulfill a convergence criterion, the leanrable weights of the machine learning model may be updated in step 709 and steps 701 to 707 may be repeated for a next entry of the training dataset 500; otherwise, the trained machine learning model may be provided in step 711. The update of the neural network weights may be performed using gradient descent.
[0115] Fig. 9 is a flowchart of a method for training a quantum machine learning model in accordance with an example of the present subject matter. The training may be performed using the training dataset as generated e.g., by the method of Fig. 1. For simplification, Fig. 9 may be described with reference to the data structure of Fig. 6. The entries 501.1 -n may be processed one by one using the method. In this example, the location information stored in each entry of the training dataset may be the feature vector of the image, e.g., L( = {F^, F }. [0116] In step 801 , the location information of the current i-th entry 501 .i may be input to the quantum machine learning model. In response, the quantum machine learning model may generate using the encoding layer a quantum state representing the feature vector in step 803. The quantum state may be changed in step 805 using the learning layer of the quantum machine learning model. The set of measurements of the qubits may be provided in step 807 as an indication whether the image represented by the location information L( is morphed or not. A loss function may be evaluated by a classical computer in step 809 using the output of the model and the label labelt of the current entry 501. i. In case (811 ) the loss function does not fulfill a convergence criterion the learnable parameters of the learning layer may be updated in step 813 and steps 801 to 811 may be repeated for a next entry of the training dataset 500, otherwise, the trained quantum machine learning model may be provided in step 815. The update of the learnable parameters may be performed using gradient descent.
[0117] Fig. 10 is a flowchart of a method for training a hybrid classical-quantum machine learning model in accordance with an example of the present subject matter. The training may be performed using the training dataset as generated e.g., by the method of Fig. 1 . For simplification, Fig. 10 may be described with reference to the data structure of Fig. 6. The entries 501.1 -n may be processed one by one using the method. In this example, the location information stored in each entry of the training dataset may be the set of locations of the set of landmarks of the image, e.g., Li = {LOC , ■■■ ■■ LOC™1}.
[0118] In step 901 , the location information Lt of the current i-th entry 501 .i may be input to a first neural network. The first neural network may output a feature vector in response to receiving the location information. In step 902, the feature vector may be input to the quantum machine learning model. In response, the quantum machine learning model may generate using the encoding layer a quantum state representing the feature vector in step 903. The quantum state may be changed in step 905 using the learning layer of the quantum machine learning model. The set of measurements of the qubits may be provided in step 906. A second neural network may receive as input the set of measurements and output in step 907 a value Vi indicating whether the image represented by the location information Lt is morphed or not. A loss function may be evaluated in step 909 using the value and the label labelt of the current entry 501 .i. In case (911 ) the loss function does not fulfill a convergence criterion the learnable parameters of the quantum machine learning model and the weights of the two neural networks may be updated in step 913 and steps 901 to 911 may be repeated for a next entry of the training dataset 500; otherwise, the trained quantum machine learning model as well as the two trained networks may be provided in step 915. The update of the neural network weights and the learnable parameters may be performed using gradient descent.
[0119] Fig. 11 is a flowchart of a method for morphed image detection in accordance with an example of the present subject matter.
[0120] An image of an object may be received in step 1001 . A set of landmarks of the object my be identified in step 1003 in accordance with a landmark pattern. Locations of the set of landmarks may be determined in step 1005 with respect to a coordinate system defined relative to the object. Location information may be input in step 1007 to a trained machine learning model (e.g., which is obtained in Fig. 8, 9 or 10). The location information indicates the set of locations. An output of the trained machine learning model may be received in step 1009 from the trained machine learning model. The output indicates whether the received image is a morphed image or non-morphed image.
[0121] The method of Fig. 11 may, for exmaple, be performed by the computer system of Fig. 7.
[0122] As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as an apparatus, method, computer program or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit,” “module” or “system.” Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer executable code embodied thereon. A computer program comprises the computer executable code or "program instructions".
[0123] The term “computer system” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example, a programmable processor, a computer, or multiple processors or com-puters. The apparatus can also be or further include special purpose logic circuitry, e.g., a central processing unit (CPU), a FPGA (field programmable gate array), or an ASIC (application specific integrated circuit). In some implementations, the data pro-cessing apparatus and/or special purpose logic circuitry may be hardware-based and/or software-based. The apparatus can optionally include code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The present disclosure contemplates the use of data processing apparatuses with or without conventional operating systems, for example LINUX, UNIX, WINDOWS, MAC OS, ANDROID, IOS or any other suitable conventional operating system.
[0124] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable storage medium. A ‘computer-readable storage medium’ as used herein encompasses any tangible storage medium which may store instructions which are executable by a processor of a computing device. The computer-readable storage medium may be referred to as a computer-readable non-transitory storage medium. The computer- readable storage medium may also be referred to as a tangible computer readable medium. In some embodiments, a computer-readable storage medium may also be able to store data which is able to be accessed by the processor of the computing device.
[0125] ‘Computer memory’ or ‘memory’ is an example of a computer-readable storage medium. Computer memory is any memory which is directly accessible to a processor. ‘Computer storage’ or ‘storage’ is a further example of a computer- readable storage medium. Computer storage is any non-volatile computer-readable storage medium. In some embodiments computer storage may also be computer memory or vice versa.
[0126] A ‘processor’ as used herein encompasses an electronic component which is able to execute a program or machine executable instruction or computer executable code. References to the computing device comprising “a processor” should be interpreted as possibly containing more than one processor or processing core. The processor may for instance be a multi-core processor. A processor may also refer to a collection of processors within a single computer system or distributed amongst multiple computer systems. The term computing device should also be interpreted to possibly refer to a collection or network of computing devices each comprising a processor or processors. The computer executable code may be executed by multiple processors that may be within the same computing device or which may even be distributed across multiple computing devices.
[0127] Computer executable code may comprise machine executable instructions or a program which causes a processor to perform an aspect of the present invention. Computer executable code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages and compiled into machine executable instructions. In some instances the computer executable code may be in the form of a high level language or in a pre-compiled form and be used in conjunction with an interpreter which generates the machine executable instructions on the fly.
[0128] Generally, the program instructions can be executed on one processor or on several processors. In the case of multiple processors, they can be distributed over several different entities. Each processor could execute a portion of the instructions intended for that entity. Thus, when referring to a system or process involving multiple entities, the computer program or program instructions are understood to be adapted to be executed by a processor associated or related to the respective entity. [0129] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed examples.
REFERENCE SIGNS LIST
101-107 method steps
130 image
131 local coordinate system
132 origin
135 absolute coordinate system
150 image
151 local coordinate system
152 origin
155 absolute coordinate system
201-207 method steps
301-311 method steps
401-409 method steps
500 data structure
501.1-n entries
603 processing units
605 memory unit
607 bus
609 network adapter
611 storage system
613 display
619 I/O interface
701-711 method steps
801-815 method steps
901-915 method steps
1001-1009 method steps

Claims

1. A method of generating a training dataset (500) for training a machine learning model for morphed image detection, the method comprising: repeatedly performing the following: receiving (101 ) an image (130, 150) of an object (134, 154); identifying (103) a set of landmarks of the object (134, 154) in accordance with a landmark pattern; determining (105) locations of the set of landmarks with respect to a coordinate system (131 , 151 ) defined relative to the object (134, 154); adding an entry (501.1 -n) to the training dataset (500), the entry indicating the set of locations and a label, wherein the label indicates whether the received image (130, 150) is a real image or morphed image.
2. The method of claim 1 , the receiving of the image further comprising: detecting (403) the object in the image; cropping (405) an area of the image defined by the detected object, resulting in a cropped image, wherein the landmarks are identified in the cropped image and the locations are determined using the cropped image.
3. The method of any of the preceding claims, further comprising: representing the set of locations with a feature vector in a k- dimensional feature space having a dimension k smaller than the number of set of landmarks, wherein the entry comprises the feature vector.
4. The method of any of the preceding claims 1 to 3, further comprising representing the set of locations with a feature vector using a linear transformation of a vector, y, representing the set of locations into the feature vector, x, using a weight matrix, W, where y = x x W, wherein the entry comprises the feature vector.
5. The method of any of the preceding claims 1 to 3, further comprising: representing the set of locations with a feature vector using a trained neural network comprising a fully connected layer having nodes representing the set of locations and an output layer representing the feature vector; the neural network being configured to receive the set of locations and to output the feature vector, wherein the entry comprises the feature vector.
6. The method of any of the preceding claims 1 to 5, the machine learning model comprising a quantum machine learning model, the method further comprising: representing the set of locations with a feature vector having a size that is smaller than or equal to the number of qubits of a quantum processing unit, wherein the entry comprises the feature vector, wherein the entry comprises the feature vector.
7. The method of any of the preceding claims, the landmark location being two- dimensional, 2D, coordinates.
8. The method of any of the preceding claims 1 to 6, the landmark location being three-dimensional 3D coordinates.
9. The method of any of the preceding claims, the object being a human face (134).
10. The method of any of the preceding claims, further comprising using the training dataset to train the machine learning model to detect morphed images.
11 . The method of any of the preceding claims, the method being repeated for each image of a set of images, wherein the set of images represent objects of the same object type.
12. The method of any of the preceding claims 1 to 10, the method being repeated for each image of a set of images, wherein the set of images represent objects of different object types (134, 154).
13. A computer program comprising machine executable instructions, wherein execution of the machine executable instructions causes a computer system to perform at least the following: repeatedly performing the following: receiving (101 ) an image (130, 150) of an object (134, 154); identifying (103) a set of landmarks of the object (134, 154) in accordance with a landmark pattern; determining (105) locations of the set of landmarks with respect to a coordinate system defined relative to the object (134, 154); adding (107) an entry to a training dataset (500), the entry indicating the set of locations and a label, wherein the label indicates whether the received image (130, 150) is a real image or morphed image.
14. A computer system (602) of generating a training dataset (500) for training a machine learning model for morphed image detection, the computer system (602) being configured for: repeatedly performing the following: receiving (101 ) an image (130, 150) of an object (134, 154); identifying (103) a set of landmarks of the object (134, 154) in accordance with a landmark pattern; determining (105) locations of the set of landmarks with respect to a coordinate system defined relative to the object (134, 154); adding (107) an entry to the training dataset, the entry indicating the set of locations and a label, wherein the label indicates whether the received image (130, 150) is a real image or morphed image.
15. A computer implemented data structure (500) comprising training data for training a machine learning model for morphed image detection, the data structure (500) comprising entries (501.1-n), wherein the entry comprises location information (L1 -Ln) indicating a set of landmark locations of an imaged object (134, 154), the landmark locations being relative locations, and a label indicating a real image or morphed image.
16. A method for morphed image detection, the method comprising: receiving (1001 ) an image of an object; identifying (1003) a set of landmarks of the object in accordance with a landmark pattern; determining (1005) locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting (1007) location information to a trained machine learning model, the location information indicating the set of locations; receiving (1009) an output of the trained machine learning model indicating whether the received image is a morphed image or nonmorphed image.
17. The method of claim 16, further comprising: representing the set of locations with a feature vector in a k- dimensional feature space having a dimension k smaller than the number of set of landmarks, wherein the location information comprises the feature vector.
18. The method of claim 16, wherein the location information comprises the set of locations.
19. The method of any of the preceding claims 16 to 18, wherein the machine learning model is the model of claim 10.
20. The method of any of the preceding claims 16 to 18, further comprising capturing an image of the object from an identity token, wherein the received image is the captured image.
21. The method of any of the preceding claims 16 to 20, further comprising generating a control signal for enabling access to services based on the output.
22. A computer program comprising machine executable instructions, wherein execution of the machine executable instructions causes a computer system to perform at least the following: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting location information to a trained machine learning model, the location information indicating the set of locations; receiving an output of the trained machine learning model indicating whether the received image is a morphed image or non-morphed image.
23. A computer system for morphed image detection, the computer system being configured for: receiving an image of an object; identifying a set of landmarks of the object in accordance with a landmark pattern; determining locations of the set of landmarks with respect to a coordinate system defined relative to the object; inputting location information to a trained machine learning model, the location information indicating the set of locations; receiving an output of the trained machine learning model indicating whether the received image is a morphed image or non-morphed image.
EP24728518.2A 2023-06-02 2024-05-17 Training data for image morphing detection Pending EP4710306A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
DE102023114634.3A DE102023114634B4 (en) 2023-06-02 2023-06-02 TRAINING DATA FOR IMAGE MORPHING DETECTION
PCT/EP2024/063799 WO2024245792A1 (en) 2023-06-02 2024-05-17 Training data for image morphing detection

Publications (1)

Publication Number Publication Date
EP4710306A1 true EP4710306A1 (en) 2026-03-18

Family

ID=91274548

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24728518.2A Pending EP4710306A1 (en) 2023-06-02 2024-05-17 Training data for image morphing detection

Country Status (3)

Country Link
EP (1) EP4710306A1 (en)
DE (1) DE102023114634B4 (en)
WO (1) WO2024245792A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11238271B2 (en) * 2017-06-20 2022-02-01 Hochschule Darmstadt Detecting artificial facial images using facial landmarks
GB201710560D0 (en) * 2017-06-30 2017-08-16 Norwegian Univ Of Science And Tech (Ntnu) Detection of manipulated images

Also Published As

Publication number Publication date
WO2024245792A1 (en) 2024-12-05
DE102023114634A1 (en) 2024-12-05
DE102023114634B4 (en) 2026-01-15

Similar Documents

Publication Publication Date Title
EP4320606B1 (en) Personalized biometric anti-spoofing protection using machine learning and enrollment data
CN115050064B (en) Human face liveness detection method, device, equipment and medium
KR102294574B1 (en) Face Recognition System For Real Image Judgment Using Face Recognition Model Based on Deep Learning
US11238271B2 (en) Detecting artificial facial images using facial landmarks
Chopparapu et al. A hybrid facial features extraction-based classification framework for typhlotic people
CN110222573B (en) Face recognition method, device, computer equipment and storage medium
CN111241989A (en) Image recognition method and device and electronic equipment
KR102137329B1 (en) Face Recognition System for Extracting Feature Vector Using Face Recognition Model Based on Deep Learning
CN114972016B (en) Image processing method, apparatus, computer device, storage medium, and program product
Pan et al. Hierarchical support vector machine for facial micro-expression recognition
KR20200075063A (en) Apparatus for Extracting Face Image Based on Deep Learning
US11694480B2 (en) Method and apparatus with liveness detection
Lakshmi et al. Off-line signature verification using Neural Networks
Kumar et al. A comparative study of various techniques of image segmentation for the identification of hand gesture used to guide the slide show navigation
Vinayagam et al. A two-step verification-based multimodal-biometric authentication system using KCP-DCNN and QR code generation
Sujana et al. An effective CNN based feature extraction approach for iris recognition system
WO2022217294A1 (en) Personalized biometric anti-spoofing protection using machine learning and enrollment data
Arora et al. Cryptography and Tay-Grey wolf optimization based multimodal biometrics for effective security
EP4710306A1 (en) Training data for image morphing detection
Kumar et al. ResUNet: an automated deep learning model for image splicing localization
WO2025016877A1 (en) Image morphing detection
CN115188082B (en) Training method, device, equipment and storage medium of face fake identification model
Ramkumar et al. IRIS DETECTION FOR BIOMETRIC PATTERN IDENTIFICATION USING DEEP LEARNING.
Yogarajan et al. Robust deepfake detection using multi-scale feature fusion
EP1810216A1 (en) 3d object recognition

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251209

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR