WO2017136060A1 - Improving distance metric learning with n-pair loss - Google Patents

Improving distance metric learning with n-pair loss Download PDF

Info

Publication number
WO2017136060A1
WO2017136060A1 PCT/US2016/067946 US2016067946W WO2017136060A1 WO 2017136060 A1 WO2017136060 A1 WO 2017136060A1 US 2016067946 W US2016067946 W US 2016067946W WO 2017136060 A1 WO2017136060 A1 WO 2017136060A1
Authority
WO
WIPO (PCT)
Prior art keywords
pairs
classes
training examples
computer
examples
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2016/067946
Other languages
French (fr)
Inventor
Kihyuk SOHN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Laboratories America Inc
Original Assignee
NEC Laboratories America Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Laboratories America Inc filed Critical NEC Laboratories America Inc
Priority to DE112016006360.1T priority Critical patent/DE112016006360T5/en
Priority to JP2018540162A priority patent/JP2019509551A/en
Publication of WO2017136060A1 publication Critical patent/WO2017136060A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F17/00Digital computing or data processing equipment or methods, specially adapted for specific functions
    • G06F17/10Complex mathematical operations
    • G06F17/11Complex mathematical operations for solving equations, e.g. nonlinear equations, general mathematical optimization problems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N7/00Computing arrangements based on specific mathematical models
    • G06N7/01Probabilistic graphical models, e.g. probabilistic networks

Definitions

  • the present invention relates t computer learning and more particularly- to improving distance metric learning with N-pair loss.
  • a computer-implemented method includes receiving, by a processor, N pairs of training examples and class labels tor th training examples that correspond to a- ' plurality of classes.
  • Each of the N pairs includes a respective anchor example and further includes a respecti ve non-anchor example capable of being a positive training example or a negative training example.
  • the method further includes extracting., by the processor., features of the N pairs by applying a deep convolutional neural network to the N pairs and to the class labels.
  • the method also includes calculating, by the processor for each of the N pairs ' ased OK ' the features, a respective similarly measure between the respective ancho example and the respective non-anchor example.
  • the method additionally includes calculating, by the processor, a similarity score based o the respective similarity measure for each of the pairs.
  • the similarity score represents one or more similarities between all anchor points and all positive training examples in the N pairs relative to one or more similarities betwee all of the anchor points and all negati ve training examples in the H pairs.
  • the method further includes maximizing, by the processor, the similarity score for the respective anchor example for each of the N pairs to pull together in a distribution space the training examples from a same one of the pluralit of classes while pushing apart in the distribution space the training examples from different ones of the plurality of classes.
  • a system includes a processor.
  • the processor is configured to receive N pairs of training examples and class labels for the training examples that correspond to a plurality o classes.
  • Each of the ⁇ pairs includes a respective anchor example and further includes a respective non-anchor example capable of being a positi ve training exampl e or a negati ve training example.
  • the processor is further configured to extract features of the pairs by applying a deep convolutional neural network to the N pairs and to the class labels.
  • the processor is also configured to calculate, for each of the N pairs based on the features, a respective similarly measure between the respective anchor example and the respective no» ⁇ a»chor example.
  • the processor is additionally configured to calculate a similarity score based on the respective similarity measure for each of the N pairs.
  • the similarity score represents one or more similarities between ail anchor points and all positive training examples in the N pairs relative to one or more ⁇ similarities between all of the anchor points and all negative training examples in the pairs,
  • the processor is further configured to maximize the similarity score for the respective anchor exaaiple for each of the N pairs to pull together in a distribution space the training examples from a same one of the plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurality of classes.
  • FIG, I shows a block diagram of an exemplary processing system 100 to which the present invention ma be applied, in accordance with ' an embodiment of the present invention
  • FIG. 2 shows an exemplary environment 200 to whic the present invention can be applied, in accordance with an embodiment of the present invention ;
  • FIG. 3 shows a high-level block/flow diagram of an exemplary system/method
  • FIG. 4 further shows step 310 of method 300 of FIG. 3, in accordance with an embodiment of the present invention.
  • FIG. 4 further shows distance metric learning with N-pair loss 400, in accordance with an embodiment of the present invention
  • FIG. 5 is a diagram graphically showing the N-pair loss 400 of FIG. 4 in accordance with an embodiment of the present invention versus a conventional triplet loss 599 id accordance with the prior art;
  • FIGs. 6-8 show a flow diagram of a method 600 for deep metric learning with N-pair loss, in accordance with an embodiment of the present invention
  • the present invention is directed to improving distance metric learning with N- pair loss.
  • the present invention solves the tundameiital machine learning problem of distance metric learning when the num e of output classes is extremely large, the total nomber of output classes is unknown or the distribution of output classes is variable over ⁇ time using deep learning.
  • the present invention considers N pairs of examples from N different classes at once. [0017] ⁇ » an embodiment, the present invention introduces a new objective function for deep metric learning. The objective ' function . allows faster convergence to better local optimum,
  • the present invention provides an N-patr loss for deep metric learning.
  • the present invention allows training of deep neural networks such that it trains to pull examples from the same class together while pushiag those from different classes apart.
  • the present invention pushes not just one negative example at each update but N - 1 negative examples from all. different classes based on their relative distances to the reference example.
  • FIG. 1 shows a block diagram of an exemplary processing system 100 to which the invention principles may be applied, i accordance with an embodiment of the present invention.
  • the processing system 100 includes at least one processor (CPU) 104 operatively coupled to other components via a system bus 10:2.
  • a cache .106, a Read Only Memory (ROM) 108, a Random Access Memory (RAM) 1 10, an input output (I/O) adapter 120, a sound adapter 130, a network adapter 140, a user interface adapter 150, and a display adapter 160, are operatively coupled to the system bus 102.
  • a first storage device 122 and a second storage device 124 are operatively coupled to system bus 102 by the I/O. adapter 1 0.
  • the storage devices 122 aid 124 can be any of a disk storage device (e.g., a magnetic or optical disk storage device), a solid state magnetic device, and so forth.
  • the storage devices 122 and 124 can be the same type of storage device or different types of storage devices.
  • a speaker 132 is operati vely coupled to system bus 102 by the sound adapter 130
  • a transceiver 1 2 is operativel coupled to system bits 102 b network adapter 1
  • a display device 162 is operaitvely eoispled to system bus 102 by display adapter 160.
  • A. first user input device 152, a second user input device 154, and a third user input device .156 are operative! ⁇ ' coupled .to sy stem bus 102 by user interface adapter 150.
  • the user input devices 152, 154, and 156 can be any of a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device incorporating the functionality of at least two of the preceding devices, and so forth. Of coarse, other types of input devices can also be used, while maintaining the spirit of the present invention.
  • the user input devices 152, 154, and 156 can be the same ' type of user input device or different types of user input devices.
  • the use input devices 152, 154, and 156 are used to input and output information ' to and from system 100.
  • the processing system 100 may also include oilier elements (not shown), as readily contemplated by one of skill in the art, as well as omit certain elements.
  • various other input devices and/or output devices can be included in processing system 100, depending upon the particular implementation of the same, as readily understood by one of ordinary skill in the art.
  • various types of wireless and/or wired input and/or output devices can be used.
  • additional processors, controllers, memories, and so forth, in various configurations can also be utilized as readily appreciated by one of ordinary skill in the art.
  • environment 200 described below with respect to FIG. 2 is an environment tor implementing respective embodiments of the present invention.
  • Part or -all .of processing system 100 may be implem nted, in one or more of the elements of environment 200,
  • processing system ⁇ 00 may perform at least part of the method described herein including, for example , at least part of method 300 of FIG. 3 and/or at least part of method 400 of FIG. 4 and or at least part of method 600 of FlGs. 6-8.
  • pari or all of environment 200 may be used to perform at least part of method 300 of FIG. 3 and/or at least part of method 400 of FIG. 4 and'or at least part of method 600 of FIGs . 6-8.
  • FIG. 2 shows an exemplary environment 200 to which the present invention can b applied, in accordance with an embodiment of the present invention.
  • the environmen 200 is representative of a computer network to which the present invention can be applied.
  • the elements shown relative to FIG. 2 are set forth for the sake of illustration. However, it is to be appreciated that the present inventio can be applied to other network configurations as readily contemplated by one of ordinary skill in the art given, the teachings of the present invention provided herein, while maintaining the spirit of the present invention.
  • the environment 200 at, least includes a set of computer processing systems 210.
  • the computer processing systems 210 can be any type of computer processing system including, but not limited to, servers, desktops, laptops, tablets, smart phones, media playback devices, and so forth.
  • the computer processing systems 210 include server -210A, server 21.0B, and server 2 I OC.
  • an emlx diment
  • the present invention improves distance metric learning with N-pair loss.
  • the present invention can employ any of the computer processing systems 210 to perform distance metric learning with deep learning as described, herein, in an embodiment, one of the -computer processing systems 210 can classify information recei ved by other ones of the com uter processing systems.
  • FIG. 2 the elements thereof are interconnected by a nefwork(s) 201.
  • DSP Digital Signal Processing
  • ASIC Application Specific- Integrated Circuit
  • FPGA Field Programmable Gate Array
  • C LD Complex Programmable Logic Devices
  • FIG. 3 shows a high-level block/flow diagram of an exemplary sysiem/method 300 for deep metric learning with N-pair loss, in accordance with an embodiment of the present in vers don.
  • step 3 i perform distance metric learning with deep learning.
  • step 310 includes steps 1. OA, 31 OB, and 310C.
  • the images include N pairs of examples from N different classes at once,
  • step 310B extract features from the images.
  • step 3 I OC perform distance metric learning with N ⁇ pair loss on the features and form a classifier 370.
  • step 320 test the system on image verification.
  • step 320 includes steps 320A, 320B, 320C, 320D, 320E, and 320F.
  • step 320A receive a first image (image 1).
  • step 32033 receive a second image (image 2).
  • step 320C extract features using a trained deep convoiutional neural network 35QA.
  • the deep convoiutional neural network 350 is trained to become trained deep convoiutional neural network 350A.
  • step 320D output a first feature (feature 1).
  • step 320E output a second feature (feature .2).
  • step 320F input the features (feature 1 and feature 2) and into the classifier
  • the classifier 370 can be used to generate predictions, based on which, certain actions can be taken (e.g., see. FIG. 6).
  • step 10 it is to be appreciated that the same differs from previous approaches in at least using N pairs of examples from N different classes at once.
  • step 320 the N-pair loss can be viewed as a form of neighborhood component analysis.
  • FIG. 4 further shows step 310 of method 300 of FIG . 3, in accordance with an embodiment of the present invention.
  • FIG. 4 further shows distance metric learning with ' N-pair loss 400. in. accordance with an embodiment of the present invention.
  • the deep convolutiona! neural network 350 receives N pairs of images 421 from different classes at once.
  • the reference .numeral 401 shows features before training with N-pair loss
  • the reference numeral 402 shows the feature after training with N-pair loss.
  • N-pair loss can be defined as follows:
  • FIG. 5 is a diagram graphically showing the N-pair loss 400 of FIG. 4 in accordance wit an embodiment of the present inventio versas a conventional triplet loss 599 in accordance with the prior art.
  • the conventional triplet loss 599 is equivalent to a 2-pair loss.
  • the 2-pai loss is a generalization to a N-pair loss for N 2.
  • FiGs. 6-8 sho a flow diagram of a method 600 for deep metric learning with N-pair loss, in accordance with an embodiment of the present invention.
  • each of the N pairs includes a respecti ve anchor example and further includes a respective non-anchor example capable of being a positive training example or a negative training example.
  • each of the N pairs of the training examples can correspond to a different one of the plurality of classes.
  • the plurality of classes can be randomly sejected as a subset from a set of classes, wherein the set of classes includes the plurality of classes and on or more other classes.
  • the total number of the pluralit of classes at least one of (i) changes over time, (ii) is larger than a threshold amount, and (ni) is unknown.
  • step 620 extract features of the N pairs by applying a deep convolutional neural network to the N pairs and to the class labels.
  • step 630 calculate, for each of the N pairs based oft the features, a respective similarly measure between the respective anchor example and the respective non-anchor example.
  • step 640 calculate a similarity score based on the respective similarity measure for each of the pairs.
  • the similarity score represents one or more similarities between all anchor points and all positive training examples in the pairs relative to one or more .similarities between all of the anchor points and all negative training examples in the N pairs.
  • ste 640 includes one or more of steps 640 A , 640B, and 640C.
  • step 640A ' bound a variable (p ( ) used to calculate the respective similarity score of each of the pairs of training examples by at least one of a lower limit and an upper limit the variable representing a relative similarity between the anchor point and the positive training examples with respect to the anchor point and the negative training examples.
  • step 640B compute a gradient of a logarithm of the similarity score.
  • step 640C maximize aiV objecti ve function for deep metric learning.
  • step 640C inc! udes step 640C I .
  • step 640CL maximize a portion of the objection function that relates to the anchor points, wherein the objective function include the portion relating to the anchor points and at least one other portion relating to the non-anchor points.
  • step 650 maximize the similarity score for the respective anchor example for each of the N pairs to pull together in a distribution space the training examples from a same one of th plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurality- of classes, in an.
  • step 650 is capable of simultaneously pushing N-i examples away from a single reference sample from among the N pairs of training examples, in the distribution, space.
  • step 650 is capable of simultaneously pushing N-I examples towards a single reference sample from among the N pairs of training examples, in the distribution space.
  • ste 660 generate prediction using the deep convomtionaS neural network. For example, generate a facial recognition prediction, a speech recognitio prediction, a speaker recognition prediction, and so forth.
  • step 670 perform an action responsive to the prediction.
  • the actio taken is dependent upon the implementation. For example, access to an entity including, but not limited to, device, a system, or a facility, cm be granted responsive to the prediction. It is to be appreciated that the preceding actions are merely illustrative and, thus, other actions can also be performed! ' as readily appreciated by one of ordinary skill in the art, while maintaining the spirit of the present ' invention.
  • step 670 includes step 670A.
  • step 670A verify a user and provide the user access to an entity, based on the prediction.
  • Supervised deep metric learning aims to learn an embedding vector representation of the data using deep neural networks that preserves the distance between examples from the same class small and those from different classes large.
  • the contrastive loss and the triplet loss functions have been used to trai deep embedding networks:
  • ⁇ W is a embedding ' kernel defined by deep neural networks, and yi e ⁇ l 5 , ..,L ⁇ 's are the label of the data x ( € ⁇ .
  • x + and x ⁇ are used to represent positive and negative examples of x, i.e., y + — y and y ⁇ ⁇ y, respectively.
  • /TMj(x) is- used to denote embedding vector representation of * while inheriting all superscripts, and subscripts when exist.
  • Two objective functions ' are similar in the sense that they both optimize embedding kernels to preserve the distance between examples in the label space to the embedding space, but the triplet loss can be considered as a relaxation of th contrastjve loss since it only cares the relative margin of distances between positive and negative pairs, not their absolute values.
  • the loss functions are differentiable with respect to the kernel parameters and therefore they can be readily used as an objective function for training dee neural networks.
  • a pa is a normalized self-similarity, Xe.
  • Not thai i is bounded by (0, .1 ) and it represents the relative similarity between anchor and positive points to the similarities between anchor and negative points. Maximizing the score of all anchor points in N-pair training subset pulls the examples from the same class together, but at the same time, it pushes the examples from different classes away based on their relative dissimilarity, i.e., negative examples in the proximity of anchor point will be pushed away than those already far enough, as illustrated in FIG, 4. After all, the A-pair loss is defined as follows; mpair ( ⁇ (3 ⁇ 4 **) ⁇ f ) ⁇ l - lOgPi (5)
  • TABLE 1 shows a comparison of Joss functions for deep metric learning. 2- pair loss is equivalent to triplet loss under 6 convergence criteria, while its score function is an approximation to that of .A/ ⁇ pak loss ft >rA ⁇ > 2
  • iV-pair loss is described with respect to triplet loss and sofhnax loss.
  • T-f-pair ⁇ «3 ⁇ 4 equivalent when € TM ⁇ — ff ( ⁇ ).
  • Equation (15) where details in Equation (15) are omitted as it repeats Equation (1.1) - - ( ⁇ 4) baekwardly. Finally, this proves / 6 33 ⁇ 4 .
  • the exact partition function can be approximated by randomly selecting a small subset of N templates including a ground-truth template as follows:
  • the -pair loss is an approximation of ' A'-pair loss .forM ⁇ A 7 .
  • the score function in Equation (4) is not designed t he invariant to the norm of embedding vectors.
  • the score- function can be made to he arbitrarily close to I or 0 by re-sealing embedding vectors.
  • the self- similarity score function can be maxim ized by increasing the norm of embedding vectors rather than finding right direction, and it is important to regularize the norm of the embedding vector io avoid such situation, e.g., 12 normalization on embedding vectors to compute triplet loss.
  • 12 normalization makes optimization very difficult since the self-similarity score is upper bounded by (for example, the upper bound is 0.88 whea N ⁇ 2, but it decreases
  • the present invention allows efficient training by (!) removing hard negative data mimng, (2) removing computationally and parameter heavy softmax layer, and (3) faster convergence than previous deep metric learning approaches.
  • the present invention is effective for technologies such as face recognition where the number of output classes (e.g., identity) is extremely large.
  • the present invention is effective for online learning where the number of output classes is unknow or changi g over time.
  • N pairs of exanrples are used from a random subset of classes that enables pushing examples from different classes apart quickly,
  • Embodiments described herein may be entirely hardware, entirely software or including both hardware and software" elements.
  • the present invention is implemented in software, which includes but is not iimited to firmware, resident software, microcode, etc.
  • Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer o any ins traction execution system.
  • a computer-usable or computer readable medium may include any apparatus that stores, conunu-iicates, propagates, or transports the program for use by or m connection with the instruction execution system, apparatus,- of device.
  • the medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
  • the medium may include a computer-readable storage medium such as a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk and an optical disk, etc.
  • Each computer program may be tangibly stored in a machine- readable storage media or device (e.g., program memory or magnetic disk) readable by a general or special purpose programmable computer, for configuring and ' controlling operation, of a computer when the storage medi or device is read by the computer to perform the procedures described herein.
  • the inventive system may also be considered to be embodied in a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
  • a data processing system suitable for storin and/or executing program code may include at least one processor coupled directly or indirectly to memory elements through a system bus.
  • the memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution.
  • I/O devices including but not limited to keyboards, displays, pointing devices, etc.
  • I/O controllers including but not limited to keyboards, displays, pointing devices, etc.
  • Network adapters m y also he coupled to the system to enable the data processing; system to become eonpied to other dat processing systems or remote printer or storage -devices through intervening private or public networks.
  • Modems, cable motlera and Ethernet cards are just a few of the currently available types of network adapters.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Artificial Intelligence (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Mathematical Optimization (AREA)
  • Mathematical Analysis (AREA)
  • Computational Mathematics (AREA)
  • Pure & Applied Mathematics (AREA)
  • Algebra (AREA)
  • Operations Research (AREA)
  • Databases & Information Systems (AREA)
  • Probability & Statistics with Applications (AREA)
  • Image Analysis (AREA)
  • Machine Translation (AREA)

Abstract

A method includes receiving N pairs of training examples and class labels therefor. Each pair includes a respective anchor example, and a respective non-anchor example capable of being a positive or a negative training example. The method further includes extracting features of the pairs by applying a DHCNN, and calculating, for each pair based on the features, a respective similarly measure between the respective anchor and non-anchor example. The method additionally includes calculating a similarity score based on the respective similarity measure for each pair. The score represents similarities between all anchor points and positive training examples in the pairs relative to similarities between all anchor points and negative training examples in the pairs. The method further includes maximizing the similarity score for the anchor example for each pair to pull together the training examples from a same class while pushing apart the training examples from different classes.

Description

IMPROVING DISTANCE METRIC LEARNING WITH N-PAIR LOSS
RELATED APPLICATION INFORMATION
[0001 ] This application claims priority to CIS. Provisional Pat. App. Ser, No. 62/2 1 ,025 filed on February 4, 2016, incorporated herein by reference in its entirety.
BACKGROUND
Technical Field
[0002] The present invention relates t computer learning and more particularly- to improving distance metric learning with N-pair loss.
Description of the Related Art
[0003] Deep metric learning has been tackled in many ways but most notably, -conirastive loss and triplet loss have been .used for training objectives, of dee learning. Previous approaches considered pair/wise relationship between two different classes and suffered from slo convergence to an unsatisfactory local n nrnurn. Thus, there is a need for improved metric learning,
SUMMARY
[0004] According to an aspect of the present invention, a computer-implemented method is provided. The method includes receiving, by a processor, N pairs of training examples and class labels tor th training examples that correspond to a- 'plurality of classes. Each of the N pairs includes a respective anchor example and further includes a respecti ve non-anchor example capable of being a positive training example or a negative training example. The method further includes extracting., by the processor., features of the N pairs by applying a deep convolutional neural network to the N pairs and to the class labels. The method also includes calculating, by the processor for each of the N pairs' ased OK 'the features, a respective similarly measure between the respective ancho example and the respective non-anchor example. The method additionally includes calculating, by the processor, a similarity score based o the respective similarity measure for each of the pairs. The similarity score represents one or more similarities between all anchor points and all positive training examples in the N pairs relative to one or more similarities betwee all of the anchor points and all negati ve training examples in the H pairs. The method further includes maximizing, by the processor, the similarity score for the respective anchor example for each of the N pairs to pull together in a distribution space the training examples from a same one of the pluralit of classes while pushing apart in the distribution space the training examples from different ones of the plurality of classes.
[0005] According to another aspect, of the present invention, a system is provided. The system includes a processor. The processor is configured to receive N pairs of training examples and class labels for the training examples that correspond to a plurality o classes. Each of the Ή pairs includes a respective anchor example and further includes a respective non-anchor example capable of being a positi ve training exampl e or a negati ve training example. The processor is further configured to extract features of the pairs by applying a deep convolutional neural network to the N pairs and to the class labels. The processor is also configured to calculate, for each of the N pairs based on the features, a respective similarly measure between the respective anchor example and the respective no»~a»chor example. The processor is additionally configured to calculate a similarity score based on the respective similarity measure for each of the N pairs. The similarity score represents one or more similarities between ail anchor points and all positive training examples in the N pairs relative to one or moresimilarities between all of the anchor points and all negative training examples in the pairs, The processor is further configured to maximize the similarity score for the respective anchor exaaiple for each of the N pairs to pull together in a distribution space the training examples from a same one of the plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurality of classes.
[0006] These and othe features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which is to e read in connection with the accompanying drawings:
BRIEF DESCRIPTION OF DRAWINGS
[0007] The disclosure will provide details in the following description of preferred embodiments with reference to the following figures wherein:
[0008] FIG, I shows a block diagram of an exemplary processing system 100 to which the present invention ma be applied, in accordance with' an embodiment of the present invention;
[0009] FIG. 2 shows an exemplary environment 200 to whic the present invention can be applied, in accordance with an embodiment of the present invention ; [0010] FIG. 3 shows a high-level block/flow diagram of an exemplary system/method
300 for deep metric learning with N-pair loss, in accordance with ars .embodiment of the present invention;
[0011] FIG. 4 further shows step 310 of method 300 of FIG. 3, in accordance with an embodiment of the present invention. In particular, FIG. 4 further shows distance metric learning with N-pair loss 400, in accordance with an embodiment of the present invention;
[0012] FIG. 5 is a diagram graphically showing the N-pair loss 400 of FIG. 4 in accordance with an embodiment of the present invention versus a conventional triplet loss 599 id accordance with the prior art; and
[0013] FIGs. 6-8 show a flow diagram of a method 600 for deep metric learning with N-pair loss, in accordance with an embodiment of the present invention,
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0014] The present invention is directed to improving distance metric learning with N- pair loss.
[0015] The present invention solves the tundameiital machine learning problem of distance metric learning when the num e of output classes is extremely large, the total nomber of output classes is unknown or the distribution of output classes is variable over¬ time using deep learning.
[0016] In an embodiment, and in contrast to prior ait approaches, the present invention considers N pairs of examples from N different classes at once. [0017] ί» an embodiment, the present invention introduces a new objective function for deep metric learning. The objective' function . allows faster convergence to better local optimum,
[0018] The present invention provides an N-patr loss for deep metric learning. The present invention allows training of deep neural networks such that it trains to pull examples from the same class together while pushiag those from different classes apart. The present invention pushes not just one negative example at each update but N - 1 negative examples from all. different classes based on their relative distances to the reference example.
[0019] FIG. 1 shows a block diagram of an exemplary processing system 100 to which the invention principles may be applied, i accordance with an embodiment of the present invention. The processing system 100 includes at least one processor (CPU) 104 operatively coupled to other components via a system bus 10:2. A cache .106, a Read Only Memory (ROM) 108, a Random Access Memory (RAM) 1 10, an input output (I/O) adapter 120, a sound adapter 130, a network adapter 140, a user interface adapter 150, and a display adapter 160, are operatively coupled to the system bus 102.
[0020] A first storage device 122 and a second storage device 124 are operatively coupled to system bus 102 by the I/O. adapter 1 0. The storage devices 122 aid 124 can be any of a disk storage device (e.g., a magnetic or optical disk storage device), a solid state magnetic device, and so forth. The storage devices 122 and 124 can be the same type of storage device or different types of storage devices. [0021] A speaker 132 is operati vely coupled to system bus 102 by the sound adapter 130, A transceiver 1 2 is operativel coupled to system bits 102 b network adapter 1 0, A display device 162 is operaitvely eoispled to system bus 102 by display adapter 160.
[0022] A. first user input device 152, a second user input device 154, and a third user input device .156 are operative!}' coupled .to sy stem bus 102 by user interface adapter 150. The user input devices 152, 154, and 156 can be any of a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device incorporating the functionality of at least two of the preceding devices, and so forth.. Of coarse, other types of input devices can also be used, while maintaining the spirit of the present invention. The user input devices 152, 154, and 156 can be the same 'type of user input device or different types of user input devices. The use input devices 152, 154, and 156 are used to input and output information 'to and from system 100.
[0023] Of course, the processing system 100 may also include oilier elements (not shown), as readily contemplated by one of skill in the art, as well as omit certain elements. For example, various other input devices and/or output devices can be included in processing system 100, depending upon the particular implementation of the same, as readily understood by one of ordinary skill in the art. For example, various types of wireless and/or wired input and/or output devices can be used. Moreover, additional processors, controllers, memories, and so forth, in various configurations can also be utilized as readily appreciated by one of ordinary skill in the art. These and other variations of the processing system 100 are readily contemplated, by one of ordinary skill in the art given the teachings of the present invention provided herein. [0024] Moreover, it is to be appreciated that environment 200 described below with respect to FIG. 2 is an environment tor implementing respective embodiments of the present invention. Part or -all .of processing system 100 may be implem nted, in one or more of the elements of environment 200,
[0O25J Further, it is to be appreciated that processing system Ϊ00 may perform at least part of the method described herein including, for example , at least part of method 300 of FIG. 3 and/or at least part of method 400 of FIG. 4 and or at least part of method 600 of FlGs. 6-8. Similarly, pari or all of environment 200 may be used to perform at least part of method 300 of FIG. 3 and/or at least part of method 400 of FIG. 4 and'or at least part of method 600 of FIGs . 6-8.
[0026] FIG. 2 shows an exemplary environment 200 to which the present invention can b applied, in accordance with an embodiment of the present invention. The environmen 200 is representative of a computer network to which the present invention can be applied. The elements shown relative to FIG. 2 are set forth for the sake of illustration. However, it is to be appreciated that the present inventio can be applied to other network configurations as readily contemplated by one of ordinary skill in the art given, the teachings of the present invention provided herein, while maintaining the spirit of the present invention.
[0027] The environment 200 at, least includes a set of computer processing systems 210. The computer processing systems 210 can be any type of computer processing system including, but not limited to, servers, desktops, laptops, tablets, smart phones, media playback devices, and so forth. For the sake of illustration, the computer processing systems 210 include server -210A, server 21.0B, and server 2 I OC. [0028] ί» an emlx diment, the present invention improves distance metric learning with N-pair loss. The present invention can employ any of the computer processing systems 210 to perform distance metric learning with deep learning as described, herein, in an embodiment, one of the -computer processing systems 210 can classify information recei ved by other ones of the com uter processing systems.
[0029] In the embodiment shown in FIG. 2, the elements thereof are interconnected by a nefwork(s) 201. However, in other embodiments, other types of connections can also be used. Additionally, one or more elements i FIG. 2 .may be implemented by a variety of devices, which include but are not limited to, Digital Signal Processing (DSP) circuits, programmable processors, Application Specific- Integrated Circuits (ASIC's), Field Programmable Gate Arrays (FPGAs), Complex Programmable Logic Devices (C LDs), and so forth . These and other var iations of the elements of envi ronment 200 are readily determined by one of ordinary ski ll in the art, given the teachings of the present invention provided herein, while maintaining the spirit of the present invention.
[0030] FIG. 3 shows a high-level block/flow diagram of an exemplary sysiem/method 300 for deep metric learning with N-pair loss, in accordance with an embodiment of the present in vers don.
[0031 J At step 3 i 0, perform distance metric learning with deep learning.
[0032] In an embodiment step 310 includes steps 1. OA, 31 OB, and 310C.
[0033] At ste 31 OA, provide images to a deep convolutions!, neural network 350, The images include N pairs of examples from N different classes at once,
[0034] At step 310B. extract features from the images. [0035] At step 3 I OC, perform distance metric learning with N~pair loss on the features and form a classifier 370.
[0036] At step 320, test the system on image verification.
[0037] In an embodiment, step 320 includes steps 320A, 320B, 320C, 320D, 320E, and 320F.
[0038] At step 320A, receive a first image (image 1).
[0039] At step 32033, receive a second image (image 2).
[0040] At step 320C, extract features using a trained deep convoiutional neural network 35QA. The deep convoiutional neural network 350 is trained to become trained deep convoiutional neural network 350A.
[0041 ] At step 320D, output a first feature (feature 1).
[0042] At step 320E, output a second feature (feature .2).
[0043] At step 320F, input the features (feature 1 and feature 2) and into the classifier
370.
[0044] The classifier 370 can be used to generate predictions, based on which, certain actions can be taken (e.g., see. FIG. 6).
[0045] Regarding step 10, it is to be appreciated that the same differs from previous approaches in at least using N pairs of examples from N different classes at once.
[0046] Regarding step 320. it is to be appreciated that, the N-pair loss can be viewed as a form of neighborhood component analysis.
[0047] FIG. 4 further shows step 310 of method 300 of FIG . 3, in accordance with an embodiment of the present invention. In particular, FIG. 4 further shows distance metric learning with 'N-pair loss 400. in. accordance with an embodiment of the present invention.
[0048] The deep convolutiona! neural network 350 receives N pairs of images 421 from different classes at once. In FIG; 4, the reference .numeral 401 shows features before training with N-pair loss, and the reference numeral 402 shows the feature after training with N-pair loss.
[0049] In FIG. 4, the following notations apply:
x: input image;
f: output feature;
fi: example from i-th pair;
fT: positive example from i-th pair with
lis having different class labels.
[0050] In an embodiment, N-pair loss can be defined as follows:
[0051] FIG. 5 is a diagram graphically showing the N-pair loss 400 of FIG. 4 in accordance wit an embodiment of the present inventio versas a conventional triplet loss 599 in accordance with the prior art.
[0052] The conventional triplet loss 599 is equivalent to a 2-pair loss.
[0053] The 2-pai loss is a generalization to a N-pair loss for N 2.
[0054] The following equations apply
[0055] FiGs. 6-8 sho a flow diagram of a method 600 for deep metric learning with N-pair loss, in accordance with an embodiment of the present invention.
[0056] At step 610, receive N pairs of training examples and class labels for the 'framing examples that correspond to a plurality of classes . Each of the N pairs includes a respecti ve anchor example and further includes a respective non-anchor example capable of being a positive training example or a negative training example. In an embodiment, each of the N pairs of the training examples can correspond to a different one of the plurality of classes. In. an embodiment, the plurality of classes can be randomly sejected as a subset from a set of classes, wherein the set of classes includes the plurality of classes and on or more other classes. In an embodiment, the total number of the pluralit of classes at least one of (i) changes over time, (ii) is larger than a threshold amount, and (ni) is unknown.
[0057] At step 620. extract features of the N pairs by applying a deep convolutional neural network to the N pairs and to the class labels.
[0O58J At step 630, calculate, for each of the N pairs based oft the features, a respective similarly measure between the respective anchor example and the respective non-anchor example.
[0059] At step 640, calculate a similarity score based on the respective similarity measure for each of the pairs. The similarity score represents one or more similarities between all anchor points and all positive training examples in the pairs relative to one or more .similarities between all of the anchor points and all negative training examples in the N pairs.
[0060] In an embodiment, ste 640 includes one or more of steps 640 A , 640B, and 640C.
[0061] At step 640A, 'bound a variable (p() used to calculate the respective similarity score of each of the pairs of training examples by at least one of a lower limit and an upper limit the variable representing a relative similarity between the anchor point and the positive training examples with respect to the anchor point and the negative training examples.
[0062] At step 640B, compute a gradient of a logarithm of the similarity score.
[0063] At step 640C, maximize aiV objecti ve function for deep metric learning.
[0064] In an embodiment, step 640C inc!udes step 640C I .
[0065] At step 640CL maximize a portion of the objection function that relates to the anchor points, wherein the objective function include the portion relating to the anchor points and at least one other portion relating to the non-anchor points.
[0066] At step 650. maximize the similarity score for the respective anchor example for each of the N pairs to pull together in a distribution space the training examples from a same one of th plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurality- of classes, in an. embodiment, step 650 is capable of simultaneously pushing N-i examples away from a single reference sample from among the N pairs of training examples, in the distribution, space. In. an embodiment, step 650 is capable of simultaneously pushing N-I examples towards a single reference sample from among the N pairs of training examples, in the distribution space.
[0067] At ste 660, generate prediction using the deep convomtionaS neural network. For example, generate a facial recognition prediction, a speech recognitio prediction, a speaker recognition prediction, and so forth.
[0068] At step 670, perform an action responsive to the prediction. As appreciated by one of ordinary skill in the art, the actio taken is dependent upon the implementation. For example, access to an entity including, but not limited to, device, a system, or a facility, cm be granted responsive to the prediction. It is to be appreciated that the preceding actions are merely illustrative and, thus, other actions can also be performed!' as readily appreciated by one of ordinary skill in the art, while maintaining the spirit of the present' invention.
[0069] In a embodiment, step 670 includes step 670A.
[0070] At step 670A, verify a user and provide the user access to an entity, based on the prediction.
[0071] A description will now be given regarding supervised deep metric learning, in accordance with an embodiment of the present invention.
[0072] The description regarding supervised deep metric learning will 'commence with a description regarding contrastive and triplet loss,
[0073] Supervised deep metric learning aims to learn an embedding vector representation of the data using deep neural networks that preserves the distance between examples from the same class small and those from different classes large. The contrastive loss and the triplet loss functions have been used to trai deep embedding networks:
Figure imgf000014_0001
where /'(; $): χ→ W is a embedding 'kernel defined by deep neural networks, and yi e{l 5, ..,L} 's are the label of the data x(€ χ. Herein, x+and x~ are used to represent positive and negative examples of x, i.e., y+— y and y~≠ y, respectively. [<£]+- m x {0; d) and m > 0 Is a toning parameter for margin. For simplicity,/™j(x) is- used to denote embedding vector representation of * while inheriting all superscripts, and subscripts when exist. Two objective functions' are similar in the sense that they both optimize embedding kernels to preserve the distance between examples in the label space to the embedding space, but the triplet loss can be considered as a relaxation of th contrastjve loss since it only cares the relative margin of distances between positive and negative pairs, not their absolute values. The loss functions are differentiable with respect to the kernel parameters and therefore they can be readily used as an objective function for training dee neural networks.
[0074] Although it 'sounds- straightforward., applying contrastiv loss or triplet loss functions to train dee neural networks that can provide highly discriminative embedding vector is non-trivial because the margin constraints of the above loss functions can be easily satisfied for most of the training pairs or triplets after few epochs of training. To avoid bad local minima, different data selection methods have been explored such as an online triplet selection algorithm that, selects {semi-) hard negative but all positive examples within each mini-batch containing few thousands of exemplars. Although the data selection step is essential, it becomes more inefficient for deep metric learning since each data sample should go through the forward pass of deep neural networks to compute the distance.
[0075] A description will now be given regarding Ar~pair Loss for Deep Metric Learning, in accordance with an embodiment of the present invention. Also, theoretical insight is provided regarding why A-pair loss is better than other existing loss functions for deep metric learning by showing relations to those loss functions, such as triplet loss and sof nsa loss.
[0076] A description of <V-pair loss will now be given. Consider N pairs of training examples £{% x† )}£L¾ and labels {( ^*)}!^*' By ' definition, >¾ = y * and it is presumed that none of the pairs of examples are from the same class, i.e., >·',·≠ yj Vi≠ . The similarity measure between the anchor point xt and the positive or negative points {xfy -i is defined as follows:
pij ^ expire r }- and the score ( A pa is a normalized self-similarity, Xe.,
Figure imgf000016_0001
[0077] Not thai i is bounded by (0, .1 ) and it represents the relative similarity between anchor and positive points to the similarities between anchor and negative points. Maximizing the score of all anchor points in N-pair training subset pulls the examples from the same class together, but at the same time, it pushes the examples from different classes away based on their relative dissimilarity, i.e., negative examples in the proximity of anchor point will be pushed away than those already far enough,, as illustrated in FIG, 4. After all, the A-pair loss is defined as follows; mpair ({(¾ **)}f ) ~∑ l - lOgPi (5)
[0078] The gradient of log pt w.r.t fj can be derived, as follows:
Figure imgf000016_0002
Figure imgf000017_0001
and the gradient wxt Θ can be computed b chain-rule.
[0079] TABLE 1 shows a comparison of Joss functions for deep metric learning. 2- pair loss is equivalent to triplet loss under 6 convergence criteria, while its score function is an approximation to that of .A/~pak loss ft >rA< > 2
TA BLE 1
triplet loss )if - n\2 2 - wr -r «2 '+HL
exp( T +)
2-pair loss 0gLe p( T f+) + e p( /-) exp(/! /'+)
~pair l ss — t QPi
eS∞p( +) +∑&i ex { /i+)
[0080] A description will now be given regarding a comparison of iV-pair loss t triplet loss.
[0081] To illustrate the present invention, iV-pair loss is described with respect to triplet loss and sofhnax loss.
[0082] Regarding the comparison of iV-pair loss to triplet loss, a description of triplet loss and 2-pair loss will now be given,
[0083 J Relation between loss functions can be demonstrated by showing the equivalence between two sets of optimal embedding kernels with respect to each loss function (although the optima! sets of embedding kemeis for two loss ictions are equivalent). To proceed, optimaiity conditions for loss functions are defined as ioilows:
¾ = (f\m^t(xtx+,x-,'f) - 0,V(x,x+,x-)} {9}
Figure imgf000018_0001
~ V ( χ, X{ , Xz, X )} (10) where £ _pair = ~∑£=i— togpi— 6']+ and. embedding kemeis are restricted to have unit 12 norm for both 2-pair and triplet losses. In the following, it is shown that 3¾¾ and
T-f-pair ί«¾ equivalent when€ ·— ff (~).
^tn' c: ?2- αίτ / e ^n-i a d consider any valid 2-pair sample
Figure imgf000018_0002
forms a valid triplet sample, we have the followin ;
Figure imgf000018_0003
** ΙΙΛ-Λ+ΪΙ^-ΗΛ-Λ Ι^™ on
Figure imgf000018_0004
^ fflfi -flfz >^) ( 3)
Figure imgf000018_0005
and this proves€ z-w r-
■^2- air -¾¾■ Similarly, let * 6 -FjLpairand consider any valid -triplet sample
For any x2 with y2 - ^, a 2-pair sample {O&t,**), (x2,¾^)} can be built that satisfies the following:
Figure imgf000018_0006
— log pi≤ — log σ (~~) ♦«< (15)
Figure imgf000019_0001
where details in Equation (15) are omitted as it repeats Equation (1.1) - - (Ϊ4) baekwardly. Finally, this proves / 6 3¾ .
[0084] A descriptio will now be given regarding insight from soflraax loss,
[0085] The softmax loss with I classes is written as follows:
£-sa ftm.ax (% >¾)— ~* > 0;f I xi) ( 17)
Figure imgf000019_0002
where is a weight vector or a template for class /, It is often inefficient of impractical to compute the exact partition function
Figure imgf000019_0003
at training when L is very large. For such cases, the exact partition function can be approximated by randomly selecting a small subset of N templates including a ground-truth template as follows:
Figure imgf000019_0004
where 5 cr {1, .., , Lj, jSj— · iV and ¾ e S. The local partition function Zs (x) less than ^for any S, and the approximation becomes more accurate with larger (noting that advanced subset sampling methods, such as importance sampling and hashing can be used to reduce approximation error with small /V), This provides a valuable insight when i¥-pair loss is compared to 2-pair loss (or A/-pair loss for M < N) since the self-similarity score of 2-pair loss can be viewed as an approximation to that of JV-pair loss, hi other words, afty self-Similarity score of A-pair loss can be approximated with those of 2-pair loss, but none of them is tight:'
∑fLi expCtftf") expCf^) + expCtf )
V/ £ {!, -·- , {i}. This implies that the actual score ofN-pair loss could be concealed behind the overvaiiied scor of the model when it is trained with 2-pair loss, and therefore the model is likely to be sab-opiimai. It. has been determined that the 2-pair loss significantly underilts to the training data compared to the iV-pair loss with /V > 2 or softmax loss models.
[0086] A description will now be given regarding implications of the present invention with respect to various relations.
[0087] The impl ications of these relations are summarised below:
! . The optimal set of embedding kernels for 2-pair loss and triplet loss are equivalent and the performance of the models trained -with these loss functions would he similar,
2. The -pair loss is an approximation of'A'-pair loss .forM < A7.
[0088] A description will now be given regarding L2-Norm Reguiarization.
[0089] Note that the score function in Equation (4) is not designed t he invariant to the norm of embedding vectors. In other words, the score- function can be made to he arbitrarily close to I or 0 by re-sealing embedding vectors. This implies that the self- similarity score function can be maxim ized by increasing the norm of embedding vectors rather than finding right direction, and it is important to regularize the norm of the embedding vector io avoid such situation, e.g., 12 normalization on embedding vectors to compute triplet loss. However, for A-pair loss, applying 12 normalization makes optimization very difficult since the self-similarity score is upper bounded by (for example, the upper bound is 0.88 whea N ~ 2, but it decreases
Figure imgf000021_0001
to 0.105 when N ~ 64). Instead, we regularise b adding following penalty term ~∑LiSf/ilL + ! i+I to objective function that promotes the 12-norm of embedding vectors t be small,
[0090] A description will now be given regarding competitive/commercial values of the solution achieved by the present invention.
[00 1] The present invention allows efficient training by (!) removing hard negative data mimng, (2) removing computationally and parameter heavy softmax layer, and (3) faster convergence than previous deep metric learning approaches.
[0092] The present invention is effective for technologies such as face recognition where the number of output classes (e.g., identity) is extremely large.
[0093 The present invention is effective for online learning where the number of output classes is unknow or changi g over time.
[0094] Rather than using two pairs of examples with hard negative mining, N pairs of exanrples are used from a random subset of classes that enables pushing examples from different classes apart quickly,
[0095] Embodiments described herein may be entirely hardware, entirely software or including both hardware and software" elements. 111 a preferred embodiment, the present invention is implemented in software, which includes but is not iimited to firmware, resident software, microcode, etc.
[0096] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer o any ins traction execution system. A computer-usable or computer readable medium may include any apparatus that stores, conunu-iicates, propagates, or transports the program for use by or m connection with the instruction execution system, apparatus,- of device. The medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. The medium may include a computer-readable storage medium such as a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk and an optical disk, etc.
[0097] Each computer program may be tangibly stored in a machine- readable storage media or device (e.g., program memory or magnetic disk) readable by a general or special purpose programmable computer, for configuring and' controlling operation, of a computer when the storage medi or device is read by the computer to perform the procedures described herein. The inventive system may also be considered to be embodied in a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0098] A data processing system suitable for storin and/or executing program code may include at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution. Input/output or I/O devices (including but not limited to keyboards, displays, pointing devices, etc.) ma be coupled to the -system either directl or through intervening I/O controllers,
[0099] Network adapters m y also he coupled to the system to enable the data processing; system to become eonpied to other dat processing systems or remote printer or storage -devices through intervening private or public networks. Modems, cable motlera and Ethernet cards are just a few of the currently available types of network adapters.
[00100] Reference in the specification to "one embodiment" or "an embodiment:" of the present invention, as well as other variations thereof, means that a particular feature, 'structure, characteristic, and so forth described in connection with the embodiment is included in at least erne embodiment of the present invention. Thus, the appearances of the -phrase "in one embodiment'5 or ¾ an embodiment", as well any other -variations, appearing in various places throughout the specification are not necessaril all referring to the same embodiment.
[00101] It is to be appreciated that the use of any of the following "/". "and/or", and "at least one of, for example, in the cases of "A/B", "A and/or B" and "at least one of A and B", is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A aid B). As a further example, in the cases of "A, B, and/or C" and. "at least one of A, B, and C", such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and 6) only, or the Selection of the first and third listed options (A and C) only, or the- selection of the second and third listed options (B and G) onl.y5-or the selection of all three options (A and B and C). This may be extended, as readil apparent by one of ordinary skill in this and related arts, for as many items listed,
[00102] The foregoing is to be understood as bein in ever respect illustrative and exemplary, but not restrictive, and the scope of the invention disclosed herein is not to be determined from the Detailed Description, but rather from the claims as interpreted according to the full breadt permitted by the patent laws. It is to be understood that the embodiments shown and described herein are only illustrative of the principles of tire present invention and that those skilled in the art may implement various modifications without departin from the scope and spirit of the invention. Those skilled in the art could implement various other feature combinations without departing from the scope and spirit of the invention. Having thus described aspects of the invention, with the details and particularity required by die patent laws, what is claimed and desired protected by Letters Patent is set forth in the appended claims.

Claims

WHAT IS CLAIMED IS:
1 , A computer-implemented method, comprising:
receiving, b a processor, N pairs of traming examples nd class labels for the training examples that correspond to a plurality of c lasses, w herein each of th e N pairs includes a respective anchor example and further includes a respective non-anchor example capable of being a positive training example or a negative traming example;
extracting, by the processor, features of the N pairs by applying a deep convolutional neural network to the N pairs and to the class labels;
calculating, by the processor for each of the N pairs based on the features, a respective similarly measure between the respective anchor example and the respective non-anchor example;
calculating, by the processor, a similarity score based on the respective similarit measure for each of the N pairs, the similarity score representing one or more similarities betwee all anchor points and all positive training examples in the N pairs relative to one or more similarities between all of the anchor points and all negative training examples in the N pairs; and
maximizing, by the processor, the similarity score for the respective anchor example for each of the N pairs to pull togethe in a distribution space the training examples from a same one of the plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurali ty of classes.
2. The computer-implemented method of claim 1 , wherein each of the pairs of training examples corresponds' to a differen one of the plurality of classes,
3. The computer-implemented method of claim 2, wherein the plurality Of classes are randomly selected as a subset from a set of classes, and wherein the set of classes includes the plurality of classes and one or more other classes.
4. The computer-implemented method of claim 1. wherem said maximizing step is capable of simultaneously pushing N-1 examples away from, a single reference sample from among the N pairs of training examples, in the distribution space .
5. The computer -implemented method of claim 1 , wherein 'said maximizing step is capable of simultaneously pushing N-1 examples towards a single reference sample from among the pairs of training examples, in the distribution space.
6. The computer-implemented method of claim 1, wherein the deep convolutions! neural network is configured to include embedding vectors that are tramed to satisfy a set of constraints on each loss function in a set of loss functions, wherein the deep convolutions! neural network is trained using the set of loss functions.
7. The computer-implemented method of claim L wherein said maximizing step comprising computing a gradient of a logarithm of the similarity score.
8. The computer implemented method of claim 1 , wherein said maxim zing s ep'inae&imizes.-aa'-obj^tiye -function for deep metric learning.
9. The computer im ktrsen ted method of claim 1 , wherein a total number of the plurality of classes at least one of (i) changes over time, (ii) is larger than a threshold amount, and (in) is unknown.
10. The computer-implemented method of claim 1. further comprising verifying a user and providing the user access to an entity, based on a prediction generated using tie deep convoiniioiial neural network.
11. A non-transitory article of manufacture tangibly embodying a -computer- readable program which when executed causes a computer to perform the steps of claim 1.
12. A system, comprising:
a processor, configured to;
receive, N pairs of training examples and class labels for the training examples that correspond to a plurality of classes, wherein eac of the N pairs includes a respective anchor example and further includes a respective non-anchor example capable of being a positive training example or a negative training example; extract features of the N pairs by applying a deep convolutions! neural network to the I pairs and to the class labels;
calculate, for each of the N pairs based on the features, a respective similarly measure between the respective anchor example and the respective non-aftdior example;
calculate a similarity score based o the respective similarity measure for each of the pairs, the similarity score .representing one or more similarities between all anchor points and all positive training examples in the N pairs relative to one or more similarities between all of the anchor points and all negati ve training examples in the N pairs; and
maximize the similarity score for the respective anchor example for each of the N pairs to pull together in a distribution space the training examples from same one of the plurality of classes while pushing apart in the distribution space the training examples from different ones of the plurality of classes.
13, The system of claim 12, wherein each of the pairs of the training examples corresponds to a different one of the plurality of classes and
14. The system of claim 13, wherein the processor is configured to randomly selected the plurality of classes as a subset from a set of classes, and wherein the set of classes includes the plurality of classes and one or more other classes.
15. The system of claim 12, wherein said processor is configured to
simultaneously push, in the distribution space, N-l examples away from a single
reference sample from among the N pairs of trai ing examples, responsive to a
.maximization of the similarity score,
1 . The system of claim 12, -wherein said processor is configured to simultaneously push, in the distribution space, N-l examples towards a single reference sample from among the N pairs of training examples, responsive to a maximization of the similarity score.
17. The system of claim 12, wherein the deep conventional neural network is configured to include embedding vectors that are trained to satisfy a set of constraints on. each loss function in a set of loss functions,' wherein the deep convokiiional neural network is trained using the set of loss functions.
18. The system, of claim 1.2, wherein said processor is configured to maximize the similarity score by computing 'a gradient of a logarithm of the similarity score .
19. The computer-implemented method, of claim 12, wherein a total number of the plurality of classes at least one of (i) changes over time, fii) is larger than a threshold amount, and (iii) is unknown. 20, The system of claim 12, w'iierem said processor i s further configured to verify a user and provide the user access to an entity, based on a prediction generated using the deep coiivoiiitioiia! neural network.
PCT/US2016/067946 2016-02-04 2016-12-21 Improving distance metric learning with n-pair loss Ceased WO2017136060A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
DE112016006360.1T DE112016006360T5 (en) 2016-02-04 2016-12-21 IMPROVING LEARNING OF DISTANCE METHOD WITH AN N-PAIR LOSS
JP2018540162A JP2019509551A (en) 2016-02-04 2016-12-21 Improvement of distance metric learning by N pair loss

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US201662291025P 2016-02-04 2016-02-04
US62/291,025 2016-02-04
US15/385,283 US10565496B2 (en) 2016-02-04 2016-12-20 Distance metric learning with N-pair loss
US15/385,283 2016-12-20

Publications (1)

Publication Number Publication Date
WO2017136060A1 true WO2017136060A1 (en) 2017-08-10

Family

ID=59497846

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2016/067946 Ceased WO2017136060A1 (en) 2016-02-04 2016-12-21 Improving distance metric learning with n-pair loss

Country Status (4)

Country Link
US (1) US10565496B2 (en)
JP (1) JP2019509551A (en)
DE (1) DE112016006360T5 (en)
WO (1) WO2017136060A1 (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12287848B2 (en) 2021-06-11 2025-04-29 International Business Machines Corporation Learning Mahalanobis distance metrics from data
US12530873B2 (en) 2022-02-28 2026-01-20 Panasonic Automotive Systems Co., Ltd. Training method

Families Citing this family (40)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10649970B1 (en) 2013-03-14 2020-05-12 Invincea, Inc. Methods and apparatus for detection of functionality
US9690938B1 (en) 2015-08-05 2017-06-27 Invincea, Inc. Methods and apparatus for machine learning based malware detection
US10115032B2 (en) * 2015-11-04 2018-10-30 Nec Corporation Universal correspondence network
WO2017223294A1 (en) 2016-06-22 2017-12-28 Invincea, Inc. Methods and apparatus for detecting whether a string of characters represents malicious activity using machine learning
KR102648770B1 (en) * 2016-07-14 2024-03-15 매직 립, 인코포레이티드 Deep neural network for iris identification
GB2555192B (en) * 2016-08-02 2021-11-24 Invincea Inc Methods and apparatus for detecting and identifying malware by mapping feature data into a semantic space
US10740596B2 (en) * 2016-11-08 2020-08-11 Nec Corporation Video security system using a Siamese reconstruction convolutional neural network for pose-invariant face recognition
US10540961B2 (en) * 2017-03-13 2020-01-21 Baidu Usa Llc Convolutional recurrent neural networks for small-footprint keyword spotting
US10387749B2 (en) * 2017-08-30 2019-08-20 Google Llc Distance metric learning using proxies
EP3688750B1 (en) * 2017-10-27 2024-03-13 Google LLC Unsupervised learning of semantic audio representations
KR102535411B1 (en) * 2017-11-16 2023-05-23 삼성전자주식회사 Apparatus and method related to metric learning based data classification
CN109815971B (en) * 2017-11-20 2023-03-10 富士通株式会社 Information processing method and information processing device
CN108922542B (en) * 2018-06-01 2023-04-28 平安科技(深圳)有限公司 Sample triplet acquisition method and device, computer equipment and storage medium
CN109256139A (en) * 2018-07-26 2019-01-22 广东工业大学 A kind of method for distinguishing speek person based on Triplet-Loss
US11537872B2 (en) 2018-07-30 2022-12-27 International Business Machines Corporation Imitation learning by action shaping with antagonist reinforcement learning
US11734575B2 (en) 2018-07-30 2023-08-22 International Business Machines Corporation Sequential learning of constraints for hierarchical reinforcement learning
US11501157B2 (en) 2018-07-30 2022-11-15 International Business Machines Corporation Action shaping from demonstration for fast reinforcement learning
US11636123B2 (en) * 2018-10-05 2023-04-25 Accenture Global Solutions Limited Density-based computation for information discovery in knowledge graphs
CN111325223B (en) * 2018-12-13 2023-10-24 中国电信股份有限公司 Training method, device and computer-readable storage medium for deep learning model
CN110032645B (en) * 2019-04-17 2021-02-09 携程旅游信息技术(上海)有限公司 Text emotion recognition method, system, device and medium
JP7262290B2 (en) * 2019-04-26 2023-04-21 株式会社日立製作所 A system for generating feature vectors
KR102522894B1 (en) * 2019-05-22 2023-04-18 한국전자통신연구원 Method of training image deep learning model and device thereof
US11720790B2 (en) 2019-05-22 2023-08-08 Electronics And Telecommunications Research Institute Method of training image deep learning model and device thereof
CN110532880B (en) * 2019-07-29 2022-11-22 深圳大学 Sample screening and expression recognition method, neural network, device and storage medium
KR102635606B1 (en) * 2019-11-21 2024-02-13 고려대학교 산학협력단 Brain-computer interface apparatus based on feature extraction reflecting similarity between users using distance learning and task classification method using the same
CN111339891A (en) * 2020-02-20 2020-06-26 苏州浪潮智能科技有限公司 Target detection method of image data and related device
CN111400591B (en) * 2020-03-11 2023-04-07 深圳市雅阅科技有限公司 Information recommendation method and device, electronic equipment and storage medium
CN111667050B (en) * 2020-04-21 2021-11-30 佳都科技集团股份有限公司 Metric learning method, device, equipment and storage medium
KR20220166355A (en) * 2020-04-21 2022-12-16 구글 엘엘씨 Supervised and Controlled Learning Using Multiple Positive Examples
CN111797893B (en) * 2020-05-26 2021-09-14 华为技术有限公司 Neural network training method, image classification system and related equipment
CN113742288B (en) * 2020-05-29 2024-09-24 伊姆西Ip控股有限责任公司 Method, electronic device and computer program product for data indexing
US12468952B2 (en) * 2020-06-02 2025-11-11 Salesforce, Inc. Systems and methods for noise-robust contrastive learning
JP7425445B2 (en) * 2020-07-17 2024-01-31 日本電信電話株式会社 Feature learning device, feature extraction device, feature learning method and program
JP7420278B2 (en) 2020-09-29 2024-01-23 日本電気株式会社 Information processing device, information processing method, and recording medium
US20220101117A1 (en) * 2020-09-29 2022-03-31 International Business Machines Corporation Neural network systems for abstract reasoning
CN112329833B (en) * 2020-10-28 2022-08-12 浙江大学 An Image Metric Learning Method Based on Spherical Embedding
KR102577342B1 (en) * 2021-01-20 2023-09-11 네이버 주식회사 Computer system for learning with memory-based virtual class for metric learning, and method thereof
CN113408299B (en) 2021-06-30 2022-03-25 北京百度网讯科技有限公司 Training method, apparatus, device and storage medium for semantic representation model
US20240211796A1 (en) * 2022-12-22 2024-06-27 Microsoft Technology Licensing, Llc Explanation of emergent semantics in embedding spaces via analogy
US12518523B2 (en) * 2023-10-12 2026-01-06 Vinbrain Joint Stock Company System and method for performing distance metric learning using proxies

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20080319973A1 (en) * 2007-06-20 2008-12-25 Microsoft Corporation Recommending content using discriminatively trained document similarity
US20110219012A1 (en) * 2010-03-02 2011-09-08 Yih Wen-Tau Learning Element Weighting for Similarity Measures
US20120323968A1 (en) * 2011-06-14 2012-12-20 Microsoft Corporation Learning Discriminative Projections for Text Similarity Measures
US20150186495A1 (en) * 2013-12-31 2015-07-02 Quixey, Inc. Latent semantic indexing in application classification

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5706402A (en) * 1994-11-29 1998-01-06 The Salk Institute For Biological Studies Blind signal processing system employing information maximization to recover unknown signals through unsupervised minimization of output redundancy
US10115032B2 (en) * 2015-11-04 2018-10-30 Nec Corporation Universal correspondence network

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20080319973A1 (en) * 2007-06-20 2008-12-25 Microsoft Corporation Recommending content using discriminatively trained document similarity
US20110219012A1 (en) * 2010-03-02 2011-09-08 Yih Wen-Tau Learning Element Weighting for Similarity Measures
US20120323968A1 (en) * 2011-06-14 2012-12-20 Microsoft Corporation Learning Discriminative Projections for Text Similarity Measures
US20150186495A1 (en) * 2013-12-31 2015-07-02 Quixey, Inc. Latent semantic indexing in application classification

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
GUOSHENG HU ET AL.: "When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition", ICCV, vol. 2015, 2015, pages 142 - 150, XP032864988, Retrieved from the Internet <URL:http://www.cv-foundation.org/openaccess/content_iccv_2015_workshops/wll/html/Hu_When_Face_Recognition_ICCV_2015_paper.html> *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12287848B2 (en) 2021-06-11 2025-04-29 International Business Machines Corporation Learning Mahalanobis distance metrics from data
US12530873B2 (en) 2022-02-28 2026-01-20 Panasonic Automotive Systems Co., Ltd. Training method

Also Published As

Publication number Publication date
US20170228641A1 (en) 2017-08-10
DE112016006360T5 (en) 2018-10-11
JP2019509551A (en) 2019-04-04
US10565496B2 (en) 2020-02-18

Similar Documents

Publication Publication Date Title
WO2017136060A1 (en) Improving distance metric learning with n-pair loss
JP7024515B2 (en) Learning programs, learning methods and learning devices
Huang et al. Federated learning with long-tailed data via representation unification and classifier rectification
JP7124427B2 (en) Multi-view vector processing method and apparatus
US12254250B2 (en) Mask estimation device, mask estimation method, and mask estimation program
US8266083B2 (en) Large scale manifold transduction that predicts class labels with a neural network and uses a mean of the class labels
Frei et al. Self-training converts weak learners to strong learners in mixture models
CN116976461A (en) Federated learning methods, devices, equipment and media
EP4095757A1 (en) Learning deep latent variable models by short-run mcmc inference with optimal transport correction
CN114048729A (en) Medical literature evaluation methods, electronic devices, storage media and program products
KR102026226B1 (en) Method for extracting signal unit features using variational inference model based deep learning and system thereof
US10614343B2 (en) Pattern recognition apparatus, method, and program using domain adaptation
CN114467095A (en) Locally Interpretable Models Based on Reinforcement Learning
CN110442733A (en) A kind of subject generating method, device and equipment and medium
Wu et al. Acoustic to articulatory mapping with deep neural network
CN115019128A (en) Image generation model training method, image generation method and related device
Wang et al. Federated multi-phase curriculum learning to synchronously correlate user heterogeneity
JP7103235B2 (en) Parameter calculation device, parameter calculation method, and parameter calculation program
US20230136609A1 (en) Method and apparatus for unsupervised domain adaptation
CN112560982A (en) CNN-LDA-based semi-supervised image label generation method
JP6193823B2 (en) Sound source number estimation device, sound source number estimation method, and sound source number estimation program
JP6636973B2 (en) Mask estimation apparatus, mask estimation method, and mask estimation program
Mengara The last Dance: Robust backdoor attack via diffusion models and bayesian approach
CN115700788A (en) Method, apparatus and computer program product for image recognition
Nguyen et al. An adaptive EM accelerator for unsupervised learning of Gaussian mixture models

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16889657

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2018540162

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 112016006360

Country of ref document: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16889657

Country of ref document: EP

Kind code of ref document: A1