WO2016141972A1 - Two-factor authentication based on ambient sound - Google Patents

Two-factor authentication based on ambient sound Download PDF

Info

Publication number
WO2016141972A1
WO2016141972A1 PCT/EP2015/054937 EP2015054937W WO2016141972A1 WO 2016141972 A1 WO2016141972 A1 WO 2016141972A1 EP 2015054937 W EP2015054937 W EP 2015054937W WO 2016141972 A1 WO2016141972 A1 WO 2016141972A1
Authority
WO
WIPO (PCT)
Prior art keywords
ambient sound
sound sample
user
sample
software application
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2015/054937
Other languages
French (fr)
Inventor
Claudio Soriente
Nikolaos KARAPANOS
Claudio MARFORIO
Srdjan Capkun
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Eidgenoessische Technische Hochschule Zurich ETHZ
Original Assignee
Eidgenoessische Technische Hochschule Zurich ETHZ
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Eidgenoessische Technische Hochschule Zurich ETHZ filed Critical Eidgenoessische Technische Hochschule Zurich ETHZ
Priority to PCT/EP2015/054937 priority Critical patent/WO2016141972A1/en
Publication of WO2016141972A1 publication Critical patent/WO2016141972A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F21/00Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
    • G06F21/30Authentication, i.e. establishing the identity or authorisation of security principals
    • G06F21/31User authentication
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/06Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being correlation coefficients
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L63/00Network architectures or network communication protocols for network security
    • H04L63/08Network architectures or network communication protocols for network security for authentication of entities
    • H04L63/0853Network architectures or network communication protocols for network security for authentication of entities using an additional device, e.g. smartcard, SIM or a different communication terminal

Definitions

  • the present invention pertains to a method for
  • the present invention further pertains to a software application for a mobile device, such as a pocketable or wearable device, for assisting authentication of a user based on such a method.
  • Authentication is the process of verifying the validity of the identity of an entity, e.g. of a person.
  • An important use of authentication is access control. For instance a computer system that is supposed to be accessed only by authorised users must attempt to detect and exclude non- authorised persons. Access to the computer system is therefore usually controlled by means of an authentication procedure to establish with some degree of confidence the identity of the user, and then granting privileges
  • buildings, vehicles, safes, vaults (i.e. physical entry control) or automatic teller machines (ATMs) are further examples where authentication is typically desired.
  • ATMs automatic teller machines
  • the present invention provides a method for authenticating a user, comprising the steps of:
  • a user credential such as a user password, pass phrase, personal identification number (PIN), challenge response or biometric identifier, for instance together with a username, as input in a first device, in particular as input in a web browser running on the first device;
  • PIN personal identification number
  • a second device located in close proximity of the first device, in particular within a distance of 10 m, more particularly within a distance of 5 m, most particularly within a distance of 2 m, e.g. by means of a microphone of the second device;
  • the first device may for instance be a computer
  • the second device may for instance be a mobile phone, such as a smartphone, or other personal portable, pocketable or wearable device of the user, such as a personal digital assistant, media player, activity tracker, headset, hearing device or smartwatch.
  • the proposed method is essentially transparent to the user in that just like in a password-only authentication
  • the user only has to enter his username and password, e.g. into a web browser running on the first device.
  • password e.g.
  • Existing two-factor authentication methods transfer a one-time code from a token to the browser via the user.
  • the user is relieved from this task, and instead it is verified that the first device, e.g. a laptop from which the user is trying to log in, and a second personal device of the user, e.g. his smartphone, are co- located by comparing ambient sound samples (i.e. brief recordings) captured by their microphones.
  • the proposed method is essentially transparent to the user because he is not asked to interact with the second device (e.g. his smartphone) in any way. Moreover, it is possible to determine that the first device and the second device are co-located even if the second device (e.g. the smartphone).
  • smartphone is for instance in the user's pocket or purse.
  • the method further comprises determining a sound level, such as a sound pressure level (SPL) , of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold.
  • SPL sound pressure level
  • the method further comprises requesting the user to make a sound, such as to speak (e.g. an arbitrary word) , cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table.
  • a sound such as to speak (e.g. an arbitrary word) , cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table.
  • capturing of the first ambient sound sample and capturing of the second ambient sound sample is time synchronised, i.e. is offset by less than a predefined maximum value, such as for instance 150 milliseconds or more particularly 75
  • the method further comprises shifting the first ambient sound sample with respect to the second ambient sound sample in order to achieve time- alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
  • determining the measure of similarity comprises determining a cross- correlation of the first ambient sound sample and the second ambient sound sample.
  • determining the measure of similarity comprises decomposing the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, for instance by means of band-pass filtering, determining a per- component cross-correlation of the first ambient sound sample and the second ambient sound sample for each
  • the method further comprises sending the first ambient sound sample from the first device to the second device, in particular via a server.
  • the method further comprises encrypting the first ambient sound sample with a key, such as a public key, in particular of (i.e. associated with) the second device.
  • a key such as a public key, in particular of (i.e. associated with) the second device.
  • the method further comprises a server providing the key, in particular of the second device, to the first device.
  • determining the measure of similarity of the first ambient sound sample and the second ambient sound sample is performed by the second device .
  • the first ambient sound sample and/or the second ambient sound sample are identical to each other.
  • the method further comprises applying lossy audio data compression to the first ambient sound sample and/or the second ambient sound sample.
  • the method further comprises applying lossy audio data compression to the first ambient sound sample and/or the second ambient sound sample.
  • the method further comprises down-sampling the first ambient sound sample and/or the second ambient sound sample, i.e. reducing the sampling rate of either or both of these signals by means of decimation.
  • down-sampling the first ambient sound sample and/or the second ambient sound sample i.e. reducing the sampling rate of either or both of these signals by means of decimation.
  • the method further comprises the second device indicating to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
  • the method further comprises repeating (at least once but preferably not more than a predefined maximum number of times) the steps of capturing a first ambient sound sample, capturing a second ambient sound sample, and determining a measure of similarity of the first ambient sound sample and the second ambient sound sample when the measure of similarity is within a
  • determined measure of similarity is in a critical range (i.e. borderline to being acceptably similar, thus possibly leading to a false positive authentication of the user) .
  • a software application for a mobile device is proposed, in particular a native mobile application, more particularly a mobile client application, comprising program code
  • the software application being adapted to:
  • the software application is further adapted to determine a sound level, such as a sound pressure level (SPL), of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold.
  • a sound level such as a sound pressure level (SPL)
  • SPL sound pressure level
  • the software application is further adapted to reguest the user to make a sound, such as to speak (e.g. an arbitrary word), cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table.
  • a sound such as to speak (e.g. an arbitrary word), cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table.
  • the software application is further adapted to shift the first ambient sound sample with respect to the second ambient sound sample in order to achieve time-alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
  • the software application is further adapted to determine a cross-correlation of the first ambient sound sample and the second ambient sound sample as part of determining the measure of similarity.
  • the software application is further adapted to decompose the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, for instance by means of band-pass
  • filtering to determine a per-component cross-correlation of the first ambient sound sample and the second ambient sound sample for each frequency band component, and to determine an aggregate statistical value, such as an average, based on the per-component cross-correlations as part of determining the measure of similarity.
  • the software application is further adapted to decrypt the first ambient sound sample, in particular encrypted with a key, such as a public key of (i.e. associated with) the mobile device.
  • the software application is further adapted to provide the key, in particular of the mobile device, to a server or to the other device.
  • the first ambient sound sample and/or the second ambient sound sample have/has a length in the range from 2 to 5 seconds, in particular in the range from 3 to 4 seconds.
  • the software application is further adapted to apply lossy audio data decompression to the first ambient sound sample or to up-sample the first ambient sound sample, i.e. to increase its sampling rate by means of interpolation.
  • the software application is further adapted to indicate to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
  • - generating a sound by the mobile device e.g. by means of a loudspeaker.
  • the software application is further adapted to repeat (at least once but preferably not more than a predefined maximum number of times) the steps of capturing a first ambient sound sample, capturing a second ambient sound sample, and determining a measure of
  • the mobile device is a mobile phone, such as a smartphone, or other personal portable, pocketable or wearable device of the user, such as a personal digital assistant, media player, activity tracker, headset, hearing device or smartwatch .
  • FIG. 1 a high-level block diagram of an exemplary setup for performing the method according to the present invention
  • Fig. 2 a high-level block diagram of an exemplary setup for performing the method according to the present invention
  • Fig. 3a a high-level flow diagram of an exemplary
  • Fig. 3b a high-level flow diagram of an exemplary
  • Fig. 1 depicts a high-level block diagram of a general setup for (bowser-based) remote authentication.
  • the user has a username and password to authenticate to a server 1.
  • the server 1 implements a two-factor authentication mechanism that leverages the user' s mobile phone 3 as a software token.
  • the user visits the server's web-page using a browser installed on his laptop 2.
  • the user enters his credentials, e.g. username and
  • the server 1 verifies the validity of the password and challenges the user to prove possession of the second authentication factor, i.e. the user's mobile phone 3. Authentication is successful only if the user convinces the server 1 that he indeed possesses the second authentication factor
  • the proposed two-factor authentication mechanism determines proximity of a user' s mobile phone 3 and the laptop 2 where he is attempting to authenticate from by computing a similarity score between the ambient sound captured by their respective microphones 4 & 5. For privacy reasons clear text sound samples are not uploaded to the server 1. This approach would allow the server 1 to arbitrarily query the user' s mobile phone 3 for a recording of the mobile phone's ambient sound. The laptop 2 encrypts its audio sample under the public key of the mobile phone 3. The mobile phone 3 receives the encrypted ambient sound sample from the laptop 2, decrypts it, and compares it against the ambient sound sample recorded locally. Finally, the mobile phone 3 tells the server 1 whether the two devices (i.e. the laptop 2 and the mobile phone 3) are co-located or not. All the communication between the laptop 2 and the mobile phone 3 goes through the server 1.
  • Comparison of the two ambient sound samples is for instance achieved by determining a measure of their similarity (i.e. by calculating a similarity score) e.g. in terms of the cross-correlation r x , y (l):
  • x(i), y(i) denote the two ambient sound samples, and 1 is the lag applied to y.
  • the cross- correlation may be normalised as:
  • the normalisation maps f x y (l to the range [-1, 1].
  • a value of x y ( ) 1 indicates that at the lag 1, the two ambient sound samples have the exact same shape even if the amplitudes may be different.
  • Two ambient sound samples with a normalised cross- correlation value above a predefined threshold Tcc are deemed as "legitimate", i.e. as provided by two co-located devices.
  • the adversary must also know when the sound happens with a time error of less than lmax.
  • the average power of the two ambient sound samples is computed and pairs are discarded where either ambient sound sample has an average power below a predetermined threshold Td B - This is done to avoid comparing sound recordings of silence or very quiet noises, like the buzz of a fridge, the noise of a fan or the ticks of a clock. Moreover, this prevents an adversary from exploiting the quietness in a victim' s surroundings to improve the chances of a fraudulent login being accepted.
  • An alternative measure of the similarity of the two ambient sound samples takes into consideration both time domain and frequency domain information by employing both band-pass filtering, e.g. one-third octave band filtering, and per-band cross- correlation. Splitting a signal in one-third octave bands provides high frequency resolution information of the original signal while maintaining its time-domain
  • the audible range of frequencies (i.e. from 20 Hz to 20 kHz) is split into eleven non-overlapping octave bands where the ratio of the highest in-band
  • Each octave is represented by its centre frequency, where the centre frequency of a particular octave is twice the centre frequency of the previous octave.
  • One-third octave bands split the first ten octave bands in three and the last octave band in two, for a total of 32 bands.
  • One-third octave bands are widely used in acoustics and their
  • the centre frequency of the lowest band is 16 Hz (covering from 14.1 Hz to 17.8 Hz) while the centre frequency of the highest band is 20 kHz (covering from 17780 Hz to 22390 Hz) .
  • B [LB, HB] denotes a set of contiguous one- third octave bands, from the band that has its centre frequency at LB to the band that has its centre frequency at HB.
  • Fig. 2 shows a high-level block diagram of a function for determining such an alternative measure of similarity or similarity score.
  • Each audio signal is input to a bank of band-pass filters to obtain n signal components, one per each of the one-third octave bands that are taken into account.
  • Xi is the signal component for the i-th one-third octave band of the signal x.
  • the similarity score is the average of the maximum cross-correlation x . y . (i) over the pairs of signal components xi, yi :
  • Fig. 3a shows an overview of the proposed two-factor authentication process.
  • the user points the browser to the uniform resource locator (URL) of the server 1 and enters his username and password.
  • the server 1 retrieves the public key of the user' s mobile phone 3 and sends it to the browser running on the laptop 2.
  • Both the browser and the mobile phone 3 start recording ambient sound through their local microphones 4 & 5 for t seconds (with 2 s ⁇ t ⁇ 5 s) .
  • the browser encrypts the sound sample under the mobile phone' s public key and sends it to the mobile phone 3, using the server 1 as a proxy.
  • the mobile phone 3 decrypts the sound sample recorded by the browser and compares it against the one recorded locally.
  • the mobile phone 3 concludes that it is co-located with the laptop 2 from which the authentication attempt originated and informs the server 1 that the login attempt is
  • the mobile phone application Upon enrolment, the mobile phone application generates a fresh key-pair (2048 bit RSA) and sends the public key to the server 1 which associates the public key to the user' s account (i.e. with the mobile phone 3) . Every time the user logs in from a browser, the server 1 initiates the proposed two-factor authentication mechanism, as
  • the browser sends the username and password to the server 1 (step 1), which in turn triggers a push notification for the mobile phone application (step 2) .
  • step 1) sends the username and password to the server 1
  • step 2 triggers a push notification for the mobile phone application.
  • the server 1 then notifies the browser to start recording (step 4) and sends it the mobile phone's public key.
  • the browser encrypts the sound sample/file and uploads the cipher text to the server 1 (step 5) .
  • the recording is encrypted with AES256 using a fresh key and then the key is encrypted with the public key using RSA2048.
  • the mobile phone 3 fetches the encrypted sound sample recorded by the laptop 2 (step 6) , decrypts it with its private key, and finally compares it with the sound sample recorded locally. If the
  • the proposed solution relies on a lag-bounded cross-correlation between the two ambient sound samples to discriminate between legitimate and fraudulent logins.
  • the bound on the lag reguires the two recordings to be closely synchronised.
  • a simple time- synchronisation protocol based on the standard NTP (Network Time Protocol) can be used.
  • the protocol can be
  • each device 2, 3 runs the time- synchronisation protocol with the server 1 while it is recording sound via its microphone 4, 5 (as shown in Fig.
  • each device 2, 3 adjusts the timestamp of its sound sample taking into account the clock difference with the server 1.
  • the present invention proposes a usable two-factor
  • the two sound recordings are compared on the mobile phone 3.
  • An approach where samples are compared at the server 1 allows the server 1 to ask for a recording of the mobile phone's ambient sound at any time. This is because the user does not interact with the mobile phone 3 so he cannot allow or deny the recording by, for example, pressing a button. Recordings of the sound captured by the mobile phone 3 are preferably not uploaded to the server 1.
  • the laptop 2 encrypts its recording under the mobile phone's public key and the two sound samples are compared on the mobile phone 3.
  • the mobile phone 3 only outputs a binary answer in respect of the laptop 2 being co-located with the mobile phone 3.
  • the application of the present invention can be applied within the context of a very broad range of access control mechanisms.
  • the application of the present invention is not confined to computer systems, e.g. for achieving access to a server from a client via the World Wide Web, but also applicable in conjunction with physical entry control systems, e.g. for gaining access to buildings, vehicles, safes or vaults, or for user authentication at automatic teller machines

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computer Security & Cryptography (AREA)
  • Signal Processing (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Computer Hardware Design (AREA)
  • Human Computer Interaction (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Telephone Function (AREA)

Abstract

The present invention proposes a method for two-factor authenticating a user based a user credential, such as a user password, as well as ambient sound samples captured by a first device (2), e.g. a laptop running a web browser, and a co-located second device (3), e.g. a smartphone. A measure of similarity (or similarity score) of the first ambient sound sample and the second ambient sound sample is determined, and the user is positively authenticated if both the received user credential is valid and the measure of similarity is above a predefined threshold, otherwise authentication of the user fails. In a further aspect of the present invention a software application ("app") for a mobile device (3) is proposed for assisting authentication of a user based on the provided method.

Description

TWO-FACTOR AUTHENTICATION BASED ON AMBIENT SOUND
TECHNICAL FIELD
The present invention pertains to a method for
authenticating a user, in particular a scheme based on two- factor authentication, for instance for achieving access to a server from a client via the World Wide Web. The present invention further pertains to a software application for a mobile device, such as a pocketable or wearable device, for assisting authentication of a user based on such a method.
BACKGROUND OF THE INVENTION
Authentication is the process of verifying the validity of the identity of an entity, e.g. of a person. An important use of authentication is access control. For instance a computer system that is supposed to be accessed only by authorised users must attempt to detect and exclude non- authorised persons. Access to the computer system is therefore usually controlled by means of an authentication procedure to establish with some degree of confidence the identity of the user, and then granting privileges
established for that identity. Access control to
buildings, vehicles, safes, vaults (i.e. physical entry control) or automatic teller machines (ATMs) are further examples where authentication is typically desired. Common authentication procedures require a user to enter a
username and a password, the latter being a "secret" that preferably only the user knows. However, such an
authentication scheme is vulnerable to password "leakage", when a user unintentionally exposes his password to an attacker/adversary or a password database is breached. In order to improve security two-factor authentication is therefore commonly employed, whereby user identification is achieved by means of the combination of two different components, for instance by something the user knows (e.g. a password, pass phrase, personal identification number (PIN), or challenge response where the user must answer a question) and something the user has (e.g. a dedicated security token or mobile phone capable of providing a software token such as a mobile transaction authentication number (mTAN) ) . Smartphones are rapidly replacing
dedicated security tokens in two-factor authentication mechanisms. Using a mobile phone in place of a hardware token dramatically improves deployability and also betters the usability of two-factor authentication. From the service provider' s point of view, two-factor authentication based on mobile phones results in a substantial reduction of manufacturing and shipping costs. From the user's perspective, there is no hardware to carry (since users always carry their mobile phones with them anyway) and a mobile phone can provide/accommodate tokens for multiple service providers. However, a vast majority of users prefers password-only authentication mechanisms for services where two-factor authentication is not mandatory. One of the reasons behind two-factor authentication being rather unpopular is the extra steps a user must complete to log in. In current two-factor authentication schemes the user must for
instance copy a one-time code from a dedicated hardware token to a web browser.
Hence, there is a need for improved two-factor
authentication schemes, which do not place an extra burden on a user, i.e. which are essentially transparent to a user, and therefore result in increased user acceptance.
SUMMARY OF THE INVENTION
It is an object of the present invention to provide an improved method for two-factor authentication which
substantially relieves a user from having to perform additional steps over those required by commonly employed single-factor, e.g. password-only authentication schemes. This object is achieved by the method according to claim 1.
It is a further object of the present invention to provide means for assisting authentication of a user based on the proposed method. This further object is achieved by the software application according to claim 15. Various specific embodiments of the method and software application according to the present invention are given in the dependent claims .
The present invention provides a method for authenticating a user, comprising the steps of:
- receiving a user credential, such as a user password, pass phrase, personal identification number (PIN), challenge response or biometric identifier, for instance together with a username, as input in a first device, in particular as input in a web browser running on the first device;
- capturing/recording a first ambient sound sample by the first device, e.g. by means of a microphone of the first device ;
- capturing/recording a second ambient sound sample by a second device located in close proximity of the first device, in particular within a distance of 10 m, more particularly within a distance of 5 m, most particularly within a distance of 2 m, e.g. by means of a microphone of the second device;
- determining a measure of similarity (or similarity
score) of the first ambient sound sample and the second ambient sound sample; and
- positively authenticating the user if both the received user credential is valid/correct and the measure of similarity is above a predefined first threshold, otherwise failing authentication of the user.
The first device may for instance be a computer,
workstation, laptop, tablet, safe, vault, automatic teller machine (ATM), building or vehicle entry terminal (e.g. a door unlocking device) . The second device may for instance be a mobile phone, such as a smartphone, or other personal portable, pocketable or wearable device of the user, such as a personal digital assistant, media player, activity tracker, headset, hearing device or smartwatch.
The proposed method is essentially transparent to the user in that just like in a password-only authentication
session, the user only has to enter his username and password, e.g. into a web browser running on the first device. Likewise, when the user wishes to enter a building or a vehicle, or when wanting to retrieve valuables from a safe or to withdraw money from an ATM the user is only required to enter his password or PIN. Existing two-factor authentication methods transfer a one-time code from a token to the browser via the user. According to the proposed method the user is relieved from this task, and instead it is verified that the first device, e.g. a laptop from which the user is trying to log in, and a second personal device of the user, e.g. his smartphone, are co- located by comparing ambient sound samples (i.e. brief recordings) captured by their microphones. Dependent on the outcome of determining whether the first device is in the same environment as the second device (i.e. the two devices are co-located) it is decided whether the login attempt is legitimate or fraudulent. As mentioned above the proposed method is essentially transparent to the user because he is not asked to interact with the second device (e.g. his smartphone) in any way. Moreover, it is possible to determine that the first device and the second device are co-located even if the second device (e.g. the
smartphone) is for instance in the user's pocket or purse.
In an embodiment the method further comprises determining a sound level, such as a sound pressure level (SPL) , of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold. Only accepting sound samples that have enough power and discarding low-power sound samples ensures a certain minimal loudness of the sound in the user's
surroundings (for instance silence or the buzz of a fridge are not accepted) , which guarantees robustness against adversaries that try to guess the second authentication factor (i.e. the sound around the victim's second device, e.g. his smartphone). Otherwise, an adversary could exploit the guietness in a victim's surroundings, e.g. when the victim is sleeping, to improve the chances of a
fraudulent login being accepted. In a further embodiment the method further comprises requesting the user to make a sound, such as to speak (e.g. an arbitrary word) , cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table. In this way it is ensured that the proposed method reliably works even in quiet surroundings .
In a further embodiment of the method capturing of the first ambient sound sample and capturing of the second ambient sound sample is time synchronised, i.e. is offset by less than a predefined maximum value, such as for instance 150 milliseconds or more particularly 75
milliseconds. Synchronising the starting time of capturing (i.e. recording) the sound samples at the first and second devices drastically reduces the chances of an impersonation attack even if the adversary correctly guesses the sound in the victim's surroundings.
In a further embodiment the method further comprises shifting the first ambient sound sample with respect to the second ambient sound sample in order to achieve time- alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds. By time-aligning the two sound samples it is possible to extract relevant information from their time domain representations, thus making the process of
comparing the two sound samples, e.g. by determining a measure of their similarity, simpler. In a further embodiment of the method determining the measure of similarity comprises determining a cross- correlation of the first ambient sound sample and the second ambient sound sample.
In a further embodiment of the method determining the measure of similarity comprises decomposing the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, for instance by means of band-pass filtering, determining a per- component cross-correlation of the first ambient sound sample and the second ambient sound sample for each
frequency band component, and determining an aggregate statistical value, such as an average, based on the per- component cross-correlations. In this way both time as well as frequency domain information of the captured ambient sound samples is taken into account in determined the measure of similarity.
In a further embodiment the method further comprises sending the first ambient sound sample from the first device to the second device, in particular via a server.
In a further embodiment the method further comprises encrypting the first ambient sound sample with a key, such as a public key, in particular of (i.e. associated with) the second device. In a further embodiment the method further comprises a server providing the key, in particular of the second device, to the first device.
In a further embodiment of the method determining the measure of similarity of the first ambient sound sample and the second ambient sound sample is performed by the second device .
In a further embodiment of the method the first ambient sound sample and/or the second ambient sound sample
have/has a length in the range from 2 to 5 seconds, in particular in the range from 3 to 4 seconds.
In a first alternative of a further embodiment the method further comprises applying lossy audio data compression to the first ambient sound sample and/or the second ambient sound sample. In a second alternative of a further
embodiment the method further comprises down-sampling the first ambient sound sample and/or the second ambient sound sample, i.e. reducing the sampling rate of either or both of these signals by means of decimation. In this way it is possible to reduce the time it takes to carry out the authentication process by decreasing the transmission time of the first and/or the second sound sample to the location whether the two sound samples are compared (i.e. where the measure of similarity is determined) . Moreover, it is more time efficient to transfer the recorded sound sample from the source to the destination by streaming the sound data as soon as the recording starts, instead of waiting for the recording to finish and then transferring the sound sample (or a compressed version thereof), e.g. as a single batch of data.
In a further embodiment the method further comprises the second device indicating to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
- vibration of the second device;
- lighting up a display or optical indicator of the second device ;
- displaying a message on a display of the second device; - generating a sound by the second device, e.g. by means of a loudspeaker.
Because sound recording and comparison is completely transparent to the user, the user has no means to tell him if he is being victim of an impersonation attack. By alerting the user that an authentication attempt is taking place the chances that the user catches an attempt to impersonate him (i.e. to become aware of a fraudulent authentication attempt) is increased.
In a further embodiment the method further comprises repeating (at least once but preferably not more than a predefined maximum number of times) the steps of capturing a first ambient sound sample, capturing a second ambient sound sample, and determining a measure of similarity of the first ambient sound sample and the second ambient sound sample when the measure of similarity is within a
predefined range [Tcc,min, Tec] . In this way additional sound recordings are triggered in cases where the
determined measure of similarity is in a critical range (i.e. borderline to being acceptably similar, thus possibly leading to a false positive authentication of the user) .
It is pointed out that combinations of the above-mentioned embodiments of the proposed method can yield even further, more specific embodiments of the method according to the present invention.
In a further aspect of the present invention a software application ("app") for a mobile device is proposed, in particular a native mobile application, more particularly a mobile client application, comprising program code
executable by the mobile device for assisting
authentication of a user of the mobile device, the software application being adapted to:
- receive a first ambient sound sample from another
device ; - capture a second ambient sound sample by the mobile
device; - determine a measure of similarity (or similarity score) of the first ambient sound sample and the second ambient sound sample; and
- provide an affirmation that the mobile device and the other device are co-located when the measure of
similarity is above a predefined first threshold, otherwise providing a negation, in particular to an authentication entity, such as a server.
In an embodiment the software application is further adapted to determine a sound level, such as a sound pressure level (SPL), of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold.
In a further embodiment the software application is further adapted to reguest the user to make a sound, such as to speak (e.g. an arbitrary word), cough, clear his throat, whistle or hum, or to knock or scratch, e.g. on a table.
In a further embodiment the software application is further adapted to shift the first ambient sound sample with respect to the second ambient sound sample in order to achieve time-alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
In a further embodiment the software application is further adapted to determine a cross-correlation of the first ambient sound sample and the second ambient sound sample as part of determining the measure of similarity.
In a further embodiment the software application is further adapted to decompose the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, for instance by means of band-pass
filtering, to determine a per-component cross-correlation of the first ambient sound sample and the second ambient sound sample for each frequency band component, and to determine an aggregate statistical value, such as an average, based on the per-component cross-correlations as part of determining the measure of similarity.
In a further embodiment the software application is further adapted to decrypt the first ambient sound sample, in particular encrypted with a key, such as a public key of (i.e. associated with) the mobile device.
In a further embodiment the software application is further adapted to provide the key, in particular of the mobile device, to a server or to the other device. In a further embodiment of the software application the first ambient sound sample and/or the second ambient sound sample have/has a length in the range from 2 to 5 seconds, in particular in the range from 3 to 4 seconds.
In a further embodiment the software application is further adapted to apply lossy audio data decompression to the first ambient sound sample or to up-sample the first ambient sound sample, i.e. to increase its sampling rate by means of interpolation.
In a further embodiment the software application is further adapted to indicate to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
- vibration of the mobile device;
- lighting up a display or optical indicator of the mobile device ; - displaying a message on a display of the mobile device;
- generating a sound by the mobile device, e.g. by means of a loudspeaker.
In a further embodiment the software application is further adapted to repeat (at least once but preferably not more than a predefined maximum number of times) the steps of capturing a first ambient sound sample, capturing a second ambient sound sample, and determining a measure of
similarity of the first ambient sound sample and the second ambient sound sample when the measure of similarity is within a predefined (critical) range [TCc, min , Tcc] .
In a further embodiment of the software application the mobile device is a mobile phone, such as a smartphone, or other personal portable, pocketable or wearable device of the user, such as a personal digital assistant, media player, activity tracker, headset, hearing device or smartwatch .
It is pointed out that combinations of the above-mentioned embodiments of the proposed software application can yield even further, more specific embodiments of the software application according to the present invention.
BRIEF DESCRIPTION OF THE DRAWINGS
The present invention is further explained below by means of non-limiting specific embodiments and with reference to the accompanying drawings, which show:
Fig. 1 a high-level block diagram of an exemplary setup for performing the method according to the present invention; Fig. 2 a high-level block diagram of an exemplary
function for determining a measure of similarity or similarity score of the ambient sound samples;
Fig. 3a a high-level flow diagram of an exemplary
embodiment of the method according to the present invention; and
Fig. 3b a high-level flow diagram of an exemplary
implementation of an embodiment of the method according to the present invention.
DETAILED DESCRIPTION OF THE INVENTION
Fig. 1 depicts a high-level block diagram of a general setup for (bowser-based) remote authentication. The user has a username and password to authenticate to a server 1. The server 1 implements a two-factor authentication mechanism that leverages the user' s mobile phone 3 as a software token. The user visits the server's web-page using a browser installed on his laptop 2. At this point the user enters his credentials, e.g. username and
password. The server 1 verifies the validity of the password and challenges the user to prove possession of the second authentication factor, i.e. the user's mobile phone 3. Authentication is successful only if the user convinces the server 1 that he indeed possesses the second
authentication factor. The proposed two-factor authentication mechanism determines proximity of a user' s mobile phone 3 and the laptop 2 where he is attempting to authenticate from by computing a similarity score between the ambient sound captured by their respective microphones 4 & 5. For privacy reasons clear text sound samples are not uploaded to the server 1. This approach would allow the server 1 to arbitrarily query the user' s mobile phone 3 for a recording of the mobile phone's ambient sound. The laptop 2 encrypts its audio sample under the public key of the mobile phone 3. The mobile phone 3 receives the encrypted ambient sound sample from the laptop 2, decrypts it, and compares it against the ambient sound sample recorded locally. Finally, the mobile phone 3 tells the server 1 whether the two devices (i.e. the laptop 2 and the mobile phone 3) are co-located or not. All the communication between the laptop 2 and the mobile phone 3 goes through the server 1.
Comparison of the two ambient sound samples is for instance achieved by determining a measure of their similarity (i.e. by calculating a similarity score) e.g. in terms of the cross-correlation rx,y(l):
Figure imgf000018_0001
where x(i), y(i) denote the two ambient sound samples, and 1 is the lag applied to y. To accommodate for different amplitudes of the two ambient sound samples, the cross- correlation may be normalised as: The normalisation maps fx y (l to the range [-1, 1]. A value of x y( ) = 1 indicates that at the lag 1, the two ambient sound samples have the exact same shape even if the amplitudes may be different. To speed up the computation the cross-correlation theorem can be leveraged, which states that rx,y(l) = F_1(F(x(i) ) * · F ( y ( i ) ) ) , where F() denotes the discrete Fourier transform (DFT) and the asterisk denotes the complex conjugate. Thereby, the DFT may be efficiently and rapidly computed by means of the fast Fourier transform (FFT).
Two ambient sound samples with a normalised cross- correlation value above a predefined threshold Tcc are deemed as "legitimate", i.e. as provided by two co-located devices. The lag 1 of the cross-correlation is bounded to lmax. This means that the two ambient sound samples must be closely synchronised, i.e. captured/recorded at essentially the same time, specifically they are offset by less than lmax = 150 ms (or more particularly less than 75 ms) . This is in order to thwart attacks where the adversary
successfully guesses the sound in the victim's (user's) surroundings at the time of the attack. For example, even if the adversary correctly guesses that in the victim' s surroundings somebody is coughing or is clinking a metal spoon against a cup and is able to produce a very similar sound, the adversary must also know when the sound happens with a time error of less than lmax. Moreover, the average power of the two ambient sound samples is computed and pairs are discarded where either ambient sound sample has an average power below a predetermined threshold TdB- This is done to avoid comparing sound recordings of silence or very quiet noises, like the buzz of a fridge, the noise of a fan or the ticks of a clock. Moreover, this prevents an adversary from exploiting the quietness in a victim' s surroundings to improve the chances of a fraudulent login being accepted.
An alternative measure of the similarity of the two ambient sound samples (i.e. an alternative similarity score) takes into consideration both time domain and frequency domain information by employing both band-pass filtering, e.g. one-third octave band filtering, and per-band cross- correlation. Splitting a signal in one-third octave bands provides high frequency resolution information of the original signal while maintaining its time-domain
representation .
As an example, the audible range of frequencies (i.e. from 20 Hz to 20 kHz) is split into eleven non-overlapping octave bands where the ratio of the highest in-band
frequency to the lowest in-band frequency is 2 to 1. Each octave is represented by its centre frequency, where the centre frequency of a particular octave is twice the centre frequency of the previous octave. One-third octave bands split the first ten octave bands in three and the last octave band in two, for a total of 32 bands. One-third octave bands are widely used in acoustics and their
frequency ranges have been standardised. The centre frequency of the lowest band is 16 Hz (covering from 14.1 Hz to 17.8 Hz) while the centre frequency of the highest band is 20 kHz (covering from 17780 Hz to 22390 Hz) . In the following B = [LB, HB] denotes a set of contiguous one- third octave bands, from the band that has its centre frequency at LB to the band that has its centre frequency at HB.
Fig. 2 shows a high-level block diagram of a function for determining such an alternative measure of similarity or similarity score. Each audio signal is input to a bank of band-pass filters to obtain n signal components, one per each of the one-third octave bands that are taken into account. Xi is the signal component for the i-th one-third octave band of the signal x. The similarity score is the average of the maximum cross-correlation x.y. (i) over the pairs of signal components xi, yi :
Figure imgf000021_0001
where 1 is bounded between 0 and lmax . The set B of one- third octave bands to consider for determining this
similarity score is an important parameter. A spectral region should be selected that i) includes most common sounds and ii) is robust to attenuation and directionality of the sounds to be recorded by the laptop 2 and the mobile phone 3. The frequencies below 50 Hz should be disregarded in order to remove very low-frequency noise. Moreover frequencies above 8 kHz should be disregarded because sounds at these frequencies are attenuated by fabric and therefore are not suitable for scenarios where the mobile phone 3 is in a pocket or purse. The similarity score Sx,y is preferably determined over the set of one-third octave bands B = [50 Hz, 4 kHz], which provides a good performance in respect of both usability and security in terms of false rejection rate (i.e. a legitimate login being rejected) and false acceptance rate (i.e. a fraudulent login being accepted), respectively.
Fig. 3a shows an overview of the proposed two-factor authentication process. The user points the browser to the uniform resource locator (URL) of the server 1 and enters his username and password. The server 1 retrieves the public key of the user' s mobile phone 3 and sends it to the browser running on the laptop 2. Both the browser and the mobile phone 3 start recording ambient sound through their local microphones 4 & 5 for t seconds (with 2 s ≤ t ≤ 5 s) . After recording, the browser encrypts the sound sample under the mobile phone' s public key and sends it to the mobile phone 3, using the server 1 as a proxy. The mobile phone 3 decrypts the sound sample recorded by the browser and compares it against the one recorded locally. If the average power of both sound samples is above the predefined threshold TdB and the cross-correlation between the two sound samples is above the predefined threshold TCc, the mobile phone 3 concludes that it is co-located with the laptop 2 from which the authentication attempt originated and informs the server 1 that the login attempt is
legitimate. A correlation threshold value Tcc in the range between 0.1 and 0.13, in particular a threshold value of Tcc = 0.13 was determined to be a good trade-off between usability and security in terms of false rejection rate and false acceptance rate, respectively. An average power threshold value of TdB = 40 dB is suitable to identify silence or in general very quiet surroundings and then discard such ambient sound samples.
Upon enrolment, the mobile phone application generates a fresh key-pair (2048 bit RSA) and sends the public key to the server 1 which associates the public key to the user' s account (i.e. with the mobile phone 3) . Every time the user logs in from a browser, the server 1 initiates the proposed two-factor authentication mechanism, as
illustrated in Fig. 3b. The browser sends the username and password to the server 1 (step 1), which in turn triggers a push notification for the mobile phone application (step 2) . As soon as the mobile phone 3 receives the
notification it replies to the server 1 to notify reception (step 3) and starts recording. The server 1 then notifies the browser to start recording (step 4) and sends it the mobile phone's public key. After recording, the browser encrypts the sound sample/file and uploads the cipher text to the server 1 (step 5) . The recording is encrypted with AES256 using a fresh key and then the key is encrypted with the public key using RSA2048. The mobile phone 3 fetches the encrypted sound sample recorded by the laptop 2 (step 6) , decrypts it with its private key, and finally compares it with the sound sample recorded locally. If the
comparison yields a value above Tcc and the average power of both sound samples is above TdB the mobile phone 3 sends a positive message to the server 1 (step 7), otherwise it sends a negative one.
As already mentioned, the proposed solution relies on a lag-bounded cross-correlation between the two ambient sound samples to discriminate between legitimate and fraudulent logins. The bound on the lag reguires the two recordings to be closely synchronised. For this reason a simple time- synchronisation protocol based on the standard NTP (Network Time Protocol) can be used. The protocol can be
implemented over HTTP and allows each device to compute the difference between the local clock and the one of the server. To save time, each device 2, 3 runs the time- synchronisation protocol with the server 1 while it is recording sound via its microphone 4, 5 (as shown in Fig.
3b). When recording is completed, each device 2, 3 adjusts the timestamp of its sound sample taking into account the clock difference with the server 1.
The present invention proposes a usable two-factor
authentication mechanism with the goal of increasing the adoption rate over presently known two-factor
authentication schemes. In order to protect the privacy of the user from a prying server 1 the two sound recordings are compared on the mobile phone 3. An approach where samples are compared at the server 1 allows the server 1 to ask for a recording of the mobile phone's ambient sound at any time. This is because the user does not interact with the mobile phone 3 so he cannot allow or deny the recording by, for example, pressing a button. Recordings of the sound captured by the mobile phone 3 are preferably not uploaded to the server 1. The laptop 2 encrypts its recording under the mobile phone's public key and the two sound samples are compared on the mobile phone 3. The mobile phone 3 only outputs a binary answer in respect of the laptop 2 being co-located with the mobile phone 3.
It is expressly pointed out that the proposed method for authenticating a user and the supporting software
application can be applied within the context of a very broad range of access control mechanisms. The application of the present invention is not confined to computer systems, e.g. for achieving access to a server from a client via the World Wide Web, but also applicable in conjunction with physical entry control systems, e.g. for gaining access to buildings, vehicles, safes or vaults, or for user authentication at automatic teller machines
(ATMs) .

Claims

1. A method for authenticating a user, comprising the steps of:
- receiving a user credential, such as a user password, pass phrase, personal identification number, challenge response or biometric identifier, as input in a first device (2 ) ;
- capturing a first ambient sound sample by the first
device ( 2 ) ;
- capturing a second ambient sound sample by a second
device (3) located in close proximity of the first device (2), in particular within a distance of 10 m, more particularly within a distance of 5 m, most
particularly within a distance of 2 m;
- determining a measure of similarity of the first ambient sound sample and the second ambient sound sample; and
- positively authenticating the user if both the received user credential is valid and the measure of similarity is above a predefined first threshold, otherwise failing authentication of the user.
2. The method of claim 1, further comprising determining a sound level of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold.
3. The method of claim 1 or 2, further comprising
requesting the user to make a sound.
4. The method of one of claims 1 to 3, wherein capturing of the first ambient sound sample and capturing of the second ambient sound sample is time synchronised, in particular is offset by less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
5. The method of one of claims 1 to 4, further comprising shifting the first ambient sound sample with respect to the second ambient sound sample in order to achieve time- alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
6. The method of one of claims 1 to 5, wherein determining the measure of similarity comprises determining a cross- correlation of the first ambient sound sample and the second ambient sound sample.
7. The method of one of claims 1 to 5, wherein determining the measure of similarity comprises decomposing the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, determining a per-component cross-correlation of the first ambient sound sample and the second ambient sound sample for each frequency band component, and determining an aggregate statistical value, such as an average, based on the per- component cross-correlations.
8. The method of one of claims 1 to 7, further comprising sending the first ambient sound sample from the first device (2) to the second device (3), in particular via a server ( 1 ) .
9. The method of one of claims 1 to 8, further comprising encrypting the first ambient sound sample, in particular with a key of the second device (3) .
10. The method of claim 9, further comprising a server (1) providing the key, in particular of the second device (3), to the first device (2).
11. The method of one of claims 1 to 10, wherein
determining the measure of similarity of the first ambient sound sample and the second ambient sound sample is performed by the second device (3) .
12. The method of one of claims 1 to 11, wherein the first ambient sound sample and/or the second ambient sound sample have/has a length in the range from 2 to 5 seconds, in particular in the range from 3 to 4 seconds.
13. The method of one of claims 1 to 12, further comprising applying lossy audio data compression to or down-sampling the first ambient sound sample and/or the second ambient sound sample.
14. The method of one of claims 1 to 13, further comprising the second device (3) indicating to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
- vibration of the second device (3);
- lighting up a display or optical indicator (7) of the second device (3);
- displaying a message on a display of the second device (3) ;
- generating a sound by the second device (3), e.g. by
means of a loudspeaker (6).
15. A software application for a mobile device (3), in particular a native mobile application, more particularly a mobile client application, comprising program code
executable by the mobile device (3) for assisting
authentication of a user of the mobile device (3), the software application being adapted to: - receive a first ambient sound sample from another device (2) ;
- capture a second ambient sound sample by the mobile
device ( 3 ) ; - determine a measure of similarity of the first ambient sound sample and the second ambient sound sample; and
- provide an affirmation that the mobile device (3) and the other device (2) are co-located when the measure of similarity is above a predefined first threshold, otherwise providing a negation, in particular to an authentication entity, such as a server (1).
16. The software application of claim 15, being further adapted to determine a sound level of the first ambient sound sample and/or of the second ambient sound sample, and discarding the first ambient sound sample and/or the second ambient sound sample if the determined sound level of the first ambient sound sample and/or the second ambient sound sample is below a predefined second threshold.
17. The software application of claim 15 or 16, being further adapted to request the user to make a sound.
18. The software application of one of claims 15 to 17, being further adapted to shift the first ambient sound sample with respect to the second ambient sound sample in order to achieve time-alignment, in particular with an offset of less than 150 milliseconds, more particularly with an offset of less than 75 milliseconds.
19. The software application of one of claims 15 to 18, being further adapted to determine a cross-correlation of the first ambient sound sample and the second ambient sound sample as part of determining the measure of similarity.
20. The software application of one of claims 15 to 18, being further adapted to decompose the first ambient sound sample and the second ambient sound sample into a plurality of frequency band components, for instance by means of band-pass filtering, to determine a per-component cross- correlation of the first ambient sound sample and the second ambient sound sample for each frequency band
component, and to determine an aggregate statistical value, such as an average, based on the per-component cross- correlations as part of determining the measure of
similarity.
21. The software application of one of claims 15 to 20, being further adapted to decrypt the first ambient sound sample, in particular encrypted with a key of the mobile device ( 3 ) .
22. The software application of claim 21, being further adapted to provide the key, in particular of the mobile device (3) to a server (1) or to the other device (2) .
23. The software application of one of claims 15 to 22, wherein the first ambient sound sample and/or the second ambient sound sample have/has a length in the range from 2 to 5 seconds, in particular in the range from 3 to 4 seconds .
24. The software application of one of claims 15 to 23, being further adapted to apply lossy audio data
decompression to the first ambient sound sample or to up- sample the first ambient sound sample
25. The software application of one of claims 15 to 24, being further adapted to indicate to the user when an authentication attempt is taking place, in particular by means of at least one of the following:
- vibration of the mobile device (3);
- lighting up a display or optical indicator (7) of the mobile device (3); - displaying a message on a display of the mobile device (3) ;
- generating a sound by the mobile device (3), e.g. by means of a loudspeaker (6) .
26. The software application of one of claims 15 to 25, wherein the mobile device (3) is a mobile phone, such as a smartphone, or other personal portable, pocketable or wearable device of the user, such as a personal digital assistant, media player, activity tracker, headset, hearing device or smartwatch.
PCT/EP2015/054937 2015-03-10 2015-03-10 Two-factor authentication based on ambient sound Ceased WO2016141972A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/EP2015/054937 WO2016141972A1 (en) 2015-03-10 2015-03-10 Two-factor authentication based on ambient sound

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/EP2015/054937 WO2016141972A1 (en) 2015-03-10 2015-03-10 Two-factor authentication based on ambient sound

Publications (1)

Publication Number Publication Date
WO2016141972A1 true WO2016141972A1 (en) 2016-09-15

Family

ID=52727085

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2015/054937 Ceased WO2016141972A1 (en) 2015-03-10 2015-03-10 Two-factor authentication based on ambient sound

Country Status (1)

Country Link
WO (1) WO2016141972A1 (en)

Cited By (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10375083B2 (en) 2017-01-25 2019-08-06 International Business Machines Corporation System, method and computer program product for location verification
US20190386984A1 (en) * 2018-06-14 2019-12-19 Paypal, Inc. Two-factor authentication through ultrasonic audio transmissions
US10623403B1 (en) 2018-03-22 2020-04-14 Pindrop Security, Inc. Leveraging multiple audio channels for authentication
US10665244B1 (en) 2018-03-22 2020-05-26 Pindrop Security, Inc. Leveraging multiple audio channels for authentication
US10873461B2 (en) 2017-07-13 2020-12-22 Pindrop Security, Inc. Zero-knowledge multiparty secure sharing of voiceprints
US11482231B2 (en) 2020-01-06 2022-10-25 Vmware, Inc. Skill redirections in a voice assistant
US11509479B2 (en) 2019-06-04 2022-11-22 Vmware, Inc. Service authentication through a voice assistant
US11570165B2 (en) 2019-12-09 2023-01-31 Vmware, Inc. Single sign-on service authentication through a voice assistant
US11765595B2 (en) 2019-06-26 2023-09-19 Vmware, Inc. Proximity based authentication of a user through a voice assistant device
US11830098B2 (en) 2020-01-02 2023-11-28 Vmware, Inc. Data leak prevention using user and device contexts
US12063214B2 (en) * 2020-01-02 2024-08-13 VMware LLC Service authentication through a voice assistant
US12088585B2 (en) 2020-01-06 2024-09-10 VMware LLC Voice skill session lifetime management
CN119094159A (en) * 2024-08-01 2024-12-06 易君刚 A multi-factor mobile phone security authentication method based on integrated voice mobile network
CN119485306A (en) * 2025-01-10 2025-02-18 北京久佳信通科技有限公司 A login method and system based on 5G message voiceprint verification code

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013125414A1 (en) * 2012-02-23 2013-08-29 日本電気株式会社 Mutual authentication system, mutual authentication server, mutual authentication method, and mutual authentication program
US8949958B1 (en) * 2011-08-25 2015-02-03 Amazon Technologies, Inc. Authentication using media fingerprinting

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8949958B1 (en) * 2011-08-25 2015-02-03 Amazon Technologies, Inc. Authentication using media fingerprinting
WO2013125414A1 (en) * 2012-02-23 2013-08-29 日本電気株式会社 Mutual authentication system, mutual authentication server, mutual authentication method, and mutual authentication program

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
TRUONG HIEN THI THU ET AL: "Comparing and fusing different sensor modalities for relay attack resistance in Zero-Interaction Authentication", 2014 IEEE INTERNATIONAL CONFERENCE ON PERVASIVE COMPUTING AND COMMUNICATIONS (PERCOM), IEEE, 24 March 2014 (2014-03-24), pages 163 - 171, XP032594197, DOI: 10.1109/PERCOM.2014.6813957 *
TZIPORA HALEVI ET AL: "Secure Proximity Detection for NFC Devices Based on Ambient Sensor Data", 10 September 2012, COMPUTER SECURITY ESORICS 2012, SPRINGER BERLIN HEIDELBERG, BERLIN, HEIDELBERG, PAGE(S) 379 - 396, ISBN: 978-3-642-33166-4, XP047014032 *

Cited By (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10375083B2 (en) 2017-01-25 2019-08-06 International Business Machines Corporation System, method and computer program product for location verification
US10873461B2 (en) 2017-07-13 2020-12-22 Pindrop Security, Inc. Zero-knowledge multiparty secure sharing of voiceprints
US10623403B1 (en) 2018-03-22 2020-04-14 Pindrop Security, Inc. Leveraging multiple audio channels for authentication
US10665244B1 (en) 2018-03-22 2020-05-26 Pindrop Security, Inc. Leveraging multiple audio channels for authentication
US20190386984A1 (en) * 2018-06-14 2019-12-19 Paypal, Inc. Two-factor authentication through ultrasonic audio transmissions
US11509479B2 (en) 2019-06-04 2022-11-22 Vmware, Inc. Service authentication through a voice assistant
US11765595B2 (en) 2019-06-26 2023-09-19 Vmware, Inc. Proximity based authentication of a user through a voice assistant device
US11570165B2 (en) 2019-12-09 2023-01-31 Vmware, Inc. Single sign-on service authentication through a voice assistant
US11830098B2 (en) 2020-01-02 2023-11-28 Vmware, Inc. Data leak prevention using user and device contexts
US12063214B2 (en) * 2020-01-02 2024-08-13 VMware LLC Service authentication through a voice assistant
US11482231B2 (en) 2020-01-06 2022-10-25 Vmware, Inc. Skill redirections in a voice assistant
US12088585B2 (en) 2020-01-06 2024-09-10 VMware LLC Voice skill session lifetime management
CN119094159A (en) * 2024-08-01 2024-12-06 易君刚 A multi-factor mobile phone security authentication method based on integrated voice mobile network
CN119485306A (en) * 2025-01-10 2025-02-18 北京久佳信通科技有限公司 A login method and system based on 5G message voiceprint verification code

Similar Documents

Publication Publication Date Title
Karapanos et al. {Sound-Proof}: Usable {Two-Factor} authentication based on ambient sound
US9712526B2 (en) User authentication for social networks
US10665244B1 (en) Leveraging multiple audio channels for authentication
US8862888B2 (en) Systems and methods for three-factor authentication
Han et al. Proximity-proof: Secure and usable mobile two-factor authentication
US10762181B2 (en) System and method for user confirmation of online transactions
US8516562B2 (en) Multi-channel multi-factor authentication
US11436311B2 (en) Method and apparatus for secure and usable mobile two-factor authentication
US9730001B2 (en) Proximity based authentication using bluetooth
US20180151182A1 (en) System and method for multi-factor authentication using voice biometric verification
US20080313707A1 (en) Token-based system and method for secure authentication to a service provider
US10623403B1 (en) Leveraging multiple audio channels for authentication
US20190213306A1 (en) System and method for identity authentication
US9882719B2 (en) Methods and systems for multi-factor authentication
Shrestha et al. Listening watch: Wearable two-factor authentication using speech signals resilient to near-far attacks
US10425407B2 (en) Secure transaction and access using insecure device
US9853971B2 (en) Proximity based authentication using bluetooth
US9461987B2 (en) Audio authentication system
Zhu et al. Quickauth: Two-factor quick authentication based on ambient sound
Alattar et al. Privacy‐preserving hands‐free voice authentication leveraging edge technology
Shrestha et al. Sound-based two-factor authentication: Vulnerabilities and redesign
EP2560122B1 (en) Multi-Channel Multi-Factor Authentication
Prasad A comparative study of passwordless authentication
Prasad Breaking Barriers: Passwordless Authentication as the Future of Security
Jayaram et al. Improved Mobile Security Based on Two‐Factor Authentication

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15711674

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15711674

Country of ref document: EP

Kind code of ref document: A1