WO2014196971A1 - Voice recognition and identification - Google Patents

Voice recognition and identification Download PDF

Info

Publication number
WO2014196971A1
WO2014196971A1 PCT/US2013/044413 US2013044413W WO2014196971A1 WO 2014196971 A1 WO2014196971 A1 WO 2014196971A1 US 2013044413 W US2013044413 W US 2013044413W WO 2014196971 A1 WO2014196971 A1 WO 2014196971A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
speaker
pattern
sample
identifying information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2013/044413
Other languages
French (fr)
Inventor
Mark Alan Schultz
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Thomson Licensing SAS
Original Assignee
Thomson Licensing SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Thomson Licensing SAS filed Critical Thomson Licensing SAS
Priority to PCT/US2013/044413 priority Critical patent/WO2014196971A1/en
Publication of WO2014196971A1 publication Critical patent/WO2014196971A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques

Definitions

  • a voice recognition method for identifying a person on a display screen based on when the person's voice is recognized is disclosed.
  • the method can be implemented in a number of systems where such identification is desirable, such as
  • the method comprises obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier
  • FFT Transform
  • the system comprises a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; a microphone for receiving a second voice sample; a display screen; and a processor in communication with the microphone, the memory, and the display screen.
  • FFT Fast Fourier Transform
  • the processor is configured to perform the following process: convert the second voice sample into a voice pattern using FFT; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 ⁇ R ⁇ 1 ; determine whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, instruct the display screen to display the identifying information associated with the speaker profile.
  • FIG. 1 is a high-level flow chart showing a voice sample collection method accordance with an embodiment of the present invention
  • FIG. 2a is an FFT image of a sample of a first speaker's voice to be used in connection with a profile of the first speaker;
  • FIG. 2b is an FFT image of a sample of a second speaker's voice to be used in connection with a profile of the second speaker;
  • FIG. 2c is an FFT image of a sample of a third speaker's voice to be used in connection with a profile of the third speaker;
  • FIG. 2d is an FFT image of a sample of a fourth speaker's voice to be used in connection with a profile of the fourth speaker;
  • FIG. 3 is an FFT image of the captured audio of a conversation where multiple speakers spoke
  • FIG. 4 is a schematic view of a voice recognition system cross correlating the
  • FIG. 5a is a schematic view of a display screen showing a first person's identifying information
  • FIG. 5b is a schematic view of a display screen showing a second person's identifying information.
  • the present disclosure generally relates to a voice-recognition system and method where a person speaking through an audio communication medium is identified to the receiving party by populating a profile unique to the speaker on a display screen on the receiving party's end.
  • the system does so by sampling a speaker's voice while he or she is speaking and, through spectral analysis and the use of sliding Fast Fourier Transforms ("FFTs"), cross-correlates each sample to a number of voice patterns stored in the system, where each voice pattern is associated with a profile.
  • FFTs sliding Fast Fourier Transforms
  • the system sends a signal to an electronic device with a display screen on the listening party's end, the signal indicating that the profile associated with the matched voice pattern should appear on the display screen.
  • This profile will identify who the speaker is to the listening party so as to remove any confusion on the listening party's part as to who may be speaking at that given time.
  • these elements are implemented in various forms of hardware, software or combinations thereof.
  • these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input/output interfaces.
  • general-purpose devices which may include a processor, memory and input/output interfaces.
  • Other elements can be implemented through the use of specifically- purposed devices, such as microphones, audio speakers, and electronic display screens.
  • FIG. 1 a high-level flow chart of voice sample collection method for teleconferencing 100 in accordance with an embodiment of the present invention is shown.
  • this embodiment involves at least two speaking persons.
  • multiple persons participate in the call by using their own telephone receivers to call in to a teleconferencing service.
  • one set of multiple persons are using one telephone receiver to speak to a second set of multiple persons on a second telephone receiver.
  • each person on the call provides a sample of his or her voice to the system (steps 104(a)-(n)) before connecting to the call. In one embodiment, this is done at a stage where each person is asked to state his or her name before being connected.
  • the system Upon receiving a sample of each person's voice, the system performs a spectral analysis of the voice sample via a Fast Fourier Transform ("FFT") to create a spectral pattern showing the dominant frequencies in the person's voice (steps 106(a)-(n)).
  • FFT Fast Fourier Transform
  • Each spectral pattern is then stored in a memory as a signature pattern for each individual speaker, respectively (steps 108(a)-(n)).
  • Each signature pattern is then associated with an individual unique speaker profile (steps 110(a)-(n)).
  • the signature patterns and the speaker profiles are then stored in the system's memory (step 112) and the speakers on the call are then connected to each other to conduct the conference call (step 114).
  • FIGs. 2a-d show four examples of speaker profiles 200a-d that were collected from four individual speakers A-D, respectively, using the embodiment illustrated in FIG. 1.
  • each of speaker profiles 200a-d includes a signature pattern (e.g., one of signature patterns 202a-d, respectively) which illustrates three dominant vocal frequencies that
  • signature pattern 202a for individual speaker A shows three dominant frequencies of 1 kHz, 500 Hz, and 250 Hz that characterize the voice of individual speaker A
  • signature pattern 202b for individual speaker B shows three dominant frequencies of 800 Hz, 600 Hz, and 200 Hz that characterize the voice of individual speaker B.
  • signature patterns 202a-d are used as bases upon which voice identification can take place during a conference call between individual speakers A-D. This voice identification method is discussed further below.
  • each of the speaker profiles 200a-d also includes speaker identification information (not shown) particular to the speaker associated with the speaker profile.
  • This speaker identification information is used to inform others on the conference call who is speaking once the voice identification method has identified a particular speaker as speaking at a given time during the call.
  • the identification information includes the speaker's name.
  • the identification information can include, but is not limited to, the speaker's name and a picture of the speaker's face and the like.
  • FIG. 3 illustrates an FFT profile 300 of a captured conversation during such a conference call.
  • the FFT profile 300 exhibits where windowed filters captured samples of the conversation at a-q instances in time.
  • the a-q instances represent the intervals of time at which a voice sample was captured during the conversation. In one embodiment, the intervals are 0.5 seconds long. In another embodiment, the intervals are 0.25 seconds long.
  • a spectral pattern is given representing the dominant frequencies, if any at all, that were present during the conversation at that given instance.
  • the spectral graph 300 shows nine instances where a voice sample was taken and converted via FFT into a voice pattern (e.g., voice patterns 302- 318).
  • FIG. 4 illustrates how speakers are identified from voice samples taken during the conversation.
  • each of the signature patterns 202a-d found in the speaker profiles 200a-d is slid across the FFT profile of the conversation to determine whether the person associated with the selected signature pattern is speaking.
  • the system then cross correlates each of the signature patterns 202a-d against the converted voice sample captured at that instance of time.
  • Each cross correlation produces a correlation coefficient R having a value between 0 and 1 , where 1 indicates a perfect match and a 0 indicates no match.
  • the correlation efficient R produces a value above a predetermined threshold, the system sends a signal indicating that the individual speaker associated with the selected signature pattern is currently speaking.
  • signature pattern 202d from speaker profile 200d is slid across the conversation's FFT profile 300 and cross correlated with the captured sample at each of the a-q instances in time during the conversation.
  • the cross correlation coefficient R will be 0, since no one was speaking during those instances of time.
  • the cross correlation of voice pattern 302 with signature pattern 202d results in an R value of 1 , indicating that speaker D is speaking at that instance in time. This is also the case for instances o and q, where correlations with voice patterns 316 and 318 also result in an R value of 1.
  • a voice sample produces a voice pattern that is not the same as that of speaker D (such as voice patterns 304, 306, 308, 310, 312, and 314)
  • the cross correlation with signature pattern 202d reveals an R value between 0 and 1.
  • a predetermined threshold for R is set high enough such that other voice patterns correlated with signature pattern 202d do not cause the system to signal false positive identifications.
  • a predetermined threshold for R is 0.8.
  • FIGS. 5a and 5b illustrate examples of a display screen 500 displaying identification information 502a, 502b upon receiving a signal from the system to do so. To use the examples shown in FIGS.
  • the display screen 500 displays the identification information 502a from speaker profile 200a when it receives a signal indicating a match with signature pattern 202a and displays the identification information 502b from speaker profile 200b when it receives a signal indicating a match with signature pattern 202b. This allows those on the call within view of the display screen 500 to see that speakers A and B are speaking on the call when their respective identification information 502a, 502b appears on the display screen 500.
  • the system can also send identification information to multiple devices such as mobile devices including, but not limited to, laptops, tablets and/or mobile phones and the like. This can be accomplished via wired (e.g., Ethernet cables, etc.) and/or wireless means (e.g., Bluetooth, WiFi, etc.). This allows easier access to the identification information. It also permits each user to customize how the information is displayed on their device. For example, when a senior person with high authority speaks, the information can be shown in red or bordered by a red border to gain the user's attention more quickly.
  • the system can also take the voice patterns which have been identified as a match to one of the signature patterns and average them together with the identified signature pattern as the conversation progresses. This allows the system to learn the characteristics of each person's voice as the conversation progresses and can potentially lead to more accurate identifications of the various speakers during the call.
  • the various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof.
  • the software is preferably implemented as an application program tangibly embodied on a program storage unit or computer readable medium.
  • the application program may be uploaded to, and executed by, a machine comprising any suitable architecture.
  • the machine is implemented on a computer platform having hardware such as one or more central processing units ("CPUs"), a memory, and input/output interfaces.
  • CPUs central processing units
  • the computer platform may also include an operating system and microinstruction code.
  • the various processes and functions described herein may be either part of the microinstruction code or part of the application program, or any combination thereof, which may be executed by a CPU, whether or not such computer or processor is explicitly shown.
  • various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Telephonic Communication Services (AREA)

Abstract

A voice recognition method and system for identifying a person on a display screen based on when the person's voice is recognized is disclosed. The method includes obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; receiving a second voice sample; converting the second voice sample into a voice pattern using FFT; cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 ≤ R ≤ 1; and sending the identifying information associated with the speaker profile when the cross correlation coefficient R is above a threshold.

Description

VOICE RECOGNITION AND IDENTIFICATION
BACKGROUND
[0001] As teleconferencing and video conferencing increase in usage, knowing who is speaking at any given time is important. Many times, the persons on a teleconferencing call may know each other's names but have never met before and are not familiar with each other's voices. In such circumstances, when one person speaks during a teleconferencing call, the other persons on the call may not know who is speaking without the person self-identifying him or herself during the call. However, not every person on a call may remember to self-identify during the call, which may cause confusion for the other persons on the call and cause them to ask who was speaking, which wastes time.
SUMMARY
[0002] In view of the foregoing, a voice recognition method for identifying a person on a display screen based on when the person's voice is recognized is disclosed. The method can be implemented in a number of systems where such identification is desirable, such as
teleconferencing, video conferencing, multiplayer video game platforms, and the like.
[0003] In one embodiment, the method comprises obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier
Transform (FFT) of a first voice sample of the speaker's voice; receiving a second voice sample; converting the second voice sample into a voice pattern using FFT; cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 < R < 1; determining whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, communicating with the display screen to display the identifying information associated with the speaker profile.
[0004] Also disclosed is a system for identifying a person on a display screen based on voice recognition. In one embodiment, the system comprises a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; a microphone for receiving a second voice sample; a display screen; and a processor in communication with the microphone, the memory, and the display screen. In this embodiment, the processor is configured to perform the following process: convert the second voice sample into a voice pattern using FFT; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 < R < 1 ; determine whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, instruct the display screen to display the identifying information associated with the speaker profile.
BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For a more complete understanding of the present invention, reference is made to the following detailed description of an embodiment considered in conjunction with the accompanying drawings, in which:
[0006] FIG. 1 is a high-level flow chart showing a voice sample collection method accordance with an embodiment of the present invention;
[0007] FIG. 2a is an FFT image of a sample of a first speaker's voice to be used in connection with a profile of the first speaker;
[0008] FIG. 2b is an FFT image of a sample of a second speaker's voice to be used in connection with a profile of the second speaker;
[0009] FIG. 2c is an FFT image of a sample of a third speaker's voice to be used in connection with a profile of the third speaker;
[0010] FIG. 2d is an FFT image of a sample of a fourth speaker's voice to be used in connection with a profile of the fourth speaker;
[0011] FIG. 3 is an FFT image of the captured audio of a conversation where multiple speakers spoke;
[0012] FIG. 4 is a schematic view of a voice recognition system cross correlating the
FFT image of FIG. 2d against the FFT image of FIG. 3;
[0013] FIG. 5a is a schematic view of a display screen showing a first person's identifying information; and [0014] FIG. 5b is a schematic view of a display screen showing a second person's identifying information.
DETAILED DESCRIPTION
[0015] The present disclosure generally relates to a voice-recognition system and method where a person speaking through an audio communication medium is identified to the receiving party by populating a profile unique to the speaker on a display screen on the receiving party's end. The system does so by sampling a speaker's voice while he or she is speaking and, through spectral analysis and the use of sliding Fast Fourier Transforms ("FFTs"), cross-correlates each sample to a number of voice patterns stored in the system, where each voice pattern is associated with a profile. When the cross-correlation of the speaker's voice against the stored voice patterns results in a match with one of the stored voice patterns, the system sends a signal to an electronic device with a display screen on the listening party's end, the signal indicating that the profile associated with the matched voice pattern should appear on the display screen. This profile will identify who the speaker is to the listening party so as to remove any confusion on the listening party's part as to who may be speaking at that given time.
[0016] It should be understood that the elements shown in the figures can be
implemented in various forms of hardware, software or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input/output interfaces. Other elements can be implemented through the use of specifically- purposed devices, such as microphones, audio speakers, and electronic display screens.
[0017] Turning now to FIG. 1, a high-level flow chart of voice sample collection method for teleconferencing 100 in accordance with an embodiment of the present invention is shown. As is typical during a teleconference call, this embodiment involves at least two speaking persons. In one embodiment, multiple persons participate in the call by using their own telephone receivers to call in to a teleconferencing service. In another embodiment, one set of multiple persons are using one telephone receiver to speak to a second set of multiple persons on a second telephone receiver.
[0018] At the beginning of the conference call 102, each person on the call provides a sample of his or her voice to the system (steps 104(a)-(n)) before connecting to the call. In one embodiment, this is done at a stage where each person is asked to state his or her name before being connected. Upon receiving a sample of each person's voice, the system performs a spectral analysis of the voice sample via a Fast Fourier Transform ("FFT") to create a spectral pattern showing the dominant frequencies in the person's voice (steps 106(a)-(n)). Each spectral pattern is then stored in a memory as a signature pattern for each individual speaker, respectively (steps 108(a)-(n)). Each signature pattern is then associated with an individual unique speaker profile (steps 110(a)-(n)). The signature patterns and the speaker profiles are then stored in the system's memory (step 112) and the speakers on the call are then connected to each other to conduct the conference call (step 114).
[0019] FIGs. 2a-d show four examples of speaker profiles 200a-d that were collected from four individual speakers A-D, respectively, using the embodiment illustrated in FIG. 1. As can be seen, each of speaker profiles 200a-d includes a signature pattern (e.g., one of signature patterns 202a-d, respectively) which illustrates three dominant vocal frequencies that
characterize the voice of the individual speaker who gave the voice sample. For example, signature pattern 202a for individual speaker A shows three dominant frequencies of 1 kHz, 500 Hz, and 250 Hz that characterize the voice of individual speaker A, while signature pattern 202b for individual speaker B shows three dominant frequencies of 800 Hz, 600 Hz, and 200 Hz that characterize the voice of individual speaker B. These signature patterns 202a-d are used as bases upon which voice identification can take place during a conference call between individual speakers A-D. This voice identification method is discussed further below.
[0020] Still referring to FIGs. 2a-d, in addition to the signature patterns 202a-d, each of the speaker profiles 200a-d also includes speaker identification information (not shown) particular to the speaker associated with the speaker profile. This speaker identification information is used to inform others on the conference call who is speaking once the voice identification method has identified a particular speaker as speaking at a given time during the call. In one embodiment, the identification information includes the speaker's name. In other embodiments, the identification information can include, but is not limited to, the speaker's name and a picture of the speaker's face and the like.
[0021] During a conference call, the system takes samples of the conversation in real time and converts each sample as it is taken into an FFT image and/or pattern highlighting the dominant frequencies in the sample. FIG. 3 illustrates an FFT profile 300 of a captured conversation during such a conference call. The FFT profile 300 exhibits where windowed filters captured samples of the conversation at a-q instances in time. The a-q instances represent the intervals of time at which a voice sample was captured during the conversation. In one embodiment, the intervals are 0.5 seconds long. In another embodiment, the intervals are 0.25 seconds long. At each of the a-q instances in time, a spectral pattern is given representing the dominant frequencies, if any at all, that were present during the conversation at that given instance. In the example shown in FIG. 3, the spectral graph 300 shows nine instances where a voice sample was taken and converted via FFT into a voice pattern (e.g., voice patterns 302- 318).
[0022] FIG. 4 illustrates how speakers are identified from voice samples taken during the conversation. During the sampled conversation, each of the signature patterns 202a-d found in the speaker profiles 200a-d is slid across the FFT profile of the conversation to determine whether the person associated with the selected signature pattern is speaking. At each instance of time a-q, the system then cross correlates each of the signature patterns 202a-d against the converted voice sample captured at that instance of time. Each cross correlation produces a correlation coefficient R having a value between 0 and 1 , where 1 indicates a perfect match and a 0 indicates no match. When the correlation efficient R produces a value above a predetermined threshold, the system sends a signal indicating that the individual speaker associated with the selected signature pattern is currently speaking.
[0023] In the example shown in FIG. 4, signature pattern 202d from speaker profile 200d is slid across the conversation's FFT profile 300 and cross correlated with the captured sample at each of the a-q instances in time during the conversation. At instances a and b, the cross correlation coefficient R will be 0, since no one was speaking during those instances of time. However, at instance c, the cross correlation of voice pattern 302 with signature pattern 202d results in an R value of 1 , indicating that speaker D is speaking at that instance in time. This is also the case for instances o and q, where correlations with voice patterns 316 and 318 also result in an R value of 1. At instances of time where a voice sample produces a voice pattern that is not the same as that of speaker D (such as voice patterns 304, 306, 308, 310, 312, and 314), the cross correlation with signature pattern 202d reveals an R value between 0 and 1. In such instances, a predetermined threshold for R is set high enough such that other voice patterns correlated with signature pattern 202d do not cause the system to signal false positive identifications. In one embodiment, a predetermined threshold for R is 0.8.
[0024] Once an R value associated with a cross correlation of a particular signature pattern reaches or exceeds a predetermined threshold, the systems sends a signal to a display screen to display the identification information from the speaker profile associated with the particular signature pattern. The display screen then displays the identification information to inform those listening to the call who is speaking at that instance in time. FIGS. 5a and 5b illustrate examples of a display screen 500 displaying identification information 502a, 502b upon receiving a signal from the system to do so. To use the examples shown in FIGS. 2a-d, 3, and 4, the display screen 500 displays the identification information 502a from speaker profile 200a when it receives a signal indicating a match with signature pattern 202a and displays the identification information 502b from speaker profile 200b when it receives a signal indicating a match with signature pattern 202b. This allows those on the call within view of the display screen 500 to see that speakers A and B are speaking on the call when their respective identification information 502a, 502b appears on the display screen 500.
[0025] The system can also send identification information to multiple devices such as mobile devices including, but not limited to, laptops, tablets and/or mobile phones and the like. This can be accomplished via wired (e.g., Ethernet cables, etc.) and/or wireless means (e.g., Bluetooth, WiFi, etc.). This allows easier access to the identification information. It also permits each user to customize how the information is displayed on their device. For example, when a senior person with high authority speaks, the information can be shown in red or bordered by a red border to gain the user's attention more quickly.
[0026] Many variations can be made to the systems and methods discussed above. For example, in one embodiment, as the system performs cross-correlations of the voice patterns in the conversation against the stored signature patterns, the system can also take the voice patterns which have been identified as a match to one of the signature patterns and average them together with the identified signature pattern as the conversation progresses. This allows the system to learn the characteristics of each person's voice as the conversation progresses and can potentially lead to more accurate identifications of the various speakers during the call.
[0027] The various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof. Moreover, the software is preferably implemented as an application program tangibly embodied on a program storage unit or computer readable medium. The application program may be uploaded to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units ("CPUs"), a memory, and input/output interfaces. The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be either part of the microinstruction code or part of the application program, or any combination thereof, which may be executed by a CPU, whether or not such computer or processor is explicitly shown. In addition, various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit.
[0028] All examples and conditional language recited herein are intended for pedagogical purposes to aid the reader in understanding the principles of the invention and the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof.
Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0029] It will be understood that the embodiments described herein are merely exemplary and that a person skilled in the art may make many variations and modifications without departing from the spirit and scope of the invention. All such variations and modifications are intended to be included within the scope of the invention as defined in the appended claims.

Claims

CLAIMS:
1. A voice recognition method for a communication system, the method comprising: obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern created from a spectral analysis of a first voice sample of the speaker's voice;
receiving a second voice sample;
converting the second voice sample into a voice pattern using spectral analysis;
cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 < R < 1 ; and
sending the identifying information associated with the speaker profile to at least one device when the cross correlation coefficient R is above a threshold.
2. The voice recognition method according to claim 1, wherein the identifying information includes the speaker's name.
3. The voice recognition method according to claim 1, wherein the identifying information includes an image associated with the speaker.
4. The voice recognition method according to claim 1 , wherein the method is repeated X times over X instances of time, and wherein the second voice sample is a new voice sample at each of the X instances of time, wherein X is an integer.
5. The voice recognition method according to claim 1, wherein obtaining the speaker profile includes receiving the first voice sample from the speaker, converting the first voice sample into the signature pattern using FFT, and associating the identifying information with the signature pattern to create the speaker profile.
6. The voice recognition method according to claim 1, wherein a threshold is 0.8.
7. A system for identifying a person based on voice recognition, the system comprising:
a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern created from a spectral analysis of a first voice sample of the speaker's voice; a microphone for receiving a second voice sample; and
a processor in communication with the microphone, the memory, and the display device, the processor configured to perform the following process:
convert the second voice sample into a voice pattern using spectral analysis; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 < R < 1; and
send the identifying information associated with the speaker profile to at least one device when the cross correlation coefficient R is above a threshold,.
8. The system according to claim 7, wherein the identifying information includes the speaker's name.
9. The system according to claim 7, wherein the identifying information includes an image associated with the speaker.
10. The system according to claim 7, wherein the processor is further configured to repeat the process times over X instances of time, and wherein the second voice sample is a new voice sample at each of the X instances of time, wherein X is an integer.
11. The system according to claim 7, wherein a threshold is 0.8.
PCT/US2013/044413 2013-06-06 2013-06-06 Voice recognition and identification Ceased WO2014196971A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/US2013/044413 WO2014196971A1 (en) 2013-06-06 2013-06-06 Voice recognition and identification

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2013/044413 WO2014196971A1 (en) 2013-06-06 2013-06-06 Voice recognition and identification

Publications (1)

Publication Number Publication Date
WO2014196971A1 true WO2014196971A1 (en) 2014-12-11

Family

ID=48670830

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2013/044413 Ceased WO2014196971A1 (en) 2013-06-06 2013-06-06 Voice recognition and identification

Country Status (1)

Country Link
WO (1) WO2014196971A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6718306B1 (en) * 1999-10-21 2004-04-06 Casio Computer Co., Ltd. Speech collating apparatus and speech collating method
US6931113B2 (en) * 2002-11-08 2005-08-16 Verizon Services Corp. Facilitation of a conference call

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6718306B1 (en) * 1999-10-21 2004-04-06 Casio Computer Co., Ltd. Speech collating apparatus and speech collating method
US6931113B2 (en) * 2002-11-08 2005-08-16 Verizon Services Corp. Facilitation of a conference call

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
DHANANJAYA N ET AL: "Correlation-Based Similarity Between Signals for Speaker Verification with Limited Amount of Speech Data", 1 January 2006, MULTIMEDIA CONTENT REPRESENTATION, CLASSIFICATION AND SECURITY LECTURE NOTES IN COMPUTER SCIENCE;;LNCS, SPRINGER, BERLIN, DE, PAGE(S) 17 - 25, ISBN: 978-3-540-39392-4, XP019040403 *
RAKESH D R ET AL: "SPEAKER RECOGNITION AND AUTHENTICATION", INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MOBILE COMPUTING, 30 May 2013 (2013-05-30), pages 402 - 407, XP055114301, Retrieved from the Internet <URL:http://ijcsmc.com/docs/papers/May2013/V2I52013136.pdf> *
ZALEWSKI J ET AL: "Cross Correlation of Long-Term Speech Spectra as a Speaker Identification Technique", vol. 34, no. 1, 1 November 1975 (1975-11-01), pages 20 - 24, XP008168845, ISSN: 1610-1928, Retrieved from the Internet <URL:http://www.ingentaconnect.com/content/dav/aaua/1975/00000034/00000001/art00005> *

Similar Documents

Publication Publication Date Title
US10573318B2 (en) Voice information control method and terminal device
CN105979197B (en) Teleconference control method and device based on sound automatic identification of uttering long and high-pitched sounds
US20240007344A1 (en) Network device maintenance
US10645214B1 (en) Identical conversation detection method and apparatus
CN107749313B (en) A kind of method of automatic transcription and generation Telemedicine Consultation record
KR101528086B1 (en) System and method for providing conference information
US7672844B2 (en) Voice processing apparatus
WO2021118686A1 (en) Communication of transcriptions
WO2011090411A1 (en) Meeting room participant recogniser
EP2973559B1 (en) Audio transmission channel quality assessment
WO2016173132A1 (en) Method and device for voice recognition, and user equipment
WO2016198132A1 (en) Communication system, audio server, and method for operating a communication system
US9843683B2 (en) Configuration method for sound collection system for meeting using terminals and server apparatus
CN113971956A (en) Information processing method and device, electronic equipment and readable storage medium
JP2014149571A (en) Content search device
JP2019153099A (en) Conference assisting system, and conference assisting program
CN111210810A (en) Model training method and device
CN111199751A (en) Microphone shielding method, device and electronic device
CN116153328A (en) Audio data processing method, system, storage medium and electronic equipment
CN107197404B (en) A sound effect automatic adjustment method, device and a recording and broadcasting system
JP2007241130A (en) Systems and devices that use voiceprint recognition
CN110556114B (en) Caller identification method and device based on attention mechanism
CN111951809B (en) Multi-person voiceprint identification method and system
JP2013197906A (en) Administrator automatic notification system of call support status
CN103929532A (en) Information processing method and electronic equipment

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13730749

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13730749

Country of ref document: EP

Kind code of ref document: A1