WO2014196971A1 - Voice recognition and identification - Google Patents
Voice recognition and identification Download PDFInfo
- Publication number
- WO2014196971A1 WO2014196971A1 PCT/US2013/044413 US2013044413W WO2014196971A1 WO 2014196971 A1 WO2014196971 A1 WO 2014196971A1 US 2013044413 W US2013044413 W US 2013044413W WO 2014196971 A1 WO2014196971 A1 WO 2014196971A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- speaker
- pattern
- sample
- identifying information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
Definitions
- a voice recognition method for identifying a person on a display screen based on when the person's voice is recognized is disclosed.
- the method can be implemented in a number of systems where such identification is desirable, such as
- the method comprises obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier
- FFT Transform
- the system comprises a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; a microphone for receiving a second voice sample; a display screen; and a processor in communication with the microphone, the memory, and the display screen.
- FFT Fast Fourier Transform
- the processor is configured to perform the following process: convert the second voice sample into a voice pattern using FFT; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 ⁇ R ⁇ 1 ; determine whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, instruct the display screen to display the identifying information associated with the speaker profile.
- FIG. 1 is a high-level flow chart showing a voice sample collection method accordance with an embodiment of the present invention
- FIG. 2a is an FFT image of a sample of a first speaker's voice to be used in connection with a profile of the first speaker;
- FIG. 2b is an FFT image of a sample of a second speaker's voice to be used in connection with a profile of the second speaker;
- FIG. 2c is an FFT image of a sample of a third speaker's voice to be used in connection with a profile of the third speaker;
- FIG. 2d is an FFT image of a sample of a fourth speaker's voice to be used in connection with a profile of the fourth speaker;
- FIG. 3 is an FFT image of the captured audio of a conversation where multiple speakers spoke
- FIG. 4 is a schematic view of a voice recognition system cross correlating the
- FIG. 5a is a schematic view of a display screen showing a first person's identifying information
- FIG. 5b is a schematic view of a display screen showing a second person's identifying information.
- the present disclosure generally relates to a voice-recognition system and method where a person speaking through an audio communication medium is identified to the receiving party by populating a profile unique to the speaker on a display screen on the receiving party's end.
- the system does so by sampling a speaker's voice while he or she is speaking and, through spectral analysis and the use of sliding Fast Fourier Transforms ("FFTs"), cross-correlates each sample to a number of voice patterns stored in the system, where each voice pattern is associated with a profile.
- FFTs sliding Fast Fourier Transforms
- the system sends a signal to an electronic device with a display screen on the listening party's end, the signal indicating that the profile associated with the matched voice pattern should appear on the display screen.
- This profile will identify who the speaker is to the listening party so as to remove any confusion on the listening party's part as to who may be speaking at that given time.
- these elements are implemented in various forms of hardware, software or combinations thereof.
- these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input/output interfaces.
- general-purpose devices which may include a processor, memory and input/output interfaces.
- Other elements can be implemented through the use of specifically- purposed devices, such as microphones, audio speakers, and electronic display screens.
- FIG. 1 a high-level flow chart of voice sample collection method for teleconferencing 100 in accordance with an embodiment of the present invention is shown.
- this embodiment involves at least two speaking persons.
- multiple persons participate in the call by using their own telephone receivers to call in to a teleconferencing service.
- one set of multiple persons are using one telephone receiver to speak to a second set of multiple persons on a second telephone receiver.
- each person on the call provides a sample of his or her voice to the system (steps 104(a)-(n)) before connecting to the call. In one embodiment, this is done at a stage where each person is asked to state his or her name before being connected.
- the system Upon receiving a sample of each person's voice, the system performs a spectral analysis of the voice sample via a Fast Fourier Transform ("FFT") to create a spectral pattern showing the dominant frequencies in the person's voice (steps 106(a)-(n)).
- FFT Fast Fourier Transform
- Each spectral pattern is then stored in a memory as a signature pattern for each individual speaker, respectively (steps 108(a)-(n)).
- Each signature pattern is then associated with an individual unique speaker profile (steps 110(a)-(n)).
- the signature patterns and the speaker profiles are then stored in the system's memory (step 112) and the speakers on the call are then connected to each other to conduct the conference call (step 114).
- FIGs. 2a-d show four examples of speaker profiles 200a-d that were collected from four individual speakers A-D, respectively, using the embodiment illustrated in FIG. 1.
- each of speaker profiles 200a-d includes a signature pattern (e.g., one of signature patterns 202a-d, respectively) which illustrates three dominant vocal frequencies that
- signature pattern 202a for individual speaker A shows three dominant frequencies of 1 kHz, 500 Hz, and 250 Hz that characterize the voice of individual speaker A
- signature pattern 202b for individual speaker B shows three dominant frequencies of 800 Hz, 600 Hz, and 200 Hz that characterize the voice of individual speaker B.
- signature patterns 202a-d are used as bases upon which voice identification can take place during a conference call between individual speakers A-D. This voice identification method is discussed further below.
- each of the speaker profiles 200a-d also includes speaker identification information (not shown) particular to the speaker associated with the speaker profile.
- This speaker identification information is used to inform others on the conference call who is speaking once the voice identification method has identified a particular speaker as speaking at a given time during the call.
- the identification information includes the speaker's name.
- the identification information can include, but is not limited to, the speaker's name and a picture of the speaker's face and the like.
- FIG. 3 illustrates an FFT profile 300 of a captured conversation during such a conference call.
- the FFT profile 300 exhibits where windowed filters captured samples of the conversation at a-q instances in time.
- the a-q instances represent the intervals of time at which a voice sample was captured during the conversation. In one embodiment, the intervals are 0.5 seconds long. In another embodiment, the intervals are 0.25 seconds long.
- a spectral pattern is given representing the dominant frequencies, if any at all, that were present during the conversation at that given instance.
- the spectral graph 300 shows nine instances where a voice sample was taken and converted via FFT into a voice pattern (e.g., voice patterns 302- 318).
- FIG. 4 illustrates how speakers are identified from voice samples taken during the conversation.
- each of the signature patterns 202a-d found in the speaker profiles 200a-d is slid across the FFT profile of the conversation to determine whether the person associated with the selected signature pattern is speaking.
- the system then cross correlates each of the signature patterns 202a-d against the converted voice sample captured at that instance of time.
- Each cross correlation produces a correlation coefficient R having a value between 0 and 1 , where 1 indicates a perfect match and a 0 indicates no match.
- the correlation efficient R produces a value above a predetermined threshold, the system sends a signal indicating that the individual speaker associated with the selected signature pattern is currently speaking.
- signature pattern 202d from speaker profile 200d is slid across the conversation's FFT profile 300 and cross correlated with the captured sample at each of the a-q instances in time during the conversation.
- the cross correlation coefficient R will be 0, since no one was speaking during those instances of time.
- the cross correlation of voice pattern 302 with signature pattern 202d results in an R value of 1 , indicating that speaker D is speaking at that instance in time. This is also the case for instances o and q, where correlations with voice patterns 316 and 318 also result in an R value of 1.
- a voice sample produces a voice pattern that is not the same as that of speaker D (such as voice patterns 304, 306, 308, 310, 312, and 314)
- the cross correlation with signature pattern 202d reveals an R value between 0 and 1.
- a predetermined threshold for R is set high enough such that other voice patterns correlated with signature pattern 202d do not cause the system to signal false positive identifications.
- a predetermined threshold for R is 0.8.
- FIGS. 5a and 5b illustrate examples of a display screen 500 displaying identification information 502a, 502b upon receiving a signal from the system to do so. To use the examples shown in FIGS.
- the display screen 500 displays the identification information 502a from speaker profile 200a when it receives a signal indicating a match with signature pattern 202a and displays the identification information 502b from speaker profile 200b when it receives a signal indicating a match with signature pattern 202b. This allows those on the call within view of the display screen 500 to see that speakers A and B are speaking on the call when their respective identification information 502a, 502b appears on the display screen 500.
- the system can also send identification information to multiple devices such as mobile devices including, but not limited to, laptops, tablets and/or mobile phones and the like. This can be accomplished via wired (e.g., Ethernet cables, etc.) and/or wireless means (e.g., Bluetooth, WiFi, etc.). This allows easier access to the identification information. It also permits each user to customize how the information is displayed on their device. For example, when a senior person with high authority speaks, the information can be shown in red or bordered by a red border to gain the user's attention more quickly.
- the system can also take the voice patterns which have been identified as a match to one of the signature patterns and average them together with the identified signature pattern as the conversation progresses. This allows the system to learn the characteristics of each person's voice as the conversation progresses and can potentially lead to more accurate identifications of the various speakers during the call.
- the various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof.
- the software is preferably implemented as an application program tangibly embodied on a program storage unit or computer readable medium.
- the application program may be uploaded to, and executed by, a machine comprising any suitable architecture.
- the machine is implemented on a computer platform having hardware such as one or more central processing units ("CPUs"), a memory, and input/output interfaces.
- CPUs central processing units
- the computer platform may also include an operating system and microinstruction code.
- the various processes and functions described herein may be either part of the microinstruction code or part of the application program, or any combination thereof, which may be executed by a CPU, whether or not such computer or processor is explicitly shown.
- various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Telephonic Communication Services (AREA)
Abstract
A voice recognition method and system for identifying a person on a display screen based on when the person's voice is recognized is disclosed. The method includes obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; receiving a second voice sample; converting the second voice sample into a voice pattern using FFT; cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 ≤ R ≤ 1; and sending the identifying information associated with the speaker profile when the cross correlation coefficient R is above a threshold.
Description
VOICE RECOGNITION AND IDENTIFICATION
BACKGROUND
[0001] As teleconferencing and video conferencing increase in usage, knowing who is speaking at any given time is important. Many times, the persons on a teleconferencing call may know each other's names but have never met before and are not familiar with each other's voices. In such circumstances, when one person speaks during a teleconferencing call, the other persons on the call may not know who is speaking without the person self-identifying him or herself during the call. However, not every person on a call may remember to self-identify during the call, which may cause confusion for the other persons on the call and cause them to ask who was speaking, which wastes time.
SUMMARY
[0002] In view of the foregoing, a voice recognition method for identifying a person on a display screen based on when the person's voice is recognized is disclosed. The method can be implemented in a number of systems where such identification is desirable, such as
teleconferencing, video conferencing, multiplayer video game platforms, and the like.
[0003] In one embodiment, the method comprises obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier
Transform (FFT) of a first voice sample of the speaker's voice; receiving a second voice sample; converting the second voice sample into a voice pattern using FFT; cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 < R < 1; determining whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, communicating with the display screen to display the identifying information associated with the speaker profile.
[0004] Also disclosed is a system for identifying a person on a display screen based on voice recognition. In one embodiment, the system comprises a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern having been created from a Fast Fourier Transform (FFT) of a first voice sample of the speaker's voice; a microphone
for receiving a second voice sample; a display screen; and a processor in communication with the microphone, the memory, and the display screen. In this embodiment, the processor is configured to perform the following process: convert the second voice sample into a voice pattern using FFT; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 < R < 1 ; determine whether the cross correlation coefficient R is above a predetermined threshold; and when the cross correlation coefficient R is above a predetermined threshold, instruct the display screen to display the identifying information associated with the speaker profile.
BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For a more complete understanding of the present invention, reference is made to the following detailed description of an embodiment considered in conjunction with the accompanying drawings, in which:
[0006] FIG. 1 is a high-level flow chart showing a voice sample collection method accordance with an embodiment of the present invention;
[0007] FIG. 2a is an FFT image of a sample of a first speaker's voice to be used in connection with a profile of the first speaker;
[0008] FIG. 2b is an FFT image of a sample of a second speaker's voice to be used in connection with a profile of the second speaker;
[0009] FIG. 2c is an FFT image of a sample of a third speaker's voice to be used in connection with a profile of the third speaker;
[0010] FIG. 2d is an FFT image of a sample of a fourth speaker's voice to be used in connection with a profile of the fourth speaker;
[0011] FIG. 3 is an FFT image of the captured audio of a conversation where multiple speakers spoke;
[0012] FIG. 4 is a schematic view of a voice recognition system cross correlating the
FFT image of FIG. 2d against the FFT image of FIG. 3;
[0013] FIG. 5a is a schematic view of a display screen showing a first person's identifying information; and
[0014] FIG. 5b is a schematic view of a display screen showing a second person's identifying information.
DETAILED DESCRIPTION
[0015] The present disclosure generally relates to a voice-recognition system and method where a person speaking through an audio communication medium is identified to the receiving party by populating a profile unique to the speaker on a display screen on the receiving party's end. The system does so by sampling a speaker's voice while he or she is speaking and, through spectral analysis and the use of sliding Fast Fourier Transforms ("FFTs"), cross-correlates each sample to a number of voice patterns stored in the system, where each voice pattern is associated with a profile. When the cross-correlation of the speaker's voice against the stored voice patterns results in a match with one of the stored voice patterns, the system sends a signal to an electronic device with a display screen on the listening party's end, the signal indicating that the profile associated with the matched voice pattern should appear on the display screen. This profile will identify who the speaker is to the listening party so as to remove any confusion on the listening party's part as to who may be speaking at that given time.
[0016] It should be understood that the elements shown in the figures can be
implemented in various forms of hardware, software or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input/output interfaces. Other elements can be implemented through the use of specifically- purposed devices, such as microphones, audio speakers, and electronic display screens.
[0017] Turning now to FIG. 1, a high-level flow chart of voice sample collection method for teleconferencing 100 in accordance with an embodiment of the present invention is shown. As is typical during a teleconference call, this embodiment involves at least two speaking persons. In one embodiment, multiple persons participate in the call by using their own telephone receivers to call in to a teleconferencing service. In another embodiment, one set of multiple persons are using one telephone receiver to speak to a second set of multiple persons on a second telephone receiver.
[0018] At the beginning of the conference call 102, each person on the call provides a sample of his or her voice to the system (steps 104(a)-(n)) before connecting to the call. In one
embodiment, this is done at a stage where each person is asked to state his or her name before being connected. Upon receiving a sample of each person's voice, the system performs a spectral analysis of the voice sample via a Fast Fourier Transform ("FFT") to create a spectral pattern showing the dominant frequencies in the person's voice (steps 106(a)-(n)). Each spectral pattern is then stored in a memory as a signature pattern for each individual speaker, respectively (steps 108(a)-(n)). Each signature pattern is then associated with an individual unique speaker profile (steps 110(a)-(n)). The signature patterns and the speaker profiles are then stored in the system's memory (step 112) and the speakers on the call are then connected to each other to conduct the conference call (step 114).
[0019] FIGs. 2a-d show four examples of speaker profiles 200a-d that were collected from four individual speakers A-D, respectively, using the embodiment illustrated in FIG. 1. As can be seen, each of speaker profiles 200a-d includes a signature pattern (e.g., one of signature patterns 202a-d, respectively) which illustrates three dominant vocal frequencies that
characterize the voice of the individual speaker who gave the voice sample. For example, signature pattern 202a for individual speaker A shows three dominant frequencies of 1 kHz, 500 Hz, and 250 Hz that characterize the voice of individual speaker A, while signature pattern 202b for individual speaker B shows three dominant frequencies of 800 Hz, 600 Hz, and 200 Hz that characterize the voice of individual speaker B. These signature patterns 202a-d are used as bases upon which voice identification can take place during a conference call between individual speakers A-D. This voice identification method is discussed further below.
[0020] Still referring to FIGs. 2a-d, in addition to the signature patterns 202a-d, each of the speaker profiles 200a-d also includes speaker identification information (not shown) particular to the speaker associated with the speaker profile. This speaker identification information is used to inform others on the conference call who is speaking once the voice identification method has identified a particular speaker as speaking at a given time during the call. In one embodiment, the identification information includes the speaker's name. In other embodiments, the identification information can include, but is not limited to, the speaker's name and a picture of the speaker's face and the like.
[0021] During a conference call, the system takes samples of the conversation in real time and converts each sample as it is taken into an FFT image and/or pattern highlighting the dominant frequencies in the sample. FIG. 3 illustrates an FFT profile 300 of a captured
conversation during such a conference call. The FFT profile 300 exhibits where windowed filters captured samples of the conversation at a-q instances in time. The a-q instances represent the intervals of time at which a voice sample was captured during the conversation. In one embodiment, the intervals are 0.5 seconds long. In another embodiment, the intervals are 0.25 seconds long. At each of the a-q instances in time, a spectral pattern is given representing the dominant frequencies, if any at all, that were present during the conversation at that given instance. In the example shown in FIG. 3, the spectral graph 300 shows nine instances where a voice sample was taken and converted via FFT into a voice pattern (e.g., voice patterns 302- 318).
[0022] FIG. 4 illustrates how speakers are identified from voice samples taken during the conversation. During the sampled conversation, each of the signature patterns 202a-d found in the speaker profiles 200a-d is slid across the FFT profile of the conversation to determine whether the person associated with the selected signature pattern is speaking. At each instance of time a-q, the system then cross correlates each of the signature patterns 202a-d against the converted voice sample captured at that instance of time. Each cross correlation produces a correlation coefficient R having a value between 0 and 1 , where 1 indicates a perfect match and a 0 indicates no match. When the correlation efficient R produces a value above a predetermined threshold, the system sends a signal indicating that the individual speaker associated with the selected signature pattern is currently speaking.
[0023] In the example shown in FIG. 4, signature pattern 202d from speaker profile 200d is slid across the conversation's FFT profile 300 and cross correlated with the captured sample at each of the a-q instances in time during the conversation. At instances a and b, the cross correlation coefficient R will be 0, since no one was speaking during those instances of time. However, at instance c, the cross correlation of voice pattern 302 with signature pattern 202d results in an R value of 1 , indicating that speaker D is speaking at that instance in time. This is also the case for instances o and q, where correlations with voice patterns 316 and 318 also result in an R value of 1. At instances of time where a voice sample produces a voice pattern that is not the same as that of speaker D (such as voice patterns 304, 306, 308, 310, 312, and 314), the cross correlation with signature pattern 202d reveals an R value between 0 and 1. In such instances, a predetermined threshold for R is set high enough such that other voice patterns correlated with
signature pattern 202d do not cause the system to signal false positive identifications. In one embodiment, a predetermined threshold for R is 0.8.
[0024] Once an R value associated with a cross correlation of a particular signature pattern reaches or exceeds a predetermined threshold, the systems sends a signal to a display screen to display the identification information from the speaker profile associated with the particular signature pattern. The display screen then displays the identification information to inform those listening to the call who is speaking at that instance in time. FIGS. 5a and 5b illustrate examples of a display screen 500 displaying identification information 502a, 502b upon receiving a signal from the system to do so. To use the examples shown in FIGS. 2a-d, 3, and 4, the display screen 500 displays the identification information 502a from speaker profile 200a when it receives a signal indicating a match with signature pattern 202a and displays the identification information 502b from speaker profile 200b when it receives a signal indicating a match with signature pattern 202b. This allows those on the call within view of the display screen 500 to see that speakers A and B are speaking on the call when their respective identification information 502a, 502b appears on the display screen 500.
[0025] The system can also send identification information to multiple devices such as mobile devices including, but not limited to, laptops, tablets and/or mobile phones and the like. This can be accomplished via wired (e.g., Ethernet cables, etc.) and/or wireless means (e.g., Bluetooth, WiFi, etc.). This allows easier access to the identification information. It also permits each user to customize how the information is displayed on their device. For example, when a senior person with high authority speaks, the information can be shown in red or bordered by a red border to gain the user's attention more quickly.
[0026] Many variations can be made to the systems and methods discussed above. For example, in one embodiment, as the system performs cross-correlations of the voice patterns in the conversation against the stored signature patterns, the system can also take the voice patterns which have been identified as a match to one of the signature patterns and average them together with the identified signature pattern as the conversation progresses. This allows the system to learn the characteristics of each person's voice as the conversation progresses and can potentially lead to more accurate identifications of the various speakers during the call.
[0027] The various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof. Moreover, the software is preferably
implemented as an application program tangibly embodied on a program storage unit or computer readable medium. The application program may be uploaded to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units ("CPUs"), a memory, and input/output interfaces. The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be either part of the microinstruction code or part of the application program, or any combination thereof, which may be executed by a CPU, whether or not such computer or processor is explicitly shown. In addition, various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit.
[0028] All examples and conditional language recited herein are intended for pedagogical purposes to aid the reader in understanding the principles of the invention and the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof.
Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0029] It will be understood that the embodiments described herein are merely exemplary and that a person skilled in the art may make many variations and modifications without departing from the spirit and scope of the invention. All such variations and modifications are intended to be included within the scope of the invention as defined in the appended claims.
Claims
1. A voice recognition method for a communication system, the method comprising: obtaining a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern created from a spectral analysis of a first voice sample of the speaker's voice;
receiving a second voice sample;
converting the second voice sample into a voice pattern using spectral analysis;
cross correlating the voice pattern with the signature pattern, resulting in a cross correlation coefficient R, where 0 < R < 1 ; and
sending the identifying information associated with the speaker profile to at least one device when the cross correlation coefficient R is above a threshold.
2. The voice recognition method according to claim 1, wherein the identifying information includes the speaker's name.
3. The voice recognition method according to claim 1, wherein the identifying information includes an image associated with the speaker.
4. The voice recognition method according to claim 1 , wherein the method is repeated X times over X instances of time, and wherein the second voice sample is a new voice sample at each of the X instances of time, wherein X is an integer.
5. The voice recognition method according to claim 1, wherein obtaining the speaker profile includes receiving the first voice sample from the speaker, converting the first voice sample into the signature pattern using FFT, and associating the identifying information with the signature pattern to create the speaker profile.
6. The voice recognition method according to claim 1, wherein a threshold is 0.8.
7. A system for identifying a person based on voice recognition, the system comprising:
a memory configured to store a speaker profile associated with a speaker, the speaker profile including a signature pattern and identifying information associated with the speaker, the signature pattern created from a spectral analysis of a first voice sample of the speaker's voice; a microphone for receiving a second voice sample; and
a processor in communication with the microphone, the memory, and the display device, the processor configured to perform the following process:
convert the second voice sample into a voice pattern using spectral analysis; cross correlate the voice pattern with the signature pattern from the memory; calculate a cross correlation coefficient R associated with the cross correlation of the voice pattern with the signature pattern, where 0 < R < 1; and
send the identifying information associated with the speaker profile to at least one device when the cross correlation coefficient R is above a threshold,.
8. The system according to claim 7, wherein the identifying information includes the speaker's name.
9. The system according to claim 7, wherein the identifying information includes an image associated with the speaker.
10. The system according to claim 7, wherein the processor is further configured to repeat the process times over X instances of time, and wherein the second voice sample is a new voice sample at each of the X instances of time, wherein X is an integer.
11. The system according to claim 7, wherein a threshold is 0.8.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/US2013/044413 WO2014196971A1 (en) | 2013-06-06 | 2013-06-06 | Voice recognition and identification |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/US2013/044413 WO2014196971A1 (en) | 2013-06-06 | 2013-06-06 | Voice recognition and identification |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2014196971A1 true WO2014196971A1 (en) | 2014-12-11 |
Family
ID=48670830
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2013/044413 Ceased WO2014196971A1 (en) | 2013-06-06 | 2013-06-06 | Voice recognition and identification |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2014196971A1 (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6718306B1 (en) * | 1999-10-21 | 2004-04-06 | Casio Computer Co., Ltd. | Speech collating apparatus and speech collating method |
| US6931113B2 (en) * | 2002-11-08 | 2005-08-16 | Verizon Services Corp. | Facilitation of a conference call |
-
2013
- 2013-06-06 WO PCT/US2013/044413 patent/WO2014196971A1/en not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6718306B1 (en) * | 1999-10-21 | 2004-04-06 | Casio Computer Co., Ltd. | Speech collating apparatus and speech collating method |
| US6931113B2 (en) * | 2002-11-08 | 2005-08-16 | Verizon Services Corp. | Facilitation of a conference call |
Non-Patent Citations (3)
| Title |
|---|
| DHANANJAYA N ET AL: "Correlation-Based Similarity Between Signals for Speaker Verification with Limited Amount of Speech Data", 1 January 2006, MULTIMEDIA CONTENT REPRESENTATION, CLASSIFICATION AND SECURITY LECTURE NOTES IN COMPUTER SCIENCE;;LNCS, SPRINGER, BERLIN, DE, PAGE(S) 17 - 25, ISBN: 978-3-540-39392-4, XP019040403 * |
| RAKESH D R ET AL: "SPEAKER RECOGNITION AND AUTHENTICATION", INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MOBILE COMPUTING, 30 May 2013 (2013-05-30), pages 402 - 407, XP055114301, Retrieved from the Internet <URL:http://ijcsmc.com/docs/papers/May2013/V2I52013136.pdf> * |
| ZALEWSKI J ET AL: "Cross Correlation of Long-Term Speech Spectra as a Speaker Identification Technique", vol. 34, no. 1, 1 November 1975 (1975-11-01), pages 20 - 24, XP008168845, ISSN: 1610-1928, Retrieved from the Internet <URL:http://www.ingentaconnect.com/content/dav/aaua/1975/00000034/00000001/art00005> * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10573318B2 (en) | Voice information control method and terminal device | |
| CN105979197B (en) | Teleconference control method and device based on sound automatic identification of uttering long and high-pitched sounds | |
| US20240007344A1 (en) | Network device maintenance | |
| US10645214B1 (en) | Identical conversation detection method and apparatus | |
| CN107749313B (en) | A kind of method of automatic transcription and generation Telemedicine Consultation record | |
| KR101528086B1 (en) | System and method for providing conference information | |
| US7672844B2 (en) | Voice processing apparatus | |
| WO2021118686A1 (en) | Communication of transcriptions | |
| WO2011090411A1 (en) | Meeting room participant recogniser | |
| EP2973559B1 (en) | Audio transmission channel quality assessment | |
| WO2016173132A1 (en) | Method and device for voice recognition, and user equipment | |
| WO2016198132A1 (en) | Communication system, audio server, and method for operating a communication system | |
| US9843683B2 (en) | Configuration method for sound collection system for meeting using terminals and server apparatus | |
| CN113971956A (en) | Information processing method and device, electronic equipment and readable storage medium | |
| JP2014149571A (en) | Content search device | |
| JP2019153099A (en) | Conference assisting system, and conference assisting program | |
| CN111210810A (en) | Model training method and device | |
| CN111199751A (en) | Microphone shielding method, device and electronic device | |
| CN116153328A (en) | Audio data processing method, system, storage medium and electronic equipment | |
| CN107197404B (en) | A sound effect automatic adjustment method, device and a recording and broadcasting system | |
| JP2007241130A (en) | Systems and devices that use voiceprint recognition | |
| CN110556114B (en) | Caller identification method and device based on attention mechanism | |
| CN111951809B (en) | Multi-person voiceprint identification method and system | |
| JP2013197906A (en) | Administrator automatic notification system of call support status | |
| CN103929532A (en) | Information processing method and electronic equipment |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 13730749 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 13730749 Country of ref document: EP Kind code of ref document: A1 |