US8036901B2 - Systems and methods of performing speech recognition using sensory inputs of human position - Google Patents

Systems and methods of performing speech recognition using sensory inputs of human position Download PDF

Info

Publication number
US8036901B2
US8036901B2 US11973140 US97314007A US8036901B2 US 8036901 B2 US8036901 B2 US 8036901B2 US 11973140 US11973140 US 11973140 US 97314007 A US97314007 A US 97314007A US 8036901 B2 US8036901 B2 US 8036901B2
Authority
US
Grant status
Grant
Patent type
Prior art keywords
recognition
set
speech
position
human
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active, expires
Application number
US11973140
Other versions
US20090094032A1 (en )
Inventor
Todd F. Mozer
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sensory Inc
Original Assignee
Sensory Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Grant date

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS OR SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS OR SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/226Taking into account non-speech caracteristics
    • G10L2015/227Taking into account non-speech caracteristics of the speaker; Human-factor methodology

Abstract

Embodiments of the present invention improve methods of performing speech recognition using sensory inputs of human position. In one embodiment, the present invention includes a speech recognition method comprising sensing a change in position of at least one part of a human body, selecting a recognition set based on the change of position, receiving a speech input signal, and recognizing the speech input signal in the context of the first recognition set.

Description

BACKGROUND

The present invention relates to speech recognition, and more particularly, to systems and methods of performing speech recognition using sensory inputs of human position.

Electronic devices have become more readily available to the public and people find themselves interfacing with many different electronic devices during their daily lives. Historically, the adoption of an electronic device required the user to spend considerable time learning to interface with the device. The advent of menu driven interfaces helped to alleviate some of the tedium of learning to interface with an electronic device, but this method of interfacing with an electronic device still required a person to learn where the menus were and how to use them. More recently, the tactile and motion interfaces have attempted to make the experience of using an electronic device more intuitive. Although the advancements in tactile and motion devices have improved the experience of interfacing with an electronic device, the user is still constrained by the use of visual cues to maneuver through the options and functions of the electronic device. Due to this limitation, the user may still be required to spend a great deal of time learning to interface with the electronic device. Speech recognition would help improve the interface immensely by allowing the user to tell the device what task was desired. Historically however, effective speech recognition requires large amounts of memory and uses a considerable time to “recognize” a given utterance. In this way, the historical speech recognition methods may simply add to the frustration of interfacing with an electronic device rather than facilitate its use. These factors, as well as many others, have prevented the use of speech recognition in electronic devices.

The present invention solves these and other problems with systems and methods of performing speech recognition using sensory inputs of human position.

SUMMARY

Embodiments of the present invention improve methods of performing speech recognition using sensory inputs of human position. In one embodiment, the present invention includes a speech recognition method comprising sensing a change in position of at least one part of a human body, selecting a recognition set based on the change of position, receiving a speech input signal, and recognizing the speech input signal in the context of the first recognition set.

In one embodiment, the change of position includes a change in orientation.

In one embodiment, the change of position includes a change in direction.

In one embodiment, the change of position includes a portion of a human hand proximate with a surface.

In one embodiment, the change of position includes motion.

In one embodiment, the speech input signal is a portion of an utterance.

In one embodiment, the recognizing includes choosing an element from the first recognition set.

In one embodiment, the recognizing includes using the first recognition set to weight a set of likelihoods, wherein the first recognition set includes segments of speech.

In one embodiment, the method further comprises initiating a state of a computer program, and selecting a state recognition set based on the state of the computer program, wherein the selecting of the first recognition set includes finding a subset of the state recognition set.

In one embodiment, the method further comprises changing the state of the computer program according to the recognition result.

In one embodiment, the speech input signal includes digital data.

In one embodiment, the sensing includes the use of a tactile sensor.

In one embodiment, the sensing includes the use of a gyroscope.

BRIEF DESCRIPTION OF THE DRAWINGS

FIG. 1 illustrates a system for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention.

FIG. 2 illustrates a method for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention.

FIG. 3 illustrates another method for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention.

FIG. 4 illustrates another method for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention.

FIGS. 5A and 5B illustrates an example of how the position of a human hand may be used for performing speech recognition according to one embodiment of the present invention.

FIGS. 6A and 6B illustrates an example of how the position of a human head may be used for performing speech recognition according to another embodiment of the present invention.

FIG. 7 illustrates another example of how the position of parts of a human body on a tactile screen may be used for performing speech recognition according to another embodiment of the present invention.

DETAILED DESCRIPTION

Described herein are techniques for a content selection systems and methods using speech recognition. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present invention. It will be evident, however, to one skilled in the art that the present invention as defined by the claims may include some or all of the features in these examples alone or in combination with other features described below, and may further include obvious modifications and equivalents of the features and concepts described herein.

FIG. 1 illustrates a system for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention. System 100 includes an output device 101, a sensor A 102, a sensor B 103, a controller 104, a repository of recognition sets 105, an audio user interface 106, and a speech recognizer 107. The sensor A 102, sensor B 103, or both sensors receive signals indicating a change of position of at least one part of the human body. This change in position may include a change in orientation, direction, posture, or position, for example. The sensor may be sensing the direction a person's head is facing, for example. Also the sensor may be sensing the orientation of a person's hand, arm or leg, as another example. The change in position may also include motion. For example, the sensor may sense the speed at which the hand moves or how fast the feet are engaging a surface during a running exercise. The sensor may be a tactile sensor, a motion sensor, or any other sensor which can sense a change of position of at least one part of the human body. For example, a surface computing device may sense the arms on the surface and at the same time a finger moving substantially along the surface. Surface computing is the use of a specialized computer GUI (Graphic User Interface) in which traditional GUI elements are replaced by intuitive, everyday objects. Instead of a keyboard and mouse, the user interacts directly with a touch-sensitive screen, replicating the familiar hands-on experience of everyday object manipulation. In this example, the information of the relative positions of the arms, hands, and fingers in contact with the surface may be one type of input given by the sensor, and a finger motion on the surface may be another type of input given by the sensor. The change of position may be a portion of the human hand proximate with a surface (eg. touching the surface). The sensor B 103 may be an active sensor which may be configured to retrieve the information desired. For example, sensor B 103 may be an image sensor which may be configured to a bright light environment or a dark environment.

Controller 104 is coupled to sensor A 102 and/or sensor B 103. The controller receives the signals regarding changes in position of at least one part of a human body from at least one sensor. The controller is also coupled to select a first recognition set from repository of recognition sets 105 based on the change of position. The repository of recognition sets may be on a local hard drive, be distributed in a local network, or may be distributed across the internet, for example. The first recognition set may be a subset of another recognition set dictated by the state of the program running on the system. For example the, controller may be a game controller and may already have a recognition set selected corresponding to the state of the game. The recognition sets 105 may be word sets, segments of sounds, snippets of sounds, or representations of sounds. The controller is also coupled to a recognizer 107. The controller loads the first recognition set into the recognizer 107. A speech input is provided to the audio user interface 106. The audio user interface 106 converts the speech input into a speech input signal appropriate to be processed. This may be a conversion of the audio signal to a digital format, for example. The audio user interface 106 output is coupled to the recognizer 107. The speech recognizer 107 recognizes the speech input signal in the context of the first recognition set 108. The recognizer 107 may come up with a list of probable elements and weight the elements based on the segments of sounds in the first recognition set, for example. Alternately, the recognizer may simply use the elements in the first recognition set to compare to the speech input signal for a best match, for example. The smaller the first recognition set, the faster and more accurate the recognition can be. The controller receives the recognition result and processes the command or request. The controller is also coupled to an output device 101. The controller conveys the changes in the program flow to the output device 101. The output device 101 may be a video display which the controller may elect to depict a new state of a video game, for example. Also, the output device may be a control mechanism in a manufacturing line which the controller may command to change configuration, for example. Also, the output device 101 may be an audio output which the controller communicates alternatives to the user, for example.

FIG. 2 illustrates a method 200 for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention. At 201, a change of position of at least on part of a human body is sensed. As mentioned earlier, the change of position may include a change in the direction a human is facing or an orientation of one or more parts of the human body, for example. At 202, a first recognition set is selected based on the change in position. The change in position may indicate that a certain command set be used, for example. In one example, a human being raises his hand while playing a video game and the program selects a first recognition set. The first recognition set may look like the following.

{shield, stop, duck, jump, map}

At 203, a speech input is received. The speech input signal may be a digital signal representing an utterance which was spoken by the user. At 204, the speech input signal is recognized in the context of the first recognition set. For example, a human being playing the video game described above may have said “shield up”. This speech may be compared against the first recognition set to recognize the phrase, for example. In this example the element “shield” may be the recognition result which may allow the character in the video game to be protected by a shield.

FIG. 3 illustrates another method 300 for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention. At 301, a state of a computer program is initiated. This may be a program sequence which is based on several previous inputs. For example, the video program mentioned above may initiate a state of the video game in which the user's character is making his way through a virtual forest. At 302, a state recognition set is selected based on the state of the computer program. The program may be at a point when only a limited number of options are available and in this way a recognition set may correspond to these options, for example. In the video game example, the user's character is making his way through a virtual forest and may have a state recognition set as follows.

{shield, stop, duck, jump, map, run, walk, left, right, sword, lance,
climb tree, talk, ride, mount, borrow, steal, fire, call}

At 303, a change of position of least one part of a human body is sensed. This includes the examples of sensing a change of position mentioned previously. In the video game example, the user may change the position of his hand while holding a game sensor (eg. a game sensor 502 illustrated in FIG. 5 below). He may hold the game sensor 502 up so that his fingers are substantially vertical to one another (510). At 304, a first recognition set is selected based on the change in position, wherein the first recognition set includes finding a subset of the state recognition set. In the video game example, the subset of the state recognition set above would be selected based on the change of position. In this case, the subset may look as follows.

{shield, stop, duck, jump, map}

This would be called the first recognition set in this example. At 305, a speech input is received. The speech input signal may be a digital signal representing a portion of an utterance which was spoken by the user. At 306, the speech input signal is recognized in the context of the first recognition set. For example, a human being playing the video game, described above, may have said “stop”. This speech may be compared against the first recognition set to recognize the phrase, for example. In this example, the element “stop” may be the recognition result which may allow the character in the video game to stop walking or running within the virtual forest.

FIG. 4 illustrates another method 400 for performing speech recognition using sensory inputs of human position according to one embodiment of the present invention. At 401, a state of a computer is initiated. This may be a program sequence which is based on several previous inputs. For example, the video program mentioned above may initiate a state of the video game in which the user's character is making his way through the courtyard of a virtual castle. At 402, a state recognition set is selected based on the state of the computer program. The program may be at a point when only a limited number of options are available and in this way a recognition set may correspond to these options, for example. In the video game example, the user's character is making his way through the courtyard of a virtual castle and may have a state recognition set as follows.

{shield, stop, duck, jump, map, run, walk, crawl, climb, up stairs, enter
door, close door, enter window, close window, open crate, close crate,
draw bridge, left, right, sword, knife, key, climb, talk, smile, sell,
buy, call, lift veil}

At 403, a change of position of least one part of a human body is sensed. This includes the examples of sensing a change of position mentioned previously. In the video game example, the user may change the position of his hand while holding a game controller. He may hold the game sensor flat so that his palm is substantially facing down. At 404, a first recognition set is selected based on the change in position, wherein the first recognition set includes finding a subset of the state recognition set. In the video game example, the subset of the state recognition set above would be selected based on the change of position. In this case, the subset may look as follows.

{enter door, enter window, open crate, lift veil}

This may be called the first recognition set in this example. At 405, a speech input is received. The speech input signal may be a digital signal representing an utterance which was spoken by the user. At 406, the speech input signal is recognized in the context of the first recognition set. For example, a human being playing the video game described above may have said “enter door”. This speech may be compared against the first recognition set to recognize the phrase, for example. At 407, a new state of the computer is selected based on the recognition result. In the video game example, the element “enter door” may be the recognition result which may prompt the video game to select a new state of the video game program based on the recognition result “enter door”. This new state of the video game may be initiated and the program may display the room entered on the video screen. Also the video game program may select a new state recognition set based on the new state of the video game program and continue the method all over again.

FIGS. 5A and 5B illustrates an example of how the position of a human hand may be used for performing speech recognition according to one embodiment of the present invention. FIGS. 5A and 5B includes human arm 501, game sensor 502, an indication of the direction of the change of position 503, a first orientation of a human hand 504, a second orientation of a human hand 514, and a microphone 505. When the hand has changed position from the first orientation of the human hand 504 to the second orientation of the human hand 514, the game sensor senses the change and the system may user this information to select a first recognition set. This sensor may be a gyroscope, for example. Again referring to the video game example, a first recognition set may look like the following.

{sword, lance, fire}

This first recognition set may be a subset of a state recognition set. This state recognition set may look like the following.

{shield, stop, duck, jump, map, run, walk, left, right, sword, lance,
climb tree, talk, ride, mount, borrow, steal, fire, call}

This selection of a first recognition set has been described in methods 200, 300, and 400. The selection occurs at 202, 304, and 404, respectfully, in regards to this example. The speech enters microphone 505. This would be used at 203, 305, and 405 in the aforementioned methods along with appropriate electronic circuitry (amplifier, analog-to-digital converter, etc. . . . ) in order to receive a speech input signal. Next, the speech input signal is recognized in the context of the first recognition set. This has been described previously.

FIGS. 6A and 6B illustrates an example of how the position of a human head may be used for performing speech recognition according to another embodiment of the present invention. FIGS. 6A and 6B includes a microphone 601, a human head 602, a direction of rotation 603, a position sensing device 604 attached to said human head 602, instrument panel A 605, instrument panel B 606, instrument panel C 607, a first facing direction 608, and a second facing direction 618. In this embodiment, a human user having a human head 608 in a position 600 in which the human head is in the first facing direction 608 is speaking commands. Instrument panel A, B, and C may be panels which perform different functions in an automated manufacturing line, a monitoring station in a power plant, or a control room for a television broadcasting studio, for example. In this embodiment, position 600 is part of a first change in position. And according to this embodiment, a panel A recognition set has been selected based on a change to the first facing direction 608, and panel A recognition set includes commands associated with instrument panel A 605. In one example, instrument panel A 605 may be an instrument for controlling the lighting in a television studio and the panel A recognition set may look as follows.

{back lights, left lights, right lights, fade, up, down, sequence, one, two,
three}

The human associated with said head 602 may say “back lights fade sequence two”. Each segment of speech would be processed in the context of the panel A recognition set and the controls on the panel would command the lighting system to fade the back lights according to sequence two. This sequence may be a pre-programmed sequence which gives some desired affect in terms of how fast the back lights fade.

At some time the human associated with said head 602 may rotate in the direction 603 until said human head 602 is facing in the second facing direction 618. Sensor 604 would sense the change of position. And according to this example, a panel C recognition set may be selected based on a change to the second facing direction 618, and panel C recognition set includes commands associated with instrument panel C 607. In the television studio example, instrument panel C 607 may be an instrument for controlling the cameras in the television studio and the panel C recognition set may look as follows.

{front, camera, up, down, tilt, up, down, pan, left, right, sequence, one,
two, three}

The human associated with said head 602 may say “left camera pan left left left”. Each segment of speech would be processed in the context of the “panel A” recognition set and the controls on the panel may command the left camera to pan left three increments.

FIG. 7 illustrates another example of how the position of parts of a human body on a tactile screen may be used for performing speech recognition according to another embodiment of the present invention. This example, illustrates a surface computing application. FIG. 7 illustrates a scenario 700 which includes a human user 701, a tactile computer interface surface 704, and a microphone 709 attached to a headset 712. The human user 701 includes a left hand 702, a right hand 703, and a head 711. The human user 701 is wearing the headset 712 on his head 711. The tactile computer interface surface 704 may be integrated into a table or an office desk. Visible on the computer interface surface 704 is a virtual filing cabinet 705, a telephone symbol 706, a virtual stack of files 707, a virtual open file 708, and a shape 710 which denotes the sensing of a change of position of a portion of a human hand moving substantially along a portion of the tactile surface 704. This portion of the hand happens to be the index finger of the right hand 703 of the human user 701. In this example, the portion of the right hand is proximate with the computer interface surface 704.

In the scenario 700, human user 701 is manipulating the virtual stack of files 707 and reading the contents of the virtual open file 708. She realizes that something within the virtual open file 708 needs to be clarified and decides to call her client Carl Fiyel. She moves her right hand 703 to a portion of the tactile surface 704 proximate to the phone symbol 706. The tactile computer interface surface 704 senses the change of position. Since the change is in the vicinity of the phone symbol 706, a computer program selects a telephone recognition set. In this example the telephone recognition set may look as follows.

{call, redial, voicemail, line 1, line 2, hold, address book, home, Wilma,
Peter Henry, Mom, Sister, Karl Nale, George Smith, Tom Jones,
George Martinez, Carl Fiyel, Daniel Fihel, Marlo Stiles, Camp Fire
Girls, Marlo Stiles, Sara Chen, Larry Popadopolis, Nina Nbeheru,
Macey Epstein, etc . . . }

In this example, the people in the user's virtual telephone directory are part of the recognition set. If the human user 701 has many contacts, this set may contain hundred's of entries. The human user 701 now moves the index finger of her right hand 703 proximate with the tactile surface in a manner resembling the letter “C” 710. This change of position selects a snippet recognition set. This set may represent snippets or other types of segments of speech which correspond to the letter “C”, for example. This set may look as follows.

{ca, ce, co, cu}

The human user 701 now utters the phrase, “Call Carl Fiyel” into microphone 709. A recognizer recognizes the resulting speech input signal in the context of the telephone recognition set and the snippet recognition set. First the recognizer may come up with an initial guess at the utterance, “Call Carl Fiyel”. A possible first guess may be a set of ordered entries followed by the likelihoods associated with those entries. The top 5 entries may be part of this initial guess. The set may look as follows.

{Call Marlo Stiles, 452,
Call Carl Fiyel, 233,
Call Camp Fire Girls, 230,
Call Daniel Fiyel, 52,
Call Karl Nale, 34}

Next, the snippet recognition set may be used to weight this set of likelihoods according to how well they match the sound segments associated with the letter “C”. The resulting reordered set may look as follows.

{Call Carl Fiyel, 723,
Call Camp Fire Girls, 320,
Call Marlo Stiles, 270,
Call Karl Nale, 143,
Call Daniel Fiyel, 125}

In this example, the recognizing of the speech input signal includes using the snippet recognition set to weight a set of likelihoods, wherein the snippet recognition set includes segments of speech. The highest weighted likelihood is taken as the recognition result and the system dials the number of Carl Fiyel located within a database accessible to the computer.

The human user talks to Carl Fiyel on headset 712 and clarifies the issue. The human user edits the virtual open file 708, closes the virtual file, and adds the file to the virtual stack of files 707. Since all the files relate to “D Industries”, the human user 701 chooses to file them together. She moves the virtual stack of files 707 to the virtual filing cabinet 705. The tactile surface senses the change of position of the left hand 702 across the screen. The file manager program selects a file cabinet recognition set based on the motion of the left hand 702 and the state of the program running the system. The file cabinet recognition set may look as follows.

{file, other file, find, index, retrieve, copy, paste, subject, company,
modified, created, owner, Walsh Industries, Company A,
Company B, Firm C, D Industries, Miscellaneous Notes,
Research Project A, Research Project B, George Smith, Tom
Jones, George Martinez, Carl Fiyel, Daniel Fihel, Marlo
Stiles, George Form, Sara Chen, Larry Popadopolis, Macey
pstein, etc . . . }

In this example, a file manager program has included all the information associated with the different operations a user may wish to perform with the virtual filing cabinet 705. The human user 701 utters “file D Industries”. The system converts this utterance into two smaller portions. A first speech input signal includes “file” and a second speech input signal includes “D Industries”. In this example, the first speech input signal is considered a command. A subset of the file cabinet recognition set is selected based on the fact that this is the first speech input signal. This command recognition set may look as follows.

{file, find, index, retrieve, copy, paste}

The first speech input signal is recognized in the context of the command recognition set for the file manager program. This results in the “file” element being chosen as the recognition result. The file manager program now initiates an inventory of other states within other programs running within the system. Since the command “file” has been recognized, the system determines which files have been removed from the filing cabinet and which files have been created during this session of use and a scenario recognition set is selected which would include the possible choices for what the user may want to file. This scenario recognition set may look as follows.

{D Industries, Miscellaneous Notes, Carl Fiyel, Daniel Fihel, new, other}

There are six possibilities in filing this stack of files 707. The second speech input signal is recognized in the context of the scenario recognition set. “D Industries” element is chosen as the recognition result. This may result in selecting a new the state of a file manager program and initiating a routine within this new state. This may include storing the stack of files 707 into the “D Industries” folder. Initiating a new state of the file manager program may begin the method of selecting another recognition set based on the new state of the file manager program.

The above description illustrates various embodiments of the present invention along with examples of how aspects of the present invention may be implemented. The above examples and embodiments should not be deemed to be the only embodiments, and are presented to illustrate the flexibility and advantages of the present invention as defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations and equivalents will be evident to those skilled in the art and may be employed without departing from the spirit and scope of the invention as defined by the claims. The terms and expressions that have been employed here are used to describe the various embodiments and examples. These terms and expressions are not to be construed as excluding equivalents of the features shown and described, or portions thereof, it being recognized that various modifications are possible within the scope of the appended claims.

Claims (15)

1. A speech recognition method comprising:
sensing a change of position of at least one part of a human body, the sensing being performed by a gyroscope, wherein the gyroscope senses said change in position through physical contact with said at least one part of the human body;
selecting, by a controller, a speech recognition set based on the change of position, the speech recognition set specifying a set of words to be recognized, wherein when the gyroscope indicates that at least one part of the human body is in a first position, a first speech recognition set is selected to recognize a first set of words, and wherein when the gyroscope indicates that the at least one part of the human body is in a second position, a second speech recognition set is selected to recognize a second set of words different from the first set of words;
receiving a speech input signal; and
recognizing the speech input signal in the context of the selected speech recognition set, the recognizing resulting in a recognition result.
2. The method of claim 1 wherein the change of position includes a change in orientation.
3. The method of claim 1 wherein the change of position includes a change in direction.
4. The method of claim 1 wherein the change of position includes motion.
5. The method of claim 1 wherein the second speech recognition set is a subset of the first recognition set.
6. The method of claim 1 wherein selecting the speech recognition set comprises selecting a subset of words from a state recognition set for a computer program, the state recognition set comprising a set of words corresponding to available options of the computer program when the computer program is in a particular state.
7. The method of claim 1 further comprising:
sensing a second change in position of at least one part of a human body; and
modifying the recognition result based on the second change in position.
8. The method of claim 1 further comprising:
initiating a state of a computer program; and
selecting a state recognition set based on the state of the computer program,
wherein the selecting of the first recognition set includes finding a subset of the state recognition set.
9. The method of claim 8 wherein the change of position includes a change in orientation.
10. The method of claim 8 wherein the change of position includes a change in a direction.
11. The method of claim 8 wherein the change of position includes motion.
12. The method of claim 8 wherein the recognizing includes choosing a select element from the first recognition set.
13. The method of claim 8 further comprising changing the state of the computer program according to the recognition result.
14. The method of claim 1 wherein the speech input signal includes digital data.
15. The method of claim 1 wherein the first recognition set includes segments of sound data.
US11973140 2007-10-05 2007-10-05 Systems and methods of performing speech recognition using sensory inputs of human position Active 2030-07-14 US8036901B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US11973140 US8036901B2 (en) 2007-10-05 2007-10-05 Systems and methods of performing speech recognition using sensory inputs of human position

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US11973140 US8036901B2 (en) 2007-10-05 2007-10-05 Systems and methods of performing speech recognition using sensory inputs of human position
PCT/US2008/077742 WO2009045861A1 (en) 2007-10-05 2008-09-25 Systems and methods of performing speech recognition using gestures
US12238287 US8321219B2 (en) 2007-10-05 2008-09-25 Systems and methods of performing speech recognition using gestures

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US12238287 Continuation-In-Part US8321219B2 (en) 2007-10-05 2008-09-25 Systems and methods of performing speech recognition using gestures

Publications (2)

Publication Number Publication Date
US20090094032A1 true US20090094032A1 (en) 2009-04-09
US8036901B2 true US8036901B2 (en) 2011-10-11

Family

ID=40524026

Family Applications (1)

Application Number Title Priority Date Filing Date
US11973140 Active 2030-07-14 US8036901B2 (en) 2007-10-05 2007-10-05 Systems and methods of performing speech recognition using sensory inputs of human position

Country Status (1)

Country Link
US (1) US8036901B2 (en)

Cited By (91)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090166122A1 (en) * 2007-12-27 2009-07-02 Byd Co. Ltd. Hybrid Vehicle Having Power Assembly Arranged Transversely In Engine Compartment
US8296383B2 (en) 2008-10-02 2012-10-23 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US8311838B2 (en) 2010-01-13 2012-11-13 Apple Inc. Devices and methods for identifying a prompt corresponding to a voice input in a sequence of prompts
US8345665B2 (en) 2001-10-22 2013-01-01 Apple Inc. Text to speech conversion of text messages from mobile communication devices
US8352272B2 (en) 2008-09-29 2013-01-08 Apple Inc. Systems and methods for text to speech synthesis
US8352268B2 (en) 2008-09-29 2013-01-08 Apple Inc. Systems and methods for selective rate of speech and speech preferences for text to speech synthesis
US8355919B2 (en) 2008-09-29 2013-01-15 Apple Inc. Systems and methods for text normalization for text to speech synthesis
US8380507B2 (en) 2009-03-09 2013-02-19 Apple Inc. Systems and methods for determining the language to use for speech generated by a text to speech engine
US8396714B2 (en) 2008-09-29 2013-03-12 Apple Inc. Systems and methods for concatenation of words in text to speech synthesis
US8458278B2 (en) 2003-05-02 2013-06-04 Apple Inc. Method and apparatus for displaying information during an instant messaging session
US8527861B2 (en) 1999-08-13 2013-09-03 Apple Inc. Methods and apparatuses for display and traversing of links in page character array
US8543407B1 (en) 2007-10-04 2013-09-24 Great Northern Research, LLC Speech interface system and method for control and interaction with applications on a computing system
US8583418B2 (en) 2008-09-29 2013-11-12 Apple Inc. Systems and methods of detecting language and natural language strings for text to speech synthesis
US8600743B2 (en) 2010-01-06 2013-12-03 Apple Inc. Noise profile determination for voice-related feature
US8614431B2 (en) 2005-09-30 2013-12-24 Apple Inc. Automated response to and sensing of user activity in portable devices
US8620662B2 (en) 2007-11-20 2013-12-31 Apple Inc. Context-aware unit selection
US8639516B2 (en) 2010-06-04 2014-01-28 Apple Inc. User-specific noise suppression for voice quality improvements
US8645137B2 (en) 2000-03-16 2014-02-04 Apple Inc. Fast, language-independent method for user authentication by voice
US8660849B2 (en) 2010-01-18 2014-02-25 Apple Inc. Prioritizing selection criteria by automated assistant
US8677377B2 (en) 2005-09-08 2014-03-18 Apple Inc. Method and apparatus for building an intelligent automated assistant
US8682667B2 (en) 2010-02-25 2014-03-25 Apple Inc. User profiling for selecting user specific voice input processing information
US8682649B2 (en) 2009-11-12 2014-03-25 Apple Inc. Sentiment prediction from textual data
US8688446B2 (en) 2008-02-22 2014-04-01 Apple Inc. Providing text input using speech data and non-speech data
US8706472B2 (en) 2011-08-11 2014-04-22 Apple Inc. Method for disambiguating multiple readings in language conversion
US8712776B2 (en) 2008-09-29 2014-04-29 Apple Inc. Systems and methods for selective text to speech synthesis
US8713021B2 (en) 2010-07-07 2014-04-29 Apple Inc. Unsupervised document clustering using latent semantic density analysis
US8719014B2 (en) 2010-09-27 2014-05-06 Apple Inc. Electronic device with text error correction based on voice recognition data
US8719006B2 (en) 2010-08-27 2014-05-06 Apple Inc. Combined statistical and rule-based part-of-speech tagging for text-to-speech synthesis
US8762156B2 (en) 2011-09-28 2014-06-24 Apple Inc. Speech recognition repair using contextual information
US8768702B2 (en) 2008-09-05 2014-07-01 Apple Inc. Multi-tiered voice feedback in an electronic device
US8775442B2 (en) 2012-05-15 2014-07-08 Apple Inc. Semantic search using a single-source semantic model
US8781836B2 (en) 2011-02-22 2014-07-15 Apple Inc. Hearing assistance system for providing consistent human speech
US8812294B2 (en) 2011-06-21 2014-08-19 Apple Inc. Translating phrases from one language into another using an order-based set of declarative rules
US8862252B2 (en) 2009-01-30 2014-10-14 Apple Inc. Audio user interface for displayless electronic device
US8898568B2 (en) 2008-09-09 2014-11-25 Apple Inc. Audio user interface
US8935167B2 (en) 2012-09-25 2015-01-13 Apple Inc. Exemplar-based latent perceptual modeling for automatic speech recognition
US8977255B2 (en) 2007-04-03 2015-03-10 Apple Inc. Method and system for operating a multi-function portable electronic device using voice-activation
US8996376B2 (en) 2008-04-05 2015-03-31 Apple Inc. Intelligent text-to-speech conversion
US9053089B2 (en) 2007-10-02 2015-06-09 Apple Inc. Part-of-speech tagging using latent analogy
US9262612B2 (en) 2011-03-21 2016-02-16 Apple Inc. Device access using voice authentication
US9280610B2 (en) 2012-05-14 2016-03-08 Apple Inc. Crowd sourcing information to fulfill user requests
US9300784B2 (en) 2013-06-13 2016-03-29 Apple Inc. System and method for emergency calls initiated by voice command
US9311043B2 (en) 2010-01-13 2016-04-12 Apple Inc. Adaptive audio feedback system and method
US9330381B2 (en) 2008-01-06 2016-05-03 Apple Inc. Portable multifunction device, method, and graphical user interface for viewing and managing electronic calendars
US9330720B2 (en) 2008-01-03 2016-05-03 Apple Inc. Methods and apparatus for altering audio output signals
US9338493B2 (en) 2014-06-30 2016-05-10 Apple Inc. Intelligent automated assistant for TV user interactions
US9368114B2 (en) 2013-03-14 2016-06-14 Apple Inc. Context-sensitive handling of interruptions
US9431006B2 (en) 2009-07-02 2016-08-30 Apple Inc. Methods and apparatuses for automatic speech recognition
US9430463B2 (en) 2014-05-30 2016-08-30 Apple Inc. Exemplar-based natural language processing
US9483461B2 (en) 2012-03-06 2016-11-01 Apple Inc. Handling speech synthesis of content for multiple languages
US9495129B2 (en) 2012-06-29 2016-11-15 Apple Inc. Device, method, and user interface for voice-activated navigation and browsing of a document
US9502031B2 (en) 2014-05-27 2016-11-22 Apple Inc. Method for supporting dynamic grammars in WFST-based ASR
US9535906B2 (en) 2008-07-31 2017-01-03 Apple Inc. Mobile device having human language translation capability with positional feedback
US9547647B2 (en) 2012-09-19 2017-01-17 Apple Inc. Voice-based media searching
US9576574B2 (en) 2012-09-10 2017-02-21 Apple Inc. Context-sensitive handling of interruptions by intelligent digital assistant
US9582608B2 (en) 2013-06-07 2017-02-28 Apple Inc. Unified ranking with entropy-weighted information for phrase-based semantic auto-completion
US9606986B2 (en) 2014-09-29 2017-03-28 Apple Inc. Integrated word N-gram and class M-gram language models
US9620105B2 (en) 2014-05-15 2017-04-11 Apple Inc. Analyzing audio input for efficient speech and music recognition
US9620104B2 (en) 2013-06-07 2017-04-11 Apple Inc. System and method for user-specified pronunciation of words for speech synthesis and recognition
US9633674B2 (en) 2013-06-07 2017-04-25 Apple Inc. System and method for detecting errors in interactions with a voice-based digital assistant
US9633004B2 (en) 2014-05-30 2017-04-25 Apple Inc. Better resolution when referencing to concepts
US9646609B2 (en) 2014-09-30 2017-05-09 Apple Inc. Caching apparatus for serving phonetic pronunciations
US9668121B2 (en) 2014-09-30 2017-05-30 Apple Inc. Social reminders
US9697822B1 (en) 2013-03-15 2017-07-04 Apple Inc. System and method for updating an adaptive speech recognition model
US9697820B2 (en) 2015-09-24 2017-07-04 Apple Inc. Unit-selection text-to-speech synthesis using concatenation-sensitive neural networks
US9711141B2 (en) 2014-12-09 2017-07-18 Apple Inc. Disambiguating heteronyms in speech synthesis
US9715875B2 (en) 2014-05-30 2017-07-25 Apple Inc. Reducing the need for manual start/end-pointing and trigger phrases
US9721566B2 (en) 2015-03-08 2017-08-01 Apple Inc. Competing devices responding to voice triggers
US9721563B2 (en) 2012-06-08 2017-08-01 Apple Inc. Name recognition system
US9733821B2 (en) 2013-03-14 2017-08-15 Apple Inc. Voice control to diagnose inadvertent activation of accessibility features
US9760559B2 (en) 2014-05-30 2017-09-12 Apple Inc. Predictive text input
US9785630B2 (en) 2014-05-30 2017-10-10 Apple Inc. Text prediction using combined word N-gram and unigram language models
US9798393B2 (en) 2011-08-29 2017-10-24 Apple Inc. Text correction processing
US9818400B2 (en) 2014-09-11 2017-11-14 Apple Inc. Method and apparatus for discovering trending terms in speech requests
US9830913B2 (en) 2013-10-29 2017-11-28 Knowles Electronics, Llc VAD detection apparatus and method of operation the same
US9842101B2 (en) 2014-05-30 2017-12-12 Apple Inc. Predictive conversion of language input
US9842105B2 (en) 2015-04-16 2017-12-12 Apple Inc. Parsimonious continuous-space phrase representations for natural language processing
US9858925B2 (en) 2009-06-05 2018-01-02 Apple Inc. Using context information to facilitate processing of commands in a virtual assistant
US9865280B2 (en) 2015-03-06 2018-01-09 Apple Inc. Structured dictation using intelligent automated assistants
US9886953B2 (en) 2015-03-08 2018-02-06 Apple Inc. Virtual assistant activation
US9886432B2 (en) 2014-09-30 2018-02-06 Apple Inc. Parsimonious handling of word inflection via categorical stem + suffix N-gram language models
US9899019B2 (en) 2015-03-18 2018-02-20 Apple Inc. Systems and methods for structured stem and suffix language models
US9922642B2 (en) 2013-03-15 2018-03-20 Apple Inc. Training an at least partial voice command system
US9934775B2 (en) 2016-05-26 2018-04-03 Apple Inc. Unit-selection text-to-speech synthesis based on predicted concatenation parameters
US9946706B2 (en) 2008-06-07 2018-04-17 Apple Inc. Automatic language identification for dynamic text processing
US9959870B2 (en) 2008-12-11 2018-05-01 Apple Inc. Speech recognition involving a mobile device
US9966068B2 (en) 2013-06-08 2018-05-08 Apple Inc. Interpreting and acting upon commands that involve sharing information with remote devices
US9966065B2 (en) 2014-05-30 2018-05-08 Apple Inc. Multi-command single utterance input method
US9972304B2 (en) 2016-06-03 2018-05-15 Apple Inc. Privacy preserving distributed evaluation framework for embedded personalized systems
US9977779B2 (en) 2013-03-14 2018-05-22 Apple Inc. Automatic supplementation of word correction dictionaries
US9986419B2 (en) 2017-05-26 2018-05-29 Apple Inc. Social reminders

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8568189B2 (en) * 2009-11-25 2013-10-29 Hallmark Cards, Incorporated Context-based interactive plush toy
US9421475B2 (en) 2009-11-25 2016-08-23 Hallmark Cards Incorporated Context-based interactive plush toy

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6292158B1 (en) * 1997-05-08 2001-09-18 Shimadzu Corporation Display system
US20030046087A1 (en) 2001-08-17 2003-03-06 At&T Corp. Systems and methods for classifying and representing gestural inputs
US20060013440A1 (en) 1998-08-10 2006-01-19 Cohen Charles J Gesture-controlled interfaces for self-service machines and other applications

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6292158B1 (en) * 1997-05-08 2001-09-18 Shimadzu Corporation Display system
US20060013440A1 (en) 1998-08-10 2006-01-19 Cohen Charles J Gesture-controlled interfaces for self-service machines and other applications
US20030046087A1 (en) 2001-08-17 2003-03-06 At&T Corp. Systems and methods for classifying and representing gestural inputs

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
International Search Report (from corresponding PCT application), PCT/US08/77742, mailed Dec. 2, 2008.

Cited By (125)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8527861B2 (en) 1999-08-13 2013-09-03 Apple Inc. Methods and apparatuses for display and traversing of links in page character array
US8645137B2 (en) 2000-03-16 2014-02-04 Apple Inc. Fast, language-independent method for user authentication by voice
US9646614B2 (en) 2000-03-16 2017-05-09 Apple Inc. Fast, language-independent method for user authentication by voice
US8345665B2 (en) 2001-10-22 2013-01-01 Apple Inc. Text to speech conversion of text messages from mobile communication devices
US8718047B2 (en) 2001-10-22 2014-05-06 Apple Inc. Text to speech conversion of text messages from mobile communication devices
US8458278B2 (en) 2003-05-02 2013-06-04 Apple Inc. Method and apparatus for displaying information during an instant messaging session
US9501741B2 (en) 2005-09-08 2016-11-22 Apple Inc. Method and apparatus for building an intelligent automated assistant
US8677377B2 (en) 2005-09-08 2014-03-18 Apple Inc. Method and apparatus for building an intelligent automated assistant
US8614431B2 (en) 2005-09-30 2013-12-24 Apple Inc. Automated response to and sensing of user activity in portable devices
US9619079B2 (en) 2005-09-30 2017-04-11 Apple Inc. Automated response to and sensing of user activity in portable devices
US9958987B2 (en) 2005-09-30 2018-05-01 Apple Inc. Automated response to and sensing of user activity in portable devices
US9389729B2 (en) 2005-09-30 2016-07-12 Apple Inc. Automated response to and sensing of user activity in portable devices
US9117447B2 (en) 2006-09-08 2015-08-25 Apple Inc. Using event alert text as input to an automated assistant
US8942986B2 (en) 2006-09-08 2015-01-27 Apple Inc. Determining user intent based on ontologies of domains
US8930191B2 (en) 2006-09-08 2015-01-06 Apple Inc. Paraphrasing of user requests and results by automated digital assistant
US8977255B2 (en) 2007-04-03 2015-03-10 Apple Inc. Method and system for operating a multi-function portable electronic device using voice-activation
US9053089B2 (en) 2007-10-02 2015-06-09 Apple Inc. Part-of-speech tagging using latent analogy
US8543407B1 (en) 2007-10-04 2013-09-24 Great Northern Research, LLC Speech interface system and method for control and interaction with applications on a computing system
US8620662B2 (en) 2007-11-20 2013-12-31 Apple Inc. Context-aware unit selection
US20090166122A1 (en) * 2007-12-27 2009-07-02 Byd Co. Ltd. Hybrid Vehicle Having Power Assembly Arranged Transversely In Engine Compartment
US9330720B2 (en) 2008-01-03 2016-05-03 Apple Inc. Methods and apparatus for altering audio output signals
US9330381B2 (en) 2008-01-06 2016-05-03 Apple Inc. Portable multifunction device, method, and graphical user interface for viewing and managing electronic calendars
US9361886B2 (en) 2008-02-22 2016-06-07 Apple Inc. Providing text input using speech data and non-speech data
US8688446B2 (en) 2008-02-22 2014-04-01 Apple Inc. Providing text input using speech data and non-speech data
US9865248B2 (en) 2008-04-05 2018-01-09 Apple Inc. Intelligent text-to-speech conversion
US8996376B2 (en) 2008-04-05 2015-03-31 Apple Inc. Intelligent text-to-speech conversion
US9626955B2 (en) 2008-04-05 2017-04-18 Apple Inc. Intelligent text-to-speech conversion
US9946706B2 (en) 2008-06-07 2018-04-17 Apple Inc. Automatic language identification for dynamic text processing
US9535906B2 (en) 2008-07-31 2017-01-03 Apple Inc. Mobile device having human language translation capability with positional feedback
US8768702B2 (en) 2008-09-05 2014-07-01 Apple Inc. Multi-tiered voice feedback in an electronic device
US9691383B2 (en) 2008-09-05 2017-06-27 Apple Inc. Multi-tiered voice feedback in an electronic device
US8898568B2 (en) 2008-09-09 2014-11-25 Apple Inc. Audio user interface
US8712776B2 (en) 2008-09-29 2014-04-29 Apple Inc. Systems and methods for selective text to speech synthesis
US8355919B2 (en) 2008-09-29 2013-01-15 Apple Inc. Systems and methods for text normalization for text to speech synthesis
US8352272B2 (en) 2008-09-29 2013-01-08 Apple Inc. Systems and methods for text to speech synthesis
US8352268B2 (en) 2008-09-29 2013-01-08 Apple Inc. Systems and methods for selective rate of speech and speech preferences for text to speech synthesis
US8583418B2 (en) 2008-09-29 2013-11-12 Apple Inc. Systems and methods of detecting language and natural language strings for text to speech synthesis
US8396714B2 (en) 2008-09-29 2013-03-12 Apple Inc. Systems and methods for concatenation of words in text to speech synthesis
US8296383B2 (en) 2008-10-02 2012-10-23 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US8762469B2 (en) 2008-10-02 2014-06-24 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US9412392B2 (en) 2008-10-02 2016-08-09 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US8713119B2 (en) 2008-10-02 2014-04-29 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US8676904B2 (en) 2008-10-02 2014-03-18 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US9959870B2 (en) 2008-12-11 2018-05-01 Apple Inc. Speech recognition involving a mobile device
US8862252B2 (en) 2009-01-30 2014-10-14 Apple Inc. Audio user interface for displayless electronic device
US8751238B2 (en) 2009-03-09 2014-06-10 Apple Inc. Systems and methods for determining the language to use for speech generated by a text to speech engine
US8380507B2 (en) 2009-03-09 2013-02-19 Apple Inc. Systems and methods for determining the language to use for speech generated by a text to speech engine
US9858925B2 (en) 2009-06-05 2018-01-02 Apple Inc. Using context information to facilitate processing of commands in a virtual assistant
US9431006B2 (en) 2009-07-02 2016-08-30 Apple Inc. Methods and apparatuses for automatic speech recognition
US8682649B2 (en) 2009-11-12 2014-03-25 Apple Inc. Sentiment prediction from textual data
US8600743B2 (en) 2010-01-06 2013-12-03 Apple Inc. Noise profile determination for voice-related feature
US9311043B2 (en) 2010-01-13 2016-04-12 Apple Inc. Adaptive audio feedback system and method
US8311838B2 (en) 2010-01-13 2012-11-13 Apple Inc. Devices and methods for identifying a prompt corresponding to a voice input in a sequence of prompts
US8670985B2 (en) 2010-01-13 2014-03-11 Apple Inc. Devices and methods for identifying a prompt corresponding to a voice input in a sequence of prompts
US8706503B2 (en) 2010-01-18 2014-04-22 Apple Inc. Intent deduction based on previous user interactions with voice assistant
US9548050B2 (en) 2010-01-18 2017-01-17 Apple Inc. Intelligent automated assistant
US8799000B2 (en) 2010-01-18 2014-08-05 Apple Inc. Disambiguation based on active input elicitation by intelligent automated assistant
US8892446B2 (en) 2010-01-18 2014-11-18 Apple Inc. Service orchestration for intelligent automated assistant
US8670979B2 (en) 2010-01-18 2014-03-11 Apple Inc. Active input elicitation by intelligent automated assistant
US8903716B2 (en) 2010-01-18 2014-12-02 Apple Inc. Personalized vocabulary for digital assistant
US9318108B2 (en) 2010-01-18 2016-04-19 Apple Inc. Intelligent automated assistant
US8731942B2 (en) 2010-01-18 2014-05-20 Apple Inc. Maintaining context information between user interactions with a voice assistant
US8660849B2 (en) 2010-01-18 2014-02-25 Apple Inc. Prioritizing selection criteria by automated assistant
US9633660B2 (en) 2010-02-25 2017-04-25 Apple Inc. User profiling for voice input processing
US8682667B2 (en) 2010-02-25 2014-03-25 Apple Inc. User profiling for selecting user specific voice input processing information
US9190062B2 (en) 2010-02-25 2015-11-17 Apple Inc. User profiling for voice input processing
US8639516B2 (en) 2010-06-04 2014-01-28 Apple Inc. User-specific noise suppression for voice quality improvements
US8713021B2 (en) 2010-07-07 2014-04-29 Apple Inc. Unsupervised document clustering using latent semantic density analysis
US8719006B2 (en) 2010-08-27 2014-05-06 Apple Inc. Combined statistical and rule-based part-of-speech tagging for text-to-speech synthesis
US8719014B2 (en) 2010-09-27 2014-05-06 Apple Inc. Electronic device with text error correction based on voice recognition data
US9075783B2 (en) 2010-09-27 2015-07-07 Apple Inc. Electronic device with text error correction based on voice recognition data
US8781836B2 (en) 2011-02-22 2014-07-15 Apple Inc. Hearing assistance system for providing consistent human speech
US9262612B2 (en) 2011-03-21 2016-02-16 Apple Inc. Device access using voice authentication
US8812294B2 (en) 2011-06-21 2014-08-19 Apple Inc. Translating phrases from one language into another using an order-based set of declarative rules
US8706472B2 (en) 2011-08-11 2014-04-22 Apple Inc. Method for disambiguating multiple readings in language conversion
US9798393B2 (en) 2011-08-29 2017-10-24 Apple Inc. Text correction processing
US8762156B2 (en) 2011-09-28 2014-06-24 Apple Inc. Speech recognition repair using contextual information
US9483461B2 (en) 2012-03-06 2016-11-01 Apple Inc. Handling speech synthesis of content for multiple languages
US9953088B2 (en) 2012-05-14 2018-04-24 Apple Inc. Crowd sourcing information to fulfill user requests
US9280610B2 (en) 2012-05-14 2016-03-08 Apple Inc. Crowd sourcing information to fulfill user requests
US8775442B2 (en) 2012-05-15 2014-07-08 Apple Inc. Semantic search using a single-source semantic model
US9721563B2 (en) 2012-06-08 2017-08-01 Apple Inc. Name recognition system
US9495129B2 (en) 2012-06-29 2016-11-15 Apple Inc. Device, method, and user interface for voice-activated navigation and browsing of a document
US9576574B2 (en) 2012-09-10 2017-02-21 Apple Inc. Context-sensitive handling of interruptions by intelligent digital assistant
US9971774B2 (en) 2012-09-19 2018-05-15 Apple Inc. Voice-based media searching
US9547647B2 (en) 2012-09-19 2017-01-17 Apple Inc. Voice-based media searching
US8935167B2 (en) 2012-09-25 2015-01-13 Apple Inc. Exemplar-based latent perceptual modeling for automatic speech recognition
US9733821B2 (en) 2013-03-14 2017-08-15 Apple Inc. Voice control to diagnose inadvertent activation of accessibility features
US9368114B2 (en) 2013-03-14 2016-06-14 Apple Inc. Context-sensitive handling of interruptions
US9977779B2 (en) 2013-03-14 2018-05-22 Apple Inc. Automatic supplementation of word correction dictionaries
US9697822B1 (en) 2013-03-15 2017-07-04 Apple Inc. System and method for updating an adaptive speech recognition model
US9922642B2 (en) 2013-03-15 2018-03-20 Apple Inc. Training an at least partial voice command system
US9966060B2 (en) 2013-06-07 2018-05-08 Apple Inc. System and method for user-specified pronunciation of words for speech synthesis and recognition
US9633674B2 (en) 2013-06-07 2017-04-25 Apple Inc. System and method for detecting errors in interactions with a voice-based digital assistant
US9582608B2 (en) 2013-06-07 2017-02-28 Apple Inc. Unified ranking with entropy-weighted information for phrase-based semantic auto-completion
US9620104B2 (en) 2013-06-07 2017-04-11 Apple Inc. System and method for user-specified pronunciation of words for speech synthesis and recognition
US9966068B2 (en) 2013-06-08 2018-05-08 Apple Inc. Interpreting and acting upon commands that involve sharing information with remote devices
US9300784B2 (en) 2013-06-13 2016-03-29 Apple Inc. System and method for emergency calls initiated by voice command
US9830913B2 (en) 2013-10-29 2017-11-28 Knowles Electronics, Llc VAD detection apparatus and method of operation the same
US9620105B2 (en) 2014-05-15 2017-04-11 Apple Inc. Analyzing audio input for efficient speech and music recognition
US9502031B2 (en) 2014-05-27 2016-11-22 Apple Inc. Method for supporting dynamic grammars in WFST-based ASR
US9966065B2 (en) 2014-05-30 2018-05-08 Apple Inc. Multi-command single utterance input method
US9760559B2 (en) 2014-05-30 2017-09-12 Apple Inc. Predictive text input
US9842101B2 (en) 2014-05-30 2017-12-12 Apple Inc. Predictive conversion of language input
US9715875B2 (en) 2014-05-30 2017-07-25 Apple Inc. Reducing the need for manual start/end-pointing and trigger phrases
US9430463B2 (en) 2014-05-30 2016-08-30 Apple Inc. Exemplar-based natural language processing
US9633004B2 (en) 2014-05-30 2017-04-25 Apple Inc. Better resolution when referencing to concepts
US9785630B2 (en) 2014-05-30 2017-10-10 Apple Inc. Text prediction using combined word N-gram and unigram language models
US9668024B2 (en) 2014-06-30 2017-05-30 Apple Inc. Intelligent automated assistant for TV user interactions
US9338493B2 (en) 2014-06-30 2016-05-10 Apple Inc. Intelligent automated assistant for TV user interactions
US9818400B2 (en) 2014-09-11 2017-11-14 Apple Inc. Method and apparatus for discovering trending terms in speech requests
US9606986B2 (en) 2014-09-29 2017-03-28 Apple Inc. Integrated word N-gram and class M-gram language models
US9646609B2 (en) 2014-09-30 2017-05-09 Apple Inc. Caching apparatus for serving phonetic pronunciations
US9886432B2 (en) 2014-09-30 2018-02-06 Apple Inc. Parsimonious handling of word inflection via categorical stem + suffix N-gram language models
US9668121B2 (en) 2014-09-30 2017-05-30 Apple Inc. Social reminders
US9711141B2 (en) 2014-12-09 2017-07-18 Apple Inc. Disambiguating heteronyms in speech synthesis
US9865280B2 (en) 2015-03-06 2018-01-09 Apple Inc. Structured dictation using intelligent automated assistants
US9721566B2 (en) 2015-03-08 2017-08-01 Apple Inc. Competing devices responding to voice triggers
US9886953B2 (en) 2015-03-08 2018-02-06 Apple Inc. Virtual assistant activation
US9899019B2 (en) 2015-03-18 2018-02-20 Apple Inc. Systems and methods for structured stem and suffix language models
US9842105B2 (en) 2015-04-16 2017-12-12 Apple Inc. Parsimonious continuous-space phrase representations for natural language processing
US9697820B2 (en) 2015-09-24 2017-07-04 Apple Inc. Unit-selection text-to-speech synthesis using concatenation-sensitive neural networks
US9934775B2 (en) 2016-05-26 2018-04-03 Apple Inc. Unit-selection text-to-speech synthesis based on predicted concatenation parameters
US9972304B2 (en) 2016-06-03 2018-05-15 Apple Inc. Privacy preserving distributed evaluation framework for embedded personalized systems
US9986419B2 (en) 2017-05-26 2018-05-29 Apple Inc. Social reminders

Also Published As

Publication number Publication date Type
US20090094032A1 (en) 2009-04-09 application

Similar Documents

Publication Publication Date Title
US8165886B1 (en) Speech interface system and method for control and interaction with applications on a computing system
US7358956B2 (en) Method for providing feedback responsive to sensing a physical presence proximate to a control of an electronic device
Schmandt et al. Augmenting a window system with speech input
US20030028382A1 (en) System and method for voice dictation and command input modes
US7602382B2 (en) Method for displaying information responsive to sensing a physical presence proximate to a computer input device
US20040196400A1 (en) Digital camera user interface using hand gestures
US20020018051A1 (en) Apparatus and method for moving objects on a touchscreen display
US20070182595A1 (en) Systems to enhance data entry in mobile and fixed environment
US20090166098A1 (en) Non-visual control of multi-touch device
US20100318366A1 (en) Touch Anywhere to Speak
US20100009719A1 (en) Mobile terminal and method for displaying menu thereof
US20130297319A1 (en) Mobile device having at least one microphone sensor and method for controlling the same
US20080246778A1 (en) Controlling image and mobile terminal
US20080148149A1 (en) Methods, systems, and computer program products for interacting simultaneously with multiple application programs
US20100306718A1 (en) Apparatus and method for unlocking a locking mode of portable terminal
US20050195159A1 (en) Keyboardless text entry
US20120242584A1 (en) Method and apparatus for providing sight independent activity reports responsive to a touch gesture
US20090306980A1 (en) Mobile terminal and text correcting method in the same
US20120110518A1 (en) Translation of directional input to gesture
US20090199092A1 (en) Data entry system
US8325214B2 (en) Enhanced interface for voice and video communications
US20080189658A1 (en) Terminal and menu display method
US20100105364A1 (en) Mobile terminal and control method thereof
US20080189614A1 (en) Terminal and menu display method
Shafer et al. Interaction issues in context-aware intelligent environments

Legal Events

Date Code Title Description
AS Assignment

Owner name: SENSORY, INCORPORATED, CALIFORNIA

Free format text: ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNOR:MOZER, TODD F.;REEL/FRAME:019989/0707

Effective date: 20070824

AS Assignment

Owner name: ENABLENCE TECHNOLOGIES, INC, CANADA

Free format text: SECURITY AGREEMENT;ASSIGNOR:WAVE7 OPTICS, INC;REEL/FRAME:020817/0818

Effective date: 20080414

AS Assignment

Owner name: WAVE7 OPTICS, INC., GEORGIA

Free format text: RELEASE BY SECURED PARTY;ASSIGNOR:ENABLENCE;REEL/FRAME:020976/0779

Effective date: 20080520

AS Assignment

Owner name: SENSORY, INC., CALIFORNIA

Free format text: CORRECTIVE ASSIGNMENT BY DECLARATION TO REMOVE APPLICATION NUMBER 11973140 PREVIOUSLY RECORDED ON REEL 020976 / FRAME 0779;ASSIGNOR:SENSORY, INC.;REEL/FRAME:029317/0770

Effective date: 20121114

FPAY Fee payment

Year of fee payment: 4