KR101683083B1

KR101683083B1 - Using context information to facilitate processing of commands in a virtual assistant

Info

Publication number: KR101683083B1
Application number: KR1020120109552A
Authority: KR
Inventors: 토마스 로버트 그루버; 크리스토퍼 딘 브리검; 다니엘 에스. 킨; 그레고리. 노빅; 벤자민 에스. 핍스
Original assignee: 애플 인크.
Priority date: 2011-09-30
Filing date: 2012-09-28
Publication date: 2016-12-07
Also published as: JP2022119908A; BR102012024861A2; KR102048375B1; GB2495222B; JP2023169360A; CA3023918A1; KR20220136304A; NL2009544B1; KR20160142802A; KR102622737B1; EP2575128A2; DE102012019178A1; BR102012024861B1; CN103226949B; EP3200185A1; AU2012232977A1; RU2012141604A; JP6740162B2; GB201217449D0; CN103226949A

Abstract

가상 비서는 자연어 또는 사용자로부터의 제스처 입력을 대체하기 위하여 컨텍스트 정보를 이용한다. 컨텍스트는 사용자의 의도를 해명하고 사용자 입력의 해석 후보의 수를 감소시키는데 도움이 되며, 사용자가 과도한 해명 입력을 제공할 필요를 감소시킨다. 컨텍스트는 정보 처리 문제를 제한하기 위해 그리고/또는 결과를 맞춤화하기 위하여 비서가 명확한 사용자 입력을 대체할 수 없는 임의의 이용가능한 정보를 포함할 수 있다. 컨텍스트는 예컨대 음성 인식, 자연어 처리, 태스크 플로우 처리, 및 대화 생성을 포함하는 다양한 양상의 처리 동안 솔루션을 제한하는데 사용될 수 있다.The virtual assistant uses context information to replace gesture input from natural language or from the user. The context helps to clarify the intent of the user and reduce the number of interpretation candidates of the user input, and reduces the need for the user to provide excessive clarification input. The context may include any available information that the secretary can not substitute for clear user input to limit information processing problems and / or to customize the results. The context may be used to limit the solution during processing of various aspects including, for example, speech recognition, natural language processing, task flow processing, and dialog generation.

Description

BACKGROUND OF THE INVENTION 1. Field of the Invention < RTI ID = 0.0 > [0001] < / RTI > A method of using context information to facilitate command processing in a virtual assistant,

관련 출원들에 대한 상호 참조Cross reference to related applications

본 출원은, 2009년 6월 5일에 "Contextual Voice Commands" 라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 실용 출원 제 12/479,477 호(대리인 문서 번호 P7393US1)의 부분 계속 출원으로서의 우선권을 주장한다.This application claims the benefit of US Provisional Application No. 12 / 479,477 (Attorney Docket No. P7393US1) filed on June 5, 2009, entitled " Contextual Voice Commands ", the entire disclosure of which is incorporated herein by reference. Claim priority as an application.

본 출원은, 2011년 1월 10일에 "Intelligent Automated Assistant" 라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 실용 출원 제 12/987,982 호(대리인 문서 번호 P10575US1)의 부분 계속 출원으로서의 우선권을 더 주장한다.This application claims the benefit of US Provisional Application No. 12 / 987,982 (Attorney Docket No. P10575US1) filed on January 10, 2011, entitled " Intelligent Automated Assistant ", the entire disclosure of which is incorporated herein by reference. Further claim priority as an application.

미국 실용 출원 제 12/987,982 호는, 2010년 1월 18일에 "Intelligent Automated Assistant" 라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 가특허 출원 제 61/295,774 호(대리인 문서 번호 SIRIP003P)로부터의 우선권을 주장한다.U.S. Provisional Patent Application No. 61 / 295,774, filed on January 18, 2010, entitled " Intelligent Automated Assistant ", which is incorporated herein by reference in its entirety, Document No. SIRIP003P).

본 출원은, 2011년 6월 3일에 "Generating and Processing Data Items That Represent Tasks to Perform" 이라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 가출원 제 61/493,201 호(대리인 문서 번호 P11337P1)로부터의 우선권을 더 주장한다.This application is a continuation-in-part of US Provisional Application No. 61 / 493,201, filed on June 3, 2011, entitled " Generating and Processing Data Items That Represent Tasks to Perform, " No. P11337P1).

본 출원은, 본 출원과 동일자에 "Generating and Processing Task Items that Represent Tasks to Perform" 이라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 실용 출원 제 / 호(대리인 문서 번호 P11337US1)와 관련된 것이다.This application is a continuation-in-part of United States Provisional Patent Application, filed with the present application entitled "Generating and Processing Task Items that Represent Tasks to Perform ", the entire disclosure of which is incorporated herein by reference, / (Attorney Docket No. P11337US1).

본 출원은, 본 출원과 동일자에 "Automatically Adapting User Interfaces for Hands-Free Interaction" 이라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 실용 출원 제 / 호(대리인 문서 번호 P11357US1)와 관련된 것이다.This application is a continuation-in-part of U. S. Provisional Patent Application, entitled " Automatically Adapting User Interfaces for Hands-Free Interaction, " filed on even date herewith and incorporated herein by reference in its entirety, / (Attorney Docket No. P11357US1).

본 발명은 가상 비서(virtual assistant)에 관한 것으로서, 보다 구체적으로는, 그러한 비서에 제공된 커맨드들의 해석 및 처리를 개선시키기 위한 메카니즘들에 관한 것이다.The present invention relates to a virtual assistant, and more particularly, to mechanisms for improving the interpretation and processing of commands provided to such a secretary.

오늘날의 전자 디바이스들은, 인터넷을 통해서 및 다른 소스들로부터, 크고, 성장하는, 그리고 다양한 양의 기능들, 서비스들 및 정보에 액세스할 수 있다. 그러한 디바이스들의 기능은, 많은 소비자 디바이스, 스마트폰, 태블릿 컴퓨터 등이 다양한 태스크들을 수행하기 위한 소프트웨어 애플리케이션들을 실행시킬 수 있으며, 상이한 유형의 정보를 제공할 수 있게 됨에 따라, 빠르게 증가하고 있다. 때때로, 각각의 애플리케이션, 기능, 웹사이트 또는 특징은, 그 자신의 사용자 인터페이스 및 그 자신의 동작 패러다임들을 가지며, 그 중 많은 것이 배우기 부담스럽거나 또는 사용자들을 압도하는 것이다. 또한, 많은 사용자들은 그들의 전자 디바이스들 또는 다양한 웹사이트들에 대해 어떤 기능 및/또는 정보가 이용가능한지를 찾는 것조차도 어려울 수 있으므로, 그러한 사용자들은 좌절감을 느끼거나 압도당할 수 있으며, 또는 단순히 그들에게 이용가능한 자원들을 효율적인 방식으로 이용하지 못할 수 있다.Today's electronic devices are capable of accessing large, growing, and diverse amounts of functions, services and information over the Internet and from other sources. The functionality of such devices is rapidly increasing as many consumer devices, smartphones, tablet computers, etc. can run software applications to perform various tasks and provide different types of information. Sometimes, each application, function, website, or feature has its own user interface and its own operating paradigms, many of which are either too burdensome to learn or overwhelm users. In addition, many users may find it difficult to find out what functions and / or information are available to their electronic devices or various web sites, and therefore such users may feel frustrated or overwhelmed, or simply use them The available resources may not be available in an efficient manner.

특히, 초보 사용자들, 또는 정상적으로 기능하지 못하거나 장애를 갖고/갖거나, 나이가 많고, 바쁘고, 집중을 못하고/못하거나, 차량을 동작중인 개인들은 그들의 전자 디바이스들과 효율적으로 인터페이싱하고/하거나, 온라인 서비스들에 효율적으로 참여하는데 어려움을 가질 수 있다. 특히, 그러한 사용자들은 그들의 사용을 위해 이용가능할 수 있는 다수의 다양하고 일관성이 없는 기능들, 애플리케이션들 및 웹사이트들에 어려움을 가질 것이다.In particular, novice users, or individuals who are not functioning normally, have / have a disability, are old, busy, unable to concentrate / fail, or are in the running of a vehicle can efficiently interfere with their electronic devices and / It can be difficult to participate efficiently in online services. In particular, such users will have difficulties with a large number of diverse and inconsistent functions, applications and websites that may be available for their use.

따라서, 현존하는 시스템들은 때때로 사용 및 네비게이션이 어렵고, 때로는 사용자들에게 일관성이 없고 압도적인 인터페이스들을 제공함으로써, 사용자들이 기술을 효율적으로 이용하는 것을 방해한다.Thus, existing systems sometimes have difficulty in use and navigation, and sometimes provide inconsistent and overwhelming interfaces to users, hindering users from using technology efficiently.

본 명세서에서 가상 비서 라고도 지칭되는 지능형 자동 비서(intelligent automated assistant)가 사람과 컴퓨터 간에 개선된 인터페이스를 제공할 수 있다. 2011년 1월 10일에 "Intelligent Automated Assistant" 라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 관련 미국 실용 출원 제 12/987,982 호(대리인 문서 번호 P10575US1)에 기술된 바와 같이 구현될 수 있는 그러한 비서는, 사용자들이 말로 표현된 및/또는 텍스트 형태의 자연어를 이용하여 디바이스 또는 시스템과 상호작용할 수 있도록 한다. 그러한 비서는 사용자 입력들을 해석하고, 사용자의 의도를 태스크들 및 그러한 태스크들에 대한 파라미터들로 조작할 수 있게 하고, 그러한 태스크들을 지원하기 위한 서비스들을 실행시키고, 사용자가 이해할 수 있는 출력을 생성한다.An intelligent automated assistant, also referred to herein as a virtual secretary, can provide an improved interface between a person and a computer. (Attorney Docket No. P10575US1) filed on January 10, 2011, entitled " Intelligent Automated Assistant ", the entire disclosure of which is incorporated herein by reference. Such a secretary, which may be, allows users to interact with a device or system using a natural language in the form of words and / or text. Such a secretary interprets user inputs, enables the user to manipulate the intent of the user with the tasks and parameters for those tasks, executes services to support those tasks, and generates output that the user can understand .

가상 비서는, 예를 들면, 지식 기반(knowledge base)들, 모델들 및/또는 데이터를 포함하는 사용자 입력을 처리하기 위해, 다수의 정보 소스들 중 임의의 것을 이끌어 낼 수 있다. 많은 경우에 있어서, 사용자의 입력 하나만으로는 사용자의 의도 및 수행될 태스크를 명확하게 정의하기에 충분하지 않다. 이것은 입력 스트림, 사용자들 사이의 개인적인 차이 및/또는 자연어 고유의 모호성 때문일 수 있다. 예를 들어, 전화 상에서의 텍스트 메시징 애플리케이션의 사용자는 가상 비서를 호출(invoke)하여, "그녀에게 전화하라(call her)"는 커맨드를 말할 수 있다. 그러한 커맨드는 완전하게 합리적인 영어이지만, 이러한 요청에 대해 많은 해석들 및 가능한 해결책들이 존재할 수 있기 때문에, 그것은 정확하고, 실행가능한 표현이 아니다. 따라서, 추가의 정보를 갖지 않고서, 가상 비서는 그러한 입력을 올바르게 해석 및 처리하지 못할 수도 있다. 이러한 유형의 모호성은 에러들, 부정확한 동작들의 수행 및/또는 입력을 명확하게 하라는 요청들로 사용자에게 과도한 부담을 주는 것을 초래할 수 있다.The virtual assistant can derive any of a number of information sources, for example, to process user input including knowledge bases, models and / or data. In many cases, the user's input alone is not sufficient to clearly define the user's intent and the task to be performed. This may be due to the input stream, individual differences between users, and / or natural language ambiguity. For example, a user of a text messaging application on a telephone may invoke a virtual assistant to say a command "call her ". Such a command is perfectly reasonable in English, but it is not an accurate, executable expression because there can be many interpretations and possible solutions to such a request. Thus, without having additional information, the virtual assistant may not be able to correctly interpret and process such input. This type of ambiguity can result in errors, incomplete operations, and / or excessive burden on the user with requests to clarify input.

본 발명의 다양한 실시예에 따르면, 가상 비서는 사용자로부터의 자연어 또는 제스처(gestural) 입력을 보충하기 위해, (본 명세서에서 "컨텍스트(context)" 이라고도 지칭되는) 컨텍스트 정보를 이용한다. 이것은 사용자의 의도를 명확하게 하고, 사용자 입력의 후보 해석들의 수를 감소시키는데 도움을 주며, 사용자가 과도한 해명(clarification) 입력을 제공할 필요성을 감소시킨다. 컨텍스트는 정보-처리 문제를 제한하고/하거나 결과들을 개인화하기 위해 명시적인 사용자 입력을 보충하도록 비서에 의해 이용될 수 있는 임의의 이용가능한 정보를 포함할 수 있다. 예를 들어, 사용자로부터의 입력이 ("그녀에게 전화하라(call her)는 커맨드"에서의 "그녀(her)"와 같은) 대명사를 포함한다면, 가상 비서는 컨텍스트를 이용하여 그러한 대명사의 지시 대상을 추론, 예를 들면, 전화를 받게 될 개인의 아이덴티티(identity) 및/또는 이용할 전화 번호를 확인할 수 있다. 컨텍스트의 다른 이용들이 본 명세서에서 기술된다.According to various embodiments of the present invention, the virtual assistant uses contextual information (also referred to herein as a " context ") to supplement natural language or gestural input from the user. This helps clarify the user's intention, helps reduce the number of candidate interpretations of user input, and reduces the need for the user to provide excessive clarification input. The context may include any available information that can be used by a secretary to supplement explicit user input to limit information-processing problems and / or personalize results. For example, if an input from a user includes a pronoun (such as " her "in a " call her" command), the virtual assistant uses the context to indicate that pronoun For example, the identity of the individual to receive the call and / or the telephone number to use. Other uses of the context are described herein.

본 발명의 다양한 실시예에 따르면, 전자 디바이스 상에서 구현된 가상 비서에서의 계산(computation)들을 수행하기 위한 컨텍스트 정보를 획득 및 적용하기 위해 임의의 수의 메카니즘들이 구현될 수 있다. 다양한 실시예에서, 가상 비서는 2011년 1월 10일에 "Intelligent Automated Assistant" 라는 제목으로 출원되었으며 그 전체 개시 내용이 본 명세서에서 참조로 인용되는 미국 실용 출원 제 12/987,982 호(대리인 문서 번호 P10575US1)에 기술된 바와 같은 지능형 자동 비서다. 그러한 비서는 자연어 대화를 이용한 통합적인 대화의 방식으로 사용자와 관계를 형성하며, 정보를 얻거나 또는 다양한 동작들을 수행하는 것이 적절할 경우 외부의 서비스들을 호출한다. 본 명세서에 서술된 기술들에 따르면, 컨텍스트 정보가 그러한 비서에서 이용되어, 예를 들면, 음성 인식, 자연어 처리, 태스크 플로우 처리 및 대화 생성과 같은 정보 처리 기능들을 수행할 때의 모호성을 감소시킨다.According to various embodiments of the present invention, any number of mechanisms may be implemented to obtain and apply context information for performing computations in a virtual assistant implemented on an electronic device. In various embodiments, the virtual assistant was filed on January 10, 2011 under the heading "Intelligent Automated Assistant ", and the entire disclosure of which is hereby incorporated by reference in its entirety into U.S. Provisional Application No. 12 / 987,982 (Attorney Docket No. P10575US1 ) Is an intelligent automatic secretary as described in U.S. Pat. Such a secretary forms a relationship with the user in an integrated way of communicating using natural language conversation, and invokes external services when it is appropriate to obtain information or perform various actions. According to the techniques described herein, context information is used in such secretaries to reduce ambiguity in performing information processing functions such as speech recognition, natural language processing, task flow processing, and dialog generation.

본 발명의 다양한 실시예에 따르면, 가상 비서는 다양하고 상이한 종류의 동작들, 기능들 및/또는 피처(feature)들의 수행에 컨텍스트를 이용하고, 및/또는, 그것이 설치되는 전자 디바이스의 복수의 피처들, 동작들 및 애플리케이션들을 조합하도록 구성, 설계 및/또는 동작가능하게 될 수 있다. 일부 실시예에 있어서, 본 발명의 가상 비서는 사용자로부터 입력을 적극적으로 유도해 내는 것, 사용자의 의도를 해석하는 것, 경합 해석들의 차이를 명확히 하는 것, 필요에 따라 명확한 정보를 요구 및 수신하는 것 및/또는 파악한 의도에 기초하여 액션들을 수행(또는 개시)하는 것 중 어느 하나 또는 모두를 수행할 때에 컨텍스트를 이용할 수 있다.According to various embodiments of the present invention, the virtual assistant uses contexts to perform a variety of different types of operations, functions and / or features, and / or uses a plurality of features of the electronic device in which it is installed May be configured, designed, and / or operable to combine, manipulate, and operate applications, systems, operations, and applications. In some embodiments, the virtual assistant of the present invention may be used to actively derive input from a user, interpret user intent, clarify differences in contention analysis, request and receive clear information as needed And / or performing (or initiating) actions based on the identified intent.

예컨대, 전자 디바이스에서 이용할 수 있는 애플리케이션들 또는 서비스들뿐만 아니라, 인터넷과 같은 전자 네트워크를 통해 이용할 수 있는 서비스들을 활성화 및/또는 인터페이싱함으로써, 동작들을 수행할 수 있다. 다양한 실시예에 있어서, 그러한 외부 서비스들의 활성화는 애플리케이션 프로그래밍 인터페이스(API)들이나 어떤 다른 적합한 메커니즘에 의해 수행될 수 있다. 이와 같이, 본 발명의 다양한 실시예에 따라 구현된 가상 비서는 전자 디바이스의 많은 상이한 애플리케이션들 및 기능들에 대한, 그리고 인터넷을 통해 이용할 수 있는 서비스들에 대한, 사용자의 경험을 통합, 간소화 및 개선할 수 있다. 이로써, 사용자는 디바이스에서 그리고 웹 접속 서비스들에서 이용할 수 있는 기능이 무엇인지, 원하는 것을 얻기 위해 어떻게 그러한 서비스들과 인터페이싱할지, 그리고 어떻게 그러한 서비스들로부터 수신한 출력을 해석할지, 학습해야 하는 짐을 덜 수 있다. 오히려 본 발명의 가상 비서는 사용자와 그러한 다양한 서비스들 사이에서 중개 역할을 할 수 있다.For example, operations may be performed by activating and / or interfacing services and / or services available on an electronic network, such as the Internet, as well as applications or services available on an electronic device. In various embodiments, activation of such external services may be performed by application programming interfaces (APIs) or some other suitable mechanism. As such, the virtual assistant implemented in accordance with various embodiments of the present invention can be used to integrate, simplify, and improve the user experience for many different applications and functions of the electronic device, and for services available over the Internet can do. This allows the user to know what functions are available on the device and in web access services, how to interface with those services to get what they want, how to interpret the output received from those services, less burden to learn . Rather, the inventive virtual assistant can act as an intermediary between a user and such a variety of services.

또한, 다양한 실시예에 있어서, 본 발명의 가상 비서는 사용자가 종래의 그래픽 사용자 인터페이스에 비해 더 사용하기 쉽고 덜 부담스러운 것으로 여길 수 있는 대화식 인터페이스를 제공한다. 사용자는 예컨대 음성, 그래픽 사용자 인터페이스들(버튼들 및 링크들), 텍스트 기입 등과 같은, 다수의 이용가능한 입력 및 출력 메커니즘들 중 어느 하나를 이용하여 가상 비서와 일상 대화에서 쓰이는 대화의 형태로 관계를 맺을 수 있다. 시스템은 디바이스 API들, 웹, 이메일 등과 같이, 다수의 상이한 플랫폼들 중 어느 하나를 이용하여 구현될 수 있다. 추가의 입력에 대한 요구들은 그러한 대화의 컨텍스트로 사용자에게 제시될 수 있다. 정해진 세션 내의 이전 이벤트들 및 통신들뿐만 아니라 사용자에 관한 히스토리 및 프로파일 정보를 고려하여 적절한 컨텍스트로 사용자 입력을 해석할 수 있도록 단기 및 장기 메모리를 사용할 수 있다.In addition, in various embodiments, the virtual assistant of the present invention provides an interactive interface that a user may find more convenient and less burdensome than a conventional graphical user interface. A user may use any of a number of available input and output mechanisms, such as voice, graphical user interfaces (buttons and links), text entry, etc., to create a relationship in virtual conversation Can be concluded. The system may be implemented using any of a number of different platforms, such as device APIs, web, email, and the like. Requests for further input may be presented to the user in the context of such a conversation. Long-term memory can be used to interpret user input in an appropriate context, taking into account historical events and communications within a given session as well as history and profile information about the user.

또한, 다양한 실시예에 있어서, 디바이스 상의 피처, 동작 또는 애플리케이션과 사용자의 상호 작용으로부터 얻는 컨텍스트 정보는 그 디바이스 또는 다른 디바이스 상의 다른 피처들, 동작들 또는 애플리케이션들의 동작을 간소화하는 데에 이용될 수 있다. 예컨대, 가상 비서는 (전화한 사람과 같은) 전화 호출의 컨텍스트를 이용하여 텍스트 메시지의 개시를 간소화할 수 있다(예컨대, 사용자가 텍스트 메시지의 수신인을 명백히 지정하지 않아도 텍스트 메시지를 동일한 사람에게 보내야 한다고 결정할 수 있다). 이로써, 본 발명의 가상 비서는 "그에게 텍스트 메시지를 보내라"와 같은 커맨드들을 해석할 수 있고, 여기서 "그"는 현재 전화 호출로부터 및/또는 디바이스 상의 임의의 피처, 동작 또는 애플리케이션으로부터 얻는 컨텍스트 정보에 따라 해석된다. 다양한 실시예에 있어서, 가상 비서는 다양한 종류의 이용가능한 컨텍스트 데이터를 고려하여, 어떤 어드레스 북 연락처를 이용할지, 어떤 연락처 데이터를 이용할지, 연락처에 어떤 전화 번호를 이용할지 등을 결정함으로써, 사용자가 그러한 정보를 수동으로 재지정할 필요가 없게 한다.Further, in various embodiments, context information from a feature, an action, or an interaction of an application with a user on a device may be used to simplify the operation of other features, operations, or applications on the device or other device . For example, a virtual assistant can simplify the initiation of a text message using the context of a telephone call (such as the person making the call) (e.g., the user must send a text message to the same person without explicitly specifying the recipient of the text message Can be determined). Thus, the virtual assistant of the present invention can interpret commands such as "Send him a text message ", where" that " includes context information from the current telephone call and / or from any feature, . In various embodiments, the virtual assistant considers the various types of available context data to determine which address book contacts to use, which contact data to use, which phone number to use with the contact, Thereby eliminating the need to manually reassign such information.

컨텍스트 정보 소스는 현재 시간, 위치, 애플리케이션 또는 데이터 객체와 같이, 가상 비서에 대한 인터페이스로서 이용되는 디바이스의 현재 상태; 사용자의 어드레스 북, 캘린더 및 애플리케이션 사용 히스토리와 같은 개인 데이터; 및 최근 언급된 사람 및/또는 장소와 같이, 사용자와 가상 비서 간의 대화의 상태를 포함하며, 이것은 예시적인 것이지 제한적인 것이 아니다.The context information source may be the current state of the device used as an interface to the virtual secretary, such as the current time, location, application or data object; Personal data such as a user's address book, calendar, and application usage history; And the state of the conversation between the user and the virtual assistant, such as the recently mentioned person and / or place, and this is illustrative and not limiting.

가상 비서의 동작에서 다양한 계산 및 추론에 컨텍스트를 적용할 수 있다. 예컨대, 사용자 입력을 처리할 때 애매함을 줄이거나 솔루션들의 수를 제약하는 데에 컨텍스트를 이용할 수 있다. 따라서, 다양한 처리 단계 동안 솔루션들을 제약하는 데에 컨텍스트를 이용할 수 있으며, 이것은 예시적인 것이지 제한적인 것이 아니다.The context can be applied to various calculations and inferences in the operation of the virtual assistant. For example, contexts can be used to reduce ambiguity or limit the number of solutions when processing user input. Thus, contexts can be used to constrain solutions during various processing steps, and this is illustrative and not limiting.

ㆍ음성 인식 - 보이스 입력을 수신하고 후보 해석을 텍스트로, 예컨대 "call her", "collar" 및 "call Herb"로 생성한다. 음성 인식 모듈에 의해 어떤 단어들 및 어구들을 고려할지, 어떻게 그것들을 서열화할지, 고려시 임계치 위로서 어떤 것을 수락할지를 제약하는 데에 컨텍스트를 이용할 수 있다. 예컨대, 사용자의 어드레스 북은 음성의 다른 언어-일반 모델에 개인 이름들을 추가하여, 이들 이름들이 인식되고 우선순위를 가질 수 있게 한다. Voice recognition - Receives voice input and generates candidate interpretations as text, for example, "call her", "collar" and "call Herb". Contexts can be used to constrain which words and phrases are to be considered by the speech recognition module, how to order them, and what to accept as thresholds in consideration. For example, a user's address book may add personal names to other language-generic models of speech, allowing these names to be recognized and prioritized.

ㆍ자연어 처리( NLP ) - 텍스트를 파싱하고 단어들을 구문론적 및 의미론적 규칙들과 연관시키는데, 예컨대 사용자 입력이 대명사 "그녀"가 나타내는 사람에게 전화 호출을 행하는 것에 관한 것임을 결정하고, 이 사람에 대한 특정 데이터 표현을 찾는 것이다. 예컨대, 텍스트 메시징 애플리케이션의 컨텍스트는 "그녀"의 해석을 "내가 텍스트로 대화하고 있는 사람"을 의미하는 것으로 제약하는 것을 도울 수 있다. Natural Language Processing ( NLP ) - parses text and associates words with syntactic and semantic rules, such as determining that the user input is about making a telephone call to the person the pronoun "she" is about, It is looking for a particular data representation. For example, the context of a text messaging application may help constrain the interpretation of "her " to mean" the person who is speaking with text. &Quot;

ㆍ태스크 플로우 처리 - 사용자 태스크, 태스크 단계들, 및 그 태스크를 돕는 데에 이용되는 태스크 파라미터들, 예컨대 "그녀"가 나타내는 사람에 대해 어떤 전화 번호를 이용할지를 식별한다. 또한, 텍스트 메시징 애플리케이션의 컨텍스트는 시스템이 텍스트 메시징 대화에 현재 또는 최근 이용한 번호를 이용해야 한다는 것을 나타내도록 전화 번호의 해석을 제약할 수 있다. Task flow processing - Identifies which telephone number to use for the user task, the task steps, and the task parameters used to assist the task, such as the person she represents. In addition, the context of the text messaging application may constrain the interpretation of the telephone number to indicate that the system should use a current or recently used number in a text messaging conversation.

ㆍ대화 생성 - 그들 태스크에 관하여 사용자와의 대화의 일부로서 비서 응답들을 생성하는데, 예컨대 사용자의 의도를 "응, 내가 레베카에게 그녀의 모바일에 전화할께..."라는 응답으로 바꾸어 표현한다(paraphrase). 다변(verbosity) 및 격식없는 어조(informal tone)의 레벨은 컨텍스트 정보에 의해 안내될 수 있는 선택권이다. Generate Dialog - Create secretary responses as part of a conversation with the user about their tasks, for example by changing the user's intent in response to "I'll call Rebecca to her mobile ..." (paraphrase ). The level of verbosity and informal tone is an option that can be guided by contextual information.

다양한 실시예에 있어서, 본 발명의 가상 비서는 전자 디바이스의 다양한 피처들 및 동작들을 제어할 수 있다. 예컨대, 가상 비서는 API들이나 다른 수단에 의해 디바이스 상의 기능 및 애플리케이션과 인터페이싱하는 서비스들을 호출하여, 그 디바이스 상의 종래의 사용자 인터페이스를 이용하여 개시되어야 했던 기능들 및 동작들을 수행할 수 있다. 그러한 기능들 및 동작들은 예컨대, 알람 설정, 전화 걸기, 텍스트 메시지 또는 이메일 메시지 전송, 캘린더 이벤트 추가 등을 포함할 수 있다. 그러한 기능들 및 동작들은 사용자와 가상 비서 간의 일상 대화에서 쓰이는 대화의 컨텍스트에서 부가 기능들로서 수행될 수 있다. 그러한 기능들 및 동작들은 그러한 대화의 컨텍스트에서 사용자에 의해 지정되거나, 그 대화의 컨텍스트에 기초하여 자동으로 수행될 수 있다. 당업자라면, 이로써 가상 비서를 전자 디바이스 상의 다양한 동작들을 개시 및 제어하는 제어 메커니즘으로서 이용할 수 있고, 따라서 버튼 또는 그래픽 사용자 인터페이스와 같은 종래의 메커니즘의 대안으로서 이용할 수 있다는 것을 인식할 것이다. 본 명세서에 기재한 바와 같이, 그러한 가상 비서의 제어 메커니즘으로서의 이용을 알리고 개선하는 데에 컨텍스트 정보를 이용할 수 있다.In various embodiments, the virtual assistant of the present invention may control various features and operations of the electronic device. For example, the virtual assistant can invoke services and functions that interface with the functions and applications on the device by APIs or other means to perform functions and operations that had to be initiated using a conventional user interface on the device. Such functions and actions may include, for example, setting an alarm, dialing, sending a text or email message, adding calendar events, and the like. Such functions and actions may be performed as add-ons in the context of the conversation used in the daily conversation between the user and the virtual assistant. Such functions and actions may be specified by the user in the context of such conversations, or may be performed automatically based on the context of the conversation. Those skilled in the art will appreciate that the virtual assistant can be used as a control mechanism to initiate and control various operations on an electronic device and thus can be used as an alternative to conventional mechanisms such as buttons or a graphical user interface. As described herein, context information can be used to inform and improve the use of such virtual assistant as a control mechanism.

첨부 도면은 본 발명의 여러 실시예들을 도시하며, 설명과 함께, 실시예들에 따라 본 발명의 원리를 설명하는 기능을 한다.　 본 기술 분야의 당업자는 도면에 도시된 특정 실시예들이 단지 예시적이고 본 발명의 범위를 제한하도록 의도되지 않음을 인식할 것이다.
도 1은 일 실시예에 따라 가상 비서 및 그 동작에 영향을 줄 수 있는 컨텍스트의 소스들의 일부 예들을 도시하는 블록도.
도 2는 일 실시예에 따라, 가상 비서에서의 다양한 처리 스테이지들에서 컨텍스트를 이용하는 방법을 도시하는 흐름도.
도 3은 일 실시예에 따라, 음성 도출 및 해석에서 컨텍스트를 이용하는 방법을 도시하는 흐름도.
도 4는 일 실시예에 따라, 자연어 처리에서 컨텍스트를 이용하는 방법을 도시하는 흐름도.
도 5는 일 실시예에 따라, 태스크 플로우 처리에서 컨텍스트를 이용하는 방법을 도시하는 흐름도.
도 6은 일 실시예에 따라, 클라이언트와 서버 사이에서 분산되는 컨텍스트의 소스들의 예를 도시하는 블록도.
도 7a 내지 도 7d는 다양한 실시예들에 따라 컨텍스트 정보를 획득하고 조정하기 위한 메카니즘들의 예들을 도시하는 이벤트 다이어그램.
도 8a 내지 도 8d는 본 발명의 다양한 실시예들과 관련하여 이용될 수 있는 컨텍스트 정보의 다양한 표현의 예를 도시하는 도면.
도 9는 일 실시예에 따라, 다양한 컨텍스트 정보 소스에 대한 정책들을 캐싱하고 통신을 지정하는 설정 테이블의 예를 도시하는 도면.
도 10은 일 실시예에 따라, 상호작용 시퀀스의 처리 동안 도 9에 설성된 컨텍스트 정보 소스들에 액세스하는 예를 도시하는 이벤트 다이어그램.
도 11 내지 도 13은 일 실시예에 따라, 대명사에 대한 지시 대상을 도출하기 위해 텍스트 메시징 도메인에서 애플리케이션 컨텍스트를 이용하는 예를 도시하는 일련의 스크린샷.
도 14는 일 실시예에 따라, 이름 명확화(name disambiguation)를 프롬프트하는 가상 비서를 도시하는 스크린샷.
도 15는 일 실시예에 따라, 커맨드에 대한 위치를 추론하기 위해 대화 컨텍스트를 이용하는 가상 비서를 도시하는 스크린샷.
도 16은 일 실시예에 따라, 컨텍스트의 소스로서 전화 선호도 리스트를 이용하는 예를 도시하는 스크린샷.
도 17은 내지 도 20은 일 실시예에 따라, 커맨드를 해석하고 동작화하기 위해 현재 애플리케이션 컨텍스트를 이용하는 예를 도시하는 일련의 스크린 샷들.
도 21은 상이한 애플리케이션을 호출하는 커맨드를 해석하기 위해 현재 애플리케이션 컨텍스트를 이용하는 예를 도시하는 스크린샷.
도 22 내지 도 24는 일 실시예에 따라, 착신 텍스트 메시지의 형태로 이벤트 컨텍스트를 이용하는 예를 도시하는 일련의 스크린샷.
도 25a 및 도 25b는 일 실시예에 따라, 이전 대화 텍스트를 이용하는 예를 도시하는 일련의 스크린 샷.
도 26a 및 도 26b는 일 실시예에 따라, 후보 해석들 중에서 하나를 선택하기 위한 사용자 인터페이스의 예를 도시하는 스크린 샷들.
도 27은 가상 비서 시스템의 일 실시예의 예를 도시하는 블록도.
도 28은 적어도 일 실시예에 따라 가상 비서의 적어도 일부를 구현하기에 적합한 컴퓨팅 디바이스를 도시하는 블록도.
도 29는 적어도 일 실시예에 따라, 독립형 컴퓨팅 시스템 상의 가상 비서의 적어도 일부를 구현하기 위한 아키텍쳐를 도시하는 블록도.
도 30은 적어도 일 실시예에 따라, 분산형 컴퓨팅 네트워크 상의 가상 비서의 적어도 일부를 구현하기 위한 아키텍쳐를 도시하는 블록도.
도 31은 여러 상이한 유형의 클라이언트들 및 동작 모드들을 도시하는 시스템 아키텍쳐를 도시하는 블록도.
도 32는 일 실시예에 따라 본 발명을 구현하기 위해 서로 통신하는 클라이언트와 서버를 도시하는 블록도.The accompanying drawings illustrate several embodiments of the invention and, together with the description, serve to explain the principles of the invention in accordance with the embodiments. Those skilled in the art will recognize that the specific embodiments shown in the figures are illustrative only and are not intended to limit the scope of the invention.
1 is a block diagram illustrating some examples of sources of context that may affect a virtual assistant and its operation in accordance with one embodiment;
2 is a flow diagram illustrating a method for using a context in various processing stages in a virtual assistant, in accordance with one embodiment;
3 is a flow diagram illustrating a method for using context in speech derivation and interpretation, in accordance with one embodiment;
4 is a flow diagram illustrating a method for using context in natural language processing, according to one embodiment.
5 is a flow diagram illustrating a method for using a context in task flow processing, in accordance with one embodiment;
6 is a block diagram illustrating examples of sources of context that are distributed between a client and a server, according to one embodiment.
7A-7D are event diagrams illustrating examples of mechanisms for obtaining and adjusting context information in accordance with various embodiments.
Figures 8A-8D illustrate examples of various representations of contextual information that may be used in connection with various embodiments of the present invention.
9 illustrates an example of a settings table for caching policies and specifying communications for various context information sources, according to one embodiment;
10 is an event diagram illustrating an example of accessing the context information sources depicted in FIG. 9 during processing of an interaction sequence, in accordance with one embodiment;
11-13 are a series of screenshots illustrating an example of using an application context in a text messaging domain to derive an indicative target for a pronoun, in accordance with one embodiment.
Figure 14 is a screen shot illustrating a virtual secretary prompting for name disambiguation, in accordance with one embodiment.
15 is a screen shot illustrating a virtual secretary using a conversation context to infer a location for a command, in accordance with one embodiment.
16 is a screen shot illustrating an example of using a telephone preference list as a source of context, according to one embodiment;
17 through 20 are a series of screen shots illustrating an example of utilizing the current application context to interpret and operate the command, according to one embodiment.
21 is a screen shot illustrating an example of utilizing the current application context to interpret a command that invokes a different application.
Figures 22-24 are a series of screen shots illustrating an example of using an event context in the form of an incoming text message, according to one embodiment.
25A and 25B are a series of screen shots illustrating an example using previous conversation text, according to one embodiment.
26A and 26B are screen shots illustrating an example of a user interface for selecting one of the candidate interpretations, according to one embodiment.
27 is a block diagram showing an example of an embodiment of a virtual secretarial system;
28 is a block diagram illustrating a computing device suitable for implementing at least a portion of a virtual secretary in accordance with at least one embodiment;
29 is a block diagram illustrating an architecture for implementing at least a portion of a virtual assistant on a standalone computing system, in accordance with at least one embodiment.
30 is a block diagram illustrating an architecture for implementing at least a portion of a virtual assistant on a distributed computing network, according to at least one embodiment.
31 is a block diagram illustrating a system architecture illustrating various different types of clients and modes of operation.
32 is a block diagram illustrating a client and server communicating with one another to implement the invention in accordance with one embodiment.

본 발명의 각종 실시예들에 따라, 가상 비서의 동작들의 지원시 정보 처리 기능들을 수행하기 위해 다양한 컨텍스트 정보가 획득되고 적용된다.　 설명을 위해, "가상 비서"라는 용어는 "지능형 자동화 비서"라는 용어와 동등하고, 둘 다는 아래의 기능들 중 하나 이상을 수행하는 임의의 정보 처리 시스템을 지칭한다.In accordance with various embodiments of the present invention, various context information is obtained and applied to perform information processing functions in support of operations of the virtual assistant. For purposes of illustration, the term "virtual secretary" is equivalent to the term "intelligent automation secretary ", and both refer to any information handling system that performs one or more of the following functions.

음성(spoken) 또는 텍스트 형태로, 인간 언어 입력을 해석

Interpret human language input in spoken or text form

단계들 및/또는 파라미터들로 태스크를 표현하는 것과 같은, 실행될 수 있는 형태로 사용자 의도를 표현하는 것의 동작화.

Activation of expressing user intent in an executable form, such as expressing a task with steps and / or parameters.

프로그램들, 메소드들, 서비스들, API들 등을 호출함으로써 태스크 표현들을 실행; 및

Executing task expressions by calling programs, methods, services, APIs, etc.; And

언어 및/또는 그래픽 형태로 사용자에 대한 출력 응답들을 생성.

Generate output responses for users in language and / or graphical form.

이러한 가상 비서의 예는 2011년 1월 10에 출원된 "지능형 자동화 비서"(참조 번호 P10575US1)라는 제목의 관련 미국 실용신안 제12/987,982호에 개시되어 있고, 그 전체 개시물은 참조로 본원에 포함된다.An example of such a virtual assistant is disclosed in related U.S. Utility Model 12 / 987,982 entitled " Intelligent Automation Assistant "(reference P10575US1) filed on January 10, 2011, the entire disclosure of which is incorporated herein by reference .

이제, 첨부 도면에 도시된 바와 같은 실시예들을 참조하여 각종 기술들이 상세히 설명될 것이다.　 이하의 설명에서는, 본원에서 기술된 하나 이상의 양태들 및/또는 특징들 또는 참조내용의 전반적인 이해를 제공하기 위해 다수의 특정 상세들이 개시된다.　 그러나, 당업자들에게, 본원에서 기술된 하나 이상의 양태들 및/또는 특징들 또는 참조내용은 이들 특정 상세들의 전부 또는 일부 없이도 실시될 수 있음이 자명할 것이다.　 다른 예들에서, 본원에서 기술된 양태들 및/또는 특징들 또는 참조내용의 일부가 모호해지지 않도록 공지의 프로세스 단계 및/또는 구조는 상세히 기술되지 않았다.Various techniques will now be described in detail with reference to the embodiments as shown in the accompanying drawings. In the following description, numerous specific details are set forth in order to provide a thorough understanding of one or more aspects and / or features described herein or the content of the references. However, it will be apparent to one skilled in the art that one or more aspects and / or features or references described herein may be practiced without some or all of these specific details. In other instances, well-known process steps and / or structures have not been described in detail so as not to obscure aspects and / or features or parts of the references described herein.

본 출원서에서 하나 이상의 상이한 발명들이 설명될 수 있다. 또한, 본원에서 설명된 하나 이상의 발명(들)에 대해 다양한 실시예들이 본 특허 출원서에서 설명될 수 있고, 오직 예시적인 목적으로 제시된다. 설명된 실시예들은 어떤 의미로든 한정하도록 의도되지 않는다. 본 발명(들) 중 하나 이상은, 개시 내용으로부터 명백하듯이, 다양한 실시예들에 널리 적용될 수 있다. 이러한 실시예들은 본 기술분야의 당업자가 본 발명(들) 중 하나 이상을 실시하는 것이 가능하도록 충분히 상세히 설명되며, 그외의 실시예들이 이용될 수 있고, 구조적, 논리적, 소프트웨어, 전기적 및 그외의 변경들이 본 발명(들)의 범주를 벗어나지 않으면서 이루어질 수 있다는 것이 이해될 것이다. 따라서, 본 기술분야의 당업자는 본 발명(들) 중 하나 이상이 다양한 변형들 및 대안들로 실시될 수 있음을 인지할 것이다. 본 발명(들) 중 하나 이상의 특정 피처들은 하나 이상의 특정 실시예들 또는 본 개시내용의 일부를 형성하며, 본 발명(들) 중 하나 이상의 특정 실시예들이 예시의 방법으로써 도시되는 도면들을 참조하여 설명될 수 있다. 그러나, 그러한 피처들은 그들이 참조하여 설명되는 하나 이상의 특정 실시예들 또는 도면들에서의 사용으로 한정되는 것은 아님을 이해해야 한다. 본 개시내용은 본 발명(들) 중 하나 이상의 모든 실시예들의 문자적 기술 또는 모든 실시예들에 존재해야 하는 본 발명(들) 중 하나 이상의 피처들의 나열 어느 것도 아니다.In the present application, one or more different inventions may be described. In addition, various embodiments for one or more of the inventions (s) described herein may be described in this patent application and are presented for illustrative purposes only. The described embodiments are not intended to be limiting in any sense. One or more of the present invention (s) can be widely applied to various embodiments, as will be apparent from the disclosure. These embodiments are described in sufficient detail to enable those skilled in the art to practice one or more of the invention (s), and other embodiments may be utilized, and structural, logical, software, electrical and other changes (S) can be made without departing from the scope of the present invention (s). Accordingly, those skilled in the art will recognize that one or more of the inventions may be practiced with various modifications and alternatives. It will be appreciated that one or more of the specific features of the present invention may form one or more specific embodiments or portions of the disclosure, and that one or more specific embodiments of the invention (s) may be described with reference to the drawings, . It is to be understood, however, that such features are not limited to use with one or more specific embodiments or drawings described in the context of the reference. This disclosure is not intended to be a literal description of all or any one or more of the embodiments of the present invention or to list one or more features of the invention (s) that must be present in all embodiments.

본 특허 출원서에 제공되는 섹션들의 주제(heading)들 및 본 특허 출원서의 제목은 편의를 위한 것일 뿐, 임의의 방식으로 본 개시내용을 한정하는 것으로 받아들여서는 아니된다.The headings of the sections provided in this patent application and the title of the present patent application are for convenience only and are not to be construed as limiting the present disclosure in any way.

서로 통신하는 디바이스들은, 명백히 달리 특정되지 않는 한, 서로 연속적인 통신을 할 필요는 없다. 또한, 서로 통신하는 디바이스들은 하나 이상의 중계물들(intermediaries)을 통해 직접적으로 또는 간접적으로 통신할 수 있다.Devices that communicate with each other do not need to be in continuous communication with each other unless explicitly specified otherwise. Also, devices that communicate with each other can communicate directly or indirectly through one or more intermediaries.

서로 통신하는 몇몇 컴포넌트들을 이용하는 실시예들의 설명이 모든 그러한 컴포넌트들이 요구된다는 것을 의미하지는 않는다. 반대로, 다양한 선택적 컴포넌트들이 설명되어 본 발명(들) 중 하나 이상의 광범위한 가능한 실시예들을 예시한다.The description of embodiments utilizing several components that communicate with each other does not imply that all such components are required. Conversely, various optional components are described to illustrate one or more of the broadest possible embodiments of the present invention (s).

또한, 프로세스 단계들, 방법 단계들, 알고리즘 등이 순차적인 순서로 설명될 수 있으나, 그러한 프로세스들, 메소드들 및 알고리즘들은 임의의 적절한 순서로 동작하도록 구성될 수 있다. 다시 말해서, 본 특허 출원서에서 설명될 수 있는 단계들의 임의의 시퀀스 또는 순서는, 그 자체로는, 그 단계들이 그 순서로 수행되어야 한다는 요건을 나타내는 것은 아니다. 또한, (예를 들어, 하나의 단계가 다른 단계 후에 설명되기 때문에) 비동시적으로 발생하는 것으로 설명되거나 시사되었지만, 일부 단계들은 동시에 수행될 수 있다. 또한, 도면에서의 그것의 묘사에 의한 프로세스의 예시는 예시된 프로세스가 그에 대한 다른 변형들 및 변경들을 배제한다는 것을 의미하지 않으며, 예시된 프로세스 또는 그것의 단계들 중 임의의 단계가 본 발명(들) 중 하나 이상에 필요하다는 것을 의미하지 않으며, 예시된 프로세스가 바람직하다는 것을 의미하지는 않는다.Also, process steps, method steps, algorithms, and the like may be described in a sequential order, but such processes, methods, and algorithms may be configured to operate in any appropriate order. In other words, any sequence or sequence of steps that may be described in this patent application does not by itself indicate a requirement that the steps be performed in that order. Also, although some aspects are described or suggested as occurring asynchronously (e.g., because one step is described after another), some steps may be performed concurrently. Also, an example of a process by its description in the figures does not mean that the illustrated process excludes other variations and modifications thereto, and it is understood that any of the steps of the illustrated process, or steps thereof, Quot;), and does not imply that the illustrated process is preferred.

단일 디바이스 또는 물품이 설명되는 경우, 하나보다 많은 디바이스/물품(그들이 공조하든 아니든)은 단일 디바이스/물품을 대신하여 이용될 수 있다. 마찬가지로, 하나보다 많은 디바이스 또는 물품이 설명되는 경우(그들이 공조하든 아니든), 단일 디바이스/물품이 하나보다 많은 디바이스 또는 물품을 대신하여 이용될 수 있음이 명백할 것이다.Where a single device or article is described, more than one device / article (whether or not they cooperate) may be used in place of a single device / article. Likewise, it will be apparent that if more than one device or article is described (whether they cooperate or not), a single device / article may be used in place of more than one device or article.

디바이스의 기능성 및/또는 피처들은 그러한 기능성/피처들을 갖는 것으로 명시적으로 설명되지 않은 하나 이상의 다른 디바이스들에 의해 대안적으로 실시될 수있다. 따라서, 본 발명(들) 중 하나 이상의 다른 실시예들은 디바이스 자체를 필요로 하지 않는다.The functionality and / or features of a device may alternatively be implemented by one or more other devices that are not explicitly described as having such functionality / features. Accordingly, one or more other embodiments of the present invention do not require the device itself.

본원에서 설명되거나 참조된 기술들 및 메커니즘들은 때로는 명료함을 위해 하나의 형식(singular form)으로 설명될 것이다. 그러나, 특정 실시예들은 달리 언급되지 않는 한, 메커니즘의 다수의 예시화 또는 다수의 기법의 반복을 포함한다는 것을 유의해야 한다.The techniques and mechanisms described or referenced herein will sometimes be described in singular form for clarity. It should be noted, however, that certain embodiments, unless otherwise stated, include multiple instantiations of a mechanism or repetition of multiple techniques.

가상 비서(virtual assistant)라고도 알려진, 지능형 자동 비서를 구현하기 위한 기술의 맥락 내에서 설명되었으나, 본 명세서에서 설명된 다양한 양태들 및 기술들이 또한 채용되고/되거나, 인간 및/또는 소프트웨어와의 컴퓨터화된 상호작용을 수반하는 그외의 기술 분야에 적용될 수 있다.Although described in the context of techniques for implementing an intelligent automatic secretary, also known as a virtual assistant, various aspects and techniques described herein may also be employed and / or computerized with human and / or software Lt; RTI ID = 0.0 > interactions. &Lt; / RTI >

가상 비서 기술(예를 들어, 본 명세서에서 설명된 하나 이상의 가상 비서 시스템 실시예들에 의해 활용되고, 제공되고, 및/또는 구현된)이 이하의 하나 이상에서 개시되며, 그 전체 개시내용이 본 명세서에 참조로서 포함된다.Virtual secretarial techniques (e.g., utilized, provided, and / or implemented by one or more virtual secretarial system embodiments described herein) are disclosed in one or more of the following, Are incorporated herein by reference.

미국 실용 출원 번호 제12/987,982호, "Intelligent Automated Assistant", 대리인 정리 번호 P10575US1, 2011.1.10 제출.

U.S. Utility Application No. 12 / 987,982, "Intelligent Automated Assistant", Attorney Docket No. P10575US1, submitted on January 10, 2011.

미국 가특허 출원 번호 제61/295,774호, "Intelligent Automated Assistant", 대리인 정리 번호 SIRIP003P, 2010.1.18 제출.

U.S. Provisional Patent Application No. 61 / 295,774, "Intelligent Automated Assistant", Attorney Docket No. SIRIP003P, submitted on January 18, 2010.

미국 특허 출원 번호 제11/518,292호, "Method And Apparatus for Building an Intelligent Automated Assistant", 2006년 9월 8일 제출.

U.S. Patent Application Serial No. 11 / 518,292 entitled " Method And Apparatus for Building Intelligent Automated Assistant ", filed on September 8, 2006.

미국 가특허 출원 번호 제61/186,414호, "System and Method for Semantic Auto-Completion", 2009.6.12 제출.

U.S. Provisional Patent Application No. 61 / 186,414 entitled "System and Method for Semantic Auto-Completion"

하드웨어 아키텍처Hardware architecture

일반적으로, 본 명세서에서 개시된 가상 비서 기술들은 하드웨어 또는 소프트웨어와 하드웨어의 조합 상에서 구현될 수 있다. 예를 들어, 그것들은 특별하게 구성된 머신 상에서 및/또는 네트워크 인터페이스 카드 상에서 운영 체제 커널(operating system kernel)로, 개별적인 사용자 프로세스로, 네트워크 애플리케이션들에 바운드된 라이브러리로 구현될 수 있다. 특정 실시예에서, 본 명세서에서 개시된 기술들은 운영 체제 또는 운영 체제 상에서 실행하는 애플리케이션과 같은 소프트웨어로 구현될 수 있다.In general, the virtual secretarial techniques disclosed herein may be implemented on hardware or on a combination of software and hardware. For example, they may be implemented as a library bound to network applications, on a specially configured machine and / or as an operating system kernel on a network interface card, as a separate user process. In certain embodiments, the techniques disclosed herein may be implemented in software such as an operating system or an application running on an operating system.

본 명세서에서 개시된 가상 비서 실시예(들)의 적어도 일부의 소프트웨어/하드웨어 하이브리드 구현(들)은 메모리에 저장된 컴퓨터 프로그램에 의해 선택적으로 활성화되거나 재구성된 프로그램가능한 머신 상에서 구현될 수 있다. 그러한 네트워크 디바이스들은 상이한 유형들의 네트워크 통신 프로토콜들을 이용하도록 구성되거나 설계될 수 있는 다수의 네트워크 인터페이스들을 가질 수 있다. 이들 머신들의 일부에 대한 일반적인 아키텍처는 본 명세서에 개시된 설명으로부터 명백할 수 있다. 특정 실시예들에 따라, 본 명세서에 개시된 다양한 가상 비서 실시예들의 피처들 및/또는 기능성들의 적어도 일부는, 최종 사용자 컴퓨터 시스템, 컴퓨터, 네트워크 서버 또는 서버 시스템과 같은 하나 이상의 범용 네트워크 호스트 머신들, 모바일 컴퓨팅 디바이스(예를 들어, PDA(personal digital assistant), 모바일 폰, 스마트폰, 랩톱, 태블릿 컴퓨터 등), 가전 디바이스, 음악 재생기 또는 라우터, 스위치 등과 같은 임의의 그외의 적절한 전자 디바이스 또는 그 조합으로 구현될 수 있다. 적어도 일부 실시예들에서, 본 명세서에서 개시된 다양한 가상 비서 실시예들의 피처들 및/또는 기능들의 적어도 일부는 하나 이상의 가상화된 컴퓨팅 환경들(예를 들어, 네트워크 컴퓨팅 클라우드들 등)에서 구현될 수 있다.The software / hardware hybrid implementation (s) of at least a portion of the virtual assistant embodiment (s) disclosed herein may be implemented on a programmable machine selectively activated or reconfigured by a computer program stored in memory. Such network devices may have multiple network interfaces that may be configured or designed to use different types of network communication protocols. The general architecture for some of these machines may be apparent from the description disclosed herein. According to certain embodiments, at least some of the features and / or functionality of the various virtual assistant embodiments disclosed herein may be implemented in one or more general purpose network host machines, such as an end user computer system, a computer, a network server or a server system, Any other suitable electronic device, such as a mobile computing device (e.g., a personal digital assistant (PDA), a mobile phone, a smartphone, a laptop, a tablet computer, etc.), a consumer device, a music player or router, Can be implemented. In at least some embodiments, at least some of the features and / or functions of the various virtual assistant embodiments disclosed herein may be implemented in one or more virtualized computing environments (e.g., network computing clouds, etc.) .

이제 도 28을 참조하면, 본 명세서에 개시된 가상 비서(virtual assistant) 피처들 및/또는 기능들의 적어도 일부를 구현하기에 적합한 컴퓨팅 디바이스(60)를 나타내는 블록도가 도시된다. 컴퓨팅 디바이스(60)는, 예를 들면, 엔드 유저 컴퓨터 시스템, 네트워크 서버 또는 서버 시스템, 모바일 컴퓨팅 디바이스 (예를 들면, PDA(personal digital assistant), 모바일 폰, 스마트폰, 랩탑, 테블릿 컴퓨터 등), 가전 디바이스, 음악 재생기, 또는 임의의 다른 적당한 전자 디바이스, 또는 그들의 임의의 조합 또는 부분이 될 수 있다. 컴퓨팅 디바이스(60)는 클라이언트들 및/또는 서버들과 같은 다른 컴퓨팅 디바이스와 인터넷과 같은 통신 네트워크를 통하여 그러한 통신을 위한 공지의 프로토콜을 사용하여 무선 또는 유선으로 통신하도록 구성될 수 있다.Referring now to FIG. 28, a block diagram illustrating a computing device 60 suitable for implementing at least some of the virtual assistant features and / or functions described herein is shown. The computing device 60 may be, for example, an end-user computer system, a network server or server system, a mobile computing device (e.g., a personal digital assistant (PDA), a mobile phone, a smartphone, a laptop, , A consumer device, a music player, or any other suitable electronic device, or any combination or portion thereof. The computing device 60 may be configured to communicate wirelessly or wiredly with other computing devices, such as clients and / or servers, using a known protocol for such communication over a communication network such as the Internet.

일 실시예에서, 컴퓨팅 디바이스(60)는 중앙 처리 유닛(CPU)(62), 인터페이스들(68), 및 (PCI(peripheral component interconnect) 등의) 버스(67)를 포함한다. 적절한 소프트웨어 또는 펌웨어의 제어 하에서 동작할 때, CPU(62)는 특정하게 구성된 컴퓨팅 디바이스 또는 기계의 기능들과 연관된 특정 기능들을 구현하는 것을 책임질 수 있다. 예를 들면, 적어도 일 실시예에서, 사용자의 개인 디지털 비서(PDA) 또는 스마트폰이 CPU(62), 메모리(61, 65), 및 인터페이스(들)(68)을 활용하는 가상 비서 시스템으로서 기능하도록 구성되거나 또는 디자인될 수 있다. 적어도 일 실시예에서, CPU(62)가 소프트웨어 모듈들/컴포넌트들 - 이들은, 예를 들어, 운영 시스템 및 임의의 적절한 응용 프로그램 소프트웨어, 드라이버들 등을 포함할 수 있음 - 의 제어 하에서 하나 이상의 상이한 타입의 가상 비서 기능들 및/또는 작업들을 수행하도록 초래될 수 있다.In one embodiment, the computing device 60 includes a central processing unit (CPU) 62, interfaces 68, and a bus 67 (such as a peripheral component interconnect (PCI)). When operating under the control of appropriate software or firmware, the CPU 62 may be responsible for implementing specific functions associated with the functions of the specifically configured computing device or machine. For example, in at least one embodiment, a user's personal digital assistant (PDA) or smart phone may function as a virtual assistant system utilizing CPU 62, memory 61, 65, and interface Or may be designed. In at least one embodiment, the CPU 62 is configured to execute one or more different types of software modules under the control of software modules / components-which may include, for example, an operating system and any suitable application software, drivers, To perform virtual assistant functions and / or tasks.

CPU(62)는 예를 들어, 모토롤라의 프로세서 또는 인텔 마이크로프로세서 패밀리 또는 MIPS 마이크로프로세서 패밀리와 같은 하나 이상의 프로세서(들)(63)을 포함할 수 있다. 일부 실시예에서, 프로세서(들)(63)은 컴퓨팅 디바이스(60)의 동작들을 제어하기 위해 특별히 디자인된 하드웨어(예를 들면, ASIC(application-specific integrated circuit)들, EEPROM(electrically erasable programmable read-only memory)들, FPGA(field-programmable gate array)들 등)를 포함할 수 있다. 특정 실시예에서, (비휘발성 RAM(random access memory) 및/또는 ROM(read-only memory) 등의) 메모리(61)가 또한 CPU(62)의 부분을 형성한다. 그러나, 메모리가 시스템과 결합할 수 있는 많은 상이한 방식들이 있다. 메모리 블록(61)은 예를 들면, 데이터, 프로그래밍 명령어들 등의 캐싱(caching) 및/또는 저장과 같은 다양한 목적을 위해 사용될 수 있다.CPU 62 may include one or more processor (s) 63, such as, for example, a Motorola processor or Intel microprocessor family or MIPS microprocessor family. In some embodiments, processor (s) 63 may comprise hardware (e.g., application-specific integrated circuits (ASICs), electrically erasable programmable read- only memories, field-programmable gate arrays (FPGAs), etc.). In certain embodiments, memory 61 (such as non-volatile random access memory (RAM) and / or read-only memory (ROM)) also forms part of CPU 62. However, there are many different ways in which memory can be combined with a system. The memory block 61 may be used for various purposes such as, for example, caching and / or storage of data, programming instructions, and the like.

본 명세서에서 사용되는 바와 같이, "프로세서"라는 용어는 단지 당해 기술 분야에서 프로세서라고 지칭되는 그러한 집적 회로들에 한정되는 것이 아니라, 마이크로컨트롤러, 마이크로컴퓨터, 프로그램가능 로직 컨트롤러(programmable logic controller), ASIC(application-specific integrated circuit), 및 임의의 다른 프로그램가능 회로를 널리 지칭한다.As used herein, the term "processor" is not limited to such integrated circuits referred to in the art as just a processor in the art, but may be a microcontroller, a microcomputer, a programmable logic controller, an ASIC (application-specific integrated circuit), and any other programmable circuit.

일 실시예에서, 인터페이스들(68)은 (때때로 "라인 카드들"로 지칭되는) 인터페이스 카드들로서 제공된다. 일반적으로, 그것들은 컴퓨팅 네트워크 상의 데이터 패킷들의 전송 및 수신을 제어하고 때때로 컴퓨팅 디바이스(60)와 함께 사용되는 다른 주변기기들을 지원한다. 제공될 수 있는 인터페이스들 중에는 이더넷 인터페이스들, 프레임 릴레이 인터페이스들, 케이블 인터페이스들, DSL 인터페이스들, 토큰 링 인터페이스들 등이 있다. 또한, 예를 들어 유니버설 시리얼 버스(USB), 시리얼, 이더넷, 화이어와이어, PCI, 패러랠, 무선 주파수(RF), 블루투스^TM, (예를 들면, 근거리 자기장 마그네틱들을 이용하는) 근거리 자기장 통신, 802.11(WiFi), 프레임 릴레이, TCP/IP, ISDN, 패스트 이더넷 인터페이스들, 기가비트 이더넷 인터페이스들, 비동기 전달 모드(ATM) 인터페이스들, 고속 시리얼 인터페이스(HSSI) 인터페이스들, POS(Point of Sale) 인터페이스들, FDDI(fiber data distributed interface)들 등과 같은 다양한 타입의 인터페이스들이 제공될 수 있다. 일반적으로, 이러한 인터페이스들(68)은 적절한 매체와의 통신에 적절한 포트들을 포함할 수 있다. 일부 경우들에서, 그것들은 또한 독립적인 프로세서를 포함할 수 있으며, 일부 사례에서 휘발성 및/또는 비휘발성 메모리(예를 들면, RAM)를 포함할 수 있다.In one embodiment, the interfaces 68 are provided as interface cards (sometimes referred to as "line cards"). In general, they control the transmission and reception of data packets on the computing network and sometimes support other peripherals used with the computing device 60. Among the interfaces that may be provided are Ethernet interfaces, frame relay interfaces, cable interfaces, DSL interfaces, token ring interfaces, and the like. In addition, it is also possible to use, for example, a universal serial bus (USB), serial, Ethernet, wirewire, PCI, parallel, radio frequency (RF), Bluetooth ^TM , near field communication (e.g. using near field magnetic) ), Frame Relay, TCP / IP, ISDN, Fast Ethernet interfaces, Gigabit Ethernet interfaces, Asynchronous Transfer Mode (ATM) interfaces, High Speed Serial Interface (HSSI) interfaces, POS (Point of Sale) fiber data distributed interfaces), and the like. In general, these interfaces 68 may include ports suitable for communication with the appropriate media. In some cases, they may also include independent processors and may include volatile and / or nonvolatile memory (e.g., RAM) in some instances.

도 28에 도시된 시스템이 본 명세서에 기술된 발명의 기술을 구현하기 위한 컴퓨팅 디바이스(60)에 대한 하나의 특정 아키텍쳐를 도시하지만, 이는 절대로 본 명세서에 기술된 특징들 및 기술들의 적어도 일부가 구현될 수 있는 유일한 디바이스 아키텍쳐는 아니다. 예를 들면, 하나 또는 임의의 수의 프로세서들(63)을 가지는 아키텍쳐들이 사용될 수 있으며, 그러한 프로세서들(63)은 하나의 디바이스 내에 존재하거나 임의의 수의 디바이스들 사이에 분포할 수 있다. 일 실시예에서, 하나의 프로세서(63)는 라우팅 계산과 함께 통신을 취급한다. 다양한 실시예들에서, 상이한 타입의 가상 비서 특징들 및/또는 기능들이 (개인 디지털 비서 또는 클라이언트 소프트웨어를 실행 중인 스마트폰과 같은) 클라이언트 디바이스 및 (이하에서 더 상세히 설명될 서버 시스템과 같은) 서버 시스템(들)을 포함하는 가상 비서 시스템에서 구현될 수 있다.Although the system shown in FIG. 28 illustrates one particular architecture for computing device 60 for implementing the inventive techniques described herein, it is by no means that at least some of the features and techniques described herein It is not the only device architecture that can be. For example, architectures having one or any number of processors 63 may be used, and such processors 63 may be within one device or distributed among any number of devices. In one embodiment, one processor 63 handles communication with the routing computation. In various embodiments, different types of virtual assistant features and / or functions may be implemented on a client device (such as a smartphone running personal digital assistant or client software) and a server system (such as a server system, (S) in the virtual secretarial system.

네트워크 디바이스 구성과 무관하게, 본 발명의 시스템은 데이터, 범용 네트워크 연산들을 위한 프로그램 명령어들 및/또는 본 명세서에 기술된 가상 비서 기술들의 기능과 관련된 다른 정보를 저장하도록 구성된 하나 이상의 메모리들 또는 (예를 들면, 메모리 블록(65)과 같은) 메모리 모듈들을 채용할 수 있다. 프로그램 명령어들은 예를 들어 운영 시스템의 동작 및/또는 하나 이상의 응용 프로그램들을 제어할 수 있다. 메모리 또는 메모리들은 또한 데이터 구조들, 키워드 분류 정보, 광고 정보, 사용자 클릭 및 노출(click and impression) 정보, 및/또는 본 명세서에 설명된 다른 특정한 비 프로그램 정보를 저장하도록 구성될 수 있다.Regardless of the network device configuration, the system of the present invention may include one or more memories (e. G., &Lt; RTI ID = 0.0 > For example, a memory block 65). The program instructions may, for example, control the operation of the operating system and / or one or more application programs. The memory or memories may also be configured to store data structures, keyword classification information, advertising information, user click and impression information, and / or other specific non-program information described herein.

그러한 정보 및 프로그램 명령어들이 본 명세서에 설명된 시스템들/방법들을 구현하기 위해 채용될 수 있기 때문에, 적어도 일부의 네트워크 디바이스 실시예들은 비일시적(nontransitory) 기계 판독가능 저장 매체를 포함할 수 있는데, 비일시적 기계 판독가능 저장 매체는, 예를 들어, 본 명세서에서 설명된 다양한 작업들을 실행하기 위한 프로그램 명령어들, 상태 정보 등을 저장하도록 구성되거나 설계될 수 있다. 그러한 비일시적 기계 판독가능 매체의 예들은 하드 디스크들, 플로피 디스크들, 및 마그네틱 테이프와 같은 마그네틱 매체; CD-ROM 디스크들과 같은 광학 매체; 플롭티컬 디스크(floptical disk)들과 같은 마그네토-광학 매체, 및 프로그램 명령어들을 저장하고 실행하도록 특별히 구성된 판독전용 메모리 디바이스들(ROM), 플래시 메모리, 멤리스터(memristor) 메모리, 랜덤 액세스 메모리(RAM) 등과 같은 하드웨어 디바이스들을 포함하지만, 이에 한정되는 것은 아니다. 프로그램 명령어들의 예들은 컴파일러에 의해 생성된 것과 같은 기계 코드(machine code), 및 해석기를 사용하는 컴퓨터에 의해 실행될 수 있는 더 높은 레벨의 코드를 포함하는 파일들 양쪽을 포함한다.Because such information and program instructions may be employed to implement the systems / methods described herein, at least some of the network device embodiments may include a nontransitory machine-readable storage medium, Temporary machine-readable storage media may be configured or designed to store, for example, program instructions, status information, and the like, for performing the various tasks described herein. Examples of such non-volatile machine-readable media include magnetic media such as hard disks, floppy disks, and magnetic tape; Optical media such as CD-ROM disks; Read only memory devices (ROM), flash memory, memristor memory, random access memory (RAM), and magneto-optical media, such as magneto-optical media, such as floppy disks, , And the like, but are not limited thereto. Examples of program instructions include both machine code as generated by the compiler and files containing higher level code that can be executed by a computer using the interpreter.

일 실시예에서, 본 발명의 시스템은 독립형 컴퓨팅 시스템 상에 구현된다. 이제 도 29를 참조하면, 적어도 일 실시예에 따라 독립형 컴퓨팅 시스템 상에 가상 비서의 적어도 부분을 구현하는 아키텍쳐를 도시하는 블록도가 도시된다. 컴퓨팅 디바이스(60)는 가상 비서(1002)를 구현하는 소프트웨어를 실행하는 프로세스(들)(63)을 포함한다. 입력 디바이스(1206)는 예를 들어 키보드, 터치스크린, (예를 들면, 음성 입력용의) 마이크로폰, 마우스, 터치 패드, 트랙볼, 5 웨이 스위치(five-way switch), 조이스틱, 및/또는 그것들의 결합을 포함하여, 사용자 입력을 수신하기에 적당한 임의의 타입이 될 수 있다. 출력 디바이스(1207)는 스크린, 스피커, 프린터, 및/또는 그들의 조합이 될 수 있다. 메모리(1210)는 소프트웨어를 실행하는 중에 프로세서(들)(63)에 의해 사용되는, 당해 기술 분야에 알려진 것과 같은 구조 및 아키텍쳐를 가지는 랜덤 액세스 메모리가 될 수 있다. 스토리지 디바이스(1208)는 데이터를 디지털 형태로 저장하기 위한 임의의 마그네틱, 광학, 및/또는 전기적 스토리지 디바이스가 될 수 있다; 예로는 플래시 메모리, 마그네틱 하드 디스크, CD-ROM 등이 포함된다.In one embodiment, the system of the present invention is implemented on a standalone computing system. Referring now to FIG. 29, a block diagram illustrating an architecture for implementing at least a portion of a virtual secretary on a standalone computing system in accordance with at least one embodiment is shown. Computing device 60 includes process (s) 63 for executing software implementing virtual secretary 1002. The input device 1206 may be, for example, a keyboard, a touch screen, a microphone (e.g., for voice input), a mouse, a touchpad, a trackball, a five-way switch, a joystick, and / May be any type suitable for receiving user input, including combining. Output device 1207 may be a screen, a speaker, a printer, and / or a combination thereof. Memory 1210 may be a random access memory having the same architecture and architecture as known in the art, used by processor (s) 63 during execution of software. The storage device 1208 may be any magnetic, optical, and / or electrical storage device for storing data in digital form; Examples include flash memory, magnetic hard disk, CD-ROM, and the like.

다른 실시예에서, 본 발명의 시스템은 임의의 수의 클라이언트들 및/또는 서버들을 갖는 것과 같은 분산된 컴퓨팅 네트워크 상에 구현된다. 이제 도 30을 참조하면, 적어도 일 실시예에 따라, 분산된 컴퓨팅 네트워크 상에 가상의 비서의 적어도 일부분을 구현하기 위한 아키텍처를 도시하는 블록도가 도시되어 있다.In another embodiment, the system of the present invention is implemented on a distributed computing network, such as having any number of clients and / or servers. Referring now to FIG. 30, there is shown a block diagram illustrating an architecture for implementing at least a portion of a virtual secretary on a distributed computing network, according to at least one embodiment.

도 30에 도시된 배열에서, 임의의 수의 클라이언트들(1304)이 제공된다; 각각의 클라이언트(1304)는 본 발명의 클라이언트측 부분들을 구현하기 위한 소프트웨어를 실행할 수 있다. 또한, 클라이언트들(1304)로부터 수신된 요청들을 핸들링하기 위해 임의의 수의 서버들(1340)이 제공될 수 있다. 클라이언트들(1304) 및 서버들(1340)은 인터넷과 같은 전자 네트워크(1361)를 통해 서로 통신할 수 있다. 네트워크(1361)는 예를 들어 유선 및/또는 무선 프로토콜들을 포함하는 임의의 공지된 네트워크 프로토콜들을 사용하여 구현될 수 있다.In the arrangement shown in FIG. 30, any number of clients 1304 are provided; Each client 1304 may execute software for implementing the client-side portions of the present invention. In addition, any number of servers 1340 may be provided to handle requests received from clients 1304. Clients 1304 and servers 1340 may communicate with each other via an electronic network 1361, such as the Internet. The network 1361 may be implemented using any known network protocols including, for example, wired and / or wireless protocols.

또한, 일 실시예에서, 서버들(1340)은 추가 정보의 취득 또는 특정한 사용자들과의 이전 상호작용들에 관한 데이터의 저장이 필요한 경우 외부 서비스들(1360)을 호출할 수 있다. 외부 서비스들(1360)과의 통신은 예를 들어 네트워크(1361)를 통해 발생할 수 있다. 다양한 실시예들에서, 외부 서비스들(1360)은 웹-인에이블된 서비스들 및/또는 하드웨어 디바이스 자체에 관한 또는 그것에 설치된 기능을 포함한다. 예를 들어, 스마트폰 또는 다른 전자 디바이스에 비서(1002)가 구현되는 실시예에서, 비서(1002)는 캘린더 애플리케이션("앱(app)"), 연락처 및/또는 다른 소스들에 저장된 정보를 취득할 수 있다.In addition, in one embodiment, the servers 1340 may invoke external services 1360 when it is necessary to acquire additional information or to store data relating to previous interactions with particular users. Communication with external services 1360 may occur via network 1361, for example. In various embodiments, the external services 1360 include functionality related to or installed on the web-enabled services and / or the hardware device itself. For example, in an embodiment where a secretary 1002 is implemented on a smartphone or other electronic device, the secretary 1002 may obtain information stored in a calendar application ("app"), contacts and / can do.

다양한 실시예들에서, 비서(1002)는 그것이 설치된 전자 디바이스의 다수의 피처들 및 동작들을 제어할 수 있다. 예를 들어, 비서(1002)는, 달리 디바이스의 종래의 사용자 인터페이스를 사용하여 개시될 수 있는 기능들 및 동작들을 수행하기 위해, API들을 통해 또는 다른 수단에 의해 디바이스의 기능 및 애플리케이션들과 인터페이스하는 외부 서비스들(1360)을 호출할 수 있다. 이러한 기능들 및 동작들은, 예를 들어, 알람을 설정하고, 전화하고, 텍스트 메시지 또는 이메일 메시지를 전송하고, 캘린더 이벤트를 추가하는 것 등을 포함할 수 있다. 이러한 기능들 및 동작들은 사용자와 비서(1002) 사이의 일상 대화에서 쓰이는 대화의 컨텍스트에서 추가 기능들로서 수행될 수 있다. 이러한 기능들 및 동작들은 이러한 대화의 컨텍스트에서 사용자에 의해 특정될 수 있고, 또는 대화의 컨텍스트에 기초하여 자동으로 수행될 수 있다. 당업자는 비서(1002)가 그에 의해 전자 디바이스의 다양한 동작들을 개시하고 제어하기 위한 제어 메커니즘으로서 사용될 수 있고, 전자 디바이스는 버튼들 또는 그래픽 사용자 인터페이스들 등의 종래의 메커니즘들에 대한 대안으로서 사용될 수 있다는 것을 인식할 것이다.In various embodiments, the secretary 1002 can control a number of features and operations of the electronic device in which it is installed. For example, the secretary 1002 may interact with the functions and applications of the device via APIs or by other means to perform functions and operations that may otherwise be initiated using the device's conventional user interface May call external services 1360. These functions and operations may include, for example, setting up an alarm, calling, sending a text or email message, adding calendar events, and the like. These functions and actions may be performed as additional functions in the context of the conversation used in the daily conversation between the user and the secretary 1002. [ These functions and actions may be specified by the user in the context of such a conversation, or may be performed automatically based on the context of the conversation. Those skilled in the art will appreciate that the secretary 1002 can thereby be used as a control mechanism for initiating and controlling various operations of the electronic device and the electronic device can be used as an alternative to conventional mechanisms such as buttons or graphical user interfaces &Lt; / RTI >

예를 들어, 사용자는 비서(1002)에 "나는 내일 아침 8시에 일어날 필요가 있다"라는 입력을 제공할 수 있다. 비서(1002)가 본 명세서에 기술된 기술들을 이용하여 사용자의 의도를 판정하면, 비서(1002)는 외부 서비스들(1340)을 호출하여 디바이스의 알람 시계 기능 또는 애플리케이션과 인터페이싱한다. 비서(1002)는 사용자를 대신하여 알람을 설정한다. 이러한 방식으로, 사용자는 알람을 설정하거나 또는 디바이스의 다른 기능들을 수행하기 위한 종래의 메커니즘들에 대한 대체물로서 비서(1002)를 사용할 수 있다. 사용자의 요청들이 모호하거나 또는 추가의 명료화를 필요로 하는 경우, 비서(1002)는 적극적인 유도(elicitation), 바꾸어 말하기(paraphrasing), 제안들 등을 포함하며 취득 컨텍스트 정보를 포함하는 본 명세서에 기술된 다양한 기술들을 사용할 수 있어 정확한 서비스들(1340)이 호출되고 의도된 동작이 취해진다. 일 실시예에서, 비서(1002)는 서비스(1340)를 호출하여 기능을 수행하기 전에 임의의 적절한 소스로부터의 추가의 컨텍스트 정보의 요청 및/또는 확인을 사용자에게 재촉할 수 있다. 일 실시예에서, 사용자는 특정한 서비스들(1340)을 호출하는 비서(1002)의 능력을 선택적으로 디스에이블할 수 있거나 또는 원한다면 이러한 모든 서비스 호출을 디스에이블할 수 있다.For example, the user may provide the secretary 1002 with an entry saying "I need to get up at 8am tomorrow ". If the secretary 1002 determines the intent of the user using the techniques described herein, the secretary 1002 invokes external services 1340 to interface with the device's alarm clock function or application. Secretary 1002 sets an alarm on behalf of the user. In this manner, the user can use the secretary 1002 as an alternative to conventional mechanisms for setting alarms or performing other functions of the device. If the user's requests are ambiguous or require additional clarification, the secretary 1002 may include an active elicitation, paraphrasing, suggestions, etc., Various techniques may be used, so that the correct services 1340 are invoked and the intended operation is taken. In one embodiment, the secretary 1002 may prompt the user to request and / or confirm additional contextual information from any suitable source before invoking the service 1340 to perform the function. In one embodiment, the user may selectively disable the ability of the secretary 1002 to call specific services 1340, or may disable all such service calls, if desired.

본 발명의 시스템은 동작 모드들 및 클라이언트들(1304)의 다수의 상이한 유형들 중 임의의 것을 이용하여 구현될 수 있다. 이제, 도 31을 참조하면, 동작 모드들 및 클라이언트들(1304)의 몇몇의 상이한 유형들을 예시하는 시스템 아키텍처를 도시하는 블록도가 도시된다. 당업자는, 도 31에 도시된 동작 모드들 및 클라이언트들(1304)의 다양한 유형들이 단순히 예시적이며, 본 발명의 시스템은 도시된 것들 이외의 동작 모드들 및/또는 클라이언트들(1304)을 이용하여 구현될 수 있다는 것을 인식할 것이다. 또한, 시스템은 동작 모드들 및/또는 이러한 클라이언트들(1304) 중 임의의 것 또는 모두를 단독으로 또는 임의의 조합으로 포함할 수 있다. 도시된 예시들은 다음을 포함한다: The system of the present invention may be implemented using any of a number of different types of operating modes and clients 1304. Referring now to FIG. 31, a block diagram illustrating a system architecture illustrating operating modes and several different types of clients 1304 is shown. Those skilled in the art will appreciate that the various types of operating modes and clients 1304 shown in Figure 31 are merely exemplary and that the system of the present invention may be implemented using operating modes and / As will be appreciated by those skilled in the art. In addition, the system may include any or all of the operating modes and / or such clients 1304, either alone or in any combination. Illustrative examples include:

입력/출력 디바이스들 및/또는 센서들(1402)을 갖는 컴퓨터 디바이스들. 클라이언트 컴포넌트는 임의의 이러한 컴퓨터 디바이스(1402)에 배치될 수 있다. 적어도 하나의 실시예는 네트워크(1361)를 통한 서버들(1340)과의 통신을 가능하게 하기 위해 웹 브라우저(1304A) 또는 다른 소프트웨어 애플리케이션을 사용하여 구현될 수 있다. 입력 및 출력 채널들은 예를 들어 시각 및/또는 청각 채널들을 포함하는 임의의 유형일 수 있다. 예를 들어, 일 실시예에서, 본 발명의 시스템은, 웹 브라우저의 동등물이 음성에 의해 구동되고 출력을 위해 음성을 사용하는 맹인용 비서의 실시예를 가능하게 하는, 음성 기반 통신 방법들을 사용하여 구현될 수 있다.

Computer devices having input / output devices and / or sensors 1402. The client components may be located in any such computer device (1402). At least one embodiment may be implemented using a web browser 1304A or other software application to enable communication with the servers 1340 over the network 1361. [ The input and output channels may be of any type, including, for example, visual and / or auditory channels. For example, in one embodiment, the system of the present invention may be implemented using voice-based communication methods that enable embodiments of the blind secretary using the voice for output and the equivalent of a web browser driven by voice Can be implemented.

클라이언트가 모바일 디바이스(1304B)의 애플리케이션으로서 구현될 수 있는 I/O 및 센서들(1406)을 갖는 모바일 디바이스들. 이것은 휴대폰, 스마트폰, PDA, 태블릿 디바이스, 네트워크 게임 콘솔 등을 포함하지만 이에 한정되는 것은 아니다.

Mobile devices having I / O and sensors 1406 that a client can be implemented as an application of mobile device 1304B. This includes, but is not limited to, cell phones, smart phones, PDAs, tablet devices, network game consoles, and the like.

클라이언트가 기기(1304C)의 임베딩된 애플리케이션으로서 구현될 수 있는 I/O 및 센서들(1410)을 갖는 소비자 가전들.

Consumer appliances having I / O and sensors 1410, wherein the client may be implemented as an embedded application of device 1304C.

클라이언트가 임베딩된 시스템 애플리케이션(1304D)으로서 구현될 수 있는 대시보드 인터페이스들 및 센서들(1414)을 갖는 승용차들 및 다른 차량들. 이것은 네비게이션 시스템, 음성 제어 시스템, 차량내 엔터테인먼트 시스템 등을 포함하지만 이에 한정되는 것은 아니다.

Passenger cars and other vehicles having dashboard interfaces and sensors 1414 that can be implemented as a client embedded system application 1304D. This includes, but is not limited to, navigation systems, voice control systems, in-vehicle entertainment systems, and the like.

클라이언트가 디바이스 상주 애플리케이션(1304E)으로서 구현될 수 있는, 라우터들(1418) 또는 네트워크에 상주하거나 네트워크와 인터페이스하는 임의의 다른 디바이스 등의 네트워킹된 컴퓨팅 디바이스들.

Networked computing devices such as routers 1418, or any other device that resides on or interfaces with a network, where the client may be implemented as a device resident application 1304E.

비서의 실시예가 이메일 모달리티 서버(1426)를 통해 접속되는 이메일 클라이언트들(1424). 이메일 모달리티 서버(1426)는, 예를 들어 사용자로부터의 입력을 비서로 전송된 이메일 메시지들로서 수취하고 응답들로서 비서로부터의 출력을 사용자에게 전송하는, 통신 브리지로서 동작한다.

Email clients 1424 in which embodiments of the secretary are connected via email modality server 1426. The email modality server 1426 acts as a communication bridge, for example receiving input from a user as email messages sent to a secretary and sending the output from a secretary as a response to the user.

비서의 실시예가 메시지 모달리티 서버(1430)를 통해 접속되는 인스턴트 메시지 클라이언트들(1428). 메시지 모달리티 서버(1430)는, 사용자로부터의 입력을 비서로 전송된 메시지들로서 수취하고 비서로부터의 출력을 응답 시의 메시지들로서 사용자에게 전송하는, 통신 브리지로서 동작한다.

Instant messaging clients 1428 in which embodiments of a secretary are connected through a message modality server 1430. The message modality server 1430 acts as a communication bridge that receives input from the user as messages sent to the secretary and sends the output from the secretary as messages in response to the user.

비서의 실시예가 VoIP(Voice over Internet Protocol) 모달리티 서버(1430)를 통해 접속되는 음성 전화들(1432). VoIP 모달리티 서버(1430)는, 사용자로부터의 입력을 비서에게 말한 음성으로서 수취하고 비서로부터의 출력을 예를 들어 응답 시의 합성된 음성으로서 사용자에게 전송하는, 통신 브리지로서 동작한다.

Voice telephones 1432 to which embodiments of the secretary are connected via a Voice over Internet Protocol (VoIP) modality server 1430. The VoIP modality server 1430 acts as a communication bridge that receives the input from the user as a voice spoken to the secretary and transmits the output from the secretor as a synthesized voice, for example, in response.

이메일, 인스턴트 메시징, 토론 포럼들, 그룹 채팅 세션들, 실시간 도움 또는 고객 지원 세션들, 등을 포함하지만 이에 제한되지는 않는 메시징 플랫폼들에 대해, 비서(1002)는 대화들 내의 참가자로서 기능할 수 있다. 비서(1002)는 대화를 모니터링하고, 일대일 상호작용들에 대한 본원에 개시된 하나 이상의 기법들 및 방법들을 이용하는 개인들 또는 그룹에 응답할 수 있다.For messaging platforms, including but not limited to e-mail, instant messaging, discussion forums, group chat sessions, real-time help or customer support sessions, secretary 1002 can act as a participant in conversations have. Secretary 1002 may monitor the conversation and respond to individuals or groups using one or more of the techniques and methods described herein for one-to-one interactions.

다양한 실시예들에서, 본 발명의 기법들을 구현하기 위한 기능성은 임의의 수의 클라이언트 및/또는 서버 컴포넌트들 중에 분산될 수 있다. 예를 들면, 다양한 소프트웨어 모듈들은 본 발명과 관련된 다양한 기능들을 수행하기 위해 구현될 수 있고, 그러한 모듈들은 서버 및/또는 클라이언트 컴포넌트들 상에서 구동하기 위해 다양하게 구현될 수 있다. 그러한 구성에 대한 추가적인 상세는, 그 전체가 참조로서 본원에 통합되는, 2011년 1월 10일에 출원된, 대리인 문서 번호 P10575US1인, 관련된 미국 특허 출원 제12/987,982호, "Intelligent Automated Assistant"에 제공된다.In various embodiments, functionality for implementing the techniques of the present invention may be distributed among any number of client and / or server components. For example, various software modules may be implemented to perform various functions related to the present invention, and such modules may be variously implemented to operate on a server and / or client components. Further details on such arrangements may be found in the relevant U.S. Patent Application No. 12 / 987,982, "Intelligent Automated Assistant ", Attorney Docket No. P10575US1, filed January 10, 2011, which is incorporated herein by reference in its entirety / RTI >

도 32의 예에서, 입력 유도 기능 및 출력 처리 기능은 클라이언트(1304) 및 서버(1340)에 분산되고, 입력 유도의 클라이언트 부분(2794a) 및 출력 처리의 클라이언트 부분(2792a)은 클라이언트(1304)에 위치하고, 입력 유도의 서버 부분(2794b) 및 출력 처리의 서버 부분(2792b)은 서버(1340)에 위치한다. 이하의 컴포넌트들은 서버(1340)에 위치한다:32, the input and output processing functions are distributed to the client 1304 and the server 1340, and the client portion 2794a of the input induction and the client portion 2792a of the output processing are distributed to the client 1304 And the server portion 2794b of the input induction and the server portion 2792b of the output processing are located at the server 1340. [ The following components are located at server 1340:

완성된 어휘(2758b);

Completed vocabulary (2758b);

언어 패턴 인식기들의 완성된 라이브러리(2760b);

A completed library 2760b of language pattern recognizers;

단기 개인 메모리의 마스터 버전(2752b);

A master version 2752b of short term private memory;

장기 개인 메모리의 마스터 버전(2754b).

Master version of long-term personal memory (2754b).

일 실시예에서, 클라이언트(1304)는, 응답성을 향상시키고 네트워크 통신에 대한 의존을 감소시키기 위해, 이들 컴포넌트들의 서브세트들 및/또는 부분들을 국부적으로 유지한다. 그러한 서브세트들 및/또는 부분들은 공지된 캐시 관리 기법들에 따라 유지되고 업데이트될 수 있다. 그러한 서브세트들 및/또는 부분들은, 예를 들면, 이하를 포함한다:In one embodiment, client 1304 locally maintains subsets and / or portions of these components to improve responsiveness and reduce dependence on network communications. Such subsets and / or portions may be maintained and updated in accordance with known cache management techniques. Such subsets and / or portions include, for example, the following:

어휘의 서브세트(2758a);

A subset of lexicals 2758a;

언어 패턴 인식기들의 라이브러리의 서브세트(2760a);

A subset 2760a of libraries of language pattern recognizers;

단기 개인 메모리의 캐시(2752a);

Cache 2752a of short term private memory;

장기 개인 메모리의 캐시(2754a).

A cache of long term private memory (2754a).

추가적인 컴포넌트들은, 예를 들어, 이하를 포함하는, 서버(1340)의 부분으로서 구현될 수 있다.Additional components may be implemented as part of server 1340, including, for example,

언어 해석기(2770);

A language interpreter 2770;

대화 플로우 프로세서(2780);

A conversation flow processor 2780;

출력 프로세서(2790);

An output processor 2790;

도메인 엔티티 데이터베이스들(2772);

Domain entity databases 2772;

태스크 플로우 모델들(2786);

Task flow models 2786;

서비스 편성(2782);

Service organization 2782;

서비스 능력 모델들(2788).

Service Capability Models (2788).

이들 컴포넌트들 각각은 이하에 더욱 상세히 설명될 것이다. 서버(1340)는 필요할 때 외부 서비스들(1360)과 인터페이싱함으로써 추가적인 정보를 얻는다.Each of these components will be described in more detail below. Server 1340 obtains additional information by interfacing with external services 1360 as needed.

개념상의 Conceptual 아키텍쳐Architecture

이제 도 27을 참조하면, 가상 비서(1002)의 특정 예시의 실시예의 간략화된 블록도가 도시된다. 상기에 언급한 관련된 미국 특허 출원들에 더욱 상세히 개시된 바와 같이, 가상 비서(1002)의 상이한 실시예들은, 가상 비서 기술에 일반적으로 관련되는 다양한 상이한 유형의 동작들, 기능들, 및/또는 피처들을 제공하도록 구성되고, 설계되고, 및/또는 동작가능할 수 있다. 또한, 본원에 더욱 상세히 개시되는 바와 같이, 본원에 개시된 가상 비서(1002)의 다양한 동작들, 기능들, 및/또는 특징들 중 많은 수가 가상 비서(1002)와 상호작용하는 상이한 엔티티들에 대해 상이한 유형의 장점들 및/또는 이점들을 가능하게 하거나 또는 제공할 수 있다. 도 27에 도시된 실시예는 전술한 임의의 하드웨어 아키텍쳐들을 이용하여, 또는 상이한 유형의 하드웨어 아키텍쳐를 이용하여 구현될 수 있다.Referring now to FIG. 27, a simplified block diagram of an embodiment of a specific example of virtual secretary 1002 is shown. Different embodiments of the virtual secretary 1002, as described in more detail in the above-mentioned related U.S. patent applications, may provide various different types of operations, functions, and / or features And / or may be operable. It will also be appreciated that many of the various operations, functions, and / or features of the virtual secretary 1002 disclosed herein may be different for different entities that interact with the virtual secretary 1002, Type advantages and / or advantages of the present invention. The embodiment shown in FIG. 27 may be implemented using any of the hardware architectures described above, or using different types of hardware architectures.

예를 들면, 상이한 실시예들에 따르면, 가상 비서(1002)는, 예를 들면, 이하 중 하나 이상(또는 그의 조합들)과 같은, 다양한 상이한 유형들의 동작들, 기능성들, 및/또는 특징들을 제공하도록 구성되고, 설계되고, 및/또는 동작가능할 수 있다:For example, in accordance with different embodiments, the virtual secretary 1002 may include a variety of different types of operations, functionality, and / or features, such as, for example, one or more of the following Be designed, designed, and / or operable to provide:

제품들 및 서비스들을 찾거나, 발견하거나, 그 중에서 선택하거나, 구매하거나, 예약하거나, 또는 주문하기 위해 인터넷을 통해 이용가능한 데이터 및 서비스들의 애플리케이션을 자동화함. 이들 데이터 및 서비스들을 이용하는 것의 프로세스를 자동화하는 것에 더하여, 가상 비서(1002)는 또한 한번에 데이터 및 서비스들의 몇몇 소스들의 결합된 사용을 가능하게 할 수 있다. 예를 들면, 그것은 몇몇 리뷰 사이트들로부터 제품들에 관한 정보를 결합하고, 다수의 유통업자로부터 가격 및 입수가능성을 확인하고, 그들의 위치들 및 시간 제약들을 확인하고, 사용자가 그들의 문제에 대한 개인화된 해법을 찾는 것을 도와줄 수 있다.

Automate the application of data and services available over the Internet to find, discover, select, purchase, reserve, or order products and services. In addition to automating the process of using these data and services, virtual assistant 1002 can also enable the combined use of several sources of data and services at one time. For example, it combines information about products from several review sites, identifies pricing and availability from multiple distributors, identifies their location and time constraints, and allows users to customize their problems You can help find the solution.

(영화들, 이벤트들, 공연들, 전시들, 쇼들 및 어트랙션들을 포함하나 이에 제한되지 않는) 해야할 일들; (여행 목적지들, 호텔들 및 머무르기 위한 그외의 장소들, 랜드마크들 및 관심 있는 그외의 사이트들, 등을 포함하나 이에 제한되지 않는) 가야할 장소들; (레스토랑 및 바와 같은) 먹거나 또는 마시기 위한 장소들, 다른 사람들을 만나기 위한 시간들 및 장소들, 및 인터넷 상에서 찾을 수 있는 엔터테인먼트 또는 소셜 상호작용의 임의의 그외의 소스들에 대해 찾거나, 조사하거나, 그 중에서 선택하거나, 예약하거나, 그렇지 않은 경우 배우기 위해 인터넷을 통해 이용가능한 데이터 및 서비스들의 이용을 자동화함.

(Including, but not limited to, movies, events, performances, exhibitions, shows and attractions); (Including but not limited to travel destinations, hotels and other places to stay, landmarks and other sites of interest, etc.); Find or search for places to eat or drink (such as restaurants and bars), times and places for meeting other people, and any other sources of entertainment or social interaction that can be found on the Internet, Automate the use of data and services available over the Internet to select, reserve, or otherwise learn from them.

(위치 기반 탐색을 포함하는) 탐색; 네비게이션(지도들 및 방향들); (이름 또는 그외의 속성들에 의해 비지니스들 또는 사람들을 찾는 것과 같은) 데이터베이스 검색; 기상 상태 및 예보를 얻는 것, 마켓 아이템들의 가격 또는 금융 거래의 상태를 확인하는 것; 교통 상황 또는 항공편의 상태를 모니터링하는 것; 캘린더들 및 스케쥴들을 액세스하고 업데이트하는 것; 리마인더들, 경보들, 태스크들, 및 프로젝트들을 관리하는 것; 이메일 또는 그외의 메시징 플랫폼들을 통해 통신하는 것; 및 국부적으로 또는 원격으로 디바이스들을 동작시키는 것(예를 들면, 전화를 거는 것, 전등 및 온도를 제어하는 것, 가정 보안 디바이스들을 제어하는 것, 음악 또는 비디오를 플레이하는 것, 등)을 포함하는 그래픽 사용자 인터페이스들을 갖는 전용 애플리케이션들에 의해 다르게 제공되는 자연어 대화를 통해 애플리케이션들 및 서비스들의 동작을 가능하게 함. 일 실시예에서, 가상 비서(1002)는 디바이스 상에서 이용가능한 많은 기능들 및 앱들을 시작하고, 동작시키고, 제어하는 데에 이용될 수 있다.

(Including location based search); Navigation (maps and directions); Database search (such as finding businesses or people by name or other attributes); Obtaining weather conditions and forecasts, confirming the prices of market items or the status of financial transactions; Monitoring the status of traffic or flight; Accessing and updating calendars and schedules; Managing reminders, alerts, tasks, and projects; Communicating via email or other messaging platforms; And operating devices locally or remotely (e.g., dialing, controlling light and temperature, controlling home security devices, playing music or video, etc.) Enabling the operation of applications and services through natural language conversations that are otherwise provided by dedicated applications with graphical user interfaces. In one embodiment, the virtual secretary 1002 can be used to start, operate, and control many of the functions and applications available on the device.

활동들, 제품들, 서비스들, 엔터테인먼트의 소스, 시간 관리, 또는 자연어의 상호작용 대화 및 데이터 및 서비스들에 대한 자동화된 액세스로부터 이익을 얻는 임의의 그외의 유형의 추천 서비스에 대한 개인적 추천들을 제공함.

Provides personal recommendations for activities, products, services, sources of entertainment, time management, or any other type of referral service that benefits from interactive communication of natural language and automated access to data and services .

상이한 실시예들에 따르면, 가상 비서(1002)에 의해 제공되는 다양한 유형의 기능들, 동작들, 액션들 및/또는 그외의 특징들 중 적어도 일부는 하나 이상의 클라이언트 시스템(들)에서, 하나 이상의 서버 시스템(들)에서, 및/또는 그의 조합에서 구현될 수 있다.In accordance with different embodiments, at least some of the various types of functions, operations, actions, and / or other features provided by virtual secretary 1002 may be implemented in one or more client systems (s) System (s), and / or combinations thereof.

상이한 실시예들에 따라, 가상 비서(1002)에 의해 제공되는 다양한 유형의 기능, 동작, 액션, 및/또는 다른 특징들의 적어도 일부는 본원에서 더 자세히 설명되는 바와 같이, 사용자 입력을 해석하고 동작화하는 컨텍스트 관련 정보를 사용할 수 있다.In accordance with different embodiments, at least some of the various types of functions, operations, actions, and / or other features provided by the virtual assistant 1002 can be used to interpret user input, Context-related information can be used.

예를 들면, 적어도 하나의 실시예에서, 가상 비서(1002)는 구체적인 태스크 및/또는 동작을 수행할 때 다양한 상이한 유형들의 데이터 및/또는 다른 유형들의 정보를 이용 및/또는 발생시키도록 동작할 수 있다. 이는 예를 들면, 입력 데이터/정보 및/또는 출력 데이터/정보를 포함할 수 있다. 예를 들면, 적어도 하나의 실시예에서, 가상 비서(1002)는 예를 들면, 하나 이상의 로컬 및/또는 원격 메모리와 같은 하나 이상의 상이한 유형의 소스들, 디바이스 및/또는 시스템들로부터의 정보를 액세스, 프로세스, 및/또는 이용하도록 동작할 수 있다. 또한, 적어도 하나의 실시예에서, 가상 비서(1002)는, 예를 들면, 하나 이상의 로컬 및/또는 원격 디바이스의 메모리 및/또는 시스템에 저장될 수 있는 하나 이상의 상이한 유형의 출력 데이터/정보를 발생시키도록 동작할 수 있다.For example, in at least one embodiment, the virtual secretary 1002 may be operable to utilize and / or to generate various different types of data and / or other types of information when performing a specific task and / have. This may include, for example, input data / information and / or output data / information. For example, in at least one embodiment, virtual secretary 1002 may access information from one or more different types of sources, devices, and / or systems, such as, for example, one or more local and / Process, and / or utilize the functions described herein. Further, in at least one embodiment, the virtual secretary 1002 may generate one or more different types of output data / information that may be stored in, for example, memory and / or systems of one or more local and / or remote devices .

가상 비서(1002)에 의해 액세스 및/또는 이용될 수 있는 입력 데이터/정보의 상이한 유형들의 예시는 하기의 하나 이상(또는 그들의 조합)을 포함할 수 있으나 이에 한정되지는 않는다.Examples of different types of input data / information that may be accessed and / or used by virtual secretary 1002 may include, but are not limited to, one or more of the following (or combinations thereof).

이동 전화 및 태블릿과 같은 이동 디바이스, 마이크를 구비한 컴퓨터, 블루투스 헤드셋, 자동차 음성 제어 시스템, 전화를 통한 시스템, 응답 서비스 상의 기록, 통합 메시징 서비스 상의 오디오 음성 메일, 클록 라디오와 같은 음성 입력을 갖는 소비자 애플리케이션, 전화국, 홈 엔터테인먼트 제어 시스템, 및 게임 컨솔로부터의 : 음성 입력

Consumers with voice inputs such as mobile devices such as mobile phones and tablets, computers with microphones, Bluetooth headsets, automotive voice control systems, telephony systems, records on answering services, audio voicemail on unified messaging services, From applications, phone stations, home entertainment control systems, and gaming consoles: voice input

컴퓨터 또는 이동 디바이스 상의 키보드, 원격 제어 또는 다른 소비자 전자 디바이스 상의 키패드, 비서에 보내지는 이메일 메시지, 비서에 보내지는 인스턴트 메시지 또는 유사한 단문 메시지, 멀티유저 게임 환경에서의 플레이어들로부터 수신한 텍스트, 및 메시지 피드에서 스트리밍되는 텍스트로부터의 텍스트 입력

A keyboard on a computer or mobile device, a keypad on a remote control or other consumer electronic device, an email message sent to a secretary, an instant or similar short message sent to a secretary, text received from players in a multi- Text input from text streamed from the feed

센서 또는 위치 기반 시스템으로부터 들어오는 위치 정보. 예시는 글로벌 위치확인 시스템(GPS) 및 이동 전화 상의 어시스트티드 GPS(A-GPS)를 포함한다. 일 실시예에서, 위치 정보는 명확한 사용자 입력과 결합된다. 일 실시예에서, 본 발명의 시스템은 알려진 어드레스 정보 및 현재 위치 결정에 기초하여, 사용자가 집에 있을 때를 검출할 수 있다. 이러한 방식으로, 사용자가 집 밖에 있을 때와는 달리 집에 있을 때에 흥미가 있을 수 있는 정보의 유형뿐만 아니라 사용자가 집에 있는지의 여부에 따라 사용자를 대신해 호출되어야 하는 서비스 및 액션의 유형들에 대한 특정 추론들이 행해질 수 있다.

Location information from sensors or location-based systems. Examples include Global Positioning System (GPS) and Assisted GPS (A-GPS) on mobile phones. In one embodiment, the location information is combined with explicit user input. In one embodiment, the system of the present invention can detect when the user is at home, based on known address information and current positioning. In this way, the types of services and actions that must be invoked on behalf of the user, depending on whether the user is home or not, as well as the type of information that may be of interest when the user is at home, Certain inferences can be made.

클라이언트 디바이스 상의 클록으로부터의 시간 정보. 이는 예를 들면, 로컬 시간 및 시간 존을 가리키는 전화기 또는 다른 클라이언트 디바이스로부터의 시간을 포함할 수 있다. 또한, 시간은 사용자 요청의 내용, 예를 들면 "한 시간 안에" 및 "오늘밤"과 같은 어구를 해석하는 데에 사용될 수 있다.

Time information from the clock on the client device. This may include, for example, time from a telephone or other client device pointing to local time and time zones. In addition, the time can be used to interpret the content of the user request, such as "within an hour" and "tonight ".

나침반, 가속도계, 자이로스코프, 및/또는 이동 속도 데이터, 뿐만 아니라 이동 또는 핸드헬드 디바이스로부터의 또는 자동차 제어 시스템과 같은 임베디드 시스템으로부터의 다른 센서 데이터. 이는 원격 제어로부터 응용들 및 게임 콘솔까지의 디바이스 위치확인 데이터를 포함할 수 있다.

Compass, accelerometer, gyroscope, and / or movement speed data, as well as other sensor data from a mobile or handheld device or from an embedded system such as a car control system. This may include device location data from remote control to applications and game consoles.

클릭킹 및 메뉴 선택 및 그래픽 유저 인터페이스(GUI)를 갖는 임의의 디바이스 상의 GUI로부터의 다른 이벤트. 추가 예시들은 터치 스크린에의 터치를 포함한다.

Other events from the GUI on any device with click-kings and menu selections and a graphical user interface (GUI). Additional examples include a touch to the touch screen.

센서들 및 알람 클록, 캘린더 알림, 가격 변동 트리거, 위치 트리거, 서버로부터의 디바이스로의 푸시 통지 등과 같은 다른 데이터-구동 트리거로부터의 이벤트.

Events from other data-driven triggers such as sensors and alarm clocks, calendar notifications, pricing triggers, location triggers, push notifications to devices from the server, and so on.

본원에서 설명되는 실시예들에의 입력은 대화 및 요청 히스토리를 포함하는 사용자 상호작용 히스토리의 내용을 또한 포함할 수 있다.Inputs to the embodiments described herein may also include the contents of a user interaction history, including dialogue and request history.

상기의 상호-참조되는 관련 미국 실용신안 출원에 기재된 바와 같이, 많은 상이한 유형들의 출력 데이터/정보가 가상 비서(1002)에 의해 생성될 수 있다. 이들은 하기의 하나 이상(또는 그들의 조합)을 포함할 수 있지만 이에 한정되는 것은 아니다.Many different types of output data / information may be generated by the virtual secretary 1002, as described in the above-referenced related US utility model application. These may include, but are not limited to, one or more of the following (or combinations thereof).

출력 디바이스에 및/또는 디바이스의 사용자 인터페이스에 직접 보내지는 텍스트 출력;

Text output sent directly to the output device and / or to the user interface of the device;

이메일을 통해 사용자에게 보내지는 텍스트 및 그래픽;

Text and graphics sent to users via email;

메시징 서비스를 통해 사용자에게 보내지는 텍스트 및 그래픽;

Text and graphics sent to users through messaging services;

하기의 하나 이상(또는 그들의 조합)을 포함할 수 있는 음성 출력:

Voice output that may include one or more of the following (or combinations thereof):

ㅇ 합성된 음성;Synthesized speech;

ㅇ 샘플링된 음성;O sampled voice;

ㅇ 기록된 음성;O recorded voice;

사진, 리치 텍스트, 비디오, 사운드, 및 하이퍼링크(예를 들면, 웹 브라우저에 렌더링되는 컨텐츠)를 갖는 정보의 그래픽 레이아웃;

Graphic layout of information with pictures, rich text, video, sound, and hyperlinks (e.g., content rendered in a web browser);

디바이스가 턴온 또는 턴오프되게 하고, 사운드를 발생시키게 하고, 컬러를 변경하고, 진동하고, 빛을 제어하게 하는 등의 디바이스상의 물리적 작용을 제어하는 액츄에이터 출력;

An actuator output to control the physical action on the device such as to cause the device to turn on or off, to generate sound, to change color, to vibrate, to control the light;

매핑 애플리케이션을 호출, 전화기를 음성 다이얼링, 이메일 또는 인스턴트 메시지를 전송, 미디어를 재생, 캘린더, 태스크 매니저, 및 노트 애플리케이션의 엔트리를 작성, 및 다른 애플리케이션들과 같은 디바이스 상의 다른 애플리케이션을 호출;

Calling a mapping application, voice dialing a phone, sending an email or an instant message, playing media, creating entries for calendars, task managers, and note applications, and calling other applications on the device, such as other applications;

원격 카메라, 휠체어의 제어, 원격 스피커 상의 음악 재생, 원격 디스플레이 상의 비디오 재생 등과 같은 디바이스에 부착 또는 제어되는 디바이스들에의 물리적 액션을 제어하는 액츄에이터 출력.

An actuator output that controls physical actions on devices attached to or controlled by a device such as a remote camera, control of a wheelchair, music playback on a remote speaker, video playback on a remote display,

도 27의 가상 비서(1002)는 구현될 수 있는 가상 비서 시스템 실시예들의 넓은 범위로부터의 하나의 예시라는 것이 이해될 수 있다. (도시되지 않은) 가상 비서 시스템의 다른 실시예들은 예를 들면, 도 27의 가상 비서 시스템 실시예에 설명된 것들보다 추가적인, 적은 및/또는 상이한 컴포넌트/특징들을 포함할 수 있다.It is to be understood that the virtual secretary 1002 of FIG. 27 is one example from a wide range of virtual assistant system embodiments that may be implemented. Other embodiments of virtual assistant systems (not shown) may include additional, fewer and / or different components / features than those described in the virtual assistant system embodiment of FIG. 27, for example.

가상 비서(1002)는 복수의 상이한 유형의 컴포넌트, 디바이스, 모듈, 프로세스, 시스템 등을 포함할 수 있는데, 이는 예를 들어 하드웨어 및/또는 하드웨어 및 소프트웨어의 조합들의 사용을 통해 구현 및/또는 예시될 수 있다. 예를 들어, 도 27의 실시예에 도시된 바와 같이, 비서(1002)는 시스템, 컴포넌트, 디바이스, 프로세스 등(또는 이들의 조합)의 다음 유형들 중 하나 이상을 포함할 수 있다.The virtual assistant 1002 may include a plurality of different types of components, devices, modules, processes, systems, etc., which may be implemented and / or illustrated, for example, through the use of hardware and / . 27, secretary 1002 may include one or more of the following types of systems, components, devices, processes, etc. (or a combination thereof).

하나 이상의 액티브 온톨로지(ontology)들(1050);

One or more active ontologies 1050;

액티브 입력 유도 컴포넌트(들)(2794)(클라이언트 부분(2794a) 및 서버 부분(2794b)을 포함할 수 있음);

Active input induction component (s) 2794 (which may include a client portion 2794a and a server portion 2794b);

단기 개인 메모리 컴포넌트(들)(2752)(마스터 버전(2752b) 및 캐시(2752a)를 포함할 수 있음);

Short term private memory component (s) 2752 (may include master version 2752b and cache 2752a);

장기 개인 메모리 컴포넌트(들)(2754)(마스터 버전(2754b) 및 캐시(2754a)를 포함할 수 있으며, 예를 들어, 개인 데이터베이스들(1058), 애플리케이션 선호도 및 사용 히스토리(1072) 등을 포함할 수 있음);

May include long term private memory component (s) 2754 (master version 2754b and cache 2754a), including, for example, personal databases 1058, application preferences and usage history 1072, Lt; / RTI >

도메인 모델 컴포넌트(들)(2756);

Domain model component (s) 2756;

어휘 컴포넌트(들)(2758)(완전한 어휘(2758b) 및 서브세트(2758a)를 포함할 수 있음);

Lexical component (s) 2758 (which may include a complete vocabulary 2758b and a subset 2758a);

언어 패턴 인식기(들) 컴포넌트(들)(2760)(전체 라이브러리(2760b) 및 서브세트(2760a)를 포함할 수 있음);

Language pattern recognizer (s) component (s) 2760 (which may include the entire library 2760b and subset 2760a);

언어 해석기 컴포넌트(들)(2770);

Language interpreter component (s) 2770;

도메인 엔티티 데이터베이스(들)(2772);

Domain entity database (s) 2772;

대화 플로우 프로세서 컴포넌트(들)(2780);

Conversation flow processor component (s) 2780;

서비스 편성 컴포넌트(들)(2782);

Service organization component (s) 2782;

서비스 컴포넌트(들)(2784);

Service component (s) 2784;

태스크 플로우 모델 컴포넌트(들)(2786);

Task flow model component (s) 2786;

대화 플로우 모델 컴포넌트(들)(2787);

Conversation flow model component (s) 2787;

서비스 모델 컴포넌트(들)(2788);

Service model component (s) 2788;

출력 프로세서 컴포넌트(들)(2790);

Output processor component (s) 2790;

특정 클라이언트/서버 기반 실시예들에서, 이 컴포넌트들 중 일부 또는 전부가 클라이언트(1304)와 서버(1340) 사이에 분산될 수 있다.In certain client / server based embodiments, some or all of these components may be distributed between client 1304 and server 1340.

일 실시예에서, 가상 비서(1002)는 예를 들어 터치스크린 입력, 키보드 입력, 음성 입력, 및/또는 이들의 임의의 조합을 포함하는, 임의의 적합한 입력 모달리티를 통해 사용자 입력(2704)을 수신한다. 본원에 더 상세하게 설명되는 바와 같이, 일 실시예에서, 비서(1002)는 이벤트 컨텍스트(2706) 및/또는 컨텍스트의 몇몇 다른 유형들 중 임의의 것을 포함할 수 있는 컨텍스트 정보(1000)도 수신한다.In one embodiment, virtual assistant 1002 receives user input 2704 via any suitable input modality, including, for example, touch screen input, keyboard input, voice input, and / do. As described in more detail herein, secretary 1002 also receives context information 1000, which may include event context 2706 and / or any of several other types of contexts .

본원에 설명된 기법들에 따라 사용자 입력(2704) 및 컨텍스트 정보(1000)를 처리하면, 가상 비서(1002)는 사용자에게 제시할 출력(2708)을 생성한다. 출력(2708)은 임의의 적합한 출력 모달리티에 따라 생성될 수 있는데, 이는 컨텍스트(1000) 뿐 아니라, 적합하다면 다른 요소들에 의해 통지될 수 있다. 출력 모달리티들의 예들은, 스크린 상에 표시되는 시각적 출력, (음성 출력 및/또는 비프음들 및 다른 소리들을 포함할 수 있는) 청각적 출력, (진동과 같은) 햅틱 출력, 및/또는 이들의 임의의 조합을 포함한다.Upon processing user input 2704 and contextual information 1000 in accordance with the techniques described herein, virtual assistant 1002 generates output 2708 to be presented to the user. Output 2708 may be generated according to any suitable output modality, which may be notified by context 1000 as well as other factors as appropriate. Examples of output modalities include visual output displayed on the screen, auditory output (which may include audio output and / or beeps and other sounds), haptic output (such as vibration), and / . &Lt; / RTI >

도 27에 도시된 다양한 컴포넌트들의 동작에 관한 추가적인 세부 사항들은, 본원에 전체가 참조로서 통합된, 2011년 1월 10일에 출원된, 변호사 도켓 번호 P10575US1인, 관련된 미국 특허 출원 번호 12/987,982, "Intelligent Automated Assistant"에 제공된다.Additional details regarding the operation of the various components shown in FIG. 27 may be found in related U.S. Patent Application Nos. 12 / 987,982, filed January 10, 2011, Attorney Docket No. P10575US1, It is provided in the "Intelligent Automated Assistant".

컨텍스트Context

위에서 설명한 바와 같이, 일 실시예에서, 가상 비서(1002)는 정보 처리 기능들을 수행하기 위해 다양한 컨텍스트 정보를 획득하고 적용한다. 다음의 설명이 제시된다:As described above, in one embodiment, virtual secretary 1002 acquires and applies various contextual information to perform information processing functions. The following description is presented:

가상 비서(1002)에 의해 사용될 컨텍스트 정보의 소스들의 범위

A range of sources of context information to be used by virtual secretary 1002

컨텍스트 정보를 나타내고, 조직하고, 탐색하기 위한 기법들

Techniques for representing, organizing, and navigating context information

컨텍스트 정보가 가상 비서의 몇몇 기능들의 동작을 지원할 수 있는 방법들

Methods in which the context information can support the operation of some functions of the virtual assistant

분산된 시스템 내에서 컨텍스트 정보를 효과적으로 획득하고, 액세스하고, 적용하기 위한 방법들

Methods for effectively acquiring, accessing, and applying context information within a distributed system

본 기술분야의 당업자는 컨텍스트 정보를 사용하기 위한 소스들, 기법들 및 방법들에 관한 뒤따르는 설명은 단지 예시적일 뿐이며, 본 발명의 필수적 특성들로부터 벗어나지 않으면서 다른 소스들, 기법들 및 방법들이 사용될 수 있다는 것을 인식할 것이다.Those skilled in the art will appreciate that the following description of sources, techniques, and methods for using context information is exemplary only and that other sources, techniques, and methods may be used without departing from the essential characteristics of the present invention As will be appreciated by those skilled in the art.

컨텍스트의Contextual 소스들 Sources

가상 비서(1002)에 의해 수행되는 정보 처리의 단계들을 통해, 사용자 입력의 가능한 해석들을 감소시키기 위해 몇몇 상이한 종류들의 컨텍스트가 사용될 수 있다. 예들은 애플리케이션 컨텍스트, 개인 데이터 컨텍스트, 및 이전 대화 히스토리을 포함한다. 본 기술분야의 당업자는, 컨텍스트의 다른 소스들도 이용 가능할 수 있다는 것을 인식할 것이다.Through the steps of information processing performed by virtual secretary 1002, several different kinds of contexts can be used to reduce possible interpretations of user input. Examples include an application context, a personal data context, and a previous conversation history. Those skilled in the art will recognize that other sources of context may also be available.

이제 도 1을 참조하면, 일 실시예에 따른 가상 비서(1002), 및 그것의 동작에 영향을 줄 수 있는 컨텍스트의 소스들의 일부 예들을 도시하는 블록도가 도시되었다. 가상 비서(1002)는 음성 또는 타이핑된 언어와 같은 사용자 입력(2704)을 취하고, 사용자에게 출력(2708)을 생성하고 그리고/또는 사용자를 대신하여 액션들을 수행한다(2710). 도 1에 도시된 가상 비서(1002)는 구현될 수 있는 광범위한 가상 비서 시스템 실시예들 중 단지 하나의 예일 뿐이다. 가상 비서 시스템의 다른 실시예들(도시되지 않음)은, 도 1에 도시된 가상 비서(1002)의 예에 예시된 것에 비해 추가적인, 더 적은 및/또는 상이한 컴포넌트들/특징들을 포함할 수 있다.Referring now to FIG. 1, there is shown a block diagram illustrating some examples of sources of context that may affect the operation of virtual assistant 1002, and in accordance with one embodiment. The virtual assistant 1002 takes user input 2704, such as a voice or typed language, creates an output 2708 to the user, and / or performs actions on behalf of the user (2710). The virtual secretary 1002 shown in Figure 1 is only one example of a wide range of virtual secretarial system embodiments that can be implemented. Other embodiments (not shown) of the virtual assistant system may include additional, fewer and / or different components / features than those illustrated in the example of virtual assistant 1002 shown in FIG.

본원에 더 상세히 설명되는 것과 같이, 가상 비서(1002)는 사전들, 도메인 모델들 및/또는 태스크 모델들과 같은, 지식 및 데이터의 다수의 상이한 소스들 중 임의의 것에 접근할 수 있다. 본 발명의 관점에서, 백그라운드 소스라 지칭되는 그러한 소스들은 비서(1002) 내부에 있다. 사용자 입력(2704) 및 백그라운드 소스들 외에, 가상 비서(1002)는 또한 예를 들어 디바이스 센서 데이터(1056), 애플리케이션 선호도 및 사용 히스토리(1072), 대화 히스토리 및 보조 메모리(1052), 개인 데이터베이스(1058), 개인 음향 컨텍스트 데이터(1080), 현재 애플리케이션 컨텍스트(1060), 및 이벤트 컨텍스트(2706)을 포함하는, 컨텍스트의 몇몇 소스들로부터의 정보에 접근할 수 있다. 이는 본원에 상세하게 설명될 것이다.As described in more detail herein, the virtual assistant 1002 can access any of a number of different sources of knowledge and data, such as dictionaries, domain models and / or task models. In the context of the present invention, such sources, referred to as background sources, are inside the secretary 1002. In addition to user input 2704 and background sources, virtual assistant 1002 may also include device sensor data 1056, application preferences and usage history 1072, conversation history and auxiliary memory 1052, personal database 1058 ), Personal audio context data 1080, current application context 1060, and event context 2706. In some embodiments of the present invention, This will be described in detail herein.

애플리케이션 Application 컨텍스트Context (1060)(1060)

애플리케이션 컨텍스트(1060)는, 사용자가 무언가를 하고 있는 애플리케이션 또는 유사한 소프트웨어의 상태를 가리킨다. 예를 들면, 사용자는 텍스트 메시징 애플리케이션을 사용하여 특정한 사람과 채팅을 할 수 있다. 가상 비서(1002)는, 텍스트 메시징 애플리케이션의 사용자 인터페이스에 특정되거나 사용자 인터페이스의 일부일 필요는 없다. 오히려, 가상 비서(1002)는 임의의 수의 애플리케이션들로부터 컨텍스트를 수신할 수 있고, 각 애플리케이션은 자신의 컨텍스트를 건네 주어 가상 비서(1002)에게 알릴 수 있다.Application context 1060 indicates the status of an application or similar software in which the user is doing something. For example, a user can use a text messaging application to chat with a particular person. The virtual assistant 1002 need not be specific to or part of the user interface of the text messaging application. Rather, the virtual secretary 1002 can receive the context from any number of applications, and each application can pass on its context and notify the virtual secretary 1002.

가상 비서(1002)가 호출된 경우에 사용자가 현재 애플리케이션을 사용하고 있으면, 그 애플리케이션의 상태는 유용한 컨텍스트 정보를 제공할 수 있다. 예를 들면, 가상 비서(1002)가 이메일 애플리케이션 내로부터 호출된 경우, 컨텍스트 정보는 송신자 정보, 수신자 정보, 보낸 날짜 및/또는 시간, 제목, 이메일 컨텐트로부터 추출된 데이터, 메일박스 또는 폴더 이름 등을 포함할 수 있다.If the user is currently using the application when the virtual assistant 1002 is invoked, the state of the application may provide useful context information. For example, when virtual assistant 1002 is invoked from within an email application, the context information may include sender information, recipient information, send date and / or time, title, data extracted from email content, mailbox or folder name, .

도 11 내지 도 13을 참조하면, 일 실시예에 따라, 텍스트 메시징 도메인에서 애플리케이션 컨텍스트를 사용하여, 대명사가 지시하는 대상을 도출하는 예를 도시하는 한 세트의 스크린 샷들이 도시된다. 도 11은, 사용자가 테스트 메시징 애플리케이션을 사용하는 동안 디스플레이될 수 있는 스크린(1150)을 도시한다. 도 12는, 가상 비서(1002)가 텍스트 메시징 애플리케이션의 컨텍스트에서 활성화된 후의 스크린(1250)을 도시한다. 이 예에서, 가상 비서(1002)는 사용자에게 프롬프트(1251)를 제시한다. 일 실시예에서, 사용자는 마이크로폰 아이콘(1252)를 탭핑함으로써 음성 입력(spoken input)을 제공할 수 있다. 다른 실시예에서, 비서(1002)는 언제라도 음성 입력을 받아들일 수 있고, 입력을 제공하기 전에 사용자에게 마이크로폰 아이콘(1252)을 탭핑할 것을 요구하지 않는다; 따라서, 아이콘(1252)은 비서(1002)가 음성 입력을 대기중이라는 리마인더일 수 있다.Referring to Figures 11-13, a set of screen shots is shown illustrating an example of deriving an object to which a pronoun is directed using an application context in a text messaging domain, according to one embodiment. 11 shows a screen 1150 that can be displayed while the user is using a test messaging application. Figure 12 shows a screen 1250 after the virtual secretary 1002 is activated in the context of a text messaging application. In this example, the virtual secretary 1002 presents a prompt 1251 to the user. In one embodiment, a user may provide a spoken input by tapping the microphone icon 1252. In another embodiment, the secretary 1002 can accept voice input at any time and does not require the user to tap the microphone icon 1252 before providing input; Thus, the icon 1252 may be a reminder that the secretary 1002 is waiting for a voice input.

도 13에서, 스크린(1253)에 도시된 바와 같이, 사용자는 가상 비서(1002)와 대화를 하고 있다. 사용자가 말한 입력 "call him"이 반향되었고(echoed back), 가상 비서(1002)는 특정 전화 번호의 특정 사람에게 전화를 걸 것이라고 응답하고 있다. 사용자의 애매한 입력을 해석하기 위해, 가상 비서(1002)는, 본원에서 더 자세히 기술되는 바와 같이, 다양한 소스의 컨텍스트를 결합해 사용하여 대명사가 지시하는 대상을 도출한다.In FIG. 13, as shown on screen 1253, the user is in conversation with the virtual secretary 1002. The input "call him" said by the user has been echoed back and the virtual secretary 1002 is responding that he will call a particular person at a particular telephone number. To interpret the ambiguous input of the user, the virtual assistant 1002 combines the contexts of the various sources, as described in more detail herein, to derive the subject to which the pronoun is pointing.

도 17 내지 도 20을 참조하면, 일 실시예에 따라, 현재 애플리케이션 컨텍스트를 사용하여 커맨드를 해석하고 조작할 수 있게 하는 다른 예를 도시한다.Referring to Figures 17-20, there is shown another example of enabling a command to be interpreted and manipulated using the current application context, according to one embodiment.

도 17에서, 사용자에게 사용자의 이메일 인박스(1750)가 제시되고, 사용자는 특정 이메일 메시지(1751)를 보기 위해 선택한다. 도 18은, 보려는 이메일 메시지가 선택된 후의 이메일 메시지(1751)를 도시한다; 이 예에서, 이메일 메시지(1751)는 이미지를 포함한다. 17, a user's email inbox 1750 is presented to the user, and the user selects to view a particular email message 1751. [ 18 shows an email message 1751 after an email message to view is selected; In this example, the email message 1751 includes an image.

도 19에서, 사용자는, 이메일 애플리케이션 내로부터의 이메일 메시지(1751)를 보는 동안 가상 비서(1002)를 활성화시켰다. 일 실시예에서, 이메일 메시지(1751)가 스크린 위쪽으로 옮겨져서 디스플레이되어, 가상 비서(1002)로부터의 프롬프트(150)를 위한 공간을 만든다. 이러한 디스플레이는, 가상 비서(1002)가 현재 보여지는 이메일 메시지(1751)의 컨텍스트에서 보조를 제공하고 있다는 관념을 강화시킨다. 따라서, 가상 비서(1002)에 대한 사용자의 입력은 보여지는 이메일 메시지(1751)의 현재 컨텍스트에서 해석될 것이다.In FIG. 19, the user activated the virtual secretary 1002 while viewing an email message 1751 from within the email application. In one embodiment, an email message 1751 is moved and displayed above the screen to create a space for the prompt 150 from the virtual secretary 1002. This display enhances the notion that the virtual assistant 1002 is providing assistance in the context of the email message 1751 currently being viewed. Thus, the user's input to the virtual assistant 1002 will be interpreted in the current context of the email message 1751 shown.

도 20에서, 사용자는 커맨드(2050)를 제공했다: "Reply let's get this to marketing right away". 이메일 메시지(1751) 및 이메일 메시지가 디스플레이되는 이메일 애플리케이션에 관한 정보를 포함하여, 컨텍스트 정보가 커맨드(2050)를 해석하는데 사용된다. 이 컨텍스트는 커맨드(2050)의 "reply"와 "this"란 단어의 의미를 판정하고, 특정 메시지 스레드 상의 특정 수신인에 대한 이메일 구성 트랜잭션을 어떻게 설정해야 할지 결정하는데 사용될 수 있다. 이 경우에, 가상 비서(1002)는 컨텍스트 정보에 액세스하여, "marketing"이 "John Applecore"라는 이름의 수신인을 가리킨다는 것을 판정할 수 있고, 수신인에 대해 사용할 이메일 어드레스를 판정할 수 있다. 따라서, 가상 비서(1002)는 사용자가 승인하고 전송할 이메일(2052)을 구성한다. 이러한 방식으로, 가상 비서(1002)는, 현재 애플리케이션의 상태를 기술하는 컨텍스트 정보와 함께 사용자 입력에 기초하여, (이메일 메시지를 구성하는) 작업을 조작할 수 있게 할 수 있다.In FIG. 20, the user has provided a command 2050: "Reply let's get this to marketing right away." Context information is used to interpret the command 2050, including the email message 1751 and information about the email application for which the email message is displayed. This context can be used to determine the meaning of the words "reply" and "this " of the command 2050 and to determine how to configure an email configuration transaction for a particular recipient on a particular message thread. In this case, the virtual assistant 1002 can access the context information to determine that "marketing" refers to a recipient named "John Applecore" and can determine the email address to use for the recipient. Thus, the virtual secretary 1002 configures the email 2052 to be approved and sent by the user. In this manner, the virtual assistant 1002 may be able to manipulate tasks (constituting an email message) based on user input with context information describing the current application state.

애플리케이션 컨텍스트는 또한, 애플리케이션들에 걸쳐 사용자 의도의 의미를 식별할 수 있게 도움을 줄 수 있다. 도 21을 참조하면, 보고 있는 이메일(예컨대, 이메일 메시지(1751))의 컨텍스트에서 사용자가 가상 비서(1002)를 호출하고, 사용자의 커멘드(2150)는 "Send him a text..."라고 말한 예를 도시한다. 커멘드(2150)는, 가상 비서(1002)에 의해 이메일보다 텍스트 메시지가 송신되야 한다는 것을 나타내는 것으로서 해석된다. 그러나, "him"이라는 단어의 사용은 동일한 수신인(John Appleseed)이 의도된다는 것을 가리킨다. 따라서, 가상 비서(1002)는, 이 통신이 이러한 수신인에게 상이한 채널(디바이스에 저장된 연락처 정보로부터 획득된 사람의 전화 번호로의 텍스트 메시지)로 가야 한다는 것으로 인식한다. 따라서, 가상 비서(1002)는 사용자가 승인하고 전송할 텍스트 메시지(2152)를 구성한다.The application context can also help to identify the meaning of the user's intent across applications. 21, a user invokes a virtual assistant 1002 in the context of a viewing email (e.g., email message 1751), and the user's command 2150 calls the "Send him a text ..." Fig. Command 2150 is interpreted by virtual correspondent 1002 as indicating that a text message should be sent rather than e-mail. However, the use of the word "him " indicates that the same recipient (John Appleseed) is intended. Thus, the virtual assistant 1002 recognizes that this communication should go to this recipient on a different channel (a text message to the telephone number of the person obtained from the contact information stored in the device). Thus, the virtual secretary 1002 constructs a text message 2152 to be accepted and transmitted by the user.

애플리케이션(들)로부터 획득될 수 있는 컨텍스트 정보의 예들에는 아래의 것이 포함되지만 이들로만 제한되지는 않는다:Examples of context information that can be obtained from the application (s) include, but are not limited to:

애플리케이션의 아이덴티티;

The identity of the application;

현재 이메일 메시지, 현재 음악 또는 재생목록 또는 재생되는 채널, 현재 책 또는 영화 또는 사진, 현재 캘린더 일/주/달, 현재 리마인더 리스트, 현재 전화 통화, 현재 텍스트 메시징 대화, 현재 지도 위치, 현재 웹 페이지 또는 검색 쿼리, 위치 감지 애플리케이션들을 위한 현재 도시 또는 다른 위치, 현재 소셜 네트워크 프로필, 또는 현재 객체들의 임의의 다른 애플리케이션 특정 개념;

Current email message, current music or playlist or channel being played, current book or movie or photo, current calendar day / week / month, current reminder list, current phone call, current text messaging conversation, current map location, A search query, a current city or other location for location sensing applications, a current social network profile, or any other application specific concept of current objects;

이름, 장소, 날짜, 및 현재 객체들로부터 추출될 수 있는 식별가능한 다른 엔티티 또는 값.

Name, place, date, and other identifiable entities or values that can be extracted from the current objects.

개인 데이터베이스들(1058)Personal databases (1058)

컨텍스트 데이터의 다른 소스는 전화기와 같은 디바이스에 있는 사용자의 개인 데이터베이스(들)(1058)이며, 그 예로는 이름들과 전화 번호들을 포함한 어드레스 북이다. 도 14를 참조하면, 일 실시예에 따라, 애매한 이름에 대해 가상 비서(1002)가 프롬프팅하는 스크린 샷(1451)의 예를 도시한다. 여기에서, 사용자는 "Call Herb"라고 말했다; 가상 비서(1002)는 사용자의 어드레스 북에서 일치하는 연락처들 중에서 사용자가 선택하도록 프롬프트한다. 따라서, 어드레스 북은 개인 데이터 컨텍스트의 소스로서 사용된다.Another source of context data is the user's personal database (s) 1058 on the device, such as a telephone, for example an address book containing names and telephone numbers. Referring to Fig. 14, there is shown an example of a screenshot 1451 in which the virtual secretary 1002 prompts for an ambiguous name, according to one embodiment. Here, the user said "Call Herb"; The virtual assistant 1002 prompts the user to select from matching contacts in the address book of the user. Thus, the address book is used as the source of the personal data context.

일 실시예에서, 사용자의 개인 정보는, 사용자의 의도 또는 가상 비서(1002)의 다른 기능들을 해석 및/또는 조작할 수 있게 하기 위한 컨텍스트로서 사용하기 위해 개인 데이터베이스(1058)로부터 획득된다. 예를 들면, 사용자의 연락처 데이터베이스 내의 데이터는, 사용자가 누군가를 이름만으로 언급했을 때, 사용자의 커맨드를 해석할 때의 애매함을 감소시키기 위해 사용될 수 있다. 개인 데이터베이스들(1058)로부터 획득될 수 있는 컨텍스트 정보의 예들에는 아래의 것이 포함되지만 이들로만 제한되지는 않는다:In one embodiment, the user's personal information is obtained from the personal database 1058 for use as a context to enable the user's intention or other functions of the virtual assistant 1002 to be interpreted and / or manipulated. For example, the data in the user's contact database can be used to reduce ambiguity in interpreting the user's commands when the user mentions someone by name alone. Examples of context information that may be obtained from personal databases 1058 include, but are not limited to, the following:

사용자의 연락처 데이터베이스(어드레스 북) -- 이름, 전화 번호, 물리 어드레스, 네트워크 어드레스, 계정 식별자, 중요한 날짜에 관한 정보가 포함됨 -- 사람, 회사, 조직, 위치, 웹 사이트, 및 사용자가 참조할 수 있는 다른 엔티티에 대한 것;

User's contact database (address book) - Includes information about name, phone number, physical address, network address, account identifier, and important date. - User, company, organization, location, It's for another entity;

사용자 자신의 이름, 선호되는 발음, 어드레스, 전화 번호, 등;

User's own name, preferred pronunciation, address, telephone number, etc.;

어머니, 아버지, 자매, 상사 등과 같은 사용자가 명명한 관계;

A relationship named by a user such as a mother, father, sister, or boss;

캘린더 이벤트, 특별한 날의 명칭, 사용자가 참조할 수 있는 명명된 다른 엔트리들을 포함하는, 사용자의 캘린더 데이터;

A calendar event of a user, a name of a special day, user's calendar data including other named entries that the user can refer to;

해야 하거나 기억해야 하거나 얻어야 할 것의 목록을 포함하여 사용자가 참조할 수 있는 사용자의 리마인더 또는 작업 목록;

A list of reminders or tasks of the user that the user can reference, including a list of what to do or need to remember or get;

노래, 장르, 재생 목록, 및 사용자가 참조할 수 있는 사용자의 뮤직 라이브러리와 연관된 다른 데이터;

Songs, genres, playlists, and other data associated with the user's music library that the user can refer to;

사람, 장소, 카테고리, 태그, 레이블, 또는 사용자의 미디어 라이브러리 내의 포토 또는 비디오 또는 다른 미디어 상의 다른 상징적 이름;

A person, place, category, tag, label, or other symbolic name on a photo or video or other media in your media library;

제목, 저자, 장르, 또는 사용자의 개인 라이브러리에 있는 다른 문헌 또는 어드레스 북의 다른 상징적 이름 .

Title, author, genre, or other symbolic name of another document or address book in your personal library.

대화 Conversation 히스토리history (1052)(1052)

컨텍스트 데이터의 다른 소스는 가상 비서(1002)와 사용자의 대화 히스토리(1052)이다. 이러한 히스토리에는, 예를 들어, 도메인, 사람, 장소 등에 대한 참조가 포함될 수 있다. 도 15를 참조하면, 일 실시예에 따라, 가상 비서(1002)가 대화 컨텍스트를 사용하여 커멘드에 대한 위치를 추론하는 예를 도시한다. 스크린(1551)에서, 사용자는 먼저 "What's the time in New York"이라고 묻는다; 가상 비서(1002)는 뉴욕 시의 현재 시간을 제공함으로써 응답한다(1552). 이후에 사용자는 "What's the weather"이라고 묻는다. 가상 비서(1002)는 이전의 대화 히스토리를 사용하여, 날씨 쿼리가 의도한 위치가 대화 히스토리에서 언급된 마지막 위치인 것으로 추론한다. 따라서, 그 응답(1553)은, 뉴욕 시에 대한 날씨 정보를 제공한다. Another source of context data is the virtual secretary 1002 and the user's conversation history 1052. Such a history may include, for example, references to domains, people, places, and the like. Referring to FIG. 15, there is shown an example in which the virtual secretary 1002 uses a conversation context to deduce a location for a command, according to one embodiment. On screen 1551, the user first asks "What's the time in New York "; The virtual secretary 1002 responds by providing the current time in New York City (1552). Later, the user asks "What's the weather?" The virtual secretary 1002 uses the previous conversation history to infer that the intended location of the weather query is the last location mentioned in the conversation history. Accordingly, the response 1553 provides weather information for New York City.

다른 예로서, 사용자가 "find camera shops near here"이라고 말하면, 결과들을 조사한 후에, "how about in San Francisco?"라고 말하고, 비서는 대화 컨텍스트를 사용하여 "how about"이 의미하는 바가 "do the same task(find camera stores)"라고 판정하고, "in San Francisco"가 의미하는 바가 "changing the locus of the search from here to San Francisco"라고 판정한다. 가상 비서(1002)는 또한, 사용자에게 제공된 이전의 출력과 같은, 이전의 대화의 세부내용들을 컨텍스트로 사용할 수 있다. 예를 들면, 가상 비서(1002)가 "Sure thing, you're the boss"와 같이 유머로 의도된 재치 있는 응답을 사용했다면, 가상 비서는 이러한 응답을 이미 말했다는 것을 기억하고, 대화 세션 내에서 그러한 문구를 반복하지 않게 할 수 있다. As another example, if a user says "find camera shops near here," after examining the results, say "how about in San Francisco?" And the secretary uses the conversation context to say "how about" the same task (find camera stores) "and the word" in San Francisco "means" changing the locus of the search from here to San Francisco ". The virtual secretary 1002 may also use context details of the previous conversation, such as the previous output provided to the user. For example, if the virtual assistant 1002 used a humorous response intended to be humorous, such as "Sure thing, you're the boss ", the virtual assistant remembers this response already, You can avoid repeating those phrases.

대화 히스토리 및 가상 보조 메모리로부터의 컨텍스트 정보의 예로는,Examples of context information from the conversation history and virtual auxiliary memory include,

대화에서 언급된 사람들;

People mentioned in the conversation;

대화에서 언급된 장소들 및 위치들;

Places and locations mentioned in the conversation;

포커스되는 현재의 시간 프레임;

The current time frame being focused;

이메일이나 캘린더와 같이, 포커스되는 현재의 애플리케이션 도메인;

The current application domain being focused, such as an email or calendar;

이메일을 읽는 것 또는 캘린더 엔트리를 생성하는 것과 같이, 포커스되는 현재의 태스크;

The current task being focused, such as reading an email or creating a calendar entry;

막 읽은 이메일 메시지 또는 막 생성한 캘린더 엔트리와 같이, 포커스되는 현재의 도메인 객체들;

The current domain objects being focused, such as the email message just read or the calendar entry just created;

질문을 하고 있는지 및 어떤 가능한 대답들이 예상되는지와 같이, 대화 또는 트랜잭션 플로우의 현재의 상태;

The current state of the conversation or transaction flow, such as whether it is asking questions and what possible responses are expected;

"good Italian restaurants"와 같이, 사용자 요청들의 히스토리;

the history of user requests, such as "good Italian restaurants ";

리턴된 레스토랑 세트와 같이, 사용자 요청들의 결과들의 히스토리;

A history of the results of user requests, such as returned restaurant sets;

대화에서 비서에 의해 이용된 어구들의 히스토리;

The history of the phrases used by the secretary in the conversation;

"my mother is Rebecca Richards" 및 "I liked that restaurant"와 같이, 사용자가 비서에게 말한 사실들

Facts the user told the secretary, such as "my mother is Rebecca Richards" and "I liked that restaurant"

이 포함되지만, 이에 제한되지는 않는다.But are not limited thereto.

이하 도 25a 및 도 25b를 참조하면, 일 실시예에 따른 이전의 대화 컨텍스트의 이용의 예를 도시한 일련의 스크린샷들이 도시되어 있다. 도 25a에서, 사용자는 John으로부터의 임의의 신규 이메일에 대한 요청(2550)을 입력하였다. 가상 비서(1002)는 John으로부터의 이메일 메시지(2551)를 디스플레이함으로써 응답한다. 도 25b에서, 사용자는 커맨드(2552) "Reply let's get this to marketing right away"를 입력한다. 가상 비서(1002)는 이전의 대화 컨텍스트를 이용하여 커맨드(2552)를 해석하는데; 상세하게는, 이 커맨드는 도 25에 디스플레이된 이메일 메시지(2551)을 언급하는 것으로 해석된다.Referring now to Figures 25A and 25B, there is shown a series of screen shots illustrating an example of the use of a prior conversation context in accordance with one embodiment. 25A, the user has entered a request 2550 for any new email from John. The virtual secretary 1002 responds by displaying an email message 2551 from John. In Fig. 25B, the user inputs a command 2552 "Reply let's get this to marketing right away ". The virtual secretary 1002 interprets the command 2552 using the previous conversation context; In particular, this command is interpreted to refer to the email message 2551 displayed in Fig.

디바이스device 센서 데이터(1056) Sensor data (1056)

일 실시예에 있어서, 가상 비서(1002)를 실행하는 물리 디바이스는 하나 이상의 센서를 가질 수 있다. 이러한 센서들은 컨텍스트 정보의 소스들을 제공할 수 있다. 이러한 정보의 예로는,In one embodiment, the physical device executing the virtual secretary 1002 may have more than one sensor. These sensors may provide sources of context information. As an example of such information,

사용자의 현재 위치;

Your current location;

사용자의 현재 위치에서의 현지 시간;

Local time at the user's current location;

디바이스의 위치, 방위 및 모션;

Device position, orientation and motion;

현재의 광 레벨; 온도 및 다른 환경 척도;

Current light level; Temperature and other environmental measures;

사용 중인 마이크로폰들 및 카메라들의 속성들;

Properties of the microphones and cameras in use;

이용되고 있는 현재의 네트워크들 및 접속된 네트워크들의 서명들 - 이 네트워크는 이더넷, Wi-Fi 및 블루투스를 포함함 -

Current networks being used and signatures of connected networks - This network includes Ethernet, Wi-Fi and Bluetooth -

이 포함되지만, 이에 제한되지는 않는다. 서명들은 네트워크 액세스 포인트들의 MAC 어드레스들, 할당된 IP 어드레스들, 블루투스 네임들과 같은 디바이스 식별자들, 주파수 채널들 및 무선 네트워크들의 다른 속성들을 포함한다.But are not limited thereto. Signatures include MAC addresses of network access points, assigned IP addresses, device identifiers such as Bluetooth names, frequency channels, and other attributes of wireless networks.

센서들은, 예를 들어 가속도계, 나침반, GPS 유닛, 고도 검출기, 광 센서, 온도계, 기압계, 클록, 네트워크 인터페이스, 배터리 테스트 회로 등을 포함한 임의의 타입의 센서일 수 있다.The sensors may be any type of sensor including, for example, an accelerometer, a compass, a GPS unit, an altitude detector, an optical sensor, a thermometer, a barometer, a clock, a network interface,

애플리케이션 선호도 및 사용 Application affinity and usage 히스토리history (1072)(1072)

일 실시예에 있어서, 다양한 애플리케이션들에 대한 사용자의 선호도들 및 설정들뿐만 아니라 사용자의 사용 히스토리(1072)를 기술하는 정보가 사용자의 의도 또는 가상 비서(1002)의 다른 기능을 해석하고/하거나 조작할 수 있게 하기 위해 컨텍스트로서 이용된다. 이러한 선호도 및 히스토리(1072)의 예로는,In one embodiment, the user's preferences and settings for various applications, as well as information describing the user's usage history 1072 may be used to interpret and / or manipulate the user's intent or other function of the virtual secretary 1002 And is used as a context in order to be able to do so. As an example of such preferences and history 1072,

단축키, 즐겨찾기, 북마크, 친구 리스트, 또는 사람들, 회사들, 어드레스들, 전화 번호들, 장소들, 웹 사이트들, 이메일 메시지들 또는 임의의 다른 참조들에 대한 사용자 데이터의 임의의 다른 집합;

Any other set of user data for shortcuts, favorites, bookmarks, friend lists, or people, companies, addresses, phone numbers, places, websites, email messages or any other references;

디바이스 상에서 이루어진 최근의 통화들;

Recent calls made on the device;

대화 상대방들을 포함하는 최근의 텍스트 메시지 대화들;

Recent text message conversations involving conversation partners;

지도 또는 방향에 대한 최근의 요청들;

Recent requests for maps or directions;

최근의 웹 검색들 및 URL들;

Recent web searches and URLs;

주식 애플리케이션에 열거된 주식들;

Stocks listed in stock applications;

재생되는 최근의 노래나 비디오 또는 다른 미디어;

Recent songs, videos or other media being played;

경보 애플리케이션들 상에 설정된 알람의 명칭들;

The names of the alarms set on the alarm applications;

디바이스 상의 애플리케이션들 또는 다른 디지털 객체들의 명칭들;

The names of applications or other digital objects on the device;

사용자의 선호 언어 또는 사용자의 위치에서 사용 중인 언어

The language you are using in your preferred language or user's location

가 포함되지만, 이에 제한되지는 않는다.But is not limited thereto.

이하 도 16을 참조하면, 일 실시예에 따라 컨텍스트의 소스로서 전화기 즐겨찾기 리스트의 이용의 일례가 도시되어 있다. 스크린(1650)에서, 즐겨찾기 연락처들(1651)의 리스트가 도시되어 있다. 사용자가 "call John"에 대한 입력을 제공하는 경우, 이러한 즐겨찾기 연락처들(1651)의 리스트가 이용되어, "John"이 John Appleseed의 모바일 번호를 언급한다고 결정할 수 있는데, 그 이유는 그 번호가 이 리스트에 나타나기 때문이다.Referring now to Figure 16, an example of the use of a telephone book favorites list as a source of context is illustrated in accordance with one embodiment. On screen 1650, a list of favorite contacts 1651 is shown. If the user provides an input for "call John ", a list of these favorite contacts 1651 may be used to determine that" John " refers to the mobile number of John Appleseed, This is because it appears in this list.

이벤트 event 컨텍스트Context (2706)(2706)

일 실시예에 있어서, 가상 비서(1002)는 가상 비서(1002)와 사용자의 상호작용에 독립적으로 일어나는 비동기 이벤트들과 연관된 컨텍스트를 이용할 수 있다. 이하 도 22 내지 도 24를 참조하면, 일 실시예에 따라 이벤트 컨텍스트 또는 경보 컨텍스트를 제공할 수 있는 이벤트가 발생한 이후의 가상 비서(1002)의 기동을 나타내는 일례가 도시되어 있다. 이 경우, 이벤트는 도 22에 도시된 바와 같이 착신 텍스트 메시지(2250)이다. 도 23에서, 가상 비서(1002)가 호출되었고, 프롬프트(1251)와 함께 텍스트 메시지(2250)가 나타난다. 도 24에서, 사용자는 커맨드 "call him"(2450)을 입력하였다. 가상 비서(1002)는 이벤트 컨텍스트를 이용하여, 착신 텍스트 메시지(2250)를 송신한 사람을 의미하는 것으로 "him"을 해석함으로써 커맨드를 명확하게 한다. 가상 비서(1002)는 발신 통화를 위해 어떤 전화 번호를 이용할지를 결정하기 위해서 이벤트 컨텍스트를 또한 이용한다. 전화를 걸고 있다는 것을 나타내기 위해서 확인 메시지(2451)가 디스플레이된다.In one embodiment, the virtual secretary 1002 may utilize the context associated with the asynchronous events that occur independently of the interaction of the user with the virtual secretary 1002. [ 22 to 24, there is shown an example of the activation of the virtual secretary 1002 after an event that can provide an event context or an alert context according to an embodiment. In this case, the event is an incoming text message 2250, as shown in FIG. 23, the virtual secretary 1002 is invoked and a text message 2250 appears with a prompt 1251. 24, the user has entered the command "call him" (2450). The virtual secretary 1002 uses the event context to clarify the command by interpreting "him " to mean the person who sent the incoming text message 2250. [ Virtual assistant 1002 also uses the event context to determine which phone number to use for outgoing calls. A confirmation message 2451 is displayed to indicate that a call is being made.

경보 컨텍스트 정보의 예로는,Examples of alert context information include,

착신 텍스트 메시지들 또는 페이지들;

Incoming text messages or pages;

착신 이메일 메시지들;

Incoming email messages;

착신 전화 통화들;

Incoming phone calls;

리마인더 통지들 또는 태스크 경보들;

Reminder notifications or task alerts;

캘린더 경보들;

Calendar alerts;

알람 클록, 타이머 또는 다른 시간-기반 경보들;

Alarm clock, timer or other time-based alarms;

게임으로부터의 스코어나 다른 이벤트의 통지들;

Notifications of scores or other events from the game;

주가 경보와 같은 금융 이벤트의 통지들;

Notifications of financial events such as stock price warnings;

뉴스 플래시들 또는 다른 방송 통지들;

News flashes or other broadcast notifications;

임의의 애플리케이션으로부터의 푸시 통지들

Push notifications from any application

개인 음향 Personal acoustics 컨텍스트Context 데이터(1080) Data (1080)

음성 입력을 해석하는 경우, 가상 비서(1002)는 음성이 입력되는 음향 환경을 또한 고려할 수 있다. 예를 들어, 조용한 사무실의 잡음 프로파일은 자동차나 공공 장소의 잡음 프로파일과 상이하다. 음성 인식 시스템이 음향 프로파일 데이터를 식별 및 저장할 수 있는 경우, 이러한 데이터가 또한 컨텍스트 정보로서 제공될 수 있다. 사용 중인 마이크로폰의 속성, 현재 위치 및 현재의 대화 상태와 같은 다른 컨텍스트 정보와 결합되는 경우, 음향 컨텍스트는 입력의 인식 및 해석을 지원할 수 있다.In interpreting the speech input, the virtual assistant 1002 may also consider the acoustic environment into which the speech is input. For example, the noise profile of a quiet office is different from the noise profile of a car or public place. If the speech recognition system can identify and store the sound profile data, such data can also be provided as context information. The acoustic context may support recognition and interpretation of the input when combined with other context information, such as the properties of the microphone in use, the current location, and the current conversation state.

컨텍스트의Contextual 표현 및 액세스 Expression and Access

전술한 바와 같이, 가상 비서(1002)는 다수의 상이한 소스들 중 임의의 소스로부터의 컨텍스트 정보를 이용할 수 있다. 가상 비서(1002)에 이용가능해질 수 있도록 컨텍스트를 표현하기 위해 다수의 상이한 메커니즘 중 임의의 메커니즘이 이용될 수 있다. 이하 도 8a 내지 도 8d를 참조하면, 본 발명의 다양한 실시예와 관련하여 이용될 수 있는 바와 같은 컨텍스트 정보의 표현의 수개의 예가 도시되어 있다.As described above, the virtual secretary 1002 may use context information from any of a number of different sources. Any of a number of different mechanisms may be used to represent the context to be available to the virtual secretary 1002. [ Referring now to Figures 8A-8D, several examples of representations of contextual information as may be utilized in connection with various embodiments of the present invention are shown.

사람들, 장소들, People, places, 시간들Time , 도메인들, 태스크들 및 객체들의 표현, Representations of domains, tasks and objects

도 8a는 사용자의 현재 위치의 지리 좌표와 같은 단순한 속성을 나타내는 컨텍스트 변수들의 예들(801-809)을 도시한다. 일 실시예에 있어서, 컨텍스트 변수들의 코어 세트에 대해 현재의 값이 유지될 수 있다. 예를 들어, 현재의 사용자, 포커스되는 현재의 위치, 포커스되는 현재의 시간 프레임, 포커스되는 현재의 애플리케이션 도메인, 포커스되는 현재의 태스크 및 포커스되는 현재의 도메인 객체가 존재할 수 있다. 도 8a에 도시된 바와 같은 데이터 구조가 이러한 표현에 이용될 수 있다.8A shows examples of context variables 801-809 that represent simple attributes such as the geographical coordinates of the user's current location. In one embodiment, the current value may be maintained for the core set of context variables. For example, there may be a current user, the current location being focused, the current time frame being focused, the current application domain being focused, the current task being focused, and the current domain object being focused. A data structure as shown in Fig. 8A may be used for such representation.

도 8b는 연락처에 대한 컨텍스트 정보를 저장하는데 이용될 수 있는 보다 복잡한 표현의 예(850)를 도시한다. 또한, 연락처에 대한 데이터를 포함하는 표현의 예(851)도 도시되어 있다. 일 실시예에 있어서, 연락처(또는 사람)는 이름, 성별, 어드레스, 전화 번호에 대한 속성, 및 연락처 데이터베이스에 유지될 수 있는 다른 속성을 갖는 객체로서 표현될 수 있다. 장소, 시간, 애플리케이션 도메인, 태스크, 도메인 객체 등에 대해 유사한 표현이 이용될 수 있다.8B shows an example of a more complex representation 850 that may be used to store context information for a contact. Also shown is an example of an expression 851 that includes data for a contact. In one embodiment, the contact (or person) may be represented as an object having a name, a gender, an address, an attribute for a telephone number, and other attributes that can be maintained in the contact database. Similar expressions can be used for location, time, application domain, task, domain object, and so on.

일 실시예에 있어서, 주어진 타입의 현재 값들의 세트가 표현된다. 이러한 세트는 현재의 사람들, 현재의 장소들, 현재의 시간들 등을 언급할 수 있다.In one embodiment, a set of current values of a given type is represented. This set can refer to the current people, current locations, current times, and so on.

일 실시예에 있어서, 컨텍스트 값은 히스토리로 배열되어, 반복 N에서, 현재의 컨텍스트 값들의 프레임이 존재하며, 또한 반복 N-1에서 현재였던 컨텍스트 값의 프레임이 존재하는데, 이는 요구되는 히스토리의 길이에 대해 소정 제한을 한다. 도 8c는 컨텍스트 값들의 히스토리를 포함하는 어레이(811)의 예를 도시한다. 상세하게는, 도 8c의 각 열은 컨텍스트 변수를 나타내는데, 여기서 행들은 상이한 시간들에 대응한다.In one embodiment, the context values are arranged into a history such that in the iteration N, there is a frame of current context values, and there is also a frame of context values that was current in the iteration N-1, . 8C shows an example of an array 811 that includes a history of context values. In particular, each column in Figure 8C represents a context variable, where the rows correspond to different times.

일 실시예에 있어서, 타이핑된 컨텍스트 변수들의 세트는 도 8d에 도시된 바와 같이 히스토리로 배열된다. 이 예에서, 사람들을 언급하는 컨텍스트 변수들의 세트(861)는 장소들을 언급하는 컨텍스트 변수들의 다른 세트(871)와 함께 도시되어 있다. 따라서, 히스토리에서의 특정 시간에 대한 관련 컨텍스트 데이터가 검색 및 적용될 수 있다.In one embodiment, the set of typed context variables are arranged into a history as shown in Figure 8D. In this example, a set of context variables 861 referring to people is shown with a different set of context variables 871 referring to places. Thus, relevant context data for a particular time in the history can be retrieved and applied.

당업자는, 도 8a 내지 도 8d에 도시된 특정한 표현들은 단지 예시일 뿐이고, 컨텍스트를 나타내기 위한 많은 다른 메커니즘 및/또는 데이터 포맷이 이용될 수 있다는 것을 인식할 것이다. 예들은 다음을 포함한다:Those skilled in the art will appreciate that the specific representations shown in Figures 8A-8D are exemplary only and that many other mechanisms and / or data formats for representing the context may be used. Examples include:

일 실시예에서는, 가상 비서(1002)가 사용자에게 어떻게 어드레스하고 사용자의 집, 업무, 이동 전화 등을 참조할지를 알 수 있도록 시스템의 현재 사용자가 일부 특수한 방식으로 표현될 수 있다.

In one embodiment, the current user of the system may be represented in some special way so that the virtual secretary 1002 knows how to address and refer to the user's home, work, mobile phone, and the like.

일 실시예에서는, 가상 비서(1002)가 "나의 엄마" 또는 "나의 상사의 집"과 같은 참조(reference)들을 이해할 수 있도록 사람들 간의 관계가 표현될 수 있다.

In one embodiment, relationships between people can be expressed such that the virtual secretary 1002 can understand references such as "my mom" or "my boss's home. &Quot;

장소는, 이름, 거리 어드레스, 지리적 좌표 등과 같은 속성들을 갖는 객체로서 표현될 수 있다.

A place may be represented as an object having attributes such as name, street address, geographical coordinates, and so on.

시간은 (년, 월, 일, 시간, 분 또는 초 등의) 유니버셜 타임, 타임존 오프셋, 레졸루션을 포함하는 속성들을 갖는 객체로서 표현될 수 있다. 시간 객체는 또한 "오늘", "이번 주", "이번 (다가오는) 주말", "다음 주", "애니의 생일" 등과 같은 심볼 시간을 나타낼 수도 있다. 또한, 시간 객체는 시기들 또는 시점들을 나타낼 수 있다.

The time can be represented as an object having properties including universal time (such as year, month, day, hour, minute or second), time zone offset, resolution. The time object may also represent symbol times such as "today", "this week", "this coming weekend", "next week", "Annie's birthday" Also, a time object may represent periods or points in time.

이메일, 텍스트 메시징, 전화, 캘린더, 연락처, 사진, 비디오, 지도, 날씨, 리마인더, 시계, 웹 브라우저, 페이스북, 판도라 등과 같은 담화(discourse)의 도메인 또는 애플리케이션 또는 서비스를 나타내는 애플리케이션 도메인의 면에서 또한 컨텍스트가 제공될 수 있다. 현재의 도메인은 어떤 도메인에 초점이 맞추어져 있는지를 표시한다.

In terms of an application domain representing a domain or application or service such as email, text messaging, phone, calendar, contact, photo, video, map, weather, reminder, clock, web browser, facebook, Pandora, Context can be provided. The current domain indicates which domain is focused.

컨텍스트는 또한 도메인 내에서 수행하기 위한 하나 이상의 태스크, 또는 동작들을 정의할 수 있다. 예를 들어, 이메일 도메인 내에서는, 이메일 메시지 판독, 이메일 검색, 새로운 이메일 작성 등과 같은 태스크들이 있다.

The context may also define one or more tasks, or actions, to perform within the domain. For example, within an email domain, there are tasks such as reading email messages, searching for emails, creating new emails, and so on.

도메인 객체들은 다양한 도메인과 연관된 데이터 객체들이다. 예를 들면, 이메일 도메인은 이메일 메시지 상에서 동작하며, 캘린더 도메인은 캘린더 이벤트 상에서 동작한다.

Domain objects are data objects associated with various domains. For example, an email domain operates on an email message, and a calendar domain operates on a calendar event.

여기에 제공되는 설명을 목적으로, 컨텍스트 정보의 표현을 주어진 유형의 컨텍스트 변수로 언급한다. 예를 들면, 현재 사용자의 표현은 타입 사람(type person)의 컨텍스트 변수이다. For purposes of explanation provided herein, the expression of context information is referred to as a context variable of a given type. For example, the expression of the current user is a context variable of the type person.

컨텍스트Context 도출의 표현 Expression of derivation

일 실시예에서는, 컨텍스트 변수의 도출이 정보 처리에 사용될 수 있도록 명시적으로 표현된다. 컨텍스트 정보의 도출는 정보를 결론짓고 검색하게 되어 있는 간섭들의 소스 및/또는 셋트들을 특징으로 한다. 예를 들면, 도 8b에 도시되어 있는 개인(person) 컨텍스트 값(851)이 이벤트 컨텍스트(2706)로부터 획득된 텍스트 메시지 도메인 객체로부터 도출될 수 있다. 이러한 컨텍스트 값(851)의 소스가 표현될 수 있다.In one embodiment, derivation of the context variable is explicitly expressed so that it can be used for information processing. Derivation of the context information is characterized by a source and / or set of interferences that are to conclude and retrieve the information. For example, the person context value 851 shown in FIG. 8B may be derived from the text message domain object obtained from the event context 2706. The source of this context value 851 may be represented.

사용자 요청 및/또는 의도의 User requests and / or intentional 히스토리의History 표현 expression

일 실시예에서는, 사용자 요청의 히스토리가 저장될 수 있다. 일 실시예에서는, (자연어 처리로부터 도출된) 사용자 의도의 심오한 구조적 표현의 히스토리가 또한 저장될 수 있다. 이는 가상 비서(1002)로 하여금 이전에 해석된 입력의 컨텍스트에서 새로운 입력이 의미가 맞게끔 되게 해준다. 예를 들면, 사용자가 "뉴욕 날씨는 어때?"라고 물으면, 언어 해석기(2770)는 뉴욕의 위치를 언급하는 것으로 질문을 해석할 수 있다. 그러면, 사용자가 "이번 주는 어때?"라고 말하면, 가상 비서(1002)는 이전 해석을 참조하여, "어때"를 "날씨가 어때"를 의미하는 것으로 해석해야 하는 것으로 판단할 수 있다.In one embodiment, a history of user requests may be stored. In one embodiment, a history of profound structural representations of user intent (derived from natural language processing) can also be stored. This allows the virtual secretary 1002 to make the new input meaningful in the context of the previously interpreted input. For example, if the user asks "What is New York weather? &Quot;, the language interpreter 2770 can interpret the question by referring to the location of New York. Then, if the user says, "What about this week?", The virtual assistant 1002 can refer to the previous interpretation and determine that "what" should be interpreted as meaning "what is the weather?"

결과 result 히스토리의History 표현 expression

일 실시예에서는, 사용자 요청의 결과의 히스토리가 도메인 객체의 형태로 저장될 수 있다. 예를 들어, "좋은 이탈리아 레스토랑을 찾아주세요."라는 사용자의 요청에 대해, 레스토랑을 표시하는 도메인 객체들의 셋트가 반환될 수 있다. 그 후, 사용자가 "아밀리오스를 불러주세요."와 같은 커맨드를 입력하면, 가상 비서(1002)는, 검색 결과 내에서, 불러올 수 있는 모든 가능한 장소들 중에서 가장 작은 셋트인 아밀리오스라는 이름의 레스토랑에 대한 결과를 검색할 수 있다.In one embodiment, a history of the results of a user request may be stored in the form of a domain object. For example, for a user request of "Find a good Italian restaurant ", a set of domain objects representing a restaurant may be returned. Thereafter, when the user inputs a command such as "Please call Amilio", the virtual assistant 1002 will search for a restaurant named Amilio, the smallest of all possible places that can be recalled within the search results You can search for results for.

컨텍스트Context 변수의 지연된 바인딩 Delayed binding of variables

일 실시예에서는, 컨텍스트 변수가 필요에 따라(on demand) 검색되거나 유도되는 정보를 나타낼 수 있다. 예를 들어, 현재 위치를 나타내는 컨텍스트 변수가, 액세스시에, 디바이스로부터 현재 위치 데이터를 검색하는 API를 호출한 다음, 예를 들어, 거리 어드레스를 계산하기 위한 다른 처리를 할 수 있다. 그 컨텍스트 변수의 값은 캐싱 정책에 따라서, 특정 시간 기간 동안 유지될 수 있다.In one embodiment, the context variable may represent information that is retrieved or derived on demand. For example, a context variable representing the current location may, upon access, call the API to retrieve the current location data from the device and then do other processing to calculate the street address, for example. The value of the context variable may be maintained for a specific time period, depending on the caching policy.

컨텍스트Context 검색 Search

가상 비서(1002)는, 정보 처리 문제를 해결하기 위해, 관련 컨텍스트 정보를 검색하기 위한 많은 서로 다른 접근법 중 임의의 방법을 사용할 수 있다. 서로 다른 유형의 검색의 예는, 이에 제한되는 바는 아니지만, 다음을 포함한다:The virtual assistant 1002 may use any of a number of different approaches for retrieving relevant context information to solve information processing problems. Examples of different types of searches include, but are not limited to, the following:

컨텍스트 변수 이름에 의한 검색. "현재 사용자의 이름(first name)"과 같이, 요구된 컨텍스트 변수의 이름이 알려져 있는 경우, 가상 비서(1002)는 그 이름의 인스턴스를 검색할 수 있다. 히스토리가 유지되고 있는 경우, 가상 비서(1002)는 먼저 현재 값들을 검색한 다음, 매칭이 발견될 때까지 이전 데이터를 컨설팅한다.

Search by context variable name . If the name of the requested context variable is known, such as "first user name ", virtual assistant 1002 can retrieve an instance of that name. If the history is maintained, the virtual assistant 1002 first retrieves the current values and then consults the previous data until a match is found.

컨텍스트 변수 유형에 의한 검색. 개인과 같은, 요청된 컨텍스트 변수의 유형이 알려져 있는 경우, 가상 비서(1002)는 이러한 유형의 컨텍스트 변수의 인스턴스들을 검색할 수 있다. 히스토리가 유지되는 경우, 가상 비서(1002)는 먼저 현재 값들을 검색한 다음, 매칭이 발견될 때까지 이전 데이터를 컨설팅할 수 있다.

Retrieval by context variable type . If the type of the requested context variable, such as an individual, is known, the virtual assistant 1002 can retrieve instances of this type of context variable. If the history is maintained, the virtual secretary 1002 may first retrieve the current values and then consult the previous data until a match is found.

일 실시예에서는, 현재 정보 처리 문제가 단 하나의 매칭을 요구하는 경우, 매칭이 일단 발견되면 검색은 종료한다. 여러 개의 매칭이 허용되는 경우에는, 몇몇 제한에 도달할 때까지 매칭 결과가 검색될 수 있다.In one embodiment, if the current information processing problem requires only one match, the search is terminated once a match is found. If multiple matches are allowed, matching results can be retrieved until some limit is reached.

일 실시예에서는, 적절하다면, 가상 비서(1002)가 특정 유도를 갖는 데이터에 대한 검색을 제한할 수 있다. 예를 들어, 이메일을 보내기 위해 태스크 플로우 내에서 사람 객체를 찾고 있는 경우, 가상 비서(1002)는 그 유도가 그 도메인과 연관된 애플리케이션인 컨텍스트 변수를 고려하기만 하면 될 수 있다. In one embodiment, if appropriate, the virtual secretary 1002 may limit the search for data with a particular derivation. For example, if a person is looking for a person object in a task flow to send an e-mail, the virtual secretary 1002 only has to consider the context variable whose derivation is the application associated with that domain.

일 실시예에서, 가상 비서(1002)는 컨텍스트 변수의 임의의 이용 가능한 속성들을 이용한 발견법(heuristics)에 따라 매칭을 서열화하는 규칙을 이용한다. 예를 들어, "그녀에게 늦겠다고 전해줘"라는 커맨드를 포함하는 사용자 입력을 처리하는 경우, 가상 비서(1002)는 컨텍스트를 참조하여 "그녀"를 해석한다. 그러는 동안, 가상 비서(1002)는, 그 도출이 텍스트 메시징 및 이메일과 같은 통신 애플리케이션을 위한 애플리케이션 사용 히스토리인 사람 객체에 대한 기본설정을 표시하도록 서열화를 적용할 수 있다. 또 다른 예로서, "그녀에게 전화를 해"라는 커맨드를 해석하는 경우, 가상 비서(1002)는 전화 번호가 알려져 있지 않은 사람 객체들보다는 전화 번호가 있는 사람 객체들을 선호하는 것으로 서열화를 적용할 수 있다. 일 실시예에서, 서열화 규칙은 도메인들과 연관될 수 있다. 예를 들어, 이메일 및 전화 도메인에 대한 개인 변수의 서열화를 위해 서로 다른 서열화 규칙들이 이용될 수 있다. 당업자는, 이러한 임의의 서열화 규칙(들)이, 필요한 컨텍스트 정보에 대한 특수한 표현 및 액세스에 따라서 창출 및/또는 적용될 수 있다는 것을 인식할 것이다.In one embodiment, the virtual secretary 1002 uses rules to rank the matches according to heuristics using any available attributes of the context variable. For example, when processing a user input that includes a command "Tell her to be late ", the virtual assistant 1002 interprets" her " Meanwhile, the virtual secretary 1002 may apply the sequencing so that its derivation represents a preference for a person object that is an application usage history for a communication application, such as text messaging and e-mail. As another example, when interpreting the command "Call her ", the virtual assistant 1002 can apply the sequencing by preferring the person objects with the telephone number rather than the person objects whose telephone number is not known have. In one embodiment, the sequencing rules may be associated with domains. For example, different ranking rules may be used for ranking personal variables for e-mail and telephone domains. Those skilled in the art will recognize that any such sorting rule (s) may be created and / or applied in accordance with a particular representation and access to the required context information.

가상 비서 처리를 향상시키기 위한 To improve virtual secretary processing 컨텍스트의Contextual 이용 Use

전술한 바와 같이, 컨텍스트는 가상 비서(1002)의 동작과 관련하여 각종 계산 및 추론에 적용될 수 있다. 지금 도 2를 참조하면, 일 실시예에 따른, 가상 비서(1002)에서의 다양한 처리 단계들에서 컨텍스트를 사용하는 방법(10)을 도시하는 흐름도가 도시된다. As described above, the context can be applied to various calculations and inferences with respect to the operation of the virtual secretary 1002. [ Referring now to FIG. 2, there is shown a flow diagram illustrating a method 10 of using a context in various processing steps in a virtual assistant 1002, in accordance with one embodiment.

방법(10)은 가상 비서(1002)의 하나 이상의 실시예들과 관련되어 구현될 수 있다.The method 10 may be implemented in association with one or more embodiments of the virtual assistant 1002.

적어도 일 실시예에서는, 방법(10)이 다양한 유형의 기능들, 동작들, 액션들, 및/또는 예를 들면 이하의 하나 이상의 단계들(또는 이들의 조합)과 같은 다른 특징들을 수행하도록 동작될 수 있다:In at least one embodiment, the method 10 may be implemented to perform various types of functions, operations, actions, and / or other features such as, for example, one or more of the following steps (or combinations thereof) Can:

사용자와 가상 비서(1002) 간의 대화 인터페이스의 인터페이스 제어 흐름 루프를 실행한다. 방법(10)의 적어도 한 번의 반복은 대화에서 플라이(ply)로서 역할을 할 수 있다. 대화 인터페이스는, 대화 방식으로, 발성을 앞뒤로 함으로써 사용자와 비서(1002)가 통신하는 인터페이스이다.

And executes the interface control flow loop of the conversation interface between the user and the virtual secretary 1002. [ At least one iteration of method (10) can act as a ply in the conversation. The conversation interface is an interface through which the user and the secretary 1002 communicate with each other in a conversation manner by moving the talk back and forth.

가상 비서(1002)에 대한 실행적인 제어 흐름을 제공한다. 즉, 절차는 입력의 수집, 입력의 처리, 출력의 생성, 및 출력의 사용자에의 제시를 제어한다.

And provides an executive control flow to the virtual secretary 1002. That is, the procedure controls the collection of inputs, processing of inputs, generation of outputs, and presentation of the outputs to the user.

가상 비서(1002)의 컴포넌트들 간의 통신을 조정한다. 즉, 하나의 컴포넌트의 출력이 다른 하나의 컴포넌트에 피딩되는 장소 그리고 환경으로부터의 전체 입력 및 환경에 대한 액션이 발생할 수 있는 장소를 지시할 수 있다.

And coordinates the communication between the components of the virtual secretary 1002. That is, it can indicate where the output of one component is fed to another component and where all actions from the environment and environment can occur.

적어도 일부 실시예들에서는, 방법(10)의 일부가 또한 컴퓨터 네트워크의 다른 디바이스들 및/또는 시스템들에서 구현될 수 있다.In at least some embodiments, portions of the method 10 may also be implemented in other devices and / or systems of a computer network.

특정한 실시예에 따르면, 방법(10)의 다수의 인스턴스들 또는 스레드들이 동시에 구현되고/되거나 하나 이상의 프로세서(63) 및/또는 하드웨어의 다른 조합 및/또는 하드웨어와 소프트웨어의 이용에 의해 동시에 개시될 수 있다. 적어도 일 실시예에서는, 방법(10)의 하나 이상의 또는 선택된 일부가 하나 이상의 클라이언트(들)(1304), 하나 이상의 서버(들)(1340), 및/또는 이들의 조합에서 구현될 수 있다.According to a particular embodiment, multiple instances or threads of the method 10 may be implemented simultaneously and / or concurrently with the use of one or more processors 63 and / or other combinations of hardware and / or hardware and software have. In at least one embodiment, one or more selected portions of method 10 may be implemented in one or more client (s) 1304, one or more server (s) 1340, and / or combinations thereof.

예를 들면, 적어도 일부 실시예들에서는, 방법(10)의 다양한 양태, 특징 및/또는 기능성들이 소프트웨어 컴포넌트, 네트워크 서비스, 데이터베이스, 및/또는 기타, 또는 이들의 임의의 조합에 의해 수행, 구현 및/또는 개시될 수 있다. For example, in at least some embodiments, various aspects, features, and / or functionality of the method 10 may be implemented, implemented, and / or provided by software components, network services, databases, and / or the like, / &Lt; / RTI >

상이한 실시예들에 따르면, 방법(10)의 하나 이상의 스레드 또는 인스턴스는 방법(10)의 적어도 하나의 인스턴스의 개시를 트리거하기 위하여 (예를 들어, 최소 임계값 기준과 같은) 하나 이상의 상이한 유형의 기준(criteria)을 만족시키는 하나 이상의 조건 또는 이벤트의 검출에 응답하여 개시될 수 있다. 방법의 하나 이상의 상이한 스레드 또는 인스턴스의 예들의 개시 및/또는 구현을 트리거할 수 있는 다양한 유형의 조건 또는 이벤트의 예는 다음 중 하나 이상(또는 그 결합)(이에 제한되지 않음)을 포함할 수 있다:According to different embodiments, one or more threads or instances of the method 10 may be used to trigger the initiation of at least one instance of the method 10 (e. G., A minimum threshold criterion) May be initiated in response to the detection of one or more conditions or events that satisfy the criteria. Examples of various types of conditions or events that can trigger the initiation and / or implementation of examples of one or more different threads or instances of a method may include (but are not limited to) one or more of the following :

예를 들어, 다음 중 하나 이상(이에 제한되지 않음)과 같은 가상 비서(1002)의 인스턴스를 갖는 사용자 세션:

For example, a user session with an instance of a virtual secretary 1002, such as but not limited to one or more of the following:

ㅇ 예를 들어, 가상 비서(1002)의 실시예를 구현 중인 모바일 디바이스 애플리케이션을 시동 중인(starting up) 모바일 디바이스 애플리케이션;For example, a mobile device application starting up a mobile device application implementing an embodiment of virtual secretary 1002;

ㅇ 예를 들어, 가상 비서(1002)의 실시예를 구현 중인 애플리케이션을 시동 중인 컴퓨터 애플리케이션;A computer application that is running an application that is implementing an embodiment of virtual assistant 1002, for example;

ㅇ "음성 입력 버튼"과 같은, 눌려지는 모바일 디바이스 상의 전용 버튼;Dedicated buttons on the pressed mobile device, such as "voice input buttons";

ㅇ 헤드셋, 전화 핸드셋 또는 기지국, GPS 네비게이션 시스템, 소비자 기기, 리모트 컨트롤 또는 보조자를 호출하는 것과 연관될 수 있는 버튼을 갖는 임의의 기타 디바이스와 같은 컴퓨터 또는 모바일 디바이스에 부착된 주변 디바이스 상의 버튼;Buttons on a peripheral device attached to a computer or mobile device, such as a headset, a telephone handset or base station, a GPS navigation system, a consumer device, any other device having a button that can be associated with calling a remote control or assistant;

ㅇ 웹브라우저로부터 가상 비서(1002)를 구현하는 웹사이트로 시작된 웹세션;A web session initiated from a web browser to a web site that implements the virtual secretary 1002;

ㅇ 기존의 웹브라우저 세션 내로부터 예를 들어, 가상 비서(1002) 서비스가 요청되는 가상 비서(1002)를 구현하는 웹사이트로 시작된 상호작용;An interaction initiated from within an existing web browser session, for example, a web site that implements a virtual secretary 1002 where a virtual secretary 1002 service is requested;

ㅇ 가상 비서(1002)의 구현물과의 통신을 조정하는 모달리티 서버(1426)로 전송된 이메일 메시지;An email message sent to the modality server 1426 that coordinates communications with the implementation of the virtual secretary 1002;

ㅇ 텍스트 메시지가 가상 비서(1002)의 구현물과의 통신을 조정하는 모달리티 서버(1426)로 전송됨;A text message is sent to the modality server 1426 that coordinates communication with the implementation of the virtual secretary 1002;

ㅇ 가상 비서(1002)의 구현물과의 통신을 조정하는 모달리티 서버(1434)에 전화 통화가 행해짐;A phone call is made to the modality server 1434 that coordinates communications with the implementation of the virtual secretary 1002;

ㅇ 가상 비서(1002)의 구현을 제공하는 애플리케이션에 경보 또는 통지와 같은 이벤트가 전송됨.An event, such as an alert or a notification, is sent to the application providing the implementation of the virtual secretary (1002).

가상 비서(1002)를 제공하는 디바이스가 턴온 및/또는 시작될 때.

When the device providing the virtual secretary 1002 is turned on and / or started.

상이한 실시예들에 따르면, 방법(10)의 하나 이상의 상이한 스레드 또는 인스턴스는 수동으로, 자동으로, 통계적으로, 동적으로, 동시에 및/또는 그 조합으로 개시 및/또는 구현될 수 있다. 또한, 방법(10)의 상이한 인스턴스들 및/또는 실시예들은 하나 이상의 상이한 시구간에서 개시될 수 있다(예를 들어, 특정 시구간 동안, 규칙적인 주기적 구간에, 불규칙적인 주기적 구간에, 수요 시 등).According to different embodiments, one or more different threads or instances of method 10 may be initiated and / or implemented manually, automatically, statistically, dynamically, concurrently, and / or in combination. Further, different instances and / or embodiments of the method 10 may be initiated in one or more different time periods (e.g., during a specific time period, at regular periodic intervals, at irregular periodic intervals, Etc).

적어도 일 실시예에서, 방법(10)의 주어진 인스턴스는 전술한 바와 같은 컨텍스트 데이터를 포함한 특정 태스크들 및/또는 동작들을 수행할 때, 다양한 상이한 유형의 데이터 및/또는 기타 유형의 정보를 사용 및/또는 생성할 수 있다. 데이터는 또한 임의의 다른 유형의 입력 데이터/정보 및/또는 출력 데이터/정보를 포함할 수 있다. 예를 들어, 적어도 일 실시예에서, 방법(10)의 적어도 하나의 인스턴스는 예를 들어, 하나 이상의 데이터베이스와 같은 하나 이상의 상이한 유형의 소스로부터의 정보를 액세스, 처리 및/또는 그렇지 않으면 사용할 수 있다. 적어도 일 실시예에서, 적어도 데이터베이스 정보의 일부분은 하나 이상의 로컬 및/또는 리모트 메모리 디바이스와의 통신을 통해 액세스될 수 있다. 또한, 방법(10)의 적어도 하나의 인스턴스는 예를 들어, 로컬 메모리 및/또는 리모트 메모리 디바이스들에 저장될 수 있는 하나 이상의 상이한 유형의 출력 데이터/정보를 생성할 수 있다.In at least one embodiment, a given instance of the method 10 may use and / or manipulate various different types of data and / or other types of information when performing particular tasks and / or actions, including contextual data as described above. Or < / RTI > The data may also include any other type of input data / information and / or output data / information. For example, in at least one embodiment, at least one instance of method 10 may access, process, and / or otherwise use information from one or more different types of sources, such as, for example, one or more databases . In at least one embodiment, at least a portion of the database information may be accessed through communication with one or more local and / or remote memory devices. Also, at least one instance of the method 10 may generate one or more different types of output data / information that may be stored, for example, in local memory and / or remote memory devices.

적어도 일 실시예에서, 방법(10)의 주어진 인스턴스의 초기 구성은 하나 이상의 상이한 유형의 초기화 파라미터를 사용하여 수행될 수 있다. 적어도 일 실시예에서, 초기화 파라미터들 중 적어도 일부는 하나 이상의 로컬 및/또는 리모트 메모리 디바이스와의 통신을 통해 액세스될 수 있다. 적어도 일 실시예에서, 방법(10)의 인스턴스에 제공된 초기화 파라미터의 적어도 일부는 입력 데이터/정보에 대응할 수 있고, 및/또는 입력 데이터/정보로부터 파생될 수 있다.In at least one embodiment, the initial configuration of a given instance of the method 10 may be performed using one or more different types of initialization parameters. In at least one embodiment, at least some of the initialization parameters may be accessed through communication with one or more local and / or remote memory devices. In at least one embodiment, at least some of the initialization parameters provided to the instance of method 10 may correspond to input data / information, and / or may be derived from input data / information.

도 2의 특정 예에서, 단일 사용자가 음성 입력 기능들을 갖는 클라이언트 애플리케이션으로부터 네트워크를 통해 가상 비서(1002)의 인스턴스에 액세스하고 있다고 가정한다.In the particular example of FIG. 2, assume that a single user is accessing an instance of virtual secretary 1002 over the network from a client application having voice input capabilities.

음성 입력은 유도(elicitation)되어 해석된다(100). 유도는 임의의 적절한 모드에서 프롬프트들을 제시하는 것을 포함할 수 있다. 다양한 실시예들에서, 클라이언트의 사용자 인터페이스는 입력의 여러 모드들을 제공한다. 이들은 예를 들어, 다음을 포함할 수 있다;The speech input is interpreted by elicitation (100). Induction may include presenting prompts in any suitable mode. In various embodiments, the user interface of the client provides various modes of input. These may include, for example:

액티브 타이핑된 입력(active typed-input) 유도 절차를 호출할 수 있는, 타이핑된 입력을 위한 인터페이스;

An interface for a typed input that can invoke an active typed-input derivation procedure;

액티브 음성 입력 유도 절차를 호출할 수 있는, 음성 입력을 위한 인터페이스;

An interface for voice input, capable of calling an active voice input inducing procedure;

액티브 GUI 기반 입력 유도를 호출할 수 있는, 메뉴로부터 입력들을 선택하기 위한 인터페이스.

An interface for selecting inputs from a menu that can invoke an active GUI-based input derivation.

이들 각각을 수행하기 위한 기술들이 상기 참조된 관련 특허 출원들에서 개시된다. 당업자는 다른 입력 모드들이 제공될 수 있다는 것을 인식할 것이다. 단계(100)의 출력은 입력 음성의 후보 해석의 집합(190)이다.Techniques for performing each of these are disclosed in the above referenced related patent applications. Those skilled in the art will recognize that other input modes may be provided. The output of step 100 is a set 190 of candidate interpretations of the input speech.

후보 해석의 집합(190)은 언어 해석기(2770)(자연어 프로세서 또는 NLP라고도 칭함)에 의해 처리되며(200), 언어 해석기(2770)는 텍스트 입력을 파싱하고, 사용자 의도의 가능한 해석의 집합(290)을 생성한다.The set of candidate interpretations 190 is processed 200 by a language interpreter 2770 (also referred to as a natural language processor or NLP) and the language interpreter 2770 parses the text input and generates a set of possible interpretations 290 ).

단계(300)에서, 사용자의 의도의 표시(들)(290)는 도 5와 관련하여 기술된 플로우 분석 절차 및 대화의 실시예를 구현하는 대화 플로우 프로세서(2780)로 전달된다. 대화 플로우 프로세서(2780)는 어느 의도의 해석이 가장 가능성 있는지를 결정하고, 이 해석을 도메인 모델들의 인스턴스들 및 태스크 모델의 파라미터들에 맵핑하고, 태스크 플로우에서 다음 플로우 단계를 결정한다.At step 300, the indication (s) 290 of the user's intent is conveyed to a dialog flow processor 2780 which implements an embodiment of the flow analysis procedure and dialogue described with respect to Fig. Conversation flow processor 2780 determines which interpretation is most likely, maps this interpretation to the instances of the domain models and the parameters of the task model, and determines the next flow step in the task flow.

단계(400)에서, 식별된 플로우 단계가 실행된다. 일 실시예에서, 플로우 단계의 호출은 사용자의 요청 대신에 서비스들의 집합을 호출하는 서비스 편성 컴포넌트(2782)에 의해 수행된다. 일 실시예에서, 이들 서비스들은 일부 데이터를 공통 결과로 참가시킨다(contribute).In step 400, the identified flow step is executed. In one embodiment, a call to the flow step is performed by a service organizing component 2782 that invokes a set of services instead of a user request. In one embodiment, these services contribute some data as a common result.

단계(500)에서, 대화 응답이 생성된다. 단계(700)에서, 응답이 출력을 위하여 클라이언트 디바이스로 전송된다. 디바이스 상의 클라이언트 소프트웨어는 이것을 클라이언트 디바이스의 스크린(또는 다른 출력 디바이스) 상에 렌더링한다.At step 500, a conversation response is generated. At step 700, a response is sent to the client device for output. The client software on the device renders it on the screen (or other output device) of the client device.

응답을 본 후, 사용자가 완료하면(790), 방법은 종료한다. 사용자가 종료하지 않으면, 단계(100)로 리턴함으로써 루프의 다른 반복이 개시된다.After viewing the response, if the user completes (790), the method ends. If the user does not terminate, another iteration of the loop is initiated by returning to step 100.

컨텍스트 정보(1000)가 방법(10)의 다양한 포인트들에서 시스템의 다양한 컴포넌트들에 의해 사용될 수 있다. 예를 들어, 도 2에 도시된 바와 같이, 컨텍스트(1000)는 단계(100, 200, 300 및 500)에서 사용될 수 있다. 이들 단계들에서의 컨텍스트(1000)의 사용의 추가적인 설명이 이하에서 제공된다. 그러나, 당업자는 본 발명의 핵심적인 특징들로부터 벗어나지 않고, 컨텍스트 정보의 사용이 이들 특정 단계들에 제한되지 않고, 또한 시스템이 다른 포인트들에서 컨텍스트 정보를 사용할 수 있다는 것을 인식할 것이다.Context information 1000 may be used by various components of the system at various points in the method 10. For example, as shown in FIG. 2, a context 1000 may be used in steps 100, 200, 300, and 500. Additional explanations of the use of contexts 1000 in these steps are provided below. However, those skilled in the art will recognize that the use of context information is not limited to these specific steps, and that the system can use context information at other points without departing from the essential features of the present invention.

또한, 당업자는 방법(10)의 상이한 실시예들이 도 2에 도시된 특정 실시예에 도시된 것 외의 추가의 특징들 및/또는 동작들을 포함할 수 있고, 도 2의 특정 실시예에 도시된 것과 같은 방법(10)의 특징들 및/또는 동작들 중 적어도 일부를 생략할 수 있다는 것을 인식할 것이다.Further, those skilled in the art will appreciate that the different embodiments of method 10 may include additional features and / or operations other than those shown in the specific embodiment shown in FIG. 2, It will be appreciated that at least some of the features and / or operations of the same method 10 may be omitted.

음성 유도 및 해석에서의 In speech induction and interpretation 컨텍스트Context 사용 use

이제 도 3을 참조하면, 일 실시예에 따른 음성 인식을 향상시키도록, 음성 유도 및 해석(100)에서 컨텍스트를 사용하기 위한 방법을 도시하는 흐름도가 도시되어 있다. 컨텍스트(1000)는 예를 들어, 음소(phoneme)들을 단어들에 매칭시키는 후보 가설들의 생성, 서열화 및 필터링을 가이드하기 위하여 음성 인식에서의 명확성(disambiguation)을 위하여 사용될 수 있다. 상이한 음성 인식 시스템들은 생성, 서열 및 필터의 다양한 혼합들을 사용하지만, 컨텍스트(1000)는 일반적으로 임의의 스테이지에서 가설 공간을 감소시키기 위하여 적용될 수 있다.Referring now to FIG. 3, there is shown a flow chart illustrating a method for using context in speech induction and interpretation 100 to improve speech recognition in accordance with an embodiment. Context 1000 may be used for disambiguation in speech recognition, for example, to guide the generation, sequencing, and filtering of candidate hypotheses that match phonemes to words. Different speech recognition systems use various mixtures of productions, sequences and filters, but the context 1000 can generally be applied to reduce the hypothetical space at any stage.

방법이 시작된다(100). 가상 비서(1002)는 청각 신호의 형태로 보이스 또는 음성 입력을 수신한다(121). 음성-텍스트 변환 서비스(122) 또는 프로세서는 청각 신호의 후보 텍스트 해석(124)의 집합을 생성한다. 일 실시예에서, 음성-텍스트 변환 서비스(122)는 예를 들어, Massachusetts주의 Burlington의 Nuance Communications, Inc.에서 이용가능한 Nuance Recognizer를 사용하여 구현된다.The method begins (100). The virtual secretary 1002 receives a voice or voice input in the form of an auditory signal (121). The speech-to-text conversion service 122 or processor generates a set of candidate text interpretations 124 of the auditory signals. In one embodiment, the voice-to-text conversion service 122 is implemented using a Nuance Recognizer available, for example, from Nuance Communications, Inc. of Burlington, Massachusetts.

일 실시예에서, 가상 비서(1002)는 음성 입력(121)의 후보 텍스트 해석(124)을 생성하기 위하여 통계 언어 모델들(1029)을 사용한다. 일 실시예에서, 컨텍스트(1000)는 음성-텍스트 변환 서비스(122)에 의해 생성되는 후보 해석들(124)의 생성, 필터링 및/또는 서열화를 바이어스하기 위하여 적용된다. 예를 들어,In one embodiment, virtual secretary 1002 uses statistical language models 1029 to generate candidate text interpretation 124 of speech input 121. In one embodiment, the context 1000 is applied to bias the generation, filtering, and / or sequencing of the candidate interpretations 124 generated by the voice-to-text conversion service 122. E.g,

음성-텍스트 변환 서비스(122)는 통계 언어 모델들(1029)을 바이어스하기 위하여 사용자 퍼스널 데이터베이스(들)(1058)로부터의 어휘를 사용할 수 있다.

The voice-to-text conversion service 122 may use the vocabulary from the user personal database (s) 1058 to bias the statistical language models 1029.

음성-텍스트 변환 서비스(122)는 커스텀 통계 언어 모델(1029)을 선택하기 위하여 대화 상태 컨텍스트를 사용할 수 있다. 예를 들어, 예/아니오 질문을 할 때, 이들 단어들을 듣는 쪽으로 바이어스되는 통계 언어 모델(1029)이 선택될 수 있다.

The speech-to-text conversion service 122 may use the dialog state context to select a custom statistical language model 1029. [ For example, when asking yes / no questions, a statistical language model 1029 that is biased towards listening to these words may be selected.

음성-텍스트 변환 서비스(122)는 적절한 단어들 쪽으로 바이어스하기 위하여 현재 애플리케이션 컨텍스트를 사용할 수 있다. 예를 들어, "그녀에게 전화하는 것(call her)"는 텍스트 메시지 애플리케이션 컨텍스트에서 "칼라(collar)"보다 선호(prefer)될 수 있는데, 왜냐하면 이러한 컨텍스트는 전화걸 수 있는 사람 객체(Person Object)들을 제공하기 때문이다.

The voice-to-text conversion service 122 may use the current application context to bias towards the appropriate words. For example, "call her" may be preferred over a "collar" in the context of a text message application because such a context may contain a Person object, As shown in FIG.

예를 들어, 주어진 음성 입력은 해석들 "그녀에게 전화하는 것" 및 "칼라"를 생성하기 위하여 음성-텍스트 변환 서비스(122)를 이끌어낼 수 있다. 통계적 언어 모델(SLM; 1029)들에 의해 가이드된, 음성-텍스트 변환 서비스(speech-to-text service; 122)는 문법적 제약사항에 의해 "call"을 들은 후에 이름들을 듣는 것으로 튜닝(tune)될 수 있다. 음성-텍스트 변환 서비스(122)는 또한 컨텍스트(1000)에 기초하여 튜닝될 수 있다. 예를 들어, "Herb"가 사용자 어드레스 북의 이름이라면, 이 컨텍스트는 "Herb"를 두 번째 음절의 해석으로서 간주하기 위하여 임계값을 낮추는 데에 이용될 수 있다. 즉, 사용자의 개인 데이터 컨텍스트에서의 이름들의 존재(presence)는 가설(hypotheses)을 세우는데 사용되는 통계적 언어 모델(1029)의 조정 및 선택에 영향을 미칠 수 있다. 이름 "Herb"는 일반 SLM(1029)의 일부일 수 있거나, 컨텍스트(1000)에 의해 직접 추가될 수 있다. 일 실시예에서, 이것은 추가적인 SLM(1029)로서 추가될 수 있는데, 이는 컨텍스트(1000)에 기초하여 조정된다. 일 실시예에서, 이것은 기존의 SLM(1029)의 조정일 수 있는데, 이는 컨텍스트(1000)에 기초하여 조정된다.For example, a given speech input may lead to a speech-to-text conversion service 122 to generate interpretations "call her" and "color". A speech-to-text service 122, guided by statistical language models (SLMs) 1029, is tune to hearing names after hearing a "call " . The voice-to-text conversion service 122 may also be tuned based on the context 1000. For example, if "Herb" is the name of the user address book, then this context can be used to lower the threshold to regard "Herb " as an interpretation of the second syllable. That is, the presence of names in the user's personal data context may affect the adjustment and selection of the statistical language model 1029 used to construct the hypotheses. The name "Herb" may be part of the generic SLM 1029, or may be added directly by the context 1000. In one embodiment, this may be added as an additional SLM 1029, which is adjusted based on the context 1000. In one embodiment, this may be an adjustment of the existing SLM 1029, which is adjusted based on the context 1000.

일 실시예에서, 통계적 언어 모델(1029)들은 또한, 장기 개인 메모리(2754)에 저장될 수 있는 애플리케이션 선호도 및 사용 히스토리(1072; application preferences and usage history) 및/또는 개인 데이터베이스(1058)로부터 단어, 이름, 및 어구를 찾기 위해서 조정된다. 예를 들어, 통계적 언어 모델(1029)에는 할일(to-do) 항목들, 리스트 항목들, 개인 메모들, 캘린더 엔트리들, 연락처/어드레스 북의 사람 이름들, 이메일 어드레스들, 연락처/어드레스 북에 언급된 거리 또는 도시명, 및 기타 등등이 주어질 수 있다.In one embodiment, statistical language models 1029 may also be used to store words or phrases from application preferences and usage history (s) 1072 and / or personal database (s) 1058 that may be stored in long term personal memory 2754. [ Names, and phrases. For example, the statistical language model 1029 may include to-do items, list items, personal notes, calendar entries, person names of contact / address book, email addresses, contact / The street or city name mentioned, and so on.

서열화 컴포넌트(ranking component)는 후보 해석들(candidate interpretations; 124)을 분석하고 이들이 가상 비서(virtual assistant; 1002)의 구문론적 및/또는 의미론적 모델에 얼마나 부합한지에 따라 이 후보 해석들을 서열화한다(126). 사용자 입력에 대한 제약사항의 소스들은 어느 것이라도 이용될 수 있다. 예를 들어, 일 실시예에서, 비서(1002)는 해석들이 구문론적 및/또는 의미론적인 관점, 도메인 모델, 태스크 플로우 모델, 및/또는 대화 모델, 및/또는 기타 등등으로 얼마나 잘 파싱(parse)하는지에 따라 음성-텍스트 해석기(interpreter)의 출력을 서열화할 수 있다: 이는, 앞서 참조된 관련 U.S. 실용신안 출원 명세서에서 기술된 바와 같이, 후보 해석(124) 내의 단어들의 여러 가지 조합들이 개념들, 관계들, 엔터티들, 액티브 온톨로지(active ontology)의 속성들 및 그 관련 모델들에 얼마나 부합할 것인지를 평가한다.The ranking component analyzes the candidate interpretations 124 and orders these candidate interpretations according to how they fit into the syntactic and / or semantic model of the virtual assistant 1002 126). Any source of constraints on user input can be used. For example, in one embodiment, the secretary 1002 can parse how well the interpretations are in a syntactic and / or semantic perspective, a domain model, a task flow model, and / or an interactive model, The output of the voice-to-text interpreter can be sequenced according to whether it is a related US As described in the utility model specification, how many combinations of words in the candidate interpretation 124 fit into the concepts, relationships, entities, properties of the active ontology and their associated models .

후보 해석들의 서열화(126)는 컨텍스트(1000)에 의해 또한 영향받을 수 있다. 예를 들어, 사용자가 현재 가상 비서(1002)가 호출될 때 텍스트 메시징 애플리케이션으로 대화를 전달하고 있는 경우, 어구 "call her"은 단어 "collar"보다 올바른 해석이 될 가능성이 더 높은데, 그 이유는 이 컨텍스트에서 전화를 할(call) 잠재적인 "her"가 있기 때문이다. 이러한 바이어스(bias)는 현재 애플리케이션 컨텍스트가 "전화걸 수 있는 엔터티(callable entities)"를 제공할 수 있는 애플리케이션을 지시할 때 "call her" 또는 "call <contact name>"과 같이 우세한 어구들로 가설의 서열화(126)를 조정함으로써 이루어질 수 있다. The sequencing 126 of candidate interpretations may also be influenced by the context 1000. For example, if the user is currently delivering a conversation to the text messaging application when the virtual assistant 1002 is invoked, the phrase " call her "is more likely to be a better interpretation than the word " collar" This is because there is a potential "her" to call in this context. This bias may be used to pre-empt the dominant phrases such as "call her" or "call <contact name>" when the current application context indicates an application capable of providing "callable entities & (126). &Lt; / RTI >

각종 실시예들에서, 도 3에 도시된 자연어 처리 절차의 임의의 실시예를 포함하는 텍스트 입력들의 해석을 위하여 비서(1002)에 의해 사용되는 알고리즘 또는 절차는 음성-텍스트 변환 서비스(122)에 의해 생성된 후보 텍스트 해석들(124)을 서열화하고 스코어링(score)하는 데에 이용될 수 있다.In various embodiments, the algorithm or procedure used by the secretary 1002 for the interpretation of textual inputs, including any embodiment of the natural language processing procedure shown in FIG. 3, is performed by the voice-to-text translation service 122 May be used to sequence and score the generated candidate text interpretations (124).

컨텍스트(1000)는 후보 해석들(124)의 서열화에 영향을 주거나 그 생성을 제약하는 대신에 또는 그에 더하여, 후보 해석들(124)을 필터링하는 데에도 사용될 수 있다. 예를 들어, 필터링 규칙은 "Herb"에 대한 어드레스 북 엔트리의 컨텍스트가, 이 컨텍스트를 포함하는 어구는, 비록 다른 경우에 필터링 임계값 이하에 있을 경우더라도, 최상위 후보(130)로 간주되어야 함을 충분히 나타낸다고 규정할 수 있다. 사용되고 있는 특정 음성 인식 기술에 따라서, 컨텍스트의 바이어스에 기초한 제약들이 생성, 서열화, 및/또는 필터링 단계에 적용될 수 있다.Context 1000 may also be used to filter candidate interpretations 124 instead of or in addition to affecting or limiting the sequencing of candidate interpretations 124. For example, the filtering rule should be considered the highest candidate 130, even if the context of the address book entry for "Herb ", and the phrase containing this context, is otherwise below the filtering threshold It can be stipulated to be sufficient. Depending on the particular speech recognition technique being used, constraints based on the bias of the context may be applied to the generation, sequencing, and / or filtering steps.

일 실시예에서, 서열화 컴포넌트(126)는 해석들(124)로부터 최고 서열화 음성 해석이 특정된 임계값을 초과하여 서열화된다고 결정하면(128), 이 최고 서열화 해석은 자동으로 선택될 수 있다(130). 어떠한 해석도 특정된 임계값을 초과하여 서열화되지 않으면, 음성의 가능한 후보 해석들(134)이 사용자에게 제공된다(132). 그러면 사용자가 디스플레이된 선택사항들(choices) 중에서 선택을 할 수 있다(136).In one embodiment, if the ranking component 126 determines 128 that the highest ranking speech interpretation from the interpretations 124 is sequenced above a specified threshold, then this highest ranking interpretation may be automatically selected 130 ). If no interpretation is sequenced beyond a specified threshold, possible candidate interpretations 134 of speech are provided to the user (132). The user can then select among the displayed choices (136).

이제 도 26a 및 26b를 참조하여 보면, 일 실시예에 따른, 후보 해석들 중에서 선택을 하기 위한 사용자 인터페이스의 일례를 도시하는 스크린 샷이 도시된다. 도 26a는 모호한 해석(ambiguous interpretation; 2651) 아래에 점선이 있는 사용자의 음성의 표시를 도시한다. 사용자가 텍스트 상으로 태핑(tap)한다면, 도 26b에 도시된 바와 같이, 대안적인 해석들(2652A, 2652B)을 보여준다. 일 실시예에서, 컨텍스트(1000)는 후보 해석들(2652A, 2652B) 중 어떤 것이 바람직한 해석인지(도 26a에서 초기 디폴트로 도시됨)에 그리고 도 26b에서와 같이 제시한 대안들의 유한 집합에서의 선택에도 영향을 미칠 수 있다.Referring now to Figures 26A and 26B, there is shown a screen shot illustrating an example of a user interface for selection among candidate analyzes, in accordance with one embodiment. 26A shows an indication of a user's voice with a dotted line under ambiguous interpretation 2651. FIG. If the user taps on the text, alternative interpretations 2652A and 2652B are shown, as shown in Figure 26B. In one embodiment, the context 1000 includes a selection of candidate interpretations 2652A and 2652B in a finite set of proposed alternatives as shown in Figure 26 (b) (which is shown in the initial default in Figure 26a) .

각종 실시예에서, 디스플레이된 선택사항 중에서의 사용자 선택(136)은 예컨데, 다모드 입력을 포함하는, 임의의 입력 모드에 의해 이루어질 수 있다. 이러한 입력 모드는 적극 유도형 입력(actively elicited typed input), 적극 유도형 음성 입력, 입력에 대한 적극 제시형 GUI, 및/또는 기타 등등을 포함하지만, 이에 제한되지 않는다. 일 실시예에서, 사용자는, 예컨테 태핑 또는 스피킹에 의하여, 후보 해석들(134) 중에서 선택을 할 수 있다. 스피킹의 경우, 새로운 음성 입력의 가능한 해석은 제공된 선택사항의 소형 세트(134)로 크게 제약된다.In various embodiments, the user selection 136 among the displayed options may be accomplished by any input mode, including, for example, multi-mode inputs. Such input modes include, but are not limited to, actively elicited typed input, aggressive inductive voice input, aggressive GUI for input, and / or the like. In one embodiment, the user may make a choice among the candidate interpretations 134, by example, tapping or speaking. In the case of speech, the possible interpretation of the new speech input is highly constrained by the small set of choices 134 provided.

입력이 자동으로 선택되든지(130) 사용자에 의해 선택되는지(136) 간에, 그 결과의 하나 이상의 텍스트 해석(들)(190)이 리턴된다. 적어도 하나의 실시예에서, 이 리턴된 입력에 주석이 달려서, 단계 136에서 어떠한 선택이 이루어졌는지에 대한 정보가 컨텍스트의 입력과 함께 보유된다. 이는, 예를 들어, 스트링의 기초를 이루는 의미론적 개념 또는 엔터티가 이 스트링 리턴시에 스트링과 연관될 수 있도록 해줌으로써, 후속 언어 해석의 정확도를 높여준다.The resulting one or more text interpretations (s) 190 are returned between whether the input is automatically selected (130) or selected (136) by the user. In at least one embodiment, this returned input is annotated so that information about what choice was made in step 136 is retained along with the input of the context. This improves the accuracy of subsequent language interpretation, for example, by allowing a semantic concept or entity underlying the string to be associated with the string at this string return.

도 1에 관련하여 기술된 소스들은 어느 것이라도 도 3에 도시된 음성 유도 및 해석 방법에 컨텍스트(1000)를 제공할 수 있다. 예를 들자면, 다음과 같다:Any of the sources described in connection with FIG. 1 may provide the context 1000 to the speech derivation and interpretation method shown in FIG. For example:

개인 음향 컨텍스트 데이터(Personal Acoustic Context Data; 1080)가, 가능한 SLM(1029)들로부터 선택을 하거나 다른 경우 이들 SLM을 인식된 음향 컨텍스트들에 최적화시키도록 튜닝하는 데에 이용될 수 있다.

Personal Acoustic Context Data 1080 can be used to select from possible SLMs 1029 or otherwise tune to optimize these SLMs to recognized acoustic contexts.

사용 중인 마이크로폰 및/또는 카메라의 속성들을 기술하는 디바이스 센서 데이터(1056)가, 가능한 SLM(1029)들로부터 선택을 하거나 다른 경우 이들 SLM을 인식된 음향 컨텍스트들에 최적화시키도록 튜닝하는 데에 이용될 수 있다.

Device sensor data 1056 describing the properties of the microphone and / or camera in use may be used to tune to select from possible SLMs 1029 or otherwise optimize these SLMs to recognized acoustic contexts .

개인 데이터베이스(1058) 및 애플리케이션 선호도 및 사용 히스토리(1072)로부터의 어휘(Vocabulary)가 컨텍스트(1000)로서 이용될 수 있다. 예를 들어, 미디어 명칭 및 아티스트의 이름이 언어 모델(1029)을 튜닝하는 데에 이용될 수 있다.

The personal database 1058 and the vocabulary from the application preferences and usage history 1072 may be used as the context 1000. For example, the name of the media and the name of the artist may be used to tune the language model 1029.

대화 히스토리 및 보조 메모리(1052)의 일부인 현재의 대화 상태(Current dialog state)가 음성-텍스트 변환 서비스(122)에 의해 후보 해석들(124)의 생성/필터링/서열화를 바이어스하는 데에 이용될 수 있다. 예를 들어, 대화 상태의 한 종류는 예/아니오 질문을 하는 것이다. 이러한 상태일 때, 절차(100)는 이들 단어를 듣는 방향으로 바이어싱하는 SLM(1029)을 선택할 수 있거나, 122에서 컨텍스트-특정 조정으로 이들 단어의 서열화 및 필터링을 바이어싱할 수 있다.

The current dialog state that is part of the conversation history and auxiliary memory 1052 can be used to bias generation / filtering / sequencing of candidate interpretations 124 by the voice-to-text conversion service 122 have. For example, one type of conversation state is asking yes / no questions. In this state, the procedure 100 may select the SLM 1029 which biases these words in the listening direction, or may bias the ranking and filtering of these words in a context-specific adjustment at 122.

자연어 처리에서의 In natural language processing 컨텍스트Context 사용 use

컨텍스트(1000)는 자연어 처리(NLP) - 텍스트 입력을 가능한 파스(parse)들을 나타내는 의미론적 구조로 파싱하는 것 - 을 용이하게 하는 데에 이용될 수 있다. 이제 도 4를 참조하여 보면, 일 실시예에 따른, 언어 해석기(2770)에 의해 수행될 수 있는, 자연어 처리에 있어 컨텍스트를 사용하기 위한 방법을 도시하는 플로우 차트가 도시된다. Context 1000 can be used to facilitate natural language processing (NLP) - parsing text input into a semantic structure representing possible parse. Referring now to FIG. 4, there is shown a flowchart illustrating a method for using a context in natural language processing, which may be performed by a language interpreter 2770, in accordance with an embodiment.

방법은 200에서 시작한다. 입력 텍스트(202)가 수신된다. 일 실시예에서, 입력 텍스트(202)는 패턴 인식기(2760), 어휘 데이터베이스(2758), 온톨로지들 및 기타 모델들(1050)을 이용하여 단어들 및 어구들에 대해 매칭되어(210), 사용자 입력과 개념들 간의 연관성(associations)을 식별한다. 단계 210는 후보 구문론적 파스들(212) 세트를 산출하는데, 이 파스들은 후보 의미론적 파스들(222)을 산출하는 의미론적 관련성(semantic relevance; 220)에 대하여 매치된다. 그 다음 후보 파스들은 230에서 모호한 대안들을 제거하는 처리가 수행되고, 필터링되고, 관련성에 의해 분류된 다음(232) 리턴된다.The method starts at 200. The input text 202 is received. In one embodiment, input text 202 is matched 210 to words and phrases using pattern recognizer 2760, lexicon database 2758, ontologies and other models 1050, And associations between concepts. Step 210 produces a set of candidate syntactic parses 212 that match for semantic relevance 220 that yields candidate semantic puzzles 222. The candidate parses are then processed at 230 to remove ambiguous alternatives, filtered, sorted by relevance (232), and returned.

자연어 처리 전반에 걸쳐, 가설 공간 및 제약 가능 파스들을 줄이기 위해 컨텍스트 정보(1000)가 적용될 수 있다. 예를 들어, 언어 해석기(2770)가 2개의 후보 "call her" 및 "call Herb"를 수신하였다면, 언어 해석기(2770)는 단어 "call", "her", 및 "Herb"에 대한 바인딩(212)들을 찾을 것이다. 애플리케이션 컨텍스트(1060)는 "call"에 대한 가능한 단어 뜻을 "phone call"을 의미하는 것으로 제약하는 데에 이용될 수 있다. 컨텍스트는 또한 "her" 및 "Herb"에 대한 지시 대상(referents)을 찾는 데에도 이용될 수 있다. "her"에 대해서는, 컨텍스트 소스(1000)들이 전화걸 수 있는 엔터티들의 소스를 찾기 위해 검색될 수 있다. 이 예에서, 텍스트 메시징 대화의 대상자가 전화걸 수 있는 엔터티이고, 이 정보는 텍스트 메시징 애플리케이션으로부터 오는 컨텍스트의 일부이다. "Herb"의 경우에는, 사용자의 어드레스 북이, (도메인 엔터티 데이터베이스(2772)로부터의 가장 선호하는 숫자들과 같은) 애플리케이션 선호도 및 (도메인 엔터티 데이터베이스(2772)로부터의 최근 건 전화와 같은) 애플리케이션 사용 히스토리와 같은 다른 개인 데이터에서와 같이, 명확히 하는(disambiguating) 컨텍스트의 소스이다. 이 예에서, 현재 텍스트 메시징을 하는 자는 Rebec-caRichards이고, 이 사용자의 어드레스 북에 HerbGowen이 있는 경우, 언어 해석기(2770)에 의해 생성된 2개의 파스들은 "Phone-Call(RebeccaRichards)" 및 "PhoneCall (HerbGowen)"를 나타내는 의미론적 구조일 것이다. Throughout the natural language processing, context information 1000 can be applied to reduce hypothesis space and constraintable parses. For example, if the language interpreter 2770 has received two candidates "call her" and "call Herb", the language interpreter 2770 receives a binding 212 for the words "call", "her", and "Herb" ). The application context 1060 can be used to constrain the possible word meaning for "call " to mean" phone call. &Quot; The context can also be used to find referents for "her" and "Herb ". For "her, " context sources 1000 may be searched to find sources of entities that can dial. In this example, the subject of the text messaging conversation is an entity that can dial, and this information is part of the context coming from the text messaging application. In the case of "Herb ", the address book of the user is used for application preferences (such as the most preferred numbers from the domain entity database 2772) and application usage (such as recent calls from the domain entity database 2772) It is the source of a disambiguating context, as in other personal data such as history. In this example, if the current text messaging is Rebec-caRichards and there is a HerbGowen in the address book of this user, the two parses generated by the language interpreter 2770 are "Phone-Call (Rebecca Richards)" and "PhoneCall (HerbGowen) ".

애플리케이션 선호도 및 사용 히스토리(1072)로부터의 데이터, 대화 히스토리 및 보조 메모리(1052), 및/또는 개인 데이터베이스(1058)는 언어 해석기(2770)에 의해 후보 구문론적 파스들(212)을 생성하는 데에 또한 사용될 수 있다. 이러한 데이터는, 예를 들어, 단기 메모리 및/또는 장기 메모리(2752, 2754)로부터 획득될 수 있다. 이런 식으로, 성능을 향상시키고, 모호함을 줄이며, 인터렉션의 대화적인 성격을 강화시키기 위해 동일한 세션에서 사전에 제공되었던 입력 및/또는 사용자에 대하여 알려진 정보가 이용될 수 있다. 또한, 유효한 후보 구문론적 파스들(212)을 결정함에 있어 명확한 추론(evidential reasoning)을 구현하는 데에 액티브 온톨로지(1050), 도메인 모델들(2756), 및 태스크 플로우 모델들(2786)로부터의 데이터가 이용될 수 있다.Data from the application preferences and usage history 1072, conversation history and auxiliary memory 1052, and / or personal database 1058 are used by language interpreter 2770 to generate candidate syntactic parses 212 It can also be used. Such data may be obtained, for example, from short term memory and / or long term memory 2752, 2754. In this manner, information known to the user and / or to the input that was previously provided in the same session may be used to enhance performance, reduce ambiguity, and enhance the interactive nature of the interaction. In addition, data from active ontology 1050, domain models 2756, and task flow models 2786 to implement evidential reasoning in determining valid candidate syntactic parses 212 Can be used.

의미론적 매칭(220)에서, 언어 해석기(2770)는 가능한 파싱 결과들의 조합을, 이들이 도메인 모델들 및 데이터베이스들과 같은 의미론적 모델들에 얼마나 잘 부합하는지에 따라 고려한다. 의미론적 매칭(220)은, 예를 들어, 액티브 온톨로지(1050), 단기 개인 메모리(2752), 및 장기 개인 메모리(2754)로부터의 데이터를 사용할 수 있다. 예를 들어, 의미론적 매칭(220)은 대화의 장소들 또는 로컬 이벤트(대화 히스토리 및 보조 메모리(1052)) 또는 개인의 가장 선호하는 장소들(애플리케이션 선호도 및 사용 히스토리(1072))에 대한 이전 참조로부터의 데이터를 이용할 수 있다. 의미론적 매칭(220) 단계는 또한 어구들을 도메인 의도 구조(domain intent structure)들로 해석하는 데에 컨텍스트(1000)를 이용한다. 후보, 또는 잠재적인 의미론적 파스 결과들의 세트가 생성된다(222).In semantic matching 220, language interpreter 2770 considers the combination of possible parsing results according to how well they match semantic models such as domain models and databases. Semantic matching 220 may use data from, for example, active ontology 1050, short-term private memory 2752, and long-term private memory 2754. For example, the semantic matching 220 can be used to identify the locations of the conversations or previous references to the local events (conversation history and auxiliary memory 1052) or the individual's most preferred locations (application preferences and usage history 1072) Can be used. The semantic matching step 220 also uses the context 1000 to interpret the phrases into domain intent structures. A set of candidates, or potential semantic parsing results is generated 222.

명확화 단계(disambiguation step)(230)에서, 언어 해석기(2770)는 후보 의미론적 파스 결과들(222)의 입증 강도에 중점을 둔다. 명확화(230)는 바라지 않거나 중복의 대안들을 제거함으로써 후보 의미론적 파스(222)의 수를 줄이는 것에 관련된다. 명확화(230)는, 예를 들어, 액티브 온톨로지(active ontology)(1050)의 구조로부터의 데이터를 이용할 수 있다. 적어도 일 실시예에서, 액티브 온톨로지 내의 노드들 사이의 접속들은 후보 의미론적 파스 결과들(222) 중에서 명확화를 위한 입증 지지를 제공한다. 일 실시예에서, 컨텍스트(1000)는 명확화와 같은 것을 돕는데 사용된다. 그러한 명확화의 예시들은: 동일한 이름을 갖는 몇몇 사람들 중 하나를 결정하는 단계; "응답(reply)"(이메일 또는 텍스트 메세지)과 같은 명령어에 대한 지시 대상(referent)을 결정하는 단계; 대명사의 역참조(pronoun dereferencing) 등을 포함한다.In the disambiguation step 230, the language interpreter 2770 focuses on the proof strength of the candidate semantic pars results 222. Clarification 230 relates to reducing the number of candidate semantic parses 222 by eliminating unwanted or redundant alternatives. Clarification 230 may utilize data from, for example, the structure of an active ontology 1050. In at least one embodiment, connections between nodes in the active ontology provide proof support for clarification among the candidate semantic pars results 222. [ In one embodiment, the context 1000 is used to assist with such things as clarification. Examples of such clarification include: determining one of several persons having the same name; Determining a referent to a command such as "reply" (e-mail or text message); And pronoun dereferencing of pronouns.

예를 들어, "허브 호출(call Herb)"와 같은 입력은 "허브"와 매칭하는 임의의 엔티티(entity)를 잠재적으로 나타낸다. 그러한 엔티티들은 사용자의 어드레스 북(개인 데이터베이스들(1058)) 뿐만 아니라, 개인 데이터베이스들(1058) 및/또는 도메인 엔티티 데이터베이스(2772)로부터의 사업자의 이름들의 데이터베이스 내에 임의의 수만큼 있을 수 있다. 컨텍스트의 몇몇 소스들은 단계(232)에서 "허브들"을 매칭하는 세트를 제한하고/하거나, 그것들을 순위를 매기고 필터링할 수 있다. 예를 들어:For example, an input such as "call Herb" potentially represents any entity that matches the "hub ". Such entities may be any number in the database of business names from personal databases 1058 and / or domain entity database 2772, as well as the address book (personal databases 1058) of the user. Some sources in the context may limit and / or rank and match the sets that match the "hubs " at step 232. E.g:

인기 전화 번호 리스트 상에 있거나 최근에 호출된 허브와 같은 다른 애플리케이션 선호들 및 사용 히스토리(1072), 또는 텍스트 메세지 대화 또는 이메일 스레드(email thread)에 대한 최근 단체(party);

Recent application preferences and usage history 1072, such as a hub that is on the popular phone number list or recently called, or a recent party to a text message conversation or email thread;

아버지 또는 형제와 같이, 관계로서 이름지어진 허브같은, 개인 데이터베이스(1058)에서 언급된 허브, 또는 최근 캘린더 이벤트의 리스트된 참석자. 태스크가 전화 호출 대신에 미디어 재생인 경우, 미디어 제목, 제작자 등으로부터의 이름들은 제한의 소스들이 될 수 있다;

A hub referred to in a personal database 1058, such as a hub named as a relationship, such as a father or a sibling, or a listed attendee of a recent calendar event. If the task is media playback instead of a phone call, names from the media title, author, etc. may be sources of restrictions;

요청이나 결과들 내의 대화(1052)의 최근 플라이(ply). 예를 들어, 도 25a 내지 25b에 연계하여 위에서 설명된 바와 같이, 존(John)으로부터의 이메일을 검색한 후에, 대화 컨텍스트 내에 여전히 있는 검색 결과를 이용하여, 사용자는 답변을 작성할 수 있다. 비서(1002)는 특정한 애플리케니션 도메인 객체 컨텍스트를 식별하기 위해 대화 컨텍스트를 사용할 수 있다.

Recent ply of conversation 1052 in request or results. For example, after retrieving an email from John, as described above in connection with Figures 25a-25b, the user can create a reply using the search results that are still in the conversation context. Secretary 1002 may use the conversation context to identify a particular application domain object context.

또한, 컨텍스트(1000)는 적절한 이름들 외의 단어들에서의 모호성을 감소시키는 것을 도울 수 있다. 예를 들어, (도 20에 도시된 바와 같이) 이메일 애플리케이션의 사용자가 비서(1002)에게 "응답(reply)"이라고 말하는 경우, 애플리케이션의 컨텍스트는 단어가 텍스트 메세지 응답의 반대에 있는 것처럼 이메일 응답과 연관되어야 한다는 것을 결정하는 것을 돕는다. Also, the context 1000 may help to reduce ambiguity in words other than proper names. For example, if the user of the email application (as shown in FIG. 20) speaks to the secretary 1002 as "reply, " the context of the application is the email response Helping to determine that they should be associated.

단계(232)에서, 언어 해석기(2770)는 사용자 의도(290)의 표현으로서의 최고 의미론적 파스들을 필터링하고 분류한다(232). 컨텍스트(1000)는 그러한 필터링과 분류(232)를 공지하는데 사용될 수 있다. 결과는 사용자 의도(290)의 표현이다. At step 232, the language interpreter 2770 filters and classifies the highest semantic pars as a representation of the user intent 290 (232). Context 1000 may be used to announce such filtering and classification 232. The result is a representation of the user intent 290.

태스크 task 플로우Flow 처리에서의 In treatment 컨텍스트Context 사용 use

이제 도 5를 참조하면, 일 실시예에 따라, 대화 플로우 프로세서(2780)에 의해 수행될 수 있는 것과 같이. 태스크 플로우 처리에서의 컨텍스트를 사용하는 방법을 도시하는 플로우 다이어그램이 도시된다. 태스크 플로우 처리에서, 도 4의 방법으로부터 발생된 후보 파스들은 실행될 수 있는 동작의 태스크 설명들을 만들어 내기 위해 서열화되고 예시 되어진다.Referring now to FIG. 5, and in accordance with one embodiment, as may be done by the conversation flow processor 2780, A flow diagram illustrating a method of using a context in task flow processing is shown. In task flow processing, the candidate fragments generated from the method of FIG. 4 are sequenced and illustrated to produce task descriptions of the actions that can be performed.

방법이 시작된다(300). 사용자 의도(290)의 다수의 후보 표현이 수신된다. 일 실시예에서, 도 4와 연계하여 설명된 바와 같이, 사용자 의도(290)의 표현들은 의미론적 파스들의 세트를 포함한다. The method begins (300). A number of candidate representations of user intent 290 are received. In one embodiment, as described in connection with FIG. 4, the expressions of user intent 290 comprise a set of semantic pars.

단계(312)에서, 대화 플로우 프로세서(2780)는 사용자 의도의 결정에 기초하여, 수행하기 위한 태스크 및 그것의 파라미터를 결정하기 위한 다른 정보와 함께 의미론적 파스(들)의 선호되는 해석을 결정한다. 정보는, 예를 들어, 도메인 모델들(2756), 태스크 플로우 모델들(2786) 및/또는 대화 플로우 모델들(2787), 또는 그것들의 조합으로부터 획득될 수 있다. 예를 들어, 태스크는 전화 걸기일 수 있고, 태스크 파라미터는 전화하기 위한 전화번호이다. In step 312, the conversation flow processor 2780 determines a preferred interpretation of the semantic pars (s), along with other information for determining the task and its parameters to perform, based on the determination of the user's intention . The information may be obtained, for example, from domain models 2756, task flow models 2786 and / or conversation flow models 2787, or a combination thereof. For example, the task may be a telephone call, and the task parameter is a telephone number for calling.

일 실시예에서, 컨텍스트(1000)는 초기값들을 추론하여 모호성을 해결함으로써 파라미터들(312)의 결합을 가이드하기 위해 수행 단계(312)에서 사용된다. 예를 들어, 컨텍스트(1000)는 태스크 설명과 사용자 의도의 최고의 해석이 있는지를 결정하는 것의 예시를 가이드할 수 있다. In one embodiment, the context 1000 is used in an execution step 312 to guide combinations of parameters 312 by deducing initial values to resolve ambiguities. For example, the context 1000 may guide an example of a task description and determining whether there is a best interpretation of the user's intent.

예를 들어, 의도 입력들(290)이 "전화걸기(RebeccaRichards)" 및 전호 걸기(HerbGowen)"라고 가정해 보라. 전화 걸기 태스크는 PhoneNumber 파라미터를 요구한다. 몇몇의 컨텍스트(100)의 소스들은 Rebecca 및 Herb에 대한 어떤 전화번호가 작동할지를 결정하도록 구성될 수 있다. 이 예시에서, 연락처 데이터베이스 내의 Rebecca를 위한 어드레스 북 엔트리는 두 개의 전화번호들을 가지고, Herb를 위한 엔트리는 전화번호들은 가지지 않고, 하나의 이메일 어드레스를 가진다. 연락처 데이터베이스와 같은 개인 데이터베이스(1058)로부터의 컨텍스트 정보(1000)는 가상 비서(1002)로 하여금 Herb보다 Rebecca를 선호하게 하는데, 이는 Rebecca의 전화번호는 있지만, Herb의 전화번호는 없기 때문이다. Rebecca를 위해 어떤 전화번호를 사용할지를 결정하기 위해, 애플리케이션 컨텍스트(1060)는 Rebecca와 텍스트 메세지 대화를 수행하는데 사용된 번호를 선택하도록 컨설팅할 수 있다. 따라서, Rebecca Rechards와 의 텍스트 메세지 대화의 컨텍스트 내의 "call her"는 Rebecca가 텍스트 메세지용으로 사용하고 있는 이동 전화에 전화를 걸라는 것임을 가상 비서(1002)는 결정할 수 있다. 이 특별한 정보는 단계(390)에서 리턴된다. For example, suppose the Intent inputs 290 are "RebeccaRichards" and "HerbGowen." The dialing task requires the PhoneNumber parameter. Some of the sources of the context 100 are Rebecca And an address book entry for Rebecca in the contact database has two telephone numbers, entries for Herb do not have telephone numbers, and one for Herb. Context information 1000 from a personal database 1058, such as a contact database, allows virtual assistant 1002 to favor Rebecca over Herb, which is Rebecca's phone number, In order to determine which phone number to use for Rebecca, the application context 1060, You can consult with Rebecca to choose the number used to carry out the text messaging conversation, so "call her" in the context of a text messaging conversation with Rebecca Rechards calls the mobile phone Rebecca uses for text messaging The virtual secretary 1002 can determine that it is a girl. This particular information is returned at step 390.

컨텍스트(1000)는 전화 번호의 모호성을 줄이는 것 이외의 것을 위해 사용될 수 있다. 그것은, 태스크 파라미터를 위한 값을 갖는 컨텍스트(1000)의 임의의 소스가 이용 가능한 한, 태스크 파라미터를 위한 다수의 가능한 값들이 있을 때는 언제든지 사용될 수 있다. 컨텍스트(1000)가 모호성을 줄일 수 있는(그리고 사용자로 하여금 후보들 중에서 선택하도록 하는 것을 피할 수 있는) 다른 예시들은; 이메일 어드레스들; 물리적 어드레스들; 시간과 날짜들; 장소; 리스트 이름; 미디어 제목들; 아티스트 이름들; 비지니스 이름들; 또는 임의의 다른 값 공간을, 제한 없이, 포함한다. Context 1000 may be used for anything other than reducing the ambiguity of the telephone number. It can be used whenever there are a number of possible values for a task parameter, as long as any source of the context 1000 having a value for the task parameter is available. Other examples in which the context 1000 can reduce ambiguity (and avoid having the user choose between candidates); Email addresses; Physical addresses; Time and dates; Place; List name; Media titles; Artist names; Business names; Or any other value space, without limitation.

태스크 플로우 처리(300)를 위해 필요한 다른 종류의 추론들은 컨텍스트(1000)로부터 이득을 취할수도 있다. 예를 들어, 초기값 추론은 현재 위치 시간 및 다른 현재의 값들을 사용할 수 있다. 초기값 추론은 사용자의 요청에서 암시된 태스크 파라미터의 값들을 결정하는데 유용할 수 있다. 예를 들어, 누군가가 "날씨가 어때요?"라고 말하면, 그것은 이 근처의 현재 날씨가 어떤지를 암시적으로 의미한다. Other types of reasoning needed for task flow processing 300 may also benefit from context 1000. For example, the initial value inference can use the current position time and other current values. The initial value inference can be useful in determining the values of the task parameters implied in the user's request. For example, if someone says, "How is the weather?", That implies what the current weather is in the vicinity.

단계(310)에서, 대화 플로우 프로세서(2780)는 사용자 의도의 이 해석이 진행하기에 충분하도록 강하게 지지 되는지 및/또는 대안의 모호한 파스들보다 더 지지 되는지를 결정한다. 만약, 경쟁적인 모호성 또는 충분한 불확실성이 있다면, 그 후에, 대화 플로우 단계를 설정하여 실행 단계가 대화로 하여금 사용자로부터 더 많은 정보에 대한 프롬프트를 출력하도록 하기 위해, 단계(322)가 수행된다. 사용자가 모호성을 해결하도록 하게 하는 스크린 샷이 도 14에 도시된다. 컨텍스트(1000)는, 사용자가 선택하는 후보 아이템들의 디스플레이된 메뉴를 분류하고 주석을 다는 단계(322)에서 사용될 수 있다.In step 310, the conversation flow processor 2780 determines whether this interpretation of user intent is strongly supported and / or supported by alternative ambiguous parcels to proceed. If there is a compelling ambiguity or sufficient uncertainty, then step 322 is performed to set up a dialog flow step so that the action step causes the dialog to output a prompt for more information from the user. A screen shot is shown in FIG. 14 that allows the user to resolve ambiguities. Context 1000 may be used in step 322 to classify and annotate the displayed menu of candidate items that the user selects.

단계(320)에서, 태스크 플로우 모델은 적절한 다음 단계를 결정하도록 컨설팅된다. 정보는, 예를 들어, 도메인 모델들(2756), 태스크 플로우 모델들(2786) 및/또는 대화 플로우 모델(2787), 또는 그것들의 임의의 조합으로부터 획득될 수 있다. At step 320, the task flow model is consulted to determine an appropriate next step. The information may be obtained, for example, from domain models 2756, task flow models 2786 and / or conversation flow model 2787, or any combination thereof.

단계(320) 또는 단계(322)의 결과는 사용자의 요청(390)의 표현이며, 이는 적절한 서비스에 보내기 위한 대화 플로우 프로세서(2780) 및 서비스 편성(2782)에 대해 충분한 태스크 파라미터들을 포함할 수 있다. The result of step 320 or step 322 is a representation of the user's request 390 which may include sufficient task parameters for the conversation flow processor 2780 and service combination 2782 to send to the appropriate service .

대화 생성을 개선하기 위한 To improve conversation creation 컨텍스트의Contextual 사용 use

대화 응답 생성(500) 중에, 비서(1002)는 그것의 사용자의 의도와 태스크에서 그것이 어떻게 동작하는지의 이해를 다시 다른 말로 바꾸어 표현(paraphrase)할 수 있다. 그러한 출력의 예시는 "오케이, Rebecca에게 그녀의 이동전화로 전화를 걸겠습니다...". 이는 사용자로 하여금 비서(1002)가 호출을 위치시키는 것과 같은 연관된 태스크 자동화를 수행하도록 인증하게 하는 것을 허용한다. 대화 생성 단계(500)에서, 비서(1002)는 사용자의 의도의 이해를 다른 말로 표현하는 것에 있어서 사용자에게 얼마나 상세히 전달할 수 있는지를 결정한다. During the conversation response creation 500, the secretary 1002 may again paraphrase an understanding of its intent to the user and how it behaves in the task. An example of such an output is "Okay, I'll call Rebecca on her mobile phone ...". This allows the user to authenticate the secretary 1002 to perform associated task automation such as placing a call. In the dialog creation step 500, the secretary 1002 determines how much detail the user can convey to the user in expressing the understanding of the user's intention.

일 실시예에서, 또한, 컨텍스트(1000)는 대화에 있어서의 상세의 적절한 레벨의 선택을 가이드 하는 것뿐 아니라, (정보를 반복하는 것을 피하기 위해) 이전의 출력에 기초하여 필터링하는 것에 사용될 수 있다. 예를 들어, 비서(1002)는 사람과 전화 번호가 이름을 언급할지에 대한 여부와 레벨의 상세한 정도를 결정하기 위한 컨텍스트(1000)로부터 추론될 수 있다는 사실을 이용할 수 있다. 적용될 수 있는 규칙들의 예시는, 제한 없이 다음 사항들을 포함한다: In one embodiment, context 1000 can also be used to filter based on previous output (to avoid repeating information) as well as to guide selection of an appropriate level of detail in conversation . For example, the secretary 1002 can take advantage of the fact that the person and phone number can be deduced from the context 1000 to determine whether to mention the name and the level of detail. Examples of rules that may be applied include, without limitation, the following:

컨텍스트에 의해 대명사가 해결되었을 때, 전화 걸 사람의 이름을 언급하라.

When the pronoun is resolved by context, mention the name of the person to call.

사람이 텍스트 메세지와 같은 유사 컨텍스트로부터 추론되는 경우, 이름(first name)만을 사용하라.

If a person is inferred from a similar context, such as a text message, use only the first name.

전화 번호가 애플리케이션 또는 개인 데이터 컨텍스트로부터 추론되는 경우, 다이얼 할 실제 번호보다 "이동 전화"와 같은 전화 번호의 상징적 이름을 사용하라.

If the phone number is deduced from the application or personal data context, use the symbolic name of the phone number, such as "mobile phone", rather than the actual number to dial.

상세의 적절한 레벨을 가이드하는 단계에 부가하여, 또한, 컨텍스트(1000)는, 예를 들어, 반복을 피하도록 이전의 발언(utterance)들을 필터링하기 위해, 그리고 대화중의 이전에 언급된 엔티티들을 나타내기 위해 대화 생성 단계(500)에서 사용될 수 있다. In addition to guiding the appropriate level of detail, the context 1000 may also include, for example, filtering out previous utterances to avoid repetition, and displaying previously mentioned entities in a conversation May be used in the dialogue generation step 500.

본 기술분야의 숙련자는 컨텍스트(1000)가 다른 방식들로 사용될 수도 있다는 것을 알게 될 것이다. 예를 들어, 이곳에 설명된 기술들과 연계하여, 컨텍스트(1000)는 전체 명세서가 참조로서 이곳에 통합되고, 2009년 6월 5일에 출원된, 대리인 문서 번호 P7393US1인 "Contextual Voice Commands"에 관한 관련 U.S 유틸리티 출원 번호 제12/479,477호에 설명된 메카니즘에 따라 이용될 수 있다. Those skilled in the art will recognize that the context 1000 may be used in other ways. For example, in conjunction with the techniques described herein, the context 1000 may be incorporated into the "Contextual Voice Commands ", Attorney Docket No. P7393US1, filed June 5, 2009, May be utilized in accordance with the mechanism described in related US utility application No. 12 / 479,477.

컨텍스트Context 수집 및 통신 Collection and communication 메카니즘Mechanism

다양한 실시예들에서, 상이한 메카니즘들이 가상 비서(1002) 내의 컨텍스트 정보를 수집하고 통신하는데 이용된다. 예를 들어, 일 실시예에서, 가상 비서(1002)는 클라이언트/서버 환경에서 구현되어 그것의 서비스들이 클라이언트와 서버 사이에 분포되며, 컨텍스트(1000)의 소스들이 분포될 수도 있다. In various embodiments, different mechanisms are used to collect and communicate contextual information within virtual secretary 1002. [ For example, in one embodiment, the virtual secretary 1002 is implemented in a client / server environment and its services are distributed between the client and the server, and the sources of the context 1000 may be distributed.

이제 도 6을 참조하면, 일 실시예에 따른, 클라이언트(1304)와 서버(1340) 사이의 컨텍스트(1000)의 소스들의 분포의 예시가 도시된다. 이동 컴퓨팅 디바이스 또는 다른 디바이스일 수 있는 클라이언트 디바이스(1304)는 디바이스 센서 데이터(1056), 현재 애플리케이션 컨텍스트(1060), 이벤트 컨텍스트(2706) 등과 같은 컨텍스트 정보(1000)의 소스가 될 수 있다. 컨텍스트(1000)의 다른 소스들은 클라이언트(1304) 또는 서버(1340) 또는 양쪽 모두의 일부 조합 상에 분포될 수 있다. 예시들은 애플리케이션 선호도 및 사용 히스토리(1072c, 1072s); 대화 히스토리 및 보조 메모리(1052c, 1052s); 개인 데이터베이스들(1058c, 1058s); 및 개인 음향 컨텍스트 데이터(1080c, 1080s)를 포함한다. 이들 예시들의 각각에서, 컨텍스트(1000)의 소스들은 서버(1340), 클라이언트(1304), 또는 둘 모두 상에 존재할 수 있다. 더욱이, 위에서 설명된 바와 같이, 도 2에서 도시된 다양한 단계들은 클라이언트(1304) 또는 서버(1340), 또는 둘 모두의 일부 조합에 의해 수행될 수 있다.6, an example of the distribution of the sources of context 1000 between client 1304 and server 1340 is shown, according to one embodiment. Client device 1304, which may be a mobile computing device or other device, may be the source of context information 1000, such as device sensor data 1056, current application context 1060, event context 2706, and the like. Other sources of context 1000 may be distributed over some combination of client 1304 or server 1340 or both. Examples include application preferences and usage histories 1072c and 1072s; Conversation history and auxiliary memory 1052c and 1052s; Personal databases 1058c and 1058s; And personal acoustic context data 1080c, 1080s. In each of these examples, the sources of context 1000 may reside on server 1340, client 1304, or both. Moreover, as described above, the various steps shown in FIG. 2 may be performed by some combination of client 1304 or server 1340, or both.

일 실시예에서, 컨텍스트(1000)는 클라이언트(1304) 및 서버(1340)와 같은 분포된 컴포넌트들 사이에서 통신될 수 있다. 그러한 통신은 로컬 API 또는 분산 네트워크, 또는 일부 다른 수단들에 의한 것일 수 있다. In one embodiment, context 1000 may be communicated between distributed components, such as client 1304 and server 1340. Such communication may be by a local API or a distributed network, or some other means.

이제 도 7a 내지 도 7d를 참조하면, 다양한 실시예들에 따라, 컨텍스트 정보(1000)를 획득하고 코디네이트하기 위한 메카니즘들의 예시들을 도시하는 이벤트 다이어그램들이 도시된다. 다양한 기술들이 컨텍스트를 로딩 또는 통신하기 위해 존재하여, 필요할 때 또는 유용할 때, 가상 비서(1002)에 대해 이용 가능하게 된다. 이들 메카니즘들 각각은 가상 비서(1002)의 동작; 디바이스 또는 애플리케이션 초기화(601); 초기 사용자 입력(602); 초기 입력 처리(603); 및 컨텍스트에 의존하는 처리(604)에 관해 위치할 수 있는 4 개의 이벤트들에 관하여 설명된다. Referring now to Figures 7A-7D, there are illustrated event diagrams illustrating examples of mechanisms for obtaining and coordinating contextual information 1000, in accordance with various embodiments. Various techniques are available for loading or communicating contexts and are available to the virtual secretary 1002 when needed or useful. Each of these mechanisms includes an operation of the virtual secretary 1002; Device or application initialization 601; Initial user input 602; Initial input processing 603; And the context-dependent processing 604. [0060]

도 7a는 사용자 입력(602)이 시작되는 경우, 컨텍스트 정보(1000)가 "풀(pull)" 메카니즘을 이용하여 로딩되는 접근법을 도시한다. 사용자가 가상 비서(1002)를 호출하고, 적어도 일부 입력(602)을 제공하는 경우, 가상 비서(1002)는 컨텍스트(1000)를 로딩한다(610). 로딩(610)은 적절한 소스로부터 컨텍스트 정보(1000)를 요청하거나 검색하는 것에 의해 수행될 수 있다. 입력 처리(603)는 컨텍스트(1000)가 로딩된(610) 경우 시작한다.7A illustrates an approach in which context information 1000 is loaded using a "pull" mechanism when user input 602 is initiated. If the user calls virtual assistant 1002 and provides at least some input 602, virtual assistant 1002 loads context 1000 (610). Loading 610 may be performed by requesting or retrieving context information 1000 from an appropriate source. The input process 603 begins when the context 1000 is loaded (610).

도 7b는 디바이스 또는 애플리케이션이 초기화되었을 때(601), 일부 컨텍스트 정보(1000)가 로딩되며(620), 사용자 입력이 시작되는 경우(602), 부가적인 컨텍스트 정보(1000)가 풀 메카니즘을 이용하여 로딩되는 접근법을 도시한다. 일 실시예에서, 초기화에서 로딩된(620) 컨텍스트 정보(1000)는 정적 컨텍스트(즉, 자주 바뀌지 않는 컨텍스트)를 포함하며; 사용자 입력이 시작된 경우(602), 로딩되는(621) 컨텍스트 정보(1000)는 동적 컨텍스트(즉, 정적 컨텍스트가 로딩 되었기(620) 때문에 바뀔 수 있는 컨텍스트)를 포함한다. 그러한 접근법은 시스템의 런타임 성능으로부터 정적 컨텍스트 정보(1000)를 로딩하는 것의 비용을 제거함으로써 성능을 개선시킬 수 있다.7B illustrates an exemplary case where the context information 1000 is loaded (620) when a device or application is initialized (601), when the user input is initiated (602), additional context information (1000) Lt; / RTI > In one embodiment, the context information 1000 loaded in initialization 620 includes a static context (i.e., a context that does not change often); When user input is initiated (602), the context information (1000) that is loaded (621) includes a dynamic context (i.e., a context that may be changed due to the static context being loaded (620)). Such an approach can improve performance by eliminating the cost of loading static context information 1000 from the runtime capabilities of the system.

도 7c는 도 7b의 접근법의 변형예를 도시한다. 본 예에서, 동적 컨텍스트 정보(1000)는 입력 처리가 시작(603)된 후 로딩(621)의 지속을 허용한다. 따라서, 로딩(621)은 입력 처리와 병렬로 일어날 수 있다. 가상 비서(1002) 절차는 처리가 수신된 컨텍스트 정보(1000)에 의존할 때 단계(604)에서 단지 차단된다.Figure 7c shows a variation of the approach of Figure 7b. In this example, the dynamic context information 1000 allows the continuation of the loading 621 after the input process has begun (603). Thus, the loading 621 may occur in parallel with the input processing. The virtual secretary 1002 procedure is only blocked at step 604 when the process is dependent on the received context information 1000.

도 7d는 아래와 같은 5가지 상이한 방식 중 어느 하나로 컨텍스트를 다루는 완전 구성가능한 버전을 도시한다:Figure 7d shows a fully configurable version that deals with the context in any of the following five different ways:

정적 컨텍스트(static contextual) 정보(1000)는 컨텍스트 소스에서 가상 비서(1002)를 실행하는 환경 또는 디바이스로의 일 방향으로 동기화(640)된다. 데이터가 컨텍스트 소스에서 변경될 때, 그 변경은 가상 비서(1002)에게 푸시(push)된다. 예컨대, 어드레스 북은 초기에 생성 또는 인에이블될 때 가상 비서(1002)에 동기화된다. 어드레스 북이 수정될 때마다, 변경은 가상 비서(1002)에게 즉시 또는 일괄 접근 방식으로 푸시된다. 도 7d에 도시된 바와 같이, 이런 동기화(640)는 사용자 입력이 시작(602)되기 전을 포함하는 어느 때나 일어날 수 있다.

Static contextual information 1000 is synchronized 640 in one direction to the environment or device executing the virtual secretary 1002 at the context source. When the data is changed at the context source, the change is pushed to the virtual secretary 1002. For example, the address book is synchronized to the virtual secretary 1002 when it is initially created or enabled. Each time the address book is modified, the changes are pushed to the virtual secretary 1002 either immediately or in a batch approach. As shown in FIG. 7D, this synchronization 640 may occur at any time, including before the user input begins (602).

일 실시예에서, 사용자 입력이 시작(602)될 때, 정적 컨텍스트 소스는 동기화 상태를 체크할 수 있다. 필요한 경우, 나머지 정적 컨텍스트 정보(1000)를 동기화하는 프로세스가 시작된다(641).

In one embodiment, when user input is initiated (602), the static context source may check the synchronization state. If necessary, the process of synchronizing the remaining static context information 1000 begins (641).

사용자 입력이 시작(602)될 때, 일부 동적 컨텍스트(1000)는 사실상 610 및 621에서 로딩(642)된다. 컨텍스트(1000)를 소비하는 절차들은 이들이 필요로 하는 아직까지 로딩되지 않은 컨텍스트 정보(1000)를 대기하도록 단지 차단된다.

When the user input is started 602, some dynamic context 1000 is actually loaded (642) at 610 and 621. Procedures consuming the context 1000 are simply blocked to wait for context information 1000 that has not yet been loaded that they need.

다른 컨텍스트 정보(1000)는 이들이 필요로 할 때 프로세스에 의해 온디맨드(On demand)(643)식으로 로딩된다.

Other context information 1000 is loaded on demand 643 by the process when they are needed.

이벤트 컨텍스트(2706)는 이벤트가 일어날 때 소스에서 가상 비서(1002)를 실행하는 디바이스로 전송(644)된다. 이벤트 컨텍스트(2706)를 소비하는 프로세스는 준비될 이벤트의 캐시를 단지 대기하며, 그 이후 어떠한 시간도 차단함이 없이 처리될 수 있다. 이런 식으로 로딩된 이벤트 컨텍스트(2706)는 다음 중 어느 하나를 포함할 수 있다:

The event context 2706 is transmitted (644) from the source to the device executing the virtual secretary 1002 when the event occurs. The process consuming the event context 2706 may just be waiting for a cache of events to be prepared and then processed without blocking any time thereafter. The event context 2706 loaded in this way may include any of the following:

사용자 입력이 시작(602)되기 전에 로딩된 이벤트 컨텍스트(2706), 예컨대, 비판독 메시지 통지. 이런 정보는 예컨대 동기화된 캐시를 이용하여 유지될 수 있다.

A loaded event context 2706, e.g., a non-read message notification, before user input is initiated (602). This information can be maintained, for example, using a synchronized cache.

사용자 입력의 시작(602)과 동시 또는 그 이후에 로딩된 이벤트 컨텍스트(2706). 예컨대, 사용자가 가상 비서(1002)와 대화중, 텍스트 메시지가 도착할 수 있다; 비서(1002)에게 이런 이벤트를 통지하는 이벤트 컨텍스트는 비서(1002) 처리에 병렬로 푸시될 수 있다.

Event context 2706 loaded at or after the start (602) of user input. For example, a text message may arrive while the user is talking to the virtual secretary 1002; The event context for notifying the secretary 1002 of such an event can be pushed in parallel to the secretary 1002 process.

일 실시예에서, 컨텍스트 정보(1000)를 획득하고 조정하는 유연성은, 각각의 컨텍스트 정보(1000)의 소스에 대해, 모든 요청에 이용가능한 정보를 갖는 값에 대한 통신 비용을 밸런싱하는(balance) 액세스 API 및 통신 정책(policy)을 규정함으로써 달성된다. 예컨대, 모든 음성-텍스트 변환 요청에 관련된 변수들, 즉 마이크로폰의 파라미터를 기술하는(describe) 디바이스 센서 데이터(1056) 또는 개인용 음향 컨텍스트 데이터(1080)가 모든 요청에 대해 로딩될 수 있다. 이런 통신 정책은 예컨대 구성 테이블에서 특정될 수 있다.In one embodiment, the flexibility to obtain and adjust the context information 1000 is such that, for each source of context information 1000, the balance of communication costs for values with information available to all requests API, and communication policy. For example, device sensor data 1056 or personal acoustic context data 1080 describing the parameters associated with all voice-to-text conversion requests, i.e., the parameters of the microphone, may be loaded for every request. Such a communication policy can be specified, for example, in a configuration table.

도 9를 참고하면, 일 실시예에 따르는, 컨텍스트 정보(1000)의 여러 소스에 대한 캐싱 정책(caching policy) 및 통신을 특정하는데 사용될 수 있는 구성 테이블(900)의 예가 도시된다. 사용자 이름, 어드레스 북 이름, 어드레스 북 번호, SMS 이벤트 컨텍스트, 및 캘린더 데이터베이스를 포함하는 다수의 상이한 컨텍스트 소스들 각각에 대해서, 컨텍스트 로딩의 특정 타입은 도 2의 단계들 각각에서 특정된다: 음성 정보를 유도하고 해석한다(100); 자연어를 해석한다(200); 태스크를 식별한다(300); 대화 응답을 생성한다(500). 테이블(900) 내의 각각의 엔트리는 다음 중 하나를 나타낸다:Referring to FIG. 9, an example of a configuration table 900 that may be used to specify a caching policy and communication for various sources of context information 1000, according to one embodiment, is shown. For each of a number of different context sources, including a user name, an address book name, an address book number, an SMS event context, and a calendar database, the specific type of context loading is specified in each of the steps of Figure 2: Induce and interpret (100); Interpret natural language (200); Identify a task (300); A dialog response is generated (500). Each entry in the table 900 represents one of the following:

동기화(Sync): 컨텍스트 정보(1000)는 디바이스 상에서 동기화된다.

Sync: The context information 1000 is synchronized on the device.

온 디맨드(On demand): 컨텍스트 정보(1000)는 가상 비서(1002)의 요청에 응답하여 제공된다.

On demand: Context information 1000 is provided in response to a request of the virtual secretary 1002.

푸시(Push): 컨텍스트 정보(1000)는 디바이스에 푸시된다.

Push: Context information 1000 is pushed to the device.

완전 구성가능한 방법에서는 잠재적으로 관련된 컨텍스트 정보(1000)의 큰 공간이 인간과 기계 사이의 자연어 대화(interaction)를 스트림라인(streamline)하는 것을 가능하게 한다. 비호율성을 초래할 수 있는 모든 시간에 이런 정보 모두를 로딩하기보다는, 일부 정보는 컨텍스트 소스 및 가상 비서(1002) 모두에서 유지되고, 반면에 다른 정보는 온디맨드식으로 쿼리(query)된다. 예컨대, 전술한 바와 같이, 음성 인식과 같은 실시간 동작에 사용되는 이름과 같은 정보는 국부적으로 유지되며, 반면에 사용자의 개인용 캘린더와 같은 일부 가능한 요청에 의해서만 사용되는 정보는 주문형으로 쿼리된다. 수신 SMS 이벤트와 같은 사용자가 비서를 호출(invoking)할 때 기대할 수 없는 데이터는 이벤트가 발생될 때 푸시된다.In a fully configurable manner, a large space of potentially related context information 1000 enables streamline natural language interaction between a human and a machine. Rather than loading all of this information at all times that may result in incompatibility, some information is maintained in both the context source and the virtual secretary 1002, while other information is queried on demand. For example, as described above, information such as names used in real-time operations such as voice recognition is maintained locally, while information used only by some possible requests, such as a user's personal calendar, is queried on demand. Data that can not be expected when a user invokes a secretary, such as an incoming SMS event, is pushed when the event occurs.

도 10을 참고하면, 일 실시예에 따르는, 비서(1002)가 사용자와 대화하는 대화형 시퀀스(interaction sequence)의 처리 동안에, 도 9에 구성된 컨텍스트 정보 소스를 액세싱하는 예를 도시한 이벤트도(950)가 도시된다.10, an event diagram illustrating an example of accessing the context information source configured in FIG. 9 during the processing of an interactive sequence in which a secretary 1002 interacts with a user, according to one embodiment 950 are shown.

도 10에 도시된 시퀀스는 다음의 대화형 시퀀스를 나타낸다:The sequence shown in Figure 10 represents the following interactive sequence:

T₁: 비서(1002): "안녕 스티브, 제가 무엇을 도와드릴까요?"

T ₁ : Secretary (1002): "Hi Steve, what can I do for you?"

T₂: 사용자: "나의 다음 미팅은 언제인가?"

T ₂ : User: "When is my next meeting?"

T₃: 비서(1002): "너의 다음 미팅은 중역 회의실에서 오후 1시입니다"

T ₃ : Secretary (1002): "Your next meeting is at 1:00 pm in the boardroom."

T₄: [수신 SMS 메시지의 사운드]

T ₄ : [Sound of incoming SMS message]

T₅: 사용자:"그 메시지를 나에게 읽어라"

T ₅ : User: "Read the message to me"

T₆: 비서(1002): "존슨으로부터의 너의 메시지는 '점심 어떻게 할까'"입니다"

T ₆ : Secretary (1002): "Your message from Johnson is" how to do lunch "

T₇: 사용자: "존슨에게 오늘은 점심 같이 할 수 없다고 말해라"

T ₇ : User: "Tell Johnson I can not do lunch today"

T₈: 비서(1002): "OK, 그에게 말할 것이다"

T ₈ : Secretary (1002): "OK, I will tell him"

시간 T₀에서, 대화가 시작되기 전에, 사용자 이름이 동기화되며(770), 어드레스 북 이름이 동기화된다(771). 이것은 도 7d의 엘리먼트(640)에 도시된 바와 같이 초기화 시간에 로딩되는 정적 컨텍스트의 예이다. 이는 비서(1002)가 사용자를 그의 성("스티브")으로 부르는 것을 가능하게 한다.At time T _0, before the conversation is started, the user name is synchronized (770), and the address book name is synchronized (771). This is an example of a static context that is loaded at initialization time as shown in element 640 of Figure 7d. This enables the secretary 1002 to call the user with his or her last name ("Steve").

시간 T₁에서, 동기화 단계(770 및 771)가 완료된다. 시간 T₂에서, 사용자는 도 2의 단계 100, 200 및 300에 따라 처리되는 요청을 말한다. 태스크 식별 단계(300)에서, 가상 비서(1002)는 컨텍스트(1000)의 소스로서 사용자 개인용 데이터베이스(1058)에게 쿼리한다(774): 특히, 가상 비서(1002)는 테이블(900)에 따르는 온디맨드 액세스에 대해 구성되는 사용자의 캘린더 데이터베이스로부터 정보를 요청한다. 시간 T₃에서, 단계 500가 수행되고 대화 응답이 생성된다.At time T ₁ , synchronization steps 770 and 771 are complete. At time T ₂ , the user refers to a request being processed in accordance with steps 100, 200 and 300 of FIG. In the task identification step 300, the virtual secretary 1002 queries 774 the user personal database 1058 as the source of the context 1000: in particular, the virtual secretary 1002 sends an on demand And requests information from the user's calendar database configured for access. At time T _3, the step 500 is performed a dialog response is generated.

시간 T₄에서, SMS 메시지가 수신된다: 이는 이벤트 컨텍스트(2706)의 예이다. 테이블(900)의 구성에 기초하여, 이벤트의 통지가 가상 비서(1002)에게 푸시된다(773).At time T ₄ , an SMS message is received: this is an example of an event context 2706. Based on the configuration of the table 900, a notification of the event is pushed to the virtual secretary 1002 (773).

시간 T₅에서, 사용자는 가상 비서(1002)에게 SMS 메시지를 읽도록 요청한다. 이벤트 컨텍스트(2706)의 존재는 수행 단계(200)에서 NLP 컴포넌트에게 "그 메시지"를 새로운 SMS 메시지로서 해석하도록 가이드한다. 시간 T₆에서, 단계 300이 태스크 컴포넌트에 의해 수행될 수 있어 SMS 메시지를 사용자에게 읽어주도록 API를 호출한다. 시간 T₇에서, 사용자는 애매한 동사("tell") 및 이름("Johnny")으로 요청을 행한다. NLP 컴포넌트는 단계(773)에서 수신되는 이벤트 컨텍스트(2706)를 포함하는 컨텍스트(1000)의 다양한 소스를 이용하여 이러한 모호함을 해결함으로써 자연어(200)를 해석한다: 이는 NLP 컴포넌트에게 말을 하여, 커맨드가 개인 이름 조니(Johney)로부터의 SMS 메시지를 참고하게 하는 것이다. 단계 T₇에서, 수신된 이벤트 컨텍스트 객체로부터 사용하는 번호를 탐색함에 의해 이름을 매칭하는 단계(771)를 포함하는 플로우 단계(400)의 실행이 수행된다. 비서(1002)는 따라서 새로운 SMS 메시지를 작성할 수 있으며, 이를 단계 T₈에서 확인한 바와 같이 조니(Johney)에게 전송한다.At time T ₅ , the user requests the virtual secretary 1002 to read the SMS message. The presence of the event context 2706 guides the NLP component in step 200 to interpret the "message" as a new SMS message. At time T ₆ , step 300 may be performed by the task component to invoke the API to read the SMS message to the user. At time T ₇ , the user makes a request with an ambiguous verb ("tell") and a name ("Johnny"). The NLP component interprets the natural language 200 by resolving this ambiguity using various sources of the context 1000 including the event context 2706 received at step 773: it speaks to the NLP component, To refer to an SMS message from his personal name Johnney. T in step _7, the execution of a step 771 for matching the name by the numbers used as the search from the received event context object flow step 400 is performed. Secretary 1002, and thus to create a new SMS message, and transmits it to Johnny (Johney) as identified in step T _8.

본 발명은 가능한 실시예에 대해서는 특별히 상세히 개시한다. 당업자에게는 다른 실시예도 실시가능함을 이해할 것이다. 먼저, 컴포넌트의 특정 명명, 용어의 대문자, 속성, 데이터 구조, 또는 임의의 다른 프로그래밍 또는 구조적 양상은 강제적이거나 중요한 것이 아니고, 발명 또는 그 특징을 구현하는 메카니즘은 다른 이름, 포맷 또는 프로토콜을 가질 수 있다. 더욱이, 시스템은 전술한 바와 같이 하드웨어 및 소프트웨어의 조합으로 또는 전체가 하드웨어 소자로 또는 전체가 소프트웨어 소자로 구현될 수 있다. 또한, 전술한 여러 시스템 컴포넌트들 간의 기능성 특정 분할은 단지 예시적이며 강제적이 아니다: 단일 시스템 컴포넌트에 의해 수행되는 기능은 오히려 다중 컴포넌트에 의해 수행될 수 있으며, 다중 컴포넌트에 의해 수행되는 기능은 오히려 단일 컴포넌트에 의해 수행될 수 있다.The present invention will be described in detail with respect to possible embodiments. It will be understood by those skilled in the art that other embodiments are possible. First, the particular naming of a component, the capitalization of a term, an attribute, a data structure, or any other programming or structural aspect is not mandatory or important, and the mechanism implementing the invention or its features may have a different name, format or protocol . Moreover, the system may be implemented in hardware or software combination as described above, or as a whole as a hardware element or as a whole as a software element. In addition, the functionality specific partitioning between the various system components described above is merely exemplary and not mandatory: the functionality performed by a single system component may rather be performed by multiple components, and the functionality performed by multiple components is rather simple Component. &Lt; / RTI >

여러 실시예에서, 본 발명은 전술한 기술을 수행하기 위한 시스템 또는 방법으로서 단일 또는 임의의 조합으로 구현될 수 있다. 다른 실시예에서, 본 발명은 컴퓨팅 디바이스 또는 다른 전자 디바이스에서의 프로세서로 하여금 전술한 기술을 수행하게 야기하는 비일시적인 컴퓨터 판독가능한 저장 매체 및 이 매체에 인코딩된 컴퓨터 프로그램 코드를 포함하는 컴퓨터 프로그램 제품으로서 구현될 수 있다. In various embodiments, the present invention may be implemented in a single or in any combination as a system or method for performing the techniques described above. In another embodiment, the invention is a computer program product comprising a non-transitory computer-readable storage medium and computer program code encoded on the medium, causing the processor in the computing device or other electronic device to perform the above-described techniques Can be implemented.

명세서에서 "일 실시예" 또는 "하나의 실시예"를 참고하는 것은 실시예와 결합해서 설명된 특정 특징, 구조 또는 특성이 본 발명의 적어도 하나의 실시예에 포함되는 것을 의미한다. 명세서의 여러 곳에서 "하나의 실시예"라는 문구의 출현은 반드시 동일 실시예를 모두 참고할 필요는 없다.Reference in the specification to "one embodiment" or "one embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. The appearances of the phrase "one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.

전술한 설명 중 일부는 컴퓨팅 디바이스의 메모리 내 데이터 비트상의 동작의 알고리즘 및 심볼의 관점에서 제시된다. 이들 알고리즘 설명 및 표현은 다른 당업자에게 작업의 실체를 가장 효율적으로 전달하기 위해 당업자에 의해 사용되는 의미이다. 알고리즘은 이하 일반적으로 소정의 결과를 가져오는 단계(지시)의 일관성있는 시퀀스인 것으로 고려된다. 이 단계는 물리량의 물리적 조작을 요구하는 단계이다. 통상, 반드시 필요치는 않지만, 이들 양은 저장, 전달, 조합, 비교 및 그밖에 조작될 수 있는 전기, 자기, 또는 광학 신호의 형태를 취한다. 이들 신호를 비트, 값, 소자, 심볼, 문자, 용어, 수 등으로 언급하는 것은 통상의 용법의 이유에서 항상 편리하다. 더욱이, 일반성의 상실 없이 모듈 또는 코드 디바이스로서 물리량의 물리적 조작을 요구하는 단계의 소정 구성을 언급하는 것은 언제나 편리하다.Some of the foregoing descriptions are presented in terms of algorithms and symbols of operation on in-memory data bits of a computing device. These algorithmic descriptions and representations are those used by those skilled in the art to most effectively convey the substance of an operation to others skilled in the art. The algorithm is generally considered to be a consistent sequence of steps (instructions) that result in a predetermined result. This step is a step requiring physical manipulation of the physical quantity. Typically, though not necessarily, these quantities take the form of electrical, magnetic, or optical signals that can be stored, transferred, combined, compared, and otherwise manipulated. It is always convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. for reasons of common usage. Moreover, it is always convenient to mention a certain configuration of a step requiring physical manipulation of physical quantities as a module or code device without loss of generality.

그러나, 이들 및 유사한 용어 모두가 적당한 물리량에 연관되며 이들 양에 적용되는 단순한 편리한 라벨임에 명심해야 한다. 특별히 언급되지 않는한 전술한 설명에서 명확한 바와 같이, 전체 명세서에서 "처리", "컴퓨팅", "계산", "표시", 또는 "결정" 등과 같은 용어를 활용하는 논의는 컴퓨터 시스템 메모리 또는 레지스터 또는 이런 정보 저장, 전송 또는 디스플레이 디바이스와 같은 다른 디바이스 내에서 물리(전자)량으로서 표현되는 데이터를 조작 및 변환하는 컴퓨터 시스템, 또는 유사한 전자 컴퓨팅 모듈 및/또는 디바이스의 동작 및 프로세스를 참고함을 이해해야 한다. It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, the discussion utilizing terms such as "processing", "computing", "computing", "display", or " It should be understood that reference is made to the operations and processes of a computer system or similar electronic computing module and / or device that manipulates and transforms data represented as physical (electronic) quantities within other devices, such as an information storage, transmission or display device .

본 발명의 소정 양상은 알고리즘의 형태로 개시된 프로세스 단계 및 지시를 포함한다. 본 발명의 프로세스 단계 및 지시가 소프트웨어, 펌웨어 및/또는 하드웨어로 구현될 수 있으며, 소프트웨어로 구현될 때 각종 운영 체제에 의해 사용되는 상이한 플랫폼상에 상주하도록 다운로드되고 동작할 수 있음에 유의해야 한다.Certain aspects of the invention include process steps and instructions disclosed in the form of algorithms. It should be noted that process steps and instructions of the present invention may be implemented in software, firmware, and / or hardware, and may be downloaded and operated to reside on different platforms used by various operating systems when implemented in software.

본 발명은 또한 전술한 동작을 수행하는 장치에 관한 것이다. 이런 장치는 요구되는 목적을 위해 특별히 구성될 수 있거나, 또는 컴퓨팅 디바이스에 저장된 컴퓨터 프로그램에 의해 선택적으로 활성화 또는 재구성되는 범용 컴퓨팅 디바이스를 포함할 수 있다. 이런 컴퓨터 프로그램은 제한적이지 않게 플로피 디스크, 광 디스크, CD-ROM, 자기 광학 디스크, 판독 전용 메모리(ROM), RAM, EPROM, EEPROM, 자기 또는 광학 카드, ASIC 또는 전자 지시를 저장하는데 적합한 소정 타입의 매체, 및 컴퓨터 시스템 버스에 연결된 각각의 매체를 포함하는 소정 타입의 디스크와 같은 컴퓨터 판독가능한 저장 매체에 저장될 수 있다. 더욱이, 본 명세서에서 언급된 컴퓨팅 디바이스는 단일 프로세서를 포함할 수 있거나, 또는 증가된 컴퓨팅 능력을 위한 다중 프로세서 설계를 채용하는 아키텍쳐를 가질 수 있다. The present invention also relates to an apparatus for performing the above-mentioned operations. Such a device may be specially configured for the required purpose, or it may comprise a general purpose computing device selectively activated or reconfigured by a computer program stored on the computing device. Such computer programs include, but are not limited to, any type of computer readable medium, such as a floppy disk, optical disk, CD-ROM, magneto-optical disk, read only memory (ROM), RAM, EPROM, EEPROM, magnetic or optical card, Readable storage medium, such as any type of disk, including a medium, and a respective medium coupled to a computer system bus. Moreover, the computing devices referred to herein may comprise a single processor, or may have an architecture employing a multiprocessor design for increased computing capability.

본 명세서에 개시된 알고리즘 및 디스플레이는 임의의 특정 컴퓨팅 디바이스, 가상화 시스템 또는 다른 장치에만 고유하게 관련되지 않는다. 다양한 범용 시스템은 개시된 교시에 따라 프로그램으로 사용될 수 있고, 또는 필요한 방법 단계를 수행하기 위해 보다 특별한 장치를 구성하는 것이 더 편리하다고 판명될 수도 있다. 각종 이러한 시스템에 요구되는 구조는 전술한 설명에서 자명할 것이다. 또한, 본 발명은 어떤 특정 프로그래밍 언어를 참고해서 설명되지 않았다. 각종 프로그래밍 언어가 전술한 본 발명의 교시를 구현하는데 사용될 수 있으며 특정 언어의 참고는 본 발명의 가능한 최선의 모드를 개시하고자 제공된 것임을 이해해야 한다.The algorithms and displays disclosed herein are not inherently related to any particular computing device, virtualization system, or other device. A variety of general purpose systems may be used as a program in accordance with the teachings disclosed or it may prove more convenient to construct a more specific apparatus to perform the required method steps. The structures required for various such systems will be apparent from the foregoing description. Furthermore, the present invention has not been described with reference to any particular programming language. It should be appreciated that various programming languages may be used to implement the teachings of the invention described above, and that references to particular languages are provided to disclose the best mode possible of the invention.

따라서, 여러 실시예에서, 본 발명은 컴퓨터 시스템, 컴퓨팅 디바이스, 또는 다른 전자 디바이스 또는 이들의 임의의 조합을 제어하기 위한 소프트웨어, 하드웨어 및/또는 다른 소자로서 구현될 수 있다. 이런 전자 디바이스는, 본 기술 분야에서 공지된 기술에 따르는, 예컨대 프로세서, 입력 디바이스(예를 들어, 키보드, 마우스, 터치패드, 트랙패드, 조이스틱, 트랙볼, 마이크로폰, 및/또는 이들의 임의의 조합), 출력 디바이스(예를 들어, 스크린, 스피커 등), 메모리, 장기 스토리지(예를 들어, 자기 스토리지, 광 스토리지 등), 및/또는 네트워크 연결성을 포함한다. 이런 전자 디바이스는 휴대용 또는 비휴대용일 수 있다. 본 발명을 구현하는데 사용되는 전자 디바이스의 예는, 모바일 폰, 개인 휴대 단말기, 스마트폰, 키오스크(kiosk), 데스크톱 컴퓨터, 랩톱 컴퓨터, 태블릿 컴퓨터, 가전 제품, 가정용 오락 디바이스, 음악 플레이어, 카메라, 텔레비젼, 셋탑 박스, 전자 게임 유닛 등을 포함한다. 본 발명을 구현하는 전자 디바이스는 예컨대 미국 캘리포니아 쿠퍼티노의 애플(사)에서 구입가능한 iOS 또는 MacOS, 또는 디바이스의 사용에 적합한 임의의 다른 운영 체제와 같은 임의의 운영 체제를 사용할 수 있다.Thus, in various embodiments, the invention may be implemented as software, hardware, and / or other components for controlling a computer system, a computing device, or other electronic device, or any combination thereof. Such an electronic device may be, for example, a processor, an input device (e.g., a keyboard, a mouse, a touchpad, a trackpad, a joystick, a trackball, a microphone, and / or any combination thereof) in accordance with techniques known in the art. , Output devices (e.g., a screen, speakers, etc.), memory, long-term storage (e.g., magnetic storage, optical storage, etc.), and / or network connectivity. Such electronic devices may be portable or non-portable. Examples of electronic devices that may be used to implement the invention include mobile phones, personal digital assistants, smart phones, kiosks, desktop computers, laptop computers, tablet computers, home appliances, home entertainment devices, music players, , A set-top box, an electronic game unit, and the like. An electronic device embodying the invention may use any operating system, such as iOS or MacOS available from Apple, Inc. of Cupertino, CA, or any other operating system suitable for use with the device.

본 발명이 제한된 수의 실시예로 설명된다 할지라도, 당업자에게는 개시된 본 발명의 범위를 벗어남이 없이 다른 실시예가 고안될 수 있음을 이해할 것이다. 또한, 본 명세서에서 사용되는 언어는 판독가능성 및 도구적인 목적에서 선택된 것이고, 본 발명의 주제를 제한하고 억제하고자 선택된 것은 아니다. 따라서, 본 발명의 개시는 예시적인 것으로 의도된 것이고 이하 특허청구범위 개시된 발명의 범위를 제한하는 것이 아니다.Although the present invention has been described with a limited number of embodiments, it will be understood by those skilled in the art that other embodiments may be devised without departing from the scope of the invention disclosed herein. Also, the language used herein is selected for readability and tooling purposes, and is not intended to limit or inhibit the subject matter of the present invention. Accordingly, the disclosure of the present invention is intended to be illustrative and not to limit the scope of the invention as hereinafter claimed.

Claims

43. A computer readable storage medium having encoded computer program code for implementing a method for interpreting user input to perform a task on a computing device having at least one processor,
The computer program code comprising at least one processor,
Causing the output device to prompt a user for input;
Receiving spoken user input via an input device;
The method comprising: receiving context information from a context source, the context information including acoustic environment data describing an acoustic environment in which the voice user input is received;
Analyzing the received voice user input to derive a representation of a user's intent, the interpreting step comprising: based on the acoustic environment data of the received context information to derive an expression of the user's intent Disambiguating the received voice user input;
Identifying at least one task and at least one parameter for the task based at least in part on the derived representation of the user intent;
Executing at least one task using the at least one parameter to derive a result;
Generating an interactive response based on the derived result; And
And causing the output device to output the generated dialog response, the method comprising:
Wherein the computer program code further comprises at least one processor configured to cause the output device to prompt a user for input using the received context information, interpret the received voice user input, Identifying at least one parameter for the task and the task, and generating an interactive response. &Lt; Desc / Clms Page number 22 >

A system for interpreting user input to perform a task,
An output device configured to provide a prompt to the user for input;
An input device configured to receive a voice user input;
At least one processor communicatively coupled to the output device and the input device,
/ RTI >
Wherein the at least one processor comprises:
Receiving context information from a context source, the context information including acoustic environment data describing an acoustic environment in which the voice user input is received;
Analyzing the received voice user input to derive a representation of the user's intent, the interpretation including interpreting the received voice user input based on the acoustic environment data of the received context information, Including disambiguating input;
Identifying at least one task and at least one parameter for the task based at least in part on the derived representation of the user intent;
Executing at least one task using at least one parameter to derive a result; And
And generate an interactive response based on the derived result,
The output device is further configured to output a generated conversation response,
Using the received context information to provide the user with a prompt for input, interpreting the received voice user input, identifying at least one task and at least one parameter for the task, And generating a conversation response is performed.

delete

3. The method of claim 2,
Wherein the output device is configured to provide a prompt to the user via an interactive interface,
Wherein the input device is configured to receive the voice user input via the interactive interface,
Wherein the at least one processor is configured to convert the speech user input to a textual representation.

5. The apparatus of claim 4, wherein the at least one processor is further configured to generate a plurality of candidate text interpretations of the voice user input and at least rank the subset of the generated candidate text interpretations, And configured to convert user input to a textual representation,
Wherein at least one of generating and sequencing is performed using the received context information.

6. The method of claim 5, wherein the received context information used in at least one of generating and sequencing comprises: the acoustic environment data describing an acoustic environment in which the speech user input is received; A vocabulary acquired from a database associated with the user, a vocabulary associated with an application preference, a vocabulary obtained from a usage history, and a current conversation state.

3. The system of claim 2, wherein the output device is configured to provide a prompt to a user by generating at least one prompt based at least in part on the received context information.

3. The method of claim 2, wherein the at least one processor is configured to disambiguate the received voice user input based on the acoustic environment data of the received context information to determine the received context information based at least in part on the received context information A user input interpretation system configured to derive a user's intent by performing natural language processing on a voice user input.

9. The method of claim 8, wherein the received context information used to disambiguate the received voice user input includes data describing an event, an application context, an input previously provided by the user, And at least one selected from the group consisting of location, date, environmental condition, and history.

3. The method of claim 2, wherein the at least one processor is configured to determine at least one task and at least one parameter for the task based at least in part on the received context information to determine at least one task and at least one And to identify parameters of the user input.

11. The method of claim 10, wherein the received context information used to identify at least one task and at least one parameter for the task comprises at least one of: data describing an event; data from a database associated with the user; At least one selected from the group consisting of received data, application context, input previously provided by the user, known information about the user, location, date, environmental conditions, and history.

3. The system of claim 2, wherein the at least one processor is configured to generate an interaction response based at least in part upon the received context information.

13. The method of claim 12, wherein the received context information used to generate the conversation response comprises data from a database associated with the user, an application context, an input previously provided by the user, known information about the user, Condition, and history of the user input.

3. The system of claim 2, wherein the received context information includes at least one selected from the group consisting of context information stored in a server and context information stored in a client.

3. The system of claim 2, wherein the at least one processor is configured to receive context information from a context source by requesting the context information from a context source and receiving the context information in response to the request.

3. The system of claim 2, wherein the at least one processor is configured to receive context information from a context source by receiving at least a portion of the context information prior to receiving the voice user input.

3. The system of claim 2, wherein the at least one processor is configured to receive context information from a context source by receiving at least a portion of the context information after receiving the voice user input.

3. The method of claim 2, wherein the at least one processor receives static context information as part of an initialization step and receives context information from a context source by receiving additional context information after receiving the speech user input A user input interpretation system.

3. The method of claim 2, wherein the at least one processor is configured to receive context information from a context source by receiving a push notification of the change in the context information and updating locally stored context information in response to the push notification Configured user input interpretation system.

3. The method of claim 2, wherein the output device, the input device, and the at least one processor are selected from the group consisting of a phone, a smartphone, a tablet computer, a laptop computer, a personal digital assistant, a desktop computer, a kiosk, Wherein the at least one component is implemented as at least one component selected from the group consisting of a music player, a camera, a television, an electronic game unit, and a set-top box.

3. The system of claim 2, wherein the received context information further comprises an application context.

3. The system of claim 2, wherein the received context information further comprises personal data associated with the user.

3. The system of claim 2, wherein the received context information further comprises data from a database associated with the user.

3. The system of claim 2, wherein the received context information further comprises data obtained from the conversation history.

3. The system of claim 2, wherein the received context information further comprises data received from at least one sensor.

3. The system of claim 2, wherein the received context information further comprises an application preference.

3. The system of claim 2, wherein the received context information further comprises an application usage history.

3. The system of claim 2, wherein the received context information further comprises data describing an event.

3. The system of claim 2, wherein the received context information further comprises a current conversation state.

3. The system of claim 2, wherein the received context information further comprises an input previously provided by the user.

3. The system of claim 2, wherein the received context information further comprises a location.

3. The system of claim 2, wherein the received context information further comprises a local time.

3. The system of claim 2, wherein the received context information further comprises an environmental condition.