EP4728351A1 - User interfaces and techniques for changing how an object is displayed - Google Patents

User interfaces and techniques for changing how an object is displayed

Info

Publication number
EP4728351A1
EP4728351A1 EP24790048.3A EP24790048A EP4728351A1 EP 4728351 A1 EP4728351 A1 EP 4728351A1 EP 24790048 A EP24790048 A EP 24790048A EP 4728351 A1 EP4728351 A1 EP 4728351A1
Authority
EP
European Patent Office
Prior art keywords
content
user
computer system
displaying
detecting
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24790048.3A
Other languages
German (de)
French (fr)
Inventor
Agatha Y. YU
Ji Chen Jason YUAN
Tuhin Kumar
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Apple Inc
Original Assignee
Apple Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Apple Inc filed Critical Apple Inc
Publication of EP4728351A1 publication Critical patent/EP4728351A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/012Head tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

The present disclosure generally relates to user interfaces.

Description

USER INTERFACES AND TECHNIQUES FOR CHANGING HOW AN OBJECT IS
DISPLAYED
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to U.S. Provisional Patent Application Serial No. 63/541,811, filed September 30, 2023, and to U.S. Provisional Patent Application Serial No. 63/541,832, filed September 30, 2023, which are hereby incorporated by reference in their entireties for all purposes.
BACKGROUND
[0002] Users often use computer systems to display objects. Such objects include videos, animations, and pictures.
SUMMARY
[0003] Existing techniques for changing display of an object using electronic devices are generally cumbersome and inefficient. In some embodiments, some existing techniques use a complex and time-consuming user interface, which may include multiple key presses or keystrokes. Some existing techniques require more time than necessary, wasting user time and device energy. This latter consideration is particularly important in battery-operated devices.
[0004] Accordingly, the present technique provides electronic devices with faster, more efficient methods and interfaces for changing display of an object is directed, and/or for displaying an overlay. Such methods and interfaces optionally complement or replace other methods for changing display of an object is directed, and/or for displaying an overlay. Such methods and interfaces reduce the cognitive burden on a user and produce a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power and increase the time between battery charges. Such methods and interfaces may complement or replace other methods for changing display of an object.
[0005] In some embodiments, a method that is performed at a computer system that is in communication with a display component and the one or more input devices is described. In some embodiments, the method comprises: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
[0006] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of- detection of the one or more input devices.
[0007] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
[0008] In some embodiments, a computer system that is in communication with a display component and the one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
[0009] In some embodiments, a computer system that is in communication with a display component and the one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
[0010] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices. In some embodiments, the one or more programs include instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
[0011] In some embodiments, a method that is performed at a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the method comprises: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0012] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0013] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0014] In some embodiments, a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0015] In some embodiments, a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0016] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
[0017] In some embodiments, a method that is performed at a computer system that is in communication with a display component, a camera, and one or more input devices is described. In some embodiments, the method comprises: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0018] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0019] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0020] In some embodiments, a computer system that is in communication with a display component, a camera, and one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: while detecting a first entity in the field- of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field- of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0021] In some embodiments, a computer system that is in communication with a display component, a camera, and one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0022] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices. In some embodiments, the one or more programs include instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of- detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
[0023] In some embodiments, a method that is performed at a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the method comprises: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
[0024] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar. [0025] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
[0026] In some embodiments, a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar. [0027] In some embodiments, a computer system that is in communication with a display component and one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that the third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
[0028] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
[0029] In some embodiments, a method that is performed at a computer system that is in communication with a display component, an audio generation component, and a movement component is described. In some embodiments, the method comprises: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0030] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component is described. In some embodiments, the one or more programs includes instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0031] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component is described. In some embodiments, the one or more programs includes instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0032] In some embodiments, a computer system that is in communication with a display component, an audio generation component, and a movement component is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0033] In some embodiments, a computer system that is in communication with a display component, an audio generation component, and a movement component is described. In some embodiments, the computer system comprises means for performing each of the following steps: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0034] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component. In some embodiments, the one or more programs include instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
[0035] In some embodiments, a method that is performed at a computer system that is in communication with a one or more output devices and one or more input devices is described. In some embodiments, the method comprises: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0036] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0037] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0038] In some embodiments, a computer system that is in communication with a one or more output devices and one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0039] In some embodiments, a computer system that is in communication with a one or more output devices and one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0040] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
[0041] In some embodiments, a method that is performed at a computer system that is in communication with one or more output devices and one or more input devices is described. In some embodiments, the method comprises: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0042] In some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0043] In some embodiments, a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices is described. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that the attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0044] In some embodiments, a computer system that is in communication with one or more output devices and one or more input devices is described. In some embodiments, the computer system comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs includes instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0045] In some embodiments, a computer system that is in communication with one or more output devices and one or more input devices is described. In some embodiments, the computer system comprises means for performing each of the following steps: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0046] In some embodiments, a computer program product is described. In some embodiments, the computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
[0047] Executable instructions for performing these functions are, optionally, included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are, optionally, included in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
DESCRIPTION OF THE FIGURES
[0048] For a better understanding of the various described embodiments, reference should be made to the Detailed Description below, in conjunction with the following drawings in which like reference numerals refer to corresponding parts throughout the figures.
[0049] FIG. l is a block diagram illustrating a computer system in accordance with some embodiments.
[0050] FIGS. 2A-2C are diagrams illustrating exemplary components and user interfaces of device 200 in accordance with some embodiments.
[0051] FIG. 3 is a block diagram illustrating exemplary components of a device in accordance with some embodiments.
[0052] FIG. 4 is a functional diagram of an exemplary actuator device in accordance with some embodiments.
[0053] FIG. 5 is a functional diagram of an exemplary agent system in accordance with some embodiments.
[0054] FIGS. 6A-6G illustrate exemplary user interfaces for changing display of an object in accordance with some embodiments.
[0055] FIG. 7 is a flow diagram illustrating methods for displaying an object facing a direction in accordance with some embodiments.
[0056] FIG. 8 is a flow diagram illustrating methods for displaying an indication of eye contact of an object in accordance with some embodiments.
[0057] FIG. 9 is a flow diagram illustrating methods for de-emphasizing an object in accordance with some embodiments. [0058] FIGS. 10A-10F illustrate exemplary user interfaces for providing content in accordance with some embodiments.
[0059] FIG. 11 is a flow diagram illustrating methods for displaying a system avatar in accordance with some embodiments.
[0060] FIG. 12 is a flow diagram illustrating methods for selectively moving a portion of a computer system in accordance with some embodiments.
[0061] FIG. 13 is a flow diagram illustrating methods for navigating content in accordance with some embodiments.
[0062] FIG. 14 is a flow diagram illustrating methods for pausing content in accordance with some embodiments.
DETAILED DESCRIPTION
[0063] The description to follow sets forth exemplary methods, components, parameters, and the like. While specific examples are set out below, it should be recognized that such embodiments should not be understood as limiting the scope of the present disclosure to the explicit descriptions of the examples set forth herein but instead should be understood as providing illustrative examples.
[0064] Each of the identified modules and applications herein corresponds to a set of executable instructions for performing one or more functions described above and the methods described in this application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) optionally need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), procedures, or modules, and thus various subsets of these modules are, optionally, combined or otherwise rearranged in various embodiments. In some embodiments, a video player module is, optionally, combined with a music player module into a single module. In some embodiments, memory optionally stores a subset of the modules and data structures identified above. Furthermore, memory optionally stores additional modules and data structures not described above.
[0065] One or more steps of the methods described herein can rely on (be contingent on) one or more conditions being satisfied. In some embodiments, a method is performed by iterating a process multiple times. In some embodiments, contingent steps can be satisfied on different iterations of the same process and still be within the scope of the methods described herein. In some embodiments, for a given method that includes two steps that are contingent on different conditions, one of ordinary skill in the art would understand that the given method is considered performed even when a process is repeated multiple times until the contingent steps are satisfied. In some embodiments, multiple iterations of a process are not required to in order to practice claims as presented herein. In some embodiments, electronic device, system, or computer readable medium claims can be performed without iteratively repeating a process. In some embodiments, the electronic device, system, or computer readable medium claims include instructions for performing one or more steps that are contingent upon one or more conditions being satisfied. Because such instructions are stored in one or more processors and/or at one or more memory locations, the electronic device, system, or computer readable medium claims can include logic that determines whether the one or more conditions have been satisfied without needing to repeat steps of a process.
[0066] Although elements are described below using numerical descriptors, such as “a first” and/or “a second,” these elements do not correspond to order or distinct representations and should not be limited to the stated numerical term. In some embodiments, these terms simply used as prefix to distinguish a reference to one element from a reference to another element. In some embodiments, a “first” device and a “second” device can be two separate references to the same device. In contrast, in some embodiments, a “first” device and a “second” device can be a reference to two different devices (e.g., not the same device and/or not the same type of device). In some embodiments, a first computer system and a second computer system do not correspond to a first and a second in time, and merely are used to distinguish between two computer systems. As such, the first computer system can be termed a second computer system, and the second computer system can be termed a first computer system without departing from the scope of the various described embodiments.
[0067] For description of various elements and examples, the use of certain terminology is used to provide productive descriptions of the subject matter below and should not be read as limiting. As used to describe various examples herein, the singular forms of “a,” “an,” and “the” should not be interpreted as precluding or excluding the plural forms as well, unless the context clearly indicates otherwise. As well, “and/or” is used to encompasses any and all possible combinations of one or more associated listed items. In some embodiments, “x and/or y” should be interpreted as including “x,” or “y,” as well as “x and y” as possible permutations. Further, the use of the terms “includes,” “including,” “comprises,” and/or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and/or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and/or groups thereof.
[0068] When describing choices and/or logical possibilities, the term “if’ is, optionally, construed to mean “when,” “upon,” “in response to determining,” “in response to detecting,” or “in accordance with a determination that” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” is, optionally, construed to mean “upon determining,” “in response to determining,” “upon detecting [the stated condition or event],” “in response to detecting [the stated condition or event],” or “in accordance with a determination that [the stated condition or event]” depending on the context.
[0069] The processes described below enhance the operability of the devices and make the user-device interfaces more efficient (e.g., by helping the user and/or user to provide proper inputs and reducing user mistakes when operating/interacting with the device) through various techniques, including by providing improved feedback (e.g., visual, haptic, audible, and/or tactile feedback) to the user, reducing the number of inputs needed to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, performing an operation when a set of conditions has been met without requiring further input (e.g., input by a user), and/or additional techniques, such as increasing the security and/or privacy of the computer system and reducing burn-in of one or more portions of a user interface of a display. These techniques also reduce power usage and improve battery life of the device by enabling the user to use the device more quickly and efficiently.
[0070] Below, FIGS. 1, 2A-2C, and 3-5 provide a description of exemplary devices for performing the techniques described herein. FIGS. 6A-6G illustrate exemplary user interfaces for changing display of an object in accordance with some embodiments. FIG. 7 is a flow diagram illustrating methods for displaying an object facing a direction in accordance with some embodiments. FIG. 8 is a flow diagram illustrating methods for displaying an indication of eye contact of an object in accordance with some embodiments. FIG. 9 is a flow diagram illustrating methods for de-emphasizing an object in accordance with some embodiments. The user interfaces in FIGS. 6A-6G are used to illustrate the processes described below, including the processes in FIGS. 7, 8, and 9. FIGS. 10A-10F illustrate exemplary user interfaces for providing content in accordance with some embodiments. FIG. 11 is a flow diagram illustrating methods for displaying a system avatar in accordance with some embodiments. FIG. 12 is a flow diagram illustrating methods for selectively moving a portion of a computer system in accordance with some embodiments. FIG. 13 is a flow diagram illustrating methods for navigating content in accordance with some embodiments. FIG. 14 is a flow diagram illustrating methods for pausing content in accordance with some embodiments. The user interfaces in FIGS. 10A-10F are used to illustrate the processes described below, including the processes in FIGS. 11, 12, 13, and 14.
[0071] FIG. 1 depicts a block diagram of computer system 100 (e.g., electronic device and/or electronic system) including a set of electronic components in communication with (e.g., connected to) (e.g., wired or wirelessly) to each other. It should be understood that computer system 100 is merely one example of a computer system that can be used to perform functionality described below and that one or more other computer systems can be used to perform the functionality described below. Additionally, while FIG. 1 depicts a computer architecture of computer system 100, other computer architectures (e.g., including more components, similar components, and/or fewer components) of a computer system can be used to perform functionality described herein.
[0072] In some embodiments, computer system 100 can correspond to (e.g., be and/or include) a system on a chip, a server system, a personal computer system, a smart phone, a smart watch, a wearable device, a tablet, a laptop computer, a fitness tracking device, a headmounted display (HMD) device, a desktop computer, a communal device (e.g., smart speaker, connected thermostat, and/or additional home based computer systems), an accessory (e.g., switch, light, speaker, air conditioner, heater, window cover, fan, lock, media playback device, television, and so forth), a controller, a hub, and/or a sensor.
[0073] In some embodiments, a sensor includes one or more hardware components capable of detecting (e.g., sensing, generating, and/or processing) information about a physical environment in proximity to the sensor. In some embodiments, a sensor can be configured to detect information surrounding the sensor, detect information in one or more directions casting away from the sensor, and/or detect information based on contact of the sensor with an element of the physical environment. In some embodiments, a hardware component of a sensor includes a sensing component (e.g., a temperature and/or image sensor), a transmitting component (e.g., a radio and/or laser transmitter), and/or a receiving component (e.g., a laser and/or radio receiver). In some embodiments, a sensor includes an angle sensor, a breakage sensor,, a flow sensor, a force sensor, a gas sensor, a humidity or moisture sensor, a glass breakage sensor, a chemical sensor, a contact sensor, a non-contact sensor, an image sensor (e.g., a RGB camera and/or an infrared sensor), a particle sensor, a photoelectric sensor (e.g., ambient light and/or solar), a position sensor (e.g., a global positioning system), a precipitation sensor, a pressure sensor, a proximity sensor, a radiation sensor, an inertial measurement unit, a leak sensor, a level sensor, a metal sensor, a microphone, a motion sensor, a range or depth sensor (e.g., RADAR, LiDAR), a speed sensor, a temperature sensor, a time-of-flight sensor, a torque sensor, and an ultrasonic sensor, a vacancy sensor, a presence sensor, a voltage and/or current sensor, a conductivity sensor, a resistivity sensor, a capacitive sensor, and/or a water sensor. While only a single computer system is depicted in FIG. 1, functionality described below can be implemented with two or more computer systems operating together. Additionally, in some embodiments, computer system 100 includes one or more sensors as described above, and information about the physical environment is captured by combining data from one sensor with data from one or more additional sensors (e.g., that are part of the computer and/or one or more additional computer systems).
[0074] As illustrated in FIG. 1, computer system 100 consists of processor subsystem 110, memory 120, and VO interface 130. Memory 120 corresponds to system memory in communication with processor subsystem 110. The electronic components making up computer system 100 are electrically connected through interconnect 150, which allows communication between the components of computer system 100. In some embodiments, interconnect 150 can be a system bus, one or more memory locations, and/or additional electrical channels for connective multiple components of computer system 100. Also, I/O interface 130 is connected to, via a wired and/or wireless connection, I/O device 140. In some embodiments, computer system 100 includes a component made up of I/O interface 130 and I/O device 140 such that the functionality of the individual components is included in the component. Additionally, it should be understood that computer system 100 can include one or more I/O interfaces, communicating with one or more I/O devices. In some embodiments, computer system 100 consists of multiple processor subsystem 100s, each electrically connected through interconnect 150.
[0075] In some embodiments, processor subsystem 110 includes one or more processors or individual processing units capable of executing instructions (e.g., program, system, and/or interrupt) to perform functionality described herein. In some embodiments, operating system level and/or application-level instructions executed by processor subsystem 110. In some embodiments, processor subsystem 110 includes one or more components (e.g., implemented as hardware, software, and/or a combination thereof) capable of supporting, interpreting, and/or performing machine learning instructions and/or operations. In some embodiments, computer system 100 can perform operations according to a machine learning model locally. Alternatively, or in addition, computer system 100 can communicate with (e.g., performing calculations on and/or executing instructions corresponding to) a remote interactive knowledge base (e.g., a processing resource that implements a machine learning model, artificial intelligence model, and/or large language model) to perform operations that can be otherwise outside a set of capabilities of computer system 100. In some embodiments, computer system 100 can determine a set of inputs (e.g., instructions, data, and/or parameters) to the interactive knowledge base for performing desired machine learning operations.
[0076] Memory 120 in communication with processor subsystem 110 can be implemented by a variety of different physical, non-transitory memory media. In some embodiments, computer system 100 includes multiple memory components and/or multiple types of memory components, each connected to processor subsystem 110 directly and/or via interconnect 150. In some embodiments, memory 120 can be implemented using a removable flash drive, storage array, a storage area network (e.g., SAN), flash memory, hard disk storage, optical drive storage, floppy disk storage, removable disk storage, random access memory (e g., SDRAM, DDR SDRAM, RAM-SRAM, EDO RAM, and/or RAMBUS RAM), and/or read only memory (e.g., PROM and/or EEPROM). Additionally, in some embodiments, processor subsystem 110 and/or interconnect 150 is connected to a memory controller that is electrically connected to memory 120.
[0077] In some embodiments, instructions can be executed by processor subsystem 110. In this example, memory 120 can include a computer readable medium (e.g., non-transitory or transitory computer readable medium) usable to store (e.g., configured to store, assigned to store, and/or that stores) instructions to be executable by processor subsystem 110. In some embodiments each instruction stored by memory 120 and executed by processor subsystem 110 corresponds to an operation for completing the functionality described herein. In some embodiments, memory 120 can store program instructions to implement the functionality associated with the methods described below including 700, 800, and 900 (FIGS. 7, 8, and 9).
[0078] As mentioned above, I/O interface 130 can be one or more types of interfaces enabling computer system 100 to communicate with other devices. In some embodiments, I/O interface 130 includes a bridge chip (e.g., Southbridge) from a front-side bus to one or more back-side buses. In some embodiments, I/O interface 130 enables communication with one or more I/O devices, illustrated as I/O device 140, via one or more corresponding buses or other interfaces. In some embodiments, an I/O device can include one or more: a physical user-interface devices (e.g., a physical keyboard, a mouse, and/or a joystick), storage devices (e.g., as described above with respect to memory 120), network interface devices (e.g., to a local or wide-area network), sensor devices (e.g., as described above with respect to sensors), and/or auditory and/or visual output devices (e.g., screen, speaker, light, and/or projector). In some embodiments, the visual output device is referred to as a display component. In some embodiments, the display component can be configured to provide visual output, such as displaying images on a physically viewable medium via an LED display or image projection. As used herein, “displaying” content includes causing to display the content (e.g., video data rendered and/or decoded by a display controller) by transmitting, via a wired or wireless connection, data (e.g., image data and/or video data) to an integrated or external display component to visually produce the content.
[0079] In some embodiments, computer system 100 includes a component that integrates EO device 140 with other components (e.g., a component that includes EO interface 130 and EO device 140). In some embodiments, EO device 140 is separate from other components of computer system 100 (e.g., is a discrete component). In some embodiments, EO device 140 includes a network interface device that permits computer system 100 to connect to (e.g., communicate with) a network or other computer systems, in a wired or wireless manner. In some embodiments, a network interface device can include Wi-Fi, Bluetooth, NFC, USB, Thunderbolt, Ethernet, and so forth. In some embodiments, computer system 100 can utilize an NFC connection to facilitate a bank, credit, financial, token (e.g., fungible or non-fungible token), and/or cryptocurrency transaction between computer system 100 and another computer system within proximity.
[0080] In some embodiments, I/O device 140 includes components for detecting a user (e.g., a user, a person, an animal, another computer system different from the computer system, and/or an object) and/or an input (e.g., a tap input and/or a non-tap input (e.g., a verbal input, an audible request, an audible command, an audible statement, a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) from a detected user. In some embodiments, I/O device 140 enables computer system 100 to identify users associated with and/or without an account within an environment. In some embodiments, computer system 100 can detect a known user (e.g., a user that corresponds to an account) and access information about the user using the known user’s account. In some embodiments, as part of computer system 100 detecting a user, computer system 100 detects that the user’s account is associated with (e.g., is included in and/or identified with respect to) a group of users. In some embodiments, computer system 100 can access information associated with a family of accounts in response to detecting a member of the family that is defined as a group of accounts. In some embodiments, as account corresponding to a user can be connected with additional accounts and/or additional computer systems. In some embodiments, computer system 100 can detect such additional computer systems and/or detect such computer systems for detecting the user. In some embodiments, computer system 100 detects unknown users and enables guest accounts for the unknown users to utilize computer system 100.
[0081] In some embodiments, I/O device 140 includes one or more cameras. In some embodiments, a camera includes an image sensor (e.g., one or more optical sensors and/or one or more depth camera sensors) that provides computer system 100 with the ability to detect a user and/or a user’s gestures (e.g., hand gestures and/or air gestures) as input. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independently of an input element that is a part of the device) and is based on detected motion of a portion of the user’s body through the air including motion of the user’s body relative to an absolute reference (e.g., an angle of the user’s arm relative to the ground or a distance of the user’s hand relative to the ground), relative to another portion of the user’s body (e.g., movement of a hand of the user relative to a shoulder of the user, movement of one hand of the user relative to another hand of the user, and/or movement of a finger of the user relative to another finger or portion of a hand of the user), and/or absolute motion of a portion of the user’s body (e.g., a tap gesture that includes movement of a hand in a predetermined pose by a predetermined amount and/or speed, or a shake gesture that includes a predetermined speed or amount of rotation of a portion of the user’s body). In some embodiments, the one or more cameras enable computer system 100 to transmit pictorial and/or video information to an application. In some embodiments, image data captured by a camera can enable computer system 100 to complete a video phone call by transmitting video data to an application for performing the video phone call.
[0082] In some embodiments, I/O device 140 includes one or more microphones. In some embodiments, a microphone can be used by 100 to obtain data and/or information from a user without a contact input. In some embodiments, a microphone enables computer system 100 to detect verbal and/or speech input from a user. In some embodiments, computer system 100 utilizes speech input to enable personal assistant functionality. In some embodiments, a user eliciting a request to computer system 100 to perform an action and/or obtain information for the user. In some embodiments, computer system 100 utilizes speech input (e.g., along with one or more other input and/or output techniques) to request and/or detect information from a user without requiring the user to make physical contact with computer system 100.
[0083] In some embodiments, I/O device 140 includes physical input mediums for a user to interact directly with computer system 100. In some embodiments, a physical input medium includes one or more physical buttons (e.g., tactile depressible button and/or touch sensitive non-depressible component) on computer system 100 and/or connected to computer system 100, a mouse and keyboard input method (e.g., connected to computer system 100 together and/or separately with one or more I/O interfaces), and/or a touch sensitive display component.
[0084] In some embodiments, I/O device 140 includes one or more components for outputting information (e.g., a display component, an audio generation component, a speaker, a haptic output device, a display screen, a projector, and/or a touch-sensitive display). In some embodiments, computer system 100 uses I/O device 140 to convey information and/or a state of computer system 100. In some embodiments, I/O device 140 includes a tactile output component. In some embodiments, a tactile output component can be a haptic generation component that enables computer system 100 to convey information to a user in contact with (e.g., holding, touching, and/or nearby) computer system 100. In some embodiments, I/O device 140 includes one or more components for outputting visual outputs (e.g., video, image, animation, 3D rendering, augmented reality overlay, motion graphics, data visualization, digital art, etc.). In some embodiments, displaying content from one or more applications and/or system applications, and/or displaying a widget (e.g., a control that displays real-time information and/or data) corresponding to one or more applications.
[0085] In some embodiments, I/O device 140 includes one or more components for outputting audio (e.g., smart speakers, home theater system, soundbars, headphones, earphones, earbuds, speakers, television speakers, augmented reality headset speakers, audio jacks, optical audio output, Bluetooth audio outputs, HDMI audio outputs, audio sensors, etc.). In some embodiments, computer system 100 is able to output audio through the one or more speakers. In some embodiments, computer system 100 outputting audio-based content and/or information to a user. In some embodiments, the one or more speakers enable spatial audio (e.g., an audio output corresponding to an environment (e.g., computer system 100 detecting materials and/or objects within the environment and/or computer system 100 altering the audio pattern, intensity, and/or waveform to compensate for varying characteristics of an environment)).
[0086] FIGS. 2-5 illustrate exemplary components and user interfaces of device 200 in accordance with some embodiments. Device 200 (sometimes referred to herein as device 200) can include one or more features of computer system 100. In the examples described with respect to FIGS. 2-5, device 200 is a laptop computer. In some embodiments, device 200 is not limited to being a laptop computer and one of ordinary skill in the art should recognize that device 200 can be one or more other devices (e.g., as described herein and/or that include one or more of the components and/or functions described herein with respect to device 200). In some embodiments, device 200 can be a communal device (such as a smart display, a smart speaker, and/or a television) and/or a personal device (such as a smart phone, a smart watch, a tablet, a desktop computer, a fitness tracking device, and/or a head mounted display device). In some embodiments, a communal device is configured to provide functionality to multiple users (e.g., at the same time and/or at different times). In such embodiments, the communal device can be administered and/or set up by a single user. In some embodiments, a personal device is configured to provide functionality to a single user (e.g., at a time, such as when the single user is logged into the personal device).
[0087] FIGS. 2A-2C illustrate device 200 in three different physical positions. As illustrated in FIG. 2A, device 200 is a laptop computer (also referred to herein as a “laptop”) that includes base portion 200-2 (e.g., that rests on a surface, such as a desk, horizontally as shown in FIG. 2A) and display portion 200-1 that is connected to base portion 200-2 at connection 200-3 (e.g., one or more connection points, a motorized arm, a hinge, and/or a joint) that enables display portion 200-1 to pivot and/or change orientation with respect to base portion 200-2. In some embodiments, device 200 can pivot at connection 200-3 to rotate display portion 200-1 and/or device 200 to one or more positions corresponding to an “OFF” internal state (e.g., as further described below in relation to FIG. 2C). In some embodiments, a position corresponding to an “OFF” internal state is a position in which device 200 is in a predetermined pose. In some embodiments, a predetermined pose can include display portion 200-1 positioned parallel to base portion 200-2 or display portion 200-1 forming a predetermined angle (e.g., 60-degree angle) with respect to base portion 200-2. In some embodiments, in the “OFF” internal state, an area in which content is displayed by device 200 is positioned in a manner that corresponds to (e.g., represents, is associated with, and/or is configured to accompany) the “OFF” internal state (e.g., facing down, not visible, and/or obscuring the area in which content is displayed). In some embodiments, in the “OFF” internal state, an area in which content is displayed by device 200 is not positioned in a manner that corresponds to (e.g., represents, is associated with, and/or is configured to accompany) the “OFF” internal state (e.g., instead is positioned in a manner that corresponds to an “ON” internal state). In some embodiments, when not in the “OFF” internal state, device 200 can be positioned within a range of different open positions (e.g., in which display portion 200-1 is not parallel to base portion 200-2 and the area in which content is displayed by device 200 is visible and/or not obscured). It should be recognized that display portion 200-1 being parallel to base portion 200-2 is an example of a position corresponding to an “OFF” internal state (e.g., a closed position) of device 200. In some embodiments, another configuration could set another orientation of display portion 200-1 with respect to base portion 200-2 as the closed position of device 200, such as illustrated in FIG. 2C.
[0088] FIG. 2A illustrates display screen 200-4 (representing the area in which content is displayed by device 200) on the left and device 200 in a corresponding pose on the right. As illustrated in FIG. 2A, device 200 is in a first position (e.g., display portion 200-1 is perpendicular to base portion 200-2 forming a 90-degree angle). In FIG. 2A, display screen 200-4 represents what is currently being displayed (e.g., via a display component) by device 200 while open in the first position. In FIG. 2A, display screen 200-4 illustrates an internal state in which device 200 is “ON” (e.g., operational, powered on, awake, a higher powered and/or more resource intensive state than the “OFF” state, and/or activated). In some embodiments, device 200 displays (e.g., via display screen 200-4) one or more user interfaces (e.g., user interface objects, windows, application user interfaces, system user interfaces, controls, and/or other visual content). In some embodiments, device 200 displays (e.g., via display screen 200-4) the one or more user interfaces while in the “ON” internal state. In some embodiments, in FIG. 2A, device 200 is in the “ON” internal state and display screen 200-4 displays a desktop user interface 200-5 that includes an application window. In some embodiments, a user interface includes (and/or is) one or more user interface objects (e.g., windows, icons, and/or other graphical objects). In some embodiments, a user interface (e.g., 200-5) can include one or more graphical objects different than, and/or the same as, an application window.
[0089] FIG. 2B illustrates display screen 200-4 on the left and device 200 in a corresponding pose on the right. As illustrated in FIG. 2B, device 200 is in a second position (e.g., display portion 200-1 is angled (e.g., via connection 200-3) with respect to base portion 200-2 forming at a 120-degree angle (e.g., a larger angle than in FIG. 2 A)). In FIG. 2B, display screen 200-4 represents what is being displayed by device 200 while in the second position. Display screen 200-4 illustrates an internal state in which device 200 is “ON” (e.g., the same internal state as the top diagram of FIG. 2A). In FIG. 2B, device 200 displays (e.g., via display screen 200-4) desktop user interface 200-5 (e.g., and is the same as displayed in FIG. 2A). In some embodiments, device 200 displays a different user interface (e.g., other than desktop user interface 200-5). In some embodiments, although FIG. 2B illustrates device 200 displaying the same desktop user interface 200-5 as in FIGS. 2A while in a different position than in FIG. 2A, device 200 can display a different user interface. In some embodiments, device 200 displays a user interface that corresponds to (e.g., is based on, due to, caused by, related to, and/or configured to accompany) a physical state (e.g., position, location, and/or orientation), including content that is specific to a particular angle or specific to a current context.
[0090] FIG. 2C illustrates display screen 200-4 on the left and device 200 in a corresponding pose on the right. As illustrated in FIG. 2C, device 200 is in a third position (e.g., display portion 200-1 is angled (e.g., via connection 200-3) with respect to base portion 200-2 forming at a 60-degree angle (e.g., a smaller angle than in FIG. 2A and FIG. 2B)). In FIG. 2C, display screen 200-4 represents what is being displayed by device 200 while in the third position. In FIG. 2C, display screen 200-4 illustrates an internal state in which device 200 is “OFF” (e.g., not operational, not powered on, not awake, not activated, powered off, asleep, hibernating, inactive, and/or deactivated). In some embodiments, device 200 does not display (e.g., via display screen 200-4) (e.g., forgoes displaying) the one or more user interfaces while in the “OFF” internal state (e.g., does not display any visual content). In some embodiments, device 200 displays (e.g., via display screen 200-4) one or more user interfaces while in the “OFF” internal state (e.g., the same and/or different from one or more user interfaces displayed while in the “ON” internal state) (e.g., a user interface specific to the “OFF” state and/or a manner of displaying a user interface that is not specific to the “OFF” internal state). In FIG. 2C, display screen 200-4 is blank because nothing is being displayed on the display of device 200 (e.g., display screen 200-4 is off and/or not displaying a user interface) (e.g., desktop user interface 200-5 is not displayed on display screen 200-4).
[0091] In some embodiments, device 200 includes one or more components (also referred to herein as “movement components”) that enable device 200 to perform (e.g., cause and/or control) movement (and/or be moved). In some embodiments, performing movement can include moving a portion of device 200 (e.g., less than or all components of the device move), moving all of device 200 (e.g., the entire device (including all of its components) moves, such as by changing location), and/or moving one or more other devices and/or components (e.g., that are in communication with device 200 and/or movement components of device 200). In some embodiments, device 200 can automatically move (e.g., pivot), cause, and/or control movement of display portion 200-1 relative to base portion 200-2, such as to any of the positions illustrated in FIGS. 2A-2C. In some embodiments, device 200 performs movement based on an internal state of device 200. Performing movement based on an internal state can enable new (e.g., otherwise unavailable) interactions by device 200. In some embodiments, such new interactions of device 200 can be configured using special features, functions, modes, and/or programs that take advantage of the ability of device 200 to perform movement. Examples of such interaction include using movement to communicate (e.g., to a user) an internal state (e.g., on, off, sleeping, and/or hibernating) of the device, to assist with user input (e.g., reduce distance to a user), and/or to augment interaction behavior of the device (e.g., moving in particular ways, during an interaction with a user, that convey information such as importance and/or direction of attention). In some embodiments, the movement performed corresponds to (e.g., is caused by, is in response to, and/or is determined and/or performed based on) one or more of: detected input, detected context (e.g., environmental context and/or user context), and/or an internal state of device 200 (e.g., an internal state and/or a set of multiple internal states). In some embodiments, device 200 can perform a movement of the display portion such that device 200 moves from being in the first position illustrated in FIG. 2A to being in the second position illustrated in FIG. 2B. In this example, device 200 can detect that a user has repositioned with respect to device 200 (e.g., the user stood up), and in response, device 200 can perform the movement to the second position so that the display is at an optimized viewing angle based on the repositioned height and/or angle of the user’s eyes with respect to the display of device 200. As another example, device 200 can perform a movement such that device 200 moves from being in the first position illustrated in FIG. 2A to being in the third position illustrated in FIG. 2C. In this example, device 200 can perform the movement to the third position in response to detecting an internal state with reduced activity (e.g., the “OFF” internal state as described above). In this way, the movement of device 200 to one or more positions can indicate an internal state of device 200.
[0092] FIGS. 2A-2C illustrate device 200 having a display portion that is able to move with one degree of freedom via connection 200-3 (e.g., a hinge) connecting display portion 200-1 to base portion 200-2. In some embodiments, device 200 includes one or more components that have one or more degrees of freedom. In some embodiments, a movement component (e.g., an output component that causes and/or allows movement) (e.g., 200-26C of FIG. 5) of device 200 can include multiple degrees of freedom (e.g., six degrees of freedom including three components of translation and three components of rotation). In some embodiments, device 200 can be implemented to be able to move the display portion in a telescoping forward or backward motion (e.g., display portion 200-1 moves forward while base portion 200-2 remains stationary in space relative to the base portion (e.g., to reduce and/or extend viewing distance for a user)). As yet another example, device 200 can be implemented to be able to move the display portion to rotate about an axis that is perpendicular to the hinge such that the display portion can turn to position the display to follow a user as they walk around device 200. While the examples shown in FIGS. 2A-2C illustrate a hinge, other movement components can be included in device 200, such as an actuator (e.g., a pneumatic actuator, hydraulic actuator and/or an electric actuator), a movable base, a rotatable component, and/or a rotatable base. In some embodiments, one or more movement components can cause device 200 to move in different ways, such as to rotate (e.g., 0-360 degrees), to move laterally (e.g., right, left, down, up, and/or any combination thereof), and/or to tilt (e.g., 0-360 degrees).
[0093] FIG. 3 illustrates exemplary block diagram of device 200. In some embodiments, device 200 includes some or all of the components described with respect to FIGS. 1 A, IB, 3, and 5B. As illustrated in FIG. 3, device 200 has bus 200-13 that operatively couples VO section 200-12 (also referred to as an I/O subsection and/or an I/O interface) with processors 200-11 and memory 200-10. As illustrated in FIG. 3, I/O section 200-12 is connected to output devices 200-16 (also referred to herein as “output components”). In some embodiments, output devices 200-16 include one or more visual output devices (e.g., a display component, such as a display, a display screen, a projector, and/or a touch-sensitive display), one or more haptic output devices (e.g., a device that causes vibration and/or other tactile output), one or more audio output devices (e.g., a speaker), and/or one or more movement components (e.g., an actuator, a motor, a mechanical linkage, devices that cause and/or allow movement, and/or one or more movement components as described above). As illustrated in FIG. 3, output devices 200-16 include two exemplary movement components (e.g., movement controller 200-17 and actuator 200-18). Actuator 200-18 can be any component that performs physical movement (e.g., of a portion and/or of the entirety) of a device (e.g., device 200 and/or a device coupled to and/or in contact with device 200). Movement controller 200-17 can be any component (e.g., a control device) that controls (e.g., provides control signals to) actuator 200-18. In some embodiments, movement controller 200-17 can provide control signals that cause actuator 200-18 to actuate (e.g., cause physical movement). In some embodiments, movement controller 200-17 includes one or more logic component (e.g., a processor), one or more feedback component (e.g., sensor), and/or one or more control components (e.g., for applying control signals, such as a relay, a switch, and/or a control line). In some embodiments, movement controller 200-17 and actuator 200-18 are embodied in the same device and/or component as each other (e.g., a dedicated onboard movement controller 200-17 that is affixed to actuator 200-18). In some embodiments, movement controller 200-17 and actuator 200-18 are embodied in different devices and/or components from each other (e.g., one or more processors 200-11 can function as the movement controller 200-17 of actuator 200-18). In some embodiments, movement controller 200-17 and/or actuator 200-18 are embodied in a device (or one or more devices) other than device 200 (e.g., device 200 is coupled to (e.g., temporarily and/or removably) another device and can instruct movement controller 200-17 and/or control actuator 200-18 of the other device). Actuator 200-18 can function to cause one or more types of mechanical movement (e.g., linear and/or rotational) in one or more manners (e.g., using electric, magnetic, hydraulic, and/or pneumatic power). Examples of actuator 200-18 can include electromechanical actuators, linear actuators, and/or rotary actuators.
[0094] As illustrated in FIG. 3, VO section 200-12 is connected to input devices 200-14. In some embodiments, input devices 200-14 include one or more visual input devices (e.g., a camera and/or a light sensor), one or more physical input devices (e.g., a button, a slider, a switch, a touch-sensitive surface, and/or a rotatable input mechanism), one or more audio input devices (e.g., a microphone), and/or other input devices (e.g., accelerometer, a pressure sensor (e.g., contact intensity sensor), a ranging sensor, a temperature sensor, a GPS sensor, an accelerometer, a directional sensor (e.g., compass), a gyroscope, a motion sensor, and/or a biometric sensor). In addition, VO section 200-12 can be connected with communication unit 200-15 for receiving application and operating system data, using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and/or other wireless (and/or wired) communication techniques.
[0095] Memory 200-10 of device 200 can include one or more non-transitory computer- readable storage mediums, for storing computer-executable instructions, which, when executed by one or more computer processors 200-11, in some embodiments, cause the computer processors to perform the techniques described below, including processes 700, 800, 900, 1100, 1200, 1300, and 1400 (FIGS. 7, 8, 9, 11, 12, 13, and 14). A computer- readable storage medium can be any medium that can tangibly contain or store computerexecutable instructions for use by or in connection with the instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transitory computer- readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can include, but is not limited to, magnetic, optical, and/or semiconductor storages. Examples of such storage include magnetic disks, optical discs based on CD, DVD, and Blu-ray technologies, as well as persistent solid-state memory such as flash and solid-state drives. Device 200 is not limited to the components and configuration of FIG. 3, but can include other and/or additional components in a multitude of possible configurations, all of which are intended to be within the scope of this disclosure. [0096] FIG. 4 illustrates a functional diagram of actuator 200- 18B in accordance with some embodiments. As described above, actuator 200-18B can be any component that performs physical movement. In some embodiments, actuator 200- 18B operates using input that includes control signal 200-18A and/or energy source 200-18B. In some embodiments, actuator 200-18 can be a rotary actuator that converts electric energy into rotational movement. This rotational movement can cause the movement of the display portion of device 200 described above with respect to FIGS. 2A-2C (e.g., a counterclockwise rotational movement of the actuator causes device 200 to move to a position having a larger angle (e.g., the second position illustrated in FIG. 2B) and a clockwise (e.g., opposite) rotational movement of the actuator causes device 200 to move to a position having a smaller angle (e.g., the third position illustrated in FIG. 2C)). Control signal 200-18A can indicate one or more start and/or stop instructions, a movement and/or actuation direction, a movement and/or actuation speed, an amount of time to move and/or actuate, a goal position (e.g., pose and/or location) for movement and/or actuation, and/or one or more other characteristics of movement and/or actuation. In some embodiments, the control signal and the energy source are the same signal and/or input. In some embodiments, one or more additional components (e.g., mechanical and/or electric) are coupled (e.g., removably or permanently) to actuator 200-18B for affecting movement and/or actuation (e.g., mechanical linkage such as a lead screw, gears, and/or other component for changing (e.g., converting) a characteristic of movement and/or actuation). In some embodiments, actuator 200-18B includes one or more feedback components (e.g., position sensor, encoder, overcurrent sensor, and/or force sensor) that form part of a feedback loop for modifying and/or ceasing movement and/or actuation (e.g., slowing actuation as a goal position is reached and/or ceasing actuation if physical resistance to actuation is detected via a sensor). In some embodiments, the one or more feedback components are included (e.g., partially and/or wholly) in a movement controller (e.g., movement controller 200-13) operatively coupled to the actuator.
[0097] Attention is now turned to functionality (e.g., features and/or capabilities) of one or more devices (e.g., computer system 100 and/or device 200). One such functionality is implementing an “agent,” which can alternatively be referred to as a software agent, an intelligent agent, an interactive agent, a virtual assistant, an intelligent virtual assistant, an interactive virtual assistant, a personal assistant, an intelligent personal assistant, an interactive personal assistant, an intelligent interactive personal assistant, and/or an artificial intelligence (Al) assistant. In some embodiments, an agent refers to a set of one or more functions implemented in hardware and/or software (e.g., locally and/or remotely) on an agent system (e.g., a single device and/or multiple devices). In some embodiments, an agent performs operations to perceive an environment, acquire knowledge, retrieve knowledge, learn skills, interact with users, and/or perform tasks. The agent can, in some embodiments, perform these (and/or other) operations in response to user input and/or automatically (e.g., at an appropriate time determined based on a perceived context). A non-exhaustive list of exemplary operations that an agent can be used for and/or with includes: tracking a user’s eyes, face, and/or body (e.g., to move with the user and/or identify an intent and/or activity of the user); detecting, recognizing, and/or classifying a user in the environment; detecting and/or responding to input (e.g., verbal input, air gestures, and/or physical input, such as touch input and/or force inputs to physical hardware components (e.g., button, knobs, and/or sliders)); detecting context (e.g., user context, operating context, and/or environmental context); moving (e.g., changing pose, position, orientation, and/or location); performing one or more operations in response to input, context, and/or stimulus (e.g., an object or event (e.g., external and/or internal to a device) that causes one or more responsive operations by a device); providing intelligent interaction capabilities (e.g., due to in part to one or more machine learning (“ML”) models such as a large language model (“LLM”)) for responding and/or causing operations to be performed; and/or performing tasks (e.g., a set of operations for achieving a particular goal) (e.g., automatically and/or intelligently). In some embodiments, an agent performs operations in response to non-contact inputs (e.g., air gestures and/or natural language commands). The preceding list is meant to be illustrative of operations that can be performed using an agent but is not meant to be an exhaustive list. Other operations fall within the intended scope of the capabilities of an agent. Additionally, for the purposes of this disclosure, an agent does not need to include all of the functionality mentioned herein but can include less functionality or more functionality (e.g., an agent can be implemented on an agent system that does not have movement functionality but that otherwise includes an intelligent personal assistant that can interact with a user).
[0098] In some embodiments, a user is (e.g., represents, includes, and/or is included in) one or more of a user, person, object, and/or animal in an environment (e.g., a physical and/or virtual environment) (e.g., of the device). In some embodiments, a user is (e.g., represents, includes, and/or is included in) an entity that is perceived (e.g., detected by the device, one or more other devices, and/or one or more components thereof). In some embodiments, an entity is something that is distinguished from surrounding entities (e.g., pieces of environments and/or other users) and/or that is considered as a discrete logical construct via one or more components (e.g., perception components and/or other components). In some embodiments, a user is physical and/or virtual. In some embodiments, a physical user can represent a user standing in front of, and being perceived by, the device. As another example, a virtual user can represent an avatar in a virtual scene perceived by the device (e.g., the avatar is detected in a media stream received by the device and/or captured by a camera of the device). Although presented above as examples of a “user,” the terms and/or concepts referred to as “user,” “person,” “object,” and/or “animal” can be interchanged with “user” throughout this disclosure, unless explicitly indicated otherwise.
[0099] As an example, and referring back to FIGS. 2A-2C, an agent implemented at least partially on device 200 can perform operations that cause display portion 200-1 of device 200 to move with respect to base portion 200-2. In some embodiments, the agent detects (e.g., perceives and determines the occurrence of) a context that includes the user standing up (e.g., based on facial detection and tracking); and, in response, the agent causes device 200 to open and/or device 200 opens display portion 200-1 to the larger angle. As another example, the agent can detect verbal input that corresponds to (e.g., is interpreted as and/or that refers to an operation that includes) a request to move the display (e.g., “Please move my display,” or “Please enter sleep mode.”); and, in response, the agent causes device 200 to move and/or device 200 moves display portion 200-1.
[0100] FIG. 5 illustrates a functional diagram of an exemplary agent system 200-20A. As illustrated in FIG. 5, agent system 200-20A has a dotted box boundary that encloses input components 200-22, agent components 200-24, and output components 200-26. In some embodiments, agent system 200-20A includes fewer, more, and/or different components than illustrated in FIG. 5. In some embodiments, agent system 200-20 is implemented on a single device (e.g., computer system 100 and/or device 200). In some embodiments, agent system 200-20 is implemented on multiple devices. In some embodiments, one or more components of agent system 200-20 illustrated in and/or described with respect to FIG. 5 are external to but operatively coupled to agent system 200-20 (e.g., an accessory, an external device, an external sensor, an external actuator, an external display component, an external speaker, and/or an external database). In some embodiments, one or more components of agent system 200-20 are local to one or more other components of agent system 200-20. In some embodiments, one or more components of agent system 200-20 are remote from one or more other components of agent system 200-20.
[0101] In some embodiments, input components 200-22 includes components for performing sensing and/or communications functions of agent system 200-20. As illustrated in FIG. 5, input components 200-22 includes one or more sensors 200-22A. One or more sensors 200-22A can include any component that functions to detect data corresponding to a physical environment. Examples of one or more sensors 200-22A can include: a camera, a light sensor, a microphone, an accelerometer, a position sensor, a pressure sensor, a temperature sensor, olfactory sensor, and/or a contact sensor. This list is not intended to be exhaustive, and one or more sensors 200-22A can include other sensors not explicitly identified herein that detect, generate, and/or otherwise provide data that can be used (e.g., processed, stored, and/or transformed) for detecting data corresponding to a physical environment. As illustrated in FIG. 5, input components 200-22 includes one or more communications components 200-22B. One or more communications components 200-22B can include any component that functions to send and/or receive communications (e.g., an antenna, a modem, a network interface component, an encoder, a decoder, and/or a communication protocol stack) internal and/or external to agent system 200-20. Communications components 200-22B can be between different devices and/or between components of the same device. The communications can include control signals and/or data (e.g., messages, instructions, files, application data, and/or media streams). In some embodiments, input components 200-22 includes fewer, more, and/or different components than those illustrated in FIG. 5. In some embodiments, input components 200-22 is implemented in hardware and/or software.
[0102] In some embodiments, agent components 200-24 includes components that manage and/or carry out functions of an agent of agent system 200-20. As illustrated in FIG. 5, agent components 200-24 includes the following functional components: task flow, coordination, and/or orchestration component 200-24A, administration component 200-24B, perception component 200-24C, evaluation component 200-24D, interaction component 200- 24E, policy and decision component 200-24F, knowledge component 200-24G, learning component 200-24H, models component 200-241, and APIs component 200-24J. Each of these components is described briefly below. Notably, this list of agent components 200-24 is not intended to be exhaustive, and agent components 200-24 can include other functional components not explicitly identified herein that can be used (e.g., processed, stored, and/or transformed) for performing any function of an agent, such as those described herein. In some embodiments, agent components 200-24 includes fewer, more, and/or different components than those illustrated in FIG. 5. In some embodiments, agent components 200-24 is implemented in hardware and/or software.
[0103] In some embodiments, task flow, coordination, and/or orchestration component 200-24A performs operations that enable an agent to handle coordination between various components. In some embodiments, operations can include handling a data processing task flow to move from perception component 200-24C (e.g., that detects speech input) to models component 200-241 (e.g., for processing the detected speech input using a large language model to determine content and/or intent of the speech input). In some embodiments, task flow, coordination, and/or orchestration component 200-24A performs operations that enable an agent to handle coordination between one or more external components (e.g., resources). In some embodiments, FIG. 5 illustrates examples of external components, such as external database 200-30. In some embodiments, administration component 200-24B includes functionality performed by an operating system of a device implementing agent system 200- 20. In some embodiments, administration component 200-24B includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0104] In some embodiments, administration component 200-24B performs operations that enable an agent system to handle administrative tasks like managing system and/or component updates, managing user accounts, managing system settings, and/or managing component settings. In some embodiments, administration component 200-24B includes functionality performed by an operating system of a device implementing agent system 200- 20. In some embodiments, administration component 200-24B includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0105] In some embodiments, perception component 200-24C performs operations that enable an agent to perceive environmental input. In some embodiments, operations can include detecting that a context and/or environmental condition has occurred, detecting the presence of a user (e.g., user, person, object, and/or animal in an environment), detecting an input that includes speech, detecting an input that includes an air gesture, detecting facial expressions, detecting characteristics (e.g., visible and/or non-visible) of a user, and/or detecting verbal and/or physical cues. In some embodiments, perception component 200-24C includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, perception component 200-24C includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0106] In some embodiments, evaluation component 200-24D performs operations that enable an agent to process evaluate data (e.g., to determine a context such as a user context, an environmental context, and/or an operating context). In some embodiments, operations can include evaluating data gathered from perception component 200-24C, knowledge component 200-24G, external database 200-30, and/or remote processing resource 200-32. In some embodiments, evaluation component 200-24D includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, evaluation component 200-24D includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0107] Reference is made herein to environmental context (also referred to herein as a “context of an environment” and/or “a context corresponding to an environment”). In some embodiments, an environmental context is a context based on one or more characteristics of the environment (e.g., users, locations, time, weather, and/or lighting). In some embodiments, an environmental context can include that it is raining outside, that it is daytime, and/or that a device is currently located in a park. In some embodiments, a device (e.g., using an agent) determines an environmental context (e.g., to be currently true, occurring, and/or applicable) using one or more of detecting input (e.g., via one or more input components) and/or receiving data (e.g., from one or more other devices and/or components in communication with the device).
[0108] Reference is made herein to user context (also referred to herein as a “context of a user” and/or “a context corresponding to a user”) (and/or a user context). In some embodiments, a user context is a context based on one or more characteristics of the user. In some embodiments, a user context can include the user’s appearance and/or clothing, personality, actions, behavior, movement, location, and/or pose. In some embodiments, a device (e.g., using an agent) determines a user context (e.g., to be currently true, occurring, and/or applicable) using one or more of detecting input (e.g., via one or more input components) and/or receiving data (e.g., from one or more other devices and/or components in communication with the device). In some embodiments, a device determines user context based on historical context and/or learned characteristics of the user, where one or more characteristics of the user are learned and/or stored over a period of time by the device.
[0109] Reference is made herein to operational context (also referred to herein as a “context of operation” and/or an “operating context”). In some embodiments, an operational context is a context based on one or more characteristics of the operation of a device (e.g., the device determining and/or accessing the operational context and/or one or more other devices). In some embodiments, an operational context can include the internal state of the device (and/or of one or more components of the device), an internal dialogue of the device (e.g., the device’s understanding of a context), operations being performed by the device, applications and/processes that are executing (e.g., running and/or open) on the device. In some embodiments, a device (e.g., using an agent) determines an operational context (e.g., to be currently true, occurring, and/or applicable) using one or more of detecting input (e.g., via one or more input components) and/or receiving data (e.g., from one or more other devices and/or components in communication with the device). In some embodiments, a device (e.g., using an agent) determines an operational context (e.g., to be currently true, occurring, and/or applicable) using one or more internal states (e.g., accessed, retrieved, and/or queried by a process of the device).
[0110] In some embodiments, interaction component 200-24E performs operations that enable an agent to manage and/or perform interactions with users. In some embodiments, operations can include determining an appropriate interaction model for a particular context and/or in response to a particular input. In some embodiments, interaction component 200- 24E includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, interaction component 200-24E includes functionality performed by one or more applications of a device implementing agent system 200-20.
[OHl] In some embodiments, policy and decision component 200-24F performs operations that enable an agent to take actions in view of available data. In some embodiments, operations can include determining which operations to perform and/or which functional components to utilize in response to a detected context. In some embodiments, policy and decision component 200-24F includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, policy and decision component 200-24F includes functionality performed by one or more applications of a device implementing agent system 200-20. [0112] In some embodiments, knowledge component 200-24G performs operations that enable an agent to access and use stored knowledge. In some embodiments, operations can include indexing, storing, and/or retrieving data from a data store, a database, and/or other resource. In some embodiments, knowledge component 200-24G includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, knowledge component 200-24G includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0113] In some embodiments, learning component 200-24H performs operations that enable an agent to learn through experiences. In some embodiments, operations can include observing and/or keeping track of data that includes preferences, routines, user characteristics, and/or environmental characteristics in a manner in which such data can be used to inform future operation by the agent and/or a component thereof (e.g., such as when performing tasks and/or interactions with users). In some embodiments, learning component 200-24H includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, learning component 200-24H includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0114] In some embodiments, models component 200-241 performs operations that enable an agent to apply ML models (e.g., such as a large language model (LLM)) to process data. In some embodiments, operations can include storing ML models, executing ML models, training and/or re-training ML models, and/or otherwise managing aspects of implementing ML models. In some embodiments, models component 200-241 includes functionality performed by an operating system of a device implementing agent system 200- 20. In some embodiments, models component 200-241 includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0115] In some embodiments, agent system 200-20 responds to natural language input. For example, agent system 200-20 responds to a natural language input that is in the form of a statement, a question, a command, and/or a request. In some embodiments, agent system 200-20 outputs text and/or speech output that is provided in a natural language or mimicking a natural language style. For example, agent system 200-20 can process the natural language question “How hot is it outside?” with a speech response that indicates the current temperature outside at the user’s location (e.g., “It is 18 degrees outside.”). In some embodiments, agent system 200-20 responds to natural language input by providing information (e.g., weather, travel, and/or calendar information) and/or performing a task (e.g., opening a document, searching a database, and/or opening an application).
[0116] In some embodiments, agent system 200-20 includes and/or relies on one or more data models to process input (e.g., natural language input, gesture input, visual input, and/or other data input) and/or provide output (e.g., output of information via natural language output, visual output, audio output, and/or textual output). Such data models can include and/or be trained using user data (e.g., based on particular interactions and/or data from the user being interacted with) and/or global data (e.g., general data based on interactions and/or data from many users). For example, user data (e.g., preferences, previous use of language and/or phrases, calendar entries, a contact list, and/or activity data) can be used to better infer user intent and/or provide responses that are more likely to address a user’s request. In some embodiments, data models used by agent system 200-20 include, are used by, and/or are implemented using one or more machine learning components (e.g., hardware and/or software) (e.g., one or more neural networks). Such machine learning components can be used to process verbal input to determine words and/or phrases therein, one or more contexts that correspond to the words, a user intent corresponding to the words, one or more confidence scores, and/or a set of one or more actions to take in response to the verbal input. Analogous operations can be performed to process other types of inputs, such as visual input, data input, and/or textual input. Such data models can include machine learning and/or data processing models, including, but not limited to, natural language processing models, language models, speech recognition models, object recognition models, visual processing models, ontologies, task flow models, and/or intent recognition models (e.g., used to determine user intent).
[0117] In some embodiments, Application Programming Interfaces (APIs) component 200-24J performs operations that enable an agent to interface with services, devices, and/or components. In some embodiments, operations can include relaying data (e.g., requests, responses, and/or other messages) between data interfaces (e.g., between software programs, between a system process and application process, between system processes, between application processes, between communication protocols, between a client and a server, between file systems, and/or between components on different sides of a trust boundary). In some embodiments, the data interfaces served by APIs component 200-24 J are local (e.g., to the device, such as two application processes exchanging data) and/or remote (e.g., from the device, such as interfacing with a web service via a remote server). In some embodiments, APIs component 200-24J includes functionality performed by an operating system of a device implementing agent system 200-20. In some embodiments, APIs component 200-24J includes functionality performed by one or more applications of a device implementing agent system 200-20.
[0118] In some embodiments, output components 200-26 includes components for performing output functions of agent system 200-20. The exemplary output components illustrated in FIG. 5 are described briefly below. In some embodiments, output components 200-26 include fewer components, more, and/or different components than those illustrated in FIG. 5. In some embodiments, input components are implemented in hardware and/or software.
[0119] As illustrated in FIG. 5, output components 200-26 includes one or more visual output components 200-26 A. One or more visual output components 200-26 A can include any component that functions to output (e.g., generate, create, and/or display), and/or cause output of, a visual output (e.g., an output that is visually perceptible, such as graphical user interface, playback of visual media content, and/or lighting). Examples of one or more visual output components 200-26A can include: a display component, a projector, a head mounted display (HMD), a light-emitting diode (“LED”), and/or a component that creates visually perceptible effects (e.g., movement). This list is not intended to be exhaustive, and one or more visual output components 200-26 A can include other visual output components not explicitly identified herein that detect, generate, and/or otherwise provide data that can be used (e.g., processed, stored, and/or transformed) for outputting visual output.
[0120] As illustrated in FIG. 5, output components 200-26 include one or more audio output components 200-26B. One or more audio output components 200-26B can include any component that functions to output (e.g., generate and/or create), and/or cause output of, an audio output (e.g., an output that is audibly perceptible, such as a sound, music, speech, and/or audio media content). Examples of one or more audio output components 200-26B can include: a speaker, an audio amplifier, a tone generator, and/or a component that creates audibly perceptible effects (e.g., movement such as vibrations). This list is not intended to be exhaustive, and one or more audio output components 200-26B can include other audio output components not explicitly identified herein that detect, generate, and/or otherwise provide data that can be used (e.g., processed, stored, and/or transformed) for outputting audio output.
[0121] As illustrated in FIG. 5, output components 200-26 include one or more movement output components 200-26C (also referred to herein as a “movement component”). One or more movement output components 200-26C can include any component that functions to output (e.g., generate and/or create), and/or cause output of, a movement output (e.g., an output that includes physical movement of the device and/or another device/component). Examples of one or more movement output components 200- 26C can include: a movement controller, an actuator, a mechanical linkage, an electromechanical device, and/or a component that creates physical movement. This list is not intended to be exhaustive, and one or more movement output components 200-26C can include other movement output components not explicitly identified herein that detect, generate, and/or otherwise provide data that can be used (e.g., processed, stored, and/or transformed) for outputting movement output. As illustrated in FIG. 5, output components 200-26 include one or more haptic output components 200-26D. One or more haptic output components 200-26D can include any component that functions to output (e.g., generate, create, and/or display), and/or cause output of, a haptic output (e.g., an output that is physically perceptible using tactile sensation, such as a vibration, pressure, texture, and/or shape). Examples of one or more haptic output components 200-26D can include: a speaker, a component that generates vibrations, a component that generates texture changes, a component that generates pressure changes, and/or a component that creates perceivable tactile effects. This list is not intended to be exhaustive, and one or more haptic output components 200-26D can include other haptic output components not explicitly identified herein that detect, generate, and/or otherwise provide data that can be used (e.g., processed, stored, and/or transformed) for outputting haptic output.
[0122] As illustrated in FIG. 5, output components 200-26 include one or more communications components 200-26E. One or more communications components 200-26E can include any component that functions to send and/or receive communications (e.g., an antenna, a modem, a network interface component, an encoder, a decoder, and/or a communication protocol stack) internal and/or external to agent system 200-20. In some embodiments, the communications can be between different devices and/or between components of the same device. In some embodiments, the communications can include control signals and/or data (e.g., messages, instructions, files, application data, and/or media streams). In some embodiments, one or more communications components 200-26E includes one or more features of one or more communications components 200-22B (e.g., as described above). In some embodiments, one or more communications components 200-26E are the same as one or more communications components 200-22B (e.g., one or more components that handle communication inputs and outputs and thus be considered as either and/or both an input component and an output component).
[0123] Throughout this disclosure, reference can be made to movement output (e.g., referred to in various forms such as: movement, device movement, output of movement, device motion, output of motion, and/or motion output). In some embodiments, outputting (e.g., causing output of) movement refers to movement of an electronic device (e.g., a portion or component thereof relative to another portion and/or of the whole electronic device). In some embodiments, referring back to FIG. 2B, movement output can refer to device 200 actuating movement component 200-3 to move display portion 200-1 to the position illustrated in FIG. 2B (e.g., from the position in FIG. 2A). In some embodiments, movement output is not (e.g., does not include and/or does not only include) haptic output (e.g., haptic movement output). In some embodiments, movement output is not (e.g., does not include and/or does not only include) vibration output. In some embodiments, movement output is not (e.g., does not include and/or does not only include) oscillating movement (e.g., movement of an actuator that merely causes vibration by moving a component repeatedly along a path that is internal to the device). In some embodiments, movement output includes (e.g., requires and/or results in) changing a location and/or pose of at least a portion of (and/or the entirety of) a component or the electronic device. In some embodiments, movement output includes output that moves at least a portion of (and/or the entirety of) a component or the electronic device from a first location and/or first pose to a second location and/or second pose. In some embodiments, with respect to FIGS. 2A-2C, display portion 200-1 is shown in a different location (e.g., in space) and pose (e.g., relative to base portion 200-2) in each of FIGS. 2A, 2B, and 2C. In some embodiments, movement output includes output that moves at least a portion (and/or the entirety of) a component or the electronic device to a third location and/or third pose (e.g., from the first location and/or first pose and/or from the second location and/or the second pose). In some embodiments, the third location and/or the third pose is the same as the first location and/or first pose and/or as the second location and/or the second pose. In some embodiments, movement output can include device 200 in FIG. 2A beginning from the first position illustrated in FIG. 2A, moving to the second position illustrated in FIG. 2B, and moving to return to the first position illustrated in FIG. 2A. In some embodiments, movement output can include device 200 in FIG. 2A beginning from the first position illustrated in FIG. 2A, moving to the second position illustrated in FIG. 2B, and continuing movement to come to rest at the third position illustrated in FIG. 2C.
[0124] Throughout this disclosure, an electronic device can be illustrated in (and/or described as being in) different locations and/or poses at different times. In some embodiments, in FIG. 2A illustrates device 200 in the first position, FIG. 2B illustrates device 200 in the second position, and FIG. 2A illustrates device 200 in the third position. In some embodiments, the electronic device moves itself between such locations and/or poses (e.g., using movement output). In some embodiments, device 200 moves from the first position to the second position under its own power (e.g., using a power source and one or more actuators to cause movement). In particular, any example herein that illustrates and/or describes an electronic device being at different locations and/or poses (e.g., at different times) should be understood to cover a scenario in which the device moved itself between such locations and/or poses (e.g., unless otherwise clearly indicated).
[0125] Throughout this disclosure, reference can be made to “performing output,” “causing output,” and/or “outputting” (e.g., by one or more output generation devices and/or by one or more output generation components) (and/or similar such phrases). In some embodiments, outputting (e.g., or the aforementioned variants) includes (and/or is) outputting movement (e.g., movement output as described above).
[0126] Throughout this disclosure, reference can be made to “displaying,” “causing display of,” and/or “outputting visual content” (e.g., by one or more display components) (and/or similar such phrases). In some embodiments, displaying (e.g., or the aforementioned variants) includes displaying visual content in connection with outputting movement (e.g., movement output as described above).
[0127] Throughout this disclosure, reference can be made to “outputting audio,” “causing output of audio,” and/or “providing audio output” (e.g., by one or more audio generation components and/or by one or more audio output devices) (and/or similar such phrases). In some embodiments, outputting audio (e.g., or the aforementioned variants) includes outputting audio content in connection with outputting movement (e.g., movement output as described above).
[0128] Throughout this disclosure, reference can be made to movement of an avatar (e.g., or other representation of a user, an agent and/or a character that is displayed) (e.g., by one or more display components) (and/or similar such phrases). In some embodiments, moving an avatar (e.g., or the aforementioned variants) includes displaying movement of visual content in connection with outputting movement (e.g., movement output as described above). In some embodiments, displaying an avatar nodding in agreement can include movement of the electronic device in a similar manner as the avatar movement (e.g., mimicking nodding). In some embodiments, moving an avatar (e.g., or the aforementioned variants) includes outputting movement (e.g., movement output as described above) without displaying movement of visual content. In some embodiments, a device can perform movement output that mimics nodding without moving a displayed avatar (e.g., the avatar does not move relative to the display). As illustrated in FIG. 5, agent system 200-20 can optionally interface with external components such as external database 200-30, remote processing component 200-32, and/or remote administration component 200-34. In some embodiments, external database 200-30 represents one or more functions that provide data storage resources accessible to agent system 200-20. In some embodiments, access to the data of external database 200-30 is provided directly to agent system 200-20 (e.g., the agent system manages the database) and/or indirectly to agent system 200-20 (e.g., a database is managed by a different system, but data stored therein can be provided and/or stored for use by agent system 200-20). In some embodiments, external database 200-30 is dedicated to (e.g., only for use by) agent system 200-20, is not dedicated to agent system 200-20 (e.g., is a database of a web service accessible to different agent systems), and/or is a combination of both dedicated and non-dedicated database resources. In some embodiments, remote processing component 200-32 represents one or more components that function as a data processing resource that is accessible to agent system 200-20. In some embodiments, access to remote processing component 200-32 is provided directly to agent system 200-20 (e.g., the agent system manages the processing resources) and/or indirectly to agent system 200-20 (e.g., a processing resource managed by a different system, but that can provide data processing for the benefit of agent system 200-20). In some embodiments, remote processing component 200-32 is dedicated to (e.g., only for use by) agent system 200-20, is not dedicated to agent system 200-20 (e.g., is a processing resource of a web service accessible to different agent systems), and/or is a combination of both dedicated and non-dedicated processing resources. Examples of data processing include processing image data (e.g., for feature extraction and/or object detection), processing audio data (e.g., for processing natural language speech input via a large language model), and/or training a machine learning algorithm and/or model. In some embodiments, remote administration component 200-34 represents functions that include and/or are related to administrative functions. In some embodiments, such administrative functions can include providing component updates to agent system 200-30 (e.g., software and/or firmware updates), managing accounts (e.g., permissions, access control, and/or preferences associated therewith), synchronizing between different agent systems and/or components thereof (e.g., such that an agent accessible via multiple devices of a user can provide a consistent user experience between such devices), managing cooperation with other services and/or agent systems, error reporting, managing backup resources to maintain agent system reliability and/or agent availability, and/or other functions required by agent system 200-20 to perform operations, such as those described herein.
[0129] The various components of agent system 200-20 described above with respect to FIG. 5 represent functional blocks that represent functionality. This functionality can be implemented on the same and/or different hardware (e.g., physical components) and/or by the same and/or different software. In some embodiments, the functional blocks can be implemented using one or more physical components, devices (e.g., computer system 100 and/or device 200), and/or software programs. In other words, each functional block does not necessarily represent a single, discrete physical component, device, and/or software program, but can be implemented using one or more of these. Further, agent system 200-20 can include multiple implementations of functionality represented by a respective functional block. In some embodiments, agent system 200-20 can include multiple different model components representing ML models that are used in different contexts, can include multiple different API components representing different APIs that are used for different services, and/or can include multiple different visual output components that are used for outputting different types of visual output.
[0130] Attention is now turned to discussion of concepts that can arise with respect to operation of an agent.
[0131] As discussed throughout, an agent can be capable of interacting with a user. In some embodiments, this capability includes the ability to process explicit requests, commands, and/or statements. In some embodiments, explicit requests, commands, and/or statements include and/or are interpreted as instructions directed to accomplishing a task (e.g., display X, complete task Y, and/or perform operation Z). In some embodiments, an agent includes the ability to process implicit requests, commands, and/or statements. In some embodiments, an implicit request, command, and/or statement does not include an explicit request, command, and/or statement. In some embodiments, “I like going to Europe,” can be interpreted as an implicit request, command, and/or statement which, in response to detecting, device 200 displays an itinerary in response to the statement. As another example, “This picture is for my grandmother,” can be interpreted as an implicit request, command, and/or statement which, in response to detecting, device 200 displays suggestions for modifying the picture). As another example, “I’m so tired,” can be interpreted as an implicit request, command, and/or statement which, in response to detecting, device 200 causes a sleep meditation application to begin a meditation session. As yet another example, “I miss my grandad” can be interpreted as an implicit request, command, and/or statement when, in response to detecting, device 200 can initiate a live communication session (e.g., telephone call, video call, and/or text messaging session) with grandad. In some embodiments, an implicit request is more likely to be processed according to one or more current environmental context, operational context, and/or user context, while an explicit request is less likely to be processed according to one or more current environmental context, operational context, and/or user context. In some embodiments, the phrase, “call my grandad,” can be an explicit request, and in response to detecting the request, device 200 will initiate a live communication session with grandad, irrespective of one or more current environmental context, operational context, and/or user context. However, the phrase, “I miss my grandad,” can be an implicit request, and in response to detecting the request, device 200 can display a list of gifts to buy for grandad if a user has been recently talking about buying gifts or could call grandad in another context that does not include the user recently discussing buying gifts. In some embodiments, a request can include one or more explicit requests and one or more implicit requests. In some embodiments, an implicit request is responded to independently from an explicit request; and in other embodiments, a response to an implicit request is dependent on an explicit request.
[0132] Reference can be made herein to a response by an agent that is output by a device. In some embodiments, a response includes an audio portion (e.g., audio output, audible output, sound, and/or speech) (also referred to herein as a “verbal response,” an “audio response,” and/or an “audible response) and/or a visual portion (e.g., display and/or movement of a representation and/or avatar). In some embodiments, a response includes a movement portion (e.g., movement of the device). In some embodiments, a response includes a haptic portion (e.g., touch and/or vibration).
[0133] Reference can be made herein to an internal dialogue, internal context, and/or an operational context, which can refer to a dynamic context or dynamic decision-making process of the device, an internal state of device 200, and/or internal data the device is partially basing its decision on. In some embodiments, an internal dialogue includes a set of one or more rules, characteristics, detections, and/or observations that the computer system uses to generate a response to one or more commands, questions, and/or statements). In some embodiments, the set of one or more rules, characteristics, detections, and/or observations are learned and/or generated via deep learning and/or one or more machine learning algorithms, and/or using one or more machine learning and/or system agents. In some embodiments, an internal dialogue is generated in real-time. In some embodiments, an internal dialogue is locally stored and/or stored via the cloud. In some embodiments, an internal dialogue can be modified, updated, and/or deleted. In some embodiments, an internal dialogue is generated based on other internal dialogues.
[0134] Reference can be made herein to personality and/or behavior (or a representation of personality /behavior) (e.g., of an agent, user, and/or character). In some embodiments, personality and/or behavior refers to a set of one or more characteristics that the device detects, has knowledge of, conforms to, applies, and/or tracks. In some embodiments, the personality or behavior is used as basis to perform operations. In some embodiments, an agent can detect a user’s personality and respond in a manner based on the personality (e.g., output different responses in response to different user personalities). As another example, the agent can output a response having characteristics that correspond to one or more characteristics that correspond to the personality and/or behavior (e.g., output a response in different ways that depend on personality of the agent). In some embodiments, such characteristics represent and/or mimic personality of a user, such as how the user acts and/or speaks. In some embodiments, such characteristics approximate a user’s personality.
[0135] In some embodiments, an agent is a system agent. In some embodiments, a system agent is an agent that corresponds to a process that originates from and/or is controlled by an operating system of the device (e.g., the device implementing the agent). In some embodiments, an agent is an application agent. In some embodiments, an application agent is an agent that corresponds to a process that originates from and/or is controlled by an application of (e.g., installed on and/or executed by) the device (e.g., the device implementing the agent).
[0136] Reference can be made herein to a representation (e.g., an avatar and/or avatar representation) of an agent (e.g., and/or of a user (e.g., person, object, and/or an animal) and/or a user interface object (e.g., an animated character)). In some embodiments, a representation of an agent refers to a set of output characteristics (e.g., visual and/or audio) of the agent (and/or the user and/or the user interface object). In some embodiments, a representation of an agent can include (and/or correspond to) a set of one or more visual characteristics (e.g., facial features of an animated face) and/or one or more audio characteristics (e.g., language and voice characteristics of audio output). In some embodiments, a representation (e.g., of an agent) is used to represent output by the agent. In some embodiments, a device implementing an interactive agent outputs audio in a voice of the agent and displays an animated face of the agent moving in a manner to simulate the agent speaking the audio output. In this way, a user can feel like they are having a normal conversation with the agent. In some embodiments, a representation of an agent is (or is not) inclusive of personality and/or behavior characteristics (e.g., as described above). In some embodiments, a representation of an agent can include (and/or correspond to) a set of visual characteristics (e.g., facial features of an animated face) and also a set of personality characteristics. In some embodiments, a representation of an agent includes a set of user characteristics that correspond to visual representation of a user (e.g., representations of a user’s appearance, voice, and/or personality are used as an avatar that appears to move and/or speak). In some embodiments, a representation is a representation of a face (e.g., a user interface object that is output having features that simulate a face and/or facial expressions of a person (e.g., for conveying information to a viewer)).
[0137] In some embodiments, a character (e.g., of an agent and/or avatar) refers to a particular set of characteristics of a representation. In some embodiments, an avatar can take on (e.g., use, apply, interact with, and/or output according to) characteristics of a fictional and/or non-fictional character (e.g., from a movie, a show, a book, a series, and/or popular culture). [0138] In some embodiments, a voice (e.g., of an agent and/or avatar) refers to a set of one or more characteristics corresponding to sound output that resembles (e.g., represents, mimics, and/or recreates) vocal utterance (e.g., attributable and/or simulated as being output by an agent and/or avatar). In some embodiments, device 200 can output a sentence that sounds different depending on a voice used. In some embodiments, a particular character and/or avatar can be configured to use a particular voice (e.g., have a corresponding voice). In some embodiments, the particular voice can mimic a user’s voice.
[0139] In some embodiments, an appearance (e.g., of an agent and/or avatar) refers to a set of one or more characteristics corresponding to visual output that represents an avatar (and/or an agent). In some embodiments, device 200 can output an avatar that has a set of facial features forming an appearance that resembles a particular character from a movie.
[0140] In some embodiments, an expression of an avatar refers to a set of one or more characteristics corresponding to a particular visual appearance of a user, an avatar, and/or an agent. In some embodiments, device 200 can output an avatar that has a set of facial features arranged in a particular way to give the appearance of a facial expression (e.g., which can be used as a form of non-verbal communication to a user) (e.g., a frown is an expression of sadness, a smile is an expression of happiness, and/or wide open eyes is an expression of surprise). As another example, device 200 can output an avatar that has a set of body features (e.g., arms and/or legs) arranged in a particular way to give the appearance of a body expression (e.g., which can be used as a form of non-verbal communication to a user) (e.g., a hand gesture is an expression of approval, covering eyes is an expression of fear, and/or shrugging shoulders is an expression of lack of knowledge). In some embodiments, an expression includes movement (e.g., a head nod is an expression of agreement and/or disagreement) of the avatar. In some embodiments, device 200 can move, via the movement component, to indicate an expression with or without the avatar moving. In some embodiments, an agent performs one or more operations that depend on a user’s expression (e.g., detects if a person is sad and responds with a kind statement or question). In some embodiments, expressions (e.g., whether and/or how they are used and/or how they are output) depends on personality. In some embodiments, a first personality can use a particular expression more than a second personality. As another example, an expression (e.g., frown, smile, and/or how wide eyes are opened) for the first personality can appear different from the expression (and/or a similar and/or equivalent expression) for a second personality (e.g., the first personality smiles in a manner that reveals teeth, but the second personality smiles without revealing teeth).
[0141] In some embodiments, an agent (e.g., an avatar of the agent and/or an agent system (e.g., hardware and/or software) implementing the agent) mimics characteristics of another user, agent, and/or character (e.g., in personality, behavior, expressions, and/or voice). In some embodiments, mimicking includes mirroring a user (e.g., copying use of a phrase and/or movement detected from a user interacting with the agent). In some embodiments, mimicking characteristics of a user includes attempting to reproduce the characteristics of the user (e.g., in the exact same manner and/or in manner that resembles the characteristics but is not an exact reproduction of the characteristics). In some embodiments, an agent mimicking voice and/or expressions does not require the agent have the exact same voice and/or expressions as the user being mimicked (e.g., but rather simply resembles the user’s voice and/or expressions).
[0142] In some embodiments, a component and/or device uses (e.g., performs operations, makes decisions, and/or determines context based on) learned characteristics (e.g., characteristics of a context, user, and/or environment that the device has learned over time (e.g., via detection, prior experience, and/or feedback (e.g., from one or more users)). In some embodiments, characteristics learned over time can include a user’s routine. In such example, if a particular user asks an agent for a summary of any new messages for the user at the same time every day, the agent can learn to perform operations automatically based on the learned characteristics of the routine (e.g., what data is needed, when the data is needed, and/or for which user). In some embodiments, use of learned characteristics enables an agent (and/or device) to improve understanding of (and/or responses to) a context, user, and/or environment, and/or to understand a context, user, and/or environment that otherwise was not (and/or would not be) understood (e.g., not responded to or responded to incorrectly). In some embodiments, learned characteristics are formed (e.g., by and/or for an agent) using reinforcement learning. In some embodiments, learned characteristics correspond to one or more levels of confidence, certainty, and/or reward (e.g., that are shaped by one or more reward functions). In some embodiments, learned characteristics (and/or how they are used to affect output of an agent and/or device) can change over time (e.g., levels confidence, certainty, and/or reward change over time). In some embodiments, output of a device before learning a set of learned characteristics can be different from output of the device after learning the set of learned characteristics. In some embodiments, a component and/or device uses learned knowledge. In some embodiments, similar to described above with respect to learned characteristics, learned knowledge can refer to information used to update (e.g., enhance, add to, and/or augment) a knowledge base of a device (e.g., for use by an agent implemented thereon). In some embodiments, multiple sets of learned characteristics for a user can be stored and/or used. In some embodiments, different sets of learned characteristics for different users can be stored and/or used.
[0143] Reference can be made herein to interaction with an agent (and/or a device). In some embodiments, an interaction refers to a set of one or more inputs and/or outputs of a device implementing the agent and one or more users. In some embodiments, an interaction can be an input by a user (e.g., “Please turn on the lights”) and a corresponding output (e.g., causing the lights to turn on and/or a response by the device of “Okay”). In some embodiments, interaction can include multiple inputs/outputs by one or more of the parties to the interaction (e.g., device and/or users). In some embodiments, an interaction can include a first input by a user (e.g., “Please turn on the lights”) and a corresponding first output (e.g., “Which lights?”), and also include a second input by the user (e.g., “Kitchen lights”) and a second output from the device (e.g., “Okay”). In some embodiments, which inputs and/or outputs are considered together as an interaction is based on a logical and/or contextual grouping (e.g., interactions within the previous thirty (30) seconds and/or interactions relating to turning on the lights). As one of skill will appreciate, an interaction can be considered in a manner that depends on the implementation (e.g., determining when an interaction is complete can involve determining if the user still present (e.g., speaking at all) and/or if the user still talking about the lights or has moved onto a different topic). In some embodiments, an interaction is a current interaction (e.g., ongoing, presently occurring, and/or active). In some embodiments, an interaction is a previous interaction. The examples above describe a device having a conversation with a user. In some embodiments, a conversation is between two or more users (e.g., users in an environment). In some embodiments, a device can detect a conversation between to users (e.g., the users are directing speech and responses to each other, rather than to the device).
[0144] In some embodiments an agent (and/or device) determines and/or performs an operation based on an intent corresponding to a user. In some embodiments, a device detects user input and outputs a response that depends on an intent of the user input. In some embodiments, a device detects user input that includes a pointing gesture detected together with verbal instruction to “turn on that light,” and in response, the device turns on the light that is determined to correspond to the intent of the input (e.g., the light toward which the pointing gesture directed). In some embodiments, intent is determined (e.g., by the device that detects input and/or by one or more other devices) using one or more of: one or more inputs, knowledge (e.g., learned knowledge about a user based on a history of observed behavior, personality, and interactions), learned characteristics, and/or context. In some embodiments, intent is determined from one or more types of input (e.g., verbal input, visual input via a camera, and/or contextual input).
[0145] Attention is now directed towards embodiments of user interfaces (“UI”) and associated processes that are implemented on an electronic device, such as computer system 100 and/or device 200.
[0146] FIGS. 6A-6G illustrate exemplary user interfaces for modifying an appearance, location, and/or size of a user interface object in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes in FIGS. 7, 8, and 9.
[0147] FIGS. 6A-6G illustrate a computer system 600 (e.g., a tablet). It should be recognized that computer system 600 can be other types of computer systems such as a smart phone, a smart watch, a laptop, a communal device, a smart speaker, an accessory, a personal gaming system, a desktop computer, a fitness tracking device, and/or a head-mounted display (HMD) device. In some embodiments, computer system 600 includes and/or is in communication with one or more input devices and/or sensors (e.g., a camera, a lidar detector, a motion sensor, an infrared sensor, a touch-sensitive surface, a physical input mechanism (such as a button or a slider), and/or a microphone). Such input devices and/or sensors can be used to detect presence of, attention of, statements from, inputs corresponding to, requests from, and/or instructions from a user in an environment. It should be recognized that, while some embodiments described herein refer to inputs being voice inputs, other types of inputs can be used with techniques described herein, such as touch inputs via a touch- sensitive surface and air gestures detected via a camera. In some embodiments, computer system 600 includes and/or is in communication with one or more output devices (e.g., a display screen, a projector, a touch-sensitive display, speaker, and/or a movement component). Such output devices can be used to present information and/or cause different visual changes of computer system 600. In some embodiments, computer system 600 includes and/or is in communication with one or more movement components (e.g., an actuator, a moveable base, a rotatable component, and/or a rotatable base). Such movement components, as discussed above, can be used to change a position (e.g., location and/or orientation) of computer system 600 and/or a portion (e.g., including one or more sensors, input components, and/or output components) of computer system 600. In some embodiments, computer system 600 includes one or more components and/or features described above in relation to computer system 100 and/or device 200. In some embodiments, computer system 600 includes one or more agents and/or functions of an agent as described above with respect to FIG. 5. In some embodiments, computer system 600 is, includes, implements, and/or is in communication with one or more agent systems, as described above with respect to FIG. 5, for performing (and/or causing performance of) one or more operations of an agent.
[0148] As illustrated in FIGS. 6A-6G, computer system 600 displays a user interface object (e.g., user interface object 604). In some embodiments, user interface object 604 is a representation of an animated face and/or the agent. The animated face can include eyes, a mouth, a nose, a body, and/or other elements. In some embodiments, user interface object 604 corresponds to an application and/or a system process of computer system 600. In such embodiments, it should be recognized that user interface object 604 can have different appearances. In some embodiments, user interface object 604 can be used with multiple different applications. In other embodiments, different user interface objects can be used with different applications. In some embodiments, user interface object 604 discussed below and illustrated in FIGS. 6A-6G can be used for a set of one or more applications while another user interface object is used for another set of one or more applications.
[0149] FIGS. 6A-6D illustrate a process of computer system 600 modifying an appearance (e.g., modifying a portion, altering a color, changing a shape, modifying a size, and/or changing a location) of user interface object 604 in response to computer system 600 agreeing with an input (e.g., detected via the one or more input devices described above) from a user. In some embodiments, computer system 600 can cause user interface object 604 to appear to (1) initiate and/or maintain eye contact with the user when computer system agrees with the input and/or (2) cease and/or avoid eye contact with the user when computer system does not agree with the input. It should be recognized that other movements of a portion (e.g., including an input device and/or output device) of computer system 600 and/or user interface object 604 can be used in response to computer system 600 agreeing with an input from the user (e.g., moving the portion up and down and/or moving user interface object 604 up and down (e.g., to simulate the nodding of a head)) or disagreeing with an input from the user (e.g., moving the portion to the left and right and/or moving user interface object 604 to the left and right (e.g., to simulate the shaking of a head)).
[0150] In some embodiments, computer system 600 agrees with an input when the input includes a correct and/or truthful statement, as determined by computer system 600 and/or another computer system in communication with computer system 600. In some embodiments, the input can include an indication that “The sky is blue.” In response to detecting the input computer system 600 can determine that the sky is blue and therefore that computer system 600 agrees with the input. It should be recognized that other criteria can be used to determine whether computer system 600 agrees with an input, such as whether the input is consistent with one or more user preferences known by computer system 600 (e.g., previously input by the user and/or detected by computer system 600 in a previous interaction with the user). In some embodiments, a user preference of the user can be that the user likes blue based on the user previously telling computer system 600 such. Accordingly, when computer system 600 detects a question or statement from the user corresponding to the color blue (e.g., “Does this blue shirt look good on me?”), computer system 600 initiates eye contact between user interface object 604 and the user (and/or performs some other output indicating agreement, such as outputting audio of “Yes” and/or making user interface object 604 appear to nod its head up and down). When computer system 600 detects a question or statement from the user corresponding to the color red (e.g., “Does this red shirt look good on me?”), computer system 600 does not cause user interface object 604 to initiate eye contact with the user (and/or causes user interface object 604 to perform some other output indicating disagreement, such as outputting audio of “No” and/or making user interface object 604 appear to shake its head from left to right). In some embodiments, it is not required for the user to use the word “blue” in a question or statement but instead computer system 600 can detect, via a camera, a characteristic of an object, such as a color of a shirt.
[0151] FIG. 6A illustrates computer system 600 displaying user interface 602. User interface 602 includes user interface object 604, as discussed above. At FIG. 6A, computer system 600 detects voice input 606 (e.g., “The sky is currently blue”) from a user. Accordingly, voice input 606 does not include an explicit indication to change user interface object 604. That is, computer system 600 does not alter the appearance of user interface object 604 based on a command from the user (e.g., the user telling user interface object 604 “Look at me”). Rather, computer system 600 determines to alter the appearance of user interface object 604 based on agreement or disagreement with voice input 606. In some embodiments, instead of a statement, computer system 600 detects an input from the user in the form of a question. In such embodiments, computer system 600 can change the appearance of user interface object 604 based on the question, such as causing user interface object 604 to look in a direction of and/or make eye contact with the user when computer system 600 agrees with the question and/or is answering the question with a positive answer (e.g., “Yes”).
[0152] At FIG. 6B, computer system 600 determines that computer system 600 agrees with voice input 606. As illustrated in FIG. 6B, in response to detecting voice input 606 and determining that computer system 600 agrees with voice input 606, computer system 600 changes an appearance of user interface object 604. In some embodiments, computer system 600 displays user interface object 604 as appearing to make eye contact with the user (e.g., looking up in a direction of the user) by changing the position of the eyes of user interface object 604. In some embodiments, while computer system 600 causes user interface object 604 to change appearance, computer system 600 maintains user interface object 604 at a particular location in user interface 602 and/or maintains the appearance of one or more other user interface objects (e.g., a clock, an icon, a background, and/or information) in user interface 602.
[0153] In some embodiments, computer system 600 maintains user interface object 604 in a way that appears to make eye contact with the user for a predetermined period of time. That is, after the predetermined period of time has passed since making eye contact, computer system 600 ceases to display user interface object 604 as making eye contact with the user (e.g., as illustrated in FIG. 6C and described further below).
[0154] In some embodiments, computer system 600 ceases to display user interface object 604 as making eye contact with the user in response to detecting another voice input (e.g., from the user or another user different from the user. That is, in some embodiments, while computer system 600 displays user interface object 604 moving in and/or positioned in a way to make eye contact with the user and computer system 600 detects another voice input, computer system 600 ceases displaying user interface object 604 as appearing to make eye contact with the user until the other voice input is completed. In some embodiments, computer system 600 ceases displaying eye contact with the user irrespective of an agreement or disagreement with respect to the other voice input when detecting the other voice input to show to the user that user interface object 604 is restarting before acknowledging the other voice input. In some embodiments, if user interface object 604 is making eye contact with the user in response to computer system 600 detecting a voice input of “The sky is blue,” when computer system 600 detects a question or statement from the user saying, “Birds have feathers,” computer system 600 will cease eye contact between user interface object 604 and the user before again displaying user interface object 604 making eye contact with the user as a result of agreeing with the second question or statement.
[0155] As illustrated in FIG. 6C, computer system 600 displays user interface object 604 as looking forward and no longer having eye contact with the user. At FIG. 6C, computer system 600 detects voice input 608 (e.g., “The sky is currently green”) from the user.
[0156] At FIG. 6D, computer system 600 does not agree with voice input 608 (e.g., based on weather information indicating sky color and/or other information known by computer system 600). Accordingly, in response to detecting voice input 608, computer system 600 displays user interface object 604 as continuing to look straight ahead in the same position as illustrated in FIG. 6C.
[0157] In some embodiments, in response to detecting a question or statement from the user with which computer system 600 disagrees, computer system 600 does not change an appearance of user interface object 604 (e.g., user interface object 604 continues to look forward). In some embodiments, in response to detecting a question or statement from the user with which computer system 600 disagrees, computer system 600 displays user interface object 604 moving in a different manner as compared to when computer system 600 detects a question or statement with which it agrees (e.g., looking down, shaking head, and/or closing eyes).
[0158] FIGS. 6D-6G illustrate a process of computer system 600 displaying user interface object 604 interacting with content that computer system 600 displays. As illustrated in FIG. 6D, computer system 600 displays user interface 602 with user interface object 604 looking straight ahead (e.g., in a direction other than within user interface 602) and/or in a direction of a user. FIG. 6D also illustrates user interface object 604 as being the same size as illustrated in FIGS. 6A-6C and taking up a majority of user interface 602. At FIG. 6D, computer system 600 detects voice input 610 from the user (e.g., “What is the weather?”). Voice input 610 represents a request for computer system 600 to tell the user a current state of the weather.
[0159] In some embodiments, the request to interact with computer system 600 can be alternative types of input (e.g., an air gesture, a touch input, and/or a gaze input of the user). In some embodiments, the request to interact with computer system 600 does not include an explicit request. In some embodiments, the user can make the statement “It’s hot outside” and, in response, computer system 600 displays content that tells the user the temperature outside. In some embodiments, an explicit request includes directly asking computer system 600 a question as a request for information. An explicit request can also include a direct question for a voice output from computer system 600, such as “Please answer this question” or “Can you help me?”
[0160] As illustrated in FIG. 6E, in response to detecting voice input 610, computer system 600 shrinks user interface object 604 from its size as illustrated in FIG. 6D into a smaller size in the bottom left corner of user interface 602. In some embodiments, computer system 600 displays user interface object 604 as shrinking into the bottom left comer of user interface 602 because the user is standing on the left side of computer system 600. In some embodiments, if the user was standing on the right side of computer system 600, computer system 600 would display user interface object 604 as shrining into the bottom right corner of user interface 602. In some embodiments, computer system 600 displays user interface object 604 as shrinking into the bottom left comer of user interface 602 because content that computer system 600 will display (e.g., as illustrated in FIG. 6F) on a right side of user interface 602. In some embodiments, if the content would be displayed on the left side of user interface 602, computer system 600 would display user interface object 604 as shrining into the bottom right corner of user interface 602.
[0161] As illustrated in FIG. 6E, computer system 600 displays the eyes of user interface object 604 as maintaining the same direction as illustrated in FIG. 6D. In some embodiments, after (e.g., or while) shrinking user interface object 604, computer system displays user interface object 604 as beginning to look in a direction of a location in which computer system 600 will display content (e.g., as illustrated in FIG. 6F and discussed further below). [0162] As illustrated in FIG. 6F, in response to detecting voice input 610, computer system 600 displays content 612, represented in this example as a temperature value of 70 degrees. In some embodiments, computer system 600 displays multiple sets of content (e.g., images, icons, and/or words).
[0163] In some embodiments, computer system 600 displays user interface object 604 as looking at the user while the user is interacting with computer system 600 and looking at content when content is being output by computer system 600 and/or viewed by the user. As illustrated in FIG. 6F, in response detecting voice input 610 and/or displaying content 612, computer system 600 can display user interface object 604 as moving its gaze from one set of content to another as computer system 600 detects that the process of the user receiving and/or looking at each set of content is complete. In some embodiments, computer system 600 displays user interface object 604 as moving its gaze from one set of content to another in order to direct the attention of the user from one set of content to another. In some embodiments, in response to detecting voice input 610 and/or displaying content 612, computer system 600 does not alter a position or gaze of user interface object 604 from its position as illustrated in FIG. 6E.
[0164] Notably, computer system 600 displays content 612 in a portion of user interface 602 in which computer system 600 previously displayed user interface object 604 as illustrated in FIG. 6D. That is, in FIG. 6F, content 612 occupies the majority of user interface 602. Also illustrated in FIG. 6F, in response to detecting voice input 610, computer system 600 alters user interface object 604 from looking out of user interface 602 to looking in a direction of content 612. In this example, computer system 600 displays user interface object 604 as looking in the direction that is a target direction for the user to look. That is, computer system 600 directs the user’s gaze to certain areas of user interface 602 that are of importance to the user’s request for information.
[0165] In some embodiments, when computer system 600 is not waiting on a response (e.g., input) from the user (e.g., is in the process of performing a task or series of tasks) (e.g., after looking at content 612), computer system 600 displays user interface object 604 as looking in another direction other than at the user. In some embodiments, the other direction is content 612 in response to determining that the user is looking at content 612. In some embodiments, the other direction is another user within the environment. In some embodiments, the other direction is an object within the environment, such as a clock, a couch, and/or a television. In some embodiments, the other direction is in a direction that is not in a particular direction of something.
[0166] In some embodiments, computer system 600 can alter the gaze of user interface object 604 from being directed at content 612 to being directed at the user in response to detecting that the user ceases to look at content 612. In some embodiments, computer system 600 displays user interface object 604 as looking at content 612 as an indication to the user that the user should be looking at content 612 (e.g., computer system 600 has displayed content 612 for a period of time and the user has not yet looked at it) (e.g., computer system 600 was outputting audio such as audio response 614, so the user was not focused on and/or looking at content 612).
[0167] As discussed above, computer system 600 can detect multiple people within an environment. In some embodiments, computer system 600 detects a condition (e.g., a completed task) in response to which computer system 600 ceases displaying user interface object 604 as gazing at content 612. After computer system 600 ceases to display user interface object 604 as gazing at content 612, computer system 600 displays user interface object 604 as gazing at the user in the environment that initiated voice input 610. That is, in some embodiments, computer system 600 returns the gaze of user interface object 604 to the user whose voice computer system 600 detected and not another user that did not initiate voice input 610. In some embodiments, computer system 600 detects an interaction from the user and, in response, initiates a process of determining which user to display user interface object 604 as looking at. In some embodiments, a first user can ask computer system 600 to display a video to a second user and, in response, computer system 600 displays the video and directs the gaze of user interface object 604 to the second user. For another example, a first user can make a voice input directed to computer system 600, computer system 600 can display user interface object 604 as turning its gaze from the first user that made the voice input to another user to indicate that user interface object 604 is awaiting a voice input from the other user.
[0168] Also illustrated in FIG. 6F, in response to detecting voice input 610 and/or computer system 600 shrinking user interface object 604 and/or displaying content 612, computer system 600 outputs audio response 614 (e.g., “The temperature is 70 degrees”) from user interface object 604. Audio response 614 tells the user in words the information that computer system 600 displays in the form of content 612. That is, audio response 614 tells the user that the temperature is 70 degrees, which is the temperature that content 612 displays. In some embodiments, computer system 600 uses audio response 614 to convey information (to the user) that is contextually related to content 612. That is, computer system 600 can use audio response 614 to provide additional context to content 612 that might help the user absorb the meaning of content 612. In some embodiments, computer system 600 replaces previously displayed content (e.g., icons and/or photos) with content 612. That is, in some embodiments, content 612 replaces content that computer system 600 displayed first with content relevant to the user’s request for information.
[0169] As illustrated in FIG. 6G, computer system 600 detects that computer system 600 has finished its task of displaying content 612 and outputting audio response 614. In some embodiments, in response to detecting that computer system 600 has finished its task and/or that a set of criteria is met (e.g., computer system 600 detects that an input from the user is needed, that a predetermined period of time has passed, and/or that the user is looking at content 612), computer system 600 displays user interface object 604 as returning its gaze toward the user. In some embodiments, in response to detecting that the user has finished looking at content 612 (e.g., the user has looked at content 612 for a predetermined period of time and then looked in another direction), computer system 600 ceases displaying content 612 and redisplays user interface object 604 at its original size as illustrated in FIGS. 6A-6D. In some embodiments, in response to detecting that the user has finished looking at content 612, computer system 600 shrinks content 612 to a smaller size.
[0170] FIG. 7 is a flow diagram illustrating a method for displaying an object facing a direction using a computer system in accordance with some embodiments. Process 700 is performed at a computer system (e.g., 100, 200, and/or 600). Some operations in process 700 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
[0171] As described below, process 700 provides an intuitive way for displaying an object facing a direction. The method reduces the cognitive burden on a user for displaying an object facing a direction, thereby creating a more efficient human-machine interface. For battery operated computing devices, enabling a user to display an object facing a direction faster and more efficiently conserves power and increases the time between battery charges. [0172] In some embodiments, process 700 is performed at a computer system (e.g., 600) that is in communication with a display component (e.g., a display screen, a projector, and/or a touch-sensitive display) and the one or more input devices (e.g., a camera, a depth sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a fitness tracking device, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0173] While (and/or after and/or in response to) displaying, via the display component, content (e.g., 612) (e.g., sound, media, visual content, audio content, a topic, a user, a discussion, and/or a relationship), the computer system detects (702) (e.g., via one or more input devices) a first interaction condition (e.g., speech, touch, context (e.g., what is being discussed, talked about, and/or output) and/or movement of the computer system and/or a user) (e.g., via a verbal input (e.g., an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) corresponding to the content (e.g., as described above with respect to FIGS. 6D-6F). In some embodiments, the first interaction condition is detected via the one or more input devices.
[0174] In response to (704) detecting the first interaction condition corresponding to the content, in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, the computer system displays (706), via the display component, a representation of a face (e.g., 604) looking (e.g., in a manner that indicates eye contact, in a manner that indicates gaze, appearing to look at, directed to, and/or eyes, mouth, face, and/or noise appear to be looking at) in the direction of (e.g., corresponding to and/or at) the content (e.g., 612) (e.g., as described above with respect to FIG. 6F). In some embodiments, the first set of one or more criteria includes a criterion that is satisfied when the content is being pointed to, discussed, reviewed, introduced, and/or interacted with.
[0175] In response to (704) detecting the first interaction condition corresponding to the content, in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria (e.g., without satisfying the first set of one or more criteria), the computer system displays (708), via the display component, the representation of the face (e.g., 604) looking (e.g., in a manner that indicates eye contact, in a manner that indicates gaze, appearing to look at, directed to, and/or eyes, mouth, face, and/or noise appear to be looking at) in the direction of (e.g., corresponding to and/or at) a first user (e.g., in a manner that indicates eye contact and/or eye connection with a user) (e.g., a user, a person, an animal, and/or an object) detected in a field-of-detection of the one or more input devices (e.g., as described above with respect to FIG. 6G). In some embodiments, the second set of one or more criteria includes a criterion that is satisfied when the content is not being pointed to, not being discussed, not being reviewed, not being introduced, and/or not being interacted with. In some embodiments, the second set of one or more criteria includes a criterion that is satisfied when a determination is made that a response is needed from the first user, that a portion of the content is directed to the first user, that agreement is needed, required, and/or desired from the first user, and/or the portion of content being output concerns the first user. Displaying a representation of a face in response to interaction enables the computer system to provide the user with a confirmation that an interaction has been received, thereby providing improved visual feedback. Displaying a representation of a face via the computer system in a direction based on which set of one or more criteria is met enables the computer system to selectively display the representation of the face in the proper direction based on the type of interaction and allows the computer system to display an indication for a particular use dependent on the interaction, thereby providing improved visual feedback, performing an operation when a set of conditions has been met without requiring further input, and/or reducing the number of user inputs needed to perform an operation.
[0176] In some embodiments, the first set of one or more criteria includes a first criterion that is satisfied when a determination is made that a portion of the content (e.g., 612) has been displayed for at least a predetermined period of time (e.g., as described above with respect to FIG. 6F) (e.g. 0.1-10 seconds) (e.g., the first criterion is satisfied when the content is initially displayed, new portion of content is initially displayed and/or new content is initially displayed) (e.g., the first interaction condition is a request to display the content when the content was not previously (e.g., immediately previously) displayed, a new portion of the content is displayed and/or new content is displayed). Displaying the representation of the face in the direction of the content based on a criterion that a portion of the content is displayed for a least a predetermined amount of time enables the computer system to confirm content is on the display component in order to indicate the position of the content to the user and/or reduce distractions on the display component, thereby providing improved visual feedback, performing an operation when a set of conditions has been met without requiring further input, and/or reducing the number of inputs needed to perform an operation.
[0177] In some embodiments, the first set of one or more criteria includes a second criterion that is satisfied when a determination is made that the first user is looking in (e.g., for at least a threshold amount of time (0.1-10 seconds) and/or is gazing at and/or in) the direction (e.g., based on the detection (e.g., sound, gaze, and/or eyes) of the user in the field- of-detection of the one or more input devices) (e.g., left, right, up, down, and/or any combination thereof) of the content (e.g., 612) (e.g., as described above with respect to FIG. 6F) (e.g., the first interaction condition is detected and/or occurs after and/or while displaying the content). In some embodiments, an interaction is detected and/or occurs before displaying the content that satisfies the first set of one or more criteria or the second set of one or more criteria. Displaying the representation of the face in the direction of the content based on a criterion that a user is looking in the direction of the content enables the computer system to selectively display the representation of the face in the proper direction based on the detection of the user via one or more input systems and/or reduce distractions on the display component, thereby providing improved visual feedback, performing an operation when a set of conditions has been met without requiring further input, and/or reducing the number of inputs needed to perform an operation.
[0178] In some embodiments, the first set of one or more criteria includes a third criterion that is satisfied when a determination is made that the first user should look at the content (e.g., 612) (e.g., as described above with respect to FIG. 6F) (e.g., based on a pattern of use (e.g., based on historical use and/or based on a historical pattern) of the first user corresponding to one or more previous interactions of the first user with content that includes an indication that the first user looked at content being displayed) (e.g., the first user is associated with (e.g., corresponds to, historically known to be connected to, detected as previously having) a pattern of use of looking at new content after a previous reaction condition, such as when: an application is initially launched, when one or more extra inputs are requested with respect to content being displayed (e.g., to verify content is correct, to verify moving to the next (e.g., new and/or a different portion of) content), when a question is outputted, when asking to send a text, and/or when to review text). In some embodiments, the determination that the first user should look at the content includes detecting that the first user has looked at the content being displayed one or more times after the content has been displayed. Displaying the representation of the face in the direction of the content based on a criterion that a user should look at the content enables to the computer system to confirm that content is being displayed to the user and/or allows the computer system to direct the user’s attention to the content, thereby providing improved visual feedback, performing an operation when a set of conditions has been met without requiring further input, and/or reducing the number of inputs needed to perform an operation.
[0179] In some embodiments, the second set of one or more criteria includes a fourth criterion that is satisfied when a determination is made that an input (e.g., 610) (e.g., an interaction, a command and/or a request) (e.g., corresponding to the first user and/or made by the first user) (e.g., a verbal input (e.g., an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) has been received(e.g., as described above with respect to FIGS. 6D-6E). In some embodiments, after displaying the representation of the face looking in the direction of the content for a first predefined period of time, the computer system displays, via the display component, the representation of the face looking in the direction of the first user detected in the field-of-detection of the one or more input devices. In some embodiments, after detecting, via the one or more input devices, the first user looking (e.g., for at least a second predefined period of time (0.1-10 seconds) and/or is gazing at and/or in) in the direction of the content, the computer system displays, via the display component, the representation of the face looking in the direction of the first user detected in the field-of-detection of the one or more input devices. Displaying the representation of the face in the direction of the user based on the criterion that an input has been received enables the computer system to confirm that an input has been received, thereby providing improved visual feedback to the user and performing an operation when a set of conditions has been met without requiring further input.
[0180] In some embodiments, the second set of one or more criteria includes a fifth criterion that is satisfied when a determination is made that the computer system (e.g., 600) is waiting for a response (e.g., as described above with respect to FIGS 6D-6G) (e.g., from the first user and/or another user) (e.g., waiting on (e.g., and/or is waiting on) an input and/or a particular type of input (e.g., tap input, an air gesture, a voice command, and/or a mouse click)) (e.g., the computer system requires one or more additional inputs in order to proceed with an operation) (e.g., first user needs to provide a second interaction condition (e.g., confirming a text message, choosing a contact, choosing between different applications and/or verification of content)). Displaying the representation of the face in the direction of the user based on the criterion that the computer system is waiting for a response allows the user to infer that a new input is needed to continue and/or provide the user with feedback that computer system needs an input (e.g., without displaying additional information that the computer system is waiting for a response), thereby providing improved visual feedback to the user, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0181] In some embodiments, in response to detecting the first interaction condition corresponding to the content and in accordance with a determination that the first interaction condition corresponding to the content satisfies a set of one or more criteria that includes a criterion that is satisfied when a determination is made that the computer system (e.g., 600) is not waiting for the response, the computer system displays, via the display component, the representation of the face(e.g., 604) looking in a second direction that is different from the direction of the first user detected in the field-of-detection of the one or more input devices (e.g., as described above with respect to FIG. 6B ) (and, in some embodiments, the direction of the content). In some embodiments, the second direction is in the field-of-view of the one or more input device but not in the field-of-detection of the first user. In some embodiments, the second direction is in the direction of the user interface but not in the direction of the content. In some embodiments, the second direction is away from the direction of the first user. Displaying the representation of the face in a second direction away from the user based on a criterion that the computer system is not waiting for a response enables the computer system to provide the user with feedback that the computer system does not need an input without displaying additional information on the user interface, thereby providing improved visual feedback to the user and performing an operation when a set of conditions has been met without requiring further input.
[0182] In some embodiments, the second direction corresponds to a second user (e.g., a second user, a person, an animal, and/or an object) detected in the field-of-detection (e.g., within the field-of-view of one or more input systems (e.g., camera, depth sensor, microphone)) of the one or more input devices. In some embodiments, the second user is different from the first user (e.g., as described above with respect to FIGS. 6D-6G). In some embodiments, while displaying, via the display component, the representation of the face looking in the second direction, representation of the face looks at the second user. Displaying the representation of the face in a second direction of a different user different from the first user based on a criterion that the computer system is not waiting for a response enables the computer system to confirm the detection of a second user in the field-of- detection of the one or more input devices, thereby providing improved visual feedback and performing an operation when a set of conditions has been met without requiring further input.
[0183] In some embodiments, the second direction corresponds to the content (e.g., 612) (e.g., as described above with respect to FIGS. 6F). In some embodiments, while displaying, via the display component, the representation of the face looking in the second direction, the representation of the face looks at the content (and/or at least a portion of the content). Displaying the representation of the face in a second direction corresponding to the content based on a criterion that the computer system is not waiting for a response provides the user with feedback that content is being displayed and allows the computer system to bring the user’s attention to the content, thereby providing improved visual feedback to the user and performing an operation when a set of conditions has been met without requiring further input.
[0184] In some embodiments, the second direction corresponds to an object (e.g., a physical and/or virtual object) (e.g., an object located and/or detected) in the field-of- detection of the one or more input devices (e.g., as described above with respect to FIGS. 6D- 6G) (e.g., in an environment (e.g., a physical, a virtual, or a mixed-reality environment) including the first user). Displaying the representation of the face in a second direction corresponding to an object in the field-of-detection based on a criterion that the computer system is not waiting for a response enables the computer system to confirm the detection of an object in the field-of-detection of the one or more input devices (without displaying additional information on the user interface to confirm detection) and/or provide feedback that the content is done outputting and no response is needed (e.g., it is time to exit the application, and/or no new content can be provided), thereby providing improved visual feedback, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component. [0185] In some embodiments, while (and/or after) displaying the representation of the face looking (e.g., 604) in the direction of the content (e.g., 612) (and, in some embodiments, without looking in the direction of the first user), the computer system detects (e.g., via one or more inputs devices) a second interaction condition (e.g., speech, touch, context (e.g., what is being discussed, talked about, and/or output) and/or movement of the computer system and/or a user) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold- and-drag input, a gaze input, an air gesture, and/or a mouse click)) corresponding to the content (e.g., 614) (e.g., and/or corresponding to the other content that is different from the content) (e.g., as described above with respect to FIG. 6F). In some embodiments, in response to detecting the second interaction condition corresponding to the content (e.g., 614), the computer system displays, via the display component, the representation of the face (e.g., 604) looking in the direction of the first user detected in the field-of-detection of the one or more input devices (e.g., as described above with respect to FIG. 6G). In some embodiments, in response to detecting the second interaction condition corresponding to the content, the computer system changes display of the representation of the face from looking in the direction of the content to looking in the direction of the first user detected in the field-of- detection of the one or more input devices. Displaying the representation of the face from the direction of the content to the direction of the user in response to detecting a new interaction condition enables the computer system to confirm that the second interaction condition has been detected, thereby providing improved visual feedback to the user and performing an operation when a set of conditions has been met without requiring further input.
[0186] In some embodiments, while (and/or after) displaying the representation of the face (e.g., 604) looking in the direction of the user detected in the field-of-detection of the one or more input devices, the computer system detects a third interaction condition (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) corresponding to the content (e.g., 614) (e.g., as described above with respect to FIGS. 6E-6F) (e.g., and/or corresponding to the other content). In some embodiments, in response to detecting the third interaction condition corresponding to the content (e.g., 614), the computer system displays, via the display component, the representation of the face (e.g., 604) looking in the direction of the content (e.g., 612) (e.g., as described above with respect to FIGS. 6E-6F). In some embodiments, the computer system changes displaying the representation of the face from looking in the direction of the first user detected in the field-of-detection of the one or more input devices to looking in the direction of the content. Displaying the representation of the face from the direction of the user to the direction of the content in response to detecting a new interaction condition enables the computer system to confirm that the new interaction condition is received and/or indicate that the user’s attention should be directed to the content, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further input.
[0187] In some embodiments, while (and/or after) displaying, via the display component, the content (e.g., 612) (and, in some embodiments, the representation of the face (e.g., looking in the direction of the content or the first user and/or another direction not corresponding to the first user and not corresponding to the content)), the computer system detects a fourth interaction condition (e.g., via a verbal input (e.g., an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) corresponding to the content (e.g., 614) (e.g., and/or corresponding to the other content). In some embodiments, in response to detecting the fourth interaction condition corresponding to the content (e.g., 614), the computer system displays, via the display component, the representation of the face (e.g., 604) looking in a third direction that does not correspond to the first user and the content (e.g., as described above with respect to FIG. 6G). In some embodiments, the computer system changes display of the representation of the face from looking in the direction of the first user detected in the field-of-detection of the one or more input devices to looking in the third direction. In some embodiments, the computer system changes display of the representation of the face from looking in the direction of the first user to looking in the third direction. In some embodiments, while displaying, via the display component, the representation of the face looking in the third direction, the representation of the face is not looking at the content and the representation of the face is not looking at the first user. In some embodiments, the third direction is in the field-of-view of the one or more input device but not in the field-of-detection of the first user or the second user. In some embodiments, the third direction is in the direction of the user interface but not in the direction of the content. Displaying the representation of the face in a direction not corresponding to the first user and the content in response to detecting the fourth interaction enables the computer system to confirm that a new interaction is received and/or confirm that the new interaction is not related to the content or the user (e.g., an accidental interaction, an unrecognized interaction) and/or the detection of a new user and/or object, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further input.
[0188] In some embodiments, while (and/or after) displaying, via the display component, the representation of the face (e.g., 604) looking in a third direction, the computer system detects a fifth interaction condition (e.g., speech, touch, context (e.g., what is being discussed, talked about, and/or output) and/or movement of the computer system and/or a user) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) corresponding to the content (e.g., as described above with respect to FIGS. 6D) (e.g., and/or corresponding to the other content). In some embodiments, in response to detecting the fifth interaction condition, the computer system continues displaying, via the display component, the representation of the face (e.g., 604) looking in the third direction (e.g., as described above with respect to FIGS. 6D-6E) (e.g., the same direction as when, before, and/or while detecting the fifth interaction condition). In some embodiments, in response to detecting the fifth interaction condition, the computer system does not change the direction that the representation of the face is looking in response to detecting the fifth interaction condition. In some embodiments, the fifth interaction condition is a different type of interaction condition than the first interaction condition. Continuing displaying the representation of the face looking in the fourth direction in response to detecting a fifth interaction corresponding to the content enables the computer system to confirm that the new interaction condition corresponds to direction the representation of the face is already facing (e.g., of the content and/or the user) and/or confirm that the new interaction condition was ignored, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further input.
[0189] In some embodiments, while (and/or after) displaying, via the display component, the content (e.g., 612), the computer system detects a sixth interaction condition (e.g., speech, touch, context (e.g., what is being discussed, talked about, and/or output) and/or movement of the computer system and/or a user) corresponding the content (e.g., and/or corresponding to the other content) (e.g., via a verbal input (e.g., an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) (e.g., as described above with respect to FIGS. 6D-6F). In some embodiments, in response to detecting the sixth interaction condition corresponding to the content, in accordance with a determination that the sixth interaction condition corresponding to the content satisfies the second set of one or more criteria, the computer system displays, via the display component, the representation of the face (e.g., 604) looking in the direction of the first user detected in the field-of-detection of the one or more input devices (e.g., as described above with respect to FIGS. 6F-6G). In some embodiments, in response to detecting the sixth interaction condition corresponding to the content, in accordance with a determination that the sixth interaction condition corresponding to the content satisfies a fourth set of one or more criteria different from the first set of one or more criteria and the second set of one or more criteria (e.g., and/or the third set of one or more criteria), the computer system displays, via the display component, the representation of the face (e.g., 604) looking in the direction of a third user detected in the field-of-detection of the one or more input devices, wherein the third user is different from the first user (e.g., as described above with respect to FIGS. 6F-6G). Displaying the representation of the face looking in the direction of the first user based on the sixth interaction condition including the second set of one or more criterion enables the computer system to confirm that the new interaction condition is received and determined to correspond to the first user and/or confirmation that the operation being performed is catered to the third user detected in the field-of-detection of the one or more input devices, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further input. Displaying the representation of the face looking in the direction of the third user based on the sixth interaction including a fourth set of one or more criterion enables the computer system to confirm that the new interaction condition is received and determined to correspond to the third user and/or confirmation that the operation being performed is catered to the third user detected in the field-of-detection of the one or more input devices, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further input. [0190] In some embodiments, the computer system (e.g., 600) is in communication with a movement component (e.g., an actuator (e.g., a pneumatic actuator, a hydraulic actuator and/or an electric actuator), a movable base, a rotatable component, and/or a rotatable base). In some embodiments, in response to detecting the sixth interaction condition corresponding to the content and in accordance with a determination that the sixth interaction condition satisfies the second set of one or more criteria, the computer system moves, via the movement component, a portion (e.g., a hardware component (e.g., a button and/or a rotatable input mechanism), a display, and/or a center and/or another portion of the display) of the computer system (e.g., 600) from a first position in the environment to a second position in the environment (e.g., as described above with respect to FIGS. 6D-6G).
[0191] In some embodiments, while the portion of the computer system (e.g., 600) is in the second position, the computer system detects a seventh interaction condition (e.g., and/or detecting that satisfies the second set of one or more criteria is no longer satisfied) corresponding to the content (e.g., as described above with respect to FIGS. 6D-6G). In some embodiments, in response to detecting the seventh interaction condition corresponding to the content and in accordance with a determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria (e.g., as described above with respect to FIGS. 6D-6G), the computer system moves, via the movement component, the portion of the computer system (e.g., 600) from the second position in the environment to the first position in the environment (e.g., as described above with respect to FIGS. 6D-6G). In some embodiments, in response to detecting the seventh interaction condition corresponding to the content and in accordance with the determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria, the computer system displays, via the display component, the representation of the face (e.g., 604) looking in the direction of the displayed content (e.g., 612) (e.g., as described above with respect to FIGS. 6D-6G). In some embodiments, in response to detecting the seventh interaction condition corresponding to the content and in accordance with a determination that the seventh interaction corresponding to the content satisfies the second set of one or more criteria, the computer system does not move the portion of the computer system from the second position in the environment to the first position in the environment and/or display the representation of the face looking in the direction of the displayed content. [0192] In some embodiments, while the portion of the computer system (e.g., 600) is in the second position, the computer system detects a eighth interaction condition (e.g., and/or detecting that satisfies the second set of one or more criteria is no longer satisfied) corresponding to the content (e.g., 610) (e.g., as described above with respect to FIG. 6D- 6G). In some embodiments, in response to detecting the eighth interaction condition (e.g., 610) corresponding to the content (e.g., as described above with respect to FIG. 6D-6G), in accordance with a determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria, the computer system moves, via the movement component, the portion of the computer system (e.g., 600) from the second position in the environment to the first position in the environment while continuing to display, via the display component, the representation of the face (e.g., 604) looking in the direction of the user (e.g., as described above with respect to FIGS. 6D-6G) (e.g., which is a different direction than looking at the location of the content). In some embodiments, in response to detecting the eighth interaction condition corresponding to the content and in accordance with a determination that the seventh interaction corresponding to the content satisfies the first set of one or more criteria, the computer system does not move the portion of the computer system from the second position in the environment to the first position in the environment and/or continue to display, via the display component, the representation of the face looking in the direction of the user.
[0193] Note that details of the processes described above with respect to process 700 (e.g., FIG. 7) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 800 optionally includes one or more of the characteristics of the various methods described above with reference to process 700. In some embodiments, the computer system can use one or more techniques of process 700 to determine an agreement was made with respect to an input and display the object in a manner that indicates eye contact with a user using one or more techniques of process 800. For brevity, these details are not repeated below.
[0194] FIG. 8 is a flow diagram illustrating a method for displaying an indication of eye contact of an object using a computer system in accordance with some embodiments. Process 800 is performed at a computer system (e.g., 100, 200, 600). Some operations in process 800 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted. [0195] As described below, process 800 provides an intuitive way for displaying an indication of eye contact of an object. The method reduces the cognitive burden on a user for displaying an indication of eye contact of an object, thereby creating a more efficient humanmachine interface. For battery operated computing devices, enabling a user to display an indication of eye contact of an object faster and more efficiently conserves power and increases the time between battery charges.
[0196] In some embodiments, process 800 is performed at a computer system (e.g., 600) that is in communication with a display component (e.g., a display screen, a projector, and/or a touch-sensitive display) and one or more input devices (e.g., a camera, a depth sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a fitness tracking device, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0197] While (e.g., after and/or in response to) displaying, via the display component, a user interface (e.g., 602) that includes a user interface object (e.g., a shape, an icon, an emoji, and/or an avatar) representing a portion (e.g., face, eyes, lips, hands, legs, and/or mouth) of a person (e.g., 604), the computer system detects (802), via the one or more input devices, a first input (e.g., a voice input, a tap input, a swipe input, and/or an air gesture) (e.g., as described above with respect to FIG. 6A). In some embodiments, the voice input may be received through a text to speech device within the computer system and/or via a second computer system.
[0198] In response to (804) detecting the first input, in accordance with a determination that an agreement was made with respect to the first input (e.g., 606) (e.g., agreement that a statement is (1) correct and/or (2) consistent with one or more user preferences) (e.g., agreement that a response to a question is (1) that the question is correct, (2) that a response to the question is positive (e.g., “Yes,” “I agree,” and/or “I like it”)), the computer system continues (806) displaying, via the display component, the user interface (e.g., 602) while changing the user interface object in a manner (e.g., a pose, a way, and/or a visual manner) that indicates eye contact with a user (e.g., a person, an animal, and/or an object) (e.g., as described above with respect to FIGS. 6A-6C). [0199] In response to (804) detecting the first input, in accordance with a determination that an agreement was not made with respect to the first input (e.g., 606) (e.g., disagreement to the statement being correct and/or disagreement that the question is factually correct/accurate), the computer system continues (808) displaying, via the display component, the user interface (e.g., 602) without changing the user interface object in the manner that indicates eye contact with the user (e.g., as described above with respect to FIGS 6B-6F). Displaying the user interface object in a manner that indicates eye contact with the user when an agreement is made with respect to the first input provides the user with feedback that the computer system detected the input and agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component. Displaying the user interface object without changing the user interface object in a manner that indicates eye contact with the user when an agreement is not made provides the user with feedback that the computer system does not agree with the detected input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0200] In some embodiments, continuing displaying the user interface (e.g., 602) without changing the user interface object in the manner that indicates eye contact with the user includes displaying, via the display component, the user interface object (e.g., 604) in a manner that does not indicate eye contact with (e.g., corresponding to and/or at) the user (e.g., as described above with respect to FIGS. 6B-6F). In some embodiments, the computer system changes a portion of the user interface object from looking in the direction of the user to another direction not directed at the user, in response to detecting the first input. Displaying the user interface object in a manner that does not indicate eye contact with the user when an agreement is not made provides the user with feedback that the computer system does not agree with the detected input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component. [0201] In some embodiments, changing the user interface object (e.g., 604) in the manner that indicates eye contact with the user includes: detecting, via the one or more input devices, a position of the user (e.g., a person, an animal, and/or an object) in a field-of-detection of the computer system (e.g., 600) (e.g., as described above with respect to FIGS. 6A-6C); and in some embodiments, the field-of-detection is a field-of-view of a camera, a field-of-detection of a depth sensor, and/or a field-of-detection of a microphone, changing a first portion (e.g., the eye portion, the face portion, the nose portion, and/or the mouth portion) of the user interface object (e.g., 604) to be directed to (e.g., in the direction of, looking at, and/or pointing at) the detected position of the user in the field-of-detection of the computer system (e.g., 600) (e.g., as described above with respect to FIGS. 6A-6G). Changing a first portion of the user interface object to be directed to the detected position of the user in the field of detection of the computer system when an agreement is made with respect to the first input provides the user with feedback that the computer system detected the input and agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0202] In some embodiments, changing the first portion of the user interface object (e.g., 604) to be directed to the detected position of the user in the field-of-detection includes moving, via the display component, the first portion of the user interface object from a position that is more than a predetermined distance (e.g., 0.1-10 meters) from the detected position of the user to a position that is no more than the predetermined distance from the detected position of the user in the field-of-detection (e.g., as described above with respect to FIGS. 6A-6G) (e.g., tilting the user interface object, moving the angle of at least a portion (e.g., the eyes, the mouth, the ear, and/or the nose) of the user interface object in a direction closer to the position of the user., changing a portion of the user interface object (e.g., the eyes, the mouth, the ear, and/or the nose) to be directed in the direction closer to the position of the user, and/or changing the eyes of the user interface object to look closer to the position of the user). Moving the first portion of the user interface object from a position that is more than a predetermined distance from the detected position of the user to a position that is no more than the predetermined distance from the detected position of the user in the field-of- detection of the computer system when an agreement is made with respect to the first input provides the user with feedback that the computer system detected the input and agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0203] In some embodiments, changing the first portion of the user interface object (e.g., 604) to be directed to the detected position of the user in the field-of-detection includes displaying, via the display component, the first portion of the user interface object pointing to (e.g., appearing to point to and/or protruding towards) the position (e.g., in the direction) of the user detected in the field-of-detection (e.g., as described above with respect to FIGS. 6A- 6G). Displaying the first portion of the user interface object pointing to the position of the user detected in the field-of-detection of the computer system when an agreement is made with respect to the first input provides the user with feedback that the computer system detected the input and agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0204] In some embodiments, the first input (e.g., 606) does not include an explicit indication to change the user interface object (e.g., the explicit indication is not an explicit request to move the user interface object) (e.g., the explicit indication is not an explicit request to change a portion of the user interface object) (e.g., the first input includes a question, a statement and/or any other voice input that corresponds the user interface object without an explicit request and/or indication) (e.g., as described above with respect to FIG. 6A).
[0205] In some embodiments, displaying the user interface (e.g., 602) while changing the user interface object (e.g., 604) in the manner that indicates eye contact with the user includes moving, via the display component, a second portion of the user interface object (e.g., as described above with respect to FIGS. 6E-6F) (e.g., moving a portion of the user interface object in the direction of the user when an agreement is made) (e.g., moving a portion of the user interface object in a direction away from the user when an agreement is not made) (e.g., moving a portion of the user interface from a first respective position to a second respective position different from the first respective position). Moving a second portion of the user interface object when an agreement is made with respect to the first input provides the user with feedback that the computer system detected the input and agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0206] In some embodiments, while (and/or after) displaying the user interface object (e.g., 604) in the manner that indicates eye contact with the user, the computer system detects, via the one or more input devices, a second input (e.g., 608) (e.g., a verbal input (e.g., an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) (e.g., as described above with respect to FIGS. 6A-6C). In some embodiments, in response to detecting the second input (e.g., 608), the computer system continues to display the user interface (e.g., 602) while changing, via the display component, the user interface object (e.g., 604) in a first manner (e.g., a pose, a way, and/or a visual manner) that does not indicate eye contact with the user (e.g., as described above with respect to FIGS. 6A-6C) (e.g., irrespective of the determination of agreement) (e.g., before the determination of an agreement is made) (e.g., immediately after input is detected). Changing the user interface object in the first manner that does not indicate eye contact with the user in response to detecting a second input allows the computer system to suggest a change whether disagreement and/or agreement is being made with respect to the second input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component. While user interface object is changed to indicate eye contact, detect another voice input; and in response, change or do not change user interface object to indicate eye contact based on whether there is agreement with the use.
[0207] In some embodiments, while (and/or after) displaying the user interface object (e.g., 604) in the manner that indicates eye contact with the user, the computer system detects, via the one or more input devices, a third input (e.g., 608) (e.g., as described above with respect to FIGS. 6A-6C) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)). In some embodiments, in response to detecting the third input, in accordance with a determination that an agreement was not made with respect to the third input (e.g., 608), the computer system continues displaying, via the display component, the user interface (e.g., 602) while changing the user interface object (e.g., 604) in a second manner that does not indicate eye contact with the user (e.g., as described above with respect to FIGS. 6A-6C). In some embodiments, an agreement being made includes the computer system verifying and/or validating that content of the voice input is true (e.g., factually true based on publicly available data and/or factually true based on private data (e.g., user data, data associated with a user account, and/or data associated with one or more groups and/or one or more groups that the user belongs to)), content of the voice input corresponds to and/or is in agreement with one or more characteristics (e.g., facts, personality characteristics, likes, dislikes, and/or favorites) associated with, related to, and/or of the user. In some embodiments, in response to detecting the third input, in accordance with a determination that the agreement was made with respect to the third input (e.g., 608), the computer system continues displaying, via the display component, the user interface (e.g., 602) without changing the user interface object (e.g., 604) in the second manner that does not indicate eye contact with the user (e.g., as described above with respect to FIGS. 6A-6C). Displaying the user interface object without and/or while changing the user interface object in the second manner that does not indicate eye contact with the user in accordance with a determination an agreement being made or not made with respect to the third input provides the user with feedback that the computer system detected the input and allow the computer system to provide feedback on whether or not it agrees with the input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0208] In some embodiments, while (and/or after) displaying the use interface object (e.g., 604) in the manner that indicates eye contact with the user, the computer system detects, via the one or more input devices, a fourth input (e.g., 608) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) (e.g., as described above with respect to FIG. 6C). In some embodiments, in response to detecting the fourth input (e.g., 608) and in accordance with a determination that a disagreement was made with respect to the fourth input (e.g., the user is wrong) (e.g., the user gave a false statement) (e.g., the user gave a negative answer and/or statement in response to a positive input (e.g., an answer and/or a statement) is needed), the computer system continues displaying, via the display component, the user interface object (e.g., 604) in the manner that indicates eye contact with the user (and/or performs an operation indicating the disagreement, such as representing shaking a head of the user interface object and/or outputting a negative statement (e.g., “No” and/or “I do not agree”). described above with respect to FIG. 6C). In some embodiments, an agreement being made includes the computer system verifying and/or validating that content of the voice input is true (e.g., factually true based on publicly available data and/or factually true based on private data (e.g., user data, data associated with a user account, and/or data associated with one or more groups and/or one or more groups that the user belongs to)), content of the voice input corresponds to and/or is in agreement with one or more characteristics (e.g., facts, personality characteristics, likes, dislikes, and/or favorites) associated with, related to, and/or of the user. In some embodiments, in response to detecting the fourth input and in accordance with a determination that a disagreement was not made with respect to the fourth input, the computer system does not continue displaying, via the display component, the user interface object in the manner that indicates eye contact with the user and/or displays the user interface object in a manner that does not indicate eye contact with the user. Continuing displaying the user interface object in the manner that indicates eye contact with the user when an agreement is not made with respect to the fourth input provides the user with a consistent and/or on-going interaction with the user interface object, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0209] In some embodiments, while (and/or after) displaying the user interface object (e.g., 604) in the manner that indicates eye contact with the user, the computer system detects, via the one or more input devices, a fifth input (e.g., as described above with respect to FIGS. 6A-6C). In some embodiments, in response to detecting the fifth input (e.g., 608) in accordance with a determination that a disagreement was made with respect to the fifth input, the computer system displays, via the display component, the user interface object in a third manner (e.g., 604) that does not indicate eye contact with the user (e.g., as described above with respect to FIGS. 6B). In some embodiments, in response to detecting the fifth input in accordance with a determination that a disagreement was made with respect to the fifth input, the computer system modifies display of the user interface object from the manner that indicates eye contact with the user to the third manner that does not indicate eye contact with the user. Displaying the user interface object while changing the user interface object in a third manner that does not indicate eye contact with the user when an agreement is not made with respect to the fifth input provides the user with feedback that the computer system does not agree with the detected input, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0210] In some embodiments, in response to detecting the first input (e.g., 606), the computer system continues to display, via the display component, a portion of the user interface (e.g., 602) that does not include the user interface object (e.g., 604) (e.g., changing a portion of the user interface object while maintaining the current content being displayed) (e.g., changing a portion of the user interface object while maintaining the current display of everything else being displayed on the user interface) (e.g., changing the user interface object from a manner that indicates eye contact to a manner that does not indicate eye contact includes continuing to display the portion of the user interface object) while changing the user interface object in the manner that indicates eye contact with the user (e.g., as described above with respect to FIGS. 6A-6G). Continuing to display a portion of the user interface that does not include the user interface object while changing the user interface object in a manner that indicated eye contact with the user enables the computer system to maintain the user interface while still providing feedback that the computer system agrees with the detected input and reduces visual distractions from displaying an entirely different user interface, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0211] In some embodiments, after changing the user interface object (e.g., 604) in the manner that indicates eye contact with the user (e.g., as described above with respect to FIGS. 6A-6C), in accordance with a determination that a predetermined period of time (e.g., 0.1-10 seconds) has not passed since the user interface object (e.g., 604) was changed in the manner that indicates eye contact with the user, the computer system continues to display, via the display component, the user interface object in the manner that indicates eye contact with the user. In some embodiments, after changing the user interface object in the manner that indicates eye contact with the user, in accordance with a determination that the predetermined period of time has passed since the user interface object (e.g., 604) was changed in the manner that indicates eye contact with the user, the computer system forgoes continuing to display, via the display component, the user interface object in the manner that indicates eye contact with the user (e.g., as described above with respect to FIGS. 6A-6C). In some embodiments, after the predetermined period of time passes, the computer system ceases to display the user interface object in the pose (e.g., returns to a previous pose) In some embodiments, after changing the user interface object in the manner that indicates eye contact with the user and in accordance with a determination that the predetermined period of time has passed since the user interface object was changed in the manner that indicates eye contact with the user, the computer system changes the user interface object from the manner that indicates eye contact with the user to a different manner that does not indicate eye contact with the user. Continuing to display the user interface object in a manner that indicates eye contact with the user in accordance with a determination that the predetermined period of time has not passed since user interface object was changed in a manner that indicates eye contact and/or forgoing continuing to display the user interface object in a manner that indicates eye contact with the user in accordance with a determination that the predetermined period of time has passed since the user interface object was changed in a manner that indicates eye contact enables the first computer system to automatically change the display of the user interface object based on the recency of the detected input determined to be in agreement by the computer system, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0212] In some embodiments, the first input (e.g., 606) includes a question (e.g., as described above with respect to FIG. 6A). In some embodiments, the determination that the agreement was made with respect to the first input (e.g., 606) includes a determination that the question corresponds to a positive response (e.g., as described above with respect to FIGS. 6A-6C) (e.g., the answer and/or response is a confirmation, an approval, and/or an acceptance to the question). In some embodiments, the determination that the agreement was not made with respect to the first input (e.g., 606) includes a determination that the question corresponds to a negative response (e.g., as described above with respect to FIGS. 6A-6C) (e.g., the answer and/or response is a disapproval, a non-acceptance, and/or a nonconfirmation of the question). Displaying the user interface object in a manner that indicates eye contact with the user when a determination is made that a first input is a question that has a positive response provides the user with feedback that the computer system detected the input and that the computer system determines the question is correct, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component. Displaying the user interface object without changing the user interface object in the manner that indicates eye contact with the user when a determination is made that a first input that is a question that has a negative response provides the user with feedback that the computer system detected the input and that the computer system determines the question is wrong, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0213] In some embodiments, the first input (e.g., 606) includes a statement (e.g., as described above with respect to FIG. 6A). In some embodiments, the determination that the agreement was made with respect to the first input (e.g., 606) includes a determination that the statement corresponds to a factual statement (e.g., a statement that is factually correct (e.g., based on data, such as public, private, and/or bespoke data)) (e.g., as described above with respect to FIGS. 6A-6C). In some embodiments, the determination that the agreement was not made with respect to the first input (e.g., 606) includes a determination that the statement does not correspond to the factual statement (e.g., as described above with respect to FIGS. 6A-6C) (e.g., the answer and/or response is a disapproval, non-acceptance, and/or non-confirmation of the question). Displaying the user interface object in a manner that indicates eye contact with the user when a determination is made that a first input is a factual statement provides the user with feedback that the computer system detected the input and that the computer system determines the statement is factual, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component. Displaying the user interface object without changing the user interface object in a manner that indicates eye contact with the user when a determination is made that a first input is not a factual statement provides the user with feedback that the computer system detected the input and that the computer system determines the statement is not factual, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0214] In some embodiments, the determination that the agreement was made with respect to the first input (e.g., 606) includes a determination that the first input satisfies a user preference (e.g., a user preference, a known preference, a preference learned over time, a historical preference, and/or a preference based on a pattern of use (e.g., how the user interacts with the computer system and/or how the user interacts with one or more applications and/or external devices (e.g., fitness tracking devices, smart devices (e.g., smart lights, smart locks, and/or smart blinds)), and/or speakers)) of the first user (e.g., user likes blue, eye contact made with the user when color blue is within the phrase (e.g., picked and/or chosen)). In some embodiments, the determination that the agreement was not made with respect to the first input includes a determination that the first input does not satisfy the user preference of the first user (e.g., as described above with respect to FIGS. 6A-6C) (e.g., user likes blue, user interface object does not make eye contact with user when red is within the phrase (e.g., picked and/or chosen)).
[0215] In some embodiments, in response to detecting the first input (e.g., 606) via the one or more input devices (e.g., as described above with respect to FIGS. 6A-6C), the computer system provides, via one or more output devices (e.g., speakers, haptic output devices, and/or audio generation components) in communication with the computer system (e.g., 600), an output (e.g., an output different from changing a user interface object, an audio output, and/or a haptic output) corresponding to the first input (e.g., 610), wherein: in accordance with a determination that an agreement was made with respect to the first input (e.g., 610), the output includes a first phrase (e.g., “Yes,” “I agree,” and/or “You are right”) that indicates agreement with a user preference (e.g., and the display component continues displaying the user interface while changing the user interface object in a manner that indicates eye contact with the user) (e.g., the user likes blue, the user interface object makes eye contact with user when describing a blue colored wall) (e.g., as described above with respect to FIGS. 6A-6C; and in accordance with a determination that an agreement was not made with respect to the first input (e.g., 610), the output includes a second phrase (e.g., “No,” “I disagree,” and/or “I don’t think that’s right”) that indicates disagreement with the user preference (e.g., as described above with respect to FIGS. 6A-6C) (e.g., and the display component continues displaying the user interface without changing the user interface object in a manner that does not indicate eye contact with the user) (e.g., the user likes blue, the user interface object does not make eye contact with the user when describing a yellow colored wall). Outputting a first phrase and/or outputting a second phrase with a determination that an agreement when prescribed conditions are met allows the computer system to cater the operation to the user based on a determination of the user’s preferences and provide audible feedback to confirm the determination of the first input (e.g., as agreeing and/or disagreeing to a user preference and/or as being a saved user preference or not), thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0216] In some embodiments, continuing displaying, via the display component, the user interface (e.g., 602) without changing the user interface object (e.g., 604) in the manner that indicates eye contact with the user does not include outputting the second phrase (e.g., as described above with respect to FIGS. 6A-6C).
[0217] Note that details of the processes described above with respect to process 800 (e.g., FIG. 8) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 700 optionally includes one or more of the characteristics of the various methods described above with reference to process 800. In some embodiments, the computer system can use one or more techniques of process 800 to determine an interaction condition was satisfied and display the object facing in the direction of content using one or more techniques of process 700. For brevity, these details are not repeated below.
[0218] FIG. 9 is a flow diagram illustrating a method for de-emphasizing an object using a computer system in accordance with some embodiments. Process 900 is performed at a computer system (e.g., 100, 200, and/or 600). Some operations in process 900 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted. [0219] As described below, process 900 provides an intuitive way for de-emphasizing an object. The method reduces the cognitive burden on a user for de-emphasizing an object, thereby creating a more efficient human -machine interface. For battery operated computing devices, enabling a user to de-emphasize an object faster and more efficiently conserve power and increases the time between battery charges.
[0220] In some embodiments, process 900 is performed at a computer system (e.g., 600) that is in communication with a display component (e.g., a display screen, a projector, and/or a touch-sensitive display), a camera (e.g., one or more wide-angle cameras, telephoto cameras, and/or ultra-wide-angle cameras), and one or more input devices (e.g., a camera, a depth sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a fitness tracking device, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0221] While (and/or after and/or in response to) detecting a first entity (e.g., a user, an animal, and/or an object) in the field-of-view of the camera, the computer system displays (902) a user interface object (e.g., 604) representing a portion of a user (e.g., a person, an animal, and/or an object) in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system (e.g., 600) (e.g., field-of- detection of a microphone, field-of-view of a camera, and/or a field-of-detection of a depth sensor), wherein the user interface object is displayed at a first size (e.g., as described above with respect to FIG. 6D).
[0222] While displaying a user interface object indicating the portion of the user (e.g., 604), the computer system detects (904) a request to interact with content (e.g., 610) (e.g., voice input that includes an indication of the content, one or more inputs made by a portion of an entity’s body, and/or an air gesture) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)).
[0223] In response to (906) detecting the request (e.g., 610) to interact with the content, the computer system displays (908) at least a first portion of the content (e.g., 612) (e.g., as described above with respect to FIG. 6F). [0224] In response to (906) detecting the request to interact with the content, the computer system displays (910) the user interface object representing the portion of the user (e.g., 604) at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content (e.g., 612) (e.g., as described above with respect to FIG.S 6E-6F) (and is not directed to the first entity in the field-of- detection of the camera). Detecting the first entity and displaying a user interface object representing a portion of a user in a first manner directed to the first user in the field of detection of the computer system allows the computer system to confirm the detection of the first user, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component. Displaying at least the first portion of the content in response to detecting the request to interact with the content enables the computer system to provide interactive content to the user, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component. Displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content in response to detecting the request to interact with the content reduces visual pollution while adding the content to the user interface and/or provides the user with an indication that content is displayed, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0225] In some embodiments, before detecting the request to interact with the content (e.g., 610), displaying a portion of the user interface object representing the portion of the user (e.g., 604) at a first location (e.g., of the display component and/or of a user interface) (e.g., as described above with respect to FIG. 6D). In some embodiments, displaying the user interface object representing the portion of the user (e.g., 604) at the second size smaller than the first size in response to detecting the request (e.g., 610) includes ceasing to display the portion of the user interface object at the first location (e.g., as described above with respect to FIGS. 6D-6F). In some embodiments, the first portion of the content is displayed at the first location in response to detecting the request to interact with the content (e.g., 610) (e.g., as described above with respect to FIG. 6F). In some embodiments, in response to detecting the request to interact with the content, the computer system replaces display of the portion of the user interface object at the first location with display of the first portion of the content at the first location. Ceasing displaying the user interface object representing the portion of the user at the first location and displaying the portion of the content at the first location in response to detecting the request to interact with the content reduces visual pollution while adding the content to the user interface and/or without displaying an entirely different user interface, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0226] In some embodiments, before detecting the request to interact with the content (e.g., 610), the computer system forgoes displaying at least the portion of the content (e.g., as described above with respect to FIG. 6D) (e.g., while displaying the user interface object representing the portion of the user in the first manner and the first size). In some embodiments, at least the first portion of the content is not displayed before the request to interact with the content is detected. Not displaying at least the portion of the content before detecting the request to interact with the content reduces visual pollution of the user interface, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0227] In some embodiments, the one or more input devices includes a microphone. In some embodiments, detecting the request to interact with the content (e.g., 610) includes receiving, via the microphone, a voice input (e.g., a voice command, a voice request, an audible question and/or an audible statement) (e.g., as described above in relation to process 700 and/or as described above in relation to process 800) corresponding to the content (e.g., as described above with respect to FIG. 6D) (e.g., a voice input to initiates a request that involves an interaction with the content). In some embodiments, the voice input may be received through a text to speech device within the computer system and/or via a second computer system. [0228] In some embodiments, the request to interact with the content does not include an explicit request (e.g., 610) (e.g., as described above with respect to FIG. 6D) (e.g., to interact with content) (e.g., explicit request does not include the need to verify information (text, contact, etc.)) (e.g., explicit request includes a request to open content with a need to interact with content, such as “open content” and/or “display content”) (e.g., the request does not include a command to interact with content (e.g., “open content,” “move content,” and/or “show content”)).
[0229] In some embodiments, the request to interact with content (e.g., 610) is a first request to interact with the first portion of the content (e.g., as described above with respect to FIG. 6D). In some embodiments, while (and/or after) displaying at least the first portion of the content and the user interface object representing the portion of the user (e.g., 604) at the second size, the computer system detects, via the one or more input devices, a request (e.g., a swipe, tap, press and hold and/or a voice input) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) to interact with a second portion of the content different from the first portion of the content (e.g., 612) (e.g., the same content but a different portion and/or new content) (e.g., as described above with respect to FIG. 6F). In some embodiments, based on determination via the computer system that an interaction is needed that is on the second portion of the content and/or a user input to interact with the second portion of the content not shown in the first portion of the content. In some embodiments, in response to detecting the request to interact with the second portion of content, the computer system displays, via the display component, at least the second portion of the content, wherein the second portion of the content replaces the first portion of the content (e.g., as described above with respect to FIG. 6F). In some embodiments, in response to the detecting the request to interact with the second portion of content, the computer system ceases displaying, via the display component, at least the first portion of the content. In some embodiments, in response to detecting the request to interact with the second portion of content, the computer system continues to display, via the display component, the user interface object representing the portion of the user (e.g., 604) at the second size and in the second manner (e.g., as described above with respect to FIG. 6F). Displaying at least the second portion of the content, where the second portion of the content replaces the first portion of the content when a request to interact with the second portion of the content is detected allows the computer system to transition to new content without displaying an entirely different user interface, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0230] In some embodiments, after displaying the user interface object representing the portion of the user (e.g., 604) at the second size and in the second manner to indicate that the user interface object is directed to the portion of the content (e.g., 612) and in accordance with a determination that a first set of one or more criteria is satisfied (e.g., agreement criteria (e.g., as described above in relation to process 800) and/or important criteria (e.g., whether or not the user interface object and/or the request is important), waiting on a user input, and/or outputting the second content is completed) (e.g., with respect to the request), the computer system displays, via the display component, the user interface object representing the portion of the user at the second size and in a third manner (e.g., the first manner and/or not the second manner) to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system (e.g., 600) (e.g., as described above with respect to FIGS. 6F-6G). In some embodiments, after displaying the user interface object representing the portion of the person at the second size and in the second manner and in accordance with a determination that the first set of one or more criteria is satisfied, the computer system ceases to display in the second manner to indicate that the user interface object is directed to the portion of the content. In some embodiments, the portion of content is the first portion of content. Displaying the user interface object representing the portion of the user at the second size and in a third manner (e.g., the first manner and/or not the second manner) to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system based on first set of one or more criteria being met allows the computer system to confirm to the user that one or more criterion is satisfied (e.g., agreement criteria and/or important criteria), thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bumin of the display component.
[0231] In some embodiments, after displaying the user interface object representing the portion of the user (e.g., 604) at the second size and in the second manner to indicate that the user interface object is directed to the portion of content (e.g., 612) and in accordance to the determination that the first set of one or more criteria is not satisfied (e.g., no agreement is made, a request for interaction is detected, and/or entity is still interacting with content), the computer system continues displaying, via the display component, the user interface object representing the user at the second size and in the second manner (e.g., the second manner and/or not the first manner) to indicate that the user interface object is directed to the portion of the content (e.g., as described above with respect to FIG. 6F). Displaying the user interface object representing the portion of the user at the second size and in the second manner to indicate that the user interface object is directed to the portion of the content based on first set of one or more criteria not being met allows the computer system to confirm to the user that one or more criterion is not satisfied, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0232] In some embodiments, displaying the user interface object representing the portion of the user (e.g., 604) at the second size in the second manner to indicate that the user interface object is directed to the portion of content (e.g., 612) includes changing the user interface object representing the portion of the user to the second size before displaying the user interface object representing the portion of the user in the second manner to indicate that the user interface object is directed to the portion of content (e.g., as described above with respect to FIGS. 6E-6F). Changing the user interface object representing the portion of the user to the second size before displaying the user interface object representing the portion of the user in the second manner to indicate that the user interface object is directed to the portion of content allows the computer system to transition smoothly to the addition of the content without displaying an entirely different user interface, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid bum-in of the display component.
[0233] In some embodiments, displaying, via the display component, the user interface object representing the portion of the user (e.g., 604) in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection includes displaying, via the display component, a representation of eyes displayed with a first set of one or more characteristics (e.g., in a first position, in a first shape, and/or in a first orientation and/or distance apart) (e.g., as described above with respect to FIG. 6D-6G). In some embodiments, displaying, via the display component, the user interface object representing the portion of the user (e.g., 604) in the second manner to indicate that the user interface object is directed to the portion of content (e.g., 612) includes displaying, via the display component, the representation of the eyes with a second set of characteristics (e.g., in a first position, in a first shape, and/or in a first orientation and/or distance apart) different from the first set of one or more characteristics (e.g., as described above with respect to FIG. 6F). Displaying a representation of eyes displayed with a first set of one or more characteristics when displaying the user interface object representing the portion of the user in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection and/or the representation of the eyes with a second set of characteristics when displaying the user interface object representing the portion of the user in the second manner to indicate that the user interface object is directed to the portion of content allows the computer system to provide an indication of where the user interface object representing the portion of the user is being directed, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0234] In some embodiments, after displaying at least the first portion of the content (e.g., 612) and in accordance with a determination that content should be removed (e.g., a request to close content, no more content can be displayed (e.g., outputting content is completed) and/or the computer system is waiting on an input) (and, in some embodiments, while displaying the user interface object representing the portion of the user at the second size) (e.g., as described above with respect to FIGS. 6F-6G), the computer system ceases displaying at least the first portion of the content (e.g., 604) (e.g., as described above with respect to FIGS. 6D-6E). In some embodiments, after displaying at least the first portion of the content and in accordance with the determination that content should be removed, the computer system re-displays, via the display component, the user interface object representing the portion of the user (e.g., 604) at the first size (and, in some embodiments, in the first manner) (e.g., as described above with respect to FIGS. 6D-6E). Ceasing displaying at least the first portion of the content re-displaying the user interface object representing the portion of the user at the first size when a determination is made that the content should be removed allows the computer system to use the available space in the user interface properly, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0235] In some embodiments, after displaying at least the first portion of the content (e.g., 612) and in accordance with a determination that content should be removed and a third set of one or more criteria is satisfied (e.g., agreement criteria (e.g., as described above in relation to process 800) and/or important criteria (e.g., whether or not the user interface object and/or the request is important), waiting on a user input, and/or outputting the second content is completed) (e.g., with respect to the request), the user interface object (e.g., 604) is displayed in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection of the computer system (e.g., 600) (e.g., as described above with respect to FIGS. 6D-6G). Displaying the user interface object in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection of the computer system when a determination is made that content should be removed and a third set of one or more criteria is satisfied allows the user interface object to confirm that the third set of one or more criteria has been satisfied, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0236] In some embodiments, after displaying at least the first portion of the content (e.g., 612) and in accordance with a determination that content should be removed and a fourth set of one or more criteria is satisfied (e.g., agreement criteria (e.g., as described above in relation to process 800) and/or importunate criteria (e.g., whether or not the user interface object and/or the request is important), waiting on a user input, and/or outputting the second content is completed) (e.g., with respect to the request) (e.g., different from the third set of one or more criteria), the user interface object (e.g., 604) is displayed a fourth manner to indicate that the user interface object is not directed to the first entity(e.g., as described above with respect to FIGS. 6D-6G) (e.g., the fourth manner includes displaying the user interface object to be directed is in the field-of-detection of the one or more input devices but not the position where the first user is detected) (e.g., the fourth manner includes displaying the user interface object to be directed is in the direction of the first user interface but not directed at the content). Displaying the user interface object in a fourth manner to indicate that the user interface object is not directed to the first entity when a determination is made that content should be removed and a third set of one or more criteria is not satisfied allows the user interface object to confirm that the third set of one or more criteria has not been satisfied, thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0237] In some embodiments, while displaying the user interface object representing the portion of the user (e.g., 604) in the second manner, the computer system detects an interaction condition (e.g., as described above in relation to process 800) corresponding to the content (e.g., entity makes a new request to interact with new content, entity makes an explicit request to stop looking at content, and/or entity makes a request to close content, a second entity is detected in the field-of-detection) (e.g., an interaction condition as described above in relation to process 800) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input (e.g., a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click)) (e.g., as described above with respect to FIG. 6F). In some embodiments, in response to detecting the interaction condition, in accordance with a determination that the interaction condition satisfies a fifth set of one or more criteria (e.g., the first entity is interacting with content and/or the first entity makes a new request), the computer system displays, via the display component, the user interface object representing the portion of the user (e.g., 604) in the first manner to indicate that the user interface object is directed to (e.g., in the direction corresponding to and/or at) the first entity in the field-of-view of the camera, wherein the first entity initiated the interaction condition (e.g., as described above with respect to FIGS. 6D- 6G). In some embodiments, in response to detecting the interaction condition, in accordance with a determination that the interaction condition corresponding to the content satisfies a sixth set of one or more criteria different from the fifth set of one or more criteria (e.g., without satisfying the fifth set of one or more criteria) (e.g., an entity other than the first entity is interacting with content and/or the second entity makes a new request), the computer system displays, via the display component, the user interface object representing the portion of the user (e.g., 604) in a manner to indicate that the user interface object is directed to (e.g., in the direction corresponding to and/or at) a second entity detected in the field-of-view of the computer system (e.g., 600), wherein the second entity initiated the interaction condition (e.g., as described above with respect to FIGS. 6D-6G). Displaying the user interface object representing the portion of the user in the first manner to indicate that user interface object is directed to the first entity in the field-of-view of the camera when detecting the first entity initiated the interaction condition and/or displaying the user interface object representing the portion of the user in a manner to indicate that the user interface object is directed to a second entity detected in the field-of-view of the computer system when the second entity initiated the interaction condition allows the computer system to confirm the user who initiated the request to interact with content), thereby providing improved feedback, reducing the number of inputs needed to perform an operation, performing an operation when a set of conditions has been met without requiring further input, and/or allowing the computer system to avoid burn-in of the display component.
[0238] Note that details of the processes described above with respect to process 900 (e.g., FIG. 9) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 700 optionally includes one or more of the characteristics of the various methods described above with reference to process 900. In some embodiments, the computer system can use one or more techniques of process 900 to determine an interaction condition was satisfied to display the object facing in the direction of content using one or more techniques of process 700. For brevity, these details are not repeated below.
[0239] FIGS. 10A-10F illustrate exemplary user interface for providing content in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes in FIGS. 11-14.
[0240] FIGS. 10A-10F illustrate computer system 1000 as a tablet. It should be recognized that computer system 1000 can be other types of computer systems, such as a smart phone, a smart watch, a laptop, a movable device, a rotatable device, a smart display device, a communal device, a smart speaker, an accessory, a personal gaming system, a desktop computer, a fitness tracking device, and/or a head-mounted display (HMD) device. In some embodiments, computer system 1000 includes and/or is in communication with one or more sensors (e.g., camera sensor, lidar detector, motion sensor, infrared sensor, and/or microphone). In some embodiments, computer system 1000 includes and/or is in communication with one or more output devices (e.g., a display screen, a projector, a touch- sensitive display, a haptic output device, and/or a speaker). In some embodiments, computer system 1000 includes and/or is in communication with one or more movement components (e.g., an actuator, a moveable base, a rotatable component, and/or a rotatable base). In some embodiments, the one or more movement components are configured to move one or more physical portions (e.g., of computer system 1000). In some embodiments, computer system 1000 includes one or more components and/or features described above in relation to computer system 100 and/or device 200.
[0241] As illustrated in FIGS. 10A-10F, computer system 1000 displays user interface 1002. In some embodiments, user interface 1002 is a home screen user interface that includes one or more application controls and/or a representation of a virtual assistant (e.g., 1004, 1014, 1022, and/or 1026). In some embodiments, user interface 1002 is an educational application user interface. In some embodiments, user interface 1002 can be an application used to teach users how to read. In some embodiments, user interface 1002 is an office assistant user interface. In some embodiments, user interface 1002 can be an application user interface that provides updates to the user of the latest in-office reports. In some embodiments, user interface 1002 is an entertainment application user interface. In some embodiments, user interface 1002 can be an electronic book application user interface used to read content to a user.
[0242] As illustrated in FIGS. 10A-10F, user interface 1002 includes different system avatars, including user interface object 1004 (e.g., in FIGS. 10A-10B), second system avatar 1014 (e.g., in FIGS. 10C and 1 OF), third system avatar 1022 (e.g., in FIG. 10D), and fourth system avatar 1026 (e.g., in FIG. 10E), at different times. In the examples in FIGS. 10A-10F, the different system avatars are anthropomorphic visual representations of an artificial intelligence application and/or a virtual assistant that visually changes based on content output and/or input and reacts and/or responds to input from a user in the vicinity of an input device of computer system 1000. In some embodiments, a system avatar is an avatar that corresponds to (e.g., managed, controlled, output, created, and/or requested by) a system process of a computer system (e.g., 1000). In some embodiments, a system avatar can be displayed while other applications run in the background. In some embodiments, a system process is a process corresponding to an operating system of the computer system. In some embodiments, an avatar corresponds to an application process, which corresponds to a software application installed and/or executed by the computer system. In some embodiments, a system process is different from an application process. In some embodiments, computer system 1000 has access to and/or outputs content from one or more system processes and/or one or more application processes in a manner that appears as if one or more avatars (e.g., user interface object 1004, second system avatar 1014, third system avatar 1022, and/or fourth system avatar 1026) is outputting the content. In some embodiments, an application corresponding to a system avatar can have access to a teaching application and an electronic book application, allowing the system avatar to appear to teach using electronic books. For another example, an application corresponding to a system avatar can have access to a news application and an internet application, allowing the system avatar to appear to answer user questions about information in news articles.
[0243] In some embodiments, computer system 1000 displays a system avatar with a visual appearance and/or facial expression (e.g., happy, sad, confused, scared, shocked, and/or neutral) that corresponds to content being output by computer system 1000. In some embodiments, a system avatar can have the appearance of a dragon if the content being output by computer system 1000 includes a dragon. For another example, a system avatar can have the appearance of laughing if the content being output by computer system 1000 is intended to be funny. In some embodiments, computer system 1000 displays different system avatars as having different appearances (e.g., different colors (e.g., sets of colors, flesh tones, reds, oranges, yellows, greens, blues, and/or purples), textures (e.g., skin, hair, fur, scales, plastic, glass, feathers, and/or wood), accessories (e.g., hat, glasses, monocle, wand, book, collar, bow, wings, halo, and/or crown), and/or face types (e.g., human, animal, anthropomorphized object, alien, non-descript face, fantasy creature, and/or a collection of objects that resemble a face)). In some embodiments, computer system 1000 outputs audio through one or more output devices (e.g., speakers) included in and/or in communication with computer system 1000 that corresponds to a voice for a system avatar. In some embodiments, computer system 1000 can output audio of a squeaky voice if the displayed system avatar has the appearance of a mouse. For another example, computer system 1000 can output audio of a booming voice if the displayed system avatar has the appearance of a giant. In some embodiments, computer system 1000 outputs audio with different voices (e.g., male, female, high/low pitch, with/without an accent, soft, and/or loud) to correspond to different system avatars, depending on which system avatar and/or what appearance a system avatar has that is displayed. [0244] In some embodiments, computer system 1000 displays a system avatar with a set of one or more movement patterns via the display. In some embodiments, a system avatar can be displayed as having bouncy movements (e.g., a system avatar can bounce up and down when content corresponding to the system avatar is output by computer system 1000) if the system avatar has the appearance of a rabbit. For another example, a system avatar can be displayed as shaking if content being output by computer system 1000 is intended to by scary. In some embodiments, computer system 1000 displays different system avatars with different sets of one or more movement patterns via the display (e.g., bouncing, wiggling, zooming, and/or dancing). In some embodiments, a system avatar corresponds to a set of one or more movement patterns of computer system 1000 (via one or more movement components). In some embodiments, a portion of computer system 1000 can physically move side to side, giving the impression of a system avatar slithering if the system avatar has the appearance of a snake. For another example, a portion of computer system 1000 can move up and down, giving the impression of a system avatar nodding if the system avatar is responding in the affirmative to input from the user. In some embodiments, different system avatars correspond to different movement patterns of a portion of computer system 1000 (e.g., swaying, bowing, rotating, and/or tilting).
[0245] In the example illustrated in FIGS. 10A-10F, there are four system avatars illustrated. In some embodiments, there are more or less than four system avatars. At FIGS. 10A-10F, only one system avatar is displayed at a time. In some embodiments, multiple system avatars are displayed at a time. In some embodiments, computer system 1000 can display multiple system avatars if the content being output by computer system 1000 includes a conversation between multiple characters, allowing for each system avatar to appear to perform a different part of the conversation and appear to interact together.
[0246] The right side of FIGS. 10A-10F include diagram 1006. Diagram 1006 is a visual aid that is a representation of a physical environment that includes computer system 1000. Diagram 1006 includes computer system representation 1008, field of view 1008a, and user representation 1010. Field of view 1008a represents the field of view of one or more front facing sensors of computer system 1000. The positioning of user representation 1010 and computer system representation 1008 within diagram 1006 is representative of real -world positioning of the user with respect to computer system 1000. [0247] As illustrated in FIG. 10A, computer system 1000 displays user interface 1002 with user interface object 1004. In some embodiments, user interface object 1004 is a default system avatar of computer system 1000. In some embodiments, the default system avatar is an avatar that is preset (e.g., previously configured) by a publisher of user interface 1002 (e.g., or otherwise selected by default, rather than based on user input and/or content being output). In some embodiments, the default system avatar is an avatar that is preset by the user. In this example, computer system 1000 displays the default system avatar (e.g., user interface object 1004) when content output by computer system 1000 does not correspond to a specific and/or different system avatar. In some embodiments, movement by user interface object 1004 is synchronized with content output by computer system 1000 while user interface object 1004 is displayed. In some embodiments, computer system 1000 can synchronize visual movements of facial features of user interface object 1004 with audio output to give the impression that user interface object 1004 is talking to the user.
[0248] As illustrated in FIG. 10A, computer system 1000 displays user interface object 1004 at a center location of user interface 1002, taking up a majority of user interface 1002. In other examples, user interface object 1004 can be displayed in other locations and/or take up more or less of user interface 1002. In this example, computer system 1000 displays user interface object 1004 as a representation of a human face. In some embodiments, user interface object 1004 is displayed as a representation of a different human face. In some embodiments, user interface object 1004 is displayed as a representation of a face that is not human (e.g., an animal, an anthropomorphized object, an alien, a non-descript face, fantasy creature, and/or a collection of objects that resemble a face). It should be recognized that other representations of first system user interface object 1004 can be used with techniques described herein and that a representation of a face is one example.
[0249] At FIG. 10 A, as indicated by the positioning of user representation 1010 within field of view 1008a, a user is within the field of view of computer system 1000. At FIG. 10A, computer system 1000 detects the user. In some embodiments, detecting the user causes computer system 1000 to display user interface object 1004 as having a different facial expression, such as going from a default or neutral expression to a smiling expression. At FIG. 10 A, the user moves to a new location within the field of view of computer system 1000. In some embodiments, computer system 1000 moves a portion (e.g., a display component and/or other component) of computer system 1000 to keep the user within the field of view. In some embodiments, as the user physically moves through the space, the portion of computer system 1000 can rotate via one or more movement components to keep the user within the field of view. At FIG. 10 A, computer system 1000 detects audio input 1012. Audio input 1012 corresponds to a voice input request from the user to have computer system 1000 output requested content (e.g., Let’s read “the dog”).
[0250] As illustrated in FIG. 10B, in response to detecting audio input 1012, computer system 1000 provides output of a first portion of the requested content (e.g., “the dog”). In this example, the requested content includes audio output (e.g., reading “the dog”) and visual output (e.g., displaying text of “the dog”). In some embodiments, as part of providing the audio output of a first portion of the requested content, computer system 1000 displays first text 1016, which corresponds to the first portion of the requested content. As illustrated in FIG. 10B, in response to detecting audio input 1012, computer system 1000 decreases the size of user interface object 1004 and moves user interface object 1004 down and to the left to make room for first text 1016. In some embodiments, computer system 1000 does not alter the size and/or location of user interface object 1004. In some embodiments, computer system 1000 does not display first text 1016. In some embodiments, computer system 1000 does not output audio in the form of reading the requested content. In some embodiments, computer system 1000 outputs audio content that enhances the requested content that is not the requested content. In some embodiments, computer system 1000 can output the sounds of waves and/or seagulls if the requested content corresponds to a boat. In some embodiments, the portion of computer system 1000 physically moves in sync with the audio content. In some embodiments, the portion of computer system 1000 can move in a way that imitates the motion of a boat in sync with the sound of the waves. At FIG. 10B, as indicated by user representation 1010 being at a different location within field of view 1008a than user representation 1010 within field of view 1008a in FIG. 10A, the user is at a different location within the field of view of computer system 1000. At FIG. 10B, even though the user has moved, the user is still within the field of view of computer system 1000, causing computer system 1000 to detect the user within the field of view. As described below, if computer system 1000 is already outputting content, in response to continuing to detect the user within the field of view, computer system 1000 continues output of content.
[0251] In some embodiments, if computer system 1000 is not facing the user (e.g., the user is not within the field of view), in response to detecting audio input 1012, computer system 1000 moves to face the user (e.g., moves the portion until the user is within the field of view). In some embodiments, if computer system 1000 is moving to face the user in response to detecting audio input 1012, computer system 1000 outputs a first portion of the requested content before ceasing movement (e.g., in response to detecting the user within the field of view). In some embodiments, in response to detecting audio input 1012, computer system 1000 can start outputting “the dog” as it is rotating to face the user. In some embodiments, if computer system 1000 is moving to face the user in response to detecting audio input 1012, computer system 1000 outputs a first portion of the requested content after ceasing movement (e.g., in response to detecting the user within the field of view). In some embodiments, in response to detecting audio input 1012, computer system 1000 can rotate to face the user, cease movement, and then begin outputting “the dog”. In some embodiments, if computer system 1000 is moving (e.g., moving to face the user, and/or moving as part of outputting content), in response to detecting the user within a predefined distance to computer system 1000, computer system 1000 ceases movement. In some embodiments, if computer system 1000 is moving in sync to music content, in response to detecting the user within a predefined distance (e.g., user approaches and/or user reaches for computer system 1000), computer system 1000 ceases movement.
[0252] As illustrated in FIG. 10C, in response to continuing to detect the user within the field of view, after providing output of the first portion of the requested content, computer system 1000 continues to provide output of a second portion of the requested content that immediately follows the first portion (e.g., if no other input is provided). As illustrated in FIG. 10C, as part of providing output of the second portion of the requested content, computer system 1000 displays second text 1018 that corresponds to the second portion of the requested content. In some embodiments, if computer system 1000 is moving the portion of computer system 1000 (e.g., as part of outputting content), in response to outputting the second portion of the requested content, computer system 1000 ceases movement to provide the user a better view of the displayed content after changing displayed content. In some embodiments, at a predetermined time after ceasing movement in response to outputting the second portion of the requested content, computer system 1000 physically moves the portion of computer system 1000 as part of outputting the second portion of the requested content. In some embodiments, in response to outputting the second portion of the requested content, computer system 1000 can pause movement for a period of time to allow the user time to read the newly displayed text. In some embodiments, if computer system 1000 is moving the portion of computer system 1000 (e.g., as part of outputting content), computer system 1000 displays an item (e.g., an alert, a notification, and/or a call) unrelated to the requested content, causing computer system 1000 to cease movement to allow the user the ability to notice and/or interact with the item. In some embodiments, computer system 1000 can pause movement in response to displaying a calendar alert to so the user can view the calendar alert.
[0253] At FIG. 10C, computer system 1000 determines the second portion of the requested content corresponds to second system avatar 1014. As illustrated in FIG. 10C, in response to the determination that the second portion of the requested content corresponds to second system avatar 1014, computer system 1000 ceases to display user interface object 1004 and displays second system avatar 1014. In some embodiments, second system avatar 1014 is displayed with one or more visual characteristics that are the same size and/or the same facial expressions as user interface object 1004. In some embodiments, as illustrated in FIG. 10C, computer system 1000 displays second system avatar 1014 as having the same facial expression as user interface object 1004 in FIG. 10B. In other examples, second system avatar 1014 is displayed with one or more visual characteristics that are a different size and/or different facial expressions than user interface object 1004. In some embodiments, as illustrated in FIG. 10C, computer system 1000 displays second system avatar 1014 as having larger ears than user interface object 1004 in FIG. 10B. As illustrated in FIG. 10C, second system avatar 1014 is displayed to appear as a dog to correspond to the second portion of the requested content (e.g., second text 1018). In other examples, the requested content corresponds to a different system avatar, causing computer system 1000 to display the different system avatar instead of second system avatar 1014. In some embodiments, computer system 1000 synchronizes the movements (e.g., visually via the display and/or physically via a portion of computer system 1000) of second system avatar 1014 with the second portion of the requested content. In some embodiments, since the second portion of the requested content corresponds to “the dog started all alone,” second system avatar 1014 can appear to be looking around (e.g., visually looking side to side via the display and/or physically rotating side to side via movement of a portion of computer system 1000) as part of the content that corresponds to the second portion of the requested content. In some embodiments, computer system outputs audio of a voice corresponding to second system avatar 1014 that is different than the audio of the voice corresponding to user interface object 1004. [0254] In some embodiments, computer system 1000 outputting the second portion of the requested content includes outputting audio corresponding to reading second text 1018, followed by outputting audio that enhances the requested content. In some embodiments, computer system 1000 can output content to seem like second system avatar 1014 is reading a portion of the story then barking. In some embodiments, in response to outputting audio corresponding to reading second text 1018, computer system 1000 moves the portion of computer system 1000 a first degree of movement. In some embodiments, in response to outputting audio corresponding to reading second text 1018, computer system 1000 can lean towards the user (e.g., a small amount). In some embodiments, in response to outputting audio that enhances the story, computer system 1000 moves the portion of computer system 1000 a second degree of movement different from the first degree of movement. In some embodiments, in response to outputting audio of a dog barking, computer system can bounce the portion of computer system 1000 to imitate the movements of an excited dog. In some embodiments, the second degree of movement is greater than the first degree of movement.
[0255] At FIG. 10C, computer system 1000 detects audio input 1020. Audio input 1020 corresponds to a voice input request from the user for computer system 1000 to switch to providing output of a third portion of the requested content (e.g., “Change to when the dog is safe”). The voice input request does not include an explicit indication of the third portion of the requested content (e.g., does not include the name of a chapter/section, the number associated with a chapter/section of the requested content, and/or a page number of a chapter/section). The voice input includes a description of one or more attributes of a scene. In some embodiments, the description of one or more attributes of a scene includes description of a portion of the plot of the requested content (e.g., when the dog is safe, when she confronts the witch, and/or when the dragon burns the town). In some embodiments, the user can say “Jump to the part where they find the magic key,” causing computer system 1000 to change to outputting the portion of the requested content that contains the point when the characters find the magic key. In some embodiments, the description of one or more attributes of a scene includes description of a portion of the theme of the requested content (e.g., action, calm, violence, love, and/or musical). In some embodiments, the user can say “I only want to watch the dance numbers,” causing computer system 1000 to only output requested content portions with dancing. In some embodiments, the description of one or more attributes of a scene includes description of a portion of a character’s arch in the requested content (e.g., when the character turns good, when the character betrays the best friend, and/or when the character realizes they had the power the whole time). In some embodiments, the user can say “I want to see the training montage,” causing computer system 1000 to output the portion of the requested content with the training montage. In some embodiments, the description of one or more attributes of a scene includes description of one or more characteristics in the requested content (e.g., everyone is wearing red, the scene in the police station, and/or the part with the elephants). In some embodiments, the user can say “Go back to the scene on the train,” causing computer system 1000 to jump back to outputting the portion of the requested content that contains the scene on the train.
[0256] In some embodiments, computer system 1000 performs different operations in response to different voice inputs. In some embodiments, as described above, in response to a first voice input, computer system 1000 can change what part of the requested content is being output. In response to a second voice input (e.g., different from the first voice input), computer system 1000 can modify aspects of the content playback (e.g., volume, speed, language, font size, tempo, and/or font contrast). In some embodiments, while detecting a voice input that does not correspond to a request to change to outputting a third portion of requested content while providing the second portion of the requested content, computer system 1000 continues to provide output corresponding to the second portion of the requested content. In some embodiments, if computer system 1000 detects a voice input that corresponds to a request to increase the volume of the audio content while computer system 1000 is providing output corresponding to the second portion of the requested content, computer system 1000 can continue to provide output corresponding to the second portion of the requested content while simultaneously increasing the volume of the audio content. In some embodiments, in response to detecting a voice input, computer system 1000 ceases providing output corresponding to the second portion of the requested content, allowing for more focused user interaction. In other examples, computer system 1000 only ceases providing output corresponding to the second portion of the requested content when computer system 1000 determines that the voice request is going to exceed a threshold amount of time, likely to require stopping output, and/or change what is output in response to the voice request.
[0257] As illustrated in FIG. 10D, in response to detecting audio input 1020, computer system 1000 changes the portion of the requested content being output based on the user description of one or more attributes of a scene (e.g., “Change to when the dog is safe”). Computer system 1000 changing the portion of the requested content being output includes computer system 1000 ceasing to provide output of the second portion of the requested content and providing output of a third portion of the requested content. In some embodiments, ceasing to provide output of the second portion of the requested content includes computer system 1000 ceasing movement of computer system 1000 and/or second system avatar 1014. As illustrated in FIG. 10D, outputting the third portion of the requested content includes computer system 1000 displaying third text 1024 which corresponds to the third portion of the requested content. In some embodiments, outputting the third portion of the content includes computer system 1000 outputting audio that corresponds to third text 1024. In some embodiments, in response to outputting audio that corresponds to third text 1024, computer system 1000 physically moves a portion of computer system 1000. In some embodiments, in response to outputting audio that corresponds with third text 1024, computer system 1000 can lift and drop a portion of computer system 1000 in a motion imitating a sigh while simultaneously outputting audio of a “sigh.” In some embodiments, in response to outputting audio that corresponds to third text 1024, computer system 1000 does not physically move, allowing the user to better focus on the content after changing from one scene to another in response to a request from the user (e.g., whereas if the third portion was output after a portion immediately before (e.g., without the user causing a jump to the third portion), computer system 1000 would physically move the portion of computer system 1000 in response to outputting audio that corresponds to third text 1024).
[0258] In this example, the third portion of the requested content is after the second portion of the requested content. In some embodiments, the third portion of the requested content is before the second portion of the requested content. In some embodiments, the second portion and the third portion are in the same chapter/section of the requested content. In some embodiments, the second portion and the third portion are in different chapters/sections of the requested content.
[0259] If the voice input includes a description of one or more attributes of a scene that concerns scenes before the scene of the second portion of the requested content, the third portion of the requested content is before the second portion. If the voice input includes a description of one or more attributes of a scene that concerns scenes after the scene of the second portion of the requested content, the third portion of the requested content is after the second portion. Computer system 1000 selects the third portion of the requested content based on one or more scene attributes detected within the requested content. In some embodiments, after providing output of a third portion, computer system 1000 continues to provide output of a portion that immediately follows the third portion (e.g., if no other input is provided).
[0260] In some embodiments, a fourth portion of the requested content is between the second portion and the third portion of the requested content. In this example, the fourth portion of the requested content does not contain one or more attributes of a scene requested by the user, causing computer system 1000 to not output the fourth portion of the requested content. In some embodiments, if the user requested to see the scene with the elephants and the fourth portion of the requested content contained horses instead of elephants, computer system 1000 would not output the fourth portion of the requested content and would instead output the next portion of the requested content that contained elephants. In some embodiments, the user does not request to change the portion of the requested content being output and computer system 1000 outputs the second portion, the fourth portion, and the third portion of the requested content in chronological order. At FIG. 10D, in response to detecting audio input 1020, computer system 1000 does not output the fourth portion of content.
[0261] In some embodiments, the voice input corresponds to a request to output portions of the requested content with one or more attributes of a scene (e.g., “skip to the rescue scene,” “only show the musical numbers,” and/or “go to the part in the forest”). When the voice input corresponds to a request to output portions of the requested content with one or more attributes of a scene, computer system 1000 selects the third portion of requested content based on that portion containing one or more of the attributes of a scene requested to be output. In some embodiments, the voice input corresponds to a request to not output portions of the requested content with one or more attributes of a scene (e.g., “I don’t like violence,” “don’t show the scenes with smoking,” and/or “skip the part with the butler”). When the voice input corresponds to a request to not output portions of the requested content with one or more attributes of a scene, computer system 1000 selects the third portion of requested content based on that portion not containing one or more attributes of a scene requested to not be output. In some embodiments, the user can say “I don’t like snakes,” causing computer system 1000 to skip the requested content portions with snakes and output only the requested content portions without snakes. [0262] At FIG. 10D, computer system 1000 determines the third portion of the requested content corresponds to third system avatar 1022. As illustrated in FIG. 10D, in response to the determination that the third portion of the requested content corresponds to third system avatar 1022, computer system 1000 ceases displaying second system avatar 1014 and displays third system avatar 1022. In some embodiments, third system avatar 1022 is displayed with one or more visual characteristics that are the same size and/or has the same facial expression as user interface object 1004 and/or second system avatar 1014. In some embodiments, as illustrated in FIG. 10D, computer system 1000 displays third system avatar 1022 as having the same size eyes as second system avatar 1014 in FIG. 10C. In other examples, third system avatar 1022 is displayed with one or more visual characteristics that are a different size and/or different facial expression than user interface object 1004 and/or second system avatar 1014. In some embodiments, as illustrated in FIG. 10D, computer system 1000 displays third system avatar 1022 as having longer hair than user interface object 1004 in FIG. 10B. In this example, computer system 1000 displays third system avatar 1022 as having a different appearance than user interface object 1004 and second system avatar 1014. As illustrated in FIG. 10D, third system avatar 1022 is displayed as a woman to correspond to the third portion of the content. In other examples, computer system 1000 determines that the third portion of the requested content corresponds to a different system avatar, causing computer system 1000 to display the different system avatar instead of third system avatar 1022. In some embodiments, computer system 1000 synchronizes movements (e.g., visually via the display and/or physically via a portion of computer system 1000) of third system avatar 1022 with the third portion of the content. In some embodiments, third system avatar 1022 can appear to be celebrating the part of the story that corresponds to the third portion of the content. In some embodiments, computer system 1000 outputs audio of a voice corresponding to third system avatar 1022 that is different from the audio of voices corresponding to user interface object 1004 and second system avatar 1014.
[0263] At FIG. 10E, the user moves to a location outside of the field of view of computer system 1000. In some embodiments, being at a location outside the field of view of computer system 1000 includes being too far away from computer system 1000 or being too close to computer system 1000. In other examples, being at a location outside the field of view of computer system 1000 includes being at a location that is father than computer system 1000 can turn and/or too far to one side of computer system 1000. In some embodiments, if the user walks to the side of the room to look out the window, causing the user to be at an angle
I l l farther than computer system 1000 can rotate to keep the user within the field of view, the user will be outside the field of view of computer system 1000. In some embodiments, being at a location outside the field of view of computer system 1000 includes being behind an object between the user and computer system 1000. In some embodiments, a user can be outside the field of view of computer system 1000 if the user is behind a piece of furniture, such as a cabinet, a couch, and/or a bookshelf. In some embodiments, being at a location outside the field of view of computer system 1000 includes leaving the physical environment computer system 1000 is located in (e.g., the user leaves the room). In some embodiments, being outside the field of view of computer system 1000 includes the user moving too fast for the sensors of computer system 1000 to detect. It should be recognized that, while discussed as being at a location outside of the field of view and/or not being detected within the field of view, techniques described herein can instead be performed when attention of the user is detected to not correspond to computer system 1000 (e.g., detecting that the attention of the user does not correspond to computer system 1000 can result in the same or similar operations as described below for when the user has not been detected within the field of view.
[0264] At FIG. 10E, computer system 1000 determines that the user has not been detected within the field of view for a predetermined amount of time, such as longer than 5, 10, or 30 seconds. At FIG. 10E, as indicated by the user representation 1010 not being located within diagram 1006, the user is not in field of view 1008a of computer system 1000. In some embodiments, user representation 1010 is located within diagram 1006 but is not located within field of view 1008a, still indicating that the user is not within the field of view of computer system 1000.
[0265] As illustrated in FIG. 10E, in response to the determination that the user has not been detected within the field of view for a predetermined amount of time, computer system 1000 pauses the output of requested content. In some embodiments, pausing the output of requested content includes pausing and/or stopping visual output. As illustrated in FIG. 10E, as part of pausing the output of visual content, computer system 1000 ceases displaying third system avatar 1022 and displays fourth system avatar 1026. As illustrated in FIG. 10E, fourth system avatar 1026 is the same as user interface object 1004 illustrated in FIG. 10B. Also as illustrated in FIG. 10E, as part of pausing the output of visual content, computer system 1000 displays third text 1024 as deemphasized (e.g., greyed out, blurred, decreased opacity and/or lowered contrast). In some embodiments, as part of pausing the output of visual content, third text 1024 is overlaid by a translucent color overlay. In other examples, as part of pausing the output of visual content, computer system 1000 ceases displaying third text 1024. In some embodiments, as part of pausing the output of visual content, computer system 1000 deemphasizes (e.g., greys out, blurs, decreases the opacity of and/or lowers the contrast of) the entirety of user interface 1002. In some embodiments, as part of pausing the output of visual content, computer system 1000 can grey out all of user interface 1002 including fourth system avatar 1026 and third text 1024. In some embodiments, as part of pausing the output of visual content, computer system 1000 maintains the display of the previous system avatar (e.g., third system avatar 1022). In some embodiments, as part of pausing the output of visual content, as the user moves away from computer system 1000. In some embodiments, as the user (and/or the attention of the user) moves away from computer system 1000 (e.g., as determined by one or more sensors connected to and/or in communication with computer system 1000), computer system 1000 can gradually deemphasizes third text 1024 until third text 1025 reaches a predetermined contrast level.
[0266] In some embodiments, pausing the output of requested content includes pausing and/or stopping audio output. In some embodiments, as part of pausing the output of audio output, computer system 1000 gradually deemphasizes audio content as the user moves away from computer system 1000. In some embodiments, as the user (and/or the attention of the user) moves away from computer system 1000 (e.g., as determined by one or more sensors connected to and/or in communication with computer system 1000), computer system 1000 can gradually decrease the volume of the audio output until computer system 1000 is no longer outputting audio. In some embodiments, pausing output of requested content includes ceasing movement of computer system 1000 and/or a system avatar. In some embodiments, pausing the output of content includes a combination of pausing and/or stopping of audio output, visual output, and/or movement (e.g., physical movement of computer system 1000 and/or a system avatar via the display). At FIG. 10F, the user moves to a location within the field of view of computer system 1000 (and/or directs their attention to computer system 1000).
[0267] In some embodiments, there is more than one user (e.g., person) within the physical environment. In some embodiments, one user is a primary user (e.g., the user with the most control, privileges and/or rights within the systems and/or programs of computer system 1000). In some embodiments, no user is the primary user while one user moves to a location outside the field of view of computer system 1000 and at least one other user remains within the field of view of computer system 1000. In response to detecting at least one user remaining within the field of view of computer system 1000 (and/or the attention of one user remaining corresponding to computer system 1000) while no user is the primary user, computer system 1000 continues to provide output of content. In some embodiments, one user (e.g., not the primary user) moves to a location outside the field of view of computer system 1000 while the primary user remains within the field of view of computer system 1000 (and/or remains with their attention on computer system 1000). In response to detecting the primary user remaining within the field of view (and/or the attention of the primary user corresponding to computer system 1000), computer system 1000 continues to provide output of content. In some embodiments, the primary user moves to a location outside the field of view of computer system 1000 while one or more other users remain within the field of view of computer system 1000 (and/or their attention corresponding to computer system 1000). In response to not detecting the primary user for a predetermined amount of time within the field of view, while continuing to detect one or more other users within the field of view, computer system 1000 pauses output of content.
[0268] At FIG. 10F, computer system 1000 detects the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing output of requested content. At FIG. 10F, as indicated by user representation 1010 being within field of view 1008a, the user is within the field of view of computer system 1000. As illustrated in FIG. 10F, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing output of requested content, computer system 1000 resumes output of requested content. As illustrated in FIG. 10F, in response to detecting the user within the field of view for a period of time (and/or the attention of the user corresponding to computer system 1000) after pausing output of requested content, computer system 1000 ceases providing output of the third portion of the requested content and provides output of a fifth portion of the requested content, which is before the third portion of the requested content. In this example, computer system 1000 outputs a portion of the requested content that is before the paused content to provide content context to the user, similar to turning back one page in a novel to refresh the memory of what happened in the story. As illustrated in FIG. 10F, as part of providing output of the fifth portion of the content, computer system 1000 displays fourth text 1030, which corresponds to the fifth portion of the content. At FIG. 10F, computer system 1000 determines that the fifth portion of the content corresponds to second system avatar 1014. As illustrated in FIG. 10F, in response to the determination that the fifth portion of the content corresponds to second system avatar 1014, computer system 1000 ceases displaying fourth system avatar 1026 and displays second system avatar 1014.
[0269] In some embodiments, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing requested content and before providing output of a fifth portion of the requested content, computer system 1000 continues to display fourth system avatar 1026. In some embodiments, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing requested content and before providing output of a fifth portion of the requested content, computer system 1000 changes the appearance of fourth system avatar 1026 and visually moves fourth system avatar 1026 via the display. In some embodiments, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing requested content and before providing output of a fifth portion of the requested content, computer system 1000 can display fourth system avatar 1026 in a way that it appears fourth system avatar 1026 is visually smiling and nodding in acknowledgement before continuing the requested content. In some embodiments, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing requested content and before providing output of a fifth portion of the requested content, computer system 1000 physically moves to the portion of computer system 1000 to appear as though fourth system avatar 1026 physically moves via computer system 1000. In some embodiments, in response to detecting the user within the field of view (and/or the attention of the user corresponding to computer system 1000) for a period of time after pausing requested content and before providing output of a fifth portion of the requested content, computer system 1000 can move up and down so that fourth system avatar 1026 appears to be bowing to the user before continuing the requested content. Such movements (e.g., via the display and/or via a portion of computer system 1000) are not part of the requested content but rather are used to signal to the user that the requested content is beginning again. In some embodiments, in response to detecting the user in the field of view (and/or the attention of the user corresponding to computer system 1000), fourth system avatar 1026 visually moves via the display and/or physically moves via computer system 1000.
[0270] In some embodiments, the third portion of the content is the beginning of a chapter/section and the fifth portion of the content is the end of the previous chapter/section. In some embodiments, if computer system 1000 pauses requested content at the beginning of a chapter, computer system 1000 can begin output of requested content at the end of the previous chapter to refresh the memory of the user in a way similar to how shows will play a recap of the previous episode before starting the current episode. In some embodiments, the third portion of the content is soon after the beginning of a chapter/section and the fifth portion of the content is the beginning of the same chapter/section. In some embodiments, if computer system 1000 pauses the requested content in the middle of a chapter, computer system 1000 can begin output of the requested content at the beginning of the same chapter, so the user can begin the chapter anew and be fully informed. In some embodiments, the third portion and the fifth portion of the content are in the middle of the same chapter/section, with the fifth portion being before the third portion. In some embodiments, if computer system 1000 pauses the requested content in the middle of a chapter, computer system 1000 can begin output of the requested content at a point in the chapter shortly before the pause, similar to turning back a page in a book to refresh the reader’s memory of the events. In some embodiments, the fifth portion was provided while the presence of the user was detected (and/or the attention of the user was detected to correspond to computer system 1000) and before initially providing, via one or more output devices, one or more outputs corresponding to a third portion of the content. In some embodiments, computer system 1000 can output the requested content illustrated in FIG. 10F before outputting the requested content illustrated in FIG. 10E, then pause while the user is not detected in the field of view (and/or the attention of the user corresponding to computer system 1000 is not detected), then computer system 1000 can output the requested content illustrated in FIG. 10F a second time once the user and/or the attention of the user is again detected. The third portion of content is provided after providing the fifth portion of content in response to detecting the user and/or the attention of the user within a predetermined period of time.
[0271] In the examples described above, computer system 1000 outputs content in the form of a story. For example, computer system 1000 outputs other forms of content. In some embodiments, computer system 1000 outputs content that pertains to communication (e.g., email, text messages, and/or social media feeds). In some embodiments, computer system 1000 can read the user an email in the voice of the person who sent the email while displaying the avatar of the person who sent the email. In the example where computer system 1000 is reading an email, the user can say “Jump to the conclusion,” causing computer system 1000 to start reading the email at the portion where the email reads “In conclusion.”
[0272] FIG. 11 is a flow diagram illustrating a method for displaying a system avatar using a computer system in accordance with some embodiments. Process 1100 is performed at a computer system (e.g., 100, 200, and/or 1000). Some operations in process 1100 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
[0273] As described below, process 1100 provides an intuitive way for displaying a system avatar. The method reduces the cognitive burden on a user for displaying a system avatar, thereby creating a more efficient human-machine interface. For battery operated computing devices, enabling a user to display a system avatar faster and more efficiently conserves power and increases the time between battery charges.
[0274] In some embodiments, process 1100 is performed at a computer system (e.g., 1000) that is in communication with a display component (e.g., a display screen, a projector, and/or a touch-sensitive display) and one or more input devices (e.g., a touch-sensitive surface, an input mechanism (e.g., a physical input mechanism, such as a button and/or a rotational input mechanism), a camera, a depth sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0275] The computer system displays (1102), via the display component, a first system avatar (e.g., 1004, 1014, 1022, and/or 1026), wherein the first system avatar corresponds to a first character. In some embodiments, a system avatar is a representation of a character and/or user. In some embodiments, the system avatar is a predefined avatar. In some embodiments, the system avatar is generated based on one or more characteristics and/or a description of a particular character and/or user. In some embodiments, the system avatar includes a graphical representation. In some embodiments, the system avatar includes a voice output. In some embodiments, the system avatar is one of many audio-visual representations of one of many characters and/or users. In some embodiments, the system avatar is an intelligent audio-visual representation of a particular character and/or users. In some embodiments, the system avatar is a predefined audio-visual representation of a particular character and/or user. In some embodiments, the system avatar is an animated representation of a character that includes a plurality of movement patterns. In some embodiments, the system avatar is universal and adaptable to most (and/or all) computer system applications. In some embodiments, the system avatar is a dynamic representation capable of assuming different forms throughout the content playback, adapting its audio-visual representation through multiple ways including but not limited to: transforming, shapeshifting, ceasing to display and reappearing, disguising, morphing and/or camouflaging among others. In some embodiments, a character is an underlying set of features within a content. In some embodiments, the system avatar serves as a representative embodiment of one of the many characters and/or users. In some embodiments, the character is an entity represented in the audio-visual embodiment of the system avatar. In some embodiments, the character includes a set of defining traits such as appearance, voice, behavior, movement patterns, purpose, personality, actions, reactions, gestures, mannerisms, unique abilities, powers, skills, perspectives, motivations, and/or backstories within a context and/or narrative and/or a set of information in the content. The character is a figure within the content that can take multiple forms, including but not limited to human, user, object, car, robot, animal, mythical creature, alien, supernatural being, plant, fantasy creature, historical figure, supernatural entity, imaginary creature, spirit, cyborg, and/or mechanical entity, among others. In some embodiments, the character conveys a narrative and/or information. In some embodiments, the character engages in dialogue and/or communicates with other characters of viewers. In some embodiments, the character is an entity exhibiting various movement patterns and behaviors tailored to the specific content being presented.
[0276] While displaying the first system avatar (e.g., 1004, 1014, 1022, and/or 1026), the computer system outputs (1104) (e.g., via one or more output devices (e.g., a speaker, a display component, and/or a haptic output device) in communication with the computer system) first content (e.g., 1016, 1018, 1024, and/or 1030). In some embodiments, a first content is a distinct portion of the content being output. In some embodiments, the first content represents a cohesive segment within the overall content. In some embodiments, the first content is characterized by a limitation, which can be established by a plurality of factors such as a new scene, a transition in the story, a shift in the narrative, and/or a change in the one or more characters displayed among others. In some embodiments, the first content is characterized by a defined limitation that separates it from subsequent contents within the overall content. In some embodiments, the limitation between the end of the first content and the start of a subsequent content is broad, allowing for a flexibility in defining the transition point based on the narrative or other features within the content. In some embodiments, the limitation between the first content and subsequent contents varies depending on the audiovisual characteristics of the overall content.
[0277] After (1106) outputting the first content (e.g., 1016, 1018, 1024, and/or 1030) and without detecting input via the one or more input devices (e.g., corresponding to a request to change the first system avatar), in accordance with a determination that second content (e.g., 1016, 1018, 1024, and/or 1030) different from the first content (e.g., 1016, 1018, 1024, and/or 1030) is to be output and that the second content corresponds to a second system avatar (e.g., 1004, 1014, 1022, and/or 1026) different from the first system avatar, the computer system displays (1108), via the display component, the second system avatar (e.g., without displaying, via the display component, the first system avatar), wherein the second system avatar corresponds to a second character different from the first character (e.g., as described above at FIGS. 10A-10D). In some embodiments, the first system avatar represents a first character, and the second system avatar represents a second character within the content. In some embodiments, the first character is different from the second character. In some embodiments, the first system avatar includes a first audio-visual appearance, and the second system avatar includes a second audio-visual appearance. In some embodiments, the first audio-visual appearance is different from the second audio-visual appearance. In some embodiments, the first system avatar has a first purpose in the context of the content, and the second system avatar has a second purpose in the context of the content. In some embodiments, the first purpose is different from the second purpose. In some embodiments, the second system avatar is displayed while the second content is output.
[0278] After (1106) outputting the first content and without detecting input via the one or more input devices, in accordance with a determination that third content (e.g., 1016, 1018, 1024, and/or 1030) different from the first content (e.g., 1016, 1018, 1024, and/or 1030) (and/or the second content) is to be output and that the third content corresponds to the first system avatar (e.g., 1004, 1014, 1022, and/or 1026), the computer system continues (1110) displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., as described above at FIGS. 10A-10D). In some embodiments, the first system avatar is displayed while the third content is output. In some embodiments, while displaying the first system avatar, the computer system displays, via the display component, a third system avatar corresponding to a third character. In some embodiments, the third system avatar is different from the first system avatar and/or the second system avatar. In some embodiments, the third character is different from the first character and/or the second character. In some embodiments, after outputting the first content, without detecting input via one or more input devices, and in accordance with a determination that the second content is to be output and that the second content corresponds to the second system avatar and the third system avatar, the computer system displays, via the display component, the third system avatar with the second system avatar. In some embodiments, after outputting the first content, without detecting input via one or more input devices, and in accordance with a determination that the third content is to be output and that the third content corresponds to the third system avatar, the computer system continues displaying the third system avatar and the first system avatar (e.g., without displaying the second system avatar). Displaying a different system avatar depending on which corresponding content is being output allows the computer system to enhance output by automatically providing a matching system avatar character to the content, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0279] In some embodiments, after outputting the first content (e.g., 1016, 1018, 1024, and/or 1030) and without detecting input via the one or more input devices, in accordance with the determination that the second content (e.g., 1016, 1018, 1024, and/or 1030) is to be output and that the second content corresponds to the second system avatar (e.g., 1004, 1014, 1022, and/or 1026), the computer system ceases displaying, via the display component, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., as described above at FIGS. 10A-10D) (e.g., while the second system avatar is displayed). In some embodiments, the first system avatar ceases to display in conjunction with (e.g., immediately before, immediately after, and/or contemporaneously in time with) the start of display of the second system avatar. In some embodiments, the computer system ceases to display the first system avatar within a predetermined time interval before or after displaying the second system avatar. In some embodiments, the computer system ceases to display the first system avatar gradually as the second system avatar progressively appears. In some embodiments, the computer system ceases to display the first system avatar in synchronization with the appearance of the second system avatar. In some embodiments, first system avatar transitions (e.g., changes, morphs, and/or transforms visually into) the second system avatar. Ceasing displaying the first system avatar when displaying the second system avatar allows the computer system to provide a focused and enhanced output by automatically displaying one or more system avatars at a time depending on which corresponding content is being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, allowing the computer system to avoid bum-in of the display component, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0280] In some embodiments, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., at a first time) includes (and/or has and/or is generated with) a first expression. In some embodiments, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., and outputting the second content) (e.g., at a second time different from the first time) includes the first expression (e.g., as described above at FIGS. 10A-10D). In some embodiments, an expression includes one or more characteristics of a system avatar, such as an appearance (e.g., a visual appearance, shape, color, orientation, and/or size) of one or more user interface elements of the system avatar and/or a tone of voice attributed to the first avatar. In some embodiments, a user interface element (of the one or more user interface elements) of the system avatar includes features (e.g., facial features (e.g., eyes, lips, and/or nose), body features (e.g., arms, legs and/or hands), accessories and/or articles of clothing). In some embodiments, an expression is a configuration of one or more user interface elements of a system avatar (e.g., the first system avatar and/or the second system avatar). In some embodiments, the expression is an animated face with exaggerated mouth movements, mimicking speech patterns, and gestures associated with storytelling. In some embodiments, the expression is a focused and/or serious look, conveying the system avatar’s role as a narrator or storyteller. In some embodiments, the expression is a raised eyebrow and/or an open-mouthed expression of surprise, representing the system avatar’s response to plot twists or unexpected events in a story. In some embodiments, the expression is a thoughtful and focused look, indicating the system avatar’s engagement in providing educational information or guidance. In some embodiments, the expression is an animated and enthusiastic face, conveying excitement while teaching or explaining a concept or telling a story. In some embodiments, the expression is a raised index finger or a gesture of emphasizing a point or providing instruction. In some embodiments, the expression is a wide- eyed and/or open-mouthed look of wonder, representing the system avatar’s excitement during an adventurous experience. In some embodiments, the expression is a determined and confident smirk, suggesting the system avatar’s readiness for a thrilling adventure. In some embodiments, the expression is an exhilarated grin and raised fist, signifying triumph or accomplishment after overcoming a challenge. In some embodiments, the expression is a gentle smile with soft, soothing eyes, conveying a sense of comfort and reassurance. In some embodiments, the expression is a compassionate and understanding look, indicating empathy and support. In some embodiments, the expression is a mischievous smirk and/or raised eyebrow, indicating the system avatar’s playful teasing or taunting behavior. In some embodiments, the expression is a finger placed over the mouth in a shushing gesture, implying a secretive or teasing demeanor. In some embodiments, the expression is a wink accompanied by a smile, conveying a teasing intent. In some embodiments, the expression is a furrowed brow and a concerned look, indicating the system avatar’s empathy. In some embodiments, the expression is a nod with a compassionate smile, indicating the system avatar’s recognition and acknowledgment of a certain experience. In some embodiments, the expression is a focused and attentive look, representing the avatar’s readiness to assist and provide helpful information or guidance. In some embodiments, the expression is a friendly smile and/or welcoming posture, conveying the system avatar’s approachability and willingness to support the user’s needs. In some embodiments, the expression is a thumbs-up gesture or a nod of approval, indicating the system avatar’s positive response to an action or a request. In some embodiments, the expression is a wide smile and/or raised eyebrows and/or wide-open eyes, indicating happiness or joy. In some embodiments, the expression is a furrowed brow and downtumed mouth, indicating sadness or disappointment. In some embodiments, the expression is a scrunched-up face and/or clenched fists, representing anger or frustration. In some embodiments, the expression is an exaggerated open-mouthed laugh with the avatar holding its stomach, conveying uncontrollable amusement. In some embodiments, the expression is a comically surprised face with widened eyes and dropped jaw, capturing a humorous reaction to a situation. In some embodiments, the expression is a timid smile with lowered eyes, indicating the system avatar’s shyness or bashfulness. In some embodiments, the expression is a blush on the cheeks, suggesting a sense of shyness or embarrassment. Having the first system avatar include the first expression and the second system avatar also include the first expression allows the computer system to have system avatars that have the same expressions, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0281] In some embodiments, the computer system (e.g., 1000) is in communication with one or more audio output devices (e.g., speakers). In some embodiments, outputting the first content (e.g., 1016, 1018, 1024, and/or 1030) includes outputting, via the one or more audio output devices, audio corresponding to the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) in a first voice. In some embodiments, outputting the second content (e.g., 1016, 1018, 1024, and/or 1030) includes outputting, via the one or more audio output devices, audio corresponding to the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) in a second voice different from the first voice (e.g., as described above at FIGS. 10A-10D). In some embodiments, a voice of the system avatar based on (e.g., is affected by, is determined based on, is defined by, and/or is created using) one or more properties such as: a set of one or more audio frequencies (e.g., range, pitch, timbre, and/or tone), a speed or pace (and/or variation thereof), and/or a volume (and/or variation thereof) of verbal output. In some embodiments, the system avatar's voice has a fast and energetic cadence, suggesting enthusiasm, urgency, or excitement in its speech. In some embodiments, the system avatar's voice has a rhythmic and poetic cadence, enhancing storytelling or creating a mesmerizing effect on the listener. In some embodiments, the system avatar's voice is warm and gentle, expressing kindness, empathy, or comfort to the listener. In some embodiments, the system avatar's voice is stern or authoritative, suggesting seriousness or the giving of instructions. In some embodiments, the system avatar's voice takes on a playful and lighthearted tone, evoking humor, joy, or a sense of fun. In some embodiments, the system avatar's voice is robotic or mechanical, reflecting a futuristic or technological persona. In some embodiments, the system avatar's voice includes an echo or reverberation, creating an otherworldly or ethereal effect. In some embodiments, the system avatar's voice is deep and resonant, conveying strength, authority, or wisdom. In some embodiments, the system avatar speaks at a slow and deliberate pace. Outputting the first system avatar including the first voice and outputting the second system avatar including outputting the second voice that is different from the first voice enables the computer system to adapt audio output depending on the system avatar being output, thereby reducing the number of inputs needed to perform an operation and/or performing an operation when a set of conditions has been met without requiring further user input. [0282] In some embodiments, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) has a first appearance (e.g., a first visual appearance). In some embodiments, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) has a second appearance (e.g., a second visual appearance) (e.g., when having the same expression as the first system avatar) different from the first appearance (e.g., as described above at FIGS. 10A-10D). In some embodiments, the first appearance includes the same types of features as the second appearance (e.g., both include ears, eyes, and/or noses). In some embodiments, the same types of features have different appearances in the first appearance and the second appearance (e.g., the ears in the first appearance are different in appearance than the ears in the second appearance). In some embodiments, the first appearance does not include the same types of features as the second appearance (e.g., both appearances do not include the same set of features) (e.g., the first appearance includes a nose and/or arms and the second appearance does not). In some embodiments, a system avatar has a stylized human appearance (e.g., including exaggerated or unique facial proportions, colorful hair, and/or distinct fashion choices). In some embodiments, a system avatar has an appearance of an animal, such as a dog, cat, bird, or mythical creature. In some embodiments, a system avatar has hybrid characteristics, combining human and animal features (e.g., an anthropomorphized animal). In some embodiments, a system avatar has an abstract and/or non-representational appearance, including one or more geometric shapes, patterns, and/or colors. In some embodiments, a system avatar is represented by a symbolic icon and/or logo, representing a brand, concept, and/or organization. In some embodiments, a system avatar embodies a fictional character from literature, movies, and/or games. In some embodiments, a system avatar has a fantastical appearance, such as being an elf, an alien, and/or a mythical creature. In some embodiments, a system avatar has a customizable appearance, allowing a user to personalize one or more features such as facial structure, hairstyle, clothing, accessories, and/or body proportions. In some embodiments, a system avatar adapts its appearance dynamically based on media content and/or on user preferences, context, and/or emotions. The first system avatar having the first appearance and the second system avatar having the second appearance that is different from the first appearance enables the computer system to provide a customized visual appearance to each system avatar being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input. [0283] In some embodiments, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) is a first size (e.g., at a first time, such as when the first content and/or the third content is output, and/or after outputting the first content and without detecting input via the one or more input devices). In some embodiments, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) is the first size (e.g., as described above at FIGS. 10B-10D) (e.g., at a second time, such as when the second content is output, and/or after outputting the first content and without detecting input via the one or more input devices) (e.g., the first system avatar is the same size as the second system avatar). The first system avatar being the first size and the second system avatar being the first size allows the computer system to provide a consistent visual output of the system avatar, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0284] In some embodiments, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) is a second size (e.g., at a third time, such as when the first content and/or the third content is output, and/or after outputting the first content and without detecting input via the one or more input devices). In some embodiments, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) is a third size (e.g., at a fourth time, such as when the second content is output, and/or after outputting the first content and without detecting input via the one or more input devices) different from the second size (e.g., as described above at FIGS. 10A-10B) (e.g., the first system avatar and the second system avatar are different sizes). The first system avatar being the second size and the second system avatar being the third size allows the computer system to output each system avatar with a customized size, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0285] In some embodiments, the first content (e.g., 1016, 1018, 1024, and/or 1030) corresponds (and/or is determined to correspond) to the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., as described above at FIGS. 10B-10D). In some embodiments, the first content corresponds to a system avatar based on one or more of a preexisting setting (e.g., configured automatically, by default, and/or by user input), contextual information (e.g., included in the content), and/or user input. In some embodiments, one or more characteristics of the first content is determined based on the system avatar and/or the system avatar is determined based on one or more characteristics of the first content. The first content corresponding to the first system avatar enables the computer system to determine which system avatar to output depending on the content being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0286] In some embodiments, the first content (e.g., 1016, 1018, 1024, and/or 1030) corresponds to a first application. In some embodiments, while displaying the first system avatar (e.g., 1004, 1014, 1022, and/or 1026), the computer system outputs (e.g., via one or more output devices (e.g., a speaker, the display component, and/or a haptic output device) in communication with the computer system) fourth content (e.g., 1016, 1018, 1024, and/or 1030) corresponding to a second application different from the first application. In some embodiments, while outputting the fourth content, the first system avatar is modified such that it appears like the fourth content is being output by the first system avatar. In some embodiments, the first system avatar is a global system avatar that outputs content with respect to multiple different applications of (e.g., installed on, executing on, and/or managed by) the system. Outputting the fourth content corresponding to the second application that is different from the first application allows the computer system to use the first system avatar to output different content from more than one application, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0287] In some embodiments, the second content (e.g., 1016, 1018, 1024, and/or 1030) corresponds to a third application. In some embodiments, while displaying the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) (and/or after outputting the second content while displaying the second system avatar), the computer system outputs (e.g., via one or more output devices (e.g., a speaker, the display component, and/or a haptic output device) in communication with the computer system) fifth content (e.g., 1016, 1018, 1024, and/or 1030), corresponding to a fourth application different from the third application. In some embodiments, while outputting the fifth content, the second system avatar is modified such that it appears like the fifth content is being output by the second system avatar. In some embodiments, the second system avatar is a global system avatar that outputs content with respect to multiple different applications of (e.g., installed on, executing on, and/or managed by) the system. In some embodiments, the second content is output with an indication of the third application without being output with an indication of the fourth application, and the fifth content is output with an indication of the fifth application without being output with an indication of the fifth application. Outputting the fifth content corresponding to the fourth application that is different from the third application allows the computer system to use the second system avatar to output different content for more than one application, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0288] In some embodiments, after (and/or while) displaying the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., and without detecting input via the one or more input devices (e.g., corresponding to a request to change to the first system avatar)), in accordance with a determination that sixth content (e.g., 1016, 1018, 1024, and/or 1030) different from the first content (e.g., 1016, 1018, 1024, and/or 1030) is to be output and that the sixth content does not correspond to the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., and corresponds to the first system avatar) (e.g., and does not correspond to any and/or a particular system avatar), the computer system displays, via the display component, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., as described above at FIGS. 10B-10D). In some embodiments, in accordance with a determination that the sixth content corresponding to the second system avatar, the computer system maintains displaying, via the display component, the second system avatar. Displaying the first system avatar in accordance with a determination that the sixth content that is different from the first content is to be output and that the sixth content does not correspond to the second system avatar enables the computer system to switch back to the system avatar that corresponds to the content being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0289] In some embodiments, after outputting the first content (e.g., 1016, 1018, 1024, and/or 1030) and without detecting input via the one or more input devices, in accordance with a determination that seventh content (e.g., 1016, 1018, 1024, and/or 1030) different from the first content (e.g., 1016, 1018, 1024, and/or 1030) (and/or the second content, third content, fourth content, fifth content, and/or sixth content) is to be output and that the seventh content corresponds to a third system avatar (e.g., 1004, 1014, 1022, and/or 1026) different from the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) and the second system avatar (e.g., 1004, 1014, 1022, and/or 1026), the computer system displays, via the display component, the third system avatar (e.g., as described above at FIGS. 10B-10D). Displaying the third system avatar in accordance with a determination that the seventh content that is different from the first content is to be output and that the seventh content corresponds to the third system avatar that is different from the first system avatar and the second system avatar enables the computer system to output as many system avatars as needed that the different contents correspond to, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0290] In some embodiments, the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) includes (e.g., has and/or is generated with) a first set of movement patterns (e.g., motion sequences and/or animation patterns). In some embodiments, the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) includes (e.g., has and/or is generated with) a second set of movement patterns different from the first set of movement patterns (e.g., as described above at FIGS. 10B-10D). In some embodiments, the computer system detects (e.g., via the one or more input devices and/or an instruction executed by the computer system) an input. In some embodiments, in response to detecting the input and in accordance with the first system avatar currently being displayed, the computer system causes the first system avatar to move according to a first movement pattern of the first set of one or more movement patterns. In some embodiments, in response to detecting the input and in accordance with the second system avatar currently being displayed, the computer system causes the second system avatar to move according to a second movement pattern of the second set of one or more movement patterns (e.g., the same input causes different movement patterns to be performed depending on which system avatar is currently displayed). The first system avatar including the first set of movement patterns and the second system avatar including the second set of movement patterns that is different from the first set of movement patterns allows the computer system to provide customized movement patterns to each system avatar being output, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input. [0291] In some embodiments, while displaying the first system avatar (e.g., 1004, 1014, 1022, and/or 1026) and outputting the first content (e.g., 1016, 1018, 1024, and/or 1030), the computer system synchronizes movement (e.g., via the display component) of (e.g., the computer system moves) the first system avatar (and/or a portion of the first system avatar) with (e.g., according to and/or based on) the first content (e.g., as described above at FIGS. 10B-10D) (e.g., the first system avatar is synchronized with the first content and/or output of the first content) (e.g., a mouth of the first system avatar moves such that it appears that the mouth of the first system avatar is outputting the first content). In some embodiments, while displaying the first system avatar and outputting the first content, the computer system adjusts the position of the first system avatar (and/or a portion of the first system avatar) dynamically to ensure that it does not obstruct the content being displayed (e.g., the first system avatar moves to a location where it does not cover the words that are being dictated). In some embodiments, while displaying the first system avatar and outputting the first content, the computer system synchronizes the movements of the first system avatar (and/or a portion of the first system avatar, such as the mouth) in accordance with the output of the first content (e.g., the first system avatar's mouth moves in a way that simulates it speaking the words from the first content). In some embodiments, while displaying the first system avatar and outputting the first content, the computer system adjusts the position and movement of the first system avatar to reflect those events and/or elements (e.g., the first system avatar reacts by showing surprise when a surprising event is described in the first content) of the first content. Synchronizing movement of the first system avatar with the first content while displaying the first system avatar and outputting the first content allows the computer system to enhance the first content by making the first system avatar’s behavior reflect the first content that is being output, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0292] In some embodiments, while displaying the second system avatar (e.g., 1004, 1014, 1022, and/or 1026) and outputting the second content (e.g., 1016, 1018, 1024, and/or 1030), the computer system synchronizes movement (e.g., via the display component) of (e.g., the computer system moves) the second system avatar with (e.g., according to and/or based on) the second content (e.g., as described above at FIGS. 10B-10D) (e.g., the second system avatar is synchronized with the second content and/or output of the second content) (e.g., a mouth of the second system avatar moves such that it appears that the mouth of the second system avatar is outputting the second content). In some embodiments, while displaying the second system avatar and outputting the second content, the computer system dynamically adjusts the movements of the second system avatar to emphasize important points in the content, (e.g., the avatar's gestures become more pronounced and/or its posture changes to draw attention to key information being conveyed). In some embodiments, while displaying the second system avatar and outputting the second content, the computer system synchronizes the movements of the second system avatar to convey emotions and/or moods (e.g., dispositions) in response to the content, (e.g., changes in facial expressions, body language, or overall demeanor to reflect happiness, sadness, excitement, or any other relevant emotion). In some embodiments, while displaying the second system avatar and outputting the second content, the computer system coordinates the movements of the second system avatar to simulate physical actions described in the content, (e.g., mimicking gestures, performing actions, and/or manipulating objects). Synchronizing movement of the second system avatar with the second content while displaying the second system avatar and outputting the second content allows the computer system to enhance the second content by making the second system avatar’s behavior reflect the second content that is being output, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0293] Note that details of the processes described above with respect to process 1100 (e.g., FIG. 11) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 1200 optionally includes one or more of the characteristics of the various methods described above with reference to process 1100. In some embodiments, the visual content of process 1200 can be the first system avatar of process 1100. For brevity, these details are not repeated below.
[0294] FIG. 12 is a flow diagram illustrating a method for selectively moving a portion of a computer system using a computer system in accordance with some embodiments. Process 1200 is performed at a computer system (e.g., 100, 200, and/or 1000). Some operations in process 1200 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
[0295] As described below, process 1200 provides an intuitive way for selectively moving a portion of a computer system. The method reduces the cognitive burden on a user for selectively moving a portion of a computer system, thereby creating a more efficient human-machine interface. For battery operated computing devices, enabling a user to selectively move a portion of a computer system faster and more efficiently conserves power and increases the time between battery charges.
[0296] In some embodiments, process 1200 is performed at a computer system (e.g., 1000) that is in communication with a display component (e.g., a display screen, a projector, and/or a touch-sensitive display), an audio generation component (e.g., a speaker and/or a set of speakers), and a movement component (e.g., an actuator (e.g., a pneumatic actuator, hydraulic actuator and/or an electric actuator), a movable base, a rotatable component, and/or a rotatable base). In some embodiments, the computer system is a watch, a phone, a tablet, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0297] The computer system outputs (1202), via the audio generation component, audio content (e.g., as described above at FIGS. 10B-10D) (e.g., using spatial audio and/or stereo audio) (e.g., audio corresponding to a response to detecting an input, audio corresponding to one or more instructions, operations, and/or information).
[0298] While outputting the audio content, the computer system physically moves (1204) (e.g., rotating (e.g., 0-360 degrees, tilting (e.g., 0-360 degrees), and/or moving laterally (e.g., right, left, up, and/or down))), via the movement component, a portion (e.g., a display, a center of a display, and/or a part of the computer system) of the computer system (e.g., 1000) (e.g., as described above at FIGS. 10B-10D).
[0299] While physically moving, via the movement component, the portion of the computer system (e.g., 1000), the computer system detects (1206) a request (e.g., 1012 and/or 1020) to display visual content (e.g., 1016, 1018, 1024, and/or 1030) (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input, a swipe input, a hold-and-drag input, a gaze input, an air gesture, and/or a mouse click). In some embodiments, the visual content is a representation and/or include one or more representations of the audio content. In some embodiments, the visual content is not the representation and/or does not include one or more representations of the audio content.
[0300] In response to (1208) detecting the request (e.g., 1012 and/or 1020) to display the visual content (e.g., 1016, 1018, 1024, and/or 1030), the computer system ceases (1210) (e.g., stopping, forgoing, and/or not) physically moving, via the movement component, the portion of the computer system (e.g., 1000) (e.g., as described above at FIGS. 10B-10D).
[0301] In response to (1208) detecting the request to display the visual content, the computer system displays (1212), via the display component, the visual content (e.g., 1016, 1018, 1024, and/or 1030). In some embodiments, in response to detecting the request to display the visual content, the computer system does not physical move the portion of the computer system. Ceasing physical moving the portion of the computer system and displaying the visual content in response to detecting the request to display the visual content allows the computer system to reduce visual distraction by stopping movement while displaying certain content, thereby providing improved visual feedback to the user and reducing the number of inputs needed to perform an operation.
[0302] In some embodiments, while outputting a portion of the audio content and in accordance with a determination that the portion of the audio content (e.g., audio content corresponding to a story, a narrative, an essay, a poem, a work of literature, a movie, and/or a video) is being output and that the portion corresponds to a first user (e.g., a first character of a story and/or of the audio content), the portion of the computer system (e.g., 1000) physically moves in a first manner (e.g., as described above at FIGS. 10A-10D) (e.g., with a first degree (e.g., speed, velocity, and/or distance) of movements and/or with a first type of movements and/or sets of movements). In some embodiments, while outputting the portion of the audio content and in accordance with a determination that the portion of the audio content is being output and that the portion does not correspond to the first user (and/or any user), the portion of the computer system (e.g., 1000) physically moves in a second manner different from the first manner (e.g., as described above at FIGS. 10A-10D) (and/or does not physical move in the first manner). In some embodiments, while outputting a portion of the audio content, in accordance with a determination that the portion of the audio content is being output and that the portion corresponds to a second user, different from the first user, the portion of the computer system physically moves in a third manner (e.g., with a second degree (e.g., speed, velocity, and/or distance) of movements and/or with a second type and/or sets of movements different from the first type and/or sets of movements), different from the first manner and/or the second manner. Moving the portion of the computer system in a particular manner based on whether or not a portion of the audio content corresponds to a particular user allows the computer system to move in a certain manner when audio output corresponds to a particular user but to not move in a manner when the audio output does not correspond to a user (or the particular user), thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further user input.
[0303] In some embodiments, the audio content is first audio content. In some embodiments, after (and/or before) ceasing physically moving, via the movement component, the portion of the computer system (e.g., 1000), the computer system outputs, via the audio generation component, second audio content that is different from the first audio content (e.g., without outputting the first audio content) (e.g., as described above at FIGS. 10B-10D). In some embodiments, while outputting, via the audio generation component, the second audio content, the computer system forgoes physically moving, via the movement component, the portion (and/or any portion) of the computer system (e.g., 1000) (e.g., as described above at FIGS. 10B-10D). In some embodiments, the second audio content is a different type (e.g., a work of art versus a non-work-of-art, a work of literature versus a non- work-of-literature, and/or a story versus a non-story) of audio content than the first audio content. Outputting, via the audio generation component, the second audio content without moving the portion of the computer system allows the computer system to not move when certain audio content is being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further user input.
[0304] In some embodiments, the audio content is third audio content. In some embodiments, after (and/or before) ceasing physically moving, via the movement component, the portion of the computer system (e.g., 1000), the computer system outputs, via the audio generation component, fourth audio content that is different from the third audio content (e.g., without outputting the third audio content). In some embodiments, while outputting, via the audio generation component, the fourth audio content, the computer system physically moves, via the movement component, the portion (and/or any portion) of the computer system (e.g., 1000) (e.g., as described above at FIGS. 10B-10D). In some embodiments, the fourth audio content is the same type (e.g., a story versus a non-story and/or a work of art versus a non-work of art) of audio content as the third audio content. Outputting, via the audio generation component, the second audio content while moving the portion of the computer system allows the computer system to move when certain audio content is being output, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further user input.
[0305] In some embodiments, the visual content (e.g., 1016, 1018, 1024, and/or 1030) is first visual content. In some embodiments, while physically moving, via the movement component, the portion of the computer system (e.g., 1000) (and, in some embodiments, while detecting the request to display visual content), the computer system displays, via the display component, second visual content (e.g., 1016, 1018, 1024, and/or 1030) different from the first visual content (e.g., 1016, 1018, 1024, and/or 1030). In some embodiments, the second visual content replaces the first visual content and/or vice-versa. Displaying, via the display component, second visual content different from the first visual content while physically moving, via the movement component, the portion of the computer system allows the computer system to provide visual content and physical move in certain situations (e.g., to animate a story), thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further user input.
[0306] In some embodiments, the request (e.g., 1012 and/or 1020) to display the visual content (e.g., 1016, 1018, 1024, and/or 1030) is a request to change from displaying the second visual content (e.g., 1016, 1018, 1024, and/or 1030) to displaying different visual content (e.g., as described above at FIGS. 10B-10D). In some embodiments, the second visual content has different information (e.g., text, images, and/or format) than the different visual content. Ceasing physical moving the portion of the computer system and displaying the visual content in response to detecting the request to change from displaying the second visual content to displaying different visual content allows the computer system to reduce visual distraction by stopping movement while display certain content, thereby providing improved visual feedback to the user and reducing the number of inputs needed to perform an operation.
[0307] In some embodiments, the portion of the computer system (e.g., 1000) is moved (e.g., rotated, tilted, and/or laterally moved) according to one or more characteristics of (e.g., in synchronization with and/or based on) the audio content output by the computer system (e.g., as described above at FIGS. 10B-10D). [0308] In some embodiments, the audio content is fifth audio content. In some embodiments, the visual content (e.g., 1016, 1018, 1024, and/or 1030) is third visual content (e.g., 1016, 1018, 1024, and/or 1030). In some embodiments, after (and/or before) ceasing physically moving, via the movement component, the portion of the computer system (e.g., 1000), the computer system outputs, via the audio generation component, sixth audio content different from the fifth audio content while physically moving, via the movement component, the portion of the computer system. In some embodiments, while outputting, via the audio generation component, the sixth audio content and physically moving the portion of the computer system (e.g., 1000), the computer system detects a request (e.g., 1012 and/or 1020) to display fourth visual content (e.g., 1016, 1018, 1024, and/or 1030), different from the third visual content (e.g., via a verbal input (e.g., a verbal input, an audible request, an audible command, and/or an audible statement) and/or a non-verbal input, a swipe input, a hold-and- drag input, a gaze input, an air gesture, and/or a mouse click). In some embodiments, in response to detecting the request (e.g., 1012 and/or 1020) to display the fourth visual content, in accordance with a determination that the fourth visual content is a first type of content (e.g., a work of art, a non-work-of-art, a work of literature, a non-work-of-literature, a story, and/or a non-story), the computer system ceases moving, via the movement component, the portion of the computer system (e.g., 1000). In some embodiments, in response to detecting the request to display the fourth visual content, in accordance with a determination that the fourth visual content is a second type of content, different from the first type of content, the computer system continues physically moving, via the movement component, the portion of the computer system (e.g., 1000).
[0309] In some embodiments, the visual content (e.g., 1016, 1018, 1024, and/or 1030) is a representation of an avatar (e.g., 1004, 1014, 1022, and/or 1026) (e.g., a system avatar (e.g., an avatar that has been assigned to, designed by, and/or selected by a user), a representation of a user, a representation of a personality, and/or a representation of a face). Displaying an avatar while physically moving, via the movement component, the portion of the computer system allows the computer system to provide a system avatar that interacts with user to establish a presence of a computer-generated user, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and performing an operation when a set of conditions has been met without requiring further user input. [0310] Note that details of the processes described above with respect to process 1200 (e.g., FIG. 12) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 1300 optionally includes one or more of the characteristics of the various methods described above with reference to process 1200. In some embodiments, the visual content of process 1200 can be the first portion of content of process 1300. For brevity, these details are not repeated below.
[0311] FIG. 13 is a flow diagram illustrating a method for navigating content using a computer system in accordance with some embodiments. Process 1300 is performed at a computer system (e.g., 100, 200, and/or 1000). Some operations in process 1300 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
[0312] As described below, process 1300 provides an intuitive way for navigating content. The method reduces the cognitive burden on a user for navigating content, thereby creating a more efficient human -machine interface. For battery operated computing devices, enabling a user to navigate content faster and more efficiently conserves power and increases the time between battery charges.
[0313] In some embodiments, process 1300 is performed at a computer system (e.g., 1000) that is in communication with a one or more output devices (e.g., a display screen, a projector, a touch-sensitive display, a speaker, a movement component (e.g., an actuator, a movable base, a rotatable component, and/or a rotatable base), and/or a haptic output device) and one or more input devices (e.g., a camera, a depth sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a fitness tracking device, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0314] While providing, via the one or more output devices, one or more outputs (e.g., text, images, audio, and/or movements) corresponding to a first portion of content (e.g., 1016, 1018, 1024, and/or 1030) (e.g., a book, a story, a poem, a piece of literature, a podcast, a movie and/or a video) (e.g., recorded content, written content, spoken content, video content, and/or animated content), the computer system detects (1302), via one or more input devices, voice input (e.g., 1012 and/or 1020) that includes a description of one or more attributes of content (e.g., attributes of a scene (e.g., attributes of a plot, theme, setting, and/or character) and/or attributes of a story) (e.g., literary attributes, plots, characters, themes, settings, imagery, allusion, dictions, symbolisms, connotations, and/or rhymes) corresponding to (e.g., described in and/or about) the content.
[0315] In response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030), the computer system provides (1304), via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content. In some embodiments, in accordance with a determination that one or more attributes of content correspond to the second portion, the computer system provides, via the one or more output devices, the one or more outputs corresponding to the second portion of content (e.g., the second portion does not overlap with the first portion). In some embodiments, in accordance with a determination that one or more attributes of content correspond to the first portion, the computer system provides, via the one or more output devices, the one or more outputs corresponding to the first portion of content (e.g., replay the first portion, the second portion includes the first portion and/or the second portion is the first portion). In some embodiments, the description includes characteristics of the second portion of content that are not location and/or positional in nature (e.g., not the location of the second portion within the content, such as “chapter 1”, “book 2”, “verse 3”, “page 4”, “paragraph 5”, and/or “line 6”) rather the characteristics of the second portion of content is another type of description (e.g., “take me to the place where the character saves another character, “skip to the part where the character is in the forest”, “go to the part where they were singing in the rain,” and/or “go back to where they showed the principle of forgiveness), such as a description of a subset of the plot, theme, a character, a scene, and/or a setting of the content that corresponds to the second portion of content. In some embodiments, in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, the computer system does not provide (e.g., while outputting the second portion of content), via the one or more output devices, one or more outputs corresponding to a third portion of content, wherein the third portion of content does not concern the one or more attributes of content. In some embodiments, in response to a user describing a specific action scene, the computer system skips a portion of the content that does not concern the action scene (e.g., instead includes background information and/or setup leading up to the scene) and begins outputting the action scene, (e.g., if a user describes a car chase, the system skips the initial dialogue and character introductions, jumping straight to the chase sequence) (e.g., when a user describes an action scene, the computer system skips slower-paced or mundane activities that occur before (or after) the action, such as scenes showing everyday routines, casual conversations, and/or non-essential activities that do not contribute significantly to the action.). In some embodiments, in response to a user’s description focusing on a particular action scene and/or storyline, the computer system may skip any subplots and/or secondary storylines that are unrelated to the described action and/or storyline. In some embodiments, in response to a user describing an action scene in the present timeline, the computer system skips one or more flashbacks and/or non-linear narrative elements that interrupt the flow of the action scene. Not providing one or more outputs corresponding to the third portion of content that does not concern the one or more attributes of content in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content allows the computer system to adjust its output and skip portions of content that do not correspond to the voice request describing a specific portion of content, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0316] In some embodiments, the first portion of content is at a first position in the content. In some embodiments, the second portion of content is at a second position in the content. In some embodiments, the third portion of content is at a third position in the content. In some embodiments, the third position is between (e.g., in time and/or according to a timeline corresponding to the content) the first position and the second position. In some embodiments, a position is a location (e.g., placement and/or point) of an event in time relative to other attributes of content in the timeline of the content. In some embodiments, the third position does not overlap with the first position and/or the second position. In some embodiments, the second portion of content is located at (e.g., towards and/or near) the end of the content, the third portion is located before the end of the content, and the first portion is located before the third portion. In some embodiments, the second portion of content is located in (e.g., towards and/or near) the middle of the content, the third portion is located before the middle of the content, and the first portion is located before the third portion. In some embodiments, the second portion of content represents a significant event (e.g., a turning point and/or a highlight) in the content. In some embodiments, the third portion of content (a skipped portion, in some embodiments) includes one or more scenes and/or descriptions leading up to the key event. In some embodiments, while the content is in the first portion, a user describes the denouement (the second portion, in some embodiments) that occurs in the final moments of the content. In some embodiments, the skipped portion (third portion) includes the build-up, character development, and/or plot twists leading up to the denouement (the finale and/or end of attributes of content). In some embodiments, while the content is in the first portion, a user describes a specific match-winning goal (second portion) in a soccer game. In some embodiments, the skipped portion (third portion, in some embodiments) includes the preceding gameplay, strategies, and/or chances created by both teams before the decisive goal. In some embodiments, while the content is in the first portion, a user describes the second portion involving a major revelation (twist, surprise) in the content. In some embodiments, the skipped portion includes the elements that set the stage for the revelation. In some embodiments, while the content is in the first portion, a user describes the plot twist (second portion, in some embodiments) in a mystery novel. In some embodiments, the skipped portion includes the clues, investigations, and character interactions that lead up to the surprising revelation. The first portion of content being at the first position in the content and the second portion of content being at the second position in the content and the third portion of content being at the third position in the content and the third position being between the first position and the second position enables the computer system to skip a portion of content leading up to the target portion of content described in the voice request, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, and/or performing an operation when a set of conditions has been met without requiring further user input. Providing one or more outputs corresponding to the second portion of content that concerns the one or more attributes of content corresponding to the content in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content allows the computer system to adjust its output and provide a specific portion of content that corresponds to the voice request describing the specific portion of content, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, performing an operation when a set of conditions has been met without requiring further user input, and/or allowing the computer system to avoid bum-in of the display component.
[0317] In some embodiments, the one or more output devices includes an audio output device (e.g., smart speakers, home theater system, soundbars, headphones, speakers, television speakers, and/or augmented reality headset speakers). In some embodiments, at least one of the one or more outputs corresponding to the second portion of content is audio output provided via the audio output device. Having at least one of the one or more outputs corresponding to the second portion of content be audio output provided via the audio output device enables the computer system to (1) provide auditory feedback and/or (2) increase engagement based on audio output, thereby performing an operation when a set of conditions has been met without requiring further user input.
[0318] In some embodiments, the one or more output devices includes a first display component. In some embodiments, providing the one or more outputs corresponding to the second portion of content includes displaying, via the first display component, a representation (e.g., video, image, animation, 3D rendering, augmented reality overlay, motion graphics, data visualization, and/or digital art) corresponding to the second portion of content. In some embodiments, displaying the representation corresponding to the second portion of content includes changing one or more color characteristics (e.g., hue, saturation, tone, and/or brightness) and/or lighting effects. In some embodiments, the representation corresponding to the second portion of content includes visual effects (e.g., Computer Generated Imagery (CGI) and/or practical effects) and/or animations. In some embodiments, displaying the representation corresponding to the second portion of content includes transitioning between scenes (e.g., fade-ins, fade-outs, crossfades, and/or wipes) and/or animations. In some embodiments, displaying the representation corresponding to the second portion of content includes capturing and generating the representation based on camera movements (e.g., panning, tracking shots, and/or zooming). In some embodiments, the representation corresponding to the second portion of content includes animated text and/or typography that transforms and/or transitions while being displayed. Providing the one or more outputs corresponding to the second portion of content including displaying the representation corresponding to the second portion of content allows the computer system to (1) provide visual feedback and/or (2) increase engagement with visual output, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0319] In some embodiments, the one or more output devices includes a movement component (e.g., an actuator (e.g., a pneumatic actuator, hydraulic actuator and/or an electric actuator), a movable base, a rotatable component, and/or a rotatable base). In some embodiments, providing the one or more outputs includes moving via the movement component (e.g., as described above at FIGS. 10A-10D). In some embodiments, providing the one or more outputs includes moving a representation of content. In some embodiments, moving the representation of content illustrates transitions of characters and/or objects moving between different locations on the display. In some embodiments, moving the representation of content includes moving characters in a scene and/or objects appearing, disappearing, and/or transitioning from one position to another, (e.g., characters and/or objects performing actions such as walking, dancing, fighting, ball bouncing, water flowing, crowd cheering, flag waving, and/or vehicle speeding). In some embodiments, movement output includes the embodiment of characters and their interaction in a displayed environment. In some embodiments, moving the representation of content represents contextually relevant interactions of characters with the displayed environment, (e.g., characters’ physical actions and responses in the displayed environment), (e.g., characters navigating (traversing) the displayed environment).
[0320] In some embodiments, the first portion of content (e.g., 1016, 1018, 1024, and/or 1030) is at a fourth position in the content. In some embodiments, the second portion of content (e.g., 1004, 1014, 1022, and/or 1026) is at a fifth position in the content that is before the fourth position. In some embodiments, in response to a user’s description, the computer system rewinds the media playback to the second portion of the content and/or to a portion that is immediately before the second portion. The first portion of content being at the fourth position in the content and the second portion of content being at the fifth position in the content that is before the fourth position allows the computer system to increase engagement by providing the ability to rewind the media playback to a previous scene, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input. [0321] In some embodiments, the first portion of content (e.g., 1016, 1018, 1024, and/or 1030) is at a sixth position in the content. In some embodiments, the second portion of content is at a seventh position in the content that is after the sixth position. In some embodiments, in response to a user’s description, the computer system fast forwards the media playback to the second portion of the content and/or to a portion that is immediately after the second portion. The first portion of content being at the sixth position in the content and the second portion of content being at the seventh position in the content that is after the sixth position allows the computer system to increase engagement by providing the ability to fast-forward the media playback to an upcoming scene, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0322] In some embodiments, the voice input does not include an explicit indication (e.g., the voice input does not include any specific mention of chapter names, section numbers, page numbers, video time indications (e.g., “30 seconds into the video” and/or “at 1 hour and 59 minutes”), and/or any other delimiters, particularly numerical and/or sectional locations in the content) of the second portion of content (e.g., 1016, 1018, 1024, and/or 1030). The voice input not including an explicit indication of the second portion of content enables the computer system to locate the second portion without an explicit indication of its position in the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0323] In some embodiments, the first portion is in a first subset (e.g., chapter, segment, page, section, and/or other predefined indication of a segmentation) of the content and the second portion is in the first subset of the content (e.g., 1016, 1018, 1024, and/or 1030). The first portion being in the first subset of the content and the second portion being in the first subset of the content enables the computer system to locate the second portion in a particular subset of content that is being played, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0324] In some embodiments, the first portion is in a second section (e.g., chapter, segment, page, section, and/or other predefined indication of a segmentation) of the content (e.g., 1016, 1018, 1024, and/or 1030) and the second portion is in a third section of the content different from the second section of the content. The first portion being in the second section of the content and the second portion being in the third section of the content that is different from the second section of the content allows the computer system to determine the second portion in different subsets of the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0325] In some embodiments, the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a description of a portion of a plot of the content. In some embodiments, the output corresponding to the second portion of the content concerns the portion of the plot of the content. In some embodiments, another portion of the plot of the content is in the first portion of the content (e.g., replay of portion of the plot). In some embodiments, a plot (story) includes a series of attributes of content collectively forming the narrative structure of a story (storyline) in the content. In some embodiments, the plot shapes the flow (development) of the story. In some embodiments, a plot includes a set of subsets (chapters, segments, pages, sections, and/or other predefined indication of segmentation) including and not limited to an introduction (e.g., the introduction of the story provides background information and/or context for the following attributes of content), a rising action (e.g., series of attributes of content building upon one another, increasing the complexity, and/or challenges of the story), a culmination (e.g., pivotal moment and/or movement of tension in the plot where conflicts reach their peak intensity), a falling action (e.g., the consequences and/or aftermath of the culmination) and/or a resolution (e.g., brings closure to the story, provides outcomes to the main conflicts in the story). The description of one or more attributes of content corresponding to the content including a description of the portion of the plot of the content and the output corresponding to the second portion of the content concerning the portion of the plot of the content allows the computer system to (1) seamlessly integrate descriptions in a request for content and/or (2) tailor its output based on implicit descriptions of the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0326] In some embodiments, the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a description of a portion of a theme of the content. In some embodiments, the output corresponding to the second portion of content concerns the portion of the theme of the content. In some embodiments, the portion of the theme of the content is in the first portion of the content (e.g., the replay of the portion of the theme). In some embodiments, the portion of the theme of the content is in the second portion of the content. In some embodiments, the first portion of content includes the second portion of content. In some embodiments, the content includes a plurality of themes (e.g., overarching themes, recurring themes, subplot themes, overlapping and/or distinct themes). In some embodiments, a theme manifests in a plurality of ways in the audiovisual content (e.g., narrative themes, visual themes, audio themes, and/or subtextual themes]). The description of one or more attributes of content corresponding to the content including a description of the portion of the theme of the content and the output corresponding to the second portion of the content concerning the portion of the plot of the content allows the computer system to tailor its output depending on thematic characteristics of portions of the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0327] In some embodiments, the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a description of a character’s arc in the content. In some embodiments, the output corresponding to the second portion of content concerns the character’s arc in the content. In some embodiments, the development of the character is in the first portion of the content. In some embodiments, the development of the character is in the second portion of the content. In some embodiments, the first portion of content includes the second portion of content. In some embodiments, the development of the character manifests in a plurality of themes (e.g., motivation, conflict, growth, transformation, relationships, decision-making, and/or navigating challenges). The description of one or more attributes of content corresponding to the content including the description of the character’s arc in the content and the output corresponding to the second portion of content concerning the character’s arc in the content allows the computer system to navigate its output to a particular point in the character’s development, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input. [0328] In some embodiments, the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a description of one or more characteristics (e.g., genre, visual and auditory storytelling, dramatic structure, cinematic techniques, language, narrative structure, character’s development, and/or theme) in the content. In some embodiments, the output corresponding to the second portion of content concerns the one or more characteristics in the content. In some embodiments, the one or more characteristics in the content is in the first portion of content. In some embodiments, the one or more characteristics in the content is in the second portion of content. In some embodiments, the first portion of content includes the second portion of content. The description of one or more attributes of content corresponding to the content including the description of one or more characteristics in the content and the output corresponding to the second portion of content concerning the one or more characteristics in the content allows the computer system to tailor its output depending on particular characteristics of portions of the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0329] In some embodiments, after outputting, via the one or more output devices, one or more outputs corresponding to the second portion of content (e.g., 1016, 1018, 1024, and/or 1030) (and, in some embodiments, in accordance with a determination that an input was not detected while and/or after outputting providing, via the one or more output devices, one or more outputs corresponding to the second portion of content and before one or more outputs corresponding to another portion of content), the computer system provides, via the one or more input devices, one or more outputs corresponding to a portion of content that immediately follows the second portion of content. In some embodiments, the portion of content that immediately follows the second portion of content is different from the second portion of content. In some embodiments, the one or more outputs corresponding to a portion of content that immediately follows the second portion of content is different from the one or more outputs corresponding to the second portion of content. In some embodiments, continuously providing output of a portion that immediately follows the second portion, in the absence of detecting, via one or more input devices, voice input that includes a description of attributes of content corresponding to the content (e.g., play content as normal without skipping any portions). Providing one or more outputs corresponding to the portion of content that immediately follows the second portion of content allows the computer system to provide a seamless experience by automatically resuming output from the desired portion of content, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0330] In some embodiments, in accordance with the determination that the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) included in the voice input is before one or more attributes of content to which the first portion of content concerns, the second portion of content is at a position in the content that is before a position of the first portion of content in the content. In some embodiments, in accordance with the determination that the description of one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) included in the voice input is after one or more attributes of content to which the first portion of content concerns, the second portion of content is at a position in the content that is after the position of the first portion of content in the content. In some embodiments, a position is a location (e.g., placement and/or point) of an event in time relative to other attributes of content in the timeline of the content. The description of one or more attributes of content corresponding to the content being before one or more attributes of content in the first portion of content, causing displaying the second portion before the first portion, and the description of one or more attributes of content corresponding to the content being after one or more attributes of content in the first portion of content, causing displaying the second portion of content after the first portion, enables the computer system to intelligently transition to the desired portion of content regardless of its position relative to the current portion of content being output, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0331] In some embodiments, in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030), in accordance with a determination that the one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a first event, the computer system selects a first respective portion of the content as the second portion of content. In some embodiments, in accordance with a determination that the one or more attributes of content corresponding to the content includes the first event, providing, via the one or more output devices, one or more outputs corresponding to the second portion of content includes selecting the first respective portion of the content as the second portion of content. In some embodiments, in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, in accordance with a determination that the one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030) includes a second event, different from the first event, the computer system selects a second respective portion of the content, different from the first respective portion of the content, as the second portion of content. In some embodiments, in accordance with a determination that the one or more attributes of content corresponding to the content includes the second event, providing, via the one or more output devices, one or more outputs corresponding to the second portion of content includes selecting the second respective portion of the content as the second portion of content. Selecting the first respective portion of the content as the second portion of content if the one or more attributes of content corresponding to the content includes the first event, and selecting the second respective portion of content as the second portion of content if the one or more attributes of content corresponding to the content includes the second event allows the computer system to methodically navigate to the desired portion of content depending on the attributes of content occurring in the content, thereby reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0332] In some embodiments, the voice output is first voice output (e.g., as described above at FIGS. 10A-10D). In some embodiments, while providing, via the one or more output devices, one or more outputs corresponding to the first portion of content (e.g., 1016, 1018, 1024, and/or 1030), the computer system detects, via one or more input devices, a second voice input (e.g., 1020) (e.g., volume input, brightness control) different from the first voice input. In some embodiments, in response to detecting the second voice input, the computer system continues to provide output corresponding to the first portion of content. In some embodiments, the second voice input is voice input that does not include a description of one or more attributes of content corresponding to the content. In some embodiments, in response to detecting third voice input that does not include a description of one or more attributes of content and in accordance with a determination that the third voice input includes a delimiter or marker (e.g., “page 4”, “chapter I”, and/or “section A”) of a position in the content, the computer system provides output of a portion of content corresponding to the third voice input (e.g., portion of the story and/or portion of the story that includes the first portion). In some embodiments, in response to detecting third voice input that does not include a description of one or more attributes of content and in accordance with a determination that the third voice input does not include the delimiter or marker (e.g., page 4, chapter I, and/or section A) of a position in the content, the computer system does not provide the portion of content corresponding to the third voice input (e.g., portion of the story and/or portion of the story that includes the first portion). Continuing to provide output corresponding to the first portion of content in response to detecting the second voice input enables the computer system to ignore unrelated input to the content and continue its output without interruption, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0333] In some embodiments, in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content (e.g., 1016, 1018, 1024, and/or 1030), the computer system ceases to provide, via the one or more output devices, the one or more outputs corresponding to the first portion of content. In some embodiments, the first portion of content does not include the second portion of content. Ceasing to provide the one or more outputs corresponding to the first portion of content in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content allows the computer system to seamlessly transition its output to the desired portion of content, thereby providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0334] In some embodiments, the voice input (e.g., 1012 and/or 1020) corresponds to a request to skip the third portion of content (e.g., a request to skip a scene). In some embodiments, the voice input (e.g., 1012 and/or 1020) corresponds to a request to jump to (and/or immediately output) the second portion of content (e.g., a request to jump a scene). In some embodiments, the voice input (e.g., 1012 and/or 1020) corresponds to a request to jump from the first portion of content (and/or the third portion of content) (e.g., a request to jump from a scene). [0335] Note that details of the processes described above with respect to process 1300 (e.g., FIG. 13) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 1400 optionally includes one or more of the characteristics of the various methods described above with reference to process 1300. In some embodiments, the first portion of content of process 1300 can be the first portion of content of process 1400. For brevity, these details are not repeated below.
[0336] FIG. 14 is a flow diagram illustrating a method for pausing content using a computer system in accordance with some embodiments. Process 1400 is performed at a computer system (e.g., 100, 200, 1000). Some operations in process 1400 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
[0337] As described below, process 1400 provides an intuitive way for pausing content. The method reduces the cognitive burden on a user for pausing content, thereby creating a more efficient human-machine interface. For battery operated computing devices, enabling a user to pause content faster and more efficiently conserves power and increases the time between battery charges.
[0338] In some embodiments, process 1400 is performed at a computer system (e.g., 1000) that is in communication with one or more output devices (e.g., a display screen, a projector, a touch-sensitive display, a speaker, a movement component (e.g., an actuator, a movable base, a rotatable component, and/or a rotatable base), and/or a haptic output device) and one or more input devices (e.g., a camera, a depth sensor, a proximity sensor, a motion sensor, and/or a microphone). In some embodiments, the computer system is a watch, a phone, a tablet, a processor, a head-mounted display (HMD) device, a communal device, a media device, a speaker, a television, and/or a personal computing device.
[0339] While providing, via the one or more output devices, one or more outputs corresponding to a first portion of content (e.g., 1016, 1018, 1024, and/or 1030) (e.g., a book, a story, a poem, a piece of literature, a podcast, a movie and/or a video) (e.g., recorded content, written content, spoken content, video content, and/or animated content), the computer system detects (1402), via the one or more inputs devices, that an attention of a user (e.g., 1010) (e.g., a primary user and/or a user that started the content being output) no longer corresponds to the computer system (e.g., 1000) (e.g., as described above at FIGS. 10D-10E) (e.g., for a predetermined period of time (e.g., 0-10 seconds)) (and/or detecting that a presence of the user is no longer detected (e.g., for the predetermined period of time)) (e.g., the user is leaving and/or left an area) (e.g., no user is in the area after the user leaves the area) (e.g., that the user is no longer in a field of detection of the computer system, that the user is no longer in a field of view of one or more cameras, and/or that the attention of the user corresponds to and/or is directed in a direction other than the computer system).
[0340] In response to detecting that the attention of the user (e.g., 1010) no longer corresponds to the computer system (e.g., 1000), the computer system ceases (1404) to provide the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030).
[0341] While the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) are not being provided, the computer system detects (1406), via the one or more inputs devices, that the attention of the user (e.g., 1010) corresponds to the computer system (e.g., 1000) (and/or detecting the presence of the user) (e.g., that the user is within the field of detection of the computer system and/or the field of view of the one or more cameras).
[0342] In response to detecting that the attention of the user (e.g., 1010) corresponds to the computer system (e.g., 1000), the computer system provides (1408), via the one or more output devices, one or more outputs corresponding to a second portion of the content (e.g., 1016, 1018, 1024, and/or 1030) that is at or before (e.g., immediately before) the first portion of the content (e.g., as described above at FIGS. 10E-10F). In some embodiments, the second portion of the content is before the first portion of the content. In some embodiments, the second portion of the content is not at the first portion of the content or the same as and/or positioned in the content within the first portion of the content. In some embodiments, the second portion of the content is at the first portion of the content, the same as the first portion of the content, and/or is positioned in the content within the first portion of the content.
[0343] In some embodiments, the one or more output devices includes an audio output device (e.g., smart speakers, home theater system, soundbars, headphones, speakers, television speakers, and/or augmented reality headset speakers). In some embodiments, providing the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) (and/or corresponding to the second portion of the content) includes outputting, via the audio output device, audio (e.g., as described above at FIGS. 10E-10F). Providing the one or more outputs corresponding to the content including outputting audio enables the computer system to (1) provide auditory feedback and/or (2) increase engagement based on audio output, thereby performing an operation when a set of conditions has been met without requiring further user input.
[0344] In some embodiments, the one or more output devices includes a first display component. In some embodiments, providing the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) includes displaying, via the first display component, a representation (e.g., video, image, animation, 3D rendering, augmented reality overlay, motion graphics, data visualization, and/or digital art) (e.g., corresponding to the content). In some embodiments, displaying the representation corresponding to the content includes changing one or more color characteristics (e.g., hue, saturate, tone, and/or brightness) and/or lighting effects. In some embodiments, the representation corresponding to the content includes visual effects (e.g., Computer Generated Imagery (CGI) and/or practical effects) and/or animations. In some embodiments, displaying the representation corresponding to the content includes transitioning between scenes (e.g., fade-ins, fade-outs, crossfades, or wipes) and/or animations. In some embodiments, displaying the representation corresponding to the content includes capturing and generating the representation based on camera movements (e.g., panning, tracking shots, and/or zooming). In some embodiments, the representation corresponding to the content includes animated text and/or typography that transforms and/or transitions while being displayed. Providing the one or more outputs corresponding to the content including displaying a representation allows the computer system to (1) provide visual feedback and/or (2) increase engagement with visual output, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0345] In some embodiments, the one or more output devices includes a movement component (e.g., an actuator (e.g., a pneumatic actuator, hydraulic actuator and/or an electric actuator), a movable base, a rotatable component, and/or a rotatable base). In some embodiments, providing the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) includes moving, via the movement component, a first portion (e.g., a hardware portion, a display screen, a hardware button, and/or a center of the display and/or another portion of the display) of the computer system (e.g., as described above at FIGS. 10A-10F). In some embodiments, the computer system moves a representation of the content. In some embodiments, the computer system provides movement output that illustrates the transitions of characters and/or objects moving between different locations on the display. In some embodiments, movement output includes movement of characters in a scene and/or objects appearing, disappearing, and/or transitioning from one position to another, (e.g., characters or objects performing actions such as walking, dancing, fighting, ball bouncing, water flowing, crowd cheering, flag waving, and/or vehicle speeding). In some embodiments, movement output represents contextually relevant interactions of characters with the displayed environment (e.g., characters’ physical actions and responses in the displayed environment) (e.g., characters navigating and/or traversing the displayed environment). Providing the one or more outputs corresponding to the first portion of the content including moving the first portion of the computer system allows the computer system to (1) provide visual feedback and/or (2) increase engagement with movement output, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0346] In some embodiments, the second portion of the content (e.g., 1016, 1018, 1024, and/or 1030) is positioned (e.g., as described above in relation to process 800) (e.g., in time and/or according to a timeline corresponding to the content) in the content at a position (e.g., as described above in relation to process 800) that is not at the beginning of a first subset (e.g., a predefined subset) (e.g., as described above in relation to process 800) of the content. The second portion of the content being positioned in the content at the position that is not at the beginning of the first subset of the content enables the computer system to (1) preserve contextual understanding of the output and/or (2) provide an enhanced user experience, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0347] In some embodiments, the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) is positioned (e.g., as described above in relation to process 800) in the content at a position (e.g., as described above in relation to process 800) that is at the beginning of a second subset (e.g., a predefined subset) (e.g., as described above in relation to process 800) of the content. The first portion of the content being positioned in the content at the position that is at the beginning of the second subset of the content allows the computer system to (1) preserve contextually assimilation of the output and/or (2) provide an enhanced user experience, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0348] In some embodiments, the first portion is positioned (e.g., in time and/or according to a timeline corresponding to the content) (e.g., as described above in relation to process 800) in the content at a first position (e.g., as described above in relation to process 800) that is in a third subset (e.g., a predefined subset) (e.g., as described above in relation to process 800) in the content. In some embodiments, the second portion is positioned in the content (e.g., 1016, 1018, 1024, and/or 1030) at a second position that is in the third subset of the content. In some embodiments, the first position is different from the second position. In some embodiments, the first position and the second position are not at a terminal position (e.g., such as a beginning position and an ending position) of the third subset of the content (e.g., 1016, 1018, 1024, and/or 1030). The first position and the second portion being in the third subset in the content and neither position being at the terminal position of the third subset of the content allows the computer system to (1) provide an enhanced user experience and/or (2) preserve contextual understanding of each subset of the content, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0349] In some embodiments, before providing the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030), the computer system provides, via the one or more output devices, the one or more outputs corresponding to the second portion of the content (e.g., 1016, 1018, 1024, and/or 1030) while detecting that the attention of the user (e.g., 1010) corresponds to the computer system (e.g., 1000). In some embodiments, the second portion of the content is before (e.g., as described above in relation to process 800) the first portion of the content in the content. In some embodiments, in response to detecting that the user will be detected within the predetermined period of time, the computer system replays and/or re-outputs the one or more outputs corresponding to the second portion of the content while detecting that the attention of the user corresponds to the computer system. Providing the one or more outputs corresponding to the second portion of the content while detecting that the attention of the user corresponds to the computer system before providing the one or more outputs corresponding to the first portion of the content allows the computer system to provide an improved user experience by adapting its output depending on the user’s presence, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0350] In some embodiments, before detecting that the attention of the user (e.g., 1010) no longer corresponds to the computer system (e.g., 1000), the computer system detects, via the one or more input devices, that the attention of the user corresponds to the computer system. In some embodiments, in response to detecting that the attention of the user (e.g., 1010) corresponds to the computer system (e.g., 1000), the computer system provides, via the one or more output devices, the one or more one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) after providing (and, in some embodiments, immediately after, where no portion is between the first portion of content and the second position of content) the one or more outputs corresponding to the second portion of the content (e.g., 1016, 1018, 1024, and/or 1030). Providing the one or more one or more outputs corresponding to the first portion of the content after providing the one or more outputs corresponding to the second portion of the content enables the computer system to (1) provide an enhanced user experience and/or (2) incorporate contextual information after an interruption in its output, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0351] In some embodiments, the computer system (e.g., 1000) is in communication with a movement component. In some embodiments, while outputting audio corresponding to the content, the computer system detects a condition. In some embodiments, in response to detecting the condition, in accordance with a determination that the condition includes detecting voice input while outputting audio corresponding to a respective scene, the computer system moves, via the movement component, (e.g., a portion of the computer system (e.g., as described above)) in a first manner (e.g., with a first pattern of movement, speed of movement, and/or direction of movement). In some embodiments, in response to detecting the condition and in accordance with the determination that the condition includes detecting voice input while outputting audio corresponding to the respective scene, the computer system moves, via a display component of the one or more output devices, a representation (e.g., as described above in relation to process 800) in a first manner (e.g., with a first pattern of movement, speed of movement, and/or direction of movement). In some embodiments, in response to detecting the condition, in accordance with a determination that the condition does not include detecting voice input while outputting audio corresponding to the respective scene, the computer system moves, via the movement component, (e.g., a portion of the computer system (e.g., as described above)) in a second manner (e.g., with a second pattern of movement, speed of movement, and/or direction of movement) different from the first manner (e.g., as described above at FIGS. 10A-10F) (e.g., without moving in the first manner). Moving in the first manner when a determination is made the condition includes detecting voice input while outputting audio corresponding to the respective scene and moving in the second manner when a determination is made the condition does not include detecting voice input while outputting audio corresponding to the respective scene allows the computer system to enhance engagement by outputting content flexibly in conditional scenarios, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0352] In some embodiments, the user (e.g., 1010) is a first user. In some embodiments, while providing, via one or more output devices, one or more outputs corresponding to the content (e.g., the first portion of the content, the second portion of the content, and/or another portion of the content), the computer system detects that an attention of a second user (e.g., 1010) (e.g., the first user or another user different from the first user) no longer corresponds to the computer system (e.g., 1000) while an attention of a third user (e.g., 1010) corresponds to the computer system, wherein the third user is different from the second user. In some embodiments, in response to detecting that the attention of the second user (e.g., 1010) no longer corresponds to the computer system (e.g., 1000) while the attention of the third user (e.g., 1010) corresponds to the computer system, the computer system continues to provide one or more outputs corresponding to the content (e.g., as described above at FIGS. 10D-10F) (and/or forgoing ceasing to provide one or more outputs corresponding to the content). In some embodiments, the second user is the first user, and the third user is not the first user. In some embodiments, in response to detecting that the attention of the second user and the attention of the third user no longer corresponds to the computer system, the computer system ceases to provide one or more outputs corresponding to the content. In some embodiments, in response to detecting that the attention of the third user no longer corresponds to the computer system while the attention of the second user corresponds to the computer system, the computer system ceases to provide one or more outputs corresponding to the content. In some embodiments, in response to detecting that the attention of the third user no longer corresponds to the computer system while the attention of the second user corresponds to the computer system, the computer system continues to provide one or more outputs corresponding to the content. Continuing to provide one or more outputs corresponding to the content in response to detecting that the attention of the second user no longer corresponds to the computer system while the attention of the third user corresponds to the computer system allows the computer system to accommodate more than one user in its output, thereby providing improved visual feedback to the user, and/or performing an operation when a set of conditions has been met without requiring further user input.
[0353] In some embodiments, the user (e.g., 1010) is a fourth user (e.g., 1010). In some embodiments, while providing, via one or more output devices, one or more outputs corresponding to the content (e.g., the first portion of the content, the second portion of the content, and/or another portion of the content), the computer system detects that an attention of a fifth user (e.g., 1010) no longer corresponds to the computer system (e.g., 1000) while detecting that an attention of a sixth user (e.g., 1010) corresponds to the computer system, wherein the sixth user (e.g., 1010) is different from the fifth user (e.g., 1010). In some embodiments, in response to detecting that the attention of the fifth user (e.g., 1010) no longer corresponds to the computer system (e.g., 1000) while detecting that the attention of the sixth user (e.g., 1010) corresponds to the computer system, in accordance with a determination that the fifth user (e.g., 1010) is a first type of user (e.g., a primary user, a controlling user, a user with a first amount of privileges with respect to operating the computer system, and/or a user who is registered to and/or who the computer system belongs to), the computer system continues to provide one or more outputs corresponding to the content. In some embodiments, in response to detecting that the attention of the fifth user no longer corresponds to the computer system while detecting that the attention of the sixth user corresponds to the computer system, in accordance with a determination that the fifth user (e.g., 1010) is a second type of user (e.g., a guest user, not a primary user, a user with a second amount of privileges with respective to operating the computer system that are less than the first amount of privileges with respective to operating the computer system, not a controlling user, and/or not a user who is registered to and/or not who the computer system belongs to) different from the first type of user, the computer system forgoes continuing to provide one or more outputs corresponding to the content (and/or ceasing to provide one or more outputs corresponding to the content). Continuing to provide one or more outputs corresponding to the content when a determination is made that the fifth user is a primary user and not continuing to provide one or more outputs corresponding to the content when a determination is made that the fifth user is not the primary user allows the computer system to improve the user experience for the primary user and prevent other users from creating confusion for the primary user, thereby providing improved visual feedback to the user, performing an operation when a set of conditions has been met without requiring further user input, and/or increasing security.
[0354] In some embodiments, ceasing to provide the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) includes gradually deemphasizing (e.g., deceasing audio, dimming display, fading out background music, and/or removing highlighting, changing the color, decreasing the size, and/or removing highlighting from one or more user interface elements and/or representations corresponding to), via the one or more output devices, the one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030). Ceasing to provide the one or more outputs corresponding to the first portion of the content including gradually de-emphasizing the one or more outputs corresponding to the first portion of the content allows the computer system to provide a seamless experience by gradually transitioning its output, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0355] In some embodiments, ceasing to provide one or more outputs corresponding to the first portion of the content (e.g., 1016, 1018, 1024, and/or 1030) includes pausing (e.g., freezing the display, freezing a position of the computer system, and/or pausing playback and/or the playing of the content), via the one or more output devices, the one or more outputs corresponding to the first portion of the content (e.g., as described above at FIG. 10E). Ceasing to provide one or more outputs corresponding to the content including pausing the one or more outputs corresponding to the content allows the computer system to promptly control its output, thereby providing improved visual feedback to the user and/or performing an operation when a set of conditions has been met without requiring further user input.
[0356] Note that details of the processes described above with respect to process 1400 (e.g., FIG. 14) are also applicable in an analogous manner to the methods described below/above. In some embodiments, process 1400 optionally includes one or more of the characteristics of the various methods described above with reference to process 1100. In some embodiments, the one or more outputs of process 1100 can be the first content of process 1400. For brevity, these details are not repeated below.
[0357] The description above has been described with reference to specific examples for the purpose of explanation. Such specific examples can be in the form of textual description above and/or in the accompanying drawings. However, such embodiments should not be interpreted as being exhaustive or limiting to the disclosure (e.g., limiting to the explicit manners described herein). Many modifications and variations are possible in view of the above teachings by one of ordinary skill in the art without departing from the scope of the present disclosure.
[0358] Aspects of the technology described above can include gathering and/or using data from various sources. Such data can include demographic data, telephone numbers, email addresses, location and/or location-related data, home addresses, work addresses, and/or any other identifying information. In some scenarios, such data can include personal information that is usable to uniquely identify a specific person. Such data can be used to improve interactions that a device has with its environment (e.g., interactions with users). The use of such data can require one or more entities handling such data. These entities can be involved in collecting, processing, disclosing, transferring, storing, or other functions that support the technologies described herein. The present disclosure expects that (e.g., does not preclude) that all use of such data complies with well-established privacy policies and/or privacy practices by such entities. As a general matter, such policies and practices should meet or exceed generally recognized industry standards and comply with all applicable data privacy and security-related governmental requirements. In particular, for example, entities should receive informed consent from users to collect and/or use such data, and such collection and/or use should only be for legitimate and reasonable uses. Further, such data should not be shared, disclosed, sold, and/or provided for uses other than legitimate and/or reasonable uses. Various scenarios can arise in which such data is not available, such as when a user selects not to share such data. For example, the user can withhold consent for collection and/or use of such data (e.g., “opt out” of sharing such data and/or not explicitly “opt in” during a registration process). The user can also employ the use of any of various hardware and/or software components that prevent collection and/or use of such data. While the use of such data can benefit a user by improving the operation of the device, the present disclosure contemplates that embodiments of the present technology can be used without such data. For example, operations of the device can use other data (e.g., instead of and/or in place of such data). Other techniques include making inferences based on other data or a minimal amount of such data. The use of such data can be utilized for the benefit of users of the device. For example, such data can be used to improve interactions that the device engages in with the user. Other benefits from the use for such data are also possible and within the scope of the present disclosure.

Claims

CLAIMS What is claimed is:
1. A method, comprising: at a computer system that is in communication with a display component and the one or more input devices: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
2. The method of claim 1, wherein the first set of one or more criteria includes a first criterion that is satisfied when a determination is made that a portion of the content has been displayed for at least a predetermined period of time.
3. The method of any one of claims 1-2, wherein the first set of one or more criteria includes a second criterion that is satisfied when a determination is made that the first user is looking in the direction of the content.
4. The method of any one of claims 1-3, wherein the first set of one or more criteria includes a third criterion that is satisfied when a determination is made that the first user should look at the content.
5. The method of any one of claims 1-4, wherein the second set of one or more criteria includes a fourth criterion that is satisfied when a determination is made that an input has been received.
6. The method of any one of claims 1-5, wherein the second set of one or more criteria includes a fifth criterion that is satisfied when a determination is made that the computer system is waiting for a response.
7. The method of claim 6, further comprising: in response to detecting the first interaction condition corresponding to the content and in accordance with a determination that the first interaction condition corresponding to the content satisfies a set of one or more criteria that includes a criterion that is satisfied when a determination is made that the computer system is not waiting for the response, displaying, via the display component, the representation of the face looking in a second direction that is different from the direction of the first user detected in the field-of-detection of the one or more input devices.
8. The method of claim 7, wherein the second direction corresponds to a second user detected in the field-of-detection of the one or more input devices, and wherein the second user is different from the first user.
9. The method of claim 7, wherein the second direction corresponds to the content.
10. The method of claim 7, wherein the second direction corresponds to an object in the field-of-detection of the one or more input devices.
11. The method of any one of claims 1-10, further comprising: while displaying the representation of the face looking in the direction of the content, detecting a second interaction condition corresponding to the content; and in response to detecting the second interaction condition corresponding to the content, displaying, via the display component, the representation of the face looking in the direction of the first user detected in the field-of-detection of the one or more input devices.
12. The method of any one of claims 1-11, further comprising: while displaying the representation of the face looking in the direction of the user detected in the field-of-detection of the one or more input devices, detecting a third interaction condition corresponding to the content; and in response to detecting the third interaction condition corresponding to the content, displaying, via the display component, the representation of the face looking in the direction of the content.
13. The method of any one of claims 1-12, further comprising: while displaying, via the display component, the content, detecting a fourth interaction condition corresponding to the content; and in response to detecting the fourth interaction condition corresponding to the content, displaying, via the display component, the representation of the face looking in a third direction that does not correspond to the first user and the content.
14. The method of any one of claims 1-13, further comprising: while displaying, via the display component, the representation of the face looking in a third direction, detecting a fifth interaction condition corresponding to the content; and in response to detecting the fifth interaction condition, continuing displaying, via the display component, the representation of the face looking in the third direction.
15. The method of any one of claims 1-14, further comprising: while displaying, via the display component, the content, detecting a sixth interaction condition corresponding the content; and in response to detecting the sixth interaction condition corresponding to the content: in accordance with a determination that the sixth interaction condition corresponding to the content satisfies the second set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of the first user detected in the field-of-detection of the one or more input devices; and in accordance with a determination that the sixth interaction condition corresponding to the content satisfies a fourth set of one or more criteria different from the first set of one or more criteria and the second set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a third user detected in the field-of-detection of the one or more input devices, wherein the third user is different from the first user.
16. The method of claim 15, wherein the computer system is in communication with a movement component, the method further comprising: in response to detecting the sixth interaction condition corresponding to the content and in accordance with a determination that the sixth interaction condition satisfies the second set of one or more criteria, moving, via the movement component, a portion of the computer system from a first position in the environment to a second position in the environment.
17. The method of claim 16, further comprising: while the portion of the computer system is in the second position, detecting a seventh interaction condition corresponding to the content; and in response to detecting the seventh interaction condition corresponding to the content and in accordance with a determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria: moving, via the movement component, the portion of the computer system from the second position in the environment to the first position in the environment; and displaying, via the display component, the representation of the face looking in the direction of the displayed content.
18. The method of claim 16, further comprising: while the portion of the computer system is in the second position, detecting an eighth interaction condition corresponding to the content; and in response to detecting the eighth interaction condition corresponding to the content: in accordance with a determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria, moving, via the movement component, the portion of the computer system from the second position in the environment to the first position in the environment while continuing to display, via the display component, the representation of the face looking in the direction of the user.
19. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices, the one or more programs including instructions for performing the method of any one of claims 1-18.
20. A computer system that is in communication with a display component and the one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1-18.
21. A computer system that is in communication with a display component and the one or more input devices, comprising: means for performing the method of any one of claims 1-18.
22. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices, the one or more programs including instructions for performing the method of any one of claims 1-18.
23. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices, the one or more programs including instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
24. A computer system that is in communication with a display component and the one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
25. A computer system that is in communication with a display component and the one or more input devices, comprising: means for, while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: means for, in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and means for, in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of- detection of the one or more input devices.
26. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and the one or more input devices, the one or more programs including instructions for: while displaying, via the display component, content, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: in accordance with a determination that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying, via the display component, a representation of a face looking in the direction of the content; and in accordance with a determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying, via the display component, the representation of the face looking in the direction of a first user detected in a field-of-detection of the one or more input devices.
27. A method, comprising: at a computer system that is in communication with a display component and one or more input devices: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
28. The method of claim 27, wherein continuing displaying the user interface without changing the user interface object in the manner that indicates eye contact with the user includes displaying, via the display component, the user interface object in a manner that does not indicate eye contact with the user.
29. The method of any one of claims 27-28, wherein changing the user interface object in the manner that indicates eye contact with the user includes: detecting, via the one or more input devices, a position of the user in a field-of- detection of the computer system; and changing a first portion of the user interface object to be directed to the detected position of the user in the field-of-detection of the computer system.
30. The method of claim 29, wherein changing the first portion of the user interface object to be directed to the detected position of the user in the field-of-detection includes moving, via the display component, the first portion of the user interface object from a position that is more than a predetermined distance from the detected position of the user to a position that is no more than the predetermined distance from the detected position of the user in the field-of-detection.
31. The method of any one of claims 29-30, wherein changing the first portion of the user interface object to be directed to the detected position of the user in the field-of-detection includes displaying, via the display component, the first portion of the user interface object pointing to the position of the user detected in the field-of-detection.
32. The method of any one of claims 27-31, wherein the first input does not include an explicit indication to change the user interface object.
33. The method of any one of claims 27-32, wherein displaying the user interface while changing the user interface object in the manner that indicates eye contact with the user includes moving, via the display component, a second portion of the user interface object.
34. The method of any one of claims 27-33, further comprising: while displaying the user interface object in the manner that indicates eye contact with the user, detecting, via the one or more input devices, a second input; and in response to detecting the second input, continuing to display the user interface while changing, via the display component, the user interface object in a first manner that does not indicate eye contact with the user.
35. The method of any one of claims 27-34, further comprising: while displaying the user interface object in the manner that indicates eye contact with the user, detecting, via the one or more input devices, a third input; and in response to detecting the third input: in accordance with a determination that an agreement was not made with respect to the third input, continuing displaying, via the display component, the user interface while changing the user interface object in a second manner that does not indicate eye contact with the user; and in accordance with a determination that the agreement was made with respect to the third input, continuing displaying, via the display component, the user interface without changing the user interface object in the second manner that does not indicate eye contact with the user.
36. The method of any one of claims 27-35, further comprising: while displaying the use interface object in the manner that indicates eye contact with the user, detecting, via the one or more input devices, a fourth input; and in response to detecting the fourth input and in accordance with a determination that a disagreement was made with respect to the fourth input, continuing displaying, via the display component, the user interface object in the manner that indicates eye contact with the user.
37. The method of any one of claims 27-36, further comprising: while displaying the user interface object in the manner that indicates eye contact with the user, detecting, via the one or more input devices, a fifth input; and in response to detecting the fifth input in accordance with a determination that a disagreement was made with respect to the fifth input, displaying, via the display component, the user interface object in a third manner that does not indicate eye contact with the user.
38. The method of any one of claims 27-37, further comprising: in response to detecting the first input, continuing to display, via the display component, a portion of the user interface that does not include the user interface object while changing the user interface object in the manner that indicates eye contact with the user.
39. The method of any one of claims 27-38, further comprising: after changing the user interface object in the manner that indicates eye contact with the user: in accordance with a determination that a predetermined period of time has not passed since the user interface object was changed in the manner that indicates eye contact with the user, continuing to display, via the display component, the user interface object in the manner that indicates eye contact with the user; and in accordance with a determination that the predetermined period of time has passed since the user interface object was changed in the manner that indicates eye contact with the user, forgoing continuing to display, via the display component, the user interface object in the manner that indicates eye contact with the user.
40. The method of any one of claims 27-39, wherein: the first input includes a question; the determination that the agreement was made with respect to the first input includes a determination that the question corresponds to a positive response; and the determination that the agreement was not made with respect to the first input includes a determination that the question corresponds to a negative response.
41. The method of any one of claims 27-40, wherein: the first input includes a statement; the determination that the agreement was made with respect to the first input includes a determination that the statement corresponds to a factual statement; and the determination that the agreement was not made with respect to the first input includes a determination that the statement does not correspond to the factual statement.
42. The method of any one of claims 27-41, wherein the determination that the agreement was made with respect to the first input includes a determination that the first input satisfies a user preference of the first user, and wherein the determination that the agreement was not made with respect to the first input includes a determination that the first input does not satisfy the user preference of the first user.
43. The method of any one of claims 27-42, further comprising: in response to detecting the first input via the one or more input devices, providing, via one or more output devices in communication with the computer system, an output corresponding to the first input, wherein: in accordance with a determination that an agreement was made with respect to the first input, the output includes a first phrase that indicates agreement with a user preference; and in accordance with a determination that an agreement was not made with respect to the first input, the output includes a second phrase that indicates disagreement with the user preference.
44. The method of claim 43, wherein continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user does not include outputting the second phrase.
45. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 27-44.
46. A computer system that is in communication with a display component and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 27-44.
47. A computer system that is in communication with a display component and one or more input devices, comprising: means for performing the method of any one of claims 27-44.
48. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 27-44.
49. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
50. A computer system that is in communication with a display component and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
51. A computer system that is in communication with a display component and one or more input devices, comprising: means for, while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: means for, in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and means for, in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
52. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for: while displaying, via the display component, a user interface that includes a user interface object representing a portion of a person, detecting, via the one or more input devices, a first input; and in response to detecting the first input: in accordance with a determination that an agreement was made with respect to the first input, continuing displaying, via the display component, the user interface while changing the user interface object in a manner that indicates eye contact with a user; and in accordance with a determination that an agreement was not made with respect to the first input, continuing displaying, via the display component, the user interface without changing the user interface object in the manner that indicates eye contact with the user.
53. A method, comprising: at a computer system that is in communication with a display component, a camera, and one or more input devices: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
54. The method of claim 53, wherein: before detecting the request to interact with the content, displaying a portion of the user interface object representing the portion of the user at a first location; displaying the user interface object representing the portion of the user at the second size smaller than the first size in response to detecting the request includes ceasing to display the portion of the user interface object at the first location; and the first portion of the content is displayed at the first location in response to detecting the request to interact with the content.
55. The method of any one of claims 53-54, further comprising: before detecting the request to interact with the content, forgoing displaying at least the portion of the content.
56. The method of any one of claims 53-55, wherein the one or more input devices includes a microphone, and wherein detecting the request to interact with the content includes receiving, via the microphone, a voice input corresponding to the content.
57. The method of any one of claims 53-56, wherein the request to interact with the content does not include an explicit request.
58. The method of any one of claims 53-57, wherein the request to interact with content is a first request to interact with the first portion of the content, the method further comprising: while displaying at least the first portion of the content and the user interface object representing the portion of the user at the second size, detecting, via the one or more input devices, a request to interact with a second portion of the content different from the first portion of the content; and in response to detecting the request to interact with the second portion of content: displaying, via the display component, at least the second portion of the content, wherein the second portion of the content replaces the first portion of the content; and continuing to display, via the display component, the user interface object representing the portion of the user at the second size and in the second manner.
59. The method of claim 53-58, further comprising: after displaying the user interface object representing the portion of the user at the second size and in the second manner to indicate that the user interface object is directed to the portion of the content and in accordance with a determination that a first set of one or more criteria is satisfied, displaying, via the display component, the user interface object representing the portion of the user at the second size and in a third manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system.
60. The method of claim 59, further comprising: after displaying the user interface object representing the portion of the user at the second size and in the second manner to indicate that the user interface object is directed to the portion of content and in accordance to the determination that the first set of one or more criteria is not satisfied, continuing displaying, via the display component, the user interface object representing the user at the second size and in the second manner to indicate that the user interface object is directed to the portion of the content.
61. The method of any one of claims 53-60, wherein displaying the user interface object representing the portion of the user at the second size in the second manner to indicate that the user interface object is directed to the portion of content includes changing the user interface object representing the portion of the user to the second size before displaying the user interface object representing the portion of the user in the second manner to indicate that the user interface object is directed to the portion of content.
62. The method of any one of claims 53-61, wherein: displaying, via the display component, the user interface object representing the portion of the user in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection includes displaying, via the display component, a representation of eyes displayed with a first set of one or more characteristics; and displaying, via the display component, the user interface object representing the portion of the user in the second manner to indicate that the user interface object is directed to the portion of content includes displaying, via the display component, the representation of the eyes with a second set of characteristics different from the first set of one or more characteristics.
63. The method of any one of claims 53-63, further comprising: after displaying at least the first portion of the content and in accordance with a determination that content should be removed: ceasing displaying at least the first portion of the content; and re-displaying, via the display component, the user interface object representing the portion of the user at the first size.
64. The method of claim 63, wherein, after displaying at least the first portion of the content and in accordance with a determination that content should be removed and a third set of one or more criteria is satisfied, the user interface object is displayed in the first manner to indicate that the user interface object is directed to the first entity in the field-of-detection of the computer system.
65. The method of any one of claims 62-63, wherein, after displaying at least the first portion of the content and in accordance with a determination that content should be removed and a fourth set of one or more criteria is satisfied, the user interface object is displayed a fourth manner to indicate that the user interface object is not directed to the first entity.
66. The method of any one of claims 53-65, further comprising: while displaying the user interface object representing the portion of the user in the second manner, detecting an interaction condition corresponding to the content; and in response to detecting the interaction condition: in accordance with a determination that the interaction condition satisfies a fifth set of one or more criteria, displaying, via the display component, the user interface object representing the portion of the user in the first manner to indicate that the user interface object is directed to the first entity in the field-of-view of the camera, wherein the first entity initiated the interaction condition; and in accordance with a determination that the interaction condition corresponding to the content satisfies a sixth set of one or more criteria different from the fifth set of one or more criteria, displaying, via the display component, the user interface object representing the portion of the user in a manner to indicate that the user interface object is directed to a second entity detected in the field-of-view of the computer system, wherein the second entity initiated the interaction condition.
67. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 53- 66.
68. A computer system that is in communication with a display component, a camera, and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 53-66.
69. A computer system that is in communication with a display component, a camera, and one or more input devices, comprising: means for performing the method of any one of claims 53-66.
70. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 53-66.
71. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
72. A computer system that is in communication with a display component, a camera, and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
73. A computer system that is in communication with a display component, a camera, and one or more input devices, comprising: means for, while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; means for, while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: means for displaying at least a first portion of the content; and means for displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
74. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for: while detecting a first entity in the field-of-view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed to the first user in the field-of-detection of the computer system, wherein the user interface object is displayed at a first size; while displaying a user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed to the portion of the content.
75. A method, comprising: at a computer system that is in communication with a display component and one or more input devices: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
76. The method of claim 75, further comprising: after outputting the first content and without detecting input via the one or more input devices: in accordance with the determination that the second content is to be output and that the second content corresponds to the second system avatar, ceasing displaying, via the display component, the first system avatar.
77. The method of any one of claims 75-76, wherein the first system avatar includes a first expression, and wherein the second system avatar includes the first expression.
78. The method of any one of claims 75-77, wherein: the computer system is in communication with one or more audio output devices; outputting the first content includes outputting, via the one or more audio output devices, audio corresponding to the first system avatar in a first voice; and outputting the second content includes outputting, via the one or more audio output devices, audio corresponding to the second system avatar in a second voice different from the first voice.
79. The method of any one of claims 75-78, wherein the first system avatar has a first appearance, and wherein the second system avatar has a second appearance different from the first appearance.
80. The method of any one of claims 75-79, wherein the first system avatar is a first size, and wherein the second system avatar is the first size.
81. The method of any one of claims 75-79, wherein the first system avatar is a second size, and wherein the second system avatar is a third size different from the second size.
82. The method of any one of claims 75-81, wherein the first content corresponds to the first system avatar.
83. The method of any one of claims 75-82, wherein the first content corresponds to a first application, the method further comprising: while displaying the first system avatar, outputting fourth content corresponding to a second application different from the first application.
84. The method of any one of claims 75-83, wherein the second content corresponds to a third application, the method further comprising: while displaying the second system avatar, outputting fifth content, corresponding to a fourth application different from the third application.
85. The method of any one of claims 75-84, further comprising: after displaying the second system avatar: in accordance with a determination that sixth content different from the first content is to be output and that the sixth content does not correspond to the second system avatar, displaying, via the display component, the first system avatar.
86. The method of any one of claims 75-85, further comprising: after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that seventh content different from the first content is to be output and that the seventh content corresponds to a third system avatar different from the first system avatar and the second system avatar, displaying, via the display component, the third system avatar.
87. The method of any one of claims 75-86, wherein: the first system avatar includes a first set of movement patterns; and the second system avatar includes a second set of movement patterns different from the first set of movement patterns.
88. The method of any one of claims 75-87, further comprising: while displaying the first system avatar and outputting the first content, synchronizing movement of the first system avatar with the first content.
89. The method of any one of claims 75-88, further comprising: while displaying the second system avatar and outputting the second content, synchronizing movement of the second system avatar with the second content.
90. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 75-89.
91. A computer system that is in communication with a display component and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 75-89.
92. A computer system that is in communication with a display component and one or more input devices, comprising: means for performing the method of any one of claims 75-89.
93. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 75-89.
94. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
95. A computer system that is in communication with a display component and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
96. A computer system that is in communication with a display component and one or more input devices, comprising: means for, displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; means for, while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: means for, in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and means for, in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
97. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component and one or more input devices, the one or more programs including instructions for: displaying, via the display component, a first system avatar, wherein the first system avatar corresponds to a first character; while displaying the first system avatar, outputting first content; and after outputting the first content and without detecting input via the one or more input devices: in accordance with a determination that second content different from the first content is to be output and that the second content corresponds to a second system avatar different from the first system avatar, displaying, via the display component, the second system avatar, wherein the second system avatar corresponds to a second character different from the first character; and in accordance with a determination that third content different from the first content is to be output and that the third content corresponds to the first system avatar, continuing displaying, via the display component, the first system avatar without displaying, via the display component, the second system avatar.
98. A method, comprising: at a computer system that is in communication with a display component, an audio generation component, and a movement component: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
99. The method of claim 98, wherein: while outputting a portion of the audio content: in accordance with a determination that the portion of the audio content is being output and that the portion corresponds to a first user, the portion of the computer system physically moves in a first manner; and in accordance with a determination that the portion of the audio content is being output and that the portion does not correspond to the first user, the portion of the computer system physically moves in a second manner different from the first manner.
100. The method of any one of claims 98-99, wherein the audio content is first audio content, the method further comprising: after ceasing physically moving, via the movement component, the portion of the computer system, outputting, via the audio generation component, second audio content that is different from the first audio content; and while outputting, via the audio generation component, the second audio content, forgoing physically moving, via the movement component, the portion of the computer system.
101. The method of any one of claims 98-100, wherein the audio content is third audio content, the method further comprising: after ceasing physically moving, via the movement component, the portion of the computer system, outputting, via the audio generation component, fourth audio content that is different from the third audio content; and while outputting, via the audio generation component, the fourth audio content, physically moving, via the movement component, the portion of the computer system.
102. The method of any one of claims 98-101, wherein the visual content is first visual content, the method further comprising: while physically moving, via the movement component, the portion of the computer system, displaying, via the display component, second visual content different from the first visual content.
103. The method of claim 102, wherein the request to display the visual content is a request to change from displaying the second visual content to displaying different visual content.
104. The method of any one of claims 98-103, wherein the portion of the computer system is moved according to one or more characteristics of the audio content output by the computer system.
105. The method of any one of claims 98-104, wherein the audio content is fifth audio content, wherein the visual content is third visual content, the method further comprising: after ceasing physically moving, via the movement component, the portion of the computer system, outputting, via the audio generation component, sixth audio content different from the fifth audio content while physically moving, via the movement component, the portion of the computer system; while outputting, via the audio generation component, the sixth audio content and physically moving the portion of the computer system, detecting a request to display fourth visual content, different from the third visual content; and in response to detecting the request to display the fourth visual content: in accordance with a determination that the fourth visual content is a first type of content, ceasing moving, via the movement component, the portion of the computer system; and in accordance with a determination that the fourth visual content is a second type of content, different from the first type of content, continuing physically moving, via the movement component, the portion of the computer system.
106. The method of any one of claims 98-105, wherein the visual content is a representation of an avatar.
107. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component, the one or more programs including instructions for performing the method of any one of claims 98-106.
108. A computer system that is in communication with a display component, an audio generation component, and a movement component, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 98-106.
109. A computer system that is in communication with a display component, an audio generation component, and a movement component, comprising: means for performing the method of any one of claims 98-106.
110. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component, the one or more programs including instructions for performing the method of any one of claims 98- 106.
111. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component, the one or more programs including instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
112. A computer system that is in communication with a display component, an audio generation component, and a movement component, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
113. A computer system that is in communication with a display component, an audio generation component, and a movement component, comprising: means for, outputting, via the audio generation component, audio content; means for, while outputting the audio content, physically moving, via the movement component, a portion of the computer system; means for, while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: means for ceasing physically moving, via the movement component, the portion of the computer system; and means for, displaying, via the display component, the visual content.
114. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display component, an audio generation component, and a movement component, the one or more programs including instructions for: outputting, via the audio generation component, audio content; while outputting the audio content, physically moving, via the movement component, a portion of the computer system; while physically moving, via the movement component, the portion of the computer system, detecting a request to display visual content; and in response to detecting the request to display the visual content: ceasing physically moving, via the movement component, the portion of the computer system; and displaying, via the display component, the visual content.
115. A method, comprising: at a computer system that is in communication with a one or more output devices and one or more input devices: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
116. The method of claim 115, wherein the one or more output devices includes an audio output device, and wherein at least one of the one or more outputs corresponding to the second portion of content is audio output provided via the audio output device.
117. The method of any one of claims 115-116, wherein the one or more output devices includes a first display component, and wherein providing the one or more outputs corresponding to the second portion of content includes displaying, via the first display component, a representation corresponding to the second portion of content.
118. The method of any one of claims 115-117, wherein the one or more output devices includes a movement component, and wherein providing the one or more outputs includes moving via the movement component.
119. The method of any one of claims 115-118, wherein the first portion of content is at a fourth position in the content, and wherein the second portion of content is at a fifth position in the content that is before the fourth position.
120. The method of any one of claims 115-119, wherein the first portion of content is at a sixth position in the content, and wherein the second portion of content is at a seventh position in the content that is after the sixth position.
121. The method of any one of claims 115-120, wherein the voice input does not include an explicit indication of the second portion of content.
122. The method of any one of claims 115-121, wherein the first portion is in a first subset of the content and the second portion is in the first subset of the content.
123. The method of any one of claims 115-122, wherein the first portion is in a second section of the content and the second portion is in a third section of the content different from the second section of the content.
124. The method of any one of claims 115-123, wherein the description of one or more attributes of content corresponding to the content includes a description of a portion of a plot of the content, and wherein the output corresponding to the second portion of the content concerns the portion of the plot of the content.
125. The method of any one of claims 115-124, wherein the description of one or more attributes of content corresponding to the content includes a description of a portion of a theme of the content, and wherein the output corresponding to the second portion of content concerns the portion of the theme of the content.
126. The method of any one of claims 115-125, wherein the description of one or more attributes of content corresponding to the content includes a description of a character’s arc in the content, and wherein the output corresponding to the second portion of content concerns the character’s arc in the content.
127. The method of any one of claims 115-126, wherein the description of one or more attributes of content corresponding to the content includes a description of one or more characteristics in the content, and wherein the output corresponding to the second portion of content concerns the one or more characteristics in the content.
128. The method of any one of claims 115-127, further comprising: after outputting, via the one or more output devices, one or more outputs corresponding to the second portion of content, providing, via the one or more input devices, one or more outputs corresponding to a portion of content that immediately follows the second portion of content.
129. The method of any one of claims 115-128, wherein: in accordance with the determination that the description of one or more attributes of content corresponding to the content included in the voice input is before one or more attributes of content to which the first portion of content concerns, the second portion of content is at a position in the content that is before a position of the first portion of content in the content; and in accordance with the determination that the description of one or more attributes of content corresponding to the content included in the voice input is after one or more attributes of content to which the first portion of content concerns, the second portion of content is at a position in the content that is after the position of the first portion of content in the content.
130. The method of any one of claims 115-129, further comprising: in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content: in accordance with a determination that the one or more attributes of content corresponding to the content includes a first event, selecting a first respective portion of the content as the second portion of content; and in accordance with a determination that the one or more attributes of content corresponding to the content includes a second event, different from the first event, selecting a second respective portion of the content, different from the first respective portion of the content, as the second portion of content.
131. The method of any one of claims 115-130, wherein the voice output is first voice output: while providing, via the one or more output devices, one or more outputs corresponding to the first portion of content, detecting, via one or more input devices, second voice input different from the first voice input; and in response to detecting the second voice input, continuing to provide output corresponding to the first portion of content.
132. The method of any one of claims 115-131, further comprising: in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, ceasing to provide, via the one or more output devices, the one or more outputs corresponding to the first portion of content.
133. The method of any one of claims 115-132, wherein the voice input corresponds to a request to skip the third portion of content.
134. The method of any one of claims 115-133, wherein the voice input corresponds to a request to jump to the second portion of content.
135. The method of any one of claims 115-134, wherein the voice input corresponds to a request to jump from the first portion of content.
136. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 115- 135.
137. A computer system that is in communication with a one or more output devices and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 115-135.
138. A computer system that is in communication with a one or more output devices and one or more input devices, comprising: means for performing the method of any one of claims 115-135.
139. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 115-135.
140. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
141. A computer system that is in communication with a one or more output devices and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
142. A computer system that is in communication with a one or more output devices and one or more input devices, comprising: means for, while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and means for, in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
143. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with a one or more output devices and one or more input devices, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via one or more input devices, voice input that includes a description of one or more attributes of content corresponding to the content; and in response to detecting the voice input that includes the description of the one or more attributes of content corresponding to the content, providing, via the one or more output devices, one or more outputs corresponding to a second portion of content without providing output corresponding to a third portion of content that is between the first portion of content and the second portion of content, wherein the second portion of content is different from the first portion of content.
144. A method, comprising: at a computer system that is in communication with one or more output devices and one or more input devices: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
145. The method of claim 144, wherein the one or more output devices includes an audio output device, and wherein providing the one or more outputs corresponding to the first portion of the content includes outputting, via the audio output device, audio.
146. The method of any one of claims 144-145, wherein the one or more output devices includes a first display component, and wherein providing the one or more outputs corresponding to the first portion of the content includes displaying, via the first display component, a representation.
147. The method of any one of claims 144-146, wherein the one or more output devices includes a movement component, and wherein providing the one or more outputs corresponding to the first portion of the content includes moving, via the movement component, a first portion of the computer system.
148. The method of any one of claims 144-147, wherein the second portion of the content is positioned in the content at a position that is not at the beginning of a first subset of the content.
149. The method of any one of claims 144-148, wherein the first portion of the content is positioned in the content at a position that is at the beginning of a second subset of the content.
150. The method of any one of claims 144-149, wherein: the first portion is positioned in the content at a first position that is in a third subset in the content; the second portion is positioned in the content at a second position that is in the third subset of the content; the first position is different from the second position; and the first position and the second position are not at a terminal position of the third subset of the content.
151. The method of any one of claims 144-150, further comprising: before providing the one or more outputs corresponding to the first portion of the content, providing, via the one or more output devices, the one or more outputs corresponding to the second portion of the content while detecting that the attention of the user corresponds to the computer system.
152. The method of any one of claims 144-151, further comprising: before detecting that the attention of the user no longer corresponds to the computer system, detecting, via the one or more input devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, the one or more one or more outputs corresponding to the first portion of the content after providing the one or more outputs corresponding to the second portion of the content.
153. The method of any one of claims 144-152, wherein the computer system is in communication with a movement component, the method further comprising: while outputting audio corresponding to the content, detecting a condition; and in response to detecting the condition: in accordance with a determination that the condition includes detecting voice input while outputting audio corresponding to a respective scene, moving, via the movement component, in a first manner; and in accordance with a determination that the condition does not include detecting voice input while outputting audio corresponding to the respective scene, moving, via the movement component, in a second manner different from the first manner.
154. The method of any one of claims 144-153, wherein the user is a first user, the method further comprising: while providing, via one or more output devices, one or more outputs corresponding to the content, detecting that an attention of a second user no longer corresponds to the computer system while an attention of a third user corresponds to the computer system, wherein the third user is different from the second user; and in response to detecting that the attention of the second user no longer corresponds to the computer system while the attention of the third user corresponds to the computer system, continuing to provide one or more outputs corresponding to the content.
155. The method of any one of claims 144-154, wherein the user is a fourth user, the method further comprising: while providing, via one or more output devices, one or more outputs corresponding to the content, detecting that an attention of a fifth user no longer corresponds to the computer system while detecting that an attention of a sixth user corresponds to the computer system, wherein the sixth user is different from the fifth user; and in response to detecting that the attention of the fifth user no longer corresponds to the computer system while detecting that the attention of the sixth user corresponds to the computer system: in accordance with a determination that the fifth user is a first type of user, continuing to provide one or more outputs corresponding to the content; and in accordance with a determination that the fifth user is a second type of user different from the first type of user, forgoing continuing to provide one or more outputs corresponding to the content.
156. The method of any one of claims 144-155, wherein ceasing to provide the one or more outputs corresponding to the first portion of the content includes gradually de- emphasizing, via the one or more output devices, the one or more outputs corresponding to the first portion of the content.
157. The method of any one of claims 144-156, wherein ceasing to provide one or more outputs corresponding to the first portion of the content includes pausing, via the one or more output devices, the one or more outputs corresponding to the first portion of the content.
158. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 144- 157.
159. A computer system that is in communication with one or more output devices and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 144-157.
160. A computer system that is in communication with one or more output devices and one or more input devices, comprising: means for performing the method of any one of claims 144-157.
161. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 144-157.
162. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
163. A computer system that is in communication with one or more output devices and one or more input devices, comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
164. A computer system that is in communication with one or more output devices and one or more input devices, comprising: means for, while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; means for, in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; means for, while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and means for, in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
165. A computer program product, comprising one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more output devices and one or more input devices, the one or more programs including instructions for: while providing, via the one or more output devices, one or more outputs corresponding to a first portion of content, detecting, via the one or more inputs devices, that an attention of a user no longer corresponds to the computer system; in response to detecting that the attention of the user no longer corresponds to the computer system, ceasing to provide the one or more outputs corresponding to the first portion of the content; while the one or more outputs corresponding to the first portion of the content are not being provided, detecting, via the one or more inputs devices, that the attention of the user corresponds to the computer system; and in response to detecting that the attention of the user corresponds to the computer system, providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content that is at or before the first portion of the content.
EP24790048.3A 2023-09-30 2024-09-25 User interfaces and techniques for changing how an object is displayed Pending EP4728351A1 (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US202363541832P 2023-09-30 2023-09-30
US202363541811P 2023-09-30 2023-09-30
PCT/US2024/048481 WO2025072385A1 (en) 2023-09-30 2024-09-25 User interfaces and techniques for changing how an object is displayed

Publications (1)

Publication Number Publication Date
EP4728351A1 true EP4728351A1 (en) 2026-04-22

Family

ID=93100430

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24790048.3A Pending EP4728351A1 (en) 2023-09-30 2024-09-25 User interfaces and techniques for changing how an object is displayed

Country Status (2)

Country Link
EP (1) EP4728351A1 (en)
WO (1) WO2025072385A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020017981A1 (en) * 2018-07-19 2020-01-23 Soul Machines Limited Machine interaction
JP7805785B2 (en) * 2020-05-13 2026-01-26 エヌビディア コーポレーション Conversational AI platform that utilizes rendered graphical output

Also Published As

Publication number Publication date
WO2025072385A1 (en) 2025-04-03

Similar Documents

Publication Publication Date Title
US11398067B2 (en) Virtual reality presentation of body postures of avatars
US10489960B2 (en) Virtual reality presentation of eye movement and eye contact
EP3381175B1 (en) Apparatus and method for operating personal agent
RU2749000C2 (en) Systems and methods for animated head of a character
US20180133900A1 (en) Embodied dialog and embodied speech authoring tools for use with an expressive social robot
WO2024080135A1 (en) Display control device, display control method, and display control program
JP2019124855A (en) Apparatus and program and the like
US20260056604A1 (en) User interfaces and techniques for responding to notifications
EP4728351A1 (en) User interfaces and techniques for changing how an object is displayed
KR20260057173A (en) User interfaces and techniques for changing how objects are displayed
US20260072639A1 (en) User interfaces for updating an indication of an activity
US20260072954A1 (en) User interfaces and techniques for interactions
US20260056603A1 (en) User interfaces and techniques for moving a computer system
WO2025072328A1 (en) User interfaces and techniques for performing an operation based on learned characteristics
US20260064236A1 (en) User interfaces and techniques for managing content
KR20260055473A (en) User interfaces and techniques for interactions
US12598363B2 (en) Systems and methods for providing sexual entertainment by monitoring target elements
JP7760554B2 (en) Control System
KR20260061263A (en) User interfaces and techniques for performing actions based on learned characteristics
US20260050322A1 (en) User interfaces and techniques for presenting content
WO2025188634A1 (en) Techniques for capturing media
KR20260055463A (en) User interfaces and techniques for managing content
WO2025260106A2 (en) Techniques for outputting content
CN120960793A (en) Information processing method and device in game, electronic equipment and readable storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20260119

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR