WO2025249802A1 - 멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 - Google Patents
멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체Info
- Publication number
- WO2025249802A1 WO2025249802A1 PCT/KR2025/006520 KR2025006520W WO2025249802A1 WO 2025249802 A1 WO2025249802 A1 WO 2025249802A1 KR 2025006520 W KR2025006520 W KR 2025006520W WO 2025249802 A1 WO2025249802 A1 WO 2025249802A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- electronic device
- audio signal
- data
- processor
- image
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/048—Interaction techniques based on graphical user interfaces [GUI]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/08—Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W4/00—Services specially adapted for wireless communication networks; Facilities therefor
- H04W4/80—Services using short range communication, e.g. near-field communication [NFC], radio-frequency identification [RFID] or low energy communication
Definitions
- the following descriptions relate to electronic devices, methods, and computer-readable storage media for utilizing multimodal models.
- An electronic device may include a microphone.
- the electronic device may acquire an audio signal through the microphone.
- the audio signal may include speech or voice uttered by a speaker.
- the electronic device may provide a function for recognizing the voice within the audio signal.
- the electronic device may include a communication circuit.
- the electronic device may transmit data to and/or receive data from an external electronic device via the communication circuit.
- the electronic device may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for driving the image sensor.
- the instructions, when individually or collectively executed by the at least one processor may cause the electronic device to acquire a first image through the image sensor driven in response to the event, and to acquire a first audio signal through the microphone, based on the detection.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, a signal to an external electronic device, wherein the signal causes a wake-up of a second multimodal model within the external electronic device based on identifying, using a first multimodal model within the electronic device, the first image representing a reference gesture of the user and the first audio signal including a reference voice command.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device, after transmitting the signal, to acquire, via the microphone, a second audio signal and to acquire, via the image sensor, a second image.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, first data acquired using the second audio signal and second data regarding a visual object within the second image corresponding to the user's lips, to the external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information for the second audio signal, obtained using the second multimodal model while in a wake-up state in response to the signal, from the external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.
- a method is provided.
- the method can be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit.
- the method can include detecting an event for driving the image sensor.
- the method can include acquiring a first image through the image sensor driven according to the event and acquiring a first audio signal through the microphone, based on the detection.
- the method can include transmitting a signal to an external electronic device via the communication circuit, the signal causing a wake-up of a second multimodal model within the external electronic device, based on identifying the first image expressing a reference gesture of a user and the first audio signal including a reference voice command, using a first multimodal model within the electronic device.
- the method can include, after transmitting the signal, acquiring a second audio signal through the microphone and acquiring a second image through the image sensor.
- the method may include an operation of transmitting, to the external electronic device, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips, via the communication circuit.
- the method may include an operation of receiving, from the external electronic device, response information regarding the second audio signal obtained using the second multimodal model in a wake-up state according to the signal, via the communication circuit.
- the method may include an operation of performing a function according to the response information.
- a non-transitory computer-readable storage medium may store one or more programs.
- the one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit, cause the electronic device to detect an event to drive the image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image through the image sensor driven according to the event, and to acquire a first audio signal through the microphone, based on the detection.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a second multimodal model within the external electronic device based on identifying, using a first multimodal model within the electronic device, the first image representing a reference gesture of the user and the first audio signal including a reference voice command.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, after transmitting the signal, acquire a second audio signal through the microphone and acquire a second image through the image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuitry, first data acquired using the second audio signal and second data regarding a visual object within the second image corresponding to the user's lips, to the external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, response information for the second audio signal obtained using the second multimodal model in a wake-up state in response to the signal from the external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function in response to the response information.
- the electronic device may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first audio signal via the microphone.
- the instructions, when individually or collectively executed by the at least one processor may cause the electronic device to identify a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to the external electronic device, in accordance with a reference voice command included in the first audio signal, to wake up a first multimodal model within the external electronic device based on identifying the type of the environment as a first type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor based on identifying the type of the environment as a second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image via the driven image sensor.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a second audio signal via the microphone.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, to the external electronic device, the signal causing a wake-up of the first multimodal model based on identifying, using a second multimodal model within the electronic device, the first image representing a reference gesture of the user and the second audio signal including the reference voice command.
- a method is provided.
- the method may be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry.
- the method may include acquiring a first audio signal via the microphone.
- the method may include identifying a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device.
- the method may include transmitting, via the communication circuitry, a signal to an external electronic device, based on identifying that the type of environment is the first type, a wake-up signal for a first multimodal model within the external electronic device, in accordance with a reference voice command included in the first audio signal.
- the method may include driving the image sensor based on identifying that the type of environment is the second type.
- the method may include acquiring a first image via the driven image sensor.
- the method may include acquiring a second audio signal via the microphone.
- the method may include an operation of transmitting, to the external electronic device through the communication circuit, a signal causing a wake-up of the first multimodal model based on identifying the first image representing the user's reference gesture and the second audio signal including the reference voice command using a second multimodal model within the electronic device.
- a non-transitory computer-readable storage medium may store one or more programs.
- the one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry, cause the electronic device to acquire a first audio signal via the microphone.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the first audio signal, a type of environment in which the electronic device is located, using a model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, via the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a first multimodal model within the external electronic device, based on identifying that the type of environment is the first type, in accordance with a reference voice command included in the first audio signal.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to drive the image sensor based on identifying that the type of the environment is a second type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image via the driven image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a second audio signal via the microphone.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, via the communication circuit, to the external electronic device, the signal that causes a wake-up of the first multimodal model based on identifying, using a second multimodal model within the electronic device, the first image representing a reference gesture of the user and the second audio signal including the reference voice command.
- the electronic device may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for a call connection with an external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first audio signal through the microphone based on performing the call connection with the external electronic device based on the detection, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to, based on identifying that the type of the environment is a first type, acquire a second audio signal via the microphone, and transmit the second audio signal, via the communication circuit, to the external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to, based on identifying that the type of the environment is a second type, drive the image sensor, acquire an image via the driven image sensor, acquire a third audio signal via the microphone, and provide first data acquired using the third audio signal and second data regarding a visual object in the image corresponding to the user's lips to a multimodal model within the electronic device, thereby acquiring response information regarding the third audio signal and the visual object, and performing a function according to the response information.
- a method is provided.
- the method can be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit.
- the method can include detecting an event for a call connection with an external electronic device.
- the method can include, based on the detection, performing the call connection with the external electronic device, acquiring a first audio signal through the microphone, and using a model within the electronic device, identifying a type of environment in which the electronic device is located from the first audio signal through the microphone.
- the method can include, based on identifying that the type of the environment is the first type, acquiring a second audio signal through the microphone, and transmitting the second audio signal to the external electronic device through the communication circuit.
- the method may include an operation of driving the image sensor based on identifying that the type of the environment is the second type, acquiring an image through the driven image sensor, acquiring a third audio signal through the microphone, and providing first data acquired using the third audio signal and second data about a visual object in the image corresponding to the user's lips to a multimodal model in the electronic device, thereby acquiring response information about the third audio signal and the visual object, and performing a function according to the response information.
- a non-transitory computer-readable storage medium may store one or more programs.
- the one or more programs may include instructions that, when executed by an electronic device having an image sensor, a microphone, and a communication circuit facing a front side of the electronic device, cause the electronic device to detect an event for a call connection with an external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first audio signal through the microphone based on the detection of performing the call connection with the external electronic device, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, based on identifying that the type of the environment is a first type, acquire a second audio signal through the microphone, and transmit the second audio signal to the external electronic device through the communication circuit.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, based on identifying that the type of the environment is a second type, drive the image sensor, acquire an image through the driven image sensor, acquire a third audio signal through the microphone, and provide first data acquired using the third audio signal and second data regarding a visual object in the image corresponding to the user's lips to a multimodal model within the electronic device, thereby acquiring response information regarding the third audio signal and the visual object, and performing a function according to the response information.
- a wearable device may include at least one display.
- the wearable device may include a microphone.
- the wearable device may include at least one sensor.
- the wearable device may include a memory including one or more storage media storing instructions.
- the wearable device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through pass-through.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on the first direction of the gaze of the user and the second direction of the gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, text generated from the voice signal, superimposed on the screen, based on identifying that the user and the other user corresponding to the visual object are viewing each other.
- a method is provided.
- the method can be performed in a wearable device having at least one display, a microphone, and at least one sensor.
- the method can include an operation of displaying a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through a pass-through through the at least one display.
- the method can include an operation of identifying a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen based on an audio signal acquired through the microphone while the screen is displayed.
- the method can include an operation of identifying a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal through the at least one sensor.
- the method can include an operation of identifying whether the user and a speaker corresponding to the visual object are looking at each other based on the first direction of the gaze of the user and the second direction of the gaze of the visual object.
- the method may include an action of displaying text generated from the voice signal by overlaying it on the screen through the at least one display based on identifying that the user and the other user corresponding to the visual object are seeing each other.
- a non-transitory computer-readable storage medium may store one or more programs.
- the one or more programs may include instructions that, when executed by a wearable device having at least one display, a microphone, and at least one sensor, cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through pass-through.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on the first direction of the gaze of the user and the second direction of the gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to display, through the at least one display, a text generated from the voice signal, overlapping it on the screen, based on identifying that the user and the other user corresponding to the visual object are looking at each other.
- Figure 1 illustrates an example of an environment including an electronic device and an external electronic device.
- Figure 2a illustrates an example of an environment where the volume of noise including electronic devices is greater than the user's voice volume.
- Figure 2b illustrates an example of an environment in which a user including an electronic device must whisper utterances.
- Figure 3a is a simplified block diagram of an exemplary electronic device.
- FIG. 3b is a simplified block diagram of another exemplary electronic device and external electronic device.
- Figure 4a shows an example of learned actions of the first multimodal model.
- Figure 4b shows an example of learned actions of the fourth multimodal model.
- Figure 5 illustrates examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model.
- Figures 6a and 6b illustrate examples of events for driving at least one sensor.
- Figure 7 illustrates examples of reference gestures and reference voice commands recognized by the first multimodal model.
- Figures 8a and 8b illustrate examples of operations in which an electronic device transmits a signal to cause a wake-up of a second multimodal model.
- Figures 9a and 9b illustrate examples of operations in which an electronic device transmits data to an external electronic device and receives response information from the external electronic device to utilize a second multimodal model.
- FIGS 10a and 10b illustrate examples of electronic devices that perform functions according to response information.
- Figure 11 illustrates an example in which an electronic device performs a call connection with an external electronic device.
- FIG. 12 is a block diagram of an electronic device within a network environment according to various embodiments.
- Figure 13a illustrates an example of a perspective view of a wearable device.
- FIG. 13b illustrates an example of one or more hardware devices arranged within a wearable device.
- FIG. 13c illustrates an example of a wearable device according to one embodiment.
- Figures 14a and 14b illustrate an example of the appearance of a wearable device.
- Figure 15 illustrates an example of the appearance of a wearable device.
- Figure 16 illustrates an example of a block diagram of a wearable device.
- Figure 17 illustrates an example block diagram of a wearable device for displaying images in virtual space.
- Figure 18 illustrates an example of a wearable device that recognizes audio signals emitted from external objects.
- Figure 19 illustrates an example of a wearable device that adaptively provides a response to an audio signal based on user input.
- FIGS. 20A and 20B illustrate examples of wearable devices that adaptively provide responses to audio signals based on a user's gaze.
- Figure 21 illustrates an example of a wearable device that provides a response in a second language to an audio signal spoken in a first language.
- terms referring to data e.g., data, information, sensing data, signal
- terms referring to values e.g., threshold value, reference value
- terms for operational states e.g., operation, process
- terms referring to objects e.g., visual objects, emoji graphical objects
- terms referring to network entities e.g., terms referring to components of devices, etc.
- terms such as '... part', '... device', '... object', '... body', etc. used below may mean at least one shape structure or a unit that processes a function.
- expressions such as “more than” or “less than” may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as “more than” or “less than.”
- a condition described as “more than” may be replaced with “more than”
- a condition described as “less than” may be replaced with “less than”
- a condition described as “more than and less than” may be replaced with “more than and less than.”
- “A” to “B” mean at least one of elements from A (including A) to B (including B).
- C and/or “D” mean at least one of "C” or “D,” that is, including ⁇ "C", “D", “C” and “D” ⁇ .
- Figure 1 illustrates an example of an environment including an electronic device and an external electronic device.
- the electronic device (120) may include a microphone.
- the microphone may operate in a state for low power consumption.
- the microphone may detect an audio signal (or sound signal) while in the state for low power consumption.
- the electronic device (120) may acquire an audio signal (or sound signal) through the microphone while the microphone is in the state for low power consumption.
- the electronic device (120) may acquire an audio signal from an external environment through the microphone.
- the audio signal may include a voice spoken by a user (110).
- the audio signal may include noise.
- the electronic device (120) can recognize the audio signal.
- the electronic device (120) can recognize a voice signal within the audio signal.
- the electronic device (120) can provide a voice recognition function.
- the electronic device (120) can obtain data using the recognized audio signal.
- the data can correspond to the audio signal.
- the data can include an audio signal that has undergone post-processing on the audio signal.
- the post-processing can be performed to input the audio signal to a model (e.g., an artificial intelligence model).
- the data can include text representing the audio signal.
- the data can include text representing a voice signal (or voice signal) included in the audio signal.
- the electronic device (120) may use a model for a voice recognition function.
- the electronic device (120) may train (or learn) the model.
- the voice signal may be recognized more accurately than when the electronic device (120) recognizes the voice signal without using the trained model.
- the electronic device (120) may be connected to another model in an external electronic device (130) to use the voice recognition function.
- the electronic device (120) may include a communication circuit.
- the electronic device (120) may transmit the audio signal to the external electronic device (130) through the communication circuit.
- the electronic device (120) may transmit data acquired using the audio signal to the external electronic device (130) through the communication circuit.
- the electronic device (120) may include a smartphone.
- the electronic device (120) may include a wearable device.
- the electronic device (120) may include a smartwatch.
- the electronic device (120) may include a portable electronic device such as a smartphone, a tablet, a laptop computer, or a smartwatch.
- the electronic device (120) may be described as a multi-function device or a user device.
- the external electronic device (130) can recognize a voice signal.
- the external electronic device (130) can use a model for recognizing the voice signal.
- the external electronic device (130) can train the model.
- the external electronic device (130) can recognize the voice signal more accurately by using the trained model.
- the model of the external electronic device (130) can be more complex than the model of the electronic device (120).
- the external electronic device (130) can output (or obtain) response information for the recognized voice signal.
- the external electronic device (130) can recognize the voice signal using the model, generate a prompt based on the voice signal, and provide the prompt to a large language model within the external electronic device (130).
- the external electronic device (130) can output or obtain the response information for the voice signal using the large language model.
- the external electronic device (130) may include a communication circuit.
- the external electronic device (130) may receive a signal (or data) from the electronic device (120) through the communication circuit.
- the external electronic device (130) may receive a signal from the electronic device (120) that causes a wake-up of a model within the external electronic device (130).
- the signal may change the model in a state for low power consumption to an active state.
- the signal may be referred to as a signal that operates the model in the state for low power consumption.
- the external electronic device (130) can receive or obtain a voice signal from the electronic device (120).
- the external electronic device (130) can receive or obtain data obtained using the voice signal from the electronic device (120).
- the external electronic device (130) can recognize the obtained voice signal using the model and generate a prompt.
- the external electronic device (130) can obtain response information by providing the prompt to the large language model.
- a user (110) may utter a reference voice command (or wake word) to activate a voice recognition function of an electronic device (120).
- the electronic device (120) may acquire an audio signal including the reference voice command via the microphone.
- the electronic device (120) may cause the operation of the voice recognition function based on noise within the acquired audio signal and the reference voice command.
- Figure 2a illustrates an example of an environment where the volume of noise including electronic devices is greater than the user's voice volume.
- the environment (200) can be described as an environment with a relatively high noise volume.
- the environment (200) may include an indoor environment.
- the environment (200) may include an outdoor environment.
- the environment (200) may include a cafe, a restaurant, a subway, a construction site, a plaza, or a concert hall, but the embodiments are not limited thereto.
- the environment (200) may represent an environment in which the signal-to-noise ratio (SNR) of an audio signal acquired through the microphone of the electronic device (120) is below a reference value.
- SNR signal-to-noise ratio
- the sound quality of the audio signal in the environment (200) may be relatively low.
- the environment (200) may represent an environment in which the volume of the user's (110) voice is low in relation to the noise volume.
- Figure 2b illustrates an example of an environment in which a user including an electronic device must whisper utterances.
- the environment (210) may represent an environment in which the user (110) cannot speak at a relatively high volume.
- the environment (210) may represent a library.
- the environment (210) is not limited to a library.
- the environment (210) may include, but is not limited to, a conference room, a movie theater, or an art gallery.
- a user (110) may whisper to an electronic device (120).
- the voice volume of the user (110) may be lower than a reference volume that the electronic device (120) can recognize.
- the recognition quality of the electronic device (120) may deteriorate.
- the electronic device (120) may not recognize the voice of the user (110).
- a method for removing noise from an audio signal received through a microphone of an electronic device (120) and/or a method for enhancing a voice signal within the audio signal may be required.
- a method for assisting voice recognition using at least one sensor of the electronic device (120) may be required.
- a method for recognizing a reference gesture may be required to operate the voice recognition function of the electronic device (120).
- a method for recognizing the shape of the lips of the user (110) may be required.
- a method for identifying whether sensing data acquired through at least one sensor of the electronic device (120) corresponds to the reference sensing data may be required.
- the electronic device exemplified below may include components (or hardware components) for providing these methods. These components are described and exemplified in more detail with reference to FIG. 3A.
- FIG. 3A is a simplified block diagram of an exemplary electronic device.
- the electronic device (301) may be an example of the electronic device (101) of FIG. 1.
- the electronic device (301) may include at least one processor (300), a communication circuit (310), a memory (320), a microphone (330), at least one sensor (340), a display (311), and/or a speaker (312).
- the at least one processor (300), the communication circuit (310), the memory (320), the microphone (330), at least one sensor (340), the display (311), and/or the speaker (312) may be electronically and/or operably coupled with each other by a communication bus.
- operably coupled hardware components may mean that a direct connection or an indirect connection is established between the hardware components, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components.
- FIG. 3A are illustrated based on different blocks, the present disclosure is not limited thereto.
- some of the hardware components illustrated in FIG. 3A e.g., at least one processor (300), a communication circuit (310), and at least a portion of the memory (320)
- SoC system on chip
- SIP system in package
- the type and/or number of hardware components included in the electronic device (301) are not limited to those illustrated in FIG. 3A.
- the electronic device (301) may include only some of the hardware components illustrated in FIG. 3A.
- At least one processor (300) may include a hardware component for processing data based on executing instructions.
- the hardware component for processing data may include, for example, a central processing unit (CPU) (e.g., including processing circuitry).
- the hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry).
- the hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry).
- the hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry).
- At least one processor (300) may include one or more cores.
- at least one processor (300) may have a multi-core processor architecture, such as a dual core, quad core, or hexa core.
- the content of the processor (1220) of FIG. 12 may be substantially identically applied to at least one processor (300).
- the communication circuit (310) may include hardware components for supporting transmission and/or reception of signals between the electronic device (301) and the external electronic device (302).
- the communication circuit (310) may include, for example, at least one of a modem (modulator and demodulator), an antenna, and an optical/electronic (O/E) converter.
- the communication circuit (310) may support transmission and/or reception of electrical signals based on various types of protocols, such as Ethernet, a local area network (LAN), a wide area network (WAN), wireless fidelity (WiFi), Bluetooth, Bluetooth low energy (BLE), zigbee, long term evolution (LTE), and 5G new radio (NR).
- Specific details regarding the communication circuit (310) of FIG. 3A may be substantially identically applied to the communication module (1290) and/or the antenna module (1297) of FIG. 12.
- the memory (320) may include a hardware component for storing data and/or instructions input to and/or output from at least one processor (300).
- the instructions may represent operations and/or actions to be performed on data by at least one processor (300) of the electronic device (301).
- the memory (320) may include, for example, volatile memory such as random-access memory (RAM) and/or non-volatile memory such as read-only memory (ROM).
- RAM random-access memory
- ROM read-only memory
- the volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM).
- DRAM dynamic RAM
- SRAM static RAM
- PSRAM pseudo SRAM
- Non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disk, and an embedded multimedia card (EMMC).
- PROM programmable ROM
- EPROM erasable PROM
- EEPROM electrically erasable PROM
- flash memory a hard disk, a compact disk, and an embedded multimedia card (EMMC).
- EMMC embedded multimedia card
- the microphone (330) may be configured to acquire audio signals generated around the electronic device (301).
- the microphone (330) may be used to acquire voice signals within the audio signals.
- the specific details of the microphone (330) of FIG. 3A may be substantially identical to those of the input module (1250) of FIG. 12.
- At least one sensor (340) may include a hardware component of an electronic device (301) used to acquire an external signal.
- at least one sensor (340) may detect an external signal and generate an electrical signal or data value corresponding to a detected state.
- at least one sensor (340) may include an image sensor.
- at least one sensor (340) may include a heart rate sensor.
- at least one sensor (340) may include an acceleration sensor.
- at least one sensor (340) may include a gyro sensor.
- the present invention is not limited thereto.
- the specific details of at least one sensor (340) of FIG. 3A may be substantially identically applied to the details of the sensor module (1276) of FIG. 12.
- the electronic device (301) may obtain sensing data (e.g., an image (e.g., obtained via an image sensor), heart rate data (e.g., obtained via a heart rate sensor), and/or movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340).
- sensing data e.g., an image (e.g., obtained via an image sensor), heart rate data (e.g., obtained via a heart rate sensor), and/or movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)
- sensing data e.g., an image (e.g., obtained via an image sensor)
- heart rate data e.g., obtained via a heart rate sensor
- movement data e.g., obtained via an acceleration sensor and/or a gyro sensor
- the display (311) may include hardware components of the electronic device (301) used to display a screen.
- the display (311) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light.
- each of the light-emitting elements may include an organic light emitting diode (OLED) or a micro LED.
- the display (311) may include a liquid crystal display (LCD).
- LCD liquid crystal display
- the speaker (312) can be used to output audio signals to the outside of the electronic device (301).
- the speaker (312) can be used for general purposes, such as multimedia playback or recording playback.
- the specific details of the speaker (312) of FIG. 3A can be substantially identically applied to the details of the audio output module (1255) of FIG. 12.
- At least one processor (300) may execute operations within the electronic device (301) to wake up a multimodal model within the external electronic device (302) using sensing data acquired through at least one sensor (340) and/or an audio signal acquired through a microphone (330). For example, at least one processor (300) may execute operations within the electronic device (301) to identify a type of environment in which the electronic device (301) is located. For example, at least one processor (300) may execute operations within the electronic device (301) to perform noise canceling on another audio signal acquired through the microphone (330) according to the type of the environment. The operations will be exemplified within the descriptions of FIGS. 5 to 10B.
- FIG. 3b is a simplified block diagram of another exemplary electronic device and external electronic device.
- the electronic device (301) may include a first multimodal model (350), a third model (370), and/or a fourth multimodal model (380).
- the first multimodal model (350) can be used to identify a voice signal included in an audio signal acquired through the microphone (330). For example, the first multimodal model (350) can be used to identify whether a reference voice command is included in the audio signal acquired through the microphone (330).
- the first multimodal model (350) can identify whether sensing data acquired through at least one sensor (340) (e.g., an image (e.g., acquired through an image sensor), heart rate data (e.g., acquired through a heart rate sensor), and/or movement data (e.g., acquired through an acceleration sensor and/or a gyro sensor)) represents (or includes) reference sensing data.
- the reference sensing data can include a reference gesture of a user.
- the first multimodal model (350) can identify whether an image acquired through an image sensor represents a reference gesture.
- the first multimodal model (350) can identify whether information acquired through a heart rate sensor represents reference heart rate information.
- the first multimodal model (350) can identify whether information acquired through an acceleration sensor and/or a gyro sensor represents reference movement.
- the first multimodal model (350) can be trained (or learned) using one or more modalities.
- a modality can be referred to as a type of data.
- the type of data can include, but is not limited to, text, images, and/or audio signals.
- the first multimodal model (350) may provide an automatic speech recognition (ASR) function.
- ASR automatic speech recognition
- the first multimodal model (350) may be trained to recognize or identify an audio signal acquired through a microphone (330).
- the first multimodal model (350) may be trained to perform natural language processing on an audio signal acquired through the microphone (330).
- the first multimodal model (350) may be trained to recognize or identify sensing data acquired through at least one sensor (340).
- the first multimodal model (350) may be used to acquire data representing a voice signal (or speech signal) included in an audio signal based on an audio signal and/or sensing data.
- the data representing the voice signal may include text representing the voice signal.
- FIG. 3B illustrates that the external electronic device (302) provides the output signal of the first multimodal model (350) transmitted from the electronic device (301) to the second multimodal model (360) within the external electronic device (302)
- the external electronic device (302) may provide the output signal of the first multimodal model (350) transmitted from the electronic device (301) to the large language model (390) within the external electronic device (302).
- the output signal of the first multimodal model (350) may include response information for an audio signal and/or sensing data.
- the electronic device (301) may provide a response to the audio signal and/or sensing data by executing a function according to the response information using the response information for the audio signal and/or sensing data.
- the response information for the audio signal and/or sensing data may be used to allow the electronic device (301) to provide a response based on the audio signal and/or sensing data.
- the response information for the audio signal and/or sensing data may be used to display text representing a voice signal included in the audio signal through a display (e.g., the display (311)).
- the response information for the audio signal and/or sensing data may be used to output an audio signal representing a voice signal included in the audio signal through a speaker (e.g., the speaker (312)).
- the response information for the audio signal and/or sensing data may be used to execute a function corresponding to (or linked to) a reference voice command included in the audio signal.
- response information to audio signals and/or sensing data may be used to execute a function corresponding to (or linked to) a reference gesture represented by the sensing data.
- At least one processor (300) may provide an audio signal acquired through a microphone (330) and sensing data acquired through at least one sensor (340) (e.g., an image (e.g., acquired through an image sensor), heart rate data (e.g., acquired through a heart rate sensor), and/or movement data (e.g., acquired through an acceleration sensor and/or a gyro sensor)) to the first multimodal model (350).
- at least one processor (300) may obtain response information for the audio signal and the sensing data by using an ASR function of the first multimodal model (350).
- At least one processor (300) may provide an audio signal acquired through a microphone (330) and an image acquired through at least one sensor (340) to a first multimodal model (350).
- the acquired image may include a visual object corresponding to the lips of a user of the electronic device (301).
- at least one processor (300) may use an ASR function of the first multimodal model (350) to acquire response information for the audio signal and the visual object in the image corresponding to the lips of the user.
- At least one processor (300) may transmit the acquired response information to an external electronic device (302) via a communication circuit (310).
- the response information may include data for input to a large language model (390) within the external electronic device (302).
- the data for input to the large language model (390) may be displayed or used as a prompt.
- At least one processor (300) may execute a function for acquired response information within the electronic device (301).
- the function for acquired response information may include an operation for displaying text corresponding to the acquired response information on the display (311).
- the function for acquired response information may include an operation for displaying an emoji corresponding to the acquired response information on the display (311).
- the present invention is not limited thereto.
- the third model (370) can be used to identify (or classify) the type of environment in which the electronic device (301) is located, from the audio signal acquired through the microphone (330). For example, the third model (370) can identify the type of the environment as the first type if the signal-to-noise ratio (SNR) of the audio signal is equal to or greater than a reference value. For example, the third model (370) can identify the type of the environment as the second type if the SNR of the audio signal is less than the reference value.
- SNR signal-to-noise ratio
- the third model (370) can identify the type of the environment.
- the third model (370) can provide data about the type of the identified environment to the first multimodal model (350).
- the third model (370) can provide data about the second type to the fourth multimodal model (380).
- the third model (370) can be trained using a conformer algorithm, but is not limited thereto.
- the type of environment in which the electronic device (301) is located may include, but is not limited to, an indoor environment, an outdoor environment, the first type of environment, and/or the second type of environment.
- the fourth multimodal model (380) can be used to perform noise cancellation on an audio signal acquired through the microphone (330).
- the fourth multimodal model (380) can generate an audio signal from which noise is removed from the input audio signal.
- the fourth multimodal model (380) can perform noise cancellation on the acquired audio signal using the audio signal acquired through the microphone (330), sensing data acquired through at least one sensor (340), and/or data on the type of the environment provided from the third model.
- the fourth multimodal model (380) may be trained to perform noise cancellation on the acquired audio signal.
- the fourth multimodal model (380) may represent a trained artificial intelligence model. The training of the fourth multimodal model (380) will be described and illustrated in more detail with reference to FIG. 4B.
- the external electronic device (302) may include a second multimodal model (360) and/or a large language model (390).
- the second multimodal model (360) can provide the same functionality as the first multimodal model (350).
- the second multimodal model (360) can provide an ASR function.
- the second multimodal model (360) can correspond to the first multimodal model (350).
- the second multimodal model (360) can be trained (or learned) using one or more modalities.
- the second multimodal model (360) can be learned using the same algorithm as the first multimodal model (350).
- the second multimodal model (360) can be more complex than the first multimodal model (350).
- the second multimodal model (360) can have more layers than the first multimodal model (350).
- the second multimodal model (360) may have a larger number of parameters than the first multimodal model (350).
- a large language model (390) may be referred to as a language model composed of an artificial neural network that has been pre-trained with a large amount of text data.
- the large language model (390) may include more than 10 times as many parameters (e.g., more than 100 billion parameters) as a conventional general language model.
- the large language model (390) may use a transformer artificial neural network structure based on an attention mechanism.
- the attention mechanism is a technology that helps an artificial intelligence model to focus on important parts within input data.
- the attention mechanism may be used to predict output data by predicting the degree to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of some layers of a neural network) contributes to the intermediate or final output of the neural network.
- the recurrent neural network (RNN) structure which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration within the overall (or partial) context of the input data.
- a large language model may include a transformer with an encoder-decoder structure.
- the encoder may process input data to output compressed information (e.g., an attention mechanism), and the decoder may process the compressed information to output output data in token units.
- compressed information e.g., an attention mechanism
- Each of the encoder and decoder may include an independent attention network, and may include a cross-attention network connecting the encoder and decoder.
- a large language model (390) can be trained in two stages: pre-training and fine-tuning.
- Pre-training is the process of allowing a large language model (390) to process a large amount of text data and acquire general linguistic knowledge. For example, it can include self-supervised learning to predict the next word using a previous word sequence in a text sequence.
- Fine-tuning is the process of training a large language model (390) to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. Based on a pre-trained model, additional supervised learning (or adaptive learning) can be performed using a dataset suitable for the domain purpose.
- a large language model (390) can perform a task with a text input containing natural language called a prompt.
- a large language model (390) can include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer).
- LLM Large Language Model
- the term "LLM (Large Language Model)” can refer to the neural network model itself, but can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation).
- an LLM-based chatbot such as chatGPT can also be referred to as an LLM.
- “LLM” can also include an inference engine that utilizes the LLM neural network model. For example, "entering an input prompt into an LLM” can be referred to as "entering an input prompt into an LLM-based inference engine.”
- the large language model (390) is an artificial intelligence model trained with text data and can be used to provide response information for a voice signal within an audio signal.
- the audio signal may include an audio signal received from an electronic device (301).
- the voice signal within the audio signal may include a voice signal acquired by a microphone (330) of the electronic device (301).
- the large language model (390) may be referred to as a large language model and a large language model (LLM).
- LLM large language model
- the large language model (390) may perform natural language processing.
- the large language model (390) may perform natural language understanding.
- the external electronic device (302) can generate a prompt based on an audio signal and/or sensed data using the second multimodal model (360).
- the external electronic device (302) can input or provide the generated prompt to a large language model (390).
- the large language model (390) can generate response information based on the input of the generated prompt.
- the external electronic device (302) can obtain response information for the audio signal using the large language model (390).
- Figure 4a shows an example of learned actions of the first multimodal model.
- the operations may be performed sequentially, but are not necessarily sequential.
- the order of the operations may be changed, and at least two operations may be performed in parallel.
- operations 411 to 415 may be understood to be performed by a processor (e.g., at least one processor (300) of FIG. 3A) of an electronic device (e.g., electronic device (301) of FIG. 3A).
- a processor e.g., at least one processor (300) of FIG. 3A
- an electronic device e.g., electronic device (301) of FIG. 3A.
- an audio signal and sensing data may be input to a first multimodal model (e.g., the first multimodal model (350)).
- the audio signal may represent a learning audio signal.
- the sensing data may represent learning sensing data.
- the electronic device (301) may include a feature extractor (or encoder).
- the first multimodal model (350) may include the feature extractor.
- the feature extractor may extract features of data of the audio signal from the audio signal and the sensing data.
- the feature extractor may extract features of the sensing data from the sensing data.
- the first multimodal model (350) may extract an embedding vector from the audio signal and the sensing data using the feature extractor.
- the training of the first multimodal model (350) may be performed based on supervised learning and/or unsupervised learning.
- the first multimodal model (350) may compare the similarity between the extracted embedding vector and the ground truth.
- the first multimodal model (350) may obtain a loss function.
- the first multimodal model (350) may be trained using the loss function.
- the first multimodal model (350) may change the connection weights between nodes included in each of the layers (e.g., an input layer, one or more hidden layers, and an output layer) during training.
- the first multimodal model (350) can adjust the probability of each modality for dropout.
- the dropout may indicate partially excluding neurons in the neural network of the first multimodal model (350) from learning.
- the first multimodal model (350) can avoid overfitting through the dropout.
- the first multimodal model (350) can adjust the first probability for dropout of the audio signal.
- the first multimodal model (350) can adjust the second probability for dropout of the sensing data.
- the first multimodal model (350) can adjust the third probability for dropout of the audio signal and the sensing data.
- the first multimodal model (350) may perform dropout and fine tuning.
- the first multimodal model (350) may perform dropout and fine tuning using the first probability, the second probability, and the third probability.
- the first multimodal model (350) may change the connection weights between nodes included in each layer during tuning.
- the learned motion of the first multimodal model (350) may correspond to the learned motion of the second multimodal model (360) of FIG. 3b.
- the second multimodal model (360) of FIG. 3b may be learned with the motion illustrated in FIG. 4a.
- Figure 4b shows an example of learned actions of the fourth multimodal model.
- the operations may be performed sequentially, but are not necessarily sequential.
- the order of the operations may be changed, and at least two operations may be performed in parallel.
- operations 421 to 425 may be understood to be performed by a processor (e.g., at least one processor (300) of FIG. 3A) of an electronic device (e.g., electronic device (301) of FIG. 3A).
- a processor e.g., at least one processor (300) of FIG. 3A
- an electronic device e.g., electronic device (301) of FIG. 3A.
- the electronic device (301) may include a feature extractor (not shown).
- a fourth multimodal model e.g., the fourth multimodal model (380)
- the fourth multimodal model (380) may extract features of the learning data from the learning data using the feature extractor.
- the fourth multimodal model (380) may extract an embedding vector from the learning data using the feature extractor.
- the training data may be stored in a memory (e.g., memory (320)).
- the training data may include an audio signal.
- the feature extractor may extract features of the training data by applying a short time Fourier transform (STFT).
- STFT short time Fourier transform
- voice features may be extracted from the audio signal.
- the fourth multimodal model (380) can classify (or identify) the environment in which the training audio signal is acquired from the training audio signal.
- the fourth multimodal model (380) can classify (or identify) the environment in which the training data is acquired from the embedding vector.
- the fourth multimodal model (380) can classify (or identify) the environment using a third model (e.g., the third model (370)).
- the third model (370) can classify the environment and provide data about the environment to the fourth multimodal model (380).
- the fourth multimodal model (380) can use the data about the classified environment as a token.
- the voice features extracted from the audio signal, the data about the classified environment, and the features of the learning sensing data may be input to a fourth multimodal model (380).
- the voice features and the data about the classified environment may be input to the fourth multimodal model (380) after a data concatenation operation is performed.
- the fourth multimodal model (380) may include an encoder.
- the fourth multimodal model (380) may use the encoder to perform fusion of the speech features, the data about the environment, and the features of the learning sensing data.
- the fusion may refer to combining different types of modalities into a single data.
- the features of a single multimodal model may be acquired.
- the fourth multimodal model (380) may perform noise cancellation on the audio signal by utilizing the characteristics of the single multimodal model.
- the noise cancellation may be performed by utilizing an activation function (e.g., a sigmoid function).
- the fourth multimodal model (380) may obtain a loss function based on the audio signal on which the noise cancellation was performed and the ground truth.
- the fourth multimodal model (380) may be trained through the loss function.
- the fourth multimodal model (380) may change the connection weights between nodes included in each of the layers (e.g., an input layer, one or more hidden layers, and an output layer) while being trained.
- FIG. 5 illustrates examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model.
- the operations of the electronic device (e.g., electronic device (301)) illustrated in FIG. 5 may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)).
- the operations may be performed sequentially, but are not necessarily sequential.
- the order of the operations may be changed, and at least two operations may be performed in parallel.
- At least one processor (300) may detect an event that activates at least one sensor.
- the event may be referenced as an input for driving at least one sensor (340).
- at least one processor (300) may activate at least one sensor (340) by detecting the event.
- At least one sensor (340) may be in an inactive state before operation 511 is performed by at least one processor (300).
- the image sensor may not be driven before at least one processor (300) performs operation 511.
- activating at least one sensor (340) may include at least one processor (300) driving the image sensor.
- At least one sensor (340) may be running before operation 511 is executed by at least one processor (300).
- activating at least one sensor (340) may include utilizing sensed data acquired via the at least one sensor (340) in connection with audio signal processing.
- a heart rate sensor, a motion sensor, and/or an acceleration sensor may be running before operation 511 is executed.
- activating at least one sensor (340) may include utilizing sensed data acquired via the at least one processor (300) via the heart rate sensor, the motion sensor, and/or the acceleration sensor in connection with audio signal processing. The above events are described and illustrated in more detail with reference to FIGS. 6A and 6B .
- Figures 6a and 6b illustrate examples of events for driving at least one sensor.
- the electronic device (301) may include an input element (610).
- the input element (610) may be exposed through a portion of the housing of the electronic device (301).
- the event may include at least one processor (300) receiving an input signal through the input element (610).
- at least one processor (300) may detect the input signal received through the input element (610).
- at least one processor (300) may activate at least one sensor (340) based on the detection.
- the input element (610) may be pressable.
- the input element (610) may be a physical button.
- the input element (610) may represent a pressable input button.
- the event may include at least one processor (300) receiving a push input via the input element (610).
- the input element (610) may be rotatable.
- the input element (610) may be rotatable relative to the housing of the electronic device (301).
- the event may include at least one processor (300) receiving an input that rotates the input component via the input element (610).
- the input element (610) may include a touch sensor.
- the event may include, but is not limited to, at least one processor (300) receiving a touch input through the input element (610).
- the electronic device (301) is not limited to an electronic device (301) configured in the shape of a watch.
- the electronic device (301) may include a portable electronic device such as a smartphone, a tablet, a laptop computer, or a smartwatch.
- an event for activating at least one sensor (340) may include detecting a reference motion. For example, based on the detection of the event, at least one sensor (340) in an inactive state may be driven.
- the electronic device (301) may include a motion sensor and/or an acceleration sensor.
- the motion sensor and/or the acceleration sensor may be in an activated state.
- the motion sensor and/or the acceleration sensor may be in a state for low power consumption.
- at least one processor (300) may detect the reference motion using the motion sensor and/or the acceleration sensor.
- the reference movement may include a movement of the electronic device (301) changing from state (620) to state (630).
- state (620) may represent a state in which the direction of the front side of the electronic device (301) is not facing the user.
- state (630) may represent a state in which the direction of the front side of the electronic device (301) is facing the user.
- the front side may be referred to as the front surface of the electronic device (301).
- the front side may be referred to as an area that includes the display (311) of the electronic device (301).
- the reference movement may be related to the speed at which the electronic device (301) changes from state (620) to state (630).
- the reference movement may be related to the difference between the direction of the front side of the electronic device (301) in state (620) and the direction of the front side of the electronic device (301) in state (630).
- At least one processor (300) may acquire an audio signal via a microphone (330) based on detecting the event.
- at least one processor (300) may acquire sensing data (e.g., an image (e.g., acquired via an image sensor), heart rate data (e.g., acquired via a heart rate sensor), and/or movement data (e.g., acquired via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340) activated according to the event, based on detecting the event.
- at least one processor (300) may acquire an image via an image sensor driven according to the event, based on detecting the event.
- At least one processor (300) may identify whether the sensing data represents reference sensing data. For example, at least one processor (300) may identify whether the audio signal includes a reference voice command. For example, at least one processor (300) may identify the sensing data representing the reference sensing data of the user and the audio signal including the reference voice command using a first multimodal model (e.g., the first multimodal model (350)). For example, the reference sensing data may be preset by the user.
- a first multimodal model e.g., the first multimodal model (350)
- At least one processor (300) may identify whether the image represents a reference gesture of the user. For example, at least one processor (300) may identify the image representing the reference gesture of the user and the audio signal including the reference voice command using a first multimodal model (350).
- the reference gesture may be preset by the user. For example, the reference gesture is described and exemplified in more detail with reference to FIG. 7.
- Figure 7 illustrates examples of reference gestures and reference voice commands recognized by the first multimodal model.
- the user may indicate the reference gesture.
- the reference gesture may indicate a gesture in which the user places the user's finger to the user's mouth.
- the user may utter a reference voice command (720) while performing the reference gesture.
- the reference voice command (720) may include a single syllable of speech.
- the reference voice command (720) may include a two-syllable speech.
- the reference voice command (720) may indicate 'shh'.
- the present invention is not limited thereto.
- the reference gesture and reference voice command (720) illustrated in FIG. 7 are merely examples.
- At least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the audio signal including the sensing data representing the reference sensing data and the reference voice command using the first multimodal model (350).
- the signal may include a signal causing a wake-up of a second multimodal model (e.g., the second multimodal model (360)) within the external electronic device (302).
- At least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the image representing the reference gesture of the user and the audio signal including the reference voice command using the first multimodal model (350).
- the signal may include a signal causing a wake-up of a second multimodal model (360) within the external electronic device (302).
- the external electronic device (302) may receive the signal via the communication circuit. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to wake up. For example, in response to receiving the signal, the external electronic device (302) may cause the activation of the second multimodal model (360).
- the electronic device (301) can cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where the volume of noise is relatively large by having at least one processor (300) perform the operations illustrated in FIG. 5.
- the electronic device (301) can utilize a voice recognition function through the second multimodal model (360) in an environment where the volume of noise is relatively large.
- the electronic device (301) may cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where a user must whisper an utterance, by having at least one processor (300) perform the exemplary operations of FIG. 5 .
- the electronic device (301) may utilize a voice recognition function through the second multimodal model (360) in an environment where a user must whisper an utterance.
- the quality of the voice recognition function of the electronic device (301) may be enhanced.
- FIGS. 8A and 8B illustrate examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model.
- the operations of the electronic device (e.g., electronic device (301)) illustrated in FIGS. 8A and 8B may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)).
- the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
- At least one processor (300) may obtain a first audio signal.
- at least one processor (300) may obtain the first audio signal through a microphone (330).
- At least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, in operation 812, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the first audio signal. For example, at least one processor (300) may identify the type of environment using a third model (e.g., the third model (370)). As a non-limiting example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).
- a third model e.g., the third model (370)
- at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).
- the type of the environment may include the first type and the second type.
- the first type and the second type reference may be made to the descriptions of the third model (370) of FIG. 3B.
- at least one processor (300) may identify an environment as the first type if the SNR of the environment in which the electronic device (301) is located is equal to or greater than a reference value.
- at least one processor (300) may identify an environment as the second type if the SNR of the environment in which the electronic device (301) is located is less than a reference value.
- At least one processor (300) may execute operation 813 based on identifying that the type of the environment is the first type.
- at least one processor (300) may execute operation 815 based on identifying that the type of the environment is the second type.
- At operation 813 at least one processor (300) may identify that the reference voice command is included in the first audio signal based on identifying that the type of the environment is the first type. For example, at least one processor (300) may identify that the reference voice command is included in the first audio signal using the first multimodal model (350). For example, at operation 813, at least one processor (300) may identify whether the reference voice command is included in the first audio signal.
- At least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) to cause a wake-up of the second multimodal model (360) according to the reference voice command included in the first audio signal.
- At operation 815 at least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, at least one processor (300) may drive at least one sensor (340) based on identifying that the type is the second type. For example, at least one sensor (340) may be in an inactive state prior to operation 815 being executed by at least one processor (300).
- the image sensor may not be running before at least one processor (300) executes operation 815.
- activating at least one sensor (340) may include at least one processor (300) driving the image sensor.
- At least one sensor (340) may be running before operation 815 is executed by at least one processor (300).
- activating at least one sensor (340) may include utilizing sensed data acquired via the at least one sensor (340) in connection with audio signal processing.
- a heart rate sensor, a motion sensor, and/or an acceleration sensor may be running before operation 815 is executed.
- activating at least one sensor (340) may include utilizing sensed data acquired via the at least one processor (300) via the heart rate sensor, the motion sensor, and/or the acceleration sensor in connection with audio signal processing.
- At least one processor (300) may display a screen including content notifying the activation of at least one sensor (340) through the display (311) before activating at least one sensor (340).
- At least one processor (300) may obtain the first sensing data (e.g., a first image (e.g., obtained via an image sensor), first heartbeat data (e.g., obtained via a heartbeat sensor), and/or first movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340).
- a first image e.g., obtained via an image sensor
- first heartbeat data e.g., obtained via a heartbeat sensor
- first movement data e.g., obtained via an acceleration sensor and/or a gyro sensor
- At least one processor (300) may identify reference sensing data and reference voice command using a first multimodal model (350). For example, at least one processor (300) may identify that the reference voice command is included in the second audio signal. For example, at least one processor (300) may identify whether the reference voice command is included in the second audio signal, in operation 817. For example, at least one processor (300) may identify that the reference sensing data is indicated in the first sensing data. For example, at least one processor (300) may identify whether the reference sensing data is indicated in the first sensing data. For example, at least one processor (300) may identify the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350).
- a first multimodal model 350
- At least one processor (300) may identify whether the reference gesture is included in the first image. For example, at least one processor (300) may identify the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350). For example, operation 817 may correspond to operation 513 of FIG. 5 .
- At least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350).
- the signal may be transmitted to the external electronic device (302) via the communication circuit (310) based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350).
- the signal may represent a signal that causes a wake-up of the second multimodal model (360).
- operation 818 may correspond to operation 514 of FIG. 5.
- the external electronic device (302) may receive the signal via the communication circuit.
- the external electronic device (302) may cause the second multimodal model (360) to wake up.
- the external electronic device (302) may cause the activation of the second multimodal model (360).
- operation 819 may correspond to operation 515 of FIG. 5 .
- At least one processor (300) may activate a timer.
- at least one processor (300) may execute the timer.
- operation 821 may be executed based on at least one processor (300) identifying in operation 812 of FIG. 8A that the type of the environment is the second type.
- At least one processor (300) may activate the timer based on identifying that the type of the environment is the second type.
- the timer may be associated with at least one sensor (340).
- the timer may be associated with the image sensor.
- at least one processor (300) may activate at least one sensor (340) while the timer is activated.
- at least one processor (300) may drive the image sensor while the timer is activated.
- At operation 822 at least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, operation 822 may correspond to operation 815.
- At least one processor (300) may obtain the first sensing data through at least one sensor (340).
- at least one processor (300) may obtain a first image through an image sensor.
- at least one processor (300) may obtain a second audio signal through a microphone (330).
- operation 823 may correspond to operation 816.
- At least one processor (300) may identify whether the reference voice command is included in the second audio signal. For example, at least one processor (300) may identify whether the reference gesture is included in the first sensing data. For example, at least one processor (300) may identify the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350).
- At least one processor (300) may identify whether the reference gesture is included in the first image. For example, at least one processor (300) may identify the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350).
- At least one processor (300) may execute operations 825 and 827 based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command.
- at least one processor (300) may execute operation 828 based on identifying the first sensing data not representing the reference sensing data and/or the second audio signal not including the reference voice command.
- At least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350). For example, at least one processor (300) may transmit the signal to the external electronic device (302) via the communication circuit (310) based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using a first multimodal model (350).
- the signal may be represented as a signal that causes the wake-up of the second multimodal model (360).
- operation 825 may correspond to operation 818.
- the external electronic device (302) may receive the signal via the communication circuit. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to wake up. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to be activated. For example, operation 826 may correspond to operation 819.
- At operation 827 at least one processor (300) can maintain the activation of at least one sensor (340) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, the at least one processor (300) can control the at least one sensor (340) to maintain the activation of the at least one sensor (340). For example, the at least one processor (300) can maintain operating the at least one sensor (340) independently of the expiration of the timer.
- At least one processor (300) can maintain the activation of at least one sensor (340) independently of the expiration of the timer based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, at least one processor (300) can maintain the operation of the image sensor independently of the expiration of the timer.
- At least one processor (300) may extend the remaining time of the timer based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, as the remaining time of the timer is extended, the activation state of at least one sensor (340) may be maintained. For example, as the remaining time of the timer is extended, the operation of the image sensor may be maintained.
- At step 828, at least one processor (300) may deactivate at least one sensor (340) based on identifying the first sensing data that does not represent the reference sensing data and/or the second audio signal that does not include the reference voice command.
- the at least one processor (300) may control the at least one sensor (340) to deactivate the at least one sensor (340).
- the at least one processor (300) may suspend an activated state of the at least one sensor (340).
- the suspending may be executed in response to the expiration of the timer.
- the suspending may include suspending operation of the image sensor.
- the suspending may include suspending use of data acquired via a heart rate sensor, an acceleration sensor, and/or a gyro sensor for audio signal processing.
- At least one processor (300) may change the state of at least one sensor (340) from an activated state to a deactivated state based on identifying the first reference sensing data that does not represent the reference sensing data and/or the second audio signal that does not include the reference voice command. For example, the change may be performed in response to the expiration of the timer.
- At least one processor (300) may stop operation of the image sensor based on identifying the first image that does not express the reference gesture and/or the second audio signal that does not include the reference voice command. For example, the stopping may be performed in response to the expiration of the timer.
- the electronic device (301) can cause the wake-up of the second multimodal model (360) within the external electronic device (302) to be efficiently performed according to the environment in which the electronic device (301) is located by having at least one processor (300) perform the exemplary operations of FIGS. 8A to 8B.
- the electronic device (301) may cause a wake-up of the second multimodal model (360) within the external electronic device (302) in an environment where the noise volume is relatively large.
- the electronic device (301) may use the second multimodal model (360) for a voice recognition function in an environment where the noise volume is relatively large.
- the electronic device (301) can cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where a user must whisper speech, by having at least one processor (300) perform the exemplary operations of FIGS. 8A and 8B .
- the electronic device (301) can utilize the second multimodal model (360) for a speech recognition function in an environment where a user must whisper speech.
- the quality of the speech recognition function of the electronic device (301) can be enhanced.
- FIGS. 9A and 9B illustrate examples of operations in which an electronic device transmits data to an external electronic device and receives response information from the external electronic device to utilize a second multimodal model.
- the operations of the electronic device (e.g., electronic device (301)) illustrated in FIGS. 9A and 9B may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)).
- the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
- operation 911 may represent an operation subsequent to operation 514 of FIG. 5.
- operation 911 may represent an operation subsequent to operation 818 of FIG. 8A.
- operation 911 may represent an operation subsequent to operation 827 of FIG. 8B.
- operation 911 may be described as an operation after at least one processor (300) transmits, via the communication circuit (310), the signal causing the second multimodal model (360) to wake up.
- At least one processor (300) may obtain a third audio signal via a microphone (330).
- at least one processor (300) may obtain second sensing data (e.g., a second image (e.g., obtained via an image sensor), second heartbeat data (e.g., obtained via a heartbeat sensor), and/or second movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340).
- second sensing data e.g., a second image (e.g., obtained via an image sensor), second heartbeat data (e.g., obtained via a heartbeat sensor), and/or second movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)
- at least one processor (300) may obtain a second image via the image sensor.
- At least one processor (300) may obtain first data using the third audio signal.
- the first data may represent data regarding an enhanced signal of the third audio signal.
- the first data may represent data regarding an amplified signal of a user's voice signal within the third audio signal.
- the first data may represent data regarding a noise-removed signal within the third audio signal.
- the first data may represent data regarding feature extraction performed on the third audio signal for use in the second multimodal model (360).
- At least one processor (300) may acquire second data using the second sensing data.
- the second data may represent data regarding a visual object within the second image.
- the visual object may correspond to the lips of a user of the electronic device (301).
- the second data may represent data related to the movement and/or shape of the visual object corresponding to the lips. However, this is not limited thereto.
- the electronic device (301) may include a face detection model (not shown).
- the face detection model may identify a user's face from an image acquired through an image sensor.
- the face detection model may identify an area containing another visual object corresponding to the user's face within the second image.
- at least one processor (300) may obtain data regarding the area.
- at least one processor (300) may use the data regarding the area to obtain second data regarding a visual object corresponding to the lips.
- At least one processor (300) can transmit the first data and the second data to an external electronic device (302) via a communication circuit (310).
- the external electronic device (302) may operate (or utilize) the second multimodal model (360).
- the external electronic device (302) may receive the first data and the second data via a communication circuit.
- the external electronic device (302) may provide (or input) the first data and the second data to the second multimodal model (360).
- the second multimodal model (360) may be in a wake-up state.
- the external electronic device (302) can obtain data for input into the large language model (390) using the second multimodal model (360).
- the data for input into the large language model (390) can be represented as a prompt.
- the external electronic device (302) can generate a prompt for input into the large language model (390) by providing the first data and the second data to the second multimodal model (360).
- the external electronic device (302) may operate the large language model (390).
- the external electronic device (302) may obtain response information by inputting the prompt into the large language model.
- the response information may be based on the first data and/or the second data.
- the response information may represent a response to the third audio signal.
- the response information may include a response to a visual object within the second image. The response information will be described and exemplified in more detail with reference to FIGS. 10A and 10B .
- the external electronic device (302) can transmit the response information to the electronic device (301) via the communication circuit.
- At least one processor (300) may perform or execute a function according to the response information. For example, at least one processor (300) may apply TTS (text to speech) to the response information. For example, at least one processor (300) may obtain an audio signal by applying TTS to the response information. For example, at least one processor (300) may output the audio signal obtained by applying TTS to the response information to the outside through a speaker (312).
- TTS text to speech
- at least one processor (300) may obtain an audio signal by applying TTS to the response information.
- at least one processor (300) may output the audio signal obtained by applying TTS to the response information to the outside through a speaker (312).
- At least one processor (300) may display text within the response information through a display (311) by performing the function according to the response information.
- the text within the response information may be based on the first data and/or the second data.
- the text within the response information may be generated by a second multimodal model (360) based on the first data and/or the second data.
- At least one processor (300) may obtain the third audio signal via the microphone (330).
- at least one processor (300) may obtain the second sensing data (e.g., a second image (e.g., obtained via an image sensor), second heartbeat data (e.g., obtained via a heartbeat sensor), and/or second movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340).
- at least one processor (300) may obtain the second image via the image sensor.
- Operation 921 may correspond to operation 911.
- operation 921 may represent a subsequent operation of operation 814 of FIG. 8A.
- At operation 922, at least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, at operation 922, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the third audio signal. For example, at least one processor (300) may identify the type of environment using the third model (370). As a non-limiting example, at least one processor (300) may identify the type of environment using the second sensing data acquired through at least one sensor (340).
- At least one processor (300) may identify an environment in which an electronic device (301) is located as the first type of environment if the SNR of the environment is greater than or equal to a reference value. For example, at least one processor (300) may identify an environment in which an electronic device (301) is located as the second type of environment if the SNR of the environment is less than or equal to a reference value.
- At least one processor (300) may execute operation 923 based on identifying that the type of the environment is the first type.
- at least one processor (300) may execute operation 924 and/or operation 925 based on identifying that the type of the environment is the second type.
- At least one processor (300) may transmit data regarding the third audio signal and data regarding the second sensing data. For example, at least one processor (300) may obtain the second data using the second sensing data. For example, at least one processor (300) may transmit data regarding the third audio signal and the second data regarding the visual object within the second image.
- At least one processor (300) may obtain data for the third audio signal as the first data based on identifying that the type of the environment is the first type. For example, at least one processor (300) may transmit the first data and the second data to an external electronic device (302) via a communication circuit (310). For example, operation 923 may correspond to operation 912. For example, after at least one processor (300) executes operation 923, the external electronic device (302) may execute operations 913 to 915. For example, after the external electronic device (302) executes operations 913 to 915, the at least one processor (300) may execute operation 916.
- At step 924, at least one processor (300) may perform noise cancellation on the third audio signal. For example, at least one processor (300) may perform noise cancellation on the third audio signal based on identifying that the type of the environment is the second type. For example, at least one processor (300) may perform noise cancellation using the fourth multimodal model (380).
- At least one processor (300) can perform the noise canceling by providing (or inputting) the second data, the data for the second type, and/or the data for the third audio signal to the fourth multimodal model (380).
- at least one processor (300) can obtain a fourth audio signal from which noise is removed from the third audio signal by providing the second data, the data for the second type, and/or the data for the third audio signal to the fourth multimodal model (380).
- the fourth audio signal can be generated by performing noise canceling on the third audio signal by the fourth multimodal model (380).
- the fourth audio signal may represent an audio signal in which noise within the third audio signal is reduced.
- the fourth audio signal may represent an audio signal in which a voice signal within the third audio signal is enhanced.
- At least one processor (300) may obtain data for the fourth audio signal on which noise cancellation is performed on the third audio signal.
- at least one processor (300) may obtain data for the fourth audio signal on which noise cancellation is performed on the third audio signal as the first data.
- the data for the fourth audio signal may represent data on which feature extraction is performed on the fourth audio signal for use by the second multimodal model (360).
- the SNR of the fourth audio signal may be higher than the SNR of the third audio signal.
- the ASR function of the second multimodal model (360) may be enhanced when data for the fourth audio signal is input to the second multimodal model (360) compared to when data for the third audio signal is input to the second multimodal model (360).
- At least one processor (300) may transmit the first data and the second data to the external electronic device (302) via the communication circuit (310).
- operation 925 may correspond to operation 912.
- the external electronic device (302) may execute operations 913 to 915.
- the at least one processor (300) may execute operation 916.
- the electronic device (301) can determine whether to perform noise canceling on an audio signal acquired through a microphone (330) according to the environment.
- the electronic device (301) can efficiently utilize the second multimodal model (360) within the external electronic device (302) depending on the environment in which the electronic device (301) is located. For example, the electronic device (301) can reduce the current consumed for utilizing the second multimodal model (360) by determining whether to perform noise canceling depending on the environment.
- the electronic device (301) can enhance the quality of the voice recognition function by performing noise cancellation on the acquired audio signal using the fourth multimodal model (380).
- the electronic device (301) can enhance the quality of the voice recognition function by transmitting sensing data acquired through at least one sensor (340) to an external electronic device (302).
- FIGS 10a and 10b illustrate examples of electronic devices that perform functions according to response information.
- At least one processor (300) may obtain the second sensing data through at least one sensor (340).
- at least one processor (300) may obtain the third audio signal through a microphone (330).
- At least one processor (300) may perform a function according to response information in operation 916.
- at least one processor (300) may perform the function according to the response information, thereby displaying an emoji graphical object (1021) appearing in the response information among emoji graphical objects available in the electronic device (301) through the display (311).
- the emoji graphical object (1021) may be selected by the second multimodal model (360).
- the emoji graphical object (1021) may be selected by the second multimodal model (360) based on the first data and/or the second data.
- the emoji graphical object (1021) may be pre-stored in the memory (320) in relation to the first data and/or the second data.
- the emoji graphical object (1021) may be pre-stored in the external electronic device (302) in relation to the first data and/or the second data.
- the emoji graphical object (1021) illustrated in FIG. 10A is merely exemplary.
- the emoji graphical object (1021) may include other emoji graphical objects (1021) than the city of FIG. 10A.
- At least one processor (300) may perform the operations exemplified in FIGS. 9a and 9b and perform a function according to response information.
- at least one processor (300) may perform the function according to the response information, thereby displaying text (1031) indicated by the response information among texts pre-stored in the electronic device (301) and/or the external electronic device (302) through the display (311).
- text (1031) may be selected by the second multimodal model (360).
- text (1031) may be selected by the second multimodal model (360) based on the first data and/or the second data.
- text (1031) may be pre-stored in the memory (320) in relation to the first data and/or the second data.
- text (1031) may be pre-stored in the external electronic device (302) in relation to the first data and/or the second data.
- the text (1031) shown in FIG. 10b is merely exemplary.
- the text (1031) may include text (1031) different from that shown in FIG. 10b.
- FIG. 11 illustrates an example of an electronic device performing a call connection with an external electronic device.
- the operations of the electronic device (e.g., electronic device (301)) illustrated in FIG. 11 may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)).
- the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
- At least one processor (300) may detect an event for a call connection with an external electronic device.
- the event may include receiving a touch input through a display (e.g., display (311)).
- the event may include receiving a touch input for an executable object displayed through the display (311).
- the event may include receiving a scroll input for an executable object displayed through the display (311).
- the present invention is not limited thereto.
- the event may include an event exemplified in FIG. 6A.
- At least one processor (300) may obtain a fifth audio signal via a microphone (e.g., microphone (330)).
- a microphone e.g., microphone (330)
- At least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the fifth audio signal. For example, at least one processor (300) may identify the type of environment using a third model (e.g., the third model (370)). As a non-limiting example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).
- a third model e.g., the third model (370)
- at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).
- the type of the environment may include the first type and the second type.
- at least one processor (300) may identify the environment as the first type if the SNR of the environment in which the electronic device (301) is located is equal to or greater than a reference value.
- at least one processor (300) may identify the environment as the second type if the SNR of the environment in which the electronic device (301) is located is less than a reference value.
- At least one processor (300) may execute operation 1104 based on identifying that the type of the environment is the first type.
- at least one processor (300) may execute operation 1106 based on identifying that the type of the environment is the second type.
- operation 1103 may correspond to operation 812 of FIG. 8A.
- operation 1103 may correspond to operations 922 of FIG. 9B.
- At least one processor (300) may obtain a sixth audio signal via a microphone (330).
- the sixth audio signal may include a user's voice.
- At least one processor (300) may transmit the sixth audio signal to an external electronic device via a communication circuit (310).
- the external electronic device may include a server and a base station.
- the external electronic device may include an electronic device of another user.
- the electronic device (301) may indicate a state in which it is in a call with the external electronic device.
- At least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, at least one processor (300) may drive at least one sensor (340) based on identifying that the type is the second type. For example, at least one sensor (340) may be in a disabled state prior to operation 1106 being executed by the at least one processor (300). For example, operation 1106 may correspond to operation 815 of FIG. 8A.
- At least one processor (300) may obtain third sensing data (e.g., a third image (e.g., obtained via an image sensor), third heartbeat data (e.g., obtained via a heartbeat sensor), and/or third movement data (e.g., obtained via an acceleration sensor and/or a gyro sensor)) via at least one sensor (340).
- third sensing data e.g., a third image (e.g., obtained via an image sensor
- third heartbeat data e.g., obtained via a heartbeat sensor
- third movement data e.g., obtained via an acceleration sensor and/or a gyro sensor
- operation 1107 may correspond to operation 816 of FIG. 8A.
- At least one processor (300) may provide the seventh audio signal and the third sensing data to a first multimodal model (e.g., the first multimodal model (350)).
- a first multimodal model e.g., the first multimodal model (350)
- at least one processor (300) may obtain response information for the seventh audio signal and the third sensing data.
- the response information may represent a response to the seventh audio signal.
- the response information may include a response to a user's utterance within the seventh audio signal.
- the response information may represent a response to the third sensing data.
- At least one processor (300) may perform a function according to response information for the seventh audio signal and the third sensing data.
- the response information may cause the at least one processor (300) to display an emoji graphical object (e.g., an emoji graphical object (1021)) corresponding to the seventh audio signal through the display (311).
- the emoji graphical object (1021) may be selected by the first multimodal model (350).
- the emoji graphical object (1021) may be generated by the first multimodal model (350) based on the seventh audio signal and/or the third sensing data.
- the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310).
- the external electronic device receiving the signal may display an emoji graphical object (1021) corresponding to the seventh audio signal through a display of the external electronic device.
- the response information may cause at least one processor (300) to display an emoji graphical object (1021) corresponding to the third sensing data through the display (311).
- the third sensing data may include data related to a visual object corresponding to the user's lips in an image acquired through the image sensor.
- the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310).
- the external electronic device receiving the signal may display an emoji graphical object (1021) corresponding to the seventh audio signal through a display of the external electronic device.
- the response information may cause at least one processor (300) to display text (e.g., text (1031)) corresponding to the seventh audio signal via the display (311).
- the text (1031) may be selected by the first multimodal model (350).
- the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310).
- the external electronic device receiving the signal may display text (1031) corresponding to the seventh audio signal through a display of the external electronic device.
- text (1031) may be selected by the first multimodal model (350).
- text (1031) may be generated by the first multimodal model (350) based on the seventh audio signal and/or the third sensing data.
- the response information may cause at least one processor (300) to display text (1031) corresponding to third sensing data through the display (311).
- the third sensing data may include data related to a visual object corresponding to the user's lips in an image acquired through the image sensor.
- the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310).
- the external electronic device receiving the signal may display text (1031) corresponding to the seventh audio signal through a display of the external electronic device.
- FIG. 12 is a block diagram of an electronic device (1201) within a network environment (1200) according to various embodiments.
- the electronic device (1201) may include an electronic device (301).
- an electronic device (1201) may communicate with an electronic device (1202) via a first network (1298) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1204) or a server (1208) via a second network (1299) (e.g., a long-range wireless communication network).
- the electronic device (1201) may communicate with the electronic device (1204) via the server (1208).
- the electronic device (1201) may include a processor (1220), a memory (1230), an input module (1250), an audio output module (1255), a display module (1260), an audio module (1270), a sensor module (1276), an interface (1277), a connection terminal (1278), a haptic module (1279), a camera module (1280), a power management module (1288), a battery (1289), a communication module (1290), a subscriber identification module (1296), or an antenna module (1297).
- the electronic device (1201) may omit at least one of these components (e.g., the connection terminal (1278)), or may have one or more other components added.
- some of these components e.g., sensor module (1276), camera module (1280), or antenna module (1297) may be integrated into a single component (e.g., display module (1260)).
- the processor (1220) may control at least one other component (e.g., hardware or software component) of the electronic device (1201) connected to the processor (1220) by executing, for example, software (e.g., program (1240)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1220) may store commands or data received from other components (e.g., sensor module (1276) or communication module (1290)) in volatile memory (1232), process the commands or data stored in volatile memory (1232), and store result data in non-volatile memory (1234).
- other components e.g., sensor module (1276) or communication module (1290)
- the processor (1220) may include a main processor (1221) (e.g., a central processing unit or an application processor) or an auxiliary processor (1223) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1221).
- a main processor (1221) e.g., a central processing unit or an application processor
- an auxiliary processor (1223) e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor
- the auxiliary processor (1223) may be configured to use less power than the main processor (1221) or to be specialized for a given function.
- the auxiliary processor (1223) may be implemented separately from the main processor (1221) or as a part thereof.
- the auxiliary processor (1223) may control at least a portion of functions or states associated with at least one component (e.g., a display module (1260), a sensor module (1276), or a communication module (1290)) of the electronic device (1201), for example, on behalf of the main processor (1221) while the main processor (1221) is in an inactive (e.g., sleep) state, or together with the main processor (1221) while the main processor (1221) is in an active (e.g., application execution) state.
- the auxiliary processor (1223) e.g., an image signal processor or a communication processor
- the auxiliary processor (1223) may include a hardware structure specialized for processing artificial intelligence models.
- the artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (1201) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1208)).
- the learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above.
- the artificial intelligence model can include a plurality of artificial neural network layers.
- the artificial neural network can be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, or a combination of two or more of the above, but is not limited to the examples described above.
- the artificial intelligence model can additionally or alternatively include a software structure.
- the memory (1230) can store various data used by at least one component (e.g., the processor (1220) or the sensor module (1276)) of the electronic device (1201).
- the data can include, for example, software (e.g., the program (1240)) and input data or output data for commands related thereto.
- the memory (1230) can include a volatile memory (1232) or a non-volatile memory (1234).
- the program (1240) may be stored as software in memory (1230) and may include, for example, an operating system (1242), middleware (1244), or an application (1246).
- the input module (1250) can receive commands or data to be used in a component of the electronic device (1201) (e.g., a processor (1220)) from an external source (e.g., a user) of the electronic device (1201).
- the input module (1250) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
- the audio output module (1255) can output audio signals to the outside of the electronic device (1201).
- the audio output module (1255) can include, for example, a speaker or a receiver.
- the speaker can be used for general purposes, such as multimedia playback or recording playback.
- the receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
- the display module (1260) can visually provide information to an external party (e.g., a user) of the electronic device (1201).
- the display module (1260) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device.
- the display module (1260) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
- the audio module (1270) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1270) can acquire sound through the input module (1250), output sound through the sound output module (1255), or an external electronic device (e.g., electronic device (1202)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1201).
- an external electronic device e.g., electronic device (1202)
- electronic device (1202) e.g., speaker or headphone
- the sensor module (1276) can detect the operating status (e.g., power or temperature) of the electronic device (1201) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status.
- the sensor module (1276) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
- the interface (1277) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1201) with an external electronic device (e.g., the electronic device (1202)).
- the interface (1277) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
- HDMI high definition multimedia interface
- USB universal serial bus
- SD card interface Secure Digital Card
- the connection terminal (1278) may include a connector through which the electronic device (1201) may be physically connected to an external electronic device (e.g., the electronic device (1202)).
- the connection terminal (1278) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
- the haptic module (1279) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations.
- the haptic module (1279) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
- the camera module (1280) can capture still images and videos.
- the camera module (1280) may include one or more lenses, image sensors, image signal processors, or flashes.
- the power management module (1288) can manage the power supplied to the electronic device (1201).
- the power management module (1288) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
- PMIC power management integrated circuit
- a battery (1289) may power at least one component of the electronic device (1201).
- the battery (1289) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
- the communication module (1290) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1201) and an external electronic device (e.g., electronic device (1202), electronic device (1204), or server (1208)), and the performance of communication through the established communication channel.
- the communication module (1290) may operate independently from the processor (1220) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication.
- the communication module (1290) may include a wireless communication module (1292) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1294) (e.g., a local area network (LAN) communication module, or a power line communication module).
- a wireless communication module (1292) e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module
- GNSS global navigation satellite system
- wired communication module (1294) e.g., a local area network (LAN) communication module, or a power line communication module.
- any of these communication modules may communicate with an external electronic device (1204) via a first network (1298) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1299) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).
- a first network e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)
- a second network (1299) e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)
- a first network e.g.,
- the wireless communication module (1292) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1296) to verify or authenticate the electronic device (1201) within a communication network such as the first network (1298) or the second network (1299).
- subscriber information e.g., an international mobile subscriber identity (IMSI)
- IMSI international mobile subscriber identity
- the wireless communication module (1292) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology).
- the NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)).
- eMBB enhanced mobile broadband
- mMTC massive machine type communications
- URLLC ultra-reliable and low-latency communications
- the wireless communication module (1292) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate.
- a high-frequency band e.g., mmWave band
- the wireless communication module (1292) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna.
- the wireless communication module (1292) may support various requirements specified in the electronic device (1201), an external electronic device (e.g., the electronic device (1204)), or a network system (e.g., the second network (1299)).
- the wireless communication module (1292) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
- a peak data rate e.g., 20 Gbps or more
- a loss coverage e.g., 164 dB or less
- U-plane latency e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip
- the antenna module (1297) can transmit or receive signals or power to or from an external device (e.g., an external electronic device).
- the antenna module (1297) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB).
- the antenna module (1297) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1298) or the second network (1299), may be selected from the plurality of antennas, for example, by the communication module (1290). A signal or power may be transmitted or received between the communication module (1290) and an external electronic device via the selected at least one antenna.
- another component e.g., a radio frequency integrated circuit (RFIC)
- RFIC radio frequency integrated circuit
- the antenna module (1297) may form a mmWave antenna module.
- the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
- a first side e.g., a bottom side
- a plurality of antennas e.g., an array antenna
- At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
- peripheral devices e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
- commands or data may be transmitted or received between the electronic device (1201) and an external electronic device (1204) via a server (1208) connected to a second network (1299).
- Each of the external electronic devices (1202 or 1204) may be the same or a different type of device as the electronic device (1201).
- all or part of the operations executed in the electronic device (1201) may be executed in one or more of the external electronic devices (1202, 1204, or 1208). For example, when the electronic device (1201) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1201) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service.
- One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1201).
- the electronic device (1201) may process the result as is or additionally and provide it as at least a portion of a response to the request.
- cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example.
- the electronic device (1201) may provide an ultra-low latency service using distributed computing or mobile edge computing, for example.
- the external electronic device (1204) may include an Internet of Things (IoT) device.
- the server (1208) may be an intelligent server utilizing machine learning and/or a neural network.
- the external electronic device (1204) or the server (1208) may be included in the second network (1299).
- the electronic device (1201) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
- an electronic device may be an electronic device for displaying an image in a virtual space.
- an electronic device e.g., electronic device (301) of FIG. 1 for displaying an image in a virtual space
- a wearable device e.g., wearable device (1301) to be described below
- a wearable device may include a head-mounted display (HMD) that is wearable on a user's head.
- HMD head-mounted display
- a wearable device may be referred to as a head-mounted electronic device.
- a wearable device may be referred to as a head-mounted device (HMD), a headgear electronic device, a glasses-type electronic device, a video see-through (VST) device, an extended reality (XR) device, a virtual reality (VR) device, and/or an augmented reality (AR) device.
- HMD head-mounted device
- VST video see-through
- XR extended reality
- VR virtual reality
- AR augmented reality
- An example of a hardware configuration included in a wearable device is exemplarily described with reference to FIG. 16.
- An example of a structure of a wearable device that can be worn on a user's head is described with reference to FIGS. 13A, 13B, 13C, 14A, 14B, and/or 15.
- a wearable device may be referred to as an electronic device.
- a wearable device may be combined with an accessory (e.g., a strap) to attach to a user's head to form an HMD.
- an accessory e.g.
- a wearable device may perform functions related to augmented reality (AR) and/or mixed reality (MR).
- AR augmented reality
- MR mixed reality
- the wearable device may include at least one lens positioned adjacent to the user's eyes.
- the wearable device may combine ambient light passing through the lens with light emitted from a display of the wearable device.
- a display area of the display may be formed within the lens through which the ambient light passes. Because the wearable device combines the ambient light and the light emitted from the display, the user may see an image that is a mixture of a real object perceived by the ambient light and a virtual object formed by the light emitted from the display.
- XR extended reality
- a wearable device may perform functions related to video see-through (VST) and/or virtual reality (VR).
- VST video see-through
- VR virtual reality
- the wearable device may include a housing that covers the user's eyes.
- the wearable device may include a display disposed on a first side of the housing facing the eyes.
- the wearable device may include a camera disposed on a second side opposite the first side. Using the camera, the wearable device may acquire images and/or videos representing ambient light.
- the wearable device may output the images and/or videos within the display disposed on the first side, thereby allowing the user to perceive the ambient light through the display.
- a displaying area (or displaying region) (or active area (or active region)) of the display disposed on the first side may be formed by one or more pixels included in the display.
- the wearable device can synthesize a virtual object into an image and/or video output through the display, thereby allowing the user to recognize the virtual object together with a real object recognized by ambient light.
- a wearable device can identify or recognize a position (or location) and/or direction (or orientation) of the wearable device based on an image (and/or video) obtained or acquired using a camera.
- the wearable device can obtain information about the external space using one or more cameras and/or one or more sensors.
- the information can include a geographic location (e.g., global positioning system (GPS) coordinates) of the external space identified from one or more sensors.
- GPS global positioning system
- the information can include images and/or videos of the external space identified from one or more cameras.
- the wearable device can perform object recognition on the images and/or videos to identify external objects included in the external space from the images and/or videos.
- FIG. 13A illustrates an example of a perspective view of a wearable device.
- FIG. 13B illustrates an example of one or more hardware components arranged within the wearable device.
- FIG. 13C illustrates an example of a wearable device according to an embodiment.
- the wearable device (1301) may have the form of glasses that can be worn on a body part (e.g., head) of a user.
- the wearable device (1301) of FIGS. 13A to 13C may be an example of the electronic device (301) of FIG. 3A.
- the wearable device (1301) may include a head-mounted display (HMD).
- HMD head-mounted display
- the housing of the wearable device (1301) may include a flexible material, such as rubber and/or silicone, that is configured to fit closely to a portion of the user's head (e.g., a portion of the face surrounding both eyes).
- the housing of the wearable device (1301) may include one or more straps capable of being twined around the user's head, and/or one or more temples attachable to the ears of the head.
- a wearable device may include at least one display (1350) and a frame (1300) supporting at least one display (1350).
- the at least one display (1350) may be an example of the display (311) of FIG. 3A.
- a wearable device (1301) can be worn on a part of a user's body.
- the wearable device (1301) can provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to a user wearing the wearable device (1301).
- AR augmented reality
- VR virtual reality
- MR mixed reality
- the wearable device (1301) can display a virtual reality image provided from at least one optical device (1382, 1384) of FIG. 13B on at least one display (1350) in response to a user's designated gesture acquired through the motion recognition cameras (1360-2, 1360-3) of FIG. 13B.
- At least one display (1350) may provide visual information to a user.
- at least one display (1350) may include a transparent or translucent lens.
- At least one display (1350) may include a first display (1350-1) and/or a second display (1350-2) spaced apart from the first display (1350-1).
- the first display (1350-1) and the second display (1350-2) may be positioned at positions corresponding to the user's left and right eyes, respectively.
- At least one display (1350) can provide a user with visual information transmitted from external light and other visual information distinct from the visual information through a lens included in the at least one display (1350).
- the lens can be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens.
- the at least one display (1350) can include a first surface (1331) and a second surface (1332) opposite to the first surface (1331).
- a display area can be formed on the second surface (1332) of the at least one display (1350).
- external light can be transmitted to the user by being incident on the first surface (1331) and transmitted through the second surface (1332).
- at least one display (1350) can display an augmented reality image combined with a virtual reality image provided from at least one optical device (1382, 1384) on a real screen transmitted through external light, in a display area formed on the second surface (1332).
- At least one display (1350) may include at least one waveguide (1333, 1334) that diffracts light emitted from at least one optical device (1382, 1384) and transmits the diffracted light to a user.
- the at least one waveguide (1333, 1334) may be formed based on at least one of glass, plastic, or polymer.
- a nano-pattern may be formed on at least a portion of the exterior or interior of the at least one waveguide (1333, 1334).
- the nano-pattern may be formed based on a grating structure having a polygonal and/or curved shape.
- At least one waveguide (1333, 1334) may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)), or at least one reflective element (e.g., a reflective mirror).
- at least one waveguide (1333, 1334) may be arranged within the wearable device (1301) to guide a screen displayed by at least one display (1350) to the user's eyes.
- the screen may be transmitted to the user's eyes based on total internal reflection (TIR) occurring within the at least one waveguide (1333, 1334).
- TIR total internal reflection
- the wearable device (1301) can analyze an object included in a real image collected through a shooting camera (1360-4), combine a virtual object corresponding to an object to be provided with augmented reality among the analyzed objects, and display the virtual object on at least one display (1350).
- the virtual object can include at least one of text and an image regarding various information related to the object included in the real image.
- the wearable device (1301) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the wearable device (1301) can perform spatial recognition (e.g., simultaneous localization and mapping (SLAM)) using the multi-camera and/or time-of-flight (ToF).
- SLAM simultaneous localization and mapping
- ToF time-of-flight
- a user wearing the wearable device (1301) can view an image displayed on at least one display (1350).
- the frame (1300) may be configured as a physical structure that allows the wearable device (1301) to be worn on the user's body.
- the frame (1300) may be configured so that, when the user wears the wearable device (1301), the first display (1350-1) and the second display (1350-2) can be positioned corresponding to the user's left and right eyes.
- the frame (1300) may support at least one display (1350).
- the frame (1300) may support the first display (1350-1) and the second display (1350-2) to be positioned corresponding to the user's left and right eyes.
- the frame (1300) may include a region (1320) that at least partially contacts a portion of the user's body when the user wears the wearable device (1301).
- the region (1320) of the frame (1300) that contacts a portion of the user's body may include a region that contacts a portion of the user's nose, a portion of the user's ear, and a portion of the side of the user's face that the wearable device (1301) comes into contact with.
- the frame (1300) may include a nose pad (1310) that contacts a portion of the user's body. When the wearable device (1301) is worn by the user, the nose pad (1310) may contact a portion of the user's nose.
- the frame (1300) may include a first temple (1304) and a second temple (1305) that contact a different part of the user's body than the part of the user's body.
- the frame (1300) may include a first rim (1302-1) that surrounds at least a portion of the first display (1350-1), a second rim (1302-2) that surrounds at least a portion of the second display (1350-2), a bridge (1303) that is disposed between the first rim (1302-1) and the second rim (1302-2), a first pad (1311) that is disposed along a portion of the edge of the first rim (1302-1) from one end of the bridge (1303), a second pad (1312) that is disposed along a portion of the edge of the second rim (1302-2) from the other end of the bridge (1303), a first temple (1304) that extends from the first rim (1302-1) and is fixed to a portion of the wearer's ear, and a second temple (1304) that extends from the second rim (1302-2) and is fixed to the ear opposite the ear.
- a second temple (1305) may be fixed to a portion.
- the first pad (1311) and the second pad (1312) may be in contact with a portion of the user's nose, and the first temple (1304) and the second temple (1305) may be in contact with a portion of the user's face and a portion of the user's ear.
- the temples (1304, 1305) may be rotatably connected to the rim via the hinge units (1306, 1307) of FIG. 13B.
- the first temple (1304) may be rotatably connected to the first rim (1302-1) via the first hinge unit (1306) disposed between the first rim (1302-1) and the first temple (1304).
- the second temple (1305) can be rotatably connected to the second rim (1302-2) via a second hinge unit (1307) disposed between the second rim (1302-2) and the second temple (1305).
- the wearable device (1301) can identify an external object (e.g., a user's fingertip) touching the frame (1300) and/or a gesture performed by the external object by using a touch sensor, a grip sensor, and/or a proximity sensor formed on at least a portion of a surface of the frame (1300).
- the wearable device (1301) may include hardwares that perform various functions (e.g., hardwares to be described later based on the block diagram of FIG. 16).
- the hardwares may include a battery module (1370), an antenna module (1375), at least one optical device (1382, 1384), speakers (e.g., speakers 1355-1, 1355-2), a microphone (e.g., microphones 1365-1, 1365-2, 1365-3), a light-emitting module (not shown), and/or a printed circuit board (PCB) (1390) (e.g., a printed circuit board).
- the various hardwares may be arranged within the frame (1300).
- the microphones (1365-1, 1365-2, 1365-3) may be examples of the microphone (330) of FIG. 3A.
- microphones e.g., microphones 1365-1, 1365-2, 1365-3) of a wearable device (1301) may be disposed on at least a portion of a frame (1300) to acquire sound signals.
- a first microphone (1365-1) disposed on a bridge (1303), a second microphone (1365-2) disposed on a second rim (1302-2), and a third microphone (1365-3) disposed on the first rim (1302-1) are illustrated in FIG. 13B , but the number and arrangement of the microphones (1365) are not limited to the embodiment of FIG. 13B .
- the wearable device (1301) can identify the direction of a sound signal by using a plurality of microphones arranged on different parts of the frame (1300).
- At least one optical device may project a virtual object onto at least one display (1350) to provide various image information to a user.
- at least one optical device (1382, 1384) may be a projector.
- At least one optical device (1382, 1384) may be disposed adjacent to at least one display (1350) or may be included within at least one display (1350) as a part of at least one display (1350).
- the wearable device (1301) may include a first optical device (1382) corresponding to a first display (1350-1) and a second optical device (1384) corresponding to a second display (1350-2).
- At least one optical device may include a first optical device (1382) disposed at an edge of a first display (1350-1) and a second optical device (1384) disposed at an edge of a second display (1350-2).
- the first optical device (1382) may transmit light to a first waveguide (1333) disposed on the first display (1350-1), and the second optical device (1384) may transmit light to a second waveguide (1334) disposed on the second display (1350-2).
- the camera (1360) may include a recording camera (1360-4), an eye tracking camera (ET CAM) (1360-1), and/or a motion recognition camera (1360-2, 1360-3).
- the recording camera (1360-4), the eye tracking camera (1360-1), and the motion recognition cameras (1360-2, 1360-3) may be positioned at different locations on the frame (1300) and may perform different functions.
- the eye tracking camera (1360-1) may output data indicating the position or gaze of the eyes of a user wearing the wearable device (1301). For example, the wearable device (1301) may detect the gaze from an image including the user's pupils obtained through the eye tracking camera (1360-1).
- the wearable device (1301) can identify an object (e.g., a real object and/or a virtual object) focused on by the user using the user's gaze acquired through the gaze tracking camera (1360-1).
- the wearable device (1301) that has identified the focused object can execute a function (e.g., gaze interaction) for interaction between the user and the focused object.
- the wearable device (1301) can express a part corresponding to the eye of an avatar representing the user in a virtual space using the user's gaze acquired through the gaze tracking camera (1360-1).
- the wearable device (1301) can render an image (or screen) displayed on at least one display (1350) based on the position of the user's eyes.
- the visual quality of a first area related to the gaze within the image and the visual quality (e.g., resolution, brightness, saturation, grayscale, PPI (pixels per inch)) of a second area distinguished from the first area may be different from each other.
- the term "resolution” is used to refer to the density of pixels of an image and/or display (1350). The density and/or resolution of pixels may be measured or parameterized based on units of PPI and/or dpi (dots per inch).
- the wearable device (1301) may acquire an image having a visual quality of a first area matching the user's gaze and a visual quality of a second area using foveated rendering.
- a gaze tracking camera (1360-1) positioned toward the user's right eye is shown in FIG. 13B, the embodiment is not limited thereto, and the gaze tracking camera (1360-1) may be positioned solely toward the user's left eye, or may be positioned toward both eyes.
- the capturing camera (1360-4) can capture an actual image or background to be aligned with a virtual image to implement augmented reality or mixed reality content.
- the capturing camera (1360-4) can be used to obtain a high-resolution image based on HR (high resolution) or PV (photo video).
- the capturing camera (1360-4) can capture an image of a specific object existing at a location viewed by the user and provide the image to at least one display (1350).
- the at least one display (1350) can display a single image in which information about an actual image or background including an image of the specific object obtained using the capturing camera (1360-4) and a virtual image provided through at least one optical device (1382, 1384) are superimposed.
- the wearable device (1301) can compensate for depth information (e.g., the distance between the wearable device (1301) and an external object acquired through a depth sensor) using an image acquired through the capture camera (1360-4).
- the wearable device (1301) can perform object recognition using an image acquired through the capture camera (1360-4).
- the wearable device (1301) can perform a function of focusing on an object (or subject) in an image (e.g., auto focus (AF)) and/or an optical image stabilization (OIS) function (e.g., anti-shake function) using the capture camera (1360-4).
- AF auto focus
- OIS optical image stabilization
- the wearable device (1301) can perform a pass-through function to display an image acquired through the capture camera (1360-4) by overlapping at least a portion of a screen representing a virtual space on at least one display (1350).
- the shooting camera (1360-4) may be positioned on a bridge (1303) positioned between the first rim (1302-1) and the second rim (1302-2).
- the gaze tracking camera (1360-1) can implement more realistic augmented reality by tracking the gaze of a user wearing a wearable device (1301) and matching the user's gaze with visual information provided to at least one display (1350). For example, when the wearable device (1301) looks straight ahead, the wearable device (1301) can naturally display environmental information related to the user's front at a location where the user is located on at least one display (1350).
- the gaze tracking camera (1360-1) can be configured to capture an image of the user's pupil to determine the user's gaze.
- the gaze tracking camera (1360-1) can receive gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light.
- the gaze tracking camera (1360-1) can be positioned at positions corresponding to the user's left and right eyes.
- the gaze tracking camera (1360-1) may be positioned within the first rim (1302-1) and/or the second rim (1302-2) to face the direction in which the user wearing the wearable device (1301) is positioned.
- the gesture recognition cameras (1360-2, 1360-3) can recognize the movement of the user's entire body, such as the user's torso, hands, or face, or a part of the body, and thereby provide a specific event on a screen provided on at least one display (1350).
- the gesture recognition cameras (1360-2, 1360-3) can recognize the user's gesture (gesture recognition), obtain a signal corresponding to the gesture, and provide a display corresponding to the signal on at least one display (1350).
- the processor can identify the signal corresponding to the gesture, and perform a designated function based on the identification.
- the gesture recognition cameras (1360-2, 1360-3) can be used to perform a spatial recognition function using SLAM and/or a depth map for 6 degrees of freedom pose (6 dof pose).
- the processor may perform gesture recognition and/or object tracking functions using the motion recognition cameras (1360-2, 1360-3).
- the motion recognition cameras (1360-2, 1360-3) may be positioned on the first rim (1302-1) and/or the second rim (1302-2).
- the camera (1360) included in the wearable device (1301) is not limited to the above-described gaze tracking camera (1360-1) and motion recognition cameras (1360-2, 1360-3).
- the wearable device (1301) can identify an external object included in the user's field of view (FoV) using a camera positioned toward the FoV.
- the wearable device (1301) identifying an external object can be performed based on a sensor for identifying the distance between the wearable device (1301) and the external object, such as a depth sensor and/or a time of flight (ToF) sensor.
- the camera (1360) positioned toward the FoV can support an autofocus (AF) function and/or an optical image stabilization (OIS) function.
- the wearable device (1301) may include a camera (1360) (e.g., a face tracking (FT) camera) positioned toward the face to obtain an image including the face of a user wearing the wearable device (1301).
- FT face tracking
- the wearable device (1301) may further include a light source (e.g., an LED) that emits light toward a subject (e.g., a user's eyes, face, and/or an external object within the FoV) being captured using the camera (1360).
- the light source may include an infrared wavelength LED.
- the light source may be disposed on at least one of the frame (1300) and the hinge units (1306, 1307).
- the battery module (1370) may supply power to electronic components of the wearable device (1301).
- the battery module (1370) may be disposed within the first temple (1304) and/or the second temple (1305).
- the battery module (1370) may be a plurality of battery modules (1370).
- the plurality of battery modules (1370) may be disposed within each of the first temple (1304) and the second temple (1305).
- the battery module (1370) may be disposed at an end of the first temple (1304) and/or the second temple (1305).
- the antenna module (1375) can transmit signals or power to the outside of the wearable device (1301), or receive signals or power from the outside.
- the antenna module (1375) can be positioned within the first temple (1304) and/or the second temple (1305).
- the antenna module (1375) can be positioned close to one surface of the first temple (1304) and/or the second temple (1305).
- the speaker (1355) can output an acoustic signal to the outside of the wearable device (1301).
- the acoustic output module may be referred to as a speaker.
- the speaker (1355) may be positioned within the first temple (1304) and/or the second temple (1305) so as to be positioned adjacent to the ear of a user wearing the wearable device (1301).
- the speaker (1355) may include a second speaker (1355-2) positioned within the first temple (1304) and thus positioned adjacent to the user's left ear, and a first speaker (1355-1) positioned within the second temple (1305) and thus positioned adjacent to the user's right ear.
- the speaker (1355) may be an example of the speaker (312) of FIG. 3A.
- the light-emitting module may include at least one light-emitting element.
- the light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state in order to visually provide information regarding a specific state of the wearable device (1301) to the user. For example, when the wearable device (1301) requires charging, it may emit red light at a regular cycle.
- the light-emitting module may be disposed on the first rim (1302-1) and/or the second rim (1302-2).
- a wearable device may include a printed circuit board (PCB) (1390).
- the PCB (1390) may be included in at least one of the first temple (1304) or the second temple (1305).
- the PCB (1390) may include an interposer disposed between at least two sub-PCBs.
- One or more hardwares included in the wearable device (1301) e.g., hardwares illustrated by different blocks in FIG. 16
- the wearable device (1301) may include a flexible PCB (FPCB) for interconnecting the hardwares.
- FPCB flexible PCB
- a wearable device (1301) may include at least one of a gyro sensor, a gravity sensor, and/or an acceleration sensor for detecting a posture of the wearable device (1301) and/or a posture of a body part (e.g., a head) of a user wearing the wearable device (1301).
- Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and/or acceleration based on mutually perpendicular designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis).
- the gyro sensor may measure an angular velocity of each of the designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyro sensor may be referred to as an inertial measurement unit (IMU).
- IMU inertial measurement unit
- the wearable device (1301) may identify a user's motion and/or gesture performed to execute or terminate a specific function of the wearable device (1301) based on the IMU.
- the wearable device (1301) of FIG. 13C may have the form of glasses. At least some of the hardware (or components) of the wearable device (1301) illustrated in FIGS. 13A and 13B may be applied to or included in the wearable device (1301) of FIG. 13C. Accordingly, overlapping content may be omitted.
- the wearable device (1301) may include a first rim (1302-1) and/or a second rim (1302-2).
- the wearable device (1301) may include a display (1350).
- the first display (1350-1) may be disposed on the first rim (1302-1).
- the first display (1350-1) may be disposed within the first rim (1302-1) so as to face a face of a user wearing the wearable device (1301).
- the first display (1350-1) may be disposed within the first rim (1302-1) so as to face an eye of a user wearing the wearable device (1301).
- the first display (1350-1) may be disposed at a location corresponding to the user's left eye.
- the first display (1350-1) may be positioned at the upper portion within the first rim (1302-1).
- the first display (1350-1) may display the first screen (1393-1).
- the wearable device (1301) is worn by the user, the user may view the first screen (1393-1).
- the user may view the first screen (1393-1) while looking at the upper portion of the field of view.
- the second display (1350-2) may be positioned on the second rim (1302-2).
- the second display (1350-2) may be positioned within the second rim (1302-2) to face the face of a user wearing the wearable device (1301).
- the second display (1350-2) may be positioned within the second rim (1302-2) to face the eye of a user wearing the wearable device (1301).
- the second display (1350-2) may be positioned at a position corresponding to the user's right eye.
- the second display (1350-2) may be positioned at an upper portion within the second rim (1302-2).
- the second display (1350-2) may display a second screen (1393-2).
- the wearable device (1301) is worn by the user, the user can view the second screen (1393-2).
- the user can view the second screen (1393-2) while looking at the upper part of the field of view.
- the wearable device (1301) may include a nose pad (1310).
- the nose pad (1310) may include a bridge (1303), a first pad (1311), and/or a second pad (1312).
- the first pad (1311) and the second pad (1312) may come into contact with a portion of the user's nose.
- the first display (1350-1) may be positioned on the first rim (1350-1) so as to face the user's face.
- the first display (1350-1) may be included in the user's field of view.
- the first screen (1393-1) may be included in the field of view of the user's left eye.
- the second display (1350-2) may be positioned on the second rim (1350-2) so as to face the user's face.
- the second display (1350-2) may be included in the field of view of the user.
- the second screen (1393-2) may be included in the field of view of the user's right eye.
- Figures 14a and 14b illustrate an example of an exterior appearance of a wearable device.
- the wearable device (1301) of Figures 14a and 14b may be an example of the electronic device (301) of Figure 3a.
- an example of an exterior appearance of a first side (1410) of a housing of a wearable device (1301) is illustrated in Figure 14a
- an example of an exterior appearance of a second side (1420) opposite to the first side (1410) may be illustrated in Figure 14b.
- a first surface (1410) of a wearable device (1301) may have a form attachable to a body part of a user (e.g., the face of the user).
- the wearable device (1301) may further include a strap for fixing to a body part of a user, and/or one or more temples (e.g., the first temple (1304) and/or the second temple (1305) of FIGS. 13A to 13C).
- a first display (1350-1) for outputting an image to a left eye among the user's two eyes, and a second display (1350-2) for outputting an image to a right eye among the two eyes, may be disposed on the first surface (1410).
- the wearable device (1301) may be formed on the first surface (1410) and may further include a rubber or silicone packing to prevent interference from light (e.g., ambient light) different from the light emitted from the first display (1350-1) and the second display (1350-2).
- light e.g., ambient light
- a wearable device may include cameras (1360-1) for photographing and/or tracking both eyes of a user adjacent to each of the first display (1350-1) and the second display (1350-2).
- the cameras (1360-1) may be referred to as the gaze tracking camera (1360-1) of FIG. 13B.
- a wearable device (1301) may include cameras (1360-5, 1360-6) for photographing and/or recognizing a face of a user.
- the cameras (1360-5, 1360-6) may be referred to as FT cameras.
- the wearable device (1301) can control an avatar representing the user in a virtual space based on the facial motion of the user identified using cameras (1360-5, 1360-6).
- the wearable device (1301) can change the texture and/or shape of a part of the avatar (e.g., a part of the avatar representing a human face) using information obtained by cameras (1360-5, 1360-6) (e.g., an FT camera) and representing the facial expression of the user wearing the wearable device (1301).
- a part of the avatar e.g., a part of the avatar representing a human face
- cameras (1360-5, 1360-6) e.g., an FT camera
- a camera e.g., cameras (1360-7, 1360-8, 1360-9, 1360-10, 1360-11, 1360-12)
- a sensor e.g., a depth sensor (1430)
- the cameras (1360-7, 1360-8, 1360-9, 1360-10) may be disposed on the second surface (1420) to recognize external objects.
- Cameras (1360-7, 1360-8, 1360-9, 1360-10) may be referenced to the motion recognition cameras (1360-2, 1360-3) of FIG. 13b.
- the wearable device (1301) can acquire images and/or videos to be transmitted to each of the user's eyes.
- the camera (1360-11) can be positioned on the second face (1420) of the wearable device (1301) to acquire an image to be displayed through the second display (1350-2) corresponding to the right eye among the two eyes.
- the camera (1360-12) can be positioned on the second face (1420) of the wearable device (1301) to acquire an image to be displayed through the first display (1350-1) corresponding to the left eye among the two eyes.
- the cameras (1360-11, 1360-12) can be referred to as the shooting camera (1360-4) of FIG. 13B.
- a wearable device (1301) may include a depth sensor (1430) disposed on a second face (1420) to identify a distance between the wearable device (1301) and an external object. Using the depth sensor (1430), the wearable device (1301) may obtain spatial information (e.g., a depth map) for at least a portion of a field of view (FoV) of a user wearing the wearable device (1301).
- a microphone may be disposed on the second face (1420) of the wearable device (1301) to obtain a sound output from an external object. The number of microphones may be one or more, depending on the embodiment.
- Fig. 15 illustrates an example of the appearance of a wearable device.
- the wearable device (1301) may be an example of the electronic device (301) of Fig. 3a.
- the wearable device (1301) illustrated in Fig. 15 may be understood as an embodiment of the wearable device (1301) illustrated in Figs. 13a to 14b.
- the wearable device (1301) may have a headset form and/or a headgear form.
- the wearable device (1301) may include a speaker (1355), a microphone (1365), a strap (1530), an ear pad (1540), and/or a camera (1560).
- the descriptions of the speaker (1355) in FIGS. 13A and 13B may be referred to for the speaker (1355).
- the speaker (1355) may be an example of the speaker (312) in FIG. 3A.
- Each of the first speaker (1355-1) and the second speaker (1355-2) may be positioned to face the user's ear while the wearable device (1301) is worn by the user.
- the microphone (1365) may be an example of the microphone (330) in FIG. 3A.
- the microphone (1365) may be positioned adjacent to the user's mouth while the wearable device (1301) is worn by the user.
- the microphone (1365) is illustrated in FIG. 13C as being positioned below the first camera (1560-1), the embodiment is not limited thereto.
- the microphone (1365) may be positioned above the first camera (1560-1) and/or next to the first camera (1560-1).
- the microphone (1365) may be positioned around the second camera (1560-2).
- the strap (1530) may be referred to as an accessory to be attached to the user's head.
- the ear pad (1540) may be referred to as a member that comes into contact with the user's ear.
- the ear pad (1540) may be used to improve the wearing comfort of the wearable device (1301).
- the ear pad (1540) may be composed of a foam material, memory foam, fabric, and/or artificial leather.
- the ear pad (1540) may be used to prevent leakage of an audio signal output by the speaker (1355).
- the first ear pad (1540-1) may be positioned between the first speaker (1355-1) and the user's ear while the wearable device (1301) is worn by the user.
- the second ear pad (1540-2) can be positioned between the second speaker (1355-2) and the user's ear while the wearable device (1301) is worn by the user.
- a camera (1560) can be used to acquire an image.
- the camera (1560) can be an example of the camera (1360) of FIGS. 13A to 14B .
- the camera (1560) can be used to acquire an image representing an external environment.
- the wearable device (1301) can identify an external object included in an image acquired through the camera (1560).
- the first camera (1560-1) can be positioned to face the direction in which the user's face faces while the wearable device (1301) is worn by the user.
- the first camera (1560-1) can be positioned on a housing of the wearable device (1301) that includes the first speaker (1255-1).
- the second camera (1560-2) may be positioned to face the direction in which the user's face is facing while the wearable device (1301) is worn by the user.
- the second camera (1560-2) may be positioned on the housing of the wearable device (1301) including the second speaker (1255-2).
- Fig. 16 illustrates an example of a block diagram of a wearable device.
- the wearable device (1301) of Fig. 16 may be an example of the electronic device (301) of Fig. 3A.
- the wearable device (1301) may be an example of the wearable device (1301) of Figs. 13A to 15.
- a wearable device (1301) may include a processor (1610), a memory (1615), a display (1350) (e.g., the first display (1350-1) and/or the second display (1350-2) of FIGS. 13A, 13B, 13C, 14A, and 14B), a speaker (1355), a microphone (1365), and/or at least one sensor (1620).
- the processor (1610), the memory (1615), the display (1350), the speaker (1355), the microphone (1365), and/or the at least one sensor (1620) may be electrically and/or operatively connected to each other by electronic components such as a communication bus (402).
- the operational connection of the electronic components may include a direct connection established between the electronic components and/or an indirect connection established between the electronic components such that a first electronic component among the electronic components is controlled by a second electronic component among the electronic components.
- the type and/or number of electronic components included in the wearable device (1301) is not limited to those illustrated in FIG. 16.
- the wearable device (1301) may include only some of the electronic components illustrated in FIG. 16.
- the wearable device (1301) may include the communication circuit (310) of FIG. 3A.
- the wearable device (1301) may include the first multimodal model (350), the third model (370), and/or the fourth multimodal model (380) of FIG. 3B.
- a processor (1610) of a wearable device (1301) may include a circuit (e.g., a processing circuit) for processing data based on one or more instructions.
- the processor (1610) may be an example of at least one processor (300) of FIG. 3A. Since the processor (1610) may be substantially the same as at least one processor (300) of FIG. 3A, a redundant description thereof will be omitted.
- the structure of the processor (1610) is not limited to one embodiment of the present disclosure, and at least one circuit may be formed as a separate processor that is physically separated from the processor.
- operations and/or functions of the present disclosure may be individually or collectively performed by one or more cores included in the processor (1610).
- the memory (1615) of the wearable device (1301) may include electronic components for storing data and/or instructions input to and/or output from the processor (1610).
- the memory (1615) may be an example of the memory (320) of FIG. 3A. Since the processor (1610) may be substantially the same as the memory (320) of FIG. 3A, a redundant description thereof will be omitted.
- the memory (1615) may be referred to as storage.
- a display (1350) of a wearable device (1301) may output visualized information to a user of the wearable device (1301).
- the display (1350) may be an example of the display (311) of FIG. 3A. Since the display (1350) may be substantially the same as the display (311) of FIG. 3A, a redundant description thereof will be omitted.
- a display (1350) arranged in front of the eyes of a user wearing a wearable device (1301) may be disposed on at least a portion of a housing of the wearable device (1301) (e.g., the first display (1350-1) and/or the second display (1350-2) of FIGS. 13A, 13B, 13C, 14A, and 14B).
- the display (1350) may be included in a display assembly.
- the display (1350) may be controlled by a processor (1610) to output visualized information to the user.
- the display (1350) may include a flexible display, a flat panel display (FPD), and/or electronic paper.
- the display (1350) may include a liquid crystal display (LCD), a plasma display panel (PDP), and/or one or more light emitting diodes (LEDs).
- the LED may include an OLED (organic LED).
- the embodiment is not limited thereto, and for example, if the wearable device (1301) includes a lens for transmitting external light (or ambient light), the display (1350) may include a projector (or projection assembly) for projecting light onto the lens.
- the display (1350) may be referred to as a display panel and/or a display module.
- the pixels included in the display (1350) may be arranged to face one of the user's two eyes when the wearable device (1301) is worn by the user.
- the display (1350) may include display areas (or active areas) corresponding to each of the user's two eyes.
- At least one sensor (1620) of the wearable device (1301) may generate electrical information that may be processed by the processor (1610) and/or the memory (1615) from non-electronic information related to the wearable device (1301).
- the at least one sensor (1620) may be an example of the at least one sensor (340) of FIG. 3A.
- the at least one sensor (1620) may be substantially the same as the at least one sensor (340) of FIG. 3A, and thus, a redundant description thereof will be omitted.
- the at least one sensor (1620) may include a global positioning system (GPS) sensor for detecting a geographic location of the wearable device (1301).
- GPS global positioning system
- At least one sensor (1620) may generate information indicating the geographic location of the wearable device (1301) based on a global navigation satellite system (GNSS) such as, for example, Galileo or Beidou (compass).
- GNSS global navigation satellite system
- the information may be stored in the memory (1615), processed by the processor (1610), and/or transmitted to another electronic device distinct from the wearable device (1301) via a communication circuit.
- the image sensor (1621) may include one or more light sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals representing the color and/or brightness of light.
- the image sensor (1621) may be referred to as a camera.
- the image sensor (1621) may include one or more cameras.
- the image sensor (1621) may include the camera (1360) of FIGS. 13A to 14B and/or the camera (1560) of FIG. 15.
- a plurality of light sensors included in the image sensor (1621) may be arranged in the form of a two-dimensional grid (2-dimensional array).
- the image sensor (1621) may acquire electrical signals of each of the plurality of light sensors substantially simultaneously, and generate two-dimensional frame data corresponding to light reaching the light sensors of the two-dimensional grid.
- an image (or photograph data) captured using the image sensor (1621) may mean one (a) two-dimensional frame data acquired from the image sensor (1621).
- video data captured using the image sensor (1621) may mean a sequence of a plurality of two-dimensional frame data acquired from the image sensor (1621) according to a frame rate.
- the image sensor (1621) may further include a flash light that is arranged to face a direction in which the image sensor (1621) receives light and outputs light toward the direction.
- a wearable device may include a plurality of image sensors, as examples of image sensors (1621), arranged in different directions.
- the plurality of image sensors may include gaze tracking cameras (e.g., gaze tracking camera 1360-1 of FIGS. 13B and 14A) configured to be arranged toward the eyes of a user wearing the wearable device (1301).
- the plurality of image sensors may include outward cameras.
- the processor (1610) may identify the direction of the user's gaze using images and/or videos acquired from the gaze tracking cameras.
- the gaze tracking cameras may include infrared (IR) sensors.
- the gaze tracking cameras may be referred to as eye sensors and/or eye trackers.
- the external camera may be positioned facing the front of a user wearing the wearable device (1301) (e.g., in a direction in which both eyes may face).
- the wearable device (1301) may include multiple external cameras. The embodiment is not limited thereto, and the external camera may be positioned facing an external space.
- the processor (1610) may identify external objects. For example, the processor (1610) may identify the position, shape, and/or gesture (e.g., hand gesture) of a hand of a user wearing the wearable device (1301) based on images and/or videos acquired from the external cameras. Using images and/or videos of the external environment acquired from the external cameras, the processor (1610) may recognize or track one or more objects within the external environment.
- the motion sensor (1622) may output electrical signals representing gravitational accelerations, accelerations, and/or angular velocities of a plurality of axes (e.g., x-axis, y-axis, and z-axis) that are perpendicular to each other and relative to a designated origin within the wearable device (1301) and/or the motion sensor (1622).
- the processor (1610) may repeatedly receive or acquire sensor data from the motion sensor (1622) that includes accelerations, angular velocities, and/or magnitudes of magnetic fields of the plurality of axes based on a designated period (e.g., 1 millisecond).
- the motion sensor (1622) may be referred to as an inertial measurement unit (IMU).
- IMU inertial measurement unit
- At least one sensor (1620) included in the wearable device (1301) is not limited to those described above and may include a grip sensor, a proximity sensor, a heart rate sensor, a fingerprint sensor, an ambient light sensor, and/or a ToF sensor.
- the processor (1610) can detect motion of the wearable device (1301) (e.g., motion of the wearable device (1301) caused by a user wearing the wearable device (1301).
- one or more instructions (or commands) representing data to be processed, calculations to be performed, and/or operations to be performed by the processor (1610) of the wearable device (1301) may be stored within the memory (1615) of the wearable device (1301).
- a set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and/or a software application (hereinafter, “application”).
- the wearable device (1301) and/or the processor (1610) may perform at least one of the operations of FIG. 3b, FIG. 4a, FIG. 4b, FIG. 5, FIG. 6a, FIG. 6b, FIG. 7, FIG. 8a, FIG. 8b, FIG. 9a, FIG.
- a software application is installed in a wearable device (1301) may mean that one or more instructions provided in the form of a software application (or package) are stored in a memory (1615), and that the one or more applications are stored in a format executable by the processor (1610) (e.g., a file having an extension specified by the operating system of the wearable device (1301)).
- the application may include a program and/or a library related to a service provided to a user.
- programs installed in the wearable device (1301) may be included in any one of different layers, including an application layer (1640), a framework layer (1650), and/or a hardware abstraction layer (HAL) (1680), based on the target.
- programs e.g., modules or drivers
- the hardware e.g., the display (1350), and/or at least one sensor (1620)
- the hardware abstraction layer (1680) e.g., the android system HAL, and/or the XR HAL
- the framework layer (1650) may be referred to as an XR framework layer in the sense that it includes one or more programs for providing an XR (extended reality) service.
- the layers illustrated in FIG. 16 may be logically (or for convenience of explanation) separated, and may not mean that the address space of the memory (1615) is separated by the layers.
- programs designed to target at least one of the hardware abstraction layer (1680) and/or the application layer (1640) may be included.
- the programs included in the framework layer (1650) may provide an API (application programming interface) that is executable (or callable) based on other programs.
- the application layer (1640) may include programs designed to target users of the wearable device (1301). Examples of programs included in the application layer (1640) include, but are not limited to, an extended reality (XR) system user interface (UI) (1641) and/or an XR application (1642). For example, programs (e.g., software applications) included in the application layer (1640) may call APIs to cause execution of functions supported by programs included in the framework layer (1650).
- XR extended reality
- UI system user interface
- XR application (1642 e.g., software applications
- the wearable device (1301) may display one or more visual objects on the display (1350) for performing interaction with the user based on the execution of the XR system UI (1641).
- a visual object may refer to an object that can be placed on a screen for transmitting and/or interacting with information, such as text, an image, an icon, a video, a button, a checkbox, a radio button, a text box, a slider, and/or a table.
- a visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and/or a view element.
- the wearable device (1301) may provide the user with functions available within a virtual space based on the execution of the XR system UI (1641).
- a lightweight renderer (1643) and/or an XR plug-in (1644) are illustrated to be included within the XR system UI (1641), but are not limited thereto.
- the processor (1610) may execute a lightweight renderer (1643) and/or an XR plug-in (1644) within the framework layer (1650).
- the wearable device (1301) may acquire resources (e.g., APIs, system processes, and/or libraries) used to define, create, and/or execute a rendering pipeline that allows partial changes based on the execution of a lightweight renderer (1643).
- the lightweight renderer (1643) may be referred to as a lightweight render pipeline in terms of defining a rendering pipeline that allows partial changes.
- the lightweight renderer (1643) may include a renderer built prior to the execution of a software application (e.g., a prebuilt renderer).
- the wearable device (1301) may acquire resources (e.g., APIs, system processes, and/or libraries) used to define, create, and/or execute an entire rendering pipeline based on the execution of an XR plug-in (1644).
- the XR plugin (1644) can be referred to as an open XR native client from the perspective of defining (or configuring) the entire rendering pipeline.
- the wearable device (1301) may display a screen representing at least a portion of a virtual space on the display (1350) based on the execution of the XR application (1642).
- the XR plug-in (1644-1) included in the XR application (1642) may include instructions that support functions similar to those of the XR plug-in (1644) of the XR system UI (1641). Descriptions of the XR plug-in (1644-1) that overlap with those of the XR plug-in (1644) may be omitted.
- the wearable device (1301) may cause the execution of the virtual space manager (1651) based on the execution of the XR application (1642).
- the wearable device (1301) can display an image on the display (1350) in a virtual space based on the execution of the application (16165).
- the application (16165) can be configured to output image information for displaying a two-dimensional image.
- the wearable device (1301) can cause the execution of the virtual space manager (1651) based on the execution of the application (16165).
- the wearable device (1301) can generate dual image information to display the two-dimensional image in a three-dimensional virtual space based on the execution of the application (16165).
- the dual image information can include first image information for the left eye and second image information for the right eye, taking into account binocular disparity.
- the wearable device (1301) can generate the dual image information based on the image information for displaying the two-dimensional image.
- the wearable device (1301) can provide a virtual space service based on the execution of the virtual space manager (1651).
- the virtual space manager (1651) can include a platform for supporting the virtual space service.
- the wearable device (1301) can identify a virtual space formed based on the user's location indicated by data acquired through at least one sensor (1620), and can display at least a portion of the virtual space on the display (1350).
- the virtual space manager (1651) can be referred to as a composition presentation manager (CPM).
- the virtual space manager (1651) may include a runtime service (1652).
- the runtime service (1652) may be referred to as an OpenXR runtime module (or an OpenXR runtime program).
- the wearable device (1301) may execute at least one of a user's pose prediction function, a frame timing function, and/or a spatial input function based on the execution of the runtime service (1652).
- the wearable device (1301) may perform rendering for a virtual space service for the user based on the execution of the runtime service (1652).
- a function related to a virtual space, executable by the application layer (1640) may be supported based on the execution of the runtime service (1652).
- the virtual space manager (1651) may include a pass-through manager (1653). Based on the execution of the pass-through manager (1653), the wearable device (1301) may display an image and/or video representing an actual space acquired through an external camera on at least a portion of the screen while displaying a screen representing a virtual space on the display (1350).
- the virtual space manager (1651) may include an input manager (1654).
- the wearable device (1301) may identify data (e.g., sensor data) acquired by executing one or more programs included in the recognition service layer (1670) based on the execution of the input manager (1654).
- the wearable device (1301) may use the acquired data to identify user input related to the wearable device (1301).
- the user input may be related to a motion (e.g., a hand gesture), gaze, and/or speech of the user identified by at least one sensor (1620) (e.g., an image sensor (1621) such as an external camera).
- the user input may be identified based on an external electronic device connected (or paired) via a communication circuit.
- the perception abstract layer (1660) can be used for data exchange between the virtual space manager (1651) and the perception service layer (1670). From the perspective of being used for data exchange between the virtual space manager (1651) and the perception service layer (1670), the perception abstract layer (1660) can be referred to as an interface. For example, the perception abstract layer (1660) can be referenced as OpenPX. The perception abstract layer (1660) can be used for a perception client and a perception service.
- the recognition service layer (1670) may include one or more programs for processing data acquired from at least one sensor (1620).
- the one or more programs may include at least one of a position tracker (1671), a spatial recognizer (1672), a gesture tracker (1673), an eye tracker (1674), a face tracker (1675), and/or a renderer (1690).
- the type and/or number of the one or more programs included in the recognition service layer (1670) are not limited to those illustrated in FIG. 16.
- the wearable device (1301) can identify the pose of the wearable device (1301) using at least one sensor (1620) based on the execution of the position tracker (1671).
- the wearable device (1301) can identify the 6 degrees of freedom pose (6 dof pose) of the wearable device (1301) using data acquired using an external camera (e.g., an image sensor (1621)) and/or an IMU (e.g., a motion sensor (1622) including a gyro sensor, an acceleration sensor, and/or a geomagnetic sensor) based on the execution of the position tracker (1671).
- the position tracker (1671) can be referred to as a head tracking (HeT) module (or head tracker, a head tracking program).
- HeT head tracking
- the wearable device (1301) can obtain information for providing a three-dimensional virtual space corresponding to the surrounding environment (e.g., external space) of the wearable device (1301) (or the user of the wearable device (1301)) based on the execution of the space recognizer (1672).
- the wearable device (1301) can reproduce the surrounding environment of the wearable device (1301) in three dimensions using data obtained using an external camera (e.g., an image sensor (1621)) based on the execution of the space recognizer (1672).
- the wearable device (1301) can identify at least one of a plane, a slope, and stairs based on the surrounding environment of the wearable device (1301) reproduced in three dimensions based on the execution of the space recognizer (1672).
- the spatial recognizer (1672) may be referred to as a scene understanding (SU) module (or scene recognition program).
- SU scene understanding
- the wearable device (1301) can identify (or recognize) a pose and/or gesture of a hand of a user of the wearable device (1301) based on the execution of the gesture tracker (1673). For example, the wearable device (1301) can identify a pose and/or gesture of a hand of a user using data acquired from an external camera (e.g., an image sensor (1621)) based on the execution of the gesture tracker (1673). For example, the wearable device (1301) can identify a pose and/or gesture of a hand of a user based on data (or images) acquired using an external camera based on the execution of the gesture tracker (1673).
- the gesture tracker (1673) may be referred to as a hand tracking (HaT) module (or hand tracking program) and/or a gesture tracking module.
- HaT hand tracking
- a gesture tracking module or hand tracking program
- the wearable device (1301) can identify (or track) eye movements of a user of the wearable device (1301) based on the execution of the gaze tracker (1674). For example, the wearable device (1301) can identify eye movements of the user using data acquired from a gaze tracking camera (e.g., an image sensor (1621)) based on the execution of the gaze tracker (1674).
- the gaze tracker (1674) may be referred to as an eye tracking (ET) module (or eye tracking program) and/or a gaze tracking module.
- the recognition service layer (1670) of the wearable device (1301) may further include a face tracker (1675) for tracking the user's face.
- the wearable device (1301) may identify (or track) the movement of the user's face and/or the user's expression based on the execution of the face tracker (1675).
- the wearable device (1301) may estimate the user's expression based on the movement of the user's face based on the execution of the face tracker (1675).
- the wearable device (1301) may identify the movement of the user's face and/or the user's expression based on data (e.g., images and/or videos) acquired using an FT camera (e.g., a camera facing at least a portion of the user's face, an image sensor (1621)) based on the execution of the face tracker (1675).
- the face tracker (1675) may be referred to as a face tracking (FT) (or face tracking program), and/or a face tracking module.
- the renderer (1690) may include instructions for rendering images in a three-dimensional virtual space.
- the processor (1610) e.g., DPU
- the renderer (1690) may obtain at least one image to be at least partially displayed in the display area of the display (1350) from a software application (e.g., a software application executed by the CPU and/or GPU).
- a software application e.g., a software application executed by the CPU and/or GPU.
- the processor (1610) executing the renderer (1690) may determine the location of the area in which the application (e.g., XR application (1642), application (16165)) is to be rendered.
- the processor (1610) executing the renderer (1690) may generate an image of the application to be displayed on the display (1350).
- the renderer (1690) may synthesize images to generate a composite image to be displayed on the display (1350).
- the processor (1610) executing the renderer (1690) can divide the display area of the display (1350) into a foveated portion (or may be referred to as a foveated area) and a peripheral portion (or may be referred to as a residual area) using the gaze position calculated using the position tracker (1671) and/or the gaze tracker (1674). For example, the processor (1610) detecting the coordinate values of the gaze position can determine the portion of the display area including the coordinate values as the foveated area.
- the DPU executing the renderer (1690) can obtain at least one image corresponding to each of the foveated area and the residual area, and having a size smaller than the size of the entire display area of the display (1350) or a resolution smaller than the resolution of the display area.
- the processor (1610) executing the renderer (1690) may obtain or generate a composite image to be displayed on the display (1350) by synthesizing an image corresponding to the foveated area and an image corresponding to the surrounding area. For example, the processor (1610) may perform upscaling to enlarge the image corresponding to the surrounding area to the size of the entire display area of the display (1350). On the enlarged image, the processor (1610) may combine the image corresponding to the foveated area to generate a composite image to be displayed on the display (1350). Along the boundary line of the image corresponding to the foveated area, the processor (1610) may apply a visual effect, such as blur, to blend the enlarged image and the image corresponding to the foveated area.
- a visual effect such as blur
- Fig. 17 shows an example of a block diagram of a wearable device for displaying an image in a virtual space.
- the wearable device of Fig. 17 (e.g., wearable device (1301)) may be an example of the electronic device (301) of Fig. 3A.
- Fig. 17 describes an example of executing multiple programs (or instructions) for displaying an image in a virtual space.
- the multiple programs (or instructions) may all be executed in one processor (e.g., AP) or may be executed by multiple processors (e.g., AP, GPU (graphics processing unit), NPU (neural processing unit)).
- the meaning of being executed by the multiple processors means that some programs (or instructions) may be executed by a first processor and other programs (or instructions) may be executed by a second processor different from the first processor.
- the wearable device (1301) may execute a virtual space manager (1750) (e.g., the virtual space manager (1651) of FIG. 16, CPM) to render an image in a virtual space.
- a virtual space manager (1750) e.g., the virtual space manager (1651) of FIG. 16, CPM
- the virtual space manager (1750) may include a platform for supporting a virtual space service.
- the virtual space manager (1750) may include a runtime service (1751) (e.g., OpenXR Runtime), a panel rendering (1752) (e.g., 2D Panel Render), and an XR compositor (1753).
- the wearable device (1301) may execute at least one of a user's pose prediction function, a frame timing function, and/or a spatial input function based on the execution of the runtime service (1751). For the runtime service (1751), at least some of the descriptions of the runtime service (1652) of FIG. 16 may be referred to.
- the wearable device (1301) may display at least one image (video) on a panel (e.g., a 2D panel) to implement a virtual space through the display (1350) based on the execution of the panel rendering (1752). For example, the wearable device (1301) may display a rendering image corresponding to RGB information (1766) for the panel from the spatialization manager (1740) described below through the display (e.g., the display (1350)).
- the wearable device (1301) may synthesize an image of an actual area captured by a camera in the virtual space (hereinafter, a pass-through image) with an image of a virtual area based on the execution of the XR synthesis unit (1753) (XR Compositor). For example, the wearable device (1301) can generate a composite image by merging the pass-through image and the virtual area image based on the execution of the XR synthesis unit (1753). The wearable device (1301) can transmit the generated composite image to a display buffer so that the composite image is displayed.
- the wearable device (1301) can identify a virtual space through a virtual space manager (1750) and display at least a portion of the virtual space on the display (1350).
- the virtual space manager (1750) can be referred to as a CPM.
- the wearable device (1301) can execute the virtual space manager (1750) to render an image corresponding to at least a portion of the virtual space.
- the wearable device (1301) may execute a spatialization manager (1740).
- the spatialization manager (1740) may perform processes for displaying an image in a three-dimensional virtual space.
- the wearable device (1301) may perform preprocessing based on the execution of the spatialization manager (1740) so that the image can be rendered in a three-dimensional virtual space through the virtual space manager (1750).
- the wearable device (1301) may perform at least some of the functions of the renderer (1690) of FIG. 16 based on the execution of the spatialization manager (1740).
- the wearable device (1301) can process image information provided by an application (e.g., an XR application (1710), an application providing a general 2D screen other than XR (1720), an application providing a system UI (1730)) based on the execution of the spatialization manager (1740).
- the spatialization manager (1740) e.g., Space Flinger
- the spatialization manager (1740) can include a system scene manager (1741) (e.g., System scene), an input manager (1742) (e.g., Input Routing), and a lightweight rendering engine (1743) (e.g., Impress Engine).
- the system scene manager (1741) can be executed to display the system UI (1730).
- System UI-related information (1764) can be transmitted to the system scene manager (1741) from a program (e.g., API) providing the system UI (1730).
- System UI related information (1764) can be obtained through a spatializer API and/or a same-process private API.
- the spatialization manager (1740) can determine the layout (e.g., location, display order) of the screen of the system UI (1730) in a three-dimensional space through pre-allocated resources.
- the system screen manager (1741) can transmit image information (1767) for rendering the screen of the system UI (1730) according to the layout to the virtual space manager (1750).
- the input manager (1742) can be configured to process user input (e.g., user input on a system screen or an app screen).
- the input manager (1742) can map user input recognized by at least one sensor (1620) of the wearable device (1301) to at least one of one or more software applications (e.g., an XR application (1710), an application (1720) that provides a general 2D screen that is not XR, and an application that provides a system UI (1730)) mapped to a virtual space by the spatialization manager (1740).
- mapping of the user input may include an operation of executing instructions (e.g., a sub-routine and/or an event handler) of the software application for processing the user input.
- the lightweight rendering engine (1743) may be a renderer for generating an image (e.g., a lightweight renderer (1643)).
- the lightweight rendering engine (1743) may be used to display the system UI (1730).
- the spatialization manager (1740) may include a lightweight rendering engine (1743) for rendering the system UI.
- a lightweight rendering engine (1743) for rendering the system UI.
- at least one external rendering engine may be used.
- an external rendering engine support module may be added within the spatialization manager (1740).
- the electronic device can execute an application.
- an XR application (1710) e.g., an XR application (1642), a 3D game, an XR map, or other immersive application
- the electronic device can execute a virtual space manager (1750).
- the wearable device (1301) can provide dual image information (1761) provided from the XR application (1710) to the virtual space manager (1750).
- the dual image information (1761) can include two pieces of image information that take binocular parallax into account.
- the dual image information (1761) can include first image information for the user's left eye and second image information for the user's right eye for rendering in a three-dimensional virtual space.
- dual image information is used to refer to image information for displaying images for both eyes in a three-dimensional space.
- the above dual image information may also include binocular image information, dual image data, dual images, binocular image data, stereoscopic image information, 3D image information, spatial image information, spatial image data, 2D-3D conversion data, dimensional conversion image data, binocular parallax image data, and/or equivalent technical terms.
- the wearable device (1301) can generate a composite image by merging image layers through a virtual space manager (1750).
- the wearable device (1301) can transmit the generated composite image to a display buffer.
- the composite image can be displayed on the display (1350) of the wearable device (1301).
- the electronic device can execute at least one of an XR application (1710) and other applications (1720) (e.g., a first application (1720-1), a second application (1720-2), ..., an Nth application (1720-N)).
- the application (1720) can be configured to output image information for displaying a two-dimensional (2D) image.
- the application (1720) can provide a two-dimensional (2D) image (e.g., a window and/or an activity).
- the application (1720) can be a video application, a schedule application, or an Internet browser application.
- the image information (1762) provided from the application (1720) is provided to the virtual space manager (1750)
- the image information (1762) only has x-coordinates and y-coordinates within a two-dimensional plane, so it may be difficult to consider the chronological relationship (i.e., distance from the user) between other applications centered on the user.
- the wearable device (1301) may execute the spatialization manager (1740) to provide dual image information to the virtual space manager (1750) even when displaying the application (1720) that provides a general 2D screen. For example, based on the execution of the spatialization manager (1740), the wearable device (1301) may receive application-related information (1763) from the first application (1720-1).
- the application-related information (1763) may include image information representing a two-dimensional image of the first application (1720-1) (e.g., information including RGB for each pixel) and/or content information (e.g., characteristics of content executed in the first application, type of content) in the first application (1720-1).
- the application-related information (1763) may be obtained through a spatializer API.
- the wearable device (1301) may identify information (hereinafter, “location information”) about the location of the area to be rendered by the first application (1720-1) and the size of the area to be rendered.
- the wearable device (1301) may generate dual image information (1765, e.g., RGBx2) that takes into account the user's binocular disparity through the image information and the location information. Based on the execution of the spatialization manager (1740), the wearable device (1301) can provide dual image information (1765) to the virtual space manager (1750). By converting a simple two-dimensional image into dual image information (1765), a problem that occurs when the image information (1762) is directly transmitted to the virtual space manager (1750) can be resolved. In addition, since at least some of the functions for displaying images in a virtual space are performed by the spatialization manager (1740) instead of the virtual space manager (1750), the burden on the virtual space manager (1750) can be reduced.
- dual image information (1765 e.g., RGBx2
- Fig. 18 illustrates an example of a wearable device that recognizes audio signals emitted from an external object.
- the operations of the wearable device (1301) illustrated in Fig. 18 may be executed, performed, or controlled by a processor (e.g., processor (1610)).
- a user (1801) may wear a wearable device (1301).
- the wearable device (1301) may acquire an audio signal through a microphone (e.g., microphone (1365)).
- the audio signal may include a voice signal (1810) and/or noise.
- the size of the voice signal (1810) may be relatively small within the audio signal.
- the size of the noise within the audio signal may be relatively large.
- the SNR of the audio signal may be relatively small.
- a method for increasing the recognition rate of the voice signal (1810) may be required.
- the wearable device (1301) can detect a change in the face of the speaker (1802) of the voice signal (1810).
- the change in the face of the speaker (1802) can be referred to as a change in the shape of the face caused by the speaker (1802) speaking.
- the speaker (1802) can be referred to as an utterer.
- the wearable device (1301) can detect a change in the lips of the speaker (1802) of the voice signal (1810).
- the change in the lips of the speaker (1802) can be referred to as a change in the shape of the lips caused by the speaker (1802) speaking.
- the wearable device (1301) can obtain sensing data through at least one sensor (e.g., at least one sensor (1620)).
- the sensing data can include an image obtained through an image sensor (e.g., the image sensor (1621)).
- the image can be an image of an external environment (or a physical environment) captured.
- the image can include a visual object (1821) corresponding to a speaker (1802).
- a part of the visual object (1821) can correspond to the face (e.g., lips) of the speaker (1802).
- the wearable device (1301) can detect a change in the face (e.g., lips) of the speaker (1802) by identifying a part of the visual object (1821) in the image obtained through the image sensor (1621).
- the wearable device (1301) can provide an audio signal (e.g., including a voice signal (1810)) acquired through a microphone (1365) and/or sensing data through at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)).
- a first multimodal model e.g., the first multimodal model (350)
- the wearable device (1301) can provide an audio signal and/or data indicating changes in the lips of a speaker (1802) to the first multimodal model (350).
- the wearable device (1301) can identify a voice signal (1810) included in an audio signal even in an environment where the SNR of the audio signal is relatively low by using the first multimodal model (350).
- the wearable device (1301) can identify the voice signal (1810) of a speaker (1802) even if the voice signal is relatively noisy and/or the speaker (1802) is located relatively far away.
- the wearable device (1301) can obtain response information for the audio signal and/or sensing data by providing the audio signal and/or sensing data to the first multimodal model (350).
- the response information can be generated by the first multimodal model (350) according to the audio signal and/or sensing data.
- the response information can be used to display text representing a voice signal (1810) included in the audio signal through a display (e.g., the display (1350)).
- the response information can be used to output an audio signal representing the voice signal (1810) through a speaker (e.g., the speaker (1355)).
- the audio signal provided to the first multimodal model (350) can be an audio signal on which noise cancellation is performed according to a fourth multimodal model (e.g., the fourth multimodal model (380)).
- the wearable device (1301) can provide a pass-through function for the external environment by using an image acquired through the image sensor (1621).
- the wearable device (1301) can provide a screen (1820) representing at least a portion of a virtual space (or a three-dimensional space) corresponding to a physical environment through a display (e.g., a display (1350)).
- the screen (1820) can include a visual object (1821).
- the visual object (1821) can correspond to a speaker (1802).
- the visual object (1821) can represent the speaker (1802).
- the wearable device (1301) may use the response information to display text (1827) in an area (1825) within the screen (1820).
- the text (1827) may correspond to the voice signal (1810).
- the text (1827) may represent the voice signal (1810).
- the location of the area (1825) may be adjacent to the visual object (1821).
- the area (1825) may be parallel to the visual object (1821) on the screen (1820).
- the area (1825) may have a predefined location on the screen (1820).
- the area (1825) may overlap at least a portion of the visual object (1821).
- the wearable device (1301) may use the response information to output an audio signal (1830) through the speaker (1355).
- the audio signal (1830) may correspond to the voice signal (1810).
- the SNR of the audio signal (1830) may be higher than the SNR of an audio signal (e.g., including the voice signal (1810)) acquired through the microphone (1365).
- the audio signal (1830) may be acquired by performing noise canceling on the audio signal acquired through the microphone (1365).
- the fourth multimodal model (380) may be used for noise canceling.
- the audio signal (1830) may be generated by the electronic device (1301).
- an audio signal (1830) may be generated using the response information.
- the wearable device (1301) may generate or obtain the audio signal (1830) by performing text-to-speech (TTS) on text (1827) representing the voice signal (1810).
- the audio signal (1830) may be generated by the first multimodal model (350).
- the voice feature of the generated audio signal (1830) may be a predetermined voice feature.
- the voice feature of the generated audio signal (1830) may be substantially the same as an audio signal of an audio signal obtained through a microphone (1365) (e.g., including the voice signal (1810)).
- the electronic device (1301) may generate the audio signal (1830) using the voice feature of the audio signal obtained through the microphone (1365).
- the electronic device (1301) may generate an audio signal (1830) using pre-stored voice characteristics of another person (e.g., a celebrity).
- the voice characteristics may include at least one of the gender, voice pattern, timbre, pitch, or pronunciation of the speaker (1802).
- the wearable device (1301) may perform inference on the region.
- the wearable device (1301) may perform inference on the region using an artificial intelligence model.
- the wearable device (1301) may obtain text (1827) and/or an audio signal (1830) by performing inference on the region.
- the wearable device (1301) may display a visual object (1829) through the screen (1820).
- the visual object (1829) may indicate that the wearable device (1301) provides a function to display text (1827) through the display (1350) based on an audio signal (e.g., a voice signal (1810)) acquired through the microphone (1365).
- the visual object (1829) may indicate that the wearable device (1301) provides a function to output an audio signal (1830) through the speaker (1355) based on an audio signal (e.g., a voice signal (1810)) acquired through the microphone (1365).
- the wearable device (1301) may change the state of the wearable device (1301) to a state that does not provide a function of displaying text (1827) through the display (1350) based on a user input.
- the wearable device (1301) may change the state of the wearable device (1301) to a state that does not provide a function of outputting an audio signal (1830) through the speaker (1355) based on a user input.
- the user input may include an input to a visual object (1829), but the embodiment is not limited thereto.
- the user input may include a predetermined gesture input.
- the wearable device (1301) may provide a function of displaying text (1827) based on an audio signal through a display (1350) and/or a function of outputting an audio signal (1830) based on the audio signal through a speaker (1355) while the wearable device (1301) is positioned in an environment where noise is greater than a reference noise.
- the wearable device (1301) may provide a function of displaying text (1827) based on an audio signal through a display (1350) and/or a function of outputting an audio signal (1830) based on the audio signal through a speaker (1355) when the quality of the audio signal acquired through the microphone (1365) is relatively low.
- the wearable device (1301) may identify the size of a voice signal (1810) included in an audio signal acquired through the microphone (1365). For example, the wearable device (1301) may provide a function to display text (1827) through a display (1350) based on an audio signal, based on identifying that the size of the voice signal (1810) is smaller than a reference size. For example, the wearable device (1301) may provide a function to output an audio signal (1830) through a speaker (1355) based on an audio signal, based on identifying that the size of the voice signal (1810) is smaller than a reference size.
- the wearable device (1301) may refrain from, stop, or bypass providing a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the size of the voice signal (1810) is greater than a reference size.
- the wearable device (1301) may refrain from, stop, or bypass providing a function to output an audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the size of the voice signal (1810) is greater than a reference size.
- the wearable device (1301) can identify the SNR of an audio signal (e.g., including a voice signal (1810)) acquired via the microphone (1365). For example, the wearable device (1301) can provide a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the SNR of the audio signal is less than a reference value. For example, the wearable device (1301) can provide a function to output the audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the SNR is less than a reference value.
- an audio signal e.g., including a voice signal (1810)
- the wearable device (1301) can provide a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the SNR of the audio signal is less than a reference value.
- the wearable device (1301) may refrain from, stop, or bypass providing a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the SNR of the audio signal of the voice signal (1810) is greater than a reference value.
- the wearable device (1301) may refrain from, stop, or bypass providing a function to output the audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the SNR of the audio signal is greater than a reference value.
- the wearable device (1301) can communicate with an external electronic device (1838) (e.g., a smartphone).
- the external electronic device (1838) may be an example of the electronic device (301) of FIG. 3A and/or FIG. 3B .
- the wearable device (1301) may be referred to as a companion device of the external electronic device (1838).
- the wearable device (1301) may establish a communication link between the wearable device (1301) and the external electronic device (1838).
- the wearable device (1301) may transmit an audio signal (e.g., including a voice signal (1810)) acquired via a microphone (1365) to the external electronic device (1838).
- an audio signal e.g., including a voice signal (1810)
- the wearable device (1301) may transmit sensing data acquired through at least one sensor (1620) to an external electronic device (1838).
- the sensing data may include an image acquired through an image sensor (1621).
- the image may include a portion of a visual object (1821) corresponding to the lips of a speaker (1802).
- the external electronic device (1838) may include the first multimodal model (350).
- the external electronic device (1838) may display a screen (1840) through the display of the external electronic device (1838) based on an audio signal transmitted from the wearable device (1301) and/or sensing data transmitted from the wearable device (1301).
- the screen (1840) may correspond to the screen (1820).
- the screen (1840) may be substantially identical to the screen (1820), and thus, a redundant description will be omitted.
- the visual object (1841) may correspond to the visual object (1821).
- the area (1845) may correspond to the area (1825).
- the descriptions for the area (1825) may be referenced for the area (1845).
- text (1847) may correspond to text (1827).
- text (1827) reference may be made to descriptions of text (1847).
- the external electronic device (1838) may output an audio signal (1850) through a speaker of the external electronic device (1838) based on the audio signal and/or the transmitted sensing data.
- the audio signal (1850) may correspond to audio signal (1830).
- audio signal (1830) reference may be made to descriptions of audio signal (1850).
- the external electronic device (1838) can transmit data representing text (1847) and/or data representing audio signals (1850) to the wearable device (1301).
- the wearable device (1301) can use the data representing text (1847) to display text (1827) through the display (1350).
- the wearable device (1301) can use the data representing audio signals (1850) to output audio signals (1830) through the speaker (1355).
- the wearable device (1301) is illustrated as providing a pass-through function, but the embodiment is not limited thereto.
- the wearable device (1301) may perform the operations illustrated in FIG. 18 while displaying or playing a video through the display (1350).
- FIG. 19 illustrates an example of a wearable device that adaptively provides a response to an audio signal based on the gaze of an external object.
- the operations of the wearable device (1301) illustrated in FIG. 19 may be executed, performed, or controlled by a processor (e.g., processor (1610)).
- processor e.g., processor (1610)
- any content that overlaps with the description of FIG. 18 may be omitted.
- a wearable device (1301) can provide a pass-through function to a user (1801).
- the wearable device (1301) can obtain an audio signal through a microphone (e.g., microphone (1365)).
- the audio signal can include a voice signal (1810), a voice signal (1905), and/or noise.
- the wearable device (1301) can obtain sensing data through at least one sensor (e.g., at least one sensor (1620)).
- the sensing data can include an image obtained through an image sensor (e.g., image sensor (1621)).
- the image can be an image of an external environment (or physical environment) captured.
- the image may include a visual object (1921) corresponding to the speaker (1802) and/or a visual object (1923) corresponding to the speaker (1903).
- the visual object (1921) may be an example of the visual object (1821) of FIG. 18.
- the image may include a portion of the visual object (1921) corresponding to the lips of the speaker (1802) and/or a portion of the visual object (1923) corresponding to the lips of the speaker (1903).
- the wearable device (1301) can provide audio signals acquired via a microphone (1365) and/or sensing data via at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (1810) and/or a voice signal (1905) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.
- a first multimodal model e.g., the first multimodal model (350)
- the wearable device (1301) can identify a voice signal (1810) and/or a voice signal (1905) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.
- the wearable device (1301) can obtain response information for the audio signal and/or the sensing data by providing the audio signal and/or the sensing data to the first multimodal model (350).
- the response information can be generated by the first multimodal model (350) according to the audio signal and/or the sensing data.
- the response information can be used to display text representing the voice signal (1810) and/or text representing the voice signal (1905) through a display (e.g., the display (1350)).
- the response information can be used to output an audio signal representing the voice signal (1810) and/or an audio signal representing the voice signal (1905) through a speaker (e.g., the speaker (1355)).
- the audio signal provided to the first multimodal model (350) may be an audio signal on which noise cancellation has been performed according to the fourth multimodal model (e.g., the fourth multimodal model (380)).
- the wearable device (1301) may display a screen (1920) via the display (1350).
- the screen (1920) may include a visual object (1921) corresponding to a speaker (1802) and/or a visual object (1923) corresponding to a speaker (1903).
- the wearable device (1301) may use the response information to display text (1927) in an area (1925) within the screen (1920).
- the text (1927) may represent a voice signal (1810).
- the wearable device (1301) may use the response information to display text (1928) in an area (1926) within the screen (1920).
- the text (1928) may represent a voice signal (1905).
- each of area (1925) and area (1926) may be substantially identical to area (1825) of FIG. 18, so overlapping content may be omitted.
- the wearable device (1301) may use the response information to output an audio signal (1930) and/or an audio signal (1931) through a speaker (e.g., speaker (1355)).
- the audio signal (1930) may represent the voice signal (1810).
- the quality of the audio signal (1930) may be higher than the quality of the voice signal (1810).
- the SNR of the audio signal (1930) may be higher than the SNR of the voice signal (1810).
- the size (e.g., volume) of the audio signal (1930) may be greater than the size of the voice signal (1810).
- the audio signal (1931) may represent the voice signal (1905).
- the quality of the audio signal (1931) may be higher than the quality of the voice signal (1905).
- the SNR of the audio signal (1931) may be higher than the SNR of the voice signal (1905).
- the magnitude (e.g., loudness) of the audio signal (1931) may be greater than the magnitude of the voice signal (1905).
- the wearable device (1301) can identify the location of the speaker (1802) and/or the location of the speaker (1903) using sensing data acquired through at least one sensor (1620).
- the sensing data may include an image acquired through an image sensor (1621).
- the image may include a visual object (1921) and/or a visual object (1923).
- the sensing data may include data acquired through a ToF sensor.
- the wearable device (1301) can identify the distance between the wearable device (1301) and the speaker (1802).
- the wearable device (1301) can identify the direction from the wearable device (1301) toward the speaker (1802).
- the wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1802) is facing. For example, the wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1802) is facing by identifying the direction in which a portion corresponding to the face (e.g., lips) of a visual object (1921) is facing.
- the wearable device (1301) can identify the distance between the wearable device (1301) and the speaker (1903). For example, the wearable device (1301) can identify the direction from the wearable device (1301) toward the speaker (1903). The wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1903) is facing. For example, the wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1903) is facing by identifying the direction in which a portion corresponding to the face (e.g., lips) of a visual object (1923) is facing.
- the wearable device (1301) can identify a voice signal (e.g., voice signal (1810), voice signal (1905)) included in the audio signal by using an audio signal acquired through a microphone (1365).
- the wearable device (1301) can identify the speech location of the voice signal included in the audio signal.
- the microphone (1365) can include a plurality of microphones.
- each of the plurality of microphones can acquire a voice signal.
- the phase of the acquired voice signal can be different depending on the time difference of the voice signal reaching each of the plurality of microphones.
- the wearable device (1301) can identify the speech location of the voice signal by using the phase of the acquired voice signal.
- the wearable device (1301) can identify the location of speech of a voice signal by utilizing the difference in sound pressure of the voice signal reaching each of a plurality of microphones.
- the wearable device (1301) can identify the speaker of each of the voice signals (e.g., voice signal (1810), voice signal (1905)).
- the wearable device (1301) can determine the speaker of the voice signal based on the similarity between the location of the identified voice signal and the location of the identified speaker (e.g., speaker (1802), speaker (1903)).
- the wearable device (1301) can identify that the voice signal (1810) is spoken by the speaker (1802) based on the similarity between the location of the voice signal (1810) and the location of the speaker (1802).
- the wearable device (1301) can identify that the voice signal (1905) is uttered by the speaker (1903) based on the similarity between the location of the utterance of the voice signal (1905) and the location of the speaker (1903).
- the wearable device (1301) can identify who the speaker of the voice signal (e.g., voice signal (1810), voice signal (1905)) is, it can provide a function according to the voice signal according to each speaker of the voice signal.
- the function according to the voice signal can include a function of displaying text representing the voice signal (e.g., text (1927), text (1928)) and/or a function of outputting an audio signal representing the voice signal (e.g., audio signal (1930), audio signal (1931)).
- the wearable device (1301) can determine a speaker.
- the wearable device (1301) can determine a speaker (1802) among speakers by receiving a user input (1922) for a visual object (1921) within a screen (1920).
- the user input (1922) for the visual object (1921) can change from example (1901) to example (1902).
- the user input (1922) for the visual object (1921) can include at least one of a gesture input for the visual object (1921), a gaze input for the visual object (1921), or a voice input for the visual object (1921).
- the wearable device (1301) can continue to provide a function according to a voice signal (1810) in response to a user input (1922) for a visual object (1921).
- the wearable device (1301) can continue to display text (1927) in an area (1925) within the screen (1920).
- the wearable device (1301) can continue to output an audio signal (1930) through a speaker (1355).
- the wearable device (1301) may stop, skip, refrain from, or bypass providing a function according to a voice signal (1905) based on a user input (1922) to a visual object (1921). For example, the wearable device (1301) may stop, skip, refrain from, or bypass displaying text (1928) based on a user input (1922) to a visual object (1921). For example, the text (1928) may not be displayed on the screen (1920). The wearable device (1301) may stop, skip, refrain from, or bypass outputting an audio signal (1931) through a speaker (1355).
- the wearable device (1301) may perform cancellation on a voice signal (1905) acquired through a microphone (1365) according to a user input (1922) for a visual object (1921).
- the wearable device (1301) may output an audio signal (1930) among the audio signal (1931) and the audio signal (1930) according to a user input (1922) for a visual object (1921).
- the wearable device (1301) may output a voice signal (1905) and an audio signal (1930) acquired through a microphone (1365) through a speaker (1355) in response to a user input (1922) for a visual object (1921).
- the wearable device (1301) can track the visual object (1921) based on a user input (1922) for the visual object (1921). For example, the wearable device (1301) can identify whether the visual object (1921) is positioned within the screen (1920) based on the user input (1922) for the visual object (1921). For example, the wearable device (1301) can continue to provide the function according to the voice signal (1810) while the visual object (1921) is positioned within the screen (1920). For example, the wearable device (1301) can stop providing the function according to the voice signal (1810) based on identifying that at least a portion of the visual object (1921) is not positioned within the screen (1920). For example, the wearable device (1301) may temporarily disable providing functionality in response to a voice signal (1810).
- the wearable device (1301) may activate a timer based on identifying that at least a portion of a visual object (1921) is not positioned within the screen (1920). For example, the wearable device (1301) may provide a function according to the voice signal (1810) based on identifying that the visual object (1921) is positioned within the screen (1920) before the timer expires. For example, the wearable device (1301) may terminate the function according to the voice signal (1810) in response to the expiration of the timer. As a non-limiting example, the wearable device (1301) may not terminate the function according to the voice signal (1810) based on identifying that the visual object (1921) is included within one or more images acquired via the image sensor (1621).
- the wearable device (1301) may display a visual effect for a visual object (1921) through a screen (1920) in response to a user input (1922) for the visual object (1921).
- the visual effect may be used to distinguish the visual object (1921) that is the target of the user input (1922) from other visual objects.
- FIGS. 20A and 20B illustrate examples of a wearable device that adaptively provides responses to audio signals based on a user's gaze.
- the operations of the wearable device (1301) illustrated in FIGS. 20A and 20B may be executed, performed, or controlled by a processor (e.g., processor (1610)). Any details that overlap with the descriptions of FIGS. 18 and 19 may be omitted.
- a wearable device (1301) can provide a pass-through function to a user (1801).
- the wearable device (1301) can obtain an audio signal via a microphone (e.g., microphone (1365)).
- the audio signal can include a voice signal (1810), a voice signal (2005), and/or noise.
- the wearable device (1301) can obtain sensing data via at least one sensor (e.g., at least one sensor (1620)).
- the sensing data can include an image obtained via an image sensor (e.g., image sensor (1621)).
- the image can be an image of an external environment (or physical environment) captured.
- the image may include a visual object (2021) corresponding to the speaker (1802) and/or a visual object (2023) corresponding to the speaker (2003).
- the image may include a portion of a visual object (2021) corresponding to the lips of the speaker (1802).
- the wearable device (1301) can provide audio signals acquired via a microphone (1365) and/or sensing data via at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (1810) and/or a voice signal (2005) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.
- a first multimodal model e.g., the first multimodal model (350)
- the wearable device (1301) can identify a voice signal (1810) and/or a voice signal (2005) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.
- the wearable device (1301) may identify the direction of the gaze of the user (1801).
- the wearable device (1301) may identify the direction of the gaze of the speaker (1802).
- the wearable device (1301) may identify the direction of the gaze of a visual object (2021) included in an image acquired through the image sensor (1621).
- the direction of the gaze of the speaker (1802) may correspond to the direction of the gaze of the visual object (2021).
- the wearable device (1301) may identify the direction of the gaze of the speaker (1802) using the direction of the gaze of the visual object (2021).
- the wearable device (1301) can identify that the user (1801) and the speaker (1802) are looking at each other based on the direction of the user's (1801) gaze and the direction of the speaker's (1802) gaze.
- the wearable device (1301) can identify that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction, thereby identifying that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction.
- the wearable device (1301) can provide a function according to the voice signal (1810) of the speaker (1802) based on identifying that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction. For example, for the function according to the voice signal (1810), reference may be made to the descriptions of FIG. 19.
- a function according to a voice signal (1810) may include a function of displaying text (2027) representing the voice signal (1810) through an area (2025) within a screen (2020) and/or a function of outputting an audio signal (2030) representing the voice signal (1810).
- the wearable device (1301) may identify the direction of the gaze of the user (1801).
- the wearable device (1301) may identify the direction of the gaze of the speaker (2003).
- the wearable device (1301) may identify the direction of the gaze of a visual object (2023) included in an image acquired through the image sensor (1621).
- the direction of the gaze of the speaker (2003) may correspond to the direction of the gaze of the visual object (2023).
- the wearable device (1301) may identify the direction of the gaze of the speaker (2003) using the direction of the gaze of the visual object (2023).
- the wearable device (1301) can identify that the user (1801) and the speaker (2003) are not looking at each other based on the direction of the user's (1801) gaze and the direction of the speaker's (2003) gaze. As the wearable device (1301) identifies that the direction of the user's (1801) gaze does not correspond to the direction of the speaker's (2003) gaze, the wearable device (1301) can identify that the user (1801) and the speaker (2003) are not looking at each other. Based on identifying that the user (1801) and the speaker (2003) are not looking at each other, the wearable device (1301) can refrain from, stop, skip, or bypass providing a function according to the speaker's (2003) voice signal (1810).
- the example (2001) may change to the example (2002).
- the user (1801) and the speaker (2003) may be looking at each other.
- the wearable device (1301) may identify the direction of the gaze of the user (1801).
- the wearable device (1301) may identify the direction of the gaze of the speaker (2003).
- the wearable device (1301) may identify the direction of the gaze of a visual object (2023) included in an image acquired through the image sensor (1621).
- the direction of the gaze of the speaker (2003) may correspond to the direction of the gaze of the visual object (2023).
- the wearable device (1301) can identify the direction of the gaze of the speaker (2003) using the direction of the gaze of the visual object (2023). Based on the direction of the gaze of the user (1801) and the direction of the gaze of the speaker (2003), the wearable device (1301) can identify that the user (1801) and the speaker (2003) are looking at each other. As the wearable device (1301) identifies that the direction of the gaze of the user (1801) corresponds to the direction of the gaze of the speaker (2003), the wearable device can identify that the user (1801) and the speaker (2003) are looking at each other.
- the wearable device (1301) can provide a function based on the voice signal (1810) of the speaker (2003) based on identifying that the user (1801) and the speaker (2003) are looking at each other.
- the wearable device (1301) can further display text (2028) in an area (2026) within the screen (2020).
- the text (2028) can represent the voice signal (2005).
- the wearable device (1301) can output an audio signal (2040) through the speaker (1355).
- the audio signal (2040) can represent the voice signal (1810) and/or the voice signal (2005).
- the audio signal (2040) may correspond to a signal that is a combination of the audio signal (2030) and the voice signal (2005).
- Fig. 20b the operations illustrated in Fig. 20b may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
- the wearable device (1301) may display a screen representing at least a portion of a three-dimensional space through a display (e.g., display (1350)).
- the three-dimensional space may be provided according to a pass-through function of the wearable device (1301).
- the three-dimensional space may correspond to a physical environment.
- the wearable device (1301) may display a screen representing at least a portion of the three-dimensional space using an image acquired through an image sensor (e.g., image sensor (1361)).
- the screen representing at least a portion of the three-dimensional space may include one or more visual objects.
- each visual object within the one or more visual objects may correspond to a speaker.
- the wearable device (1301) may identify a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in a screen based on an audio signal acquired through a microphone (e.g., microphone (1365)).
- the audio signal may be acquired while the screen is displayed through the display (1350).
- the wearable device (1301) may identify a speaker for the voice signal included in the audio signal.
- the wearable device (1301) may identify a visual object corresponding to a speaker for the voice signal.
- the wearable device (1301) may identify a speaker corresponding to a visual object among one or more speakers using the audio signal.
- the wearable device (1301) can identify a first direction of a gaze of a user wearing the wearable device (1301) and a second direction of a gaze of a visual object corresponding to a voice signal, through at least one sensor (e.g., at least one sensor (1620)).
- the wearable device (1301) can identify the gaze of the user through an image sensor (1621) (e.g., a gaze tracking camera (1360-1)).
- the gaze of a visual object corresponding to a voice signal can be related to the gaze of a speaker corresponding to the visual object.
- the gaze of the visual object can correspond to the gaze of the speaker.
- the wearable device (1301) can identify whether the user wearing the wearable device (1301) and the speaker corresponding to the visual object are looking at each other based on a first direction of the gaze of the user wearing the wearable device (1301) and a second direction of the gaze of the visual object. For example, the wearable device (1301) can identify the location of the user wearing the wearable device (1301) using at least one sensor (1620). For example, the wearable device (1301) can identify the location of the speaker corresponding to the visual object using at least one sensor (1620).
- the wearable device (1301) can identify whether a first condition that the speaker is located in the first direction from the location of the user and a second condition that the user is located in the second direction from the location of the speaker are each satisfied.
- the first condition may be referenced as the user's gaze being directed toward the speaker.
- the second condition may be referenced as the speaker's gaze being directed toward the user.
- the wearable device (1301) can identify that the user and the speaker are looking at each other based on the fulfillment of the first condition and the fulfillment of the second condition.
- the wearable device (1301) can identify that the user and the speaker are not looking at each other based on the failure (or non-fulfillment) of the first condition and/or the failure (or non-fulfillment) of the second condition.
- the wearable device (1301) can execute operation 2059 based on identifying that the user and the speaker corresponding to the visual object are looking at each other.
- the wearable device (1301) can execute operation 2061 based on identifying that the user and the speaker corresponding to the visual object are not looking at each other.
- the wearable device (1301) may display, on the display (1350), a text generated from the speaker's voice signal, superimposed on the screen, based on identifying that the user and the speaker corresponding to the visual object are looking at each other.
- the text may represent the voice signal.
- the text generated from the voice signal may include text converted from the voice signal.
- the wearable device (1301) may generate the text by applying STT (speech to text) to the voice signal.
- the text may be displayed in conjunction with a visual object.
- the text may be displayed next to the visual object.
- the wearable device (1301) may output an audio signal representing a voice signal through a speaker (e.g., speaker (1355)) based on identifying that a user and a speaker corresponding to a visual object are looking at each other.
- a speaker e.g., speaker (1355)
- the wearable device (1301) may refrain from, stop, skip, or bypass displaying text generated from the speaker's voice signal through the display (1350) based on the wearable device (1301) identifying that the user and the speaker corresponding to the visual object are not looking at each other. For example, the wearable device (1301) may continue to display the screen of operation 2051 through the display (1350).
- the wearable device (1301) may refrain from, stop, skip, or bypass outputting an audio signal representing a voice signal through the speaker (1355) based on identifying that the user and the speaker corresponding to the visual object are not looking at each other.
- Figure 21 illustrates an example of a wearable device that provides a response in a second language to an audio signal spoken in a first language.
- the operations of the wearable device (1301) illustrated in Figure 21 may be executed, performed, or controlled by a processor (e.g., processor (1610)).
- the wearable device (1301) can acquire an audio signal via a microphone (e.g., microphone (1365)).
- the audio signal can include a voice signal (2110) and/or noise.
- the voice signal (2110) can be represented by a first language.
- the voice signal (2110) can include an utterance in the first language.
- the wearable device (1301) can acquire sensing data via at least one sensor (e.g., at least one sensor (1620)).
- the sensing data can include an image acquired via an image sensor (e.g., image sensor (1621)).
- the image can be an image of an external environment (or physical environment) captured.
- the image can include a visual object (2121) corresponding to a speaker (1802).
- the image may include a portion of a visual object (2121) corresponding to the lips of the speaker (1802).
- the wearable device (1301) can provide audio signals acquired through a microphone (1365) and/or sensing data through at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (2110) even in an environment where the SNR of the audio signal is relatively low by using the first multimodal model (350).
- a first multimodal model e.g., the first multimodal model (350)
- the wearable device (1301) can identify a voice signal (2110) even in an environment where the SNR of the audio signal is relatively low by using the first multimodal model (350).
- the first multimodal model (350) may include one or more multimodal models.
- each of the one or more multimodal models may be trained using a different language.
- each of the one or more multimodal models may be trained using images representing lips speaking a different language.
- the wearable device (1301) can identify that the language of the voice signal (2110) is a first language. For example, the wearable device (1301) can determine a multimodal model corresponding to the first language from among one or more multimodal models. For example, if the first language is English, the wearable device (1301) can determine a multimodal model trained using English. For example, if the first language is Arabic, the wearable device (1301) can determine a multimodal model trained using Arabic.
- the wearable device (1301) can display the screen (2120) through a display (e.g., display (1350)).
- the wearable device (1301) can display text (2125) corresponding to the voice signal (2110) through an area (2125) within the screen (2120).
- the text (2125) can be represented by a second language.
- the context of the voice signal (2110) and the context of the text (2125) can be substantially the same.
- the text (2125) can be a text representing the voice signal (2110) translated from a first language to a second language.
- the wearable device (1301) can output an audio signal (2130) corresponding to the voice signal (2110) through a speaker (e.g., speaker (1355)).
- the language of the audio signal (2130) may be a second language.
- the audio signal (2130) may indicate that the voice signal (2110) has been translated from a first language into a second language.
- the second language may be different from the first language of the voice signal (2110).
- the second language may be a language set in the wearable device (1301).
- the second language may be set because the language most frequently spoken by the user (1801) within a given time period is the second language.
- the second language may be set by the user (1801).
- An electronic device as described above may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for driving the image sensor.
- the instructions, when individually or collectively executed by the at least one processor may cause the electronic device to acquire a first image through the image sensor driven in response to the event, and to acquire a first audio signal through the microphone, based on the detection.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to an external electronic device, wherein the signal causes a wake-up of a second multimodal model (e.g., a second multimodal model (360)) within the external electronic device based on identifying, using a first multimodal model (e.g., a first multimodal model (350)) within the electronic device, the first image and the first audio signal including the reference voice command representing the reference gesture of the user.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device, after transmitting the signal, to acquire a second audio signal via the microphone and to acquire a second image via the image sensor.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, to the external electronic device via the communication circuit, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device via the communication circuit, response information regarding the second audio signal, obtained using the second multimodal model while in a wake-up state in response to the signal.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from the second audio signal among the second image and the second audio signal, a type of environment in which the electronic device is located, using a third model within the electronic device (e.g., the third model (370)).
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain data for a third audio signal as the first data through noise canceling on the second audio signal, which is performed by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type.
- a fourth multimodal model e.g., the fourth multimodal model (380)
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device through the communication circuit, the response information representing a text selected by the second multimodal model from among texts pre-stored in the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.
- the electronic device may include a speaker.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information.
- TTS text-to-speech
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to output the audio signal through the speaker.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second data using the third data.
- a method performed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit as described above may include an operation of detecting an event for driving the image sensor.
- the method may include an operation of acquiring a first image through the image sensor driven in response to the event based on the detection, and an operation of acquiring a first audio signal through the microphone.
- the method may include an operation of transmitting, through the communication circuit, a signal to an external electronic device, based on identifying the first image expressing a reference gesture of a user and the first audio signal including a reference voice command, using a first multimodal model within the electronic device (e.g., a first multimodal model (350)), the signal causing a wake-up of a second multimodal model within the external electronic device (e.g., a second multimodal model (360)).
- the method may include an operation of acquiring a second audio signal through the microphone, and a second image through the image sensor.
- the method may include an operation of transmitting, to the external electronic device, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips, via the communication circuit.
- the method may include an operation of receiving, from the external electronic device, response information regarding the second audio signal obtained using the second multimodal model in a wake-up state according to the signal, via the communication circuit.
- the method may include an operation of performing a function according to the response information.
- the method may include an operation of identifying a type of an environment in which the electronic device is located from the second audio signal among the second audio signal and the second image using a third model within the electronic device (e.g., the third model (370)).
- the method may include an operation of obtaining data for a third audio signal as the first data through noise cancellation of the second audio signal by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type.
- the method may include an operation of obtaining the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type.
- the method may further include an operation of transmitting the first data and the second data to the external electronic device through the communication circuit.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device, via the communication circuit.
- the method may include performing the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, through the communication circuit, response information indicating a text selected by the second multimodal model from among texts pre-stored within the electronic device.
- the method may include performing the function according to the response information by displaying the text indicated by the response information through the display.
- the electronic device may include a speaker.
- the method may include an operation of obtaining an audio signal by applying text-to-speech (TTS) to the received response information.
- TTS text-to-speech
- the method may include an operation of outputting the audio signal through the speaker.
- the electronic device may include a display.
- the method may include an operation of receiving, via the communication circuit, the response information, which includes text generated based on the first data and the second data.
- the method may include an operation of performing the function according to the response information by displaying the text within the response information through the display.
- the method may include an operation of obtaining third data for an area including another visual object corresponding to the user's face within the second image using a face detection model within the electronic device.
- the method may include an operation of obtaining the second data using the third data.
- the one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit, cause the electronic device to detect an event to drive the image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image through the image sensor driven according to the event, and to acquire a first audio signal through the microphone, based on the detection.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuit, a signal to the external electronic device, wherein the signal causes a wake-up of a second multimodal model (e.g., a second multimodal model (360)) within the external electronic device based on identifying, using a first multimodal model (e.g., a first multimodal model (350)) within the electronic device, the first image representing the user's reference gesture and the first audio signal including the reference voice command.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, after transmitting the signal, acquire a second audio signal via the microphone and to acquire a second image via the image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, to the external electronic device via the communication circuit, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, from the external electronic device via the communication circuit, response information regarding the second audio signal obtained using the second multimodal model in a wake-up state in response to the signal.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function according to the response information.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify a type of an environment in which the electronic device is located from the second audio signal among the second image and the second audio signal using a third model within the electronic device (e.g., the third model (370)).
- a third model within the electronic device e.g., the third model (370)
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain data for a third audio signal as the first data through noise canceling on the second audio signal, which is performed by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type, and obtain the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, from an external electronic device, response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the one or more programs may include instructions that cause the electronic device to receive, through the communication circuit, from the external electronic device, the response information indicating a text selected by the second multimodal model from among texts pre-stored in the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text indicated by the response information through the display.
- the electronic device may include a speaker.
- the one or more programs may include instructions that cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to output the audio signal through the speaker.
- TTS text-to-speech
- the electronic device may include a display.
- the one or more programs may include instructions that cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the second data using the third data.
- the electronic device may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first audio signal via the microphone.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to identify a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device (e.g., the third model (370)).
- a model within the electronic device e.g., the third model (370)
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to the external electronic device, in accordance with a reference voice command included in the first audio signal, a wake-up of a first multimodal model (e.g., a second multimodal model (360)) within the external electronic device based on identifying the type of the environment as a first type.
- a first multimodal model e.g., a second multimodal model (360)
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor based on identifying the type of the environment as a second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image via the driven image sensor.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a second audio signal via the microphone.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuit, to the external electronic device, the signal causing a wake-up of the first multimodal model based on identifying the first image representing the user's reference gesture and the second audio signal including the reference voice command using a second multimodal model (e.g., the first multimodal model (350)) within the electronic device.
- a second multimodal model e.g., the first multimodal model (350)
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to activate a timer associated with the image sensor based on identifying that the type of the environment is the second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to acquire the first image via the driven image sensor prior to expiration of the timer based on identifying that the type of the environment is the second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, the signal to cause a wake-up of the first multimodal model based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using the second multimodal model, based on identifying the type of the environment to be the second type, and to continue driving the image sensor independently of the expiration of the timer.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to stop driving the image sensor in response to the expiration of the timer, based on identifying the type of the environment to be the second type, and based on identifying the first image not representing the reference gesture and/or the second audio signal not including the reference voice command using the second multimodal model.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the signal based on identifying that the type of the environment is the second type, and then acquire a third audio signal via the microphone and acquire a second image via the image sensor.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, to the external electronic device via the communication circuit, first data acquired using the third audio signal and second data regarding a visual object in the second image corresponding to the user's lips.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device via the communication circuit, response information for the third audio signal, the first data acquired using the first multimodal model in a wake-up state in response to the signal.
- the above instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using the model within the electronic device, the type of the environment in which the electronic device is located from the third audio signal among the third audio signal and the second image.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for the third audio signal as the first data based on identifying that the type of the environment is the first type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain data for a fourth audio signal as the first data by providing the second data, fourth data for the second type, and the third data to a third multimodal model (e.g., fourth multimodal model (380)) based on identifying that the type of the environment is the second type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device through the communication circuit, response information representing text selected by the first multimodal model from among texts pre-stored in the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.
- the electronic device may include a speaker.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information.
- TTS text-to-speech
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to output the audio signal through the speaker.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second data using the third data.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a third audio signal via the microphone after transmitting the signal based on identifying the type of the environment as the first type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using the model within the electronic device, from the third audio signal, the type of the environment in which the electronic device is located.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, data obtained from the third audio signal to the external electronic device based on identifying, from the third audio signal, the type of the environment as the first type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information for the third audio signal, obtained using the first multimodal model while in a wake-up state in response to the signal, from the external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor based on identifying, from the third audio signal, that the type of the environment is the second type, acquire a second image through the driven image sensor, acquire a fourth audio signal through the microphone, and obtain fourth data for a fifth audio signal through noise canceling on the fourth audio signal by providing first data for a visual object in the second image corresponding to the user's lips, second data for the second type, and third data for the fourth audio signal to a third multimodal model (e.g., a fourth multimodal model (380)), and transmit the first data and the fourth data to the external electronic device through the communication circuit.
- a third multimodal model e.g., a fourth multimodal model (380)
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information for the fourth audio signal, obtained using the first multimodal model while in a wake-up state in response to the signal, from the external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.
- the electronic device may include a display.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- a method performed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry as described above may include an operation of acquiring a first audio signal via the microphone.
- the method may include an operation of identifying a type of an environment in which the electronic device is located from the first audio signal using a model within the electronic device (e.g., a third model (370)).
- the method may include an operation of transmitting, via the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a first multimodal model (e.g., a second multimodal model (360)) within the external electronic device, based on identifying that the type of the environment is the first type, according to a reference voice command included in the first audio signal.
- a first multimodal model e.g., a second multimodal model (360)
- the method may include an operation of driving the image sensor based on identifying that the type of the environment is the second type.
- the method may include an operation of acquiring a first image via the driven image sensor.
- the method may include an action of acquiring a second audio signal through the microphone.
- the method may include an action of transmitting, to the external electronic device through the communication circuit, a signal that causes a wake-up of the first multimodal model based on identifying the second audio signal including the first image representing the user's reference gesture and the reference voice command using a second multimodal model (e.g., the first multimodal model (350)) within the electronic device.
- a second multimodal model e.g., the first multimodal model (350)
- the method may include activating a timer associated with the image sensor based on identifying that the type of the environment is the second type.
- the method may include acquiring the first image via the driven image sensor before the timer expires based on identifying that the type of the environment is the second type.
- the method may include transmitting, via the communication circuitry, to the external electronic device the signal causing a wake-up of the first multimodal model based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using the second multimodal model based on identifying that the type of the environment is the second type, and maintaining driving the image sensor independently of the expiration of the timer.
- the method may include an action of ceasing to drive the image sensor in response to the expiration of the timer, based on identifying, using the second multimodal model, the first image that does not express the reference gesture and/or the second audio signal that does not include the reference voice command, based on identifying that the type of the environment is the second type.
- the method may include an operation of transmitting the signal based on identifying that the type of the environment is the second type, and then acquiring a third audio signal through the microphone and acquiring a second image through the image sensor.
- the method may include an operation of transmitting, to the external electronic device through the communication circuit, first data acquired using the third audio signal and second data regarding a visual object in the second image corresponding to the lips of the user.
- the method may include an operation of receiving, from the external electronic device through the communication circuit, response information for the third audio signal, the response information being acquired using the first multimodal model in a wake-up state according to the signal.
- the method may include an operation of performing a function according to the response information.
- the method may include an operation of identifying the type of the environment in which the electronic device is located from the third audio signal among the third audio signal and the second image using the model within the electronic device.
- the method may include an operation of obtaining third data for the third audio signal as the first data based on identifying that the type of the environment is the first type.
- the method may include an operation of obtaining data for a fourth audio signal as the first data through noise canceling for the third audio signal by providing the second data, fourth data for the second type, and the third data to a third multimodal model (e.g., a fourth multimodal model (380)) based on identifying that the type of the environment is the second type.
- the method may include an operation of transmitting the first data and the second data to the external electronic device through the communication circuit.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device, via the communication circuit.
- the method may include performing the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, through the communication circuit, response information indicating a text selected by the first multimodal model from among texts pre-stored within the electronic device.
- the method may include performing the function according to the response information by displaying the text indicated by the response information through the display.
- the electronic device may include a speaker.
- the method may include an operation of obtaining an audio signal by applying text-to-speech (TTS) to the received response information.
- TTS text-to-speech
- the method may include an operation of outputting the audio signal through the speaker.
- the electronic device may include a display.
- the method may include an operation of receiving, via the communication circuit, the response information, which includes text generated based on the first data and the second data.
- the method may include an operation of performing the function according to the response information by displaying the text within the response information through the display.
- the method may include an operation of obtaining third data for an area including another visual object corresponding to the user's face within the second image using a face detection model within the electronic device.
- the method may include an operation of obtaining the second data using the third data.
- the method may include an operation of transmitting the signal based on identifying the type of the environment as the first type, and then acquiring a third audio signal through the microphone.
- the method may include an operation of identifying the type of the environment in which the electronic device is located, from the third audio signal, using the model within the electronic device.
- the method may include an operation of transmitting data acquired from the third audio signal to the external electronic device through the communication circuit, based on identifying the type of the environment as the first type, from the third audio signal.
- the method may include an operation of receiving, from the external electronic device through the communication circuit, response information for the third audio signal, the response information being acquired using the first multimodal model in a wake-up state according to the signal.
- the method may include an operation of performing a function according to the response information.
- the method may include an operation of driving the image sensor based on identifying, from the third audio signal, that the type of the environment is the second type, acquiring a second image through the driven image sensor, acquiring a fourth audio signal through the microphone, and providing first data for a visual object in the second image corresponding to the user's lips, second data for the second type, and third data for the fourth audio signal to a third multimodal model to thereby acquire fourth data for a fifth audio signal through noise cancellation on the fourth audio signal, and transmitting the first data and the fourth data to the external electronic device through the communication circuit.
- the method may include an operation of receiving, from the external electronic device through the communication circuit, response information for the fourth audio signal, the response information being acquired using the first multimodal model in a wake-up state according to the signal.
- the method may include an operation of performing a function according to the response information.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device, via the communication circuit.
- the method may include performing the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the method may include receiving, from the external electronic device, through the communication circuit, response information indicating a text selected by the first multimodal model from among texts pre-stored within the electronic device.
- the method may include performing the function according to the response information by displaying the text indicated by the response information through the display.
- the one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit, cause the electronic device to obtain a first audio signal through the microphone.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device (e.g., a third model (370)).
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuitry, a signal to the external electronic device, based on a reference voice command included in the first audio signal, a wake-up signal of a first multimodal model (e.g., a second multimodal model (360)) within the external electronic device, based on identifying that the type of the environment is a first type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to drive the image sensor based on identifying that the type of the environment is a second type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image via the driven image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a second audio signal via the microphone.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuit, to the external electronic device, a signal that causes a wake-up of the first multimodal model based on identifying the first image representing the user's reference gesture and the second audio signal including the reference voice command using a second multimodal model (e.g., the first multimodal model (350)) within the electronic device.
- a second multimodal model e.g., the first multimodal model (350)
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to activate a timer associated with the image sensor based on identifying that the type of the environment is the second type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire the first image via the driven image sensor prior to expiration of the timer.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, via the communication circuitry, to the external electronic device the signal that causes a wake-up of the first multimodal model based on identifying, using the second multimodal model, the first image representing the reference gesture and the second audio signal including the reference voice command, and to continue driving the image sensor independently of the expiration of the timer.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to stop driving the image sensor in response to the expiration of the timer based on identifying, using the second multimodal model, the first image that does not express the reference gesture and/or the second audio signal that does not include the reference voice command.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit the signal based on identifying that the type of the environment is the second type, and then acquire a third audio signal through the microphone and acquire a second image through the image sensor.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, to the external electronic device via the communication circuit, first data acquired using the third audio signal and second data regarding a visual object in the second image corresponding to the user's lips.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, from the external electronic device via the communication circuit, response information for the third audio signal, the first data acquired using the first multimodal model in a wake-up state in response to the signal.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function according to the response information.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the third audio signal among the second image and the third audio signal, the type of the environment in which the electronic device is located, using the model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third data for the third audio signal as the first data based on identifying that the type of the environment is the first type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain data for a fourth audio signal as the first data through noise canceling on the third audio signal, which is performed by providing the second data, fourth data for the second type, and the third data to a third multimodal model based on identifying that the type of the environment is the second type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, from an external electronic device, response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, from the external electronic device through the communication circuit, response information representing text selected by the first multimodal model from among texts pre-stored in the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.
- the electronic device may include a speaker.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to output the audio signal through the speaker.
- TTS text-to-speech
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the second data using the third data.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a third audio signal through the microphone after transmitting the signal based on identifying that the type of the environment is the first type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the third audio signal, the type of the environment in which the electronic device is located, using the model within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuit, data obtained from the third audio signal to the external electronic device based on identifying, from the third audio signal, that the type of the environment is the first type.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, response information for the third audio signal obtained using the first multimodal model in a wake-up state in response to the signal from the external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function in response to the response information.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to drive the image sensor based on identifying, from the third audio signal, that the type of the environment is the second type, acquire a second image through the driven image sensor, acquire a fourth audio signal through the microphone, obtain fourth data for a fifth audio signal through noise canceling on the fourth audio signal by providing first data for a visual object in the second image corresponding to the user's lips, second data for the second type, and third data for the fourth audio signal to a third multimodal model, and transmit the first data and the fourth data to the external electronic device through the communication circuit.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, response information for the fourth audio signal obtained using the first multimodal model in a wake-up state according to the signal from the external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function according to the response information.
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, from an external electronic device, response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.
- the electronic device may include a display.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, from the external electronic device through the communication circuit, response information representing text selected by the first multimodal model from among texts pre-stored in the electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.
- the electronic device may include an image sensor facing a front side of the electronic device.
- the electronic device may include a microphone.
- the electronic device may include communication circuitry.
- the electronic device may include a memory storing instructions and including one or more storage media.
- the electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for a call connection with an external electronic device.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first audio signal through the microphone, based on the detection of performing the call connection with the external electronic device, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone (330) using a model within the electronic device (e.g., a third model (370)).
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second audio signal through the microphone, and to transmit the second audio signal, through the communication circuit, to the external electronic device, based on identifying that the type of the environment is the first type.
- the instructions when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor, acquire an image through the driven image sensor, acquire a third audio signal through the microphone, and provide first data acquired using the third audio signal and second data about a visual object in the image corresponding to the user's lips to a multimodal model (e.g., a first multimodal model (350)) within the electronic device, based on identifying that the type of the environment is a second type, thereby acquiring response information about the third audio signal and the visual object, and performing a function according to the response information.
- a multimodal model e.g., a first multimodal model (350)
- a method performed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit as described above may include an operation of detecting an event for a call connection with an external electronic device.
- the method may include an operation of obtaining a first audio signal through the microphone based on performing the call connection with the external electronic device according to the detection, and identifying a type of an environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device (e.g., a third model (370)).
- the method may include an operation of obtaining a second audio signal through the microphone based on identifying that the type of the environment is the first type, and transmitting the second audio signal to the external electronic device through the communication circuit.
- the method may include an operation of driving the image sensor based on identifying that the type of the environment is the second type, acquiring an image through the driven image sensor, acquiring a third audio signal through the microphone, and providing first data acquired using the third audio signal and second data for a visual object in the image corresponding to the user's lips to a multimodal model (e.g., a first multimodal model (350)) within the electronic device, thereby acquiring response information for the third audio signal and the visual object, and performing a function according to the response information.
- a multimodal model e.g., a first multimodal model (350)
- the one or more programs may include instructions that, when executed by the electronic device having an image sensor, a microphone, and a communication circuit facing a front side of the electronic device, cause the electronic device to detect an event for a call connection with an external electronic device.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first audio signal through the microphone based on performing the call connection with the external electronic device in response to the detection, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device (e.g., a third model (370)).
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, based on identifying that the type of the environment is a first type, acquire a second audio signal through the microphone, and transmit the second audio signal to the external electronic device through the communication circuit.
- the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, based on identifying that the type of the environment is a second type, drive the image sensor, acquire an image through the driven image sensor, acquire a third audio signal through the microphone, and provide first data acquired using the third audio signal and second data regarding a visual object in the image corresponding to the user's lips to a multimodal model (e.g., a first multimodal model (350)) within the electronic device, thereby acquiring response information regarding the third audio signal and the visual object, and performing a function according to the response information.
- a multimodal model e.g., a first multimodal model (350)
- a wearable device may include at least one display (e.g., display (1350)).
- the wearable device may include a microphone (e.g., microphone (1365)).
- the wearable device may include at least one sensor (e.g., at least one sensor 1620).
- the wearable device may include a memory including one or more storage media for storing instructions.
- the wearable device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through pass-through.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal.
- the instructions may cause the wearable device to identify, via the at least one sensor, a first direction of gaze of a user wearing the wearable device and a second direction of gaze of the visual object corresponding to the voice signal.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on the first direction of gaze of the user and the second direction of gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.
- the instructions when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, text generated from the voice signal, superimposed on the screen, based on identifying that the user and the other user corresponding to the visual object are looking at each other.
- a method performed in a wearable device having at least one display (e.g., display (1350)), a microphone (e.g., microphone (1365)), and at least one sensor (e.g., at least one sensor (1620)), as described above, may include an operation of displaying a screen representing at least a part of a three-dimensional space corresponding to a physical environment through a pass-through through the at least one display.
- the method may include an operation of identifying, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen.
- the method may include an operation of identifying, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal.
- the method may include an operation of identifying whether the user and a speaker corresponding to the visual object are looking at each other based on a first direction of the user's gaze and a second direction of the visual object's gaze.
- the method may include an operation of displaying a text generated from the voice signal by overlapping it on the screen through the at least one display based on identifying that the user and the other user corresponding to the visual object are looking at each other.
- the one or more programs may include instructions that, when executed by a wearable device (e.g., a wearable device (1301)) having at least one display (e.g., a display (1350)), a microphone (e.g., a microphone (1365)), and at least one sensor (e.g., at least one sensor (1620)), cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through a pass-through.
- a wearable device e.g., a wearable device (1301)
- at least one display e.g., a display (1350)
- a microphone e.g., a microphone (1365)
- at least one sensor e.g., at least one sensor (1620)
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on the first direction of the gaze of the user and the second direction of the gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.
- the one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to display, through the at least one display, a text generated from the voice signal, overlapping it on the screen, based on identifying that the user and the other user corresponding to the visual object are looking at each other.
- At least one of the components described in one or more of the preceding drawings may be configured to perform one or more operations, techniques, processes, and/or methods as described herein.
- a processor e.g., a baseband processor
- circuitry associated with a user equipment (UE), a base station, a network element, etc., as described above with respect to one or more of the preceding drawings, may be configured to operate according to one or more examples described herein.
- Electronic devices may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
- first,” “second,” or “first” or “second” may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order).
- a component e.g., a first component
- another e.g., a second component
- functionally e.g., a third component
- module used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit.
- a module may be an integral component, or a minimum unit or part of such a component that performs one or more functions.
- a module may be implemented in the form of an application-specific integrated circuit (ASIC).
- ASIC application-specific integrated circuit
- Various embodiments of the present document may be implemented as software (e.g., a program (1240)) including one or more instructions stored in a storage medium (e.g., an internal memory (1236) or an external memory (1238)) readable by a machine (e.g., an electronic device (1201)).
- a processor e.g., a processor (1220) of the machine (e.g., an electronic device (1201)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction.
- the one or more instructions may include code generated by a compiler or code executable by an interpreter.
- the machine-readable storage medium may be provided in the form of a non-transitory storage medium.
- ‘non-transitory’ simply means that the storage medium is a tangible device and does not contain signals (e.g. electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
- the method according to various embodiments disclosed in this document may be provided as included in a computer program product.
- the computer program product may be traded as a product between a seller and a buyer.
- the computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play StoreTM) or directly between two user devices (e.g., smart phones).
- an application store e.g., Play StoreTM
- at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
- each component e.g., a module or a program of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components.
- one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added.
- a plurality of components e.g., a module or a program
- the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration.
- the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Multimedia (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Acoustics & Sound (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Psychiatry (AREA)
- Social Psychology (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Computer Networks & Wireless Communication (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
전자 장치는, 이미지 센서, 마이크로폰, 통신 회로, 인스트럭션들을 저장하는 메모리, 및 프로세싱 회로를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 실행될 시, 이벤트를 검출하고, 제1 이미지를 획득하고, 제1 오디오 신호를 획득하고, 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 제1 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 외부 전자 장치에게 송신하고, 상기 신호를 송신한 후, 제2 오디오 신호를 획득하고, 제2 이미지를 획득하고, 상기 제2 오디오 신호를 이용하여 획득된 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 데이터를, 상기 외부 전자 장치에게 송신하고, 상기 제2 오디오 신호에 대한 응답 정보를 수신하고, 및 기능을 수행하도록, 상기 전자 장치를 야기할 수 있다.
Description
아래의 설명들은, 멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체에 관한 것이다.
전자 장치는 마이크로폰(microphone)을 포함할 수 있다. 상기 전자 장치는, 상기 마이크로폰을 통해 오디오 신호를 획득할 수 있다. 예를 들어, 상기 오디오 신호는, 발화자(speaker)가 발화한(utter) 음성(speech or voice)을 포함할 수 있다. 상기 전자 장치는, 상기 오디오 신호 내의 음성을 인식하는 기능을 제공할 수 있다.
상기 전자 장치는 통신 회로를 포함할 수 있다. 상기 전자 장치는, 상기 통신 회로를 통해 외부 전자 장치에게 데이터를 송신 및/또는 상기 외부 전자 장치로부터 데이터를 수신할 수 있다.
상술한 정보는 본 개시에 대한 이해를 돕기 위한 목적으로 하는 배경 기술(related art)로 제공될 수 있다. 상술한 내용 중 어느 것도 본 개시와 관련된 종래 기술(prior art)로서 적용될 수 있는지에 대하여 어떠한 주장이나 결정이 제기되지 않는다.
전자 장치가 제공된다. 상기 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
방법이 제공된다. 상기 방법은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치 내에서 실행될 수 있다. 상기 방법은, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하는 동작을 포함할 수 있다. 상기 방법은, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
비일시적 컴퓨터 판독가능 저장 매체가 제공된다. 상기 비일시적 컴퓨터 판독가능 저장 매체는, 하나 이상의 프로그램들을 저장할 수 있다. 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
전자 장치가 제공된다. 상기 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 모델을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다.
방법이 제공된다. 상기 방법은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치 내에서 실행될 수 있다. 상기 방법은, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 전자 장치 내의 모델을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하는 동작을 포함할 수 있다. 상기 방법은, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다.
비일시적 컴퓨터 판독가능 저장 매체가 제공된다. 상기 비일시적 컴퓨터 판독가능 저장 매체는, 하나 이상의 프로그램들을 저장할 수 있다. 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 모델을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
전자 장치가 제공된다. 상기 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델을 이용하여, 상기 마이크로폰을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
방법이 제공된다. 상기 방법은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치 내에서 실행될 수 있다. 상기 방법은, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하는 동작을 포함할 수 있다. 상기 방법은, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델을 이용하여, 상기 마이크로폰을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
비일시적 컴퓨터 판독가능 저장 매체가 제공된다. 상기 비일시적 컴퓨터 판독가능 저장 매체는, 하나 이상의 프로그램들을 저장할 수 있다. 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델을 이용하여, 상기 마이크로폰을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
웨어러블 장치가 제공된다. 상기 웨어러블 장치는 적어도 하나의 디스플레이를 포함할 수 있다. 상기 웨어러블 장치는 마이크로폰를 포함할 수 있다. 상기 웨어러블 장치는 적어도 하나의 센서를 포함할 수 있다. 상기 웨어러블 장치는, 인스트럭션들(instructions)을 저장하는 하나 이상의 저장 매체(storage medium)들을 포함하는 메모리를 포함할 수 있다. 상기 웨어러블 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하도록 상기 웨어러블 장치를 야기할 수 있다.
방법이 제공된다. 상기 방법은, 적어도 하나의 디스플레이, 마이크로폰, 및 적어도 하나의 센서를 가지는 웨어러블 장치에서 수행될 수 있다. 상기 방법은, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하는 동작을 포함할 수 있다. 상기 방법은, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하는 동작을 포함할 수 있다.
비일시적 컴퓨터 판독가능 저장 매체가 제공된다. 상기 비일시적 컴퓨터 판독가능 저장 매체는, 하나 이상의 프로그램들을 저장할 수 있다. 상기 하나 이상의 프로그램들은, 적어도 하나의 디스플레이, 마이크로폰, 및 적어도 하나의 센서를 가지는 웨어러블 장치에 의해 실행될 시, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다.
도면의 설명과 관련하여, 동일 또는 유사한 구성 요소에 대해서는 동일 또는 유사한 참조 부호가 사용될 수 있다.
도 1은 전자 장치 및 외부 전자 장치를 포함하는 환경의 예를 도시한다.
도 2a는 전자 장치를 포함하는 노이즈의 볼륨이 사용자의 음성 볼륨보다 큰 환경의 예를 도시한다.
도 2b는 전자 장치를 포함하는 사용자가 발화를 속삭여야 하는 환경의 예를 도시한다.
도 3a는 예시적인 전자 장치의 간소화된 블록도이다.
도 3b는 예시적인 전자 장치 및 외부 전자 장치의 간소화된 다른 블록도이다.
도 4a는 제1 멀티모달 모델의 학습되는 동작의 예를 도시한다.
도 4b는 제4 멀티모달 모델의 학습되는 동작의 예를 도시한다.
도 5는 전자 장치가 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를 송신하는 동작들의 예를 도시한다.
도 6a 및 6b는 적어도 하나의 센서를 구동하기 위한 이벤트의 예를 도시한다.
도 7은 제1 멀티모달 모델이 인식하는 기준 제스쳐 및 기준 음성 명령의 예를 도시한다.
도 8a 및 8b는 전자 장치가 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를 송신하는 동작들의 예를 도시한다.
도면 9a 및 9b는 전자 장치가 제2 멀티모달 모델을 이용하기 위해 외부 전자 장치에게 데이터를 송신하고, 외부 전자 장치로부터 응답 정보를 수신하는 동작들의 예를 도시한다.
도 10a 및 10b는 응답 정보에 따른 기능을 수행하는 전자 장치의 예를 도시한다.
도 11은, 전자 장치가 외부 전자 장치와 통화 연결을 수행하는 예를 도시한다.
도 12는, 다양한 실시예들에 따른, 네트워크 환경 내의 전자 장치의 블록도이다.
도 13a는 웨어러블 장치의 사시도(perspective view)의 일 예를 도시한다.
도 13b는 웨어러블 장치 내에 배치된 하나 이상의 하드웨어들의 일 예를 도시한다.
도 13c는 일 실시예에 따른 웨어러블 장치의 예를 도시한다.
도 14a 내지 도 14b는 웨어러블 장치의 외관의 일 예를 도시한다.
도 15는, 웨어러블 장치의 외관의 일 예를 도시한다.
도 16은 웨어러블 장치의 블록도의 일 예를 도시한다.
도 17은 가상 공간에서 이미지를 표시하기 위한 웨어러블 장치의 블록도의 예를 도시한다.
도 18은 외부 객체로부터 발화된 오디오 신호를 인식하는 웨어러블 장치의 예를 도시한다.
도 19는 사용자 입력에 따라 적응적으로 오디오 신호에 대한 응답을 제공하는 웨어러블 장치의 예를 도시한다.
도 20a 내지 20b는 사용자의 시선에 따라 적응적으로 오디오 신호에 대한 응답을 제공하는 웨어러블 장치의 예를 도시한다.
도 21은 제1 언어로 발화된 오디오 신호에 대하여 제2 언어로 응답을 제공하는 웨어러블 장치의 예를 도시한다.
이하에서는 도면을 참조하여 본 개시의 실시예에 대하여 본 개시가 속하는 기술 분야에서 통상의 지식을 가진 자가 용이하게 실시할 수 있도록 상세히 설명한다. 그러나 본 개시는 여러 가지 상이한 형태로 구현될 수 있으며 여기에서 설명하는 실시예에 한정되지 않는다. 도면의 설명과 관련하여, 동일하거나 유사한 구성요소에 대해서는 동일하거나 유사한 참조 부호가 사용될 수 있다. 또한, 도면 및 관련된 설명에서는, 잘 알려진 기능 및 구성에 대한 설명이 명확성과 간결성을 위해 생략될 수 있다.
본 개시에서 사용되는 용어들은 단지 특정한 실시예를 설명하기 위해 사용된 것으로, 다른 실시예의 범위를 한정하려는 의도가 아닐 수 있다. 단수의 표현은 문맥상 명백하게 다르게 뜻하지 않는 한, 복수의 표현을 포함할 수 있다. 기술적이거나 과학적인 용어를 포함해서 여기서 사용되는 용어들은 본 개시에 기재된 기술 분야에서 통상의 지식을 가진 자에 의해 일반적으로 이해되는 것과 동일한 의미를 가질 수 있다. 본 개시에 사용된 용어들 중 일반적인 사전에 정의된 용어들은, 관련 기술의 문맥상 가지는 의미와 동일 또는 유사한 의미로 해석될 수 있으며, 본 개시에서 명백하게 정의되지 않는 한, 이상적이거나 과도하게 형식적인 의미로 해석되지 않는다. 경우에 따라서, 본 개시에서 정의된 용어일지라도 본 개시의 실시예들을 배제하도록 해석될 수 없다.
이하에서 설명되는 본 개시의 다양한 실시예들에서는 하드웨어적인 접근 방법을 예시로서 설명한다. 하지만, 본 개시의 다양한 실시예들에서는 하드웨어와 소프트웨어를 모두 사용하는 기술을 포함하고 있으므로, 본 개시의 다양한 실시예들이 소프트웨어 기반의 접근 방법을 제외하는 것은 아니다.
이하 설명에서 사용되는 데이터를 지칭하는 용어(예: 데이터, 정보, 센싱 데이터, 신호), 값을 지칭하는 용어(예: 임계값, 기준 값), 연산 상태를 위한 용어(예: 동작(operation), 프로세스), 객체를 지칭하는 용어(예: 시각적 객체, 이모지 그래픽적 객체), 네트워크 객체(network entity)들을 지칭하는 용어, 장치의 구성 요소를 지칭하는 용어 등은 설명의 편의를 위해 예시된 것이다. 따라서, 본 개시가 후술되는 용어들에 한정되는 것은 아니며, 동등한 기술적 의미를 가지는 다른 용어가 사용될 수 있다. 또한, 이하 사용되는 '...부', '...기', '...물', '...체' 등의 용어는 적어도 하나의 형상 구조를 의미하거나 또는 기능을 처리하는 단위를 의미할 수 있다.
또한, 본 개시에서, 특정 조건의 만족(satisfied), 충족(fulfilled) 여부를 판단하기 위해, 초과 또는 미만의 표현이 사용될 수 있으나, 이는 일 예를 표현하기 위한 기재일 뿐 이상 또는 이하의 기재를 배제하는 것이 아니다. '이상'으로 기재된 조건은 '초과', '이하'로 기재된 조건은 '미만', '이상 및 미만'으로 기재된 조건은 '초과 및 이하'로 대체될 수 있다. 또한, 이하, 'A' 내지 'B'는 A부터(A 포함) B까지의(B 포함) 요소들 중 적어도 하나를 의미한다. 이하, 'C' 및/또는 'D'는 'C' 또는 'D' 중 적어도 하나, 즉, {'C', 'D', 'C'와 'D'}를 포함하는 것을 의미한다.
도 1은 전자 장치 및 외부 전자 장치를 포함하는 환경의 예를 도시한다.
도 1을 참고하면, 전자 장치(120)는 마이크로폰을 포함할 수 있다. 예를 들면, 상기 마이크로폰은 저 전력 소비를 위한 상태 내에서 동작할 수 있다. 예를 들면, 상기 마이크로폰은, 상기 저 전력 소비를 위한 상기 상태 내에서 있는 동안, 오디오 신호(또는 사운드 신호)를 검출할 수 있다. 예를 들면, 전자 장치(120)는, 상기 마이크로폰이 상기 저 전력 소비를 위한 상기 상태 내에서 있는 동안, 상기 마이크로폰을 통해, 오디오 신호(또는 사운드 신호)를 획득할 수 있다. 예를 들면, 전자 장치(120)는 상기 마이크로폰을 통해 외부 환경으로부터 오디오 신호를 획득할 수 있다. 예를 들면, 상기 오디오 신호는 사용자(110)에 의해 발화된 음성을 포함할 수 있다. 예를 들면, 상기 오디오 신호는 노이즈를 포함할 수 있다.
예를 들면, 전자 장치(120)는 상기 오디오 신호를 인식할 수 있다. 예를 들면, 전자 장치(120)는, 상기 오디오 신호 내의 음성 신호를 인식할 수 있다. 예를 들면, 전자 장치(120)는 음성 인식 기능을 제공할 수 있다. 예를 들면, 전자 장치(120)는 인식된 오디오 신호를 이용하여 데이터를 획득할 수 있다. 예를 들면, 상기 데이터는 상기 오디오 신호에 대응할 수 있다. 예를 들면, 상기 데이터는 상기 오디오 신호에 후처리가 수행된 오디오 신호를 포함할 수 있다. 예를 들면, 상기 후처리는, 모델(예: 인공지능 모델)에게 상기 오디오 신호를 입력하기 위하여 수행될 수 있다. 예를 들면, 상기 데이터는 상기 오디오 신호를 나타내는 텍스트를 포함할 수 있다. 예를 들면, 상기 데이터는 상기 오디오 신호에 포함된 음성 신호(또는 보이스 신호)를 나타내는 텍스트를 포함할 수 있다.
예를 들면, 전자 장치(120)는 음성 인식 기능을 위하여 모델(model)을 이용할 수 있다. 예를 들면, 전자 장치(120)는, 상기 모델을 훈련(또는 학습)할 수 있다. 예를 들면, 전자 장치(120)가 상기 훈련된 모델을 이용하여 음성 신호를 인식하는 경우에, 전자 장치(120)가 상기 훈련된 모델을 이용하지 않고 음성 신호를 인식하는 경우 보다 정확하게 음성 신호를 인식할 수 있다. 예를 들면, 전자 장치(120)는, 음성 인식 기능을 이용하기 위하여, 외부 전자 장치(130)내의 다른 모델과 연결될 수 있다. 예를 들면, 전자 장치(120)는 통신 회로를 포함할 수 있다. 예를 들면, 전자 장치(120)는 상기 오디오 신호를, 상기 통신 회로를 통해, 외부 전자 장치(130)에게 송신할 수 있다. 예를 들면, 전자 장치(120)는 상기 오디오 신호를 이용하여 획득된 데이터를, 상기 통신 회로를 통해, 외부 전자 장치(130)에게 송신할 수 있다.
예를 들면, 전자 장치(120)는 스마트폰을 포함할 수 있다. 예를 들면, 전자 장치(120)는 웨어러블 장치를 포함할 수 있다. 예를 들면, 전자 장치(120)는 스마트 워치를 포함할 수 있다. 제한하지 않는 예로, 전자 장치(120)는, 스마트폰, 테블릿, 랩탑 컴퓨터, 스마트 워치와 같은 포터블 전자 장치를 포함할 수 있다. 예를 들면, 전자 장치(120)는, 다기능 장치 또는 사용자 장치로 설명될 수 있다.
예를 들면, 외부 전자 장치(130)는, 음성 신호를 인식할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 음성 신호를 인식하기 위한 모델을 이용할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 모델을 훈련할 수 있다. 예를 들면, 외부 전자 장치(130)는, 훈련된 상기 모델을 이용하여, 음성 신호를 보다 정확하게 인식할 수 있다. 예를 들면, 외부 전자 장치(130)의 모델은, 전자 장치(120)의 모델 보다 복잡도가 클 수 있다.
예를 들면, 외부 전자 장치(130)는, 상기 인식된 음성 신호에 대한 응답 정보를 출력(또는 획득)할 수 있다. 예를 들면, 외부 전자 장치(130)는 상기 모델을 이용하여 음성 신호를 인식하고, 상기 음성 신호에 기반하는 프롬프트를 생성하여, 외부 전자 장치(130) 내의 대형 언어 모델에게 제공할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 대형 언어 모델을 이용하여 상기 음성 신호에 대한 상기 응답 정보를 출력하거나 획득할 수 있다.
예를 들면, 외부 전자 장치(130)는 통신 회로를 포함할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 통신 회로를 통해, 전자 장치(120)로부터 신호(또는 데이터)를 수신할 수 있다. 예를 들면, 외부 전자 장치(130)는, 전자 장치(120)로부터 외부 전자 장치(130) 내의 모델의 웨이크-업을 야기하는 신호를 수신할 수 있다. 예를 들면, 상기 신호는, 저 전력 소비를 위한 상태 내에서 있는 상기 모델을 활성화 상태로 변경할 수 있다. 예를 들면, 상기 신호는, 상기 저 전력 소비를 위한 상기 상태 내에서 있는 상기 모델을 동작시키는 신호로 참조될 수 있다.
예를 들면, 외부 전자 장치(130)는, 전자 장치(120)로부터 음성 신호를 수신하거나 획득할 수 있다. 예를 들면, 외부 전자 장치(130)는, 전자 장치(120)로부터, 상기 음성 신호를 이용하여 획득된 데이터를, 수신하거나 획득할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 획득된 음성 신호를 상기 모델을 이용하여 인식하고, 프롬프트를 생성할 수 있다. 예를 들면, 외부 전자 장치(130)는, 상기 프롬프트를 상기 대형 언어 모델에게 제공함으로써, 응답 정보를 획득할 수 있다.
예를 들면, 상태(100) 내에서, 사용자(110)는 전자 장치(120)의 음성 인식 기능을 동작하기 위하여, 기준 음성 명령(또는 호출어)을 발화할 수 있다. 예를 들면, 전자 장치(120)는 상기 마이크로폰을 통해 상기 기준 음성 명령을 포함하는 오디오 신호를 획득할 수 있다. 예를 들면, 전자 장치(120)는, 획득된 오디오 신호 내의 노이즈 및 상기 기준 음성 명령에 기반하여, 상기 음성 인식 기능의 동작을 야기할 수 있다.
도 2a는 전자 장치를 포함하는 노이즈의 볼륨이 사용자의 음성 볼륨보다 큰 환경의 예를 도시한다.
도 2a를 참고하면, 환경(200)은, 상대적으로 노이즈의 볼륨이 큰 환경으로 설명될 수 있다. 예를 들면 환경(200)은 실내 환경을 포함할 수 있다. 예를 들면 환경(200)은 실외 환경을 포함할 수 있다. 제한하지 않는 예로, 환경(200)은, 카페, 식당, 지하철, 공사장, 광장, 콘서트 홀을 포함할 수 있으나, 실시예가 제한되는 것은 아니다.
예를 들면, 환경(200)은, 전자 장치(120)의 상기 마이크로폰을 통해 획득된 오디오 신호의 SNR(signal to noise ratio)이 기준 값 이하인 환경을 나타낼 수 있다. 예를 들면, 환경(200)에서 상기 오디오 신호의 음질은 상대적으로 낮을 수 있다. 예를 들면, 환경(200)은, 노이즈 볼륨에 관계하여 사용자(110)의 음성의 볼륨이 낮은 환경을 나타낼 수 있다. 예를 들면, 환경(200)에서 노이즈의 볼륨이 클수록, 전자 장치(120)의 음성 인식 기능의 품질은 떨어질 수 있다.
도 2b는 전자 장치를 포함하는 사용자가 발화를 속삭여야 하는 환경의 예를 도시한다.
도 2b를 참고하면, 환경(210)은, 사용자(110)가 상대적으로 큰 볼륨의 음성을 발화할 수 없는 환경을 나타낼 수 있다. 예를 들면, 환경(210)은, 도서관을 나타낼 수 있다. 예를 들면 환경(210)은 도서관으로 제한되는 것은 아니다. 환경(210)은 회의실, 영화관, 미술관을 포함할 수 있으나, 실시예가 제한되는 것은 아니다.
예를 들면, 환경(210)에서, 사용자(110)는 전자 장치(120)에게 속삭일 수 있다. 예를 들면, 환경(210)에서, 사용자(110)의 음성 볼륨은 전자 장치(120)가 인식할 수 있는 기준 볼륨보다 낮을 수 있다. 예를 들면, 사용자(110)의 상기 음성 볼륨이 상기 기준 볼륨보다 낮기 때문에, 전자 장치(120)의 인식 품질이 떨어질 수 있다. 예를 들면, 환경(210)에서, 전자 장치(120)는 사용자(110)의 음성을 인식하지 못할 수 있다.
예를 들면, 도 2a의 환경(200)에서 음성 인식 기능의 품질이 상대적으로 떨어지기 때문에, 전자 장치(120)의 마이크로폰을 통해 수신된 오디오 신호의 노이즈를 제거하는 방법 및/또는 오디오 신호 내의 음성 신호를 강화하는 방법이 요구될 수 있다. 예를 들면, 전자 장치(120)의 적어도 하나의 센서를 이용하여 음성 인식을 보조하는 방법이 요구될 수 있다.
예를 들면, 도 2b의 환경(210)에서, 사용자(110)의 음성 볼륨이 전자 장치(120)가 인식할 수 있는 상기 기준 볼륨 보다 낮기 때문에, 전자 장치(120)의 음성 인식 기능을 동작하기 위하여, 기준 제스쳐를 인식하는 방법이 요구될 수 있다. 예를 들면, 환경(210)에서, 사용자(110)의 입술의 모양을 인식하는 방법이 요구될 수 있다. 예를 들면, 환경(210)에서, 전자 장치(120)의 적어도 하나의 센서를 통해 획득된 센싱 데이터가 기준 센싱 데이터에 대응하는지 여부를 식별하는 방법이 요구될 수 있다.
이러한 방법들은, 아래에서 예시되는 전자 장치 내에서 실행될 수 있다. 예를 들면, 아래에서 예시되는 전자 장치는, 이러한 방법들을 제공하기 위한 구성요소들(또는 하드웨어 구성요소들)을 포함할 수 있다. 상기 구성요소들은, 도 3a를 참조하여 보다 상세히 설명되고 예시된다.
도 3a는 예시적인 전자 장치의 간소화된 블록도이다. 전자 장치(301)는 도 1의 전자 장치(101)의 일 예일 수 있다.
도 3a를 참고하면, 전자 장치(301)는, 적어도 하나의 프로세서(300), 통신 회로(310), 메모리(320), 마이크로폰(330), 적어도 하나의 센서(340), 디스플레이(311), 및/또는 스피커(312)를 포함할 수 있다. 예를 들면, 적어도 하나의 프로세서(300), 통신 회로(310), 메모리(320), 마이크로폰(330), 적어도 하나의 센서(340), 디스플레이(311), 및/또는 스피커(312)는 통신 버스(a communication bus)에 의해 서로 전기적으로 및/또는 작동적으로 연결될 수 있다(electronically and/or operably coupled with each other). 이하에서, 하드웨어 컴포넌트들이 작동적으로 결합된 것은, 하드웨어 컴포넌트들 중 제1 하드웨어 컴포넌트에 의해 제2 하드웨어 컴포넌트가 제어되도록, 하드웨어 컴포넌트들 사이의 직접적인 연결, 또는 간접적인 연결이 유선으로, 또는 무선으로 수립된 것을 의미할 수 있다. 도 3a에 도시된 하드웨어 컴포넌트들은 상이한 블록들에 기반하여 도시되었으나, 본 개시가 이에 제한되는 것은 아니다. 예를 들면, 도 3a에 도시된 하드웨어 컴포넌트 중 일부분(예: 적어도 하나의 프로세서(300), 통신 회로(310), 및 메모리(320)의 적어도 일부분)이 SoC(system on chip) 또는 SIP(system in package)와 같은 단일 집적 회로(single integrated circuit)에 포함될 수 있다. 전자 장치(301)에 포함되는 하드웨어 컴포넌트의 타입 및/또는 개수는 도 3a에 도시된 바에 제한되지 않는다. 예를 들면, 전자 장치(301)는, 도 3a에 도시된 하드웨어 컴포넌트들 중 일부만 포함할 수 있다.
적어도 하나의 프로세서(300)는 인스트럭션들을 실행하는 것에 기반하여 데이터를 처리하기 위한 하드웨어 컴포넌트를 포함할 수 있다. 데이터를 처리하기 위한 하드웨어 컴포넌트는, 예를 들면, CPU(central processing unit)(예: 프로세싱 회로를 포함함)를 포함할 수 있다. 예를 들면, 데이터를 처리하기 위한 하드웨어 컴포넌트는, GPU(graphic processing unit)(예: 프로세싱 회로를 포함함)를 포함할 수 있다. 예를 들면, 데이터를 처리하기 위한 하드웨어 컴포넌트는, DPU(display processing unit)(예: 프로세싱 회로를 포함함)를 포함할 수 있다. 예를 들면, 데이터를 처리하기 위한 하드웨어 컴포넌트는, NPU(neural processing unit)(예: 프로세싱 회로를 포함함)를 포함할 수 있다.
적어도 하나의 프로세서(300)는 하나 이상의 코어(core)들을 포함할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 듀얼 코어(dual core), 쿼드 코어(quad core) 또는 헥사 코어(hexa core)와 같은 멀티-코어 프로세서의 구조를 가질 수 있다. 적어도 하나의 프로세서(300)는, 도 12의 프로세서(1220)에 대한 내용이 실질적으로 동일하게 적용될 수도 있다.
통신 회로(310)는 전자 장치(301)와 외부 전자 장치(302) 사이에서 신호의 송신 및/또는 수신을 지원하기 위한 하드웨어 컴포넌트를 포함할 수 있다. 통신 회로(310)는, 예를 들면, 모뎀(MODEM, modulator and demodulator), 안테나, O/E(optic/electronic) 변환기 중 적어도 하나를 포함할 수 있다. 통신 회로(310)는, 이더넷(ethernet), LAN(local area network), WAN(wide area network), WiFi(wireless fidelity), Bluetooth, BLE(Bluetooth low energy), zigbee, LTE(long term evolution), 5G NR(new radio)와 같은 다양한 타입의 프로토콜에 기반하여 전기 신호의 송신 및/또는 수신을 지원할 수 있다. 도 3a의 통신 회로(310)에 대한 구체적인 내용은, 도 12의 통신 모듈(1290), 및/또는 안테나 모듈(1297)이 실질적으로 동일하게 적용될 수 있다.
메모리(320)는 적어도 하나의 프로세서(300)에 입력되거나 및/또는 적어도 하나의 프로세서(300)로부터 출력되는 데이터, 및/또는 인스트럭션들을 저장하기 위한 하드웨어 컴포넌트를 포함할 수 있다. 예를 들면, 인스트럭션들은 전자 장치(301)의 적어도 하나의 프로세서(300)가 데이터에 수행할 연산 및/또는 동작을 나타낼 수 있다. 메모리(320)는, 예를 들면, RAM(random-access memory)과 같은 휘발성 메모리(volatile memory) 및/또는 ROM(read-only memory)과 같은 비휘발성 메모리(non-volatile memory)를 포함할 수 있다. 휘발성 메모리는, 예를 들면, DRAM(dynamic RAM), SRAM(static RAM), cache RAM, PSRAM (pseudo SRAM) 중 적어도 하나를 포함할 수 있다. 비휘발성 메모리는, 예를 들면, PROM(programmable ROM), EPROM (erasable PROM), EEPROM (electrically erasable PROM), 플래시 메모리, 하드디스크, 컴팩트 디스크, EMMC(embedded multimedia card) 중 적어도 하나를 포함할 수 있다. 도 3a의 메모리(320)에 대한 구체적인 내용은, 도 12의 메모리(1230)에 대한 내용이 실질적으로 동일하게 적용될 수 있다.
마이크로폰(330)은, 전자 장치(301) 주변에서 야기된 오디오 신호를 획득하도록 구성될 수 있다. 예를 들면, 마이크로폰(330)은, 오디오 신호 내의 음성 신호를 획득하기 위해 이용될 수 있다. 도 3a의 마이크로폰(330)에 대한 구체적인 내용은, 도 12의 입력 모듈(1250)에 대한 내용이 실질적으로 동일하게 적용될 수 있다.
적어도 하나의 센서(340), 외부 신호를 획득하기 위해 이용되는 전자 장치(301)의 하드웨어 컴포넌트를 포함할 수 있다. 예를 들면, 적어도 하나의 센서(340)는 외부 신호를 감지하고, 감지된 상태에 대응하는 전기 신호 또는 데이터 값을 생성할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 이미지 센서를 포함할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 심박 센서를 포함할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 가속도 센서를 포함할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 자이로 센서를 포함할 수 있다. 하지만 이에 제한되지 않는다. 도 3a의 적어도 하나의 센서(340)에 대한 구체적인 내용은, 도 12의 센서 모듈(1276)에 대한 내용이 실질적으로 동일하게 적용될 수 있다.
예를 들면, 전자 장치(301)는, 적어도 하나의 센서(340)를 통해 센싱 데이터(예: 이미지(예: 이미지 센서를 통해 획득됨), 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다.
디스플레이(311)는, 화면을 표시하기 위해 이용되는 전자 장치(301)의 하드웨어 컴포넌트를 포함할 수 있다. 예를 들면, 디스플레이(311)는, 발광 소자들 및 광을 발광하도록 상기 발광 소자들을 제어하는 회로들(예: 트랜지스터들)을 포함할 수 있다. 예를 들면, 상기 발광 소자들 각각은, OLED(organic light emitting diode) 또는 마이크로 LED를 포함할 수 있다. 하지만, 이에 제한되지 않는다. 예를 들면, 디스플레이(311)는, LCD(liquid crystal display)를 포함할 수 있다. 도 3a의 디스플레이(311)에 대한 구체적인 내용은, 도 12의 디스플레이 모듈(1260)에 대한 내용이 실질적으로 동일하게 적용될 수 있다.
스피커(312)는, 음성 신호를 전자 장치(301)의 외부로 출력하기 위해 이용될 수 있다. 예를 들면, 스피커(312)는 멀티미디어 재생 또는 녹음 재생과 같이 일반적인 용도로 사용될 수 있다. 도 3a의 스피커(312)에 대한 구체적인 내용은, 도 12의 음향 출력 모듈(1255)에 대한 내용이 실질적으로 동일하게 적용될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)는 통해 획득된 센싱 데이터 및/또는 마이크로폰(330)을 통해 획득된 오디오 신호를 이용하여 외부 전자 장치(302) 내의 멀티모달 모델을 웨이크-업하는 동작들을 전자 장치(301) 내에서 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 유형을 식별하는 동작들을 전자 장치(301) 내에서 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형에 따라, 마이크로폰(330)을 통해 획득된 다른 오디오 신호에 대한 노이즈 캔슬링을 수행하는 동작들을 전자 장치(301) 내에서 실행할 수 있다. 상기 동작들은 도 5 내지 도 10b의 설명 내에서 예시될 것이다.
도 3b는 예시적인 전자 장치 및 외부 전자 장치의 간소화된 다른 블록도이다.
도 3b를 참고하면, 전자 장치(301)는, 제1 멀티모달 모델(350), 제3 모델(370), 및/또는 제4 멀티모달 모델(380)을 포함할 수 있다.
제1 멀티모달 모델(350)은, 마이크로폰(330)을 통해 획득된 오디오 신호 내에 포함된 보이스 신호를 식별하기 위해 이용될 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 마이크로폰(330)을 통해 획득된 오디오 신호 내의 기준 음성 명령이 포함되어 있는지 여부를 식별하기 위해 이용될 수 있다. 제1 멀티모달 모델(350)은, 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터(예: 이미지(예: 이미지 센서를 통해 획득됨), 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))가 기준 센싱 데이터를 나타내는지(또는 포함되어 있는지) 여부를 식별할 수 있다. 예를 들면, 기준 센싱 데이터는 사용자의 기준 제스쳐를 포함할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 이미지 센서를 통해 획득된 이미지가 기준 제스쳐를 나타내는지 여부를 식별할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 심박 센서를 통해 획득된 정보가 기준 심박 정보를 나타내는지 여부를 식별할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 가속도 센서 및/또는 자이로 센서를 통해 획득된 정보가 기준 움직임을 나타내는지 여부를 식별할 수 있다.
예를 들면, 제1 멀티모달 모델(350)은, 하나 이상의 모달리티(modality)를 이용하여 훈련(또는 학습)될 수 있다. 예를 들면, 모달리티는 데이터의 유형으로 참조될 수 있다. 예를 들면, 데이터의 유형은, 텍스트, 이미지, 및/또는 오디오 신호를 포함할 수 있다. 하지만 이에 제한되지 않는다.
예를 들면, 제1 멀티모달 모델(350)은, ASR(automatic speech recognition)기능을 제공할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 마이크로폰(330)을 통해 획득된 오디오 신호를 인식하거나 식별하기 위하여 학습되어 있을 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 마이크로폰(330)을 통해 획득된 오디오 신호에 대한 자연어 처리를 위해 학습되어 있을 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터를 인식하거나 식별하기 위하여 학습되어 있을 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 오디오 신호 및/또는 센싱 데이터에 기반하여, 오디오 신호에 포함된 보이스 신호(또는 음성 신호)를 나타내는 데이터를 획득하기 위해 이용될 수 있다. 예를 들면, 보이스 신호를 나타내는 데이터는, 보이스 신호를 나타내는 텍스트를 포함할 수 있다. 예를 들면, 보이스 신호를 나타내는 데이터는, 보이스 신호를 나타내는 오디오 신호를 포함할 수 있다. 제1 멀티모달 모델(350)의 학습은, 도 4a를 참조하여 보다 상세히 설명되고 예시될 것이다.
도 3b에서 외부 전자 장치(302)가, 전자 장치(301)로부터 송신된 제1 멀티모달 모델(350)의 출력 신호를 외부 전자 장치(302) 내의 제2 멀티 모달 모델(360)에게 제공하는 것으로 도시하고 있으나, 이는 단지 예시적인 것이다. 예를 들면, 외부 전자 장치(302)는, 전자 장치(301)로부터 송신된 제1 멀티모달 모델(350)의 출력 신호를 외부 전자 장치(302) 내의 대형 언어 모델(390)에게, 제공할 수 있다. 예를 들면, 제1 멀티모달 모델(350)의 출력 신호는, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보를 포함할 수 있다.
제한되지 않는 예로, 전자 장치(301)는, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보를 이용하여, 상기 응답 정보에 따른 기능을 실행함으로써, 오디오 신호 및/또는 센싱 데이터에 대한 응답을 제공할 수 있다. 예를 들면, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보는, 오디오 신호 및/또는 센싱 데이터에 기반하여, 전자 장치(301)가 응답을 제공하기 위해 이용될 수 있다. 예를 들면, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보는, 오디오 신호에 포함된 보이스 신호를 나타내는 텍스트를 디스플레이(예: 디스플레이(311))를 통해 표시하기 위해 이용될 수 있다. 예를 들면, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보는, 오디오 신호에 포함된 보이스 신호를 나타내는 오디오 신호를 스피커(예: 스피커(312))를 통해 출력하기 위해 이용될 수 있다. 예를 들면, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보는, 오디오 신호 내에 포함된 기준 음성 명령에 대응하는(또는 연계된) 기능을 실행하기 위해 이용될 수 있다. 예를 들면, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보는, 센싱 데이터에 의해 표현되는 기준 제스쳐에 대응하는(또는 연계된) 기능을 실행하기 위해 이용될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 획득된 오디오 신호 및 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터(예: 이미지(예: 이미지 센서를 통해 획득됨), 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 제1 멀티모달 모델(350)에게 제공할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제1 멀티모달 모델(350)의 ASR 기능을 이용하여 오디오 신호 및 센싱 데이터에 대한 응답 정보를 획득할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 획득된 오디오 신호 및 적어도 하나의 센서(340)를 통해 획득된 이미지를 제1 멀티모달 모델(350)에게 제공할 수 있다. 예를 들면, 획득된 이미지는, 전자 장치(301)의 사용자의 입술에 대응하는 시각적 객체를 포함할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제1 멀티모달 모델(350)의 ASR 기능을 이용하여, 오디오 신호 및 사용자의 입술에 대응하는 이미지 내의 시각적 객체에 대한 응답 정보를 획득할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 획득된 응답 정보를, 통신 회로(310)를 통해 외부 전자 장치(302)에게 송신할 수 있다. 예를 들면, 응답 정보는, 외부 전자 장치(302) 내의 대형 언어 모델(390)에게 입력하기 위한 데이터를 포함할 수 있다. 예를 들면, 대형 언어 모델(390)에게 입력하기 위한 데이터는, 프롬프트로 나타내어지거나 이용될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 획득된 응답 정보에 대한 기능을 전자 장치(301) 내에서 실행할 수 있다. 예를 들면, 획득된 응답 정보에 대한 기능은, 획득된 응답 정보에 대응하는 텍스트를 디스플레이(311)에 표시하는 동작을 포함할 수 있다. 예를 들면, 획득된 응답 정보에 대한 기능은, 획득된 응답 정보에 대응하는 이모지를 디스플레이(311)에 표시하는 동작을 포함할 수 있다. 하지만 이에 제한하지 않는다.
제3 모델(370)은 마이크로폰(330)을 통해 획득된 오디오 신호로부터, 전자 장치(301)가 위치된 환경의 유형을 식별(또는 분류)하기 위해 이용될 수 있다. 예를 들면, 제3 모델(370)은, 상기 오디오 신호의 SNR(signal to noise ratio)이 기준 값 이상인 경우, 상기 환경의 상기 유형을 제1 유형으로 식별할 수 있다. 예를 들면, 제3 모델(370)은, 상기 오디오 신호의 SNR이 상기 기준 값 미만인 경우, 상기 환경의 상기 유형을 제2 유형으로 식별할 수 있다.
예를 들면, 제3 모델(370)은, 상기 환경의 상기 유형을 식별할 수 있다. 예를 들면, 제3 모델(370)은, 식별된 상기 환경의 상기 유형에 대한 데이터를 제1 멀티모달 모델(350)에게 제공할 수 있다. 예를 들면, 제3 모델(370)은, 상기 환경의 상기 유형을 제2 유형으로 식별하는 것에 기반하여, 상기 제2 유형에 대한 데이터를 제4 멀티모달 모델(380)에게 제공할 수 있다. 예를 들면, 제3 모델(370)은, 컨포머(conformer) 알고리즘으로 학습될 수 있다. 하지만 이에 제한하지 않는다.
예를 들면, 전자 장치(301)가 위치된 상기 환경의 상기 유형은 실내 환경, 실외 환경, 상기 제1 유형의 환경, 및/또는 상기 제2 유형의 환경을 포함할 수 있다. 하지만 이에 제한하지 않는다.
제4 멀티모달 모델(380)은, 마이크로폰(330)을 통해 획득된 오디오 신호에 대하여 노이즈 캔슬링을 수행하기 위해 이용될 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 입력된 상기 오디오 신호 내의 노이즈를 제거한 오디오 신호를 생성할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 마이크로폰(330)을 통해 획득된 오디오 신호, 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터, 및/또는 제3 모델로부터 제공받은 상기 환경의 상기 유형에 대한 데이터를 이용하여 상기 획득된 오디오 신호에 대한 노이즈 캔슬링을 수행할 수 있다.
예를 들면, 제4 멀티모달 모델(380)은, 상기 획득된 오디오 신호에 대한 노이즈 캔슬링을 수행하기 위하여 학습되어 있을 수 있다. 예를 들면, 제4 멀티모달 모델(380)은, 학습된 인공 지능 모델을 나타낼 수 있다. 제4 멀티모달 모델(380)의 학습은, 도 4b를 참조하여 보다 상세히 설명되고 예시될 것이다.
외부 전자 장치(302)는 제2 멀티모달 모델(360), 및/또는 대형 언어 모델(390)을 포함할 수 있다.
제2 멀티모달 모델(360)은 제1 멀티모달 모델(350)과 동일한 기능을 제공할 수 있다. 예를 들면, 제2 멀티모달 모델(360)은 ASR기능을 제공할 수 있다. 예를 들면, 제2 멀티모달 모델(360)은 제1 멀티모달 모델(350)과 대응할 수 있다. 예를 들면, 제2 멀티모달 모델(360)은 하나 이상의 모달리티들을 이용하여 훈련(또는 학습)될 수 있다. 예를 들면, 제2 멀티모달 모델(360)은, 제1 멀티모달 모델(350)과 같은 알고리즘으로 학습될 수 있다. 예를 들면, 제2 멀티모달 모델(360)은, 제1 멀티모달 모델(350) 보다 복잡도가 클 수 있다. 예를 들면, 제2 멀티모달 모델(360)은, 제1 멀티모달 모델(350) 보다 레이어의 개수가 많을 수 있다. 예를 들면, 제2 멀티모달 모델(360)은 제1 멀티모달 모델(350)보다 파라미터의 개수가 많을 수 있다.
대형 언어 모델(390)은, 방대한 양의 텍스트 데이터로 사전 학습된, 인공신경망으로 구성된 언어 모델로 참조될 수 있다. 대형 언어 모델(390)은, 기존의 일반 언어 모델에 비해 10배 이상 많은 매개 변수(예를 들어, 1000억 개 이상의 매개 변수)를 포함할 수 있다. 대형 언어 모델(390)은, 집중 메커니즘(attention mechanism)을 기반으로 하는 트랜스포머(transformer) 인공 신경망 구조를 사용할 수 있다. 집중 메커니즘(attention mechanism)은, 인공지능 모델이 입력 데이터 내에서, 중요한 부분에 집중(attention)할 수 있게 돕는 기술이다. 집중 메커니즘(attention mechanism)은, 시계열 입력 데이터(예: 음성, 동영상과 같은 입력 데이터 또는 신경망 일부 층의 입력 데이터)의 적어도 일부에 신경망 중간 또는 최종 출력에 기여하는 정도를 예측하여, 출력 데이터 예측에 이용될 수 있다. 시퀀스의 각 요소를 순차적으로 처리하는 순환 신경망(RNN) 구조는 긴 시계열 거리 간 정보 의존도(dependency)가 있는 경우에 대한 예측 성능이 저하되지만, 집중 메커니즘(attention mechanism)은 입력 데이터의 전체적인(또는 일부) 컨텍스트(context) 내에서 가중치 집중(attention) 정도를 제어함으로써 긴 시계열 거리 간 정보 의존도를 고려할 수 있다.
예를 들어, 대형 언어 모델(390)은, 트랜스포머는 인코더-디코더 구조를 포함할 수 있다. 인코더는 입력 데이터를 처리하여 압축 정보(예: 집중 메커니즘(attention mechanism))를 출력하고, 디코더는 압축 정보를 처리하여 출력 데이터를 토큰 단위로 출력할 수 있다. 인코더, 디코더 각각은 독립적인 집중 네트워크(attention network)를 포함할 수 있고, 인코더-디코더를 연결하는 교차 집중 네트워크(cross-attention network)를 포함할 수 있다.
예를 들어, 대형 언어 모델(390)은, 사전 학습과 파인 튜닝(fine-tuning)이라는 두 가지 단계로 학습될 수 있다. 사전 학습은, 대형 언어 모델(390)이 방대한 양의 텍스트 데이터를 처리하고 일반적인 언어 지식을 습득하도록 하는 과정으로, 예를 들어, 텍스트 열(text sequence)의 이전 단어 열을 이용하여 다음 단어를 예측하도록 자기 지도 학습(self-supervised learning)을 포함할 수 있다. 파인 튜닝은 대형 언어 모델(390)이 특정 도메인(예: chatbot, 번역, 요약, Q&A)이나 작업에 적합하도록 훈련하는 과정으로, 사전 학습된 모델을 기반으로 도메인 목적에 맞는 데이터셋을 이용하여 추가적으로 지도 학습(또는 적응 학습)될 수 있다. 대형 언어 모델(390)은 프롬프트라는 자연어를 포함하는 텍스트 입력으로 작업을 수행할 수 있다. 예를 들어, 대형 언어 모델(390)은, BERT(bidirectional encoder representations from transformer), GPT(generative pre-trained transformer)를 포함할 수 있다. 'LLM(lagrge language model)'이라는 표현은 신경망 모델 자체를 참조할 수 있으나, LLM 기반의 어플리케이션(예: chatbot, 번역, 요약, 텍스트 분류, 문장 생성)의 모델을 의미할 수도 있다. 예를 들어, chatGPT와 같은 LLM 기반의 chatbot 또한 LLM으로 참조될 수 있다. 'LLM'은 또한 LLM 신경망 모델을 이용한 추론 엔진을 포함하는 것일 수도 있다. 예를 들어, "입력 프롬프트를 LLM에 입력한다"는 것은 "입력 프롬프트를 LLM에 기반한 추론 엔진에 입력한다"는 것으로 참조될 수 있다."
대형 언어 모델(390)은, 텍스트 데이터로 학습된 인공 지능 모델로서, 오디오 신호 내의 음성 신호에 대한 응답 정보를 제공하기 위해 이용될 수 있다. 예를 들면, 상기 오디오 신호는, 전자 장치(301)로부터 수신한 오디오 신호를 포함할 수 있다. 예를 들면, 상기 오디오 신호 내의 상기 음성 신호는, 전자 장치(301)의 마이크로폰(330)로 획득된 음성 신호를 포함할 수 있다. 대형 언어 모델(390)은 대규모 언어 모델, 및 LLM(large language model)로 참조될 수 있다. 예를 들면, 대형 언어 모델(390)은 자연어 처리(natural language processing)를 수행할 수 있다. 예를 들면, 대형 언어 모델(390)은 자연어 이해(natural language understanding)를 수행할 수 있다.
예를 들면, 외부 전자 장치(302)는, 제2 멀티모달 모델(360)을 이용하여 오디오 신호 및/또는 센싱 데이터에 기반하는 프롬프트를 생성할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 생성된 프롬프트를 대형 언어 모델(390)에 입력하거나 제공할 수 있다. 예를 들면, 대형언어 모델(390)은 상기 생성된 프롬프트의 입력에 기반하여, 응답 정보를 생성할 수 있다. 예를 들면, 외부 전자 장치(302)는 대형 언어 모델(390)을 이용하여 상기 오디오 신호에 대한 응답 정보를 획득할 수 있다.
도 4a는 제1 멀티모달 모델의 학습되는 동작의 예를 도시한다.
이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
일 실시 예에 따르면, 동작 411 내지 415는 전자 장치(예: 도 3a의 전자 장치(301))의 프로세서(예: 도 3a의 적어도 하나의 프로세서(300))에 의해 수행되는 것으로 이해될 수 있다.
도 4a를 참고하면, 동작 411에서, 예를 들면, 오디오 신호 및 센싱 데이터는 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 입력될 수 있다. 예를 들면, 상기 오디오 신호는 학습 오디오 신호를 나타낼 수 있다. 예를 들면, 상기 센싱 데이터는 학습 센싱 데이터를 나타낼 수 있다.
동작 412에서, 전자 장치(301)는 특징 추출기(또는 인코더)를 포함할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은 상기 특징 추출기를 포함할 수 있다. 예를 들면, 상기 특징 추출기는, 상기 오디오 신호 및 상기 센싱 데이터로부터, 상기 오디오 신호의 데이터의 특징을 추출할 수 있다. 예를 들면, 상기 특징 추출기는, 상기 센싱 데이터로부터, 상기 센싱 데이터의 특징을 추출할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은 상기 오디오 신호 및 상기 센싱 데이터로부터, 상기 특징 추출기를 이용하여 임베딩 벡터를 추출할 수 있다.
동작 413에서, 제1 멀티모달 모델(350)의 학습은, 지도 학습(supervised learning) 및/또는 비지도 학습(unsupervised learning)에 기반하여 수행될 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 상기 추출된 임베딩 벡터와 실제 값(ground truth) 사이의 유사성을 비교할 수 있다. 예를 들면 상기 비교의 결과에 따라, 제1 멀티모달 모델(350)은 손실 함수(loss function)를 획득할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 상기 손실 함수를 통해 학습될 수 있다. 예를 들면, 제1 멀티모달 모델(350)은 학습되는 동안, 레이어들(예: 입력 레이어, 하나 이상의 히든 레이어들 및 출력 레이어) 각각에 포함된 노드들 사이의 연결 가중치를 변경할 수 있다.
동작 414에서, 제1 멀티모달 모델(350)은, 드롭아웃(dropout)에 대한 모달리티 각각의 확률을 조정할 수 있다. 예를 들면, 상기 드롭아웃은, 제1 멀티모달 모델(350)의 신경망의 뉴런을 부분적으로 학습에서 제외하는 것을 나타낼 수 있다. 예를 들면, 제1 멀티모달 모델(350)은, 상기 드롭아웃을 통해, 과적합(overfitting)을 삼가할 수 있다.
예를 들면, 제1 멀티모달 모델(350)에게 상기 오디오 신호에 대한 임베딩 벡터를 입력함에 따라, 제1 멀티모달 모델(350)은 상기 오디오 신호의 드롭아웃에 대한 제1 확률을 조정할 수 있다. 예를 들면, 제1 멀티모달 모델(350)에게 상기 센싱 데이터에 대한 임베딩 벡터를 입력함에 따라, 제1 멀티모달 모델(350)은 센싱 데이터의 드롭아웃에 대한 제2 확률을 조정할 수 있다. 예를 들면, 제1 멀티모달 모델(350)에게 상기 오디오 신호에 대한 임베딩 벡터 및 상기 센싱 데이터에 대한 임베딩 벡터를 입력함에 따라, 제1 멀티모달 모델(350)은 상기 오디오 신호 및 상기 센싱 데이터의 드롭아웃에 대한 제3 확률을 조정할 수 있다.
동작 415에서, 제1 멀티모달 모델(350)은 드롭 아웃을 수행하고, 튜닝(fine)을 수행할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은 상기 제1 확률, 상기 제2 확률, 및 상기 제3 확률을 이용하여 드롭아웃을 수행하고, 상기 튜닝을 수행할 수 있다. 예를 들면, 제1 멀티모달 모델(350)은 튜닝되는 동안, 레이어들 각각에 포함된 노드들 사이의 연결 가중치를 변경할 수 있다.
예를 들면, 도 4a에서는 도시하지 않았으나, 제1 멀티모달 모델(350)의 학습되는 동작은 도 3b의 제2 멀티모달 모델(360)의 학습되는 동작에 대응할 수 있다. 예를 들면, 도 3b의 제2 멀티모달 모델(360)은 도 4a에서 예시되는 동작으로 학습될 수 있다.
도 4b는 제4 멀티모달 모델의 학습되는 동작의 예를 도시한다.
이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
일 실시 예에 따르면, 동작 421 내지 425는 전자 장치(예: 도 3a의 전자 장치(301))의 프로세서(예: 도 3a의 적어도 하나의 프로세서(300))에 의해 수행되는 것으로 이해될 수 있다.
도 4b를 참고하면, 동작 421에서, 전자 장치(301)는 특징 추출기(미도시)를 포함할 수 있다. 예를 들면, 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))은 상기 특징 추출기를 포함할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 학습 데이터로부터, 상기 특징 추출기를 이용하여 학습 데이터의 특징을 추출할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 학습 데이터로부터, 상기 특징 추출기를 이용하여 임베딩 벡터를 추출할 수 있다.
예를 들면, 상기 학습 데이터는 메모리(예: 메모리(320))에 저장될 수 있다. 예를 들면, 상기 학습 데이터는 오디오 신호를 포함할 수 있다. 예를 들면, 상기 특징 추출기는, STFT(short time fourier transform)을 적용하여 학습 데이터의 특징을 추출할 수 있다. 예를 들면, 상기 특징 추출기를 이용하여, 음성 특징이 상기 오디오 신호로부터 추출될 수 있다.
예를 들면, 동작 422에서, 제4 멀티모달 모델(380)은 학습 오디오 신호로부터 학습 오디오 신호가 획득된 환경을 분류(또는 식별)할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 상기 임베딩 벡터로부터 상기 학습 데이터가 획득된 환경을 분류(또는 식별)할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은, 제3 모델(예: 제3 모델(370))을 이용하여, 상기 환경을 분류(또는 식별)할 수 있다. 예를 들면, 제3 모델(370)은, 상기 환경을 분류하고, 상기 환경에 대한 데이터를 제4 멀티모달 모델(380)에게 제공할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 상기 분류된 환경에 대한 데이터를 토큰으로 이용할 수 있다.
동작 423에서, 상기 오디오 신호로부터 추출된 상기 음성 특징, 상기 분류된 환경에 대한 상기 데이터, 및 학습 센싱 데이터의 특징은 제4 멀티모달 모델(380)에게 입력될 수 있다. 예를 들면, 상기 음성 특징, 및 상기 분류된 환경에 대한 상기 데이터는 데이터를 연결하는 작업(concatenate)이 수행되고, 제4 멀티모달 모델(380)에게 입력될 수 있다.
동작 424에서, 제4 멀티모달 모델(380)은 인코더(encoder)를 포함할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 상기 인코더를 이용하여 상기 음성 특징, 상기 환경에 대한 상기 데이터, 및 상기 학습 센싱 데이터의 특징의 퓨전(fusion)을 수행할 수 있다. 예를 들면, 상기 퓨전은 다른 유형의 모달리티를 하나의 데이터로 합치는 것을 나타낼 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 상기 퓨전을 수행함에 따라, 하나의 멀티모달 모델의 특징이 획득될 수 있다.
동작 425에서, 제4 멀티모달 모델(380)은, 상기 하나의 멀티모달 모델의 특징을 이용하여, 상기 오디오 신호에 대한 노이즈 캔슬링을 수행할 수 있다. 예를 들면, 상기 노이즈 캔슬링은, 활성화 함수(예: sigmoid function)를 이용하여 수행될 수 있다. 예를 들면, 제4 멀티모달 모델(380)은, 상기 노이즈 캔슬링이 수행된 오디오 신호 및 실제 값(ground truth)에 기반하여 손실 함수(loss function)를 획득할 수 있다. 예를 들면, 제4 멀티모달 모델(380)은, 상기 손실 함수를 통해 학습될 수 있다. 예를 들면, 제4 멀티모달 모델(380)은 학습되는 동안, 레이어들(예: 입력 레이어, 하나 이상의 히든 레이어들 및 출력 레이어) 각각에 포함된 노드들 사이의 연결 가중치를 변경할 수 있다.
도 5는 전자 장치가 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를 송신하는 동작들의 예를 도시한다. 도 5에서 예시되는 전자 장치(예: 전자 장치(301))의 동작들은, 적어도 하나의 프로세서(예: 적어도 하나의 프로세서(300))에 의해 실행되거나 수행되거나 제어될 수 있다.
이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
도 5를 참고하면, 동작 511에서, 적어도 하나의 프로세서(300)는 적어도 하나의 센서를 활성화하는 이벤트를 검출할 수 있다. 예를 들면, 상기 이벤트는, 적어도 하나의 센서(340)를 구동하기 위한 입력으로 참조될 수 있다. 예를 들면, 적어도 하나의 프로세서(300)가 상기 이벤트를 검출함으로써 적어도 하나의 센서(340)를 활성화할 수 있다.
예를 들면, 적어도 하나의 센서(340)는, 적어도 하나의 프로세서(300)에 의해 동작 511이 실행되기 전 비활성화 상태에 놓여 있을 수 있다. 예를 들면, 이미지 센서는 적어도 하나의 프로세서(300)가 동작 511을 실행하기 전에 구동되고 있지 않을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은, 적어도 하나의 프로세서(300)가 이미지 센서를 구동하는 것을 포함할 수 있다.
제한되지 않는 예로, 적어도 하나의 센서(340)는, 적어도 하나의 프로세서(300)에 의해 동작 511이 실행되기 전에 구동되고 있을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터를 오디오 신호 처리와 관련하여 이용하는 것을 포함할 수 있다. 예를 들면, 심박 센서, 모션 센서, 및/또는 가속도 센서는 동작 511이 실행되기 전 구동되고 있을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은, 적어도 하나의 프로세서(300)가 심박 센서, 모션 센서, 및/또는 가속도 센서를 통해 획득된 센싱 데이터를 오디오 신호 처리와 관련하여 이용하는 것을 포함할 수 있다. 상기 이벤트는, 도 6a 및 6b를 참조하여 보다 상세히 설명되고 예시된다.
도 6a 및 6b는 적어도 하나의 센서를 구동하기 위한 이벤트의 예를 도시한다.
도 6a를 참고하면, 전자 장치(301)는, 입력 요소(610)(input element)를 포함할 수 있다. 예를 들면, 입력 요소(610)는, 전자 장치(301)의 하우징의 일부를 통해 노출될 수 있다. 예를 들면, 상기 이벤트는, 적어도 하나의 프로세서(300)가 입력 요소(610)를 통해 입력 신호를 수신하는 것을 포함할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 입력 요소(610)를 통해 수신된 입력 신호를 검출할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 검출에 기반하여, 적어도 하나의 센서(340)를 활성화할 수 있다.
예를 들면, 입력 요소(610)는 누름 가능할(pressable) 수 있다. 예를 들면, 입력 요소(610)는 물리적 버튼일 수 있다. 예를 들면, 입력 요소(610)는 누름 가능한 입력 버튼을 나타낼 수 있다. 예를 들면, 상기 이벤트는, 적어도 하나의 프로세서(300)가, 입력 요소(610)를 통해 누름 입력(push input)을 수신하는 것을 포함할 수 있다.
예를 들면, 입력 요소(610)는 회전 가능할 수 있다. 예를 들면, 입력 요소(610)는 전자 장치(301)의 하우징에 대하여 회전 가능할 수 있다. 예를 들면, 상기 이벤트는, 적어도 하나의 프로세서(300)가, 입력 요소(610)를 통해 입력 구성요소를 회전시키는 입력을 수신하는 것을 포함할 수 있다.
예를 들면, 입력 요소(610)는 터치 센서를 포함할 수 있다. 예를 들면, 상기 이벤트는, 적어도 하나의 프로세서(300)가, 입력 요소(610)를 통해 터치 입력을 수신하는 것을 포함할 수 있다. 하지만 이에 제한되지 않는다.
예를 들면, 전자 장치(301)는 워치 형상으로 구성된 전자 장치(301)에 제한되는 것은 아니다. 제한하지 않는 예로, 전자 장치(301)는, 스마트폰, 테블릿, 랩탑 컴퓨터, 스마트 워치와 같은 포터블 전자 장치를 포함할 수 있다.
도 6b를 참고하면, 적어도 하나의 센서(340)를 활성화하기 위한 이벤트는, 기준 움직임을 검출하는 것을 포함할 수 있다. 예를 들면, 상기 이벤트의 검출에 기반하여, 비활성화 상태 내에서 있는 적어도 하나의 센서(340)를 구동할 수 있다. 예를 들면, 전자 장치(301)는 모션 센서 및/또는 가속도 센서를 포함할 수 있다. 예를 들면, 상기 모션 센서 및/또는 상기 가속도 센서는 활성화 상태 내에서 있을 수 있다. 예를 들면, 상기 모션 센서 및/또는 상기 가속도 센서는 저 전력 소비를 위한 상태 내에서 있을 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 모션 센서 및/또는 상기 가속도 센서를 이용하여 상기 기준 움직임을 검출할 수 있다.
예를 들면, 상기 기준 움직임은, 전자 장치(301)가 상태(620)에서 상태(630)으로 변하는 움직임을 포함할 수 있다. 예를 들면, 상태(620)는 전자 장치(301)의 프론트 사이드의 방향이 사용자를 향하고 있지 않는 상태를 나타낼 수 있다. 예를 들면, 상태(630)는 전자 장치(301)의 프론트 사이드의 방향이 사용자를 향하고 있는 상태를 나타낼 수 있다.
예를 들면, 프론트 사이드는 전자 장치(301)의 전면으로 참조될 수 있다. 예를 들면, 프론트 사이드는, 전자 장치(301)의 디스플레이(311)를 포함하는 영역으로 참조될 수 있다.
예를 들면, 상기 기준 움직임은, 전자 장치(301)가 상태(620)에서 상태(630)로 변하는 동안의 속도와 관련될 수 있다. 예를 들면, 상기 기준 움직임은, 상태(620)에서 전자 장치(301)의 프론트 사이드의 방향과, 상태(630)에서 전자 장치(301)의 상기 프론트 사이드의 방향과의 차이와 관련될 수 있다.
다시 도 5를 참고하면, 동작 512에서, 적어도 하나의 프로세서(300)는 상기 이벤트를 검출하는 것에 기반하여, 마이크로폰(330)을 통해 오디오 신호를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 이벤트를 검출하는 것에 기반하여, 상기 이벤트에 따라 활성화된 적어도 하나의 센서(340)를 통해 센싱 데이터(예: 이미지(예: 이미지 센서를 통해 획득됨), 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 이벤트를 검출하는 것에 기반하여, 상기 이벤트에 따라 구동된 이미지 센서를 통해 이미지를 획득할 수 있다.
동작 513에서, 적어도 하나의 프로세서(300)는, 상기 센싱 데이터가 기준 센싱 데이터를 나타내는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 오디오 신호가 기준 음성 명령을 포함하는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 상기 사용자의 상기 기준 센싱 데이터를 표현하는 상기 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 오디오 신호를 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별할 수 있다. 예를 들면, 상기 기준 센싱 데이터는 상기 사용자에 의해 미리 설정될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 이미지가 사용자의 기준 제스쳐를 표현하는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 상기 사용자의 상기 기준 제스쳐를 표현하는 상기 이미지 및 상기 기준 음성 명령을 포함하는 상기 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다. 예를 들면, 상기 기준 제스쳐는 상기 사용자가 미리 설정할 수 있다. 예를 들면, 상기 기준 제스쳐는 도 7을 참조하여 보다 상세히 설명되고 예시된다.
도 7은 제1 멀티모달 모델이 인식하는 기준 제스쳐 및 기준 음성 명령의 예를 도시한다.
도 7을 참고하면, 상태(710)에서, 사용자는 상기 기준 제스쳐를 나타낼 수 있다. 예를 들면, 상기 기준 제스쳐는 사용자가 사용자의 손가락을 사용자의 입에 대는 제스쳐를 나타낼 수 있다.
예를 들면, 상태(710)에서, 사용자는 상기 기준 제스쳐를 행동하는 중에 기준 음성 명령(720)을 발화할 수 있다. 예를 들면, 상기 기준 음성 명령(720)은 1음절의 음성을 포함할 수 있다. 예를 들면 상기 기준 음성 명령(720)은, 2음절의 음성을 포함할 수 있다. 예를 들면, 상기 기준 음성 명령(720)은 '쉿(shh)'을 나타낼 수 있다. 하지만 이에 제한되지 않는다. 도 7에서 도시된 상기 기준 제스쳐 및 기준 음성 명령(720)은 단지 예시일 뿐이다.
다시 도 5를 참고하면, 동작 514에서, 적어도 하나의 프로세서(300)는 상기 기준 센싱 데이터를 나타내는 상기 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 외부 전자 장치(302)에게 신호를 통신 회로(310)를 통해 송신할 수 있다. 예를 들면, 상기 신호는 외부 전자 장치(302) 내의 제2 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를 포함할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는 상기 사용자의 상기 기준 제스쳐를 표현하는 상기 이미지 및 상기 기준 음성 명령을 포함하는 상기 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 외부 전자 장치(302)에게 신호를 통신 회로(310)를 통해 송신할 수 있다. 예를 들면, 상기 신호는 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크-업을 야기하는 신호를 포함할 수 있다.
예를 들면, 동작 515에서, 외부 전자 장치(302)는, 통신 회로를 통해 상기 신호를 수신할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)을 웨이크-업하도록 야기할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)의 활성화를 야기할 수 있다.
예를 들면, 전자 장치(301)는 적어도 하나의 프로세서(300)가 도 5의 예시된 동작들을 수행함으로써, 노이즈의 볼륨이 상대적으로 큰 환경에서, 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크 업을 야기할 수 있다. 예를 들면, 전자 장치(301)는 노이즈의 볼륨이 상대적으로 큰 환경에서, 제2 멀티모달 모델(360)을 통해 음성 인식 기능을 이용할 수 있다.
예를 들면, 전자 장치(301)는 적어도 하나의 프로세서(300)가 도 5의 예시된 동작들을 수행함으로써, 사용자가 발화를 속삭여야 하는 환경에서, 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크 업을 야기할 수 있다. 예를 들면, 전자 장치(301)는 사용자가 발화를 속삭여야 하는 환경에서, 제2 멀티모달 모델(360)을 통해 음성 인식 기능을 이용할 수 있다. 예를 들면, 전자 장치(301)의 음성 인식 기능의 품질이 강화될 수 있다.
도 8a 및 8b는 전자 장치가 제2 멀티모달 모델의 웨이크-업을 야기하는 신호를 송신하는 동작들의 예를 도시한다. 도 8a 및 8b에서 예시되는 전자 장치(예: 전자 장치(301))의 동작들은, 적어도 하나의 프로세서(예: 적어도 하나의 프로세서(300))에 의해 실행되거나 수행되거나 제어될 수 있다. 이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
도 8a를 참고하면, 동작 811에서, 적어도 하나의 프로세서(300)는, 제1 오디오 신호를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제1 오디오 신호를 획득할 수 있다.
예를 들면, 동작 812에서, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 동작 812에서, 적어도 하나의 프로세서(300)는, 상기 제1 오디오 신호로부터 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제3 모델(예: 제3 모델(370))을 이용하여 상기 환경의 유형을 식별할 수 있다. 제한하지 않는 예로, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(예: 적어도 하나의 센서(340))를 통해 획득된 센싱 데이터를 통해, 전자 장치(301)가 위치된 상기 환경의 상기 유형을 식별할 수 있다.
예를 들면, 상기 환경의 상기 유형은 상기 제1 유형 및 상기 제2 유형을 포함할 수 있다. 예를 들면, 상기 제1 유형 및 상기 제2 유형을 위해, 도 3b의 제3 모델(370)에 대한 설명들이 참조될 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 이상인 경우, 제1 유형의 환경으로 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 미만인 경우, 제2 유형의 환경으로 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 동작 813을 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 동작 815를 실행할 수 있다.
예를 들면, 동작 813에서, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 상기 기준 음성 명령이 포함됨을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제1 멀티모달 모델(350)을 이용하여 상기 기준 음성 명령이 상기 제1 오디오 신호 내에 포함됨을 식별할 수 있다. 예를 들면, 동작 813에서, 적어도 하나의 프로세서(300)는 상기 제1 오디오 신호 내에 상기 기준 음성 명령이 포함되었는지 여부를 식별할 수 있다.
예를 들면, 동작 814에서, 적어도 하나의 프로세서(300)는, 상기 제1 오디오 신호 내에 포함된 상기 기준 음성 명령에 따라, 제2 멀티모달 모델(360)의 웨이크-업을 야기하는 신호를, 통신 회로(310)를 통해 외부 전자 장치(302)에게 송신할 수 있다.
예를 들면, 동작 815에서, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 활성화할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 구동할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 적어도 하나의 프로세서(300)에 의해 동작 815가 실행되기 전 비활성화 상태에 놓여 있을 수 있다.
예를 들면, 이미지 센서는 적어도 하나의 프로세서(300)가 동작 815를 실행하기 전 구동되고 있지 않을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은, 적어도 하나의 프로세서(300)가 이미지 센서를 구동하는 것을 포함할 수 있다.
제한되지 않는 예로, 적어도 하나의 센서(340)는, 적어도 하나의 프로세서(300)에 의해 동작 815가 실행되기 전 구동되고 있을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터를 오디오 신호 처리와 관련하여 이용하는 것을 포함할 수 있다. 예를 들면, 심박 센서, 모션 센서, 및/또는 가속도 센서는 동작 815가 실행되기 전 구동되고 있을 수 있다. 예를 들면, 적어도 하나의 센서(340)를 활성화하는 것은, 적어도 하나의 프로세서(300)가 심박 센서, 모션 센서, 및/또는 가속도 센서를 통해 획득된 센싱 데이터를 오디오 신호 처리와 관련하여 이용하는 것을 포함할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 활성화하기 전, 적어도 하나의 센서(340)의 활성화를 알리는 컨텐츠를 포함하는 화면을 디스플레이(311)를 통해 표시할 수 있다.
예를 들면, 동작 816에서, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 통해 상기 제1 센싱 데이터(예: 제1 이미지(예: 이미지 센서를 통해 획득됨), 제1 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 제1 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 이미지 센서를 통해 제1 이미지를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제2 오디오 신호를 획득할 수 있다.
예를 들면, 동작 817에서, 적어도 하나의 프로세서(300)는, 기준 센싱 데이터 및 기준 음성 명령을 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제2 오디오 신호 내에 상기 기준 음성 명령이 포함됨을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 동작 817에서, 상기 제2 오디오 신호 내에 상기 기준 음성 명령이 포함되었는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 센싱 데이터 내에 상기 기준 센싱 데이터가 지시됨을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 센싱 데이터 내에 상기 기준 센싱 데이터가 지시되는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 이미지 내에 상기 기준 제스쳐가 포함되었는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다. 예를 들면, 동작 817은, 도 5의 동작 513에 대응할 수 있다.
예를 들면, 동작 818에서, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 통신 회로(310)를 통해 외부 전자 장치(302)에게 신호를 송신할 수 있다. 예를 들면, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 통신 회로(310)를 통해 외부 전자 장치(302)에게 신호를 송신할 수 있다. 예를 들면, 상기 신호는 제2 멀티모달 모델(360)의 웨이크-업을 야기하는 신호를 나타낼 수 있다. 예를 들면, 동작 818은, 도 5의 동작 514에 대응할 수 있다.
예를 들면, 동작 819에서, 외부 전자 장치(302)는 상기 신호를 통신 회로를 통해 수신할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)을 웨이크-업하도록 야기할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)의 활성화를 야기할 수 있다. 예를 들면, 동작 819은, 도 5의 동작 515에 대응할 수 있다.
도 8b를 참고하면, 동작 821에서, 적어도 하나의 프로세서(300)는, 타이머를 활성화할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 타이머를 실행할 수 있다. 예를 들면, 동작 821은, 도 8a의 동작 812에서 적어도 하나의 프로세서(300)가 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여 실행될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 타이머를 활성화할 수 있다. 예를 들면, 상기 타이머는 적어도 하나의 센서(340)와 관련될 수 있다. 예를 들면, 상기 타이머는 상기 이미지 센서와 관련될 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 타이머가 활성화되는 동안, 적어도 하나의 센서(340)를 활성화할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 타이머가 활성화되는 동안, 이미지 센서를 구동할 수 있다.
동작 822에서, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 활성화할 수 있다. 예를 들면, 동작 822는, 동작 815에 대응할 수 있다.
동작 823에서, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 통해 상기 제1 센싱 데이터를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 이미지 센서를 통해 제1 이미지를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제2 오디오 신호를 획득할 수 있다. 예를 들면, 동작 823은, 동작 816에 대응할 수 있다.
동작 824에서, 적어도 하나의 프로세서(300)는, 상기 제2 오디오 신호 내에 상기 기준 음성 명령이 포함되었는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 센싱 데이터 내에 상기 기준 제스쳐가 포함되었는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 이미지 내에 상기 기준 제스쳐가 포함되었는지 여부를 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 식별하는 것에 기반하여, 동작 825 및 동작 827을 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내지 않는 상기 제1 센싱 데이터 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 식별하는 것에 기반하여, 동작 828을 실행할 수 있다.
동작 825에서, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 표현하는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 통신 회로(310)를 통해 외부 전자 장치(302)에게 신호를 송신할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 통신 회로(310)를 통해 외부 전자 장치(302)에게 상기 신호를 송신할 수 있다.
예를 들면, 상기 신호는 제2 멀티모달 모델(360)의 웨이크-업을 야기하는 신호로 나타내어질 수 있다. 예를 들면, 동작 825는, 동작 818에 대응할 수 있다.
동작 826에서, 외부 전자 장치(302)는 상기 신호를 통신 회로를 통해 수신할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)을 웨이크-업하도록 야기할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 신호를 수신하는 것에 응답하여, 제2 멀티모달 모델(360)의 활성화를 야기할 수 있다. 예를 들면, 동작 826은, 동작 819에 대응할 수 있다.
동작 827에서, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 적어도 하나의 센서(340)의 활성화를 유지할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)의 활성화를 유지하도록 적어도 하나의 센서(340)를 제어할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 타이머의 상기 만료와 독립적으로 적어도 하나의 센서(340)를 구동하는 것을 유지할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 표현하는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 상기 타이머의 상기 만료와 독립적으로 적어도 하나의 센서(340)의 활성화를 유지할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 타이머의 상기 만료와 독립적으로 상기 이미지 센서의 구동을 유지할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 표현하는 상기 제1 센싱 데이터 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 상기 타이머의 남은 시간을 연장할 수 있다. 예를 들면, 상기 타이머의 남은 시간이 연장됨에 따라, 적어도 하나의 센서(340)의 활성화 상태는 유지될 수 있다. 예를 들면, 상기 타이머의 남은 시간이 연장됨에 따라, 상기 이미지 센서의 구동은 유지될 수 있다.
동작 828에서, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 나타내지 않는 상기 제1 센싱 데이터 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 비활성화할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 비활성화하도록 적어도 하나의 센서(340)를 제어할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)의 활성화 상태를 중단할 수 있다. 예를 들면, 상기 중단은, 상기 타이머의 상기 만료에 응답하여 실행될 수 있다. 예를 들면, 상기 중단은, 상기 이미지 센서의 구동을 중단하는 것을 포함할 수 있다. 예를 들면, 상기 중단은 심박 센서, 가속도 센서, 및/또는 자이로 센서를 통해서 획득된 데이터를 오디오 신호 처리에 사용하는 것을 중단하는 것을 포함할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 센싱 데이터를 표현하지 않는 상기 제1 기준 센싱 데이터 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 식별하는 것에 기반하여, 적어도 하나의 센서(340)의 상태를 활성화 상태에서 비활성화 상태로 변경할 수 있다. 예를 들면, 상기 변경은, 상기 타이머의 상기 만료에 응답하여 실행될 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 기준 제스쳐를 표현하지 않는 상기 제1 이미지 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 식별하는 것에 기반하여, 상기 이미지 센서의 구동을 중단할 수 있다. 예를 들면, 상기 중단은, 상기 타이머의 상기 만료에 응답하여 실행될 수 있다.
예를 들면, 전자 장치(301)는 적어도 하나의 프로세서(300)가 도 8a 내지 8b의 예시된 동작들을 수행함으로써, 전자 장치(301)가 위치된 환경에 따라 효율적으로 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크-업을 야기할 수 있다.
예를 들면, 전자 장치(301)는, 노이즈의 볼륨이 상대적으로 큰 환경에서, 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크 업을 야기할 수 있다. 예를 들면, 전자 장치(301)는 노이즈의 볼륨이 상대적으로 큰 환경에서, 음성 인식 기능을 위하여 제2 멀티모달 모델(360)을 이용할 수 있다.
예를 들면, 전자 장치(301)는 적어도 하나의 프로세서(300)가 도 8a 내지 8b의 예시된 동작들을 수행함으로써, 사용자가 발화를 속삭여야 하는 환경에서, 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크 업을 야기할 수 있다. 예를 들면, 전자 장치(301)는 사용자가 발화를 속삭여야 하는 환경에서, 음성 인식 기능을 위하여 제2 멀티모달 모델(360)을 이용할 수 있다. 예를 들면, 전자 장치(301)의 음성 인식 기능의 품질은 강화될 수 있다.
도면 9a 및 9b는 전자 장치가 제2 멀티모달 모델을 이용하기 위해 외부 전자 장치에게 데이터를 송신하고, 외부 전자 장치로부터 응답 정보를 수신하는 동작들의 예를 도시한다. 도 9a 및 9b에서 예시되는 전자 장치(예: 전자 장치(301))의 동작들은, 적어도 하나의 프로세서(예: 적어도 하나의 프로세서(300))에 의해 실행되거나 수행되거나 제어될 수 있다. 이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
도 9a를 참고하면, 동작 911은, 도 5의 동작 514의 이후 동작을 나타낼 수 있다. 예를 들면, 동작 911은, 도 8a의 동작 818의 이후 동작을 나타낼 수 있다. 예를 들면, 동작 911은, 도 8b의 동작 827의 이후 동작을 나타낼 수 있다. 예를 들면, 동작 911은, 적어도 하나의 프로세서(300)가, 통신 회로(310)를 통해, 제2 멀티모달 모델(360)의 웨이크 업을 야기하는 상기 신호를 송신한 후의 동작으로 설명될 수 있다.
예를 들면, 동작 911에서, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제3 오디오 신호를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 적어도 하나의 센서(340)를 통해 제2 센싱 데이터(예: 제2 이미지(예: 이미지 센서를 통해 획득됨), 제2 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 제2 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 이미지 센서를 통해 제2 이미지를 획득할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제3 오디오 신호를 이용하여 제1 데이터를 획득할 수 있다. 예를 들면, 상기 제1 데이터는 상기 제3 오디오 신호가 강화된 신호에 대한 데이터를 나타낼 수 있다. 예를 들면, 상기 제1 데이터는, 상기 제3 오디오 신호 내의 사용자의 음성 신호가 증폭된 신호에 대한 데이터를 나타낼 수 있다. 예를 들면, 상기 제1 데이터는, 상기 제3 오디오 신호 내의 노이즈가 제거된 신호에 대한 데이터를 나타낼 수 있다. 예를 들면, 상기 제1 데이터는 제2 멀티모달 모델(360)에서 이용되기 위하여, 상기 제3 오디오 신호에 대한 특징 추출이 수행된 데이터로 나타내어질 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제2 센싱 데이터를 이용하여 제2 데이터를 획득할 수 있다. 예를 들면, 상기 제2 데이터는, 상기 제2 이미지 내의 시각적 객체에 대한 데이터를 나타낼 수 있다. 예를 들면, 상기 시각적 객체는, 전자 장치(301)의 사용자의 입술에 대응할 수 있다. 예를 들면, 상기 제2 데이터는, 상기 입술에 대응하는 시각적 객체의 움직임 및/또는 모양과 관계되는 데이터를 나타낼 수 있다. 하지만 이제 제한되지 않는다.
예를 들면, 전자 장치(301)는 얼굴 검출 모델(미도시)을 포함할 수 있다. 예를 들면, 상기 얼굴 검출 모델은 이미지 센서를 통해 획득된 이미지로부터 사용자의 얼굴을 식별할 수 있다. 예를 들면, 상기 얼굴 검출 모델은 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 영역에 대한 데이터를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 영역에 대한 상기 데이터를 이용하여 상기 입술에 대응하는 시각적 객체에 대한 상기 제2 데이터를 획득할 수 있다.
동작 912에서, 적어도 하나의 프로세서(300)는, 상기 제1 데이터 및 상기 제2 데이터를 통신 회로(310)를 통해 외부 전자 장치(302)에게 송신할 수 있다.
동작 913에서, 외부 전자 장치(302)는 제2 멀티모달 모델(360)을 동작(또는 이용)할 수 있다. 예를 들면, 외부 전자 장치(302)는, 통신 회로를 통해, 상기 제1 데이터 및 상기 제2 데이터를 수신할 수 있다. 예를 들면, 외부 전자 장치(302)는 상기 제1 데이터 및 상기 제2 데이터를 제2 멀티모달 모델(360)에게 제공(또는 입력)할 수 있다. 예를 들면, 제2 멀티모달 모델(360)은 웨이크 업 상태에 있을 수 있다.
예를 들면, 외부 전자 장치(302)는, 제2 멀티모달 모델(360)을 이용하여 대형 언어 모델(390)에 입력하기 위한 데이터를 획득할 수 있다. 예를 들면, 대형 언어 모델(390)에 입력하기 위한 상기 데이터는, 프롬프트로 나타내어질 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 제1 데이터 및 상기 제2 데이터를 제2 멀티모달 모델(360)에게 제공함으로써, 대형 언어 모델(390)에 입력하기 위한 프롬프트를 생성할 수 있다.
동작 914에서, 외부 전자 장치(302)는 대형 언어 모델(390)을 동작 할 수 있다. 예를 들면, 외부 전자 장치(302)는, 상기 프롬프트를 대형 언어 모델에 입력함으로써, 응답 정보를 획득할 수 있다. 예를 들면, 상기 응답 정보는, 상기 제1 데이터 및/또는 상기 제2 데이터에 기반할 수 있다. 예를 들면, 상기 응답 정보는 상기 제3 오디오 신호에 대한 응답을 나타낼 수 있다. 예를 들면, 상기 응답 정보는, 상기 제2 이미지 내의 시각적 객체에 대한 응답을 포함할 수 있다. 상기 응답 정보는, 도 10a 및 10b를 참조하여 보다 상세히 설명되고 예시될 것이다.
동작 915에서, 외부 전자 장치(302)는 상기 응답 정보를 통신 회로를 통해 전자 장치(301)에게 송신할 수 있다.
동작 916에서, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 따른 기능을 수행하거나 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 TTS(text to speech)를 적용할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 TTS를 적용함으로써 오디오 신호를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 TTS를 적용함으로써 획득된 상기 오디오 신호를 스피커(312)를 통해 외부로 출력할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 따른 상기 기능을 수행함으로써, 상기 응답 정보 내의 텍스트를 디스플레이(311)를 통해 표시할 수 있다. 예를 들면, 상기 응답 정보 내의 상기 텍스트는, 상기 제1 데이터 및/또는 상기 제2 데이터에 기반할 수 있다. 예를 들면, 상기 응답 정보 내의 상기 텍스트는 상기 제1 데이터 및/또는 상기 제2 데이터에 기반하여 제2 멀티모달 모델(360)에 의해 생성될 수 있다.
상기 응답 정보에 따른 상기 기능은, 도 10a 내지 10b를 참조하여 보다 상세히 설명되고 예시될 것이다.
도 9b를 참고하면, 동작 921에서, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 상기 제3 오디오 신호를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 적어도 하나의 센서(340)를 통해 상기 제2 센싱 데이터(예: 제2 이미지(예: 이미지 센서를 통해 획득됨), 제2 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 제2 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 이미지 센서를 통해 상기 제2 이미지를 획득할 수 있다. 동작 921은, 동작 911에 대응할 수 있다. 예를 들면, 동작 921은, 도 8a의 동작 814의 이후 동작을 나타낼 수 있다.
동작 922에서, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 동작 922에서, 적어도 하나의 프로세서(300)는, 상기 제3 오디오 신호로부터 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제3 모델(370)을 이용하여 상기 환경의 유형을 식별할 수 있다. 제한하지 않는 예로, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 통해 획득된 상기 제2 센싱 데이터를 이용하여 상기 환경의 유형을 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 이상인 경우, 상기 제1 유형의 환경으로 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 미만인 경우, 상기 제2 유형의 환경으로 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 동작 923을 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 동작 924 및/또는 동작 925를 실행할 수 있다.
동작 923에서, 적어도 하나의 프로세서(300)는 상기 제3 오디오 신호에 대한 데이터 및 상기 제2 센싱 데이터에 대한 데이터를 송신할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제2 센싱 데이터를 이용하여 상기 제2 데이터를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는 상기 제3 오디오 신호에 대한 데이터 및 상기 제2 이미지 내의 상기 시각적 객체에 대한 상기 제2 데이터를 송신할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제1 데이터 및 상기 제2 데이터를 통신 회로(310)를 통해 외부 전자 장치(302)에게 송신할 수 있다. 예를 들면, 동작 923은 동작912에 대응할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)가 동작 923을 실행한 후에, 외부 전자 장치(302)는 동작 913 내지 동작915를 실행할 수 있다. 예를 들면, 외부 전자 장치(302)가 동작 913 내지 동작 915를 실행한 후에, 적어도 하나의 프로세서(300)는 동작 916을 실행할 수 있다.
동작 924에서, 적어도 하나의 프로세서(300)는, 상기 제3 오디오 신호에 대한 노이즈 캔슬링을 수행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 상기 노이즈 캔슬링을 수행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제4 멀티모달 모델(380)을 이용하여 상기 노이즈 캔슬링을 수행할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제2 데이터, 상기 제2 유형에 대한 데이터, 및/또는 상기 제3 오디오 신호에 대한 데이터를 제4 멀티모달 모델(380)에게 제공(또는 입력)함으로써, 상기 노이즈 캔슬링을 수행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제2 데이터, 상기 제2 유형에 대한 데이터, 및/또는 상기 제3 오디오 신호에 대한 데이터를 제4 멀티모달 모델(380)에게 제공함으로써, 상기 제3 오디오 신호 중 노이즈가 제거된 제4 오디오 신호를 획득할 수 있다. 예를 들면, 상기 제4 오디오 신호는, 제4 멀티모달 모델(380)에 의해 상기 제3 오디오 신호에 대한 노이즈 캔슬링이 수행됨으로써, 생성될 수 있다.
예를 들면, 상기 제4 오디오 신호는, 상기 제3 오디오 신호 내의 노이즈가 감소된 오디오 신호를 나타낼 수 있다. 예를 들면, 상기 제4 오디오 신호는, 상기 제3 오디오 신호 내의 음성 신호가 강화된 오디오 신호를 나타낼 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 제3 오디오 신호에 대한 노이즈 캔슬링이 수행된 상기 제4 오디오 신호에 대한 데이터를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제3 오디오 신호에 대한 노이즈 캔슬링이 수행된 상기 제4 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득할 수 있다. 예를 들면, 상기 제4 오디오 신호에 대한 상기 데이터는, 제2 멀티모달 모델(360)이 이용하기 위하여, 상기 제4 오디오 신호에 대한 특징 추출이 수행된 데이터를 나타낼 수 있다.
예를 들면, 상기 제4 오디오 신호의 SNR은 상기 제3 오디오 신호의 SNR 보다 높을 수 있다. 예를 들면, 상기 제4 오디오 신호의 SNR은 상기 제3 오디오 신호의 SNR 보다 높기 때문에, 상기 제3 오디오 신호에 대한 데이터를 제2 멀티모달 모델(360)에 입력할 때 보다 상기 제4 오디오 신호에 대한 데이터를 제2 멀티모달 모델(360)에 입력할 때 제2 멀티모달 모델(360)의 ASR기능이 강화될 수 있다.
동작 925에서, 적어도 하나의 프로세서(300)는, 상기 제1 데이터 및 상기 제2 데이터를 외부 전자 장치(302)에게 통신 회로(310)를 통해 송신할 수 있다. 예를 들면, 동작 925는 동작 912에 대응할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)가 동작 925를 실행한 후에, 외부 전자 장치(302)는 동작 913 내지 동작 915를 실행할 수 있다. 예를 들면, 외부 전자 장치(302)가 동작 913 내지 동작 915를 실행한 후에, 적어도 하나의 프로세서(300)는 동작 916을 실행할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)가 도 9a 내지 9b의 예시된 동작들을 수행함으로써, 전자 장치(301)는 상기 환경에 따라, 마이크로폰(330)을 통해 획득된 오디오 신호에 대한 노이즈 캔슬링의 수행 여부를 결정할 수 있다.
예를 들면, 전자 장치(301)는 전자 장치(301)가 위치된 환경에 따라 효율적으로 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)을 이용할 수 있다. 예를 들면, 전자 장치(301)는 상기 환경에 따라 노이즈 캔슬링의 수행여부를 결정함으로써, 제2 멀티모달 모델(360)의 이용을 위해 소모되는 전류를 감소시킬 수 있다.
예를 들면, 전자 장치(301)는 획득된 오디오 신호에 대한 노이즈 캔슬링을 제4 멀티모달 모델(380)을 이용하여 수행함으로써, 음성 인식 기능의 품질을 강화시킬 수 있다. 예를 들면, 전자 장치(301)는 적어도 하나의 센서(340)를 통해 획득된 센싱 데이터를 외부 전자 장치(302)에게 송신함으로써, 음성인식 기능의 품질을 강화할 수 있다.
도 10a 및 10b는 응답 정보에 따른 기능을 수행하는 전자 장치의 예를 도시한다.
도 10a를 참고하면, 사용자가 전자 장치(301)에게 음성 신호를 발화하는 상태로 설명될 수 있다. 예를 들면, 상태(1010)에서, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 통해 상기 제2 센싱 데이터를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 상기 제3 오디오 신호를 획득할 수 있다.
예를 들면, 상태(1020)에서, 적어도 하나의 프로세서(300)는, 동작 916에서, 응답 정보에 따른 기능을 수행할 수 있다. 예를 들면, 상태(1020)에서 예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 따른 상기 기능을 수행함으로써, 전자 장치(301) 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 응답 정보에 나타나는 이모지 그래픽적 객체(1021)를 디스플레이(311)를 통해 표시할 수 있다.
예를 들면, 이모지 그래픽적 객체(1021)는, 제2 멀티모달 모델(360)에 의해 선택될 수 있다. 예를 들면, 이모지 그래픽적 객체(1021)는, 상기 제1 데이터 및/또는 상기 제2 데이터에 기반하여 제2 멀티모달 모델(360)에 의해 선택될 수 있다. 예를 들면, 이모지 그래픽적 객체(1021)는, 상기 제1 데이터 및/또는 상기 제2 데이터와 관련되어 메모리(320)에 미리 저장되어 있을 수 있다. 예를 들면, 이모지 그래픽적 객체(1021)는, 상기 제1 데이터 및/또는 상기 제2 데이터와 관련되어 외부 전자 장치(302)에 미리 저장되어 있을 수 있다.
도 10a에 의해 나타내어지는 이모지 그래픽적 객체(1021)는 단지 예시적인 것이다. 이모지 그래픽적 객체(1021)는, 도 10a의 도시와 다른 이모지 그래픽적 객체(1021)를 포함할 수 있다.
도 10b를 참고하면, 사용자가 전자 장치(301)에게 음성 신호를 발화하는 상태로 설명될 수 있다. 예를 들면, 상태(1030)에서, 적어도 하나의 프로세서(300)는, 도 9a 내지 도 9b에서 예시되는 동작들을 수행하고, 응답 정보에 따른 기능을 수행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 응답 정보에 따른 상기 기능을 수행함으로써, 전자 장치(301) 및/또는 외부 전자 장치(302) 내에 미리 저장된 텍스트들 중 상기 응답 정보에 의해 나타나는 텍스트(1031)를 디스플레이(311)를 통해 표시할 수 있다.
예를 들면, 텍스트(1031)는, 제2 멀티모달 모델(360)에 의해 선택될 수 있다. 예를 들면, 텍스트(1031)는, 상기 제1 데이터 및/또는 상기 제2 데이터에 기반하여 제2 멀티모달 모델(360)에 의해 선택될 수 있다. 예를 들면, 텍스트(1031)는, 상기 제1 데이터 및/또는 상기 제2 데이터와 관련되어 메모리(320)에 미리 저장되어 있을 수 있다. 예를 들면, 텍스트(1031)는, 상기 제1 데이터 및/또는 상기 제2 데이터와 관련되어 외부 전자 장치(302)에 미리 저장되어 있을 수 있다.
도 10b에 의해 나타내어지는 텍스트(1031)는 단지 예시적인 것이다. 텍스트(1031)는, 도 10b의 도시와 다른 텍스트(1031)를 포함할 수 있다.
도 11은, 전자 장치가 외부 전자 장치와 통화 연결을 수행하는 예를 도시한다. 도 11에서 예시되는 전자 장치(예: 전자 장치(301))의 동작들은, 적어도 하나의 프로세서(예: 적어도 하나의 프로세서(300))에 의해 실행되거나 수행되거나 제어될 수 있다. 이하 실시예에서 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
예를 들면, 동작 1101에서, 적어도 하나의 프로세서(300)는, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출할 수 있다. 예를 들면, 상기 이벤트는, 디스플레이(예: 디스플레이(311))를 통해 터치 입력을 수신하는 것을 포함할 수 있다. 예를 들면, 상기 이벤트는, 디스플레이(311)를 통해 표시되는 실행가능한 객체에 대하여 터치 입력을 수신하는 것을 포함할 수 있다. 예를 들면, 상기 이벤트는, 디스플레이(311)를 통해 표시되는 실행가능한 객체에 대하여 스크롤 입력을 수신하는 것을 포함할 수 있다. 하지만 이에 제한하지 않는다. 예를 들면, 상기 이벤트는, 도 6a에서 예시된 이벤트를 포함할 수 있다.
예를 들면, 동작 1102에서, 적어도 하나의 프로세서(300)는, 마이크로폰(예: 마이크로폰(330))을 통해 제5 오디오 신호를 획득할 수 있다.
예를 들면, 동작 1103에서, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 제5 오디오 신호로부터 전자 장치(301)가 위치된 환경의 유형을 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제3 모델(예: 제3 모델(370))을 이용하여 상기 환경의 유형을 식별할 수 있다. 제한하지 않는 예로, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(예: 적어도 하나의 센서(340))를 통해 획득된 센싱 데이터를 통해, 전자 장치(301)가 위치된 상기 환경의 상기 유형을 식별할 수 있다.
예를 들면, 상기 환경의 상기 유형은 상기 제1 유형 및 상기 제2 유형을 포함할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 이상인 경우, 제1 유형의 환경으로 식별할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 전자 장치(301)가 위치된 환경의 SNR이 기준 값 미만인 경우, 제2 유형의 환경으로 식별할 수 있다.
예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 동작 1104을 실행할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 동작 1106을 실행할 수 있다. 예를 들면, 동작 1103은 도 8a의 동작 812에 대응할 수 있다. 예를 들면, 동작 1103은, 도 9b의 동작 922들과 대응할 수 있다.
예를 들면, 동작 1104에서, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제6 오디오 신호를 획득할 수 있다. 예를 들면, 제6 오디오 신호는, 사용자의 음성을 포함할 수 있다.
예를 들면, 동작 1105에서, 적어도 하나의 프로세서(300)는, 상기 제6 오디오 신호를, 통신 회로(310)를 통해, 외부 전자 장치에게 송신할 수 있다. 예를 들면, 상기 외부 전자 장치는, 서버 및 기지국을 포함할 수 있다. 예를 들면, 상기 외부 전자 장치는, 다른 사용자의 전자 장치를 포함할 수 있다. 예를 들면, 동작 1105에서, 전자 장치(301)는, 외부 전자 장치와 통화 중인 상태를 나타낼 수 있다.
예를 들면, 동작 1106에서, 적어도 하나의 프로세서(300)는, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 활성화할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 적어도 하나의 센서(340)를 구동할 수 있다. 예를 들면, 적어도 하나의 센서(340)는, 적어도 하나의 프로세서(300)에 의해 동작 1106가 실행되기 전 비활성화 상태에 놓여 있을 수 있다. 예를 들면, 동작 1106은 도 8a의 동작 815에 대응할 수 있다.
예를 들면, 동작 1107에서, 적어도 하나의 프로세서(300)는, 적어도 하나의 센서(340)를 통해 제3 센싱 데이터(예: 제3 이미지(예: 이미지 센서를 통해 획득됨), 제3 심박 데이터(예: 심박 센서를 통해 획득됨), 및/또는 제3 움직임 데이터(예: 가속도 센서 및/또는 자이로 센서를 통해 획득됨))를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 이미지 센서를 통해 제3 이미지를 획득할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 마이크로폰(330)을 통해 제7 오디오 신호를 획득할 수 있다. 예를 들면, 동작 1107은, 도 8a의 동작 816에 대응할 수 있다.
예를 들면, 동작 1108에서, 적어도 하나의 프로세서(300)는, 상기 제7 오디오 신호 및 제3 센싱 데이터를 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 적어도 하나의 프로세서(300)는, 제7 오디오 신호 및 제3 센싱 데이터에 대한 응답 정보를 획득할 수 있다. 예를 들면, 상기 응답 정보는, 제7 오디오 신호에 대한 응답을 나타낼 수 있다. 예를 들면, 상기 응답 정보는, 제7 오디오 신호 내의 사용자의 발화에 대한 응답을 포함할 수 있다. 예를 들면, 상기 응답 정보는, 제3 센싱 데이터에 대한 응답을 나타낼 수 있다.
예를 들면, 동작 1109에서, 적어도 하나의 프로세서(300)는, 제7 오디오 신호 및 제3 센싱 데이터에 대한 응답 정보에 따른 기능을 수행할 수 있다. 예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 제7 오디오 신호에 대응하는 이모지 그래픽적 객체(예: 이모지 그래픽적 객체(1021))를, 디스플레이(311)를 통해 표시하도록 야기할 수 있다. 예를 들면, 이모지 그래픽적 객체(1021)는, 제1 멀티모달 모델(350)에 의해 선택될 수 있다. 예를 들면, 이모지 그래픽적 객체(1021)는, 제7 오디오 신호 및/또는 제3 센싱 데이터에 기반하여, 제1 멀티모달 모델(350)에 의해 생성될 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 통신 회로(310)를 통해, 외부 전자 장치에게 신호를 송신하도록 야기할 수 있다. 예를 들면, 상기 신호를 수신한 외부 전자 장치는, 제7 오디오 신호에 대응하는 이모지 그래픽적 객체(1021)를, 외부 전자 장치의 디스플레이를 통해 표시할 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 제3 센싱 데이터에 대응하는 이모지 그래픽적 객체(1021)를 디스플레이(311)를 통해 표시하도록 야기할 수 있다. 예를 들면, 제3 센싱 데이터는, 이미지 센서를 통해 획득된 이미지 내의 사용자의 입술에 대응하는 시각적 객체와 관련된 데이터를 포함할 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 통신 회로(310)를 통해, 외부 전자 장치에게 신호를 송신하도록 야기할 수 있다. 예를 들면, 상기 신호를 수신한 외부 전자 장치는, 제7 오디오 신호에 대응하는 이모지 그래픽적 객체(1021)를, 외부 전자 장치의 디스플레이를 통해 표시할 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 제7 오디오 신호에 대응하는 텍스트(예: 텍스트(1031))를, 디스플레이(311)를 통해 표시하도록 야기할 수 있다. 예를 들면, 텍스트(1031)는, 제1 멀티모달 모델(350)에 의해 선택될 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 통신 회로(310)를 통해, 외부 전자 장치에게 신호를 송신하도록 야기할 수 있다. 예를 들면, 상기 신호를 수신한 외부 전자 장치는, 제7 오디오 신호에 대응하는 텍스트(1031)를, 외부 전자 장치의 디스플레이를 통해 표시할 수 있다.
예를 들면, 텍스트(1031)는, 제1 멀티모달 모델(350)에 의해 선택될 수 있다. 예를 들면, 텍스트(1031)는, 제7 오디오 신호 및/또는 제3 센싱 데이터에 기반하여, 제1 멀티모달 모델(350)에 의해 생성될 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 제3 센싱 데이터에 대응하는 텍스트(1031)를 디스플레이(311)를 통해 표시하도록 야기할 수 있다. 예를 들면, 제3 센싱 데이터는, 이미지 센서를 통해 획득된 이미지 내의 사용자의 입술에 대응하는 시각적 객체와 관련된 데이터를 포함할 수 있다.
예를 들면, 상기 응답 정보는, 적어도 하나의 프로세서(300)가 통신 회로(310)를 통해, 외부 전자 장치에게 신호를 송신하도록 야기할 수 있다. 예를 들면, 상기 신호를 수신한 외부 전자 장치는, 제7 오디오 신호에 대응하는 텍스트(1031)를, 외부 전자 장치의 디스플레이를 통해 표시할 수 있다.
도 12는, 다양한 실시예들에 따른, 네트워크 환경(1200) 내의 전자 장치(1201)의 블록도이다. 예를 들면, 전자 장치(1201)는 전자 장치(301)를 포함할 수 있다.
도 12를 참고하면, 네트워크 환경(1200)에서 전자 장치(1201)는 제 1 네트워크(1298)(예: 근거리 무선 통신 네트워크)를 통하여 전자 장치(1202)와 통신하거나, 또는 제 2 네트워크(1299)(예: 원거리 무선 통신 네트워크)를 통하여 전자 장치(1204) 또는 서버(1208) 중 적어도 하나와 통신할 수 있다. 일실시예에 따르면, 전자 장치(1201)는 서버(1208)를 통하여 전자 장치(1204)와 통신할 수 있다. 일실시예에 따르면, 전자 장치(1201)는 프로세서(1220), 메모리(1230), 입력 모듈(1250), 음향 출력 모듈(1255), 디스플레이 모듈(1260), 오디오 모듈(1270), 센서 모듈(1276), 인터페이스(1277), 연결 단자(1278), 햅틱 모듈(1279), 카메라 모듈(1280), 전력 관리 모듈(1288), 배터리(1289), 통신 모듈(1290), 가입자 식별 모듈(1296), 또는 안테나 모듈(1297)을 포함할 수 있다. 어떤 실시예에서는, 전자 장치(1201)에는, 이 구성요소들 중 적어도 하나(예: 연결 단자(1278))가 생략되거나, 하나 이상의 다른 구성요소가 추가될 수 있다. 어떤 실시예에서는, 이 구성요소들 중 일부들(예: 센서 모듈(1276), 카메라 모듈(1280), 또는 안테나 모듈(1297))은 하나의 구성요소(예: 디스플레이 모듈(1260))로 통합될 수 있다.
프로세서(1220)는, 예를 들면, 소프트웨어(예: 프로그램(1240))를 실행하여 프로세서(1220)에 연결된 전자 장치(1201)의 적어도 하나의 다른 구성요소(예: 하드웨어 또는 소프트웨어 구성요소)를 제어할 수 있고, 다양한 데이터 처리 또는 연산을 수행할 수 있다. 일실시예에 따르면, 데이터 처리 또는 연산의 적어도 일부로서, 프로세서(1220)는 다른 구성요소(예: 센서 모듈(1276) 또는 통신 모듈(1290))로부터 수신된 명령 또는 데이터를 휘발성 메모리(1232)에 저장하고, 휘발성 메모리(1232)에 저장된 명령 또는 데이터를 처리하고, 결과 데이터를 비휘발성 메모리(1234)에 저장할 수 있다. 일실시예에 따르면, 프로세서(1220)는 메인 프로세서(1221)(예: 중앙 처리 장치 또는 어플리케이션 프로세서) 또는 이와는 독립적으로 또는 함께 운영 가능한 보조 프로세서(1223)(예: 그래픽 처리 장치, 신경망 처리 장치(NPU: neural processing unit), 이미지 시그널 프로세서, 센서 허브 프로세서, 또는 커뮤니케이션 프로세서)를 포함할 수 있다. 예를 들어, 전자 장치(1201)가 메인 프로세서(1221) 및 보조 프로세서(1223)를 포함하는 경우, 보조 프로세서(1223)는 메인 프로세서(1221)보다 저전력을 사용하거나, 지정된 기능에 특화되도록 설정될 수 있다. 보조 프로세서(1223)는 메인 프로세서(1221)와 별개로, 또는 그 일부로서 구현될 수 있다.
보조 프로세서(1223)는, 예를 들면, 메인 프로세서(1221)가 인액티브(예: 슬립) 상태에 있는 동안 메인 프로세서(1221)를 대신하여, 또는 메인 프로세서(1221)가 액티브(예: 어플리케이션 실행) 상태에 있는 동안 메인 프로세서(1221)와 함께, 전자 장치(1201)의 구성요소들 중 적어도 하나의 구성요소(예: 디스플레이 모듈(1260), 센서 모듈(1276), 또는 통신 모듈(1290))와 관련된 기능 또는 상태들의 적어도 일부를 제어할 수 있다. 일실시예에 따르면, 보조 프로세서(1223)(예: 이미지 시그널 프로세서 또는 커뮤니케이션 프로세서)는 기능적으로 관련 있는 다른 구성요소(예: 카메라 모듈(1280) 또는 통신 모듈(1290))의 일부로서 구현될 수 있다. 일실시예에 따르면, 보조 프로세서(1223)(예: 신경망 처리 장치)는 인공지능 모델의 처리에 특화된 하드웨어 구조를 포함할 수 있다. 인공지능 모델은 기계 학습을 통해 생성될 수 있다. 이러한 학습은, 예를 들어, 인공지능 모델이 수행되는 전자 장치(1201) 자체에서 수행될 수 있고, 별도의 서버(예: 서버(1208))를 통해 수행될 수도 있다. 학습 알고리즘은, 예를 들어, 지도형 학습(supervised learning), 비지도형 학습(unsupervised learning), 준지도형 학습(semi-supervised learning) 또는 강화 학습(reinforcement learning)을 포함할 수 있으나, 전술한 예에 한정되지 않는다. 인공지능 모델은, 복수의 인공 신경망 레이어들을 포함할 수 있다. 인공 신경망은 심층 신경망(DNN: deep neural network), CNN(convolutional neural network), RNN(recurrent neural network), RBM(restricted boltzmann machine), DBN(deep belief network), BRDNN(bidirectional recurrent deep neural network), 심층 Q-네트워크(deep Q-networks) 또는 상기 중 둘 이상의 조합 중 하나일 수 있으나, 전술한 예에 한정되지 않는다. 인공지능 모델은 하드웨어 구조 이외에, 추가적으로 또는 대체적으로, 소프트웨어 구조를 포함할 수 있다.
메모리(1230)는, 전자 장치(1201)의 적어도 하나의 구성요소(예: 프로세서(1220) 또는 센서 모듈(1276))에 의해 사용되는 다양한 데이터를 저장할 수 있다. 데이터는, 예를 들어, 소프트웨어(예: 프로그램(1240)) 및, 이와 관련된 명령에 대한 입력 데이터 또는 출력 데이터를 포함할 수 있다. 메모리(1230)는, 휘발성 메모리(1232) 또는 비휘발성 메모리(1234)를 포함할 수 있다.
프로그램(1240)은 메모리(1230)에 소프트웨어로서 저장될 수 있으며, 예를 들면, 운영 체제(1242), 미들 웨어(1244) 또는 어플리케이션(1246)을 포함할 수 있다.
입력 모듈(1250)은, 전자 장치(1201)의 구성요소(예: 프로세서(1220))에 사용될 명령 또는 데이터를 전자 장치(1201)의 외부(예: 사용자)로부터 수신할 수 있다. 입력 모듈(1250)은, 예를 들면, 마이크, 마우스, 키보드, 키(예: 버튼), 또는 디지털 펜(예: 스타일러스 펜)을 포함할 수 있다.
음향 출력 모듈(1255)은 음향 신호를 전자 장치(1201)의 외부로 출력할 수 있다. 음향 출력 모듈(1255)은, 예를 들면, 스피커 또는 리시버를 포함할 수 있다. 스피커는 멀티미디어 재생 또는 녹음 재생과 같이 일반적인 용도로 사용될 수 있다. 리시버는 착신 전화를 수신하기 위해 사용될 수 있다. 일실시예에 따르면, 리시버는 스피커와 별개로, 또는 그 일부로서 구현될 수 있다.
디스플레이 모듈(1260)은 전자 장치(1201)의 외부(예: 사용자)로 정보를 시각적으로 제공할 수 있다. 디스플레이 모듈(1260)은, 예를 들면, 디스플레이, 홀로그램 장치, 또는 프로젝터 및 해당 장치를 제어하기 위한 제어 회로를 포함할 수 있다. 일실시예에 따르면, 디스플레이 모듈(1260)은 터치를 감지하도록 설정된 터치 센서, 또는 상기 터치에 의해 발생되는 힘의 세기를 측정하도록 설정된 압력 센서를 포함할 수 있다.
오디오 모듈(1270)은 소리를 전기 신호로 변환시키거나, 반대로 전기 신호를 소리로 변환시킬 수 있다. 일실시예에 따르면, 오디오 모듈(1270)은, 입력 모듈(1250)을 통해 소리를 획득하거나, 음향 출력 모듈(1255), 또는 전자 장치(1201)와 직접 또는 무선으로 연결된 외부 전자 장치(예: 전자 장치(1202))(예: 스피커 또는 헤드폰)를 통해 소리를 출력할 수 있다.
센서 모듈(1276)은 전자 장치(1201)의 작동 상태(예: 전력 또는 온도), 또는 외부의 환경 상태(예: 사용자 상태)를 감지하고, 감지된 상태에 대응하는 전기 신호 또는 데이터 값을 생성할 수 있다. 일실시예에 따르면, 센서 모듈(1276)은, 예를 들면, 제스처 센서, 자이로 센서, 기압 센서, 마그네틱 센서, 가속도 센서, 그립 센서, 근접 센서, 컬러 센서, IR(infrared) 센서, 생체 센서, 온도 센서, 습도 센서, 또는 조도 센서를 포함할 수 있다.
인터페이스(1277)는 전자 장치(1201)가 외부 전자 장치(예: 전자 장치(1202))와 직접 또는 무선으로 연결되기 위해 사용될 수 있는 하나 이상의 지정된 프로토콜들을 지원할 수 있다. 일실시예에 따르면, 인터페이스(1277)는, 예를 들면, HDMI(high definition multimedia interface), USB(universal serial bus) 인터페이스, SD카드 인터페이스, 또는 오디오 인터페이스를 포함할 수 있다.
연결 단자(1278)는, 그를 통해서 전자 장치(1201)가 외부 전자 장치(예: 전자 장치(1202))와 물리적으로 연결될 수 있는 커넥터를 포함할 수 있다. 일실시예에 따르면, 연결 단자(1278)는, 예를 들면, HDMI 커넥터, USB 커넥터, SD 카드 커넥터, 또는 오디오 커넥터(예: 헤드폰 커넥터)를 포함할 수 있다.
햅틱 모듈(1279)은 전기적 신호를 사용자가 촉각 또는 운동 감각을 통해서 인지할 수 있는 기계적인 자극(예: 진동 또는 움직임) 또는 전기적인 자극으로 변환할 수 있다. 일실시예에 따르면, 햅틱 모듈(1279)은, 예를 들면, 모터, 압전 소자, 또는 전기 자극 장치를 포함할 수 있다.
카메라 모듈(1280)은 정지 영상 및 동영상을 촬영할 수 있다. 일실시예에 따르면, 카메라 모듈(1280)은 하나 이상의 렌즈들, 이미지 센서들, 이미지 시그널 프로세서들, 또는 플래시들을 포함할 수 있다.
전력 관리 모듈(1288)은 전자 장치(1201)에 공급되는 전력을 관리할 수 있다. 일실시예에 따르면, 전력 관리 모듈(1288)은, 예를 들면, PMIC(power management integrated circuit)의 적어도 일부로서 구현될 수 있다.
배터리(1289)는 전자 장치(1201)의 적어도 하나의 구성요소에 전력을 공급할 수 있다. 일실시예에 따르면, 배터리(1289)는, 예를 들면, 재충전 불가능한 1차 전지, 재충전 가능한 2차 전지 또는 연료 전지를 포함할 수 있다.
통신 모듈(1290)은 전자 장치(1201)와 외부 전자 장치(예: 전자 장치(1202), 전자 장치(1204), 또는 서버(1208)) 간의 직접(예: 유선) 통신 채널 또는 무선 통신 채널의 수립, 및 수립된 통신 채널을 통한 통신 수행을 지원할 수 있다. 통신 모듈(1290)은 프로세서(1220)(예: 어플리케이션 프로세서)와 독립적으로 운영되고, 직접(예: 유선) 통신 또는 무선 통신을 지원하는 하나 이상의 커뮤니케이션 프로세서를 포함할 수 있다. 일실시예에 따르면, 통신 모듈(1290)은 무선 통신 모듈(1292)(예: 셀룰러 통신 모듈, 근거리 무선 통신 모듈, 또는 GNSS(global navigation satellite system) 통신 모듈) 또는 유선 통신 모듈(1294)(예: LAN(local area network) 통신 모듈, 또는 전력선 통신 모듈)을 포함할 수 있다. 이들 통신 모듈 중 해당하는 통신 모듈은 제 1 네트워크(1298)(예: 블루투스, WiFi(wireless fidelity) direct 또는 IrDA(infrared data association)와 같은 근거리 통신 네트워크) 또는 제 2 네트워크(1299)(예: 레거시 셀룰러 네트워크, 5G 네트워크, 차세대 통신 네트워크, 인터넷, 또는 컴퓨터 네트워크(예: LAN 또는 WAN)와 같은 원거리 통신 네트워크)를 통하여 외부의 전자 장치(1204)와 통신할 수 있다. 이런 여러 종류의 통신 모듈들은 하나의 구성요소(예: 단일 칩)로 통합되거나, 또는 서로 별도의 복수의 구성요소들(예: 복수 칩들)로 구현될 수 있다. 무선 통신 모듈(1292)은 가입자 식별 모듈(1296)에 저장된 가입자 정보(예: 국제 모바일 가입자 식별자(IMSI))를 이용하여 제 1 네트워크(1298) 또는 제 2 네트워크(1299)와 같은 통신 네트워크 내에서 전자 장치(1201)를 확인 또는 인증할 수 있다.
무선 통신 모듈(1292)은 4G 네트워크 이후의 5G 네트워크 및 차세대 통신 기술, 예를 들어, NR 접속 기술(new radio access technology)을 지원할 수 있다. NR 접속 기술은 고용량 데이터의 고속 전송(eMBB(enhanced mobile broadband)), 단말 전력 최소화와 다수 단말의 접속(mMTC(massive machine type communications)), 또는 고신뢰도와 저지연(URLLC(ultra-reliable and low-latency communications))을 지원할 수 있다. 무선 통신 모듈(1292)은, 예를 들어, 높은 데이터 전송률 달성을 위해, 고주파 대역(예: mmWave 대역)을 지원할 수 있다. 무선 통신 모듈(1292)은 고주파 대역에서의 성능 확보를 위한 다양한 기술들, 예를 들어, 빔포밍(beamforming), 거대 배열 다중 입출력(massive MIMO(multiple-input and multiple-output)), 전차원 다중입출력(FD-MIMO: full dimensional MIMO), 어레이 안테나(array antenna), 아날로그 빔형성(analog beam-forming), 또는 대규모 안테나(large scale antenna)와 같은 기술들을 지원할 수 있다. 무선 통신 모듈(1292)은 전자 장치(1201), 외부 전자 장치(예: 전자 장치(1204)) 또는 네트워크 시스템(예: 제 2 네트워크(1299))에 규정되는 다양한 요구사항을 지원할 수 있다. 일실시예에 따르면, 무선 통신 모듈(1292)은 eMBB 실현을 위한 Peak data rate(예: 20Gbps 이상), mMTC 실현을 위한 손실 Coverage(예: 164dB 이하), 또는 URLLC 실현을 위한 U-plane latency(예: 다운링크(DL) 및 업링크(UL) 각각 0.5ms 이하, 또는 라운드 트립 1ms 이하)를 지원할 수 있다.
안테나 모듈(1297)은 신호 또는 전력을 외부(예: 외부의 전자 장치)로 송신하거나 외부로부터 수신할 수 있다. 일실시예에 따르면, 안테나 모듈(1297)은 서브스트레이트(예: PCB) 위에 형성된 도전체 또는 도전성 패턴으로 이루어진 방사체를 포함하는 안테나를 포함할 수 있다. 일실시예에 따르면, 안테나 모듈(1297)은 복수의 안테나들(예: 어레이 안테나)을 포함할 수 있다. 이런 경우, 제 1 네트워크(1298) 또는 제 2 네트워크(1299)와 같은 통신 네트워크에서 사용되는 통신 방식에 적합한 적어도 하나의 안테나가, 예를 들면, 통신 모듈(1290)에 의하여 상기 복수의 안테나들로부터 선택될 수 있다. 신호 또는 전력은 상기 선택된 적어도 하나의 안테나를 통하여 통신 모듈(1290)과 외부의 전자 장치 간에 송신되거나 수신될 수 있다. 어떤 실시예에 따르면, 방사체 이외에 다른 부품(예: RFIC(radio frequency integrated circuit))이 추가로 안테나 모듈(1297)의 일부로 형성될 수 있다.
다양한 실시예에 따르면, 안테나 모듈(1297)은 mmWave 안테나 모듈을 형성할 수 있다. 일실시예에 따르면, mmWave 안테나 모듈은 인쇄 회로 기판, 상기 인쇄 회로 기판의 제 1 면(예: 아래 면)에 또는 그에 인접하여 배치되고 지정된 고주파 대역(예: mmWave 대역)을 지원할 수 있는 RFIC, 및 상기 인쇄 회로 기판의 제 2 면(예: 윗 면 또는 측 면)에 또는 그에 인접하여 배치되고 상기 지정된 고주파 대역의 신호를 송신 또는 수신할 수 있는 복수의 안테나들(예: 어레이 안테나)을 포함할 수 있다.
상기 구성요소들 중 적어도 일부는 주변 기기들간 통신 방식(예: 버스, GPIO(general purpose input and output), SPI(serial peripheral interface), 또는 MIPI(mobile industry processor interface))을 통해 서로 연결되고 신호(예: 명령 또는 데이터)를 상호간에 교환할 수 있다.
일실시예에 따르면, 명령 또는 데이터는 제 2 네트워크(1299)에 연결된 서버(1208)를 통해서 전자 장치(1201)와 외부의 전자 장치(1204)간에 송신 또는 수신될 수 있다. 외부의 전자 장치(1202, 또는 1204) 각각은 전자 장치(1201)와 동일한 또는 다른 종류의 장치일 수 있다. 일실시예에 따르면, 전자 장치(1201)에서 실행되는 동작들의 전부 또는 일부는 외부의 전자 장치들(1202, 1204, 또는 1208) 중 하나 이상의 외부의 전자 장치들에서 실행될 수 있다. 예를 들면, 전자 장치(1201)가 어떤 기능이나 서비스를 자동으로, 또는 사용자 또는 다른 장치로부터의 요청에 반응하여 수행해야 할 경우에, 전자 장치(1201)는 기능 또는 서비스를 자체적으로 실행시키는 대신에 또는 추가적으로, 하나 이상의 외부의 전자 장치들에게 그 기능 또는 그 서비스의 적어도 일부를 수행하라고 요청할 수 있다. 상기 요청을 수신한 하나 이상의 외부의 전자 장치들은 요청된 기능 또는 서비스의 적어도 일부, 또는 상기 요청과 관련된 추가 기능 또는 서비스를 실행하고, 그 실행의 결과를 전자 장치(1201)로 전달할 수 있다. 전자 장치(1201)는 상기 결과를, 그대로 또는 추가적으로 처리하여, 상기 요청에 대한 응답의 적어도 일부로서 제공할 수 있다. 이를 위하여, 예를 들면, 클라우드 컴퓨팅, 분산 컴퓨팅, 모바일 에지 컴퓨팅(MEC: mobile edge computing), 또는 클라이언트-서버 컴퓨팅 기술이 이용될 수 있다. 전자 장치(1201)는, 예를 들어, 분산 컴퓨팅 또는 모바일 에지 컴퓨팅을 이용하여 초저지연 서비스를 제공할 수 있다. 다른 실시예에 있어서, 외부의 전자 장치(1204)는 IoT(internet of things) 기기를 포함할 수 있다. 서버(1208)는 기계 학습 및/또는 신경망을 이용한 지능형 서버일 수 있다. 일실시예에 따르면, 외부의 전자 장치(1204) 또는 서버(1208)는 제 2 네트워크(1299) 내에 포함될 수 있다. 전자 장치(1201)는 5G 통신 기술 및 IoT 관련 기술을 기반으로 지능형 서비스(예: 스마트 홈, 스마트 시티, 스마트 카, 또는 헬스 케어)에 적용될 수 있다.
본 개시의 실시예들에서, 전자 장치(예: 전자 장치(301))는 가상 공간에 이미지를 표시하기 위한 전자 장치일 수 있다. 예를 들면, 가상 공간에 이미지를 표시하기 위한 전자 장치(예: 도 1의 전자 장치(301))는 웨어러블 장치일 수 있다. 이하에서, 전자 장치의 일 예로 웨어러블 장치(예: 후술될 웨어러블 장치(1301))에 대한 설명들이 서술될 수 있다. 웨어러블 장치는 사용자의 머리에 착용 가능한(wearable on) HMD(head-mounted display)를 포함할 수 있다. 예를 들어, 웨어러블 장치는 머리-착용 전자 장치로 지칭될 수 있다. 웨어러블 장치는 HMD(head-mount device), 헤드기어(headgear) 전자 장치, 안경형(glasses-type) 전자 장치, VST(video see-through 또는 visible see-through) 디바이스, XR(extended reality) 디바이스, VR(virtual reality) 디바이스 및/또는 AR(augmented reality) 디바이스로 지칭될(referred) 수 있다. 비록 안경의 형태를 가지는 웨어러블 장치의 외형이 도시되지만, 실시예가 이에 제한되는 것은 아니다. 웨어러블 장치 내에 포함된 하드웨어 구성(hardware configuration)의 일 예가, 도 16를 참고하여 예시적으로 설명된다. 사용자의 머리에 착용가능한 웨어러블 장치의 구조의 일 예가, 도 13a, 도 13b, 도 13c, 도 14a, 도 14b, 및/또는 도 15를 참고하여 설명된다. 웨어러블 장치는 전자 장치로 지칭될(referred) 수 있다. 예를 들어, 웨어러블 장치는, 사용자의 머리에 부착되기 위한 액세서리(예: 스트랩(strap))과 결합되어, HMD를 형성할 수 있다.
일 실시예에 따른, 웨어러블 장치는 증강 현실(augmented reality, AR) 및/또는 혼합 현실(mixed reality, MR)과 관련된 기능을 실행할 수 있다. 예를 들어, 사용자가 웨어러블 장치를 착용한 상태 내에서, 웨어러블 장치는 사용자의 눈에 인접하게 배치된 적어도 하나의 렌즈를 포함할 수 있다. 웨어러블 장치는 렌즈를 통과하는 주변 광에, 웨어러블 장치의 디스플레이로부터 방사된 광을 결합할 수 있다. 상기 디스플레이의 표시 영역은, 주변 광이 통과되는 렌즈 내에서 형성될 수 있다. 웨어러블 장치가 상기 주변 광 및 상기 디스플레이로부터 방사된 상기 광을 결합하기 때문에, 사용자는 상기 주변 광에 의해 인식되는 실제 객체(real object) 및 상기 디스플레이로부터 방사된 상기 광에 의해 형성된 가상 객체(virtual object)가 혼합된 상(image)을 볼 수 있다. 상술된 증강 현실, 혼합 현실, 및/또는 가상 현실은, 확장 현실(extended reality, XR)로 지칭될(referred) 수 있다.
일 실시예에 따른, 웨어러블 장치는 VST(video see-through 또는 visible see-through) 및/또는 가상 현실(virtual reality, VR)과 관련된 기능을 실행할 수 있다. 예를 들어, 사용자가 웨어러블 장치를 착용한 상태 내에서, 웨어러블 장치는 사용자의 눈을 덮는 하우징을 포함할 수 있다. 웨어러블 장치는, 상기 상태 내에서, 상기 눈을 향하는 상기 하우징의 제1 면에 배치된 디스플레이를 포함할 수 있다. 웨어러블 장치는 상기 제1 면과 반대인 제2 면 상에 배치된 카메라를 포함할 수 있다. 상기 카메라를 이용하여, 웨어러블 장치는 주변 광(ambient light)을 표현하는(representing) 이미지 및/또는 비디오를 획득할 수 있다. 웨어러블 장치는 상기 제1 면에 배치된 디스플레이 내에, 상기 이미지 및/또는 비디오를 출력하여, 사용자가 상기 디스플레이를 통해 상기 주변 광을 인식하게 할 수 있다. 상기 제1 면에 배치된 디스플레이의 표시 영역(displaying area 또는 displaying region)(또는 활성 영역(active area 또는 active region))은, 상기 디스플레이에 포함된 하나 이상의 픽셀들에 의해 형성될 수 있다. 웨어러블 장치는 상기 디스플레이를 통해 출력되는 이미지 및/또는 비디오에 가상 객체를 합성하여, 사용자가 주변 광에 의해 인식되는 실제 객체와 함께 상기 가상 객체를 인식하게 할 수 있다.
일 실시예에 따른, 웨어러블 장치는, 카메라를 이용하여 획득된(obtained 또는 acquired) 이미지(및/또는 비디오)에 기반하여, 웨어러블 장치의 위치(position 또는 location) 및/또는 방향(direction 또는 orientation)을 식별할 수 있거나, 또는 인식(recognize)할 수 있다. 웨어러블 장치는 하나 이상의 카메라들 및/또는 하나 이상의 센서들을 이용하여 상기 외부 공간에 대한 정보를 획득할 수 있다. 상기 정보는, 하나 이상의 센서들로부터 식별된, 외부 공간의 지리적 위치(geographic location)(예: GPS(global positioning system) 좌표)를 포함할 수 있다. 상기 정보는, 하나 이상의 카메라들로부터 식별된, 외부 공간에 대한 이미지 및/또는 비디오를 포함할 수 있다. 웨어러블 장치는 이미지 및/또는 비디오에 대한 객체 인식을 수행하여, 상기 이미지 및/또는 상기 비디오로부터, 상기 외부 공간에 포함된 외부 객체들을 식별할 수 있다.
이하에서는, 도 13a, 도 13b, 도 13c, 도 14a, 도 14b, 도 15, 및/또는 도 16를 참고하여, 웨어러블 장치의 하드웨어 구성(hardware configuration)의 일 예가 설명된다.
도 13a는 웨어러블 장치의 사시도(perspective view)의 일 예를 도시한다. 도 13b는 웨어러블 장치 내에 배치된 하나 이상의 하드웨어들의 일 예를 도시한다. 도 13c는 일 실시예에 따른 웨어러블 장치의 예를 도시한다. 웨어러블 장치(1301)는, 사용자의 신체 부위(예: 머리) 상에 착용 가능한(wearable on), 안경의 형태를 가질 수 있다. 도 13a 내지 도 13c의 웨어러블 장치(1301)는, 도 3a의 전자 장치(301)의 일 예일 수 있다. 웨어러블 장치(1301)는, HMD(head-mounted display)를 포함할 수 있다. 예를 들어, 웨어러블 장치(1301)의 하우징은 사용자의 머리의 일부분(예를 들어, 두 눈을 감싸는 얼굴의 일부분)에 밀착되는 형태를 가지는 고무, 및/또는 실리콘과 같은 유연성 소재(flexible material)를 포함할 수 있다. 예를 들어, 웨어러블 장치(1301)의 하우징은 사용자의 머리에 감길 수 있는(able to be twined around) 하나 이상의 스트랩들, 및/또는 상기 머리의 귀로 탈착 가능한(attachable to) 하나 이상의 템플들(temples)을 포함할 수 있다.
도 13a를 참고하면, 일 실시예에 따른, 웨어러블 장치(1301)는, 적어도 하나의 디스플레이(1350), 및 적어도 하나의 디스플레이(1350)를 지지하는 프레임(1300)을 포함할 수 있다. 예를 들면, 적어도 하나의 디스플레이(1350)는, 도 3a의 디스플레이(311)의 일 예일 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)는 사용자의 신체의 일부 상에 착용될 수 있다. 웨어러블 장치(1301)는, 웨어러블 장치(1301)를 착용한 사용자에게, 증강 현실(AR), 가상 현실(VR), 또는 증강 현실과 가상 현실을 혼합한 혼합 현실(MR)을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 도 13b의 동작 인식 카메라(1360-2, 1360-3)를 통해 획득된 사용자의 지정된 제스처에 응답하여, 도 13b의 적어도 하나의 광학 장치(1382, 1384)에서 제공되는 가상 현실 영상을 적어도 하나의 디스플레이(1350)에 표시할 수 있다.
일 실시예에 따르면, 적어도 하나의 디스플레이(1350)는, 사용자에게 시각 정보를 제공할 수 있다. 예를 들면, 적어도 하나의 디스플레이(1350)는, 투명 또는 반투명한 렌즈를 포함할 수 있다. 적어도 하나의 디스플레이(1350)는, 제1 디스플레이(1350-1) 및/또는 제1 디스플레이(1350-1)로부터 이격된 제2 디스플레이(1350-2)를 포함할 수 있다. 예를 들면, 제1 디스플레이(1350-1), 및 제2 디스플레이(1350-2)는, 사용자의 좌안과 우안에 각각 대응되는 위치에 배치될 수 있다.
도 13b를 참고하면, 적어도 하나의 디스플레이(1350)는, 적어도 하나의 디스플레이(1350)에 포함되는 렌즈를 통해 사용자에게 외부 광으로부터 전달되는 시각적 정보와, 상기 시각적 정보와 구별되는 다른 시각적 정보를 제공할 수 있다. 상기 렌즈는, 프레넬(fresnel) 렌즈, 팬케이크(pancake) 렌즈, 또는 멀티-채널 렌즈 중 적어도 하나에 기반하여 형성될 수 있다. 예를 들면, 적어도 하나의 디스플레이(1350)는, 제1 면(surface)(1331), 및 제1 면(1331)에 반대인 제2 면(1332)을 포함할 수 있다. 적어도 하나의 디스플레이(1350)의 제2 면(1332) 상에, 표시 영역이 형성될 수 있다. 사용자가 웨어러블 장치(1301)를 착용하였을 때, 외부 광은 제1 면(1331)으로 입사되고, 제2 면(1332)을 통해 투과됨으로써, 사용자에게 전달될 수 있다. 다른 예를 들면, 적어도 하나의 디스플레이(1350)는, 외부 광을 통해 전달되는 현실 화면에, 적어도 하나의 광학 장치(1382, 1384)에서 제공되는 가상 현실 영상이 결합된 증강 현실 영상을, 제2 면(1332) 상에 형성된 표시 영역에 표시할 수 있다.
일 실시예에서, 적어도 하나의 디스플레이(1350)는, 적어도 하나의 광학 장치(1382, 1384)에서 송출된 광을 회절시켜, 사용자에게 전달하는, 적어도 하나의 웨이브가이드(waveguide)(1333, 1334)를 포함할 수 있다. 적어도 하나의 웨이브가이드(1333, 1334)는, 글래스, 플라스틱, 또는 폴리머 중 적어도 하나에 기반하여 형성될 수 있다. 적어도 하나의 웨이브가이드(1333, 1334)의 외부, 또는 내부의 적어도 일부분에, 나노 패턴이 형성될 수 있다. 상기 나노 패턴은, 다각형, 및/또는 곡면 형상의 격자 구조(grating structure)에 기반하여 형성될 수 있다. 적어도 하나의 웨이브가이드(1333, 1334)의 일 단으로 입사된 광은, 상기 나노 패턴에 의해 적어도 하나의 웨이브가이드(1333, 1334)의 타 단으로 전파될 수 있다. 적어도 하나의 웨이브가이드(1333, 1334)는 적어도 하나의 회절 요소(예: DOE(diffractive optical element), HOE(holographic optical element)), 반사 요소(예: 반사 거울) 중 적어도 하나를 포함할 수 있다. 예를 들어, 적어도 하나의 웨이브가이드(1333, 1334)는, 적어도 하나의 디스플레이(1350)에 의해 표시되는 화면을, 사용자의 눈으로 가이드하기 위하여, 웨어러블 장치(1301) 내에 배치될 수 있다. 예를 들어, 상기 화면은, 적어도 하나의 웨이브가이드(1333, 1334) 내에서 발생되는 전반사(total internal reflection, TIR)에 기반하여, 사용자의 눈으로 송신될 수 있다.
웨어러블 장치(1301)는, 촬영 카메라(1360-4)를 통해 수집된 현실 영상에 포함된 오브젝트(object)를 분석하고, 분석된 오브젝트 중에서 증강 현실 제공의 대상이 되는 오브젝트에 대응되는 가상 오브젝트(virtual object)를 결합하여, 적어도 하나의 디스플레이(1350)에 표시할 수 있다. 가상 오브젝트는, 현실 영상에 포함된 오브젝트에 관련된 다양한 정보에 대한 텍스트, 및 이미지 중 적어도 하나를 포함할 수 있다. 웨어러블 장치(1301)는, 스테레오 카메라와 같은 멀티-카메라에 기반하여, 오브젝트를 분석할 수 있다. 상기 오브젝트 분석을 위하여, 웨어러블 장치(1301)는 멀티-카메라, 및/또는, ToF(time-of-flight)를 이용하여, 공간 인식(예: SLAM(simultaneous localization and mapping))을 실행할 수 있다. 웨어러블 장치(1301)를 착용한 사용자는, 적어도 하나의 디스플레이(1350)에 표시되는 영상을 시청할 수 있다.
일 실시예에 따르면, 프레임(1300)은, 웨어러블 장치(1301)가 사용자의 신체 상에 착용될 수 있는 물리적인 구조로 이루어질 수 있다. 일 실시예에 따르면, 프레임(1300)은, 사용자가 웨어러블 장치(1301)를 착용하였을 때, 제1 디스플레이(1350-1) 및 제2 디스플레이(1350-2)가 사용자의 좌안 및 우안에 대응되는 위치할 수 있도록, 구성될 수 있다. 프레임(1300)은, 적어도 하나의 디스플레이(1350)를 지지할 수 있다. 예를 들면, 프레임(1300)은, 제1 디스플레이(1350-1) 및 제2 디스플레이(1350-2)를 사용자의 좌안 및 우안에 대응되는 위치에 위치되도록 지지할 수 있다.
도 13a를 참고하면, 프레임(1300)은, 사용자가 웨어러블 장치(1301)를 착용한 경우, 적어도 일부가 사용자의 신체의 일부분과 접촉되는 영역(1320)을 포함할 수 있다. 예를 들면, 프레임(1300)의 사용자의 신체의 일부분과 접촉되는 영역(1320)은, 웨어러블 장치(1301)가 접하는 사용자의 코의 일부분, 사용자의 귀의 일부분 및 사용자의 얼굴의 측면 일부분과 접촉하는 영역을 포함할 수 있다. 일 실시예에 따르면, 프레임(1300)은, 사용자의 신체의 일부 상에 접촉되는 노즈 패드(1310)를 포함할 수 있다. 웨어러블 장치(1301)가 사용자에 의해 착용될 시, 노즈 패드(1310)는, 사용자의 코의 일부 상에 접촉될 수 있다. 프레임(1300)은, 상기 사용자의 신체의 일부와 구별되는 사용자의 신체의 다른 일부 상에 접촉되는 제1 템플(temple)(1304) 및 제2 템플(1305)을 포함할 수 있다.
예를 들면, 프레임(1300)은, 제1 디스플레이(1350-1)의 적어도 일부를 감싸는 제1 림(rim)(1302-1), 제2 디스플레이(1350-2)의 적어도 일부를 감싸는 제2 림(1302-2), 제1 림(1302-1)과 제2 림(1302-2) 사이에 배치되는 브릿지(bridge)(1303), 브릿지(1303)의 일단으로부터 제1 림(1302-1)의 가장자리 일부를 따라 배치되는 제1 패드(1311), 브릿지(1303)의 타단으로부터 제2 림(1302-2)의 가장자리 일부를 따라 배치되는 제2 패드(1312), 제1 림(1302-1)으로부터 연장되어 착용자의 귀의 일부분에 고정되는 제1 템플(1304), 및 제2 림(1302-2)으로부터 연장되어 상기 귀의 반대측 귀의 일부분에 고정되는 제2 템플(1305)을 포함할 수 있다. 제1 패드(1311), 및 제2 패드(1312)는, 사용자의 코의 일부분과 접촉될 수 있고, 제1 템플(1304) 및 제2 템플(1305)은, 사용자의 안면의 일부분 및 귀의 일부분과 접촉될 수 있다. 템플(1304, 1305)은, 도 13b의 힌지 유닛들(1306, 1307)을 통해 림과 회전 가능하게(rotatably) 연결될 수 있다. 제1 템플(1304)은, 제1 림(1302-1)과 제1 템플(1304)의 사이에 배치된 제1 힌지 유닛(1306)을 통해, 제1 림(1302-1)에 대하여 회전 가능하게 연결될 수 있다. 제2 템플(1305)은, 제2 림(1302-2)과 제2 템플(1305)의 사이에 배치된 제2 힌지 유닛(1307)을 통해 제2 림(1302-2)에 대하여 회전 가능하게 연결될 수 있다. 일 실시예에 따른, 웨어러블 장치(1301)는 프레임(1300)의 표면의 적어도 일부분 상에 형성된, 터치 센서, 그립 센서, 및/또는 근접 센서를 이용하여, 프레임(1300)을 터치하는 외부 객체(예: 사용자의 손끝(fingertip)), 및/또는 상기 외부 객체에 의해 수행된 제스처를 식별할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는, 다양한 기능들을 수행하는 하드웨어들(예: 도 16의 블록도에 기반하여 후술될 하드웨어들)을 포함할 수 있다. 예를 들면, 상기 하드웨어들은, 배터리 모듈(1370), 안테나 모듈(1375), 적어도 하나의 광학 장치(1382, 1384), 스피커들(예: 스피커들(1355-1, 1355-2)), 마이크로폰(예: 마이크로폰들(1365-1, 1365-2, 1365-3)), 발광 모듈(미도시), 및/또는 PCB(printed circuit board)(1390)(예: 인쇄 회로 기판)을 포함할 수 있다. 다양한 하드웨어들은, 프레임(1300) 내에 배치될 수 있다. 예를 들면, 마이크로폰들(1365-1, 1365-2, 1365-3)은 도 3a의 마이크로폰(330)의 일 예일 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)의 마이크로폰(예: 마이크로폰들(1365-1, 1365-2, 1365-3))는, 프레임(1300)의 적어도 일부분에 배치되어, 소리 신호를 획득할 수 있다. 브릿지(1303) 상에 배치된 제1 마이크로폰(1365-1), 제2 림(1302-2) 상에 배치된 제2 마이크로폰(1365-2), 및 제1 림(1302-1) 상에 배치된 제3 마이크로폰(1365-3)이 도 13b 내에 도시되지만, 마이크로폰(1365)의 개수, 및 배치가 도 13b의 일 실시예에 제한되는 것은 아니다. 웨어러블 장치(1301) 내에 포함된 마이크로폰(1365)의 개수가 두 개 이상인 경우, 웨어러블 장치(1301)는 프레임(1300)의 상이한 부분들 상에 배치된 복수의 마이크들을 이용하여, 소리 신호의 방향을 식별할 수 있다.
일 실시예에 따르면, 적어도 하나의 광학 장치(1382, 1384)는, 다양한 이미지 정보를 사용자에게 제공하기 위하여, 적어도 하나의 디스플레이(1350)에 가상 오브젝트를 투영할 수 있다. 예를 들면, 적어도 하나의 광학 장치(1382, 1384)는, 프로젝터일 수 있다. 적어도 하나의 광학 장치(1382, 1384)는, 적어도 하나의 디스플레이(1350)에 인접하여 배치되거나, 적어도 하나의 디스플레이(1350)의 일부로써, 적어도 하나의 디스플레이(1350) 내에 포함될 수 있다. 일 실시예에 따르면, 웨어러블 장치(1301)는, 제1 디스플레이(1350-1)에 대응되는, 제1 광학 장치(1382) 및 제2 디스플레이(1350-2)에 대응되는, 제2 광학 장치(1384)를 포함할 수 있다. 예를 들면, 적어도 하나의 광학 장치(1382, 1384)는, 제1 디스플레이(1350-1)의 가장자리에 배치되는 제1 광학 장치(1382) 및 제2 디스플레이(1350-2)의 가장자리에 배치되는 제2 광학 장치(1384)를 포함할 수 있다. 제1 광학 장치(1382)는, 제1 디스플레이(1350-1) 상에 배치된 제1 웨이브가이드(1333)로 광을 송출할 수 있고, 제2 광학 장치(1384)는, 제2 디스플레이(1350-2) 상에 배치된 제2 웨이브가이드(1334)로 광을 송출할 수 있다.
일 실시예에서, 카메라(1360)는, 촬영 카메라(1360-4), 시선 추적 카메라(eye tracking camera, ET CAM)(1360-1), 및/또는 동작 인식 카메라(1360-2, 1306-3)를 포함할 수 있다. 촬영 카메라(1360-4), 시선 추적 카메라(1360-1) 및 동작 인식 카메라(1360-2, 1360-3)는, 프레임(1300) 상에서 서로 다른 위치에 배치될 수 있고, 서로 다른 기능을 수행할 수 있다. 시선 추적 카메라(1360-1)는, 웨어러블 장치(1301)를 착용한 사용자의 눈의 위치 또는 시선(gaze)을 나타내는 데이터를 출력할 수 있다. 예를 들어, 웨어러블 장치(1301)는 시선 추적 카메라(1360-1)를 통하여 획득된, 사용자의 눈동자가 포함된 이미지로부터, 상기 시선을 탐지할 수 있다. 웨어러블 장치(1301)는 시선 추적 카메라(1360-1)를 통해 획득된 사용자의 시선을 이용하여, 사용자에 의해 포커스 된 객체(예: 실제 객체, 및/또는 가상 객체)를 식별할 수 있다. 포커스된 객체를 식별한 웨어러블 장치(1301)는, 사용자 및 포커스 된 객체 사이의 인터랙션을 위한 기능(예: gaze interaction)을 실행할 수 있다. 웨어러블 장치(1301)는 시선 추적 카메라(1360-1)를 통해 획득된 사용자의 시선을 이용하여, 가상 공간 내 사용자를 나타내는 아바타의 눈에 대응하는 부분을 표현할 수 있다. 웨어러블 장치(1301)는 사용자의 눈의 위치에 기반하여, 적어도 하나의 디스플레이(1350) 상에 표시되는 이미지(또는 화면)를 렌더링할 수 있다. 예를 들어, 이미지 내에서 시선과 관련된 제1 영역의 시각적 품질 및 상기 제1 영역과 구분되는 제2 영역의 시각적 품질(예: 해상도, 밝기, 채도, 그레이스케일, PPI(pixels per inch))은 서로 다를 수 있다. 본 개시에서, "해상도" 용어는, 이미지 및/또는 디스플레이(1350)의 픽셀들의 밀도를 지칭하기 위하여 이용된다. 픽셀들의 밀도 및/또는 해상도는, PPI 및/또는 dpi(dots per inch)의 단위에 기반하여 측정될 수 있거나, 또는 매개변수화될 수 있다. 웨어러블 장치(1301)는, 포비티드 렌더링(foveated rendering)을 이용하여, 사용자의 시선에 매칭되는 제1 영역의 시각적 품질 및 상기 제2 영역의 시각적 품질을 가지는 이미지를 획득할 수 있다. 예를 들어, 웨어러블 장치(1301)가 홍채 인식 기능을 지원하는 경우, 시선 추적 카메라(1360-1)를 이용하여 획득된 홍채 정보에 기반하여, 사용자 인증을 수행할 수 있다. 시선 추적 카메라(1360-1)가 사용자의 우측 눈을 향하여 배치된 일 예가 도 13b 내에 도시되지만, 실시예가 이에 제한되는 것은 아니며, 시선 추적 카메라(1360-1)는, 사용자의 좌측 눈을 향하여 단독으로 배치되거나, 또는 양 눈들 전부를 향하여 배치될 수 있다.
일 실시예에서, 촬영 카메라(1360-4)는, 증강 현실 또는 혼합 현실 콘텐츠를 구현하기 위해서 가상의 이미지와 정합될 실제의 이미지나 배경을 촬영할 수 있다. 촬영 카메라(1360-4)는 HR(high resolution) 또는 PV(photo video)에 기반하여, 고해상도를 가지는 이미지를 획득하기 위해 이용될 수 있다. 촬영 카메라(1360-4)는, 사용자가 바라보는 위치에 존재하는 특정 사물의 이미지를 촬영하고, 그 이미지를 적어도 하나의 디스플레이(1350)로 제공할 수 있다. 적어도 하나의 디스플레이(1350)는, 촬영 카메라(1360-4)를 이용해 획득된 상기 특정 사물의 이미지를 포함하는 실제의 이미지나 배경에 관한 정보와, 적어도 하나의 광학 장치(1382, 1384)를 통해 제공되는 가상 이미지가 겹쳐진 하나의 영상을 표시할 수 있다. 웨어러블 장치(1301)는 촬영 카메라(1360-4)통해 획득된 이미지를 이용하여, 깊이 정보(예: 깊이 센서를 통해 획득된 웨어러블 장치(1301) 및 외부 객체 사이의 거리)를 보상할 수 있다. 웨어러블 장치(1301)는 촬영 카메라(1360-4)를 이용하여 획득된 이미지를 통해, 객체 인식을 수행할 수 있다. 웨어러블 장치(1301)는 촬영 카메라(1360-4)를 이용하여 이미지 내 객체(또는 피사체)에 초점을 맞추는 기능(예: AF(auto focus)) 및/또는 OIS(optical image stabilization) 기능(예: 손떨림 방지 기능)을 수행할 수 있다. 웨어러블 장치(1301)는 적어도 하나의 디스플레이(1350) 상에, 가상 공간을 나타내는 화면을 표시하는 동안, 촬영 카메라(1360-4)를 통해 획득된 이미지를, 상기 화면의 적어도 일부분에 중첩하여 표시하기 위한 패스-쓰루(pass through) 기능을 수행할 수 있다. 일 실시예에서, 촬영 카메라(1360-4)는, 제1 림(1302-1) 및 제2 림(1302-2) 사이에 배치되는 브릿지(1303) 상에 배치될 수 있다.
시선 추적 카메라(1360-1)는, 웨어러블 장치(1301)를 착용한 사용자의 시선(gaze)을 추적함으로써, 사용자의 시선과 적어도 하나의 디스플레이(1350)에 제공되는 시각 정보를 일치시켜 보다 현실적인 증강 현실을 구현할 수 있다. 예를 들어, 웨어러블 장치(1301)는, 사용자가 정면을 바라볼 때, 사용자가 위치한 장소에서 사용자의 정면에 관련된 환경 정보를 자연스럽게 적어도 하나의 디스플레이(1350)에 표시할 수 있다. 시선 추적 카메라(1360-1)는, 사용자의 시선을 결정하기 위하여, 사용자의 동공의 이미지를 캡쳐 하도록, 구성될 수 있다. 예를 들면, 시선 추적 카메라(1360-1)는, 사용자의 동공에서 반사된 시선 검출 광을 수신하고, 수신된 시선 검출 광의 위치 및 움직임에 기반하여, 사용자의 시선을 추적할 수 있다. 일 실시예에서, 시선 추적 카메라(1360-1)는, 사용자의 좌안과 우안에 대응되는 위치에 배치될 수 있다. 예를 들면, 시선 추적 카메라(1360-1)는, 제1 림(1302-1) 및/또는 제2 림(1302-2) 내에서, 웨어러블 장치(1301)를 착용한 사용자가 위치하는 방향을 향하도록 배치될 수 있다.
동작 인식 카메라(1360-2, 1360-3)는, 사용자의 몸통, 손, 또는 얼굴 등 사용자의 신체 전체 또는 일부의 움직임을 인식함으로써, 적어도 하나의 디스플레이(1350)에 제공되는 화면에 특정 이벤트를 제공할 수 있다. 동작 인식 카메라(1360-2, 1360-3)는, 사용자의 동작을 인식(gesture recognition)하여 상기 동작에 대응되는 신호를 획득하고, 상기 신호에 대응되는 표시를 적어도 하나의 디스플레이(1350)에 제공할 수 있다. 프로세서는, 상기 동작에 대응되는 신호를 식별하고, 상기 식별에 기반하여, 지정된 기능을 수행할 수 있다. 동작 인식 카메라(1360-2, 1360-3)는, 6 자유도 자세(6 degrees of freedom pose, 6 dof pose)를 위한 SLAM 및/또는 깊이 맵을 이용한 공간 인식 기능을 수행하기 위해 이용될 수 있다. 프로세서는 동작 인식 카메라(1360-2, 1360-3)을 이용하여, 제스처 인식 기능 및/또는 객체 추적(object tracking) 기능을 수행할 수 있다. 일 실시예에서, 동작 인식 카메라(1360-2, 1360-3)는, 제1 림(1302-1) 및/또는 제2 림(1302-2)상에 배치될 수 있다.
웨어러블 장치(1301) 내에 포함된 카메라(1360)는, 상술된 시선 추적 카메라(1360-1), 동작 인식 카메라(1360-2, 1360-3)에 제한되지 않는다. 예를 들어, 웨어러블 장치(1301)는 사용자의 FoV(field of view)를 향하여 배치된 카메라를 이용하여, 상기 FoV 내에 포함된 외부 객체를 식별할 수 있다. 웨어러블 장치(1301)가 외부 객체를 식별하는 것은, 깊이 센서, 및/또는 ToF(time of flight) 센서와 같이, 웨어러블 장치(1301), 및 외부 객체 사이의 거리를 식별하기 위한 센서에 기반하여 수행될 수 있다. 상기 FoV를 향하여 배치된 상기 카메라(1360)는, 오토포커스(AF) 기능, 및/또는 OIS(optical image stabilization) 기능을 지원할 수 있다. 예를 들어, 웨어러블 장치(1301)는, 웨어러블 장치(1301)를 착용한 사용자의 얼굴을 포함하는 이미지를 획득하기 위하여, 상기 얼굴을 향하여 배치된 카메라(1360)(예: FT(face tracking) 카메라)를 포함할 수 있다.
비록 도시되지 않았지만, 일 실시예에 따른, 웨어러블 장치(1301)는, 카메라(1360)를 이용하여 촬영되는 피사체(예: 사용자의 눈, 얼굴, 및/또는 FoV 내 외부 객체)를 향하여 빛을 방사하는 광원(예: LED)을 더 포함할 수 있다. 상기 광원은 적외선 파장의 LED를 포함할 수 있다. 상기 광원은, 프레임(1300), 힌지 유닛들(1306, 1307) 중 적어도 하나에 배치될 수 있다.
일 실시예에 따르면, 배터리 모듈(1370)은, 웨어러블 장치(1301)의 전자 부품들에 전력을 공급할 수 있다. 일 실시예에서, 배터리 모듈(1370)은, 제1 템플(1304) 및/또는 제2 템플(1305) 내에 배치될 수 있다. 예를 들면, 배터리 모듈(1370)은, 복수의 배터리 모듈(1370)들일 수 있다. 복수의 배터리 모듈(1370)들은, 각각 제1 템플(1304)과 제2 템플(1305) 각각에 배치될 수 있다. 일 실시예에서, 배터리 모듈(1370)은 제1 템플(1304) 및/또는 제2 템플(1305)의 단부에 배치될 수 있다.
안테나 모듈(1375)은, 신호 또는 전력을 웨어러블 장치(1301)의 외부로 송신하거나, 외부로부터 신호 또는 전력을 수신할 수 있다. 일 실시예에서, 안테나 모듈(1375)은, 제1 템플(1304) 및/또는 제2 템플(1305) 내에 배치될 수 있다. 예를 들면, 안테나 모듈(1375)은, 제1 템플(1304), 및/또는 제2 템플(1305)의 일면에 가깝게 배치될 수 있다.
스피커(1355)는, 음향 신호를 웨어러블 장치(1301)의 외부로 출력할 수 있다. 음향 출력 모듈은, 스피커로 참조될 수 있다. 일 실시예에서, 스피커(1355)는, 웨어러블 장치(1301)를 착용한 사용자의 귀에 인접하게 배치되기 위하여, 제1 템플(1304), 및/또는 제2 템플(1305) 내에 배치될 수 있다. 예를 들면, 스피커(1355)는, 제1 템플(1304) 내에 배치됨으로써 사용자의 좌측 귀에 인접하게 배치되는, 제2 스피커(1355-2), 및 제2 템플(1305) 내에 배치됨으로써 사용자의 우측 귀에 인접하게 배치되는, 제1 스피커(1355-1)를 포함할 수 있다. 예를 들면, 스피커(1355)는 도 3a의 스피커(312)의 일 예일 수 있다.
발광 모듈(미도시)은, 적어도 하나의 발광 소자를 포함할 수 있다. 발광 모듈은, 웨어러블 장치(1301)의 특정 상태에 관한 정보를 사용자에게 시각적으로 제공하기 위하여, 특정 상태에 대응되는 색상의 빛을 방출하거나, 특정 상태에 대응되는 동작으로 빛을 방출할 수 있다. 예를 들면, 웨어러블 장치(1301)가, 충전이 필요한 경우, 적색 광의 빛을 일정한 주기로 방출할 수 있다. 일 실시예에서, 발광 모듈은, 제1 림(1302-1) 및/또는 제2 림(1302-2) 상에 배치될 수 있다.
도 13b를 참고하면, 일 실시예에 따른, 웨어러블 장치(1301)는 PCB(printed circuit board)(1390)을 포함할 수 있다. PCB(1390)는, 제1 템플(1304), 또는 제2 템플(1305) 중 적어도 하나에 포함될 수 있다. PCB(1390)는, 적어도 두 개의 서브 PCB들 사이에 배치된 인터포저를 포함할 수 있다. PCB(1390) 상에서, 웨어러블 장치(1301)에 포함된 하나 이상의 하드웨어들(예: 도 16의 상이한 블록들에 의하여 도시된 하드웨어들)이 배치될 수 있다. 웨어러블 장치(1301)는, 상기 하드웨어들을 상호연결하기 위한, FPCB(flexible PCB)를 포함할 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)는, 웨어러블 장치(1301)의 자세, 및/또는 웨어러블 장치(1301)를 착용한 사용자의 신체 부위(예: 머리)의 자세를 탐지하기 위한 자이로 센서, 중력 센서, 및/또는 가속도 센서 중 적어도 하나를 포함할 수 있다. 중력 센서, 및 가속도 센서 각각은, 서로 수직인 지정된 3차원 축들(예: x축, y축 및 z축)에 기반하여 중력 가속도, 및/또는 가속도를 측정할 수 있다. 자이로 센서는 지정된 3차원 축들(예: x축, y축 및 z축) 각각의 각속도를 측정할 수 있다. 상기 중력 센서, 상기 가속도 센서, 및 상기 자이로 센서 중 적어도 하나가, IMU(inertial measurement unit)로 참조될 수 있다. 일 실시예에 따른, 웨어러블 장치(1301)는 IMU에 기반하여 웨어러블 장치(1301)의 특정 기능을 실행하거나, 또는 중단하기 위해 수행된 사용자의 모션, 및/또는 제스처를 식별할 수 있다.
도 13c를 참고하면, 웨어러블 장치(1301)의 일 실시예가 도시된다. 제한되지 않는 예로, 도 13c의 웨어러블 장치(1301)는 안경의 형태를 가질 수 있다. 도 13a 내지 도 13b에서 예시된 웨어러블 장치(1301)의 하드웨어들(또는 구성요소들)의 적어도 일부가 도 13c의 웨어러블 장치(1301)에 적용되거나 포함될 수 있다. 이에 따라, 중복되는 내용은 생략될 수 있다.
웨어러블 장치(1301)는, 제1 림(1302-1) 및/또는 제2 림(1302-2)을 포함할 수 있다. 웨어러블 장치(1301)는 디스플레이(1350)를 포함할 수 있다. 제1 디스플레이(1350-1)는 제1 림(1302-1)에 배치될 수 있다. 예를 들면, 제1 디스플레이(1350-1)는 제1 림(1302-1) 내에서, 웨어러블 장치(1301)를 착용한 사용자의 얼굴을 향하도록 배치될 수 있다. 예를 들면, 제1 디스플레이(1350-1)는 제1 림(1302-1) 내에서, 웨어러블 장치(1301)를 착용한 사용자의 눈을 향하도록 배치될 수 있다. 예를 들면, 제1 디스플레이(1350-1)는 사용자의 좌안에 대응되는 위치에 배치될 수 있다. 예를 들면, 제1 디스플레이(1350-1)는 제1 림(1302-1) 내에서 상단부에 배치될 수 있다. 예를 들면, 제1 디스플레이(1350-1)는 제1 화면(1393-1)을 표시할 수 있다. 예를 들면, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자는 제1 화면(1393-1)을 볼 수 있다. 예를 들면, 사용자는, 시야의 상단부를 바라보는 동안, 제1 화면(1393-1)을 볼 수 있다.
제2 디스플레이(1350-2)는 제2 림(1302-2)에 배치될 수 있다. 예를 들면, 제2 디스플레이(1350-2)는 제2 림(1302-2) 내에서, 웨어러블 장치(1301)를 착용한 사용자의 얼굴을 향하도록 배치될 수 있다. 예를 들면, 제2 디스플레이(1350-2)는 제2 림(1302-2) 내에서, 웨어러블 장치(1301)를 착용한 사용자의 눈을 향하도록 배치될 수 있다. 예를 들면, 제2 디스플레이(1350-2)는 사용자의 우안에 대응되는 위치에 배치될 수 있다. 예를 들면, 제2 디스플레이(1350-2)는 제2 림(1302-2) 내에서 상단부에 배치될 수 있다. 예를 들면, 제2 디스플레이(1350-2)는 제2 화면(1393-2)을 표시할 수 있다. 예를 들면, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자는 제2 화면(1393-2)을 볼 수 있다. 예를 들면, 사용자는, 시야의 상단부를 바라보는 동안, 제2 화면(1393-2)을 볼 수 있다.
웨어러블 장치(1301)는 노즈 패드(1310)를 포함할 수 있다. 예를 들면, 노즈 패드(1310)는, 브릿지(1303), 제1 패드(1311), 및/또는 제2 패드(1312)를 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 제1 패드(1311) 및 제2 패드(1312)는 사용자의 코의 일부분과 접촉될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제1 디스플레이(1350-1)가 사용자의 얼굴을 향하도록 제1 림(1350-1) 상에 배치될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제1 디스플레이(1350-1)가 사용자의 시야에 포함될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제1 화면(1393-1)이 사용자의 좌안의 시야에 포함될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제2 디스플레이(1350-2)가 사용자의 얼굴을 향하도록 제2 림(1350-2) 상에 배치될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제2 디스플레이(1350-2)가 사용자의 시야에 포함될 수 있다. 예를 들면, 제1 패드(1311) 및 제2 패드(1312)가 사용자의 코의 일부분과 접촉되는 동안, 제2 화면(1393-2)이 사용자의 우안의 시야에 포함될 수 있다.
도 14a 내지 도 14b는 웨어러블 장치의 외관의 일 예를 도시한다. 도 14a 내지 도 14b의 웨어러블 장치(1301)는, 도 3a의 전자 장치(301)의 일 예일 수 있다. 일 실시예에 따른, 웨어러블 장치(1301)의 하우징의 제1 면(1410)의 외관의 일 예가 도 14a에 도시되고, 상기 제1 면(1410)의 반대되는(opposite to) 제2 면(1420)의 외관의 일 예가 도 14b에 도시될 수 있다.
도 14a를 참고하면, 일 실시예에 따른, 웨어러블 장치(1301)의 제1 면(1410)은, 사용자의 신체 부위(예: 상기 사용자의 얼굴) 상에 부착가능한(attachable) 형태를 가질 수 있다. 비록 도시되지 않았지만, 웨어러블 장치(1301)는, 사용자의 신체 부위 상에 고정되기 위한 스트랩, 및/또는 하나 이상의 템플들(예: 도 13a 내지 도 13c의 제1 템플(1304), 및/또는 제2 템플(1305))을 더 포함할 수 있다. 사용자의 양 눈들 중에서 좌측 눈으로 이미지를 출력하기 위한 제1 디스플레이(1350-1), 및 상기 양 눈들 중에서 우측 눈으로 이미지를 출력하기 위한 제2 디스플레이(1350-2)가 제1 면(1410) 상에 배치될 수 있다. 웨어러블 장치(1301)는 제1 면(1410) 상에 형성되고, 상기 제1 디스플레이(1350-1), 및 상기 제2 디스플레이(1350-2)로부터 방사되는 광과 상이한 광(예: 외부 광(ambient light))에 의한 간섭을 방지하기 위한, 고무, 또는 실리콘 패킹(packing)을 더 포함할 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)는, 상기 제1 디스플레이(1350-1), 및 상기 제2 디스플레이(1350-2) 각각에 인접한 사용자의 양 눈들을 촬영, 및/또는 추적하기 위한 카메라들(1360-1)을 포함할 수 있다. 상기 카메라들(1360-1)은, 도 13b의 시선 추적 카메라(1360-1)에 참조될 수 있다. 일 실시예에 따른, 웨어러블 장치(1301)는, 사용자의 얼굴을 촬영, 및/또는 인식하기 위한 카메라들(1360-5, 1360-6)을 포함할 수 있다. 상기 카메라들(1360-5, 1360-6)은, FT 카메라로 참조될 수 있다. 웨어러블 장치(1301)는 카메라들(1360-5, 1360-6)을 이용하여 식별된 사용자의 얼굴의 움직임(motion)에 기반하여, 가상 공간 내 상기 사용자를 표현하는 아바타를 제어할 수 있다. 예를 들어, 웨어러블 장치(1301)는, 카메라들(1360-5, 1360-6)(예: FT 카메라)에 의해 획득되고 웨어러블 장치(1301)를 착용한 사용자의 얼굴 표정을 나타내는 정보를 이용하여, 아바타의 일부분(예: 사람의 얼굴을 표현하는 아바타의 일부분)의 텍스쳐 및/또는 형태를 변경할 수 있다.
도 14b를 참고하면, 도 14a의 제1 면(1410)과 반대되는 제2 면(1420) 상에, 웨어러블 장치(1301)의 외부 환경과 관련된 정보를 획득하기 위한 카메라(예: 카메라들(1360-7, 1360-8, 1360-9, 1360-10, 1360-11, 1360-12)), 및/또는 센서(예: 깊이 센서(1430))가 배치될 수 있다. 예를 들어, 카메라들(1360-7, 1360-8, 1360-9, 1360-10)은, 외부 객체를 인식하기 위하여, 제2 면(1420) 상에 배치될 수 있다. 카메라들(1360-7, 1360-8, 1360-9, 1360-10)은, 도 13b의 동작 인식 카메라(1360-2, 1360-3)에 참조될 수 있다.
카메라들(1360-11, 1360-12)을 이용하여, 웨어러블 장치(1301)는 사용자의 양 눈들 각각으로 송신될 이미지, 및/또는 비디오를 획득할 수 있다. 카메라(1360-11)는, 상기 양 눈들 중에서 우측 눈에 대응하는 제2 디스플레이(1350-2)를 통해 표시될 이미지를 획득하도록, 웨어러블 장치(1301)의 제2 면(1420) 상에 배치될 수 있다. 카메라(1360-12)는, 상기 양 눈들 중에서 좌측 눈에 대응하는 제1 디스플레이(1350-1)를 통해 표시될 이미지를 획득하도록, 웨어러블 장치(1301)의 제2 면(1420) 상에 배치될 수 있다. 카메라들(1360-11, 1360-12)은 도 13b의 촬영 카메라(1360-4)에 참조될 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)는, 웨어러블 장치(1301), 및 외부 객체 사이의 거리를 식별하기 위하여 제2 면(1420) 상에 배치된 깊이 센서(1430)를 포함할 수 있다. 깊이 센서(1430)를 이용하여, 웨어러블 장치(1301)는, 웨어러블 장치(1301)를 착용한 사용자의 FoV(field of view)의 적어도 일부분에 대한 공간 정보(spatial information)(예: 깊이 맵(depth map))를 획득할 수 있다. 비록 도시되지 않았지만, 웨어러블 장치(1301)의 제2 면(1420) 상에, 외부 객체로부터 출력된 소리를 획득하기 위한 마이크가 배치될 수 있다. 마이크의 개수는, 실시예에 따라 하나 이상일 수 있다.
도 15는, 웨어러블 장치의 외관의 일 예를 도시한다. 웨어러블 장치(1301)는, 도 3a의 전자 장치(301)의 일 예일 수 있다. 예를 들면, 도 15에서 도시되는 웨어러블 장치(1301)는, 도 13a 내지 도 14b에서 도시된 웨어러블 장치(1301)의 일 실시예로 이해될 수 있다.
도 15를 참고하면, 웨어러블 장치(1301)는, 헤드셋 형태, 및/또는 헤드 기어 형태를 가질 수 있다. 예를 들면, 웨어러블 장치(1301)는, 스피커(1355), 마이크로폰(1365), 스트랩(1530), 이어패드(1540), 및/또는 카메라(1560)를 포함할 수 있다. 예를 들면, 스피커(1355)를 위하여 도 13a 및 도 13b의 스피커(1355)에 대한 설명들이 참조될 수 있다. 예를 들면, 스피커(1355)는, 도 3a의 스피커(312)의 일 예일 수 있다. 제1 스피커(1355-1) 및 제2 스피커(1355-2) 각각은, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자의 귀를 향하도록 배치될 수 있다. 예를 들면, 마이크로폰(1365)을 위하여 도 13a 및 도 13b의 마이크로폰(1365)에 대한 설명들이 참조될 수 있다. 예를 들면, 마이크로폰(1365)은, 도 3a의 마이크로폰(330)의 일 예일 수 있다. 예를 들면, 마이크로폰(1365)은 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자의 입에 인접하도록 배치될 수 있다. 도 13c에서 마이크로폰(1365)이 제1 카메라(1560-1) 아래에 배치되는 것으로 도시되었으나, 실시예가 제한되는 것은 아니다. 예를 들면, 마이크로폰(1365)은 제1 카메라(1560-1) 위에 배치되거나 및/또는 제1 카메라(1560-1) 옆에 배치될 수 있다. 제한되지 않는 예로, 마이크로폰(1365)은, 제2 카메라(1560-2) 주위에 배치될 수 있다.
스트랩(1530)은 사용자의 머리에 부착되기 위한 액세서리로 참조될 수 있다. 예를 들면, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 스트랩(1530)이 사용자의 머리에 접할 수 있다. 이어패드(1540)는, 사용자의 귀에 접촉되는 부재로 참조될 수 있다. 이어패드(1540)는 웨어러블 장치(1301)의 착용감을 향상시키기 위해 이용될 수 있다. 이어패드(1540)는 폼 재질, 메모리폼, 패브릭, 및/또는 인조가죽의 소재로 구성될 수 있다. 이어패드(1540)는, 스피커(1355)에 의해 출력되는 오디오 신호의 누출을 방지하기 위해 이용될 수 있다. 제1 이어패드(1540-1)는, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 제1 스피커(1355-1)와 사용자의 귀 사이에 위치될 수 있다. 제2 이어패드(1540-2)는, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 제2 스피커(1355-2)와 사용자의 귀 사이에 위치될 수 있다.
카메라(1560)는, 이미지를 획득하기 위해 이용될 수 있다. 카메라(1560)는, 도 13a 내지 도 14b의 카메라(1360)의 일 예일 수 있다. 예를 들면, 카메라(1560)를 위해 도 13a 내지 도 14b의 카메라(1360)에 대한 설명들이 참조될 수 있다. 예를 들면, 카메라(1560)는 외부 환경을 나타내는 이미지를 획득하기 위해 이용될 수 있다. 예를 들면, 웨어러블 장치(1301)는, 카메라(1560)를 통해 획득된 이미지 내에 포함된 외부 객체를 식별할 수 있다. 제1 카메라(1560-1)는, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자의 얼굴이 향하는 방향을 향하도록 배치될 수 있다. 제1 카메라(1560-1)는, 제1 스피커(1255-1)를 포함하는 웨어러블 장치(1301)의 하우징 상에 배치될 수 있다. 제2 카메라(1560-2)는, 웨어러블 장치(1301)가 사용자에 의해 착용되는 동안, 사용자의 얼굴이 향하는 방향을 향하도록 배치될 수 있다. 제2 카메라(1560-2)는, 제2 스피커(1255-2)를 포함하는 웨어러블 장치(1301)의 하우징 상에 배치될 수 있다.
이하, 도 16를 참조하여, 웨어러블 장치(1301)의 하드웨어 또는 소프트웨어 구성이 후술된다.
도 16은, 웨어러블 장치의 블록도의 일 예를 도시한다. 도 16의 웨어러블 장치(1301)는, 도 3a의 전자 장치(301)의 일 예일 수 있다. 웨어러블 장치(1301)는 도 13a 내지 도 15의 웨어러블 장치(1301)의 일 예일 수 있다.
도 16을 참고하면, 일 실시예에 따른 웨어러블 장치(1301)는 프로세서(1610), 메모리(1615), 디스플레이(1350)(예: 도 13a, 도 13b, 도 13c, 도 14a, 및 도 14b의 제1 디스플레이(1350-1) 및/또는 제2 디스플레이(1350-2)), 스피커(1355), 마이크로폰(1365), 및/또는 적어도 하나의 센서(1620)를 포함할 수 있다. 프로세서(1610), 메모리(1615), 디스플레이(1350), 스피커(1355), 마이크로폰(1365), 및/또는 적어도 하나의 센서(1620)는 통신 버스(402)와 같은 전자 부품에 의해 서로 전기적으로 및/또는 작동적으로 연결될 수 있다. 본 개시에서, 전자 부품들의 작동적인 연결은, 상기 전자 부품들 중 제1 전자 부품이 상기 전자 부품들 중 제2 전자 부품에 의해 제어되도록, 상기 전자 부품들 사이에 수립된 직접적인 연결 및/또는 상기 전자 부품들 사이에 수립된 간접적인 연결을 포함할 수 있다. 웨어러블 장치(1301)에 포함된 전자 부품의 타입 및/또는 개수는 도 16에 도시된 바에 제한되지 않는다. 예를 들어, 웨어러블 장치(1301)는 도 16에 도시된 전자 부품들 중 일부만 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 도 3a의 통신 회로(310)를 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)는 도 3b의 제1 멀티모달 모델(350), 제3 모델(370), 및/또는 제4 멀티모달 모델(380)을 포함할 수 있다.
일 실시예에 따른 웨어러블 장치(1301)의 프로세서(1610)는 하나 이상의 인스트럭션들에 기반하여 데이터를 처리하기 위한 회로(예: 처리 회로)를 포함할 수 있다. 프로세서(1610)는, 도 3a의 적어도 하나의 프로세서(300)의 일 예일 수 있다. 프로세서(1610)는, 도 3a의 적어도 하나의 프로세서(300)와 실질적으로 동일할 수 있으므로, 중복되는 설명은 생략하도록 한다. 일 실시예에 따르면, 프로세서(1610)의 구조는 본 개시의 일 실시예에 제한되지 않으며, 적어도 하나의 회로는 프로세서의 외부에 물리적으로 분리된 별도의 프로세서로 형성될 수 있다. 멀티-코어 프로세서 구조를 가지는 프로세서(1610)를 포함하는 일 실시예에서, 본 개시의 동작들, 및/또는 기능들은, 프로세서(1610)에 포함된 하나 이상의 코어들에 의해 개별적으로 또는 집합적으로 수행될 수 있다.
일 실시예에 따른 웨어러블 장치(1301)의 메모리(1615)는 프로세서(1610)로 입력되거나, 및/또는 프로세서(1610)로부터 출력되는 데이터 및/또는 인스트럭션들을 저장하기 위한 전자 부품을 포함할 수 있다. 메모리(1615)는, 도 3a의 메모리(320)의 일 예일 수 있다. 프로세서(1610)는, 도 3a의 메모리(320)와 실질적으로 동일할 수 있으므로, 중복되는 설명은 생략하도록 한다. 일 실시예에서, 메모리(1615)는 스토리지로 지칭될 수 있다.
일 실시예에서, 웨어러블 장치(1301)의 디스플레이(1350)는 웨어러블 장치(1301)의 사용자에게 시각화된 정보를 출력할 수 있다. 디스플레이(1350)는, 도 3a의 디스플레이(311)의 일 예일 수 있다. 디스플레이(1350)는, 도 3a의 디스플레이(311)와 실질적으로 동일할 수 있으므로, 중복되는 설명은 생략하도록 한다. 또한, 디스플레이(1350)를 위해, 도 13a 내지 도 14b의 디스플레이(1350)에 대한 설명들이 참조될 수 있다. 웨어러블 장치(1301)를 착용한 사용자의 눈 앞에 배열되는 디스플레이(1350)는, 웨어러블 장치(1301)의 하우징의 적어도 일부에 배치될 수 있다(예: 도 13a, 도 13b, 도 13c, 도 14a, 및 도 14b의 제1 디스플레이(1350-1) 및/또는 제2 디스플레이(1350-2)). 예를 들어, 디스플레이(1350)는, 디스플레이 어셈블리 내에 포함될 수 있다. 예를 들어, 디스플레이(1350)는, 프로세서(1610)에 의해 제어되어, 사용자에게 시각화된 정보(visualized information)를 출력할 수 있다. 디스플레이(1350)는 플렉서블 디스플레이, FPD(flat panel display) 및/또는 전자 종이(electronic paper)를 포함할 수 있다. 디스플레이(1350)는 LCD(liquid crystal display), PDP(plasma display panel) 및/또는 하나 이상의 LED(light emitting diode)를 포함할 수 있다. 상기 LED는 OLED(organic LED)를 포함할 수 있다. 실시예가 이에 제한되는 것은 아니며, 예를 들어, 웨어러블 장치(1301)가 외부 광(external light 또는 ambient light)을 투과하기 위한 렌즈를 포함하는 경우, 디스플레이(1350)는 상기 렌즈로 광을 투사하기 위한(projecting onto) 프로젝터(또는 프로젝션 어셈블리)를 포함할 수 있다. 일 실시예에서, 디스플레이(1350)는, 디스플레이 패널 및/또는 디스플레이 모듈로 지칭될 수 있다. 디스플레이(1350)에 포함된 픽셀들은, 웨어러블 장치(1301)의 사용자에 의해 착용될 시에, 상기 사용자의 두 눈들 중 어느 하나를 향해 배치될 수 있다. 예를 들어, 디스플레이(1350)는 사용자의 두 눈들 각각에 대응하는 표시 영역들(또는 활성 영역들)을 포함할 수 있다.
일 실시예에서, 웨어러블 장치(1301)의 적어도 하나의 센서(1620)는, 웨어러블 장치(1301)와 관련된 비-전기적 정보(non-electronic information)로부터 프로세서(1610) 및/또는 메모리(1615)에 의해 처리될 수 있는 전기적 정보를 생성할 수 있다. 적어도 하나의 센서(1620)는, 도 3a의 적어도 하나의 센서(340)의 일 예일 수 있다. 적어도 하나의 센서(1620)는, 도 3a의 적어도 하나의 센서(340)와 실질적으로 동일할 수 있으므로, 중복되는 설명은 생략하도록 한다. 예를 들어, 적어도 하나의 센서(1620)는 웨어러블 장치(1301)의 지리적 위치(geographic location)를 탐지하기 위한 GPS(global positioning system) 센서를 포함할 수 있다. 상기 GPS 방식 외에도, 적어도 하나의 센서(1620)는, 예를 들어, 갈릴레오(galileo), 베이더우(beidou, compass)와 같은 GNSS(global navigation satellite system)에 기반하여 웨어러블 장치(1301)의 지리적 위치를 나타내는 정보를 생성할 수 있다. 상기 정보는 메모리(1615)에 저장되거나, 프로세서(1610)에 의해 처리되거나, 및/또는 통신 회로를 통해 웨어러블 장치(1301)와 구별되는 다른 전자 장치로 송신될 수 있다.
도 16을 참고하면, 웨어러블 장치(1301)에 포함된 적어도 하나의 센서(1620)의 일 예로, 이미지 센서(1621) 및/또는 모션 센서(1622)가 도시된다. 이미지 센서(1621)는, 빛의 색상 및/또는 밝기를 나타내는 전기 신호를 생성하는 하나 이상의 광 센서들(예: CCD(charged coupled device) 센서, CMOS(complementary metal oxide semiconductor) 센서)을 포함할 수 있다. 이미지 센서(1621)는, 카메라로 지칭될 수 있다. 이미지 센서(1621)는 하나 이상의 카메라들을 포함할 수 있다. 예를 들면, 이미지 센서(1621)는, 도 13a 내지 도 14b의 카메라(1360), 및/또는 도 15의 카메라(1560)를 포함할 수 있다. 이미지 센서(1621)를 위하여, 도 13a 내지 도 14b의 카메라(1360), 및/또는 도 15의 카메라(1560)에 대한 설명들이 참조될 수 있다.
이미지 센서(1621)에 포함된 복수의 광 센서들은 2차원 격자(2 dimensional array)의 형태로 배치될 수 있다. 이미지 센서(1621)는 복수의 광 센서들 각각의 전기 신호들을 실질적으로 동시에 획득하여, 2차원 격자의 광 센서들에 도달한 빛에 대응하는 2차원 프레임 데이터를 생성할 수 있다. 예를 들어, 이미지 센서(1621)를 이용하여 캡쳐한 이미지(또는 사진 데이터)는, 이미지 센서(1621)로부터 획득된 하나의(a) 2차원 프레임 데이터를 의미할 수 있다. 예를 들어, 이미지 센서(1621)를 이용하여 캡쳐한 비디오 데이터는, 프레임 레이트를 따라 이미지 센서(1621)로부터 획득된, 복수의 2차원 프레임 데이터의 시퀀스(sequence)를 의미할 수 있다. 이미지 센서(1621)는, 이미지 센서(1621)가 광을 수신하는 방향을 향하여 배치되고, 상기 방향을 향하여 광을 출력하기 위한 플래시 라이트를 더 포함할 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)는 이미지 센서(1621)의 일 예로, 상이한 방향들을 향하여 배치된 복수의 이미지 센서들을 포함할 수 있다. 도 13a, 도 13b, 도 13c, 도 14a, 도 14b, 도 15를 참고하여 상술한 바와 같이, 상기 복수의 이미지 센서들은, 웨어러블 장치(1301)를 착용한 사용자의 눈을 향하여 배열되도록 구성된, 시선 추적 카메라(예: 도 13b 및 도 14a의 시선 추적 카메라(1360-1))를 포함할 수 있다. 상기 복수의 이미지 센서들은, 외부(outward) 카메라를 포함할 수 있다. 프로세서(1610)는 시선 추적 카메라로부터 획득된 이미지 및/또는 비디오를 이용하여, 사용자의 시선의 방향을 식별할 수 있다. 시선 추적 카메라는 IR(infrared) 센서를 포함할 수 있다. 시선 추적 카메라는, 안구 센서 및/또는 안구 추적기로 지칭될 수 있다.
외부 카메라는, 웨어러블 장치(1301)를 착용한 사용자의 전 방(예: 두 눈들이 향할 수 있는 방향)을 향하여 배치될 수 있다. 웨어러블 장치(1301)는 복수의 외부 카메라들을 포함할 수 있다. 실시예가 이에 제한되는 것은 아니며, 외부 카메라는 외부 공간을 향하여 배치될 수 있다. 외부 카메라로부터 획득된 이미지 및/또는 비디오를 이용하여, 프로세서(1610)는 외부 객체를 식별할 수 있다. 예를 들어, 프로세서(1610)는 외부 카메라로부터 획득된 이미지 및/또는 비디오에 기반하여, 웨어러블 장치(1301)를 착용한 사용자의 손의 위치, 형태 및/또는 제스처(예: 손 제스처)를 식별할 수 있다. 외부 카메라로부터 획득된, 외부 환경에 대한 이미지 및/또는 비디오를 이용하여, 프로세서(1610)는 상기 외부 환경 내 하나 이상의 물체들을 인식할 수 있거나, 또는 추적할 수 있다.
일 실시예에 따른, 모션 센서(1622)는, 서로 수직이고(perpendicular to each other), 웨어러블 장치(1301) 및/또는 모션 센서(1622) 내 지정된 원점에 기반하는, 복수의 축들(예: x축, y축 및 z축)의 중력 가속도들, 가속도들, 및/또는 각속도들을 나타내는 전기 신호를 출력할 수 있다. 예를 들어, 프로세서(1610)는 모션 센서(1622)로부터, 지정된 주기(예: 1 밀리초)에 기반하여, 상기 복수의 축들의 개수의 가속도들, 각속도들, 및/또는 자계의 크기들을 포함하는 센서 데이터를 반복적으로 수신할 수 있거나, 또는 획득할 수 있다. 일 실시예에서, 모션 센서(1622)는 IMU(inertial measurement unit)로 지칭될 수 있다. 웨어러블 장치(1301)에 포함된 적어도 하나의 센서(1620)는 상술한 바에 제한되지 않으며, 그립 센서, 근접 센서, 심박 센서, 지문 센서, 조도 센서 및/또는 ToF 센서를 포함할 수 있다. 모션 센서(1622)를 이용하여, 프로세서(1610)는 웨어러블 장치(1301)의 모션(예: 웨어러블 장치(1301)를 착용한 사용자에 의해 야기되는, 웨어러블 장치(1301)의 모션)을 탐지할 수 있다.
일 실시예에 따른, 웨어러블 장치(1301)의 메모리(1615) 내에서, 웨어러블 장치(1301)의 프로세서(1610)에 의해 처리될 데이터, 수행될 계산 및/또는 동작을 나타내는 하나 이상의 인스트럭션들(또는 명령어들)이 저장될 수 있다. 하나 이상의 인스트럭션들의 집합은, 프로그램, 펌웨어, 운영 체제, 프로세스, 루틴, 서브-루틴 및/또는 소프트웨어 어플리케이션(이하, 어플리케이션)으로 참조될 수 있다. 예를 들어, 웨어러블 장치(1301), 및/또는 프로세서(1610)는, 운영체제, 펌웨어, 드라이버, 프로그램, 및/또는 소프트웨어 어플리케이션 형태로 배포된 복수의 인스트럭션들의 집합(set of a plurality of instructions)이 실행될 시에, 도 3b, 도 4a, 도 4b, 도 5, 도 6a, 도 6b, 도 7, 도 8a, 도 8b, 도 9a, 도 9b, 도 10a, 도 10b, 도 11, 도 17, 도 18, 도 19, 도 20a, 도 20b, 및 도 21의 동작들 중 적어도 하나를 수행할 수 있다. 이하에서, 소프트웨어 어플리케이션이 웨어러블 장치(1301) 내에 설치되었다는 것은, 소프트웨어 어플리케이션(또는 패키지)의 형태로 제공된 하나 이상의 인스트럭션들이 메모리(1615) 내에 저장된 것으로써, 상기 하나 이상의 어플리케이션들이 프로세서(1610)에 의해 실행 가능한(executable) 포맷(예: 웨어러블 장치(1301)의 운영 체제에 의해 지정된 확장자를 가지는 파일)으로 저장된 것을 의미할 수 있다. 일 예로, 어플리케이션은, 사용자에게 제공되는 서비스와 관련된 프로그램, 및/또는 라이브러리를 포함할 수 있다.
도 16를 참고하면, 웨어러블 장치(1301)에 설치된 프로그램들은, 타겟에 기반하여, 어플리케이션 레이어(1640), 프레임워크 레이어(1650) 및/또는 하드웨어 추상화 레이어(hardware abstraction layer, HAL)(1680)를 포함하는 상이한 레이어들 중 어느 한 레이어에 포함될 수 있다. 예를 들어, 하드웨어 추상화 레이어(1680)(예: android system HAL, 및/또는 XR HAL) 내에, 웨어러블 장치(1301)의 하드웨어(예: 디스플레이(1350), 및/또는 적어도 하나의 센서(1620))를 타겟으로 설계된 프로그램들(예: 모듈, 또는 드라이버)이 포함될 수 있다. 프레임워크 레이어(1650)는, XR(extended reality) 서비스를 제공하기 위한 하나 이상의 프로그램들을 포함하는 관점에서, XR 프레임워크 레이어로 지칭될 수 있다. 예를 들어, 도 16에 도시된 레이어들은, 논리적으로(또는 설명의 편의를 위하여) 구분된 것으로 메모리(1615)의 주소 공간이 상기 레이어들에 의해 구분되는 것을 의미하지 않을 수 있다.
프레임워크 레이어(1650) 내에, 하드웨어 추상화 레이어(1680) 및/또는 어플리케이션 레이어(1640) 중 적어도 하나를 타겟으로 설계된 프로그램들(예: 위치 추적기(1671), 공간 인식기(1672), 제스처 추적기(1673), 시선 추적기(1674), 얼굴 추적기(1675), 및/또는 렌더러(1690))이 포함될 수 있다. 프레임워크 레이어(1650)에 포함된 프로그램들은, 다른 프로그램에 기반하여 실행가능한(또는 호출(invoke 또는 call)가능한) API(application programming interface)를 제공할 수 있다.
어플리케이션 레이어(1640) 내에, 웨어러블 장치(1301)의 사용자를 타겟으로 설계된 프로그램이 포함될 수 있다. 어플리케이션 레이어(1640)에 포함된 프로그램들의 일 예로, XR(extended reality) 시스템 UI(user interface)(1641), 및/또는 XR 어플리케이션(1642)이 예시되지만, 실시예가 이에 제한되는 것은 아니다. 예를 들어, 어플리케이션 레이어(1640)에 포함된 프로그램들(예: 소프트웨어 어플리케이션)은, API를 호출하여, 프레임워크 레이어(1650)에 포함된 프로그램들에 의해 지원되는 기능의 실행을 야기할 수 있다.
웨어러블 장치(1301)는 XR 시스템 UI(1641)의 실행에 기반하여, 사용자와 상호작용을 수행하기 위한 하나 이상의 시각적 객체들을 디스플레이(1350) 상에 표시할 수 있다. 시각적 객체는, 텍스트, 이미지, 아이콘, 비디오, 버튼, 체크박스, 라디오버튼, 텍스트 박스, 슬라이더 및/또는 테이블과 같이, 정보의 송신 및/또는 상호작용을 위해 화면 내에 배치될 수 있는 객체를 의미할 수 있다. 시각적 객체는 시각적 가이드, 가상 객체, 시각 요소, UI 요소, 뷰 객체, 및/또는 뷰 요소로 지칭될 수 있다. 웨어러블 장치(1301)는, XR 시스템 UI(1641)의 실행에 기반하여, 사용자에게, 가상 공간 내에서 이용가능한 기능들을 제공할 수 있다.
도 16를 참고하면, XR 시스템 UI(1641) 내에 경량 렌더러(1643), 및/또는 XR 플러그인(1644)이 포함되도록 도시되어 있지만, 이에 제한되는 것은 아니다. 예를 들어, XR 시스템 UI(1641)에 기반하여, 프로세서(1610)는 프레임워크 레이어(1650) 내 경량 렌더러(1643), 및/또는 XR 플러그인(1644)을 실행할 수 있다.
웨어러블 장치(1301)는 경량 렌더러(lightweight renderer)(1643)의 실행에 기반하여, 부분적인 변경이 허용된, 렌더링 파이프라인을 정의, 생성 및/또는 실행하기 위해 이용되는 자원(예: API, 시스템 프로세스 및/또는 라이브러리)을 획득할 수 있다. 경량 렌더러(1643)는 부분적인 변경이 허용된, 렌더링 파이프라인을 정의하는 관점에서, lightweight render pipeline으로 지칭될 수 있다. 경량 렌더러(1643)는, 소프트웨어 어플리케이션의 실행 이전에 빌드된 렌더러(예: 프리-빌드(prebuilt) 렌더러)를 포함할 수 있다. 예를 들어, 웨어러블 장치(1301)는 XR 플러그인(1644)의 실행에 기반하여, 렌더링 파이프라인 전체를 정의, 생성 및/또는 실행하기 위해 이용되는 자원(예: API, 시스템 프로세스 및/또는 라이브러리)을 획득할 수 있다. XR 플러그인(1644)은 렌더링 파이프라인 전체를 정의(또는 설정)하는 관점에서, open XR native client로 지칭될 수 있다.
웨어러블 장치(1301)는, XR 어플리케이션(1642)의 실행에 기반하여, 가상 공간의 적어도 일부분을 나타내는 화면을 디스플레이(1350) 상에 표시할 수 있다. XR 어플리케이션(1642)에 포함된 XR 플러그인(1644-1)은 XR 시스템 UI(1641)의 XR 플러그인(1644)과 유사한 기능을 지원하는 인스트럭션들을 포함할 수 있다. XR 플러그인(1644-1)의 설명 중 XR 플러그인(1644)의 설명과 중복되는 설명은 생략될 수 있다. 웨어러블 장치(1301)는, XR 어플리케이션(1642)의 실행에 기반하여, 가상 공간 매니저(1651)의 실행을 야기할 수 있다.
웨어러블 장치(1301)는, 어플리케이션(16165)의 실행에 기반하여, 가상 공간 상에 이미지를 디스플레이(1350) 상에 표시할 수 있다. 어플리케이션(16165)은 2차원 이미지를 표시하기 위한 이미지 정보를 출력하도록 구성될 수 있다. 웨어러블 장치(1301)는, 어플리케이션(16165)의 실행에 기반하여, 가상 공간 매니저(1651)의 실행을 야기할 수 있다. 웨어러블 장치(1301)는, 어플리케이션(16165)의 실행에 기반하여, 3차원의 가상 공간에서 상기 2차원 이미지를 나타내기 위하여, 이중 이미지 정보를 생성할 수 있다. 여기서, 이중 이미지 정보는 양안 시차를 고려하여, 좌안을 위한 제1 이미지 정보와 우안을 위한 제2 이미지 정보를 포함할 수 있다. 상기 2차원 이미지를 3차원의 가상 공간에서 나타내기 위해, 웨어러블 장치(1301)는 상기 2차원 이미지를 표시하기 위한 이미지 정보에 기반하여, 상기 이중 이미지 정보를 생성할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 가상 공간 매니저(1651)의 실행에 기반하여 가상 공간 서비스를 제공할 수 있다. 예를 들어, 가상 공간 매니저(1651)는 가상 공간 서비스를 지원하기 위한 플랫폼을 포함할 수 있다. 웨어러블 장치(1301)는 가상 공간 매니저(1651)의 실행에 기반하여, 적어도 하나의 센서(1620)를 통해 획득된 데이터에 의해 나타나는 사용자의 위치를 기준으로 형성된 가상 공간을 식별할 수 있고, 상기 가상 공간의 적어도 일부분을 디스플레이(1350) 상에 표시할 수 있다. 가상 공간 매니저(1651)는, CPM(composition presentation manager)로 참조될 수 있다.
가상 공간 매니저(1651)는, 런타임 서비스(1652)를 포함할 수 있다. 일 예로, 런타임 서비스(1652)는 OpenXR runtime 모듈(또는 OpenXR runtime 프로그램)로 참조될 수 있다. 웨어러블 장치(1301)는 런타임 서비스(1652)의 실행에 기반하여, 사용자의 포즈 예측 기능, 프레임 타이밍 기능, 및/또는 공간 입력 기능 중 적어도 하나를 실행할 수 있다. 일 예로, 웨어러블 장치(1301)는 런타임 서비스(1652)의 실행에 기반하여 사용자에게 가상 공간 서비스를 위한 렌더링을 수행할 수 있다. 예를 들어, 런타임 서비스(1652)의 실행에 기반하여, 어플리케이션 레이어(1640)에 의해 실행 가능한, 가상 공간과 관련된 기능이 지원될 수 있다.
가상 공간 매니저(1651)는 pass-through 매니저(1653)룰 포함할 수 있다. 웨어러블 장치(1301)는, pass-through 매니저(1653)의 실행에 기반하여, 가상 공간을 나타내는 화면을 디스플레이(1350)상에 표시하는 동안, 외부 카메라를 통해 획득된 실제 공간을 나타내는 이미지 및/또는 비디오를 상기 화면의 적어도 일부분에 중첩하여 표시할 수 있다.
가상 공간 매니저(1651)는, 입력 매니저(1654)를 포함할 수 있다. 웨어러블 장치(1301)는 입력 매니저(1654)의 실행에 기반하여, 인식 서비스 레이어(1670) 내에 포함된 하나 이상의 프로그램들을 실행하여 획득된 데이터(예: 센서 데이터)를 식별할 수 있다. 웨어러블 장치(1301)는 상기 획득된 데이터를 이용하여, 웨어러블 장치(1301)와 관련된 사용자 입력을 식별할 수 있다. 상기 사용자 입력은, 적어도 하나의 센서(1620) (예: 외부 카메라와 같은 이미지 센서(1621))에 의해 식별된 사용자의 모션(예: 손 제스처), 시선 및/또는 발언과 관련될 수 있다. 상기 사용자 입력은, 통신 회로를 통해 연결된(또는 페어링된) 외부 전자 장치에 기반하여 식별될 수 있다.
인식 추상화 레이어(1660)(perception abstract layer)는 가상 공간 매니저(1651)와 인식 서비스 레이어(perception service layer)(1670) 사이의 데이터 교환을 위해 사용될 수 있다. 가상 공간 매니저(1651)와 인식 서비스 레이어(perception service layer)(1670) 사이의 데이터 교환을 위해 사용되는 관점에서, 인식 추상화 레이어(1660)(perception abstract layer)는 인터페이스(interface)로 지칭될 수 있다. 일 예로, 인식 추상화 레이어(1660)는 OpenPX로 참조될 수 있다. 인식 추상화 레이어(1660)는 perception client 및 perception service를 위해 사용될 수 있다.
일 실시예에 따르면, 인식 서비스 레이어(1670)는 적어도 하나의 센서(1620)로부터 획득된 데이터를 처리하기 위한 하나 이상의 프로그램들을 포함할 수 있다. 하나 이상의 프로그램들은 위치 추적기(1671), 공간 인식기(1672), 제스처 추적기(1673), 시선 추적기(1674), 얼굴 추적기(1675), 및/또는 렌더러(1690) 중 적어도 하나를 포함할 수 있다. 인식 서비스 레이어(1670)에 포함된 하나 이상의 프로그램들의 타입 및/또는 개수는 도 16에 도시된 바에 제한되지 않는다.
웨어러블 장치(1301)는, 위치 추적기(1671)의 실행에 기반하여, 적어도 하나의 센서(1620)를 이용하여, 웨어러블 장치(1301)의 자세를 식별할 수 있다. 웨어러블 장치(1301)는 위치 추적기(1671)의 실행에 기반하여, 외부 카메라(예: 이미지 센서(1621)) 및/또는 IMU(예: 자이로 센서, 가속도 센서 및/또는 지자기 센서를 포함하는 모션 센서(1622))를 이용하여 획득된 데이터를 이용하여, 웨어러블 장치(1301)의 6 자유도 자세(6 degrees of freedom pose, 6 dof pose)를 식별할 수 있다. 위치 추적기(1671)는, 헤드 트래킹(head tracking, HeT) 모듈(또는 헤드 트래커, 헤드 트래킹 프로그램)로 참조될 수 있다.
웨어러블 장치(1301)는 공간 인식기(1672)의 실행에 기반하여, 웨어러블 장치(1301)(또는, 웨어러블 장치(1301)의 사용자)의 주변 환경(예: 외부 공간)에 대응하는 3 차원의 가상 공간을 제공하기 위한 정보를 획득할 수 있다. 웨어러블 장치(1301)는 공간 인식기(1672)의 실행에 기반하여, 외부 카메라(예: 이미지 센서(1621))를 이용하여 획득된 데이터를 이용하여, 웨어러블 장치(1301)의 주변 환경을 3 차원으로 재현할 수 있다. 웨어러블 장치(1301)는 공간 인식기(1672)의 실행에 기반하여 3 차원으로 재현된 웨어러블 장치(1301)의 주변 환경에 기반하여, 평면, 경사, 계단 중 적어도 하나를 식별할 수 있다. 공간 인식기(1672)는, 장면 인식(scene understanding, SU) 모듈(또는 장면 인식 프로그램)로 참조될 수 있다.
웨어러블 장치(1301)는 제스처 추적기(1673)의 실행에 기반하여, 웨어러블 장치(1301)의 사용자의 손의 포즈 및/또는 제스처를 식별(또는 인식)할 수 있다. 일 예로, 웨어러블 장치(1301)는, 제스처 추적기(1673)의 실행에 기반하여, 외부 카메라(예: 이미지 센서(1621))로부터 획득된 데이터를 이용하여, 사용자의 손의 포즈 및/또는 제스처를 식별할 수 있다. 일 예로, 웨어러블 장치(1301)는 제스처 추적기(1673)의 실행에 기반하여, 외부 카메라를 이용하여 획득된 데이터(또는, 이미지)에 기반하여, 사용자의 손의 포즈 및/또는 제스처를 식별할 수 있다. 제스처 추적기(1673)는, 핸드 트래킹(hand tracking, HaT) 모듈(또는 핸드 트래킹 프로그램), 및/또는 제스처 트레킹 모듈로 참조될 수 있다.
웨어러블 장치(1301)는 시선 추적기(1674)의 실행에 기반하여, 웨어러블 장치(1301)의 사용자의 눈의 움직임을 식별(또는 추적(tracking))할 수 있다. 일 예로, 웨어러블 장치(1301)는 시선 추적기(1674)의 실행에 기반하여 시선 추적 카메라(예: 이미지 센서(1621))로부터 획득된 데이터를 이용하여, 사용자의 눈의 움직임을 식별할 수 있다. 시선 추적기(1674)는, 아이 트래킹(eye tracking, ET) 모듈(또는 아이 트래킹 프로그램), 및/또는 시선(gaze) 트래킹 모듈로 참조될 수 있다.
웨어러블 장치(1301)의 인식 서비스 레이어(1670)는, 사용자의 얼굴을 추적하기 위한 얼굴 추적기(1675)를 더 포함할 수 있다. 예를 들어, 웨어러블 장치(1301)는 얼굴 추적기(1675)의 실행에 기반하여, 사용자의 얼굴의 움직임 및/또는 사용자의 표정을 식별(또는 추적)할 수 있다. 웨어러블 장치(1301)는 얼굴 추적기(1675)의 실행에 기반하여, 사용자의 얼굴의 움직임에 기반하여, 사용자의 표정을 추정할 수 있다. 일 예로, 웨어러블 장치(1301)는, 얼굴 추적기(1675)의 실행에 기반하여, FT 카메라(예, 사용자의 얼굴의 적어도 일부분을 향하는 카메라, 이미지 센서(1621))를 이용하여 획득된 데이터(예, 이미지 및/또는 비디오)에 기반하여, 사용자의 얼굴의 움직임 및/또는 사용자의 표정을 식별할 수 있다. 얼굴 추적기(1675)는 페이스 트래킹(FT(face tracking))(또는 페이스 트래킹 프로그램), 및/또는 페이스 트래킹 모듈로 참조될 수 있다.
렌더러(1690)는, 3차원 가상 공간에서 이미지들을 렌더링하기 위한 인스트럭션들을 포함할 수 있다. 렌더러(1690)를 실행한 프로세서(1610)(예: DPU))는, 소프트웨어 어플리케이션(예: CPU 및/또는 GPU에 의해 실행되는 소프트웨어 어플리케이션)에서 디스플레이(1350)의 표시 영역에 적어도 부분적으로 표시될, 적어도 하나의 이미지를 획득할 수 있다. 예를 들어, 렌더러(1690)를 실행한 프로세서(1610)는, 어플리케이션(예: XR 어플리케이션(1642), 어플리케이션(16165))이 렌더링 될 영역의 위치를 결정할 수 있다. 렌더러(1690)를 실행한 프로세서(1610)는 디스플레이(1350) 상에 표시될 상기 어플리케이션의 이미지를 생성할 수 있다. 렌더러(1690)는 이미지들을 합성하여, 디스플레이(1350) 상에 표시될 합성 이미지를 생성할 수 있다.
렌더러(1690)를 실행한 프로세서(1610)는, 위치 추적기(1671) 및/또는 시선 추적기(1674)를 이용하여 계산된, 시선 위치를 이용하여, 디스플레이(1350)의 표시 영역을, 포비티드(foveated) 부분(혹은 포비티드 영역으로 지칭될 수 있음) 및 주변 부분(혹은 잔여 영역으로 지칭될 수 있음)으로 구분할 수 있다. 예를 들어, 시선 위치의 좌표 값들을 탐지한 프로세서(1610)는, 상기 좌표 값들을 포함하는 표시 영역의 부분을, 포비티드 영역으로 결정할 수 있다. 렌더러(1690)를 실행한 DPU는, 상기 포비티드 영역 및 상기 잔여 영역 각각에 대응하고, 디스플레이(1350)의 전체 표시 영역의 사이즈 보다 작은 사이즈를 가지거나, 표시 영역의 해상도 보다 적은 해상도를 가지는, 적어도 하나의 이미지를 획득할 수 있다.
렌더러(1690)를 실행한 프로세서(1610)는, 포비티드 영역에 대응하는 이미지 및 주변 부분에 대응하는 이미지를 합성하여, 디스플레이(1350) 상에 표시될 합성 이미지를 획득할 수 있거나, 또는 생성할 수 있다. 예를 들어, 프로세서(1610)는 업스케일링을 수행하여, 디스플레이(1350)의 전체 표시 영역의 사이즈로, 주변 부분에 대응하는 이미지를 확대할 수 있다. 확대된 이미지 상에, 프로세서(1610)는 포비티드 영역에 대응하는 이미지를 결합하여, 디스플레이(1350) 상에 표시될 합성 이미지를 생성할 수 있다. 포비티드 영역에 대응하는 이미지의 경계 선을 따라, 프로세서(1610)는 블러와 같은 시각 효과를 적용하여, 확대된 이미지 및 포비티드 영역에 대응하는 이미지를 혼합할 수 있다.
도 17은 가상 공간에서 이미지를 표시하기 위한 웨어러블 장치의 블록도의 예를 나타낸다. 도 17의 웨어러블 장치(예: 웨어러블 장치(1301))는 도 3a의 전자 장치(301)의 일 예일 수 있다. 도 17에서는, 가상 공간에서 이미지를 표시하기 위한 다수의 프로그램들(또는 인스트럭션들)이 실행되는 예가 서술된다. 상기 다수의 프로그램들(또는 인스트럭션들)은 하나의 프로세서(예: AP)에서 모두 실행되거나, 다수의 프로세서들(예: AP, GPU(graphic processing unit), NPU(neural processing unit))에 의해 실행될 수 있다. 상기 다수의 프로세서들에 의해 실행될 수 있다는 것의 의미는, 일부 프로그램(또는 인스트럭션)은 제1 프로세서에 의해 실행되고 다른 일부 프로그램(또는 인스트럭션)은 상기 제1 프로세서와 다른 제2 프로세서에 의해 실행될 수 있음을 나타낸다.
도 17를 참고하면, 웨어러블 장치(1301)는, 가상 공간에서 이미지를 렌더링하기 위해 가상 공간 매니저(1750)(예: 도 16의 가상 공간 매니저(1651), CPM)를 실행할 수 있다. 가상 공간 매니저(1750)를 위해, 도 16의 가상 공간 매니저(1651)에 대한 설명들이 적어도 일부 참조될 수 있다. 가상 공간 매니저(1750)는 가상 공간 서비스를 지원하기 위한 플랫폼을 포함할 수 있다. 가상 공간 매니저(1750)는 런타임 서비스(1751)(예: OpenXR Runtime), 패널 렌더링(1752)(예: 2D Panel Render), XR 합성부(1753)(XR Compositor)를 포함할 수 있다. 웨어러블 장치(1301)는 런타임 서비스(1751)의 실행에 기반하여, 사용자의 포즈 예측 기능, 프레임 타이밍 기능, 및/또는 공간 입력 기능 중 적어도 하나를 실행할 수 있다. 런타임 서비스(1751)를 위해, 도 16의 런타임 서비스(1652)에 대한 설명들이 적어도 일부 참조될 수 있다. 웨어러블 장치(1301)는 패널 렌더링(1752)의 실행에 기반하여, 디스플레이(1350)를 통해 가상 공간을 구현할 수 있도록 패널(예: 2D 패널)에 적어도 하나의 이미지(영상)를 표시할 수 있다. 예를 들어, 웨어러블 장치(1301)는, 후술하는 공간화 매니저(1740)로부터의 패널을 위한 RGB 정보(1766)에 대응하는 렌더링 이미지를 디스플레이(예: 디스플레이(1350))를 통해 표시할 수 있다. 웨어러블 장치(1301)는 XR 합성부(1753)(XR Compositor)의 실행에 기반하여, 가상 공간 상에서 카메라를 통해 촬영된 실제 영역에 대한 이미지(이하, 패스-쓰루 이미지)와 가상 영역 이미지를 합성할 수 있다. 예를 들어, 웨어러블 장치(1301)는, XR 합성부(1753)의 실행에 기반하여, 상기 패스-쓰루 이미지와 상기 가상 영역 이미지를 병합함으로써, 합성 이미지를 생성할 수 있다. 웨어러블 장치(1301)는, 상기 합성 이미지가 표시되도록, 상기 생성된 합성 이미지를 디스플레이 버퍼에게 전송할 수 있다. 웨어러블 장치(1301)는 가상 공간 매니저(1750)를 통해 가상 공간을 식별할 수 있고, 가상 공간의 적어도 일부분을 디스플레이(1350) 상에 표시할 수 있다. 가상 공간 매니저(1750)는, CPM으로 지칭될 수 있다. 웨어러블 장치(1301)는 상기 가상 공간의 적어도 일부분에 대응하는 이미지를 렌더링하기 위해, 가상 공간 매니저(1750)를 실행할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 공간화 매니저(1740)를 실행할 수 있다. 공간화 매니저(1740)는 3차원의 가상 공간에 이미지를 표시하기 위한 처리들을 수행할 수 있다. 웨어러블 장치(1301)는, 가상 공간 매니저(1750)를 통해 3차원의 가상 공간에 이미지가 렌더링될 수 있도록, 공간화 매니저(1740)의 실행에 기반하여 사전 처리를 수행할 수 있다. 예를 들어, 웨어러블 장치(1301)는, 공간화 매니저(1740)의 실행에 기반하여, 도 16의 렌더러(1690)의 기능들 중 적어도 일부를 수행할 수 있다. 웨어러블 장치(1301)는, 공간화 매니저(1740)의 실행에 기반하여, 어플리케이션(예: XR 어플리케이션(1710), XR이 아닌 일반 2D 화면을 제공하는 어플리케이션(1720), 시스템 UI(1730)를 제공하는 어플리케이션)에 의해 제공되는 이미지 정보를 처리할 수 있다. 공간화 매니저(1740)(예: Space Flinger)는 시스템 화면 매니저(1741)(예: System scene), 입력 매니저(1742)(예: Input Routing), 및 경량 렌더링 엔진(1743)(예: Impress Engine)을 포함할 수 있다. 시스템 화면 매니저(1741)는, 시스템 UI(1730)를 표시하기 위해, 실행될 수 있다. 시스템 UI(1730)를 제공하는 프로그램(예: API)으로부터, 시스템 UI 관련 정보(1764)가 시스템 화면 매니저(1741)에게 전송될 수 있다. 시스템 UI 관련 정보(1764)는 공간화(spatializer) API 및/또는 Same-process private API를 통해 획득될 수 있다. 공간화 매니저(1740)는, 미리 할당된 리소스들을 통해, 3차원 공간에서 시스템 UI(1730)의 화면의 레이아웃(예: 위치, 표시 순서)를 결정할 수 있다. 시스템 화면 매니저(1741)는, 상기 레이아웃에 따라, 시스템 UI(1730)의 화면을 렌더링하기 위한 이미지 정보(1767)를 가상 공간 매니저(1750)에게 전송할 수 있다. 입력 매니저(1742)는 사용자 입력(예: 시스템 화면이나 앱 화면 상에서의 사용자 입력)을 처리하도록 구성될 수 있다. 입력 매니저(1742)는, 웨어러블 장치(1301)의 적어도 하나의 센서(1620)에 의해 인식된 사용자 입력을, 공간화 매니저(1740)에 의해 가상 공간에 매핑된 하나 이상의 소프트웨어 어플리케이션들(예: XR 어플리케이션(1710), XR이 아닌 일반 2D 화면을 제공하는 어플리케이션(1720), 시스템 UI(1730)를 제공하는 어플리케이션) 중 적어도 하나로 매핑할 수 있다. 예를 들어, 사용자 입력의 매핑은, 사용자 입력을 처리하기 위한 소프트웨어 어플리케이션의 인스트럭션들(예, 서브-루틴 및/또는 이벤트 핸들러)을 실행하는 동작을 포함할 수 있다. 경량 렌더링 엔진(1743)은 이미지 생성을 위한 렌더러(예: 경량 렌더러(1643))일 수 있다. 예를 들어, 경량 렌더링 엔진(1743)은 시스템 UI(1730)를 표시하기 위해 이용될 수 있다.
일 실시예에 따르면, 공간화 매니저(1740)에는 시스템 UI를 렌더링(rendering) 하기 위한 경량 렌더링 엔진(1743)을 포함할 수 있다. 일 실시예에 따르면, 경량 렌더링 엔진(1743)이 HMD에서 사용되는 아바타(avatar)를 렌더링하기에 리소스가 충분하지 않을 경우, 적어도 하나의 외부 랜더링 엔진이 사용될 수도 있다. 이때, 외부 렌더링(예: 3rd party 엔진)과의 호환성 이슈를 해결하기 위해, 공간화 매니저(1740) 내부에 외부 렌더링 엔진 지원 모듈이 추가될 수 있다.
일 실시예에 따르면, 전자 장치는 어플리케이션을 실행할 수 있다. 예를 들어, XR 어플리케이션(1710)(예: XR 어플리케이션(1642), 3D 게임, XR 맵, 기타 몰입형(immersive) 어플리케이션)의 실행에 응답하여, 가상 공간 매니저(1750)를 실행할 수 있다. 웨어러블 장치(1301)는 XR 어플리케이션(1710)으로부터 제공되는 이중 이미지 정보(1761)를 가상 공간 매니저(1750)에게 제공할 수 있다. 3차원 공간에서 이미지를 표시하기 위해, 이중 이미지 정보(1761)는 양안 시차를 고려한 2개의 이미지 정보를 포함할 수 있다. 예를 들어, 이중 이미지 정보(1761)는 3차원의 가상 공간에서 렌더링하기 위해, 사용자의 좌안을 위한 제1 이미지 정보와 사용자의 우안을 위한 제2 이미지 정보를 포함할 수 있다. 이하, 본 개시에서 3차원 공간에서 양안을 위한 이미지들을 나타내기 위한 이미지 정보를 지칭하는 용어로서, 이중 이미지 정보가 사용된다. 상기 이중 이미지 정보는 이중 이미지 정보 외에도 양안 이미지 정보, 이중 이미지 데이터, 이중 이미지, 양안 이미지 데이터, 입체 이미지 정보, 3D 이미지 정보, 공간 이미지 정보, 공간 이미지 데이터, 2D-3D 변환 데이터, 차원 변환 이미지 데이터, 양안 시차 이미지 데이터, 및/또는 이와 동등한 기술적 용어가 이용될 수 있다. 웨어러블 장치(1301)는 가상 공간 매니저(1750)를 통해 이미지 레이어들을 병합함으로써, 합성 이미지를 생성할 수 있다. 웨어러블 장치(1301)는 상기 생성된 합성 이미지를 디스플레이 버퍼에게 전송할 수 있다. 상기 합성 이미지는 웨어러블 장치(1301)의 디스플레이(1350) 상에 표시될 수 있다.
일 실시예에 따르면, 전자 장치는 XR 어플리케이션(1710)과 다른 어플리케이션(1720)(예: 제1 어플리케이션(1720-1), 제2 어플리케이션(1720-2), ..., N번째 어플리케이션(1720-N)) 중 적어도 하나의 어플리케이션을 실행할 수 있다. 일 실시예에 따르면, 어플리케이션(1720)은 2차원 이미지를 표시하기 위한 이미지 정보를 출력하도록 구성될 수 있다. 다시 말해, 어플리케이션(1720)은 2차원(2D) 이미지(예: 윈도우, 및/또는 액티비티)를 제공할 수 있다. 일 예로, 어플리케이션(1720)은 영상 어플리케이션, 일정 어플리케이션, 또는 인터넷 브라우저 어플리케이션일 수 있다. 만약, 어플리케이션(1720)의 실행에 응답하여, 어플리케이션(1720)으로부터 제공되는 이미지 정보(1762)가 가상 공간 매니저(1750)에게 제공되는 경우, 이미지 정보(1762)는 2차원 평면 내에서의 x좌표와 y좌표만을 가지므로, 사용자를 중심으로 다른 어플리케이션들 간의 선후 관계(즉, 사용자로부터 이격된 거리)가 고려되기 어려울 수 있다. 웨어러블 장치(1301)는, 일반적인 2D 화면을 제공하는 어플리케이션(1720)을 표시할 때에도, 이중 이미지 정보를 가상 공간 매니저(1750)에게 제공하기 위해, 공간화 매니저(1740)를 실행할 수 있다. 예를 들어, 공간화 매니저(1740)의 실행에 기반하여, 웨어러블 장치(1301)는 제1 어플리케이션(1720-1)으로부터 어플리케이션 관련 정보(1763)를 수신할 수 있다. 예를 들어, 어플리케이션 관련 정보(1763)는 제1 어플리케이션(1720-1)의 2차원 이미지를 나타내는 이미지 정보(예: 픽셀 별 RGB를 포함하는 정보) 및/또는 제1 어플리케이션(1720-1)에서의 컨텐츠 정보(예: 제1 어플리케이션에서 실행되는 컨텐츠의 특성, 컨텐츠의 유형)를 포함할 수 있다. 어플리케이션 관련 정보(1763)는 공간화(spatializer) API를 통해 획득될 수 있다. 공간화 매니저(1740)의 실행에 기반하여, 웨어러블 장치(1301)는 제1 어플리케이션(1720-1)이 렌더링될 영역의 위치 및 렌더링될 영역의 크기에 대한 정보(이하, 위치 정보)를 식별할 수 있다. 공간화 매니저(1740)의 실행에 기반하여, 웨어러블 장치(1301)는 상기 이미지 정보 및 상기 위치 정보를 통해, 사용자의 양안 시차가 고려된, 이중 이미지 정보(1765, 예: RGBx2)를 생성할 수 있다. 공간화 매니저(1740)의 실행에 기반하여, 웨어러블 장치(1301)는 이중 이미지 정보(1765)를 가상 공간 매니저(1750)에게 제공할 수 있다. 단순한 2차원 이미지를 이중 이미지 정보(1765)로 변환함으로써, 이미지 정보(1762)가 가상 공간 매니저(1750)에게 직접 전달됨으로써 발생하는 문제는 해소될 수 있다. 또한, 가상 공간에서의 이미지 표시를 위한 기능들 중 적어도 일부가 가상 공간 매니저(1750) 대신 공간화 매니저(1740)에 의해 수행됨에 따라, 가상 공간 매니저(1750)의 부담이 감소할 수 있다.
도 18은, 외부 객체로부터 발화된 오디오 신호를 인식하는 웨어러블 장치의 예를 도시한다. 도 18에서 예시되는 웨어러블 장치(1301)의 동작들은, 프로세서(예: 프로세서(1610))에 의해 실행되거나 수행되거나 제어될 수 있다.
도 18을 참고하면, 사용자(1801)는 웨어러블 장치(1301)를 착용할 수 있다. 웨어러블 장치(1301)는 마이크로폰(예: 마이크로폰(1365))을 통해 오디오 신호를 획득할 수 있다. 예를 들면, 오디오 신호는, 보이스 신호(1810) 및/또는 노이즈를 포함할 수 있다. 예를 들면, 오디오 신호 내에 보이스 신호(1810)의 크기가 상대적으로 작을 수 있다. 예를 들면, 오디오 신호 내에 노이즈의 크기가 상대적으로 클 수 있다. 예를 들면, 오디오 신호의 SNR이 상대적으로 작을 수 있다. 예를 들면, 오디오 신호의 SNR이 상대적으로 작은 경우에, 보이스 신호(1810)의 인식률을 높이기 위한 방안이 요구될 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 보이스 신호(1810)의 발화자(speaker)(1802)의 얼굴의 변화를 검출할 수 있다. 예를 들면, 발화자(1802)의 얼굴의 변화는, 발화자(1802)가 발화함에 따라 야기되는 얼굴의 모양의 변화로 참조될 수 있다. 예를 들면, 발화자(1802)는 발언자(utterer)로 참조될 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)의 발화자(1802)의 입술의 변화를 검출할 수 있다. 예를 들면, 발화자(1802)의 입술의 변화는, 발화자(1802)가 발화함에 따라 야기되는 입술의 모양의 변화로 참조될 수 있다. 예를 들면, 웨어러블 장치(1301)는, 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 통해 센싱 데이터를 획득할 수 있다. 예를 들면, 센싱 데이터는, 이미지 센서(예: 이미지 센서(1621))를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는, 외부 환경(또는 물리 환경)이 촬영된 이미지일 수 있다. 예를 들면, 이미지는, 발화자(1802)에 대응하는 시각적 객체(1821)를 포함할 수 있다. 예를 들면, 시각적 객체(1821)의 일부는, 발화자(1802)의 얼굴(예: 입술)에 대응할 수 있다. 예를 들면, 웨어러블 장치(1301)는 이미지 센서(1621)를 통해 획득된 이미지에서 시각적 객체(1821)의 일부를 식별하는 것에 따라, 발화자(1802)의 얼굴(예: 입술)의 변화를 검출할 수 있다.
웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810)를 포함함) 및/또는 적어도 하나의 센서(1620)를 통해 센싱 데이터를, 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호 및/또는 발화자(1802)의 입술의 변화를 나타내는 데이터를, 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 제1 멀티모달 모델(350)을 이용하여, 오디오 신호의 SNR이 상대적으로 작은 환경에서도, 오디오 신호에 포함된 보이스 신호(1810)를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 상대적으로 노이즈가 많거나 및/또는 발화자(1802)가 상대적으로 먼 위치에 있더라도, 발화자(1802)의 보이스 신호(1810)를 식별할 수 있다.
웨어러블 장치(1301)는 오디오 신호 및/또는 센싱 데이터를 제1 멀티모달 모델(350)에게 제공함으로써, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보를 획득할 수 있다. 예를 들면, 응답 정보는, 오디오 신호 및/또는 센싱 데이터에 따라, 제1 멀티모달 모델(350)에 의해 생성될 수 있다. 예를 들면, 응답 정보는, 오디오 신호에 포함된 보이스 신호(1810)를 나타내는 텍스트를 디스플레이(예: 디스플레이(1350))를 통해 표시하기 위해 이용될 수 있다. 예를 들면, 응답 정보는, 보이스 신호(1810)를 나타내는 오디오 신호를 스피커(예: 스피커(1355))를 통해 출력하기 위해 이용될 수 있다. 제한되지 않는 예로, 제1 멀티모달 모델(350)에게 제공되는 오디오 신호는, 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))에 따라 노이즈 캔슬링이 수행된 오디오 신호일 수 있다.
웨어러블 장치(1301)는, 이미지 센서(1621)를 통해 획득된 이미지를 이용하여, 외부 환경에 대한 패스-쓰루 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 물리 환경에 대응하는 가상 공간(또는 3차원 공간)의 적어도 일부를 표현하는 화면(1820)을 디스플레이(예: 디스플레이(1350))를 통해 제공할 수 있다. 예를 들면, 화면(1820)은 시각적 객체(1821)를 포함할 수 있다. 예를 들면, 시각적 객체(1821)는 발화자(1802)에 대응할 수 있다. 예를 들면, 시각적 객체(1821)는 발화자(1802)를 표현할 수 있다.
웨어러블 장치(1301)는 상기 응답 정보를 이용하여, 화면(1820) 내의 영역(1825)에서 텍스트(1827)를 표시할 수 있다. 텍스트(1827)는, 보이스 신호(1810)에 대응할 수 있다. 예를 들면, 텍스트(1827)는, 보이스 신호(1810)를 나타낼 수 있다. 영역(1825)의 위치는, 시각적 객체(1821)에 인접할 수 있다. 예를 들면, 영역(1825)은, 화면(1820)에서 시각적 객체(1821)와 나란할 수 있다. 예를 들면, 영역(1825)은 화면(1820)에서 미리 지정된 위치를 가질 수 있다. 제한되지 않는 예로, 영역(1825)은, 시각적 객체(1821)의 적어도 일부에 중첩될 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 상기 응답 정보를 이용하여, 오디오 신호(1830)를 스피커(1355)를 통해 출력할 수 있다. 예를 들면, 오디오 신호(1830)는, 보이스 신호(1810)에 대응할 수 있다. 예를 들면, 오디오 신호(1830)의 SNR은, 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810)를 포함함)의 SNR 보다 높을 수 있다. 일 실시예로, 오디오 신호(1830)는, 마이크로폰(1365)을 통해 획득된 오디오 신호에 대한 노이즈 캔슬링을 수행함으로써, 획득될 수 있다. 예를 들면, 노이즈 캔슬링을 위해, 제4 멀티모달 모델(380)이 이용될 수 있다. 일 실시예로, 오디오 신호(1830)는 전자 장치(1301)에 의해 생성될 수 있다. 예를 들면, 오디오 신호(1830)는 상기 응답 정보를 이용하여 생성될 수 있다. 예를 들면, 웨어러블 장치(1301)는, 보이스 신호(1810)를 나타내는 텍스트(1827)에 대한 TTS(text-to-speech)를 수행함으로써, 오디오 신호(1830)를 생성하거나 획득할 수 있다. 예를 들면, 오디오 신호(1830)는 제1 멀티모달 모델(350)에 의해 생성될 수 있다. 예를 들면, 생성된 오디오 신호(1830)의 음성 특징은, 미리 결정된 음성 특징일 수 있다. 예를 들면, 생성된 오디오 신호(1830)의 음성 특징은, 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810)를 포함함)의 오디오 신호와 실질적으로 동일할 수 있다. 예를 들면, 전자 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호의 음성 특징을 이용하여, 오디오 신호(1830)를 생성할 수 있다. 제한되지 않는 예로, 전자 장치(1301)는, 미리 저장된 타인(예: 연예인)의 음성 특징을 이용하여, 오디오 신호(1830)를 생성할 수 있다. 예를 들면, 음성 특징은, 발화자(1802)의 성별, 성문, 음색, 피치, 또는 발음 중 적어도 하나를 포함할 수 있다.
제한되지 않는 예로, 마이크로폰(1365)을 통해 획득된 오디오 신호에서 웨어러블 장치(1301)에 의해 식별되지 않는 보이스 신호(1810)의 일 영역이 있을 수 있다. 예를 들면, 웨어러블 장치(1301)는 상기 일 영역에 대한 추론을 수행할 수 있다. 예를 들면, 웨어러블 장치(1301)는 인공지능 모델을 이용하여, 상기 일 영역에 대한 추론을 수행할 수 있다. 예를 들면, 웨어러블 장치(1301)는 일 영역에 대한 추론을 수행함으로써, 텍스트(1827) 및/또는 오디오 신호(1830)를 획득할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 시각적 객체(1829)를 화면(1820)을 통해 표시할 수 있다. 예를 들면, 시각적 객체(1829)는 웨어러블 장치(1301)가 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810))에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공하고 있음을 나타낼 수 있다. 또한, 시각적 객체(1829)는, 웨어러블 장치(1301)가 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810))에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공하고 있음을 나타낼 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자 입력에 따라, 웨어러블 장치(1301)의 상태를 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공하지 않는 상태로 변경할 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자 입력에 따라, 웨어러블 장치(1301)의 상태를 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공하지 않는 상태로 변경할 수 있다. 예를 들면, 상기 사용자 입력은, 시각적 객체(1829)에 대한 입력을 포함할 수 있으나, 실시예가 제한되는 것은 아니다. 예를 들면, 상기 사용자 입력은, 미리 정해진 제스쳐 입력을 포함할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 노이즈가 기준 노이즈 보다 큰 환경에 웨어러블 장치(1301)가 위치되는 동안, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능 및/또는 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공할 수 있다. 웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호의 품질이 상대적으로 낮은 경우에, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능 및/또는 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호에 포함된 보이스 신호(1810)의 크기를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)의 크기가 기준 크기 보다 작은 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)의 크기가 기준 크기 보다 작은 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공할 수 있다.
웨어러블 장치(1301)는 보이스 신호(1810)의 크기가 기준 크기 보다 큰 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공하는 것을 삼가거나 중단하거나 바이패스할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)의 크기가 기준 크기 보다 큰 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공하는 것을 삼가거나 중단하거나 바이패스할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810)를 포함함)의 SNR을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호의 SNR이 기준 값 보다 작은 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는 SNR이 기준 값 보다 작은 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공할 수 있다.
웨어러블 장치(1301)는 보이스 신호(1810)의 오디오 신호의 SNR이 기준 값 보다 큰 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 텍스트(1827)를 디스플레이(1350)를 통해 표시하는 기능을 제공하는 것을 삼가거나 중단하거나 바이패스할 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호의 SNR이 기준 값 보다 큰 것을 식별하는 것에 기반하여, 오디오 신호에 기반하여 오디오 신호(1830)를 스피커(1355)를 통해 출력하는 기능을 제공하는 것을 삼가거나 중단하거나 바이패스할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 외부 전자 장치(1838)(예: 스마트폰)와 통신을 수행할 수 있다. 예를 들면, 외부 전자 장치(1838)는, 도 3a 및/또는 도 3b의 전자 장치(301)의 일 예일 수 있다. 예를 들면, 웨어러블 장치(1301)는 외부 전자 장치(1838)의 컴패니언(companion) 기기로 참조될 수 있다. 예를 들면, 웨어러블 장치(1301)는 웨어러블 장치(1301)와 외부 전자 장치(1838) 사이에 통신 링크를 수립할 수 있다. 예를 들면, 웨어러블 장치(1301)는 마이크로폰(1365)을 통해 획득된 오디오 신호(예: 보이스 신호(1810)를 포함함)를 외부 전자 장치(1838)에게 송신할 수 있다. 예를 들면, 웨어러블 장치(1301)는 적어도 하나의 센서(1620)를 통해 획득된 센싱 데이터를 외부 전자 장치(1838)에게 송신할 수 있다. 예를 들면, 센싱 데이터는, 이미지 센서(1621)를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는, 발화자(1802)의 입술에 대응하는 시각적 객체(1821)의 일부를 포함할 수 있다.
일 실시예에 따르면, 외부 전자 장치(1838)는, 제1 멀티모달 모델(350)을 포함할 수 있다. 예를 들면, 외부 전자 장치(1838)는, 웨어러블 장치(1301)로부터 송신된 오디오 신호 및/또는 웨어러블 장치(1301)로부터 송신된 센싱 데이터에 기반하여, 화면(1840)을 외부 전자 장치(1838)의 디스플레이를 통해 표시할 수 있다. 예를 들면, 화면(1840)은 화면(1820)에 대응할 수 있다. 예를 들면, 화면(1840)은 화면(1820)에 실질적으로 동일할 수 있으므로, 중복되는 설명은 생략하도록 한다. 예를 들면, 시각적 객체(1841)는 시각적 객체(1821)에 대응할 수 있다. 예를 들면, 영역(1845)은 영역(1825)에 대응할 수 있다. 예를 들면, 영역(1845)을 위해 영역(1825)에 대한 설명들이 참조될 수 있다. 예를 들면, 텍스트(1847)는 텍스트(1827)에 대응할 수 있다. 텍스트(1827)를 위해 텍스트(1847)에 대한 설명들이 참조될 수 있다. 외부 전자 장치(1838)는, 오디오 신호 및/또는 송신된 센싱 데이터에 기반하여, 오디오 신호(1850)를 외부 전자 장치(1838)의 스피커를 통해 출력할 수 있다. 예를 들면, 오디오 신호(1850)는 오디오 신호(1830)에 대응할 수 있다. 오디오 신호(1830)를 위해 오디오 신호(1850)에 대한 설명들이 참조될 수 있다.
일 실시예에 따르면, 외부 전자 장치(1838)는, 텍스트(1847)를 나타내는 데이터 및/또는 오디오 신호(1850)를 나타내는 데이터를 웨어러블 장치(1301)에게 송신할 수 있다. 예를 들면, 웨어러블 장치(1301)는 텍스트(1847)를 나타내는 데이터를 이용하여, 텍스트(1827)를 디스플레이(1350)를 통해 표시할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 오디오 신호(1850)를 나타내는 데이터를 이용하여, 오디오 신호(1830)를 스피커(1355)를 통해 출력할 수 있다.
도 18에서, 웨어러블 장치(1301)는 패스 쓰루 기능을 제공하고 있는 상태로 도시되고 있으나, 실시예가 제한되는 것은 아니다. 예를 들면, 웨어러블 장치(1301)는, 비디오를 디스플레이(1350)를 통해 표시하거나 재생하는 동안, 도 18에서 예시된 동작들을 실행할 수 있다.
도 19는 외부 객체의 시선에 따라 적응적으로 오디오 신호에 대한 응답을 제공하는 웨어러블 장치의 예를 도시한다. 도 19에서 예시되는 웨어러블 장치(1301)의 동작들은, 프로세서(예: 프로세서(1610))에 의해 실행되거나 수행되거나 제어될 수 있다. 도 19에 대한 설명 중에서 도 18의 설명과 중복되는 내용은 생략될 수 있다.
도 19를 참고하면, 예(1901)에서, 웨어러블 장치(1301)는 사용자(1801)에게 패스-쓰루 기능을 제공할 수 있다. 웨어러블 장치(1301)는 마이크로폰(예: 마이크로폰(1365))을 통해 오디오 신호를 획득할 수 있다. 예를 들면, 오디오 신호는, 보이스 신호(1810), 보이스 신호(1905), 및/또는 노이즈를 포함할 수 있다. 웨어러블 장치(1301)는, 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 통해 센싱 데이터를 획득할 수 있다. 예를 들면, 센싱 데이터는, 이미지 센서(예: 이미지 센서(1621))를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는, 외부 환경(또는 물리 환경)이 촬영된 이미지일 수 있다. 예를 들면, 이미지는, 발화자(1802)에 대응하는 시각적 객체(1921) 및/또는, 발화자(1903)에 대응하는 시각적 객체(1923)를 포함할 수 있다. 예를 들면, 시각적 객체(1921)는 도 18의 시각적 객체(1821)의 일 예일 수 있다. 예를 들면, 이미지는, 발화자(1802)의 입술에 대응하는 시각적 객체(1921)의 일부, 및/또는 발화자(1903)의 입술에 대응하는 시각적 객체(1923)의 일부를 포함할 수 있다.
웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호 및/또는 적어도 하나의 센서(1620)를 통해 센싱 데이터를, 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 제1 멀티모달 모델(350)을 이용하여, 오디오 신호의 SNR이 상대적으로 작은 환경에서도, 보이스 신호(1810), 및/또는 보이스 신호(1905)를 식별할 수 있다.
웨어러블 장치(1301)는 오디오 신호 및/또는 센싱 데이터를 제1 멀티모달 모델(350)에게 제공함으로써, 오디오 신호 및/또는 센싱 데이터에 대한 응답 정보를 획득할 수 있다. 예를 들면, 응답 정보는, 오디오 신호 및/또는 센싱 데이터에 따라, 제1 멀티모달 모델(350)에 의해 생성될 수 있다. 예를 들면, 응답 정보는, 보이스 신호(1810)를 나타내는 텍스트 및/또는 보이스 신호(1905)를 나타내는 텍스트를, 디스플레이(예: 디스플레이(1350))를 통해, 표시하기 위해 이용될 수 있다. 예를 들면, 응답 정보는, 보이스 신호(1810)를 나타내는 오디오 신호, 및/또는 보이스 신호(1905)를 나타내는 오디오 신호를, 스피커(예: 스피커(1355))를 통해 출력하기 위해 이용될 수 있다. 제한되지 않는 예로, 제1 멀티모달 모델(350)에게 제공되는 오디오 신호는, 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))에 따라 노이즈 캔슬링이 수행된 오디오 신호일 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 디스플레이(1350)를 통해 화면(1920)을 표시할 수 있다. 화면(1920)은 발화자(1802)에 대응하는 시각적 객체(1921) 및/또는 발화자(1903)에 대응하는 시각적 객체(1923)를 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)는 상기 응답 정보를 이용하여, 텍스트(1927)를 화면(1920) 내의 영역(1925)에 표시할 수 있다. 예를 들면, 텍스트(1927)는 보이스 신호(1810)를 나타낼 수 있다. 예를 들면, 웨어러블 장치(1301)는 상기 응답 정보를 이용하여, 텍스트(1928)를 화면(1920) 내의 영역(1926)에 표시할 수 있다. 예를 들면, 텍스트(1928)는 보이스 신호(1905)를 나타낼 수 있다. 예를 들면, 영역(1925) 및 영역(1926) 각각은, 도 18의 영역(1825)에 실질적으로 동일할 수 있으므로, 중복되는 내용이 생략될 수 있다.
웨어러블 장치(1301)는 상기 응답 정보를 이용하여, 오디오 신호(1930), 및/또는 오디오 신호(1931)를 스피커(예: 스피커(1355))를 통해 출력할 수 있다. 예를 들면, 오디오 신호(1930)는 보이스 신호(1810)를 나타낼 수 있다. 예를 들면, 오디오 신호(1930)의 품질은 보이스 신호(1810)의 품질 보다 높을 수 있다. 예를 들면, 오디오 신호(1930)의 SNR은 보이스 신호(1810)의 SNR 보다 높을 수 있다. 예를 들면, 오디오 신호(1930)의 크기(예: 음량)는 보이스 신호(1810)의 크기 보다 클 수 있다. 예를 들면, 오디오 신호(1931)는 보이스 신호(1905)를 나타낼 수 있다. 예를 들면, 오디오 신호(1931)의 품질은 보이스 신호(1905)의 품질 보다 높을 수 있다. 예를 들면, 오디오 신호(1931)의 SNR은 보이스 신호(1905)의 SNR 보다 높을 수 있다. 예를 들면, 오디오 신호(1931)의 크기(예: 음량)는 보이스 신호(1905)의 크기 보다 클 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는, 적어도 하나의 센서(1620)를 통해 획득된 센싱 데이터를 이용하여, 발화자(1802)의 위치 및/또는 발화자(1903)의 위치를 식별할 수 있다. 예를 들면, 센싱 데이터는 이미지 센서(1621)를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는 시각적 객체(1921) 및/또는 시각적 객체(1923)를 포함할 수 있다. 제한되지 않는 예로, 센싱 데이터는 ToF 센서를 통해 획득된 데이터를 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 웨어러블 장치(1301)와 발화자(1802) 사이의 거리를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 웨어러블 장치(1301)로부터 발화자(1802)를 향하는 방향을 식별할 수 있다. 웨어러블 장치(1301)는 발화자(1802)의 얼굴(예: 입술)(또는 시선)이 향하는 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)의 얼굴(예: 입술)에 대응하는 부분이 향하는 방향을 식별함으로써, 발화자(1802)의 얼굴(예: 입술)(또는 시선)이 향하는 방향을 식별할 수 있다.
웨어러블 장치(1301)는, 웨어러블 장치(1301)와 발화자(1903) 사이의 거리를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 웨어러블 장치(1301)로부터 발화자(1903)를 향하는 방향을 식별할 수 있다. 웨어러블 장치(1301)는 발화자(1903)의 얼굴(예: 입술)(또는 시선)이 향하는 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1923)의 얼굴(예: 입술)에 대응하는 부분이 향하는 방향을 식별함으로써, 발화자(1903)의 얼굴(예: 입술)(또는 시선)이 향하는 방향을 식별할 수 있다.
웨어러블 장치(1301)는 마이크로폰(1365)을 통해 획득된 오디오 신호를 이용하여, 오디오 신호에 포함된 보이스 신호(예: 보이스 신호(1810), 보이스 신호(1905))를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호에 포함된 보이스 신호의 발화 위치를 식별할 수 있다. 예를 들면, 마이크로폰(1365)은 복수 개의 마이크로폰들을 포함할 수 있다. 예를 들면, 복수 개의 마이크로폰들 각각은 보이스 신호를 획득할 수 있다. 예를 들면, 복수 개의 마이크로폰들 각각에 도달하는 보이스 신호의 시간 차이에 따라, 획득된 보이스 신호의 위상은 다를 수 있다. 예를 들면, 웨어러블 장치(1301)는 획득된 보이스 신호의 위상을 이용하여, 보이스 신호의 발화 위치를 식별할 수 있다. 제한되지 않는 예로, 웨어러블 장치(1301)는 복수 개의 마이크로폰들 각각에 도달하는 보이스 신호의 음압 차이를 이용하여, 보이스 신호의 발화 위치를 식별할 수 있다.
웨어러블 장치(1301)는 보이스 신호들(예: 보이스 신호(1810), 보이스 신호(1905)) 각각의 발화자가 누구인지를 식별할 수 있다. 웨어러블 장치(1301)는 식별된 보이스 신호의 발화 위치와 식별된 발화자(예: 발화자(1802), 발화자(1903))의 위치 사이의 유사도에 따라, 해당 보이스 신호의 발화자가 누구인지를 결정할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)의 발화 위치 및 발화자(1802)의 위치 사이의 유사도에 따라, 보이스 신호(1810)가 발화자(1802)로부터 발화됨을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1905)의 발화 위치 및 발화자(1903)의 위치 사이의 유사도에 따라, 보이스 신호(1905)가 발화자(1903)로부터 발화됨을 식별할 수 있다.
웨어러블 장치(1301)는 보이스 신호(예: 보이스 신호(1810), 보이스 신호(1905))의 발화자가 누구인지 식별할 수 있기 때문에, 보이스 신호들의 발화자들 각각에 따라, 보이스 신호에 따른 기능을 제공할 수 있다. 예를 들면, 보이스 신호에 따른 기능은, 보이스 신호를 나타내는 텍스트(예: 텍스트(1927), 텍스트(1928))를 표시하는 기능 및/또는 보이스 신호를 나타내는 오디오 신호(예: 오디오 신호(1930), 오디오 신호(1931))를 출력하는 기능을 포함할 수 있다.
웨어러블 장치(1301)는 발화자를 결정할 수 있다. 예를 들면, 웨어러블 장치(1301)는 화면(1920) 내의 시각적 객체(1921)에 대한 사용자 입력(1922)을 수신함엔 따라, 발화자들 중에서 발화자(1802)를 결정할 수 있다. 예를 들면, 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 예(1901)에서 예(1902)로 변경될 수 있다. 예를 들면, 시각적 객체(1921)에 대한 사용자 입력(1922)은, 시각적 객체(1921)에 대한 제스쳐 입력, 시각적 객체(1921)에 대한 시선 입력, 또는 시각적 객체(1921)에 대한 음성 입력 중 적어도 하나를 포함할 수 있다.
예(1902)에서, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 웨어러블 장치(1301)는 보이스 신호(1810)에 따른 기능을 제공하는 것을 유지할 수 있다. 웨어러블 장치(1301)는, 화면(1920) 내의 영역(1925)에서 텍스트(1927)를 표시하는 것을 유지할 수 있다. 웨어러블 장치(1301)는 오디오 신호(1930)를 스피커(1355)를 통해 출력하는 것을 유지할 수 있다.
웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 웨어러블 장치(1301)는 보이스 신호(1905)에 따른 기능을 제공하는 것을 중지하거나 스킵하거나 삼가하거나 바이패스할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 텍스트(1928)를 표시하는 것을 중지하거나 스킵하거나 삼가하거나 바이패스할 수 있다. 예를 들면, 화면(1920)에 텍스트(1928)가 표시되지 않을 수 있다. 웨어러블 장치(1301)는 오디오 신호(1931)를 스피커(1355)를 통해 출력하는 것을 중지하거나 스킵하거나 삼가하거나 바이패스할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 마이크로폰(1365)을 통해 획득된 보이스 신호(1905)에 대한 캔슬링을 수행할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 오디오 신호(1930) 및 오디오 신호(1931) 중 오디오 신호(1930)를 출력할 수 있다.
제한되지 않는 예로, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 마이크로폰(1365)을 통해 획득된 보이스 신호(1905) 및 오디오 신호(1930)를 스피커(1355)를 통해 출력할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 시각적 객체(1921)를 추적할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 시각적 객체(1921)가 화면(1920) 내에 위치되는지 여부를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)가 화면(1920) 내에 위치되는 동안, 보이스 신호(1810)에 따른 기능을 제공하는 것을 유지할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)의 적어도 일부가 화면(1920) 내에 위치되지 않음을 식별하는 것에 기반하여, 보이스 신호(1810)에 따른 기능을 제공하는 것을 유지하는 것을 중지할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(1810)에 따른 기능을 제공하는 것을 일시적으로 비활성화할 수 있다.
예를 들면, 웨어러블 장치(1301)는 시각적 객체(1921)의 적어도 일부가 화면(1920) 내에 위치되지 않음을 식별하는 것에 기반하여, 타이머를 활성화할 수 있다. 예를 들면, 웨어러블 장치(1301)는 타이머가 만료되기 전 시각적 객체(1921)가 화면(1920) 내에 위치되는 것을 식별하는 것에 기반하여, 보이스 신호(1810)에 따른 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는 타이머의 만료에 응답하여, 보이스 신호(1810)에 따른 기능을 종료할 수 있다. 제한되지 않는 예로, 웨어러블 장치(1301)는 이미지 센서(1621)를 통해 획득된 하나 이상의 이미지들 내에 시각적 객체(1921)가 포함됨을 식별하는 것에 기반하여, 보이스 신호(1810)에 따른 기능을 종료하지 않을 수 있다.
도 19에서 도시되지 않았지만, 일 실시예에 따르면, 웨어러블 장치(1301)는 시각적 객체(1921)에 대한 사용자 입력(1922)에 따라, 시각적 객체(1921)에 대한 시각적 효과를 화면(1920)을 통해 표시할 수 있다. 예를 들면, 상기 시각적 효과는 사용자 입력(1922)의 대상이 된 시각적 객체(1921)를 다른 시각적 객체와 구별하기 위해 이용될 수 있다.
도 20a 내지 20b는 사용자의 시선에 따라 적응적으로 오디오 신호에 대한 응답을 제공하는 웨어러블 장치의 예를 도시한다. 도 20a 내지 도 20b에서 예시되는 웨어러블 장치(1301)의 동작들은, 프로세서(예: 프로세서(1610))에 의해 실행되거나 수행되거나 제어될 수 있다. 도 18 및 도 19의 설명과 중복되는 내용이 생략될 수 있다.
도 20a를 참고하면, 예(2001)에서, 웨어러블 장치(1301)는 사용자(1801)에게 패스-쓰루 기능을 제공할 수 있다. 웨어러블 장치(1301)는 마이크로폰(예: 마이크로폰(1365))을 통해 오디오 신호를 획득할 수 있다. 예를 들면, 오디오 신호는, 보이스 신호(1810), 보이스 신호(2005), 및/또는 노이즈를 포함할 수 있다. 웨어러블 장치(1301)는, 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 통해 센싱 데이터를 획득할 수 있다. 예를 들면, 센싱 데이터는, 이미지 센서(예: 이미지 센서(1621))를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는, 외부 환경(또는 물리 환경)이 촬영된 이미지일 수 있다. 예를 들면, 이미지는, 발화자(1802)에 대응하는 시각적 객체(2021) 및/또는, 발화자(2003)에 대응하는 시각적 객체(2023)를 포함할 수 있다. 예를 들면, 이미지는, 발화자(1802)의 입술에 대응하는 시각적 객체(2021)의 일부를 포함할 수 있다.
웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호 및/또는 적어도 하나의 센서(1620)를 통해 센싱 데이터를, 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 제1 멀티모달 모델(350)을 이용하여, 오디오 신호의 SNR이 상대적으로 작은 환경에서도, 보이스 신호(1810), 및/또는 보이스 신호(2005)를 식별할 수 있다.
일 실시예에 따르면, 예(2001)에서, 사용자(1801)와 발화자(1802)는 서로 바라보고 있을 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 발화자(1802)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 이미지 센서(1621)를 통해 획득된 이미지에 포함된 시각적 객체(2021)의 시선의 방향을 식별할 수 있다. 예를 들면, 발화자(1802)의 시선의 방향은 시각적 객체(2021)의 시선의 방향에 대응할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(2021)의 시선의 방향을 이용하여 발화자(1802)의 시선의 방향을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향 및 발화자(1802)의 시선의 방향에 기반하여, 사용자(1801)와 발화자(1802)는 서로 바라보고 있음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향이 발화자(1802)의 시선의 방향에 대응함을 식별함에 따라, 사용자(1801)와 발화자(1802)는 서로 바라보고 있음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)와 발화자(1802)는 서로 바라보고 있음을 식별하는 것에 기반하여, 발화자(1802)의 보이스 신호(1810)에 따른 기능을 제공할 수 있다. 예를 들면, 보이스 신호(1810)에 따른 기능을 위해, 도 19의 설명들이 참조될 수 있다. 예를 들면, 보이스 신호(1810)에 따른 기능은, 보이스 신호(1810)를 나타내는 텍스트(2027)를 화면(2020) 내의 영역(2025)을 통해 표시하는 기능 및/또는 보이스 신호(1810)를 나타내는 오디오 신호(2030)를 출력하는 기능을 포함할 수 있다.
일 실시예에 따르면, 사용자(1801)와 발화자(2003)는 서로 바라보고 있지 않을 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 발화자(2003)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 이미지 센서(1621)를 통해 획득된 이미지에 포함된 시각적 객체(2023)의 시선의 방향을 식별할 수 있다. 예를 들면, 발화자(2003)의 시선의 방향은 시각적 객체(2023)의 시선의 방향에 대응할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(2023)의 시선의 방향을 이용하여 발화자(2003)의 시선의 방향을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향 및 발화자(2003)의 시선의 방향에 기반하여, 사용자(1801)와 발화자(2003)는 서로 바라보고 있지 않음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향이 발화자(2003)의 시선의 방향에 대응하지 않음을 식별함에 따라, 사용자(1801)와 발화자(2003)는 서로 바라보고 있지 않음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)와 발화자(2003)는 서로 바라보고 있지 않음을 식별하는 것에 기반하여, 발화자(2003)의 보이스 신호(1810)에 따른 기능을 제공하는 것을 삼가하거나 중지하거나 스킵하거나 바이패스할 수 있다.
일 실시예에 따르면, 발화자(2003)의 자세의 변경에 따라, 예(2001)에서 예(2002)로 변경될 수 있다. 예(2002)에서, 사용자(1801)와 발화자(2003)는 서로 바라보고 있을 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 발화자(2003)의 시선의 방향을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 이미지 센서(1621)를 통해 획득된 이미지에 포함된 시각적 객체(2023)의 시선의 방향을 식별할 수 있다. 예를 들면, 발화자(2003)의 시선의 방향은 시각적 객체(2023)의 시선의 방향에 대응할 수 있다. 예를 들면, 웨어러블 장치(1301)는 시각적 객체(2023)의 시선의 방향을 이용하여 발화자(2003)의 시선의 방향을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향 및 발화자(2003)의 시선의 방향에 기반하여, 사용자(1801)와 발화자(2003)는 서로 바라보고 있음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)의 시선의 방향이 발화자(2003)의 시선의 방향에 대응함을 식별함에 따라, 사용자(1801)와 발화자(2003)는 서로 바라보고 있음을 식별할 수 있다. 웨어러블 장치(1301)는 사용자(1801)와 발화자(2003)는 서로 바라보고 있음을 식별하는 것에 기반하여, 발화자(2003)의 보이스 신호(1810)에 따른 기능을 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는 화면(2020) 내의 영역(2026)에 텍스트(2028)를 더 표시할 수 있다. 예를 들면, 텍스트(2028)는 보이스 신호(2005)를 나타낼 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호(2040)를 스피커(1355)를 통해 출력할 수 있다. 예를 들면, 오디오 신호(2040)는 보이스 신호(1810) 및/또는 보이스 신호(2005)를 나타낼 수 있다. 예를 들면, 오디오 신호(2040)는, 오디오 신호(2030)와 보이스 신호(2005)가 결합된 신호에 대응할 수 있다.
도 20b를 참고하면, 도 20b에서 예시되는 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다.
동작 2051에서, 웨어러블 장치(1301)는 3차원 공간의 적어도 일부를 표현하는 화면을 디스플레이(예: 디스플레이(1350))를 통해 표시할 수 있다. 예를 들면, 3차원 공간은 웨어러블 장치(1301)의 패스-쓰루 기능에 따라 제공될 수 있다. 예를 들면, 3차원 공간은 물리 환경에 대응할 수 있다. 예를 들면, 웨어러블 장치(1301)는 이미지 센서(예: 이미지 센서(1361))를 통해 획득된 이미지를 이용하여, 3차원 공간의 적어도 일부를 표현하는 화면을 표시할 수 있다. 예를 들면, 3차원 공간의 적어도 일부를 표현하는 화면은, 하나 이상의 시각적 객체들을 포함할 수 있다. 예를 들면, 하나 이상의 시각적 객체들 내의 각 시각적 객체는, 발화자에 대응할 수 있다.
동작 2053에서, 웨어러블 장치(1301)는 마이크로폰(예: 마이크로폰(1365))을 통해 획득된 오디오 신호에 기반하여, 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별할 수 있다. 예를 들면, 오디오 신호는, 디스플레이(1350)를 통해 화면이 표시되는 동안 획득될 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호에 포함된 보이스 신호에 대한 발화자를 식별할 수 있다. 오디오 신호에 포함된 보이스 신호에 대한 발화자를 식별하는 방법을 위해, 도 19의 설명들이 참조될 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호에 대한 발화자에 대응하는 시각적 객체를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 오디오 신호를 이용하여, 시각적 객체에 대응하는 발화자를, 하나 이상의 발화자들 중에서, 식별할 수 있다.
동작 2055에서, 웨어러블 장치(1301)는 웨어러블 장치(1301)를 착용한 사용자의 시선의 제1 방향 및 보이스 신호에 대응하는 시각적 객체의 시선의 제2 방향을, 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 통해, 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자의 시선을 이미지 센서(1621)(예: 시선 추적 카메라(1360-1))를 통해 식별할 수 있다. 예를 들면, 보이스 신호에 대응하는 시각적 객체의 시선은 시각적 객체에 대응하는 발화자의 시선과 관련될 수 있다. 예를 들면, 시각적 객체의 시선은 발화자의 시선에 대응할 수 있다.
동작 2057에서, 웨어러블 장치(1301)는 웨어러블 장치(1301)를 착용한 사용자의 시선의 제1 방향 및 시각적 객체의 시선의 제2 방향에 기반하여, 웨어러블 장치(1301)를 착용한 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 적어도 하나의 센서(1620)를 이용하여, 웨어러블 장치(1301)를 착용한 사용자의 위치를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 적어도 하나의 센서(1620)를 이용하여, 시각적 객체에 대응하는 발화자의 위치를 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 상기 사용자의 위치로부터 상기 제1 방향에 상기 발화자가 위치되는 제1 조건 및 상기 발화자의 위치로부터 상기 제2 방향에 상기 사용자가 위치되는 제2 조건 각각의 충족 여부를 식별할 수 있다. 예를 들면, 제1 조건은, 상기 사용자의 시선이 상기 발화자를 향하는 것으로 참조될 수 있다. 예를 들면, 제2 조건은, 상기 발화자의 시선이 상기 사용자를 향하는 것으로 참조될 수 있다.
웨어러블 장치(1301)는, 상기 제1 조건의 충족 및 상기 제2 조건의 충족에 기반하여, 상기 사용자와 상기 발화자가 서로를 보고 있음을 식별할 수 있다. 웨어러블 장치(1301)는, 상기 제1 조건의 실패(또는 불충족) 및/또는 상기 제2 조건의 실패(또는 불충족)에 기반하여, 상기 사용자와 상기 발화자가 서로를 보고 있지 않음을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있음을 식별하는 것에 기반하여, 동작 2059를 실행할 수 있다. 예를 들면, 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있지 않음을 식별하는 것에 기반하여, 동작 2061를 실행할 수 있다.
동작 2059에서, 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있음을 식별하는 것에 기반하여, 발화자의 보이스 신호로부터 생성된 텍스트를, 화면에 중첩하여, 디스플레이(1350)를 통해 표시할 수 있다. 예를 들면, 텍스트는, 보이스 신호를 나타낼 수 있다. 예를 들면, 보이스 신호로부터 생성된 텍스트는, 보이스 신호로부터 변환된 텍스트를 포함할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호에 STT(speech to text)를 적용함으로써, 텍스트를 생성할 수 있다. 예를 들면, 텍스트는 시각적 객체와 연계로 표시될 수 있다. 예를 들면, 텍스트는 시각적 객체 옆에서 표시될 수 있다.
제한되지 않는 예로, 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있음을 식별하는 것에 기반하여, 보이스 신호를 나타내는 오디오 신호를, 스피커(예: 스피커(1355))를 통해 출력할 수 있다.
동작 2061에서, 웨어러블 장치(1301)는 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있지 않음을 식별하는 것에 기반하여, 발화자의 보이스 신호로부터 생성된 텍스트를 디스플레이(1350)를 통해 표시하는 것을 삼가하거나 중지하거나 스킵하거나 바이패스할 수 있다. 예를 들면, 웨어러블 장치(1301)는 동작 2051의 화면을 디스플레이(1350)를 통해 표시하는 것을 유지할 수 있다.
제한되지 않는 예로, 웨어러블 장치(1301)는 사용자와 시각적 객체에 대응하는 발화자가 서로를 보고 있지 않음을 식별하는 것에 기반하여, 보이스 신호를 나타내는 오디오 신호를, 스피커(1355)를 통해 출력하는 것을 삼가하거나 중지하거나 스킵하거나 바이패스할 수 있다.
도 21은 제1 언어로 발화된 오디오 신호에 대하여 제2 언어로 응답을 제공하는 웨어러블 장치의 예를 도시한다. 도 21에서 예시되는 웨어러블 장치(1301)의 동작들은, 프로세서(예: 프로세서(1610))에 의해 실행되거나 수행되거나 제어될 수 있다.
도 21을 참고하면, 웨어러블 장치(1301)는 마이크로폰(예: 마이크로폰(1365))을 통해 오디오 신호를 획득할 수 있다. 예를 들면, 오디오 신호는, 보이스 신호(2110) 및/또는 노이즈를 포함할 수 있다. 예를 들면, 보이스 신호(2110)는 제1 언어에 의해 나타내어질 수 있다. 예를 들면, 보이스 신호(2110)는 제1 언어로 된 발화를 포함할 수 있다. 웨어러블 장치(1301)는, 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 통해 센싱 데이터를 획득할 수 있다. 예를 들면, 센싱 데이터는, 이미지 센서(예: 이미지 센서(1621))를 통해 획득된 이미지를 포함할 수 있다. 예를 들면, 이미지는, 외부 환경(또는 물리 환경)이 촬영된 이미지일 수 있다. 예를 들면, 이미지는, 발화자(1802)에 대응하는 시각적 객체(2121)를 포함할 수 있다. 예를 들면, 이미지는, 발화자(1802)의 입술에 대응하는 시각적 객체(2121)의 일부를 포함할 수 있다.
웨어러블 장치(1301)는, 마이크로폰(1365)을 통해 획득된 오디오 신호 및/또는 적어도 하나의 센서(1620)를 통해 센싱 데이터를, 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공할 수 있다. 예를 들면, 웨어러블 장치(1301)는, 제1 멀티모달 모델(350)을 이용하여, 오디오 신호의 SNR이 상대적으로 작은 환경에서도, 보이스 신호(2110)를 식별할 수 있다.
일 실시예에 따르면, 제1 멀티모달 모델(350)은, 하나 이상의 멀티모달 모델을 포함할 수 있다. 예를 들면, 하나 이상의 멀티모달 모델 각각은 다른 언어를 이용하여 학습될 수 있다. 예를 들면, 하나 이상의 멀티모달 모델 각각은, 각각의 다른 언어를 발화하는 입술을 표현하는 이미지를 이용하여 학습된 모델일 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 보이스 신호(2110)의 언어가 제1 언어임을 식별할 수 있다. 예를 들면, 웨어러블 장치(1301)는 하나 이상의 멀티모달 모델들 중에서 제1 언어에 대응하는 멀티모달 모델을 결정할 수 있다. 예를 들면, 제1 언어가 영어인 경우, 웨어러블 장치(1301)는 영어를 이용하여 학습된 멀티모달 모델을 결정할 수 있다. 예를 들면, 제1 언어가 아랍어인 경우, 웨어러블 장치(1301)는 아랍어를 이용하여 학습된 멀티모달 모델을 결정할 수 있다.
일 실시예에 따르면, 웨어러블 장치(1301)는 화면(2120)을 디스플레이(예: 디스플레이(1350))를 통해 표시할 수 있다. 예를 들면, 웨어러블 장치(1301)는 보이스 신호(2110)에 대응하는 텍스트(2125)를 화면(2120) 내의 영역(2125)을 통해 표시할 수 있다. 예를 들면, 텍스트(2125)는 제2 언어에 의해 나타내어질 수 있다. 예를 들면, 보이스 신호(2110)의 컨텍스트와 텍스트(2125)의 컨텍스트는 실질적으로 동일할 수 있다. 예를 들면, 텍스트(2125)는 보이스 신호(2110)를 나타내는 텍스트가 제1 언어에서 제2 언어로 번역된 것일 수 있다. 일 실시예에 따르면, 웨어러블 장치(1301)는 보이스 신호(2110)에 대응하는 오디오 신호(2130)를 스피커(예: 스피커(1355))를 통해 출력할 수 있다. 예를 들면, 오디오 신호(2130)의 언어는 제2 언어일 수 있다. 예를 들면, 오디오 신호(2130)는 보이스 신호(2110)가 제1 언어에서 제2 언어로 번역된 것을 나타낼 수 있다. 예를 들면, 제2 언어는, 보이스 신호(2110)의 제1 언어와 다를 수 있다. 예를 들면, 제2 언어는, 웨어러블 장치(1301)에 설정된 언어일 수 있다. 예를 들면, 지정된 시간 내에 사용자(1801)에 의해 가장 많이 발화된 언어가 제2 언어이기 때문에, 제2 언어가 설정될 수 있다. 제한되지 않는 예로, 제2 언어는 사용자(1801)에 의해 설정될 수 있다.
본 개시에서 이루고자 하는 기술적 과제는 이상에서 언급한 기술적 과제로 제한되지 않으며, 언급되지 않은 또 다른 기술적 과제들은 본 개시에 속하는 기술분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
상술한 바와 같은 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 제3 모델(예: 제3 모델(370))을 이용하여, 상기 제2 오디오 신호 및 상기 제2 이미지 중 상기 제2 오디오 신호로부터, 상기 전자 장치가 위치된 환경의 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제1 유형에 대한 제3 데이터, 및 상기 제2 오디오 신호에 대한 제4 데이터를 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제2 오디오 신호에 대한 노이즈 캔슬링을 통해, 제3 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 제2 오디오 신호에 대한 상기 제4 데이터를 상기 제1 데이터로 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록, 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제2 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제2 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 스피커를 통해 상기 오디오 신호를 출력하도록, 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하도록, 상기 전자 장치를 야기할 수 있다.
상술한 바와 같은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 수행되는 방법은, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하는 동작을 포함할 수 있다. 상기 방법은, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 전자 장치 내의 제3 모델(예: 제3 모델(370))을 이용하여, 상기 제2 오디오 신호 및 상기 제2 이미지 중 상기 제2 오디오 신호로부터, 상기 전자 장치가 위치된 환경의 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제1 유형에 대한 제3 데이터, 및 상기 제2 오디오 신호에 대한 제4 데이터를 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제2 오디오 신호에 대한 노이즈 캔슬링을 통해, 제3 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 제2 오디오 신호에 대한 상기 제4 데이터를 상기 제1 데이터로 획득하는 동작을 포함할 수 있다. 상기 방법은, 및 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제2 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제2 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 방법은, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 스피커를 통해 상기 오디오 신호를 출력하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하는 동작을 포함할 수 있다.
상술한 바와 같은, 하나 이상의 프로그램들이 저장된 컴퓨터 판독 가능 저장 매체에 있어서, 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 상기 이미지 센서를 구동하기 위한, 이벤트를 검출하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서를 통해 제1 이미지를 획득하고, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치 내의 제1 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 외부 전자 장치 내의 제2 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 제3 모델(예: 제3 모델(370))을 이용하여, 상기 제2 오디오 신호 및 상기 제2 이미지 중 상기 제2 오디오 신호로부터, 상기 전자 장치가 위치된 환경의 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제1 유형에 대한 제3 데이터, 및 상기 제2 오디오 신호에 대한 제4 데이터를 제4 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제2 오디오 신호에 대한 노이즈 캔슬링을 통해, 제3 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하고, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 제2 오디오 신호에 대한 상기 제4 데이터를 상기 제1 데이터로 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제2 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제2 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 스피커를 통해 상기 오디오 신호를 출력하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
상술한 바와 같은, 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서와 관련된 타이머를 활성화하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 타이머의 만료 전, 상기 구동된 이미지 센서를 통해 상기 제1 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하고, 및 상기 타이머의 상기 만료와 독립적으로 상기 이미지 센서를 구동하는 것을 유지하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 기준 제스쳐를 표현하지 않는 상기 제1 이미지 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 타이머의 상기 만료에 응답하여 상기 이미지 센서를 구동하는 것을 중단하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호 및 상기 제2 이미지 중 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 제3 데이터를 상기 제1 데이터로 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제2 유형에 대한 제4 데이터, 및 상기 제3 데이터를 제3 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제3 오디오 신호에 대한 노이즈 캔슬링을 통해, 제4 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 스피커를 통해 상기 오디오 신호를 출력하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호로부터 획득된 데이터를 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 제2 이미지를 획득하고, 상기 마이크로폰을 통해 제4 오디오 신호를 획득하고, 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제1 데이터, 상기 제2 유형에 대한 제2 데이터, 및 상기 제4 오디오 신호에 대한 제3 데이터를 제3 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제4 오디오 신호에 대한 노이즈 캔슬링을 통해, 제5 오디오 신호에 대한 제4 데이터를 획득하고, 및 상기 제1 데이터 및 상기 제4 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제4 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
상술한 바와 같은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 수행되는 방법은, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하는 동작을 포함할 수 있다. 상기 방법은, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서와 관련된 타이머를 활성화하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 타이머의 만료 전, 상기 구동된 이미지 센서를 통해 상기 제1 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하고, 및 상기 타이머의 상기 만료와 독립적으로 상기 이미지 센서를 구동하는 것을 유지하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 기준 제스쳐를 표현하지 않는 상기 제1 이미지 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 타이머의 상기 만료에 응답하여 상기 이미지 센서를 구동하는 것을 중단하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호 및 상기 제2 이미지 중 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 제3 데이터를 상기 제1 데이터로 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제2 유형에 대한 제4 데이터, 및 상기 제3 데이터를 제3 멀티모달 모델(예: 제4 멀티모달 모델(380))에게 제공함으로써 수행되는 상기 제3 오디오 신호에 대한 노이즈 캔슬링을 통해, 제4 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 방법은, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 스피커를 통해 상기 오디오 신호를 출력하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하는 동작을 포함할 수 있다. 상기 방법은, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호로부터 획득된 데이터를 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 방법은, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 제2 이미지를 획득하고, 상기 마이크로폰을 통해 제4 오디오 신호를 획득하고, 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제1 데이터, 상기 제2 유형에 대한 제2 데이터, 및 상기 제4 오디오 신호에 대한 제3 데이터를 제3 멀티모달 모델에게 제공함으로써 수행되는 상기 제4 오디오 신호에 대한 노이즈 캔슬링을 통해, 제5 오디오 신호에 대한 제4 데이터를 획득하고, 및 상기 제1 데이터 및 상기 제4 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제4 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 방법은, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하는 동작을 포함할 수 있다. 상기 방법은, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하는 동작을 포함할 수 있다.
상술한 바와 같은, 하나 이상의 프로그램들이 저장된 컴퓨터 판독 가능 저장 매체에 있어서, 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 상기 마이크로폰을 통해, 제1 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치 내의 제1 멀티모달 모델(예: 제2 멀티모달 모델(360))의 웨이크-업을 야기하는 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 구동된 이미지 센서를 통해 제1 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치 내의 제2 멀티모달 모델(예: 제1 멀티모달 모델(350))을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서와 관련된 타이머를 활성화하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 타이머의 만료 전, 상기 구동된 이미지 센서를 통해 상기 제1 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하고, 및 상기 타이머의 상기 만료와 독립적으로 상기 이미지 센서를 구동하는 것을 유지하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 기준 제스쳐를 표현하지 않는 상기 제1 이미지 및/또는 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델을 이용하여 식별하는 것에 기반하여, 상기 타이머의 상기 만료에 응답하여 상기 이미지 센서를 구동하는 것을 중단하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 이미지 센서를 통해 제2 이미지를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호 및 상기 제2 이미지 중 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 제3 데이터를 상기 제1 데이터로 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제2 유형에 대한 제4 데이터, 및 상기 제3 데이터를 제3 멀티모달 모델에게 제공함으로써 수행되는 상기 제3 오디오 신호에 대한 노이즈 캔슬링을 통해, 제4 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 스피커를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 스피커를 통해 상기 오디오 신호를 출력하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보 내의 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 전자 장치 내의 상기 모델을 이용하여, 상기 제3 오디오 신호로부터, 상기 전자 장치가 위치된 상기 환경의 상기 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호로부터 획득된 데이터를 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 제2 이미지를 획득하고, 상기 마이크로폰을 통해 제4 오디오 신호를 획득하고, 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제1 데이터, 상기 제2 유형에 대한 제2 데이터, 및 상기 제4 오디오 신호에 대한 제3 데이터를 제3 멀티모달 모델에게 제공함으로써 수행되는 상기 제4 오디오 신호에 대한 노이즈 캔슬링을 통해, 제5 오디오 신호에 대한 제4 데이터를 획득하고, 및 상기 제1 데이터 및 상기 제4 데이터를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게, 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델을 이용하여 획득된, 상기 제4 오디오 신호에 대한 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델에 의해 선택된 이모지 그래픽적 객체를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 이모지 그래픽적 객체를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는, 디스플레이를 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 외부 전자 장치로부터, 상기 전자 장치 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델에 의해 선택된 텍스트를 나타내는 상기 응답 정보를, 상기 통신 회로를 통해, 수신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 응답 정보에 의해 나타내어진 상기 텍스트를 상기 디스플레이를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
상술한 바와 같은, 전자 장치는, 상기 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서를 포함할 수 있다. 상기 전자 장치는, 마이크로폰을 포함할 수 있다. 상기 전자 장치는, 통신 회로를 포함할 수 있다. 상기 전자 장치는, 인스트럭션들(instructions)을 저장하고, 하나 이상의 저장 매체(storage medium)들을 포함하는, 메모리를 포함할 수 있다. 상기 전자 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 마이크로폰(330)을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록 상기 전자 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하도록 상기 전자 장치를 야기할 수 있다.
상술한 바와 같은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 수행되는 방법은, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하는 동작을 포함할 수 있다. 상기 방법은, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 마이크로폰을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하는 동작을 포함할 수 있다. 상기 방법은, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하는 동작을 포함할 수 있다.
상술한 바와 같은, 하나 이상의 프로그램들이 저장된 컴퓨터 판독 가능 저장 매체에 있어서, 상기 하나 이상의 프로그램들은, 전자 장치의 프론트 사이드의 방향을 향하는 이미지 센서, 마이크로폰, 및 통신 회로를 가지는 상기 전자 장치에 의해 실행될 시, 외부 전자 장치와 통화 연결을 위한 이벤트를 검출하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 검출에 따라, 상기 외부 전자 장치와 상기 통화 연결을 수행하는 것에 기반하여, 상기 마이크로폰을 통해 제1 오디오 신호를 획득하고, 및 상기 전자 장치 내의 모델(예: 제3 모델(370))을 이용하여, 상기 마이크로폰을 통해 상기 제1 오디오 신호로부터 상기 전자 장치가 위치된 환경의 유형을 식별하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 마이크로폰을 통해 제2 오디오 신호를 획득하고, 및 상기 제2 오디오 신호를, 상기 통신 회로를 통해, 상기 외부 전자 장치에게 송신하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 전자 장치에 의해 실행될 시, 상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 이미지 센서를 구동하고, 상기 구동된 이미지 센서를 통해 이미지를 획득하고, 상기 마이크로폰을 통해 제3 오디오 신호를 획득하고, 상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 사용자의 입술에 대응하는 상기 이미지 내의 시각적 객체에 대한 제2 데이터를 상기 전자 장치 내의 멀티모달 모델(예: 제1 멀티모달 모델(350))에게 제공함으로써, 상기 제3 오디오 신호 및 상기 시각적 객체에 대한 응답 정보를, 획득하고, 및 상기 응답 정보에 따른 기능을 수행하도록, 상기 전자 장치를 야기하는 인스트럭션들을 포함할 수 있다.
상술한 바와 같은, 웨어러블 장치(예: 웨어러블 장치(1301))는 적어도 하나의 디스플레이(예: 디스플레이(1350))를 포함할 수 있다. 상기 웨어러블 장치는 마이크로폰(예: 마이크로폰(1365))을 포함할 수 있다. 상기 웨어러블 장치는 적어도 하나의 센서(예: 적어도 하나의 센서(1620)를 포함할 수 있다. 상기 웨어러블 장치는, 인스트럭션들(instructions)을 저장하는 하나 이상의 저장 매체(storage medium)들을 포함하는 메모리를 포함할 수 있다. 상기 웨어러블 장치는, 프로세싱 회로(processing circuitry)를 포함하는 적어도 하나의 프로세서를 포함할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하도록 상기 웨어러블 장치를 야기할 수 있다. 상기 인스트럭션들은, 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 시, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하도록 상기 웨어러블 장치를 야기할 수 있다.
상술한 바와 같은, 적어도 하나의 디스플레이(예: 디스플레이(1350)), 마이크로폰(예: 마이크로폰(1365)), 및 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 가지는 웨어러블 장치(예: 웨어러블 장치(1301))에서 수행되는 방법은, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하는 동작을 포함할 수 있다. 상기 방법은, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하는 동작을 포함할 수 있다. 상기 방법은, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하는 동작을 포함할 수 있다.
상술한 바와 같은, 하나 이상의 프로그램들이 저장된 컴퓨터 판독 가능 저장 매체에 있어서, 상기 하나 이상의 프로그램들은, 적어도 하나의 디스플레이(예: 디스플레이(1350)), 마이크로폰(예: 마이크로폰(1365)), 및 적어도 하나의 센서(예: 적어도 하나의 센서(1620))를 가지는 웨어러블 장치(예: 웨어러블 장치(1301))에 의해 실행될 시, 패스-쓰루(pass-through)를 통해 물리 환경에 대응하는 3차원 공간의 적어도 일부를 표현하는 화면을, 상기 적어도 하나의 디스플레이를 통해 표시하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 화면을 표시하는 동안 상기 마이크로폰을 통해 획득된 오디오 신호에 기반하여, 상기 화면 내에 포함된 하나 이상의 시각적 객체들 중에서 상기 오디오 신호에 포함된 보이스 신호에 대응하는 시각적 객체를 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 웨어러블 장치를 착용한 사용자의 시선의 제1 방향 및 상기 보이스 신호에 대응하는 상기 시각적 객체의 시선의 제2 방향을, 상기 적어도 하나의 센서를 통해, 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 사용자의 상기 시선의 제1 방향 및 상기 시각적 객체의 상기 시선의 상기 제2 방향에 기반하여, 상기 사용자와 상기 시각적 객체에 대응하는 발화자가 서로를 보고 있는지 여부를 식별하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다. 상기 하나 이상의 프로그램들은, 상기 웨어러블 장치에 의해 실행될 시, 상기 사용자와 상기 시각적 객체에 대응하는 상기 다른 사용자가 서로를 보고 있음을 식별하는 것에 기반하여, 상기 보이스 신호로부터 생성된 텍스트를, 상기 화면에 중첩하여, 상기 적어도 하나의 디스플레이를 통해 표시하도록, 상기 웨어러블 장치를 야기하는 인스트럭션들을 포함할 수 있다.
본 개시에서 얻을 수 있는 효과는 이상에서 언급한 효과들로 제한되지 않으며, 언급하지 않은 또 다른 효과들은 본 개시가 속하는 기술 분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
하나 이상의 실시예들에 대해, 선행 도면 중 하나 이상에 기재된 구성요소 중 적어도 하나는 본 개시에서 기재된 바와 같은 하나 이상의 동작, 기술, 프로세스 및/또는 방법을 수행하도록 구성될 수 있다. 예를 들어, 선행 도면 중 하나 이상과 관련하여 본 개시에 기술된 프로세서(예: 베이스밴드 프로세서)는 본 개시에서 기재된 하나 이상의 예들에 따라 동작하도록 구성될 수 있다. 다른 예를 들어, 이전 도면 중 하나 이상과 관련하여 위에서 설명된 바와 같은 UE(user equipment), 기지국(base station), 네트워크 요소(network element) 등과 연관된 회로는 여기에 설명된 하나 이상의 예에 따라 작동하도록 구성될 수 있다
위에서 설명된 실시예들 중에서 임의의 것은 달리 명시적으로 언급되지 않는 한 임의의 다른 실시예(또는 실시예의 조합)와 조합될 수 있다. 하나 이상의 구현에 대한 전술한 설명은 예시 및 설명을 제공하지만, 개시된 정확한 형태로 실시예의 범위를 제한하거나 철저하게 하려는 의도는 아니다. 위의 가르침에 비추어 수정 및 변형이 가능하거나 다양한 실시예의 실시로부터 얻어질 수 있다.
본 문서에 개시된 다양한 실시예들에 따른 전자 장치는 다양한 형태의 장치가 될 수 있다. 전자 장치는, 예를 들면, 휴대용 통신 장치(예: 스마트폰), 컴퓨터 장치, 휴대용 멀티미디어 장치, 휴대용 의료 기기, 카메라, 웨어러블 장치, 또는 가전 장치를 포함할 수 있다. 본 문서의 실시예에 따른 전자 장치는 전술한 기기들에 한정되지 않는다.
본 문서의 다양한 실시예들 및 이에 사용된 용어들은 본 문서에 기재된 기술적 특징들을 특정한 실시예들로 한정하려는 것이 아니며, 해당 실시예의 다양한 변경, 균등물, 또는 대체물을 포함하는 것으로 이해되어야 한다. 도면의 설명과 관련하여, 유사한 또는 관련된 구성요소에 대해서는 유사한 참조 부호가 사용될 수 있다. 아이템에 대응하는 명사의 단수 형은 관련된 문맥상 명백하게 다르게 지시하지 않는 한, 상기 아이템 한 개 또는 복수 개를 포함할 수 있다. 본 문서에서, "A 또는 B", "A 및 B 중 적어도 하나", "A 또는 B 중 적어도 하나", "A, B 또는 C", "A, B 및 C 중 적어도 하나", 및 "A, B, 또는 C 중 적어도 하나"와 같은 문구들 각각은 그 문구들 중 해당하는 문구에 함께 나열된 항목들 중 어느 하나, 또는 그들의 모든 가능한 조합을 포함할 수 있다. "제 1", "제 2", 또는 "첫째" 또는 "둘째"와 같은 용어들은 단순히 해당 구성요소를 다른 해당 구성요소와 구분하기 위해 사용될 수 있으며, 해당 구성요소들을 다른 측면(예: 중요성 또는 순서)에서 한정하지 않는다. 어떤(예: 제 1) 구성요소가 다른(예: 제 2) 구성요소에, "기능적으로" 또는 "통신적으로"라는 용어와 함께 또는 이런 용어 없이, "커플드" 또는 "커넥티드"라고 언급된 경우, 그것은 상기 어떤 구성요소가 상기 다른 구성요소에 직접적으로(예: 유선으로), 무선으로, 또는 제 3 구성요소를 통하여 연결될 수 있다는 것을 의미한다.
본 문서의 다양한 실시예들에서 사용된 용어 "모듈"은 하드웨어, 소프트웨어 또는 펌웨어로 구현된 유닛을 포함할 수 있으며, 예를 들면, 로직, 논리 블록, 부품, 또는 회로와 같은 용어와 상호 호환적으로 사용될 수 있다. 모듈은, 일체로 구성된 부품 또는 하나 또는 그 이상의 기능을 수행하는, 상기 부품의 최소 단위 또는 그 일부가 될 수 있다. 예를 들면, 일실시예에 따르면, 모듈은 ASIC(application-specific integrated circuit)의 형태로 구현될 수 있다.
본 문서의 다양한 실시예들은 기기(machine)(예: 전자 장치(1201)) 의해 읽을 수 있는 저장 매체(storage medium)(예: 내장 메모리(1236) 또는 외장 메모리(1238))에 저장된 하나 이상의 명령어들을 포함하는 소프트웨어(예: 프로그램(1240))로서 구현될 수 있다. 예를 들면, 기기(예: 전자 장치(1201))의 프로세서(예: 프로세서(1220))는, 저장 매체로부터 저장된 하나 이상의 명령어들 중 적어도 하나의 명령을 호출하고, 그것을 실행할 수 있다. 이것은 기기가 상기 호출된 적어도 하나의 명령어에 따라 적어도 하나의 기능을 수행하도록 운영되는 것을 가능하게 한다. 상기 하나 이상의 명령어들은 컴파일러에 의해 생성된 코드 또는 인터프리터에 의해 실행될 수 있는 코드를 포함할 수 있다. 기기로 읽을 수 있는 저장 매체는, 비일시적(non-transitory) 저장 매체의 형태로 제공될 수 있다. 여기서, ‘비일시적’은 저장 매체가 실재(tangible)하는 장치이고, 신호(signal)(예: 전자기파)를 포함하지 않는다는 것을 의미할 뿐이며, 이 용어는 데이터가 저장 매체에 반영구적으로 저장되는 경우와 임시적으로 저장되는 경우를 구분하지 않는다.
일실시예에 따르면, 본 문서에 개시된 다양한 실시예들에 따른 방법은 컴퓨터 프로그램 제품(computer program product)에 포함되어 제공될 수 있다. 컴퓨터 프로그램 제품은 상품으로서 판매자 및 구매자 간에 거래될 수 있다. 컴퓨터 프로그램 제품은 기기로 읽을 수 있는 저장 매체(예: compact disc read only memory(CD-ROM))의 형태로 배포되거나, 또는 어플리케이션 스토어(예: 플레이 스토어™)를 통해 또는 두 개의 사용자 장치들(예: 스마트 폰들) 간에 직접, 온라인으로 배포(예: 다운로드 또는 업로드)될 수 있다. 온라인 배포의 경우에, 컴퓨터 프로그램 제품의 적어도 일부는 제조사의 서버, 어플리케이션 스토어의 서버, 또는 중계 서버의 메모리와 같은 기기로 읽을 수 있는 저장 매체에 적어도 일시 저장되거나, 임시적으로 생성될 수 있다.
다양한 실시예들에 따르면, 상기 기술한 구성요소들의 각각의 구성요소(예: 모듈 또는 프로그램)는 단수 또는 복수의 개체들을 포함할 수 있으며, 복수의 개체 중 일부는 다른 구성요소에 분리 배치될 수도 있다. 다양한 실시예들에 따르면, 전술한 해당 구성요소들 중 하나 이상의 구성요소들 또는 동작들이 생략되거나, 또는 하나 이상의 다른 구성요소들 또는 동작들이 추가될 수 있다. 대체적으로 또는 추가적으로, 복수의 구성요소들(예: 모듈 또는 프로그램)은 하나의 구성요소로 통합될 수 있다. 이런 경우, 통합된 구성요소는 상기 복수의 구성요소들 각각의 구성요소의 하나 이상의 기능들을 상기 통합 이전에 상기 복수의 구성요소들 중 해당 구성요소에 의해 수행되는 것과 동일 또는 유사하게 수행할 수 있다. 다양한 실시예들에 따르면, 모듈, 프로그램 또는 다른 구성요소에 의해 수행되는 동작들은 순차적으로, 병렬적으로, 반복적으로, 또는 휴리스틱하게 실행되거나, 상기 동작들 중 하나 이상이 다른 순서로 실행되거나, 생략되거나, 또는 하나 이상의 다른 동작들이 추가될 수 있다.
Claims (15)
- 전자 장치(301)에 있어서,상기 전자 장치(301)의 프론트 사이드의 방향을 향하는 이미지 센서(340);마이크로폰(330);통신 회로(310);인스트럭션들을 저장하는 하나 이상의 저장 매체들을 포함하는 메모리(320); 및프로세싱 회로를 포함하는 적어도 하나의 프로세서(300)를 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 이미지 센서(340)를 구동하기 위한, 이벤트를 검출하고,상기 검출에 기반하여, 상기 이벤트에 따라 구동된 상기 이미지 센서(340)를 통해 제1 이미지를 획득하고, 상기 마이크로폰(330)을 통해 제1 오디오 신호를 획득하고,사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 기준 음성 명령을 포함하는 상기 제1 오디오 신호를 상기 전자 장치(301) 내의 제1 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 외부 전자 장치(302) 내의 제2 멀티모달 모델(360)의 웨이크-업을 야기하는 신호를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고,상기 신호를 송신한 후, 상기 마이크로폰(330)을 통해 제2 오디오 신호를 획득하고, 상기 이미지 센서(340)를 통해 제2 이미지를 획득하고,상기 제2 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고,상기 외부 전자 장치(302)로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제2 멀티모달 모델(360)을 이용하여 획득된, 상기 제2 오디오 신호에 대한 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 따른 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 전자 장치(301) 내의 제3 모델(370)을 이용하여, 상기 제2 오디오 신호 및 상기 제2 이미지 중 상기 제2 오디오 신호로부터, 상기 전자 장치(301)가 위치된 환경의 유형을 식별하고,상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제1 유형에 대한 제3 데이터, 및 상기 제2 오디오 신호에 대한 제4 데이터를 제4 멀티모달 모델(380)에게 제공함으로써 수행되는 상기 제2 오디오 신호에 대한 노이즈 캔슬링을 통해, 제3 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하고,상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여, 상기 제2 오디오 신호에 대한 상기 제4 데이터를 상기 제1 데이터로 획득하고, 및상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게, 송신하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,디스플레이(311)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 외부 전자 장치(302)로부터, 상기 전자 장치(301) 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제2 멀티모달 모델(360)에 의해 선택된 이모지 그래픽적 객체(1021)를 나타내는 상기 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 이모지 그래픽적 객체(1021)를 상기 디스플레이(311)를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,디스플레이(311)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 외부 전자 장치(302)로부터, 상기 전자 장치(301) 내에 미리 저장된 텍스트들 중 상기 제2 멀티모달 모델(360)에 의해 선택된 텍스트(1031)를 나타내는 상기 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 의해 나타내어진 상기 텍스트(1031)를 상기 디스플레이(311)를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,스피커(312)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 수신된 응답 정보에 TTS(text to speech)를 적용함으로써 오디오 신호를 획득하고, 및상기 스피커(312)를 통해 상기 오디오 신호를 출력하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,디스플레이(311)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 제1 데이터 및 상기 제2 데이터에 기반하여 생성된 텍스트를 포함하는 상기 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보 내의 상기 텍스트를 상기 디스플레이(311)를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 1에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 전자 장치(301) 내의 얼굴 검출 모델을 이용하여, 상기 제2 이미지 내의 상기 사용자의 얼굴에 대응하는 다른 시각적 객체를 포함하는 영역에 대한 제3 데이터를 획득하고, 및상기 제3 데이터를 이용하여 상기 제2 데이터를 획득하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 전자 장치(301)에 있어서,상기 전자 장치(301)의 프론트 사이드의 방향을 향하는 이미지 센서(340);마이크로폰(330);통신 회로(310);인스트럭션들을 저장하고, 하나 이상의 저장 매체들을 포함하는 메모리(320); 및프로세싱 회로를 포함하는 적어도 하나의 프로세서(300)를 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 마이크로폰(330)을 통해, 제1 오디오 신호를 획득하고;상기 전자 장치(301) 내의 모델(370)을 이용하여, 상기 제1 오디오 신호로부터 상기 전자 장치(301)가 위치된 환경의 유형을 식별하고;상기 환경의 상기 유형이 제1 유형임을 식별하는 것에 기반하여, 상기 제1 오디오 신호 내에 포함된 기준 음성 명령에 따라, 외부 전자 장치(302) 내의 제1 멀티모달 모델(360)의 웨이크-업을 야기하는 신호를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고; 및상기 환경의 상기 유형이 제2 유형임을 식별하는 것에 기반하여:상기 이미지 센서(340)를 구동하고,상기 구동된 이미지 센서(340)를 통해 제1 이미지를 획득하고,상기 마이크로폰(330)을 통해 제2 오디오 신호를 획득하고, 및사용자의 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 전자 장치(301) 내의 제2 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 8에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여:상기 이미지 센서(340)와 관련된 타이머를 활성화하고,상기 타이머의 만료 전, 상기 구동된 이미지 센서(340)를 통해 상기 제1 이미지를 획득하고,상기 기준 제스쳐를 표현하는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 상기 제1 멀티모달 모델의 웨이크-업을 야기하는 상기 신호를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고, 및 상기 타이머의 상기 만료와 독립적으로 상기 이미지 센서(340)를 구동하는 것을 유지하고, 및상기 기준 제스쳐를 표현하지 않는 상기 제1 이미지 및 상기 기준 음성 명령을 포함하지 않는 상기 제2 오디오 신호를 상기 제2 멀티모달 모델(350)을 이용하여 식별하는 것에 기반하여, 상기 타이머의 상기 만료에 응답하여 상기 이미지 센서(340)를 구동하는 것을 중단하도록,상기 전자 장치(301)를 야기하는.전자 장치(301).
- 청구항 8에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰(330)을 통해 제3 오디오 신호를 획득하고, 상기 이미지 센서(340)를 통해 제2 이미지를 획득하고,상기 제3 오디오 신호를 이용하여 획득된 제1 데이터 및 상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제2 데이터를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고,상기 외부 전자 장치(302)로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델(360)을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 따른 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 10에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 전자 장치(301) 내의 상기 모델(370)을 이용하여, 상기 제3 오디오 신호 및 상기 제2 이미지 중 상기 제3 오디오 신호로부터, 상기 전자 장치(301)가 위치된 상기 환경의 상기 유형을 식별하고,상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호에 대한 제3 데이터를 상기 제1 데이터로 획득하고,상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여, 상기 제2 데이터, 상기 제2 유형에 대한 제4 데이터, 및 상기 제3 데이터를 제3 멀티모달 모델(380)에게 제공함으로써 수행되는 상기 제3 오디오 신호에 대한 노이즈 캔슬링을 통해, 제4 오디오 신호에 대한 데이터를 상기 제1 데이터로 획득하고, 및상기 제1 데이터 및 상기 제2 데이터를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게, 송신하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 10에 있어서,디스플레이(311)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 외부 전자 장치(302)로부터, 상기 전자 장치(301) 내에서 이용가능한 이모지 그래픽적 객체들 중 상기 제1 멀티모달 모델(360)에 의해 선택된 이모지 그래픽적 객체(1021)를 나타내는 상기 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 이모지 그래픽적 객체(1021)를 상기 디스플레이(311)를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 10에 있어서,디스플레이(311)를 더 포함하고,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 외부 전자 장치(302)로부터, 상기 전자 장치(301) 내에 미리 저장된 텍스트들 중 상기 제1 멀티모달 모델(360)에 의해 선택된 텍스트(1031)를 나타내는 상기 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 의해 나타내어진 상기 텍스트(1031)를 상기 디스플레이(311)를 통해 표시함으로써 상기 응답 정보에 따른 상기 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 8에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여 상기 신호를 송신한 후, 상기 마이크로폰(330)을 통해 제3 오디오 신호를 획득하고,상기 전자 장치(301) 내의 상기 모델을 이용하여, 상기 제3 오디오 신호로부터, 상기 전자 장치(301)가 위치된 상기 환경의 상기 유형을 식별하고,상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제1 유형임을 식별하는 것에 기반하여, 상기 제3 오디오 신호로부터 획득된 데이터를 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게 송신하고,상기 외부 전자 장치(302)로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델(360)을 이용하여 획득된, 상기 제3 오디오 신호에 대한 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 따른 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
- 청구항 14에 있어서,상기 인스트럭션들은, 상기 적어도 하나의 프로세서(300)에 의해 개별적으로 또는 집합적으로 실행될 시,상기 제3 오디오 신호로부터, 상기 환경의 상기 유형이 상기 제2 유형임을 식별하는 것에 기반하여:상기 이미지 센서(340)를 구동하고,상기 구동된 이미지 센서(340)를 통해 제2 이미지를 획득하고,상기 마이크로폰(330)을 통해 제4 오디오 신호를 획득하고,상기 사용자의 입술에 대응하는 상기 제2 이미지 내의 시각적 객체에 대한 제1 데이터, 상기 제2 유형에 대한 제2 데이터, 및 상기 제4 오디오 신호에 대한 제3 데이터를 제3 멀티모달 모델(380)에게 제공함으로써 수행되는 상기 제4 오디오 신호에 대한 노이즈 캔슬링을 통해, 제5 오디오 신호에 대한 제4 데이터를 획득하고, 및상기 제1 데이터 및 상기 제4 데이터를, 상기 통신 회로(310)를 통해, 상기 외부 전자 장치(302)에게, 송신하고,상기 외부 전자 장치(302)로부터, 상기 신호에 따라 웨이크 업 상태 내에서 있는 상기 제1 멀티모달 모델(360)을 이용하여 획득된, 상기 제4 오디오 신호에 대한 응답 정보를, 상기 통신 회로(310)를 통해, 수신하고, 및상기 응답 정보에 따른 기능을 수행하도록,상기 전자 장치(301)를 야기하는,전자 장치(301).
Applications Claiming Priority (8)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2024-0071970 | 2024-05-31 | ||
| KR20240071970 | 2024-05-31 | ||
| KR20240083583 | 2024-06-26 | ||
| KR10-2024-0083583 | 2024-06-26 | ||
| KR10-2025-0048941 | 2025-04-15 | ||
| KR20250048941 | 2025-04-15 | ||
| KR1020250060054A KR20250172385A (ko) | 2024-05-31 | 2025-05-08 | 멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 |
| KR10-2025-0060054 | 2025-05-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025249802A1 true WO2025249802A1 (ko) | 2025-12-04 |
Family
ID=97871019
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2025/006520 Pending WO2025249802A1 (ko) | 2024-05-31 | 2025-05-14 | 멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025249802A1 (ko) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20170186432A1 (en) * | 2015-12-29 | 2017-06-29 | Google Inc. | Speech Recognition With Selective Use Of Dynamic Language Models |
| KR20190042918A (ko) * | 2017-10-17 | 2019-04-25 | 삼성전자주식회사 | 전자 장치 및 그의 동작 방법 |
| KR20190098859A (ko) * | 2018-02-01 | 2019-08-23 | 삼성전자주식회사 | 컨텍스트에 따라 이벤트의 출력 정보를 제공하는 전자 장치 및 이의 제어 방법 |
| KR20210042520A (ko) * | 2019-10-10 | 2021-04-20 | 삼성전자주식회사 | 전자 장치 및 이의 제어 방법 |
| KR20220045116A (ko) * | 2019-07-11 | 2022-04-12 | 사운드하운드, 인코포레이티드 | 시각 보조 음성 처리 |
-
2025
- 2025-05-14 WO PCT/KR2025/006520 patent/WO2025249802A1/ko active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20170186432A1 (en) * | 2015-12-29 | 2017-06-29 | Google Inc. | Speech Recognition With Selective Use Of Dynamic Language Models |
| KR20190042918A (ko) * | 2017-10-17 | 2019-04-25 | 삼성전자주식회사 | 전자 장치 및 그의 동작 방법 |
| KR20190098859A (ko) * | 2018-02-01 | 2019-08-23 | 삼성전자주식회사 | 컨텍스트에 따라 이벤트의 출력 정보를 제공하는 전자 장치 및 이의 제어 방법 |
| KR20220045116A (ko) * | 2019-07-11 | 2022-04-12 | 사운드하운드, 인코포레이티드 | 시각 보조 음성 처리 |
| KR20210042520A (ko) * | 2019-10-10 | 2021-04-20 | 삼성전자주식회사 | 전자 장치 및 이의 제어 방법 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2025249802A1 (ko) | 멀티모달 모델을 이용하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 | |
| WO2023191314A1 (ko) | 정보를 제공하는 방법 및 이를 지원하는 전자 장치 | |
| WO2025230159A1 (ko) | 가상 공간 내에서 대화를 위한 웨어러블 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 | |
| WO2025258879A1 (ko) | 데이터를 입력으로 인식하는 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2025216422A1 (ko) | 픽셀들의 깊이 값을 설정하기 위한 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2025150689A1 (ko) | 콘텐트를 생성하는 전자 장치 및 방법 | |
| WO2026023845A1 (ko) | 외부 전자 장치와 연결하기 위한 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2024063302A1 (ko) | 개체에 다른 개체와의 인터랙션을 적용하는 가상 공간을 제공하기 위한 방법 및 장치 | |
| WO2025225926A1 (ko) | 아바타의 전역 움직임을 생성하기 위한 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2026034748A1 (ko) | 전자 장치의 데이터를 입력으로 인식하기 위한 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2025239571A1 (ko) | 미디어 콘텐트를 재생하기 위한 장치, 방법, 및 저장 매체 | |
| WO2025023479A1 (ko) | 가상 공간 내에서 어플리케이션에 관한 시각적 객체를 표시하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 | |
| WO2024228472A1 (ko) | 가상 환경 내 아바타를 표시하기 위한 전자 장치 및 방법 | |
| WO2025070993A1 (ko) | 제스처 입력을 위한 웨어러블 장치, 방법, 및 비-일시적 컴퓨터 판독 가능 기록 매체 | |
| WO2026054491A1 (ko) | 헤드 마운트 디스플레이 장치 및 디스플레이 제어 방법 | |
| WO2025150657A1 (ko) | 3인칭 시점의 콘텐트를 제공하기 위한 전자 장치 및 방법 | |
| WO2026059040A1 (ko) | 사용자의 시력에 기반하여 화면을 표시하기 위한 웨어러블 장치, 방법, 및 비-일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2025263831A1 (ko) | 제스처를 이용한 번역 장치, 그 동작 방법과 기록매체 | |
| WO2025023480A1 (ko) | 가상 공간의 전환에 기반하여 화면을 변경하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체 | |
| WO2026095304A1 (ko) | 센서 데이터를 이용하여 텍스트를 요약하기 위한 전자 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2026095355A1 (ko) | 디스플레이를 통해 표시되는 화면을 제어하기 위한 전자 장치, 방법, 및 비-일시적 컴퓨터 판독가능 저장 매체 | |
| WO2025154914A1 (ko) | 소스 객체 및 타겟 객체에 기초하여 추가 객체를 획득하는 방법 및 장치 | |
| WO2025244289A1 (ko) | 사용자 인터페이스를 표시하기 위한 웨어러블 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2024029699A1 (ko) | 가상 오브젝트를 표시하는 웨어러블 전자 장치 및 이의 제어 방법 | |
| WO2026054296A1 (ko) | 웨어러블 장치를 통해 사용자를 인증하는 전자 장치, 방법, 및 비-일시적 컴퓨터 판독 가능 기록 매체 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25816047 Country of ref document: EP Kind code of ref document: A1 |