WO2026005387A1 - 모달리티 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치 - Google Patents
모달리티 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치Info
- Publication number
- WO2026005387A1 WO2026005387A1 PCT/KR2025/008529 KR2025008529W WO2026005387A1 WO 2026005387 A1 WO2026005387 A1 WO 2026005387A1 KR 2025008529 W KR2025008529 W KR 2025008529W WO 2026005387 A1 WO2026005387 A1 WO 2026005387A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- electronic device
- camera
- view
- user
- field
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/18—Eye characteristics, e.g. of the iris
Definitions
- Embodiments disclosed in this document relate to a method for providing a response based on determination of modality and an electronic device therefor.
- a variety of services utilizing generative AI are being developed. Early services supported generating text-based responses or generating related images based on images. With the advancement of generative AI, services supporting various input formats are being introduced. For example, a service is being developed that performs image editing using a specified image and text data specifying the desired modification. By utilizing generative AI services that support various input formats, users can generate results that align with their intended intent.
- An electronic device may include at least one sensor, a camera, a microphone, a memory, and at least one processor communicatively connected to the at least one sensor, the camera, the microphone, and the memory.
- the memory may store instructions that, when individually or in combination, are executed by the at least one processor, cause the electronic device to obtain a voice input from a user using the microphone and determine whether a field of view of the camera and a field of view of the user correspond to each other.
- the instructions when individually or in combination, are executed by the at least one processor, cause the electronic device to generate at least one first prompt based on an image obtained by the camera and the voice input, input the at least one first prompt into an artificial intelligence model to obtain first result data, and provide a response associated with the voice input based on the first result data.
- a method for providing a response based on modality determination may include an operation of acquiring a voice input and an operation of determining whether a field of view of a camera and a field of view of a user correspond to each other.
- the method for providing a response based on modality determination may include an operation of generating at least one first prompt based on an image acquired by the camera and the voice input when the field of view of the camera and the field of view of the user correspond to each other, an operation of acquiring first result data by inputting the at least one first prompt into an artificial intelligence model, and an operation of providing a response associated with the voice input based on the first result data.
- a computer-readable storage medium can store instructions that, when executed by a processor of an electronic device, cause the electronic device to perform a method for providing a response based on the modality determination.
- Figure 1a illustrates examples of electronic devices.
- FIG. 1b illustrates a block diagram of an electronic device according to one embodiment.
- Figure 2 illustrates a block diagram of an artificial intelligence system according to one embodiment.
- Figure 3 illustrates the structure of an artificial intelligence model according to one embodiment.
- Figure 4 is a flowchart of a task performing method according to one embodiment.
- Figure 5 is a flowchart of a method for determining an operation mode according to one embodiment.
- Figure 6a is a flowchart of a method for determining an operation mode according to an embodiment.
- Figure 6b is a flowchart of a method for determining an operation mode according to one embodiment.
- FIG. 7A illustrates a first mode operating environment of an electronic device according to one embodiment.
- FIG. 7b illustrates a second mode operating environment of an electronic device according to one embodiment.
- FIG. 7c illustrates a second mode operating environment of an electronic device according to one embodiment.
- FIG. 7d illustrates a third mode operating environment of an electronic device according to one embodiment.
- FIG. 8b illustrates an electronic device according to one embodiment.
- FIG. 10 is a flowchart of a method for providing a response based on modality determination according to one embodiment.
- FIG. 11 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.
- Figure 1a illustrates examples of electronic devices.
- an electronic device may support multiple modalities.
- the electronic device may be configured to receive and process various types of input.
- “modality” may refer to a channel for interaction between the electronic device and a user.
- voice input and text input via an interface e.g., a virtual keyboard
- voice input and text input via an interface may be considered different modalities.
- “modality” may refer to the format of data input to an artificial intelligence model (e.g., a generative artificial intelligence model).
- a user’s voice input may be converted into text data via STT (speech-to-text) and NLU (natural language understanding) and input to the artificial intelligence model.
- voice input and text input via an interface may be considered the same modality.
- modality may refer to a channel between a user and the electronic device and/or the format of input data to the artificial intelligence model.
- modality may be referred to as ‘input type’.
- the first electronic device (10a) may be smart glasses.
- the first electronic device (10a) may obtain a user's voice input using a microphone.
- the first electronic device (10a) may obtain an image input using a camera.
- the first electronic device (10a) may obtain a user input using an interface (e.g., a button and/or a touchpad).
- the first electronic device (10a) may be configured to obtain a voice input, an image input, and/or a user input, and to process the obtained input.
- the second electronic device (10b) may be a wearable device configured to be attached to the user's body or clothing.
- the second electronic device (10b) may include an attachment structure that can be attached to the user's body or clothing.
- the second electronic device (10b) may obtain input via, for example, a microphone, a camera, and/or an interface.
- the third electronic device (10c) may be a mobile device.
- the third electronic device (10c) may be a handheld device.
- the third electronic device (10c) may, for example, obtain input through a microphone, a camera, and/or an interface.
- the first electronic device (10a), the second electronic device (10b), and the third electronic device (10c) are examples of the electronic devices (10) described with reference to FIG. 2.
- the electronic devices (10) of the present disclosure are not limited to the first, second, and third electronic devices (10a, 10b, 10c) of FIG. 1.
- the electronic devices (10) may be any user devices that support multiple modalities.
- the electronic devices (10) may include a video see-through (VST) device, a watch-type electronic device, a head mounted device (HMD), and/or a vehicle infotainment system.
- FIG. 1b illustrates a block diagram of an electronic device according to one embodiment.
- the electronic device (10) may include a processor (120), a memory (130), a sensor circuit (140), a display (160), a camera (170), an interface (180), and/or a communication circuit (190).
- the electronic device (10) may correspond to the first electronic device (10a), the second electronic device (10b), the third electronic device (10c) of FIG. 1, and/or the electronic device (1100) of FIG. 11.
- the electronic device (10) may include a configuration similar to the electronic device (1100) described below with reference to FIG. 11.
- the processor (120) may correspond to at least one processor (1110) of FIG. 11.
- the memory (130) may correspond to the memory (1120) of FIG. 11.
- the sensor circuit (140) may correspond to the sensor interface (1119) and/or the sensor (1170) of FIG. 11.
- the display (160) may correspond to the display (1140) of FIG. 11.
- the camera (170) may include the image sensor (1150) of FIG. 11.
- the communication circuit (190) may correspond to the communication circuit (1160) of FIG. 11.
- the configuration of the electronic device (10) illustrated in FIG. 1B is exemplary, and the configuration of the electronic device (10) is not limited thereto.
- the electronic device (10) may further include a configuration not illustrated in FIG. 1B (e.g., at least one of the configurations of the electronic device (1100) of FIG. 11).
- the electronic device (10) may not include at least one of the components illustrated in FIG. 1B (e.g., the second microphone (183), the speaker (184), and/or the sensor circuit (140)).
- the processor (120) may be communicatively, electrically, operatively, or functionally connected to the memory (130), the sensor circuit (140), the display (160), the camera (170), the interface (180), and/or the communication circuit (190).
- a component when a component is “operatively” connected to another component, it may mean that the component is connected so as to be able to operate the other component. For example, the component may operate the other component by transmitting a control signal to the other component, either directly or via another component.
- a component when a component is “functionally” connected to another component, it may mean that the component is connected so as to be able to execute a function of the other component. For example, the component may execute a function of the other component by transmitting a control signal to the other component, either directly or via another component.
- the processor (120) may include at least one processor.
- the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and/or a communication processor (CP).
- the processor (120) may include at least one chip or one chipset.
- the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit.
- the processor (120) may be disposed on a substrate (e.g., a printed circuit board) located within the electronic device (10) and may communicate with other components of the electronic device (10) through at least one conductive path formed on the substrate.
- the memory (130) can store instructions. When executed by the processor (120), the instructions can cause the electronic device (10) to perform various operations. For example, the instructions can be individually or collectively executed by at least one processor to cause the electronic device (10) to perform various operations. In various embodiments of the present disclosure, the operation of the electronic device (10) can be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130).
- the memory (130) can be referred to as a hardware component for data storage.
- the sensor circuit (140) may be configured to detect information (e.g., context information) associated with the electronic device (10).
- the context information may include optical information surrounding the electronic device (10), movement information of the electronic device (10), attachment status information of the electronic device (10), and/or information of adjacent objects.
- the sensor circuit (140) may include an image sensor (141), a motion sensor (143), an attachment sensor (145), and/or a proximity sensor (147).
- the configuration of the sensor circuit (140) illustrated in FIG. 1B is exemplary, and embodiments of the present disclosure are not limited thereto.
- the image sensor (141) may be configured to detect optical signals.
- the image sensor (141) may be configured to detect light signals, infrared signals, and/or ultraviolet signals.
- the image sensor (141) may be configured to acquire depth information.
- the image sensor (141) may include a light detection and ranging (LiDAR).
- the motion sensor (143) may be configured to detect movement of the electronic device (10).
- the motion sensor (143) may include an inertial measurement unit (IMU) configured to detect orientation, magnetism, angular rate, and/or specific force of the electronic device (10).
- the motion sensor (143) may include an accelerometer, a gyroscope, and/or a magnetometer.
- the attachment sensor (145) may be configured to detect the attachment status of the electronic device (10).
- the attachment sensor (145) may detect the attachment status of the electronic device (10) by detecting an attachment of the electronic device.
- the attachment sensor (145) may include a Hall sensor configured to detect the magnetic force of the attachment. The operation of the attachment sensor (145) may be described below with reference to FIGS. 8A and 8B .
- the proximity sensor (147) may be configured to detect an external object in proximity to the electronic device (10).
- the electronic device (10) may detect a blockage state of the electronic device (10) using the proximity sensor (147).
- the proximity sensor (147) may include a capacitive sensor, a photoelectric sensor, a touch switch, an illuminance sensor, and/or an inductive sensor.
- the display (160) may include at least one pixel configured to display an image.
- the display (160) may include multiple displays.
- the display (160) may include a left-eye display and a right-eye display.
- the display (160) may include a front display and/or a rear display.
- the display (160) may include at least one of a see-through display, a flexible display, a rollable display, a foldable display, and/or a rigid display.
- the display (160) may include at least one projector for projecting an image.
- the camera (170) may include at least one camera.
- each of the multiple cameras may have at least one different facing direction, magnification, or field of view.
- the electronic device (10) may select a camera for acquiring an image from the multiple cameras.
- the electronic device (10) may select a camera for acquiring an image based on user context information (e.g., utterance and/or movement direction).
- user context information e.g., utterance and/or movement direction
- the electronic device (10) may identify an image including an object corresponding to a user utterance from among the multiple cameras through image recognition.
- the electronic device (10) may use the camera that acquired the identified image to acquire an image corresponding to a voice input.
- the electronic device (10) may select a camera facing the movement direction of the electronic device (10) from among the multiple cameras.
- the electronic device (10) may select one camera from among the multiple cameras based on a user input.
- the interface (180) may include at least one device configured to receive an input.
- the interface (180) may include a touch circuit (181) configured to receive a touch input (e.g., a touch screen display).
- the interface (180) may include at least one microphone (e.g., a first microphone (182) and/or a second microphone (183)) configured to receive a voice input.
- the electronic device (10) may perform beamforming using the first microphone (182) and the second microphone (183) to identify the location of a speaker of the voice input (e.g., a relative direction of the speaker with respect to the electronic device (10)).
- the interface (180) may include at least one device for output.
- the interface (180) may include a haptic module for tactile output, at least one speaker (e.g., speaker (184)) for sound output, and/or an indicator.
- the interface (180) may include any human interface device (HID).
- the interface (180) may include a button (185).
- the processor (120) may be configured to receive input using the interface (180) and process the received input.
- the communication circuit (190) may be configured to perform short-range wireless communication and/or long-range wireless communication.
- the communication circuit (190) may include a network interface card (NIC).
- NIC network interface card
- the processor (120) may communicate with other external electronic devices based on wireless communication and/or wired communication, for example, using the communication circuit (190).
- the processor (120) may communicate with an external device via an IP (Internet Protocol) network, for example, using the communication circuit (190).
- IP Internet Protocol
- the electronic device (10) may obtain an input (e.g., a request to perform a task) and process the input to generate a result. For example, the electronic device (10) may generate a result by processing the input using a generative artificial intelligence model. The electronic device (10) may generate a result using a single-modal input and/or a multi-modal input. The electronic device (10) may generate a result using, for example, an artificial intelligence system (200) described below with reference to FIG. 2. The electronic device (10) may provide a response to the input using the generated result. The electronic device (10) may provide the response using the interface (180) and/or the display (160). The electronic device (10) may transmit information corresponding to the response to an external electronic device using a communication circuit (190), thereby causing the external electronic device (not shown) to provide a response.
- an input e.g., a request to perform a task
- the electronic device (10) may generate a result by processing the input using a generative artificial intelligence model.
- the electronic device (10) may generate
- an electronic device (10) supporting multi-modal input may generate a result by inputting the user input and an image acquired using the camera (170) into an artificial intelligence model. If the user input is unrelated to the image acquired using the camera (170), the correlation between the response and the user input may be reduced. For example, a response unrelated to the user's intention may be provided. Furthermore, if the user specifies the modality through a separate input, the user experience may be degraded. Furthermore, to process multi-modal input, the electronic device (10) may constantly activate the camera (170). In this case, the power consumption of the electronic device (10) may increase.
- the electronic device (10) may determine an input modality based on the reception of a user input. For example, the electronic device (10) may determine the input modality using context information of the electronic device (10) (e.g., information related to the operating environment). The electronic device (10) may obtain the context information using the sensor circuit (140), the camera (170), and/or the interface (180). If processing of a user input by a single modal input is determined, the electronic device (10) may generate a response by processing the user input using an artificial intelligence model. If processing of a user input by a multimodal input is determined, the electronic device (10) may generate a response by processing the user input and context information using an artificial intelligence model. By determining the input modality, the electronic device (10) may increase the correlation between the response and the user input and reduce the power consumption of the electronic device (10).
- context information of the electronic device (10) e.g., information related to the operating environment.
- the electronic device (10) may obtain the context information using the sensor circuit (140), the camera (170), and/or the interface (
- Figure 2 illustrates a block diagram of an artificial intelligence system according to one embodiment.
- the artificial intelligence system (200) may be configured to process received input using an artificial intelligence model and provide a result generated using the artificial intelligence model.
- the artificial intelligence system (200) may be implemented by an electronic device (10).
- Components of the artificial intelligence system (200) may be software modules (e.g., threads, functions, databases, and/or programs) implemented by the electronic device (10) executing instructions stored in the memory (130) using the processor (120).
- a part of the artificial intelligence system (200) may be implemented by an external device (e.g., a server).
- An instruction database (260) and/or an artificial intelligence model database (280) may be implemented by the external device.
- the electronic device (10) may transmit and receive information by communicating with the external device using a communication circuit (190).
- the artificial intelligence system (200) may include an I/O (input/output) interface (210), an artificial intelligence framework (220), a knowledge DB (260), an application (270), an artificial intelligence model DB (280), and/or an operation model determination module (290).
- I/O input/output
- the configurations of the artificial intelligence system (200) illustrated in FIG. 2 are examples, and at least some of the configurations may be implemented as a single software module.
- the I/O (input/output) interface (210) may provide user input and/or context information to the artificial intelligence framework (220) and/or the action model determination module (290).
- the user input may include a voice input (e.g., natural language input) obtained using the first microphone (182) and/or the second microphone (183).
- the user input may include a verbal input and/or a non-verbal input (e.g., an input for selecting a menu).
- the user input may include a text input obtained by performing STT for the voice input, an input entered through a button (185), a keyboard, and/or a text input obtained through an interface of a virtual keyboard.
- the user input may include an input obtained from any human interface device (HID) or external device communicatively connected to the electronic device (10).
- the user input may include visual information obtained using a camera (170) and/or an image sensor (141).
- the user input may include each of the above-described pieces of information or any combination of the above-described pieces of information.
- Context information may include information related to the environment of the electronic device (10).
- the context information may include a battery SoC (state of charge) of the electronic device (10), background sound acquired using the first microphone (182) and/or the second microphone (183), movement information acquired using the motion sensor (143), attachment information acquired using the attachment sensor (145), and/or shielding status information of the electronic device (10).
- the shielding status information may be acquired by detecting an external object using at least one of the camera (170), the communication circuit (190), the image sensor (141), and/or the proximity sensor (147).
- the context information may include a surrounding image of the electronic device (10) acquired using the camera (170) and/or the image sensor (141).
- the context information may include information on an application currently running on the electronic device (10) and/or location information of the electronic device (10).
- the context information may include each of the above-described pieces of information or any combination of the above-described pieces of information.
- the I/O (input/output) interface (210) may be configured to output results generated by an artificial intelligence model.
- the I/O interface (210) may receive results from the AI framework (220).
- the I/O interface (210) may output the results in the form of natural language or content in any form.
- the I/O interface (210) may visually output the results using the display (160).
- the I/O interface (210) may output the results as audio content using the speaker (184).
- the I/O interface (210) may transmit the results to an external device using the communication circuit (190), thereby causing the external device to output the results.
- the artificial intelligence framework (220) may be configured to receive user input and/or context information from the I/O interface (210), and control components to perform an action corresponding to an intention (e.g., performing a task) corresponding to a user query based on the user query included in the user input.
- the artificial intelligence framework (220) may include a prompt manager (230), an application manager (240), and/or an output manager (250).
- the prompt manager (230) can generate prompts for input into an artificial intelligence model from user input and/or context information.
- the prompt manager (230) can process the user input and/or context information into a form of prompts for input into a large language model (LLM), a large multi-modal model (LMM), and/or a large vision model (LVM).
- LLM large language model
- LMM large multi-modal model
- LMM large vision model
- the prompt manager (230) can generate prompts using text information, feature information, and/or object information extracted from the voice input.
- the prompt manager (230) can utilize a trained machine learning algorithm or an artificial intelligence neural network to generate prompts.
- the knowledge database (260) can store history information related to the electronic device (10). For example, the knowledge database (260) can store user preference data, a prompt library, and/or prompt example data generated based on user input.
- the prompt manager (230) can generate prompts from user input and/or context information using data stored in the knowledge database (260).
- the prompt manager (230) can transmit the generated prompts to an artificial intelligence model (e.g., LLM, LVM, or LMM).
- an artificial intelligence model e.g., LLM, LVM, or LMM
- the application manager (240) may provide an interface between external components (e.g., knowledge DB (260), application (270), artificial intelligence model DB (280), and/or operation mode determination module (290)) and the artificial intelligence framework (220). For example, the application manager (240) may establish a channel between the external components and the artificial intelligence framework (220) using an application programming interface (API) or a plug-in.
- external components e.g., knowledge DB (260), application (270), artificial intelligence model DB (280), and/or operation mode determination module (290)
- API application programming interface
- the application manager (240) can obtain the additional information by communicating with the application (270).
- the application manager (240) can provide a notification requesting additional information using the application (270) and can receive the additional information from the application (270).
- the application manager (240) can obtain the additional information by accessing an external database (e.g., a knowledge database (260).
- the application manager (240) can transmit the additional information to the prompt manager (230) and/or the artificial intelligence model.
- user input may include an intent to perform a specified task.
- the output generated by the AI model may include an action or a sequence of actions for performing the task corresponding to the intent.
- the application manager (240) may execute the action or sequence of actions included in the output using at least one application or service associated with the task.
- the output manager (250) can perform fine-tuning on the output from the AI model.
- fine-tuning may include tuning the output based on content policies (e.g., user terms of use, harmfulness policies, or prohibited content policies) and/or relevance (e.g., relevance between user input and the output).
- content policies e.g., user terms of use, harmfulness policies, or prohibited content policies
- relevance e.g., relevance between user input and the output.
- the output manager (250) may determine whether the output contains prohibited content (e.g., racially or politically biased content). For example, the output manager (250) may determine whether the output contains harmful content (e.g., content requiring age verification or content that violates public order and morals). For example, the output manager (250) may determine whether content that violates the Terms of Service exists. If prohibited content, harmful content, and/or content that violates the Terms of Service exists, the output manager (250) may exclude the content from the output or replace the content with other content. In one example, if prohibited content, harmful content, and/or content that violates the Terms of Service exists, the output manager (250) may provide a notification informing the user that the content cannot be provided. In one example, the output manager (250) may provide guidance information to the user to prevent prohibited content, harmful content, and/or content that violates the Terms of Service from being output.
- prohibited content e.g., racially or politically biased content
- the output manager (250) may determine whether the output contains harmful content (e
- the output manager (250) can verify the correlation between the output and the user input (e.g., the intent of the user input). For example, the output manager (250) can utilize an artificial intelligence model to extract information about the output. The output manager (250) can identify the degree of correlation between the extracted information and the user input by comparing the extracted information and the user input. If the degree of correlation is below a threshold, the output manager (250) can perform additional actions. For example, the output manager (250) can modify the user input and provide the modified user input to the prompt manager (230). The artificial intelligence system (200) can process the prompt generated based on the modified user input using the artificial intelligence model, thereby generating an output that matches the intent of the user input.
- the output manager (250) can process the prompt generated based on the modified user input using the artificial intelligence model, thereby generating an output that matches the intent of the user input.
- the application (270) may include any application installed on the electronic device (10).
- the application (270) may function as a user interface (e.g., front-end) and/or a client-end for the artificial intelligence framework (220).
- the application (270) may provide a user interface for obtaining input.
- the application (270) may be configured to process actions directed by the application manager (240).
- the artificial intelligence model DB (280) can store at least one artificial intelligence model.
- the at least one artificial intelligence model can include an LMM, an LVM, and/or an LMM.
- the artificial intelligence model can include at least one generative artificial intelligence model.
- the artificial intelligence model can include an artificial intelligence model trained to generate images and/or language.
- the artificial intelligence model can include a generative adversarial network (GAN), a variational autoencoder (VAE), and/or a diffusion-based generative model for generating images.
- the diffusion-based generative model can utilize a VAE and a transformer structure.
- the artificial intelligence model can include a model (e.g., CHAT-GPT 3, CHAT-GPT 4) trained to generate statistical output values based on input values for generating text data.
- the operation mode determination module (290) may be configured to determine the operation mode of the electronic device (10) based on user input and/or context information. For example, the operation mode determination module (290) may determine the operation mode of the electronic device (10) based on a field of view (FoV) of the camera (170), an image acquired using the camera (170), information acquired using the sensor circuit (140), and/or information acquired through the interface (180).
- FoV field of view
- the operation mode determination module (290) can determine the operation mode based on the direction of the user's gaze and the direction of the camera's gaze (170). For example, the operation mode determination module (290) can determine the operation mode based on whether the electronic device (10) is shielded. The operation mode determination module (290) can determine the operation mode based on the attachment state of the attachment of the electronic device (10).
- the operation mode determination method of the operation mode determination module (290) is described in more detail in the examples described below with respect to FIGS. 4 to 10.
- the operation mode determination module (290) can control the input to be used according to the determined operation mode.
- the first mode and the second mode may be modes that support multi-modal input.
- the third mode may be a mode that supports single-modal input.
- determining the operation mode may be referred to as determining the input modality or input type.
- the operation mode determination module (290) may control the I/O interface (210) to obtain input according to the determined operation mode. For example, in the case of an operation mode that only uses user input, the operation mode determination module (290) may control the I/O interface (210) to obtain only user input. In this case, a module for obtaining context information (e.g., a sensor circuit (140) and/or a camera (170)) may not be activated.
- a module for obtaining context information e.g., a sensor circuit (140) and/or a camera (170) may not be activated.
- the operation mode determination module (290) can control the prompt manager (230) to use only inputs according to the determined operation mode. For example, in the case of an operation mode that uses only user input, the operation mode determination module (290) can control the prompt manager (230) to use only the user input among the acquired user input and context information.
- Figure 3 illustrates the structure of an artificial intelligence model according to one embodiment.
- the artificial intelligence model DB (280) of FIG. 2 may include at least one artificial neural network model (300).
- the artificial neural network model (300) may include a plurality of hidden layers (320).
- the hidden layers (320) are positioned between the input layer (310) and the output layer (330), and may include at least one layer learned while transmitting data (x1, x2, x3, ..., xn) (n is an integer greater than or equal to 4) transmitted from the input layer (310) to the output layer (330).
- the hidden layers (320) may include a first hidden layer (320-1), a second hidden layer (320-2), and third and Mth hidden layers (320-M) (M is an integer greater than or equal to 3).
- the number of hidden layers illustrated in FIG. 3 is an example, and embodiments of the present disclosure are not limited thereto.
- the artificial neural network model (300) may include more hidden layers than the number of hidden layers illustrated in FIG. 3, or may include fewer hidden layers than the number of hidden layers illustrated in FIG. 3.
- the artificial neural network model (300) may correspond to a multi-layer perception (MLP) model.
- Each of the hidden layers (320) may include a plurality of nodes (N1, N2, ..., Nk). The weight values of each node may be learned using input data.
- the electronic device (10) of FIG. 1B can input the generated prompt as input data to an artificial neural network model (300).
- the input data is calculated by the artificial neural network model (300), and the artificial neural network model (300) can output output data (Y).
- the artificial neural network model (300) is an example of an artificial intelligence model, and the embodiments of the present disclosure are not limited thereto.
- the term "artificial intelligence model” may include a generative artificial intelligence model.
- the artificial intelligence model may include a large language model (LLM), a large multi-modal model (LMM), and/or a large vision model (LVM).
- Figure 4 is a flowchart of a task performing method according to one embodiment.
- the electronic device (10) may perform a task based on user input. For example, the electronic device (10) may obtain a user input requesting the performance of a task for a specified project.
- the operations described below with reference to FIG. 4 may be referred to as the operations of the electronic device (10) of FIG. 1B.
- the order of the operations described below with reference to FIG. 4 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order than that of FIG. 4, or may be executed substantially simultaneously with other operations of FIG. 4.
- the electronic device (10) may detect a trigger event.
- a trigger event may be referred to as any event that causes the electronic device (10) to initiate performance of a task corresponding to a user's request.
- a trigger event may include a user's request to perform the task (e.g., an input instructing the task), an input subsequent to the task performance request (e.g., a task execution input subsequent to the input instructing the task), or an input preceding the task performance request.
- a trigger event may include a voice input, a touch input, and/or a button input.
- the electronic device (10) may detect a trigger event when a voice input is received from the user using a microphone (e.g., a first microphone (182) and/or a second microphone (183)).
- the electronic device (10) may detect a trigger event when data corresponding to a voice input is received from an external device using a communication circuit (190).
- the electronic device (10) may detect a trigger event when a touch input or a button input is received through
- the user input may be referred to as an input that instructs a task.
- the input that instructs a task may include an intention for a designated task or an input that describes a designated task. If the trigger event is an input that instructs a task, the electronic device (10) may obtain the user input through operation 405. If the trigger event is an input that follows the input that instructs a task, the electronic device (10) may obtain the user input prior to operation 405. If the trigger event is an input that precedes the input that instructs a task, the electronic device (10) may obtain the user input after operation 410.
- the electronic device (10) may identify an operating mode. For example, the electronic device (10) may identify an operating mode based on detection of a trigger event. The electronic device (10) may identify an operating mode when a trigger event is detected. In one example, the electronic device (10) may identify an operating mode using information about the operating mode stored in the memory (130). In this case, the operating mode may be predetermined prior to operation 410. In one example, the electronic device (10) may identify an operating mode by determining an operating mode when a trigger event is detected. A method for determining an operating mode may be described below with reference to FIG. 5.
- the operating mode may include a first mode, a second mode, and a third mode.
- the first mode and the second mode may be referred to as operating modes that support multiple modalities.
- the third mode may be referred to as an operating mode that supports a single modality (e.g., an operating mode that does not support multiple modalities).
- an “operating mode” may be referred to as an input modality or an input type.
- identifying an operating mode of an electronic device (10) may be referred to as identifying an input modality or an input type of the electronic device (10).
- the electronic device (10) can detect user input according to the predetermined operation mode.
- the predetermined operation mode is a first mode or a second mode
- the electronic device (10) can activate multiple components to detect multi-modal user input.
- the electronic device (10) can keep the microphone (e.g., the first microphone (182) and/or the second microphone (183)) and the camera (170) in an activated state.
- the electronic device (10) can obtain multi-modal user input using the microphone and the camera (170).
- the predetermined operation mode is a third mode
- the electronic device (10) can keep the microphone in an activated state to detect a single-modal user input.
- the electronic device (10) can keep the camera (170) in an inactive state or an idle state (e.g., a state in which image processing is not performed).
- the electronic device (10) may determine whether the identified operating mode supports multi-modality. If the identified operating mode is the first mode or the second mode, the electronic device (10) may determine that the operating mode supports multi-modality. If the identified operating mode is the third mode, the electronic device (10) may determine that the operating mode does not support multi-modality.
- the electronic device (10) may generate a prompt based on user input and context information. For example, the electronic device (10) may generate a prompt using the prompt manager (230) described above with respect to FIG. 2. For example, the electronic device (10) may obtain a user input from a voice input obtained using a microphone. The electronic device (10) may obtain at least one image obtained using a camera (170) as context information. In this case, the electronic device (10) may generate a prompt (e.g., at least one first prompt) using the voice input and at least one image.
- a prompt e.g., at least one first prompt
- the electronic device (10) may generate a user input-based prompt.
- the electronic device (10) may generate the prompt without using context information.
- the electronic device (10) may obtain a user input from a voice input obtained using a microphone, and may generate a prompt (e.g., a second prompt) from the obtained user input.
- the electronic device (10) can generate a prompt-based result and output the generated result.
- the electronic device (10) can generate the result using the prompt of operation 420 (e.g., at least one first prompt) or the prompt of operation 425 (e.g., a second prompt).
- the electronic device (10) can generate the result by inputting the prompt into at least one artificial intelligence model (e.g., the artificial intelligence model DB (280) of FIG. 2).
- the electronic device (10) can generate the result by adjusting (e.g., fine tuning) the output of the artificial intelligence model.
- the electronic device (10) can output the generated result using the display (160) and/or the speaker (184).
- the electronic device (10) can transmit information about the generated result to an external device using the communication circuit (190) and cause the external device to output the result.
- Figure 5 is a flowchart of a method for determining an operation mode according to one embodiment.
- the electronic device (10) can determine an operating mode of the electronic device (10).
- the operations described below with reference to FIG. 5 may be referred to as operations of the electronic device (10) of FIG. 1B.
- the order of the operations described below with reference to FIG. 5 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order from that of FIG. 5, or may be executed substantially simultaneously with other operations of FIG. 5.
- the electronic device (10) may detect a trigger event.
- the trigger event of operation 505 may be referred to as an event that causes the electronic device (10) to determine an operation mode.
- the trigger event may include a turn-on, activation, boot-up, reset, attachment state change, a specified period, a movement, or reception of a user input of the electronic device (10).
- the electronic device (10) may execute operation 510, which will be described later, upon turn-on, activation, reset, or boot-up.
- the electronic device (10) may execute operation 510 when movement information detected using the motion sensor (143) exceeds a threshold value.
- the electronic device (10) may execute operation 510, which will be described later, when a change in attachment state is detected using the attachment sensor (145).
- the electronic device (10) may execute operation 510 at every specified period.
- the electronic device (10) can execute operation 510 when user input is received.
- the electronic device (10) may determine the operating mode as a mode that supports multi-modality (e.g., a first mode or a second mode) when, for example, the field of view of the camera (170) corresponds to the field of view of the user.
- the electronic device (10) may determine the operating mode as a mode that does not support multi-modality (e.g., a third mode) when the field of view of the camera (170) does not correspond to the field of view of the user.
- the first mode may correspond to an operating mode of the electronic device (10) when the user's field of view and the camera's (170) field of view are oriented in the same direction.
- the first mode at least a portion of the user's field of view and at least a portion of the camera's (170) field of view may overlap each other.
- the camera (170) may capture an image of at least a portion of the field of view in the direction the user is looking or an image larger than the user's field of view.
- the electronic device (10) may obtain visual context information corresponding to the user's field of view using the camera (170). By utilizing the user input and the visual context, the electronic device (10) may provide a response that is more consistent with the intent indicated by the user input.
- the first mode may correspond to the operating mode of the electronic device (10) when the content of the user's speech corresponds to the field of view of the camera (170).
- the electronic device (10) may perform voice recognition on the user's voice input and identify at least one entity (e.g., a keyword referring to an object) from the voice input.
- the electronic device (10) may acquire an image using the camera (170) based on reception of the voice input and perform object recognition on the image. If the object recognized from the image matches the entity identified from the voice input, the electronic device (10) may determine the operating mode of the electronic device (10) as the first mode.
- the second mode may correspond to the operating mode of the electronic device (10) when the user is positioned within the field of view of the camera (170).
- the field of view of the camera (170) and the field of view of the user may face each other, and a portion of the field of view of the camera (170) and the field of view of the user may overlap.
- the electronic device (10) may be mounted at a certain location.
- the electronic device (10) may be placed at a designated location or may be held by the user. With the electronic device (10) mounted at the designated location, the user may look at the camera (170) of the electronic device (10) and speak a voice corresponding to the user input.
- the electronic device (10) may identify an image corresponding to the user from images acquired using the camera (170) and acquire visual context information using the identified image. For example, in the second mode, the user may speak, “How do I look today?” The electronic device (10) may provide a response corresponding to a user input using a user image acquired using a camera (170) and a user input corresponding to a voice input. For example, the electronic device (10) may provide a response to a user input based on the user's appearance, outfit, and/or clothing identified from the user's image.
- the third mode may correspond to an operation mode in which context information is not required for processing user input.
- the user's field of view and the field of view of the camera (170) may not correspond to each other.
- the electronic device (10) may determine whether a camera (170) corresponding to the user's field of view exists among a plurality of cameras based on the user's context information. For example, if there is no camera corresponding to the user's field of view, the electronic device (10) may determine that the field of view of the camera (170) and the user's field of view do not correspond.
- the electronic device (10) in the third mode may process user input without acquiring an image using the camera (170).
- the electronic device (10) may not receive input using the camera (170) or may not process the received input (e.g., by transmitting it to an artificial intelligence model).
- the electronic device (10) may be in a shielded state (e.g., in a state placed in a user's pocket or bag) or attached in a state where the user cannot be identified.
- camera input may include a task for an image acquired using the camera (170).
- the electronic device (10) may output guide information suggesting a change in the position and/or mounting state of the electronic device (10) for changing the operation mode.
- the electronic device (10) can activate the microphone (182, 183) and the camera (170) to receive multi-modal input.
- the electronic device (10) can continuously activate the microphone (182, 183) and the camera (170).
- the electronic device (10) can provide a response using context information acquired prior to the user input.
- the user input can be “Isn’t that David who passed by earlier?”
- the electronic device (10) can provide a response to the user input by analyzing images acquired using the camera (170) prior to the user input.
- the electronic device (10) may constantly activate only the microphones (182, 183) to reduce power consumption. For example, when a user input is detected, the electronic device (10) may activate the camera (170) based on the detection of the user input. In one example, the electronic device (10) may activate the camera (170) when a speech is detected. By activating the camera (170) before the recognition of the voice input is completed, the delay between the user's speech and context information (e.g., an image acquired using the camera (170)) can be reduced.
- context information e.g., an image acquired using the camera (170
- the electronic device (10) may activate the microphone (182, 183) and the camera (170) when a specified input (e.g., input via a button (185), a gesture input, or a touch input) is received.
- the electronic device (10) may determine an operating mode after receiving the specified input.
- the electronic device (10) may determine an operation mode using a plurality of images acquired using the camera (170) and information from the motion sensor (143). For example, if the motion information based on the images acquired using the camera (170) corresponds to the motion information of the electronic device (10) acquired using the motion sensor (143), the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other. In this case, the electronic device (10) may determine the operation mode of the electronic device (10) as the first mode.
- the electronic device (10) may determine the operation mode as the second mode if a user image exists in the images acquired using the camera (170). If the motion information based on the images does not correspond to the motion information of the electronic device (10) obtained using the motion sensor (143), and if there is no user image in the image obtained using the camera (170), the electronic device (10) may determine the operation mode as the third mode.
- the determination of the operation mode using the image and motion sensor (143) may be described later with reference to FIG. 6A.
- the electronic device (10) may determine an operation mode based on attachment information acquired using the attachment sensor (145). For example, the electronic device (10) may determine whether the electronic device (10) is worn by the user. If the electronic device (10) is worn by the user, the electronic device (10) may determine the operation mode as the first mode. If the electronic device (10) is not worn by the user, the electronic device (10) may determine the operation mode as the second mode if a user image exists in an image acquired using the camera (170). If the electronic device (10) is not worn by the user, the electronic device (10) may determine the operation mode as the third mode if a user image does not exist in an image acquired using the camera (170). Determination of an operation mode based on an attachment state may be described later with reference to FIGS. 8A and 8B .
- the electronic device (10) can determine an operation mode based on information acquired using microphones (182, 183).
- the electronic device (10) can perform beamforming using the first microphone (182) and the second microphone (183).
- the electronic device (10) can identify a relative position (e.g., relative direction) of a speaker with respect to the electronic device (10). If the position of the speaker is located at the rear of the electronic device (10) (e.g., facing in the opposite direction from the direction in which the camera (170) of the electronic device (10) faces), the electronic device (10) can determine the operation mode as the first mode. Determining the operation mode using a microphone can be described later with reference to FIG. 8A.
- the electronic device (10) may determine an operating mode based on whether it is shielded.
- the electronic device (10) may use a communication circuit (190), an image sensor (141), and/or a proximity sensor (147) to determine whether the electronic device (10) is in a shielded state (e.g., the electronic device (10) is positioned in a user's pocket or bag). If the electronic device (10) is in a shielded state, the electronic device (10) may determine the operating mode to be a third mode.
- the electronic device may set an operating mode mapped to the designated input.
- a designated input e.g., a button press, a designated gesture, or a designated number of button presses
- Figure 6a is a flowchart of a method for determining an operation mode according to one embodiment.
- the electronic device (10) may determine an operating mode based on movement information.
- the electronic device (10) may determine an operating mode based on information acquired using a camera (170) and information acquired using a sensor circuit (140).
- the electronic device (10) may acquire a plurality of images using the camera (170).
- the electronic device (10) may acquire a plurality of images when a trigger event for determining an operation mode is detected (e.g., operation 505 of FIG. 5).
- the electronic device (10) may be configured to continuously store images for a certain period of time using the camera (170), and when a trigger event is detected, the electronic device (10) may use images within a certain time range from the time at which the trigger event is detected as a plurality of images.
- the electronic device (10) may include a plurality of cameras. In this case, the electronic device (10) may acquire a plurality of images using a camera (170) determined based on context information of a user of the plurality of cameras.
- the electronic device (10) can obtain movement information of the electronic device (10) using the motion sensor (143).
- the electronic device (10) can obtain movement information of the electronic device (10) that is time-synchronized with a plurality of images.
- the electronic device (10) can obtain movement information of the electronic device (10) that is time-synchronized with a plurality of images by substantially simultaneously activating the camera (170) and the motion sensor (143).
- the electronic device (10) may determine the operation mode as the first mode.
- the electronic device (10) may determine whether a user image is identified from a plurality of images (e.g., operation 625). If the user image is identified, the electronic device (10) may determine the operation mode as the second mode. If the user image is not identified, the electronic device (10) may determine the operation mode as the first mode.
- the electronic device (10) may determine the operating mode to be the third mode. In one example, if the electronic device (10) determines that the electronic device (10) is in a shielded state, the electronic device (10) may determine the operating mode to be the third mode. In this case, the electronic device (10) may not perform the operations illustrated in FIG. 6A.
- Figure 6b is a flowchart of a method for determining an operation mode according to one embodiment.
- the operations described below with respect to FIG. 6B may be referred to as operations of the electronic device (10) of FIG. 1B.
- the order of the operations described below with respect to FIG. 6B is merely an example, and embodiments of the present disclosure are not limited thereto.
- at least some operations may be executed differently from the order of FIG. 6B, or may be executed substantially simultaneously with other operations of FIG. 6B.
- the operations described below with respect to FIG. 6B may correspond to operation 510 of FIG. 5.
- the electronic device (10) may acquire at least one image using the camera (170).
- the electronic device (10) may acquire at least one image when a trigger event for determining an operation mode is detected (e.g., operation 505 of FIG. 5).
- the electronic device (10) may be configured to continuously store images for a certain period of time using the camera (170), and when a trigger event is detected, the electronic device (10) may use images within a certain time range from the time at which the trigger event is detected.
- the electronic device (10) may include a plurality of cameras. In this case, the electronic device (10) may acquire at least one image using the camera (170) determined based on context information of a user of the plurality of cameras.
- the electronic device (10) may perform image recognition on at least one image. For example, the electronic device (10) may identify at least one object from the image using any algorithm for image recognition.
- the electronic device (10) may determine whether a voice-responsive object is identified. For example, the electronic device (10) may acquire a voice input through a trigger event (e.g., operation 505 of FIG. 5 ). The electronic device (10) may identify at least one keyword included in the voice input through voice recognition of the voice input. The electronic device (10) may determine that the voice-responsive object is identified if an object corresponding to the identified keyword is identified from at least one image.
- a trigger event e.g., operation 505 of FIG. 5
- the electronic device (10) may identify at least one keyword included in the voice input through voice recognition of the voice input.
- the electronic device (10) may determine that the voice-responsive object is identified if an object corresponding to the identified keyword is identified from at least one image.
- the operating mode may be determined as the first mode in operation 670.
- the electronic device (10) may also determine the operating mode as the second mode. For example, if a voice-responsive object is identified and an image of the user is also identified from the image, the electronic device (10) may determine the operating mode as the second mode.
- the electronic device (10) may determine the operation mode based on whether or not there is eye contact. For example, the electronic device (10) may determine the operation mode according to the method described above with respect to FIG. 6A. Without being limited to operation 675, the electronic device (10) may determine the operation mode according to any operation mode determination method described in the present disclosure.
- methods for providing a response according to the operation mode of the electronic device (10) may be described with reference to FIGS. 7A to 7D . In the examples of FIGS. 7A to 7D , examples are described using the second electronic device (10b) of FIG. 1A , but one of ordinary skill in the art will understand that the embodiments of the present disclosure may also be applied to other types of electronic devices (e.g., the electronic device (10), the first electronic device (10a), or the third electronic device (10c)).
- FIG. 7A illustrates a first mode operating environment of an electronic device according to one embodiment.
- the second electronic device (10b) may be attached to the clothing of the user (701).
- the field of view (705) of the user (701) and the field of view (710) of the second electronic device (10b) e.g., the field of view of the camera of the second electronic device (10b)
- the second electronic device (10b) may acquire an image of at least a portion of an area within the field of view (705).
- the second electronic device (10b) may generate a prompt using information about the surrounding environment (e.g., context information) that the user (701) may be looking at.
- the electronic device (10) may include multiple cameras. In this case, the electronic device (10) can use a camera among multiple cameras determined based on the user's context information.
- a user (701) may utter a voice input such as “It’s really cool.”
- the second electronic device (10b) may analyze an image captured using a camera. By performing image analysis on the captured image, the second electronic device (10b) may identify an image of the sea in the evening. In this case, the second electronic device (10b) may generate a prompt such as “I’m looking at the calm sea with the sunset” as a context-based prompt. In addition, the second electronic device (10b) may generate a prompt such as “A response to my exclamation of “It’s really cool” based on the voice input of the user (701). The second electronic device (10b) may generate an output by inputting the context-based prompt and the voice input-based prompt into an artificial intelligence model.
- the second electronic device (10b) may use information on an image captured using a camera as a context-based prompt.
- the second electronic device (10b) can generate a result by processing an image using the artificial intelligence model together with a voice input-based prompt.
- FIG. 7b illustrates a second mode operating environment of an electronic device according to one embodiment.
- a user (702) may be moving in a second direction (795).
- the user (702) may move a second electronic device (10b) in a first direction (790).
- the motion information of the second electronic device (10b) detected by the motion sensor of the second electronic device (10b) may be a value in which the movement in the second direction (795) is offset from the movement in the first direction (790).
- the motion information based on a plurality of images acquired using a camera of the second electronic device (10b) e.g., a camera selected based on user context information
- the motion information of the second electronic device (10b) and the image-based motion information may not correspond to each other.
- the second electronic device (10b) may determine the operation mode as the second mode.
- the second electronic device (10b) may determine the operating mode of the second electronic device (10b) as the second mode based on the identification of an image corresponding to the user (702) within an image using the camera.
- a user (702) can hold a second electronic device (10b) in his/her hand and alternately capture images of his/her own face and another user's face (not shown) using the second electronic device (10b). While capturing, the user (702) can utter a voice input such as "Who looks older?" In this case, when capturing the user (702), the electronic device (10) can set the operation mode to the second mode and tag the image captured in the second mode with information indicating that the image was captured in the second mode.
- the information indicating that the image was captured in the second mode can include information indicating that the person in the image is the user (702).
- the electronic device (10) can set the operation mode to the first mode and tag the image captured in the first mode with information indicating that the image was captured in the first mode.
- the information indicating that the image was captured in the first mode can include information indicating that the person in the image is not the user (702).
- a user (702) may utter, “My back hurts.”
- a second electronic device (10b) in a second mode may analyze an image of the user (702).
- the second electronic device (10b) may generate a context-based prompt, “User sitting in a chair and reading a book,” from the images of the user (702).
- the second electronic device (10b) may generate a user input-based prompt, “User having a back hurt.”
- the second electronic device (10b) may process the context-based prompt and the user input-based prompt using an artificial intelligence model to generate a result.
- the second electronic device (10b) may provide a response based on the result.
- the second electronic device (10b) may also provide an image of the user associated with the context-based prompt (e.g., an image of the user sitting in a chair).
- FIG. 7c illustrates a second mode operating environment of an electronic device according to one embodiment.
- the second electronic device (10b) in the second mode may provide a proactive suggestion.
- the second electronic device (10b) may provide a proactive suggestion based on changes in the user (703), facial expressions of the user (703), poses of the user (703), actions of the user (703), changes in the surrounding environment of the user (703), or gestures of the user (703) without a voice input from the user (703).
- the electronic device (10) may include multiple cameras. In this case, the electronic device (10) may use a camera determined based on context information of the user among the multiple cameras. For example, the electronic device (10) may use a camera among the multiple cameras that is used to acquire an image in which the user (703) is identified.
- the second electronic device (10b) can analyze an image of the user (703) acquired using a camera.
- the second electronic device (10b) can provide suggestions to the user (703) based on the analyzed information.
- the second electronic device (10b) can provide a suggestion such as, “User, you look tired. How about going to the bedroom?”
- the second electronic device (10b) can provide a notification based on contextual information (e.g., an image acquired using a camera and/or information acquired using a sensor circuit).
- the second electronic device (10b) can be configured to provide suggestions for a specified situation based on user preferences or user settings.
- FIG. 7d illustrates a third mode operating environment of an electronic device according to one embodiment.
- the second electronic device (10b) may be stored in the bag of the user (704).
- the second electronic device (10b) may detect that the second electronic device (10b) is in a shielded state. Based on the detection of the shielded state, the second electronic device (10b) may determine the operating mode of the second electronic device (10b) as a third mode. In the third mode, the second electronic device (10b) may operate based on the voice input of the user (704). For example, the second electronic device (10b) may keep the camera in an inactive or idle state.
- FIG. 8A illustrates an electronic device according to one embodiment.
- the second electronic device (10b) may include a first microphone (882) (e.g., the first microphone (182) of FIG. 1B), a second microphone (883) (e.g., the second microphone (183) of FIG. 1B), a camera (870) (e.g., the camera (170) of FIG. 1B), and a display (860) (e.g., the display (160) of FIG. 1B).
- the second electronic device (10b) may include a main body (801) and an attachment part (802).
- the main body (801) and the attachment part (802) may be coupled based on magnetic force.
- FIG. 8A examples are described focusing on the second electronic device (10b) for convenience of explanation, but those skilled in the art will understand that the same examples may be applied to any electronic device having a similar structure.
- the attachment portion (802) may include at least one magnet for coupling with the main body portion (801).
- the attachment portion (802) may include a battery and/or a processing circuit.
- the attachment portion (802) may supply power to the main body portion (801) by electromagnetically communicating with the main body portion (801).
- the attachment portion (802) may include at least some of the components of the second electronic device (10b).
- the attachment portion (802) may implement the functions of the second electronic device (10b) together with the main body portion (801) through electromagnetic coupling with the main body portion (801). For example, when the attachment portion (802) is coupled with the main body portion (801), the attachment portion (802) may supply power to the main body portion (801) using a battery.
- the second electronic device (10b) may be booted.
- the second electronic device (10b) can perform an operation for determining the operation mode (e.g., operation 510 of FIG. 5) at boot time.
- the second electronic device (10b) can perform an operation for determining the operation mode when a trigger event for determining the operation mode is detected (e.g., operation 505 of FIG. 5).
- the second electronic device (10b) may be mounted on the user's clothing, body, or object in various forms.
- the second electronic device (10b) may detect the attachment state of the attachment portion (802).
- the second electronic device (10b) may detect the attachment state using an attachment sensor (e.g., the attachment sensor (145) of FIG. 1B).
- an attachment sensor e.g., the attachment sensor (145) of FIG. 1B.
- the main body (801) and the attachment portion (802) may be in close contact (e.g., not worn).
- a gap may be generated between the main body (801) and the attachment portion (802) due to another object (e.g., clothes, a bag).
- the second electronic device (10b) can detect the attachment state by detecting the difference in magnetic force generated in the wearing state and the non-wearing state. For example, the second electronic device (10b) can determine the attachment state as the non-wearing state if the detected magnetic force exceeds a threshold value. The second electronic device (10b) can determine the attachment state as the wearing state if the detected magnetic force is a value within a specified range below the threshold value.
- the second electronic device (10b) may include an attachment sensor that may cause mechanical deformation depending on the attachment state.
- the second electronic device (10b) may identify the attachment state by detecting the mechanical deformation.
- the second electronic device (10b) may determine the operation mode based on the attachment state. For example, if the attachment state of the second electronic device (10b) is a worn state, the second electronic device (10b) may determine the operation mode as the first mode. Even if the attachment state is a worn state, if the second electronic device (10b) is shielded, the second electronic device (10b) may determine the operation mode as the third mode. If the attachment state is a non-wearing state, the second electronic device (10b) may determine the operation mode as the second mode or the third mode. For example, the second electronic device (10b) may determine the operation mode according to the method described above with respect to FIGS. 4 to 6b.
- the second electronic device (10b) can perform a life-log operation.
- the life-log operation the second electronic device (10b) can continuously record audio and video (e.g., for a certain period of time) even without user input.
- the second electronic device (10b) can perform the life-log operation according to user settings.
- the second electronic device (10b) can generate a context-based prompt using the acquired image and generate a user input-based prompt based on the user's voice.
- the second electronic device (10b) can use a motion sensor to determine whether the second electronic device (10b) is in a carry state or a stand state.
- the carry state may be referred to as a state in which the second electronic device (10b) is carried by the user. If the second electronic device (10b) detects a movement exceeding a threshold value using the motion sensor in the non-wearing state, the second electronic device (10b) may determine that the second electronic device (10b) is in a carry state.
- the stand state may be referred to as a state in which the second electronic device (10b) is placed in a fixed position. If the second electronic device (10b) does not detect a movement exceeding a threshold value in the non-wearing state, the second electronic device (10b) may be determined to be in a stand state.
- the second electronic device (10b) can determine the operation mode based on whether the user image is recognized. If the user image is recognized from an image acquired using the camera (870), the second electronic device (10b) can determine the operation mode as the first mode. If the user image is not recognized from an image acquired using the camera (870), the second electronic device (10b) can determine the operation mode as the second mode.
- the second electronic device (10b) can determine the operating mode based on the attachment state and the attachment location. For example, the second electronic device (10b) can detect the attached location of the second electronic device (10b) in a worn state. The second electronic device (10b) can detect the attached location of the second electronic device (10b) using a sensor capable of detecting magnetic force, a proximity sensor, and/or a sensor configured to detect a physical connection state. The second electronic device (10b) can determine the operating mode based on the operating mode set for the attached location. In one example, the second electronic device (10b) can determine that the attached location is the user's body. In this case, the second electronic device (10b) can control the settings of the camera (870) to acquire an image in a state of high movement. For example, the second electronic device (10b) can adjust image stabilization (e.g., optical image stabilization, video digital image stabilization), auto focus, white balance, exposure value, and/or shutter speed.
- image stabilization e.g., optical image stabilization, video digital image stabilization
- the second electronic device (10b) when in a stand state, may determine an operation mode (e.g., operation 510 of FIG. 5) based on detection of movement, passage of a specified period of time, or user input. If movement greater than a specified value is detected in the stand state, the second electronic device (10b) may determine an operation mode. If a specified period of time passes without detecting movement greater than a specified value or user input in the stand state, the second electronic device (10b) may determine an operation mode. If a user input is detected in the stand state, the second electronic device (10b) may determine an operation mode. In the stand state, the second electronic device (10b) may wait with the microphones (882, 883) activated without activating the camera (870). If movement is detected, a specified period of time passes, or a user input is received during the standby, the second electronic device (10b) may activate the camera (870) to determine an operation mode.
- an operation mode e.g., operation 510 of FIG. 5
- the second electronic device (10b) may determine an
- the second electronic device (10b) may determine the operation mode based on the position of the speaker. For example, the second electronic device (10b) may receive the user's speech using the first microphone (882) and the second microphone (883). Based on beamforming technology, the second electronic device (10b) may identify the relative position of the speaker (e.g., the user) with respect to the second electronic device (10b). If the speaker is located at the rear of the second electronic device (10b) (e.g., in the opposite direction from the direction in which the camera (870) faces), the second electronic device (10b) may determine the operation mode as the first mode.
- the second electronic device (10b) may determine the operation mode as the first mode.
- the second electronic device (10b) may determine that the field of view of the camera (870) and the field of view of the user correspond to each other when the speaker is located at the rear when the second electronic device (10b) is worn.
- the second electronic device (10b) can determine the operation mode as the third mode.
- the second electronic device (10b) may determine an operating mode based on the speaker's location based on the attachment location of the second electronic device (10b). For example, the second electronic device (10b) may determine an operating mode based on the speaker's location when mounted at a designated location, mounted on a designated accessory, or mounted on a designated body part.
- FIG. 8b illustrates an electronic device according to one embodiment.
- the second electronic device (10b) may have a cylindrical shape. Since the main body (801) and the attachment portion (802) are formed in a cylindrical shape, a user can rotate the attachment portion (802) relative to the main body (801) in a wearing or non-wearing state.
- the second electronic device (10b) may receive user input based on the rotation of the attachment portion (802). For example, the second electronic device (10b) may set an operation mode based on the rotation direction.
- FIG. 9 is a flowchart of a method for providing a response based on modality determination according to one embodiment.
- the electronic device (10) may provide a response based on a user input.
- the electronic device (10) may provide a response using an input based on a determined modality.
- the operations described below with reference to FIG. 9 may be referred to as operations of the electronic device (10) of FIG. 1B.
- the order of the operations described below with reference to FIG. 9 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order than that of FIG. 9, or may be executed substantially simultaneously with other operations of FIG. 9.
- the electronic device (10) may include at least one sensor (e.g., a sensor circuit (140), a camera (170), a microphone (e.g., a first microphone (182) and/or a second microphone (183)), a memory (130), and at least one processor (e.g., a processor (120).
- the at least one processor may be communicatively connected to at least one sensor, a camera (170), a microphone, and a memory (130).
- the memory (130) may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device (10) to perform the operations described below.
- the electronic device (10) can obtain a voice input.
- the electronic device (10) can obtain a voice input from a user using a microphone.
- the electronic device (10) may determine whether the field of view of the camera (170) corresponds to the field of view of the user. For example, the electronic device (10) may identify the relative position of the user with respect to the electronic device (10) from a voice input acquired using a plurality of microphones (e.g., a first microphone (182) and a second microphone (183)). If the relative position is located in the opposite direction to the direction corresponding to the field of view of the camera (170) (e.g., toward the rear of the electronic device (10)), the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other.
- a plurality of microphones e.g., a first microphone (182) and a second microphone (183)
- the electronic device (10) may acquire movement information of the electronic device (10) using the motion sensor (143) and acquire a plurality of images using the camera (170).
- the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other.
- the electronic device (10) may acquire at least one image using the camera (170) based on a voice input.
- the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other.
- the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user do not correspond to each other.
- the camera (170) may include a camera selected from a plurality of cameras.
- Operation 910 may correspond to an operation mode determination operation.
- the electronic device (10) may determine the operation mode according to the operation mode determination operation described above with reference to FIGS. 4 to 8B. For example, if the operation mode is determined to be the first mode or the second mode, the electronic device (10) may determine that the field of view of the camera (170) corresponds to the user's field of view. For example, if the operation mode is determined to be the third mode, the electronic device (10) may determine that the field of view of the camera (170) does not correspond to the user's field of view.
- the electronic device (10) can perform operation 910 based on a trigger event. For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is a wearing state, the electronic device (10) can determine whether the field of view of the camera (170) corresponds to the field of view of the user. For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is a non-wearing state, the electronic device (10) can acquire at least one image using the camera (170). If an image corresponding to the user is identified from the at least one image, the electronic device (10) can determine that the field of view of the camera (170) and the field of view of the user correspond.
- a trigger event For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is a wearing state, the electronic device (10) can determine whether the field of view of the camera (170) corresponds to the field of view of the user. For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is
- the electronic device (10) may include a main body (e.g., main body (801) of FIGS. 8A and 8B) and an attachment portion (e.g., attachment portion (802) of FIGS. 8A and 8B).
- the attachment portion may be configured to be coupled to the main body based on magnetic force.
- the attachment sensor (145) may be configured to detect the attachment state of the electronic device (10) based on magnetic force.
- the operating mode of the electronic device (10) may be determined prior to acquisition of the voice input (e.g., operation 905). In this case, operation 910 may be omitted from the method for providing a response based on modality determination.
- the electronic device (10) may generate a prompt using a multi-modal input or a single-modal input depending on the determined operating mode.
- the electronic device (10) may perform operations according to reference point A. Operations according to reference point A may be described later with reference to FIG. 10.
- the electronic device (10) may generate at least one prompt (e.g., at least one first prompt) based on an image (e.g., an image acquired using the camera (170)) and a voice input.
- the at least one prompt may include a user input-based prompt and a context information-based prompt.
- the electronic device (10) may generate a user input-based prompt based on a voice input.
- the electronic device (10) may generate a context information-based prompt (e.g., text data generated by recognizing the image or the image data itself) using an image.
- the electronic device (10) may perform image recognition on the acquired image and generate at least one prompt based on a combination of a result of the image recognition (e.g., text data) and a voice input (e.g., text data recognized from the voice input).
- the electronic device (10) may generate at least one prompt according to operation 420 of FIG. 4.
- the electronic device (10) may determine the operating mode as a multi-modal operating mode (e.g., the first mode or the second mode) even when the field of view of the camera (170) does not correspond to the field of view of the user (e.g., operation 910-NO).
- the electronic device (10) may determine the operating mode according to the method described above with respect to FIG. 6B.
- the operating mode when the field of view of the camera (170) and the field of view of the user do not correspond may be referred to as a separate fourth mode.
- the fourth mode may be referred to as a multi-modal mode that utilizes both voice input and camera input.
- the electronic device (10) may generate a prompt that includes a description of an image related to the user's utterance.
- the electronic device (10) may generate a response by inputting the prompt into a generative artificial intelligence model and provide the generated response.
- the electronic device (10) can obtain result data (e.g., first result data).
- the electronic device (10) can obtain the result data by inputting at least one generated prompt into an artificial intelligence model (e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2).
- an artificial intelligence model e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2.
- the electronic device (10) may provide a response associated with the voice input based on the result data.
- the electronic device (10) may provide the response according to the methods described above with respect to the I/O interface (210) of FIG. 2.
- the electronic device (10) may perform fine tuning on the result data.
- FIG. 10 is a flowchart of a method for providing a response based on modality determination according to one embodiment.
- the electronic device (10) may provide a response according to the second mode or the third mode.
- the operations described below with respect to FIG. 10 may be referred to as operations of the electronic device (10) of FIG. 1B .
- the order of the operations described below with respect to FIG. 10 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 10 , or may be executed substantially simultaneously with other operations of FIG. 10 .
- the electronic device (10) may perform the operations of FIG. 10 .
- the electronic device (10) may generate at least one prompt (e.g., a second prompt) based on a voice input (e.g., the voice input of operation 905 of FIG. 9 ).
- the electronic device (10) may generate a user input-based prompt based solely on the voice input, without using contextual information (e.g., an image).
- the electronic device (10) may generate a prompt according to operation 425.
- the electronic device (10) can obtain result data (e.g., second result data).
- the electronic device (10) can obtain the result data by inputting at least one generated prompt into an artificial intelligence model (e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2).
- an artificial intelligence model e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2.
- the electronic device (10) may provide a response associated with the voice input based on the result data.
- the electronic device (10) may provide the response according to the methods described above with respect to the I/O interface (210) of FIG. 2.
- the electronic device (10) may perform fine tuning on the result data.
- the operations of the electronic device (10) described above with reference to FIGS. 9 and 10 are part of the embodiments described above with reference to FIGS. 1A to 8B , and a person skilled in the art will understand that the electronic device (10) can perform the embodiments described above with reference to FIGS. 1A to 8B .
- the electronic device (10) does not distinguish between the first mode and the second mode, but the electronic device (10) can perform an operation to determine the first mode or the second mode when the field of view of the camera (170) corresponds to the field of view of the user.
- the prompt generated in the second mode may include information indicating that the image corresponds to the user's image. Accordingly, the result data in the first mode may be different from the result data in the second mode.
- FIG. 11 is a block diagram of an exemplary electronic device (1100) capable of performing the operations described in this document.
- the electronic device (1100) may be one of various forms of electronic devices, such as a notebook (1190), smartphones (1191) having various form factors (e.g., a bar-type smartphone (1191-1), a foldable-type smartphone (1191-2), or a sliderable (or rollable) type smartphone (1191-3)), a tablet (1192), a cellular phone (not shown), and other similar computing devices (not shown).
- a notebook (1190) smartphones (1191) having various form factors (e.g., a bar-type smartphone (1191-1), a foldable-type smartphone (1191-2), or a sliderable (or rollable) type smartphone (1191-3)), a tablet (1192), a cellular phone (not shown), and other similar computing devices (not shown).
- the components, their relationships, and their functions illustrated in FIG. 11 are exemplary only and do not limit the implementations described or claimed in this document.
- the electronic device (1100) may be referred to as a mobile device, a user device, a multi-function device, a portable
- the electronic device (1100) may include components including at least one processor (1110) (hereinafter referred to as processor (1110)), at least one memory (1120) (hereinafter referred to as memory (1120)), at least one display (1140) (hereinafter referred to as display (1140)), at least one image sensor (1150) (hereinafter referred to as image sensor (1150)), at least one communication circuit (1160) (hereinafter referred to as communication circuit (1160)), and/or at least one sensor (1170) (hereinafter referred to as sensor (1170)).
- processor 1110
- memory (1120) hereinafter referred to as memory (1120)
- display (1140) at least one image sensor (1150)
- image sensor (1150) hereinafter referred to as image sensor (1150
- communication circuit (1160) hereinafter referred to as communication circuit (1160)
- sensor (1170) hereinafter referred to as sensor (1170)
- the electronic device (1100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input/output interface).
- PMIC power management integrated circuitry
- audio processing circuitry e.g., audio processing circuitry, an antenna, a rechargeable battery, or an input/output interface.
- some components may be omitted from the electronic device (1100).
- several components can be combined into one component.
- the processor (1110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing.
- the processor (1110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data, etc.) stored in the memory (1120).
- the processor (1110) may include a processor assembly including one or more processing circuits.
- the processor (1110) may include any processing circuit operative to control the performance and operations of one or more components of the electronic device (1100) (e.g., the memory (1120), the display (1140), the image sensor (1150), the communication circuit (1160), and/or the sensor (1170)).
- the processor (1110) e.g., an application processor (AP)
- SoC system on chip
- the processor (1110) may be implemented as multiple cores (or at least one core circuit), multiple chips, or multiple chipsets.
- the processor (1110) may include one or more processing circuits.
- the processor (1110) may include one or more processing circuits configured to individually and/or collectively perform various functions of the present disclosure.
- At least a portion of the processor (1110) may be included in a first chip of the electronic device (1100), and at least another portion of the processor (1110) may be included in a second chip of the electronic device (1100) that is different from the first chip of the electronic device (1100).
- the processor (1110) may include a central processing unit (CPU) (1111), a graphics processing unit (GPU) (1112), a neural processing unit (NPU) (1113), an image signal processor (ISP) (1114), a display controller (1115), a memory controller (1116), a storage controller (1117), a communication processor (CP) (1118), and/or a sensor interface (1119).
- CPU central processing unit
- GPU graphics processing unit
- NPU neural processing unit
- ISP image signal processor
- the processor (1110) may further include other components.
- some components of the processor (1110) may be omitted from the processor (1110).
- some components of the processor (1110) may be included as separate components of the electronic device (1100) outside the processor (1110).
- processors e.g., memory controller (1116)
- memory controller e.g., memory controller (1116)
- interface e.g., available for connection to at least one component of the electronic device (100)
- display e.g., a display
- image sensor e.g., image sensor
- the processor (1110) may cause other components of the electronic device (1100) to perform various operations by executing instructions stored in the memory (1120).
- the CPU (1111) (or central processing circuit) may be configured to control components of the processor (1110) based on the execution of instructions stored in the memory (1120) (e.g., volatile memory (1121) and/or non-volatile memory (1122)).
- the GPU (1112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering).
- the NPU (1113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation).
- AI artificial intelligence
- the ISP (1114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (1150) into a format suitable for a component within the electronic device (1100) or a component of the processor (1110).
- the display controller (1115) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (1111), the GPU (1112), the ISP (1114), or the memory (1120) (e.g., the volatile memory (1121)) into a format suitable for the display (1140).
- the memory controller (1116) (or memory control circuit) may be configured to control reading data from the volatile memory (1121) and writing data to the volatile memory (1121).
- the storage controller (1117) (or storage control circuit) may be configured to control reading data from the non-volatile memory (1122) and writing data to the non-volatile memory (1122).
- the CP (1118) (communication processing circuit) may be configured to process data obtained from a component of the processor (1110) into a format suitable for transmitting to another electronic device via the communication circuit (1160), or to process data obtained from another electronic device via the communication circuit (1160) into a format suitable for processing by the component of the processor (1110).
- the communication circuit (1160) may include one or more communication circuits.
- the sensor interface (1119) (or sensing data processing circuit, sensor hub) may be configured to process data on the state of the electronic device (1100) and/or the state of the surroundings of the electronic device (1100), obtained via the sensor (1170), into a format suitable for the component of the processor (1110).
- the memory (1120) may include one or more storage media (or one or more storage devices).
- the memory (1120) may include a memory assembly including one or more storage media.
- the one or more storage media may include permanent memory (e.g., non-volatile memory (1122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (1121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof.
- the memory (1120) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (1100). As a non-limiting example, the cache memory may be included within the processor (1110).
- the memory (1120) may be fixedly embedded within the electronic device (1100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and/or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (1100).
- SIM subscriber identity module
- SD secure digital
- the memory (1120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and/or applet) software application, and/or any other suitable software applications.
- the one or more software applications may include instructions executable by the processor (1110).
- the memory (1120) may store instructions callable by an application programming interface (API).
- API application programming interface
- the memory (1120) may store instructions within a library.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Ophthalmology & Optometry (AREA)
- Multimedia (AREA)
- Audiology, Speech & Language Pathology (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
적어도 하나의 센서, 카메라, 마이크, 메모리, 및 프로세서를 포함하는 전자 장치가 개시된다. 전자 장치는, 마이크를 이용하여 음성 입력을 획득하고, 카메라의 시야와 사용자의 시야가 대응하는지 결정할 수 있다. 카메라의 시야와 사용자의 시야가 대응하는 경우, 전자 장치는 이미지 및 음성 입력에 기반한 적어도 하나의 제1 프롬프트를 생성하고, 적어도 하나의 제1 프롬프트를 인공지능 모델에 입력함으로써 생성된 제1 결과 데이터에 기반한 응답을 제공할 수 있다.
Description
본 문서에서 개시되는 실시 예들은, 모달리티(modality)의 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치에 관련된다.
생성형 인공지능을 이용한 다양한 서비스들이 개발되고 있다. 초기의 서비스는 텍스트에 기반한 응답을 생성하거나, 이미지에 기반한 관련 이미지를 생성하는 것을 지원하였다. 생성형 인공지능의 발달에 따라서, 다양한 형태의 입력을 지원하는 서비스들이 도입되고 있다. 예를 들어, 지정된 이미지와 지정된 이미지에 대한 수정을 지시하는 텍스트 데이터를 이용하여 이미지에 대한 수정을 수행하는 서비스가 제공된다. 다양한 형태의 입력을 지원하는 생성형 인공지능 서비스들을 이용하여, 사용자는 의도에 부합하는 결과물을 생성할 수 있다.
상술한 정보는 본 개시에 대한 이해를 돕기 위한 목적으로 하는 배경 기술(related art)로 제공될 수 있다. 상술한 내용 중 어느 것도 본 개시와 관련하여 종래 기술(prior art)로서 적용될 수 있는지에 관해서는 어떠한 주장이나 결정이 제기되지 않는다.
본 문서에 개시되는 일 실시 예에 따른 전자 장치는, 적어도 하나의 센서, 카메라, 마이크, 메모리, 및 상기 적어도 하나의 센서, 상기 카메라, 상기 마이크, 및 상기 메모리와 통신적으로(communicatively) 연결된 적어도 하나의 프로세서를 포함할 수 있다. 상기 메모리는 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가, 상기 마이크를 이용하여 사용자로부터 음성 입력을 획득하고, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는지 결정하도록 하는 인스트럭션들을 저장할 수 있다. 상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 경우, 상기 카메라에 의하여 획득된 이미지 및 상기 음성 입력에 기반한 적어도 하나의 제1 프롬프트를 생성하고, 상기 적어도 하나의 제1 프롬프트를 인공지능 모델에 입력함으로써 제1 결과 데이터를 획득하고, 상기 제1 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하도록 할 수 있다.
또한, 본 문서에 개시되는 일 실시 예에 따른 모달리티 결정에 기반한 응답 제공 방법 방법은, 음성 입력을 획득하는 동작 및 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작을 포함할 수 있다. 상기 모달리티 결정에 기반한 응답 제공 방법 방법은, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 경우, 상기 카메라에 의하여 획득된 이미지 및 상기 음성 입력에 기반한 적어도 하나의 제1 프롬프트를 생성하는 동작, 상기 적어도 하나의 제1 프롬프트를 인공지능 모델에 입력함으로써 제1 결과 데이터를 획득하는 동작, 및 상기 제1 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하는 동작을 포함할 수 있다.
본 문서에 개시되는 일 실시 예에 따른 컴퓨터 판독가능 저장 매체는, 전자 장치의 프로세서에 의하여 실행되었을 때에 상기 모달리티 결정에 기반한 응답 제공 방법을 전자 장치가 수행하도록 하는 인스트럭션들을 저장할 수 있다.
도 1a는 전자 장치의 예시들을 도시한다.
도 1b는 일 실시 예에 따른 전자 장치의 블록도를 도시한다.
도 2는 일 실시 예에 따른 인공지능 시스템의 블록도를 도시한다.
도 3은 일 실시 예에 따른 인공지능 모델의 구조를 도시한다.
도 4는 일 실시 예에 따른 태스크 수행 방법의 흐름도이다.
도 5는 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 6a는은 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 6b은 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 7a는 일 실시 예에 따른 전자 장치의 제1 모드 동작 환경을 도시한다.
도 7b는 일 실시 예에 따른 전자 장치의 제2 모드 동작 환경을 도시한다.
도 7c는 일 실시 예에 따른 전자 장치의 제2 모드 동작 환경을 도시한다.
도 7d는 일 실시 예에 따른 전자 장치의 제3 모드 동작 환경을 도시한다.
도 8a는 일 실시 예에 따른 전자 장치를 도시한다.
도 8b는 일 실시 예에 따른 전자 장치를 도시한다.
도 9는 일 실시 예에 따른 모달리티 결정에 기반한 응답 제공 방법의 흐름도이다.
도 10는 일 실시 예에 따른 모달리티 결정에 기반한 응답 제공 방법의 흐름도이다.
도 11은 본 문서 내에서 설명된 동작들을 수행할 수 있는 예시적인(exemplary) 전자 장치의 블록도이다.
도면의 설명과 관련하여, 동일 또는 유사한 구성요소에 대해서는 동일 또는 유사한 참조 부호가 사용될 수 있다.
이하, 본 발명의 다양한 실시 예가 첨부된 도면을 참조하여 기재된다. 그러나, 이는 본 발명을 특정한 실시 형태에 대해 한정하려는 것이 아니며, 본 발명의 실시 예의 다양한 변경(modification), 균등물(equivalent), 및/또는 대체물(alternative)을 포함하는 것으로 이해되어야 한다.
도 1a는 전자 장치의 예시들을 도시한다.
도 1a를 참조하여, 일 실시 예에 따르면, 전자 장치는 다중 모달리티(modality)를 지원할 수 있다. 예를 들어, 전자 장치는 다양한 유형의 입력을 수신하고 이를 처리하도록 설정될 수 있다. 예를 들어, ‘모달리티’는 전자 장치와 사용자 사이의 상호작용을 위한 채널을 의미할 수 있다. 이 경우, 음성 입력과 인터페이스(예: 가상 키보드)를 통한 텍스트 입력은 서로 다른 모달리티로 간주될 수 있다. 예를 들어, ‘모달리티’는 인공지능 모델(예: 생성형 인공지능 모델)에 입력되는 데이터의 포맷을 의미할 수 있다. 예를 들어, 사용자의 음성 입력은 STT(speech to text) 및 NLU(natural language understanding)을 통하여 텍스트 데이터로 변환되어 인공지능 모델에 입력될 수 있다. 이 경우, 인공지능 모델의 관점에서, 음성 입력과 인터페이스(예: 가상 키보드)를 통한 텍스트 입력은 동일한 모달리티로 간주될 수 있다. 이하에서, ‘모달리티’는 사용자와 전자 장치 사이의 채널 및/또는 인공지능 모델의 입력 데이터의 포맷을 의미할 수 있다. 본 개시에서, ‘모달리티’는 ‘입력 유형’으로 참조될(referred as) 수 있다.
일 예에서, 제1 전자 장치(10a)는 스마트 글래스일 수 있다. 제1 전자 장치(10a)는 마이크를 이용하여 사용자의 음성 입력을 획득할 수 있다. 제1 전자 장치(10a)는 카메라를 이용하여 이미지 입력을 획득할 수 있다. 제1 전자 장치(10a)는 인터페이스(예: 버튼 및/또는 터치 패드)를 이용하여 사용자 입력을 획득할 수 있다. 제1 전자 장치(10a)는 음성 입력, 이미지 입력 및/또는 사용자 입력을 획득하고, 획득된 입력을 처리하도록 설정될 수 있다.
일 예에서, 제2 전자 장치(10b)는 사용자의 신체 또는 옷에 부착되도록 설정된 웨어러블 장치일 수 있다. 예를 들어, 제2 전자 장치(10b)는 사용자의 신체 또는 의류에 부착할 수 있는 부착 구조를 포함할 수 있다. 제2 전자 장치(10b)는, 예를 들어, 마이크, 카메라, 및/또는 인터페이스를 통하여 입력을 획득할 수 있다.
일 예에서, 제3 전자 장치(10c)는 모바일 장치일 수 있다. 제3 전자 장치(10c)는 핸드헬드(handheld) 장치일 수 있다. 제3 전자 장치(10c)는 예를 들어, 제3 전자 장치(10c)는 마이크, 카메라, 및/또는 인터페이스를 통하여 입력을 획득할 수 있다.
제1 전자 장치(10a), 제2 전자 장치(10b), 및 제3 전자 장치(10c)는 도 2와 관련하여 설명되는 전자 장치(10)의 일 예시이다. 본 개시의 전자 장치(10)는 도 1의 제1, 제2, 및 제3 전자 장치(10a, 10b, 10c)에 제한되지 않는다. 예를 들어, 전자 장치(10)는 다중 모달리티를 지원하는 임의의 사용자 장치일 수 있다. 전자 장치(10)는, VST(video see-through) 장치, 시계형 전자 장치, HMD(head mounted device), 및/또는 비히클 인포테인먼트 시스템(vehicle infotainment system)을 포함할 수 있다.
도 1b는 일 실시 예에 따른 전자 장치의 블록도를 도시한다.
도 1b를 참조하여, 일 실시 예에 따르면, 전자 장치(10)는 프로세서(120), 메모리(130), 센서 회로(140), 디스플레이(160), 카메라(170), 인터페이스(180), 및/또는 통신 회로(190)를 포함할 수 있다. 전자 장치(10)는 도 1의 제1 전자 장치(10a), 제2 전자 장치(10b), 제3 전자 장치(10c), 및/또는 도 11의 전자 장치(1100)에 대응할 수 있다. 전자 장치(10)는 도 11과 관련하여 후술되는 전자 장치(1100)와 유사한 구성을 포함할 수 있다. 예를 들어, 프로세서(120)는 도 11의 적어도 하나의 프로세서(1110)에 대응할 수 있다. 예를 들어, 메모리(130)는 도 11의 메모리(1120)에 대응할 수 있다. 예를 들어, 센서 회로(140)는 도 11의 센서 인터페이스(1119) 및/또는 센서(1170)에 대응할 수 있다. 예를 들어, 디스플레이(160)는 도 11의 디스플레이(1140)에 대응할 수 있다. 예를 들어, 카메라(170)는 도 11의 이미지 센서(1150)를 포함할 수 있다. 예를 들어, 통신 회로(190)는 도 11의 통신 회로(1160)에 대응할 수 있다. 도 1b에 도시된 전자 장치(10)의 구성은 예시적인 것으로서, 전자 장치(10)의 구성이 이에 제한되는 것은 아니다. 예를 들어, 전자 장치(10)는 도 1b에 미도시된 구성(예: 도 11의 전자 장치(1100)의 구성 중 적어도 하나)을 더 포함할 수 있다. 예를 들어, 전자 장치(10)는 도 1b에 도시된 구성 중 적어도 하나(예: 제2 마이크(183), 스피커(184), 및/또는 센서 회로(140))를 포함하지 않을 수 있다.
프로세서(120)는 메모리(130), 센서 회로(140), 디스플레이(160), 카메라(170), 인터페이스(180), 및/또는 통신 회로(190)와 통신적으로(communicatively), 전기적으로(electrically), 작동적으로(operatively), 또는 기능적으로(functionally) 연결될 수 있다. 본 개시의 다양한 실시 예들에서, 일 구성이 타 구성과 “작동적으로” 연결된 경우, 일 구성은 타 구성을 작동시킬 수 있도록 연결된 것을 의미할 수 있다. 예를 들어, 일 구성은 직접 또는 다른 구성을 거쳐서 타 구성에 제어 신호를 전달함으로써 타 구성을 작동시킬 수 있다. 본 개시의 다양한 실시 예들에서 일 구성이 타 구성과 “기능적으로” 연결된 경우, 일 구성은 타 구성의 기능을 실행할 수 있도록 연결된 것을 의미할 수 있다. 예를 들어, 일 구성은 직접 또는 다른 구성을 거쳐서 타 구성에 제어 신호를 전달함으로써 타 구성의 기능을 실행시킬 수 있다.
프로세서(120)는 적어도 하나의 프로세서를 포함할 수 있다. 예를 들어, 프로세서(120)는 AP(application processor), CPU(central processing unit), ISP(image signal processor), GPU(graphics processing unit), NPU(neural processing unit), TPU(tensor processing unit), 및/또는 CP(communication processor)를 포함할 수 있다. 프로세서(120)는 적어도 하나의 칩 또는 하나의 칩셋을 포함할 수 있다. 본 개시에서, 프로세서(120)는 적어도 하나의 프로세싱 회로에 의한 구조(architecture)를 갖는 하드웨어 컴포넌트로 참조될 수 있다. 예를 들어, 프로세서(120)는 전자 장치(10) 내부에 위치된 기판(예: 인쇄 회로 기판)에 배치되고, 기판에 형성된 적어도 하나의 도전성 경로를 통하여 전자 장치(10)의 다른 컴포넌트와 통신할 수 있다.
메모리(130)는 인스트럭션들을 저장할 수 있다. 인스트럭션들은 프로세서(120)에 의하여 실행되었을 때, 전자 장치(10)로 하여금 다양한 동작들을 수행하도록 할 수 있다. 예를 들어, 인스트럭션들은 적어도 하나의 프로세서에 의하여 개별적으로(individually) 또는 조합적으로(collectively) 실행됨으로써 전자 장치(10)가 다양한 동작을 수행하도록 할 수 있다. 본 개시의 다양한 실시 예들에서, 전자 장치(10)의 동작은 메모리(130)에 저장된 인스트럭션들을 실행함으로써 프로세서(120)에 의하여 수행되는 동작으로 참조될 수 있다. 메모리(130)는 데이터 저장을 위한 하드웨어 컴포넌트로 참조될 수 있다.
센서 회로(140)는 전자 장치(10)에 연관된 정보(예: 컨텍스트 정보)를 감지하도록 설정될 수 있다. 컨텍스트 정보는 전자 장치(10)의 주변의 광학적 정보, 전자 장치(10)의 움직임 정보, 전자 장치(10)의 부착 상태 정보, 및/또는 인접 오브젝트의 정보를 포함할 수 있다. 예를 들어, 센서 회로(140)는 이미지 센서(141), 모션 센서(143), 부착 센서(145), 및/또는 근접 센서(147)를 포함할 수 있다. 도 1b에 도시된 센서 회로(140)의 구성은 예시적인 것으로서, 본 개시의 실시 예들이 이에 제한되는 것은 아니다.
이미지 센서(141)는 광학적 신호를 감지하도록 설정될 수 있다. 예를 들어, 이미지 센서(141)는 광신호, 적외선 신호, 및/또는 자외선 신호를 감지하도록 설정될 수 있다. 이미지 센서(141)는 뎁스(depth) 정보를 획득하도록 설정될 수 있다. 예를 들어, 이미지 센서(141)는 라이다(light detection and ranging, LiDAR)를 포함할 수 있다.
모션 센서(143)는 전자 장치(10)의 움직임(movement)을 감지하도록 설정될 수 있다. 예를 들어, 모션 센서(143)는 전자 장치(10)의 배향(orientation), 자력, 각율(angular rate), 및/또는 비력(specific force)을 감지하도록 설정된 IMU(inertial measurement unit)을 포함할 수 있다. 모션 센서(143)는 가속도계(accelerometer), 자이로스코프(gyroscope), 및/또는 자기계(magnetometer)를 포함할 수 있다.
부착 센서(145)는 전자 장치(10)의 부착 상태를 감지하도록 설정될 수 있다. 예를 들어, 부착 센서(145)는 전자 장치의 부착물을 감지함으로써 전자 장치(10)의 부착 상태를 감지할 수 있다. 부착 센서(145)는 부착물의 자력을 감지하도록 설정된 홀(hall) 센서를 포함할 수 있다. 부착 센서(145)의 동작은 도 8a 및 8b와 관련하여 후술될 수 있다.
근접 센서(147)는 전자 장치(10)와 근접한 외부 오브젝트를 감지하도록 설정될 수 있다. 전자 장치(10)는 근접 센서(147)를 이용하여 전자 장치(10)의 차폐(blockage) 상태를 감지할 수 있다. 예를 들어, 근접 센서(147)는 정전식(capacitive) 센서, 광전(photoelectric) 센서, 터치 스위치, 조도 센서, 및/또는 유도형(inductive) 센서를 포함할 수 있다.
디스플레이(160)는 이미지를 디스플레이하도록 설정된 적어도 하나의 픽셀을 포함할 수 있다. 일 예에서, 디스플레이(160)는 복수의 디스플레이들을 포함할 수 있다. 예를 들어, 디스플레이(160)는 좌안 디스플레이 및 우안 디스플레이를 포함할 수 있다. 디스플레이(160)는 전면 디스플레이 및/또는 후면 디스플레이를 포함할 수 있다. 디스플레이(160)는 씨쓰루(see-through) 디스플레이, 플렉서블(flexible) 디스플레이, 롤러블(rollable) 디스플레이, 폴더블(foldable) 디스플레이 및/또는 리지드(rigid) 디스플레이 중 적어도 하나를 포함할 수 있다. 일 예에서, 디스플레이(160)는 이미지를 프로젝팅(projecting)하기 위한 적어도 하나의 프로젝터를 포함할 수 있다.
카메라(170)는 적어도 하나의 카메라를 포함할 수 있다. 전자 장치(10)가 복수의 카메라들을 포함하는 경우, 복수의 카메라들 각각은 향하는 방향, 배율, 또는 시야 중 적어도 하나가 상이할 수 있다. 일 예에서, 전자 장치(10)는 이미지를 획득하기 위한 카메라를 복수의 카메라들로부터 선택할 수 있다. 전자 장치(10)는 사용자의 컨텍스트 정보(예: 발화 및/또는 이동 방향)에 기반하여 이미지를 획득하기 위한 카메라를 선택할 수 있다. 예를 들어, 전자 장치(10)는 복수의 카메라들 중, 사용자 발화에 대응하는 오브젝트를 포함하는 이미지를 이미지 인식을 통하여 식별할 수 있다. 전자 장치(10)는 식별된 이미지를 획득한 카메라를 이용하여 음성 입력에 대응하는 이미지를 획득할 수 있다. 예를 들어, 전자 장치(10)는 복수의 카메라들 중, 전자 장치(10)의 이동 방향을 향하는 카메라를 선택할 수 있다. 일 예에서, 전자 장치(10)는 사용자 입력에 따라서 복수의 카메라들 중 하나의 카메라를 선택할 수 있다.
인터페이스(180)는 입력을 수신하도록 설정된 적어도 하나의 장치를 포함할 수 있다. 예를 들어, 인터페이스(180)는 터치 입력을 수신하도록 설정된 터치 회로(181)(예: 터치 스크린 디스플레이)를 포함할 수 있다. 인터페이스(180)는 음성 입력을 수신하도록 설정된 적어도 하나의 마이크(예: 제1 마이크(182) 및/또는 제2 마이크(183))를 포함할 수 있다. 일 예에서, 전자 장치(10)는 제1 마이크(182) 및 제2 마이크(183)를 이용하여 빔포밍을 수행함으로써, 음성 입력의 화자의 위치(예: 전자 장치(10)에 대한 화자의 상대적인 방향)를 식별할 수 있다. 인터페이스(180)는 출력을 위한 적어도 하나의 장치를 포함할 수 있다. 예를 들어, 인터페이스(180)는 촉각 출력을 위한 햅틱 모듈, 소리 출력을 위한 적어도 하나의 스피커(예: 스피커 (184)), 및/또는 인디케이터를 포함할 수 있다. 인터페이스(180)는 임의의 HID(human interface device)를 포함할 수 있다. 예를 들어, 인터페이스(180)는 버튼(185)을 포함할 수 있다. 일 실시 예에 따르면, 프로세서(120)는 인터페이스(180)를 이용하여 입력을 수신하고, 수신된 입력을 처리하도록 설정될 수 있다.
통신 회로(190)는 근거리 무선 통신 및/또는 원거리 무선 통신을 수행하도록 설정될 수 있다. 통신 회로(190)는 NIC(network interface card)를 포함할 수 있다. 프로세서(120)는, 예를 들어, 통신 회로(190)를 이용하여 무선 통신 및/또는 유선 통신에 기반하여 다른 외부 전자 장치와 통신할 수 있다. 프로세서(120)는, 예를 들어, 통신 회로(190)를 이용하여 IP(internet protocol) 네트워크를 통하여 외부 장치와 통신할 수 있다.
일 실시 예에 따르면, 전자 장치(10)는 입력(예: 태스크(task) 수행 요청)을 획득하고, 입력을 처리하여 결과물을 생성할 수 있다. 예를 들어, 전자 장치(10)는 입력을 생성형 인공지능 모델을 이용하여 입력을 처리함으로써 결과물을 생성할 수 있다. 전자 장치(10)는 단일 모달 입력 및/또는 다중 모달 입력을 이용하여 결과물을 생성할 수 있다. 전자 장치(10)는, 예를 들어, 도 2와 관련하여 후술되는 인공지능 시스템(200)을 이용하여 결과물을 생성할 수 있다. 전자 장치(10)는 생성된 결과물을 이용하여 입력에 대한 응답을 제공할 수 있다. 전자 장치(10)는 인터페이스(180) 및/또는 디스플레이(160)를 이용하여 응답을 제공할 수 있다. 전자 장치(10)는 응답에 대응하는 정보를 통신 회로(190)를 이용하여 외부 전자 장치로 송신함으로써, 외부 전자 장치(미도시)로 하여금 응답을 제공하도록 할 수 있다.
본 개시의 실시예와 무관한 예시로서, 사용자 입력(예: 음성 입력)이 수신되면, 다중 모달 입력을 지원하는 전자 장치(10)는 사용자 입력 및 카메라(170)를 이용하여 획득된 이미지를 인공지능 모델에 입력함으로써 결과물을 생성할 수 있다. 사용자 입력이 카메라(170)를 이용하여 획득된 이미지와 무관한 경우, 응답과 사용자 입력 사이의 연관성이 감소될 수 있다. 예를 들어, 사용자의 의도와 무관한 응답이 제공될 수 있다. 또한, 사용자가 별도 입력에 의하여 모달리티를 지정하는 경우, 사용자의 경험이 열화될 수 있다. 또한, 다중 모달 입력의 처리를 위하여, 전자 장치(10)는 카메라(170)를 상시적으로 활성화시킬 수 있다. 이 경우, 전자 장치(10)의 소비 전력이 증가할 수 있다.
본 개시의 일 실시 예에 따르면, 전자 장치(10)는 사용자 입력의 수신에 기반하여, 입력 모달리티를 결정할 수 있다. 예를 들어, 전자 장치(10)는 전자 장치(10)의 컨텍스트 정보(예: 동작 환경에 연관된 정보)를 이용하여 입력 모달리티를 결정할 수 있다. 전자 장치(10)는 센서 회로(140), 카메라(170), 및/또는 인터페이스(180)를 이용하여 컨텍스트 정보를 획득할 수 있다. 단일 모달 입력에 의한 사용자 입력의 처리가 결정된 경우, 전자 장치(10)는 사용자 입력을 인공지능 모델을 이용하여 처리함으로써 응답을 생성할 수 있다. 다중 모달 입력에 의한 사용자 입력의 처리가 결정된 경우, 전자 장치(10)는 사용자 입력 및 컨텍스트 정보를 인공지능 모델을 이용하여 처리함으로써 응답을 생성할 수 있다. 입력 모달리티의 결정을 통하여, 전자 장치(10)는 응답과 사용자 입력 사이의 연관성을 증가시키고, 전자 장치(10)의 소모 전력을 감소시킬 수 있다.
도 2는 일 실시 예에 따른 인공지능 시스템의 블록도를 도시한다.
도 1b 및 도 2를 참조하여, 일 실시 예에 따르면, 인공지능 시스템(200)은 수신된 입력을 인공지능 모델을 이용하여 처리하고, 인공지능 모델을 이용하여 생성된 결과물을 제공하도록 설정될 수 있다. 일 예에서, 인공지능 시스템(200)은 전자 장치(10)에 의하여 구현될 수 있다. 인공지능 시스템(200)의 구성들은 전자 장치(10)가 메모리(130)에 저장된 인스트럭션들을 프로세서(120)를 이용하여 실행함으로써 구현되는 소프트웨어 모듈(예: 쓰레드, 함수, 데이터베이스 및/또는 프로그램)일 수 있다. 일 예에서, 인공지능 시스템(200)의 일부는 외부 장치(예: 서버)에 의하여 구현될 수 있다. 지시 DB(database)(260) 및/또는 인공지능 모델 DB(280)는 외부 장치에 의하여 구현될 수 있다. 이 경우, 전자 장치(10)는 통신 회로(190)를 이용하여 외부 장치와 통신함으로써 정보를 송수신할 수 있다.
도 2의 예시에서, 인공지능 시스템(200)은 I/O(input/output) 인터페이스(210), 인공지능 프레임워크(framework)(220), 지식 DB(260), 어플리케이션(270), 인공지능 모델 DB(280), 및/또는 동작 모들 결정 모듈(290)을 포함할 수 있다. 도 2에 도시된 인공지능 시스템(200)의 구성들은 일 예시로서, 구성들의 적어도 일부는 하나의 소프트웨어 모듈로서 구현될 수 있다.
일 실시 예에 따르면, I/O(input/output) 인터페이스(210)는 사용자 입력 및/또는 컨텍스트 정보를 인공지능 프레임워크(220) 및/또는 동작 모들 결정 모듈(290)에 사용자 입력 및/또는 컨텍스트 정보를 제공할 수 있다.
사용자 입력은, 제1 마이크(182) 및/또는 제2 마이크(183)를 이용하여 획득된 음성 입력(예: 자연어 입력)을 포함할 수 있다. 사용자 입력은 언어적 입력 및/또는 비언어적 입력(예: 메뉴를 선택하는 입력)을 포함할 수 있다. 사용자 입력은 음성 입력에 대하여 STT를 수행함으로써 획득된 텍스트 입력, 버튼(185)을 통하여 입력된 입력, 키보드, 및/또는 가상 키보드의 인터페이스를 통하여 획득된 텍스트 입력을 포함할 수 있다. 사용자 입력은 전자 장치(10)와 통신적으로 연결된 임의의 HID(human interface device) 또는 외부 장치로부터 획득된 입력을 포함할 수 있다. 사용자 입력은 카메라(170) 및/또는 이미지 센서(141)를 이용하여 획득된 시각적 정보를 포함할 수 있다. 사용자 입력은 상술된 정보들 각각 또는 상술된 정보들의 임의의 조합을 포함할 수 있다.
컨텍스트 정보는 전자 장치(10)의 환경에 연관된 정보를 포함할 수 있다. 예를 들어, 컨텍스트 정보는 전자 장치(10)의 배터리 SoC(state of Charge), 제1 마이크(182) 및/또는 제2 마이크(183)를 이용하여 획득된 배경음, 모션 센서(143)를 이용하여 획득된 움직임 정보, 부착 센서(145)를 이용하여 획득된 부착 정보, 및/또는 전자 장치(10) 차폐 상태 정보를 포함할 수 있다. 차폐 상태 정보는 카메라(170), 통신 회로(190), 이미지 센서(141), 및/또는 근접 센서(147) 중 적어도 하나를 이용하여 외부 오브젝트를 감지함으로써 획득될 수 있다. 컨텍스트 정보는 카메라(170) 및/또는 이미지 센서(141)를 이용하여 획득된 전자 장치(10)의 주변 이미지를 포함할 수 있다. 컨텍스트 정보는, 전자 장치(10)에서 현재 실행 중인 어플리케이션 정보 및/또는 전자 장치(10)의 위치 정보를 포함할 수 있다. 컨텍스트 정보는 상술된 정보들 각각 또는 상술된 정보들의 임의의 조합을 포함할 수 있다.
I/O(input/output) 인터페이스(210)는 인공지능 모델에 의하여 생성된 결과물을 출력하도록 설정될 수 있다. 예를 들어, I/O 인터페이스(210)는 AI 프레임워크(220)로부터 결과물을 수신할 수 있다. I/O 인터페이스(210)는 결과물을 자연어 또는 임의의 형태의 컨텐츠로 출력할 수 있다. 예를 들어, I/O 인터페이스(210)는 디스플레이(160)를 이용하여 결과물을 시각적으로 출력할 수 있다. 예를 들어, I/O 인터페이스(210)는 스피커(184)를 이용하여 결과물을 오디오 컨텐츠로 출력할 수 있다. 예를 들어, I/O 인터페이스(210)는 통신 회로(190)를 이용하여 외부 장치에 송신함으로써 외부 장치가 결과물을 출력하도록 할 수 있다.
인공지능 프레임워크(220)는 I/O 인터페이스(210)로부터 수신된 사용자 입력 및/또는 컨텍스트 정보를 수신하고, 사용자 입력에 포함된 사용자의 질의에 기초하여 사용자 질의에 대응하는 의도(예: 태스크의 수행)에 대응하는 동작을 수행하기 위하여 구성 요소들을 제어하도록 설정될 수 있다. 예를 들어, 인공지능 프레임워크(220)는 프롬프트 관리자(230), 어플리케이션 관리자(240), 및/또는 출력 관리자(250)를 포함할 수 있다.
프롬프트 관리자(230)는 사용자 입력 및/또는 컨텍스트 정보로부터 인공지능 모델에 입력하기 위한 프롬프트를 생성할 수 있다. 예를 들어, 프롬프트 관리자(230)는 사용자 입력 및/또는 컨텍스트 정보를 대형 언어 모델(large language model, LLM), 대형 다중 모달 모델(large multi-modal model, LMM), 및/또는 대형 비전 모델(large vision model, LVM)에 입력하기 위한 형태의 프롬프트로 가공할 수 있다. 예를 들어, 입력이 음성 입력인 경우, 음성 입력에 대한 텍스트 변환 및 자연어 이해가 수행될 수 있다. 예를 들어, 입력이 이미지를 포함하는 경우, 이미지에 대한 특징점 추출 또는 오브젝트 식별이 수행될 수 있다. 프롬프트 관리자(230)는 음성 입력으로부터 추출된 텍스트 정보, 특징점 정보, 및/또는 오브젝트 정보를 이용하여 프롬프트를 생성할 수 있다. 일 예에서, 프롬프트 관리자(230)는 프롬프트의 생성을 위하여 학습된 머신 러닝 알고리즘 또는 인공지능 신경망을 이용할 수 있다.
지식 DB(260)는 전자 장치(10)와 관련된 이력 정보를 저장할 수 있다. 예를 들어, 지식 DB(260)는 사용자 입력에 기반하여 생성된 사용자 선호도 데이터, 프롬프트 라이브러리(library), 및/또는 프롬프트 예제 데이터를 저장할 수 있다. 프롬프트 관리자(230)는 지식 DB(260)에 저장된 데이터를 이용하여 사용자 입력 및/또는 컨텍스트 정보로부터 프롬프트를 생성할 수 있다. 프롬프트 관리자(230)는 생성된 프롬프트를 인공지능 모델(예: LLM, LVM, 또는 LMM)에 전달할 수 있다.
어플리케이션 관리자(240)는 외부 구성 요소(예: 지식 DB(260), 어플리케이션(270), 인공지능 모델 DB(280), 및/또는 동작 모드 결정 모듈(290))과 인공지능 프레임워크(220) 사이의 인터페이스를 제공할 수 있다. 예를 들어, 어플리케이션 관리자(240)는 API(application programming interface) 또는 플러그인(plug-in)을 이용하여 외부 구성 요소와 인공지능 프레임워크(220) 사이의 채널을 구축할 수 있다.
예를 들어, 사용자 입력이 인공지능 모델에 전달될 때 또는 프롬프트를 생성할 때, 추가 정보가 요구될 수 있다. 어플리케이션 관리자(240)는 어플리케이션(270)과 통신함으로써 추가 정보를 획득할 수 있다. 어플리케이션 관리자(240)는 어플리케이션(270)을 이용하여 추가 정보를 요청하는 알림을 제공하고, 어플리케이션(270)으로부터 추가 정보를 수신할 수 있다. 어플리케이션 관리자(240)는 외부 DB(예: 지식 DB(260)에 접근함으로써 추가 정보를 획득할 수 있다. 어플리케이션 관리자(240)는 추가 정보를 프롬프트 관리자(230) 및/또는 인공지능 모델에 송신할 수 있다.
일 예에서, 사용자 입력이 지정된 태스크 수행에 대한 의도(intent)를 포함할 수 있다. 이 경우, 인공지능 모델에 의하여 생성된 결과물은 의도에 대응하는 태스크를 수행하기 위한 액션(action) 또는 액션의 시퀀스(sequence)를 포함할 수 있다. 어플리케이션 관리자(240)는 태스크 수행에 연관된 적어도 하나의 어플리케이션 또는 서비스를 이용하여 결과물에 포함된 액션 또는 액션의 시퀀스를 실행할 수 있다.
출력 관리자(250)는 인공지능 모델로부터 출력된 결과물에 대한 세부 튜닝(fine-tuning)을 수행할 수 있다. 예를 들어, 세부 튜닝은 컨텐츠에 대한 정책(예: 사용자 약관, 유해성에 대한 정책, 또는 금지된 컨텐츠에 대한 정책) 및/또는 연관성(예: 사용자 입력과 결과물 사이의 연관성)에 기반한 결과물의 튜닝을 포함할 수 있다.
예를 들어, 출력 관리자(250)는 결과물에 금지된 컨텐츠(예: 인종 또는 정치적으로 편향된 컨텐츠)가 존재하는지 결정할 수 있다. 예를 들어, 출력 관리자(250)는 결과물에 유해한 컨텐츠(예: 성인 인증이 요구되는 컨텐츠 또는 공서양속에 반하는 컨텐츠)가 존재하는지 결정할 수 있다. 예를 들어, 출력 관리자(250)는 서비스 약관에 반하는 컨텐츠가 존재하는지 결정할 수 있다. 금지된 컨텐츠, 유해한 컨텐츠, 및/또는 서비스 약관에 반하는 컨텐츠가 존재하는 경우, 출력 관리자(250)는 해당 컨텐츠를 결과물로부터 제외시키거나, 해당 컨텐츠를 다른 컨텐츠로 대체할 수 있다. 일 예에서, 출력 관리자(250)는 금지된 컨텐츠, 유해한 컨텐츠, 및/또는 서비스 약관에 반하는 컨텐츠가 존재하는 경우, 해당 컨텐츠가 제공될 수 없음을 안내하는 알림을 제공할 수 있다. 일 예에서, 출력 관리자(250)는 금지된 컨텐츠, 유해한 컨텐츠, 및/또는 서비스 약관에 반하는 컨텐츠가 출력되지 않도록 하는 가이드 정보를 사용자에게 제공할 수 있다.
예를 들어, 출력 관리자(250)는 결과물과 사용자 입력(예: 사용자 입력의 의도) 사이의 연관성을 검증(verify)할 수 있다. 예를 들어, 출력 관리자(250)는 결과물에 대한 정보를 추출하기 위하여, 인공지능 모델을 이용할 수 있다. 출력 관리자(250)는 추출된 정보와 사용자 입력을 비교함으로써 두 정보 사이의 연관 정도를 식별할 수 있다. 연관 정도가 임계값 미만인 경우, 출력 관리자(250)는 추가적인 동작을 수행할 수 있다. 예를 들어, 출력 관리자(250)는 사용자 입력을 수정하고, 수정된 사용자 입력을 프롬프트 관리자(230)에 제공할 수 있다. 인공지능 시스템(200)은 수정된 사용자 입력에 기반하여 생성된 프롬프트를 인공지능 모델을 이용하여 처리함으로써, 사용자 입력의 의도에 부합하는 결과물을 생성할 수 있다.
어플리케이션(270)은 전자 장치(10)에 설치된 임의의 어플리케이션을 포함할 수 있다. 어플리케이션(270)은 사용자 인터페이스(예: front-end) 및/또는 인공지능 프레임워크(220)에 대한 클라이언트 엔드(client-end)로 동작할 수 있다. 어플리케이션(270)은 입력을 획득하기 위한 사용자 인터페이스를 제공할 수 있다. 일 예에서, 어플리케이션(270)은 어플리케이션 관리자(240)에 의하여 지시된 동작을 처리하도록 설정될 수 있다.
인공지능 모델 DB(280)는 적어도 하나의 인공지능 모델을 저장할 수 있다. 적어도 하나의 인공지능 모델은, LMM, LVM, 및/또는 LMM을 포함할 수 있다. 예를 들어, 인공지능 모델은, 적어도 하나의 생성형 인공지능 모델을 포함할 수 있다. 인공지능 모델은, 이미지 및/또는 언어를 생성하도록 학습된 인공지능 모델을 포함할 수 있다. 예를 들어, 인공지능 모델은 이미지의 생성을 위한 GAN(generative adversarial network), VAE(variational auto encoder), 및/또는 디퓨전(diffusion) 기반 생성형 모델을 포함할 수 있다. 디퓨전 기반 생성형 모델은 VAE와 트랜스포머(transformer)구조를 이용할 수 있다. 예를 들어, 인공지능 모델은 텍스트 데이터의 생성을 위하여, 입력 값을 기반하여 통계학적인 출력 값을 생성하도록 학습된 모델(예: CHAT-GPT 3, CHAT-GPT 4)을 포함할 수 있다.
동작 모드 결정 모듈(290)은 사용자 입력 및/또는 컨텍스트 정보에 기반하여 전자 장치(10)의 동작 모드를 결정하도록 설정될 수 있다. 예를 들어, 동작 모드 결정 모듈(290)은 카메라(170)의 시야(field of view, FoV), 카메라(170)를 이용하여 획득된 이미지, 센서 회로(140)를 이용하여 획득된 정보, 및/또는 인터페이스(180)를 통하여 획득된 정보에 기반하여 전자 장치(10)의 동작 모드를 결정할 수 있다.
예를 들어, 동작 모드 결정 모듈(290)은 사용자의 시선의 방향과 카메라(170)의 시선의 방향에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 동작 모드 결정 모듈(290)은 전자 장치(10)의 차폐 여부에 기반하여 동작 모드를 결정할 수 있다. 동작 모드 결정 모듈(290)은 전자 장치(10)의 부착물의 부착 상태에 기반하여 동작 모드를 결정할 수 있다. 동작 모드 결정 모듈(290)의 동작 모드 결정 방법은 도 4 내지 도 10과 관련하여 후술되는 예시들에서 보다 구체적으로 설명된다.
동작 모드 결정 모듈(290)은 결정된 동작 모드에 따라서 이용될 입력을 제어할 수 있다. 예를 들어, 제1 모드 및 제2 모드는 다중 모달 입력을 지원하는 모드일 수 있다. 예를 들어, 제3 모드는 단일 모달 입력을 지원하는 모드일 수 있다. 본 개시에서, “동작 모드의 결정”은 입력 모달리티 또는 입력 유형의 결정으로 참조될 수 있다.
일 실시 예에 따르면, 동작 모드 결정 모듈(290)은 결정된 동작 모드에 따른 입력을 획득하도록 I/O 인터페이스(210)를 제어할 수 있다. 예를 들어, 사용자 입력만을 이용하는 동작 모드인 경우, 동작 모드 결정 모듈(290)은 I/O 인터페이스(210)가 사용자 입력만을 획득하도록 I/O 인터페이스(210)를 제어할 수 있다. 이 경우, 컨텍스트 정보의 획득을 위한 모듈(예: 센서 회로(140) 및/또는 카메라(170))이 활성화되지 않을 수 있다.
일 실시 예에 따르면, 동작 모드 결정 모듈(290)은 결정된 동작 모드에 따른 입력만을 이용하도록 프롬프트 관리자(230)를 제어할 수 있다. 예를 들어, 사용자 입력만을 이용하는 동작 모드인 경우, 동작 모드 결정 모듈(290)은 프롬프트 관리자(230)가 획득된 사용자 입력과 컨텍스트 정보 중 사용자 입력만을 이용하도록 프롬프트 관리자(230)를 제어할 수 있다.
도 3은 일 실시 예에 따른 인공지능 모델의 구조를 도시한다.
도 3을 참조하여, 일 실시 예에 따르면, 도 2의 인공지능 모델 DB(280)는 적어도 하나의 인공 신경망 모델(300)을 포함할 수 있다. 인공 신경망 모델(300)은 복수의 은닉층(hidden layer)들(320)을 포함할 수 있다. 은닉층들(320)은 입력층(310)과 출력층(330) 사이에 위치되고, 입력층(310)으로부터 전달된 데이터(x1, x2, x3, ... , xn)(n은 4 이상의 정수)를 출력층(330)으로 전달하면서 학습된 적어도 하나의 계층을 포함할 수 있다. 예를 들어, 은닉층들(320)은 제1 은닉층(320-1), 제2 은닉층(320-2), 제3 및 제M 은닉층(320-M)(M은 3 이상의 정수)을 포함할 수 있다. 도 3에 도시된 은닉층의 개수는 일 예시로서, 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 인공 신경망 모델(300)은 도 3에 도시된 은닉층들보다 더 많은 은닉층들을 포함하거나, 도 3에 도시된 은닉층들보다 적은 은닉층들을 포함할 수 있다. 일 예에서, 인공 신경망 모델(300)은 MLP(multi-layer perception) 모델에 대응할 수 있다. 각각의 은닉층들(320)은 복수의 노드들(N1, N2, ... , Nk)을 포함할 수 있다. 각각의 노드들의 가중치의 값들은 입력 데이터를 이용하여 학습될 수 있다.
예를 들어, 도 1b의 전자 장치(10)는 생성된 프롬프트를 입력 데이터로서 인공 신경망 모델(300)에 입력할 수 있다. 입력 데이터는 인공 신경망 모델(300)에 의하여 연산되고, 인공 신경망 모델(300)은 출력 데이터(Y)를 출력할 수 있다.
인공 신경망 모델(300)은 인공지능 모델의 일 예로서, 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 본 개시에서, 용어 “인공지능 모델”은 생성형 인공지능 모델을 포함할 수 있다. 예를 들어, 인공지능 모델은 대형 언어 모델(large language model, LLM), 대형 다중 모달 모델(large multi-modal model, LMM), 및/또는 대형 비전 모델(large vision model, LVM)을 포함할 수 있다.
도 4는 일 실시 예에 따른 태스크 수행 방법의 흐름도이다.
도 1b 및 도 4를 참조하여, 일 실시예에 따르면, 전자 장치(10)는 사용자 입력에 기반하여 태스크를 수행할 수 있다. 예를 들어, 전자 장치(10)는 지정된 프로젝트에 대한 태스크의 수행을 요청하는 사용자 입력을 획득할 수 있다.
도 4와 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 4와 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 4의 순서와 다르게 실행되거나, 도 4의 다른 동작과 실질적으로 동시에 실행될 수 있다.
동작 405에서, 전자 장치(10)는 트리거 이벤트를 감지할 수 있다. 트리거 이벤트는 전자 장치(10)가 사용자의 요청에 대응하는 태스크를 수행을 개시하도록 하는 임의의 이벤트로 참조될 수 있다. 트리거 이벤트는 사용자의 태스크 수행 요청(예: 태스크를 지시하는 입력) 자체, 태스크 수행 요청에 후속하는 입력(예: 태스크를 지시하는 입력에 후행하는 태스크 실행 입력) 또는 태스크 수행 요청에 선행하는 입력을 포함할 수 있다. 예를 들어, 트리거 이벤트는 음성 입력, 터치 입력, 및/또는 버튼 입력을 포함할 수 있다. 전자 장치(10)는 마이크(예: 제1 마이크(182) 및/또는 제2 마이크(183))를 이용하여 사용자로부터 음성 입력이 수신된 경우에 트리거 이벤트를 감지할 수 있다. 전자 장치(10)는 통신 회로(190)를 이용하여 외부 장치로부터 음성 입력에 대응하는 데이터를 수신한 경우에 트리거 이벤트를 감지할 수 있다. 전자 장치(10)는 인터페이스(180)를 통한 터치 입력 또는 버튼 입력이 수신된 경우에 트리거 이벤트를 감지할 수 있다.
도 4의 예시에서, 사용자 입력은 태스크를 지시하는 입력으로 참조될 수 있다. 태스크를 지시하는 입력은, 지정된 태스크에 대한 의도를 포함하거나 지정된 태스크를 설명하는 입력을 포함할 수 있다. 트리거 이벤트가 태스크를 지시하는 입력인 경우, 전자 장치(10)는 동작 405를 통하여 사용자 입력을 획득할 수 있다. 트리거 이벤트가 태스크를 지시하는 입력에 후행하는 입력인 경우, 전자 장치(10)는 동작 405에 앞서서 사용자 입력을 획득할 수 있다. 트리거 이벤트가 태스크 지시하는 입력에 선행하는 입력인 경우, 전자 장치(10)는 동작 410 이후에 사용자 입력을 획득할 수 있다.
동작 410에서, 전자 장치(10)는 동작 모드를 식별할 수 있다. 일 예어서, 전자 장치(10)는 트리거 이벤트의 감지에 기반하여 동작 모드를 식별할 수 있다. 전자 장치(10)는 트리거 이벤트가 감지되면 동작 모드를 식별할 수 있다. 일 예에서, 전자 장치(10)는 메모리(130)에 저장된 동작 모드의 정보를 이용하여 동작 모드를 식별할 수 있다. 이 경우, 동작 모드는 동작 410의 이전에 기결정된 것일 수 있다. 일 예에서, 전자 장치(10)는 트리거 이벤트가 감지되면 동작 모드를 결정함으로써 동작 모드를 식별할 수 있다. 동작 모드의 결정 방법은 도 5와 관련하여 후술될 수 있다.
일 예에 따르면, 동작 모드는 제1 모드, 제2 모드 및 제3 모드를 포함할 수 있다. 제1 모드와 제2 모드는 다중 모달을 지원하는 동작 모드로 참조될 수 있다. 제3 모드는 단일 모달을 지원하는 동작 모드(예: 다중 모달을 지원하지 않는 동작 모드)로 참조될 수 있다. 본 개시에서, “동작 모드”는 입력 모달리티 또는 입력 유형으로 참조될 수 있다. 예를 들어, 전자 장치(10)의 동작 모드 식별은 전자 장치(10)의 입력 모달리티 또는 입력 유형의 식별로 참조될 수 있다.
동작 모드가 기결정된 경우, 전자 장치(10)는 사용자 입력을 기결정된 동작 모드에 따라서 감지할 수 있다. 기결정된 동작 모드가 제1 모드 또는 제2 모드인 경우, 전자 장치(10)는 다중 모달 사용자 입력의 감지를 위하여 다중 구성요소들을 활성화할 수 있다. 예를 들어, 제1 모드 또는 제2 모드에서, 전자 장치(10)는 마이크(예: 제1 마이크(182) 및/또는 제2 마이크(183)) 및 카메라(170)를 활성화 상태로 유지할 수 있다. 이 경우, 전자 장치(10)는 마이크 및 카메라(170)를 이용하여 다중 모달 사용자 입력을 획득할 수 있다. 기결정된 동작 모드가 제3 모드인 경우, 전자 장치(10)는 단일 모달 사용자 입력의 감지를 위하여 마이크를 활성화 상태로 유지할 수 있다. 이 경우, 전자 장치(10)는 카메라(170)를 비활성화 상태 또는 유휴(idle) 상태(예: 이미지 프로세싱을 수행하지 않는 상태)로 유지할 수 있다.
동작 415에서, 전자 장치(10)는 식별된 동작 모드가 다중 모달을 지원하는지 결정할 수 있다. 식별된 동작 모드가 제1 모드 또는 제2 모드인 경우, 전자 장치(10)는 동작 모드가 다중 모달을 지원하는 것으로 결정할 수 있다. 식별된 동작 모드가 제3 모드인 경우, 전자 장치(10)는 동작 모드가 다중 모달을 지원하지 않는 것으로 결정할 수 있다.
동작 모드가 다중 모달을 지원하는 경우(예: 동작 415-YES), 동작 420에서, 전자 장치(10)는 사용자 입력 및 컨텍스트 정보 기반 프롬프트를 생성할 수 있다. 예를 들어, 전자 장치(10)는 도 2와 관련하여 상술된 프롬프트 관리자(230)를 이용하여 프롬프트를 생성할 수 있다. 예를 들어, 전자 장치(10)는 마이크를 이용하여 획득된 음성 입력으로부터 사용자 입력을 획득할 수 있다. 전자 장치(10)는 카메라(170)를 이용하여 획득된 적어도 하나의 이미지를 컨텍스트 정보로 획득할 수 있다. 이 경우, 전자 장치(10)는 음성 입력 및 적어도 하나의 이미지를 이용하여 프롬프트(예: 적어도 하나의 제1 프롬프트)를 생성할 수 있다.
동작 모드가 다중 모달을 지원하지 않는 경우(예: 동작 415-NO), 동작 425에서, 전자 장치(10)는 사용자 입력 기반 프롬프트를 생성할 수 있다. 동작 425에서, 전자 장치(10)는 컨텍스트 정보를 이용하지 않고, 프롬프트를 생성할 수 있다. 전자 장치(10)는 마이크를 이용하여 획득된 음성 입력으로부터 사용자 입력을 획득하고, 획득된 사용자 입력으로부터 프롬프트(예: 제2 프롬프트)를 생성할 수 있다.
동작 430에서, 전자 장치(10)는 프롬프트 기반 결과물을 생성하고 생성된 결과물을 출력할 수 있다. 전자 장치(10)는 동작 420의 프롬프트(예: 적어도 하나의 제1 프롬프트) 또는 동작 425의 프롬프트(예: 제2 프롬프트)를 이용하여 결과물을 생성할 수 있다. 전자 장치(10)는 프롬프트를 적어도 하나의 인공지능 모델(예: 도 2의 인공지능 모델 DB(280))에 입력함으로써 결과물을 생성할 수 있다. 도 2와 관련하여 상술된 바와 같이, 전자 장치(10)는 인공지능 모델의 출력에 대한 조정(예: 파인 튜닝)을 수행함으로써 결과물을 생성할 수 있다. 전자 장치(10)는 생성된 결과물을 디스플레이(160) 및/또는 스피커(184)를 이용하여 출력할 수 있다. 전자 장치(10)는 생성된 결과물에 대한 정보를 통신 회로(190)를 이용하여 외부 장치에 송신하고, 외부 장치로 하여금 결과물을 출력하도록 할 수 있다.
도 5는 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 1b 및 도 5를 참조하여, 일 실시예에 따르면, 전자 장치(10)는 전자 장치(10)의 동작 모드를 결정할 수 있다.
도 5와 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 5와 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 5의 순서와 다르게 실행되거나, 도 5의 다른 동작과 실질적으로 동시에 실행될 수 있다.
동작 505에서, 전자 장치(10)는 트리거 이벤트를 감지할 수 있다. 동작 505의 트리거 이벤트는 전자 장치(10)가 동작 모드를 결정하도록 하는 이벤트로 참조될 수 있다. 예를 들어, 트리거 이벤트는 전자 장치(10)의 턴온(turn-on), 활성화, 붓업(boot-up), 리셋(reset), 부착 상태 변경, 지정된 주기, 움직임, 또는 사용자 입력 수신을 포함할 수 있다. 전자 장치(10)는 턴온, 활성화, 리셋, 또는 붓업 시에 후술되는 동작 510을 실행할 수 있다. 전자 장치(10)는 모션 센서(143)를 이용하여 감지된 움직임 정보가 임계값을 초과하면 동작 510을 실행할 수 있다. 전자 장치(10)는 부착 센서(145)를 이용하여 부착 상태 변경이 감지되면 후술되는 동작 510을 실행할 수 있다. 전자 장치(10)는 지정된 주기 마다 동작 510을 실행할 수 있다. 전자 장치(10)는 사용자 입력이 수신되면 동작 510을 실행할 수 있다.
동작 510에서, 전자 장치(10)는 카메라(170) 및/또는 센서 회로(140)를 이용하여 동작 모드를 결정할 수 있다. 전자 장치(10)는 결정된 동작 모드의 정보를 메모리(130)에 저장할 수 있다. 동작 510이 태스크를 지시하는 사용자 입력의 수신 전에 수행되는 경우, 전자 장치(10)는 결정된 동작 모드에 따라서 사용자 입력(예: 다중 모달 입력 또는 단일 모달 입력)을 수신할 수 있다.
전자 장치(10)는, 예를 들어, 카메라(170)의 시야와 사용자의 시야가 대응하는 경우에 동작 모드를 멀티 모달을 지원하는 모드(예: 제1 모드 또는 제2 모드)로 결정할 수 있다. 전자 장치(10)는 카메라(170)의 시야와 사용자의 시야가 대응하지 않는 경우에 동작 모드를 멀티 모달을 지원하지 않는 모드(예: 제3 모드)로 결정할 수 있다. 카메라(170)의 시야와 사용자의 시야가 대응하는 경우는, 카메라(170)의 시야의 방향과 사용자의 시야의 방향이 실질적으로 동일한 경우 또는 카메라(170)의 시야와 사용자의 시야가 서로를 향하는 경우를 포함할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)는 복수의 카메라들 중 사용자의 시야와 대응하는 카메라(170)가 존재하는지를 사용자의 컨텍스트 정보에 기반하여 결정할 수 있다. 예를 들어, 사용자의 시야에 대응하는 카메라가 존재하는 경우, 전자 장치(10)는 카메라(170)시야와 사용자의 시야가 대응하는 것으로 결정할 수 있다.
제1 모드는 사용자의 시야와 카메라(170)의 시야가 같은 방향을 향하는 경우의 전자 장치(10)의 동작 모드에 대응할 수 있다. 제1 모드에서, 사용자의 시야의 적어도 일부와 카메라(170)의 시야의 적어도 일부가 서로 중첩될 수 있다. 제1 모드에서, 카메라(170)는 사용자가 보는 방향의 시야의 적어도 일부 또는 사용자 시야 이상의 이미지를 촬영할 수 있다. 이 경우, 전자 장치(10)는 사용자의 시야에 대응하는 시각적 컨텍스트 정보를 카메라(170)를 이용하여 획득할 수 있다. 사용자 입력과 시각적 컨텍스트를 이용함으로써, 전자 장치(10)는 보다 사용자 입력에 의하여 지시된 의도에 부합하는 응답을 제공할 수 있다.
일 예에서, 제1 모드는 사용자의 발화의 내용이 카메라(170)의 시야에 대응하는 경우의 전자 장치(10)의 동작 모드에 대응할 수 있다. 예를 들어, 전자 장치(10)는 사용자의 음성 입력에 대한 음성 인식을 수행하고, 음성 입력으로부터 적어도 하나의 엔티티(예: 오브젝트를 지칭하는 키워드)를 식별할 수 있다. 전자 장치(10)는 음성 입력의 수신에 기반하여 카메라(170)를 이용하여 이미지를 획득하고, 이미지에 대한 오브젝트 인식을 수행할 수 있다. 이미지로부터 인식된 오브젝트와 음성 입력으로부터 식별된 엔티티가 서로 매칭되는 경우, 전자 장치(10)는 전자 장치(10)의 동작 모드를 제1 모드로 결정하 수 있다.
제2 모드는 카메라(170)의 시야 내에 사용자가 위치된 경우의 전자 장치(10)의 동작 모드에 대응할 수 있다. 제2 모드에서, 카메라(170)의 시야와 사용자의 시야는 서로를 향하고, 카메라(170)의 시야와 사용자의 시야의 일부가 중첩될 수 있다. 예를 들어, 전자 장치(10)가 일정한 위치에 거치될 수 있다. 전자 장치(10)는 지정된 위치에 놓여 있거나, 사용자에 의하여 파지된 상태일 수 있다. 전자 장치(10)가 지정된 위치에 거치된 상태에서, 사용자는 전자 장치(10)의 카메라(170)를 바라보면서 사용자 입력에 대응하는 음성을 발화할 수 있다. 이 경우, 전자 장치(10)는 카메라(170)를 이용하여 획득된 이미지로부터 사용자에 해당하는 이미지를 식별하고, 식별된 이미지를 이용하여 시각적 컨텍스트 정보를 획득할 수 있다. 예를 들어, 제2 모드에서, 사용자는 “오늘 나 어때 보여?”라는 음성을 발화할 수 있다. 전자 장치(10)는 카메라(170)를 이용하여 획득된 사용자의 이미지와 음성 입력에 대응하는 사용자 입력을 이용하여 사용자 입력에 대응하는 응답을 제공할 수 있다. 예를 들어, 전자 장치(10)는 사용자의 이미지로부터 식별되는 사용자의 외관(outfit), 외모(appearance), 및/또는 옷에 기반하여 사용자 입력에 대한 응답을 제공할 수 있다.
제3 모드는 사용자 입력의 처리에 있어서, 컨텍스트 정보 가 요구되지 않는 동작 모드에 대응할 수 있다. 제3 모드에서, 사용자의 시야와 카메라(170)의 시야는 서로 대응하지 않을 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)는 복수의 카메라들 중 사용자의 시야와 대응하는 카메라(170)가 존재하는지를 사용자의 컨텍스트 정보에 기반하여 결정할 수 있다. 예를 들어, 사용자의 시야에 대응하는 카메라가 존재하지 않는 경우, 전자 장치(10)는 카메라(170)시야와 사용자의 시야가 대응하지 않는 것으로 결정할 수 있다.
예를 들어, 제3 모드의 전자 장치(10)는 카메라(170)를 이용하여 이미지를 획득하지 않고, 사용자 입력을 처리할 수 있다. 제3 모드에서, 전자 장치(10)는 카메라(170)를 이용하여 입력을 수신하지 않거나, 수신된 입력을 처리(예: 인공지능 모델에 전달함으로써)하지 않을 수 있다. 제3 모드에서, 전자 장치(10)는 차폐된 상태(예: 사용자의 주머니 또는 가방에 넣어진 상태)이거나, 사용자를 확인할 수 없는 위치에 부착된 상태일 수 있다. 제3 모드에서, 사용자 입력의 처리에 카메라 입력이 추가되어야 하는 경우가 있을 수 있다. 예를 들어, 사용자 입력이 카메라(170)를 이용하여 획득되는 이미지에 대한 태스크를 포함할 수 있다. 이 경우, 전자 장치(10)는 동작 모드의 변경을 위한 전자 장치(10)의 위치 변경 및/또는 장착 상태 변경을 제안하는 가이드 정보를 출력할 수 있다.
제1 모드 및 제2 모드에서, 전자 장치(10)는 다중 모달 입력의 수신을 위하여, 마이크(182, 183) 및 카메라(170)를 활성화할 수 있다. 예를 들어, 전자 장치(10)는 마이크(182, 183) 및 카메라(170)를 상시적으로 활성화할 수 있다. 이 경우, 전자 장치(10)는 사용자 입력의 이전에 획득된 컨텍스트 정보를 이용하여 응답을 제공할 수 있다. 예를 들어, 사용자 입력은 “아까 지나간 사람 데이비드 아니야?”일 수 있다. 전자 장치(10)는 사용자 입력의 이전에 카메라(170)를 이용하여 획득된 이미지들을 분석함으로써, 사용자의 입력에 대한 응답을 제공할 수 있다.
제1 모드 및 제2 모드에서, 전자 장치(10)는 전력 소모를 감소시키기 위하여, 마이크(182, 183)만을 상시적으로 활성화할 수 있다. 예를 들어, 전자 장치(10)는 사용자 입력이 검출되면, 사용자 입력의 검출에 기반하여 카메라(170)를 활성화시킬 수 있다. 일 예시에서, 전자 장치(10)는 발화가 감지되면 카메라(170)를 활성화시킬 수 있다. 음성 입력에 대한 인식이 완료되기 전에 카메라(170)를 활성화시킴으로써, 사용자의 발화와 컨텍스트 정보(예: 카메라(170)를 이용하여 획득된 이미지) 사이의 지연을 감소시킬 수 있다.
일 예에서, 전자 장치(10)는 지정된 입력(예: 버튼(185)을 통한 입력, 제스처 입력, 또는 터치 입력)이 수신되면 마이크(182, 183) 및 카메라(170)를 활성화시킬 수 있다. 전자 장치(10)는 지정된 입력의 수신 후에 동작 모드를 결정할 수 있다.
일 실시 예에서, 전자 장치(10)는 카메라(170)를 이용하여 획득된 복수의 이미지들과 모션 센서(143)의 정보를 이용하여 동작 모드를 결정할 수 있다. 예를 들어, 카메라(170)를 이용하여 획득된 이미지들에 기반한 움직임 정보가 모션 센서(143)를 이용하여 획득된 전자 장치(10)의 움직임 정보가 서로 대응하는 경우, 전자 장치(10)는 카메라(170)의 시야와 사용자의 시야가 서로 대응하는 것으로 결정할 수 있다. 이 경우, 전자 장치(10)는 전자 장치(10)의 동작 모드를 제1 모드로 결정할 수 있다. 이미지들에 기반한 움직임 정보가 모션 센서(143)를 이용하여 획득된 전자 장치(10)의 움직임 정보가 서로 대응하지 않는 경우, 전자 장치(10)는 카메라(170)를 이용하여 획득된 이미지 내의 사용자 이미지가 존재하면 동작 모드를 제2 모드로 결정할 수 있다. 이미지들에 기반한 움직임 정보가 모션 센서(143)를 이용하여 획득된 전자 장치(10)의 움직임 정보가 서로 대응하지 않는 경우, 카메라(170)를 이용하여 획득된 이미지 내의 사용자 이미지가 존재하지 않으면 전자 장치(10)는 동작 모드를 제3 모드로 결정할 수 있다. 이미지 및 모션 센서(143)를 이용한 동작 모드의 결정은 도 6a와 관련하여 후술될 수 있다.
일 실시 예에서, 전자 장치(10)는 부착 센서(145)를 이용하여 획득된 부착 정보에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 전자 장치(10)는 전자 장치(10)가 사용자에게 착용된 상태인지를 결정할 수 있다. 전자 장치(10)는 전자 장치(10)가 사용자에 의하여 착용된 상태인 경우에 동작 모드를 제1 모드로 결정할 수 있다. 전자 장치(10)가 사용자에 의하여 착용된 상태가 아닌 경우, 전자 장치(10)는 카메라(170)를 이용하여 획득된 이미지 내의 사용자 이미지가 존재하면 동작 모드를 제2 모드로 결정할 수 있다. 전자 장치(10)가 사용자에 의하여 착용된 상태가 아닌 경우, 카메라(170)를 이용하여 획득된 이미지 내의 사용자 이미지가 존재하지 않으면 전자 장치(10)는 동작 모드를 제3 모드로 결정할 수 있다. 부착 상태에 기반한 동작 모드의 결정은 도 8a 및 8b와 관련하여 후술될 수 있다.
일 실시 예에서, 전자 장치(10)는 마이크(182, 183)를 이용하여 획득된 정보에 기반하여 동작 모드를 결정할 수 있다. 전자 장치(10)는 제1 마이크(182) 및 제2 마이크(183)를 이용하여 빔포밍을 수행할 수 있다. 빔포밍 기술을 이용하여, 전자 장치(10)는 발화자의 전자 장치(10)에 대한 상대적인 위치(예: 상대적인 방향)를 식별할 수 있다. 발화자의 위치가 전자 장치(10)의 뒤쪽(예: 전자 장치(10)의 카메라(170)가 향하는 방향의 반대 방향을 향하는 쪽)에 위치된 경우, 전자 장치(10)는 동작 모드를 제1 모드로 결정할 수 있다. 마이크를 이용한 동작 모드의 결정은 도 8a와 관련하여 후술될 수 있다.
일 실시 예에서, 전자 장치(10)는 차폐 여부에 기반하여 동작 모드를 결정할 수 있다. 전자 장치(10)는 통신 회로(190), 이미지 센서(141), 및/또는 근접 센서(147)를 이용하여 전자 장치(10)가 차폐된 상태(예: 전자 장치(10)가 사용자의 주머니 또는 가방 안에 위치된 상태)인지 결정할 수 있다. 전자 장치(10)가 차폐된 상태인 경우, 전자 장치(10)는 동작 모드를 제3 모드로 결정할 수 있다.
일 예에서, 동작 모드의 결정을 위한 정보가 부족할 수 있다. 일 예에서, 전자 장치(10)는 부족한 정보의 획득을 위하여 사용자에게 피드백을 제공할 수 있다. 예를 들어, 전자 장치(10)는 사용자에게 “다시 보여주세요” 또는 “전자 장치의 카메라를 시선과 같은 방향에 위치시켜 주세요”와 같은 가이드를 제공함으로써, 전자 장치(10)의 동작 모드 결정을 위한 정보의 획득을 도모할 수 있다. 일 예에서, 전자 장치(10)는 동작 모드의 결정을 위한 정보가 부족한 경우, 지정된 모드(예: 다중 모달 입력 지원 모드)로 전자 장치(10)의 동작 모드를 결정할 수 있다.
통상의 기술자는 상술된 방법 외에도 동작 모드를 설정하기 위한 다양한 방법이 이용될 수 있음을 이해할 수 있을 것이다. 예를 들어, 전자 장치(10)는 지정된 입력(예: 버튼 입력, 지정된 제스쳐, 또는 지정된 횟수의 버튼 입력)이 수신되면, 지정된 입력에 매핑된 동작 모드를 설정할 수 있다.
도 6a는 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 1b 및 도 6a를 참조하여, 일 실시예에 따르면, 전자 장치(10)는 움직임 정보에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 전자 장치(10)는 카메라(170)를 이용하여 획득된 정보와 센서 회로(140)를 이용하여 획득된 정보에 기반하여 동작 모드를 결정할 수 있다.
도 6a와 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 6a와 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 6a의 순서와 다르게 실행되거나, 도 6a의 다른 동작과 실질적으로 동시에 실행될 수 있다. 예를 들어, 도 6a와 관련하여 후술되는 동작들은 도 5의 동작 510에 대응할 수 잇다.
동작 605에서, 전자 장치(10)는 카메라(170)를 이용하여 복수의 이미지들을 획득할 수 있다. 예를 들어, 전자 장치(10)는 동작 모드를 결정하기 위한 트리거 이벤트가 감지(예: 도 5의 동작 505)되었을 때 복수의 이미지들을 획득할 수 있다. 예를 들어, 전자 장치(10)는 상시적으로 카메라(170)를 이용하여 일정 시간 구간 동안의 이미지들을 저장하도록 설정되고, 트리거 이벤트가 감지되면 트리거 이벤트가 감지된 시점으로부터 일정 시간 범위 내의 이미지들을 복수의 이미지들로 이용할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)가 복수의 카메라들을 포함할 수 있다. 이 경우, 전자 장치(10)는 복수의 카메라들 사용자의 컨텍스트 정보에 기반하여 결정된 카메라(170)를 이용하여 복수의 이미지들을 획득할 수 있다.
동작 610에서, 전자 장치(10)는 모션 센서(143)를 이용하여 전자 장치(10)의 움직임 정보를 획득할 수 있다. 전자 장치(10)는 복수의 이미지들과 시간 동기화된 전자 장치(10)의 움직임 정보를 획득할 수 있다. 예를 들어, 전자 장치(10)는 카메라(170)와 모션 센서(143)를 실질적으로 동시에 활성화시킴으로써, 복수의 이미지들과 시간 동기화된 전자 장치(10)의 움직임 정보를 획득할 수 있다.
동작 615에서, 전자 장치(10)는 복수의 이미지들로부터 식별된 이미지 기반 움직임이 움직임 정보에 대응하는지 결정할 수 있다. 전자 장치(10)는 연속하는 이미지들을 이용하여 이미지들 사이의 픽셀의 이동을 나타내는 벡터를 식별할 수 있다. 전자 장치(10)는 이미지들의 벡터값을 이용하여 이미지 기반 움직임 정보를 식별할 수 있다. 이미지 기반 움직임 정보가 전자 장치(10)의 움직임 정보(예: 모션 센서(143)를 이용하여 획득된 움직임 정보)에 대응(예: 실질적으로 동일)하는 경우, 전자 장치(10)는 카메라(170)의 시야와 사용자의 시야가 대응하는 것으로 결정할 수 있다.
이미지 기반 움직임이 움직임 정보에 대응하는 경우(예: 동작 615-YES), 동작 620에서, 전자 장치(10)는 동작 모드를 제1 모드로 결정할 수 있다. 일 예에서, 사용자가 전자 장치(10)를 손에 든 상태(예: 전자 장치(10)를 캐리(carry)하는 상태)에서 이동하는 경우, 이미지 기반 움직임과 전자 장치(10)의 움직임 정보가 대응할 수 있다. 이 경우, 제1 모드와 제2 모드의 구분을 위하여, 전자 장치(10)는 사용자 이미지가 복수의 이미지들로부터 식별되는지 결정(예: 동작 625)할 수 있다. 전자 장치(10)는 사용자 이미지가 식별되면, 동작 모드를 제2 모드로 결정할 수 있다. 전자 장치(10)는 사용자 이미지가 식별되지 않으면 동작 모드를 제1 모드로 결정할 수 있다.
이미지 기반 움직임이 움직임 정보에 대응하지 않는 경우(예: 동작 615-NO), 동작 625에서, 전자 장치(10)는 복수의 이미지들로부터 사용자 이미지가 식별되는지 결정할 수 있다. 예를 들어, 전자 장치(10)는 복수의 이미지들에 대하여 이미지 인식을 수행할 수 있다. 이미지 인식의 결과, 복수의 이미지들 중 적어도 일부로부터 사용자에 대응하는 이미지(예: 저장된 사용자의 얼굴 이미지)가 식별되면, 전자 장치(10)는 복수의 이미지들로부터 사용자 이미지가 식별된다고 결정할 수 있다.
사용자 이미지가 식별된 경우(예: 동작 625-YES), 동작 630에서, 전자 장치(10)는 동작 모드를 제2 모드로 결정할 수 있다. 사용자 이미지가 식별되지 않은 경우(예: 동작 625-NO), 동작 635에서, 전자 장치(10)는 동작 모드를 제3 모드로 결정할 수 있다.
일 예에서, 전자 장치(10)는 동작 모드가 제1 모드 및 제2 모드가 아닌 것으로 결정되면, 동작 모드를 제3 모드로 결정할 수 있다. 일 예에서, 전자 장치(10)는 전자 장치(10)가 차폐된 상태인 것으로 결정되면, 동작 모드를 제3 모드로 결정할 수 있다. 이 경우, 전자 장치(10)는 도 6a에 도시된 동작들을 수행하지 않을 수 있다.
도 6a에서는, 동작 625가 동작 615에 후속하여 수행되는 것으로 도시되어 있으나, 본 개시의 실시예들이 이에 제한되는 것은 아니다. 예를 들어, 동작 625가 동작 615에 앞서 수행될 수 있다. 사용자 이미지가 식별되면, 전자 장치(10)는 동작 모드를 제2 모드로 결정할 수 있다. 사용자 이미지가 식별되지 않으면, 전자 장치(10)는 동작 615에 따라서 제1 모드 또는 제3 모드로 동작 모드를 결정할 수 있다. 이미지 기반 움직임이 움직임 정보에 대응하면, 전자 장치(10)는 동작 모드를 제1 모드로 결정할 수 있다. 이미지 기반 움직임이 움직임 정보에 대응하지 않으면, 전자 장치(10)는 동작 모드를 제3 모드로 결정할 수 있다.
도 6b는 일 실시 예에 따른 동작 모드 결정 방법의 흐름도이다.
도 1b 및 도 6b를 참조하여, 일 실시예에 따르면, 전자 장치(10)는 음성과 이미지 입력의 대응 여부에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 전자 장치(10)는 카메라(170)를 이용하여 획득된 정보와 음성 입력에 기반하여 동작 모드를 결정할 수 있다.
도 6b와 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 6b와 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 6b의 순서와 다르게 실행되거나, 도 6b의 다른 동작과 실질적으로 동시에 실행될 수 있다. 예를 들어, 도 6b와 관련하여 후술되는 동작들은 도 5의 동작 510에 대응할 수 잇다. 예를 들어, 전자 장치(10)는 동작 505를 통하여 음성 입력을 획득한 것으로 가정할 수 있다.
동작 655에서, 전자 장치(10)는 카메라(170)를 이용하여 적어도 하나의 이미지를 획득할 수 있다. 예를 들어, 전자 장치(10)는 동작 모드를 결정하기 위한 트리거 이벤트가 감지(예: 도 5의 동작 505)되었을 때 적어도 하나의 이미지를 획득할 수 있다. 예를 들어, 전자 장치(10)는 상시적으로 카메라(170)를 이용하여 일정 시간 구간 동안의 이미지들을 저장하도록 설정되고, 트리거 이벤트가 감지되면 트리거 이벤트가 감지된 시점으로부터 일정 시간 범위 내의 이미지들을 이용할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)가 복수의 카메라들을 포함할 수 있다. 이 경우, 전자 장치(10)는 복수의 카메라들 사용자의 컨텍스트 정보에 기반하여 결정된 카메라(170)를 이용하여 적어도 하나의 이미지를 획득할 수 있다.
동작 660에서, 전자 장치(10)는 적어도 하나의 이미지에 대한 이미지 인식을 수행할 수 있다. 예를 들어, 전자 장치(10)는 이미지 인식을 위한 임의의 알고리즘을 이용하여 이미지로부터 적어도 하나의 오브젝트를 식별할 수 있다.
동작 665에서, 전자 장치(10)는 음성 대응 오브젝트가 식별되는지 결정할 수 있다. 예를 들어, 전자 장치(10)는 트리거 이벤트(예: 도 5의 동작 505)를 통하여 음성 입력을 획득할 수 있다. 전자 장치(10)는 음성 입력에 대한 음성 인식을 통하여 음성 입력에 포함된 적어도 하나의 키워드를 식별할 수 있다. 전자 장치(10)는 식별된 키워드에 대응하는 오브젝트가 적어도 하나의 이미지로부터 식별되는 경우 음성 대응 오브젝트가 식별되는 것으로 결정할 수 있다.
음성 대응 오브젝트가 식별되는 경우(예: 동작 665-YES), 동작 670에서, 동작 모드를 제1 모드로 결정할 수 있다. 동작 670에서, 전자 장치(10)는 동작 모드를 제2 모드로 결정할 수도 있다. 예를 들어, 전자 장치(10)는 음성 대응 오브젝트가 식별되고, 이미지로부터 사용자의 이미지도 함께 식별되는 경우, 동작 모드를 제2 모드로 결정할 수 있다.
음성 대응 오브젝트가 식별되지 않는 경우(예: 동작 665-NO), 동작 675에서, 전자 장치(10)는 시선 대응 여부에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 전자 장치(10)는 도 6a와 관련하여 상술된 방법에 따라서 동작 모드를 결정할 수 있다. 동작 675에 제한되지 아니하고, 전자 장치(10)는 본 개시에서 설명된 임의의 동작 모드 결정 방법에 따라서 동작 모드를 결정할 수 있다. 이하에서, 도 7a 내지 도 7d를 참조하여, 전자 장치(10)의 동작 모드에 따른 응답 제공 방법들이 설명될 수 있다. 도 7a 내지 도 7d의 예시에서, 도 1a의 제2 전자 장치(10b)를 이용하여 예시들이 설명되나, 통상의 기술자는 본 개시의 실시예들이 다른 형태의 전자 장치들(예: 전자 장치(10), 제1 전자 장치(10a), 또는 제3 전자 장치(10c))에도 적용될 수 있음을 이해할 수 있을 것이다.
도 7a는 일 실시 예에 따른 전자 장치의 제1 모드 동작 환경을 도시한다.
도 7a를 참조하여, 제2 전자 장치(10b)는 사용자(701)의 옷에 부착된 상태일 수 있다. 이 경우, 사용자(701)의 시야(705)와 제2 전자 장치(10b)의 시야(710)(예: 제2 전자 장치(10b)의 카메라의 시야)는 같은 방향을 향하고, 사용자(701)의 시야(705)의 적어도 일부와 제2 전자 장치(10b)의 시야(710)의 적어도 일부가 중첩될 수 있다. 제2 전자 장치(10b)는 시야(705) 내의 적어도 일부 영역에 대한 이미지를 획득할 수 있다. 이 경우, 제2 전자 장치(10b)는 사용자(701)로부터 음성 입력이 수신되면, 사용자(701)가 바라보고 있을 수 있는 주변 환경 정보(예: 컨텍스트 정보)를 이용하여 프롬프트를 생성할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)가 복수의 카메라들을 포함할 수 있다. 이 경우, 전자 장치(10)는 복수의 카메라들 중 사용자의 컨텍스트 정보에 기반하여 결정된 카메라를 이용할 수 있다.
예를 들어, 사용자(701)가 “정말 멋지다”라는 음성 입력을 발화할 수 있다. 이 경우, 제2 전자 장치(10b)는 카메라를 이용하여 획득된 이미지를 분석할 수 있다. 제2 전자 장치(10b)는 획득된 이미지에 대한 영상 분석을 수행함으로써, 저녁 시간 대의 바다의 이미지를 식별할 수 있다. 이 경우, 제2 전자 장치(10b)는 컨텍스트 기반 프롬프트로써, “노을 지는 잔잔한 바다를 바라보고 있는 나”라는 프롬프트를 생성할 수 있다. 또한, 제2 전자 장치(10b)는 사용자(701)의 음성 입력에 기반하여 “정말 멋지다고 한 나의 감탄에 대한 응답”이라는 프롬프트를 생성할 수 있다. 제2 전자 장치(10b)는 컨텍스트 기반 프롬프트 및 음성 입력 기반 프롬프트를 인공지능 모델에 입력함으로써 결과물을 생성할 수 있다. 일 예에서, 제2 전자 장치(10b)는 카메라를 이용하여 획득된 이미지의 정보를 컨텍스트 기반 프롬프트로 이용할 수 있다. 이미지 입력 및 텍스트 입력을 지원하는 인공지능 모델의 경우, 제2 전자 장치(10b)는 이미지를 음성 입력 기반 프롬프트와 함께 인공지능 모델을 이용하여 처리함으로써, 결과물을 생성할 수 있다.
도 7b는 일 실시 예에 따른 전자 장치의 제2 모드 동작 환경을 도시한다.
도 7b를 참조하여, 사용자(702)는 제2 방향(795)으로 이동 중일 수 있다. 동시에, 사용자(702)는 제2 전자 장치(10b)를 제1 방향(790)으로 이동시킬 수 있다. 이 경우, 제2 전자 장치(10b)의 모션 센서에 의하여 감지된 제2 전자 장치(10b)의 움직임 정보는 제2 방향(795)의 이동이 제1 방향(790)의 이동이 상쇄된 값을 수 있다. 반면, 제2 전자 장치(10b)의 카메라(예: 사용자 컨텍스트 정보에 기반하여 선택된 카메라)를 이용하여 획득된 복수의 이미지들에 기반한 움직임 정보는 제1 방향(790)의 이동에 대응할 수 있다. 이 경우, 제2 전자 장치(10b)의 움직임 정보와 이미지 기반 움직임 정보가 서로 대응하지 않을 수 있다. 제2 전자 장치(10b)는 동작 모드를 제2 모드로 결정할 수 있다.
일 예에서, 제2 전자 장치(10b)는 카메라를 이용한 이미지 내에서 사용자(702)에 대응하는 이미지가 식별됨에 기반하여 제2 전자 장치(10b)의 동작 모드를 제2 모드로 결정할 수 있다.
예를 들어, 사용자(702)는 제2 전자 장치(10b)를 손에 든 상태에서 사용자(702)의 얼굴과 다른 사용자(미도시)의 얼굴을 제2 전자 장치(10b)를 이용하여 번갈아 촬영할 수 있다. 사용자(702)는 촬영 중에, “누가 더 나이가 많아 보여”라는 음성 입력을 발화할 수 있다. 이 경우, 전자 장치(10)는 사용자(702)를 촬영할 때 동작 모드를 제2 모드로 설정하고, 제2 모드에서 촬영된 이미지에 제2 모드에서 촬영된 이미지임을 나타내는 정보를 태깅(tag)할 수 있다. 예를 들어, 제2 모드에서 촬영된 이미지임을 나타내는 정보는, 이미지 내의 인물이 사용자(702)임을 나타내는 정보를 포함할 수 있다. 전자 장치(10)는 다른 사용자를 촬영할 때 동작 모드를 제1 모드로 설정하고, 제1 모드에서 촬영된 이미지에 제1 모드에서 촬영된 이미지임을 나타내는 정보를 태깅할 수 있다. 예를 들어, 제1 모드에서 촬영된 이미지임을 나타내는 정보는, 이미지 내의 인물이 사용자(702)가 아님을 나타내는 정보를 포함할 수 있다.
예를 들어, 사용자(702)는 “허리가 아프다”라고 발화할 수 있다. 이 경우, 제2 모드의 제2 전자 장치(10b)는 사용자(702)의 이미지를 분석할 수 있다. 제2 전자 장치(10b)는 사용자(702)의 이미지들로부터 “의자에 앉아서 책을 읽고 있는 사용자”라는 컨텍스트 기반 프롬프트를 생성할 수 있다. 제2 전자 장치(10b)는 사용자(702)의 발화에 기반하여, “허리가 아픈 사용자”라는 사용자 입력 기반 프롬프트를 생성할 수 있다. 제2 전자 장치(10b)는 컨텍스트 기반 프롬프트와 사용자 입력 기반 프롬프트를 인공지능 모델을 이용하여 처리함으로써 결과물을 생성할 수 있다. 제2 전자 장치(10b)는 결과물에 기반하여 응답을 제공할 수 있다. 응답을 제공할 때, 제2 전자 장치(10b)는 컨텍스트 기반 프롬프트에 연관된 사용자의 이미지(예: 의자에 앉아 있는 사용자의 이미지)를 함께 제공할 수 있다.
도 7c는 일 실시 예에 따른 전자 장치의 제2 모드 동작 환경을 도시한다.
도 7c를 참조하여, 일 실시 예에 따르면, 제2 모드의 제2 전자 장치(10b)는 선제적인 제안을 제공할 수 있다. 제2 모드에서, 제2 전자 장치(10b)는 사용자(703)의 음성 입력이 없는 상태에서, 사용자(703)의 변화, 사용자(703)의 표정, 사용자(703)의 포즈, 사용자(703)의 행동, 사용자(703)의 주변 환경의 변화, 또는 사용자(703)의 제스처에 기반하여 선제적인 제안을 제공할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 전자 장치(10)가 복수의 카메라들을 포함할 수 있다. 이 경우, 전자 장치(10)는 복수의 카메라들 중 사용자의 컨텍스트 정보에 기반하여 결정된 카메라를 이용할 수 있다. 예를 들어, 전자 장치(10)는 복수의 카메라들 중 사용자(703)가 식별되는 이미지를 획득하는 데 이용된 카메라를 이용할 수 있다.
제2 모드에서, 제2 전자 장치(10b)는 카메라를 이용하여 획득된 사용자(703)의 이미지를 분석할 수 있다. 제2 전자 장치(10b)는 분석된 정보에 기반하여 사용자(703)에 제안을 제공할 수 있다. 예를 들어, 제2 전자 장치(10b)는 “사용자님 피곤해 보이시는데 침실로 가시는 건 어떨까요?”와 같은 제안을 제공할 수 있다. 제2 전자 장치(10b)는 컨텍스트 정보(예: 카메라를 이용하여 획득된 이미지 및/또는 센서 회로를 이용하여 획득된 정보)에 기반한 알림을 제공할 수 있다. 제2 전자 장치(10b)는 사용자 선호 또는 사용자 설정에 따라서 지정된 상황에 대한 제안을 제공하도록 설정될 수 있다.
도 7d는 일 실시 예에 따른 전자 장치의 제3 모드 동작 환경을 도시한다.
도 7d를 참조하여, 제2 전자 장치(10b)는 사용자(704)의 가방 안에 수납될 수 있다. 이 경우, 제2 전자 장치(10b)는 제2 전자 장치(10b)가 차폐 상태임을 감지할 수 있다. 차폐 상태의 감지에 기반하여, 제2 전자 장치(10b)는 제2 전자 장치(10b)의 동작 모드를 제3 모드로 결정할 수 있다. 제3 모드에서, 제2 전자 장치(10b)는 사용자(704)의 음성 입력에 기반하여 동작할 수 있다. 예를 들어, 제2 전자 장치(10b)는 카메라를 비활성화 상태 또는 유휴 상태로 유지할 수 있다.
도 8a는 일 실시 예에 따른 전자 장치를 도시한다.
도 8a를 참조하여, 제2 전자 장치(10b)는 제1 마이크(882)(예: 도 1b의 제1 마이크(182)), 제2 마이크(883)(예: 도 1b의 제2 마이크(183)), 카메라(870)(예: 도 1b의 카메라(170)), 및 디스플레이(860)(예: 도 1b의 디스플레이(160))를 포함할 수 있다. 예를 들어, 제2 전자 장치(10b)는 본체부(801) 및 부착부(802)를 포함할 수 있다. 예를 들어, 본체부(801)와 부착부(802)는 자력에 기반하여 결합될 수 있다. 도 8a에서, 설명의 편의를 위하여 제2 전자 장치(10b)를 중심으로 예시들이 설명되나, 통상의 기술자는 유사한 구조를 갖는 임의 전자 장치에 동일한 예시들이 적용될 수 있음을 이해할 수 있을 것이다.
예를 들어, 부착부(802)는 본체부(801)와의 결합을 위한 적어도 하나의 자석을 포함할 수 있다. 부착부(802)는 배터리 및/또는 프로세싱 회로를 포함할 수 있다. 부착부(802)는 본체부(801)와 전자기적으로 통신함으로써, 본체부(801)에 전력을 공급할 수 있다. 제2 전자 장치(10b)의 구성들 중 적어도 일부를 포함할 수 있다. 부착부(802)는 본체부(801)와의 전자기적인 결합을 통하여, 본체부(801)와 함께 제2 전자 장치(10b)의 기능들을 구현할 수 있다. 예를 들어, 부착부(802)는 본체부(801)와 결합되었을 때, 배터리를 이용하여 본체부(801)에 전력을 공급할 수 있다. 부착부(802)가 결합됨에 기반하여, 제2 전자 장치(10b)는 부팅될 수 있다. 제2 전자 장치(10b)는 부팅 시에 동작 모드의 결정 동작(예: 도 5의 동작 510)을 수행할 수 있다. 일 예에서, 제2 전자 장치(10b)는 본체부(801)와 부착부(802)가 결합된 후에, 동작 모들의 결정을 위한 트리거 이벤트가 감지(예: 도 5의 동작 505)되면, 동작 모드의 결정 동작을 수행할 수 있다.
일 실시 예에 따르면, 제2 전자 장치(10b)는 다양한 형태로 사용자의 의류, 신체, 또는 사물에 장착될 수도 있다. 예를 들어, 제2 전자 장치(10b)는 부착부(802)의 부착 상태를 감지할 수 있다. 제2 전자 장치(10b)는 부착 센서(예: 도 1b의 부착 센서(145))를 이용하여 부착 상태를 감지할 수 있다. 제2 전자 장치(10b)가 다른 사물에 부착되지 않은 경우, 본체부(801)와 부착부(802)는 밀착된 상태(예: 비착용 상태)일 수 있다. 제2 전자 장치(10b)가 다른 사물에 부착(예: 착용 상태)된 경우, 본체부(801)와 부착부(802) 사이에는 다른 사물(예: 옷, 가방)에 의하여 갭(gap)이 발생될 수 있다. 본체부(801)와 부착부(802) 사이의 거리(d)가 변경됨에 따라서, 본체부(801)에서 감지할 수 있는 자력의 크기가 변경될 수 있다. 따라서, 제2 전자 장치(10b)는 착용 상태와 비착용 상태에서 발생하는 자력의 차이를 감지함으로써, 부착 상태를 감지할 수 있다. 예를 들어, 제2 전자 장치(10b)는 감지된 자력이 임계값을 초과하면 부착 상태를 비착용 상태로 결정할 수 있다. 제2 전자 장치(10b)는 감지된 자력이 상기 임계값 이하의 지정된 범위 이내의 값이면 부착 상태를 착용 상태로 결정할 수 있다.
도 8a와 관련하여, 제2 전자 장치(10b)가 자력에 기반하여 본체부(801)와 부착부(802) 사이의 부착 상태를 결정하는 동작들이 설명되었으나, 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 제2 전자 장치(10b)는 부착 상태에 따라서 기구적인 변형이 발생할 수 있는 부착 센서를 포함할 수 있다. 제2 전자 장치(10b)는 기구적인 변형을 감지함으로써 부착 상태를 식별할 수 있다.
일 실시 예에 따르면, 제2 전자 장치(10b)는 부착 상태에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 제2 전자 장치(10b)의 부착 상태가 착용 상태인 경우, 제2 전자 장치(10b)는 동작 모드를 제1 모드로 결정할 수 있다. 부착 상태가 착용 상태인 경우라도, 제2 전자 장치(10b)가 차폐된 경우, 제2 전자 장치(10b)는 동작 모드를 제3 모드로 결정할 수 있다. 부착 상태가 비착용 상태인 경우, 제2 전자 장치(10b)는 동작 모드를 제2 모드 또는 제3 모드로 결정할 수 있다. 예를 들어, 제2 전자 장치(10b)는 도 4 내지 도 6b와 관련하여 상술된 방법에 따라서 동작 모드를 결정할 수 있다.
착용 상태의 제1 모드에서, 제2 전자 장치(10b)는 라이프 로그(life-log) 동작을 수행할 수 있다. 라이프 로그 동작에 따라서, 제2 전자 장치(10b)는 사용자의 입력이 없더라도, 음성 및 영상을 상시 녹화(예: 일정 기간 동안)할 수 있다. 일 예에서, 제2 전자 장치(10b)는 사용자 설정에 따라서 라이프 로그 동작을 수행할 수 있다. 라이프 로그 동작의 수행 중에, 제2 전자 장치(10b)는 획득된 이미지를 이용하여 컨텍스트 기반 프롬프트를 생성하고, 사용자 음성에 기반하여 사용자 입력 기반 프롬프트를 생성할 수 있다.
비착용 상태에서, 제2 전자 장치(10b)는 모션 센서를 이용하여 제2 전자 장치(10b)가 캐리(carry) 상태인지 또는 스탠드(stand) 상태인지 결정할 수 있다. 캐리 상태는, 제2 전자 장치(10b)가 사용자에 의하여 캐리되는 상태로 참조될 수 있다. 제2 전자 장치(10b)는 비착용 상태에서 모션 센서를 이용하여 임계값을 초과하는 움직임을 감지하면, 제2 전자 장치(10b)가 캐리 상태인 것으로 결정할 수 있다. 스탠드 상태는, 제2 전자 장치(10b)가 고정된 위치에 거치된 상태로 참조될 수 있다. 제2 전자 장치(10b)는 비착용 상태에서 임계값을 초과하는 움직임이 감지되지 않으면, 제2 전자 장치(10b)가 스탠드 상태인 것으로 결정될 수 있다.
캐리 상태인 경우, 제2 전자 장치(10b)는 사용자 이미지의 인식 여부에 기반하여 동작 모드를 결정할 수 있다. 카메라(870)를 이용하여 획득된 이미지로부터 사용자 이미지가 인식되면, 제2 전자 장치(10b)는 동작 모드를 제1 모드로 결정할 수 있다. 카메라(870)를 이용하여 획득된 이미지로부터 사용자 이미지가 인식되지 않으면, 제2 전자 장치(10b)는 동작 모드를 제2 모드로 결정할 수 있다.
일 실시 예에 따르면, 제2 전자 장치(10b)는 부착 상태 및 부착 위치에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 제2 전자 장치(10b)는 착용 상태에서 제2 전자 장치(10b)가 부착된 위치를 감지할 수 있다. 제2 전자 장치(10b)는 자력을 감지할 수 있는 센서, 근접 센서, 및/또는 물리적인 체결 상태를 감지하도록 설정된 센서를 이용하여 제2 전자 장치(10b)가 부착된 위치를 감지할 수 있다. 제2 전자 장치(10b)는 부착된 위치에 대하여 설정된 동작 모드로 동작 모드를 결정할 수 있다. 일 예에서, 제2 전자 장치(10b)는 부착 위치가 사용자의 신체라고 결정할 수 있다. 이 경우, 제2 전자 장치(10b)는 움직임 많은 상태에서 이미지를 획득하기 위하여 카메라(870)의 설정을 제어할 수 있다. 예를 들어, 제2 전자 장치(10b)는 이미지 안정화(예: optical image stabilization, video digital image stabilization), 오토 포커스(auto focus), 화이트밸런스(white balance), 노출 값, 및/또는 셔터 스피드(shutter speed)를 조정할 수 있다.
일 실시 예에 따르면, 스탠드 상태인 경우, 제2 전자 장치(10b)는 움직임의 감지, 지정된 시간의 경과, 또는 사용자 입력에 기반하여 동작 모드를 결정(예: 도 5의 동작 510)할 수 있다. 스탠드 상태에서 지정된 값 이상의 움직임이 감지되면, 제2 전자 장치(10b)는 동작 모드를 결정될 수 있다. 스탠드 상태에서 지정된 값 이상의 움직임이나 사용자 입력이 감지되지 않은 상태로 지정된 시간이 경과되는 경우, 제2 전자 장치(10b)는 동작 모드를 결정할 수 있다. 스탠드 상태에서 사용자 입력이 감지되는 경우, 제2 전자 장치(10b)는 동작 모드를 결정할 수 있다. 스탠드 상태에서, 제2 전자 장치(10b)는 카메라(870)를 활성화하지 않고 마이크(882, 883)를 활성화한 상태로 대기할 수 있다. 대기 중에 움직임의 감지, 지정된 시간의 경과, 또는 사용자 입력이 수신되는 경우, 제2 전자 장치(10b)는 카메라(870)를 활성화하여 동작 모드를 결정할 수 있다.
일 실시 예에 따르면, 제2 전자 장치(10b)는 화자의 위치에 기반하여 동작 모드를 결정할 수 있다. 예를 들어, 제2 전자 장치(10b)는 제1 마이크(882) 및 제2 마이크(883)를 이용하여 사용자의 발화를 수신할 수 있다. 빔포밍 기술에 기반하여, 제2 전자 장치(10b)는 제2 전자 장치(10b)에 대한 발화자(예: 사용자)의 상대적인 위치를 식별할 수 있다. 발화자가 제2 전자 장치(10b)의 뒤쪽(예: 카메라(870)가 향하는 방향의 반대 방향쪽)에 위치된 경우, 제2 전자 장치(10b)는 동작 모드를 제1 모드로 결정할 수 있다. 예를 들어, 제2 전자 장치(10b)는 착용 상태에서 발화자가 뒤쪽에 위치된 경우, 동작 모드를 제1 모드로 결정할 수 있다. 제2 전자 장치(10b)는 착용 상태에서 발화자가 뒤쪽에 위치된 경우에, 카메라(870)의 시야와 사용자의 시야가 대응하는 것으로 결정할 수 있다. 부착 상태에서 발화자가 제2 전자 장치(10b)의 뒤쪽에 위치되지 않은 경우, 제2 전자 장치(10b)는 동작 모드를 제3 모드로 결정할 수 있다.
일 예에서, 제2 전자 장치(10b)는 화자의 위치에 기반한 동작 모드의 결정을 제2 전자 장치(10b)의 부착 위치에 기반하여 수행할 수 있다. 예를 들어, 제2 전자 장치(10b)는 지정된 위치에 장착된 경우, 지정된 악세서리에 장착된 경우, 또는 지정된 신체 부위에 장착된 경우에, 화자의 위치에 기반한 동작 모드의 결정을 수행할 수 있다.
도 8b는 일 실시 예에 따른 전자 장치를 도시한다.
도 8b를 참조하여, 일 실시 예에 따르면, 제2 전자 장치(10b)는 원통형 형상을 가질 수 있다. 본체부(801)와 부착부(802)가 원통형으로 형성되기 때문에, 사용자는 착용 상태 또는 비착용 상태에서 부착부(802)를 본체부(801)에 대하여 상대적으로 회전시킬 수 있다. 제2 전자 장치(10b)는 부착부(802)의 회전에 기반하여 사용자 입력을 수신할 수 있다. 예를 들어, 제2 전자 장치(10b)는 회전 방향에 기반하여 동작 모드를 설정할 수 있다.
도 9는 일 실시 예에 따른 모달리티 결정에 기반한 응답 제공 방법의 흐름도이다.
도 1b 및 도 9를 참조하여, 일 실시예에 따르면, 전자 장치(10)는 사용자 입력에 기반하여 응답을 제공할 수 있다. 예를 들어, 전자 장치(10)는 결정된 모달리티에 기반한 입력을 이용하여 응답을 제공할 수 있다.
도 9와 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 9와 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 9의 순서와 다르게 실행되거나, 도 9의 다른 동작과 실질적으로 동시에 실행될 수 있다.
예를 들어, 전자 장치(10)는 적어도 하나의 센서(예: 센서 회로(140), 카메라(170), 마이크(예: 제1 마이크(182) 및/또는 제2 마이크(183)), 메모리(130), 및 적어도 하나의 프로세서(예: 프로세서(120)를 포함할 수 있다. 적어도 하나의 프로세서는 적어도 하나의 센서, 카메라(170), 마이크, 및 메모리(130)와 통신적으로 연결될 수 있다. 메모리(130)는 적어도 하나의 프로세서에 의하여 개별적으로(individually) 또는 조합적으로(collectively) 실행되었을 때, 전자 장치(10)가 후술되는 동작들을 수행하도록 하는 인스트럭션들을 저장할 수 있다.
동작 905에서, 전자 장치(10)는 음성 입력을 획득할 수 있다. 예를 들어, 전자 장치(10)는 마이크를 이용하여 사용자로부터 음성 입력을 획득할 수 있다.
동작 910에서, 전자 장치(10)는 카메라(170)의 시야가 사용자의 시야에 대응하는지 결정할 수 있다. 예를 들어, 전자 장치(10)는 복수의 마이크들(예: 제1 마이크(182) 및 제2 마이크(183))을 이용하여 획득된 음성 입력으로부터 전자 장치(10)에 대한 사용자의 상대적인 위치를 식별할 수 있다. 상대적인 위치가 카메라(170)의 시야에 대응하는 방향의 반대 방향(예: 전자 장치(10)의 뒤쪽 방향)에 위치된 경우, 전자 장치(10)는 카메라(170)의 시야와 사용자의 시야가 서로 대응하는 것으로 결정할 수 있다. 예를 들어, 전자 장치(10)는 모션 센서(143)를 이용하여 전자 장치(10)의 움직임 정보를 획득하고, 카메라(170)를 이용하여 복수의 이미지들을 획득할 수 있다. 전자 장치(10)의 움직임 정보와 복수의 이미지들로부터 획득된 이미지 기반 움직임 정보가 서로 대응하는 경우, 전자 장치(10)는 카메라(170)의 시야와 사용자의 시야가 서로 대응하는 것으로 결정할 수 있다. 일 예에서, 전자 장치(10)는 음성 입력에 기반하여, 카메라(170)를 이용하여 적어도 하나의 이미지를 획득할 수 있다. 전자 장치(10)는 적어도 하나의 이미지로부터 사용자에 대응하는 이미지가 식별되면 카메라(170)의 시야와 사용자의 시야가 서로 대응하는 것으로 결정할 수 있다. 예를 들어, 전자 장치(10)는 카메라(170)의 차폐가 감지되는 경우, 카메라(170)의 시야와 사용자의 시야가 서로 대응하지 않는 것으로 결정할 수 있다. 도 1b와 관련하여 상술된 바와 같이, 카메라(170)는 복수의 카메라들로부터 선택된 카메라를 포함할 수 있다.
동작 910은 동작 모드의 결정 동작에 대응할 수 있다. 예를 들어, 전자 장치(10)는 도 4 내지 도 8b와 관련하여 상술된 동작 모드 결정 동작에 따라서 동작 모드를 결정할 수 있다. 예를 들어, 동작 모드가 제1 모드 또는 제2 모드로 결정되는 경우, 전자 장치(10)는 카메라(170)의 시야가 사용자의 시야에 대응하는 것으로 결정할 수 있다. 예를 들어, 동작 모드가 제3 모드로 결정되는 경우, 전자 장치(10)는 카메라(170)의 시야가 사용자의 시야에 대응하지 않는 것으로 결정할 수 있다.
상술된 바와 같이, 전자 장치(10)는 트리거 이벤트에 기반하여 동작 910을 수행할 수 있다. 예를 들어, 전자 장치(10)는 부착 센서(145)를 이용하여 감지된 전자 장치(10)의 부착 상태가 착용 상태인 경우, 카메라(170)의 시야가 사용자의 시야에 대응하는지 결정할 수 있다. 예를 들어, 전자 장치(10)는 부착 센서(145)를 이용하여 감지된 전자 장치(10)의 부착 상태가 비착용 상태인 경우, 카메라(170)를 이용하여 적어도 하나의 이미지를 획득할 수 있다. 전자 장치(10)는 적어도 하나의 이미지로부터 사용자에 대응하는 이미지가 식별되면 상기 카메라(170)의 시야와 사용자의 시야가 대응하는 것으로 결정할 수 있다.
예를 들어, 전자 장치(10)는 본체부(예: 도 8a 및 8b의 본체부(801))와 부착부(예: 도 8a 및 8b의 부착부(802))를 포함할 수 있다. 부착부는 본체부와 자력에 기반하여 결합되도록 설정될 수 있다. 부착 센서(145)는 자력에 기반하여 전자 장치(10)의 부착 상태를 감지하도록 설정될 수 있다.
상술된 바와 같이, 음성 입력의 획득(예: 동작 905) 전에 전자 장치(10)의 동작 모드가 결정된 상태일 수 있다. 이 경우, 동작 910은 모달리티 결정에 기반한 응답 제공 방법으로부터 생략될 수 있다. 전자 장치(10)는 결정된 동작 모드에 따라서 다중 모달 입력 또는 싱글 모달 입력을 이용하여 프롬프트를 생성할 수 있다.
카메라(170)의 시야가 사용자의 시야에 대응하지 않는 경우(예: 동작 910-NO), 전자 장치(10)는 참조 포인트 A에 따른 동작들을 수행할 수 있다. 참조 포인트 A에 따른 동작들은 도 10과 관련하여 후술될 수 있다.
카메라(170)의 시야가 사용자의 시야에 대응하는 경우(예: 동작 910-YES), 동작 915에서, 전자 장치(10)는 이미지(예: 카메라(170)를 이용하여 획득된 이미지) 및 음성 입력에 기반하여 적어도 하나의 프롬프트(예: 적어도 하나의 제1 프롬프트)를 생성할 수 있다. 적어도 하나의 프롬프트는 사용자 입력 기반 프롬프트 및 컨텍스트 정보 기반 프롬프트를 포함할 수 있다. 예를 들어, 전자 장치(10)는 음성 입력에 기반하여 사용자 입력 기반 프롬프트를 생성할 수 있다. 전자 장치(10)는 이미지를 이용하여 컨텍스트 정보 기반 프롬프트(예: 이미지를 인식하여 생성된 텍스트 데이터 또는 이미지 데이터 자체)를 생성할 수 있다. 예를 들어, 전자 장치(10)는 획득된 이미지에 대한 이미지 인식을 수행하고, 이미지 인식의 결과(예: 텍스트 데이터) 및 음성 입력(예: 음성 입력으로부터 인식된 텍스트 데이터)의 조합에 기반하여 적어도 하나의 프롬프트를 생성할 수 있다. 전자 장치(10)는 도 4의 동작 420에 따라서 적어도 하나의 프롬프트를 생성할 수 있다.
일 예시에서, 전자 장치(10)는 카메라(170)의 시야가 사용자의 시야에 대응하지 않는 경우(예: 동작 910-NO)에도, 동작 모드를 멀티 모달을 이용하는 동작 모드로(예: 제1 모드 또는 제2 모드) 결정할 수 있다. 예를 들어, 전자 장치(10)는 도 6b와 관련하여 상술된 방법에 따라서 동작 모드를 결정할 수 있다. 카메라(170)의 시야와 사용자의 시야가 대응하지 않는 경우의 동작 모드는 별개의 제4 모드로 참조될 수 있다. 이 경우, 제4 모드는 음성 입력과 카메라 입력을 모두 이용하는 다중 모달 모드로 참조될 수 있다. 제4 모드에서, 전자 장치(10)는 사용자의 발화에 관련된 이미지에 대한 설명을 포함하는 프롬프트를 생성할 수 있다. 전자 장치(10)는 프롬프트를 생성형 인공지능 모델에 입력함으로써 응답을 생성하고, 생성된 응답을 제공할 수 있다.
동작 920에서, 전자 장치(10)는 결과 데이터(예: 제1 결과 데이터)를 획득할 수 있다. 전자 장치(10)는 생성된 적어도 하나의 프롬프트를 인공지능 모델(예: 도 2의 인공지능 모델 DB(280)의 인공지능 모델)에 입력함으로써 결과 데이터를 획득할 수 있다.
동작 925에서, 전자 장치(10)는 결과 데이터에 기반하여 음성 입력에 연관된 응답을 제공할 수 있다. 예를 들어, 전자 장치(10)는 도 2의 I/O 인터페이스(210)와 관련하여 상술된 방법들에 따라서 응답을 제공할 수 있다. 도 2의 출력 관리자(250)와 관련하여 상술된 바와 같이, 전자 장치(10)는 결과 데이터에 대한 미세 조정(fine tuning)을 수행할 수 있다.
도 10는 일 실시 예에 따른 모달리티 결정에 기반한 응답 제공 방법의 흐름도이다.
도 1b 및 도 10을 참조하여, 일 실시예에 따르면, 전자 장치(10)는 제2 모드 또는 제3 모드에 따라서 응답을 제공할 수 있다. 도 10과 관련하여 후술되는 동작은 도 1b의 전자 장치(10)의 동작으로 참조될 수 있다. 도 10과 관련하여 후술되는 동작들의 순서는 일 예시로서 본 개시의 실시 예들이 이에 제한되는 것은 아니다. 예를 들어, 적어도 일부의 동작은 도 10의 순서와 다르게 실행되거나, 도 10의 다른 동작과 실질적으로 동시에 실행될 수 있다. 카메라(170)의 시야가 사용자의 시야에 대응하지 않는 경우(예: 도 9의 동작 910-NO), 전자 장치(10)는 도 10의 동작들을 수행할 수 있다.
동작 1010에서, 전자 장치(10)는 음성 입력(예: 도 9의 동작 905의 음성 입력)에 기반하여 적어도 하나의 프롬프트(예: 제2 프롬프트)를 생성할 수 있다. 이 경우, 전자 장치(10)는 컨텍스트 정보(예: 이미지)를 이용하지 않고, 음성 입력에만 기반하여 사용자 입력 기반 프롬프트를 생성할 수 있다. 예를 들어, 전자 장치(10)는 동작 425에 따라서 프롬프트를 생성할 수 있다.
동작 1015에서, 전자 장치(10)는 결과 데이터(예: 제2 결과 데이터)를 획득할 수 있다. 전자 장치(10)는 생성된 적어도 하나의 프롬프트를 인공지능 모델(예: 도 2의 인공지능 모델 DB(280)의 인공지능 모델)에 입력함으로써 결과 데이터를 획득할 수 있다.
동작 1020에서, 전자 장치(10)는 결과 데이터에 기반하여 음성 입력에 연관된 응답을 제공할 수 있다. 예를 들어, 전자 장치(10)는 도 2의 I/O 인터페이스(210)와 관련하여 상술된 방법들에 따라서 응답을 제공할 수 있다. 도 2의 출력 관리자(250)와 관련하여 상술된 바와 같이, 전자 장치(10)는 결과 데이터에 대한 미세 조정(fine tuning)을 수행할 수 있다.
도 9 및 도 10과 관련하여 상술된 전자 장치(10)의 동작들은 도 1a 내지 도 8b와 관련하여 상술된 실시 예들 중의 일부로서, 통상의 기술자는 전자 장치(10)가 도 1a 내지 도 8b와 관련하여 상술된 실시예들을 수행할 수 있음을 이해할 수 있을 것이다. 예를 들어, 도 9 및 도 10의 실시예에 있어서, 전자 장치(10)는 제1 모드와 제2 모드를 구별하고 있지 아니하나, 전자 장치(10)는 카메라(170)의 시야가 사용자의 시야에 대응하는 경우, 제1 모드 또는 제2 모드를 결정하기 위한 동작을 수행할 수 있다. 예를 들어, 제1 모드에서 생성된 프롬프트와 달리, 제2 모드에서 생성된 프롬프트에는 이미지가 사용자의 이미지에 해당함을 나타내는 정보가 포함될 수 있다. 따라서, 제1 모드에서의 결과 데이터와 제2 모드에서의 결과 데이터가 상이할 수 있다.
도 11은, 본 문서 내에서 설명된 동작들을 수행할 수 있는 예시적인(exemplary) 전자 장치(1100)의 블록도이다.
도 11을 참조하면, 전자 장치(1100)는, 노트북(1190), 다양한 폼 팩터들을 가지는 스마트폰들(1191)((예: 바 타입의 스마트폰(1191-1), 폴더블 타입의 스마트폰(1191-2), 또는 슬라이더블(또는 롤러블) 타입의 스마트폰(1191-3)), 태블릿(1192), 셀룰러 전화(미도시), 및 기타 유사 컴퓨팅 장치들(미도시)과 같은, 다양한 형태들의 전자 장치들 중 하나일 수 있다. 도 11 내에서 도시된 구성요소들, 그들의 관계들, 및 그들의 기능들은, 예시적일 뿐이며, 본 문서 내에서 설명되거나 청구된 구현들을 제한하는 것이 아니다. 전자 장치(1100)는, 모바일 장치, 사용자 장치, 다기능 장치, 휴대용 장치 또는 서버로 참조될 수 있다.
전자 장치(1100)는, 적어도 하나의 프로세서(1110)(이하, 프로세서(1110)으로 참조), 적어도 하나의 메모리(1120)(이하, 메모리(1120)으로 참조), 적어도 하나의 디스플레이(1140)(이하, 디스플레이(1140)으로 참조), 적어도 하나의 이미지 센서(1150)(이하, 이미지 센서(1150)으로 참조), 적어도 하나의 통신 회로(1160)(이하, 통신 회로(1160)으로 참조), 및/또는 적어도 하나의 센서(1170)(이하, 센서(1170)으로 참조)를 포함하는 구성요소들을 포함할 수 있다. 상기 구성요소들은 단지 예시적인 것이다. 예를 들어, 전자 장치(1100)는 다른 구성요소들(예: PMIC(power management integrated circuitry), 오디오 처리 회로, 안테나, 재충전가능한(rechargeable) 배터리 또는 입출력 인터페이스)을 포함할 수 있다. 예를 들어, 일부 구성요소들은 전자 장치(1100)로부터 생략될 수 있다. 예를 들어, 몇몇 구성요소들은 하나의 구성요소로 통합될 수 있다.
프로세서(1110)는, 하나 이상의 IC(integrated circuit(또는 circuitry)) 칩으로 구현될 수 있고, 다양한 데이터 처리들을 실행할 수 있다. 프로세서(1110)는 적어도 하나의 전기적 회로를 포함할 수 있고, 메모리(1120) 내에 저장된 인스트럭션들(또는 프로그램, 데이터, 기타 등등)을 개별적으로 또는 집합적으로 분산 처리할 수 있다. 프로세서(1110)는 하나 이상의 프로세싱 회로들을 포함하는 프로세서 집합체를 포함할 수 있다. 프로세서(1110)는, 전자 장치(1100)의 하나 이상의 구성요소들(예: 메모리(1120), 디스플레이(1140), 이미지 센서(1150), 통신 회로(1160), 및/또는 센서(1170))의 수행(performance) 및 동작들을 제어하기 위해 작동적인(operative) 어느(any) 프로세싱 회로를 포함할 수 있다. 예를 들어, 프로세서(1110)(예: AP(application processor))는, SoC(system on chip)(예를 들어, 하나의 칩 또는 칩셋)로 구현될 수 있다. 예를 들어, 프로세서(1110)는, 다수의 코어들(또는 적어도 하나의 코어 회로), 다수의 칩들 또는 다수의 칩셋들로 구현될 수 있다. 예를 들어, 프로세서(1110)는, 하나 이상의 프로세싱 회로들을 포함할 수 있다. 예를 들어, 프로세서(1110)는 본 개시의 여러 기능들을 개별적으로, 및/또는 집합적으로(collectively) 수행하도록 구성된 하나 이상의 프로세싱 회로를 포함할 수 있다. 제한하지 않는 예로, 프로세서(1110)의 적어도 일부분은 전자 장치(1100)의 제1 칩에 포함되고, 프로세서(1110)의 적어도 다른 부분은 전자 장치(1100)의 상기 제1 칩과 다른 전자 장치(1100)의 제2 칩에 포함될 수 있다.
예를 들어, 프로세서(1110)는, CPU(central processing unit)(1111), GPU(graphics processing unit)(1112), NPU(neural processing unit)(1113), ISP(image signal processor)(1114), 디스플레이 컨트롤러(1115), 메모리 컨트롤러(1116), 스토리지(storage) 컨트롤러(1117), CP(communication processor)(1118), 및/또는 센서 인터페이스(1119)를 포함할 수 있다. 프로세서(1110)의 이러한 구성요소들은, 단지 예시적인 것이다. 예를 들어, 프로세서(1110)는, 다른 구성요소들을 더 포함할 수 있다. 예를 들어, 프로세서(1110)의 몇몇 구성요소들은, 프로세서(1110)로부터 생략될 수 있다. 예를 들어, 프로세서(1110)의 몇몇 구성요소들은, 프로세서(1110) 외부에서 전자 장치(1100)의 별도의 구성요소로 포함될 수 있다. 예를 들어, 프로세서(1110)의 일부 구성요소들(예: 메모리 컨트롤러(1116))은, 다른 구성요소들(예: 메모리(1120)의 적어도 일부, 인터페이스(예: 전자 장치(100)의 적어도 하나의 구성요소에 연결하기 위해 이용가능함), 디스플레이(1140) 및/또는 이미지 센서(1150)) 내에 포함될 수 있다.
프로세서(1110)는, 메모리(1120) 내에 저장된 인스트럭션들을 실행함으로써 다양한 동작들을 수행하도록 전자 장치(1100)의 다른 구성요소들을 야기할 수 있다. CPU(1111)(또는 중앙 처리 회로)는 메모리(1120)(예: 휘발성 메모리(1121) 및/또는 비휘발성 메모리(1122)) 내에 저장된 인스트럭션들의 실행에 기반하여 프로세서(1110)의 구성요소들을 제어하도록 구성될 수 있다. GPU(1112)(또는 그래픽 처리 회로)는 병렬 연산(예: 렌더링(rendering))들을 실행하도록 구성될 수 있다. NPU(1113)(또는 뉴럴 처리 회로, 또는 AI(artificial intelligence) 칩)는 인공 지능 모델을 위한 연산(예: 합성곱 연산(convolution computation))들을 실행하도록 구성될 수 있다. ISP(1114)(또는 이미지 신호 처리 회로)는 이미지 센서(1150)를 통해 획득된 원시 이미지(raw image)를 전자 장치(1100) 내의 구성요소 또는 프로세서(1110)의 구성요소를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 디스플레이 컨트롤러(1115)(또는 디스플레이 제어 회로, 또는 DPU(display processing unit))는 CPU(1111), GPU(1112), ISP(1114), 또는 메모리(1120)(예: 휘발성 메모리(1121))로부터 획득된 이미지를 디스플레이(1140)를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 메모리 컨트롤러(1116)(또는 메모리 제어 회로)는 휘발성 메모리(1121)로부터 데이터를 읽고 데이터를 휘발성 메모리(1121)에 기록하는 것을 제어하도록 구성될 수 있다. 스토리지 컨트롤러(1117)(또는 스토리지 제어 회로)는, 비휘발성 메모리(1122)로부터 데이터를 읽고 데이터를 비휘발성 메모리(1122)에 기록하는 것을 제어하도록 구성될 수 있다. CP(1118)(통신 처리 회로)는 프로세서(1110)의 구성요소로부터 획득된 데이터를 통신 회로(1160)를 통해 다른 전자 장치에게 송신하는 것을 위해 적합한 포맷으로 처리하거나, 다른 전자 장치로부터 통신 회로(1160)를 통해 획득된 데이터를 프로세서(1110)의 구성요소의 처리를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 예를 들어, 통신 회로(1160)는 하나 이상의 통신 회로들을 포함할 수 있다. 센서 인터페이스(1119)(또는 센싱 데이터 처리 회로, 센서 허브)는 센서(1170)를 통해 획득된, 전자 장치(1100)의 상태 및/또는 전자 장치(1100) 주변의 상태에 대한 데이터를 프로세서(1110)의 구성요소를 위해 적합한 포맷으로 처리하도록 구성될 수 있다.
메모리(1120)는, 하나 이상의 저장 매체들(storage medium)(또는 하나 이상의 저장 장치들)을 포함할 수 있다. 예를 들면, 메모리(1120)는, 하나 이상의 저장 매체들을 포함하는 메모리 집합체를 포함할 수 있다. 예를 들면, 상기 하나 이상의 저장 매체들은, 하드 드라이브, 플래시 메모리, ROM(read-only memory)과 같은 영구 메모리(permanent memory)(예, 비휘발성 메모리(1122)), RAM(random access memory)과 같은 반-영구 메모리(semi-permanent memory)(예, 휘발성 메모리(1121)), 어느 다른 적합한 유형(any other suitable type)의 저장소(또는 저장 집합체(storage assembly)), 또는 이들의 어떤 조합(any combination thereof)을 포함할 수 있다. 메모리(1120)는, 전자 장치(1100)의 기능(function or feature)을 위한 데이터를 일시적으로(temporarily) 저장하기 위해 이용되는 하나 이상의 다른 유형들(one or more different types)의 메모리인 캐시 메모리를 포함할 수 있다. 제한되지 않는 예로, 상기 캐시 메모리는 프로세서(1110) 내에 포함될 수 있다. 메모리(1120)는, 전자 장치(1100) 안에 고정적으로(fixedly) 임베디드될 수 있거나, 반복적으로(repeatedly) 전자 장치(1100) 안으로 삽입되고 전자 장치(1100)로부터 제거될 수 있는 하나 이상의 적합한 유형의 구성요소들(onto one or more suitable types of components)(예: SIM(subscriber identity module) 카드 및/또는 SD(secure digital) 카드)에 통합될(incorporated) 수 있다.
예를 들어, 메모리(1120)는, 운영 체제(operating system)(또는 시스템)소프트웨어 어플리케이션, 펌웨어 소프트웨어 어플리케이션, 드라이버 소프트웨어 어플리케이션, 플러그인(예, 애드-인, 애드-온, 및/또는 어플릿) 소프트웨어 어플리케이션, 및/또는 어느 다른(any other) 적합한(suitable) 소프트웨어 어플리케이션들과 같은 하나 이상의 소프트웨어 어플리케이션들을 저장할 수 있다. 예를 들어, 상기 하나 이상의 소프트웨어 어플리케이션들은, 프로세서(1110)에 의해 실행가능한 인스트럭션들을 포함할 수 있다. 예를 들어, 메모리(1120)는, API(application programming interface)에 의해 호출가능한 인스트럭션들을 저장할 수 있다. 예를 들어, 메모리(1120)는, 라이브러리 내에서 인스트럭션들을 저장할 수 있다.
Claims (15)
- 전자 장치(10)에 있어서,적어도 하나의 센서(140);카메라(170);마이크(182, 183);메모리(130); 및상기 적어도 하나의 센서, 상기 카메라, 상기 마이크, 및 상기 메모리와 통신적으로(communicatively) 연결된 적어도 하나의 프로세서(120)를 포함하고,상기 메모리는 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 마이크를 이용하여 사용자로부터 음성 입력을 획득하고,상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는지 결정하고,상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 경우, 상기 카메라에 의하여 획득된 이미지 및 상기 음성 입력에 기반한 적어도 하나의 제1 프롬프트를 생성하고,상기 적어도 하나의 제1 프롬프트를 인공지능 모델에 입력함으로써 제1 결과 데이터를 획득하고,상기 제1 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하도록 하는 인스트럭션들을 저장하는, 전자 장치.
- 제 1 항에 있어서,복수의 마이크들을 포함하고,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 복수의 마이크들을 이용하여 상기 음성 입력을 획득함으로써 상기 전자 장치에 대한 상기 사용자의 상대적인 위치를 식별하고,상기 상대적인 위치가 상기 카메라의 시야에 대응하는 방향의 반대 방향에 위치된 경우에 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하도록 하는, 전자 장치.
- 제 1 항에 있어서,모션 센서를 더 포함하고,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 모션 센서를 이용하여 상기 전자 장치의 움직임(movement) 정보를 획득하고,상기 카메라를 이용하여 복수의 이미지들을 획득함으로써 상기 복수의 이미지들로부터 이미지 기반 움직임 정보를 획득하고,상기 전자 장치의 움직임 정보와 상기 이미지 기반 움직임 정보가 서로 대응하는 경우, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하도록 하는, 전자 장치.
- 제 1 항에 있어서,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 음성 입력에 기반하여, 상기 카메라를 이용하여 적어도 하나의 이미지를 획득하고,상기 적어도 하나의 이미지로부터 상기 사용자에 대응하는 이미지가 식별되면 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하도록 하는, 전자 장치.
- 제 1 항에 있어서,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하지 않는 경우, 상기 음성 입력에 기반한 제2 프롬프트를 생성하고,상기 제2 프롬프트를 인공지능 모델에 입력함으로써 제2 결과 데이터를 획득하고,상기 제2 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하도록 하는, 전자 장치.
- 제 1 항에 있어서,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 카메라의 차폐가 감지되는 경우, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하지 않는 것으로 결정하도록 하는, 전자 장치.
- 제 1 항에 있어서,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 획득된 이미지에 대한 이미지 인식을 수행하고,상기 이미지 인식의 결과 및 상기 음성 입력의 조합에 기반하여 상기 적어도 하나의 제1 프롬프트를 생성하도록 하는, 전자 장치.
- 제 1 항에 있어서,상기 전자 장치의 부착 상태를 감지하도록 설정된 부착 센서(145)를 더 포함하고,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 부착 센서를 이용하여 감지된 부착 상태가 착용 상태인 경우, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하도록 하는, 전자 장치.
- 제 8 항에 있어서,상기 인스트럭션들은 상기 적어도 하나의 프로세서에 의하여 개별적으로 또는 조합적으로 실행되었을 때 상기 전자 장치가:상기 부착 센서를 이용하여 감지된 상기 부착 상태가 비착용 상태인 경우, 상기 카메라를 이용하여 적어도 하나의 이미지를 획득하고,상기 적어도 하나의 이미지로부터 상기 사용자에 대응하는 이미지가 식별되면 상기 카메라의 시야와 상기 사용자의 시야가 대응하는 것으로 결정하도록 하는, 전자 장치.
- 전자 장치의 모달리티 결정에 기반한 응답 제공 방법 방법으로서,음성 입력을 획득하는 동작; 및상기 전자 장치의 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작을 포함하고,상기 방법은, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 경우:상기 카메라에 의하여 획득된 이미지 및 상기 음성 입력에 기반한 적어도 하나의 제1 프롬프트를 생성하는 동작;상기 적어도 하나의 제1 프롬프트를 인공지능 모델에 입력함으로써 제1 결과 데이터를 획득하는 동작; 및상기 제1 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하는 동작을 더 포함하는, 방법.
- 제 10 항에 있어서,상기 전자 장치의 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작은:상기 전자 장치의 복수의 마이크들을 이용하여 상기 음성 입력을 획득함으로써 상기 전자 장치에 대한 상기 사용자의 상대적인 위치를 식별하는 동작; 및상기 상대적인 위치가 상기 카메라의 시야에 대응하는 방향의 반대 방향에 위치된 경우에 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하는 동작을 포함하는, 방법.
- 제 10 항에 있어서,상기 전자 장치의 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작은:상기 전자 장치의 모션 센서를 이용하여 상기 전자 장치의 움직임(movement) 정보를 획득하는 동작;상기 카메라를 이용하여 복수의 이미지들을 획득함으로써 상기 복수의 이미지들로부터 이미지 기반 움직임 정보를 획득하는 동작; 및상기 전자 장치의 움직임 정보와 상기 이미지 기반 움직임 정보가 서로 대응하는 경우, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하는 동작을 포함하는, 방법.
- 제 10 항에 있어서,상기 전자 장치의 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작은:상기 음성 입력에 기반하여, 상기 카메라를 이용하여 적어도 하나의 이미지를 획득하는 동작; 및상기 적어도 하나의 이미지로부터 상기 사용자에 대응하는 이미지가 식별되면 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하는 것으로 결정하는 동작을 포함하는, 방법.
- 제 10 항에 있어서,상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하지 않는 경우, 상기 음성 입력에 기반한 제2 프롬프트를 생성하는 동작;상기 제2 프롬프트를 인공지능 모델에 입력함으로써 제2 결과 데이터를 획득하는 동작; 및상기 제2 결과 데이터에 기반하여 상기 음성 입력에 연관된 응답을 제공하는 동작을 더 포함하는, 방법.
- 제 10 항에 있어서,상기 전자 장치의 카메라의 시야와 사용자의 시야가 서로 대응하는지 결정하는 동작은:상기 카메라의 차폐가 감지되는 경우, 상기 카메라의 시야와 상기 사용자의 시야가 서로 대응하지 않는 것으로 결정하는 동작을 포함하는, 방법.
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2024-0084956 | 2024-06-28 | ||
| KR20240084956 | 2024-06-28 | ||
| KR10-2024-0097453 | 2024-07-23 | ||
| KR1020240097453A KR20260002075A (ko) | 2024-06-28 | 2024-07-23 | 모달리티 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026005387A1 true WO2026005387A1 (ko) | 2026-01-02 |
Family
ID=98222265
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2025/008529 Pending WO2026005387A1 (ko) | 2024-06-28 | 2025-06-19 | 모달리티 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026005387A1 (ko) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20210020219A (ko) * | 2019-08-13 | 2021-02-24 | 삼성전자주식회사 | 대용어(Co-reference)를 이해하는 전자 장치 및 그 제어 방법 |
| KR20210041131A (ko) * | 2018-09-28 | 2021-04-14 | 애플 인크. | 응시 정보를 사용한 디바이스 제어 |
| KR20220002066A (ko) * | 2020-06-30 | 2022-01-06 | 베이징 바이두 넷컴 사이언스 테크놀로지 컴퍼니 리미티드 | 이미지 문답 방법, 장치, 컴퓨터 장비, 컴퓨터 판독가능 저장 매체 및 컴퓨터 프로그램 |
| US11323665B2 (en) * | 2017-03-31 | 2022-05-03 | Ecolink Intelligent Technology, Inc. | Method and apparatus for interaction with an intelligent personal assistant |
| KR20230088436A (ko) * | 2020-10-20 | 2023-06-19 | 로비 가이드스, 인크. | 눈 움직임에 기반한 확장된 현실 환경 상호작용의 방법 및 시스템 |
-
2025
- 2025-06-19 WO PCT/KR2025/008529 patent/WO2026005387A1/ko active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11323665B2 (en) * | 2017-03-31 | 2022-05-03 | Ecolink Intelligent Technology, Inc. | Method and apparatus for interaction with an intelligent personal assistant |
| KR20210041131A (ko) * | 2018-09-28 | 2021-04-14 | 애플 인크. | 응시 정보를 사용한 디바이스 제어 |
| KR20210020219A (ko) * | 2019-08-13 | 2021-02-24 | 삼성전자주식회사 | 대용어(Co-reference)를 이해하는 전자 장치 및 그 제어 방법 |
| KR20220002066A (ko) * | 2020-06-30 | 2022-01-06 | 베이징 바이두 넷컴 사이언스 테크놀로지 컴퍼니 리미티드 | 이미지 문답 방법, 장치, 컴퓨터 장비, 컴퓨터 판독가능 저장 매체 및 컴퓨터 프로그램 |
| KR20230088436A (ko) * | 2020-10-20 | 2023-06-19 | 로비 가이드스, 인크. | 눈 움직임에 기반한 확장된 현실 환경 상호작용의 방법 및 시스템 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020171568A1 (en) | Electronic device and method for changing magnification of image using multiple cameras | |
| WO2020130691A1 (en) | Electronic device and method for providing information thereof | |
| WO2019107981A1 (en) | Electronic device recognizing text in image | |
| WO2018009029A1 (en) | Electronic device and operating method thereof | |
| WO2021025509A1 (en) | Apparatus and method for displaying graphic elements according to object | |
| WO2019156480A1 (ko) | 시선에 기반한 관심 영역 검출 방법 및 이를 위한 전자 장치 | |
| WO2020091538A1 (ko) | 저전력 상태에서 디스플레이를 통해 화면을 표시하기 위한 전자 장치 및 그의 동작 방법 | |
| WO2020091248A1 (ko) | 음성 명령에 응답하여 컨텐츠를 표시하기 위한 방법 및 그 전자 장치 | |
| WO2018026142A1 (ko) | 홍채 센서의 동작을 제어하는 방법 및 이를 위한 전자 장치 | |
| WO2021080360A1 (en) | Electronic device and method for controlling display operation thereof | |
| WO2020032524A1 (ko) | 전자 장치와 연결된 스타일러스 펜과 연관된 정보를 출력하기 위한 장치 및 그에 관한 방법 | |
| WO2018135750A1 (ko) | 전자 장치 및 전자 장치 제어 방법 | |
| WO2021096219A1 (ko) | 카메라를 포함하는 전자 장치 및 그의 방법 | |
| WO2020027614A1 (ko) | 스타일러스 펜의 입력을 표시하기 위한 방법 및 그 전자 장치 | |
| WO2020171333A1 (ko) | 이미지 내의 오브젝트 선택에 대응하는 서비스를 제공하기 위한 전자 장치 및 방법 | |
| WO2022010193A1 (ko) | 이미지 개선을 위한 전자장치 및 그 전자장치의 카메라 운용 방법 | |
| WO2022114548A1 (ko) | 플렉서블 디스플레이의 사용자 인터페이스 제어 방법 및 장치 | |
| WO2022245037A1 (ko) | 이미지 센서 및 동적 비전 센서를 포함하는 전자 장치 및 그 동작 방법 | |
| WO2021118229A1 (en) | Information providing method and electronic device for supporting the same | |
| WO2026005387A1 (ko) | 모달리티 결정에 기반한 응답 제공 방법 및 이를 위한 전자 장치 | |
| WO2019035504A1 (ko) | 이동 단말기 및 그 제어 방법 | |
| WO2025023764A1 (ko) | 가상 현실 공간을 제공하는 전자 장치 및 전자 장치에서 가상 진동 소리를 출력하기 위한 방법 및 비 일시적 저장 매체 | |
| WO2021107200A1 (ko) | 이동 단말기 및 이동 단말기 제어 방법 | |
| WO2018110786A1 (en) | Mobile terminal and method for controlling the same | |
| WO2020022829A1 (ko) | 사용자 입력을 지원하는 전자 장치 및 전자 장치의 제어 방법 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25827352 Country of ref document: EP Kind code of ref document: A1 |