WO2026005331A1 - 영상 인식 방법 및 이를 지원하는 전자 장치 - Google Patents
영상 인식 방법 및 이를 지원하는 전자 장치Info
- Publication number
- WO2026005331A1 WO2026005331A1 PCT/KR2025/007732 KR2025007732W WO2026005331A1 WO 2026005331 A1 WO2026005331 A1 WO 2026005331A1 KR 2025007732 W KR2025007732 W KR 2025007732W WO 2026005331 A1 WO2026005331 A1 WO 2026005331A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- electronic device
- image data
- data
- objects
- processor
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F1/00—Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
- G06F1/16—Constructional details or arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/732—Query formulation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/738—Presentation of query results
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
Definitions
- Embodiments disclosed in this document relate to an image recognition method and an electronic device supporting the same.
- electronic devices can provide image recognition capabilities that recognize objects within image data and apply them to various services.
- the electronic device can provide information (e.g., object attribute information) about a target object indicated by a user among multiple objects contained in image data.
- an electronic device can identify the direction in which a user's gaze or finger is pointed, and identify an object corresponding to the direction in which the user's gaze or finger is pointed in image data as a target object.
- electronic devices can identify a single object as the target object, whether the user's gaze is focused on it or the user's finger is pointing at it. For example, electronic devices have difficulty simultaneously identifying multiple target objects within image data and providing relevant information about them.
- an electronic device includes at least one processor, a camera, a microphone, and a memory (313) operatively connected to the at least one processor, the camera, and the microphone and storing at least one command, wherein the at least one command, when individually or collectively executed by the at least one processor, causes the electronic device to: acquire image data through the camera while speech data including a designated word related to designation of an object is acquired through the microphone; identify, among a plurality of objects included in the image data, at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered; and provide associated information for the at least two objects.
- An operating method of an electronic device may include an operation of acquiring speech data including a designated word related to designation of an object, an operation of acquiring image data while the speech data is acquired, an operation of identifying at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and an operation of providing associated information about the at least two objects.
- a computer-readable recording medium may be configured such that when executed by an electronic device, the electronic device obtains speech data including a designated word related to designation of an object, obtains image data while the speech data is being obtained, identifies at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provides associated information about the at least two objects.
- An image recognition system includes a first electronic device and a second electronic device, wherein the first electronic device is configured to acquire image data through a camera while acquiring speech data including a designated word related to designation of an object through a microphone, identify at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provide associated information about the at least two objects, and the second electronic device is configured to provide information related to a posture of the second electronic device acquired through a sensor to the first electronic device, and the first electronic device may be configured to identify the at least two objects corresponding to a gesture identified in the image data when the information related to the posture of the second electronic device satisfies a designated condition.
- Electronic devices can reduce resource waste and enable selection of multiple objects within image data and provision of associated information therefor.
- FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.
- FIG. 2A is a drawing for explaining an image recognition function of an electronic device according to various embodiments.
- FIG. 2b is a diagram for explaining information provided through an image recognition function according to various embodiments.
- FIG. 2c is a drawing for explaining an image recognition function of an electronic device according to various embodiments.
- FIG. 2D is a diagram illustrating a configuration for an electronic device according to various embodiments to determine whether a target object is suitable as a comparison target.
- FIG. 2e is a diagram illustrating a configuration for guiding an electronic device to reselect a comparison target according to various embodiments.
- FIG. 2f is a diagram illustrating an operation of providing information about a target object according to various embodiments.
- FIG. 3a is a diagram illustrating an image recognition system according to various embodiments.
- FIG. 3b is a diagram schematically illustrating the configuration of an image recognition system according to various embodiments.
- Figure 3c is a drawing for explaining motion data used for object identification.
- FIG. 4 is a drawing for explaining an image recognition function of an electronic device according to various embodiments.
- FIG. 5A is a drawing for explaining an image recognition function of a first electronic device according to various embodiments.
- FIG. 5b is a drawing for explaining an operation of recognizing a user's gesture in a first electronic device according to various embodiments.
- FIG. 5c is a diagram illustrating the operation of a first electronic device controlled based on gestures according to various embodiments.
- Figure 6 is a diagram illustrating the configuration of an information provision model according to various embodiments.
- FIG. 7 is a diagram illustrating a procedure for processing input data of an object identification model according to various embodiments.
- FIG. 8A is a diagram illustrating an image recognition system according to various embodiments.
- FIG. 8b is a diagram illustrating an image recognition system according to various embodiments.
- FIG. 9A is a flowchart illustrating the operation of an electronic device according to various embodiments.
- FIG. 9b is a diagram for explaining related information according to various embodiments.
- FIG. 10 is a flowchart illustrating a motion data acquisition operation of an electronic device according to various embodiments.
- FIG. 11 is a flowchart illustrating a target object identification operation of an electronic device according to various embodiments.
- FIG. 12 is a flowchart illustrating a field of view correction operation of an electronic device according to various embodiments.
- Figure 13 is a drawing for explaining a field of view correction process according to various embodiments.
- FIG. 14 is a flowchart illustrating an operation of providing related information in an electronic device according to various embodiments.
- FIG. 15 is another flowchart illustrating the operation of an electronic device according to various embodiments.
- FIGS. 16A to 16C are drawings for explaining the operation of the first electronic device in various embodiments.
- FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.
- the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown).
- a notebook 190
- smartphones (191) having various form factors e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable) type smartphone (191-3)
- a tablet (192) e.g., a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown).
- the components, their relationships, and their functions illustrated in FIG. 1 are exemplary only and do not limit the implementations described or claimed in this document.
- the electronic device (100) may be referred to as a mobile device, a user device, a multi-function
- the electronic device (100) may include components including at least one processor (110) (hereinafter referred to as processor (110)), at least one memory (120) (hereinafter referred to as memory (120)), at least one display (140) (hereinafter referred to as display (140)), at least one image sensor (150) (hereinafter referred to as image sensor (150)), at least one communication circuit (160) (hereinafter referred to as communication circuit (160)), and/or at least one sensor (170) (hereinafter referred to as sensor (170)).
- processor (110) processor
- memory (120) hereinafter referred to as memory (120)
- display (140) at least one image sensor (150)
- image sensor (150) hereinafter referred to as image sensor (150)
- communication circuit (160) hereinafter referred to as communication circuit (160)
- sensor (170 sensor (170
- the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input/output interface).
- PMIC power management integrated circuitry
- audio processing circuitry e.g., audio processing circuitry, an antenna, a rechargeable battery, or an input/output interface.
- some components may be omitted from the electronic device (100).
- some components may be integrated into one component.
- the processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing.
- the processor (110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data) stored in the memory (120).
- the processor (110) may include a processor assembly including one or more processing circuits.
- the processor (110) may include any processing circuit operative to control the performance and operations of one or more components (e.g., the memory (120), the display (140), the image sensor (150), the communication circuit (160), and/or the sensor (170)) of the electronic device (100).
- the processor (110) e.g., the application processor (AP)
- SoC system on chip
- the processor (110) may be implemented with multiple cores (or at least one core circuit), multiple chips, or multiple chipsets.
- the processor (110) may include one or more processing circuits.
- the processor (110) may include one or more processing circuits configured to individually and/or collectively perform various functions of the present disclosure.
- At least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) that is different from the first chip of the electronic device (100).
- the processor (110) may include a central processing unit (CPU) (111), a graphics processing unit (GPU) (112), a neural processing unit (NPU) (113), an image signal processor (ISP) (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (CP) (118), and/or a sensor interface (119).
- CPU central processing unit
- GPU graphics processing unit
- NPU neural processing unit
- ISP image signal processor
- 114 image signal processor
- a display controller 115
- a memory controller 116
- storage controller 117
- a communication processor (CP) 118
- a sensor interface 119
- these components of the processor (110) are merely exemplary.
- the processor (110) may further include other components.
- some components of the processor (110) may be omitted from the processor (110).
- some components of the processor (110) may be included as separate components of the electronic device (100) outside the processor (110).
- some components of the processor (110) may be included within other components (e.g., at least a portion of the memory (120), an interface (e.g., available for connection to at least one component of the electronic device (100)), a display (140) and/or an image sensor (150)).
- the processor (110) can control the operations of the electronic device (100) by executing instructions stored in the memory (120).
- the processor (110) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.
- the processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in the memory (120).
- the CPU (111) (or central processing circuit) may be configured to control components of the processor (110) based on the execution of instructions stored in the memory (120) (e.g., volatile memory (121) and/or non-volatile memory (122)).
- the GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering).
- the NPU (113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation).
- the ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110).
- the display controller (115) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (111), the GPU (112), the ISP (114), or the memory (120) (e.g., the volatile memory (121)) into a format suitable for the display (140).
- the memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121).
- the storage controller (117) (or storage control circuit) may be configured to control reading data from the nonvolatile memory (122) and writing data to the nonvolatile memory (122).
- the CP (118) (communication processing circuit) may be configured to process data acquired from a component of the processor (110) into a format suitable for transmitting to another electronic device via the communication circuit (160), or to process data acquired from another electronic device via the communication circuit (160) into a format suitable for processing by the component of the processor (110).
- the communication circuit (160) may include one or more communication circuits.
- the sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data about the state of the electronic device (100) and/or the state of the surroundings of the electronic device (100), acquired via the sensor (170), into a format suitable for the component of the processor (110).
- the memory (120) may include one or more storage media (or one or more storage devices).
- the memory (120) may include a memory assembly including one or more storage media.
- the one or more storage media may include permanent memory (e.g., non-volatile memory (122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof.
- the memory (120) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As a non-limiting example, the cache memory may be included within the processor (110).
- the memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and/or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (100).
- SIM subscriber identity module
- SD secure digital
- the memory (120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and/or applet) software application, and/or any other suitable software applications.
- the one or more software applications may include instructions executable by the processor (110).
- the memory (120) may store instructions callable by an application programming interface (API).
- API application programming interface
- the memory (120) may store instructions within a library.
- the aforementioned electronic device (100) may provide an image recognition function that recognizes (or identifies) objects within image data and applies the recognition function to various services. This will be described in detail with reference to FIGS. 2A to 15 below. Furthermore, at least one of the various embodiments described with reference to FIGS. 2A to 15 below may be combined with other embodiments.
- FIG. 2A is a diagram for explaining an image recognition function of an electronic device according to various embodiments.
- FIG. 2B is a diagram for explaining information provided through an image recognition function according to various embodiments.
- FIG. 2C is a diagram for explaining an image recognition function of an electronic device according to various embodiments.
- FIG. 2D is a diagram for explaining a configuration for an electronic device according to various embodiments to determine whether a target object is appropriate as a comparison target.
- FIG. 2E is a diagram for explaining a configuration for guiding an electronic device according to various embodiments to reselect a comparison target.
- FIG. 2F is a diagram for explaining an operation for providing information on a target object according to various embodiments.
- 200 of FIG. 2a represents a situation in which a user speaks while sequentially pointing to a plurality of specific objects with an index finger (e.g., repeating a pointing gesture), and 230 of FIG. 2a represents image data (210) acquired by an electronic device (100) while the user speaks.
- index finger e.g., repeating a pointing gesture
- an electronic device (100) may identify a target object (or comparison target) indicated by a user among a plurality of objects included in the image data (210) based on speech data (or speech input) (201) and image data (210). According to one embodiment, the electronic device (100) may identify a plurality of target objects in the image data (210) based on speech data (201) and provide related information thereon.
- An image recognition function according to various embodiments related thereto will be described in more detail below.
- the electronic device (100) can obtain speech data (201) and identify a first part, a second part, and a third part thereof.
- the first part (and the second part) of the utterance data (201) may correspond to designated words (e.g., demonstrative pronouns indicating a single object such as this, that, here, there, he, she, you) that designate the first target object (and the second target object), and the third part of the utterance data (201) may correspond to a designated query.
- designated words e.g., demonstrative pronouns indicating a single object such as this, that, here, there, he, she, you
- the third part of the utterance data (201) may correspond to a designated query.
- the electronic device (100) may identify a first word (e.g., this child) (201-1) designating a first target object (or one object) as a first part, a second word (e.g., that child) (201-2) designating a second target object (e.g., another object) as a second part, and a sentence corresponding to the query (e.g., which child will be taller when they grow up?) (201-3) as a third part for the utterance data (201).
- a first word e.g., this child
- a second word e.g., that child
- a sentence corresponding to the query e.g., which child will be taller when they grow up
- 201-3 a third part for the utterance data
- the electronic device (100) may acquire image data (210) while acquiring speech data (201). According to one embodiment, the electronic device (100) may recognize (or extract) an object included in the image data (210) prior to identifying the target object.
- the electronic device (100) can recognize a first object (211), a second object (213), a third object (215), and a fourth object (217) based on feature data (e.g., feature points) extracted from image data (210).
- feature data e.g., feature points
- the electronic device (100) can extract feature data using various known techniques such as histogram of oriented gradient (HOG), scale invariant feature transform (SIFT), local binary pattern (LBP), and modified census transform (MCT).
- HOG histogram of oriented gradient
- SIFT scale invariant feature transform
- LBP local binary pattern
- MCT modified census transform
- the electronic device (100) can identify a target object indicated by a user among recognized objects (e.g., a first object (211) to a fourth object (217)) based on the first part and the second part identified from the speech data (201).
- recognized objects e.g., a first object (211) to a fourth object (217)
- the electronic device (100) may identify an object (e.g., a second object (213)) corresponding to a first location (221) of an index finger identified in image data (210) as a first target object at a first time point (or after the first time point) when a first portion (e.g., a first word (201-1)) is identified, as illustrated in 230 of FIG. 2A.
- an object e.g., a second object (213)
- a first location (221) of an index finger identified in image data (210) a first target object at a first time point (or after the first time point) when a first portion (e.g., a first word (201-1)) is identified, as illustrated in 230 of FIG. 2A.
- the electronic device (100) may identify a second object (213) existing within a certain range (or a certain distance) based on the first location (221) as the first target object.
- the electronic device (100) can identify an object (e.g., a third object (215)) corresponding to the second position (223) of the index finger identified in the image data (210) as the second target object at a second time point (or after the second time point) when the second part (e.g., the second word (201-2)) is identified. For example, when the index finger located at the first position (221) moves to the second position (223) within a certain range based on the third object (215) at the second time point, the electronic device (100) can identify the third object (215) as the second target object.
- an object e.g., a third object (215)
- the electronic device (100) can identify the third object (215) as the second target object.
- the electronic device (100) may provide associated information about identified target objects (e.g., a first target object (e.g., a second object (213)) and a second target object (e.g., a third object (215))) based on the third portion (201-3) of the speech data (201).
- identified target objects e.g., a first target object (e.g., a second object (213)) and a second target object (e.g., a third object (215)
- a first target object e.g., a second object (213)
- a second target object e.g., a third object (215)
- information (250) about an identified target object may include at least one of first information (251) related to a first target object (e.g., a poodle), second information (253) related to a second target object (e.g., a Welsh corgi), and third information (255) which is detailed information about them.
- the related information may include related information (e.g., comparison information) about the first target object and the second target objects.
- the electronic device (100) may use the first target object, the second target object, and the third part (201-3) of the data (201) as a search term (e.g., which child will grow up to be bigger, a poodle or a Welsh corgi?), and obtain a search result therefor from an external source (e.g., a search server).
- a search term e.g., which child will grow up to be bigger, a poodle or a Welsh corgi?
- an external source e.g., a search server
- the associated information may include information about each of the first target object and the second target object.
- the electronic device (100) may use queries inquiring about the types of the first target object, the second target object, and each target object as search terms (e.g., "Describe the types of poodles and Welsh corgis"), and may obtain search results (e.g., "Description of poodles” and "Description of Welsh corgis”) from an external source.
- search terms e.g., "Describe the types of poodles and Welsh corgis”
- search results e.g., "Description of poodles” and “Description of Welsh corgis
- the electronic device (100) may output the associated information (250) regarding the identified target object through the electronic device (100) or may output it through another electronic device.
- the associated information (250) regarding the identified target object may be output as visual information through the display (140) of the electronic device (100).
- the associated information (250) regarding the identified target object may be output as auditory information through the speaker of the electronic device (100), as illustrated in 295 of FIG. 2F, or may be output through another electronic device (e.g., wireless earphones (298), a smartwatch, another smart phone, or an AI speaker) that is connected to the electronic device (100) through communication, as illustrated in 297 of FIG. 2F.
- the electronic device (100) may also identify a target object from a plurality of image data.
- the electronic device (100) may obtain first image data (260) corresponding to a first field of view (e.g., 11 o'clock direction) at a first point in time when a first portion (e.g., a first word (201-1)) is identified, as illustrated in FIG. 2C.
- the electronic device (100) may obtain second image data (270) corresponding to a second field of view (e.g., 1 o'clock direction) different from the first field of view at a second point in time when a second portion (e.g., a second word (201-2)) is identified.
- the electronic device (100) may identify an object (e.g., a second object (213)) corresponding to a first position (261) of an index finger identified in the first image data (260) as a first target object at a first time point when the first portion is identified.
- the electronic device (100) may identify an object (e.g., a third object (215)) corresponding to a second position (271) of an index finger identified in the second image data (270) as a second target object at a second time point when the second portion is identified.
- the electronic device (100) may compare attributes (e.g., type) of the first target object with attributes of the second target object in providing association information (e.g., comparison information) for identified target objects (e.g., first target object and second target object).
- attributes e.g., type
- association information e.g., comparison information
- the electronic device (100) can determine whether a first target object and a second target object are appropriate as comparison objects through attribute comparison. For example, as illustrated in 208-1 of FIG. 2D, when a first target object (e.g., a first smart phone) (281) and a second target object (e.g., a second smart phone) (282) indicated by a user are identified, the electronic device (100) can determine whether the first target object (281) and the second target object (282) have attributes that are related to each other, and determine that the first target object (281) and the second target object (282) having attributes that are related to each other (e.g., a mobile device) are appropriate as comparison objects. In addition, the electronic device (100) can provide association information (250) for target objects (e.g., the first target object and the second target object) that are determined to be appropriate as comparison objects.
- association information 250
- the electronic device (100) may check an attribute (e.g., a mobile device) related to the second target object (282) and an attribute (e.g., a computer device) related to the third target object (283), and may determine that the second target object (282) and the third target object (283) having different attributes are not appropriate as comparison objects.
- an attribute e.g., a mobile device
- an attribute e.g., a computer device
- the electronic device (100) may provide guide information (291) for reselecting a comparison object or guide information for re-inputting a query, as illustrated in FIG. 2e.
- the electronic device (100) may also provide information related to the cause of the inappropriateness of the comparison target (e.g., the properties of the specified target objects are different from each other) as at least part of the guide information (291).
- the image recognition function may be provided through the independent operation of the electronic device (100).
- the electronic device (100) may also provide the image recognition function through collaboration with at least one other device. This will be described in detail with reference to FIGS. 3A to 3C below.
- FIG. 3a is a diagram illustrating an image recognition system according to various embodiments.
- FIG. 3b is a diagram schematically illustrating the configuration of an image recognition system according to various embodiments, and
- FIG. 3c is a diagram illustrating motion data utilized for object identification.
- an image recognition system (300) may be configured with a first electronic device (310) and a second electronic device (320).
- the first electronic device (310) and the second electronic device (320) may communicate with each other via a network (e.g., a short-range communication network or a long-range communication network).
- a network e.g., a short-range communication network or a long-range communication network.
- the first electronic device (310) may be referred to as a smartphone
- the second electronic device (320) may be referred to as a wearable device (e.g., a ring-shaped electronic device) worn on a part of the body (e.g., an index finger).
- the first electronic device (310) may be referred to as a foldable type smartphone (191-2) that can be attached to clothing worn by a user in a folded state (e.g., an outside pocket of clothing), or a pin (or clip) type wearable device that can be attached to clothing (e.g., a shirt) or accessories (e.g., a tie, a necklace) worn by the user.
- a foldable type smartphone (191-2) that can be attached to clothing worn by a user in a folded state (e.g., an outside pocket of clothing), or a pin (or clip) type wearable device that can be attached to clothing (e.g., a shirt) or accessories (e.g., a tie, a necklace) worn by the user.
- the first electronic device (310) and the second electronic device (320) may be referred to as a smartphone, and at least one of the first electronic device (310) and the second electronic device (320) may be referred to as a wearable device worn on another part of the body (e.g., a watch-type electronic device wearable on the wrist, a necklace-type electronic device wearable on the neck, eyeglass-type electronic device wearable on the face, or an augmented reality or virtual reality device wearable on the head).
- a wearable device worn on another part of the body e.g., a watch-type electronic device wearable on the wrist, a necklace-type electronic device wearable on the neck, eyeglass-type electronic device wearable on the face, or an augmented reality or virtual reality device wearable on the head.
- the first electronic device (310) can provide an image recognition function through collaboration with the second electronic device (320).
- the first electronic device (310) may acquire (or collect) speech data (201) and image data (210), and the second electronic device (320) may acquire motion data (341) related to the posture of the second electronic device (320) while it is worn on the body (e.g., motion data for a part of the body on which the second electronic device is worn (e.g., an index finger)).
- the first electronic device (310) may identify a target object indicated by a user among a plurality of objects (e.g., a first object (211) to a fourth object (217)) included in the image data (210) based on the speech data (201) and image data (210) acquired by the first electronic device (310) and the motion data (341) acquired by the second electronic device (320).
- the configuration of the first electronic device (310) according to various embodiments related thereto will be described in more detail below.
- a first electronic device (310) may be configured with at least one first processor (311) (hereinafter referred to as a first processor (311)), at least one first communication circuit (312) (hereinafter referred to as a first communication circuit (312)), at least one first memory (313) (hereinafter referred to as a first memory (313)), at least one image sensor (314) (or at least one camera) (hereinafter referred to as an image sensor (314)), at least one microphone (315) (hereinafter referred to as a microphone (315)), and at least one output device (316) (hereinafter referred to as an output device (316)).
- a first processor 311
- a first communication circuit (312) hereinafter referred to as a first communication circuit (312)
- at least one first memory (313) hereinafter referred to as a first memory (313)
- at least one image sensor (314) or at least one camera
- an output device (316) hereinafter referred to as an output device (316)
- the first electronic device (310) may be implemented to have more or fewer components than the aforementioned components.
- the first electronic device (310) may correspond to the electronic device (100) illustrated in FIG. 1.
- the first communication circuit (312) may support wireless communication with the second electronic device (320).
- the first communication circuit (312) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (310) and the second electronic device (320).
- the first communication circuit (312) may communicate with the second electronic device (320) via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).
- a first network e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)
- a second network e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)
- a first network e.g., a short-range communication network
- the first memory (313) may store commands or data related to components of the first electronic device (310).
- the first memory (313) may store programs, algorithms, routines, and commands related to image recognition.
- the image sensor (314) may acquire (or capture) image data (210).
- the image data (210) may be stored in the first memory (313).
- the first electronic device (310) has an output device (316) configured to output visual information, the image data (210) may also be output through the output device (316).
- the microphone (315) may be configured to convert external sounds into electrical audio signals and output them. According to one embodiment, the microphone (315) may be configured to acquire speech data (201). Depending on the embodiment, the microphone (315) may be applied with various noise removal algorithms to remove noise generated during the process of receiving external sounds.
- the output device (316) may provide various information related to the operation of the first electronic device (310). At least some of the various information may be related to the image recognition function.
- the output device (316) may include a display configured to provide visual information (e.g., text, images, videos, icons, or symbols) to the user and receive user input (e.g., touch input).
- visual information e.g., text, images, videos, icons, or symbols
- user input e.g., touch input
- the output device (316) may be configured to provide auditory information, tactile information, or a combination thereof.
- the first processor (311) may be operatively connected to a first communication circuit (312), a first memory (313), an image sensor (314), a microphone (315), and an output device (316), and may control various components (e.g., hardware or software components) of the first electronic device (310).
- the first processor (311) can identify a target object indicated by a user in image data (210) and provide related information (250) about the identified target object.
- the target object may be part or all of a plurality of objects included in the image data (210).
- the first processor (311) may acquire image data (210) while acquiring speech data (201). For example, the first processor (311) may identify a target object (e.g., the first target object and the second target object of FIG. 2A) among objects included in the image data (210) based on a first portion (e.g., a designated first word (201-1) of FIG. 2A) and a second portion (e.g., a designated second word (201-2) of FIG. 2A) identified from the speech data (201). In addition, the first processor (311) may provide related information (250) about the target object based on a third portion (e.g., a sentence (201-3) corresponding to a query of FIG. 2A) identified from the speech data (201). A specific description of the first processor (311) related to this may refer to the operation of the electronic device (100) described through FIGS. 2a and 2b.
- a target object e.g., the first target object and the second target object of FIG. 2A
- a first portion e.g
- the first processor (311) may utilize motion data (341) acquired through the second electronic device (320) to identify a target object.
- the first processor (311) may acquire motion data (341) from the second electronic device (320) while acquiring speech data (201) (and/or while acquiring image data (210)).
- the motion data (341) acquired from the second electronic device (320) may include position information, velocity information, acceleration information, direction information, or a combination thereof.
- the first processor (311) may obtain first motion data acquired through the second electronic device (320) at a first time point when a first portion of the speech data (201) is identified, and utilize the same to identify a target object (e.g., a first target object).
- the first processor (311) may obtain second motion data acquired through the second electronic device (320) at a second time point when a second portion of the speech data (201) is identified, and utilize the same to identify another target object (e.g., a second target object).
- the first processor (311) may acquire (or extract) a portion (341-1) corresponding to a first time point (t1) at which a first portion (e.g., a designated first word (e.g., this child) (201-1)) of the speech data (201) is identified, as first motion data, from the motion data (341) acquired from the second electronic device (320).
- the first processor (311) may acquire another portion (341-2) of the motion data (341) corresponding to a second time point (t2) at which a second portion (e.g., a designated second word (e.g., that child) (201-2)) of the speech data (201) is identified, as second motion data.
- the first processor (311) may determine whether the first motion data satisfies a specified condition when identifying the first target object. For example, if the first motion data acquired at the first time point (t1) satisfies the specified condition, the first processor (311) may identify an object corresponding to the first position of a body (e.g., an index finger) identified in the image data (210) as the first target object.
- a body e.g., an index finger
- the first processor (311) can determine whether the second motion data satisfies a specified condition when identifying the second target object. For example, if the second motion data acquired at the second point in time satisfies the specified condition, the first processor (311) can identify an object corresponding to the second position of the body identified in the image data (210) as the second target object.
- the first processor (311) can identify a target object in image data (210) when motion data (341) obtained through the second electronic device (320) satisfies a specified condition.
- the first processor (311) may determine that a specified condition is satisfied if the first motion data (and/or the second motion data) corresponds to a specified gesture.
- the specified gesture may include a gesture in which the user indicates a specific object using a part of the body (e.g., a pointing gesture).
- the specified gesture may also occur when the second electronic device (320) (e.g., a touch sensor) is at least partially touched by a part of the body.
- various gestures that can indicate a specific object (e.g., drawing a circle with the index finger, drawing with the hand while pointing to an object with the finger, forming a finger frame, or forming a circle using the thumb and index finger) can be used under specified conditions.
- a specific object e.g., drawing a circle with the index finger, drawing with the hand while pointing to an object with the finger, forming a finger frame, or forming a circle using the thumb and index finger
- the first processor (311) may acquire second motion data if a second portion of the speech data (201) is identified within a predetermined time (e.g., 5 seconds) after the first portion of the speech data (201) is identified. For example, the first processor (311) may exclude a second portion that is identified after a predetermined time has elapsed since the first portion was identified from the identification of the second target object.
- a predetermined time e.g. 5 seconds
- the first electronic device (310) can determine whether motion data (341) satisfying a specified condition is obtained from the second electronic device (320). However, in this case, a synchronization problem or a transmission speed problem for the motion data (341) may occur.
- the second processor (321) may determine whether the first motion data satisfies a specified condition when identifying the first target object.
- the first processor (311) may also obtain motion data (341) that satisfies the specified condition from the second electronic device (320).
- the second electronic device (320) may obtain motion data (341) related to a posture of the second electronic device (320) at the time of identifying the first part (and/or the second part) of the speech data (201) or receiving information representing the first part (and/or the second part) of the speech data (201) from the first electronic device (310).
- the second electronic device (320) may provide the motion data (341) to the first electronic device (310) when the obtained motion data (341) satisfies the specified condition.
- the second electronic device (320) may extract feature data for the motion data (341) and provide it to the first electronic device (310) instead of the motion data (341) that satisfies the specified conditions.
- the first electronic device (310) may identify the target object in the image data (210) based on the feature data provided from the second electronic device (320).
- the second electronic device (320) may extract feature data related to a motion feature of a body part (e.g., a finger) (or the second electronic device (320)).
- the feature data may include at least one of feature data corresponding to a moving body part, feature data corresponding to a stationary body part, feature data corresponding to a body part that has stopped moving, and feature data corresponding to a body part that has stopped moving.
- the first processor (311) can obtain motion data (341) used to identify a target object through the second electronic device (320).
- the configuration of an exemplary second electronic device (320) related thereto will be described in more detail below.
- a second electronic device (320) may be configured with at least one second processor (321) (hereinafter referred to as a second processor (321)), a second communication circuit (322) (hereinafter referred to as a second communication circuit (322)), at least one second memory (323) (hereinafter referred to as a second memory (323)), and at least one sensor (324) (hereinafter referred to as a sensor (224)), as illustrated in FIG. 3B.
- the second electronic device (320) may be implemented to have more or fewer components than the aforementioned components.
- the second communication circuit (322) and the second memory (323) described above may be similar to or identical to the first communication circuit (312) and the first memory (313) of the first electronic device (210) described above, and thus a detailed description thereof may be omitted.
- the second communication circuit (322) may support wireless communication with the first electronic device (310).
- the second communication circuit (322) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (310) and the second electronic device (320).
- the second memory (323) may store commands or data related to at least one other component of the second electronic device (320).
- the senor (324) may be configured to obtain motion data related to the posture of the second electronic device (320).
- the sensor (324) may include at least one of an acceleration sensor, a gyro sensor, a gesture sensor, or a barometric pressure sensor.
- the second processor (321) may be operatively connected to the second communication circuit (322), the second memory (323), and the sensor (324), and may control various components (e.g., hardware or software components) of the second electronic device (320).
- the second processor (321) may provide motion data collected through the sensor (324) to the first electronic device (310).
- the second processor (321) may control the sensor (324) so that motion data (341) is collected while the first electronic device (310) and the second electronic device (320) are connected to each other through communication.
- the second processor (321) may control the sensor (324) to collect motion data (341) based on the occurrence of a specified event. For example, if the second electronic device (320) is configured to receive a user input (e.g., a touch input) (e.g., if a touch sensor is provided in the second electronic device (320), the first processor (321) may process the motion data (341) to be collected after (or while) the user input is detected.
- a user input e.g., a touch input
- the first processor (321) may process the motion data (341) to be collected after (or while) the user input is detected.
- the first electronic device (310) can identify a target object in the image data (210) based on the speech data (201) and image data (210) collected by the first electronic device (310) and the motion data (314) collected by the second electronic device (320).
- the first electronic device (310) may perform an operation of scanning a beam of an optical pointer, such as a laser pointer, to identify a target object.
- the first electronic device (310) or the second electronic device (320) may include a light emitting unit configured to irradiate a laser point, and the first electronic device (310) may identify an object corresponding to the direction in which the laser point is pointed among a plurality of objects as a target object, thereby further improving object identification accuracy.
- FIG. 4 is a drawing for explaining an image recognition function of an electronic device according to various embodiments.
- the image recognition method illustrated in FIG. 4 differs from the image recognition methods illustrated in FIGS. 2a and 2b in that it utilizes a gesture that simultaneously designates multiple specific objects.
- the image recognition functions according to various related embodiments will be described in more detail below.
- a first electronic device (310) can obtain utterance data (411) and identify a fourth part and a third part thereof.
- the fourth part of the utterance data (411) may correspond to a designated word (e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and they) that designates multiple target objects (e.g., a first target object and a second target object) at once, and the third part of the utterance data (411) may correspond to a designated query.
- a designated word e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and they
- multiple target objects e.g., a first target object and a second target object
- the first electronic device (310) can identify a word (e.g., these) designating multiple target objects as a fourth part for the speech data (411) (e.g., which of these is the tallest when it grows up?), and can identify a sentence corresponding to the query (e.g., which is taller when it grows up?) as a third part.
- a word e.g., these
- the speech data (411) e.g., which of these is the tallest when it grows up?
- a sentence corresponding to the query e.g., which is taller when it grows up?)
- the first electronic device (310) can identify a target object in the image data (210) based on the fourth part of the speech data (411).
- the first electronic device (310) can recognize (or extract) objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) after acquiring the image data (210) while the speech data (411) is acquired.
- the first electronic device (310) can identify a plurality of target objects indicated by the user among the recognized objects (e.g., the first object (211) to the fourth object (217)) based on the fourth part of the speech data (411).
- the first electronic device (310) can identify a target object in the image data (210) when the motion data (341) obtained through the second electronic device (320) satisfies a specified condition.
- the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) corresponding to the gesture (413) identified in the image data (210) as target objects at a fourth time point when the fourth portion is identified, as illustrated.
- the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) included in a circle (413-1) drawn by the gesture (413) as target objects.
- the first electronic device (310) may provide association information (250) for a plurality of target objects based on the third part of the speech data (411).
- FIG. 5A is a diagram illustrating an image recognition function of a first electronic device according to various embodiments.
- FIG. 5B is a diagram illustrating an operation of recognizing a user's gesture in a first electronic device according to various embodiments.
- FIG. 5C is a diagram illustrating an operation of a first electronic device controlled based on a gesture according to various embodiments.
- the image recognition method illustrated in FIG. 5A differs from the image recognition method illustrated in FIG. 4 in that motion data is acquired through a plurality of second electronic devices (320-1 and 320-2).
- one (320-1) of the plurality of second electronic devices (320-1 and 320-2) may be worn on a first part of the body (e.g., the left hand (511-1)), and the other (320-2) of the plurality of second electronic devices (320-1 and 320-2) may be worn on a second part of the body (e.g., the right hand (511-2)).
- the image recognition function according to various embodiments related thereto will be described in more detail below.
- a first electronic device (310) can obtain utterance data (513) and identify a fourth part and a third part thereof.
- the fourth part of the utterance data (513) may correspond to a designated word (e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and she) that designates multiple target objects (e.g., a first target object and a second target object) at once, and the third part of the utterance data (513) may correspond to a designated query.
- a designated word e.g., a demonstrative pronoun indicating multiple objects such as these, those, we, they, you, and she
- multiple target objects e.g., a first target object and a second target object
- the first electronic device (310) can identify a word (these) designating multiple target objects as a fourth part for the user's speech data (513) (e.g., "Which of these is the tallest?") and identify a sentence corresponding to the query (e.g., "Who is taller?") as a third part.
- the first electronic device (310) can identify a target object in the image data (210) based on the fourth part of the speech data (411).
- the first electronic device (310) can recognize (or extract) objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) after acquiring the image data (210) while acquiring the speech data (513).
- the first electronic device (310) can identify a plurality of target objects indicated by the user among the recognized objects (e.g., the first object (211) to the fourth object (217)) based on the fourth part of the speech data (411).
- the first electronic device (310) can identify a target object in the image data (210) when the motion data (341) obtained through the plurality of second electronic devices (320-1 and 320-2) satisfies a specified condition.
- the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) corresponding to a gesture (e.g., a hand frame gesture) identified in the image data (210) as target objects at a fourth time point when the fourth part is identified, as illustrated.
- the first electronic device (310) can identify a plurality of objects (e.g., a second object (213) and a third object (215)) included in a hand frame formed by a gesture (511) using the first part (511-1) of the body and the second part (511-2) of the body as target objects.
- the first electronic device (310) may utilize predetermined signals transmitted and received with a plurality of second electronic devices (320-1 and 320-2) via wireless communication (e.g., ultra wideband (UWB) communication) for gesture identification.
- the first electronic device (310) may determine a first distance (B) with one second electronic device (320-1), a second distance (C) with another second electronic device (320-2), and a third distance (A) between the plurality of second electronic devices (320-1 and 320-2) based on the transmitted and received predetermined signals.
- the first electronic device (310) may identify a gesture in which the third distance (A) between the plurality of second electronic devices (320-1 and 320-2) falls within a specified range (e.g., within 15 cm).
- the first electronic device (310) may transmit and receive a predetermined signal with a plurality of second electronic devices (320-1 and 320-2) through an algorithm related to ranging.
- the algorithm related to ranging may include at least one of ToF (time of flight), TWR (two way ranging), DS-TWR (double sided-TWR), SS-TWR (single sided-TWR), TDoA (time difference of arrival), or AoA (angle of arrival).
- the first electronic device (310) may provide association information (250) for a plurality of target objects based on the third portion (201-3) of the speech data (513).
- the first electronic device (310) can identify a target object from the image data (210) based on speech data (201, 411, 513) collected by the first electronic device (310), image data (210), and motion data (341) collected by the second electronic device (320).
- the first electronic device (310) may utilize an artificial intelligence model generated through machine learning to identify a target object.
- the first electronic device (310) e.g., the first memory (313)
- an artificial intelligence model e.g., an information provision model (3131)
- related information e.g., comparison information
- the first electronic device (310) may perform a designated action based on a gesture (e.g., a hand frame gesture) identified in the image data (210).
- the first electronic device (310) may activate an image sensor that is in an inactive state in response to the gesture identification.
- the first electronic device (310) may obtain an image corresponding to a field of view of the first electronic device (310) (e.g., an image sensor), as illustrated in 540 of FIG. 5C .
- the first electronic device (310) may obtain an image that includes only target objects (e.g., a second object (213) and a third object (215)) identified by the gesture (e.g., a hand frame gesture), as illustrated in 550 of FIG.
- target objects e.g., a second object (213) and a third object (215)
- the gesture e.g., a hand frame gesture
- Figure 6 is a diagram illustrating the configuration of an information provision model according to various embodiments.
- the information providing model (3131) may be an artificial intelligence model generated through machine learning.
- the information providing model (3131) may input motion data (610) (e.g., motion data (341)), image data (620) (e.g., image data (210)), and speech data (630) (e.g., speech data (201, 411, 513)), and output related information (e.g., comparison information) regarding a target object identified in the image data (620).
- the information providing model (3131) may include a plurality of artificial neural network layers.
- the artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, a transformer network, or a combination of two or more thereof, but is not limited to the examples described above.
- the information providing model (3131) may additionally or alternatively include a hardware structure. The configuration of the information providing model (3131) according to various embodiments related thereto will be described in more detail below.
- an information provision model (3131) may include an object identification model (605), an information retrieval model (607), and an information generation model (609).
- the object identification model (605) can identify a target object from image data (620).
- the object identification model (605) can receive motion data (610), image data (620), and speech data (630) as inputs, and output identification information (615) about the target object identified in the image data (620).
- the object identification model (605) can identify the target object based on a gesture identified in the image data (620) while speech data (630) is input.
- the information retrieval model (607) can retrieve information related to a target object identified in image data (620).
- the information retrieval model (607) can receive speech data (630) and information about the target object (e.g., identification result (615)) as inputs, and output a search result (617).
- the information retrieval model (607) can generate a search term based on the speech data (630) and the identification result (615), and obtain a search result (617) based on the search term from an external source (e.g., a search server).
- an external source e.g., a search server
- the information generation model (609) can convert the search result (617) of the information retrieval model (607) into a form that can be recognized by the user and output information (619) (e.g., related information (250)) about the target object.
- the information generation model (609) can convert the text-type search result (617) output by the information retrieval model (607) into an auditory form.
- the information (619) about the target object may also be output in a visual form.
- the information generation model (609) may include a generative model configured to generate new output data (e.g., image data) based on the search results (617).
- a generative model configured to generate new output data (e.g., image data) based on the search results (617).
- new output data e.g., image data
- the information generation model (609) may be comprised of various types of models other than the generative model.
- the information provision model (3131) may be configured with fewer components than the aforementioned components.
- the information generation model (609) may be omitted from the information provision model (3131).
- the search results (617) obtained by the information retrieval model (607) may be provided as information (619) regarding the target object.
- the information provision model (3131) may be configured to have more components than the aforementioned components.
- a pattern recognition model (603) may be included in the information provision model (3131).
- the object identification model (605) may identify the user's gesture based on pattern information (613) of motion data (610) output by the pattern recognition model (603).
- the configuration of the pattern recognition model (603) according to various embodiments related thereto will be described in more detail below.
- the pattern recognition model (603) can identify pattern information (613) based on motion data (610).
- the pattern information can be related to the repetition of motion, the direction of motion, the speed of motion, or the magnitude of motion.
- the pattern recognition model (603) can receive motion data (610) as input and identify a specific pattern in the motion data (610).
- the pattern recognition model (603) can output pattern information (613) related to the identified specific pattern. This pattern information (613) is input to the object identification model (605), and the object identification model (605) can utilize the pattern information (613) to identify a user's gesture.
- the pattern recognition model (603) may identify a specific pattern in the motion data (610) based on feature data (611) extracted from the motion data (610).
- the information provision model (3131) may further include a feature extraction model (601). The configuration of the feature extraction model (601) according to various related embodiments will be described in more detail below.
- the feature extraction model (601) can extract feature data (611) from motion data (610).
- the feature data (611) may be data that can be utilized to identify a specific pattern in the motion data (610).
- the feature extraction model (601) can extract feature data (611) related to appearance features and motion features of a body part (e.g., a finger) from the motion data (610). This feature data (611) is input to a pattern recognition model (603), and the pattern recognition model (603) can utilize the feature data (611) to identify the pattern.
- FIG. 7 is a diagram illustrating a procedure for processing input data of an object identification model according to various embodiments.
- the object identification model (605) receives image data (620) and speech data (630) as inputs and can identify a target object from the image data (620).
- the information provision model (3131) can input the image data (620) and speech data (630) into the object identification model (605).
- the information provision model (3131) can perform an operation of converting one-dimensional data into two-dimensional data when merging (650) image data (620) and speech data (630).
- the information provision model (3131) can obtain refined image data by preprocessing (621) the image data (620) in a manner such as filtering and sampling, and convert it into two-dimensional image data expressed in terms of the relationship between time and frequency (623).
- the information provision model (3131) can obtain refined speech input by preprocessing (631) the speech input (630) in a manner such as filtering and sampling, and convert it into two-dimensional speech data expressed in terms of the relationship between time and frequency (633).
- These two-dimensional image data (623) and speech data (633) are merged (650) and provided as input to an object identification model (605), and the object identification model (605) can output identification information (615) for the target object based on the merged image data (623) and speech data (633).
- motion data (610) (or feature data (611) or pattern information (613)) may be further utilized in addition to image data (620) and speech data (630).
- the information provision model (3131) may convert motion data (610) into two-dimensional motion data, merge it with two-dimensional image data (623) and speech data (633), and output it as an object identification model (605).
- FIG. 8A is a diagram illustrating an image recognition system according to various embodiments.
- an image recognition system (81) may be composed of a first electronic device (810), a second electronic device (820), and a third electronic device (830).
- the first electronic device (810) may communicate with the second electronic device (820) and the third electronic device (830) through a network (e.g., a short-range communication network or a long-range communication network).
- a network e.g., a short-range communication network or a long-range communication network.
- the first electronic device (810) may provide an image recognition function through collaboration with the second electronic device (820) and the third electronic device (830).
- a first electronic device (810) may obtain (or collect) image data (210), a second electronic device (820) may obtain motion data (341) related to a posture of the second electronic device (820) while it is worn on a body, and a third electronic device (830) may obtain speech data (201).
- the first electronic device (810) may identify a target object indicated by a user among a plurality of objects (e.g., a first object (211) to a fourth object (217)) included in the image data (210) based on the image data (210) obtained by the first electronic device (810), the motion data (341) obtained by the second electronic device (820), and the speech data (201) obtained by the third electronic device (830).
- the configuration of the first electronic device (810) according to various embodiments related thereto will be described in more detail below.
- a first electronic device (810) may be composed of a first processor (811), a first communication circuit (812), a first memory (813), an image sensor (814), and an output device (816).
- the configurations of the first electronic device (810) illustrated in FIG. 8a may be similar or identical to the configurations of the first electronic device (320) described above through FIG. 3b, and thus a detailed description thereof may be omitted.
- the first processor (811) can identify a target object indicated by a user from image data (210) acquired through an image sensor (814). For example, the first processor (811) can utilize motion data (341) acquired by a second electronic device (820) and speech data (201) acquired by a third electronic device (830) to identify the target object.
- a second electronic device (820) may be composed of a second processor (821), a second communication circuit (822), a second memory (823), and a sensor (824).
- the configurations of the second electronic device (820) illustrated in FIG. 8a may be similar or identical to the configurations of the second electronic device (320) described above through FIG. 3b, and thus a detailed description thereof may be omitted.
- the second processor (821) may provide motion data (341) collected through a sensor (824) to the first electronic device (810).
- a third electronic device (830) may be composed of a third processor (831), a third communication circuit (832), a third memory (833), and a microphone (834).
- the configurations of the third electronic device (830) illustrated in FIG. 8a may be similar or identical to the configurations of the first electronic device (810) or the second electronic device (820) described above through FIG. 3b, and thus a detailed description thereof may be omitted.
- the third processor (831) may provide speech data (201) collected through a microphone (834) to the first electronic device (810).
- FIG. 8b is a diagram illustrating an image recognition system according to various embodiments.
- an image recognition system (82) may be composed of a first electronic device (810), a second electronic device (820), a third electronic device (830), and a fourth electronic device (840).
- the first electronic device (810) may communicate with the second electronic device (820), the third electronic device (830), and the fourth electronic device (840) through a network (e.g., a short-range communication network or a long-range communication network).
- a network e.g., a short-range communication network or a long-range communication network.
- the first electronic device (810) may provide an image recognition function through collaboration with the second electronic device (820), the third electronic device (830), and the fourth electronic device (840).
- the first electronic device (810) can identify a target object indicated by a user among a plurality of objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) based on motion data (341) acquired by the second electronic device (820), speech data (201) acquired by the third electronic device (830), and image data (210) acquired by the fourth electronic device (840).
- a target object indicated by a user among a plurality of objects (e.g., the first object (211) to the fourth object (217)) included in the image data (210) based on motion data (341) acquired by the second electronic device (820), speech data (201) acquired by the third electronic device (830), and image data (210) acquired by the fourth electronic device (840).
- the configuration of the first electronic device (810) according to various embodiments related thereto will be described in more detail below.
- a first electronic device (810) may be composed of a first processor (811), a first communication circuit (812), a first memory (813), and an output device (816).
- the first processor (811) can utilize motion data (341) acquired by the second electronic device (820), speech data (201) acquired by the third electronic device (830), and image data (210) acquired by the fourth electronic device (840) to identify a target object.
- a second electronic device (820) may be composed of a second processor (821), a second communication circuit (822), a second memory (823), and a sensor (824).
- the second processor (821) may provide motion data (341) collected through a sensor (824) to the first electronic device (810).
- a third electronic device (830) may be composed of a third processor (831), a third communication circuit (832), a third memory (833), and a microphone (834).
- At least one third processor (831) may provide speech data (201) collected via a microphone (834) to the first electronic device (810).
- a fourth electronic device (840) may be composed of a fourth processor (841), a fourth communication circuit (842), a fourth memory (843), and an image sensor (844).
- the fourth processor (841) may provide image data (210) collected through the image sensor (844) to the first electronic device (810).
- an image recognition system may be configured with a first electronic device (810) configured to acquire image data (210) and a second electronic device (820) configured to collect motion data (341) and speech data (201).
- an image recognition system may be configured with a first electronic device (810) configured to acquire speech data (201) and a second electronic device (820) configured to collect motion data (341) and image data (210).
- An electronic device (310) may include at least one processor (311), a camera (e.g., an image sensor) (314), a microphone (315), and a memory (313) operatively connected to the at least one processor, the camera, and the microphone and storing at least one command.
- processors 311
- a camera e.g., an image sensor
- microphone 315
- memory 313
- the at least one command when individually or collectively executed by the at least one processor (311), may cause the electronic device (310) to: acquire image data (210) through the camera (314) while utterance data (201) including a designated word related to the designation of an object is acquired through the microphone (315), identify at least two objects corresponding to a gesture identified in the image data at the time when the designated word is uttered among a plurality of objects (211 to 217) included in the image data, and provide associated information (250) for the at least two objects.
- the utterance data may further include a query (201-3).
- the at least one command when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: provide the associated information related to the query.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: generate a search term based on the at least two objects and at least a portion of the query, and provide the associated information related to the generated search term.
- the electronic device (310) may further include a communication circuit (312).
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: obtain the associated information from an external device (320) via the communication circuit (312).
- the designated word may include a first word designating a plurality of objects.
- the at least one command when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify, among a plurality of objects included in the image data, a first object and a second object corresponding to a first gesture (413, 511) identified in the image data at the time when the first word is uttered, and provide the associated information for the first object and the second object.
- the designated word may include a second word (201-1) and a third word (201-2) designating a single object.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify, among a plurality of objects included in the image data, a first object corresponding to a second gesture (221) identified in the image data at a time when the second word is uttered; identify, among a plurality of objects included in the image data, a second object corresponding to a third gesture (223) identified in the image data at a time when the third word is uttered; and provide the association information for the first object and the second object.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: identify a second object in the image data if the third word is uttered within a predetermined time after the second word is uttered.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: acquire first image data (260) and second image data (270) corresponding to different fields of view through the camera (314), identify a first object corresponding to a second gesture (221) identified in the first image data at a time when the second word is uttered among a plurality of objects included in the first image data, and identify a second object corresponding to a third gesture (223) identified in the second image data at a time when the third word is uttered among a plurality of objects included in the second image data.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: acquire motion data (341) from an external device (320) via the communication circuit (312) while the speech data is acquired, and identify the at least two objects if the motion data satisfies a specified condition.
- the at least one instruction when individually or collectively executed by the at least one processor (311), may be configured to cause the electronic device (310) to: input the speech data and the image data into an artificial intelligence model (3131) stored in the electronic device (310) (e.g., memory (313)), and identify the at least two objects based on an output of the artificial intelligence model.
- an artificial intelligence model 3131
- the electronic device (310) e.g., memory (313)
- Figure 9a is a flowchart illustrating the operation of an electronic device according to various embodiments.
- Figure 9b is a diagram for explaining related information according to various embodiments.
- the operations in the following embodiments may be performed sequentially, but are not necessarily performed sequentially.
- the order of the operations may be changed, and at least two operations may be performed in parallel.
- at least one of the aforementioned operations may be omitted depending on the embodiment.
- operations 910 to 950 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
- a processor e.g., processor (110) of FIG. 1
- an electronic device e.g., electronic device (100) of FIG. 1).
- an electronic device (100) may obtain image data (210) in operation 910.
- the image data (210) may be obtained through an image sensor (e.g., an image sensor (314)) of the electronic device (100) (e.g., the first electronic device (310)) or may be obtained from an external source (e.g., a second electronic device (320)).
- an image sensor e.g., an image sensor (314)
- an external source e.g., a second electronic device (320)
- An electronic device (100) (e.g., a first electronic device (310)) may obtain motion data (341) in operation 920.
- the electronic device (100) may obtain motion data (341) while obtaining image data (210).
- the motion data (341) may be obtained from an external source (e.g., a second electronic device (320)).
- an electronic device (100) may obtain a speech input (or speech data) (201) in operation 930.
- the speech input (201) may be obtained through a microphone (e.g., a microphone (315)) of the electronic device (100) or may be obtained externally.
- the speech may be a speech requesting related information (e.g., comparison information) about an object.
- the speech may be a speech requesting information about each object, and according to an embodiment, the speech may be a speech indicating various requests, such as taking pictures of the object.
- An electronic device (100) may, at operation 940, identify at least two objects (e.g., two target objects) included in image data (210) based on speech input (201) and motion data (341).
- identify at least two objects e.g., two target objects included in image data (210) based on speech input (201) and motion data (341).
- the electronic device (100) can identify a target object based on the location of a body part (e.g., an index finger) identified in the image data (210) when motion data satisfying a specified condition is acquired while a speech input is acquired.
- a body part e.g., an index finger
- An electronic device (100) may provide association information on a target object corresponding to a speech input at operation 950.
- the electronic device (100) may generate a search term based on the speech input (201) and the identified target object, and obtain search results based on the search term from an external source (e.g., a search server).
- the association information may include at least one of the visual association information described above with reference to FIG. 2b or the auditory association information described above with reference to FIG. 2f.
- the electronic device (100) when the electronic device (100) obtains a gesture indicating a plurality of objects (971, 972) and a user utterance requesting a comparison of the plurality of objects (971, 972) (e.g., "Tell me where to buy the cheaper object among the two objects") (970), as illustrated in FIG. 9b, the electronic device (100) may provide the price comparison result and information related to the purchase location (e.g., the seller's homepage address) to the plurality of objects (971, 972) (980).
- the gestures indicating the plurality of objects (971, 972) may be the same or different from each other.
- one object (971) may be designated by a first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing at an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and index finger), and another object (972) may also be designated by the first gesture.
- a first gesture e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing at an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and index finger
- another object (972) may also be designated by the first gesture.
- one object (971) may be designated by a first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing to an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and an index finger), and another object (972) may be designated by a second gesture different from the first gesture (e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing to an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and an index finger).
- a first gesture e.g., one of a gesture of drawing a circle with a finger, a gesture of drawing with a hand while pointing to an object with a finger, a gesture of making a finger frame, or a gesture of making a circle using a thumb and an index finger
- FIG. 10 is a flowchart illustrating motion data acquisition operations of an electronic device according to various embodiments. The operations of FIG. 10 described below may represent various embodiments of operation 910 of FIG. 9.
- an electronic device (100) (e.g., a first electronic device (310)) may obtain motion data (341) in operation 1010.
- the electronic device (100) may obtain motion data (341) in a state in which the operation of an image sensor (e.g., an image sensor (314)) that consumes relatively much power is deactivated.
- an image sensor e.g., an image sensor (314)
- an electronic device (100) may determine, in operation 1020, whether motion data (341) corresponding to a specified gesture is acquired.
- the electronic device (100) may determine whether motion data (341) collected by another electronic device (e.g., a second electronic device (320)) that is connected to the electronic device (100) through communication corresponds to the specified gesture.
- the specified gesture may be a gesture of drawing a specified shape using a body on which another electronic device is worn (e.g., a gesture of repeatedly drawing a circle a certain number of times).
- the specified gesture may also be a gesture of generating a specified input (e.g., a touch input) to another electronic device worn on the body using another part of the body.
- An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may repeatedly perform operations 1010 and 1020 if motion data (341) corresponding to a specified gesture is not obtained.
- an electronic device (100) may activate an image sensor in operation 1030 when motion data (341) corresponding to a specified gesture is obtained.
- the electronic device (100) may reduce power consumption of the electronic device (100) by controlling the operation of the image sensor (314) using a sensor (e.g., a motion sensor) that consumes relatively little power.
- An electronic device (100) may obtain image data (210) in operation 1040.
- the image data (210) may be obtained through an image sensor of the electronic device (100) or may be obtained from an external source (e.g., a second electronic device (320)).
- FIG. 11 is a flowchart illustrating a target object identification operation of an electronic device according to various embodiments. The operations of FIG. 10 described below may represent various embodiments of operation 940 of FIG. 9.
- an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments can, in operation 1110, check first motion data corresponding to a first portion of a speech input (201).
- the first part of the speech input (201) may correspond to a designated word designating a first target object (e.g., a demonstrative pronoun indicating a single object such as this, that, here, there, he, she, you), as described above through FIG. 2a.
- a designated word designating a first target object e.g., a demonstrative pronoun indicating a single object such as this, that, here, there, he, she, you
- the electronic device (100) can acquire, as described above through FIG. 3c, some motion data corresponding to the first time point (t1) at which the first part of the speech data (201) is identified from the acquired motion data (341), as first motion data.
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1120, identify second motion data corresponding to a second portion of a speech input.
- the second part of the speech input (201) may correspond to a designated word designating a second target object, as described above with reference to FIG. 2a.
- the electronic device (100) can acquire, as described above through FIG. 3c, some motion data corresponding to the second time point (t2) at which the second part of the speech data (201) is identified from the acquired motion data (341), as second motion data.
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1130, identify a first object (e.g., a first target object) corresponding to first motion data among a plurality of objects included in image data (210).
- a first object e.g., a first target object
- the electronic device (100) may determine whether the first motion data satisfies a specified condition when identifying the first object. For example, if the first motion data acquired at the first time point (t1) satisfies the specified condition, the electronic device (100) may identify an object corresponding to the position of a body (e.g., an index finger) identified in the image data (210) as the first target object.
- a body e.g., an index finger
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1140, identify a second object (e.g., a second target object) corresponding to second motion data among a plurality of objects included in image data (210).
- a second object e.g., a second target object
- the electronic device (100) may determine whether the second motion data satisfies a specified condition when identifying the second object. For example, if the second motion data acquired at the second time point (t2) satisfies the specified condition, the electronic device (100) may identify an object corresponding to the position of a body (e.g., an index finger) identified in the image data (210) as the second target object.
- a body e.g., an index finger
- the electronic device (100) can identify a target object by recognizing a user's gesture that sequentially points to a plurality of specific objects in image data (210). However, if the image data (210) corresponding to the user's field of view is not acquired, the target object identified by the electronic device (100) and the object indicated by the user may not match each other. In this regard, the electronic device (100) according to various embodiments can acquire the image data (210) corresponding to the user's field of view and improve the identification performance for the target object by matching the acquisition range of the image data (210) to the user's field of view. This will be described in more detail with reference to FIGS. 12 and 13 below.
- FIG. 12 is a flowchart illustrating a field of view correction operation of an electronic device according to various embodiments.
- FIG. 13 is a diagram for explaining a field of view correction process according to various embodiments.
- each operation in the following embodiments may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. In addition, at least one of the above-described operations may be omitted depending on the embodiment.
- operations 1210 to 1250 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
- a processor e.g., processor (110) of FIG. 1
- an electronic device e.g., electronic device (100) of FIG. 1).
- an electronic device (100) may obtain image data (1301) in operation 1210.
- the image data (1301) may be obtained through an image sensor (e.g., an image sensor (314)) of the electronic device (100) or may be obtained from an external source (e.g., a second electronic device (320)).
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1220, select a reference object from image data (1301).
- the reference object may include at least one of the objects included in the image data (1301).
- the electronic device (100) may recognize a first object (e.g., a sofa) (1303), a second object (e.g., a table) (1305), a third object (e.g., a vase) (1307), and a fourth object (e.g., a picture frame) (1309) based on feature data (e.g., feature points) extracted from image data (1301), and select at least one of the recognized objects (e.g., the fourth object (1309)) as a reference object.
- a first object e.g., a sofa
- a second object e.g., a table
- a third object e.g., a vase
- a fourth object e.g., a picture frame
- an electronic device (100) may output guide information that guides selection of a reference object in operation 1230.
- the electronic device (100) may output guide information in an auditory form (e.g., “Point to the picture frame,” “Point to the right corner of the picture frame”) that guides a user to point to the reference object.
- an auditory form e.g., “Point to the picture frame,” “Point to the right corner of the picture frame”
- the electronic device (100) may provide guide information in a visual form, and according to an embodiment, may provide guide information in an auditory form and guide information in a visual form together.
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1240, identify a first location (1321) corresponding to a gesture related to selection of a reference object in image data (1301).
- An electronic device (100) may, in operation 1250, correct a field of view of an image sensor (314) based on a first location (1321) and a second location corresponding to a reference object.
- the electronic device (100) may correct a field of view of the image sensor based on a distance and direction difference (1323) between the first location and the second location.
- the electronic device (100) may select an object (e.g., a third object (e.g., a vase) (1307)) having a size smaller than a specified size among a plurality of objects (e.g., a first object (e.g., a sofa) (1303), a second object (e.g., a table) (1305), a third object (e.g., a vase) (1307), and a fourth object (e.g., a picture frame) (1309)) recognized from image data (1301) as a reference object in order to improve the field of view correction performance of the image sensor.
- a third object e.g., a vase
- a fourth object e.g., a picture frame
- FIG. 14 is a flowchart illustrating operations for providing related information in an electronic device according to various embodiments. The operations of FIG. 14 described below may represent various embodiments of operation 950 of FIG. 9.
- an electronic device (100) (e.g., a first electronic device (310)) according to various embodiments can check a third part of a speech input (201) in operation 1410.
- the third part of the speech input (201) may correspond to a specified query, as described above with reference to FIG. 2a.
- An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may perform a search operation in operations 1420 and 1430.
- the electronic device may perform a search operation using the third part of the speech data (201), the first target object, and the second target object as search words (operation 1420).
- the electronic device (100) may generate a search word based on at least a portion of the third part of the speech data (201), the first target object, and the second target object.
- the electronic device (100) may obtain a search result based on the generated search word from an external source (e.g., a search server) (operation 1430).
- an external source e.g., a search server
- Figure 15 is another flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
- operations 1510 to 1580 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
- a processor e.g., processor (110) of FIG. 1
- an electronic device e.g., electronic device (100) of FIG. 1).
- an electronic device (100) (e.g., a first electronic device (310)) may, in operation 1510, acquire image data (210) through an image sensor (e.g., an image sensor (314)) of the electronic device (100) or may acquire it from the outside (e.g., a second electronic device (320)).
- an image sensor e.g., an image sensor (314)
- a second electronic device e.g., a second electronic device (320)
- An electronic device (100) (e.g., a first electronic device (310)) may detect a designated first gesture in operation 1520.
- the designated first gesture may be a gesture in which a user points to a specific object with an index finger.
- An electronic device (100) may determine, in operation 1530, whether a first utterance for object designation is detected.
- the first utterance may correspond to a designated word designating an object (e.g., a demonstrative pronoun indicating a single object such as this, that, here, there, he, she, you).
- An electronic device (100) (e.g., a first electronic device (310)) according to various embodiments may repeatedly perform operations related to operations 1510 to 1530 if the first ignition is not detected.
- an electronic device (100) may, when a first utterance is detected, identify a first object in image data (210) in operation 1540.
- the electronic device (100) may identify a first object corresponding to a first gesture among a plurality of objects included in the image data (210).
- An electronic device (100) may detect a designated second gesture at operation 1550.
- the designated second gesture may be similar to the first gesture described above.
- the electronic device (100) may detect a second gesture indicating another object.
- An electronic device (100) may determine, at operation 1560, whether a second utterance for object designation is detected.
- the second utterance may be similar to the first utterance described above.
- the electronic device (100) may detect a second utterance indicating another object after detecting the first utterance.
- An electronic device (100) may repeatedly perform operations related to operations 1510 to 1560 if a second ignition is not detected.
- an electronic device (100) may, when a second utterance is detected, identify a second object in the image data (210) at operation 1570.
- the electronic device (100) may identify a second object corresponding to the second gesture among a plurality of objects included in the image data (210).
- An electronic device (100) may detect an information request utterance at operation 1580.
- the information request utterance may be an utterance requesting related information (e.g., comparison information) about a first object and a second object.
- the information request utterance may be an utterance requesting information about each of the first object and the second object.
- An electronic device (100) (e.g., a first electronic device (310)) may, in operation 1590, provide related information about a first object and a second object based on an information request utterance.
- the electronic device (100) can obtain related information about the first object and the second object from the outside.
- the electronic device (100) when an information request utterance requesting information about each of a first object and a second object is detected, the electronic device (100) can obtain information about the first object and information about the second object from the outside.
- FIGS. 16A to 16C are drawings for explaining the operation of the first electronic device in various embodiments.
- a first electronic device (310) may acquire a first gesture (1603) and a first utterance (1605) while acquiring image data (1601).
- the first gesture (1603) may be a gesture identified within the image data (1601) that indicates at least one specific object included in the image data (1601) (e.g., a gesture of making a finger frame).
- the first utterance (1605) may be an utterance that indicates processing (e.g., capturing or storing) at least a portion within the image data (1601) (e.g., capturing here).
- a first electronic device (310) may acquire a second gesture (1623) and a second utterance (1625) after acquiring a first gesture (1603) and a first utterance (1605).
- the second gesture (1623) may be another gesture identified in the image data (1601) that indicates at least one other specific object included in the image data (1601) (e.g., a gesture pointing to a target with a finger).
- the second utterance may be an utterance that indicates another processing (e.g., a search) for at least a portion of the image data (e.g., “What is this?
- the first electronic device (310) can acquire a second gesture (1623) and a second utterance (1625) within a certain time after the first gesture (1603) and the first utterance (1605) are acquired.
- a first electronic device (310) may perform first processing on image data (1601) based on a first gesture (1603) and a first utterance (1605).
- the first electronic device (310) may designate the entire image data (1601) as a first processing target based on the first gesture (1603), and store (1633) the image data (1601) designated as the first processing target based on the first utterance (1605).
- the first electronic device (310) may perform second processing on the image data (1601) based on the second gesture (1623) and the second utterance (1625).
- the first electronic device (310) may designate a portion of the image data (1601) as a second processing target based on the second gesture (1623), perform a search operation on the second processing target designated as the processing target based on the second utterance (1625), and use at least a portion of the search result to create a storage album (or folder) for the image data (1601).
- the first electronic device (310) may set at least a portion of the result as a storage album name (1631) based on the second utterance (1625).
- the first electronic device (310) may provide a processing result for at least one of the first processing and the second processing while performing the first processing (e.g., storing the image data (1601)) and the second processing (e.g., setting a name for a storage album after a search) on the image data (1601).
- the first electronic device (310) may provide visual information (1635) indicating a search result (e.g., search content) for the second processing target (e.g., Gwanghwamun. Gwanghwamun is the main gate of Gyeongbokgung Palace, and so on).
- a search result e.g., search content
- Gwanghwamun is the main gate of Gyeongbokgung Palace, and so on.
- the processing result for at least one of the first processing and the second processing may be output as auditory information through an external electronic device (e.g., wireless earphones (1641)) connected to the first electronic device (310) through communication, as illustrated in 1640 of FIG. 16c, or may be output as auditory information through a speaker of the first electronic device (310), as illustrated in 1650 of FIG. 16c.
- an external electronic device e.g., wireless earphones (1641)
- the processing result for at least one of the first processing and the second processing may be output as auditory information through an external electronic device (e.g., wireless earphones (1641)) connected to the first electronic device (310) through communication, as illustrated in 1640 of FIG. 16c, or may be output as auditory information through a speaker of the first electronic device (310), as illustrated in 1650 of FIG. 16c.
- the first electronic device (310) may designate the entire image data (1601) as a processing target based on the first gesture (1603).
- the first electronic device (310) may designate only a specific object included in the hand frame formed by the first gesture (1603) as a processing target.
- the first electronic device (310) may store a portion (1634) of the image data (1601) corresponding to the first gesture (1603) stored based on the first utterance (1605) in a storage album set based on the second utterance (1625).
- An operating method of an electronic device (310) may include an operation of acquiring speech data (201) including a designated word related to designation of an object, an operation of acquiring image data (210) while the speech data is being acquired, an operation of identifying at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects (211 to 217) included in the image data, and an operation of providing associated information (250) for the at least two objects.
- the utterance data may further include a query (201-3).
- the operating method of the electronic device (310) may include an operation of providing the related information related to the query.
- the method of operating the electronic device (310) may include generating a search term based on the at least two objects and at least a portion of the query, and providing the associated information related to the generated search term.
- the method of operating the electronic device (310) may include an operation of obtaining the related information from an external device (320).
- the designated word may include a first word designating a plurality of objects.
- the operating method of the electronic device (310) may include an operation of identifying a first object and a second object corresponding to a first gesture (413, 511) identified in the image data at a time when the first word is uttered, among a plurality of objects included in the image data, and an operation of providing the associated information for the first object and the second object.
- the designated word may include a second word (201-1) and a third word (201-2) designating a single object.
- the operating method of the electronic device (310) may include an operation of identifying a first object corresponding to a second gesture (221) identified in the image data at a time when the second word is uttered among a plurality of objects included in the image data, an operation of identifying a second object corresponding to a third gesture (223) identified in the image data at a time when the third word is uttered among a plurality of objects included in the image data, and an operation of providing the association information for the first object and the second object.
- the operating method of the electronic device (310) may include an operation of identifying a second object in the image data when the third word is uttered within a certain time after the second word is uttered.
- the operating method of the electronic device (310) may include an operation of acquiring first image data (260) and second image data (270) corresponding to different fields of view, an operation of identifying a first object corresponding to a second gesture (221) identified in the first image data at a time when the second word is uttered among a plurality of objects included in the first image data, and an operation of identifying a second object corresponding to a third gesture (223) identified in the second image data at a time when the third word is uttered among a plurality of objects included in the second image data.
- the operating method of the electronic device (310) may include an operation of acquiring motion data (341) from an external device (320) while the speech data is acquired, and an operation of identifying at least two objects if the motion data satisfies a specified condition.
- the operating method of the electronic device (310) may include an operation of inputting the speech data and the image data into an artificial intelligence model (3131) stored in the electronic device (310) (e.g., memory (313)) and an operation of identifying the at least two objects based on an output of the artificial intelligence model.
- an artificial intelligence model 3131
- the electronic device (310) e.g., memory (313)
- a computer-readable recording medium may be configured such that when executed by an electronic device, the electronic device obtains speech data including a designated word related to designation of an object, obtains image data while the speech data is being obtained, identifies at least two objects corresponding to a gesture identified in the image data at a time when the designated word is uttered among a plurality of objects included in the image data, and provides associated information about the at least two objects.
- An image recognition system may include a first electronic device (310) and a second electronic device (320).
- the first electronic device (310) includes at least one first processor (311), a camera (314), a microphone (315), and a first memory (313) operatively connected to the at least one first processor, the camera, and the microphone and storing at least one first command, wherein the at least one first command, when executed by the at least one first processor, causes the first electronic device to: acquire image data (210) through the camera while utterance data (201) including a designated word related to designation of an object is acquired through the microphone; identify at least two objects corresponding to a gesture identified in the image data at the time when the designated word is uttered, among a plurality of objects (211 to 217) included in the image data; and provide associated information (250) for the at least two objects.
- the at least one first command when executed by the at least one first processor, causes the first electronic device to: acquire image data (210) through the camera while utterance data (201) including a designated word related to designation of an object is acquired through the microphone; identify at least two objects corresponding to a gesture identified in the image data
- the second electronic device (320) includes at least one second processor (321), a sensor (324), and a second memory (323) operatively connected to the at least one second processor and the sensor and storing at least one second instruction, wherein the at least one second instruction, when executed by the at least one second processor, is configured to cause the second electronic device to provide information related to a posture of the second electronic device obtained through the sensor to the first electronic device.
- the at least one first instruction when executed by the at least one first processor, may be configured to cause the first electronic device to: identify at least two objects corresponding to a gesture identified in the image data if information related to a posture of the second electronic device provided from the second electronic device satisfies a specified condition.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Library & Information Science (AREA)
- Computational Linguistics (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Evolutionary Computation (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Psychiatry (AREA)
- Social Psychology (AREA)
- Artificial Intelligence (AREA)
- Mathematical Physics (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
다양한 실시 예에 따른 전자 장치는 적어도 하나의 프로세서, 카메라, 마이크 및 상기 적어도 하나의 프로세서, 상기 카메라 및 상기 마이크와 동작적으로 연결되며 적어도 하나의 명령어를 저장하는 메모리(313)를 포함하고, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터가 상기 마이크를 통해 획득되는 동안, 상기 카메라를 통해 영상 데이터를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하도록 구성될 수 있다.
Description
본 문서에서 개시되는 실시 예들은 영상 인식 방법 및 이를 지원하는 전자 장치와 관련된다.
디지털 기술의 발달과 함께 이동통신 단말기, 전자 수첩, 스마트 폰, 태블릿 PC, 웨어러블 디바이스와 같이 이동하면서 통신 및 개인정보 처리가 가능한 전자 장치들이 다양하게 출시되고 있다. 이러한 전자 장치들은 단순한 음성 통화 및 메시지 전송 기능에서 영상 통화, 전자 수첩 기능, 문서 기능, 이메일 기능, 인터넷 기능, 촬영 기능과 같은 다양한 기능을 구비하게 되었다.
또한, 전자 장치는 영상 데이터 내의 객체를 인식하여 다양한 서비스에 적용하는 영상 인식 기능을 제공할 수 있다. 예를 들어, 전자 장치는 영상 데이터에 포함된 다수의 객체 중, 사용자가 지시하는 대상 객체에 대한 정보(예: 객체의 속성 정보)를 제공할 수 있다.
상술한 정보는 본 개시에 대한 이해를 돕기 위한 목적으로 하는 배경 기술(related art)로 제공될 수 있다. 상술한 내용 중 어느 것도 본 개시와 관련된 종래 기술(prior art)로서 적용될 수 있는지에 대하여 어떠한 주장이나 결정이 제기되지 않는다.
일반적으로, 전자 장치는 사용자의 시선이나 손가락이 가리키는 방향을 식별하고, 영상 데이터에서 사용자의 시선이나 손가락이 가리키는 방향에 대응되는 객체를 대상 객체로 식별할 수 있다.
하지만, 전자 장치는 대상 객체를 식별함에 있어서, 사용자의 시선이나 손가락이 가리키는 방향을 지속적으로 모니터링해야 하는 문제점이 있다. 이러한 모니터링은 전자 장치의 리소스 낭비를 야기할 수 있다.
또한, 전자 장치는 사용자의 시선이 유지되거나 손가락이 가리키고 있는 하나의 객체를 대상 객체로 식별할 수 있다. 예를 들어, 전자 장치는 영상 데이터 내의 복수의 대상 객체들을 동시에 식별하고 이들에 대한 연관 정보를 제공하는데 어려움이 있다.
본 문서에서 이루고자 하는 기술적 과제는 이상에서 언급한 기술적 과제로 제한되지 않으며, 언급되지 않은 또 다른 기술적 과제들은 아래의 기재로부터 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
다양한 실시 예에 따른 전자 장치는 적어도 하나의 프로세서, 카메라, 마이크 및 상기 적어도 하나의 프로세서, 상기 카메라 및 상기 마이크와 동작적으로 연결되며 적어도 하나의 명령어를 저장하는 메모리(313)를 포함하고, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치로 하여금: 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터가 상기 마이크를 통해 획득되는 동안, 상기 카메라를 통해 영상 데이터를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하도록 구성될 수 있다.
다양한 실시 예에 따른 전자 장치의 동작 방법은 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터를 획득하는 동작, 상기 발화 데이터가 획득되는 동안, 영상 데이터를 획득하는 동작, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하는 동작 및 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따른 컴퓨터로 판독 가능한 기록 매체는, 전자 장치에 의해 실행되었을 때, 상기 전자 장치가, 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터를 획득하고, 상기 발화 데이터가 획득되는 동안, 영상 데이터를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하도록 설정될 수 있다.
다양한 실시 예에 따른 영상 인식 시스템은 제1 전자 장치와 제2 전자 장치를 포함하며, 상기 제1 전자 장치는 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터가 마이크를 통해 획득되는 동안, 카메라를 통해 영상 데이터를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하도록 구성되며, 상기 제2 전자 장치는 센서를 통해 획득되는 상기 제2 전자 장치의 자세와 관련된 정보를 상기 제1 전자 장치로 제공하도록 구성되고, 상기 제1 전자 장치는 상기 제2 전자 장치의 자세와 관련된 정보가 지정된 조건을 만족하는 경우, 상기 영상 데이터에서 식별되는 제스처에 대응되는 상기 적어도 두 개의 객체를 식별하도록 구성될 수 있다.
본 문서에 개시되는 다양한 실시 예들에 따른 전자 장치는 리소스 낭비를 감소하고, 영상 데이터 내의 복수의 객체들 선택 및 이들에 대한 연관 정보를 제공을 가능하게 할 수 있다.
본 개시에서 얻을 수 있는 효과는 이상에서 언급한 효과들로 제한되지 않으며, 언급하지 않은 또 다른 효과들은 아래의 기재로부터 본 개시가 속하는 기술 분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
도면의 설명과 관련하여, 동일 또는 유사한 구성 요소에 대해서는 동일 또는 유사한 참조 부호가 사용될 수 있다.
도 1은, 본 문서 내에서 설명된 동작들을 수행할 수 있는 예시적인 전자 장치의 블록도이다.
도 2a는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다.
도 2b는 다양한 실시 예에 따른 영상 인식 기능을 통해 제공되는 정보를 설명하기 위한 도면이다.
도 2c는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다.
도 2d는 다양한 실시 예에 따른 전자 장치가 대상 객체가 비교 대상으로 적절한지 판단하는 구성을 설명하기 위한 도면이다.
도 2e는 다양한 실시 예에 따른 전자 장치가 비교 대상을 다시 선택하도록 가이드 하는 구성을 설명하기 위한 도면이다.
도 2f는 다양한 실시 예에 따른 대상 객체에 대한 정보를 제공하는 동작을 설명하기 위한 도면이다.
도 3a는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면이다.
도 3b는 다양한 실시 예에 따른 영상 인식 시스템의 구성을 개략적으로 도시한 도면이다.
도 3c는 객체 식별에 활용되는 모션 데이터를 설명하기 위한 도면이다.
도 4는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다.
도 5a는 다양한 실시 예에 따른 제1 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다.
도 5b는 다양한 실시 예에 따른 제1 전자 장치에서 사용자의 제스처를 인식하는 동작을 설명하기 위한 도면이다.
도 5c는 다양한 실시 예에 따른 제스처에 기반하여 제어되는 제1 전자 장치의 동작을 설명하기 위한 도면이다.
도 6은 다양한 실시 예에 따른 정보 제공 모델의 구성을 도시한 도면이다.
도 7은 다양한 실시 예에 따른 객체 식별 모델의 입력 데이터를 처리하는 절차를 설명하기 위한 도면이다.
도 8a는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면들이다.
도 8b는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면들이다.
도 9a는 다양한 실시 예에 따른 전자 장치의 동작을 도시한 흐름도이다.
도 9b는 다양한 실시 예에 따른 연관 정보를 설명하기 위한 도면이다.
도 10은 다양한 실시 예에 따른 전자 장치의 모션 데이터 획득 동작을 도시한 흐름도이다.
도 11은 다양한 실시 예에 따른 전자 장치의 대상 객체 식별 동작을 도시한 흐름도이다.
도 12는 다양한 실시 예에 따른 전자 장치의 시야 보정 동작을 도시한 흐름도이다.
도 13은 다양한 실시 예에 따른 시야 보정 과정을 설명하기 위한 도면이다.
도 14는 다양한 실시 예에 따른 전자 장치에서 연관 정보를 제공하는 동작을 도시한 흐름도이다.
도 15는 다양한 실시 예에 따른 전자 장치의 동작을 도시한 다른 흐름도이다.
도 16a 내지 도 16c는 다양한 실시 예에 제1 전자 장치의 동작을 설명하기 위한 도면이다.
이하, 본 문서의 다양한 실시 예가 첨부된 도면을 참조하여 기재된다. 그러나, 이는 본 문서에 기재된 기술을 특정한 실시 형태에 대해 한정하려는 것이 아니며, 본 문서의 실시 예의 다양한 변경(modifications), 균등물(equivalents), 및/또는 대체물(alternatives)을 포함하는 것으로 이해되어야 한다. 도면의 설명과 관련하여, 유사한 구성 요소에 대해서는 유사한 참조 부호가 사용될 수 있다.
도 1은, 본 문서 내에서 설명된 동작들을 수행할 수 있는 예시적인(exemplary) 전자 장치(100)의 블록도이다.
도 1을 참조하면, 전자 장치(100)는, 노트북(190), 다양한 폼 팩터들을 가지는 스마트폰들(191)((예: 바 타입의 스마트폰(191-1), 폴더블 타입의 스마트폰(191-2), 또는 슬라이더블(또는 롤러블) 타입의 스마트폰(191-3)), 태블릿(192), 셀룰러 전화(미도시), 및 기타 유사 컴퓨팅 장치들(미도시)과 같은, 다양한 형태들의 전자 장치들 중 하나일 수 있다. 도 1 내에서 도시된 구성 요소들, 그들의 관계들, 및 그들의 기능들은, 예시적일 뿐이며, 본 문서 내에서 설명되거나 청구된 구현들을 제한하는 것이 아니다. 전자 장치(100)는, 모바일 장치, 사용자 장치, 다기능 장치, 휴대용 장치 또는 서버로 참조될 수 있다.
전자 장치(100)는, 적어도 하나의 프로세서(110)(이하, 프로세서(110)으로 참조), 적어도 하나의 메모리(120)(이하, 메모리(120)으로 참조), 적어도 하나의 디스플레이(140)(이하, 디스플레이(140)으로 참조), 적어도 하나의 이미지 센서(150)(이하, 이미지 센서(150)으로 참조), 적어도 하나의 통신 회로(160)(이하, 통신 회로(160)으로 참조), 및/또는 적어도 하나의 센서(170)(이하, 센서(170)으로 참조)를 포함하는 구성 요소들을 포함할 수 있다. 상기 구성 요소들은 단지 예시적인 것이다. 예를 들어, 전자 장치(100)는 다른 구성 요소들(예: PMIC(power management integrated circuitry), 오디오 처리 회로, 안테나, 재충전가능한(rechargeable) 배터리 또는 입출력 인터페이스)을 포함할 수 있다. 예를 들어, 일부 구성 요소들은 전자 장치(100)로부터 생략될 수 있다. 예를 들어, 몇몇 구성 요소들은 하나의 구성 요소로 통합될 수 있다.
프로세서(110)는, 하나 이상의 IC(integrated circuit(또는 circuitry)) 칩으로 구현될 수 있고, 다양한 데이터 처리들을 실행할 수 있다. 프로세서(110)는 적어도 하나의 전기적 회로를 포함할 수 있고, 메모리(120) 내에 저장된 인스트럭션들(또는 프로그램, 데이터)을 개별적으로 또는 집합적으로 분산 처리할 수 있다. 프로세서(110)는 하나 이상의 프로세싱 회로들을 포함하는 프로세서 집합체를 포함할 수 있다. 프로세서(110)는, 전자 장치(100)의 하나 이상의 구성 요소들(예: 메모리(120), 디스플레이(140), 이미지 센서(150), 통신 회로(160), 및/또는 센서(170))의 수행(performance) 및 동작들을 제어하기 위해 작동적인(operative) 어느(any) 프로세싱 회로를 포함할 수 있다. 예를 들어, 프로세서(110)(예: AP(application processor))는, SoC(system on chip)(예를 들어, 하나의 칩 또는 칩셋)로 구현될 수 있다. 예를 들어, 프로세서(110)는, 다수의 코어들(또는 적어도 하나의 코어 회로), 다수의 칩들 또는 다수의 칩셋들로 구현될 수 있다. 예를 들어, 프로세서(110)는, 하나 이상의 프로세싱 회로들을 포함할 수 있다. 예를 들어, 프로세서(110)는 본 개시의 여러 기능들을 개별적으로, 및/또는 집합적으로(collectively) 수행하도록 구성된 하나 이상의 프로세싱 회로를 포함할 수 있다. 제한하지 않는 예로, 프로세서(110)의 적어도 일부분은 전자 장치(100)의 제1 칩에 포함되고, 프로세서(110)의 적어도 다른 부분은 전자 장치(100)의 상기 제1 칩과 다른 전자 장치(100)의 제2 칩에 포함될 수 있다.
예를 들어, 프로세서(110)는, CPU(central processing unit)(111), GPU(graphics processing unit)(112), NPU(neural processing unit)(113), ISP(image signal processor)(114), 디스플레이 컨트롤러(115), 메모리 컨트롤러(116), 스토리지(storage) 컨트롤러(117), CP(communication processor)(118), 및/또는 센서 인터페이스(119)를 포함할 수 있다. 프로세서(110)의 이러한 구성 요소들은, 단지 예시적인 것이다. 예를 들어, 프로세서(110)는, 다른 구성 요소들을 더 포함할 수 있다. 예를 들어, 프로세서(110)의 몇몇 구성 요소들은, 프로세서(110)로부터 생략될 수 있다. 예를 들어, 프로세서(110)의 몇몇 구성 요소들은, 프로세서(110) 외부에서 전자 장치(100)의 별도의 구성 요소로 포함될 수 있다. 예를 들어, 프로세서(110)의 일부 구성 요소들(예: 메모리 컨트롤러(116))은, 다른 구성 요소들(예: 메모리(120)의 적어도 일부, 인터페이스(예: 전자 장치(100)의 적어도 하나의 구성 요소에 연결하기 위해 이용 가능함), 디스플레이(140) 및/또는 이미지 센서(150)) 내에 포함될 수 있다.
프로세서(110)는 메모리(120)에 저장된 명령어들을 실행함으로써 전자 장치(100)의 동작들을 제어할 수 있다. 예를 들면, 프로세서(110)는 복수의 동작들을 프로세서들 사이에서 분할하여 집합적으로 수행하는 복수의 프로세서들에 대응될 수 있다.
프로세서(110)는, 메모리(120) 내에 저장된 인스트럭션들을 실행함으로써 다양한 동작들을 수행하도록 전자 장치(100)의 다른 구성 요소들을 야기할 수 있다. CPU(111)(또는 중앙 처리 회로)는 메모리(120)(예: 휘발성 메모리(121) 및/또는 비휘발성 메모리(122)) 내에 저장된 인스트럭션들의 실행에 기반하여 프로세서(110)의 구성 요소들을 제어하도록 구성될 수 있다. GPU(112)(또는 그래픽 처리 회로)는 병렬 연산(예: 렌더링(rendering))들을 실행하도록 구성될 수 있다. NPU(113)(또는 뉴럴 처리 회로, 또는 AI(artificial intelligence) 칩)는 인공 지능 모델을 위한 연산(예: 합성곱 연산(convolution computation))들을 실행하도록 구성될 수 있다. ISP(114)(또는 이미지 신호 처리 회로)는 이미지 센서(150)를 통해 획득된 원시 이미지(raw image)를 전자 장치(100) 내의 구성 요소 또는 프로세서(110)의 구성 요소를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 디스플레이 컨트롤러(115)(또는 디스플레이 제어 회로, 또는 DPU(display processing unit))는 CPU(111), GPU(112), ISP(114), 또는 메모리(120)(예: 휘발성 메모리(121))로부터 획득된 이미지를 디스플레이(140)를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 메모리 컨트롤러(116)(또는 메모리 제어 회로)는 휘발성 메모리(121)로부터 데이터를 읽고 데이터를 휘발성 메모리(121)에 기록하는 것을 제어하도록 구성될 수 있다. 스토리지 컨트롤러(117)(또는 스토리지 제어 회로)는, 비휘발성 메모리(122)로부터 데이터를 읽고 데이터를 비휘발성 메모리(122)에 기록하는 것을 제어하도록 구성될 수 있다. CP(118)(통신 처리 회로)는 프로세서(110)의 구성 요소로부터 획득된 데이터를 통신 회로(160)를 통해 다른 전자 장치에게 송신하는 것을 위해 적합한 포맷으로 처리하거나, 다른 전자 장치로부터 통신 회로(160)를 통해 획득된 데이터를 프로세서(110)의 구성 요소의 처리를 위해 적합한 포맷으로 처리하도록 구성될 수 있다. 예를 들어, 통신 회로(160)는 하나 이상의 통신 회로들을 포함할 수 있다. 센서 인터페이스(119)(또는 센싱 데이터 처리 회로, 센서 허브)는 센서(170)를 통해 획득된, 전자 장치(100)의 상태 및/또는 전자 장치(100) 주변의 상태에 대한 데이터를 프로세서(110)의 구성 요소를 위해 적합한 포맷으로 처리하도록 구성될 수 있다.
메모리(120)는, 하나 이상의 저장 매체들(storage medium)(또는 하나 이상의 저장 장치들)을 포함할 수 있다. 예를 들면, 메모리(120)는, 하나 이상의 저장 매체들을 포함하는 메모리 집합체를 포함할 수 있다. 예를 들면, 상기 하나 이상의 저장 매체들은, 하드 드라이브, 플래시 메모리, ROM(read-only memory)과 같은 영구 메모리(permanent memory)(예, 비휘발성 메모리(122)), RAM(random access memory)과 같은 반-영구 메모리(semi-permanent memory)(예, 휘발성 메모리(121)), 어느 다른 적합한 유형(any other suitable type)의 저장소(또는 저장 집합체(storage assembly)), 또는 이들의 어떤 조합(any combination thereof)을 포함할 수 있다. 메모리(120)는, 전자 장치(100)의 기능(function or feature)을 위한 데이터를 일시적으로(temporarily) 저장하기 위해 이용되는 하나 이상의 다른 유형들(one or more different types)의 메모리인 캐시 메모리를 포함할 수 있다. 제한되지 않는 예로, 상기 캐시 메모리는 프로세서(110) 내에 포함될 수 있다. 메모리(120)는, 전자 장치(100) 안에 고정적으로(fixedly) 임베디드될 수 있거나, 반복적으로(repeatedly) 전자 장치(100) 안으로 삽입되고 전자 장치(100)로부터 제거될 수 있는 하나 이상의 적합한 유형의 구성 요소들(onto one or more suitable types of components)(예: SIM(subscriber identity module) 카드 및/또는 SD(secure digital) 카드)에 통합될(incorporated) 수 있다.
예를 들어, 메모리(120)는, 운영 체제(operating system)(또는 시스템)소프트웨어 어플리케이션, 펌웨어 소프트웨어 어플리케이션, 드라이버 소프트웨어 어플리케이션, 플러그인(예, 애드-인, 애드-온, 및/또는 어플릿) 소프트웨어 어플리케이션, 및/또는 어느 다른(any other) 적합한(suitable) 소프트웨어 어플리케이션들과 같은 하나 이상의 소프트웨어 어플리케이션들을 저장할 수 있다. 예를 들어, 상기 하나 이상의 소프트웨어 어플리케이션들은, 프로세서(110)에 의해 실행가능한 인스트럭션들을 포함할 수 있다. 예를 들어, 메모리(120)는, API(application programming interface)에 의해 호출가능한 인스트럭션들을 저장할 수 있다. 예를 들어, 메모리(120)는, 라이브러리 내에서 인스트럭션들을 저장할 수 있다.
다양한 실시 예에 따르면, 전술한 전자 장치(100)는 영상 데이터 내의 객체를 인식(또는 식별)하여 다양한 서비스에 적용하는 영상 인식 기능을 제공할 수 있다. 이와 관련하여는 이하의 도 2a 내지 도 15를 통해 상세히 설명하도록 한다. 또한, 이하의 도 2a 내지 도 15를 통해 설명되는 다양한 실시 예들 중 적어도 하나의 실시 예는 다른 실시 예와 조합될 수도 있다.
도 2a는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다. 그리고, 도 2b는 다양한 실시 예에 따른 영상 인식 기능을 통해 제공되는 정보를 설명하기 위한 도면이고, 도 2c는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이고, 도 2d는 다양한 실시 예에 따른 전자 장치가 대상 객체가 비교 대상으로 적절한지 판단하는 구성을 설명하기 위한 도면이고, 도 2e는 다양한 실시 예에 따른 전자 장치가 비교 대상을 다시 선택하도록 가이드 하는 구성을 설명하기 위한 도면이며, 도 2f는 다양한 실시 예에 따른 대상 객체에 대한 정보를 제공하는 동작을 설명하기 위한 도면이다.
도 2a의 200은 사용자가 검지 손가락(index finger)으로 복수의 특정 객체들을 순차적으로 가리키면서(예: 포인팅 제스처를 반복하면서) 발화하는 상황을 나타내며, 도 2a의 230은 사용자가 발화하는 동안 전자 장치(100)에 의해 획득되는 영상 데이터(210)를 나타낸다.
도 2a를 참조하면, 다양한 실시 예에 따른 전자 장치(100)는 발화 데이터(또는 발화 입력)(201) 및 영상 데이터(210)에 기반하여, 영상 데이터(210)에 포함된 다수의 객체들 중 사용자가 지시하는 대상 객체(또는 비교 대상)를 식별할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 발화 데이터(201)에 기반하여 영상 데이터(210) 내의 복수의 대상 객체들을 식별할 수 있으며, 이들에 대한 연관 정보를 제공할 수 있다. 이와 관련된 다양한 실시 예에 따른 영상 인식 기능을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따르면, 전자 장치(100)는 발화 데이터(201)를 획득하고, 이에 대한 제1 부분, 제2 부분 및 제3 부분을 식별할 수 있다.
일 실시 예에 따르면, 발화 데이터(201)의 제1 부분(및 제2 부분)은 제1 대상 객체(및 제2 대상 객체)를 지정하는 지정된 단어(예: 이것, 저것, 이곳, 저곳, 그는, 그녀는, 너는 같은 단일의 대상을 지시하는 지시 대명사)에 대응될 수 있으며, 발화 데이터(201)의 제3 부분은 지정된 질의에 대응될 수 있다.
예를 들어, 전자 장치(100)는 발화 데이터(201)(예: 이 아이와 저 아이 중 누가 더 크지?)에 대하여, 제1 대상 객체(또는 하나의 객체)를 지정하는 지정된 제1 단어(예: 이 아이)(201-1)를 제1 부분으로 식별하고, 제2 대상 객체(예: 다른 하나의 객체)를 지정하는 지정된 제2 단어(예: 저 아이)(201-2)를 제2 부분으로 식별하고, 질의에 해당되는 문장(예: 다 자라면 어떤 아이가 더 커지지?)(201-3)을 제3 부분으로 식별할 수 있다.
다양한 실시 예에 따르면, 전자 장치(100)는 발화 데이터(201)가 획득되는 동안 영상 데이터(210)를 획득할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는, 대상 객체를 식별하기에 앞서, 영상 데이터(210)에 포함된 객체를 인식(또는 추출)할 수 있다.
예를 들어, 전자 장치(100)는 영상 데이터(210)로부터 추출되는 특징 데이터(예: 특징점)에 기반하여 제1 객체(211), 제2 객체(213), 제3 객체(215) 및 제4 객체(217)를 인식할 수 있다. 이와 관련하여, 전자 장치(100)는 HOG(histogram of oriented gradient), SIFT(scale invariant feature transform), LBP(local binary pattern), MCT(modified census transform)와 같은 공지된 다양한 기술을 이용하여 특징 데이터를 추출할 수 있다.
다양한 실시 예에 따르면, 전자 장치(100)는 발화 데이터(201)로부터 식별되는 제1 부분과 제2 부분에 기반하여, 인식된 객체들(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 대상 객체를 식별할 수 있다.
실시 예에 따라, 전자 장치(100)는, 도 2a의 230에 도시된 바와 같이, 제1 부분(예: 제1 단어(201-1))이 식별된 제1 시점에(또는 제1 시점 이후에), 영상 데이터(210)에서 식별되는 검지 손가락의 제1 위치(221)에 대응되는 객체(예: 제2 객체(213))를 제1 대상 객체로 식별할 수 있다. 예를 들어, 전자 장치(100)는 제1 위치(221)를 기준으로 일정 범위(또는 일정 거리) 내에 존재하는 제2 객체(213)를 제1 대상 객체로 식별할 수 있다.
이와 유사하게, 전자 장치(100)는 제2 부분(예: 제2 단어(201-2))이 식별된 제2 시점에(또는 제2 시점 이후에), 영상 데이터(210)에서 식별되는 검지 손가락의 제2 위치(223)에 대응되는 객체(예: 제3 객체(215))를 제2 대상 객체로 식별할 수 있다. 예를 들어, 제1 위치(221)에 위치한 검지 손가락이, 제2 시점에 제3 객체(215)를 기준으로 일정 범위 내에 존재하는 제2 위치(223)로 이동되면, 전자 장치(100)는 제3 객체(215)를 제2 대상 객체로 식별할 수 있다.
다양한 실시 예에 따르면, 전자 장치(100)는, 발화 데이터(201)의 제3 부분(201-3)에 기반하여, 식별된 대상 객체들(예: 제1 대상 객체(예: 제2 객체(213))와 제2 대상 객체(예: 제3 객체(215)))에 대한 연관 정보를 제공할 수 있다.
일 실시 예에 따르면, 도 2b에 도시된 바와 같이, 식별된 대상 객체에 대한 정보(250)는 제1 대상 객체(예: 푸들)와 관련된 제1 정보(251), 제2 대상 객체(예: 웰시코기(welsh corgi))와 관련된 제2 정보(253) 그리고 이들에 대한 상세 정보인 제3 정보(255) 중 적어도 하나를 포함할 수 있다. 연관 정보는 제1 대상 객체와 제2 대상 객체들에 대한 연관 정보(예: 비교 정보)를 포함할 수 있다. 이와 관련하여, 전자 장치(100)는 제1 대상 객체, 제2 대상 객체 및 데이터(201)의 제3 부분(201-3)을 검색어(예: 푸들과 웰시코기 중 다 자라면 어떤 아이가 더 커지지?)로 활용하고, 이에 대한 검색 결과를 외부(예: 검색 서버)로부터 획득할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 연관 정보는 제1 대상 객체와 제2 대상 객체 각각에 대한 정보를 포함할 수 있다. 이와 관련하여, 전자 장치(100)는 제1 대상 객체, 제2 대상 객체 및 각 대상 객체의 종류를 문의하는 질의를 검색어(예: 푸들의 종류와 웰시코기의 종류에 대해 설명해줘)로 활용하고, 이에 대한 검색 결과(예: 푸들에 대한 설명과 웰시코기에 대한 설명)를 외부로부터 획득할 수도 있다.
다양한 실시 예에 따르면, 전자 장치(100)는, 식별된 대상 객체에 대한 연관 정보(250)를 전자 장치(100)를 통해 출력하거나 다른 전자 장치를 통해 출력할 수도 있다. 일 실시 예에 따르면, 식별된 대상 객체에 대한 연관 정보(250)는, 전자 장치(100)의 디스플레이(140)를 통해 시각적 정보로 출력될 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 식별된 대상 객체에 대한 연관 정보(250)는, 도 2f의 295에 도시된 바와 같이, 전자 장치(100)의 스피커를 통해 청각적 정보로 출력되거나, 도 2f의 297에 도시된 바와 같이, 전자 장치(100)와 통신으로 연결된 다른 전자 장치(예: 무선 이어폰(298), 스마트 워치, 다른 스마트 폰 또는 AI 스피커)를 통해 출력될 수도 있다.
추가적으로 또는 선택적으로, 다양한 실시 예에 따른 전자 장치(100)는 복수의 영상 데이터들에서 대상 객체를 식별할 수도 있다.
실시 예에 따라, 전자 장치(100)는, 도 2c에 도시된 바와 같이, 제1 부분(예: 제1 단어(201-1))이 식별된 제1 시점에, 제1 시야(field of view)(예: 11시 방향)에 대응하는 제1 영상 데이터(260)를 획득할 수 있다. 더하여, 전자 장치(100)는 제2 부분(예: 제2 단어(201-2))이 식별된 제2 시점에, 제1 시야와 다른 제2 시야(예: 1시 방향)에 대응하는 제2 영상 데이터(270)를 획득할 수 있다.
이와 관련하여, 전자 장치(100)는 제1 부분이 식별된 제1 시점에, 제1 영상 데이터(260)에서 식별되는 검지 손가락의 제1 위치(261)에 대응되는 객체(예: 제2 객체(213))를 제1 대상 객체로 식별할 수 있다. 더하여, 전자 장치(100)는 제2 부분이 식별된 제2 시점에, 제2 영상 데이터(270)에서 식별되는 검지 손가락의 제2 위치(271)에 대응되는 객체(예: 제3 객체(215))를 제2 대상 객체로 식별할 수 있다.
추가적으로 또는 선택적으로, 다양한 실시 예에 따른 전자 장치(100)는, 식별된 대상 객체들(예: 제1 대상 객체와 제2 대상 객체)에 대한 연관 정보(예: 비교 정보)를 제공함에 있어서, 제1 대상 객체의 속성(예: 종류)과 제2 대상 객체의 속성을 비교할 수도 있다.
일 실시 예에 따르면, 전자 장치(100)는 속성 비교를 통해 제1 대상 객체와 제2 대상 객체가 비교 대상으로 적절한지를 판단할 수 있다. 예를 들어, 도 2d의 208-1에 도시된 바와 같이, 사용자가 지시하는 제1 대상 객체(예: 제1 스마트 폰)(281)와 제2 대상 객체(예: 제2 스마트 폰)(282)가 확인되는 경우, 전자 장치(100)는 제1 대상 객체(281)와 제2 대상 객체(282)가 서로 관련된 속성을 가지는지를 판단하고, 서로 관련된 속성(예: 모바일 디바이스)을 가지는 제1 대상 객체(281)와 제2 대상 객체(282)에 대하여 비교 대상으로 적절하다고 판단할 수 있다. 더하여, 전자 장치(100)는 비교 대상으로 적절하다고 판단된 대상 객체들(예: 제1 대상 객체와 제2 대상 객체)에 대한 연관 정보(250)를 제공할 수 있다.
예를 들어, 도 2d의 208-2에 도시된 바와 같이, 사용자가 지시하는 제2 대상 객체(예: 제2 스마트 폰)(282)와 제3 대상 객체(예: 노트북)(283)가 확인되는 경우, 전자 장치(100)는 제2 대상 객체(282)와 관련된 속성(예: 모바일 디바이스)과 제3 대상 객체(283)와 관련된 속성(예: 컴퓨터 디바이스)을 확인하고, 서로 다른 속성을 가지는 제2 대상 객체(282)와 제3 대상 객체(283)에 대하여는 비교 대상으로 적절하지 않다고 판단할 수 있다. 더하여, 전자 장치(100)는, 대상 객체들(예: 제2 대상 객체(282) 및 제3 대상 객체(283))이 비교 대상으로 적절하지 않다고 판단하는 경우, 도 2e에 도시된 바와 같이, 비교 대상을 다시 선택하도록 하는 가이드 정보(291) 또는 질의를 다시 입력하도록 하는 가이드 정보를 제공할 수도 있다. 실시 예에 따라, 전자 장치(100)는 비교 대상으로 적절하지 않은 원인과 관련된 정보(예: 지정된 대상 객체의 속성이 서로 상이합니다.)를 가이드 정보(291)의 적어도 일부로 제공할 수도 있다.
전술한 바와 같이, 다양한 실시 예에 따른 영상 인식 기능은 전자 장치(100)의 독자적인 동작에 의해 제공될 수 있다. 실시 예에 따라, 전자 장치(100)는 적어도 하나의 다른 장치와의 협업을 통해 영상 인식 기능을 제공할 수도 있다. 이와 관련하여는 이하의 도 3a 내지 도 3c를 통해 상세히 설명하도록 한다.
도 3a는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면이다. 그리고, 도 3b는 다양한 실시 예에 따른 영상 인식 시스템의 구성을 개략적으로 도시한 도면이고, 도 3c는 객체 식별에 활용되는 모션 데이터를 설명하기 위한 도면이다.
도 3a 내지 도 3c를 참조하면, 다양한 실시 예에 따른 영상 인식 시스템(300)은 제1 전자 장치(310)와 제2 전자 장치(320)로 구성될 수 있다. 일 실시 예에 따르면, 제1 전자 장치(310)와 제2 전자 장치(320)는 네트워크(예: 근거리 통신 네트워크 또는 원거리 통신 네트워크)를 통하여 서로 통신할 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 스마트폰으로 참조될 수 있으며, 제2 전자 장치(320)는 신체의 일부(예: 검지 손가락)에 착용되는 웨어러블 장치(예: 반지 형태의 전자 장치)로 참조될 수 있다.
그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예를 들어, 제1 전자 장치(310)는 접혀진 상태로 사용자가 착용한 의복(예: 의복의 외부 주머니(outside pocket)에 부착될 수 있는 폴더블 타입의 스마트폰(191-2) 또는 사용자가 착용한 의복(예: 셔츠) 또는 액세서리(예: 넥타이, 목걸이)에 부착이 가능한 핀(또는 클립) 타입의 웨어러블 장치로 참조될 수도 있다. 실시 예에 따라, 제1 전자 장치(310)와 제2 전자 장치(320)가 스마트폰으로 참조될 수 있으며, 제1 전자 장치(310)와 제2 전자 장치(320) 중 적어도 하나가 신체의 다른 일부에 착용되는 웨어러블 장치(예: 손목에 착용 가능한 시계 형태의 전자 장치, 목에 착용 가능한 목걸이 형태의 전자 장치, 얼굴에 착용 가능한 안경 형태의 전자 장치, 두부에 착용 가능한 증강현실 또는 가상현실 장치)로 참조될 수도 있다.
다양한 실시 예에 따르면, 제1 전자 장치(310)는 제2 전자 장치(320)와의 협업을 통해 영상 인식 기능을 제공할 수 있다.
일 실시 예에 따르면, 제1 전자 장치(310)는 발화 데이터(201)와 영상 데이터(210)를 획득(또는 수집)하고, 제2 전자 장치(320)는 신체에 착용된 상태에서 제2 전자 장치(320)의 자세와 관련된 모션 데이터(341)(예: 제2 전자 장치가 착용된 신체의 일부(예: 검지 손가락)에 대한 모션 데이터)를 획득할 수 있다. 또한, 제1 전자 장치(310)는 제1 전자 장치(310)에 의해 획득되는 발화 데이터(201)와 영상 데이터(210) 그리고, 제2 전자 장치(320)에 의해 획득되는 모션 데이터(341)에 기반하여, 영상 데이터(210)에 포함된 다수의 객체들(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 대상 객체를 식별할 수 있다.
이와 관련된 다양한 실시 예에 따른 제1 전자 장치(310)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
도 3b를 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)는 적어도 하나의 제1 프로세서(311)(이하, 제1 프로세서(311)로 칭함), 적어도 하나의 제1 통신 회로(312)(이하, 제1 통신 회로(312)로 칭함), 적어도 하나의 제1 메모리(313)(이하, 제1 메모리(313)로 칭함), 적어도 하나의 이미지 센서(314)(또는 적어도 하나의 카메라)(이하, 이미지 센서(314)로 칭함), 적어도 하나의 마이크(315)(이하, 마이크(315)로 칭함) 및 적어도 하나의 출력 장치(316)(이하, 출력 장치(316)로 칭함)로 구성될 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 전술한 구성 요소들 보다 많은 구성 요소들을 가지거나, 또는 그 보다 적은 구성 요소를 가지는 것으로 구현될 수 있다. 예컨대, 제1 전자 장치(310)는 도 1에 도시된 전자 장치(100)에 해당될 수 있다.
다양한 실시 예에 따르면, 제1 통신 회로(312)는 제2 전자 장치(320)와의 무선 통신 수행을 지원할 수 있다. 일 실시 예에 따르면, 제1 통신 회로(312)는 제1 전자 장치(310)와 제2 전자 장치(320) 사이의 신호(예: 명령 또는 데이터)를 송수신하기 위한 하드웨어 및 소프트웨어를 포함하는 장치일 수 있다.
예를 들어, 제1 통신 회로(312)는 제1 네트워크(예: 블루투스, WiFi(wireless fidelity) direct 또는 IrDA(infrared data association)와 같은 근거리 통신 네트워크) 또는 제2 네트워크(예: 레거시 셀룰러 네트워크, 5G 네트워크, 차세대 통신 네트워크, 인터넷, 또는 컴퓨터 네트워크(예: LAN 또는 WAN)와 같은 원거리 통신 네트워크)를 통하여 제2 전자 장치(320)와 통신할 수 있다.
다양한 실시 예에 따르면, 제1 메모리(313)는 제1 전자 장치(310)의 구성 요소에 관계된 명령 또는 데이터를 저장할 수 있다. 예를 들어, 제1 메모리(313)는 영상 인식과 관련된 프로그램, 알고리즘, 루틴, 및 명령어를 저장할 수 있다.
다양한 실시 예에 따르면, 이미지 센서(314)는, 영상 데이터(210)를 획득(또는 캡쳐)할 수 있다. 일 실시 예에 따르면, 영상 데이터(210)는 제1 메모리(313)에 저장될 수 있다. 더하여, 제1 전자 장치(310)가 시각적 정보를 출력하도록 구성된 출력 장치(316)를 구비하는 경우, 영상 데이터(210)는 출력 장치(316)를 통해 출력될 수도 있다.
다양한 실시 예에 따르면, 마이크(315)는 외부의 소리를 전기적인 오디오 신호로 변환하여 출력하도록 구성될 수 있다. 일 실시 예에 따르면, 마이크(315)는 발화 데이터(201)를 획득하도록 구성될 수 있다. 실시 예에 따라, 마이크(315)는 외부의 소리를 입력 받는 과정에서 발생되는 잡음(noise)을 제거하기 위한 다양한 잡음 제거 알고리즘이 적용될 수도 있다.
다양한 실시 예에 따르면, 출력 장치(316)는, 제1 전자 장치(310)의 동작과 관련된 다양한 정보를 제공할 수 있다. 다양한 정보 중 적어도 일부는 영상 인식 기능과 관련될 수 있다. 일 실시 예에 따르면, 출력 장치(316)는 시각적 정보(예: 텍스트, 이미지, 비디오, 아이콘, 또는 심볼)를 사용자에게 제공하고 사용자 입력(예: 터치 입력)을 수신하도록 구성된 디스플레이를 포함할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 실시 예에 따라, 출력 장치(316)는 청각적 정보, 촉각적 정보 또는 이들의 조합을 제공하도록 구성될 수도 있다.
다양한 실시 예에 따르면, 제1 프로세서(311)는 제1 통신 회로(312), 제1 메모리(313), 이미지 센서(314), 마이크(315) 및 출력 장치(316)와 작동적으로 연결될 수 있으며, 제1 전자 장치(310)의 다양한 구성 요소(예: 하드웨어 또는 소프트웨어 구성 요소)들을 제어할 수 있다.
일 실시 예에 따르면, 제1 프로세서(311)는 영상 데이터(210)에서 사용자가 지시하는 대상 객체를 식별하고, 식별된 대상 객체에 대한 연관 정보(250)를 제공할 수 있다. 대상 객체는 영상 데이터(210)에 포함된 다수의 객체들 중 일부이거나 전체일 수 있다.
이와 관련하여, 제1 프로세서(311)는, 발화 데이터(201)가 획득되는 동안 영상 데이터(210)를 획득할 수 있다. 예를 들어, 제1 프로세서(311)는 발화 데이터(201)로부터 식별되는 제1 부분(예: 도 2a의 지정된 제1 단어(201-1))과 제2 부분(예: 도 2a의 지정된 제2 단어(201-2))에 기반하여, 영상 데이터(210)에 포함된 객체들 중 대상 객체(예: 도 2a의 제1 대상 객체와 제2 대상 객체)를 식별할 수 있다. 더하여, 제1 프로세서(311)는 발화 데이터(201)로부터 식별되는 제3 부분(예: 도 2a의 질의에 해당되는 문장(201-3))에 기반하여, 대상 객체에 대한 연관 정보(250)를 제공할 수 있다. 이와 관련된 제1 프로세서(311)의 구체적인 설명은 도 2a 내지 도 2b를 통해 설명한 전자 장치(100)의 동작이 참조될 수 있다.
추가적으로, 다양한 실시 예에 따른 제1 프로세서(311)는 제2 전자 장치(320)를 통해 획득된 모션 데이터(341)를 대상 객체를 식별하는데 활용할 수 있다. 이와 관련하여, 제1 프로세서(311)는 발화 데이터(201)가 획득되는 동안(및/또는 영상 데이터(210)가 획득되는 동안), 제2 전자 장치(320)로부터 모션 데이터(341)를 획득할 수 있다. 예를 들어, 제2 전자 장치(320)로부터 획득되는 모션 데이터(341)는 위치 정보, 속도 정보, 가속도 정보, 방향 정보 또는 이들의 조합을 포함할 수 있다.
일 실시 예에 따르면, 제1 프로세서(311)는 발화 데이터(201)의 제1 부분이 식별된 제1 시점에, 제2 전자 장치(320)를 통해 획득된 제1 모션 데이터를 획득하고, 이를 대상 객체(예: 제1 대상 객체)를 식별하는데 활용할 수 있다. 더하여, 제1 프로세서(311)는 발화 데이터(201)의 제2 부분이 식별된 제2 시점에, 제2 전자 장치(320)를 통해 획득된 제2 모션 데이터를 획득하고, 이를 다른 대상 객체(예: 제2 대상 객체)를 식별하는데 활용할 수 있다.
예를 들어, 도 3c에 도시된 바와 같이, 제1 프로세서(311)는, 제2 전자 장치(320)로부터 획득된 모션 데이터(341)에서, 발화 데이터(201)의 제1 부분(예: 지정된 제1 단어(예: 이 아이)(201-1))이 식별된 제1 시점(t1)에 대응하는 일부(341-1)를 제1 모션 데이터로 획득(또는 추출)할 수 있다. 이와 유사하게, 제1 프로세서(311)는, 발화 데이터(201)의 제2 부분(예: 지정된 제2 단어(예: 저 아이)(201-2))이 식별된 제2 시점(t2)에 대응하는 모션 데이터(341)의 다른 일부(341-2)를 제2 모션 데이터로 획득할 수 있다.
일 실시 예에 따르면, 제1 프로세서(311)는 제1 대상 객체를 식별함에 있어서, 제1 모션 데이터가 지정된 조건을 만족하는지를 판단할 수 있다. 예를 들어, 제1 시점(t1)에서 획득된 제1 모션 데이터가 지정된 조건을 만족하면, 제1 프로세서(311)는 영상 데이터(210)에서 식별되는 신체(예: 검지 손가락)의 제1 위치에 대응되는 객체를 제1 대상 객체로 식별할 수 있다.
이와 유사하게, 제1 프로세서(311)는 제2 대상 객체를 식별함에 있어서, 제2 모션 데이터가 지정된 조건을 만족하는지를 판단할 수 있다. 예를 들어, 제2 시점에 획득된 제2 모션 데이터가 지정된 조건을 만족하면, 제1 프로세서(311)는 영상 데이터(210)에서 식별되는 신체의 제2 위치에 대응되는 객체를 제2 대상 객체로 식별할 수 있다.
전술한 바와 같이, 다양한 실시 예에 따른 제1 프로세서(311)는 제2 전자 장치(320)를 통해 획득된 모션 데이터(341)가 지정된 조건을 만족하는 경우 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
이와 관련하여, 제1 프로세서(311)는 제1 모션 데이터(및/또는 제2 모션 데이터)가 지정된 제스처에 대응되는 경우, 지정된 조건을 만족하였다고 판단할 수 있다. 예를 들어, 지정된 제스처는 사용자가 신체 일부를 이용하여 특정 객체를 지시하는 제스처(예: 포인팅 제스처)를 포함할 수 있다. 실시 예에 따라, 지정된 제스처는 제2 전자 장치(320)(예: 터치 센서)를 신체의 일부로 적어도 일부 터치한 상태에서 발생될 수도 있다.
그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 특정 객체를 지시할 수 있는 다양한 제스처(예: 검지 손가락으로 원을 그리는 제스처, 손가락으로 대상을 지칭하면서 손으로 그리는 제스처, 손 액자(finger frame)를 만드는 제스처 또는 엄지 손가락과 검지 손가락을 이용하여 원을 만드는 제스처)를 지정된 조건으로 사용될 수 있다.
추가적으로 또는 선택적으로, 제1 프로세서(311)는 발화 데이터(201)의 제1 부분이 식별된 이후 일정 시간(예: 5 초) 이내에 발화 데이터(201)의 제2 부분이 식별되는 경우, 제2 모션 데이터를 획득할 수 있다. 예를 들어, 제1 프로세서(311)는 제1 부분이 식별된 후 일정 시간이 경과된 이후에 식별되는 제2 부분은 제2 대상 객체의 식별에서 배제할 수 있다.
전술한 바와 같이, 제1 전자 장치(310)는 제2 전자 장치(320)로부터 지정된 조건을 만족하는 모션 데이터(341)가 획득되는지를 판단할 수 있다. 하지만, 이러한 경우, 모션 데이터(341)에 대한 동기화 문제 또는 모션 데이터(341)에 대한 전송 속도 문제가 발생될 수 있다.
이와 관련하여, 일 실시 예에 따르면, 제2 프로세서(321)는 제1 대상 객체를 식별함에 있어서, 제1 모션 데이터가 지정된 조건을 만족하는지를 판단할 수 있다. 제1 프로세서(311)는 제2 전자 장치(320)로부터 지정된 조건을 만족하는 모션 데이터(341)를 획득할 수도 있다. 예를 들어, 제2 전자 장치(320)는 발화 데이터(201)의 제1 부분(및/또는 제2 부분)을 식별하거나 또는 제1 전자 장치(310)로부터 발화 데이터(201)의 제1 부분(및/또는 제2 부분)을 나타내는 정보를 수신하는 시점에, 제2 전자 장치(320)의 자세와 관련된 모션 데이터(341)를 획득할 수 있다. 더하여, 제2 전자 장치(320)는 획득된 모션 데이터(341)가 지정된 조건을 만족하는 경우 모션 데이터(341)를 제1 전자 장치(310)로 제공할 수도 있다.
추가적으로 또는 선택적으로, 제2 전자 장치(320)는 지정된 조건을 만족하는 모션 데이터(341) 대신에, 모션 데이터(341)에 대한 특징 데이터를 추출하여 제1 전자 장치(310)로 제공할 수도 있다. 이러한 경우, 제1 전자 장치(310)는 제2 전자 장치(320)로부터 제공받은 특징 데이터에 기반하여, 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
실시 예에 따라, 제2 전자 장치(320)는 신체 일부(예: 손가락)(또는 제2 전자 장치(320))의 움직임 특징(motion feature)과 관련된 특징 데이터를 추출할 수 있다. 예를 들어, 특징 데이터는 움직이고 있는 신체 일부에 대응되는 특징 데이터, 멈춰있는 신체 일부에 대응되는 특징 데이터, 신체 일부가 움직이다가 멈춤에 대응되는 특징 데이터, 신체 일부가 멈춤 후 움직임에 대응되는 특징 데이터 중 적어도 하나를 포함할 수 있다.
전술한 바와 같이, 제1 프로세서(311)는 대상 객체 식별에 활용되는 모션 데이터(341)를 제2 전자 장치(320)를 통해 획득할 수 있다. 이와 관련된 예시적인 제2 전자 장치(320)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따른 제2 전자 장치(320)는 도 3b에 도시된 바와 같이, 적어도 하나의 제2 프로세서(321)(이하, 제2 프로세서(321)로 칭함), 제2 통신 회로(322)(이하, 제2 통신 회로(322)로 칭함), 적어도 하나의 제2 메모리(323)(이하, 제2 메모리(323)로 칭함) 및 적어도 하나의 센서(324)(이하, 센서(224)로 칭함)로 구성될 수 있다. 실시 예에 따라, 제2 전자 장치(320)는 전술한 구성 요소들 보다 많은 구성 요소들을 가지거나, 또는 그 보다 적은 구성 요소를 가지는 것으로 구현될 수 있다.
전술한 제2 통신 회로(322) 및 제2 메모리(323)는 전술한 제1 전자 장치(210)의 제1 통신 회로(312) 및 제1 메모리(313)와 유사하거나 동일할 수 있으며, 이에 그에 대한 구체적인 설명은 생략될 수 있다.
다양한 실시 예에 따르면, 제2 통신 회로(322)는 제1 전자 장치(310)와의 무선 통신 수행을 지원할 수 있다. 일 실시 예에 따르면, 제2 통신 회로(322)는 제1 전자 장치(310)와 제2 전자 장치(320) 사이의 신호(예: 명령 또는 데이터)를 송수신하기 위한 하드웨어 및 소프트웨어를 포함하는 장치일 수 있다.
다양한 실시 예에 따르면, 제2 메모리(323)는 제2 전자 장치(320)의 적어도 하나의 다른 구성 요소에 관계된 명령 또는 데이터를 저장할 수 있다.
다양한 실시 예에 따르면, 센서(324)는 제2 전자 장치(320)의 자세와 관련된 모션 데이터를 획득하도록 구성될 수 있다. 일 실시 예에 따르면, 센서(324)는 가속도 센서, 자이로 센서, 제스처 센서 또는 기압 센서 중 적어도 하나를 포함할 수 있다.
다양한 실시 예에 따르면, 제2 프로세서(321)는 제2 통신 회로(322), 제2 메모리(323) 및 센서(324)와 작동적으로 연결될 수 있으며, 제2 전자 장치(320)의 다양한 구성 요소(예: 하드웨어 또는 소프트웨어 구성 요소)들을 제어할 수 있다.
일 실시 예에 따르면, 제2 프로세서(321)는 센서(324)를 통해 수집되는 모션 데이터를 제1 전자 장치(310)로 제공할 수 있다. 이와 관련하여, 제2 프로세서(321)는 제1 전자 장치(310)와 제2 전자 장치(320)가 통신으로 연결된 상태에서 모션 데이터(341)가 수집되도록 센서(324)를 제어할 수 있다.
실시 예에 따라, 제2 프로세서(321)는 지정된 이벤트의 발생에 기반하여 모션 데이터(341)가 수집되도록 센서(324)를 제어할 수 있다. 예를 들어, 제2 전자 장치(320)가 사용자 입력(예: 터치 입력)을 수신하도록 구성된 경우(예: 터치 센서가 제2 전자 장치(320)에 구비된 경우), 제1 프로세서(321)는 사용자의 입력이 감지된 후(또는 감지되는 동안) 모션 데이터(341)가 수집되도록 처리할 수 있다.
전술한 바와 같이, 다양한 실시 예에 따른 제1 전자 장치(310)는 제1 전자 장치(310)에 의해 수집되는 발화 데이터(201)와 영상 데이터(210) 그리고, 제2 전자 장치(320)에 의해 수집되는 모션 데이터(314)에 기반하여, 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 대상 객체를 식별함에 있어서 레이저 포인터(pointer)와 같은 광 포인터의 빔을 스캔하는 동작을 수행할 수 있다. 이와 관련하여, 제1 전자 장치(310) 또는 제2 전자 장치(320)는 레이저 포인트를 조사하도록 구성된 발광부를 포함할 수 있으며, 제1 전자 장치(310)는 복수의 객체들 중 레이저 포인트가 향하는 방향에 대응되는 객체를 대상 객체로 식별함으로써, 객체 식별 정확도를 보다 향상시킬 수도 있다.
도 4는 다양한 실시 예에 따른 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다.
도 4에 도시된 영상 인식 방법은 복수의 특정 객체들을 한번에 지정하는 제스처를 영상 인식 방법에 활용하는 면에서, 도 2a 및 도 2b에 도시된 영상 인식 방법과 차이가 있다. 이와 관련된 다양한 실시 예에 따른 영상 인식 기능을 이하를 통해 보다 구체적으로 설명하도록 한다.
도 4를 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)는 발화 데이터(411)를 획득하고, 이에 대한 제4 부분 및 제3 부분을 식별할 수 있다.
일 실시 예에 따르면, 발화 데이터(411)의 제4 부분은 복수의 대상 객체들(예: 제1 대상 객체와 제2 대상 객체)을 한번에 지정하는 지정된 단어(예: 이것들, 저것들, 우리는, 그들은, 너희들, 그녀들 같은 복수의 대상들을 지시하는 지시 대명사)에 대응될 수 있으며, 발화 데이터(411)의 제3 부분은 지정된 질의에 대응될 수 있다.
예를 들어, 제1 전자 장치(310)는 발화 데이터(411)(예: 이 들 중 자라면 누가 제일 크지?)에 대하여, 복수의 대상 객체들을 지정하는 단어(예: 이 들)을 제4 부분으로 식별하고, 질의에 해당되는 문장(예: 자라면 누가 더 크지?)을 제3 부분으로 식별할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(310)는 발화 데이터(411)의 제4 부분에 기반하여, 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
이와 관련하여, 제1 전자 장치(310)는 발화 데이터(411)가 획득되는 동안의 영상 데이터(210)를 획득한 후, 영상 데이터(210)에 포함된 객체들(예: 제1 객체(211) 내지 제4 객체(217))을 인식(또는 추출)할 수 있다. 더하여, 제1 전자 장치(310)는 발화 데이터(411)의 제4 부분에 기반하여, 인식된 객체(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 복수의 대상 객체들을 식별할 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 제2 전자 장치(320)를 통해 획득된 모션 데이터(341)가 지정된 조건을 만족하는 경우 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
예를 들어, 제1 전자 장치(310)는, 도시된 바와 같이, 제4 부분이 식별된 제4 시점에, 영상 데이터(210)에서 식별되는 제스처(413)에 대응되는 복수의 객체들(예: 제2 객체(213)와 제3 객체(215))을 대상 객체로 식별할 수 있다. 예컨대, 제1 전자 장치(310)는 제스처(413)에 의해 그려지는 원(413-1)에 포함되는 복수의 객체들(예: 제2 객체(213)와 제3 객체(215))을 대상 객체로 식별할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(310)는, 발화 데이터(411)의 제3 부분에 기반하여, 복수의 대상 객체들에 대한 연관 정보(250)를 제공할 수 있다.
도 5a는 다양한 실시 예에 따른 제1 전자 장치의 영상 인식 기능을 설명하기 위한 도면이다. 그리고, 도 5b는 다양한 실시 예에 따른 제1 전자 장치에서 사용자의 제스처를 인식하는 동작을 설명하기 위한 도면이며, 도 5c는 다양한 실시 예에 따른 제스처에 기반하여 제어되는 제1 전자 장치의 동작을 설명하기 위한 도면이다.
도 5a에 도시된 영상 인식 방법은 복수의 제2 전자 장치들(320-1 및 320-2)을 통해 모션 데이터를 획득하는 면에서, 도 4에 도시된 영상 인식 방법과 차이가 있다. 예를 들어, 복수의 제2 전자 장치들(320-1 및 320-2) 중 하나(320-1)는 신체의 제1 부분(예: 왼 손(511-1))에 착용되고 복수의 제2 전자 장치들(320-1 및 320-2) 중 다른 하나(320-2)는 신체의 제2 부분(예: 오른 손(511-2))에 착용될 수 있다. 이와 관련된 다양한 실시 예에 따른 영상 인식 기능을 이하를 통해 보다 구체적으로 설명하도록 한다.
도 5a를 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)는 발화 데이터(513)를 획득하고, 이에 대한 제4 부분 및 제3 부분을 식별할 수 있다.
일 실시 예에 따르면, 발화 데이터(513)의 제4 부분은 복수의 대상 객체들(예: 제1 대상 객체와 제2 대상 객체)을 한번에 지정하는 지정된 단어(예: 이것들, 저것들, 우리는, 그들은, 너희들, 그녀들 같은 복수의 대상들을 지시하는 지시 대명사)에 대응될 수 있으며, 발화 데이터(513)의 제3 부분은 지정된 질의에 대응될 수 있다.
예를 들어, 제1 전자 장치(310)는 사용자의 발화 데이터(513)(예: 이 들 중 자라면 누가 제일 크지?)에 대하여, 복수의 대상 객체들을 지정하는 단어(이 들)을 제4 부분으로 식별하고, 질의(예: 누가 더 크지?)에 해당되는 문장을 제3 부분으로 식별할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(310)는 발화 데이터(411)의 제4 부분에 기반하여, 영상 데이터(210)에서 대상 객체를 식별할 수 있다
이와 관련하여, 제1 전자 장치(310)는 발화 데이터(513)가 획득되는 동안 영상의 데이터(210)를 획득한 후, 영상 데이터(210)에 포함된 객체들(예: 제1 객체(211) 내지 제4 객체(217))을 인식(또는 추출)할 수 있다. 더하여, 제1 전자 장치(310)는 발화 데이터(411)의 제4 부분에 기반하여, 인식된 객체(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 복수의 대상 객체들을 식별할 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 복수의 제2 전자 장치들(320-1 및 320-2)을 통해 획득된 모션 데이터(341)가 지정된 조건을 만족하는 경우 영상 데이터(210)에서 대상 객체를 식별할 수 있다.
예를 들어, 제1 전자 장치(310)는, 도시된 바와 같이, 제4 부분이 식별된 제4 시점에, 영상 데이터(210)에서 식별되는 제스처(예: 손 액자 제스처)에 대응되는 복수의 객체들(예: 제2 객체(213)와 제3 객체(215))을 대상 객체로 식별할 수 있다. 예를 들어, 제1 전자 장치(310)는 신체의 제1 부분(511-1)과 신체의 제2 부분(511-2)을 이용한 제스처(511)에 의해 형성되는 손 액자에 포함된 복수의 객체들(예: 제2 객체(213)와 제3 객체(215))을 대상 객체로 식별할 수 있다.
일 실시 예에 따르면, 제1 전자 장치(310)는 무선 통신(예: UWB(ultra wideband) 통신)을 통해, 복수의 제2 전자 장치들(320-1 및 320-2)과 송수신하는 소정의 신호를 제스처 식별에 활용할 수 있다. 예를 들어, 도 5b에 도시된 바와 같이, 제1 전자 장치(310)는 송수신된 소정의 신호를 기반하여 하나의 제2 전자 장치(320-1)와의 제 1 거리(B), 다른 제2 전자 장치(320-2)와의 제 2 거리(C) 그리고 복수의 제2 전자 장치들(320-1 및 320-2) 사이의 제3 거리(A)를 확인할 수 있다. 실시 예에 따라, 제1 전자 장치(310)는 복수의 제2 전자 장치들(320-1 및 320-2) 사이의 제3 거리(A)가 지정된 범위(예: 15cm 이내)에 해당되는 제스처를 식별할 수 있다.
이와 관련하여, 제1 전자 장치(310)는 레인징과 관련된 알고리즘을 통해, 복수의 제2 전자 장치들(320-1 및 320-2)과 소정의 신호를 송수신할 수 있다. 예컨대, 레인징과 관련된 알고리즘은, ToF(time of flight), TWR(two way ranging), DS-TWR(double side- TWR), SS-TWR(single sided-TWR), TDoA(time difference of arrival) 또는 AoA(angle of arrival) 중 적어도 하나를 포함할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(310)는, 발화 데이터(513)의 제3 부분(201-3)에 기반하여, 복수의 대상 객체들에 대한 연관 정보(250)를 제공할 수 있다.
전술한 바와 같이, 다양한 실시 예에 따른 제1 전자 장치(310)는 제1 전자 장치(310)에 의해 수집되는 발화 데이터(201, 411, 513), 영상 데이터(210) 그리고, 제2 전자 장치(320)에 의해 수집되는 모션 데이터(341)에 기반하여, 영상 데이터(210)로부터 대상 객체를 식별할 수 있다.
실시 예에 따라, 제1 전자 장치(310)는 대상 객체를 식별함에 있어서, 기계 학습을 통해 생성된 인공지능 모델을 활용할 수 있다. 이와 관련하여, 제1 전자 장치(310)(예: 제1 메모리(313))에는 영상 데이터(210)로부터 대상 객체를 식별하고, 대상 객체에 대한 연관 정보(예: 비교 정보)를 제공하도록 구성된 인공지능 모델(예: 정보 제공 모델(3131))이 저장될 수 있다. 이에 대한 구체적인 설명은 이하의 도 6을 통해 보다 상세히 설명하도록 한다.
추가적으로 또는 선택적으로, 다양한 실시 예에 따른 제1 전자 장치(310)는 영상 데이터(210)에서 식별되는 제스처(예: 손 액자 제스처)에 기반하여 지정된 동작을 수행할 수도 있다. 일 실시 예에 따르면, 제1 전자 장치(310)는 제스처 식별에 대응하여 비활성화 상태의 이미지 센서를 활성화할 수 있다. 예를 들어, 제1 전자 장치(310)는, 도 5c의 540에 도시된 바와 같이, 제1 전자 장치(310)(예: 이미지 센서)의 촬영 범위(field of view)에 대응되는 이미지를 획득할 수 있다. 실시 예에 따라, 제1 전자 장치(310)는, 도 5c의 550에 도시된 바와 같이, 제스처(예: 손 액자 제스처)에 의해 식별된 대상 객체들(예: 제2 객체(213)와 제3 객체(215))만 포함하는 이미지를 획득할 수도 있다.
도 6은 다양한 실시 예에 따른 정보 제공 모델의 구성을 도시한 도면이다.
다양한 실시 예에 따른 정보 제공 모델(3131)은 기계 학습을 통해 생성된 인공지능 모델일 수 있다. 실시 예에 따라, 정보 제공 모델(3131)은 모션 데이터(610)(예: 모션 데이터(341)), 영상 데이터(620)(예: 영상 데이터(210)) 및 발화 데이터(630)(예: 발화 데이터(201, 411, 513))를 입력으로 하고, 영상 데이터(620)에서 식별되는 대상 객체에 대한 연관 정보(예: 비교 정보)를 출력할 수 있다.
일 실시 예에 따르면, 정보 제공 모델(3131)은, 복수의 인공 신경망 레이어들을 포함할 수 있다. 인공 신경망은 심층 신경망(DNN: deep neural network), CNN(convolutional neural network), RNN(recurrent neural network), RBM(restricted boltzmann machine), DBN(deep belief network), BRDNN(bidirectional recurrent deep neural network), 심층 Q-네트워크(deep Q-networks), 트랜스포머 네트워크(transformer network) 또는 이들 중 둘 이상의 조합 중 하나일 수 있으나, 전술한 예에 한정되지 않는다. 정보 제공 모델(3131)은 소프트웨어 구조 이외에, 추가적으로 또는 대체적으로 하드웨어 구조를 포함할 수 있다. 이와 관련된 다양한 실시 예에 따른 정보 제공 모델(3131)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
도 6을 참조하면, 다양한 실시 예에 따른 정보 제공 모델(3131)은 객체 식별 모델(605), 정보 검색 모델(607) 및 정보 생성 모델(609)을 포함할 수 있다.
다양한 실시 예에 따르면, 객체 식별 모델(605)은 영상 데이터(620)로부터 대상 객체를 식별할 수 있다. 일 실시 예에 따르면, 객체 식별 모델(605)은 모션 데이터(610), 영상 데이터(620) 및 발화 데이터(630)를 입력으로 제공받고, 영상 데이터(620)에서 식별되는 대상 객체에 대한 식별 정보(615)를 출력할 수 있다. 예를 들어, 객체 식별 모델(605)은 발화 데이터(630)가 입력되는 동안, 영상 데이터(620)에서 식별되는 제스처에 기반하여 대상 객체를 식별할 수 있다.
다양한 실시 예에 따르면, 정보 검색 모델(607)은 영상 데이터(620)에서 식별된 대상 객체와 관련된 정보를 검색할 수 있다. 일 실시 예에 따르면, 정보 검색 모델(607)은 발화 데이터(630)와 대상 객체에 대한 정보(예: 식별 결과(615))를 입력으로 제공받고, 검색 결과(617)를 출력할 수 있다. 예를 들어, 정보 검색 모델(607)은 발화 데이터(630)와 식별 결과(615)에 기반하여 검색어를 생성하고, 이를 기반으로 하는 검색 결과(617)를 외부(예: 검색 서버)로부터 획득할 수 있다.
다양한 실시 예에 따르면, 정보 생성 모델(609)은 정보 검색 모델(607)의 검색 결과(617)를 사용자가 인식할 수 있는 형태로 변환하여 대상 객체에 대한 정보(619)(예: 연관 정보(250))로 출력할 수 있다. 일 실시 예에 따르면, 정보 생성 모델(609)은 정보 검색 모델(607)에 의해 출력되는 텍스트 형태의 검색 결과(617)를 청각적 형태로 변환할 수 있다. 실시 예에 따라, 대상 객체에 대한 정보(619)는 시각적 형태로 출력될 수도 있다.
추가적으로 또는 선택적으로, 정보 생성 모델(609)은 검색 결과(617)에 기반하여 새로운 출력 데이터(예: 이미지 데이터)를 생성하도록 구성된 생성형 모델(generative model)을 포함할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 정보 생성 모델(609)은 생성형 모델 외에 다양한 종류의 모델로 구성될 수도 있다.
실시 예에 따라, 정보 제공 모델(3131)은 전술한 구성 요소보다 적은 구성 요소를 가지는 것으로 구성될 수 있다. 예컨대, 정보 생성 모델(609)은 정보 제공 모델(3131)에서 생략될 수 있다. 이러한 경우, 정보 검색 모델(607)에 의해 획득된 검색 결과(617)가 대상 객체에 대한 정보(619)로 제공될 수 있다.
실시 예에 따라, 정보 제공 모델(3131)은 전술한 구성 요소보다 더 많은 구성 요소를 가지는 것으로 구성될 수 있다. 예컨대, 패턴 인식 모델(603)이 정보 제공 모델(3131)에 포함될 수 있다. 이러한 경우, 객체 식별 모델(605)은 패턴 인식 모델(603)에 의해 출력되는 모션 데이터(610)의 패턴 정보(613)에 기반하여, 사용자의 제스처를 식별할 수 있다. 이와 관련된 다양한 실시 예에 따른 패턴 인식 모델(603)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따르면, 패턴 인식 모델(603)은 모션 데이터(610)에 기반하여 패턴 정보(613)를 식별할 수 있다. 패턴 정보는 모션의 반복성, 모션의 방향, 모션의 속도 또는 모션의 크기와 관련될 수 있다. 일 실시 예에 따르면, 패턴 인식 모델(603)은 모션 데이터(610)를 입력을 제공받고 모션 데이터(610) 속의 특정 패턴을 식별할 수 있다. 더하여, 패턴 인식 모델(603)은 식별된 특정 패턴과 관련된 패턴 정보(613)를 출력할 수 있다. 이러한 패턴 정보(613)는 객체 식별 모델(605)로 입력되며, 객체 식별 모델(605)은 패턴 정보(613)를 사용자의 제스처 식별에 활용할 수 있다.
추가적으로 또는 선택적으로, 패턴 인식 모델(603)은 모션 데이터(610)로부터 추출된 특징 데이터(611)에 기반하여, 모션 데이터(610) 속의 특정 패턴을 식별할 수 있다. 이러한 경우, 정보 제공 모델(3131)은 특징 추출 모델(601)을 더 포함할 수도 있다. 이와 관련된 다양한 실시 예에 따른 특징 추출 모델(601)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따르면, 특징 추출 모델(601)은 모션 데이터(610)로부터 특징 데이터(611)를 추출할 수 있다. 특징 데이터(611)는 모션 데이터(610) 속의 특정 패턴을 식별하는데 활용될 수 있는 데이터일 수 있다. 일 실시 예에 따르면, 특징 추출 모델(601)은 모션 데이터(610)로부터 신체 일부(예: 손가락)의 외형 특징(appearance feature)과 움직임 특징(motion feature)과 관련된 특징 데이터(611)를 추출할 수 있다. 이러한 특징 데이터(611)는 패턴 인식 모델(603)로 입력되며, 패턴 인식 모델(603)은 특징 데이터(611)를 패턴 식별에 활용할 수 있다.
도 7은 다양한 실시 예에 따른 객체 식별 모델의 입력 데이터를 처리하는 절차를 설명하기 위한 도면이다.
다양한 실시 예에 따르면, 객체 식별 모델(605)은 영상 데이터(620)와 발화 데이터(630)를 입력으로 제공받고, 영상 데이터(620)에서 대상 객체를 식별할 수 있다.
실시 예에 따라, 정보 제공 모델(3131)은 영상 데이터(620)와 발화 데이터(630)를 병합(concatenation)하여 객체 식별 모델(605)로 입력할 수 있다.
도 7을 참조하면, 정보 제공 모델(3131)은 영상 데이터(620)와 발화 데이터(630)를 병합(650)함에 있어서 1차원의 데이터를 2차원의 데이터로 변환하는 동작을 수행할 수 있다.
예를 들어, 정보 제공 모델(3131)은 영상 데이터(620)를 필터링 및 샘플링과 같은 방식으로 전처리(621)하여 정제된 영상 데이터를 획득하고, 이를 시간과 주파수의 관계로 표현되는 2차원의 영상 데이터로 변환(623)할 수 있다.
이와 유사하게, 정보 제공 모델(3131)은 발화 입력(630)을 필터링 및 샘플링과 같은 방식으로 전처리(631)하여 정제된 발화 입력을 획득하고, 이를 시간과 주파수의 관계로 표현되는 2차원의 발화 데이터로 변환(633)할 수 있다.
이러한 2차원의 영상 데이터(623) 및 발화 데이터(633)는 병합(650)되어 객체 식별 모델(605)의 입력으로 제공되며, 객체 식별 모델(605)은 병합된 영상 데이터(623) 및 발화 데이터(633)에 기반하여 대상 객체에 대한 식별 정보(615)를 출력할 수 있다.
그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 전술한 바와 같이, 대상 객체를 식별함에 있어서, 영상 데이터(620)와 발화 데이터(630) 외에 모션 데이터(610)(또는 특징 데이터(611) 또는 패턴 정보(613))가 더 활용될 수 있다. 이와 관련하여, 정보 제공 모델(3131)은 모션 데이터(610)를 2차원의 모션 데이터로 변환하고, 이를 2차원의 영상 데이터(623) 및 발화 데이터(633)와 병합하여 객체 식별 모델(605)로 출력할 수도 있다.
도 8a는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면들이다.
도 8a를 참조하면, 다양한 실시 예에 따른 영상 인식 시스템(81)은 제1 전자 장치(810), 제2 전자 장치(820) 및 제3 전자 장치(830)로 구성될 수 있다. 일 실시 예에 따르면, 제1 전자 장치(810)는 제2 전자 장치(820) 및 제3 전자 장치(830)와 네트워크(예: 근거리 통신 네트워크 또는 원거리 통신 네트워크)를 통하여 서로 통신할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(810)는 제2 전자 장치(820) 및 제3 전자 장치(830)와의 협업을 통해 영상 인식 기능을 제공할 수 있다.
일 실시 예에 따르면, 제1 전자 장치(810)는 영상 데이터(210)를 획득(또는 수집)하고, 제2 전자 장치(820)는 신체에 착용된 상태에서 제2 전자 장치(820)의 자세와 관련된 모션 데이터(341)를 획득하고, 제3 전자 장치(830)은 발화 데이터(201)를 획득할 수 있다. 또한, 제1 전자 장치(810)는 제1 전자 장치(810)에 의해 획득되는 영상 데이터(210), 제2 전자 장치(820)에 의해 획득되는 모션 데이터(341) 그리고, 제3 전자 장치(830)에 의해 획득되는 발화 데이터(201)에 기반하여, 영상 데이터(210)에 포함된 다수의 객체들(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 대상 객체를 식별할 수 있다.
이와 관련된 다양한 실시 예에 따른 제1 전자 장치(810)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따른 제1 전자 장치(810)는 제1 프로세서(811), 제1 통신 회로(812), 제1 메모리(813), 이미지 센서(814) 및 출력 장치(816)로 구성될 수 있다.
실시 예에 따라, 도 8a에 도시된 제1 전자 장치(810)의 구성들은 도 3b를 통해 전술한 제1 전자 장치(320)의 구성과 유사하거나 동일할 수 있으며, 이에 그에 대한 구체적인 설명은 생략될 수 있다.
다양한 실시 예에 따르면, 제1 프로세서(811)는 이미지 센서(814)를 통해 획득되는 영상 데이터(210)에서 사용자가 지시하는 대상 객체를 식별할 수 있다. 예를 들어, 제1 프로세서(811)는 제2 전자 장치(820)에 의해 획득되는 모션 데이터(341) 그리고, 제3 전자 장치(830)에 의해 획득되는 발화 데이터(201)를 대상 객체 식별에 활용할 수 있다.
이와 관련된 다양한 실시 예에 따른 제2 전자 장치(820)의 구성과 제3 전자 장치(830)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따른 제2 전자 장치(820)는 제2 프로세서(821), 제2 통신 회로(822), 제2 메모리(823) 및 센서(824)로 구성될 수 있다.
실시 예에 따라, 도 8a에 도시된 제2 전자 장치(820)의 구성들은 도 3b를 통해 전술한 제2 전자 장치(320)의 구성과 유사하거나 동일할 수 있으며, 이에 그에 대한 구체적인 설명은 생략될 수 있다.
일 실시 예에 따르면, 제2 프로세서(821)는 센서(824)를 통해 수집되는 모션 데이터(341)를 제1 전자 장치(810)로 제공할 수 있다.
다양한 실시 예에 따른 제3 전자 장치(830)는 제3 프로세서(831), 제3 통신 회로(832), 제3 메모리(833) 및 마이크(834)로 구성될 수 있다.
실시 예에 따라, 도 8a에 도시된 제3 전자 장치(830)의 구성들은 도 3b를 통해 전술한 제1 전자 장치(810) 또는 제2 전자 장치(820)의 구성과 유사하거나 동일할 수 있으며, 이에 그에 대한 구체적인 설명은 생략될 수 있다.
일 실시 예에 따르면, 제3 프로세서(831)는 마이크(834)를 통해 수집되는 발화 데이터(201)를 제1 전자 장치(810)로 제공할 수 있다.
도 8b는 다양한 실시 예에 따른 영상 인식 시스템을 도시한 도면들이다.
도 8b를 참조하면, 다양한 실시 예에 따른 영상 인식 시스템(82)은 제1 전자 장치(810), 제2 전자 장치(820), 제3 전자 장치(830) 및 제4 전자 장치(840)로 구성될 수 있다. 일 실시 예에 따르면, 제1 전자 장치(810)는 제2 전자 장치(820), 제3 전자 장치(830) 및 제4 전자 장치(840)와 네트워크(예: 근거리 통신 네트워크 또는 원거리 통신 네트워크)를 통하여 서로 통신할 수 있다.
다양한 실시 예에 따르면, 제1 전자 장치(810)는 제2 전자 장치(820), 제3 전자 장치(830) 및 제4 전자 장치(840)와의 협업을 통해 영상 인식 기능을 제공할 수 있다.
일 실시 예에 따르면, 제1 전자 장치(810)는 제2 전자 장치(820)에 의해 획득되는 모션 데이터(341), 제3 전자 장치(830)에 의해 획득되는 발화 데이터(201) 및 제4 전자 장치(840)에 의해 획득되는 영상 데이터(210)에 기반하여, 영상 데이터(210)에 포함된 다수의 객체들(예: 제1 객체(211) 내지 제4 객체(217)) 중 사용자가 지시하는 대상 객체를 식별할 수 있다.
이와 관련된 다양한 실시 예에 따른 제1 전자 장치(810)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따른 제1 전자 장치(810)는 제1 프로세서(811), 제1 통신 회로(812), 제1 메모리(813) 및 출력 장치(816)로 구성될 수 있다.
일 실시 예에 따르면, 제1 프로세서(811)는 제2 전자 장치(820)에 의해 획득되는 모션 데이터(341), 제3 전자 장치(830)에 의해 획득되는 발화 데이터(201) 및 제4 전자 장치(840)에 의해 획득되는 영상 데이터(210)를 대상 객체 식별에 활용할 수 있다.
이와 관련된 다양한 실시 예에 따른 제2 전자 장치(820) 내지 제4 전자 장치(840)의 구성을 이하를 통해 보다 구체적으로 설명하도록 한다.
다양한 실시 예에 따른 제2 전자 장치(820)는 제2 프로세서(821), 제2 통신 회로(822), 제2 메모리(823) 및 센서(824)로 구성될 수 있다.
일 실시 예에 따르면, 제2 프로세서(821)는 센서(824)를 통해 수집되는 모션 데이터(341)를 제1 전자 장치(810)로 제공할 수 있다.
다양한 실시 예에 따른 제3 전자 장치(830)는 제3 프로세서(831), 제3 통신 회로(832), 제3 메모리(833) 및 마이크(834)로 구성될 수 있다.
일 실시 예에 따르면, 적어도 하나의 제3 프로세서(831)는 마이크(834)를 통해 수집되는 발화 데이터(201)를 제1 전자 장치(810)로 제공할 수 있다.
다양한 실시 예에 따른 제4 전자 장치(840)는 제4 프로세서(841), 제4 통신 회로(842), 제4 메모리(843) 및 이미지 센서(844)로 구성될 수 있다.
일 실시 예에 따르면, 제4 프로세서(841)는 이미지 센서(844)를 통해 수집되는 영상 데이터(210)를 제1 전자 장치(810)로 제공할 수 있다.
전술한 도 8a 및 도 8b를 통해 설명한 영상 인식 시스템(81, 82)은 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 다양한 실시 예에 따른 영상 인식 시스템은 영상 데이터(210)를 획득하도록 구성된 제1 전자 장치(810)와 모션 데이터(341)와 발화 데이터(201)를 수집하도록 구성된 제2 전자 장치(820)로 구성될 수도 있다. 또한, 다양한 실시 예에 따른 영상 인식 시스템은 발화 데이터(201)를 획득하도록 구성된 제1 전자 장치(810)와 모션 데이터(341)와 영상 데이터(210)를 수집하도록 구성된 제2 전자 장치(820)로 구성될 수도 있다.
다양한 실시 예에 따른 전자 장치(310)는, 적어도 하나의 프로세서(311), 카메라(예: 이미지 센서)(314), 마이크(315) 및 상기 적어도 하나의 프로세서, 상기 카메라 및 상기 마이크와 동작적으로 연결되며 적어도 하나의 명령어를 저장하는 메모리(313)를 포함할 수 있다. 일 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터(201)가 상기 마이크(315)를 통해 획득되는 동안, 상기 카메라(314)를 통해 영상 데이터(210)를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들(211 내지 217) 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보(250)를 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 발화 데이터는, 질의(201-3)를 더 포함할 수 있다. 일 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 질의와 관련된 상기 연관 정보를 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 적어도 두 개의 객체와 상기 질의의 적어도 일부에 기반하여 검색어를 생성하고, 상기 생성된 검색어와 관련된 상기 연관 정보를 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)는 통신 회로(312)를 더 포함할 수 있다. 일 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 통신 회로(312)를 통해 외부 장치(320)로부터 상기 연관 정보를 획득하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 지정된 단어는, 복수의 대상들을 지정하는 제1 단어를 포함할 수 있다. 일 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제1 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제1 제스처(413, 511)에 대응되는 제1 객체와 제2 객체를 식별하고, 상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 지정된 단어는, 단일의 대상을 지정하는 제2 단어(201-1) 및 제3 단어(201-2)를 포함할 수 있다. 일 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하고, 상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 제2 단어가 발화된 이후 일정 시간 내에 상기 제3 단어가 발화되면, 상기 영상 데이터에서 제2 객체를 식별하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 카메라(314)를 통해 서로 다른 시야에 해당되는 제1 영상 데이터(260)와 제2 영상 데이터(270)를 획득하고, 상기 제1 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 제1 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하고, 상기 제2 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 제2 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 발화 데이터가 획득되는 동안, 상기 통신 회로(312)를 통해 외부 장치(320)로부터 모션 데이터(341)를 획득하고, 상기 모션 데이터가 지정된 조건을 만족하면, 상기 적어도 두 개의 객체를 식별하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금: 상기 발화 데이터와 상기 영상 데이터를 상기 전자 장치(310)(예: 메모리(313))에 저장된 인공 지능 모델(3131)로 입력하고, 상기 인공 지능 모델의 출력에 기반하여 상기 적어도 두 개의 객체를 식별하도록 구성될 수 있다.
도 9a는 다양한 실시 예에 따른 전자 장치의 동작을 도시한 흐름도이다. 그리고, 도 9b는 다양한 실시 예에 따른 연관 정보를 설명하기 위한 도면이다. 실시 예에 따라, 이하의 실시 예에서의 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다. 또한, 전술한 동작들 중 적어도 하나의 동작은 실시 예에 따라 생략될 수도 있다.
일 실시 예에 따르면, 동작 910 내지 950은 전자 장치(예: 도 1의 전자 장치(100))의 프로세서(예: 도 1의 프로세서(110))에서 수행되는 것으로 이해될 수 있다.
도 9a를 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 910에서, 영상 데이터(210)를 획득할 수 있다. 일 실시 예에 따르면, 영상 데이터(210)는 전자 장치(100)(예: 제1 전자 장치(310))의 이미지 센서(예: 이미지 센서(314))를 통해 획득되거나 외부(예: 제2 전자 장치(320))로부터 획득될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 920에서, 모션 데이터(341)를 획득할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 영상 데이터(210)가 획득되는 동안 모션 데이터(341)를 획득할 수 있다. 예컨대, 모션 데이터(341)는 외부(예: 제2 전자 장치(320))로부터 획득될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 930에서, 발화 입력(또는 발화 데이터)(201)을 획득할 수 있다. 일 실시 예에 따르면, 발화 입력(201)은 전자 장치(100)의 마이크(예: 마이크(315))를 통해 획득되거나 외부로부터 획득될 수 있다. 발화는 객체에 대한 연관 정보(예: 비교 정보)를 요청하는 발화일 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 발화는 객체 각각에 대한 정보를 요청하는 발화일 수 있으며, 실시 예에 따라, 객체에 대한 촬영과 같은 다양한 요청을 지시하는 발화일 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 940에서, 발화 입력(201) 및 모션 데이터(341)에 기반하여, 영상 데이터(210)에 포함된 적어도 두 개의 객체(예: 두 개의 대상 객체)를 식별할 수 있다.
일 실시 예에 따르면, 전자 장치(100)는 발화 입력이 획득되는 동안 지정된 조건을 만족하는 모션 데이터가 획득되면, 영상 데이터(210)에서 식별되는 신체 일부(예: 검지 손가락)의 위치에 기반하여 대상 객체를 식별할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 950에서, 발화 입력에 대응되는, 대상 객체에 대한 연관 정보를 제공할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 발화 입력(201)과 식별된 대상 객체에 기반하여 검색어를 생성하고, 이를 기반으로 하는 검색 결과를 외부(예: 검색 서버)로부터 획득할 수 있다. 예를 들어, 연관 정보는 도 2b를 통해 전술한 시각적 형태의 연관 정보 또는 도 2f를 통해 전술한 청각적 형태의 연관 정보 중 적어도 하나를 포함할 수 있다.
실시 예에 따라, 전자 장치(100)는 도 9b에 도시된 바와 같이, 복수의 객체들(971, 972)을 지시하는 제스처와 복수의 객체들(971, 972)에 대한 비교를 요청하는 사용자 발화(예: 두 객체 중 저렴한 객체에 대한 구매처를 알려줘)를 획득하는 경우(970), 복수의 객체들(971, 972)에 가격 비교 결과와 구매처 관련 정보(예: 판매자 홈페이지 주소)를 제공할 수도 있다(980).일 실시 예에 따르면, 복수의 객체들(971, 972)을 지시하는 제스처는 서로 동일하거나 또는 상이할 수 있다.
예를 들어, 하나의 객체(971)는 제1 제스처(예: 손가락으로 원을 그리는 제스처, 손가락으로 대상을 지칭하면서 손으로 그리는 제스처, 손 액자(finger frame)를 만드는 제스처 또는 엄지 손가락과 검지 손가락을 이용하여 원을 만드는 제스처 중 하나)에 의해 지정되고, 다른 하나의 객체(972)도 제1 제스처에 의해 지정될 수 있다.
실시 예에 따라, 하나의 객체(971)는 제1 제스처(예: 손가락으로 원을 그리는 제스처, 손가락으로 대상을 지칭하면서 손으로 그리는 제스처, 손 액자(finger frame)를 만드는 제스처 또는 엄지 손가락과 검지 손가락을 이용하여 원을 만드는 제스처 중 하나)에 의해 지정되고, 다른 하나의 객체(972)는 제1 제스처와 다른 제2 제스처(예: 손가락으로 원을 그리는 제스처, 손가락으로 대상을 지칭하면서 손으로 그리는 제스처, 손 액자(finger frame)를 만드는 제스처 또는 엄지 손가락과 검지 손가락을 이용하여 원을 만드는 제스처 중 다른 하나)에 의해 지정될 수 있다.
도 10은 다양한 실시 예에 따른 전자 장치의 모션 데이터 획득 동작을 도시한 흐름도이다. 이하에서 설명되는 도 10의 동작들은, 도 9의 동작 910에 대한 다양한 실시 예를 나타낸 것일 수 있다.
도 10을 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1010에서, 모션 데이터(341)를 획득할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 상대적으로 전력 소모가 많은 이미지 센서(예: 이미지 센서(314))의 동작이 비활성화된 상태에서 모션 데이터(341)를 획득할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1020에서, 지정된 제스처에 대응되는 모션 데이터(341)가 획득되는지 판단할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 전자 장치(100)와 통신으로 연결된 다른 전자 장치(예: 제2 전자 장치(320))에 의해 수집되는 모션 데이터(341)가 지정된 제스처에 대응되는지 판단할 수 있다. 예를 들어, 지정된 제스처는 다른 전자 장치가 착용된 신체를 이용하여 지정된 형상을 그리는 제스처(예: 원을 일정 횟수를 반복해서 그리는 제스처)일 수 있다. 실시 예에 따라, 지정된 제스처는 신체의 다른 일부를 이용하여 신체에 착용된 다른 전자 장치에 대해 지정된 입력(예: 터치 입력)을 발생시키는 제스처일 수도 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 지정된 제스처에 대응되는 모션 데이터(341)가 획득되지 않으면, 동작 1010 및 동작 1020를 반복 수행할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 지정된 제스처에 대응되는 모션 데이터(341)가 획득되면, 동작 1030에서, 이미지 센서를 활성화할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 상대적으로 전력 소모가 적은 센서(예: 모션 센서)를 이용하여 이미지 센서(314)의 동작을 제어함으로써 전자 장치(100)의 전력 소모를 감소시킬 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1040에서, 영상 데이터(210)를 획득할 수 있다. 도 9의 동작 910에서 설명한 바와 같이, 영상 데이터(210)는 전자 장치(100)의 이미지 센서를 통해 획득되거나 외부(예: 제2 전자 장치(320))로부터 획득될 수 있다.
도 11은 다양한 실시 예에 따른 전자 장치의 대상 객체 식별 동작을 도시한 흐름도이다. 이하에서 설명되는 도 10의 동작들은, 도 9의 동작 940에 대한 다양한 실시 예를 나타낸 것일 수 있다.
도 11을 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1110에서, 발화 입력(201)의 제1 부분에 대응되는 제1 모션 데이터를 확인할 수 있다.
일 실시 예에 따르면, 발화 입력(201)의 제1 부분은, 도 2a를 통해 전술한 바와 같이, 제1 대상 객체를 지정하는 지정된 단어(예: 이것, 저것, 이곳, 저곳, 그는, 그녀는, 너는 같은 단일의 대상을 지시하는 지시 대명사)에 대응될 수 있다.
이와 관련하여, 전자 장치(100)는, 도 3c를 통해 전술한 바와 같이, 획득된 모션 데이터(341)에서, 발화 데이터(201)의 제1 부분이 식별된 제1 시점(t1)에 대응하는 일부 모션 데이터를 제1 모션 데이터로 획득할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1120에서, 발화 입력의 제2 부분에 대응되는 제2 모션 데이터를 확인할 수 있다.
일 실시 예에 따르면, 발화 입력(201)의 제2 부분은, 도 2a를 통해 전술한 바와 같이, 제2 대상 객체를 지정하는 지정된 단어에 대응될 수 있다.
이와 관련하여, 전자 장치(100)는, 도 3c를 통해 전술한 바와 같이, 획득된 모션 데이터(341)에서, 발화 데이터(201)의 제2 부분이 식별된 제2 시점(t2)에 대응하는 일부 모션 데이터를 제2 모션 데이터로 획득할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1130에서, 영상 데이터(210)에 포함된 복수의 객체들 중, 제1 모션 데이터에 대응되는 제1 객체(예: 제1 대상 객체)를 식별할 수 있다.
일 실시 예에 따르면, 전자 장치(100)는 제1 객체를 식별함에 있어서, 제1 모션 데이터가 지정된 조건을 만족하는지를 판단할 수 있다. 예를 들어, 제1 시점(t1)에서 획득된 제1 모션 데이터가 지정된 조건을 만족하면, 전자 장치(100)는 영상 데이터(210)에서 식별되는 신체(예: 검지 손가락)의 위치에 대응되는 객체를 제1 대상 객체로 식별할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1140에서, 영상 데이터(210)에 포함된 복수의 객체들 중, 제2 모션 데이터에 대응되는 제2 객체(예: 제2 대상 객체)를 식별할 수 있다.
일 실시 예에 따르면, 전자 장치(100)는 제2 객체를 식별함에 있어서, 제2 모션 데이터가 지정된 조건을 만족하는지를 판단할 수 있다. 예를 들어, 제2 시점(t2)에서 획득된 제2 모션 데이터가 지정된 조건을 만족하면, 전자 장치(100)는 영상 데이터(210)에서 식별되는 신체(예: 검지 손가락)의 위치에 대응되는 객체를 제2 대상 객체로 식별할 수 있다.
전술한 바와 같이, 다양한 실시 예에 따른 전자 장치(100)는 영상 데이터(210)에서 복수의 특정 객체들을 순차적으로 가리키는 사용자의 제스처를 인식함으로써 대상 객체를 식별할 수 있다. 하지만, 사용자의 시야에 대응되는 영상 데이터(210)가 획득되지 않으면, 전자 장치(100)에서 식별된 대상 객체와 사용자가 지시한 객체가 서로 일치되지 않을 수 있다. 이와 관련하여, 다양한 실시 예에 따른 전자 장치(100)는 영상 데이터(210)의 획득 범위를 사용자의 시야에 일치시킴으로써, 사용자의 시야에 대응되는 영상 데이터(210)를 획득하고 대상 객체에 대한 식별 성능을 향상시킬 수 있다. 이와 관련하여는 이하의 도 12 및 13을 통해 보다 상세히 설명하도록 한다.
도 12는 다양한 실시 예에 따른 전자 장치의 시야 보정 동작을 도시한 흐름도이다. 그리고, 도 13은 다양한 실시 예에 따른 시야 보정 과정을 설명하기 위한 도면이다. 또한, 이하의 실시 예에서의 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다. 또한, 전술한 동작들 중 적어도 하나의 동작은 실시 예에 따라 생략될 수도 있다.
일 실시 예에 따르면, 동작 1210 내지 1250은 전자 장치(예: 도 1의 전자 장치(100))의 프로세서(예: 도 1의 프로세서(110))에서 수행되는 것으로 이해될 수 있다.
도 12 및 도 13을 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1210에서, 영상 데이터(1301)를 획득할 수 있다. 일 실시 예에 따르면, 영상 데이터(1301)는 전자 장치(100)의 이미지 센서(예: 이미지 센서(314))를 통해 획득되거나 외부(예: 제2 전자 장치(320))로부터 획득될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1220에서, 영상 데이터(1301)에서 기준 객체를 선택할 수 있다. 기준 객체는 영상 데이터(1301)에 포함된 객체 중 적어도 하나를 포함할 수 있다.
일 실시 예에 따르면, 전자 장치(100)는 영상 데이터(1301)로부터 추출되는 특징 데이터(예: 특징점)에 기반하여 제1 객체(예: 소파)(1303), 제2 객체(예: 테이블)(1305), 제3 객체(예: 꽃병)(1307) 및 제4 객체(예: 액자)(1309)를 인식하고, 인식된 객체 중 적어도 하나(예: 제4 객체(1309))를 기준 객체로 선택할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1230에서, 기준 객체의 선택을 유도하는 가이드 정보를 출력할 수 있다. 예를 들어, 전자 장치(100)는 사용자가 기준 객체를 가리키도록 유도하는 청각적 형태의 가이드 정보(예: 액자를 가리키세요, 액자의 우측 모서리를 가리키세요)를 출력할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예를 들어, 전자 장치(100)는 시각적 형태의 가이드 정보를 제공할 수도 있으며, 실시 예에 따라, 청각적 형태의 가이드 정보와 시각적 형태의 가이드 정보를 함께 제공할 수도 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1240에서, 기준 객체 선택과 관련된 제스처에 대응하는 제1 위치(1321)를 영상 데이터(1301)에서 식별할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1250에서, 제1 위치(1321)와 기준 객체에 대응되는 제2 위치에 기반하여, 이미지 센서(314)의 시야를 보정할 수 있다. 예를 들어, 전자 장치(100)는 제1 위치와 제2 위치의 거리 및 방향 차이(1323)에 기반하여 이미지 센서의 시야를 보정할 수 있다.
추가적으로 또는 선택적으로, 다양한 실시 예에 따른 전자 장치(100)는, 이미지 센서의 시야 보정 성능을 향상시키기 위하여, 영상 데이터(1301)로부터 인식되는 복수의 객체(예: 제1 객체(예: 소파)(1303), 제2 객체(예: 테이블)(1305), 제3 객체(예: 꽃병)(1307) 및 제4 객체(예: 액자)(1309))들 중 지정된 일정 크기보다 작은 크기를 가지는 객체(예: 제3 객체(예: 꽃병)(1307))를 기준 객체로 선택할 수도 있다.
도 14는 다양한 실시 예에 따른 전자 장치에서 연관 정보를 제공하는 동작을 도시한 흐름도이다. 이하에서 설명되는 도 14의 동작들은, 도 9의 동작 950에 대한 다양한 실시 예를 나타낸 것일 수 있다.
도 14를 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1410에서, 발화 입력(201)의 제3 부분을 확인할 수 있다.
일 실시 예에 따르면, 발화 입력(201)의 제3 부분은, 도 2a를 통해 전술한 바와 같이, 지정된 질의에 대응될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1420 및 동작 1430에서, 검색 동작을 수해할 수 있다.
일 실시 예에 따르면, 전자 장치는 발화 데이터(201)의 제3 부분, 제1 대상 객체 및 제2 대상 객체를 검색어로 하여 검색 동작을 수행할 수 있다(동작 1420). 이와 관련하여, 전자 장치(100)는 발화 데이터(201)의 제3 부분, 제1 대상 객체 및 제2 대상 객체의 적어도 일부에 기반하여 검색어를 생성할 수 있다. 더하여, 전자 장치(100)는 생성된 검색어를 기반으로 하는 검색 결과를 외부(예: 검색 서버)로부터 획득할 수 있다(동작 1430).
도 15는 다양한 실시 예에 따른 전자 장치의 동작을 도시한 다른 흐름도이다. 그리고, 이하의 실시 예에서의 각 동작들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 동작들의 순서가 변경될 수도 있으며, 적어도 두 동작들이 병렬적으로 수행될 수도 있다. 또한, 전술한 동작들 중 적어도 하나의 동작은 실시 예에 따라 생략될 수도 있다.
일 실시 예에 따르면, 동작 1510 내지 1580은 전자 장치(예: 도 1의 전자 장치(100))의 프로세서(예: 도 1의 프로세서(110))에서 수행되는 것으로 이해될 수 있다.
도 15를 참조하면, 다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1510에서, 영상 데이터(210)는 전자 장치(100)의 이미지 센서(예: 이미지 센서(314))를 통해 획득되거나 외부(예: 제2 전자 장치(320))로부터 획득될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1520에서, 지정된 제1 제스처를 감지할 수 있다. 일 실시 예에 따르면, 지정된 제1 제스처는 사용자가 검지 손가락(index finger)으로 특정 객체를 지시하는 제스처일 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1530에서, 객체 지정을 위한 제1 발화가 감지되는지 판단할 수 있다. 제1 발화는 객체를 지정하는 지정된 단어(예: 이것, 저것, 이곳, 저곳, 그는, 그녀는, 너는 같은 단일의 대상을 지시하는 지시 대명사)에 대응될 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 제1 발화가 감지되지 않으면, 동작 1510 내지 동작 1530과 관련된 동작을 반복 수행할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 제1 발화가 감지되면, 동작 1540에서, 영상 데이터(210)에서 제1 객체를 식별할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 영상 데이터(210)에 포함된 복수의 객체들 중, 제1 제스처에 대응되는 제1 객체를 식별할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1550에서, 지정된 제2 제스처를 감지할 수 있다. 지정된 제2 제스처는 전술한 제1 제스처와 유사할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 제1 제스처를 감지한 후 다른 객체를 지시하는 제2 제스처를 감지할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1560에서, 객체 지정을 위한 제2 발화가 감지되는지 판단할 수 있다. 제2 발화는 전술한 제1 발화와 유사할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 제1 발화를 감지한 후 다른 객체를 지시하는 제2 발화를 감지할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 제2 발화가 감지되지 않으면, 동작 1510 내지 동작 1560과 관련된 동작을 반복 수행할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 제2 발화가 감지되면, 동작 1570에서, 영상 데이터(210)에서 제2 객체를 식별할 수 있다. 일 실시 예에 따르면, 전자 장치(100)는 영상 데이터(210)에 포함된 복수의 객체들 중, 제2 제스처에 대응되는 제2 객체를 식별할 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1580에서, 정보 요청 발화를 감지할 수 있다. 예를 들어, 정보 요청 발화는 제1 객체와 제2 객체에 대한 연관 정보(예: 비교 정보)를 요청하는 발화일 수 있다. 예를 들어, 정보 요청 발화는 제1 객체와 제2 객체 각각에 대한 정보를 요청하는 발화일 수 있다.
다양한 실시 예에 따른 전자 장치(100)(예: 제1 전자 장치(310))는, 동작 1590에서, 정보 요청 발화에 기반하여, 제1 객체와 제2 객체에 대한 연관 정보를 제공할 수 있다.
일 실시 예에 따르면, 제1 객체와 제2 객체에 대한 연관 정보를 요청하는 정보 요청 발화가 감지되면, 전자 장치(100)는 제1 객체와 제2 객체에 대한 연관 정보를 외부로부터 획득할 수 있다.
일 실시 예에 따르면, 제1 객체와 제2 객체 각각에 대한 정보를 요청하는 정보 요청 발화가 감지되면, 전자 장치(100)는 제1 객체에 대한 정보와 제2 객체에 대한 정보를 외부로 부터 획득할 수 있다.
도 16a 내지 도 16c는 다양한 실시 예에 제1 전자 장치의 동작을 설명하기 위한 도면이다.
도 16a의 1610을 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)(예: 전자 장치(100))는, 영상 데이터(1601)가 획득되는 동안 제1 제스처(1603)와 제1 발화(1605)를 획득할 수 있다. 예를 들어, 제1 제스처(1603)는 영상 데이터(1601) 내에서 식별되는 제스처로, 영상 데이터(1601)에 포함된 적어도 하나의 특정 객체를 지시하는 제스처(예: 손 액자(finger frame)를 만드는 제스처)일 수 있다. 예를 들어, 제1 발화(1605)는 영상 데이터(1601) 내의 적어도 일부에 대한 처리(예: 촬영 또는 저장)를 지시하는 발화(예: 여기 촬영하고)일 수 있다.
도 16a의 1620을 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)(예: 전자 장치(100))는, 제1 제스처(1603)와 제1 발화(1605)가 획득된 후, 제2 제스처(1623)와 제2 발화(1625)를 획득할 수 있다. 예를 들어, 제2 제스처(1623)는 영상 데이터(1601) 내에서 식별되는 다른 제스처로, 영상 데이터(1601)에 포함된 적어도 하나의 다른 특정 객체를 지시하는 제스처(예: 손가락으로 대상을 지칭하는 제스처)일 수 있다. 예를 들어, 제2 발화는 영상 데이터 내의 적어도 일부에 대한 다른 처리(예: 검색)를 지시하는 발화(예: 이거 뭐야? 한국어로 번역해서 앨범 이름으로 만들어서 촬영한 사진 저장해줘)일 수 있다. 일 실시 예에 따르면, 제1 전자 장치(310)는 제1 제스처(1603)와 제1 발화(1605)가 획득된 후 일정 시간 내에, 제2 제스처(1623)와 제2 발화(1625)를 획득할 수 있다.
도 16b의 1630을 참조하면, 다양한 실시 예에 따른 제1 전자 장치(310)(예: 전자 장치(100))는, 제1 제스처(1603) 및 제1 발화(1605)에 기반하여 영상 데이터(1601)에 대한 제1 처리를 수행할 수 있다. 일 실시 예에 따르면, 제1 전자 장치(310)는 제1 제스처(1603)에 기반하여 영상 데이터(1601)를 전체를 제1 처리 대상으로 지정하고, 제1 발화(1605)에 기반하여 제1 처리 대상으로 지정된 영상 데이터(1601)를 저장(1633)할 수 있다.
추가적으로, 다양한 실시 예에 따른 제1 전자 장치(310)(예: 전자 장치(100))는, 제2 제스처(1623) 및 제2 발화(1625)에 기반하여 영상 데이터(1601)에 대한 제2 처리를 수행할 수 있다. 일 실시 예에 따르면, 제1 전자 장치(310)는 제2 제스처(1623)에 기반하여 영상 데이터(1601)를 일부를 제2 처리 대상으로 지정하고, 제2 발화(1625)에 기반하여 처리 대상으로 지정된 제2 처리 대상에 대한 검색 동작을 수행하고, 검색 결과의 적어도 일부를 영상 데이터(1601)의 저장 앨범(또는 폴더) 생성에 이용할 수 있다. 예를 들어, 도시된 바와 같이, 제1 전자 장치(310)는 제2 발화(1625)에 기반하여, 결과의 적어도 일부를 저장 앨범 이름(1631)으로 설정할 수 있다.
추가적으로 또는 선택적으로, 다양한 실시 예에 따른 제1 전자 장치(310)(예: 전자 장치(100))는, 영상 데이터(1601)에 대한 제1 처리(예: 영상 데이터(1601) 저장) 및 제2 처리(예: 검색 후 저장 앨범 이름 설정)를 수행하는 동안, 제1 처리 및 제2 처리 중 적어도 하나에 대한 처리 결과를 제공할 수 있다. 예를 들어, 제1 전자 장치(310)는 제2 처리 대상에 대한 검색 결과(예: 검색 내용)를 나타내는 시각적 정보(1635)(예: 광화문입니다. 광화문은 경복궁의 정문(正門)이며…)를 제공할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 제1 처리 및 제2 처리 중 적어도 하나에 대한 처리 결과는, 도 16c의 1640에 도시된 바와 같이 제1 전자 장치(310)와 통신으로 연결된 외부 전자 장치(예: 무선 이어폰(1641))를 통해 청각적 정보로 출력되거나 도 16c의 1650에 도시된 바와 같이, 제1 전자 장치(310)의 스피커를 통해 청각적 정보로 출력될 수도 있다.
전술한 바와 같이, 다양한 실시 예에 따른 제1 전자 장치(310)는 제1 제스처(1603)에 기반하여 영상 데이터(1601)를 전체를 처리 대상으로 지정할 수 있다. 그러나, 이는 예시적일 뿐, 다양한 실시 예들이 이에 한정되는 것은 아니다. 예컨대, 도 16b의 1635에 도시된 바와 같이, 제1 전자 장치(310)는 제1 제스처(1603)에 의해 형성되는 손 액자에 포함된 특정 객체만 처리 대상으로 지정할 수도 있다. 예를 들어, 제1 전자 장치(310)는 제1 발화(1605)에 기반하여 저장된 제1 제스처(1603)에 대응되는 영상 데이터(1601)의 일부(1634)를 제2 발화(1625)에 기반하여 설정된 저장 앨범에 저장할 수 있다.
다양한 실시 예에 따른 전자 장치(310)의 동작 방법은 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터(201)를 획득하는 동작, 상기 발화 데이터가 획득되는 동안, 영상 데이터(210)를 획득하는 동작, 상기 영상 데이터에 포함된 복수의 객체들(211 내지 217) 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하는 동작 및 상기 적어도 두 개의 객체에 대한 연관 정보(250)를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 발화 데이터는 질의(201-3)를 더 포함할 수 있다. 일 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 질의와 관련된 상기 연관 정보를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 적어도 두 개의 객체와 상기 질의의 적어도 일부에 기반하여 검색어를 생성하는 동작 및 상기 생성된 검색어와 관련된 상기 연관 정보를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 연관 정보를 외부 장치(320)로부터 획득하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 지정된 단어는, 복수의 대상들을 지정하는 제1 단어를 포함할 수 있다. 일 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제1 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제1 제스처(413, 511)에 대응되는 제1 객체와 제2 객체를 식별하는 동작 및 상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 지정된 단어는, 단일의 대상을 지정하는 제2 단어(201-1) 및 제3 단어(201-2)를 포함할 수 있다. 일 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하는 동작, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하는 동작 및 상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 제2 단어가 발화된 이후 일정 시간 내에 상기 제3 단어가 발화되면, 상기 영상 데이터에서 제2 객체를 식별하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 서로 다른 시야에 해당되는 제1 영상 데이터(260)와 제2 영상 데이터(270)를 획득하는 동작, 상기 제1 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 제1 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하는 동작 및 상기 제2 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 제2 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 발화 데이터가 획득되는 동안, 외부 장치(320)로부터 모션 데이터(341)를 획득하는 동작 및 상기 모션 데이터가 지정된 조건을 만족하면, 상기 적어도 두 개의 객체를 식별하는 동작을 포함할 수 있다.
다양한 실시 예에 따르면, 상기 전자 장치(310)의 동작 방법은 상기 발화 데이터와 상기 영상 데이터를 상기 전자 장치(310)(예: 메모리(313))에 저장된 인공 지능 모델(3131)로 입력하는 동작 및 상기 인공 지능 모델의 출력에 기반하여 상기 적어도 두 개의 객체를 식별하는 동작을 포함할 수 있다.
다양한 실시 예에 따른 컴퓨터로 판독 가능한 기록 매체는, 전자 장치에 의해 실행되었을 때, 상기 전자 장치가, 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터를 획득하고, 상기 발화 데이터가 획득되는 동안, 영상 데이터를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보를 제공하도록 설정될 수 있다.
다양한 실시 예에 따른 영상 인식 시스템은 제1 전자 장치(310)와 제2 전자 장치(320)를 포함할 수 있다.
일 실시 예에 따르면, 상기 제1 전자 장치(310)는 적어도 하나의 제1 프로세서(311), 카메라(314), 마이크(315) 및 상기 적어도 하나의 제1 프로세서, 상기 카메라 및 상기 마이크와 동작적으로 연결되며 적어도 하나의 제1 명령어를 저장하는 제1 메모리(313)를 포함하고, 상기 적어도 하나의 제1 명령어는 상기 적어도 하나의 제1 프로세서에 의해 실행될 때, 상기 제1 전자 장치로 하여금: 객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터(201)가 상기 마이크를 통해 획득되는 동안, 상기 카메라를 통해 영상 데이터(210)를 획득하고, 상기 영상 데이터에 포함된 복수의 객체들(211 내지 217) 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고, 상기 적어도 두 개의 객체에 대한 연관 정보(250)를 제공하도록 구성될 수 있다.
일 실시 예에 따르면, 상기 제2 전자 장치(320)는 적어도 하나의 제2 프로세서(321), 센서(324) 및 상기 적어도 하나의 제2 프로세서 및 상기 센서와 동작적으로 연결되며 적어도 하나의 제2 명령어를 저장하는 제2 메모리(323)를 포함하고, 상기 적어도 하나의 제2 명령어는 상기 적어도 하나의 제2 프로세서에 의해 실행될 때, 상기 제2 전자 장치로 하여금 상기 센서를 통해 획득되는 상기 제2 전자 장치의 자세와 관련된 정보를 상기 제1 전자 장치로 제공하도록 구성될 수 있다.
다양한 실시 예에 따르면, 상기 적어도 하나의 제1 명령어는 상기 적어도 하나의 제1 프로세서에 의해 실행될 때, 상기 제1 전자 장치로 하여금: 상기 제2 전자 장치로부터 제공되는 상기 제2 전자 장치의 자세와 관련된 정보가 지정된 조건을 만족하는 경우, 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하도록 구성될 수 있다.
Claims (15)
- 전자 장치(310)에 있어서,적어도 하나의 프로세서(311);카메라(314);마이크(315); 및상기 적어도 하나의 프로세서(311), 상기 카메라(314) 및 상기 마이크(315)와 동작적으로 연결되며 적어도 하나의 명령어를 저장하는 메모리(313)를 포함하고,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로(individually or collectively) 실행될 때, 상기 전자 장치(310)로 하여금:객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터(201)가 상기 마이크(315)를 통해 획득되는 동안, 상기 카메라(314)를 통해 영상 데이터(210)를 획득하고,상기 영상 데이터에 포함된 복수의 객체들(211 내지 217) 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하고,상기 적어도 두 개의 객체에 대한 연관 정보(250)를 제공하도록 구성된 전자 장치.
- 제1항에 있어서,상기 발화 데이터는, 질의(201-3)를 더 포함하며,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 질의와 관련된 상기 연관 정보를 제공하도록 구성된 전자 장치.
- 제2항에 있어서,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 적어도 두 개의 객체와 상기 질의의 적어도 일부에 기반하여 검색어를 생성하고,상기 생성된 검색어와 관련된 상기 연관 정보를 제공하도록 구성된 전자 장치.
- 제1항에 있어서,통신 회로(312)를 더 포함하며,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 통신 회로(312)를 통해 외부 장치(320)로부터 상기 연관 정보를 획득하도록 구성된 전자 장치.
- 제1항에 있어서,상기 지정된 단어는, 복수의 대상들을 지정하는 제1 단어를 포함하며,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제1 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제1 제스처(413, 511)에 대응되는 제1 객체와 제2 객체를 식별하고,상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하도록 구성된 전자 장치.
- 제1항에 있어서,상기 지정된 단어는, 단일의 대상을 지정하는 제2 단어(201-1) 및 제3 단어(201-2)를 포함하며,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하고,상기 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하고,상기 제1 객체와 상기 제2 객체에 대한 상기 연관 정보를 제공하도록 구성된 전자 장치.
- 제6항에 있어서,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 제2 단어가 발화된 이후 일정 시간 내에 상기 제3 단어가 발화되면, 상기 영상 데이터에서 제2 객체를 식별하도록 구성된 전자 장치.
- 제6항에 있어서,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 카메라(314)를 통해 서로 다른 시야에 해당되는 제1 영상 데이터(250)와 제2 영상 데이터(270)를 획득하고,상기 제1 영상 데이터에 포함된 복수의 객체들 중, 상기 제2 단어가 발화되는 시점에 상기 제1 영상 데이터에서 식별되는 제2 제스처(221)에 대응되는 제1 객체를 식별하고,상기 제2 영상 데이터에 포함된 복수의 객체들 중, 상기 제3 단어가 발화되는 시점에 상기 제2 영상 데이터에서 식별되는 제3 제스처(223)에 대응되는 제2 객체를 식별하도록 구성된 전자 장치.
- 제1항에 있어서,통신 회로(312)를 더 포함하며,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 발화 데이터가 획득되는 동안, 상기 통신 회로(312)를 통해 외부 장치(320)로부터 모션 데이터(341)를 획득하고,상기 모션 데이터가 지정된 조건을 만족하면, 상기 적어도 두 개의 객체를 식별하도록 구성된 전자 장치.
- 제1항 내지 제9항 중 어느 하나에 있어서,상기 적어도 하나의 명령어는 상기 적어도 하나의 프로세서(311)에 의해 개별적으로 또는 집합적으로 실행될 때, 상기 전자 장치(310)로 하여금:상기 발화 데이터와 상기 영상 데이터를 상기 전자 장치(310)에 저장된 인공 지능 모델(3131)로 입력하고,상기 인공 지능 모델의 출력에 기반하여 상기 적어도 두 개의 객체를 식별하도록 구성된 전자 장치.
- 전자 장치(310)의 동작 방법에 있어서,객체의 지정과 관련된 지정된 단어가 포함된 발화 데이터(201)를 획득하는 동작;상기 발화 데이터가 획득되는 동안, 영상 데이터(210)를 획득하는 동작;상기 영상 데이터에 포함된 복수의 객체들(211 내지 217) 중, 상기 지정된 단어가 발화되는 시점에 상기 영상 데이터에서 식별되는 제스처에 대응되는 적어도 두 개의 객체를 식별하는 동작; 및상기 적어도 두 개의 객체에 대한 연관 정보(250)를 제공하는 동작을 포함하는 방법.
- 제11항에 있어서,상기 발화 데이터는 질의(201-3)를 더 포함하며,상기 질의와 관련된 상기 연관 정보를 제공하는 동작을 포함하는 방법.
- 제12항에 있어서,상기 적어도 두 개의 객체와 상기 질의의 적어도 일부에 기반하여 검색어를 생성하는 동작; 및상기 생성된 검색어와 관련된 상기 연관 정보를 제공하는 동작을 포함하는 방법.
- 제11항에 있어서,상기 연관 정보를 외부 장치(320)로부터 획득하는 동작을 포함하는 방법.
- 제11항 내지 제14항 중 어느 하나에 있어서,상기 발화 데이터와 상기 영상 데이터를 상기 전자 장치(310)에 저장된 인공 지능 모델(3131)로 입력하는 동작; 및상기 인공 지능 모델의 출력에 기반하여 상기 적어도 두 개의 객체를 식별하는 동작을 포함하는 방법.
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR20240084955 | 2024-06-28 | ||
| KR10-2024-0084955 | 2024-06-28 | ||
| KR1020240105108A KR20260002099A (ko) | 2024-06-28 | 2024-08-07 | 영상 인식 방법 및 이를 지원하는 전자 장치 |
| KR10-2024-0105108 | 2024-08-07 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026005331A1 true WO2026005331A1 (ko) | 2026-01-02 |
Family
ID=98222402
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2025/007732 Pending WO2026005331A1 (ko) | 2024-06-28 | 2025-06-05 | 영상 인식 방법 및 이를 지원하는 전자 장치 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026005331A1 (ko) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20190013390A (ko) * | 2017-08-01 | 2019-02-11 | 삼성전자주식회사 | 전자 장치 및 이의 검색 결과 제공 방법 |
| US20230169744A1 (en) * | 2020-04-14 | 2023-06-01 | Worldpay Limited | Methods and systems for displaying virtual objects from an augmented reality environment on a multimedia device |
| KR20230118334A (ko) * | 2022-02-04 | 2023-08-11 | 한국전자통신연구원 | 사용자 인터랙션에 기초하여 레이블링 및 캡셔닝을 수행하는 전자 장치 |
| JP7373068B2 (ja) * | 2020-05-18 | 2023-11-01 | 株式会社Nttドコモ | 情報処理システム |
| US20240037956A1 (en) * | 2022-07-29 | 2024-02-01 | Faurecia Clarion Electronics Co., Ltd. | Data processing system, data processing method, and information providing system |
-
2025
- 2025-06-05 WO PCT/KR2025/007732 patent/WO2026005331A1/ko active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20190013390A (ko) * | 2017-08-01 | 2019-02-11 | 삼성전자주식회사 | 전자 장치 및 이의 검색 결과 제공 방법 |
| US20230169744A1 (en) * | 2020-04-14 | 2023-06-01 | Worldpay Limited | Methods and systems for displaying virtual objects from an augmented reality environment on a multimedia device |
| JP7373068B2 (ja) * | 2020-05-18 | 2023-11-01 | 株式会社Nttドコモ | 情報処理システム |
| KR20230118334A (ko) * | 2022-02-04 | 2023-08-11 | 한국전자통신연구원 | 사용자 인터랙션에 기초하여 레이블링 및 캡셔닝을 수행하는 전자 장치 |
| US20240037956A1 (en) * | 2022-07-29 | 2024-02-01 | Faurecia Clarion Electronics Co., Ltd. | Data processing system, data processing method, and information providing system |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2016085173A1 (en) | Device and method of providing handwritten content in the same | |
| WO2019107981A1 (en) | Electronic device recognizing text in image | |
| WO2020032563A1 (en) | System for processing user voice utterance and method for operating same | |
| WO2016117836A1 (en) | Apparatus and method for editing content | |
| WO2020076014A1 (en) | Electronic apparatus and method for controlling the electronic apparatus | |
| WO2020071858A1 (en) | Electronic apparatus and assistant service providing method thereof | |
| WO2016089079A1 (en) | Device and method for outputting response | |
| WO2020032564A1 (en) | Electronic device and method for providing one or more items in response to user speech | |
| WO2021029582A1 (en) | Co-reference understanding electronic apparatus and controlling method thereof | |
| WO2019194426A1 (en) | Method for executing application and electronic device supporting the same | |
| WO2016108407A1 (ko) | 주석 제공 방법 및 장치 | |
| EP3381180A1 (en) | Photographing device and method of controlling the same | |
| WO2026005331A1 (ko) | 영상 인식 방법 및 이를 지원하는 전자 장치 | |
| WO2025005544A1 (ko) | 개인화된 이미지를 제공하는 방법 및 전자 장치 | |
| WO2019168208A1 (ko) | 이동 단말기 및 그 제어 방법 | |
| WO2022131578A1 (ko) | 증강 현실 환경을 제공하기 위한 방법 및 전자 장치 | |
| WO2022177224A1 (ko) | 전자 장치 및 전자 장치의 동작 방법 | |
| WO2019124775A1 (ko) | 전자 장치 및 전자 장치에서 방송 콘텐트와 관련된 서비스 정보 제공 방법 | |
| WO2025089769A1 (ko) | 전자 장치 및 이미지 편집 방법 | |
| WO2025165060A1 (ko) | 전자 장치 및 이미지 편집 방법 | |
| WO2025258824A1 (ko) | 비디오 프레임의 추가 영역을 생성하기 위한 전자 장치, 방법, 및 비일시적 컴퓨터 판독 가능 저장 매체 | |
| WO2026014703A1 (ko) | 콘텐츠와 사용자 입력을 연동하는 전자 장치, 방법, 및 비-일시적 컴퓨터 판독 가능 기록 매체 | |
| WO2024228577A1 (ko) | 고도화된 추론 및 추정 기능 기반의 지능적 응답 에이전트 제공 방법 및 그 시스템 | |
| WO2026005187A1 (ko) | 전자 장치 및 전자 장치에서 외부 ai모델과 온디바이스 ai 모델을 이용하여 데이터를 생성하는 방법 | |
| WO2026054381A1 (ko) | 메시지를 제공하는 전자 장치, 방법, 및 저장 매체 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25827297 Country of ref document: EP Kind code of ref document: A1 |