WO2026005261A1 - 추론 컨텍스트 확장 방법 및 장치 - Google Patents

추론 컨텍스트 확장 방법 및 장치

Info

Publication number
WO2026005261A1
WO2026005261A1 PCT/KR2025/005923 KR2025005923W WO2026005261A1 WO 2026005261 A1 WO2026005261 A1 WO 2026005261A1 KR 2025005923 W KR2025005923 W KR 2025005923W WO 2026005261 A1 WO2026005261 A1 WO 2026005261A1
Authority
WO
WIPO (PCT)
Prior art keywords
inference
electronic device
prompt
user input
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/005923
Other languages
English (en)
French (fr)
Inventor
박재현
김형석
송가진
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020240101521A external-priority patent/KR20260001025A/ko
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Publication of WO2026005261A1 publication Critical patent/WO2026005261A1/ko
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/332Query formulation
    • G06F16/3329Natural language query formulation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/3331Query processing
    • G06F16/334Query execution
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/338Presentation of query results
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/953Querying, e.g. by the use of web search engines
    • G06F16/9538Presentation of query results
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/04Inference or reasoning models

Definitions

  • the present disclosure relates to a method and device for extending an inference context.
  • Embodiments of the present disclosure relate to a method and device capable of extending an inference context by sharing inference results between different inference models to process user input prompts (e.g., user utterances).
  • Generative AI can generate natural, high-quality content by understanding the context of a conversation based on natural language input prompts from the user.
  • Generative AI based on a Large Language Model (LLM) learns the relationships between words, phrases, and sentences in large-scale text data collected from the internet and elsewhere, enabling it to generate contextually appropriate responses to user input prompts.
  • LLM Large Language Model
  • Generative AI can perform a variety of functions, including answering user questions, engaging in conversational interactions, summarizing documents, and generating creative writing, across diverse fields such as education, information provision, entertainment, and technical support. Furthermore, generative AI can enhance educational utility and information accessibility by answering questions on a wide range of topics and assisting with complex problem-solving.
  • the present disclosure relates to a method and device for extending an inference context.
  • Embodiments of the present disclosure relate to a method and device capable of extending an inference context by sharing inference results between different inference models to process user input prompts.
  • an electronic device may include: a communication circuit; at least one processor including a processing circuit; and a memory including one or more storage media storing instructions.
  • the electronic device may cause the electronic device to determine an inference model from among at least one inference model; determine a context vector based on a user input prompt; selectively determine at least one reference vector based on the user input prompt based on a quality of the determined inference model; retrieve related data from an inference database based on the context vector and the reference vector; generate an integrated prompt including the user input prompt and the related data; obtain an inference result from the determined inference model based on the integrated prompt; and update the inference database based on the user input prompt and the inference result.
  • a method for expanding an inference context of an electronic device may include an operation of determining an inference model from among at least one inference model; an operation of determining a context vector based on a user input prompt; an operation of selectively determining at least one reference vector based on the user input prompt based on a quality of the determined inference model; an operation of searching for related data in an inference database based on the context vector and the reference vector; an operation of generating an integrated prompt including the user input prompt and the related data; an operation of obtaining an inference result from the determined inference model based on the integrated prompt; and an operation of updating the inference database based on the user input prompt and the inference result.
  • a computer-readable recording medium having recorded thereon a program for performing the method may be included.
  • responses optimized for the individual can be generated based on the user's preferences, interests, previous questions, etc., thereby improving the user experience and providing the user with the information he or she needs more quickly and accurately.
  • conversational continuity can be maintained between services based on different inference models or across different sessions, allowing users to continue conversations without losing the context of previous conversations even when using different sessions or different inference model services.
  • This conversational continuity can further enhance the user experience, particularly in activities requiring continuous information exchange and interaction, such as project work, education, and research.
  • personalized recommendations can be provided based on the user's previous activities and interactions. If the user previously inquired about meditation techniques to help manage stress, the system can then provide more personalized meditation guidance or health advice, taking into account the user's stress level and lifestyle.
  • the platform can reference their previous questions and interactions to provide a more in-depth and personalized learning experience.
  • the platform can consider their previous learning progress and interests to suggest questions that can guide them to the next step, or provide additional resources for areas they struggled to understand. This allows learners to tailor their learning to their own pace and style, creating personalized learning paths.
  • Various embodiments of the present disclosure can improve information accessibility by making it easier to access related information on previously encountered topics or questions.
  • Various embodiments of the present disclosure can eliminate the need for users to repeatedly request the same information, and can provide more in-depth and relevant, personalized responses to users' questions.
  • shopping applications utilizing AR technology allow users to virtually experience products while shopping.
  • personalized product recommendations or virtual demonstrations can be provided based on previous conversations about product types or brands the user has shown interest in. For example, if a user has previously expressed interest in luxury watches, an AR shopping application could offer the user the opportunity to virtually try on a new collection of luxury timepieces, providing a more personalized shopping experience.
  • AR vision sensors can identify items of interest and record them along with tag information. When matching information is found, the application can provide not only text information but also image information recorded with the tag.
  • the inference response quality can be improved even for an inference model with low response quality.
  • FIGS. 1a, 1b, 1c and 1d are conceptual diagrams of an inference context extension method according to one embodiment of the present disclosure.
  • Figure 3 illustrates the inference result according to the prior art without expanding the inference context.
  • FIGS. 6A and 6B are schematic flowcharts of a method for expanding an inference context in an electronic device according to one embodiment of the present disclosure.
  • FIG. 7 is a schematic flowchart of a method for storing and updating an inference database when the data processing level is 0 according to one embodiment of the present disclosure.
  • FIG. 8 is a schematic flowchart of a method for storing and updating an inference database through inference expansion when the data processing level is 1 or higher according to one embodiment of the present disclosure.
  • FIGS. 9A and 9B illustrate UI screens displaying inference results based on some of the relevant data retrieved from the inference database selected by the user according to one embodiment of the present disclosure.
  • FIG. 10 is a block diagram of an electronic device within a network environment according to various embodiments.
  • FIGS. 1a, 1b, 1c and 1d are conceptual diagrams of an inference context extension method according to one embodiment of the present disclosure.
  • a method for expanding an inference context can obtain an inference result based on a previous user context by generating an integrated prompt including a previous conversation content (or related data) having a predetermined degree of similarity or higher with a user input prompt in an inference database.
  • an electronic device that extends an inference context may receive a prompt from a user (111).
  • the prompt may be data input for an inference model to generate output.
  • Components of the prompt may include at least one of an instruction that instructs a specific task or instruction to be performed by the inference model, external information that can adjust the inference model, context information that refers to additional context, input data corresponding to a question for which an answer is sought, and output data that refers to the type or format of the output.
  • the prompt does not have to include all components, and may include information such as instructions or questions to be conveyed to the inference model, as well as other detailed information such as input or examples.
  • the electronic device may determine an inference model from among at least one available inference model (112).
  • the electronic device may determine the inference model from among the at least one inference model based on the quality of the inference model and status information of the electronic device.
  • the quality of the inference model may be determined based on at least one of the size of the inference model and an inference result benchmarking score.
  • the status information of the electronic device may include at least one of a network status and a resource status.
  • the resource status may include at least one of memory availability, performance of at least one processor of the electronic device, and battery availability. For example, if the electronic device is currently in a state where it is difficult to use a network, an on-device AI that can be used without a network connection may be determined as the inference model from among the at least one inference model.
  • the electronic device can determine a context vector based on a user input prompt.
  • the electronic device can determine the context vector by converting the user input prompt into an embedding vector (114).
  • the context vector can include an embedding vector corresponding to a word or combination of words within the user input prompt that best describes the user input prompt, a sentence or phrase that summarizes the user input prompt, or the like.
  • the electronic device may selectively determine at least one reference vector based on the quality of the determined inference model, and optionally based on a user input prompt (115). For example, if the determined inference model is small in size, such as an on-device AI, the electronic device may additionally determine at least one reference vector based on the user input prompt to improve the quality of the inference response. The electronic device may determine the at least one reference vector by generating at least one keyword based on the user input prompt and generating an embedding vector corresponding to the at least one keyword.
  • the electronic device can search for related data (123) in the inference database (122) based on the context vector and the reference vector.
  • the electronic device can search for related data (123) (or previous conversations) having a similarity (or correlation criterion) or higher in the inference database (122) based on the context vector and the reference vector.
  • the electronic device can search for context data and reference data having an embedding vector having a similarity or higher than the predetermined level of embedding vector of each of the context vector and the reference vector.
  • the related data (123) can include at least one of context data searched for based on the context vector and reference data searched for based on the reference vector in the inference database (122).
  • the inference database (122) may include at least one database record.
  • Each database record may include text keywords (key in FIG. 1b) corresponding to all or part of keywords in conversations with all inference models (e.g., LLM, on-device AI) previously used by the user, embedding vectors (embedding in FIG. 1b) corresponding to the keywords, and inference results (value in FIG. 1b) obtained from the inference models.
  • the inference database (122) may store previous inference results without modification, or may store them after modifying them, such as by extracting, adding, and summarizing them, according to the user's settings or system conditions.
  • the context data and the reference data may be selected based on user selection to select only the necessary data.
  • the correlation criterion may be preset by user input or system settings, or may be dynamically set based on the amount of the context data and the reference data.
  • the correlation criterion may be updated based on the user's partial selection of the context data and the reference data, thereby reflecting user-specific tendencies.
  • the electronic device may generate an integrated prompt including at least one of the user input prompt and the related data (131).
  • the electronic device may further generate a soft prompt to generate the integrated prompt.
  • the soft prompt may include information preset for each inference model for further explanation of context data or reference data.
  • the soft prompt may be determined based on suitability among frequently used values generated in advance, or may be determined based on artificial intelligence learning, but is not limited thereto.
  • a UI screen for displaying an inference result based on a user selection from among related data searched from an inference database will be described below with reference to FIG. 9.
  • the electronic device may obtain an inference result from the determined inference model based on the integrated prompt.
  • the electronic device may display the user input prompt and the inference result on a UI screen (132).
  • the electronic device may additionally display the context data and reference data used by the inference model for inference on the UI screen.
  • the electronic device can update the inference database (122) based on the user input prompt and the inference result (141, 142).
  • a recursive inference process can be further performed based on the data processing level. This allows for obtaining an expanded inference result and updating the inference database (122).
  • the electronic device (400) may determine whether an identical database record exists in the inference database based on the embedding vector. If an identical database record exists, the electronic device may proceed to operation 730, and if an identical database record does not exist, the electronic device may proceed to operation 750.
  • the electronic device (400) may post-process the inference result and the existing database record. If there is an existing database record that matches the embedding vector by a predetermined degree of similarity or higher, the existing inference result and the newly updated inference result may be merged into one. At this time, in order to remove duplication in the merged inference result, the merged inference result may be transmitted to a predetermined inference model (e.g., a large-scale language model) to obtain an inference result with semantic duplication removed.
  • a predetermined inference model e.g., a large-scale language model
  • the electronic device (400) may store the user input prompt, the embedding vector, and the inference result as one database record in an inference database.
  • the electronic device (400) can expand the inference result based on the data processing level and update the expanded inference result in the inference database.
  • the data processing level may be a level that determines how much the current inference result is expanded and stored when storing the inference result in the inference database. The higher the data processing level, the more expanded inference result can be obtained through a proportional number of recursive inference processes.
  • the electronic device (400) can obtain each second inference result based on at least one text from a predetermined inference model.
  • the electronic device (400) can convert at least one text into at least one embedding vector.
  • the electronic device (400) may determine whether an identical database record exists in the inference database based on the embedding vector. If an identical database record exists, the electronic device may proceed to operation 850, and if an identical database record does not exist, the electronic device may proceed to operation 870.
  • the electronic device (400) can update the text, embedding vector, and post-processing results in the inference database.
  • FIGS. 9A and 9B illustrate UI screens displaying inference results based on some of the relevant data retrieved from the inference database selected by the user according to one embodiment of the present disclosure.
  • the electronic device (400) can obtain a user input prompt based on user input.
  • the electronic device (400) can determine a context vector based on a user input prompt.
  • the electronic device (400) can determine the context vector by converting the user input prompt into an embedding vector.
  • the electronic device (400) can determine at least one reference vector by extracting at least one keyword based on the user input prompt and converting it into an embedding vector.
  • the electronic device (400) can search for related data in the inference database (122) based on the context vector and the reference vector.
  • the electronic device (400) can search for related data having a similarity (or correlation reference value) or higher in the inference database based on the context vector and the reference vector.
  • the electronic device (400) can search for context data and reference data having an embedding vector having a similarity or higher than a predetermined level with respect to the embedding vectors of each of the context vector and the reference vector.
  • the related data can include at least one of context data searched based on the context vector and reference data searched based on the reference vector in the inference database (122).
  • the electronic device (400) can obtain and display an inference result from the inference model based on the generated integrated prompt (930).
  • the electronic device (400) can generate an integrated prompt by adding user information associated with a user account. Based on the integrated prompt with the added user information, the electronic device (400) can obtain and display customized inference results based on the user information from the inference model.
  • the electronic device (400) can generate an integrated prompt by adding a user profile generated in each conversation session. Based on the integrated prompt with the added user profile, the electronic device (400) can obtain and display inference results based on previous conversations accumulated within a single session from an inference model.
  • FIG. 10 is a block diagram of an electronic device within a network environment according to various embodiments.
  • FIG. 10 is a block diagram of an electronic device (1001) within a network environment (1000) according to various embodiments.
  • the electronic device (1001) may communicate with the electronic device (1002) via a first network (1098) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1004) or the server (1008) via a second network (1099) (e.g., a long-range wireless communication network).
  • the electronic device (1001) may communicate with the electronic device (1004) via the server (1008).
  • the electronic device (1001) may include a processor (1020), a memory (1030), an input module (1050), an audio output module (1055), a display module (1060), an audio module (1070), a sensor module (1076), an interface (1077), a connection terminal (1078), a haptic module (1079), a camera module (1080), a power management module (1088), a battery (1089), a communication module (1090), a subscriber identification module (1096), or an antenna module (1097).
  • the electronic device (1001) may omit at least one of these components (e.g., the connection terminal (1078)), or may have one or more other components added.
  • some of these components e.g., sensor module (1076), camera module (1080), or antenna module (1097) may be integrated into a single component (e.g., display module (1060)).
  • the processor (1020) may, for example, execute software (e.g., a program (1040)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1001) connected to the processor (1020) and perform various data processing or operations.
  • the processor (1020) may store commands or data received from other components (e.g., a sensor module (1076) or a communication module (1090)) in a volatile memory (1032), process the commands or data stored in the volatile memory (1032), and store result data in a non-volatile memory (1034).
  • the processor (1020) may include a main processor (1021) (e.g., a central processing unit or an application processor) or an auxiliary processor (1023) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1021).
  • a main processor (1021) e.g., a central processing unit or an application processor
  • an auxiliary processor (1023) e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor
  • the auxiliary processor (1023) may be configured to use less power than the main processor (1021) or to be specialized for a given function.
  • the auxiliary processor (1023) may be implemented separately from the main processor (1021) or as a part thereof.
  • the auxiliary processor (1023) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1060), the sensor module (1076), or the communication module (1090)) of the electronic device (1001), for example, on behalf of the main processor (1021) while the main processor (1021) is in an inactive (e.g., sleep) state, or together with the main processor (1021) while the main processor (1021) is in an active (e.g., application execution) state.
  • the auxiliary processor (1023) e.g., an image signal processor or a communication processor
  • the auxiliary processor (1023) may include a hardware structure specialized for processing artificial intelligence models.
  • the artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (1001) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1008)).
  • the learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above.
  • the artificial intelligence model can include a plurality of artificial neural network layers.
  • the artificial neural network can be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, or a combination of two or more of the above, but is not limited to the examples described above.
  • the artificial intelligence model can additionally or alternatively include a software structure.
  • the memory (1030) can store various data used by at least one component (e.g., the processor (1020) or the sensor module (1076)) of the electronic device (1001).
  • the data can include, for example, software (e.g., the program (1040)) and input data or output data for commands related thereto.
  • the memory (1030) can include volatile memory (1032) or non-volatile memory (1034).
  • the program (1040) may be stored as software in memory (1030) and may include, for example, an operating system (1042), middleware (1044), or an application (1046).
  • the input module (1050) can receive commands or data to be used in a component of the electronic device (1001) (e.g., a processor (1020)) from an external source (e.g., a user) of the electronic device (1001).
  • the input module (1050) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
  • the audio output module (1055) can output audio signals to the outside of the electronic device (1001).
  • the audio output module (1055) can include, for example, a speaker or a receiver.
  • the speaker can be used for general purposes, such as multimedia playback or recording playback.
  • the receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
  • the display module (1060) can visually provide information to an external party (e.g., a user) of the electronic device (1001).
  • the display module (1060) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device.
  • the display module (1060) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
  • the audio module (1070) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1070) can acquire sound through the input module (1050), output sound through the sound output module (1055), or an external electronic device (e.g., electronic device (1002)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1001).
  • an external electronic device e.g., electronic device (1002)
  • an external electronic device e.g., electronic device (1002)
  • speaker or headphone directly or wirelessly connected to the electronic device (1001).
  • the sensor module (1076) can detect the operating status (e.g., power or temperature) of the electronic device (1001) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status.
  • the sensor module (1076) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
  • the interface (1077) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1001) to an external electronic device (e.g., the electronic device (1002)).
  • the interface (1077) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
  • HDMI high definition multimedia interface
  • USB universal serial bus
  • SD card interface Secure Digital Card
  • the connection terminal (1078) may include a connector through which the electronic device (1001) may be physically connected to an external electronic device (e.g., the electronic device (1002)).
  • the connection terminal (1078) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
  • the haptic module (1079) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations.
  • the haptic module (1079) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
  • the camera module (1080) can capture still images and videos.
  • the camera module (1080) may include one or more lenses, image sensors, image signal processors, or flashes.
  • the power management module (1088) can manage power supplied to the electronic device (1001).
  • the power management module (1088) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
  • PMIC power management integrated circuit
  • a battery (1089) may power at least one component of the electronic device (1001).
  • the battery (1089) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
  • the communication module (1090) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1001) and an external electronic device (e.g., electronic device (1002), electronic device (1004), or server (1008)), and the performance of communication through the established communication channel.
  • the communication module (1090) may operate independently from the processor (1020) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication.
  • the communication module (1090) may include a wireless communication module (1092) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1094) (e.g., a local area network (LAN) communication module, or a power line communication module).
  • a wireless communication module (1092) e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module
  • GNSS global navigation satellite system
  • a wired communication module (1094) e.g., a local area network (LAN) communication module, or a power line communication module.
  • any of these communication modules may communicate with an external electronic device (1004) via a first network (1098) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1099) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).
  • a first network e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)
  • a second network (1099) e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)
  • a first network e.g.,
  • the wireless communication module (1092) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1096) to verify or authenticate the electronic device (1001) within a communication network such as the first network (1098) or the second network (1099).
  • subscriber information e.g., an international mobile subscriber identity (IMSI)
  • IMSI international mobile subscriber identity
  • the wireless communication module (1092) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology).
  • NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)).
  • eMBB enhanced mobile broadband
  • mMTC massive machine type communications
  • URLLC ultra-reliable and low-latency communications
  • the wireless communication module (1092) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate.
  • a high-frequency band e.g., mmWave band
  • the wireless communication module (1092) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna.
  • the wireless communication module (1092) may support various requirements specified in the electronic device (1001), an external electronic device (e.g., the electronic device (1004)), or a network system (e.g., the second network (1099)).
  • the wireless communication module (1092) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
  • a peak data rate e.g., 20 Gbps or more
  • a loss coverage e.g., 164 dB or less
  • U-plane latency e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip
  • the antenna module (1097) can transmit or receive signals or power to or from an external device (e.g., an external electronic device).
  • the antenna module (1097) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB).
  • the antenna module (1097) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1098) or the second network (1099), may be selected from the plurality of antennas, for example, by the communication module (1090). A signal or power may be transmitted or received between the communication module (1090) and an external electronic device via the at least one selected antenna.
  • another component e.g., a radio frequency integrated circuit (RFIC)
  • RFIC radio frequency integrated circuit
  • the antenna module (1097) may form a mmWave antenna module.
  • the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
  • a first side e.g., a bottom side
  • a plurality of antennas e.g., an array antenna
  • At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
  • peripheral devices e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
  • commands or data may be transmitted or received between the electronic device (1001) and an external electronic device (1004) via a server (1008) connected to a second network (1099).
  • Each of the external electronic devices (1002 or 1004) may be the same or a different type of device as the electronic device (1001).
  • all or part of the operations executed in the electronic device (1001) may be executed in one or more of the external electronic devices (1002, 1004, or 1008). For example, when the electronic device (1001) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1001) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service.
  • One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1001).
  • the electronic device (1001) may process the result as is or additionally and provide it as at least a portion of a response to the request.
  • cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example.
  • the electronic device (1001) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example.
  • the external electronic device (1004) may include an Internet of Things (IoT) device.
  • the server (1008) may be an intelligent server utilizing machine learning and/or a neural network.
  • the external electronic device (1004) or the server (1008) may be included in the second network (1099).
  • the electronic device (1001) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
  • an electronic device may include: a communication circuit; at least one processor including a processing circuit; and a memory including one or more storage media storing instructions.
  • the instructions may cause the electronic device to determine an inference model from among at least one inference model; determine a context vector based on a user input prompt; selectively determine at least one reference vector based on a quality of the determined inference model, based on the user input prompt; retrieve related data from an inference database based on the context vector and the reference vector; generate an integrated prompt including the user input prompt and the related data; obtain an inference result from the determined inference model based on the integrated prompt; and update the inference database based on the user input prompt and the inference result.
  • the electronic device may cause: the electronic device to determine an inference model from among the at least one inference model based on a quality of the inference model and status information of the electronic device.
  • the quality of the inference model is determined based on at least one of a size of the inference model and an inference result benchmarking score;
  • the status information of the electronic device includes at least one of a network status and a resource status;
  • the resource status may include at least one of a memory availability, a performance of the at least one processor, and a battery availability.
  • the electronic device may cause: if the quality of the determined inference model is below a predetermined standard, to extract at least one keyword from the user input prompt; and to convert each of the at least one keyword into at least one embedding vector to determine the at least one reference vector.
  • the related data may include at least one context data retrieved based on the context vector and at least one reference data retrieved based on the reference vector in the inference database.
  • the electronic device may cause: the electronic device to generate the integrated prompt by selecting a portion of the related data that has a predetermined degree of similarity or greater with the user input prompt, summarizing the related data, or selecting a portion of the related data based on a user input, taking into account the size of the integrated prompt that the determined inference model can process.
  • the electronic device may cause: to segment the user prompt and the inference result into at least one text of a predetermined unit based on a data processing level; to obtain each second inference result from one of the at least one inference model based on each of the at least one text; to convert each of the at least one text into at least one embedding vector; and to update the text, the embedding vector, and the second inference result in the inference database.
  • the data processing level may be set by user input, preset in the electronic device, or set based on the quality of the determined inference model and the quality of at least one currently available inference model.
  • the electronic device further comprises at least one sensor; and when the instructions are individually or collectively executed by the at least one processor, the electronic device may cause the electronic device to: obtain system data from the at least one sensor; and further include tag information corresponding to the system data to generate the integrated prompt.
  • the system data may include at least one of location data, day of the week data, and time data; and the tag information may correspond to the system data.
  • a method for expanding an inference context of an electronic device may include an operation of determining an inference model from among at least one inference model; an operation of determining a context vector based on a user input prompt; an operation of selectively determining at least one reference vector based on the user input prompt based on a quality of the determined inference model; an operation of searching for related data in an inference database based on the context vector and the reference vector; an operation of generating an integrated prompt including the user input prompt and the related data; an operation of obtaining an inference result from the determined inference model based on the integrated prompt; and an operation of updating the inference database based on the user input prompt and the inference result.
  • the operation of selectively determining at least one reference vector based on the quality of the determined inference model and based on a user input prompt may include: extracting at least one keyword from the user input prompt when the quality of the inference model is below a predetermined standard; and converting each of the at least one keyword into at least one embedding vector to determine the at least one reference vector.
  • Electronic devices may take various forms. Electronic devices may include, for example, display devices, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
  • phrases such as “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B, or C”, “at least one of A, B, and C”, and “at least one of A, B, or C” can each include any one of the items listed together in that phrase, or all possible combinations thereof.
  • Terms such as “first”, “second”, or “first” or “second” may be used merely to distinguish the corresponding element from other corresponding elements and do not limit the corresponding elements in any other respect (e.g., importance or order).
  • part or module used in various embodiments of this document may include units implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit.
  • the "part” or “module” may be an integrally formed component or a minimum unit or part of the component that performs one or more functions.
  • the “part” or “module” may be implemented in the form of an application-specific integrated circuit (ASIC).
  • ASIC application-specific integrated circuit
  • the program executed by the electronic device (400) described in this document may be implemented as hardware components, software components, and/or a combination of hardware components and software components.
  • the program may be executed by any system capable of executing computer-readable instructions.
  • Software may include a computer program, code, instructions, or a combination of one or more of these, which can configure a processing device to perform a desired operation or command the processing device, either independently or collectively.
  • Software may be implemented as a computer program including instructions stored on a computer-readable storage medium. Examples of the computer-readable storage medium include magnetic storage media (e.g., read-only memory (ROM), random-access memory (RAM), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROMs, digital versatile discs (DVDs)).
  • the computer-readable storage medium may be distributed across network-connected computer systems so that the computer-readable code can be stored and executed in a distributed manner.
  • Computer programs can be distributed online (e.g., by download or upload) through an application store (e.g., Play StoreTM) or directly between two user devices (e.g., smartphones).
  • an application store e.g., Play StoreTM
  • two user devices e.g., smartphones
  • at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
  • each component e.g., a module or a program of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components.
  • one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added.
  • a plurality of components e.g., a module or a program
  • the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration.
  • the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Databases & Information Systems (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Human Computer Interaction (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Software Systems (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

본 개시의 다양한 실시예는 추론 컨텍스트 확장 방법 및 장치에 관한 것이다. 이를 위한 전자 장치는 적어도 하나의 추론 모델 중 추론 모델을 결정하고; 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하고; 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하고; 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하고; 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하고; 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하고; 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기반하여 상기 추론 데이터베이스를 갱신할 수 있다.

Description

추론 컨텍스트 확장 방법 및 장치
본 개시는 추론 컨텍스트 확장 방법 및 장치에 관한 것이다. 본 개시의 실시예들은 사용자 입력 프롬프트(예. 사용자 발화)를 처리하기 위해, 상이한 추론 모델들 간에 추론 결과를 공유함으로써 추론 컨텍스트를 확장할 수 있는 방법 및 장치에 관한 것이다.
생성형 AI는 사용자가 입력하는 자연어 입력 프롬프트를 기반으로 대화의 컨텍스트를 이해함으로써 자연스럽고 고품질의 컨텐츠를 생성할 수 있다. 대규모 언어 모델(Large Language Model; LLM)에 기반한 생성형 AI는 인터넷 등에서 수집한 대규모 텍스트 데이터의 단어, 구, 문장 간의 관계를 학습함으로써, 사용자 입력 프롬프트에 대해 문맥적으로 적절한 답변을 생성할 수 있다.
생성형 AI는 교육, 정보 제공, 엔터테인먼트, 기술 지원 등 다양한 분야에서 사용자의 질문에 대한 답변, 대화 형식의 상호작용, 문서 요약, 창작 글쓰기 등 다양한 기능을 수행하고 컨텐츠를 생성할 수 있다. 또한, 생성형 AI는 다양한 주제의 질문에 답변하고, 복잡한 문제 해결을 지원하는 등 교육적 활용도 및 정보 접근성을 향상시킬 수 있다.
본 개시는 추론 컨텍스트 확장 방법 및 장치에 관한 것이다. 본 개시의 실시예들은 사용자 입력 프롬프트를 처리하기 위해, 상이한 추론 모델들 간에 추론 결과를 공유함으로써 추론 컨텍스트를 확장할 수 있는 방법 및 장치에 관한 것이다.
상술한 기술적 과제를 해결하기 위하여 본 개시의 일 실시예에 따르면 전자 장치는, 통신 회로; 프로세싱 회로를 포함하는 적어도 하나의 프로세서; 및 명령어들을 저장하는 하나 이상의 저장 매체를 포함하는, 메모리를 포함할 수 있다. 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금, 적어도 하나의 추론 모델 중 추론 모델을 결정하고; 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하고; 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하고; 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하고; 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하고; 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하고; 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기초하여 상기 추론 데이터베이스를 갱신하도록 야기할 수 있다.
또한, 본 개시의 일 실시예에 따르면 전자 장치의 추론 컨텍스트 확장 방법은, 적어도 하나의 추론 모델 중 추론 모델을 결정하는 동작; 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하는 동작; 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하는 동작; 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하는 동작; 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하는 동작; 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하는 동작; 및 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기반하여 상기 추론 데이터베이스를 갱신하는 동작을 포함할 수 있다.
또한, 본 개시의 일 실시예에 따르면 상기 방법을 수행하기 위한 프로그램이 기록된 컴퓨터로 읽을 수 있는 기록매체를 포함할 수 있다.
본 개시의 다양한 실시예들에 따르면, 확장형 추론 데이터베이스를 통해 서로 다른 추론 모델 간에 사용자 대화 내용을 공유함으로써, 사용자의 선호도, 관심사, 이전 질문 등에 기반하여 개인에게 최적화된 응답을 생성할 수 있으므로 사용자 경험을 개선할 수 있고, 사용자가 필요한 정보를 더 신속하고 정확하게 제공할 수 있다.
본 개시의 다양한 실시예들에 따르면, 상이한 추론 모델 기반의 서비스들 사이에서나 상이한 세션 사이에서 대화의 연속성을 유지할 수 있으므로, 사용자가 다른 세션 또는 다른 추론 모델 서비스를 사용하더라도 이전 대화의 맥락(context)을 잃지 않고 이어갈 수 있다. 특히, 프로젝트 작업, 교육, 연구 등 연속적인 정보 교환과 상호 작용이 필요한 활동에서 이러한 대화 연속성은 사용자 경험을 더 향상시킬 수 있다.
예를 들어, 사용자가 건강 및 웰니스 추적 애플리케이션을 통해 운동 목표 설정, 식단 관리, 스트레스 관리 등에 관한 조언을 구하는 경우, 사용자의 이전 활동과 상호작용을 기반으로 개인화된 추천을 제공할 수 있다. 사용자가 이전에 스트레스 관리에 도움이 되는 명상 방법을 문의했다면, 시스템은 그 이후 사용자의 스트레스 수준과 생활 패턴을 고려하여 보다 맞춤화된 명상 가이드나 건강 조언을 제공할 수 있다.
예를 들어, 온라인 교육 플랫폼에서 학습자가 특정 주제에 대해 질문하고 학습하는 과정에서, 학습자의 이전 질문과 상호작용을 참조하여 보다 깊이 있고 맞춤형 학습 경험을 제공할 수 있다. 학습자가 특정 과학 주제에 대해 궁금증을 제시했을 때, 학습자의 이전 학습 진행 상황과 관심사를 고려하여 다음 단계로 나아갈 수 있는 질문을 제시하거나, 학습자가 이해하기 어려워한 부분에 대해 추가 자료를 제공할 수 있다. 따라서, 학습자가 자신의 학습 속도와 스타일에 맞게 학습을 조정할 수 있으며, 개인화된 학습 경로를 생성할 수 있다.
본 개시의 다양한 실시예들에 따르면, 이전에 접했던 주제나 질문에 대한 연관 정보를 더 쉽게 접근할 수 있으므로 정보의 접근성을 향상시킬 수 있다. 본 개시의 다양한 실시예들에 따르면, 사용자는 여러 번 같은 정보를 요청하지 않아도 되며, 사용자의 질문에 더 깊이 있고 관련성 높은 맞춤화된 응답을 제공할 수 있다.
예를 들어, 증강 현실 기술을 활용한 쇼핑 애플리케이션에서 사용자는 가상으로 제품을 체험하며 쇼핑을 할 수 있다. 특히 사용자가 이전에 관심을 보였던 제품 유형이나 브랜드에 대한 대화 내용을 기반으로, 사용자에게 맞춤형 제품 추천이나 가상 시연을 제공할 수 있다. 사용자가 이전에 고급 시계에 대한 관심을 표현했다면, AR(Augmented Reality) 쇼핑 애플리케이션은 사용자에게 새로운 컬렉션의 고급 시계를 가상으로 착용해볼 수 있는 기회를 제공함으로써, 보다 개인화된 쇼핑 경험을 제공할 수 있다. 또한 증강 현실 비전 센서를 통해 사용자의 관심 물품을 파악하고 이를 태그 정보와 함께 기록함으로써, 이후 사용자가 원하는 정보와 일치할 때 텍스트 정보뿐만 아니라 해당 태그와 함께 기록된 이미지 정보를 추가로 제공해 줄 수 있다.
본 개시의 다양한 실시예들에 따르면, 확장형 추론 데이터베이스로부터 맥락 정보를 획득할 뿐만 아니라 상기 맥락 정보와 연관된 참조 정보를 추가로 획득하여, 사용자 입력 프롬프트를 증강시킨 통합 프롬프트를 생성하고 추론 모델에 제공함으로써, 응답 품질이 낮은 추론 모델이더라도 추론 응답 품질을 개선시킬 수 있다.
본 개시의 예시적 실시예들에서 얻을 수 있는 효과는 이상에서 언급한 효과들로 제한되지 아니하며, 언급되지 아니한 다른 효과들은 이하의 기재로부터 본 개시의 예시적 실시예들이 속하는 기술분야에서 통상의 지식을 가진 자에게 명확하게 도출되고 이해될 수 있다. 즉, 본 개시의 예시적 실시예들을 실시함에 따른 의도하지 아니한 효과들 역시 본 개시의 예시적 실시예들로부터 당해 기술분야의 통상의 지식을 가진 자에 의해 도출될 수 있다.
도 1a, 1b, 1c 및 1d는 본 개시의 일 실시예에 따른 추론 컨텍스트 확장 방법의 개념도이다.
도 2a 및 도 2b는 본 개시의 일 실시예에 따른 추론 컨텍스트를 확장한 경우 획득할 수 있는 추론 결과를 도시한다.
도 3은 추론 컨텍스트를 확장하지 않은 종래 기술에 따른 추론 결과를 도시한다.
도 4는 본 개시의 일 실시예에 따른 전자 장치의 블록도이다.
도 5a 및 도 5b는 본 개시의 일 실시예에 따른 데이터 처리 레벨에 따른 추론 확장을 개략적으로 도시한다.
도 6a 및 6b는 본 개시의 일 실시예에 따른 전자 장치에서의 추론 컨텍스트 확장 방법의 개략적인 흐름도이다.
도 7은 본 개시의 일 실시예에 따른 데이터 처리 레벨이 0인 경우 추론 데이터베이스 저장 및 갱신 방법의 개략적인 흐름도이다.
도 8은 본 개시의 일 실시예에 따른 데이터 처리 레벨이 1이상인 경우 추론 확장을 통한 추론 데이터베이스 저장 및 갱신 방법의 개략적인 흐름도이다.
도 9a 및 도 9b는 본 개시의 일 실시예에 따른 추론 데이터베이스에서 검색된 관련 데이터 중 사용자가 선택한 일부에 기초한 추론 결과를 디스플레이하는 UI 화면을 도시한다.
도 10은 다양한 실시예들에 따른, 네트워크 환경 내의 전자 장치의 블록도이다.
이하에서는 도면을 참조하여 본 개시의 실시예에 대하여 본 개시가 속하는 기술 분야에서 통상의 지식을 가진 자가 용이하게 실시할 수 있도록 상세히 설명한다. 그러나 본 개시는 여러 가지 상이한 형태로 구현될 수 있으며 여기에서 설명하는 실시예에 한정되지 않는다. 도면의 설명과 관련하여, 동일하거나 유사한 구성요소에 대해서는 동일하거나 유사한 참조 부호가 사용될 수 있다. 또한, 도면 및 관련된 설명에서는, 잘 알려진 기능 및 구성에 대한 설명이 명확성과 간결성을 위해 생략될 수 있다.
도 1a, 1b, 1c 및 1d는 본 개시의 일 실시예에 따른 추론 컨텍스트 확장 방법의 개념도이다.
일 실시예에 따른 추론 컨텍스트 확장 방법은, 추론 데이터베이스에서 사용자 입력 프롬프트와 소정 유사도 이상인 이전 대화 내용(또는 관련 데이터)을 포함하는 통합 프롬프트를 생성함으로써, 이전 사용자 컨텍스트에 기초한 추론 결과를 획득할 수 있다.
도 1a를 참조하면, 추론 컨텍스트를 확장하는 전자 장치는 사용자로부터 프롬프트를 입력받을 수 있다(111). 상기 프롬프트는 추론 모델이 출력을 생성하기 위해 입력되는 데이터일 수 있다. 상기 프롬프트의 구성 요소는 추론 모델이 수행할 특정 작업 또는 지침을 지시하는 명령(Instruction), 추론 모델을 조정할 수 있는 외부 정보, 추가 맥락을 일컫는 맥락 정보(Context), 답변을 획득하고자 하는 입력, 질문에 대응하는 입력 데이터(Input Data), 출력의 유형 또는 형식을 의미하는 출력 데이터(Output Data)가 적어도 하나 이상 포함될 수 있다. 상기 프롬프트에 모든 구성 요소가 포함되어야 하는 것은 아니며 추론 모델에 전달할 지시 사항이나 질문 같은 정보가 포함될 수 있으며, 입력이나 예제 같은 기타 세부 정보도 포함될 수 있다.
일 실시예에 따르면, 상기 전자 장치는 이용 가능한 적어도 하나의 추론 모델 중 추론 모델을 결정할 수 있다(112). 상기 전자 장치는 추론 모델의 품질 및 상기 전자 장치의 상태 정보에 기초하여 상기 적어도 하나의 추론 모델 중 추론 모델을 결정할 수 있다. 상기 추론 모델의 품질은 상기 추론 모델의 사이즈 및 추론 결과 벤치마킹 점수 중 적어도 하나에 기초하여 결정될 수 있다. 상기 전자 장치의 상태 정보는 네트워크 상태 및 리소스 상태 중 적어도 하나를 포함할 수 있다. 상기 리소스 상태는 메모리 가용량, 상기 전자 장치의 적어도 하나의 프로세서의 성능 및 배터리 가용량 중 적어도 하나를 포함할 수 있다. 예를 들어, 상기 전자 장치가 현재 네트워크 사용이 어려운 상태라면, 상기 적어도 하나의 추론 모델 중 네트워크 연결 없이도 사용 가능한 온 디바이스 AI를 추론 모델로 결정할 수 있다.
일 실시예에 따르면, 상기 전자 장치는 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정할 수 있다. 상기 전자 장치는 사용자 입력 프롬프트를 임베딩 벡터로 변경함으로써 컨텍스트 벡터를 결정할 수 있다(114). 상기 컨텍스트 벡터는 사용자 입력 프롬프트를 가장 잘 설명할 수 있는, 사용자 입력 프롬프트 내의 단어나 단어들의 조합, 사용자 입력 프롬프트를 요약한 문장이나 구 등에 대응되는 임베딩 벡터를 포함할 수 있다.
일 실시예에 따르면, 상기 전자 장치는 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정할 수 있다(115). 예를 들어, 상기 결정된 추론 모델이 온 디바이스 AI와 같이 추론 모델의 사이즈가 작은 경우, 상기 전자 장치는 추론 응답 품질을 높이기 위해 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 추가로 결정할 수 있다. 상기 전자 장치는 상기 사용자 입력 프롬프트에 기초하여 적어도 하나의 키워드를 생성하고, 적어도 하나의 키워드에 대응하는 임베딩 벡터를 생성함으로써 적어도 하나의 참조 벡터를 결정할 수 있다.
도 1b를 참조하면, 상기 전자 장치는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스(122)에서 관련 데이터(123)를 검색할 수 있다. 상기 전자 장치는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스(122)에서 소정 유사도(또는 상관도 기준값) 이상인 관련 데이터(123)(또는 이전 대화)를 검색할 수 있다. 상기 전자 장치는 상기 컨텍스트 벡터 및 상기 참조 벡터 각각의 임베딩 벡터와 소정 유사도 이상의 임베딩 벡터를 갖는 컨텍스트 데이터 및 참조 데이터를 각각 검색할 수 있다. 따라서, 관련 데이터(123)는 추론 데이터베이스(122)에서, 상기 컨텍스트 벡터에 기초하여 검색된 컨텍스트 데이터 및 상기 참조 벡터에 기초하여 검색된 참조 데이터를 적어도 하나 포함할 수 있다.
일 실시예에 따르면, 추론 데이터베이스(122)는 적어도 하나의 데이터베이스 레코드를 포함할 수 있다.
각각의 데이터베이스 레코드는 사용자가 이전에 사용한 모든 추론 모델들(예. LLM, 온 디바이스 AI)과의 대화의 전부 또는 일부 키워드에 대응하는 텍스트 키워드(도 1b의 key), 상기 키워드에 대응하는 임베딩 벡터(도 1b의 embedding), 추론 모델로부터 획득한 추론 결과(도 1b의 value)를 포함할 수 있다. 추론 데이터베이스(122)는 이전 추론 결과를 변형 없이 저장하거나, 사용자의 설정 또는 시스템 상황에 따라 추출, 추가 및 요약 등 변형하여 저장할 수 있다. 상기 컨텍스트 데이터 및 상기 참조 데이터는 사용자 선택에 기초하여 필요한 데이터만 선택될 수 있다. 상기 상관도 기준값은 사용자 입력 또는 시스템 설정에 의해 사전 설정될 수 있거나, 상기 컨텍스트 데이터 및 상기 참조 데이터의 양에 기초하여 동적으로 설정될 수 있다. 상기 상관도 기준값은 상기 컨텍스트 데이터 및 상기 참조 데이터를 사용자가 일부 선택함에 기초하여 갱신될 수 있고, 이를 통해 사용자별 성향이 반영될 수 있다.
도 1c를 참조하면, 상기 전자 장치는 상기 사용자 입력 프롬프트 및 상기 관련 데이터 중 적어도 하나를 포함하는 통합 프롬프트를 생성할 수 있다(131). 상기 전자 장치는 소프트 프롬프트를 더 추가하여 통합 프롬프트를 생성할 수 있다. 상기 소프트 프롬프트는 컨텍스트 데이터나 참조 데이터에 대한 부연 설명을 위해 추론 모델별로 사전 설정된 정보를 포함할 수 있다. 소프트 프롬프트는 자주 사용되는 값들을 사전에 생성해서 이 중 적합도에 따라 결정되거나, 인공지능 학습에 기초하여 결정될 수도 있으나 이에 제한되지 않는다. 추론 데이터베이스에서 검색된 관련 데이터 중 사용자 선택에 기초한 추론 결과를 디스플레이하는 UI 화면을 이하 도 9를 참조하여 후술한다.
일 실시예에 따르면, 상기 전자 장치는 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득할 수 있다. 상기 전자 장치는 상기 사용자 입력 프롬프트 및 상기 추론 결과를 UI 화면으로 구성하여 디스플레이할 수 있다(132). 상기 전자 장치는 추론 모델이 추론에 사용한 컨텍스트 데이터 및 참조 데이터를 상기 UI 화면에 추가 구성하여 디스플레이할 수 있다.
도 1d를 참조하면, 상기 전자 장치는 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기초하여 추론 데이터베이스(122)를 갱신할 수 있다(141, 142). 이 경우, 데이터 처리 레벨에 기초하여, 재귀적 추론 과정을 더 수행할 수 있다. 이를 통해 확장된 추론 결과를 획득할 수 있고, 확장된 추론 결과를 추론 데이터베이스(122)에 갱신할 수 있다.
일 실시예에 따르면, 추론 모델과의 대화 내용을 추론 데이터베이스에 저장 및 갱신하고, 추론 데이터베이스를 여러 추론 모델 간에 공유함으로써 사용자가 다른 세션 또는 다른 추론 모델을 사용할 때도 이전 대화의 컨텍스트를 연속적으로 유지할 수 있다. 사용자는 추론 모델이 달라지거나, 세션이 달라지더라도 대화의 연속성을 보장받을 수 있고, 이전 컨텍스트 전달을 위해 사용자가 이전 대화에서 입력했던 내용을 추가 입력하지 않더라도 사용자 맞춤형 응답을 생성할 수 있다. 일 실시예에 따르면, 추론 데이터베이스를 통해 사용자가 이전에 표현한 관심사, 선호하는 답변 스타일, 과거 질문 주제들을 분석할 수 있고, 사용자의 성향과 필요에 맞춘 맞춤형 응답을 생성할 수 있다.
도 2a 및 도 2b는 본 개시의 일 실시예에 따른 추론 컨텍스트를 확장한 경우 획득할 수 있는 추론 결과를 도시한다.
일 실시예에 따르면, 이전 대화 내용 중 현재 사용자 입력 프롬프트와 소정 유사도 이상인 대화 내용을 추론 데이터베이스에서 획득하여 현재 사용자 입력 프롬프트를 증강시킨 통합 프롬프트를 생성할 수 있다. 일 실시예에 따르면, 상기 통합 프롬프트에 기초하여 사용자의 이전 대화와 컨텍스트가 일치하는 추론 결과를 획득할 수 있다. 상기 추론 데이터베이스는 이전 대화 내용에 대응하는 단어, 문장뿐만 아니라 해당 대화가 생성된 시점 정보를 더 포함할 수 있다. 대화 생성 시점 정보를 통해 사용자 입력 프롬프트와의 유사도를 판단할 때, 대화 생성 시점에 기초하여 가중치를 부여할 수 있다. 예를 들어, 최근에 생성된 대화 내용은 유사도 판단에 있어 더 높은 가중치를 부여받을 수 있다.
예를 들어, 도 2a는 사용자의 취향이나 관심사를 반영한 추론 모델의 추론 결과를 도시한다. 사용자가 최근에 음식, 요리, 식단과 같은 주제의 대화를 나눈 경험이 있다면 해당 대화 내용이 추론 데이터베이스에 저장되고, 현재 사용자 입력 프롬프트와 소정 유사도 이상인 대화 내용을 추론 데이터베이스에서 검색할 수 있다.
예를 들어 사용자가 최근에 식단, 단백질 및 체중 감량에 대한 대화를 나누었다면 도 2b에서와 같이, 식단, 단백질 및 체중 감량에 대응되는 데이터베이스 레코드들이 추론 데이터베이스에 저장될 수 있다. 저장된 대화 내용과 관련된 사용자 입력 프롬프트가 입력된 경우 추론 데이터베이스에서 대응되는 데이터베이스 레코드들이 검색될 수 있다. 이를 통해 사용자 입력 프롬프트 및 추론 데이터베이스 검색 결과가 포함된 통합 프롬프트가 생성될 수 있고, 상기 통합 프롬프트에 기초하여 추론 모델로부터 추론 결과를 획득할 수 있다. 따라서, 상기 추론 결과는 사용자의 관심사나 취미 등 이전 대화 내용에 기초한 응답을 포함할 수 있다.
도 3은 추론 컨텍스트를 확장하지 않은 종래 기술에 따른 추론 결과를 도시한다.
도 2a의 추론 결과와 상이하게, 도 3의 추론 결과는 현재 사용자 입력 프롬프트에만 기초하여 생성된 것을 보여준다.
도 4는 본 개시의 일 실시예에 따른 전자 장치의 블록도이다.
일 실시예에 따르면, 도 4를 참조하여 설명되는 전자 장치(400)의 동작에 있어서, 전술한 도면들(도 1a 내지 도3)에서 설명된 부분과 중복되는 부분은 생략될 수 있다. 전자 장치(400)는 도시된 구성요소 외에 추가적인 구성요소를 포함하거나, 도시된 구성요소 중 적어도 하나를 생략할 수 있다.
일 실시예에 따르면, 전자 장치(400)는 통신 회로(미도시), 프로세싱 회로를 포함하는 적어도 하나의 프로세서(미도시); 및 명령어들을 저장하는 하나 이상의 저장 매체를 포함하는 메모리(미도시)를 포함할 수 있다. 상기 적어도 하나의 프로세서는 추론 컨텍스트 매니저(Inference Context Manager)(420), 추론 데이터베이스 매니저(Inference DB Manager)(430), 컨텍스트 스플리터(Context Splitter)(440), 시스템 데이터 매니저(System Data Manager)(450), 추론 요청 핸들러(460), 온 디바이스 AI(470), LLM(480) 각각의 모듈을 구성하는 적어도 하나의 명령어들을 실행할 수 있다.
일 실시예에 따르면, 추론 컨텍스트 매니저(Inference Context Manager)(420)는 프롬프트 애널라이저(Prompt Analyzer)(421), 임베딩 핸들러(Embedding Handler)(422), 프롬프트 제너레이터(Prompt Generator)(424), 데이터베이스 컨트롤러(DB Controller)(423)를 포함할 수 있다.
일 실시예에 따르면, 추론 컨텍스트 매니저(Inference Context Manager)(420)는 사용자 입력에 기초하여 사용자 입력 프롬프트를 획득할 수 있다.
일 실시예에 따르면, 프롬프트 애널라이저(Prompt Analyzer)(421)는 사용자 입력 프롬프트를 단어, 구, 문장으로 구분하여 적어도 하나의 키워드를 추출할 수 있다.
일 실시예에 따르면, 임베딩 핸들러(Embedding Handler)(422)는 사용자 입력 프롬프트를 임베딩 벡터로 변경함으로써 컨텍스트 벡터를 결정할 수 있다. 임베딩 핸들러(Embedding Handler)(422)는 상기 적어도 하나의 키워드 각각을 임베딩 벡터로 변경함으로써 적어도 하나의 참조 벡터를 결정할 수 있다. 임베딩 핸들러(Embedding Handler)(422)는 단어 및 문장과 같은 텍스트 데이터, 다른 유형의 원시 데이터를 전자 장치(400)가 이해하고 처리할 수 있는 형태인 수치 벡터로 변환할 수 있다. 임베딩 핸들러(Embedding Handler)(422)는 고차원 데이터를 저차원의 밀집 벡터로 표현함으로써 데이터 간의 관계를 수학적으로 계산할 수 있도록 하고, 텍스트 데이터 간의 의미의 유사도를 수학적으로 계산할 수 있다.
일 실시예에 따르면, 데이터베이스 컨트롤러(DB Controller)(423)는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스(Inference DB)(431)에서 관련 데이터를 검색할 수 있다. 데이터베이스 컨트롤러(DB Controller)(423)는 추론 데이터베이스 매니저(Inference DB Manager)(430)를 통해 추론 데이터베이스(Inference DB)(431)에서 관련 데이터를 검색할 수 있다. 전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터 각각의 임베딩 벡터와 소정 유사도 이상의 임베딩 벡터를 갖는 컨텍스트 데이터 및 참조 데이터를 각각 검색할 수 있다. 도 4를 참조하면 추론 데이터베이스(Inference DB)(431)는 전자 장치(400) 내부에 있으나, 클라우드와 같이 전자 장치(400) 외부에 존재할 수 있다. 추론 데이터베이스(122)의 위치나 데이터베이스 관리 방법이 다양할 수 있음은 당업자에게 이해될 것이다.
일 실시예에 따르면, 시스템 데이터 매니저(System Data Manager)(450)는 전자 장치(400)에 포함된 적어도 하나의 센서로부터 시스템 데이터를 획득할 수 있다. 상기 시스템 데이터는 위치 데이터, 요일 데이터 및 시간 데이터 중 적어도 하나를 포함할 수 있다. 프롬프트 제너레이터(Prompt Generator)(424)는 상기 시스템 데이터에 대응되는 태그 정보를 더 포함하여 상기 통합 프롬프트를 생성할 수 있다. 상기 태그 정보는 상기 시스템 데이터에 대응하는 텍스트 정보를 포함할 수 있다. 예를 들어, 사용자가 '여기에서~~'와 같이 장소와 관련된 질의를 하는 경우, 대응되는 위치 정보를 획득할 수 있고, 상기 위치 정보에 대응되는 장소를 지시하는 텍스트 정보를 태그 정보로서 획득할 수 있다. 시스템 데이터 매니저(System Data Manager)(450)에 의해 지속적으로 획득되는 상기 태그 정보가 기존의 사용자 컨텍스트와 중복되는 경우, 예를 들어 GPS 센서에 수집된 정보가 사용자의 집과 같이 이미 시스템 데이터 매니저(System Data Manager)(450)에 의해 획득된 태그 정보와 같을 경우, 중복된 태그 정보는 삭제될 수 있다. 또한 기존에 수집된 복수의 시스템 데이터가 집합적으로 하나의 의미를 가리키는 경우, 예를 들어 일정 시간 이상 수집된 GPS 데이터가 하나의 의미를 가지는 장소(예. 회사)를 나타내는 경우, 기존에 수집된 복수의 GPS 데이터는 하나의 대표값으로 표현될 수 있고, 상기 대표값에 대응하는 태그 정보가 획득될 수 있다. 임베딩 핸들러(Embedding Handler)(422)는 상기 태그 정보에 대응되는 임베딩 벡터를 결정할 수 있다. 상기 태그 정보에 대응되는 임베딩 벡터는 데이터베이스 컨트롤러(DB Controller)(423)에게 전달될 수 있다. 데이터베이스 컨트롤러(DB Controller)(423)는 상기 태그 정보에 대응하는 임베딩 벡터에 기초하여 추론 데이터베이스(Inference DB)(431)에서 관련 데이터를 검색할 수 있다.
일 실시예에 따르면, 프롬프트 제너레이터(Prompt Generator)(424)는 상기 사용자 입력 프롬프트 및 상기 관련 데이터 중 적어도 하나를 포함하는 통합 프롬프트를 생성할 수 있다. 프롬프트 제너레이터(Prompt Generator)(424)는 상기 결정된 추론 모델이 처리할 수 있는 통합 프롬프트의 크기를 고려하여, 상기 관련 데이터 중 상기 사용자 입력 프롬프트와 소정 유사도 이상인 관련 데이터를 선택하거나, 상기 관련 데이터를 요약하거나, 사용자 입력에 기초하여 상기 관련 데이터 중 일부를 선택함으로써 상기 통합 프롬프트를 생성할 수 있다. 프롬프트 제너레이터(Prompt Generator)(424)는 소프트 프롬프트를 더 추가하여 통합 프롬프트를 생성할 수 있다. 상기 소프트 프롬프트는 컨텍스트 데이터나 참조 데이터에 대한 부연 설명을 위해 추론 모델별로 사전 설정된 정보를 포함할 수 있다.
일 실시예에 따르면, 추론 요청 핸들러(Inference Request Handler)(460)는 상기 통합 프롬프트를 온 디바이스 AI(470) 또는 LLM(480)에 전달함으로써 추론 결과를 요청할 수 있다. 적어도 하나의 추론 모델 중 온 디바이스 AI(470) 또는 LLM(480)을 결정하는 방법은 도 1a를 참조하여 전술한 바와 같으나 이에 제한되지 않는다. 추론 요청 핸들러(Inference Request Handler)(460)는 온 디바이스 AI(470) 또는 LLM(480)으로부터 상기 통합 프롬프트에 대응하는 추론 결과를 획득할 수 있다. 온 디바이스 AI(470)는 전자 장치(400) 내에 존재하는 추론 모델로서, 네트워크 연결이 없더라도 전자 장치(400) 내에서 추론 결과를 제공할 수 있다. LLM(480)은 클라우드 기반으로 추론을 수행하는 추론 모델로서, 네트워크를 통해 전자 장치(400)에게 추론 결과를 제공할 수 있다. 상기 사용자 입력 프롬프트 및 상기 추론 결과는 UI 화면으로 구성되어 디스플레이될 수 있다. 상기 UI 화면은 추론 모델이 추론에 사용한 컨텍스트 데이터 및 참조 데이터를 추가로 디스플레이할 수 있다. 상기 추론 결과는 추론 컨텍스트 매니저(Inference Context Manager)(420)에게도 전달될 수 있다.
일 실시예에 따르면, 임베딩 핸들러(Embedding Handler)(422)는 사용자 입력 프롬프트 및 상기 추론 결과에 대응하는 적어도 하나의 임베딩 벡터를 결정할 수 있다. 상기 추론 결과는 단어, 구, 문장 등으로 구분되어 적어도 하나의 임베딩 벡터로 결정될 수 있다. 사용자 입력 프롬프트 및 상기 추론 결과에 대응하는 상기 적어도 하나의 임베딩 벡터들은 데이터베이스 컨트롤러(DB Controller)(423)에게 전달될 수 있다.
일 실시예에 따르면, 데이터베이스 컨트롤러(DB Controller)(423)는 추론 데이터베이스 매니저(Inference DB Manager)(430)를 통해, 추론 데이터베이스(Inference DB)(431)에 레코드를 추가, 삭제 및 갱신할 수 있다. 데이터베이스 컨트롤러(DB Controller)(423)는 상기 적어도 하나의 임베딩 벡터들에 기초하여, 추론 데이터베이스(Inference DB)(431)를 갱신할 수 있다. 일 실시예에 따르면, 데이터베이스 컨트롤러(DB Controller)(423)는 추론 데이터베이스 매니저(Inference DB Manager)(430)를 통해 추론 데이터베이스를 갱신할 수 있다. 상기 임베딩 벡터와 소정 유사도 이상 일치하는 기존 데이터베이스 레코드가 존재하는 경우, 기존 추론 결과와 새로 갱신할 추론 결과를 하나로 합병하여 추론 데이터베이스(Inference DB)(431)를 갱신할 수 있다. 이때 합병된 추론 결과의 중복을 제거하기 위해 소정 추론 모델(예. 대규모 언어 모델)에게 합병된 추론 결과를 전달하고 의미상 중복이 제거된 추론 결과를 획득할 수 있고 추론 데이터베이스(Inference DB)(431)에 갱신할 수 있다. 또는, 합병된 추론 결과를 적어도 하나의 키워드로 나누어 임베딩 벡터를 결정한 후, 임베딩 벡터의 유사도에 기초하여 합병된 추론 결과에서 의미상 중복이 제거된 추론 결과를 획득할 수 있고 추론 데이터베이스(Inference DB)(431)에 갱신할 수 있다. 합병된 추론 결과의 중복을 제거하는 방법이 이에 제한되지 않음은 당업자에게 이해될 것이다.
일 실시예에 따르면, 컨텍스트 스플리터(Context Splitter)(440)는 데이터 처리 레벨에 기초하여 추론 결과를 확장하고, 확장된 추론 결과를 추론 데이터베이스(Inference DB)(431)에 갱신할 수 있다. 상기 데이터 처리 레벨은 추론 결과를 추론 데이터베이스(Inference DB)(431)에 저장할 때 현재의 추론 결과를 얼마나 확장시켜 저장할지를 결정하는 레벨일 수 있다. 데이터 처리 레벨이 높을수록, 그에 비례하는 재귀적 횟수의 추론 과정을 통해 더 확장된 추론 결과를 획득할 수 있다.
예를 들어, 데이터 처리 레벨이 0인 경우, 추론 모델로부터 획득한 추론 결과를 하나의 데이터베이스 레코드로 추론 데이터베이스(Inference DB)(431)에 저장할 수 있다. 예를 들어, 데이터 처리 레벨이 1 이상인 경우, 상기 사용자 프롬프트 및 상기 추론 결과를 데이터 처리 레벨에 기초한 소정 단위의 적어도 하나의 텍스트로 구분할 수 있다. 컨텍스트 스플리터(Context Splitter)(440)는 상기 적어도 하나의 텍스트 각각에 기초하여 적어도 하나의 추론 모델 중 하나의 추론 모델로부터 각각의 제2 추론 결과를 획득할 수 있다. 컨텍스트 스플리터(Context Splitter)(440)는 상기 적어도 하나의 텍스트를 적어도 하나의 임베딩 벡터로 각각 변환할 수 있다. 컨텍스트 스플리터(Context Splitter)(440)는 상기 텍스트, 상기 임베딩 벡터, 상기 제2 추론 결과를 추론 데이터베이스(Inference DB)(431)에 갱신할 수 있다.
상기 데이터 처리 레벨은 사용자 입력에 의해 설정되거나, 전자 장치(400)에서 사전 설정되거나, 이용중인 추론 모델의 품질 및 현재 이용가능한 적어도 하나의 추론 모델의 품질에 기초하여 설정될 수 있다. 예를 들어, 현재 이용중인 추론 모델이 온 디바이스 AI와 같이 품질이 낮고, 현재 LLM 추론 모델을 이용할 수 있는 경우, 데이터 처리 레벨을 1이상으로 설정함으로써, 현재의 추론 결과를 더 확장시켜 추론 데이터베이스(Inference DB)(431)에 저장할 수 있다.
도 5a 및 도 5b는 본 개시의 일 실시예에 따른 데이터 처리 레벨에 따른 추론 확장을 개략적으로 도시한다.
도 5a 및 도 5b를 참조하면, 데이터 처리 레벨이 0인 경우(520), 사용자 입력 프롬프트를 하나의 임베딩 벡터로 생성하고, 사용자 입력 프롬프트, 상기 임베딩 벡터 및 추론 모델로부터 획득한 추론 결과를 포함하는 하나의 데이터베이스 레코드를 추론 데이터베이스에 저장할 수 있다.
도 5a 및 도 5b를 참조하면, 데이터 처리 레벨이 2인 경우(530), 먼저 데이터 처리 레벨 0에 대응하는 추론 결과에서 적어도 하나의 키워드를 추출하고, 상기 적어도 하나의 키워드에 기초하여 새로운 추론 결과(즉, 데이터 처리 레벨 1에 대응하는 추론 결과)를 획득할 수 있다. 이때 새로운 추론 결과를 획득하기 위해 사용하는 추론 모델은 현재 사용자가 사용하는 추론 모델이거나, 시스템 설정, 응용 프로그램 설정 또는 사용자 입력에 기초하여 선택된 추론 모델이 사용될 수 있으나 이에 제한되지 않는다. 다음으로, 데이터 처리 레벨 1에 대응하는 추론 결과에서 적어도 하나의 키워드를 추출하고, 상기 적어도 하나의 키워드에 기초하여 또 다른 새로운 추론 결과(즉, 데이터 처리 레벨 2에 대응하는 추론 결과)를 획득할 수 있다.
일 실시예에 따르면, 데이터 처리 레벨에 기초하여 재귀적 추론 과정이 수행될 수 있다. 이를 통해 확장된 추론 결과를 획득할 수 있고, 확장된 추론 결과를 추론 데이터베이스에 갱신할 수 있다. 예를 들어, 추론 능력이 상대적으로 떨어지는 온 디바이스 AI를 통해 추론 결과를 확보한 경우라도, 데이터 처리 레벨에 기초한 재귀적 추론 과정을 통해 확장된 추론 결과를 획득할 수 있다.
도 6a 및 6b는 본 개시의 일 실시예에 따른 전자 장치에서의 추론 컨텍스트 확장 방법의 개략적인 흐름도이다.
일 실시예에 따르면, 상기 전자 장치는 도 4에 도시한 전자 장치(400)에 상응하는 전자 장치일 수 있다. 도 6a 및 6b에 설명되는 전자 장치의 동작에 있어서, 도 4에서 설명된 부분과 중복되는 부분은 생략될 수 있다. 도 6a 및 6b에 도시되는 동작들의 일부가 생략될 수 있으며, 도 6a 및 6b에 도시되지 않은 동작이 추가될 수 있다.
도 6a를 참조하면, 일 실시예에 따른 동작 611에서, 전자 장치(400)는 사용자 입력에 기초하여 사용자 입력 프롬프트를 획득할 수 있다.
일 실시예에 따른 동작 612에서, 전자 장치(400)는 적어도 하나의 추론 모델 중 추론 모델을 결정할 수 있다. 전자 장치(400)는 추론 모델의 품질 및 전자 장치(400)의 상태 정보에 기초하여 상기 적어도 하나의 추론 모델 중 추론 모델을 결정할 수 있다. 상기 추론 모델의 품질은 상기 추론 모델의 사이즈 및 추론 결과 벤치마킹 점수 중 적어도 하나에 기초하여 결정될 수 있다. 전자 장치(400)의 상태 정보는 네트워크 상태 및 리소스 상태 중 적어도 하나를 포함할 수 있다. 상기 리소스 상태는 메모리 가용량, 전자 장치(400)의 적어도 하나의 프로세서의 성능 및 배터리 가용량 중 적어도 하나를 포함할 수 있다.
일 실시예에 따른 동작 613에서, 전자 장치(400)는 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정할 수 있다. 전자 장치는 사용자 입력 프롬프트를 임베딩 벡터로 변경함으로써 컨텍스트 벡터를 결정할 수 있다. 상기 컨텍스트 벡터는 사용자 입력 프롬프트를 가장 잘 설명할 수 있는 사용자 입력 프롬프트 내의 단어나 단어들의 조합, 사용자 입력 프롬프트를 요약한 문장이나 구 등에 대응되는 임베딩 벡터를 포함할 수 있다.
일 실시예에 따른 동작 614에서, 전자 장치(400)는 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정할 수 있다. 전자 장치(400)는 상기 사용자 입력 프롬프트에 기초하여 적어도 하나의 키워드를 추출하고 임베딩 벡터로 변경함으로써 적어도 하나의 참조 벡터를 결정할 수 있다.
일 실시예에 따른 동작 615에서, 전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색할 수 있다. 상기 전자 장치는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 소정 유사도(또는 상관도 기준값) 이상인 관련 데이터를 검색할 수 있다. 전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터 각각의 임베딩 벡터와 소정 유사도 이상의 임베딩 벡터를 갖는 컨텍스트 데이터 및 참조 데이터를 각각 검색할 수 있다.
도 6b를 참조하면, 일 실시예에 따른 동작 621에서, 전자 장치(400)는 추론 데이터베이스에서 관련 데이터를 검색한 결과가 존재하는지 여부를 판단할 수 있다. 관련 데이터가 검색된 경우 동작 622로 이동하고, 관련 데이터가 검색되지 않은 경우 동작 625로 이동할 수 있다.
일 실시예에 따른 동작 622에서, 전자 장치(400)는 관련 데이터가 추론 모델의 입력 범위를 초과하는지 여부를 판단할 수 있다. 추론 모델의 입력 범위를 초과한 경우 동작 623으로 이동하고, 추론 모델의 입력 범위를 초과하지 않은 경우 동작 624로 이동할 수 있다.
일 실시예에 따른 동작 623에서, 전자 장치(400)는 추론 모델이 처리할 수 있는 통합 프롬프트의 크기를 고려하여, 관련 데이터를 축소할 수 있다. 관련 데이터를 축소하는 방법은 사용자 입력 프롬프트와 소정 유사도 이상인 관련 데이터를 선택하거나, 상기 관련 데이터를 요약하거나, 사용자 입력에 기초하여 상기 관련 데이터 중 일부를 선택할 수 있으나, 이에 제한되지 않는다.
일 실시예에 따른 동작 624에서, 전자 장치(400)는 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성할 수 있다.
일 실시예에 따른 동작 625에서, 전자 장치(400)는 사용자 입력 프롬프트 또는 통합 프롬프트를 추론 모델에 전달할 수 있다. 동작 621에서 관련 데이터가 검색되지 않은 경우, 전자 장치(400)는 사용자 입력 프롬프트를 추론 모델에 전달할 수 있다. 동작 621에서 관련 데이터가 검색된 경우, 전자 장치(400)는 통합 프롬프트를 추론 모델에 전달할 수 있다.
일 실시예에 따른 동작 626에서, 전자 장치(400)는 추론 모델로부터 추론 결과를 획득할 수 있다.
일 실시예에 따른 동작 627에서, 전자 장치(400)는 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기초하여 추론 데이터베이스를 갱신할 수 있다. 이 경우, 데이터 처리 레벨에 기초하여, 재귀적 추론 과정을 더 수행할 수 있다. 이를 통해 확장된 추론 결과를 획득할 수 있고, 확장된 추론 결과를 추론 데이터베이스에 갱신할 수 있다. 데이터 처리 레벨에 기초한 확장된 추론 결과 획득 및 추론 데이터베이스 갱신은 이하 도 7 및 도 8을 참조하여 후술한다.
도 7은 본 개시의 일 실시예에 따른 전자 장치에서의 데이터 처리 레벨이 0인 경우 추론 데이터베이스 저장 및 갱신 방법의 개략적인 흐름도이다. 도 8은 본 개시의 일 실시예에 따른 전자 장치에서의 데이터 처리 레벨이 1이상인 경우 데이터 확장을 통한 추론 데이터베이스 저장 및 갱신 방법의 개략적인 흐름도이다.
일 실시예에 따르면, 상기 전자 장치는 도 4에 도시한 전자 장치(400)에 상응하는 전자 장치일 수 있다. 도 7 및 도 8에 설명되는 전자 장치의 동작에 있어서, 도 4에서 설명된 부분과 중복되는 부분은 생략될 수 있다. 도 7 및 도 8에 도시되는 동작들의 일부가 생략될 수 있으며, 도 7 및 도 8에 도시되지 않은 동작이 추가될 수 있다.
도 7을 참조하면, 데이터 처리 레벨이 0인 경우, 사용자 입력 프롬프트에 대한 임베딩 벡터를 획득하고, 사용자 입력 프롬프트, 상기 임베딩 벡터 및 추론 모델로부터 획득한 추론 결과를 포함하는 하나의 데이터베이스 레코드를 추론 데이터베이스에 저장할 수 있다.
일 실시예에 따른 동작 710에서, 전자 장치(400)는 사용자 입력 프롬프트에 대한 임베딩 벡터를 획득할 수 있다.
일 실시예에 따른 동작 720에서, 전자 장치(400)는 상기 임베딩 벡터에 기초하여 추론 데이터베이스에 동일한 데이터베이스 레코드가 존재하는지 여부를 판단할 수 있다. 동일한 데이터베이스 레코드가 존재하는 경우 동작 730으로 이동할 수 있고, 동일한 데이터베이스 레코드가 존재하지 않는 경우 동작 750으로 이동할 수 있다.
일 실시예에 따른 동작 730에서, 전자 장치(400)는 추론 결과 및 기존 데이터베이스 레코드를 후처리할 수 있다. 상기 임베딩 벡터와 소정 유사도 이상 일치하는 기존 데이터베이스 레코드가 존재하는 경우, 기존 추론 결과와 새로 갱신할 추론 결과를 하나로 합병할 수 있다. 이때 합병된 추론 결과의 중복을 제거하기 위해 소정 추론 모델(예. 대규모 언어 모델)에게 합병된 추론 결과를 전달하고 의미상 중복이 제거된 추론 결과를 획득할 수 있다. 또는, 합병된 추론 결과를 적어도 하나의 키워드로 나누어 임베딩 벡터를 결정한 후, 임베딩 벡터의 유사도에 기초하여 합병된 추론 결과에서 의미상 중복이 제거된 추론 결과를 획득할 수 있다. 이에 제한되지 않고 중복을 제거하기 위한 다양한 후처리 방법이 있을 수 있음은 당업자에게 이해될 것이다.
일 실시예에 따른 동작 740에서, 전자 장치(400)는 사용자 입력 프롬프트, 임베딩 벡터, 후처리 결과를 하나의 데이터베이스 레코드로 추론 데이터베이스에 갱신할 수 있다.
일 실시예에 따른 동작 750에서, 전자 장치(400)는 사용자 입력 프롬프트, 임베딩 벡터, 추론 결과를 하나의 데이터베이스 레코드로 추론 데이터베이스에 저장할 수 있다.
도 8을 참조하면, 전자 장치(400)는 데이터 처리 레벨에 기초하여 추론 결과를 확장하고, 확장된 추론 결과를 추론 데이터베이스에 갱신할 수 있다. 상기 데이터 처리 레벨은 추론 결과를 추론 데이터베이스에 저장할 때 현재의 추론 결과를 얼마나 확장시켜 저장할지를 결정하는 레벨일 수 있다. 데이터 처리 레벨이 높을수록, 그에 비례하는 재귀적 횟수의 추론 과정을 통해 더 확장된 추론 결과를 획득할 수 있다.
일 실시예에 따른 동작 810에서, 전자 장치(400)는 사용자 입력 프롬프트 및 추론 결과를 데이터 처리 레벨에 기초한 소정 단위의 적어도 하나의 텍스트로 구분할 수 있다.
일 실시예에 따른 동작 820에서, 전자 장치(400)는 소정 추론 모델로부터 적어도 하나의 텍스트 각각에 기초한 각각의 제2 추론 결과를 획득할 수 있다.
일 실시예에 따른 동작 830에서, 전자 장치(400)는 적어도 하나의 텍스트를 적어도 하나의 임베딩 벡터로 각각 변환할 수 있다.
일 실시예에 따른 동작 840에서, 전자 장치(400)는 상기 임베딩 벡터에 기초하여 추론 데이터베이스에 동일한 데이터베이스 레코드가 존재하는지 여부를 판단할 수 있다. 동일한 데이터베이스 레코드가 존재하는 경우 동작 850으로 이동할 수 있고, 동일한 데이터베이스 레코드가 존재하지 않는 경우 동작 870으로 이동할 수 있다.
일 실시예에 따른 동작 850에서, 전자 장치(400)는 제2 추론 결과 및 기존 데이터베이스 레코드를 후처리할 수 있다. 상기 임베딩 벡터와 소정 유사도 이상 일치하는 기존 데이터베이스 레코드가 존재하는 경우, 기존 추론 결과와 새로 갱신할 상기 제2 추론 결과를 하나로 합병할 수 있다. 이때 합병된 추론 결과의 중복을 제거하기 위해 소정 추론 모델(예. 대규모 언어 모델)에게 합병된 추론 결과를 전달하고 의미상 중복이 제거된 추론 결과를 획득할 수 있다. 또는, 합병된 추론 결과를 적어도 하나의 키워드로 나누어 임베딩 벡터를 결정한 후, 임베딩 벡터의 유사도에 기초하여 합병된 추론 결과에서 의미상 중복이 제거된 추론 결과를 획득할 수 있다. 이에 제한되지 않고 중복을 제거하기 위한 다양한 후처리 방법이 있을 수 있음은 당업자에게 이해될 것이다.
일 실시예에 따른 동작 860에서, 전자 장치(400)는 텍스트, 임베딩 벡터, 후처리 결과를 추론 데이터베이스에 갱신할 수 있다.
일 실시예에 따른 동작 870에서, 전자 장치(400)는 텍스트, 임베딩 벡터, 상기 제2 추론 결과를 추론 데이터베이스에 저장할 수 있다.
동작 810 내지 동작 870은 데이터 처리 레벨을 1씩 증가하면서 상기 데이터 처리 레벨까지 재귀적으로 반복될 수 있고, 이를 통해 더 확장된 추론 결과를 획득할 수 있고 추론 데이터베이스에 저장할 수 있다.
도 9a 및 도 9b는 본 개시의 일 실시예에 따른 추론 데이터베이스에서 검색된 관련 데이터 중 사용자가 선택한 일부에 기초한 추론 결과를 디스플레이하는 UI 화면을 도시한다.
일 실시예에 따르면, 전자 장치(400)는 사용자 입력에 기초하여 사용자 입력 프롬프트를 획득할 수 있다.
일 실시예에 따르면, 전자 장치(400)는 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정할 수 있다. 전자 장치(400)는 사용자 입력 프롬프트를 임베딩 벡터로 변경함으로써 컨텍스트 벡터를 결정할 수 있다. 전자 장치(400)는 상기 사용자 입력 프롬프트에 기초하여 적어도 하나의 키워드를 추출하고 임베딩 벡터로 변경함으로써 적어도 하나의 참조 벡터를 결정할 수 있다.
전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스(122)에서 관련 데이터를 검색할 수 있다. 전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 소정 유사도(또는 상관도 기준값) 이상인 관련 데이터를 검색할 수 있다. 전자 장치(400)는 상기 컨텍스트 벡터 및 상기 참조 벡터 각각의 임베딩 벡터와 소정 유사도 이상의 임베딩 벡터를 갖는 컨텍스트 데이터 및 참조 데이터를 각각 검색할 수 있다. 따라서, 상기 관련 데이터는 추론 데이터베이스(122)에서, 상기 컨텍스트 벡터에 기초하여 검색된 컨텍스트 데이터 및 상기 참조 벡터에 기초하여 검색된 참조 데이터를 적어도 하나 포함할 수 있다.
도 9a를 참조하면, 추론 데이터베이스(122)에서 관련 데이터가 검색되면, 전자 장치(400)는 관련 데이터(922)가 검색되었음을 지시하는 소정 아이콘(911)을 디스플레이할 수 있다. 사용자가 소정 아이콘(911)을 클릭하는 경우, 전자 장치(400)는 관련 데이터(922) 각각의 항목 및 선택 컨트롤러 GUI 리소스(예. 체크 박스)를 포함하여 구성된 UI(922)를 디스플레이할 수 있다. 전자 장치(400)는 사용자 입력 기초하여 선택된 관련 데이터 및 사용자 입력 프롬프트를 포함하는 통합 프롬프트를 생성할 수 있다.
도 9b를 참조하면, 전자 장치(400)는 상기 생성된 통합 프롬프트에 기초하여 추론 모델로부터 추론 결과를 획득하고 디스플레이할 수 있다(930).
도시되지는 않았으나 일 실시예에 따르면, 전자 장치(400)는 사용자 계정과 연동된 사용자 정보를 추가하여 통합 프롬프트를 생성할 수 있다. 전자 장치(400)는 사용자 정보가 추가된 통합 프롬프트에 기초하여 추론 모델로부터 사용자 정보에 기반한 맞춤형 추론 결과를 획득하고 디스플레이할 수 있다.
일 실시예에 따르면, 전자 장치(400)는 각각의 대화 세션에서 생성되는 사용자 프로필을 추가하여 통합 프롬프트를 생성할 수 있다. 전자 장치(400)는 사용자 프로필이 추가된 통합 프롬프트에 기초하여 추론 모델로부터 단일 세션 내에서 축적된 이전 대화에 기반한 추론 결과를 획득하고 디스플레이할 수 있다.
도 10은 다양한 실시예들에 따른, 네트워크 환경 내의 전자 장치의 블록도이다.
도 10은, 다양한 실시예들에 따른, 네트워크 환경(1000) 내의 전자 장치(1001)의 블록도이다. 도 10을 참조하면, 네트워크 환경(1000)에서 전자 장치(1001)는 제 1 네트워크(1098)(예: 근거리 무선 통신 네트워크)를 통하여 전자 장치(1002)와 통신하거나, 또는 제 2 네트워크(1099)(예: 원거리 무선 통신 네트워크)를 통하여 전자 장치(1004) 또는 서버(1008) 중 적어도 하나와 통신할 수 있다. 일실시예에 따르면, 전자 장치(1001)는 서버(1008)를 통하여 전자 장치(1004)와 통신할 수 있다. 일실시예에 따르면, 전자 장치(1001)는 프로세서(1020), 메모리(1030), 입력 모듈(1050), 음향 출력 모듈(1055), 디스플레이 모듈(1060), 오디오 모듈(1070), 센서 모듈(1076), 인터페이스(1077), 연결 단자(1078), 햅틱 모듈(1079), 카메라 모듈(1080), 전력 관리 모듈(1088), 배터리(1089), 통신 모듈(1090), 가입자 식별 모듈(1096), 또는 안테나 모듈(1097)을 포함할 수 있다. 어떤 실시예에서는, 전자 장치(1001)에는, 이 구성요소들 중 적어도 하나(예: 연결 단자(1078))가 생략되거나, 하나 이상의 다른 구성요소가 추가될 수 있다. 어떤 실시예에서는, 이 구성요소들 중 일부들(예: 센서 모듈(1076), 카메라 모듈(1080), 또는 안테나 모듈(1097))은 하나의 구성요소(예: 디스플레이 모듈(1060))로 통합될 수 있다.
프로세서(1020)는, 예를 들면, 소프트웨어(예: 프로그램(1040))를 실행하여 프로세서(1020)에 연결된 전자 장치(1001)의 적어도 하나의 다른 구성요소(예: 하드웨어 또는 소프트웨어 구성요소)를 제어할 수 있고, 다양한 데이터 처리 또는 연산을 수행할 수 있다. 일실시예에 따르면, 데이터 처리 또는 연산의 적어도 일부로서, 프로세서(1020)는 다른 구성요소(예: 센서 모듈(1076) 또는 통신 모듈(1090))로부터 수신된 명령 또는 데이터를 휘발성 메모리(1032)에 저장하고, 휘발성 메모리(1032)에 저장된 명령 또는 데이터를 처리하고, 결과 데이터를 비휘발성 메모리(1034)에 저장할 수 있다. 일실시예에 따르면, 프로세서(1020)는 메인 프로세서(1021)(예: 중앙 처리 장치 또는 어플리케이션 프로세서) 또는 이와는 독립적으로 또는 함께 운영 가능한 보조 프로세서(1023)(예: 그래픽 처리 장치, 신경망 처리 장치(NPU: neural processing unit), 이미지 시그널 프로세서, 센서 허브 프로세서, 또는 커뮤니케이션 프로세서)를 포함할 수 있다. 예를 들어, 전자 장치(1001)가 메인 프로세서(1021) 및 보조 프로세서(1023)를 포함하는 경우, 보조 프로세서(1023)는 메인 프로세서(1021)보다 저전력을 사용하거나, 지정된 기능에 특화되도록 설정될 수 있다. 보조 프로세서(1023)는 메인 프로세서(1021)와 별개로, 또는 그 일부로서 구현될 수 있다.
보조 프로세서(1023)는, 예를 들면, 메인 프로세서(1021)가 인액티브(예: 슬립) 상태에 있는 동안 메인 프로세서(1021)를 대신하여, 또는 메인 프로세서(1021)가 액티브(예: 어플리케이션 실행) 상태에 있는 동안 메인 프로세서(1021)와 함께, 전자 장치(1001)의 구성요소들 중 적어도 하나의 구성요소(예: 디스플레이 모듈(1060), 센서 모듈(1076), 또는 통신 모듈(1090))와 관련된 기능 또는 상태들의 적어도 일부를 제어할 수 있다. 일실시예에 따르면, 보조 프로세서(1023)(예: 이미지 시그널 프로세서 또는 커뮤니케이션 프로세서)는 기능적으로 관련 있는 다른 구성요소(예: 카메라 모듈(1080) 또는 통신 모듈(1090))의 일부로서 구현될 수 있다. 일실시예에 따르면, 보조 프로세서(1023)(예: 신경망 처리 장치)는 인공지능 모델의 처리에 특화된 하드웨어 구조를 포함할 수 있다. 인공지능 모델은 기계 학습을 통해 생성될 수 있다. 이러한 학습은, 예를 들어, 인공지능 모델이 수행되는 전자 장치(1001) 자체에서 수행될 수 있고, 별도의 서버(예: 서버(1008))를 통해 수행될 수도 있다. 학습 알고리즘은, 예를 들어, 지도형 학습(supervised learning), 비지도형 학습(unsupervised learning), 준지도형 학습(semi-supervised learning) 또는 강화 학습(reinforcement learning)을 포함할 수 있으나, 전술한 예에 한정되지 않는다. 인공지능 모델은, 복수의 인공 신경망 레이어들을 포함할 수 있다. 인공 신경망은 심층 신경망(DNN: deep neural network), CNN(convolutional neural network), RNN(recurrent neural network), RBM(restricted boltzmann machine), DBN(deep belief network), BRDNN(bidirectional recurrent deep neural network), 심층 Q-네트워크(deep Q-networks) 또는 상기 중 둘 이상의 조합 중 하나일 수 있으나, 전술한 예에 한정되지 않는다. 인공지능 모델은 하드웨어 구조 이외에, 추가적으로 또는 대체적으로, 소프트웨어 구조를 포함할 수 있다.
메모리(1030)는, 전자 장치(1001)의 적어도 하나의 구성요소(예: 프로세서(1020) 또는 센서 모듈(1076))에 의해 사용되는 다양한 데이터를 저장할 수 있다. 데이터는, 예를 들어, 소프트웨어(예: 프로그램(1040)) 및, 이와 관련된 명령에 대한 입력 데이터 또는 출력 데이터를 포함할 수 있다. 메모리(1030)는, 휘발성 메모리(1032) 또는 비휘발성 메모리(1034)를 포함할 수 있다.
프로그램(1040)은 메모리(1030)에 소프트웨어로서 저장될 수 있으며, 예를 들면, 운영 체제(1042), 미들 웨어(1044) 또는 어플리케이션(1046)을 포함할 수 있다.
입력 모듈(1050)은, 전자 장치(1001)의 구성요소(예: 프로세서(1020))에 사용될 명령 또는 데이터를 전자 장치(1001)의 외부(예: 사용자)로부터 수신할 수 있다. 입력 모듈(1050)은, 예를 들면, 마이크, 마우스, 키보드, 키(예: 버튼), 또는 디지털 펜(예: 스타일러스 펜)을 포함할 수 있다.
음향 출력 모듈(1055)은 음향 신호를 전자 장치(1001)의 외부로 출력할 수 있다. 음향 출력 모듈(1055)은, 예를 들면, 스피커 또는 리시버를 포함할 수 있다. 스피커는 멀티미디어 재생 또는 녹음 재생과 같이 일반적인 용도로 사용될 수 있다. 리시버는 착신 전화를 수신하기 위해 사용될 수 있다. 일실시예에 따르면, 리시버는 스피커와 별개로, 또는 그 일부로서 구현될 수 있다.
디스플레이 모듈(1060)은 전자 장치(1001)의 외부(예: 사용자)로 정보를 시각적으로 제공할 수 있다. 디스플레이 모듈(1060)은, 예를 들면, 디스플레이, 홀로그램 장치, 또는 프로젝터 및 해당 장치를 제어하기 위한 제어 회로를 포함할 수 있다. 일실시예에 따르면, 디스플레이 모듈(1060)은 터치를 감지하도록 설정된 터치 센서, 또는 상기 터치에 의해 발생되는 힘의 세기를 측정하도록 설정된 압력 센서를 포함할 수 있다.
오디오 모듈(1070)은 소리를 전기 신호로 변환시키거나, 반대로 전기 신호를 소리로 변환시킬 수 있다. 일실시예에 따르면, 오디오 모듈(1070)은, 입력 모듈(1050)을 통해 소리를 획득하거나, 음향 출력 모듈(1055), 또는 전자 장치(1001)와 직접 또는 무선으로 연결된 외부 전자 장치(예: 전자 장치(1002))(예: 스피커 또는 헤드폰)를 통해 소리를 출력할 수 있다.
센서 모듈(1076)은 전자 장치(1001)의 작동 상태(예: 전력 또는 온도), 또는 외부의 환경 상태(예: 사용자 상태)를 감지하고, 감지된 상태에 대응하는 전기 신호 또는 데이터 값을 생성할 수 있다. 일실시예에 따르면, 센서 모듈(1076)은, 예를 들면, 제스처 센서, 자이로 센서, 기압 센서, 마그네틱 센서, 가속도 센서, 그립 센서, 근접 센서, 컬러 센서, IR(infrared) 센서, 생체 센서, 온도 센서, 습도 센서, 또는 조도 센서를 포함할 수 있다.
인터페이스(1077)는 전자 장치(1001)가 외부 전자 장치(예: 전자 장치(1002))와 직접 또는 무선으로 연결되기 위해 사용될 수 있는 하나 이상의 지정된 프로토콜들을 지원할 수 있다. 일실시예에 따르면, 인터페이스(1077)는, 예를 들면, HDMI(high definition multimedia interface), USB(universal serial bus) 인터페이스, SD카드 인터페이스, 또는 오디오 인터페이스를 포함할 수 있다.
연결 단자(1078)는, 그를 통해서 전자 장치(1001)가 외부 전자 장치(예: 전자 장치(1002))와 물리적으로 연결될 수 있는 커넥터를 포함할 수 있다. 일실시예에 따르면, 연결 단자(1078)는, 예를 들면, HDMI 커넥터, USB 커넥터, SD 카드 커넥터, 또는 오디오 커넥터(예: 헤드폰 커넥터)를 포함할 수 있다.
햅틱 모듈(1079)은 전기적 신호를 사용자가 촉각 또는 운동 감각을 통해서 인지할 수 있는 기계적인 자극(예: 진동 또는 움직임) 또는 전기적인 자극으로 변환할 수 있다. 일실시예에 따르면, 햅틱 모듈(1079)은, 예를 들면, 모터, 압전 소자, 또는 전기 자극 장치를 포함할 수 있다.
카메라 모듈(1080)은 정지 영상 및 동영상을 촬영할 수 있다. 일실시예에 따르면, 카메라 모듈(1080)은 하나 이상의 렌즈들, 이미지 센서들, 이미지 시그널 프로세서들, 또는 플래시들을 포함할 수 있다.
전력 관리 모듈(1088)은 전자 장치(1001)에 공급되는 전력을 관리할 수 있다. 일실시예에 따르면, 전력 관리 모듈(1088)은, 예를 들면, PMIC(power management integrated circuit)의 적어도 일부로서 구현될 수 있다.
배터리(1089)는 전자 장치(1001)의 적어도 하나의 구성요소에 전력을 공급할 수 있다. 일실시예에 따르면, 배터리(1089)는, 예를 들면, 재충전 불가능한 1차 전지, 재충전 가능한 2차 전지 또는 연료 전지를 포함할 수 있다.
통신 모듈(1090)은 전자 장치(1001)와 외부 전자 장치(예: 전자 장치(1002), 전자 장치(1004), 또는 서버(1008)) 간의 직접(예: 유선) 통신 채널 또는 무선 통신 채널의 수립, 및 수립된 통신 채널을 통한 통신 수행을 지원할 수 있다. 통신 모듈(1090)은 프로세서(1020)(예: 어플리케이션 프로세서)와 독립적으로 운영되고, 직접(예: 유선) 통신 또는 무선 통신을 지원하는 하나 이상의 커뮤니케이션 프로세서를 포함할 수 있다. 일실시예에 따르면, 통신 모듈(1090)은 무선 통신 모듈(1092)(예: 셀룰러 통신 모듈, 근거리 무선 통신 모듈, 또는 GNSS(global navigation satellite system) 통신 모듈) 또는 유선 통신 모듈(1094)(예: LAN(local area network) 통신 모듈, 또는 전력선 통신 모듈)을 포함할 수 있다. 이들 통신 모듈 중 해당하는 통신 모듈은 제 1 네트워크(1098)(예: 블루투스, WiFi(wireless fidelity) direct 또는 IrDA(infrared data association)와 같은 근거리 통신 네트워크) 또는 제 2 네트워크(1099)(예: 레거시 셀룰러 네트워크, 5G 네트워크, 차세대 통신 네트워크, 인터넷, 또는 컴퓨터 네트워크(예: LAN 또는 WAN)와 같은 원거리 통신 네트워크)를 통하여 외부의 전자 장치(1004)와 통신할 수 있다. 이런 여러 종류의 통신 모듈들은 하나의 구성요소(예: 단일 칩)로 통합되거나, 또는 서로 별도의 복수의 구성요소들(예: 복수 칩들)로 구현될 수 있다. 무선 통신 모듈(1092)은 가입자 식별 모듈(1096)에 저장된 가입자 정보(예: 국제 모바일 가입자 식별자(IMSI))를 이용하여 제 1 네트워크(1098) 또는 제 2 네트워크(1099)와 같은 통신 네트워크 내에서 전자 장치(1001)를 확인 또는 인증할 수 있다.
무선 통신 모듈(1092)은 4G 네트워크 이후의 5G 네트워크 및 차세대 통신 기술, 예를 들어, NR 접속 기술(new radio access technology)을 지원할 수 있다. NR 접속 기술은 고용량 데이터의 고속 전송(eMBB(enhanced mobile broadband)), 단말 전력 최소화와 다수 단말의 접속(mMTC(massive machine type communications)), 또는 고신뢰도와 저지연(URLLC(ultra-reliable and low-latency communications))을 지원할 수 있다. 무선 통신 모듈(1092)은, 예를 들어, 높은 데이터 전송률 달성을 위해, 고주파 대역(예: mmWave 대역)을 지원할 수 있다. 무선 통신 모듈(1092)은 고주파 대역에서의 성능 확보를 위한 다양한 기술들, 예를 들어, 빔포밍(beamforming), 거대 배열 다중 입출력(massive MIMO(multiple-input and multiple-output)), 전차원 다중입출력(FD-MIMO: full dimensional MIMO), 어레이 안테나(array antenna), 아날로그 빔형성(analog beam-forming), 또는 대규모 안테나(large scale antenna)와 같은 기술들을 지원할 수 있다. 무선 통신 모듈(1092)은 전자 장치(1001), 외부 전자 장치(예: 전자 장치(1004)) 또는 네트워크 시스템(예: 제 2 네트워크(1099))에 규정되는 다양한 요구사항을 지원할 수 있다. 일실시예에 따르면, 무선 통신 모듈(1092)은 eMBB 실현을 위한 Peak data rate(예: 20Gbps 이상), mMTC 실현을 위한 손실 Coverage(예: 164dB 이하), 또는 URLLC 실현을 위한 U-plane latency(예: 다운링크(DL) 및 업링크(UL) 각각 0.5ms 이하, 또는 라운드 트립 1ms 이하)를 지원할 수 있다.
안테나 모듈(1097)은 신호 또는 전력을 외부(예: 외부의 전자 장치)로 송신하거나 외부로부터 수신할 수 있다. 일실시예에 따르면, 안테나 모듈(1097)은 서브스트레이트(예: PCB) 위에 형성된 도전체 또는 도전성 패턴으로 이루어진 방사체를 포함하는 안테나를 포함할 수 있다. 일실시예에 따르면, 안테나 모듈(1097)은 복수의 안테나들(예: 어레이 안테나)을 포함할 수 있다. 이런 경우, 제 1 네트워크(1098) 또는 제 2 네트워크(1099)와 같은 통신 네트워크에서 사용되는 통신 방식에 적합한 적어도 하나의 안테나가, 예를 들면, 통신 모듈(1090)에 의하여 상기 복수의 안테나들로부터 선택될 수 있다. 신호 또는 전력은 상기 선택된 적어도 하나의 안테나를 통하여 통신 모듈(1090)과 외부의 전자 장치 간에 송신되거나 수신될 수 있다. 어떤 실시예에 따르면, 방사체 이외에 다른 부품(예: RFIC(radio frequency integrated circuit))이 추가로 안테나 모듈(1097)의 일부로 형성될 수 있다.
다양한 실시예에 따르면, 안테나 모듈(1097)은 mmWave 안테나 모듈을 형성할 수 있다. 일실시예에 따르면, mmWave 안테나 모듈은 인쇄 회로 기판, 상기 인쇄 회로 기판의 제 1 면(예: 아래 면)에 또는 그에 인접하여 배치되고 지정된 고주파 대역(예: mmWave 대역)을 지원할 수 있는 RFIC, 및 상기 인쇄 회로 기판의 제 2 면(예: 윗 면 또는 측 면)에 또는 그에 인접하여 배치되고 상기 지정된 고주파 대역의 신호를 송신 또는 수신할 수 있는 복수의 안테나들(예: 어레이 안테나)을 포함할 수 있다.
상기 구성요소들 중 적어도 일부는 주변 기기들간 통신 방식(예: 버스, GPIO(general purpose input and output), SPI(serial peripheral interface), 또는 MIPI(mobile industry processor interface))을 통해 서로 연결되고 신호(예: 명령 또는 데이터)를 상호간에 교환할 수 있다.
일실시예에 따르면, 명령 또는 데이터는 제 2 네트워크(1099)에 연결된 서버(1008)를 통해서 전자 장치(1001)와 외부의 전자 장치(1004)간에 송신 또는 수신될 수 있다. 외부의 전자 장치(1002, 또는 1004) 각각은 전자 장치(1001)와 동일한 또는 다른 종류의 장치일 수 있다. 일실시예에 따르면, 전자 장치(1001)에서 실행되는 동작들의 전부 또는 일부는 외부의 전자 장치들(1002, 1004, 또는 1008) 중 하나 이상의 외부의 전자 장치들에서 실행될 수 있다. 예를 들면, 전자 장치(1001)가 어떤 기능이나 서비스를 자동으로, 또는 사용자 또는 다른 장치로부터의 요청에 반응하여 수행해야 할 경우에, 전자 장치(1001)는 기능 또는 서비스를 자체적으로 실행시키는 대신에 또는 추가적으로, 하나 이상의 외부의 전자 장치들에게 그 기능 또는 그 서비스의 적어도 일부를 수행하라고 요청할 수 있다. 상기 요청을 수신한 하나 이상의 외부의 전자 장치들은 요청된 기능 또는 서비스의 적어도 일부, 또는 상기 요청과 관련된 추가 기능 또는 서비스를 실행하고, 그 실행의 결과를 전자 장치(1001)로 전달할 수 있다. 전자 장치(1001)는 상기 결과를, 그대로 또는 추가적으로 처리하여, 상기 요청에 대한 응답의 적어도 일부로서 제공할 수 있다. 이를 위하여, 예를 들면, 클라우드 컴퓨팅, 분산 컴퓨팅, 모바일 에지 컴퓨팅(MEC: mobile edge computing), 또는 클라이언트-서버 컴퓨팅 기술이 이용될 수 있다. 전자 장치(1001)는, 예를 들어, 분산 컴퓨팅 또는 모바일 에지 컴퓨팅을 이용하여 초저지연 서비스를 제공할 수 있다. 다른 실시예에 있어서, 외부의 전자 장치(1004)는 IoT(internet of things) 기기를 포함할 수 있다. 서버(1008)는 기계 학습 및/또는 신경망을 이용한 지능형 서버일 수 있다. 일실시예에 따르면, 외부의 전자 장치(1004) 또는 서버(1008)는 제 2 네트워크(1099) 내에 포함될 수 있다. 전자 장치(1001)는 5G 통신 기술 및 IoT 관련 기술을 기반으로 지능형 서비스(예: 스마트 홈, 스마트 시티, 스마트 카, 또는 헬스 케어)에 적용될 수 있다.
본 개시의 일 실시예에 따르면 전자 장치는, 통신 회로; 프로세싱 회로를 포함하는 적어도 하나의 프로세서; 및 명령어들을 저장하는 하나 이상의 저장 매체를 포함하는, 메모리를 포함할 수 있다. 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금, 적어도 하나의 추론 모델 중 추론 모델을 결정하고; 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하고; 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하고; 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하고; 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하고; 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하고; 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기초하여 상기 추론 데이터베이스를 갱신하도록 야기할 수 있다.
일 실시예에 따르면, 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 추론 모델의 품질 및 상기 전자 장치의 상태 정보에 기초하여 상기 적어도 하나의 추론 모델 중 추론 모델을 결정하도록 야기할 수 있다. 상기 추론 모델의 품질은 상기 추론 모델의 사이즈 및 추론 결과 벤치마킹 점수 중 적어도 하나에 기초하여 결정되고; 상기 전자 장치의 상태 정보는 네트워크 상태 및 리소스 상태 중 적어도 하나를 포함하고; 상기 리소스 상태는 메모리 가용량, 상기 적어도 하나의 프로세서의 성능 및 배터리 가용량 중 적어도 하나를 포함할 수 있다.
일 실시예에 따르면, 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 상기 결정된 추론 모델의 품질이 소정 기준 이하인 경우, 상기 사용자 입력 프롬프트에서 적어도 하나의 키워드를 추출하고; 상기 적어도 하나의 키워드 각각을 적어도 하나의 임베딩 벡터로 변환하여 상기 적어도 하나의 참조 벡터를 결정하도록 야기할 수 있다.
일 실시예에 따르면, 상기 관련 데이터는 상기 추론 데이터베이스에서, 상기 컨텍스트 벡터에 기초하여 검색된 컨텍스트 데이터 및 상기 참조 벡터에 기초하여 검색된 참조 데이터를 적어도 하나 포함할 수 있다.
일 실시예에 따르면, 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 상기 결정된 추론 모델이 처리할 수 있는 통합 프롬프트의 크기를 고려하여, 상기 관련 데이터 중 상기 사용자 입력 프롬프트와 소정 유사도 이상인 일부를 선택하거나, 상기 관련 데이터를 요약하거나, 사용자 입력에 기초하여 상기 관련 데이터 중 일부를 선택함으로써 상기 통합 프롬프트를 생성하도록 야기할 수 있다.
일 실시예에 따르면, 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 상기 사용자 프롬프트 및 상기 추론 결과를 데이터 처리 레벨에 기초한 소정 단위의 적어도 하나의 텍스트로 구분하고; 상기 적어도 하나의 텍스트 각각에 기초하여 적어도 하나의 추론 모델 중 하나의 추론 모델로부터 각각의 제2 추론 결과를 획득하고; 상기 적어도 하나의 텍스트를 적어도 하나의 임베딩 벡터로 각각 변환하고; 상기 텍스트, 상기 임베딩 벡터, 상기 제2 추론 결과를 상기 추론 데이터베이스에 갱신하도록 야기할 수 있다.
일 실시예에 따르면, 상기 데이터 처리 레벨은 사용자 입력에 의해 설정되거나, 상기 전자 장치에서 사전 설정되거나, 상기 결정된 추론 모델의 품질 및 현재 이용가능한 적어도 하나의 추론 모델의 품질에 기초하여 설정될 수 있다.
일 실시예에 따르면, 적어도 하나의 센서를 더 포함하고; 상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금: 상기 적어도 하나의 센서로부터 시스템 데이터를 획득하고; 상기 시스템 데이터에 대응되는 태그 정보를 더 포함하여 상기 통합 프롬프트를 생성하도록 야기할 수 있다. 상기 시스템 데이터는 위치 데이터, 요일 데이터 및 시간 데이터 중 적어도 하나를 포함하고; 상기 태그 정보는 상기 시스템 데이터에 대응할 수 있다.
또한, 본 개시의 일 실시예에 따르면 전자 장치의 추론 컨텍스트 확장 방법은, 적어도 하나의 추론 모델 중 추론 모델을 결정하는 동작; 사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하는 동작; 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하는 동작; 상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하는 동작; 상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하는 동작; 상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하는 동작; 및 상기 사용자 입력 프롬프트 및 상기 추론 결과에 기반하여 상기 추론 데이터베이스를 갱신하는 동작을 포함할 수 있다.
일 실시예에 따르면 상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하는 동작은 상기 추론 모델의 품질이 소정 기준 이하인 경우, 상기 사용자 입력 프롬프트에서 적어도 하나의 키워드를 추출하는 동작; 및 상기 적어도 하나의 키워드 각각을 적어도 하나의 임베딩 벡터로 변환하여 상기 적어도 하나의 참조 벡터를 결정하는 동작을 포함할 수 있다.
본 문서에 개시된 다양한 실시예들에 따른 전자 장치는 다양한 형태의 장치가 될 수 있다. 전자 장치는, 예를 들면, 디스플레이 장치, 휴대용 통신 장치(예: 스마트폰), 컴퓨터 장치, 휴대용 멀티미디어 장치, 휴대용 의료 기기, 카메라, 웨어러블 장치, 또는 가전 장치를 포함할 수 있다. 본 문서의 실시예에 따른 전자 장치는 전술한 기기들에 한정되지 않는다.
본 문서의 다양한 실시예들 및 이에 사용된 용어들은 본 문서에 기재된 기술적 특징들을 특정한 실시예들로 한정하려는 것이 아니며, 해당 실시예의 다양한 변경, 균등물, 또는 대체물을 포함하는 것으로 이해되어야 한다. 예를 들면, 단수로 표현된 구성요소는 문맥상 명백하게 단수만을 의미하지 않는다면 복수의 구성요소를 포함하는 개념으로 이해되어야 한다. 본 문서에서 사용되는 '및/또는'이라는 용어는, 열거되는 항목들 중 하나 이상의 항목에 의한 임의의 가능한 모든 조합들을 포괄하는 것임이 이해되어야 한다. 본 개시에서 사용되는 '포함하다,' '가지다,' '구성되다' 등의 용어는 본 개시 상에 기재된 특징, 구성 요소, 부분품 또는 이들을 조합한 것이 존재함을 지정하려는 것일 뿐이고, 이러한 용어의 사용에 의해 하나 또는 그 이상의 다른 특징들이나 구성 요소, 부분품 또는 이들을 조합한 것들의 존재 또는 부가 가능성을 배제하려는 것은 아니다. 본 문서에서, "A 또는 B", "A 및 B 중 적어도 하나", "A 또는 B 중 적어도 하나", "A, B 또는 C", "A, B 및 C 중 적어도 하나", 및 "A, B, 또는 C 중 적어도 하나"와 같은 문구들 각각은 그 문구들 중 해당하는 문구에 함께 나열된 항목들 중 어느 하나, 또는 그들의 모든 가능한 조합을 포함할 수 있다. "제 1", "제 2", 또는 "첫째" 또는 "둘째"와 같은 용어들은 단순히 해당 구성요소를 다른 해당 구성요소와 구분하기 위해 사용될 수 있으며, 해당 구성요소들을 다른 측면(예: 중요성 또는 순서)에서 한정하지 않는다.
본 문서의 다양한 실시예들에서 사용된 용어 "~부" 또는 "~모듈"은 하드웨어, 소프트웨어 또는 펌웨어로 구현된 유닛을 포함할 수 있으며, 예를 들면, 로직, 논리 블록, 부품, 또는 회로와 같은 용어와 상호 호환적으로 사용될 수 있다. "~부" 또는 "~모듈"은, 일체로 구성된 부품 또는 하나 또는 그 이상의 기능을 수행하는, 상기 부품의 최소 단위 또는 그 일부가 될 수 있다. 예를 들면, 일 실시예에 따르면, "~부" 또는 "~모듈"은 ASIC(application-specific integrated circuit)의 형태로 구현될 수 있다.
본 문서의 다양한 실시예들에서 사용된 용어 “~할 경우”는 문맥에 따라 “~할 때”, 또는 “~할 시” 또는 “결정하는 것에 응답하여” 또는 “검출하는 것에 응답하여”를 의미하는 것으로 해석될 수 있다. 유사하게, “~라고 결정되는 경우” 또는 “~이 검출되는 경우”는 문맥에 따라 “결정 시” 또는 “결정하는 것에 응답하여”, 또는 “검출 시” 또는 “검출하는 것에 응답하여”를 의미하는 것으로 해석될 수 있다.
본 문서를 통해 설명된 전자 장치(400)에 의해 실행되는 프로그램은 하드웨어 구성요소, 소프트웨어 구성요소, 및/또는 하드웨어 구성요소 및 소프트웨어 구성요소의 조합으로 구현될 수 있다. 프로그램은 컴퓨터로 읽을 수 있는 명령어들을 수행할 수 있는 모든 시스템에 의해 수행될 수 있다.
소프트웨어는 컴퓨터 프로그램(computer program), 코드(code), 명령어(instruction), 또는 이들 중 하나 이상의 조합을 포함할 수 있으며, 원하는 대로 동작하도록 처리 장치를 구성하거나 독립적으로 또는 결합적으로 (collectively) 처리 장치를 명령할 수 있다. 소프트웨어는, 컴퓨터로 읽을 수 있는 저장 매체(computer-readable storage media)에 저장된 명령어를 포함하는 컴퓨터 프로그램으로 구현될 수 있다. 컴퓨터가 읽을 수 있는 저장 매체로는, 예를 들어 마그네틱 저장 매체(예컨대, ROM(Read-Only Memory), RAM(Random-Access Memory), 플로피 디스크, 하드 디스크 등) 및 광학적 판독 매체(예컨대, 시디롬(CD-ROM), 디브이디(DVD: Digital Versatile Disc)) 등이 있다. 컴퓨터가 읽을 수 있는 저장 매체는 네트워크로 연결된 컴퓨터 시스템들에 분산되어, 분산 방식으로 컴퓨터가 판독 가능한 코드가 저장되고 실행될 수 있다. 컴퓨터 프로그램은 어플리케이션 스토어(예: 플레이 스토어TM)를 통해 또는 두 개의 사용자 장치들(예: 스마트 폰들) 간에 직접, 온라인으로 배포(예: 다운로드 또는 업로드)될 수 있다. 온라인 배포의 경우에, 컴퓨터 프로그램 제품의 적어도 일부는 제조사의 서버, 어플리케이션 스토어의 서버, 또는 중계 서버의 메모리와 같은 기기로 읽을 수 있는 저장 매체에 적어도 일시 저장되거나, 임시적으로 생성될 수 있다.
다양한 실시예들에 따르면, 상기 기술한 구성요소들의 각각의 구성요소(예: 모듈 또는 프로그램)는 단수 또는 복수의 개체를 포함할 수 있으며, 복수의 개체 중 일부는 다른 구성요소에 분리 배치될 수도 있다. 다양한 실시예들에 따르면, 전술한 해당 구성요소들 중 하나 이상의 구성요소들 또는 동작들이 생략되거나, 또는 하나 이상의 다른 구성요소들 또는 동작들이 추가될 수 있다. 대체적으로 또는 추가적으로, 복수의 구성요소들(예: 모듈 또는 프로그램)은 하나의 구성요소로 통합될 수 있다. 이런 경우, 통합된 구성요소는 상기 복수의 구성요소들 각각의 구성요소의 하나 이상의 기능들을 상기 통합 이전에 상기 복수의 구성요소들 중 해당 구성요소에 의해 수행되는 것과 동일 또는 유사하게 수행할 수 있다. 다양한 실시예들에 따르면, 모듈, 프로그램 또는 다른 구성요소에 의해 수행되는 동작들은 순차적으로, 병렬적으로, 반복적으로, 또는 휴리스틱하게 실행되거나, 상기 동작들 중 하나 이상이 다른 순서로 실행되거나, 생략되거나, 또는 하나 이상의 다른 동작들이 추가될 수 있다.

Claims (10)

  1. 전자 장치에 있어서,
    통신 회로;
    프로세싱 회로를 포함하는 적어도 하나의 프로세서; 및
    명령어들을 저장하는 하나 이상의 저장 매체를 포함하는, 메모리를 포함하고;
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    적어도 하나의 추론 모델 중 추론 모델을 결정하고;
    사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하고;
    상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하고;
    상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하고;
    상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하고;
    상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하고;
    상기 사용자 입력 프롬프트 및 상기 추론 결과에 기초하여 상기 추론 데이터베이스를 갱신하도록 야기하는, 전자 장치.
  2. 제1항에 있어서,
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    추론 모델의 품질 및 상기 전자 장치의 상태 정보에 기초하여 상기 적어도 하나의 추론 모델 중 추론 모델을 결정하도록 야기하고;
    상기 추론 모델의 품질은 상기 추론 모델의 사이즈 및 추론 결과 벤치마킹 점수 중 적어도 하나에 기초하여 결정되고;
    상기 전자 장치의 상태 정보는 네트워크 상태 및 리소스 상태 중 적어도 하나를 포함하고;
    상기 리소스 상태는 메모리 가용량, 상기 적어도 하나의 프로세서의 성능 및 배터리 가용량 중 적어도 하나를 포함하는, 전자 장치.
  3. 제1항에 있어서,
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    상기 결정된 추론 모델의 품질이 소정 기준 이하인 경우, 상기 사용자 입력 프롬프트에서 적어도 하나의 키워드를 추출하고;
    상기 적어도 하나의 키워드 각각을 적어도 하나의 임베딩 벡터로 변환하여 상기 적어도 하나의 참조 벡터를 결정하도록 야기하는, 전자 장치.
  4. 제1항에 있어서,
    상기 관련 데이터는 상기 추론 데이터베이스에서, 상기 컨텍스트 벡터에 기초하여 검색된 컨텍스트 데이터 및 상기 참조 벡터에 기초하여 검색된 참조 데이터를 적어도 하나 포함하는, 전자 장치.
  5. 제4항에 있어서,
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    상기 결정된 추론 모델이 처리할 수 있는 통합 프롬프트의 크기를 고려하여, 상기 관련 데이터 중 상기 사용자 입력 프롬프트와 소정 유사도 이상인 일부를 선택하거나, 상기 관련 데이터를 요약하거나, 사용자 입력에 기초하여 상기 관련 데이터 중 일부를 선택함으로써 상기 통합 프롬프트를 생성하도록 야기하는, 전자 장치.
  6. 제1항에 있어서,
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    상기 사용자 프롬프트 및 상기 추론 결과를 데이터 처리 레벨에 기초한 소정 단위의 적어도 하나의 텍스트로 구분하고;
    상기 적어도 하나의 텍스트 각각에 기초하여 적어도 하나의 추론 모델 중 하나의 추론 모델로부터 각각의 제2 추론 결과를 획득하고;
    상기 적어도 하나의 텍스트를 적어도 하나의 임베딩 벡터로 각각 변환하고;
    상기 텍스트, 상기 임베딩 벡터, 상기 제2 추론 결과를 상기 추론 데이터베이스에 갱신하도록 야기하는, 전자 장치.
  7. 제6항에 있어서,
    상기 데이터 처리 레벨은 사용자 입력에 의해 설정되거나, 상기 전자 장치에서 사전 설정되거나, 상기 결정된 추론 모델의 품질 및 현재 이용가능한 적어도 하나의 추론 모델의 품질에 기초하여 설정되는, 전자 장치.
  8. 제1항에 있어서,
    적어도 하나의 센서를 더 포함하고;
    상기 명령어들이 개별적으로 또는 집합적으로 상기 적어도 하나의 프로세서에 의해 실행될 때, 상기 전자 장치로 하여금:
    상기 적어도 하나의 센서로부터 시스템 데이터를 획득하고;
    상기 시스템 데이터에 대응되는 태그 정보를 더 포함하여 상기 통합 프롬프트를 생성하도록 야기하고;
    상기 시스템 데이터는 위치 데이터, 요일 데이터 및 시간 데이터 중 적어도 하나를 포함하고;
    상기 태그 정보는 상기 시스템 데이터에 대응하는 텍스트 정보인, 전자 장치.
  9. 전자 장치의 추론 컨텍스트 확장 방법에 있어서,
    적어도 하나의 추론 모델 중 추론 모델을 결정하는 동작;
    사용자 입력 프롬프트에 기초하여 컨텍스트 벡터를 결정하는 동작;
    상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하는 동작;
    상기 컨텍스트 벡터 및 상기 참조 벡터에 기초하여 추론 데이터베이스에서 관련 데이터를 검색하는 동작;
    상기 사용자 입력 프롬프트 및 상기 관련 데이터를 포함하는 통합 프롬프트를 생성하는 동작;
    상기 통합 프롬프트에 기초하여 상기 결정된 추론 모델로부터 추론 결과를 획득하는 동작; 및
    상기 사용자 입력 프롬프트 및 상기 추론 결과에 기반하여 상기 추론 데이터베이스를 갱신하는 동작을 포함하는, 방법.
  10. 제9항에 있어서,
    상기 결정된 추론 모델의 품질에 기초하여 선택적으로, 사용자 입력 프롬프트에 기초하여 적어도 하나의 참조 벡터를 결정하는 동작은
    상기 추론 모델의 품질이 소정 기준 이하인 경우, 상기 사용자 입력 프롬프트에서 적어도 하나의 키워드를 추출하는 동작; 및
    상기 적어도 하나의 키워드 각각을 적어도 하나의 임베딩 벡터로 변환하여 상기 적어도 하나의 참조 벡터를 결정하는 동작을 포함하는, 방법.
PCT/KR2025/005923 2024-06-26 2025-04-30 추론 컨텍스트 확장 방법 및 장치 Pending WO2026005261A1 (ko)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
KR20240083824 2024-06-26
KR10-2024-0083824 2024-06-26
KR1020240101521A KR20260001025A (ko) 2024-06-26 2024-07-31 추론 컨텍스트 확장 방법 및 장치
KR10-2024-0101521 2024-07-31

Publications (1)

Publication Number Publication Date
WO2026005261A1 true WO2026005261A1 (ko) 2026-01-02

Family

ID=98222481

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/005923 Pending WO2026005261A1 (ko) 2024-06-26 2025-04-30 추론 컨텍스트 확장 방법 및 장치

Country Status (1)

Country Link
WO (1) WO2026005261A1 (ko)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20210087384A (ko) * 2020-01-02 2021-07-12 삼성전자주식회사 자연어 이해 모델의 학습을 위한 서버, 클라이언트 디바이스 및 그 동작 방법
KR102354716B1 (ko) * 2014-04-14 2022-01-21 마이크로소프트 테크놀로지 라이센싱, 엘엘씨 딥 러닝 모델을 이용한 상황 의존 검색 기법
KR20230079595A (ko) * 2021-11-29 2023-06-07 주식회사 리얼타임테크 다중 모델 근사 질의 처리 시스템 및 방법
KR102672166B1 (ko) * 2023-10-12 2024-06-07 (주)아스트론시큐리티 생성형 ai에 대한 프롬프트 정보 최적화 방법

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102354716B1 (ko) * 2014-04-14 2022-01-21 마이크로소프트 테크놀로지 라이센싱, 엘엘씨 딥 러닝 모델을 이용한 상황 의존 검색 기법
KR20210087384A (ko) * 2020-01-02 2021-07-12 삼성전자주식회사 자연어 이해 모델의 학습을 위한 서버, 클라이언트 디바이스 및 그 동작 방법
KR20230079595A (ko) * 2021-11-29 2023-06-07 주식회사 리얼타임테크 다중 모델 근사 질의 처리 시스템 및 방법
KR102672166B1 (ko) * 2023-10-12 2024-06-07 (주)아스트론시큐리티 생성형 ai에 대한 프롬프트 정보 최적화 방법

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
JING ZHI, SU YONGYE, HAN YIKUN, YUAN BO, XU HAIYUN, LIU CHUNJIANG, CHEN KEHAI, ZHANG MIN: "When Large Language Models Meet Vector Databases: A Survey", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, ARXIV.ORG, 1 June 2025 (2025-06-01), XP093383123, Retrieved from the Internet <URL:https://arxiv.org/pdf/2402.01763> *

Similar Documents

Publication Publication Date Title
WO2021025397A1 (en) Method and electronic device for quantifying user interest
WO2020180008A1 (en) Method for processing plans having multiple end points and electronic device applying the same method
WO2023277380A1 (ko) 입력 필드를 기반으로 사용자 인터페이스를 구성하는 방법 및 전자 장치
WO2019203494A1 (ko) 문자를 입력하기 위한 전자 장치 및 그의 동작 방법
WO2023101159A1 (ko) 장애인을 위한 시청각 콘텐츠 제공 장치 및 방법
WO2023054913A1 (ko) 포스 터치를 확인하는 전자 장치와 이의 동작 방법
WO2022182038A1 (ko) 음성 명령 처리 장치 및 방법
WO2026005261A1 (ko) 추론 컨텍스트 확장 방법 및 장치
WO2024262868A1 (ko) 전자 장치 및 사용자 발화 처리 방법
WO2023059000A1 (ko) 학습을 보조하기 위한 방법 및 장치
KR20260001025A (ko) 추론 컨텍스트 확장 방법 및 장치
WO2022191395A1 (ko) 사용자 명령을 처리하는 장치 및 그 동작 방법
WO2022075621A1 (ko) 전자 장치 및 전자 장치의 동작 방법
WO2026059355A1 (ko) 사용자 프로필을 생성하기 위한 방법, 전자 장치, 프로그램 및 저장 매체
WO2025135551A1 (ko) 화면에 표시된 하나 이상의 요소들에 기반하여 가상 키보드 또는 추천 단어를 제공하기 위한 전자 장치, 방법, 및 컴퓨터 판독 가능 저장 매체
WO2023106607A1 (ko) 콘텐트를 검색하는 전자 장치 및 그 방법
WO2026043250A1 (ko) 전자 장치 및 전자 장치의 이미지 편집 방법
WO2025121815A1 (ko) 음성 명령을 처리하기 위한 방법, 서버, 및 저장 매체
WO2022191418A1 (ko) 전자 장치 및 미디어 콘텐츠의 재생구간 이동 방법
WO2026071536A1 (ko) 사용자 발화에 대한 결과를 제공하는 전자 장치, 그 동작 방법과, 비일시적 저장 매체
WO2021261882A1 (ko) 전자 장치 및 전자 장치에서 신조어 기반 문장 변환 방법
WO2025150684A1 (ko) 텍스트 컨텐츠를 포맷에 따라 요약하는 전자 장치 및 그 동작 방법
WO2025023722A1 (ko) 전자 장치 및 사용자 발화 처리 방법
WO2024029850A1 (ko) 언어 모델에 기초하여 사용자 발화를 처리하는 방법 및 전자 장치
WO2025042145A1 (ko) 발화 로그를 기반으로 결정된 발화 카테고리를 제공하는 전자 장치 및 그 제어 방법

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25827270

Country of ref document: EP

Kind code of ref document: A1