WO2022196897A1 - 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법 - Google Patents

편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법 Download PDF

Info

Publication number
WO2022196897A1
WO2022196897A1 PCT/KR2021/019326 KR2021019326W WO2022196897A1 WO 2022196897 A1 WO2022196897 A1 WO 2022196897A1 KR 2021019326 W KR2021019326 W KR 2021019326W WO 2022196897 A1 WO2022196897 A1 WO 2022196897A1
Authority
WO
WIPO (PCT)
Prior art keywords
low
prediction error
environment
baseline
bias
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2021/019326
Other languages
English (en)
French (fr)
Inventor
이상완
김동재
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Korea Advanced Institute of Science and Technology KAIST
Original Assignee
Korea Advanced Institute of Science and Technology KAIST
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Korea Advanced Institute of Science and Technology KAIST filed Critical Korea Advanced Institute of Science and Technology KAIST
Publication of WO2022196897A1 publication Critical patent/WO2022196897A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B13/00Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
    • G05B13/02Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
    • G05B13/0265Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric the criterion being a learning criterion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/015Input arrangements based on nervous system activity detection, e.g. brain waves [EEG] detection, electromyograms [EMG] detection, electrodermal response detection
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B13/00Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
    • G05B13/02Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
    • G05B13/0205Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric not using a model or a simulator of the controlled system
    • G05B13/024Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric not using a model or a simulator of the controlled system in which a parameter or coefficient is automatically adjusted to optimise the performance
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B13/00Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
    • G05B13/02Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
    • G05B13/0205Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric not using a model or a simulator of the controlled system
    • G05B13/026Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric not using a model or a simulator of the controlled system using a predictor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models

Definitions

  • Various embodiments relate to an electronic device and a method thereof for brain-simulated adaptive control that solves a bias-dispersion dilemma.
  • bias-variance tradeoff is one of the most fundamental unresolved issues in engineering control and learning system design.
  • a high-complexity intelligent system is advantageous for solving a specific problem situation (low bias error), but shows significant performance degradation even with small changes in the environment (high variance error).
  • a low-complexity intelligent system has a small performance difference according to changes in the environment (low variance error), but shows low performance overall (high bias error).
  • the existing mainstream methodology optimally selects the system of the lesser evil that has the minimum sum of bias error and variance error.
  • an eclectic methodology does not respond promptly to changes in the context (also referred to as context) of the real environment, so there is a risk of performance degradation.
  • context also referred to as context
  • humans show a learning pattern that quickly adapts to context changes.
  • Various embodiments are for deriving a new type of adaptive control system from a computational mechanism of the brain that always maintains a low error value through appropriate fluid control between an intelligent system with low bias error and an intelligent system with low variance error.
  • Various embodiments provide an electronic device and method thereof for brain-simulated adaptive control that solves a bias-dispersion dilemma.
  • a method of an electronic device is based on a first prediction error of a low variance intelligent system and a second prediction error of a low bias intelligent system with respect to an environment. estimating a prediction error baseline for the environment, and based on the estimated prediction error baseline, combining the low-dispersive intelligence system and the low-bias intelligence system to form an adaptive control system It may include implementing steps.
  • an electronic device includes a memory and a processor coupled to the memory and configured to execute at least one instruction stored in the memory, the processor comprising: a low-distributed intelligence system for an environment based on the first prediction error of and the second prediction error of the low-bias intelligent system, estimate a prediction error baseline for the environment, and based on the estimated prediction error baseline, the low-dispersive intelligent system and the low-bias By combining the intelligent system, it may be configured to implement an adaptive control system.
  • estimating a prediction error baseline for the environment based on the first prediction error of the low-dispersion intelligent system for the environment and the second prediction error of the low-bias intelligent system for the environment, estimating a prediction error baseline for the environment, and the estimated one or more programs for executing a method comprising implementing an adaptive control system by combining the low-dispersion intelligent system and the low-bias intelligent system based on a prediction error baseline.
  • the electronic device may implement a brain-simulating adaptive control system that solves the bias-dispersion dilemma.
  • the electronic device flexibly combines a low-dispersion intelligence system and a low-bias intelligence system based on the prediction error baseline for the environment, and through this, the bias-dispersion dilemma through the characteristics of the natural intelligence system based on the human brain. can solve
  • the electronic device may track the total prediction error and maintain a low prediction error by updating the prediction error baseline in response to a change in the environment. Therefore, the adaptive control system can maintain a low prediction error while simultaneously having a low variance error and a low bias error.
  • FIG. 1 is a diagram illustrating an electronic device according to various embodiments of the present disclosure
  • 2, 3, 4, and 5 are diagrams for describing features of an electronic device according to various embodiments.
  • FIG. 6 is a diagram illustrating a method of an electronic device according to various embodiments of the present disclosure.
  • 7, 8, 9A, and 9B are diagrams for explaining the performance of an electronic device according to various embodiments.
  • 1 is a diagram illustrating an electronic device 100 according to various embodiments.
  • 2, 3, 4, and 5 are diagrams for explaining features of the electronic device 100 according to various embodiments.
  • an electronic device 100 may include at least one of an input module 110 , an output module 120 , a memory 130 , and a processor 140 .
  • at least one of the components of the electronic device 100 may be omitted, and at least one other component may be added.
  • at least two of the components of the electronic device 100 may be implemented as one integrated circuit.
  • the input module 110 may input a signal to be used in at least one component of the electronic device 100 .
  • the input module 110 is configured to receive a signal from an input device configured to allow a user to directly input a signal to the electronic device 100 , a sensor device configured to generate a signal by sensing a change in the environment, or an external device It may include at least one of the receiving devices.
  • the sensor device may comprise an inertial measurement unit (IMU).
  • the inertial measurement unit includes a gyroscope, an accelerometer, and a geomagnetic sensor, and the accelerometer may detect roll, yaw, pitch, and the like.
  • the input device may include at least one of a microphone, a mouse, and a keyboard.
  • the input device may include at least one of a touch circuitry configured to sense a touch or a sensor circuit configured to measure an intensity of a force generated by the touch.
  • the output module 120 may output information to the outside of the electronic device 100 .
  • the output module 120 may include at least one of a display device configured to visually output information, an audio output device capable of outputting information as an audio signal, or a transmission device capable of wirelessly transmitting information.
  • the display device may include at least one of a display, a hologram device, and a projector.
  • the display device may be implemented as a touch screen by being assembled with at least one of a touch circuit and a sensor circuit of the input module 110 .
  • the audio output device may include at least one of a speaker and a receiver.
  • the receiving device and the transmitting device may be implemented as a communication module.
  • the communication module may communicate with an external device in the electronic device 100 .
  • the communication module may establish a communication channel between the electronic device 100 and an external device, and communicate with the external device through the communication channel.
  • the external device may include at least one of a vehicle, a satellite, a base station, a server, or another electronic device.
  • the communication module may include at least one of a wired communication module and a wireless communication module.
  • the wired communication module may be connected to an external device by wire and communicate via wire.
  • the wireless communication module may include at least one of a short-range communication module and a long-distance communication module.
  • the short-distance communication module may communicate with an external device in a short-distance communication method.
  • the short-range communication method may include at least one of Bluetooth, WiFi direct, and infrared data association (IrDA).
  • the remote communication module may communicate with an external device in a remote communication method.
  • the remote communication module may communicate with an external device through a network.
  • the network may include at least one of a cellular network, the Internet, or a computer network such as a local area network (LAN) or a wide area network (WAN).
  • the memory 130 may store various data used by at least one component of the electronic device 100 .
  • the memory 130 may include at least one of a volatile memory and a non-volatile memory.
  • the data may include at least one program and input data or output data related thereto.
  • the program may be stored in the memory 130 as software including at least one instruction, and may include at least one of an operating system, middleware, or an application.
  • the processor 140 may execute a program in the memory 130 to control at least one component of the electronic device 100 . Through this, the processor 140 may process data or perform an operation. In this case, the processor 140 may execute a command stored in the memory 130 .
  • the processor 140 may implement an optimal intelligent system by combining a low variance intelligent system and a low bias intelligent system.
  • the low-dispersion intelligent system may indicate an intelligent system having a low variance error
  • the low-biased intelligent system may indicate an intelligent system having a low bias error.
  • the processor 140 is an adaptive control system as an optimal intelligent system that maintains a low prediction error (PE) by flexibly combining a low-distributed intelligent system and a low-biased intelligent system according to a change in the environment. can be implemented.
  • PE prediction error
  • an adaptive control system can be implemented as an intelligent system in which the information processing process of the human brain that solves the bias-dispersion dilemma is transplanted into a model.
  • the low-distributed intelligence system includes a model-free (MF) reinforcement learning (hereinafter, may also be referred to as MF) algorithm
  • the low-biased intelligent system includes a model It may include a model-based (MB) reinforcement learning (hereinafter, may also be referred to as MB) algorithm.
  • the MF algorithm, the MB algorithm, and an algorithm in which the MF algorithm and the MB algorithm are combined (MF+MB) according to various embodiments may have characteristics as illustrated in FIG. 2 . That is, the MF algorithm may have high bias error and low variance error, and the MB algorithm may have high bias error or low bias error, and high variance error.
  • the combined (MF+MB) algorithm according to various embodiments may have low bias error and low variance error.
  • the processor 140 is configured to generate a prediction error baseline (PE baseline) for the environment based on the first prediction error of the low-dispersion intelligent system for the environment and the second prediction error of the low-bias intelligent system for the environment.
  • PE baseline may be variable in a dynamic environment, that is, in response to a change in the environment, and accordingly, the processor 140 may update the prediction error baseline according to a change in the environment.
  • the prediction error baseline may represent a minimum value within the prediction error range achieved by the combination of the low-dispersive intelligent system and the low-biased intelligent system.
  • the first prediction error may be a reward prediction error (RPE)
  • the second prediction error may be a state prediction error (SPE).
  • the processor 140 estimates the prediction error for the environment through value learning based on the first prediction error and the second prediction error, while estimating the strategy control (strategy control). ), it is possible to estimate the prediction error baseline to minimize the prediction error.
  • the processor 140 may implement an adaptive control system by combining a low-dispersive intelligent system and a low-biased intelligent system based on a prediction error baseline.
  • the processor 140 may control the combination ratio of the low-dispersive intelligent system and the low-biased intelligent system based on the prediction error baseline.
  • the processor 140 determines a coupling ratio between the low-dispersive intelligent system and the low-bias intelligent system for achieving the prediction error baseline through adaptive control, and the coupling ratio Therefore, it is possible to combine a low-distributed intelligent system and a low-biased intelligent system.
  • an adaptive control system can be implemented as an optimal intelligent system adaptive to the environment. Therefore, the adaptive control system can maintain a low prediction error while simultaneously having a low variance error and a low bias error.
  • the processor 140 may implement an adaptive control system by combining a model-free (MF) reinforcement learning algorithm and a model-based (MB) reinforcement learning algorithm.
  • MF model-free
  • MB model-based
  • the processor 140 as shown in FIG. 4, the model-free (MF) for the environment compensation prediction error (RPE) of the reinforcement learning algorithm and the model-based (MB) state prediction error of the reinforcement learning algorithm ( SPE) based on value learning, while estimating the prediction error for the environment, it is possible to estimate the prediction error baseline to minimize the prediction error through strategy control.
  • RPE environment compensation prediction error
  • SPE model-based
  • the processor 140 determines the combination ratio of the model-free (MF) reinforcement learning algorithm and the model-based (MB) reinforcement learning algorithm to achieve the prediction error baseline, and the combination ratio It is possible to combine a model-free (MF) reinforcement learning algorithm and a model-based (MB) reinforcement learning algorithm.
  • MF model-free
  • MB model-based
  • the model-free (MF) reinforcement learning algorithm and the model-based (MB) may also be variable in response to changes in the environment. That is, as shown in (b) of FIG. 5 , the processor 140 uses a model-free (MF) reinforcement learning algorithm and model-based (MB) reinforcement to achieve a variable prediction error baseline through adaptive control.
  • the combination ratio of the learning algorithm can be adaptively determined. Through this, an adaptive control system can be implemented as an optimal intelligent system adaptive to the environment.
  • FIG. 6 is a diagram illustrating a method of the electronic device 100 according to various embodiments.
  • the electronic device 100 sets a prediction error baseline for the environment based on the first prediction error of the low-dispersion intelligent system and the second prediction error of the low-biased intelligent system for the environment.
  • the prediction error baseline may represent a minimum value within the prediction error range achieved by the combination of the low-dispersive intelligent system and the low-biased intelligent system.
  • the low-dispersion intelligence system may include a model-free (MF) reinforcement learning algorithm
  • the low-bias intelligence system may include a model-based (MB) reinforcement learning algorithm.
  • the first prediction error may be a compensated prediction error (RPE)
  • the second prediction error may be a state prediction error (SPE).
  • the processor 140 estimates the prediction error for the environment through value learning based on the first prediction error and the second prediction error, and predicts the prediction through strategy control.
  • a prediction error baseline can be estimated to minimize the error.
  • the electronic device 100 may combine the low-dispersive intelligence system and the low-bias intelligence system based on the prediction error baseline.
  • the processor 140 may control the combination ratio of the low-dispersive intelligent system and the low-biased intelligent system based on the prediction error baseline.
  • an adaptive control system can be implemented as an optimal intelligent system adaptive to the environment.
  • the processor 140 may implement an adaptive control system by combining a model-free (MF) reinforcement learning algorithm and a model-based (MB) reinforcement learning algorithm.
  • MF model-free
  • MB model-based
  • the processor 140 determines the combination ratio of the low-dispersive intelligent system and the low-biased intelligent system to achieve the prediction error baseline through adaptive control, as shown in FIG. 3 or FIG. 5B , and , it is possible to combine a low-dispersive intelligent system and a low-biased intelligent system according to the combination ratio.
  • the electronic device 100 may repeatedly perform steps 610 and 620 .
  • the prediction error baseline may be variable in a dynamic environment, that is, in response to a change in the environment, and accordingly, the processor 140 may update the prediction error baseline according to a change in the environment in operation 610 .
  • the processor 140 may adaptively determine a combination ratio of the low-dispersion intelligent system and the low-bias intelligent system to achieve a variable prediction error baseline.
  • an adaptive control system can be implemented as an optimal intelligent system adaptive to the environment. Therefore, the adaptive control system can maintain a low prediction error while simultaneously having a low variance error and a low bias error.
  • FIG 7 and 8 are diagrams for explaining the performance of the electronic device 100 according to various embodiments.
  • the adaptive control system based on the variable prediction error baseline has superior performance compared to other adaptive control systems.
  • the MB algorithm, the MF algorithm, the fixed MB+MF algorithm, and the variable MB+MF algorithm were compared.
  • the excess probability of the adaptively combined (variable MB+MF) algorithm was greater than 0.99, indicating that the performance of the adaptively combined (variable MB+MF) algorithm is superior to that of the rest of the algorithms.
  • the fit indices of the adaptively combined (variable MB+MF) algorithm such as the Bayesian information criterion (BIC) and the Akaike information criterion (AIC), were smaller than the fit indices of the remaining algorithms, which are variable MB+MF) indicates that the performance of the algorithm is superior to that of the rest of the algorithms.
  • BIC Bayesian information criterion
  • AIC Akaike information criterion
  • parameter recovery analysis was performed to examine overfitting to a problem situation in the learning process.
  • the influence of the context of the environment on the selection operation of the MB algorithm, the MF algorithm, the fixed MB+MF algorithm, and the variably, that is, the adaptively coupled (variable MB+MF) algorithm was compared.
  • the adaptively combined (variable MB+MF) algorithm performed an appropriate selection operation for the context of the environment, compared to the rest of the algorithms.
  • 9A and 9B are diagrams for explaining the performance of the electronic device 100 according to various embodiments.
  • the prediction error baseline estimated reflects the information processing process of the human brain for solving the bias-variance dilemma. That is, when compared with the neural activity patterns of the prefrontal cortex in the human brain, the reliability of the MB algorithm and the MF algorithm and the prediction error baseline estimated based on them is very high. This indicates that the electronic device 100 is capable of brain-simulating adaptive control that solves the bias-dispersion dilemma through appropriate fluid control between an intelligent system having a low bias error and an intelligent system having a low variance error.
  • the electronic device 100 may implement a brain-simulating adaptive control system that solves the bias-dispersion dilemma. That is, the electronic device 100 flexibly combines the low-dispersion intelligence system and the low-bias intelligence system based on the prediction error baseline for the environment, and through this, biases through the characteristics of the natural intelligence system based on the human brain. - Solve the dispersion dilemma. In this case, the electronic device 100 may track the total prediction error and maintain a low prediction error by updating the prediction error baseline in response to a change in the environment. Therefore, the adaptive control system can maintain a low prediction error while simultaneously having a low variance error and a low bias error.
  • various embodiments may be applied or applied to various fields.
  • the fields include the field of control systems through sensors, the field of human-robot/computer interaction, the field of smart Internet-of-things (IoT), the field of expert profiling and smart education, and the field of user-targeted AD. may, but is not limited thereto.
  • the first field is the field of control systems through sensors. Control systems, which were previously controlled by automobiles and mechanical equipment, are being replaced by electronic equipment in recent years. This electronization process simply simulates a mechanical process, and thus does not enjoy ancillary benefits through the characteristics of electronization. This is because, when a function failure or error occurs, the damage is large. Control through an intelligent system that can minimize both bias and variance errors can significantly lower the occurrence of errors, so it is important to develop a low-cost-high-performance control system. can be applied
  • the second field is the field of human-robot/computer interaction. Because the basis of all behaviors of natural intelligence is based on higher-order cognitive functions that seek to minimize bias-variance errors, this system can predict human behavior more precisely.
  • the purpose of reading emotions which is one of the types of human cognitive states, is to assist human actions according to the situation. According to various embodiments, it is possible to efficiently respond in assisting human behavior through prediction of environmental changes (eg, arousal and non-awakening) that are contextually similar to emotions that a computer can recognize, beyond simply reading emotions. You can build systems to help humans achieve great results.
  • the third field is the smart IoT field.
  • the cognitive functions used to control each device may vary.
  • the versatility of various embodiments can assist humans by recognizing changes in the environment in controlling each device and predicting functions that humans want to use more efficiently without overfitting.
  • the fourth area is expert profiling and smart education. Resolving the bias-dispersion dilemma by recognizing changes in the environment means that optimal learning is achieved in the human learning process.
  • Various embodiments may clarify in which part of the human learning process (1) immaturity in recognizing environmental changes, and (2) bias and variance error. Therefore, by using various embodiments, it is possible to build an education system that cultivates the ability to perform tasks for judges, doctors, financial experts, military operations commanders, etc., where effective and efficient decision-making is important. In addition, it is possible to pre-profiling which part of the inexperienced part is the key for a customized system for such smart education.
  • the fifth field is a user-targeted AD field.
  • Current ad suggestion technology recommends new ads based on human past search history.
  • this advertisement suggestion technology does not fully consider the changes in the environment that humans experience every moment. This results in a poorly performing system that offers advertisements that are completely out of line with the user's interests.
  • it is possible to implement a more precise user-targeted advertisement by constructing a system that proposes an advertisement having a lower rate (error) not to make a choice than in an environment experienced by humans.
  • the method of the electronic device 100 estimates a prediction error baseline for the environment based on a first prediction error of the low-dispersion intelligent system and a second prediction error of the low-biased intelligent system with respect to the environment. and implementing (step 620) an adaptive control system by combining the low-dispersive intelligent system and the low-biased intelligent system based on the estimated prediction error baseline (step 610).
  • the step of implementing the adaptive control system may include controlling a combination ratio of the low-dispersive intelligent system and the low-biased intelligent system based on the prediction error baseline.
  • the prediction error baseline may be variable, in response to a change in the environment.
  • the prediction error baseline may represent a minimum value within a prediction error range achieved by the combination of the low-dispersive intelligent system and the low-biased intelligent system.
  • the low-dispersion intelligence system may include a model-free (MF) reinforcement learning algorithm
  • the low-bias intelligence system may include a model-based (MB) reinforcement learning algorithm
  • the first prediction error may be a compensated prediction error (RPE), and the second prediction error may be a state prediction error (SPE).
  • RPE compensated prediction error
  • SPE state prediction error
  • the method may be repeatedly performed according to a change in the environment.
  • the electronic device 100 may include a memory 130 and a processor 140 connected to the memory 130 and configured to execute at least one instruction stored in the memory 130 . have.
  • the processor 140 estimates a prediction error baseline for the environment based on the first prediction error of the low-dispersion intelligent system and the second prediction error of the low-bias intelligent system with respect to the environment, Based on the estimated prediction error baseline, the low-dispersion intelligent system and the low-bias intelligent system may be combined to implement an adaptive control system.
  • the processor 140 may be configured to control a combination ratio of the low-dispersive intelligent system and the low-biased intelligent system based on the prediction error baseline.
  • the prediction error baseline may be variable, in response to a change in the environment.
  • the prediction error baseline may represent a minimum value within a prediction error range achieved by the combination of the low-dispersive intelligent system and the low-biased intelligent system.
  • the low-dispersion intelligence system may include a model-free (MF) reinforcement learning algorithm
  • the low-bias intelligence system may include a model-based (MB) reinforcement learning algorithm
  • the first prediction error may be a compensated prediction error (RPE), and the second prediction error may be a state prediction error (SPE).
  • RPE compensated prediction error
  • SPE state prediction error
  • the device described above may be implemented as a hardware component, a software component, and/or a combination of the hardware component and the software component.
  • the devices and components described in the embodiments may include a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), and a programmable logic unit (PLU).
  • ALU arithmetic logic unit
  • FPGA field programmable gate array
  • PLU programmable logic unit
  • It may be implemented using one or more general purpose or special purpose computers, such as a logic unit, microprocessor, or any other device capable of executing and responding to instructions.
  • the processing device may execute an operating system (OS) and one or more software applications running on the operating system.
  • the processing device may also access, store, manipulate, process, and generate data in response to execution of the software.
  • OS operating system
  • the processing device may also access, store, manipulate, process, and generate data in response to execution of the software.
  • the processing device includes a plurality of processing elements and/or a plurality of types of processing elements. It can be seen that may include For example, the processing device may include a plurality of processors or one processor and one controller. Other processing configurations are also possible, such as parallel processors.
  • the software may comprise a computer program, code, instructions, or a combination of one or more thereof, which configures a processing device to operate as desired or is independently or collectively processed You can command the device.
  • the software and/or data may be embodied in any type of machine, component, physical device, computer storage medium or device to be interpreted by or provide instructions or data to the processing device. have.
  • the software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored in one or more computer-readable recording media.
  • the method according to various embodiments may be implemented in the form of program instructions that may be executed by various computer means and recorded in a computer-readable medium.
  • the medium may continue to store the program executable by the computer, or may be temporarily stored for execution or download.
  • the medium may be various recording means or storage means in the form of a single or several hardware combined, it is not limited to a medium directly connected to any computer system, and may exist distributed on a network. Examples of the medium include a hard disk, a magnetic medium such as a floppy disk and a magnetic tape, an optical recording medium such as CD-ROM and DVD, a magneto-optical medium such as a floppy disk, and those configured to store program instructions, including ROM, RAM, flash memory, and the like.
  • examples of other media may include recording media or storage media managed by an app store that distributes applications, sites that supply or distribute other various software, and servers.
  • an (eg, first) component is referred to as being “(functionally or communicatively) connected” or “connected” to another (eg, second) component, that component is It may be directly connected to the component, or may be connected through another component (eg, a third component).
  • module includes a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as, for example, logic, logic block, component, or circuit.
  • a module may be an integrally formed part or a minimum unit or a part of one or more functions.
  • the module may be configured as an application-specific integrated circuit (ASIC).
  • ASIC application-specific integrated circuit
  • each component eg, a module or a program of the described components may include a singular or a plurality of entities.
  • one or more components or steps among the above-described corresponding components may be omitted, or one or more other components or steps may be added.
  • a plurality of components eg, a module or a program
  • the integrated component may perform one or more functions of each component of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to integration.
  • steps performed by a module, program, or other component are executed sequentially, in parallel, iteratively, or heuristically, or one or more of the steps are executed in a different order, omitted, or , or one or more other steps may be added.

Landscapes

  • Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Automation & Control Theory (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biomedical Technology (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Molecular Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Dermatology (AREA)
  • Neurology (AREA)
  • Neurosurgery (AREA)
  • Human Computer Interaction (AREA)
  • Feedback Control In General (AREA)
  • Magnetic Resonance Imaging Apparatus (AREA)
  • Apparatus For Radiation Diagnosis (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

다양한 실시예들은 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법에 관한 것으로, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류(prediction error) 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 환경에 대한 예측 오류 기준선(prediction error baseline)을 추정하고, 추정되는 예측 오류 기준선을 기반으로, 저 편향성 지능 시스템과 저 분산성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하도록 구성된다.

Description

편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법
다양한 실시예들은 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법에 관한 것이다.
편향-분산 상충의 딜레마(bias-variance tradeoff)는 공학적 제어 및 학습 시스템 디자인에서 아직 해결되지 않은 가장 근본적인 이슈 중 하나이다. 높은 복잡도의 지능 시스템은 특정 문제 상황의 해결에 유리하지만(낮은 편향 오류(bias error)), 환경의 조그마한 변화에도 현격한 성능 저하를 보인다(높은 분산 오류(variance error)). 한편, 낮은 복잡도의 지능 시스템은 환경의 변화에 따른 성능차가 적으나(낮은 분산 오류), 전반적으로 낮은 성능을 보인다(높은 편향 오류).
최적의 지능 시스템의 개발을 위해, 기존의 주류 방법론은 절충적으로 편향 오류와 분산 오류의 합이 최소가 되는 차악(次惡)의 시스템을 최적으로 선정하였다. 그러나, 이러한 절충적 방법론은 실제 환경의 컨텍스트(context)(문맥으로도 지칭됨)의 변화에 기민하게 대처하지 못해 성능이 저하될 우려가 있다. 반면, 인간은 문맥 변화에 빠르게 적응하는 학습패턴을 보인다.
다양한 실시예들은, 낮은 편향 오류를 갖는 지능 시스템과 낮은 분산 오류를 갖는 지능 시스템 사이에서 적절한 유동적 제어를 통해, 항상 낮은 오류 값을 유지하는 뇌의 계산적 메커니즘으로부터 새로운 형태의 적응 제어 시스템을 도출하기 위한 것이다.
다양한 실시예들은, 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법을 제공한다.
다양한 실시예들에 따르면, 전자 장치의 방법은, 환경에 대한 저 분산성(low variance) 지능 시스템의 제 1 예측 오류(prediction error) 및 저 편향성(low bias) 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선(prediction error baseline)을 추정하는 단계, 및 상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하는 단계를 포함할 수 있다.
다양한 실시예들에 따르면, 전자 장치는, 메모리, 및 상기 메모리와 연결되고, 상기 메모리에 저장된 적어도 하나의 명령을 실행하도록 구성되는 프로세서를 포함하고, 상기 프로세서는, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선을 추정하고, 상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하도록 구성될 수 있다.
다양한 실시예들에 따르면, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선을 추정하는 단계, 및 상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하는 단계를 포함하는 방법을 실행하기 위한 하나 이상의 프로그램들을 저장할 수 있다.
다양한 실시예들에 따르면, 전자 장치는 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어 시스템을 구현할 수 있다. 즉, 전자 장치는 환경에 대한 예측 오류 기준선을 기반으로 저 분산성 지능 시스템과 저 편향성 지능 시스템을 유동적으로 결합하고, 이를 통해 인간의 뇌를 기반으로 하는 자연 지능 시스템의 특성을 통해 편향-분산 딜레마를 해결할 수 있다. 이 때, 전자 장치는 환경의 변화에 대응하여 예측 오류 기준선을 업데이트함으로써, 총 예측 오류를 추적하고, 낮은 예측 오류를 유지할 수 있다. 따라서, 적응 제어 시스템은 낮은 분산 오류 및 낮은 편향 오류를 동시에 가지면서, 낮은 예측 오류를 유지할 수 있다.
도 1은 다양한 실시예들에 따른 전자 장치를 도시하는 도면이다.
도 2, 도 3, 도 4, 및 도 5는 다양한 실시예들에 따른 전자 장치의 특징을 설명하기 위한 도면들이다.
도 6은 다양한 실시예들에 따른 전자 장치의 방법을 도시하는 도면이다.
도 7, 도 8, 도 9a, 및 도 9b는 다양한 실시예들에 따른 전자 장치의 성능을 설명하기 위한 도면들이다.
이하, 본 문서의 다양한 실시예들이 첨부된 도면을 참조하여 설명된다.
문제 상황의 해결을 위한 지능 시스템의 개발은 필연적으로 편향-분산 딜레마(bias-variance tradeoff)를 발생시킨다. 높은 복잡도를 갖는 지능 시스템은 학습 과정에서 경험한 문제 상황에 과적합(overfitting)되어 환경의 조그마한 변화에도 제대로 기능하지 못하고 현격한 성능의 저하를 겪는다(낮은 편향 오류 및 높은 분산 오류). 이에 반해, 너무 낮은 복잡도를 갖는 지능 시스템은 학습이 충분히 이루어지지 못해 낮은 성능을 갖으나(underfitting) 조그마한 변화에 대응할 수 있는 유연성을 갖는다(높은 편향 오류 및 낮은 분산 오류). 이런 편향-분산 딜레마 속에서 최적의 지능 시스템을 선정하기 위해 일반적으로 쓰이는 방법은 총 오류(편향 오류와 분산 오류의 합)가 낮은, 즉 적절한 복잡도를 갖는 지능 시스템을 이용하는 것이다. 그러나, 이는 절충적 대안에 불과하고 여전히 높은 오류를 갖는다는 문제점을 가져 편향-분산 딜레마의 해결 방법이라고 볼 수 없으며, 특히 실제 상황과 같이 환경이 크게, 그리고 다양한 형태로 변화한다면 더욱 높은 오류를 가질 수밖에 없다.
이러한 문제를 해결하기 위해서는, (1) 다양한 환경의 변화에도 불구하고 총 오류를 최소화할 수 있으며, (2) 낮은 분산 오류를 갖는 시스템과 낮은 편향 오류를 갖는 시스템 사이의 적절하고도 효과적인 적응 제어를 통해서 항상 낮은 오류를 유지할 수 있는 지능 시스템이 개발되어야 한다.
다양한 환경의 변화에서도 총 오류를 최소화하기 위해서는, 환경의 변화에 따라 달라지는 오류의 분포에 제대로 대응하는 지능 시스템일 필요가 있다. 작은 수준의 환경 변화는 해당 분포에 큰 변화가 없기 때문에 적절히 낮은 복잡도를 갖는 지능 시스템을 이용함으로써 대응할 수 있지만, 환경의 변화가 기존과는 완전히 다른 오류 분포를 만들어 내는 경우에는 제대로 대응할 수 없다. 이에 대응하기 위해서는 지능 시스템 자체가 환경의 변화에 따라 변화하는 오류를 추적해야 할 필요가 있다. 예측 오류 기준선을 업데이트하여 변화하는 오류를 추적한다면 환경의 오류 분포와 지능 시스템이 예측하는 오류 분포 간 편차가 줄어듦으로써 환경이 크게 변화해도 총 오류를 최소화할 수 있다.
즉, 단일 시스템이 아닌 여러 개의 지능 시스템을 활용, 상황에 따라 낮은 오류를 갖도록 유동적으로 지능 시스템을 제어하는 것으로, 인간을 포함한 자연 지능 시스템은 이러한 유동적 제어를 통해 환경의 변화에 대응할 수 있다. 특히, 모델-기반과 모델-프리로 구분되는 생물의 두 강화학습 알고리즘들은 각각 낮은 편향 오류와 높은 복잡도, 그리고 낮은 분산 오류와 낮은 복잡도를 갖는 것으로 알려져 있다. 따라서, 이 두 강화학습 알고리즘들 간의 적절한 제어는 총 오류를 낮춰 편형-분산 오류를 해결하는 방법이 될 수 있다.
도 1은 다양한 실시예들에 따른 전자 장치(100)를 도시하는 도면이다. 도 2, 도 3, 도 4, 및 도 5는 다양한 실시예들에 따른 전자 장치(100)의 특징을 설명하기 위한 도면들이다.
도 1을 참조하면, 다양한 실시예들에 따른 전자 장치(100)는, 입력 모듈(110), 출력 모듈(120), 메모리(130), 또는 프로세서(140) 중 적어도 하나를 포함할 수 있다. 어떤 실시예에서, 전자 장치(100)의 구성 요소들 중 적어도 하나가 생략될 수 있으며, 적어도 하나의 다른 구성 요소가 추가될 수 있다. 어떤 실시예에서, 전자 장치(100)의 구성 요소들 중 적어도 두 개가 하나의 통합된 회로로 구현될 수 있다.
입력 모듈(110)은 전자 장치(100)의 적어도 하나의 구성 요소에 사용될 신호를 입력할 수 있다. 입력 모듈(110)은, 사용자가 전자 장치(100)에 직접적으로 신호를 입력하도록 구성되는 입력 장치, 주변의 변화를 감지하여 신호를 발생하도록 구성되는 센서 장치, 또는 외부 기기로부터 신호를 수신하도록 구성되는 수신 장치 중 적어도 하나를 포함할 수 있다. 예를 들면, 센서 장치는 관성 측정 유닛(inertial measurement unit; IMU)을 포함할 수 있다. 관성 측정 유닛은 자이로스코프(gyroscope), 가속도계 및 지자계 센서를 포함하며, 가속도계가 롤(roll), 요(yaw), 피치(pitch) 등을 감지할 수 있다. 예를 들면, 입력 장치는 마이크로폰(microphone), 마우스(mouse) 또는 키보드(keyboard) 중 적어도 하나를 포함할 수 있다. 어떤 실시예에서, 입력 장치는 터치를 감지하도록 설정된 터치 회로(touch circuitry) 또는 터치에 의해 발생되는 힘의 세기를 측정하도록 설정된 센서 회로 중 적어도 하나를 포함할 수 있다.
출력 모듈(120)은 전자 장치(100)의 외부로 정보를 출력할 수 있다. 출력 모듈(120)은, 정보를 시각적으로 출력하도록 구성되는 표시 장치, 정보를 오디오 신호로 출력할 수 있는 오디오 출력 장치, 또는 정보를 무선으로 송신할 수 있는 송신 장치 중 적어도 하나를 포함할 수 있다. 예를 들면, 표시 장치는 디스플레이, 홀로그램 장치 또는 프로젝터 중 적어도 하나를 포함할 수 있다. 일 예로, 표시 장치는 입력 모듈(110)의 터치 회로 또는 센서 회로 중 적어도 하나와 조립되어, 터치 스크린으로 구현될 수 있다. 예를 들면, 오디오 출력 장치는 스피커 또는 리시버 중 적어도 하나를 포함할 수 있다.
일 실시예에 따르면, 수신 장치와 송신 장치는 통신 모듈로 구현될 수 있다. 통신 모듈은 전자 장치(100)에서 외부 기기와 통신을 수행할 수 있다. 통신 모듈은 전자 장치(100)와 외부 기기 간 통신 채널을 수립하고, 통신 채널을 통해, 외부 기기와 통신을 수행할 수 있다. 여기서, 외부 기기는 차량, 위성, 기지국, 서버 또는 다른 전자 장치 중 적어도 하나를 포함할 수 있다. 통신 모듈은 유선 통신 모듈 또는 무선 통신 모듈 중 적어도 하나를 포함할 수 있다. 유선 통신 모듈은 외부 기기와 유선으로 연결되어, 유선으로 통신할 수 있다. 무선 통신 모듈은 근거리 통신 모듈 또는 원거리 통신 모듈 중 적어도 하나를 포함할 수 있다. 근거리 통신 모듈은 외부 기기와 근거리 통신 방식으로 통신할 수 있다. 예를 들면, 근거리 통신 방식은, 블루투스(Bluetooth), 와이파이 다이렉트(WiFi direct), 또는 적외선 통신(IrDA; infrared data association) 중 적어도 하나를 포함할 수 있다. 원거리 통신 모듈은 외부 기기와 원거리 통신 방식으로 통신할 수 있다. 여기서, 원거리 통신 모듈은 네트워크를 통해 외부 기기와 통신할 수 있다. 예를 들면, 네트워크는 셀룰러 네트워크, 인터넷, 또는 LAN(local area network)이나 WAN(wide area network)과 같은 컴퓨터 네트워크 중 적어도 하나를 포함할 수 있다.
메모리(130)는 전자 장치(100)의 적어도 하나의 구성 요소에 의해 사용되는 다양한 데이터를 저장할 수 있다. 예를 들면, 메모리(130)는 휘발성 메모리 또는 비휘발성 메모리 중 적어도 하나를 포함할 수 있다. 데이터는 적어도 하나의 프로그램 및 이와 관련된 입력 데이터 또는 출력 데이터를 포함할 수 있다. 프로그램은 메모리(130)에 적어도 하나의 명령을 포함하는 소프트웨어로서 저장될 수 있으며, 운영 체제, 미들 웨어 또는 어플리케이션 중 적어도 하나를 포함할 수 있다.
프로세서(140)는 메모리(130)의 프로그램을 실행하여, 전자 장치(100)의 적어도 하나의 구성 요소를 제어할 수 있다. 이를 통해, 프로세서(140)는 데이터 처리 또는 연산을 수행할 수 있다. 이 때, 프로세서(140)는 메모리(130)에 저장된 명령을 실행할 수 있다.
다양한 실시예들에 따르면, 프로세서(140)는 저 분산성(low variance) 지능 시스템과 저 편향성(low bias) 지능 시스템을 결합하여, 최적의 지능 시스템을 구현할 수 있다. 여기서, 저 분산성 지능 시스템은 낮은 분산 오류(variance error)를 갖는 지능 시스템을 나타내고, 저 편향성 지능 시스템은 낮은 편향 오류(bias error)를 갖는 지능 시스템을 나타낼 수 있다. 이 때, 프로세서(140)는 환경의 변화에 따라, 저 분산성 지능 시스템과 저 편향성 지능 시스템을 유동적으로 결합하여, 낮은 예측 오류(prediction error; PE)을 유지하는 최적의 지능 시스템으로서 적응 제어 시스템을 구현할 수 있다. 이를 통해, 편향-분산 딜레마를 해결하는 인간의 뇌의 정보 처리 과정이 모델로 이식된 지능 시스템으로서, 적응 제어 시스템이 구현될 수 있다.
일 실시예에 따르면, 저 분산성 지능 시스템은 모델-프리(model-free; MF) 강화학습(reinforcement learning)(이하에서, MF로도 지칭될 수 있음) 알고리즘을 포함하고, 저 편향성 지능 시스템은 모델-기반(model-based; MB) 강화학습(이하에서, MB로도 지칭될 수 있음) 알고리즘을 포함할 수 있다. MF 알고리즘, MB 알고리즘, 및 다양한 실시예들에 따라 MF 알고리즘과 MB 알고리즘이 결합된(MF+MB) 알고리즘은, 도 2에 도시된 바와 같은 특성들을 가질 수 있다. 즉, MF 알고리즘은 높은 편향 오류 및 낮은 분산 오류를 가지며, MB 알고리즘은 높은 편향 오류 또는 낮은 편향 오류, 및 높은 분산 오류를 가질 수 있다. 이에 비해, 다양한 실시예들에 따라 결합된(MF+MB) 알고리즘은 낮은 편향 오류 및 낮은 분산 오류를 가질 수 있다.
이를 위해, 프로세서(140)는, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 환경에 대한 예측 오류 기준선(prediction error baseline; PE baseline)을 추정할 수 있다. 여기서, 예측 오류 기준선은, 동적 환경에서, 즉 환경의 변화에 대응하여, 가변적일 수 있으며, 이에 따라, 프로세서(140)는 환경의 변화에 따라 예측 오류 기준선을 업데이트할 수 있다. 이 때, 예측 오류 기준선은, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타낼 수 있다. 일 실시예에 따르면, 제 1 예측 오류는 보상 예측 오류(reward prediction error; RPE)이고, 제 2 예측 오류는 상태 예측 오류(state prediction error; SPE)일 수 있다. 여기서, 프로세서(140)는 도 3에 도시된 바와 같이, 제 1 예측 오류 및 제 2 예측 오류를 기반으로 가치 학습(value learning)을 통해, 환경에 대한 예측 오류를 추정하면서, 전략 제어(strategy control)를 통해, 예측 오류를 최소화하기 위한 예측 오류 기준선을 추정할 수 있다.
그리고, 프로세서(140)는 예측 오류 기준선(prediction error baseline)을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현할 수 있다. 이 때, 프로세서(140)는 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 제어할 수 있다. 여기서, 프로세서(140)는 도 3에 도시된 바와 같이, 적응 제어(adaptive control)를 통해, 예측 오류 기준선을 달성하기 위한 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 결정하고, 결합 비율에 따라 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합할 수 있다. 이를 통해, 환경에 대해 적응적인 최적의 지능 시스템으로서, 적응 제어 시스템이 구현될 수 있다. 따라서, 적응 제어 시스템은 낮은 분산 오류 및 낮은 편향 오류를 동시에 가지면서, 낮은 예측 오류를 유지할 수 있다.
일 실시예에 따르면, 프로세서(140)는 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘을 결합하여, 적응 제어 시스템을 구현할 수 있다. 이 때, 프로세서(140)는 도 4에 도시된 바와 같이, 환경에 대한 모델-프리(MF) 강화학습 알고리즘의 보상 예측 오류(RPE) 및 모델-기반(MB) 강화학습 알고리즘의 상태 예측 오류(SPE)를 기반으로 가치 학습을 통해, 환경에 대한 예측 오류를 추정하면서, 전략 제어를 통해, 예측 오류를 최소화하기 위한 예측 오류 기준선을 추정할 수 있다. 그리고, 프로세서(140)는 도 5에 도시된 바와 같이, 예측 오류 기준선을 달성하기 위한 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘의 결합 비율을 결정하고, 결합 비율에 따라 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘을 결합할 수 있다. 여기서, 도 5의 (a)에 도시된 바와 같이 예측 오류 기준선이 환경의 변화에 대해 고정적으로 활용되는 경우(fixed PE baseline), 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘의 결합 비율이 환경의 변화에도 유지될 수 있다. 이에 비해, 도 5의 (b)에 도시된 바와 같이 예측 오류 기준선이 환경의 변화에 대응하여 가변적으로 활용되는 경우(variable PE baseline), 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘의 결합 비율도 환경의 변화에 대응하여 가변적일 수 있다. 즉, 프로세서(140)는 도 5의 (b)에 도시된 바와 같이, 적응 제어를 통해, 가변적인 예측 오류 기준선을 달성하기 위한 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘의 결합 비율을 적응적으로 결정할 수 있다. 이를 통해, 환경에 대해 적응적인 최적의 지능 시스템으로서, 적응 제어 시스템이 구현될 수 있다.
도 6은 다양한 실시예들에 따른 전자 장치(100)의 방법을 도시하는 도면이다.
도 6을 참조하면, 전자 장치(100)는 610 단계에서, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 환경에 대한 예측 오류 기준선을 추정할 수 있다. 이 때, 예측 오류 기준선은, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타낼 수 있다. 일 실시예에 따르면, 저 분산성 지능 시스템은 모델-프리(MF) 강화학습 알고리즘을 포함하고, 저 편향성 지능 시스템은 모델-기반(MB) 강화학습 알고리즘을 포함할 수 있다. 이러한 경우, 제 1 예측 오류는 보상 예측 오류(RPE)이고, 제 2 예측 오류는 상태 예측 오류(SPE)일 수 있다. 여기서, 프로세서(140)는 도 3 또는 도 4에 도시된 바와 같이, 제 1 예측 오류 및 제 2 예측 오류를 기반으로 가치 학습을 통해, 환경에 대한 예측 오류를 추정하면서, 전략 제어를 통해, 예측 오류를 최소화하기 위한 예측 오류 기준선을 추정할 수 있다.
다음으로, 전자 장치(100)는 620 단계에서 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합할 수 있다. 이 때, 프로세서(140)는 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 제어할 수 있다. 이를 통해, 환경에 대해 적응적인 최적의 지능 시스템으로서, 적응 제어 시스템이 구현될 수 있다. 일 실시예에 따르면, 프로세서(140)는 모델-프리(MF) 강화학습 알고리즘과 모델-기반(MB) 강화학습 알고리즘을 결합하여, 적응 제어 시스템을 구현할 수 있다. 여기서, 프로세서(140)는 도 3 또는 도 5의 (b)에 도시된 바와 같이, 적응 제어를 통해, 예측 오류 기준선을 달성하기 위한 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 결정하고, 결합 비율에 따라 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합할 수 있다.
다양한 실시예들에 따르면, 전자 장치(100)는 610 단계 및 620 단계를 반복하여 수행할 수 있다. 여기서, 예측 오류 기준선은, 동적 환경에서, 즉 환경의 변화에 대응하여, 가변적일 수 있으며, 이에 따라, 프로세서(140)는 610 단계에서 환경의 변화에 따라 예측 오류 기준선을 업데이트할 수 있다. 그리고, 프로세서(140)는 가변적인 예측 오류 기준선을 달성하기 위한 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 적응적으로 결정할 수 있다. 이를 통해, 환경에 대해 적응적인 최적의 지능 시스템으로서, 적응 제어 시스템이 구현될 수 있다. 따라서, 적응 제어 시스템은 낮은 분산 오류 및 낮은 편향 오류를 동시에 가지면서, 낮은 예측 오류를 유지할 수 있다.
도 7 및 도 8은 다양한 실시예들에 따른 전자 장치(100)의 성능을 설명하기 위한 도면들이다.
도 7 및 도 8을 참조하면, 가변적인 예측 오류 기준선을 기반으로 하는 적응 제어 시스템은, 다른 적응 제어 시스템과 비교하여, 우수한 성능을 갖는다. 도 7의 (a)에 도시된 바와 같이, MB 알고리즘, MF 알고리즘, 고정적으로 결합된(fixed MB+MF) 알고리즘, 가변적으로, 즉 적응적으로 결합되는(variable MB+MF) 알고리즘이 비교되었다. 결과적으로, 적응적으로 결합되는(variable MB+MF) 알고리즘의 초과 확률(exceedance probability)만이 0.99 보다 컸으며, 이는 적응적으로 결합되는(variable MB+MF) 알고리즘의 성능이 나머지 알고리즘들의 성능 보다 우수함을 나타낸다. 아울러, 적응적으로 결합되는(variable MB+MF) 알고리즘의 적합 지수, 예컨대 BIC(Bayesian information criterion) 및 AIC(Akaike information criterion)가 나머지 알고리즘들의 적합 지수 보다 작았으며, 이는 적응적으로 결합되는(variable MB+MF) 알고리즘의 성능이 나머지 알고리즘들의 성능 보다 우수함을 나타낸다. 한편, 도 7의 (b)에 도시된 바와 같이, 적응적으로 결합되는(variable MB+MF) 알고리즘의 선택 동작에 대한 환경의 컨텍스트의 영향이 평가되었다. 결과적으로, 적응적으로 결합되는(variable MB+MF) 알고리즘은 환경의 컨텍스트에 대해 적절한 선택 동작을 수행하였다. 한편, 도 8에 도시된 바와 같이, 학습 과정에서의 문제 상황에 대한 과적합을 검사하기 위해, 파라미터 리커버리 분석(parameter recovery analysis)이 수행되었다. 이 때, MB 알고리즘, MF 알고리즘, 고정적으로 결합된(fixed MB+MF) 알고리즘, 가변적으로, 즉 적응적으로 결합되는(variable MB+MF) 알고리즘의 선택 동작에 대한 환경의 컨텍스트의 영향이 비교되었다. 결과적으로, 적응적으로 결합되는(variable MB+MF) 알고리즘은 나머지 알고리즘들에 비해, 환경의 컨텍스트에 대해 적절한 선택 동작을 수행하였다.
도 9a 및 도 9b는 다양한 실시예들에 따른 전자 장치(100)의 성능을 설명하기 위한 도면들이다.
도 9a 및 도 9b를 참조하면, 다양한 실시예들에 따라 추정되는 예측 오류 기준선은 편향-분산 딜레마를 해결하는 인간의 뇌의 정보 처리 과정을 반영한다. 즉, 인간의 뇌 영역에서 측전전두피질의 신경활성 패턴과 비교하면, MB 알고리즘과 MF 알고리즘, 및 이들을 기반으로 추정되는 예측 오류 기준선에 대한 신뢰성은 매우 높다. 이는, 전자 장치(100)가 낮은 편향 오류를 갖는 지능 시스템과 낮은 분산 오류를 갖는 지능 시스템 사이에서 적절한 유동적 제어를 통해, 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어가 가능함을 나타낸다.
다양한 실시예들에 따르면, 전자 장치(100)는 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어 시스템을 구현할 수 있다. 즉, 전자 장치(100)는 환경에 대한 예측 오류 기준선을 기반으로 저 분산성 지능 시스템과 저 편향성 지능 시스템을 유동적으로 결합하고, 이를 통해 인간의 뇌를 기반으로 하는 자연 지능 시스템의 특성을 통해 편향-분산 딜레마를 해결할 수 있다. 이 때, 전자 장치(100)는 환경의 변화에 대응하여 예측 오류 기준선을 업데이트함으로써, 총 예측 오류를 추적하고, 낮은 예측 오류를 유지할 수 있다. 따라서, 적응 제어 시스템은 낮은 분산 오류 및 낮은 편향 오류를 동시에 가지면서, 낮은 예측 오류를 유지할 수 있다.
현재 개발되는 지능 시스템은 실패의 리스크를 줄이는 보수적인 방식으로 구성되어 있기 때문에 낮은 복잡도를 갖기에 효과적인 성능의 향상이 제한되어 왔다. 그러나, 다양한 실시예들에 따르면, 적응 제어 시스템은 편향-분산 딜레마를 해결할 수 있기 때문에 기존의 지능 시스템의 획기적인 성능 향상을 도모할 수 있다. 이에 따라, 다양한 실시예들은 다양한 분야들에 적용 또는 응용될 수 있다. 예를 들어, 그 분야들로는 센서를 통한 제어 시스템 분야, 인간-로봇/컴퓨터 상호작용 분야, 스마트 IoT(Internet-of-things) 분야, 전문가 프로파일링 및 스마트 교육 분야, 유저 타겟형 AD 분야 등이 있을 수 있으며, 이에 제한되지 않는다.
첫 번째 분야는 센서를 통한 제어 시스템 분야이다. 자동차를 필두로, 과거에는 기계장비를 통해서 제어되던 제어 시스템은 최근 전자 장비로 대체되고 있다. 이러한 전자화 과정은 단순히 기계적 과정을 모사하는 방식으로 이루어져 전자화의 특성을 통해 부수적인 이득을 누리고 있지 못하다. 이는 기능 실패, 혹은 오류가 발생하는 경우 그 손해가 크기 때문인데, 편향-분산 오류를 모두 최소화할 수 있는 지능 시스템을 통한 제어는 그 오류의 발생을 현격히 낮출 수 있기에 저비용-고성능의 제어 시스템 개발에 응용될 수 있다.
두 번째 분야는 인간-로봇/컴퓨터 상호작용 분야이다. 자연 지능의 모든 행동의 기반에는 편향-분산 오류를 최소화하고자 하는 고차원적인 인지 기능에 근거하여 일어나므로, 이 시스템을 통해 인간의 행동을 보다 정밀히 예측할 수 있다. 대표적인 예로, 감정 컴퓨팅(affective computing) 분야에서는 인간의 인지 상태의 종류 중 하나인 감정을 읽어 내어 상황에 맞게 인간의 행동을 보조하는 것을 목적으로 한다. 다양한 실시예들에 따르면, 단순히 감정을 읽어내는 것을 넘어서 컴퓨터가 인식할 수 있는 감정과 맥락적으로 유사한 환경 변화(예: 각성과 비각성)의 예측을 통해서 인간 행동의 보조에 있어서 효율적으로 대응하는 시스템을 구축하여 인간이 훌륭한 성과를 거둘 수 있도록 보조할 수 있다.
세 번째 분야는 스마트 IoT 분야이다. 특히, IoT 분야에서는 다양한 기기를 컨트롤 해야 하므로 각 기기의 컨트롤에 활용되는 인지 기능이 다양할 수 있다. 이때 다양한 실시예들의 범용성은 각 기기를 제어함에 있어서 환경의 변화를 인식하여 보다 효율적으로 인간이 이용하고자 하는 기능을 과적합없이 예측함으로써 인간을 보조할 수 있다.
네 번째 분야는 전문가 프로파일링 및 스마트 교육 분야이다. 환경 변화를 인식함으로써 이뤄지는 편향-분산 딜레마의 해결은 인간의 학습 과정에서 최적의 학습이 이뤄졌을 때를 의미한다. 다양한 실시예들은 인간의 학습 과정이 (1) 환경 변화의 인식의 미숙함, (2) 편형과 분산 오류 중 어느 부분에서의 미숙함이 있는지를 명확히 할 수 있다. 때문에 다양한 실시예들을 이용함으로써 효과적이고도 효율적인 의사결정이 중요한 판사, 의사, 금융 전문가, 군사 작전 지휘관 등에 대한 작업 수행능력을 함양하는 교육 시스템을 구축할 수 있다. 또한 이러한 스마트 교육을 위한 맞춤형 시스템을 위해 어느 부분의 미숙함이 핵심이 되는지 사전 프로파일링이 가능하다.
다섯 번째 분야는 유저 타겟형 AD 분야이다. 현행 광고 제안 기술은 인간의 과거 검색 기록을 바탕으로 새로운 광고를 추천하고 있다. 그러나 이러한 광고 제안 기술은 매 순간 인간이 경험하는 환경의 변화에 대해서는 완전히 고려하지 않고 있다. 이는 사용자의 관심사와 완전히 동떨어진 광고를 제안하는, 성능이 낮은 시스템의 원인이 된다. 다양한 실시예들을 활용하면 인간이 경험하는 환경에서 보다 선택을 하지 않을 비율(오류)이 낮은 광고를 제안하는 시스템을 구축함으로써 보다 정밀한 유저 타겟형 광고를 구현할 수 있다.
지능형 시스템은 그 유용함에도 불구하고 기능의 실패가 발생할 시 치명적인 결과로 이어지는 분야에서는 활용이 보수적으로 이루어질 수밖에 없었다. 그러나, 최근의 시장에서는 제어 시스템에 지능형 시스템의 도입을 인공지능을 기반으로 증대시키는 추세이다. 이런 추세에서 편향-분산의 딜레마를 해결하는 지능 시스템은 기존의 보수적 제어 시스템(낮은 복잡도 및 낮은 성능) 인공지능 기반의 최신 제어 시스템(높은 복잡도 및 높은 성능) 사이에서 효율적이고도 효과적인 제어 시스템으로 기능할 것이다(낮은 복잡도 및 높은 성능). 편향-분산 딜레마는 모든 지능 시스템이 필연적으로 경험하게 된다. 다양한 실시예들은 이를 해결할 뿐만 아니라 환경의 높은 변동성에도 성공적으로 기능할 수 있기에 최신의 딥 러닝 기반으로 개발된 지능 시스템에 비해서는 낮은 복잡도를 갖으며, 고전적인 지능 시스템에 비해서는 높은 성능을 가짐으로써 지능 시스템을 주요 사업군으로 하는 직종에 모두 적용될 수 있을 것이다.
다양한 실시예들에 따른 전자 장치(100)의 방법은, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 환경에 대한 예측 오류 기준선을 추정하는 단계(610 단계), 및 추정되는 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하는 단계(620 단계)를 포함할 수 있다.
다양한 실시예들에 따르면, 적응 제어 시스템을 구현하는 단계(620 단계)는, 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 제어하는 단계를 포함할 수 있다.
다양한 실시예들에 따르면, 예측 오류 기준선은, 환경의 변화에 대응하여, 가변적일 수 있다.
다양한 실시예들에 따르면, 예측 오류 기준선은, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타낼 수 있다.
다양한 실시예들에 따르면, 저 분산성 지능 시스템은 모델-프리(MF) 강화학습 알고리즘을 포함하고, 저 편향성 지능 시스템은 모델-기반(MB) 강화학습 알고리즘을 포함할 수 있다.
다양한 실시예들에 따르면, 제 1 예측 오류는 보상 예측 오류(RPE)이고, 제 2 예측 오류는 상태 예측 오류(SPE)일 수 있다.
다양한 실시예들에 따르면, 상기 방법은, 환경의 변화에 따라, 반복적으로 수행될 수 있다.
다양한 실시예들에 따른 전자 장치(100)는, 메모리(130), 및 메모리(130)와 연결되고, 메모리(130)에 저장된 적어도 하나의 명령을 실행하도록 구성되는 프로세서(140)를 포함할 수 있다.
다양한 실시예들에 따르면, 프로세서(140)는, 환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 환경에 대한 예측 오류 기준선을 추정하고, 추정되는 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하도록 구성될 수 있다.
다양한 실시예들에 따르면, 프로세서(140)는, 예측 오류 기준선을 기반으로, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합 비율을 제어하도록 구성될 수 있다.
다양한 실시예들에 따르면, 예측 오류 기준선은, 환경의 변화에 대응하여, 가변적일 수 있다.
다양한 실시예들에 따르면, 예측 오류 기준선은, 저 분산성 지능 시스템과 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타낼 수 있다.
다양한 실시예들에 따르면, 저 분산성 지능 시스템은 모델-프리(MF) 강화학습 알고리즘을 포함하고, 저 편향성 지능 시스템은 모델-기반(MB) 강화학습 알고리즘을 포함할 수 있다.
다양한 실시예들에 따르면, 제 1 예측 오류는 보상 예측 오류(RPE)이고, 제 2 예측 오류는 상태 예측 오류(SPE)일 수 있다.
이상에서 설명된 장치는 하드웨어 구성 요소, 소프트웨어 구성 요소, 및/또는 하드웨어 구성 요소 및 소프트웨어 구성 요소의 조합으로 구현될 수 있다. 예를 들어, 실시예들에서 설명된 장치 및 구성 요소는, 프로세서, 컨트롤러, ALU(arithmetic logic unit), 디지털 신호 프로세서(digital signal processor), 마이크로컴퓨터, FPGA(field programmable gate array), PLU(programmable logic unit), 마이크로프로세서, 또는 명령(instruction)을 실행하고 응답할 수 있는 다른 어떠한 장치와 같이, 하나 이상의 범용 컴퓨터 또는 특수 목적 컴퓨터를 이용하여 구현될 수 있다. 처리 장치는 운영 체제(OS) 및 상기 운영 체제 상에서 수행되는 하나 이상의 소프트웨어 어플리케이션을 수행할 수 있다. 또한, 처리 장치는 소프트웨어의 실행에 응답하여, 데이터를 접근, 저장, 조작, 처리 및 생성할 수도 있다. 이해의 편의를 위하여, 처리 장치는 하나가 사용되는 것으로 설명된 경우도 있지만, 해당 기술분야에서 통상의 지식을 가진 자는, 처리 장치가 복수 개의 처리 요소(processing element) 및/또는 복수 유형의 처리 요소를 포함할 수 있음을 알 수 있다. 예를 들어, 처리 장치는 복수 개의 프로세서 또는 하나의 프로세서 및 하나의 컨트롤러를 포함할 수 있다. 또한, 병렬 프로세서(parallel processor)와 같은, 다른 처리 구성(processing configuration)도 가능하다.
소프트웨어는 컴퓨터 프로그램(computer program), 코드(code), 명령(instruction), 또는 이들 중 하나 이상의 조합을 포함할 수 있으며, 원하는 대로 동작하도록 처리 장치를 구성하거나 독립적으로 또는 결합적으로(collectively) 처리 장치를 명령할 수 있다. 소프트웨어 및/또는 데이터는, 처리 장치에 의하여 해석되거나 처리 장치에 명령 또는 데이터를 제공하기 위하여, 어떤 유형의 기계, 구성 요소(component), 물리적 장치, 컴퓨터 저장 매체 또는 장치에 구체화(embody)될 수 있다. 소프트웨어는 네트워크로 연결된 컴퓨터 시스템 상에 분산되어서, 분산된 방법으로 저장되거나 실행될 수도 있다. 소프트웨어 및 데이터는 하나 이상의 컴퓨터 판독 가능 기록 매체에 저장될 수 있다.
다양한 실시예들에 따른 방법은 다양한 컴퓨터 수단을 통하여 수행될 수 있는 프로그램 명령 형태로 구현되어 컴퓨터-판독 가능 매체에 기록될 수 있다. 이 때, 매체는 컴퓨터로 실행 가능한 프로그램을 계속 저장하거나, 실행 또는 다운로드를 위해 임시 저장하는 것일 수도 있다. 그리고, 매체는 단일 또는 수 개의 하드웨어가 결합된 형태의 다양한 기록수단 또는 저장수단일 수 있는데, 어떤 컴퓨터 시스템에 직접 접속되는 매체에 한정되지 않고, 네트워크 상에 분산 존재하는 것일 수도 있다. 매체의 예시로는, 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체, CD-ROM 및 DVD와 같은 광기록 매체, 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical medium), 및 ROM, RAM, 플래시 메모리 등을 포함하여 프로그램 명령어가 저장되도록 구성된 것이 있을 수 있다. 또한, 다른 매체의 예시로, 어플리케이션을 유통하는 앱 스토어나 기타 다양한 소프트웨어를 공급 내지 유통하는 사이트, 서버 등에서 관리하는 기록매체 내지 저장매체도 들 수 있다.
본 문서의 다양한 실시예들 및 이에 사용된 용어들은 본 문서에 기재된 기술을 특정한 실시 형태에 대해 한정하려는 것이 아니며, 해당 실시 예의 다양한 변경, 균등물, 및/또는 대체물을 포함하는 것으로 이해되어야 한다. 도면의 설명과 관련하여, 유사한 구성 요소에 대해서는 유사한 참조 부호가 사용될 수 있다. 단수의 표현은 문맥상 명백하게 다르게 뜻하지 않는 한, 복수의 표현을 포함할 수 있다. 본 문서에서, "A 또는 B", "A 및/또는 B 중 적어도 하나", "A, B 또는 C" 또는 "A, B 및/또는 C 중 적어도 하나" 등의 표현은 함께 나열된 항목들의 모든 가능한 조합을 포함할 수 있다. "제 1", "제 2", "첫째" 또는 "둘째" 등의 표현들은 해당 구성 요소들을, 순서 또는 중요도에 상관없이 수식할 수 있고, 한 구성 요소를 다른 구성 요소와 구분하기 위해 사용될 뿐 해당 구성 요소들을 한정하지 않는다. 어떤(예: 제 1) 구성 요소가 다른(예: 제 2) 구성 요소에 "(기능적으로 또는 통신적으로) 연결되어" 있다거나 "접속되어" 있다고 언급된 때에는, 상기 어떤 구성 요소가 상기 다른 구성 요소에 직접적으로 연결되거나, 다른 구성 요소(예: 제 3 구성 요소)를 통하여 연결될 수 있다.
본 문서에서 사용된 용어 "모듈"은 하드웨어, 소프트웨어 또는 펌웨어로 구성된 유닛을 포함하며, 예를 들면, 로직, 논리 블록, 부품, 또는 회로 등의 용어와 상호 호환적으로 사용될 수 있다. 모듈은, 일체로 구성된 부품 또는 하나 또는 그 이상의 기능을 수행하는 최소 단위 또는 그 일부가 될 수 있다. 예를 들면, 모듈은 ASIC(application-specific integrated circuit)으로 구성될 수 있다.
다양한 실시예들에 따르면, 기술한 구성 요소들의 각각의 구성 요소(예: 모듈 또는 프로그램)는 단수 또는 복수의 개체를 포함할 수 있다. 다양한 실시예들에 따르면, 전술한 해당 구성 요소들 중 하나 이상의 구성 요소들 또는 단계들이 생략되거나, 또는 하나 이상의 다른 구성 요소들 또는 단계들이 추가될 수 있다. 대체적으로 또는 추가적으로, 복수의 구성 요소들(예: 모듈 또는 프로그램)은 하나의 구성 요소로 통합될 수 있다. 이런 경우, 통합된 구성 요소는 복수의 구성 요소들 각각의 구성 요소의 하나 이상의 기능들을 통합 이전에 복수의 구성 요소들 중 해당 구성 요소에 의해 수행되는 것과 동일 또는 유사하게 수행할 수 있다. 다양한 실시예들에 따르면, 모듈, 프로그램 또는 다른 구성 요소에 의해 수행되는 단계들은 순차적으로, 병렬적으로, 반복적으로, 또는 휴리스틱하게 실행되거나, 단계들 중 하나 이상이 다른 순서로 실행되거나, 생략되거나, 또는 하나 이상의 다른 단계들이 추가될 수 있다.

Claims (20)

  1. 전자 장치의 방법에 있어서,
    환경에 대한 저 분산성(low variance) 지능 시스템의 제 1 예측 오류(prediction error) 및 저 편향성(low bias) 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선(prediction error baseline)을 추정하는 단계; 및
    상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하는 단계
    를 포함하는,
    전자 장치의 방법.
  2. 제 1 항에 있어서,
    상기 적응 제어 시스템을 구현하는 단계는,
    상기 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합 비율을 제어하는 단계
    를 포함하는,
    전자 장치의 방법.
  3. 제 1 항에 있어서,
    상기 예측 오류 기준선은,
    환경의 변화에 대응하여, 가변적인,
    전자 장치의 방법.
  4. 제 1 항에 있어서,
    상기 예측 오류 기준선은,
    상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타내는,
    전자 장치의 방법.
  5. 제 1 항에 있어서,
    상기 저 분산성 지능 시스템은,
    모델-프리(model-free; MF) 강화학습 알고리즘을 포함하고,
    상기 저 편향성 지능 시스템은,
    모델-기반(model-based; MB) 강화학습 알고리즘을 포함하는,
    전자 장치의 방법.
  6. 제 1 항에 있어서,
    상기 제 1 예측 오류는,
    보상 예측 오류(reward prediction error; RPE)이고,
    상기 제 2 예측 오류는,
    상태 예측 오류(state prediction error; SPE)인,
    전자 장치의 방법.
  7. 제 3 항에 있어서,
    상기 방법은,
    상기 환경의 변화에 따라, 반복적으로 수행되는,
    전자 장치의 방법.
  8. 전자 장치에 있어서,
    메모리; 및
    상기 메모리와 연결되고, 상기 메모리에 저장된 적어도 하나의 명령을 실행하도록 구성되는 프로세서
    를 포함하고,
    상기 프로세서는,
    환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선을 추정하고,
    상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하도록 구성되는,
    전자 장치.
  9. 제 8 항에 있어서,
    상기 프로세서는,
    상기 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합 비율을 제어하도록 구성되는,
    전자 장치.
  10. 제 8 항에 있어서,
    상기 예측 오류 기준선은,
    환경의 변화에 대응하여, 가변적인,
    전자 장치.
  11. 제 8 항에 있어서,
    상기 예측 오류 기준선은,
    상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타내는,
    전자 장치.
  12. 제 8 항에 있어서,
    상기 저 분산성 지능 시스템은,
    모델-프리(MF) 강화학습 알고리즘을 포함하고,
    상기 저 편향성 지능 시스템은,
    모델-기반(MB) 강화학습 알고리즘을 포함하는,
    전자 장치.
  13. 제 8 항에 있어서,
    상기 제 1 예측 오류는,
    보상 예측 오류(RPE)이고,
    상기 제 2 예측 오류는,
    상태 예측 오류(SPE)인,
    전자 장치.
  14. 비-일시적인 컴퓨터-판독 가능 저장 매체에 있어서,
    환경에 대한 저 분산성 지능 시스템의 제 1 예측 오류 및 저 편향성 지능 시스템의 제 2 예측 오류를 기반으로, 상기 환경에 대한 예측 오류 기준선을 추정하는 단계; 및
    상기 추정되는 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템을 결합하여, 적응 제어 시스템을 구현하는 단계
    를 포함하는 방법을 실행하기 위한 하나 이상의 프로그램들을 저장하기 위한,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  15. 제 14 항에 있어서,
    상기 적응 제어 시스템을 구현하는 단계는,
    상기 예측 오류 기준선을 기반으로, 상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합 비율을 제어하는 단계
    를 포함하는,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  16. 제 14 항에 있어서,
    상기 예측 오류 기준선은,
    환경의 변화에 대응하여, 가변적인,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  17. 제 14 항에 있어서,
    상기 예측 오류 기준선은,
    상기 저 분산성 지능 시스템과 상기 저 편향성 지능 시스템의 결합에 의해 달성되는 예측 오류 범위 내에서의 최소값을 나타내는,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  18. 제 14 항에 있어서,
    상기 저 분산성 지능 시스템은,
    모델-프리(MF) 강화학습 알고리즘을 포함하고,
    상기 저 편향성 지능 시스템은,
    모델-기반(MB) 강화학습 알고리즘을 포함하는,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  19. 제 14 항에 있어서,
    상기 제 1 예측 오류는,
    보상 예측 오류(RPE)이고,
    상기 제 2 예측 오류는,
    상태 예측 오류(SPE)인,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
  20. 제 16 항에 있어서,
    상기 방법은,
    상기 환경의 변화에 따라, 반복적으로 수행되는,
    비-일시적인 컴퓨터-판독 가능 저장 매체.
PCT/KR2021/019326 2021-03-18 2021-12-17 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법 Ceased WO2022196897A1 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020210035103A KR102578377B1 (ko) 2021-03-18 2021-03-18 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법
KR10-2021-0035103 2021-03-18

Publications (1)

Publication Number Publication Date
WO2022196897A1 true WO2022196897A1 (ko) 2022-09-22

Family

ID=83284717

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2021/019326 Ceased WO2022196897A1 (ko) 2021-03-18 2021-12-17 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법

Country Status (4)

Country Link
US (1) US12099333B2 (ko)
KR (1) KR102578377B1 (ko)
CN (1) CN115113726B (ko)
WO (1) WO2022196897A1 (ko)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200118017A1 (en) * 2018-10-12 2020-04-16 Adobe Inc. Cohort Event Prediction in a Digital Medium Environment using Regularization

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9959664B2 (en) * 2016-08-05 2018-05-01 Disney Enterprises, Inc. Adaptive polynomial rendering
CN108121204A (zh) * 2017-11-30 2018-06-05 上海航天控制技术研究所 一种组合体航天器姿态无模型的自适应控制方法和系统
KR102111857B1 (ko) * 2018-07-31 2020-05-15 한국과학기술원 인공지능 기반 게임 전략 유도 시스템 및 방법
US11480972B2 (en) * 2018-11-13 2022-10-25 Qualcomm Incorporated Hybrid reinforcement learning for autonomous driving
US20200272905A1 (en) * 2019-02-26 2020-08-27 GE Precision Healthcare LLC Artificial neural network compression via iterative hybrid reinforcement learning approach
EP3734507A1 (en) * 2019-05-03 2020-11-04 Essilor International Apparatus for machine learning-based visual equipment selection

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200118017A1 (en) * 2018-10-12 2020-04-16 Adobe Inc. Cohort Event Prediction in a Digital Medium Environment using Regularization

Non-Patent Citations (4)

* Cited by examiner, † Cited by third party
Title
ALI MAHMOUDZADEH; SOPHIA LIU; SOL SADEGHI; PAUL LUO LI; SOMIT GUPTA: "Bias Variance Tradeoff in Analysis of Online Controlled Experiments", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 10 September 2020 (2020-09-10), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081760154 *
CHENG RICHARD, VERMA ABHINAV, OROSZ GÁBOR, CHAUDHURI SWARAT, YUE YISONG, BURDICK JOEL W: "Control Regularization for Reduced Variance Reinforcement Learning", ARXIV:1905.05380V1, 14 May 2019 (2019-05-14), XP055967649, Retrieved from the Internet <URL:https://arxiv.org/pdf/1905.05380v1.pdf> [retrieved on 20221004] *
JOSE BLANCHET; FERNANDO HERNANDEZ; VIET ANH NGUYEN; MARKUS PELGER; XUHUI ZHANG: "Time-Series Imputation with Wasserstein Interpolation for Optimal Look-Ahead-Bias and Variance Tradeoff", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 25 February 2021 (2021-02-25), 201 Olin Library Cornell University Ithaca, NY 14853 , XP081891829 *
KIM DONGJAE, JEONG JAESEUNG, LEE SANG WAN: "Prefrontal solution to the bias-variance tradeoff during reinforcement learning", BIORXIV, 24 December 2020 (2020-12-24), XP055967648, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2020.12.23.424258v1.full.pdf> [retrieved on 20221004], DOI: 10.1101/2020.12.23.424258 *

Also Published As

Publication number Publication date
CN115113726B (zh) 2025-07-29
KR20220130343A (ko) 2022-09-27
US20220299948A1 (en) 2022-09-22
US12099333B2 (en) 2024-09-24
CN115113726A (zh) 2022-09-27
KR102578377B1 (ko) 2023-09-15

Similar Documents

Publication Publication Date Title
Kallimani et al. TinyML: Tools, applications, challenges, and future research directions
US7941392B2 (en) Scheduling system and method in a hierarchical temporal memory based system
US10474934B1 (en) Machine learning for computing enabled systems and/or devices
US10102449B1 (en) Devices, systems, and methods for use in automation
KR102785546B1 (ko) 기기에 탑재된 다수의 연합 학습 모델을 관리하는 방법, 시스템, 및 컴퓨터 프로그램
CA3221550A1 (en) Systems and methods for operating an autonomous system
US20260010773A1 (en) Temporal dynamics simulation in matmul-free neural architectures
Sato Context sensitive interactive systems design: A framework for representation of contexts
CN114981820A (zh) 用于在边缘设备上评估和选择性蒸馏机器学习模型的系统和方法
EP4558905A1 (en) Learning to combine explicit diversity conditions for effective question answer generation
KR20230134809A (ko) 언어 모델 압축을 위한 언어 컨텍스트 지식 증류의 방법 및 그를 수행하는 컴퓨터 시스템
WO2023017884A1 (ko) 디바이스에서 딥러닝 모델의 레이턴시를 예측하는 방법 및 시스템
US20240303505A1 (en) Federated learning method and device using device clustering
US11423225B2 (en) On-device lightweight natural language understanding (NLU) continual learning
KR102578377B1 (ko) 편향-분산 딜레마를 해결하는 뇌모사형 적응 제어를 위한 전자 장치 및 그의 방법
Kotilainen et al. The programmable world and its emerging privacy nightmare
Landauer et al. Model-based cooperative system engineering and integration
CN119724605A (zh) 医学大数据的可解释性分析和增量学习方法及系统
Batista et al. A new intelligent scheduler to improve reactive OpenFlow communication in SDN-based IoT data streams
Hafez Human digital twins: Two-layer machine learning architecture for intelligent human-machine collaboration
KR102590113B1 (ko) 인간의 불확실성 추론을 위한 컴퓨터 시스템 및 그의 방법
Vinogradov et al. Patterns in smart wireless sensor network nodes
Nielsen et al. Assistive and adaptive dialog management
Noguchi et al. Multi-modal shared module that enables the bottom-up formation of map representation and top-down map reading
Hörnle et al. Companion-Systems: A Reference Architecture

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21931838

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21931838

Country of ref document: EP

Kind code of ref document: A1