WO2022014861A1 - 강화학습을 이용한 약물 주입 조절 장치 및 방법 - Google Patents
강화학습을 이용한 약물 주입 조절 장치 및 방법 Download PDFInfo
- Publication number
- WO2022014861A1 WO2022014861A1 PCT/KR2021/006850 KR2021006850W WO2022014861A1 WO 2022014861 A1 WO2022014861 A1 WO 2022014861A1 KR 2021006850 W KR2021006850 W KR 2021006850W WO 2022014861 A1 WO2022014861 A1 WO 2022014861A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- state information
- patient
- anesthesia
- drug
- anesthesia state
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/168—Means for controlling media flow to the body or for metering media to the body, e.g. drip meters, counters ; Monitoring media flow to the body
- A61M5/172—Means for controlling media flow to the body or for metering media to the body, e.g. drip meters, counters ; Monitoring media flow to the body electrical or electronic
- A61M5/1723—Means for controlling media flow to the body or for metering media to the body, e.g. drip meters, counters ; Monitoring media flow to the body electrical or electronic using feedback of body parameters, e.g. blood-sugar, pressure
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/48—Other medical applications
- A61B5/4821—Determining level or depth of anaesthesia
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/168—Means for controlling media flow to the body or for metering media to the body, e.g. drip meters, counters ; Monitoring media flow to the body
- A61M5/16877—Adjusting flow; Devices for setting a flow rate
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
- G16H20/10—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to drugs or medications, e.g. for ensuring correct administration to patients
- G16H20/17—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to drugs or medications, e.g. for ensuring correct administration to patients delivered via infusion or injection
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/142—Pressure infusion, e.g. using pumps
- A61M2005/14208—Pressure infusion, e.g. using pumps with a programmable infusion control system, characterised by the infusion program
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/142—Pressure infusion, e.g. using pumps
- A61M2005/14288—Infusion or injection simulation
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/142—Pressure infusion, e.g. using pumps
- A61M2005/14288—Infusion or injection simulation
- A61M2005/14292—Computer-based infusion planning or simulation of spatio-temporal infusate distribution
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M5/00—Devices for bringing media into the body in a subcutaneous, intra-vascular or intramuscular way; Accessories therefor, e.g. filling or cleaning devices, arm-rests
- A61M5/14—Infusion devices, e.g. infusing by gravity; Blood infusion; Accessories therefor
- A61M5/142—Pressure infusion, e.g. using pumps
- A61M2005/14288—Infusion or injection simulation
- A61M2005/14296—Pharmacokinetic models
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61M—DEVICES FOR INTRODUCING MEDIA INTO, OR ONTO, THE BODY; DEVICES FOR TRANSDUCING BODY MEDIA OR FOR TAKING MEDIA FROM THE BODY; DEVICES FOR PRODUCING OR ENDING SLEEP OR STUPOR
- A61M2202/00—Special media to be introduced, removed or treated
- A61M2202/04—Liquids
- A61M2202/0468—Liquids non-physiological
- A61M2202/048—Anaesthetics
Definitions
- the present invention relates to an apparatus and method for controlling drug injection using reinforcement learning, and more particularly, to an apparatus and method for controlling drug injection for maintaining anesthesia of a patient during surgery.
- an anesthesiologist adjusts the injection rate of a drug in order to induce a patient's condition into a constant anesthetic state.
- the injection rate of the drug is directly controlled so that the patient's condition is kept constant.
- the anesthesiologist uses an infusion pump to inject the drug into the patient.
- PK Pharmacokinetics
- PD Pharmacodynamics
- maintaining the patient's state in a constant anesthetic state is controlled by the anesthetist, and accordingly, it may be difficult to maintain the anesthetic state depending on the health state and skill level of the anesthetist. For example, if the drug injected into the patient is insufficient, the patient may regain consciousness during surgery, and if the drug is injected into the patient excessively, hemodynamic instability and other side effects may occur .
- the technical problem to be solved by the present invention is to provide a drug injection control device and method for stably maintaining anesthesia state of a patient during surgery using reinforcement learning.
- anesthesia state information calculating unit for calculating the patient's anesthesia state information; Set target anesthesia state information set in advance for the anesthesia state information, and learn a compensation value according to a change in the anesthesia state information by a drug injection rate set so that the anesthesia state information follows the target anesthesia state information
- the policy model learning unit for generating the policy model
- a predictive model learning unit for generating a predictive model by learning the anesthesia state information that changes according to the change in the drug injection rate
- a control unit configured to set the drug injection rate from the anesthesia state information based on the policy model
- a predictor for predicting expected anesthesia state information from a set drug infusion rate and a previously set drug infusion rate, based on the predictive model.
- the anesthesia state information calculation unit calculates the effect site concentration and the plasma concentration, respectively, based on the drug model prepared in advance for the lean mass of the patient and the drug injected into the patient, and the effect site concentration and the plasma concentration Accordingly, the anesthesia state information may be calculated to indicate the patient's state.
- control unit the difference between the anesthetic state information and the target anesthetic state information calculated from the patient's condition, the injection rate of remifentanil injected into the patient during a preset time interval, and the injection rate of the remifentanil to the patient during the time interval
- the injection rate of the remifentanil and the injection rate of the propofol may be controlled according to the injection rate of propofol.
- the policy model learning unit is configured to control the injection rate of the remifentanil and the propofol to control the anesthesia state information and the target anesthesia state information calculated from the changed state of the patient after the time interval has elapsed.
- a compensation value can be calculated according to the difference.
- the policy model learning unit based on a plurality of compensation values calculated from changes in different anesthesia state information, from arbitrary anesthesia state information, a preset time interval for the change of the anesthesia state information elapses several times.
- an expected value may be calculated according to a plurality of compensation values matching each change in anesthesia state information.
- the policy model learning unit may generate the policy model according to the drug injection rate in the change of the anesthesia state information matching the compensation value selected so that the expected value is calculated as the maximum value.
- Another aspect of the present invention provides a drug injection control method using a drug injection control device using reinforcement learning, the method comprising: calculating anesthesia state information of a patient; Set target anesthesia state information set in advance with respect to the anesthesia state information, and learn a compensation value according to a change in the anesthesia state information by a drug injection rate set so that the anesthesia state information follows the target anesthesia state information to create a policy model; generating a predictive model by learning the anesthesia state information that changes according to the change in the drug injection rate; setting the drug infusion rate from the anesthesia state information based on the policy model; and predicting the expected anesthesia state information from the set drug infusion rate and the previously set drug infusion rate, based on the prediction model.
- the calculating of the anesthesia state information includes calculating the effect site concentration and plasma concentration, respectively, based on the drug model prepared in advance for the patient's lean mass and the drug injected into the patient, and the effect site concentration and The anesthesia state information may be calculated to indicate the patient's state according to the plasma concentration.
- the setting of the drug infusion rate includes: a difference between the anesthesia state information calculated from the patient's state and the target anesthetic state information, the injection rate and the time of remifentanil injected into the patient during a preset time interval
- the infusion rate of the remifentanil and the propofol infusion rate may be controlled according to the propofol infusion rate injected into the patient during the interval.
- the injection rate of remifentanil and the injection rate of propofol are controlled, and after the time interval elapses, the anesthesia state information and the target anesthesia calculated from the changed state of the patient A compensation value may be calculated according to a difference in state information.
- a preset time interval for a change in the anesthetic state information is obtained from arbitrary anesthesia state information.
- an expected value may be calculated according to a plurality of compensation values matching each change of anesthesia state information.
- the generating of the policy model may include generating the policy model according to a drug injection rate in a change in the anesthesia state information matching a compensation value selected so that the expected value is calculated as a maximum value.
- FIG. 1 is a schematic diagram of a drug infusion control device according to an embodiment of the present invention.
- FIG. 2 is a control block diagram of a device for controlling drug injection according to an embodiment of the present invention.
- FIG. 3 is a block diagram illustrating a process of generating a policy model in the policy model learning unit of FIG. 2 .
- FIG. 4 is a block diagram illustrating a process of generating a predictive model in the predictive model learning unit of FIG. 2 .
- FIG. 5 is a flowchart of a method for controlling drug injection according to an embodiment of the present invention.
- FIG. 6 is a detailed flowchart of the step of generating the policy model of FIG.
- FIG. 1 is a schematic diagram of a drug infusion control device according to an embodiment of the present invention.
- the drug injection control apparatus 100 may be provided to adjust the injection rate of the drug injected into the patient to constantly maintain the anesthesia state of the patient undergoing surgery.
- the anesthesia state of the patient may be indicated using a Bispectral Index (BIS), and in general, it may be recommended to maintain a BIS of 50 of the patient being anesthetized during a surgical operation.
- BIOS Bispectral Index
- the drug infusion control device 100 is provided to maintain the patient's BIS constant by adjusting the injection rate of remifentanil and propofol injected into the patient, respectively.
- the drug injection control apparatus 100 may calculate the patient's anesthesia state information.
- the anesthesia state information may refer to a BIS calculated or measured from the patient's state.
- the drug injection control device 100 may measure the depth of anesthesia of the patient by using a depth of anesthesia monitoring device or an anesthesia depth monitoring device provided to measure the patient's anesthesia state, and in this case, the patient's
- the anesthesia state information may mean a depth of anesthesia of a patient measured by an anesthesia depth monitoring device or an anesthesia depth monitoring device.
- the drug injection control device 100 may calculate the patient's lean body mass according to the patient's height and weight, and the drug injection control device 100 determines the patient's lean body mass and the drug injected into the patient in advance. Based on the prepared drug model, an effect site concentration and a plasma concentration may be calculated, respectively.
- Equations 1 to 5 below can be understood as equations used to calculate the concentration of the effect site and the plasma concentration of the patient.
- lbm_male may mean the lean mass of a man
- lbm_female may mean the lean mass of a woman
- Weight may mean the weight of the patient
- Height may mean the height of the patient.
- PPF may mean propofol
- h_1 to h_17 may be understood as variable values by the Schinider model, which is a drug model defined for propofol
- lbm may mean lean mass of the patient
- age is the patient can mean the age of
- Vol_1 ⁇ PPF to Vol_3 ⁇ PPF can be understood as the amount of drug for each component of propofol
- C_l1 ⁇ PPF to C_13 ⁇ PPF can be understood as representing the rate at which the drug is removed from the body for each component of propofol.
- k_ij ⁇ PPF can be understood as a variable indicating the change or modification of the drug in the body for each component of propofol.
- Table 1 is a table showing the variables for each component of propofol in any patient condition according to the Schinider model.
- the Schinider model is 'The influence of method of administration and covariates on the pharmacokinetics of propofol in adult volunteers (T. Schnider, C. Minto, P. Gambus, C. Andresen, D. Goodale, S. Shafer, and E. Youngs, Anesthesiology, vol. 88, no. 5, pp. 1170-1182, 1998.)', a detailed description thereof will be omitted.
- RFTN may mean remifentanil
- f_1 to f_17 may be understood as variable values by the Minto model, which is a drug model defined for remifentanil
- lbm may mean lean mass of a patient
- age may mean the age of the patient.
- Vol_1 ⁇ RFTN to Vol_3 ⁇ RFTN can be understood as the amount of drug for each component of remifentanil
- C_11 ⁇ RFTN to C_13 ⁇ RFTN represent the rate at which the drug is removed from the body for each component of remifentanil.
- k_ij ⁇ RFTN can be understood as a variable representing the change or modification of the drug in the body for each component of remifentanil.
- Table 2 is a table showing the parameters for each component of remifentanil in any patient condition, according to the Minto model.
- Minto model is 'Influence of Age and Gender on the Pharmacokinetics and Pharmacodynamics of Remifentanil.
- I. Model development C. Minto, T. Schnider, T. Egan, E. Youngs, H. Lemmens, P. Gambus, V. Billard, J. Hoke, K. Moore, D. Hermann et al., Anesthesiology, vol. 86, no. 1, pp. 10-23, 1997.
- xi(t) may be understood as a variable representing the amount of each component of the drug over time
- k_ij may be understood as a variable representing the change or transformation of the drug in the body with respect to each component of the drug.
- x_e(t) can be understood as a variable representing the amount of drug at the site of effect of the drug over time
- C_p(t) can mean the plasma concentration of the drug
- C_e(t) is the may refer to an effect site concentration indicating the concentration at the effect site.
- the drug injection control device 100 calculates the effect site concentration and plasma concentration based on a drug model prepared in advance for the patient's lean mass and the drug injected into the patient according to Equations 1 to 5. each can be calculated.
- the drug infusion control device 100 may calculate the respective effective site concentration and plasma concentration for propofol or remifentanil. You can also input information such as
- the drug injection control apparatus 100 may calculate anesthesia state information to indicate the patient's state according to the calculated effect site concentration and plasma concentration.
- Equation 6 can be understood as a formula for calculating anesthesia state information from the concentration of the effect site and the plasma concentration.
- BIS may mean a patient's BIS value used as anesthesia state information.
- C_e ⁇ PPF may mean a calculated effect site concentration for propofol
- C_e ⁇ RFTN may mean a calculated effect site concentration for remifentanil
- Epsilon may mean Gaussian noise for BIS can do.
- the drug injection control apparatus 100 may calculate the patient's anesthesia state information according to a preset time interval.
- the drug injection control apparatus 100 may set target anesthesia state information set in advance with respect to the patient's anesthetic state information, and the drug injection control apparatus 100 may set the target anesthetic state information calculated according to the patient's state of anesthesia state information.
- a compensation value may be calculated according to a change in anesthesia state information due to a drug injection rate set to follow the information, and the drug injection control apparatus 100 may learn the calculated compensation value to generate a policy model.
- the target anesthesia state information may mean a numerical value of the anesthesia state information provided so that the patient can maintain a stable anesthesia state, for example, the target anesthesia state information is 50 of the BIS values between 0 and 100. can be set to
- the drug infusion rate may include an infusion rate of propofol injected into the patient and an infusion rate of remifentanil injected into the patient, wherein the propofol injection rate and remifentanil injection rate are different from each other. It can be set to be
- Equation 7 may be understood as an expression representing a compensation value calculated according to a change in anesthesia state information.
- R may mean a compensation value calculated according to a change in anesthesia state information due to the patient's state and drug injection rate
- s_t is a state prepared to indicate the patient's anesthetic state information and drug injection rate It may mean a variable
- a_t may mean an operation variable provided to indicate the injection rate of the drug controlled by the drug injection control device 100 .
- g_t may mean target anesthesia state information prepared so that the patient can maintain a stable anesthesia state
- the anesthesia state information may mean anesthesia state information calculated from the patient's state
- BIS_error ⁇ hist is It may mean a change amount of anesthesia state information changed according to a set time interval.
- Alpha may mean a preset maximum compensation value.
- the drug injection control device 100 calculates a compensation value expressed as the reciprocal of the result of adding the maximum compensation value set in advance to the absolute value of the difference between the anesthesia state information calculated from the patient's condition and the target anesthetic state information can do.
- the state variables are the difference between the anesthetic state information calculated from the patient's state and the target anesthetic state information, the infusion rate of remifentanil to be injected into the patient during a preset time interval, and to the patient during the preset time interval.
- the rate of infusion of propofol to be injected may be included.
- the operating variable may be provided to indicate an infusion rate of a drug to be injected into the patient during a preset time interval.
- the drug infusion control device 100 may set a preset time interval to 10 seconds, and the drug infusion control device 100 may set the maximum amount of propofol injected into the patient for 10 seconds to 27.8 mg, and , the drug injection control device 100 may set the maximum amount of remifentanil injected into the patient for 10 seconds to 27.8 ug.
- the operating parameter for propofol can be set at an infusion rate at which an amount of propofol is infused between 0 and 27.8 mg per 10 sec, and the operating parameter for remifentanil is that the parameter is between 0 and 27.8 ⁇ g of remi per 10 sec.
- the amount of fentanyl can be set to the infusion rate at which it is injected.
- the compensation value is the difference between the anesthesia state information calculated from the patient's condition and the target anesthetic state information, the infusion rate of remifentanil to be injected into the patient during a preset time interval, and injection into the patient during a preset time interval.
- it may be a value provided to indicate compensation for the result of controlling the infusion rate of remifentanil injected into the patient and the infusion rate of propofol injected into the patient.
- the result of controlling the injection rate of the drug can be understood as meaning a state according to the state variable after a preset time interval elapses after controlling the injection rate of remifentanil and propofol.
- the drug injection control device 100 controls the injection rate of remifentanil and propofol, and after a preset time interval elapses, anesthesia state information and target anesthesia state calculated from the changed patient state A compensation value can be calculated according to the difference in information.
- the drug injection control apparatus 100 performs the preset time interval for the change of the anesthetic state information from arbitrary anesthesia state information based on a plurality of compensation values calculated from the different changes in the anesthetic state information several times.
- an expected value may be calculated according to a plurality of compensation values that match each change of anesthesia state information.
- Equation 8 is an equation prepared to calculate an expected value from the compensation value.
- J(Phi_Theta) may mean an expected value prepared to indicate a compensation value calculated when a preset time interval for a change in anesthesia state information elapses several times
- Phi_Theta) is It may mean a probability value provided to indicate a probability that an arbitrary operation variable occurs in an arbitrary state variable
- R(Tau) may mean a compensation value.
- the drug injection control device 100 calculates the expected value represented by the sum of the product of the probability value and the compensation value at each time interval. can be calculated.
- Equation 9 is an equation for calculating a probability value.
- Phi_Theta) may mean a probability value provided to indicate the probability that an arbitrary operation variable occurs in an arbitrary state variable
- Rho(S_0) is anesthesia state information measured first from a patient and It may mean an initial state variable prepared to include a difference in target anesthetic state information, an infusion rate of remifentanil initially injected into a patient, and an infusion rate of propofol initially injected into a patient.
- s_t, a_t) may mean a probability value that a state variable changes to an arbitrary state variable after a time interval preset by an arbitrary state variable and an arbitrary operation variable has elapsed.
- s_t, g_t) may refer to a policy model prepared to select a specific operation variable according to an arbitrary state variable and target anesthesia state information.
- the drug injection control apparatus 100 may calculate a probability value expressed as a value obtained by multiplying a value obtained by multiplying a value obtained by multiplying a product of a policy model and a probability value at which a state variable changes at different times is multiplied by an initial state variable.
- the drug injection control apparatus 100 may generate a policy model according to the drug injection rate in the change of the anesthesia state information matching the compensation value selected so that the expected value is calculated as the maximum value.
- the drug injection control device 100 may learn a compensation value according to a change in anesthesia state information according to an arbitrary drug injection rate, and select a compensation value in which an expected value is calculated as a maximum value, and the drug injection control device ( 100 ) may create a policy model such that the drug injection rate for a change in the anesthesia state information matching the selected compensation value appears as an action variable.
- the drug injection control apparatus 100 learns a compensation value from a plurality of data sets prepared in advance to include a change in anesthesia state information according to a preset time interval for an arbitrarily set drug injection rate, thereby learning a policy model. can also create
- the drug injection control apparatus 100 performs an arbitrary operation in an arbitrary state, learns a reward according to a changing state, generates a policy, and performs an operation in an arbitrary state according to the generated policy.
- a policy model can be created using a reinforcement learning technique of your choice.
- the drug injection control apparatus 100 may generate a predictive model by learning a change in anesthesia state information according to a drug injection rate.
- the drug injection control device 100 provides information such as age, sex, weight and height of the patient, the difference between the anesthesia state information calculated from the patient's state according to a preset time interval and the target anesthetic state information, A change in the infusion rate of remifentanil to be injected into the patient during a preset time interval and the infusion rate of propofol injected into the patient during a preset time interval may be learned.
- the drug injection control device 100 records information such as the patient's age, sex, weight, and height, and a change in a state variable according to a preset time interval in a Deep Neural Network (DNN) or Long Short-Term Memory (LSTM), respectively. models) can be provided.
- DNN Deep Neural Network
- LSTM Long Short-Term Memory
- the DNN is a technique for outputting feature information from information input to the input layer by providing a plurality of hidden layers between the input layer and the output layer. It is a technique of adjusting the amount of output information based on currently input information.
- the drug injection control device 100 may input information output from the DNN or LSTM into one DNN to fuse the information. It may be provided to input the output information to another DNN which is provided to output the expected anesthesia state information.
- the drug injection control device 100 provides information on the patient's age, sex, weight and height, and the difference between the anesthetic state information and the target anesthetic state information calculated from the patient's state according to a preset time interval, Expected anesthesia information may be generated by using the change in the infusion rate of remifentanil to be injected into the patient during the time interval set in , and the infusion rate of propofol injected into the patient during the preset time interval.
- the drug injection control apparatus 100 may set the drug injection rate from the anesthesia state information based on the policy model.
- the drug injection control device 100 determines the difference between the anesthesia state information calculated from the patient's state and the target anesthetic state information, the injection rate of remifentanil injected into the patient during a preset time interval, and a preset time interval.
- the infusion rate of remifentanil and propofol infusion rate can be controlled according to the propofol infusion rate to be injected into the patient during the period.
- the drug injection control apparatus 100 may select an operation variable set by the policy model according to the current state variable, and control the injection speed of the drug injected into the patient according to the operation variable.
- the drug injection control apparatus 100 may predict the expected anesthetic state information from the set drug infusion rate and the previously set drug infusion rate, based on the prediction model.
- the drug injection control device 100 may output the predicted expected anesthesia state information in the form of a graph or a numerical value so that the user can check it, and for this, the drug injection control device 100 is a separate display device may include
- the drug injection control apparatus 100 calculates a compensation value according to the patient's state according to the expected anesthetic state information from the patient's state according to the anesthetic state information calculated from the patient's state by using the predicted expected anesthetic state information Also, in this case, the drug injection control apparatus 100 may generate a policy model by learning a compensation value based on the expected anesthetic state information.
- the drug injection control device 100 may measure the depth of anesthesia of the patient by using a depth of anesthesia monitoring device or an anesthesia depth monitoring device provided to measure the anesthesia state of the patient, in this case, the drug injection
- the control apparatus 100 may generate a policy model according to the depth of anesthesia measured from the patient and the target depth of anesthesia set so that the anesthesia state of the patient can be stably maintained.
- the apparatus for measuring the patient's anesthesia state such as the depth of anesthesia monitoring apparatus or the depth of anesthesia monitoring apparatus
- the apparatus for measuring the patient's anesthesia state may be an apparatus according to any apparatus, apparatus, and method using known techniques.
- FIG. 2 is a control block diagram of a device for controlling drug injection according to an embodiment of the present invention.
- the drug injection control apparatus 100 may include an anesthesia state information calculator 110 , a policy model learner 120 , a predictive model learner 130 , a controller 140 , and a predictor 150 .
- the anesthesia state information calculator 110 may calculate the patient's anesthesia state information.
- the anesthesia state information calculator 110 may receive information such as the patient's age, sex, weight and height. .
- the anesthesia state information calculating unit 110 may calculate the patient's lean body mass according to the patient's height and weight, and the anesthetic state information calculating unit 110 calculates the patient's lean body mass and the drug injected into the patient. Based on the drug model prepared in advance, the effect site concentration and the plasma concentration can be calculated, respectively. At this time, the anesthesia state information calculating unit 110 may calculate the concentration and plasma concentration of each effect site for propofol or remifentanil.
- the anesthesia state information calculator 110 may calculate the anesthesia state information to indicate the patient's state according to the calculated effect site concentration and plasma concentration. According to the set time interval, it is possible to calculate the patient's anesthesia state information.
- the policy model learning unit 120 may set target anesthesia state information set in advance with respect to the patient's anesthesia state information, and the policy model learning unit 120 sets the target anesthesia state information calculated according to the patient's state to the target anesthetic state information.
- a compensation value may be calculated according to a change in anesthesia state information due to a drug injection rate set to follow the information, and the policy model learning unit 120 may learn the calculated compensation value to generate a policy model.
- the policy model learning unit 120 calculates a compensation value expressed as the reciprocal of the result of adding the maximum compensation value set in advance to the absolute value of the difference between the anesthesia state information calculated from the patient's condition and the target anesthesia state information.
- the policy model learning unit 120 controls the injection rate of remifentanil and propofol, and after a preset time interval elapses, anesthesia state information and target anesthesia state calculated from the changed patient state A compensation value can be calculated according to the difference in information.
- the policy model learning unit 120 is based on a plurality of compensation values calculated from changes in different anesthesia state information, from arbitrary anesthesia state information, a preset time interval for changes in anesthesia state information has elapsed several times.
- an expected value may be calculated according to a plurality of compensation values that match each change in anesthesia state information.
- the policy model learning unit 120 calculates an expected value represented by the sum of the product of the probability value and the compensation value at each time interval. can be calculated.
- the policy model learning unit 120 may calculate a probability value expressed as a value obtained by multiplying a value obtained by multiplying a product of a policy model and a probability value at which a state variable changes at different times is multiplied by an initial state variable.
- the policy model learning unit 120 may generate the policy model according to the drug injection rate in the change of the anesthesia state information matching the compensation value selected so that the expected value is calculated as the maximum value.
- the policy model learning unit 120 may learn a compensation value according to a change in anesthesia state information by an arbitrary drug injection rate, and select a compensation value in which an expected value is calculated as a maximum value, and the policy model learning unit 120 may create a policy model so that the drug injection rate for a change in the anesthesia state information matching the selected compensation value appears as an operation variable.
- the policy model learning unit 120 learns a compensation value from a plurality of data sets prepared in advance to include a change in anesthesia state information according to a preset time interval with respect to an arbitrarily set drug injection rate to learn the policy model. can also create
- the policy model learning unit 120 performs an arbitrary operation in an arbitrary state to learn a reward according to a changing state, generates a policy, and performs an operation performed in an arbitrary state according to the generated policy.
- a policy model can be created using a reinforcement learning technique of your choice.
- the policy model learning unit 120 calculates a compensation value according to the patient's state according to the expected anesthetic state information from the patient's state according to the anesthetic state information calculated from the patient's state by using the predicted expected anesthetic state information.
- the policy model learning unit 120 may generate a policy model by learning a compensation value based on the expected anesthesia state information.
- the predictive model learning unit 130 may generate a predictive model by learning a change in anesthesia state information according to a drug injection rate.
- the predictive model learning unit 130 for information such as the patient's age, gender, weight and height, the difference between the anesthesia state information and the target anesthesia state information calculated from the patient's state according to a preset time interval, A change in the infusion rate of remifentanil to be injected into the patient during a preset time interval and the infusion rate of propofol injected into the patient during a preset time interval may be learned.
- the controller 140 may set the drug injection rate from the anesthesia state information based on the policy model.
- control unit 140 controls the difference between the anesthesia state information calculated from the patient's condition and the target anesthetic state information, the injection rate of remifentanil to be injected into the patient for a preset time interval, and the patient for a preset time interval.
- the injection rate of remifentanil and the injection rate of propofol can be controlled according to the injection rate of propofol to be injected.
- the controller 140 may select an operation variable set by the policy model according to the current state variable, and control the injection speed of the drug injected into the patient according to the operation variable.
- the prediction unit 150 may predict the expected anesthetic state information from the set drug injection rate and the previously set drug injection rate, based on the prediction model.
- the prediction unit 150 may output the predicted predicted anesthetic state information in the form of a graph or a numerical value so that the user can check it, and for this, the drug injection control device 100 includes a separate display device can do.
- FIG. 3 is a block diagram illustrating a process of generating a policy model in the policy model learning unit of FIG. 2 .
- the anesthesia state information calculator 110 may calculate the patient's anesthesia state information.
- the anesthesia state information calculator 110 calculates the patient's age, gender, weight and height You can also enter information.
- the anesthesia state information calculating unit 110 may calculate the patient's lean body mass according to the patient's height and weight, and the anesthetic state information calculating unit 110 calculates the patient's lean body mass and the drug injected into the patient. Based on the drug model prepared in advance, the effect site concentration and the plasma concentration can be calculated, respectively. At this time, the anesthesia state information calculating unit 110 may calculate the concentration and plasma concentration of each effect site for propofol or remifentanil.
- the anesthesia state information calculator 110 may calculate the anesthesia state information to indicate the patient's state according to the calculated effect site concentration and plasma concentration. According to the set time interval, it is possible to calculate the patient's anesthesia state information.
- the policy model learning unit 120 may set target anesthesia state information set in advance with respect to the patient's anesthesia state information, and the policy model learning unit 120 determines that the patient's anesthetic state information follows the target anesthetic state information.
- a policy model can be created by learning a compensation value according to a change in anesthesia state information according to a drug injection rate set to do so.
- the policy model learning unit 120 calculates a compensation value expressed as the reciprocal of the result of adding the maximum compensation value set in advance to the absolute value of the difference between the anesthesia state information calculated from the patient's condition and the target anesthesia state information.
- the policy model learning unit 120 controls the injection rate of remifentanil and propofol, and after a preset time interval elapses, anesthesia state information and target anesthesia state calculated from the changed patient state A compensation value can be calculated according to the difference in information.
- the policy model learning unit 120 is based on a plurality of compensation values calculated from changes in different anesthesia state information, and a preset time interval for changes in anesthesia state information from arbitrary anesthesia state information elapses several times.
- an expected value may be calculated according to a plurality of compensation values matching each change in anesthesia state information.
- the policy model learning unit 120 calculates an expected value represented by the sum of the product of the probability value and the compensation value at each time interval. can be calculated.
- the policy model learning unit 120 may calculate a probability value expressed as a value obtained by multiplying a value obtained by multiplying a product of a policy model and a probability value at which a state variable changes at different times, and an initial state variable.
- the policy model learning unit 120 may generate the policy model according to the drug injection rate in the change of the anesthesia state information matching the compensation value selected so that the expected value is calculated as the maximum value.
- the policy model learning unit 120 may learn a compensation value according to a change in anesthesia state information by an arbitrary drug injection rate, and select a compensation value in which an expected value is calculated as a maximum value, and the policy model learning unit 120 may create a policy model such that the drug injection rate for a change in the anesthesia state information matching the selected compensation value appears as an operation variable.
- the policy model learning unit 120 learns a compensation value from a plurality of data sets prepared in advance to include a change in anesthesia state information according to a preset time interval with respect to an arbitrarily set drug injection rate to learn the policy model. can also create
- the controller 140 may set the drug injection rate from the anesthesia state information based on the policy model.
- control unit 140 controls the difference between the anesthesia state information calculated from the patient's condition and the target anesthetic state information, the injection rate of remifentanil to be injected into the patient for a preset time interval, and the patient for a preset time interval.
- the injection rate of remifentanil and the injection rate of propofol can be controlled according to the injection rate of propofol to be injected.
- the controller 140 may select an operation variable set by the policy model according to the current state variable, and control the injection speed of the drug injected into the patient according to the operation variable.
- FIG. 4 is a block diagram illustrating a process of generating a predictive model in the predictive model learning unit of FIG. 2 .
- the anesthesia state information calculator 110 may calculate the patient's anesthesia state information.
- the anesthesia state information calculator 110 calculates the patient's age, gender, weight and height, etc. You can also enter information.
- the anesthesia state information calculating unit 110 may calculate the patient's lean body mass according to the patient's height and weight, and the anesthetic state information calculating unit 110 calculates the patient's lean body mass and the drug injected into the patient. Based on the drug model prepared in advance, the effect site concentration and the plasma concentration can be calculated, respectively. At this time, the anesthesia state information calculator 110 may calculate the concentration and plasma concentration of each effect site for propofol or remifentanil.
- the anesthesia state information calculator 110 may calculate the anesthesia state information to indicate the patient's state according to the calculated effect site concentration and plasma concentration. According to the set time interval, it is possible to calculate the patient's anesthesia state information.
- the predictive model learning unit 130 may generate a predictive model by learning the anesthesia state information that changes according to a change in the drug injection rate.
- the predictive model learning unit 130 for information such as the patient's age, gender, weight and height, the difference between the anesthesia state information and the target anesthesia state information calculated from the patient's state according to a preset time interval, A change in the infusion rate of remifentanil to be injected into the patient during a preset time interval and the infusion rate of propofol injected into the patient during a preset time interval may be learned.
- the prediction unit 150 may predict the expected anesthetic state information from the set drug injection rate and the previously set drug injection rate, based on the prediction model.
- the prediction unit 150 may output the predicted predicted anesthetic state information in the form of a graph or a numerical value so that the user can check it, and for this, the drug injection control device 100 includes a separate display device can do.
- the policy model learning unit 120 calculates a compensation value according to the patient's state according to the expected anesthetic state information from the patient's state according to the anesthetic state information calculated from the patient's state by using the predicted expected anesthetic state information. Also, in this case, the policy model learning unit 120 may generate a policy model by learning a compensation value based on the expected anesthesia state information.
- FIG. 5 is a flowchart of a method for controlling drug injection according to an embodiment of the present invention.
- the drug injection control method according to an embodiment of the present invention proceeds in substantially the same configuration as the drug injection control device 100 shown in FIG. 1, for the same components as the drug injection control device 100 of FIG.
- the same reference numerals are given, and repeated descriptions will be omitted.
- the drug injection control method includes calculating anesthesia state information (600), generating a policy model (610), generating a predictive model (620), setting a drug injection rate (630), and expected anesthesia state predicting the information ( 640 ).
- the step 600 of calculating the anesthesia state information may be a step in which the anesthesia state information calculating unit 110 calculates the anesthesia state information of the patient.
- step 610 of generating the policy model the policy model learning unit 120 sets target anesthesia state information set in advance for the anesthesia state information, and the calculated anesthesia state information is set to follow the target anesthesia state information. It may be a step of calculating a compensation value according to a change in anesthesia state information according to the drug injection rate, learning the compensation value, and generating a policy model.
- the step 620 of generating the predictive model may be a step in which the predictive model learning unit 130 learns the change in the anesthesia state information according to the drug injection rate to generate the predictive model.
- the step 630 of setting the drug injection rate may be a step in which the controller 140 sets the drug injection rate from the anesthesia state information based on the policy model.
- Predicting the expected anesthesia state information 640 may be a step in which the prediction unit 150 predicts the expected anesthetic state information from a set drug injection rate and a previously set drug injection rate based on the prediction model.
- FIG. 6 is a detailed flowchart of the step of generating the policy model of FIG.
- Step 610 of generating the policy model includes the steps of calculating a change in anesthesia state information (611), calculating a compensation value (612), calculating an expected value (613), and selecting a compensation value ( 614) may be included.
- the policy model learning unit 120 relates to the difference between the anesthetic state information calculated from the patient's state and the target anesthetic state information, the injection rate of remifentanil and propofol injection After the speed is controlled and a preset time interval has elapsed, it may be a step of calculating a change in the anesthetic state information expressed as a difference between the anesthetic state information calculated from the changed patient state and the target anesthetic state information.
- step 612 of calculating the compensation value the policy model learning unit 120 adds the maximum compensation value set in advance to the absolute value of the difference between the anesthesia state information calculated from the patient's condition and the target anesthesia state information. It may be a step of calculating a compensation value expressed as a reciprocal number.
- the policy model learning unit 120 controls the injection rate of remifentanil and propofol, and after a preset time interval elapses, the changed patient's state It may be a step of calculating a compensation value according to a difference between the anesthesia state information calculated from the , and the target anesthesia state information.
- the policy model learning unit 120 responds to a change in anesthesia state information from arbitrary anesthesia state information, based on a plurality of compensation values calculated from different anesthesia state information changes. It may be a step of calculating an expected value according to a plurality of compensation values matching each change of anesthesia state information in a process in which a preset time interval elapses several times.
- the policy model learning unit 120 learns a compensation value according to a change in anesthesia state information due to an arbitrary drug injection rate, and selects a compensation value in which an expected value is calculated as the maximum value , and thus, the policy model learning unit 120 may generate a policy model such that the drug injection rate for a change in the anesthesia state information matching the selected compensation value appears as an operation variable.
- control unit 140 control unit
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Public Health (AREA)
- General Health & Medical Sciences (AREA)
- Veterinary Medicine (AREA)
- Heart & Thoracic Surgery (AREA)
- Biomedical Technology (AREA)
- Animal Behavior & Ethology (AREA)
- Anesthesiology (AREA)
- Hematology (AREA)
- Vascular Medicine (AREA)
- Medical Informatics (AREA)
- Physics & Mathematics (AREA)
- Surgery (AREA)
- Biophysics (AREA)
- Pathology (AREA)
- Molecular Biology (AREA)
- Fluid Mechanics (AREA)
- Diabetes (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Medicinal Chemistry (AREA)
- Epidemiology (AREA)
- Primary Health Care (AREA)
- Infusion, Injection, And Reservoir Apparatuses (AREA)
Abstract
환자의 마취상태 정보가 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 마취상태 정보의 변화를 학습하여, 정책 모델을 생성하고, 약물 주입 속도의 변화에 따른 마취상태 정보의 변화를 학습하여, 예측 모델을 생성하며, 정책 모델에 기초하여, 마취상태 정보로부터 약물 주입 속도를 설정하고, 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는, 약물 주입 조절 장치를 제공한다. [대표도] 도 1
Description
본 발명은 강화학습을 이용한 약물 주입 조절 장치 및 방법에 관한 것으로, 보다 상세하게는, 수술 중에 환자의 마취 상태를 유지하기 위한 약물 주입 조절 장치 및 방법에 관한 것이다.
일반적으로, 마취 전문의는 환자의 상태를 일정한 마취 상태로 유도하기 위해 약물의 주입 속도를 조절하게 된다. 현대에는, 환자의 BIS(Bispectral Index) 또는 뇌전도(EEG: Electroencephalogram)를 측정하여, 이러한 환자의 상태가 일정하게 유지되도록 약물의 주입 속도를 직접 조절하게 된다. 이때, 마취 전문의는 약물을 환자에게 주입하기 위해 주입 펌프를 이용하게 된다.
이때, 환자를 최면 및 무통을 포함하는 마취 상태 또는 진정 상태로 유도하는 과정에서, 마취 전문의는 프로포폴(Propofol) 및 레미펜타닐(Remifentanil) 등의 신속하게 작용되는 약물을 이용하게 되며, 이 과정에서, 약동학(PK: Pharmacokinetics) - 약력학(PD: Pharmacodynamics)이 이용된다.
한편, 환자의 상태를 일정한 마취 상태로 유지하는 것은 마취 전문의에 의해 제어되며, 이에 따라, 마취 전문의의 건강 상태, 숙련도 등에 따라 마취 상태를 유지하는데 어려움이 발생할 수 있다. 예를 들어, 환자에게 주입되는 약물이 부족한 경우에, 환자는 수술 중에 의식을 회복하게 될 수 있으며, 환자에게 약물이 과다하게 주입되는 경우에는, 환자의 혈역학적 불안정성 및 다른 부작용 등을 유발할 수 있다.
이에 따라, 환자의 마취 상태가 일정하게 유지되도록 환자에게 주입되는 약물을 적절히 조절하는 방안이 요구되는 실정이다.
본 발명이 해결하고자 하는 기술적 과제는 강화학습을 이용하여 수술 중에 환자의 마취 상태를 안정적으로 유지하는 약물 주입 조절 장치 및 방법을 제공하는 것이다.
본 발명의 일측면은, 환자의 마취상태 정보를 산출하는 마취상태 정보 산출부; 상기 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정하고, 상기 마취상태 정보가 상기 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 상기 마취상태 정보의 변화에 따른 보상 값을 학습하여, 정책 모델을 생성하는 정책 모델 학습부; 상기 약물 주입 속도의 변화에 따라 변하는 상기 마취상태 정보를 학습하여, 예측 모델을 생성하는 예측 모델 학습부; 상기 정책 모델에 기초하여, 상기 마취상태 정보로부터 상기 약물 주입 속도를 설정하는 제어부; 및 상기 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는 예측부를 포함할 수 있다.
또한, 상기 마취상태 정보 산출부는, 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출하고, 상기 효과 부위 농도와 상기 혈장 농도에 따라 환자의 상태를 나타내도록 상기 마취상태 정보를 산출할 수 있다.
또한, 상기 제어부는, 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 상기 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도를 제어할 수 있다.
또한, 상기 정책 모델 학습부는, 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도가 제어되어, 상기 시간 간격이 경과된 이후, 변화된 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이에 따라 보상 값을 산출할 수 있다.
또한, 상기 정책 모델 학습부는, 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 상기 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출할 수 있다.
또한, 상기 정책 모델 학습부는, 상기 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 상기 마취상태 정보의 변화에서의 약물 주입 속도에 따라 상기 정책 모델을 생성할 수 있다.
본 발명의 다른 일측면은, 강화학습을 이용한 약물 주입 조절 장치를 이용하는 약물 주입 조절 방법에 있어서, 환자의 마취상태 정보를 산출하는 단계; 상기 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정하고, 상기 마취상태 정보가 상기 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 상기 마취상태 정보의 변화에 따른 보상 값을 학습하여, 정책 모델을 생성하는 단계; 상기 약물 주입 속도의 변화에 따라 변하는 상기 마취상태 정보를 학습하여, 예측 모델을 생성하는 단계; 상기 정책 모델에 기초하여, 상기 마취상태 정보로부터 상기 약물 주입 속도를 설정하는 단계; 및 상기 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는 단계를 포함할 수 있다.
또한, 상기 마취상태 정보를 산출하는 단계는, 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출하고, 상기 효과 부위 농도와 상기 혈장 농도에 따라 환자의 상태를 나타내도록 상기 마취상태 정보를 산출할 수 있다.
또한, 상기 약물 주입 속도를 설정하는 단계는, 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 상기 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도를 제어할 수 있다.
또한, 상기 정책 모델을 생성하는 단계는, 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도가 제어되어, 상기 시간 간격이 경과된 이후, 변화된 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이에 따라 보상 값을 산출할 수 있다.
또한, 상기 정책 모델을 생성하는 단계는, 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 상기 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출할 수 있다.
또한, 상기 정책 모델을 생성하는 단계는, 상기 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 상기 마취상태 정보의 변화에서의 약물 주입 속도에 따라 상기 정책 모델을 생성할 수 있다.
상술한 본 발명의 일측면에 따르면, 강화학습을 이용한 약물 주입 조절 장치 및 방법을 제공함으로써, 강화학습을 이용하여 수술 중에 환자의 마취 상태를 안정적으로 유지할 수 있다.
도1은 본 발명의 일 실시예에 따른 약물 주입 조절 장치의 개략도이다.
도2는 본 발명의 일 실시예에 따른 약물 주입 조절 장치의 제어블록도이다.
도3은 도2의 정책 모델 학습부에서 정책 모델을 생성하는 과정을 나타낸 블록도이다.
도4는 도2의 예측 모델 학습부에서 예측 모델을 생성하는 과정을 나타낸 블록도이다.
도5는 본 발명의 일 실시예에 따른 약물 주입 조절 방법의 순서도이다.
도6은 도5의 정책 모델을 생성하는 단계의 세부 순서도이다.
후술하는 본 발명에 대한 상세한 설명은, 본 발명이 실시될 수 있는 특정 실시예를 예시로서 도시하는 첨부 도면을 참조한다. 이들 실시예는 당업자가 본 발명을 실시할 수 있기에 충분하도록 상세히 설명된다. 본 발명의 다양한 실시예는 서로 다르지만 상호 배타적일 필요는 없음이 이해되어야 한다. 예를 들어, 여기에 기재되어 있는 특정 형상, 구조 및 특성은 일 실시예와 관련하여 본 발명의 정신 및 범위를 벗어나지 않으면서 다른 실시예로 구현될 수 있다. 또한, 각각의 개시된 실시예 내의 개별 구성요소의 위치 또는 배치는 본 발명의 정신 및 범위를 벗어나지 않으면서 변경될 수 있음이 이해되어야 한다. 따라서, 후술하는 상세한 설명은 한정적인 의미로서 취하려는 것이 아니며, 본 발명의 범위는, 적절하게 설명된다면, 그 청구항들이 주장하는 것과 균등한 모든 범위와 더불어 첨부된 청구항에 의해서만 한정된다. 도면에서 유사한 참조부호는 여러 측면에 걸쳐서 동일하거나 유사한 기능을 지칭한다.
이하, 도면들을 참조하여 본 발명의 바람직한 실시예들을 보다 상세하게 설명하기로 한다.
도1은 본 발명의 일 실시예에 따른 약물 주입 조절 장치의 개략도이다.
약물 주입 조절 장치(100)는 수술 중인 환자의 마취 상태를 일정하게 유지하도록 환자에게 주입되는 약물의 주입 속도를 조절하도록 마련될 수 있다.
이때, 환자의 마취 상태는 BIS(Bispectral Index)를 이용하여 나타낼 수 있으며, 일반적으로, 외과 수술 중에 마취되는 환자의 BIS는 50을 유지하도록 권장될 수 있다.
이를 위해, 약물 주입 조절 장치(100)는 환자에게 주입되는 레미펜타닐(Remifentanil)의 주입 속도와 프로포폴(Propofol)의 주입 속도를 각각 조절하여 환자의 BIS를 일정하게 유지하도록 마련된다.
약물 주입 조절 장치(100)는 환자의 마취상태 정보를 산출할 수 있다. 여기에서, 마취상태 정보는 환자의 상태로부터 산출되거나, 또는 측정되는 BIS를 의미할 수 있다.
이때, 약물 주입 조절 장치(100)는 환자의 마취 상태를 측정할 수 있도록 마련되는 마취 심도 모니터링 장치 또는 마취 심도 감시 장치 등을 이용하여 환자의 마취 심도를 측정할 수도 있으며, 이러한 경우에, 환자의 마취상태 정보는 마취 심도 모니터링 장치 또는 마취 심도 감시 장치 등에 의해 측정된 환자의 마취 심도를 의미할 수 있다.
이를 위해, 약물 주입 조절 장치(100)는 환자의 신장과 체중에 따라 환자의 제지방량을 산출할 수 있으며, 약물 주입 조절 장치(100)는 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출할 수 있다.
아래의 수학식 1 내지 수학식 5는 환자의 효과 부위 농도와 혈장 농도를 산출하는데 이용되는 수식으로 이해할 수 있다.
[수학식 1]
여기에서, lbm_male은 남성의 제지방량을 의미할 수 있고, lbm_female은 여성의 제지방량을 의미할 수 있다. 또한, Weight는 환자의 체중을 의미할 수 있고, Height는 환자의 신장을 의미할 수 있다.
[수학식 2]
여기에서, PPF는 프로포폴을 의미할 수 있으며, h_1 내지 h_17은 프로포폴에 대해 정의된 약물 모델인 Schinider 모델에 의한 변수 값으로 이해할 수 있으며, lbm은 환자의 제지방량을 의미할 수 있고, age는 환자의 나이를 의미할 수 있다. 또한, Vol_1^PPF 내지 Vol_3^PPF는 프로포폴의 각 성분에 대한 약물의 양으로 이해할 수 있고, C_l1^PPF 내지 C_l3^PPF는 프로포폴의 각 성분에 대한 체내의 약물이 제거되는 속도를 나타내는 것으로 이해할 수 있으며, k_ij^PPF는 프로포폴의 각 성분에 대한 체내에서의 약물의 변화 또는 변형을 나타내는 변수인 것으로 이해할 수 있다.
[표 1]
표 1은 Schinider 모델에 따라 임의의 환자 상태에서의 프로포폴의 각 성분에 대한 변수를 나타내는 표이다.
이때, Schinider 모델은 'The influence of method of administration and covariates on the pharmacokinetics of propofol in adult volunteers(T. Schnider, C. Minto, P. Gambus, C. Andresen, D. Goodale, S. Shafer, and E. Youngs, Anesthesiology, vol. 88, no. 5, pp. 1170-1182, 1998.)'에 공지된 바, 자세한 설명은 생략하도록 한다.
[수학식 3]
여기에서, RFTN은 레미펜타닐을 의미할 수 있으며, f_1 내지 f_17은 레미펜타닐에 대해 정의된 약물 모델인 Minto 모델에 의한 변수 값으로 이해할 수 있으며, lbm은 환자의 제지방량을 의미할 수 있고, age는 환자의 나이를 의미할 수 있다. 또한, Vol_1^RFTN 내지 Vol_3^RFTN는 레미펜타닐의 각 성분에 대한 약물의 양으로 이해할 수 있고, C_l1^RFTN 내지 C_l3^RFTN는 레미펜타닐의 각 성분에 대한 체내의 약물이 제거되는 속도를 나타내는 것으로 이해할 수 있으며, k_ij^RFTN는 레미펜타닐의 각 성분에 대한 체내에서의 약물의 변화 또는 변형을 나타내는 변수인 것으로 이해할 수 있다.
[표 2]
표 2는 Minto 모델에 따라, 임의의 환자 상태에서의 레미펜타닐의 각 성분에 대한 변수를 나타내는 표이다.
이때, Minto 모델은 'Influence of Age and Gender on the Pharmacokinetics and Pharmacodynamics of Remifentanil. I. Model development (C. Minto, T. Schnider, T. Egan, E. Youngs, H. Lemmens, P. Gambus, V. Billard, J. Hoke, K. Moore, D. Hermann et al., Anesthesiology, vol. 86, no. 1, pp. 10-23, 1997.)'에 공지된 바, 자세한 설명은 생략하도록 한다.
[수학식 4]
여기에서, xi(t)는 시간에 따른 약물의 각 성분의 양을 나타내는 변수로 이해할 수 있으며, k_ij는 약물의 각 성분에 대한 체내에서의 약물의 변화 또는 변형을 나타내는 변수인 것으로 이해할 수 있다.
[수학식 5]
여기에서, x_e(t)는 시간에 따른 약물의 효과 부위에서의 약물의 양을 나타내는 변수로 이해할 수 있으며, C_p(t)는 약물의 혈장 농도를 의미할 수 있고, C_e(t)는 약물의 효과 부위에서의 농도를 나타내는 효과 부위 농도를 의미할 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 수학식 1 내지 수학식 5에 따라, 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출할 수 있다.
이때, 약물 주입 조절 장치(100)는 프로포폴 또는 레미펜타닐에 대한 각각의 효과 부위 농도와 혈장 농도를 산출할 수 있으며, 이를 위해, 약물 주입 조절 장치(100)는 환자의 나이, 성별, 체중 및 신장 등의 정보를 입력 받을 수도 있다.
한편, 약물 주입 조절 장치(100)는 산출된 효과 부위 농도와 혈장 농도에 따라 환자의 상태를 나타내도록 마취상태 정보를 산출할 수 있다.
아래의 수학식 6은 효과 부위 농도와 혈장 농도로부터 마취상태 정보를 산출하는 수식으로 이해할 수 있다.
[수학식 6]
여기에서, BIS는 마취상태 정보로써 이용되는 환자의 BIS 수치를 의미할 수 있다. C_e^PPF는 프로포폴에 대해 산출된 효과 부위 농도를 의미할 수 있고, C_e^RFTN은 레미펜타닐에 대해 산출된 효과 부위 농도를 의미할 수 있으며, Epsilon은 BIS에 대한 가우시안 잡음(Gaussian noise)를 의미할 수 있다.
이와 관련하여, 약물 주입 조절 장치(100)는 사전에 설정되는 시간 간격에 따라 환자의 마취상태 정보를 산출할 수 있다.
약물 주입 조절 장치(100)는 환자의 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정할 수 있고, 약물 주입 조절 장치(100)는 환자의 상태에 따라 산출된 마취상태 정보가 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 마취상태 정보의 변화에 따라 보상 값을 산출할 수 있으며, 약물 주입 조절 장치(100)는 산출된 보상 값을 학습하여, 정책 모델을 생성할 수 있다.
여기에서, 목표 마취상태 정보는 환자가 안정적인 마취 상태를 유지할 수 있도록 마련되는 마취상태 정보의 수치를 의미할 수 있으며, 예를 들어, 목표 마취상태 정보는 0과 100 사이의 값으로 나타나는 BIS 중 50으로 설정될 수 있다.
또한, 약물 주입 속도는 환자에게 주입되는 프로포폴의 주입 속도와 환자에게 주입되는 레미펜타닐의 주입 속도를 포함할 수 있으며, 이때, 프로포폴의 주입 속도와 레미펜타닐의 주입 속도는 서로 다른 속도로 환자에게 주입되도록 설정될 수 있다.
한편, 아래의 수학식 7은 마취상태 정보의 변화에 따라 산출되는 보상 값을 나타내는 수식으로 이해할 수 있다.
[수학식 7]
여기에서, R은 환자의 상태와 약물의 주입 속도에 의한 마취상태 정보의 변화에 따라 산출되는 보상 값을 의미할 수 있고, s_t는 환자의 마취상태 정보와 약물의 주입 속도를 나타내도록 마련되는 상태 변수를 의미할 수 있으며, a_t는 약물 주입 조절 장치(100)에 의해 제어되는 약물의 주입 속도를 나타내도록 마련되는 동작 변수를 의미할 수 있다.
또한, g_t는 환자가 안정적인 마취 상태를 유지할 수 있도록 마련되는 목표 마취상태 정보를 의미할 수 있고, 마취상태 정보는 환자의 상태로부터 산출되는 마취상태 정보를 의미할 수 있으며, BIS_error^hist는 사전에 설정되는 시간 간격에 따라 변화된 마취상태 정보의 변화량을 의미할 수 있다. 또한, Alpha는 사전에 설정되는 최대 보상 값을 의미할 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 환자의 상태로부터 산출된 마취상태 정보와 목표 마취상태 정보의 차이의 절대 값에 사전에 설정되는 최대 보상 값을 합한 결과 값의 역수로 나타나는 보상 값을 산출할 수 있다.
이와 관련하여, 상태 변수는 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도를 포함할 수 있다.
또한, 동작 변수는 사전에 설정되는 시간 간격 동안 환자에게 주입되는 약물의 주입 속도를 나타내도록 마련될 수 있다.
예를 들어, 약물 주입 조절 장치(100)는 사전에 설정되는 시간 간격을 10초로 설정할 수 있으며, 약물 주입 조절 장치(100)는 10초 동안 환자에게 주입되는 프로포폴의 최대량을 27.8mg으로 설정할 수 있고, 약물 주입 조절 장치(100)는 10초 동안 환자에게 주입되는 레미펜타닐의 최대량을 27.8ug로 설정할 수 있다. 이러한 경우에, 프로포폴에 대한 동작 변수는 10초당 0과 27.8mg 사이의 프로포폴의 양이 주입되는 주입 속도로 설정될 수 있고, 레미펜타닐에 대한 동작 변수는 변수는 10초당 0과 27.8ug 사이의 레미펜타닐의 양이 주입되는 주입 속도로 설정될 수 있다.
이에 따라, 보상 값은 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 나타나는 상태에서, 환자에게 주입되는 레미펜타닐의 주입 속도와 환자에게 주입되는 프로포폴의 주입 속도를 제어한 결과에 대한 보상을 나타내도록 마련되는 값일 수 있다.
이때, 약물의 주입 속도를 제어한 결과는 레미펜타닐의 주입 속도와 프로포폴의 주입 속도를 제어한 후, 사전에 설정된 시간 간격이 경과된 이후의 상태 변수에 따른 상태를 의미하는 것으로 이해할 수 있다.
다시 말해서, 약물 주입 조절 장치(100)는 레미펜타닐의 주입 속도와 프로포폴의 주입 속도가 제어되어, 사전에 설정된 시간 간격이 경과된 이후에, 변화된 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이에 따라 보상 값을 산출할 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출할 수 있다.
아래의 수학식 8은 보상 값으로부터 기대 값을 산출하도록 마련되는 수식이다.
[수학식 8]
여기에서, J(Phi_Theta)는 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 경우에 산출되는 보상 값을 나타내도록 마련되는 기대 값을 의미할 수 있고, P(Tau|Phi_Theta)는 임의의 상태 변수에서 임의의 동작 변수가 발생하는 확률을 나타내도록 마련되는 확률 값을 의미할 수 있으며, R(Tau)는 보상 값을 의미할 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 경우에, 각각의 시간 간격에서의 확률 값과 보상 값의 곱의 총합으로 나타나는 기대 값을 산출할 수 있다.
이와 관련하여, 아래의 수학식 9는 확률 값을 산출하는 수식이다.
[수학식 9]
여기에서, P(Tau|Phi_Theta)는 임의의 상태 변수에서 임의의 동작 변수가 발생하는 확률을 나타내도록 마련되는 확률 값을 의미할 수 있고, Rho(S_0)는 환자로부터 최초로 측정되는 마취상태 정보와 목표 마취상태 정보의 차이, 환자에게 최초로 주입되는 레미펜타닐의 주입 속도 및 환자에게 최초로 주입되는 프로포폴의 주입 속도를 포함하도록 마련되는 초기 상태 변수를 의미할 수 있다. 또한, P(S_t+1|s_t, a_t)는 임의의 상태 변수와 임의의 동작 변수에 의해 사전에 설정된 시간 간격이 경과된 이후에 상태 변수가 임의의 상태 변수로 변화하는 확률 값을 의미할 수 있고, Phi_Theta(a_t|s_t, g_t)는 임의의 상태 변수와 목표 마취상태 정보에 따라 특정한 동작 변수를 선택하도록 마련되는 정책 모델을 의미할 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 서로 다른 시간 상에서의 상태 변수가 변화하는 확률 값과 정책 모델의 곱을 각각 곱한 값에 초기 상태 변수를 곱한 값으로 나타나는 확률 값을 산출할 수 있다.
이때, 약물 주입 조절 장치(100)는 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 마취상태 정보의 변화에서의 약물 주입 속도에 따라 정책 모델을 생성할 수 있다.
다시 말해서, 약물 주입 조절 장치(100)는 임의의 약물 주입 속도에 의한 마취상태 정보의 변화에 따른 보상 값을 학습하여, 기대 값이 최대 값으로 산출되는 보상 값을 선택할 수 있으며, 약물 주입 조절 장치(100)는 선택된 보상 값에 매칭되는 마취상태 정보의 변화에 대한 약물 주입 속도가 동작 변수로 나타나도록 정책 모델을 생성할 수 있다.
여기에서, 약물 주입 조절 장치(100)는 임의로 설정된 약물의 주입 속도에 대해 사전에 설정된 시간 간격에 따른 마취상태 정보의 변화를 포함하도록 사전에 마련되는 복수개의 데이터 세트로부터 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
이와 관련하여, 약물 주입 조절 장치(100)는 임의의 상태에서 임의의 동작을 수행하여 변화하는 상태에 따른 보상을 학습하여, 정책을 생성하고, 생성된 정책에 따라 임의의 상태에서 수행하는 동작을 선택하는 강화 학습(Reinforcement Learning) 기법을 이용하여 정책 모델을 생성할 수 있다.
약물 주입 조절 장치(100)는 약물 주입 속도에 따른 마취상태 정보의 변화를 학습하여, 예측 모델을 생성할 수 있다.
이를 위해, 약물 주입 조절 장치(100)는 환자의 나이, 성별, 체중 및 신장 등의 정보에 대해, 사전에 설정된 시간 간격에 따른 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도의 변화를 학습할 수 있다.
이때, 약물 주입 조절 장치(100)는 환자의 나이, 성별, 체중 및 신장 등의 정보와 사전에 설정된 시간 간격에 따른 상태 변수의 변화를 각각 DNN(Deep Neural Network) 또는 LSTM(Long Short-Term Memory models)에 입력하도록 마련될 수 있다.
여기에서, DNN은 입력층과 출력층 사이에 다수개의 은닉층을 마련하여, 입력층에 입력되는 정보로부터 특징 정보를 출력하는 기법이며, LSTM은 시간적 순서에 따라 정보를 입력 받으며, 이전에 입력된 정보와 현재 입력되는 정보에 기초하여 출력되는 정보의 양을 조절하는 기법이다.
이에 따라, 약물 주입 조절 장치(100)는 DNN 또는 LSTM으로부터 출력되는 정보를 하나의 DNN에 입력하여 정보를 융합할 수 있으며, 이때, 약물 주입 조절 장치(100)는 정보의 융합에 이용된 DNN으로부터 출력되는 정보를 예상 마취상태 정보를 출력하도록 마련되는 다른 DNN에 입력하도록 마련될 수 있다.
이에 따라, 약물 주입 조절 장치(100)는 환자의 나이, 성별, 체중 및 신장 등의 정보와, 사전에 설정된 시간 간격에 따른 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도의 변화를 이용하여 예상 마취상태 정보를 생성할 수 있다.
약물 주입 조절 장치(100)는 정책 모델에 기초하여, 마취상태 정보로부터 약물 주입 속도를 설정할 수 있다.
이때, 약물 주입 조절 장치(100)는 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 레미펜타닐의 주입 속도와 프로포폴의 주입 속도를 제어할 수 있다.
다시 말해서, 약물 주입 조절 장치(100)는 현재의 상태 변수에 따라 정책 모델에 의해 설정된 동작 변수를 선택하여, 동작 변수에 따라 환자에게 주입되는 약물의 주입 속도를 제어할 수 있다.
약물 주입 조절 장치(100)는 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측할 수 있다.
이때, 약물 주입 조절 장치(100)는 예측된 예상 마취상태 정보를 사용자가 확인할 수 있도록 그래프, 또는 수치 등의 형상으로 출력할 수 있으며, 이를 위해, 약물 주입 조절 장치(100)는 별도의 디스플레이 장치를 포함할 수 있다.
한편, 약물 주입 조절 장치(100)는 예측된 예상 마취상태 정보를 이용하여, 환자의 상태로부터 산출되는 마취상태 정보에 따른 환자의 상태로부터 예상 마취상태 정보에 따른 환자의 상태에 따라 보상 값을 산출할 수도 있으며, 이러한 경우에, 약물 주입 조절 장치(100)는 예상 마취상태 정보에 의한 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
또한, 약물 주입 조절 장치(100)는 환자의 마취 상태를 측정할 수 있도록 마련되는 마취 심도 모니터링 장치 또는 마취 심도 감시 장치 등을 이용하여 환자의 마취 심도를 측정할 수도 있으며, 이러한 경우에, 약물 주입 조절 장치(100)는 환자로부터 측정되는 마취 심도와 환자의 마취 상태가 안정적으로 유지될 수 있도록 설정되는 목표 마취 심도에 따라 정책 모델을 생성할 수도 있다.
여기에서, 마취 심도 모니터링 장치 또는 마취 심도 감시 장치 등의 환자의 마취 상태를 측정하는 장치는 공지된 기술이 이용된 어느 기구, 장치 및 방법에 따른 장치일 수 있다.
도2는 본 발명의 일 실시예에 따른 약물 주입 조절 장치의 제어블록도이다.
약물 주입 조절 장치(100)는 마취상태 정보 산출부(110), 정책 모델 학습부(120), 예측 모델 학습부(130), 제어부(140) 및 예측부(150)를 포함할 수 있다.
마취상태 정보 산출부(110)는 환자의 마취상태 정보를 산출할 수 있으며, 이와 관련하여, 마취상태 정보 산출부(110)는 환자의 나이, 성별, 체중 및 신장 등의 정보를 입력 받을 수도 있다.
이를 위해, 마취상태 정보 산출부(110)는 환자의 신장과 체중에 따라 환자의 제지방량을 산출할 수 있으며, 마취상태 정보 산출부(110)는 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출할 수 있다. 이때, 마취상태 정보 산출부(110)는 프로포폴 또는 레미펜타닐에 대한 각각의 효과 부위 농도와 혈장 농도를 산출할 수 있다.
이에 따라, 마취상태 정보 산출부(110)는 산출된 효과 부위 농도와 혈장 농도에 따라 환자의 상태를 나타내도록 마취상태 정보를 산출할 수 있으며, 이때, 마취상태 정보 산출부(110)는 사전에 설정되는 시간 간격에 따라 환자의 마취상태 정보를 산출할 수 있다.
정책 모델 학습부(120)는 환자의 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정할 수 있고, 정책 모델 학습부(120)는 환자의 상태에 따라 산출된 마취상태 정보가 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 마취상태 정보의 변화에 따라 보상 값을 산출할 수 있으며, 정책 모델 학습부(120)는 산출된 보상 값을 학습하여, 정책 모델을 생성할 수 있다.
이때, 정책 모델 학습부(120)는 환자의 상태로부터 산출된 마취상태 정보와 목표 마취상태 정보의 차이의 절대 값에 사전에 설정되는 최대 보상 값을 합한 결과 값의 역수로 나타나는 보상 값을 산출할 수 있다.
다시 말해서, 정책 모델 학습부(120)는 레미펜타닐의 주입 속도와 프로포폴의 주입 속도가 제어되어, 사전에 설정된 시간 간격이 경과된 이후에, 변화된 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이에 따라 보상 값을 산출할 수 있다.
또한, 정책 모델 학습부(120)는 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출할 수 있다.
여기에서, 정책 모델 학습부(120)는 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 경우에, 각각의 시간 간격에서의 확률 값과 보상 값의 곱의 총합으로 나타나는 기대 값을 산출할 수 있다.
이와 관련하여, 정책 모델 학습부(120)는 서로 다른 시간 상에서의 상태 변수가 변화하는 확률 값과 정책 모델의 곱을 각각 곱한 값에 초기 상태 변수를 곱한 값으로 나타나는 확률 값을 산출할 수 있다.
이때, 정책 모델 학습부(120)는 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 마취상태 정보의 변화에서의 약물 주입 속도에 따라 정책 모델을 생성할 수 있다.
다시 말해서, 정책 모델 학습부(120)는 임의의 약물 주입 속도에 의한 마취상태 정보의 변화에 따른 보상 값을 학습하여, 기대 값이 최대 값으로 산출되는 보상 값을 선택할 수 있으며, 정책 모델 학습부(120)는 선택된 보상 값에 매칭되는 마취상태 정보의 변화에 대한 약물 주입 속도가 동작 변수로 나타나도록 정책 모델을 생성할 수 있다.
여기에서, 정책 모델 학습부(120)는 임의로 설정된 약물의 주입 속도에 대해 사전에 설정된 시간 간격에 따른 마취상태 정보의 변화를 포함하도록 사전에 마련되는 복수개의 데이터 세트로부터 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
이와 관련하여, 정책 모델 학습부(120)는 임의의 상태에서 임의의 동작을 수행하여 변화하는 상태에 따른 보상을 학습하여, 정책을 생성하고, 생성된 정책에 따라 임의의 상태에서 수행하는 동작을 선택하는 강화 학습(Reinforcement Learning) 기법을 이용하여 정책 모델을 생성할 수 있다.
한편, 정책 모델 학습부(120)는 예측된 예상 마취상태 정보를 이용하여, 환자의 상태로부터 산출되는 마취상태 정보에 따른 환자의 상태로부터 예상 마취상태 정보에 따른 환자의 상태에 따라 보상 값을 산출할 수도 있으며, 이러한 경우에, 정책 모델 학습부(120)는 예상 마취상태 정보에 의한 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
예측 모델 학습부(130)는 약물 주입 속도에 따른 마취상태 정보의 변화를 학습하여, 예측 모델을 생성할 수 있다.
이를 위해, 예측 모델 학습부(130)는 환자의 나이, 성별, 체중 및 신장 등의 정보에 대해, 사전에 설정된 시간 간격에 따른 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도의 변화를 학습할 수 있다.
제어부(140)는 정책 모델에 기초하여, 마취상태 정보로부터 약물 주입 속도를 설정할 수 있다.
이때, 제어부(140)는 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 레미펜타닐의 주입 속도와 프로포폴의 주입 속도를 제어할 수 있다.
다시 말해서, 제어부(140)는 현재의 상태 변수에 따라 정책 모델에 의해 설정된 동작 변수를 선택하여, 동작 변수에 따라 환자에게 주입되는 약물의 주입 속도를 제어할 수 있다.
예측부(150)는 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측할 수 있다.
이때, 예측부(150)는 예측된 예상 마취상태 정보를 사용자가 확인할 수 있도록 그래프, 또는 수치 등의 형상으로 출력할 수 있으며, 이를 위해, 약물 주입 조절 장치(100)는 별도의 디스플레이 장치를 포함할 수 있다.
도3은 도2의 정책 모델 학습부에서 정책 모델을 생성하는 과정을 나타낸 블록도이다.
도3을 참조하면, 마취상태 정보 산출부(110)는 환자의 마취상태 정보를 산출할 수 있으며, 이와 관련하여, 마취상태 정보 산출부(110)는 환자의 나이, 성별, 체중 및 신장 등의 정보를 입력 받을 수도 있다.
이를 위해, 마취상태 정보 산출부(110)는 환자의 신장과 체중에 따라 환자의 제지방량을 산출할 수 있으며, 마취상태 정보 산출부(110)는 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출할 수 있다. 이때, 마취상태 정보 산출부(110)는 프로포폴 또는 레미펜타닐에 대한 각각의 효과 부위 농도와 혈장 농도를 산출할 수 있다.
이에 따라, 마취상태 정보 산출부(110)는 산출된 효과 부위 농도와 혈장 농도에 따라 환자의 상태를 나타내도록 마취상태 정보를 산출할 수 있으며, 이때, 마취상태 정보 산출부(110)는 사전에 설정되는 시간 간격에 따라 환자의 마취상태 정보를 산출할 수 있다.
이때, 정책 모델 학습부(120)는 환자의 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정할 수 있으며, 정책 모델 학습부(120)는 환자의 마취상태 정보가 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 마취상태 정보의 변화에 따른 보상 값을 학습하여, 정책 모델을 생성할 수 있다.
이때, 정책 모델 학습부(120)는 환자의 상태로부터 산출된 마취상태 정보와 목표 마취상태 정보의 차이의 절대 값에 사전에 설정되는 최대 보상 값을 합한 결과 값의 역수로 나타나는 보상 값을 산출할 수 있다.
다시 말해서, 정책 모델 학습부(120)는 레미펜타닐의 주입 속도와 프로포폴의 주입 속도가 제어되어, 사전에 설정된 시간 간격이 경과된 이후에, 변화된 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이에 따라 보상 값을 산출할 수 있다.
또한, 정책 모델 학습부(120)는 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출할 수 있다.
여기에서, 정책 모델 학습부(120)는 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 경우에, 각각의 시간 간격에서의 확률 값과 보상 값의 곱의 총합으로 나타나는 기대 값을 산출할 수 있다.
이와 관련하여, 정책 모델 학습부(120)는 서로 다른 시간 상에서의 상태 변수가 변화하는 확률 값과 정책 모델의 곱을 각각 곱한 값에 초기 상태 변수를 곱한 값으로 나타나는 확률 값을 산출할 수 있다.
이때, 정책 모델 학습부(120)는 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 마취상태 정보의 변화에서의 약물 주입 속도에 따라 정책 모델을 생성할 수 있다.
다시 말해서, 정책 모델 학습부(120)는 임의의 약물 주입 속도에 의한 마취상태 정보의 변화에 따른 보상 값을 학습하여, 기대 값이 최대 값으로 산출되는 보상 값을 선택할 수 있으며, 정책 모델 학습부(120)는 선택된 보상 값에 매칭되는 마취상태 정보의 변화에 대한 약물 주입 속도가 동작 변수로 나타나도록 정책 모델을 생성할 수 있다.
여기에서, 정책 모델 학습부(120)는 임의로 설정된 약물의 주입 속도에 대해 사전에 설정된 시간 간격에 따른 마취상태 정보의 변화를 포함하도록 사전에 마련되는 복수개의 데이터 세트로부터 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
이에 따라, 제어부(140)는 정책 모델에 기초하여, 마취상태 정보로부터 약물 주입 속도를 설정할 수 있다.
이때, 제어부(140)는 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 레미펜타닐의 주입 속도와 프로포폴의 주입 속도를 제어할 수 있다.
다시 말해서, 제어부(140)는 현재의 상태 변수에 따라 정책 모델에 의해 설정된 동작 변수를 선택하여, 동작 변수에 따라 환자에게 주입되는 약물의 주입 속도를 제어할 수 있다.
도4는 도2의 예측 모델 학습부에서 예측 모델을 생성하는 과정을 나타낸 블록도이다.
도4를 참조하면, 마취상태 정보 산출부(110)는 환자의 마취상태 정보를 산출할 수 있으며, 이와 관련하여, 마취상태 정보 산출부(110)는 환자의 나이, 성별, 체중 및 신장 등의 정보를 입력 받을 수도 있다.
이를 위해, 마취상태 정보 산출부(110)는 환자의 신장과 체중에 따라 환자의 제지방량을 산출할 수 있으며, 마취상태 정보 산출부(110)는 환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출할 수 있다. 이때, 마취상태 정보 산출부(110)는 프로포폴 또는 레미펜타닐에 대한 각각의 효과 부위 농도와 혈장 농도를 산출할 수 있다.
이에 따라, 마취상태 정보 산출부(110)는 산출된 효과 부위 농도와 혈장 농도에 따라 환자의 상태를 나타내도록 마취상태 정보를 산출할 수 있으며, 이때, 마취상태 정보 산출부(110)는 사전에 설정되는 시간 간격에 따라 환자의 마취상태 정보를 산출할 수 있다.
이때, 예측 모델 학습부(130)는 약물 주입 속도의 변화에 따라 변하는 마취상태 정보를 학습하여, 예측 모델을 생성할 수 있다.
이를 위해, 예측 모델 학습부(130)는 환자의 나이, 성별, 체중 및 신장 등의 정보에 대해, 사전에 설정된 시간 간격에 따른 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 사전에 설정되는 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도의 변화를 학습할 수 있다.
이에 따라, 예측부(150)는 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측할 수 있다.
이때, 예측부(150)는 예측된 예상 마취상태 정보를 사용자가 확인할 수 있도록 그래프, 또는 수치 등의 형상으로 출력할 수 있으며, 이를 위해, 약물 주입 조절 장치(100)는 별도의 디스플레이 장치를 포함할 수 있다.
한편, 정책 모델 학습부(120)는 예측된 예상 마취상태 정보를 이용하여, 환자의 상태로부터 산출되는 마취상태 정보에 따른 환자의 상태로부터 예상 마취상태 정보에 따른 환자의 상태에 따라 보상 값을 산출할 수도 있으며, 이러한 경우에, 정책 모델 학습부(120)는 예상 마취상태 정보에 의한 보상 값을 학습하여 정책 모델을 생성할 수도 있다.
도5는 본 발명의 일 실시예에 따른 약물 주입 조절 방법의 순서도이다.
본 발명의 일 실시예에 따른 약물 주입 조절 방법은 도 1에 도시된 약물 주입 조절 장치(100)와 실질적으로 동일한 구성 상에서 진행되므로, 도 1의 약물 주입 조절 장치(100)와 동일한 구성요소에 대해 동일한 도면 부호를 부여하고, 반복되는 설명은 생략하기로 한다.
약물 주입 조절 방법은 마취상태 정보를 산출하는 단계(600), 정책 모델을 생성하는 단계(610), 예측 모델을 생성하는 단계(620), 약물 주입 속도를 설정하는 단계(630) 및 예상 마취상태 정보를 예측하는 단계(640)를 포함할 수 있다.
마취상태 정보를 산출하는 단계(600)는 마취상태 정보 산출부(110)가 환자의 마취상태 정보를 산출하는 단계일 수 있다.
정책 모델을 생성하는 단계(610)는 정책 모델 학습부(120)가 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정하고, 산출된 마취상태 정보가 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 마취상태 정보의 변화에 따라 보상 값을 산출하며, 보상 값을 학습하여, 정책 모델을 생성하는 단계일 수 있다.
예측 모델을 생성하는 단계(620)는 예측 모델 학습부(130)가 약물 주입 속도에 따른 마취상태 정보의 변화를 학습하여, 예측 모델을 생성하는 단계일 수 있다.
약물 주입 속도를 설정하는 단계(630)는 제어부(140)가 정책 모델에 기초하여, 마취상태 정보로부터 약물 주입 속도를 설정하는 단계일 수 있다.
예상 마취상태 정보를 예측하는 단계(640)는 예측부(150)가 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는 단계일 수 있다.
도6은 도5의 정책 모델을 생성하는 단계의 세부 순서도이다.
정책 모델을 생성하는 단계(610)는 마취상태 정보의 변화를 산출하는 단계(611), 보상 값을 산출하는 단계(612), 기대 값을 산출하는 단계(613) 및 보상 값을 선택하는 단계(614)를 포함할 수 있다.
마취상태 정보의 변화를 산출하는 단계(611)는 정책 모델 학습부(120)가 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이와 관련하여, 레미펜타닐의 주입 속도와 프로포폴의 주입 속도가 제어되어, 사전에 설정된 시간 간격이 경과된 이후에, 변화된 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이로 나타나는 마취상태 정보의 변화를 산출하는 단계일 수 있다.
보상 값을 산출하는 단계(612)는 정책 모델 학습부(120)가 환자의 상태로부터 산출된 마취상태 정보와 목표 마취상태 정보의 차이의 절대 값에 사전에 설정되는 최대 보상 값을 합한 결과 값의 역수로 나타나는 보상 값을 산출하는 단계일 수 있다.
다시 말해서, 보상 값을 산출하는 단계(612)는 정책 모델 학습부(120)가 레미펜타닐의 주입 속도와 프로포폴의 주입 속도가 제어되어, 사전에 설정된 시간 간격이 경과된 이후에, 변화된 환자의 상태로부터 산출되는 마취상태 정보와 목표 마취상태 정보의 차이에 따라 보상 값을 산출하는 단계일 수 있다.
기대 값을 산출하는 단계(613)는 정책 모델 학습부(120)가 서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출하는 단계일 수 있다.
보상 값을 선택하는 단계(614)는 정책 모델 학습부(120)가 임의의 약물 주입 속도에 의한 마취상태 정보의 변화에 따른 보상 값을 학습하여, 기대 값이 최대 값으로 산출되는 보상 값을 선택하는 단계일 수 있으며, 이에 따라, 정책 모델 학습부(120)는 선택된 보상 값에 매칭되는 마취상태 정보의 변화에 대한 약물 주입 속도가 동작 변수로 나타나도록 정책 모델을 생성할 수 있다.
이상에서는 실시예들을 참조하여 설명하였지만, 해당 기술 분야의 숙련된 당업자는 하기의 특허 청구범위에 기재된 본 발명의 사상 및 영역으로부터 벗어나지 않는 범위 내에서 본 발명을 다양하게 수정 및 변경시킬 수 있음을 이해할 수 있을 것이다.
[부호의 설명]
100: 약물 주입 조절 장치
110: 마취상태 정보 산출부
120: 정책 모델 학습부
130: 예측 모델 학습부
140: 제어부
150: 예측부
Claims (12)
- 환자의 마취상태 정보를 산출하는 마취상태 정보 산출부;상기 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정하고, 상기 마취상태 정보가 상기 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 상기 마취상태 정보의 변화에 따라 보상 값을 산출하며, 상기 보상 값을 학습하여, 정책 모델을 생성하는 정책 모델 학습부;상기 약물 주입 속도에 따른 상기 마취상태 정보의 변화를 학습하여, 예측 모델을 생성하는 예측 모델 학습부;상기 정책 모델에 기초하여, 상기 마취상태 정보로부터 상기 약물 주입 속도를 설정하는 제어부; 및상기 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는 예측부를 포함하는, 약물 주입 조절 장치.
- 제1항에 있어서, 상기 마취상태 정보 산출부는,환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출하고, 상기 효과 부위 농도와 상기 혈장 농도에 따라 환자의 상태를 나타내도록 상기 마취상태 정보를 산출하는, 약물 주입 조절 장치.
- 제1항에 있어서, 상기 제어부는,환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 상기 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도를 제어하는, 약물 주입 조절 장치.
- 제3항에 있어서, 상기 정책 모델 학습부는,상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도가 제어되어, 상기 시간 간격이 경과된 이후, 변화된 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이에 따라 보상 값을 산출하는, 약물 주입 조절 장치.
- 제1항에 있어서, 상기 정책 모델 학습부는,서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 상기 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출하는, 약물 주입 조절 장치.
- 제5항에 있어서, 상기 정책 모델 학습부는,상기 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 상기 마취상태 정보의 변화에서의 약물 주입 속도에 따라 상기 정책 모델을 생성하는, 약물 주입 조절 장치.
- 강화학습을 이용한 약물 주입 조절 장치를 이용하는 약물 주입 조절 방법에 있어서,환자의 마취상태 정보를 산출하는 단계;상기 마취상태 정보에 대해 사전에 설정되는 목표 마취상태 정보를 설정하고, 상기 마취상태 정보가 상기 목표 마취상태 정보를 추종하도록 설정되는 약물 주입 속도에 의한 상기 마취상태 정보의 변화에 따라 보상 값을 산출하며, 상기 보상 값을 학습하여, 정책 모델을 생성하는 단계;상기 약물 주입 속도에 따른 상기 마취상태 정보의 변화를 학습하여, 예측 모델을 생성하는 단계;상기 정책 모델에 기초하여, 상기 마취상태 정보로부터 상기 약물 주입 속도를 설정하는 단계; 및상기 예측 모델에 기초하여, 설정된 약물 주입 속도와 이전에 설정된 약물 주입 속도로부터 예상 마취상태 정보를 예측하는 단계를 포함하는, 약물 주입 조절 방법.
- 제7항에 있어서, 상기 마취상태 정보를 산출하는 단계는,환자의 제지방량과 환자에게 주입되는 약물에 대해 사전에 마련되는 약물 모델에 기초하여, 효과 부위 농도와 혈장 농도를 각각 산출하고, 상기 효과 부위 농도와 상기 혈장 농도에 따라 환자의 상태를 나타내도록 상기 마취상태 정보를 산출하는, 약물 주입 조절 방법.
- 제7항에 있어서, 상기 약물 주입 속도를 설정하는 단계는,환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이, 사전에 설정되는 시간 간격 동안 환자에게 주입되는 레미펜타닐의 주입 속도 및 상기 시간 간격 동안 환자에게 주입되는 프로포폴의 주입 속도에 따라 상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도를 제어하는, 약물 주입 조절 방법.
- 제9항에 있어서, 상기 정책 모델을 생성하는 단계는,상기 레미펜타닐의 주입 속도와 상기 프로포폴의 주입 속도가 제어되어, 상기 시간 간격이 경과된 이후, 변화된 환자의 상태로부터 산출되는 상기 마취상태 정보와 상기 목표 마취상태 정보의 차이에 따라 보상 값을 산출하는, 약물 주입 조절 방법.
- 제7항에 있어서, 상기 정책 모델을 생성하는 단계는,서로 다른 마취상태 정보의 변화로부터 산출된 복수개의 보상 값에 기초하여, 임의의 마취상태 정보로부터, 상기 마취상태 정보의 변화에 대해 사전에 설정된 시간 간격이 수회 경과되는 과정에서 각각의 마취상태 정보의 변화에 매칭되는 복수개의 보상 값에 따라 기대 값을 산출하는, 약물 주입 조절 방법.
- 제11항에 있어서, 상기 정책 모델을 생성하는 단계는,상기 기대 값이 최대 값으로 산출되도록 선택된 보상 값에 매칭되는 상기 마취상태 정보의 변화에서의 약물 주입 속도에 따라 상기 정책 모델을 생성하는, 약물 주입 조절 방법.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/016,507 US20230293099A1 (en) | 2020-07-14 | 2021-06-02 | Drug injection adjusting apparatus and method using reinforcement learning |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR1020200086957A KR102234007B1 (ko) | 2020-07-14 | 2020-07-14 | 강화학습을 이용한 약물 주입 조절 장치 및 방법 |
| KR10-2020-0086957 | 2020-07-14 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022014861A1 true WO2022014861A1 (ko) | 2022-01-20 |
Family
ID=75237712
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2021/006850 Ceased WO2022014861A1 (ko) | 2020-07-14 | 2021-06-02 | 강화학습을 이용한 약물 주입 조절 장치 및 방법 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20230293099A1 (ko) |
| KR (1) | KR102234007B1 (ko) |
| WO (1) | WO2022014861A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115299881A (zh) * | 2022-08-08 | 2022-11-08 | 中科院成都信息技术股份有限公司 | 一种麻醉维持期的监测与调控系统及方法 |
Families Citing this family (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102234007B1 (ko) * | 2020-07-14 | 2021-03-31 | 고려대학교 산학협력단 | 강화학습을 이용한 약물 주입 조절 장치 및 방법 |
| KR102481888B1 (ko) * | 2020-09-01 | 2022-12-28 | 고려대학교 산학협력단 | 딥러닝을 이용한 수면유도제 투여량 예측방법 및 예측 장치 |
| FR3142280B1 (fr) * | 2022-11-23 | 2024-10-11 | Diabeloop | Method for determining an uncertainty level of deep reinforcement learning network and device implementing such method. Procédé de détermination d'un niveau d'incertitude d'un réseau d'apprentissage |
| CN117017234B (zh) * | 2023-10-09 | 2024-01-16 | 深圳市格阳医疗科技有限公司 | 多参数集成麻醉监测与分析系统 |
| CN117797359B (zh) * | 2024-01-12 | 2024-07-05 | 江苏朗沁科技有限公司 | 一种透皮给药用注射泵控制系统 |
| CN119560180B (zh) * | 2024-11-21 | 2025-10-03 | 武汉大学 | 基于mems微流控麻醉弹的药物释放系统 |
| CN119405958A (zh) * | 2024-11-21 | 2025-02-11 | 武汉大学 | 一种麻醉弹的流量调节系统及方法 |
| CN120199414B (zh) * | 2025-05-26 | 2025-11-11 | 铜川市人民医院 | 一种基于人工智能的麻醉药物剂量优化方法 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100461947B1 (ko) * | 2001-07-06 | 2004-12-17 | 주식회사 바이오넷 | 목표농도 조절 주입기 및 그 제어방법 |
| KR101246667B1 (ko) * | 2011-08-31 | 2013-03-25 | 전자부품연구원 | 통합 마취 관리 시스템 및 방법 |
| KR20130125086A (ko) * | 2012-05-08 | 2013-11-18 | 주식회사 바이오넷 | 폐루프 형태의 액체약물 목표농도주입 시스템 및 그 동작방법 |
| JP2016516514A (ja) * | 2013-04-24 | 2016-06-09 | フレゼニウス カービ ドイチュラント ゲーエムベーハー | 薬剤注入装置を制御する制御装置を操作する方法 |
| CN110349642A (zh) * | 2019-07-09 | 2019-10-18 | 泰康保险集团股份有限公司 | 智能麻醉实施方法、装置、设备及存储介质 |
| KR102234007B1 (ko) * | 2020-07-14 | 2021-03-31 | 고려대학교 산학협력단 | 강화학습을 이용한 약물 주입 조절 장치 및 방법 |
-
2020
- 2020-07-14 KR KR1020200086957A patent/KR102234007B1/ko active Active
-
2021
- 2021-06-02 US US18/016,507 patent/US20230293099A1/en active Pending
- 2021-06-02 WO PCT/KR2021/006850 patent/WO2022014861A1/ko not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100461947B1 (ko) * | 2001-07-06 | 2004-12-17 | 주식회사 바이오넷 | 목표농도 조절 주입기 및 그 제어방법 |
| KR101246667B1 (ko) * | 2011-08-31 | 2013-03-25 | 전자부품연구원 | 통합 마취 관리 시스템 및 방법 |
| KR20130125086A (ko) * | 2012-05-08 | 2013-11-18 | 주식회사 바이오넷 | 폐루프 형태의 액체약물 목표농도주입 시스템 및 그 동작방법 |
| JP2016516514A (ja) * | 2013-04-24 | 2016-06-09 | フレゼニウス カービ ドイチュラント ゲーエムベーハー | 薬剤注入装置を制御する制御装置を操作する方法 |
| CN110349642A (zh) * | 2019-07-09 | 2019-10-18 | 泰康保险集团股份有限公司 | 智能麻醉实施方法、装置、设备及存储介质 |
| KR102234007B1 (ko) * | 2020-07-14 | 2021-03-31 | 고려대학교 산학협력단 | 강화학습을 이용한 약물 주입 조절 장치 및 방법 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115299881A (zh) * | 2022-08-08 | 2022-11-08 | 中科院成都信息技术股份有限公司 | 一种麻醉维持期的监测与调控系统及方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| KR102234007B1 (ko) | 2021-03-31 |
| US20230293099A1 (en) | 2023-09-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020098013A1 (zh) | 电视节目推荐方法、终端、系统及存储介质 | |
| WO2020125251A1 (zh) | 基于联邦学习的模型参数训练方法、装置、设备及介质 | |
| WO2018076865A1 (zh) | 数据分享方法、装置、存储介质及电子设备 | |
| WO2021085742A1 (ko) | 서보 모터 제어 시스템 및 그 제어 방법 | |
| WO2019124921A1 (ko) | 미디어 컨텐츠의 재생 시점을 제어하는 방법 및 장치 | |
| WO2019039808A1 (ko) | 저혈당 예측 장치, 방법 및 프로그램과, 저혈당 예측 모델 생성 장치, 방법 및 프로그램 | |
| WO2022065763A1 (en) | Display apparatus and method for controlling thereof | |
| WO2018223520A1 (zh) | 面向儿童的学习方法、学习设备及存储介质 | |
| WO2017067375A1 (zh) | 一种视频背景设置方法及终端设备 | |
| WO2020062615A1 (zh) | 显示面板的伽马值调节方法、装置及显示设备 | |
| WO2012070910A2 (ko) | 대표값 산출 장치 및 방법. | |
| WO2017111197A1 (ko) | 학습 분석에서 빅데이터 시각화 시스템 및 방법 | |
| WO2011115315A1 (ko) | 인지 재활 훈련 프로그램을 위한 시스템 및 그를 이용한 서비스 방법. | |
| WO2021029658A1 (ko) | 처음입력문자에 기반한 한자 입력 장치 및 그 제어 방법 | |
| WO2012064161A2 (en) | Method and apparatus for generating community | |
| WO2022158847A1 (ko) | 멀티 모달 데이터를 처리하는 전자 장치 및 그 동작 방법 | |
| WO2020056952A1 (zh) | 显示面板的控制方法、显示面板及存储介质 | |
| WO2020138542A1 (ko) | 액션 로봇용 콘텐츠 판매 서비스 운영 장치 및 그의 동작 방법 | |
| WO2016106820A1 (zh) | 显示面板的参数设置方法及装置 | |
| WO2022060167A1 (ko) | 약물과민반응 예방 시스템 | |
| WO2020141745A1 (ko) | 집단 지성 기반의 상호 검증을 이용한 인공지능 기반 객체인식 신뢰도 향상 방법 | |
| WO2012036363A1 (ko) | 송전선로 선로정수 측정장치 및 그 측정방법 | |
| WO2012169675A1 (ko) | 누적 이동 평균에 기반하여 다원 탐색 트리의 노드를 분할하는 방법 및 장치 | |
| WO2025183532A1 (ko) | 예측 궤적을 디스플레이하는 디스플레이 장치 및 그 제어 방법 | |
| WO2026023821A1 (ko) | 주사기에 주사액을 충진하는 방법, 및 이를 위한 주사액 충진 장치 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21843085 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21843085 Country of ref document: EP Kind code of ref document: A1 |






