WO2018117453A1 - 대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템 - Google Patents
대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템 Download PDFInfo
- Publication number
- WO2018117453A1 WO2018117453A1 PCT/KR2017/013633 KR2017013633W WO2018117453A1 WO 2018117453 A1 WO2018117453 A1 WO 2018117453A1 KR 2017013633 W KR2017013633 W KR 2017013633W WO 2018117453 A1 WO2018117453 A1 WO 2018117453A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- log
- event
- storage system
- log file
- replay
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/30—Monitoring
- G06F11/34—Recording or statistical evaluation of computer activity, e.g. of down time, of input/output operation ; Recording or statistical evaluation of user activity, e.g. usability assessment
- G06F11/3466—Performance evaluation by tracing or monitoring
- G06F11/3476—Data logging
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/18—File system types
- G06F16/182—Distributed file systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/14—Error detection or correction of the data by redundancy in operations
- G06F11/1471—Error detection or correction of the data by redundancy in operations involving logging of persistent data for recovery
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/30—Monitoring
- G06F11/3065—Monitoring arrangements determined by the means or processing involved in reporting the monitored data
- G06F11/3072—Monitoring arrangements determined by the means or processing involved in reporting the monitored data where the reporting involves data filtering, e.g. pattern matching, time or event triggered, adaptive or policy-based reporting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/30—Monitoring
- G06F11/32—Monitoring with visual or acoustical indication of the functioning of the machine
- G06F11/323—Visualisation of programs or trace data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/16—File or folder operations, e.g. details of user interfaces specifically adapted to file systems
- G06F16/168—Details of user interfaces specifically adapted to file systems, e.g. browsing and visualisation, 2d or 3d GUIs
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/17—Details of further file system functions
- G06F16/1734—Details of monitoring file system events, e.g. by the use of hooks, filter drivers, logs
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/18—File system types
- G06F16/1805—Append-only file systems, e.g. using logs or journals to store data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/903—Querying
- G06F16/9038—Presentation of query results
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/54—Interprogram communication
- G06F9/542—Event management; Broadcasting; Multicasting; Notifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/54—Interprogram communication
- G06F9/547—Remote procedure calls [RPC]; Web services
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2201/00—Indexing scheme relating to error detection, to error correction, and to monitoring
- G06F2201/86—Event-based monitoring
Definitions
- the present invention relates to a technique for reducing the processing time of an event log in a log replay process for reproducing an event log used for process mining.
- Log regeneration is a technique of visualizing and reproducing this log record through a process model, and is used in a process mining field in which knowledge is extracted from an event log recorded in a log file as a process is performed.
- An embodiment of the present invention uses a MapReduce algorithm to divide a large event log into small partitions based on a distributed API web service, thereby improving the processing speed for reproducing the event log, thereby eliminating bottlenecks and deviations. It aims to solve.
- a method of reproducing a large amount of event logs comprising: counting an event log generated before the next step performed following the step, and maintaining the log in the storage system as a log file; Generating a plurality of divided log files by identifying a log file corresponding to a process included in the access command, and dividing the checked log file into a predetermined size as a connection command is generated. To obtain from the storage system.
- the large-capacity event log reproducing system includes a process, a recording unit for counting the event log generated before the next process performed subsequent to the process and maintaining the log in the storage system as a log file, and log reproducing.
- a connection command for a page is generated, a log unit corresponding to a process included in the access command is divided into a confirmation unit for confirming in the storage system, and the identified log file is divided into a predetermined size, and a plurality of divided logs are provided.
- an acquisition unit for generating a file and obtaining the file from the storage system.
- the event log in the log replay process of reproducing the event log, by using a MapReduce algorithm, by partitioning the large event log into a small partition based on a distributed API web service, the event log
- a large event log Split and distributed processing can improve log replay performance.
- FIG. 1 is a diagram showing the configuration of a large-scale event log playback system according to an embodiment of the present invention.
- FIG. 2 is a block diagram showing the internal configuration of a large-scale event log playback system according to an embodiment of the present invention.
- FIG. 3 is a diagram illustrating an example of an event log and a log file in the large-capacity event log playback system according to an embodiment of the present invention.
- FIG. 4 is a diagram illustrating an example of a process model generated from an event log in a large event log reproducing system according to an embodiment of the present invention.
- 5A and 5B are flowcharts illustrating a series of processes for providing a log replay page in a large event log replay system according to an embodiment of the present invention.
- FIG. 6 is a detailed flowchart illustrating step 517 shown in FIG. 5B in detail.
- FIG. 7 is a detailed flowchart showing step 524 shown in FIG. 5B in detail.
- FIG. 8 is a diagram illustrating an example of visualizing a log play page in a large event log play system according to an embodiment of the present invention.
- FIG. 9 is a flowchart illustrating a procedure of a method for reproducing a large event log according to an embodiment of the present invention.
- FIG. 1 is a diagram showing the configuration of a large-scale event log playback system according to an embodiment of the present invention.
- the mass event log playing system 100 may include a web browser 110, an HDFS 120, and a SPARK 130.
- the web browser 110 is one of web applications that are executed in the administrator terminal. For example, Google's 'Chrome' may be used.
- the web browser 110 checks the log file corresponding to the process (process) included in the access command on the HDFS 120 as the access command for the log play page is generated in the administrator terminal.
- the plurality of divided log files obtained by dividing the file into predetermined sizes may be sequentially obtained from the HDFS 120.
- the web browser 110 may process a predetermined number of event logs included in each of the plurality of divided log files based on REST_API, and may implement a log replay page indicating an event occurrence level for each process in the manager terminal.
- the web browser 110 may process each of the plurality of split log files based on REST_API, and may visualize and implement the log play page in stages.
- the web browser 110 may wait for generation of an event log according to a process performed in the SPARK 130.
- the SPARK 130 may count the event log generated before the process and the next process performed subsequent to the process, and maintain the SPARK 130 in the HDFS 120 as a log file.
- the SPARK 130 is an example of a general-purpose high performance partition processing platform.
- the SPARK 130 may perform a function of distributing data by partitioning a process on a memory basis according to a MapReduce function.
- MapReduce is designed to handle massive data processing in distributed parallel computing, and is a software framework released in 2004. It can be composed of functions called Map and Reduce, which are generally used in functional programming.
- the large-capacity event log reproducing system 100 of the present invention can improve the processing speed of reproducing the event log by dividing the large-capacity event log into small partitions in the log reproducing process of reproducing the event log.
- the Representational state transfer API (REST_API) provides a one-way information processing interoperability among computer systems on the Internet, a web service that allows access requests to the systems or controls the textual display of web resources. It may refer to one of the APIs.
- the HDFS 120 is one of a storage system that distributes a large amount of data.
- the HDFS 120 may divide and store a large file to facilitate data processing.
- FIG. 2 is a block diagram showing the internal configuration of a large-scale event log playback system according to an embodiment of the present invention.
- the large-capacity event log reproducing system 200 includes a recording unit 210, a confirmation unit 220, an acquisition unit 230, and an implementation unit 240. can do.
- the recording unit 210 counts the event log generated before the process and the next process performed subsequent to the process, and maintains the log in the storage system as a log file.
- the recording unit 210 may record an event log generated for various work actions in a process performed by a split processing platform (eg, 'SPARK') as a log file in the order of occurrence.
- a split processing platform eg, 'SPARK'
- the recording unit 210 may distribute and store log files in a storage system (for example, 'HDFS') that distributes a large amount of data, thereby facilitating processing of large event logs during log reproduction.
- a storage system for example, 'HDFS'
- the recording unit 210 may process a process according to a log reproducing algorithm, and may be implemented by SPARK that maintains the processing result in HDFS.
- the event log is a set of events (eg, start, end, cancel, etc.) of unit tasks constituting a process (process), which is related to one case and indicates who performed what, when and where.
- the information may be recorded as a log file containing various additional information.
- the event log may be a resource (eg, performer, system, equipment, etc.) that initiates or performs a task, a timestamp of the event (time of event occurrence), or data associated with the event (eg, 'order size'). It may include various additional information of at least one of activity of the event.
- the activity may be information indicating the performance of a task, etc. as a basic unit constituting the process.
- the identification unit 220 checks the log file corresponding to the process included in the access command in the storage system.
- the identification unit 220 may be implemented by a web browser (for example, 'Chrome' of Google Inc.) that is executed in the manager terminal generating the access command.
- a web browser for example, 'Chrome' of Google Inc.
- the verification unit 220 When the log file associated with the access command is not maintained in the storage system (eg, 'HDFS'), the verification unit 220 generates an event log according to a process performed in a split processing platform (eg, 'SPARK'). Can wait.
- a split processing platform eg, 'SPARK'
- the acquirer 230 when the log file associated with the access command is not maintained in the storage system, the acquirer 230 generates an event log as a process is processed in the division processing platform, and records the event log. If this is maintained in the storage system, the generated log file may be obtained from the storage system.
- SPARK stores the result of processing the process (log file) in HDFS according to a log replay algorithm.
- HDFS can then return the saved HDFS file to the App Server & REST API and respond with JASON to the web browser.
- the split processing platform SPARK applies a map function to all event logs in each case. Map to a transition, apply a reduce function to the event transition, sort the event transition by creation time using a sort function, initialize a partition, and use an insert function By inserting a predetermined number of event transitions (eg, '5,000') into each of the partitions, maintaining each partition in the storage system, and maintaining summary information for each partition in the storage system. have.
- a predetermined number of event transitions eg, '5,000'
- the acquirer 230 divides the checked log file into a predetermined size (partition), generates a plurality of divided log files, and obtains the plurality of divided log files from the storage system (eg, 'HDFS').
- the acquirer 230 may transmit a partition command to the storage system, and generate the plurality of partition log files by dividing the log file in partition units in the storage system according to the partition command. .
- the acquirer 230 determines the size of the divided log file according to the type of the process or the capacity of the log file, and acquires each of the divided log files from the storage system in proportion to the determined size.
- the interval can be adjusted.
- the storage system may divide the log file in consideration of at least one of the number of event logs per partition and the number of partitions, which are designated by the partitioning command.
- the acquirer 230 may sequentially obtain each of the plurality of divided log files in correspondence with the order of being located in the log file. For example, the acquirer 230 may sequentially obtain a plurality of divided log files obtained by dividing the log file checked from the HDFS in units of partitions from the HDFS.
- the event log is recorded in the log file in response to the occurrence time, and while the obtaining unit 230 processes a predetermined number of event logs included in the first divided log file among the plurality of divided log files, A second partitioned log file located in the log file next to the first partitioned log file may be obtained from the storage system.
- the acquirer 230 may transmit a partition command to the storage system and obtain a specific partition log file identified by the token included in the partition command from the storage system.
- the implementation unit 240 implements a log reproduction page indicating the degree of occurrence of events for each process by rendering a predetermined number of event logs included in each of the divided log files.
- the implementation unit 240 may visualize the log play page by processing the plurality of split log files through a REST_API based web application ('web browser').
- the implementation unit 240 reproduces a log from a storage system (for example, Hadoop file system 'HDFS') by rendering a predetermined number of event logs included in each of a plurality of split log files obtained sequentially.
- a storage system for example, Hadoop file system 'HDFS'
- the page may be implemented in the manager terminal as shown in FIG. 8.
- the implementation unit 240 includes activities ('QW', 'PP', 'RC', 'SP', 'AN', 'PR', 'RW' and 'WS' included in the event log). ) As a node, connecting each node with an arc, representing the flow of the process through a circle moving along the arc between nodes, and indicating the amount of flow of the process through the size of the circle, thereby logging on the administrator terminal. You can visualize the playback page.
- a large event log Split and distributed processing can improve log replay performance.
- FIG. 3 is a diagram illustrating an example of a configuration of an event log and a log file in a large event log reproducing system according to an exemplary embodiment of the present invention.
- the event log represents a set of events used as an input for process mining
- an event is a unit task constituting a process, which includes a start of a task, It can indicate the actions recorded in the log file, such as termination and cancellation.
- process mining relates to a technique for extracting knowledge by analyzing an event log generated by a machine or an information related system.
- An event is related to one case, and a case is a sequence of events, which may refer to a process instance, which is objects processed by a process to be analyzed.
- the event log may include, for each event, various additional information available for process mining techniques.
- an event log may include information about who performed what, when, where and what resources (eg, performer, system, equipment, etc.), timestamp of the event, or event occurrence. Time), data associated with the event (eg, 'order size'), and activity of the event.
- the activity may be information indicating the performance of a task, etc. as a basic unit constituting the process.
- the event log 'W [4]' includes a set 'E' of an event 'e' and a set 'ET' of an event type 'et'. ), A set (A) of activities ('a'), a set ('R') of resources ('r'), a set ('C') of cases ('c'), and events A function that assigns a timestamp to a function ('t: E ⁇ iR 0 + '), a function that assigns an event type to each event ('et: E ⁇ ET'), and a function that associates each event with an activity (' a: E ⁇ EA '), a function that associates each event with a resource (' r: E ⁇ R ⁇ ⁇ '), and a function that associates each event with a case (' c: E ⁇ C ').
- the event log is generated for various work actions during the process execution, and can be recorded in the log file in the order of occurrence.
- the table shown in (ii) of FIG. 3 shows an example of a log file in which an event log is recorded.
- the event log includes an event identification information (ID), a timestamp (event occurrence time), an activity (action of an event), a resource (e.g., an operator, a system, Equipment, etc.) and cost, and the additional information may be recorded in a log file.
- ID event identification information
- timestamp event occurrence time
- activity action of an event
- resource e.g., an operator, a system, Equipment, etc.
- the event log (events) may be recorded in log files scattered in different databases, not in one log file.
- FIG. 4 is a diagram illustrating an example of a process model generated from an event log in a large event log reproducing system according to an embodiment of the present invention.
- the large-capacity event log replay system of the present invention may generate a process model through analysis of the event log according to a process mining technique.
- the large-scale event log replay system of the present invention focuses on process flows such as the sequence of tasks and grasps all possible path characteristics, for example the Petri net model, EPC, BPMN, UML activity diagram.
- a process model can be derived.
- the large-scale event log replay system of the present invention focuses on information about resources hidden in the event log, which actors (e.g., workers, systems, roles, departments) are involved in performing tasks and how they are connected. By identifying people, you can classify people by role or function to create an organizational structure or to derive a process model that shows social networks between people.
- actors e.g., workers, systems, roles, departments
- the large-capacity event log reproduction system of the present invention can derive a process model, focusing on the characteristics of the case. That is, the large-capacity event log replay system of the present invention may analyze the characteristics of the case through a path or a worker participating in the case, or may analyze the characteristics of the case according to the value of data associated with the case. For example, if there is a case representing a replenishment order, the quantity and supplier of the ordered product may be useful information.
- the large-capacity event log reproducing system of the present invention may derive a process model by analyzing the time and frequency of the event through a timestamp included in the event log.
- the large-capacity event log reproduction system of the present invention may perform bottleneck derivation, service level measurement, resource utilization monitoring, and prediction of remaining time of a case being performed.
- the process model 'G' includes a set of nodes 'N' and a set of arcs connecting the nodes ('ENN ⁇ N'). And a function (na: N ⁇ P (A)) that associates one node with a set of activities.
- FIG. 4 (ii) shows a graph of the process model derived based on the log file shown in FIG. 3 (ii).
- the large-capacity event log reproducing system may include activities (eg, 'register request', 'examine thoroughly', 'check ticket', 'decide', ' reject request ',' examine casually ', and' pay compensation ') may be generated as nodes, and the nodes may be connected with an arc to derive a graph of the process model as shown.
- activities eg, 'register request', 'examine thoroughly', 'check ticket', 'decide', ' reject request ',' examine casually ', and' pay compensation '
- 5A and 5B are flowcharts illustrating a series of processes for providing a log replay page in a large event log replay system according to an embodiment of the present invention.
- steps 501 to 502 when a web browser receives an open command for a log play page from an administrator terminal, the web browser requests a log play page to an app server & REST API.
- steps 503 to 505 when the web browser receives the response from the App Server & REST API in HTML, it calls a heuristic miner API to request an HDFS file (log file).
- step 506 the HDFS checks if valid data exists in the HDFS when an HDFS file request is received.
- HDFS If there is valid data in HDFS as a result of the check in step 506, then in step 507, HDFS returns the HDFS file to the App Server & REST API, and in step 508 the App Server & REST API, Answer JASON to the web browser.
- step 506 If there is no valid data in the HDFS as a result of the checking in step 506, in steps 509 to 510, the SPARK processes the process according to the heuristic minor algorithm, and sends the processed result (log file) to the HDFS. Save it.
- step 507 HDFS returns the stored HDFS file to the App Server & REST API, and in step 508 the App Server & REST API responds with JASON to the web browser.
- JASON JavaScript Object Notation
- JavaScript Object Notation is a lightweight data exchange format that is easy to read and write to humans, easy to parse and generate on the machine, and is based on a subset of the JavaScript programming language.
- step 511 the web browser renders the process model and initializes a user interaction function.
- the web browser calls a log play page to request an HDFS file.
- step 514 the HDFS, upon receiving an HDFS file request, verifies that valid data exists in the HDFS.
- step 514 If the verification in step 514 indicates that valid data exists in HDFS, then in step 515 HDFS returns the HDFS file to the App Server & REST API, and in step 516 the App Server & REST API, Answer JASON to the web browser.
- step 514 If there is no valid data in the HDFS as a result of the check in step 514, in steps 517 to 518, the SPARK processes the process according to the log replay algorithm and sends the processed result (log file) to the HDFS. Save it.
- step 515 HDFS returns the stored HDFS file to the App Server & REST API, and in step 516 the App Server & REST API responds with JASON to the web browser.
- step 519 the web browser initializes the animation and lists all partitions.
- the web browser initializes a parent timeline ('tmp') for animation handling.
- steps 520 through 521 the web browser calls the partition API to obtain token data of partition P i , and in step 522, HDFS returns the HDFS file to the App Server & REST API using the token data. In step 523, the App Server & REST API responds with JASON to the web browser.
- the web browser initializes the animation for the token of the partition 'P i ', plays the animation, and then obtains the next partition 'P i + 1 '.
- FIG. 6 is a detailed flowchart illustrating step 517 shown in FIG. 5B in detail.
- FIG. 6 shows a detailed process of processing a process according to a log replay algorithm in SPARK and storing the processed result in HDFS.
- SPARK retrieves a list of cases 'C' from the system.
- SPARK begins by mapping a list of event transitions ('k') in the event log ('E').
- the list of cases 'C' is retrieved from the system, and each case 'C i ' contains a sequence of events 'e'.
- step 602 the SPARK checks whether a next case 'C i + 1 ' exists.
- step 603 SPARK lists the event 'e' in the case.
- the SPARK repeats each case 'C i ' to examine the sequence of events 'e'.
- step 604 the SPARK checks whether a next event ('e i + 1 ') exists.
- step 605 SPARK applies an MAP function to adjust the event transition K i + 1 attribute value.
- an event transition may indicate a change between two consecutive events.
- the SPARK repeats the sequence of events 'e ij ' in the case 'C i ' to set the event transition 'K ij ' attribute.
- step 604 when the next event 'e i + 1 ' does not exist, the SPARK moves to step 602 to re-check whether the next case 'C i + 1 ' exists. do.
- step 606 SPARK applies the REDUCE function to the event transition (K).
- the SPARK executes the reduce function to collect the entire event transition data K from the distributed system.
- step 607 SPARK applies a SORT function to sort the event transition 'K' according to the start time.
- the result value is formed as a sequence of event transition 'K', and the list of event transition 'K' data is sorted according to the start time.
- step 608 SPARK initializes the list of partitions 'P', and in step 609, SPARK checks whether the next partition 'P i + 1 ' exists.
- SPARK repeatedly initializes the list of partitions 'P' including the limited number of empty slots per partition 'I'.
- step 609 when the next partition 'P i + 1 ' exists, in step 610, SPARK initializes the partition 'P i ', and in step 611, Check for an empty slot.
- step 611 if an empty slot exists, in step 612, the SPARK inputs an event transition 'K' to the empty slot, and then moves to step 611 to display the empty slot. Double check that it exists.
- step 611 if no empty slot exists, in steps 613 to 614, SPARK sets a partition ('P i ') attribute value, and partitions ('P i ') to HDFS. Save the data.
- SPARK needs to set the partition ('P i ') attribute value and store partition data in HDFS before repeating the next partition ('P i + 1 ').
- the SPARK is the Partitions Summary ('Ps') attribute value. And store partition summary ('Ps') data in HDFS.
- FIG. 7 is a detailed flowchart showing step 524 shown in FIG. 5B in detail.
- FIG. 7 illustrates a process of initializing an animation for a token of a partition 'P i ' in a web browser.
- steps 701 to 702 the web browser stops the parent timeline 'tmp' and handles child time, for animation handling for partition 'P i '. 'tmc i ') is initialized.
- the web browser lists the tokens K in the partition 'P i ' and groups the tokens K according to the same source, target, start time and completion time. .
- the web browser lists the grouped tokens Kg in the partition 'P i ', generates an SVG for the grouped tokens Kg, and timelines the generated Kg i SVG. to (tmc i )
- SVG Scalable Vector Graphics
- W3C World Wide Web Consortium
- step 708 the web browser checks whether the next grouped token Kg i + 1 exists, and if it is found that the next grouped token Kg i + 1 exists, go to step 706 to the next step. Generates an SVG for the grouped tokens (Kg).
- step 709 the web browser resumes the parent timeline 'tmp'.
- FIG. 8 is a diagram illustrating an example of visualizing a log play page in a large event log play system according to an embodiment of the present invention.
- a log file corresponding to a process included in the access command is determined. By dividing by size, a plurality of split log files can be generated.
- the large-capacity event log reproducing system sequentially acquires each of the plurality of divided log files from a storage system (for example, Hadoop file system 'HDFS'), and generates a predetermined number of event logs included in each of the plurality of divided log files.
- a storage system for example, Hadoop file system 'HDFS'
- the log play page may be implemented in the manager terminal as shown in FIG. 8.
- the large-capacity event log reproducing system includes activities ('QW', 'PP', 'RC', 'SP', 'AN', 'PR', 'RW' and 'WS) included in the event log.
- activities 'QW', 'PP', 'RC', 'SP', 'AN', 'PR', 'RW' and 'WS.
- each node represents the flow of the process through a circle moving along the arc between nodes, and indicating the amount of flow of the process through the size of the circle. You can visualize the log replay page.
- FIG. 9 will be described in detail the workflow of the large-capacity event log playback system 200 according to the embodiments of the present invention.
- FIG. 9 is a flowchart illustrating a procedure of a method for reproducing a large event log according to an embodiment of the present invention.
- the large-capacity event log reproducing method according to the present embodiment may be performed by the large-capacity event log reproducing system 200 described above.
- step 910 the large-capacity event log reproducing system 200 counts the event log occurring before the process and the next process performed subsequent to the process, and maintains it in the storage system as a log file. .
- the large-capacity event log reproducing system 200 records the event log generated for various work actions in the process of executing the process by the split processing platform (eg, 'SPARK') as a log file in the order of occurrence, and logs the log file By distributing and storing data in distributed storage systems (eg, 'HDFS'), it is easier to process large event logs during log replay.
- the split processing platform eg, 'SPARK'
- 'HDFS' distributed storage systems
- step 920 the mass event log reproducing system 200 determines whether an access command for a log reproducing page occurs in the manager terminal.
- step 930 the mass event log reproducing system 200 generates a log file corresponding to the process included in the access command. Check on the storage system.
- step 940 the large-capacity event log reproducing system 200 divides the identified log file into a predetermined size (partition), generates a plurality of divided log files, and stores the plurality of divided log files from the storage system (eg, 'HDFS'). Acquire.
- the mass event log reproducing system 200 transmits a split command to the storage system, and generates the plurality of split log files by dividing the log file in partition units in the storage system according to the split command. You can do that.
- the large-capacity event log reproducing system 200 determines the size of the split log file according to the type of the process or the capacity of the log file, and in proportion to the determined size, the split log from the storage system. You can adjust the acquisition interval for each file.
- the large-capacity event log reproducing system 200 may obtain a plurality of split log files sequentially from the HDFS. That is, the large-capacity event log reproducing system 200 may be arranged in the log file, after the first split log file, while a predetermined number of event logs included in the first split log file are processed.
- the second split log file located may be obtained from the storage system.
- the large-capacity event log reproducing system 200 visualizes the log replay page on a web basis through processing of each of the divided log files.
- the large-capacity event log reproducing system 200 may process the plurality of divided log files through a REST_API-based web application ('web browser') to visualize a log reproducing page indicating an event occurrence level for each process. have.
- a REST_API-based web application 'web browser'
- the mass event log reproducing system 200 may divide the processing platform (eg, 'SPARK'). ) Can wait for the generation of the event log.
- the large-capacity event log reproducing system 200 when the log file associated with the access command is not maintained in the storage system, the large-capacity event log reproducing system 200 generates an event log as the process is processed in the divided processing platform, thereby generating the event log. If the log file recorded therein is maintained in the storage system, the generated log file may be obtained from the storage system.
- the split processing platform SPARK applies a map function to all event logs in each case. Map to a transition, apply a reduce function to the event transition, sort the event transition by creation time using a sort function, initialize a partition, and use an insert function By inserting a predetermined number of event transitions (eg, '5,000') into each of the partitions, maintaining each partition in the storage system, and maintaining summary information for each partition in the storage system. have.
- a predetermined number of event transitions eg, '5,000'
- the method according to an embodiment of the present invention can be implemented in the form of program instructions that can be executed by various computer means and recorded in a computer readable medium.
- the computer readable medium may include program instructions, data files, data structures, etc. alone or in combination.
- the program instructions recorded on the media may be those specially designed and constructed for the purposes of the embodiments, or they may be of the kind well-known and available to those having skill in the computer software arts.
- Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tape, optical media such as CD-ROMs, DVDs, and magnetic disks, such as floppy disks.
- Examples of program instructions include not only machine code generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
- the hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Software Systems (AREA)
- Quality & Reliability (AREA)
- Computational Linguistics (AREA)
- Computer Hardware Design (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Human Computer Interaction (AREA)
- Debugging And Monitoring (AREA)
Abstract
대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템이 개시된다. 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 방법은, 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지하는 단계와, 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인하는 단계, 및 상기 확인된 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템으로부터 획득하는 단계를 포함한다.
Description
본 발명은 프로세스 마이닝(Process Mining)에 사용되는 이벤트 로그(Event Log)를 재생하는 로그 재생(Log Replay) 과정에서, 이벤트 로그의 처리 시간을 줄이는 기술에 관한 것이다.
로그 재생이란, 이번의 로그기록을 프로세스 모델을 통해 시각화하여 재생하는 기술로서, 프로세스 수행에 따라 로그파일에 기록되는 이벤트 로그로부터 지식을 추출하는 프로세스 마이닝 분야에서 사용되고 있다.
기존의 로그 재생 기술에 따르면, 프로세스 수행에 따라 대량의 이벤트 로그가 발생되면, 이벤트 로그를 처리하여 토큰 애니메이션을 생성하고 real KPI를 계산하는 과정에서 병목과 편차가 발생하는 문제점이 생길 수 있다.
이에 따라, 로그 재생 과정에서 발생하는 병목과 편차 문제를 해결하고, 연산 처리 속도를 향상시킬 수 있는 기술이 요구되고 있다.
본 발명의 실시예는 맵 리듀스(MapReduce) 알고리즘을 이용하여, 분산 API 웹 서비스를 기반으로 대용량 이벤트 로그를 작은 파티션으로 분할 함으로써, 이벤트 로그를 재생하는 처리 속도를 향상시켜, 병목과 편차 문제를 해결하는 것을 목적으로 한다.
본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 방법은, 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지하는 단계와, 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인하는 단계, 및 상기 확인된 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템으로부터 획득하는 단계를 포함한다.
또한, 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템은, 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지하는 기록부와, 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인하는 확인부, 및 상기 확인된 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템으로부터 획득하는 획득부를 포함한다.
본 발명의 일실시예에 따르면, 이벤트 로그를 재생하는 로그 재생 과정에서, 맵 리듀스(MapReduce) 알고리즘을 이용하여, 분산 API 웹 서비스를 기반으로 대용량 이벤트 로그를 작은 파티션으로 분할 함으로써, 이벤트 로그를 재생하는 처리 속도를 향상시켜, 병목과 편차 문제를 해결할 수 있다.
또한, 본 발명의 일실시예에 따르면, 대용량 이벤트 로그를 분할하여 분산 API 웹 서비스 형태로 애니메이션을 생성해 Real KPI를 계산하고, 분할된 이벤트 로그를 순차적으로 클라이언트 웹 어플리케이션에 재생 함으로써, 대용량 이벤트 로그의 분할 및 분산 처리를 통해 로그 재생 처리 성능을 향상시킬 수 있다.
도 1은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템의 구성을 도시한 도면이다.
도 2는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템의 내부 구성을 도시한 블록도이다.
도 3은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 이벤트 로그 및 로그파일의 일례를 도시한 도면이다.
도 4는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 이벤트 로그로부터 생성되는 프로세스 모델의 일례를 도시한 도면이다.
도 5a 및 5b는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 로그 재생 페이지를 제공하는 일련의 과정을 도시한 흐름도이다.
도 6은, 도 5b에 도시한 단계(517)를 세부적으로 나타낸 상세흐름도이다.
도 7은, 도 5b에 도시한 단계(524)를 세부적으로 나타낸 상세흐름도이다.
도 8은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 로그 재생 페이지를 시각화 하는 일례를 도시한 도면이다.
도 9는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 방법의 순서를 도시한 흐름도이다.
이하, 첨부된 도면들을 참조하여 본 발명의 일실시예에 따른 응용프로그램 업데이트 장치 및 방법에 대해 상세히 설명한다. 그러나, 본 발명이 실시예들에 의해 제한되거나 한정되는 것은 아니다. 각 도면에 제시된 동일한 참조 부호는 동일한 부재를 나타낸다.
도 1은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템의 구성을 도시한 도면이다.
도 1을 참조하면, 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템(100)은, 웹 브라우저(110), HDFS(120) 및 SPARK(130)를 포함하여 구성될 수 있다.
웹 브라우저(110)는 관리자 단말에서 실행되는 웹 어플리케이션 중 하나로서, 일례로, 구글 사의 '크롬(Chrome)'을 예로 들 수 있다.
웹 브라우저(110)는 관리자 단말에서 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정(프로세스)에 대응되는 로그파일을, HDFS(120)에서 확인하고, 확인된 로그파일을 정해진 크기로 분할한 복수의 분할 로그파일을, HDFS(120)으로부터 순차적으로 획득할 수 있다.
웹 브라우저(110)는 REST_API 기반으로 상기 복수의 분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그를 처리하여, 공정 별 이벤트 발생 정도를 나타내는 로그 재생 페이지를 관리자 단말에 구현할 수 있다.
즉, 웹 브라우저(110)는 REST_API 기반으로 복수의 분할 로그파일 각각을 처리하여, 상기 로그 재생 페이지를 단계적으로 시각화하여 구현할 수 있다.
실시예에 따라, 웹 브라우저(110)는 HDFS(120)에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 경우, SPARK(130)에서의 공정 수행에 따른 이벤트 로그의 발생을 대기할 수 있다.
여기서, SPARK(130)는 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 HDFS(120)에 유지할 수 있다.
SPARK(130)는 범용 고성능 분할 처리 플랫폼의 일례로서, 맵리듀스(MapReduce) 함수에 따라, 메모리 기반으로 프로세스를 분할 수행하여 데이터를 분산 처리하는 기능을 할 수 있다.
여기서, 맵리듀스(MapReduce)는 대용량 데이터 처리를 분산 병렬 컴퓨팅에서 처리하기 위한 목적으로 제작되었으며, 2004년 발표한 소프트웨어 프레임워크로서 함수형 프로그래밍에서 일반적으로 사용되는 Map과 Reduce라는 함수로 구성될 수 있다.
이에 따라, 본 발명의 대용량 이벤트 로그 재생 시스템(100)은 이벤트 로그를 재생하는 로그 재생 과정에서, 대용량 이벤트 로그를 작은 파티션으로 분할 함으로써, 이벤트 로그를 재생하는 처리 속도를 향상시킬 수 있다.
여기서, REST_API(Representational state transfer API)는, 인터넷 상의 컴퓨터 시스템들 사이에서 일방향의 정보 처리 상호 운용(interoperability)을 제공하며, 시스템들에 접근 요청을 허용하거나, 웹 리소스의 텍스트 표시를 제어하는 웹 서비스 API의 하나를 지칭할 수 있다.
HDFS(Hadoop File System)(120)는, 대량의 자료를 분산 처리하는 스토리지 시스템의 하나로서, 대용량 파일을 나눠서 저장하여 데이터 처리를 용이하게 할 수 있다.
도 2는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템의 내부 구성을 도시한 블록도이다.
도 2를 참조하면, 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템(200)은, 기록부(210), 확인부(220), 획득부(230) 및 구현부(240)를 포함하여 구성할 수 있다.
기록부(210)는 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지한다.
즉, 기록부(210)는 분할 처리 플랫폼(예, 'SPARK')에 의한 프로세스 수행 과정에서 다양한 작업행위에 대해 발생되는 이벤트 로그를 발생순서에 따라 로그파일로서 기록할 수 있다.
기록부(210)는 로그파일을, 대량의 자료를 분산 처리하는 스토리지 시스템(예, 'HDFS')에 분산하여 저장 함으로써, 로그 재생 시 대용량 이벤트 로그의 처리가 보다 용이해지도록 할 수 있다.
예를 들어, 도 5a 및 5b를 참조하면, 기록부(210)는 로그 재생 알고리즘에 따라 프로세스를 처리하고, 처리한 결과를, HDFS에 유지하는 SPARK에 의해 구현될 수 있다.
여기서, 이벤트 로그는, 프로세스(공정)를 구성하는 단위 작업인 이벤트(예컨대, 작업의 시작, 종료, 취소 등)들의 집합으로서, 하나의 케이스와 관련을 가지며, 누가 언제 어디서 어떤 작업을 수행하였는지에 대한 정보를 나타낼 수 있도록, 다양한 부가적인 정보를 포함하여 로그파일로서 기록될 수 있다.
일례로, 이벤트 로그는 작업을 시작하거나 수행하는 리소스(예, 작업수행자, 시스템, 장비 등), 이벤트의 타임스탬프(timestamp, 이벤트 발생시간) 및 이벤트와 연관된 데이터(예컨대, '주문의 규모'), 이벤트의 액티비티(Activity) 중 적어도 하나의 여러 가지 부가적인 정보를 포함할 수 있다. 여기서, 액티비티는 프로세스를 구성하는 기본 단위로서 작업의 수행 등을 나타내는 정보일 수 있다.
확인부(220)는 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인한다.
확인부(220)는, 상기 접속 명령을 발생하는 관리자 단말에서 실행되는 웹 브라우저(예컨대, 구글 사의 '크롬(Chrome)')에 의해 구현될 수 있다.
확인부(220)는 상기 스토리지 시스템(예, 'HDFS')에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 경우, 분할 처리 플랫폼(예, 'SPARK')에서의 공정 수행에 따른 이벤트 로그의 발생을 대기할 수 있다.
구체적으로, 획득부(230)는 상기 스토리지 시스템에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 경우, 상기 분할 처리 플랫폼에서, 프로세스가 처리 됨에 따라 이벤트 로그가 생성되어, 상기 이벤트 로그를 기록한 로그파일이, 상기 스토리지 시스템 내에 유지되면, 상기 스토리지 시스템으로부터, 상기 생성된 로그파일을 획득할 수 있다.
예를 들어, 도 5b의 단계(514 내지 518)를 참조하면, HDFS에 유효한 데이터가 존재하지 않는 것으로 확인되면, SPARK는, 로그 재생 알고리즘에 따라 프로세스를 처리한 결과(로그파일)를 HDFS에 저장하고, HDFS는 저장된 HDFS 파일을 앱 서버 & REST API로 반환하여, 웹 브라우저에 JASON으로 응답할 수 있다.
보다 구체적으로, 도 6을 참조하면, 상기 분할 처리 플랫폼(SPARK)은, 이벤트 로그의 시퀀스로 구성된 케이스에 대한 리스트가 입력되면, 상기 각 케이스 내 모든 이벤트 로그를 맵(Map) 함수를 적용하여 이벤트 트랜지션으로 매핑하고, 상기 이벤트 트랜지션에 리듀스(Reduce) 함수를 적용하고, 소트(Sort) 함수를 이용하여 상기 이벤트 트랜지션을 생성시점에 따라 정렬하고, 파티션을 초기화하고, 인서트(Insert) 함수를 이용하여, 상기 파티션 각각에 일정 개수(예를 들어, '5,000개')의 이벤트 트랜지션을 삽입하고, 상기 각 파티션을 상기 스토리지 시스템에 유지하고, 상기 각 파티션에 대한 요약정보를 상기 스토리지 시스템에 유지할 수 있다.
획득부(230)는 상기 확인된 로그파일을 정해진 크기(파티션)로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템(예, 'HDFS')으로부터 획득한다.
즉, 획득부(230)는 상기 스토리지 시스템으로 분할 명령을 전송하고, 상기 분할 명령에 따라, 상기 스토리지 시스템에서, 상기 로그파일을 파티션 단위로 분할하여 상기 복수의 분할 로그파일을 생성하도록 할 수 있다.
이때, 획득부(230)는 상기 공정의 종류 또는 상기 로그파일의 용량에 따라, 분할 로그파일이 갖는 크기를 정하고, 정해진 상기 크기에 비례하여, 상기 스토리지 시스템으로부터의, 상기 분할 로그파일 각각의 획득 간격을 조정할 수 있다.
상기 스토리지 시스템은, 상기 분할 명령에 의해 지정되는, 파티션 당 이벤트 로그의 개수 및 파티션의 개수 중 적어도 하나를 고려하여, 상기 로그파일을 분할할 수 있다.
일례로, 획득부(230)는 상기 로그파일 내에 위치하는 순서에 상응하여, 상기 복수의 분할 로그파일 각각을 순차적으로 획득할 수 있다. 예를 들어, 획득부(230)는 HDFS로부터 확인한 로그파일을 파티션 단위로 분할한 복수의 분할 로그파일을, HDFS로부터 순차적으로 획득할 수 있다.
또한, 상기 이벤트 로그는, 발생시점에 대응하여 상기 로그파일에 기록되고, 획득부(230)는 상기 복수의 분할 로그파일 중 제1 분할 로그파일에 포함되는 일정 개수의 이벤트 로그가 처리되는 동안, 상기 로그파일 내, 상기 제1 분할 로그파일 다음에 위치하는 제2 분할 로그파일을, 상기 스토리지 시스템으로부터 획득할 수도 있다.
또한, 획득부(230)는 상기 스토리지 시스템으로 분할 명령을 전송하고, 상기 분할 명령에 포함되는 토큰에 의해 식별되는 특정의 분할 로그파일을 상기 스토리지 시스템으로부터 획득할 수도 있다.
구현부(240)는 분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그에 대한 렌더링을 통해, 공정 별 이벤트 발생 정도를 나타내는 로그 재생 페이지를 구현한다.
즉, 구현부(240)는 REST_API 기반의 웹 어플리케이션('웹 브라우저')을 통해, 상기 복수의 분할 로그파일을 처리하여, 상기 로그 재생 페이지를 시각화 할 수 있다.
일례로, 구현부(240)는 스토리지 시스템(일례로, 하둡 파일 시스템 'HDFS')으로부터, 순차적으로 획득한 복수의 분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그에 대한 렌더링을 통해, 로그 재생 페이지를 도 8과 같이 관리자 단말에 구현할 수 있다.
도 8을 참조하면, 구현부(240)는 이벤트 로그에 포함된 액티비티('QW', 'PP', 'RC', 'SP', 'AN', 'PR', 'RW' 및 'WS')를 각각 노드로서 표시하고, 각 노드를 아크로 연결하고, 노드 사이에서 아크를 따라 움직이는 원을 통해 프로세스의 흐름을 표현하고, 원의 크기를 통해 프로세스의 흐름의 양을 나타냄으로써, 관리자 단말에서 로그 재생 페이지를 시각화 할 수 있다.
이와 같이, 본 발명의 일실시예에 따르면, 이벤트 로그를 재생하는 로그 재생 과정에서, 맵 리듀스(MapReduce) 알고리즘을 이용하여, 분산 API 웹 서비스를 기반으로 대용량 이벤트 로그를 작은 파티션으로 분할 함으로써, 이벤트 로그를 재생하는 처리 속도를 향상시켜, 병목과 편차 문제를 해결할 수 있다. 또한, 본 발명의 일실시예에 따르면, 대용량 이벤트 로그를 분할하여 분산 API 웹 서비스 형태로 애니메이션을 생성해 Real KPI를 계산하고, 분할된 이벤트 로그를 순차적으로 클라이언트 웹 어플리케이션에 재생 함으로써, 대용량 이벤트 로그의 분할 및 분산 처리를 통해 로그 재생 처리 성능을 향상시킬 수 있다.
도 3은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 이벤트 로그의 구성 및 로그파일의 일례를 도시한 도면이다.
본 발명에서, 이벤트 로그(Event Log)는, 프로세스 마이닝(Process Mining)을 위한 투입으로 사용되는 이벤트들의 집합을 나타내고, 이벤트(Event)는 프로세스(공정)를 구성하는 단위 작업으로서, 작업의 시작, 종료, 취소 등과 같이 로그파일에 기록되는 행위를 나타낼 수 있다.
참고로, 프로세스 마이닝은 기계나 정보 관련 시스템에서 생성되는 이벤트 로그(transaction dataset)를 분석하여 지식을 추출하는 기법에 관한 것이다.
이벤트(이벤트 로그)는 하나의 케이스와 관련을 가지며, 케이스(Case)는 이벤트들의 시퀀스로서, 분석할 프로세스에 의해 처리된 개체들인 프로세스 인스턴스(Process Instance)를 지칭할 수 있다.
이벤트 로그는, 각 이벤트에 대해, 프로세스 마이닝 기법에 활용 가능한 다양한 부가적인 정보를 포함하여 구성될 수 있다.
일례로, 이벤트 로그는, 누가 언제 어디서 어떤 작업을 수행하였는지에 대한 정보를 나타낼 수 있도록, 작업을 시작하거나 수행하는 리소스(예, 작업수행자, 시스템, 장비 등), 이벤트의 타임스탬프(timestamp, 이벤트 발생시간) 및 이벤트와 연관된 데이터(예컨대, '주문의 규모'), 이벤트의 액티비티(Activity) 중 적어도 하나의 여러 가지 부가적인 정보를 포함할 수 있다. 여기서, 액티비티는 프로세스를 구성하는 기본 단위로서 작업의 수행 등을 나타내는 정보일 수 있다.
도 3의 (i)은 이벤트 로그('W[4]')의 구성의 일례를 나타내고 있다.
도 3의 (i)을 참조하면, 이벤트 로그('W[4]')는, 이벤트('e')의 집합('E')과, 이벤트 타입('et')의 집합('ET')과, 액티비티('a')의 집합('A')과, 리소스('r')의 집합('R')과, 케이스('c')의 집합('C')과, 각 이벤트에 타임스탬프를 할당하는 함수('t:E→iR0
+')와, 각 이벤트에 이벤트 타입을 할당하는 함수('et:E→ET')와, 각 이벤트를 액티비티에 연관시키는 함수('a:E→EA')와, 각 이벤트를 리소스에 연관시키는 함수('r:E→R∪{⊥}'), 및 각 이벤트를 케이스에 연관시키는 함수('c:E→C')로 구성될 수 있다.
이벤트 로그는, 프로세스 수행 과정에서 다양한 작업행위에 대해 발생되며, 발생순서에 따라 로그파일에 기록될 수 있다.
도 3의 (ii)에 도시한 테이블은 이벤트 로그를 기록한 로그파일의 일례를 나타내고 있다.
도 3의 (ii)를 참조하면, 이벤트 로그는, 각 이벤트를, 이벤트 식별정보(ID)와, 타임스탬프(이벤트 발생시간), 액티비티(이벤트의 작업), 리소스(예, 작업수행자, 시스템, 장비 등) 및 코스트 중 적어도 하나의 부가정보에 연관시켜 로그파일에 기록할 수 있다.
이때, 이벤트 로그(이벤트들)은 하나의 로그파일이 아닌, 서로 다른 데이터베이스에 흩어져 있는 로그파일에 기록될 수도 있다.
도 4는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 이벤트 로그로부터 생성되는 프로세스 모델의 일례를 도시한 도면이다.
본 발명의 대용량 이벤트 로그 재생 시스템은, 프로세스 마이닝 기법에 따라 이벤트 로그에 대한 분석을 통해 프로세스 모델(Process Model)을 생성할 수 있다.
일례로, 본 발명의 대용량 이벤트 로그 재생 시스템은, 작업의 순서와 같은 프로세스의 흐름에 중점을 두고, 모든 가능한 경로의 특징을 파악 함으로써, 예를 들어 Petri net 모델이나, EPC, BPMN, UML activity diagram과 같은 프로세스 모델을 도출할 수 있다.
또한, 본 발명의 대용량 이벤트 로그 재생 시스템은, 이벤트 로그에 숨겨진 리소스에 대한 정보에 초점을 맞추어, 어떤 행위자(예, 작업자, 시스템, 역할, 부서)가 업무 수행에 관여하고, 그들이 어떻게 연결되어 있는지 파악하여, 역할이나 기능에 따라 사람들을 분류하여 조직 체계를 만들거나 사람들 사이의 소셜 네트워크를 보여주는 프로세스 모델을 도출할 수 있다.
또한, 본 발명의 대용량 이벤트 로그 재생 시스템은, 케이스의 특징에 초점을 맞추어, 프로세스 모델을 도출할 수 있다. 즉, 본 발명의 대용량 이벤트 로그 재생 시스템은, 프로세스 내에서의 경로 또는 케이스에 참여하는 작업자를 통해 케이스의 특성을 분석하거나, 케이스와 연관된 데이터의 값에 따라 케이스의 특성을 분석할 수도 있다. 예를 들어, 보충 주문을 나타내는 케이스가 있을 경우, 주문된 상품의 수량과 공급자는 유용한 정보가 될 수 있다.
또한, 본 발명의 대용량 이벤트 로그 재생 시스템은, 이벤트 로그에 포함되는 타임스탬프를 통해 이벤트의 시간과 빈도를 분석하여, 프로세스 모델을 도출할 수도 있다. 일례로, 본 발명의 대용량 이벤트 로그 재생 시스템은, 병목점 도출, 서비스의 레벨 측정, 리소스 활용도 모니터링 및 수행 중인 케이스의 잔여 시간 예측 등을 수행할 수 있다.
도 4의 (i)은 프로세스 모델('G')의 구성의 일례를 나타내고 있다.
도 4의 (i)을 참조하면, 프로세스 모델('G')은 노드(node)의 집합('N')과, 노드들을 연결하는 아크(arc)의 집합('E⊆N×N') 및 하나의 노드를 액티비티의 집합에 연관시키는 함수(na:N→P(A))로 구성될 수 있다.
도 4의 (ii)는, 도 3의 (ii)에 도시한 로그파일에 근거하여 도출한 프로세스 모델의 그래프를 나타내고 있다.
도 4의 (ii)를 참조하면, 대용량 이벤트 로그 재생 시스템은, 도 3의 (ii)에 도시한 액티비티(예컨대, 'register request', 'examine thoroughly', 'check ticket', 'decide', 'reject request', 'examine casually', 'pay compensation')를 각각 노드로서 생성하고, 상기 각 노드를 아크로 연결하여, 도시한 것과 같은 프로세스 모델의 그래프를 도출할 수 있다.
도 5a 및 5b는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 로그 재생 페이지를 제공하는 일련의 과정을 도시한 흐름도이다.
도 5a를 참조하면, 단계(501 내지 502)에서, 웹 브라우저는, 관리자 단말로부터 로그 재생 페이지에 대한 오픈(open) 명령이 수신되면, 로그 재생 페이지를 앱 서버 & REST API로 요청한다.
단계(503 내지 505)에서, 웹 브라우저는, 앱 서버 & REST API로부터 HTML로 응답을 수신하면, 휴리스틱 마이너 API(heuristic miner API)를 호출하여, HDFS 파일(로그파일)을 요청한다.
단계(506)에서 HDFS는, HDFS 파일 요청이 수신되면, HDFS에 유효한 데이터가 존재하는지를 확인한다.
단계(506)에서의 확인 결과, HDFS에 유효한 데이터가 존재하는 경우, 단계(507)에서, HDFS는 HDFS 파일을 앱 서버 & REST API로 반환하고, 단계(508)에서 앱 서버 & REST API는, 웹 브라우저에 JASON으로 응답한다.
단계(506)에서의 확인 결과, HDFS에 유효한 데이터가 존재하지 않는 경우, 단계(509 내지 510)에서, SPARK는, 휴리스틱 마이너 알고리즘에 따라 프로세스를 처리하고, 처리한 결과(로그파일)를 HDFS에 저장한다. 단계(507)에서, HDFS는 저장된 HDFS 파일을 앱 서버 & REST API로 반환하고, 단계(508)에서, 앱 서버 & REST API는, 웹 브라우저에 JASON으로 응답한다.
여기서, JASON(JavaScript Object Notation)은 경량 데이터 교환 포맷으로서, 사람에게 있어서 읽고 쓰기가 쉽고, 기계에 있어서 파싱과 생성이 용이하며, 자바스크립트 프로그래밍 언어의 서브셋에 기반을 두고 있다.
단계(511)에서, 웹 브라우저는, 프로세스 모델을 렌더링하고, 사용자 인터액션 함수(User Interaction function)를 초기화 한다.
도 5b를 참조하면, 단계(512 내지 513)에서, 웹 브라우저는 로그 재생 페이지를 호출하여 HDFS 파일을 요청한다.
단계(514)에서 HDFS는, HDFS 파일 요청이 수신되면, HDFS에 유효한 데이터가 존재하는지를 확인한다.
단계(514)에서의 확인 결과, HDFS에 유효한 데이터가 존재하는 경우, 단계(515)에서, HDFS는 HDFS 파일을 앱 서버 & REST API로 반환하고, 단계(516)에서 앱 서버 & REST API는, 웹 브라우저에 JASON으로 응답한다.
단계(514)에서의 확인 결과, HDFS에 유효한 데이터가 존재하지 않는 경우, 단계(517 내지 518)에서, SPARK는, 로그 재생 알고리즘에 따라 프로세스를 처리하고, 처리한 결과(로그파일)를 HDFS에 저장한다. 단계(515)에서, HDFS는 저장된 HDFS 파일을 앱 서버 & REST API로 반환하고, 단계(516)에서, 앱 서버 & REST API는, 웹 브라우저에 JASON으로 응답한다.
단계(519)에서, 웹 브라우저는, 애니매이션을 초기화 하고, 모든 파티션을 리스팅한다.
구체적으로, 웹 브라우저는, 애니메이션 핸들링을 위해, 모 타임라인(parent timeline, 'tmp')을 초기화 한다.
단계(520 내지 521)에서, 웹 브라우저는, 파티션 API를 호출하여 파티션 Pi의 토큰 데이터를 획득하고, 단계(522)에서, HDFS는 상기 토큰 데이터를 이용해 HDFS 파일을 앱 서버 & REST API로 반환하고, 단계(523)에서, 앱 서버 & REST API는, 웹 브라우저에 JASON으로 응답한다.
단계(524 내지 526)에서, 웹 브라우저는, 파티션('Pi')의 토큰에 대한 애니메이션을 초기화 하고, 애니메이션을 재생한 후, 차기 파티션('Pi+1')을 획득한다.
도 6은, 도 5b에 도시한 단계(517)를 세부적으로 나타낸 상세흐름도이다.
도 6에는, SPARK에서 로그 재생 알고리즘에 따라 프로세스를 처리하고, 처리한 결과를 HDFS에 저장하는 상세한 과정이 도시되어 있다.
도 6을 참조하면, 단계(601)에서, SPARK는 케이스('C')의 리스트를 시스템으로부터 검색한다.
SPARK는 이벤트 로그('E')에서 이벤트 트랜지션('k')의 리스트를 매핑하는 것으로 시작한다. 케이스('C')의 리스트는 시스템으로부터 검색되고, 각 케이스('Ci')는 이벤트('e')의 시퀀스를 포함한다.
단계(602)에서, SPARK는 차기 케이스('Ci+1')가 존재하는지 확인한다.
상기 단계(602)에서의 확인 결과, 차기 케이스('Ci+1')가 존재하는 경우, 단계(603)에서, SPARK는, 케이스 내에 이벤트('e')를 리스팅 한다.
본 단계(602 내지 603)에서, SPARK는 이벤트('e')의 시퀀스를 검사하기 위해 각 케이스('Ci')를 반복한다.
단계(604)에서, SPARK는, 차기 이벤트('ei+1')가 존재하는지 확인한다.
상기 단계(604)에서의 확인 결과, 차기 이벤트('ei+1')가 존재하는 경우, 단계(605)에서, SPARK는, MAP 함수를 적용하여 이벤트 트랜지션(Ki+1) 속성 값을 설정한다. 여기서, 이벤트 트랜지션(Event Transition)은, 연속된 2개의 이벤트 간 변화를 나타낼 수 있다.
본 단계(604 내지 605)에서, SPARK는 이벤트 트랜지션('Kij') 속성을 설정하기 위해 케이스('Ci')에 이벤트('eij')의 시퀀스를 반복한다.
상기 단계(604)에서의 확인 결과, 차기 이벤트('ei+1')가 존재하지 않는 경우, SPARK는 단계(602)로 이동하여, 차기 케이스('Ci+1')가 존재하는지 재확인한다.
상기 단계(602)에서의 확인 결과, 차기 케이스('Ci+1')가 존재하지 않는 경우, 단계(606)에서, SPARK는, 이벤트 트랜지션('K')에 REDUCE 함수를 적용한다.
본 단계(606)에서, SPARK는 맵 함수를 완료한 후, 리듀스 함수를 실행하여, 전체 이벤트 트랜지션 데이터 K를 분산 시스템으로부터 수집한다.
단계(607)에서, SPARK는, 소트(SORT) 함수를 적용하여 이벤트 트랜지션('K')을 시작시간(start time)에 따라 정렬한다. 여기서, 결과값은 이벤트 트랜지션('K')의 시퀀스로서 형성되며, 이벤트 트랜지션('K') 데이터의 리스트는 시작시간에 따라 정렬된다.
단계(608)에서, SPARK는, 파티션('P')의 리스트를 초기화 하고, 단계(609)에서, SPARK는 차기 파티션('Pi+1')이 존재하는지 확인한다.
여기서, SPARK는 파티션('I') 당 제한되는 빈 슬롯 수를 포함하여 파티션('P')의 리스트를 반복적으로 초기화 한다.
상기 단계(609)에서의 확인 결과, 차기 파티션('Pi+1')이 존재하는 경우, 단계(610)에서, SPARK는 파티션('Pi')를 초기화 하고, 단계(611)에서, 빈 슬롯이 존재하는지 확인한다.
상기 단계(611)에서의 확인 결과, 빈 슬롯이 존재하는 경우, 단계(612)에서, SPARK는 빈 슬롯에 이벤트 트랜지션('K')를 입력한 후, 단계(611)로 이동하여 빈 슬롯이 존재하는지 재확인한다.
본 단계(611 내지 612)에서, SPARK는 각 파티션은 빈 슬롯을 포함하기 때문에, 이벤트 트랜지션('K') 속성을 빈 슬롯에 반복적으로 입력한다.
상기 단계(611)에서의 확인 결과, 빈 슬롯이 존재하지 않는 경우, 단계(613 내지 614)에서, SPARK는 파티션('Pi') 속성 값을 설정하고, HDFS로 파티션('Pi') 데이터를 저장한다.
여기서, SPARK는 차기 파티션('Pi+1')을 반복하기 전에, 파티션('Pi') 속성 값을 설정한 후 파티션 데이터를 HDFS에 저장할 필요가 있다.
상기 단계(609)에서의 확인 결과, 차기 파티션('Pi+1')이 존재하지 않는 경우, 단계(615 내지 616)에서, SPARK는 파티션 요약정보(Partitions Summary)('Ps') 속성 값을 설정하고, 파티션 요약정보('Ps') 데이터를 HDFS에 저장한다.
여기서, SPARK는 마지막 파티션에서 이벤트 트랜지션의 처리를 완료한 후에, 파티션 요약정보('Ps') 속성 값을 설정한 뒤, 그 데이터를 HDFS에 저장한다.
도 7은, 도 5b에 도시한 단계(524)를 세부적으로 나타낸 상세흐름도이다.
도 7에는 웹 브라우저에서 파티션('Pi')의 토큰에 대한 애니메이션을 초기화 하는 과정이 도시되어 있다.
도 7을 참조하면, 단계(701 내지 702)에서, 웹 브라우저는 모 타임라인('tmp')을 중지하고, 파티션('Pi')을 위한 애니메이션 핸들링을 위해, 자 타임라인(child time, 'tmci')을 초기화 한다.
단계(703 내지 704)에서, 웹 브라우저는 파티션('Pi')에서 토큰(K)을 리스팅하고, 같은 소스(source), 타겟, 개시 시간 및 완료 시간에 따라, 토큰(K)을 그룹핑 한다.
단계(705 내지 707)에서, 웹 브라우저는 파티션('Pi')에 그룹핑된 토큰(Kg)을 리스팅하고, 그룹핑된 토큰(Kg)을 위한 SVG를 생성하고, 생성한 KgiSVG를 타임라인(tmci)에 추가한다.
여기서, SVG(Scalable Vector Graphics)는 상호성(interactivity)과 애니매이션(animation)을 지원하는 2차원 그래픽 XML 기반의 벡터 이미지 포맷으로서, W3C(World Wide Web Consortium)에 의해 개발된 공개 표준으로서, 동적이고 상호적이라는 특징을 가진다.
단계(708)에서, 웹 브라우저는 차기 그룹핑된 토큰(Kgi+1)이 존재하는지 확인하고, 차기 그룹핑된 토큰(Kgi+1)이 존재하는 것으로 확인되면, 단계(706)으로 이동하여 차기 그룹핑된 토큰(Kg)을 위한 SVG를 생성한다.
차기 그룹핑된 토큰(Kgi+1)이 존재하지 않는 것으로 확인되면, 단계(709)에서, 웹 브라우저는 모 타임라인('tmp')을 재개한다.
도 8은 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 시스템에서, 로그 재생 페이지를 시각화 하는 일례를 도시한 도면이다.
도 8을 참조하면, 본 발명의 실시예들에 따른 대용량 이벤트 로그 재생 시스템은, 로그 재생 페이지에 대한 접속 명령이 발생 함에 따라, 상기 접속 명령에 포함되는 공정(Process)에 대응되는 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성할 수 있다.
대용량 이벤트 로그 재생 시스템은, 복수의 분할 로그파일 각각을 스토리지 시스템(일례로, 하둡 파일 시스템 'HDFS')으로부터 순차적으로 획득하고, 상기 복수의 분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그에 대한 렌더링을 통해, 로그 재생 페이지를 도 8과 같이 관리자 단말에 구현할 수 있다.
도 8을 참조하면, 대용량 이벤트 로그 재생 시스템은, 이벤트 로그에 포함된 액티비티('QW', 'PP', 'RC', 'SP', 'AN', 'PR', 'RW' 및 'WS')를 각각 노드로서 표시하고, 각 노드를 아크로 연결하고, 노드 사이에서 아크를 따라 움직이는 원을 통해 프로세스의 흐름을 표현하고, 원의 크기를 통해 프로세스의 흐름의 양을 나타냄으로써, 관리자 단말에서 로그 재생 페이지를 시각화 할 수 있다.
이하, 도 9에서는 본 발명의 실시예들에 따른 대용량 이벤트 로그 재생 시스템(200)의 작업 흐름을 상세히 설명한다.
도 9는 본 발명의 일실시예에 따른 대용량 이벤트 로그 재생 방법의 순서를 도시한 흐름도이다.
본 실시예에 따른 대용량 이벤트 로그 재생 방법은 상술한 대용량 이벤트 로그 재생 시스템(200)에 의해 수행될 수 있다.
도 9를 참조하면, 단계(910)에서, 대용량 이벤트 로그 재생 시스템(200)은, 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지한다.
대용량 이벤트 로그 재생 시스템(200)은 분할 처리 플랫폼(예, 'SPARK')에 의한 프로세스 수행 과정에서 다양한 작업행위에 대해 발생되는 이벤트 로그를 발생순서에 따라 로그파일로서 기록하고, 로그파일을, 대량의 자료를 분산 처리하는 스토리지 시스템(예, 'HDFS')에 분산하여 저장 함으로써, 로그 재생 시 대용량 이벤트 로그의 처리가 보다 용이해지도록 할 수 있다.
단계(920)에서, 대용량 이벤트 로그 재생 시스템(200)은, 관리자 단말에서 로그 재생 페이지에 대한 접속 명령이 발생하는지 판단한다.
단계(920)에서의 판단 결과, 로그 재생 페이지에 대한 접속 명령이 발생하는 경우, 단계(930)에서, 대용량 이벤트 로그 재생 시스템(200)은, 상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인한다.
단계(940)에서, 대용량 이벤트 로그 재생 시스템(200)은, 확인된 로그파일을 정해진 크기(파티션)로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템(예, 'HDFS')으로부터 획득한다.
일례로, 대용량 이벤트 로그 재생 시스템(200)은 상기 스토리지 시스템으로 분할 명령을 전송하고, 상기 분할 명령에 따라, 상기 스토리지 시스템에서, 상기 로그파일을 파티션 단위로 분할하여 상기 복수의 분할 로그파일을 생성하도록 할 수 있다.
이때, 대용량 이벤트 로그 재생 시스템(200)은, 상기 공정의 종류 또는 상기 로그파일의 용량에 따라, 분할 로그파일이 갖는 크기를 정하고, 정해진 상기 크기에 비례하여, 상기 스토리지 시스템으로부터의, 상기 분할 로그파일 각각의 획득 간격을 조정할 수 있다.
대용량 이벤트 로그 재생 시스템(200)은, 복수의 분할 로그파일을 HDFS로부터 순차적으로 획득할 수 있다. 즉, 대용량 이벤트 로그 재생 시스템(200)은, 상기 복수의 분할 로그파일 중 제1 분할 로그파일에 포함되는 일정 개수의 이벤트 로그가 처리되는 동안, 상기 로그파일 내, 상기 제1 분할 로그파일 다음에 위치하는 제2 분할 로그파일을, 상기 스토리지 시스템으로부터 획득할 수 있다.
단계(950)에서, 대용량 이벤트 로그 재생 시스템(200)은, 분할 로그파일 각각에 대한 처리를 통해, 로그 재생 페이지를 웹 기반으로 시각화 한다.
즉, 대용량 이벤트 로그 재생 시스템(200)은, REST_API 기반의 웹 어플리케이션('웹 브라우저')을 통해, 상기 복수의 분할 로그파일을 처리하여, 공정 별 이벤트 발생 정도를 나타내는 로그 재생 페이지를 시각화 할 수 있다.
실시예에 따라, 대용량 이벤트 로그 재생 시스템(200)은, 상기 스토리지 시스템(예, 'HDFS')에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 것으로 확인되면, 분할 처리 플랫폼(예, 'SPARK')에서의 공정 수행에 따른 이벤트 로그의 발생을 대기할 수 있다.
구체적으로, 대용량 이벤트 로그 재생 시스템(200)은, 상기 스토리지 시스템에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 경우, 상기 분할 처리 플랫폼에서, 프로세스가 처리 됨에 따라 이벤트 로그가 생성되어, 상기 이벤트 로그를 기록한 로그파일이, 상기 스토리지 시스템 내에 유지되면, 상기 스토리지 시스템으로부터, 상기 생성된 로그파일을 획득할 수 있다.
보다 구체적으로, 도 6을 참조하면, 상기 분할 처리 플랫폼(SPARK)은, 이벤트 로그의 시퀀스로 구성된 케이스에 대한 리스트가 입력되면, 상기 각 케이스 내 모든 이벤트 로그를 맵(Map) 함수를 적용하여 이벤트 트랜지션으로 매핑하고, 상기 이벤트 트랜지션에 리듀스(Reduce) 함수를 적용하고, 소트(Sort) 함수를 이용하여 상기 이벤트 트랜지션을 생성시점에 따라 정렬하고, 파티션을 초기화하고, 인서트(Insert) 함수를 이용하여, 상기 파티션 각각에 일정 개수(예를 들어, '5,000개')의 이벤트 트랜지션을 삽입하고, 상기 각 파티션을 상기 스토리지 시스템에 유지하고, 상기 각 파티션에 대한 요약정보를 상기 스토리지 시스템에 유지할 수 있다.
본 발명의 실시예에 따른 방법은 다양한 컴퓨터 수단을 통하여 수행될 수 있는 프로그램 명령 형태로 구현되어 컴퓨터 판독 가능 매체에 기록될 수 있다. 상기 컴퓨터 판독 가능 매체는 프로그램 명령, 데이터 파일, 데이터 구조 등을 단독으로 또는 조합하여 포함할 수 있다. 상기 매체에 기록되는 프로그램 명령은 실시예를 위하여 특별히 설계되고 구성된 것들이거나 컴퓨터 소프트웨어 당업자에게 공지되어 사용 가능한 것일 수도 있다. 컴퓨터 판독 가능 기록 매체의 예에는 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체(magnetic media), CD-ROM, DVD와 같은 광기록 매체(optical media), 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical media), 및 롬(ROM), 램(RAM), 플래시 메모리 등과 같은 프로그램 명령을 저장하고 수행하도록 특별히 구성된 하드웨어 장치가 포함된다. 프로그램 명령의 예에는 컴파일러에 의해 만들어지는 것과 같은 기계어 코드뿐만 아니라 인터프리터 등을 사용해서 컴퓨터에 의해서 실행될 수 있는 고급 언어 코드를 포함한다. 상기된 하드웨어 장치는 실시예의 동작을 수행하기 위해 하나 이상의 소프트웨어 모듈로서 작동하도록 구성될 수 있으며, 그 역도 마찬가지이다.
이상과 같이 실시예들이 비록 한정된 실시예와 도면에 의해 설명되었으나, 해당 기술분야에서 통상의 지식을 가진 자라면 상기의 기재로부터 다양한 수정 및 변형이 가능하다. 예를 들어, 설명된 기술들이 설명된 방법과 다른 순서로 수행되거나, 및/또는 설명된 시스템, 구조, 장치, 회로 등의 구성요소들이 설명된 방법과 다른 형태로 결합 또는 조합되거나, 다른 구성요소 또는 균등물에 의하여 대치되거나 치환되더라도 적절한 결과가 달성될 수 있다.
그러므로, 다른 구현들, 다른 실시예들 및 특허청구범위와 균등한 것들도 후술하는 특허청구범위의 범위에 속한다.
Claims (15)
- 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지하는 단계;로그 재생 페이지에 대한 접속 명령이 발생 함에 따라,상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인하는 단계; 및상기 확인된 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템으로부터 획득하는 단계를 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 복수의 분할 로그파일을 이용하여, 공정 별 이벤트 발생 정도를 나타내는 상기 로그 재생 페이지를 구현하는 단계로서, 분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그에 대한 렌더링을 통해, 상기 로그 재생 페이지를 구현하는 단계를 더 포함하는 대용량 이벤트 로그 재생 방법.
- 제2항에 있어서,상기 로그 재생 페이지를 구현하는 단계는,REST_API 기반의 웹 어플리케이션을 통해, 상기 복수의 분할 로그파일을 처리하여, 상기 로그 재생 페이지를 시각화 하는 단계를 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 공정의 종류 또는 상기 로그파일의 용량에 따라, 분할 로그파일이 갖는 크기를 정하는 단계; 및정해진 상기 크기에 비례하여, 상기 스토리지 시스템으로부터의, 상기 분할 로그파일 각각의 획득 간격을 조정하는 단계를 더 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 스토리지 시스템으로 분할 명령을 전송하는 단계; 및상기 분할 명령에 따라, 상기 스토리지 시스템에서, 상기 로그파일을 파티션 단위로 분할하여 상기 복수의 분할 로그파일을 생성하는 단계를 더 포함하는 대용량 이벤트 로그 재생 방법.
- 제5항에 있어서,상기 복수의 분할 로그파일을 생성하는 단계는,상기 분할 명령에 의해 지정되는, 파티션 당 이벤트 로그의 개수 및 파티션의 개수 중 적어도 하나를 고려하여, 상기 로그파일을 분할하는 단계를 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 스토리지 시스템으로 분할 명령을 전송하는 단계; 및상기 분할 명령에 포함되는 토큰에 의해 식별되는 특정의 분할 로그파일을 상기 스토리지 시스템으로부터 획득하는 단계를 더 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 획득하는 단계는,상기 로그파일 내에 위치하는 순서에 상응하여, 상기 복수의 분할 로그파일 각각을 순차적으로 획득하는 단계를 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 이벤트 로그는, 발생시점에 대응하여 상기 로그파일에 기록되고,상기 획득하는 단계는,상기 복수의 분할 로그파일 중 제1 분할 로그파일에 포함되는 일정 개수의 이벤트 로그가 처리되는 동안, 상기 로그파일 내, 상기 제1 분할 로그파일 다음에 위치하는 제2 분할 로그파일을, 상기 스토리지 시스템으로부터 획득하는 단계를 포함하는 대용량 이벤트 로그 재생 방법.
- 제1항에 있어서,상기 스토리지 시스템에 상기 접속 명령과 관련된 로그파일이 유지되지 않는 경우,분할 처리 플랫폼에서의 공정 수행에 따른 이벤트 로그의 발생을 대기하는 단계를 더 포함하는 대용량 이벤트 로그 재생 방법.
- 공정과, 상기 공정에 연이어 수행되는 차기 공정 전에 발생하는 이벤트 로그를 카운트하여, 로그파일로서 스토리지 시스템에 유지하는 기록부;로그 재생 페이지에 대한 접속 명령이 발생 함에 따라,상기 접속 명령에 포함되는 공정에 대응되는 로그파일을, 상기 스토리지 시스템에서 확인하는 확인부; 및상기 확인된 로그파일을 정해진 크기로 분할하여, 복수의 분할 로그파일을 생성하여, 상기 스토리지 시스템으로부터 획득하는 획득부를 포함하는 대용량 이벤트 로그 재생 시스템.
- 제11항에 있어서,분할 로그파일 각각에 포함되는 일정 개수의 이벤트 로그에 대한 렌더링을 통해, 공정 별 이벤트 발생 정도를 나타내는 로그 재생 페이지를 구현하는 구현부를 더 포함하는 대용량 이벤트 로그 재생 시스템.
- 제11항에 있어서,상기 획득부는,상기 공정의 종류 또는 상기 로그파일의 용량에 따라, 분할 로그파일이 갖는 크기를 정하고, 정해진 상기 크기에 비례하여, 상기 스토리지 시스템으로부터의, 상기 분할 로그파일 각각의 획득 간격을 조정하는대용량 이벤트 로그 재생 시스템.
- 제11항에 있어서,상기 획득부는,상기 스토리지 시스템으로 분할 명령을 전송하고,상기 분할 명령에 따라, 상기 스토리지 시스템에서, 상기 로그파일을 파티션 단위로 분할하여 상기 복수의 분할 로그파일을 생성하도록 하는대용량 이벤트 로그 재생 시스템.
- 제11항에 있어서,상기 이벤트 로그는, 발생시점에 대응하여 상기 로그파일에 기록되고,상기 획득부는,상기 복수의 분할 로그파일 중 제1 분할 로그파일에 포함되는 일정 개수의 이벤트 로그가 처리되는 동안, 상기 로그파일 내, 상기 제1 분할 로그파일 다음에 위치하는 제2 분할 로그파일을, 상기 스토리지 시스템으로부터 획득하는대용량 이벤트 로그 재생 시스템.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/947,734 US10146659B2 (en) | 2016-12-23 | 2018-04-06 | Large event log replay method and system |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2016-0177345 | 2016-12-23 | ||
| KR1020160177345A KR101914347B1 (ko) | 2016-12-23 | 2016-12-23 | 대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US15/947,734 Continuation US10146659B2 (en) | 2016-12-23 | 2018-04-06 | Large event log replay method and system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018117453A1 true WO2018117453A1 (ko) | 2018-06-28 |
Family
ID=62626677
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2017/013633 Ceased WO2018117453A1 (ko) | 2016-12-23 | 2017-11-28 | 대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US10146659B2 (ko) |
| KR (1) | KR101914347B1 (ko) |
| WO (1) | WO2018117453A1 (ko) |
Families Citing this family (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102007126B1 (ko) * | 2018-12-27 | 2019-08-02 | 부산대학교 산학협력단 | 결손된 운영 데이터의 복원 방법 및 복원 장치 |
| US12190142B2 (en) * | 2019-12-30 | 2025-01-07 | UiPath, Inc. | Visual conformance checking of processes |
| GB2596502B (en) * | 2020-01-06 | 2023-01-04 | British Telecomm | Crypto-jacking detection |
| CN111639059A (zh) * | 2020-05-28 | 2020-09-08 | 深圳壹账通智能科技有限公司 | 日志信息的存储及定位方法、电子设备及存储介质 |
| US11586532B2 (en) * | 2021-07-02 | 2023-02-21 | Grammatech, Inc. | Software test environment with automated configurable harness capabilities, and associated methods and computer readable storage media |
| CN116561138A (zh) * | 2022-01-28 | 2023-08-08 | 马上消费金融股份有限公司 | 数据处理方法和装置 |
| CN117194566B (zh) * | 2023-08-21 | 2024-04-19 | 泽拓科技(深圳)有限责任公司 | 多存储引擎数据复制方法、系统、计算机设备 |
| CN118733384B (zh) * | 2024-06-25 | 2025-01-28 | 北京科杰科技有限公司 | 一种日志管理方法 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003033958A (ja) * | 2001-07-26 | 2003-02-04 | Sumitomo Heavy Ind Ltd | 射出成形機の過去の動作状態の記憶・出力方法 |
| KR20070015871A (ko) * | 2005-08-01 | 2007-02-06 | 소니 가부시끼 가이샤 | 정보 처리장치, 콘텐츠 재생장치, 정보 처리방법, 이벤트로그 기록 방법, 및 컴퓨터 프로그램 |
| KR20100027836A (ko) * | 2008-09-03 | 2010-03-11 | 충남대학교산학협력단 | 룰기반 웹아이디에스 시스템용 웹로그 전처리방법 및 시스템 |
| KR20150063233A (ko) * | 2013-11-29 | 2015-06-09 | 건국대학교 산학협력단 | 로그 데이터 처리 방법 및 이를 수행하는 시스템 |
| KR101556541B1 (ko) * | 2014-10-17 | 2015-10-02 | 한국과학기술정보연구원 | 고부하 경로 기반의 복합 이벤트 처리 장치 및 그 방법 |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100528001B1 (ko) | 2002-12-20 | 2005-11-09 | 삼성에스디에스 주식회사 | Mpeg 파일의 분산 관리 시스템 및 방법 |
| KR100912868B1 (ko) | 2006-08-29 | 2009-08-19 | 삼성전자주식회사 | 서비스 분산을 위한 장치 및 방법 |
| CN101192227B (zh) * | 2006-11-30 | 2011-05-25 | 阿里巴巴集团控股有限公司 | 一种基于分布式计算网络的日志文件分析方法和系统 |
| CN101582064B (zh) * | 2008-05-15 | 2011-12-21 | 阿里巴巴集团控股有限公司 | 一种大数据量数据处理方法及系统 |
| US8666818B2 (en) * | 2011-08-15 | 2014-03-04 | Logobar Innovations, Llc | Progress bar is advertisement |
| KR101341441B1 (ko) | 2011-11-04 | 2013-12-13 | 방한민 | 멀티미디어 콘텐츠 분할 및 분산 방법 |
| US9069704B2 (en) * | 2011-11-07 | 2015-06-30 | Sap Se | Database log replay parallelization |
| US8880479B2 (en) * | 2011-12-29 | 2014-11-04 | Bmc Software, Inc. | Database recovery progress report |
| US8977909B2 (en) * | 2012-07-19 | 2015-03-10 | Dell Products L.P. | Large log file diagnostics system |
| US9842053B2 (en) * | 2013-03-15 | 2017-12-12 | Sandisk Technologies Llc | Systems and methods for persistent cache logging |
| US9098453B2 (en) * | 2013-07-11 | 2015-08-04 | International Business Machines Corporation | Speculative recovery using storage snapshot in a clustered database |
| US9734021B1 (en) * | 2014-08-18 | 2017-08-15 | Amazon Technologies, Inc. | Visualizing restoration operation granularity for a database |
-
2016
- 2016-12-23 KR KR1020160177345A patent/KR101914347B1/ko not_active Expired - Fee Related
-
2017
- 2017-11-28 WO PCT/KR2017/013633 patent/WO2018117453A1/ko not_active Ceased
-
2018
- 2018-04-06 US US15/947,734 patent/US10146659B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003033958A (ja) * | 2001-07-26 | 2003-02-04 | Sumitomo Heavy Ind Ltd | 射出成形機の過去の動作状態の記憶・出力方法 |
| KR20070015871A (ko) * | 2005-08-01 | 2007-02-06 | 소니 가부시끼 가이샤 | 정보 처리장치, 콘텐츠 재생장치, 정보 처리방법, 이벤트로그 기록 방법, 및 컴퓨터 프로그램 |
| KR20100027836A (ko) * | 2008-09-03 | 2010-03-11 | 충남대학교산학협력단 | 룰기반 웹아이디에스 시스템용 웹로그 전처리방법 및 시스템 |
| KR20150063233A (ko) * | 2013-11-29 | 2015-06-09 | 건국대학교 산학협력단 | 로그 데이터 처리 방법 및 이를 수행하는 시스템 |
| KR101556541B1 (ko) * | 2014-10-17 | 2015-10-02 | 한국과학기술정보연구원 | 고부하 경로 기반의 복합 이벤트 처리 장치 및 그 방법 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20180225189A1 (en) | 2018-08-09 |
| KR20180073861A (ko) | 2018-07-03 |
| US10146659B2 (en) | 2018-12-04 |
| KR101914347B1 (ko) | 2018-11-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2018117453A1 (ko) | 대용량 이벤트 로그 재생 방법 및 대용량 이벤트 로그 재생 시스템 | |
| Fan et al. | Microservices vs Serverless: A Performance Comparison on a Cloud-native Web Application. | |
| US11789846B2 (en) | Method and system for using stacktrace signatures for bug triaging in a microservice architecture | |
| US11288245B2 (en) | Telemetry definition system | |
| US20110172963A1 (en) | Methods and Apparatus for Predicting the Performance of a Multi-Tier Computer Software System | |
| CN111796809A (zh) | 接口文档生成方法、装置、电子设备及介质 | |
| JP2006031109A (ja) | 管理システム及び管理方法 | |
| WO2023140642A1 (ko) | 라이브 커머스 플랫폼에서의 실시간 인스펙터를 위한 방법, 컴퓨터 장치, 및 컴퓨터 프로그램 | |
| US20140365614A1 (en) | Monitoring similar data in stream computing | |
| CN110719215A (zh) | 虚拟网络的流信息采集方法及装置 | |
| CN113230661A (zh) | 数据同步方法、装置、计算机可读介质及电子设备 | |
| Gonzalez et al. | Net2vec: Deep learning for the network | |
| CN113806416B (zh) | 实时数据服务的实现方法、装置及电子设备 | |
| CN103514044B (zh) | 一种动态行为分析系统的资源优化方法、装置和系统 | |
| US11037297B2 (en) | Image analysis method and device | |
| US8140318B2 (en) | Method and system for generating application simulations | |
| CN118449752A (zh) | 一种基于请求的无服务器的攻击溯源方法及系统 | |
| US20040111706A1 (en) | Analysis of latencies in a multi-node system | |
| CN114175067A (zh) | 安全事故调查工作空间生成和调查控制 | |
| US11921603B2 (en) | Automated interoperational tracking in computing systems | |
| CN119995938B (zh) | 自动化细粒度日志标注方法 | |
| WO2021112308A1 (ko) | 채용 면접 서비스를 제공하는 방법 및 서버 | |
| US20210224417A1 (en) | Non-transitory computer-readable recording medium having stored therein information processing program, information processing method, and information processing apparatus | |
| US20210224408A1 (en) | Non-transitory computer-readable recording medium having stored therein screen displaying program, method for screen displaying, and screen displaying apparatus | |
| CN109086125B (zh) | 图片分析方法、装置及系统、计算机设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17882315 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17882315 Country of ref document: EP Kind code of ref document: A1 |