WO2020111489A1 - 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 - Google Patents
주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 Download PDFInfo
- Publication number
- WO2020111489A1 WO2020111489A1 PCT/KR2019/012737 KR2019012737W WO2020111489A1 WO 2020111489 A1 WO2020111489 A1 WO 2020111489A1 KR 2019012737 W KR2019012737 W KR 2019012737W WO 2020111489 A1 WO2020111489 A1 WO 2020111489A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- reliability
- document
- sentence
- distribution
- user
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/3332—Query translation
- G06F16/3335—Syntactic pre-processing, e.g. stopword elimination, stemming
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
Definitions
- the present invention relates to a technique for enhancing the reliability of a document, and more specifically, to perform an automatic reliability distribution prediction for a possible candidate candidate for a given document, and to see the mean and standard deviation of each of the original document and the candidate candidate reliability distribution It relates to a method and a system capable of performing reliability enhancement based on a change.
- the technology of one embodiment of the automatic document proofing system researches and develops an automatic document proofing system using a learning model utilizing neural machine translation (NMT), and the corresponding system is a CoNLL for automatic grammar proofing. -It was experimentally proved that it showed a high performance of F0.5 score corresponding to about 40% for the 2014 Shared task test set.
- the technique of another embodiment of the automatic document proofing system explores an evaluation index that can best match the verbal judgment of a native language user in evaluating the automatic grammar proofing method and system, and evaluates the existing automatic grammar proofing method and system.
- We proposed the GLEU index which is a variation of the BLEU index that has been utilized for the purpose.
- the technology of another embodiment of the automatic document proofing system has researched and developed ERRANT, an annotation assist tool for effectively constructing the training data and evaluation data required to build an automatic document proofing system, and another example of the automatic document proofing system.
- the technology of one embodiment built a large learning assessment data corpus for training and evaluating an automatic deep correction system for technical documents.
- Another embodiment technique for an automatic document proofing system is a neural sequence-to-sequence trained based on the EF-Cambridge Open Language Database (EFCAMDAT), a large learning assessment data corpus for automatic proofreading of grammar.
- EFCAMDAT EF-Cambridge Open Language Database
- Google's G-Suite recently added automated grammatical error correction proposals to services such as Google Docs and Gmail.
- Another typical commercialization example is an add-on program for the Google Docs service, which has researched and developed a program that automatically provides feedback on document persuasiveness, topic development, consistency, grammatical errors, and word selection.
- the persuasiveness, consistency, etc. of documents analyzed as related elements in the existing methodology for automatic document proofreading are about the characteristics of the sentence arrangement and components in the document, and the public consensus rate is high. high.
- Reliability for a given sentence is defined in relation to the fact that the specific event described by the given sentence actually occurred, or in relation to whether a particular speaker/content is spoken by a specific speaker. Reliability studies conducted in relation to the fact that a specific event actually occurred confirmed that the rate of consensus was relatively high.
- the public consensus rate that a particular document has high/low persuasiveness, that it has high/low consistency, and that it has high/low sentiment strength is that certain documents contain reliable/non-confidential content. It is higher than the public consensus rate.
- Persuasion/consistency is about the structure of the argument, and in the case of a document with a clear argument structure, it can be evaluated that the corresponding document has high persuasion/consistency even if there is a disagreement with its content.
- the intensity of emotions is also limited to emotion-inducing words and emotion-inducing expressions, a consensus rate for evaluation that a document that uses a lot of corresponding words and expressions has a high emotional intensity may be high.
- reliability is about an individual's subjective beliefs, what beliefs about the content are the most important factors in determining reliability, which can be greatly influenced by demographic characteristics and demographic characteristics.
- the rate of public consensus on reliability evaluation is relatively low because it can be influenced by the subjectivity of individuals who cannot be defined as.
- automatic reliability distribution prediction for possible candidate candidates for a given document may be performed, and reliability enhancement may be performed based on a change in the mean and standard deviation in which the reliability distributions of the original document and the candidate candidates are visible.
- the document reliability enhancement method includes: when a document is input from a user, generating a candidate group for correcting each sentence of the input document; Predicting a reliability distribution for each of the input document and the generated correction candidate group; And performing reliability enhancement for the input document through a plurality of augmentation methods based on the average and standard deviation of the predicted reliability distribution.
- the document reliability enhancement method further includes providing the user with the reliability enhancement document for the input document and the reliability distribution change information for before and after the reliability enhancement for the user. Can be.
- the step of providing to the user may further provide the user with a document correction result most similar to the aspect of the reliability distribution desired by the user through interaction with the user.
- the plurality of augmentation schemes is a first item corresponding to a reliability enhancement of a scheme that directs an average of more than a preset reference average of the reliability distribution, a reliability enhancement of a scheme that directs a standard deviation below a preset reference standard deviation of the reliability distribution. It may include at least one of the corresponding second item and at least one of the third items corresponding to the enhancement of reliability in a manner of oriented ⁇ average/standard deviation ⁇ above a predetermined reference ⁇ average/standard deviation ⁇ of the reliability distribution.
- the generating step includes receiving control information for the plurality of augmentation methods from the user; Pre-processing each sentence of the input document; And generating a plurality of sets of corrected candidate sentences for each sentence of the input document, wherein the predicting step is a plurality of combinations of each sentence of the input document and each sentence in the corrected candidate sentence set.
- the reliability distribution for each sentence in the document set is predicted, and the step of performing is performed by selecting a reliability enhancement document based on the reliability distribution for each sentence in the predicted plurality of document sets and the control information, thereby inputting the input. It is possible to increase the reliability of documents.
- Document reliability enhancement system when the user inputs a document generating unit for generating a candidate group for correction for each sentence of the input document; A prediction unit for predicting a reliability distribution for each of the input document and the generated correction candidate group; And an augmentation unit that increases the reliability of the input document through a plurality of augmentation methods based on the average and standard deviation of the predicted reliability distribution.
- the document reliability enhancement system further includes an output unit for providing the user with the reliability enhancement enhancement document for the input document and the reliability distribution change information for before and after the reliability enhancement for the user. Can be.
- the output unit may further provide the user with a document correction result that is most similar to the aspect of the reliability distribution desired by the user through interaction with the user.
- the plurality of augmentation schemes is a first item corresponding to a reliability enhancement of a scheme that directs an average of more than a preset reference average of the reliability distribution, a reliability enhancement of a scheme that directs a standard deviation below a preset reference standard deviation of the reliability distribution. It may include at least one of the corresponding second item and at least one of the third items corresponding to the enhancement of reliability in a manner of oriented ⁇ average/standard deviation ⁇ above a predetermined reference ⁇ average/standard deviation ⁇ of the reliability distribution.
- the generating unit receives control information for the multiple augmentation methods from the user, performs pre-processing for each sentence of the input document, generates a plurality of candidate candidate sentence sets for each sentence of the input document, ,
- the prediction unit predicts a reliability distribution for each sentence of a plurality of document sets that can be combined through each sentence of the input document and each sentence in the corrected candidate sentence set, and the augmentation unit of the predicted plurality of document sets Reliability enhancement for the input document may be performed by selecting a reliability enhancement document based on the reliability distribution for each sentence and the control information.
- automatic reliability distribution prediction of possible candidate candidates for a given document is performed, and reliability enhancement is performed based on a change in the mean and standard deviation of each of the original documents and the reliability distribution of the candidate candidates. Can be.
- the present invention is to improve the reliability of a method in which a high average of the reliability distribution is oriented, and to improve the reliability of the document to be trusted by more people. Reliability is improved so that there is no more controversy about whether or not to trust the documents to be targeted for reliability enhancement by proceeding with the enhancement of reliability in a way that aims for low standard deviation, and (3) aiming for high ⁇ average/standard deviation ⁇ of the reliability distribution It is possible to improve the reliability of a method in which a document that is a target of reliability enhancement is combined with the two methods.
- a user inputs an aspect of the desired reliability distribution through a drag and drop method and shows a reliability distribution aspect most similar to that of the corresponding reliability distribution.
- a sentence correction candidate can be provided to the user. Therefore, the present invention enables document correction for reliability enhancement, which is difficult to be applied to an automated correction system, because it is greatly influenced by the subjectivity of an individual, through the introduction of a distribution concept, thereby enabling a document to be more trusted by an unspecified group of others. Detailed information on how to be modified can be provided in various forms.
- the present invention may perform automatic proofreading of public documents, news, and technical documents in which the reader's reliability is important, or may produce documents with high reliability based on personal subjectivity.
- FIG. 1 shows a configuration for a document reliability enhancement system according to an embodiment of the present invention.
- FIG. 2 shows a configuration of an embodiment of the reliability enhancement unit illustrated in FIG. 1.
- FIG. 3 illustrates an exemplary configuration of the output unit illustrated in FIG. 1.
- Figure 4 shows an exemplary view for explaining the output result of the document reliability enhancement system of the present invention.
- FIG. 5 shows an example of a sentence correction candidate corpus included in the document reliability enhancement system.
- FIG. 6 shows an example of a reliability distribution corpus included in the document reliability enhancement system.
- FIG. 7 is a flowchart illustrating an operation of an embodiment by the reliability enhancement unit illustrated in FIG. 1.
- FIG. 8 is a flowchart illustrating an operation of an embodiment by the output unit illustrated in FIG. 1.
- Embodiments of the present invention perform automatic reliability distribution prediction for possible candidate candidates for a given document, and enhance reliability based on changes in the mean and standard deviation of each of the original documents and the candidate distributions of the candidate candidates.
- the present invention is the first item corresponding to the enhancement of the reliability of the method that aims for a high average of the reliability distribution
- the second item corresponding to the enhancement of the reliability of the method of directing the low standard deviation of the reliability distribution
- the high ⁇ average of the reliability distribution Reliability enhancement may be performed according to a plurality of augmentation methods based on an average and a standard deviation of a reliability distribution including at least one item of a third item corresponding to a reliability enhancement of a method for /standard deviation ⁇ .
- the present invention provides a statistical index related to the augmented document to the user in a graph form after increasing the reliability of the document input by the user, and a document input by the user by an unspecified group of others through interaction with the user Preliminary monitoring can be performed as to what reliability change will occur in accordance with the fine tuning of.
- FIG. 1 shows a configuration for a document reliability enhancement system according to an embodiment of the present invention.
- the document reliability enhancement system 100 includes a reliability enhancement unit 110, a sentence correction candidate corpus 120, a reliability distribution corpus 130, and an output unit 140 ) And the control unit 150.
- the document reliability enhancement system 100 of the present invention performs reliability enhancement for each sentence of a document (input text) received from a user, and provides a document with enhanced reliability to a user.
- an original reliability distribution corresponding to before/after reliability enhancement can be output to the user for each sentence.
- the reliability enhancement unit 110 performs reliability enhancement for the received document according to a plurality of augmentation methods based on the average and standard deviation of the reliability distribution.
- the multiple augmentation method based on the mean and standard deviation of the reliability distribution is a high average of the reliability distribution, for example, the first item corresponding to the enhancement of the reliability of the method that orients the average above the preset reference average, the low standard of the reliability distribution Deviation, for example, the second item corresponding to the enhancement of the reliability in the way to direct the standard deviation below the preset reference standard deviation and the high ⁇ average/standard deviation ⁇ of the reliability distribution, for example, the preset reference ⁇ average/standard deviation ⁇ It may include at least one or more items from the third item corresponding to enhancement of reliability in the manner of oriented ⁇ average/standard deviation ⁇ .
- increasing the reliability of a method that aims for a high reliability average can be interpreted as editing a document so that more people can trust a document that is a target of increasing the reliability, and increasing the reliability of a method that aims for a low reliability standard deviation.
- the reliability enhancement unit 110 receives a document, which is an object of reliability enhancement, from the user as an input, and selectively outputs information corresponding to any one of a plurality of augmentation methods based on the mean and standard deviation of the reliability distribution from the user. It receives information on whether or not (hereinafter referred to as "control information"), transmits the received control information to the control unit 130, generates a candidate candidate group for each sentence in a given document, and returns control information to the control unit 130 ), the documents that match the control information are selected from documents corresponding to multiple augmentation methods based on the average and standard deviation of the reliability distribution among documents that can be generated by combining candidate candidates for correction.
- control information information on whether or not
- the documents corresponding to the multiple augmentation methods based on the mean and standard deviation of the reliability distribution are (1) the document after the sharpening with the highest mean of the reliability distribution predicted for the entire document, and (2) the document predicted for the overall document. It may include a post-adjustment document with the lowest standard deviation of the reliability distribution, and (3) a post-adjustment document with the highest ⁇ average/standard deviation ⁇ value of the predicted reliability distribution for the entire document.
- the reliability enhancement unit 110 may output information corresponding to both the average and standard deviation of the reliability distribution as control information.
- the sentence correction candidate corpus 120 is a sentence in which the sentence structure is modified by the language expert to manually change the sentence structure while maintaining the meaning of the sentence (hereinafter referred to as a “sentence after sentence”), the original sentence, and a language expert By storing the difference between the meaning of the original sentence and the sentence after correction, it is judged as a score within the range of 0 points (that is, there is no difference in meaning) to 5 points (that is, large difference in meaning).
- the sentence correction candidate corpus 120 may be used as a learning criterion for supervised learning of a corresponding model when the reliability enhancer 110 uses a model for generating a candidate candidate group for each sentence in a given document. Can be.
- the sentence correction candidate corpus stores a plurality of example sentence sets and a plurality of sentence correction sentences for each example sentence, For each corrected sentence, the similarity to the original sentence is stored before the corrected sentence.
- the corpus 130 stores the distribution of the reader's reliability for each sentence collected through a direct reliability questionnaire.
- the reliability distribution corpus 130 is used as a learning criterion for supervised learning the corresponding model when the reliability enhancer 110 includes a reliability distribution prediction model in predicting the reliability distribution of the document after the correction. Can be.
- FIG. 6 shows an example of the reliability distribution corpus included in the document reliability enhancement system, and as shown in FIG. 6, the reliability distribution corpus shows the distribution of reliability that the reader groups for each document set and each document actually showed through a questionnaire. To save.
- the plurality of document sets stored by the reliability distribution corpus 130 should be able to include various types of documents on various topics encountered in everyday life.
- the subject of the document can include life, health, politics, policy, economy, environment, and the types of documents include SNS posts, blog posts, online news, online forum posts, research papers, and books. It can contain.
- the subject and type of each of the documents subject to direct questioning across the corpus should not be uniform. This is because the reliability distribution corpus 130 collected through a questionnaire in the system according to an embodiment of the present invention is used as a learning criterion for automatically predicting the reliability distribution.
- the predictive model learned through the corresponding prediction criteria corpus is the reliability of the document on politics/policy. It will be inadequate to predict the distribution, and you can expect the prediction results to differ from the distribution of confidence that real readers will see.
- the output unit 140 outputs the document with the enhanced reliability selected by the reliability enhancement unit 110 to the user.
- the output unit 140 parallelizes the predicted reliability distribution for each sentence in the original document input by the user and the predicted reliability distribution for each sentence in the document enhanced by the reliability enhancer 110 in a graph form. Compare and provide output to the user by comparison. For example, the output unit 140 compares the predicted reliability distribution for each sentence in the original document and the predicted reliability distribution for each sentence in the augmented document in parallel, as shown in the output result shown in FIG. Output.
- the output unit 140 provides an output of the user's confidence distribution corresponding to before and after the confidence increase to each user for each sentence, and by providing user-interacting detailed information about the confidence enhancement step by step, the user Can perform preliminary monitoring of how the reliability response changes depending on the micro-deformation of a document input by a user.
- the output unit 140 may interact with the user through a drag-and-drop method for the displayed graph, and when the user inputs an aspect of the desired reliability distribution through the drag-and-drop method, the output is the closest to the received reliability distribution. It selects the candidates with a reliability distribution and provides the re-output of the sentence that is the object of the reliability distribution to the corresponding candidate, and displays the graph that the user has previously interacted with by drag and drop with the reliability distribution of the candidate. Changes can be reprinted and provided.
- the controller 150 determines which kind of reliability enhancement is to be performed among a plurality of augmentation methods based on the mean and standard deviation of the reliability distribution.
- control unit 150 may determine which type of reliability enhancement to perform among the multiple enhancement methods based on the average and standard deviation of the reliability distribution based on the control information received by the reliability enhancement unit 110.
- FIG. 2 shows a configuration of an embodiment of the reliability enhancement unit illustrated in FIG. 1.
- the reliability enhancement unit 110 includes an input unit 111, a pre-processing unit 112, a correction candidate generation unit 113, a reliability distribution prediction unit 114, and an enhancement document selection unit 115. Includes.
- the input unit 111 receives a document, which is an object of enhancement of reliability, from the user as an input, and optionally receives control information from the user and transmits it to the control unit 130.
- the input unit 111 may provide information corresponding to both the average of the reliability distribution and the multiple augmentation methods based on the standard deviation as control information.
- control information when the control information is not input from the user, the control information may output information corresponding to both a plurality of augmentation methods based on the mean and standard deviation of the reliability distribution as control information.
- the pre-processing unit 112 receives the document input from the user from the input unit 111, and performs semantic role labeling, dependency parsing, and discourse parsing.
- semantic domain analysis can be performed through a semantic extractor such as DeepSemanticRoleLabeling or PathLSTM Semantic Role Labeler
- dependency analysis can be performed through a parser such as StanfordCoreNLP
- discourse structure analysis is a discourse structure such as PDTB-style discourse parser. It can be done through the analyzer.
- the candidate candidate generating unit 113 receives (1) a document input from a user from the pre-processing unit 112, (2) a semantic domain analysis result, (3) a dependency analysis result, and (4) a discourse structure analysis result. For each sentence in the original document, a plurality of candidate candidate sentences are generated.
- the correction candidate generation unit 113 may select L sentences having the smallest difference in meaning from the original sentence among the multiple correction candidate sentences that can be generated for each sentence and generate them as a set of multiple correction candidate sentences.
- the correction candidate generation unit 113 may generate a correction candidate sentence through a supervised learning-based correction sentence generation model, and the corresponding model may be supervised through the sentence correction candidate corpus 120.
- the learning model loads original sentences from a set of multiple sentences from the candidate candidate corpus 120 during learning, performs semantic analysis, dependency analysis, and discourse structure analysis.
- (1) Original sentences and (2) semantic analysis Results, (3) dependency analysis results, (4) discourse structure analysis results are used as input criteria, and the corresponding differences between each original sentence and the meaning difference between the original sentence and the sentence after the correction are called and used as an output criterion.
- semantic domain analysis, dependency analysis, and discourse structure analysis in the addition candidate generation unit 113 may also be performed in the same manner as that performed by the pre-processing unit 112.
- the reliability distribution prediction unit 114 includes (1) sentences in the document input from the user from the correction candidate generation unit 113, (2) semantic domain analysis results, (3) dependency analysis results, and (4) discourse structure. As a result of the analysis, (5) a plurality of sets of candidate sentences are received or transmitted, and each sentence in a set of multiple documents that can be combined is transmitted through a plurality of sets of sentences in a document received from a user and corresponding sentences in a set of multiple candidates for correction. Predict the distribution of confidence that a group of readers would expect to see.
- the set of multiple documents that can be combined can be made by traversing the random selection and replacement method described above M times to generate M corrected candidate documents.
- M is a natural number set as the initial value of the system, and M may be set in consideration of the calculation speed of the system and user satisfaction.
- the reliability distribution prediction unit 114 predicts the reliability distribution for each sentence by using a model trained through supervised learning from the reliability distribution corpus 130.
- the learning model when each sentence in a plurality of document sets stored in the reliability distribution corpus 130 during learning is used as a prediction target for the reliability distribution, for example, a predetermined number of sentences appearing in front of the sentence for the reliability distribution prediction, for example, N sentences, Based on the semantic analysis result, dependency analysis result, and discourse structure analysis result for the 2N+1 sentence sets that include both N sentences and the corresponding sentences, the actual questionnaire for the reliability distribution prediction sentence The reliability distribution may be called from the reliability distribution corpus 130 as an output criterion in the measured multiple sentence set.
- N is a natural number that is set as the system initial value, and if there are fewer than N sentences before or after the sentence to predict the reliability distribution in a given document, the non-existing sentence is replaced with a null sentence. Can be used as an input criterion.
- semantic domain analysis, dependency analysis, and discourse structure analysis in the reliability distribution prediction unit 114 may be performed in the same manner as that performed by the pre-processing unit 113.
- the augmented document selection unit 115 shows a plurality of document sets that can be combined through each sentence in a plurality of corrected candidate sentence sets corresponding to a plurality of sentence sets in a document input from a user from the reliability distribution prediction unit 114 and a reader group for each Receiving the reliability distribution expected to be received, and receiving control information for a plurality of augmentation methods based on the average and standard deviation of the reliability distribution from the control unit 130, selects documents matching the control information among the preset documents as the reliability enhancement documents .
- the preset document includes (1) the document after the correction with the highest average value for all the sentences in the document that averages the predicted reliability distribution for each sentence, and (2) the predicted reliability distribution for each sentence.
- FIG. 3 illustrates an exemplary configuration of the output unit illustrated in FIG. 1.
- the output unit 140 includes an augmented text output unit 141, a graph output unit 142, and a user interaction unit 143.
- the augmented text output unit 141 receives the reliability enhancement document selected by the augmented document selection unit 115 and outputs it to a user.
- the graph output unit 142 receives the selected reliability enhancement document from the augmented document sorting unit 115, (1) the result of the reliability distribution prediction, (2) the original document, and (3) the reliability distribution prediction of the original document Receiving the result from the reliability prediction unit 114, as shown in the example shown in Figure 4, the predicted reliability distribution for each sentence in the original document input by the user and the document enhanced by the reliability enhancement unit 110 The predicted reliability distribution for each sentence in the graph is compared and compared in parallel to provide the output to the user.
- the user interaction unit 143 allows the user to interact with the user through a drag and drop method for the reliability distribution graph output to the user by the graph output unit 142, and the user desires the reliability distribution through the drag and drop method
- the correction candidate having the reliability distribution closest to the input reliability distribution is selected, and the sentence targeted for the reliability distribution is changed to the corresponding correction candidate and re-printed.
- Reliability distribution of candidates for correction is provided by re-printing the graph that the user had previously interacted with by drag and drop.
- the user interaction unit 143 may allow the user to input an aspect of the reliability distribution desired by the user through a drag and drop method in the reliability distribution graph output by the graph output unit 142.
- each of the correction candidates and correspondence among the set of correction candidates of the corresponding sentence The predicted reliability distribution is received from the reliability distribution prediction unit 114, and a correction candidate having a prediction reliability distribution having the highest similarity to the mode of the reliability distribution desired by the user is selected from the correction candidates through the distribution similarity index.
- the distribution similarity may be defined as Kullback-Leibler divergence or Levy-Prokhorov metric.
- the method according to an embodiment of the present invention performs automatic reliability distribution prediction for possible correction candidate groups for a given document, and increases reliability based on a change in the mean and standard deviation of each of the original document and the reliability distributions of the candidate candidates. You can do
- the method according to an embodiment of the present invention enables document correction for a reliability concept including a large individual difference because the requirement of introducing a distribution concept is satisfied. That is, the method according to the present invention includes (1) improving the reliability of a method in which a high average of the reliability distribution is oriented, and adhering the document so that more people can trust the document that is the target of the reliability enhancement, (2) Promoting reliability enhancement in a way that aims for a low standard deviation of the distribution of reliability, so that there is no more controversy about whether or not to trust the documents that are the target of the enhancement, (3) high ⁇ average/standard deviation of the distribution of reliability By proceeding with the enhancement of the reliability of the method directed toward ⁇ , it is possible to add a document that is the object of the enhancement of the reliability in a combined manner of the two methods.
- the present invention enables document correction for reliability enhancement, which is difficult to be applied to an automated correction system, because it is greatly influenced by the subjectivity of an individual, through the introduction of a distribution concept, thereby enabling a document to be more trusted by an unspecified group of others.
- Detailed information on how to be modified can be provided in various forms.
- FIG. 7 is a flowchart illustrating an operation of an embodiment by the reliability enhancement unit illustrated in FIG. 1.
- the method by the reliability enhancement unit 110 includes an input step (S310), a pre-processing step (S320), a reliability distribution prediction step (S330), a correction candidate generation step (S340), and an evidence document selection step (S350). ).
- the input step (S310) is a step of receiving a document, which is a target for increasing reliability, from a user as input, and selectively receiving control information for a plurality of augmentation methods based on a mean and standard deviation of the reliability distribution from the user.
- the pre-processing step S320 is a step of pre-processing each sentence in the document received from the user.
- the step of generating a candidate candidate for correction (S330) is a step of generating a set of multiple candidate candidates for each sentence in the original document.
- the reliability distribution prediction step (S340) is a prediction of the reliability distribution that the reader group is expected to see for each sentence in the multiple document set combinable through each sentence in the multiple document set in the document received from the user and the corresponding multiple correction candidate sentence set.
- Evidence document selection step is an analysis of the reliability distribution that is expected to be seen by the reader group for each of the multiple document sets that can be combined through each sentence in the multiple-sentence candidate sentence set corresponding to the multiple sentence sets in the document received from the user. This is a step of selecting a document for enhancing reliability according to control information.
- FIG. 8 is a flowchart illustrating an operation of an embodiment by the output unit illustrated in FIG. 1.
- the method by the output unit 140 includes an augmented text output step (S410 ), a graph output step (S420 ), and a user interaction step (S430 ).
- the augmented text output step (S410) is a step of providing and outputting a reliability-enhanced document to a user.
- the graph output step (S420) is a step in which the reliability distribution predicted for each sentence in the original document input by the user and the predicted reliability distribution for each sentence in the augmented document are compared and compared in parallel in a graph form to provide the output to the user to be.
- the user interaction step (S350) enables the user to interact with the user by drag and drop method for the reliability distribution graph output to the user, and is received when the user inputs the desired reliability distribution mode through the drag and drop method.
- the candidates with the closest confidence distribution are selected, and the sentence that is the object of the confidence distribution is changed and re-printed with the corresponding candidate, and the reliability distribution of the corresponding candidate is mutually recognized by dragging and dropping. This is a step to reprint the changed graph and provide it.
- each step constituting FIGS. 7 and 8 may include all the contents described in FIGS. 1 to 6, which is to those skilled in the art.
- the system or device described above may be implemented with hardware components, software components, and/or combinations of hardware components and software components.
- the systems, devices, and components described in the embodiments include, for example, processors, controllers, arithmetic logic units (ALUs), digital signal processors (micro signal processors), microcomputers, and field programmable arrays (FPAs). ), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions, may be implemented using one or more general purpose computers or special purpose computers.
- the processing device may run an operating system (OS) and one or more software applications running on the operating system.
- the processing device may access, store, manipulate, process, and generate data in response to the execution of the software.
- OS operating system
- the processing device may access, store, manipulate, process, and generate data in response to the execution of the software.
- a processing device may be described as one being used, but a person having ordinary skill in the art, the processing device may include a plurality of processing elements and/or a plurality of types of processing elements. It can be seen that may include.
- the processing device may include a plurality of processors or a processor and a controller.
- other processing configurations such as parallel processors, are possible.
- the software may include a computer program, code, instruction, or a combination of one or more of these, and configure the processing device to operate as desired, or process independently or collectively You can command the device.
- Software and/or data may be interpreted by a processing device, or to provide instructions or data to a processing device, of any type of machine, component, physical device, virtual equipment, computer storage medium or device. , Or may be permanently or temporarily embodied in the transmitted signal wave.
- the software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored in one or more computer-readable recording media.
- the method according to the embodiments may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer readable medium.
- the computer-readable medium may include program instructions, data files, data structures, or the like alone or in combination.
- the program instructions recorded in the medium may be specially designed and configured for the embodiments or may be known and usable by those skilled in computer software.
- Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs, DVDs, and magnetic media such as floptical disks.
- -Hardware devices specifically configured to store and execute program instructions such as magneto-optical media, and ROM, RAM, flash memory, and the like.
- program instructions include high-level language code that can be executed by a computer using an interpreter, etc., as well as machine language codes produced by a compiler.
- the hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Databases & Information Systems (AREA)
- Probability & Statistics with Applications (AREA)
- Data Mining & Analysis (AREA)
- Machine Translation (AREA)
Abstract
주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템이 개시된다. 본 발명의 일 실시예에 따른 문서 신뢰도 증강 방법은 사용자로부터 문서가 입력되면 상기 입력된 문서의 각 문장에 대한 첨삭 후보군을 생성하는 단계; 상기 입력된 문서와 상기 생성된 첨삭 후보군 각각에 대한 신뢰도 분포를 예측하는 단계; 및 상기 예측된 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식을 통해 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 단계를 포함한다.
Description
본 발명은 문서의 신뢰도를 증강시키는 기술에 관한 것으로, 보다 구체적으로는 주어진 문서에 대해 가능한 첨삭 후보군에 대한 자동 신뢰도 분포 예측을 수행하고, 원본 문서와 첨삭 후보군의 신뢰도 분포 각각이 보이는 평균과 표준편차 변화에 기반하여 신뢰도 증강을 수행할 수 있는 방법 및 그 시스템에 관한 것이다.
자연언어처리 기술이 발전하고 의미론적/화용론적 특징에 대한 자동 분석 기술의 성능이 향상함에 따라 자동 문서 교정 시스템에 대한 연구 기술 개발과 상용화가 활발히 진행되고 있다.
자동 문서 교정 시스템에 대한 일 실시예의 기술은 신경망 기반 기계 번역 기법(neural machine translation, NMT)을 활용한 학습 모델을 활용한 자동 문서 교정 시스템을 연구 개발하고, 해당하는 시스템이 문법 자동 교정에 대한 CoNLL-2014 Shared task test set에 대해 약 40%에 해당하는 F0.5 score의 높은 성능을 보임을 실험적으로 입증하였다. 자동 문서 교정 시스템에 대한 다른 일 실시예의 기술은 자동 문법 교정 방법 및 시스템의 평가에 있어 모국어 사용자의 언어적 판단과 가장 일치할 수 있는 평가 지표를 탐색하고, 기존에 자동 문법 교정 방법 및 시스템의 평가를 위해 활용되어온 BLEU 지표의 변형인 GLEU 지표를 제안하였다. 자동 문서 교정 시스템에 대한 또 다른 일 실시예의 기술은 자동문서 교정 시스템을 구축하기 위해 필요한 학습 데이터 및 평가 데이터를 효과적으로 구축하기 위한 주석 보조 도구인 ERRANT를 연구 개발하였으며, 자동 문서 교정 시스템에 대한 또 다른 일 실시예의 기술은 기술 문서에 대한 자동 심층 교정 시스템을 학습시키고 평가시키기 위한 대규모 학습 평가 데이터 코퍼스를 구축하였다. 자동 문서 교정 시스템에 대한 또 다른 일 실시예의 기술은 문법 자동 교정을 위한 대규모 학습 평가 데이터 코퍼스인 EF-Cambridge Open Language Database (EFCAMDAT)에 기반하여 학습시킨 뉴럴 시퀀스 대 시퀀스(neural sequence-to-sequence) 모델을 활용하여 다양한 주제의 문서에 대해 높은 성능을 보이는 자동 문법 교정 시스템을 연구 개발하였다.
자동 문서 교정 시스템에 대한 대표적인 상용화 사례로, Grammarly는 사용자가 입력한 문서에 대한 자동 문법 교정 및 어조 교정 서비스를 제공한다. 이와 관련하여 미국 특허 US9465793B2 (granted, 2016-10-11), "Systems and methods for advanced grammar checking"는 자동화된 문법 오류 교정 및 첨삭 제안 및 문법 오류 교정 제안 후보에 대한 우선 순위 설정 방법을 제안하였다. 또한, Grammarly는 최근, 단순 문법 교정 및 어조 교정뿐만 아니라, 문서의 의도(예를 들어, 새로운 정보 제공, 특정 주제에 대한 설명, 설득, 스토리텔링), 독자의 특징(예를 들어, 일반, 지식인, 전문가), 문서의 스타일(예를 들어, 격식 있는, 격식 없는), 문서의 감정 세기(예를 들어, 약함, 강함), 문서 도메인(예를 들어, 일반, 학술, 사업, 기술, 창의, 캐쥬얼)에 맞추어 문장 스타일에 대한 자동 첨삭 제안 기능을 추가하였다.
다른 대표적인 상업화 사례로, Google의 G-Suite는 최근 Google Docs, Gmail 등의 서비스에 자동화된 문법 오류 교정 제안 기능을 추가하였다. 또 다른 대표적인 상업화 사례는 Google Docs 서비스에 대한 추가 기능 프로그램으로서, 문서의 설득력, 주제 전개, 일관성, 문법 오류, 단어 선택 등에 대한 피드백을 자동으로 제공하는 프로그램을 연구 개발하였다.
자동 문서 교정을 위한 기존의 방법론은 대중적인 합의가 비교적 높은 개념들에 국한된 것으로, 신뢰도 개념과 같이 개인차가 큰 값에 대해서 적용되기 힘들다는 단점이 있다.
예를 들어, 자동 문서 교정을 위한 기존의 방법론에서 관련 요소로서 분석하는 문서의 설득력, 일관성 등은 문서 내의 문장 배열 및 구성 요소의 특징에 대한 것으로 대중적 합의율이 높으며, 감정 세기 역시 대중적 합의율이 높다.
주어진 문장에 대한 신뢰도는 주어진 문장이 서술하는 특정 사건이 실제로 일어났는지(factuality)와 관련하여 정의되거나, 특정 화자가 말하는 특정 주장/내용을 신뢰하는지와 관련되어 정의된다. 특정 사건이 실제로 일어났는지(factuality)와 관련하여 진행된 신뢰도 연구들은 이에 대해 비교적 대중적 합의율이 높다는 것을 확인하였다.
그러나, 특정 화자가 말하는 특정 주장/내용에 대한 신뢰도 개념은 개인의 주관과 매우 밀접한 관련이 있기 때문에 어떤 문장/문서를 신뢰하는지에 대한 것은 합의율이 낮다는 점이 여러 연구를 통해 확인되었다.
다르게 말해, 특정 문서가 높은/낮은 설득력을 갖는다는 것, 높은/낮은 일관성을 갖는다는 것, 높은/낮은 감정 세기를 갖는다는 것에 대한 대중적 합의율은 특정 문서가 신뢰할 수 있는/없는 내용을 담고 있다는 것에 대한 대중적 합의율보다 높다. 이에 대한 이유는 다음과 같다. 설득력/일관성은 논지의 구조에 대한 것이며, 명확한 논지 구조를 가진 문서의 경우 그 내용에 대한 반감이 있을 경우에도, 해당하는 문서가 높은 설득력/일관성을 갖는다는 평가가 가능하다. 또한, 감정의 세기 역시, 감정 유발 단어 및 감정 유발 표현이 제한적이기 때문에 해당하는 단어와 표현을 많이 사용하는 문서는 높은 감정 세기를 갖는다는 평가에 대한 합의율이 높을 수 있다.
그러나, 신뢰도의 경우 개인의 주관적인 신념에 대한 것이기 때문에 그 내용에 대해 어떤 신념을 가지고 있는 지가 신뢰도를 결정짓는 가장 주요한 요소이며, 이는 인구 통계학적 특징에 의해 큰 영향을 받을 수 있으며, 인구 통계학적 특징으로 정의될 수 없는 개인의 주관에 의한 영향도 받을 수 있기 때문에 신뢰도 평가에 대한 대중적 합의율이 비교적 낮다.
따라서, 자동 문서 교정을 위한 기존의 방법론이 단순히 일관성 점수, 설득력 점수, 감정 세기 점수 등의 한 숫자로 정의된 지표를 통해 교정 제안 혹은 문서 첨삭 후보군의 적절성을 평가하였다면, 신뢰도의 증강을 위한 문서 교정 및 첨삭을 위해서는 개인차를 고려한 신뢰도 분포 개념에 입각하여 교정 제안 혹은 문서 첨삭 후보군의 적절성을 평가해야 한다.
본 발명의 실시예들은, 주어진 문서에 대해 가능한 첨삭 후보군에 대한 자동 신뢰도 분포 예측을 수행하고, 원본 문서와 첨삭 후보군의 신뢰도 분포 각각이 보이는 평균과 표준편차 변화에 기반하여 신뢰도 증강을 수행할 수 있는 방법 및 그 시스템을 제공한다.
본 발명의 일 실시예에 따른 문서 신뢰도 증강 방법은 사용자로부터 문서가 입력되면 상기 입력된 문서의 각 문장에 대한 첨삭 후보군을 생성하는 단계; 상기 입력된 문서와 상기 생성된 첨삭 후보군 각각에 대한 신뢰도 분포를 예측하는 단계; 및 상기 예측된 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식을 통해 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 단계를 포함한다.
나아가, 본 발명의 일 실시예에 따른 문서 신뢰도 증강 방법은 상기 입력된 문서에 대해 신뢰도 증강된 신뢰도 증강 문서 및 상기 신뢰도 증강의 전과 후에 대한 신뢰도 분포 변화 정보를 상기 사용자에게 제공하는 단계를 더 포함할 수 있다.
상기 사용자에게 제공하는 단계는 상기 사용자와의 상호작용을 통해 상기 사용자가 희망하는 신뢰도 분포의 양태에 가장 유사한 문서 첨삭 결과를 상기 사용자에게 더 제공할 수 있다.
상기 복수 증강 방식은 상기 신뢰도 분포의 미리 설정된 기준 평균 이상의 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 상기 신뢰도 분포의 미리 설정된 기준 표준편차 이하의 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 상기 신뢰도 분포의 미리 설정된 기준 {평균/표준편차} 이상의 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함할 수 있다.
상기 생성하는 단계는 상기 사용자로부터 상기 복수 증강 방식에 대한 제어 정보를 입력 받는 단계; 상기 입력된 문서의 각 문장에 대한 전처리를 수행하는 단계; 및 상기 입력된 문서의 각 문장에 대한 복수의 첨삭 후보 문장 집합을 생성하는 단계를 포함하고, 상기 예측하는 단계는 상기 입력된 문서의 각 문장과 상기 첨삭 후보 문장 집합 내의 각 문장을 통해 조합 가능한 복수의 문서 집합의 각 문장에 대한 신뢰도 분포를 예측하며, 상기 수행하는 단계는 상기 예측된 복수의 문서 집합의 각 문장에 대한 신뢰도 분포와 상기 제어 정보에 기초하여 신뢰도 증강 문서를 선별함으로써, 상기 입력된 문서에 대한 신뢰도 증강을 수행할 수 있다.
본 발명의 일 실시예에 따른 문서 신뢰도 증강 시스템은 사용자로부터 문서가 입력되면 상기 입력된 문서의 각 문장에 대한 첨삭 후보군을 생성하는 생성부; 상기 입력된 문서와 상기 생성된 첨삭 후보군 각각에 대한 신뢰도 분포를 예측하는 예측부; 및 상기 예측된 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식을 통해 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 증강부를 포함한다.
나아가, 본 발명의 일 실시예에 따른 문서 신뢰도 증강 시스템은 상기 입력된 문서에 대해 신뢰도 증강된 신뢰도 증강 문서 및 상기 신뢰도 증강의 전과 후에 대한 신뢰도 분포 변화 정보를 상기 사용자에게 제공하는 출력부를 더 포함할 수 있다.
상기 출력부는 상기 사용자와의 상호작용을 통해 상기 사용자가 희망하는 신뢰도 분포의 양태에 가장 유사한 문서 첨삭 결과를 상기 사용자에게 더 제공할 수 있다.
상기 복수 증강 방식은 상기 신뢰도 분포의 미리 설정된 기준 평균 이상의 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 상기 신뢰도 분포의 미리 설정된 기준 표준편차 이하의 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 상기 신뢰도 분포의 미리 설정된 기준 {평균/표준편차} 이상의 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함할 수 있다.
상기 생성부는 상기 사용자로부터 상기 복수 증강 방식에 대한 제어 정보를 입력 받고, 상기 입력된 문서의 각 문장에 대한 전처리를 수행하며, 상기 입력된 문서의 각 문장에 대한 복수의 첨삭 후보 문장 집합을 생성하고, 상기 예측부는 상기 입력된 문서의 각 문장과 상기 첨삭 후보 문장 집합 내의 각 문장을 통해 조합 가능한 복수의 문서 집합의 각 문장에 대한 신뢰도 분포를 예측하며, 상기 증강부는 상기 예측된 복수의 문서 집합의 각 문장에 대한 신뢰도 분포와 상기 제어 정보에 기초하여 신뢰도 증강 문서를 선별함으로써, 상기 입력된 문서에 대한 신뢰도 증강을 수행할 수 있다.
본 발명의 실시예들에 따르면, 주어진 문서에 대해 가능한 첨삭 후보군에 대한 자동 신뢰도 분포 예측을 수행하고, 원본 문서와 첨삭 후보군의 신뢰도 분포 각각이 보이는 평균과 표준편차 변화에 기반하여 신뢰도 증강을 수행할 수 있다.
본 발명의 실시예들에 따르면, 분포 개념의 도입이라는 필요조건이 만족되기 때문에 큰 개인차를 포함하는 신뢰도 개념에 대한 문서 첨삭을 가능하게 한다. 즉, 본 발명은 (1) 신뢰도 분포의 높은 평균을 지향하는 방식의 신뢰도 증강을 진행하여 보다 많은 사람들이 신뢰도 증강의 대상이 되는 문서를 신뢰할 수 있도록 문서를 첨삭하는 것, (2) 신뢰도 분포의 낮은 표준편차를 지향하는 방식의 신뢰도 증강을 진행하여 신뢰도 증강의 대상이 되는 문서에 대한 신뢰 여부가 보다 논란의 여지가 없도록 첨삭하는 것, (3) 신뢰도 분포의 높은 {평균/표준편차}를 지향하는 방식의 신뢰도 증강을 진행하여 상기 두 방식의 결합된 방식으로 신뢰도 증강의 대상이 되는 문서를 첨삭하는 것을 가능하게 한다.
본 발명의 실시예들에 따르면, 교정 대상이 되는 문서 내의 각 문장에 대해 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 입력하고 해당하는 신뢰도 분포의 양태와 가장 유사한 신뢰도 분포 양태를 보이는 문장 첨삭 후보를 사용자에게 제공할 수 있다. 따라서 본 발명은 개인의 주관에 대해 큰 영향을 받기 때문에 자동화된 교정 시스템에 적용되기 어려웠던 신뢰도 증강에 대한 문서 교정을 분포 개념의 도입을 통해 가능하도록 하며, 이를 통해 불특정 타인 집단으로부터 보다 신뢰받기 위해 문서가 어떻게 변형되어야 하는 지에 대한 상세 정보를 다양한 형태로 제공받을 수 있다.
이러한 본 발명은 독자의 신뢰도가 중요한 공문서, 뉴스, 기술 문서의 자동 교정을 수행할 수도 있고, 개인의 주관에 기반하되 신뢰성이 높은 문서를 생산할 수도 있다.
도 1은 본 발명의 일 실시예에 따른 문서 신뢰도 증강 시스템에 대한 구성을 나타낸 것이다.
도 2는 도 1에 도시된 신뢰도 증강부의 일 실시예 구성을 나타낸 것이다.
도 3은 도 1에 도시된 출력부의 일 실시예 구성을 나타낸 것이다.
도 4는 본 발명의 문서 신뢰도 증강 시스템의 출력 결과를 설명하기 위한 일 예시도를 나타낸 것이다.
도 5는 문서 신뢰도 증강 시스템에 포함되는 문장 첨삭 후보 코퍼스의 예시를 나타낸 것이다.
도 6은 문서 신뢰도 증강 시스템에 포함되는 신뢰도 분포 코퍼스의 예시를 나타낸 것이다.
도 7은 도 1에 도시된 신뢰도 증강부에 의한 일 실시예의 동작 흐름도를 나타낸 것이다.
도 8은 도 1에 도시된 출력부에 의한 일 실시예의 동작 흐름도를 나타낸 것이다.
본 발명의 이점 및 특징, 그리고 그것들을 달성하는 방법은 첨부되는 도면과 함께 상세하게 후술되어 있는 실시예들을 참조하면 명확해질 것이다. 그러나, 본 발명은 이하에서 개시되는 실시예들에 한정되는 것이 아니라 서로 다른 다양한 형 태로 구현될 것이며, 단지 본 실시예들은 본 발명의 개시가 완전하도록 하며, 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 발명의 범주를 완전하게 알려주기 위해 제공되는 것이며, 본 발명은 청구항의 범주에 의해 정의될 뿐이다.
본 명세서에서 사용된 용어는 실시예들을 설명하기 위한 것이며, 본 발명을 제한하고자 하는 것은 아니다. 본 명세서에서, 단수형은 문구에서 특별히 언급하지 않는 한 복수형도 포함한다. 명세서에서 사용되는 "포함한다(comprises)" 및/또는 "포함하는(comprising)"은 언급된 구성요소, 단계, 동작 및/또는 소자는 하나 이상 의 다른 구성요소, 단계, 동작 및/또는 소자의 존재 또는 추가를 배제하지 않는다.
다른 정의가 없다면, 본 명세서에서 사용되는 모든 용어(기술 및 과학적 용어를 포함)는 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 공통적으로 이해될 수 있는 의미로 사용될 수 있을 것이다. 또한, 일반적으로 사용되는 사전에 정의되어 있는 용어들은 명백하게 특별히 정의되어 있지 않는 한 이상적으로 또는 과도하게 해석되지 않는다.
이하, 첨부한 도면들을 참조하여, 본 발명의 바람직한 실시예들을 보다 상세하게 설명하고자 한다. 도면 상의 동일한 구성요소에 대해서는 동일한 참조 부호를 사용하고 동일한 구성요소에 대해서 중복된 설명은 생략한다.
본 발명의 실시예들은, 주어진 문서에 대해 가능한 첨삭 후보군에 대한 자동 신뢰도 분포 예측을 수행하고, 원본 문서와 첨삭 후보군의 신뢰도 분포 각각이 보이는 평균과 표준편차 변화에 기반하여 신뢰도를 증강시키는 것을 그 요지로 한다.
여기서, 본 발명은 신뢰도 분포의 높은 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 신뢰도 분포의 낮은 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 신뢰도 분포의 높은 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함하는 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 따라 신뢰도 증강을 수행할 수 있다.
나아가, 본 발명은 사용자가 입력한 문서에 대한 신뢰도 증강을 진행한 뒤 사용자에게 증강된 문서와 관련된 통계 지표를 그래프 형태로 제공하며, 사용자와의 상호작용을 통해 불특정 타인 집단이 사용자가 입력한 문서의 미세 조정에 따라 어떤 신뢰도 변화를 보일 지에 대하여 사전 모니터링을 수행할 수 있게 한다.
도 1은 본 발명의 일 실시예에 따른 문서 신뢰도 증강 시스템에 대한 구성을 나타낸 것이다.
도 1에 도시된 바와 같이, 본 발명의 일 실시예에 따른 문서 신뢰도 증강 시스템(100)은 신뢰도 증강부(110), 문장 첨삭 후보 코퍼스(120), 신뢰도 분포 코퍼스(130), 출력부(140) 및 제어부(150)를 포함한다.
여기서, 본 발명의 문서 신뢰도 증강 시스템(100)은 도 4에 도시된 바와 같이, 사용자로부터 입력 받은 문서(입력 텍스트)에 대해 각 문장 별로 신뢰도 증강을 수행하며 신뢰도가 증강된 문서를 사용자에게 제공함과 동시에 신뢰도 증강 전/후에 해당하는 독자 신뢰도 분포를 각 문장에 대해 사용자에게 출력 제공할 수 있다.
신뢰도 증강부(110)는 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 따라 입력 받은 문서에 대한 신뢰도 증강을 수행한다.
이 때, 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식은 신뢰도 분포의 높은 평균 예를 들어, 미리 설정된 기준 평균 이상의 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 신뢰도 분포의 낮은 표준편차 예를 들어, 미리 설정된 기준 표준편차 이하의 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 신뢰도 분포의 높은 {평균/표준편차} 예를 들어, 미리 설정된 기준 {평균/표준편차} 이상의 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함할 수 있다.
구체적으로, 높은 신뢰도 평균을 지향하는 방식의 신뢰도 증강은 보다 많은 사람들이 신뢰도 증강의 대상이 되는 문서를 신뢰할 수 있도록 문서를 첨삭하는 것으로 해석할 수 있으며, 낮은 신뢰도 표준편차를 지향하는 방식의 신뢰도 증강은 신뢰도 증강의 대상이 되는 문서에 대한 신뢰 여부가 보다 논란의 여지가 없도록 첨삭하는 것으로 해석할 수 있고, 높은 {평균/표준편차} 값을 지향하는 방식의 신뢰도 증강은 상기 두 방식의 결합으로 해석될 수 있다.
즉, 신뢰도 증강부(110)는 사용자로부터 신뢰도 증강의 대상이 되는 문서를 입력으로 받고, 선택적으로 사용자로부터 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 중 어떤 방식에 해당하는 정보를 출력 제공할 것인지 대한 정보(이하, "제어 정보"라 칭함)를 입력 받고, 입력 받은 제어 정보를 제어부(130)로 전달하며, 주어진 문서 내의 각 문장에 대해 첨삭 후보군을 생성하고, 제어 정보를 다시 제어부(130)로부터 전달받아 첨삭 후보군을 조합하여 생성 가능한 문서 중 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 각각 해당되는 문서 중 상기 제어 정보와 일치하는 문서를 선별한다.
여기서, 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 각각 해당되는 문서는 (1) 문서 전반에 대해 예측된 신뢰도 분포의 평균이 가장 높은 첨삭 후 문서와, (2) 문서 전반에 대해 예측된 신뢰도 분포의 표준편차가 가장 낮은 첨삭 후 문서와, (3) 문서 전반에 대해 예측된 신뢰도 분포의 {평균/표준편차} 값이 가장 높은 첨삭 후 문서를 포함할 수 있다.
나아가, 신뢰도 증강부(110)는 제어 정보가 사용자로부터 입력되지 않았을 경우 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 모두에 해당하는 정보를 제어 정보로 출력 제공할 수 있다.
문장 첨삭 후보 코퍼스(120)는 언어 전문가에 의해 수동적으로 문장의 의미를 유지한 채로 문장 구조를 변형하는 방식의 첨삭을 진행한 문장(이하, "첨삭 후 문장"이라 칭함), 원본 문장, 언어 전문가에 의해 0점(즉, 의미 차이 없음)에서 5점(즉, 큰 의미 차이) 사이의 범위 내 점수로 판단된 원본 문장과 첨삭 후 문장의 의미 차이를 저장한다.
여기서, 문장 첨삭 후보 코퍼스(120)는 신뢰도 증강부(110)가 주어진 문서 내의 각 문장에 대한 첨삭 후보군을 생성하는 모델을 활용할 때 해당하는 모델을 지도 학습(supervised learning) 시키기 위한 학습 기준으로 활용될 수 있다.
도 5는 문서 신뢰도 증강 시스템에 포함되는 문장 첨삭 후보 코퍼스의 예시를 나타낸 것으로, 도 5에 도시된 바와 같이, 문장 첨삭 후보 코퍼스는 복수 예시 문장 집합과 각 예시 문장에 대해서 복수 첨삭된 문장들을 저장하며, 첨삭된 문장마다 첨삭 전 원본 문장과의 유사도를 저장한다.
신뢰도 분포 코퍼스(130)는 직접적인 신뢰도 설문을 통해 수집한 각 문장에 대한 독자들의 신뢰도 분포를 저장한다.
여기서, 신뢰도 분포 코퍼스(130)는 신뢰도 증강부(110)가 첨삭 후 문서의 신뢰도 분포를 예측함에 있어 신뢰도 분포 예측 모델을 포함할 때 해당하는 모델을 지도 학습(supervised learning)시키기 위한 학습 기준으로 활용될 수 있다.
도 6은 문서 신뢰도 증강 시스템에 포함되는 신뢰도 분포 코퍼스의 예시를 나타낸 것으로, 도 6에 도시된 바와 같이, 신뢰도 분포 코퍼스는 복수 문서 집합과 각 문서에 대해 독자 그룹들이 각각 실제로 설문을 통해 보인 신뢰도 분포를 저장한다.
나아가, 신뢰도 분포 코퍼스(130)가 저장하는 복수 문서 집합은 일상생활에서 접할 수 있는 다양한 주제들에 대한 다양한 종류의 문서들을 포함할 수 있어야 한다. 상세한 설명을 위한 예시로, 문서의 주제는 생활, 건강, 정치, 정책, 경제, 환경을 포함할 수 있으며, 문서의 종류는 SNS 게시글, 블로그 게시글, 온라인 뉴스, 온라인 포럼 게시글, 연구 논문, 도서를 포함할 수 있다. 바람직하게, 코퍼스 전반에 걸쳐 직접적인 설문의 대상이 되는 문서들 각각이 갖는 주제와 종류는 획일화되지 않아야 한다. 이는 본 발명의 일 실시예에 따른 시스템에서 설문을 통해 수집한 신뢰도 분포 코퍼스(130)가 신뢰도 분포를 자동 예측하기 위한 학습 기준으로 활용되기 때문이다. 신뢰도 분포 코퍼스(130)가 생활/건강에 대한 문서들과 이 문서들에 대해 독자들이 보이는 신뢰도 설문 결과만을 포함할 경우 해당하는 예측 기준 코퍼스를 통해 학습된 예측 모델은 정치/정책에 대한 문서의 신뢰도 분포를 예측하기에 부적절할 것이며 예측 결과가 실제 독자들이 보일 신뢰도 분포와는 상이할 것으로 예상할 수 있다.
출력부(140)는 신뢰도 증강부(110)에 의해 선별된 신뢰도 증강된 문서를 사용자에게 출력 제공한다. 동시에, 출력부(140)는 사용자에 의해 입력된 원본 문서 내 각 문장에 대해 예측된 신뢰도 분포와 신뢰도 증강부(110)에 의해 증강된 문서 내 각 문장에 대해 예측된 신뢰도 분포를 그래프 형태로 병렬 대조 비교하여 사용자에게 출력 제공한다. 예컨대, 출력부(140)는 도 4에 도시된 출력 결과와 같이 원본 문서 내 각 문장에 대해 예측된 신뢰도 분포와 증강된 문서 내 각 문장에 대해 예측된 신뢰도 분포를 그래프 형태로 병렬 대조 비교하여 사용자에게 출력 제공할 수 있다.
이 때, 출력부(140)는 신뢰도 증강 전/후에 해당되는 독자 신뢰도 분포를 각 문장에 대해 사용자에게 출력 제공하며, 사용자 상호작용을 통해 신뢰도 증강에 대한 자세한 정보를 선택적, 단계별로 제공함으로써, 사용자는 불특정 타인 집단이 사용자가 입력한 문서의 미세 변형에 따라 어떤 신뢰도 반응 변화를 보일 지에 대한 사전 모니터링을 수행할 수 있다.
또한, 출력부(140)는 사용자와 출력된 그래프에 대한 드래그 앤 드랍 방식으로 상호작용할 수 있으며, 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 입력하였을 때 입력 받은 신뢰도 분포와 가장 가까운 신뢰도 분포를 갖는 첨삭 후보를 선별하여 신뢰도 분포의 대상이 되는 문장을 해당하는 첨삭 후보로 변경 재출력 제공하며, 해당하는 첨삭 후보의 신뢰도 분포로 기존에 사용자가 드래그 앤 드랍 방식으로 상호작용했던 그래프를 변경 재출력하여 제공할 수 있다.
제어부(150)는 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 중 어떤 종류의 신뢰도 증강을 수행할 것인지를 결정한다.
이 때, 제어부(150)는 신뢰도 증강부(110)로 수신된 제어 정보에 기초하여 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 중 어떤 종류의 신뢰도 증강을 수행할 것인지를 결정할 수 있다.
도 2는 도 1에 도시된 신뢰도 증강부의 일 실시예 구성을 나타낸 것이다.
도 2에 도시된 바와 같이, 신뢰도 증강부(110)는 입력부(111), 전처리부(112), 첨삭 후보 생성부(113), 신뢰도 분포 예측부(114) 및 증강 문서 선별부(115)를 포함한다.
입력부(111)는 사용자로부터 신뢰도 증강의 대상이 되는 문서를 입력으로 받고, 선택적으로 사용자로부터 제어 정보를 입력 받아 제어부(130)로 전달한다.
이 때, 입력부(111)는 제어 정보가 사용자로부터 입력되지 않았을 경우 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 모두에 해당하는 정보를 제어 정보로 제공할 수 있다.
바람직하게, 상기 제어 정보가 사용자로부터 입력되지 않았을 경우, 상기 제어 정보는 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식 모두에 해당하는 정보를 제어 정보로 출력 제공할 수 있다.
전처리부(112)는 사용자로부터 입력 받은 문서를 입력부(111)로부터 전달받고, 이에 대해 의미역 분석(semantic role labeling), 의존성 분석(dependency parsing), 담화 구조 분석(discourse parsing)을 수행한다.
여기서, 의미역 분석은 DeepSemanticRoleLabeling 또는 PathLSTM Semantic Role Labeler와 같은 의미역 추출기를 통해 진행될 수 있고, 의존성 분석은 StanfordCoreNLP와 같은 구문 분석기를 통해 진행될 수 있으며, 담화 구조 분석은 PDTB-style discourse parser와 같은 담화 구조 분석기를 통해 진행될 수 있다.
첨삭 후보 생성부(113)는 전처리부(112)로부터 (1) 사용자로부터 입력 받은 문서, 이에 대한 (2) 의미역 분석 결과, (3) 의존성 분석 결과, (4) 담화 구조 분석 결과를 수신하여 원본 문서 내 각 문장에 대해 복수의 첨삭 후보 문장 집합을 생성한다.
이 때, 첨삭 후보 생성부(113)는 각 문장에 대해 생성 가능한 복수 첨삭 후보 문장 중 원본 문장과의 의미 차이가 가장 적은 L개의 문장을 선별하여 복수의 첨삭 후보 문장 집합으로 생성할 수 있다.
바람직하게, 첨삭 후보 생성부(113)는 지도 학습(supervised learning) 기반 첨삭 문장 생성 모델을 통해 첨삭 후보 문장을 생성할 수 있으며, 해당하는 모델은 문장 첨삭 후보 코퍼스(120)를 통해 지도 학습될 수 있다. 여기서, 학습 모델은 학습 중에 첨삭 후보 코퍼스(120)로부터 복수 문장 집합 중 원본 문장들을 불러와 의미역 분석, 의존성 분석, 담화 구조 분석을 수행하며, (1) 원본 문장들과 (2) 의미역 분석 결과, (3) 의존성 분석 결과, (4) 담화 구조 분석 결과를 입력 기준으로 삼고, 각 원본 문장마다 상응하는 첨삭 후 문장들 그리고 원본 문장과 첨삭 후 문장의 의미 차이들을 불러와 출력 기준으로 삼을 수 있다. 물론, 첨삭 후보 생성부(113)에서의 의미역 분석, 의존성 분석, 담화 구조 분석 또한 전처리부(112)에서 수행하는 것과 같은 방식으로 진행될 수 있다.
신뢰도 분포 예측부(114)는 첨삭 후보 생성부(113)로부터 (1) 사용자로부터 입력 받은 문서 내 문장들, 이에 대한 (2) 의미역 분석 결과, (3) 의존성 분석 결과, (4) 담화 구조 분석 결과, (5) 복수 첨삭 후보 문장 집합을 수신하고 또는 전달받고, 사용자로부터 입력 받은 문서 내의 복수 문장 집합과 이에 상응하는 복수 첨삭 후보 문장 집합 내의 각 문장을 통해, 조합 가능한 복수 문서 집합 내의 각 문장에 대해 독자 그룹이 보일 것으로 예상하는 신뢰도 분포를 예측한다.
이 때, 첨삭 후보 문장 집합 내 각 문장을 통해 복수 문서 집합을 조합하는 과정은, 사용자로부터 입력 받은 문서에 존재하는 복수 문장 집합의 문장에 상응하는 복수 첨삭 후보 중 하나의 첨삭 후보를 무작위로 선정하여 대체하는 과정을 복수 문장 집합 전체에 대해 진행하는 방식으로 할 수 있다.
바람직하게, 조합 가능한 복수 문서 집합은 상술한 무작위 선정 대체 방식을 M번 순회하여 M개의 첨삭 후보 문서를 생성하는 방식으로 할 수 있다. 여기서, M은 시스템 초기값으로 설정되는 자연수이며, 시스템의 계산 속도와 사용자 만족도를 고려하여 M이 설정될 수 있다.
바람직하게, 신뢰도 분포 예측부(114)는 신뢰도 분포 코퍼스(130)로부터 지도 학습(supervised learning)을 통해 학습한 모델을 이용하여 각 문장에 대한 신뢰도 분포를 예측한다. 여기서, 학습 모델은 학습 중에 신뢰도 분포 코퍼스(130)에 저장된 복수 문서 집합 내 각 문장을 신뢰도 분포 예측 대상 문장으로 삼았을 때, 신뢰도 분포 예측 대상 문장의 앞에 등장한 일정 개수 예를 들어, N개 문장, 뒤에 등장한 N개 문장과 해당 문장을 모두 포함하는 2N+1개 문장 집합에 대한 의미역 분석 결과, 의존성 분석 결과, 담화 구조 분석 결과를 입력 기준으로 삼고, 신뢰도 분포 예측 대상 문장에 대해 실제 설문을 통해 측정된 복수 문장 집합에 신뢰도 분포를 신뢰도 분포 코퍼스(130)로부터 불러와 출력 기준으로 삼을 수 있다. 여기서, N은 시스템 초기값으로 설정되는 자연수이며, 주어진 문서 내에서 신뢰도 분포 예측 대상 문장의 앞 혹은 뒤에 N개보다 적은 수의 문장이 존재할 경우에는 존재하지 않는 문장을 공백 문장(null sentence)으로 대체하여 입력 기준으로 삼을 수 있다. 물론, 신뢰도 분포 예측부(114)에서의 의미역 분석, 의존성 분석, 담화 구조 분석은 전처리부(113)에서 수행하는 것과 같은 방식으로 진행될 수 있다.
증강 문서 선별부(115)는 신뢰도 분포 예측부(114)에서 사용자로부터 입력 받은 문서 내 복수 문장 집합에 상응하는 복수 첨삭 후보 문장 집합 내 각 문장을 통해 조합 가능한 복수 문서 집합과 각각에 대해 독자 그룹이 보일 것으로 예상하는 신뢰도 분포를 전달받고, 제어부(130)로부터 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 대한 제어 정보를 전달받아 미리 설정된 문서 중 제어 정보와 일치하는 문서를 신뢰도 증강 문서로서 선별한다.
여기서, 미리 설정된 문서는 (1) 각 문장에 대해 예측된 신뢰도 분포의 평균을 해당하는 문서 내 모든 문장에 대해 평균한 값이 가장 높은 첨삭 후 문서와, (2) 각 문장에 대해 예측된 신뢰도 분포의 표준편차를 해당하는 문서 내 모든 문장에 대해 평균한 값이 가장 낮은 첨삭 후 문서와, (3) 각 문장에 대해 예측된 신뢰도 분포의 {평균/표준편차} 값을 해당하는 문서 내 모든 문장에 대해 평균한 값이 가장 높은 첨삭 후 문서를 포함할 수 있다.
도 3은 도 1에 도시된 출력부의 일 실시예 구성을 나타낸 것이다.
도 3에 도시된 바와 같이, 출력부(140)는 증강 텍스트 출력부(141), 그래프 출력부(142) 및 사용자 상호작용부(143)를 포함한다.
증강 텍스트 출력부(141)는 증강 문서 선별부(115)에 의해 선별된 신뢰도 증강 문서를 전달받고, 이를 사용자에게 출력 제공한다.
그래프 출력부(142)는 증강 문서 선별부(115)로부터 선별된 신뢰도 증강 문서를 전달받고, (1) 이에 대한 신뢰도 분포 예측 결과, (2) 원본 문서, (3) 원본 문서에 대한 신뢰도 분포 예측 결과를 신뢰도 예측부(114)로부터 전달받아, 도 4에 도시된 일 예와 같이, 사용자에 의해 입력된 원본 문서 내 각 문장에 대해 예측된 신뢰도 분포와 신뢰도 증강부(110)에 의해 증강된 문서 내 각 문장에 대해 예측된 신뢰도 분포를 그래프 형태로 병렬 대조 비교하여 사용자에게 출력 제공한다.
사용자 상호작용부(143)는 그래프 출력부(142)에 의해 사용자에게 출력된 신뢰도 분포 그래프에 대한 드래그 앤 드랍 방식으로 사용자와 상호작용할 수 있도록 하며, 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 사용자 상호작용부(143)에 입력하였을 때 입력 받은 신뢰도 분포와 가장 가까운 신뢰도 분포를 갖는 첨삭 후보를 선별하여 신뢰도 분포의 대상이 되는 문장을 해당하는 첨삭 후보로 변경 재출력하며, 해당하는 첨삭 후보의 신뢰도 분포로 기존에 사용자가 드래그 앤 드랍 방식으로 상호작용했던 그래프를 변경 재출력하여 제공한다.
이 때, 사용자 상호작용부(143)는 그래프 출력부(142)에 의해 출력된 신뢰도 분포 그래프에서 사용자가 드래그 앤 드랍 방식을 통해 사용자가 희망하는 신뢰도 분포의 양태를 입력할 수 있도록 할 수 있다.
즉, 사용자 상호작용부(143)는 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 사용자 상호작용부(143)에 입력하였을 때, 해당하는 문장의 첨삭 후보 집합 중 각 첨삭 후보 및 상응하는 예측 신뢰도 분포를 신뢰도 분포 예측부(114)로부터 전달받고, 분포 유사도 지표를 통해 첨삭 후보 중 사용자가 희망하는 신뢰도 분포의 양태와 가장 유사도가 높은 예측 신뢰도 분포를 갖는 첨삭 후보를 선별한다. 여기서, 분포 유사도는 Kullback-Leibler divergence나 Levy-Prokhorov metric으로 정의될 수 있다.
이와 같이, 본 발명의 실시예에 따른 방법은 주어진 문서에 대해 가능한 첨삭 후보군에 대한 자동 신뢰도 분포 예측을 수행하고, 원본 문서와 첨삭 후보군의 신뢰도 분포 각각이 보이는 평균과 표준편차 변화에 기반하여 신뢰도 증강을 수행할 수 있다.
또한, 본 발명의 실시예에 따른 방법은 분포 개념의 도입이라는 필요조건이 만족되기 때문에 큰 개인차를 포함하는 신뢰도 개념에 대한 문서 첨삭을 가능하게 한다. 즉, 본 발명에 따른 방법은 (1) 신뢰도 분포의 높은 평균을 지향하는 방식의 신뢰도 증강을 진행하여 보다 많은 사람들이 신뢰도 증강의 대상이 되는 문서를 신뢰할 수 있도록 문서를 첨삭하는 것, (2) 신뢰도 분포의 낮은 표준편차를 지향하는 방식의 신뢰도 증강을 진행하여 신뢰도 증강의 대상이 되는 문서에 대한 신뢰 여부가 보다 논란의 여지가 없도록 첨삭하는 것, (3) 신뢰도 분포의 높은 {평균/표준편차}를 지향하는 방식의 신뢰도 증강을 진행하여 상기 두 방식의 결합된 방식으로 신뢰도 증강의 대상이 되는 문서를 첨삭하는 것을 가능하게 한다.
또한, 본 발명의 실시예에 따른 방법은 교정 대상이 되는 문서 내의 각 문장에 대해 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 입력하고 해당하는 신뢰도 분포의 양태와 가장 유사한 신뢰도 분포 양태를 보이는 문장 첨삭 후보를 사용자에게 제공할 수 있다. 따라서 본 발명은 개인의 주관에 대해 큰 영향을 받기 때문에 자동화된 교정 시스템에 적용되기 어려웠던 신뢰도 증강에 대한 문서 교정을 분포 개념의 도입을 통해 가능하도록 하며, 이를 통해 불특정 타인 집단으로부터 보다 신뢰받기 위해 문서가 어떻게 변형되어야 하는 지에 대한 상세 정보를 다양한 형태로 제공받을 수 있다.
도 7은 도 1에 도시된 신뢰도 증강부에 의한 일 실시예의 동작 흐름도를 나타낸 것이다.
도 7을 참조하면, 신뢰도 증강부(110)에 의한 방법은 입력 단계(S310), 전처리 단계(S320), 신뢰도 분포 예측 단계(S330), 첨삭 후보 생성 단계(S340), 증거 문서 선별 단계(S350)를 포함한다.
입력 단계(S310)는 사용자로부터 신뢰도 증강의 대상이 되는 문서를 입력으로 받고, 선택적으로 사용자로부터 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식에 대한 제어 정보를 입력 받는 단계이다. 전처리 단계(S320)는 사용자로부터 입력 받은 문서 내 각 문장에 대한 전처리를 진행하는 단계이다. 첨삭 후보 생성 단계(S330)는 원본 문서 내 각 문장에 대해 복수 첨삭 후보 문장 집합을 생성하는 단계이다. 신뢰도 분포 예측 단계(S340)는 사용자로부터 입력 받은 문서 내 복수 문서 집합과 이에 상응하는 복수 첨삭 후보 문장 집합 내 각 문장을 통해 조합 가능한 복수 문서 집합 내 각 문장에 대해 독자 그룹이 보일 것으로 예상되는 신뢰도 분포를 예측하는 단계이다. 증거 문서 선별 단계(S350)는 사용자로부터 입력 받은 문서 내 복수 문장 집합에 상응하는 복수 첨삭 후보 문장 집합 내 각 문장을 통해 조합 가능한 복수 문서 집합과 각각에 대해 독자 그룹이 보일 것으로 예상되는 신뢰도 분포에 대한 분석을 통해 제어 정보에 따라 신뢰도 증강 문서를 선별하는 단계이다.
도 8은 도 1에 도시된 출력부에 의한 일 실시예의 동작 흐름도를 나타낸 것이다.
도 8을 참조하면, 출력부(140)에 의한 방법은 증강 텍스트 출력 단계(S410), 그래프 출력 단계(S420) 및 사용자 상호작용 단계(S430)를 포함한다.
증강 텍스트 출력 단계(S410)는 신뢰도 증강 문서를 사용자에게 출력 제공하는 단계이다. 그래프 출력 단계(S420)는 사용자에 의해 입력된 원본 문서 내 각 문장에 대해 예측된 신뢰도 분포와 증강된 문서 내 각 문장에 대해 예측된 신뢰도 분포를 그래프 형태로 병렬 대조 비교하여 사용자에게 출력 제공하는 단계이다. 사용자 상호작용 단계(S350)는 사용자에게 출력된 신뢰도 분포 그래프에 대한 드래그 앤 드랍 방식으로 사용자와 상호작용할 수 있도록 하며, 사용자가 드래그 앤 드랍 방식을 통해 희망하는 신뢰도 분포의 양태를 입력하였을 때 입력 받은 신뢰도 분포와 가장 가까운 신뢰도 분포를 갖는 첨삭 후보를 선별하여 신뢰도 분포의 대상이 되는 문장을 해당하는 첨삭 후보로 변경 재출력하며, 해당하는 첨삭 후보의 신뢰도 분포로 기존에 사용자가 드래그 앤 드랍 방식으로 상호작용했던 그래프를 변경 재출력하여 제공하는 단계이다.
비록, 도 7과 8의 방법에서 그 설명이 생략되었더라도, 도 7과 8을 구성하는 각 단계는 도 1 내지 도 6에서 설명한 모든 내용을 포함할 수 있으며, 이는 이 기술 분야에 종사하는 당업자에게 있어서 자명하다.
이상에서 설명된 시스템 또는 장치는 하드웨어 구성요소, 소프트웨어 구성요소, 및/또는 하드웨어 구성요소 및 소프트웨어 구성요소의 조합으로 구현될 수 있다. 예를 들어, 실시예들에서 설명된 시스템, 장치 및 구성요소는, 예를 들어, 프로세서, 콘트롤러, ALU(arithmetic logic unit), 디지털 신호 프로세서(digital signal processor), 마이크로컴퓨터, FPA(field programmable array), PLU(programmable logic unit), 마이크로프로세서, 또는 명령(instruction)을 실행하고 응답할 수 있는 다른 어떠한 장치와 같이, 하나 이상의 범용 컴퓨터 또는 특수 목적 컴퓨터를 이용하여 구현될 수 있다. 처리 장치는 운영 체제(OS) 및 상기 운영 체제 상에서 수행되는 하나 이상의 소프트웨어 애플리케이션을 수행할 수 있다. 또한, 처리 장치는 소프트웨어의 실행에 응답하여, 데이터를 접근, 저장, 조작, 처리 및 생성할 수도 있다. 이해의 편의를 위하여, 처리 장치는 하나가 사용되는 것으로 설명된 경우도 있지만, 해당 기술분야에서 통상의 지식을 가진 자는, 처리 장치가 복수 개의 처리 요소(processing element) 및/또는 복수 유형의 처리 요소를 포함할 수 있음을 알 수 있다. 예를 들어, 처리 장치는 복수 개의 프로세서 또는 하나의 프로세서 및 하나의 콘트롤러를 포함할 수 있다. 또한, 병렬 프로세서(parallel processor)와 같은, 다른 처리 구성(processing configuration)도 가능하다.
소프트웨어는 컴퓨터 프로그램(computer program), 코드(code), 명령(instruction), 또는 이들 중 하나 이상의 조합을 포함할 수 있으며, 원하는 대로 동작하도록 처리 장치를 구성하거나 독립적으로 또는 결합적으로(collectively) 처리 장치를 명령할 수 있다. 소프트웨어 및/또는 데이터는, 처리 장치에 의하여 해석되거나 처리 장치에 명령 또는 데이터를 제공하기 위하여, 어떤 유형의 기계, 구성요소(component), 물리적 장치, 가상 장치(virtual equipment), 컴퓨터 저장 매체 또는 장치, 또는 전송되는 신호 파(signal wave)에 영구적으로, 또는 일시적으로 구체화(embody)될 수 있다. 소프트웨어는 네트워크로 연결된 컴퓨터 시스템 상에 분산되어서, 분산된 방법으로 저장되거나 실행될 수도 있다. 소프트웨어 및 데이터는 하나 이상의 컴퓨터 판독 가능 기록 매체에 저장될 수 있다.
실시예들에 따른 방법은 다양한 컴퓨터 수단을 통하여 수행될 수 있는 프로그램 명령 형태로 구현되어 컴퓨터 판독 가능 매체에 기록될 수 있다. 상기 컴퓨터 판독 가능 매체는 프로그램 명령, 데이터 파일, 데이터 구조 등을 단독으로 또는 조합하여 포함할 수 있다. 상기 매체에 기록되는 프로그램 명령은 실시예를 위하여 특별히 설계되고 구성된 것들이거나 컴퓨터 소프트웨어 당업자에게 공지되어 사용 가능한 것일 수도 있다. 컴퓨터 판독 가능 기록 매체의 예에는 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체(magnetic media), CD-ROM, DVD와 같은 광기록 매체(optical media), 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical media), 및 롬(ROM), 램(RAM), 플래시 메모리 등과 같은 프로그램 명령을 저장하고 수행하도록 특별히 구성된 하드웨어 장치가 포함된다. 프로그램 명령의 예에는 컴파일러에 의해 만들어지는 것과 같은 기계어 코드뿐만 아니라 인터프리터 등을 사용해서 컴퓨터에 의해서 실행될 수 있는 고급 언어 코드를 포함한다. 상기된 하드웨어 장치는 실시예의 동작을 수행하기 위해 하나 이상의 소프트웨어 모듈로서 작동하도록 구성될 수 있으며, 그 역도 마찬가지이다.
이상과 같이 실시예들이 비록 한정된 실시예와 도면에 의해 설명되었으나, 해당 기술분야에서 통상의 지식을 가진 자라면 상기의 기재로부터 다양한 수정 및 변형이 가능하다. 예를 들어, 설명된 기술들이 설명된 방법과 다른 순서로 수행되거나, 및/또는 설명된 시스템, 구조, 장치, 회로 등의 구성요소들이 설명된 방법과 다른 형태로 결합 또는 조합되거나, 다른 구성요소 또는 균등물에 의하여 대치되거나 치환되더라도 적절한 결과가 달성될 수 있다.
그러므로, 다른 구현들, 다른 실시예들 및 특허청구범위와 균등한 것들도 후술하는 특허청구범위의 범위에 속한다.
Claims (10)
- 사용자로부터 문서가 입력되면 상기 입력된 문서의 각 문장에 대한 첨삭 후보군을 생성하는 단계;상기 입력된 문서와 상기 생성된 첨삭 후보군 각각에 대한 신뢰도 분포를 예측하는 단계; 및상기 예측된 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식을 통해 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 단계를 포함하는 문서 신뢰도 증강 방법.
- 제1항에 있어서,상기 입력된 문서에 대해 신뢰도 증강된 신뢰도 증강 문서 및 상기 신뢰도 증강의 전과 후에 대한 신뢰도 분포 변화 정보를 상기 사용자에게 제공하는 단계를 더 포함하는 것을 특징으로 하는 문서 신뢰도 증강 방법.
- 제2항에 있어서,상기 사용자에게 제공하는 단계는상기 사용자와의 상호작용을 통해 상기 사용자가 희망하는 신뢰도 분포의 양태에 가장 유사한 문서 첨삭 결과를 상기 사용자에게 더 제공하는 것을 특징으로 하는 문서 신뢰도 증강 방법.
- 제1항에 있어서,상기 복수 증강 방식은상기 신뢰도 분포의 미리 설정된 기준 평균 이상의 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 상기 신뢰도 분포의 미리 설정된 기준 표준편차 이하의 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 상기 신뢰도 분포의 미리 설정된 기준 {평균/표준편차} 이상의 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함하는 것을 특징으로 하는 문서 신뢰도 증강 방법.
- 제1항에 있어서,상기 생성하는 단계는상기 사용자로부터 상기 복수 증강 방식에 대한 제어 정보를 입력 받는 단계;상기 입력된 문서의 각 문장에 대한 전처리를 수행하는 단계; 및상기 입력된 문서의 각 문장에 대한 복수의 첨삭 후보 문장 집합을 생성하는 단계를 포함하고,상기 예측하는 단계는상기 입력된 문서의 각 문장과 상기 첨삭 후보 문장 집합 내의 각 문장을 통해 조합 가능한 복수의 문서 집합의 각 문장에 대한 신뢰도 분포를 예측하며,상기 수행하는 단계는상기 예측된 복수의 문서 집합의 각 문장에 대한 신뢰도 분포와 상기 제어 정보에 기초하여 신뢰도 증강 문서를 선별함으로써, 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 것을 특징으로 하는 문서 신뢰도 증강 방법.
- 사용자로부터 문서가 입력되면 상기 입력된 문서의 각 문장에 대한 첨삭 후보군을 생성하는 생성부;상기 입력된 문서와 상기 생성된 첨삭 후보군 각각에 대한 신뢰도 분포를 예측하는 예측부; 및상기 예측된 신뢰도 분포의 평균과 표준편차에 기반한 복수 증강 방식을 통해 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 증강부를 포함하는 문서 신뢰도 증강 시스템.
- 제6항에 있어서,상기 입력된 문서에 대해 신뢰도 증강된 신뢰도 증강 문서 및 상기 신뢰도 증강의 전과 후에 대한 신뢰도 분포 변화 정보를 상기 사용자에게 제공하는 출력부를 더 포함하는 것을 특징으로 하는 문서 신뢰도 증강 시스템.
- 제7항에 있어서,상기 출력부는상기 사용자와의 상호작용을 통해 상기 사용자가 희망하는 신뢰도 분포의 양태에 가장 유사한 문서 첨삭 결과를 상기 사용자에게 더 제공하는 것을 특징으로 하는 문서 신뢰도 증강 시스템.
- 제6항에 있어서,상기 복수 증강 방식은상기 신뢰도 분포의 미리 설정된 기준 평균 이상의 평균을 지향하는 방식의 신뢰도 증강에 해당하는 제1 항목, 상기 신뢰도 분포의 미리 설정된 기준 표준편차 이하의 표준편차를 지향하는 방식의 신뢰도 증강에 해당하는 제2 항목 및 상기 신뢰도 분포의 미리 설정된 기준 {평균/표준편차} 이상의 {평균/표준편차}를 지향하는 방식의 신뢰도 증강에 해당하는 제3 항목 중 적어도 하나 이상의 항목을 포함하는 것을 특징으로 하는 문서 신뢰도 증강 시스템.
- 제6항에 있어서,상기 생성부는상기 사용자로부터 상기 복수 증강 방식에 대한 제어 정보를 입력 받고, 상기 입력된 문서의 각 문장에 대한 전처리를 수행하며, 상기 입력된 문서의 각 문장에 대한 복수의 첨삭 후보 문장 집합을 생성하고,상기 예측부는상기 입력된 문서의 각 문장과 상기 첨삭 후보 문장 집합 내의 각 문장을 통해 조합 가능한 복수의 문서 집합의 각 문장에 대한 신뢰도 분포를 예측하며,상기 증강부는상기 예측된 복수의 문서 집합의 각 문장에 대한 신뢰도 분포와 상기 제어 정보에 기초하여 신뢰도 증강 문서를 선별함으로써, 상기 입력된 문서에 대한 신뢰도 증강을 수행하는 것을 특징으로 하는 문서 신뢰도 증강 시스템.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/297,627 US20220004717A1 (en) | 2018-11-30 | 2019-09-30 | Method and system for enhancing document reliability to enable given document to receive higher reliability from reader |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2018-0151722 | 2018-11-30 | ||
| KR1020180151722A KR101983517B1 (ko) | 2018-11-30 | 2018-11-30 | 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020111489A1 true WO2020111489A1 (ko) | 2020-06-04 |
Family
ID=66672351
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2019/012737 Ceased WO2020111489A1 (ko) | 2018-11-30 | 2019-09-30 | 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20220004717A1 (ko) |
| KR (1) | KR101983517B1 (ko) |
| WO (1) | WO2020111489A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20250148557A1 (en) * | 2023-11-06 | 2025-05-08 | Fevr Llc | Real estate listing evaluation engine |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR101983517B1 (ko) * | 2018-11-30 | 2019-05-29 | 한국과학기술원 | 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 |
| US12387720B2 (en) | 2020-11-20 | 2025-08-12 | SoundHound AI IP, LLC. | Neural sentence generator for virtual assistants |
| US12223948B2 (en) | 2022-02-03 | 2025-02-11 | Soundhound, Inc. | Token confidence scores for automatic speech recognition |
| US12394411B2 (en) * | 2022-10-27 | 2025-08-19 | SoundHound AI IP, LLC. | Domain specific neural sentence generator for multi-domain virtual assistants |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20080021771A (ko) * | 2005-06-15 | 2008-03-07 | 각코호진 와세다다이가쿠 | 문장 평가 장치 및 문장 평가 프로그램 |
| KR20090017830A (ko) * | 2007-08-16 | 2009-02-19 | 한국과학기술원 | 신뢰도를 향상시킨 문서 구조 기반 군집 장치 및 방법 |
| US20170177563A1 (en) * | 2010-09-24 | 2017-06-22 | National University Of Singapore | Methods and systems for automated text correction |
| KR101752679B1 (ko) * | 2015-11-18 | 2017-09-26 | 라이팅박스(주) | 자연어 처리를 이용한 온라인 첨삭 방법 및 그 장치 |
| KR101983517B1 (ko) * | 2018-11-30 | 2019-05-29 | 한국과학기술원 | 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP5596649B2 (ja) * | 2011-09-26 | 2014-09-24 | 株式会社東芝 | 文書マークアップ支援装置、方法、及びプログラム |
| JP5870790B2 (ja) * | 2012-03-19 | 2016-03-01 | 富士通株式会社 | 文章校正装置、及び文章校正方法 |
| US9245015B2 (en) * | 2013-03-08 | 2016-01-26 | Accenture Global Services Limited | Entity disambiguation in natural language text |
| US20170178528A1 (en) * | 2015-12-16 | 2017-06-22 | Turnitin, Llc | Method and System for Providing Automated Localized Feedback for an Extracted Component of an Electronic Document File |
| US10389677B2 (en) * | 2016-12-23 | 2019-08-20 | International Business Machines Corporation | Analyzing messages in social networks |
| JP6815899B2 (ja) * | 2017-03-02 | 2021-01-20 | 東京都公立大学法人 | 出力文生成装置、出力文生成方法および出力文生成プログラム |
| US10650097B2 (en) * | 2018-09-27 | 2020-05-12 | International Business Machines Corporation | Machine learning from tone analysis in online customer service |
-
2018
- 2018-11-30 KR KR1020180151722A patent/KR101983517B1/ko not_active Expired - Fee Related
-
2019
- 2019-09-30 US US17/297,627 patent/US20220004717A1/en not_active Abandoned
- 2019-09-30 WO PCT/KR2019/012737 patent/WO2020111489A1/ko not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20080021771A (ko) * | 2005-06-15 | 2008-03-07 | 각코호진 와세다다이가쿠 | 문장 평가 장치 및 문장 평가 프로그램 |
| KR20090017830A (ko) * | 2007-08-16 | 2009-02-19 | 한국과학기술원 | 신뢰도를 향상시킨 문서 구조 기반 군집 장치 및 방법 |
| US20170177563A1 (en) * | 2010-09-24 | 2017-06-22 | National University Of Singapore | Methods and systems for automated text correction |
| KR101752679B1 (ko) * | 2015-11-18 | 2017-09-26 | 라이팅박스(주) | 자연어 처리를 이용한 온라인 첨삭 방법 및 그 장치 |
| KR101983517B1 (ko) * | 2018-11-30 | 2019-05-29 | 한국과학기술원 | 주어진 문서가 독자에게 보다 높은 신뢰를 받을 수 있도록 하는 문서 신뢰도 증강 방법 및 그 시스템 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20250148557A1 (en) * | 2023-11-06 | 2025-05-08 | Fevr Llc | Real estate listing evaluation engine |
Also Published As
| Publication number | Publication date |
|---|---|
| KR101983517B1 (ko) | 2019-05-29 |
| US20220004717A1 (en) | 2022-01-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Yan et al. | Practical and ethical challenges of large language models in education: A systematic scoping review | |
| Ji et al. | DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome | |
| Zhang et al. | Ontochat: a framework for conversational ontology engineering using language models | |
| US10904072B2 (en) | System and method for recommending automation solutions for technology infrastructure issues | |
| Casillo et al. | Detecting privacy requirements from User Stories with NLP transfer learning models | |
| Niu et al. | Evaluation of linguistic features useful in extraction of interactions from PubMed; application to annotating known, high-throughput and predicted interactions in I2D | |
| US8301628B2 (en) | Predictive analytic method and apparatus | |
| Araki et al. | Joint event trigger identification and event coreference resolution with structured perceptron | |
| Mirza et al. | Classifying temporal relations with simple features | |
| Sánchez-Vega et al. | Paraphrase plagiarism identification with character-level features | |
| US20220004717A1 (en) | Method and system for enhancing document reliability to enable given document to receive higher reliability from reader | |
| Kawahara et al. | Rapid development of a corpus with discourse annotations using two-stage crowdsourcing | |
| Cong | Manner implicatures in large language models | |
| Ghosh et al. | Deep cascaded multitask framework for detection of temporal orientation, sentiment and emotion from suicide notes | |
| Alotaibi et al. | Weakly supervised deep learning for arabic tweet sentiment analysis on education reforms: Leveraging pre-trained models and llms with snorkel | |
| Esposito et al. | Bridging auditory perception and natural language processing with semantically informed deep neural networks | |
| Yin et al. | Asl stem wiki: Dataset and benchmark for interpreting stem articles | |
| Koehler et al. | Towards intelligent process support for customer service desks: Extracting problem descriptions from noisy and multi-lingual texts | |
| Goel et al. | Social media analysis: A tool for popularity prediction using machine learning classifiers | |
| Marulli et al. | Tuning syntaxnet for pos tagging italian sentences | |
| Zhao et al. | LLM-powered Topic Modeling for Discovering Public Mental Health Trends in Social Media | |
| Yan et al. | Collaborative stance detection via small-large language model consistency verification | |
| Kwak et al. | Transferring Legal Natural Language Inference Model from a US State to Another: What Makes It So Hard? | |
| Chaid et al. | Comparative analysis of innovative machine learning algorithms: Advancements in natural language processing | |
| Liebold et al. | Transcription factor prediction using protein 3D secondary structures |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19889940 Country of ref document: EP Kind code of ref document: A1 |