WO2022166808A1 - 文本还原方法、装置及电子设备 - Google Patents
文本还原方法、装置及电子设备 Download PDFInfo
- Publication number
- WO2022166808A1 WO2022166808A1 PCT/CN2022/074583 CN2022074583W WO2022166808A1 WO 2022166808 A1 WO2022166808 A1 WO 2022166808A1 CN 2022074583 W CN2022074583 W CN 2022074583W WO 2022166808 A1 WO2022166808 A1 WO 2022166808A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target
- candidate word
- character group
- text
- character
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
Definitions
- the present application belongs to the technical field of language recognition, and in particular relates to a text restoration method, device and electronic device.
- the Western character group (such as English character group) at the end of a certain line of the text cannot be displayed in the whole line
- the Western character group can be disconnected from the position of automatic line wrapping, And add a separator at the position of the line break, such as mark 1, mark 2, mark 3, mark 4, mark 5, mark 6 in Figure 1.
- the character groups can be automatically restored based on these delimiters.
- the separator at the end of the text line can be directly removed, so that the character groups before and after the separator form a character group, which is displayed in the copied text.
- the text shown in Figure 2 is the text shown in Figure 1 The resulting text after copying.
- the purpose of the embodiments of the present application is to provide a text restoration method, device and electronic device, which can solve the problem of inaccurate text restoration by existing electronic devices.
- an embodiment of the present application provides a method for text restoration, the method comprising: obtaining a first candidate word and a second candidate word according to a first character group, where the first character group is in the target text to be restored The character group at the end of the Nth line and ending with a separator, the first candidate word is a word obtained by combining the first character group and the second character group, and the second candidate word is a combination of the third character group and the second character group.
- the second character group is the first character group of the N+1th line in the target text to be restored, and the third character group is the character group obtained after removing the separator from the first character group; determine the first perplexity degree and the second perplexity degree, the first perplexity degree is the perplexity degree corresponding to the first sentence obtained by replacing the first character group and the second character group in the target sentence with the first candidate word, and the second perplexity degree is the second candidate word replacement target.
- an embodiment of the present application provides a text restoration device, where the text restoration device includes an acquisition module, a determination module, and a restoration module.
- the acquisition module is used to acquire the first candidate word and the second candidate word according to the first character group, where the first character group is the character group at the end of the Nth line in the target text to be restored and ending with a separator,
- the first candidate word is the word obtained by combining the first character group and the second character group
- the second candidate word is the word obtained by combining the third character group and the second character group
- the second character group is the target text to be restored.
- the first character group in the N+1 line, the third character group is the character group obtained by removing the separator from the first character group; the determination module is used to determine the first degree of confusion and the second degree of confusion, the first degree of confusion The confusion degree corresponding to the first sentence obtained by replacing the first character group and the second character group in the target sentence for the first candidate word, and the second confusion degree is the second candidate word replacing the first character group and the second character group in the target sentence.
- the perplexity degree corresponding to the second sentence obtained by the character group; the restoration module is used to obtain the restored target text according to the first candidate word when the first perplexity degree is less than the second perplexity degree; or in the second perplexity degree In the case of less than the first perplexity degree, the restored target text is obtained according to the second candidate word.
- an embodiment of the present application provides an electronic device, the electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor.
- the program or instruction is executed by the processor, Implement the steps of the text restoration method as in the first aspect above.
- an embodiment of the present application provides a readable storage medium, on which a program or an instruction is stored, and when the program or instruction is executed by a processor, the text restoration method as described in the first aspect above is implemented. step.
- an embodiment of the present application provides a chip, the chip includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run a program or an instruction to implement the text restoration method in the first aspect above. step.
- the first candidate word and the second candidate word may be obtained according to the first character group, where the first character group is the line at the end of the Nth line in the target text to be restored and ends with a delimiter Character group, the first candidate word is the word obtained by combining the first character group and the second character group, the second candidate word is the word obtained by combining the third character group and the second character group, and the second character group is the target to be restored
- the first character group of the N+1th line in the text, and the third character group is the character group obtained by removing the separator from the first character group; determine the first perplexity degree and the second perplexity degree, and the first perplexity degree is the first perplexity degree.
- the restored target text is obtained.
- the confusion degree corresponding to the first sentence and the basis of The confusion degree corresponding to the second sentence obtained by the second candidate word can determine which of the first candidate word and the second candidate word is correct, that is, it can be determined that the correct combination of the first character group and the second character group in the target text is words, so that the text can be accurately restored.
- FIG. 1 is a schematic diagram of a text to be restored provided by an embodiment of the present application.
- FIG. 2 is a schematic diagram of a restored text provided by an embodiment of the present application.
- FIG. 3 is a schematic flowchart of a text restoration method provided by an embodiment of the present application.
- FIG. 4 is a schematic structural diagram of a text restoration device provided by an embodiment of the present application.
- FIG. 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
- FIG. 6 is a schematic hardware diagram of an electronic device provided by an embodiment of the present application.
- first, second and the like in the description and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It is to be understood that the data so used are interchangeable under appropriate circumstances so that the embodiments of the present application can be practiced in sequences other than those illustrated or described herein, and distinguish between “first”, “second”, etc.
- the objects are usually of one type, and the number of objects is not limited.
- the first object may be one or more than one.
- “and/or” in the description and claims indicates at least one of the connected objects, and the character “/" generally indicates that the associated objects are in an "or” relationship.
- words such as “exemplary” or “for example” are used to represent examples, illustrations or illustrations. Any embodiments or designs described in the embodiments of the present application as “exemplary” or “such as” should not be construed as preferred or advantageous over other embodiments or designs. Rather, the use of words such as “exemplary” or “such as” is intended to present the related concepts in a specific manner.
- an embodiment of the present application provides a text restoration method, and the method includes the following steps 201 - 204 , or steps 201 - 203 and 205 .
- the execution body of the text restoration method provided by the embodiment of the present application may be a text restoration apparatus, or a control module in the text restoration apparatus for executing the text restoration method, or an electronic device.
- the text restoration method provided by the embodiments of the present application will be exemplarily described below by taking a text restoration apparatus as an example.
- the electronic device when the execution body of the text restoration method provided by the embodiment of the present application is an electronic device, the electronic device may include the text restoration apparatus provided in the embodiment of the present application, or externally connect the text restoration apparatus. Specifically, it can be determined according to actual use requirements, and is not limited in the embodiments of the present application.
- Step 201 the electronic device obtains the first candidate word and the second candidate word according to the first character group.
- the above-mentioned first character group may be a character group at the end of the Nth line in the target text to be restored and ending with a separator
- the first candidate word is a word obtained by combining the first character group and the second character group
- the second candidate word is the word obtained by combining the third character group and the second character group
- the second character group is the first character group of the N+1th line in the target text to be restored
- the third character group is the A character group obtained after removing the separator
- N is a positive integer.
- the electronic device may acquire the first candidate word and the second candidate word according to the first character group, so that the first candidate word and the second candidate word can be The correct word in the two candidate words restores the target text.
- the text restoration method provided in this embodiment of the present application may be applied to the following two possible scenarios:
- Scenario 1 The electronic device copies the target text from one location to another, for example, copying the target text from one document to another document.
- Scenario 2 The target text is the text in the target image, and the electronic device recognizes the text in the target image through optical character recognition (OCR) technology.
- OCR optical character recognition
- the text in the target image may be typeset horizontally or vertically.
- the above-mentioned first character group may be the character group at the end of the M-th column in the target text to be restored and ending with a separator
- the second character group is the character group to be restored
- the text restoration method provided by the embodiments of the present application may also be applied to any other possible scenarios, which may be determined according to actual usage requirements, which are not limited in the embodiments of the present application.
- the character group involved in the embodiment of the present application may be a Western character group, such as an English character group, a French character group, a German character group, a Russian character group, or a Portuguese character group, etc. It is confirmed that the embodiments of the present application are not limited. Wherein, the embodiment of the present application is exemplified by taking an English character group as an example.
- the electronic device can detect line by line whether the end of each line of text in the target text ends with a separator or a specific separator (such as "-"), If yes, the electronic device may regard the character group including the separator as the above-mentioned first character group. If not, then the electronic device can proceed to detect the next line of text.
- a separator or a specific separator such as "-"
- the manner in which the electronic device obtains the first candidate word and the second candidate word may be:
- Step 1 The electronic device forms a unit of the character group before the separator at the end of the current line (that is, the above-mentioned third character group), the separator (such as "-"), and the first character group of the next line of the current line. (hereinafter referred to as the processing candidate set).
- the electronic device can obtain the candidate set to be processed ⁇ representa,-,tion ⁇ from the first row, and the candidate set to be processed ⁇ repre,-,sentation ⁇ from the fourth row.
- the candidate set to be processed ⁇ pre,-,train ⁇ is obtained from the sixth line
- the candidate set to be processed ⁇ re,-,sult ⁇ is obtained from the ninth line
- the candidate set to be processed ⁇ fine,-,tuned is obtained from the tenth line ⁇ , get the candidate set ⁇ task,-,specific ⁇ to be processed from the fourteenth line.
- Step 2 For all candidates in each candidate set to be processed, combine the character groups before and after the separator to obtain candidate words, such as ⁇ representation ⁇ , ⁇ representation ⁇ , ⁇ pretrain ⁇ , ⁇ result ⁇ , ⁇ finetuned ⁇ and ⁇ taskspecific ⁇ ⁇ , that is, the first candidate word can be obtained.
- candidate words such as ⁇ representation ⁇ , ⁇ representation ⁇ , ⁇ pretrain ⁇ , ⁇ result ⁇ , ⁇ finetuned ⁇ and ⁇ taskspecific ⁇ ⁇ , that is, the first candidate word can be obtained.
- Step 3 For all candidates in each candidate set to be processed, generate compound words with reserved separators, such as ⁇ representa-tion ⁇ , ⁇ repre-sentation ⁇ , ⁇ pre-train ⁇ , ⁇ re-sult ⁇ , ⁇ fine-tuned ⁇ and ⁇ task-specific ⁇ , that is, the above-mentioned second candidate word can be obtained.
- the second candidate word is a compound word
- the electronic device can use the word after the combination of the characters before and after the separator and the compound word formed by the separator, respectively. detection, so as to ensure the accuracy of the restored target text.
- Step 202 the electronic device determines a first degree of confusion and a second degree of confusion.
- the first degree of confusion may be the degree of confusion corresponding to the first sentence obtained by replacing the first character group and the second character group in the target sentence with the first candidate word
- the second degree of confusion is the degree of confusion in the target sentence that is replaced by the second candidate word.
- the perplexity corresponding to the second sentence obtained from the first character group and the second character group.
- the electronic device after the electronic device acquires the first candidate word and the second candidate word, the electronic device can determine the first degree of confusion and the second degree of confusion, so that the first candidate word and the second candidate word can be obtained from the first candidate word and the second candidate word. to determine the correct words to restore the above target text.
- the electronic device may perform the following steps 202a and 202b respectively on the above-mentioned first candidate word and the second candidate word, so as to determine the above-mentioned first degree of confusion and the first degree of confusion. Second confusion.
- steps 202a and 202b are exemplified by one candidate word (for example, the target candidate word in the embodiment of the present application) among the above-mentioned first candidate word and second candidate word.
- Step 202a the electronic device determines the target parameter based on the probability that each character in the target candidate word appears in the target text.
- the target candidate word may be the first candidate word or the second candidate word.
- Step 202b the electronic device determines the confusion degree corresponding to the target candidate word according to the target parameter.
- the above-mentioned target parameters may include: the legitimacy value of the target candidate word, the fluency value of the target phrase, and the fluency value of the target sentence.
- the target phrase may include target candidate words, a fourth character group and a fifth character group, the fourth character group may be the character group located before the first character group in the target text, and the fifth character group is the second character group in the target text. character group after.
- the electronic device may determine the legitimacy value of the target candidate word and the fluency of the target phrase based on the probability that each character in the target candidate word (the first candidate word or the second candidate word) appears in the target text
- the degree value and the fluency value of the target sentence can be obtained, so that the above target parameter can be obtained, and then the electronic device can determine the degree of confusion corresponding to the target candidate word (for example, the above-mentioned first degree of confusion or the second degree of confusion) according to the target parameter.
- the electronic device can input the target candidate word (the first candidate word or the second candidate word) and the target text into the language model, and then the language model can calculate the validity value of the target candidate word and the fluency value of the target phrase and the calculation of the fluency value of the target sentence, so that the above target parameters can be obtained.
- the validity value of the above target candidate word may be the probability of the target candidate word appearing in the target text in the target text (denoted as Score_1).
- the validity value of the target candidate word may be the product of the probabilities that each character in the target candidate word appears in the target text.
- the probability that the Kth character in the target candidate word appears in the target text refers to the probability that the Kth character appears when the sixth character group appears in the target text, and the sixth character group is determined by the target candidate word. It consists of the first character to the (K-1)th character of , where K is an integer greater than 1.
- the validity value of the target candidate word can be expressed as:
- P(W) represents the validity value of the target candidate word
- p(C 1 ) represents the probability that the first character in the target candidate word appears in the target text
- C 1 ,C 2 ,... C K-1 ) represents the product of the probabilities of the K-th character appearing when the sixth character group appears in the target text
- the sixth character group consists of the 1st character to the (K-1)th character in the target candidate word composed of characters.
- the language model is shown in the following formula (1), W represents a candidate word, C 1 represents the first character in the candidate word, C k represents the last character in the candidate word, Determine whether W is a valid word by calculating the probability that the candidate word W is composed of characters from C 1 to C k .
- the probability formula for calculating the word is shown in the following formula (2), where p(C 1 ) represents the probability that the character C 1 appears in the target text, and the calculation formula is shown in the following formula (3).
- C 1 represents character r
- C 1 ) indicates that the occurrence of C 2 is related to C 1 , that is, the probability that C 2 appears under the condition that C 1 appears.
- C 1 represents the character "w”
- C 2 represents the character “e”
- the probability that the character "e” appears under the condition that the character "w” appears is: P(e
- w) P( we)/P(w).
- the fluency value of the target phrase may be the probability that the phrase composed of the target candidate word, the fourth character group and the fifth character group appears in the target text (referred to as Score_2).
- the fluency value of the target phrase can be calculated according to the following formula (5).
- S represents a sentence or phrase consisting of words W 1 ...W N.
- W 1 ...W N the perplexity
- the fluency value of the target sentence may be the probability of the target sentence appearing in the target text (referred to as Score_3).
- the fluency value of the above target sentence may be calculated according to the above formula (5).
- step 202b may be specifically implemented by the following step 202b1.
- Step 202b1 the electronic device obtains the target according to the product of the legitimacy value of the target candidate word and the first coefficient, the product of the fluency value of the target phrase and the second coefficient, and the product of the fluency value of the target sentence and the third coefficient.
- the perplexity corresponding to the candidate word is the perplexity corresponding to the candidate word.
- the sum of the first coefficient, the second coefficient and the third coefficient is equal to 1.
- the electronic device can calculate the legality value of the target candidate word and the first coefficient. (denoted as ⁇ ), the sum of the product of the fluency value of the target phrase and the second coefficient (denoted as ⁇ ), and the sum of the product of the fluency value of the target sentence and the third coefficient (denoted as ⁇ ), so that the target can be obtained.
- the perplexity corresponding to the candidate word (denoted as Score).
- Score ⁇ Score_1+ ⁇ Score_2+ ⁇ Score_3.
- the values of the first coefficient, the second coefficient and the third coefficient may be any possible positive numbers, and the sum of the first coefficient, the second coefficient and the third coefficient is equal to 1.
- Step 203 the electronic device determines whether the first degree of confusion is less than the second degree of confusion.
- the electronic device may compare the magnitudes of the first degree of confusion and the second degree of confusion. Thereby, it is determined which candidate word among the above-mentioned first candidate word and second candidate word is correct.
- the electronic device can obtain the restored target text according to the first candidate word, that is, when the first degree of confusion is less than the second degree of confusion, the electronic device can obtain the restored target text according to the first candidate word.
- the device may perform step 204 described below. If the second degree of confusion is less than the first degree of confusion, the electronic device can obtain the restored target text according to the second candidate word, that is, when the second degree of confusion is less than the first degree of confusion, the electronic device can perform the following Step 205.
- Step 204 the electronic device obtains the restored target text according to the first candidate word.
- the electronic device when the first degree of confusion is less than the second degree of confusion, the electronic device can restore the target text according to the first candidate word, so that the restored target text can be obtained.
- the electronic device can directly use the first candidate word to replace the first character group and the second character group in the target text, so that the restored target text can be obtained.
- the electronic device can use the above-mentioned first sentence (the first sentence includes the first candidate word) to replace the target sentence in the target text, so that the restored target text can be obtained.
- Step 205 the electronic device obtains the restored target text according to the second candidate word.
- the electronic device when the second degree of confusion is less than the first degree of confusion, the electronic device can restore the target text according to the second candidate word, so that the restored target text can be obtained.
- the electronic device can directly use the second candidate word to replace the first character group and the second character group in the target text, so that the restored target text can be obtained.
- the electronic device can use the above-mentioned second sentence (the second sentence includes the second candidate word) to replace the target sentence in the target text, so that the restored target text can be obtained.
- the smoother the sentence is that is, the smaller the degree of confusion corresponding to the sentence is, the more accurate the sentence is. Therefore, by comparing the first sentence obtained according to the first candidate word
- the corresponding perplexity degree and the perplexity degree corresponding to the second sentence obtained from the second candidate word can determine which of the first candidate word and the second candidate word is correct, that is, the first character group and the second word group in the target text can be determined.
- the correct word composed of character groups, so that the text can be accurately restored.
- the text restoration method provided by the embodiment of the present application may further include the following step 206 .
- Step 206 the electronic device acquires the keywords of the restored target text based on the keyword recognition model.
- the content type of the above keyword may be the same as the content type preset in the keyword identification model.
- the electronic device may input the restored target text into the keyword recognition model, so that the key words in the restored target text can be obtained based on the keyword recognition model. In this way, accurate keywords can be obtained, and then the accuracy of keyword recognition can be improved.
- the keyword identification model may output a keyword list to the electronic device.
- the keyword list may include all keywords in the restored target text.
- the keyword recognition model can Extract and output all words related to "place name” from the text, so as to obtain the above keywords.
- the keyword recognition model can perform keyword recognition on the restored target text, so as to obtain the keywords in the restored target text , and output the list of these keywords to the electronic device, so that the keywords in the target text can be accurately obtained.
- the text restoration apparatus provided by the embodiment of the present application will be described below by taking the text restoration method performed by the text restoration apparatus in the embodiment of the present application as an example.
- an embodiment of the present application provides a text restoration apparatus 300 .
- the text restoration apparatus 300 includes an acquisition module 301 , a determination module 302 and a restoration module 303 .
- the obtaining module 301 is used to obtain the first candidate word and the second candidate word according to the first character group, where the first character group is the character group at the end of the Nth line in the target text to be restored and ending with a delimiter , the first candidate word is the word obtained by combining the first character group and the second character group, the second candidate word is the word obtained by combining the third character group and the second character group, and the second character group is the target text to be restored.
- the first character group in the N+1th line of the The perplexity degree is the perplexity degree corresponding to the first sentence obtained by the first candidate word replacing the first character group and the second character group in the target sentence
- the second perplexity degree is the second candidate word replacing the first character group and the second character group in the target sentence.
- the perplexity degree corresponding to the second sentence obtained by the second character group; the restoration module 303 is used to obtain the restored target text according to the first candidate word when the first perplexity degree is less than the second perplexity degree; When the second perplexity degree is less than the first perplexity degree, the restored target text is obtained according to the second candidate word.
- the determination module is specifically configured to perform the following steps respectively on the first candidate word and the second candidate word: based on the probability that each character in the target candidate word appears in the target text, determine the target parameter, and the target candidate word is The first candidate word or the second candidate word; according to the target parameter, determine the degree of confusion corresponding to the target candidate word; wherein, the target parameter includes: the legitimacy value of the target candidate word, the fluency value of the target phrase and the fluency value of the target sentence ;
- the target phrase includes a target candidate word, the fourth character group and the fifth character group, the fourth character group is the character group located before the first character group in the target text, and the fifth character group is located in the target text after the second character group character group.
- the determination module is specifically used for the product of the validity value of the target candidate word and the first coefficient, the product of the fluency value of the target phrase and the second coefficient, and the product of the fluency value of the target sentence and the third coefficient.
- the sum is obtained to obtain the perplexity corresponding to the target candidate word; wherein, the sum of the first coefficient, the second coefficient and the third coefficient is equal to 1.
- the legitimacy value of the target candidate word is the probability that the target candidate word appears in the target text in the target text
- the fluency value of the target phrase is the phrase consisting of the target candidate word, the fourth character group and the fifth character group.
- the probability of appearing in the target text; the fluency value of the target sentence is the probability that the target sentence appears in the target text.
- the legitimacy value of the target candidate word is the product of the probabilities that each character in the target candidate word appears in the target text; wherein, the probability that the Kth character in the target candidate word appears in the target text Refers to: the probability that the Kth character appears when the sixth character group appears in the target text.
- the sixth character group is composed of the first character to the (K-1)th character in the target candidate word, and K is Integer greater than 1.
- the determining module is further configured to obtain the restored keyword of the target text based on the keyword recognition model, where the content type of the keyword is the same as the preset content type in the keyword recognition model.
- An embodiment of the present application provides a text restoration device. Since the lower the degree of confusion corresponding to the sentence, the smoother the sentence is, that is, the smaller the degree of confusion corresponding to the sentence, the more accurate the sentence. Therefore, by comparing the first candidate word obtained by comparing the first candidate word The confusion degree corresponding to the sentence and the confusion degree corresponding to the second sentence obtained from the second candidate word can determine which of the first candidate word and the second candidate word is correct, that is, the first character group and the first character group in the target text can be determined. The correct word composed of two-character groups, so that the text can be accurately restored.
- the text restoration apparatus in this embodiment of the present application may be an apparatus, or may be a component, an integrated circuit, or a chip in an electronic device.
- the apparatus may be a mobile electronic device or a non-mobile electronic device.
- the mobile electronic device may be a mobile phone, a tablet computer, a notebook computer, a palmtop computer, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (personal digital assistant).
- UMPC ultra-mobile personal computer
- netbook or a personal digital assistant (personal digital assistant). assistant, PDA), etc.
- the non-mobile electronic device may be a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which are not specifically limited in the embodiments of the present application.
- the text restoration apparatus in this embodiment of the present application may be an apparatus having an operating system.
- the operating system may be an Android (Android) operating system, an ios operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
- the text restoration device provided in the embodiment of the present application can implement each process implemented by the foregoing method embodiment, which is not repeated here to avoid repetition.
- an embodiment of the present application further provides an electronic device 500, including a processor 501, a memory 502, a program or instruction stored in the memory 502 and executable on the processor 501, the program Or, when the instruction is executed by the processor 501, each process of the foregoing text restoration method embodiment can be implemented, and the same technical effect can be achieved. In order to avoid repetition, details are not repeated here.
- the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
- FIG. 6 is a schematic diagram of a hardware structure of an electronic device implementing an embodiment of the present application.
- the electronic device 100 includes but is not limited to: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, etc. part.
- the electronic device 100 may also include a power source (such as a battery) for supplying power to various components, and the power source may be logically connected to the processor 110 through a power management system, so as to manage charging, discharging, and power management through the power management system. consumption management and other functions.
- a power source such as a battery
- the structure of the electronic device shown in FIG. 6 does not constitute a limitation on the electronic device, and the electronic device may include more or less components than those shown in the figure, or combine some components, or arrange different components, which will not be repeated here. .
- the processor 110 may be configured to obtain the first candidate word and the second candidate word according to the first character group, where the first character group is at the end of the Nth line in the target text to be restored and ends with a delimiter character group, the first candidate word is the word obtained by combining the first character group and the second character group, the second candidate word is the word obtained by combining the third character group and the second character group, and the second character group is the word to be restored
- the first character group of the N+1th line in the target text, and the third character group is the character group obtained by removing the separator from the first character group; and determine the first perplexity degree and the second perplexity degree, the first perplexity degree
- the confusion degree corresponding to the first sentence obtained by replacing the first character group and the second character group in the target sentence for the first candidate word, and the second confusion degree is the second candidate word replacing the first character group and the second character group in the target sentence.
- the confusion degree corresponding to the second sentence obtained by the character group and when the first confusion degree is less than the second confusion degree, obtain the restored target text according to the first candidate word; or when the second confusion degree is smaller than the first confusion degree In the case of degree, the restored target text is obtained according to the second candidate word.
- the processor 110 is specifically configured to perform the following steps on the first candidate word and the second candidate word respectively: based on the probability that each character in the target candidate word appears in the target text, determine the target parameter, the target candidate word is the first candidate word or the second candidate word; according to the target parameter, determine the degree of confusion corresponding to the target candidate word; wherein, the target parameter includes: the legitimacy value of the target candidate word, the fluency value of the target phrase and the fluency of the target sentence value; the target phrase includes the target candidate word, the fourth character group and the fifth character group, the fourth character group is the character group located before the first character group in the target text, and the fifth character group is the second character group in the target text. character group after.
- the processor 110 is specifically configured according to the product of the legitimacy value of the target candidate word and the first coefficient, the product of the fluency value of the target phrase and the second coefficient, and the fluency value of the target sentence and the third coefficient.
- the sum of the products is used to obtain the perplexity corresponding to the target candidate word; wherein, the sum of the first coefficient, the second coefficient and the third coefficient is equal to 1.
- the legitimacy value of the target candidate word is the probability that the target candidate word appears in the target text in the target text
- the fluency value of the target phrase is the phrase consisting of the target candidate word, the fourth character group and the fifth character group.
- the probability of appearing in the target text; the fluency value of the target sentence is the probability that the target sentence appears in the target text.
- the legitimacy value of the target candidate word is the product of the probabilities that each character in the target candidate word appears in the target text; wherein, the probability that the Kth character in the target candidate word appears in the target text Refers to: the probability that the Kth character appears when the sixth character group appears in the target text.
- the sixth character group is composed of the first character to the (K-1)th character in the target candidate word, and K is Integer greater than 1.
- the processor 110 is further configured to acquire, based on the keyword recognition model, the keyword of the restored target text, where the content type of the keyword is the same as the preset content type in the keyword recognition model.
- the embodiment of the present application provides an electronic device. Since the lower the degree of confusion corresponding to the sentence, the smoother the sentence is, that is, the smaller the degree of confusion corresponding to the sentence is, the more accurate the sentence is. Therefore, by comparing the first sentence obtained according to the first candidate word The corresponding perplexity degree and the perplexity degree corresponding to the second sentence obtained from the second candidate word can determine which of the first candidate word and the second candidate word is correct, that is, the first character group and the second word group in the target text can be determined. The correct word composed of the character group, so that the text can be accurately restored.
- the acquisition module, the determination module, the restoration module, and the input module in the above-mentioned text restoration apparatus may all be implemented by the above-mentioned processor 110 .
- the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, and the like.
- the electronic device provides the user with wireless broadband Internet access through the network module 102, such as helping the user to send and receive emails, browse web pages, and access streaming media.
- the audio output unit 103 may include a speaker, a buzzer, a receiver, and the like.
- the input unit 104 may include a graphics processor (Graphics Processing Unit, GPU) 1041 and a microphone 1042, and the graphics processor 1041 captures images of still pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode data is processed.
- the display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light emitting diode, or the like.
- the user input unit 107 includes a touch panel 1071 and other input devices 1072 .
- the touch panel 1071 is also called a touch screen.
- the touch panel 1071 may include two parts, a touch detection device and a touch controller.
- Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, and joysticks, which will not be described herein again.
- Memory 109 may be used to store software programs as well as various data including, but not limited to, application programs and operating systems.
- the processor 110 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and an application program, and the like, and the modem processor mainly processes wireless communication. It can be understood that, the above-mentioned modulation and demodulation processor may not be integrated into the processor 110 .
- Embodiments of the present application further provide a readable storage medium, where a program or an instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, each process of the foregoing text restoration method embodiment can be achieved, and can achieve the same The technical effect, in order to avoid repetition, will not be repeated here.
- the above-mentioned processor is the processor in the electronic device in the above-mentioned embodiment.
- the readable storage medium may include a computer-readable storage medium, such as a computer read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and the like.
- the embodiment of the present application further provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used for running a program or an instruction to implement each process of the above text restoration method embodiment, and can achieve the same In order to avoid repetition, the technical effect will not be repeated here.
- the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, a system-on-chip, a system-on-a-chip, or a system-on-a-chip, or the like.
- the method of the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is better implementation.
- the technical solution of the present application can be embodied in the form of a software product in essence or in a part that contributes to the prior art, and the computer software product is stored in a storage medium (such as ROM/RAM, magnetic disk, CD-ROM), including several instructions to make an electronic device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
- a storage medium such as ROM/RAM, magnetic disk, CD-ROM
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Machine Translation (AREA)
Abstract
Description
Claims (15)
- 一种文本还原方法,包括:根据第一字符组,获取第一候选词和第二候选词,所述第一字符组为处于待还原的目标文本中的第N行的行末、且以分隔符结尾的字符组,所述第一候选词为所述第一字符组与第二字符组组合得到的词,所述第二候选词为第三字符组与所述第二字符组组合得到的词,所述第二字符组为待还原的所述目标文本中的第N+1行的第一个字符组,所述第三字符组为所述第一字符组除去所述分隔符后得到的字符组;确定第一困惑度和第二困惑度,所述第一困惑度为所述第一候选词替换目标语句中的所述第一字符组和所述第二字符组得到的第一语句对应的困惑度,所述第二困惑度为所述第二候选词替换目标语句中的所述第一字符组和所述第二字符组得到的第二语句对应的困惑度;在所述第一困惑度小于所述第二困惑度的情况下,根据所述第一候选词,得到还原后的所述目标文本;或在所述第二困惑度小于所述第一困惑度的情况下,根据所述第二候选词,得到还原后的所述目标文本。
- 根据权利要求1所述的方法,其中,所述确定第一困惑度和第二困惑度,包括:对所述第一候选词和所述第二候选词分别执行以下步骤:基于目标候选词中的每个字符在所述目标文本中出现的概率,确定目标参数,所述目标候选词为所述第一候选词或所述第二候选词;根据所述目标参数,确定所述目标候选词对应的困惑度;其中,所述目标参数包括:目标候选词的合法性值、目标词组的流畅度值和所述目标语句的流畅度值;所述目标词组包括所述目标候选词、第四字符组和第五字符组,所述第四字符组为所述目标文本中位于所述第一字符组之前的字符组,所述第五字符组为所述目标文本中位于所述第二字符组之后的字符组。
- 根据权利要求2所述的方法,其中,所述根据所述目标参数,确定所述 目标候选词对应的困惑度,包括:根据所述目标候选词的合法性值与第一系数的乘积、所述目标词组的流畅度值与第二系数的乘积、所述目标语句的流畅度值与第三系数的乘积之和,得到所述目标候选词对应的困惑度;其中,所述第一系数、第二系数和第三系数之和等于1。
- 根据权利要求2或3所述的方法,其中,所述目标候选词的合法性值为所述目标文本中,所述目标候选词在所述目标文本中出现的概率;所述目标词组的流畅度值为所述目标候选词、所述第四字符组和所述第五字符组组成的词组在所述目标文本中出现的概率;所述目标语句的流畅度值为所述目标语句在所述目标文本中出现的概率。
- 根据权利要求4所述的方法,其中,所述目标候选词的合法性值为目标候选词中的每个字符在所述目标文本中出现的概率之间的乘积;其中,所述目标候选词中的第K个字符在所述目标文本中出现的概率是指:在所述目标文本中出现第六字符组的情况下出现第K个字符的概率,所述第六字符组由所述目标候选词中的第1个字符至第(K-1)个字符组成,K为大于1的整数。
- 一种文本还原装置,包括获取模块,确定模块和还原模块;获取模块,用于根据第一字符组,获取第一候选词和第二候选词,所述第一字符组为处于待还原的目标文本中的第N行的行末、且以分隔符结尾的字符组,所述第一候选词为所述第一字符组与第二字符组组合得到的词,所述第二候选词为第三字符组与所述第二字符组组合得到的词,所述第二字符组为待还原的所述目标文本中的第N+1行的第一个字符组,所述第三字符组为所述第一字符组除去所述分隔符后得到的字符组;确定模块,用于确定第一困惑度和第二困惑度,所述第一困惑度为所述第一候选词替换目标语句中的所述第一字符组和所述第二字符组得到的第一语句对应的困惑度,所述第二困惑度为所述第二候选词替换目标语句中的所述第一字符组和所述第二字符组得到的第二语句对应的困惑度;还原模块,用于在所述第一困惑度小于所述第二困惑度的情况下,根据所述第一候选词,得到还原后的所述目标文本;或在所述第二困惑度小于所述第一困惑度的情况下,根据所述第二候选词,得到还原后的所述目标文本。
- 根据权利要求6所述的装置,其中,所述确定模块,具体用于对所述第一候选词和所述第二候选词分别执行以下步骤:基于目标候选词中的每个字符在所述目标文本中出现的概率,确定目标参数,所述目标候选词为所述第一候选词或所述第二候选词;根据所述目标参数,确定所述目标候选词对应的困惑度;其中,所述目标参数包括:目标候选词的合法性值、目标词组的流畅度值和所述目标语句的流畅度值;所述目标词组包括所述目标候选词、第四字符组和第五字符组,所述第四字符组为所述目标文本中位于所述第一字符组之前的字符组,所述第五字符组为所述目标文本中位于所述第二字符组之后的字符组。
- 根据权利要求7所述的装置,其中,所述确定模块,具体用于根据所述目标候选词的合法性值与第一系数的乘积、所述目标词组的流畅度值与第二系数的乘积、所述目标语句的流畅度值与第三系数的乘积之和,得到所述目标候选词对应的困惑度;其中,所述第一系数、第二系数和第三系数之和等于1。
- 根据权利要求7或8所述的装置,其中,所述目标候选词的合法性值为所述目标文本中,所述目标候选词在所述目标文本中出现的概率;所述目标词组的流畅度值为所述目标候选词、所述第四字符组和所述第五字符组组成的词组在所述目标文本中出现的概率;所述目标语句的流畅度值为所述目标语句在所述目标文本中出现的概率。
- 根据权利要求9所述的装置,其中,所述目标候选词的合法性值为目标候选词中的每个字符在所述目标文本中出现的概率之间的乘积;其中,所述目标候选词中的第K个字符在所述目标文本中出现的概率是指:在所述目标文本中出现第六字符组的情况下出现第K个字符的概率,所述 第六字符组由所述目标候选词中的第1个字符至第(K-1)个字符组成,K为大于1的整数。
- 一种电子设备,包括处理器,存储器及存储在所述存储器上并可在所述处理器上运行的程序或指令,所述程序或所述指令被所述处理器执行时实现如权利要求1-5中任一项所述的文本还原方法的步骤。
- 一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或所述指令被处理器执行时实现如权利要求1-5中任一项所述的文本还原方法的步骤。
- 一种计算机程序产品,所述计算机程序产品被至少一个处理器执行以实现如权利要求1-5中任一项所述的文本还原方法的步骤。
- 一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如权利要求1-5任一项所述的文本还原方法的步骤。
- 一种电子设备,所述电子设备被配置成用于执行如权利要求1-5任一项所述的文本还原方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202110158872.0 | 2021-02-04 | ||
| CN202110158872.0A CN112949261A (zh) | 2021-02-04 | 2021-02-04 | 文本还原方法、装置及电子设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022166808A1 true WO2022166808A1 (zh) | 2022-08-11 |
Family
ID=76244023
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/074583 Ceased WO2022166808A1 (zh) | 2021-02-04 | 2022-01-28 | 文本还原方法、装置及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112949261A (zh) |
| WO (1) | WO2022166808A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115690806A (zh) * | 2022-10-11 | 2023-02-03 | 杭州瑞成信息技术股份有限公司 | 一种基于图像数据处理的非结构化文档格式识别方法 |
| CN116127949A (zh) * | 2022-11-18 | 2023-05-16 | 马上消费金融股份有限公司 | 语言发展的影响词确定方法、语言发展过程的展示方法 |
| CN116882391A (zh) * | 2023-07-18 | 2023-10-13 | 北京方寸无忧科技发展有限公司 | 一种轻量化的文本纠错方法及装置 |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112949261A (zh) * | 2021-02-04 | 2021-06-11 | 维沃移动通信有限公司 | 文本还原方法、装置及电子设备 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108899016A (zh) * | 2018-08-02 | 2018-11-27 | 科大讯飞股份有限公司 | 一种语音文本规整方法、装置、设备及可读存储介质 |
| US10402490B1 (en) * | 2015-08-14 | 2019-09-03 | Shutterstock, Inc. | Edit distance based spellcheck |
| CN111401004A (zh) * | 2020-03-28 | 2020-07-10 | 苏州机数芯微科技有限公司 | 一种基于机器学习的文章断句方法 |
| CN112949261A (zh) * | 2021-02-04 | 2021-06-11 | 维沃移动通信有限公司 | 文本还原方法、装置及电子设备 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111626065A (zh) * | 2019-02-26 | 2020-09-04 | 株式会社理光 | 神经机器翻译模型的训练方法、装置及存储介质 |
| CN110852087B (zh) * | 2019-09-23 | 2022-02-22 | 腾讯科技(深圳)有限公司 | 中文纠错方法和装置、存储介质及电子装置 |
| CN112269475A (zh) * | 2020-10-23 | 2021-01-26 | 维沃移动通信有限公司 | 文字显示方法、装置和电子设备 |
-
2021
- 2021-02-04 CN CN202110158872.0A patent/CN112949261A/zh active Pending
-
2022
- 2022-01-28 WO PCT/CN2022/074583 patent/WO2022166808A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10402490B1 (en) * | 2015-08-14 | 2019-09-03 | Shutterstock, Inc. | Edit distance based spellcheck |
| CN108899016A (zh) * | 2018-08-02 | 2018-11-27 | 科大讯飞股份有限公司 | 一种语音文本规整方法、装置、设备及可读存储介质 |
| CN111401004A (zh) * | 2020-03-28 | 2020-07-10 | 苏州机数芯微科技有限公司 | 一种基于机器学习的文章断句方法 |
| CN112949261A (zh) * | 2021-02-04 | 2021-06-11 | 维沃移动通信有限公司 | 文本还原方法、装置及电子设备 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115690806A (zh) * | 2022-10-11 | 2023-02-03 | 杭州瑞成信息技术股份有限公司 | 一种基于图像数据处理的非结构化文档格式识别方法 |
| CN116127949A (zh) * | 2022-11-18 | 2023-05-16 | 马上消费金融股份有限公司 | 语言发展的影响词确定方法、语言发展过程的展示方法 |
| CN116882391A (zh) * | 2023-07-18 | 2023-10-13 | 北京方寸无忧科技发展有限公司 | 一种轻量化的文本纠错方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112949261A (zh) | 2021-06-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022166808A1 (zh) | 文本还原方法、装置及电子设备 | |
| EP3881296B1 (en) | Method and apparatus for augmented reality | |
| CN105976818B (zh) | 指令识别的处理方法及装置 | |
| CN112631437B (zh) | 信息推荐方法、装置及电子设备 | |
| US8843493B1 (en) | Document fingerprint | |
| US20170109435A1 (en) | Apparatus and method for searching for information | |
| WO2022199543A1 (zh) | 消息处理方法、装置和电子设备 | |
| CN111880668A (zh) | 输入显示方法、装置及电子设备 | |
| WO2022223002A1 (zh) | 页面显示方法、装置及电子设备 | |
| CN110069624A (zh) | 文本处理方法及装置 | |
| CN112148135A (zh) | 输入法处理方法、装置和电子设备 | |
| CN112231507A (zh) | 识别方法、装置及电子设备 | |
| CN108255939A (zh) | 一种跨语言搜索方法和装置、一种用于跨语言搜索的装置 | |
| CN113359999A (zh) | 候选词更新方法、装置和电子设备 | |
| CN111538830B (zh) | 法条检索方法、装置、计算机设备及存储介质 | |
| CN111858966A (zh) | 知识图谱的更新方法、装置、终端设备及可读存储介质 | |
| RU2608470C2 (ru) | Способ и устройство для обновления пользовательских данных | |
| WO2022105754A1 (zh) | 输入文字的方法、装置和电子设备 | |
| CN111966267A (zh) | 应用评论方法、装置和电子设备 | |
| CN111638831A (zh) | 一种内容融合方法、装置及电子设备 | |
| WO2024179519A1 (zh) | 语义识别方法及其装置 | |
| CN112328096B (zh) | 候选词显示方法、装置和电子设备 | |
| CN113190125B (zh) | 输入矫正方法和装置 | |
| WO2022161307A1 (zh) | 文本翻译方法、装置、设备及介质 | |
| CN111782060B (zh) | 对象显示方法、装置和电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22749087 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22749087 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 240124) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22749087 Country of ref document: EP Kind code of ref document: A1 |