WO2020175806A1 - 글자 인식 장치 및 이에 의한 글자 인식 방법 - Google Patents

글자 인식 장치 및 이에 의한 글자 인식 방법 Download PDF

Info

Publication number
WO2020175806A1
WO2020175806A1 PCT/KR2020/001333 KR2020001333W WO2020175806A1 WO 2020175806 A1 WO2020175806 A1 WO 2020175806A1 KR 2020001333 W KR2020001333 W KR 2020001333W WO 2020175806 A1 WO2020175806 A1 WO 2020175806A1
Authority
WO
WIPO (PCT)
Prior art keywords
character
data
character recognition
score map
input data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2020/001333
Other languages
English (en)
French (fr)
Inventor
백영민
이활석
신승
이영무
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Naver Corp
Original Assignee
Naver Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Naver Corp filed Critical Naver Corp
Priority to JP2021549641A priority Critical patent/JP7297910B2/ja
Publication of WO2020175806A1 publication Critical patent/WO2020175806A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/24Character recognition characterised by the processing or recognition method
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • G06T11/60Creating or editing images; Combining images with text
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/24Aligning, centring, orientation detection or correction of the image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/14Image acquisition
    • G06V30/146Aligning or centring of the image pick-up or image-field

Definitions

  • This disclosure relates to the field of data processing. More specifically, this disclosure relates to a character recognition apparatus and method for recognizing characters in data such as images.
  • a character recognition device and a character recognition method according to an embodiment make it a technical task to accurately and quickly recognize characters in data such as images.
  • a character recognition device and a character recognition method are a technical task to contribute to the development of the fintech industry by accurately recognizing characters within the image of a real card.
  • a character recognition method includes the steps of: inputting input data into a character detection model; acquiring location information of a word area in the input data based on output data output from the character detection model; Extracting partial data corresponding to the acquired location information from the input data; And inputting the partial data into a character recognition model to recognize a character in the partial data.
  • the character recognition device and the character recognition method according to the embodiment can accurately and quickly recognize characters in data such as images.
  • the character recognition device and the character recognition method according to the embodiment can contribute to the development of the fintech industry by accurately recognizing characters within the image of a real card.
  • FIG. 1 is a view showing a character recognition device according to an embodiment.
  • FIG. 2 is a flow chart for explaining a character recognition method according to an embodiment.
  • FIG. 3 is a diagram illustrating a process of recognizing a character through a character recognition device according to an exemplary embodiment.
  • FIG. 4 is an exemplary diagram showing output data output by a character detection model.
  • FIG. 5 is a view for explaining a method of acquiring position information of a word area in input data based on output data output from a character detection model.
  • FIG. 6 is a diagram for explaining the binarization process and the merger process shown in FIG. 5.
  • 7 is a view for explaining the word box determination process shown in FIG. 8 is a view for explaining the structure of a feature extraction model according to an embodiment.
  • 9 is a view for explaining the structure of a character recognition model according to an embodiment.
  • Fig. is a flow chart for explaining a training method of a character detection model according to an embodiment.
  • FIG. 11 is a diagram for explaining a method of generating a first (31 score map).
  • FIG. 12 is a view for explaining a method of generating a second (31 score map).
  • Figure 13 shows a method of determining a connection box between adjacent character boxes
  • FIG. 14 is a block diagram showing the configuration of a character recognition device according to an embodiment.
  • 15 is a server device to which a character recognition device according to an embodiment can be applied, and
  • a character recognition method includes the steps of: inputting input data into a character detection model; acquiring location information of a word region within the input data based on output data output from the character detection model; Extracting partial data corresponding to the acquired location information from the input data; And Recognizing a character in the partial data by inputting the partial data into a character recognition model.
  • a character recognition device is a processor; and at least one
  • the one component When referred to as "connected” or the like, the one component may be directly connected to the other component or may be directly connected, but unless there is a particularly contrary device, it is connected through another component in the middle. Or it should be understood that it can be accessed.
  • the components expressed as' ⁇ sub (unit)','module', etc. are two or more components combined into one component, or one component is divided into a more subdivided function.
  • each of the components to be described below may additionally perform some or all of the functions that other components are responsible for, in addition to the main functions that are in charge of them, and each component is responsible for each component.
  • some of the main functions may be dedicated and performed by other components.
  • characters'letter' may refer to the basic unit of characters constituting a word or sentence.
  • each alphabet may correspond to a letter, and in the case of numbers, '0
  • Each of the numbers' ⁇ '9' can correspond to a letter, and in the case of Korean, a character combined with a consonant and a vowel (for example,'A'), a character combined with a consonant, a vowel and a consonant 'River'), consonants written solely (for example,'ki'), vowels written solely (for example, me') can correspond to this letter.
  • the letter may correspond to a symbol (for example,'/ ','- , etc.) can also be included.
  • 'word' may mean a unit of characters containing at least one letter.
  • the characters constituting the'word' may not be separated by more than a predetermined distance from each other.
  • 'Word' is one word It can also consist of the letters of, for example,'negatives' in English can be made up of a single letter, but may correspond to'words' if they are separated by more than a predetermined distance from the surrounding letters.
  • FIG. 1 is a view showing a character recognition apparatus 100 according to an embodiment.
  • the character recognition apparatus 100 acquires the input data 10 and recognizes the character 50 in the input data 10.
  • the input data 10 is a check card or a credit card. It may include an image taken of a real card such as, or, as described later, a feature output from the feature extraction model 800 based on the image taken of a real card, etc. 11 ]3) may be included.
  • the character recognition device 00) can recognize and store card information such as the card number and expiration date from the input data (10).
  • the card information recognized and stored by the character recognition device 00) is recognized and stored for purchase of goods, etc. It can be used to pay for the bill.
  • Fig. 2 is a flow chart for explaining a character recognition method according to an embodiment
  • Fig. 3 is a view for explaining a process in which a character is recognized through the character recognition device 100 according to an embodiment.
  • step 8210 the character recognition device 100 detects the input data 10
  • the character recognition device 100 may store the character detection model 410 in advance.
  • the character detection model (company 0) may be trained based on the data for learning.
  • step 8220 the character recognition device 100 acquires the location information of the word area in the input data 10 based on the output data 30 output from the character detection model 410.
  • the character recognition device (0) acquires the location information of the word area containing at least one character in the input data (10) based on the output data (30).
  • step 8230 the character recognition device 100 extracts the partial data 40 corresponding to the location information of the word area from the input data 10.
  • the location information of the word area is plural. If acquired, multiple parts corresponding to each location information
  • Data 40 can be extracted from input data 10.
  • step 8240 the character recognition device 100 recognizes the partial data 40
  • the character recognition device 100 By inputting into the model 420, the character 50 included in the partial data 40 is recognized.
  • the character recognition device 100 writes each of the plurality of partial data 40.
  • the recognition model 420 By inputting into the recognition model 420, the characters 50 included in each of the plurality of partial data 40 can be recognized.
  • the character recognition device 100 outputs the character detection model 410
  • Data (30) can also be entered as a character recognition model (420) along with partial data (40).
  • the output data (30) of the character detection model (new 0) is the position of the individual character in the input data (10). 2020/175806 1»(:1 ⁇ 1 ⁇ 2020/001333
  • the accuracy of the character recognition of the character recognition model 420 can be further improved.
  • the character recognition device 100 stores the recognized characters or
  • the output data 30 is a first score map 31 that indicates the probability that a character within the input data 10 will exist on the data space (eg, image space) corresponding to the input data 10, and Input data (10)
  • the connectivity between my characters ( ⁇ : 011116( : 11 ⁇ ) 0 can be included in a second score map (33) that shows up in the data space corresponding to the input data (10).
  • a value (e.g., a pixel value) stored in each location in the first score map 31 may indicate the probability that a character will exist in the input data 10 corresponding to the location.
  • the second score map (33) A value stored in each location (eg, pixel value) can indicate the probability that a plurality of characters will be adjacent to each other within the input data corresponding to the location (10).
  • the size of the first score map (31) and the second score map (33) can be the same as the input data (10) in order to facilitate the calculation of the position correspondence relationship.
  • the character detection model 410 is generated in response to the learning data.
  • the score map and the second (a first score map 31 and a second score map 33 similar to the 31 score map) can be trained to be output.
  • the character recognition device 100 can determine the location information of the word area in the input data 10 based on the first score map 31 and the second score map 33. For this, refer to Figs. 5 to 7 Explain with reference.
  • FIG. 10 is a diagram for explaining a method of acquiring the location information of the word region
  • FIG. 6 is a diagram for explaining the binarization process and the merging process shown in FIG. 5
  • FIG. 7 is a word box determination shown in FIG. This is a drawing to explain the process.
  • steps 8510 and 8520 the character recognition device 100 is within the first score map (31).
  • the first score map (31) is obtained by comparing the data values with the threshold values.
  • Binarization is performed ( 1131 urine 011) , and the second score map 33 is binarized by comparing the data values in the second score map 33 with a threshold value.
  • the character recognition device 00 is the first score map 31
  • the data values in the second score map (33) data values greater than or equal to the threshold value may be changed to the first value, and data values less than the threshold value may be changed to the second value.
  • data having a value greater than or equal to the threshold value in the first score map 31 and the second score map 33 are binarized first score map 601 and binarized second score map.
  • Data having a value less than the threshold value in the first score map (31) and the second score map (33) are changed to have a first value at (603). It can be changed to have a second value in the map 601 and the binarized second score map 603.
  • the threshold for binarization of the first score map (31) and the threshold for binarization of the second score map (33) may be the same or different.
  • step S530 the character recognition device 100 and the binary first score map 601
  • the binary coded second score map 603 is merged.
  • the character recognition device 100 stores data values in the binarized first score map 601 and the binarized second score map 603.
  • the merged map 605 can be created by adding or performing the or operation. For example, as shown in FIG. 6, the binary first score map 601 and the binarized second score map 603 have a first value. Data may be included in the merged map 605. In this way, the merged map 605 can be divided into areas 606 in which characters in the input data 10 are likely to exist and areas that do not.
  • step S540 the character recognition apparatus 100 determines a word box 6W indicating an area including a character by using the merged map 605.
  • the character recognition device W0 may label each of the word regions in order to distinguish the word regions in the merged map 605.
  • the character recognition device 100 is recognized using the merged map 605
  • each of the regions 606 contains the actual word
  • additional checks can be made, specifically, for example, in the merged map 605, with the same (or within the same range) values, adjacently connected regions ( 606) as a word candidate area, and if at least one of the values of the first score map 601 corresponding to each data in the word candidate area exists, the corresponding word candidate area can be determined as the word area.
  • the character recognition apparatus 100 may determine a word box (6 W) of the smallest size including an area of data verified to correspond to the word area.
  • the character recognition device 100 stores the location information of the determined word box 610 (for example, the input data 10 or the position values of the corners of the word box 6W on the merged map 605) in a word area. It can be determined by the location information of the determined word box 610 (for example, the input data 10 or the position values of the corners of the word box 6W on the merged map 605) in a word area. It can be determined by the location information of the determined word box 610 (for example, the input data 10 or the position values of the corners of the word box 6W on the merged map 605) in a word area. It can be determined by the location information of
  • the character recognition device (W0) extracts the partial data 40 corresponding to the location information from the input data 10, and recognizes the extracted partial data 40 as a character. By inputting into the model 420, the character within the partial data 40 can be recognized.
  • the input data 10 input to the character detection model 4W may include a feature map output from the feature detection model 800 based on the original image.
  • FIG. 8 Is a diagram for explaining the structure of the feature detection model 800.
  • the original image 20 can be input as a feature detection model 800.
  • the original image 20 can be input as a feature detection model 800.
  • Image (20) refers to the image input to the feature detection model (800), and does not mean that it is not a copied or modified image of the image taken by the first card, etc.
  • the original image (20) is the first convolution layer (805), the second convolution layer (810), the third
  • Convolutional layer 815 fourth convolutional layer 820, fifth convolutional layer 825, and sixth
  • Convolution processing is performed in the convolutional layer 830.
  • Output task of the sixth convolutional layer 830 5 The convolutional force of the convolutional symptom 825 is concatenated (concatenation) and the first up
  • the values input to the convolution layer 835 and input to the first up-convolution layer 835 are convolutional processing (836), batch normalization (837), convolution processing (838), and batch normalization ( 839) is input to the first up-sampling layer 840.
  • the output of the sampling layer 840 is concatenated with the output of the fourth convolutional layer 820 and processed in the second up-sampling layer 845 and the second up-sampling layer 850.
  • the output of the sampling layer 850 is concatenated with the output of the third convolutional layer 815 to be processed in the third up-convolution layer 855 and the third up-sampling layer 860, and the processing result is the second convolution. It is concatenated with the output of the layer 810 and inputted to the fourth upconvolution layer 865. Then, the result output from the fourth upconvolution layer 865 can be used as the input data 10.
  • the horizontal size and vertical size of the input data 10 may be 1/2 of the horizontal size and vertical size of the original image 20, but are not limited thereto.
  • the structure of the feature detection model 800 shown in Fig. 8 is only one example,
  • the number and processing order of the convolution layer, the up-convolution layer, and the up-sampling layer can be varied in various ways.
  • Figure 9 is for explaining the structure of the character recognition model 420 according to an embodiment
  • the character recognition model 420 uses partial data 40 extracted from the input data 10.
  • the character recognition model 420 includes a convolutional network 421, a recurrent neural network (RNN) 423, and
  • Convolutional network 421 contains at least one convolutional layer
  • the data 40 is convolved to extract a feature map.
  • the convolutional network 421 may include known VGG, ResNet, etc., but in one embodiment, the character recognition model 420 is an original image. Partial data (40) extracted from the feature map of (20) (i.e., input data) can be input, so the required convolution layer The number can be small.
  • RNN (423) is a feature vector from the feature map corresponding to the partial data (40).
  • the RNN 423 can grasp the context relationship of consecutive feature vectors through bi-LSTM (bidirectional long-short-term memory).
  • the decoder 425 extracts a character from the sequence information of feature vectors.
  • the decoder 425 may perform an attention step and a generation step.
  • the decoder 425 calculates a weight indicating whether information is to be extracted from which sequence, and a weight value in the generation step. It can be applied to a sequence, and individual characters can be given through long-short-term memory (LSTM).
  • LSTM long-short-term memory
  • the character recognition device 100 may classify the character groups recognized in each of the plurality of pieces of data 40 according to a predetermined standard.
  • the character recognition device 100 may have any of the character groups. If the character group recognized in the partial data (40) contains a predetermined symbol (for example, /), the character group can be determined by the first type of information. The validity period in the card is divided into year and month. Since it is common to include predefined symbols for the following, the character recognition device 100
  • the character group recognized in data (40) contains a predetermined symbol, the character group can be determined as the expiration date information.
  • the character recognition device 100 uses a character group with a large number (for example, a number located on the right of the symbol) corresponding to the year. If the card contains the validity period and the issuance date, the year included in the validity period will be larger than the year included in the issuance date, so the character recognition device 100 A group of letters with a large number can be determined by the expiration date information.
  • a character group with a large number for example, a number located on the right of the symbol
  • the character recognition apparatus 100 may determine character groups that do not contain a predetermined symbol among character groups recognized in each of the plurality of partial data 40 as the second type of information.
  • the second type of information may include, for example, card number information.
  • the character recognition device 100 includes a plurality of partial data 40
  • the character groups recognized in each can be sorted according to the position of the plurality of partial data 40 in the input data 10. For example, the character recognition device W0 is based on the upper left of the input data 10. Character groups can be sorted using the Z scan method.
  • This character recognition device 100 can determine whether or not re-recognition of characters is necessary based on the number of characters included in a predetermined number of character groups arranged in succession among the arranged character groups. For example, character recognition. If there is a predetermined number of character groups arranged in succession while each of the arranged character groups contains a predetermined number of characters, the character recognition is correctly performed and re-recognition of characters is not required. 2020/175806 1»(:1 ⁇ 1 ⁇ 2020/001333
  • the card number contains 16 numbers
  • the character recognition device 00 does not require re-recognition of characters when 4 character groups including 4 numbers among the aligned character groups are arranged in succession. It can be decided not.
  • the character recognition device 100 is a multi-part data (40)
  • the character recognition device 100 is capable of retaking images.
  • Information that is necessary can be output through a speaker, monitor, etc., or notified to an external device through a network.
  • the character recognition device 100 recognizes characters from the preview image of the camera, re-recognition of characters is required. If it is determined that it is determined, it is also possible to re-recognize the characters from the preview image that is continuously shot through the camera.
  • Fig. is a flow chart for explaining a training method of the character detection model (Y0) according to an embodiment.
  • step 81010 the character recognition device 100 displays the learning data 60, the first ( 31 score map) indicating the probability of the existence of my character in the data space, and the learning data 60 between my characters.
  • a second (31 score map 73) is obtained that indicates the connectivity of the data in the data space.
  • the horizontal and vertical dimensions of the training data 60 may be the same as the horizontal and vertical dimensions of the input data 10.
  • the horizontal and vertical dimensions of the learning data (60) are the same as the horizontal and vertical dimensions of the first ( 31 score map (base)), and the horizontal and vertical dimensions of the second (31' score map (73)) can do.
  • the learning data 60 may include an image photographed of an object such as a card or a feature map extracted based on the image, similar to the original image 20 described above.
  • the character recognition device 100 can directly generate at least one of the first ( 31 score map (Q) and the second (31 score map (73)) from the learning data (60), or It is also possible to receive at least one of the first ( 31 score map (base)) and the second ( 31 score map (73).
  • the values in the first (31 score map (group)) can indicate the probability that a character will be located in the learning data (60) at the point.
  • the values in the second ( 31 score map (73)) are multiple at that point. It can indicate the probability that the letters of the letter will be adjacent to each other.
  • step 81020 the character recognition device 100 detects the text for learning data 60
  • step 81030 in the character detection model (410) in response to the learning data (60)
  • Task 1 31 score map (year) and 2nd score map, respectively
  • the internal weight of the character detection model (4W) can be updated.
  • the loss value can be calculated according to the comparison result of the score map 73.
  • the loss value can correspond to the L2 Loss value, for example.
  • there are various methods such as LI loss and smooth LI loss.
  • the calculated loss value is input to the character detection model (4W), and the character detection model (4W) can update the internal weight value according to the loss value.
  • FIG. 11 is a diagram for explaining a method of generating a first GT score map (base)
  • FIG. 12 is a diagram for explaining a method of generating a second GT score map 73.
  • FIG. It is a drawing for explaining how to determine the connection box 63a between the adjacent character boxes 62a and 62b.
  • the learning data 60 including at least one character
  • the word boxes (61a, 61b, 61c, 61d, 61e) are determined for the word areas, and the word box (61a) according to the number of characters included in the word box (61a, 61b, 61c, 61d, 61e). , 61b, 61c, 61d, 61e) are divided into at least one text box (62a, 62b, 62c, 62d), for example, if four letters are included in one word box, the corresponding word box Can be divided into a total of 4 text boxes.
  • a predetermined image (1100), for example, a 2D Gaussian image can be combined to generate the first GT score map (qi).
  • Connection boxes (63a, 63b, 63c) located on the boundary line (L) between adjacent text boxes are determined, and a predetermined image (1100) in the connection boxes (63a, 63b, 63c), for example, a 2D Gaussian image Can be synthesized to generate a second GT score map 73.
  • connection boxes 63a, 63b, 63c can be determined by connecting a plurality of points set in the inner space of the adjacent character boxes. Specifically, as shown in FIG. 13, the adjacent character boxes ( 62a, 62b) A connection box 63a connecting two points in the left text box 62a and two points in the right text box 62b can be determined.
  • connection box 63a may be determined by connecting the lower left corner and the upper right corner among the corners, connecting the upper left corner and the lower right corner to determine upper and lower triangles, and connecting the midpoints of the corresponding triangles.
  • FIG. 14 is a block diagram showing the configuration of a character recognition apparatus 100 according to an embodiment.
  • the character recognition device 100 may include a memory 1410, a communication module 1430, and a processor 1450.
  • the memory 1410 includes at least one
  • Instructions can be stored, and the processor 1450 can control the character detection and training of the character detection model (4W) according to at least one instruction.
  • the character recognition device 100 may include a plurality of memories and/or a plurality of processors.
  • This memory 1410 can store a character detection model 410 and a character recognition model 420. In addition, the memory 1410 may further store the feature extraction model 800.
  • the processor 1450 inputs the input data 10 to the character detection model 410, and based on the output data output from the character detection model (company 0), the location information of the word area in the input data 10 Then, the processor 1450 inputs partial data corresponding to the acquired location information into the character recognition model 420, and stores the character information output from the character recognition model 420 into the memory 1410 or other Can be saved to storage device.
  • the processor 1450 may train at least one of the character detection model 410, the character recognition model 420, and the feature extraction model 800 based on the learning data 60.
  • the communication module 1430 transmits and receives data to and from an external device through a network.
  • the communication module 1430 transmits/receives an image to and from the external device, or transmits text information recognized in the input data 10. It can transmit and receive external devices.
  • the character recognition device 100 is a server to which the character recognition device 100 according to an embodiment can be applied
  • the character recognition device 100 is implemented as a server device 1510 or a client
  • the server device 1510 may receive an image from the client device 1520, recognize and store the character in the received image.
  • the server device 1510 is a client
  • Character information recognized in the image received from the device 1520 may be transmitted to the client device 1520.
  • the server device 1510 may be transmitted to the client device 1520.
  • the client device 1520 is a character in the image captured by the camera of the client device 1520 or the image stored in the client device 1520 Can be recognized and saved.
  • the client device 1520 is a character detection model (company 0), character recognition
  • Data for execution of at least one of the model 420 and feature extraction model 800 can be received from the server device 1510.
  • the client device 1520 is an image captured through a camera module, an image stored in the internal memory, or An image received from an external device may be input into at least one of a character detection model 410, a character recognition model 420, and a feature extraction model 800 to recognize characters.
  • the client device 1520 receives at least one of a character detection model (4W), a character recognition model (420), and a feature extraction model (800) by receiving learning data from an external device or using the training data stored inside. You can also control your training.
  • a server device 1510 providing data for execution of at least one of the character detection model (4W), the character recognition model 420, and the feature extraction model 800 to the client device 1520 is used for learning data. Based on the character detection model (4W), the character recognition model (420), and the feature extraction model (800), at least one of the training can be controlled. In this case, the server device 1510 can control only the weight information updated as a result of the training. Transmitted to the device 1520, the client device 1520 may update at least one of the character detection model 410, the character recognition model 420, and the feature extraction model 800 according to the received information.
  • the client device 15 is a client device 1520, which shows a desktop PC, but is not limited thereto, and the client device 1520 includes a laptop, a smartphone, a tablet PC, an artificial intelligence (AI) robot, an AI speaker, and a wearable. It may include equipment, etc.
  • AI artificial intelligence
  • the above-described embodiments of the present disclosure can be written as a program that can be executed on a computer, and the created program can be stored on a medium.
  • the medium continues to store, execute, or
  • the medium may be a variety of recording means or storage means in the form of a single or a combination of several pieces of hardware, not limited to media directly connected to any computer system, but distributed over the network.
  • Examples of media include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, and magnetic-optical media such as floptical disks.
  • -optical medium may be configured to store program instructions, including ROM, RAM, flash memory, etc.
  • a site that supplies or distributes an app store or other various software that distributes applications. , Recording media and storage media managed by servers, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Software Systems (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Data Mining & Analysis (AREA)
  • Medical Informatics (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Character Discrimination (AREA)
  • Character Input (AREA)
  • Image Analysis (AREA)

Abstract

글자 인식 장치에 의한 입력 데이터 내 글자를 인식하는 방법에 있어서, 입력 데이터를 글자 검출 모델에 입력하는 단계; 글자 검출 모델에서 출력되는 출력 데이터에 기초하여 입력 데이터 내 단어 영역의 위치 정보를 획득하는 단계; 획득한 위치 정보에 대응하는 부분 데이터를 입력 데이터로부터 추출하는 단계; 및 부분 데이터를 글자 인식 모델에 입력하여 부분 데이터 내 글자를 인식하는 단계를 포함하는 것을 특징으로 하는 일 실시예에 따른 글자 인식 방법이 개시된다.

Description

2020/175806 1»(:1/10公020/001333 명세서
발명의 명칭:글자인식장치 및이에의한글자인식방법 기술분야
[1] 본개시는데이터처리분야에관한것이다.보다구체적으로,본개시는이미지 등의데이터에서글자를인식하는글자인식장치및방법에관한것이다.
배경기술
四 핀테크 1^11)기술의발전에따라,휴대폰등에카드정보를저장하여놓고 간편하게결제할수있게하는서비스가제공되고있다.신용카드,체크카드 등의실물카드이미지에서카드번호및유효기간등의정보를인식및 저장하는기술은간편결제서비스를위한핵심이되는기술중하나이다.
[3] 그러나,카드이미지에서글자를인식함에 있어 ,카드내양각으로인쇄된
글자가다수존재하고,카드배경이다양하므로,카드번호및유효기간을 정확하게인식하는것에기술장벽이존재한다.
발명의상세한설명
기술적과제
[4] 일실시예에따른글자인식장치및이에의한글자인식방법은이미지등의 데이터에서글자를정확하고신속하게인식하는것을기술적과제로한다.
[5] 또한,일실시예에따른글자인식장치및이에의한글자인식방법은실물 카드의이미지내에서글자를정확하게인식하여핀테크산업의발전에 기여하는것을기술적과제로한다.
과제해결수단
[6] 일실시예에따른글자인식방법은,입력데이터를글자검출모델에입력하는 단계 ;상기글자검출모델에서출력되는출력데이터에기초하여상기입력 데이터내단어영역의위치정보를획득하는단계;상기획득한위치정보에 대응하는부분데이터를상기입력데이터로부터추출하는단계 ;및상기부분 데이터를글자인식모델에입력하여상기부분데이터내글자를인식하는 단계를포함할수있다.
발명의효과
[7] 일실시예에따른글자인식장치및이에의한글자인식방법은이미지등의 데이터에서글자를정확하고신속하게인식할수있다.
[8] 또한,일실시예에따른글자인식장치및이에의한글자인식방법은실물 카드의이미지내에서글자를정확하게인식하여핀테크산업의발전에기여할 수있다.
[9] 다만,일실시예에따른글자인식장치및이에의한글자인식방법이달성할 수있는효과는이상에서언급한것들로제한되지않으며,언급하지않은또 다른효과들은아래의기재로부터본개시가속하는기술분야에서통상의 2020/175806 1»(:1/10公020/001333
2 지식을가진자에게 명확하게 이해될수있을것이다.
도면의간단한설명
본명세서에서 인용되는도면을보다충분히 이해하기위하여 각도면의 간단한설명이제공된다.
도 1은일실시예에 따른글자인식장치를도시하는도면이다.
도 2는일실시예에 따른글자인식방법을설명하기위한순서도이다.
도 3은일실시예에 따른글자인식장치를통해글자가인식되는과정을 설명하기 위한도면이다.
도 4는글자검출모델에의해출력되는출력 데이터를도시하는예시적인 도면이다.
도 5는글자검출모델에서출력된출력 데이터에 기초하여 입력 데이터내 단어 영역의위치 정보를획득하는방법을설명하기 위한도면이다.
도 6은도 5에도시된이진화과정 및병합과정을설명하기위한도면이다. 도 7은도 5에도시된단어 박스결정과정을설명하기 위한도면이다. 도 8은일실시예에 따른특징추출모델의구조를설명하기 위한도면이다. 도 9는일실시예에 따른글자인식모델의구조를설명하기 위한도면이다. 도 은일실시예에 따른글자검출모델의훈련방법을설명하기위한 순서도이다.
[21] 도 11은제 1 (31스코어 맵을생성하는방법을설명하기위한도면이다.
[22] 도 12는제 2 (31스코어 맵을생성하는방법을설명하기위한도면이다.
[23] 도 13은서로인접한글자박스들사이에서 연결박스를결정하는방법을
설명하기 위한도면이다.
도 14는일실시예에 따른글자인식장치의구성을도시하는블록도이다. 도 15는일실시예에 따른글자인식장치가적용될수있는서버장치 및
024790 4356581
22 27 클라이언트장치를도시하는도면이다.
발명의실시를위한최선의형태
[26] 일실시예에 따른글자인식방법은,입력 데이터를글자검출모델에 입력하는 단계 ;상기글자검출모델에서출력되는출력 데이터에기초하여상기 입력 데이터 내단어 영역의위치 정보를획득하는단계;상기 획득한위치 정보에 대응하는부분데이터를상기 입력 데이터로부터추출하는단계 ;및상기부분 데이터를글자인식모델에 입력하여상기부분데이터내글자를인식하는 단계를포함할수있다.
다른실시예에 따른글자인식장치는,프로세서 ;및적어도하나의
인스트럭션을저장하는메모리를포함하되,상기프로세서는상기 적어도 하나의 인스트럭션에따라,입력 데이터를글자검출모델에 입력하고,상기 글자검출모델에서출력되는출력 데이터에기초하여상기 입력 데이터내단어 영역의 위치정보를획득하고,상기 획득한위치정보에 대응하는부분데이터를 2020/175806 1»(:1^1{2020/001333
3 상기 입력 데이터로부터추출하고,상기부분데이터를글자인식모델에 입력하여상기부분데이터내글자를인식할수있다.
발명의실시를위한형태
[28] 본개시는다양한변경을가할수있고여러 가지실시예를가질수있는바, 특정실시예들을도면에 예시하고,이를상세한설명을통해설명하고자한다. 그러나,이는본개시를특정한실시 형태에 대해 한정하려는것이 아니며,본 개시의사상및기술범위에포함되는모든변경,균등물내지 대체물을 포함하는것으로이해되어야한다.
[29] 실시예를설명함에 있어서,관련된공지 기술에 대한구체적인설명이요지를 불필요하게흐릴수있다고판단되는경우그상세한설명을생략한다.또한, 실시예의 설명과정에서 이용되는숫자(예를들어,제 1,제 2등)는하나의 구성요소를다른구성요소와구분하기위한식별기호에불과하다.
[3이 또한,본명세서에서 일구성요소가다른구성요소와 "연결된다”거나
"접속된다”등으로언급된때에는,상기 일구성요소가상기다른구성요소와 직접 연결되거나또는직접접속될수도있지만,특별히 반대되는기재가 존재하지 않는이상,중간에또다른구성요소를매개하여 연결되거나또는 접속될수도있다고이해되어야할것이다.
[31] 또한,본명세서에서’〜부(유닛)’,’모듈’등으로표현되는구성요소는 2개 이상의 구성요소가하나의구성요소로합쳐지거나또는하나의구성요소가보다 세분화된기능별로 2개 이상으로분화될수도있다.또한,이하에서 설명할 구성요소각각은자신이 담당하는주기능이외에도다른구성요소가담당하는 기능중일부또는전부의 기능을추가적으로수행할수도있으며,구성요소 각각이 담당하는주기능중일부기능이다른구성요소에의해 전담되어수행될 수도있음은물론이다.
[32] 또한,본명세서에서’글자’는단어나문장을구성하는기본문자단위를의미할 수있다,예를들어,영어의 경우에는각각의 알파벳이글자에 해당할수있고, 숫자의 경우에는’ 0’내지’9’의숫자각각이글자에해당할수있고,한국어의 경우에는자음과모음이 결합된문자(예를들어,’가’),자음,모음및자음이 결합된문자(예를들어,’강’),단독으로기재된자음(예를들어,’기’),단독으로 기재된모음(예를들어,나’)이글자에 해당할수있다.또한,글자는기호(예를 들어,’/’, '-,등)를포함할수도있다.
[33] 또한,본명세서에서’단어’는적어도하나의글자를포함하는문자단위를 의미할수있다.’단어’를구성하는글자들은서로간에소정간격 이상이격되어 있지 않을수있다.’단어’는하나의글자로이루어질수도있다.예를들어, 영어의부정사’ 는하나의글자로이루어졌지만주변글자와소정 거리 이상 이격되어 있는경우’단어’에해당할수있다.
[34] 또한,본명세서에서’글자그룹’은후술하는어느하나의부분데이터에서 2020/175806 1»(:1^1{2020/001333
4 인식된적어도하나의글자들을의미할수있다.
[35] 이하,본개시의기술적사상에의한실시예들을차례로상세히설명한다.
[36] 도 1은일실시예에따른글자인식장치 (100)를도시하는도면이다.
[37] 일실시예에따른글자인식장치 (100)는입력데이터 (10)를획득하고,입력 데이터 (10)내에서글자 (50)를인식한다.입력데이터 (10)는체크카드,신용카드 등의실물카드를촬영한이미지를포함할수있으며,또는후술하는바와같이, 실물카드등을촬영한이미지에기초하여특징추출모델 (800)에서출력된특징
Figure imgf000006_0001
11 ]3)을포함할수도있다.
[38] 글자인식장치 00)는입력데이터 (10)에서카드번호,유효기간등의카드 정보를인식및저장할수있다.글자인식장치 00)에의해인식및저장된카드 정보는물건등의구매를위해대금을지불하는데이용될수있다.
[39] 이하에서는,도 2및도 3을참조하여글자인식장치 00)의동작에대해
설명한다.
[4이 도 2는일실시예에따른글자인식방법을설명하기위한순서도이고,도 3은 일실시예에따른글자인식장치 (100)를통해글자가인식되는과정을설명하기 위한도면이다.
[41] 8210단계에서,글자인식장치 (100)는입력데이터 (10)를글자검출
모델 (410)에입력한다.글자인식장치 (100)는글자검출모델 (410)을미리 저장할수있다.글자검출모델 (사 0)은학습용데이터에기초하여훈련될수 있다.
[42] 8220단계에서,글자인식장치 (100)는글자검출모델 (410)에서출력되는출력 데이터 (30)에기초하여입력데이터 (10)내단어영역의위치정보를획득한다.
[43] 글자검출모델 (410)에서출력되는출력데이터 (30)는,입력데이터 (10)내
글자가존재할것으로예상되는지점의위치를나타낸다.글자인식장치 ( 0)는 출력데이터 (30)에기초하여입력데이터 (10)내적어도하나의글자를포함하는 단어영역의위치정보를획득한다.
[44] 8230단계에서 ,글자인식장치 (100)는단어영역의위치정보에대응하는부분 데이터 (40)를입력데이터 (10)로부터추출한다.일실시예에서,단어영역의위치 정보가복수개로획득된경우,각위치정보에대응하는복수의부분
데이터들 (40)이입력데이터 (10)로부터추출될수있다.
[45] 8240단계에서 ,글자인식장치 (100)는부분데이터 (40)를글자인식
모델 (420)에입력하여부분데이터 (40)에포함된글자 (50)를인식한다.부분 데이터 (40)가복수개인경우,글자인식장치 (100)는복수의부분데이터 (40) 각각을글자인식모델 (420)에입력하여복수의부분데이터 (40)각각에포함된 글자 (50)들을인식할수있다.
[46] 일실시예에서,글자인식장치 (100)는글자검출모델 (410)의출력
데이터 (30)를부분데이터 (40)와함께글자인식모델 (420)로입력할수도있다. 글자검출모델 (신 0)의출력데이터 (30)는입력데이터 (10)내개별글자의위치 2020/175806 1»(:1^1{2020/001333
5 정보를포함할수있으므로,글자인식모델 (420)의글자인식의정확도가보다 향상될수있다.
[47] 글자인식장치 (100)는인식된글자를저장하거나,네트워크를통해외부
장치로전송할수있다.
[48] 도 4는글자검출모델 (410)에의해출력되는출력데이터 (30)의일예를
도시하는예시적인도면이다.
[49] 출력데이터 (30)는입력데이터 (10)내글자가존재할확률을입력데이터 (10)에 대응되는데이터공간 (예를들어,이미지공간)상에나타내는제 1스코어맵 (31), 및입력데이터 (10)내글자들사이의연결성 (<:011116(:11\ )0을입력데이터 (10)에 대응되는데이터공간상에나타내는제 2스코어맵 (33)을포함할수있다.
[5이 제 1스코어맵 (31)내각위치에저장된값 (예를들어,픽셀값)은해당위치에 대응하는입력데이터 (10)에글자가존재할확률을나타낼수있다.또한,제 2 스코어맵 (33)내각위치에저장된값 (예를들어 ,픽셀값)은해당위치에 대응하는입력데이터 (10)내에서복수의글자들이서로인접할확률을나타낼 수있다.
[51] 위치대응관계에대한계산을용이하게하기위하여제 1스코어맵 (31)및제 2 스코어맵 (33)의크기는입력데이터 (10)와동일하게할수있다.
[52] 후술하는바와같이,글자검출모델 (410)은학습용데이터에대응하여생성된
Figure imgf000007_0001
스코어맵및제 2 (31스코어맵과유사한제 1스코어 맵 (31)및제 2스코어맵 (33)이출력되도록훈련될수있다.
[53] 글자인식장치 (100)는제 1스코어맵 (31)및제 2스코어맵 (33)에기초하여 입력데이터 (10)내단어영역의위치정보를결정할수있는데,이에대해서는도 5내지도 7을참조하여설명한다.
[54] 도 5는글자검출모델 (410)에서출력된출력데이터 (30)에기초하여입력
데이터 (10)내단어영역의위치정보를획득하는방법을설명하기위한 도면이고,도 6은도 5에도시된이진화과정및병합과정을설명하기위한 도면이고,도 7은도 5에도시된단어박스결정과정을설명하기위한도면이다.
[55] 8510단계및 8520단계에서,글자인식장치 (100)는제 1스코어맵 (31)내
데이터값들을임계값과비교하여제 1스코어맵 (31)을
이진화 ( 1131뇨止011)하고,제 2스코어맵 (33)내데이터값들을임계값과 비교하여제 2스코어맵 (33)을이진화한다.일예에서,글자인식장치 00)는제 1스코어맵 (31)및제 2스코어맵 (33)내데이터값들중임계값이상의데이터 값들을제 1값으로변경하고,임계값미만의데이터값들을제 2값으로변경할 수있다.
[56] 도 6에도시된바와같이,제 1스코어맵 (31)및제 2스코어맵 (33)에서임계값 이상의값을갖는데이터들은이진화된제 1스코어맵 (601)및이진화된제 2 스코어맵 (603)에서제 1값을갖도록변경되고,제 1스코어맵 (31)및제 2 스코어맵 (33)에서임계값미만의값을갖는데이터들은이진화된제 1스코어 맵 (601)및이진화된제 2스코어맵 (603)에서제 2값을갖도록변경될수있다.
[57] 제 1스코어맵 (31)의이진화를위한임계값과제 2스코어맵 (33)의이진화를 위한임계값은서로동일하거나서로상이할수있다.
[58] S530단계에서,글자인식장치 (100)는이진화된제 1스코어맵 (601)와
이진화된제 2스코어맵 (603)을병합 (merge)한다.예를들어,글자인식 장치 (100)는이진화된제 1스코어맵 (601)과이진화된제 2스코어맵 (603)내 데이터값들을더하거나, or연산을하여병합맵 (605)을생성할수있다.예를 들어도 6에도시된바와같이 ,이진화된제 1스코어맵 (601)및이진화된제 2 스코어맵 (603)내제 1값을갖는데이터들이병합맵 (605)에함께포함될수 있다.이러한방법으로병합맵 (605)은입력데이터 (10)내글자가존재할 가능성이높은영역 (606)들과그렇지않은영역들로구분될수있다.
[59] S540단계에서,글자인식장치 (100)는병합맵 (605)을이용하여글자가포함된 영역을나타내는단어박스 (6W)를결정하게된다.
[6이 예를들면,병합맵 (605)내에서동일한 (내지는동일한범위의)값을가지고 서로인접하게연결된영역 (606)들의적어도일부를단어영역으로결정하고, 결정된단어영역을포함하는단어박스 (6W)를결정할수있다.일실시예에서, 글자인식장치 (W0)는병합맵 (605)내단어영역들의구분을위해단어영역 각각에대해레이블링 (labeling)을할수도있다.
[61] 일실시예에서,글자인식장치 (100)는병합맵 (605)을이용하여인식된
영역 (606)각각이실제단어를포함하는지를검증하기위해,추가확인을할수 있다.구체적으로예를들어,병합맵 (605)내에서동일한 (내지는동일한범위의) 값을가지고서로인접하게연결된영역 (606)을단어후보영역으로두고,단어 후보영역내의각데이터에대응하는제 1스코어맵 (601)의값들중에정해진 임계치보다높은것이하나이상존재하면해당단어후보영역을단어영역으로 결정할수있다.즉,각단어후보영역에대응하는제 1스코어맵 (601)의값들중 가장큰값과임계값을비교하여각단어후보영역이단어영역에해당하는지 여부를검증할수있다.
[62] 이렇게하면글자와유사한배경이있어단어후보영역으로결정된경우들을 필터링할수있게된다.
[63] 일실시예에서,글자인식장치 (100)는단어영역에해당하는것으로검증된 데이터들의영역을포함하는최소크기의단어박스 (6 W)를결정할수있다.
[64] 글자인식장치 (100)는결정된단어박스 (610)의위치정보 (예를들어,입력 데이터 (10)또는병합맵 (605)상에서의단어박스 (6W)의모서리들의위치값)를 단어영역의위치정보로결정할수있다.
[65] 단어영역의위치정보가결정되면,글자인식장치 (W0)는해당위치정보에 대응하는부분데이터 (40)를입력데이터 (10)로부터추출하고,추출된부분 데이터 (40)를글자인식모델 (420)에입력하여부분데이터 (40)내글자를인식할 수있다. [66] 앞서설명한바와같이,글자검출모델 (4W)로입력되는입력데이터 (10)는 원본이미지에기초하여특징검출모델 (800)에서출력되는특징맵 (feature map)을포함할수있다.도 8은특징검출모델 (800)의구조를설명하기위한 도면이다.
[67] 원본이미지 (20)는특징검출모델 (800)로입력될수있다.여기서,원본
이미지 (20)는특징검출모델 (800)로입력되는이미지를의미하는것이며,최초 카드등을촬영한이미지를복사한이미지또는변형한이미지가아님을 의미하는것은아니다.
[68] 원본이미지 (20)는제 1컨볼루션층 (805),제 2컨볼루션층 (810),제 3
컨볼루션층 (815),제 4컨볼루션층 (820),제 5컨볼루션층 (825)및제 6
컨볼루션층 (830)에서컨볼루션처리가된다.제 6컨볼루션층 (830)의출력과제 5 컨볼루션증 (825)의줄력이연접 (concatenation)연산되어제 1업
컨볼루션층 (835)으로입력되고,제 1업컨볼루션층 (835)으로입력된값들은 컨볼루션처리 (836),배치정규화 (normalization)(837),컨볼루션처리 (838)및배치 정규화 (839)를통해제 1업샘플링층 (840)으로입력된다.제 1업
샘플링층 (840)의출력은제 4컨볼루션층 (820)의출력과연접연산되어제 2업 컨볼루션층 (845)및제 2업샘플링층 (850)에서처리된다.제 2업
샘플링층 (850)의출력은제 3컨볼루션층 (815)의출력과연접연산되어제 3업 컨볼루션층 (855)과제 3업샘플링층 (860)에서처리되고,처리결과는제 2 컨볼루션층 (810)의출력과연접연산되어제 4업컨볼루션층 (865)에입력된다. 그리고,제 4업컨볼루션층 (865)으로부터출력된결과를입력데이터 (10)로 사용할수있다.
[69] 일실시예에서,입력데이터 (10)의가로크기및세로크기는원본이미지 (20)의 가로크기및세로크기의 1/2일수있으나,이에한정되는것은아니다.
P이 도 8에도시된특징검출모델 (800)의구조는하나의예시일뿐이며,
컨볼루션층,업컨볼루션층,업샘플링층의개수및처리순서는다양하게 변형될수있다.
1] 도 9는일실시예에따른글자인식모델 (420)의구조를설명하기위한
도면이다.
[72] 글자인식모델 (420)은입력데이터 (10)로부터추출된부분데이터 (40)를
입력받고,부분데이터 (40)내글자 (50)들을인식한다.글자인식모델 (420)은 컨볼루션네트워크 (421), RNN (recurrent neural network)(423)및
디코더 (decoder)(425)를포함할수있다.
3] 컨볼루션네트워크 (421)는적어도하나의컨볼루션층을포함하고,부분
데이터 (40)를컨볼루션처리하여특징맵을추출한다.일예시에서,컨볼루션 네트워크 (421)는공지된 VGG, ResNet등을포함할수있지만,일실시예에서 글자인식모델 (420)은원본이미지 (20)의특징맵 (즉,입력데이터)으로부터 추출된부분데이터 (40)를입력받을수있으므로,필요로하는컨볼루션층의 개수는적을수있다.
[74] RNN(423)은부분데이터 (40)에대응하는특징맵으로부터특징벡터의
시퀀스를주줄한다. RNN(423)은 bi-LSTM (bidirectional Long- short-term memory)을통해연속되는특징벡터들의컨텍스트 (context)관계를파악할수 있다.
5] 디코더 (425)는특징벡터들의시퀀스정보에서글자를추출한다.
디코더 (425)는어텐션 (attention)단계및생성 (generation)단계를수행할수 있는데,어텐션단계에서,디코더 (425)는어떤시퀀스에서정보를뽑을지여부를 나타내는웨이트를계산하고,생성단계에서가중치를시퀀스에적용하고, LSTM(Long-short-term memory)을통해개별글자를주줄할수있다.
[76] 한편,일실시예에서 ,글자인식장치 (100)는여러부분데이터 (40)들각각에서 인식된글자그룹들을소정기준에따라분류할수있다.일예에서,글자인식 장치 (100)는어느부분데이터 (40)에서인식된글자그룹에소정의기호 (예를 들어 , /)가포함되어 있으면,해당글자그룹을제 1종류의정보로결정할수 있다.카드내유효기간에는년도와월을구분하기위한소정의기호가 포함되어 있는것이일반적이므로,글자인식장치 (100)는어느부분
데이터 (40)에서인식된글자그룹에소정기호가포함되어있으면해당글자 그룹을유효기간정보로결정할수있는것이다.
7] 만약,소정의기호가포함되어있는글자그룹의개수가복수개인경우,글자 인식장치 (100)는년도에해당하는숫자 (예를들어,기호를기준으로우측에 위치하는숫자)가큰글자그룹을유효기간정보로결정할수있다.카드에유효 기간과발급일자가포함되어있는경우,유효기간에포함된년도가발급 일자에포함된년도보다클것이므로,글자인식장치 (100)는년도에해당하는 숫자가큰글자그룹을유효기간정보로결정할수있는것이다.
8] 또한,일실시예에서 ,글자인식장치 (100)는복수의부분데이터 (40)각각에서 인식된글자그룹들중소정의기호를포함하고있지않은글자그룹들을제 2 종류의정보로결정할수있다.제 2종류의정보는예를들어,카드번호정보를 포함할수있다.
[79] 또한,일실시예에서 ,글자인식장치 (100)는복수의부분데이터 (40)들
각각에서인식된글자그룹들을,입력데이터 (10)내에서의복수의부분 데이터 (40)들의위치에따라정렬할수있다.일예로,글자인식장치 (W0)는 입력데이터 (10)내좌상단을기준으로 Z스캔방식으로글자그룹들을정렬할 수있다.
[8이 글자인식장치 (100)는정렬된글자그룹들중연속으로정렬된소정개수의 글자그룹에포함된글자의개수에기초하여글자의재인식이필요한지여부를 결정할수있다.일예로,글자인식장치 (100)는정렬된글자그룹들중소정 개수의숫자를각각포함하면서연속으로정렬된소정개수의글자그룹이 존재하는경우,글자인식이정확히수행되어글자의재인식이필요하지않은 2020/175806 1»(:1^1{2020/001333
9 것으로결정할수있다.일반적으로카드번호는 16개의숫자들을포함하되,
4개의숫자끼리하나의글자그룹을이룬다는면에서,글자인식장치 00)는 정렬된글자그룹들중 4개의숫자를포함하는 4개의글자그룹이연속으로 정렬되어 있는경우,글자의재인식이필요하지않은것으로결정할수있다.
[81] 또한,일실시예에서 ,글자인식장치 (100)는여러부분데이터 (40)들에서
인식된글자그룹들에소정의기호가존재하지않으면,글자의재인식이필요한 것으로결정할수있다.
[82] 글자의재인식이필요한경우,글자인식장치 (100)는이미지의재촬영이
필요하다는정보를스피커,모니터등을통해출력하거나네트워크를통해외부 장치로알릴수있다.일실시예에서,글자인식장치 (100)가카메라의프리뷰 이미지로부터글자를인식하는중에,글자의재인식이필요한것으로결정된 경우,카메라를통해연속적으로촬영되고있는프리뷰이미지로부터글자를 재인식할수도있다.
[83] 이하에서는,도 10내지도 13을참조하여글자검출모델 (410)을훈련시키는 방법에대해설명한다.
[84] 도 은일실시예에따른글자검출모델 (사 0)의훈련방법을설명하기위한 순서도이다.
[85] 81010단계에서 ,글자인식장치 (100)는학습용데이터 (60)내글자가존재할 확률을데이터공간상에나타내는제 1 (31스코어맵 (기),및학습용데이터 (60) 내글자들사이의연결성을데이터공간상에나타내는제 2 (31스코어맵 (73)을 획득한다.학습용데이터 (60)의가로크기및세로크기는입력데이터 (10)의 가로크기및세로크기와동일할수있다.또한,학습용데이터 (60)의가로크기 및세로크기는제 1 (31스코어맵 (기)의가로크기및세로크기와동일하고,제 2 (31'스코어맵 (73)의가로크기및세로크기와도동일할수있다.
[86] 일실시예에서 ,학습용데이터 (60)는전술한원본이미지 (20)와마찬가지로 카드등의대상체를촬영한이미지또는해당이미지에기초하여추출된특징 맵을포함할수있다.
[87] 글자인식장치 (100)는학습용데이터 (60)로부터제 1 (31스코어맵 (기)및제 2 (31스코어맵 (73)중적어도하나를직접생성할수도있고,또는네트워크나 외부관리자를통해제 1 (31스코어맵 (기)및제 2 (31스코어맵 (73)중적어도 하나를수신할수도있다.
[88] 제 1 (31스코어맵 (기)내값들은해당지점에서학습용데이터 (60)에글자가 위치할확률을나타낼수있다.또한,제 2 (31스코어맵 (73)내값들은해당 지점에서복수의글자들이서로인접할확률을나타낼수있다.
[89] 81020단계에서 ,글자인식장치 (100)는학습용데이터 (60)를글자검출
모델 (사 0)에입력한다.
[9이 81030단계에서 ,학습용데이터 (60)에대응하여글자검출모델 (410)에서
출력되는제 1스코어맵및제 2스코어맵각각과제 1 (31스코어맵 (기)및제 2 GT스코어맵 (73)의비교결과에따라글자검출모델 (4W)의내부가중치가 갱신될수있다.
[91] 제 1스코어맵및제 2스코어맵각각과제 1 GT스코어맵 (기)및제 2 GT
스코어맵 (73)의비교결과에따라로스 (loss)값이산출될수있다.로스값은 예를들어 , L2 Loss값에해당할수있다.로스값은그외에도, LI loss, smooth LI loss등다양한방법을이용할수있다.산출된로스값은글자검출모델 (4W)에 입력되고,글자검출모델 (4W)은로스값에따라내부가중치를갱신할수있다.
[92] 도 11은제 1 GT스코어맵 (기)을생성하는방법을설명하기위한도면이고,도 12는제 2 GT스코어맵 (73)을생성하는방법을설명하기위한도면이다.또한, 도 13은서로인접한글자박스들 (62a, 62b)사이에서연결박스 (63a)를결정하는 방법을설명하기위한도면이다.
[93] 도 11을참조하면,학습용데이터 (60)내적어도하나의글자들을포함하는
단어영역들에대해단어박스들 (61a, 61b, 61c, 61d, 61e)이결정된다.그리고, 단어박스 (61a, 61b, 61c, 61d, 61e)내포함된글자들의개수에따라단어 박스 (61a, 61b, 61c, 61d, 61e)가적어도하나의글자박스 (62a, 62b, 62c, 62d)로 분할된다.예를들어,어느하나의단어박스내에 4개의글자들이포함되어 있는 경우,해당단어박스는총 4개의글자박스로분할될수있다.글자박스 (62a,
62b, 62c, 62d)각각에소정의이미지 (1100),예를들어 , 2D가우시안이미지가 합성되어제 1 GT스코어맵 (기)이생성될수있다.
[94] 도 12및도 13을참조하면,복수의글자박스들 (62a, 62b, 62c, 62d)중서로
인접한글자박스들사이의경계선 (L)상에위치하는연결박스 (63a, 63b, 63c)가 결정되고,연결박스 (63a, 63b, 63c)에소정이미지 (1100),예를들어, 2D가우시안 이미지가합성되어제 2 GT스코어맵 (73)이생성될수있다.
[95] 연결박스 (63a, 63b, 63c)는,서로인접한글자박스들의내부공간에설정된 복수의지점들을연결함으로써결정될수있다.구체적으로,도 13에도시된 바와같이,서로인접한글자박스들 (62a, 62b)중좌측글자박스 (62a)내 2개의 지점및우측글자박스 (62b)내 2개의지점을연결한연결박스 (63a)가결정될수 있다.
[96] 일예에서 ,서로인접한좌측글자박스 (62a)및우측글자박스 (62b)의
모서리들중좌측하단모서리와우측상단모서리를연결하고,좌측상단 모서리와우측하단모서리를연결하여상부및하부의삼각형들을결정하고, 해당삼각형들의중점들을연결함으로써연결박스 (63a)가결정될수도있다.
[97] 도 14는일실시예에따른글자인식장치 (100)의구성을도시하는블록도이다.
[98] 도 14를참조하면,글자인식장치 (100)는메모리 (1410),통신모듈 (1430)및 프로세서 (1450)를포함할수있다.메모리 (1410)에는적어도하나의
인스트럭션이저장될수있고,프로세서 (1450)는적어도하나의인스트럭션에 따라글자검출및글자검출모델 (4W)의훈련을제어할수있다.
[99] 도 14는하나의메모리 (1410)와하나의프로세서 (1450)만을도시하고있으나, 2020/175806 1»(:1^1{2020/001333
11 글자인식장치 (100)는복수의 메모리 및/또는복수의프로세서를포함할수도 있다.
[10이 메모리 (1410)는글자검출모델 (410)및글자인식모델 (420)을저장할수있다. 또한,메모리 (1410)는특징추출모델 (800)을더 저장할수있다.
[101] 프로세서 (1450)는글자검출모델 (410)로입력 데이터 (10)를입력하고,글자 검출모델 (사 0)에서출력되는출력 데이터에기초하여 입력 데이터 (10)내단어 영역의 위치정보를획득할수있다.그리고,프로세서 (1450)는획득한위치 정보에 대응하는부분데이터를글자인식모델 (420)에 입력하고,글자인식 모델 (420)에서출력된글자정보를메모리 (1410)또는기타저장장치에 저장할 수있다.
[102] 일실시예에서,프로세서 (1450)는학습용데이터 (60)에기초하여글자검출 모델 (410),글자인식모델 (420)및특징추출모델 (800)중적어도하나를 훈련시킬수있다.
[103] 통신모듈 (1430)은네트워크를통해외부장치와데이터를송수신한다.예를 들어,통신모듈 (1430)은외부장치와이미지를송수신하거나,입력 데이터 (10) 내에서 인식된글자정보를외부장치와송수신할수있다.
[104] 도 15는일실시예에 따른글자인식장치 (100)가적용될수있는서버
장치 (1510)및클라이언트장치 (1520)를도시하는도면이다.
[105] 글자인식장치 (100)는서버장치 (1510)로구현되거나또는클라이언트
장치 (1520)로구현될수있다.
[106] 글자인식장치 (100)가서버장치 (1510)로구현되는경우,서버장치 (1510)는 클라이언트장치 (1520)로부터 이미지를수신하고,수신된이미지내에서글자를 인식하여 저장할수있다.일예에서,서버장치 (1510)는클라이언트
장치 (1520)로부터수신된이미지내에서 인식된글자정보를클라이언트 장치 (1520)로전송할수도있다.또한,서버장치 (1510)는클라이언트
장치 (1520)를포함한외부장치로부터 학습용데이터를수신하거나,또는내부에 저장된학습용데이터를이용하여글자검출모델 (사 0),글자인식모델 (420)및 특징추출모델 (800)중적어도하나의훈련을제어할수도있다.
[107] 글자인식장치 (100)가클라이언트장치 (1520)로구현되는경우,클라이언트 장치 (1520)는클라이언트장치 (1520)의카메라에 의해촬영된이미지또는 클라이언트장치 (1520)에 저장된이미지 내에서글자를인식하여 저장할수 있다.
[108] 일실시예에서,클라이언트장치 (1520)는글자검출모델 (사 0),글자인식
모델 (420)및특징추출모델 (800)중적어도하나의실행을위한데이터를서버 장치 (1510)로부터수신할수있다.클라이언트장치 (1520)는카메라모듈을통해 촬영된이미지,내부메모리에 저장된이미지또는외부장치로부터수신된 이미지를글자검출모델 (410),글자인식모델 (420)및특징추출모델 (800)중 적어도하나에 입력시켜글자를인식할수있다. [109] 클라이언트장치 (1520)는외부장치로부터학습용데이터를수신하거나,또는 내부에저장된학습용데이터를이용하여글자검출모델 (4W),글자인식 모델 (420)및특징추출모델 (800)중적어도하나의훈련을제어할수도있다. 구현예에따라,글자검출모델 (4W),글자인식모델 (420)및특징추출모델 (800) 중적어도하나의실행을위한데이터를클라이언트장치 (1520)로제공한서버 장치 (1510)가학습용데이터에기초하여글자검출모델 (4W),글자인식 모델 (420)및특징추출모델 (800)중적어도하나의훈련을제어할수도있다.이 경우,서버장치 (1510)는훈련결과갱신된가중치정보만을클라이언트 장치 (1520)로전송하고,클라이언트장치 (1520)는수신된정보에따라글자검출 모델 (410),글자인식모델 (420)및특징추출모델 (800)중적어도하나를갱신할 수있다.
[110] 도 15는클라이언트장치 (1520)로서 ,데스크탑 PC를도시하고있으나,이에 한정되는것은아니고클라이언트장치 (1520)는노트북,스마트폰,태블릿 PC, AI(artificial intelligence)로봇, AI스피커 ,웨어러블기기등을포함할수있다.
[111] 한편,상술한본개시의실시예들은컴퓨터에서실행될수있는프로그램으로 작성가능하고,작성된프로그램은매체에저장될수있다.
[112] 매체는컴퓨터로실행가능한프로그램을계속저장하거나,실행또는
다운로드를위해임시저장하는것일수도있다.또한,매체는단일또는수개 하드웨어가결합된형태의다양한기록수단또는저장수단일수있는데,어떤 컴퓨터시스템에직접접속되는매체에한정되지않고,네트워크상에분산 존재하는것일수도있다.매체의예시로는,하드디스크,플로피디스크및자기 테이프와같은자기매체 , CD-ROM및 DVD와같은광기록매체 ,플롭티컬 디스크 (floptical disk)와같은자기 -광매체 (magneto-optical medium),및 ROM, RAM,플래시메모리등을포함하여프로그램명령어가저장되도록구성된것이 있을수있다.또한,다른매체의예시로,애플리케이션을유통하는앱스토어나 기타다양한소프트웨어를공급내지유통하는사이트,서버등에서관리하는 기록매체내지저장매체도들수있다.
[113] 이상,본개시의기술적사상을바람직한실시예를들어상세하게
설명하였으나,본개시의기술적사상은상기실시예들에한정되지않고,본 개시의기술적사상의범위내에서당분야에서통상의지식을가진자에의하여 여러가지변형및변경이가능하다.

Claims

2020/175806 1»(:1/10公020/001333 13 청구범위
[청구항 1] 글자인식장치에의한입력데이터내글자를인식하는방법에있어서, 입력데이터를글자검출모델에입력하는단계;
상기글자검출모델에서출력되는출력데이터에기초하여상기입력 데이터내단어영역의위치정보를획득하는단계;
상기획득한위치정보에대응하는부분데이터를상기입력 데이터로부터추출하는단계;및
상기부분데이터를글자인식모델에입력하여상기부분데이터내 글자를인식하는단계를포함하는것을특징으로하는글자인식방법 .
[청구항 2] 제 1항에 있어서,
상기줄력데이터는,
상기입력데이터내글자가존재할확률을상기입력데이터에대응되는 데이터공간상에나타내는제 1스코어맵,및상기입력데이터내글자들 사이의연결성을상기입력데이터에대응되는데이터공간상에 나타내는제 2스코어맵을포함하는것을특징으로하는글자인식방법 .
[청구항 3] 제 2항에 있어서,
상기단어영역의위치정보를획득하는단계는,
상기제 1스코어맵및상기제 2스코어맵내값들과임계값의비교 결과에따라상기제 1스코어맵및상기제 2스코어맵을이진화하는 단계;
이진화된상기제 1스코어맵과이진화된상기제 2스코어맵을 병합하는단계 ;
병합맵내에서소정값을갖는영역을결정하는단계 ;및
상기결정된영역을포함하는단어영역의위치정보를결정하는단계를 포함하는것을특징으로하는글자인식방법.
[청구항 4] 제 3항에 있어서,
상기단어영역의위치정보를결정하는단계는,
상기결정된영역을포함하는최소크기의단어박스를결정하는단계;및 상기결정된단어박스의위치정보를상기단어영역의위치정보로 결정하는단계를포함하는것을특징으로하는글자인식방법.
[청구항 5] 제 2항에 있어서,
상기글자인식방법은,
학습용데이터내글자가존재할확률을데이터공간상에나타내는제 1 (31스코어맵,및상기학습용데이터내글자들사이의연결성을데이터 공간상에나타내는제 2 (31스코어맵을획득하는단계 ;및 상기학습용데이터를상기글자검출모델에입력하는단계를더 포함하되, 2020/175806 1»(:1^1{2020/001333
14 상기학습용데이터에대응하여상기글자검출모델에서출력되는제 1 스코어맵및제 2스코어맵각각과상기제 1 (31스코어맵및상기제 2 (31스코어맵의비교결과에따라상기글자검출모델의내부가중치가 갱신되는것을특징으로하는글자인식방법.
[청구항 6] 제 5항에있어서,
상기제 1 (31스코어맵을획득하는단계는,
상기학습용데이터내단어를포함하는단어박스를결정하는단계; 상기결정된단어박스에포함된글자의개수에따라상기단어박스를 복수의글자박스로분할하는단계 ;및
상기복수의글자박스각각에소정의이미지를합성하여상기제 1 01 스코어맵을생성하는단계를포함하는것을특징으로하는글자인식 방법.
[청구항 7] 제 6항에있어서,
상기제 2 (31스코어맵을생성하는단계는,
상기복수의글자박스들중서로인접한글자박스들사이의경계선상에 위치하는연결박스를결정하는단계 ;및
상기연결박스에소정의이미지를합성하여상기제 2 (31'스코어맵을 생성하는단계를포함하는것을특징으로하는글자인식방법 .
[청구항 8] 제 1항에있어서,
상기글자인식방법은,
상기부분데이터내에서인식된글자그룹에소정의기호가포함되어 있는경우,해당글자그룹을제 1종류의정보로결정하는단계를더 포함하는것을특징으로하는글자인식방법.
[청구항 9] 제 1항에있어서,
상기입력데이터로부터추출된부분데이터의개수는복수개이되 , 상기글자인식방법은,
복수의부분데이터들각각에서인식된글자그룹들을,상기입력데이터 내에서의상기복수의부분데이터들의위치에따라정렬하는단계를더 포함하는것을특징으로하는글자인식방법.
[청구항 10] 제 9항에있어서,
상기글자인식방법은,
상기정렬된글자그룹들중연속으로정렬된소정개수의글자그룹에 포함된글자의개수에기초하여글자의재인식이필요한지여부를 결정하는단계를더포함하는것을특징으로하는글자인식방법.
[청구항 11] 제 1항에있어서,
상기글자를인식하는단계는,
상기글자검출모델에서출력되는출력데이터를상기글자인식모델로 더입력시켜상기부분데이터내글자를인식하는단계를포함하는것을 2020/175806 1»(:1^1{2020/001333
15 특징으로하는글자인식방법 .
[청구항 12] 제 1항에 있어서,
상기 입력 데이터는,
원본이미지에 대응하여특징추출모델로부터출력된특징 맵을 포함하는것을특징으로하는글자인식방법.
[청구항 13] 하드웨어와결합하여제 1항의글자인식 방법을실행하기 위하여 매체에 저장된프로그램.
[청구항 14] 프로세서;및
적어도하나의 인스트럭션을저장하는메모리를포함하되,
상기프로세서는상기 적어도하나의 인스트럭션에따라, 입력 데이터를글자검출모델에 입력하고,
상기글자검출모델에서출력되는출력 데이터에 기초하여상기 입력 데이터 내단어 영역의위치 정보를획득하고,
상기 획득한위치 정보에 대응하는부분데이터를상기 입력
데이터로부터추출하고,
상기부분데이터를글자인식모델에 입력하여상기부분데이터내 글자를인식하는것을특징으로하는글자인식장치.
PCT/KR2020/001333 2019-02-25 2020-01-29 글자 인식 장치 및 이에 의한 글자 인식 방법 Ceased WO2020175806A1 (ko)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2021549641A JP7297910B2 (ja) 2019-02-25 2020-01-29 文字認識装置及び文字認識装置による文字認識方法

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020190022102A KR102206604B1 (ko) 2019-02-25 2019-02-25 글자 인식 장치 및 이에 의한 글자 인식 방법
KR10-2019-0022102 2019-02-25

Publications (1)

Publication Number Publication Date
WO2020175806A1 true WO2020175806A1 (ko) 2020-09-03

Family

ID=72240107

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2020/001333 Ceased WO2020175806A1 (ko) 2019-02-25 2020-01-29 글자 인식 장치 및 이에 의한 글자 인식 방법

Country Status (3)

Country Link
JP (1) JP7297910B2 (ko)
KR (1) KR102206604B1 (ko)
WO (1) WO2020175806A1 (ko)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPWO2023058082A1 (ko) * 2021-10-04 2023-04-13
JP7618458B2 (ja) 2021-02-08 2025-01-21 株式会社東芝 文字認識装置、文字認識方法、及びプログラム

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102386162B1 (ko) * 2020-11-13 2022-04-15 주식회사 와들 이미지로부터 상품 정보 데이터를 생성하기 위한 시스템 및 그에 관한 방법
KR102548826B1 (ko) * 2020-12-11 2023-06-28 엔에이치엔클라우드 주식회사 딥러닝 기반의 메뉴판 제공 방법 및 그 시스템
CN119522448A (zh) * 2022-07-13 2025-02-25 株式会社东芝 字符识别装置、字符识别方法以及程序

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101295000B1 (ko) * 2013-01-22 2013-08-09 주식회사 케이지모빌리언스 카드 번호의 영역 특성을 이용하는 신용 카드의 번호 인식 시스템 및 신용 카드의 번호 인식 방법
KR20160065174A (ko) * 2013-10-03 2016-06-08 마이크로소프트 테크놀로지 라이센싱, 엘엘씨 텍스트 예측용 이모지
KR20170029947A (ko) * 2015-09-08 2017-03-16 에스케이플래닛 주식회사 단말장치를 이용한 신용카드 번호 및 유효기간 인식 시스템 및 방법
KR101805318B1 (ko) * 2016-11-01 2017-12-06 포항공과대학교 산학협력단 텍스트 영역 식별 방법 및 장치
KR20180112590A (ko) * 2017-04-04 2018-10-12 한국전자통신연구원 멀티미디어 지식 베이스 구축 시스템 및 방법

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3840206B2 (ja) * 2003-06-23 2006-11-01 株式会社東芝 複写機における翻訳方法及びプログラム
JP2010191724A (ja) * 2009-02-18 2010-09-02 Seiko Epson Corp 画像処理装置および制御プログラム
KR101727137B1 (ko) * 2010-12-14 2017-04-14 한국전자통신연구원 텍스트 영역의 추출 방법, 추출 장치 및 이를 이용한 번호판 자동 인식 시스템
US9449239B2 (en) * 2014-05-30 2016-09-20 Apple Inc. Credit card auto-fill
JP2017058950A (ja) * 2015-09-16 2017-03-23 大日本印刷株式会社 認識装置、撮像システム、撮像装置並びに認識方法及び認識用プログラム

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101295000B1 (ko) * 2013-01-22 2013-08-09 주식회사 케이지모빌리언스 카드 번호의 영역 특성을 이용하는 신용 카드의 번호 인식 시스템 및 신용 카드의 번호 인식 방법
KR20160065174A (ko) * 2013-10-03 2016-06-08 마이크로소프트 테크놀로지 라이센싱, 엘엘씨 텍스트 예측용 이모지
KR20170029947A (ko) * 2015-09-08 2017-03-16 에스케이플래닛 주식회사 단말장치를 이용한 신용카드 번호 및 유효기간 인식 시스템 및 방법
KR101805318B1 (ko) * 2016-11-01 2017-12-06 포항공과대학교 산학협력단 텍스트 영역 식별 방법 및 장치
KR20180112590A (ko) * 2017-04-04 2018-10-12 한국전자통신연구원 멀티미디어 지식 베이스 구축 시스템 및 방법

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7618458B2 (ja) 2021-02-08 2025-01-21 株式会社東芝 文字認識装置、文字認識方法、及びプログラム
JPWO2023058082A1 (ko) * 2021-10-04 2023-04-13
JP7643579B2 (ja) 2021-10-04 2025-03-11 日本電気株式会社 情報処理装置、情報処理システム、情報処理方法、及び、プログラム

Also Published As

Publication number Publication date
KR102206604B1 (ko) 2021-01-22
KR20200106110A (ko) 2020-09-11
JP7297910B2 (ja) 2023-06-26
JP2022522425A (ja) 2022-04-19

Similar Documents

Publication Publication Date Title
WO2020175806A1 (ko) 글자 인식 장치 및 이에 의한 글자 인식 방법
CN109858555B (zh) 基于图像的数据处理方法、装置、设备及可读存储介质
US8744196B2 (en) Automatic recognition of images
RU2661750C1 (ru) Распознавание символов с использованием искусственного интеллекта
CN111967286B (zh) 信息承载介质的识别方法、识别装置、计算机设备和介质
CN110929573A (zh) 基于图像检测的试题检查方法及相关设备
CN108304835A (zh) 文字检测方法和装置
JP7198350B2 (ja) 文字検出装置、文字検出方法及び文字検出システム
CN112818852A (zh) 印章校验方法、装置、设备及存储介质
CN114299509A (zh) 一种获取信息的方法、装置、设备及介质
CN117523586A (zh) 支票印章的验证方法、装置、电子设备和介质
Le et al. Deep learning approach for receipt recognition
US11699297B2 (en) Image analysis based document processing for inference of key-value pairs in non-fixed digital documents
KR102351578B1 (ko) 글자 인식 장치 및 이에 의한 글자 인식 방법
CN114418124A (zh) 生成图神经网络模型的方法、装置、设备及存储介质
US12462589B2 (en) Text line detection
Dat et al. An improved CRNN for Vietnamese Identity Card Information Recognition.
Zhao et al. DetectGAN: GAN-based text detector for camera-captured document images
US20250225804A1 (en) Method of extracting information from an image of a document
Pham et al. A deep learning approach for text segmentation in document analysis
Chang et al. Re-Attention is all you need: Memory-efficient scene text detection via re-attention on uncertain regions
Ou et al. ERCS: An efficient and robust card recognition system for camera-based image
CN115482539A (zh) 一种基于图像识别的收货单生成方法和装置
Dugar et al. From pixels to words: A scalable journey of text information from product images to retail catalog
Shinde et al. Using CRNN to Perform OCR over Forms

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20763568

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021549641

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20763568

Country of ref document: EP

Kind code of ref document: A1