TH105604B - A process for semi-automatic character image recognition. - Google Patents

A process for semi-automatic character image recognition.

Info

Publication number
TH105604B
TH105604B TH901003430A TH0901003430A TH105604B TH 105604 B TH105604 B TH 105604B TH 901003430 A TH901003430 A TH 901003430A TH 0901003430 A TH0901003430 A TH 0901003430A TH 105604 B TH105604 B TH 105604B
Authority
TH
Thailand
Prior art keywords
character
image
database
character image
string
Prior art date
Application number
TH901003430A
Other languages
Thai (th)
Other versions
TH103248A (en
TH105604A (en
Inventor
วัชรบุศราคำ นางสาวศรินทร์
ดูเบ นายเปรมนาถ
โยชิดะ นายฮิโตกิ คานากาวะ นายโยชินาริ
สินธุภิญโญ นายวศิน
มฤคทัต นายสรรพฤทธิ์
Original Assignee
นิตโต เดนโก คอร์ปอเรชั่น
สำนักงานพัฒนาวิทยาศาตร์และเทคโนโลยีแห่งชาติ
Filing date
Publication date
Application filed by นิตโต เดนโก คอร์ปอเรชั่น, สำนักงานพัฒนาวิทยาศาตร์และเทคโนโลยีแห่งชาติ filed Critical นิตโต เดนโก คอร์ปอเรชั่น
Publication of TH103248A publication Critical patent/TH103248A/en
Publication of TH105604A publication Critical patent/TH105604A/en
Publication of TH105604B publication Critical patent/TH105604B/en

Links

Abstract

------21/02/2563------(OCR) หน้าที่ 1 ของจำนวน 1 หน้า บทสรุปการประดิษฐ์ การประดิษฐ์นี้นำเสนอระเบียบวิธีการรู้จำภาพตัวอักษรกึ่งอัตโนมัติ ซึ่งเป็นการรู้จำตัวอักษรโดยที่ผู้ใช้ต้องป้อนข้อมูลตัวอักษรก่อนการใช้งาน จากนั้นระบบจะทำการรู้จำภาพตัวอักษร โดยใช้การคำนวณค่าความใกล้เคียงของภาพตัวอักษรที่อินพุตกับตัวอักษรที่ผู้ใช้กำหนดให้ในตอนแรก และถ้าตัวอักษรมีการเชื่อมติดกัน 2 ตัวขึ้นไปหรือ เป็นตัวอักษรที่ไม่สมบูรณ์เนื่องจากการขาดของเส้นตัวอักษร ก็จะใช้ลักษณะเฉพาะของตัวอักษร เช่น ความกว้างของตัวอักษรมาช่วยกำหนดขนาดของตัวอักษรและอาศัยค่าความน่าจะเป็นของคู่ตัวอักษรช่วยคาดเดาตัวอักษรที่ติดกันว่าควรเป็นตัวอักษรใด เพื่อช่วยเพิ่มความถูกต้องในการรู้จำมากขึ้น นอกจากนี้ระเบียบวิธีที่นำเสนอยังสามารถนำไปประยุกต์ใช้กับภาษาใดก็ได้ ------------ DC60 การประดิษฐ์นี้นำเสนอระเบียบวิธีการรู้จำภาพตัวอักษรกึ่งอัตโนมัติ ซึ่งเป็นการรู้จำ ตัวอักษร โดยที่ผู้ใช้ต้องป้อนข้อมูลตัวอักษรก่อนการใช้งาน จากนั้นระบบจะทำการรู้จำภาพ ตัวอักษร โดยใช้การคำนวณค่าความใกล้เคียงของภาพตัวอักษรที่อินพุตกับตัวอักษรที่ผู้ใช้ กำหนดให้ในตอนแรก และถ้าตัวอักษรมีการเชื่อมติดกัน 2 ตัวขึ้นไปหรือ เป็นตัวอักษรที่ไม่สมบูรณ์ เนื่องจากการขาดของเส้นตัวอักษร ก็จะใช้ลักษณะเฉพาะของตัวอักษร เช่น ความกว้างของ ตัวอักษรมาช่วยกำหนดขนาดของตัวอักษรและอาศัยค่าความน่าจะเป็นของคู่ตัวอักษรช่วยคาดเดา ตัวอักษรที่ติดกันว่าควรเป็นตัวอักษรใด เพื่อช่วยเพิ่มความถูกต้องในการรู้จำมากขึ้น นอกจากนี้ ระเบียบวิธีที่นำเสนอยังสามารถนำไปประยุกต์ใช้กับภาษาใดก็ได้------February 21, 2020------(OCR) Page 1 of 1 Abstract of the Invention This invention presents a semi-automatic character image recognition method. This method recognizes characters where the user must input character data before use. The system then recognizes the character image by calculating the proximity of the input character image to the character initially specified by the user. If two or more characters are connected or are incomplete due to broken character strokes, it uses character characteristics such as character width to determine the character size and relies on the probability of character pairs to predict which adjacent character should be identified, thereby improving recognition accuracy. Furthermore, the proposed method can be applied to any language. ------------ DC60 This invention presents a semi-automatic character image recognition method. This method recognizes characters where the user must input character data before use. The system then recognizes the character image by calculating the proximity of the input character image to the character initially specified by the user. If two or more characters are connected or are incomplete due to broken character strokes, it uses character characteristics such as character width to determine the character size and relies on the probability of character pairs to predict which adjacent character should be identified, thereby improving recognition accuracy. In addition, the proposed method can be applied to any language. This method uses unique characteristics of characters, such as character width, to determine character size and relies on the probability of character pairs to predict which adjacent characters should be, thereby improving recognition accuracy. Furthermore, the proposed method can be applied to any language.

Claims (1)

------21/02/2563------(OCR) หน้าที่ 1 ของจำนวน 2 หน้า ข้อถือสิทธิ1. วิธีการรู้จำภาพตัวอักษรกึ่งอัตโนมัติเพื่อแปลงภาพตัวอักษรให้เป็นสายตัวอักษร ในระบบที่ ประกอบด้วย ฐานข้อมูล ไบแกรม (2-grams) และฐานข้อมูลรูปแบบตัวอักษร (template) ซึ่งวิธีการดังกล่าวประกอบด้วยขั้นตอนของ i) การรับข้อมูลภาพที่ประกอบด้วยภาพตัวอักษรจำนวนหนึ่ง ii) การเตรียมข้อมูลภาพที่ประกอบด้วย การตัดแบ่งภาพตัวอักษรดังกล่าวออกเป็น สายตัวอักษร โดยที่แต่ละสายตัวอักษรประกอบรวมด้วยตัวอักษรอย่างน้อยหนึ่ง ตัว และที่ซึ่งสายตัวอักษรแต่ละสายจะแยกจากกันด้วยช่องว่าง (space) iii) การตรวจสอบขนาดของสายตัวอักษรตังกล่าว ถ้าขนาดของสายตัวอักษรมีขนาด มากกว่าที่กำหนดไว้ แสดงว่าในสายตัวอักษรดังกล่าวประกอบด้วยตัวอักษร มากกว่าหนึ่งตัว ทำการแบ่งสายตัวอักษรออกเป็นส่วนภาพตัวอักษรที่หนึ่งและ ส่วนภาพตัวอักษรที่สอง iv) ทำการเปรียบเทียบส่วนภาพตัวอักษรตัวที่หนึ่งดังกล่าวของสายตัวอักษรดังกล่าว กับรูปแบบตัวอักษรในฐานข้อมูลรูปแบบตัวอักษร (template) และเลือกตัวอักษรที่ ใกล้เคียงที่สุดเพื่อกำหนดให้เป็นตัวอักษรที่หนึ่ง v) ทำการค้นหาคู่ตัวอักษรตัวที่สองจากฐานข้อมูลไบแกรม (2-grams) ดังกล่าว โดย อยู่บนพื้นฐานของตัวอักษรที่หนึ่งที่เลือกไว้ดังกล่าว โดยเลือกคู่ตัวอักษรที่มีค่า ความน่าจะเป็นสูงที่สุดในฐานข้อมูลไบแกรมก่อน เพื่อนำภาพตัวอักษรจาก ฐานข้อมูลรูปแบบตัวอักษรที่สอดคล้องกันมาเปรียบเทียบกับส่วนภาพตัวอักษรที่ สองดังกล่าว และ vi) ทำซํ้าขั้นตอนที่ (v) โดยใช้คู่ตัวอักษรที่มีค่าความน่าจะเป็นสูงในอันดับถัดไป จาก ฐานข้อมูลไบแกรมดังกล่าว จนกว่าจะพบภาพตัวอักษรที่ตรงกันกับตัวอักษรใน ฐานข้อมูลรูปแบบตัวอักษรดังกล่าว หรือ ตัวอักษรในฐานข้อมูลรูปแบบภาพ ตัวอักษรที่ให้ค่าความเหมือนกันมากที่สุด2. วิธีการตามข้อถือสิทธิ 1 ที่ยังประกอบเพิ่มเติมด้วย ขั้นตอนการคำนวณหาระยะห่างระหว่าง ตัวอักษรเฉลี่ย(space) และความกว้างของตัวอักษรเฉลี่ย3. วิธีการตามข้อถือสิทธิ 1 ที่ซึ่ง ขั้นตอนการรับข้อมูลภาพดังกล่าว ดำเนินการโดยการสแกน ภาพ หน้าที่ 2 ของจำนวน 2 หน้า4. วิธีการตามข้อถือสิทธิ 1 ที่ซึ่ง ขั้นตอนการเตรียมข้อมูลภาพดังกล่าวยังประกอบเพิ่มเติมด้วย ทำการหาเส้นบรรทัดโดยเทคนิคโปรเจคชันในแนวนอน เพื่อแบ่งรูปภาพตัวอักษรออกเป็น แถว จำนวนหนึ่ง5. วิธีการตามข้อถือสิทธิ 1 ที่ซึ่ง ขั้นตอนการคำนวณความใกล้เคียงกันของภาพตัวอักษรกับ รูปแบบตัวอักษร (template) ใช้สมการดังนี้ MAX Mi = (สัญลักษณ์)(Tij-Cij) ซึ่ง Mi หมายถึง ค่าความใกล้เคียง ณ ตำแหน่งแถวนอน i Tij หมายถึง ค่าสีของภาพ Template ณ ตำแหน่ง แถวนอน i และ แถวตั้ง j Cij หมายถึง ค่าสีของภาพตัวอักษรที่ต้องการรู้จำ ณ ตำแหน่ง แถว นอน i และ แถวตั้ง j ------------------February 21, 2020------(OCR) Page 1 of 2. Claim 1. A semi-automatic character image recognition method to convert character images into character strings in a system consisting of a bigram (2-grams) database and a font template database. The method consists of the following steps: i) Receiving image data containing a number of character images; ii) Preparing the image data by dividing the character image into character strings, each string containing at least one character, and separated by spaces; iii) Checking the size of the character string. If the string size is greater than the specified size, it means that the string contains more than one character. The string is then divided into a first character image segment and a second character image segment; iv) Comparing the first character image segment of the character string with the font template in the font template database and selecting the closest character to be the first character; v) Searching for the second character pair from the bigram (2-grams) database based on the selected first character, selecting the character pair with the closest value. The highest probability in the first bigram database is used to compare the character image from the corresponding character style database with the second character image, and vi) repeat step (v) using the next highest probability character pairs from the bigram database until a character image that matches the character in the character style database or the character in the character image database that gives the highest similarity is found. 2. The method according to claim 1, which also includes a step to calculate the average character spacing (space) and the average character width. 3. The method according to claim 1, where the image data acquisition step is performed by scanning the second page of two pages. 4. The method according to claim 1, where the image data preparation step also includes performing horizontal projection to divide the character image into a number of rows. 5. The method according to claim 1, where the calculation of the similarity of the character image with The font style (template) uses the following equation: MAX Mi = (symbol)(Tij-Cij), where Mi represents the proximity value at row i, Tij represents the color value of the template image at row i and column j, and Cij represents the color value of the character image to be recognized at row i and column j.
TH901003430A 2009-02-26 A process for semi-automatic character image recognition. TH105604B (en)

Publications (3)

Publication Number Publication Date
TH103248A TH103248A (en) 2010-08-11
TH105604A TH105604A (en) 2010-12-30
TH105604B true TH105604B (en) 2024-12-18

Family

ID=

Similar Documents

Publication Publication Date Title
CN110114776B (en) System and method for character recognition using fully convolutional neural networks
Firmani et al. Towards Knowledge Discovery from the Vatican Secret Archives. In Codice Ratio-Episode 1: Machine Transcription of the Manuscripts.
Bai et al. Keyword spotting in document images through word shape coding
US20090304282A1 (en) Recognition of tabular structures
JP2014106961A (en) Method executed by computer for automatically recognizing text in arabic, and computer program
CN111630521A (en) Image processing method and image processing system
US20050182760A1 (en) Apparatus and method for searching for digital ink query
Reffle et al. Unsupervised profiling of OCRed historical documents
TWI567569B (en) Natural language processing systems, natural language processing methods, and natural language processing programs
CN109815452A (en) Text comparative approach, device, storage medium and electronic equipment
Okamoto et al. Performance evaluation of a robust method for mathematical expression recognition
US9934429B2 (en) Storage medium, recognition method, and recognition apparatus
Ramakrishnan et al. Global and local features for recognition of online handwritten numerals and Tamil characters
Springmann et al. Automatic quality evaluation and (semi-) automatic improvement of OCR models for historical printings
US9384304B2 (en) Document search apparatus, document search method, and program product
JP2012043385A (en) Character recognition device and character recognition method
Li et al. A fast keyword-spotting technique
US20150063698A1 (en) Assisted OCR
TH103248A (en) Process for semi-automatic character image recognition
Tariq et al. Softconverter: A novel approach to construct OCR for printed Urdu isolated characters
Madhavaraj et al. Improved recognition of aged Kannada documents by effective segmentation of merged characters
KR20160053544A (en) Method for extracting candidate character
KR20160053587A (en) Method for minimizing database size of n-gram language model
JP5853488B2 (en) Information processing apparatus and program
US11270153B2 (en) System and method for whole word conversion of text in image