WO2014050481A1 - 文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム - Google Patents

文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム Download PDF

Info

Publication number
WO2014050481A1
WO2014050481A1 PCT/JP2013/073886 JP2013073886W WO2014050481A1 WO 2014050481 A1 WO2014050481 A1 WO 2014050481A1 JP 2013073886 W JP2013073886 W JP 2013073886W WO 2014050481 A1 WO2014050481 A1 WO 2014050481A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
character
ruby
character image
combined
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2013/073886
Other languages
English (en)
French (fr)
Inventor
浩教 矢野
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujifilm Corp
Original Assignee
Fujifilm Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujifilm Corp filed Critical Fujifilm Corp
Publication of WO2014050481A1 publication Critical patent/WO2014050481A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • G06T11/60Creating or editing images; Combining images with text

Definitions

  • the present invention relates to a document image processing apparatus, its operation control method, and its operation control program.
  • Non-Patent Document 1 “development of document image layout reconstruction technology“ GT-Layout ”for portable terminals” (Non-Patent Document 1) is known. This GT-Layout makes it possible to view documents by scrolling in one direction by rearranging the character positions according to the display screen from the document image and character position information, and configuring the document image to match the display screen size. To do.
  • Patent Document 1 it is possible to rearrange character patterns without performing character recognition (Patent Document 1), but not for a document image, but prohibiting separation of ruby and parent characters (Patent Document) 2, 3)
  • the document image processing apparatus includes a character image cutout unit that cuts out a character image representing a character included in an image from a document image obtained by imaging the document, and a character represented by the character image cut out by the character image cutout unit.
  • Ruby determination means for determining whether or not it is ruby
  • parent character image detection for detecting a parent character image representing a parent character to which ruby determined by the ruby determination means is shaken from the character image cut out by the character image cutout means
  • a combined character image generating unit that generates a combined character image by combining the parent character image detected by the parent character image detecting unit and the ruby image determined to represent ruby by the ruby determining unit. It is characterized by being.
  • the present invention also provides an operation control method suitable for a document image processing apparatus. That is, in this method, the character image cutout means cuts out the character image representing the character included in the image from the document image obtained by imaging the document, and the ruby determination means uses the character image cut out in the character image cutout means. It is determined whether or not the character to be represented is ruby, and the parent character image detecting unit represents the parent character to which the ruby determined by the ruby determining unit is assigned from the character image cut out by the character image cutting unit. And the combined character image generation unit generates a combined character image by combining the parent character image detected by the parent character image detection unit and the ruby image determined to represent ruby by the ruby determination unit. Is.
  • the present invention also provides a computer readable program for executing the operation control method of the document image processing apparatus.
  • a recording medium storing such a program may also be provided.
  • a character image is cut out from a document image, and it is determined whether or not the cut out character image represents ruby.
  • a parent character image to which ruby is shaken is also detected. Then, the detected parent character image and the ruby character image are combined to generate a combined character image. Since the ruby image is combined with the parent character image to form one combined character image, when the document image is displayed again in the desired display area, the combined character image may be positioned, and the ruby image is positioned. There is no need. It can be displayed relatively easily.
  • the ruby determination means includes determination means for determining whether the character image cut out by the character image cutout means constitutes a row or a column, for example. In this case, the number of characters is smaller than the number of characters in the other rows or columns and the size of the character image in the other rows in response to the determination unit determining that the ruby image constitutes a row or column. Is determined to be a ruby image, the ruby image exists between rows or columns according to the determination that the ruby image does not constitute a row or column, and is smaller than the size of other character images. In this case, it will be determined as a ruby image.
  • the combined character image generation means for example, has a plurality of combined character images in which the parent character image and the ruby image are combined, and when these combined character images are adjacent, the adjacent combined character images are displayed. Furthermore, it is what is combined.
  • the character image clipped by the character image cutout means and the combined character image generated by the combined character image generation means are positioned in the display area of the display screen according to the character arrangement in the document image, and cut out by the character image cutout means.
  • the center line of the displayed character image is preferably positioned in the display area of the display screen so that the center line of the combined character image generated by the combined character image generating means matches.
  • FIG. 1 shows an embodiment of the present invention, in which an electric image of a document image processing apparatus 1 for shaping a document image (an imaged document will be referred to as a document image) so that it can be displayed in a desired display area.
  • a document image an imaged document will be referred to as a document image
  • FIG. 1 shows an embodiment of the present invention, in which an electric image of a document image processing apparatus 1 for shaping a document image (an imaged document will be referred to as a document image) so that it can be displayed in a desired display area.
  • a document image processing apparatus 1 for shaping a document image an imaged document will be referred to as a document image
  • the overall operation of the document image processing apparatus 1 is controlled by the control apparatus 2.
  • the document image processing apparatus 1 includes an input device 3 such as a keyboard for inputting various commands, a communication device 4 for communicating with other client terminal devices, mobile phones, etc., a display device 5 for displaying a document image, etc. A memory 6 for storing the data is provided.
  • the document image processing apparatus 1 is provided with a CD (compact disk) driver 7. When a compact disk 8 storing a program for controlling operations to be described later is loaded into the CD driver 7 and the program stored in the compact disk 8 is read by the CD driver 7, the program is processed by the document image processing apparatus. 1 installed.
  • the communication device 4 may be used to receive a program, and the received program may be installed in the document image processing device 1.
  • the document image processing device 1 includes a character region acquisition device 11, a ruby extraction device 12, a region synthesis device 13, and a shaped image creation device 14.
  • the character area acquisition device 11 detects and extracts a character image area from a document image. Extraction of character images can use the function of OCR (Optical Character Reader). The coordinate position of the character image in the document image, the character type represented by the character image, the order of the characters, and whether the character is written horizontally or vertically are also detected.
  • OCR Optical Character Reader
  • the ruby extraction device 12 extracts a ruby image from the character image acquired by the character area acquisition device 11.
  • the extraction of ruby images can also use the OCR function.
  • the area composition device 13 combines the ruby image with the parent character image representing the parent character to which the ruby is shaken to generate a combined character image. Whether the image is a parent character image can use the OCR function. In addition, when viewed from the ruby character image, a vertical line is drawn downward for horizontal writing and to the left for vertical writing, and the closest character image in the closest row can be determined as the parent character image. The parent character image and the ruby image are combined to generate a combined character image. When the combined character images generated in this way are adjacent to each other, the adjacent combined character images are also combined. However, it goes without saying that the combined character images need not be combined.
  • the shaped image creation device 14 positions and shapes the character image obtained by the character region acquisition device 11 and the combined character image obtained by the region synthesis device 13 so as to be displayed on a display screen having a desired display area. An image is generated. The center line of the parent character is matched to balance the characters.
  • FIG. 2 is an example of an imaged document image 20.
  • the document image 20 includes characters 21 to 29, 31 to 39, and 41 and 42 represented by the image. These letters 21 to 29, 31 to 39 and 41 and 42 are represented by circles. These characters 21 to 29, 31 to 39 and 41 and 42 are not represented by text data, but are represented by images.
  • the document image 20 is shaped so as to be displayed on a display screen (display window) having a desired display area.
  • Characters 41 and 42 are ruby characters that are waved to characters 37 and 38.
  • Ruby is a character with a role such as phonetics, explanations, and different readings for any character in a sentence, usually written on the right side of a character in vertical writing and above the character in horizontal writing. Ruby is often given in Japanese or Chinese, but ruby may be given in other languages.
  • FIG. 3 is a flowchart showing the processing procedure of the document image processing apparatus 1.
  • document image data representing the document image 20 is stored in the memory 6.
  • processing such as extraction of a character image from the document image 20 is performed (step 81).
  • FIG. 4 shows a state in which character images 51 to 59, 61 to 69, 71 and 72 representing characters 21 to 29, 31 to 39, 41 and 42 are extracted from the document image 20.
  • the extraction of the character images 51 to 59, 61 to 69, and 71 and 72 uses the OCR function as described above.
  • the extracted character images 51 to 59, 61 to 69, 71 and 72 are surrounded by a rectangle.
  • the upper left coordinates of these rectangles are the coordinate positions of the character images 51 to 59, 61 to 69, 71 and 72.
  • the position of the character image 51 is represented by coordinates (x1, y1).
  • the position of the character image 61 is represented by coordinates (x11, y11).
  • the positions of the ruby images 71 and 72 are represented by coordinates (x17, y17) and (x18, y18), respectively.
  • Other character image positions are also represented by coordinates.
  • the widths and heights of the character images 51 to 59, 61 to 69, 71 and 72 are also detected.
  • the detected coordinates of the character images 51 to 59, 61 to 69, 71 and 72, etc. are stored in the character information table.
  • FIG. 6 is an example of a character information table.
  • the character information table shown in FIG. 6 is for the document image 20.
  • the X-coordinate, Y-coordinate, width, height of the character image, the type of character represented by the character image, and the parent character Stores data indicating ruby.
  • the ID stored in the character information table identifies the character images 51 to 59, 61 to 69, 71 and 72, respectively.
  • ID is the arrangement order of the document image 20 except for the ruby images 71 and 72.
  • the ID of the character image 21 is ID1
  • the X coordinate is x1
  • the Y coordinate is y1
  • the width is w
  • the height is h.
  • the type of character detected by the OCR function (what character is represented) is also included.
  • character image 51 (character 21) is neither a parent character nor ruby. It can be seen that character images 67 and 68 (characters 37 and 38) specified by ID16 and ID17 are parent characters, and character images 71 and 72 (characters 41 and 42) specified by ID19 and ID20 are ruby.
  • ruby determination processing is performed on the extracted character image (step 82).
  • the ruby determination process can use the OCR function as described above, and can use other methods described above. If the extracted character image includes a ruby image (YES in step 83), a parent image corresponding to the ruby image is detected (step 84). As described above, the OCR function can be used to detect the parent image, and other methods described above can also be used.
  • the ruby image and the parent image are detected, the ruby image and the parent image are combined to generate a combined character image (step 85).
  • FIG. 5 shows how a combined character image is generated.
  • the ruby image 71 is combined with the parent character image 67 to generate a combined character image 81. Since the base character of the ruby 42 is the base character 38, the ruby image 72 is combined with the base character image 68 to generate a combined character image 82.
  • the combined character images 81 and 82 in which the ruby and the parent character are integrated are generated.
  • the combined character images 81 and 82 generated in this way are adjacent, these combined character images 81 and 82 are combined.
  • a combined character image 80 is generated.
  • the parent character image representing a plurality of parent characters to which ruby is shaken is displayed, the parent character image is separated.
  • the combined character images 81 and 82 need not be combined to generate the combined character image 80.
  • FIG. 7 shows an example of the corrected character information table.
  • the ruby image 71 and the parent image 67 are combined to generate a combined character image 81, and the ruby image 71 and the parent image 68 are combined to generate a combined character image 82. It is an example of the character information table after correction in the case of being done.
  • the combined character image 81 is generated by combining the ruby image 71 with the parent image 67, the character information about the ruby image 71 is deleted, and the character information about the parent image 67 indicated by ID16 is replaced with the combined character image 81.
  • the combined character image 81 is a combination of the ruby image 71 and the parent image 67, the new coordinate position is set to x31 for the X coordinate and y31 for the Y coordinate, the width does not change, and the height is 6h / Has been changed to 5.
  • Data (parent + ruby) indicating that a parent character and a ruby character are included is also recorded in the character information table.
  • the ruby image 72 is combined with the parent image 68 to generate the combined character image 82
  • the character information about the ruby image 72 is deleted, and the character information about the parent image 68 indicated by ID17 is combined.
  • the character information about the character image 82 is used.
  • the new coordinate position of the combined character image 82 is set to x32 for the X coordinate and y32 for the Y coordinate, the width is not changed, and the height is changed to 6h / 5.
  • Data (parent + ruby) indicating that a parent character and a ruby character are included is also recorded in the character information table.
  • the corrected character information table is as shown in FIG.
  • the corrected character information table is as shown in FIG.
  • the character information about the combined character image 82 is deleted, and the character information about the combined character image 81 indicated by ID16 is changed. , Character information about the new combined character image 80.
  • the width of the new combined character image 80 is changed from w to 2w.
  • a ruby image and a parent image are detected and a combined character image is generated, or if a ruby image is not detected, the detected character image is positioned in the display area of the display screen to be displayed. (Step 86). As a result, processing for creating a shaped image is performed.
  • FIG. 9 shows a state in which the character image is positioned in the display area 50 corresponding to the desired display screen.
  • the width of the display area 50 is narrower than the width of the document image 20. Assuming that ruby 41 and 42 are not included in the number of lines, the document image 20 displays all of the character images 51 to 59 and 61 to 69 in two lines. Images 51 through 59 and 61 through 69 cannot all be displayed.
  • Character images 51 to 55 are positioned on the first line of the display area 50, character images 56 to 59 and 61 are positioned on the second line of the display area 50, and the third line of the display area 50 is positioned on the third line.
  • Character images 62 to 66 are positioned, and a combined character image 80 and a character image 69 are positioned in the fourth line of the display area 50. Positioning of these character images 51 to 59, 61 to 65 and 69 and the combined character image 80 is performed using the character information table shown in FIG. 8 so as to fit in the display area according to the character arrangement of the document image 20. Needless to say.
  • the combined character image 80 obtained by combining the combined character images 81 and 82 is positioned at the end of the line, and the combined character image 80 does not fit in the line, the combined character image 80 is Positioning may be performed so as to be positioned at the end of the line, and the character image of the entire line may be reduced at a predetermined reduction rate.
  • the character image positioned in the display area 50 in this way is displayed on the display screen 6 of the display device 5 (step 87).
  • FIG. 10 is a flowchart showing the ruby determination processing procedure, and shows the processing procedure of step 82 in FIG.
  • step 91 it is confirmed whether or not the extracted character image constitutes a line (step 91). Whether or not the character image forms a line can be determined by whether or not a plurality of character images having the same detected Y-coordinate position are arranged and the character images are arranged in the line direction at regular intervals. If it is determined that the character image constitutes a line (YES in step 91), the number of characters per line is counted (step 92). The number of characters per line can be easily determined from the character information table. When the number of characters per line is obtained, the average value of the number of characters per line is calculated, and the average value is used as the character number threshold value (step 93).
  • Ruby is rarely applied to all characters that make up a line, and the ruby for one line is generally less than the average number of characters in the line, so it is greater than the character count threshold. It is determined that the character in the character number line is not ruby (NO in step 94).
  • Step 95 it is further checked whether the size of the character image is smaller than the average character image size obtained from the character information table. Since the ruby is smaller than the average size of the characters included in the document image 20, if the size of the character image is smaller than the average size of the character image (YES in step 95), the character image is the ruby image. Is determined (step 96). If it is larger than the average character image size (NO in step 95), the character image is not a ruby image.
  • Whether or not a character image exists between lines can be determined from the sequence of Y coordinate positions of the character image obtained from the character information table. For example, when there are a plurality of character images having the same Y coordinate, they are considered to be character images constituting a line, and a character at a position sandwiched between such character images in the line direction (Y-axis direction). It can be determined that the image is a character image existing between lines.
  • FIG. 11 and FIG. 12 show a method for positioning the character image and the combined character image.
  • FIG. 11 shows how the combined character images 81 and 82 and the character image 69 are positioned in a state where the combined character image 81 and the combined character image 82 are not combined.
  • FIG. 12 shows a state in which the combined character image 80 and the character image 69 obtained by combining the combined character image 81 and the combined character image 82 are positioned.
  • the centers of the parent characters (parent character images) 37 and 38 included in the combined character image 80 and the centers of the characters 39 (character image 69) included in the character image 69 are displayed. Through a center line C extending in the horizontal direction. If the center of the parent character 37 and the center of the parent character 38 do not coincide with each other in the Y direction, the average Y coordinate of those centers and the Y coordinate of the center of the character image 60 are aligned.
  • one character image or combined character image is sequentially positioned in the display area 50 one by one.
  • a plurality of character images (including combined character images) corresponding to the width of the display area 50 are included. May be cut out from the document image, and the cut out character image group may be positioned in the display area 50.
  • the document image processing apparatus 1 performs processing for extracting a character image from the document image, processing for determining whether the extracted character image is ruby or a parent character, processing for generating a combined character image, and formatting.
  • Image creation processing and display processing on the display device 5 are performed.
  • Data representing the created shaped image is transmitted from the document image processing device 1 to another terminal device such as a mobile phone, and the terminal device.
  • the display process may be performed at.
  • the shaping image creation process may be performed in another terminal device.
  • the processing in the document image processing apparatus 1 may be executed by software using a server instead of a dedicated apparatus, or may be executed by a mobile phone such as a smartphone.
  • the horizontally written document image has been described.
  • the embodiment can be similarly applied to a vertically written document image instead of horizontally written.
  • vertical writing it may be read as a column instead of a row.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Document Processing Apparatus (AREA)

Abstract

 画像化された文字画像にルビが振られていた場合でも,所望の表示領域に比較的簡単に表示し直すようにする。 文書画像から文字画像が抽出され(ステップ81),抽出された文字画像についてルビ判定処理が行われる(ステップ82)。文字画像がルビ画像であると,そのルビ画像の親文字画像も検出される(ステップ84)。ルビ画像と親文字画像とが結合され,結合文字画像が生成される(ステップ85)。検出された文字画像,生成された結合文字画像が,所望の表示領域に位置決めされ(ステップ86),表示される(ステップ87)。ルビ画像と親文字画像とが結合されるので,ルビ画像と親文字画像とが分離されてしまうことが未然に防止される。

Description

文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム
 この発明は,文書画像処理装置ならびにその動作制御方法およびその動作制御プログラムに関する。
 文書画像や固定型レイアウトの文書ファイルを携帯端末で閲覧した場合,文書内の段落のサイズが携帯端末の表示画面サイズよりも大きいため,継続した閲覧には段落領域のスクロールを必要とする。このため携帯端末における文書閲覧では閲覧行為と端末操作行為とを交互に意識する必要があり,閲覧行為のみを継続することで得られる快適な文書閲覧ができなくなる。このような問題を解決するために「携帯端末向け文書画像レイアウト再構成技術「GT-Layout」の開発」(非特許文献1)が知られている。このGT-Layoutは,文書画像と文字位置情報とから,表示画面にあわせて文字位置を並べ替え,表示画面サイズに合った文書画像を構成することで一方向のスクロールにより文書の閲覧を可能にするものである。
 また,文字認識を行うことなく文字パターンの再配置を行うことができるもの(特許文献1),文書画像についてのものではないが,ルビと親文字との分離を禁止しているもの(特許文献2,3)などもある
http://www.fujifilm.co.jp/rd/report/rd057/pack/pdf/ff_rd057_009.pdf 特開平5-266168号公報 特開2000-305552号公報 特開2000-305926号公報
 文書画像を所望の表示領域に表示しなおす場合において,文書画像に含まれる文字画像にルビが振られていた場合,ルビ画像を親文字画像に対応させて表示させる必要がある。しかしながら,親文字画像とルビ画像とのバランスをとるためには,親文字画像との位置関係を考慮してルビ画像の位置を決定しなければならないので,表示までの時間が長くなってしまうことがある。
 この発明は,文書画像を所望の表示領域に表示しなおす場合において文書画像に含まれている文字画像にルビが振られていたときでも比較的簡単に表示しなおすことができるようにすることを目的とする。
 この発明による文書画像処理装置は,文書が画像化された文書画像から,画像に含まれる文字を表わす文字画像を切り出す文字画像切り出し手段,文字画像切り出し手段において切り出された文字画像によって表わされる文字がルビかどうかを判定するルビ判定手段,上記文字画像切り出し手段において切り出された文字画像から,ルビ判定手段において判定されたルビが振られている親文字を表わす親文字画像を検出する親文字画像検出手段,および親文字画像検出手段によって検出された親文字画像とルビ判定手段においてルビを表わしていると判定されたルビ画像とを結合して結合文字画像を生成する結合文字画像生成手段を備えていることを特徴とする。
 この発明は,文書画像処理装置に適した動作制御方法も提供している。すなわち,この方法は,文字画像切り出し手段が,文書が画像化された文書画像から,画像に含まれる文字を表わす文字画像を切り出し,ルビ判定手段が,文字画像切り出し手段において切り出された文字画像によって表わされる文字がルビかどうかを判定し,親文字画像検出手段が,文字画像切り出し手段において切り出された文字画像から,ルビ判定手段において判定されたルビが振られている親文字を表わす親文字画像を検出し,結合文字画像生成手段が,親文字画像検出手段によって検出された親文字画像とルビ判定手段においてルビを表わしていると判定されたルビ画像とを結合して結合文字画像を生成するものである。
 この発明は,文書画像処理装置の動作制御方法を実施するためのコンピュータが読み取り可能なプログラムも提供している。そのようなプログラムを格納した記録媒体も提供するようにしてもよい。
 この発明によると,文書画像から文字画像が切り出され,切り出された文字画像がルビを表わすものかどうか判定される。また,ルビが振られている親文字画像も検出される。すると,検出された親文字画像とルビ文字画像とが結合され,結合文字画像が生成される。ルビ画像は親文字画像と結合され,一つの結合文字画像とされるので,所望の表示領域に文書画像を表示しなおす場合には,その結合文字画像を位置決めすればよく,ルビ画像を位置決めする必要は無い。比較的簡単に表示しなおすことができる。
 ルビ判定手段は,たとえば,文字画像切り出し手段によって切り出された文字画像が行または列を構成しているかどうかを判定する判定手段を備えている。この場合,判定手段によってルビ画像が行または列を構成していると判定されたことに応じて,他の行または列の文字数と比べて文字数が少なく,かつ他の行の文字画像の大きさよりも小さい場合にルビ画像と判定し,判定手段によってルビ画像が行または列を構成していないと判定されたことに応じて行間または列間に存在し,かつ他の文字画像の大きさよりも小さい場合にルビ画像と判定するものとなろう。
 結合文字画像生成手段は,たとえば,親文字画像とルビ画像とが結合された結合文字画像が複数あり,それらの複数の結合文字画像が隣接している場合には隣接している結合文字画像をさらに結合するものである。
 文字画像切り出し手段において切り出された文字画像および結合文字画像生成手段によって生成された結合文字画像を文書画像における文字配列にしたがって表示画面の表示領域に位置決めするものであって,文字画像切り出し手段において切り出された文字画像の中心線と結合文字画像生成手段によって生成された結合文字画像の中心線とが一致するように,表示画面の表示領域に位置決めすることが好ましい。
文書画像処理装置の電気的構成を示すブロック図である。 文書画像の一例である。 文書画像処理装置の処理手順を示すフローチャートである。 文書画像の一例である。 結合文字画像が生成される様子を示している。 文字情報テーブルの一例である。 文字情報テーブルの一例である。 文字情報テーブルの一例である。 表示領域に文字画像が位置決めされた様子を示している。 ルビ判定処理手順を示すフローチャートである。 結合文字画像が位置決めされる様子を示している。 結合文字画像が位置決めされる様子を示している。
 図1は,この発明の実施例を示すもので,文書画像(画像化された文書を文書画像ということにする)を,所望の表示領域に表示できるように整形する文書画像処理装置1の電気的構成を示すブロック図である。
 文書画像処理装置1の全体の動作は,制御装置2によって統括される。
 文書画像処理装置1には,種々の指令を入力するキーボードなどの入力装置3,他のクライアント端末装置,携帯電話などと通信するための通信装置4,文書画像等を表示する表示装置5,所定のデータを記憶するメモリ6などが設けられている。また,文書画像処理装置1には,CD(コンパクト・ディスク)ドライバ7が設けられている。後述する動作を制御するプログラムが格納されているコンパクト・ディスク8がCDドライバ7に装填され,コンパクト・ディスク8に格納されているプログラムがCDドライバ7によって読み取られると,そのプログラムが文書画像処理装置1にインストールされる。もっとも,通信装置4を利用してプログラムを受信し,受信したプログラムが文書画像処理装置1にインストールされるようにしてもよい。
 さらに,文書画像処理装置1には,文字領域取得装置11,ルビ抽出装置12,領域合成装置13および整形画像作成装置14が含まれている。
 文字領域取得装置11は,文書画像から文字画像の領域を検出して抽出するものである。文字画像の抽出は,OCR(Optical Character Reader)の機能を利用できる。文書画像における文字画像の座標位置,文字画像によって表わされる文字の種類,文字の並び順,文字が横書きか,縦書きかどうかも検出される。
 ルビ抽出装置12は,文字領域取得装置11によって取得された文字画像のうちルビ画像を抽出するものである。ルビ画像の抽出もOCRの機能を利用することができる。また,抽出された文字画像が行を構成しているかどうかを判定し,抽出された文字画像が行を構成していると判定された場合には,他の行または列の文字数と比べて文字数が少なく,かつ他の行の文字画像の大きさよりも小さいときにルビ画像と判定し,文字画像が行を構成していないと判定されたことに応じて行間に存在し,かつ他の文字画像の大きさよりも小さい場合にルビ画像と判定するようにしてもよい。さらに,平均行間隔を検出し,その平均行間隔以下で文字画像に隣接している場合であって,他の行の平均的な文字数わりも少ないときにルビ画像と判定するようにしてもよい。
 領域合成装置13は,ルビ画像を,そのルビが振られている親文字を表わす親文字画像に結合して結合文字画像を生成するものである。親文字画像かどうかはOCR機能を利用することができる。また,ルビ文字画像から見て横書きなら下,縦書きなら左に垂線を引き,最近接の行内でかつ最近接の文字画像を親文字画像と判断することもできる。親文字画像とルビ画像とが結合されて結合文字画像が生成される。このようにして生成された結合文字画像が隣接している場合には,それらの隣接している結合文字画像同士も結合される。もっとも,結合文字画像同士を結合しなくともよいのはいうまでもない。
 整形画像作成装置14は,所望の表示領域をもつ表示画面に表示できるように,文字領域取得装置11において得られた文字画像および領域合成装置13において得られた結合文字画像の位置決めをして整形画像を生成するものである。文字間のバランスをとるために,親文字の中心線が一致させられる。
 図2は,画像化された文書画像20の一例である。
 文書画像20には,画像によって表わされている文字21から29,31から39ならびに41および42が含まれている。これらの文字21から29,31から39ならびに41および42は,丸印によって表現されている。これらの文字21から29,31から39ならびに41および42は,テキスト・データによって表わされているものではなく,画像によって表わされているものである。この実施例では,この文書画像20が所望の表示領域をもつ表示画面(表示ウインドウ)に表示されるように整形される。
 文字41および42は,文字37および38に振られているルビである。ルビは,文章内の任意の文字に対しふりがな,説明,異なる読み方といった役割の文字をより小さな文字で,通常,縦書きの際は文字の右側に横書きの際は文字の上側に記される。日本語や中国語にルビが振られることが多いが,その他の言語にルビが振られてもよい。
 図3は,文書画像処理装置1の処理手順を示すフローチャートである。
 メモリ6には,文書画像20を表わす文書画像データが格納されているものとする。図2に示したように,文書画像20から,文字画像の抽出等の処理が行われる(ステップ81)。
 図4は,文書画像20から,文字21から29,31から39ならびに41および42を表わす文字画像51から59,61から69ならびに71および72が抽出された様子を示している。文字画像51から59,61から69ならびに71および72の抽出は,上述のようにOCRの機能が利用される。抽出された文字画像51から59,61から69ならびに71および72は,矩形で囲まれている。文書画像20の左上の頂点を原点(X0,Y0)としたときに,これらの矩形の左上の座標が,文字画像51から59,61から69ならびに71および72の座標位置となる。たとえば,文字画像51の位置は,座標(x1,y1)で表わされる。同様に,文字画像61の位置は,座標(x11,y11)で表わされる。さらに,ルビ画像71および72の位置は,それぞれ座標(x17,y17)および(x18,y18)で表わされる。その他の文字画像の位置についても座標で表わされる。また,文字画像51から59,61から69ならびに71および72の幅および高さも検出される。検出された文字画像51から59,61から69ならびに71および72の座標等は,文字情報テーブルに格納される。
 図6は,文字情報テーブルの一例である。
 図6に示す文字情報テーブルは,文書画像20についてのものである。
 文字情報テーブルには,検出された文字画像を識別するためのIDごとに,文字画像のX座標,Y座標,幅,高さ,ならびに文字画像によって表わされている文字の種類および親文字かルビかを示すデータが格納されている。文字情報テーブルに格納されているIDは,文字画像51から59,61から69ならびに71および72を,それぞれ識別するものである。IDは,ルビ画像71および72を除いて文書画像20の配列順である。たとえば,文字画像21のIDは,ID1であり,X座標はx1,Y座標はy1,幅はw,高さはhである。また,OCR機能によって検出された文字の種類(どのような文字を表わしているか)も含まれている。さらに,文字画像51(文字21)は親文字でもルビでもないということが分る。ID16およびID17で特定される文字画像67および68(文字37および38)は親文字であり,ID19およびID20で特定される文字画像71および72(文字41および42)はルビであることがわかる。
 図3に戻って,抽出された文字画像についてルビ判定処理が行われる(ステップ82)。ルビ判定処理は,上述したようにOCRの機能を利用できるし,上述したその他の方法も利用できる。抽出された文字画像にルビ画像が含まれていると(ステップ83でYES),そのルビ画像に対応する親画像が検出される(ステップ84)。親画像の検出も,上述したようにOCRの機能を利用できるし,上述したその他の方法も利用できる。
 ルビ画像と親画像とが検出されると,それらのルビ画像と親画像とが結合されて,結合文字画像が生成される(ステップ85)。
 図5は,結合文字画像が生成される様子を示している。
 ルビ41の親文字は親文字37であることから,ルビ画像71は親文字画像67と結合されて結合文字画像81が生成される。また,ルビ42の親文字は親文字38であることから,ルビ画像72は親文字画像68と結合されて結合文字画像82が生成される。ルビと親文字とが一体となった結合文字画像81および82が生成されることとなる。
 さらに,この実施例では,このようにして生成された結合文字画像81および82が隣接している場合には,それらの結合文字画像81と82とが結合される。結合文字画像81と結合文字画像82とが結合されることにより,結合文字画像80が生成される。ルビが振られている複数の親文字を表わす親文字画像を表示するときに親文字画像が分離されてしまうことが未然にされる。もっとも,結合文字画像81と82とを結合して結合文字画像80を生成しなくともよい。
 文字画像にルビ画像が含まれていない場合には(ステップ83でNO),ステップ84および85の処理はスキップされる。
 上述したように結合文字画像が生成されると,上述した文字情報テーブルが修正される。
 図7は,修正後の文字情報テーブルの一例である。
 図7は,図5に示すように,ルビ画像71と親画像67とが結合されて結合文字画像81が生成され,かつルビ画像71と親画像68とが結合されて結合文字画像82が生成された場合の修正後の文字情報テーブルの一例である。
 ルビ画像71が親画像67と結合されて結合文字画像81が生成されたことにより,ルビ画像71についての文字情報が削除され,ID16で示される親画像67についての文字情報が,結合文字画像81についての文字情報とされている。結合文字画像81は,ルビ画像71と親画像67とが結合されたものであるので,新たな座標位置はX座標がx31,Y座標がy31とされ,幅は変わらず,高さは6h/5と変更されている。また,親文字とルビ文字を含むことを示すデータ(親+ルビ)も文字情報テーブルに記録される。
 同様に,ルビ画像72が親画像68と結合されて結合文字画像82が生成されたことにより,ルビ画像72についての文字情報が削除され,ID17で示される親画像68についての文字情報が,結合文字画像82についての文字情報とされている。結合文字画像82の新たな座標位置はX座標がx32,Y座標がy32とされ,幅は変わらず,高さは6h/5と変更されている。親文字とルビ文字を含むことを示すデータ(親+ルビ)も文字情報テーブルに記録される。
 結合文字画像81と82とが結合されない場合には,修正後の文字情報テーブルは図7に示すものとなる。これに対して,結合文字画像81と82とが結合される場合には,修正後の文字情報テーブルは図8に示すものとなる。
 結合文字画像81と結合文字画像82とが結合されて結合文字画像80が生成されたことにより,結合文字画像82についての文字情報が削除され,ID16で示される結合文字画像81についての文字情報が,新たな結合文字画像80についての文字情報とされている。新たな結合文字画像80の幅はwから2wとなる。
 図3に戻って,ルビ画像および親画像が検出され,結合文字画像が生成される,あるいはルビ画像が検出されないと,検出された文字画像等が表示しようとする表示画面の表示領域に位置決めされる(ステップ86)。これにより整形画像の作成処理が行われることとなる。
 図9は,所望の表示画面に対応する表示領域50に文字画像が位置決めされた様子を示している。
 表示領域50の横幅は,文書画像20の横幅よりも狭い。ルビ41および42を行数に含めないと考えると,文書画像20では2行に文字画像51から59および61から69のすべてが表示されているが,表示領域50の2行にはそれらの文字画像51から59および61から69のすべてを表示することはできない。
 表示領域50の第1行目には文字画像51から55が位置決めされ,表示領域50の第2行目には文字画像56から59および61が位置決めされ,表示領域50の第3行目には文字画像62から66が位置決めされ,表示領域50の第4行目には結合文字画像80および文字画像69が位置決めされている。これらの文字画像51から59,61から65および69ならびに結合文字画像80の位置決めは,図8に示す文字情報テーブルを利用して文書画像20の文字配列どおりに表示領域に収まるように行われるのはいうまでもない。結合文字画像81と82とを結合して得られた結合文字画像80が行の終わりの部分に位置決めされることにより,その行内に結合文字画像80が収まらない場合には,結合文字画像80が行の終わりに位置するように位置決めし,かつその行全体の文字画像を所定の縮小率で縮小するようにしてもよい。
 このようにして表示領域50内に位置決めされた文字画像が表示装置5の表示画面6に表示されることとなる(ステップ87)。
 図10は,ルビ判定処理手順を示すフローチャートであり,図3のステップ82の処理手順を示している。
 まず,抽出された文字画像が行を構成しているかどうかが確認される(ステップ91)。文字画像が行を構成しているかどうかは,検出されたY座標位置が共通する文字画像が複数個並んでおり,かつ一定間隔で文字画像が行方向に並んでいるかどうかで判断できる。文字画像が行を構成していると判断されると(ステップ91でYES),行ごとの文字数がカウントされる(ステップ92)。行ごとの文字数は,文字情報テーブルから容易に分る。行ごとの文字数が得られると,行ごとの文字数の平均値が算出され,その平均値が文字数しきい値とされる(ステップ93)。行を構成するすべての文字についてルビが振られることはほとんど無く,1行分のルビは行の文字数の平均値よりも少ないことが一般的であると考えられるので,文字数しきい値よりも多い文字数の行の文字はルビではないと判断される(ステップ94でNO)。
 文字数しきい値以下の文字数の行の場合には(ステップ94でYES),さらに文字画像の大きさが,文字情報テーブルから得られる平均的な文字画像の大きさよりも小さいかどうかが確認される(ステップ95)。ルビは文書画像20に含まれる文字の平均的な大きさよりも小さいので,平均的な文字画像の大きさよりも小さい文字画像の大きさであれば(ステップ95でYES),その文字画像がルビ画像と決定される(ステップ96)。平均的な文字画像の大きさ以上であれば(ステップ95でNO),その文字画像はルビ画像とはされない。
 文字画像が行を構成していると判断されない場合であっても(ステップ91でNO),文字画像が行間に存在するものであり(ステップ97でYES),かつ文字画像の大きさが小さければ(ステップ95でYES),ルビ画像と決定される(ステップ96)。文字画像が行間に存在するかどうかは,文字情報テーブルから得られる文字画像のY座標位置の並びから判断できる。たとえば,同一のY座標をもつ文字画像が複数ある場合には,それが行を構成する文字画像と考えられ,行方向(Y軸方向)において,そのような文字画像に挟まれる位置にある文字画像が行間に存在する文字画像と判断できる。
 図11および図12は,文字画像と結合文字画像との位置決めの方法を示している。
 図11は,結合文字画像81と結合文字画像82とを結合していない状態での結合文字画像81および82ならびに文字画像69を位置決めする様子を示している。
 結合文字画像81および82ならびに文字画像69が位置決めされる場合,結合文字画像81および82のそれぞれに含まれる親文字(親文字画像)37および38の中心と文字画像69に含まれる文字39(文字画像69)の中心と水平方向に伸びた中心線Cを通るようにされる(中心のY座標が一致させられる)。ルビ41および42を含む結合文字画像81および82のそれぞれの中心と文字画像69の中心とが水平方向に伸びた中心線Cを通るように位置決めされると,親文字37および38の位置が文字画像69に含まれる文字39の位置よりも下がってしまい,バランスが悪くなってしまう。そのようなアンバランスが未然に防止される。結合文字画像81および82ならびに文字画像69の底辺が一致するように位置決めされてもよい。
 図12は,結合文字画像81と結合文字画像82とが結合された結合文字画像80と文字画像69とを位置決めする様子を示している。
 結合文字画像80および文字画像69が位置決めされる場合も,結合文字画像80に含まれる親文字(親文字画像)37および38の中心と文字画像69に含まれる文字39(文字画像69)の中心とが水平方向に伸びた中心線Cを通るようにされる。親文字37の中心と親文字38の中心がY方向の位置が一致しない場合には,それらの中心のY座標の平均と文字画像60の中心のY座標とが一致するように位置決めされる。
 上述した実施例では,一つの文字画像または結合文字画像を表示領域50に一つずつ順に位置決めしていくものであるが,表示領域50の横幅に対応した複数の文字画像(結合文字画像が含まれる場合もあるし,含まれない場合もある)から構成される一つの文字画像群を文書画像から切り出し,その切り出された文字画像群を表示領域50に位置決めするようにしてもよい。
 また,上述の実施例では,文書画像処理装置1において,文書画像から文字画像を抽出する処理,抽出した文字画像がルビまたは親文字かどうかを判定する処理,結合文字画像を生成する処理,整形画像の作成処理および表示装置5への表示処理が行われているが,作成された整形画像を表わすデータを文書画像処理装置1から,携帯電話などの他の端末装置に送信し,その端末装置において表示処理が行われるようにしてもよい。また,整形画像の作成処理も他の端末装置において行われるようにしてもよい。さらに,文書画像処理装置1における処理は,専用装置でなく,サーバを利用したソフトウエアによって実行するようにしてもよいし,スマートフォンのような携帯電話において実行するようにしてもよい。
 さらに,上述の実施例では,横書きの文書画像について説明したが,横書きでなく,縦書きの文書画像についても同様に実施できる。縦書きの場合,行の変わりに列と読み替えればよいこととなろう。
1 文書画像処理装置
2 制御装置
11 文字領域取得装置
12 ルビ抽出装置
13 領域合成装置
14 整形画像作成装置
20 文書画像

Claims (6)

  1.  文書が画像化された文書画像から,画像に含まれる文字を表わす文字画像を切り出す文字画像切り出し手段,
     上記文字画像切り出し手段において切り出された文字画像によって表わされる文字がルビかどうかを判定するルビ判定手段,
     上記文字画像切り出し手段において切り出された文字画像から,上記ルビ判定手段において判定されたルビが振られている親文字を表わす親文字画像を検出する親文字画像検出手段,および
     上記親文字画像検出手段によって検出された親文字画像と上記ルビ判定手段においてルビを表わしていると判定されたルビ画像とを結合して結合文字画像を生成する結合文字画像生成手段,
     を備えた文書画像処理装置。
  2.  上記ルビ判定手段は,
     上記文字画像切り出し手段によって切り出された文字画像が行または列を構成しているかどうかを判定する判定手段,
     上記判定手段によってルビ画像が行または列を構成していると判定されたことに応じて,他の行または列の文字数と比べて文字数が少なく,かつ他の行の文字画像の大きさよりも小さい場合にルビ画像と判定し,上記判定手段によってルビ画像が行または列を構成していないと判定されたことに応じて行間または列間に存在し,かつ他の文字画像の大きさよりも小さい場合にルビ画像と判定するものである,
     請求項1に記載の文書画像処理装置。
  3. 上記結合文字画像生成手段は,
     親文字画像とルビ画像とが結合された結合文字画像が複数あり,それらの複数の結合文字画像が隣接している場合には隣接している結合文字画像をさらに結合するものである,
     請求項1または2に記載の文書画像処理装置。
  4.  上記文字画像切り出し手段において切り出された文字画像および上記結合文字画像生成手段によって生成された結合文字画像を文書画像における文字配列にしたがって表示画面の表示領域に位置決めするものであって,上記文字画像切り出し手段において切り出された文字画像の中心線と上記結合文字画像生成手段によって生成された結合文字画像の中心線とが一致するように,表示画面の表示領域に位置決めするものである,
     請求項1から3のうち,いずれか一項に記載の文書画像処理装置。
  5.  文字画像切り出し手段が,文書が画像化された文書画像から,画像に含まれる文字を表わす文字画像を切り出し,
     ルビ判定手段が,上記文字画像切り出し手段において切り出された文字画像によって表わされる文字がルビかどうかを判定し,
     親文字画像検出手段が,上記文字画像切り出し手段において切り出された文字画像から,上記ルビ判定手段において判定されたルビが振られている親文字を表わす親文字画像を検出し,
     結合文字画像生成手段が,上記親文字画像検出手段によって検出された親文字画像と上記ルビ判定手段においてルビを表わしていると判定されたルビ画像とを結合して結合文字画像を生成する,
     文書画像処理装置の動作制御方法。
  6.  文書画像処理装置のコンピュータを制御するコンピュータが読み取り可能なプログラムであって,
     文書が画像化された文書画像から,画像に含まれる文字を表わす文字画像を切り出させ,
     切り出された文字画像によって表わされる文字がルビかどうかを判定させ,
     切り出された文字画像から,判定されたルビが振られている親文字を表わす親文字画像を検出させ,
     検出された親文字画像とルビを表わしていると判定されたルビ画像とを結合して結合文字画像を生成させるように文書画像処理装置のコンピュータを制御するプログラム。
PCT/JP2013/073886 2012-09-26 2013-09-05 文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム Ceased WO2014050481A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2012-211632 2012-09-26
JP2012211632 2012-09-26

Publications (1)

Publication Number Publication Date
WO2014050481A1 true WO2014050481A1 (ja) 2014-04-03

Family

ID=50387889

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2013/073886 Ceased WO2014050481A1 (ja) 2012-09-26 2013-09-05 文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム

Country Status (1)

Country Link
WO (1) WO2014050481A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018142286A (ja) * 2017-02-28 2018-09-13 シナノケンシ株式会社 電子図書製作用プログラム

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0594511A (ja) * 1991-10-02 1993-04-16 Ricoh Co Ltd 画像処理装置
JPH1031716A (ja) * 1996-05-13 1998-02-03 Matsushita Electric Ind Co Ltd 文字行抽出方法および装置
JP2004005453A (ja) * 2002-03-01 2004-01-08 Xerox Corp 文書画像レイアウトの解体と再表示の方法およびシステム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0594511A (ja) * 1991-10-02 1993-04-16 Ricoh Co Ltd 画像処理装置
JPH1031716A (ja) * 1996-05-13 1998-02-03 Matsushita Electric Ind Co Ltd 文字行抽出方法および装置
JP2004005453A (ja) * 2002-03-01 2004-01-08 Xerox Corp 文書画像レイアウトの解体と再表示の方法およびシステム

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018142286A (ja) * 2017-02-28 2018-09-13 シナノケンシ株式会社 電子図書製作用プログラム

Similar Documents

Publication Publication Date Title
US12412150B2 (en) Information processing apparatus, control method, and program
US20160154579A1 (en) Handwriting input apparatus and control method thereof
US20130187923A1 (en) Legend indicator for selecting an active graph series
US20140164963A1 (en) User configurable subdivision of user interface elements and full-screen access to subdivided elements
JP2006079220A (ja) 画像検索装置および方法
JP5654851B2 (ja) 文書画像表示装置ならびにその動作制御方法およびその制御プログラム
US20140157116A1 (en) Method and Device for Determining a Display Mode of Electronic Documents
EP3125087A1 (en) Terminal device, display control method, and program
JP6109688B2 (ja) 帳票読取装置およびプログラム
JP6287498B2 (ja) 電子ホワイトボード装置、電子ホワイトボードの入力支援方法、及びプログラム
US20160209988A1 (en) Information Input Device, Control Method and Storage Medium
US8824806B1 (en) Sequential digital image panning
WO2014050481A1 (ja) 文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム
US9377853B2 (en) Information processing apparatus and information processing method
CN107797736A (zh) 信息显示方法和装置
CN109032476B (zh) 一种在图形用户界面中显示大数据集的方法
JP5794154B2 (ja) 画像処理プログラム、画像処理方法、及び画像処理装置
JP6160359B2 (ja) 情報処理装置、プログラム、および情報提示方法
CN104463153B (zh) 一种提高版式文档中字符识别率的方法和系统
JP6146222B2 (ja) 手書き入力装置およびプログラム
JP2009258972A (ja) 帳票イメージ表示装置
CN108346126B (zh) 基于内存拷贝方式绘制手机图片的方法及装置
JP2016081426A (ja) 文書処理装置、その制御方法、およびプログラム
KR102419696B1 (ko) 스크롤 제어 방법, 장치, 프로그램 및 컴퓨터 판독가능 기록매체
WO2014050480A1 (ja) 文書画像処理装置ならびにその動作制御方法およびその動作制御プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13841061

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13841061

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: JP