WO2017071061A1 - 区域识别方法及装置 - Google Patents

区域识别方法及装置 Download PDF

Info

Publication number
WO2017071061A1
WO2017071061A1 PCT/CN2015/099286 CN2015099286W WO2017071061A1 WO 2017071061 A1 WO2017071061 A1 WO 2017071061A1 CN 2015099286 W CN2015099286 W CN 2015099286W WO 2017071061 A1 WO2017071061 A1 WO 2017071061A1
Authority
WO
WIPO (PCT)
Prior art keywords
predetermined
edge
region
predetermined edge
target
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/099286
Other languages
English (en)
French (fr)
Inventor
龙飞
张涛
陈志军
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Xiaomi Inc
Original Assignee
Xiaomi Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Xiaomi Inc filed Critical Xiaomi Inc
Priority to JP2017547044A priority Critical patent/JP6392467B2/ja
Priority to KR1020167005597A priority patent/KR101758580B1/ko
Priority to RU2016109047A priority patent/RU2633184C2/ru
Priority to MX2016003318A priority patent/MX374325B/es
Publication of WO2017071061A1 publication Critical patent/WO2017071061A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/40Document-oriented image-based pattern recognition
    • G06V30/41Analysis of document content
    • G06V30/414Extracting the geometrical structure, e.g. layout tree; Block segmentation, e.g. bounding boxes for graphics or text
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/40Document-oriented image-based pattern recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/13Edge detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/25Determination of region of interest [ROI] or a volume of interest [VOI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/74Image or video pattern matching; Proximity measures in feature spaces
    • G06V10/75Organisation of the matching processes, e.g. simultaneous or sequential comparisons of image or video features; Coarse-fine approaches, e.g. multi-scale approaches; using context analysis; Selection of dictionaries
    • G06V10/758Involving statistics of pixels or of feature values, e.g. histogram matching
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20016Hierarchical, coarse-to-fine, multiscale or multiresolution image processing; Pyramid transform
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20024Filtering details
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30176Document
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/172Classification, e.g. identification
    • G06V40/173Classification, e.g. identification face re-identification, e.g. recognising unknown faces across different face tracks

Definitions

  • the present disclosure relates to the field of image processing, and in particular, to an area recognition method and apparatus.
  • the automatic identification technology of the ID card is a technology for recognizing the text information on the ID card through image processing.
  • the related art provides an automatic identification method for an identity card, which scans an ID card according to a fixed relative position by an ID card scanning device, and obtains a scanned image of the ID card; and performs character recognition on n predetermined regions in the scanned image. At least one of name information, gender information, ethnic information, date of birth information, address information, and citizenship number information.
  • ID images obtained directly from the camera there is still a large difficulty in recognition.
  • the present disclosure provides an area identification method and apparatus.
  • the technical solution is as follows:
  • a region identification method comprising:
  • the predetermined edge being an edge located in a predetermined direction of the document
  • a predetermined edge of one of the predetermined edges of the n candidates is determined as a target predetermined edge, n ⁇ 2;
  • At least one information area is identified in the document image according to the target predetermined edge.
  • determining a predetermined edge of one of the predetermined edges of the n candidates as the target predetermined edge includes:
  • the first relative positional relationship is a relative positional relationship between the target predetermined edge and the target information region.
  • the predetermined edge of the i-th candidate is determined as the target predetermined edge
  • the predetermined edges of the n candidates are ordered, including:
  • the predetermined edge For each candidate predetermined edge, the predetermined edge is intersected with the foreground pixel at the same position in the processed document image to obtain the intersection number corresponding to the predetermined edge; the processed document image is subjected to Sobel horizontal filtering and two Valued image
  • the predetermined edges of the n candidates are sorted in order of the number of intersections in descending order.
  • using the predetermined edge of the i-th candidate and the first relative positional relationship to attempt to identify the target information area in the document image includes:
  • a character area conforming to a predetermined feature is identified in the region of interest, the predetermined feature being a feature of the character region in the target information region.
  • identifying whether there is a character area in the region of interest that meets the predetermined feature includes:
  • the height of the continuous row set formed by the row of the foreground pixel in the first histogram is greater than the predetermined threshold, and the accumulated value of the foreground pixel in the second histogram is greater than the second threshold.
  • the accumulated value of the foreground pixel in the first histogram is greater than the continuous line of the first threshold If the height of the set does not meet the predetermined height interval, or the number of consecutive column sets formed by the column of the foreground pixel pixel in the second histogram is greater than the predetermined number, the number of consecutive column sets does not meet the predetermined number. A character area that conforms to a predetermined feature is identified.
  • identifying a predetermined edge in the document image includes:
  • determining at least one information area in the document image according to the target predetermined edge comprises:
  • the second relative positional relationship is a relative positional relationship between the target predetermined edge and the information area.
  • an area identifying apparatus comprising:
  • An identification module configured to identify a predetermined edge in the document image, the predetermined edge being an edge located in a predetermined direction of the document;
  • a determining module configured to determine, as a target predetermined edge, a predetermined edge of one of the predetermined edges of the n candidates when the predetermined edges of the n candidates are identified, n ⁇ 2;
  • the area identification module is configured to identify at least one information area in the document image according to the target predetermined edge.
  • the determining module comprises:
  • a first sorting sub-module configured to sort predetermined edges of the n candidates
  • a first identification sub-module configured to use the predetermined edge of the i-th candidate and the first relative positional relationship to attempt to identify the target information region in the document image, 1 ⁇ i ⁇ n; the first relative positional relationship is the target predetermined edge and The relative positional relationship between the target information areas.
  • a second identification submodule configured to determine a predetermined edge of the i th candidate as a target predetermined edge when the target information region is successfully identified
  • the first sorting submodule includes:
  • the processed document image is an image that has been Souter horizontally filtered and binarized
  • the second sorting sub-module is configured to sort the predetermined edges of the n candidates in order of the number of intersections.
  • the first identification submodule includes:
  • Intercepting a sub-module configured to use the predetermined edge of the i-th candidate and the first relative positional relationship to intercept the region of interest in the document image;
  • the fourth identification submodule is configured to identify whether there is a character region in the region of interest that conforms to the predetermined feature, and the predetermined feature is a feature of the character region in the target information region.
  • the fourth identification submodule includes:
  • a binarization sub-module configured to binarize the region of interest to obtain a binarized region of interest
  • a first calculation submodule configured to calculate a first histogram according to a horizontal direction for the binarized region of interest, the first histogram comprising: a vertical coordinate of each row of pixels and a foreground pixel in each row of pixels The accumulated value of the point;
  • a second calculation sub-module configured to calculate a second histogram in the vertical direction for the binarized region of interest, the second histogram comprising: an abscissa of each column of pixels and a foreground pixel in each column of pixels Cumulative value
  • a character recognition sub-module configured to: in a first histogram, a height of a continuous row set consisting of rows in which the accumulated value of the foreground color pixel is greater than the first threshold meets a predetermined height interval, and the foreground pixel in the second histogram When the number of consecutive column sets formed by the columns whose accumulated value of the point is greater than the second threshold meets a predetermined number, the character region meeting the predetermined feature is successfully identified in the region of interest;
  • a fifth identification sub-module configured to: in a first histogram, a height of a continuous row set consisting of rows in which the accumulated value of the foreground color pixel is greater than the first threshold does not meet the predetermined height interval, or in the second histogram When the number of consecutive column sets composed of the columns of the scene pixels whose accumulated value is larger than the second threshold does not satisfy the predetermined number, the character region meeting the predetermined feature is not recognized in the region of interest.
  • the identification module comprises:
  • a filtering sub-module configured to perform Sobel level filtering and binarization on the document image to obtain a processed document image
  • a detecting submodule configured to perform line detection on a predetermined area in the processed document image to obtain at least one straight line
  • the edge recognition sub-module is configured to recognize n lines as predetermined edges of n candidates when the line is n, n ⁇ 2.
  • the area identification module is configured to determine at least one information area according to the target predetermined edge and the second relative positional relationship; the second relative positional relationship is a relative between the target predetermined edge and the information area Positional relationship.
  • an area identifying apparatus comprising:
  • a memory for storing processor executable instructions
  • processor is configured to:
  • the predetermined edge being an edge located in a predetermined direction of the document
  • a predetermined edge of one of the predetermined edges of the n candidates is determined as a target predetermined edge, n ⁇ 2;
  • At least one information area is identified in the document image according to the target predetermined edge.
  • the predetermined edge is an edge located in a predetermined direction of the document; when the predetermined edge of the n candidates is identified, the predetermined edge of one of the predetermined edges of the n candidates is determined as the target reservation Edge, n ⁇ 2; at least one information area is identified in the document image according to the target predetermined edge; and the identification of certain information areas in the document image obtained by direct shooting and the positioning of certain information areas are solved in the related art. Inaccurate problem; the target predetermined edge is determined by the predetermined edges of the n candidates in the document image, and at least one information region is determined based on the target predetermined edge, thereby accurately positioning the information region.
  • FIG. 1 is a flowchart of an area identification method according to an exemplary embodiment
  • FIG. 2 is a flowchart of an area identification method according to another exemplary embodiment
  • FIG. 3A is a flowchart of an area identification method according to another exemplary embodiment
  • FIG. 3B is a schematic diagram of a document image binarization according to an exemplary embodiment
  • FIG. 3C is a schematic diagram of line detection in a document image according to an exemplary embodiment
  • FIG. 3D is a schematic diagram showing n candidate predetermined edges in a document image according to an exemplary embodiment
  • FIG. 4 is a flowchart of an area identification method according to another exemplary embodiment
  • FIG. 5A is a flowchart of an area identification method according to another exemplary embodiment
  • FIG. 5B is a schematic diagram of determining a target information area according to an exemplary embodiment
  • FIG. 6A is a flowchart of an area identification method according to another exemplary embodiment
  • FIG. 6B is a schematic diagram of calculating a first histogram according to a horizontal direction, according to an exemplary embodiment
  • FIG. 6C is a schematic diagram of calculating a second histogram in a vertical direction, according to an exemplary embodiment
  • 6D is a schematic diagram of a continuous row set, according to an exemplary embodiment
  • 6E is a schematic diagram of a continuous column set, according to an exemplary embodiment
  • FIG. 7 is a flowchart of an area identification method according to another exemplary embodiment.
  • FIG. 8 is a block diagram of an area identifying apparatus according to an exemplary embodiment
  • FIG. 9 is a block diagram of an area identifying apparatus according to another exemplary embodiment.
  • FIG. 10 is a block diagram of an area identifying apparatus according to another exemplary embodiment.
  • FIG. 11A is a block diagram of a first sorting sub-module in an area identifying apparatus according to an exemplary embodiment
  • FIG. 11B is a block diagram of a first identification sub-module in an area identification apparatus according to an exemplary embodiment
  • FIG. 12 is a block diagram of a fourth identification sub-module in an area identification apparatus according to an exemplary embodiment
  • FIG. 13 is a block diagram of an area identifying apparatus, according to an exemplary embodiment.
  • FIG. 1 is a flowchart of an area identification method according to an exemplary embodiment. As shown in FIG. 1, the area identification method includes the following steps.
  • a predetermined edge in the document image is identified, the predetermined edge being an edge located in a predetermined direction of the document;
  • the ID image is an image obtained directly from the document, such as an ID card image, a social security card image, and the like.
  • the predetermined edge may be any one of the upper edge, the lower edge, the left edge, and the right edge of the document.
  • the following edges are exemplified in the embodiments of the present disclosure. For the case where the predetermined edge is the upper edge, the left edge, or the right edge, it will not be described again.
  • step 104 when the predetermined edges of the n candidates are identified, the predetermined edge of one of the predetermined edges of the n candidates is determined as the target predetermined edge, n ⁇ 2;
  • the photographing effect of the document image is affected by factors such as shooting angle, background, lighting conditions, and shooting parameters, when a predetermined edge is recognized, it is possible to recognize more than one predetermined edge, that is, a predetermined edge of n candidates.
  • one of the predetermined edges of the n candidates is determined as the target predetermined edge of the document in the document image.
  • the target predetermined edge may be considered to be a true predetermined edge, a correct predetermined edge, or a predetermined edge having a higher accuracy.
  • step 106 at least one information area is identified in the document image based on the target predetermined edge.
  • the position of the predetermined edge of the target in the document image is relatively fixed, and each information area on the document can be determined in the document image according to the target predetermined edge.
  • the information area refers to the area in the document image that carries the text information, such as: the name information area, the date of birth information area, the gender area, the address information area, the citizenship number information area, the number information area, the information area of the issuing authority, and the effective date. At least one of an information area and the like information area.
  • the area recognition method provided in the embodiment of the present disclosure is obtained by identifying the image in the ID image.
  • a predetermined edge the predetermined edge being an edge located in a predetermined direction of the document; when the predetermined edge of the n candidates is identified, one of the predetermined edges of the n candidates is determined as the target predetermined edge, n ⁇ 2;
  • the predetermined edge identifies at least one information area in the document image; and solves the problem that the identification of certain information areas in the document image obtained by direct shooting is difficult and the positioning of some information areas is inaccurate in the related art;
  • the predetermined edge of the n candidates in the document image determines a target predetermined edge, and at least one information region is determined based on the target predetermined edge, thereby accurately positioning the information region.
  • FIG. 2 is a flowchart of an area identification method according to another exemplary embodiment. As shown in FIG. 2, the area identification method includes the following steps.
  • a predetermined edge in the document image is identified, the predetermined edge being an edge located in a predetermined direction of the document;
  • a rectangular area for guiding shooting is set in the shooting interface, and when the user aligns the rectangular area with the document, the document image is captured.
  • the predetermined edge of the document in the document image is identified by a line detection technique according to the image of the obtained document.
  • the predetermined edge of the identification is determined as the target predetermined edge of the document in the document image, and step 212 is directly performed;
  • the line detection technology recognizes that there are n predetermined edges of the documents in the document image, and performs steps 204 to 210 to process the predetermined edges of the n candidates, n ⁇ 2.
  • step 204 the predetermined edges of the n candidates are sorted
  • the likelihood that the predetermined edges of the n candidates are the target predetermined edges is sorted from large to small.
  • step 206 using the predetermined edge of the i-th candidate and the first relative positional relationship, attempting to identify the target information area in the document image, 1 ⁇ i ⁇ n;
  • the difficulty of identifying the target information area is usually low.
  • the predetermined edge of the i-th candidate is the target predetermined edge, and the target information region is attempted to be identified by the predetermined edge of the i-th candidate; if the target information region can be successfully identified, the predetermined edge of the i-th candidate is confirmed as the target Predetermined edge; if not recognized When the target information area is out, it is confirmed that the predetermined edge of the i-th candidate is not the target predetermined edge.
  • the attempt is sequentially performed in accordance with the predetermined edge sorted in step 204.
  • the first relative positional relationship is a relative positional relationship between the target predetermined edge and the target information area.
  • step 208 if the target information area is successfully identified, the predetermined edge of the i-th candidate is determined as the target predetermined edge;
  • the predetermined edge of the i-th candidate is confirmed as the target predetermined edge.
  • step 212 at least one information area is identified in the document image based on the target predetermined edge.
  • At least one information region is determined according to the target predetermined edge and the second relative positional relationship.
  • the information area such as the name information area, the date of birth information area, the gender area, the address information area, the citizenship number information area, the number information area, the issuing document authority information area, the effective date information area, and the like are determined.
  • the second relative positional relationship is a relative positional relationship between the target predetermined edge and the information area.
  • the first relative positional relationship is a subset of the second relative positional relationship.
  • the second relative positional relationship may include: a relative positional relationship between the target predetermined edge and the name information area, a relative positional relationship between the target predetermined edge and the date of birth information region, or between the target predetermined edge and the gender region. Relative positional relationship and so on.
  • the area identifying method by identifying a predetermined edge in the document image, the predetermined edge is an edge located in a predetermined direction of the document; sorting the predetermined edges of the n candidates; using the i a predetermined edge of the candidate and a first relative positional relationship, attempting to identify the target information region in the document image, 1 ⁇ i ⁇ n; the first relative positional relationship is a relative positional relationship between the target predetermined edge and the target information region;
  • the predetermined edge identifies at least one information area in the document image; and solves the knowledge of certain information areas in the document image obtained by direct shooting in the related art The problem that the difficulty is large and the positioning of some information areas is inaccurate; the predetermined edge of the target of the n candidates in the document image is determined, and at least one information area is determined according to the predetermined edge of the target, thereby The effect of accurate regional positioning.
  • step 202 identifies a predetermined edge in the document image. It can alternatively be implemented as the following steps 202a to 202c, as shown in FIG. 3A:
  • step 202a Sobel level filtering and binarization are performed on the document image to obtain a processed document image
  • Binarization refers to comparing the gray value of the pixel in the image of the document with the preset gray threshold, and dividing the pixel in the image into two parts: a pixel group larger than the preset gray threshold and less than the preset gray level.
  • the threshold pixel group displays two different color colors of black and white in the image of the two parts, and obtains the image of the binary image, as shown in FIG. 3B.
  • the pixel of one color in the foreground is called the foreground pixel, that is, the white pixel of FIG. 3B; the pixel of one color of the background is called the background pixel, that is, FIG. 3B Black pixel points.
  • step 202b performing linear detection on a predetermined area in the processed document image to obtain at least one straight line
  • the predetermined area is an area located in a predetermined direction of the document.
  • the predetermined area is an area in which the lower edge of the document is in the document image, or the predetermined area is an area in which the upper edge of the document is in the document image, and the like.
  • the processed document image is subjected to straight line detection, and the line detection includes straight line fitting or Hough transform, thereby obtaining at least one straight line, as shown in FIG. 3C.
  • step 202c when the straight line is n, n ⁇ 2, and n straight lines are identified as predetermined edges of n candidates.
  • the n straight lines are identified as predetermined edges of n candidates, where n ⁇ 2;
  • the line of the lower edge of the document is detected in the image of the document, and after straight line fitting or Hough transform, the predetermined edge of n candidates is obtained, then n
  • the lower edge of the candidate document is in the area of the binarized document image, as shown in Figure 3D.
  • step 212 is directly executed.
  • the region identification method obtains the processed image of the certificate by performing Sobel horizontal filtering and binarization on the image of the document, and performs line detection on the predetermined area in the processed document image. At least one straight line identifies the n straight lines as predetermined edges of the n candidates, so that the detection of the target predetermined edge of the document image is more accurate, and the accuracy of the subsequent information area recognition can be improved.
  • the step 204 sorting the predetermined edges of the n candidates may be implemented as the following steps 204a and 204b, as shown in FIG. 4:
  • step 204a for a predetermined edge of each candidate, the predetermined edge is intersected with the foreground pixel at the same position in the processed document image to obtain a number of intersection points corresponding to the predetermined edge;
  • the processed document image is an image that is subjected to Sobel horizontal filtering and binarization
  • the sobel image is first subjected to sobel horizontal filtering, that is, the sobel operator is used to filter in the horizontal direction. Then, the filtered document image is binarized.
  • the predetermined predetermined edge of each candidate is intersected at the foreground pixel at the same location in the processed document image. That is, the number of foreground pixel points at the same position in the binarized document image is obtained for each candidate's predetermined edge.
  • step 204b the predetermined edges of the n candidates are sorted in order of the number of intersections in descending order.
  • the number of intersection numbers corresponding to the predetermined edge of each candidate is obtained, the number of intersection numbers is sorted in order of the number of intersection points, and the predetermined edges of the sorted n candidates are obtained.
  • the region identification method provided by the embodiment improves the speed of determining a predetermined edge of the target by sorting the predetermined edges of the n candidates, thereby facilitating accurate positioning of the predetermined edge of the target, and improving subsequent information region recognition. Time accuracy.
  • step 206 attempts to identify the target information area in the document image using the predetermined edge of the i-th candidate and the first relative positional relationship.
  • it can be implemented as the following steps 206a and 206b, as shown in FIG. 5A:
  • step 206a using the predetermined edge of the i-th candidate and the first relative positional relationship, the region of interest is truncated in the document image;
  • the approximate positions of the upper edge, the lower edge, the left edge, and the right edge of the region of interest may be determined, and thus, according to the determined upper and lower edges of the region of interest, The left and right edges intercept the region of interest in the document image.
  • the predetermined edge of the i-th candidate in FIG. 3C is determined as the lower edge 51 of the ID card, according to the relationship between the lower edge 51 of the ID card and the citizenship number. Relative positional relationship, the approximate positions of the upper edge 52, the lower edge 53, the left edge 54, and the right edge 55 of the citizenship number are determined as shown in Fig. 5B.
  • the region of interest refers to an area determined according to a predetermined edge of the i-th candidate and a first relative positional relationship.
  • step 206b it is identified whether there is a character area in the region of interest that conforms to the predetermined feature, the predetermined feature being a feature of the character region in the target information region.
  • the region of interest After the region of interest is truncated, it is identified whether there is a character region in the region of interest that conforms to the predetermined feature according to the predetermined feature.
  • the predetermined feature is a feature of the character region in the target information region.
  • the feature may include a continuous 18-character area (or 18-digit area), and the word spacing between adjacent two character areas is small, and each The height of the character area conforms to the predetermined interval.
  • the target information region is successfully identified
  • the target information region is not recognized.
  • Step 206b may alternatively be implemented as step 301 to step 305, as shown in FIG. 6A.
  • step 301 the region of interest is binarized to obtain a binarized region of interest
  • the area of interest is a citizenship number area.
  • the area of interest is pre-processed.
  • the pre-processing may include: operations such as denoising, filtering, and extracting edges; and binarizing the pre-processed region of interest.
  • the first histogram is calculated according to the horizontal direction for the binarized region of interest, and the first histogram includes: a vertical coordinate of each row of pixels and an accumulated value of foreground pixels in each row of pixels. ;
  • the binarized region of interest is calculated in a horizontal direction by a first histogram indicating a vertical coordinate of each row of pixels in a vertical direction and a foreground pixel in each row of pixels in a horizontal direction.
  • the foreground color pixel refers to a white pixel in the binarized image, as shown in FIG. 6B.
  • the second histogram is calculated according to the vertical direction for the binarized region of interest, and the second histogram includes: an abscissa of each column of pixels and an accumulated value of foreground pixels in each column of pixels;
  • step 304 if the height of the continuous row set formed by the row of the foreground pixel points in the first histogram is greater than the first threshold, the height of the continuous row set meets the predetermined height interval, and the sum of the foreground pixel points in the second histogram is accumulated. If the number of consecutive column sets formed by the column whose value is greater than the second threshold meets a predetermined number, the character region meeting the predetermined feature is successfully identified in the region of interest;
  • an accumulated value of the foreground pixel in each row of pixels can be obtained, and the accumulated value of the foreground pixel in each row of pixels is compared with the first threshold to obtain the first histogram.
  • the accumulated value of the scene pixel is greater than the height of the contiguous set of rows of the first threshold.
  • the continuous row set means that the row whose foreground value of the pixel is larger than the first threshold is a continuous m row, and the set of the consecutive m rows of pixels is as shown in FIG. 6D, for the m rows of pixels in the figure. Point, the accumulated value of the foreground color pixel in the left histogram is greater than the first threshold.
  • the m-row pixel corresponds to the citizenship number line "0421199" in the document image, and the height of the m-row pixel is the height of the continuous line set.
  • the continuous column set means that the column whose accumulated value of the foreground color pixel is larger than the second threshold is a continuous p column.
  • the set of consecutive p columns of pixel points is a set of consecutive columns of p, i.e., a continuous white region formed in the second histogram.
  • the accumulated values of the foreground color pixel points located in the lower side histogram are all greater than the second threshold.
  • the p column of pixels corresponds to the character area "3" in the document image.
  • step 305 if the height of the continuous line set formed by the line of the foreground pixel in the first histogram is greater than the predetermined threshold, or the foreground pixel in the second histogram If the number of consecutive column sets consisting of the columns whose accumulated value is greater than the second threshold does not satisfy the predetermined number, the character region meeting the predetermined feature is not recognized in the region of interest.
  • the region identification method performed in this embodiment performs binarization on the region of interest, and calculates the first histogram and the second histogram according to the horizontal direction and the vertical direction respectively.
  • the figure determines whether the character region conforming to the predetermined feature is recognized according to the height of the continuous row set in the first histogram and the number of consecutive column sets in the second histogram, so that the positioning of the character region is more accurate.
  • the characters in the information area can be identified according to the following steps. As shown in Figure 7:
  • step 701 the information area is binarized to obtain a binarized information area
  • the information area is a citizenship number area.
  • the information area is pre-processed.
  • the pre-processing may include: operations such as denoising, filtering, and extracting edges; and binarizing the pre-processed information region.
  • the first histogram is calculated according to the horizontal direction for the binarized information region, and the first histogram includes: a vertical coordinate of each row of pixels and an accumulated value of foreground pixels in each row of pixels. ;
  • step 703 according to the continuous line set formed by the line in which the accumulated value of the foreground color pixel in the first histogram is greater than the first threshold, a line of text is recognized, and a is a positive integer;
  • the accumulated value of the foreground color pixel in each row of pixels can be obtained, and the accumulated value of the foreground color pixel in each row of pixels is compared with the first threshold, and the first histogram is A set of consecutive rows consisting of rows of scene pixels having a cumulative value greater than the first threshold is determined to be the row in which the text region is located.
  • the text area may be two or more lines. At this time, each successive line set is recognized as a line of text area, and a consecutive line set is recognized as a line of text areas.
  • the second histogram is calculated according to the vertical direction, and the second histogram includes: the abscissa of each column of pixels and the accumulated value of the foreground pixels in each column of pixels, a ⁇ i ⁇ 1, i is a positive integer;
  • a second histogram is calculated in a vertical direction, the second histogram indicating the abscissa of each column of pixels in the horizontal direction and the foreground pixel in each column of pixels in the vertical direction.
  • the cumulative number of the number is the cumulative number of the number.
  • step 705 a i character regions are identified according to a continuous column set consisting of columns in which the accumulated value of the foreground pixels in the second histogram is greater than the second threshold.
  • an accumulated value of the foreground color pixel points in each column of pixels can be obtained, and the accumulated value of the foreground color pixel points in each column of pixels is compared with a second threshold value, and the foreground color pixel in the second histogram is obtained.
  • a contiguous set of columns consisting of columns whose accumulated value is greater than the second threshold is determined to be the column in which the character region is located.
  • Each successive column set is identified as a character region, and b consecutive column sets are identified as b character regions.
  • b character regions In Fig. 6E, 18 character regions can be identified.
  • step 701 and step 705 are executed once for each line of text area, and are performed a total of times.
  • the characters contained in the character region can also be identified by character recognition technology.
  • the text can be a single character of a Chinese character, an English letter, a number, or other language.
  • the a-line text area in the information area is determined, and then the a-line is respectively
  • the text area calculates the second histogram in the vertical direction, identifies the character area corresponding to each character, and can improve the accuracy of the character area in the identification information area.
  • FIG. 8 is a block diagram of an area identifying apparatus according to an exemplary embodiment. As shown in FIG. 8, the area identifying apparatus includes, but is not limited to:
  • the identification module 810 is configured to identify a predetermined edge in the document image, the predetermined edge being an edge located in a predetermined direction of the document;
  • the ID image is an image obtained directly from the document, such as an ID card image, a social security card image, and the like.
  • the predetermined edge may be any one of the upper edge, the lower edge, the left edge, and the right edge of the document.
  • a determining module 820 configured to determine a predetermined edge of one of the predetermined edges of the n candidates as the target predetermined edge, n ⁇ 2, when the predetermined edge of the n candidates is identified;
  • the determination module 820 determines one of the predetermined edges of the n candidates as the target predetermined edge of the document in the document image.
  • the target predetermined edge may be considered to be a true predetermined edge, a correct predetermined edge, or a predetermined edge having a higher accuracy.
  • the area identification module 830 is configured to identify at least one information area in the document image according to the target predetermined edge.
  • the information area refers to the area in the document image that carries the text information, such as: the name information area, the date of birth information area, the gender area, the address information area, the citizenship number information area, the number information area, the information area of the issuing authority, and the effective date. At least one of an information area and the like information area.
  • the area identifying apparatus by recognizing a predetermined edge in the document image, the predetermined edge is an edge located in a predetermined direction of the document; when the predetermined edge of the n candidates is identified, n One of the predetermined edges of the candidate candidates is determined as the target predetermined edge, n ⁇ 2; at least one information region is identified in the document image according to the target predetermined edge; and one of the related art images for direct shooting is solved in the related art.
  • the problem that the identification of the information areas is difficult and the positioning of some information areas is inaccurate; the predetermined edge of the target is determined by the predetermined edges of the n candidates in the document image, and at least one information area is determined based on the predetermined edge of the target. , thus the effect of accurately positioning the information area.
  • the determining module 820 may include the following sub-modules, as shown in FIG. 9:
  • a first sorting sub-module 821 configured to sort predetermined edges of the n candidates
  • the first sorting sub-module 821 sorts the likelihood that the predetermined edges of the n candidates are the target predetermined edges from large to small.
  • the first identification sub-module 822 is configured to use the predetermined edge of the i-th candidate and the first relative positional relationship to try to identify the target information area in the document image, 1 ⁇ i ⁇ n;
  • the first identifying sub-module 822 attempts to identify the target information region with the predetermined edge of the i-th candidate; if the target information region can be successfully identified, the reservation of the i-th candidate is confirmed The edge is the target predetermined edge; if the target information area is not recognized, it is confirmed that the predetermined edge of the i-th candidate is not the target predetermined edge.
  • the first identifying sub-module 822 sequentially attempts to use the predetermined edge sorted in the first sorting sub-module 821 in the predetermined edge identifying target information region of the i-th candidate, assuming that the predetermined edge of the i-th candidate is the target predetermined edge. .
  • the first relative positional relationship is a relative positional relationship between the target predetermined edge and the target information area.
  • a second identification sub-module 823 configured to determine a predetermined edge of the i-th candidate as a target predetermined edge when the target information region is successfully identified
  • the second identification sub-module 823 confirms the predetermined edge of the i-th candidate as The target is scheduled to edge.
  • the area identification module 830 is further configured to determine at least one information area according to the target predetermined edge and the second relative position relationship; the second relative positional relationship is a relative positional relationship between the target predetermined edge and the information area.
  • the area identifying apparatus by identifying a predetermined edge in the document image, the predetermined edge is an edge located in a predetermined direction of the document; sorting the predetermined edges of the n candidates; using the i-th a predetermined edge of the candidate and a first relative positional relationship, attempting to identify the target information region in the document image, 1 ⁇ i ⁇ n; the first relative positional relationship is a relative positional relationship between the target predetermined edge and the target information region;
  • the predetermined edge identifies at least one information area in the document image; and solves the problem that the identification of certain information areas in the document image obtained by direct shooting is difficult and the positioning of some information areas is inaccurate in the related art; n in the ID image
  • the predetermined edge of the candidate bar determines the target predetermined edge, and determines at least one information region based on the predetermined edge of the target, thereby accurately positioning the information region.
  • the identification module 810 can include the following sub-modules, as shown in FIG.
  • the filtering sub-module 811 is configured to perform Sobel level filtering and binarization on the ID image to obtain a processed ID image
  • the filtering sub-module 811 performs sobel horizontal filtering on the document image, that is, the sobel operator is used to filter in the horizontal direction. Then, the filtered document image is binarized.
  • Binarization refers to comparing the gray value of the pixel in the image of the document with the preset gray threshold, and dividing the pixel in the image into two parts: a pixel group larger than the preset gray threshold and less than the preset gray level.
  • the pixel group of the threshold value presents two different color colors of black and white in the image of the two parts, and obtains the image of the binary image after binarization.
  • the detecting sub-module 812 is configured to perform line detection on a predetermined area in the processed document image to obtain at least one straight line;
  • the predetermined area is an area located in a predetermined direction of the document.
  • the image of the certificate processed by the filtering sub-module 811 is obtained, and the detecting sub-module 812 performs line detection on the processed image of the certificate, and the line detection includes a straight line fitting or a Hough transform, thereby obtaining at least one straight line.
  • the edge recognition sub-module 813 is configured to recognize n lines as predetermined edges of n candidates when the line is n, n ⁇ 2.
  • the edge recognition sub-module 813 identifies n straight lines as predetermined edges of n candidates, where n ⁇ 2;
  • the straight line is directly recognized as the target predetermined edge in the document image, and the function of the area identifying module 830 is directly executed.
  • the area identification device obtains the processed image of the certificate by performing Sobel level filtering and binarization on the image of the document, and performs line detection on the predetermined area in the processed document image. At least one straight line identifying n straight lines as predetermined edges of n candidates The detection of the target predetermined edge of the document image is made more accurate, and the accuracy of the subsequent information area recognition can be improved.
  • the first sorting sub-module 821 can include the following sub-modules, as shown in FIG. 11A:
  • intersection sub-module 1110 is configured to, for each predetermined edge of the candidate, intersect the foreground image with the foreground pixel at the same position in the processed document image to obtain a number of intersection points corresponding to the predetermined edge;
  • the processed document image is an image processed by the filtering sub-module 811;
  • the sobe image is first subjected to sobel horizontal filtering, that is, the sobel operator is used to perform filtering in the horizontal direction. Then, the filtered document image is binarized.
  • intersection sub-module 1110 intersects the predetermined edge of each candidate with the foreground pixel at the same location in the processed document image. That is, the number of foreground pixel points at the same position in the binarized document image is obtained for each candidate's predetermined edge.
  • the second sorting sub-module 1120 is configured to sort the predetermined edges of the n candidates in order of the number of intersections.
  • the intersection module 1110 After the intersection module 1110 obtains the number of intersection points corresponding to the predetermined edge of each candidate, the second sorting sub-module 1120 sorts according to the order of the number of intersection points, and obtains the scheduled reservation of the n candidates. edge.
  • the area identifying apparatus improves the speed of determining the predetermined edge of the target by sorting the predetermined edges of the n candidates, thereby facilitating accurate positioning of the predetermined edge of the target, and improving subsequent information area recognition. Time accuracy.
  • the first identification sub-module 822 can include the following sub-modules, as shown in FIG. 11B:
  • the intercepting sub-module 1130 is configured to use the predetermined edge of the i-th candidate and the first relative positional relationship to intercept the region of interest in the document image;
  • the approximate positions of the upper edge, the lower edge, the left edge, and the right edge of the region of interest may be determined, and therefore, the intercepting sub-module 1130 determines according to the determination.
  • the upper edge, the lower edge, the left edge, and the right edge of the region of interest intercept the region of interest in the document image.
  • the region of interest refers to an area determined according to a predetermined edge of the i-th candidate and a first relative positional relationship.
  • the fourth identification sub-module 1140 is configured to identify whether there is a character area in the region of interest that conforms to a predetermined feature, the predetermined feature being a feature of the character region in the target information region.
  • the fourth identifying sub-module 1140 identifies, based on the predetermined feature, whether there is a character region in the region of interest that conforms to the predetermined feature.
  • the predetermined feature is a feature of the character region in the target information region.
  • the area identifying apparatus improves the speed of determining the target information area by using the predetermined edge of the i-th candidate and the first relative positional relationship to improve the speed of determining the target information area. Accurate positioning of the target information area can improve the accuracy of subsequent information area recognition.
  • the fourth identification sub-module 1140 may include the following sub-modules, as shown in FIG. 12:
  • the binarization sub-module 1141 is configured to perform binarization on the region of interest to obtain a binarized region of interest
  • the interest area is a citizenship number area.
  • the binarization sub-module 1141 pre-processes the area of interest.
  • the pre-processing may include: operations such as denoising, filtering, and extracting edges; and binarizing the pre-processed region of interest.
  • the first calculation sub-module 1142 is configured to calculate a first histogram in the horizontal direction for the binarized region of interest, the first histogram comprising: a vertical coordinate of each row of pixels and a foreground color in each row of pixels The accumulated value of the pixel;
  • the first calculation sub-module 1142 calculates the first histogram according to the horizontal direction of the region of interest processed by the binarization sub-module 1141, and the first histogram represents the vertical coordinate of each row of pixel points in the vertical direction, in the horizontal direction. Represents the cumulative value of the number of foreground pixels in each row of pixels.
  • the second calculation sub-module 1143 is configured to calculate a second histogram in the vertical direction for the binarized region of interest, the second histogram comprising: an abscissa of each column of pixels and a foreground pixel in each column of pixels The accumulated value of the point;
  • the second calculation sub-module 1143 calculates the second histogram according to the vertical direction of the region of interest processed by the binarization sub-module 1141.
  • the second histogram represents the abscissa of each column of pixels in the horizontal direction, and represents the vertical direction.
  • the character recognition sub-module 1144 is configured to, in the first histogram, the height of the continuous line set formed by the rows of the foreground pixel points whose accumulated value is greater than the first threshold meets the predetermined height interval, and the foreground color in the second histogram When the number of consecutive column sets consisting of the columns whose accumulated value of the pixel is greater than the second threshold meets a predetermined number, the character region meeting the predetermined feature is successfully identified in the region of interest;
  • the accumulated value of the foreground pixel in each row of pixels can be obtained according to the first histogram, and the character recognition sub-module 1144 compares the accumulated value of the foreground pixel in each row of pixels with the first threshold to obtain the first The height of the contiguous set of rows formed by the rows of foreground pixels in the histogram that are greater than the first threshold.
  • the continuous row set refers to a set in which the accumulated value of the foreground color pixel is larger than the first threshold is a continuous m row, and the continuous m row pixel is composed of a set.
  • an accumulated value of the foreground pixel in each column of pixels may be acquired, and the character recognition sub-module 1144 compares the accumulated value of the foreground pixel in each column of pixels with a second threshold to obtain a second histogram.
  • the accumulated value of the foreground pixel is greater than the number of consecutive column sets composed of the second threshold.
  • the continuous column set means that the column whose accumulated value of the foreground color pixel is larger than the second threshold is a continuous p column, and the continuous p column pixel group is composed of a set.
  • the fifth identification sub-module 1145 is configured to: in the first histogram, the height of the continuous row set formed by the row of the foreground pixel points whose accumulated value is greater than the first threshold does not meet the predetermined height interval, or in the second histogram When the number of consecutive column sets composed of the columns of the foreground pixels having the accumulated value greater than the second threshold does not satisfy the predetermined number, the character region meeting the predetermined feature is not recognized in the region of interest.
  • the fifth identifying sub-module 1145 considers that the character region meeting the predetermined feature is not recognized in the region of interest.
  • the area identifying apparatus performs binarization on the region of interest, and calculates the first histogram and the second histogram according to the horizontal direction and the vertical direction respectively.
  • the figure determines whether the character region conforming to the predetermined feature is recognized according to the height of the continuous row set in the first histogram and the number of consecutive column sets in the second histogram, so that the positioning of the character region is more accurate.
  • An exemplary embodiment of the present disclosure provides an area identifying apparatus capable of implementing the area identifying method provided by the present disclosure, the area identifying apparatus comprising: a processor, a memory for storing processor executable instructions;
  • processor is configured to:
  • the predetermined edge being an edge located in a predetermined direction of the document
  • a predetermined edge of one of the predetermined edges of the n candidates is determined as a target predetermined edge, n ⁇ 2;
  • At least one information area is identified in the document image according to the target predetermined edge.
  • FIG. 13 is a block diagram of an apparatus for using a region identification method according to an exemplary embodiment.
  • device 1300 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a gaming console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
  • apparatus 1300 can include one or more of the following components: processing component 1302, memory 1304, power component 1306, multimedia component 1308, audio component 1310, input/output (I/O) interface 1312, sensor component 1314, and Communication component 1316.
  • Processing component 1302 typically controls the overall operation of device 1300, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations.
  • Processing component 1302 can include one or more processors 1318 to execute instructions to perform all or part of the steps described above.
  • processing component 1302 can include one or more modules to facilitate interaction between component 1302 and other components.
  • processing component 1302 can include a multimedia module to facilitate interaction between multimedia component 1308 and processing component 1302.
  • Memory 1304 is configured to store various types of data to support operation at device 1300. Examples of such data include instructions for any application or method operating on device 1300, contact data, phone book data, messages, pictures, videos, and the like.
  • Memory 1304 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable Programmable read only memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), Magnetic Memory, Flash Memory, Disk or Optical Disk.
  • SRAM static random access memory
  • EEPROM electrically erasable programmable read only memory
  • EPROM erasable Programmable read only memory
  • PROM Programmable Read Only Memory
  • ROM Read Only Memory
  • Magnetic Memory Flash Memory
  • Disk Disk or Optical Disk.
  • Power component 1306 provides power to various components of device 1300.
  • Power component 1306 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for device 1300.
  • the multimedia component 1308 includes a screen between the device 1300 and the user that provides an output interface.
  • the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user.
  • the touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can sense not only the boundaries of the touch or sliding action, but also the duration and pressure associated with the touch or slide operation.
  • the multimedia component 1308 includes a front camera and/or a rear camera. When the device 1300 is in an operation mode, such as a shooting mode or a video mode, the front camera and/or the rear camera can receive external multimedia data. Each front and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
  • the audio component 1310 is configured to output and/or input an audio signal.
  • the audio component 1310 includes a microphone (MIC) that is configured to receive an external audio signal when the device 1300 is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode.
  • the received audio signal may be further stored in memory 1304 or transmitted via communication component 1316.
  • the audio component 1310 also includes a speaker for outputting an audio signal.
  • the I/O interface 1312 provides an interface between the processing component 1302 and the peripheral interface module, which may be a keyboard, a click wheel, a button, or the like. These buttons may include, but are not limited to, a home button, a volume button, a start button, and a lock button.
  • Sensor assembly 1314 includes one or more sensors for providing device 1300 with a status assessment of various aspects.
  • the sensor assembly 1314 can detect an open/closed state of the device 1300, the relative positioning of the components, such as a display and a keypad of the device 1300, and the sensor component 1314 can also detect a change in position of a component of the device 1300 or device 1300, the user The presence or absence of contact with device 1300, device 1300 orientation or acceleration/deceleration and temperature variation of device 1300.
  • Sensor assembly 1314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact.
  • Sensor assembly 1314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications.
  • the sensor component 1314 can also be packaged Including acceleration sensor, gyro sensor, magnetic sensor, pressure sensor or temperature sensor.
  • Communication component 1316 is configured to facilitate wired or wireless communication between device 1300 and other devices.
  • the device 1300 can access a wireless network based on a communication standard, such as Wi-Fi, 2G or 3G, or a combination thereof.
  • communication component 1316 receives broadcast signals or broadcast associated information from an external broadcast management system via a broadcast channel.
  • communication component 1316 also includes a near field communication (NFC) module to facilitate short range communication.
  • NFC near field communication
  • the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IRDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
  • RFID radio frequency identification
  • IRDA infrared data association
  • UWB ultra-wideband
  • Bluetooth Bluetooth
  • apparatus 1300 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable A gate array (FPGA), controller, microcontroller, microprocessor or other electronic component implementation for performing the above-described region identification method.
  • ASICs application specific integrated circuits
  • DSPs digital signal processors
  • DSPDs digital signal processing devices
  • PLDs programmable logic devices
  • FPGA field programmable A gate array
  • controller microcontroller, microprocessor or other electronic component implementation for performing the above-described region identification method.
  • non-transitory computer readable storage medium comprising instructions, such as a memory 1304 comprising instructions executable by processor 1318 of apparatus 1300 to perform the area identification method described above.
  • the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Databases & Information Systems (AREA)
  • Evolutionary Computation (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Computer Graphics (AREA)
  • Geometry (AREA)
  • Image Analysis (AREA)
  • Character Input (AREA)
  • Image Processing (AREA)

Abstract

本公开揭示了一种区域识别方法及装置,属于图像处理领域。所述区域识别方法包括:通过识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。

Description

区域识别方法及装置
本申请基于申请号为201510726012.7、申请日为2015年10月30日的中国专利申请提出,并要求该中国专利申请的优先权,该中国专利申请的全部内容在此引入本申请作为参考。
技术领域
本公开涉及图像处理领域,特别涉及一种区域识别方法及装置。
背景技术
身份证的自动识别技术是一种通过图像处理对身份证上的文字信息进行识别的技术。
相关技术提供了一种身份证的自动识别方法,通过身份证扫描设备按照固定的相对位置对身份证进行扫描,得到身份证的扫描图像;对扫描图像中的n个预定区域进行文字识别,得到姓名信息、性别信息、民族信息、出生日期信息、地址信息和公民身份号码信息中的至少一种。但是对于直接拍摄得到的身份证图像,仍然有较大的识别难度。
发明内容
为了解决相关技术中的问题,本公开提供一种区域识别方法及装置。所述技术方案如下:
根据本公开实施例的第一方面,提供一种区域识别方法,该方法包括:
识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
根据目标预定边缘在证件图像中识别出至少一个信息区域。
在一个可选的实施例中,在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,包括:
对n条候选的预定边缘进行排序;
使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系;
若成功识别出目标信息区域,则将第i条候选的预定边缘确定为目标预定边缘;
若未能识别出目标信息区域,则令i=i+1,重新执行使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域的步骤。
在一个可选的实施例中,对n条候选的预定边缘进行排序,包括:
对于每一条候选的预定边缘,将预定边缘与处理后的证件图像中相同位置处的前景色像素点求交,得到预定边缘对应的交点数;处理后的证件图像是经过索贝尔水平滤波和二值化的图像;
按照交点数由多到少的顺序,将n条候选的预定边缘进行排序。
在一个可选的实施例中,使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,包括:
使用第i条候选的预定边缘和第一相对位置关系,在证件图像中截取出兴趣区域;
识别兴趣区域中是否存在符合预定特征的字符区域,预定特征是目标信息区域中的字符区域所具有的特征。
在一个可选的实施例中,识别兴趣区域中是否存在符合预定特征的字符区域,包括:
对兴趣区域进行二值化,得到二值化后的兴趣区域;
对二值化后的兴趣区域按照水平方向计算第一直方图,第一直方图包括:每行像素点的竖坐标和每行像素点中前景色像素点的累加值;
对二值化后的兴趣区域按照竖直方向计算第二直方图,第二直方图包括:每列像素点的横坐标和每列像素点中前景色像素点的累加值;
若第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数,则在兴趣区域中成功识别出符合预定特征的字符区域;
若第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行 集合的高度不符合预定高度区间,或者第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数,则在兴趣区域中未能识别出符合预定特征的字符区域。
在一个可选的实施例中,识别证件图像中的预定边缘,包括:
对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像;
对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线;
在直线为n条时,n≥2,将n条直线识别为n条候选的预定边缘。
在一个可选的实施例中,根据目标预定边缘在证件图像中确定出至少一个信息区域,包括:
根据目标预定边缘和第二相对位置关系,确定出至少一个信息区域;第二相对位置关系是目标预定边缘与信息区域之间的相对位置关系。
根据本公开实施例的第二方面,提供一种区域识别装置,该装置包括:
识别模块,被配置为识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
确定模块,被配置为在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
区域识别模块,被配置为根据目标预定边缘在证件图像中识别出至少一个信息区域。
在一个可选的实施例中,确定模块,包括:
第一排序子模块,被配置为对n条候选的预定边缘进行排序;
第一识别子模块,被配置为使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系。
第二识别子模块,被配置为在成功识别出目标信息区域时,将第i条候选的预定边缘确定为目标预定边缘;
第三识别子模块,被配置为在未能识别出目标信息区域时,令i=i+1,重新返回第一识别子模块。
在一个可选的实施例中,第一排序子模块,包括:
求交子模块,被配置为对于每一条候选的预定边缘,将预定边缘与处理后的证件图像中相同位置处的前景色像素点求交,得到预定边缘对应的交点数; 处理后的证件图像是经过索贝尔水平滤波和二值化的图像;
第二排序子模块,被配置为按照交点数由多到少的顺序,将n条候选的预定边缘进行排序。
在一个可选的实施例中,第一识别子模块,包括:
截取子模块,被配置为使用第i条候选的预定边缘和第一相对位置关系,在证件图像中截取出兴趣区域;
第四识别子模块,被配置为识别兴趣区域中是否存在符合预定特征的字符区域,预定特征是目标信息区域中的字符区域所具有的特征。
在一个可选的实施例中,第四识别子模块,包括:
二值化子模块,被配置为对兴趣区域进行二值化,得到二值化后的兴趣区域;
第一计算子模块,被配置为对二值化后的兴趣区域按照水平方向计算第一直方图,第一直方图包括:每行像素点的竖坐标和每行像素点中前景色像素点的累加值;
第二计算子模块,被配置为对二值化后的兴趣区域按照竖直方向计算第二直方图,第二直方图包括:每列像素点的横坐标和每列像素点中前景色像素点的累加值;
字符识别子模块,被配置为在第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数时,在兴趣区域中成功识别出符合预定特征的字符区域;
第五识别子模块,被配置为在第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度不符合预定高度区间,或者第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数时,在兴趣区域中未能识别出符合预定特征的字符区域。
在一个可选的实施例中,识别模块,包括:
滤波子模块,被配置为对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像;
检测子模块,被配置为对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线;
边缘识别子模块,被配置为在直线为n条时,n≥2,将n条直线识别为n条候选的预定边缘。
在一个可选的实施例中,区域识别模块,被配置为根据目标预定边缘和第二相对位置关系,确定出至少一个信息区域;第二相对位置关系是目标预定边缘与信息区域之间的相对位置关系。
根据本公开实施例的第三方面,提供一种区域识别装置,该装置包括:
处理器;
用于存储处理器可执行指令的存储器;
其中,处理器被配置为:
识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
根据目标预定边缘在证件图像中识别出至少一个信息区域。
本公开的实施例提供的技术方案可以包括以下有益效果:
通过识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性的,并不能限制本公开。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并于说明书一起用于解释本公开的原理。
图1是根据一示例性实施例示出的一种区域识别方法的流程图;
图2是根据另一示例性实施例示出的一种区域识别方法的流程图;
图3A是根据另一示例性实施例示出的一种区域识别方法的流程图;
图3B是根据一示例性实施例示出的一种证件图像二值化的示意图;
图3C是根据一示例性实施例示出的一种证件图像中直线检测的示意图;
图3D是根据一示例性实施例示出的一种证件图像中n条候选预定边缘的示意图;
图4是根据另一示例性实施例示出的一种区域识别方法的流程图;
图5A是根据另一示例性实施例示出的一种区域识别方法的流程图;
图5B是根据一示例性实施例示出的一种确定目标信息区域的示意图;
图6A是根据另一示例性实施例示出的一种区域识别方法的流程图;
图6B是根据一示例性实施例示出的一种按照水平方向计算第一直方图的示意图;
图6C是根据一示例性实施例示出的一种按照竖直方向计算第二直方图的示意图;
图6D是根据一示例性实施例示出的一种连续行集合的示意图;
图6E是根据一示例性实施例示出的一种连续列集合的示意图;
图7是根据另一示例性实施例示出的一种区域识别方法的流程图;
图8是根据一示例性实施例示出的一种区域识别装置的框图;
图9是根据另一示例性实施例示出的一种区域识别装置的框图;
图10是根据另一示例性实施例示出的一种区域识别装置的框图;
图11A是根据一示例性实施例示出的一种区域识别装置中第一排序子模块的框图;
图11B是根据一示例性实施例示出的一种区域识别装置中第一识别子模块的框图;
图12是根据一示例性实施例示出的一种区域识别装置中第四识别子模块的框图;
图13是根据一示例性实施例示出的一种区域识别装置的框图。
具体实施方式
这里将详细地对示例性实施例进行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。 以下示例性实施例中所描述的实施方式并不代表与本公开相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本公开的一些方面相一致的装置和方法的例子。
图1是根据一示例性实施例示出的一种区域识别方法的流程图。如图1所示,该区域识别方法包括如下步骤。
在步骤102中,识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
证件图像是对证件直接拍摄得到的图像,比如:身份证图像、社会保障卡图像等。
预定边缘可以是证件的上边缘、下边缘、左边缘和右边缘中的任意一个。本公开实施例中均以下边缘来举例说明。对于预定边缘是上边缘、左边缘或右边缘的情况,不再赘述。
在步骤104中,在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
由于证件图像的拍摄效果受拍摄角度、背景、光照条件和拍摄参数等因素的多重影响,在识别预定边缘时,有可能会识别出不止一条预定边缘,也即n条候选的预定边缘。
当n≥2时,将n条候选的预定边缘中的某一条预定边缘确定为证件在证件图像中的目标预定边缘。
目标预定边缘可以认为是真实的预定边缘、正确的预定边缘或者准确度较高的预定边缘。
在步骤106中,根据目标预定边缘在证件图像中识别出至少一个信息区域。
证件图像中的目标预定边缘的位置是相对固定的,根据目标预定边缘能够在证件图像中确定出证件上的各个信息区域。
信息区域是指证件图像中携带有文字信息的区域,比如:姓名信息区域、出生日期信息区域、性别区域、地址信息区域、公民身份号码信息区域、编号信息区域、颁发证件机关信息区域、有效日期信息区域等等信息区域中的至少一种。
综上所述,本公开实施例中提供的区域识别方法,通过识别证件图像中的 预定边缘,预定边缘是位于证件的预定方向上的边缘;在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条预定边缘确定为目标预定边缘,n≥2;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。
图2是根据另一示例性实施例示出的一种区域识别方法的流程图。如图2所示,该区域识别方法包括以下几个步骤。
在步骤202中,识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
可选地,在拍摄证件图像时,拍摄界面中设置有用于引导拍摄的矩形区域,用户将矩形区域对准证件时,拍摄得到证件图像。
可选的,根据拍摄得到的证件图像,通过直线检测技术识别出证件图像中证件的预定边缘。
当通过直线检测技术识别出证件图像中证件的预定边缘仅有一条时,将识别出的该条预定边缘确定为证件在证件图像中的目标预定边缘,直接执行步骤212;
由于拍摄角度、背景、光照条件和拍摄参数等因素,使得直线检测技术识别出证件图像中证件的预定边缘有n条时,执行步骤204至步骤210对n条候选的预定边缘进行处理,n≥2。
在步骤204中,对n条候选的预定边缘进行排序;
在获取到n条候选的预定边缘后,根据n条候选的预定边缘是目标预定边缘的可能性由大到小进行排序。
在步骤206中,使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;
目标信息区域的识别难度通常较低。本步骤中假设第i条候选的预定边缘是目标预定边缘,尝试用第i条候选的预定边缘识别目标信息区域;若能够成功识别出目标信息区域,则确认第i条候选的预定边缘是目标预定边缘;若未能识别 出目标信息区域,则确认第i条候选的预定边缘不是目标预定边缘。
在假定第i条候选的预定边缘是目标预定边缘,尝试用第i条候选的预定边缘识别目标信息区域的步骤中,按照步骤204中排序的预定边缘依次进行尝试。
第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系。
在步骤208中,若成功识别出目标信息区域,则将第i条候选的预定边缘确定为目标预定边缘;
若根据第i条候选的预定边缘和第一相对位置关系成功识别出证件图像中的目标信息区域,则将第i条候选的预定边缘确认为目标预定边缘。
在步骤210中,若未能识别出目标信息区域,则令i=i+1,重新执行步骤206;
若根据第i条候选的预定边缘和第一相对位置关系未能识别出证件图像中的目标信息区域,则令i=i+1,将第i+1条候选的预定边缘确定为目标预定边缘,使用第i+1条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域。
在步骤212中,根据目标预定边缘在证件图像中识别出至少一个信息区域。
根据步骤208中确定的目标预定边缘,根据目标预定边缘和第二相对位置关系,确定出至少一个信息区域。
比如:确定出姓名信息区域、出生日期信息区域、性别区域、地址信息区域、公民身份号码信息区域、编号信息区域、颁发证件机关信息区域、有效日期信息区域等等信息区域。
其中,第二相对位置关系是目标预定边缘与信息区域之间的相对位置关系。第一相对位置关系是第二相对位置关系的子集。
比如:第二相对位置关系可以包括:目标预定边缘与姓名信息区域之间的相对位置关系、目标预定边缘与出生日期信息区域之间的相对位置关系,或,目标预定边缘与性别区域之间的相对位置关系等等。
综上所述,本公开实施例中提供的区域识别方法,通过识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;对n条候选的预定边缘进行排序;使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识 别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。
同时,通过对n条候选的预定边缘进行排序,达到了提高确定目标预定边缘的速度和对目标预定边缘的准确定位的效果。
在基于图2实施例提供的可选实施例中,步骤202识别证件图像中的预定边缘。可替代实现为如下步骤202a至步骤202c,如图3A所示:
在步骤202a中,对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像;
首先对证件图像进行sobel水平滤波,也即采用sobel算子沿水平方向进行滤波。然后,对滤波后的证件图像进行二值化。二值化是指将证件图像中的像素点的灰度值与预设灰度阈值比较,将证件图像中的像素点分成两部分:大于预设灰度阈值的像素群和小于预设灰度阈值的像素群,将两部分像素群在证件图像中分别呈现出黑和白两种不同的颜色,得到二值化后的证件图像,如图3B所示。其中,位于前景的一种颜色的像素点称之为前景色像素点,也即图3B的白色像素点;位于背景的一种颜色的像素点称之为背景色像素点,也即图3B中的黑色像素点。
在步骤202b中,对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线;
该预定区域是位于证件的预定方向的区域。比如:该预定区域是证件的下边缘在证件图像中的区域,或者,该预定区域是证件的上边缘在证件图像中的区域等等。
在获取到索贝尔水平滤波和二值化的证件图像后,对处理后的证件图像进行直线检测,该直线检测包括直线拟合或Hough变换,从而得到至少一条直线,如图3C所示。
在步骤202c中,在直线为n条时,n≥2,将n条直线识别为n条候选的预定边缘。
当获取到的直线为n条时,将n条直线识别为n条候选的预定边缘,其中,n≥2;
比如:针对索贝尔水平滤波和二值化的证件图像,对证件的下边缘在证件图像中的区域进行直线检测,通过直线拟合或Hough变换后,得到n条候选的预定边缘,则n条候选的证件的下边缘在二值化证件图像中的区域,如图3D所示。
当获取到的直线只有一条时,将该条直线直接识别为证件图像中的目标预定边缘,直接执行步骤212。
综上所述,本实施例提供的区域识别方法,通过对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像,对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线,将n条直线识别为n条候选的预定边缘,使得在对证件图像的目标预定边缘的检测更加准确,能够提高后续信息区域识别时的准确度。
在基于图2实施例提供的可选实施例中,步骤204对n条候选的预定边缘进行排序可替代实现为如下步骤204a和步骤204b,如图4所示:
在步骤204a中,对于每一条候选的预定边缘,将预定边缘与处理后的证件图像中相同位置处的前景色像素点求交,得到预定边缘对应的交点数;
其中,处理后的证件图像是经过索贝尔水平滤波和二值化的图像;
在获取到n条候选的预定边缘后,首先对证件图像进行sobel水平滤波,也即采用sobel算子沿水平方向进行滤波。然后,对滤波后的证件图像进行二值化。
将每一条候选的预定边缘于处理后的证件图像中相同位置处的前景色像素点求交。也即求取每一条候选的预定边缘在二值化后的证件图像中相同位置处属于前景色像素点的个数。
在步骤204b中,按照交点数由多到少的顺序,将n条候选的预定边缘进行排序。
在获取到每一条候选的预定边缘对应的交点数后,按照交点数的个数有多到少的顺序进行排序,获取排序后的n条候选的预定边缘。
综上所述,本实施例提供的区域识别方法,通过对n条候选的预定边缘进行排序,提高了确定目标预定边缘的速度,有利于对目标预定边缘的准确定位,能够提高后续信息区域识别时的准确性。
在基于图2实施例提供的可选实施例中,步骤206使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域。可替代实现为如下步骤206a和步骤206b,如图5A所示:
在步骤206a中,使用第i条候选的预定边缘和第一相对位置关系,在证件图像中截取出兴趣区域;
根据第i条候选的预定边缘和第一相对位置关系,可以确定出兴趣区域的上边缘、下边缘、左边缘和右边缘的大概位置,因此,根据确定的兴趣区域的上边缘、下边缘、左边缘和右边缘,在证件图像中截取出兴趣区域。
比如:以目标预定边缘为身份证的下边缘为例,将图3C中的第i条候选的预定边缘确定为身份证的下边缘51,根据身份证的下边缘51与公民身份号码之间的相对位置关系,确定出公民身份号码上边缘52、下边缘53、左边缘54和右边缘55的大概位置,如图5B所示。
其中,兴趣区域是指根据第i条候选的预定边缘和第一相对位置关系确定出的区域。
在步骤206b中,识别兴趣区域中是否存在符合预定特征的字符区域,预定特征是目标信息区域中的字符区域所具有的特征。
在截取出兴趣区域后,根据预定特征识别兴趣区域中是否存在符合预定特征的字符区域。
其中,预定特征是目标信息区域中的字符区域所具有的特征。比如,目标信息区域是公民身份号码区域,则该特征可以是包括有连续的18个字符区域(或称18个数字区域),相邻的两个字符区域之间的字间距较小,且每个字符区域的高度符合预定区间。
当识别出兴趣区域中存在符合预定特征的字符区域时,即成功识别出目标信息区域;
当识别出兴趣区域中不存在符合预定特征的字符区域时,即未能识别出目标信息区域。
可选地,识别兴趣区域中是否存在符合预定特征的字符区域。步骤206b可替代实现成为步骤301至步骤305,如图6A所示。
在步骤301中,对兴趣区域进行二值化,得到二值化后的兴趣区域;
以兴趣区域是公民身份号码区域为例,可选地,先对兴趣区域进行预处理。其中,预处理可以包括:去噪、滤波、提取边缘等操作;将预处理后的兴趣区域进行二值化。
在步骤302中,对二值化后的兴趣区域按照水平方向计算第一直方图,第一直方图包括:每行像素点的竖坐标和每行像素点中前景色像素点的累加值;
将二值化后的兴趣区域按照水平方向计算第一直方图,该第一直方图在竖直方向表示每行像素点的竖坐标,在水平方向表示每行像素点中前景色像素点的个数累加值。前景色像素点是指二值化后的图像中的白色像素点,如图6B所示。
在步骤303中,对二值化后的兴趣区域按照竖直方向计算第二直方图,第二直方图包括:每列像素点的横坐标和每列像素点中前景色像素点的累加值;
将二值化后的兴趣区域按照竖直方向计算第二直方图,该第二直方图在水平方向表示每列像素点的横坐标,在竖直方向表示每列像素点中前景色像素点的个数累加值,如图6C所示。
在步骤304中,若第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数,则在兴趣区域中成功识别出符合预定特征的字符区域;
根据第一直方图可以获取到每一行像素点中前景色像素点的累加值,将每一行像素点中前景色像素点的累加值与第一阈值进行比较,获取第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度。
连续行集合是指:前景色像素点的累加值大于第一阈值的行是连续的m行,该连续的m行像素点所组成的集合,如图6D所示,对于图中的m行像素点,在位于左侧直方图中的前景色像素点的累加值均大于第一阈值。而该m行像素点在证件图像中对应公民身份号码行“0421199”,则m行像素点的高度即为连续行集合的高度。
根据第二直方图可以获取到每一列像素点中前景色像素点的累加值,将每一列像素点中前景色像素点的累加值与第二阈值进行比较,获取第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数。
连续列集合是指:前景色像素点的累加值大于第二阈值的列是连续的p列, 该连续的p列像素点所组成的集合,如图6E所示,连续列集合为p,也即第二直方图中形成的连续白色区域。对于图中的p列像素点,在位于下侧直方图中的前景色像素点的累加值均大于第二阈值。而该p列像素点在证件图像中对应字符区域“3”。
当连续行集合的高度符合预定高度区间,且,连续列集合的个数符合预定个数时,则认为在兴趣区域中成功识别出符合预定特征的字符区域。
在步骤305中,若第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度不符合预定高度区间,或者第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数,则在兴趣区域中未能识别出符合预定特征的字符区域。
当连续行集合的高度不符合预定高度区间,或者,连续列集合的个数不符合预定个数时,则认为在兴趣区域中未能识别出符合预定特征的字符区域。
综上所述,本实施例提供的区域识别方法,通过对兴趣区域进行二值化,并将二值化后的兴趣区域分别按照水平方向和竖直方向计算第一直方图和第二直方图,根据第一直方图中连续行集合的高度和第二直方图中连续列集合的个数确定是否识别出符合预定特征的字符区域,使得对字符区域的定位更加准确。
在基于图2实施例提供的可选实施例中,在根据目标预定边缘在证件图像中识别出至少一个信息区域后,可以根据如下步骤对信息区域中的字符进行识别。如图7所示:
在步骤701中,对信息区域进行二值化,得到二值化后的信息区域;
以信息区域是公民身份号码区域为例,可选地,先对信息区域进行预处理。其中,预处理可以包括:去噪、滤波、提取边缘等操作;将预处理后的信息区域进行二值化。
在步骤702中,对二值化后的信息区域按照水平方向计算第一直方图,第一直方图包括:每行像素点的竖坐标和每行像素点中前景色像素点的累加值;
在步骤703中,根据第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合,识别得到a行文字区域,a为正整数;
根据第一直方图可以获取到每一行像素点中前景色像素点的累加值,将每一行像素点中前景色像素点的累加值与第一阈值进行比较,将第一直方图中前 景色像素点的累加值大于第一阈值的行所组成的连续行集合,确定为文字区域所在的行。
当然,若该信息区域是地址信息区域或者其他信息区域,文字区域可能为两行或者两行以上。此时,每个连续行集合识别为一行文字区域,a个连续行集合识别为a行文字区域。
在步骤704中,对于第i行文字区域,按照竖直方向计算第二直方图,第二直方图包括:每列像素点的横坐标和每列像素点中前景色像素点的累加值,a≥i≥1,i为正整数;
对于识别出的公民身份号码行,按照竖直方向计算第二直方图,该第二直方图在水平方向表示每列像素点的横坐标,在竖直方向表示每列像素点中前景色像素点的个数累加值。
在步骤705中,根据第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合,识别得到ai个字符区域。
根据第二直方图可以获取到每一列像素点中前景色像素点的累加值,将每一列像素点中前景色像素点的累加值与第二阈值进行比较,将第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合,确定为字符区域所在的列。
每个连续列集合识别为一个字符区域,b个连续列集合识别为b个字符区域。在图6E中,能够识别出18个字符区域。
若文字区域有a行,则步骤701和步骤705会针对每一行文字区域执行一次,共执行a次。
对于识别出的每个字符区域,还可以通过字符识别技术,识别出该字符区域中包含的文字。文字可以是汉字、英文字母、数字或其它语种的单个字符。
综上所述,本实施例通过对信息区域二值化,并将二值化后的信息区域按照水平方向计算第一直方图,确定信息区域中a行文字区域,再通过分别对a行文字区域按照竖直方向计算第二直方图,识别出每个文字对应的字符区域,能够提高识别信息区域中字符区域的准确度。
下述为本公开装置实施例,可以用于执行本公开方法实施例。对于本公开装置实施例中未披露的细节,请参照本公开方法实施例。
图8是根据一示例性实施例示出的一种区域识别装置的框图,如图8所示,该区域识别装置包括但不限于:
识别模块810,被配置为识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
证件图像是对证件直接拍摄得到的图像,比如:身份证图像、社会保障卡图像等。
预定边缘可以是证件的上边缘、下边缘、左边缘和右边缘中的任意一个。
确定模块820,被配置为在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
当n≥2时,确定模块820将n条候选的预定边缘中的某一条预定边缘确定为证件在证件图像中的目标预定边缘。
目标预定边缘可以认为是真实的预定边缘、正确的预定边缘或者准确度较高的预定边缘。
区域识别模块830,被配置为根据目标预定边缘在证件图像中识别出至少一个信息区域。
信息区域是指证件图像中携带有文字信息的区域,比如:姓名信息区域、出生日期信息区域、性别区域、地址信息区域、公民身份号码信息区域、编号信息区域、颁发证件机关信息区域、有效日期信息区域等等信息区域中的至少一种。
综上所述,本公开实施例中提供的区域识别装置,通过识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条预定边缘确定为目标预定边缘,n≥2;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。
在基于图8实施例提供的可选实施例中,确定模块820可以包括如下子模块,如图9所示:
第一排序子模块821,被配置为对n条候选的预定边缘进行排序;
在获取到n条候选的预定边缘后,第一排序子模块821根据n条候选的预定边缘是目标预定边缘的可能性由大到小进行排序。
第一识别子模块822,被配置为使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;
假设第i条候选的预定边缘是目标预定边缘,第一识别子模块822尝试用第i条候选的预定边缘识别目标信息区域;若能够成功识别出目标信息区域,则确认第i条候选的预定边缘是目标预定边缘;若未能识别出目标信息区域,则确认第i条候选的预定边缘不是目标预定边缘。
第一识别子模块822在假定第i条候选的预定边缘是目标预定边缘,尝试用第i条候选的预定边缘识别目标信息区域中,按照第一排序子模块821中排序的预定边缘依次进行尝试。
第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系。
第二识别子模块823,被配置为在成功识别出目标信息区域时,将第i条候选的预定边缘确定为目标预定边缘;
若根据第i条候选的预定边缘和第一相对位置关系,第一识别子模块822成功识别出证件图像中的目标信息区域,则第二识别子模块823将第i条候选的预定边缘确认为目标预定边缘。
第三识别子模块824,被配置为在未能识别出目标信息区域时,令i=i+1,重新返回第一识别子模块822,执行第一识别子模块822的功能。
可选的,区域识别模块830,还被配置为根据目标预定边缘和第二相对位置关系,确定出至少一个信息区域;第二相对位置关系是目标预定边缘与信息区域之间的相对位置关系。
综上所述,本公开实施例中提供的区域识别装置,通过识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;对n条候选的预定边缘进行排序;使用第i条候选的预定边缘和第一相对位置关系,在证件图像中尝试识别目标信息区域,1≤i≤n;第一相对位置关系是目标预定边缘与目标信息区域之间的相对位置关系;根据目标预定边缘在证件图像中识别出至少一个信息区域;解决了相关技术中对于直接拍摄得到的证件图像中的某些信息区域的识别难度大和对某些信息区域的定位不准确的问题;达到了通过证件图像中的n 条候选的预定边缘确定出目标预定边缘,根据目标预定边缘为基准确定出至少一个信息区域,从而对信息区域准确定位的效果。
同时,通过对n条候选的预定边缘进行排序,达到了提高确定目标预定边缘的速度和对目标预定边缘的准确定位的效果。
在基于图9实施例提供的可选实施例中,识别模块810可以包括如下子模块,如图10所示:
滤波子模块811,被配置为对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像;
首先滤波子模块811对证件图像进行sobel水平滤波,也即采用sobel算子沿水平方向进行滤波。然后,对滤波后的证件图像进行二值化。
二值化是指将证件图像中的像素点的灰度值与预设灰度阈值比较,将证件图像中的像素点分成两部分:大于预设灰度阈值的像素群和小于预设灰度阈值的像素群,将两部分像素群在证件图像中分别呈现出黑和白两种不同的颜色,得到二值化后的证件图像。
检测子模块812,被配置为对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线;
该预定区域是位于证件的预定方向的区域。
获取滤波子模块811处理后的证件图像,检测子模块812对处理后的证件图像进行直线检测,该直线检测包括直线拟合或Hough变换,从而得到至少一条直线。
边缘识别子模块813,被配置为在直线为n条时,n≥2,将n条直线识别为n条候选的预定边缘。
当获取到的直线为n条时,边缘识别子模块813将n条直线识别为n条候选的预定边缘,其中,n≥2;
当获取到的直线只有一条时,将该条直线直接识别为证件图像中的目标预定边缘,直接执行区域识别模块830的功能。
综上所述,本实施例提供的区域识别装置,通过对证件图像进行索贝尔水平滤波和二值化,得到处理后的证件图像,对处理后的证件图像中的预定区域进行直线检测,得到至少一条直线,将n条直线识别为n条候选的预定边缘, 使得在对证件图像的目标预定边缘的检测更加准确,能够提高后续信息区域识别时的准确度。
在基于图9实施例提供的可选实施例中,第一排序子模块821可以包括如下子模块,如图11A所示:
求交子模块1110,被配置为对于每一条候选的预定边缘,将预定边缘与处理后的证件图像中相同位置处的前景色像素点求交,得到预定边缘对应的交点数;
其中,处理后的证件图像是经过滤波子模块811处理后的图像;
在边缘识别子模块813获取到n条候选的预定边缘后,首先对证件图像进行sobel水平滤波,也即采用sobel算子沿水平方向进行滤波。然后,对滤波后的证件图像进行二值化。
求交子模块1110将每一条候选的预定边缘于处理后的证件图像中相同位置处的前景色像素点求交。也即求取每一条候选的预定边缘在二值化后的证件图像中相同位置处属于前景色像素点的个数。
第二排序子模块1120,被配置为按照交点数由多到少的顺序,将n条候选的预定边缘进行排序。
在求交子模块1110获取到每一条候选的预定边缘对应的交点数后,第二排序子模块1120按照交点数的个数有多到少的顺序进行排序,获取排序后的n条候选的预定边缘。
综上所述,本实施例提供的区域识别装置,通过对n条候选的预定边缘进行排序,提高了确定目标预定边缘的速度,有利于对目标预定边缘的准确定位,能够提高后续信息区域识别时的准确性。
在基于图9实施例提供的可选实施例中,第一识别子模块822可以包括如下子模块,如图11B所示:
截取子模块1130,被配置为使用第i条候选的预定边缘和第一相对位置关系,在证件图像中截取出兴趣区域;
根据第i条候选的预定边缘和第一相对位置关系,可以确定出兴趣区域的上边缘、下边缘、左边缘和右边缘的大概位置,因此,截取子模块1130根据确定 的兴趣区域的上边缘、下边缘、左边缘和右边缘,在证件图像中截取出兴趣区域。
其中,兴趣区域是指根据第i条候选的预定边缘和第一相对位置关系确定出的区域。
第四识别子模块1140,被配置为识别兴趣区域中是否存在符合预定特征的字符区域,预定特征是目标信息区域中的字符区域所具有的特征。
在截取子模块1130截取出兴趣区域后,第四识别子模块1140根据预定特征识别兴趣区域中是否存在符合预定特征的字符区域。
其中,预定特征是目标信息区域中的字符区域所具有的特征。
综上所述,本实施例提供的区域识别装置,通过使用第i条候选的预定边缘和第一相对位置关系,在证件图像中截取出兴趣区域,提高了确定目标信息区域的速度,有利于对目标信息区域的准确定位,能够提高后续信息区域识别时的准确性。
在基于图11B实施例提供的可选实施例中,第四识别子模块1140可以包括如下子模块,如图12所示:
二值化子模块1141,被配置为对兴趣区域进行二值化,得到二值化后的兴趣区域;
以兴趣区域是公民身份号码区域为例,可选地,二值化子模块1141先对兴趣区域进行预处理。其中,预处理可以包括:去噪、滤波、提取边缘等操作;将预处理后的兴趣区域进行二值化。
第一计算子模块1142,被配置为对二值化后的兴趣区域按照水平方向计算第一直方图,第一直方图包括:每行像素点的竖坐标和每行像素点中前景色像素点的累加值;
第一计算子模块1142将二值化子模块1141处理后的兴趣区域按照水平方向计算第一直方图,该第一直方图在竖直方向表示每行像素点的竖坐标,在水平方向表示每行像素点中前景色像素点的个数累加值。
第二计算子模块1143,被配置为对二值化后的兴趣区域按照竖直方向计算第二直方图,第二直方图包括:每列像素点的横坐标和每列像素点中前景色像素点的累加值;
第二计算子模块1143将二值化子模块1141处理后的兴趣区域按照竖直方向计算第二直方图,该第二直方图在水平方向表示每列像素点的横坐标,在竖直方向表示每列像素点中前景色像素点的个数累加值。
字符识别子模块1144,被配置为在第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数时,在兴趣区域中成功识别出符合预定特征的字符区域;
根据第一直方图可以获取到每一行像素点中前景色像素点的累加值,字符识别子模块1144将每一行像素点中前景色像素点的累加值与第一阈值进行比较,获取第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度。
连续行集合是指:前景色像素点的累加值大于第一阈值的行是连续的m行,该连续的m行像素点所组成的集合。
根据第二直方图可以获取到每一列像素点中前景色像素点的累加值,字符识别子模块1144将每一列像素点中前景色像素点的累加值与第二阈值进行比较,获取第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数。
连续列集合是指:前景色像素点的累加值大于第二阈值的列是连续的p列,该连续的p列像素点所组成的集合。
第五识别子模块1145,被配置为在第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度不符合预定高度区间,或者第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数时,在兴趣区域中未能识别出符合预定特征的字符区域。
当连续行集合的高度不符合预定高度区间,或者,连续列集合的个数不符合预定个数时,则第五识别子模块1145认为在兴趣区域中未能识别出符合预定特征的字符区域。
综上所述,本实施例提供的区域识别装置,通过对兴趣区域进行二值化,并将二值化后的兴趣区域分别按照水平方向和竖直方向计算第一直方图和第二直方图,根据第一直方图中连续行集合的高度和第二直方图中连续列集合的个数确定是否识别出符合预定特征的字符区域,使得对字符区域的定位更加准确。
本公开一示例性实施例提供了一种区域识别装置,能够实现本公开提供的区域识别方法,该区域识别装置包括:处理器、用于存储处理器可执行指令的存储器;
其中,处理器被配置为:
识别证件图像中的预定边缘,预定边缘是位于证件的预定方向上的边缘;
在识别出n条候选的预定边缘时,将n条候选的预定边缘中的一条候选的预定边缘确定为目标预定边缘,n≥2;
根据目标预定边缘在证件图像中识别出至少一个信息区域。
关于上述实施例中的装置,其中各个模块执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
图13是根据一示例性实施例示出的一种用与区域识别方法的装置的框图。例如,装置1300可以是移动电话,计算机,数字广播终端,消息收发设备,游戏控制台,平板设备,医疗设备,健身设备,个人数字助理等。
参照图13,装置1300可以包括以下一个或多个组件:处理组件1302,存储器1304,电源组件1306,多媒体组件1308,音频组件1310,输入/输出(I/O)接口1312,传感器组件1314,以及通信组件1316。
处理组件1302通常控制装置1300的整体操作,诸如与显示,电话呼叫,数据通信,相机操作和记录操作相关联的操作。处理组件1302可以包括一个或多个处理器1318来执行指令,以完成上述的方法的全部或部分步骤。此外,处理组件1302可以包括一个或多个模块,便于处理组件1302和其他组件之间的交互。例如,处理组件1302可以包括多媒体模块,以方便多媒体组件1308和处理组件1302之间的交互。
存储器1304被配置为存储各种类型的数据以支持在装置1300的操作。这些数据的示例包括用于在装置1300上操作的任何应用程序或方法的指令,联系人数据,电话簿数据,消息,图片,视频等。存储器1304可以由任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(SRAM),电可擦除可编程只读存储器(EEPROM),可擦除可编程只读存储器 (EPROM),可编程只读存储器(PROM),只读存储器(ROM),磁存储器,快闪存储器,磁盘或光盘。
电源组件1306为装置1300的各种组件提供电力。电源组件1306可以包括电源管理系统,一个或多个电源,及其他与为装置1300生成、管理和分配电力相关联的组件。
多媒体组件1308包括在装置1300和用户之间的提供一个输出接口的屏幕。在一些实施例中,屏幕可以包括液晶显示器(LCD)和触摸面板(TP)。如果屏幕包括触摸面板,屏幕可以被实现为触摸屏,以接收来自用户的输入信号。触摸面板包括一个或多个触摸传感器以感测触摸、滑动和触摸面板上的手势。触摸传感器可以不仅感测触摸或滑动动作的边界,而且还检测与触摸或滑动操作相关的持续时间和压力。在一些实施例中,多媒体组件1308包括一个前置摄像头和/或后置摄像头。当装置1300处于操作模式,如拍摄模式或视频模式时,前置摄像头和/或后置摄像头可以接收外部的多媒体数据。每个前置摄像头和后置摄像头可以是一个固定的光学透镜系统或具有焦距和光学变焦能力。
音频组件1310被配置为输出和/或输入音频信号。例如,音频组件1310包括一个麦克风(MIC),当装置1300处于操作模式,如呼叫模式、记录模式和语音识别模式时,麦克风被配置为接收外部音频信号。所接收的音频信号可以被进一步存储在存储器1304或经由通信组件1316发送。在一些实施例中,音频组件1310还包括一个扬声器,用于输出音频信号。
I/O接口1312为处理组件1302和外围接口模块之间提供接口,上述外围接口模块可以是键盘,点击轮,按钮等。这些按钮可包括但不限于:主页按钮、音量按钮、启动按钮和锁定按钮。
传感器组件1314包括一个或多个传感器,用于为装置1300提供各个方面的状态评估。例如,传感器组件1314可以检测到装置1300的打开/关闭状态,组件的相对定位,例如组件为装置1300的显示器和小键盘,传感器组件1314还可以检测装置1300或装置1300一个组件的位置改变,用户与装置1300接触的存在或不存在,装置1300方位或加速/减速和装置1300的温度变化。传感器组件1314可以包括接近传感器,被配置用来在没有任何的物理接触时检测附近物体的存在。传感器组件1314还可以包括光传感器,如CMOS或CCD图像传感器,用于在成像应用中使用。在一些实施例中,该传感器组件1314还可以包 括加速度传感器,陀螺仪传感器,磁传感器,压力传感器或温度传感器。
通信组件1316被配置为便于装置1300和其他设备之间有线或无线方式的通信。装置1300可以接入基于通信标准的无线网络,如Wi-Fi,2G或3G,或它们的组合。在一个示例性实施例中,通信组件1316经由广播信道接收来自外部广播管理系统的广播信号或广播相关信息。在一个示例性实施例中,通信组件1316还包括近场通信(NFC)模块,以促进短程通信。例如,在NFC模块可基于射频识别(RFID)技术,红外数据协会(IRDA)技术,超宽带(UWB)技术,蓝牙(BT)技术和其他技术来实现。
在示例性实施例中,装置1300可以被一个或多个应用专用集成电路(ASIC)、数字信号处理器(DSP)、数字信号处理设备(DSPD)、可编程逻辑器件(PLD)、现场可编程门阵列(FPGA)、控制器、微控制器、微处理器或其他电子元件实现,用于执行上述区域识别方法。
在示例性实施例中,还提供了一种包括指令的非临时性计算机可读存储介质,例如包括指令的存储器1304,上述指令可由装置1300的处理器1318执行以完成上述区域识别方法。例如,非临时性计算机可读存储介质可以是ROM、随机存取存储器(RAM)、CD-ROM、磁带、软盘和光数据存储设备等。
本领域技术人员在考虑说明书及实践这里公开的发明后,将容易想到本公开的其它实施方案。本申请旨在涵盖本公开的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本公开的一般性原理并包括本公开未公开的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本公开的真正范围和精神由下面的权利要求指出。
应当理解的是,本公开并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围进行各种修改和改变。本公开的范围仅由所附的权利要求来限制。

Claims (15)

  1. 一种区域识别方法,其特征在于,所述方法包括:
    识别证件图像中的预定边缘,所述预定边缘是位于所述证件的预定方向上的边缘;
    在识别出n条候选的所述预定边缘时,将n条候选的所述预定边缘中的一条候选的所述预定边缘确定为目标预定边缘,n≥2;
    根据所述目标预定边缘在所述证件图像中识别出至少一个信息区域。
  2. 根据权利要求1所述的方法,其特征在于,所述在识别出n条候选的所述预定边缘时,将n条候选的所述预定边缘中的一条候选的所述预定边缘确定为目标预定边缘,包括:
    对n条候选的所述预定边缘进行排序;
    使用第i条候选的所述预定边缘和第一相对位置关系,在所述证件图像中尝试识别目标信息区域,1≤i≤n;所述第一相对位置关系是所述目标预定边缘与所述目标信息区域之间的相对位置关系;
    若成功识别出所述目标信息区域,则将所述第i条候选的所述预定边缘确定为所述目标预定边缘;
    若未能识别出所述目标信息区域,则令i=i+1,重新执行所述使用第i条候选的所述预定边缘和第一相对位置关系,在所述证件图像中尝试识别目标信息区域的步骤。
  3. 根据权利要求2所述的方法,其特征在于,所述对n条候选的所述预定边缘进行排序,包括:
    对于每一条候选的所述预定边缘,将所述预定边缘与处理后的所述证件图像中相同位置处的前景色像素点求交,得到所述预定边缘对应的交点数;所述处理后的所述证件图像是经过索贝尔水平滤波和二值化的图像;
    按照所述交点数由多到少的顺序,将n条候选的所述预定边缘进行排序。
  4. 根据权利要求2所述的方法,其特征在于,所述使用第i条候选的所述 预定边缘和第一相对位置关系,在所述证件图像中尝试识别目标信息区域,包括:
    使用第i条候选的所述预定边缘和所述第一相对位置关系,在所述证件图像中截取出兴趣区域;
    识别所述兴趣区域中是否存在符合预定特征的字符区域,所述预定特征是所述目标信息区域中的字符区域所具有的特征。
  5. 根据权利要求4所述的方法,其特征在于,所述识别所述兴趣区域中是否存在符合预定特征的字符区域,包括:
    对所述兴趣区域进行二值化,得到二值化后的兴趣区域;
    对所述二值化后的兴趣区域按照水平方向计算第一直方图,所述第一直方图包括:每行像素点的竖坐标和所述每行像素点中前景色像素点的累加值;
    对所述二值化后的兴趣区域按照竖直方向计算第二直方图,所述第二直方图包括:每列像素点的横坐标和所述每列像素点中前景色像素点的累加值;
    若所述第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且所述第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数,则在所述兴趣区域中成功识别出符合所述预定特征的字符区域;
    若所述第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度不符合预定高度区间,或者所述第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数,则在所述兴趣区域中未能识别出符合所述预定特征的字符区域。
  6. 根据权利要求1至5任一所述的方法,其特征在于,所述识别证件图像中的预定边缘,包括:
    对所述证件图像进行索贝尔水平滤波和二值化,得到处理后的所述证件图像;
    对所述处理后的所述证件图像中的预定区域进行直线检测,得到至少一条直线;
    在所述直线为n条时,n≥2,将所述n条直线识别为n条候选的所述预定 边缘。
  7. 根据权利要求1至5任一所述的方法,其特征在于,所述根据所述目标预定边缘在所述证件图像中确定出至少一个信息区域,包括:
    根据所述目标预定边缘和第二相对位置关系,确定出至少一个信息区域;所述第二相对位置关系是所述目标预定边缘与所述信息区域之间的相对位置关系。
  8. 一种区域识别装置,其特征在于,所述装置包括:
    识别模块,被配置为识别证件图像中的预定边缘,所述预定边缘是位于所述证件的预定方向上的边缘;
    确定模块,被配置为在识别出n条候选的所述预定边缘时,将n条候选的所述预定边缘中的一条候选的所述预定边缘确定为目标预定边缘,n≥2;
    区域识别模块,被配置为根据所述目标预定边缘在所述证件图像中识别出至少一个信息区域。
  9. 根据权利要求8所述的装置,其特征在于,所述确定模块,包括:
    第一排序子模块,被配置为对n条候选的所述预定边缘进行排序;
    第一识别子模块,被配置为使用第i条候选的所述预定边缘和第一相对位置关系,在所述证件图像中尝试识别目标信息区域,1≤i≤n;所述第一相对位置关系是所述目标预定边缘与所述目标信息区域之间的相对位置关系。
    第二识别子模块,被配置为在成功识别出所述目标信息区域时,将所述第i条候选的所述预定边缘确定为所述目标预定边缘;
    第三识别子模块,被配置为在未能识别出所述目标信息区域时,令i=i+1,重新返回第一识别子模块。
  10. 根据权利要求9所述的装置,其特征在于,所述第一排序子模块,包括:
    求交子模块,被配置为对于每一条候选的所述预定边缘,将所述预定边缘与处理后的所述证件图像中相同位置处的前景色像素点求交,得到所述预定边 缘对应的交点数;所述处理后的所述证件图像是经过索贝尔水平滤波和二值化的图像;
    第二排序子模块,被配置为按照所述交点数由多到少的顺序,将n条候选的所述预定边缘进行排序。
  11. 根据权利要求9所述的装置,其特征在于,所述第一识别子模块,包括:
    截取子模块,被配置为使用第i条候选的所述预定边缘和所述第一相对位置关系,在所述证件图像中截取出兴趣区域;
    第四识别子模块,被配置为识别所述兴趣区域中是否存在符合预定特征的字符区域,所述预定特征是所述目标信息区域中的字符区域所具有的特征。
  12. 根据权利要求11所述的装置,其特征在于,所述第四识别子模块,包括:
    二值化子模块,被配置为对所述兴趣区域进行二值化,得到二值化后的兴趣区域;
    第一计算子模块,被配置为对所述二值化后的兴趣区域按照水平方向计算第一直方图,所述第一直方图包括:每行像素点的竖坐标和所述每行像素点中前景色像素点的累加值;
    第二计算子模块,被配置为对所述二值化后的兴趣区域按照竖直方向计算第二直方图,所述第二直方图包括:每列像素点的横坐标和所述每列像素点中前景色像素点的累加值;
    字符识别子模块,被配置为在所述第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度符合预定高度区间,且所述第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数符合预定个数时,在所述兴趣区域中成功识别出符合所述预定特征的字符区域;
    第五识别子模块,被配置为在所述第一直方图中前景色像素点的累加值大于第一阈值的行所组成的连续行集合的高度不符合预定高度区间,或者所述第二直方图中前景色像素点的累加值大于第二阈值的列所组成的连续列集合的个数不符合预定个数时,在所述兴趣区域中未能识别出符合所述预定特征的字符 区域。
  13. 根据权利要求8至12任一所述的装置,其特征在于,所述识别模块,包括:
    滤波子模块,被配置为对所述证件图像进行索贝尔水平滤波和二值化,得到处理后的所述证件图像;
    检测子模块,被配置为对所述处理后的所述证件图像中的预定区域进行直线检测,得到至少一条直线;
    边缘识别子模块,被配置为在所述直线为n条时,n≥2,将所述n条直线识别为n条候选的所述预定边缘。
  14. 根据权利要求8至12任一所述的装置,其特征在于,
    所述区域识别模块,被配置为根据所述目标预定边缘和第二相对位置关系,确定出至少一个信息区域;所述第二相对位置关系是所述目标预定边缘与所述信息区域之间的相对位置关系。
  15. 一种区域识别装置,其特征在于,所述装置包括:
    处理器;
    用于存储所述处理器可执行指令的存储器;
    其中,所述处理器被配置为:
    识别证件图像中的预定边缘,所述预定边缘是位于所述证件的预定方向上的边缘;
    在识别出n条候选的所述预定边缘时,将n条候选的所述预定边缘中的一条候选的所述预定边缘确定为目标预定边缘,n≥2;
    根据所述目标预定边缘在所述证件图像中识别出至少一个信息区域。
PCT/CN2015/099286 2015-10-30 2015-12-28 区域识别方法及装置 Ceased WO2017071061A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
JP2017547044A JP6392467B2 (ja) 2015-10-30 2015-12-28 領域識別方法及び装置
KR1020167005597A KR101758580B1 (ko) 2015-10-30 2015-12-28 영역 인식 방법 및 장치
RU2016109047A RU2633184C2 (ru) 2015-10-30 2015-12-28 Способ и устройство для идентификации области
MX2016003318A MX374325B (es) 2015-10-30 2015-12-28 Metodo y aparato para identificacion de area.

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510726012.7 2015-10-30
CN201510726012.7A CN105550633B (zh) 2015-10-30 2015-10-30 区域识别方法及装置

Publications (1)

Publication Number Publication Date
WO2017071061A1 true WO2017071061A1 (zh) 2017-05-04

Family

ID=55829816

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/099286 Ceased WO2017071061A1 (zh) 2015-10-30 2015-12-28 区域识别方法及装置

Country Status (8)

Country Link
US (1) US10095949B2 (zh)
EP (1) EP3163503A1 (zh)
JP (1) JP6392467B2 (zh)
KR (1) KR101758580B1 (zh)
CN (1) CN105550633B (zh)
MX (1) MX374325B (zh)
RU (1) RU2633184C2 (zh)
WO (1) WO2017071061A1 (zh)

Families Citing this family (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106710372A (zh) * 2017-01-11 2017-05-24 成都市极米科技有限公司 拼音卡片识别方法、装置以及终端设备
CN107358150B (zh) * 2017-06-01 2020-08-18 深圳赛飞百步印社科技有限公司 物体边框识别方法、装置和高拍仪
CN109559344B (zh) * 2017-09-26 2023-10-13 腾讯科技(上海)有限公司 边框检测方法、装置及存储介质
US10740644B2 (en) * 2018-02-27 2020-08-11 Intuit Inc. Method and system for background removal from documents
KR102063036B1 (ko) * 2018-04-19 2020-01-07 한밭대학교 산학협력단 딥러닝과 문자인식으로 구현한 시각주의 모델 기반의 문서 종류 자동 분류 장치 및 방법
JP6574921B1 (ja) * 2018-07-06 2019-09-11 楽天株式会社 画像処理システム、画像処理方法、及びプログラム
JP7018372B2 (ja) * 2018-08-30 2022-02-10 株式会社Pfu 画像処理装置及び画像処理方法
US10331966B1 (en) * 2018-10-19 2019-06-25 Capital One Services, Llc Image processing to detect a rectangular object
CN111325197B (zh) * 2018-11-29 2023-10-31 北京搜狗科技发展有限公司 数据处理方法和装置、用于数据处理的装置
CN113811924A (zh) * 2019-05-13 2021-12-17 普利斯梅德实验室有限公司 用于验证电传导安全特征的设备和方法以及用于电传导安全特征的验证设备
CN110399872B (zh) * 2019-06-20 2023-04-28 创新先进技术有限公司 图像处理方法以及装置
CN110738204B (zh) * 2019-09-18 2023-07-25 平安科技(深圳)有限公司 一种证件区域定位的方法及装置
CN111260631B (zh) * 2020-01-16 2023-05-05 成都地铁运营有限公司 一种高效刚性接触线结构光光条提取方法
CN111724346A (zh) * 2020-05-21 2020-09-29 北京配天技术有限公司 直线边缘检测方法、机器人及存储装置
CN113835582B (zh) * 2021-09-27 2024-03-15 青岛海信移动通信技术有限公司 一种终端设备、信息显示方法和存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101751568A (zh) * 2008-12-12 2010-06-23 汉王科技股份有限公司 证件号码定位和识别方法
CN104217444A (zh) * 2013-06-03 2014-12-17 支付宝(中国)网络技术有限公司 定位卡片区域的方法和设备
CN104268864A (zh) * 2014-09-18 2015-01-07 小米科技有限责任公司 卡片边缘提取方法和装置
EP2858007A1 (en) * 2012-05-31 2015-04-08 Southeast University Sift feature bag based bovine iris image recognition method

Family Cites Families (18)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08161423A (ja) * 1994-12-06 1996-06-21 Dainippon Printing Co Ltd 照明装置および文字読取装置
DE19958553A1 (de) * 1999-12-04 2001-06-07 Luratech Ges Fuer Luft Und Rau Verfahren zur Kompression von gescannten Farb- und/oder Graustufendokumenten
JP3823782B2 (ja) * 2001-08-31 2006-09-20 日産自動車株式会社 先行車両認識装置
KR100608443B1 (ko) 2004-06-30 2006-08-02 노틸러스효성 주식회사 형태복원을 통한 지로인식방법
RU2309456C2 (ru) * 2005-12-08 2007-10-27 "Аби Софтвер Лтд." Способ распознавания текстовой информации из векторно-растрового изображения
US8265393B2 (en) * 2007-05-01 2012-09-11 Compulink Management Center, Inc. Photo-document segmentation method and system
KR101291195B1 (ko) 2007-11-22 2013-07-31 삼성전자주식회사 문자인식장치 및 방법
US8345106B2 (en) * 2009-09-23 2013-01-01 Microsoft Corporation Camera-based scanning
US9639949B2 (en) * 2010-03-15 2017-05-02 Analog Devices, Inc. Edge orientation for second derivative edge detection methods
JP5742399B2 (ja) * 2011-04-06 2015-07-01 富士ゼロックス株式会社 画像処理装置及びプログラム
CN103577817B (zh) 2012-07-24 2017-03-01 阿里巴巴集团控股有限公司 表单识别方法与装置
CN102982160B (zh) 2012-12-05 2016-04-20 上海合合信息科技发展有限公司 方便电子化的专业笔记本及其电子化文档的自动分类方法
JP6161484B2 (ja) * 2013-09-19 2017-07-12 株式会社Pfu 画像処理装置、画像処理方法及びコンピュータプログラム
CN103488984B (zh) * 2013-10-11 2017-04-12 瑞典爱立信有限公司 基于智能移动设备的二代身份证识别方法及装置
CN104573616A (zh) * 2013-10-29 2015-04-29 腾讯科技(深圳)有限公司 一种信息识别方法、相关装置及系统
KR101663790B1 (ko) 2014-04-11 2016-10-10 (주) 엠티콤 얼굴인식장치 및 그 동작 방법
CN104408450A (zh) * 2014-11-21 2015-03-11 深圳天源迪科信息技术股份有限公司 身份证识别方法、装置及系统
CN105528600A (zh) 2015-10-30 2016-04-27 小米科技有限责任公司 区域识别方法及装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101751568A (zh) * 2008-12-12 2010-06-23 汉王科技股份有限公司 证件号码定位和识别方法
EP2858007A1 (en) * 2012-05-31 2015-04-08 Southeast University Sift feature bag based bovine iris image recognition method
CN104217444A (zh) * 2013-06-03 2014-12-17 支付宝(中国)网络技术有限公司 定位卡片区域的方法和设备
CN104268864A (zh) * 2014-09-18 2015-01-07 小米科技有限责任公司 卡片边缘提取方法和装置

Also Published As

Publication number Publication date
CN105550633A (zh) 2016-05-04
EP3163503A1 (en) 2017-05-03
US10095949B2 (en) 2018-10-09
KR20170061633A (ko) 2017-06-05
KR101758580B1 (ko) 2017-07-14
JP6392467B2 (ja) 2018-09-19
MX374325B (es) 2025-03-06
RU2016109047A (ru) 2017-09-19
RU2633184C2 (ru) 2017-10-11
JP2018500703A (ja) 2018-01-11
US20170124419A1 (en) 2017-05-04
MX2016003318A (es) 2018-06-22
CN105550633B (zh) 2018-12-11

Similar Documents

Publication Publication Date Title
CN105550633B (zh) 区域识别方法及装置
US10127471B2 (en) Method, device, and computer-readable storage medium for area extraction
KR101864759B1 (ko) 영역 인식 방법 및 장치
KR101782633B1 (ko) 영역 인식 방법 및 장치
RU2639668C2 (ru) Способ и устройство для идентификации области
CN106228168B (zh) 卡片图像反光检测方法和装置
WO2017071064A1 (zh) 区域提取方法、模型训练方法及装置
CN105894042B (zh) 检测证件图像遮挡的方法和装置
CN106250894A (zh) 卡片信息识别方法及装置
CN106296665A (zh) 卡片图像模糊检测方法和装置

Legal Events

Date Code Title Description
ENP Entry into the national phase

Ref document number: 2017547044

Country of ref document: JP

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 20167005597

Country of ref document: KR

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 2016109047

Country of ref document: RU

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: MX/A/2016/003318

Country of ref document: MX

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15907122

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15907122

Country of ref document: EP

Kind code of ref document: A1