WO2020143316A1 - 证件图像提取方法及终端设备 - Google Patents

证件图像提取方法及终端设备 Download PDF

Info

Publication number
WO2020143316A1
WO2020143316A1 PCT/CN2019/118133 CN2019118133W WO2020143316A1 WO 2020143316 A1 WO2020143316 A1 WO 2020143316A1 CN 2019118133 W CN2019118133 W CN 2019118133W WO 2020143316 A1 WO2020143316 A1 WO 2020143316A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
document
pixel
feature model
original image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/118133
Other languages
English (en)
French (fr)
Inventor
黄锦伦
熊冬根
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Priority to SG11202100270VA priority Critical patent/SG11202100270VA/en
Priority to JP2021500946A priority patent/JP2021531571A/ja
Publication of WO2020143316A1 publication Critical patent/WO2020143316A1/zh
Priority to US17/167,075 priority patent/US11790499B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/90Dynamic range modification of images or parts thereof
    • G06T5/92Dynamic range modification of images or parts thereof based on global image properties
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/22Matching criteria, e.g. proximity measures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0499Feedforward networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/60Image enhancement or restoration using machine learning, e.g. neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/80Geometric correction
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/19Recognition using electronic means
    • G06V30/191Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
    • G06V30/19147Obtaining sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/19Recognition using electronic means
    • G06V30/191Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
    • G06V30/19173Classification techniques
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/40Document-oriented image-based pattern recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/40Document-oriented image-based pattern recognition
    • G06V30/41Analysis of document content
    • G06V30/413Classification of content, e.g. text, photographs or tables
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/10Machine learning using kernel methods, e.g. support vector machines [SVM]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N7/00Computing arrangements based on specific mathematical models
    • G06N7/01Probabilistic graphical models, e.g. probabilistic networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10024Color image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30176Document

Definitions

  • the present application belongs to the field of computer application technology, and in particular relates to a method for extracting a certificate image, a terminal device, and a computer non-volatile readable storage medium.
  • Machine vision is a function that allows "things" to be seen, not only with information collection functions, but also with advanced functions such as processing and recognition.
  • the equipment cost of machine vision is low, and the most commonly used equipment is the camera.
  • the installation rate of public cameras and cameras of households and enterprises in major cities has greatly increased, and the installation ratio of cameras in households and enterprises is also very high. Will use a lot of cameras. With the rapid popularization of cameras, related applications of machine vision technology will develop more rapidly. With the development of the field of machine vision, the identity verification technology of document photos will also be more widely used in this wave.
  • Embodiments of the present application provide a certificate image extraction method, terminal device, and computer non-volatile readable storage medium to solve the problem of poor quality and inaccuracy of the obtained certificate image due to the influence of many external environments in the prior art problem.
  • a first aspect of an embodiment of the present application provides a method for extracting a certificate image, including:
  • the original image containing the ID image
  • the original image is captured by the camera device
  • the pre-trained document feature model determine the position of the document image from the balanced image; the document feature model is obtained by training based on the historical document image, the document image model and preset initial weights;
  • the image of the certificate is extracted from the balanced image.
  • a second aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, and the processor executes the computer
  • the following steps are realized when the instructions are readable:
  • the original image containing the ID image
  • the original image is captured by the camera device
  • the pre-trained document feature model determine the position of the document image from the balanced image; the document feature model is obtained by training based on the historical document image, the document image model and preset initial weights;
  • the image of the certificate is extracted from the balanced image.
  • a third aspect of the embodiments of the present application provides a terminal device, including:
  • An obtaining unit configured to obtain an original image including a document image; the original image is obtained by a camera device;
  • the processing unit is configured to perform white balance processing on the original image according to the component values of each pixel in the original image in the three color components of red, green, and blue to obtain a balanced image;
  • the determining unit is used to determine the position of the document image from the balanced image according to the pre-trained document feature model; the document feature model is based on the historical document image, the document image model and preset initial weights Get training;
  • the extracting unit is configured to extract the image of the document from the balanced image according to the position of the document image.
  • a fourth aspect of the embodiments of the present application provides a computer non-volatile readable storage medium, the computer storage medium storing computer readable instructions, which when executed by a processor causes the processing Implements the method of the first aspect described above.
  • the original image is obtained by a camera device; according to the component value of each pixel in the original image in the three color components of red, green, and blue, Perform white balance processing on the original image to obtain a balanced image; according to the pre-trained document feature model, determine the position of the document image from the balanced image; the document feature model is based on the historical document image, document image model and preset The initial weight of is obtained through training; according to the position of the certificate, the image of the certificate is extracted from the balanced image. According to the document feature model, locate the position of the document in the original image and extract the document image in the image, which improves the accuracy of extracting the document image from the original image.
  • FIG. 1 is a flowchart of a method for extracting a document image provided in Embodiment 1 of the present application;
  • FIG. 2 is a flowchart of a method for extracting a document image provided in Embodiment 2 of the present application;
  • FIG. 3 is a schematic diagram of a terminal device provided in Embodiment 3 of this application.
  • FIG. 4 is a schematic diagram of a terminal device provided in Embodiment 4 of the present application.
  • FIG. 1 is a flowchart of a method for extracting a document image provided in Embodiment 1 of the present application.
  • the execution subject of the certificate image extraction method is a terminal.
  • Terminals include but are not limited to mobile terminals such as smart phones, tablet computers, and wearable devices, and may also be desktop computers.
  • the method for extracting the certificate image as shown in the figure may include the following steps:
  • S101 Acquire an original image including a document image; the original image is captured by a camera device.
  • the certificate information collection system can realize the automatic extraction of certificate information and the entry of identification materials such as ID cards and passports through radio frequency identification technology and image recognition technology.
  • the richness of information acquisition means and the development of image processing technology make the document reader smaller in size, faster in information extraction, and lower in error rate of information. While improving public safety and management efficiency, it also brings great convenience to both sides of the business.
  • the certificate information collection system is also conducive to the expansion of the real-name system application. The development of the certificate information collection system has made it possible for trains, cars, subways and other places with large traffic to carry out the real-name system, which will greatly guarantee the safety of railways, highways and urban rail transit.
  • Mobile smart terminals are equipped with various operating systems like computers, but they are relatively small compared to computers, easy to carry, and have wireless Internet access. Users can download various applications corresponding to operating systems according to their needs. More common mobile smart terminals in life include smart phones, tablet computers, in-vehicle computers, and wearable mobile devices. Smart phones are currently more commonly used mobile smart terminals. Users can install applications, games, or functional programs provided by third-party service providers according to their own preferences or needs to meet users' needs for smart terminal functions. In recent years, with the continuous development of technology, all kinds of certificates are no longer a certificate, but a card similar to an ID card. With the use of certificates, the entry of certificate information has also become an important issue.
  • the traditional information entry method is to manually fill in the information in the relevant form, and then the internal staff can store the key information into the computer according to the content of the form, or scan and upload the certificate to the designated place.
  • the former method does not limit the location of information entry, each entry of information requires a lot of human and material resources, and is prone to erroneous input.
  • the latter has improved the efficiency and accuracy of information entry, the location of use is relatively fixed.
  • the emergence of mobile intelligent terminals makes it possible to enter certificate information anytime, anywhere.
  • the information recognition system on the mobile intelligent terminal can be widely used in service industry, transportation system, public security system and other parts that need to check the document information. It can complete the collection and inspection of the document information without a large number of personnel, improving the collection and inspection of the document
  • the efficiency and accuracy of information recognition have broad application prospects.
  • the user can upload the original image captured by the mobile terminal to the server or the image processing terminal. After receiving the original image, the image processing terminal processes and recognizes the original image.
  • the application scenario in this solution can be to send an image collection instruction to the user on the website where the user's ID image is acquired and verified.
  • the user takes photos through his own terminal device, such as ID cards, passports and other photos, through the application software in the mobile terminal Or, the webpage sends the captured image to the execution subject, and after acquiring the target image, the execution subject processes and recognizes the original image.
  • S102 Perform white balance processing on the original image according to the component values of each pixel in the original image in the three color components of red, green, and blue to obtain a balanced image.
  • RGB Red Green Blue
  • a table is used to store all the colors in the image, and the actual image data is no longer RGB data, but the index of RGB data in that table.
  • the color quantization process often uses two types of color palettes. One type is true color or pseudo-color image quantization to palette image; the other type is palette image to continue quantization. As the storage capacity of computers continues to increase, palette images gradually fade out of the personal computer arena, but they are still widely used in mobile phones and other special devices, especially in mobile game applications.
  • balance is an indicator that describes the accuracy of white after the three primary colors of red, green, and blue are mixed in the display.
  • White balance is a very important concept in the field of TV camera, through which it can solve a series of problems of color reproduction and tone processing.
  • White balance is produced with the true color reproduction of electronic images.
  • White balance was applied earlier in the field of professional cameras and is now widely used in household electronic products.
  • the development of technology has made white balance adjustment more and more simple and easy.
  • many users still do not understand the working principle of white balance, and there are many misunderstandings in understanding. It is to realize that the camera image can accurately reflect the color condition of the subject.
  • white balance processing can be performed on the original image through manual white balance and automatic white balance to obtain a balanced image.
  • S103 Determine the position of the document image from the balanced image according to a pre-trained document feature model; the document feature model is obtained by training based on a historical document image, a document image model, and preset initial weights.
  • the methods When recognizing the characters of documents, the methods mainly include hidden Markov model, neural network, support vector machine and template matching. All methods that use hidden Markov models require preprocessing and setting parameters based on existing knowledge. This method achieves a high recognition rate through complex preprocessing and parameterization. It can also be used to train the document feature model through a multi-layer perception neural network. The neural network uses backward feedback to train this network. It must be trained many times before it can be trained. Obtaining better results is a time-consuming process, and both the number of hidden layers and the number of hidden neurons must be obtained through experimental methods.
  • the neural network contains 24 input layer neurons, 15 hidden layer neurons, and 36 output layer neurons to identify the documents in the balanced image.
  • the neural network here is mainly to study the connection and organization of these neurons.
  • neural networks it can be divided into two types, one is layered and the other is meshed.
  • the neurons are arranged in a layer.
  • these neurons are arranged side by side to form a tight mechanism.
  • the neurons are connected, but for the neurons in each layer, they cannot be connected.
  • each neuron can be interconnected.
  • S104 Extract the image of the certificate from the balanced image according to the position of the image of the certificate.
  • the ID image is extracted from the balanced image according to the position of the ID image.
  • the extraction method can be directly cropped from the balanced image to extract the certificate image, or it can be a method of deleting the image area except the certificate image and retaining the certificate image, which is not limited here .
  • the image edge of the ID image can be detected by the method based on the edge and the gradient, and the ID image can be extracted based on the image edge.
  • the edge-based ID image method considers that there is a large difference between the ID image in the natural scene and the edge of the background. This method performs edge detection on the characters to locate the ID image through the edge information.
  • the image edge of the document image can be determined by the Sobel operator, Robert operator, and Laplace operator. Among them, the Sobel operator determines whether a certain pixel point in the ID image is greater than the threshold to determine whether the point is an edge point.
  • the Robert operator is suitable for images with a large difference between text and background, and the edge obtained after detection is thicker.
  • Laplace operator is very sensitive to noise, easy to produce bilateral effects, and is not directly used to detect edges. It is also possible to use the ID image positioning method based on the connected domain to convert the original image into a binary image to reduce the impact of noise.
  • the morphological corrosion expansion algorithm is used to connect the ID image area and the image is segmented using the discrimination between the ID image and the white background. Then, the connected domains of non-document images are excluded by various aspects of the document image, thereby obtaining the document image positioning method based on the connected domain of the document image area.
  • the document image positioning speed is faster, which can improve the recognition efficiency of the document image and the text in the document image. .
  • the original image containing the ID image is obtained; the original image is captured by a camera device; according to the component values of each pixel in the original image in the three color components of red, green, and blue, Perform white balance processing on the original image to obtain a balanced image; determine the position of the document image from the balanced image according to the pre-trained document feature model; the document feature model is based on the historical document image, the document image model and The preset initial weights are obtained through training; according to the position of the ID image, the ID image is extracted from the balanced image. According to the document feature model, locate the position of the document in the original image and extract the document image in the image, which improves the accuracy of extracting the document image from the original image.
  • FIG. 2 is a flowchart of a method for extracting a document image provided in Embodiment 2 of the present application.
  • the execution subject of the certificate image extraction method is a terminal.
  • Terminals include but are not limited to mobile terminals such as smart phones, tablet computers, and wearable devices, and may also be desktop computers.
  • the method for extracting the certificate image as shown in the figure may include the following steps:
  • S201 Collect historical document images, and filter the historical document images according to preset target image requirements to obtain a target image.
  • the document feature model can be trained based on the historical document image to extract the document image from the original image.
  • the data of the training document feature model in this solution may be historical document images, and the historical document images include historical images acquired before the formal document image extraction is performed.
  • the acquired historical document image may have an image that does not meet the requirements.
  • the acquired historical image is filtered according to the preset target image requirements to obtain the target image.
  • the preset target image requirements may be the pixel, size or shooting time requirements of the image, in addition to the completeness of the detection image, the type requirements of the certificate image, etc. These requirements can be determined by the executor, which is not limited here.
  • the acquired historical document image is matched with the target image requirements, and the historical document image whose matching degree is greater than the matching degree threshold is determined as the target image.
  • Document images obtained by acquiring various camera devices are used as initial samples for neural network training.
  • the selection of the training set will directly affect the network learning and training time, weight matrix and learning and training effect.
  • the image edge is more obvious, and the image with edges distributed in most areas of the image is used as the initial sample of the training set.
  • the image edges are clear, the particle edges are distributed in the entire image, and the texture features are rich, so that The neural network is well trained, network information such as network weights will remember more edge information, and can detect images better.
  • S202 Perform pixel identification on the target image according to a preset ID image template, and determine at least one center pixel point in the target image.
  • the pixel point with the outstanding representativeness of the image sample can be determined as the center pixel point through image recognition.
  • the processed original image is the ID card image taken by the user, according to the size and text position of the avatar in the known ID card image, it can be determined that the four corners of the avatar are center pixels, or a certain These texts are the central pixels, such as the first word in the ID card or the first word in each line, etc. Further, you can also pre-set the type of image obtained, such as ID card photos, real estate certificate photos, etc.
  • the number of pixels around the center pixel can be at least two. Preferably, 8 pixels around the center pixel can be determined for learning and training, and the pixel of each pixel in the image can be more clearly determined. happening.
  • S203 Set initial parameters of the training model, perform learning and training according to the initial parameters, pixel values of each of the central pixel and pixels around the central pixel, and obtain a credential feature model based on a neural network.
  • the learning and training in the application process are the key links.
  • the network can have the ability of association, memory and prediction.
  • the initial parameters of the network include the initial structure of the network, the weight of the connection, the threshold, and the learning rate. Different settings will affect the convergence rate of the network to a certain extent.
  • the selection of initial parameters is very important but very difficult.
  • network construction relies mainly on observation and experience.
  • the initial value of the model When training the document feature model, first determine the initial value of the model.
  • the initial weights and thresholds of the network are generally randomly selected from [-1, 1] or [0, 1], and some improved algorithms will make appropriate changes to the interval .
  • the vector is normalized. In the learning and training process, the node input should not be too large, and the adjustment of the weight value that is too small will not be conducive to network learning and training.
  • the training image is based on grayscale, so the image matrix is an integer value between [0, 255], and the feature vector dimension is relatively high. In order to improve the network training speed, the feature vector will be unified and normalized. Think of the feature vector as a row vector, expressed as:
  • X (x 0 ,x 1 ,...,x 9 ); where x 0 ,x 1 ,...,x 9 are used to represent the pixel values of the central pixel and its surrounding pixels, respectively.
  • the number of pixels around the center pixel in this embodiment may be at least two, and preferably, may be 8, to more accurately describe the situation of the pixels around the center pixel.
  • the range of degrees is [0, 255], so the normalized formula in actual processing is:
  • x 0 is used to represent the pixel value of the central pixel.
  • the idea of performing block operation on the image is adopted.
  • Each time an image sample is input in the neural network one or at least two center pixels can be determined, and the template pixels around the center pixel, that is, including the surrounding 8 pixels centered on it, can be trained and learned.
  • the gray values of the pixels are fed into the input layer in order from top to bottom and from left to right.
  • the error propagates in the reverse direction, thereby making the threshold of each neuron and the connection weight between neurons The value changes, so the network can effectively remember more edge information.
  • the training task is completed.
  • the training requirements stipulate that the network training can be stopped at any time.
  • the weights and thresholds trained are all stored in the back-end database, and finally the trained network is saved.
  • S204 Acquire an original image containing the ID image; the original image is captured by a camera device.
  • S204 is implemented in exactly the same way as S101 in the embodiment corresponding to FIG. 1.
  • S101 in the embodiment corresponding to FIG. 1
  • S205 Perform white balance processing on the original image according to the component values of each pixel in the original image in the three color components of red, green, and blue to obtain a balanced image.
  • the original image After acquiring the original image, the original image is subjected to white balance processing according to the component values of each pixel point in the red, green, and blue color components of the original image to obtain a balanced image.
  • step S205 may specifically include steps S2051-S2052:
  • S2051 Estimate the average color difference of each pixel in the original image according to the component values of each pixel in the original image in the three color components of red, green, and blue.
  • image preprocessing Before performing morphological processing or matching, recognition and other processing on the image, processing such as filtering interference information and enhancing effective information on the image is called image preprocessing.
  • the main purpose of image preprocessing is to eliminate interference or irrelevant information in the image, restore useful real information, enhance the detectability of the relevant information and simplify the data to the greatest extent, thereby improving feature extraction, image segmentation, matching and recognition Reliability.
  • the preprocessing of digital color images is generally the restoration and enhancement of brightness and color. In view of the comparison and experiment of various pretreatments, it is found that the white balance processing has a relatively large impact on the final segmentation results of the system, while other pretreatments have little impact. Therefore, the preprocessing in this scheme is mainly white balance processing.
  • color temperature in chromaticity Different light sources have different spectral components and distributions, which is called color temperature in chromaticity.
  • a white object will be reddish when exposed to light of low color temperature, and bluish when exposed to light of high color temperature.
  • the color temperature of the ambient light source will affect the image, making it inevitable that the color deviation.
  • the original color of the subject can be restored under different color temperature conditions, and color correction is required to achieve the correct color balance.
  • the YBR color model is generally used to calculate the color difference.
  • the corresponding relationship between the YBR color system and the RGB color system is as follows:
  • An area is defined in a space where Y is large enough and B and R are small enough, and all pixels in the area are regarded as white, which can participate in the calculation of color difference. Then, the average color difference of white pixels is used to represent the color difference of the entire image to obtain better accuracy.
  • S2052 Calculate the gain of each pixel in the three color components of red, green, and blue according to the average color difference of each pixel.
  • the color gain is used to indicate the freshness of the image.
  • the amount of gain is nothing more than increasing the color contrast, making the colors more vivid and saturated, causing greater visual impact, and on the other hand, it has a certain sharpening effect, making The edge lines are more distinct and clear. You can automatically adjust the contrast, color saturation, and other functions of the image through color gain. This technology of digital cameras can make photos look sharper and more eye-catching.
  • this solution performs color temperature correction on each pixel of the entire image.
  • the specific calculation formula is as follows:
  • the image can also be enhanced to eliminate noise in the image or reduce noise in the image, enhance contrast in the image, and enhance the positioning of the text area.
  • the horizontal correction of the image is to convert the original image into an image with horizontal distribution of text to enhance the positioning accuracy of the text area.
  • the image enhancement method can be Gaussian blur and sharpening.
  • Image Gaussian blur is a common method for blurring details and noise reduction.
  • Gaussian blur processing adds weighted weights to the 8 connected regions of the point, and uses the median as the pixel of the point. value.
  • Gaussian blur smoothing can smooth a lot of noise in the image and highlight the outline of the target image in the image.
  • Gaussian blur smoothing can only be applied to images with complex backgrounds, but the target contour in the image is very obvious. Smoothing can smooth the image details. While smoothing the noise, it will also smooth out some non-obvious outline details.
  • the original image can also be smoothed and filtered, and some measures can be taken to reduce the quality of the image caused during the generation of the image. Can make the image quality improved. Specifically, it is to make targeted compensation for part of the information lost in the image.
  • Another method is to process the image to highlight the image information of a certain part of the image. Further reduce some image information that is not very important. In the image processing of documents. It is often necessary to use the tool of document collection to obtain the image information of the document. In this process, some noise is often generated. Therefore, it is necessary to try to reduce noise. In this way, the image quality can be improved. Can interfere with the generated noise. Get better image information. Enhance important image information.
  • the preprocessing technique of this image is the smoothing of the image.
  • S206 Determine the position of the document image from the balanced image according to the pre-trained document feature model; the document feature model is obtained by training based on the historical document image, the document image model, and preset initial weights.
  • a multi-layer perception neural network is used to train the document feature model.
  • the neural network uses backward feedback to train this kind of network. This network must be trained many times to obtain better results. This process is relatively time-consuming. And the number of hidden layers and the number of hidden neurons must be obtained through experimental methods.
  • the neural network contains 24 input layer neurons, 15 hidden layer neurons, and 36 output layer neurons to identify the documents in the balanced image.
  • step S206 may specifically include step S2061:
  • the initial parameters of the document feature model are corrected according to the following formula: Among them: w ij (k) is used to represent the weight of the kth training; w ij (k+1) is used to represent the weight of the k+1 training; ⁇ is used to represent the learning rate and ⁇ >0 ; E(k) is used to represent the expected value of the position of the credential image obtained by the previous k trainings.
  • the error signal is obtained, and the signal is propagated back from the output terminal.
  • the weighting coefficient is continuously corrected during the propagation process to minimize the error function.
  • the network error is used Mean square error, modify the weight.
  • the adjustment formula is:
  • w ij (k) is used to represent the weight of the kth training
  • w ij (k+1) is used to represent the weight of the k+1 training
  • is used to represent the learning rate and ⁇ >0
  • E(k) is used to represent the expected value of the position of the ID image obtained by the previous k trainings, Represents the negative gradient at the kth time.
  • S207 Extract the image of the certificate from the balanced image according to the position of the image of the certificate.
  • S207 is implemented in exactly the same way as S105 in the embodiment corresponding to FIG. 1.
  • S105 in the embodiment corresponding to FIG. 1
  • S207 is implemented in exactly the same way as S105 in the embodiment corresponding to FIG. 1.
  • a target image is obtained; pixel identification of the target image is performed according to a preset document image template, and the target image Determine at least one central pixel in the image; set the initial parameters of the training model, perform learning and training based on the initial parameters, the pixel values of each of the central pixel and the pixels around the central pixel, and obtain a neural network-based The document feature model.
  • the original image containing the ID image is captured by the camera device; according to the component values of each pixel in the original image in the three color components of red, green, and blue, perform the original image White balance processing to obtain a balanced image; according to the pre-trained document feature model, determine the position of the document image from the balanced image; the document feature model is based on the historical document image, the document image model and the preset initial The weights are obtained through training; according to the position of the ID image, the ID image is extracted from the balanced image.
  • the pre-processed image is positioned according to the document feature model to locate the position of the document, and the document image in the image is extracted, which improves the accuracy of extracting the document image from the original image Sex.
  • FIG. 3 is a schematic diagram of a terminal device provided in Embodiment 3 of the present application.
  • Each unit included in the terminal device is used to execute each step in the embodiments corresponding to FIG. 1 to FIG. 2.
  • the terminal device 300 of this embodiment includes:
  • the obtaining unit 301 is used to obtain an original image containing a certificate image; the original image is obtained by a camera device;
  • the processing unit 302 is configured to perform white balance processing on the original image according to the component values of each pixel point in the red, green, and blue color components of the original image to obtain a balanced image;
  • the determining unit 303 is configured to determine the position of the document image from the balanced image according to a pre-trained document feature model; the document feature model is based on a historical document image, a document image model, and preset initial weights After training,
  • the extraction unit 304 is configured to extract the image of the certificate from the balanced image according to the position of the certificate image.
  • the terminal device may further include:
  • the screening unit is used to collect historical document images, and filter the historical document images according to preset target image requirements to obtain a target image;
  • a recognition unit configured to perform pixel recognition on the target image according to a preset document image template, and determine at least one central pixel point in the target image;
  • the training unit is used to set the initial parameters of the training model, perform learning and training according to the initial parameters, the pixel values of each of the central pixel and the pixels around the central pixel, and obtain a credential feature model based on the neural network .
  • the determining unit 303 may include:
  • a correction unit configured to correct the certificate feature model if the difference in distance between the position of the certificate obtained from the certificate feature model and the actual position of the certificate is greater than or equal to a preset difference threshold Initial parameters.
  • correction unit may include:
  • a distance calculation unit configured to determine a distance difference between the position of the certificate and the actual position of the certificate according to the characteristic model of the certificate;
  • the parameter correction unit is configured to correct the initial parameters of the document feature model according to the following formula if the distance difference is greater than or equal to the difference threshold: Among them: w ij (k) is used to represent the weight of the kth training; w ij (k+1) is used to represent the weight of the k+1 training; ⁇ is used to represent the learning rate and ⁇ >0 ; E(k) is used to represent the expected value of the position of the credential image obtained by the previous k trainings.
  • processing unit 302 may include:
  • a color difference estimation unit configured to estimate the average color difference of each pixel in the original image according to the component values of each pixel in the original image in three color components of red, green, and blue;
  • a gain calculation unit based on the average color difference of each pixel, calculating the gain of each pixel in the three color components of red, green, and blue;
  • the balance processing unit is configured to correct the color temperature of each pixel in the original image according to the gain amount to obtain the balanced image.
  • the original image containing the ID image is obtained; the original image is captured by a camera; according to the component values of each pixel in the original image in the three color components of red, green, and blue, the original The image is subjected to white balance processing to obtain a balanced image; according to the pre-trained document feature model, the position of the document image is determined from the balanced image; the document feature model is based on the historical document image, the document image model and the preset initial The weights are obtained through training; according to the position of the certificate, the image of the certificate is extracted from the balanced image. According to the document feature model, locate the position of the document in the original image and extract the document image in the image, which improves the accuracy of extracting the document image from the original image.
  • the terminal device 4 of this embodiment includes: a processor 40, a memory 41, and computer-readable instructions 42 stored in the memory 41 and executable on the processor 40.
  • the processor 40 executes the computer-readable instructions 42
  • the steps in the above embodiments of each ID image extraction method are implemented, for example, steps 101 to 104 shown in FIG. 1.
  • the processor 40 executes the computer-readable instructions 42
  • the functions of each module/unit in the foregoing device embodiments are realized, for example, the functions of the units 301 to 304 shown in FIG. 3.
  • the computer-readable instructions 42 may be divided into one or more modules/units, the one or more modules/units are stored in the memory 41, and executed by the processor 40, To complete this application.
  • the one or more modules/units may be an instruction segment of a series of computer-readable instructions capable of performing specific functions, and the instruction segment is used to describe the execution process of the computer-readable instructions 42 in the terminal device 4.
  • the terminal device 4 may be a computing device such as a desktop computer, a notebook, a palmtop computer and a cloud server.
  • the terminal device may include, but is not limited to, the processor 40 and the memory 41.
  • FIG. 4 is only an example of the terminal device 4 and does not constitute a limitation on the terminal device 4, and may include more or less components than the illustration, or a combination of certain components or different components.
  • the terminal device may further include an input and output device, a network access device, a bus, and the like.
  • the so-called processor 40 may be a central processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), Ready-made programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
  • the general-purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
  • the memory 41 may be an internal storage unit of the terminal device 4, such as a hard disk or a memory of the terminal device 4.
  • the memory 41 may also be an external storage device of the terminal device 4, for example, a plug-in hard disk equipped on the terminal device 4, a smart memory card (Smart) Media (SMC), and a secure digital (SD) Card, flash card (Flash Card, FC), etc.
  • the memory 41 may include both an internal storage unit of the terminal device 4 and an external storage device.
  • the memory 41 is used to store the computer-readable instructions and other programs and data required by the terminal device.
  • the memory 41 can also be used to temporarily store data that has been or will be output.
  • Non-volatile memory may include read-only memory (Read-Only Memory, ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • ROM Read-Only Memory
  • PROM programmable ROM
  • EPROM electrically programmable ROM
  • EEPROM electrically erasable programmable ROM
  • Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory.
  • RAM Random Access Memory
  • RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Artificial Intelligence (AREA)
  • Data Mining & Analysis (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Mathematical Physics (AREA)
  • Computing Systems (AREA)
  • Molecular Biology (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)
  • Geometry (AREA)

Abstract

一种证件图像提取方法、终端设备及计算机非易失性可读存储介质,包括:通过获取包含证件图像的原始图像(S101);所述原始图像通过摄像装置拍摄得到;根据原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像(S102);根据预先训练好的证件特征模型,从平衡图像中确定所述证件区域的位置(S103);所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;根据证件区域的位置,从所述平衡图像中提取出所述证件区域的图像(S104)。根据证件特征模型在原始图像中定位证件的位置,并提取出图像中的证件图像,提高了从原始图像中提取证件图像的精确性。

Description

证件图像提取方法及终端设备
本申请要求于2019年01月10日提交中国专利局、申请号为201910023382.2、发明名称为“证件图像提取方法及终端设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请属于计算机应用技术领域,尤其涉及一种证件图像提取方法、终端设备及计算机非易失性可读存储介质。
背景技术
机器视觉是让“物”有了看的功能,不仅具有信息采集功能,还能进行处理和识别等高级功能。另外机器视觉的设备成本低,最常使用的设备是摄像头。据统计,近几年各大城市公共摄像头和家庭、企业摄像头安装比率都大大的增加,家庭和企业摄像头安装比率也很高,随着摄像头的普及,今后在各城市的角落或者家庭和企业中都会大量使用摄像头。随着摄像头快速普及,机器视觉的技术的相关应用将更加快速发展。随着机器视觉领域的发展,证件照身份核验技术也将在这个浪潮中得到更广泛的应用。
现有技术中可以随时调取布置在城市各处的摄像头进行身份核验,想要在数以万计的人群中查找特定人的信息变得简单。但是由于很多外界环境的影响,得到的证件图像质量较差,而不能精确得到证件图像。
技术问题
本申请实施例提供了一种证件图像提取方法、终端设备及计算机非易失性可读存储介质,以解决现有技术中由于很多外界环境的影响,得到的证件图像质量较差、不精确的问题。
技术解决方案
本申请实施例的第一方面提供了一种证件图像提取方法,包括:
用于获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
用于根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
本申请实施例的第二方面提供了一种终端设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令 时实现以下步骤:
用于获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
用于根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
本申请实施例的第三方面提供了一种终端设备,包括:
获取单元,用于获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
处理单元,用于根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
确定单元,用于根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
提取单元,用于根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
本申请实施例的第四方面提供了一种计算机非易失性可读存储介质,所述计算机存储介质存储有计算机可读指令,所述计算机可读指令当被处理器执行时使所述处理器执行上述第一方面的方法。
有益效果
本申请实施例,通过获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;根据原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;根据预先训练好的证件特征模型,从平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;根据证件的位置,从所述平衡图像中提取出所述证件的图像。根据证件特征模型在原始图像中定位证件的位置,并提取出图像中的证件图像,提高了从原始图像中提取证件图像的精确性。
附图说明
图1是本申请实施例一提供的证件图像提取方法的流程图;
图2是本申请实施例二提供的证件图像提取方法的流程图;
图3是本申请实施例三提供的终端设备的示意图;
图4是本申请实施例四提供的终端设备的示意图。
本发明的实施方式
以下描述中,为了说明而不是为了限定,提出了诸如特定系统结构、技术之类的具体细节,以便透彻理解本申请实施例。然而,本领域的技术人员应当清楚,在没有这些具体细节的其它实施例中也可以实现本申请。在其它情况中,省略对众所周知的系统、装置、电路以及方法的详细说明,以免不必要的细节妨碍本申请的描述。
为了说明本申请所述的技术方案,下面通过具体实施例来进行说明。
参见图1,图1是本申请实施例一提供的证件图像提取方法的流程图。本实施例中证件图像提取方法的执行主体为终端。终端包括但不限于智能手机、平板电脑、可穿戴设备等移动终端,还可以是台式电脑等。如图所示的证件图像提取方法可以包括以下步骤:
S101:获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到。
随着社会实名制的发展,对证件信息的快速、准确的釆集成为一个越来越重要的话题,硬件性能的提高及数字图像处理技术的高速发展,大大促进了证件信息采集系统性能的提升。作为影响证件信息采集系统整体功能的证件图像处理部分,对系统的效果影响大,并且,根据系统的不同其相应的处理也有所不同。随着国家法律法规的日渐完善,社会对于公共安全的要求越来越高,故有关部门在社会民生的多个领域都推行实名制,如上网实名制、开户实名制、手机实名制等等。若个人信息的提取单纯靠人工录入及核对,必将导致低下的工作效率和较高的出错率,给业务双方带来严重不便。证件信息釆集系统通过射频识别技术和图像识别技术,可以实现对证件信息的自动提取、身份证、护照等证件资料的录入。信息获取手段的丰富和图像处理技术的发展使证件阅读仪的体积更小,信息提取的速度更快,信息的出错率更低。在提高公共安全和管理效率的同时,也给业务双方带来极大便利。此外,证件信息采集系统也有利于实名制应用的拓展。证件信息采集系统的发展使得火车、汽车、地铁等人流量较大的场合也有条件开展实名制,这将极大的保障铁路、公路和城市轨道交通的安全。
移动智能终端是指像计算机一样装有各种操作系统,但体积相对计算机来说比较小,便于携带,且拥有无线上网功能,用户可以根据自己的需求下载对应操作系统的各种应用。生活中比较常见的移动智能终端有智能手机、平板电脑、车载电脑、可穿戴移动设备等。智能手机是目前较为常用的移动智能终端,用户可以按照自己喜好或者需求安装第三方服务商提供的应用,游戏或者功能性程序等,满足用户对于智能终端功能上的需求。近年来,随着科技的不断发展,各类证件也不再是一本证书,而是类似身份证的卡片。随着证件的使用,证件信息的录入也成为一个重要问题。传统的信息录入方式是采用人工方式先填写相关表格中信息,再由内部工作人员按照表格内容把关键信息存入计算机,或者是,到指定地点进行证件的扫描上传。前一种方式虽然不限制信息录入的地点,但每一次信息的录入都需要耗费 大量的人力物力资源,并且容易出现错误的输入。后一种虽然在信息录入的效率和准确率上都有提高,但是使用地点却相对固定。移动智能终端的出现,使随时随地进行证件信息的录入成为可能。移动智能终端上的信息识别系统可以广泛的应用于服务性行业、交通系统、公安系统等需要对证件信息进行查验的部分,无需大量人员即可完成证件信息的采集查验,提高采集查验工作中证件信息识别的效率和准确率,具有广阔的应用前景。
在实际应用中,用户可以将通过移动终端拍摄的原始图像上传至服务器或者图像处理终端,图像处理终端在接收到原始图像之后,对该原始图像进行处理和识别。本方案中的应用场景可以是在用户证件图像获取并验证的网站中,向用户发送图像采集指令,用户通过自己的终端设备拍摄照片,例如身份证、护照等照片,通过移动终端中的应用软件或者网页将拍摄得到的图像发送执行主体,执行主体在获取到目标图像之后,对该原始图像进行处理和识别。
S102:根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像。
在实际应用中,表示图像最常用的颜色空间是红绿蓝(Red Green Blue,RGB)。真彩色图像三个颜色分量各有一个字节位表示,因此一个点的空间需要三个字节表示,一个1024×768的真彩图像需要1024×768×3=2.25MB。这样大的空间在早期计算机上是一笔很大开销,即使在一些内存空间相对小的环境中,比如手机,也显庞大。因此,用一个表格存放图像中的所有颜色,而实际图像数据不再是RGB数据,而是RGB数据在那个表格中的索引,为了控制索引大小,一般这个表格大小要求小于256个元素,即一个字节表示的范围,这个字节就可以表示这个图像中一个点的颜色。如果表格更小,一个点所用的索引位数就更小,这样1024×768真彩图像256色调色板仅需要1024×768×3=768.8KB。颜色量化过程往往要使用到两类调色版,一类是真彩色或者伪真彩色图像量化到调色板图像;另一类是调色板图像继续量化。随着计算机存储容量的不断提升,调色板图像逐渐淡出了个人计算机的舞台,但是在手机等一些特种设备中应用依然十分广泛,尤其是在手机游戏应用中。
随着计算机技术的不断发展,图形图像的处理已广泛应用于工业、农业、军事、医学、管理等各个领域。通过彩色扫描仪、摄像机等设备,可采集到自然界色彩斑斓的原始图像。而用计算机来显示时,由于显示设备所提供的能力和经济的原因,可表示的颜色数目总是有限的。另一方面,不同的计算机设备条件可显示的颜色数目往往不同,而同一幅图像我们希望在较低档次的机器设备条件下得到较好的再现。
在实际应用中,平衡是描述显示器中红、绿、蓝三基色混合生成后白色精确度的一项指标。白平衡是电视摄像领域一个非常重要的概念,通过它可以解决色彩还原和色调处理的 一系列问题。白平衡是随着电子影像再现色彩真实而产生的,在专业摄像领域白平衡应用的较早,现在家用电子产品中也广泛地使用,然而技术的发展使得白平衡调整变得越来越简单容易,但许多使用者还不甚了解白平衡的工作原理,理解上存在诸多误区。它是实现摄像机图像能精确反映被摄物的色彩状况。在本实施例中,可以通过手动白平衡和自动白平衡等方式对原始图像进行白平衡处理,得到平衡图像。
S103:根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到。
在对证件字符进行识别时,方法主要有隐马尔可夫模型、神经网络、支持向量机和模板匹配。所有使用隐马尔可夫模型的方法都需要进行预处理和根据己有知识设定参数。这个方法通过复杂的预处理和参数化取得较高的识别率,也可以通过多层感知神经网络来训练证件特征模型,神经网络使用后向反馈的方法来进行训练这种网络必须训练很多次才能获得比较好的结果,这个过程是比较耗时的,而且隐藏层的层数和隐藏层神经元的个数都必须通过实验的方法来获得。可选的,神经网络的包含24个输入层神经元、15个隐藏层神经元、36个输出层神经元来识别平衡图像中的证件。
在人中脑中存在着无数的神经元,对于这些神经元来说,存在着千丝万缕的联系,经过一个组织之后构成一个紧密的神经网络结构,这个神经网络结构就可以实现人脑的复杂的计算和功能。在这里的神经网络主要是研究这些神经元的连接方式和组织结构。对于神经网络来说,可以划分为二种,一个是层状的,另一个是网状的。对于第一种来说,神经元之间是一个层次的排列的方式,对于每一层来说,这些神经元是并列排列,形成一个紧密的机构,对于层与层之间来说,通过神经元进行连接,但对于每一个层内部的神经元来说则是不能进行连接,第二种的神经网络结构来说,每一个的神经元则是可以进行互联。
需要说明的是,对于神经网络来说。需要经过一定的训练之后。学习到神经网络的处理的规则和方法。并且通过这些方法进行问题的处理和解决。对于前向多层网络的结构来说。具体有如下的几个步骤实现,首先需要对于前向多层网络提供一个训练的例子。在这个例子中包括了输入和输出的模式;对于上面的设计的训练自理来说。对于输入和输出允许存在着一定的误差;对于前向多层网络的输出需要进行改变。改变输出以便使得最后的输出能够得到一个比较好的输出。满足在误差范围内。
S104:根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
在确定了平衡图像中的证件图像的位置之后,根据证件图像的位置,从平衡图像中提取该证件图像。具体的,其提取的方式可以是直接从平衡图像中裁剪的方式,将证件图像提取出来,还可以是删除掉除去证件图像之外的图像区域,保留证件的图像的方法,此处不做 限定。
除此之外,还可以通过基于边缘和梯度的方法,检测证件图像的图像边缘,基于该图像边缘,将证件图像提取出来。基于边缘的证件图像方法认为自然场景中的证件图像与背景边缘存在较大差异性,该方法对字符进行边缘检测,从而通过边缘信息定位证件图像。可选的,可以通过Sobel算子、Robert算子、Laplace算子确定证件图像的图像边缘。其中,Sobel算子通过判断证件图像中某个像素点的梯度是否大于阈值来确定该点是否为边缘点,Robert算子适合文字与背景区别较大的图像,而且检测后获得的边缘较粗,Laplace算子对噪声非常敏感,容易产生双边效果,不直接用于检测边。也可以采用基于连通域的证件图像定位方法,将原始图像转变为二值图像,减少噪声的影响,使用形态学腐蚀膨胀算法将证件图像区域连通,利用证件图像与白色背景的区分度分割图像,再通过证件图像的各方面特征排除非证件图像连通域,从而得到证件图像区域基于连通域的证件图像定位方法,证件图像定位速度较快,可以提高证件图像及其证件图像中的文字的识别效率。
上述方案,通过获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。根据证件特征模型在原始图像中定位证件的位置,并提取出图像中的证件图像,提高了从原始图像中提取证件图像的精确性。
参见图2,图2是本申请实施例二提供的证件图像提取方法的流程图。本实施例中证件图像提取方法的执行主体为终端。终端包括但不限于智能手机、平板电脑、可穿戴设备等移动终端,还可以是台式电脑等。如图所示的证件图像提取方法可以包括以下步骤:
S201:采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像。
在对原始图像进行识别和处理之前,我们需要先训练出证件特征模型来进行证件图像识别。因此本方案中可以先根据历史证件图像训练出证件特征模型,以从原始图像中提取出证件图像来。本方案中训练证件特征模型的数据可以是历史证件图像,历史证件图像包括在进行正式的证件图像提取之前,获取到的历史图像。
在实际应用中,获取到的历史证件图像可能存在不符合要求的图像,考虑到这种情况,本方案中对获取到的历史图像按照预设的目标图像要求进行筛选,得到目标图像。其中,预设的目标图像要求可以是图像的像素、大小或者拍摄时间要求,除此之外,还可以是检测图 像的完整度,证件图像的类型要求等。这些要求可以是执行者来确定,此处不做限定,在确定了目标图像要求之后,将获取到的历史证件图像与目标图像要求进行匹配,确定匹配度大于匹配度阈值的历史证件图像作为目标图像。
通过获取各种摄像装置得到的证件图像,用于作为进行神经网络训练的初始样本。在神经网络的学习训练过程中,训练集的选取将直接影响网络学习训练的时间、权值矩阵与学习训练效果等。本方案中选择图像边缘比较明显,且边缘分布在图像的大部分区域的图像作为训练即集的初始样本,图像边缘比较清晰,颗粒边缘分布在整副图像中,而且其纹理特征比较丰富,使神经网络得到很好训练,络权值等网络信息会记住更多的边缘信息,可以比较好的检测图像。
S202:根据预设的证件图像模板对所述目标图像进行像素识别,在所述目标图像中确定至少一个中心像素点。
在确定图像样本中的中心像素点时,可以通过图像识别的方式确定图像样本具有突出代表性的像素点作为中心像素点。示例性的,当处理的原始图像是用户拍摄的身份证图像时,根据已知的身份证图像中的头像的大小和文字位置,可以确定头像的四个角为中心像素点,也可以确定某些文字为中心像素点,例如身份证中的第一个字或者每一行的第一个字等;进一步的,还可以预先设定所获取的图像类型,例如身份证照片、房产证照片等,并确定每种类型的图像模板以及该模板中的每个图像元素的位置或者与证件边框的距离等信息,通过这些信息进行识别,以精确确定图像样本中的中心像素点,通过中心像素点和以其为中心的周围的像素点进行学习训练,以从原始图像中定位证件的位置。
需要说明的是,中心像素点周围的像素点的个数可以是至少两个,优选的,可以确定中心像素点周围的8个像素点来进行学习训练,更加清楚的确定图像中每个像素的情况。
S203:设置训练模型的初始参数,根据所述初始参数、每个所述中心像素点和所述中心像素点周围的像素点的像素值进行学习训练,得到基于神经网络的证件特征模型。
对于任何一个神经网络模型,其应用过程中的学习训练都是关键的环节,只有通过学习训练,网络才能够具有联想、记忆和预测的能力。通常某些参数的确定对于学习训练过程至关重要。网络的初始参数包括网络初始结构、连接的权值、阈值及学习率等,不同的设置都会在一定程度上影响网络的收敛速度。初始参数的选择非常重要却也非常困难。除去必要的技术处理,网络构建主要靠的是观察与经验。
在训练证件特征模型时,首先确定模型初始值,网络的初始权值和阈值一般都是从[-1,1]或[0,1]随机选取,某些改进算法会对区间做适当的更改。其次对向量进行归一化处理,在学习训练过程中,结点输入不宜过大,过小的权值调节将不利于网络学习训练。训练图像 是基于灰度的,因此图像矩阵均为介于[0,255]的整形数值,并且特征向量维数比较高,为了提高网络训练速度,将会对特征向量统一做归一化处理。把特征向量看作是行向量,表示为:
X=(x 0,x 1,…,x 9);其中,x 0,x 1,…,x 9分别用于表示中心像素点及其周围像素点的像素值。
本实施例中的中心像素点周围的像素点的数量可以是至少两个,优选的,可以是8个,以更加精确的说明中心像素点的周围像素点的情况,8位灰度图像的灰度值范围是[0,255],因此,实际处理中归一化公式为:
Figure PCTCN2019118133-appb-000001
其中,x 0用于表示中心像素点的像素值。
由于处理的对象是图像,图像样本集比较庞大,故采用对图像进行分块操作的思想。神经网络中每次输入一个图像样本,可以通过确定一个或者至少两个中心像素点,对这个中心像素点周围的模板像素,即包括以其为中心的周围8个像素,进行学习训练,把这些像素的灰度值以从上至下、从左至右的顺序依次送入输入层。输出层提供的期望输出像素的灰度值与实际输出层的输出像素灰度值之间存在一定的误差,误差沿反向传播,进而使每个神经元的阈值以及神经元之间的连接权值发生改变,因此网络可以有效记忆更多的边缘信息。反复进行上述过程直至误差缩小到规定的范围内,或者训练次数达到目标次数,训练任务完成。训练的要求规定可以随时停止网络的训练,同时为了方便日后利用神经网络进行检测,训练出来的权值和阈值全部存储在后端的数据库中,最后保存训练好的网络。
S204:获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到。
在本实施例中S204与图1对应的实施例中S101的实现方式完全相同,具体可参考图1对应的实施例中的S101的相关描述,在此不再赘述。
S205:根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像。
在获取到原始图像之后,根据原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像。
进一步的,步骤S205可以具体包括步骤S2051-S2052:
S2051:根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差。
在对图像进行形态学处理或匹配、识别等处理之前,对图像进行的滤除干扰信息、增强有效信息等处理称为图像的预处理。对图像进行预处理的主要目的是消除图像中的干扰或无关的信息,恢复有用的真实信息,增强有关信息的可检测性和最大限度地简化数据,从而 改进特征抽取、图像分割、匹配和识别的可靠性。对数字彩色图像进行的预处理一般是亮度、色彩的复原与增强。鉴于对各种预处理的比较及试验,发现白平衡处理对系统最后的分割结果有比较大的影响,而其它的预处理则影响不大。故本方案中预处理主要就是白平衡处理。
不同的光源具有不同的光谱成分和分布,这在色度学上称之为色温。一个白色的物体,在低色温的光线照射下会偏红,而在高色温的光线照射下会偏蓝。进行拍摄时,环境光源色温会对图像产生影响,使其不可避免地出现色彩上的偏差。为了尽可能减少外来光照对目标颜色造成的影响,在不同的色温条件下均能还原出被摄目标本来的色彩,需要进行色彩校正,以达成正确的色彩平衡。
当图像中的R、G、B三种颜色相等时,其色差为0,表现为白色。在图像处理中,一般采用YBR色彩模型来计算色差。YBR色彩系统与RGB色彩系统的对应关系如下:
Figure PCTCN2019118133-appb-000002
在Y足够大、B和R足够小的空间里定义了一个区域,并将该区域中的所有像素看作是白色的,可以参与色差的计算。然后,用白色像素的平均色差来代表整个图像的色差,以取得较好的精度。根据系统的特点,我们给出了以下约束条件:Y-|B|-|R|>180;满足该约束条件的像素都看作是白色的,得到白色像素点的平均亮度及R,G,B分量的平均值R avg、G avg、B avg
S2052:根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量。
在实际应用中色彩增益用于表示图像的鲜活程度,增益量不外乎增加色彩对比度,使颜色更鲜艳更饱和,造成视觉较大的冲击力,另一方面有一定的锐化效果,使得边缘线条更加分明清晰。可以通过色彩增益对图像进行自动调节对比度、色彩饱和度等功能类似。数码相机的这种技术可以使得照片看起来更清晰,更抢眼。
根据上一步骤计算得到的平均色差,我们可以得到白平衡各分量的增益量为:
Figure PCTCN2019118133-appb-000003
S2053:根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
根据上一步骤得到的增益量,本方案对整个图像的每个像素进行色温校正,具体计算 公式如下:
Figure PCTCN2019118133-appb-000004
可选的,还可以进行图像的增强消除图像中的噪声或降低图像中的噪声,增强图像中的对比度等增强对文本区域的定位。图像的水平校正则是将原图像转换为文本水平分布的图像,增强文本区域定位准确度。图像增强方法可以为高斯模糊和锐化处理,图像高斯模糊处理是模糊细节,降噪的常用方法,高斯模糊处理将和点的8连通区域按照一定权重加权相加,将其中值作为点的像素值。使用高斯模糊平滑处理可以将图像中很多噪声平滑,将图像中目标图像的轮廓凸显出来。高斯模糊平滑处理只能适用于背景复杂,但是图像中目标轮廓很明显的图像,平滑处理可以平滑图像细节,对于噪声有平滑的同时,也会将一些不是很明显的轮廓细节也平滑掉。
除此之外,还可以对原始图像进行图像的平滑和滤波,对一些在图像的生成的过程中造成的图像的品质的下降采取一些措施。可以使得图像的质量得到改善。具体来说就是对图像丢失的部分信息进行有针对性的补偿。另外的一个方法就是对图像进行处理将图像的某一个部分的图像信息进行突出。对一些不是很重要的图像信息进行进一步的减少。在证件的图像处理中。经常需要使用证件采集的工具获取到证件的图像信息。在这个过程中经常会产生一些噪声。因此就需要设法降低噪声。通过这种方法就可以提高图像的品质。可以对产生的噪声进行干扰。获取到比较好的图像信息。增强重要的图像的信息。这个图像的预处理的技术就是图像的平滑。对于图像的平滑技术来说。主要是通过如下的二个方法和性能要求来实现图像的增强效果。首先是对于图像的线条和边缘轮廓等重要信息需要进行保留。不能随意的破坏。其次是对于图像需要使得图像的画面清晰和图像效果。
S206:根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到。
通过多层感知神经网络来训练证件特征模型,神经网络使用后向反馈的方法来进行训练这种网络必须训练很多次才能获得比较好的结果。这个过程是比较耗时的。而且隐藏层的层数和隐藏层神经元的个数都必须通过实验的方法来获得。可选的,神经网络的包含24个输入层神经元、15个隐藏层神经元、36个输出层神经元来识别平衡图像中的证件。
进一步的,步骤S206中可以具体包括步骤S2061:
S2061:若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
在根据证件特征模型检测原始图像中的证件图像的位置时,很可能出现检测结果与其实际结果出现出入的情况,在这种情况下,我们可以对证件特征模型的参数进行调整,以使之后的检测结果能够更加精确。具体的实施方式为:
确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
Figure PCTCN2019118133-appb-000005
其中:w ij(k)用于表示第k次训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
当神经网络实际输出值与期望输出值不统一时,求取误差信号,并将该信号从输出端反向传播,同时在传播过程中不断修正加权系数,以使误差函数最小,通常网络误差采用均方差,对权值进行修改。调整公式为:
Figure PCTCN2019118133-appb-000006
式中w ij(k)用于表示第k次训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值,
Figure PCTCN2019118133-appb-000007
表示第k次时的负梯度。
S207:根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
在本实施例中S207与图1对应的实施例中S105的实现方式完全相同,具体可参考图1对应的实施例中的S105的相关描述,在此不再赘述。
上述方案,通过采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;根据预设的证件图像模板对所述目标图像进行像素识别,在所述目标图像中确定至少一个中心像素点;设置训练模型的初始参数,根据所述初始参数、每个所述中心像素点和所述中心像素点周围的像素点的像素值进行学习训练,得到基于神经网络的证件特征模型。获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。 通过对获取到的待提取图像的原始图像进行预处理,根据证件特征模型对预处理之后的图像定位证件的位置,并提取出图像中的证件图像,提高了从原始图像中提取证件图像的精确性。
参见图3,图3是本申请实施例三提供的一种终端设备的示意图。终端设备包括的各单元用于执行图1~图2对应的实施例中的各步骤。具体请参阅图1~图2各自对应的实施例中的相关描述。为了便于说明,仅示出了与本实施例相关的部分。本实施例的终端设备300包括:
获取单元301,用于获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
处理单元302,用于根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
确定单元303,用于根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
提取单元304,用于根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
进一步的,所述终端设备还可以包括:
筛选单元,用于采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;
识别单元,用于根据预设的证件图像模板对所述目标图像进行像素识别,在所述目标图像中确定至少一个中心像素点;
训练单元,用于设置训练模型的初始参数,根据所述初始参数、每个所述中心像素点和所述中心像素点周围的像素点的像素值进行学习训练,得到基于神经网络的证件特征模型。
进一步的,所述确定单元303可以包括:
修正单元,用于若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
进一步的,所述修正单元可以包括:
距离计算单元,用于确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
参数校正单元,用于若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
Figure PCTCN2019118133-appb-000008
其中:w ij(k)用于表示第k次 训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
进一步的,所述处理单元302可以包括:
色差估算单元,用于根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差;
增益计算单元,根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量;
平衡处理单元,用于根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
上述方案,通过获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;根据原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;根据预先训练好的证件特征模型,从平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;根据证件的位置,从所述平衡图像中提取出所述证件的图像。根据证件特征模型在原始图像中定位证件的位置,并提取出图像中的证件图像,提高了从原始图像中提取证件图像的精确性。
图4是本申请实施例四提供的终端设备的示意图。如图4所示,该实施例的终端设备4包括:处理器40、存储器41以及存储在所述存储器41中并可在所述处理器40上运行的计算机可读指令42。所述处理器40执行所述计算机可读指令42时实现上述各个证件图像提取方法实施例中的步骤,例如图1所示的步骤101至104。或者,所述处理器40执行所述计算机可读指令42时实现上述各装置实施例中各模块/单元的功能,例如图3所示单元301至304的功能。
示例性的,所述计算机可读指令42可以被分割成一个或多个模块/单元,所述一个或者多个模块/单元被存储在所述存储器41中,并由所述处理器40执行,以完成本申请。所述一个或多个模块/单元可以是能够完成特定功能的一系列计算机可读指令的指令段,该指令段用于描述所述计算机可读指令42在所述终端设备4中的执行过程。
所述终端设备4可以是桌上型计算机、笔记本、掌上电脑及云端服务器等计算设备。所述终端设备可包括,但不仅限于,处理器40、存储器41。本领域技术人员可以理解,图4仅仅是终端设备4的示例,并不构成对终端设备4的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件,例如所述终端设备还可以包括输入输出设备、网络接入设备、总线等。
所称处理器40可以是中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
所述存储器41可以是所述终端设备4的内部存储单元,例如终端设备4的硬盘或内存。所述存储器41也可以是所述终端设备4的外部存储设备,例如所述终端设备4上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card,FC)等。进一步地,所述存储器41还可以既包括所述终端设备4的内部存储单元也包括外部存储设备。所述存储器41用于存储所述计算机可读指令以及所述终端设备所需的其他程序和数据。所述存储器41还可用于暂时地存储已经输出或者将要输出的数据。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,所述的计算机可读指令可存储于一计算机非易失性可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(Read-Only Memory,ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(Random Access Memory,RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
以上所述实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围,均应包含在本申请的保护范围之内。

Claims (20)

  1. 一种证件图像提取方法,其特征在于,包括:
    获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
    根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
    根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
    根据所述证件图像的位置,从所述平衡图像中提取出所述证件图像。
  2. 如权利要求1所述的证件图像提取方法,其特征在于,所述根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件的位置之前,还包括:
    采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;
    根据预设的证件图像模板识别所述目标图像中的像素,在所述目标图像中确定至少一个像素作为中心像素点;
    设置训练模型的初始权值,根据所述初始权值、每个所述中心像素点和所述中心像素点周围的像素点的像素值计算证件图像的输出位置,并根据所述输出位置和预设的期望位置之间的差值调整所述初始权值得到目标权值,根据所述目标权值确定基于神经网络的证件特征模型。
  3. 如权利要求2所述的证件图像提取方法,其特征在于,所述根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件的位置,包括:
    若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
  4. 如权利要求3所述的证件图像提取方法,其特征在于,所述若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数,包括:
    确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
    若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
    Figure PCTCN2019118133-appb-100001
    其中:w ij(k)用于表示第k次训练时的权值;w ij(k+1) 用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
  5. 如权利要求1-4任一项所述的证件图像提取方法,其特征在于,所述根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像,包括:
    根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差;
    根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量;
    根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
  6. 一种终端设备,其特征在于,包括:
    获取单元,用于获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
    处理单元,用于根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
    确定单元,用于根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
    提取单元,用于根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
  7. 根据权利要求6所述的终端设备,其特征在于,所述终端设备还包括:
    筛选单元,用于采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;
    识别单元,用于根据预设的证件图像模板对所述目标图像进行像素识别,在所述目标图像中确定至少一个中心像素点;
    训练单元,用于设置训练模型的初始参数,根据所述初始参数、每个所述中心像素点和所述中心像素点周围的像素点的像素值进行学习训练,得到基于神经网络的证件特征模型。
  8. 根据权利要求7所述的终端设备,其特征在于,所述确定单元包括:
    修正单元,用于若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
  9. 根据权利要求8所述的终端设备,其特征在于,所述修正单元可以包括:
    距离计算单元,用于确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
    参数校正单元,用于若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
    Figure PCTCN2019118133-appb-100002
    其中:w ij(k)用于表示第k次训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
  10. 根据权利要求6-9任一项所述的终端设备,其特征在于,所述处理单元包括:
    色差估算单元,用于根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差;
    增益计算单元,根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量;
    平衡处理单元,用于根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
  11. 一种终端设备,其特征在于,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现如下步骤:
    获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
    根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
    根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
    根据所述证件图像的位置,从所述平衡图像中提取出所述证件的图像。
  12. 根据权利要求11所述的终端设备,其特征在于,所述根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件的位置之前,还包括:
    采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;
    根据预设的证件图像模板识别所述目标图像中的像素,在所述目标图像中确定至少一个像素作为中心像素点;
    设置训练模型的初始权值,根据所述初始权值、每个所述中心像素点和所述中心像素点周围的像素点的像素值计算证件图像的输出位置,并根据所述输出位置和预设的期望位置之间的差值调整所述初始权值得到目标权值,根据所述目标权值确定基于神经网络的证件特征模型。
  13. 根据权利要求12所述的终端设备,其特征在于,所述根据预先训练好的证件特征 模型,从所述平衡图像中确定所述证件的位置,包括:
    若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
  14. 根据权利要求13所述的终端设备,其特征在于,所述若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数,包括:
    确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
    若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
    Figure PCTCN2019118133-appb-100003
    其中:w ij(k)用于表示第k次训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
  15. 根据权利要求11-14任一项所述的终端设备,其特征在于,所述根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像,包括:
    根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差;
    根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量;
    根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
  16. 一种计算机非易失性可读存储介质,所述计算机非易失性可读存储介质存储有计算机可读指令,其特征在于,所述计算机可读指令被处理器执行时实现如下步骤:
    获取包含证件图像的原始图像;所述原始图像通过摄像装置拍摄得到;
    根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像;
    根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件图像的位置;所述证件特征模型为基于历史证件图像、证件图像模型以及预设的初始权值进行训练得到;
    根据所述证件图像的位置,从所述平衡图像中提取出所述证件图像。
  17. 根据权利要求16所述的计算机非易失性可读存储介质,其特征在于,所述根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件的位置之前,还包括:
    采集历史证件图像,并对所述历史证件图像按照预设的目标图像要求进行筛选,得到目标图像;
    根据预设的证件图像模板识别所述目标图像中的像素,在所述目标图像中确定至少一个像素作为中心像素点;
    设置训练模型的初始权值,根据所述初始权值、每个所述中心像素点和所述中心像素点周围的像素点的像素值计算证件图像的输出位置,并根据所述输出位置和预设的期望位置之间的差值调整所述初始权值得到目标权值,根据所述目标权值确定基于神经网络的证件特征模型。
  18. 根据权利要求17所述的计算机非易失性可读存储介质,其特征在于,所述根据预先训练好的证件特征模型,从所述平衡图像中确定所述证件的位置,包括:
    若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数。
  19. 根据权利要求18所述的计算机非易失性可读存储介质,其特征在于,所述若根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值大于或者等于预设的差值阈值,则修正所述证件特征模型的初始参数,包括:
    确定根据所述证件特征模型得到的所述证件的位置与所述证件的实际位置之间的距离差值;
    若所述距离差值大于或者等于所述差值阈值,则根据如下公式校正所述证件特征模型的初始参数:
    Figure PCTCN2019118133-appb-100004
    其中:w ij(k)用于表示第k次训练时的权值;w ij(k+1)用于表示第k+1次训练时的权值;η用于表示学习速率且η>0;E(k)用于表示前k次训练得到的证件图像的位置的期望值。
  20. 根据权利要求16-19任一项所述的计算机非易失性可读存储介质,其特征在于,所述根据所述原始图像中的每个像素点在红、绿、蓝三个颜色分量中的分量值,对所述原始图像进行白平衡处理,得到平衡图像,包括:
    根据所述原始图像中的每个所述像素点在红、绿、蓝三个颜色分量中的分量值,估算所述原始图像中每个所述像素点的平均色差;
    根据每个所述像素点的所述平均色差,计算每个所述像素点在红、绿、蓝三个颜色分量中增益量;
    根据所述增益量,校正所述原始图像中的每个所述像素点的色温,得到所述平衡图像。
PCT/CN2019/118133 2019-01-10 2019-11-13 证件图像提取方法及终端设备 Ceased WO2020143316A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
SG11202100270VA SG11202100270VA (en) 2019-01-10 2019-11-13 Certificate image extraction method and terminal device
JP2021500946A JP2021531571A (ja) 2019-01-10 2019-11-13 証明書画像抽出方法及び端末機器
US17/167,075 US11790499B2 (en) 2019-01-10 2021-02-03 Certificate image extraction method and terminal device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910023382.2 2019-01-10
CN201910023382.2A CN109871845B (zh) 2019-01-10 2019-01-10 证件图像提取方法及终端设备

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/167,075 Continuation-In-Part US11790499B2 (en) 2019-01-10 2021-02-03 Certificate image extraction method and terminal device

Publications (1)

Publication Number Publication Date
WO2020143316A1 true WO2020143316A1 (zh) 2020-07-16

Family

ID=66917644

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/118133 Ceased WO2020143316A1 (zh) 2019-01-10 2019-11-13 证件图像提取方法及终端设备

Country Status (5)

Country Link
US (1) US11790499B2 (zh)
JP (1) JP2021531571A (zh)
CN (1) CN109871845B (zh)
SG (1) SG11202100270VA (zh)
WO (1) WO2020143316A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112132812A (zh) * 2020-09-24 2020-12-25 平安科技(深圳)有限公司 证件校验方法、装置、电子设备及介质
CN116016883A (zh) * 2022-12-26 2023-04-25 上海闻泰信息技术有限公司 图像的白平衡处理方法、装置、电子设备及存储介质

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109871845B (zh) * 2019-01-10 2023-10-31 平安科技(深圳)有限公司 证件图像提取方法及终端设备
CN110929725B (zh) * 2019-12-06 2023-08-29 深圳市碧海扬帆科技有限公司 证件分类方法、装置及计算机可读存储介质
CN111310746B (zh) * 2020-01-15 2024-03-01 支付宝实验室(新加坡)有限公司 文本行检测方法、模型训练方法、装置、服务器及介质
SG10202001559WA (en) * 2020-02-21 2021-03-30 Alipay Labs Singapore Pte Ltd Method and system for determining authenticity of an official document
CN113570508A (zh) * 2020-04-29 2021-10-29 上海耕岩智能科技有限公司 图像修复方法及装置、存储介质、终端
CN112333356B (zh) * 2020-10-09 2022-09-20 支付宝实验室(新加坡)有限公司 一种证件图像采集方法、装置和设备
CN114140811A (zh) * 2021-11-04 2022-03-04 北京中交兴路信息科技有限公司 一种证件样本生成方法、装置、电子设备和存储介质
CN114399454B (zh) * 2022-01-18 2024-10-18 平安科技(深圳)有限公司 图像处理方法、装置、电子设备及存储介质
CN115929738B (zh) * 2022-12-28 2025-10-24 中联重科股份有限公司 用于液压系统的方法、控制模型的训练方法以及控制方法
CN116189197B (zh) * 2022-12-30 2026-01-30 北京优利绚彩科技发展有限公司 被拍证件边缘识别方法、装置、高拍仪自助机及控制系统
CN117173545B (zh) * 2023-11-03 2024-01-30 天逸财金科技服务(武汉)有限公司 一种基于计算机图形学的证照原件识别方法
WO2025192793A1 (en) * 2024-03-12 2025-09-18 Samsung Electronics Co., Ltd. A neuro-template based method and system for image correction
CN119945687B (zh) * 2025-04-03 2025-06-13 四川万网鑫成信息科技有限公司 一种批量u盾控制方法、系统及装置

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101038686A (zh) * 2007-01-10 2007-09-19 北京航空航天大学 一种基于信息融合的机读旅行证件识别方法
US20120250987A1 (en) * 2011-03-31 2012-10-04 Sony Corporation System and method for effectively performing an image identification procedure
US20130108123A1 (en) * 2011-11-01 2013-05-02 Samsung Electronics Co., Ltd. Face recognition apparatus and method for controlling the same
CN105825243A (zh) * 2015-01-07 2016-08-03 阿里巴巴集团控股有限公司 证件图像检测方法及设备
CN107844748A (zh) * 2017-10-17 2018-03-27 平安科技(深圳)有限公司 身份验证方法、装置、存储介质和计算机设备
CN109871845A (zh) * 2019-01-10 2019-06-11 平安科技(深圳)有限公司 证件图像提取方法及终端设备

Family Cites Families (18)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3213230B2 (ja) * 1996-02-23 2001-10-02 株式会社ピーエフユー 画像データ読取装置
JP3393168B2 (ja) * 1996-07-26 2003-04-07 シャープ株式会社 画像入力装置
JP2006129442A (ja) * 2004-09-30 2006-05-18 Fuji Photo Film Co Ltd 画像補正装置および方法,ならびに画像補正プログラム
JP4677226B2 (ja) * 2004-12-17 2011-04-27 キヤノン株式会社 画像処理装置及び方法
JP4227135B2 (ja) * 2005-11-25 2009-02-18 株式会社東芝 光学的文字読取装置及びカラーバランス調整方法
JP2009239323A (ja) * 2006-07-27 2009-10-15 Panasonic Corp 映像信号処理装置
CN102147860A (zh) * 2011-05-16 2011-08-10 杭州华三通信技术有限公司 一种基于白平衡的车牌识别方法和装置
JP4998637B1 (ja) * 2011-06-07 2012-08-15 オムロン株式会社 画像処理装置、情報生成装置、画像処理方法、情報生成方法、制御プログラムおよび記録媒体
US8693731B2 (en) * 2012-01-17 2014-04-08 Leap Motion, Inc. Enhanced contrast for object detection and characterization by optical imaging
JP2013197848A (ja) * 2012-03-19 2013-09-30 ▲うぇい▼強科技股▲ふん▼有限公司 スキャナのためのオートホワイトバランス調整方法
US10540564B2 (en) * 2014-06-27 2020-01-21 Blinker, Inc. Method and apparatus for identifying vehicle information from an image
WO2016207875A1 (en) * 2015-06-22 2016-12-29 Photomyne Ltd. System and method for detecting objects in an image
CN105120167B (zh) * 2015-08-31 2018-11-06 广州市幸福网络技术有限公司 一种证照相机及证照拍摄方法
JP2017059207A (ja) * 2015-09-18 2017-03-23 パナソニック インテレクチュアル プロパティ コーポレーション オブ アメリカPanasonic Intellectual Property Corporation of America 画像認識方法
CN106156767A (zh) * 2016-03-02 2016-11-23 平安科技(深圳)有限公司 行驶证有效期自动提取方法、服务器及终端
WO2018173108A1 (ja) * 2017-03-21 2018-09-27 富士通株式会社 関節位置推定装置、関節位置推定方法及び関節位置推定プログラム
US10831821B2 (en) * 2018-09-21 2020-11-10 International Business Machines Corporation Cognitive adaptive real-time pictorial summary scenes
KR102891571B1 (ko) * 2018-12-19 2025-11-26 삼성전자주식회사 중첩된 비트 표현 기반의 뉴럴 네트워크 처리 방법 및 장치

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101038686A (zh) * 2007-01-10 2007-09-19 北京航空航天大学 一种基于信息融合的机读旅行证件识别方法
US20120250987A1 (en) * 2011-03-31 2012-10-04 Sony Corporation System and method for effectively performing an image identification procedure
US20130108123A1 (en) * 2011-11-01 2013-05-02 Samsung Electronics Co., Ltd. Face recognition apparatus and method for controlling the same
CN105825243A (zh) * 2015-01-07 2016-08-03 阿里巴巴集团控股有限公司 证件图像检测方法及设备
CN107844748A (zh) * 2017-10-17 2018-03-27 平安科技(深圳)有限公司 身份验证方法、装置、存储介质和计算机设备
CN109871845A (zh) * 2019-01-10 2019-06-11 平安科技(深圳)有限公司 证件图像提取方法及终端设备

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112132812A (zh) * 2020-09-24 2020-12-25 平安科技(深圳)有限公司 证件校验方法、装置、电子设备及介质
CN112132812B (zh) * 2020-09-24 2023-06-30 平安科技(深圳)有限公司 证件校验方法、装置、电子设备及介质
CN116016883A (zh) * 2022-12-26 2023-04-25 上海闻泰信息技术有限公司 图像的白平衡处理方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
JP2021531571A (ja) 2021-11-18
CN109871845B (zh) 2023-10-31
SG11202100270VA (en) 2021-02-25
US20210166015A1 (en) 2021-06-03
CN109871845A (zh) 2019-06-11
US11790499B2 (en) 2023-10-17

Similar Documents

Publication Publication Date Title
WO2020143316A1 (zh) 证件图像提取方法及终端设备
US11354797B2 (en) Method, device, and system for testing an image
CN112651333B (zh) 静默活体检测方法、装置、终端设备和存储介质
CN111091075B (zh) 人脸识别方法、装置、电子设备及存储介质
CN112784900B (zh) 图像目标对比方法、装置、计算机设备及可读存储介质
CN112101359B (zh) 文本公式的定位方法、模型训练方法及相关装置
CN108875602A (zh) 监控环境下基于深度学习的人脸识别方法
CN111209858B (zh) 一种基于深度卷积神经网络的实时车牌检测方法
CN108090511B (zh) 图像分类方法、装置、电子设备及可读存储介质
CN111898544B (zh) 文字图像匹配方法、装置和设备及计算机存储介质
CN113033519B (zh) 活体检测方法、估算网络处理方法、装置和计算机设备
CN113537211B (zh) 一种基于非对称iou的深度学习车牌框定位方法
WO2022006829A1 (zh) 一种票据图像识别方法、系统、电子设备和存储介质
CN110298829A (zh) 一种舌诊方法、装置、系统、计算机设备和存储介质
CN111931783A (zh) 一种训练样本生成方法、机读码识别方法及装置
CN111222433A (zh) 自动人脸稽核方法、系统、设备及可读存储介质
CN114663951A (zh) 低照度人脸检测方法、装置、计算机设备及存储介质
CN115937537A (zh) 一种目标图像的智能识别方法、装置、设备及存储介质
CN116958919A (zh) 目标检测方法、装置、计算机可读介质及电子设备
CN115984712A (zh) 基于多尺度特征的遥感图像小目标检测方法及系统
CN115035313B (zh) 黑颈鹤识别方法、装置、设备及存储介质
CN116824419B (zh) 一种着装特征识别方法、识别模型的训练方法及装置
CN112686847B (zh) 身份证图像拍摄质量评价方法、装置、计算机设备和介质
CN117649358B (zh) 图像处理方法、装置、设备及存储介质
CN116883385B (zh) 一种红外人脸图像质量评价的方法及装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19908609

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021500946

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19908609

Country of ref document: EP

Kind code of ref document: A1