WO2020134229A1 - 图像处理方法、装置、电子设备及计算机可读存储介质 - Google Patents

图像处理方法、装置、电子设备及计算机可读存储介质 Download PDF

Info

Publication number
WO2020134229A1
WO2020134229A1 PCT/CN2019/107362 CN2019107362W WO2020134229A1 WO 2020134229 A1 WO2020134229 A1 WO 2020134229A1 CN 2019107362 W CN2019107362 W CN 2019107362W WO 2020134229 A1 WO2020134229 A1 WO 2020134229A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
target area
area image
target
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/107362
Other languages
English (en)
French (fr)
Inventor
杨武魁
吴立威
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to SG11202010402VA priority Critical patent/SG11202010402VA/en
Priority to US17/048,823 priority patent/US20210150745A1/en
Priority to JP2020556853A priority patent/JP7113910B2/ja
Publication of WO2020134229A1 publication Critical patent/WO2020134229A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/50Depth or shape recovery
    • G06T7/55Depth or shape recovery from multiple images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/246Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/246Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
    • G06T7/248Analysis of motion using feature-based methods, e.g. the tracking of corners or segments involving reference images or patches
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/269Analysis of motion using gradient-based methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/50Depth or shape recovery
    • G06T7/55Depth or shape recovery from multiple images
    • G06T7/593Depth or shape recovery from multiple images from stereo images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • G06V40/165Detection; Localisation; Normalisation using facial parts and geometric relationships
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/40Spoof detection, e.g. liveness detection
    • G06V40/45Detection of the body part being alive
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10004Still image; Photographic image
    • G06T2207/10012Stereo images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10048Infrared image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/07Target detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation

Definitions

  • the present disclosure relates to the field of image processing technology, and in particular, to an image processing method, device, electronic device, and computer-readable storage medium.
  • Parallax is the difference in the direction in which the observer looks at the same object at two different positions. For example, when you put a finger in front of your eyes, first close your right eye, look at it with your left eye, then close your left eye, and look at it with your right eye, you will find that the position of the finger relative to a distant object has changed See the parallax of the same point from different angles.
  • the parallax between the two images collected by the binocular camera can effectively estimate the depth, and is widely used in the fields of living body detection, identity authentication, and intelligent driving.
  • the parallax of the two images collected by the binocular camera is predicted by the binocular matching algorithm.
  • the current binocular matching algorithm generally obtains the parallax of the two images by matching all the pixels in the two images, which requires a large amount of calculation and a low matching efficiency.
  • An embodiment of the present disclosure proposes a technical solution for image processing.
  • an image processing method including: acquiring a first target area image of a target object and a second target area image of the target object, wherein the first target The area image is captured from the first image captured by the first image sensor of the binocular camera, and the second target area image is captured from the second image captured by the second image sensor of the binocular camera ; Processing the first target area image and the second target area image to determine the parallax between the first target area image and the second target area image; based on the first target area image and Displacement information between the second target area images and the parallax between the first target area image and the second target area image to obtain a parallax prediction between the first image and the second image result.
  • the acquiring the first target area image of the target object and the second target area image of the target object includes: acquiring the first image collected by the first image sensor in the binocular camera And a second image collected by a second image sensor in the binocular camera; performing target detection on the first image and the second image, respectively, to obtain a first target area image and a second target area image.
  • the acquiring the first target area image of the target object includes: performing target detection on the first image collected by the first image sensor in the binocular camera to obtain a first candidate area; Perform keypoint detection on the image of the first candidate area to obtain keypoint information; based on the keypoint information, intercept the first target area image from the first image.
  • the image sizes of the first target area image and the second target area image are the same.
  • the processing the first target area image and the second target area image to determine the parallax between the first target area image and the second target area image includes: processing the first target area image and the second target area image through a binocular matching neural network to obtain the parallax between the first target area image and the second target area image.
  • the method further includes: based on the position of the first target area image in the first image and the first The position of the second target area image in the second image determines displacement information between the first target area image and the second target area image.
  • obtaining a parallax prediction result between the first image and the second image including: adding displacement information between the first target area image and the second target area image and the parallax To obtain the disparity prediction result between the first image and the second image.
  • the method further includes: determining depth information of the target object based on disparity prediction results of the first image and the second image; determining based on depth information of the target object Biopsy results.
  • the binocular camera includes one of a same-mode binocular camera and a cross-mode binocular camera.
  • the first image sensor or the second image sensor includes one of the following image sensors: a visible light image sensor, a near infrared image sensor, and a dual-pass image sensor.
  • the target object includes a human face.
  • another image processing method including: acquiring a first target area image of a target object and a second target area image of the target object, wherein the first The target area image is captured from the first image collected from the image acquisition area at the first moment, and the second target area image is captured from the second image collected from the image acquisition area at the second moment; Processing the first target area image and the second target area image to determine optical flow information between the first target area image and the second target area image; based on the first target area image Displacement information between the second target area image and optical flow information between the first target area image and the second target area image to obtain between the first image and the second image Prediction results of optical flow information.
  • the acquiring the first target area image of the target object and the second target area image of the target object includes: acquiring the first image collected from the image acquisition area at the first moment and At the second moment, the second image collected in the image collection area; performing target detection on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • the acquiring the first target area image of the target object includes: performing target detection on the first image acquired from the image acquisition area at the first moment to obtain a first candidate area; Perform keypoint detection on the image of the first candidate area to obtain keypoint information; based on the keypoint information, intercept the first target area image from the first image.
  • the image sizes of the first target area image and the second target area image are the same.
  • the first target area image and the second target area image are processed to determine the optical flow between the first target area image and the second target area image
  • the information includes: processing the first target area image and the second target area image through a neural network to obtain optical flow information between the first target area image and the second target area image.
  • the method further includes: based on the position of the first target area image in the first image And the position of the second target area image in the second image, determining displacement information between the first target area image and the second target area image.
  • the optical flow information to obtain the optical flow information prediction result between the first image and the second image includes: the displacement information between the first target area image and the second target area image and the The optical flow information is added to obtain the optical flow information prediction result between the first image and the second image.
  • another image processing method including: acquiring a first target area image captured from a first image and a second target area image captured from a second image; Processing the first target area image and the second target area image to obtain a relative processing result of the first image and the second image; based on the displacement of the first target area image and the second target area image Information and the relative processing result of the first image and the image to obtain the final processing result of the first image and the second image.
  • the first image and the second image are images collected by two image sensors of a binocular camera at the same time.
  • the relative processing result is relative disparity
  • the final processing result is a disparity prediction result
  • the determination process of the disparity prediction result may refer to the method in the first aspect or any possible implementation manner of the first aspect.
  • the first image and the second image are images collected by the camera on the same target area at different times.
  • the relative processing result is a relative optical flow
  • the final processing result is an optical flow prediction result
  • the determination process of the optical flow prediction result may refer to the method in the second aspect or any possible implementation manner of the second aspect.
  • an image processing apparatus including: an acquiring unit configured to acquire a first target area image of a target object and a second target area image of the target object, wherein The first target area image is captured from the first image collected by the first image sensor of the binocular camera, and the second target area image is the second image collected from the second image sensor of the binocular camera Intercepted in the image; a first determining unit configured to process the first target area image and the second target area image to determine the distance between the first target area image and the second target area image Parallax; a second determination unit configured to be based on the displacement information between the first target area image and the second target area image and the parallax between the first target area image and the second target area image To obtain the disparity prediction result between the first image and the second image.
  • the acquiring unit is configured to acquire the first image collected by the first image sensor in the binocular camera and the second image collected by the second image sensor in the binocular camera; Target detection is performed on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • the acquisition unit includes a target detection unit, a key point detection unit, and an interception unit;
  • the target detection unit is configured to acquire the first image collected by the first image sensor in the binocular camera Perform target detection to obtain a first candidate area;
  • the key point detection unit is configured to perform key point detection on the image of the first candidate area to obtain key point information;
  • the interception unit is configured to be based on the key point Information, intercept the first target area image from the first image.
  • the image sizes of the first target area image and the second target area image are the same.
  • the first determining unit is configured to process the first target area image and the second target area image through a binocular matching neural network to obtain the first target area image And the parallax between the images of the second target area.
  • the device further includes a displacement determination unit configured to, between the second determination unit based on the first target area image and the second target area image Displacement information and the parallax between the first target area image and the second target area image, before obtaining the parallax prediction result between the first image and the second image, based on the first target The position of the area image in the first image and the position of the second target area image in the second image, determining displacement information between the first target area image and the second target area image .
  • the second determining unit is configured to add displacement information between the first target area image and the second target area image and the parallax to obtain the first The disparity prediction result between the image and the second image.
  • the device further includes a depth information determination unit and a living body detection determination unit; the depth information determination unit is configured to determine based on the parallax prediction results of the first image and the second image Depth information of the target object; the living body detection determination unit is configured to determine a living body detection result based on the depth information of the target object.
  • the binocular camera includes one of a same-mode binocular camera and a cross-mode binocular camera.
  • the first image sensor or the second image sensor includes one of the following image sensors: a visible light image sensor, a near infrared image sensor, and a dual-pass image sensor.
  • the target object includes a human face.
  • an image processing apparatus including: an acquiring unit configured to acquire a first target area image of a target object and a second target area image of the target object, wherein The first target area image is intercepted from the first image collected from the image acquisition area at the first moment, and the second target area image is taken from the second image collected from the image acquisition area at the second moment Intercepted; a first determining unit configured to process the first target area image and the second target area image to determine the optical flow between the first target area image and the second target area image Information; a second determination unit configured to be based on displacement information between the first target area image and the second target area image and light between the first target area image and the second target area image Flow information to obtain a prediction result of optical flow information between the first image and the second image.
  • the acquiring unit is configured to acquire the first image acquired at the image acquisition area at the first moment and the second image acquired at the image acquisition area at the second moment; Target detection is performed on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • the acquisition unit includes a target detection unit, a key point detection unit, and an interception unit;
  • the target detection unit is configured to perform the first image acquisition on the image acquisition area at the first moment Target detection to obtain a first candidate area;
  • the key point detection unit is configured to perform key point detection on the image of the first candidate area to obtain key point information;
  • the interception unit is configured to be based on the key point information , Intercepting the first target area image from the first image.
  • the image sizes of the first target area image and the second target area image are the same.
  • the first determining unit is configured to process the first target area image and the second target area image through a neural network to obtain the first target area image and the Optical flow information between images in the second target area.
  • the device further includes a displacement determination unit configured to, between the second determination unit based on the first target area image and the second target area image Displacement information and optical flow information between the first target area image and the second target area image, before obtaining the optical flow information prediction result between the first image and the second image, based on The position of the first target area image in the first image and the position of the second target area image in the second image, determining the relationship between the first target area image and the second target area image Displacement information.
  • a displacement determination unit configured to, between the second determination unit based on the first target area image and the second target area image Displacement information and optical flow information between the first target area image and the second target area image, before obtaining the optical flow information prediction result between the first image and the second image, based on The position of the first target area image in the first image and the position of the second target area image in the second image, determining the relationship between the first target area image and the second target area image Displacement information.
  • the second determining unit is configured to add displacement information between the first target area image and the second target area image and the optical flow information to obtain the The optical flow information prediction result between the first image and the second image.
  • an electronic device including: a processor; a memory for storing computer-readable instructions; wherein the processor is for invoking the computer-readable instructions stored in the memory, to Perform the image processing method described in the first aspect or the second aspect or any possible implementation manner thereof.
  • a computer-readable storage medium having computer program instructions stored thereon, which when executed by a processor implements the image processing method of the first aspect or the second aspect described above or Any possible implementation.
  • a computer program product comprising computer instructions which, when executed by a processor, implement the image processing method of the first aspect or the second aspect or any possible implementation manner thereof.
  • the computer program product includes a computer-readable storage medium storing the computer instructions.
  • acquiring a first target area image of a target object and a second target area image of the target object processing the first target area image and the second target area image to determine the The parallax between the first target area image and the second target area image; based on the displacement information between the first target area image and the second target area image and the first target area image and the The disparity between the images of the second target area, to obtain a disparity prediction result between the first image and the second image.
  • the embodiments of the present disclosure can reduce the calculation amount of the parallax prediction, increase the prediction speed of the parallax, and facilitate real-time prediction of the parallax.
  • FIG. 1 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure
  • FIG. 2 is a schematic diagram of determining parallax of a first target area image and a second target area image provided by an embodiment of the present disclosure
  • FIG. 3 is an exemplary schematic diagram of a method for determining a displacement of a target area provided by an embodiment of the present disclosure
  • FIG. 4 is another schematic flowchart of an image processing method provided by an embodiment of the present disclosure.
  • FIG. 5 is a schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • FIG. 6 is another schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • FIG. 7 is another schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • FIG 8 is another schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • FIG. 9 is a structural block diagram of an electronic device provided by an embodiment of the present disclosure.
  • the term “if” may be interpreted as “when” or “once” or “in response to a determination” or “in response to detection” depending on the context .
  • the phrase “if determined” or “if [described condition or event] is detected” may be interpreted in the context to mean “once determined” or “in response to a determination” or “once detected [described condition or event ]” or “In response to detection of [the described condition or event]”.
  • the image processing method provided by the embodiments of the present disclosure may be implemented by a terminal device or server with image processing function such as a mobile phone, a desktop computer, a laptop computer, a wearable device, or other types of electronic devices or systems, which is not limited herein.
  • a terminal device or server with image processing function such as a mobile phone, a desktop computer, a laptop computer, a wearable device, or other types of electronic devices or systems, which is not limited herein.
  • image processing device the execution subject of the image processing method will be referred to as an image processing device hereinafter.
  • FIG. 1 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure.
  • the two image sensors in the binocular camera are referred to as a first image sensor and a second image sensor.
  • the two image sensors of the binocular camera may be arranged horizontally, vertically, or other arrangements, which are not specifically limited in the embodiments of the present disclosure.
  • the first image sensor and the second image sensor may be a device with a shooting function, such as a camera.
  • the first image sensor or the second image sensor includes one of the following image sensors: a visible light image sensor, a near-infrared image sensor, and a dual-pass image sensor.
  • the first image sensor or the second image sensor in the embodiments of the present disclosure may also be other types of image sensors, and the specific type is not limited herein.
  • the visible light image sensor is an image sensor that illuminates an object with visible light to form an image.
  • the near-infrared image sensor is an image sensor that uses near infrared rays to irradiate an object to form an image.
  • the dual-pass image sensor includes an image sensor that uses the dual-channel (including R-channel) imaging principle to form an image.
  • the two image sensors in the binocular camera can be the same type of image sensor or different types of image sensors, that is, the binocular camera can be the same-mode binocular camera or the cross-modal binocular camera.
  • the two image sensors of binocular camera A are both visible light image sensors; or, the two image sensors of binocular camera B are both near infrared image sensors; or, the two image sensors of binocular camera C are both dual-pass Image sensor; or, the two image sensors of the binocular camera D are a visible light image sensor and a near-infrared image sensor; or, the two image sensors of the binocular camera E are a visible light image sensor and a two-pass image sensor; or, dual The two image sensors of the eye camera F are a near infrared image sensor and a dual-pass image sensor, and so on. You can choose the type of the two image sensors in the binocular camera according to actual needs, which has a wider adaptation range and greater scalability.
  • the target object may be a specific object such as a human body, a human face, a mask, ears, and clothing.
  • the target object may be various living objects or a part of living objects, for example, the target object may be a human, animal, human face, or the like.
  • the target object may be various types of apparel, such as headwear, tops, bottoms, and jumpsuits.
  • the target object may be a road, a building, a pedestrian, a traffic light, a designated part of a vehicle or a vehicle, etc.
  • the target object may be a bicycle, a car, a bus, a truck, a front, a car
  • the embodiments of the present disclosure do not limit the specific implementation of the target object.
  • the target object may be a human face.
  • the first target area image and the second target area image are images containing a face area or images containing a face area.
  • the target object in the embodiment of the present application is not limited to a human face, but may also be other objects.
  • the first image is acquired by the first image sensor of the binocular camera
  • the second image is acquired by the second image sensor of the binocular camera.
  • the An image and a second image may be a left view and a right view, respectively, or the first image and the second image may be a right view and a left view, respectively, which is not limited in the embodiments of the present disclosure.
  • the acquiring the first target area image of the target object and the second target area image of the target object includes: acquiring the first collected by the first image sensor of the binocular camera An image and a second image collected by the second image sensor of the binocular camera, and intercepting the first target area image of the target object from the first image, and intercepting the second target area of the target object from the second image image.
  • the binocular camera can collect static image pairs to obtain image pairs including the first image and the second image; or, the binocular camera can collect continuous video streams by selecting the video stream The frame operation results in an image pair including the first image and the second image.
  • the first image and the second image may be a static image obtained from a static image pair, or a video frame image obtained from a video stream, which is not limited in the embodiments of the present disclosure.
  • a binocular camera is provided on the image processing device, and the image processing device performs static image pair or video stream collection through the binocular camera to obtain an image pair including the first image and the second image.
  • the image processing apparatus may also receive image pairs including the first image and the second image sent by other devices.
  • the image processing apparatus acquires an image pair including the first image and the second image from a database provided at other equipment.
  • the image pair including the first image and the second image may be carried in a living body detection request, an identity authentication request, a depth prediction request, a binocular matching request, or other messages.
  • the image processing device then intercepts the first target area image and the second target area image from the first image and the second image, respectively, which is not limited in the embodiments of the present disclosure.
  • the image processing apparatus receives the image pair including the first image and the second image sent by the terminal device provided with the binocular camera, wherein, optionally, the terminal device may send the image pair including the first image to the image processing apparatus (such as a server)
  • the image pair of the image and the second image where the image pair including the first image and the second image may be a static image pair collected by the terminal device through the binocular camera or a frame selected from the video stream collected by the binocular camera Video frame image pair.
  • the terminal device sends a video sequence including the image pair to the image processing device. After receiving the video stream sent by the terminal device, the image processing device obtains the image pair including the first image and the second image through frame selection. The embodiment does not limit this.
  • the video stream may be frame-selected in various ways to obtain an image pair including the first image and the second image.
  • the video stream or video sequence collected by the first image sensor may be frame-selected to obtain a first image, and the first image may be found from the video stream or video sequence collected by the second image sensor Corresponding to the second image, an image pair including the first image and the second image is obtained.
  • the first image is selected from the multi-frame images included in the first video stream collected by the first image sensor based on the image quality, where the image quality may be based on image clarity, image brightness, image exposure, image Contrast, completeness of face, whether there is occlusion on the face, or any combination of multiple factors can be considered, based on image clarity, image brightness, image exposure, image contrast, face integrity, face Is there a combination of one or more factors, such as occlusion, to select the first image from the multi-frame images included in the first video stream collected by the first image sensor.
  • the image quality may be based on image clarity, image brightness, image exposure, image Contrast, completeness of face, whether there is occlusion on the face, or any combination of multiple factors can be considered, based on image clarity, image brightness, image exposure, image contrast, face integrity, face Is there a combination of one or more factors, such as occlusion, to select the first image from the multi-frame images included in the first video stream collected by the first image sensor.
  • the video stream may be frame-selected based on the face state and image quality of the target object included in the image to obtain the first image. For example, based on the key point information obtained through key point detection, the face state of the target object in each frame of the first video stream or at intervals of several frames in the image is determined.
  • the face state is, for example, the face orientation, and the determined Describe the image quality of each frame or several frames in the first video stream, integrate the face state and image quality of the target object in the image frame, and select the face state to meet the preset conditions (for example, face orientation is frontal orientation Or, the angle between the face orientation and the forward direction is lower than the set threshold) and one or more frames of higher image quality are used as the first image.
  • the first image can also be obtained by performing a frame selection operation based on the state of the target object included in the image.
  • the state of the target object includes any combination of one or more of the following factors: whether the face orientation in the image is face-to-face, whether it is in the closed-eye state, whether it is in the open-mouth state, and whether there is motion blur Or the focus is blurred, etc., which is not limited in the embodiments of the present disclosure.
  • the first video stream collected by the first image sensor and the second video stream collected by the second image sensor may be jointly framed to obtain an image including the first image and the second image Correct.
  • an image pair is selected from the video stream collected by the binocular camera, and the two images included in the selected image pair both satisfy the set condition.
  • the specific implementation of the set condition can be described above. For brevity, here No longer.
  • the binocular matching process is performed on the first image and the second image (for example, the first target area image is intercepted from the first image, and the second target area image is intercepted from the second image )
  • the first image and the second image may also be corrected so that the corresponding pixel points in the first image and the second image are on the same horizontal line.
  • the first image and the second image may be subjected to binocular correction processing based on the parameters of the binocular camera obtained by the calibration, for example, based on the parameters of the first image sensor, the parameters of the second image sensor, and the first The relative position parameter between the image sensor and the second image sensor performs binocular correction processing on the first image and the second image.
  • the first image and the second image may be automatically corrected without relying on the parameters of the binocular camera, for example, acquiring key point information of the target object in the first image (i.e. A key point information) and the key point information of the target object in the second image (ie, the second key point information), and based on the first key point information and the second key point information, determine the target transformation matrix (for example, using the minimum The target transformation matrix is determined by two methods), and then the first image or the second image is transformed based on the target transformation matrix to obtain the transformed first image or second image, but the embodiment of the present disclosure does not limit this.
  • the corresponding pixels in the first target area image and the second target area image are on the same horizontal line.
  • at least one of the first image and the second image may be subjected to preprocessing such as translation and/or rotation, so that the preprocessed first image and the second image
  • preprocessing such as translation and/or rotation
  • the two image sensors in the binocular camera are not calibrated.
  • the first image and the second image can be matched and detected and corrected, so that the correspondence between the corrected first image and the second image
  • the pixels are on the same horizontal line, which is not limited in the embodiments of the present disclosure.
  • the two image sensors of the binocular camera may be calibrated in advance to obtain the parameters of the first image sensor and the second image sensor.
  • the first target area image of the target object and the second target area image of the target object may be acquired in various ways.
  • the image processing apparatus may directly obtain the first target area image and the second target area image from other devices, where the first target area image and the second target area image are from the first image and the second target area image, respectively. Captured in the second image.
  • the first target area image and the second target area image may be carried in a living body detection request, an identity authentication request, a depth prediction request, a binocular matching request, or other messages, which are not limited in the embodiments of the present disclosure.
  • the image processing apparatus acquires the first target area image and the second target area image from a database provided at other equipment.
  • the image processing apparatus receives the first target area image and the second target area image sent by the terminal device provided with a binocular camera, where, optionally, the terminal device may The static image pair of the image and the second image respectively captures the first target area image and the second target area image from the first image and the second image; or, the terminal device collects the video sequence through the binocular camera and compares the collected video The sequence performs frame selection to obtain a video frame image pair including the first image and the second image.
  • the terminal device sends a video stream including an image pair of the first image and the second image to the image processing device, and then intercepts the first target area image and the second target area image from the first image and the second image respectively.
  • the disclosed embodiments do not limit this.
  • the acquiring the first target area image of the target object and the second target area image of the target object includes: acquiring the first collected by the first image sensor in the binocular camera An image and a second image collected by a second image sensor in the binocular camera; performing target detection on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • target detection can be performed on the first image and the second image respectively to obtain first position information of the target object in the first image and second position information of the target object in the second image, and The first target area image is intercepted from the first image based on the first position information, and the second target area image is intercepted from the second image based on the second position information.
  • the first image and the second image may be directly subjected to target detection, or the first image and/or the second image may be preprocessed first, and the preprocessed first image and/or the second image may be processed.
  • the pre-processing may include one or more processes such as brightness adjustment, size adjustment, translation, and rotation, which are not limited in the embodiments of the present disclosure.
  • the acquiring the first target area image of the target object includes: performing target detection on the first image collected by the first image sensor in the binocular camera to obtain a first candidate Area; performing key point detection on the image of the first candidate area to obtain key point information; based on the key point information, intercepting a first target area image from the first image.
  • target detection can be performed on the first image and the second image respectively to obtain a first candidate area in the first image and a second candidate in the second image corresponding to the first candidate area Region, based on the first candidate region to intercept the first target region image from the first image, based on the second candidate region to intercept the second target region image from the second image.
  • the image of the first candidate area may be intercepted from the first image as the first target area image.
  • the first target area is obtained by enlarging the first candidate area by a certain factor, and the image of the first target area is intercepted from the first image as the first target area image
  • the first key point information corresponding to the first candidate area is obtained, and the first target area is intercepted from the first image based on the obtained first key point information image.
  • second key point information corresponding to the second candidate area is obtained, and the second target area image is intercepted from the second image based on the obtained second key point information.
  • target detection can be performed on the first image through image processing technology (such as a convolutional neural network) to obtain the first candidate region to which the target object belongs.
  • image processing technology such as Convolutional neural network
  • image processing technology performs target detection on the second image to obtain a second candidate area to which the target object belongs; wherein, the first candidate area and the second candidate area are, for example, a first face area.
  • the target detection may be a rough positioning of the target object; accordingly, the first candidate area is a preliminary area including the target object, and the second candidate area is a preliminary area including the target object.
  • the above key point detection can be achieved through deep neural networks, such as convolutional neural networks, recurrent neural networks, etc., which can specifically be any type of neural network model such as LeNet, AlexNet, GoogLeNet, VGGNet, ResNet; or, key point detection also It may be implemented based on other machine learning methods.
  • deep neural networks such as convolutional neural networks, recurrent neural networks, etc.
  • LeNet LeNet
  • AlexNet GoogLeNet
  • VGGNet VGGNet
  • ResNet ResNet
  • key point detection also It may be implemented based on other machine learning methods.
  • the embodiments of the present disclosure do not limit the specific implementation of key point detection.
  • the key point information may include position information of each key point among multiple key points of the target object, or further include information such as confidence level, which is not limited in the embodiments of the present disclosure.
  • a face key point detection model is used to perform face key point detection on the images of the first candidate area and the second candidate area, respectively, to obtain the first
  • the image of the candidate area includes multiple key point information corresponding to key points of the face
  • the image of the second candidate area includes multiple key point information corresponding to key points of the face, based on the multiple key point information
  • the position information of the human face may be determined, and based on the position information of the human face, a first target area corresponding to the human face and a second target area corresponding to the human face may be determined.
  • the first target area and the second target area are more accurate positions of the human face, thereby helping to improve the accuracy of subsequent operations.
  • the target detection performed on the first image and the second image in the above embodiments does not need to determine the precise position of the target object or the area to which it belongs, but only needs to roughly locate the target object or the area to which it belongs, thereby reducing the target detection algorithm The accuracy of the requirements, improve the robustness and image processing speed.
  • the interception method of the second target area image and the interception method of the first target area image may or may not be the same, which is not limited in the embodiments of the present disclosure.
  • the images of the first target area image and the second target area image may have different sizes.
  • the image sizes of the first target area image and the second target area image are the same.
  • the interception parameters that characterize the same size can be used to intercept the first target area image and the second target area image from the first image and the second image, respectively, so that the first target area image and all The image size of the second target area image is the same.
  • two interception frames that completely include the target object can be obtained.
  • target detection may be performed on the first image and the second image, so that the obtained first intercept frame corresponding to the first image and the second intercept frame corresponding to the second image have the same size .
  • the first intercept frame and the second intercept frame are respectively amplified by different multiples, that is, the first intercept frame corresponds to The first interception parameter and the second interception parameter corresponding to the second interception frame are subjected to different magnification processing, so that the two interception frames obtained by the magnification processing have the same size.
  • the first target area and the second target area having the same size are determined based on the key point information of the first image and the key point information of the second image, wherein the first target area and the second target The area completely includes the target object, and so on.
  • the depth information of the image may be obtained by predicting the parallax of the image, and then determining whether the face included in the image is a living human face. Based on this, it is only necessary to focus on the face area of the image, so if only disparity prediction is performed on the face area of the image, unnecessary calculations can be avoided, thereby increasing the speed of disparity prediction.
  • the first target area image and the second target area image are processed to determine the first target area image and the second target area image
  • the parallax between the target area images includes: processing the first target area image and the second target area image through a binocular matching neural network to obtain the first target area image and the second target area Parallax of the image.
  • the first target area image and the second target area image are processed by a binocular matching neural network, and the parallax between the first target area image and the second target area image is obtained and output.
  • the first target area image and the second target area image are directly input into a binocular matching neural network for processing to obtain the parallax between the first target area image and the second target area image.
  • the first target area image and/or the second target area image may be pre-processed, such as normalization processing, etc., and then the pre-processed first target area image and The second target area image is input into a binocular matching neural network for processing to obtain the parallax between the first target area image and the second target area image.
  • the embodiments of the present disclosure do not limit this.
  • FIG. 2 is a schematic diagram of determining the parallax of the first target area image and the second target area image provided by an embodiment of the present disclosure, wherein the first target area image and the second target area image are input to the dual In the eye-matching neural network, through the binocular-matching neural network, the first feature of the first target area image (ie, feature 1 in FIG. 2) and the second feature of the second target area image (ie, image Feature 2 in 2), the matching cost calculation module in the binocular matching neural network calculates the matching cost of the first feature and the second feature, and determines the first target area image and the second target based on the obtained matching cost The disparity between the regional images; wherein, the matching cost may represent the correlation between the first feature and the second feature.
  • the determining the disparity between the first target area image and the second target area image based on the obtained matching cost includes: performing feature extraction on the matching cost, and determining the first target area based on the extracted feature data The parallax between the image and the image of the second target area.
  • the disparity between the first target area image and the second target area image may be determined by other binocular matching algorithms based on machine learning.
  • the binocular matching algorithm may be any of the following algorithms: stereo binocular vision algorithm (Sum of absolute differences SAD), bidirectional matching algorithm (bidirectional matching) BM, global matching algorithm (Semi-global block) matching (SGBM), graph cut algorithm (Graph Cuts, GC), the specific implementation of the binocular matching process in the embodiments of the present disclosure is not limited.
  • the method before performing step 103, that is, based on the displacement information between the first target area image and the second target area image and the first target area image before obtaining the parallax prediction result between the first image and the second image, the method further includes: based on the first target area image in the first image And the position of the second target area image in the second image, determining the displacement information of the first target area image and the second target area image.
  • the displacement information may include a displacement in the horizontal direction and/or a displacement in the vertical direction, where, in some embodiments, if the corresponding pixels in the first image and the second image are on the same horizontal line, then The displacement information may include only the displacement in the horizontal direction, but the embodiment of the present disclosure does not limit this.
  • the determining displacement information of the first target area image and the second target area image based on the position of the first target area image in the first image and the position of the second target area image in the second image includes: determining Determining the position of the first center point of the first target area image; determining the position of the second center point of the second target area image; based on the position of the first center point and the position of the second center point, determining The displacement information between the first target area image and the second target area image.
  • FIG. 3 is an exemplary schematic diagram of a target area displacement determination method provided by an embodiment of the present disclosure.
  • the position of the center point a of the first target area image in the first image is expressed as (x 1 , y 1 ), and
  • the position of the center point b of the second target area image in the two images is expressed as (x 2 , y 1 ), and the displacement between the center point a and the center point b is expressed as That is, displacement information between the first target area image and the first target area image.
  • the above center point may be replaced with any one of the four vertices of the target area image, which is not specifically limited in the embodiment of the present disclosure.
  • the displacement information between the first target area image and the second target area image may also be determined in other ways, which is not limited in the embodiment of the present disclosure.
  • step 103 based on the displacement information between the first target area image and the second target area image and the first target area image and the first The disparity between the two target area images to obtain the disparity prediction result between the first image and the second image includes: disparity between the first target area image and the second target area image And the displacement information are added to obtain a disparity prediction result between the first image and the second image.
  • the displacement information between the first target area image and the second target area image is x
  • the parallax of the first target area image and the second target area image is D(p).
  • the information is a result obtained by adding or subtracting x and the parallax D(p), and used as a parallax prediction result between the first image and the second image.
  • the displacement between the first target area image and the second target area image is 0, then the parallax between the first target area image and the second target area image is the difference between the first image and the second image Parallax.
  • the determination of the displacement information and the determination of the parallax between the first target area image and the second target area image may be performed in parallel, or in any sequential order.
  • the execution order of the determination of the displacement information and the determination of the parallax between the first target area image and the second target area image is not limited.
  • the method further includes: after obtaining the disparity prediction results of the first image and the second image, based on the difference between the first image and the second image
  • the parallax prediction result determines the depth information of the target object; based on the depth information of the target object, the living body detection result is determined.
  • acquiring a first target area image of a target object and a second target area image of the target object processing the first target area image and the second target area image to determine the The parallax between the first target area image and the second target area image; based on the displacement information between the first target area image and the second target area image and the first target area image and the The disparity between the images of the second target area, to obtain a disparity prediction result between the first image and the second image.
  • the embodiments of the present disclosure can reduce the calculation amount of the parallax prediction, thereby increasing the prediction speed of the parallax, which is beneficial to realize the real-time prediction of the parallax.
  • the first image and the second image are images collected by the monocular camera at different times, and so on, which is not limited in the embodiments of the present disclosure.
  • FIG. 4 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure.
  • 201 Acquire a first target area image of a target object and a second target area image of the target object, wherein the first target area image is intercepted from the first image collected from the image collection area at the first moment , The second target area image is taken from the second image collected from the image collection area at the second moment;
  • the image acquisition area may be image-collected by a monocular camera, and the first target area image and the second target area image may be obtained based on the images collected at different times.
  • the image collected at the first moment is recorded as the first image
  • the first target area image is obtained from the first image
  • the image collected at the second moment is recorded as the second image, from the second image Obtain a second target area image.
  • the acquiring the first target area image of the target object and the second target area image of the target object includes: acquiring the first An image and a second image acquired at the second moment in the image acquisition area; performing target detection on the first image and the second image respectively to obtain a first target area image and a second target area image .
  • the acquiring the first target area image of the target object includes: performing target detection on the first image acquired from the image acquisition area at the first moment to obtain a first candidate area; Perform keypoint detection on the image of the first candidate area to obtain keypoint information; based on the keypoint information, intercept the first target area image from the first image.
  • the image sizes of the first target area image and the second target area image are the same.
  • step 201 for the related description of step 201, reference may be made to the detailed description of step 101 in the foregoing embodiment, and details are not described herein again.
  • the first target area image and the second target area image are processed to determine between the first target area image and the second target area image
  • Optical flow information including: processing the first target area image and the second target area image through a neural network to obtain light between the first target area image and the second target area image ⁇ Flow information.
  • the first target area image and the second target area image can be processed through the neural network to obtain optical flow information between the first target area image and the second target area image.
  • the first target area image and the second target area image may be input into a neural network for processing to obtain optical flow information between the first target area image and the second target area image; in another In some optional embodiments, the first target area image and/or the second target area image may be pre-processed, such as normalization processing, etc., and then the pre-processed first target area image and the second The target area image is input into the neural network to obtain optical flow information between the first target area image and the second target area image.
  • the optical flow information is a relative concept, which can represent the relative optical flow information of the target object, that is, the Describe the relative motion of the target object.
  • the method further includes: based on the first target area image in the first image And the position of the second target area image in the second image, determining displacement information between the first target area image and the second target area image.
  • the optical flow information between the first image and the second image to obtain a prediction result of the optical flow information between the first image and the second image including: displacement information between the first target area image and the second target area image Add to the optical flow information to obtain a prediction result of optical flow information between the first image and the second image.
  • the positions corresponding to the first target area image and the second target area image are not absolutely unchanged, it is necessary to determine the displacement information of the first target area image and the second target area image, and then Adding or subtracting the displacement information and the optical flow information to obtain an optical flow information prediction result.
  • the prediction result of the optical flow information may represent absolute optical flow information of the target object, that is, the absolute motion of the target object.
  • the image processing method of the embodiment of the present disclosure is applied to the prediction of optical flow information, and the image processing method described in FIG. 1 is applied to the prediction of parallax information.
  • the technical realization of the two is basically the same.
  • the specific image processing method of the embodiment of the present disclosure is specific.
  • FIG. 5 is a first structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • the device 500 includes: an obtaining unit 501, a first determining unit 502, and a second determining unit 503; wherein,
  • the acquiring unit 501 is configured to acquire a first target area image of the target object and a second target area image of the target object, wherein the first target area image is acquired from the first image sensor of the binocular camera Captured in the first image of, the second target area image is captured from the second image collected by the second image sensor of the binocular camera;
  • the first determining unit 502 is configured to process the first target area image and the second target area image to determine the parallax between the first target area image and the second target area image;
  • the second determining unit 503 is configured to be based on displacement information between the first target area image and the second target area image and between the first target area image and the second target area image Parallax, to obtain a parallax prediction result between the first image and the second image.
  • the obtaining unit 501 is configured to obtain the first image collected by the first image sensor in the binocular camera and the second image collected by the second image sensor in the binocular camera Two images; performing target detection on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • the acquisition unit 501 includes a target detection unit 501-1, a key point detection unit 501-2, and an interception unit 501-3, and the target detection unit 501-1, Configured to perform target detection on the first image collected by the first image sensor in the binocular camera to obtain a first candidate area;
  • the key point detection unit 501-2 is configured to image the first candidate area Perform key point detection to obtain key point information;
  • the intercepting unit 501-3 is configured to intercept the first target area image from the first image based on the key point information.
  • the image sizes of the first target area image and the second target area image are the same.
  • the first determining unit 502 is configured to process the first target area image and the second target area image through a binocular matching neural network to obtain the first The parallax between the target area image and the second target area image.
  • the device further includes a displacement determination unit 701 configured to, based on the first target area image and Displacement information between the second target area images and the parallax between the first target area image and the second target area image to obtain a parallax prediction between the first image and the second image Before the result, based on the position of the first target area image in the first image and the position of the second target area image in the second image, determine the first target area image and the first The displacement information between the two target area images.
  • a displacement determination unit 701 configured to, based on the first target area image and Displacement information between the second target area images and the parallax between the first target area image and the second target area image to obtain a parallax prediction between the first image and the second image Before the result, based on the position of the first target area image in the first image and the position of the second target area image in the second image, determine the first target area image and the first The displacement information between the two target area images.
  • the second determining unit 503 is configured to divide the displacement information between the first target area image and the second target area image and the first target area image and The disparity between the images of the second target area is added to obtain a disparity prediction result between the first image and the second image.
  • the device further includes a depth information determination unit 702 and a living body detection determination unit 703, and the depth information determination unit 702 is configured to obtain based on the second determination unit 503 The disparity prediction results of the first image and the second image to determine the depth information of the target object; the living body detection determination unit 703 is configured to obtain the target object based on the depth information determination unit 702 Depth information to determine the biopsy results.
  • the binocular camera includes one of a same-mode binocular camera and a cross-mode binocular camera.
  • the first image sensor or the second image sensor includes one of the following image sensors: a visible light image sensor, a near infrared image sensor, and a dual-pass image sensor.
  • the target object includes a human face.
  • Embodiments of the present disclosure also provide an image processing device.
  • 8 is a fourth schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure.
  • the device 800 includes: an obtaining unit 801, a first determining unit 802, and a second determining unit 803; wherein,
  • the acquiring unit 801 is configured to acquire a first target area image of the target object and a second target area image of the target object, wherein the first target area image is acquired from the image acquisition area at the first moment Intercepted from the first image, the second target area image is intercepted from the second image collected from the image acquisition area at the second moment;
  • the first determining unit 802 is configured to process the first target area image and the second target area image to determine the optical flow between the first target area image and the second target area image information;
  • the second determining unit 803 is configured to be based on the displacement information between the first target area image and the second target area image and between the first target area image and the second target area image Optical flow information to obtain a prediction result of optical flow information between the first image and the second image.
  • the acquiring unit 801 is configured to acquire the first image acquired from the image acquisition area at the first time and the second image acquired from the image acquisition area at the second time Two images; performing target detection on the first image and the second image respectively to obtain a first target area image and a second target area image.
  • the acquisition unit 801 includes a target detection unit, a key point detection unit, and an interception unit; wherein,
  • the target detection unit is configured to perform target detection on the first image collected from the image collection area at the first moment to obtain a first candidate area
  • the key point detection unit is configured to perform key point detection on the image of the first candidate area to obtain key point information
  • the intercepting unit is configured to intercept a first target area image from the first image based on the key point information.
  • the image sizes of the first target area image and the second target area image are the same.
  • the first determining unit 802 is configured to process the first target area image and the second target area image through a neural network to obtain the first target area image And optical flow information between the images of the second target area.
  • the device further includes a displacement determination unit configured to, based on the first target area image and the second target area, in the second determination unit 803 Before the displacement information between the images and the optical flow information between the first target area image and the second target area image, before obtaining the prediction result of the optical flow information between the first image and the second image , Based on the position of the first target area image in the first image and the position of the second target area image in the second image, determining the first target area image and the second target Displacement information between regional images.
  • the second determining unit 803 is configured to add displacement information between the first target area image and the second target area image and the optical flow information, A prediction result of optical flow information between the first image and the second image is obtained.
  • the image processing device of this embodiment is applied to optical flow information prediction.
  • the functions provided by the device provided by the embodiments of the present disclosure or the modules contained therein can be used to execute the method described in the method embodiment shown in FIG. 4 The description of the embodiment of the image processing method will not be repeated here for brevity.
  • FIG. 9 is a structural block diagram of the electronic device provided by an embodiment of the present disclosure.
  • the electronic device includes: a processor 901, and a memory 904 for storing processor-executable instructions, wherein the processor 901 is configured to: execute the image of the embodiment of the present disclosure as shown in FIG. The processing method or any possible implementation manner thereof; or the implementation of the image processing method shown in FIG. 4 or any possible implementation manner of the embodiment of the present disclosure.
  • the electronic device may further include: one or more input devices 902 and one or more output devices 903.
  • the processor 901, the input device 902, the output device 903, and the memory 904 are connected through a bus 905.
  • the memory 902 is used to store instructions, and the processor 901 is used to execute the instructions stored in the memory 902.
  • the processor 901 is configured to call the program instructions to execute any one of the above image processing methods. For brevity, no more details are provided here.
  • the above device embodiments have described the technical solutions of the embodiments of the present disclosure by taking parallax prediction as an example.
  • the technical solutions of the embodiments of the present disclosure can also be applied to optical flow prediction.
  • the optical flow prediction device also belongs to the protection scope of the present disclosure.
  • the optical flow prediction device is similar to the image processing device described above. For simplicity, I won't repeat them here.
  • the so-called processor 901 may be a central processing unit (Central Processing Unit, CPU), and the processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP) , Application Specific Integrated Circuit (Application Specific Integrated Circuit, ASIC), ready-made programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
  • the general-purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
  • the input device 902 may include a mobile phone, a desktop computer, a laptop computer, a wearable device, a monitoring image sensor, etc.
  • the output device 903 may include a display (LCD, etc.).
  • the memory 904 may include a read-only memory and a random access memory, and provide instructions and data to the processor 901. A portion of the memory 904 may also include non-volatile random access memory. For example, the memory 904 may also store device type information.
  • the electronic device described in the embodiment of the present disclosure is used to perform the image processing method described above, and accordingly, the processor 901 is used to perform the steps and/or processes in the various embodiments of the image processing method provided by the embodiment of the present disclosure And will not be repeated here.
  • a computer-readable storage medium stores a computer program, and the computer program includes program instructions.
  • the program instructions are executed by a processor to implement the above.
  • the computer-readable storage medium may be an internal storage unit of the electronic device described in any of the foregoing embodiments, such as a hard disk or a memory of the terminal.
  • the computer-readable storage medium may also be an external storage device of the terminal, for example, a plug-in hard disk equipped on the terminal, a smart memory card (Smart Media (SMC), and a secure digital (SD) card) , Flash card (Flash Card), etc.
  • the computer-readable storage medium may also include both an internal storage unit of the electronic device and an external storage device.
  • the computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device.
  • the computer-readable storage medium can also be used to temporarily store data that has been or will be output.
  • the disclosed server, device, and method may be implemented in other ways.
  • the server embodiments described above are only schematic.
  • the division of the units is only a division of logical functions.
  • there may be other divisions for example, multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not implemented.
  • the displayed or discussed mutual couplings or direct couplings or communication connections may be indirect couplings or communication connections through some interfaces, devices, or units, and may also be electrical, mechanical, or other forms of connection.
  • the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments of the present disclosure.
  • each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.
  • the above integrated unit can be implemented in the form of hardware or software function unit.
  • the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium.
  • the technical solution of the present disclosure essentially or part of the contribution to the existing technology, or all or part of the technical solution can be embodied in the form of a software product
  • the computer software product is stored in a storage medium
  • several instructions are included to enable a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present disclosure.
  • the aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk and other media that can store program code .

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Human Computer Interaction (AREA)
  • Geometry (AREA)
  • Image Analysis (AREA)
  • Studio Devices (AREA)

Abstract

一种图像处理方法、装置、电子设备和存储介质,所述方法包括:获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像(101);对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差(102);基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果(103)。该方法能够减小视差预测的计算量,提高视差的预测速度。

Description

图像处理方法、装置、电子设备及计算机可读存储介质
相关申请的交叉引用
本公开基于申请号为201811647485.8、申请日为2018年12月29日的中国专利申请提出,并要求该中国专利申请的优先权,该中国专利申请的全部内容在此以引入方式并入本公开。
技术领域
本公开涉及图像处理技术领域,尤其涉及一种图像处理方法、装置、电子设备及计算机可读存储介质。
背景技术
视差是观测者在两个不同位置观看同一物体的方向之差。比如,当你伸出一个手指放在眼前,先闭上右眼,用左眼看它,再闭上左眼,用右眼看它,会发现手指相对远方的物体的位置有了变化,这就是从不同角度去看同一点的视差。
基于双目摄像头采集到的两个图像之间的视差能够有效估计深度,被广泛应用于活体检测、身份认证、智能驾驶等领域。双目摄像头采集到的两个图像的视差是通过双目匹配算法进行预测的。目前的双目匹配算法一般通过匹配两个图像中的所有像素点得到两个图像的视差,计算量较大,匹配效率较低。
发明内容
本公开实施例提出一种图像处理的技术方案。
根据本公开实施例的第一方面,提供了一种图像处理方法,该方法包括:获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从双目摄像头的第一图像传感器采集到的第一图像中截取的,所述第二目标区域图像是从所述双目摄像头的第二图像传感器采集到的第二图像中截取的;对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
在一可能的实现方式中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在一可能的实现方式中,所述获取目标对象的第一目标区域图像,包括:对所述双目摄像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;对所述第一候选区域的图像进行关键点检测,得到关键点信息;基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在一可能的实现方式中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在一可能的实现方式中,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差,包括:通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的视差。
在一可能的实现方式中,在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,所述方法还包括:基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在一可能的实现方式中,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果,包括:将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述视差相加,得到所述第一图像和所述第二图像之间的视差预测结果。
在一可能的实现方式中,所述方法还包括:基于所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;基于所述目标对象的深度信息,确定活体检测结果。
在一可能的实现方式中,所述双目摄像头包括同模态双目摄像头和跨模态双目摄像头中的一种。
在一可能的实现方式中,所述第一图像传感器或所述第二图像传感器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。
在一可能的实现方式中,所述目标对象包括人脸。
根据本公开实施例的第二方面,提供了另一种图像处理方法,该方法包括:获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息;基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
在一可能的实现方式中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在一可能的实现方式中,所述获取目标对象的第一目标区域图像,包括:对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;对所述第一候选区域的图像进行关键点检测,得到关键点信息;基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在一可能的实现方式中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在一可能的实现方式中,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息,包括:通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
在一可能的实现方式中,在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,所述方法还包括:基于所述第一目标区域图像在所述第 一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在一可能的实现方式中,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间光流信息预测结果,包括:将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
根据本公开实施例的第三方面,提供另一种图像处理方法,包括:获取从第一图像中截取的第一目标区域图像和从第二图像中截取的第二目标区域图像;通过对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一图像和所述第二图像的相对处理结果;基于所述第一目标区域图像和第二目标区域图像的位移信息以及所述第一图像和所述图像的相对处理结果,得到所述第一图像和所述第二图像的最终处理结果。
在一可能的实现方式中,所述第一图像和所述第二图像为双目摄像头的两个图像传感器在同一时刻采集到的图像。
在一可能的实现方式中,所述相对处理结果为相对视差,所述最终处理结果为视差预测结果。
可选地,所述视差预测结果的确定流程可参照第一方面或第一方面的任意可能的实现方式中的方法。
在另一可能的实现方式中,所述第一图像和所述第二图像为摄像头在不同时刻对同一目标区域采集到的图像。
在一可能的实现方式中,所述相对处理结果为相对光流,所述最终处理结果为光流预测结果。
可选地,所述光流预测结果的确定流程可参照第二方面或第二方面的任意可能的实现方式中的方法。
根据本公开实施例的第四方面,提供一种图像处理装置,该装置包括:获取单元,配置为获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从双目摄像头的第一图像传感器采集到的第一图像中截取的,所述第二目标区域图像是从所述双目摄像头的第二图像传感器采集到的第二图像中截取的;第一确定单元,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;第二确定单元,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
在一可能的实现方式中,所述获取单元配置为:获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在一可能的实现方式中,所述获取单元包括目标检测单元、关键点检测单元和截取单元;所述目标检测单元,配置为对所述双目摄像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;所述关键点检测单元,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;所述截取单元,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在一可能的实现方式中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在一可能的实现方式中,所述第一确定单元配置为,通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的视差。
在一可能的实现方式中,所述装置还包括位移确定单元,所述位移确定单元配置为,在所述第二确定单元基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区 域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在一可能的实现方式中,所述第二确定单元配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述视差相加,得到所述第一图像和所述第二图像之间的视差预测结果。
在一可能的实现方式中,所述装置还包括深度信息确定单元和活体检测确定单元;所述深度信息确定单元,配置为基于所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;所述活体检测确定单元,配置为基于所述目标对象的深度信息,确定活体检测结果。
在一可能的实现方式中,所述双目摄像头包括同模态双目摄像头和跨模态双目摄像头中的一种。
在一可能的实现方式中,所述第一图像传感器或所述第二图像传感器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。
在一可能的实现方式中,所述目标对象包括人脸。
根据本公开实施例的第五方面,提供一种图像处理装置,该装置包括:获取单元,配置为获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;第一确定单元,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息;第二确定单元,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
在一可能的实现方式中,所述获取单元配置为:获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在一可能的实现方式中,所述获取单元包括目标检测单元、关键点检测单元和截取单元;所述目标检测单元,配置为对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;所述关键点检测单元,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;所述截取单元,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在一可能的实现方式中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在一可能的实现方式中,所述第一确定单元配置为,通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
在一可能的实现方式中,所述装置还包括位移确定单元,所述位移确定单元配置为,在所述第二确定单元基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在一可能的实现方式中,所述第二确定单元配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
根据本公开实施例的第五方面,提供一种电子设备,包括:处理器;用于存储计算机可读指令的存储器;其中,所述处理器用于调用所述存储器存储的计算机可读指令,以执行上述第一方面或第二方面所述的图像处理方法或其任意可能的实现方式。
根据本公开的第六方面,提供了一种计算机可读存储介质,其上存储有计算机程序指令,所述计算机程序指令被处理器执行时实现上述第一方面或第二方面图像处理方法或其任意可能的实现方式。
根据本公开的第七方面,提供了一种计算机程序产品,包括计算机指令,所述计算机指令被处理器执行时实现上述第一方面或第二方面图像处理方法或其任意可能的实现方式。
可选地,所述计算机程序产品包括存储所述计算机指令的计算机可读存储介质。
在本公开实施例中,获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像;对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。本公开实施例能够减少视差预测的计算量,提高视差的预测速度,有利于实现视差的实时预测。
根据下面参考附图对示例性实施例的详细说明,本公开的其它特征及方面将变得清楚。
附图说明
为了更清楚地说明本公开实施例的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是本公开实施例提供的图像处理方法示意流程图;
图2是本公开实施例提供的确定第一目标区域图像和第二目标区域图像的视差的示意图;
图3是本公开实施例提供的目标区域位移确定方法的示例性示意图;
图4是本公开实施例提供的图像处理方法的另一示意流程图;
图5是本公开实施例提供的图像处理装置的结构示意图;
图6是本公开实施例提供的图像处理装置的另一结构示意图;
图7是本公开实施例提供的图像处理装置的另一结构示意图;
图8是本公开实施例提供的图像处理装置的另一结构示意图;
图9是本公开实施例提供的电子设备的结构框图。
具体实施方式
下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
应当理解,当在本说明书和所附权利要求书中使用时,术语“包括”和“包含”指示所描述特征、整体、步骤、操作、元素和/或组件的存在,但并不排除一个或多个其它特征、整体、步骤、操作、元素、组件和/或其集合的存在或添加。
还应当理解,在此本公开说明书中所使用的术语仅仅是出于描述特定实施例的目的而并不意在限制本公开。如在本公开说明书和所附权利要求书中所使用的那样,除非上下文清楚地指明其它情况,否则单数形式的“一”、“一个”及“该”意在包括复数形式。
还应当进一步理解,在本公开说明书和所附权利要求书中使用的术语“和/或”是指相关联列出的项中的一个或多个的任何组合以及所有可能组合,并且包括这些组合。
如在本说明书和所附权利要求书中所使用的那样,术语“如果”可以依据上下文被解释为“当... 时”或“一旦”或“响应于确定”或“响应于检测到”。类似地,短语“如果确定”或“如果检测到[所描述条件或事件]”可以依据上下文被解释为意指“一旦确定”或“响应于确定”或“一旦检测到[所描述条件或事件]”或“响应于检测到[所描述条件或事件]”。
本公开实施例提供的图像处理方法可以由手机、台式电脑、膝上计算机、可穿戴设备等具备图像处理功能的终端设备或服务器或其他类型的电子设备或系统实现,此处不作限定。为了便于理解,下文将图像处理方法的执行主体称为图像处理装置。
图1是本公开实施例提供的图像处理方法的示意流程图。
101、获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像。
在本公开实施例中,将双目摄像摄像头中的两个图像传感器称为第一图像传感器和第二图像传感器。双目摄像头的两个图像传感器可以是水平排列,也可以是垂直排列,还可以是其他排列形式,本公开实施例中不作具体限定。作为一种示例,上述第一图像传感器和第二图像传感器可以是具备拍摄功能的装置,如摄像头等。
在一种可能的实现方式中,所述第一图像传感器或所述第二图像传感器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。本公开实施例中的第一图像传感器或第二图像传感器也可以为其他类型的图像传感器,具体类型在此不做限定。
可见光图像传感器为利用可见光照射物体形成图像的图像传感器。近红外图像传感器为利用近红外线照射物体形成图像的图像传感器。双通图像传感器包括利用双通道(包括R通道)成像原理形成图像的图像传感器。双目摄像头中的两个图像传感器可以是相同类型的图像传感器,也可以是不同类型的图像传感器,即双目摄像头可以是同模态双目摄像头,也可以是跨模态双目摄像头。例如,双目摄像头A的两个图像传感器均为可见光图像传感器;或者,双目摄像头B的两个图像传感器均为近红外图像传感器;或者,双目摄像头C的两个图像传感器均为双通图像传感器;或者,双目摄像头D的两个图像传感器分别为可见光图像传感器和近红外图像传感器;或者,双目摄像头E的两个图像传感器分别为可见光图像传感器和双通图像传感器;或者,双目摄像头F的两个图像传感器分别为近红外图像传感器和双通图像传感器,等等。可以根据实际的需求,选择双目摄像头中两个图像传感器的类型,适应范围更广,可扩展性更强。
本公开实施例提供的技术方案可以应用于目标识别、活体检测、智能交通等领域,相应地,目标对象也随着应用领域的不同而不同。其中,在目标识别领域中,所述目标对象可以为人体、人脸、口罩、耳朵、服饰等特定物体。在活体检测领域,所述目标对象可以为各种活体对象或活体对象的一部分,例如,目标对象可以为人、动物、人脸等等。在服饰识别领域中,所述目标对象可以为各种类型的服饰,例如头饰、上衣、下装、连体装等。在智能交通领域中,所述目标对象可以为道路、建筑物、行人、交通指示灯、交通工具或交通工具的指定部位等等,例如目标对象可以为自行车、轿车、巴士、货车、车头、车尾等,本公开实施例对目标对象的具体实现不做限定。
在一些实施例中,所述目标对象可以为人脸,相应地,第一目标区域图像和第二目标区域图像为包含有人脸区域的图像或包含有面部区域的图像。当然,本申请实施例中所述目标对象不限于人脸,也可以是其他对象。
在本公开实施例中,第一图像是通过双目摄像头的第一图像传感器采集到的,第二图像是通过双目摄像头的第二图像传感器采集到的,其中,在一些实施例中,第一图像和第二图像可以分别为左视图和右视图,或者,第一图像和第二图像可以分别为右视图和左视图,本公开实施例对此不做限定。
在本公开的一些可选实施例中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:获取双目摄像头的第一图像传感器采集到的第一图像和所述双目摄像头的第二图像传感器采集到的第二图像,并从第一图像中截取目标对象的第一目标区域图像,从第二图像 中截取所述目标对象的第二目标区域图像。
在一些可选的实施方式中,双目摄像头可以采集静态的图像对,得到包括第一图像和第二图像的图像对;或者,双目摄像头可以采集连续的视频流,通过对视频流进行选帧操作得到包括第一图像和第二图像的图像对。相应地,第一图像和第二图像可以为从静态的图像对中获得的静态图像,或者从视频流中获得的视频帧图像,本公开实施例对此不做限定。
在一些可选的实现方式中,图像处理装置上设置有双目摄像头,图像处理装置通过双目摄像头进行静态图像对或视频流采集,得到包括第一图像和第二图像的图像对,本公开实施例对此不做限定。
在一些可选的实施方式中,图像处理装置还可以接收其他设备发送的包括第一图像和第二图像的图像对。例如,图像处理装置从设置在其他设备处的数据库获取包括第一图像和第二图像的图像对。其中,包括第一图像和第二图像的图像对可以携带在活体检测请求、身份认证请求、深度预测请求、双目匹配请求或其他消息中发送。图像处理装置再分别从第一图像和第二图像中截取第一目标区域图像和第二目标区域图像,本公开实施例对此不做限定。再例如,图像处理装置接收设置有双目摄像头的终端设备发送的包括第一图像和第二图像的图像对,其中,可选地,终端设备可以向图像处理装置(例如服务器)发送包括第一图像和第二图像的图像对,其中,包括第一图像和第二图像的图像对可以是终端设备通过双目摄像头采集到的静态图像对或者是从双目摄像头采集到的视频流中选帧得到的视频帧图像对。再例如,终端设备向图像处理装置发送包括该图像对的视频序列,图像处理装置在接收到终端设备发送的视频流之后,通过选帧得到包括第一图像和第二图像的图像对,本公开实施例对此不做限定。
在本公开实施例中,可以通过多种方式对视频流进行选帧操作得到包括第一图像和第二图像的图像对。
在一些实施例中,可以对第一图像传感器采集到的视频流或视频序列进行选帧处理,得到第一图像,并从第二图像传感器采集到的视频流或视频序列中查找与第一图像对应的第二图像,得到包括第一图像和第二图像的图像对。在一些示例中,基于图像质量,从第一图像传感器采集到的第一视频流包括的多帧图像中选择第一图像,其中,图像质量可以基于图像清晰度、图像亮度、图像曝光度、图像对比度、人脸完整度、人脸是否有遮挡等一种或任意多种因素的组合来进行考量,即可基于图像清晰度、图像亮度、图像曝光度、图像对比度、人脸完整度、人脸是否有遮挡等一种或任意多种因素的组合从第一图像传感器采集到的第一视频流包括的多帧图像中选择第一图像。
在一些可选的实施方式中,可基于图像中包括的目标对象的人脸状态和图像质量对视频流进行选帧操作得到第一图像。例如,基于通过关键点检测得到的关键点信息确定所述第一视频流中每一帧图像或者间隔若干帧图像中目标对象的人脸状态,所述人脸状态例如为人脸朝向,并确定所述第一视频流中每一帧图像或者间隔若干帧图像的图像质量,综合图像帧中的目标对象的人脸状态和图像质量,选择人脸状态符合预设条件(例如人脸朝向为正面朝向或者人脸朝向与正向之间的夹角低于设定阈值)且图像质量较高的一帧或者多帧图像作为所述第一图像。在一些例子中,还可基于图像中包括的目标对象的状态进行选帧操作得到第一图像。其中,可选地,目标对象的状态包括以下因素中的一种或多种因素的任意组合:图像中的人脸朝向是否正面朝向、是否处于闭眼状态、是否处于张嘴状态、是否出现运动模糊或者对焦模糊等,本公开实施例对此不做限定。
在一些可选的实施方式,还可以对第一图像传感器采集到的第一视频流和第二图像传感器采集到的第二视频流进行联合选帧,得到包括第一图像和第二图像的图像对。此时,从双目摄像头采集到的视频流中选择图像对,选择的图像对中包括的两个图像均满足设定条件,该设定条件的具体实现可以参见上文描述,为了简洁,这里不再赘述。
在本公开的一些可选实施例中,在对第一图像和第二图像进行双目匹配处理(例如从第一图像 中截取第一目标区域图像、从第二图像中截取第二目标区域图像)之前,还可以对第一图像和第二图像进行校正处理,以使得第一图像和第二图像中的对应像素点位于同一水平线上。作为一种实施方式,可以基于标定得到的双目摄像头的参数,对第一图像和第二图像进行双目校正处理,例如,基于第一图像传感器的参数、第二图像传感器的参数以及第一图像传感器和第二图像传感器之间的相对位置参数,对第一图像和第二图像进行双目校正处理。作为另一种实施方式,也可以在不依赖于双目摄像头的参数的情况下对第一图像和第二图像进行自动校正,例如,获取目标对象在第一图像中的关键点信息(即第一关键点信息)和所述目标对象在第二图像中的关键点信息(即第二关键点信息),并基于第一关键点信息和第二关键点信息,确定目标变换矩阵(例如利用最小二乘法确定目标变换矩阵),再基于目标变换矩阵对第一图像或第二图像进行变换处理,得到变换后的第一图像或第二图像,但本公开实施例对此不做限定。
在一些实施例中,第一目标区域图像和第二目标区域图像中的对应像素点位于同一水平线上。例如,可以基于第一图像传感器和第二图像传感器的参数,对第一图像和第二图像中的至少一个进行平移和/或旋转等预处理,以使得预处理后的第一图像和第二图像上的对应像素点处于同一水平线上。再例如,双目摄像头中的两个图像传感器未进行标定,此时,可以对第一图像和第二图像进行匹配检测和校正处理,以使得校正后的第一图像和第二图像中的对应像素点处于同一水平线上,本公开实施例对此不做限定。
在一些实施例中,可以预先对双目摄像头的两个图像传感器进行标定,以得到第一图像传感器和第二图像传感器的参数。
在本公开实施例中,可以通过多种方式获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像。
在一些可选实施例中,图像处理装置可从其他设备直接获得第一目标区域图像和第二目标区域图像,其中,第一目标区域图像和第二目标区域图像是分别从第一图像和第二图像中截取的。其中,第一目标区域图像和第二目标区域图像可以携带在活体检测请求、身份认证请求、深度预测请求、双目匹配请求或其他消息中发送,本公开实施例对此不做限定。例如,图像处理装置从设置在其他设备处的数据库获取第一目标区域图像和第二目标区域图像。再例如,图像处理装置(例如服务器)接收设置有双目摄像头的终端设备发送的第一目标区域图像和第二目标区域图像,其中,可选地,终端设备可以通过双目摄像头采集包括第一图像和第二图像的静态图像对,分别从第一图像和第二图像中截取第一目标区域图像和第二目标区域图像;或者,终端设备通过双目摄像头采集视频序列,对采集到的视频序列进行选帧,得到包括第一图像和第二图像的视频帧图像对。再例如,终端设备向图像处理装置发送包括第一图像和第二图像的图像对的视频流,再分别从第一图像和第二图像中截取第一目标区域图像和第二目标区域图像,本公开实施例对此不做限定。
在另一些可选实施例中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在一些实施例中,可以对第一图像和第二图像分别进行目标检测,得到目标对象在第一图像中的第一位置信息和所述目标对象在第二图像中的第二位置信息,并基于第一位置信息从第一图像中截取第一目标区域图像,基于第二位置信息从第二图像中截取第二目标区域图像。
其中,可选地,可以直接对第一图像和第二图像进行目标检测,或者先对第一图像和/或第二图像进行预处理,对预处理后的第一图像和/或第二图像进行目标检测;其中,所述预处理例如可包括亮度调整、尺寸调整、平移、旋转等一项或多项处理,本公开实施例对此不做限定。
在本公开的一些可选实施例中,所述获取目标对象的第一目标区域图像,包括:对所述双目摄 像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;对所述第一候选区域的图像进行关键点检测,得到关键点信息;基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在一些实施例中,可以对第一图像和第二图像分别进行目标检测,得到所述第一图像中的第一候选区域,以及所述第二图像中与第一候选区域对应的第二候选区域,基于第一候选区域从第一图像中截取第一目标区域图像,基于第二候选区域从第二图像中截取第二目标区域图像。
例如,可以从第一图像中截取第一候选区域的图像作为第一目标区域图像。又例如,通过对第一候选区域放大一定倍数后得到第一目标区域,并从第一图像中截取第一目标区域的图像作为第一目标区域图像
在一些实施例中,通过对第一候选区域的图像进行关键点检测,得到第一候选区域对应的第一关键点信息,基于得到的第一关键点信息从第一图像中截取第一目标区域图像。相应的,通过对第二候选区域的图像进行关键点检测,得到第二候选区域对应的第二关键点信息,基于得到的第二关键点信息从第二图像中截取第二目标区域图像。
在一种可选的实现方式中,可通过图像处理技术(例如卷积神经网络)对第一图像进行目标检测,得到目标对象所属的第一候选区域,相应的,可通过图像处理技术(例如卷积神经网络)对第二图像进行目标检测,得到目标对象所属的第二候选区域;其中,所述第一候选区域和所述第二候选区域例如为第一人脸区域。其中,所述目标检测可以是对目标对象的大致定位;相应地,所述第一候选区域为包括目标对象的初步区域,所述第二候选区域为包括所述目标对象的初步区域。
其中,上述关键点检测可以通过深度神经网络实现,例如卷积神经网络、循环神经网络等,具体可以是LeNet、AlexNet、GoogLeNet、VGGNet、ResNet等任意类型的神经网络模型;或者,关键点检测也可以是基于其他机器学习方法实现,本公开实施例对关键点检测的具体实现不作限定。
其中,关键点信息可以包括目标对象的多个关键点中每个关键点的位置信息,或者进一步包括置信度等信息,本公开实施例对此不做限定。
举例来说,在所述目标对象为人脸的情况下,利用人脸关键点检测模型,分别对所述第一候选区域和第二候选区域的图像进行人脸关键点检测,得到所述第一候选区域的图像中包括对应于人脸关键点的多个关键点信息,得到所述第二候选区域的图像中包括对应于人脸关键点的多个关键点信息,基于该多个关键点信息可确定人脸的位置信息,基于人脸的位置信息可确定对应于人脸的第一目标区域,以及对应于所述人脸的第二目标区域。与第一候选区域和第二候选区域相比,第一目标区域和第二目标区域为人脸的较为准确的位置,从而有利于提高后续操作的精确度。
在上述各个实施例中对第一图像和第二图像进行的目标检测不需要确定目标对象或其所属区域的精确位置,只需要大致定位目标对象或其所属区域即可,从而降低对目标检测算法的精确度要求,提高鲁棒性和图像处理速度。
在一些可能的实现方式中,所述第二目标区域图像的截取方式与所述第一目标区域图像的截取方式可以相同,也可以不相同,本公开实施例对此不做限定。
在本公开实施例中,可选地,所述第一目标区域图像和所述第二目标区域图像的图像可以具有不同的尺寸。或者,为了降低计算复杂度,进一步提高处理速度,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在一些实施例中,可以利用表征相同尺寸的截取框截取参数分别从第一图像和第二图像中截取第一目标区域图像和第二目标区域图像,以使得所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。例如,在上述例子中,可以基于目标对象的第一位置信息和第二位置信息,得到完全包括目标对象的两个具有相同的截取框。再例如,在上述例子中,可以对第一图像和第二图像进行目标检测,以使得到的对应于第一图像的第一截取框和对应于第二图像的第二截取框具有相同的 大小。再例如,在上述例子中,如果第一截取框和第二截取框具有不同的大小,则分别对第一截取框和第二截取框放大不同的倍数,也即分别对第一截取框对应的第一截取参数和第二截取框对应的第二截取参数进行不同倍数的放大处理,以使得放大处理得到的两个截取框具有相同的尺寸。再例如,在上述例子中,基于第一图像的关键点信息和第二图像的关键点信息,确定具有相同尺寸的第一目标区域和第二目标区域,其中,第一目标区域和第二目标区域完全包括目标对象,等等。
在本公开实施例中,通过对第一图像和第二图像进行目标对象的检测,去除目标对象或目标区域以外的无关信息,从而减少后续双目匹配算法的输入图像的大小和处理的数据量,加快了图像视差的预测速度。在一些实施方式中,在活体检测领域中,可通过预测图像的视差从而获取图像的深度信息,进而确定该图像包括的人脸是否为活体人脸。基于此,只需要关注图像的人脸区域,因此如果仅仅对图像的人脸区域进行视差预测,能避免不必要的计算,从而提高视差预测的速度。
102、对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差。
在本公开的一种可选实施例中,针对步骤102,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差,包括:通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像的视差。
本实施方式通过双目匹配神经网络对第一目标区域图像和第二目标区域图像进行处理,得到第一目标区域图像和第二目标区域图像之间的视差并输出。
在一些可选实施方式中,将第一目标区域图像和第二目标区域图像直接输入到双目匹配神经网络中进行处理,得到第一目标区域图像和第二目标区域图像之间的视差。在另一些可选实施方式中,可先对第一目标区域图像和/或第二目标区域图像进行预处理,所述预处理例如转正处理等,再将预处理后的第一目标区域图像和第二目标区域图像输入到双目匹配神经网络中进行处理,得到第一目标区域图像和第二目标区域图像之间的视差。本公开实施例对此不做限定。
参见图2,图2是本公开实施例提供的确定第一目标区域图像和第二目标区域图像的视差的的示意图,其中,将第一目标区域图像和第二目标区域图像输入到所述双目匹配神经网络中,通过所述双目匹配神经网络,分别提取所述第一目标区域图像的第一特征(即图2中的特征1)和第二目标区域图像的第二特征(即图2中的特征2),通过双目匹配神经网络中的匹配代价计算模块计算第一特征和第二特征的匹配代价,基于得到的匹配代价确定所述第一目标区域图像和所述第二目标区域图像之间的视差;其中,所述匹配代价可表示第一特征和第二特征的相关性。其中,所述基于得到的匹配代价确定所述第一目标区域图像和所述第二目标区域图像之间的视差,包括:对匹配代价进行特征提取,基于提取到的特征数据确定第一目标区域图像和第二目标区域图像之间的视差。
在另一些可选的实现方式中,针对步骤102,可以通过其他基于机器学习的双目匹配算法确定所述第一目标区域图像和所述第二目标区域图像之间的视差。实际应用中,所述双目匹配算法可以是如下算法中的任意一种:立体双目视觉算法(Sum of absolute differences SAD)、双向匹配算法(bidirectional matching BM)、全局匹配算法(Semi-global block matching SGBM)、图割算法(Graph Cuts,GC),本公开实施例中对双目匹配处理的具体实现不做限定。
103、基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
在本公开的一些可选实施例中,在执行步骤103之前,即在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,所述方法还包括:基于第一目标区域图像在第一图像中的位置和第二目标区域图像在第二图像中的位置,确定第一目标区域图像和 第二目标区域图像的位移信息。可选地,该位移信息可以包括水平方向上的位移和/或垂直方向上的位移,其中,在一些实施例中,如果第一图像和第二图像中的对应像素点位于同一水平线上,则该位移信息可以仅包括水平方向上的位移,但本公开实施例对此不做限定。
其中,所述基于第一目标区域图像在第一图像中的位置和第二目标区域图像在第二图像中的位置,确定第一目标区域图像和第二目标区域图像的位移信息,包括:确定所述第一目标区域图像的第一中心点位置,确定所述第二目标区域图像的第二中心点位置;基于所述第一中心点的位置和所述第二中心点的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
参见图3,图3是本公开实施例提供的目标区域位移确定方法的示例性示意图,第一图像中的第一目标区域图像的中心点a的位置表示为(x 1,y 1),第二图像中的第二目标区域图像的中心点b的位置表示为(x 2,y 1),中心点a和中心点b之间的位移表示为
Figure PCTCN2019107362-appb-000001
即为所述第一目标区域图像和所述第一目标区域图像之间的位移信息。在另一可能的实现方式中,上述中心点可以使用目标区域图像四个顶点中任意一个顶点来代替,本公开实施例中对此不作具体限定。
在本公开实施例中,还可以通过其他方式确定第一目标区域图像和第二目标区域图像之间的位移信息,本公开实施例对此不做限定。
在本公开的一些可选实施例中,针对步骤103,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果,包括:将所述第一目标区域图像和所述第二目标区域图像之间的视差和位移信息相加,得到所述第一图像和所述第二图像之间的视差预测结果。
例如,所述第一目标区域图像和所述第二目标区域图像之间的位移信息为x,所述第一目标区域图像和所述第二目标区域图像的视差为D(p),将位移信息为x和视差D(p)相加或相减得到的结果,作为所述第一图像和所述第二图像之间的视差预测结果。
在一些实施例中,第一目标区域图像和第二目标区域图像之间的位移为0,则第一目标区域图像和第二目标区域图像之间的视差即为第一图像和第二图像之间的视差。
在一些可能的实现方式中,所述位移信息的确定和所述第一目标区域图像和第二目标区域图像之间的视差的确定可以并行执行,或者以任意前后顺序执行,本公开实施例对位移信息的确定和所述第一目标区域图像和第二目标区域图像之间的视差的确定的执行顺序不做限定。
在本公开的一种可选实施例中,步骤103之后,所述方法还包括:在得到第一图像和第二图像的视差预测结果之后,基于所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;基于所述目标对象的深度信息,确定活体检测结果。
在本公开实施例中,获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像;对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。本公开实施例能够减少视差预测的计算量,从而提高视差的预测速度,有利于实现视差的实时预测。
应理解,上文以视差预测为例对本公开实施例的技术方案进行了描述,可选地,本公开实施例的技术方案也可以应用于其他应用场景,例如,光流预测,此时,第一图像和第二图像分别为单目摄像头在不同时刻采集到的图像,等等,本公开实施例对此不做限定。
图4是本公开实施例提供的图像处理方法的示意流程图。
201、获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;
202、对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息;
203、基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
本公开实施例中,可通过单目摄像头对所述图像采集区域进行图像采集,基于不同时刻采集的图像获得第一目标区域图像和第二目标区域图像。作为一种示例,在第一时刻采集到的图像记为第一图像,从第一图像中获得第一目标区域图像;在第二时刻采集到的图像记为第二图像,从第二图像中获得第二目标区域图像。
在本公开的一些可选实施例中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
其中,一种实施方式是,所述获取目标对象的第一目标区域图像,包括:对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;对所述第一候选区域的图像进行关键点检测,得到关键点信息;基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
本公开实施例中,可选地,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
本公开实施例中,针对步骤201的相关描述可参照前述实施例中针对步骤101的详细描述,这里不再赘述。
在本公开的一些可选实施例中,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息,包括:将通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
这样,可通过神经网络对第一目标区域图像和第二目标区域图像进行处理,得到第一目标区域图像和第二目标区域图像之间的光流信息。
在一些可选实施方式中,可将第一目标区域图像和第二目标区域图像输入至神经网络中进行处理,得到第一目标区域图像和第二目标区域图像之间的光流信息;在另一些可选实施方式中,可先对第一目标区域图像和/或第二目标区域图像进行预处理,所述预处理例如转正处理等,再将预处理后的第一目标区域图像和第二目标区域图像输入到神经网络中,得到第一目标区域图像和第二目标区域图像之间的光流信息。其中,由于第一目标区域图像和第二目标区域图像对应的位置并非绝对不变的,因此所述光流信息为一个相对的概念,可表征所述目标对象的相对光流信息,也即所述目标对象的相对运动情况。
在本公开的一些可选实施例中,在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,所述方法还包括:基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
本公开实施例中,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息的相关描述具体可参照前述实施例中所述,这里不再所述。
在本公开的一些可选实施例中,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间光流信息预测结果,包括:将所述第一目标区域图像和所述第二目标区域图像 之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
本公开实施例中,由于第一目标区域图像和第二目标区域图像对应的位置并非绝对不变的,因此需要确定所述第一目标区域图像和所述第二目标区域图像的位移信息,再将所述位移信息和所述光流信息相加或相减得到光流信息预测结果。其中,所述光流信息预测结果可表示目标对象的绝对光流信息,也即所述目标对象的绝对运动情况。
本公开实施例的图像处理方法应用于光流信息预测,图1描述的图像处理方法应用于视差信息预测,两者在技术实现上基本一致,为了简洁,本公开实施例的图像处理方法的具体实现可以参照图1描述的图像处理方法实施例的描述,这里不再赘述。
本公开实施例还提供了图像处理装置。图5是本公开实施例提供的图像处理装置的结构示意图一。该装置500包括:获取单元501、第一确定单元502和第二确定单元503;其中,
所述获取单元501,配置为获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从双目摄像头的第一图像传感器采集到的第一图像中截取的,所述第二目标区域图像是从所述双目摄像头的第二图像传感器采集到的第二图像中截取的;
所述第一确定单元502,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;
所述第二确定单元503,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
在本公开一些可选实施例中,所述获取单元501配置为,获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在本公开一些可选实施例中,参见图6,所述获取单元501包括目标检测单元501-1、关键点检测单元501-2和截取单元501-3,所述目标检测单元501-1,配置为对所述双目摄像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;所述关键点检测单元501-2,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;所述截取单元501-3,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在本公开一些可选实施例中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在本公开一些可选实施例中,所述第一确定单元502配置为,通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的视差。
在本公开一些可选实施例中,参见图7,所述装置还包括位移确定单元701,所述位移确定单元701配置为,在所述第二确定单元503基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在本公开一些可选实施例中,所述第二确定单元503配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和第二目标区域图像之间的视差相加,得到所述第一图像和所述第二图像之间的视差预测结果。
在本公开一些可选实施例中,参见图7,所述装置还包括深度信息确定单元702和活体检测确定单元703,所述深度信息确定单元702,配置为基于所述第二确定单元503获得的所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;所述活体检测确定单元703,配置 为基于所述深度信息确定单元702获得的所述目标对象的深度信息,确定活体检测结果。
在本公开一些可选实施例中,所述双目摄像头包括同模态双目摄像头和跨模态双目摄像头中的一种。
在本公开一些可选实施例中,所述第一图像传感器或所述第二图像传感器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。
在本公开一些可选实施例中,所述目标对象包括人脸。
本公开实施例提供的装置具有的功能或包含的模块可以用于执行上文图像处理方法实施例描述的方法,其具体实现可以参照上文方法实施例的描述,为了简洁,这里不再赘述。
本公开实施例还提供了图像处理装置。图8是本公开实施例提供的图像处理装置的结构示意图四。该装置800包括:获取单元801、第一确定单元802和第二确定单元803;其中,
所述获取单元801,配置为获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;
所述第一确定单元802,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息;
所述第二确定单元803,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
在本公开一些可选实施例中,所述获取单元801配置为:获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
在本公开一些可选实施例中,所述获取单元801包括目标检测单元、关键点检测单元和截取单元;其中,
所述目标检测单元,配置为对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;
所述关键点检测单元,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;
所述截取单元,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
在本公开一些可选实施例中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
在本公开一些可选实施例中,所述第一确定单元802配置为,通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
在本公开一些可选实施例中,所述装置还包括位移确定单元,所述位移确定单元配置为,在所述第二确定单元803基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
在本公开一些可选实施例中,所述第二确定单元803配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
本实施例的图像处理装置应用于光流信息预测,本公开实施例提供的装置具有的功能或包含的模块可以用于执行图4所示的方法实施例描述的方法,其具体实现可以参照图4图像处理方法实施 例的描述,为了简洁,这里不再赘述。
另外,本公开实施例提供了一种电子设备,图9是本公开实施例提供的电子设备的结构框图。如图9所示,该电子设备包括:处理器901,以及用于存储处理器可执行指令的存储器904,其中,所述处理器901被配置为:执行本公开实施例如图1所示的图像处理方法或其任意可能的实现方式;或者执行本公开实施例如图4所示的图像处理方法或其任意可能的实现方式。
可选地,所述电子设备还可以包括:一个或多个输入设备902和一个或多个输出设备903。
上述处理器901、输入设备902、输出设备903和存储器904通过总线905连接。存储器902用于存储指令,处理器901用于执行存储器902存储的指令。其中,处理器901被配置用于调用所述程序指令执行上文图像处理方法中任一实施例,为了简洁,这里不再赘述。
应理解,上文装置实施例以视差预测为例对本公开实施例的技术方案进行了描述。可选地,本公开实施例的技术方案也可以应用于光流预测,相应地,光流预测装置同样属于本公开保护范围,光流预测装置与上文描述的图像处理装置相似,为了简洁,这里不再赘述。
应当理解,在本公开实施例中,所称处理器901可以是中央处理单元(Central Processing Unit,CPU),该处理器还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
输入设备902可以包括移动手机、台式电脑、膝上计算机、可穿戴设备、监控图像传感器等,输出设备903可以包括显示器(LCD等)。
该存储器904可以包括只读存储器和随机存取存储器,并向处理器901提供指令和数据。存储器904的一部分还可以包括非易失性随机存取存储器。例如,存储器904还可以存储设备类型的信息。
本公开实施例中所描述的电子设备用于执行上文描述的图像处理方法,相应地,处理器901用于执行本公开实施例提供的图像处理方法的各个实施例中的步骤和/或流程,在此不再赘述。
在本公开的另一实施例中提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序包括程序指令,所述程序指令被处理器执行时实现上文图像处理方法中任一实施例,为了简洁,这里不再赘述。
所述计算机可读存储介质可以是前述任一实施例所述的电子设备的内部存储单元,例如终端的硬盘或内存。所述计算机可读存储介质也可以是所述终端的外部存储设备,例如所述终端上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,所述计算机可读存储介质还可以既包括所述电子设备的内部存储单元也包括外部存储设备。所述计算机可读存储介质用于存储所述计算机程序以及所述电子设备所需的其他程序和数据。所述计算机可读存储介质还可以用于暂时地存储已经输出或者将要输出的数据。
本领域普通技术人员可以意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、计算机软件或者二者的结合来实现,为了清楚地说明硬件和软件的可互换性,在上述说明中已经按照功能一般性地描述了各示例的组成及步骤。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本公开的范围。
所属领域的技术人员可以清楚地了解到,为了描述的方便和简洁,上述描述的服务器、设备和单元的具体工作过程,可以参考前述方法实施例中的对应过程,也可执行发明实施例所描述的电子设备的实现方式,在此不再赘述。
在本公开所提供的几个实施例中,应该理解到,所揭露的服务器、设备和方法,可以通过其它 的方式实现。例如,以上所描述的服务器实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另外,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口、装置或单元的间接耦合或通信连接,也可以是电的,机械的或其它的形式连接。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本公开实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以是两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分,或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本公开的具体实施方式,但本公开的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,可轻易想到各种等效的修改或替换,这些修改或替换都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应以权利要求的保护范围为准。

Claims (39)

  1. 一种图像处理方法,包括:
    获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从双目摄像头的第一图像传感器采集到的第一图像中截取的,所述第二目标区域图像是从所述双目摄像头的第二图像传感器采集到的第二图像中截取的;
    对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差;
    基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
  2. 根据权利要求1所述的方法,其中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:
    获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;
    对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
  3. 根据权利要求1或2所述的方法,其中,所述获取目标对象的第一目标区域图像,包括:
    对所述双目摄像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;
    对所述第一候选区域的图像进行关键点检测,得到关键点信息;
    基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
  4. 根据权利要求1-3任一项所述的方法,其中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
  5. 根据权利要求1所述的方法,其中,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的视差,包括:
    通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的视差。
  6. 根据权利要求1所述的方法,其中,在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,所述方法还包括:
    基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
  7. 根据权利要求1或6任一项所述的方法,其中,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果,包括:
    将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述视差相加,得到所述第一图像和所述第二图像之间的视差预测结果。
  8. 根据权利要求1-7中任一项所述的方法,其中,所述方法还包括:
    基于所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;
    基于所述目标对象的深度信息,确定活体检测结果。
  9. 根据权利要求1-8中任一项所述的方法,其中,所述双目摄像头包括同模态双目摄像头和跨模态双目摄像头中的一种。
  10. 根据权利要求1-9中任一项所述的方法,其中,所述第一图像传感器或所述第二图像传感 器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。
  11. 根据权利要求1-10中任一项所述的方法,其中,所述目标对象包括人脸。
  12. 一种图像处理方法,包括:
    获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;
    对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息;
    基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
  13. 根据权利要求12所述的方法,其中,所述获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,包括:
    获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;
    对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
  14. 根据权利要求12或13所述的方法,其中,所述获取目标对象的第一目标区域图像,包括:
    对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;
    对所述第一候选区域的图像进行关键点检测,得到关键点信息;
    基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
  15. 根据权利要求12-14任一项所述的方法,其中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
  16. 根据权利要求12所述的方法,其中,所述对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述第一目标区域图像和所述第二目标区域图像之间的光流信息,包括:
    通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
  17. 根据权利要求12所述的方法,其中,在所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,所述方法还包括:
    基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
  18. 根据权利要求12或17任一项所述的方法,其中,所述基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间光流信息预测结果,包括:
    将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
  19. 一种图像处理装置,包括:
    获取单元,配置为获取目标对象的第一目标区域图像和所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从双目摄像头的第一图像传感器采集到的第一图像中截取的,所述第二目标区域图像是从所述双目摄像头的第二图像传感器采集到的第二图像中截取的;
    第一确定单元,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述 第一目标区域图像和所述第二目标区域图像之间的视差;
    第二确定单元,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果。
  20. 根据权利要求19所述的装置,其中,所述获取单元配置为:获取所述双目摄像头中的第一图像传感器采集的第一图像和所述双目摄像头中的第二图像传感器采集的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
  21. 根据权利要求19或20所述的装置,其中,所述获取单元包括目标检测单元、关键点检测单元和截取单元,
    所述目标检测单元,配置为对所述双目摄像头中的第一图像传感器采集的第一图像进行目标检测,得到第一候选区域;
    所述关键点检测单元,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;
    所述截取单元,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
  22. 根据权利要求19-21任一项所述的装置,其中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
  23. 根据权利要求19所述的装置,其中,所述第一确定单元配置为,通过双目匹配神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的视差。
  24. 根据权利要求19所述的装置,其中,所述装置还包括位移确定单元,所述位移确定单元配置为,在所述第二确定单元基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的视差,得到所述第一图像和所述第二图像之间的视差预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
  25. 根据权利要求19或24任一项所述的装置,其中,所述第二确定单元配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述视差相加,得到所述第一图像和所述第二图像之间的视差预测结果。
  26. 根据权利要求19-25中任一项所述的装置,其中,所述装置还包括深度信息确定单元和活体检测确定单元,
    所述深度信息确定单元,配置为基于所述第一图像和所述第二图像的视差预测结果,确定所述目标对象的深度信息;
    所述活体检测确定单元,配置为基于所述目标对象的深度信息,确定活体检测结果。
  27. 根据权利要求19-26中任一项所述的装置,其中,所述双目摄像头包括同模态双目摄像头和跨模态双目摄像头中的一种。
  28. 根据权利要求19-27中任一项所述的装置,其中,所述第一图像传感器或所述第二图像传感器包括如下图像传感器中的其中一种:可见光图像传感器、近红外图像传感器、双通图像传感器。
  29. 根据权利要求19-28中任一项所述的装置,其中,所述目标对象包括人脸。
  30. 一种图像处理装置,包括:
    获取单元,配置为获取目标对象的第一目标区域图像以及所述目标对象的第二目标区域图像,其中,所述第一目标区域图像是从第一时刻对图像采集区域采集到的第一图像中截取的,所述第二目标区域图像是从第二时刻对所述图像采集区域采集到的第二图像中截取的;
    第一确定单元,配置为对所述第一目标区域图像和所述第二目标区域图像进行处理,确定所述 第一目标区域图像和所述第二目标区域图像之间的光流信息;
    第二确定单元,配置为基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果。
  31. 根据权利要求30所述的装置,其中,所述获取单元配置为:获取所述第一时刻对图像采集区域采集到的第一图像和所述第二时刻对所述图像采集区域采集到的第二图像;对所述第一图像和所述第二图像分别进行目标检测,得到第一目标区域图像和第二目标区域图像。
  32. 根据权利要求30或31所述的装置,其中,所述获取单元包括目标检测单元、关键点检测单元和截取单元;其中,
    所述目标检测单元,配置为对所述第一时刻对图像采集区域采集到的第一图像进行目标检测,得到第一候选区域;
    所述关键点检测单元,配置为对所述第一候选区域的图像进行关键点检测,得到关键点信息;
    所述截取单元,配置为基于所述关键点信息,从所述第一图像中截取第一目标区域图像。
  33. 根据权利要求30-32任一项所述的装置,其中,所述第一目标区域图像和所述第二目标区域图像的图像尺寸相同。
  34. 根据权利要求30所述的装置,其中,所述第一确定单元配置为,通过神经网络对所述第一目标区域图像和所述第二目标区域图像进行处理,得到所述第一目标区域图像和所述第二目标区域图像之间的光流信息。
  35. 根据权利要求30所述的装置,其中,所述装置还包括位移确定单元,所述位移确定单元配置为,在所述第二确定单元基于所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述第一目标区域图像和所述第二目标区域图像之间的光流信息,得到所述第一图像和所述第二图像之间的光流信息预测结果之前,基于所述第一目标区域图像在所述第一图像中的位置和所述第二目标区域图像在所述第二图像中的位置,确定所述第一目标区域图像和所述第二目标区域图像之间的位移信息。
  36. 根据权利要求30或35任一项所述的装置,其中,所述第二确定单元配置为,将所述第一目标区域图像和所述第二目标区域图像之间的位移信息和所述光流信息相加,得到所述第一图像和所述第二图像之间的光流信息预测结果。
  37. 一种电子设备,包括:
    处理器;
    用于存储计算机可读指令的存储器;
    其中,所述处理器用于调用所述存储器存储的计算机可读指令,以执行权利要求1-11中任一项所述的方法;或者,以执行权利要求12-18中任一项所述的方法。
  38. 一种计算机可读存储介质,其上存储有计算机程序指令,所述计算机程序指令被处理器执行时实现权利要求1-11中任一项所述的方法;或者,所述计算机程序指令被处理器执行时实现权利要求12-18中任一项所述的方法。
  39. 一种计算机程序产品,包括计算机指令,所述计算机指令被处理器执行时实现权利要求1-11中任一项所述的方法;或者,所述计算机指令被处理器执行时实现权利要求12-18中任一项所述的方法。
PCT/CN2019/107362 2018-12-29 2019-09-23 图像处理方法、装置、电子设备及计算机可读存储介质 Ceased WO2020134229A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
SG11202010402VA SG11202010402VA (en) 2018-12-29 2019-09-23 Image processing method, device, electronic apparatus, and computer readable storage medium
US17/048,823 US20210150745A1 (en) 2018-12-29 2019-09-23 Image processing method, device, electronic apparatus, and computer readable storage medium
JP2020556853A JP7113910B2 (ja) 2018-12-29 2019-09-23 画像処理方法及び装置、電子機器並びにコンピュータ可読記憶媒体

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201811647485.8 2018-12-29
CN201811647485.8A CN111383256B (zh) 2018-12-29 2018-12-29 图像处理方法、电子设备及计算机可读存储介质

Publications (1)

Publication Number Publication Date
WO2020134229A1 true WO2020134229A1 (zh) 2020-07-02

Family

ID=71128548

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/107362 Ceased WO2020134229A1 (zh) 2018-12-29 2019-09-23 图像处理方法、装置、电子设备及计算机可读存储介质

Country Status (5)

Country Link
US (1) US20210150745A1 (zh)
JP (1) JP7113910B2 (zh)
CN (1) CN111383256B (zh)
SG (1) SG11202010402VA (zh)
WO (1) WO2020134229A1 (zh)

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111260711B (zh) * 2020-01-10 2021-08-10 大连理工大学 一种弱监督可信代价传播的视差估计方法
WO2021158174A1 (en) * 2020-02-07 2021-08-12 Agency For Science, Technology And Research Active infrared thermography system and computer-implemented method for generating thermal image
CN112016558B (zh) * 2020-08-26 2024-05-31 大连信维科技有限公司 一种基于图像质量的介质能见度识别方法
CN114298912B (zh) * 2022-03-08 2022-10-14 北京万里红科技有限公司 图像采集方法、装置、电子设备及存储介质
CN116416292B (zh) * 2022-07-13 2025-03-25 上海砹芯科技有限公司 图像深度确定方法、装置、电子设备及存储介质
CN115713752A (zh) * 2022-11-07 2023-02-24 中汽创智科技有限公司 一种危险驾驶行为检测方法、装置、设备及存储介质
US12462562B2 (en) 2022-12-23 2025-11-04 Andy Tsz Kwan CHAN Multipurpose visual and auditory intelligent observer system

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20070063063A (ko) * 2005-12-14 2007-06-19 주식회사 제이앤에이치테크놀러지 스테레오 비전 시스템
CN103679707A (zh) * 2013-11-26 2014-03-26 西安交通大学 基于双目相机视差图的道路障碍物检测系统及检测方法
CN107545247A (zh) * 2017-08-23 2018-01-05 北京伟景智能科技有限公司 基于双目识别的立体认知方法

Family Cites Families (23)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8818024B2 (en) * 2009-03-12 2014-08-26 Nokia Corporation Method, apparatus, and computer program product for object tracking
US20170206427A1 (en) * 2015-01-21 2017-07-20 Sportstech LLC Efficient, High-Resolution System and Method to Detect Traffic Lights
JP2014099716A (ja) * 2012-11-13 2014-05-29 Canon Inc 画像符号化装置及びその制御方法
AU2014239979B2 (en) * 2013-03-15 2017-06-22 Aurora Operations, Inc. Methods, systems, and apparatus for multi-sensory stereo vision for robotics
CN104217208B (zh) * 2013-06-03 2018-01-16 株式会社理光 目标检测方法和装置
CN105095905B (zh) * 2014-04-18 2018-06-22 株式会社理光 目标识别方法和目标识别装置
US11501406B2 (en) * 2015-03-21 2022-11-15 Mine One Gmbh Disparity cache
JP6732440B2 (ja) * 2015-12-04 2020-07-29 キヤノン株式会社 画像処理装置、画像処理方法、及びそのプログラム
GB2553782B (en) * 2016-09-12 2021-10-20 Niantic Inc Predicting depth from image data using a statistical model
CN106651923A (zh) * 2016-12-13 2017-05-10 中山大学 一种视频图像目标检测与分割方法及系统
CN108537871B (zh) * 2017-03-03 2024-02-20 索尼公司 信息处理设备和信息处理方法
JP2018189443A (ja) * 2017-04-28 2018-11-29 キヤノン株式会社 距離測定装置、距離測定方法及び撮像装置
JP6922399B2 (ja) * 2017-05-15 2021-08-18 日本電気株式会社 画像処理装置、画像処理方法及び画像処理プログラム
US10244164B1 (en) * 2017-09-11 2019-03-26 Qualcomm Incorporated Systems and methods for image stitching
CN107886120A (zh) * 2017-11-03 2018-04-06 北京清瑞维航技术发展有限公司 用于目标检测跟踪的方法和装置
CN108154520B (zh) * 2017-12-25 2019-01-08 北京航空航天大学 一种基于光流与帧间匹配的运动目标检测方法
CN108335322B (zh) * 2018-02-01 2021-02-12 深圳市商汤科技有限公司 深度估计方法和装置、电子设备、程序和介质
CN108446622A (zh) * 2018-03-14 2018-08-24 海信集团有限公司 目标物体的检测跟踪方法及装置、终端
CN108520536B (zh) * 2018-03-27 2022-01-11 海信集团有限公司 一种视差图的生成方法、装置及终端
CN111444744A (zh) * 2018-12-29 2020-07-24 北京市商汤科技开发有限公司 活体检测方法、装置以及存储介质
CN109887019B (zh) * 2019-02-19 2022-05-24 北京市商汤科技开发有限公司 一种双目匹配方法及装置、设备和存储介质
CN115222782B (zh) * 2021-04-16 2025-08-08 安霸国际有限合伙企业 对单目相机立体系统中的结构光投射器的安装校准
CN115690469B (zh) * 2021-07-30 2026-02-03 北京原创世代科技有限公司 一种双目图像匹配方法、装置、设备和存储介质

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20070063063A (ko) * 2005-12-14 2007-06-19 주식회사 제이앤에이치테크놀러지 스테레오 비전 시스템
CN103679707A (zh) * 2013-11-26 2014-03-26 西安交通大学 基于双目相机视差图的道路障碍物检测系统及检测方法
CN107545247A (zh) * 2017-08-23 2018-01-05 北京伟景智能科技有限公司 基于双目识别的立体认知方法

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
LICHUAN GENG ET AL: "Optical Flow Estimation Algorithm Based on Shift Wavelet Transform", GEOMATICS SCIENCE AND TECHNOLOGY, vol. 5, no. 2, 28 April 2017 (2017-04-28), pages 40 - 46, XP055715574, ISSN: 2329-549X, DOI: 10.12677/GST.2017.52006 *
LIU YU; LIU CHANLAO; SU HAI: "Research on Matching Algorithm Based on Structured Light Binocular Vision Feature", OPTICAL INSTRUMENTS, vol. 36, no. 2, 30 April 2014 (2014-04-30), pages 161 - 166, XP009521822, DOI: 10.3969/j.issn.1005-5630.201.02.015 *

Also Published As

Publication number Publication date
JP7113910B2 (ja) 2022-08-05
US20210150745A1 (en) 2021-05-20
JP2021519983A (ja) 2021-08-12
CN111383256A (zh) 2020-07-07
CN111383256B (zh) 2024-05-17
SG11202010402VA (en) 2020-11-27

Similar Documents

Publication Publication Date Title
CN111383256B (zh) 图像处理方法、电子设备及计算机可读存储介质
US10922529B2 (en) Human face authentication method and apparatus, and storage medium
CN111160232B (zh) 正面人脸重建方法、装置及系统
TWI721786B (zh) 人臉校驗方法、裝置、伺服器及可讀儲存媒介
CN111383255B (zh) 图像处理方法、装置、电子设备及计算机可读存储介质
CN106981078B (zh) 视线校正方法、装置、智能会议终端及存储介质
WO2022121895A1 (zh) 双目活体检测方法、装置、设备和存储介质
WO2021136078A1 (zh) 图像处理方法、图像处理系统、计算机可读介质和电子设备
KR20210074333A (ko) 생체 검출 방법 및 장치, 저장 매체
US20220044039A1 (en) Living Body Detection Method and Device
WO2014180255A1 (zh) 一种数据处理方法、装置、计算机存储介质及用户终端
CN102572450A (zh) 基于sift特征与grnn网络的立体视频颜色校正方法
WO2020147346A1 (zh) 图像识别方法、系统及装置
CN111854620A (zh) 基于单目相机的实际瞳距测定方法、装置以及设备
CN111160233A (zh) 基于三维成像辅助的人脸活体检测方法、介质及系统
CN111046845A (zh) 活体检测方法、装置及系统
CN114387324A (zh) 深度成像方法、装置、电子设备和计算机可读存储介质
CN110800020B (zh) 一种图像信息获取方法、图像处理设备及计算机存储介质
CN112419399B (zh) 一种图像测距方法、装置、设备和存储介质
WO2026076996A1 (zh) 智能穿戴设备的合像距离调节方法、装置、设备及介质
CN111246116B (zh) 一种用于屏幕上智能取景显示的方法及移动终端
CN111931544B (zh) 活体检测的方法、装置、计算设备及计算机存储介质
CN112395912A (zh) 一种人脸分割方法、电子设备及计算机可读存储介质
HK40024028A (zh) 图像处理方法、电子设备及计算机可读存储介质
CN110610178A (zh) 图像识别方法、装置、终端及计算机可读存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19904114

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2020556853

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19904114

Country of ref document: EP

Kind code of ref document: A1