WO2018177337A1 - 手部三维数据确定方法、装置及电子设备 - Google Patents

手部三维数据确定方法、装置及电子设备 Download PDF

Info

Publication number
WO2018177337A1
WO2018177337A1 PCT/CN2018/080960 CN2018080960W WO2018177337A1 WO 2018177337 A1 WO2018177337 A1 WO 2018177337A1 CN 2018080960 W CN2018080960 W CN 2018080960W WO 2018177337 A1 WO2018177337 A1 WO 2018177337A1
Authority
WO
WIPO (PCT)
Prior art keywords
palm
hand image
finger
contour
determining
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/080960
Other languages
English (en)
French (fr)
Inventor
王权
钱晨
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Publication of WO2018177337A1 publication Critical patent/WO2018177337A1/zh
Priority to US16/451,077 priority Critical patent/US11120254B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/107Static hand or arm
    • G06V40/113Recognition of static hand signs
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/13Edge detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/74Image or video pattern matching; Proximity measures in feature spaces
    • G06V10/75Organisation of the matching processes, e.g. simultaneous or sequential comparisons of image or video features; Coarse-fine approaches, e.g. multi-scale approaches; using context analysis; Selection of dictionaries
    • G06V10/751Comparing pixel values or logical combinations thereof, or feature values having positional relevance, e.g. template matching
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/60Type of objects
    • G06V20/64Three-dimensional [3D] objects
    • G06V20/653Three-dimensional [3D] objects by matching three-dimensional models, e.g. conformal mapping of Riemann surfaces
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • G06V40/28Recognition of hand or arm movements, e.g. recognition of deaf sign language
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10004Still image; Photographic image
    • G06T2207/10012Stereo images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10028Range image; Depth image; 3D point clouds
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/107Static hand or arm
    • G06V40/117Biometrics derived from hands

Definitions

  • the present application relates to computer vision technology, and in particular to a method, device and electronic device for determining three-dimensional data of a hand.
  • the hand data is an important input data in the field of human-computer interaction. By capturing changes in the human hand data, it is possible to control the smart device by using gesture actions.
  • the embodiment of the present application provides a technical solution for determining three-dimensional data of a hand.
  • a method for determining three-dimensional data of a hand includes:
  • the key points and the area contours identified from the first hand image, and the key points and locations identified from the second hand image Determining a depth of the key point and depth information of the contour of the area;
  • the hand three-dimensional data is determined based at least on the key point and its depth information, and the region contour and its depth information.
  • a hand three-dimensional data determining apparatus includes:
  • An acquiring unit configured to acquire a first hand image and a second hand image captured by the binocular camera system
  • a recognition unit configured to respectively identify at least one key point from the first hand image and the second hand image and an area contour covering the key point
  • a depth determining unit configured to: according to imaging parameters of the binocular imaging system, the key points and the contour of the region recognized from the first hand image, and the image recognized from the second hand image Determining the depth information of the key point and the depth information of the area contour by the key point and the area contour;
  • a three-dimensional data determining unit configured to determine hand three-dimensional data based on at least the key point and its depth information, and the region contour and its depth information.
  • an electronic device includes: at least one processor; and a memory communicably coupled to the at least one processor; wherein the memory is stored with the at least one An instruction executed by the processor, the instruction being executed by the at least one processor to cause the at least one processor to perform an operation corresponding to the hand three-dimensional data determining method method of any one of the embodiments of the present application.
  • a computer readable storage medium wherein computer instructions are stored thereon, and when the instructions are executed, the method for determining a three-dimensional data of a hand according to any one of the embodiments of the present application is implemented. The operation of each step.
  • a computer program comprising computer readable code, when a computer readable code is run on a device, the processor in the device performs the implementation of the present application The instruction of each step in the method for determining the three-dimensional data of the hand according to an embodiment.
  • key points and region contours are identified in the first hand image and the second hand image captured by the binocular camera system, and then according to the binocular camera
  • the imaging parameters of the system, the key points and area contours identified from the first hand image, and the key points and area contours identified from the second hand image, the depth of the key point and the depth of the area contour can be determined; Determining the depth of the key points and the key points and the depth of the key points, and determining the hand three-dimensional data, the hand three-dimensional data is more accurate and richer, and the technical solution for determining the three-dimensional data of the hand has higher Efficiency and accuracy.
  • FIG. 1 is a flow chart of a method for determining three-dimensional data of a hand according to an embodiment of the present application
  • FIG. 2 is a schematic diagram of a key point and a region outline of a hand according to an embodiment of the present application
  • FIG. 3 is a schematic structural diagram of a three-dimensional hand data determining apparatus according to an embodiment of the present application.
  • FIG. 4 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
  • Embodiments of the present application can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with numerous other general purpose or special purpose computing system environments or configurations.
  • Examples of well-known terminal devices, computing systems, environments, and/or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, and the like include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients Machines, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, networked personal computers, small computer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above, and the like.
  • Electronic devices such as terminal devices, computer systems, servers, etc., can be described in the general context of computer system executable instructions (such as program modules) being executed by a computer system.
  • program modules may include routines, programs, target programs, components, logic, data structures, and the like that perform particular tasks or implement particular abstract data types.
  • the computer system/server can be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communication network.
  • program modules may be located on a local or remote computing system storage medium including storage devices.
  • FIG. 1 is a flowchart of a method for determining three-dimensional data of a hand according to an embodiment of the present application. As shown in FIG. 1, the hand three-dimensional data determining method of this embodiment includes the following steps:
  • the binocular camera system has two imaging devices, and two images of the measured object are acquired from different positions by using two imaging devices based on the parallax principle, and the two images of the acquired hand are called the first hand image. And the second hand image.
  • binocular camera systems There are a variety of binocular camera systems, and it is possible to use both binocular camera systems to capture two hand images.
  • the step S1 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by an acquisition unit operated by the processor.
  • the key points may include, but are not limited to, any one or more of the following: a fingertip, an knuckle point, and a palm.
  • the contour of the area covering the corresponding key point may include, but is not limited to, any one or more of the following: The outline of the finger area of the pointed and/or knuckle points, covering the contour of the palm area of the palm.
  • the key points are the fingertips (5) and the palm (1)
  • the outline of the area is the outline of the finger area covering the fingertips and the outline of the palm area covering the palm.
  • the above two hand images may be two-dimensional images, and there are various ways to identify a target that conforms to a predetermined feature from the two-dimensional image, and various methods are possible.
  • step S2 may be performed by the processor invoking a corresponding instruction stored in the memory or by an identification unit operated by the processor.
  • the two cameras of the binocular camera system capture the same object to obtain two images.
  • the object points in the two images can be matched, that is, the image is matched.
  • Image matching refers to the mapping of a point or an area contour in a three-dimensional space to the image points or area contours on the two imaging planes of the left and right cameras. After finding the matching object points, the depth of the object point can be determined according to the imaging parameters.
  • the key point is the palm of the hand.
  • the palms aL and aR can be respectively identified in the above two images. Since there are only one palm in each of the two images, the aL and aR corresponds; similarly, it is also possible to directly map the contours of the palm regions.
  • the area contour is composed of a large number of points, and this step can only process the edge points on the area outline.
  • the coordinate system may be established with the midpoint of the binocular connection of the binocular camera system as an origin, wherein the plane parallel to the imaging plane of the binocular camera system is the XY plane.
  • the direction perpendicular to this plane is the Z direction
  • the Z direction is the depth direction.
  • the depth value of the key point is the Z coordinate value in this coordinate system.
  • the step S3 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a depth determining unit operated by the processor.
  • S4 Determine hand three-dimensional data according to at least the at least one key point and the depth information thereof and the corresponding area contour and the depth information thereof.
  • the hand three-dimensional data may include, but is not limited to, any one or more of the following: three-dimensional data of key points, edge three-dimensional point cloud data of the area outline, hand direction data, and gesture information.
  • the step S4 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a three-dimensional data determining unit operated by the processor.
  • the hand three-dimensional data determining method identifies key points and region contours in the first hand image and the second hand image captured by the binocular camera system, and then can realize the key according to the characteristics of the human hand.
  • the point and the area contour are matched to determine the depth of the key point and the depth of the area contour; finally, according to the matching result of the key point and the depth of the key point, the matching result of the edge point on the area contour and the edge point on the area contour are projected
  • the depth determines the 3D data of the hand.
  • the depth can be determined by mutual matching, and the three-dimensional data thereof can be further determined, thereby realizing single frame-based
  • the two-dimensional image determines a large amount of three-dimensional information of the hand, and makes the determined three-dimensional data of the hand more accurate and richer, avoids unreasonable three-dimensional data of the hand, and has high efficiency and accuracy.
  • FIG. 2 is a schematic diagram of a key point and a region outline of a hand according to an embodiment of the present application.
  • the above step S2 can be implemented in various manners. For example, in an optional manner, the following steps can be included:
  • the palm 23 is identified in the first hand image and the second hand image, respectively.
  • the palm 23 can be identified based on line features, regional features in the image.
  • an infrared camera can be selected to directly obtain the grayscale image, and then the image is binarized, and all the pixels larger than a certain threshold are assigned a value of 255, and the remaining pixels are assigned a value of 0. , thus obtaining a line drawing.
  • step S21a may include the following steps:
  • the Manhattan distance transform may be performed on the maximum connected area contour from the edge of the two largest connected area contours respectively, that is, the breadth-first search is performed from the edge point, and the minimum Manhattan distance from each point to the edge point is recorded layer by layer. Get the Manhattan distance from each point to the edge within the outline of the area. The point furthest from the edge is then initially determined as the palm 23.
  • the step S21a may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm recognition unit operated by the processor.
  • the palm area contour 24 covering the palm 23 is determined based on the position of the recognized palm 23 in the first hand image and the second hand image, respectively.
  • the palm radius is determined according to the distance from the palm 23 position to the edge of the two largest connected regions, respectively.
  • the point closest to the palm 23 in the edge point of the largest connected area may be found, and the Euclidean distance between the nearest point and the palm 23 is initially determined as the palm radius;
  • the palm area contour 24 is determined according to the radius of the palm, respectively.
  • an outline of a region formed by a point having a radius of less than 1.5 times the Euclidean distance of the palm 23 can be marked as a palm region contour 24 within the outline of the hand region.
  • the maximum connected area is found in the hand image after the binarization processing, and then the palm 23 and the palm area outline 24 are determined based on the maximum connected area.
  • the present scheme finds the maximum connected area by the line feature, and then determines based on the maximum connected area.
  • the palm of the hand can reduce the effects of image color and gray scale, thereby improving processing efficiency.
  • the step S22a may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm recognition unit operated by the processor.
  • step S2 may include the following steps:
  • the palm 23 is identified in the first hand image and the second hand image, respectively.
  • the palm 23 is first determined, and then the fingertip 21 and the finger region contour 22 are determined based on the palm 23.
  • This step may be the same as step S21a above, and will not be described again here;
  • the step S21b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm recognition unit operated by the processor.
  • the maximum connected area is determined in each of the two images, whereby the fingertip 21 and the finger area outline 22 can be further sought based on the maximum connected area.
  • the step S22b may include the following steps:
  • S22b1 determines the point farthest from the palm 23 in the maximum connected area.
  • a point furthest from the palm 23 can be found in the maximum connected area, and from the farthest point, a breadth-first extension is made within the contour of the largest connected area to obtain a suspected finger area outline.
  • the fingertip 21 is determined from the point farthest from the palm 23 according to the shape of the contour of the suspected finger region and its position is acquired.
  • the step S22b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a fingertip recognition unit operated by the processor.
  • Steps S22b3 and S23b are: determining whether the shape of the contour of the suspected finger region is a finger based on the shape feature, and if the shape of the contour of the suspected finger region is a finger, the corresponding farthest point (starting point) is initially determined as the fingertip.
  • the suspect finger region contour is determined as the finger region contour 22. If the shape of the suspected finger area is not a finger, the outline of the suspected finger area is removed, and the remaining suspected finger area contours are continuously determined (this step is repeated a plurality of times) until all the suspected finger area contours are determined, and all the finger areas are found.
  • the contour 22 and the corresponding fingertip 21 are then taken to their position.
  • the palm is determined in the hand image after the binarization processing, and then the point farthest from the palm is determined in the point on the edge of the maximum connected area, and then the contour of the area suspected of the finger is found based on the starting point.
  • the outline of the area suspected of the finger further excludes the outline of the non-finger area according to the contour shape to select the outline of the finger area, thereby eliminating the influence of other areas such as the wrist on the recognition process, and determining the corresponding starting point after finding the outline of the finger area.
  • accuracy and processing efficiency can be improved.
  • the step S23b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a finger recognition unit operated by the processor.
  • S24b corrects the position of the fingertip 21 and determines the finger root position according to each finger region contour 22 in the first hand image and the second hand image, respectively.
  • Step S24b optionally includes the following steps:
  • S24b1 performs principal component analysis processing on the edge points on the at least one finger region contour 22, respectively, and takes the obtained main direction as the finger direction.
  • Principal component analysis is a multivariate statistical method. Statistical analysis is performed on two-dimensional points. Two vectors are respectively corresponding to two eigenvalues. The larger the corresponding eigenvalues, the more significant the distribution of point clouds along this direction is. The main direction corresponds to The vector with the largest eigenvalue.
  • S24b2 determines the maximum projection point in at least one finger direction as the corrected fingertip 21 position.
  • the step S24b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a fingertip correction unit operated by the processor.
  • the relative distribution of the contours 22 of the respective finger regions is determined according to the corrected position of the fingertips 21 and the position of the finger roots.
  • a straight line may be determined according to two points of the fingertip 21 and the finger root, and five fingers correspond to five straight lines, the five straight lines have a certain angle with each other, and the five straight lines Regardless of the direction of the palm 23, it is possible to determine which finger each line corresponds to based on the angle of the straight line and the direction with respect to the palm 23.
  • the missing fingers and the corresponding gestures can be further determined by the determined relative distribution of the fingers, and the missing fingers can be ignored in modeling to improve the modeling efficiency.
  • the step S25b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a distribution determining unit operated by the processor.
  • Step S25b optionally includes the following steps:
  • S25b3 respectively compares the direction of the position of the at least one finger root with respect to the position of the palm 23 with the average direction.
  • the finger root position corresponding to the direction in which the average direction is greater than the preset threshold is removed, and the finger root that deviates from the average direction exceeding a certain threshold is determined to be misdetected as the wrist of the finger, and is removed from the finger set.
  • S25b5 determines the relative distribution of the respective fingers according to the direction of the remaining position of the respective finger roots relative to the position of the palm 23.
  • the order of the respective fingers is judged according to the direction of the positions of the remaining finger roots relative to the position of the palm 23, and the retained fingers are sequentially labeled as the thumb, the index finger, the middle finger, the ring finger, and the little finger in the direction of the finger root relative to the palm 23 .
  • the above solution determines the relative distribution of the finger according to the direction of the finger relative to the palm, avoids analyzing the shape of the contour of the finger region, and reduces the shape judgment operation.
  • the scheme has high accuracy and processing efficiency, and thus the relative distribution is determined. Information can provide a basis for gesture judgment.
  • step S3 may include the following steps:
  • the depth information of the palm 23 is determined according to the imaging parameters of the binocular imaging system and the position of the palm 23 in the first hand image and the second hand image.
  • the two can be directly associated with each other, and the palm region contour 24 can be directly associated with each other.
  • the depth information of the object points can be determined according to the imaging parameters.
  • the step S31a may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm depth determining unit operated by the processor.
  • a plane can be determined based on the depth, and then the edge points of the palm region contour 24 are all projected onto the depth plane, that is, in the two hand images.
  • the palm region contours 24 are projected onto the plane, which has two palm region contours 24, and then matches each pair of closest points of the two palm region contours 24, thereby achieving contouring of the region. All edge points are matched, and then the depth information of these object points can be determined according to the imaging parameters.
  • the above embodiment first determines the depth of the object point of the palm, and then projects the points on the edge of the contour of the palm region in the two hand images onto the plane established based on the depth of the palm, and uses the projection depth as the depth of the contour of the palm region. It is avoided to match the points of the area contour one by one, thereby improving the processing efficiency.
  • the step S32a may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm area depth determining unit operated by the processor.
  • step S3 may include the following steps:
  • the step S31b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a relative position determining unit operated by the processor.
  • S32b determining at least one pair of fingertips 21 of the two images according to the mode length of the relative position vector difference, and corresponding pairs of finger region contours 22 of the two images. For example, the position of a pair of fingertips 21 that minimizes the mode length of the relative position vector difference in the first hand image and the second hand image (ie, a pair of fingertips 21 having the smallest mode length of the vector difference of the palm 23) is seen. By making the same fingertip 21, the finger region contour 22 corresponding to the position of the matched fingertip 21 can be further matched.
  • step S32b may be performed by a processor invoking a corresponding instruction stored in the memory or by a finger matching unit operated by the processor.
  • the step S33b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a fingertip depth determining unit operated by the processor.
  • the fingertip 21 depth information is first determined, then the projection plane is established based on the fingertip 21 depth, and the point on the edge of the finger region contour 22 is projected onto the plane to determine the depth information of the edge point.
  • the above embodiment firstly matches the fingers in the two images according to the relative position vector of the position of the fingertip 21 and the position of the palm 23 to determine the corresponding pairs of fingertips 21, the finger region contour 22, and then first determines the fingertip 21
  • the depth of the object point, and the points on the edge of the finger region contour 22 in the two hand images are projected onto the plane established based on the depth of the corresponding fingertip 21, and the projection depth is taken as the depth of the finger region contour 22, avoiding The points of the area contour are matched one by one, thereby improving the processing efficiency.
  • the step S34b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a finger depth determination unit operated by the processor.
  • the previous steps have already determined the depth information of the key points and the area contours.
  • the edge three-dimensional point cloud data of each area contour, and the hand direction data are three kinds of data. There are many ways to determine. Since the key points of the fingertip 21 and the palm 23 are particularly important in practical applications, in order to more accurately determine the three-dimensional data of the key points, the present embodiment is based on determining the regional point cloud data and the hand direction data. Perform the determination of the three-dimensional data of the key points.
  • the point cloud data of the area contour covering the key points may be first determined, for example, may include at least one of the following steps S41 and S42:
  • the step S41 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a finger point cloud data determining unit operated by the processor.
  • the edge 3D point cloud data of the palm area contour 24 is established according to the position of the edge point on the palm area contour 24 and the depth of the palm 23 position in the two hand images, that is, the points on the edge of the palm area contour 24 are all projected to Based on the projection plane established by the palm 23 to determine the depth information of the edge point, the set of depth information and position information is the three-dimensional point cloud data of the finger region contour 22.
  • the step S42 may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm point cloud data determining unit operated by the processor.
  • the hand direction data may be obtained based on the three-dimensional point cloud data obtained in the above steps.
  • the hand direction data may include, for example, finger direction data and/or palm normal data, and may include, for example, the following steps:
  • principal component analysis is a multivariate statistical method.
  • the analysis of three-dimensional points will result in three vectors corresponding to three eigenvalues.
  • the main direction is the vector with the largest eigenvalue.
  • the step S42a may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a finger direction determining unit operated by the processor.
  • the principal component analysis processing is performed on the edge 3D point cloud data of the palm area contour 24, and the feature direction corresponding to the obtained minimum feature value is recorded as the palm normal, and the feature direction corresponding to the obtained minimum feature value is recorded as the palm normal.
  • the step S42b may be performed by a processor invoking a corresponding instruction stored in the memory, or may be performed by a palm normal determining unit operated by the processor.
  • determining the three-dimensional data of the key points based on the hand direction data and the three-dimensional point cloud data may include the following steps:
  • S42b1 determining the three-dimensional position of the initial palm 23 according to the imaging parameters of the binocular camera system. This step can be directly determined according to the depth information and the imaging parameters of the previously determined palm 23, but the position determined based on this may be considered to be inaccurate. Further adjustments need to be made based on the initial position;
  • the above method for determining the palm three-dimensional data first determines the initial palm position by matching the palm of the two-dimensional image and combining the imaging parameters, and then the point projection on the contour of the palm region is based on the plane established by the depth of the initial palm, and is adjusted according to the effect of the projected point.
  • the initial palm position, thereby determining the palm three-digit data has higher accuracy.
  • the above embodiment first determines the regional point cloud data, and then determines the hand direction data based on the regional point cloud data, and finally determines the three-dimensional data of the key points based on the regional point cloud data and the hand direction data, thereby determining the three-dimensional data of the key points. Accurate and more stable.
  • step S4 may include the step of determining whether the number of area contours determined in step S2 is less than a preset number, and determining the missing area contour when the number of area contours is less than the preset number. If the missing area contour is the finger area outline 22, the relative distribution between the fingers can be determined in step S2, and in the case where the relative distribution is determined, which finger is missing can be determined.
  • a common finger curl gesture table can be pre-stored:
  • a 5-digit binary number represents a gesture, from the highest to the lowest, indicating the state of the thumb, forefinger, middle finger, ring finger, and little finger, 1 means finger extension, 0 means finger contraction, such as: 01000 means other than index finger
  • the other fingers are in a collapsed state. From this, the missing finger in the image can be determined, that is, the gesture information is obtained.
  • the above embodiment first determines whether the hand in the image is complete according to the number of key region contours, and determines the missing region contour according to the relative distribution of the region contours when the incompleteness is determined, thereby determining the gesture, which does not require a large number of pieces.
  • the operation and graphics analysis operations have high processing efficiency.
  • any of the hand three-dimensional data determining methods provided by the embodiments of the present application may be performed by any suitable device having data processing capabilities, including but not limited to: a terminal device, a server, and the like.
  • any one of the hand three-dimensional data determining methods provided by the embodiments of the present application may be executed by a processor, such as the processor performing the three-dimensional data determination of any hand mentioned in the embodiment of the present application by calling corresponding instructions stored in the memory. method. This will not be repeated below.
  • the foregoing program may be stored in a computer readable storage medium, and the program is executed when executed.
  • the foregoing steps include the steps of the foregoing method embodiments; and the foregoing storage medium includes: a medium that can store program codes, such as a ROM, a RAM, a magnetic disk, or an optical disk.
  • FIG. 3 is a schematic structural diagram of a three-dimensional hand data determining apparatus according to an embodiment of the present application. As shown in FIG. 3, the apparatus includes:
  • the acquiring unit 31 is configured to acquire a first hand image and a second hand image captured by the binocular camera system;
  • the identifying unit 32 is configured to respectively identify at least one key point and an area contour covering the key point from the first hand image and the second hand image;
  • a depth determining unit 33 configured to determine, according to an imaging parameter of the binocular imaging system, a key point and a corresponding area contour recognized from the first hand image, and a key point and a corresponding area contour recognized from the second hand image Depth information of at least one key point and depth information of the area outline;
  • the three-dimensional data determining unit 34 is configured to determine the hand three-dimensional data according to at least the at least one key point and the depth information thereof, and the region contour and the depth information thereof.
  • the hand three-dimensional data determining device identifies key points and region contours in the first hand image and the second hand image captured by the binocular camera system, and then according to the camera parameters of the binocular camera system, Key points and area contours identified by the first hand image, and key points and area outlines identified from the second hand image, the depth of the key point and the depth of the area outline can be determined; and then according to the key points and key points The depth, the area contour and the depth thereof determine the three-dimensional data of the hand.
  • the depth can be determined by matching each other, and further Determining the three-dimensional data thereof, thereby determining the three-dimensional information of the hand based on the two-dimensional image of the single frame, and making the determined three-dimensional data of the hand more accurate and richer, and the technical solution for determining the three-dimensional data of the hand in the embodiment of the present application With high efficiency and accuracy.
  • the above key points may include, but are not limited to, any one or more of the following: a fingertip, an knuckle point, and a palm.
  • the above-mentioned area contour covering the key points may include, but is not limited to, any one or more of the following: a finger area contour covering the fingertip and/or the knuckle point, and a palm area contour covering the palm.
  • the hand three-dimensional data may include, but is not limited to, any one or more of the following: three-dimensional data of a key point, edge three-dimensional point cloud data of the area contour, hand direction data, and gesture information.
  • the identification unit may include:
  • a palm recognition unit for identifying a palm in the first hand image and the second hand image, respectively;
  • the palm recognition unit is configured to determine a palm area contour covering the palm according to the position of the recognized palm in the first hand image and the second hand image, respectively.
  • the above optional embodiment finds the maximum connected area in the hand image after binarization processing, and further determines the palm 23 and the palm area contour 24 based on the maximum connected area.
  • the present scheme finds the maximum connected area by the line feature, and then based on the maximum connected area. To determine the palm of the hand to reduce the effects of image color and grayscale, thereby improving processing efficiency.
  • the identification unit may include:
  • a palm recognition unit for identifying a palm in the first hand image and the second hand image, respectively;
  • a fingertip recognition unit for determining at least one fingertip and/or knuckle point position according to the position of the recognized palm in the first hand image and the second hand image, respectively;
  • the finger recognition unit is configured to determine a finger region contour covering the fingertip and/or the knuckle point according to the fingertip and/or the knuckle point position in the first hand image and the second hand image, respectively.
  • the palm is determined in the hand image after the binarization processing, and then the point farthest from the palm is determined in the point on the edge of the maximum connected area, and then the contour of the area suspected of the finger is found based on the starting point.
  • the contour of the suspected finger area is further excluded according to the outline shape to exclude the outline of the finger area, thereby eliminating the influence of the wrist on the recognition process, and after finding the outline of the finger area, determining the corresponding starting point as Fingertips, which can improve accuracy and processing efficiency.
  • the hand three-dimensional data determining apparatus may further include:
  • a fingertip correction unit for correcting a fingertip position and determining a finger root position according to at least one finger region contour in the first hand image and the second hand image, respectively;
  • a distribution determining unit configured to determine a relative distribution of the at least one finger region contour according to the corrected fingertip position and the finger root position. According to the above scheme, the missing fingers and the corresponding gestures can be further determined by the determined relative distribution of the fingers, and the missing fingers can be ignored in modeling to improve the modeling efficiency.
  • the fingertip correction unit may include:
  • a finger direction determining unit configured to perform principal component analysis processing on the edge points on the contour of the at least one finger region, respectively, and use the obtained main direction as a finger direction;
  • a fingertip determining unit configured to respectively determine a maximum projection point in at least one finger direction as a corrected fingertip position
  • the finger root determining unit is configured to respectively determine a minimum projection point in at least one finger direction as a finger root position.
  • the distribution determining unit may include:
  • a root relative direction determining unit for respectively determining a direction of a position of the at least one finger root position with respect to the palm
  • An average direction determining unit for determining an average direction of a direction of at least one finger root position with respect to a position of the palm;
  • a aligning unit for respectively comparing a direction of the position of the at least one finger root with respect to a position of the palm with an average direction
  • a culling unit configured to remove a finger root position corresponding to a direction in which the deviation from the average direction is greater than a preset threshold
  • a distribution identifying unit configured to determine a relative distribution of the fingers according to a direction of the retained position of the at least one finger root relative to the position of the palm.
  • the above solution determines the relative distribution of the finger according to the direction of the finger relative to the palm, avoids analyzing the shape of the contour of the finger region, and reduces the shape judgment operation.
  • the scheme has high accuracy and processing efficiency, and thus the relative distribution is determined.
  • Information can provide a basis for gesture judgment.
  • the palm recognition unit may include:
  • a connected area determining unit configured to determine a maximum connected area contour in the first hand image and the second hand image respectively after the binarization processing
  • the palm position determining unit is configured to determine the position of the palm according to the maximum connected area contour in the first hand image and the second hand image, respectively.
  • the palm recognition unit may include:
  • a palm radius determining unit configured to determine a palm radius according to a distance from a position of the palm to an edge of the two largest connected regions, respectively;
  • the palm area contour determining unit is configured to determine the palm area contour according to the palm radius, respectively.
  • the palm recognition unit may include:
  • a connected area determining unit configured to determine a maximum connected area in the first hand image and the second hand image after the binarization process, respectively;
  • a palm position determining unit configured to determine a position of the palm according to a maximum connected area of the first hand image and the second hand image, respectively;
  • the fingertip recognition unit may include:
  • a far point determining unit for determining a point farthest from the palm position in the maximum connected area
  • a suspected area determining unit configured to determine a contour of the suspected finger area within the maximum connected area based on a point farthest from the palm position
  • a fingertip determining unit configured to determine a fingertip and obtain a position thereof from at least one point farthest from the palm position according to a shape of the at least one suspected finger region contour
  • the finger recognition unit includes:
  • a finger determining unit configured to determine a finger region contour from the at least one suspected finger region contour according to a shape of the at least one suspected finger region contour.
  • the depth determining unit may include:
  • a palm depth determining unit configured to determine depth information of the palm according to the imaging parameter of the binocular camera system, the position of the palm in the first hand image and the second hand image;
  • the palm region depth determining unit is configured to project an edge point of the palm region contour to a depth at which the palm is located, and obtain depth information of the contour of the palm region.
  • the above embodiment first determines the depth of the object point of the palm, and then projects the points on the edge of the contour of the palm region in the two hand images onto the plane established based on the depth of the palm, and uses the projection depth as the depth of the contour of the palm region. It is avoided to match the points of the area contour one by one, thereby improving the processing efficiency.
  • the depth determining unit may include:
  • a relative position determining unit configured to determine a relative position vector of the at least one fingertip position and the palm position in the first hand image and the second hand image, respectively;
  • a finger matching unit configured to determine a corresponding at least one pair of fingertips of the two images according to a mode length of the relative position vector difference, and a corresponding at least one pair of finger region contours of the two images;
  • a fingertip depth determining unit configured to determine depth information of at least one pair of fingertips according to an imaging parameter of the binocular imaging system, a position of the at least one pair of fingertips in the first hand image and the second hand image;
  • a finger depth determining unit configured to project edge points of at least one pair of finger region contours to respective depths of the fingertips to obtain depth information of at least one pair of finger region contours.
  • the above embodiment firstly matches the fingers in the two images according to the relative position vector of the fingertip position and the palm position to determine the corresponding at least one pair of fingertips, the finger region contour, and then first determines the depth of the fingertip. And projecting the points on the edge of the contour of the finger region in the two hand images onto the plane established based on the depth of the corresponding fingertip, and using the projection depth as the depth of the contour of the finger region, avoiding the points for the contour of the region one by one Matching, thereby improving processing efficiency.
  • the three-dimensional data determining unit may include:
  • a finger point cloud data determining unit configured to establish a finger region contour edge three-dimensional point cloud data according to a position of the edge point on the contour of the corresponding finger region in the two hand images and a depth of the finger tip position;
  • the palm point cloud data determining unit is configured to establish the palm area contour edge three-dimensional point cloud data according to the position of the edge point on the contour of the palm area and the depth of the palm position in the two hand images.
  • the three-dimensional data determining unit may include:
  • a finger direction determining unit configured to perform principal component analysis processing on the three-dimensional point cloud data of the contour edge of the finger region, and obtain a main direction as a finger direction;
  • the palm normal determining unit is configured to perform principal component analysis processing on the three-dimensional point cloud data of the contour edge of the palm region, and the feature direction corresponding to the obtained minimum feature value is recorded as the palm normal.
  • the three-dimensional data determining unit may include:
  • a fingertip three-dimensional data determining unit for using a maximum projection point in the direction of the finger as a fingertip and determining three-dimensional data thereof;
  • a palm three-dimensional position determining unit configured to determine an initial palm three-dimensional position according to an imaging parameter of the binocular imaging system
  • the palm three-dimensional data determining unit is configured to adjust the initial three-dimensional position of the palm, obtain the adjusted three-dimensional position of the palm and determine the three-dimensional data thereof.
  • the above method for determining the palm three-dimensional data first determines the initial palm position by matching the palm of the two-dimensional image and combining the imaging parameters, and then the point projection on the contour of the palm region is based on the plane established by the depth of the initial palm, and is adjusted according to the effect of the projected point.
  • the initial palm position, thereby determining the palm three-digit data has higher accuracy.
  • the above embodiment first determines the regional point cloud data, and then determines the hand direction data based on the regional point cloud data, and finally determines the three-dimensional data of the key points based on the regional point cloud data and the hand direction data, thereby determining the three-dimensional data of the key points. Accurate and more stable.
  • the embodiment of the present application further provides an electronic device, such as a mobile terminal, a personal computer (PC), a tablet computer, a server, and the like.
  • the electronic device includes: at least one processor; and a memory communicatively coupled to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the At least one processor performs an operation corresponding to any of the hand three-dimensional data determining methods described in the embodiments.
  • FIG. 4 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
  • the computer system 400 includes one or more processors and a communication unit.
  • the one or more processors for example: one or more central processing units (CPUs) 401, and/or one or more image processing units (GPUs) 413, etc., the processors may be stored in a read only memory (ROM)
  • the executable instructions in 402 are either loaded from executable portion 408 into executable instructions in random access memory (RAM) 403 to perform appropriate actions and processing.
  • the communication unit 412 may include, but is not limited to, a network card, which may include, but is not limited to, an IB (Infiniband) network card.
  • the processor can communicate with the ROM 402 and/or the RAM 403 to execute executable instructions, connect to the communication unit 412 via the bus 404, and communicate with other target devices via the communication unit 412, thereby completing any of the methods provided by the embodiments of the present application.
  • Corresponding operations for example, acquiring a first hand image and a second hand image captured by the binocular camera system; identifying at least one key point from the first hand image and the second hand image, respectively, and covering the An area outline of a key point; an imaging parameter of the binocular camera system, the key point and the area outline identified from the first hand image, and the image recognized from the second hand image Determining the depth information of the key point and the depth information of the area outline according to the key point and the area contour; determining the hand according to at least the key point and its depth information, and the area contour and depth information thereof Three-dimensional data.
  • the CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404.
  • ROM 402 is an optional module.
  • the RAM 403 stores executable instructions or writes executable instructions to the ROM 402 at runtime, the executable instructions causing the CPU 401 to perform operations corresponding to the above-described communication methods.
  • An input/output (I/O) interface 405 is also coupled to bus 404.
  • the communication unit 412 may be integrated or may be provided with a plurality of sub-modules (e.g., a plurality of IB network cards) and linked on the bus 404.
  • the following components are connected to the I/O interface 405: an input portion 406 including a keyboard, a mouse, etc.; an output portion 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a storage portion 408 including a hard disk or the like. And a communication portion 409 including a network interface card such as a LAN card, a modem, or the like. The communication section 409 performs communication processing via a network such as the Internet.
  • Driver 410 is also coupled to I/O interface 405 as needed.
  • a removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 410 as needed so that a computer program read therefrom is installed into the storage portion 408 as needed.
  • FIG. 4 is only an optional implementation manner.
  • the number and types of components in FIG. 4 may be selected, deleted, added, or replaced according to actual needs;
  • the function component setting may also adopt an implementation such as a separate setting or an integrated setting.
  • the GPU 413 and the CPU 401 may be separately disposed or the GPU 413 may be integrated on the CPU 401, and the communication part may be separately configured or integrated in the CPU 401. Or on GPU 413, and so on.
  • an embodiment of the present disclosure includes a computer program product comprising a computer program tangibly embodied on a machine readable medium, the computer program comprising program code for executing the method illustrated in the flowchart, the program code comprising Executing instructions corresponding to the method steps provided by the embodiments of the present application, for example, acquiring a first hand image and a second hand image captured by the binocular camera system; respectively, from the first hand image and the second hand image Identifying at least one key point and an area outline covering the key point; according to imaging parameters of the binocular camera system, the key points and the area contours identified from the first hand image, and Determining the key point and the area contour identified by the second hand image, determining depth information of the key point and depth information of the area contour; at least according to the key point and depth information thereof, and The contour of the area and its depth information determine the three-dimensional
  • the embodiment of the present application further provides a computer readable storage medium, where computer instructions are stored, and when the instructions are executed, the operations of each step in the method for determining the three-dimensional data of the hand according to any one of the embodiments of the present application are implemented. .
  • the embodiment of the present application further provides a computer program, including computer readable code, when the computer readable code is run on a device, the processor in the device executes to implement any of the embodiments of the present application.
  • the hand three-dimensional data determines instructions for each step in the method.
  • the methods, apparatus, and apparatus of the present application may be implemented in a number of ways.
  • the methods, apparatus, and apparatus of the present application can be implemented in software, hardware, firmware, or any combination of software, hardware, and firmware.
  • the above-described sequence of steps for the method is for illustrative purposes only, and the steps of the method of the present application are not limited to the order described above unless otherwise specified.
  • the present application can also be implemented as a program recorded in a recording medium, the programs including machine readable instructions for implementing the method according to the present application.
  • the present application also covers a recording medium storing a program for executing the method according to the present application.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Human Computer Interaction (AREA)
  • Software Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Databases & Information Systems (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computing Systems (AREA)
  • Artificial Intelligence (AREA)
  • Psychiatry (AREA)
  • Social Psychology (AREA)
  • Image Analysis (AREA)
  • User Interface Of Digital Computer (AREA)
  • Image Processing (AREA)

Abstract

本申请实施例提供一种手部三维数据确定方法、装置及电子设备,所述方法包括:获取双目摄像系统拍摄的第一手部图像和第二手部图像;分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。本申请实施例的技术方案具有较高的效率和准确性,且确定的手部三维数据更加准确、更丰富。

Description

手部三维数据确定方法、装置及电子设备
本申请要求在2017年03月29日提交中国专利局、申请号为CN 201710198505.7、发明名称为“手部三维数据确定方法、装置及电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及计算机视觉技术,具体涉及一种手部三维数据确定方法、装置及电子设备。
背景技术
手部数据是人机交互领域中重要的输入数据,通过捕捉人体手部数据的变化,可以实现利用手势动作对智能设备进行控制。
发明内容
本申请实施例提供了一种确定手部三维数据的技术方案。
根据本申请实施例的一个方面,提供的一种手部三维数据确定方法,包括:
获取双目摄像系统拍摄的第一手部图像和第二手部图像;
分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;
根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;
至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。
根据本申请实施例的另一方面,提供的一种手部三维数据确定装置,包括:
获取单元,用于获取双目摄像系统拍摄的第一手部图像和第二手部图像;
识别单元,用于分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;
深度确定单元,用于根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;
三维数据确定单元,用于至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。
根据本申请实施例的又一方面,提供的一种电子设备,包括:至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器执行本申请任一实施例所述手部三维数据确定方法方法对应的操作。
根据本申请实施例的再一方面,提供的一种计算机可读存储介质,其上存储有计算机指令,所述指令被执行时实现本申请任一实施例所述手部三维数据确定方法方法中各步骤 的操作。
根据本申请实施例的再一方面,提供的一种计算机程序,包括计算机可读代码,当所述计算机可读代码在设备上运行时,所述设备中的处理器执行用于实现本申请任一实施例所述手部三维数据确定方法中各步骤的指令。
根据本申请实施例提供的手部三维数据确定方法、装置及电子设备,在双目摄像系统拍摄的第一手部图像和第二手部图像中识别关键点和区域轮廓,之后根据双目摄像系统的摄像参数、从第一手部图像识别出的关键点和区域轮廓、以及从第二手部图像识别出的关键点和区域轮廓,可以确定关键点的深度和区域轮廓的深度;再根据关键点和关键点的深度、区域轮廓及其深度确定手部三维数据,由此确定的手部三维数据更加准确且更丰富,并且,本申请实施例确定手部三维数据的技术方案具有较高的效率和准确性。
下面通过附图和实施例,对本申请的技术方案做进一步的详细描述。
附图说明
为了更清楚地说明本申请实施方式或现有技术中的技术方案,下面将对本申请的实施方式或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本申请实施例的一些实施方式,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本申请实施例的手部三维数据确定方法的一个流程图;
图2为本申请实施例的手部关键点和区域轮廓的一个示意图;
图3为本申请实施例的手部三维数据确定装置的一个结构示意图;
图4为本申请实施例的电子设备的一个结构示意图。
具体实施方式
下面将结合附图对本申请实施例的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
在本申请的描述中,需要说明的是,术语“第一”、“第二”仅用于描述目的,而不能理解为指示或暗示相对重要性。此外,下面所描述的本申请不同实施方式中所涉及的技术特征只要彼此之间未构成冲突就可以相互结合。
应注意到:除非另外说明,否则在这些实施例中阐述的部件和步骤的相对布置、数字表达式和数值不限制本申请的范围。
同时,应当明白,为了便于描述,附图中所示出的各个部分的尺寸并不是按照实际的比例关系绘制的。
以下对至少一个示例性实施例的描述实际上仅仅是说明性的,不作为对本申请及其应用或使用的任何限制。
对于相关领域普通技术人员已知的技术、方法和设备可能不作详细讨论,但在适当情况下,所述技术、方法和设备应当被视为说明书的一部分。
应注意到:相似的标号和字母在下面的附图中表示类似项,因此,一旦某一项在一个 附图中被定义,则在随后的附图中不需要对其进行进一步讨论。
本申请实施例可以应用于终端设备、计算机系统、服务器等电子设备,其可与众多其它通用或专用计算系统环境或配置一起操作。适于与终端设备、计算机系统、服务器等电子设备一起使用的众所周知的终端设备、计算系统、环境和/或配置的例子包括但不限于:个人计算机系统、服务器计算机系统、瘦客户机、厚客户机、手持或膝上设备、基于微处理器的系统、机顶盒、可编程消费电子产品、网络个人电脑、小型计算机系统﹑大型计算机系统和包括上述任何系统的分布式云计算技术环境,等等。
终端设备、计算机系统、服务器等电子设备可以在由计算机系统执行的计算机系统可执行指令(诸如程序模块)的一般语境下描述。通常,程序模块可以包括例程、程序、目标程序、组件、逻辑、数据结构等等,它们执行特定的任务或者实现特定的抽象数据类型。计算机系统/服务器可以在分布式云计算环境中实施,分布式云计算环境中,任务是由通过通信网络链接的远程处理设备执行的。在分布式云计算环境中,程序模块可以位于包括存储设备的本地或远程计算系统存储介质上。
本申请实施例提供了一种手部三维数据确定方法,图1为本申请实施例的手部三维数据确定方法的一个流程图。如图1所示,该实施例的手部三维数据确定方法包括如下步骤:
S1,获取双目摄像系统拍摄的第一手部图像和第二手部图像。
其中,双目摄像系统具有2个成像设备,基于视差原理利用2个成像设备从不同的位置获取被测物体的两幅图像,获取到的手部的两幅图像即称为第一手部图像和第二手部图像。双目摄像系统有多种,利用任一种双目摄像系统对手部进行拍摄获取2个手部图像都是可行的。
在一个可选示例中,该步骤S1可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的获取单元执行。
S2,分别从第一手部图像和第二手部图像中识别至少一关键点以及覆盖关键点的区域轮廓。
其中的关键点例如可以包括但不限于以下任意一项或多项:指尖、指关节点、掌心,覆盖相应关键点的区域轮廓例如可以包括但不限于以下任意一项或多项:覆盖指尖和/或指关节点的手指区域轮廓,覆盖掌心的手掌区域轮廓。例如,关键点为指尖(5个)和掌心(1个),区域轮廓为覆盖这些指尖的手指区域轮廓和覆盖掌心的手掌区域轮廓。上述2个手部图像可以为二维图像,从二维图像中识别出符合预定特征的目标的方式有多种,利用各种方式都是可行的。
在一个可选示例中,该步骤S2可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的识别单元执行。
S3,根据双目摄像系统的摄像参数、从第一手部图像识别出的至少一个关键点和相应区域轮廓、以及从第二手部图像识别出的至少一个关键点和相应区域轮廓,确定至少一个关键点的深度信息以及相应区域轮廓的深度信息。
双目摄像系统的两个摄像头对同一物体进行拍摄得到两张图像,要确定图像中物体的深度信息,可以将两张图像中的物点匹配起来,即图像匹配。图像匹配是指将三维空间中一点或一个区域轮廓在左右摄像机的两个成像面上的像点或区域轮廓对应起来。在找到相互匹配的物点之后,即可根据摄像参数确定该物点的深度。
对于关键点而言,首先以关键点是掌心为例,经过识别处理,可以在上述2图像中分 别识别出掌心aL和aR,由于2个图像中分别只有1个掌心,因此可以直接将aL和aR对应起来;同理,也可以直接将手掌区域轮廓对应起来。关于区域轮廓的匹配,区域轮廓是由很多个点组成的,本步骤可以只对区域轮廓上的边缘点进行处理。
关于深度信息的确定,在其中一个实施例方式中,可以以双目摄像系统的双目连线的中点为原点建立坐标系,其中,与双目摄像系统的成像平面平行的平面为XY平面,与这一平面垂直的方向为Z方向,Z方向即为深度方向,关键点的深度值就是在这一坐标系下的Z坐标值。在确定了匹配结果后,则可以基于同一物点在左右相机上的投影点,借助双目摄像机参数,恢复摄像机的广角畸变,再利用摄像机焦距、左右摄像机间隔等参数,根据相机的几何成像关系可将此物点的深度信息计算出来。
在一个可选示例中,该步骤S3可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的深度确定单元执行。
S4,至少根据上述至少一个关键点及其深度信息和相应区域轮廓及其深度信息,确定手部三维数据。
上述手部三维数据例如可以包括但不限于以下任意一项或多项:关键点的三维数据、区域轮廓的边缘三维点云数据、手部方向数据、手势信息。
在一个可选示例中,该步骤S4可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的三维数据确定单元执行。
本申请实施例提供的手部三维数据确定方法,在双目摄像系统拍摄的第一手部图像和第二手部图像中识别关键点和区域轮廓,之后根据人体手部的特点可以实现对关键点和区域轮廓进行匹配,可以确定关键点的深度和区域轮廓的深度;最终根据关键点的匹配结果和关键点的深度、区域轮廓上的边缘点的匹配结果和区域轮廓上的边缘点被投影的深度确定手部三维数据。根据本申请提供的技术方案,只要是在二维图像中能够识别出的物点、区域轮廓等内容,均可以通过相互匹配确定深度,并进一步确定其三维数据,由此可以实现基于单帧的二维图像确定手部大量三维信息,并使确定的手部三维数据更加准确且更丰富,避免产生不合理的手部三维数据,并且具有较高的效率和准确性。
图2为本申请实施例的手部关键点和区域轮廓的一个示意图。如图2所示,对于仅需要识别掌心23和覆盖掌心23的手掌区域轮廓24的情况,上述步骤S2可以通过多种方式实现,例如在一种可选方式中,可以包括如下步骤:
S21a,分别在第一手部图像和第二手部图像中识别掌心23。
在其中一种实施方式中,可以根据图像中的线条特征、区域特征识别掌心23。
为了减少光照条件对图像的影响,可以选用红外摄像头,以直接得到灰度图像,然后对图像进行二值化处理,将大于某一阈值的像素点全部赋值为255,余下的像素点赋值为0,由此得到线条图。
在其中一种可选示例中,步骤S21a可以包括如下步骤:
S21a1,分别在经过二值化处理后第一手部图像和第二手部图像中确定最大连通区域轮廓,即找到赋值为255的点构成的最大连通集,将其余的点全部赋值为0,这个最大连通集被看作手在图像中的投影区域轮廓;
S21a2,分别根据第一手部图像和第二手部图像中的最大连通区域轮廓确定掌心23的位置。
可选地,可以分别从2个最大连通区域轮廓的边缘出发对最大连通区域轮廓进行曼哈 顿距离变换,即从边缘点出发做广度优先搜索,逐层记录每个点到边缘点的最小曼哈顿距离,得到区域轮廓内每个点到边缘的曼哈顿距离。然后将距边缘最远的点初步确定为掌心23。
在一个可选示例中,该步骤S21a可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的掌心识别单元执行。
S22a,分别在第一手部图像和第二手部图像中根据识别的掌心23的位置确定覆盖掌心23的手掌区域轮廓24。
在确定了掌心23的情况下,确定覆盖掌心23的区域的方式有多种,在一种可选示例中,可以包括如下步骤:
S22a1,分别根据掌心23位置到2个最大连通区域的边缘的距离确定手掌半径。
可选地,可以找出最大连通区域的边缘点中距掌心23最近的点,将最近的点与掌心23之间的欧氏距离初步定为手掌半径;
S22a2,分别根据手掌半径确定手掌区域轮廓24。
例如可以在手部区域轮廓内,将与掌心23的欧氏距离小于1.5倍半径的点构成的区域轮廓标记为手掌区域轮廓24。
上述实施方式在二值化处理后的手部图像中寻找最大连通区域,进而基于最大连通区域确定掌心23和手掌区域轮廓24,本方案通过线条特征寻找最大连通区域,然后基于最大连通区域来确定手掌,可以减少图像颜色和灰度的影响,由此可以提高处理效率。
在一个可选示例中,该步骤S22a可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手掌识别单元执行。
如图2所示,对于需要识别指尖21和覆盖指尖21的手指区域轮廓22的情况,上述步骤S2可以包括如下步骤:
S21b,分别在第一手部图像和第二手部图像中识别掌心23。
也即在识别指尖21和手指区域轮廓22之前,首先确定掌心23,然后基于掌心23确定指尖21和手指区域轮廓22。此步骤可以与上述步骤S21a相同,此次不再赘述;
在一个可选示例中,该步骤S21b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的掌心识别单元执行。
S22b,分别在第一手部图像和第二手部图像中根据识别的掌心23的位置确定至少一指尖21和/或指关节点位置。
如上,在确定掌心23的过程中,分别在两图像中确定了最大连通区域,由此可以基于该最大连通区域进一步寻找指尖21和手指区域轮廓22。在其中一个可选示例中,该步骤S22b可以包括如下步骤:
S22b1,在最大连通区域内确定距掌心23位置最远的点。
S22b2,基于距掌心23位置最远的点在最大连通区域内确定疑似手指区域轮廓。
例如,可以在最大连通区域内找到距掌心23位置最远的点,从该最远的点出发,在最大连通区域轮廓内做广度优先延伸,得到疑似手指区域轮廓。
S22b3,根据疑似手指区域轮廓的形状从距掌心23位置最远的点中确定指尖21并获取其位置。
在一个可选示例中,该步骤S22b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的指尖识别单元执行。
S23b,分别根据第一手部图像和第二手部图像中的指尖21和/或指关节点位置,确定覆盖指尖21和/或指关节点的手指区域轮廓22。
步骤S22b3和S23b操作过程为:基于形状特征判断这些疑似手指区域轮廓的形状是否为手指,若疑似手指区域轮廓的形状是手指,则将相应的最远的点(出发点)初步确定为指尖,并将该疑似手指区域轮廓判定为手指区域轮廓22。若疑似手指区域轮廓的形状不是手指,则将疑似手指区域轮廓剔除,继续判断其余疑似手指区域轮廓(将这一步骤多次重复),直至所有疑似手指区域轮廓均被判定完毕,找到全部手指区域轮廓22和相应的指尖21,然后获取其位置。上述实施方式在二值化处理后的手部图像中确定掌心,进而在最大连通区域边缘上的点中确定与掌心相距最远的点,然后基于该出发点寻找疑似手指的区域轮廓,在找到的疑似手指的区域轮廓中进一步根据轮廓形状排除非手指区域轮廓,以筛选出手指区域轮廓,由此可以排除手腕等其它区域对识别过程的影响,并在找到手指区域轮廓后,将相应的出发点确定为指尖,由此可以提高准确率和处理效率。
在一个可选示例中,该步骤S23b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手指识别单元执行。
作为另一个可选的实施方式,在上述步骤S23b之后,还可以包括如下步骤:
S24b,分别在第一手部图像和第二手部图像中根据各个手指区域轮廓22修正指尖21位置并确定指根位置。
由于之前确定的指尖21位置是根据2个图像的内容分别确定的,在确定两张图像中相应的了手指区域轮廓22之后,可以对之前的指尖21位置进行修正,即重新确定指尖21位置,修正后的指尖21位置比之前确定指尖21位置更加准确。指根位置即手指与手掌相连的位置,在已知指尖21位置和掌心23位置的情况下,确定各个指根位置的方式有多种。步骤S24b可选包括如下步骤:
S24b1,分别对至少一个手指区域轮廓22上的边缘点进行主成分分析处理,将得到的主方向作为手指方向。
主成分分析是一种多元统计方法,对二维点做统计分析,得到两个向量分别对应两个特征值,对应的特征值越大表示点云沿这个方向的分布越显著,主方向即对应的特征值最大的向量。
S24b2,分别将至少一个手指方向上的最大投影点确定为修正后的指尖21位置。
S24b3,分别将至少一个手指方向上的最小投影点确定为指根位置。
在一个可选示例中,该步骤S24b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的指尖修正单元执行。
S25b,根据修正后的指尖21位置和指根位置确定各个手指区域轮廓22的相对分布。
即识别出拇指、食指、中指、无名指、小指。可选地,对于每个手指均可根据指尖21和指根这两个点确定一条直线,5个手指则对应5条直线,这5条直线彼此之间存在一定角度,并且这5条直线相对于掌心23的方向的各不相同,根据上述直线的角度和相对于掌心23的方向即可确定出每个直线对应哪个手指。根据上述方案,通过确定的手指相对分布可以进一步确定缺少的手指以及相应的手势,并且在建模时可以忽略缺少的手指以提高建模效率。
在一个可选示例中,该步骤S25b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的分布确定单元执行。
步骤S25b可选可以包括如下步骤:
S25b1,分别确定至少一个指根位置相对于掌心23的位置的方向。
S25b2,确定至少一个指根位置相对于掌心23的位置的方向的平均方向。
S25b3,分别将至少一个指根位置相对于掌心23的位置的方向与平均方向进行比对。
S25b4,剔除偏离平均方向大于预设阈值的方向所对应的指根位置,偏离平均方向超过一定阈值的指根被判断为误检为手指的手腕,从手指集合中剔除。
S25b5,根据保留的各个指根位置相对于掌心23的位置的方向确定各个手指的相对分布。
也即根据保留的各个指根位置相对于掌心23位置的方向判断各个手指的顺序,保留的手指中,按照指根相对掌心23的方向,顺序标记为拇指、食指、中指、无名指、小指。上述方案根据手指相对于掌心的方向确定手指的相对分布,避免针对手指区域轮廓的形状进行分析,减少了形状判断操作,本方案具有较高的准确性和处理效率,由此确定出的相对分布信息可以为手势判断提供依据。
对于仅需要确定掌心23深度信息和手掌区域轮廓24的深度信息的情况,上述步骤S3可以包括如下步骤:
S31a,根据双目摄像系统的摄像参数、掌心23在第一手部图像和第二手部图像中的位置确定掌心23的深度信息。
如上,由于两张手部图像中分别只有1个掌心23,因此可以直接将二者对应起来,同理也可以直接将手掌区域轮廓24对应起来。在确定了相应的物点之后,即可根据摄像参数确定这些物点的深度信息。
在一个可选示例中,该步骤S31a可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的掌心深度确定单元执行。
S32a,将手掌区域轮廓24的边缘点投影到掌心23所在的深度,得到手掌区域轮廓24的深度信息。
在确定了掌心23的深度信息的前提下,由此即可基于该深度确定一个平面,然后将手掌区域轮廓24的边缘点全部投影到该深度平面上,也即将上述两个手部图像中的手掌区域轮廓24均投影到该平面上,该平面上则有2个手掌区域轮廓24,然后对2个手掌区域轮廓24的每对最接近的点进行匹配,由此即可实现对区域轮廓的所有的边缘点进行匹配,之后即可根据摄像参数确定这些物点的深度信息。上述实施方式首先确定掌心这一物点的深度,然后将两张手部图像中的手掌区域轮廓边缘上的点均投影到基于掌心深度建立的平面上,将投影深度作为手掌区域轮廓的深度,避免针对区域轮廓的点一一进行匹配,由此可以提高处理效率。
在一个可选示例中,该步骤S32a可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手掌区域深度确定单元执行。
对于需要确定指尖21深度信息和手指区域轮廓22的深度信息的情况,上述步骤S3可以包括如下步骤:
S31b,分别在第一手部图像和第二手部图像中确定至少一个指尖21位置与掌心23位置的相对位置向量。
在一个可选示例中,该步骤S31b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的相对位置确定单元执行。
S32b,根据相对位置向量差的模长确定两张图像中相应的至少一对指尖21,以及两张图像中相应的各对手指区域轮廓22。例如将第一手部图像和第二手部图像中的相对位置向量差的模长最小的一对指尖21位置(即相对掌心23的向量差的模长最小的一对指尖21)看作同一指尖21,进一步可以将匹配的指尖21位置对应的手指区域轮廓22进行匹配。
在一个可选示例中,该步骤S32b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手指匹配单元执行。
S33b,根据双目摄像系统的摄像参数、上述至少一对指尖21在第一手部图像和第二手部图像中的位置确定上述至少一对指尖21的深度信息。
在一个可选示例中,该步骤S33b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的指尖深度确定单元执行。
S34b,将各对手指区域轮廓22的边缘点投影到相应的至少一对指尖21所在的深度,以得到上述至少一对手指区域轮廓22的深度信息。
与上述步骤S31a和S32a类似地,即先确定指尖21深度信息,然后基于指尖21深度建立投影平面,将手指区域轮廓22边缘上的点投影到该平面上以确定边缘点的深度信息。上述实施方式首先根据指尖21位置与掌心23位置的相对位置向量对两种图像中的手指进行匹配,以确定相应的各对指尖21、手指区域轮廓22,然后首先确定指尖21这一物点的深度,并将两张手部图像中的手指区域轮廓22边缘上的点均投影到基于相应的指尖21深度建立的平面上,将投影深度作为手指区域轮廓22的深度,避免针对区域轮廓的点一一进行匹配,由此可以提高处理效率。
在一个可选示例中,该步骤S34b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手指深度确定单元执行。
关于上述步骤S4,此前的步骤已经确定了关键点和区域轮廓的深度信息,在此基础上,各关键点的三维数据、各区域轮廓的边缘三维点云数据、手部方向数据这三种数据的确定顺序有多种。由于指尖21和掌心23这种关键点在实际应用中尤为重要,为了更准确地确定关键点的三维数据,本实施例是在确定了区域点云数据和手部方向数据的基础上,才进行关键点的三维数据的确定操作。在其中一个实施方式中,可以首先确定覆盖这些关键点的区域轮廓的点云数据,例如可以包括如下步骤S41和S42中的至少一个:
S41,根据两张手部图像中的对应的至少一个手指区域轮廓22上的边缘点的位置和指尖21位置的深度建立手指区域轮廓22边缘三维点云数据,也即将手指区域轮廓22边缘上的点全部投影到基于指尖21建立的投影平面上以确定边缘点的深度信息,该深度信息与位置信息的集合即为手指区域轮廓22的三维点云数据。
在一个可选示例中,该步骤S41可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手指点云数据确定单元执行。
S42,根据两张手部图像中的手掌区域轮廓24上的边缘点的位置和掌心23位置的深度建立手掌区域轮廓24边缘三维点云数据,也即将手掌区域轮廓24边缘上的点全部投影到基于掌心23建立的投影平面上以确定边缘点的深度信息,该深度信息与位置信息的集合即为手指区域轮廓22的三维点云数据。
在一个可选示例中,该步骤S42可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手掌点云数据确定单元执行。
在另一个实施方式中,可以基于上述步骤得到的三维点云数据,得到手部方向数据, 该手部方向数据例如可以包括手指方向数据和/或手掌法向数据,例如可以包括如下步骤:
S42a,对手指区域轮廓22边缘三维点云数据进行主成分分析处理,得到主方向记为手指方向。
如前文,主成分分析是一种多元统计方法,此处是对三维点做分析,将会得到三个向量分别对应三个特征值,对应的特征值越大表示点云沿这个方向的分布越显著,主方向即对应的特征值最大的向量。
在一个可选示例中,该步骤S42a可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手指方向确定单元执行。
S42b,对手掌区域轮廓24边缘三维点云数据进行主成分分析处理,得到的最小特征值对应的特征方向记为手掌法向,即将得到的最小特征值对应的特征方向记为手掌法向。
在一个可选示例中,该步骤S42b可以由处理器调用存储器存储的相应指令执行,也可以由被处理器运行的手掌法向确定单元执行。
最后,基于手部方向数据和三维点云数据确定关键点的三维数据,可以包括如下步骤:
S42a1,将手指方向上的最大投影点作为指尖21并确定其三维数据;
S42a2,将手指方向上的最小投影点作为指根并确定其三维数据。
S42b1,根据双目摄像系统的摄像参数确定初始掌心23三维位置,此步骤可以直接根据之前确定的掌心23的深度信息和摄像参数来确定,但基于此确定的位置可能被认为是不够准确的,需要基于该初始位置进一步进行调整;
S42b2,调整初始掌心23三维位置,直至以移动后的掌心23位置为中心的三维手掌区域轮廓24投影到第一手部图像和第二手部图像上的边缘点在手掌区域轮廓24内,且投影到第一手部图像和第二手部图像上的边缘点的曼哈顿距离之和最小,将移动后的掌心23位置确定为掌心23位置,然后结合其深度信息组成掌心23的三维数据。
上述确定掌心三维数据的方式首先通过匹配二维图像中掌心并结合摄像参数确定初始掌心位置,随后将手掌区域轮廓上的点投影基于初始掌心的深度建立的平面上,根据投影的点的效果调整初始掌心位置,由此确定掌心三位数据具有更高的准确性。
上述实施方式首先确定区域点云数据,然后基于区域点云数据确定手部方向数据,最后基于区域点云数据和手部方向数据确定关键点的三维数据,由此确定的关键点的三维数据更加准确且更加稳定。
关于手势信息,步骤S4可以包括如下步骤:判断步骤S2中确定的区域轮廓的数量是否小于预设数量,当区域轮廓的数量小于预设数量时,确定缺少的区域轮廓。缺少的区域轮廓如果是手指区域轮廓22,则可以在步骤S2中确定手指间的相对分布,在确定了相对分布的情况下,即可确定缺少的是哪根手指。
可选地,当检测到的手指的个数少于5时,可以假设此时的手势是一种常用而且不难做出的手势,并在此基础上枚举可能的手势,找出最合适的一个。例如可以预先存储一个常见手指蜷缩手势表:
手指蜷缩个数     常见手势
1                01111
2                00111,01011,01101,01110,11001,11100
3                01001,01100,10001,11000
4                00001,01000,10000
5            00000
其中,一个5位二进制数表示一个手势,从最高位到最低位依次表示大拇指、食指、中指、无名指、小指的状态,1表示手指伸展,0表示手指蜷缩,如:01000表示除食指以外的其他手指都是蜷缩状态。由此即可确定出图像中缺少的手指,也即得到手势信息。上述实施方式首先根据关键区域轮廓的数量判断图像中的手部是否完整,在确定了不完整的情况下根据区域轮廓的相对分布确定缺少的区域轮廓,进而确定手势,此方式不必进行大量的枚举操作和图形分析操作,具有较高的处理效率。
本申请实施例提供的任一种手部三维数据确定方法可以由任意适当的具有数据处理能力的设备执行,包括但不限于:终端设备和服务器等。或者,本申请实施例提供的任一种手部三维数据确定方法可以由处理器执行,如处理器通过调用存储器存储的相应指令来执行本申请实施例提及的任一种手部三维数据确定方法。下文不再赘述。
本领域普通技术人员可以理解:实现上述方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成,前述的程序可以存储于一计算机可读取存储介质中,该程序在执行时,执行包括上述方法实施例的步骤;而前述的存储介质包括:ROM、RAM、磁碟或者光盘等各种可以存储程序代码的介质。
本申请的另一个实施例还提供了一种手部三维数据确定装置。图3为本申请实施例的手部三维数据确定装置的一个结构示意图,如图3所示,该装置包括:
获取单元31,用于获取双目摄像系统拍摄的第一手部图像和第二手部图像;
识别单元32,用于分别从第一手部图像和第二手部图像中识别至少一关键点以及覆盖关键点的区域轮廓;
深度确定单元33,用于根据双目摄像系统的摄像参数、从第一手部图像识别出的关键点和相应区域轮廓、以及从第二手部图像识别出的关键点和相应区域轮廓,确定上述至少一个关键点的深度信息以及区域轮廓的深度信息;
三维数据确定单元34,用于至少根据上述至少一个关键点及其深度信息、和区域轮廓及其深度信息,确定手部三维数据。
本申请实施例提供的手部三维数据确定装置,在双目摄像系统拍摄的第一手部图像和第二手部图像中识别关键点和区域轮廓,之后根据双目摄像系统的摄像参数、从第一手部图像识别出的关键点和区域轮廓、以及从第二手部图像识别出的关键点和区域轮廓,可以确定关键点的深度和区域轮廓的深度;再根据关键点和关键点的深度、区域轮廓及其深度确定手部三维数据,根据本申请提供的技术方案,只要是在二维图像中能够识别出的物点、区域轮廓等内容,均可以通过相互匹配确定深度,并进一步确定其三维数据,由此可以实现基于单帧的二维图像确定手部三维信息,并使确定的手部三维数据更加准确且更丰富,并且,本申请实施例确定手部三维数据的技术方案具有较高的效率和准确性。
可选地,上述关键点例如可以包括但不限于以下任意一项或多项:指尖、指关节点、掌心。
可选地,上述覆盖关键点的区域轮廓例如可以包括但不限于以下任意一项或多项:覆盖指尖和/或指关节点的手指区域轮廓,覆盖掌心的手掌区域轮廓。
可选地,上述手部三维数据例如可以包括但不限于以下任意一项或多项:关键点的三维数据、区域轮廓的边缘三维点云数据、手部方向数据、手势信息。
可选地,识别单元可以包括:
掌心识别单元,用于分别在第一手部图像和第二手部图像中识别掌心;
手掌识别单元,用于分别在第一手部图像和第二手部图像中根据识别的掌心的位置确定覆盖掌心的手掌区域轮廓。上述可选实施方式在二值化处理后的手部图像中寻找最大连通区域,进而基于最大连通区域确定掌心23和手掌区域轮廓24,本方案通过线条特征寻找最大连通区域,然后基于最大连通区域来确定手掌,以减少图像颜色和灰度的影响,由此可以提高处理效率。
可选地,识别单元可以包括:
掌心识别单元,用于分别在第一手部图像和第二手部图像中识别掌心;
指尖识别单元,用于分别在第一手部图像和第二手部图像中根据识别的掌心的位置确定至少一指尖和/或指关节点位置;
手指识别单元,用于分别根据第一手部图像和第二手部图像中的指尖和/或指关节点位置,确定覆盖指尖和/或指关节点的手指区域轮廓。上述可选实施方式在二值化处理后的手部图像中确定掌心,进而在最大连通区域边缘上的点中确定与掌心相距最远的点,然后基于该出发点寻找疑似手指的区域轮廓,在找到的疑似手指的区域轮廓中进一步根据轮廓形状排除非手指区域轮廓,以筛选出手指区域轮廓,由此可以排除手腕对识别过程的影响,并在找到手指区域轮廓后,将相应的出发点确定为指尖,由此可以提高准确率和处理效率。
可选地,上述手部三维数据确定装置还可以包括:
指尖修正单元,用于分别在第一手部图像和第二手部图像中根据至少一个手指区域轮廓修正指尖位置并确定指根位置;
分布确定单元,用于根据修正后的指尖位置和指根位置确定至少一个手指区域轮廓的相对分布。根据上述方案,通过确定的手指相对分布可以进一步确定缺少的手指以及相应的手势,并且在建模时可以忽略缺少的手指以提高建模效率。
可选地,指尖修正单元可以包括:
手指方向确定单元,用于分别对至少一个手指区域轮廓上的边缘点进行主成分分析处理,将得到的主方向作为手指方向;
指尖确定单元,用于分别将至少一个手指方向上的最大投影点确定为修正后的指尖位置;
指根确定单元,用于分别将至少一个手指方向上的最小投影点确定为指根位置。
可选地,分布确定单元可以包括:
指根相对方向确定单元,用于分别确定至少一个指根位置相对于掌心的位置的方向;
平均方向确定单元,用于确定至少一个指根位置相对于掌心的位置的方向的平均方向;
方向比对单元,用于分别将至少一个指根位置相对于掌心的位置的方向与平均方向进行比对;
剔除单元,用于剔除偏离平均方向大于预设阈值的方向所对应的指根位置;
分布识别单元,用于根据保留的至少一个指根位置相对于掌心的位置的方向确定个手指的相对分布。
上述方案根据手指相对于掌心的方向确定手指的相对分布,避免针对手指区域轮廓的形状进行分析,减少了形状判断操作,本方案具有较高的准确性和处理效率,由此确定出的相对分布信息可以为手势判断提供依据。
可选地,掌心识别单元可以包括:
连通区域确定单元,用于分别在经过二值化处理后第一手部图像和第二手部图像中确定最大连通区域轮廓;
掌心位置确定单元,用于分别根据第一手部图像和第二手部图像中的最大连通区域轮廓确定掌心的位置。
可选地,手掌识别单元可以包括:
手掌半径确定单元,用于分别根据掌心的位置到2个最大连通区域的边缘的距离确定手掌半径;
手掌区域轮廓确定单元,用于分别根据手掌半径确定手掌区域轮廓。
可选地,掌心识别单元可以包括:
连通区域确定单元,用于分别在经过二值化处理后第一手部图像和第二手部图像中确定最大连通区域;
掌心位置确定单元,用于分别根据第一手部图像和第二手部图像中最大连通区域确定掌心的位置;
相应地,指尖识别单元可以包括:
远点确定单元,用于在最大连通区域内确定距掌心位置最远的点;
疑似区域确定单元,用于基于距掌心位置最远的点在最大连通区域内确定疑似手指区域轮廓;
指尖确定单元,用于根据至少一个疑似手指区域轮廓的形状从至少一个距掌心位置最远的点中确定指尖并获取其位置;
手指识别单元包括:
手指确定单元,用于根据至少一个疑似手指区域轮廓的形状从至少一个疑似手指区域轮廓中确定手指区域轮廓。
可选地,深度确定单元可以包括:
掌心深度确定单元,用于根据双目摄像系统的摄像参数、掌心在第一手部图像和第二手部图像中的位置确定掌心的深度信息;
手掌区域深度确定单元,用于将手掌区域轮廓的边缘点投影到掌心所在的深度,得到手掌区域轮廓的深度信息。
上述实施方式首先确定掌心这一物点的深度,然后将两张手部图像中的手掌区域轮廓边缘上的点均投影到基于掌心深度建立的平面上,将投影深度作为手掌区域轮廓的深度,避免针对区域轮廓的点一一进行匹配,由此可以提高处理效率。
可选地,深度确定单元可以包括:
相对位置确定单元,用于分别在第一手部图像和第二手部图像中确定至少一个指尖位置与掌心位置的相对位置向量;
手指匹配单元,用于根据相对位置向量差的模长确定两张图像中相应的至少一对指尖,以及两张图像中相应的至少一对手指区域轮廓;
指尖深度确定单元,用于根据双目摄像系统的摄像参数、至少一对指尖在第一手部图像和第二手部图像中的位置确定至少一对指尖的深度信息;
手指深度确定单元,用于将至少一对手指区域轮廓的边缘点投影到相应的对指尖所在的深度,以得到至少一对手指区域轮廓的深度信息。上述实施方式首先根据指尖位置与掌心位置的相对位置向量对两种图像中的手指进行匹配,以确定相应的至少一对指尖、手指 区域轮廓,然后首先确定指尖这一物点的深度,并将两张手部图像中的手指区域轮廓边缘上的点均投影到基于相应的指尖深度建立的平面上,将投影深度作为手指区域轮廓的深度,避免针对区域轮廓的点一一进行匹配,由此可以提高处理效率。
可选地,三维数据确定单元可以包括:
手指点云数据确定单元,用于根据两张手部图像中的对应的手指区域轮廓上的边缘点的位置和指尖位置的深度建立手指区域轮廓边缘三维点云数据;和/或
手掌点云数据确定单元,用于根据两张手部图像中的手掌区域轮廓上的边缘点的位置和掌心位置的深度建立手掌区域轮廓边缘三维点云数据。
可选地,三维数据确定单元可以包括:
手指方向确定单元,用于对手指区域轮廓边缘三维点云数据进行主成分分析处理,得到主方向记为手指方向;和/或
手掌法向确定单元,用于对手掌区域轮廓边缘三维点云数据进行主成分分析处理,得到的最小特征值对应的特征方向记为手掌法向。
可选地,三维数据确定单元可以包括:
指尖三维数据确定单元,用于将手指方向上的最大投影点作为指尖并确定其三维数据;和/或
掌心三维位置确定单元,用于根据双目摄像系统的摄像参数确定初始掌心三维位置;
掌心三维数据确定单元,用于调整初始掌心三维位置,得到调整后的掌心三维位置并确定其三维数据。上述确定掌心三维数据的方式首先通过匹配二维图像中掌心并结合摄像参数确定初始掌心位置,随后将手掌区域轮廓上的点投影基于初始掌心的深度建立的平面上,根据投影的点的效果调整初始掌心位置,由此确定掌心三位数据具有更高的准确性。
上述实施方式首先确定区域点云数据,然后基于区域点云数据确定手部方向数据,最后基于区域点云数据和手部方向数据确定关键点的三维数据,由此确定的关键点的三维数据更加准确且更加稳定。
本申请实施例还提供了一种电子设备,例如可以是移动终端、个人计算机(PC)、平板电脑、服务器等。该电子设备包括:至少一个处理器;以及与上述至少一个处理器通信连接的存储器;其中,存储器存储有可被上述至少一个处理器执行的指令,该指令被至少一个处理器执行,以使上述至少一个处理器执行本申请任一是实施例所述手部三维数据确定方法对应的操作。
图4为本申请实施例的电子设备的一个结构示意图。下面参考图4,其示出了适于用来实现本申请实施例的终端设备或服务器的电子设备400的结构示意图:如图4所示,计算机系统400包括一个或多个处理器、通信部等,上述一个或多个处理器例如:一个或多个中央处理单元(CPU)401,和/或一个或多个图像处理器(GPU)413等,处理器可以根据存储在只读存储器(ROM)402中的可执行指令或者从存储部分408加载到随机访问存储器(RAM)403中的可执行指令而执行种适当的动作和处理。通信部412可包括但不限于网卡,上述网卡可包括但不限于IB(Infiniband)网卡。
处理器可与ROM 402和/或RAM 403通信以执行可执行指令,通过总线404与通信部412相连、并经通信部412与其他目标设备通信,从而完成本申请实施例提供的任一项方法对应的操作,例如,获取双目摄像系统拍摄的第一手部图像和第二手部图像;分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;根 据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。
此外,在RAM 403中,还可存储有装置操作所需的程序和数据。CPU 401、ROM 402以及RAM 403通过总线404彼此相连。在有RAM 403的情况下,ROM 402为可选模块。RAM 403存储可执行指令,或在运行时向ROM 402中写入可执行指令,可执行指令使CPU401执行上述通信方法对应的操作。输入/输出(I/O)接口405也连接至总线404。通信部412可以集成设置,也可以设置为具有多个子模块(例如多个IB网卡),并在总线404链接上。
以下部件连接至I/O接口405:包括键盘、鼠标等的输入部分406;包括诸如阴极射线管(CRT)、液晶显示器(LCD)等以及扬声器等的输出部分407;包括硬盘等的存储部分408;以及包括诸如LAN卡、调制解调器等的网络接口卡的通信部分409。通信部分409经由诸如因特网的网络执行通信处理。驱动器410也根据需要连接至I/O接口405。可拆卸介质411,诸如磁盘、光盘、磁光盘、半导体存储器等等,根据需要安装在驱动器410上,以便于从其上读出的计算机程序根据需要被安装入存储部分408。
需要说明的,如图4所示的架构仅为一种可选实现方式,在实践过程中,可根据实际需要对上述图4的部件数量和类型进行选择、删减、增加或替换;在不同功能部件设置上,也可采用分离设置或集成设置等实现方式,例如GPU 413和CPU 401可分离设置或者可将GPU 413集成在CPU 401上,通信部可分离设置,也可集成设置在CPU 401或GPU 413上,等等。这些可替换的实施方式均落入本申请公开的保护范围。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括有形地包含在机器可读介质上的计算机程序,计算机程序包含用于执行流程图所示的方法的程序代码,程序代码可包括对应执行本申请实施例提供的方法步骤对应的指令,例如,获取双目摄像系统拍摄的第一手部图像和第二手部图像;分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。在这样的实施例中,该计算机程序可以通过通信部分409从网络上被下载和安装,和/或从可拆卸介质411被安装。在该计算机程序被CPU 401执行时,执行本申请的方法中限定的上述功能。
另外,本申请实施例还提供了一种计算机可读存储介质,其上存储有计算机指令,所述指令被执行时实现本申请任一实施例所述手部三维数据确定方法中各步骤的操作。
另外,本申请实施例还提供了一种计算机程序,包括计算机可读代码,当所述计算机可读代码在设备上运行时,所述设备中的处理器执行用于实现本申请任一实施例所述手部三维数据确定方法中各步骤的指令。
可能以许多方式来实现本申请的方法和装置、设备。例如,可通过软件、硬件、固件或者软件、硬件、固件的任何组合来实现本申请的方法和装置、设备。用于方法的步骤的上述顺序仅是为了进行说明,本申请的方法的步骤不限于以上描述的顺序,除非以其它方 式特别说明。此外,在一些实施例中,还可将本申请实施为记录在记录介质中的程序,这些程序包括用于实现根据本申请的方法的机器可读指令。因而,本申请还覆盖存储用于执行根据本申请的方法的程序的记录介质。
本申请的描述是为了示例和描述起见而给出的,而并不是无遗漏的或者将本申请限于所公开的形式。很多修改和变化对于本领域的普通技术人员而言是显然的。选择和描述实施例是为了更好说明本申请的原理和实际应用,并且使本领域的普通技术人员能够理解本申请从而设计适于特定用途的带有种修改的种实施例。

Claims (37)

  1. 一种手部三维数据确定方法,其特征在于,包括:
    获取双目摄像系统拍摄的第一手部图像和第二手部图像;
    分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;
    根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;
    至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。
  2. 根据权利要求1所述的方法,其特征在于,所述关键点包括以下任意一项或多项:指尖、指关节点、掌心。
  3. 根据权利要求1或2所述的方法,其特征在于,所述覆盖所述关键点的区域轮廓包括以下任意一项或多项:覆盖指尖和/或指关节点的手指区域轮廓、覆盖掌心的手掌区域轮廓。
  4. 根据权利要求1-3任一所述方法,其特征在于,所述手部三维数据包括以下任意一项或多项:所述关键点的三维数据、所述区域轮廓的边缘三维点云数据、手部方向数据、手势信息。
  5. 根据权利要求1-4任一所述方法,其特征在于,所述分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓,包括:
    分别在所述第一手部图像和第二手部图像中识别掌心;
    分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定覆盖所述掌心的手掌区域轮廓。
  6. 根据权利要求1-4任一所述方法,其特征在于,所述分别从所述第一手部图像和第二手部图像中识别关键点以及覆盖所述关键点的区域轮廓,包括:
    分别在所述第一手部图像和第二手部图像中识别掌心;
    分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定至少一指尖和/或指关节点位置;
    分别根据所述第一手部图像和第二手部图像中的所述指尖和/或指关节点位置,确定覆盖所述指尖和/或指关节点的手指区域轮廓。
  7. 根据权利要求6所述的方法,其特征在于,在所述分别根据所述第一手部图像和第二手部图像中的所述指尖位置确定手指区域轮廓后,还包括:
    分别在所述第一手部图像和第二手部图像中根据所述手指区域轮廓修正所述指尖位置并确定指根位置;
    根据修正后的指尖位置和所述指根位置确定所述手指区域轮廓的相对分布。
  8. 根据权利要求7所述的方法,其特征在于,所述分别在所述第一手部图像和第二手部图像中根据手指区域轮廓修正所述指尖位置并确定指根位置,包括:
    分别对所述手指区域轮廓上的边缘点进行主成分分析处理,将得到的主方向作为手指方向;
    分别将所述手指方向上的最大投影点确定为修正后的指尖位置;
    分别将所述手指方向上的最小投影点确定为指根位置。
  9. 根据权利要求7所述的方法,其特征在于,所述根据修正后的指尖位置和所述指根位置确定所述手指区域轮廓的相对分布,包括:
    分别确定所述指根位置相对于所述掌心的位置的方向;
    确定所述指根位置相对于所述掌心的位置的方向的平均方向;
    分别将所述指根位置相对于所述掌心的位置的方向与所述平均方向进行比对;
    剔除偏离所述平均方向大于预设阈值的方向所对应的指根位置;
    根据保留的指根位置相对于所述掌心的位置的方向确定手指的相对分布。
  10. 根据权利要求5所述的方法,其特征在于,所述分别在所述第一手部图像和第二手部图像中识别掌心,包括:
    分别在经过二值化处理后所述第一手部图像和第二手部图像中确定最大连通区域轮廓;
    分别根据所述第一手部图像和第二手部图像中的最大连通区域轮廓确定所述掌心的位置。
  11. 根据权利要10所述的方法,其特征在于,所述分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定所述覆盖所述掌心的手掌区域轮廓,包括:
    分别根据所述掌心的位置到2个所述最大连通区域的边缘的距离确定手掌半径;
    分别根据所述手掌半径确定手掌区域轮廓。
  12. 根据权利要求6所述的方法,其特征在于,所述分别在所述第一手部图像和第二手部图像中识别掌心,包括:分别在经过二值化处理后所述第一手部图像和第二手部图像中确定最大连通区域;分别根据所述第一手部图像和第二手部图像中最大连通区域确定所述掌心的位置;
    所述分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定至少一指尖和/或指关节点位置,包括:在所述最大连通区域内确定距所述掌心位置最远的点;基于所述距所述掌心位置最远的点在所述最大连通区域内确定疑似手指区域轮廓;根据所述疑似手指区域轮廓的形状从所述距所述掌心位置最远的点中确定指尖并获取其位置;
    所述分别对所述第一手部图像和第二手部图像中的根据所述指尖位置确定手指区域轮廓,包括:
    根据所述疑似手指区域轮廓的形状从所述疑似手指区域轮廓中确定手指区域轮廓。
  13. 根据权利要求1-4任一所述方法,其特征在于,根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓以及从所述第二手部图像识别出的相应关键点和相应区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息,包括:
    根据所述双目摄像系统的摄像参数、掌心在所述第一手部图像和第二手部图像中的位置确定所述掌心的深度信息;
    将手掌区域轮廓的边缘点投影到所述掌心所在的深度,得到手掌区域轮廓的深度信息。
  14. 根据权利要求1-4任一所述方法,其特征在于,根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深 度信息,包括:
    分别在所述第一手部图像和第二手部图像中确定至少一个指尖位置与掌心位置的相对位置向量;
    根据所述相对位置向量差的模长确定两张图像中相应的至少一对指尖,以及两张图像中相应的至少一对手指区域轮廓;
    根据所述双目摄像系统的摄像参数、至少一对指尖在所述第一手部图像和第二手部图像中的位置确定所述至少一对指尖的深度信息;
    将所述至少一对手指区域轮廓的边缘点投影到相应的所述至少一对指尖所在的深度,以得到所述至少一对手指区域轮廓的深度信息。
  15. 根据权利要求4所述的方法,其特征在于,确定所述区域轮廓的边缘三维点云数据,包括:
    根据两张手部图像中的对应的所述手指区域轮廓上的边缘点的位置和所述指尖位置的深度建立手指区域轮廓边缘三维点云数据;和/或
    根据两张手部图像中的所述手掌区域轮廓上的边缘点的位置和所述掌心位置的深度建立手掌区域轮廓边缘三维点云数据。
  16. 根据权利要求15所述的方法,其特征在于,确定所述手部方向数据,包括:
    对所述手指区域轮廓边缘三维点云数据进行主成分分析处理,得到主方向记为手指方向;和/或
    对所述手掌区域轮廓边缘三维点云数据进行主成分分析处理,得到的最小特征值对应的特征方向记为手掌法向。
  17. 根据权利要求16所述的方法,其特征在于,确定所述关键点的三维数据,包括:
    将所述手指方向上的最大投影点作为指尖并确定其三维数据;和/或
    根据所述双目摄像系统的摄像参数确定初始掌心三维位置;
    调整所述初始掌心三维位置,得到调整后的掌心三维位置并确定其三维数据。
  18. 一种手部三维数据确定装置,其特征在于,包括:
    获取单元,用于获取双目摄像系统拍摄的第一手部图像和第二手部图像;
    识别单元,用于分别从所述第一手部图像和第二手部图像中识别至少一关键点以及覆盖所述关键点的区域轮廓;
    深度确定单元,用于根据所述双目摄像系统的摄像参数、从所述第一手部图像识别出的所述关键点和所述区域轮廓、以及从所述第二手部图像识别出的所述关键点和所述区域轮廓,确定所述关键点的深度信息以及所述区域轮廓的深度信息;
    三维数据确定单元,用于至少根据所述关键点及其深度信息、和所述区域轮廓及其深度信息,确定手部三维数据。
  19. 根据权利要求18所述的装置,其特征在于,所述关键点包括以下任意一项或多项:指尖、指关节点、掌心。
  20. 根据权利要求18或19所述的装置,其特征在于,所述覆盖所述关键点的区域轮廓包括以下任意一项或多项:覆盖指尖和/或指关节点的手指区域轮廓、覆盖掌心的手掌区域轮廓。
  21. 根据权利要求18-20任一所述装置,其特征在于,所述手部三维数据包括以下任意一项或多项:所述关键点的三维数据、所述区域轮廓的边缘三维点云数据、手部方向数据、 手势信息。
  22. 根据权利要求18-21任一所述装置,其特征在于,所述识别单元包括:
    掌心识别单元,用于分别在所述第一手部图像和第二手部图像中识别掌心;
    手掌识别单元,用于分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定覆盖所述掌心的手掌区域轮廓。
  23. 根据权利要求18-21任一所述装置,其特征在于,所述识别单元包括:
    掌心识别单元,用于分别在所述第一手部图像和第二手部图像中识别掌心;
    指尖识别单元,用于分别在所述第一手部图像和第二手部图像中根据识别的所述掌心的位置确定至少一指尖和/或指关节点位置;
    手指识别单元,用于分别根据所述第一手部图像和第二手部图像中的所述指尖和/或指关节点位置,确定覆盖所述指尖和/或指关节点的手指区域轮廓。
  24. 根据权利要求23所述的装置,其特征在于,还包括:
    指尖修正单元,用于分别在所述第一手部图像和第二手部图像中根据所述手指区域轮廓修正所述指尖位置并确定指根位置;
    分布确定单元,用于根据修正后的指尖位置和所述指根位置确定所述手指区域轮廓的相对分布。
  25. 根据权利要求24所述的装置,其特征在于,所述指尖修正单元包括:
    手指方向确定单元,用于分别对所述手指区域轮廓上的边缘点进行主成分分析处理,将得到的主方向作为手指方向;
    指尖确定单元,用于分别将所述手指方向上的最大投影点确定为修正后的指尖位置;
    指根确定单元,用于分别将所述手指方向上的最小投影点确定为指根位置。
  26. 根据权利要求24所述的装置,其特征在于,所述分布确定单元包括:
    指根相对方向确定单元,用于分别确定个所述指根位置相对于所述掌心的位置的方向;
    平均方向确定单元,用于确定所述指根位置相对于所述掌心的位置的方向的平均方向;
    方向比对单元,用于分别将所述指根位置相对于所述掌心的位置的方向与所述平均方向进行比对;
    剔除单元,用于剔除偏离所述平均方向大于预设阈值的方向所对应的指根位置;
    分布识别单元,用于根据保留的指根位置相对于所述掌心的位置的方向确定手指的相对分布。
  27. 根据权利要求22所述的装置,其特征在于,所述掌心识别单元包括:
    连通区域确定单元,用于分别在经过二值化处理后所述第一手部图像和第二手部图像中确定最大连通区域轮廓;
    掌心位置确定单元,用于分别根据所述第一手部图像和第二手部图像中的最大连通区域轮廓确定所述掌心的位置。
  28. 根据权利要27所述的装置,其特征在于,所述手掌识别单元包括:
    手掌半径确定单元,用于分别根据所述掌心的位置到2个所述最大连通区域的边缘的距离确定手掌半径;
    手掌区域轮廓确定单元,用于分别根据所述手掌半径确定手掌区域轮廓。
  29. 根据权利要求23所述的装置,其特征在于,所述掌心识别单元包括:
    连通区域确定单元,用于分别在经过二值化处理后所述第一手部图像和第二手部图像 中确定最大连通区域;
    掌心位置确定单元,用于分别根据所述第一手部图像和第二手部图像中最大连通区域确定所述掌心的位置;
    所述指尖识别单元包括:
    远点确定单元,用于在所述最大连通区域内确定距所述掌心位置最远的点;
    疑似区域确定单元,用于基于所述距所述掌心位置最远的点在所述最大连通区域内确定疑似手指区域轮廓;
    指尖确定单元,用于根据所述疑似手指区域轮廓的形状从所述距所述掌心位置最远的点中确定指尖并获取其位置;
    所述手指识别单元包括:
    手指确定单元,用于根据所述疑似手指区域轮廓的形状从所述疑似手指区域轮廓中确定手指区域轮廓。
  30. 根据权利要求18-21任一所述装置,其特征在于,所述深度确定单元包括:
    掌心深度确定单元,用于根据所述双目摄像系统的摄像参数、掌心在所述第一手部图像和第二手部图像中的位置确定所述掌心的深度信息;
    手掌区域深度确定单元,用于将手掌区域轮廓的边缘点投影到所述掌心所在的深度,得到手掌区域轮廓的深度信息。
  31. 根据权利要求18-21任一所述装置,其特征在于,所述深度确定单元包括:
    相对位置确定单元,用于分别在所述第一手部图像和第二手部图像中确定至少一个指尖位置与掌心位置的相对位置向量;
    手指匹配单元,用于根据所述相对位置向量差的模长确定两张图像中相应的至少一对指尖,以及两张图像中相应的至少一对手指区域轮廓;
    指尖深度确定单元,用于根据所述双目摄像系统的摄像参数、至少一对指尖在所述第一手部图像和第二手部图像中的位置确定所述至少一对指尖的深度信息;
    手指深度确定单元,用于将至少一对手指区域轮廓的边缘点投影到相应的所述至少一对指尖所在的深度,以得到所述至少一对手指区域轮廓的深度信息。
  32. 根据权利要求21所述的装置,其特征在于,所述三维数据确定单元包括:
    手指点云数据确定单元,用于根据两张手部图像中的对应的所述手指区域轮廓上的边缘点的位置和所述指尖位置的深度建立手指区域轮廓边缘三维点云数据;和/或
    手掌点云数据确定单元,用于根据两张手部图像中的所述手掌区域轮廓上的边缘点的位置和所述掌心位置的深度建立手掌区域轮廓边缘三维点云数据。
  33. 根据权利要求32所述的装置,其特征在于,所述三维数据确定单元包括:
    手指方向确定单元,用于对所述手指区域轮廓边缘三维点云数据进行主成分分析处理,得到主方向记为手指方向;和/或
    手掌法向确定单元,用于对所述手掌区域轮廓边缘三维点云数据进行主成分分析处理,得到的最小特征值对应的特征方向记为手掌法向。
  34. 根据权利要求33所述的装置,其特征在于,所述三维数据确定单元包括:
    指尖三维数据确定单元,用于将所述手指方向上的最大投影点作为指尖并确定其三维数据;和/或
    掌心三维位置确定单元,用于根据所述双目摄像系统的摄像参数确定初始掌心三维位 置;
    掌心三维数据确定单元,用于调整所述初始掌心三维位置,得到调整后的掌心三维位置并确定其三维数据。
  35. 一种电子设备,其特征在于,包括:至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器执行权利要求1-17任一所述方法对应的操作。
  36. 一种计算机可读存储介质,其上存储有计算机指令,其特征在于,所述指令被执行时实现权利要求1-17任一所述方法中各步骤的操作。
  37. 一种计算机程序,包括计算机可读代码,其特征在于,当所述计算机可读代码在设备上运行时,所述设备中的处理器执行用于实现权利要求1-17任一所述方法中各步骤的指令。
PCT/CN2018/080960 2017-03-29 2018-03-28 手部三维数据确定方法、装置及电子设备 Ceased WO2018177337A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US16/451,077 US11120254B2 (en) 2017-03-29 2019-06-25 Methods and apparatuses for determining hand three-dimensional data

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201710198505.7A CN108230383B (zh) 2017-03-29 2017-03-29 手部三维数据确定方法、装置及电子设备
CN201710198505.7 2017-03-29

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US16/451,077 Continuation US11120254B2 (en) 2017-03-29 2019-06-25 Methods and apparatuses for determining hand three-dimensional data

Publications (1)

Publication Number Publication Date
WO2018177337A1 true WO2018177337A1 (zh) 2018-10-04

Family

ID=62656518

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/080960 Ceased WO2018177337A1 (zh) 2017-03-29 2018-03-28 手部三维数据确定方法、装置及电子设备

Country Status (3)

Country Link
US (1) US11120254B2 (zh)
CN (1) CN108230383B (zh)
WO (1) WO2018177337A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115641615A (zh) * 2022-11-25 2023-01-24 湖南工商大学 复杂背景下闭合手掌感兴趣区域提取方法

Families Citing this family (27)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108230383B (zh) * 2017-03-29 2021-03-23 北京市商汤科技开发有限公司 手部三维数据确定方法、装置及电子设备
CN108197596B (zh) * 2018-01-24 2021-04-06 京东方科技集团股份有限公司 一种手势识别方法和装置
CN109063653A (zh) * 2018-08-07 2018-12-21 北京字节跳动网络技术有限公司 图像处理方法和装置
CN110909580B (zh) * 2018-09-18 2022-06-10 北京市商汤科技开发有限公司 数据处理方法及装置、电子设备及存储介质
CN113039550B (zh) * 2018-10-10 2024-08-02 深圳市道通智能航空技术股份有限公司 手势识别方法、vr视角控制方法以及vr系统
JP7280032B2 (ja) * 2018-11-27 2023-05-23 ローム株式会社 入力デバイス、自動車
CN109840500B (zh) * 2019-01-31 2021-07-02 深圳市商汤科技有限公司 一种三维人体姿态信息检测方法及装置
CN110348524B (zh) 2019-07-15 2022-03-04 深圳市商汤科技有限公司 一种人体关键点检测方法及装置、电子设备和存储介质
CN111401219B (zh) * 2020-03-10 2023-04-28 厦门熵基科技有限公司 一种手掌关键点检测方法和装置
CN111953933B (zh) * 2020-07-03 2022-07-05 北京中安安博文化科技有限公司 一种确定火灾区域的方法、装置、介质和电子设备
US20220050527A1 (en) * 2020-08-12 2022-02-17 Himax Technologies Limited Simulated system and method with an input interface
CN112013865B (zh) * 2020-08-28 2022-08-30 北京百度网讯科技有限公司 确定交通卡口的方法、系统、电子设备以及介质
CN112198962B (zh) * 2020-09-30 2023-04-28 聚好看科技股份有限公司 一种与虚拟现实设备交互的方法及虚拟现实设备
CN112183388B (zh) * 2020-09-30 2024-07-23 抖音视界有限公司 图像处理方法、装置、设备和介质
CN112215134A (zh) * 2020-10-10 2021-01-12 北京华捷艾米科技有限公司 手势跟踪方法及装置
CN112258538A (zh) * 2020-10-29 2021-01-22 深兰科技(上海)有限公司 人体三维数据的获取方法和装置
CN113393572B (zh) * 2021-06-17 2023-07-21 北京千丁互联科技有限公司 点云数据生成方法、装置、移动终端和可读存储介质
EP4361771A4 (en) * 2021-07-17 2024-07-03 Huawei Technologies Co., Ltd. GESTURE RECOGNITION METHOD AND APPARATUS, SYSTEM AND VEHICLE
CN113780201B (zh) * 2021-09-15 2022-06-10 墨奇科技(北京)有限公司 手部图像的处理方法及装置、设备和介质
CN113756059B (zh) * 2021-09-29 2023-09-29 四川虹美智能科技有限公司 物联网安全洗衣机系统
CN116193243B (zh) * 2021-11-25 2024-03-22 荣耀终端有限公司 拍摄方法和电子设备
CN117173737A (zh) * 2022-05-28 2023-12-05 深圳市光鉴科技有限公司 一种具有对齐功能的非同源双目相机
CN117173741A (zh) * 2022-05-28 2023-12-05 深圳市光鉴科技有限公司 一种基于非同源双目的刷掌识别两图对齐方法
CN117173738B (zh) * 2022-05-28 2026-05-08 深圳市光鉴科技有限公司 一种具有对齐功能的非同源双目相机
CN117234405A (zh) * 2022-06-07 2023-12-15 北京小米移动软件有限公司 信息输入方法以及装置、电子设备及存储介质
US12482171B2 (en) * 2023-01-06 2025-11-25 Snap Inc. Natural hand rendering in XR systems
TWI886443B (zh) * 2023-02-20 2025-06-11 瑞軒科技股份有限公司 電子裝置可讀取媒體、顯示器及其操作方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20040193413A1 (en) * 2003-03-25 2004-09-30 Wilson Andrew D. Architecture for controlling a computer using hand gestures
CN102799318A (zh) * 2012-08-13 2012-11-28 深圳先进技术研究院 一种基于双目立体视觉的人机交互方法及系统
CN102831407A (zh) * 2012-08-22 2012-12-19 中科宇博(北京)文化有限公司 仿生机械恐龙的视觉识别系统实现方法
CN102902355A (zh) * 2012-08-31 2013-01-30 中国科学院自动化研究所 移动设备的空间交互方法
CN102981623A (zh) * 2012-11-30 2013-03-20 深圳先进技术研究院 触发输入指令的方法及系统

Family Cites Families (32)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6417969B1 (en) * 1988-07-01 2002-07-09 Deluca Michael Multiple viewer headset display apparatus and method with second person icon display
US6198485B1 (en) * 1998-07-29 2001-03-06 Intel Corporation Method and apparatus for three-dimensional input entry
US7660458B1 (en) * 2004-12-14 2010-02-09 Google Inc. Three-dimensional model construction using unstructured pattern
US7983451B2 (en) * 2006-06-30 2011-07-19 Motorola Mobility, Inc. Recognition method using hand biometrics with anti-counterfeiting
US8005263B2 (en) * 2007-10-26 2011-08-23 Honda Motor Co., Ltd. Hand sign recognition using label assignment
US20110102570A1 (en) * 2008-04-14 2011-05-05 Saar Wilf Vision based pointing device emulation
KR101581954B1 (ko) * 2009-06-25 2015-12-31 삼성전자주식회사 실시간으로 피사체의 손을 검출하기 위한 장치 및 방법
TWI434225B (zh) * 2011-01-28 2014-04-11 Nat Univ Chung Cheng Stereo Matching Method Using Quantitative Operation of Image Intensity Value
US8971572B1 (en) * 2011-08-12 2015-03-03 The Research Foundation For The State University Of New York Hand pointing estimation for human computer interaction
US9477303B2 (en) * 2012-04-09 2016-10-25 Intel Corporation System and method for combining three-dimensional tracking with a three-dimensional display for a user interface
US9239624B2 (en) * 2012-04-13 2016-01-19 Nokia Technologies Oy Free hand gesture control of automotive user interface
US9111135B2 (en) * 2012-06-25 2015-08-18 Aquifi, Inc. Systems and methods for tracking human hands using parts based template matching using corresponding pixels in bounded regions of a sequence of frames that are a specified distance interval from a reference camera
EP2872967B1 (en) * 2012-07-13 2018-11-21 Sony Depthsensing Solutions SA/NV Method and system for detecting hand-related parameters for human-to-computer gesture-based interaction
US9870056B1 (en) * 2012-10-08 2018-01-16 Amazon Technologies, Inc. Hand and hand pose detection
TWI516093B (zh) * 2012-12-22 2016-01-01 財團法人工業技術研究院 影像互動系統、手指位置的偵測方法、立體顯示系統以及立體顯示器的控制方法
JP2014186715A (ja) * 2013-02-21 2014-10-02 Canon Inc 情報処理装置、情報処理方法
TWI649675B (zh) * 2013-03-28 2019-02-01 新力股份有限公司 Display device
JP6044426B2 (ja) * 2013-04-02 2016-12-14 富士通株式会社 情報操作表示システム、表示プログラム及び表示方法
JP2015119373A (ja) * 2013-12-19 2015-06-25 ソニー株式会社 画像処理装置および方法、並びにプログラム
WO2015139750A1 (en) * 2014-03-20 2015-09-24 Telecom Italia S.P.A. System and method for motion capture
JP6548518B2 (ja) * 2015-08-26 2019-07-24 株式会社ソニー・インタラクティブエンタテインメント 情報処理装置および情報処理方法
KR101745406B1 (ko) * 2015-09-03 2017-06-12 한국과학기술연구원 깊이 영상 기반의 손 제스처 인식 장치 및 방법
US10146375B2 (en) * 2016-07-01 2018-12-04 Intel Corporation Feature characterization from infrared radiation
US10380796B2 (en) * 2016-07-19 2019-08-13 Usens, Inc. Methods and systems for 3D contour recognition and 3D mesh generation
US11244467B2 (en) * 2016-11-07 2022-02-08 Sony Corporation Information processing apparatus, information processing mei'hod, and recording medium
CN106931910B (zh) * 2017-03-24 2019-03-05 南京理工大学 一种基于多模态复合编码和极线约束的高效三维图像获取方法
CN108230383B (zh) * 2017-03-29 2021-03-23 北京市商汤科技开发有限公司 手部三维数据确定方法、装置及电子设备
JP6955369B2 (ja) * 2017-05-22 2021-10-27 キヤノン株式会社 情報処理装置、情報処理装置の制御方法及びプログラム
CN109558000B (zh) * 2017-09-26 2021-01-22 京东方科技集团股份有限公司 一种人机交互方法及电子设备
US10956714B2 (en) * 2018-05-18 2021-03-23 Beijing Sensetime Technology Development Co., Ltd Method and apparatus for detecting living body, electronic device, and storage medium
US20200160548A1 (en) * 2018-11-16 2020-05-21 Electronics And Telecommunications Research Institute Method for determining disparity of images captured multi-baseline stereo camera and apparatus for the same
US11520409B2 (en) * 2019-04-11 2022-12-06 Samsung Electronics Co., Ltd. Head mounted display device and operating method thereof

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20040193413A1 (en) * 2003-03-25 2004-09-30 Wilson Andrew D. Architecture for controlling a computer using hand gestures
CN102799318A (zh) * 2012-08-13 2012-11-28 深圳先进技术研究院 一种基于双目立体视觉的人机交互方法及系统
CN102831407A (zh) * 2012-08-22 2012-12-19 中科宇博(北京)文化有限公司 仿生机械恐龙的视觉识别系统实现方法
CN102902355A (zh) * 2012-08-31 2013-01-30 中国科学院自动化研究所 移动设备的空间交互方法
CN102981623A (zh) * 2012-11-30 2013-03-20 深圳先进技术研究院 触发输入指令的方法及系统

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115641615A (zh) * 2022-11-25 2023-01-24 湖南工商大学 复杂背景下闭合手掌感兴趣区域提取方法
CN115641615B (zh) * 2022-11-25 2025-08-22 湖南工商大学 复杂背景下闭合手掌感兴趣区域提取方法

Also Published As

Publication number Publication date
US11120254B2 (en) 2021-09-14
CN108230383A (zh) 2018-06-29
CN108230383B (zh) 2021-03-23
US20190311190A1 (en) 2019-10-10

Similar Documents

Publication Publication Date Title
US11120254B2 (en) Methods and apparatuses for determining hand three-dimensional data
CN107958458B (zh) 图像分割方法、图像分割系统及包括其的设备
CN111178250A (zh) 物体识别定位方法、装置及终端设备
EP3113114A1 (en) Image processing method and device
US20160086017A1 (en) Face pose rectification method and apparatus
CN104317391A (zh) 一种基于立体视觉的三维手掌姿态识别交互方法和系统
JP6822482B2 (ja) 視線推定装置、視線推定方法及びプログラム記録媒体
CN108229301B (zh) 眼睑线检测方法、装置和电子设备
WO2016089529A1 (en) Technologies for learning body part geometry for use in biometric authentication
JP6071002B2 (ja) 信頼度取得装置、信頼度取得方法および信頼度取得プログラム
CN109919971B (zh) 图像处理方法、装置、电子设备及计算机可读存储介质
JP2007164720A (ja) 頭部検出装置、頭部検出方法および頭部検出プログラム
KR20150127381A (ko) 얼굴 특징점 추출 방법 및 이를 수행하는 장치
CN115239888B (zh) 用于重建三维人脸图像的方法、装置、电子设备和介质
EP3198522A1 (en) A face pose rectification method and apparatus
CN111275827A (zh) 基于边缘的增强现实三维跟踪注册方法、装置和电子设备
WO2014112346A1 (ja) 特徴点位置検出装置、特徴点位置検出方法および特徴点位置検出プログラム
WO2018198500A1 (ja) 照合装置、照合方法および照合プログラム
CN112418153A (zh) 图像处理方法、装置、电子设备和计算机存储介质
US10902628B1 (en) Method for estimating user eye orientation using a system-independent learned mapping
CN112818874B (zh) 一种图像处理方法、装置、设备及存储介质
US12230052B1 (en) System for mapping images to a canonical space
EP4660946A1 (en) Processing device and program
CN114529978A (zh) 一种运动趋势识别方法及装置
CN114627503B (zh) 一种人手识别方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18776920

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 11.02.2020)

122 Ep: pct application non-entry in european phase

Ref document number: 18776920

Country of ref document: EP

Kind code of ref document: A1