WO2018157286A1 - 一种识别方法、设备以及可移动平台 - Google Patents

一种识别方法、设备以及可移动平台 Download PDF

Info

Publication number
WO2018157286A1
WO2018157286A1 PCT/CN2017/075193 CN2017075193W WO2018157286A1 WO 2018157286 A1 WO2018157286 A1 WO 2018157286A1 CN 2017075193 W CN2017075193 W CN 2017075193W WO 2018157286 A1 WO2018157286 A1 WO 2018157286A1
Authority
WO
WIPO (PCT)
Prior art keywords
palm
point
indicating
determining
point cloud
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2017/075193
Other languages
English (en)
French (fr)
Inventor
唐克坦
周游
周谷越
郭灼
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SZ DJI Technology Co Ltd
Original Assignee
SZ DJI Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by SZ DJI Technology Co Ltd filed Critical SZ DJI Technology Co Ltd
Priority to PCT/CN2017/075193 priority Critical patent/WO2018157286A1/zh
Priority to CN201780052992.4A priority patent/CN109643372A/zh
Publication of WO2018157286A1 publication Critical patent/WO2018157286A1/zh
Priority to US16/553,680 priority patent/US11250248B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • G06V40/28Recognition of hand or arm movements, e.g. recognition of deaf sign language
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D1/00Control of position, course, altitude or attitude of land, water, air or space vehicles, e.g. using automatic pilots
    • G05D1/20Control system inputs
    • G05D1/22Command input arrangements
    • G05D1/228Command input arrangements located on-board unmanned vehicles
    • G05D1/2285Command input arrangements located on-board unmanned vehicles using voice or gesture commands
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D1/00Control of position, course, altitude or attitude of land, water, air or space vehicles, e.g. using automatic pilots
    • G05D1/20Control system inputs
    • G05D1/24Arrangements for determining position or orientation
    • G05D1/243Means capturing signals occurring naturally from the environment, e.g. ambient optical, acoustic, gravitational or magnetic signals
    • G05D1/2435Extracting 3D information
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • G06T7/246Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/50Depth or shape recovery
    • G06T7/55Depth or shape recovery from multiple images
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D2109/00Types of controlled vehicles
    • G05D2109/20Aircraft, e.g. drones
    • G05D2109/25Rotorcrafts
    • G05D2109/254Flying platforms, e.g. multicopters
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D2111/00Details of signals used for control of position, course, altitude or attitude of land, water, air or space vehicles
    • G05D2111/10Optical signals
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D2111/00Details of signals used for control of position, course, altitude or attitude of land, water, air or space vehicles
    • G05D2111/60Combination of two or more signals
    • G05D2111/63Combination of two or more signals of the same type, e.g. stereovision or optical flow
    • G05D2111/64Combination of two or more signals of the same type, e.g. stereovision or optical flow taken simultaneously from spaced apart sensors, e.g. stereovision
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10028Range image; Depth image; 3D point clouds
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person

Definitions

  • the present invention relates to the field of image processing, and in particular, to an identification method, device, and mobile platform.
  • Gesture recognition is the recognition of a user's gesture, such as a hand shape, or a movement track of a palm.
  • structured light measurement, multi-angle imaging and TOF cameras are mainly used for gesture recognition.
  • the TOF camera is widely used in gesture recognition because of its low cost and ease of miniaturization.
  • the recognition accuracy is not high when using the TOF camera for gesture recognition, especially when the TOF camera is applied to the mobile platform.
  • the embodiment of the invention provides an identification method, a device and a movable platform to improve the accuracy of gesture recognition.
  • the first aspect of the embodiments of the present invention provides an identification method, including:
  • a point cloud indicating the palm is determined from the classified point cloud
  • the gesture is determined based on the point cloud indicating the palm.
  • a second aspect of the embodiments of the present invention provides another method for identifying, including:
  • a gesture is determined based on the set of points.
  • a third aspect of the embodiments of the present invention provides an identification device, including:
  • a TOF camera for acquiring a depth image of a user
  • a processor configured to determine a point cloud corresponding to the depth image, classify the point cloud, determine a point cloud indicating the palm from the classified point cloud, and determine a gesture according to the point cloud indicating the palm.
  • a fourth aspect of the embodiments of the present invention provides an alternative identification device, including:
  • a TOF camera for acquiring a depth image of a user
  • a processor configured to determine, according to the depth information, a set of points indicating a two-dimensional image of the palm, and determine a gesture according to the set of points.
  • a fifth aspect of the embodiments of the present invention provides a mobile platform, including:
  • a processor configured to generate a corresponding control instruction according to the gesture recognized by the identification device, and control the movable platform according to the control instruction.
  • Embodiments of the present invention provide a gesture recognition method, a device, and a movable platform, which can recognize a user's gesture by acquiring a depth image of a user, and can accurately accurately obtain a depth image from a depth image when the resolution of the depth image obtained by the TOF camera is low.
  • the user's palm is extracted, and when the TOF camera captures a low frame rate, the motion trajectory of the user's palm can be accurately extracted, thereby accurately recognizing the user's gesture.
  • a control instruction corresponding to the gesture is generated, and the control platform is used to control the movable platform, thereby simplifying the control flow of the movable platform, enriching the control manner of the movable platform, and further improving The fun of manipulating mobile platforms.
  • FIG. 1 is a schematic diagram of a gesture recognition system according to an embodiment of the present invention.
  • FIG. 2 is a flow chart of a method for identifying an embodiment of the present invention
  • FIG. 3 is a schematic diagram of classifying a point cloud indicating a user according to an embodiment of the present invention
  • FIG. 4 is a schematic diagram of deleting a point set indicating a two-dimensional image of a user's arm from a point set indicating a two-dimensional image of a user's hand in an embodiment of the present invention
  • FIG. 5 is a schematic diagram of determining, according to a point set of a two-dimensional image, a distribution feature of a point set indicating a palm of a user according to an embodiment of the present invention
  • FIG. 6 is a schematic diagram of determining a speed direction corresponding to position information according to position information of a palm according to an embodiment of the present invention
  • FIG. 7 is a schematic diagram of determining a speed direction corresponding to position information according to position information of a palm according to another embodiment of the present invention.
  • FIG. 8 is a schematic diagram of determining a motion direction corresponding to position information according to a speed direction corresponding to position information of a palm according to another embodiment of the present invention.
  • FIG. 9 is a schematic diagram of identifying a tick gesture according to an embodiment of the present invention.
  • FIG. 10 is a flowchart of a method for identifying in another embodiment of the present invention.
  • FIG. 11 is a schematic diagram of an identification device according to an embodiment of the present invention.
  • FIG. 12 is a schematic diagram of a mobile platform in an embodiment of the present invention.
  • FIG. 13 is a schematic diagram of communication between a mobile platform and a control terminal according to an embodiment of the present invention.
  • Embodiments of the present invention provide a gesture recognition method device and an unmanned aerial vehicle, which recognize a user's gesture by acquiring a depth image of the user.
  • a control instruction corresponding to the gesture is generated, and the unmanned aerial vehicle is controlled by the control instruction, which enriches the control mode of the unmanned aerial vehicle and improves the interest of manipulating the unmanned aerial vehicle.
  • the role of TOF camera calibration is to sit with the camera on the coordinates of the 2D image in the depth image.
  • the three-dimensional coordinates of the camera coordinate system corresponding to each two-dimensional image coordinate that is, the three-dimensional point cloud, referred to as a point cloud.
  • the purpose of TOF camera calibration is to ensure that the relative positional relationship between the various parts of the point cloud is consistent with the real world.
  • the imaging principle of the TOF camera is the same as that of a general pinhole camera, except that the receiver of the TOF camera can only receive the modulated infrared light reflected by the target object, and the amplitude image obtained by the TOF camera is the same as the gray image obtained by the general camera.
  • the calibration method can also be used for reference.
  • R is the rotation matrix of the world coordinate system relative to the camera coordinate system
  • T is the translation vector of the world coordinate system relative to the camera coordinate system
  • is the proportional coefficient
  • the black and white checkerboard is used as the calibration pattern.
  • two corresponding points are obtained by corner detection, and one set is the coordinates of each corner point on the checkerboard coordinate system. Measured and recorded before calibration, the other group is the 2D image coordinates of the corresponding point detected by the corner point
  • the two sets of points should conform to formula (1).
  • image noise and measurement error make us only find a least squares solution:
  • the Z value is 0 in the checkerboard coordinate system and is available from equation (1).
  • the homography matrix H can be optimized by using two sets of corresponding points.
  • the optimization method is as follows:
  • a 2i*9 matrix can be written, corresponding to a system of equations consisting of 9 unknowns and 2i equations.
  • the least squares solution is the optimal solution of the objective function (2).
  • n images correspond to a linear equation of 2n equations with 6 unknowns, or you can find the least squares solution to obtain the optimal B, thus solving the camera.
  • Internal reference matrix K Internal reference matrix
  • FIG. 1 is a schematic diagram of a gesture recognition system according to an embodiment of the present invention.
  • a TOF camera-based gesture recognition system is provided in the embodiment, where the system includes a transmitter 101, wherein the transmitter 101 It can be a Light Emitting Diode (LED) or a Laser Diode (LD), where the transmitter is driven by
  • the driving module 102 is driven by the processing module 103.
  • the processing module 103 controls the driving module 102 to output a driving signal to drive the transmitter 101.
  • the frequency, duty ratio, etc. of the driving signal output by the driving module 102 can be processed.
  • the module 103 controls the driver 101 to drive the transmitter 101, and the transmitter 101 sends the modulated optical signal to the target object.
  • the target object can be the user, and the optical signal is sent to the user.
  • the optical signal will be reflected and the receiver 104 will receive the optical signal reflected by the user.
  • the receiver 104 may include a photodiode, an avalanche photodiode, and a charge coupled device.
  • the optical signal reflected by the user includes an optical signal reflected by the user's hand, and the receiver 104 converts the optical signal into an electrical signal.
  • the signal processing module 105 The signal output by the receiver 104 is processed, for example, amplified, filtered, etc., and the signal processed by the signal processing module 105 is input to the processing module 103, and the processing module 103 can convert the signal into a depth image containing the palm of the user. Location information and depth information.
  • the signal recognition module 105 may not be included in the gesture recognition system, and the receiver 104 may directly input an electrical signal to the processing module 103, or the signal processing circuit 105 may be included in the receiver 104 or Processing module 103.
  • the processing module 103 may output the depth image to the discriminating module 106, and the discriminating module 106 may recognize the gesture of the user according to the depth image; in addition, in some cases, the gesture recognition system may not include the discriminating module 106. After the processing module 103 converts the signal into a depth image, the user's gesture can be directly recognized according to the depth image.
  • FIG. 2 is a flowchart of a gesture recognition method according to an embodiment of the present invention, including:
  • S201 Acquire a depth image of the user, and determine a point cloud corresponding to the depth image
  • the user makes a special within the detection range of the TOF camera of the gesture recognition device.
  • the gesture includes a dynamic gesture of the palm, that is, a gesture formed by the user moving the palm, such as moving under the palm, moving the palm left and right, moving the palm back and forth, etc., and the gesture further includes a static gesture of the palm, that is, a user's hand, such as
  • the fist gesture recognition device includes a TOF camera, the light signal emitted by the TOF camera is directed to the user, the TOF camera receives the optical signal reflected by the user, and the TOF camera receives the signal. The incoming optical signal is processed to output a depth image of the user.
  • the TOF camera can calculate the user's point cloud based on the depth image.
  • a capture center can be set, and then the capture center is used as a center of the sphere, and a point cloud is acquired in a spherical space with a preset distance threshold value to eliminate interference.
  • the capture center can be set directly in front of the TOF camera.
  • the capture center can be set in the range of 0.4-2m directly in front of the TOF camera.
  • the capture center can be set 0.8m, 1m, 1.1m directly in front of the TOF camera.
  • the preset distance threshold can be selected by a person skilled in the art according to design requirements, for example, the preset distance threshold can be in the range of 10-70cm. Specifically, 20 cm, 25 cm, 30 cm, 35 cm, 40 cm, 45 cm, 50 cm, 55 cm, and the like can be selected.
  • S202 classify the point cloud, and determine a point cloud indicating a palm from the classified point cloud;
  • the point cloud of the user may include a point cloud of a plurality of parts of the user's body, such as a hand, a head, a torso part, etc.
  • the point cloud indicating the user's hand needs to be extracted first, and the point to the user may be The cloud is classified, and after classification, at least one point cloud cluster is obtained, and a point cloud indicating the palm of the user is determined from the cluster obtained by the classification, and the point cloud of the user's palm is extracted to extract the palm of the user, based on the palm of the user.
  • S203 Determine a gesture according to the point cloud of the palm.
  • the point cloud indicating the palm of the user may indicate location information of the palm of the user, or contour information of the palm, etc.
  • the dynamic gesture of the user may be identified by the location information contained in the point cloud
  • the static gesture of the user may be identified by the contour information of the palm.
  • the point cloud of the user is acquired, the point cloud of the user is classified, the point cloud indicating the palm is determined from the point cloud obtained by the classification, and the gesture of the user is recognized according to the point cloud indicating the palm.
  • the palm of the user when the resolution of the depth image obtained by the TOF camera is low, the palm of the user can be accurately extracted from the depth image, and the user can be accurately extracted when the frame rate of the TOF camera is low. The movement of the palm of the hand, thus accurately identifying the user's gestures, saving computing resources and high recognition rate.
  • the point cloud classification is performed to obtain a plurality of clusters of the point cloud, and the point cloud indicating the palm is determined from one of the plurality of clusters.
  • the distance between the head, the trunk portion, the hand, the foot, and the like of the user's body is different from that of the TOF camera, that is, the trunk portion of the user's body.
  • the depth information of the hand is different from that of the hand.
  • the point clouds of the same part of the user's body are generally close to each other, so that the user can distribute the various parts of the body according to the gesture.
  • the a priori information of different spatial positions classifies the body parts of the user within the detection range of the TOF camera, and obtains clusters of at least one point cloud after classification, and different clusters generally represent different parts of the user's body, and the classification may be The different parts of the body are separated. At this time, only the part belonging to the palm is determined in a specific part of the classification, that is, the clustering of a certain point cloud obtained by the classification to determine the point cloud indicating the palm of the user, so that Reduce the search range of the user's palm to improve the accuracy of recognition.
  • a clustering algorithm is utilized to classify point clouds.
  • the K-means classification (k-means) in the clustering algorithm can be used for classification.
  • K-means clustering is an unsupervised classification algorithm, and the number of clustering classes of the classification must be specified in advance, if it can be determined In the TOF detection range, only the trunk part and the hand of the human body are included, and the classification class can be classified into two categories, but in actual cases, the detection range of the TOF camera may include other objects than the user in addition to the user. Or, in the detection range of the TOF, there is only the user's hand and there is no user's trunk part, so the number of clustering classes is uncertain.
  • the clustering number of the clustering algorithm is adjustable, and the adjustment of the clustering number of the clustering algorithm is performed below.
  • the number of clustering classes may be adjusted according to the degree of dispersion between the clusters, wherein the degree of dispersion may be represented by a distance between clustering centers of the respective clusters.
  • the initial clustering class number is set to n.
  • n can be set to 3
  • n is a parameter that can be adjusted during the clustering operation, and a K-means clustering is performed, and each is obtained. Clustering the center, and then calculating the degree of dispersion of each cluster center. If the distance between two cluster centers is less than or equal to the distance threshold set in the classification algorithm, then n is decreased by 1 and re-aggregated.
  • the threshold value is also an adjustable parameter.
  • the threshold value may be set in the range of 10 to 60 cm, specifically, 10 cm, 15 cm, 20 cm, 25 cm, 30 cm, or the like. If the classification effect of the clustering algorithm is poor, increase n by 1 and re-cluster. When it is found that the distance between all the cluster centers is greater than or equal to the threshold, the clustering algorithm is stopped, and the point cloud indicating the user is classified, and the current cluster number is returned. And cluster centers.
  • a cluster indicating a point cloud of the hand is determined from the plurality of clusters according to the depth information, and a point cloud indicating the palm is determined from the cluster of the point cloud indicating the hand. Specifically, the user's point cloud is classified, and at least one cluster can be obtained.
  • FIG. 3 is a schematic diagram of classifying a point cloud indicating a user according to an embodiment of the present invention. As shown in FIG. 3, after classifying a point cloud of a user, four clusters may be acquired, and four clusters are respectively Cluster 301, cluster 302, cluster 303, cluster 304, each cluster represents a different average depth. According to a priori information, when the user gestures to the TOF camera, the hand is closest to the TOF camera.
  • the depth of the user's hand is the smallest. Therefore, the average depth of each of the clusters obtained after the classification is obtained, for example, the average depth of the cluster 301, the cluster 302, the cluster 303, and the cluster 304 is averaged.
  • the cluster with the smallest depth is determined as a cluster indicating the hand of the user, that is, the cluster 301 is determined as a cluster indicating the hand of the user, so that the point cloud of the user's hand and including the palm is determined from the point cloud of the user, and obtained
  • a point cloud indicating the user's hand can further determine a point cloud indicating the palm of the user from a point cloud indicating the user's hand.
  • the point clouds of the user are classified, and the four clusters are obtained only for the purpose of illustration, and the technical solutions of the embodiments are not limited.
  • a point cloud indicating the arm is deleted from the cluster of the point cloud indicating the hand, and the remaining point cloud in the cluster is determined as a point cloud indicating the palm.
  • the user's hand includes the user's palm and arm
  • the point cloud of the user's hand usually includes a point cloud indicating the arm.
  • it is determined from the point cloud indicating the hand.
  • the point cloud of the arm will delete the point cloud indicating the arm, and the remaining point cloud is determined as the point cloud of the palm, so that the palm is accurately extracted, and the hand is recognized according to the point cloud of the palm later. Potential.
  • the method of deleting the point cloud of the arm from the point cloud indicating the hand will be described in detail below.
  • the point with the least depth is taken in the cluster indicating the point cloud of the hand, the distance between the point cloud in the cluster and the minimum point of the depth is determined, and the point at which the distance is greater than or equal to the distance threshold is determined as the arm.
  • Point cloud Specifically, in the above-mentioned clustering of the point cloud indicating the user's hand, the arm in the hand is usually included, and the point cloud of the arm in the hand needs to be deleted before the specific gesture recognition is performed. First, a depth histogram indicating the cluster of the point cloud of the user's hand is calculated. The histogram can be used to extract the point with the smallest depth.
  • the point with the smallest depth is usually the fingertip of the finger, and the other points in the cluster are calculated to the minimum depth.
  • the distance of the points determines all the points whose distance exceeds the distance threshold as points indicating the arm, deletes them, and retains the remaining points, that is, the point whose distance is less than or equal to the distance threshold is determined as a point cloud indicating the palm of the user.
  • the distance threshold may be changed according to requirements, or determined according to the average size of the palm, such as 10 cm, 13 cm, 15 cm, 17 cm, and the like.
  • a set of points indicating a two-dimensional image of the hand is determined according to a cluster of point clouds indicating the hand, a minimum circumscribed rectangle of the set of points of the two-dimensional image of the hand is determined, the hand is determined The distance from the point in the point set of the two-dimensional image to the specified side of the minimum circumscribed rectangle, and the point at which the distance does not meet the preset distance requirement is determined as the point indicating the two-dimensional image of the arm, and the two-dimensional indication arm is determined.
  • a set of points of the image, the point cloud indicating the arm is determined according to the set of points indicating the two-dimensional image of the arm, and the point cloud indicating the arm is deleted.
  • a frame depth image is obtained, and a point cloud of the user in the depth image of the frame is determined.
  • a point cloud indicating the user's hand of the frame depth image may be determined. Since the three-dimensional coordinates of each point cloud are in one-to-one correspondence with the two-dimensional coordinates of the points on the two-dimensional image, and the two coordinates are always stored in the process of gesture recognition, the point cloud indicating the user's hand is acquired. Rear, A set of points of a two-dimensional image of the user's hand can be determined.
  • FIG. 4 is a schematic diagram of deleting a point set indicating a two-dimensional image of a user's arm from a point set indicating a two-dimensional image of a user's hand in an embodiment of the present invention, as shown in FIG. 4, acquiring the user
  • the minimum circumscribed rectangle 401 of the point set of the two-dimensional image of the hand determines the distance from the point set of the two-dimensional image of the user's hand to the designated side of the circumscribed rectangle, and determines the point that does not meet the preset distance requirement as an indication.
  • the point of the arm that is, the point to go from the point that does not meet the preset distance is deleted as the pointing arm point, and the remaining point set is the point set indicating the two-dimensional image of the palm, and the point set according to the two-dimensional image of the palm is You can get the point cloud of the palm.
  • the preset distance requirement can be determined by the side length of the minimum circumscribed rectangle 401. Specifically, the length of the long side of the rectangle is taken as w, and the length of the short side of the rectangle is h.
  • the specified side can be the lower short side, and each point of the point set indicating the two-dimensional image in the hand is calculated to the lower side.
  • the distance d i of the short side if d i ⁇ w-1.2h, the point is determined as the point indicating the arm, in this way, the point set of all the two-dimensional images indicating the arm can be deleted, and the remaining points
  • the set is based on a set of points indicating a two-dimensional image of the palm, that is, the point cloud indicating the arm can be deleted.
  • h is the width of the palm on the two-dimensional image.
  • 1.2h is determined as the length of the palm, and the difference between w and 1.2h should be The maximum distance from the point on the arm to the short side below.
  • the distance from a point in the point set of the two-dimensional image to the short side below is less than or equal to the maximum distance, it means that the point is the point belonging to the arm.
  • d ⁇ w-1.2h is only one embodiment for determining the preset distance requirement according to the side length of the minimum circumscribed rectangle, and other methods may be selected by those skilled in the art, for example, the length of the palm may be 1.1h. 1.15h, 1.25h, 1.3h, 1.35h, 1.4h, etc., and no specific limitation is made here.
  • the point set of the two-dimensional image of the palm is acquired according to the point cloud indicating the palm of the user, and the gesture is determined according to the distribution feature of the point set.
  • the user's static gestures are identified here, such as clenching a fist, stretching a palm, extending a finger, extending two fingers, and the like.
  • the point set of the two-dimensional image of the user's palm can be determined. Due to different user gestures, that is, different gestures correspond to different hand types, the distribution characteristics of the point set of the two-dimensional image of the palm will be different. For example, the distribution characteristics of the fist gesture and the distribution characteristics of the palm gesture are very different, so it can be determined.
  • the distribution feature of the point set of the two-dimensional image of the palm specifically determining the gesture made by the user in the image of the frame according to the distribution feature.
  • a distribution area of a point set indicating a two-dimensional image of the palm is determined, and a distribution characteristic of the point set is determined according to the distribution area.
  • the distribution area of the point set may be determined according to the point set indicating the two-dimensional image of the palm.
  • FIG. 5 is a distribution feature of the point set of the user's palm according to the point set of the two-dimensional image according to an embodiment of the present invention. Schematic diagram, as shown in FIG. 5, in some embodiments, a method of creating an image mask may be used to determine a distribution area of a point set, that is, an area 501 in FIG. 5, wherein the distribution area 501 is a user.
  • the area occupied by the palm of the hand on the two-dimensional image, the shape and contour of the distribution area of the different gestures are different, and the distribution feature of the point set indicating the palm can be determined by the shape and contour of the distribution area 501, and the user can be identified according to the distribution feature Gesture.
  • the polygon area is used to cover the distribution area, and the polygon area is determined.
  • a non-overlapping region between the domain and the distribution region determines a distribution feature of the set of points according to the non-overlapping region.
  • all point pixel values of the point set of the two-dimensional image of the palm may be set to 1, and other point pixels in the two-dimensional image.
  • the value is set to 0, and the distribution area is covered with a polygon, that is, all points in the point set are covered with a polygon, wherein the polygon is a convex polygon with the fewest number of sides. As shown in FIG.
  • a binocular operation may be performed on a point set of the binarized two-dimensional image indicating the palm, and the point set may be covered by the convex polygon 502 having the smallest number of sides.
  • Each vertex of the convex polygon is a point in the point set, such that the distribution area 501 of the point set of the two-dimensional image and the polygon 502 have a non-overlapping area 503, and the shape and size of the non-overlapping area 503 can be expressed.
  • the distribution feature of the point set can identify the user's gesture according to certain features of the non-overlapping area.
  • the distribution feature of the point set may be determined according to the non-overlapping area.
  • the edge of the polygon corresponding to the non-overlapping area 503 is l i
  • the non-overlapping area 503 A point is determined to be the farthest distance from the edge l i
  • the farthest distance d i is used as a distribution feature of the point set
  • the user's gesture can be identified according to the distance d i .
  • said d i may be a maximum distance, or may be a combination of a plurality of the maximum distance between the one skilled in the art can select on demand, only FIG. 5 schematic illustration.
  • the gesture is determined to be a palm.
  • the non-overlapping area between the distribution area formed by the point set indicating the two-dimensional image of the palm and the polygon is larger, and the specific expression is that the edge of the polygon surrounding the palm is away from the joint distance between the fingers.
  • Threshold when the farth distance corresponding to each edge of the polygon is less than or equal to the preset distance threshold, determining that the gesture is a fist, at least one of each corresponding maximum distance of the polygon is greater than or equal to a preset distance threshold. When you are sure, the gesture is to stretch your palm. Additionally, the second threshold can be selected based on the length of the finger.
  • obtaining a depth image of the multi-frame user determining a point cloud corresponding to the user's palm corresponding to each frame depth image in the multi-frame depth image; and classifying the point cloud corresponding to each frame depth image from the classification Determining a point cloud of the palm corresponding to each frame depth image in the point cloud; determining a position information of the palm corresponding to each frame depth image according to the point cloud corresponding to the user's palm corresponding to each frame image, according to the position
  • the sequence of information consists of the gesture of the palm.
  • the gesture of the user may be determined by using a multi-frame depth image, where the gesture refers to a gesture formed by the user by moving the palm.
  • the palm of each frame of the depth image is first extracted.
  • the point cloud of the user's palm corresponding to each frame image can be obtained according to each frame depth image, according to each
  • the point cloud of the palm of the user corresponding to the frame image can calculate the position information of the palm, wherein the position of the geometric center of the point cloud indicating the palm can be used as the position information of the palm, and the depth information in the point cloud indicating the palm can be minimized.
  • the position of the point serves as the position information of the palm.
  • the position information of the palm is determined by a person in the art in a different manner according to the point cloud indicating the palm of the user, and is not specifically limited herein.
  • the position information of the palm calculated from the multi-frame depth image may be stored in the sequence P, wherein the length of the sequence P is L, and the first-in first-out storage mode is used, and the position information of the recently acquired palm is used to replace The location information to the oldest palm.
  • This sequence P reflects the trajectory of the palm movement in a fixed time, which represents the gesture of the user, so that the user's gesture can be recognized based on the sequence P, that is, the sequence of position information of the palm.
  • the position point indicated by the position information may be used as a capture center, and when determining the position information of the palm corresponding to the next frame depth image, the capture center may be The center of the sphere acquires the user's point cloud in a spherical space with a preset distance threshold value, that is, the user's hand is extracted only in the spherical space, so that the recognition speed of the hand can be improved.
  • the Kalman filter algorithm can be used to estimate the motion model of the palm, predicting the position of the palm indicated by the depth image of the next frame, and extracting the palm of the user near the position of the predicted palm.
  • the filtering algorithm can be turned on or off at any time.
  • the moving direction of the palm motion corresponding to the position information in the sequence is determined, and the gesture is determined according to the sequence of the moving direction composition.
  • the motion direction corresponding to the position information can be calculated, wherein each corresponding motion side of the L pieces of position information can be determined.
  • the motion direction corresponding to each of the plurality of position information in the L pieces of position information may be determined, and the obtained motion direction sequence composed of the plurality of motion directions may represent a motion trajectory of the palm in the air and a motion change, and therefore, according to the motion
  • the sequence of directions can determine the gesture of the user.
  • the position information corresponding to the movement direction in the sequence P may be the speed direction corresponding to the position information, or may be a direction determined in some manner according to the speed direction.
  • determining a proportion of each of the moving directions, and determining the motion of the palm motion according to the combination of the ratios are determining a proportion of each of the moving directions, and determining the motion of the palm motion according to the combination of the ratios. Specifically, the proportion of each of the moving directions in the sequence of the moving direction is counted, so that a proportional sequence of the proportional components can be obtained, and the sequence of the proportionals is used to identify the gesture of the user. In this way, when the user gestures, no matter where the user gestures, where the starting and ending points of the palm movement are, a sequence of proportions of the same form can be obtained, which is convenient for processing.
  • the gesture recognition is performed, the sequence of the scale is input into a preset operation model, and the preset operation model identifies the gesture of the user according to the sequence of the scale.
  • the preset operation model may be a neural network, a classifier, or the like. Before performing gesture recognition, the preset operation model needs to be trained, that is, a sequence of proportions corresponding to a large number of gestures needs to be collected offline. The sequence of the proportional ratio is used as an input, and the gesture corresponding to the proportional sequence is used as an output to train the preset operation model. After the training is completed, the preset operation model can be used for gesture recognition.
  • FIG. 6 is a schematic diagram of determining a speed direction corresponding to position information according to position information of a palm according to an embodiment of the present invention.
  • P i represents a position point of a palm indicated by a frame depth image. That is, the position point indicated by the position information.
  • L is taken as 7, and the speed direction is determined using a sequence of position information of the palm, that is, according to the position point sequence of the palm, specifically, the position of the palm P
  • the speed direction of 2 is from the position point P 1 to the position point P 2
  • the speed direction of the position P 3 of the palm is from the position point P 2 to the position point P 3 , and so on, and a sequence of speed directions (V 1 , can be obtained.
  • V 2 ⁇ V 6 the sequence of the velocity direction may indicate a change in the direction of motion of the palm, and the sequence of the direction of motion may be determined according to the sequence of the velocity direction.
  • FIG. 7 is a schematic diagram of determining the speed direction corresponding to the position information according to the position information of the palm according to another embodiment of the present invention, as shown in FIG. 7 . It is shown that the moving direction of the position P 3 of the palm is from the position point P 1 to the position point P 3 , the moving direction of the position P 4 of the palm is from the position point P 2 to the position point P 4 , and so on, the speed direction can be obtained.
  • Sequence (V 1 , V 2 ⁇ V 5 ). It should be noted that after obtaining the sequence of the speed direction, the sequence input filter in the speed direction may be filtered. Specifically, the sequence in the speed direction may be input into the Kalman filter, so that the sequence in the speed direction may be Noise or abnormal speed direction filtering.
  • the speed direction corresponding to the position information in the sequence is determined, an angle between the speed direction and each of the plurality of preset directions is determined, and the moving direction is determined according to the included angle.
  • this paper only schematically illustrates how one speed direction in the velocity direction sequence determines the corresponding motion direction, and the other speed directions in the speed direction sequence determine the corresponding motion direction.
  • FIG. 8 is a schematic diagram of determining a motion direction corresponding to position information according to a speed direction corresponding to position information of a palm according to another embodiment of the present invention.
  • setting a plurality of preset directions for example, V u , V p , V l , V r , V f , V d respectively represent six preset directions of up, down, left, right, front and back, and the velocity direction V i corresponding to the position point is calculated according to the direction of the foregoing part
  • the unit vector corresponding to the velocity direction is respectively multiplied by a corresponding single-phase vector in each of the six preset directions, and ⁇ 1 to ⁇ 6 can be calculated, and the motion direction of the position information can be determined according to ⁇ 1 to ⁇ 6 .
  • the ⁇ i with the smallest angle may be determined from ⁇ 1 to ⁇ 6
  • the first preset direction corresponding to ⁇ i (for example, V r shown in FIG. 7 ) may be determined as the motion direction corresponding to the position information.
  • the six preset directions of up, down, left, right, front and back are set only for illustrative explanation. Those skilled in the art can set more preset directions when hardware conditions permit, so that the speed direction can be made. The classification is more precise, so that the motion direction error corresponding to the position information is smaller.
  • the selection of the preset direction number can be selected by those skilled in the art according to design requirements and/or hardware conditions, and is not specifically limited herein.
  • the rate corresponding to the location information is determined according to the sequence of the location information, and when the rate is less than the preset rate threshold, determining that the palm is in a stationary state when the location point indicated by the location information is used.
  • the rate corresponding to the position information in the sequence P may be determined according to the position information sequence P, wherein the rate may be calculated according to the displacement of the palm, wherein The displacement can be calculated according to the position information in the position sequence. Since the interval time between two adjacent position information in the position information sequence P is the same, the time information can be directly used to represent the rate corresponding to the position information without introducing time information. for example, P 2 P.
  • P 1 corresponding to the rate of displacement of the point P 2, wherein the displacement may be P 1
  • P 2 acquiring the position information obtained according to the same token can obtain P 3, P 4, P 5 , P 6, P 7
  • Corresponding rate when the rate is less than or equal to the rate threshold, it is considered that the palm is at rest at this time, and there is no direction of motion.
  • the location information corresponding to the rate may be calculated using other ways, for example as shown in FIG. 7, P is P 3 corresponding to the rate of displacement. 1 to point P 3, where not specifically limited.
  • a corresponding two-dimensional pattern coordinate sequence may be acquired according to the position sequence, that is, a point on the two-dimensional image is acquired.
  • the area enclosed on the two-dimensional image can be calculated.
  • the preset area threshold it is determined that the current gesture of the user is not a circular gesture. Using the area to determine, to some extent, eliminates the misidentification that may exist when different gestures are switched.
  • the preset area threshold may be selected by a person skilled in the art according to design requirements, for example, 40, 50, 60, 70, and the like.
  • the user's tick gesture is identified in accordance with the location information sequence of the aforementioned portion. Specifically, acquiring a projection sequence of the sequence indicating the position of the palm on the XY plane, traversing the points in the projection sequence, and determining from the points of the sequence that the preset is satisfied At the specific point sought, it is determined that the tick gesture is recognized.
  • the distance of the palm from the TOF camera is approximately the same, that is, the value of the Z direction is substantially unchanged in the three-dimensional space, and when determining the tick gesture, the position information may be disregarded regardless of the Z coordinate.
  • the sequence is projected onto the XY plane, wherein, according to the prior information, when the user performs the tick gesture, the motion track of the palm has a lowest point on the XY plane, and the motion trajectory of the gestures on both sides of the lowest point is approximately a straight line. And the slopes of the two approximate straight lines are opposite to each other, and the lowest point can be determined as a specific point that satisfies a preset requirement, and the specific point and the first motion trajectory formed by the point in the sequence before the specific point are determined.
  • the second motion trajectory formed by the specific point and the point after the specific point in the sequence is determined to be an approximate straight line; and the slope of the first motion trajectory is opposite to the positive and negative of the second trajectory.
  • a point in the projection sequence is obtained, the point is taken as the current point, the point in the sequence before the current point is obtained, and the current point and the point in the sequence before the current point are performed.
  • Straight line fitting obtaining a first correlation coefficient and a first slope, and determining a first motion trajectory as an approximate straight line if the correlation coefficient is greater than or equal to a correlation coefficient threshold, acquiring a point in the sequence after the current point, and the current point Performing a straight line fitting with a point in the sequence after the current point to obtain a second correlation coefficient and a second slope, and if the correlation coefficient is greater than or equal to the correlation coefficient threshold, determining that the first motion trajectory is an approximate straight line, and the second motion trajectory is approximated a straight line, and the first slope and the second slope are opposite to each other, determining that the current point is a specific point that satisfies a preset requirement, if one or both of the first correlation coefficient and the second correlation coefficient are less than or equal to the correlation coefficient
  • the threshold
  • FIG. 9 is a schematic diagram of identifying a tick gesture according to an embodiment of the present invention.
  • P 4 if P 4 is the current point, it traverses from P 4 and obtains a point before P 4 .
  • l can be related to the trajectory 2 Coefficient and slope k 2 , when the correlation coefficients of the track l 1 and the track l 2 are both greater than or equal to the correlation coefficient threshold, it is judged that the track l 1 and the track l 2 are straight lines, otherwise the next point is taken as the current point, and the above operation is repeated.
  • the slopes k 1 and k 2 of the trajectory l 1 and the trajectory l 2 are obtained. If the two slopes are opposite to each other, the current user's gesture can be judged as a ticking gesture. Thus, by traversing the point in the projection sequence until a specific point is acquired, it can be determined that the current gesture is a tick gesture.
  • the correlation coefficient can be selected by a person skilled in the art according to requirements, for example, it can be selected as 0.8.
  • the traversal from the current point to the point before the current point, the displacement sum of the point before the current point and the traversed current point is acquired, and when the displacement sum is greater than or equal to the preset displacement threshold, the traversal is stopped, Straight line fitting the current point with the point before the current point traversed; traversing from the current point to the point after the current point, obtaining the displacement sum of the point after the current point and the traversed current point, when the displacement sum is greater than or equal to When the preset displacement threshold is used, the traversal is stopped, and the current point is straight-line fitted to the point after the current point traversed. Specifically, as shown in FIG.
  • traversing the point after the current point is also performing the above operation, and for brevity, it will not be described again. If the displacement of the point before the current point and all the traversed current points is less than D, it means that the current end of the projection sequence is too close, which will cause the number of points before the current point to be insufficient or the amount of information before the current point is insufficient. The next point is the current point. Similarly, if the displacement of the point after the current point and all the traversed current points is less than D, the next point is taken as the current point.
  • Embodiments of the present invention provide a computer storage medium having stored therein program instructions, the computer storage medium storing program instructions, the program executing the identification method.
  • an embodiment of the present invention further provides an identification method, including:
  • S1001 Acquire a depth image of the user, and determine a point set indicating the two-dimensional image of the palm according to the depth information;
  • the user makes a gesture to the gesture recognition device within the detection range of the TOF camera of the gesture recognition device, wherein the gesture includes a dynamic gesture of the palm, that is, a gesture formed by the user moving the palm, such as moving under the palm, moving the palm left and right, The palm moves back and forth, etc., and the gesture further includes a static gesture of the palm, a user's hand, such as a fist, a palm, etc.
  • the gesture recognition device includes a TOF camera, the optical signal emitted by the TOF camera is directed to the user, and the TOF camera receives the user reflection. The optical signal, the TOF camera processes the received optical signal and outputs the user's depth image.
  • the distance separating the TOF cameras of the various parts of the body is different, that is, the depth is different, so the point set of the two-dimensional image of the user's palm can be determined according to the depth information. , that is, to obtain the image coordinates of all points of the palm on the two-dimensional image
  • S1002 Determine a gesture according to the set of points.
  • the gesture of the user can be identified based on the set of points.
  • a point set indicating a two-dimensional image of a palm is determined according to a depth image of the user, and a gesture of the user is recognized according to the point set.
  • the palm of the user when the resolution of the depth image obtained by the TOF camera is low, the palm of the user can be accurately extracted from the depth image, and the user can be accurately extracted when the frame rate of the TOF camera is low. The movement of the palm of the hand, thus accurately identifying the user's gestures, saving computing resources and high recognition rate.
  • determining a point indicating the palm on the two-dimensional image according to the depth information determining a point set connected to the point indicating the palm according to the preset depth range, and determining, according to the connected point set, the palm indicating the palm
  • the set of points of the dimensional image when the user gestures to the TOF camera, the distance of the palm from the TOF camera is the closest, and the depth of the palm point is the smallest, and the point with the smallest depth can be taken out as the point indicating the palm. This point is usually a fingertip. Alternatively, the three points with the lowest depth can be taken out, and the geometric center of the three points is determined as the point indicating the palm.
  • All points in communication with the point indicating the palm are taken out within a preset depth range, wherein all points of the connected connection can be achieved by a flood fill algorithm, and the preset depth range can be determined by those skilled in the art. It is selected according to actual needs (for example, the preset depth range can be selected as (0, 40 cm)), and no specific limitation is made here.
  • the point set indicating the arm is deleted from the connected point set, and the remaining point set is determined as a point set indicating the palm.
  • the set of points in the connected point usually includes a set of points indicating the arm, and the set of points indicating the arm should be deleted.
  • the minimum circumscribed rectangle 901 of the connected point set is obtained, and the distance between the connected point set point and the specified side of the circumscribed rectangle is determined, when the distance does not meet the preset
  • the point required for the distance is determined as the point indicating the arm, that is, the point at which the distance does not meet the preset distance is deleted as the pointing arm point, and the remaining point set is the point set indicating the palm.
  • the predetermined distance requirement may be determined by the side length of the minimum circumscribed rectangle. Specifically, the length of the long side of the rectangle is taken as w, and the length of the short side of the rectangle is h, and the distance from the point of the point of the two-dimensional image of the palm to the short side of the lower side is calculated. If d ⁇ w-1.2h, This point is determined as the point indicating the arm, in which way all points indicating the arm can be deleted, the remaining points being based on a set of points indicating the two-dimensional image of the palm.
  • h is the width of the palm on the two-dimensional image.
  • 1.2h is determined as the length of the palm, and the difference between w and 1.2h should be The maximum distance from the point on the arm to the short side below. If the distance from a point in the point set of the two-dimensional image to the short side below is less than or equal to the maximum distance, it means that the point is the point belonging to the arm.
  • d ⁇ w-1.2h is only one embodiment for determining the preset distance requirement according to the side length of the minimum circumscribed rectangle. Other methods may be selected by those skilled in the art, and are not specifically limited herein.
  • the gesture is determined according to a distribution feature of the set of points indicating the two-dimensional image of the palm.
  • the distribution area is determined according to a set of points indicating a two-dimensional image of the palm, and the distribution characteristics of the point set are determined according to the distribution area.
  • the distribution area is covered by using a polygon area, and a non-overlapping area between the polygon area and the distribution area is determined, and a distribution feature of the distribution area is determined according to the non-overlapping area.
  • the gesture is a fist, at least one of each corresponding maximum distance of the polygon is greater than or equal to a preset.
  • the gesture is determined to be the palm of the hand.
  • determining a point set corresponding to the user's palm two-dimensional image corresponding to each frame depth image in the multi-frame depth image and determining, according to the point set of the user's palm corresponding to each frame image, determining each frame depth image corresponding to each frame image
  • the point cloud indicating the palm determines the position information of the palm according to the point cloud indicating the palm, and determines the dynamic gesture of the palm according to the sequence composed of the position information. Specifically, since the three-dimensional coordinates of each point cloud have a one-to-one correspondence of the points on the two-dimensional image, and the two coordinates are always stored in the process of gesture recognition, the two indicating the palm of the hand are determined. After the point set of the dimension image, the point cloud indicating the palm can be determined. After the point cloud indicating the palm of the user is acquired, the gesture of the user can be identified according to the method of the foregoing section.
  • the moving direction of the palm motion corresponding to the position information in the sequence is determined, and the gesture is determined according to the sequence of the moving direction composition.
  • a sequence of motion directions is determined according to the sequence of the velocity directions.
  • determining a speed direction corresponding to the position information in the sequence determining an angle between the speed direction and each of the plurality of preset directions, and determining the moving direction according to the included angle.
  • a first preset direction that is the smallest angle with the speed direction is determined from the preset direction, and the first preset direction is determined as a motion direction corresponding to the speed direction.
  • determining the location information of the palm according to the point cloud indicating the palm, and determining the dynamic gesture of the palm according to the sequence of the location information including: determining location information of the palm according to the point cloud indicating the palm, and composing according to the location information
  • the sequence determines the tick gesture of the palm.
  • a point set indicating a two-dimensional image of the palm is acquired, and all interpretations of determining a gesture (static gesture) of the user according to the point set indicating the two-dimensional image of the palm may refer to FIG. 2
  • a point cloud indicating the palm can be acquired, and all the explanations of determining the user's gesture (dynamic gesture) according to the point cloud indicating the palm can refer to the relevant part in FIG. 2, in order to Concise, no more details here.
  • Embodiments of the present invention provide a computer storage medium having stored therein program instructions, the computer storage medium storing program instructions, the program executing the identification method.
  • an embodiment of the present invention provides a gesture recognition device, where the device 1110 includes:
  • a TOF camera 1110 configured to acquire a depth image of a user
  • the processor 1120 is configured to determine a point cloud corresponding to the depth image, classify the point cloud, determine a point cloud indicating the palm from the classified point cloud, and determine a gesture according to the point cloud indicating the palm.
  • the processor 1120 is specifically configured to obtain a plurality of clusters of the point cloud by using the point cloud classification, and determine a point cloud indicating the palm from one of the plurality of clusters.
  • the processor 1120 is specifically configured to determine a cluster of a point cloud indicating a hand from the plurality of clusters according to the depth information, and determine a palm indicating the palm from the cluster of the point cloud indicating the hand.
  • Point cloud is specifically configured to determine a cluster of a point cloud indicating a hand from the plurality of clusters according to the depth information, and determine a palm indicating the palm from the cluster of the point cloud indicating the hand. Point cloud.
  • the processor 1120 is specifically configured to acquire an average depth of each of the plurality of clusters, and determine a cluster with a minimum average depth as a cluster indicating a point cloud of the hand.
  • the processor 1120 is specifically configured to delete a point cloud indicating an arm from a cluster of point clouds indicating the hand, and determine a point cloud remaining in the cluster as a point cloud indicating a palm .
  • the processor 1120 is specifically configured to: extract a point with a minimum depth in a cluster of a point cloud indicating a hand, determine a distance between a point cloud in the cluster and a minimum point of the depth, and set the distance to be greater than or equal to the distance.
  • the point of the threshold is determined to be a point cloud indicating the arm, and the point cloud indicating the arm is deleted.
  • the processor 1120 is specifically configured to determine a point set of the two-dimensional image indicating the hand according to the cluster of the point cloud indicating the hand, and determine a minimum external point set of the two-dimensional image of the hand. rectangle;
  • the processor 1120 is specifically configured to determine a point concentration of the two-dimensional image of the hand The point of the point to the specified edge of the minimum circumscribed rectangle, and the point at which the distance does not meet the preset distance requirement is determined as a point indicating the two-dimensional image of the arm;
  • the processor 1120 is specifically configured to determine a point set indicating a two-dimensional image of the arm, determine a point cloud indicating the arm according to the point set indicating the two-dimensional image of the arm, and delete the point cloud indicating the arm.
  • the preset distance requirement may be determined by a side length of the minimum circumscribed rectangle.
  • the processor 1120 is specifically configured to acquire a point set of the two-dimensional image of the palm according to the point cloud indicating the palm of the user, determine a distribution feature of the point set, and determine a gesture according to the distribution feature.
  • the processor 1120 is specifically configured to determine a distribution area of a point set indicating a two-dimensional image of the palm, and determine a distribution feature of the point set according to the distribution area.
  • the processor 1120 is specifically configured to cover the distribution area by using a polygon area, determine a non-overlapping area between the polygon area and the distribution area, and determine a point set according to the non-overlapping area. Distribution characteristics.
  • the processor 1120 is specifically configured to cover the distribution area by using a convex polygon area with a minimum number of sides.
  • the processor 1120 is specifically configured to determine a farthest distance from a point in the non-overlapping region to an edge of the corresponding polygon, and determine the farthest distance as a distribution feature of the point set.
  • the processor 1120 is configured to determine that the gesture is a fist when the maximum distance corresponding to each non-overlapping region is less than or equal to a distance threshold.
  • the processor 1120 is configured to determine that the gesture is a palm when one or more of the farthest distances corresponding to the non-overlapping regions are greater than or equal to a distance threshold.
  • the processor 1120 is specifically configured to classify a point cloud by using a clustering algorithm.
  • the clustering class of the clustering algorithm is adjustable.
  • the processor 1120 is specifically configured to adjust the clustering class according to the degree of dispersion between the clusters.
  • the TOF camera 1110 is configured to acquire a depth image of a multi-frame user.
  • the processor 1120 is specifically configured to determine a point cloud corresponding to each frame in the multi-frame depth image
  • the processor 1120 is specifically configured to classify a point cloud corresponding to each frame depth image, and determine, from the classified point cloud, a point cloud indicating a palm corresponding to each frame depth image;
  • the processor 1120 is specifically configured to determine location information of a palm corresponding to each frame depth image according to a point cloud corresponding to the image of the user corresponding to each frame image, and determine a gesture of the palm according to the sequence of the location information.
  • the processor 1120 is specifically configured to determine a motion direction of the palm corresponding to the location information in the sequence according to a sequence indicating location information of the palm, and determine a gesture according to the sequence formed by the motion direction.
  • the processor 1120 is specifically configured to determine a proportion of each of the motion directions in the sequence of the motion direction, and determine the gesture according to the combination of the ratios.
  • the processor 1120 is specifically configured to input a combination of the ratios into a preset operation model, and use the preset operation model to determine a gesture.
  • the processor 1120 is specifically configured to determine a speed direction corresponding to the position information according to the sequence of the position information indicating the palm, and determine a moving direction of the palm according to the speed direction;
  • the processor 1120 is specifically configured to determine a speed direction corresponding to the position information in the sequence, determine an angle between the speed direction and each of the plurality of preset directions, and determine, according to the angle The direction of motion.
  • the processor 1120 is configured to determine, from a preset direction, a first preset direction that is the smallest angle with the speed direction, and determine the first preset direction to correspond to the speed direction.
  • the direction of movement is configured to determine, from a preset direction, a first preset direction that is the smallest angle with the speed direction, and determine the first preset direction to correspond to the speed direction. The direction of movement.
  • the identification device described in the embodiment of the present invention may perform the identification method provided in the embodiment of the present invention.
  • the specific explanation may be referred to the corresponding part in the identification method provided in FIG. 2, and details are not described herein again.
  • the identification device described in the embodiment of the present invention can refer to and combine the technical features in the identification method provided in FIG. 2 of the embodiment of the present invention.
  • an embodiment of the present invention provides another gesture recognition device, where the device 1110 includes:
  • a TOF camera 1110 configured to acquire a depth image of a user
  • the processor 1120 is configured to determine, according to the depth information, a set of points indicating a two-dimensional image of the palm, and determine a gesture according to the set of points.
  • the processor 1120 is specifically configured to determine, according to the depth information, a point indicating a palm on the two-dimensional image, and determine a point set that is connected to the point indicating the palm according to the preset depth range, The connected points collectively determine a set of points that indicate a two-dimensional image of the palm.
  • the processor 1120 is specifically configured to delete the finger from the connected point set.
  • a set of points of the arm is displayed, and the remaining set of points is determined as a set of points indicating the palm.
  • the processor 1120 is configured to obtain a minimum circumscribed rectangle of the connected point set, determine a distance from a point in the connected point set to a specified edge of the circumscribed rectangle, and set a distance that does not meet a preset distance requirement. Determined as the point indicating the arm, the point indicating the arm is deleted, and the remaining point set is determined as the point set indicating the palm.
  • the preset distance requirement may be determined by a side length of the minimum circumscribed rectangle.
  • the processor 1120 is specifically configured to determine a distribution feature of the point set, and determine a gesture according to the distribution feature.
  • the processor 1120 is specifically configured to determine a distribution area of a point set indicating a two-dimensional image of the palm, and determine a distribution feature of the point set according to the distribution area.
  • the processor 1120 is specifically configured to cover the distribution area by using a polygon area, determine a non-overlapping area between the polygon area and the distribution area, and determine a point set according to the non-overlapping area. Distribution characteristics.
  • the processor 1120 is specifically configured to cover the distribution area by using a convex polygon area with a minimum number of sides.
  • the processor 1120 is specifically configured to determine a farthest distance from a point in the non-overlapping region to an edge of the corresponding polygon, and determine the farthest distance as a distribution feature of the point set.
  • the processor 1120 is configured to determine that the static gesture is a fist when the distance corresponding to each non-overlapping region is less than or equal to the distance threshold.
  • the processor 1120 is specifically configured to: when the non-overlapping area corresponds to the distance When one or more of the departures is greater than or equal to the distance threshold, it is determined that the static gesture is a palm.
  • the TOF camera 1110 is configured to acquire a depth image of a multi-frame user
  • the processor 1120 is specifically configured to determine a point set of a two-dimensional image indicating a palm corresponding to each frame in the multi-frame depth image; and determine, according to a point set corresponding to the two-dimensional image of the palm corresponding to each frame.
  • the point cloud corresponding to the frame corresponding to the frame determines the position information of the palm according to the point cloud indicating the palm, and determines the gesture of the palm according to the sequence composed of the position information.
  • the processor 1120 is specifically configured to determine a motion direction of the palm corresponding to the location information in the sequence according to a sequence indicating location information of the palm, and determine a gesture according to the sequence formed by the motion direction.
  • the processor 1120 is specifically configured to determine a proportion of each of the motion directions in the sequence of the motion direction, and determine the gesture according to the combination of the ratios.
  • the processor 1120 is specifically configured to input a combination of the ratios into a preset operation model, and use the preset operation model to determine a gesture.
  • the processor 1120 is specifically configured to determine a speed direction corresponding to the position information according to the sequence of the position information indicating the palm, and determine a moving direction of the palm according to the speed direction;
  • the processor 1120 is specifically configured to determine a speed direction corresponding to the position information in the sequence, determine an angle between the speed direction and each of the plurality of preset directions, and determine, according to the angle The direction of motion.
  • the processor 1120 is configured to determine, from a preset direction, a first preset direction that is the smallest angle with the speed direction, and determine the first preset direction as The direction of motion corresponding to the speed direction.
  • the identification device described in the embodiment of the present invention may perform the identification method provided in the embodiment of the present invention.
  • the specific explanation may be referred to the corresponding part in the identification method provided in FIG. 10, and details are not described herein again.
  • the identification device described in the embodiment of the present invention can refer to and combine the technical features in the identification method provided in FIG. 10 of the embodiment of the present invention.
  • an embodiment of the present invention further provides a mobile platform, where the mobile platform 1200 includes:
  • the identification device 1100 as described above is configured to recognize a gesture of the user
  • the processor 1210 generates a corresponding control instruction according to the gesture recognized by the identification device 1100, and controls the movable platform 1200 according to the control instruction.
  • the movable platform 1200 may include an unmanned aerial vehicle, a ground robot, a remote control vehicle, etc., as shown in FIG. 12, the movable platform 1200 is schematically illustrated by using an unmanned aerial vehicle as a movable platform, wherein the following parts are mentioned. Unmanned aerial vehicles can be replaced with mobile platforms.
  • the identification device 1100 is installed at a suitable position of the unmanned aerial vehicle, wherein the identification device can be mounted on the outside of the aircraft of the unmanned aerial vehicle, or can be built in the body of the unmanned aerial vehicle, and is not specifically limited herein, for example.
  • the unmanned aerial vehicle may further include a pan/tilt head 1220 and an imaging device 1230, and the imaging device 1230 is mounted on the main body of the unmanned aerial vehicle through the pan/tilt head 1220.
  • the imaging device 1230 is configured to perform image or video capture during flight of the unmanned aerial vehicle, including but not limited to a multi-spectral imager, a hyperspectral imager,
  • the PTZ 1220 is a multi-axis transmission and stabilization system.
  • the PTZ motor compensates the shooting angle of the imaging device 1230 by adjusting the rotation angle of the rotating shaft, and prevents or by setting an appropriate buffer mechanism.
  • the jitter of the imaging device 1230 is reduced.
  • a gesture capable of generating the control command is referred to as a command gesture.
  • the mobile platform provided by the embodiment of the present invention can recognize the gesture of the user, and generate corresponding control instructions according to the gesture of the user, thereby implementing control of the movable platform.
  • the user can control the mobile platform through gestures, further enriching the control mode of the mobile platform, reducing the professional requirements for the user, and improving the fun of operating the mobile platform.
  • the processor 1210 is further configured to: after the identifying device 1100 recognizes the gesture of the user, illuminate the indicator light of the movable platform according to a preset control mode.
  • the indicator light on the UAV can be illuminated according to a preset control mode, for example, after successfully recognizing the gesture, on the UAV
  • the left navigation light and the right navigation light flash slowly, so that the user can know whether the gestures he has made are recognized by observing the flashing conditions of the left navigation light and the right navigation light, thereby avoiding the user not knowing whether his movement has been Identify and repeat the same gesture multiple times.
  • the recognition device continues to detect the user's palm and recognize.
  • the identifying device 1100 is configured to identify a confirmation gesture of the user after the indicator light of the movable platform is illuminated; the processor 1210 is configured to: after the identifying device 1100 recognizes the confirmation gesture, according to the The control commands control the mobile platform.
  • the user can know that his or her gesture has been recognized by observing the blinking of the unmanned aerial vehicle. To prevent false triggering, the user needs to confirm the previous gesture. This way After the user sees the indicator light of the unmanned aerial vehicle flashing, a confirmation gesture is made.
  • the processor After the identification device on the unmanned aerial vehicle successfully recognizes the confirmation gesture of the user, the processor generates a control instruction according to the previous command gesture, and controls the control according to the control instruction. Human aircraft. If the recognition device does not recognize the confirmation gesture within a preset time, the recognition device returns to the palm of the detection detection range to identify other command gestures of the user.
  • the mobile platform further includes a communication interface 1240, configured to receive an instruction for stopping gesture recognition, and the processor 1210 controls the identification device when the communication interface 1240 receives the instruction for stopping gesture recognition.
  • 1100 stops recognizing user gestures.
  • FIG. 13 is a schematic diagram of communication between a mobile platform and a control terminal according to an embodiment of the present invention. As shown in FIG. 13, specifically, a user may send a control command to the mobile platform through the control terminal 1300, where the control command is used. In order to cause the UAV to exit the gesture control mode, the recognition device 1100 no longer recognizes the gesture of the user.
  • the communication interface 1240 is further configured to receive an instruction to start gesture recognition, and the processor 1210 controls the identification device 1100 to start recognizing a user gesture when the communication interface 1240 receives the instruction to start gesture recognition.
  • the memory in this specification may include a volatile memory, such as a random-access memory (RAM); the memory may also include a non-volatile memory.
  • a volatile memory such as a random-access memory (RAM); the memory may also include a non-volatile memory.
  • a flash memory such as a hard disk drive (HDD), or a solid-state drive (SSD).
  • HDD hard disk drive
  • SSD solid-state drive
  • the processor may be a central processing unit (CPU).
  • the processor may further include a hardware chip.
  • the hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof.
  • ASIC application-specific integrated circuit
  • PLD programmable logic device
  • the above PLD can be Complex programmable logic device (CPLD), field-programmable gate array (FPGA), etc.
  • the steps of a method or algorithm described in connection with the embodiments disclosed herein can be implemented directly in hardware, a software module executed by a processor, or a combination of both.
  • the software module can be placed in random access memory (RAM), memory, read only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or technical field. Any other form of storage medium known.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Theoretical Computer Science (AREA)
  • Multimedia (AREA)
  • Remote Sensing (AREA)
  • Radar, Positioning & Navigation (AREA)
  • Aviation & Aerospace Engineering (AREA)
  • Automation & Control Theory (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Psychiatry (AREA)
  • Social Psychology (AREA)
  • Human Computer Interaction (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一种识别方法、设备(1100)及可移动平台(1200)。识别方法包括:获取用户的深度图像,确定与所述深度图像对应的点云(S201),对点云分类,从分类后的点云中确定指示手掌的点云(S202),根据指示手掌的点云确定手势(S203)。通过获取用户的深度图像来识别用户的手势,在TOF相机(1110)获得的深度图像的分辨率较低时,可以准确地从深度图像中提取出用户的手掌,同时在TOF相机(1110)采集帧率较低的时,可以准确地提取用户的手掌的运动轨迹,从而精准地识别用户的手势。

Description

一种识别方法、设备以及可移动平台 技术领域
本发明涉及图像处理领域,尤其涉及一种识别方法、设备以及可移动平台。
背景技术
手势识别是对用户的手势(例如手型、或手掌的运动轨迹)的识别。目前,主要利用结构光测量、多角度成像和TOF相机来进行手势识别。其中,TOF相机由于成本低廉,易于小型化,被广泛地应用在手势识别中。然而,由于从TOF相机获取的深度图像分辨率低,而且TOF相机数据采集帧率低,导致使用TOF相机来进行手势识别时,识别准确率不高,尤其是当TOF相机应用于可移动平台来进行手势识别时。
发明内容
本发明实施例提供了一种识别方法、设备和可移动平台,以提高手势识别的准确率。
例如,本发明实施例第一方面提供一种识别方法,包括:
获取用户的深度图像,确定与所述深度图像对应的点云;
对点云分类,从分类后的点云中确定指示手掌的点云;
根据指示手掌的点云来确定手势。
本发明实施例第二方面提供另一种识别方法,包括:
获取用户的深度图像,根据所述深度图像确定指示手掌的二维图 像的点集;
根据所述点集来确定手势。
本发明实施例第三方面提供一种识别设备,包括:
TOF相机,用于获取用户的深度图像;
处理器,用于确定与所述深度图像对应的点云,对点云分类,从分类后的点云中确定指示手掌的点云,根据指示手掌的点云确定手势。
本发明实施例第四方面提供一另种识别设备,包括:
TOF相机,用于获取用户的深度图像;
处理器,用于根据所述深度信息确定指示手掌的二维图像的点集,根据所述点集来确定手势。
本发明实施例第五方面提供一种可移动平台,包括:
如前所述的任一项识别设备,用于识别用户的手势;
处理器,用于根据所述识别设备识别的手势生成相应的控制指令,并根据所述控制指令控制可移动平台。
本发明实施例提供一种手势识别方法、设备以及可移动平台,通过获取用户的深度图像来识别用户的手势,在TOF相机获得的深度图像的分辨率较低时,可以准确地从深度图像中提取出用户的手掌,同时在TOF相机采集帧率较低的时,可以准确地提取用户的手掌的运动轨迹,从而精准地识别用户的手势。另外,根据识别到的手势,生成与所述手势对应的控制指令,利用所述控制指令对可移动平台进行控制,简化了可移动平台的控制流程,丰富了可移动平台的控制方式,进一步提高了操控可移动平台的趣味性。
附图说明
为了更清楚的说明本发明实施例的技术方案,下面将对本发明实施例描述中所需要使用的附图作简单的介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本发明一种实施例中手势识别系统的示意图;
图2为本发明一种实施例中识别方法的流程图;
图3为本发明一种实施例中对指示用户的点云进行分类的示意图;
图4为本发明一种实施例中从指示用户的手部的二维图像的点集中删除指示用户的手臂的二维图像的点集的示意图;
图5为本发明一种实施例中根据二维图像的点集确定指示用户的手掌的点集的分布特征的示意图;
图6为本发明一种实施例中根据手掌的位置信息确定位置信息对应的速度方向的示意图;
图7为本发明另一种实施例中根据手掌的位置信息确定位置信息对应的速度方向的示意图;
图8为本发明另一种实施例中根据手掌的位置信息对应的速度方向确定位置信息对应的运动方向的示意图;
图9为本发明一种实施例中识别打勾手势的示意图;
图10为本发明又一种实施例中识别方法的流程图;
图11为本发明一种实施例中识别设备的示意图;
图12为本发明一种实施例中可移动平台的示意图;
图13为本发明一种实施例中可移动平台与控制终端通信的示意图;
具体实施方式
本发明实施例提供一种手势识别方法设备以及无人飞行器,通过获取用户的深度图像来识别用户的手势。另外,根据识别到的手势,生成与所述手势对应的控制指令,利用所述控制指令对无人飞行器进行控制,丰富了无人飞行器的控制方式,提高了操控无人飞行器的趣味性。
下面将结合本发明实施例中的附图,对本发明实施例中的技术方案进行清楚地描述,显然,所描述的实施例仅仅是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的范围。
除非另有定义,本文所使用的所有的技术和科学术语与属于本发明的技术领域的技术人员通常理解的含义相同。在本发明的说明书中所使用的术语只是为了描述具体的实施例,不是旨在限制本发明。本文所使用的术语“及/或”包括一个或多个相关的所列项目的任意的和所有的组合。
下面通过具体实施例,分别进行详细说明。
TOF相机标定
TOF相机标定的作用是将深度图像中的二维图像坐标与相机坐 标系下的坐标对应起来,结合TOF相机获取的深度信息,则可以得到每一个二维图像坐标对应的相机坐标系下的三维坐标,即三维点云,简称点云。TOF相机标定的目的在于保证点云中各部分之间的相对位置关系与真实世界保持一致。
TOF相机的成像原理与一般的针孔型相机相同,只是TOF相机的接收器只能接收到由目标对象反射的调制红外光,TOF相机得到的幅值图像与一般相机得到的灰度图像相同,标定方式也可以借鉴。
令二维图像中的坐标为(u,v),世界坐标系的坐标为(X,Y,Z),则有
Figure PCTCN2017075193-appb-000001
其中
Figure PCTCN2017075193-appb-000002
为相机的内参矩阵,R是世界坐标系相对于相机坐标系的旋转矩阵,T是世界坐标系相对于相机坐标系的平移向量,α为比例系数。
根据张正友相机标定算法用黑白棋盘格作为标定的图案,对于每一帧标定图像,利用角点检测获得两组对应点,一组是棋盘格坐标系上每个角点的坐标
Figure PCTCN2017075193-appb-000003
在标定前测量并记录,另一组是角点检测出的对应点的二维图像坐标
Figure PCTCN2017075193-appb-000004
理论上两组点应该符合公式(1),实际上图像噪声和测量误差使得我们只能求出一个最小二乘解:
在棋盘格坐标系中Z值为0,由公式(1)可得
Figure PCTCN2017075193-appb-000005
Figure PCTCN2017075193-appb-000006
对于每一帧标定图像,设
Figure PCTCN2017075193-appb-000007
Figure PCTCN2017075193-appb-000008
Figure PCTCN2017075193-appb-000009
利用两组对应点可以优化出单应性矩阵H。优化方法如下:
Figure PCTCN2017075193-appb-000010
其中i代指了
图像中每一组对应点,则优化的目标函数为:
Figure PCTCN2017075193-appb-000011
Figure PCTCN2017075193-appb-000012
则公式(1)可以转化为如下的形式:
Figure PCTCN2017075193-appb-000013
这是一个2*9的矩阵,对应一个线性方程组,对于图像中的所有i组对应点则能写出一个2i*9矩阵,对应了9个未知数和2i个等式构成的方程组,对于这样的方程组,其最小二乘解就是目标函数(2)的最优解。
该最优解对应了一帧图像中的单应性矩阵H,而H=K[r1 r2 T],要想通过每一个H求解出相机的内参矩阵K,还需要下述约束:
Figure PCTCN2017075193-appb-000014
Figure PCTCN2017075193-appb-000015
由于r1,r2正交且都为单位向量。
设B=K-TK-1,可以将
Figure PCTCN2017075193-appb-000016
表示成
Figure PCTCN2017075193-appb-000017
的形式,其中b为B中各 个元素拉成的一列6维的向量(由于B为实对称阵,只有6个元素是待定的),则约束可以表示为一个方程的形式:
Figure PCTCN2017075193-appb-000018
对于每一帧图像都有上述方程成立,则n幅图像对应了一个2n个等式6个未知数的线性方程组,也可以对其求最小二乘解,获取最优的B,从而解出相机内参矩阵K。
利用内参矩阵K,我们可以通过TOF相机获取的某一点深度z与该点在二维图像中的坐标
Figure PCTCN2017075193-appb-000019
求得某一在相机坐标系中的实际坐标,利用公式
Figure PCTCN2017075193-appb-000020
可以求出相机系坐标
Figure PCTCN2017075193-appb-000021
其中,每一点的三维坐标与每一个二维图像坐标是一一对应的,并且在完成TOF相机标定后,存储器会始终存储所关注的点的这两种坐标。
基于TOF相机的手势识别系统
图1为本发明一种实施例中手势识别系统的示意图,如图1所示,本实施例提供的一种基于TOF相机的手势识别系统,其中所述系统包括发射器101,其中发射器101可以为发光二极管(Light Emitting Diode,简称LED)或激光二极管(Laser Diode,简称LD),其中发射器由驱 动模块102来驱动,驱动模块102由处理模块103控制,处理模块103控制驱动模块102输出驱动信号来驱动发射器101,其中驱动模块102输出的驱动信号的频率、占空比等都可以由处理模块103控制,利用驱动信号驱动发射器101,发射器101发出经过调制以后的光信号,光信号打到目标对象上,在本实施例中,目标对象可以是用户,光信号打到用户上时,光信号会发生反射,接收器104会接收到由用户反射的光信号。其中,接收器104可以包括光电二极管、雪崩光电二极管、电荷耦合元件,由用户反射的光信号包括由用户的手部反射的光信号,接收器104将光信号转化成电信号,信号处理模块105对接收器104输出的信号进行处理,例如放大、滤波等,经过信号处理模块105处理的信号输入到处理模块103中,处理模块103可以将信号转换成深度图像,深度图像中包含用户的手掌的位置信息和深度信息。需要说明书的是,在某些情况中,手势识别系统中可以不包括信号处理模块105,接收器104可以直接将电信号输入到处理模块103,或者,信号处理电路105可以包括在接收器104或处理模块103中。在某些情况中,处理模块103可以将深度图像输出,进入判别模块106,判别模块106可以根据深度图像来识别用户的手势;另外,在某些情况中,手势识别系统可能不包括判别模块106,处理模块103在将所述信号转换成深度图像后,可以直接根据深度图像来识别用户的手势。
图2为本发明实施例提供的一种手势识别方法的流程图,包括:
S201:获取用户的深度图像,确定与所述深度图像对应的点云;
具体的,用户在手势识别设备的TOF相机的探测范围内,做出特 定的手势,其中手势包括手掌的动态手势,即用户移动手掌形成的手势,例如手掌上下移动、手掌左右移动、手掌前后移动等,另外手势还包括手掌的静态手势,即用户的手型,例如握拳、伸掌、伸出一根手指、伸出两根手指等,手势识别设备包括TOF相机,所述TOF相机发射的光信号射向用户,TOF相机接收用户反射的光信号,TOF相机对接收到的光信号进行处理,输出用户的深度图像。TOF相机经过上述的标定后,可以根据深度图像计算出用户的点云。另外,在TOF相机获取到一帧深度图像时,可以设定一个捕捉中心,再以捕捉中心为球心,在以预设的距离阈值为半径的球形空间内获取点云以排除干扰。其中捕获中心可以设定在TOF相机的正前方,例如捕捉中心可以设置在TOF相机正前方的0.4-2m的范围内,具体地可以将捕获中心设置在TOF相机正前方0.8m、1m、1.1m、1.2m、1.3m、1.4m、1.5m、1.6m、1.7m,其中,预设的距离阈值本领域技术人员可以根据设计需求选取,例如预设的距离阈值可以在10-70cm的范围内,具体地可以选择20cm,25cm、30cm、35cm、40cm、45cm、50cm、55cm等。
S202:对所述点云分类,从分类后的点云中确定指示手掌的点云;
由于用户的点云中可能包括用户身体多个部分的点云,例如手部、头部、躯干部分等,为了确定手势,需要先将指示用户手部的点云提取出来,可以对用户的点云进行分类,分类后会得到至少一个点云的聚类,从分类得到的聚类中确定指示用户手掌的点云,得到用户的手掌的点云即提取出用户的手掌,基于指示用户手掌的点云来识别用户的手势。
S203:根据所述手掌的点云来确定手势。
具体地,指示用户手掌的点云可以指示用户手掌的位置信息、或手掌的轮廓信息等,通过点云包含的位置信息可以识别用户的动态手势,通过手掌的轮廓信息可以识别用户的静态手势。
本发明的实施例中,获取用户的点云,对用户的点云进行分类,从分类后得到的点云中确定出指示手掌的点云,根据指示手掌的点云来识别用户的手势。通过本发明实施例,在TOF相机获得的深度图像的分辨率较低时,可以准确地从深度图像中提取出用户的手掌,同时在TOF相机采集帧率较低的时,可以准确地提取用户的手掌的运动轨迹,从而精准地识别用户的手势,节省了运算资源,识别率高。
可选的,对点云分类获得点云的多个聚类,从多个聚类中的一个来确定指示手掌的点云。具体的,根据先验信息,用户在面对TOF相机做手势时,用户身体的头部、躯干部分、手部、脚部等与TOF相机之间的距离是不同的,即用户身体的躯干部分和手部的深度信息是不同的,另外用户在面对TOF相机做手势时,用户身体的相同部位的点云一般是相互靠近的,因此可以根据用户在做手势时,身体的各个部位分布在不同的空间位置的先验信息,将用户在TOF相机探测范围内的各个身体部分进行分类,分类后得到至少一个点云的聚类,不同聚类一般表示用户身体的不同部分,通过分类可以将身体的不同部位区分开,这时只需要在分类得到的某个特定的部分中确定属于手掌的部分,即在分类得到的某个点云的聚类来确定指示用户手掌的点云,这样可以缩小确定用户手掌的搜索范围,提高识别的准确性。
在某些实施例中,利用聚类算法来对点云分类。具体的,可以选用聚类算法中的K-均值分类(k-means)来进行分类,K-均值聚类是一种非监督的分类算法,必须事先指定分类的聚类类数,如果可以确定TOF探测范围内中只有人体的躯干部分和手部,那分类类数可以定为2类,但在实际情况中TOF相机的探测范围内除了包含用户,还可能包含了除用户以外的其他的物体,或者在TOF的探测范围内只有用户的手部而并没有用户的躯干部分,因此聚类类数是不确定的。如果分类类数大于实际类数,则会将本来应分为一类的点云分裂,反之会将本不属于一类的点归为一类。因此,在本发明的实施例中,在使用聚类算法对点云分类的过程中,聚类算法的聚类类数是可调整的,下面将对聚类算法的聚类类数的调整进行详细地说明。
具体的,可以根据聚类之间的分散程度来调整聚类类数,其中所述分散程度可以用各个聚类的聚类中心之间的距离来表示。在进行聚类算法之前,初始聚类类数设置为n,例如可以将n设置为3,n是一个在聚类运算过程中可以调整的参数,进行一次K-均值聚类,并获取每个聚类中心,然后计算各个聚类中心的分散程度,如果其中有两个聚类中心之间的距离小于或等于在分类算法中设置的距离阈值时,则将n减小1,并重新进行聚类,阈值也是可以调节的参数,例如阈值可以设置在10-60cm范围内,具体地可以为10cm、15cm、20cm、25cm、30cm等。如果聚类算法的分类效果较差,则将n增加1,并重新聚类。当发现所有聚类中心之间的距离都大于或者等于所述阈值时,停止执行聚类算法,此时对指示用户的点云分类完毕,返回当前的聚类类数 和聚类中心。
可选的,根据深度信息从多个聚类中确定指示手部的点云的聚类,从所述指示手部的点云的聚类中确定指示手掌的点云。具体的,对用户的点云进行分类,可以得到至少一个聚类。图3为本发明一种实施例中对指示用户的点云进行分类的示意图,如图3所示,在对用户的点云进行分类后,可能获取4个聚类,4个聚类分别是聚类301、聚类302、聚类303、聚类304,每一个聚类代表不同的平均深度,根据先验信息,用户在面对TOF相机做手势时,手部是离TOF相机最近的,即用户的手部的深度是最小的,因此,获取分类后得到的聚类中每一个的平均深度,例如获取聚类301、聚类302、聚类303、聚类304的平均深度,将平均深度最小的聚类确定为指示用户手部的聚类,即将聚类301确定为指示用户手部的聚类,这样从用户的点云中确定出用户手部且包含手掌的点云,在获得了指示用户手部的点云,可以进一步从指示用户手部的点云确定指示用户手掌的点云。其中,对用户的点云进行分类,得到4个聚类只是为了进行示意性说明,并不对本实施例的技术方案构成限定。
可选地,从所述指示手部的点云的聚类中删除指示手臂的点云,将所述聚类中剩余的点云确定为指示手掌的点云。具体的,用户的手部包括用户的手掌和手臂,在用户手部的点云的聚类中通常包含指示手臂的点云,为了提升手势识别的准确率,从指示手部的点云中确定出手臂的点云,将指示手臂的点云删除,剩余的点云即被确定为手掌的点云,这样就准确地提取出了手掌,后期根据手掌的点云来识别手 势。下面将详细介绍从指示手部的点云中删除手臂的点云的方法。
在某些实施例中,在指示手部的点云的聚类中取出深度最小的点,确定聚类中的点云与深度最小点的距离,将距离大于或等于距离阈值的点确定为手臂的点云。具体的,在上述指示用户手部的点云的聚类中,通常会包括手部中的手臂,在进行具体的手势识别之前,需要将手部中的手臂的点云删除。首先计算指示用户手部的点云的聚类的深度直方图,通过直方图可以取出深度最小的点,这个深度最小的点通常为手指的指尖,计算聚类中的其他点到深度最小的点的距离,将所有距离超过距离阈值的点确定为指示手臂的点,将其删除,保留剩余的点,即将距离小于或等于距离阈值的点确定为指示用户手掌的点云。其中,所述距离阈值可以根据要求改变,或者根据手掌的平均大小来确定,比如10cm、13cm、15cm、17cm等。
在某些实施例中,根据指示手部的点云的聚类确定指示手部的二维图像的点集,确定所述手部的二维图像的点集的最小外接矩形,确定所述手部的二维图像的点集中的点到所述最小外接矩形的指定边的距离,当距离不符合预设的距离要求的点确定为指示手臂的二维图像的点,确定指示手臂的二维图像的点集,根据指示手臂的二维图像的点集确定指示手臂的点云,将指示手臂的点云删除。具体的,获取一帧深度图像,确定这一帧深度图像中用户的点云,根据上述方法可以确定出这一帧深度图像的指示用户手部的点云。由于每一个点云的三维坐标与二维图像上的点的二维坐标是一一对应的,并且在手势识别的过程中会始终存储这两种坐标,在获取了指示用户手部的点云后, 可以确定出用户手部的二维图像的点集。其中,图4为本发明一种实施例中从指示用户的手部的二维图像的点集中删除指示用户的手臂的二维图像的点集的示意图,如图4所示,获取所述用户手部的二维图像的点集的最小外接矩形401,确定用户手部的二维图像的点集中的点到外接矩形指定边的距离,将距离不符合预设的距离要求的点确定为指示手臂的点,即将距离不符合预设的距离要去的点作为指示手臂点而删除,剩下的点集即为指示手掌的二维图像的点集,根据手掌的二维图像的点集即可以得到手掌的点云。
如图4所示,所述预设的距离要求可以通过最小外接矩形401的边长来确定。具体的,取矩形的长边长为w,取矩形的短边长为h,例如指定边可以为靠下的短边,计算指示手部中二维图像的点集中的每一个点到靠下的短边的距离di,若di<w-1.2h,则将该点确定为指示手臂的点,按照这种方式,可以将所有指示手臂的二维图像的点集删除,剩余的点集基于指示手掌的二维图像的点集,即可以将指示手臂的点云删除。这里使用了一个假设,即是h表示手掌在二维图像上的宽度,根据手掌的宽度和长度之间长度比例关系,将1.2h确定为手掌的长度,w与1.2h之间的差应该是手臂上的点到靠下的短边的最大距离,如果二维图像的点集中的某个点到靠下的短边的距离小于或等于这个最大距离,即表示该点是属于手臂的点。其中d<w-1.2h只是根据最小外接矩形的边长确定预设的距离要求的一种实施方式,本领域技术人员还可以选取其他的方式,例如所述的手掌的长度可以为1.1h、1.15h、1.25h、1.3h、1.35h、1.4h等,在这里不做具体的限定。
可选地,根据指示用户手掌的点云获取手掌的二维图像的点集,根据所述点集的分布特征确定手势。具体的,这里识别的是用户的静态手势,例如握拳、伸掌、伸出一根手指、伸出两根手指等。其中,获取一帧深度图像,确定这一帧深度图像中用户的点云,根据上述方法可以确定出这一帧深度图像的指示用户手掌的点云。由于每一个点云的三维坐标与二维图像上的点的二维坐标是一一对应的,并且在手势识别的过程中会始终存储这两种坐标,在确定了指示用户手掌的点云后,可以确定出用户手掌的二维图像的点集。由于用户手势的不同,即不同的手势对应不同的手型,手掌的二维图像的点集的分布特征会不一样,例如握拳手势的分布特征和伸掌手势的分布特征差别很大,因此可以确定手掌的二维图像的点集的分布特征,根据分布特征来具体判断用户在这一帧图像中所做的手势。
可选地,确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。具体的,可以根据指示手掌的二维图像的点集来确定点集的分布区域,图5为本发明一种实施例中根据二维图像的点集确定指示用户的手掌的点集的分布特征的示意图,如图5所示,在某些实施例中可以使用创建图像掩膜(mask)的方法确定点集的分布区域,即图5中的区域501,其中所述分布区域501即是用户的手掌在二维图像上所占的区域,不同的手势的分布区域的形状和轮廓不相同,通过分布区域501的形状和轮廓可以确定指示手掌的点集的分布特征,根据分布特征可以识别用户的手势。
可选的,使用多边形区域覆盖所述分布区域,确定所述多边形区 域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定所述点集的分布特征。具体的,由于分布区域是通常是不规则的形状,为了进一步来识别分布区域的特征,可以指示手掌的二维图像的点集的所有点像素值设置为1,二维图像中的其他点像素值设置为0,使用多边形覆盖所述分布区域,即使用多边形覆盖点集中的所有点,其中所述多边形为边数最少的凸多边形。如图5所示,在某些实施例中,可以对二值化后的指示手掌的二维图像的点集进行凸包操作,则可以用边数最小的凸多边形502覆盖所述点集,该凸多边形的每一个顶点都是点集中的一个点,这样二维图像的点集的分布区域501与所述多边形之间502存在不重叠区域503,不重叠区域503的形状和大小可以表现所述点集的分布特征,即可以根据不重叠区域的某些特征可以对用户的手势进行识别。
可选的,确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述距离确定为所述点集的分布特征。具体的,如图5所示,可以根据不重叠区域来确定所述点集的分布特征,例如在不重叠区域503中,不重叠区域503对应的多边形的边是li,从不重叠区域503中确定到边li距离最远的一个点,将所述最远的距离di作为所述点集的一个分布特征,根据距离di可以对用户的手势进行识别。值得注意的是,其中上述di可以是一个最远距离,也可以是多个最远距离之间的组合,本领域技术人员可以按照需求选取,图5只是进行示意性说明。
可选的,当多边形的每一条边对应的最远距离都小于或等于预设的距离阈值时,确定手势为握拳,多边形的每一个对应的最远距离中 至少有一个大于或等于预设的距离阈值时,确定手势为伸掌。具体的,当用户伸开手掌时,指示手掌的二维图像的点集形成的分布区域与多边形之间的不重叠区域较大,具体表现为包围手掌的多边形的边距离手指之间的关节距离较大,且在伸开手掌时,会形成多个这样的不重叠区域,这与握拳时形成的不重叠区域明显不同,由于在握拳时,指示手掌的二维图像的点集形成的分布区域的形状会趋向于多边形,所以,在进行凸包操作后,分布区域与多边形形成的不重区域面积小,因此多边形的每一条边对应的最远距离都比较小,这样可以设置预设的距离阈值,当多边形的每一条边对应的最远距离都小于或等于预设的距离阈值时,确定手势为握拳,多边形的每一个对应的最远距离中至少有一个大于或等于预设的距离阈值时,确定手势为伸掌。另外,第二阈值可以根据手指的长度来选定。
可选的,获取多帧用户的深度图像,确定多帧深度图像中的每一帧深度图像对应的指示用户手掌的点云;对所述每一帧深度图像对应的点云分类,从分类后的点云中确定每一帧深度图像对应的指示手掌的点云;根据所述每一帧图像对应的指示用户手掌的点云确定每一帧深度图像对应的手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
具体的,可以通过多帧深度图像来判断用户的手势,这里的手势是指用户通过移动手掌而形成的手势。为了识别手势,首先要将每一帧深度图像中的手掌提取出来,按照本文前述部分的方法,可以根据每一帧深度图像来获取每一帧图像对应的用户手掌的点云,根据每一 帧图像对应的用户手掌的点云可以计算出手掌的位置信息,其中可以将指示手掌的点云的几何中心的位置作为手掌的位置信息,另外也可以将指示手掌的点云中深度信息最小的点的位置作为手掌的位置信息。其中,本领域技术人员可以采用不同的方式来根据指示用户手掌的点云来确定手掌的位置信息,在这里不做具体的限定。在具体实现中,可以将从多帧深度图像计算得到的手掌的位置信息存储到序列P中,其中序列P的长度为L,采用先进先出的存储方式,使用最近获取的手掌的位置信息替换到最旧的手掌的位置信息。这个序列P反应了在固定时间内手掌运动的轨迹,所述轨迹即表示用户的手势,因此可以根据序列P即手掌的位置信息的序列来识别用户的手势。另外,在获得一帧深度图像对应的手掌的位置信息后,可以将该位置信息指示的位置点作为捕捉中心,在确定下一帧深度图像对应的手掌的位置信息时,可以在以捕捉中心为球心,以预设的距离阈值为半径的球形空间内获取用户的点云,即只在该球形空间内提取用户的手部,这样可以提高手部的识别速度。另外,可以使用卡尔曼滤波算法对手掌的运动模型进行估算,预测下一帧深度图像指示的手掌的位置,同时可以在预测的手掌的位置附近来提取用户的手掌。另外该滤波算法可以随时开启或者关闭。
可选的,根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌运动的运动方向,根据所述运动方向组成的序列确定手势。具体的,根据序列P中L个位置信息,可以计算出位置信息对应的运动方向,其中,可以确定L个位置信息中每一个对应的运动方 向,也可以确定L个位置信息中多个位置信息中每一个对应的运动方向,得到的多个运动方向组成的运动方向序列可表示手掌在空中的运动轨迹以及运动变化情况,因此,根据运动方向组成的序列可以确定用户的手势。值得注意的是,序列P中的位置信息对应运动方向可以是该位置信息对应的速度方向,也可以是根据所述速度方向以某种方式确定的方向。
可选的,确定所述运动方向中每一个运动方向所占的比例,根据所述比例的组合确定所述手掌运动的动作。具体地,统计所述运动方向中每一个运动方向在运动方向的序列中所占的比例,这样可以得到比例组成的比例序列,利用所述比例的序列来识别用户的手势。这样在用户做手势的,不管用户做手势时,手掌运动的起点和终点在哪里,都可以得到相同形式的比例的序列,这样便于处理。在进行手势识别时,将所述比例的序列输入预设的运算模型中,所述预设的运算模型就会根据比例的序列识别用户的手势。其中,所述预设的运算模型可以为神经网络、分类器等,在进行手势识别之前,需要对预设的运算模型进行训练,即需要在离线采集大量的手势对应的比例的序列,以所述比例的序列作为输入,比例序列对应的手势作为输出,对预设的运算模型进行训练,训练完成后,所述预设的运算模型就可以用来进行手势识别。
可选的,根据指示手掌的位置信息的序列,确定位置信息对应的速度方向,根据所述速度方向的序列确定运动方向的序列。具体的,由于TOF相机采集数据的帧率比较低,导致指示手掌的位置信息十分 离散,很难获取每一帧深度图像中手掌运动的切向速度方向。在本实施例,图6为本发明一种实施例中根据手掌的位置信息确定位置信息对应的速度方向的示意图,如图6所示,Pi代表一帧深度图像指示的手掌的位置点,即位置信息指示的位置点,此时为了进行示意性说明将L取为7,速度方向使用手掌的位置信息的序列来确定,即根据手掌的位置点序列来确定,具体的,手掌的位置P2的速度方向是从位置点P1指向位置点P2,手掌的位置P3的速度方向是从位置点P2指向位置点P3,以此类推,可以获得速度方向的序列(V1,V2¨V6),该速度方向的序列可以表示手掌的运动方向的变化情况,根据所述速度方向的序列可以确定运动方向的序列。值得注意的是,序列的长度L为7只是示意性说明,本领域技术人员可以选用需要选取长度L的值。另外,本领域技术人员可以采用其他方式计算位置点对应的运动方向,例如,图7为本发明另一种实施例中根据手掌的位置信息确定位置信息对应的速度方向的示意图,如图7所示,手掌的位置P3的运动方向是从位置点P1指向位置点P3,手掌的位置P4的运动方向是从位置点P2指向位置点P4,以此类推,可以获得速度方向的序列(V1,V2¨V5)。需要说明的是,在获得速度方向的序列后,可以对速度方向的序列输入滤波器中进行滤波,具体地,可以将速度方向的序列输入卡尔曼滤波器,这样可以将速度方向的序列中的噪声或者变化异常的速度方向滤除。
在某些实施例中,确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。为了简洁,本文只对速度方向序列中的一个速度方 向是如何确定对应的运动方向来进行示意性说明,速度方向的序列中的其他速度方向确定对应的运动方向的方法是相同的。具体的,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向,由于通过前述部分计算出的位置信息对应的速度方向十分离散,为了便于后期统计处理,非常有必要对速度方向进行归类,将相差不大的速度方向归类为同一个方向。图8为本发明另一种实施例中根据手掌的位置信息对应的速度方向确定位置信息对应的运动方向的示意图,如图8所示,设定多个预设的方向,例如Vu、Vp、Vl、Vr、Vf、Vd,分别表示上、下、左、右、前和后六个预设方向,根据前述部分的方向计算得到位置点对应的速度方向Vi,将该速度方向对应的单位向量分别于六个预设方向中每一个对应的单相向量进行点乘,可以计算出α1~α6,可以根据α1~α6来确定该位置信息的运动方向,具体的可以从α1~α6确定角度最小的αi,将与αi对应的第一预设方向(例如图7所示的Vr)确定为该位置信息对应的运动方向。其中设置上、下、左、右、前和后六个预设方向只是为了示意性说明,在硬件条件允许的情况下,本领域技术人员可以设置更多的预设方向,这样可以使速度方向的归类更加精准,使得位置信息对应的运动方向误差更小,总之,预设方向个数的选取,本领域技术人员可以根据设计需求和/或硬件条件选取,在这里不做具体的限定。
可选的,根据所述位置信息的序列确定位置信息对应的速率,当所述速率小于预设的速率阈值时,确定手掌在该位置信息指示的位置点时为静止状态。具体的,如图6所述,可以根据位置信息序列P来确 定序列P中位置信息对应的速率,即手掌在位置信息指示的位置点时的速率,其中速率可以根据手掌的位移来计算,其中位移可以根据位置序列中的位置信息计算,由于位置信息序列P中的相邻两个位置信息之间的间隔时间相同,在这里可以不用引入时间信息,直接使用位移来表示位置信息对应的速率,例如P2对应的速率为P1指向P2的位移,其中所述位移可以根据P1、P2的位置信息获取得到,同理可以获取P3、P4、P5、P6、P7对应的速率,当速率小于或等于速率阈值时,则认为手掌此时处于静止状态,没有运动方向。另外,位置信息对应的速率也可以采用其他方式计算,例如如图7所示,P3对应的速率为P1指向P3的位移,在这里不做具体的限定。
在某些实施例中,为了避免不同手势之间切换时被误判成用户的画圆手势,可以根据所述位置序列获取对应的二维图案坐标序列,即获取在二维图像上的点,对二维图像上的点对应的向量循环叉乘,即每一个点与后一个点叉乘,最后一个点与第一个点叉乘,则可以计算出在二维图像上围成的面积,当所述面积小于或等于预设的面积阈值时,则确定用户当前的手势不是画圆手势。使用面积来判定,在一定程度上消除了在不同的手势切换时可能存在的误识别。其中,预设的面积阈值,本领域技术人员可以根据设计需求选取,例如40、50、60、70等。
在某些实施例中,根据前述部分的位置信息序列来识别用户的打勾手势。具体的,获取指示手掌的位置的序列在XY平面上的投影序列,遍历所述投影序列中的点,若从所述序列的点中确定满足预设要 求的特定点,则确定识别到打勾手势。当用户在做打勾手势时,手掌距离TOF相机的距离近似不变,即在三维空间上Z方向的值基本不变,在判断打勾手势时,可以不考虑Z坐标,将所述位置信息序列投影到XY平面,其中,根据先验信息,在用户做打勾手势时,在手掌的运动轨迹在XY平面上会有一个最低点,在这最低点两侧的手势的运动轨迹近似为直线,且这两条近似直线的斜率正负相反,则可以将最低点确定为满足预设要求的特定点,所述特定点与序列中在所述特定点之前的点构成的第一运动轨迹确定为近似直线,所述特定点与在序列中所述特定点之后的点构成的第二运动轨迹确定为近似直线;且所述第一运动轨迹的斜率与所述第二轨迹的正负相反。
在某些实施例中,获取投影序列中的一个点,将所述点作为当前点,获取序列中在所述当前点之前的点,对当前点与序列中在所述当前点之前的点进行直线拟合,获取第一相关系数和第一斜率,若相关系数大于或等于相关系数阈值则确定第一运动轨迹为近似直线,获取序列中在所述当前点之后的点,对所述当前点与序列中在所述当前点之后的点进行直线拟合,获取第二相关系数和第二斜率,若相关系数大于或等于相关系数阈值则确定第一运动轨迹为近似直线,第二运动轨迹近似为直线,且第一斜率和第二斜率正负相反,则确定当前点为满足预设要求的特定点,若第一相关系数、第二相关系数中的中一个或两个小于或等于相关系数阈值,或第一斜率与第二斜率正负相同,则获取投影序列中的下一个点,将所述下一个点作为当前点。具体地,图9为本发明一种实施例中识别打勾手势的示意图,如图9所示,若 P4为当前点时,从P4点向前遍历,获取到P4点之前的点P3、P2、P1,对P4、P3、P2、P1进行直线拟合,得到轨迹l1,可以得到轨迹l1的相关系数和斜率k1,同理从P4点向后遍历,获取到P4点之后的点P5、P6、P7,对P4、P5、P6、P7进行直线拟合,得到轨迹l2,可以得到轨迹l2的相关系数和斜率k2,当轨迹l1和轨迹l2的相关系数都大于或等于相关系数阈值时,则判断轨迹l1和轨迹l2为直线,否则将下一点作为当前点,继续重复上述操作。当判断轨迹l1和轨迹l2为直线,获取轨迹l1和轨迹l2的斜率k1、k2,若两个斜率正负相反,则可以判断当前用户的手势为打勾手势。这样,通过遍历投影序列中点的方式,直到获取特定点,则可以确定当前的手势为打勾手势。其中相关系数,本领域技术人员可以根据需求选取,例如可以选为0.8。
在某些实施例中,从当前点向当前点之前的点遍历,获取当前点与遍历过的当前点之前的点的位移和,当位移和大于或等于预设的位移阈值时,停止遍历,对当前点与遍历过的当前点之前的点进行直线拟合;从当前点向当前点之后的点遍历,获取当前点与遍历过的当前点之后的点的位移和,当位移和大于或等于预设的位移阈值时,停止遍历,对当前点与遍历过的当前点之后的点进行直线拟合。具体的,如图9所示,若当前点为若P4为当前点时,从P4点向前遍历,获取到P4点之前的点P3,获取P4和P3之间的位移d1,判断位移d1是否大于或等于预设的位移阈值D,若小于D,继续遍历P2,获取P2和P3之间的位移d2,判断d1+d2是否大于预设的位移阈值D,如果d1+d2大于预设的位移阈值D,停止遍历,对P4、P3、P2进行直线拟合, 否则继续遍历P1并重复上述操作,直到位移和大于预设的位移阈值D停止遍历,对当前点和遍历过当前点之前的点进行直线拟合。同理,对当前点之后的点进行遍历也是执行上述操作,为了简便,不再赘述。若当前点与所有遍历过的当前点之前的点的位移和小于D时,则说明当前离投影序列的一端太近,会导致当前点之前点个数不足或者当前点之前的信息量不足,将下一个点作为当前点。同理,若当前点与所有遍历过的当前点之后的点的位移和小于D时,将下一个点作为当前点。
本发明的实施例提供了一种计算机存储介质,该计算机存储介质中存储有程序指令,该计算机存储介质中存储有程序指令,所述程序执行上述识别方法。
如图10所示,本发明实施例还提供一种识别方法,包括:
S1001:获取用户的深度图像,根据深度信息确定指示手掌的二维图像的点集;
具体的,用户在手势识别设备的TOF相机的探测范围内,面对手势识别设备做出手势,其中手势包括手掌的动态手势,即用户移动手掌形成的手势,例如手掌上下移动、手掌左右移动、手掌前后移动等,另外手势还包括手掌的静态手势,用户的手型,例如握拳、伸掌等,手势识别设备包括TOF相机,所述TOF相机发射的光信号射向用户,TOF相机接收用户反射的光信号,TOF相机对接收到的光信号进行处理,输出用户的深度图像。根据先验信息,用户在面对TOF相机做手势时,身体的各个部分离TOF相机的距离是不一样的,即深度不一样,因此可以根据深度信息来确定用户手掌的二维图像的点集,即获取手 掌在二维图像上的所有点的图像坐标
Figure PCTCN2017075193-appb-000022
S1002:根据所述点集来确定手势。
具体的,在获取指示手掌的二维图像的点集后,即已经成功提取用户的手掌。在提取到指示手掌的二维图像的点集后,可以根据所述点集来识别用户的手势。
本发明的实施例中,根据用户的深度图像来确定指示手掌的二维图像的点集,根据所述点集识别用户的手势。通过本发明实施例,在TOF相机获得的深度图像的分辨率较低时,可以准确地从深度图像中提取出用户的手掌,同时在TOF相机采集帧率较低的时,可以准确地提取用户的手掌的运动轨迹,从而精准地识别用户的手势,节省了运算资源,识别率高。
可选的,根据深度信息在二维图像上确定一个指示手掌的点,根据在预设的深度范围内确定与指示手掌的点连通的点集,根据所述连通的点集确定指示手掌的二维图像的点集。具体的,根据先验信息,用户在面对TOF相机做手势时,手掌离TOF相机的距离是最近的,手掌的点的深度是最小的,可以取出深度最小的一个点作为指示手掌的点,这个点通常为指尖,另外,也可以取出三个深度最小的点,将这三个点的几何中心确定为指示手掌的点。在预设的深度范围内,取出与所述指示手掌的点连通的所有点,其中获取连通的所有点可以通过漫水算法(flood fill)来实现,另外预设的深度范围本领域技术人员可以根据实际需求来选取(例如可以将预设的深度范围选取为(0,40cm)),在这里不做具体的限定。
可选的,从所述连通的点集中删除指示手臂的点集,将剩余的点集确定为指示手掌的点集。具体的,在所述连通的点集中通常会包含指示手臂的点集,应该将所述指示手臂的点集删除。如图9所示,在获取到所述连通的点集后,获取所述连通点集的最小外接矩形901,确定连通的点集中点到外接矩形指定边的距离,当距离不符合预设的距离要求的点确定为指示手臂的点,即将距离不符合预设的距离要去的点作为指示手臂点而删除,剩下的点集即为指示手掌的点集。
如图9所示,在一些实施例中,所述预设的距离要求可以通过最小外接矩形的边长来确定。具体的,取矩形的长边长为w,取矩形的短边长为h,计算指示手掌二维图像的点集中点到靠下的短边的距离,若d<w-1.2h,则将该点确定为指示手臂的点,按照这种方式,可以将所有指示手臂的点删除,剩余的点基于指示手掌的二维图像的点集。这里使用了一个假设,即是h表示手掌在二维图像上的宽度,根据手掌的宽度和长度之间长度比例关系,将1.2h确定为手掌的长度,w与1.2h之间的差应该是手臂上的点到靠下的短边的最大距离,如果二维图像的点集中的某个点到靠下的短边的距离小于或等于这个最大距离,即表示该点是属于手臂的点。其中d<w-1.2h只是根据最小外接矩形的边长确定预设的距离要求的一种实施方式,本领域技术人员还可以选取其他的方式,在这里不做具体的限定。
可选的,根据指示手掌的二维图像的点集的分布特征确定手势。
可选地,根据指示手掌的二维图像的点集确定分布区域,根据分布区域确定所述点集的分布特征。
可选的,使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定所述分布区域的分布特征。可选的,确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述距离确定为分布区域的分布特征。
可选的,当多边形的每一条边对应的最远距离都小于或等于预设的距离阈值时,确定手势为握拳,多边形的每一个对应的最远距离中至少有一个大于或等于预设的距离阈值时,确定手势为伸掌。
可选的,确定多帧深度图像中的每一帧深度图像对应的指示用户手掌二维图像的点集,根据所述每一帧图像对应的指示用户手掌的点集确定每一帧深度图像对应的指示手掌的点云,根据指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的动态手势。具体的,由于每一个点云的三维坐标都二维图像上的点的二维坐标是一一对应的,并且在手势识别的过程中会始终存储这两种坐标,在确定了指示手掌的二维图像的点集后,可以确定出指示手掌的点云。在获取到指示用户手掌的点云后,即可以根据前述部分的方法识别用户的手势。
可选的,根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌运动的运动方向,根据所述运动方向组成的序列确定手势。
可选的,确定所述运动方向中每一个运动方向所占的比例,根据所述比例的组合确定所述手掌运动的动作。
可选的,根据指示手掌的位置信息的序列,确定位置信息对应的 速度方向,根据所述速度方向的序列确定运动方向的序列。
可选的,确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
可选的,从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述速度方向对应的运动方向。
可选的,根据指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的动态手势,包括:根据指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的打勾手势。其中具体打勾手势的识别方法,请参见前述部分,在此不再赘述。
其中根据图10提供的识别方法中,获取到指示手掌的二维图像的点集,根据指示手掌的二维图像的点集确定用户的手势(静态手势)的所有解释都可以参照图2中的相关部分,另外根据指示手掌的二维图像的点集可以获取指示手掌的点云,根据指示手掌的点云确定用户的手势(动态手势)的所有解释都可以参照图2中的相关部分,为了简洁,此处不再赘述。
本发明的实施例提供了一种计算机存储介质,该计算机存储介质中存储有程序指令,该计算机存储介质中存储有程序指令,所述程序执行上述识别方法。
如图11所示,本发明实施例提供一种手势识别设备,所述设备1110包括:
TOF相机1110,用于获取用户的深度图像;
处理器1120,用于确定与所述深度图像对应的点云,对点云分类,从分类后的点云中确定指示手掌的点云,根据指示手掌的点云确定手势。
可选的,所述处理器1120,具体用于对点云分类获得点云的多个聚类,从多个聚类中的一个确定指示手掌的点云。
可选的,所述处理器1120,具体用于根据深度信息从多个聚类中确定指示手部的点云的聚类,从所述指示手部的点云的聚类中确定指示手掌的点云。
可选的,所述处理器1120,具体用于获取所述多个聚类中每一个聚类的平均深度,将平均深度最小的聚类确定为指示手部的点云的聚类。
可选的,所述处理器1120,具体用于从所述指示手部的点云的聚类中删除指示手臂的点云,将所述聚类中剩余的点云确定为指示手掌的点云。
可选的,所述处理器1120,具体用于在指示手部的点云的聚类中取出深度最小的点,确定聚类中的点云与深度最小点的距离,将距离大于或等于距离阈值的点确定为指示手臂的点云,删除所述指示手臂的点云。
可选的,所述处理器1120,具体用于根据指示手部的点云的聚类确定指示手部的二维图像的点集,确定所述手部的二维图像的点集的最小外接矩形;
所述处理器1120,具体用于确定所述手部的二维图像的点集中 的点到所述最小外接矩形的指定边的距离,当距离不符合预设的距离要求的点确定为指示手臂的二维图像的点;
所述处理器1120,具体用于确定指示手臂的二维图像的点集,根据指示手臂的二维图像的点集确定指示手臂的点云,将指示手臂的点云删除。
可选的,所述预设的距离要求可以通过最小外接矩形的边长来确定。
可选的,所述处理器1120,具体用于根据指示用户手掌的点云获取手掌的二维图像的点集,确定所述点集的分布特征,根据所述分布特征确定手势。
可选的,所述处理器1120,具体用于确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
可选的,所述处理器1120,具体用于使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
可选的,所述处理器1120,具体用于使用边数最少的凸多边形区域覆盖所述分布区域。
可选的,所述处理器1120,具体用于确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
可选的,所述处理器1120,具体用于当每一个不重叠区域对应的所述最远距离都小于或等于距离阈值时,确定手势为握拳。
可选的,所述处理器1120,具体用于当不重叠区域对应的所述最远距离中的一个或多个大于或等于距离阈值时,确定手势为伸掌。
可选的,所述处理器1120,具体用于使用聚类算法对点云分类。
可选的,在使用聚类算法对点云分类的过程中,聚类算法的聚类类数是可调整的。
可选的,所述处理器1120,具体用于根据聚类之间的分散程度来调整聚类类数。
可选的,所述TOF相机1110,用于获取多帧用户的深度图像,
所述处理器1120,具体用于确定所述多帧深度图像中每一帧对应的点云;
所述处理器1120,具体用于对所述每一帧深度图像对应的点云分类,从分类后的点云中确定每一帧深度图像对应的指示手掌的点云;
所述处理器1120,具体用于根据所述每一帧图像对应的指示用户手掌的点云确定每一帧深度图像对应的手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
可选的,所述处理器1120,具体用于根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
可选的,所述处理器1120,具体用于确定所述运动方向组成的序列中每一个运动方向所占的比例,根据所述比例的组合确定手势。
可选的,所述处理器1120,具体用于将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
可选的,所述处理器1120,具体用于根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
可选的,所述处理器1120,具体用于确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
可选的,所述处理器1120,具体用于从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述速度方向对应的运动方向。
具体实现中,本发明实施例中所描述的识别设备可执行本发明实施例图2提供的识别方法,其中具体解释可以参见图2提供的识别方法中的相应部分,在此不再赘述;另外本发明实施例中所描述的识别设备可以对本发明实施例图2提供的识别方法中技术特征进行引用与结合。
如图11所示,本发明实施例提供另一种手势识别设备,所述设备1110包括:
TOF相机1110,用于获取用户的深度图像;
处理器1120,用于根据所述深度信息确定指示手掌的二维图像的点集,根据所述点集来确定手势。
可选的,所述处理器1120,具体用于根据深度信息在二维图像上确定一个指示手掌的点,根据在预设的深度范围内确定与指示手掌的点连通的点集,从所述连通的点集中确定指示手掌的二维图像的点集。
可选的,所述处理器1120,具体用于从所述连通的点集中删除指 示手臂的点集,将剩余的点集确定为指示手掌的点集。
可选的,所述处理器1120,具体用于获取所述连通点集的最小外接矩形,确定连通的点集中的点到外接矩形指定边的距离,将距离不符合预设的距离要求的点确定为指示手臂的点,将指示手臂的点删除,将剩余的点集确定为指示手掌的点集。
可选的,所述预设的距离要求可以通过最小外接矩形的边长来确定。
可选的,所述处理器1120,具体用于确定所述点集的分布特征,根据所述分布特征确定手势。
可选的,所述处理器1120,具体用于确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
可选的,所述处理器1120,具体用于使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
可选的,所述处理器1120,具体用于使用边数最少的凸多边形区域覆盖所述分布区域。
可选的,所述处理器1120,具体用于确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
可选的,所述处理器1120,具体用于当每一个不重叠区域对应的所述距离都小于或等于距离阈值时,确定静态手势为握拳。
可选的,所述处理器1120,具体用于当不重叠区域对应的所述距 离中的一个或多个大于或等于距离阈值时,确定静态手势为伸掌。
可选的,所述TOF相机1110,用于获取多帧用户的深度图像;
所述处理器1120,具体用于确定所述多帧深度图像中每一帧对应的指示手掌的二维图像的点集;根据每一帧对应的指示手掌的二维图像的点集确定每一帧对应的指示手掌的点云,根据所述指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
可选的,所述处理器1120,具体用于根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
可选的,所述处理器1120,具体用于确定所述运动方向组成的序列中每一个运动方向所占的比例,根据所述比例的组合确定手势。
可选的,所述处理器1120,具体用于将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
可选的,所述处理器1120,具体用于根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
可选的,所述处理器1120,具体用于确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
可选的,所述处理器1120,具体用于从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述 速度方向对应的运动方向。
具体实现中,本发明实施例中所描述的识别设备可执行本发明实施例图2提供的识别方法,其中具体解释可以参见图10提供的识别方法中的相应部分,在此不再赘述;另外本发明实施例中所描述的识别设备可以对本发明实施例图10提供的识别方法中技术特征进行引用与结合。
如图12所示,本发明实施例还提供一种可移动平台,所述可移动平台1200包括:
如前所述的识别设备1100,用于识别用户的手势;
处理器1210,根据所述识别设备1100识别的手势生成相应的控制指令,并根据所述控制指令控制可移动平台1200。
其中可移动平台1200可以包括无人飞行器、地面机器人、遥控车等,如图12所示,可移动平台1200以无人飞行器作为可移动平台来进行示意性说明,其中下述部分中提到的无人飞行器均可以使用可移动平台替代。所述识别设备1100安装无人飞行器的合适位置上,其中,所述识别设备可以挂载在无人飞行器的机体外,也可以内置在无人飞行器的机体内,在这里不做具体限定,例如安装在无人飞行器的机头部分,识别设备1100对探测范围内的物体进行探测,捕获用户的手掌,识别用户的手势,每一种手势都对应不同的控制指令,处理器1210可以根据生成的控制指令来控制无人飞行器。其中,无人飞行器还可以包括云台1220以及成像设备1230,成像设备1230通过云台1220搭载于无人飞行器的主体上上。成像设备1230用于在无人飞行器的飞行过程中进行图像或视频拍摄,包括但不限于多光谱成像仪、高光谱成像仪、 可见光相机及红外相机等,云台1220为多轴传动及增稳系统,云台电机通过调整转动轴的转动角度来对成像设备1230的拍摄角度进行补偿,并通过设置适当的缓冲机构来防止或减小成像设备1230的抖动。为了便于说明,将能生成所述控制指令的手势称为命令手势。
本发明实施例提供的可移动平台能够识别用户的手势,并根据用户的手势来生成相应的控制指令,实现对可移动平台的控制。用户可以通过手势控制可移动平台,进一步丰富了可移动平台的控制方式,降低了对用户的专业性要求,提高了操作可移动平台的趣味性。
可选的,所述处理器1210,还用于在所述识别设备1100识别到用户的手势后,按照预设的控制模式点亮可移动平台的指示灯。具体的,在无人飞行器上的识别设备识别到用户的手势后,无人飞行器上的指示灯可以按照预设的控制模式来点亮,例如,在成功识别到手势后,无人飞行器上的左航灯和右航灯慢闪,这样用户通过观察左航灯和右航灯的闪灯情况,就可以知道自己所做的手势有没有被识别,避免了用户不清楚自己的动作有没有被识别,反复多次地去做同一个手势。当未成功识别用户的手势时,所述识别设备继续检测用户的手掌并去识别。
可选的,所述识别设备1100,用于在点亮可移动平台的指示灯后识别用户的确认手势;所述处理器1210,用于在所述识别设备1100识别到确认手势后,按照所述控制指令控制可移动平台。具体的,用户通过观察无人飞行器的指示灯闪烁,就可以知道自己的手势已经被识别,为了防止误触发,用户需要对之前的所做的手势进行确认。这样 用户在看到无人飞行器的指示灯闪烁后,做出确认手势,无人飞行器上的识别设备成功识别到用户的确认手势后,处理器根据之前的命令手势生成控制指令,根据控制指令控制无人飞行器。若在预设时间内,所述识别设备未识别到确认手势,所述识别设备返回检测探测范围内的手掌,去识别用户的其他的命令手势。
可选的,所述可移动平台还包括通讯接口1240,用于接收停止手势识别的指令,所述处理器1210,在通讯接口1240接收到所述停止手势识别的指令时,控制所述识别设备1100停止识别用户手势。图13为本发明一种实施例中可移动平台与控制终端通信的示意图,如图13所示,具体的,用户可以通过控制终端1300向所述可移动平台发出控制指令,所述控制指令用于使无人飞行器退出手势控制模式,此时,所述识别设备1100不再对用户的手势进行识别。另外所述通讯接口1240,还用于接收开始手势识别的指令,所述处理器1210,在通讯接口1240接收到所述开始手势识别的指令时,控制所述识别设备1100开始识别用户手势。
其中,本说明书中的存储器可以包括易失性存储器(volatile memory),例如随机存取存储器(random-access memory,RAM);所述存储器也可以包括非易失性存储器(non-volatile memory),例如快闪存储器(flash memory),硬盘(hard disk drive,HDD)或固态硬盘(solid-state drive,SSD)等。
所述处理器可以是中央处理器(central processing unit,CPU)。所述处理器还可以进一步包括硬件芯片。上述硬件芯片可以是专用集成电路(application-specific integrated circuit,ASIC),可编程逻辑器件(programmable logic device,PLD)或其组合。上述PLD可以是 复杂可编程逻辑器件(complex programmable logic device,CPLD),现场可编程逻辑门阵列(field-programmable gate array,FPGA)等。
本说明书中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其它实施例的不同之处,各个实施例之间相同或相似部分互相参见即可。对于实施例公开的设备而言,由于其与实施例公开的方法相对应,所以描述的比较简单,相关之处参见方法部分说明即可。
专业人员还可以进一步意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、计算机软件或者二者的结合来实现,为了清楚地说明硬件和软件的可互换性,在上述说明中已经按照功能一般性地描述了各示例的组成及步骤。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本发明的范围。
结合本文中所公开的实施例描述的方法或算法的步骤可以直接用硬件、处理器执行的软件模块,或者二者的结合来实施。软件模块可以置于随机存储器(RAM)、内存、只读存储器(ROM)、电可编程ROM、电可擦除可编程ROM、寄存器、硬盘、可移动磁盘、CD-ROM、或技术领域内所公知的任意其它形式的存储介质中。
以上对本发明所提供的识别的方法、设备和可移动平台以及存储介质进行了详细介绍。本文中应用了具体个例对本发明的原理及实施方式进行了阐述,以上实施例的说明只是用于帮助理解本发明的方法及其核心思想。应当指出,对于本技术领域的普通技术人员来说,在不脱离本发明原理的前提下,还可以对本发明进行若干改进和修饰,这些改进和修饰也落入本发明权利要求的保护范围。

Claims (93)

  1. 一种手势识别方法,所述方法包括:
    获取用户的深度图像;
    确定与所述深度图像对应的点云;
    对点云分类,从分类后的点云中确定指示手掌的点云;
    根据指示手掌的点云确定手势。
  2. 根据权利要求1所述的方法,其特征在于:
    所述对点云分类,从分类后的点云中确定指示手掌的点云包括:
    对点云分类获得点云的多个聚类,从多个聚类中的一个确定指示手掌的点云。
  3. 根据权利要求2所述的方法,其特征在于:
    所述从多个聚类中的一个确定指示手掌的点云包括:
    根据深度信息从多个聚类中确定指示手部的点云的聚类,从所述指示手部的点云的聚类中确定指示手掌的点云。
  4. 根据权利要求3所述的方法,其特征在于:
    所述根据深度信息从多个聚类中确定指示手部的点云的聚类包括:
    获取所述多个聚类中每一个聚类的平均深度,将平均深度最小的聚类确定为指示手部的点云的聚类。
  5. 根据权利要求3或4所述的方法,其特征在于:
    所述从所述指示手部的点云的聚类中确定指示手掌的点云包括:
    从所述指示手部的点云的聚类中删除指示手臂的点云,将所述聚 类中剩余的点云确定为指示手掌的点云。
  6. 根据权利要求5所述的方法,其特征在于:
    所述从所述指示手部的点云的聚类中删除指示手臂的点云包括:
    在指示手部的点云的聚类中取出深度最小的点,确定聚类中的点云与深度最小点的距离,将距离大于或等于距离阈值的点确定为指示手臂的点云,删除所述指示手臂的点云。
  7. 根据权利要求5所述的方法,其特征在于:
    所述从所述指示手部的点云的聚类中删除指示手臂的点云包括:
    根据指示手部的点云的聚类确定指示手部的二维图像的点集,确定所述手部的二维图像的点集的最小外接矩形;
    确定所述手部的二维图像的点集中的点到所述最小外接矩形的指定边的距离,将距离不符合预设的距离要求的点确定为指示手臂的二维图像的点;
    确定指示手臂的二维图像的点集,根据指示手臂的二维图像的点集确定指示手臂的点云,将指示手臂的点云删除。
  8. 根据权利要求7所述的方法,其特征在于:
    所述预设的距离要求可以通过最小外接矩形的边长来确定。
  9. 根据权利要求1-8任一项所述的方法,
    所述根据指示手掌的点云来确定手势包括:
    根据指示用户手掌的点云获取手掌的二维图像的点集,确定所述点集的分布特征,根据所述分布特征确定手势。
  10. 根据权利要求9所述的方法,其特征在于,
    所述确定所述点集的分布特征包括:
    确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
  11. 根据权利要求10所述的方法,其特征在于,
    所述根据分布区域确定所述点集的分布特征包括:
    使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
  12. 根据权利要求11所述的方法,其特征在于,
    所述使用多边形区域覆盖所述分布区域包括:
    使用边数最少的凸多边形区域覆盖所述分布区域。
  13. 根据权利要求12所述的方法,其特征在于,
    所述根据所述不重叠区域确定所述点集的分布特征包括:
    确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
  14. 根据权利要求13所述的方法,其特征在于,
    所述根据不重叠区域确定点集的分布特征包括:
    当每一个不重叠区域对应的所述最远距离都小于或等于距离阈值时,确定手势为握拳。
  15. 根据权利要求13所述的方法,其特征在于,
    所述不重叠区域确定点集的分布特征包括:
    当不重叠区域对应的所述最远距离中的一个或多个大于或等于距 离阈值时,确定手势为伸掌。
  16. 根据权利要求1-15任一项所述的方法,其特征在于,
    所述对所述点云进行分类包括:
    使用聚类算法对点云分类。
  17. 根据权利要16所述的方法,其特征在于,
    在使用聚类算法对点云分类的过程中,聚类算法的聚类类数是可调整的。
  18. 根据权利要求17所述的方法,其特征在于,
    根据聚类之间的分散程度来调整聚类类数。
  19. 根据权利要求1-18任一项所述的方法,其特征在于,
    所述获取用户的深度图像,确定与所述深度图像对应的点云包括;
    获取多帧用户的深度图像,确定所述多帧深度图像中每一帧对应的点云;
    所述对点云分类,从分类后的点云中确定指示手掌的点云包括;
    对所述每一帧深度图像对应的点云分类,从分类后的点云中确定每一帧深度图像对应的指示手掌的点云;
    所述根据指示手掌的点云来确定手势包括:
    根据所述每一帧图像对应的指示用户手掌的点云确定每一帧深度图像对应的手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
  20. 根据权利要求19所述的方法,其特征在于,
    所述根据所述位置信息组成的序列确定手掌的手势包括:
    根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
  21. 根据权利要求20所述的方法,其特征在于,
    所述根据所述运动方向组成的序列确定手势包括:
    确定所述运动方向组成的序列中每一个运动方向所占的比例,根据所述比例的组合确定手势。
  22. 根据权利要求21所述的方法,其特征在于,
    所述根据所述比例的组合确定手势包括:
    将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
  23. 根据权利要求20-22任一项所述的方法,其特征在于,
    所述确定所述序列中的位置信息对应的手掌运动的运动方向包括:
    根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
  24. 根据权利要求24所述的方法,其特征在于:
    所述根据所述速度方向确定手掌的运动方向包括:
    确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
  25. 根据权利要求24所述的方法,其特征在于,
    所述根据所述夹角确定所述运动方向包括
    从预设方向中确定与所述速度方向夹角最小的第一预设方向,将 所述第一预设方向确定为与所述速度方向对应的运动方向。
  26. 一种识别方法,包括:
    获取用户的深度图像,根据所述深度信息确定指示手掌的二维图像的点集;
    根据所述点集来确定手势。
  27. 根据权利要求26所述的方法,其特征在于:
    所述根据所述深度图像确定指示手掌的二维图像的点集包括:
    根据深度信息在二维图像上确定一个指示手掌的点,根据在预设的深度范围内确定与指示手掌的点连通的点集,从所述连通的点集中确定指示手掌的二维图像的点集。
  28. 根据权利要求27所述的方法,其特征在于:
    所述从所述连通的点集中确定指示手掌的二维图像的点集包括:
    从所述连通的点集中删除指示手臂的点集,将剩余的点集确定为指示手掌的点集。
  29. 根据权利要求28所述的方法,其特征在于:
    所述从所述连通的点集中删除指示手臂的点集,将剩余的点集确定为指示手掌的点集包括:
    获取所述连通点集的最小外接矩形,确定连通的点集中的点到外接矩形指定边的距离,将距离不符合预设的距离要求的点确定为指示手臂的点,将指示手臂的点删除,将剩余的点集确定为指示手掌的点集。
  30. 根据权利要求29所述的方法,其特征在于:
    所述预设的距离要求可以通过最小外接矩形的边长来确定。
  31. 根据权利要求26-30任一项所述的方法,
    所述根据所述点集来确定手势包括:
    确定所述点集的分布特征,根据所述分布特征确定手势。
  32. 根据权利要求31所述的方法,其特征在于,
    所述确定所述点集的分布特征包括:
    确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
  33. 根据权利要求32所述的方法,其特征在于,
    所述根据分布区域确定所述点集的分布特征包括:
    使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
  34. 根据权利要求33所述的方法,其特征在于,
    所述使用多边形区域覆盖所述分布区域包括:
    使用边数最少的凸多边形区域覆盖所述分布区域。
  35. 根据权利要求33或34所述的方法,其特征在于,
    所述根据所述不重叠区域确定所述点集的分布特征包括:
    确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
  36. 根据权利要求35所述的方法,其特征在于,
    所述根据不重叠区域确定点集的分布特征包括:
    当每一个不重叠区域对应的所述距离都小于或等于距离阈值时,确定静态手势为握拳。
  37. 根据权利要求35所述的方法,其特征在于,
    所述不重叠区域确定点集的分布特征包括:
    当不重叠区域对应的所述距离中的一个或多个大于或等于距离阈值时,确定静态手势为伸掌。
  38. 根据权利要求26-37任一项所述的方法,其特征在于,
    所述获取用户的深度图像,根据所述深度图像确定指示手掌的二维图像的点集包括;
    获取多帧用户的深度图像,确定所述多帧深度图像中每一帧对应的指示手掌的二维图像的点集;
    所述根据所述点集来确定手势包括:
    根据每一帧对应的指示手掌的二维图像的点集确定每一帧对应的指示手掌的点云,根据所述指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
  39. 根据权利要求38所述的方法,其特征在于,
    所述根据所述位置信息组成的序列确定手掌的手势包括:
    根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
  40. 根据权利要求39所述的方法,其特征在于,
    所述根据所述运动方向组成的序列确定手势包括:
    确定所述运动方向组成的序列中每一个运动方向所占的比例,根 据所述比例的组合确定手势。
  41. 根据权利要求40所述的方法,其特征在于,
    所述根据所述比例的组合确定手势包括:
    将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
  42. 根据权利要求39-41任一项所述的方法,其特征在于,
    所述确定所述序列中的位置信息对应的手掌运动的运动方向包括:
    根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
  43. 根据权利要求42所述的方法,其特征在于:
    所述根据所述速度方向确定手掌的运动方向包括:
    确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
  44. 根据权利要求43所述的方法,其特征在于,
    所述根据所述夹角确定所述运动方向包括
    从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述速度方向对应的运动方向。
  45. 一种手势识别设备,所述设备包括:
    TOF相机,用于获取用户的深度图像;
    处理器,用于确定与所述深度图像对应的点云,对点云分类,从分类后的点云中确定指示手掌的点云,根据指示手掌的点云确定手势。
  46. 根据权利要求45所述的设备,其特征在于:
    所述处理器,具体用于对点云分类获得点云的多个聚类,从多个聚类中的一个确定指示手掌的点云。
  47. 根据权利要求46所述的设备,其特征在于:
    所述处理器,具体用于根据深度信息从多个聚类中确定指示手部的点云的聚类,从所述指示手部的点云的聚类中确定指示手掌的点云。
  48. 根据权利要求47所述的设备,其特征在于:
    所述处理器,具体用于获取所述多个聚类中每一个聚类的平均深度,将平均深度最小的聚类确定为指示手部的点云的聚类。
  49. 根据权利要求47或48所述的设备,其特征在于:
    所述处理器,具体用于从所述指示手部的点云的聚类中删除指示手臂的点云,将所述聚类中剩余的点云确定为指示手掌的点云。
  50. 根据权利要求49所述的设备,其特征在于:
    所述处理器,具体用于在指示手部的点云的聚类中取出深度最小的点,确定聚类中的点云与深度最小点的距离,将距离大于或等于距离阈值的点确定为指示手臂的点云,删除所述指示手臂的点云。
  51. 根据权利要求49所述的设备,其特征在于:
    所述处理器,具体用于根据指示手部的点云的聚类确定指示手部的二维图像的点集,确定所述手部的二维图像的点集的最小外接矩形;
    所述处理器,具体用于确定所述手部的二维图像的点集中的点到所述最小外接矩形的指定边的距离,当距离不符合预设的距离要求的点确定为指示手臂的二维图像的点;
    所述处理器,具体用于确定指示手臂的二维图像的点集,根据指示手臂的二维图像的点集确定指示手臂的点云,将指示手臂的点云删除。
  52. 根据权利要求51所述的设备,其特征在于:
    所述预设的距离要求可以通过最小外接矩形的边长来确定。
  53. 根据权利要求45-52任一项所述的设备,
    所述处理器,具体用于根据指示用户手掌的点云获取手掌的二维图像的点集,确定所述点集的分布特征,根据所述分布特征确定手势。
  54. 根据权利要求53所述的设备,其特征在于,
    所述处理器,具体用于确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
  55. 根据权利要求54所述的设备,其特征在于,
    所述处理器,具体用于使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
  56. 根据权利要求55所述的设备,其特征在于,
    所述处理器,具体用于使用边数最少的凸多边形区域覆盖所述分布区域。
  57. 根据权利要求56所述的设备,其特征在于,
    所述处理器,具体用于确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
  58. 根据权利要求57所述的设备,其特征在于,
    所述处理器,具体用于当每一个不重叠区域对应的所述最远距离都小于或等于距离阈值时,确定手势为握拳。
  59. 根据权利要求57所述的设备,其特征在于,
    所述处理器,具体用于当不重叠区域对应的所述最远距离中的一个或多个大于或等于距离阈值时,确定手势为伸掌。
  60. 根据权利要求45-59任一项所述的设备,其特征在于,
    所述处理器,具体用于使用聚类算法对点云分类。
  61. 根据权利要60所述的设备,其特征在于,
    在使用聚类算法对点云分类的过程中,聚类算法的聚类类数是可调整的。
  62. 根据权利要求61所述的设备,其特征在于,
    所述处理器,具体用于根据聚类之间的分散程度来调整聚类类数。
  63. 根据权利要求45-62任一项所述的设备,其特征在于,
    所述TOF相机,用于获取多帧用户的深度图像,
    所述处理器,具体用于确定所述多帧深度图像中每一帧对应的点云;
    所述处理器,具体用于对所述每一帧深度图像对应的点云分类,从分类后的点云中确定每一帧深度图像对应的指示手掌的点云;
    所述处理器,具体用于根据所述每一帧图像对应的指示用户手掌的点云确定每一帧深度图像对应的手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
  64. 根据权利要求63所述的设备,其特征在于,
    所述处理器,具体用于根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
  65. 根据权利要求64所述的设备,其特征在于,
    所述处理器,具体用于确定所述运动方向组成的序列中每一个运动方向所占的比例,根据所述比例的组合确定手势。
  66. 根据权利要求65所述的设备,其特征在于,
    所述处理器,具体用于将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
  67. 根据权利要求64-66任一项所述的设备,其特征在于,
    所述处理器,具体用于根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
  68. 根据权利要求67所述的设备,其特征在于:
    所述处理器,具体用于确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
  69. 根据权利要求68所述的设备,其特征在于,
    所述处理器,具体用于从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述速度方向对应的运动方向。
  70. 一种识别设备,包括:
    TOF相机,用于获取用户的深度图像;
    处理器,用于根据所述深度信息确定指示手掌的二维图像的点集,根据所述点集来确定手势。
  71. 根据权利要求70所述的设备,其特征在于:
    所述处理器,具体用于根据深度信息在二维图像上确定一个指示手掌的点,根据在预设的深度范围内确定与指示手掌的点连通的点集,从所述连通的点集中确定指示手掌的二维图像的点集。
  72. 根据权利要求71所述的设备,其特征在于:
    所述处理器,具体用于从所述连通的点集中删除指示手臂的点集,将剩余的点集确定为指示手掌的点集。
  73. 根据权利要求72所述的设备,其特征在于:
    所述处理器,具体用于获取所述连通点集的最小外接矩形,确定连通的点集中的点到外接矩形指定边的距离,将距离不符合预设的距离要求的点确定为指示手臂的点,将指示手臂的点删除,将剩余的点集确定为指示手掌的点集。
  74. 根据权利要求73所述的设备,其特征在于:
    所述预设的距离要求可以通过最小外接矩形的边长来确定。
  75. 根据权利要求70-74任一项所述的设备,
    所述处理器,具体用于确定所述点集的分布特征,根据所述分布特征确定手势。
  76. 根据权利要求75所述的设备,其特征在于,
    所述处理器,具体用于确定指示手掌的二维图像的点集的分布区域,根据分布区域确定所述点集的分布特征。
  77. 根据权利要求76所述的设备,其特征在于,
    所述处理器,具体用于使用多边形区域覆盖所述分布区域,确定所述多边形区域和所述分布区域之间的不重叠的区域,根据所述不重叠区域确定点集的分布特征。
  78. 根据权利要求77所述的设备,其特征在于,
    所述处理器,具体用于使用边数最少的凸多边形区域覆盖所述分布区域。
  79. 根据权利要求77或78所述的设备,其特征在于,
    所述处理器,具体用于确定所述不重叠区域中的点到对应的多边形的边的最远距离,将所述最远距离确定为点集的分布特征。
  80. 根据权利要求79所述的设备,其特征在于,
    所述处理器,具体用于当每一个不重叠区域对应的所述距离都小于或等于距离阈值时,确定静态手势为握拳。
  81. 根据权利要求79所述的设备,其特征在于,
    所述处理器,具体用于当不重叠区域对应的所述距离中的一个或多个大于或等于距离阈值时,确定静态手势为伸掌。
  82. 根据权利要求70-81任一项所述的设备,其特征在于,
    所述TOF相机,用于获取多帧用户的深度图像;
    所述处理器,具体用于确定所述多帧深度图像中每一帧对应的指示手掌的二维图像的点集;根据每一帧对应的指示手掌的二维图像的点集确定每一帧对应的指示手掌的点云,根据所述指示手掌的点云确定手掌的位置信息,根据所述位置信息组成的序列确定手掌的手势。
  83. 根据权利要求82所述的设备,其特征在于,
    所述处理器,具体用于根据指示手掌的位置信息的序列,确定所述序列中的位置信息对应的手掌的运动方向,根据所述运动方向组成的序列确定手势。
  84. 根据权利要求83所述的设备,其特征在于,
    所述处理器,具体用于确定所述运动方向组成的序列中每一个运动方向所占的比例,根据所述比例的组合确定手势。
  85. 根据权利要求84所述的设备,其特征在于,
    所述处理器,具体用于将所述比例的组合输入预设的运算模型,利用所述预设的运算模型来确定手势。
  86. 根据权利要求83-85任一项所述的设备,其特征在于,
    所述处理器,具体用于根据指示手掌的位置信息的序列确定位置信息对应的速度方向,根据所述速度方向确定手掌的运动方向;
  87. 根据权利要求86所述的设备,其特征在于:
    所述处理器,具体用于确定所述序列中的位置信息对应的速度方向,确定所述速度方向与多个预设方向中每一个的夹角,根据所述夹角确定所述运动方向。
  88. 根据权利要求87所述的设备,其特征在于,
    所述处理器,具体用于从预设方向中确定与所述速度方向夹角最小的第一预设方向,将所述第一预设方向确定为与所述速度方向对应的运动方向。
  89. 一种可移动平台,包括
    如前所述的45-69或70-88任一项所述的识别设备,用于识别用户的手势;
    处理器,用于根据所述识别设备识别的手势生成相应的控制指令,并根据所述控制指令控制可移动平台。
  90. 根据权利要求89所述的可移动平台,其特征在于,
    所述处理器,还用于在所述识别设备识别到用户的手势后,按照预设的控制模式点亮可移动平台的指示灯。
  91. 根据权利要求90所述的可移动平台,其特征在于,
    所述识别设备,用于在点亮可移动平台的指示灯后识别用户的确认手势;
    所述处理器,用于在所述识别设备识别到确认手势后,根据所述控制指令控制无人飞行器。
  92. 根据权利要求91所述的方法,其特征在于,
    所述确定手势为:打勾手势、手掌在预设的时间内保持近似静止、手掌在竖直平面内画圆中的至少一种。
  93. 根据权利要求89-92任一项所述的可移动平台,其特征在于,
    所述可移动平台还包括通讯接口,所述通讯接口,用于接收停止手势识别的指令;
    所述处理器,在通讯接口接收到所述停止手势识别的指令时,控制所述识别设备停止识别用户手势。
PCT/CN2017/075193 2017-02-28 2017-02-28 一种识别方法、设备以及可移动平台 Ceased WO2018157286A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
PCT/CN2017/075193 WO2018157286A1 (zh) 2017-02-28 2017-02-28 一种识别方法、设备以及可移动平台
CN201780052992.4A CN109643372A (zh) 2017-02-28 2017-02-28 一种识别方法、设备以及可移动平台
US16/553,680 US11250248B2 (en) 2017-02-28 2019-08-28 Recognition method and apparatus and mobile platform

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2017/075193 WO2018157286A1 (zh) 2017-02-28 2017-02-28 一种识别方法、设备以及可移动平台

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US16/553,680 Continuation US11250248B2 (en) 2017-02-28 2019-08-28 Recognition method and apparatus and mobile platform

Publications (1)

Publication Number Publication Date
WO2018157286A1 true WO2018157286A1 (zh) 2018-09-07

Family

ID=63369720

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2017/075193 Ceased WO2018157286A1 (zh) 2017-02-28 2017-02-28 一种识别方法、设备以及可移动平台

Country Status (3)

Country Link
US (1) US11250248B2 (zh)
CN (1) CN109643372A (zh)
WO (1) WO2018157286A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111191687A (zh) * 2019-12-14 2020-05-22 贵州电网有限责任公司 基于改进K-means算法的电力通信数据聚类方法
WO2020150961A1 (zh) * 2019-01-24 2020-07-30 深圳市大疆创新科技有限公司 探测装置、可移动平台
CN115271200A (zh) * 2022-07-25 2022-11-01 仲恺农业工程学院 一种用于名优茶的智能连贯采摘系统

Families Citing this family (17)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI710973B (zh) * 2018-08-10 2020-11-21 緯創資通股份有限公司 手勢識別方法、手勢識別模組及手勢識別系統
CN109145803B (zh) * 2018-08-14 2022-07-22 京东方科技集团股份有限公司 手势识别方法及装置、电子设备、计算机可读存储介质
CN110874179B (zh) * 2018-09-03 2021-09-14 京东方科技集团股份有限公司 指尖检测方法、指尖检测装置、指尖检测设备及介质
WO2020130855A1 (en) * 2018-12-21 2020-06-25 Chronoptics Limited Time of flight camera data processing system
CN113711273B (zh) * 2019-04-25 2024-09-17 三菱电机株式会社 移动量估计装置、移动量估计方法及计算机可读取的记录介质
CN110888536B (zh) * 2019-12-12 2023-04-28 北方工业大学 基于mems激光扫描的手指交互识别系统
JP6708917B1 (ja) * 2020-02-05 2020-06-10 リンクウィズ株式会社 形状検出方法、形状検出システム、プログラム
KR102346294B1 (ko) * 2020-03-03 2022-01-04 주식회사 브이터치 2차원 이미지로부터 사용자의 제스처를 추정하는 방법, 시스템 및 비일시성의 컴퓨터 판독 가능 기록 매체
CN111368747A (zh) * 2020-03-06 2020-07-03 上海掌腾信息科技有限公司 基于tof技术实现掌静脉特征矫正处理的系统及其方法
CN111695420B (zh) * 2020-04-30 2024-03-08 华为技术有限公司 一种手势识别方法以及相关装置
CN111624572B (zh) * 2020-05-26 2023-07-18 京东方科技集团股份有限公司 一种人体手部与人体手势识别的方法及装置
CN111881733B (zh) * 2020-06-17 2023-07-21 艾普工华科技(武汉)有限公司 一种工人作业工步规范视觉识别判定与指导方法和系统
WO2022009091A1 (en) * 2020-07-08 2022-01-13 Cron Systems Pvt. Ltd. System and method for classification of objects by a body taper detection
CN115273220A (zh) * 2022-06-14 2022-11-01 网易有道信息技术(北京)有限公司 基于手势的查词方法、查词设备及计算机可读存储介质
CN116665245B (zh) * 2023-05-12 2026-01-09 上海数迹智能科技有限公司 多视角手势识别方法、装置、计算机设备和存储介质
CN117523382B (zh) * 2023-07-19 2024-06-04 石河子大学 一种基于改进gru神经网络的异常轨迹检测方法
TWI870007B (zh) * 2023-09-05 2025-01-11 仁寶電腦工業股份有限公司 手勢判斷方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120062736A1 (en) * 2010-09-13 2012-03-15 Xiong Huaixin Hand and indicating-point positioning method and hand gesture determining method used in human-computer interaction system
WO2014009561A2 (en) * 2012-07-13 2014-01-16 Softkinetic Software Method and system for human-to-computer gesture based simultaneous interactions using singular points of interest on a hand
CN103984928A (zh) * 2014-05-20 2014-08-13 桂林电子科技大学 基于景深图像的手指手势识别方法
CN105138990A (zh) * 2015-08-27 2015-12-09 湖北师范学院 一种基于单目摄像头的手势凸包检测与掌心定位方法
CN105389539A (zh) * 2015-10-15 2016-03-09 电子科技大学 一种基于深度数据的三维手势姿态估计方法及系统

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
IL204436A (en) * 2010-03-11 2016-03-31 Deutsche Telekom Ag A system and method for remote control of online TV by waving hands
JP5701714B2 (ja) * 2011-08-05 2015-04-15 株式会社東芝 ジェスチャ認識装置、ジェスチャ認識方法およびジェスチャ認識プログラム
RU2014108820A (ru) * 2014-03-06 2015-09-20 ЭлЭсАй Корпорейшн Процессор изображений, содержащий систему распознавания жестов с функциональными возможностями обнаружения и отслеживания пальцев

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120062736A1 (en) * 2010-09-13 2012-03-15 Xiong Huaixin Hand and indicating-point positioning method and hand gesture determining method used in human-computer interaction system
WO2014009561A2 (en) * 2012-07-13 2014-01-16 Softkinetic Software Method and system for human-to-computer gesture based simultaneous interactions using singular points of interest on a hand
CN103984928A (zh) * 2014-05-20 2014-08-13 桂林电子科技大学 基于景深图像的手指手势识别方法
CN105138990A (zh) * 2015-08-27 2015-12-09 湖北师范学院 一种基于单目摄像头的手势凸包检测与掌心定位方法
CN105389539A (zh) * 2015-10-15 2016-03-09 电子科技大学 一种基于深度数据的三维手势姿态估计方法及系统

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020150961A1 (zh) * 2019-01-24 2020-07-30 深圳市大疆创新科技有限公司 探测装置、可移动平台
CN111191687A (zh) * 2019-12-14 2020-05-22 贵州电网有限责任公司 基于改进K-means算法的电力通信数据聚类方法
CN111191687B (zh) * 2019-12-14 2023-02-10 贵州电网有限责任公司 基于改进K-means算法的电力通信数据聚类方法
CN115271200A (zh) * 2022-07-25 2022-11-01 仲恺农业工程学院 一种用于名优茶的智能连贯采摘系统

Also Published As

Publication number Publication date
US11250248B2 (en) 2022-02-15
US20190392205A1 (en) 2019-12-26
CN109643372A (zh) 2019-04-16

Similar Documents

Publication Publication Date Title
WO2018157286A1 (zh) 一种识别方法、设备以及可移动平台
US11302026B2 (en) Attitude recognition method and device, and movable platform
US11669972B2 (en) Geometry-aware instance segmentation in stereo image capture processes
CN111344644B (zh) 用于基于运动的自动图像捕获的技术
US8265425B2 (en) Rectangular table detection using hybrid RGB and depth camera sensors
US9424649B1 (en) Moving body position estimation device and moving body position estimation method
CN107688391B (zh) 一种基于单目视觉的手势识别方法和装置
CN114637023B (zh) 用于激光深度图取样的系统及方法
US8933886B2 (en) Instruction input device, instruction input method, program, recording medium, and integrated circuit
JP6125188B2 (ja) 映像処理方法及び装置
US20230020725A1 (en) Information processing apparatus, information processing method, and program
CN109241820B (zh) 基于空间探索的无人机自主拍摄方法
CN112567201A (zh) 距离测量方法以及设备
WO2018049998A1 (zh) 交通标志牌信息获取方法及装置
CN106709895A (zh) 图像生成方法和设备
CN114359714B (zh) 基于事件相机的无人体避障方法、装置及智能无人体
JP6817742B2 (ja) 情報処理装置およびその制御方法
CN110609562A (zh) 一种图像信息采集方法和装置
US20230376106A1 (en) Depth information based pose determination for mobile platforms, and associated systems and methods
Martin et al. Real time driver body pose estimation for novel assistance systems
JP2018120283A (ja) 情報処理装置、情報処理方法及びプログラム
CN112949347B (zh) 基于人体姿势的风扇调节方法、风扇和存储介质
TW202029134A (zh) 行車偵測方法、車輛及行車處理裝置
US20240312223A1 (en) Image processing apparatus
US12277646B2 (en) Method for generating point cloud data and data generating apparatus

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17898940

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17898940

Country of ref document: EP

Kind code of ref document: A1