WO2017084319A1 - 手势识别方法及虚拟现实显示输出设备 - Google Patents

手势识别方法及虚拟现实显示输出设备 Download PDF

Info

Publication number
WO2017084319A1
WO2017084319A1 PCT/CN2016/085365 CN2016085365W WO2017084319A1 WO 2017084319 A1 WO2017084319 A1 WO 2017084319A1 CN 2016085365 W CN2016085365 W CN 2016085365W WO 2017084319 A1 WO2017084319 A1 WO 2017084319A1
Authority
WO
WIPO (PCT)
Prior art keywords
gesture
information
spatial
video
plane
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/085365
Other languages
English (en)
French (fr)
Inventor
张超
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Le Holdings Beijing Co Ltd
Leshi Zhixin Electronic Technology Tianjin Co Ltd
Original Assignee
Le Holdings Beijing Co Ltd
Leshi Zhixin Electronic Technology Tianjin Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Le Holdings Beijing Co Ltd, Leshi Zhixin Electronic Technology Tianjin Co Ltd filed Critical Le Holdings Beijing Co Ltd
Priority to US15/240,571 priority Critical patent/US20170140215A1/en
Publication of WO2017084319A1 publication Critical patent/WO2017084319A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F2203/00Indexing scheme relating to G06F3/00 - G06F3/048
    • G06F2203/01Indexing scheme relating to G06F3/01
    • G06F2203/012Walk-in-place systems for allowing a user to walk in a virtual environment while constraining him to a given position in the physical environment

Definitions

  • the present invention relates to the field of virtual reality display related technologies, and in particular, to a gesture recognition method for a virtual reality display output device and a virtual reality display output device.
  • Virtual Reality (VR) technology is the use of computers or other intelligent computing devices as the core, combined with photoelectric sensing technology to generate a virtual environment within a specific range of realistic viewing, listening and touch integration.
  • the virtual reality system mainly includes an input device and an output device.
  • a typical virtual reality display output device is a Head Mount Display (HMD), which allows the user to create an independent closed immersive interactive experience with the interaction of the input device.
  • HMD of consumer products mainly has two kinds of product forms: a PC helmet display device that utilizes a personal computer (PC) computing capability access mode, and a portable helmet display device that is based on a mobile phone's computing processing capability.
  • PC personal computer
  • the main control is the handle, remote control, motion sensor and so on.
  • the scheme Based on the gesture recognition of a single common camera, the scheme has limited immersion because it can only recognize two-dimensional gestures;
  • An embodiment of the present invention provides a gesture recognition method for a virtual reality display output device, including:
  • the second plane gesture includes:
  • Separating a hand graphic from the first image of each frame in the first video Separating a hand graphic from the first image of each frame in the first video, acquiring first plane information of the hand graphic separated from the first image in each frame, and combining the plurality of first plane information into the a first plane gesture, using a timestamp of the first image corresponding to each of the first plane information as a timestamp of each of the first plane information, and a second image from each frame in the second video Separating the hand graphic, acquiring second plane information of the hand graphic separated from the second image of each frame, combining the plurality of second plane information into the second plane gesture, each of the second a timestamp of the second image corresponding to the plane information as a timestamp of each of the second plane information;
  • the method for converting the first plane information and the second plane information into spatial information by using a binocular imaging method, and generating a spatial gesture including the spatial information specifically includes:
  • the hand pattern is separated from the first image of each frame in the first video by hand detection and hand tracking, and separated from the second image of each frame in the second video by hand detection and hand tracking. Come out hand graphics.
  • first plane information includes first active part plane information of at least one active part of the hand graphic
  • second plane information includes second active part plane information of at least one active part of the hand graphic
  • the method for converting the first plane information and the second plane information into spatial information by using a binocular imaging method, and generating a spatial gesture including the spatial information specifically includes:
  • the acquiring the execution instruction corresponding to the space gesture specifically includes:
  • a gesture classification model to obtain a gesture type of the spatial gesture, and acquiring an execution instruction corresponding to the gesture type, where the gesture classification model is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • the type classification model for spatial gestures is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • Embodiments of the present invention provide a computer program comprising computer code adapted to perform all of the steps of the gesture recognition method as previously described when run on a computer.
  • the computer program is embodied on a computer readable medium.
  • An embodiment of the present invention provides a virtual reality display output device, including:
  • a video acquisition module configured to: acquire a first video from the first camera, and acquire a second video from the second camera;
  • a hand separation module configured to: separate a first plane gesture about first plane information of a hand graphic in the first video from the first video, and separate a hand graphic about the second video from the second video a second planar gesture of the second planar information;
  • a spatial information construction module configured to: use the binocular imaging method to adopt the first plane Converting the information and the second plane information into spatial information, generating a spatial gesture including the spatial information;
  • An instruction obtaining module configured to: acquire an execution instruction corresponding to the space gesture
  • An execution module is configured to execute the execution instruction.
  • the hand separation module is configured to: separate a hand graphic from a first image of each frame in the first video, and acquire first plane information of a hand graphic separated from the first image in each frame, Combining a plurality of first plane information into the first plane gesture, using a timestamp of the first image corresponding to each of the first plane information as a timestamp of each of the first plane information, Separating the hand graphic from the second image in each frame in the second video, acquiring second plane information of the hand graphic separated from the second image in each frame, and combining the plurality of second plane information into the second a plane gesture, the timestamp of the second image corresponding to each of the second plane information is used as a timestamp of each of the second plane information;
  • the spatial information constructing module is specifically configured to: calculate the first plane information and the second plane information having the same time stamp as spatial information by using a binocular imaging manner, and generate a spatial gesture including the spatial information.
  • the hand pattern is separated from the first image of each frame in the first video by hand detection and hand tracking, and separated from the second image of each frame in the second video by hand detection and hand tracking. Come out hand graphics.
  • first plane information includes first active part plane information of at least one active part of the hand graphic
  • second plane information includes second active part plane information of at least one active part of the hand graphic
  • the spatial information construction module is specifically configured to: calculate, by using a binocular imaging method, the first active part plane information and the second active part plane information of the same active part with the same time stamp to calculate the active part space of the active part
  • the information generates a spatial gesture including at least one of the active part space information.
  • instruction acquisition module is specifically configured to:
  • the gesture classification model is A type classification model for spatial gestures obtained after training using machine learning using a plurality of pre-acquired spatial gestures.
  • the hand graphics are separated from the video acquired by the two cameras, and then merged by the binocular imaging method. Since the hand graphics are separated, the interference of the external environment is avoided, and no calculation of the hand graphics is required. In the background, it is only necessary to calculate the spatial information by using the binocular imaging method of the opponent graphic, which greatly reduces the calculation amount. Therefore, the spatial information of the hand graphic can be obtained with a very small calculation amount, so that the recognition of the three-dimensional gesture can be completed by using the ordinary camera, which greatly reduces the cost and technical risk of the virtual reality display output device.
  • FIG. 1 is a flowchart of a gesture recognition method for a virtual reality display output device according to an embodiment of the present invention
  • FIG. 2 is a working flowchart of a gesture recognition method for a virtual reality display output device according to another embodiment of the present invention.
  • FIG. 3 is a structural block diagram of a virtual reality display output device according to an embodiment of the present invention.
  • FIG. 4 is a schematic structural diagram of a virtual reality display output device according to an embodiment of the present invention.
  • FIG. 1 is a flowchart of a gesture recognition method for a virtual reality display output device according to an embodiment of the present invention, including:
  • Step S101 comprising: acquiring a first video from the first camera, and acquiring a second video from the second camera;
  • Step S102 comprising: separating a first plane gesture about first plane information of a hand graphic in the first video from the first video, and separating a second graphic about a hand graphic in the second video from the second video a second plane gesture of the second plane information;
  • Step S103 comprising: using the binocular imaging method to set the first plane information and the Converting the second plane information into spatial information, and generating a spatial gesture including the spatial information;
  • Step S104 comprising: acquiring an execution instruction corresponding to the space gesture
  • Step S105 comprising: executing the execution instruction.
  • the user makes a gesture before the virtual reality display output device, and the gesture forms a hand graphic in the first video and the second video obtained from the two ordinary cameras in step S101, and then separates the first plane information and the first in step S102.
  • Two plane information refer to a planar position of the hand graphic in the first video and a planar position in the second video. Since a single camera can only obtain the plane position, if you want to obtain the three-dimensional position, you need to perform binocular imaging.
  • the main function of binocular imaging is binocular ranging, which mainly uses the target point. Here is the hand, two in the left and right.
  • the difference between the lateral coordinates of the image on the view (ie, the parallax) and the distance Z from the target point to the imaging plane are inversely proportional.
  • the target point (ie, the hand) and the camera are calculated by the parallax caused by the spacing of the two cameras. The distance, thereby determining the position of the target point (ie, the hand) in space as spatial information.
  • step S104 is executed to acquire the execution instruction and the corresponding instruction is executed in step S105.
  • the user can interact with the virtual reality display output device through gestures.
  • the embodiment of the present invention first performs step S102, and separates the video of each camera as a separate video. After the separation, step S103 is performed to perform binocular imaging, thereby avoiding background of binocular imaging. Interference, greatly reducing the amount of calculation, enabling the use of a common camera to complete the recognition of three-dimensional gestures, greatly reducing the cost and technical risk of virtual reality display output devices.
  • the step S102 includes: separating a hand graphic from the first image of each frame in the first video, and acquiring first plane information of the hand graphic separated from the first image in each frame, and multiple Combining the first plane information into the first plane gesture, a time stamp of the first image corresponding to each of the first plane information is used as a time stamp of each of the first plane information, and a hand graphic is separated from each second image in the second video.
  • the step S103 includes: calculating, by using a binocular imaging manner, the first plane information and the second plane information having the same time stamp as spatial information, and generating a spatial gesture including the spatial information.
  • the hand graphic is separated from each frame image, and a corresponding time stamp is established, and then the first plane information and the second plane information having the same time stamp are converted into spatial information in step S103, so that the space The calculation of information is more accurate.
  • the hand pattern is separated from the first image of each frame in the first video by hand detection and hand tracking, and the second frame is used from the second video by hand detection and hand tracking.
  • the hand graphic is separated from the image.
  • the hand detection employed in this embodiment includes: detection based on skin color, detection based on motion information, detection based on features, target detection based on image segmentation, and the like.
  • Hand tracking includes: tracking algorithms such as particle tracking and CamShift algorithm, and can also be combined with Kalman filtering to achieve better results.
  • the separation of the hand graphics is more accurate by the hand detection and the hand tracking manner, so that the subsequent spatial information calculation is more accurate, and a more accurate spatial gesture is recognized.
  • the first plane information includes first active part plane information of at least one active part of the hand graphic
  • the second plane information includes second active part plane information of at least one active part of the hand graphic
  • the step S103 includes: calculating, by using a binocular imaging method, the first active part plane information and the second active part plane information of the same active part with the same time stamp to calculate the active part spatial information of the active part, and generating A spatial gesture comprising at least one of the active part spatial information.
  • the active part refers to a movable part of a person's hand, such as a finger. Active part It can be pre-specified that since the hand graphic has been separated, there is no interference from other backgrounds, and the active part is generally at the edge of the hand graphic, so it can be easily identified by edge feature extraction or the like.
  • the embodiment further calculates the active part spatial information of the active part, so that more detailed gestures can be identified.
  • the step S104 specifically includes:
  • a gesture classification model to obtain a gesture type of the spatial gesture, and acquiring an execution instruction corresponding to the gesture type, where the gesture classification model is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • the type classification model for spatial gestures is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • the input of the gesture classification model is a spatial gesture
  • the output is a gesture type.
  • Machine learning can be supervised. For example, when there is supervised training, each type of spatial gesture used for training is specified. After multiple trainings, a gesture classification model is obtained. It can also be unsupervised, such as type categorization, such as using the k-Nearest Neighbor algorithm (KNN), which classifies spatial gestures for training based on their spatial position.
  • KNN k-Nearest Neighbor algorithm
  • a gesture classification model is established by using a machine learning manner, which facilitates classifying gestures, thereby increasing the robustness of gesture recognition.
  • FIG. 2 is a flowchart of a gesture recognition method for a virtual reality display output device according to another embodiment of the present invention
  • Step S201 separately using two common cameras to separately collect image data
  • the user makes a gesture before the virtual reality display output device, and the gesture forms a hand graphic in the first video and the second video obtained by the two ordinary cameras;
  • Step S202 performing hand detection and tracking on the data collected by the two cameras
  • tracking there are many methods that can be used for detection, such as skin color based detection, motion information based detection, feature based detection, image segmentation based target detection, and the like.
  • tracking algorithms such as particle tracking and CamShift algorithm can be used, and can also be combined with Kalman filtering to achieve better results;
  • Step S203 using the principle of binocular imaging for the hand that has been detected and tracked, Help the parallax caused by the distance between the two cameras to get the distance from the hand to the camera;
  • the difference between the lateral coordinates imaged on the left and right views (ie, the parallax) and the distance Z from the target point to the imaging plane are inversely proportional.
  • the target point is calculated by the parallax caused by the spacing of the two cameras (ie, Hand) the distance from the camera;
  • Step S204 the information of the hand obtained at this time includes both the color information and the position information in the space, and the hand type recognition and the gesture recognition in the three-dimensional sense can be performed at this time;
  • Step S205 the recognized gesture, the message or event that drives the response interacts with the VR system.
  • FIG. 3 is a structural block diagram of a virtual reality display output device according to an embodiment of the present invention, including:
  • the video acquisition module 301 is configured to: acquire a first video from the first camera, and acquire a second video from the second camera;
  • the hand separation module 302 is configured to: separate a first plane gesture about the first plane information of the hand graphic in the first video from the first video, and separate the second video from the second video a second planar gesture of the second planar information of the graphic;
  • the spatial information constructing module 303 is configured to: convert the first plane information and the second plane information into spatial information by using a binocular imaging manner, and generate a spatial gesture including the spatial information;
  • the instruction obtaining module 304 is configured to: acquire an execution instruction corresponding to the space gesture;
  • the executing module 305 is configured to execute the execution instruction.
  • the embodiment of the invention enables the recognition of the three-dimensional gesture by using a common camera, which greatly reduces the cost and technical risk of the virtual reality display output device.
  • the hand separation module 302 is configured to: separate a hand graphic from a first image of each frame in the first video, and acquire first plane information of a hand graphic separated from the first image in each frame, Combining a plurality of first plane information into the first plane gesture, using a timestamp of the first image corresponding to each of the first plane information as a timestamp of each of the first plane information, from Separating the hand graphic from the second image in each frame of the second video, and acquiring the hand graphic separated from the second image in each frame Second plane information, combining a plurality of second plane information into the second plane gesture, and using a timestamp of the second image corresponding to each of the second plane information as each of the second planes Timestamp of the information;
  • the spatial information constructing module 303 is specifically configured to: calculate, by using a binocular imaging manner, the first plane information and the second plane information having the same time stamp as spatial information, and generate a spatial gesture including the spatial information. .
  • This embodiment makes the calculation of spatial information more accurate.
  • the hand pattern is separated from the first image of each frame in the first video by hand detection and hand tracking, and the second frame is used from the second video by hand detection and hand tracking.
  • the hand graphic is separated from the image.
  • the separation of the hand graphics is more accurate by the hand detection and the hand tracking manner, so that the subsequent spatial information calculation is more accurate, and a more accurate spatial gesture is recognized.
  • the first plane information includes first active part plane information of at least one active part of the hand graphic
  • the second plane information includes second active part plane information of at least one active part of the hand graphic
  • the spatial information construction module is specifically configured to: calculate, by using a binocular imaging method, the first active part plane information and the second active part plane information of the same active part with the same time stamp to calculate the active part space of the active part
  • the information generates a spatial gesture including at least one of the active part space information.
  • the embodiment further calculates the active part spatial information of the active part, so that more detailed gestures can be identified.
  • the instruction acquisition module 304 is specifically configured to:
  • a gesture classification model to obtain a gesture type of the spatial gesture, and acquiring an execution instruction corresponding to the gesture type, where the gesture classification model is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • the type classification model for spatial gestures is obtained by using a plurality of pre-acquired spatial gestures by using machine learning.
  • a gesture classification model is established by using a machine learning manner, which facilitates classifying gestures, thereby increasing the robustness of gesture recognition.
  • FIG. 4 is a schematic structural diagram of a virtual reality display output device according to an embodiment of the present invention.
  • the virtual reality display output device can be accessed by using a PC computing capability.
  • the PC helmet display device, or the portable helmet display device based on the computing processing capability of the mobile phone, or the helmet display device has its own computing processing capability, and mainly includes a processor 401, a memory 402, two cameras 403, and the like.
  • the specific code for storing the foregoing method in the memory 402 is specifically executed by the processor 401, and the gesture is captured by the camera 403, and processed by the processor 401 according to the foregoing method to perform a corresponding operation.
  • the logic instructions in the memory 402 described above may be implemented in the form of a software functional unit and sold or used as a stand-alone product, and may be stored in a computer readable storage medium.
  • the technical solution of the present invention which is essential or contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product, which is stored in a storage medium, including
  • the instructions are used to cause a mobile terminal (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.
  • the foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and the like. .
  • the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, ie may be located A place, or it can be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the objectives of the embodiments of the present invention. Those of ordinary skill in the art can understand and implement without deliberate labor.

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)
  • Image Analysis (AREA)

Abstract

一种用于虚拟现实显示输出设备的手势识别方法及虚拟现实显示输出设备,识别方法包括:从第一摄像头获取第一视频,从第二摄像头获取第二视频(S101);从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势(S102);采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势(S103);获取所述空间手势对应的执行指令(S104);执行所述执行指令(S105)。所述方法能够采用普通摄像头完成三维手势的识别,大大降低了虚拟现实显示输出设备的成本及技术风险。

Description

手势识别方法及虚拟现实显示输出设备
本申请要求2015年11月18日提交的申请号为201510796509.6的中国专利申请的优先权,其全部内容通过引用被合并于此。
技术领域
本发明涉及虚拟现实显示相关技术领域,特别是一种用于虚拟现实显示输出设备的手势识别方法及虚拟现实显示输出设备。
背景技术
虚拟现实(Virtual Reality,VR)技术就是利用电脑或其他智能计算设备为核心,结合光电传感技术生成逼真的视、听、触一体化的特定范围内的虚拟环境。虚拟现实系统中主要包括输入设备和输出设备。一种典型的虚拟现实显示输出设备为头戴式显示器(Head Mount Display,HMD),配合输入设备的交互能够让用户产生独立封闭的沉浸式交互体验。目前消费品的HMD主要有两种产品形态一种是利用个人电脑(PC)计算能力接入方式的PC头盔显示设备、另一种是基于手机的计算处理能力的便携式头盔显示设备。
VR系统,操控主要有手柄、遥控器、运动传感器等。这些操作由于需要采用外部设备输入,会时刻提醒用户其所操作的是虚拟现实系统,从而严重影响VR系统的沉浸感。因此,现有技术对VR系统的输入出现采用手势输入的技术方案。
现有技术对VR系统的手势输入主要有:
基于单个普通摄像头的手势识别,该方案由于仅能识别二维手势,因此沉浸感有限;
基于双红外摄像头的三维手势识别,该方案虽然沉浸感好,但成本与技术风险都比较高。
发明内容
基于此,有必要针对现有技术在VR系统中没有成本较低且沉浸感好的手势识别的技术问题,提供一种用于虚拟现实显示输出设备的手势识别方法及虚拟现实显示输出设备。
本发明实施例提供一种用于虚拟现实显示输出设备的手势识别方法,包括:
从第一摄像头获取第一视频,从第二摄像头获取第二视频;
从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势;
采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
获取所述空间手势对应的执行指令;
执行所述执行指令。
进一步的:
所述从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势,具体包括:
从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
所述采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势,具体包括:
采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手 势。
更进一步的,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
更进一步的,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
所述采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势,具体包括:
采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一个所述活动部位空间信息的空间手势。
再进一步的,所述获取所述空间手势对应的执行指令,具体包括:
将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
本发明实施例提供一种计算机程序,包括在计算机上运行时,适合执行如前所述的手势识别方法的所有步骤的计算机代码。
进一步的,所述计算机程序收录在计算机可读媒介上。
本发明实施例提供一种虚拟现实显示输出设备,包括:
视频获取模块,用于:从第一摄像头获取第一视频,从第二摄像头获取第二视频;
手分离模块,用于:从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势;
空间信息构造模块,用于:采用双目成像方式将所述第一平面 信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
指令获取模块,用于:获取所述空间手势对应的执行指令;
执行模块,用于:执行所述执行指令。
进一步的:
所述手分离模块,具体用于:从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
所述空间信息构造模块,具体用于:采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手势。
更进一步的,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
更进一步的,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
所述空间信息构造模块,具体用于:采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一个所述活动部位空间信息的空间手势。
进一步的,所述指令获取模块,具体用于:
将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采 用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
本发明实施例通过先将手图形从两个摄像头所获取的视频中分离出来后再采用双目成像方式进行合并,由于将手图形分离出来后避免了外部环境的干扰,无需计算手图形以外的背景,只需要对手图形采用双目成像方式计算空间信息,大大减少了计算量。因此能较以非常小的计算量获得手图形的空间信息,使得能够采用普通摄像头完成三维手势的识别,大大降低了虚拟现实显示输出设备的成本及技术风险。
附图说明
图1为本发明一实施例提供的一种用于虚拟现实显示输出设备的手势识别方法的工作流程图;
图2为本发明另一实施例提供的一种用于虚拟现实显示输出设备的手势识别方法的工作流程图;
图3为本发明一实施例提供的一种虚拟现实显示输出设备的结构模块图;
图4为本发明一实施例提供的虚拟现实显示输出设备的结构示意图。
具体实施方式
下面结合附图和具体实施例对本发明做进一步详细的说明。
如图1所示为本发明一实施例提供的一种用于虚拟现实显示输出设备的手势识别方法的工作流程图,包括:
步骤S101,包括:从第一摄像头获取第一视频,从第二摄像头获取第二视频;
步骤S102,包括:从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势;
步骤S103,包括:采用双目成像方式将所述第一平面信息和所 述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
步骤S104,包括:获取所述空间手势对应的执行指令;
步骤S105,包括:执行所述执行指令。
用户在虚拟现实显示输出设备前做出手势,该手势会在步骤S101从两个普通摄像头获得的第一视频和第二视频中形成手图形,然后在步骤S102中分离出第一平面信息和第二平面信息。第一平面信息和第二平面信息指的是手图形在第一视频中的平面位置以及在第二视频中的平面位置。由于单个摄像头只能获得平面位置,因此如果要获得三维位置还需要进行双目成像,双目成像主要的功能为双目测距,主要是利用了目标点,在这里是手,在左右两幅视图上成像的横向坐标之间存在的差异(即视差)与目标点到成像平面的距离Z存在着反比例的关系借助由两个摄像头的间距导致的视差计算出目标点(即手)与摄像头的距离,从而确定目标点(即手)在空间中的位置作为空间信息。
在获得手势的空间信息后执行步骤S104获取执行指令并在步骤S105中执行相应指令。通过执行指令使得用户可以通过手势与虚拟现实显示输出设备进行交互。
现有技术之所以难以采用普通摄像头降低成本的原因,在于普通摄像头所摄录的图像会包括手以及手附近的背景,因此如果直接进行双目成像的话由于背景的干扰,很难准确识别出用户的手,为此,本发明实施例先执行步骤S102,将每个摄像头的视频作为一个单独的视频进行分离,在分离后再执行步骤S103进行双目成像,从而避免了双目成像时背景的干扰,大大的减少了计算量,使得能够采用普通摄像头完成三维手势的识别,大大降低了虚拟现实显示输出设备的成本及技术风险。
在其中一个实施例中:
所述步骤S102,具体包括:从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将 每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
所述步骤S103,具体包括:采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手势。
本实施例从每帧图像中分离出手图形,并建立对应的时间戳,然后在步骤S103中将具有相同时间戳的所述第一平面信息和所述第二平面信息转换为空间信息,使得空间信息的计算更为准确。
在其中一个实施例中,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
本实施例采用的手检测包括:基于肤色的检测,基于运动信息的检测,基于特征的检测,基于图像分割的目标检测等等。手跟踪包括:采用粒子跟踪、CamShift算法等跟踪算法,也可与卡尔曼滤波相结合达到更好的效果。
本实施例通过手检测和手跟踪方式使得手图形的分离更为准确,使得后续的空间信息计算更为准确,识别出更准确的空间手势。
在其中一个实施例中,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
所述步骤S103,具体包括:采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一个所述活动部位空间信息的空间手势。
活动部位指的是人手中可活动的部位,例如手指等。活动部分 可以预先指定,由于手图形已经被分离出来,因此没有其他背景的干扰,而活动部位一般处于手图形边缘,因此可以通过边缘特征抽取等方式很方便地识别。
本实施例进一步计算活动部位的活动部位空间信息,使得能够识别更多更细致的手势。
在其中一个实施例中,所述步骤S104,具体包括:
将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
该手势分类模型的输入为空间手势,输出为手势类型。机器学习可以是有监督方式,例如有监督训练时指定每个用于训练的空间手势的类型,通过多次训练后得出手势分类模型。也可以是无监督方式,类如类型归类,如采用K最邻近结点算法(k-Nearest Neighbor algorithm,KNN),将用于训练的空间手势根据其空间位置进行归类。
本实施例采用机器学习方式建立手势分类模型,便于将手势归类,从而增加手势识别的鲁棒性。
如图2所示为本发明另一实施例提供的一种用于虚拟现实显示输出设备的手势识别方法的工作流程图;包括:
步骤S201,利用两个普通摄像头,分别单独采集图像数据;
用户在虚拟现实显示输出设备前做出手势,该手势会在两个普通摄像头获得的第一视频和第二视频中形成手图形;
步骤S202,对两个摄像头采集的数据,分别进行手的检测与跟踪;
检测时,可以采用的方法很多,如基于肤色的检测,基于运动信息的检测,基于特征的检测,基于图像分割的目标检测等等。跟踪时,可以采用粒子跟踪、CamShift算法等跟踪算法,也可与卡尔曼滤波相结合达到更好的效果;
步骤S203,对已经检测和跟踪的手,利用双目成像的原理,借 助由两个摄像头的间距导致的视差,得到手到摄像头的距离;
在左右两幅视图上成像的横向坐标之间存在的差异(即视差)与目标点到成像平面的距离Z存在着反比例的关系借助由两个摄像头的间距导致的视差,计算出目标点(即手)与摄像头的距离;
步骤S204,此时获得的手的信息,既包含了颜色信息,也包含了其在空间中的位置信息,此时既可进行手型的识别,也可进行三维意义上的手势识别;
步骤S205,识别到的手势,驱动响应的消息或者事件与VR系统进行交互。
如图3所示为本发明一实施例提供的一种虚拟现实显示输出设备的结构模块图,包括:
视频获取模块301,用于:从第一摄像头获取第一视频,从第二摄像头获取第二视频;
手分离模块302,用于:从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势;
空间信息构造模块303,用于:采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
指令获取模块304,用于:获取所述空间手势对应的执行指令;
执行模块305,用于:执行所述执行指令。
本发明实施例使得能够采用普通摄像头完成三维手势的识别,大大降低了虚拟现实显示输出设备的成本及技术风险。
在其中一个实施例中:
所述手分离模块302,具体用于:从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形 的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
所述空间信息构造模块303,具体用于:采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手势。
本实施例使得空间信息的计算更为准确。
在其中一个实施例中,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
本实施例通过手检测和手跟踪方式使得手图形的分离更为准确,使得后续的空间信息计算更为准确,识别出更准确的空间手势。
在其中一个实施例中,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
所述空间信息构造模块,具体用于:采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一个所述活动部位空间信息的空间手势。
本实施例进一步计算活动部位的活动部位空间信息,使得能够识别更多更细致的手势。
在其中一个实施例中,所述指令获取模块304,具体用于:
将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
本实施例采用机器学习方式建立手势分类模型,便于将手势归类,从而增加手势识别的鲁棒性。
如图4所示为本发明实施例提供的虚拟现实显示输出设备的结构示意图。虚拟现实显示输出设备可以是利用PC计算能力接入方式的 PC头盔显示设备、或者基于手机的计算处理能力的便携式头盔显示设备、或者是头盔显示设备自带计算处理能力,其主要包括:处理器401、存储器402以及两个摄像头403等。
其中存储器402中存储前述方法的具体代码,由处理器401具体执行,通过摄像头403捕捉手势,并由处理器401根据前述方法进行处理后执行相应操作。
此外,上述的存储器402中的逻辑指令可以通过软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本发明的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台移动终端(可以是个人计算机,服务器,或者网络设备等)执行本发明各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本发明实施例方案的目的。本领域普通技术人员在不付出创造性的劳动的情况下,即可以理解并实施。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到各实施方式可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的理解,上述技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如ROM/RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行各个实施例或者实施例的某些部分所述的方法。
最后应说明的是:以上实施例仅用以说明本发明实施例的技术方案,而非对其限制;尽管参照前述实施例对本发明实施例进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的精神和范围。

Claims (12)

  1. 一种用于虚拟现实显示输出设备的手势识别方法,其特征在于,包括:
    从第一摄像头获取第一视频,从第二摄像头获取第二视频;
    从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势;
    采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
    获取所述空间手势对应的执行指令;
    执行所述执行指令。
  2. 根据权利要求1所述的用于虚拟现实显示输出设备的手势识别方法,其特征在于:
    所述从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于第二视频中手图形的第二平面信息的第二平面手势,具体包括:
    从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
    所述采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势,具体包括:
    采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手势。
  3. 根据权利要求2所述的用于虚拟现实显示输出设备的手势识别 方法,其特征在于,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
  4. 根据权利要求2所述的用于虚拟现实显示输出设备的手势识别方法,其特征在于,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
    所述采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势,具体包括:
    采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一个所述活动部位空间信息的空间手势。
  5. 根据权利要求1~4任一项所述的用于虚拟现实显示输出设备的手势识别方法,其特征在于,所述获取所述空间手势对应的执行指令,具体包括:
    将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
  6. 一种计算机程序,其特征在于,包括在计算机上运行时,适合执行如权利要求1~5任一项所述的手势识别方法的所有步骤的计算机代码。
  7. 根据权利要求6所述的计算机程序,其特征在于,所述计算机程序收录在计算机可读媒介上。
  8. 一种虚拟现实显示输出设备,其特征在于,包括:
    视频获取模块,用于:从第一摄像头获取第一视频,从第二摄像头获取第二视频;
    手分离模块,用于:从所述第一视频中分离出关于第一视频中手图形的第一平面信息的第一平面手势,从所述第二视频中分离出关于 第二视频中手图形的第二平面信息的第二平面手势;
    空间信息构造模块,用于:采用双目成像方式将所述第一平面信息和所述第二平面信息转换为空间信息,生成包括所述空间信息的空间手势;
    指令获取模块,用于:获取所述空间手势对应的执行指令;
    执行模块,用于:执行所述执行指令。
  9. 根据权利要求8所述的虚拟现实显示输出设备,其特征在于:
    所述手分离模块,具体用于:从所述第一视频中每帧第一图像中分离出来手图形,获取从每帧所述第一图像中分离出来的手图形的第一平面信息,将多个第一平面信息组合为所述第一平面手势,将每个所述第一平面信息所对应的所述第一图像的时间戳作为每个所述第一平面信息的时间戳,从所述第二视频中每帧第二图像中分离出来手图形,获取从每帧所述第二图像中分离出来的手图形的第二平面信息,将多个第二平面信息组合为所述第二平面手势,将每个所述第二平面信息所对应的所述第二图像的时间戳作为每个所述第二平面信息的时间戳;
    所述空间信息构造模块,具体用于:采用双目成像方式将具有相同时间戳的所述第一平面信息和所述第二平面信息计算为空间信息,生成包括所述空间信息的空间手势。
  10. 根据权利要求9所述的虚拟现实显示输出设备,其特征在于,采用手检测和手跟踪方式从所述第一视频中每帧第一图像中分离出来手图形,采用手检测和手跟踪方式从所述第二视频中每帧第二图像中分离出来手图形。
  11. 根据权利要求9所述的虚拟现实显示输出设备,其特征在于,所述第一平面信息包括手图形中至少一个活动部位的第一活动部位平面信息,所述第二平面信息包括手图形中至少一个活动部位的第二活动部位平面信息;
    所述空间信息构造模块,具体用于:采用双目成像方式将同一活动部位相同时间戳的所述第一活动部位平面信息和所述第二活动部位平面信息计算出该活动部位的活动部位空间信息,生成包括至少一 个所述活动部位空间信息的空间手势。
  12. 根据权利要求8~11任一项所述的虚拟现实显示输出设备,其特征在于,所述指令获取模块,具体用于:
    将所述空间手势输入手势分类模型,得到所述空间手势的手势类型,获取所述手势类型对应的执行指令,所述手势分类模型为采用多个预先获取的空间手势采用机器学习进行训练后得到的关于空间手势的类型分类模型。
PCT/CN2016/085365 2015-11-18 2016-06-08 手势识别方法及虚拟现实显示输出设备 Ceased WO2017084319A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US15/240,571 US20170140215A1 (en) 2015-11-18 2016-08-18 Gesture recognition method and virtual reality display output device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510796509.6 2015-11-18
CN201510796509.6A CN105892633A (zh) 2015-11-18 2015-11-18 手势识别方法及虚拟现实显示输出设备

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US15/240,571 Continuation US20170140215A1 (en) 2015-11-18 2016-08-18 Gesture recognition method and virtual reality display output device

Publications (1)

Publication Number Publication Date
WO2017084319A1 true WO2017084319A1 (zh) 2017-05-26

Family

ID=57002295

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/085365 Ceased WO2017084319A1 (zh) 2015-11-18 2016-06-08 手势识别方法及虚拟现实显示输出设备

Country Status (2)

Country Link
CN (1) CN105892633A (zh)
WO (1) WO2017084319A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10497179B2 (en) 2018-02-23 2019-12-03 Hong Kong Applied Science and Technology Research Institute Company Limited Apparatus and method for performing real object detection and control using a virtual reality head mounted display system
CN110751082A (zh) * 2019-10-17 2020-02-04 烟台艾易新能源有限公司 一种智能家庭娱乐系统手势指令识别方法
CN111367415A (zh) * 2020-03-17 2020-07-03 北京明略软件系统有限公司 一种设备的控制方法、装置、计算机设备和介质

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105892638A (zh) * 2015-12-01 2016-08-24 乐视致新电子科技(天津)有限公司 一种虚拟现实交互方法、装置和系统
CN106598235B (zh) * 2016-11-29 2019-10-22 歌尔科技有限公司 用于虚拟现实设备的手势识别方法、装置及虚拟现实设备
CN106951069A (zh) * 2017-02-23 2017-07-14 深圳市金立通信设备有限公司 一种虚拟现实界面的控制方法及虚拟现实设备
CN107087153B (zh) * 2017-04-05 2020-07-31 深圳市冠旭电子股份有限公司 3d图像生成方法、装置及vr设备
CN107272899B (zh) * 2017-06-21 2020-10-30 北京奇艺世纪科技有限公司 一种基于动态手势的vr交互方法、装置及电子设备
CN110188886B (zh) * 2018-08-17 2021-08-20 第四范式(北京)技术有限公司 对机器学习过程的数据处理步骤进行可视化的方法和系统

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6147678A (en) * 1998-12-09 2000-11-14 Lucent Technologies Inc. Video hand image-three-dimensional computer interface with multiple degrees of freedom
CN101344816A (zh) * 2008-08-15 2009-01-14 华南理工大学 基于视线跟踪和手势识别的人机交互方法及装置
CN102350700A (zh) * 2011-09-19 2012-02-15 华南理工大学 一种基于视觉的机器人控制方法
CN102789568A (zh) * 2012-07-13 2012-11-21 浙江捷尚视觉科技有限公司 一种基于深度信息的手势识别方法
US20130249786A1 (en) * 2012-03-20 2013-09-26 Robert Wang Gesture-based control system
CN103576840A (zh) * 2012-07-24 2014-02-12 上海辰戌信息科技有限公司 基于立体视觉的手势体感控制系统
US8971572B1 (en) * 2011-08-12 2015-03-03 The Research Foundation For The State University Of New York Hand pointing estimation for human computer interaction

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6147678A (en) * 1998-12-09 2000-11-14 Lucent Technologies Inc. Video hand image-three-dimensional computer interface with multiple degrees of freedom
CN101344816A (zh) * 2008-08-15 2009-01-14 华南理工大学 基于视线跟踪和手势识别的人机交互方法及装置
US8971572B1 (en) * 2011-08-12 2015-03-03 The Research Foundation For The State University Of New York Hand pointing estimation for human computer interaction
CN102350700A (zh) * 2011-09-19 2012-02-15 华南理工大学 一种基于视觉的机器人控制方法
US20130249786A1 (en) * 2012-03-20 2013-09-26 Robert Wang Gesture-based control system
CN102789568A (zh) * 2012-07-13 2012-11-21 浙江捷尚视觉科技有限公司 一种基于深度信息的手势识别方法
CN103576840A (zh) * 2012-07-24 2014-02-12 上海辰戌信息科技有限公司 基于立体视觉的手势体感控制系统

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10497179B2 (en) 2018-02-23 2019-12-03 Hong Kong Applied Science and Technology Research Institute Company Limited Apparatus and method for performing real object detection and control using a virtual reality head mounted display system
CN110751082A (zh) * 2019-10-17 2020-02-04 烟台艾易新能源有限公司 一种智能家庭娱乐系统手势指令识别方法
CN110751082B (zh) * 2019-10-17 2023-12-12 烟台艾易新能源有限公司 一种智能家庭娱乐系统手势指令识别方法
CN111367415A (zh) * 2020-03-17 2020-07-03 北京明略软件系统有限公司 一种设备的控制方法、装置、计算机设备和介质
CN111367415B (zh) * 2020-03-17 2024-01-23 北京明略软件系统有限公司 一种设备的控制方法、装置、计算机设备和介质

Also Published As

Publication number Publication date
CN105892633A (zh) 2016-08-24

Similar Documents

Publication Publication Date Title
Dubey et al. A comprehensive survey on human pose estimation approaches
WO2017084319A1 (zh) 手势识别方法及虚拟现实显示输出设备
CN113706699B (zh) 数据处理方法、装置、电子设备及计算机可读存储介质
US10674142B2 (en) Optimized object scanning using sensor fusion
US11842514B1 (en) Determining a pose of an object from rgb-d images
US11113842B2 (en) Method and apparatus with gaze estimation
US10732725B2 (en) Method and apparatus of interactive display based on gesture recognition
JP2024045273A (ja) 非制約環境において人間の視線及びジェスチャを検出するシステムと方法
RU2644520C2 (ru) Бесконтактный ввод
US20210026455A1 (en) Gesture recognition system and method of using same
WO2021093453A1 (zh) 三维表情基的生成方法、语音互动方法、装置及介质
KR20220009393A (ko) 이미지 기반 로컬화
TW201814438A (zh) 基於虛擬實境場景的輸入方法及裝置
WO2022174594A1 (zh) 基于多相机的裸手追踪显示方法、装置及系统
US11106949B2 (en) Action classification based on manipulated object movement
CN112991555B (zh) 数据展示方法、装置、设备以及存储介质
US20170140215A1 (en) Gesture recognition method and virtual reality display output device
US20160110909A1 (en) Method and apparatus for creating texture map and method of creating database
WO2023168957A1 (zh) 姿态确定方法、装置、电子设备、存储介质及程序
CN112905014A (zh) Ar场景下的交互方法、装置、电子设备及存储介质
CN105892637A (zh) 手势识别方法及虚拟现实显示输出设备
Kowalski et al. Holoface: Augmenting human-to-human interactions on hololens
WO2019085519A1 (zh) 面部跟踪方法及装置
CN117435055A (zh) 基于空间立体显示器的手势增强眼球追踪的人机交互方法
Ueng et al. Vision based multi-user human computer interaction

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16865500

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16865500

Country of ref document: EP

Kind code of ref document: A1