WO2014101606A1 - 人机交互操作的触发控制方法和装置 - Google Patents

人机交互操作的触发控制方法和装置 Download PDF

Info

Publication number
WO2014101606A1
WO2014101606A1 PCT/CN2013/087811 CN2013087811W WO2014101606A1 WO 2014101606 A1 WO2014101606 A1 WO 2014101606A1 CN 2013087811 W CN2013087811 W CN 2013087811W WO 2014101606 A1 WO2014101606 A1 WO 2014101606A1
Authority
WO
WIPO (PCT)
Prior art keywords
designated
contour
display screen
specified
area
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2013/087811
Other languages
English (en)
French (fr)
Inventor
周彬
盛森
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology Shenzhen Co Ltd
Original Assignee
Tencent Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology Shenzhen Co Ltd filed Critical Tencent Technology Shenzhen Co Ltd
Publication of WO2014101606A1 publication Critical patent/WO2014101606A1/zh
Priority to US14/750,697 priority Critical patent/US9829974B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/03Arrangements for converting the position or the displacement of a member into a coded form
    • G06F3/0304Detection arrangements using opto-electronic means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/03Arrangements for converting the position or the displacement of a member into a coded form
    • G06F3/041Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means
    • G06F3/042Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means by opto-electronic means
    • G06F3/0425Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means by opto-electronic means using a single imaging device like a video camera for tracking the absolute position of a single or a plurality of objects with respect to an imaged reference surface, e.g. video camera imaging a display or a projection screen, a table or a wall surface, on which a computer generated image is displayed or projected
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/048Interaction techniques based on graphical user interfaces [GUI]
    • G06F3/0484Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
    • G06F3/04842Selection of displayed objects or displayed text elements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T1/00General purpose image data processing
    • G06T1/0007Image acquisition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • G06V40/19Sensors therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10004Still image; Photographic image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face

Definitions

  • the invention relates to a Chinese patent application for the human-computer interaction operation.
  • the present application claims to be Chinese patent application filed on December 28, 2012, the Chinese Patent Office, the application number 2012105838196, and the invention titled “trigger control method and device for human-computer interaction operation" Priority is hereby incorporated by reference in its entirety.
  • the present application relates to the field of computer human-computer interaction technology, and in particular, to a trigger control method and apparatus for human-computer interaction operation. Background of the invention
  • Human-Computer Interaction Techniques are technologies that enable people to talk to computers in an efficient manner through computer input and output devices.
  • the technology includes: The machine provides a large amount of relevant information and prompt instructions through an output or display device, and the person inputs relevant information, operation instructions, and answers questions to the machine through the input device.
  • Human-computer interaction technology is one of the important contents in the design of computer user interface.
  • the input device may be a keyboard, a mouse, or a touch screen.
  • the computer can respond to the instruction information and perform corresponding operations.
  • the person can also click a related button on the computer interface by using a mouse.
  • the computer can respond to the instruction and perform the corresponding operation. For example, if a person clicks the "close" button with a mouse, the computer closes the window corresponding to the "close” button and the like.
  • Embodiments of the present invention provide a trigger control method and apparatus for human-computer interaction operation, so as to facilitate a person with a disability to trigger a machine operation in a non-contact manner.
  • a trigger control method for human-computer interaction operation comprising:
  • a trigger control device for human-computer interaction operation comprising:
  • a first module configured to acquire a camera image captured by the camera device, and display the camera image in a blurred manner on the display screen;
  • a second module configured to detect an interframe difference of the image capturing frame, identify a specified contour on the image capturing screen according to the interframe difference, and calculate a position of the identified specified contour on the display screen;
  • the third module is configured to determine in real time whether the position of the specified contour on the display screen intersects with the designated area displayed on the display screen, and if intersected, trigger the operation corresponding to the designated area.
  • the embodiment of the present invention obtains the camera image captured by the camera device, and displays the camera image in a blurred manner, for example, in a translucent manner on the display screen, so that the camera screen can be displayed on the display screen.
  • the other interfaces overlap, and a specified contour on the image capturing screen (for example, an outline of an eye of a person, a person's mouth, etc.) can be recognized, and the user can control the movement of the specified contour in the image capturing screen by moving the body, when When the specified outline intersects with a specified area displayed on the display (for example, a display area that can be a medium information, or a designated command area such as a button, a link, etc.), the designated area pair is triggered.
  • the operation should be. Therefore, it is possible to realize the interaction between the human and the human hands without using the hand, and it is convenient for the disabled person with the hand to trigger the machine operation in a non-contact manner.
  • FIG. 1 is a schematic flowchart of a trigger control method for human-computer interaction operation according to an embodiment of the present invention
  • FIG. 2a is a schematic diagram of a first type of machine interface in which a designated area displayed on a display screen is a display area for specifying media information according to an embodiment of the present invention
  • 2b is a schematic diagram of a second machine interface of a display area displayed on a display screen as a designated medium information according to an embodiment of the present invention
  • FIG. 3 is a schematic diagram of a first type of machine interface in which a designated area displayed on a display screen is a designated command area according to an embodiment of the present invention
  • FIG. 3b is a schematic diagram of a second machine interface in which a designated area displayed on a display screen is a designated command area according to an embodiment of the present invention
  • FIG. 4 is a schematic diagram showing the composition of a trigger control device for human-computer interaction operation according to an embodiment of the present invention
  • FIG. 5 is a schematic diagram showing the hardware structure of a trigger control device for human-computer interaction operation according to an embodiment of the present invention. detailed description
  • the embodiment of the invention provides a trigger control method and device for human-computer interaction operation.
  • an image capturing image captured by the image capturing device is acquired, and the image capturing screen is displayed in a blurred manner on the display screen; an interframe difference of the image capturing screen is detected, and the image capturing screen is identified according to the interframe difference Specifying the contour on the upper side, and calculating the position of the identified specified contour on the display screen; determining in real time whether the position of the specified contour on the display screen intersects with the designated area displayed on the display screen, and if intersected, triggering The operation corresponding to the specified area.
  • FIG. 1 is a schematic flowchart of a trigger control method for human-computer interaction operation according to an embodiment of the present invention. As shown in Figure 1, the method mainly includes:
  • Step 101 Acquire an image capturing screen captured by the camera device, and display the image capturing screen in a blurred manner on the display screen.
  • the ambiguity mode may be a specified display mode, such as a semi-transparent display mode,
  • the camera screen is displayed in a translucent manner on the display screen; or the camera screen can be converted into an animated rim screen (for example, an animated screen with a single-line rim), the animated contour screen can be superimposed on the display screen
  • the image capturing screen is displayed in a semi-transparent manner on the display screen as an example.
  • Step 102 Detect an interframe difference of the image capturing screen, identify a specified contour on the image capturing screen according to the interframe difference, and calculate a position of the identified specified contour on the display screen.
  • Step 103 Determine in real time whether the position of the designated contour on the display screen intersects with the designated area displayed on the display screen. If they intersect, perform step 104; otherwise, return to step 102.
  • Step 104 Trigger an operation corresponding to the designated area.
  • the designating the specified contour on the image capturing screen may be to identify the contour of an organ of the human body, or to identify other graphic contours.
  • a camera device such as a camera
  • the camera device usually shoots a user's head.
  • the camera device usually captures the user's head, especially It is an image of the face.
  • the specified contour on the recognition imaging screen may be an outline of the eye of the person, because the contour of the human eye It is more standard, and can further issue further operational instructions to the machine by detecting motion patterns such as blinking.
  • the designated contour on the identified camera image may also be the outline of an organ such as a person's mouth, or even a specified standard figure.
  • the user may be provided with a whiteboard pre-drawn with a specified graphic, and the user may lift the whiteboard in front of the camera to let the camera capture a specified graphic on the whiteboard (for example, a well-defined oval, etc.), the designated graphic is this The specified contour to be detected in the embodiment of the invention.
  • a designated area for example, a display area which may be a medium information, or a designated instruction area such as a button, a link, etc.
  • the operation corresponding to the designated area is triggered, so that the interaction between the human and the human machine can be triggered without using the hand, and the handicapped person with the hand can be triggered to trigger the operation of the machine in a non-contact manner.
  • the detecting the interframe difference of the imaging screen, identifying the specified contour on the imaging screen based on the interframe difference, and calculating the position of the identified designated contour on the display screen can be achieved with programming tools.
  • it can be implemented by using a specific interface function in the open source computer vision library (openCV, Open Source Computer Vision Library).
  • OpenCV is a cross-platform computer vision library based on open source distribution that runs on computer operating systems such as Linux, Windows and Mac OS.
  • LightCPC is lightweight and efficient. It consists of a series of C functions and a small number of C++ classes. It also provides a call interface for languages such as Python, Ruby, and MATLAB. It implements many common calculation methods for image processing and computer vision.
  • the cvSub interface function and the cvThreshold interface function in OpenCV can be used to detect the interframe difference of the captured picture.
  • the specific implementation code instructions are as follows:
  • the gray is the current frame of the imaging picture
  • the prev is the previous frame of the current frame
  • the diff is the interframe difference
  • the cvFindContours interface function in OpenCV may be employed to identify a specified contour based on the interframe difference, such as identifying an eye contour.
  • Specific The implementation code instructions are as follows:
  • the diff is the calculated interframe difference
  • the comp is the identified eye contour
  • the eye contour is output by the cvFindContours interface function.
  • the cvSetlmageROI interface function in OpenCV can be used to calculate the position of the identified specified contour on the display.
  • the specific implementation code instructions are as follows:
  • rect_eye is the position of the eye contour output by the interface function cvSetlmageROI in the current frame gray, and according to the position occupied by the current frame gray on the display screen, the current position of the eye contour on the display screen can be calculated.
  • the designated area displayed on the display screen in the embodiment of the present invention may have various forms, for example, a display area of electronic media information (referred to as medium information in the embodiment of the present invention), or a designated instruction area.
  • medium information in the embodiment of the present invention
  • Figure 2a is a schematic diagram of a first machine interface in which the designated area displayed on the display screen is the display area of the specified media information.
  • a medium is displayed on the machine interface 200 Information 201 and media information 202.
  • the avatar of the user captured by the camera device is displayed in the machine interface 200 in a semi-transparent manner, so that the user avatar can overlap with the information in the machine interface 200, so that the user can see the interface 200.
  • the various information, and the avatar of the person can be seen, so that the movement of the outline of the eye can be observed while moving the head of the eye, so that the contour of the eye can be moved to the designated area, that is, the display area of the medium information 201 or the medium information 202.
  • Figure 2b is a schematic diagram of a second machine interface in which the designated area displayed on the display screen is the display area of the specified media information.
  • the specified medium information 201 is triggered. The operation corresponding to the display area.
  • Corresponding operations include: recording an intersection time of the eye contour 203 and the display area of the specified media information 201, determining whether the position of the eye contour 203 on the display screen moves out of the display area of the specified media information 201, and if so Stop recording the intersection time, otherwise continue to record the intersection time.
  • the degree of attention of the user to the media information 201 can be calculated, and according to the degree of attention, other related operations, such as a charging operation, can be further performed, that is, according to the recorded specified contour (such as an eye contour) 203) Calculate the billing information corresponding to the specified medium information 201 with the intersection time of the display area of the specified medium information 201.
  • the recorded specified contour such as an eye contour
  • media information displayed on the network is based on the user's click on the media information and the number of exposures of the media information, instead of viewing the medium based on the viewer.
  • the length of the information is charged.
  • the application process of the embodiment of the present invention can calculate the intersection time of the user's eyes and the specified media information, which is equivalent to the user's attention to the media information, and can be implemented as a basis. A new way of charging media information, for example, still taking FIG.
  • the timing is started after the display area intersects, and the timing is stopped when the position of the eye contour 203 on the display screen moves out of the display area of the specified medium information 201. If the intersection time is greater than a predetermined number of seconds, the The media information 201 is charged, so that the media information can be charged based on the user's attention to the media information, so that the charging method is more refined and accurate.
  • the operation corresponding to the triggered specified medium information display area includes: detecting whether the specified contour occurs The specified motion pattern, if it is, triggers the instruction operation bound to the specified area. For example, it is detected whether the eye contour 203 has a blinking action, and if so, an instruction operation to which the designated area is bound is triggered.
  • the instruction operation of the specified media information 201 is a click action, and after the user blinks, the click action on the media information 201 can be triggered to open the web page pointed to by the media information 201.
  • Figure 3a is a schematic diagram of the first machine interface in the specified area displayed on the display for the specified command area.
  • media information 201 and media information 202 are displayed on the machine interface 200.
  • the media information 201 also has a designated command area.
  • the "change one" button 301 and the “close” button 302 are designated command areas.
  • the instruction of the "change one" button 301 is operated to switch to the next medium information, and the instruction of the "close” button 302 is operated to close the current medium information 201.
  • the avatar of the user captured by the camera device is displayed in the machine interface 200 in a semi-transparent manner, so that the user avatar can overlap with the information in the machine interface 200, so that the user can see the interface 200.
  • Figure 3b is a schematic diagram of a second machine interface in which the designated area displayed on the display screen is a designated command area.
  • the "change one" button 301 that is, the eye contour 203 and the "change one" button 301
  • the media information 201 is switched to the next media information.
  • the eye contour 203 When the user's eye contour 203 is moved to the "close” button 302, that is, the eye contour 203 intersects the "close” button 302, it can be detected whether the eye contour 203 has a specified motion pattern (such as a blinking motion), and if so The instruction operation bound by the "close” button 302 is triggered, that is, the current media information 201 is closed.
  • a specified motion pattern such as a blinking motion
  • the specified motion pattern when the specified contour is another image contour, the specified motion pattern may be an action corresponding to the image contour.
  • the specified motion pattern may be an opening and closing action of the mouth.
  • the detecting whether the specified contour has a specified motion state comprises:
  • the cvResetlmageROI(gray) interface function in OpenCV creates an eye template.
  • detecting a frame image in the template of the specified contour (such as an eye template), determining whether the change of the frame image conforms to a specified motion form; if yes, triggering an instruction operation bound by the designated area.
  • the detecting whether the specified contour occurs in a specified motion form may specifically be: detecting whether the eye contour has a blinking motion.
  • the specific method for detecting whether the eye contour has a blinking action may include: detecting a boundary value of the eye contour; detecting a maximum value and a minimum value of the boundary value; detecting whether a distance between the maximum value and the minimum value of the boundary value occurs from large to small and then from small to large, if Then, it is determined that the eye contour has a blinking motion.
  • the relevant interface parameters in OpenCV can be used to determine if an eye contour has a blinking action.
  • the specific method may include the following steps 411 ⁇ 413:
  • Step 411 Detect a boundary of an eye contour according to the cvMatchTemplate interface function.
  • Specific code instructions are as follows:
  • tpl is the eye template created by the cvResetlmageROI(gray) interface function.
  • Step 412 The cvMinMaxLoc interface function detects a maximum value and a minimum value of a boundary value of the eye contour.
  • the specific code instructions are as follows:
  • Step 413 Detect whether a distance between a maximum value and a minimum value of a boundary value of the eye contour occurs from large to small and then from small to large, that is, whether a closed motion of the eye occurs, and if yes, determine the location A blinking action occurs in the outline of the eye.
  • the embodiment of the present invention may blunt the display of the image captured by the camera, for example, the user's head image captured by the camera is in the video chat screen. Blur the display and display a web ad (ie, media information) in the video chat screen.
  • the web ad can display the ad content and can display the "change one" button and the "close” button.
  • the network advertisement can be switched to the next network advertisement, and when the eye contour is moved to the "close” button, the network advertisement can be shut down.
  • the embodiment of the invention further discloses a trigger control device for human-computer interaction operation to perform the above method.
  • the device may be implemented by a computer device, which may be a personal computer, a server, or a portable computer (such as a laptop, tablet, etc.).
  • the computer device can include at least one processor and a memory.
  • the memory stores machine readable instructions that, when executed by the processor, may implement all or part of the processes of the various method embodiments of the present invention described above.
  • the memory may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). Wait.
  • FIG. 4 is a schematic diagram of a composition of a trigger control apparatus for the human-machine interaction operation according to an embodiment of the present invention. Referring to Figure 4, the apparatus can include:
  • a first module 401 configured to acquire an image capturing screen captured by the camera device, and display the camera image in a semi-transparent manner on the display screen;
  • a second module 402 configured to detect an interframe difference of the image capturing frame, identify a specified contour on the image capturing screen according to the interframe difference, and calculate a position of the identified specified contour on the display screen;
  • the third module 403 is configured to determine in real time whether the position of the specified contour on the display screen intersects with the designated area displayed on the display screen, and if intersected, trigger the operation corresponding to the designated area.
  • the third module 403 is specifically configured to: determine in real time whether the position of the specified contour on the display screen intersects with a specified area displayed on the display screen, and if intersected, trigger the corresponding area corresponding to the designated area
  • the operation of the triggering includes: recording an intersection time of the specified contour with the designated area, determining whether a position of the specified contour on the display screen moves out of the designated area, and if so, stopping recording the intersecting time , otherwise continue to record the intersection time.
  • the charging operation may be further performed on the designated area according to the intersecting time.
  • the third module 403 is specifically configured to: determine in real time whether the position of the specified contour on the display screen intersects with a designated area displayed on the display screen, and if intersected, trigger the designated area to correspond to The operation of the triggering includes: detecting whether the specified contour has a specified motion pattern, and if so, triggering an instruction operation bound by the designated region.
  • the specified contour is an eye contour; and the third module 403 detects whether the specified contour has a specified motion pattern, specifically: detecting whether the eye contour has a blinking motion.
  • the specified contour may also be other image contours, such as the contour of a person's mouth.
  • the specified motion pattern may be an opening and closing motion of the mouth.
  • the designated area displayed on the display screen may be a display area for specifying media information, or the designated area displayed on the display screen may be a designated instruction area, or may be other designated display. Formal area.
  • the various modules in the above embodiments of the present invention may be implemented by software (for example, machine readable instructions executed by a processor stored in a computer readable medium), or may be implemented by hardware.
  • processor of the application specific integrated circuit ASIC
  • hardware for example, the processor of the application specific integrated circuit (ASIC)
  • ASIC application specific integrated circuit
  • modules in the foregoing embodiments of the present invention may be integrated into one or separately deployed; they may be combined into one module or further split into multiple sub-modules.
  • FIG. 5 is a schematic diagram showing the hardware structure of a trigger control device for human-computer interaction operation according to an embodiment of the present invention.
  • the apparatus can include: a processor 51, a memory 52, at least one port 53 and an interconnection mechanism 54.
  • the processor 51 and the memory 52 are interconnected by the interconnection mechanism 54.
  • the device can receive and transmit data information via port 53.
  • the memory 52 stores machine readable instructions;
  • the processor 51 executes the machine readable instructions to perform the following operations:
  • the processor 51 may execute machine readable instructions stored in the memory 52 to further perform all or part of the processes in the foregoing method embodiments, where Let me repeat.
  • the functions of the first module 401, the second module 402, and the third module 403 described above can be implemented when the machine readable instructions stored in the memory 52 are executed by the processor 51.
  • the embodiment of the present invention obtains the camera image captured by the camera device, and displays the camera image in a blurred manner, for example, in a translucent manner on the display screen, so that the camera screen can be displayed on the display screen.
  • the other interfaces overlap, and a specified contour on the image capturing screen (for example, an outline of an eye of a person, a person's mouth, etc.) can be recognized, and the user can control the movement of the specified contour in the image capturing screen by moving the body, when When the specified outline intersects with a specified area displayed on the display (for example, a display area that can be a medium information, or a designated command area such as a button, a link, etc.), the operation corresponding to the designated area is triggered. Therefore, it is possible to realize the interaction between the human and the human hands without using the hand, and it is convenient for the disabled person with the hand to trigger the machine operation in a non-contact manner.
  • a specified contour on the image capturing screen for example, an outline of an eye of
  • the hardware modules in the various embodiments of the present invention may be implemented mechanically or electronically.
  • An example processor such as an FPGA or ASIC, is used to perform a specific operation.
  • the hardware modules may also include programmable logic devices or circuits (such as including general purpose processors or other programmable processors) that are temporarily configured by software for performing particular operations.
  • programmable logic devices or circuits such as including general purpose processors or other programmable processors
  • Implementing hardware modules can be determined based on cost and time considerations.
  • the present invention can be implemented by means of software plus a necessary general hardware platform, that is, by machine-readable instructions to instruct related hardware, and of course It can be done through hardware, but in many cases the former is a better implementation.
  • the technical solution of the present invention which is essential or contributes to the prior art, may be embodied in the form of a software product stored in a storage medium, including a plurality of instructions for making a
  • the terminal device (which may be a cell phone, a personal computer, a server, or a network device, etc.) performs the methods described in various embodiments of the present invention.
  • the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
  • the machine readable instructions can be downloaded from the server computer by the communication network.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Psychiatry (AREA)
  • Social Psychology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Ophthalmology & Optometry (AREA)
  • User Interface Of Digital Computer (AREA)
  • Image Analysis (AREA)

Abstract

本发明实施例公开了一种人机交互操作的触发控制方法和装置,该方法包括:获取摄像装置拍摄的摄像画面,在显示屏上以虚化方式显示所述摄像画面;检测所述摄像画面的帧间差,根据所述帧间差识别所述摄像画面上的指定轮廓,并计算所识别出的指定轮廓在显示屏上的位置;实时判断所述指定轮廓在显示屏上的位置与显示屏上所显示的指定区域是否相交,如果相交,则触发所述指定区域对应的操作。

Description

人机交互操作的触发控制方法和装置 本申请要求于 2012 年 12 月 28 日提交中国专利局、 申请号为 2012105838196、 发明名称为 "人机交互操作的触发控制方法和装置" 的中国专利申请的优先权, 其全部内容通过引用结合在本申请中。 技术领域
本申请涉及计算机人机交互技术领域, 尤其涉及一种人机交互操作 的触发控制方法和装置。 发明背景
人机交互技术 ( Human- Computer Interaction Techniques )是指通过 计算机输入、 输出设备, 以有效的方式实现人与计算机对话的技术。 该 技术包括: 机器通过输出或显示设备给人提供大量有关信息及提示请示 等, 人通过输入设备给机器输入有关信息、 操作指令以及回答问题等。 人机交互技术是计算机用户界面设计中的重要内容之一。
目前的人机交互技术中, 当人通过输入设备向计算机输入有关信息 时, 通常需要用手来操作。 例如, 所述输入设备可以是键盘、 鼠标或触 摸屏等, 人使用键盘输入相关的指令信息, 则计算机可以响应该指令信 息并做出对应的操作; 人也可以使用鼠标点击计算机界面上的相关按钮 来完成指令的输入, 计算机则可以响应该指令并做出对应的操作。 例如 人用鼠标点击 "关闭" 按钮, 则计算机会关闭该 "关闭" 按钮对应的窗 口等。 发明内容
本发明实施例提供了一种人机交互操作的触发控制方法和装置, 以 方便残障人士通过非接触的方式触发机器操作。
一种人机交互操作的触发控制方法, 包括:
获取摄像装置拍摄的摄像画面, 在显示屏上以虚化方式显示所述摄 像画面;
检测所述摄像画面的帧间差, 根据所述帧间差识别所述摄像画面上 的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的位置;
实时判断该指定轮廓在显示屏上的位置与显示屏上所显示的指定区 域是否相交, 如果相交, 则触发该指定区域对应的操作。
一种人机交互操作的触发控制装置, 该装置包括:
第一模块, 用于获取摄像装置拍摄的摄像画面, 在显示屏上以虚化 方式显示所述摄像画面;
第二模块, 用于检测所述摄像画面的帧间差, 根据所述帧间差识别 所述摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的 位置;
第三模块, 用于实时判断该指定轮廓在显示屏上的位置与显示屏上 所显示的指定区域是否相交,如果相交,则触发该指定区域对应的操作。
由此可以看出, 本发明实施例通过获取摄像装置拍摄的摄像画面, 并在显示屏上以虚化方式, 例如半透明方式, 显示所述摄像画面, 从而 使摄像画面可以和显示屏上显示的其它界面相重叠, 并且可以识别出摄 像画面上的指定轮廓 (例如人的眼睛、 人的嘴巴等器官的轮廓), 用户 可以通过移动身体来控制所述摄像画面中指定轮廓的移动, 当该指定轮 廓与显示屏上所显示的指定区域(例如可以是一种媒介信息的显示区 域, 或者是指定指令区如按钮、 链接等)相交时, 则触发该指定区域对 应的操作。 因此, 可以实现不用手来触发人机之间的交互操作, 方便手 部有残疾的残障人士通过非接触的方式触发机器操作。 附图简要说明
以下附图为本发明技术方案的一些实施例, 本发明实施例并不局限 于图中示出的特征。 以下附图中, 相似的标号表示相似的元素:
图 1为本发明实施例提供的人机交互操作的触发控制方法的流程示 意图;
图 2a 为本发明实施例提供的在显示屏上所显示的指定区域为指定 媒介信息的显示区域的第一种机器界面示意图;
图 2b 为本发明实施例提供的在显示屏上所显示的指定区域为指定 媒介信息的显示区域的第二种机器界面示意图;
图 3a 为本发明实施例提供的在显示屏上所显示的指定区域为指定 指令区的第一种机器界面示意图;
图 3b 为本发明实施例提供的在显示屏上所显示的指定区域为指定 指令区的第二种机器界面示意图;
图 4为本发明实施例提供的人机交互操作的触发控制装置的组成示 意图;
图 5为本发明实施例提供的人机交互操作的触发控制装置的硬件结 构示意图。 具体实施方式
下面结合附图及具体实施例对本发明作进一步详细的说明。
为了描述上的筒洁和直观, 下文通过描述若干代表性的实施例来对 本发明的方案进行阐述。 实施例中大量的细节仅用于帮助理解本发明的 方案。 但是很明显, 本发明的技术方案实现时可以不局限于这些细节。 为了避免不必要地模糊了本发明的方案, 一些实施例没有进行细致地描 述, 而是仅给出了框架。 下文中, "包括" 是指 "包括但不限于", "根 据…… " 是指 "至少根据……, 但不限于仅根据…… "。 由于汉语的语 言习惯, 下文中没有特别指出一个成分的数量时, 意味着该成分可以是 一个也可以是多个, 或可理解为至少一个。
目前, 通过手的操作来进行人机交互的方式虽然已经被广泛接受, 但是, 对于手指有残疾的残障人士来讲, 这种通过手的操作向计算机输 入信息和指令的技术显然是不适合的, 不能实现非接触式的人机交互操 作。 虽然也出现了直接用手势手型等进行非接触式人机交互输入的技术 方案, 但是这种技术方案还是需要用手来做出相应的动作, 对于手部有 残疾的残障人士来讲还是不方便的。
本发明实施例提供了一种人机交互操作的触发控制方法和装置。 在 本发明实施例中, 获取摄像装置拍摄的摄像画面, 在显示屏上以虚化方 式显示所述摄像画面; 检测所述摄像画面的帧间差, 根据所述帧间差 识别所述摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显示 屏上的位置; 实时判断所述指定轮廓在显示屏上的位置与显示屏上所 显示的指定区域是否相交, 如果相交, 则触发所述指定区域对应的操 作。 应用本发明实施例, 可以不必用手来触发人机之间的交互操作, 方 便手部有残疾的残障人士通过非接触的方式触发机器操作。
图 1为本发明实施例提供的人机交互操作的触发控制方法的一种流 程示意图。 如图 1所示, 该方法主要包括:
步骤 101 , 获取摄像装置拍摄的摄像画面, 在显示屏上以虚化方式 显示所述摄像画面。
所述虚化方式可以是指定的一种显示方式, 如半透明显示方式, 例 如在显示屏上以半透明方式显示所述摄像画面; 或者可以将所述摄像画 面转化为动画轮靡画面 (例如具有筒单线奈轮靡的动画画面), 该动画 轮廓画面可以叠加到显示屏的原有界面之上, 用户既可以看到显示屏原 有界面又可以看到该动画轮廓画面, 从而方便用户移动摄像画面以进行 后续操作。 下面实施例中, 以在显示屏上以半透明方式显示所述摄像画 面为例进行说明。
步骤 102, 检测所述摄像画面的帧间差, 根据所述帧间差识别所述 摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的位 置。
步骤 103, 实时判断该指定轮廓在显示屏上的位置与显示屏上所显 示的指定区域是否相交,如果相交,执行步骤 104; 否则,返回步骤 102。
步骤 104, 触发该指定区域对应的操作。
本发明实施例中, 所述识别摄像画面上的指定轮廓, 可以是识别人 体某一器官的轮廓, 也可以是识别其它的图形轮廓。 通常情况下, 安装 在计算机等设备上的摄像装置(如摄像头)通常是对用户的头部进行拍 摄, 例如用户在利用视频聊天工具进行视频聊天时, 摄像装置通常都是 拍摄用户的头部尤其是面部的图像。 因此, 为了方便用户尤其是手部有 残疾的残障人士操作, 在本发明的一个实施例中, 所述识别摄像画面上 的指定轮廓可以是识别人的眼睛的轮廓, 这是因为人眼的轮廓较为标 准, 而且可以进一步通过检测眨眼等运动形态向机器发出进一步的操作 指令。
当然, 所识别的摄像画面上的指定轮廓也可以是人的嘴巴等器官的 轮廓, 甚至还可以是某一指定的标准图形。 例如可以给用户提供预先画 有指定图形的白板, 用户可以将该白板举在摄像头前让摄像头拍摄该白 板上的指定图形 (例如一个轮廓鲜明的橢圓形等), 该指定图形就是本 发明实施例中所要检测的指定轮廓。 当用户移动所述白板, 将显示屏上 所显示的该指定图形的位置与指定的区域(例如可以是一种媒介信息的 显示区域, 或者是指定指令区如按钮、 链接等)相交时, 则触发该指定 区域对应的操作, 这样可以实现不用手来触发人机之间的交互操作, 方 便手部有残疾的残障人士通过非接触的方式触发机器操作的目的。
下面以所述指定轮廓为眼睛轮廓为例对本发明实施例进行说明。 在上述步骤 102中, 所述检测所述摄像画面的帧间差, 根据所述帧 间差识别所述摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显 示屏上的位置的操作, 可以利用编程工具来实现。 例如可以采用开源计 算机视觉库 ( openCV, Open Source Computer Vision Library ) 中的针对 性的接口函数来实现。
OpenCV是一个基于开源发行的跨平台计算机视觉库, 可以运行在 Linux、 Windows和 Mac OS等计算机操作系统上。 OpenCV轻量级而且 高效, 是由一系列 C 函数和少量 C++ 类构成的, 同时提供了 Python、 Ruby, MATLAB 等语言的调用接口, 实现了图像处理和计算机视觉方 面的很多通用的计算方法。
例如在一个实施例中, 可以采用 OpenCV 中的 cvSub接口函数和 cvThreshold接口函数来检测摄像画面的帧间差。例如具体的实现代码指 令如下:
cvSub(gray, prev, diff, NULL);
cvThreshold(diff, diff, 5, 255, CV_THRESH_BINARY);
其中, 所述 gray是摄像画面的当前帧, 所述 prev是当前帧的前一 帧, 所述 diff为帧间差。
例如在一个实施例中, 可以采用 OpenCV中的 cvFindContours接口 函数来根据所述帧间差识别指定轮廓, 例如识别眼睛轮廓。 例如具体的 实现代码指令如下:
int nc = c vFindContours (
diff, /* the difference image */
storage, /* created with cvCreateMemStorage() */ &comp, /* output: connected components */ sizeof(CvContour),
CV_RETR_CCOMP,
CV_CHAIN_APPROX_SIMPLE,
cvPoint(0,0)
);
其中, 所述 diff 为上述计算出的帧间差, 所述 comp为所识别出的 眼睛轮廓, 该眼睛轮廓由 cvFindContours接口函数输出。
例如在一个实施例中, 可以采用 OpenCV中的 cvSetlmageROI接口 函数来计算所识别出的指定轮廓在显示屏上的位置, 例如具体的实现代 码指令如下:
cvSetImageROI(gray, rect_eye);
其中, rect_eye为该接口函数 cvSetlmageROI所输出的眼睛轮廓在 当前帧 gray中的位置, 再 ^据当前帧 gray在显示屏上所占据的位置, 就可以计算出眼睛轮廓在显示屏上的当前位置。 本发明实施例中在显示屏上所显示的指定区域可以有多种形态, 例 如可以为电子媒介信息(本发明实施例中筒称为媒介信息)的显示区域, 也可以为指定的指令区, 例如指定的按钮、 指定的文字链接、 指定的图 片区等等。
图 2a 为在显示屏上所显示的指定区域为指定媒介信息的显示区域 的第一种机器界面示意图。 参见图 2a, 在该机器界面 200上显示有媒介 信息 201以及媒介信息 202。 本发明实施例将摄像装置所拍摄的用户的 头像在该机器界面 200中以半透明方式显示, 这样该用户头像就可以与 该机器界面 200中的信息重叠, 使得用户既可以看清界面 200中的各种 信息, 又可以看到本人的头像, 从而可以在移动自己的头部的同时观察 眼睛轮廓的移动, 使得眼睛轮廓能够移动到指定区域, 即媒介信息 201 或媒介信息 202的显示区域。 图 2b为在显示屏上所显示的指定区域为 指定媒介信息的显示区域的第二种机器界面示意图。 参见图 2b, 当用户 的眼睛轮廓 203移动到指定的媒介信息 201的显示区域,即眼睛轮廓 203 在显示屏上的位置与该媒介信息 201的显示区域相交时, 则触发该指定 媒介信息 201的显示区域对应的操作。
在一个实施例中, 当所述指定轮廓如眼睛轮廓 203在显示屏上的位 置与显示屏上所显示的指定区域如媒介信息 201的显示区域相交时, 触 发的该指定媒介信息 201的显示区域对应的操作包括: 记录所述眼睛轮 廓 203与该指定媒介信息 201的显示区域的相交时间, 判断所述眼睛轮 廓 203在显示屏上的位置是否移出该指定媒介信息 201的显示区域, 如 果是则停止记录所述相交时间, 否则继续记录所述相交时间。 这样, 基 于该相交时间可以计算出用户对该媒介信息 201的关注程度, 根据这个 关注程度可以进一步进行其它的相关操作, 例如计费操作, 即: 根据所 记录的所述指定轮廓 (如眼睛轮廓 203 )与所述指定媒介信息 201的显 示区域的相交时间, 计算该指定媒介信息 201对应的计费信息。
一般地, 对于网络上展示的媒介信息(如网络广告就是一种媒介信 息), 都是基于用户对该媒介信息的点击和该媒介信息的曝光次数进行 计费, 而不是基于浏览者观察该媒介信息的时长来计费。 而应用本发明 实施例的上述处理步骤, 可以计算出用户眼睛与指定媒介信息的相交时 间, 这就相当于该用户对该媒介信息的关注程度, 并可以此为基础实现 新的对媒介信息进行计费的方式, 例如, 仍以图 2b 为例, 当浏览者移 动自己的头像使在显示屏幕上成像的该浏览者自己的眼睛轮廓 203与所 述指定媒介信息 201的显示区域相交后开始计时, 当所述眼睛轮廓 203 在显示屏上的位置移出该指定媒介信息 201的显示区域时就停止计时, 如果所述相交时间大于某个预定的秒数,则开始对该媒介信息 201计费, 从而可以实现基于用户对媒介信息的关注程度对媒介信息进行计费, 使 得计费方式更加细化和精确。
在另一个实施例中, 当所述指定轮廓在显示屏上的位置与显示屏上 所显示的指定区域相交时, 触发的该指定媒介信息显示区域对应的操作 包括: 检测所述指定轮廓是否发生指定的运动形态, 如果是则触发该指 定区域所绑定的指令操作。 例如, 检测所述眼睛轮廓 203是否发生眨眼 动作, 如果是则触发该指定区域所绑定的指令操作。 例如所述指定的媒 介信息 201绑定的指令操作为点击动作, 那么当用户眨眼之后, 就可以 触发对所述媒介信息 201的点击动作, 从而打开该媒介信息 201所指向 的网络页面。 图 3a 为在显示屏上所显示的指定区域为指定指令区的第一种机器 界面示意图。 参见图 3a, 在该机器界面 200上显示有媒介信息 201以及 媒介信息 202, 所述媒介信息 201上还有指定指令区, 如 "换一个" 按 钮 301和 "关闭" 按钮 302都是指定指令区。 所述 "换一个" 按钮 301 绑定的指令操作为切换到下一条媒介信息, 所述 "关闭" 按钮 302所绑 定的指令操作为关闭当前的媒介信息 201。 本发明实施例将摄像装置所 拍摄的用户的头像在该机器界面 200中以半透明方式显示, 这样该用户 头像就可以与该机器界面 200中的信息重叠, 使得用户既可以看清界面 200 中的各种信息, 又可以看到本人的头像, 从而可以在移动自己的头 部的同时观察眼睛轮廓的移动, 使得眼睛轮廓能够移动到指定指令区。 图 3b 为在显示屏上所显示的指定区域为指定指令区的第二种机器 界面示意图, 当用户的眼睛轮廓 203移动到 "换一个" 按钮 301 , 即眼 睛轮廓 203与 "换一个"按钮 301相交时,则可以检测所述眼睛轮廓 203 是否发生指定的运动形态 (如眨眼动作), 如果是则触发该 "换一个" 按钮 301所绑定的指令操作, 即将媒介信息显示区域所显示的当前媒介 信息 201切换为下一条媒介信息。 当用户的眼睛轮廓 203移动到 "关闭" 按钮 302, 即眼睛轮廓 203与 "关闭" 按钮 302相交时, 则可以检测所 述眼睛轮廓 203 是否发生指定的运动形态 (如眨眼动作), 如果是则触 发该 "关闭"按钮 302所绑定的指令操作, 即关闭当前的媒介信息 201。
当然, 在其它实施例中, 当所述指定轮廓是其它的图像轮廓时, 所 述指定的运动形态可以是该图像轮廓对应的动作。 例如当所述指定轮廓 为人的嘴巴的轮廓时, 所述指定的运动形态可以为嘴巴的张开和闭合动 作。 在一个实施例中, 所述检测所述指定轮廓是否发生指定的运动形 态, 具体包括:
首先, 创建所述指定轮廓的模板; 例如在一个实施例中, 可以采用
OpenCV中的 cvResetlmageROI(gray)接口函数来创建眼睛模板。
然后, 检测所述指定轮廓的模板(如眼睛模板) 内的帧图像, 判断 所述帧图像的变化是否符合指定的运动形态; 如果符合则触发所述指定 区域所绑定的指令操作。
例如, 当所述指定轮廓为眼睛轮廓时, 所述检测指定轮廓是否发生 指定的运动形态具体可以为: 检测所述眼睛轮廓是否发生眨眼动作。
所述检测所述眼睛轮廓是否发生眨眼动作的具体方法可包括: 检测 眼睛轮廓的边界值; 检测所述边界值的最大值和最小值; 检测所述边界 值的最大值和最小值之间的距离是否发生由大到小再由小到大的变化 过程, 如果是则判定所述眼睛轮廓发生眨眼动作。
例如在一个实施例中, 可以采用 OpenCV中的相关接口参数来判断 眼睛轮廓是否发生眨眼动作。 具体方法可包括如下步骤 411~413:
步骤 411、根据 cvMatchTemplate接口函数检测眼睛轮廓的边界。具 体的代码指令例如如下:
cvMatchTemplate(img, tpl, tm, CV—TM—CCOEFF—NORMED);
其中, tpl为所述 cvResetlmageROI(gray)接口函数创建的眼睛模板。 步骤 412、 cvMinMaxLoc接口函数检测所述眼睛轮廓的边界值的最 大值和最小值。 具体的代码指令例如如下:
cvMinMaxLoc(tm, &minval, &maxval, &minloc, &maxloc, 0);
步骤 413、 检测所述眼睛轮廓的边界值的最大值和最小值之间的距 离是否发生由大到小再由小到大的变化过程, 即判断是否发生眼睛的闭 合动作, 如果是则判定所述眼睛轮廓发生眨眼动作。 具体的代码指令例 ^口^口" f:
if (maxval < TE—THRESHOLD)
return 0;
II return the search window
*window = win;
II return eye location
*eye = cvRect(
win.x + maxloc.x,
win.y + maxloc.y,
TPL—WIDTH,
TPL HEIGHT );
if ((maxval > LB—THRESHOLD) && (maxval < UB—THRESHOLD)) return 2; //闭目艮, 即眼睛轮廓边界值的最大值和最小值之间的距离发 生由大到小的变化过程的检测代码指令。
if (maxval > OE—THRESHOLD)
return 1; //睁眼, 即眼睛轮廓边界值的最大值和最小值之间的距离发 生由小到大的变化过程的检测代码指令。
例如在一种具体应用场景中, 当用户利用网络视频即时通信工具进 行聊天时, 本发明实施例可以将摄像头拍摄的画面虚化显示, 例如, 将 摄像头拍摄的用户头部图像在视频聊天画面中虚化显示, 并在视频聊天 画面中展示一个网络广告 (即媒介信息), 该网络广告中可以显示广告 内容, 并可以显示 "换一个" 按钮和 "关闭" 按钮。 当用户移动头部, 将眼睛轮廓移动到 "换一个" 按钮上时, 则可以将该网络广告切换为下 一个网络广告, 当眼睛轮廓移动到 "关闭" 按钮上时, 则可以将该网络 广告关闭。 同时, 可以按照眼睛轮廓与该网络广告的相交时间对该网络 广告进行计费。 与上述方法对应, 本发明实施例还公开了一种人机交互操作的触发 控制装置, 以执行上述方法。 所述装置可由计算机设备实现, 所述计算 机设备可以是个人计算机, 也可以是服务器, 还可以是便携计算机(如 笔记本电脑、 平板电脑等)。 所述计算机设备可包括至少一个处理器和 存储器。 所述存储器存储有机器可读指令, 当所述机器可读指令被所述 处理器执行时, 可实现上述本发明各方法实施例中的全部或部分流程。 其中, 所述的存储器可为磁碟、 光盘、 只读存储记忆体 (Read-Only Memory, ROM )或随机存储记忆体( Random Access Memory, RAM ) 等。 图 4为本发明实施例提供的所述人机交互操作的触发控制装置的一 种组成示意图。 参见图 4, 该装置可包括:
第一模块 401 , 用于获取摄像装置拍摄的摄像画面, 在显示屏上以 半透明方式显示所述摄像画面;
第二模块 402, 用于检测所述摄像画面的帧间差, 根据所述帧间差 识别所述摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显示屏 上的位置;
第三模块 403 , 用于实时判断该指定轮廓在显示屏上的位置与显示 屏上所显示的指定区域是否相交, 如果相交, 则触发该指定区域对应的 操作。
在一个实施例中, 所述第三模块 403具体用于: 实时判断所述指定 轮廓在显示屏上的位置与显示屏上所显示的指定区域是否相交, 如果相 交, 则触发该指定区域对应的操作, 其中, 所述触发的操作包括: 记录 所述指定轮廓与该指定区域的相交时间, 判断所述指定轮廓在显示屏上 的位置是否移出该指定区域, 如果是则停止记录所述相交时间, 否则继 续记录所述相交时间。 还可以进一步根据所述相交时间对所述指定区域 进行计费操作。
在另一个实施例中, 所述第三模块 403具体用于: 实时判断所述指 定轮廓在显示屏上的位置与显示屏上所显示的指定区域是否相交, 如果 相交, 则触发该指定区域对应的操作, 其中, 所述触发的操作包括: 检 测所述指定轮廓是否发生指定的运动形态, 如果是则触发该指定区域所 绑定的指令操作。
在再一个实施例中, 所述指定轮廓为眼睛轮廓; 所述第三模块 403 检测指定轮廓是否发生指定的运动形态, 具体为: 检测所述眼睛轮廓是 否发生眨眼动作。 当然, 所述指定轮廓也可以是其它的图像轮廓, 例如人的嘴巴的轮 廓, 此时, 所述指定的运动形态可以为嘴巴的张开和闭合动作。
在又一个实施例中, 所述显示屏上所显示的指定区域可以为指定媒 介信息的显示区域, 或者所述显示屏上所显示的指定区域可以为指定指 令区, 或者可以为其它的指定显示形式区域。
上述本发明实施例中的各个模块可以由软件实现(例如存储在计算 机可读取介质中的由处理器执行的机器可读指令), 也可以由硬件实现
(例如专用集成电路(Application Specific Integrated Circuit, ASIC ) 的 处理器), 或者由软件和硬件的结合实现, 本发明实施例不作具体限定。
上述本发明实施例中的各个模块既可以集成于一体, 也可以分离部 署; 既可以合并为一个模块, 也可以进一步拆分成多个子模块。
图 5为本发明实施例提出的人机交互操作的触发控制装置的硬件结 构示意图。 如图 5所示, 该装置可包括: 处理器 51 , 存储器 52, 至少 一个端口 53以及互联机构 54。 其中, 所述处理器 51和存储器 52通过 互联机构 54互联。所述装置可通过端口 53接收和发送数据信息。其中, 所述存储器 52存储有机器可读指令;
所述处理器 51执行所述机器可读指令来执行以下操作:
获取摄像装置拍摄的摄像画面, 在显示屏上以虚化方式显示所述摄 像画面;
检测所述摄像画面的帧间差, 根据所述帧间差识别所述摄像画面上 的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的位置;
实时判断所述指定轮廓在显示屏上的位置与显示屏上所显示的指 定区域是否相交, 如果相交, 则触发所述指定区域对应的操作。
在本发明实施例中, 所述处理器 51可执行存储在存储器 52中的机 器可读指令来进一步执行前述方法实施例中的全部或部分流程, 在此不 再赘述。
由此可以看出, 当存储在存储器 52中的机器可读指令被处理器 51 执行时, 可实现前述的第一模块 401 , 第二模块 402, 以及第三模块 403 的功能。
由此可以看出, 本发明实施例通过获取摄像装置拍摄的摄像画面, 并在显示屏上以虚化方式, 例如半透明方式, 显示所述摄像画面, 从而 使摄像画面可以和显示屏上显示的其它界面相重叠, 并且可以识别出摄 像画面上的指定轮廓 (例如人的眼睛、 人的嘴巴等器官的轮廓), 用户 可以通过移动身体来控制所述摄像画面中指定轮廓的移动, 当该指定轮 廓与显示屏上所显示的指定区域(例如可以是一种媒介信息的显示区 域, 或者是指定指令区如按钮、 链接等)相交时, 则触发该指定区域对 应的操作。 因此, 可以实现不用手来触发人机之间的交互操作, 方便手 部有残疾的残障人士通过非接触的方式触发机器操作。
需要说明的是, 上述各流程和各结构图中不是所有的步骤和模块都 是必须的, 可以根据实际的需要忽略某些步骤或模块。 各步骤的执行顺 序不是固定的, 可以根据需要进行调整。 各模块的划分仅仅是为了便于 描述而采用的功能上的划分, 实际实现时, 一个模块可以由多个模块实 现, 多个模块的功能也可以由同一个模块实现, 这些模块可以位于同一 个设备中, 也可以位于不同的设备中。
本发明各实施例中的硬件模块可以以机械方式或电子方式实现。 例 处理器, 如 FPGA或 ASIC )用于完成特定的操作。 硬件模块也可以包 括由软件临时配置的可编程逻辑器件或电路(如包括通用处理器或其它 可编程处理器)用于执行特定操作。 至于具体采用机械方式, 或是采用 专用的永久性电路, 或是采用临时配置的电路(如由软件进行配置)来 实现硬件模块 , 可以根据成本和时间上的考虑来决定。
通过以上的实施例的描述, 本领域的技术人员可以清楚地了解到本 发明可借助软件加必需的通用硬件平台的方式来实现, 即通过机器可读 指令来指令相关的硬件来实现, 当然也可以通过硬件, 但很多情况下前 者是更佳的实施方式。 基于这样的理解, 本发明的技术方案本质上或者 说对现有技术做出贡献的部分可以以软件产品的形式体现出来, 该计算 机软件产品存储在一个存储介质中, 包括若干指令用以使得一台终端设 备(可以是手机, 个人计算机, 服务器, 或者网络设备等)执行本发明 各个实施例所述的方法。
本领域普通技术人员可以理解, 实现上述方法实施例中的全部或部 分流程, 是可以通过机器可读指令来指令相关的硬件模块来完成的, 所 述的机器可读指令可存储于一计算机可读取存储介质中。 当执行所述机 器可读指令时, 可实现如上述各方法的实施例的流程。 其中, 所述的存 储介质可为磁碟、 光盘、 只读存储记忆体(Read-Only Memory, ROM ) 或随机存储记忆体(Random Access Memory, RAM )等。 可选择地, 可 以由通信网络从服务器计算机上下载机器可读指令。
以上所述为本发明的实施例, 并不用以限制本发明, 凡在本发明的 精神和原则之内, 所做的任何修改、 等同替换、 改进等, 均应包含在本 发明保护的范围之内。 本发明权利要求的范围不应局限于以上描述的例 子中的实施方式, 而应当将说明书作为一个整体并给予最宽泛的解释。

Claims

权利要求书
1、 一种人机交互操作的触发控制方法, 其特征在于, 包括: 获取摄像装置拍摄的摄像画面, 在显示屏上以虚化方式显示所述摄 像画面;
检测所述摄像画面的帧间差, 根据所述帧间差识别所述摄像画面上 的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的位置;
实时判断所述指定轮廓在显示屏上的位置与显示屏上所显示的指 定区域是否相交, 如果相交, 则触发所述指定区域对应的操作。
2、 根据权利要求 1所述的方法, 其特征在于,
所述指定轮廓在显示屏上的位置与显示屏上所显示的指定区域相 交时, 触发的该指定区域对应的操作包括: 记录所述指定轮廓与所述指 定区域的相交时间, 判断所述指定轮廓在显示屏上的位置是否移出该指 定区域,如果是则停止记录所述相交时间,否则继续记录所述相交时间。
3、 根据权利要求 2所述的方法, 其特征在于, 该方法进一步包括: 根据所记录的所述指定轮廓与所述指定区域的相交时间, 计算所述 指定区域对应的计费信息。
4、 根据权利要求 1所述的方法, 其特征在于,
所述指定轮廓在显示屏上的位置与显示屏上所显示的指定区域相 交时, 触发的该指定区域对应的操作包括: 检测所述指定轮廓是否发生 指定的运动形态, 如果是则触发该指定区域所绑定的指令操作。
5、 根据权利要求 4 所述的方法, 其特征在于, 所述检测所述指定 轮廓是否发生指定的运动形态, 包括:
创建所述指定轮廓的模板;
检测所述指定轮廓的模板内的帧图像,
判断所述帧图像的变化是否符合指定的运动形态; 如果符合则触发 所述指定区域所绑定的指令操作。
6、 根据权利要求 4或 5所述的方法, 其特征在于,
所述指定轮廓为眼睛轮廓;
所述检测指定轮廓是否发生指定的运动形态, 包括: 检测所述眼睛 轮廓是否发生眨眼动作。
7、 根据权利要求 6 所述的方法, 其特征在于, 所述检测眼睛轮廓 是否发生眨眼动作, 包括:
检测眼睛轮廓的边界值;
检测所述边界值的最大值和最小值;
检测所述边界值的最大值和最小值之间的距离是否发生由大到小 再由小到大的变化过程, 如果是则判定所述眼睛轮廓发生眨眼动作。
8、 根据权利要求 1~5任一项所述的方法, 其特征在于,
所述显示屏上所显示的指定区域为指定媒介信息的显示区域, 或者 为指定指令区。
9、 一种人机交互操作的触发控制装置, 其特征在于, 该装置包括: 第一模块, 用于获取摄像装置拍摄的摄像画面, 在显示屏上以虚化 方式显示所述摄像画面;
第二模块, 用于检测所述摄像画面的帧间差, 根据所述帧间差识别 所述摄像画面上的指定轮廓, 并计算所识别出的指定轮廓在显示屏上的 位置;
第三模块, 用于实时判断所述指定轮廓在显示屏上的位置与显示屏 上所显示的指定区域是否相交, 如果相交, 则触发所述指定区域对应的 操作。
10、 根据权利要求 9所述的装置, 其特征在于, 所述第三模块用于: 在判断所述指定轮廓在显示屏上的位置与显示 屏上所显示的指定区域相交时, 触发所述指定区域对应的操作; 其中, 所述触发的操作包括: 记录所述指定轮廓与所述指定区域的相交时间, 判断所述指定轮廓在显示屏上的位置是否移出该指定区域, 如果是则停 止记录所述相交时间, 否则继续记录所述相交时间。
11、 根据权利要求 9所述的装置, 其特征在于,
所述第三模块用于: 在判断所述指定轮廓在显示屏上的位置与显示 屏上所显示的指定区域相交时, 触发所述指定区域对应的操作; 其中, 所述触发的操作包括: 检测所述指定轮廓是否发生指定的运动形态, 如 果是则触发该指定区域所绑定的指令操作。
12、 根据权利要求 11所述的装置, 其特征在于,
所述指定轮廓为眼睛轮廓;
所述第三模块还用于检测所述眼睛轮廓是否发生眨眼动作。
13、 根据权利要求 9~12任一项所述的装置, 其特征在于, 所述显示屏上所显示的指定区域为指定媒介信息的显示区域, 或者 为指定指令区。
PCT/CN2013/087811 2012-12-28 2013-11-26 人机交互操作的触发控制方法和装置 Ceased WO2014101606A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US14/750,697 US9829974B2 (en) 2012-12-28 2015-06-25 Method for controlling triggering of human-computer interaction operation and apparatus thereof

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210583819.6 2012-12-28
CN201210583819.6A CN103902192A (zh) 2012-12-28 2012-12-28 人机交互操作的触发控制方法和装置

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US14/750,697 Continuation US9829974B2 (en) 2012-12-28 2015-06-25 Method for controlling triggering of human-computer interaction operation and apparatus thereof

Publications (1)

Publication Number Publication Date
WO2014101606A1 true WO2014101606A1 (zh) 2014-07-03

Family

ID=50993543

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2013/087811 Ceased WO2014101606A1 (zh) 2012-12-28 2013-11-26 人机交互操作的触发控制方法和装置

Country Status (3)

Country Link
US (1) US9829974B2 (zh)
CN (1) CN103902192A (zh)
WO (1) WO2014101606A1 (zh)

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9652047B2 (en) * 2015-02-25 2017-05-16 Daqri, Llc Visual gestures for a head mounted device
CN106802789A (zh) * 2015-11-26 2017-06-06 中国电信股份有限公司 作用域相交检测方法和用于作用域相交检测的装置
US10387719B2 (en) * 2016-05-20 2019-08-20 Daqri, Llc Biometric based false input detection for a wearable computing device
CN107450896B (zh) * 2016-06-01 2021-11-30 上海东方传媒技术有限公司 一种使用OpenCV显示图像的方法
CN107220973A (zh) * 2017-06-29 2017-09-29 浙江中烟工业有限责任公司 基于Python+OpenCV的六边形中空滤棒快速检测方法
CN114867406A (zh) * 2019-12-19 2022-08-05 赛诺菲 眼睛跟踪装置和方法
CN112330741B (zh) * 2020-11-02 2024-10-18 重庆金山医疗技术研究院有限公司 一种图像显示方法、装置、设备和介质
CN115334222A (zh) * 2022-08-16 2022-11-11 上海研鼎信息技术有限公司 一种基于触发器的摄像头的拍摄控制系统

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101344919A (zh) * 2008-08-05 2009-01-14 华南理工大学 视线跟踪方法及应用该方法的残疾人辅助系统
CN102200881A (zh) * 2010-03-24 2011-09-28 索尼公司 图像处理装置、图像处理方法以及程序
CN102830797A (zh) * 2012-07-26 2012-12-19 深圳先进技术研究院 一种基于视线判断的人机交互方法及系统

Family Cites Families (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
DE69634913T2 (de) * 1995-04-28 2006-01-05 Matsushita Electric Industrial Co., Ltd., Kadoma Schnittstellenvorrichtung
JP2001509627A (ja) * 1997-07-09 2001-07-24 シーメンス アクチエンゲゼルシヤフト 人物の少なくとも1つの特徴領域の位置を検出し、この位置に依存する機器に対する制御信号を発生するための装置を動作するための方法
US6637883B1 (en) * 2003-01-23 2003-10-28 Vishwas V. Tengshe Gaze tracking system and method
US8292433B2 (en) * 2003-03-21 2012-10-23 Queen's University At Kingston Method and apparatus for communication between humans and devices
US8726194B2 (en) * 2007-07-27 2014-05-13 Qualcomm Incorporated Item selection using enhanced control
KR101141087B1 (ko) 2007-09-14 2012-07-12 인텔렉츄얼 벤처스 홀딩 67 엘엘씨 제스처-기반 사용자 상호작용의 프로세싱
CN101291364B (zh) * 2008-05-30 2011-04-27 华为终端有限公司 一种移动通信终端的交互方法、装置及移动通信终端
JP2011203823A (ja) * 2010-03-24 2011-10-13 Sony Corp 画像処理装置、画像処理方法及びプログラム
CN101872244B (zh) * 2010-06-25 2011-12-21 中国科学院软件研究所 一种基于用户手部运动与颜色信息的人机交互方法
US9141189B2 (en) * 2010-08-26 2015-09-22 Samsung Electronics Co., Ltd. Apparatus and method for controlling interface
US8643680B2 (en) * 2011-04-08 2014-02-04 Amazon Technologies, Inc. Gaze-based content display
CN102375542B (zh) * 2011-10-27 2015-02-11 Tcl集团股份有限公司 一种肢体遥控电视的方法及电视遥控装置
US10025381B2 (en) * 2012-01-04 2018-07-17 Tobii Ab System for gaze interaction
US9179833B2 (en) * 2013-02-28 2015-11-10 Carl Zeiss Meditec, Inc. Systems and methods for improved ease and accuracy of gaze tracking

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101344919A (zh) * 2008-08-05 2009-01-14 华南理工大学 视线跟踪方法及应用该方法的残疾人辅助系统
CN102200881A (zh) * 2010-03-24 2011-09-28 索尼公司 图像处理装置、图像处理方法以及程序
CN102830797A (zh) * 2012-07-26 2012-12-19 深圳先进技术研究院 一种基于视线判断的人机交互方法及系统

Also Published As

Publication number Publication date
CN103902192A (zh) 2014-07-02
US9829974B2 (en) 2017-11-28
US20150293589A1 (en) 2015-10-15

Similar Documents

Publication Publication Date Title
WO2014101606A1 (zh) 人机交互操作的触发控制方法和装置
Varona et al. Hands-free vision-based interface for computer accessibility
CN107430429B (zh) 化身键盘
US20180224948A1 (en) Controlling a computing-based device using gestures
CN103353935B (zh) 一种用于智能家居系统的3d动态手势识别方法
CN106462242B (zh) 使用视线跟踪的用户界面控制
US20210192192A1 (en) Method and apparatus for recognizing facial expression
CN102270035A (zh) 以非触摸方式来选择和操作对象的设备和方法
WO2021245823A1 (ja) 情報取得装置、情報取得方法及び記憶媒体
KR102431386B1 (ko) 손 제스처 인식에 기초한 인터랙션 홀로그램 디스플레이 방법 및 시스템
CN112190921A (zh) 一种游戏交互方法及装置
CN109376621A (zh) 一种样本数据生成方法、装置以及机器人
Elleuch et al. Smart tablet monitoring by a real-time head movement and eye gestures recognition system
bin Mohd Sidik et al. A study on natural interaction for human body motion using depth image data
Manresa-Yee et al. Towards hands-free interfaces based on real-time robust facial gesture recognition
US11250242B2 (en) Eye tracking method and user terminal performing same
CN118585068A (zh) 一种基于眼动追踪的增强现实交互方法及系统
Cruz Bautista et al. Hand features extractor using hand contour–a case study
Chaudhary Finger-stylus for non touch-enable systems
Islam et al. Developing a novel hands-free interaction technique based on nose and teeth movements for using mobile devices
Srinivas et al. Virtual Mouse Control Using Hand Gesture Recognition
KR20220095752A (ko) 시선 추적을 이용한 자동 스크롤링 장치
Buddhika et al. Smart photo editor for differently-abled people using assistive technology
Shin et al. Welfare interface using multiple facial features tracking
Wensheng et al. Implementation of virtual mouse based on machine vision

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13869339

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 11/11/2015)

122 Ep: pct application non-entry in european phase

Ref document number: 13869339

Country of ref document: EP

Kind code of ref document: A1