WO2020252008A1 - Fingertip tracking for touchless input device - Google Patents
Fingertip tracking for touchless input device Download PDFInfo
- Publication number
- WO2020252008A1 WO2020252008A1 PCT/US2020/036979 US2020036979W WO2020252008A1 WO 2020252008 A1 WO2020252008 A1 WO 2020252008A1 US 2020036979 W US2020036979 W US 2020036979W WO 2020252008 A1 WO2020252008 A1 WO 2020252008A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- input device
- images
- fingertip
- neural network
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/017—Gesture based interaction, e.g. based on a set of recognized hand gestures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/03—Arrangements for converting the position or the displacement of a member into a coded form
- G06F3/0304—Detection arrangements using opto-electronic means
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/03—Arrangements for converting the position or the displacement of a member into a coded form
- G06F3/041—Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means
- G06F3/042—Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means by opto-electronic means
- G06F3/0425—Digitisers, e.g. for touch screens or touch pads, characterised by the transducing means by opto-electronic means using a single imaging device like a video camera for tracking the absolute position of a single or a plurality of objects with respect to an imaged reference surface, e.g. video camera imaging a display or a projection screen, a table or a wall surface, on which a computer generated image is displayed or projected
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/048—Interaction techniques based on graphical user interfaces [GUI]
- G06F3/0487—Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser
- G06F3/0488—Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures
- G06F3/04883—Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures for inputting data by handwriting, e.g. gesture or text
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/22—Character recognition characterised by the type of writing
- G06V30/226—Character recognition characterised by the type of writing of cursive writing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
- G06V40/28—Recognition of hand or arm movements, e.g. recognition of deaf sign language
Definitions
- the present embodiments relate generally to input devices.
- Touchless input devices may control and/or operate an electronic system without any physical contact from a user.
- a touchless input device may provide the user with a more natural or organic user experience.
- a voice-enabled device provides hands-free operation by listening and responding to a user’s voice.
- the user may query the voice-enabled device for information (e.g., recipe, instructions, directions, and the like), to playback media content (e.g., music, videos, audiobooks, and the like), or to control various devices in the user’s home or office environment (e.g., lights, thermostats, garage doors, and other home automation devices).
- verbal commands may not be suitable for certain types of user input such as, for example, navigating a graphical user interface (GUI).
- GUI graphical user interface
- a method and apparatus for fingertip tracking performed by a touchless input device is disclosed.
- One innovative aspect of the subject matter of this disclosure can be implemented in a method of processing user inputs by the input device.
- the method may include steps of capturing a plurality of images of a scene; detecting a user in the plurality of images using one or more first neural network models; detecting a fingertip of the user in the plurality of images using one or more second neural network models; tracking a position of the fingertip across the plurality of images; and processing a user input based at least in part on changes in the position of the fingertip across the plurality of images.
- the memory stores instructions that, when executed by the processing system, causes the input device to detect a user in a plurality of images using one or more first neural network models; detect a fingertip of the user in the plurality of images using one or more second neural network models; track a position of the fingertip across the plurality of images; and process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
- an input device including a camera configured to capture a plurality of images of a scene, a display configured to provide a GUI, and a processing system.
- the processing system is configured to detect a user in the plurality of images using one or more first neural network models; detect a fingertip of the user in the plurality of images using one or more second neural network models; track a position of the fingertip across the plurality of images; and update the GUI based at least in part on changes in the position of the fingertip across the plurality of images.
- FIG. 1 shows an example system within which the present embodiments may be implemented.
- FIG. 2 shows a block diagram of an input device, in accordance with some embodiments.
- FIG. 3 shows an example operation of a text conversion module, in accordance with some embodiments.
- FIG. 4 shows an example environment in which the present embodiments may be implemented.
- FIG. 5 shows another block diagram of an input device, in accordance with some embodiments.
- FIG. 6 shows an illustrative flowchart depicting an example operation for processing user inputs, in accordance with some embodiments.
- circuit elements or software blocks may be shown as buses or as single signal lines.
- Each of the buses may alternatively be a single signal line, and each of the single signal lines may alternatively be buses, and a single line or bus may represent any one or more of a myriad of physical or logical mechanisms for communication between components.
- the techniques described herein may be implemented in hardware, software, firmware, or any combination thereof, unless specifically described as being implemented in a specific manner. Any features described as modules or components may also be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a non-transitory computer-readable storage medium comprising instructions that, when executed, performs one or more of the methods described above.
- the non-transitory computer-readable storage medium may form part of a computer program product, which may include packaging materials.
- the non-transitory processor-readable storage medium may comprise random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), FLASH memory, other known storage media, and the like.
- RAM synchronous dynamic random access memory
- ROM read only memory
- NVRAM non-volatile random access memory
- EEPROM electrically erasable programmable read only memory
- FLASH memory other known storage media, and the like.
- the techniques additionally, or alternatively, may be realized at least in part by a processor-readable communication medium that carries or communicates code in the form of instructions or data structures and that can be accessed, read, and/or executed by a computer or other processor.
- processors may refer to any general-purpose processor, conventional processor, controller, microcontroller, and/or state machine capable of executing scripts or instructions of one or more software programs stored in memory.
- input device may refer to any device capable of receiving and/or processing user inputs. In some aspects, an input device may be a complete electronic system. In some other aspects, an input device may process user inputs on behalf of a separate electronic system.
- Example input devices may include, but are not limited to, personal computing devices (e.g., desktop computers, laptop computers, tablets, web browsers, and personal digital assistants (PDAs)), composite input devices (e.g., physical keyboards, joysticks, and key switches), data input devices (e.g., remote controls and mice), data output devices (e.g., display screens and printers), remote terminals, kiosks, video game machines (e.g., video game consoles, portable gaming devices, and the like), communication devices (e.g., cellular phones such as smart phones), and media devices (e.g., recorders, editors, and players such as televisions, set-top boxes, music players, digital photo frames, and digital cameras).
- PDAs personal digital assistants
- composite input devices e.g., physical keyboards, joysticks, and key switches
- data input devices e.g., remote controls and mice
- data output devices e.g., display screens and printers
- remote terminals e.g., kiosks
- FIG. 1 shows an example system 100 within which the present embodiments may be implemented.
- the system 100 includes an input device 1 10 and a deep learning environment 120.
- the input device 1 10 may provide a touchless user interface configured to detect and respond to user motion. More specifically, the input device 1 10 may process motion-based instructions and/or queries without any physical contact from the user.
- the input device 1 10 includes one or more sensors 1 12, an input analysis module 1 14, and a user interface 1 16.
- the sensor 1 12 may be configured to receive user inputs and/or collect data (e.g., images, video, audio recordings, and the like) about the surrounding environment.
- Example suitable sensors include, but are not limited to: cameras, capacitive sensors, microphones, and the like.
- at least one of the sensors 1 12 e.g., a camera
- Example motion inputs may include various poses and/or movements that a user can perform with the user’s arms, legs, hands, feet, head, torso, or any other part of the user’s body.
- the input analysis module 1 14 is configured to process the received motion inputs 102. For example, the input analysis module 1 14 may convert the detected motion to one or more user inputs that can be used to control and/or operate the input device 1 10.
- Example user inputs include, but are not limited to, selection inputs (e.g., for selecting content to be displayed or played back on the input device 1 10 or an electronic system coupled to the input device 1 10), control inputs (e.g., for changing an operation mode of the input device 1 10 or an electronic system coupled to the input device 1 10), and data inputs (e.g., for entering information and/or search queries). Different types of movements may be associated with different user inputs depending on the application and/or operation mode of the input device 1 10.
- the input analysis module 1 14 may detect the motion inputs 102 from raw sensor data (e.g., images) captured by one or more of the sensors 1 12. For example, the input analysis module 1 14 may first detect the presence of a user 101 in the one or more images captured via the sensors 1 12. In some aspects, the input analysis module 1 14 may detect the user 101 based on the user’s face, and/or other body parts. Upon identifying the user 101 , the input analysis module 1 14 may then track the movements of one or more body parts (e.g., head, hands, fingers, feet and the like) across one or more images or video frames captured by the sensors 1 12. The input analysis module 1 14 may further compare the tracked movements to one or more predefined and/or recognizable motions (e.g., hand waves, finger points, finger swipes, and the like) to identify a corresponding motion input 102.
- predefined and/or recognizable motions e.g., hand waves, finger points, finger swipes, and the like
- the input analysis module 1 14 may use one or more neural network models 104 to detect and/or identify motion inputs 102 from raw sensor data.
- the neural network models 104 may be trained to infer the presence of a user (e.g., human) and/or position of the user’s body parts from an image or frame of video.
- a user e.g., human
- the neural network models 104 may be trained to detect humans based on the presence of a human face (e.g., a combination of facial features) and/or other human-specific features (e.g., hands and fingers) in a received image.
- a human face e.g., a combination of facial features
- other human-specific features e.g., hands and fingers
- one or more of the neural network models 104 may be trained to model the remaining body parts of the user based on the identified feature. For example, a user’s torso is the contiguous region directly below the user’s face, and the user’s limbs are offshoo
- the deep learning environment 120 may be configured to generate the neural network models 104 through deep learning.
- Deep learning is a particular form of machine learning in which the training phase is performed over multiple layers, generating a more abstract set of rules in each successive layer.
- Deep learning architectures are often referred to as artificial neural networks due to the way in which information is processed (e.g., similar to a biological nervous system).
- each layer of the deep learning architecture may be composed of a number of artificial neurons.
- the neurons may be interconnected across the various layers so that input data (e.g., the raw data) may be passed from one layer to another. More specifically, each layer of neurons may perform a different type of transformation on the input data that will ultimately result in the desired output (e.g., the answer).
- the interconnected framework of neurons may be referred to as a neural network model.
- the neural network models 104 may include a set of rules that can be used to describe the body parts and/or features of a user (e.g., head, neck, torso, arms, legs, hands, feet, fingers, toes, and the like).
- the deep learning environment 120 may have access to a large volume of raw data and may be trained to recognize a set of rules (e.g., certain objects, features, and/or other detectable attributes) associated with the raw data. For example, in some aspects, the deep learning environment 120 may be trained to recognize a human. During the training phase, the deep learning environment 120 may process or analyze a large number of images, videos, audio, and/or other media containing a biometric signature of a human. The deep learning environment 120 may also receive an indication that the provided data describes a human (e.g., in the form of user input from a user or operator reviewing the media and/or data or metadata provided with the media). The deep learning environment 120 may then perform statistical analysis on the images, video, audio and/or other media to determine a common set of features associated with humans. In some aspects, the determined features (or rules) may form an artificial neural network spanning multiple layers of abstraction.
- rules e.g., certain objects, features, and/or other detectable attributes
- the deep learning environment 120 may provide the learned set of rules (e.g., as the neural network models 104) to the input device 1 10 for inferencing.
- one or more of the neural network models 104 may be provided to (e.g., and stored on) the input device 1 10 at a device manufacturing stage.
- the input device 1 10 may be pre-loaded with the neural network models 104 prior to being shipped to an end user.
- the input device 1 10 may receive one or more of the neural network models 104 from the deep learning environment 120 at runtime.
- the deep learning environment 120 may be communicatively coupled to the input device 1 10 via a network (e.g., the cloud). Accordingly, the input device 1 10 may receive the neural network models 104 (including updated neural network models) from the deep learning environment 120, over the network, at any time.
- the input analysis module 1 14 may detect and process the motion inputs 102 based on the neural network models 104 provided by the deep learning environment 120. For example, during the inferencing phase, the input analysis module 1 14 may apply the neural network models 104 to the data collected from the sensors 1 12, by traversing the artificial neurons in the artificial network, to generate inferences about the user’s body parts and associated motions. For example, the input analysis module 1 14 may use the neural network models 104 to detect the user’s fingertip, track the motion of the fingertip, and/or convert the fingertip motion to a user input. The input analysis module 1 12 may further instruct the input device 1 10 to perform one or more operations associated with the detected motion input 102. By generating the inferences locally on the input device 1 10, the present embodiments may be used to perform machine learning on biometric data (e.g., images of the user 101 ) in a manner that protects user privacy.
- biometric data e.g., images of the user 101
- the input analysis module 1 14 may use the data collected from the sensors 1 12 to perform additional training on the neural network models 104. For example, the input analysis module 1 14 may refine the neural network models 104 and/or generate new neural network models based on the locally-generated sensor data. In some aspects, the neural network models 104 may be fine-tuned to detect and/or recognize the movements of a particular user. For example, such additional training may be performed based on images and/or videos of the user 101 captured by the sensors 1 12. The additional training may be initiated manually (e.g., using an independent scripted mechanism) or automatically upon detecting sensor data from the sensors 1 12. In another example, the input analysis module 1 14 may use previously-detected user movements to perform additional training on the neural network models 104 (e.g., in a feedback loop).
- the input analysis module 1 14 may provide the updated neural network models to the deep learning environment 120 to further refine the deep learning architecture.
- the deep learning environment 120 may further refine its neural network models 104 based on the sensor data captured by the input device 1 10 (e.g., combined with sensor data captured by various other input devices) without receiving or having access to the raw sensor data.
- the user interface 1 16 may provide an interface through which the user 101 can operate, interact with, or otherwise use the input device 1 10 or an electronic system (not shown for simplicity) coupled to the input device 1 10.
- the user interface 1 16 may display, render, or otherwise manifest the motion input 102 on the input device 1 10.
- the user interface 1 16 may include a graphical user interface (GUI) rendered on a display of the input device 1 10.
- GUI graphical user interface
- the user interface 1 16 may respond to motion inputs 102 by dynamically updating the display, for example, to navigate the GUI, display new content, track cursor movement, and the like.
- the user interface 1 16 may generate a search query based on the detected motion inputs 102.
- the search query may include a string of characters to be searched for or identified.
- the user interface 1 16 may search for content locally on the input device 1 10 based on the search query.
- the user interface 1 16 may transmit the search query to a network resource for further processing.
- the input analysis module 1 14 may generate the string of characters (which make up the search query) based on one or more motions detected in the raw sensor data. For example, the user 101 may draw one or more letters and/or characters in the air with the user’s hand. The input analysis module 1 14 traces the movement of the user’s hand across one or more images captured via the sensors 1 12 and converts the trace to a corresponding text string. In some aspects, the input analysis module 1 14 may use optical character recognition (OCR) techniques to perform the conversion. Aspects of the present disclosure recognize that the accuracy of the conversion may depend on the precision of the motion inputs 102 detected by the input device 1 10.
- OCR optical character recognition
- the input analysis module 1 14 may be configured to detect and track a user’s fingertip in one or more frames of sensor data.
- FIG. 2 shows a block diagram of an input device 200, in accordance with some embodiments.
- the input device 200 may be one embodiment of the input device 1 10 of FIG. 1 .
- the input device 200 may provide a touchless user interface configured to detect and respond to motion inputs that do not involve any physical contact from a user.
- the input device 200 includes a camera 210, a display 250, a network interface (l/F) 260, and a processing system 280 including a neural network 220, an input classifier 230, and a user interface 240.
- the neural network 220 and/or the input classifier 230 may be implemented by neural network acceleration hardware (which may include one or more processors configured to accelerate neural network inferencing).
- the camera 210 is configured to capture one or more images 201 of the environment surrounding the input device 200.
- the camera 210 may be one embodiment of one of the sensors 1 12 of FIG. 1.
- the camera 210 may be configured to capture images 201 (e.g., still-frame images and/or video) of a scene in front of or proximate the input device 200.
- the camera 210 may include optical sensors (e.g., photodiodes, CMOS image sensor arrays, CCD arrays, and/or any other sensors capable of detecting wavelengths of light in the visible spectrum, the infrared spectrum, and/or the ultraviolet spectrum) and/or depth sensors (e.g., time-of-flight sensors, structured light sensors, and/or any other sensors capable of measuring the depth or distances of objects).
- optical sensors e.g., photodiodes, CMOS image sensor arrays, CCD arrays, and/or any other sensors capable of detecting wavelengths of light in the visible spectrum, the infrared spectrum, and/or the ultraviolet spectrum
- depth sensors e.g., time-of-flight sensors, structured light sensors, and/or any other sensors capable of measuring the depth or distances of objects.
- NIR near-infrared
- depth-sensing cameras may capture finer details of a scene than more traditional optical cameras that detect red, green, and blue (RGB) wavelengths of
- the neural network 220 is configured to detect one or more input objects in the captured images 201 .
- the input object may include a user’s fingertip.
- the neural network 220 may identify the location of the user’s fingertips in each of the images 201 as fingerprint data 202.
- the neural network 220 may generate inferences about the user’s fingertips using one or more neural network models. For example, as described with respect to FIG. 1 , the neural network 220 may receive trained neural network models (e.g., from the deep learning environment 120) prior to receiving the images 201 from the camera 210.
- the neural network 220 may include a user detection module 222 and a fingertip detection module 224.
- the user detection module 222 may detect the presence of a human (e.g., user) in one or more of the images 201 . More specifically, the user detection module 222 may implement one or more neural network models to infer the presence of a human based on a combination of facial features and/or other human-specific features. For example, the user detection module 222 may detect the presence of a human face using any known face detection algorithms and/or techniques. In some embodiments, the user detection module 222 may be further configured to model the remaining body parts of the user based, at least in part, on the identified feature. For example, the user detection module 222 may identify the users’ torso, arms, legs, hands, and/or feet based on the relative position of the user’s face.
- a human e.g., user
- the user detection module 222 may implement one or more neural network models to infer the presence of a human based on a combination of facial features and/or other human-specific features.
- the user detection module 222 may detect the presence of a
- the fingertip detection module 224 may detect the location of a user’s fingertip in one or more of the images 201 . More specifically, the fingertip detection module 224 may implement one or more neural network models to infer the presence of fingers and locations of the fingertips from an image of the user’s hands. For example, the neural network models may be trained using images of hands with fingers in various poses and/or stages of movement. The training data may be annotated to identify the fingertip locations in each frame or image. Through training, the neural network models may be configured to detect subtle movements of a user’s fingertip (e.g., while the user’s hand remains relatively stationary) from far distances (such as across a room).
- the neural network models may be trained on images captured of users at multiple distances from an image capture device.
- the fingertip detection module 224 may crop the image 201 around the user’s hand (e.g., as detected by the user detection module 222) and apply the neural network models to the cropped image to detect the locations of the user’s fingertips as fingertip data 202.
- the input classifier 230 is configured to convert the fingertip data 202 to one or more user inputs 203.
- the fingertip data 202 may indicate the locations of the user’s fingertips in one or more images 201 captured by the camera 201 .
- the input classifier 230 may generate the user inputs 203 based on movements of the user’s fingertips. For example, the input classifier 230 may track changes in the locations of the user’s fingertips across the one or more images 201 . In some aspects, the input classifier 230 may continuously track the movements of the user’s fingertips until the movements satisfy a particular condition and/or no movement (or very little movement) is detected for at least a threshold duration.
- the input classifier 230 may include a motion input recognition module 232 and a text conversion module 234.
- the motion input recognition module 232 may compare the movements of the user’s fingertips to one or more predetermined motions to determine the associated user inputs 203.
- Example user inputs 203 include, but are not limited to, selection inputs, control inputs, and data inputs. Different types of movements may be associated with different user inputs depending on the application and/or operation mode of the input device 200. For example, a finger swiping motion may be used to scroll through content presented on the display 250, whereas a finger tapping motion may be used to select a particular content item on the display 250.
- the motion input recognition module 232 may place the input device 200 in a text-input mode in response to detecting a particular motion (such as a user raising his or her hand). While operating in the text-input mode, the input classifier 230 may classify subsequent fingerprint data 202 as text input data.
- the text conversion module 234 may convert the fingertip data 202 to text-based user inputs 203.
- the text conversion module 234 may trace the movements of one of the user’s fingertips (such as the tip of the user’s index finger) as the user draws one or more letters and/or characters in the air.
- the text conversion module 234 may further convert the trace to a character string.
- the text conversion module 234 may use OCR techniques to convert the trace to the character string.
- the text conversion module 234 may implement one or more neural network models to infer the characters to be converted based, at least in part, on samples of the user’s handwriting and/or writing style. With reference for example to FIG.
- the text conversion module 234 may trace a user’s fingertip (FT) 301 as it draws the word“Action” in the air (e.g., as fingertip trace 302). The text conversion module 234 may then convert the fingertip trace 302 to a text-based input string 303 corresponding to the word“Action.”
- FT fingertip
- the text conversion module 234 may process each character of the character string individually. For example, once a user has completed drawing a particular letter or character, the text conversion module 234 may convert the trace to a corresponding character string. With reference for example to FIG. 3, the text conversion module 234 may recognize the letters“A,”“c,”“t,”“i,”“o,” and“n,” individually, before the fingertip trace 302 is complete. In some other aspects, the text conversion module 234 may process the entire character string as a whole. For example, the text conversion module 234 may wait until the user is finished drawing a complete word or sequence of characters before converting the trace to a corresponding character string. With reference for example to FIG. 3, the text conversion module 234 may recognize the word“Action” as a whole, after the fingertip trace 302 is complete.
- the user interface 240 may generate an output 204 in response to the user inputs 203.
- the output 204 may be rendered on the display 250.
- the output 204 may include updates to a GUI presented on the display 250.
- the GUI updates may track the user inputs 203 and/or movements of the user’s fingertips (e.g., cursor movement, text insertion, and the like).
- the output 204 may include content to be presented or played back on the display 250.
- the user inputs 203 may include instructions for searching, selecting, or otherwise retrieving content to be displayed.
- the user interface 240 may include a content store 242 and a content retrieval module 244.
- the content store 242 may store or buffer content that can be rendered on the display 250 and/or a display device (not shown) coupled to the input device 200.
- the content retrieval module 244 may retrieve content from one or more network resources external to the input device 200.
- the content retrieval module 244 may be configured to generate a search query 205 based, at least in part, on the user input 203 received from the input classifier 230.
- the user input 203 may be a text-based user input comprising a string of characters.
- the search query 205 may include at least a portion of the character string included in the user input 203.
- the search query may be transmitted to a network resource (not shown for simplicity) via the network interface 260.
- the network resource may include memory and/or processing resources to generate one or more results 206 for the search query 205.
- the network resource may search one or more networked devices (e.g., the Internet, content delivery networks, and the like) for the content or information requested by the search query 205.
- the network resource may then send the results 206 (e.g., including the requested content or information) back to the input device 200.
- FIG. 4 shows an example environment 400 in which the present embodiments may be implemented.
- the environment 400 includes an input device 410, a user 420, and a non-user object 430.
- the input device 410 may be one embodiment of the input device 1 10 of FIG. 1 and/or input device 200 of FIG. 2.
- the input device 410 is depicted as a media device (e.g., a television) capable of displaying or playing back media content (e.g., images, videos, audio, and the like) to the user 420.
- a media device e.g., a television
- media content e.g., images, videos, audio, and the like
- the input device 410 includes a camera 412 and a display 414.
- the camera 412 may be one embodiment of the camera 210 of FIG. 2.
- the camera 412 may be configured to capture images (e.g., still-frame images and/or video) of a scene 401 in front of the input device 410.
- the display 414 may be one embodiment of the display 250 of FIG. 2.
- the display 414 may correspond to and/or provide a user interface through which the user 420 may interact with or use the input device 410.
- the input device 410 may provide a touchless user interface configured to detect and respond to motion inputs that do not involve any physical contact from the user
- the camera 412 continuously (or periodically) captures images of the scene 401 to detect motion inputs from the user 420.
- the input device 410 may process user inputs based, at least in part, on the movements of a user’s fingertip 421 .
- the input device 410 may first detect the presence of the user 420 in the scene 401 .
- the user 420 is shown sitting on a couch (e.g., non-user object 430) at a comfortable viewing distance from the input device 410.
- the user’s hands are shown to be resting in his or her lap. From this position, the user 420 may make small, subtle motions of the fingertip
- the user 420 may control the input device 410 in a relaxed and unobtrusive manner, allowing for a more robust user experience.
- the input device 410 may use one or more neural network models to infer the user 420 from one or more images captured via the camera 412.
- the neural network models may be trained to differentiate human users from other non-user objects 430 in the scene 401 .
- the input device 410 may generate one or more inferences about the user’s fingertip 421 .
- the input device 410 may further implement one or more neural network models to infer the location or position of the fingertip 421 .
- the neural network models may be trained to identify the user’s fingertip 421 based on the presence and/or location of one or more other body parts of the user 420 (such as the user’s hand).
- the input device 410 may detect one or more motion inputs by tracking the movements of the user’s fingertip 421 across one or more images captured via the camera 412. In some embodiments, the input device 410 may process the motion inputs to control the presentation and/or playback of media content on the display 414. For example, the motion inputs may be used to navigate, select, and/or search for media content to be played back on the display 414. In some aspects, the input device 410 may classify one or more motion inputs as text input data. For example, the input device 410 may trace the movements of the user’s fingertip 421 as the user draws one or more letters and/or characters in the air.
- the input device 410 may further convert the trace to a character string using OCR techniques (e.g., as described with respect to FIGS. 2 and 3). The input device 410 may then retrieve one or more content items associated with the character string (stored locally or on a network resource) for presentation and/or playback on the display 414.
- OCR techniques e.g., as described with respect to FIGS. 2 and 3.
- the accuracy of many OCR techniques depends on the precision of the traces from which the characters are derived (such as the fingertip trace 302 of FIG. 3).
- the input device 410 may detect and process motion inputs with a finer level of precision and granularity than would otherwise be possible by tracking the movements of any other body part.
- the present embodiments may enable and/or improve the detection of touchless text-based motion inputs.
- the fingertip tracking techniques described herein may be able to detect subtle movements of a user’s fingertip from far distances (such as shown in FIG. 4), allowing for an improved user experience.
- FIG. 5 shows another block diagram of an input device 500, in accordance with some embodiments.
- the input device 500 may be one embodiment of the input device 200 of FIG. 2 and/or the input device 1 10 of FIG. 1 .
- the input device 500 includes a sensor interface 510, a processor 520, and a memory 530.
- the sensor interface 510 may be used to communicate with one or more sensors (e.g., cameras) coupled to the input device 500.
- the sensor interface 510 may be configured to communicate with a camera of the input device 500 (e.g., camera 210 of FIG. 2).
- the sensor interface 510 may transmit signals to, and receive signals from, the camera to capture a plurality of images of a scene facing the input device 500.
- the memory 530 may also include a non-transitory computer- readable medium (e.g., one or more nonvolatile memory elements, such as EPROM, EEPROM, Flash memory, a hard drive, etc.) that may store at least the following software (SW) modules:
- SW software
- a fingertip detection SW module 534 to detect a fingertip of the user in the plurality of images using one or more second neural network models
- a fingertip tracking SW module 536 to track a position of the fingertip across the plurality of images
- a user interface SW module 538 to process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
- Each software module includes instructions that, when executed by the processor 520, cause the input device 500 to perform the corresponding functions.
- the non-transitory computer-readable medium of memory 530 thus includes instructions for performing all or a portion of the operations described below with respect to FIG. 6.
- Processor 520 may be any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in memory 530.
- the processor 520 may execute the user detection SW module 532 to detect a user in the plurality of images using one or more first neural network models.
- the processor 520 may further execute the fingertip detection SW module 534 to detect a fingertip of the user in the plurality of images using one or more second neural network models.
- the processor 520 may execute the fingertip tracking SW module 536 to track a position of the fingertip across the plurality of images.
- the processor 520 also may execute the user interface SW module 538 to process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
- FIG. 6 is an illustrative flowchart depicting an example operation 600 for processing user inputs, in accordance with some embodiments.
- the operation 600 may be performed by the input device 1 10 to control or operate a touchless user interface.
- the input device may capture a plurality of images of a scene (610).
- the images may include still-frame images and/or videos captured by a camera in, or coupled to, the input device.
- the camera may include optical sensors and/or depth sensors. In low-light conditions, NIR and/or depth-sensing cameras may capture finer details of a scene than more traditional optical cameras that detect RGB wavelengths of light.
- the input device may use RGB cameras to capture images during the day or in relatively bright conditions, and may use NIR or depth-sensing cameras to capture images at night or in low-light conditions.
- the input device detects a user in the plurality of images using one or more first neural network models (620).
- the input device may implement one or more neural network models to infer the presence of a human based on a combination of facial features and/or other human-specific features.
- the input device may detect the presence of a human face using any known face detection algorithms and/or techniques.
- the input device may further model the remaining body parts of the user based, at least in part, on the identified feature. For example, the input device may identify the users’ torso, arms, legs, hands, and/or feet based on the relative position of the user’s face.
- the input device may further detect a fingertip of the user in the plurality of images using one or more second neural network models (630).
- the input device may implement one or more neural network models to infer the presence of fingers and locations of the fingertips from an image of the user’s hands.
- the neural network models may be trained using images of hands with fingers in various poses and/or stages of movement.
- the training data may be annotated to identify the fingertip locations in each frame or image.
- the neural network models may be configured to detect subtle movements of a user’s fingertip from far distances (such as described with reference to FIG. 4).
- the input device may track a position of the fingertip across the plurality of images (640) and process a user input based at least in part on changes in the position of the fingertip across the plurality of images (650). In some implementations, the input device may continuously track the movements of the user’s fingertips until the movements satisfy a particular condition and/or no movement (or very little movement) is detected for at least a threshold duration. The input device may compare the movements of the user’s fingertips to one or more predetermined motions to determine the associated user inputs.
- Example user inputs include, but are not limited to, selection inputs, control inputs, and data inputs.
- the input device may trace the movements of one of the user’s fingertips (such as the tip of the user’s index finger) as the user draws one or more letters and/or characters in the air.
- the input device may further convert the trace to a character string (such as described with reference to FIG. 3).
- a software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
- An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Human Computer Interaction (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Data Mining & Analysis (AREA)
- Molecular Biology (AREA)
- Computational Linguistics (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Artificial Intelligence (AREA)
- Multimedia (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Social Psychology (AREA)
- Psychiatry (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
A method and apparatus for fingertip tracking performed by a touchless input device. The input device captures a plurality of images of a scene and detects a user in the plurality of images using one or more first neural network models. The input device further detects a fingertip of the user in the plurality of images using one or more second neural network models and tracks a position of the fingertip across the plurality of images. The input device further processes a user input based at least in part on changes in the position of the fingertip across the plurality of images. In some implementations, the input device may process the user inputs by updating a graphical user interface (GUI) based on the user input.
Description
FINGERTIP TRACKING FOR TOUCHLESS INPUT DEVICE
TECHNICAL FIELD
[0001 ] The present embodiments relate generally to input devices.
BACKGROUND OF RELATED ART
[0002] Touchless input devices may control and/or operate an electronic system without any physical contact from a user. In contrast with“tactile” input devices (e.g., touchpads, remote controls, keyboards, mice, etc.), a touchless input device may provide the user with a more natural or organic user experience. For example, a voice-enabled device provides hands-free operation by listening and responding to a user’s voice. The user may query the voice-enabled device for information (e.g., recipe, instructions, directions, and the like), to playback media content (e.g., music, videos, audiobooks, and the like), or to control various devices in the user’s home or office environment (e.g., lights, thermostats, garage doors, and other home automation devices). However, verbal commands may not be suitable for certain types of user input such as, for example, navigating a graphical user interface (GUI).
SUMMARY
[0003] This Summary is provided to introduce in a simplified form a selection of concepts that are further described below in the Detailed
Description. This Summary is not intended to identify key features or essential features of the claims subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0004] A method and apparatus for fingertip tracking performed by a touchless input device is disclosed. One innovative aspect of the subject matter of this disclosure can be implemented in a method of processing user inputs by the input device. In some implementations, the method may include steps of capturing a plurality of images of a scene; detecting a user in the plurality of images using one or more first neural network models; detecting a fingertip of the user in the plurality of images using one or more second neural network
models; tracking a position of the fingertip across the plurality of images; and processing a user input based at least in part on changes in the position of the fingertip across the plurality of images.
[0005] Another innovative aspect of the subject matter of this disclosure can be implemented in an input device including processing circuitry and memory. In some implementations, the memory stores instructions that, when executed by the processing system, causes the input device to detect a user in a plurality of images using one or more first neural network models; detect a fingertip of the user in the plurality of images using one or more second neural network models; track a position of the fingertip across the plurality of images; and process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
[0006] Another innovative aspect of the subject matter of this disclosure can be implemented in an input device including a camera configured to capture a plurality of images of a scene, a display configured to provide a GUI, and a processing system. In some implementations, the processing system is configured to detect a user in the plurality of images using one or more first neural network models; detect a fingertip of the user in the plurality of images using one or more second neural network models; track a position of the fingertip across the plurality of images; and update the GUI based at least in part on changes in the position of the fingertip across the plurality of images.
BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present embodiments are illustrated by way of example and are not intended to be limited by the figures of the accompanying drawings.
[0008] FIG. 1 shows an example system within which the present embodiments may be implemented.
[0009] FIG. 2 shows a block diagram of an input device, in accordance with some embodiments.
[0010] FIG. 3 shows an example operation of a text conversion module, in accordance with some embodiments.
[001 1 ] FIG. 4 shows an example environment in which the present embodiments may be implemented.
[0012] FIG. 5 shows another block diagram of an input device, in accordance with some embodiments.
[0013] FIG. 6 shows an illustrative flowchart depicting an example operation for processing user inputs, in accordance with some embodiments.
DETAILED DESCRIPTION
[0014] In the following description, numerous specific details are set forth such as examples of specific components, circuits, and processes to provide a thorough understanding of the present disclosure. The term“coupled” as used herein means connected directly to or connected through one or more intervening components or circuits. Also, in the following description and for purposes of explanation, specific nomenclature is set forth to provide a thorough understanding of the aspects of the disclosure. However, it will be apparent to one skilled in the art that these specific details may not be required to practice the example embodiments. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the present disclosure. Some portions of the detailed descriptions which follow are presented in terms of procedures, logic blocks, processing and other symbolic representations of operations on data bits within a computer memory. The interconnection between circuit elements or software blocks may be shown as buses or as single signal lines. Each of the buses may alternatively be a single signal line, and each of the single signal lines may alternatively be buses, and a single line or bus may represent any one or more of a myriad of physical or logical mechanisms for communication between components.
[0015] Unless specifically stated otherwise as apparent from the following discussions, it is appreciated that throughout the present application, discussions utilizing the terms such as“accessing,”“receiving,”“sending,” “using,”“selecting,”“determining,”“normalizing,”“multiplying,”“averaging,” “monitoring,”“comparing,”“applying,”“updating,”“measuring,”“deriving” or the like, refer to the actions and processes of a computer system, or similar
electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.
[0016] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof, unless specifically described as being implemented in a specific manner. Any features described as modules or components may also be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a non-transitory computer-readable storage medium comprising instructions that, when executed, performs one or more of the methods described above. The non-transitory computer-readable storage medium may form part of a computer program product, which may include packaging materials.
[0017] The non-transitory processor-readable storage medium may comprise random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), FLASH memory, other known storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a processor-readable communication medium that carries or communicates code in the form of instructions or data structures and that can be accessed, read, and/or executed by a computer or other processor.
[0018] The various illustrative logical blocks, modules, circuits and instructions described in connection with the embodiments disclosed herein may be executed by one or more processors. The term“processor,” as used herein, may refer to any general-purpose processor, conventional processor, controller, microcontroller, and/or state machine capable of executing scripts or instructions of one or more software programs stored in memory. The term “input device,” as used herein, may refer to any device capable of receiving and/or processing user inputs. In some aspects, an input device may be a complete electronic system. In some other aspects, an input device may
process user inputs on behalf of a separate electronic system. Example input devices may include, but are not limited to, personal computing devices (e.g., desktop computers, laptop computers, tablets, web browsers, and personal digital assistants (PDAs)), composite input devices (e.g., physical keyboards, joysticks, and key switches), data input devices (e.g., remote controls and mice), data output devices (e.g., display screens and printers), remote terminals, kiosks, video game machines (e.g., video game consoles, portable gaming devices, and the like), communication devices (e.g., cellular phones such as smart phones), and media devices (e.g., recorders, editors, and players such as televisions, set-top boxes, music players, digital photo frames, and digital cameras).
[0019] FIG. 1 shows an example system 100 within which the present embodiments may be implemented. The system 100 includes an input device 1 10 and a deep learning environment 120. In some embodiments, the input device 1 10 may provide a touchless user interface configured to detect and respond to user motion. More specifically, the input device 1 10 may process motion-based instructions and/or queries without any physical contact from the user.
[0020] The input device 1 10 includes one or more sensors 1 12, an input analysis module 1 14, and a user interface 1 16. The sensor 1 12 may be configured to receive user inputs and/or collect data (e.g., images, video, audio recordings, and the like) about the surrounding environment. Example suitable sensors include, but are not limited to: cameras, capacitive sensors, microphones, and the like. In some aspects, at least one of the sensors 1 12 (e.g., a camera) may be configured to capture and/or record a motion input 102 from the user 101 . Example motion inputs may include various poses and/or movements that a user can perform with the user’s arms, legs, hands, feet, head, torso, or any other part of the user’s body.
[0021 ] The input analysis module 1 14 is configured to process the received motion inputs 102. For example, the input analysis module 1 14 may convert the detected motion to one or more user inputs that can be used to control and/or operate the input device 1 10. Example user inputs include, but are not limited to, selection inputs (e.g., for selecting content to be displayed or
played back on the input device 1 10 or an electronic system coupled to the input device 1 10), control inputs (e.g., for changing an operation mode of the input device 1 10 or an electronic system coupled to the input device 1 10), and data inputs (e.g., for entering information and/or search queries). Different types of movements may be associated with different user inputs depending on the application and/or operation mode of the input device 1 10.
[0022] The input analysis module 1 14 may detect the motion inputs 102 from raw sensor data (e.g., images) captured by one or more of the sensors 1 12. For example, the input analysis module 1 14 may first detect the presence of a user 101 in the one or more images captured via the sensors 1 12. In some aspects, the input analysis module 1 14 may detect the user 101 based on the user’s face, and/or other body parts. Upon identifying the user 101 , the input analysis module 1 14 may then track the movements of one or more body parts (e.g., head, hands, fingers, feet and the like) across one or more images or video frames captured by the sensors 1 12. The input analysis module 1 14 may further compare the tracked movements to one or more predefined and/or recognizable motions (e.g., hand waves, finger points, finger swipes, and the like) to identify a corresponding motion input 102.
[0023] In some embodiments, the input analysis module 1 14 may use one or more neural network models 104 to detect and/or identify motion inputs 102 from raw sensor data. In some aspects, the neural network models 104 may be trained to infer the presence of a user (e.g., human) and/or position of the user’s body parts from an image or frame of video. For example, one or more of the neural network models 104 may be trained to detect humans based on the presence of a human face (e.g., a combination of facial features) and/or other human-specific features (e.g., hands and fingers) in a received image. Further, one or more of the neural network models 104 may be trained to model the remaining body parts of the user based on the identified feature. For example, a user’s torso is the contiguous region directly below the user’s face, and the user’s limbs are offshoots from the torso.
[0024] The deep learning environment 120 may be configured to generate the neural network models 104 through deep learning. Deep learning is a particular form of machine learning in which the training phase is performed
over multiple layers, generating a more abstract set of rules in each successive layer. Deep learning architectures are often referred to as artificial neural networks due to the way in which information is processed (e.g., similar to a biological nervous system). For example, each layer of the deep learning architecture may be composed of a number of artificial neurons. The neurons may be interconnected across the various layers so that input data (e.g., the raw data) may be passed from one layer to another. More specifically, each layer of neurons may perform a different type of transformation on the input data that will ultimately result in the desired output (e.g., the answer). The interconnected framework of neurons may be referred to as a neural network model. Thus, the neural network models 104 may include a set of rules that can be used to describe the body parts and/or features of a user (e.g., head, neck, torso, arms, legs, hands, feet, fingers, toes, and the like).
[0025] The deep learning environment 120 may have access to a large volume of raw data and may be trained to recognize a set of rules (e.g., certain objects, features, and/or other detectable attributes) associated with the raw data. For example, in some aspects, the deep learning environment 120 may be trained to recognize a human. During the training phase, the deep learning environment 120 may process or analyze a large number of images, videos, audio, and/or other media containing a biometric signature of a human. The deep learning environment 120 may also receive an indication that the provided data describes a human (e.g., in the form of user input from a user or operator reviewing the media and/or data or metadata provided with the media). The deep learning environment 120 may then perform statistical analysis on the images, video, audio and/or other media to determine a common set of features associated with humans. In some aspects, the determined features (or rules) may form an artificial neural network spanning multiple layers of abstraction.
[0026] The deep learning environment 120 may provide the learned set of rules (e.g., as the neural network models 104) to the input device 1 10 for inferencing. In some aspects, one or more of the neural network models 104 may be provided to (e.g., and stored on) the input device 1 10 at a device manufacturing stage. For example, the input device 1 10 may be pre-loaded with the neural network models 104 prior to being shipped to an end user. In
some other aspects, the input device 1 10 may receive one or more of the neural network models 104 from the deep learning environment 120 at runtime. For example, the deep learning environment 120 may be communicatively coupled to the input device 1 10 via a network (e.g., the cloud). Accordingly, the input device 1 10 may receive the neural network models 104 (including updated neural network models) from the deep learning environment 120, over the network, at any time.
[0027] In some embodiments, the input analysis module 1 14 may detect and process the motion inputs 102 based on the neural network models 104 provided by the deep learning environment 120. For example, during the inferencing phase, the input analysis module 1 14 may apply the neural network models 104 to the data collected from the sensors 1 12, by traversing the artificial neurons in the artificial network, to generate inferences about the user’s body parts and associated motions. For example, the input analysis module 1 14 may use the neural network models 104 to detect the user’s fingertip, track the motion of the fingertip, and/or convert the fingertip motion to a user input. The input analysis module 1 12 may further instruct the input device 1 10 to perform one or more operations associated with the detected motion input 102. By generating the inferences locally on the input device 1 10, the present embodiments may be used to perform machine learning on biometric data (e.g., images of the user 101 ) in a manner that protects user privacy.
[0028] In some embodiments, the input analysis module 1 14 may use the data collected from the sensors 1 12 to perform additional training on the neural network models 104. For example, the input analysis module 1 14 may refine the neural network models 104 and/or generate new neural network models based on the locally-generated sensor data. In some aspects, the neural network models 104 may be fine-tuned to detect and/or recognize the movements of a particular user. For example, such additional training may be performed based on images and/or videos of the user 101 captured by the sensors 1 12. The additional training may be initiated manually (e.g., using an independent scripted mechanism) or automatically upon detecting sensor data from the sensors 1 12. In another example, the input analysis module 1 14 may
use previously-detected user movements to perform additional training on the neural network models 104 (e.g., in a feedback loop).
[0029] In some aspects, the input analysis module 1 14 may provide the updated neural network models to the deep learning environment 120 to further refine the deep learning architecture. In this manner, the deep learning environment 120 may further refine its neural network models 104 based on the sensor data captured by the input device 1 10 (e.g., combined with sensor data captured by various other input devices) without receiving or having access to the raw sensor data.
[0030] The user interface 1 16 may provide an interface through which the user 101 can operate, interact with, or otherwise use the input device 1 10 or an electronic system (not shown for simplicity) coupled to the input device 1 10. In some embodiments, the user interface 1 16 may display, render, or otherwise manifest the motion input 102 on the input device 1 10. For example, the user interface 1 16 may include a graphical user interface (GUI) rendered on a display of the input device 1 10. The user interface 1 16 may respond to motion inputs 102 by dynamically updating the display, for example, to navigate the GUI, display new content, track cursor movement, and the like. In some aspects, the user interface 1 16 may generate a search query based on the detected motion inputs 102. For example, the search query may include a string of characters to be searched for or identified. In some aspects, the user interface 1 16 may search for content locally on the input device 1 10 based on the search query. In some other aspects, the user interface 1 16 may transmit the search query to a network resource for further processing.
[0031 ] The input analysis module 1 14 may generate the string of characters (which make up the search query) based on one or more motions detected in the raw sensor data. For example, the user 101 may draw one or more letters and/or characters in the air with the user’s hand. The input analysis module 1 14 traces the movement of the user’s hand across one or more images captured via the sensors 1 12 and converts the trace to a corresponding text string. In some aspects, the input analysis module 1 14 may use optical character recognition (OCR) techniques to perform the conversion. Aspects of the present disclosure recognize that the accuracy of the conversion
may depend on the precision of the motion inputs 102 detected by the input device 1 10. While larger objects (such as a user’s hand) may be easier to detect and track, smaller objects (such as a user’s fingertip) may be more finely articulated and may thus result in more granular traces when tracked by the input analysis module 1 14. Thus, in some embodiments, the input analysis module 1 14 may be configured to detect and track a user’s fingertip in one or more frames of sensor data.
[0032] FIG. 2 shows a block diagram of an input device 200, in accordance with some embodiments. The input device 200 may be one embodiment of the input device 1 10 of FIG. 1 . Thus, the input device 200 may provide a touchless user interface configured to detect and respond to motion inputs that do not involve any physical contact from a user. The input device 200 includes a camera 210, a display 250, a network interface (l/F) 260, and a processing system 280 including a neural network 220, an input classifier 230, and a user interface 240. In some embodiments, the neural network 220 and/or the input classifier 230 may be implemented by neural network acceleration hardware (which may include one or more processors configured to accelerate neural network inferencing).
[0033] The camera 210 is configured to capture one or more images 201 of the environment surrounding the input device 200. The camera 210 may be one embodiment of one of the sensors 1 12 of FIG. 1. Thus, the camera 210 may be configured to capture images 201 (e.g., still-frame images and/or video) of a scene in front of or proximate the input device 200. In some embodiments, the camera 210 may include optical sensors (e.g., photodiodes, CMOS image sensor arrays, CCD arrays, and/or any other sensors capable of detecting wavelengths of light in the visible spectrum, the infrared spectrum, and/or the ultraviolet spectrum) and/or depth sensors (e.g., time-of-flight sensors, structured light sensors, and/or any other sensors capable of measuring the depth or distances of objects). In low-light conditions, near-infrared (NIR) and/or depth-sensing cameras may capture finer details of a scene than more traditional optical cameras that detect red, green, and blue (RGB) wavelengths of light. In some aspects, the input device 200 may use RGB cameras to capture images 201 during the day or in relatively bright conditions, and may
use NIR or depth-sensing cameras to capture images 201 at night or in low-light conditions.
[0034] The neural network 220 is configured to detect one or more input objects in the captured images 201 . In some aspects, the input object may include a user’s fingertip. The neural network 220 may identify the location of the user’s fingertips in each of the images 201 as fingerprint data 202. In some embodiments, the neural network 220 may generate inferences about the user’s fingertips using one or more neural network models. For example, as described with respect to FIG. 1 , the neural network 220 may receive trained neural network models (e.g., from the deep learning environment 120) prior to receiving the images 201 from the camera 210. The neural network 220 may include a user detection module 222 and a fingertip detection module 224.
[0035] The user detection module 222 may detect the presence of a human (e.g., user) in one or more of the images 201 . More specifically, the user detection module 222 may implement one or more neural network models to infer the presence of a human based on a combination of facial features and/or other human-specific features. For example, the user detection module 222 may detect the presence of a human face using any known face detection algorithms and/or techniques. In some embodiments, the user detection module 222 may be further configured to model the remaining body parts of the user based, at least in part, on the identified feature. For example, the user detection module 222 may identify the users’ torso, arms, legs, hands, and/or feet based on the relative position of the user’s face.
[0036] The fingertip detection module 224 may detect the location of a user’s fingertip in one or more of the images 201 . More specifically, the fingertip detection module 224 may implement one or more neural network models to infer the presence of fingers and locations of the fingertips from an image of the user’s hands. For example, the neural network models may be trained using images of hands with fingers in various poses and/or stages of movement. The training data may be annotated to identify the fingertip locations in each frame or image. Through training, the neural network models may be configured to detect subtle movements of a user’s fingertip (e.g., while the user’s hand remains relatively stationary) from far distances (such as across
a room). For example, the neural network models may be trained on images captured of users at multiple distances from an image capture device. In some embodiments, the fingertip detection module 224 may crop the image 201 around the user’s hand (e.g., as detected by the user detection module 222) and apply the neural network models to the cropped image to detect the locations of the user’s fingertips as fingertip data 202.
[0037] The input classifier 230 is configured to convert the fingertip data 202 to one or more user inputs 203. As described above, the fingertip data 202 may indicate the locations of the user’s fingertips in one or more images 201 captured by the camera 201 . In some embodiments, the input classifier 230 may generate the user inputs 203 based on movements of the user’s fingertips. For example, the input classifier 230 may track changes in the locations of the user’s fingertips across the one or more images 201 . In some aspects, the input classifier 230 may continuously track the movements of the user’s fingertips until the movements satisfy a particular condition and/or no movement (or very little movement) is detected for at least a threshold duration. The input classifier 230 may include a motion input recognition module 232 and a text conversion module 234.
[0038] The motion input recognition module 232 may compare the movements of the user’s fingertips to one or more predetermined motions to determine the associated user inputs 203. Example user inputs 203 include, but are not limited to, selection inputs, control inputs, and data inputs. Different types of movements may be associated with different user inputs depending on the application and/or operation mode of the input device 200. For example, a finger swiping motion may be used to scroll through content presented on the display 250, whereas a finger tapping motion may be used to select a particular content item on the display 250. In some embodiments, the motion input recognition module 232 may place the input device 200 in a text-input mode in response to detecting a particular motion (such as a user raising his or her hand). While operating in the text-input mode, the input classifier 230 may classify subsequent fingerprint data 202 as text input data.
[0039] The text conversion module 234 may convert the fingertip data 202 to text-based user inputs 203. For example, the text conversion module
234 may trace the movements of one of the user’s fingertips (such as the tip of the user’s index finger) as the user draws one or more letters and/or characters in the air. The text conversion module 234 may further convert the trace to a character string. In some embodiments, the text conversion module 234 may use OCR techniques to convert the trace to the character string. In some other embodiments, the text conversion module 234 may implement one or more neural network models to infer the characters to be converted based, at least in part, on samples of the user’s handwriting and/or writing style. With reference for example to FIG. 3, the text conversion module 234 may trace a user’s fingertip (FT) 301 as it draws the word“Action” in the air (e.g., as fingertip trace 302). The text conversion module 234 may then convert the fingertip trace 302 to a text-based input string 303 corresponding to the word“Action.”
[0040] In some aspects, the text conversion module 234 may process each character of the character string individually. For example, once a user has completed drawing a particular letter or character, the text conversion module 234 may convert the trace to a corresponding character string. With reference for example to FIG. 3, the text conversion module 234 may recognize the letters“A,”“c,”“t,”“i,”“o,” and“n,” individually, before the fingertip trace 302 is complete. In some other aspects, the text conversion module 234 may process the entire character string as a whole. For example, the text conversion module 234 may wait until the user is finished drawing a complete word or sequence of characters before converting the trace to a corresponding character string. With reference for example to FIG. 3, the text conversion module 234 may recognize the word“Action” as a whole, after the fingertip trace 302 is complete.
[0041 ] The user interface 240 may generate an output 204 in response to the user inputs 203. The output 204 may be rendered on the display 250. In some aspects, the output 204 may include updates to a GUI presented on the display 250. For example, the GUI updates may track the user inputs 203 and/or movements of the user’s fingertips (e.g., cursor movement, text insertion, and the like). In some other aspects, the output 204 may include content to be presented or played back on the display 250. For example, the user inputs 203 may include instructions for searching, selecting, or otherwise retrieving content
to be displayed. In some implementations, the user interface 240 may include a content store 242 and a content retrieval module 244. The content store 242 may store or buffer content that can be rendered on the display 250 and/or a display device (not shown) coupled to the input device 200. The content retrieval module 244 may retrieve content from one or more network resources external to the input device 200.
[0042] In some embodiments, the content retrieval module 244 may be configured to generate a search query 205 based, at least in part, on the user input 203 received from the input classifier 230. For example, the user input 203 may be a text-based user input comprising a string of characters. In some aspects, the search query 205 may include at least a portion of the character string included in the user input 203. The search query may be transmitted to a network resource (not shown for simplicity) via the network interface 260. The network resource may include memory and/or processing resources to generate one or more results 206 for the search query 205. For example, the network resource may search one or more networked devices (e.g., the Internet, content delivery networks, and the like) for the content or information requested by the search query 205. The network resource may then send the results 206 (e.g., including the requested content or information) back to the input device 200.
[0043] FIG. 4 shows an example environment 400 in which the present embodiments may be implemented. The environment 400 includes an input device 410, a user 420, and a non-user object 430. The input device 410 may be one embodiment of the input device 1 10 of FIG. 1 and/or input device 200 of FIG. 2. In the example of FIG. 4, the input device 410 is depicted as a media device (e.g., a television) capable of displaying or playing back media content (e.g., images, videos, audio, and the like) to the user 420.
[0044] The input device 410 includes a camera 412 and a display 414. The camera 412 may be one embodiment of the camera 210 of FIG. 2. Thus, the camera 412 may be configured to capture images (e.g., still-frame images and/or video) of a scene 401 in front of the input device 410. The display 414 may be one embodiment of the display 250 of FIG. 2. Thus, the display 414 may correspond to and/or provide a user interface through which the user 420 may interact with or use the input device 410. In some embodiments, the input
device 410 may provide a touchless user interface configured to detect and respond to motion inputs that do not involve any physical contact from the user
420. Thus, in some aspects, the camera 412 continuously (or periodically) captures images of the scene 401 to detect motion inputs from the user 420.
[0045] In some embodiments, the input device 410 may process user inputs based, at least in part, on the movements of a user’s fingertip 421 . For example, the input device 410 may first detect the presence of the user 420 in the scene 401 . In the example of FIG. 4, the user 420 is shown sitting on a couch (e.g., non-user object 430) at a comfortable viewing distance from the input device 410. The user’s hands are shown to be resting in his or her lap. From this position, the user 420 may make small, subtle motions of the fingertip
421 , for example, without lifting the user’s right hand from his or her lap. Using only the fingertip 421 , the user 420 may control the input device 410 in a relaxed and unobtrusive manner, allowing for a more robust user experience.
[0046] In some embodiments, the input device 410 may use one or more neural network models to infer the user 420 from one or more images captured via the camera 412. For example, the neural network models may be trained to differentiate human users from other non-user objects 430 in the scene 401 . Upon detecting the presence of the user 420, the input device 410 may generate one or more inferences about the user’s fingertip 421 . In some embodiments, the input device 410 may further implement one or more neural network models to infer the location or position of the fingertip 421 . For example, the neural network models may be trained to identify the user’s fingertip 421 based on the presence and/or location of one or more other body parts of the user 420 (such as the user’s hand).
[0047] The input device 410 may detect one or more motion inputs by tracking the movements of the user’s fingertip 421 across one or more images captured via the camera 412. In some embodiments, the input device 410 may process the motion inputs to control the presentation and/or playback of media content on the display 414. For example, the motion inputs may be used to navigate, select, and/or search for media content to be played back on the display 414. In some aspects, the input device 410 may classify one or more motion inputs as text input data. For example, the input device 410 may trace
the movements of the user’s fingertip 421 as the user draws one or more letters and/or characters in the air. The input device 410 may further convert the trace to a character string using OCR techniques (e.g., as described with respect to FIGS. 2 and 3). The input device 410 may then retrieve one or more content items associated with the character string (stored locally or on a network resource) for presentation and/or playback on the display 414.
[0048] The accuracy of many OCR techniques depends on the precision of the traces from which the characters are derived (such as the fingertip trace 302 of FIG. 3). Aspects of the present disclosure recognize that the user’s fingertip 421 may be more finely articulated than any other part of the user’s body. Thus, by tracking the movements of the user’s fingertip 421 , the input device 410 may detect and process motion inputs with a finer level of precision and granularity than would otherwise be possible by tracking the movements of any other body part. Accordingly, the present embodiments may enable and/or improve the detection of touchless text-based motion inputs. Among other advantages, the fingertip tracking techniques described herein may be able to detect subtle movements of a user’s fingertip from far distances (such as shown in FIG. 4), allowing for an improved user experience.
[0049] FIG. 5 shows another block diagram of an input device 500, in accordance with some embodiments. The input device 500 may be one embodiment of the input device 200 of FIG. 2 and/or the input device 1 10 of FIG. 1 . The input device 500 includes a sensor interface 510, a processor 520, and a memory 530.
[0050] The sensor interface 510 may be used to communicate with one or more sensors (e.g., cameras) coupled to the input device 500. In some implementations, the sensor interface 510 may be configured to communicate with a camera of the input device 500 (e.g., camera 210 of FIG. 2). For example, the sensor interface 510 may transmit signals to, and receive signals from, the camera to capture a plurality of images of a scene facing the input device 500.
[0051 ] The memory 530 may also include a non-transitory computer- readable medium (e.g., one or more nonvolatile memory elements, such as
EPROM, EEPROM, Flash memory, a hard drive, etc.) that may store at least the following software (SW) modules:
• a user detection SW module 532 to detect a user in the plurality of
images using one or more first neural network models;
• a fingertip detection SW module 534 to detect a fingertip of the user in the plurality of images using one or more second neural network models;
• a fingertip tracking SW module 536 to track a position of the fingertip across the plurality of images; and
• a user interface SW module 538 to process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
Each software module includes instructions that, when executed by the processor 520, cause the input device 500 to perform the corresponding functions. The non-transitory computer-readable medium of memory 530 thus includes instructions for performing all or a portion of the operations described below with respect to FIG. 6.
[0052] Processor 520 may be any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in memory 530. For example, the processor 520 may execute the user detection SW module 532 to detect a user in the plurality of images using one or more first neural network models. The processor 520 may further execute the fingertip detection SW module 534 to detect a fingertip of the user in the plurality of images using one or more second neural network models. Still further, the processor 520 may execute the fingertip tracking SW module 536 to track a position of the fingertip across the plurality of images. The processor 520 also may execute the user interface SW module 538 to process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
[0053] FIG. 6 is an illustrative flowchart depicting an example operation 600 for processing user inputs, in accordance with some embodiments. With
reference for example to FIG. 1 , the operation 600 may be performed by the input device 1 10 to control or operate a touchless user interface.
[0054] The input device may capture a plurality of images of a scene (610). The images may include still-frame images and/or videos captured by a camera in, or coupled to, the input device. In some embodiments, the camera may include optical sensors and/or depth sensors. In low-light conditions, NIR and/or depth-sensing cameras may capture finer details of a scene than more traditional optical cameras that detect RGB wavelengths of light. In some aspects, the input device may use RGB cameras to capture images during the day or in relatively bright conditions, and may use NIR or depth-sensing cameras to capture images at night or in low-light conditions.
[0055] The input device detects a user in the plurality of images using one or more first neural network models (620). In some implementations, the input device may implement one or more neural network models to infer the presence of a human based on a combination of facial features and/or other human-specific features. For example, the input device may detect the presence of a human face using any known face detection algorithms and/or techniques. In some implementations, the input device may further model the remaining body parts of the user based, at least in part, on the identified feature. For example, the input device may identify the users’ torso, arms, legs, hands, and/or feet based on the relative position of the user’s face.
[0056] The input device may further detect a fingertip of the user in the plurality of images using one or more second neural network models (630). In some implementations, the input device may implement one or more neural network models to infer the presence of fingers and locations of the fingertips from an image of the user’s hands. For example, the neural network models may be trained using images of hands with fingers in various poses and/or stages of movement. The training data may be annotated to identify the fingertip locations in each frame or image. Through training, the neural network models may be configured to detect subtle movements of a user’s fingertip from far distances (such as described with reference to FIG. 4).
[0057] The input device may track a position of the fingertip across the plurality of images (640) and process a user input based at least in part on
changes in the position of the fingertip across the plurality of images (650). In some implementations, the input device may continuously track the movements of the user’s fingertips until the movements satisfy a particular condition and/or no movement (or very little movement) is detected for at least a threshold duration. The input device may compare the movements of the user’s fingertips to one or more predetermined motions to determine the associated user inputs. Example user inputs include, but are not limited to, selection inputs, control inputs, and data inputs. In some implementations, the input device may trace the movements of one of the user’s fingertips (such as the tip of the user’s index finger) as the user draws one or more letters and/or characters in the air. The input device may further convert the trace to a character string (such as described with reference to FIG. 3).
[0058] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0059] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.
[0060] The methods, sequences or algorithms described in connection with the aspects disclosed herein may be embodied directly in hardware, in a
software module executed by a processor, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor.
[0061 ] In the foregoing specification, embodiments have been described with reference to specific examples thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader scope of the disclosure as set forth in the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative sense rather than a restrictive sense.
Claims
1 . A method of processing user inputs performed by an input device, comprising:
capturing a plurality of images of a scene;
detecting a user in the plurality of images using one or more first neural network models;
detecting a fingertip of the user in the plurality of images using one or more second neural network models;
tracking a position of the fingertip across the plurality of images; and processing a user input based at least in part on changes in the position of the fingertip across the plurality of images.
2. The method of claim 1 , wherein the detecting of the user comprises:
identifying a first body part in the plurality of images; and
inferring a presence of the user in the plurality of images based on the identified first body part.
3. The method of claim 2, wherein the detecting of the fingertip comprises:
identifying a hand in the plurality of images based on a relative position of the first body part; and
inferring the position of the fingertip from the identified hand.
4. The method of claim 1 , wherein the processing of the user input comprises:
updating a graphical user interface (GUI) based on the user input.
5. The method of claim 4, wherein the updating of the GUI comprises:
moving a cursor in the GUI responsive to the changes in the position of the fingertip.
6. The method of claim 4, wherein the processing of the user input further comprises:
generating a trace based on the changes in the position of the fingertip; and
converting the trace to a character string.
7. The method of claim 6, wherein the conversion of the trace to the character string is performed using an optical character recognition (OCR) technique.
8. The method of claim 6, wherein the conversion of the trace to the character string is performed using one or more third neural network models trained to infer one or more characters of the character string based on handwriting samples of the user.
9. The method of claim 6, wherein the updating of the GUI comprises:
displaying the character string in the GUI.
10. The method of claim 1 , further comprising:
training the one or more second neural network models on images captured of users at a plurality of distances from an image capture device.
1 1 . An input device comprising:
processing circuitry; and
memory storing instructions that, when executed by the processing circuitry, causes the input device to:
detect a user in a plurality of images using one or more first neural network models;
detect a fingertip of the user in the plurality of images using one or more second neural network models;
track a position of the fingertip across the plurality of images; and
process a user input based at least in part on changes in the position of the fingertip across the plurality of images.
12. The input device of claim 1 1 , wherein execution of the instructions for detecting the user causes the input device to:
identify a first body part in the plurality of images; and
infer a presence of the user in the plurality of images based on the identified first body part.
13. The input device of claim 12, wherein execution of the instructions for detecting the fingertip causes the input device to:
identify a hand in the plurality of images based on a relative position of the first body part; and
infer the position of the fingertip from the identified hand.
14. The input device of claim 1 1 , wherein execution of the instructions for processing the user input causes the input device to:
update a graphical user interface (GUI) based on the user input.
15. The input device of claim 14, wherein execution of the instructions for updating the GUI causes the input device to:
move a cursor in the GUI responsive to the changes in the position of the fingertip.
16. The input device of claim 14, wherein execution of the instructions for processing the user input causes the input device to:
generate a trace based on the changes in the position of the fingertip; and
convert the trace to a character string.
17. The input device of claim 16, wherein the conversion of the trace to the character string is performed using an optical character recognition (OCR) technique.
18. The input device of claim 16, wherein the conversion of the trace to the character string is performed using one or more third neural network models trained to infer one or more characters of the character string based at least in part on handwriting samples of the user.
19. The input device of claim 16, wherein execution of the instructions for updating the GUI causes the input device to:
display the character string in the GUI.
20. An input device comprising:
a camera configured to capture a plurality of images of a scene;
a display configured to provide a graphical user interface (GUI); and a processing system configured to:
detect a user in the plurality of images using one or more first neural network models;
detect a fingertip of the user in the plurality of images using one or more second neural network models;
track a position of the fingertip across the plurality of images; and update the GUI based at least in part on changes in the position of the fingertip across the plurality of images.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201962860200P | 2019-06-11 | 2019-06-11 | |
| US62/860,200 | 2019-06-11 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020252008A1 true WO2020252008A1 (en) | 2020-12-17 |
Family
ID=73781258
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2020/036979 Ceased WO2020252008A1 (en) | 2019-06-11 | 2020-06-10 | Fingertip tracking for touchless input device |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2020252008A1 (en) |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20030095140A1 (en) * | 2001-10-12 | 2003-05-22 | Keaton Patricia (Trish) | Vision-based pointer tracking and object classification method and apparatus |
| US20110041100A1 (en) * | 2006-11-09 | 2011-02-17 | Marc Boillot | Method and Device for Touchless Signing and Recognition |
| JP2012059271A (en) * | 2010-09-13 | 2012-03-22 | Ricoh Co Ltd | Human-computer interaction system, hand and hand instruction point positioning method, and finger gesture determination method |
| US20170161555A1 (en) * | 2015-12-04 | 2017-06-08 | Pilot Ai Labs, Inc. | System and method for improved virtual reality user interaction utilizing deep-learning |
| KR20180130869A (en) * | 2017-05-30 | 2018-12-10 | 주식회사 케이티 | CNN For Recognizing Hand Gesture, and Device control system by hand Gesture |
| JP2018206321A (en) * | 2017-06-09 | 2018-12-27 | コニカミノルタ株式会社 | Image processing device, image processing method and image processing program |
-
2020
- 2020-06-10 WO PCT/US2020/036979 patent/WO2020252008A1/en not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20030095140A1 (en) * | 2001-10-12 | 2003-05-22 | Keaton Patricia (Trish) | Vision-based pointer tracking and object classification method and apparatus |
| US20110041100A1 (en) * | 2006-11-09 | 2011-02-17 | Marc Boillot | Method and Device for Touchless Signing and Recognition |
| JP2012059271A (en) * | 2010-09-13 | 2012-03-22 | Ricoh Co Ltd | Human-computer interaction system, hand and hand instruction point positioning method, and finger gesture determination method |
| US20170161555A1 (en) * | 2015-12-04 | 2017-06-08 | Pilot Ai Labs, Inc. | System and method for improved virtual reality user interaction utilizing deep-learning |
| KR20180130869A (en) * | 2017-05-30 | 2018-12-10 | 주식회사 케이티 | CNN For Recognizing Hand Gesture, and Device control system by hand Gesture |
| JP2018206321A (en) * | 2017-06-09 | 2018-12-27 | コニカミノルタ株式会社 | Image processing device, image processing method and image processing program |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12086323B2 (en) | Determining a primary control mode of controlling an electronic device using 3D gestures or using control manipulations from a user manipulable input device | |
| US12517589B2 (en) | Method for creating a gesture library | |
| US12282506B2 (en) | Electronic apparatus for searching related image and control method therefor | |
| US8902198B1 (en) | Feature tracking for device input | |
| US20130044912A1 (en) | Use of association of an object detected in an image to obtain information to display to a user | |
| US8897490B2 (en) | Vision-based user interface and related method | |
| WO2020078017A1 (en) | Method and apparatus for recognizing handwriting in air, and device and computer-readable storage medium | |
| CN112106042A (en) | Electronic device and control method thereof | |
| US20170220120A1 (en) | System and method for controlling playback of media using gestures | |
| CN102339125A (en) | Information equipment and control method and system thereof | |
| CN106648078B (en) | Multi-mode interaction method and system applied to intelligent robot | |
| US10423824B2 (en) | Body information analysis apparatus and method of analyzing hand skin using same | |
| KR102476619B1 (en) | Electronic device and control method thereof | |
| CN103092332A (en) | Digital image interactive method and system of television | |
| CN103135746A (en) | Non-touch control method and non-touch control system and non-touch control device based on static postures and dynamic postures | |
| WO2019134606A1 (en) | Terminal control method, device, storage medium, and electronic apparatus | |
| KR20180082950A (en) | Display apparatus and service providing method of thereof | |
| CN101110102A (en) | Game scene and character control method based on player's fist | |
| US20160054968A1 (en) | Information processing method and electronic device | |
| WO2020252008A1 (en) | Fingertip tracking for touchless input device | |
| Tiwari et al. | Volume controller using hand gestures | |
| US20250231644A1 (en) | Aggregated likelihood of unintentional touch input | |
| US11308150B2 (en) | Mobile device event control with topographical analysis of digital images inventors | |
| Mali et al. | Design and Implementation of Hand Gesture Assistant Command Control Video Player Interface for Physically Challenged People | |
| Karthikhaa Shree et al. | A Vision Based Hand Gesture Interface for Controlling VLC Media Player |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20822336 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20822336 Country of ref document: EP Kind code of ref document: A1 |