EP4626558A1 - Anwendung einer kopfbezogenen übertragungsfunktion zur klanglokalisierung in notfahrszenarien - Google Patents
Anwendung einer kopfbezogenen übertragungsfunktion zur klanglokalisierung in notfahrszenarienInfo
- Publication number
- EP4626558A1 EP4626558A1 EP23896993.5A EP23896993A EP4626558A1 EP 4626558 A1 EP4626558 A1 EP 4626558A1 EP 23896993 A EP23896993 A EP 23896993A EP 4626558 A1 EP4626558 A1 EP 4626558A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- user
- worn device
- head worn
- signal source
- sound
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- A—HUMAN NECESSITIES
- A62—LIFE-SAVING; FIRE-FIGHTING
- A62B—DEVICES, APPARATUS OR METHODS FOR LIFE-SAVING
- A62B18/00—Breathing masks or helmets, e.g. affording protection against chemical agents or for use at high altitudes or incorporating a pump or compressor for reducing the inhalation effort
- A62B18/08—Component parts for gas-masks or gas-helmets, e.g. windows, straps, speech transmitters, signal-devices
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/012—Head tracking input arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/013—Eye tracking input arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/167—Audio in a user interface, e.g. using voice commands for navigating, audio feedback
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G08—SIGNALLING
- G08B—SIGNALLING SYSTEMS, e.g. PERSONAL CALLING SYSTEMS; ORDER TELEGRAPHS; ALARM SYSTEMS
- G08B25/00—Alarm systems in which the location of the alarm condition is signalled to a central station, e.g. fire or police telegraphic systems
- G08B25/01—Alarm systems in which the location of the alarm condition is signalled to a central station, e.g. fire or police telegraphic systems characterised by the transmission medium
- G08B25/016—Personal emergency signalling and security systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
- H04S7/303—Tracking of listener position or orientation
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
- G01S2205/00—Position-fixing by co-ordinating two or more direction or position line determinations; Position-fixing by co-ordinating two or more distance determinations
- G01S2205/01—Position-fixing by co-ordinating two or more direction or position line determinations; Position-fixing by co-ordinating two or more distance determinations specially adapted for specific applications
- G01S2205/06—Emergency
-
- G—PHYSICS
- G08—SIGNALLING
- G08B—SIGNALLING SYSTEMS, e.g. PERSONAL CALLING SYSTEMS; ORDER TELEGRAPHS; ALARM SYSTEMS
- G08B21/00—Alarms responsive to a single specified undesired or abnormal condition and not otherwise provided for
- G08B21/02—Alarms for ensuring the safety of persons
Definitions
- Visibility of target items or individuals being sought may be obscured in emergency situations and environments. This obscuration may be caused by smoke, suspended matter, piles of debris, low or no light, etc., orthe target may be covered with a dusting of ash, dirt, soot, etc., making rapid recognition impossible.
- an individual could be partially hidden by machinery, electrical wiring and equipment, pipes and pipe racks, etc.
- First responders may be unable to identify the origin/location of a signal source (such as a sound generating source and/or radio beacon/signal) in such a limited visibility environment.
- a head worn device such as a face mask, goggles, and/or self-contained breathing apparatus (SCBA) may use image sensors and/or audio sensors to accurately identify the source of a sound, such as a sound emitted by a person-down alarm in an emergency environment, and to determine the gaze of the user of the head worn device.
- the head worn device may include a user interface configured to indicate the signal source location to the user of the head worn device and direct the user to the location based on the determined gaze.
- the head worn device includes at least one microphone, at least one image sensor, and processing circuitry configured to receive an audio signal detected by the at least one microphone, the audio signal originating from a signal source.
- the processing circuitry is configured to determine a location of the signal source based on the received audio signal.
- the processing circuitry is further configured to receive image data from the at least one image sensor, the image data being associated with at least one of the user’s face and at least one of the user’s eyes.
- the processing circuitry is further configured to determine a gaze direction of the user based on the received image data, and determine a user instruction based on the determined location and determined gaze direction.
- FIG. 1 is a schematic diagram of various devices and components according to some embodiments of the present invention.
- FIG. 2 is a block diagram of an example head worn device according to some embodiments of the present invention.
- FIG. 4 is an illustration of a technique for sound location using a head worn device, according to some embodiments of the present invention.
- FIG. 5 is an illustration of a technique for gaze tracking using a head worn device, according to some embodiments of the present invention.
- FIG. 6 is an illustration of another technique for gaze tracking using a head worn device, according to some embodiments of the present invention.
- FIG. 7 is an illustration of another technique for gaze tracking using a head worn device, according to some embodiments of the present invention.
- FIG. 8 is an illustration of another technique gaze tracking using a head worn device, according to some embodiments of the present invention.
- FIG. 10 is an illustration of another technique for sound location using a head worn device, according to some embodiments of the present invention.
- FIG. 11 is an illustration of another technique for sound location using a head worn device, according to some embodiments of the present invention.
- FIG. 12 is a flowchart of an example process in a head worn device according to some embodiments of the present invention.
- the joining term, “in communication with” and the like may be used to indicate electrical or data communication, which may be accomplished by physical contact, induction, electromagnetic radiation, radio signaling, infrared signaling or optical signaling, for example.
- electrical or data communication may be accomplished by physical contact, induction, electromagnetic radiation, radio signaling, infrared signaling or optical signaling, for example.
- the term “signal source” may be any detectable signal, which may, for example, indicate and/or be associated with a distress state of a person (such as a first responder) and/or a device (such as a hand held alarm device).
- a signal source may include an audible sound, such as an alarm sound emitted by a personal alarm device, which may include a piercing (e.g., high frequency) sound and/or a low frequency sound.
- a signal source may include sounds associated with an emergency/life threatening situation, such as sounds generated by a person in distress, by equipment/machinery, by running water, by an explosion, etc.
- a signal source may also/altematively include a radio signal/beacon, such as a distress radio signal emitted by a personal alarm device.
- Other detectable signal types may be employed without deviating from the scope of the present disclosure.
- Head worn device 12 may include speaker 26, e.g., integrated into head worn device 12, for providing audio indications/messages to user 14, and may include user microphone 28, e.g., integrated into head worn device 12, for receiving spoken commands from user 14.
- Microphones 18 and/or microphones 20 may be configured to detect audio signals, e.g., sound originating from signal source 30, which may be a sound-generating object/entity/event (e.g., wood breaking, a person shouting/breathing, a pressurized gas release (e.g., a jet release), etc.) in the environment/vicinity of a user 14 of head worn device 12.
- a sound-generating object/entity/event e.g., wood breaking, a person shouting/breathing, a pressurized gas release (e.g., a jet release), etc.
- Signal source location detection system 10 may include hand held device 31, which may be in communication with head worn device 12, e.g., via a wired/wireless connection.
- head worn device 12 may be a mask, such as a mask that is part of a respirator.
- head worn device 12 may include hardware 32, including microphones 18, microphones 20, display 22, speaker 26, microphone 28, accelerometer 34, light emitter 36, image sensor 38, communication interface 40, and processing circuitry 42.
- the processing circuitry 42 may include a processor 44 and a memory 46.
- the processing circuitry 42 may comprise integrated circuitry for processing and/or control, e.g., one or more processors and/or processor cores and/or FPGAs (Field Programmable Gate Array) and/or ASICs (Application Specific Integrated Circuitry) adapted to execute instructions.
- the processor 44 may be configured to access (e.g., write to and/or read from) the memory 46, which may comprise any kind of volatile and/or nonvolatile memory, e.g., cache and/or buffer memory and/or RAM (Random Access Memory) and/or ROM (Read-Only Memory) and/or optical memory and/or EPROM (Erasable Programmable Read-Only Memory).
- Hardware 32 may be removable from the mask body 15 to allow for replacement, upgrade, etc., or may be integrated as part of the head worn device 12.
- Head worn device 12 may further include software 48 stored internally in, for example, memory 46 or stored in external memory (e.g., database, storage array, network storage device, etc.) accessible by head worn device 12 via an external connection.
- the software 48 may be executable by the processing circuitry 42.
- the processing circuitry 42 may be configured to control any of the methods and/or processes described herein and/or to cause such methods, and/or processes to be performed, e.g., by head worn device 12.
- Processor 44 corresponds to one or more processors 44 for performing head worn device 12 functions described herein.
- the memory 46 is configured to store data, programmatic software code and/or other information described herein.
- the software 48 may include instructions that, when executed by the processor 44 and/or processing circuitry 42, causes the processor 44 and/or processing circuitry 42 to perform the processes described herein with respect to head worn device 12.
- head worn device 12 may include gaze tracker 50 configured to perform one or more head worn device 12 functions as described herein, such as detecting the point of gaze (i.e., where user 14 is looking), tracking the 3-dimensional line of sight of user 14, tracking the movement of user 14’s eyes, user 14’s face, etc., as described herein.
- Processing circuitry 42 of the head worn device 12 may include sound locator 52 configured to perform one or more head worn device 12 functions as described herein such as determining the location/origin of a signal source 30, as described herein.
- Processing circuitry 42 of head worn device 12 may include sound classifier 54 configured to perform one or more head worn device 12 functions as described herein such as classifying, labeling, and/or identifying the cause/type of signal source 30, as described herein.
- Processing circuitry 42 of head worn device 12 may include user interface 56 configured to perform one or more head worn device 12 functions as described herein such as displaying (e.g., using display 22) or announcing (e.g., using speaker 26) indications/messages to user 14, such as indications regarding the location of signal source 30 and/or indications regarding the distance/direction of signal source 30 relative to user 14; and/or receiving spoken commands from user 14 (e.g., using microphone 28) and/or receiving other commands from user 14 (e.g., user 14 presses a button in communication with the processing circuitry 42, user 14 interacts with a separate device, such as a smartphone or hand held device 31, which communicates the user’s interactions to the processing circuitry 42 via communication interface 40, etc.) or from other users (e.
- FIG. 1 shows two each of microphones 18 and 20, it is understood that implementations are not limited to two sets of two microphones, and that there can be different numbers of sets of microphones, each having different quantities of individual microphones.
- Display 22 may be implemented by any device, either standalone or part of head worn device 12, that is configurable for displaying indications/messages to user 14, e.g., indications regarding the location of signal source 30 and/or indications regarding the distance/direction of signal source 30 relative to user 14.
- display 22 may be configured to display an icon (e.g., an arrow), the icon indicating which direction the user should adjust his gaze to, e.g., as determined by sound locator 52 and/or gaze tracker 50.
- display 22 may be configured to display a relative distance (e.g., “5 meters”) separating the user from the signal source 30, e.g., as determined by sound locator 52 and/or gaze tracker 50.
- display 22 may be configured to display a predicted classification/type/labeFetc. of signal source 30, e.g., as determined by sound classifier 54.
- display 22 may be configured to display an indication of a location of signal source 30, e.g., as an AR overlay on lens 16, such as by drawing a circle as an augmented reality (AR) overlay on the area of the lens 16 and/or display 22 corresponding to the location of signal source 30 within user 14’s field of view.
- display 22 may be configured to instruct user 14 to change direction (e.g., turn left, turn right, look up, look down, turn around, etc.) if the location of signal source 30 is outside of user 14’s field of view.
- direction e.g., turn left, turn right, look up, look down, turn around, etc.
- Speaker 26 may be implemented by any device, either standalone or part of head worn device 12, that is configurable for generating sound that is audible to user 14 while wearing head worn device 12, and is configurable for announcing (e.g., using speaker 26) indications/messages to user 14, such as indications regarding the location of signal source 30 and/or indications regarding the distance/direction of signal source 30 relative to user 14.
- speaker 26 is configured to provide audio messages corresponding to the indications described above with respect to display 22.
- Microphone 28 may be implemented by any device, either standalone or part of head worn device 12 and/or user interface 56, that is configurable for detecting spoken commands by user 14 while user 14 is wearing head worn device 12.
- Accelerometer 34 may be implemented by any device, either standalone or part of head worn device 12, that is configurable for detecting an acceleration of head worn device 12.
- Light emitter 36 may be implemented by any device, either standalone or part of head worn device 12, that is configurable for generating light, such as infrared radiation, and directing the generated light into user 14’s eyes for detecting the location of the irises/comeas of user 14 and/or detecting the direction of user 14’s gaze.
- the direction, phase, amplitude, frequency, etc., of light emitted by light emitter 36 may be controllable by processing circuitry 42 and/or gaze tracker 50.
- Light emitter 36 may include multiple light emitters (e.g., co-located and/or mounted at different locations on head worn device 12), or may include a single light emitter.
- Image sensor 38 may be implemented by any device, either standalone or part of head worn device 12, that is configurable for detecting images, such as images of user 14’s eyes, face, and/or images of the surrounding environment of user 14, such as images of signal source 30.
- Image sensor 38 may include multiple image sensors (e.g., co-located and/or mounted at different locations on head worn device 12 and/or on other devices/equipment in communication with head worn device 12 via communication interface 40), or may include a single image sensor.
- head worn device 12 may send, via the communication interface 40, sensor readings and/or data (e.g., image data, direction data, etc.) from one or more of microphones 18, microphones 20, display 22, speaker 26, microphone 28, accelerometer 34, light emitter 36, image sensor 38, communication interface 40, and processing circuitry 42 to additional head worn devices 12 (not shown), hand held device 31, and/or remote servers (e.g., an incident command server, not shown).
- sensor readings and/or data e.g., image data, direction data, etc.
- the microphones 18 and 20 may be mounted/arranged so as to optimize sound direction reception (e.g., of sound from signal source 30) and/or to optimize strength of the molding to head worn device 12.
- microphones 18 and 20 may be built into various portions of the head worn device 12, e.g., the front, back, sides, top, bottom, etc., of the head worn device 12, to optimize sound detection from a variety of directions.
- microphones 18 and 20 may be omnidirectional/non-directional, so as to detect sound in all directions.
- the microphones 18 and 20 may be directional, so as to detect sound in a particular direction relative to the head worn device 12.
- user interface 56 and/or display 22 may be a superimposed/augmented reality (AR) overlay, which may be configured such that user 14 of head worn device 12 may see through transparent lens 16, and images/icons displayed on display 22 appear to user 14 of head worn device 12 as superimposed on the transparent/translucent field of view (FOV) through lens 16.
- AR enhanced reality
- display 22 may be separate from lens 16.
- Display 22 may be implemented using a variety of techniques known in the art, such as a liquid crystal display built into lens 16, an optical headmounted display built into head worn device 12, a retinal scan display built into head worn device 12, etc.
- Gaze tracker 50 may detect (e.g., using image sensor 38) a point on each of user 14’s eyes corresponding to the center of the pupil in each eye. Gaze tracker 50 may calculate the relative movement/distance between the pupil center and the glint position for each eye. For example, gaze tracker 50 may calculate an optical axis, which is a vector connecting the pupil center, cornea center, and eyeball center. Gaze tracker 50 may calculate a visual axis, which is a vector connecting the fovea and the center of the cornea. The visual axis and the optical axis may intersect at the cornea center (also referred to as the nodal point of the eye).
- Gaze tracker 50 may utilize preconfigured/estimated physiological data (e.g., stored in memory 46) regarding eye dimensions (e.g., cornea curvature, eye diameter, distance between pupil center and cornea center, etc.) which may be based on the demographic information of user 14 (e.g., male users and female users may have different average/estimated eye dimensions), to estimate the direction and angle of the optical axis.
- eye dimensions e.g., cornea curvature, eye diameter, distance between pupil center and cornea center, etc.
- the angle of intersection of the glint vector and the pupil center vector may be used to estimate the angle between the optical axis and visual axis.
- gaze tracker 50 may estimate the visual axis, which corresponds to the user’s estimated gaze.
- gaze tracker 50 may utilize a regression and/or machine learning model to estimate the gaze direction of user 14.
- gaze tracker 50 e.g., using image sensor 38
- gaze tracker 50 may be configured to perform a calibration procedure.
- user 14 of head worn device 12 may initiate a calibration procedure, e.g., upon first use of the device.
- the calibration procedure may include, for example, displaying reference points on display 22, instructing (e.g., using visual and/or audio commands via user interface 56) the user 14 to direct his gaze at the reference points, and adjusting one or more parameters utilized by gaze tracker 50 based thereon.
- Other calibration procedures may be used to improve the accuracy of gaze tracker 50, such as using machine learning (e.g., based on datasets of multiple users of head worn device 12), without deviating from the scope of the present disclosure.
- Gaze tracker 50 may use any technique known in the art for determining/estimating the gaze of user 14 without deviating from the scope of the invention.
- Sound locator 52 may determine the location, relative direction and/or relative distance of signal source 30 to user 14 using a variety of techniques known in the art.
- sound locator 52 may apply a head related transfer function (HRTF) to the signals received by microphones, 18, which may be used to determine the left/right/horizontal orientation of the signal source relative to user 14.
- HRTF head related transfer function
- Sound locator 52 apply a HRTF to the signals received by microphones 20, compare the HRTF result for microphones 20 with the HRTF result of the microphones 18, and based on the comparison, determine an up/down/vertical orientation of the signal source 30 relative to user 14.
- HRTF head related transfer function
- Sound locator 52 may be configured to determine a vector from a suitable point, such as the bridge of user 14’s nose, through the center of the plane formed by the four points corresponding to the locations of microphones 18 and microphones 20; this vector may point to the origin of signal source 30.
- Antenna array 60 may be implemented by any device, either standalone or part of hand held device 31, that is configurable for detecting beacon signals from a radio beacon, e.g., a radio beacon emitted by an alarm device attached to the equipment of a downed first responder.
- Antenna array 60 may include one or more directional antennas used to follow a radio beacon.
- display 22 may continually/periodically refresh, providing updated cues to user 14 regarding the location of signal source 30, until the location is reached and/or until user 14 terminates the procedure (e.g., by a voice command or toggling a button).
- user 14 may activate/deactivate the search procedure by a voice command (e.g., via microphone 28 and/or user interface 56) and/or by toggling a switch/button (e.g., in communication with user interface 56).
- a voice command e.g., via microphone 28 and/or user interface 56
- a switch/button e.g., in communication with user interface 56.
- one or more components of head worn device 12, such as gaze tracker 50, sound locator 52, sound classifier 54, and/or user interface 56 may be located in/performed by separate circuitry located elsewhere on user 14’s body, such as in a hand held device 31 or remote in communication with head worn device 12, e.g., via a wired or wireless connection with communication interface 40.
- sound locator 52 may be configured to utilize thermal imaging and/or other visual data (e.g., received from image sensor 38 and/or from other image sensors, such as an image sensor 62 in a separate hand held device 31 in communication with head worn device 12 via communication interface 40) in determining the location of signal source 30.
- visual data may include images of light outside the visible spectrum.
- hand held device 31 in communication with head worn device 12 is configured to follow a radio beacon, e.g., using antenna array 60.
- Hand held device 31 is configured to be swept back and forth, up and down, etc., by user 14, to try to identify the maximum beacon strength, e.g., by comparing measurements of signal strength detected by antenna array 60.
- hand held device 31 may determine the radio beacon’s direction by taking multiple directional measurements (e.g., with antenna array 60), which may result in a virtual conical structure as the user 14 approaches the source of the beacon.
- the locator 74 may be configured to identify the part of the conic segmented detected and calculate it backward to identify the apex of the cone, which represents the location of signal source 30.
- signal source 30 may be a personal alert/distress alarm device (e.g., a Scott Pak-Alert Personal Alert Safety System (PASS) device), worn by a downed first responder, which may give off an audible sound (e.g., a piercing sound) and/or may emit a radio signal/beacon when activated.
- PASS Personal Alert Safety System
- Detecting a radio signal in addition to/as an alternative to a detecting a sound signal may be advantageous, for example, in scenarios where detecting an audible sound signal is impractical, such as where a downed first responder is at least partially submerged underwater, where the personal alert device’s sound chambers have been occluded by debris, due to environmental conditions, etc. Detecting a radio signal in addition to detecting a sound signal may thus improve the accuracy of estimating the location/direction of signal source 30.
- PASS Personal Alert Safety System
- the hand held device 31 and/or head worn device 12 may be configured to determine the location of the personal alert device based on characteristics of the audible sound signal and/or the emitted radio signal, and the head worn device 12 may be configured to determine whether the user 14 is looking and/or facing in the direction of the location of the personal alert device and/or may be configured to direct the user 14 to the location of the signal source 30, even in scenarios where the audible sound signal cannot be detected and/or where visibility is at least partially blocked.
- Locator 74 may detect signals (e.g., sound signals/waves, radio signals/waves, etc.) at multiple various points throughout the environment (i.e., 3-dimensional space) and record characteristics of those signals, such as signal strength, power, amplitude, frequency, noise, etc. Locator 74 may construct/utilize a 3-dimensional model based on the detected signals to determine/estimate the origin of the signal source 30 and/or to guide user 14 to the signal source 30. As one non-limiting example, the signal may be modeled as a 3 -dimensional cone, and locator 74 may utilize one or more formulas known in the art, such as the equation for the curved surface of a right cone, to determine one or more characteristics of the signal.
- the hand held device 31 may detect signals as the user 14 moves through the environment, and/or user 14 may intentionally move the hand held device (e.g., in a sweeping motion) in order to gather detected signal data points at various locations relative to user 14.
- gaze tracking or gaze direction determination techniques may utilize the commonality in eye physiology, musculature, orbit, geometry, etc., in human populations to fix a point and then eye track the pupil to locate the second point.
- the two points may define a line/vector which may be defined by an equation.
- the first point may be a center of the eye’s pupil
- the second point may be a light reflection on the cornea, e.g., a reflection of the light emitted by light emitter 36 and captured by image sensor 38.
- Eye/face/image/location data associated with user 14 may be stored, e.g., in head worn device 12, and may be used to generate/refine a machine learning model, improving the accuracy of prediction over time as more data is collected. Additionally, or alternatively, head worn device 12 may use synthetic datasets, e.g., with large participant sources. Such datasets may include infrared image samples of participants’ faces and/or eyes. Using datasets/machine learning enables the gaze tracker 50 to adapt/tune/calibrate to a wide variety of users 14. In some non-limiting embodiments, gaze tracker 50 may utilize a variety of computer vision libraries, such as Python OpenCV. While gaze tracker 50 can be implemented using any suitable hardware and/or software arrangement, some embodiments may utilize/execute software code written in C++ as well as Python, making it more adaptable to being deployed on a variety of microcontrollers.
- computer vision libraries such as Python OpenCV.
- gaze tracker 50 is configured to estimate the gaze of user 14, as illustrated in FIG. 7.
- the diameter of eye 82a (or 82b) is a known and/or estimated value. The estimate may be based on population averages, demographic information of user 14, and/or machine learning techniques.
- the distance between the line representing the diameter of the eye and the line representing the light reflection may be measured, e.g., using image sensor 38.
- the angle between the center of the eye and the glint, as well as the characteristics light reflection line and the optical axis may be determined.
- a neural network may be initially trained using publicly available datasets, and may be further trained/refined for a particular group of users 14 (e.g., employees of a particular firefighting service) by gathering data from actual use and/or from an artificial training/calibration scenario, e.g., by setting up a sound target to emulate signal source 30, and instructing a user 14 to go through a pattern of movements, such as a facepiece fit sequence.
- the head worn device 12 may gather data (e.g., images of the user 14’s eyes, face, environment, signal/location data of signal source 30, etc.) to train the neural network.
- a right triangle 84 is formed, with no smaller right triangle to the right of triangle 84.
- the tangent, a is equal to the ratio of the opposite side to the adjacent side, which is equal to slope, b. of the light reflection ray.
- head worn device 12 may determine the glint/optical axis from multiple perspectives, which may improve the accuracy of the estimation.
- sound locator 52 utilizes a HRTF.
- a HRTF is a measure of the difference of hearing between the listener’s (e.g., user 14) right and left ears. Placing the microphones 18 and/or 20 on either side of the facepiece, sound locator 52 may simulate a simplified head-form hearing system without needing to account for pinna structure, ear internal inefficiencies, etc.
- the HRTF may consider two signal collection points separated by a space and use that information to determine/estimate a signal origin location. Utilizing two pairs of microphones 18 and/or 20 may further improve accuracy and/or may provide additional information, such as how user 14 ’s head is tilted and/or pointed with respect to signal source 30.
- the particular formulas utilized for the HRTF are known in the art and beyond the scope of the present disclosure.
- head worn device 12 may include two or more sets of microphones, e.g., microphones 18 and 20.
- Sound locator 52 may apply a transfer function to the top two microphones 18 to derive a left-right orientation. Sound locator 52 may apply a transfer function to the bottom two microphones 20, compare to the upper two microphones 18, and derive an up-down orientation therefrom For example, if the two microphones 18a and 20a on the left side of user 14’s head detect a comparable sound intensity that is higher than the two microphones 18b and 20b on the right side of user 14’s head, it may be predicted that the signal source 30 is to user 14’s left.
- sound locator 52 may determine the direction of signal source 30 relative to user 14.
- microphones 18 and 20 may be arranged on head worn device 12 such that the center of a plane 86 formed by the microphone 18 and 20 locations is in the general location of the bridge of user 14’s nose.
- the geometric position of plane 86 may be defined by the microphones 18 and 20.
- the location of the bridge of the nose cup may be the other point defining the line, which points to the direction of signal source 30.
- the degree of adjustment may follow a trial and error algorithm, with the process halting once the iterations of slope change less than a threshold value (e.g., 10%) after multiple trials, for example.
- a threshold value e.g. 10%
- a 2-dimensional or 3-dimensional least squares algorithm or other various geometry calculations known in the art, the particular details of which are beyond the scope of the present disclosure, may be employed to iteratively improve the accuracy.
- the ability of user 14 to adjust accurately can be assisted by providing a display 22 attached to an accelerometer 34, as well as a representation in display 22 (e.g., as an AR overlay) of the two lines in the distance. Using this display 22, user 14 may adjust his gaze to the direction of signal source 30. In some cases, the two lines may not be coincident or at least parallel.
- user 14 may have his face pointed in the correct direction of signal source 30, but may be gazing up, down, right, or left of signal source 30.
- user 14 may be confused as to the source of sound, and may be facing the wrong direction altogether.
- head worn device 12 may be detecting an echo or a sound reflection.
- Sound locator 52 may be configured to compensate for sound reflections. For example, sound locator 52 may consider the amplitude of the sound will change after the reflection, but the frequency will not change.
- Sound locator 52 may utilize additional microphones (e.g., installed on the sides/back of head worn device 12), compare the amplitudes and/or frequencies of the sounds received from the side/back microphone with that received from the front microphones (18 and 20), and determine whether the front microphones 18 and 20 are detecting a sound or an echo of a sound. If the lines share a common slope, and direction, then sound locator 52 may determine that user 14 is looking in the general vicinity of the source of the sound. Sound locator 52 may utilize vertical angles and auxiliary lines to further analyze the lines, according to geometric formulas known in the art which are beyond the scope of the present disclosure.
- FIGS. 10 and 11 illustrate another example of using HRTF to determine sound direction and gaze detection.
- the sound direction (from signal source 30) may be modeled as a straight line.
- the gaze direction may be modeled as a straight line as well.
- the head worn device 12 uses HRTF to describe/determine the sound direction/origin and the direction the user 14 is looking, and instructs the user 14 to change gaze direction in order to make the two lines parallel/converging, and this condition is indicative of the user 14 gazing in the direction of signal source 30.
- Head worn device 12 may provide user 14 (e.g., via display 22) with multiple estimates of the location of signal source 30, e.g., by displaying multiple vectors overlayed on display 22, and user 14 may determine which vector to follow, for example, user 14 may have just passed through a room and did not find the signal source 30 in that room, and user 14 may use that information to decide to ignore a vector on display 22 pointing user 14 back to that room, and instead will follow a vector pointing to a new room which user 14 has not previously entered.
- multiple head worn devices 12 associated with multiple users 14 may cooperate (e.g., by wireless transmitting data directly or indirectly with one another, by communicating with a remote server, etc.) to improve the accuracy of the estimated location of signal source 30.
- the detected signals e.g., sounds waves
- each head worn device 12 may be distributed to the other head worn devices 12 of the first responder team, each of which may utilize the additional data to improve the accuracy of the signal source 30 location detection.
- one or more users 14 in the first responder team may be equipped with hand held devices 31, which may detect radio signals emitted by signal source 30, as described herein, and the head worn devices 12 may utilize the location information generated by one or more of the multiple hand held devices 31 (e.g., as determined by locators 74).
- the location data e.g., geographical coordinates, distance/angle information, etc.
- the location data may be shared with the other head worn devices 12 of the first responder team.
- FIG. 12 is a flowchart of an example process in a head worn device 12 according to some embodiments of the invention.
- One or more blocks described herein may be performed by one or more elements of head worn device 12, such as by one or more of processing circuitry 42, microphones 18, microphones 20, display 22, speaker 26, microphone 28, accelerometer 34, light emitter 36, image sensor 38, communication interface 40, processing circuitry 42, processor 44, memory 46, software 48, gaze tracker 50, sound locator 52, sound classifier 54, and/or user interface 56.
- Head worn device 12 is configured to receive (Block S100) an audio signal detected by at least one microphone (e.g., microphones 18 and microphones 20), the audio signal originating from a signal source 30.
- a microphone e.g., microphones 18 and microphones 20
- the user instruction is determined based on at least one of a relative distance and a relative direction from the user 14 to the signal source 30. In some embodiments, the user instruction indicates a location and/or direction for the user to look.
- determining a gaze direction of the user 14 includes determining a plurality of features based on the image data using a machine learning model to predict gaze direction based on the determined plurality of features.
- the processing circuitry is further configured to receive, from the signal source, a radio signal, the determining the location of the signal source being further determined based on the received radio signal.
- the processing circuitry 42 is further configured to determine a plurality of features based on the received audio signal.
- the processing circuitry is configured to determine a sound classification based on the determined plurality of features using a machine learning model to predict a sound classification of the signal source 30 based on the determined plurality of features.
- the head worn device 12 includes a user interface 56.
- the user interface 56 is configured to display the user instruction as an augmented reality (AR) indication.
- AR augmented reality
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Software Systems (AREA)
- General Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Mathematical Physics (AREA)
- Computing Systems (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Business, Economics & Management (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Emergency Management (AREA)
- Life Sciences & Earth Sciences (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Security & Cryptography (AREA)
- Acoustics & Sound (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Medical Informatics (AREA)
- Pulmonology (AREA)
- User Interface Of Digital Computer (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263429012P | 2022-11-30 | 2022-11-30 | |
| PCT/IB2023/061747 WO2024116021A1 (en) | 2022-11-30 | 2023-11-21 | Head related transfer function application to sound location in emergengy scenarios |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4626558A1 true EP4626558A1 (de) | 2025-10-08 |
Family
ID=91323064
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23896993.5A Pending EP4626558A1 (de) | 2022-11-30 | 2023-11-21 | Anwendung einer kopfbezogenen übertragungsfunktion zur klanglokalisierung in notfahrszenarien |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4626558A1 (de) |
| CN (1) | CN120379728A (de) |
| WO (1) | WO2024116021A1 (de) |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20160015089A (ko) * | 2014-07-30 | 2016-02-12 | 쉐도우시스템즈(주) | 소방관용 스마트 보호 헬멧 시스템 및 장치 |
| US10610708B2 (en) * | 2016-06-23 | 2020-04-07 | 3M Innovative Properties Company | Indicating hazardous exposure in a supplied air respirator system |
| US10643443B2 (en) * | 2016-12-30 | 2020-05-05 | Axis Ab | Alarm masking based on gaze in video management system |
| KR102172894B1 (ko) * | 2018-01-17 | 2020-11-02 | 건국대학교 산학협력단 | 위치 기반 스마트 마스크 및 스마트 마스크 시스템 |
| CN214512317U (zh) * | 2021-02-01 | 2021-10-29 | 中航华东光电有限公司 | 消防面罩内显示组件的简易安装机构 |
-
2023
- 2023-11-21 EP EP23896993.5A patent/EP4626558A1/de active Pending
- 2023-11-21 CN CN202380081916.1A patent/CN120379728A/zh active Pending
- 2023-11-21 WO PCT/IB2023/061747 patent/WO2024116021A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| CN120379728A (zh) | 2025-07-25 |
| WO2024116021A1 (en) | 2024-06-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11514207B2 (en) | Tracking safety conditions of an area | |
| Liu et al. | BlinkListener: " Listen" to Your Eye Blink Using Your Smartphone | |
| US10957299B2 (en) | Acoustic transfer function personalization using sound scene analysis and beamforming | |
| US8995678B2 (en) | Tactile-based guidance system | |
| CN109932054B (zh) | 可穿戴式声学检测识别系统 | |
| US20060210111A1 (en) | Systems and methods for eye-operated three-dimensional object location | |
| US20050175218A1 (en) | Method and apparatus for calibration-free eye tracking using multiple glints or surface reflections | |
| US10848891B2 (en) | Remote inference of sound frequencies for determination of head-related transfer functions for a user of a headset | |
| US11729573B2 (en) | Audio enhanced augmented reality | |
| JPH07500661A (ja) | 音響探査装置 | |
| KR102713524B1 (ko) | 머리 전달 함수에 대한 헤드셋 효과의 보상 | |
| US12028419B1 (en) | Systems and methods for predictively downloading volumetric data | |
| CN110991336A (zh) | 一种基于感官替代的辅助感知方法和系统 | |
| US20240177824A1 (en) | Monitoring food consumption using an ultrawide band system | |
| US20250341636A1 (en) | Expressions from transducers and camera | |
| WO2024116021A1 (en) | Head related transfer function application to sound location in emergengy scenarios | |
| US11816886B1 (en) | Apparatus, system, and method for machine perception | |
| WO2023076824A1 (en) | Response to sounds in an environment based on correlated audio and user events | |
| CN209525006U (zh) | 可穿戴式声学检测识别系统 | |
| US20260075181A1 (en) | Electronic Devices with Gaze Tracking Circuitry | |
| WO2024161299A1 (en) | Wearable device for visual assistance, particularly for blind and/or visually impaired people | |
| WO2018084227A1 (ja) | 端末装置、動作方法及びプログラム | |
| IT202300006183A1 (it) | Sistema di tracciamento delle posizioni, delle orientazioni e delle traiettorie nello spazio, per la fruizione, la sicurezza preventiva e l’interazione assistita e inclusiva. | |
| Todd et al. | EYESEE; AN ASSISTIVE DEVICE FOR BLIND NAVIGATION WITH MULTI-SENSORY AID |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250603 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |