WO2024184232A1 - Device, method and computer program - Google Patents

Device, method and computer program Download PDF

Info

Publication number
WO2024184232A1
WO2024184232A1 PCT/EP2024/055391 EP2024055391W WO2024184232A1 WO 2024184232 A1 WO2024184232 A1 WO 2024184232A1 EP 2024055391 W EP2024055391 W EP 2024055391W WO 2024184232 A1 WO2024184232 A1 WO 2024184232A1
Authority
WO
WIPO (PCT)
Prior art keywords
user interface
virtual screen
camera
markers
holographic
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2024/055391
Other languages
French (fr)
Inventor
Anthony ANTOUN
Sebastien VANDEUN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Depthsensing Solutions NV SA
Sony Semiconductor Solutions Corp
Original Assignee
Sony Depthsensing Solutions NV SA
Sony Semiconductor Solutions Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Depthsensing Solutions NV SA, Sony Semiconductor Solutions Corp filed Critical Sony Depthsensing Solutions NV SA
Publication of WO2024184232A1 publication Critical patent/WO2024184232A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/03Arrangements for converting the position or the displacement of a member into a coded form
    • G06F3/0304Detection arrangements using opto-electronic means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/03Arrangements for converting the position or the displacement of a member into a coded form
    • G06F3/033Pointing devices displaced or positioned by the user, e.g. mice, trackballs, pens or joysticks; Accessories therefor
    • GPHYSICS
    • G03PHOTOGRAPHY; CINEMATOGRAPHY; ANALOGOUS TECHNIQUES USING WAVES OTHER THAN OPTICAL WAVES; ELECTROGRAPHY; HOLOGRAPHY
    • G03HHOLOGRAPHIC PROCESSES OR APPARATUS
    • G03H1/00Holographic processes or apparatus using light, infrared or ultraviolet waves for obtaining holograms or for obtaining an image from them; Details peculiar thereto
    • G03H1/22Processes or apparatus for obtaining an optical image from holograms
    • G03H1/2294Addressing the hologram to an active spatial light modulator
    • GPHYSICS
    • G03PHOTOGRAPHY; CINEMATOGRAPHY; ANALOGOUS TECHNIQUES USING WAVES OTHER THAN OPTICAL WAVES; ELECTROGRAPHY; HOLOGRAPHY
    • G03HHOLOGRAPHIC PROCESSES OR APPARATUS
    • G03H2226/00Electro-optic or electronic components relating to digital holography
    • G03H2226/05Means for tracking the observer

Definitions

  • a holographic user interface is a computer input method that utilizes a projected image presented to the user for the purpose of user interaction.
  • Holographic interfaces may for example utilize projectors to create a virtual 3D image in space.
  • a holographic display may for example be a 3D display that utilizes light diffraction to display a 3D image to the user.
  • the virtual image provided by a holographic user interface is constructed only of light.
  • the disclosure provides a holographic user interface comprising: a projector configured to project a virtual screen into a space; at least one first camera configured to capture images of a user of the holographic user interface; and circuitry configured to perform calibration processing by determining the position of markers on the virtual screen; and determine a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen.
  • the disclosure provides a method of calibrating a holographic user interface, the method comprising: projecting a virtual screen into a space; obtaining at least one image of a user of the holographic user interface from at least one first camera; performing calibration processing by determining the position of markers on the virtual screen; and determining a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen.
  • the disclosure provides a program comprising instructions, the instructions being configured to, when operated by a processor, perform the method mentioned above.
  • Fig.1 is a perspective view of an exemplary holographic user interface according to an embodiment of the invention
  • Fig.2 is a system overview of a holographic user interface according to an exemplary embodiment
  • Fig.3 is a system overview block diagram, which shows the overview of the holographic user interface and the internal structure of a control device of the holographic user interface according to an exemplary embodiment
  • Fig.4 is a flow diagram illustrating the process of detection of the fingertip position by the control device of the holographic user interface of Fig.3
  • Fig.5a and Fig.5b illustrate a process of determining a perceptive origin of a user of the interactive user interface according to an exemplary embodiment
  • Fig.6a and Fig.6b illustrate a process of determining the pointer location according to an exemplary embodiment
  • Fig.1 is a perspective view of an exemplary holographic user interface according to an embodiment of the invention
  • Fig.2 is a system overview of a holographic user interface according to
  • FIG. 7 illustrates a calibration process of the holographic user interface according to a first embodiment
  • Fig.8 illustrates the transformation of the calculated pointer location in the coordinate system of the depth camera to 2D coordinates in a 2D coordinate system of the visual screen
  • Fig.9 provides a schematic diagram of coordinate systems involved in the holographic user interface
  • Fig. 10 illustrates an example of a perception-based calibration process of the holographic user interface according to an exemplary embodiment.
  • Fig.11 illustrates an example of a multi-user embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS Before a detailed description of the embodiments under reference of Fig.1 to Fig.10, general explanations are made.
  • the embodiments disclose a holographic user interface comprising circuitry configured to control a projector to project a virtual screen into a space; obtain at least one image of a user of the holographic user interface from at least one first camera; perform calibration processing by determining the position of markers on the virtual screen, and determine a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen.
  • the circuitry may include a processor, a memory (RAM, ROM or the like), a storage, input means (mouse, keyboard, camera, etc.), output means (display (e.g.
  • circuitry may include sensors for sensing still image or video image data (image sensor, camera sensor, video sensor, etc.), and the like.
  • the circuitry may be configured to display the markers on the virtual screen.
  • the markers may for example be visual clues displayed on the virtual screen.
  • the circuitry may be configured to perform the calibration processing by instructing the user to point at the markers displayed on the virtual screen.
  • the holographic user interface may further comprise at least one second camera configured to acquire one or more images of the virtual screen.
  • the circuitry may be configured to detect the location of the markers on the virtual screen by analyzing the one or more images of the virtual screen obtained by the at least one second camera.
  • the markers may for example be machine readable markers.
  • the markers may be ChArUco-markers, AprilTag-markers or checker markers.
  • the markers may be displayed at the corners of the virtual screen.
  • the at least one first camera may be a depth camera.
  • the circuitry may further be configured to determine a fingertip position of the user in the space based on the at least one image obtained from the at least one first camera.
  • the circuitry may for example be configured to determine the fingertip position on the basis of a posture of a user’s hand detected by analyzing depth images obtained from a depth camera.
  • the circuitry may further be configured to determine a perceptive origin of the user based on the at least one image obtained from the first camera.
  • the perceptive origin may for example be determined based on a parameter describing the eye dominance of the user.
  • the eye dominance of the user may for example be determined by any one of: a distance hole-in-the-card test, a near convergence test or a Pediatric Eye Disease Investigator Group (PEDIG) fixation preference test.
  • the degree of eye dominance may for example be a predetermined parameter or it may be a dynamic parameter that is determined using live detection techniques. For example, eye opening percentage calculation might give a good estimation of the perceptive origin in real time.
  • the circuitry may further be configured to determine the pointer location based on the fingertip position and based on the perceptive origin of the user. In a preferred embodiment, the pointer location is an intersection between a pointing line defined by the fingertip position and the perceptive origin and the display plane of the virtual screen. The circuitry may further be configured to activate an UI element pointed by the user on the virtual screen on a basis of the pointer location.
  • the circuitry may be configured to activate the UI element pointed by the user on the virtual screen when a distance between the fingertip and the virtual screen is smaller than a predetermined threshold.
  • the embodiments disclose a method of calibrating a holographic user interface, the method comprising: projecting a virtual screen into a space; obtaining at least one image of a user of the holographic user interface from at least one first camera; performing calibration processing by determining the position of markers on the virtual screen; determining a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen.
  • the embodiments also disclose a computer program product which stores instructions, the instructions, when executed by a processor, performing the processes described above.
  • Holographic user interface Fig. 1 is a perspective view of an exemplary holographic user interface 10 which acts as a computer input method that utilizes a projected virtual image presented to the user for the purpose of user interaction.
  • the holographic user interface 10 may also be called a natural user interface (NUI), or touchless/contactless air screen, or an augmented reality pointing interface, or a digital signage display.
  • NUI natural user interface
  • a holographic user interface 10 is shown in which a spatial image Ib is generated on the display plane of a virtual (floating) screen 4 in a given space.
  • a display device (not shown; see 2 in Fig.2) displays a visible light producing image Ia on an image display surface 1A in front of the user.
  • An imaging optical system 3 forms a spatial image Ib from visible light emitted from the image display surface 1A and outputs the spatial image Ib to the display plane of the virtual screen 4.
  • the user views the virtual image Ib and designates graphical elements with which the user may interact with a finger, a pointing rod, a pen, or the like, thus enabling input without having to contact a physical display device.
  • the virtual image provided by the holographic user interface is constructed only of light. The interaction with the virtual image that provides the holographic user interface is typically detected via sensors (not shown in Fig.1; see Fig.2).
  • Fig.2 is a system overview of a holographic user interface 10 according to an exemplary embodiment.
  • the holographic user interface 10 comprises a display device 1, a control device 2, an imaging optical system 3 and a camera 5.
  • the display device 1 may be any type of a flat display panel, such as a LCD, CRT, LED, plasma display or the like, wherein visible light is emitted in the image display surface 1A, creating an image Ia.
  • the imaging optical system 3 receives visible light emitted from the image display surface 1A as an input and reflects the image Ia to form a spatial image Ib, which is free of reflected light from another light source, on the display plane of the virtual screen 4 in the given space.
  • the imaging optical system 3 deviates the visible light producing image Ia from the image display surface 1A to form a spatial image Ib within the spatial image display plane of a virtual (floating) screen 4.
  • the camera 5 is arranged to capture images of the user in which the user’s hand(s) and eyes can be seen simultaneously. In particular, the camera 5 captures a pointing gesture performed by the user’s hand(s) and fingers with respect to the display plane of the spatial image. Further, the camera 5 captures the left and right eye of the user gazing at the virtual image.
  • the camera 5 outputs image data corresponding to the captured images to the control device 2.
  • the control device 2 receives the image data from the camera 5 and calculates a pointer location as a position designated by the user in the display plane of the spatial image 4, as described in detail later with respect to Figs.4, 5a and b and 6.
  • the control device 2 further controls display of the image Ia on the image display surface 1A of the display device 1 based on the calculated pointer location.
  • the display unit 1, the control unit 2 and the imaging optical system 3 form together the projector 11.
  • the display device 1 is represented as a flat display panel.
  • the display device 1 may be curved, wherein the imaging optical system 3 is also designed to account for the curved geometries.
  • the imaging optical system 3 may comprise components similar to the image formation means as described in WO 2014/038303 A1 and WO 2014/073650 A1.
  • the image formation means has a first and a second light control panel, which are positioned either in contact or in proximity with a plurality of first and second band-shaped optical reflection parts, whereby the pathway of light is directed to preserve the image spatially from one plane to another.
  • Multiple embodiments of structures comprising surfaces for reflecting light are disclosed and are incorporated by reference.
  • Other technologies for generating a virtual screen may for example include holographic techniques.
  • the camera 5 captures images of the user in which the user’s hand(s) and eyes can be seen simultaneously.
  • the camera 5 may be a depth camera that may capture a depth image of a scene.
  • the depth image may include a two-dimensional (2-D) pixel array of the captured scene where each pixel in the 2-D pixel array may represent a depth value such as a distance of an object in the captured scene from the camera.
  • Fig.3 is a system overview block diagram, which shows the overview of the holographic user interface 10 and the internal structure of the control device 2 according to an exemplary embodiment.
  • the holographic user interface 10 of Fig.3 comprises the display device 1, the control device 2, the imaging optical system 3, and the depth camera 5 as described with reference to Figs.1 and 2.
  • the camera 5 captures images of the user in which the user’s hand(s) and eyes can be seen simultaneously.
  • the images captured by the camera show a pointing gesture of the user’s hand(s) and fingers with respect to the display plane of the virtual screen (4 in Figs.1 and 2). Further, the images captured by the camera 5 show a left eye and a right eye of the user gazing at the virtual image.
  • the camera 5 outputs image data to the control device 2.
  • Calibration cameras 7a and 7b are provided for the purpose of the calibration process of the holographic user interface described in detail with reference to Fig.7.
  • the calibration cameras 7a and 7b may for example be RGB cameras that are configured to face the virtual screen (4 in Fig.2) and to capture images of the virtual screen.
  • the calibration cameras 7a and 7b output image data to the control device 2.
  • the calibration cameras 7a and 7b are, however, optional.
  • the calibration processing is not performed based on images from calibration cameras 7a, b, but instead based on a Perception-based calibration such as described with regard to Fig. 10 below.
  • the control device 2 comprises an image receiving section 20, a calibration processing section 21, a perceptive origin detection section 22, a fingertip detection section 23, a hand skeleton model storage section 24, a pointer location determination section 25, and a display control device 26. Via the image receiving section 20, the control device 2 receives image data from the camera 5 and image data from the calibration cameras 7a and 7b as an input.
  • the image receiving section 20 receives the image data from the camera 5 and transmits the image data to the perceptive origin detection section 22 and the fingertip detection section 23.
  • the image receiving section 20 receives the image data from the calibration cameras 7a and 7b and transmits the image data to the calibration processing section 21.
  • the calibration processing section 21 based on the image data received from the calibration cameras 7a, b via the image receiving section 20 (see Fig. 7 and corresponding description) or by means of perception-based calibration (see Fig. 10 and corresponding description), performs a calibration process of the holographic user interface.
  • the fingertip detection section 23 determines a fingertip position of the user in the space as described in detail in Fig. 4 below.
  • the hand skeleton model storage section 24 stores hand skeleton model data used by the fingertip detection section 23 to determine the fingertip position.
  • the pointer location determination section 25 receives the processing results from the perceptive origin detection section 22 and the fingertip detection section 23 and calculates a pointer location based on the received processing results (see Fig. 6).
  • the pointer location determination section 25 outputs the processing result to the display control device 26.
  • the display control device 26 receives the processing result from the pointer location determination section 25 and outputs a control signal to control the display of a virtual image of a user interface (Ib in Figs.1 and 2) based on the received processing result.
  • the display control device 26 may comprise a graphic card.
  • the camera 5 is a depth camera, for example a Time of Flight (ToF) camera.
  • TOF Time of Flight
  • alternative imaging units for example a stereoscopic camera
  • the camera 5 may contain a depth image sensor that captures depth data and a depth processor that processes the depth data to generate a depth map. The processing steps performed by the depth processor are dependent upon the technique used by the depth image sensor.
  • the depth image sensor may rely on one of several different sensor technologies. Among these sensor technologies are time-of-flight (TOF) technologies, such as indirect time-of-flight (iTOF) and direct time-of-flight (dTOF), structured light, radar, stereoscopic cameras, active stereoscopic sensors, and the like.
  • TOF time-of-flight
  • iTOF indirect time-of-flight
  • dTOF direct time-of-flight
  • Fingertip detection Fig. 4 is a flow diagram illustrating the process of detection of the fingertip position by the control device 2 of the holographic user interface of Fig.3. The process may for example be implemented by the fingertip detection section (23 in Fig.3).
  • the fingertip detection section receives the image data (depth images) from the camera (5 in Fig.2).
  • the fingertip detection section processes the depth image data acquired from the camera to determine the position of the user’s hands and fingers relative to the virtual display.
  • algorithms may be applied to determine a hand skeleton model (24 in Fig. 3) based on the image data.
  • the fingertip detection section may apply an appropriate algorithm to depth image data to enable the system to utilize close-range depth data.
  • the fingertip detection section may be enabled to process depth image data in accordance with close-range optical settings and requirements to determine a presence, position, movement, etc. of an object.
  • the fingertip detection section may execute software code or algorithms for close range tracking to enable detection and/or tracking of movements of a user’s hands and finger
  • the output of the fingertip detection section can be a representation of the skeleton of the user’s hand.
  • the hand skeleton representation can include the positions of the joints of the skeleton, and may also include the rotations of the joints, relative to a center point.
  • the skeleton model (24 in Fig. 3) may be used to further improve the quality of the tracking and assign positions to joints which were not detected in the earlier stages, e.g. due to occlusions or because parts of the hand are outside the field of view of the camera.
  • the fingertip detection section (23 in Fig.3), on the basis of the hand skeleton model acquired in S2, determines the fingertip position f by using hand skeleton model data stored in the hand skeleton model storage section (24 in Fig.3).
  • Determination of a skeleton model of the user’s hand as performed in step S2 may for example be implemented according to any state-of-the-art techniques such as machine learning, or the like.
  • U.S. patent application No. 13/768,835 describes a tracking module that processes depth data of a user performing movements, for example, movements of the user’s hand and fingers.
  • US 2012/0327125 A1 describes a method for tracking a user’s hands and fingers based on depth images captured from a depth camera, and using the tracked data to control a user’s interaction with devices.
  • a fingertip position f is determined based on camera images and based on a hand skeleton model.
  • the camera 5 may detect more than one object simultaneously. The detection of multiple objects, for example two fingers of the user’s hand, may be useful, for example, in performing a “pinch-in” or “pinch-out” command.
  • Perceptive origin detection Fig.5a and 5b illustrate a process of determining a perceptive origin of a user of the interactive user interface according to an exemplary embodiment.
  • This process may for example be implemented by the perceptive origin detection section 22 of Fig.3.
  • the process calculates the perceptive origin (eye origin) of the user based on the image data received from the camera 5.
  • the perceptive origin O is located on the line connecting the left eye L and the right eye R (centroid of the pupils) of the viewer and is indicative of the user’s preference to use one eye more than the fellow eye to accomplish a task.
  • the user views items displayed on the virtual screen.
  • the camera 5 is mounted opposite the eye opening(s) so as to capture images of the user’s left eye (L) and the user’s right eye (R). It is conceivable that other embodiments could use multiple cameras.
  • the camera 5 captures images of the left and right eye of the user gazing at the virtual screen and transmits image data to the perceptive origin detection section (22 in Fig. 3).
  • the perceptive origin detection section processes the depth image data acquired from the camera 5 to determine a degree of eye dominance ⁇ ⁇ ⁇ and, based on the processing result, determines the perceptive origin O. As shown in Fig.
  • the degree of eye dominance ⁇ ( ⁇ ) can change over time and be recalculated in real time.
  • the ⁇ ( ⁇ ) must not necessarily have a time dependence. It can also be set as a predefined value ⁇ constant over time.
  • a value of ⁇ greater than 0,5 indicates an increasingly stronger dominance of the right eye and a perceptive origin that is correspondingly increasingly closer to the right eye, up to a value of 1 where only the right eye is active and the perceptive origin is located at the position of the right eye.
  • a value of ⁇ of less than 0,5 indicates a dominance of the left eye and a perceptive origin that is correspondingly closer to the left eye, a value of 0 indicating that only the left eye is active with the perceptive origin being located at the position of the left eye.
  • the parameter ⁇ which describes the degree of eye dominance may for example be known in advance and input to the holographic user interface as a predefined parameter. That is, the dominant eye may be known a priori. For example, the user may enter this information manually. Alternatively, the parameter ⁇ which describes the degree of eye dominance may be determined in a calibration stage. The determination of the dominant eye is not limited to any particular technique.
  • Eye dominance Various tests are known to determine eye dominance, such as the distance hole-in-the-card test, the near convergence test, or the Pediatric Eye Disease Investigator Group (PEDIG) fixation preference test (see, e.g., Rice M. et al., “Results of ocular dominance testing depend on assessment method”, Journal of American Association for Pediatric Ophthalmology and Strabismus. 2008;12(4):369–465).
  • PEDIG Pediatric Eye Disease Investigator Group
  • Whether an eye is dominant over the other eye may for example be determined based on how the respective eyes move and/or based on respective deviations between the points of regard for the respective eyes and a reference point when the user is prompted to look at the reference point. Eye dominance can for example lead to externally distinguishable features of the person’s eye movements.
  • one eye may exhibit a lower range of movement than the other, or lag the other, or stay still, or look in a different apparent direction to the direction of actual gaze of the user.
  • the perceptive origin detection section may determine the perceptive origin based on the eye openings.
  • the dominant eye can be defined, for example, as the eye with the wider range of movement, the eye which reaches a gaze direction appropriate to a displayed stimulus the closest and/or the most quickly, or the eye which holds a position relating to a displayed stimulus.
  • the degree of eye dominance ⁇ ( ⁇ ) can change over time and be recalculated in real time.
  • Fig. 6a and 6b illustrate a process of determining the pointer location according to an exemplary embodiment.
  • Fig. 6a schematically illustrates the positional relationship between the perceptive origin O, the fingertip position f and the pointer location p.
  • the perceptive origin O which is determined by the perceptive origin detection section (22 in Fig.3) as described above with reference to Figs.5a and 5b, is located on the line connecting the left eye L and the right eye R of the viewer.
  • the fingertip position f which is determined by the fingertip detection section (23 in Fig. 3) as described above with reference to Fig. 4, is located in the vicinity of the virtual screen 4.
  • the perceptive origin O and the fingertip position f define a pointing line PL, as shown in Fig.6a.
  • the pointer location p is determined as the intersection of this pointing line PL and the spatial image display plane of the virtual screen 4.
  • Fig.6b is a flow chart illustrating an operation process of determining the pointer location and display control according to an exemplary embodiment.
  • This process may, for example, be implemented by the pointer location determination section 25 and the display control unit 26 of the control device 2 shown in Fig. 3.
  • the process calculates the pointer location p as the position designated by the user in the display plane of the virtual screen (4 in Fig. 6a) and performs display control based on the calculated pointer location p.
  • the pointer location determination section (25 in Fig. 3) receives the fingertip f, determined as described with reference to Fig.4, from the fingertip detection section (23 in Fig.3).
  • the pointer location determination section receives the perceptive origin O, determined as described with reference to Figs.5a and b, from the perceptive origin detection section (22 in Fig.3).
  • the pointer location determination section determines the pointing line PL as a line going through the perceptive origin O and the fingertip position f.
  • the pointer location determination section calculates a pointer location p as the intersection between the pointing line PL and the spatial image display plane of the virtual screen.
  • the pointer location determination section outputs a control signal to the display control device (26 in Fig.
  • step S20 the display control device determines whether the distance between the fingertip f and the display plane of the virtual screen is smaller than a predetermined threshold. If in step S20 it is determined that the distance between the fingertip f and the display plane of the virtual screen is smaller than a predetermined threshold, the process proceeds to step S22. In step S22, the display control device generates a click event with regard to the activated UI element displayed at the calculated pointer location p.
  • a touch condition may be evaluated to trigger a click event (e.g. a hit box around the pointer). If the fingertip hits the box, a click is generated. When the finger leaves the box for a time T, this is considered as the user not touching anymore.
  • the virtual user interface may employ hysteresis to avoid accidental inputs.
  • Calibration of the holographic user interface In the embodiments described below in more detail a calibration process of the holographic user interface is presented. In general, this calibration process is based on an estimation of the location of the virtual screen location estimation in the 3D space. This estimation of the location of the screen comprises the process of inferring the normal vector of the plane of the screen and the process of estimating the boundaries of the virtual screen.
  • the boundaries of the virtual screen are determined by means of markers.
  • Calibration with external cameras A first example of a calibration process provides for an extrinsic calibration of the virtual screen.
  • a calibration is performed using cameras (e.g. one or more RGB cameras) and a specific pattern, here called markers.
  • the cameras are located external to the projector (11 in Fig. 2) of the holographic user interface which produces the virtual screen.
  • the calibration process as described below in more detail assumes that the extrinsics between the depth camera (5 in Figs.2 and 3) and the calibration cameras 7a, 7b are known a priori.
  • Fig. 7 illustrates a calibration process of the holographic user interface according to a first embodiment.
  • a virtual screen 4 of the holographic user interface is projected in the space as described with reference to Fig.1 and 2.
  • a plurality of (virtual) markers A, B, C and D are projected on the spatial image display plane of the virtual screen 4 at positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) respectively.
  • markers A, B, C, and D are displayed in the four corners of the virtual screen 4.
  • the markers A, B, C, and D may for example be ChArUco-markers, AprilTag-markers or checker markers.
  • Two calibration cameras 7a, b are positioned above the virtual screen 4 and are configured to capture images of the virtual screen 4 comprising the plurality of markers A, B, C and D.
  • the positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) of the markers are determined based on the images provided by calibration cameras 7a, b.
  • the extrinsics between the depth camera (5 in Figs. 2 and 3) and the calibration cameras 7a, 7b are known a priori.
  • the relative position and orientation of a calibration camera (7a or 7b) with respect to depth camera 5 are measured.
  • the transformation between the two camera coordinate systems is known and the position of a point ⁇ ⁇ ⁇ ⁇ in the coordinate system of a calibration camera can be transformed into the respective position ⁇ ⁇ ⁇ ⁇ ⁇ h in the coordinate system of the depth camera according to: ⁇ ⁇ ⁇ ⁇ ⁇ h ⁇ ⁇ ⁇ ⁇
  • ⁇ , ⁇ , and ⁇ denoting the rotation angles (Euler angles) that describe the relative orientation of the calibration camera with respect to the depth camera a priori known from the measurement
  • ⁇ ⁇ ⁇ are the three rotation matrices in homogeneous coordinates (where 3D transformations are represented by 4x4 matrices) and where ⁇ is the translation matrix in homogeneous coordinates: with ⁇ ⁇ , ⁇ ⁇ , and ⁇ ⁇ denoting the entries of the translation vector a priori
  • a thus calculated pointer location p in the coordinate system of the depth camera can then be transformed to 2D coordinates p x , p y in a 2D coordinate system of the visual screen.
  • the calculated pointer location p in the coordinate system of the depth camera can be transformed to a 2D screen coordinate system by a coordinate transformation that shifts and rotates, in the coordinate system of the depth camera, the virtual screen and the pointer location p into, e.g. the xy-plane, as known by the skilled person.
  • Fig.9 provides a schematic diagram of coordinate systems involved in the holographic user interface comprising a depth camera (5 in Fig. 2) for determining a fingertip position and a perceptive origin position, and three external RGB cameras which are used in the calibration process.
  • An origin of the device 32 is located in a world coordinate system 30.
  • the image sensor of the depth camera defines a camera coordinate system 34 of the depth camera.
  • the image sensor of the first RGB camera defines a camera coordinate system 36a of the first RGB camera.
  • the image sensor of the second RGB camera defines a camera coordinate system 36b of the second RGB camera.
  • the image sensor of the third RGB camera defines a camera coordinate system 36c of the third RGB camera.
  • a virtual screen is located in a virtual screen coordinate system 38.
  • the origin of the device 32 is a geometrical point that is meaningful for the device and that may serve as a reference for all sensors. It could be any position within the world space where the sensor and the holographic screen position resides. For instance, for a camera, it could be defined as the position of the sensor (3d sensing camera), e.g. the optical centre of one of its lenses (i.e. the fulcrum of the field of view of the lens).
  • a holographic device it could be defined as (but not restricted to) the top left corner of the main rectangular element, or a corner of one of the camera, etc.
  • calibration process relies on user feedback.
  • the user’s perception is used to estimate the perceived position of the screen.
  • a set of markers is projected on the virtual screen as targets.
  • the markers can be of different shapes and size (circles, square, boxes, spheres, etc).
  • the set of graphical clues allows to retrieve the position of the screen based on interactions and feedback from the user.
  • Fig. 10 illustrates an example of a perception-based calibration process of the holographic user interface.
  • a virtual screen 4 of the holographic user interface is projected in the space as described with reference to Fig.1 and 2.
  • a plurality of (virtual) markers A, B, C and D are projected on the spatial image display plane of the virtual screen 4 at positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) respectively.
  • markers A, B, C, and D are displayed in the four corners of the virtual screen 4.
  • the markers A, B, C, and D may for example be of different shapes and size, like circles, square, boxes, spheres, etc.
  • the calibration processing is performed by instructing the user to point at the markers A, B, C and D displayed on the virtual screen.
  • the depth camera 5 captures the finger of the user while pointing at the markers (A, B, C, D) as described above.
  • the user is instructed to point with the finger on the visual clue indicating the upper-left corner of the virtual screen.
  • the user is then instructed to point with the finger on the visual clue indicating the upper-right corner of the virtual screen.
  • the user is then instructed to point with the finger on the visual clue indicating the lower-right corner of the virtual screen.
  • the calibration process is not restricted to four points, and that the shape of the virtual screen does not necessarily have to be planar, but may instead be a 3D spline for instance, with each of the markers relating to one respective point of the 3D spline.
  • the calibration process may comprise e.g.
  • a physical button by which the user indicates to the system that the click is effective (i.e. that the current position of the fingertip should be recorded as indicating the position of the marker).
  • a timer-based trigger may be implemented. For example, if the user leaves the finger for 3 seconds on a marker, the system will automatically record the current fingertip position as indicating the position of the marker.
  • the user indicates the positions of the corners of the screen with its fingertip.
  • a more precise pointing (depth perception correction) such as a stylus may be used instead of the fingertip.
  • the process described above could be applicable to many users simultaneously. As illustrated in Fig.
  • the imaging apparatus described here could detect multiple users simultaneously, provided that they are in the Field of View (FoV), each of them interacting with the virtual screen at the same time and with different calibration parameters associated with each of them.
  • the multi-user support may for example be implemented by providing a level of face identification, and in addition, a body skeleton model that allows to associate the hand with one of the users in the FOV (otherwise raycasting is not possible) may be provided.
  • a non-planar virtual screen may be represented by a 3D spline.
  • the surface on which interaction is happening could be perceived as non “flat” but curved, a calibration using more points (i.e.
  • the perception-based calibration described with regard to Fig. 10 may be used in combination with the calibration with external cameras as described with regard to Fig.7.
  • the calibration with external cameras as described with regard to Fig.7 may be used as a pre-calibration
  • the perception-based calibration described with regard to Fig.10 can be used to refine calibration.
  • the description above is only an example configuration. Alternative configurations may be implemented with additional or other units, sensors, or the like.
  • a holographic user interface (10) comprising circuitry configured to: control a projector (1, 2, 3; 11) to project a virtual screen (4) into a space; obtain at least one image of a user of the holographic user interface (10) from at least one first camera (5); perform calibration processing by determining the position of markers (A, B, C, D) on the virtual screen (4); and determine a pointer location (p) based on the at least one image captured by the first camera (5) and based on the position of the markers (A, B, C, D) on the virtual screen (4).
  • a computer program comprising instructions, the instructions, when executed by a processor, performing the method of [19].

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

A holographic user interface (10) comprising circuitry configured to: control a projector (1, 2, 3; 11) to project a virtual screen (4) into a space; obtain at least one image of a user of the holographic user interface (10) from at least one first camera (5); perform calibration processing by determining the position of markers (A, B, C, D) on the virtual screen (4); and determine a pointer location (p) based on the at least one image captured by the first camera (5) and based on the position of the markers (A, B, C, D) on the virtual screen (4).

Description

DEVICE, METHOD AND COMPUTER PROGRAM TECHNICAL FIELD This present disclosure relates generally to an apparatus and method for providing user interfaces for 3D virtual environments, in particular to a natural holographic user interface and to a calibration method for such a natural holographic user interface. A holographic user interface is a computer input method that utilizes a projected image presented to the user for the purpose of user interaction. Holographic interfaces may for example utilize projectors to create a virtual 3D image in space. A holographic display may for example be a 3D display that utilizes light diffraction to display a 3D image to the user. The virtual image provided by a holographic user interface is constructed only of light. Due to the fact that the generated virtual image is immaterial, objects displayed to the user can become difficult for the user to interact with. The interaction with the virtual image that provides the holographic user interface is typically detected via sensors located in the projector that generates the image. Such sensors can record the location of the user’s hand and fingers, enabling control. In holographic user interfaces there may arise a discrepancy between finger position and the position of a projected image pointer. In particular, it may be difficult to accurately determine which user interface element should be activated and when. There is thus generally a need for a feasible technological solution that allows users to interact with projected, interactive 3D virtual environments. SUMMARY According to a first aspect, the disclosure provides a holographic user interface comprising: a projector configured to project a virtual screen into a space; at least one first camera configured to capture images of a user of the holographic user interface; and circuitry configured to perform calibration processing by determining the position of markers on the virtual screen; and determine a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen. According to a further aspect, the disclosure provides a method of calibrating a holographic user interface, the method comprising: projecting a virtual screen into a space; obtaining at least one image of a user of the holographic user interface from at least one first camera; performing calibration processing by determining the position of markers on the virtual screen; and determining a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen. According to a still further aspect, the disclosure provides a program comprising instructions, the instructions being configured to, when operated by a processor, perform the method mentioned above. Further aspects are set forth in the dependent claims, the following description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS Embodiments are explained by way of example with respect to the accompanying drawings, in which: Fig.1 is a perspective view of an exemplary holographic user interface according to an embodiment of the invention; Fig.2 is a system overview of a holographic user interface according to an exemplary embodiment; Fig.3 is a system overview block diagram, which shows the overview of the holographic user interface and the internal structure of a control device of the holographic user interface according to an exemplary embodiment; Fig.4 is a flow diagram illustrating the process of detection of the fingertip position by the control device of the holographic user interface of Fig.3; Fig.5a and Fig.5b illustrate a process of determining a perceptive origin of a user of the interactive user interface according to an exemplary embodiment; Fig.6a and Fig.6b illustrate a process of determining the pointer location according to an exemplary embodiment; Fig. 7 illustrates a calibration process of the holographic user interface according to a first embodiment; Fig.8 illustrates the transformation of the calculated pointer location in the coordinate system of the depth camera to 2D coordinates in a 2D coordinate system of the visual screen; Fig.9 provides a schematic diagram of coordinate systems involved in the holographic user interface; and Fig. 10 illustrates an example of a perception-based calibration process of the holographic user interface according to an exemplary embodiment. Fig.11 illustrates an example of a multi-user embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS Before a detailed description of the embodiments under reference of Fig.1 to Fig.10, general explanations are made. According to a first aspect, the embodiments disclose a holographic user interface comprising circuitry configured to control a projector to project a virtual screen into a space; obtain at least one image of a user of the holographic user interface from at least one first camera; perform calibration processing by determining the position of markers on the virtual screen, and determine a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen. The circuitry may include a processor, a memory (RAM, ROM or the like), a storage, input means (mouse, keyboard, camera, etc.), output means (display (e.g. liquid crystal, (organic) light emitting diode, etc.), loudspeakers, etc., a (wireless) interface, etc., as it is generally known for electronic devices (computers, smartphones, etc.). Moreover, circuitry may include sensors for sensing still image or video image data (image sensor, camera sensor, video sensor, etc.), and the like. The circuitry may be configured to display the markers on the virtual screen. The markers may for example be visual clues displayed on the virtual screen. The circuitry may be configured to perform the calibration processing by instructing the user to point at the markers displayed on the virtual screen. The holographic user interface may further comprise at least one second camera configured to acquire one or more images of the virtual screen. The circuitry may be configured to detect the location of the markers on the virtual screen by analyzing the one or more images of the virtual screen obtained by the at least one second camera. The markers may for example be machine readable markers. For example, the markers may be ChArUco-markers, AprilTag-markers or checker markers. The markers may be displayed at the corners of the virtual screen. The at least one first camera may be a depth camera. The circuitry may further be configured to determine a fingertip position of the user in the space based on the at least one image obtained from the at least one first camera. The circuitry may for example be configured to determine the fingertip position on the basis of a posture of a user’s hand detected by analyzing depth images obtained from a depth camera. The circuitry may further be configured to determine a perceptive origin of the user based on the at least one image obtained from the first camera. The perceptive origin may for example be determined based on a parameter describing the eye dominance of the user. The eye dominance of the user may for example be determined by any one of: a distance hole-in-the-card test, a near convergence test or a Pediatric Eye Disease Investigator Group (PEDIG) fixation preference test. The degree of eye dominance may for example be a predetermined parameter or it may be a dynamic parameter that is determined using live detection techniques. For example, eye opening percentage calculation might give a good estimation of the perceptive origin in real time. This may be done using a depth camera's intensity image, for instance, or with an additional RGB image, that could be optionally provided by the depth camera or a camera next to it. The circuitry may further be configured to determine the pointer location based on the fingertip position and based on the perceptive origin of the user. In a preferred embodiment, the pointer location is an intersection between a pointing line defined by the fingertip position and the perceptive origin and the display plane of the virtual screen. The circuitry may further be configured to activate an UI element pointed by the user on the virtual screen on a basis of the pointer location. In a preferred embodiment, the circuitry may be configured to activate the UI element pointed by the user on the virtual screen when a distance between the fingertip and the virtual screen is smaller than a predetermined threshold. According to a further aspect, the embodiments disclose a method of calibrating a holographic user interface, the method comprising: projecting a virtual screen into a space; obtaining at least one image of a user of the holographic user interface from at least one first camera; performing calibration processing by determining the position of markers on the virtual screen; determining a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen. The embodiments also disclose a computer program product which stores instructions, the instructions, when executed by a processor, performing the processes described above. Holographic user interface Fig. 1 is a perspective view of an exemplary holographic user interface 10 which acts as a computer input method that utilizes a projected virtual image presented to the user for the purpose of user interaction. The holographic user interface 10 may also be called a natural user interface (NUI), or touchless/contactless air screen, or an augmented reality pointing interface, or a digital signage display. A holographic user interface 10 is shown in which a spatial image Ib is generated on the display plane of a virtual (floating) screen 4 in a given space. A display device (not shown; see 2 in Fig.2) displays a visible light producing image Ia on an image display surface 1A in front of the user. An imaging optical system 3 forms a spatial image Ib from visible light emitted from the image display surface 1A and outputs the spatial image Ib to the display plane of the virtual screen 4. The user views the virtual image Ib and designates graphical elements with which the user may interact with a finger, a pointing rod, a pen, or the like, thus enabling input without having to contact a physical display device. The virtual image provided by the holographic user interface is constructed only of light. The interaction with the virtual image that provides the holographic user interface is typically detected via sensors (not shown in Fig.1; see Fig.2). Fig.2 is a system overview of a holographic user interface 10 according to an exemplary embodiment. According to the exemplary embodiment, the holographic user interface 10 comprises a display device 1, a control device 2, an imaging optical system 3 and a camera 5. The display device 1 may be any type of a flat display panel, such as a LCD, CRT, LED, plasma display or the like, wherein visible light is emitted in the image display surface 1A, creating an image Ia. The imaging optical system 3 receives visible light emitted from the image display surface 1A as an input and reflects the image Ia to form a spatial image Ib, which is free of reflected light from another light source, on the display plane of the virtual screen 4 in the given space. The imaging optical system 3 deviates the visible light producing image Ia from the image display surface 1A to form a spatial image Ib within the spatial image display plane of a virtual (floating) screen 4. The camera 5 is arranged to capture images of the user in which the user’s hand(s) and eyes can be seen simultaneously. In particular, the camera 5 captures a pointing gesture performed by the user’s hand(s) and fingers with respect to the display plane of the spatial image. Further, the camera 5 captures the left and right eye of the user gazing at the virtual image. The camera 5 outputs image data corresponding to the captured images to the control device 2. The control device 2 receives the image data from the camera 5 and calculates a pointer location as a position designated by the user in the display plane of the spatial image 4, as described in detail later with respect to Figs.4, 5a and b and 6. The control device 2 further controls display of the image Ia on the image display surface 1A of the display device 1 based on the calculated pointer location. The display unit 1, the control unit 2 and the imaging optical system 3 form together the projector 11. In Fig.2, the display device 1 is represented as a flat display panel. In another embodiment, the display device 1 may be curved, wherein the imaging optical system 3 is also designed to account for the curved geometries. The imaging optical system 3 may comprise components similar to the image formation means as described in WO 2014/038303 A1 and WO 2014/073650 A1. According to the technology described in these documents, the image formation means has a first and a second light control panel, which are positioned either in contact or in proximity with a plurality of first and second band-shaped optical reflection parts, whereby the pathway of light is directed to preserve the image spatially from one plane to another. Multiple embodiments of structures comprising surfaces for reflecting light are disclosed and are incorporated by reference. Other technologies for generating a virtual screen may for example include holographic techniques. The camera 5 captures images of the user in which the user’s hand(s) and eyes can be seen simultaneously. According to an exemplary embodiment, the camera 5 may be a depth camera that may capture a depth image of a scene. The depth image may include a two-dimensional (2-D) pixel array of the captured scene where each pixel in the 2-D pixel array may represent a depth value such as a distance of an object in the captured scene from the camera. Fig.3 is a system overview block diagram, which shows the overview of the holographic user interface 10 and the internal structure of the control device 2 according to an exemplary embodiment. The holographic user interface 10 of Fig.3 comprises the display device 1, the control device 2, the imaging optical system 3, and the depth camera 5 as described with reference to Figs.1 and 2. The camera 5 captures images of the user in which the user’s hand(s) and eyes can be seen simultaneously. In particular, the images captured by the camera show a pointing gesture of the user’s hand(s) and fingers with respect to the display plane of the virtual screen (4 in Figs.1 and 2). Further, the images captured by the camera 5 show a left eye and a right eye of the user gazing at the virtual image. The camera 5 outputs image data to the control device 2. Calibration cameras 7a and 7b are provided for the purpose of the calibration process of the holographic user interface described in detail with reference to Fig.7. The calibration cameras 7a and 7b may for example be RGB cameras that are configured to face the virtual screen (4 in Fig.2) and to capture images of the virtual screen. The calibration cameras 7a and 7b output image data to the control device 2. The calibration cameras 7a and 7b are, however, optional. According to another embodiment, the calibration processing is not performed based on images from calibration cameras 7a, b, but instead based on a Perception-based calibration such as described with regard to Fig. 10 below. The control device 2 comprises an image receiving section 20, a calibration processing section 21, a perceptive origin detection section 22, a fingertip detection section 23, a hand skeleton model storage section 24, a pointer location determination section 25, and a display control device 26. Via the image receiving section 20, the control device 2 receives image data from the camera 5 and image data from the calibration cameras 7a and 7b as an input. The image receiving section 20 receives the image data from the camera 5 and transmits the image data to the perceptive origin detection section 22 and the fingertip detection section 23. Further, the image receiving section 20 receives the image data from the calibration cameras 7a and 7b and transmits the image data to the calibration processing section 21. The calibration processing section 21, based on the image data received from the calibration cameras 7a, b via the image receiving section 20 (see Fig. 7 and corresponding description) or by means of perception-based calibration (see Fig. 10 and corresponding description), performs a calibration process of the holographic user interface. The perceptive origin detection section 22, based on the image data received from the camera 5, calculates a perceptive origin of the user as described in detail in Figs.5a and b below. The fingertip detection section 23 determines a fingertip position of the user in the space as described in detail in Fig. 4 below. The hand skeleton model storage section 24 stores hand skeleton model data used by the fingertip detection section 23 to determine the fingertip position. The pointer location determination section 25 receives the processing results from the perceptive origin detection section 22 and the fingertip detection section 23 and calculates a pointer location based on the received processing results (see Fig. 6). The pointer location determination section 25 outputs the processing result to the display control device 26. The display control device 26 receives the processing result from the pointer location determination section 25 and outputs a control signal to control the display of a virtual image of a user interface (Ib in Figs.1 and 2) based on the received processing result. The display control device 26 may comprise a graphic card. In a preferred embodiment, the camera 5 is a depth camera, for example a Time of Flight (ToF) camera. However, it will be understood that alternative imaging units, for example a stereoscopic camera, may be used. The camera 5 may contain a depth image sensor that captures depth data and a depth processor that processes the depth data to generate a depth map. The processing steps performed by the depth processor are dependent upon the technique used by the depth image sensor. The depth image sensor may rely on one of several different sensor technologies. Among these sensor technologies are time-of-flight (TOF) technologies, such as indirect time-of-flight (iTOF) and direct time-of-flight (dTOF), structured light, radar, stereoscopic cameras, active stereoscopic sensors, and the like. Most of these techniques rely on active sensors that have their own illumination source that uses modulated or pulsed light. In contrast, passive sensor techniques such as stereoscopic cameras do not have their own illumination source, but instead measure ambient light and are therefore dependent on ambient lighting. In addition to depth data, the camera 5, like conventional color cameras, can also produce color (RGB) data that can be combined with the depth data for processing. The camera 5 may also include other components, such as one or more optical lenses, an illumination source, and electronic controllers. Fingertip detection Fig. 4 is a flow diagram illustrating the process of detection of the fingertip position by the control device 2 of the holographic user interface of Fig.3. The process may for example be implemented by the fingertip detection section (23 in Fig.3). In a step S1, the fingertip detection section receives the image data (depth images) from the camera (5 in Fig.2). In a step S2, the fingertip detection section processes the depth image data acquired from the camera to determine the position of the user’s hands and fingers relative to the virtual display. In particular, algorithms may be applied to determine a hand skeleton model (24 in Fig. 3) based on the image data. For example, the fingertip detection section may apply an appropriate algorithm to depth image data to enable the system to utilize close-range depth data. The fingertip detection section may be enabled to process depth image data in accordance with close-range optical settings and requirements to determine a presence, position, movement, etc. of an object. In particular, the fingertip detection section may execute software code or algorithms for close range tracking to enable detection and/or tracking of movements of a user’s hands and finger, and the output of the fingertip detection section can be a representation of the skeleton of the user’s hand. The hand skeleton representation can include the positions of the joints of the skeleton, and may also include the rotations of the joints, relative to a center point. The skeleton model (24 in Fig. 3) may be used to further improve the quality of the tracking and assign positions to joints which were not detected in the earlier stages, e.g. due to occlusions or because parts of the hand are outside the field of view of the camera. In a step S3, the fingertip detection section (23 in Fig.3), on the basis of the hand skeleton model acquired in S2, determines the fingertip position f by using hand skeleton model data stored in the hand skeleton model storage section (24 in Fig.3). Determination of a skeleton model of the user’s hand as performed in step S2 may for example be implemented according to any state-of-the-art techniques such as machine learning, or the like. For example, U.S. patent application No. 13/768,835 describes a tracking module that processes depth data of a user performing movements, for example, movements of the user’s hand and fingers. Alternatively, US 2012/0327125 A1 describes a method for tracking a user’s hands and fingers based on depth images captured from a depth camera, and using the tracked data to control a user’s interaction with devices. In the embodiment of Fig. 4 a fingertip position f is determined based on camera images and based on a hand skeleton model. In other embodiments, the camera 5 may detect more than one object simultaneously. The detection of multiple objects, for example two fingers of the user’s hand, may be useful, for example, in performing a “pinch-in” or “pinch-out” command. Perceptive origin detection Fig.5a and 5b illustrate a process of determining a perceptive origin of a user of the interactive user interface according to an exemplary embodiment. This process may for example be implemented by the perceptive origin detection section 22 of Fig.3. The process calculates the perceptive origin (eye origin) of the user based on the image data received from the camera 5. As shown in Fig. 5a, the perceptive origin O is located on the line connecting the left eye L and the right eye R (centroid of the pupils) of the viewer and is indicative of the user’s preference to use one eye more than the fellow eye to accomplish a task. According to the present invention, the user views items displayed on the virtual screen. As illustrated in Fig.5a, the camera 5 is mounted opposite the eye opening(s) so as to capture images of the user’s left eye (L) and the user’s right eye (R). It is conceivable that other embodiments could use multiple cameras. The camera 5 captures images of the left and right eye of the user gazing at the virtual screen and transmits image data to the perceptive origin detection section (22 in Fig. 3). The perceptive origin detection section processes the depth image data acquired from the camera 5 to determine a degree of eye dominance ^ ^ ^and, based on the processing result, determines the perceptive origin O. As shown in Fig. 5b, the position of the perspective origin ^ ^^ ^^ may, for example, be represented by the following formula: ^ ^^ ^^ = ^ ^^ ^^ + ^^( ^^)(^ ^^ ^^ −^ ^^ ^^ ) where ^ ^^ ^^ is the position of the left eye, ^ ^^ ^^ is the position of the right eye, and ^^( ^^) ∈[0,1] is a parameter describing the degree of eye dominance. In general, the degree of eye dominance ^^( ^^) can change over time and be recalculated in real time. However, the ^^( ^^) must not necessarily have a time dependence. It can also be set as a predefined value ^^ constant over time. According to this embodiment, a value of ^ = 0,5 indicates a complete balance between the eyes (50/50, no dominance) and a perceptive origin O that is located in the middle between the left eye and the right eye. A value of ^ greater than 0,5 indicates an increasingly stronger dominance of the right eye and a perceptive origin that is correspondingly increasingly closer to the right eye, up to a value of 1 where only the right eye is active and the perceptive origin is located at the position of the right eye. Conversely, a value of ^ of less than 0,5 indicates a dominance of the left eye and a perceptive origin that is correspondingly closer to the left eye, a value of 0 indicating that only the left eye is active with the perceptive origin being located at the position of the left eye. The parameter ^^ which describes the degree of eye dominance may for example be known in advance and input to the holographic user interface as a predefined parameter. That is, the dominant eye may be known a priori. For example, the user may enter this information manually. Alternatively, the parameter ^^ which describes the degree of eye dominance may be determined in a calibration stage. The determination of the dominant eye is not limited to any particular technique. Various tests are known to determine eye dominance, such as the distance hole-in-the-card test, the near convergence test, or the Pediatric Eye Disease Investigator Group (PEDIG) fixation preference test (see, e.g., Rice M. et al., “Results of ocular dominance testing depend on assessment method”, Journal of American Association for Pediatric Ophthalmology and Strabismus. 2008;12(4):369–465). Whether an eye is dominant over the other eye may for example be determined based on how the respective eyes move and/or based on respective deviations between the points of regard for the respective eyes and a reference point when the user is prompted to look at the reference point. Eye dominance can for example lead to externally distinguishable features of the person’s eye movements. For example, one eye (the non-dominant eye) may exhibit a lower range of movement than the other, or lag the other, or stay still, or look in a different apparent direction to the direction of actual gaze of the user. According to an alternative embodiment, the perceptive origin detection section may determine the perceptive origin based on the eye openings. Still alternatively, the dominant eye can be defined, for example, as the eye with the wider range of movement, the eye which reaches a gaze direction appropriate to a displayed stimulus the closest and/or the most quickly, or the eye which holds a position relating to a displayed stimulus. As stated above, the degree of eye dominance ^^( ^^) can change over time and be recalculated in real time. For example, eye opening percentage calculation might give a good estimation of the perceptive origin in real time. This may be done using a depth camera's intensity image, for instance, or with an additional RGB image, that could be optionally provided by the depth camera or a camera next to it. Determination of pointer location using raycasting Fig. 6a and 6b illustrate a process of determining the pointer location according to an exemplary embodiment. Fig. 6a schematically illustrates the positional relationship between the perceptive origin O, the fingertip position f and the pointer location p. The perceptive origin O, which is determined by the perceptive origin detection section (22 in Fig.3) as described above with reference to Figs.5a and 5b, is located on the line connecting the left eye L and the right eye R of the viewer. The fingertip position f, which is determined by the fingertip detection section (23 in Fig. 3) as described above with reference to Fig. 4, is located in the vicinity of the virtual screen 4. The perceptive origin O and the fingertip position f define a pointing line PL, as shown in Fig.6a. The pointer location p is determined as the intersection of this pointing line PL and the spatial image display plane of the virtual screen 4. Fig.6b is a flow chart illustrating an operation process of determining the pointer location and display control according to an exemplary embodiment. This process may, for example, be implemented by the pointer location determination section 25 and the display control unit 26 of the control device 2 shown in Fig. 3. The process calculates the pointer location p as the position designated by the user in the display plane of the virtual screen (4 in Fig. 6a) and performs display control based on the calculated pointer location p. In a step S10, the pointer location determination section (25 in Fig. 3) receives the fingertip f, determined as described with reference to Fig.4, from the fingertip detection section (23 in Fig.3). In a step S12, the pointer location determination section receives the perceptive origin O, determined as described with reference to Figs.5a and b, from the perceptive origin detection section (22 in Fig.3). In a step S14, the pointer location determination section determines the pointing line PL as a line going through the perceptive origin O and the fingertip position f. In a step S16, the pointer location determination section calculates a pointer location p as the intersection between the pointing line PL and the spatial image display plane of the virtual screen. In a step S18, the pointer location determination section outputs a control signal to the display control device (26 in Fig. 3) based on the calculated pointer location p to activate an UI element displayed at the calculated pointer location p. Activation can take the form of highlighting the UI element, for example, to provide feedback to the user about which UI element is designated. In a step S20, the display control device determines whether the distance between the fingertip f and the display plane of the virtual screen is smaller than a predetermined threshold. If in step S20 it is determined that the distance between the fingertip f and the display plane of the virtual screen is smaller than a predetermined threshold, the process proceeds to step S22. In step S22, the display control device generates a click event with regard to the activated UI element displayed at the calculated pointer location p. Optional, a touch condition may be evaluated to trigger a click event (e.g. a hit box around the pointer). If the fingertip hits the box, a click is generated. When the finger leaves the box for a time T, this is considered as the user not touching anymore. In some embodiments, the virtual user interface may employ hysteresis to avoid accidental inputs. Calibration of the holographic user interface In the embodiments described below in more detail a calibration process of the holographic user interface is presented. In general, this calibration process is based on an estimation of the location of the virtual screen location estimation in the 3D space. This estimation of the location of the screen comprises the process of inferring the normal vector of the plane of the screen and the process of estimating the boundaries of the virtual screen. According to the embodiments, the boundaries of the virtual screen are determined by means of markers. Calibration with external cameras A first example of a calibration process provides for an extrinsic calibration of the virtual screen. According to this embodiment, a calibration is performed using cameras (e.g. one or more RGB cameras) and a specific pattern, here called markers. The cameras are located external to the projector (11 in Fig. 2) of the holographic user interface which produces the virtual screen. The calibration process as described below in more detail assumes that the extrinsics between the depth camera (5 in Figs.2 and 3) and the calibration cameras 7a, 7b are known a priori. Fig. 7 illustrates a calibration process of the holographic user interface according to a first embodiment. A virtual screen 4 of the holographic user interface is projected in the space as described with reference to Fig.1 and 2. In a calibration stage, a plurality of (virtual) markers A, B, C and D are projected on the spatial image display plane of the virtual screen 4 at positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) respectively. In the embodiment of Fig.7, markers A, B, C, and D are displayed in the four corners of the virtual screen 4. The markers A, B, C, and D may for example be ChArUco-markers, AprilTag-markers or checker markers. Two calibration cameras 7a, b (here, RGB cameras) (see also 7a, b in Fig. 2) are positioned above the virtual screen 4 and are configured to capture images of the virtual screen 4 comprising the plurality of markers A, B, C and D. According to this calibration process, the positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) of the markers are determined based on the images provided by calibration cameras 7a, b. As stated above, the extrinsics between the depth camera (5 in Figs. 2 and 3) and the calibration cameras 7a, 7b are known a priori. For example, in the calibration stage, the relative position and orientation of a calibration camera (7a or 7b) with respect to depth camera 5 are measured. As a result of this measurement, the transformation between the two camera coordinate systems is known and the position of a point ^^ ^^ ^^ ^^ in the coordinate system of a calibration camera can be transformed into the respective position ^^ ^^ ^^ ^^ ^^ℎ in the coordinate system of the depth camera according to: ^^ ^^ ^^ ^^ℎ ⊺ ^^ ^^ ⊺
Figure imgf000014_0001
where ^^, ^^, and ^^ denoting the rotation angles (Euler angles) that describe the relative orientation of the calibration camera with respect to the depth camera a priori known from the measurement, and ^^ ^^ ^^ are the three rotation matrices in homogeneous coordinates (where
Figure imgf000015_0001
3D transformations are represented by 4x4 matrices)
Figure imgf000015_0002
and where ^^ is the translation matrix in homogeneous coordinates:
Figure imgf000015_0003
with ^^ ^^ , ^^ ^^, and ^^ ^^ denoting the entries of the translation vector a priori known from the measurement. With the translation between the coordinate system of the calibration camera and the coordinate system of the depth camera a priori known, the coordinates ^^ ^^ ^^ ^^ = ( ^^ ^^, ^^ ^^, ^^ ^^), ^^ ^^ ^^ ^^ = ( ^^ ^^, ^^ ^^, ^^ ^^), ^^ ^^ ^^ ^^=( ^^ ^^, ^^ ^^, ^^ ^^), and ^^ ^^ ^^ ^^ = ( ^^ ^^, ^^ ^^, ^^ ^^) of the visual clues marking the corners of the virtual screen can be transformed from the coordinate system of a calibration camera to the coordinate system of the depth camera: ( ^^ ^^ ^^ ^^ ^^ℎ , 1)⊺ = ^^ ^^ ^^( ^^) ^^ ^^ ^^ ^^ ⊺ ^^( ^^) ^^ ^^( ^^)( ^^ , 1) ^^ ^^ ^^ ^^ℎ ⊺ ^^ ^^ ^^ ⊺
Figure imgf000015_0004
As the corner coordinates ^^ ^^ ^^ ^^ ^^ℎ , ^^ ^^ ^^ ^^ ^^ℎ ^^ ^^ ^^ ^^ ^^ℎ ^^ ^^ ^^ ^^ ^^ℎ i the position, orientation and dimensions of the virtual screen in the coordinate system of the depth camera, these coordinates ^^ ^^ ^^ ^^ ^^ℎ , ^^ ^^ ^^ ^^ ^^ℎ , ^^ ^^ ^^ ^^ ^^ℎ , and ^^ ^^ ^^ ^^ ^^ℎ can be used to determine a calibrated screen plane that can be used in the process of Fig.6 to determine a pointer location p via raycasting. As described in Fig. 8, a thus calculated pointer location p in the coordinate system of the depth camera can then be transformed to 2D coordinates px, py in a 2D coordinate system of the visual screen. Here, the x-axis and the y-axis of the 2D screen coordinate system may for example be defined by the vectors ^^ ^^ = ^^ ^^ ^^ ^^ ^^ℎ − ^^ ^^ ^^ ^^ ^^ℎ , and ^^ ^^ = ^^ ^^ ^^ ^^ ^^ℎ − ^^ ^^ ^^ ^^ ^^ℎ that describe the borders of the screen, as displayed in Fig.8. Alternatively, the calculated pointer location p in the coordinate system of the depth camera can be transformed to a 2D screen coordinate system by a coordinate transformation that shifts and rotates, in the coordinate system of the depth camera, the virtual screen and the pointer location p into, e.g. the xy-plane, as known by the skilled person. Fig.9 provides a schematic diagram of coordinate systems involved in the holographic user interface comprising a depth camera (5 in Fig. 2) for determining a fingertip position and a perceptive origin position, and three external RGB cameras which are used in the calibration process. An origin of the device 32 is located in a world coordinate system 30. The image sensor of the depth camera defines a camera coordinate system 34 of the depth camera. The image sensor of the first RGB camera defines a camera coordinate system 36a of the first RGB camera. The image sensor of the second RGB camera defines a camera coordinate system 36b of the second RGB camera. The image sensor of the third RGB camera defines a camera coordinate system 36c of the third RGB camera. A virtual screen is located in a virtual screen coordinate system 38. The origin of the device 32 is a geometrical point that is meaningful for the device and that may serve as a reference for all sensors. It could be any position within the world space where the sensor and the holographic screen position resides. For instance, for a camera, it could be defined as the position of the sensor (3d sensing camera), e.g. the optical centre of one of its lenses (i.e. the fulcrum of the field of view of the lens). For instance, for a holographic device, it could be defined as (but not restricted to) the top left corner of the main rectangular element, or a corner of one of the camera, etc. Perception-based calibration According to another embodiment, calibration process relies on user feedback. In this embodiment, the user’s perception is used to estimate the perceived position of the screen. A set of markers (graphical clues) is projected on the virtual screen as targets. The markers can be of different shapes and size (circles, square, boxes, spheres, etc). The set of graphical clues allows to retrieve the position of the screen based on interactions and feedback from the user. Fig. 10 illustrates an example of a perception-based calibration process of the holographic user interface. A virtual screen 4 of the holographic user interface is projected in the space as described with reference to Fig.1 and 2. In a calibration stage, a plurality of (virtual) markers A, B, C and D are projected on the spatial image display plane of the virtual screen 4 at positions A(ax, ay, az), B(bx, by, bz), C(cx, cy, cz) and D(dx, dy, dz) respectively. In the embodiment of Fig.10, markers A, B, C, and D are displayed in the four corners of the virtual screen 4. The markers A, B, C, and D may for example be of different shapes and size, like circles, square, boxes, spheres, etc. The calibration processing is performed by instructing the user to point at the markers A, B, C and D displayed on the virtual screen. The depth camera 5 captures the finger of the user while pointing at the markers (A, B, C, D) as described above. For example, the user is instructed to point with the finger on the visual clue indicating the upper-left corner of the virtual screen. The system will then determine this fingertip position for example as set out in Fig. 4 described above and will recognize the thus obtained fingertip position as point ^^ ^^ ^^ ^^ ^^ℎ = ( ^^ ^^, ^^ ^^, ^^ ^^) indicating the upper-left corner of the virtual screen in the coordinate system of the depth camera. The user is then instructed to point with the finger on the visual clue indicating the upper-right corner of the virtual screen. The system will then determine this fingertip position for example as set out in Fig. 4 described above and will recognize the thus obtained fingertip position as point ^^ ^^ ^^ ^^ ^^ℎ = ( ^^ ^^, ^^ ^^, ^^ ^^) indicating the upper- right corner of the virtual screen in the coordinate system of the depth camera. The user is then instructed to point with the finger on the visual clue indicating the lower-right corner of the virtual screen. The system will then determine this fingertip position for example as set out in Fig.4 described above and will recognize the thus obtained fingertip position as point ^^ ^^ ^^ ^^ ^^ℎ=( ^^ ^^, ^^ ^^, ^^ ^^), indicating the lower-right corner of the virtual screen in the coordinate system of the depth camera. The user is then instructed to point with the finger on the visual clue indicating the lower-left corner of the virtual screen. The system will then determine this fingertip position for example as set out in Fig.4 described above and will recognize the thus obtained fingertip position as point ^^ ^^ ^^ ^^ ^^ℎ = ( ^^ ^^, ^^ ^^, ^^ ^^) indicating the lower-left corner of the virtual screen in the coordinate system of the depth camera. It should be noted that the calibration process is not restricted to four points, and that the shape of the virtual screen does not necessarily have to be planar, but may instead be a 3D spline for instance, with each of the markers relating to one respective point of the 3D spline. The calibration process may comprise e.g. a physical button by which the user indicates to the system that the click is effective (i.e. that the current position of the fingertip should be recorded as indicating the position of the marker). Alternatively, a timer-based trigger may be implemented. For example, if the user leaves the finger for 3 seconds on a marker, the system will automatically record the current fingertip position as indicating the position of the marker. In the embodiment above, the user indicates the positions of the corners of the screen with its fingertip. Optionally or alternatively, a more precise pointing (depth perception correction) such as a stylus may be used instead of the fingertip. The process described above could be applicable to many users simultaneously. As illustrated in Fig. 11, the imaging apparatus described here could detect multiple users simultaneously, provided that they are in the Field of View (FoV), each of them interacting with the virtual screen at the same time and with different calibration parameters associated with each of them. The multi-user support may for example be implemented by providing a level of face identification, and in addition, a body skeleton model that allows to associate the hand with one of the users in the FOV (otherwise raycasting is not possible) may be provided. In this way, one may have custom calibration for each user that could be synchronized with a profile, whether that profile is manually loaded or automatically based on face recognition (that could be done using the same depth sensing sensor as it is already necessary for that sensor to see the eyes of the user, it will see the face).It should be noted that the perception surface might be non-planar depending on the user’s perception. The calibration process may take this into account. For example a non-planar virtual screen may be represented by a 3D spline. Depending on the user, the surface on which interaction is happening could be perceived as non “flat” but curved, a calibration using more points (i.e. 9 or 16 or 25 in a grid) would allow to model that curved surface and that model can then be used for the computation of the distance of finger to “interaction surface/holographic screen”, used for the clicking interaction for instance. It should also be noted that in alternative embodiments the perception-based calibration described with regard to Fig. 10 may be used in combination with the calibration with external cameras as described with regard to Fig.7. For example, the calibration with external cameras as described with regard to Fig.7 may be used as a pre-calibration, and the perception-based calibration described with regard to Fig.10 can be used to refine calibration. *** It should be noted that the description above is only an example configuration. Alternative configurations may be implemented with additional or other units, sensors, or the like. It should also be noted that the division of the systems into units is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. It should also be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is, however, given for illustrative purposes only and should not be construed as binding. All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example, on a chip, in FPGA, or the like, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software. In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure. Note that the present technology can also be configured as described below: [1] A holographic user interface (10) comprising circuitry configured to: control a projector (1, 2, 3; 11) to project a virtual screen (4) into a space; obtain at least one image of a user of the holographic user interface (10) from at least one first camera (5); perform calibration processing by determining the position of markers (A, B, C, D) on the virtual screen (4); and determine a pointer location (p) based on the at least one image captured by the first camera (5) and based on the position of the markers (A, B, C, D) on the virtual screen (4). [2] The holographic user interface (10) of [1], wherein the circuitry (2) is configured to display the markers (A, B, C, D) on the virtual screen (4). [3] The holographic user interface (10) of [1] or [2], wherein the markers (A, B, C, D) are visual clues displayed on the virtual screen (4). [4] The holographic user interface (10) of any one of [1] to [3], wherein the circuitry (2) is configured to perform the calibration processing by instructing the user to point at the markers (A, B, C, D) displayed on the virtual screen (4). [5] The holographic user interface (10) of any one of [1] to [4], further comprising at least one second camera (7a, 7b) configured to acquire one or more images of the virtual screen (4). [6] The holographic user interface (10) of [5], wherein the circuitry is configured to detect the location of the markers (A, B, C, D) on the virtual screen (4) by analyzing the one or more images of the virtual screen (4) obtained by the at least one second camera (7a, 7b). [7] The holographic user interface (10) of any one of [1] to [6], wherein the markers (A, B, C, D) are machine readable markers. [8] The holographic user interface (10) of any one of [1] to [7], wherein the markers (A, B, C, D) are ChArUco-markers, AprilTag-markers or checker markers. [9] The holographic user interface (10) of any one of [1] to [8], wherein the markers are displayed at the corners (A, B, C, D) of the virtual screen (4). [10] The holographic user interface (10) of any one of [1] to [9], wherein the at least one first camera (5) is a depth camera. [11] The holographic user interface (10) of any one of [1] to [10], wherein the circuitry is configured to determine a fingertip position (f) of the user in the space based on the at least one image obtained from the at least one first camera (5). [12] The holographic user interface (10) of any one of [1] to [11], wherein the circuitry is configured to determine a perceptive origin (O) of the user based on the at least one image obtained from the first camera (5). [13] The holographic user interface (10) of [12], wherein the perceptive origin (O) is further determined based on a parameter describing the eye dominance of the user. [14] The holographic user interface (10) of [12] or [13], wherein the circuitry is configured to determine the pointer location (p) based on the fingertip position (f) and based on the perceptive origin (O) of the user. [15] The holographic user interface (10) of any one of [12] to [14], wherein the pointer location (p) is an intersection between a pointing line (PL) defined by the fingertip position (f) and the perceptive origin (O) and the display plane of the virtual screen (4). [16] The holographic user interface (10) of any one of [1] to [15], wherein the circuitry (2) is further configured to activate an UI element pointed by the user on the virtual screen (4) on a basis of the pointer location (p). [17] The holographic user interface (10) of [16], wherein the circuitry is configured to activate the UI element pointed by the user on the virtual screen (4) when a distance between the fingertip and the virtual screen (4) is smaller than a predetermined threshold. [18] The holographic user interface (10) of any one of [1] to [17], further comprising the projector (11) and the at least one first camera (5). [19] A method of calibrating a holographic user interface (10), the method comprising: projecting a virtual screen (4) into a space; obtaining at least one image of a user of the holographic user interface (10) from at least one first camera (5); performing calibration processing by determining the position of markers (A, B, C, D) on the virtual screen (4); and determining a pointer location (p) based on the images captured by the first camera (5) and based on the position of the markers (A, B, C, D) on the virtual screen (4). [20] A computer program comprising instructions, the instructions, when executed by a processor, performing the method of [19]. Reference signs 1 display device 1A image display surface 2 display control device 3 imaging optical system 4 virtual screen 5 depth camera 7a, b calibration cameras (e.g. RGB cameras) 10 holographic user interface 11 projector 20 image receiving section 21 calibration processing section 22 perceptive origin detection section 23 fingertip detection section 24 hand skeleton model storage section 25 pointer location determination section 26 display control device Ia visible light producing image Ib spatial image

Claims

Claims 1. A holographic user interface comprising circuitry configured to: control a projector to project a virtual screen into a space; obtain at least one image of a user of the holographic user interface from at least one first camera; perform calibration processing by determining the position of markers on the virtual screen; and determine a pointer location based on the at least one image captured by the first camera and based on the position of the markers on the virtual screen.
2. The holographic user interface of claim 1, wherein the circuitry is configured to display the markers on the virtual screen.
3. The holographic user interface of claim 1, wherein the markers are visual clues displayed on the virtual screen.
4. The holographic user interface of claim 1, wherein the circuitry is configured to perform the calibration processing by instructing the user to point at the markers displayed on the virtual screen.
5. The holographic user interface of claim 1, further comprising at least one second camera configured to acquire one or more images of the virtual screen.
6. The holographic user interface of claim 5, wherein the circuitry is configured to detect the location of the markers on the virtual screen by analyzing the one or more images of the virtual screen obtained by the at least one second camera.
7. The holographic user interface of claim 6, wherein the markers are machine readable markers.
8. The holographic user interface of claim 6 or 7, wherein the markers are ChArUco-markers, AprilTag-markers or checker markers.
9. The holographic user interface of claim 1, wherein the markers are displayed at the corners of the virtual screen.
10. The holographic user interface of claim 1, wherein the at least one first camera is a depth camera.
11. The holographic user interface of claim 1, wherein the circuitry is configured to determine a fingertip position of the user in the space based on the at least one image obtained from the at least one first camera.
12. The holographic user interface of claim 11, wherein the circuitry is configured to determine a perceptive origin of the user based on the at least one image obtained from the first camera.
13. The holographic user interface of claim 12, wherein the perceptive origin is further determined based on a parameter describing the eye dominance of the user.
14. The holographic user interface of claim 12 or 13, wherein the circuitry is configured to determine the pointer location based on the fingertip position and based on the perceptive origin of the user.
15. The holographic user interface of claim 14, wherein the pointer location is an intersection between a pointing line defined by the fingertip position and the perceptive origin and the display plane of the virtual screen.
16. The holographic user interface of claim 1, wherein the circuitry is further configured to activate an UI element pointed by the user on the virtual screen on a basis of the pointer location.
17. The holographic user interface of claim 16, wherein the circuitry is configured to activate the UI element pointed by the user on the virtual screen when a distance between the fingertip and the virtual screen is smaller than a predetermined threshold.
18. The holographic user interface of claim 1, further comprising the projector and the at least one first camera.
19. A method of calibrating a holographic user interface, the method comprising: projecting a virtual screen into a space; obtaining at least one image of a user of the holographic user interface from at least one first camera; performing calibration processing by determining the position of markers on the virtual screen; and determining a pointer location based on the images captured by the first camera and based on the position of the markers on the virtual screen.
20. A computer program comprising instructions, the instructions, when executed by a processor, performing the method of claim 19.
PCT/EP2024/055391 2023-03-06 2024-03-01 Device, method and computer program Ceased WO2024184232A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23160260 2023-03-06
EP23160260.8 2023-03-06

Publications (1)

Publication Number Publication Date
WO2024184232A1 true WO2024184232A1 (en) 2024-09-12

Family

ID=85505542

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2024/055391 Ceased WO2024184232A1 (en) 2023-03-06 2024-03-01 Device, method and computer program

Country Status (1)

Country Link
WO (1) WO2024184232A1 (en)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110251905A1 (en) * 2008-12-24 2011-10-13 Lawrence Nicholas A Touch Sensitive Holographic Displays
US20120327125A1 (en) 2011-06-23 2012-12-27 Omek Interactive, Ltd. System and method for close-range movement tracking
WO2014038303A1 (en) 2012-09-10 2014-03-13 株式会社アスカネット Floating touch panel
WO2014073650A1 (en) 2012-11-08 2014-05-15 株式会社アスカネット Light control panel fabrication method
US20190004611A1 (en) * 2013-06-27 2019-01-03 Eyesight Mobile Technologies Ltd. Systems and methods of direct pointing detection for interaction with a digital device
US20200186786A1 (en) * 2018-12-06 2020-06-11 Novarad Corporation Calibration for Augmented Reality

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110251905A1 (en) * 2008-12-24 2011-10-13 Lawrence Nicholas A Touch Sensitive Holographic Displays
US20120327125A1 (en) 2011-06-23 2012-12-27 Omek Interactive, Ltd. System and method for close-range movement tracking
WO2014038303A1 (en) 2012-09-10 2014-03-13 株式会社アスカネット Floating touch panel
WO2014073650A1 (en) 2012-11-08 2014-05-15 株式会社アスカネット Light control panel fabrication method
US20190004611A1 (en) * 2013-06-27 2019-01-03 Eyesight Mobile Technologies Ltd. Systems and methods of direct pointing detection for interaction with a digital device
US20200186786A1 (en) * 2018-12-06 2020-06-11 Novarad Corporation Calibration for Augmented Reality

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
CAO ANDONG ET AL: "Image-based marker tracking and registration for intraoperative 3D image-guided interventions using augmented reality", PROGRESS IN BIOMEDICAL OPTICS AND IMAGING, SPIE - INTERNATIONAL SOCIETY FOR OPTICAL ENGINEERING, BELLINGHAM, WA, US, vol. 11318, 2 March 2020 (2020-03-02), pages 1131802 - 1131802, XP060129187, ISSN: 1605-7422, ISBN: 978-1-5106-0027-0, DOI: 10.1117/12.2550415 *
RICE M ET AL.: "Results of ocular dominance testing depend on assessment method", JOURNAL OF AMERICAN ASSOCIATION FOR PEDIATRIC OPHTHALMOLOGY AND STRABISMUS, vol. 12, no. 4, 2008, pages 369 - 465, XP024520001, DOI: 10.1016/j.jaapos.2008.01.017

Similar Documents

Publication Publication Date Title
EP3227760B1 (en) Pointer projection for natural user input
US20260095563A1 (en) Methods and systems for multiple access to a single hardware data stream
US9367951B1 (en) Creating realistic three-dimensional effects
US8933886B2 (en) Instruction input device, instruction input method, program, recording medium, and integrated circuit
CN112578911B (en) Apparatus and method for tracking head and eye movements
CN107004279B (en) Natural User Interface Camera Calibration
US20190266798A1 (en) Apparatus and method for performing real object detection and control using a virtual reality head mounted display system
Canessa et al. Calibrated depth and color cameras for accurate 3D interaction in a stereoscopic augmented reality environment
KR20220120649A (en) Artificial Reality System with Varifocal Display of Artificial Reality Content
US20100315414A1 (en) Display of 3-dimensional objects
US20140218281A1 (en) Systems and methods for eye gaze determination
US20150350618A1 (en) Method of and system for projecting digital information on a real object in a real environment
US20140354602A1 (en) Interactive input system and method
TW202009786A (en) Electronic apparatus operated by head movement and operation method thereof
CN110377148B (en) Computer readable medium, method of training object detection algorithm and training device
WO2014093608A1 (en) Direct interaction system for mixed reality environments
TW202025719A (en) Method, apparatus and electronic device for image processing and storage medium thereof
US20210406542A1 (en) Augmented reality eyewear with mood sharing
CN117372475A (en) Eye tracking methods and electronic devices
CN110858095A (en) Electronic device that can be controlled by head and its operation method
WO2022019976A1 (en) Systems and methods for updating continuous image alignment of separate cameras
CN110018733B (en) Method, device and memory device for determining user trigger intention
WO2024184232A1 (en) Device, method and computer program
US20200167005A1 (en) Recognition device and recognition method
US20260037100A1 (en) Virtual image sharing method and virtual image sharing system

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24708801

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 24708801

Country of ref document: EP

Kind code of ref document: A1