EP4720815A1 - System, method, and computer program for controlling a user interface - Google Patents

System, method, and computer program for controlling a user interface

Info

Publication number
EP4720815A1
EP4720815A1 EP24728210.6A EP24728210A EP4720815A1 EP 4720815 A1 EP4720815 A1 EP 4720815A1 EP 24728210 A EP24728210 A EP 24728210A EP 4720815 A1 EP4720815 A1 EP 4720815A1
Authority
EP
European Patent Office
Prior art keywords
user interface
user
point
pointing
indicator
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24728210.6A
Other languages
German (de)
French (fr)
Inventor
George Themelis
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Leica Microsystems CMS GmbH
Leica Instruments Singapore Pte Ltd
Original Assignee
Leica Microsystems CMS GmbH
Leica Instruments Singapore Pte Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from EP23175149.6A external-priority patent/EP4468121A1/en
Priority claimed from EP23175130.6A external-priority patent/EP4468118A1/en
Priority claimed from EP23175126.4A external-priority patent/EP4468117A1/en
Application filed by Leica Microsystems CMS GmbH, Leica Instruments Singapore Pte Ltd filed Critical Leica Microsystems CMS GmbH
Publication of EP4720815A1 publication Critical patent/EP4720815A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

Examples relate to a system, to a method and to a computer program for controlling a user interface. The system comprises one or more processors. It is configured to determine a posi- tion of a first point of reference based on a position of a first point on the body of a user, determine a position of a second point of reference based on a position of a pointing indicator used by the user, adjust the user interface based on the determined position in the user inter- face, generate a display signal comprising the adjusted user interface; and provide the display signal to a display device.

Description

System, Method, and Computer Program for Controlling a User Interface
Technical field
Examples relate to a system, to a method and to a computer program for controlling a user interface.
Background
Gesture-based input refers to a way of interacting with a computer, mobile device, or any other electronic device through the use of gestures or movements made by the user's body. These gestures are captured by built-in sensors, such as cameras or depth sensors, and are translated into commands that the device understands. Examples of gesture-based input include swiping, pinching, tapping, shaking, or tilting a device to navigate through applications, adjust settings, or perform specific actions. Gesture-based input is becoming more and more common because it is a natural and intuitive way of interacting with technology, and it allows for hands-free or touch-free control, making it useful for people with physical disabilities.
In some cases, touch-free gestures are used to control the position of a pointer (e.g., cursor) in a user interface. For this purpose, the position of the hand of the user is captured, using a camera, in 2D or 3D, to control the cursor. For example, Tran et al: “Real-time virtual mouse system using RGB-D images and fingertip detection” and Argyros and Lourakis: “Vision- Based Interpretation of Hand Gestures for Remote Control of a Computer Mouse” show examples of this methodology. However, known implementations of computer control via cameras that capture the user’s hand often lack intuitiveness and thus are very rarely used.
There may be a desire for an improved concept for adjusting a user interface.
Summary
This desire is addressed by the subject-matter of the independent claims. Various examples of the proposed concept are based on the finding, that touch-free gesturebased control of user interfaces is often deemed unintuitive, as a pointer being controlled via gesture control is controlled regardless of the point of view of the user. Instead, the pointer is often controlled in a linear fashion, which is not intuitive, as, from the point of view of the user, the further away the user interface is, the further the pointer travels. This becomes apparent once the user points at a portion of the user interface that is laterally offset to the user’s position in front of it. To improve the intuitiveness, and thus also confidence and speed with which the user interface can be controlled, in the proposed concept, the user’s point of view is determined by determining the position of two points of reference, one on the body (e.g., the face, such as an eye or between the eyes) and another one at a pointing indicator (such as a finger, a scalpel, or a pointing stick). From these two points of reference, a position on the user interface, an element of the user interface (such as a button or other user interface element), or a device can be determined, and a response can be determined accordingly (which may control the user interface). For example, the user interface can be adjusted based on the determined position in the user interface, e.g., a position the user is looking at. In this way, a more intuitive control of a user interface, e.g., for controlling a surgical microscope or a device within a surgical theatre, may be provided.
For this purpose, two points are determined per user - one point at the body (e.g., at an eye, or between the eyes), and one point at a pointing indicator (e.g., an index finger or a pointing stick), to determine where the respective user is pointing at. This information can be used to adjust the user interface, e.g., by highlighting the position on a digital display, by changing a perceptibility of the user interface (e.g., animate a transition, adjust an opacity, adjust a shape) and/or by annotating the position. For example, a visual representation of the user can be generated and also shown on the display, with a link between the visual representation of the user and the determined position being shown as well. In effect, a visual representation of a user pointing at a certain position in the user interface can be generated and shown on the collaborative display. Using this technique, it becomes evident where the respective users are pointing at on the respective digital display. This is particularly important in scenarios where collaboration is performed, e.g., during surgical procedures, where the main surgeon can see where one or multiple assisting surgeons or even a remote specialist is/are pointing at, e.g., while looking through the oculars of the microscope, or while using a head-mounted display. This can improve collaboration between multiple people, in particular in scenarios where it is relevant where the collaborators are pointing at the screen, such as in surgical procedures, where a surgical site is shown on the screen and discussed by the surgical personnel. For example, once the user focuses on the position of the pointing indicator, the position, e.g., a position of a user interface element, becomes ambiguous, as the eyes of the user perceive the user interface elements at two different positions. This can be overcome by determining the position of a pointing indicator, i.e., the second point of reference, and then moving, by changing the position of the user interface element in the three-dimensional user interface, the user interface element to a pre-defined position relative to the pointing indicator, e.g., just behind the pointing indicator. This leads to a more natural and intuitive user interface, through an easier alignment of the pointing indicator (e.g., fingertip) and the user interface element (e.g., a virtual button).
Some aspects of the present disclosure relate to a system for adjusting a user interface. The system comprises one or more processors. The system is configured to determine a position of a first point of reference based on a position of a first point on the body of a user. The system is configured to determine a position of a second point of reference based on a position of a pointing indicator used by the user. The system is configured to adjust the user interface based on the determined position in the user interface. The system is configured to generate a display signal comprising the adjusted user interface; and provide the display signal to a display device. This enables adjusting the user interface from the point of view of the user, and thus a more intuitive adjustment of the user interface.
In general, various pointing indicators can be used as second point of reference. For example, the system may be configured to determine the second point of reference based on a position of a finger of the hand of the user. A finger is a useful pointing indicator, as it does not require picking up another device to perform the pointing operation. Alternatively, or additionally, the system may be configured to determine the second point of reference based on a position of a pointing device, such as a pointing stick, pen, or scalpel, held by the user. This enables control of the user interface without having to drop the pointing device, e.g., without having to put down a scalpel during a surgical procedure.
In some examples, the wrist of the user can be used as an additional indicator with respect to the pointing direction. For example, the system may be configured to determine the second point of reference based on a position of a wrist and based on a position of one of an index finger of the hand of the user and a pointing device held by the user. This may improve handling of swiveling movements of the finger of the user. In general, the proposed concept may try to take into account the point of view of the user. Thus, the face, and more particularly one or both of the eyes, may be used to determine the first point of reference. For example, the system may be configured to determine the first point of reference based on the position of both eyes of the user, based on the position of a pre-defined or user-defined eye of the user, based on a position of the eye of the user being in closer proximity to the second point of reference, or based on a position of a dominant eye of the user. Thus, the point of view of the user is taken into account.
For example, the system may be configured to determine the position in the user interface or the element of the user interface by projecting an imaginary line through the first point of reference and the second point of reference towards the user interface. For example, the position or element may be determined based on where the imaginary line intersects with the user interface.
There are various ways for determining the two reference points. For example, the system may be configured to process at least one of depth sensor data of a depth sensor, imaging sensor data of a camera sensor and stereo imaging sensor data of a stereo camera to determine at least one of the position of the first point of reference and the position of the second point of reference. For example, sensors that enable determination of a 3D position of the respective points, such as the stereo imaging sensor data or the depth sensor data, may be particularly suitable for determining the position of the points of reference. To improve the quality of the determination of the respective points of reference when depth sensor data is used, the depth sensor data may be used together with (co-registered) imaging sensor data.
In some examples, the system may be configured to determine the positions of joints of a skeletal model of the body of the user by processing the respective sensor data, and to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model. There are various different pre-trained machine-learning models that enable generation of a skeletal model, which may lower the implementation complexity of the system.
Optionally or alternatively, the system may be configured to highlight the determined position in the user interface, To further improve the intuitiveness, a positional indicator may be shown that shows the user the position in the user interface that the system has determined. For example, the system may be configured to highlight the determined position in the user interface continuously using the positional indicator. For example, the system may be configured to generate the positional indicator with the shape of a hand based on a shape of the hand of the user. The shape of the hand is a useful tool for showing the user how the pointing action is being interpreted by the system.
For example the system may be configured to adjust the user interface by providing the determined position in the user interface with a link to the identity of the user. The system may be configured to obtain imaging sensor data of a camera sensor and to determine the identity of the user based on the imaging sensor data. In this way, the determined position can be linked to the user. For example, the system may be configured to provide the link to the identity by generating a visual representation of the user pointing towards the position in the user interface based on the determined identity of the user and based on the determined position in the user interface. By generating the visual representation of the user pointing towards the position in the user interface, it becomes evident where the respective users are pointing at on the respective digital display. This is particularly useful in scenarios where collaboration is useful, e.g., during surgical procedures, and can improve collaboration between multiple people, in particular in scenarios where it is relevant where the collaborators are pointing at the screen, such as in surgical procedures, where a surgical site is shown on the screen.
For example, the system may be configured to generate the visual representation of the user pointing towards the position in the user interface with a positional indicator highlighting the determined position in the user interface, an indicator representing the identity of the user and a visual link between the positional indicator and the indicator representing the identity of the user. This way, both the identity of the user and the position the user is pointing at becomes apparent, which helps other users put into context what the respective user is saying about the highlighted position in the user interface, in particular in high-stakes situations such as the aforementioned surgical procedures.
In various examples, the indicator representing the identity of the user may comprise a textual representation of the identity of the user. A textual representation has a low rendering complexity. However, the system may be required to determine the identity of the user to be able to generate such a textual representation. For example, the system may be configured to generate the visual representation of the user pointing towards the position in the user interface with a positional indicator highlighting the determined position in the user interface, an indicator representing the identity of the user and a visual link between the positional indicator and the indicator representing the identity of the user. This way, both the identity of the user and the position the user is pointing at becomes apparent, which helps other users put into context what the respective user is saying about the highlighted position in the user interface, in particular in high-stakes situations such as the aforementioned surgical procedures.
In various examples, the indicator representing the identity of the user may comprise a textual representation of the identity of the user. A textual representation has a low rendering complexity. However, the system may be required to determine the identity of the user to be able to generate such a textual representation.
Alternatively, or additionally, the indicator representing the identity of the user may comprise an image or an avatar of the user. In the case of the indicator being an image of the user, users can use the system in an ad-hoc fashion, as an image can be obtained without having to determine the actual identity of the user.
While the present concept is already useful for use by a single user, e.g., in the aforementioned surgical scenario, where an assistant is pointing at the screen and the surgeon sees where the assistant is pointing towards or in remote meetings, it becomes increasingly more useful in meetings with multiple users pointing at the screen. For example, the user may be a first user and the position in the user interface may be a first position in the user interface. The system may be further configured to determine a second position in the user interface for a second user, to determine the identity the second user. The system may be configured to generate a visual representation of the first user pointing towards the first position in the user interface and of the second user pointing towards the second position in the user interface. The system may be configured to generate the display signal comprising the user interface with the visual representation of the first user pointing towards the first position in the user interface and of the second user pointing towards the second position in the user interface. This way, multiple users pointing towards the same content may be disambiguated. In some cases, the second position in the user interface may also be determined by the system tracking the second user locally. For example, the act of determining the second position in the user interface may be based on a first point of reference and a second point of reference determined for the second user. This improves the collaboration when both users are standing in front of the same user interface.
Alternatively, or additionally, the act of determining the second position in the user interface and the act of determining the identity of the second user may comprise obtaining information on the second position in the user interface and information on the identity of the second user from a further system being remote from the system. This way, remote collaboration can be achieved.
However, the system might not only be a receiver of information on a second position, it may also provide the information on the position of the (first) user to other systems for remote collaboration. For example, the system may be configured to provide information on the position in the user interface and information on the identity of the user to a further system being remote from the system. This way, remote collaboration can be achieved.
In some cases, it may be sufficient to show images of the users pointing towards the user interface without tracking their identity. In some cases, e.g., when textual representations or avatar representations are used, their actual identity may also be determined. For example, the system may be configured to obtain imaging sensor data of a camera sensor, and to determine the identity of at least the user based on the imaging sensor data, the imaging sensor data showing at least the user. This can be done using imaging sensor data from the camera or cameras being used to track the users, for example.
Both the determination of the points of reference and the identity of the user may be determined using sensor data of various types of cameras. For example, the system may be configured to process at least one of depth sensor data of a depth sensor, imaging sensor data of a camera sensor and stereo imaging sensor data of a stereo camera to determine at least one of the position of the first point of reference and the position of the second point of reference, and an identity of the user. For example, sensors that enable determination of a 3D position of the respective points, such as the stereo imaging sensor data or the depth sensor data, may be particularly suitable for determining the position of the points of reference and/or of the devices. To improve the quality of the determination of the respective points of reference when depth sensor data is used, the depth sensor data may be used together with (co-regis- tered) imaging sensor data. The resulting sensor data may also be used to determine the identity of the users, or to extract images of the users to display in the user interface.
To determine the positions of the points of reference, which are both either at the body of the person or held by the person, a skeletal model may be used. For example, the system may be configured to determine the positions of joints of a skeletal model of the body of the user by processing the respective sensor data, and to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model. There are various different pre-trained machine-learning models that enable generation of a skeletal model, which may lower the implementation complexity of the system.
For example, the system may be configured to generate the shape of the positional indicator such, that, form the point of view of the user, the shape of the positional indicator extends beyond the shape of the hand of the user. In various examples, the system may be configured to determine the position of the positional indicator such, that, form the point of view of the user, the positional indicator may be offset by a pre-defined or user-defined offset. In both cases, the positional indicator can also be seen when the hand of the user is between the eyes of the user and the user interface.
In some examples, the system may be configured to adjust the user interface to animate a transition of the display position of the selected user interface element from an initial display position of the selected user interface element to the display position at the pre-defined distance from the pointing indicator. By including an animation, a visual affordance is provided that signals to the user that the user interface element is ready to be used.
Apart from the animation, other visual affordances may be provided as well. For example, the system may be configured to adjust the user interface to change an opacity of the selected user interface element while the user interface element is shown at the changed display position. For example, the opacity may be increased vis-a-vis other user interface elements that are not selected, highlighting to the user that this specific user interface element is ready to be used. In some examples, the system may be configured to adjust the user interface to change a shape of the selected user interface element while the user interface element is shown at the changed display position. This way, the selected user interface element may be expanded when it is selected, and contracted when it is not, which may be used to increase the options the user has for interacting with the user interface element when it is selected, while reducing the footprint of the user interface element when it is not selected, such that a larger number of user interface elements can be shown at the same time in the user interface.
In particular, the system may be configured to adjust the user interface to sub-divide or extend the selected user interface element into a group of two or more user interface elements. The system may be configured to control the user interface to set the display position of the group of two or more user interface elements such that the group of two or more user interface elements appears to be positioned at the pre-defined distance from the pointing indicator. This way, the selected user interface element may be expanded when it is selected, which can be used to increase the options the user has for interacting with the user interface element when it is selected.
According to an example, the act of determining a response may comprise controlling the user interface based on the determined position in the user interface or element of the user interface. This may enable a more intuitive control of the user interface.
For example, the system may be configured to identify a hand gesture being performed by the user, and to control the user interface based on the determined element and based on the identified hand gesture. This may allow a direct manipulation of the element of the user interface, and thus also improve the intuitiveness of the operation.
For example, the system may be configured to identify at least one of a clicking gesture, a confirmation gesture, a rotational gesture, a pick gesture, a drop gesture, and a zoom gesture. Such gestures are useful for controlling user interfaces.
In some examples, the system may be configured to display a visual representation of the identified hand gesture via a display device or via a projection device. The visual feedback may increase the confidence of the user in controlling the user interface, and thus also improve the intuitiveness of the control.
The user interface is a user interface that is shown on a screen (or projected onto a projection surface). The system is configured to generate a display signal, with the display signal comprising the adjusted user interface, and to provide the display signal to a display device. For example, the display device may be one of a display screen and a head-mounted display. Thus, the proposed concept may be used in a user interface that is shown via a display device.
In some other examples, however, the user interface is not limited to a projection or a screen. In some examples, the user interface spans a space in the real -world, such as an operating theatre, and the devices included in the space in the real -world. For example, the user interface may be a haptic user interface of a physical device. In this case, the determined element may be an element of the haptic user interface. The system may be configured to control a control functionality associated with the element of the haptic user interface. This may enable gesturebased control of devices with haptic user interfaces.
In this case, visual feedback may be provided via projection. For example, the system may be configured to generate a projection signal, the projection signal comprising at least one of a positional indicator for highlighting the determined position of the haptic user interface, an indicator for highlighting the element of the haptic user interface, and a visual representation of the identified hand gesture, and to provide the projection signal to a projection device for projecting the projection signal onto the haptic user interface. The visual feedback may increase the confidence of the user in controlling the haptic user interface via gesture control, and thus also improve the intuitiveness of the control.
In addition, or as an alternative, to visual feedback, tactile feedback may be given. For example, the system may be configured to control an emitter of a contactless tactile feedback system based on at least one of the position in the user interface and the element of the user interface, and to provide tactile feedback to the user using the emitter according to the position in the user interface or according to the element of the user interface. This may further improve the intuitiveness of using the user interface. Another aspect of the present disclosure relate to a corresponding method for controlling a user interface. The method comprises determining a position of a first point of reference based on a position of a first point on the body of a user. The method comprises determining a position of a second point of reference based on a position of a pointing indicator used by the user. The method comprises determining at least one of a position in the user interface, an element of the user interface and a device being accessible via the user interface based on the first point of reference and the second point of reference. The method comprises determining a response based on at least one of the determined position in the user interface, the determined element of the user interface and the determined device.
Another aspect of the present disclosure relates to computer program with a program code for performing the above method when the computer program is run on a processor.
Short description of the Figures
Some examples of apparatuses and/or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which:
Fig. la shows a schematic diagram of an example of a system for controlling a user interface;
Fig. lb shows a schematic diagram of examples of points of reference when the user interface is shown on a screen;
Fig. 1c shows a schematic diagram of examples of points of reference when the user interface is projected or a haptic user interface of a device;
Fig. Id shows a schematic diagram of an example of a surgical imaging system comprising the system;
Fig. 2 shows a flow chart of an example of a method for controlling a user interface;
Fig. 3 shows a schematic diagram for illustrating a misalignment between a cursor and a fingertip of a hand of a user; Fig. 4 shows a schematic illustration for illustrating the effect of a pointer moving with reference to a point of view of a user;
Fig. 5 shows a schematic diagram of examples a hand-shaped positional indicator;
Fig. 6a shows a schematic drawing of a display device with ultrasonic emitters;
Fig. 6b shows a schematic drawing of haptic zones with different tactile feedback;
Fig. 6c shows a schematic drawing of haptic feedback while moving a microscope;
Figs. 6d and 6e show schematic drawing of haptic feedback while operating a rotational user interface element via gesture control;
Fig. 7a shows a schematic diagram of an example of a selection of a device;
Fig. 7b shows a schematic diagram of an example of controlling a selected device;
Fig. 8 shows schematic drawings of gestures being used for controlling a selected device;
Fig. 9a shows a schematic diagram of a user dragging information from a vitals monitor to a screen of a surgical microscope system;
Fig. 9b shows a schematic diagram of a user dragging information from a screen of a surgical microscope system to a surgeon’s head-mounted viewer;
Fig. 10 shows a schematic diagram of a visualization of multiple users;
Fig. 11 shows a schematic diagram of a system comprising an imaging device and a computer system;
Figs. 12a and 12b shows an illustration of a user’s stereoscopic perception of a user interface element, such as a button, in a three-dimensional user interface; Fig. 13 shows use of a rotating button as user interface element, having a virtual distance to the baseline of the user interface; and
Fig. 14 shows an example of annotating the determined position in the user interface based on the determined utterance.
Detailed Description
Various examples will now be described more fully with reference to the accompanying drawings in which some examples are illustrated. In the figures, the thicknesses of lines, layers and/or regions may be exaggerated for clarity.
Fig. la shows a schematic diagram of an example of a system 110 for controlling a user interface. In general, the system 110 may be implemented as a computer system. The system 110 comprises one or more processors 114 and, optionally, one or more interfaces 112 and/or one or more storage devices 116. The one or more processors 114 are coupled to the one or more storage devices 116 and to the one or more interfaces 112. In general, the functionality of the system 110 may be provided by the one or more processors 114, in conjunction with the one or more interfaces 112 (for exchanging data/information with one or more other components, such as a depth sensor 160, a camera sensor 150, a stereo camera 170, a display device 180, a projection device 185 and/or an emitter 190 of a contactless tactile feedback system), and with the one or more storage devices 116 (for storing information, such as machine-readable instructions of a computer program being executed by the one or more processors). In general, the functionality of the one or more processors 114 may be implemented by the one or more processors 114 executing machine-readable instructions. Accordingly, any feature ascribed to the one or more processors 114 may be defined by one or more instructions of a plurality of machine-readable instructions. The system 110 may comprise the machine- readable instructions, e.g., within the one or more storage devices 116.
In the following, the functionality of the system is illustrated with respect to Figs, lb and 1c, in which various points of reference and the user interface are shown. The system 110 is configured to determine a position of a first point of reference 120 based on a position of a first point 12 on the body of a user 10. The system 110 is configured to determine a position of a second point of reference 130 based on a position of a pointing indicator 16 used by the user. The system 110 is configured to adjust the user interface based on the determined position in the user interface. The system 110 is configured to generate a display signal comprising the adjusted user interface and to provide the display signal to a display device 180.
The system 100 may adjust the user interface in different way to account for a user need and/or to improve a user experience. For example, the system 110 may be configured to determine at least one of a position in the user interface, an element of the user interface and a device being accessible via the user interface based on the first point of reference and the second point of reference. The system 110 may be configured to determine a response based on at least one of the determined position in the user interface, the determined element of the user interface and the determined device. In particular, the determined response may be to control the user interface based on the determined position in the user interface or element of the user interface.
Various examples of the proposed concept are based on the finding that, in order to improve the intuitiveness of controlling a user interface by pointing at it, not only the pointing indicator being used (e.g., the hand, or a pointing stick) is to be tracked, but also another point of reference, which is used, together with the pointing indicator, to establish where the user is pointing from the user’s point of view. Therefore, in the proposed concept, two points of reference are determined - a first point of reference at the body of the person (e.g., at the head/face, or at the sternum), and a second point of reference at a pointing indicator used by the person (e.g., at the hand/index finger, or at a pointing device, such as a pointing stick, pen, pencil, or scalpel). Using these two points of reference, not only the relative motion of the pointing indicator can be tracked, but an extrapolation can be made to determine the position (or element or device) the user is pointing at. Therefore, the proposed concept starts with determining the first and second point of reference.
The system is configured to determine the position of the first point of reference 120 based on a position of a first point 12 on the body of the user 10. As outlined above, this first point 12 on the body of the user 10 is a point that can be used as starting point for establishing the point of view of the user. Therefore, the first point 12 may preferably a point at the head/face of the user, such as the forehead, an eye, the position between two eyes, the nose, or the chin of the user. In other words, the system may be configured to determine the first point of reference based on the position of one or both eyes 12 of the user (as shown in Figs, lb, 1c. for example). As most humans use stereoscopic vision, the selection of an eye (or both eyes) as first point of reference may be supported by heuristics or machine-learning-based approaches to further improve the precision. As a compromise, the system may be configured to determine the first point of reference based on the position of both eyes 12 of the user (e.g., between the positions of both eyes). Alternatively, a (system) pre-defined eye or a user-defined eye (e.g., the left or right eye) may be used. In other words, the system may be configured to determine the first point of reference based on the position of a pre-defined or user-defined eye 12 of the user. Alternatively, the system may be configured to select one of the eyes based on the second point of reference, and in particular based on the proximity to the second point of reference. In other words, the system may be configured to determine the first point of reference based on a position of the eye 12 of the user being in closer proximity to the second point of reference. Alternatively, the system may be configured to determine a dominant eye of the user, e.g., by processing imaging sensor data or stereo imaging sensor data to determine whether the user occasionally only uses one eye (and closes the other) and selecting the eye that remains open as dominant eye. Alternatively, the dominant eye may be determined based on previous attempts of the user at controlling the user interface. For example, the system may be select an eye as dominant eye by initially calculating two positions (and intersecting elements) in the user interface based on two different first points of reference based on the positions of both eyes, and determining one of the eyes as dominant eye if a corresponding position in the user interface intersects with a user interface element. In other words, the system may learn which eye is dominant from inference and use this eye for future calculations. The system may be configured to determine the first point of reference based on a position of the dominant eye 12 of the user. In some examples, as highlighted in Fig. le, both eyes of the user may be used as first points 120a; 120b of reference.
In addition to the first point of reference, the system is configured to determine the position of the second point of reference 130 based on the position of a pointing indicator 16 used by the user (the position of the pointing indicator and the second point of reference 130 may correspond to each other). In the present disclosure, the term “pointing indicator” is used, which can both refer to a hand or finger (preferably index finder) being used to point at the user interface and to a tool that is being used for pointing at the user interface. In other words, the pointing indicator 16 is an entity (body part or tool) being used to point at the user interface. Accordingly, the system may be configured to determine the second point of reference (e.g., the position of the pointing indicator) based on a position of a finger 16a of the hand of the user. Additionally, or alternatively, the system may be configured to determine the second point of reference based on a position of a pointing device held by the user. In many cases, both may be an option - if the user uses their hand to point at the user interface, the position of the finger may be used, if, however, the user uses a tool, such as a pen, pencil, pointing stick, scalpel, endoscope etc. that can be considered an extension of the hand, the tool may be used. For example, the system may be configured to determine whether the user is holding a pointing device (e.g., tool/object) extending the hand towards the user interface, and to select the finger of the user if the user is pointing towards the user interface without such a tool/object, and to select the pointing device if the user is pointing towards the user interface with a pointing device. For example, a machine-learning model may be trained, using supervised learning and (stereo) imaging data and/or depth sensor data and corresponding labels, to perform classification as to whether the user is holding a pointing device, and to perform the selection based on the trained machine-learning model. Alternatively, or additionally, image segmentation techniques may be used. The second point of reference may be determined accordingly, based on the position of (a tip of) the finger and based on the position of (a tip of) the pointing device.
In some examples, the user may intuitively swivel (i.e., adjust the pointing direction of) their finger (or pointing device, for that matter) to make fine corrections to the position the user is pointing at. Such swiveling motions may be taken into account too. For example, the system may be configured to determine the second point of reference based on a position of a wrist 16b and based on a position of one of an index finger 16a of the hand of the user and a pointing device held by the user. For example, the position of the second point of reference may be adjusted such, that the resulting intersection point in the user interface when an extrapolated line (in the following also denoted imaginary line) passing through the first and second point of reference up to the user interface is determined, is adjusted in line with the swiveling motion of the finger (and, in some case, disproportional to the change in position of the tip of the index finger).
To determine the respective points on the body (e.g., head, sternum), hand or pointing device, sensor data may be processed. For example, as shown in Fig. la, the system may be configured to obtain sensor data from one or more of a depth sensor 160, a camera sensor 150 (which may be combined with the depth sensor 160 to provide RGB-D (Red, Green, Blue-Depth) sensor data) and a stereo camera 170 (i.e., two camera sensors forming a stereo pair), and to determine the positions on the body, hand and/or pointing device based on the respective sensor data. In other words, the system may be configured to process at least one of depth sensor data of a depth sensor 160, imaging sensor data of a camera sensor 150, combined RGB-D sensor data of a combined depth sensor and camera sensor, and stereo imaging sensor data of a stereo camera 170 to determine at least one of the position of the first point of reference and the position of the second point of reference.
As both points of reference are at least related to parts of the body of the user (with the first point of reference being on the body, and the second point of reference being part of the hand or held by a hand of the user), three-dimensional (human) pose estimation may be applied on the respective sensor data to determine the respective points of reference. There are various algorithms and machine-learning-based approaches for generating three-dimensional poseestimation data from two- and three-dimensional video data. They usually include generation of a three-dimensional skeletal model of the user based on the respective sensor data. Accordingly, the system may be configured to determine the positions of joints of a (three-dimensional) skeletal model of the body of the user by processing the respective sensor data. In the context of the present disclosure, the term “skeletal model” might not be understood in a biological sense. Instead, the skeletal model may refer to a pose-estimation skeleton, which is merely modeled after a “biological” skeleton. Thus, the generated three-dimensional poseestimation data may be defined by a position of joints of a skeleton in a three-dimensional coordinate system, with the joints of the one or more skeletons being interconnected by limbs. Similar to above, the terms “joints” and “limbs” might not be used in their strict biological sense, but with reference to the pose-estimation skeleton referred to in the context of the present disclosure. Such a skeletal model may be generated using the respective sensor data specified above. Accordingly, the system may be configured to process at least one of the depth sensor data, the imaging sensor data, the combined RGB-D sensor data, and the stereo imaging sensor data to determine the skeletal model of the user. For example, machine-learning algorithms and machine-learning models that are known from literature or that are available off the shelf may be used for this purpose. Examples of such algorithms and models are given in the following. For example, Mehta et al: "vNect: Real-time 3D Human Pose Estimation with a Single RGB Camera" (2017) provide a machine-learning based three-dimensional pose estimation algorithm for generating a three-dimensional skeletal model from imaging data. Zhang et al: "Deep Learning Methods for 3D Human Pose Estimation under Different Supervision Paradigms: A Survey" (2021) provide a machine-learning based three-dimensional pose estimation algorithm for generating a three-dimensional skeletal model from RGB-D data. Lallemand et al: "Human Pose Estimation in Stereo Images" provide a machine-learning based three-dimensional pose estimation algorithm for generating a three-dimensional skeletal model from stereo imaging sensor data. Alternatively, an application-specific machinelearning model may be trained, in line with the approaches listed above.
The resulting skeletal model may now be used to determine the points of reference. In other words, the system may be configured to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model. For example, the skeletal model may include, as joints, one or more of the position of the eyes of the user, the position of the forehead of the user, the position of the tip of the nose of the user, the position of the wrists of the user, the position of the tips of the (index) fingers of the user. In addition, the skeletal model may include an additional point representing the tip of a pointing device being held by the user. For example, the machine-learning model may be trained to include such an additional point in the skeletal model. Alternatively, the system may be configured to determine whether the user is using a pointing device to point at the user interface, to determine a position of the tip of the pointing device relative to a point on the hand of the user (based on the skeletal model and based on the respective sensor data), and to add a joint representing the tip of the pointing device to the skeletal model. For example, another machine-learning model may be trained to output the position of the tip of the pointing device relative to a point on the hand of the user, e.g., using supervised learning. For example, such a machine-learning model may be trained as regressor, using a skeletal model and the respective sensor data as input, and the relative position as desired output.
The first and second point of reference are now used to determine at least one of a position in the user interface, an element of the user interface and/or a device being accessible via the user interface. For this purpose, the system may be configured to determine the position in the user interface, the element of the user interface or the device by projecting an imaginary line 140 through the first point of reference and the second point of reference towards the user interface. The position in the user interface may correspond to the point where the imaginary line intersects the user interface. The element of the user interface may correspond to an element of the user interface that overlaps with the point in the user interface. The device may correspond to a device that is intersected by the imaginary line. However, the proposed concept is not limited to using a geometric approach that is based on projecting an imaginary line. Instead, a look-up table may be used to determine the position, element and/or device, or a machine-learning model may be trained to output the position, from which the element or device can be determined (similar to the case where an imaginary line is used). In this case, supervised learning may be used to train such a machine-learning model, using coordinates of the first and/or second point of reference, and optionally coordinates of the user interface, as training input and information on a position in the user interface as desired output.
In the present context, various positions are used, such as the point on the body/first point of reference, the second point of reference, the position in the user interface etc. These positions may be defined in a (or several) coordinate systems. For example, the first and second point of reference may be defined in a world coordinate system. If the coordinates of the user interface are known, e.g., also defined in the world coordinate system (which may be defined relative to the user interface), the imaginary line can be projected through the first and second points of reference towards the user interface, and the point of in the user interface (and element intersecting with the point) can be determined based on the point of intersection of the imaginary line. If another technique is used, such as a look-up table or trained machine-learning model, similarly, a world coordinate system or two separate coordinate systems may be used for the points of reference and the user interface. It is to be noted, that the user interface is not necessarily limited to being a two-dimensional user interface or a user interface that is being shown on a display device - in some examples, as will become clear in connection with Figs, lb and 1c, the user interface may be a three-dimensional user interface, which may be either a real-world user interface with haptic/tactile input modalities, or a three-dimensional user interface being shown on a 3D display, a three-dimensional projection or a hologram.
The system is configured to determine a response based on the determined position in the user interface, the determined element of the user interface and/or the determined device. Such a response may be used to determine, based on the position or movement of the position in the user interface, whether the user is trying to interact with, and thus control, the user interface, or whether the user is not engaged with the user interface. In the latter case, the determined position and/or element of the user interface may be used to control the user interface. In other words, the system may be configured to control the user interface based on the determined position in the user interface or element of the user interface. In this context, the “element” may be a user interface element, such as a button, a label, a dropdown menu, a radio button, a checkbox, a rotational input, or a slider of a virtual user interface (being displayed on a screen or projected), or a button, a label, a radio button, a rotational knob, or a slider of a haptic/tactile “real-world” user interface (that is arranged, as dedicated input modalities, at a device).
To control the user interface, hand gestures may be used. For example, the system may be configured to identify a hand gesture being performed by the user, and to control the user interface based on the determined el em ent/ selected user interface and based on the identified hand gesture. For example, the system may be configured to process the sensor data, using a machine-learning model, to determine the hand gesture being used. Such a machine-learning model may be trained, using supervised learning, as a classifier, using the respective sensor data as input data and a classification of the gesture as desired output. For example, examples of such a machine-learning model and algorithm can be found in Bhushan et al: "An Experimental Analysis of Various Machine Learning Algorithms for Hand Gesture Recognition" (2022). Various hand gestures may be identified and used to control the user interface. For example, some examples of gestures are discussed in connection with Fig. 8. For example, the system may be configured to identify at least one of a clicking gesture, a confirmation gesture, a rotational gesture, a pick gesture, a drop gesture, and a zoom gesture. For example, to perform the clicking gesture, the pointing indicator may be moved intentionally towards the user interface, e.g., for a pre-defined distance in at most a pre-defined time. This gesture may be used to select or actuate a user interface element. To perform the rotational gesture, the user may rotate the hand, as illustrated in Fig. 13. This gesture may be used to control a rotational input or rotational knob of the user interface. To perform the pick gesture and the drop gesture, the user may first move the pointing indicator intentionally towards an element of the user interface, draw the pointing indicator back towards the body (pick gesture), and move the pointing indicator towards another position in the user interface (drop gesture). To perform the zoom gesture, the user may move the hand, with palm pointed towards the user interface, towards or away from the user interface. Alternatively, the gestures may be implemented differently, and/or used to control different elements of the user interface.
To increase the confidence of the user controlling the user interface, visual feedback may be given to the user. In particular, the system may be configured to highlight the determined position in the user interface continuously using a positional indicator. In a screen-based user interface (as shown in Fig. lb) or projection -based user interface (as shown as one example in Fig. 1c), such a positional indicator may be included in the representation of the (virtual and optional three-dimensional) user interface itself, e.g., overlaid over the user interface. For example, the system may be configured to generate a display signal, with the display signal comprising the user interface (e.g., and the positional indicator). The system may be configured to provide the display signal to a display device 180, such as a display screen or a headmounted display. Alternatively, the system may be configured to generate a projection signal, which may comprise (in some cases, where the user interface itself is being projected) the user interface, and the positional indicator. The system may be configured to provide the projection signal to a projection device 185 for projecting the projection signal onto the haptic user interface. It is to be noted, that such a projection can also be used with “haptic” user interfaces (i.e., user interfaces that are arranged, as haptic input modalities, at the respective devices being controlled, without including a projection of the user interface.
To allow the user to grasp the mechanism, the positional indicator may be provided in a shape that conveys whether the system has determined the correct position or item. For example, the system may be configured to generate the positional indicator with the shape of a hand based on a shape of the hand of the user. For example, the system may be configured to compute the shape of the hand of the user based on the respective sensor data, and generate the positional indicator based on the shape of the hand of the user (or based on a simplified hand shape). Examples for such shapes are shown in Fig. 5, for example. In some examples, as shown on the left in Fig. 5, the system may be configured to generate the shape of the positional indicator such, that, form the point of view of the user, the shape of the positional indicator extends beyond the shape of the hand of the user, e.g., by “exploding” the shape by a pre-defined absolute margin, such that the user can see the shape even if the hand partially obscures the user interface. Alternatively, as shown on the right in Fig. 5, the system may be configured to determine the position of the positional indicator such, that, form the point of view of the user, the positional indicator is offset by a pre-defined or user-defined offset (e.g., in two dimensions).
In addition, a visual representation of the identified hand gesture may be shown, e.g., as a pictogram representation of the hand gesture or of the functionality being controlled by the hand gesture. For example, the system may be configured to display a visual representation of the identified hand gesture via a display device 180 (e.g., a 3D screen for use with shutter glasses or polarized glasses) or via a projection device 185 (e.g., for use with shutter glasses, or for projecting a hologram)..
To add a tactile dimension, an emitter of a contactless tactile feedback system can be used to give tactile feedback to the user. For example, the system may be configured to control an emitter 190 of a contactless tactile feedback system based on at least one of the position in the user interface and the element of the user interface, and to provide tactile feedback to the user using the emitter according to the position in the user interface or according to the element of the user interface. For example, as discussed in connection with Figs. 6a to 6e, the emitter 190 may be one of an ultrasound emitter or array of ultrasound emitters, one or more laser emitters, one or more air emitters, or one or more magnetic field emitters. For example, when the position in the user interface (or an imaginary line) intersects with an (user interface) element (e.g., when the user moves the pointing indicator or the interface element is thus selected), tactile feedback may be given. Similarly, tactile feedback can be given when the user interface is being controlled, e.g., if a button is actuated by a clicking gesture or a rotational input/knob is rotated using a rotational gesture.
In some examples, different levels or patterns of tactile feedback may be given, for example, a first level or first pattern of tactile feedback may be given when an element of the user interface or device is selected, and a second level or second pattern of tactile feedback may be given while the user interacts with the element of the user interface or with the device. A third level or third pattern of tactile feedback may be given when the user reaches a boundary, e.g., while controlling an element of the user interface or a device via gesture control. For example, if the user controls a virtual rotational element (e.g., rotational knob) via gesture control, the third level or third pattern of tactile feedback may be provided when the rotational element is at an extreme position (e.g., lowest, or highest allowed position). Similarly, if the user controls movement of a device, such as a microscope, the third level or third pattern of tactile feedback may be provided when the position of the device reaches a limit, e.g., due to stability restrictions or due to a limit of an actuator).
In the following, two example implementations are shown. In Fig. lb, an example is shown where the user interface is displayed on a display device (such as a 3D screen). Alternatively, the user interface 100 may be a projected user interface, e.g., a user interface that is projected onto a projection surface (not shown). Fig. lb shows a schematic diagram of examples of points of reference when the user interface is shown on a screen. Fig. lb shows a user interface 100. Two stereo cameras 170 are arranged above (or below) the user interface 100, having two partially overlapping fields of view 170a, 170b. The stereo cameras provide stereo imaging sensor data, which are used to determine the two points of reference - the first point of reference 120 at the position of an eye 12 of a user 10, and the second point of reference 130 at the position of a fingertip 16a (of the index finger) of the hand of the user. Through the first and second point of reference, an imaginary line 140 is projected, pointing at a position in the user interface 100. If a user interface element overlaps with the imaginary line, it may be selected.
In Fig. 1c, a different paradigm is used. In this case, the user interface 100 can be either a projected user interface, e.g., a user interface that is projected onto a projection surface, or the user interface can be a haptic user interface (indicated by the zig-zag shape) that is arranged at one or multiple devices being controlled by the system. Fig. 1c shows a schematic diagram of examples of points of reference when the user interface is projected or a haptic user interface of a device. In the example of Fig. 1c, a depth sensor 160 and a camera sensor 150 are used to generate RGB-D sensor data, which is processed by the system 110 to determine the first point of reference 120 and the second point of reference 130. An imaginary line 140 is projected towards the user interface. In one case, the user interface is projected by projection device 185 onto a projection surface 185a. Alternatively, only indicators might be projected onto the projection surface 185a. For example, the projection signal may comprise at least one of a positional indicator for highlighting the determined position of the haptic user interface, an indicator for highlighting the element of the haptic user interface, and a visual representation of the identified hand gesture.
As shown in one interpretation of Fig. 1c, the user interface may be a haptic, real -world user interface, as indicated by the rotational knob shown as part of the user interface 100. Alternatively, the knob may be part of the projected user interface. In the example of Fig. 1c, the position in the user interface intersects with the (real -world or projected) rotational knob, such that the rotational knob is determined as user interface element and/or such that the device comprising the rotational knob is determined as device. Further examples for the determination of a device are given in connection with Figs. 7a to 7b, for example. The rotational knob may now be controlled via a hand gesture, e.g., in the projected user interface, or by controlling the control functionality associated with the knob. In other words, if the user interface is a haptic user interface of a physical device, the determined element may be an element of the haptic user interface. Similarly, the determined device may be a device that is accessible (i.e., controllable) via the haptic user interface. In this case, the system may be configured to control a control functionality associated with the element of the haptic user interface (e.g., without turning the knob or pressing a button of the haptic user interface) and/or associated with the determined device. This way, the user can control the device or user interface element in a way that is not unlike telekinesis.
Such ways of controlling a user interface are particularly useful in scenarios, where, due to hygiene or as the user is unable to move around, a user is unable to control the user interface by touch. Such a scenario can, for example, be found in surgical theatres, where a surgeon seeks to control a surgical imaging device, such as a surgical microscope or surgical exoscope (also sometimes called an extracorporeal telescope), or other devices, such as an endoscope, an OCT (Optical Coherence Tomography) device. Exoscopes are camera-based imaging systems, and in particular camera-based 3D imaging systems, which are suitable for providing images of surgical sites with high magnification and a large depth of field. Compared to microscopes, which may be used via oculars, exoscopes are only used via display modalities, such as monitor or a head-mounted display. Such devices are usually part, or attached to, a surgical imaging system, such as a surgical microscope system or a surgical exoscope system. Thus, the system may be configured to control a user interface of the surgical microscope system (and connected devices). Fig. Id shows a schematic diagram of an example of a surgical imaging system comprising the system 110. In the case shown in Fig. Id, the surgical microscope system shown in Fig. Id comprises a base unit 20 with a haptic user interface (indicated by the knobs and buttons). Using the projection device 185, the aforementioned indicators may be projected onto the haptic user interface. During surgery, the surgeon (or an assistant) may point towards the haptic user interface, e.g., the knobs and buttons, and control the control functionality associated with the different elements of the user interface using hand gestures.
In some examples, the system 110 may be an integral part of such a surgical imaging system and may be used to control various aspects of the optical imaging system 100. In particular, the system 110 may be used to process imaging data of an imaging device (e.g., the microscope or exoscope) of the optical imaging system, and to generate a display signal based on the image data. Therefore, the system 110 may also serve as an image processing system of the surgical imaging system. In addition, the system 110 may be configured to control additional aspects of the surgical imaging system, such as an illumination provided by the surgical imaging system, a placement of the imaging device relative to a surgical site being imaged, or one or more auxiliary surgical devices being controlled via the surgical imaging device.
A surgical imaging system is an optical imaging system 100 that comprises a (surgical) imaging device, such as a digital stereo microscope, an exoscope or an endoscope, as imaging device. However, the proposed concept is not limited to such embodiments. For example, the imaging device is often also called the “optics carrier” of the surgical imaging system.
The one or more interfaces 112 of the system 110 may correspond to one or more inputs and/or outputs for receiving and/or transmitting information, which may be in digital (bit) values according to a specified code, within a module, between modules or between modules of different entities. For example, the one or more interfaces 112 may comprise interface circuitry configured to receive and/or transmit information. The one or more processors 114 of the system 110 may be implemented using one or more processing units, one or more processing devices, any means for processing, such as a processor, a computer or a programmable hardware component being operable with accordingly adapted software. In other words, the described function of the one or more processors 114 may as well be implemented in software, which is then executed on one or more programmable hardware components. Such hardware components may comprise a general -purpose processor, a Digital Signal Processor (DSP), a micro-controller, etc. The one or more storage devices 116 of the system 110 may comprise at least one element of the group of a computer readable storage medium, such as a magnetic or optical storage medium, e.g., a hard disk drive, a flash memory, Floppy- Disk, Random Access Memory (RAM), Programmable Read Only Memory (PROM), Erasable Programmable Read Only Memory (EPROM), an Electronically Erasable Programmable Read Only Memory (EEPROM), or a network storage.
In the following, the functionality of the system is illustrated with respect to Figs, le, If and 1g, in which eyes 120a; 120b of a user 10, a pointing indicator 130, one or more imaginary lines (140; 140a; 140b) and a user interface element (145; 145a-d) are shown. The system 110 may be configured to determine a position of a pointing indicator 130 used by a user relative to a three-dimensional user interface 100 (shown in the background as a screen). The system 110 may be configured to select, based on the position of the pointing indicator, a user interface element 145 of the three-dimensional user interface. The system 110 may be configured to control the user interface to change a display position of the selected user interface element in the three-dimensional user interface based on the position of the pointing indicator, such that the selected user interface elements appears to be positioned at a pre-defined distance from the pointing indicator. In this way, a perception of the user interface can be adjusted by the system 100. That is, an adjustment of the user interface may be an adjustment of the position of the user interface. Optionally or alternatively, further parameters of the user interface such like opacity, shape, number of windows can be adjusted by the system 100.
A three-dimensional user interface is a user interface that uses three-dimensional (3D) space to interact with and display information. Examples of three-dimensional user interfaces include virtual reality systems, augmented reality applications, and 3D modeling tools. In the present disclosure, contactless gestures are used to navigate through a virtual 3D environment to control the three-dimensional user interface. Various examples of the present disclosure are based on the finding, that the use of a three-dimensional user interface (i.e., a user interface shown on a 3D display, 3D projection, holographic projection, that supports showing user interface elements at different depths relative to a user) via gesture-based control is often deemed unintuitive, at is hard to reconcile the position of a user interface element with the position of a pointing indicator being used to control the three-dimensional user interface. However, when the user focuses on the pointing indicator (e.g., the user’s finger), due to stereoscopic vision, the user may not be able to unambiguously perceive the position of the user interface elements, as the position of the user interface element appears at different positions, depending on which eye the user uses. This is illustrated in Fig. 12a, which shows an illustration of a user’s stereoscopic perception of a user interface element, such as a button 140, in a three-dimensional user interface. In Fig. 12a, the upper image indicates what the screen shows, the middle image indicates what the user sees, and the lower image shows what the user perceives. In the upper image of Fig. 12a, the position of the eyes 120a; 120b are shown, together with the position of the pointing indicator 130 (e.g., finger) and two imaginary lines 140a; 140b originating from the positions of the eyes 120a; 120b and terminating at two different points of the user interface. The user interface element 145 is between the two termini of the imaginary lines. As a result, when the uses focuses on the pointing indicator 130, the user perceives the user interface element 145 at two different positions 145c; 145d, depending on the eye being used. The stereoscopic vision prevents the user from seeing the fingertip overlaying a virtual button as if they were physically at the same distance. On the contrary, the user sees either two buttons 145a; 145d or two fingertips, depending on where the user is focusing. As a result, some users tend to resolve this perception conflict by closing one eye.
In the proposed concept, this difficulty may be overcome by placing the user interface element in close proximity to the pointing indicator, e.g., directly behind (as seen from the point of view of the user) the pointing indicator, “inviting” the user to use the user interface element. This may be done by adjusting or controlling the user interface to change the display position of the selected user interface element in the three-dimensional user interface based on the position of the pointing indicator, such that the selected user interface elements appears to be positioned at a pre-defined distance from the pointing indicator. In other words, the position of the selected user interface element may be changed to a position indicated by the position of the pointing indicator, such that the user interface element appears to be near (e.g., at or directly behind) the pointing indicator, which may be achieved by varying the positioning of the user interface elements along a z-axis extending outwards from the three-dimensional user interface, and optionally along the corresponding x- and y axis. The positioning along the z- axis may be changed to reduce the perceived distance between the user interface element and the pointing indicator. For example, the system may be configured to control the user interface to change the display position of the selected user interface element in the three-dimensional user interface to a position that appears to be at a distance of at most 5 cm (or at most 4 cm, at most 3 cm, at most 2 cm, at most 1 cm) from the position of the pointing indicator. In three- dimensional user interfaces, such a positioning along the z-axis (reaching outwards from the three-dimensional user interface) can be achieved by placing the user interface element at different positions in the three-dimensional user interface, depending on which eye the image may be provided to. To differentiate which eye an image of the three-dimensional user interface may be presented to, shutter glasses or glasses admitting light of different polarizations to the different eyes may be used. Shutter glasses are a type of 3D glasses that work by isolating the left and right eyes from each other. The glasses use liquid crystal technology to rapidly alternate the view for each eye, blocking one eye while allowing the other to see an image. This process may be synchronized with the 3D content on the screen. Glasses that use different polarizations for the left and right eye have a similar effect, also isolating the left and right eyes from each other. In this case, the light of the image to be presented to the left eye has a different polarization than the light of the image to be presented to the right eye, and polarization filters being part of the glasses are used to admit the light of one polarization to the respective eye and block the other, such that the different eyes perceive different images. In both techniques, each eye may be seeing a slightly different perspective of the same image of the three-dimensional user interface, creating the illusion of depth and a three-dimensional effect. The respective glasses are usually paired with a display device or projection device that has a high refresh rate (in case of shutter glasses) or the ability to emit light separately in different polarization (in the other case), to display the three-dimensional user interface to the user.
The ability to show different images to the two eyes enables the three-dimensional effect. To change the placement of the user interface element along the z-axis, to place the user interface element (directly) behind the pointing indicator, the user interface element may be shown at different positions in the user interface, such that the user perceives the user interface element at a desired depth. In other words, the three-dimensional user interface may be based on providing a first image to a first eye of the user and a second image to a second eye of the user. The system may be configured to control the user interface to change the display position of the selected user interface element in the three-dimensional user interface, such that the selected user interface elements appears to be positioned at a pre-defined distance from the pointing indicator, by including the user interface element at different positions 145a, 145b in the first and second image. This effect is illustrated in Figs, le and If. Fig. le shows a schematic diagram of an example of possible positions of a user interface element in two channels (i.e., the first and second image shown to the first and second eye of the user). As shown in Fig. le, when the same user interface element 145 is rendered at different positions 145a; 145b in the different channels, according to imaginary lines 140a; 140b from the eyes 120a; 120b of the user through the focusing point (the pointing indicator 130) towards the respective positions 145a; 145b in the first and second image, the user interface element 145 appears to be at the position of the pointing indicator 130, as illustrated in Fig. If. Fig. If shows a schematic diagram of an example of a resulting perceived position of the user interface element 145 at the pointing indicator 130.
In many user interfaces, more than one user interface element may be present. For example, Fig. 1g shows a schematic diagram of an example of a three-dimensional user interface with five user interface elements 145-1 to 145-5 (of which user interface elements 145-4 and 145- 5 are back and forward buttons, respectively), and an indicator 105 for highlighting the position the user is pointing at in the three-dimensional user interface. These user interface elements may have different types. While the present concept has, thus far, been explained with respect to virtual buttons that are at the fingertips of the respective user, other types of user interface elements may be implemented as well. For example, the user interface element may be one of a button, a display element, a menu, a radio button, a checkbox, a rotational input (e.g., a virtual rotational knob) and a slider. To provide visual feedback to the user with respect to the selected user interface element, the selected user interface element may be differentiated from other user interface elements that are not selected. Such visual feedback may be accomplished using different techniques. For example, one type of visual feedback relates to animating the change in position of the selected user interface element. In other words, the system may be configured to control the user interface to animate a transition of the display position of the selected user interface element from an initial display position of the selected user interface element (i.e., the position of the user interface element prior to being selected) to the display position at the pre-defined distance from the pointing indicator. For example, the system may be configured to control the user interface to gradually move the selected user interface element from the initial display position to the display position at the pre-defined distance from the pointing indicator. The animation of the transition may also affect the size, shape and/or opacity of the selected user interface. For example, the user interface element may be perceived to increase in size as the selected user interface element from the initial display position to the display position at the pre-defined distance from the pointing indicator. Additionally, or alternatively, the system may be configured to control the user interface to change (e.g., increase) an opacity (i.e., decrease a transparency) of the selected user interface element while the user interface element is shown at the changed display position (i.e., at the display position at the pre-defined distance from the pointing indicator. Additionally, or alternatively, the system may be configured to control the user interface to change a shape of the selected user interface element while the user interface element is shown at the changed display position.
A variation of the latter effect is shown in Fig. Ih. In Fig. Ih, the selected user interface element 145, which initially has a circular shape, is changed into having a rectangular shape with rounded corners, and then sub-divided into four separate user interface elements 145-1 to 145-4, providing a sub-menu for controlling four different aspects via the three-dimensional user interface. Fig. Ih shows a schematic diagram of an example of a user interface element 145 being extended into four user interface elements 145-1 to 145-4. In more general terms, the system may be configured to control the user interface to sub-divide or extend the selected user interface element into a group of two or more user interface elements. In this context, sub-dividing the user interface element may occur without changing the shape of the user interface element (apart from the size) and extending the user interface element may comprise changing the shape and sub-dividing the changed shape into different user interface elements, or sub-dividing the initial magnified shape into different shapes corresponding to different user interface element. In this case, the resulting two or more user interface elements are placed at the pre-defined distance from the pointing indicator. In other words, the system may be configured to control the user interface to set the display position of the group of two or more user interface elements such that the group of two or more user interface elements appears to be positioned at the pre-defined distance from the pointing indicator.
The selection of the user interface element and positioning of the user interface element is based on the position of the pointing indicator. For this purpose, the system the system may be configured to determine the position of a pointing indicator 130 used by a user relative to a three-dimensional user interface, and to select, based on the position of the pointing indicator, the user interface element 145 of the three-dimensional user interface. This can be done by sensing the position of the pointing indicator, which may be a finger of the user, or a pointing device held by the user, such as a pen or a scalpel, using one or more suitable sensors. For example, the system may be configured to process at least one of depth sensor data of a depth sensor 160, imaging sensor data of a camera sensor 150 and stereo imaging sensor data of a stereo camera 170 to determine at least the position of the pointing indicator. Additionally, the system may be configured to process the respective sensor data to determine at least one position on the body of the person, as will become evident in the following.
In the present context, various positions are used, such as the position of the pointing indicator (which is, in the following, also denoted second point of reference), the position on the body of the person (which can be used to determine first point of reference), and the position of the user interface element. These positions may be defined in a (or several) coordinate systems. For example, the first and second point of reference may be defined in a world coordinate system. If the coordinates of the three-dimensional user interface are known, e.g., also defined in the world coordinate system (which may be defined relative to the user interface), the imaginary line can be projected through the first and second points of reference towards the user interface, and the point of in the user interface (and element intersecting with the point) can be determined based on the point of intersection of the imaginary line. If another technique is used, such as a look-up table or trained machine-learning model, similarly, a world coordinate system or two separate coordinate systems may be used for the points of reference and the user interface.
To determine the position or positions (of the first and/or second point of reference), various techniques may be used. In the context of the present disclosure, determination of the point on the body of the person, and thus of the first point of reference is optional. However, it enables a more intuitive use of the user interface, as will be outlined in the following.
More details and aspects of the system are mentioned in connection with the proposed concept, or one or more examples described above or below (e.g., Fig. 2 to 14). The system may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
Fig. 2 shows a flow chart of an example of a corresponding method for controlling a user interface, e.g., in line with the system shown in connection with Figs, la to Id. The method comprises determining 210 a position of a first point of reference based on a position of a first point on the body of the user. The method comprises determining 220 a position of a second point of reference based on a position of a pointing indicator used by the user. The method comprises determining 230 a position in the user interface based on the first point of reference and the second point of reference. The method comprises adjusting 240 the user interface based on the determined position in the user interface and generating 250 a display signal comprising the adjusted user interface. The method comprises providing 260 the display signal to a display device.
For example, the method may be performed by the system and/or surgical imaging system introduced in connection with one of the Figs, la to Id, 3 to 14. Features introduced in connection with the system and optical imaging system of one of Figs, la to Id, 3 to 14 may likewise be included in the corresponding method.
More details and aspects of the method are mentioned in connection with the proposed concept, or one or more examples described above or below (e.g., Fig. la to Id, 3 to 14). The method may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below. Various examples of the present disclosure relate to a concept for intuitive contactless control of a computer screen. The proposed concept aims to improve the functionality of contactless control of a user interface, e.g., a user interface of a computer that is shown on a screen, using the hand in the air. The proposed concept allows the user to intuitively point with their hand within their own field of view. This may be achieved by using a camera (or other sensor) to capture the 3D position of two reference points, e.g., the user’s eyes and hand and, from that, calculate where the user is pointing. As a result, the user has a touch screen-like experience, while merely pointing in the air toward the screen.
As outlined above, known implementations of computer control via cameras that capture the user’s hand often lack intuitiveness and thus are very rarely used. The lack of intuitiveness may originate from the fact that the cursor follows the movements of the user’s fingertip, but in the user’s view, the cursor and the fingertip are not aligned, as shown in Fig. 3.
Fig. 3 shows a schematic diagram for illustrating a misalignment between a cursor 340 and a fingertip of a hand 330 of a user, from the point of view of a user. Fig. 3 shows a screen 310 and a camera 320 for recording the position of the fingertip of the hand 330 of the user.
In such implementations, the user may need to train their brain to move the cursor by doing a disproportional movement of the hand, and the vision and hand operate a closed loop until the cursor is brought to the desired position on the screen. On the other side of the spectrum, touch screens allow a very natural, intuitive, and fast interaction with a computer graphics interface. However, in many cases touching the screen is not possible due to distance or risk of contamination.
Various examples of the proposed concept are based on the finding, that the lack of intuitiveness can be overcome by acquiring the position in the user interface (e.g., on the screen) the user is pointing to according to the user’s view/perspective. This may be achieved by capturing the positions of two reference points, such as the position of the user’s eyes and finger- tip/hand relative to the user interface (e.g., screen). From these positions, the position in the user interface (e.g., on the screen) may be calculated, using a straight line through the eyes and the fingertip/hand (see Fig. lb). Fig. lb shows a schematic diagram illustrating determination of the position of a position or element in a user interface from the point of view of a user. Fig. lb shows a user interface 100. Two stereo cameras 170 are arranged above (or below) the user interface 100, having two partially overlapping fields of view 170a, 170b. The stereo cameras provide stereo imaging sensor data, which can be used to determine two points of reference - a first point of reference 120 at the position of an eye 12 of a user 10, and a second point of reference 130 at the position of a fingertip 16a (of the index finger) of a hand of the user. In addition, a potential third point of reference at the position of the wrist 16b of the user is shown. Through the first and second point of reference, an imaginary line 140 can be projected, pointing at a position in the user interface 100. As Fig. lb illustrates a top view, the position at the intersection of the imaginary line 140 and the user interface 100 corresponds to the x-position pointed to by the user 10. However, the same concept can be applied in a second dimension as well to determine the y-position.
The proposed concept improves vision-based contactless control of a user interface (e.g., of a computer, of a microscope system, or of a tangible user interface) by determining the absolute position in the user interface (e.g., screen position) that the user is pointing to, instead of the relative motion offered by other systems. As a result, contactless control may become more intuitive, natural, and fast as touch screens.
As a result, the user may directly “point and click” instead of the necessity, in other systems, of having to find the cursor, and then keep moving until the cursor arrives to the intended position. In the proposed concept, the movement from one position to another, e.g., drag and drop, may happen with the speed of pointing rather than by moving the hand slowly while observing the cursor’s motion toward the intended position. Similarly, input of a path on the screen, e.g., outlining an object, may become faster if the cursor follows the fingertip instead of the cursor moving with a different speed than the fingertip (as the user-observed 2D spaces of the fingertip and the screen are not in a (perceived) 1 : 1 scale), which is illustrated in Fig. 4.
Fig. 4 shows a schematic illustration for illustrating the effect of a pointer moving with reference to a point of view 445 of a user. Fig. 4 shows a screen 410 and stereo cameras 420, 430 with corresponding overlapping fields of view 425 and 435. In the example of Fig. 4, the stereo camera captures the position of the hand 450 and eyes 440 of the user. The cursor 460 is spatially aligned with the hand 450 to appear at the fingertip from the user’s point of view.
The implementation shown in Fig. 4 is based on a stereo camera. Alternatively, any other 3D imaging sensor can be used, for example an RGB-D (Red, Green, Blue, Depth) sensor, which can be implemented as a time-of-flight camera combined with a regular RGB camera(s).
While the proposed concept primarily relates to contactless pointing, to exploit the full potential, the concept can be combined with gestures to represent actions such as mouse-click, pick and drop, rotate, zoom in and out. For example, it may be considered intuitive and fast to move the hand above a (displayed or real-life) rotating knob, and do the rotation gesture, or to pick and drop without waiting to obtain a visual confirmation about the location.
To improve confidence while using this technology, one or more of the following features may be used. For example, the hand may be visualized on screen as an outline, or a transparent mask. In perfect alignment, this visualization may be hidden behind the user’s hand. For someone other than the user, observing the screen, the misalignment will be clear. Similarly, the current recognized action or intention may be visualized. For example, pointing using the index finger may mean using the mouse/pointer function (pointing and clicking), index and thumb fingers pointing may mean ready to adjust zoom, opening the palm (similar to stop gesture) may mean pan. In some examples, the user may be allowed to set a spatial offset of the calculated position. For example, the displayed hand shadow may be shown shifted 200 pixels up-left from the hand’s position. This may allow more precise location and handling as the screen will not be hidden behind the user’s hand. Similarly, the shadow may be “exploded” (which is different from making the hand linearly bigger) so that the shadow is visible around the hand. The options listed above are illustrated in Fig. 5.
Fig. 5 shows a schematic diagram of examples a hand-shaped positional indicator. In a first example shown on the left, the shadow 510 that serves as hand-shaped position indicator is “exploded” to be visible around the hand (e.g., by extending beyond the hand, from the point of view of the user, by a pre-defined margin). In a second example shown in the middle, the shadow 520 is well-aligned and is only visible due to stereoscopic vision. In a third example shown on the right, the shadow 535 is set to an offset north-west of the hand and the precise click point is highlighted 540. While the implementation may be easier if the user were to use only one eye, human vision is stereoscopic, so adjustments may be made. For example, the midpoint of the calculations of the two eyes may be used, or the user may be enabled to choose an eye in the settings, or an eye may be chosen automatically when one eye is closed by the user.
All gestures described above may be combined with positioning in a user interface (e.g., on the screen). However, it is also possible to obtain global actions by positioning the hand closer to the eyes so that the hand looks bigger than the screen. This may be recognized as a global action, e.g., to rotate the image, or change desktop/mode.
In some examples, the contactless gesture control may be expanded to work on 3D screens. This may allow the user to navigate more precisely in the 3D view shown by a 3D display device, such as a 3D monitor.
More details and aspects of the contactless gesture control are mentioned in connection with the proposed concept, or one or more examples described above or below (e.g., Fig. la to 2, 6a to 14). The contactless gesture control may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
In the following, an example of contactless haptic (or tactile) feedback is given. Haptic feedback is valuable in medical devices like surgical microscopes because it enhances the surgeon's precision and control during delicate procedures. By providing real-time touch sensations, haptic feedback allows the surgeon to better gauge the applied force and spatial orientation, ultimately improving the accuracy and safety of the operation.
In the following, it is proposed to use contactless haptic feedback in contactless control of devices such as surgical microscopes.
For example, various technologies may be used to provide haptic feedback. For example, ultrasound-based haptic feedback may be used. By utilizing focused ultrasonic waves, this technology creates contactless haptic feedback through the generation of pressure points in mid-air, allowing users to feel tactile sensations without touching a physical surface. For example, Inami, M., Shinoda, H., & Makino, Y. (2018): “Haptoclone (Haptic-Optical Clone) for Mutual Tele-Environment” provide an example of ultrasound-based haptic feedback. As an alternative, laser-based haptic feedback may be used. Using laser-induced plasma, this method generates haptic feedback by creating a shockwave in the air, producing a touch sensation when interacting with the plasma. For example, Futami, R., Asai, Y., & Shinoda, H. (2019): “Hapbeat: Contactless haptic feedback with laser -induced plasma” shows the use of haptic feedback based on laser-induced plasma. Another technology generates air vortices. Controlled air vortices are to deliver targeted haptic feedback without physical contact, enabling users to experience a touch-like sensation from a distance. Sodhi, R., Poupyrev, I., Glisson, M., & Israr, A. (2013): “AIREAL: Interactive tactile experiences in free air” give an example of this technology. Another alternative is the user of electromagnetic haptic feedback: By manipulating magnetic fields, this technique generates contactless haptic feedback that can simulate various textures and sensations, allowing users to interact with virtual or remote objects in a more realistic manner. Wang, M., Zhang, X., & Zhang, D. (2014): “Magnetic Levitation Haptic Interface Based on Spherical Permanent Magnet” is an example of the application of this technology. Some of the above-referenced techniques may be used as part of wearable devices to provide contactless feedback. Innovations in wearable technology, such as smart gloves, utilize combinations of the aforementioned techniques to provide haptic feedback without direct skin contact, offering enhanced user experiences in virtual reality, gaming, and medical applications (see, for example, Kim, Y., Yoo, B., & Choi, S. (2018): “A Wearable Vibrotactile Haptic Device for Virtual Reality: Design and Implementation”). Another known technique is the use of vibration modules is mobile phones to provide haptic feedback that resembles pressing a physical button.
Contactless control is not widely used, and the combination of contactless control and haptic feedback is only discussed as a proof of concept. As a result, in the proof-of-concept studies, the design is still on a basic level, primarily aiming on merely simulating the haptic feedback of touching an object. However, contactless control in three-dimensional space allows a wide variety of use and implementation of haptic feedback, which is not known from the aforementioned proof-of-concept implementations. Moreover, contactless control of devices, such as a microscope, is still an unexplored space.
In the following, a concept is provided for improving the user experience of contactless control by adding contactless haptic feedback. In various examples of this concept, a 3D space map may be used that defines the type of feedback to be provided to the user by any type of contactless haptic creation technology such as ultrasound. Fig. 6a illustrates an example implementation of the concept using ultrasound emitters 190. In this implementation, a contactless control monitor is equipped with ultrasound emitters arranged perimetrically around the monitor frame. Fig. 6a shows a schematic drawing of a display device 180 with ultrasonic emitters 190. In the example of Fig. 6a, the ultrasound emitters 190 are arranged at a frame of the display device 180, surrounding a display area of the display device. When the user, shown by the first and second points of reference 120; 130 at the eyes of the user and the finger of the user uses contact-less control to control a virtual button 610 of the user interface, the ultrasound waves (shown as circles) focus at the finger 130 of the user.
Fig. 6b shows a schematic drawing of haptic zones with different tactile feedback, and in particular the concept of zones of feedback. In Fig. 6b, 3D-zones are defined where the fin- gertip/hand will receive different type of haptic feedback. In this example the virtual button defines two zones: the “button stand-by zone” 620 (when the fingertip is aligned with the button) where a week haptic feedback informs the user that the fingertip is at a position that the user can press a virtual button (i.e., the finger is aligned with the button). The other zone is the “button activation zone” 630 (when the button is pressed) which provides feedback that the virtual button is pressed (akin to confirmation of a click).
When controlling a device, such as a microscope, via gesture control, in particular alignment control is of great importance and value for the users (e.g., surgeons). However, the lack of feedback makes the contactless control counterintuitive, awkward in feeling, and as a result slows down the process. An important aspect of the use of contactless haptic feedback is to provide haptic feedback control a device and thus improve the experience and efficiency of the microscope control. Fig. 6c illustrates an example of haptic feedback in the contactless control of a microscope positioning. Fig. 6c shows a schematic drawing of haptic feedback while moving a microscope. Fig. 6c shows a microscope 640 (which is equipped with ultrasound emitters, shown as circles), which can be moved between two positions 640a; 640b, which may be considered virtual boundaries, and which define an angle adjustment range. While moving the orientation of the microscope, the microscope stays orientated towards a surgical cavity. In the example shown the user’ s hand moves near to the microscope and when enters a predefined distance, then the microscope locks and follows the user’s hand. The haptic feedback serves a guide for the alignment and as confirmation of the user’s actions. For this purpose, haptic confirmation 650a is given when the device is engaged, less pronounced haptic feedback (indicated by the short bars between the extreme positions and the middle position) is given during movement, and increasingly strong haptic feedback 650b, 650c is given as haptic warning indicating that the boundary is reached. In Figs. 6c and 6e, the intensity of the haptic feedback increases, the more the respective bars extend from the arrowed circle indicating the angle. A way to describe the concept of programming the haptic feedback is to define 3D boundaries and zones which define the type of haptic feedback. In this example the virtual boundaries 640a, 640b are derived by the microscope’s physical limitations of adjustment range, i.e., the microscope arm cannot move beyond. Alternatively, the virtual boundaries could be determined by other limitations such as the alignment with the surgical cavity, i.e., the microscope is restricted within the angles that allow to look in the surgical cavity.
In addition to the 3D boundaries and zones described above, which are fixed in space, zones and boundaries positioned relative to the microscope’s position can also be useful. For example, when the hand is used to virtually move and guide the microscope, there can be defined a “virtual grabbing zone” which is always around the microscope. This means that there is an area that when the hand is in then it can be recognized as a potential control guide. In that case it would be useful to provide haptic feedback as to if the user’s hand is at correct spot.
Figs. 6d and 6e illustrate an example of haptic feedback designed for a specific type of control, in this case a rotation knob. Figs. 6d and 6e show schematic drawing of haptic feedback while operating a rotational user interface element via gesture control. The haptic feedback pattern shown in Fig. 6e, represents the feedback during the rotation of the virtual knob shown in Fig. 6d, and provides three message types: (i) That the end of the rotation range is being reached by offering a gradually increasing haptic intensity 660b, 660c near the range end (i.e., haptic feedback that the boundary is reached), (ii) confirmation of the rotation by providing discrete haptic clicks resembling a real rotation knob (the smaller bars between the extremes and the middle, and (iii) confirmation that the knob is at the zero position (i.e., haptic feedback 660a that the device is engaged). The haptic feedback is not restricted in the rotation feedback as shown in Figs. 6d and 6e but can also assist the user to align the hand and thus have more confident in the use of a control element, and consequently use the interface easier and faster. Examples of alignment feedback is to provide haptic confirmation (i) when the hand is aligned, and (ii) when the gesture is recognized, (iii) when the hand is getting misaligned during the use of the control element, and (iv) when the gesture is getting difficult to be recognized.
For example, the haptic feedback may be used to guide the user in the user interface discussed in connection with Figs, la to 5, 7 to 14. For example, the system 110 of the discussed figures may be configured to control an emitter of a contactless tactile feedback system, such as one or more ultrasound emitters, one or more laser emitters, one or more air emitters or one or more magnetic field emitters to provide tactile feedback in response to the user controlling the selected user interface element (e.g., when the user actuates an element of the user interface or performs gesture control on an element of the user interface), or in response to the user selecting an element of the user interface or a device. For example, the haptic feedback one of the examples shown in Figs. 6a to 6e may be provided by the system controlling an emitter of a contactless tactile/haptic feedback system.
The example shown in Figs. 6d and 6e illustrates a rotating knob, but the same concept can be extended in other virtual control elements such as a sliding bar, and a switch. The goal of the haptic feedback in virtual controls is to resemble the experience of a physical control unit, so the user can have a natural, intuitive, and efficient utilization.
As a result of the contactless tactile feedback, the user may have a natural and intuitive experience. The confirmation may offer certainty to the user about whether the contactless control is activated or not. The intuitive and assertive experience may allow the user to use the virtual control faster and more efficiently. The confirmation may also reduce errors, i.e., false unintentional activation, misinterpretation of gestures, and activation failures.
More details and aspects of the haptic feedback mechanism are mentioned in connection with the proposed concept, or one or more examples described above or below (e.g., Fig. la to 5, 7a to 14). The haptic feedback mechanism may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
In the following, an application of the concept shown in Figs, la to Id, 2, is shown, in which the concept is used to select and control devices. In other words, the act of determining a response comprises selecting a target device of two or more devices based on at least one of the determined position in the user interface, the determined element of the user interface and the determined device. The act of determining a response further comprises controlling the target device. In this application of the concept, the proposed concept allows easy, fast, precise, and intuitive control of devices. The user does not need to use any additional equipment such as handheld devices or wearing eye-trackers. It also supports multiple users, limited only by the resolution and the field of view of the cameras. For example, the proposed concept may be used by multiple users simultaneously.
The presently discussed application of the proposed concept relates to contactless device control, i.e., contactless control of devices. In complex working environments such as surgical rooms, there are multiple devices, and controlling them typically requires physical contact. In the following, a contactless methodology of fast and intuitive control of devices is presented. This methodology builds upon the techniques introduced in connection with Figs, la to Id (and optionally 3 to 6e, 9a to 10), i.e., the use of finger-pointing contactless control for one or multiple devices. In Figs. 7a, 7b and 7c, additional details are shown on how the techniques introduced in connection with Figs, la to 6e, 9a to 10 can be used to select and control different devices.
For example, the platform (i.e., the system 110 introduced in connection with Figs, la to Id) may continuously analyze users’ positions and gestures until he/she puts their hand in the direction of a device (e.g., points towards a device). Fig. 7a shows a schematic diagram of an example of a selection of a device. Fig. 7a shows the first point of reference 120 (illustrated by the eyes of the user, a surgeon) and the second point of reference 130 (illustrated by the hand of the user). The platform continuously analyzes the users’ position and gesture while the user moves their hand until the user puts their hand in the direction of a device. Accordingly, the system may be configured to select the target device by determining which device of the two or more devices the is pointing at based on the first point of reference and based on the second point of reference. This may be done using the imaginary line extending through the first and second points of reference towards the respective devices, as discussed in connection with Figs, la to Id. In other words, the system may be configured to determine which device the user is pointing at by projecting an imaginary line 140 through the first point of reference and the second point of reference towards the two or more devices. The system may be configured to select one of the two or more devices as target device based on the imaginary line intersecting the respective device. The user moves his hand to select a device, and accordingly points at one of three devices- an Image-Guided Surgery (IGS) device 710, a surgical microscope 720, and an endoscope control unit 730 to select the device. If an imaginary line from the first point of reference 120 and through the second point of reference 130 is within a device recognition zone or detection area 710a; 720a; 730a around one of the devices 710; 720; 730, the respective device is selected. These device recognition zones, or detection areas, may be used to facilitate the selection of the respective devices, in particular in large settings, such as surgical theatres, where devices may be small, and distances may be large. For example, the system may be configured to determine which device the user is pointing at by projecting an imaginary line 140 through the first point of reference and the second point of reference towards the two or more devices. The system may be configured to select one of the two or more devices as target device based on the imaginary line intersecting a detection area 710a, 720a, 730a encompassing the respective device. For example, the detection area may extend beyond the device, e.g., by a pre-defined distance surrounding the respective device, or by a pre-defined number of degrees of a circle having its center point at the first point of reference, the pre-defined number of degrees extending from an imaginary line between the first point of reference and a center point of the respective device. The latter approach is shown in Fig. 7a.
To establish a spatial relationship between the different parties, i.e., the user and the devices, different approaches may be taken. In some cases, the devices have a fixed arrangement, e.g., when they are integrated in a surgical theatre. For example, the surgical microscope’s arm may be attached at a fixed point at the ceiling, and the endoscope and the IGS may be integrated or attached to the wall of the operating room. In this case, a fixed spatial relationship exists between the devices, and therefore potentially the coordinate system of the user interface. In such cases, or in cases where an initial calibration is performed to establish the fixed spatial relationship between the devices, the system may be configured to determine the position of the two or more devices relative to the user based on a fixed (spatial) relationship between the two or more devices and at least one sensor being used to determine the position of the first and second point of reference. Based on the fixed (spatial) relationship, the system may determine which device the user is pointing towards, e.g., using the imaginary line technique. Alternatively, in addition to the positions of the first and second points of reference, also the positions of the two or more devices may be determined (e.g., by reading them from a database, or using the same sensors or other sensors being used to determine the position of the first and second points of reference) and may be used to determine which device the user is pointing towards, e.g., using the imaginary line technique. In other words, the system may be configured to determine a position of the two or more devices relative to the user, and to select the target device further based on the position of the two or more devices relative to the user. For example, the system may be configured to determine the position of the two or more devices relative to the user based on sensor data of at least one sensor. For example, the two or more devices may have visual identifiers (e.g., two-dimensional codes) visibly printed or attached to their respective casings, which may be used to identify the respective devices and to include them for selection and may be used as further point of reference for determining their respective position. Similar to the determination of the position of the first and second points of reference, different sensors may be used for this purpose. For example, the system may be configured to process at least one of depth sensor data of a depth sensor 160, imaging sensor data of a camera sensor 150, and stereo imaging sensor data of a stereo camera 170 to determine at least one of the position of the first point of reference, the position of the second point of reference, and a position of one or more devices.
An indication on each device, such as a color switching light, may be used to indicate the status of each device. For example, the indication may be provided by a display or light of the respective device, or the indication may be projected onto the respective device. For example, the system may be configured to control the target device to display an indication 725 that the device is selected. Alternatively, or additionally, the system may be configured to control a projection device 185 to project, onto the selected device, an indication that the device is selected. For example, a first visual indicator may be shown if the system is selected, and a second visual indicator may be shown if a control command is being performed. A third visual indicator may be shown, when a device is available for being selected. In the following examples, a green light is used to indicate that a device is ready to be selected, a yellow light is used to indicate that a device is selected and ready to be controlled via gesture controls, and a red light is used to indicate that a device is currently being controlled.
When a device is selected by the user, the indication may change (e.g., from green to yellow) so that the user is sure about the device being selected/activated. Then each user gesture may be executed by the device (see Fig. 7b). Accordingly, an indicator light 725 of the selected device is activated or switched to a different color, e.g., from green to yellow. For example, a green indicator light may indicate that the corresponding device is available but not selected, a yellow indicator light may indicate that the corresponding device is selected and waiting for a command, and a red indicator light may indicate that a command is being executed.
Fig. 7b shows a schematic diagram of an example of controlling a selected device. In Fig. 7b, selected device 720 executes (immediately) the user’s gesture command. In other words, as outlined in connection with Figs, la to Id, the system may be configured to identify a hand gesture being performed by the user, and to control the target device based on the identified hand gesture. When a command gesture is detected the device indicator changes (e.g., the indicator light may become red, coming from yellow) and the device execute the command. In the example of Fig. 7b, the user moves the microscope 720 higher by raising the open palm upwards. Examples of gesture commands are shown in Fig. 8.
Fig. 8 shows schematic drawings of gestures being used for controlling a selected device. In Fig. 8, a click gesture 810 is shown, in which the index finger of the hand is moved towards the palm of the hand, and which may be used to activate a virtual button. In addition, a grab gesture 820 is shown, in which the index finger and the thumb of the hand are moved towards each other, and which may be used for a drag & drop operation. Moreover, a turn gesture 830 is shown, in which the hand is turned around an axis defined by the arm, and which may be used to adjust a virtual dial (rotational) button. Additionally, a scroll gesture 840 is shown, in which the open palm of the hand points upwards and is raised or lowered, which may be used to move a microscope up or down. Finally, a rotate gesture 850 is shown, in which the fingers of the hand are rotated towards the inner side of the wrist, and which may be used to rotate an image.
In some cases, a device may have more than one aspect that is controllable via gesture controls. For example, in addition to changing the working distance of a surgical microscope, as shown in Fig. 7b, the microscope’s angle towards the surgical site may be adjusted via gesture control, as shown in Fig. 6c. In this case, the system may disambiguate between the different control functionalities. For example, the system may be configured to determine an element of a user interface (i.e., an aspect of the target device that is to be controlled) of the target device based on the first point of reference and based on the second point of reference, and to control a control functionality associated with the element of the user interface based on the identified hand gesture. For example, different sections of the respective devices may be associated with different control functionalities. Depending on where the user is pointing, a different control functionality may be controlled. For example, if the user is pointing at the objective of the microscope, the focus may be controlled. If the user is pointing at the handles of the microscope, the working distance or angle may be controlled. For this purpose, the system may further select the control functionality being controlled based on the gesture being performed.
In some cases, it can help the intuitiveness of the control if additional visual feedback is given (e.g., in addition to the visual indicator indicating the state of the respective devices). For example, the system may be configured to display a visual representation of the identified hand gesture or of a functionality of the device being controlled by the hand gesture (e.g., an icon representing the identified hand gesture or an icon representing the functionality) via a display device 180 or via a projection device 185. Thus, in particular in case a device provides different control functionality that is being controllable via gesture control, the user can be sure that he is controlling the desired functionality.
The detection of the users’ positions and gestures can be done using sensors on each device, a central sensor system in the room, or a combination of the two. In other words, the system may be part of one or each of the devices, or the system may be associated with two or more devices. In various examples of the present disclosure, the system 110 is assumed to be part of the surgical microscope system, with the other devices also being controllable via the system 110 of the surgical microscope system. However, the system 110 may also be separate from the surgical microscope.
Moreover, in the present examples, it was assumed that a central set of sensors is used to control all of the devices. However, in an alternative example, each device may operate independently using embedded sensors. The devices may be interconnected using peer-to-peer connection or a central communication hub. For example, interconnected devices may support “collective control of devices”, e.g., drag and drop the output image of an endoscope to be displayed on a microscope’s monitor, as discussed in connection with Figs. 9a and 9b.
More details and aspects of the contactless gesture control are mentioned in connection with the proposed concept, or one or more examples described above or below (e.g., Fig. la to 6e, 9a to 14). The contactless gesture control may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
In the following, an application of the concept shown in Figs, la to Id, 2, is shown, in which the concept is used to transfer information between devices of an ecosystem of devices, thereby creating a contactless ecosystem, with contactless collective control of an ecosystem of devices. Using contactless control, an intelligent way to control an ecosystem of connected devices is proposed. This is based around the technique of allowing the user to configure a set of devices and displays so that the information if collected and displayed in a preferred way, using gestures as if the actual physical space is a VR (Virtual Reality) space. In a simple medical application example, a surgical room may be equipped with a surgical microscope and a vital monitor (shown in Fig. 9a). If the surgeon wants to display the vitals on the microscope screen, the surgeon can do that by dragging & dropping the vital monitor into the microscope screen (see Fig. 9a). Another medical example is when a user wants to send a video stream from a microscope or endoscope to the head-mounted digital viewer of a surgeon (see Fig. 9b). The interaction with the device ecosystem is very intuitive and fast and allows for a more agile adaptation of the ecosystem according to the dynamic shift of needs, as it is easy to add and remove devices when needed or not needed. It also enables more flexibility in the physical arrangement of devices in a room, as the user can collect all the data without the need to have optical access to each device.
Fig. 9a shows a schematic diagram of a user dragging information 915 from a vitals monitor 910 to a screen 920 of a surgical microscope system using a drag & drop gesture. On top of the display, the stereo cameras 170 for detection of the user pointing are shown. The user can easily use virtual drag & drop between devices to transfer the display of one device on the monitor as picture in picture.
Fig. 9b shows a schematic diagram of a user dragging information 925 from a screen 920 of a surgical microscope system to a surgeon’s head-mounted viewer 930. The user can easily send a video stream from a screen to a surgeon’s head mounted digital viewer.
To achieve this using the techniques illustrated in connection with Figs, la to Id, the system 110 is configured to determine the position of the first point of reference 120 based on the position of the first point 12 on the body of a user 10 for a first point of time and a second point of time. The system 110 is configured to determine the position of the second point of reference 130 based on the position of the pointing indicator 16 used by the user for the first point of time and the second point of time. The system 110 is configured to determine a first device based on the first point of reference and the second point of reference for the first point of time and a second device based on the first point of reference and the second point of reference for the second point of time. The system 110 is configured to determine the response based on the determined first and second device, the act of determining the response comprising controlling the second device to display information 915, 925 (shown in Figs. 9a and 9b) provided by the first device. In terms of the method of Fig. 2, the method comprises determining a position of a first point of reference based on a position of a first point on the body of a user for a first point of time and a second point of time. The method comprises determining a position of a second point of reference based on a position of a pointing indicator used by the user for the first point of time and the second point of time. The method comprises determining a first device based on the first point of reference and the second point of reference for the first point of time and a second device based on the first point of reference and the second point of reference for the second point of time. The method comprises determining a response based on the determined first and second device. The act of determining a response comprises controlling the second device to display information provided by the first device.
The transfer of information from the first device to the second device is based on the user pointing at both devices. In general, it was found to be more intuitive to first point at the device containing the information and then at the device, on which the information is to be shown. Accordingly, the second point of time may be later than the first point of time. The user pointing at both devices may be part of an overarching gesture, which includes pointing sequentially at both devices (to indicate which of the devices is to act as a source and which of the devices is to act as display). Accordingly, the system may be configured to identify a pre-defined gesture being performed by the user, and to determine the first and second point in time based on the pre-defined hand-gesture. This pre-defined gesture may include pointing at the first device at the first point in time and pointing at the second device at the second point in time. In other words, the user first points at the first device, and then at the second device. However, these two pointing actions may be connected by the remainder of the gesture, to indicate that the user does not merely want to select one of the devices for gesture control. In addition to the user pointing at the first and the second device, the pre-defined gesture may also include a transition between both devices. For example, during the transition, the user may keep their arm outstretched, so that, instead of two separate pointing gestures, a single pointing gesture is performed that is merely moved towards the second device. Alternatively, a grab gesture may be used, i.e., the user may point at the first device, perform the grab gesture while pointing at the first device and keep performing the grab gesture until the user points at the second device, where the user releases the grab gesture. Fig. 8 shows a schematic diagram of an example grab gesture 820, in which the index finger and the thumb of the hand are moved towards each other. This gesture may be used for a drag & drop operation, such as the one performed in the present context. Accordingly, the pre-defined gesture may be a drag-and-drop gesture. The start and end of the grab gesture may also define the first and second points in time.
In an ideal implementation, all devices are digitally interconnected. In this case, the system 110 can obtain the information provided by the first device from the first device and control the second device to display the information obtained from the first device. In other words, the system may be configured to obtain the information provided by the first device from the first device. However, it might not be always possible to have all the devices connected. In that case, an alternative implementation is to use cameras to video capture the screen of a device, e.g., a vital monitor (see Fig. 9a), and then display that information on the screen of the second device (e.g., the microscope screen) without any digital connection between the microscope and the vitals monitor. In other words, the system may be configured to obtain imaging sensor data (of a camera sensor/stereo camera sensor) showing the information provided by the first device from a camera being separate from the first device, the imaging sensor data comprising the information provided by the first device. In this case, while the first device provides the information, the providing of the information merely refers to showing the information on a screen of the first device, i.e., the information is provided by showing the information on a screen of the first device, and not digitally via a network. Accordingly, the imaging sensor data may comprise a representation of information shown on a screen of the first device. To make this information usable, the imaging sensor data may be processed, and the relevant information may be extracted. For example, the system may be configured to extract the information by the first device from the imaging sensor data, by cropping the imaging sensor data and/or geometrically transforming the imaging sensor data (e.g., to perform perspective corrections). For example, the system may be configured to process the imaging sensor data to determine a portion of the imaging sensor data showing the information, and to crop and/or geometrically transform the imaging sensor data to extract the information. The system is configured to control the second device to display information provided by the first device. For this purpose, the system may be configured to provide the information provided by the first device for displaying to the second device. For example, the system may be configured to provide a display signal for a display device of the second device, the display signal comprising a representation of the information provided by the first device. In this context, the display signal may be a signal for driving the display device of the second device, or it may be a video stream to be displayed by the second device, e.g., within a window or dedicated area on the display device of the second device. In other words, in the first option, the system drives the display device of the second device. For example, the system 110 may be part of the second device. This may be the case in the scenario shown in Figs. Id and 9a, where the second device is a surgical imaging system 100 (such as a surgical microscope system or a surgical exoscope system). For example, some examples of the present disclosure relate to a surgical imaging system, comprising the system 110 and at least one of the first device and the second device. In the second option, the second device is merely interconnected and part of the same ecosystem as the system, with the second device accepting the display signal / video stream for displaying as part of the ecosystem.
As outlined in connection with Figs, la to Id, multiple cameras may be placed at different positions and devices to ensure sufficient optical coverage in the whole room, e.g., for both determination of the first and second points of reference, and, optionally, also for providing imaging sensor data showing the screen of a device. For example, the system may be configured to process at least one of depth sensor data of a depth sensor 160, imaging sensor data of a camera sensor 150 and stereo imaging sensor data of a stereo camera 170 to determine at least one of the position of the first point of reference and the position of the second point of reference, and a position of one or more devices. In particular, the system may be configured to determine the positions of joints of a skeletal model of the body of the user by processing the respective sensor data, and to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model. For example, each device could be equipped with cameras, and also security-like cameras could have better overview in the room. The cameras could be used collectively to ensure seamless operation. More details and aspects of the concept for a contactless ecosystem of devices are mentioned in connection with the proposed concept or one or more examples described above or below (e.g., Fig. la to 8, 10 to 14). The concept for a contactless ecosystem of devices may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
In some examples, the techniques discussed in connection with Figs, la to Id, 2 may be used for contactless collaboration, e.g., to support contactless pointing on a monitor for collaboration. A technology is proposed that allows multiple users to point at the display simultaneously. The proposed concept uses a contactless detection of each user’s pointing area, as described in connection with Figs, la to Id, and visualizes the pointing areas, together with a representation of the users, on the screen (see Fig. 10). Fig. 10 shows a schematic diagram of a visualization of multiple users lOa-lOd, and in particular a visualization of pointing of multiple users at a screen. As shown in Fig. 10, the visualization may contain different components - an image or avatar 1010a- lOlOd representing the respective users, a highlighted area 1020a-1020d (e.g., highlighted by a circle, or “cursor”), and a visual link 1030 between the two image/avatar and the highlighted portion. This allows an easy, fast, precise, and intuitive communication between the users. The users do not need to use any additional equipment such as handheld devices or wearing eye-trackers. Multiple users are supported, limited only by the resolution and the field of view of the cameras. In some examples, the technique does not require user-pointer assignment, e.g., if the face of each user is shown in real time.
Such a visualization may be implemented using the techniques discussed in connection with Figs, la to Id. In particular, in terms of the system discussed in connection with Figs, la to Id, the system is configured to determine a position in the user interface based on the first point of reference and the second point of reference. Further, the system may be configured to obtain imaging sensor data of a camera sensor, determine the identity of the user based on the imaging sensor data and adjust the user interface by providing the determined position in the user interface with a link to the identity of the user. That is, the user interface can be adjusted by linking the identity of the user, e.g., by indicating the identity of a user in a displayed image.
In an example, the system may be configured to provide the link to the identity by generating a visual representation of the user pointing towards the position in the user interface based on the determined identity of the user and based on the determined position in the user interface. For example, the system may be configured to determine the response by determining an identity of the user lOa-d, and generating a visual representation lOlOa-d, 1020a-d, 1030 of the user pointing towards the position in the user interface based on the determined identity of the user and based on the determined position in the user interface. The system is configured to generate a display signal comprising the user interface with the visual representation of the user pointing towards the position in the user interface, and to provide the display signal to a display device 180. This visual representation is a visualization of the user pointing at the user interface, and therefore the collaborative screen. In other words, the user interface is shown on the display device 180, and the user is pointing towards the display device 180 on which the user interface is shown. Similarly, with respect to the method of Fig. 2, the method comprises determining a position in the user interface based on the first point of reference and the second point of reference. The method comprises determining a response based on the determined position in the user interface. The act of determining the response comprises determining an identity of the user and generating a visual representation of the user pointing towards the position in the user interface based on the determined identity of the user and based on the determined position in the user interface. The method comprises generating a display signal comprising the user interface with the visual representation of the user pointing towards the position in the user interface. The method further comprises providing the display signal to a display device.
The system is configured to generate a visual representation of the user pointing towards the position in the user interface. This visualization necessarily includes at least two components - an indication of the position in the user interface, and a representation of the user. In addition, a link between the position and the user helps other users associate the position with the position in the user interface. Accordingly, the system may be configured to generate the visual representation of the user pointing towards the position in the user interface with a positional indicator 1020a-d highlighting the determined position in the user interface, an indicator lOlOa-d representing the identity of the user and a visual link between the positional indicator and the indicator representing the identity of the user. In Fig. 10, the positional indicator is a circle 1020a-1020d highlighting the respective position in the user interface. However, alternative positional indicators are possible, such as circle with a different brightness etc. The indicator representing the identity of the user is a (real-time or stored) image 1010a- lOlOd of the user in Fig. 10. In other words, the indicator representing the identity of the user may comprise an image or an avatar of the user. For this purpose, the image may be recorded using a camera, e.g., a camera being used to determine the positions of the points of reference. In other words, the system may be configured to obtain imaging sensor data (e.g., camera sensor data) of a camera sensor 160, 170 (e.g., of a camera or stereo camera), with the imaging sensor data showing at least the user. This imaging sensor data may be used to generate the image (or the avatar) of the user. In other words, the system may be configured to generate the image, or the avatar of the user based on the imaging sensor data. For example, the system may be configured to generate (and continuously update) a real-time image of the user based on the imaging sensor data. If only the image of the user is shown, the actual identity of the user need not be determined. In this case, determining the identity of the user corresponds to linking a user with an image or avatar, with the image or avatar representing the identity of the user. In other words, the image or avatar may correspond to the identity of the user in this case, without requiring actual identification of the user.
If the identity of the user is to be shown, e.g., alongside the image/avatar or without showing the image/avatar, the identity of the user may be determined. For example, the system may be configured to determine the identity of at least the user based on the imaging sensor data. For example, the system may be configured to use machine-learning to determine the identity of the user (and of other users using the system). For example, the system may be configured to process the image of the user (from the imaging sensor data) using a machine-learning model being trained to determine a feature vector representing the user, e.g., a machine-learning model trained for person reidentification, or a deterministic algorithm for determining the feature vector. The resulting feature vector (or reidentification code) may then be compared to feature vectors/reidentification codes of known users stored in a database to identify the user. Using the identity of the user, an avatar of the user, or a textual representation (which may include the name and/or the title/job of the user) may be loaded and displayed by the system. In other words, the indicator representing the identity of the user may comprise a textual representation of the identity of the user. In other words, identification of the pointers may be done using face recognition and visualized as contact photo, avatar, or name.
As shown in Fig. 10, a visual link 1030 may be displayed between the positional indicator 1020 and the indicator 1010 representing the user. In Fig. 10, lines were used as visual links 1030 between the two items. Alternatively, the indicator 1010 representing the user may be shown in close proximity to the positional indicator 1020, thereby creating a contextual visual link between the two components of the visual representation of the user pointing towards the position in the user interface.
As previously discussed in connection with Figs, la to Id, various sensors may be used to determine the positions of the points of reference. In the present application of the concept, the same sensors (or a different sensor) may also be used for generating the imaging sensor data being used to identify the user or to generate an image of the user as indicator representing the identity of the user. For example, the system may be configured to process at least one of depth sensor data of a depth sensor 160, imaging sensor data of a camera sensor 150 and stereo imaging sensor data of a stereo camera 170 to determine at least one of the position of the first point of reference and the position of the second point of reference, the image of the user and the identity of the user. At least the positions of the points of reference may be determined using a skeletal model. For example, the system may be configured to determine the positions of joints of a skeletal model of the body of the user by processing the respective sensor data, and to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model. The skeletal model may also be used for generating the image of the user (by identifying the position of the head of the user using the skeletal model and cropping the imaging sensor data based on the position of the head of the user), and/or for identifying the user (e.g., by determining the feature vector at least partially based on the skeletal model).
The generation of a visual representation of a user pointing at a position in the user interface is particularly useful in scenarios, where multiple users point at the content at the same time, e.g., during a discussion or medical consultation. Therefore, the concept is applicable to multiple users. In other words, multiple visual representations of a user pointing at a position in the user interface may be shown at the same time as part of the user interface. In other words, the aforementioned user may be a first user 10a and the position in the user interface may be a first position in the user interface. The system may be further configured to determine a second position in the user interface for a second user 10b, and to determine the identity the second user. The system may be configured to generate a visual representation of the first user pointing towards the first position in the user interface and of the second user pointing towards the second position in the user interface, and to generate the display signal comprising the user interface with the visual representation of the first user pointing towards the first position in the user interface and of the second user pointing towards the second position in the user interface. In other words, the visual representation of the first user pointing at a first position in the user interface is extended with a visual representation of a second user (third user, fourth user) pointing at a second (third, fourth) position in the user interface.
This second user may be a local user, pointing at the same display device as the first user. In this case, the act of determining the second position in the user interface may be based on a first point of reference and a second point of reference determined for the second user. In other words, the system may be configured to determine a first point of reference and a second point of reference for the second user, and to determine the second position in the user interface based on the first point of reference and second point of reference determined for the second user.
Alternatively, or additionally, the collaboration may include multiple systems, physically separated so that users in different rooms can collaborate. In this case, the relevant information may be received from the remote system. In other words, the act of determining the second position in the user interface and the act of determining the identity of the second user may comprise obtaining information on the second position in the user interface and information on the identity of the second user from a further system being remote from the system. In other words, the system may be configured to determine the second position in the user interface and the identity of the second user by receiving information on the second position in the user interface and the identity of the second user from the further system. Similarly, the respective information of the first user may be provided to the further system as well, for display by the further system. For example, the system may be configured to provide information on the (first) position in the user interface and information on the identity of the (first) user to the further system being remote from the system.
When multiple users use the user interface, it is rare that all user point at the display device at the same time. Inactive users, i.e., users not pointing at the screen may be shown without assigning a cursor, for example.
The multi-user collaboration is based on the concepts introduced in connection with Figs, la to Id. In particular, features, such as gesture control, may be used with the present application. For example, users may use gestures to do additional actions such as cast a vote, leave a flag (with their name or face), draw a shape, move an object etc.
More details and aspects of the concept for contactless collaboration are mentioned in connection with the proposed concept or one or more examples described above or below (e.g., Fig. la to 9b, 11). The concept for contactless collaboration may comprise one or more additional optional features corresponding to one or more aspects of the proposed concept, or one or more examples described above or below.
As used herein the term “and/or” includes any and all combinations of one or more of the associated listed elements and may be abbreviated as “/”.
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
Some embodiments relate to an imaging device, such as a microscope or exoscope, comprising a system as described in connection with one or more of the Figs, la to 10. Alternatively, an imaging device, such as a microscope or exoscope, may be part of or connected to a system as described in connection with one or more of the Figs, la to 10. Fig. 11 shows a schematic illustration of a system 1100 configured to perform a method described herein. The system 1100 comprises an imaging device 1110 (such as a microscope or exoscope) and a computer system 1120. The imaging device 1110 is configured to take images and is connected to the computer system 1120. The computer system 1120 is configured to execute at least a part of a method described herein. The computer system 1120 may be configured to execute a machine learning algorithm. The computer system 1120 and imaging device 1110 may be separate entities but can also be integrated together in one common housing. The computer system 1120 may be part of a central processing system of the imaging device 1110 and/or the computer system 1120 may be part of a subcomponent of the imaging device 1110, such as a sensor, an actor, a camera, or an illumination unit, etc. of the imaging device 1110. The computer system 1120 may be a local computer device (e.g., personal computer, laptop, tablet computer or mobile phone) with one or more processors and one or more storage devices or may be a distributed computer system (e.g., a cloud computing system with one or more processors and one or more storage devices distributed at various locations, for example, at a local client and/or one or more remote server farms and/or data centers). The computer system 1120 may comprise any circuit or combination of circuits. In one embodiment, the computer system 1120 may include one or more processors which can be of any type. As used herein, processor may mean any type of computational circuit, such as but not limited to a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), multiple core processor, a field programmable gate array (FPGA), for example, of a microscope or a microscope component (e.g. camera) or any other type of processor or processing circuit. Other types of circuits that may be included in the computer system 1120 may be a custom circuit, an application-specific integrated circuit (ASIC), or the like, such as, for example, one or more circuits (such as a communication circuit) for use in wireless devices like mobile telephones, tablet computers, laptop computers, two-way radios, and similar electronic systems. The computer system 1120 may include one or more storage devices, which may include one or more memory elements suitable to the particular application, such as a main memory in the form of random-access memory (RAM), one or more hard drives, and/or one or more drives that handle removable media such as compact disks (CD), flash memory cards, digital video disk (DVD), and the like. The computer system 1120 may also include a display device, one or more speakers, and a keyboard and/or controller, which can include a mouse, trackball, touch screen, voice-recognition device, or any other device that permits a system user to input information into and receive information from the computer system 1120.
Various examples of the present disclosure relate to a concept for a contactless 3D-user interface (UI), and in particular to improved or optimized contactless control via a 3D UI.
In touch-based interface, the following input modalities and affordances are used. For example, touch-screen based user interfaces typically utilize virtual control elements such as buttons and sliders as touch controls. Moreover, in computer interfaces, e.g., on webpages, it is common to use an animation, such as a hover effect or a rollover effect, on some elements, such as buttons, when the mouse pointer is over them. Such existing UI elements are designed for touchscreens or regular screens. However, the user's view is unnatural for contactless control, where the user’ s fingertip is at a non -negligible distance from the 3D display. This is due to the stereoscopic nature of human vision, as shown in Fig. 12a. Figs. 12a and 12b shows an illustration of a user’s stereoscopic perception of a user interface element, such as a button 140, in a three-dimensional user interface. In Figs. 12a and 12b, the respective upper image indicates what the screen shows, the respective middle image indicates what the user sees, and the respective lower image shows what the user perceives. Fig. 12a shows the perception of the user without, and Fig. 12b shows the perception of the user with application of the proposed concept. In the upper image of Fig. 12a, the position of the eyes 120a; 120b are shown, together with the position of the pointing indicator 130 (e.g., finger) and two imaginary lines 140a; 140b originating from the positions of the eyes 120a; 120b and terminating at two different points of the user interface. The user interface element 145 is between the two termini of the imaginary lines. As a result, the user perceives the user interface element 145 at two different positions 145c; 145d, depending on the eye being used. This results in the perception of a virtual distance 310 between the position of the pointing element and the position of the user interface element 145. The stereoscopic vision prevents the user from seeing the fingertip overlaying a virtual button as if they were physically at the same distance. On the contrary, the user sees either two buttons 145a; 145d or two fingertips, depending on where the user is focusing. As a result, some users tend to resolve this perception conflict by closing one eye. Even if the previous limitation was not an issue, the interaction with a virtual button is not natural, as the button does not appear as it could be physically reached and pushed by the fingertip.
A major insight of the proposed disclosure is to visualize the virtual control elements, such as button, sliders, and rotating knobs, at a distance close to the finger and thus make the user experience more natural. Fig. 12b shows an illustration of the desired effect. In Fig. 12b, the user interface element 145a; 145b is shown at two different positions in the user interface. As a result, as shown in the middle image, the perceived position 145c; 145d is the same for both eyes, such that the user interface element 145 appears to be at the position of the pointing indicator 130 (shown in the lower image of Fig. lb). As a result, there is a virtual distance 320 between the user interface element and the baseline of the three-dimensional user interface (indicated by the drop shadow). This is enabled by the ability of 3D displays to easily visualize virtual buttons (or other 3D control elements) at different distances between the screen and the observer. The desired user experience is that when the user’s fingertip approaches the virtual button, the button’s disparity is increased so that the button appears to be close to the fingertip, as if the button promptly lands at reach distance to be pressed. In addition to the natural and intuitive experience for the user, this animated movement of the button serves as an unmistakable confirmation that the specific control is active. Moreover, the alignment of the fingertip and the button does not cause any stereoscopic confusion and neither the fingertip nor the button appear as “stereoscopic ghosted” (see middle image in Figs. 12a and 12b).
In the previous discussion, the user interface element was assumed to be a button. However, the virtual control element (i.e., the user interface element) may include a variety of types such as rotating knobs (see Fig. 13), sliders, buttons, switches etc. Fig. 13 shows use of a rotating button as user interface element 145, having a virtual distance 1310 to the baseline of the user interface. A rotation button can be a very intuitive way for adjustments.
The Hover effect animation of the button moving closer to the user’s hand can help the user to shape the desired perception of the virtual button being located within reach distance from the hand. Stereoscopic/3D displays suffer from limitations that limit the ability to trick the brain into believing that the displayed object is three-dimensional. For example, although the right-left eye disparity can be easily achieved, the focus distance of the human eye informs the brain that the displayed object is not at the virtual distance.
Fig. 14 shows an example of annotating the determined position in the user interface based on the determined utterance. As described with reference to Fig. 14, the system as described above, e.g., with reference to Fig. 1, may be configured to obtain audio sensor data of a microphone, determine an utterance of the user and adjust the user interface by annotating the determined position in the user interface based on the determined utterance 1420, e.g., the second clip should go here. That is, the system may be configured to add an annotation 1410 to the user interface to generate the adjusted user interface. For example, the stereo camera 170 may be configured as microphone. That is, the stereo camera 170 may be configured to acquire the imaging sensor data and the audio sensor data. Optionally or alternatively, a microphone, e.g. part of a head-mounted display may be used to acquire the imaging sensor data. Thus, the system may provide a contactless method allowing the surgeon to annotate the image in a contactless manner, directly by looking on the microscope screen. In an example, the system may be configured to highlight the determined position in the user interface. For example, as shown in Fig. 14, the system may be configured to draw a line according to a movement of the finger of the user (see also the description above). In this way, the user can indicate the determined position, e.g., to indicate the position of a surgical tool, such like a clip. That is, the system can be used to mark an area in the user interface, e.g., a live view of a sample, and to provide an annotation 1410 to the marked area. For example, another surgeon may assist the main surgeon from outside an operation room by providing the information about the marked area. For example, the system may offer recording of the surgeon’s movement to provide a more immersive impression, instead of just the drawing a line. Optionally or alternatively, in parallel a record of a voice message which will be attached to the handwritten annotation can be provided. That is, the user interface can be adjusted by highlighting the determined position, e.g., by the drawn line, and/or by the annotation 1410.
The annotation 1410 could include markers, which could also be grouped in categories, for example, anatomical: vein, artery, nerve, bone, pathological such like high grade tumor, mow grade tumor, healthy, suspicious, unknown. The annotation 1410 can be done on a still image or video, live or recorded. The annotation 1410 can, for example, comprise additional data and metadata, e.g. time, person (surgeon), imaging mode, patient number, microscope settings. Previous annotations can also be shown with little symbols informative for the annotation, e.g. tumor, surgeon X, timestamp, location, voice message.
For example, the system may be configured to perform voice transcription for analyzing a message, e.g., part of the audio sensor data, and optionally may automatically generate the annotation 1410.
The surgeon’s annotation 1410 can also be done on image guided surgery visualization and shown as part of the image guided surgery navigation. For example, preoperative or intraoperative the surgeon could make annotations about the planned surgical cavity, and this annotation 1410 could be shown as part of the image guided surgery visualization on the microscope.
Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a non- transitory storage medium such as a digital storage medium, for example a floppy disc, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier.
Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
In other words, an embodiment of the present invention is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
A further embodiment of the present invention is, therefore, a storage medium (or a data carrier, or a computer-readable medium) comprising, stored thereon, the computer program for performing one of the methods described herein when it is performed by a processor. The data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary. A further embodiment of the present invention is an apparatus as described herein comprising a processor and the storage medium. A further embodiment of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, for example, via the internet.
A further embodiment comprises a processing means, for example, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein.
A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device, or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
In some embodiments, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
Embodiments may be based on using a machine-learning model or machine-learning algorithm. Machine learning may refer to algorithms and statistical models that computer systems may use to perform a specific task without using explicit instructions, instead relying on models and inference. For example, in machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of historical and/or training data. For example, the content of images may be analyzed using a machinelearning model or using a machine-learning algorithm. In order for the machine-learning model to analyze the content of an image, the machine-learning model may be trained using training images as input and training content information as output. By training the machinelearning model with a large number of training images and/or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine-learning model "learns" to recognize the content of the images, so the content of images that are not included in the training data can be recognized using the machine-learning model. The same principle may be used for other kinds of sensor data as well: By training a machine-learning model using training sensor data and a desired output, the machine-learning model "learns" a transformation between the sensor data and the output, which can be used to provide an output based on non -training sensor data provided to the machine-learning model. The provided data (e.g., sensor data, meta data and/or image data) may be preprocessed to obtain a feature vector, which is used as input to the machine-learning model.
Machine-learning models may be trained using training input data. The examples specified above use a training method called "supervised learning". In supervised learning, the machine-learning model is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training. Apart from supervised learning, semi -supervised learning may be used. In semi -supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g., a classification algorithm, a regression algorithm or a similarity learning algorithm. Classification algorithms may be used when the outputs are restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values. Regression algorithms may be used when the outputs may have any numerical value (within a range). Similarity learning algorithms may be similar to both classification and regression algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are. Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data might be supplied, and an unsupervised learning algorithm may be used to find structure in the input data (e.g., by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data comprising a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters. Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such, that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
Furthermore, some techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
In some examples, anomaly detection (i.e., outlier detection) may be used, which is aimed at providing an identification of input values that raise suspicions by differing significantly from the majority of input or training data. In other words, the machine-learning model may at least partially be trained using anomaly detection, and/or the machine-learning algorithm may comprise an anomaly detection component.
In some examples, the machine-learning algorithm may use a decision tree as a predictive model. In other words, the machine-learning model may be based on a decision tree. In a decision tree, observations about an item (e.g., a set of input values) may be represented by the branches of the decision tree, and an output value corresponding to the item may be represented by the leaves of the decision tree. Decision trees may support both discrete values and continuous values as output values. If discrete values are used, the decision tree may be denoted a classification tree, if continuous values are used, the decision tree may be denoted a regression tree.
Association rules are a further technique that may be used in machine-learning algorithms. In other words, the machine-learning model may be based on one or more association rules. Association rules are created by identifying relationships between variables in large amounts of data. The machine-learning algorithm may identify and/or utilize one or more relational rules that represent the knowledge that is derived from the data. The rules may e.g., be used to store, manipulate, or apply the knowledge.
Machine-learning algorithms are usually based on a machine-learning model. In other words, the term "machine-learning algorithm" may denote a set of instructions that may be used to create, train, or use a machine-learning model. The term "machine-learning model" may denote a data structure and/or set of rules that represents the learned knowledge (e.g., based on the training performed by the machine-learning algorithm). In embodiments, the usage of a machine-learning algorithm may imply the usage of an underlying machine-learning model (or of a plurality of underlying machine-learning models). The usage of a machine-learning model may imply that the machine-learning model and/or the data structure/set of rules that is the machine-learning model is trained by a machine-learning algorithm.
For example, the machine-learning model may be an artificial neural network (ANN). ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain. ANNs comprise a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There are usually three types of nodes, input nodes that receiving input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node may represent an artificial neuron. Each edge may transmit information, from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g., of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and/or of edges may be adjusted in the learning process. In other words, the training of an artificial neural network may comprise adjusting the weights of the nodes and/or edges of the artificial neural network, i.e., to achieve a desired output for a given input.
Alternatively, the machine-learning model may be a support vector machine, a random forest model or a gradient boosting model. Support vector machines (i.e., support vector networks) are supervised learning models with associated learning algorithms that may be used to analyze data (e.g., in classification or regression analysis). Support vector machines may be trained by providing an input with a plurality of training input values that belong to one of two categories. The support vector machine may be trained to assign a new input value to one of the two categories. Alternatively, the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
List of reference Signs
10, lOa-d User
12 Eye
16 Hand/pointing indicator
16a Finger
16b Wrist
20 Base unit of surgical microscope system with haptic user interface
100 Screen, user interface 110 System
120 First point of reference, e.g., position on body, eye
130 Second point of reference, e.g., pointing indicator
140 Imaginary line
145 User interface element
150 Camera sensor
160 Depth sensor
170 Stereo sensor
180 Display device, screen
185 Projection device
190 Emitter
210 Determining a position of a first point of reference
220 Determining a position of a second point of reference
230 Determining at least one of a position in the user interface, an element of the user interface and a device
240 Determining a response
310 Screen
320 Camera
325 Field of view of camera
330 Hand
340 Cursor/pointer
410 Screen
420, 430 Stereo cameras
425, 425 Field of view of stereo cameras
440 Eyes
445 Point of view of user 450 Hand
460 Cursor/pointer
510, 520, 535 Shadow
530 Hand
540 Highlight
610 Virtual button
620 Button activation zone
630 Button stand-by zone
640 Microscope with ultrasound emitters
640a, 640b Virtual boundaries for microscope
650a Haptic confirmation that the device is engaged
650b, 650c Haptic warning that the boundary is reached
660a Haptic confirmation that the device is engaged
660b, 660c Haptic warning that the boundary is reached
710, 720, 730 Device
710a, 720a, 730a Detection area
725 Indicator light
810 Click gesture
820 Grab gesture
830 Turn gesture
840 Scroll gesture
850 Rotate gesture
910 Vital s monitor
915 Information
920 Screen of surgical microscope system
925 Information
930 Head-mounted display
101 Oa-d Indicator representing a user
1020a-d Positional indicator
1030 Visual link
1100 System
1110 Imaging device, microscope
1120 Computer system
1210 virtual distance 1220 virtual distance
1310 virtual distance
1410 annotation
1420 utterance of the user

Claims

Claims
1. A system (110) for adjusting a user interface (100), the system comprising one or more processors (114), wherein the system is configured to: determine a position of a first point of reference (120) based on a position of a first point (12) on the body of a user (10); determine a position of a second point of reference (130) based on a position of a pointing indicator (16) used by the user; determine a position in the user interface based on the first point of reference and the second point of reference; adjust the user interface based on the determined position in the user interface; generate a display signal comprising the adjusted user interface; and provide the display signal to a display device (180).
2. The system according to claim 1, wherein the system is configured to obtain imaging sensor data of a camera sensor (160, 170); determine the identity of the user based on the imaging sensor data; and adjust the user interface by providing the determined position in the user interface with a link to an identity of the user.
3. The system according to claim 2, wherein the system is configured to provide the link to the identity by generating a visual representation (210a-d; 220a-d; 230) of the user pointing towards the position in the user interface based on the determined identity of the user and based on the determined position in the user interface.
4. The system according to claim 1, wherein the system is configured to generate the visual representation of the user pointing towards the position in the user interface with a positional indicator (220a-d) highlighting the determined position in the user interface, an indicator (210a-d) representing the identity of the user and a visual link between the positional indicator and the indicator representing the identity of the user.
5. The system according to claim 2, wherein the indicator representing the identity of the user comprises at least one of a textual representation of the identity of the user and image or an avatar of the user.
6. The system according to any one of the preceding claims, wherein the system is configured to adjust the user interface by controlling the user interface to animate a transition of the display position of the selected user interface element from an initial display position of the selected user interface element to the display position at the predefined distance from the pointing indicator.
7. The system according any one of the preceding claims, wherein the system is configured to adjust the user interface by controlling the user interface to change an opacity of the selected user interface element while the user interface element is shown at the changed display position
8. The system according to any one of the preceding claims, wherein the system is configured to adjust the user interface by controlling the user interface to change a shape of the selected user interface element while the user interface element is shown at the changed display position.
9. The system according to any one of the preceding claims, wherein the system is configured to adjust the user interface by controlling the user interface to sub-divide or extend the selected user interface element into a group of two or more user interface elements, and to control the user interface to set the display position of the group of two or more user interface elements such that the group of two or more user interface elements appears to be positioned at the pre-defined distance from the pointing indicator.
10. The system according to one of the claims 1 to 9, wherein the system is configured to process at least one of depth sensor data of a depth sensor (160), imaging sensor data of a camera sensor (150) and stereo imaging sensor data of a stereo camera (170) to determine at least one of the position of the first point of reference and the position of the second point of reference, and an identity of the user.
11. The system according to claim 10, wherein the system is configured to determine the positions of joints of a skeletal model of the body of the user by processing the respective sensor data, and to determine at least one of the position of the first point of reference and the position of the second point of reference based on the positions of the joints of the skeletal model.
12. The system according to any one of the preceding claims, wherein the system is configured to system is configured to: obtain audio sensor data of a microphone; determine an utterance of the user; and adjust the user interface by annotating the determined position in the user interface based on the determined utterance.
13. The system according to any one of the preceding claims, wherein the system is configured to highlight the determined position in the user interface.
14. A method for controlling a user interface, the method comprising: determining (210) a position of a first point of reference based on a position of a first point on the body of a user; determining (220) a position of a second point of reference based on a position of a pointing indicator used by the user; determining (230) a position in the user interface based on the first point of reference and the second point of reference; adjusting (240) the user interface based on the determined position in the user interface; generating (250) a display signal comprising the adjusted user interface; and providing (260) the display signal to a display device.
15. Computer program with a program code for performing the method according to claim
14 when the computer program is run on a processor.
EP24728210.6A 2023-05-24 2024-05-22 System, method, and computer program for controlling a user interface Pending EP4720815A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
EP23175149.6A EP4468121A1 (en) 2023-05-24 2023-05-24 System, method, and computer program for controlling a user interface
EP23175130.6A EP4468118A1 (en) 2023-05-24 2023-05-24 System, method, and computer program for controlling a user interface
EP23175126.4A EP4468117A1 (en) 2023-05-24 2023-05-24 System, method, and computer program for controlling a user interface
PCT/EP2024/064068 WO2024240811A1 (en) 2023-05-24 2024-05-22 System, method, and computer program for controlling a user interface

Publications (1)

Publication Number Publication Date
EP4720815A1 true EP4720815A1 (en) 2026-04-08

Family

ID=91247390

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24728210.6A Pending EP4720815A1 (en) 2023-05-24 2024-05-22 System, method, and computer program for controlling a user interface

Country Status (3)

Country Link
EP (1) EP4720815A1 (en)
CN (1) CN121175644A (en)
WO (1) WO2024240811A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP4900741B2 (en) * 2010-01-29 2012-03-21 島根県 Image recognition apparatus, operation determination method, and program
KR102517425B1 (en) * 2013-06-27 2023-03-31 아이사이트 모빌 테크놀로지 엘티디 Systems and methods of direct pointing detection for interaction with a digital device
US10203762B2 (en) * 2014-03-11 2019-02-12 Magic Leap, Inc. Methods and systems for creating virtual and augmented reality

Also Published As

Publication number Publication date
CN121175644A (en) 2025-12-19
WO2024240811A1 (en) 2024-11-28

Similar Documents

Publication Publication Date Title
JP7791931B2 (en) Occlusion cursor for virtual content in mixed reality displays
JP7411133B2 (en) Keyboards for virtual reality display systems, augmented reality display systems, and mixed reality display systems
US20250349089A1 (en) Transmodal input fusion for a wearable system
Chen et al. Recent developments and future challenges in medical mixed reality
US9939914B2 (en) System and method for combining three-dimensional tracking with a three-dimensional display for a user interface
CN117178247A (en) Gestures for animating and controlling virtual and graphical elements
US20200005026A1 (en) Gesture-based casting and manipulation of virtual content in artificial-reality environments
US10896545B1 (en) Near eye display interface for artificial reality applications
KR20250056250A (en) Physical gestural interaction with objects based on intuitive design
KR20250056992A (en) Scissors hand gesture for collaboration objects
CN119968655A (en) Real-world responsiveness of collaborative objects
Hernoux et al. A seamless solution for 3D real-time interaction: design and evaluation
KR20250056993A (en) Collaboration objects associated with geographic locations
KR20250053949A (en) Optional collaboration object access
CN119816802A (en) Selective collaborative object access
CN120380507A (en) Validating selective collaboration objects
WO2024240812A1 (en) System, method, and computer program for controlling a user interface
EP4468117A1 (en) System, method, and computer program for controlling a user interface
EP4468118A1 (en) System, method, and computer program for controlling a user interface
EP4720815A1 (en) System, method, and computer program for controlling a user interface
EP4468119A1 (en) System, method, and computer program for controlling a user interface
EP4468120A1 (en) System, method, and computer program for controlling a user interface
EP4468121A1 (en) System, method, and computer program for controlling a user interface
Žagar et al. Contactless Interface for Navigation in Medical Imaging Systems
Rossol Novel Methods for Robust Real-time Hand Gesture Interfaces

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20260102

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR