WO2025257529A1 - Mixed reality system for spatial anatomy visualization - Google Patents
Mixed reality system for spatial anatomy visualizationInfo
- Publication number
- WO2025257529A1 WO2025257529A1 PCT/GB2025/051219 GB2025051219W WO2025257529A1 WO 2025257529 A1 WO2025257529 A1 WO 2025257529A1 GB 2025051219 W GB2025051219 W GB 2025051219W WO 2025257529 A1 WO2025257529 A1 WO 2025257529A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- subject
- point cloud
- registration
- model
- data set
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/30—Determination of transform parameters for the alignment of images, i.e. image registration
- G06T7/33—Determination of transform parameters for the alignment of images, i.e. image registration using feature-based methods
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B34/00—Computer-aided surgery; Manipulators or robots specially adapted for use in surgery
- A61B34/25—User interfaces for surgical systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B34/00—Computer-aided surgery; Manipulators or robots specially adapted for use in surgery
- A61B34/10—Computer-aided planning, simulation or modelling of surgical operations
- A61B2034/101—Computer-aided simulation of surgical operations
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B34/00—Computer-aided surgery; Manipulators or robots specially adapted for use in surgery
- A61B34/10—Computer-aided planning, simulation or modelling of surgical operations
- A61B2034/101—Computer-aided simulation of surgical operations
- A61B2034/105—Modelling of the patient, e.g. for ligaments or bones
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B34/00—Computer-aided surgery; Manipulators or robots specially adapted for use in surgery
- A61B34/10—Computer-aided planning, simulation or modelling of surgical operations
- A61B2034/107—Visualisation of planned trajectories or target regions
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B34/00—Computer-aided surgery; Manipulators or robots specially adapted for use in surgery
- A61B34/20—Surgical navigation systems; Devices for tracking or guiding surgical instruments, e.g. for frameless stereotaxis
- A61B2034/2072—Reference field transducer attached to an instrument or patient
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B90/00—Instruments, implements or accessories specially adapted for surgery or diagnosis and not covered by any of the groups A61B1/00 - A61B50/00, e.g. for luxation treatment or for protecting wound edges
- A61B90/36—Image-producing devices or illumination devices not otherwise provided for
- A61B2090/364—Correlation of different images or relation of image positions in respect to the body
- A61B2090/365—Correlation of different images or relation of image positions in respect to the body augmented reality, i.e. correlating a live optical image with another image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10004—Still image; Photographic image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10028—Range image; Depth image; 3D point clouds
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10072—Tomographic images
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30004—Biomedical image processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30004—Biomedical image processing
- G06T2207/30096—Tumor; Lesion
Definitions
- the present disclosure generally relates to spatial anatomy recognition and guidance through the use of spatial anatomy imaging technology.
- the present disclosure relates to marker-less imaging for providing spatial anatomy models to locate all non-visible target body segments in the spatial environment.
- Two-dimensional (2D) medical visualization techniques are often insufficient for displaying complex, three-dimensional (3D) anatomical structures.
- the visualization of medical data on a 2D screen during surgery is undesirable, because it requires a surgeon to continuously switch focus.
- augmented reality AR
- placing markers for a precise holographic overlay are costly, always have to be visible within the field of view, have the potential to increase the regulatory burden, and can disrupt the surgical workflow.
- the present disclosure provides methods and systems for marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject.
- This will provide a displayed 3D virtual model of the subject to the user, the 3D virtual model being accurate relative to the subject's anatomical composition.
- the 3D virtual model will be provided to the user in real-time. This will allow the user to better understand the relative positioning of anatomical features within the subject e.g. bones, organs, arteries etc. such that the user can better understand the anatomy of the subject and thus, make more informed decisions during surgical and therapeutic procedures.
- the method of marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject may include: scanning the subject using a time-of-flight sensor to generate a point cloud data set corresponding to a skin layer of the subject, obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the skin layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and displaying the three-dimensional virtual model aligned within the live imagery of the subject.
- An aspect of the present disclosure presents a method of marker-less registration of a preobtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, comprising: scanning the subject using a depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject; obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the anatomical layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three- dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and displaying the three-dimensional virtual model so aligned within the live imagery of the subject.
- the camera and the depth sensor are substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween. Having the camera and the depth sensor substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween improves the success rate of co-registration between the point cloud data set and the 3D virtual model.
- the processing comprises inputting the point cloud data set and the three dimensional virtual model into a trained neural network trained to find matching anatomical features in the two data sets.
- a trained neural network can accurately and efficiently find matching anatomical features in the two data sets, improving the success rate of coregistration between the point cloud data set and the 3D virtual model. This can improve the accuracy of guidance during surgery or other medical procedures, and reduce the risks to patients.
- the aligning comprises determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features. Determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features allows for accurate alignment between the two data sets.
- the aligning further comprises determining the rotation matrix R and the translation matrix t by applying a weighted singular value decompositions (SVD) to the point cloud data set. More preferably, the applying the weighted singular value decompositions, SVD, to the point cloud data set is repeated by a predetermined number of iterations, n.
- SVD weighted singular value decompositions
- the scanning of the subject using the depth sensor places emphasis on segmenting the scan to highlight surface skin of the subject.
- the one or more anatomical features are one or more of skin, organs, tissues, bones, arteries, veins, as well as the surface geography of skin, organs, and tissue planes.
- the one or more anatomical features are one or more abnormal anatomical features, the one or more abnormal anatomical features being tumours, bone fractures, bone deformities, vascular pathologies, infections, as well as pathological lumps such as malignant tumours, skin lesions benign and malignant, bone anatomy.
- the method further comprises repeating the above steps and updating the display to account for movement of the user about the subject, or movement of the anatomy of the subject, during the surgical or therapeutic procedure.
- This allows the display to track relative movement and ensure that co-registration between the point cloud data set and the corresponding anatomy in the live imagery is maintained.
- the present disclosure also provides a control system for controlling an imaging system, the imaging system comprising a depth sensor and a camera, the control system comprising one or more processors collectively configured to: scan the subject using the depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject, obtain live imagery of the subject via the camera; receive the pre-obtained medical imagery three-dimensional virtual model of a subject; process the anatomical layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; align the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in coregistration with each other; and display the three-dimensional virtual model so aligned within the live imagery of the subject.
- Figure 1 shows an example flow diagram of the hololens registration application pipeline, in accordance with the present disclosure
- Figure 2 shows a further example flow diagram of the marker-less hologram-guided registration method, in accordance with the present disclosure
- Figure 3 shows example software images of different layers of the human anatomy produced by a 3D Pre-Operative Hologram
- Figure 4 shows example software images showing different views of 3D Point Cloud Extraction using Intel Depth Camera D435
- Figure 5 shows example software images showing different views of 3D Point Cloud Extraction using Polycam mobile applications
- Figure 6 shows example software images showing different view of 3D Point Cloud Extraction using Hololens spatial mapping
- Figure 7 shows example software images showing different view of 3D Point Cloud Extraction using Hololens Depth Camera
- Figure 8 shows an example flow diagram for the data simulation pipeline for use in producing improved target point clouds
- Figure 9 shows example images of different configurations of target point cloud after applying different rotation and translation to the source point cloud
- Figure 10 shows example images of different configurations of target point cloud after occluding random areas
- Figure 11 shows example images of different configurations of target point cloud after adding a plane at the truncated region
- Figure 12 shows an example system diagram of a deep learning-based registration model, namely RPMNet, in accordance with the present disclosure
- Figure 13 shows an example system diagram for a feature extraction part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure
- Figure 14 shows an example system diagram for a parameter prediction part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure
- Figure 15 shows an example system diagram for a matrix part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure
- Figure 16 shows an example system diagram of an encoder-attention-decoder model, namely PREDATOR
- Figure 17 shows example image sample sets from the ModelNet 40 dataset
- Figure 18 shows a visualization of the registration method applied to a sample 3D model from ModelNet40 dataset for clean target point cloud, in accordance with the present disclosure
- Figure 19 shows an example flow diagram of the training method for the registration model, in accordance with the present disclosure
- Figure 20 shows an example flow diagram of the testing method of the registration model, in accordance with the present disclosure
- Figure 21 shows an example graphical user interface for the training and testing of the registration model, in accordance with the present disclosure
- Figure 22 shows a comparison of images of the registration performed using the generic registration model (on the left) versus the fine-tuned registration model (on the right), in accordance with the present disclosure
- Figure 23 shows a comparison of further images of the registration performed using the generic registration model (on the left) versus the fine-tuned registration model (on the right), in accordance with the present disclosure
- Figure 24 shows a visualization of the registration applied to data simulated using SPL/NAC Brain Atlas [Opeteb];
- Figure 25 shows a visualization of the registration applied to data simulated using SPL Head and Neck Atlas [Opetea];
- Figure 26 shows a block diagram of a computer system for use with the 3D marker-less hologram-guided registration method, in accordance with the present disclosure
- Figure 27 shows a flow diagram of the operation of the automatic registration process using deep learning model, in accordance with the present disclosure.
- the present disclosure seeks to provide methods and systems which may use 3D scene scanning technology to generate a point cloud of a subject's anatomy, and in particular in some embodiments a skin layer of the anatomy, the subject being a patient about to undergo some sort of surgical or therapeutic procedure.
- Pre-obtained medical imagery e.g. CT/MR, 3D ultrasound etc.
- 3D hologram model of the subject, which may then be registered with the point cloud so as to overlay virtually the 3D hologram model with the point cloud of the subject's anatomy in a user display.
- the marker-less registration may be performed by using neural networks to extract predefined corresponding anatomy features in each of the anatomical model and the 3D hologram model which then provide anchor points for a matching matrix transform to map the skin point cloud to the 3D hologram model across the two data sets.
- the marker-less registration and overlay may be dynamically updated as the user moves and his field of view of the subject changes or the subject position changes in the spatial environment.
- the result is an augmented view of the subject with the 3D medical imagery hologram model virtually overlaid onto the subject within the user's field of view.
- the user display is head mounted on the user, although in other embodiments the display can be on a screen such as a monitor or the like, or provided for robotics and automation purposes.
- audio sensory cues may be provided in addition to or as an alternative to the augmented view. For example, if a user's field of view strays from the intended area of anatomy, pulsing sounds may be emitted that are faster or slower depening on how far away from the intended area of anatomy the user view has strayed.
- the augmented view of the subject enables the user to receive real-time visual representation of the subject in the form of a 3D medical imagery hologram model. This enables the user to make more informed decisions during the current surgical or therapeutic procedure etc. Further, it enables procedures to be executed in a more accurate and precise manner which should reduce the time taken to complete such procedures.
- the 3D medical imagery hologram model could be used for surgical training or patient education such that both students or patients can see where the steps of a surgical procedure would occur within the subjects body in a non- invasive manner, for example.
- the present disclosure seeks to provide a method that uses a deep learning algorithm and trained Al to register the Mesh generated by depth cameras to the skin segmentation topography. Therefore, the technology locates all non-visible target body segments in the spatial environment for positional and interventional adjustment automation.
- the present disclosure may relate to a method of marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, the method may comprise: scanning the subject using a time-of-flight sensor to generate a point cloud data set corresponding to the skin layer of the subject, obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the skin layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in coregistration with each other; and displaying the three-dimensional virtual model aligned within the live imagery of the subject.
- CT/MRI scans This may be achieved through post-acquisition processing of medical images (CT/MRI scans) converting the typical 2D images into 3D segments with two specific components.
- This may provide the internal position of the part of body targeted by the intervention in relation to skin (the spatial position of the target anatomy is constant to the surrounding skin with any movement on X or Y axis and any global rotational movement.
- the deformation of the human anatomy due to respiration movement and joint movement should be considered for accurate anatomy localisation).
- This may provide the second input which depends on generating a point cloud mesh of the patient body at the time of intervention. This may be achieved by using time of flight (ToF) camera or other 3D depth cameras.
- ToF time of flight
- the main camera may be the Azure Kinect DK depth camera which implements the Amplitude Modulated Continuous Wave (AMCW) Time-of-Flight (ToF) principle. This may provide the real-time [3D] body surface scan image.
- AMCW Amplitude Modulated Continuous Wave
- TOF Time-of-Flight
- the camera may be integrated into the Hololens® device and allows for seamless integration of surgical guidance using the depth camera and display on the HoloLens.
- other cameras such as an Intel® LiDAR camera or a standalone Azure Kinect® DK depth camera may be used, and it follows the same process in the integration of the technology pipeline. Both can be used for the other clinical application in interventional radiology and radiotherapy.
- Other depth capturing imaging devices are of course available that are the equivalent of those mentioned above, any and all of which can be used in embodiments of the present disclosure.
- the Al model seeks to process the thousands/millions of geometrical points in both point clouds and produce the necessary calculation to produce an output instruction to register both inputs together through point matching. It may be referred to as the Spatial Anatomy Recognition Artificial Intelligence (SARAI) model.
- SARAI Spatial Anatomy Recognition Artificial Intelligence
- Point matching is the process of finding corresponding points between two or more sets of points. It is a technique used in computer vision, particularly in fields such as robotics, image processing, and 3D modelling.
- Point matching aids in aligning the two sets of points to create a 3D anatomy that represents the patient.
- the proposed model may use deep neural network to learn point features and can be trained to improve its precision and standardized when achieving the needed results.
- This may be further used in the tracking of any changes in patient position during the intervention and rematch the 3D anatomy to the patient.
- tissue distortion e.g. lymph node biopsy, soft tissues in extremity, limb, vascular and tumour surgeries; including both primary and metastatic lesions.
- tissue distortion e.g. lymph node biopsy, soft tissues in extremity, limb, vascular and tumour surgeries; including both primary and metastatic lesions.
- having a fixed point when there is tissue distortion is important for treatment of lymph nodes, as it is all deformable.
- the potential use cases may be, but are not limited to:
- 2- In radiotherapy it may be applied in reverse by validating and calculating the needed movement of the patient body (Realtime body surface scanning) to the radiotherapy field to ensure the non-visible target anatomy (anatomical segmentation) is in the correct position.
- the present disclosure provides a Mixed Reality solution.
- Mixed Reality is the merging of real and virtual worlds to produce new environments where both physical and digital objects co-exist and interact in real time such as the use of holograms and mixed reality glasses.
- Augmented Reality is the result of using technology to superimpose digital elements (such as sounds, images, and text) to the world we see using a tablet, smart eyeglass or a smartphone camera.
- Virtual Reality is an immersive, three-dimensional, computer-generated environment which can be explored and interacted with by a person. Instead of projecting the images and sounds on a real environment, virtual reality uses a headset to immerse users within a 360-degree environment.
- FIG. 1 An example system diagram 1 is presented in Figure 1.
- the example system diagram relates to a pipeline for 3D imaging technology; specifically, the system diagram relates to marker-less mixed reality visualization for surgical methods.
- the system diagram 1 may also be referred to as 3D medical imagery hologram model 1.
- the 3D medical imagery hologram model 1 comprises the following: pre-operative images 10, a time-of-flight sensor (such as, but not limited to, a hololens) 11, a 3D hologram 12 and a server 13.
- the time-of-flight sensor 11 enables 3D scene reconstruction 110 and hence point cloud generation of the scene.
- the server 13 enables processing of the point cloud generated 130 and the use of a registration model 131.
- the 3D medical imagery hologram model 1 initially constructs a 3D hologram model 12 of a subject by using the subjects existing image data 10.
- the existing image data 10 may be a CT scan, or any other applicable image type.
- the existing image data 10 is used to develop a simulate training set for use in training a deep learning model for the present subject.
- the user will utilize the time-of-flight sensor 11 e.g. wear the hololens, during surgery.
- a point cloud of the subject is generated 130 within the server 13 using a depth camera with the time-of-flight sensor 11.
- the point cloud is then sent to the server 13 along with the 3D hologram 12.
- Both the subject's point cloud and the 3D hologram 12 are preprocessed before being utilized as input data to a registration model 131.
- the registration model 131 performs the registration and outputs a rotation and translation matrix 132.
- the rotation and translation matrix 132 is applied to the 3D hologram 12 moving it in space to mirror the position of the patient scanned point cloud.
- a method or flow diagram of operation of the 3D medical imagery hologram model 1 is shown in the flow diagram 2 of Figure 2.
- the flow diagram 2 is split into three distinct parts; before usage of the system or model 20, during initialization of the system or model 21 and during usage of the system or model 22.
- the before usage part 20 relates to the patient or subject's pre-existing image data. This can be image data such as CT scans that have been taken prior to use of the system or model.
- the pre-existing image data is then processed by segmentation to produce 3D models of internal anatomy and 3D models of the surface. These produced 3D models are to be used in the initialization of the system or model part 21. Initially, these 3D models are used to generate a 3D point cloud of the patient's anatomy (skin layer).
- the system or model is receiving a live scene with the subject or patient somewhere in the field of view.
- This live scene is likely provided by a time-of-flight sensor.
- the depth and position of the subject in the live scene is recorded. These depth and positional values are used in the 3D point cloud registration.
- the depth and positional recording of the live scene and the 3D point cloud of the patient's anatomy are used in conjunction to create a transformation of the 3D model to the current position of the patient.
- HL2 in Figure 2 refers to a Hololens 2 device as the time-of-flight sensor, however, any other applicable time-of-flight sensor may be used.
- the transformed 3D model which utilizes the live depth and position values from the time- of-flight sensor and the 3D point cloud from the pre existing image data, is provided to the user during usage of the system 22.
- the transformed 3D model is constantly supplied to the user during use and the relative position of the time-of-flight sensor at that moment is fed back into the transformed 3D model. This ensures that the transformed 3D model is always accurately mapped to the live scene which includes the subject or patient.
- the registration model 12 is applicable for use as the registration model 131 in Figure 1.
- the registration model 12 may be a deep learning-based registration model, such as, but not limited to, a RPMNet model.
- the registration model 12 contains six main parts: a rigid transform block 120, a parameter prediction block 121, a feature extraction block 122, a compute match matrix block 123, a match matrix 124 (which may be a rotation and translation matrix) and a weighted singular value decomposition block 125.
- the feature extraction block 122 is shown in more detail in Figure 13.
- the parameter prediction block 121 is shown in more detail in Figure 14.
- the compute match matrix 123 is shown in more detail in Figure 15.
- Figure 26 is a block diagram 26 of a typical general purpose computer system 260 that can form the processing platform for the 3D marker-less hologram-guided registration method, as described, in accordance with the present disclosure.
- the general-purpose computer system 260 may also be referred to as a control system 260 for controlling an imaging system.
- the computer system 260 comprises a central processing unit (CPU), random access memory (RAM), and input/output ports (I/O) into which data can be received and output therefrom as is well known in the art and a network I/C.
- CPU central processing unit
- RAM random access memory
- I/O input/output ports
- a camera unit 262 which may be positioned to record live surgical video, a time-of-flight (ToF) sensor 263 and a network access point 264 for sending or receiving data including the receiving of preoperative images, surgical videos etc.
- ToF time-of-flight
- network access point 264 for sending or receiving data including the receiving of preoperative images, surgical videos etc.
- the computer system 260 also includes some non-volatile storage, such as a hard disk drive, solid-state drive, or Non-Volatile Memory Express (NVMe) drive.
- nonvolatile storage 261 Stored on the nonvolatile storage 261 is a number of executable computer programs together with data and data structures required for their operation or training. Overall control of the system 26 is undertaken through the execution of the Point Cloud Generation Model 2611 and the Registration Model, in conjunction with pre-existing image data 2612 and live scene image data 2614.
- the other data contained in the non-volatile storage 261 is the training data 2615 which could include surgical or pharmaceutical images or videos or the like. In some examples, the training data 2615 may be images from the ModelNet40 dataset.
- FIG. 27 An example flow diagram 27 of the operation of the 3D marker-less hologram-guided registration method is shown in Figure 27.
- the flow starts at s270 where the time-of-flight sensor e.g. the Hololens scans the surface of the subject in its field of view. This is likely to be the surface of the patient before or during a surgery or therapeutic procedure.
- the time-of-flight sensor e.g. the Hololens
- a point cloud of the subject surface is generated from the scan of the subject surface taken in s270.
- the generated point could from s271 is sent to a registration model.
- the generated point cloud of the subject surface becomes the first input to the registration model.
- a further point cloud is produced from surfaces extracted from patient preoperative images e.g. from existing subject CT scans etc.
- This produced further point cloud is sent to the registration model as a second input.
- the two flows combine as the registration model receives the first and second inputs, namely the point cloud from the scanned subject surface and the point cloud from the subject preoperative images.
- the registration model calculates a translation and transformation matrix which highlights the relative positional differences between the first and second input. This translation and transformation matrix is to be used to perform co-registration of the first and second inputs. Therefore, in s274 the translation and transformation matrix is outputted.
- the translation and transformation matrix are used to map a generated hologram (generated from the preoperative imaging) onto the real scene subject position e.g., the ToF sensor scene. This achieves co-registration.
- the operation of the 3D marker-less hologram-guided registration method may comprise:
- the subject surface skin segment of the preoperative images will be used to generate a point cloud.
- the point cloud is transferred to a registration model as a first input;
- a ToF sensor or camera e.g. a Hololens, will reconstruct the 3D scene during the real-time intervention and a point cloud will be generated from the scene reconstruction;
- the generated point cloud from the scene reconstruction is transferred to the registration model as a second input;
- the registration model calculates the needed translation and transformation matrix to perform co-registration between the first input and the second input.
- the translation and transformation matrix being calculated from the relative positional differences between the first and second input.
- the translation and transformation matrix is then outputted;
- the translation and transformation matrix is applied to the hologram generated from the pre-operative imaging.
- the translation and transformation matrix causes the hologram to be moved to match the real scene subject position. Thus, achieving co-registration.
- the flow begins by scanning the subject using a time-of-flight sensor (e.g. a Hololens) to generate a point cloud that corresponds to the skin layer of a subject.
- a time-of-flight sensor e.g. a Hololens
- live imagery of the subject will be obtained from the camera.
- the pre-obtained medical imagery 3D virtual model e.g. CT scans
- the skin layer point cloud data set and the 3D virtual model skin layer will be processed to determine anatomical features of the subject which are present in both the point cloud data set and the 3D virtual model.
- the processing may entail inputting the point cloud data set and the 3D virtual model into a trained neural network trained to find matching anatomical features in the two data sets.
- the point cloud data set and the 3D virtual model will be co-registered, which may be achieved by a registration model.
- the coregistration requires the alignment of the point cloud data set and the 3D virtual model, which is achieved by aligning the determined anatomical features of the subject which were present in both the point cloud data and the 3D virtual model.
- This alignment may be achieved by determining a rotation matrix R and a translation matrix t that aligns the skin layer point cloud data set with the 3D virtual model in dependence on the determined anatomical features.
- the 3D virtual model which is aligned with the live imagery of the subject is displayed. This may be displayed on a head
- -mounted device such as the Hololens, or other forms of screens.
- the camera and the time-of-flight sensor may be substantially aligned. This may enable them to have the same or similar field of view of the subject or have a known field of view offset therebetween. This will likely improve the success rate of coregistration between the point cloud data set and the 3D virtual model.
- the one or more anatomical features of the subject may relate to anatomical features or abnormal anatomical features.
- the anatomical features may include one or more of organs, tissues, bones, arteries, veins or any other anatomical features.
- the abnormal anatomical features may include one or more of tumours, bone fractures, bone breaks, infections, or any other abnormal anatomical features.
- Visualization tools are crucial in aiding surgeons in understanding human anatomy.
- most of these tools rely on 2D displays, which present challenges in transferring knowledge from a 2D image to a 3D patient. Switching between the screen and the patient can be inconvenient and hinder surgical accuracy and safety [AHS+21].
- Mixed Reality devices like Hololens can be used to address this issue by registering 3D models onto the patient, allowing surgeons to achieve a better field of vision during surgery and improving surgical precision and safety [zHbFwS+19].
- One more advantage of using Hololens for these applications is the shared experience, where the same MR visualization can be shared across multiple Hololens allowing multiple people simultaneous access to the information.
- This project aims to study deep learning-based models using data extracted from Hololens.
- [ILH+17] demonstrates how to run deep learning models on Hololens by connecting it to a server for request processing.
- [UBG+20] provides several examples of extracting data from different sensors on Hololens for processing.
- [LDZS18] evaluates that Hololens is capable of estimating the user's head posture accurately at low movement speeds and reconstructing the environment with high precision, particularly for flat surfaces under bright conditions.
- Hololens 2 offers several capabilities under Hololens Research mode that can be utilized for this project, such as:
- IMU sensors like the accelerometer, gyroscope, and magnetometer, which can be utilized to track the location of the headset with respect to the real world [Pol20].
- a depth camera which uses active infrared (IR) illumination to determine depth through phase-based time-of-flight.
- the camera can operate in two modes. The first mode enables high -fra me- rate (45 FPS) near-depth sensing, commonly used for hand tracking. The other mode is used for lower-frame rate (1-5 FPS) far-depth sensing, currently employed for spatial mapping [Pol20].
- the sensor streams which can either be processed or stored on the device or wirelessly transferred to another PC or the cloud for more computationally demanding tasks [Pol20] .
- Registration is the process of finding the transformation from a source point cloud to a target point cloud of the same object [LSW17].
- Manual Registration is a straightforward approach to registration where the user has to manipulate the 3D model in a way that aligns with the patient.
- Several papers have proposed various methods for executing manual registration.
- [PIL + 18] proposed a method that uses hand gestures and voice commands to execute registration. Different hand gestures for translation and rotation were used to move the 3D model until accurate alignment between the anatomical landmarks and skin was achieved. Voice commands were implemented to switch between translation and rotation modes. [MJC + 18] used a similar method for registration. [GGS + 18] attempted to manually position the scapula in a holographic mode so that the surgeon could visualize hidden parts of the scapula during surgery.
- Automated 3D registration generally includes two steps: generating a 3D point cloud for the patient and registering the 3D point cloud of the segmented preoperative CT scan's 3D triangular surface model.
- This paper proposes manually sampling points of importance to generate a 3D point cloud for the patient using a custom-made pointing device (PD).
- the PD consists of a notch, a handle, and a tip.
- the notch is equipped with a marker of known geometry, which is used for tracking.
- the PD is moved along the patient to sample the 3D points of interest. Once sufficient sample points have been collected, the ICP algorithm is used to register the 3D model onto the patient by calculating correspondences between the two 3D point clouds.
- [NCR + 20] compared three different methods of manual registration: tap-to-place, 3-point correspondence matching, and keyboard control.
- Tap to place involved using an air tap as a mouse click to place the 3D model on the patient.
- 3-point correspondence matching involved choosing two sets of 3 points from the 3D model and the patient, followed by finding the translation and rotation matrix.
- Keyboard control involved using a Bluetooth keyboard to place the 3D model. The results of the experiments concluded that the keyboard method was the most accurate, followed by the tap-to-place method, and finally the 3-point correspondence method.
- Marker based registration To address the time consumption of the manual method, marker-based methods were introduced. Hololens offers access to recorded videos and images from the front-facing camera, along with the location of the camera in the real world and the lens model of the camera. This information can then be used to identify markers using computer vision techniques and locate them using the camera coordinates. Several papers discuss the use of marker-based registration.
- [FJDV18] used Vuforia's feature detection algorithm to extract features from the input image from the camera and compare them to stored features for registration. The tracking was further enhanced by placing a known RGB cylindrical object in the environment. The paper also assumed prior knowledge of the transformation between the hologram and phantom during manual registration, using this information when performing automated registration using Vuforia.
- [MMGMGS + 18] suggested a method to create 3D printer-generated patient-specific tools that are attached with markers. The location of the marker and the known dimensions of the patient-specific tool help in the registration process.
- [AJU + 18] introduces multimodality markers that can be detected using X-ray and RGB imaging devices.
- the key advantage of this method is that because the marker can be detected in X-ray, it can become part of the 3D model created using the segmented preoperative CT scan, eliminating the need to add it externally, which can be a source of error.
- [GCJ + 20] followed a similar approach, using a series of adhesive optical codes on the patient's skin prior to the CT scan. These optical codes are visible on CT and MRI scans, allowing for automatic registration with good accuracy that responds to patient movement.
- marker based registration interferes with the sterility required in surgical fields, and also requires lengthy processing to provide the registration.
- Alternative markerless techniques can thus be preferred in some scenarios.
- ICP Iterative Closest Point
- ICP is an iterative process where given two sets of points of the same object, we first find the correspondences between these two sets using a distance-based optimization function. Once we have the corresponding points in both sets, we use them to find the rotation matrix and translation vector. Then we use these to move the source points to their new locations. This process is continued until convergence.
- This method is quite prone to get stuck in local optima, and it assumes that the source and target point sets have all the same points. To overcome these issues, several improved versions of the algorithm have been devised.
- [ZSN03] introduces the Picky-ICP method, which offers a couple of improvements over the original ICP algorithm.
- Picky-ICP uses hierarchical point selection. We choose every 2 h -th data point and run ICP until convergence, where h+ 1 is the hierarchy level. After convergence, we move to the next hierarchy. This speeds up the computation time.
- the second improvement is identifying points present in more than one pair and selecting only one such pair, the one with the shortest distance. This makes the method more robust to noise and outliers.
- [HLY + 17] introduces GF-ICP, which utilizes the geometric features of the point clouds to be registered, such as curvature, surface normal, and point cloud density. It uses these features to search for the correspondence relationship between the two point clouds and introduces the geometric features into the error function to achieve accurate registration of the two point clouds.
- PointNet addresses this by employing a symmetric function that yields the same output regardless of the input order. Additionally, the paper emphasizes the combination of local and global features. Relying solely on local features across each point is insufficient as it fails to capture the global context. Therefore, the concatenation of both local and global features offers a better representation of the overall point cloud.
- PointNet treats each point independently, without exchanging information among neighbouring points, which is crucial for understanding the local geometry.
- PointNet++ (2017) [QYSG17] introduces a CNN- like model specifically designed for point clouds.
- a CNN nearby pixels are grouped together hierarchically to extract high-level features at each level.
- PointNet++ adopts a similar approach by grouping nearby point clouds using overlapping partitions. The process begins with Farthest Point Sampling (FPS) to select centroids, and then K-nearest neighbours (KNN) are used to create partitions around these centroids. Similar to CNN, these sets of points are fed into a feature extractor, which in this case is a PointNet. A shared PointNet is applied to each partition. This process of sampling and feature extraction is repeated iteratively to ultimately obtain a global representation.
- FPS Farthest Point Sampling
- KNN K-nearest neighbours
- PointNetLK (2019) [AGSL19] introduces a Deep Recurrent Neural Network that utilizes the features extracted by PointNet and incorporates the Lucas Kanade (LK) approach for calculating the transformation matrix.
- Lucas Kanade is a classical computer vision differential method used for estimating optical flow. It tackles the optical flow problem by minimizing the discrepancy between the transformed source image and the target image, which is analogous to the 3D registration problem.
- PointNetLK takes the source and target point clouds and applies MLP and PointNet to extract features from them.
- the Lucas Kanade method is then employed to determine the transformation that minimizes the distance between the source and target features. The entire process is iterative, with the model being applied in a loop until the change in the transformation matrix obtained from the LK method falls below a specified threshold.
- DCP or Deep Closest Point employs DGCNN (Dynamic Graph Convolutional Neural Network) instead of PointNet for feature extraction because PointNet does not consider local neighbourhood information, as discussed earlier, which is crucial for efficient feature matching, a critical step in DCP.
- DCP utilizes a transformer with cross-attention to facilitate information exchange between the source and target.
- These extracted features are then used to create soft mappings between the source and target. Soft mappings are preferred over hard mappings due to their differentiability.
- the final step involves employing the Singular Value Decomposition (SVD) of the obtained soft mappings to derive a closed-form solution for the optimization problem.
- Singular Value Decomposition Singular Value Decomposition
- ICP ICP
- DCP being a global algorithm that does not require initialization, generates a reasonably accurate solution, which can be further refined using ICP. In this case, ICP will converge towards the global optimum.
- PRNet or Partial Registration Network (2019) [WS19b] follows an iterative approach similar to ICP, where it first identifies corresponding points and then calculates the transformation matrix. This iterative process continues until convergence is achieved.
- PRNet incorporates an additional step of keypoint detection to address the issue of partial overlaps between the source and target. Keypoint detection helps filter out relevant points while eliminating outliers and noise.
- PRNet utilizes Gumbel- Softmax instead of the traditional softmax employed in DCP. The advantage of using Gumbel-Softmax is that it yields sharper mappings while maintaining differentiability.
- PRNet also introduces a temperature term in the Gumbel-Softmax, which can be adjusted to control the sharpness of the matches. Instead of treating the temperature term as a hyperparameter, it is learned as a model variable based on the alignment of two shapes.
- DeepVCP (2019) [LWZ + 19] operates on source and target point clouds, as well as a prior transformation. It employs PointNet++ to extract local geometric features from the source. These features undergo a weighting stage and are filtered to focus solely on key points.
- DeepVCP introduces a novel method called Corresponding Point Generation. This approach takes the source features and applies the prior transformation matrix to transform them. The transformed points are used to generate virtual corresponding points (VCPs) by creating 3D grid voxels around them. These VCPs are then utilized to compute the correspondence matrix, followed by applying SVD to obtain the transformation matrix.
- RPM or Robust Point Matching (1998) [GRL + 98] introduced the concept of generating soft assignments as a solution to handle outliers and noise.
- Equation above describes the method of creating a soft correspondence between the source point cloud ( ; ) and the target point cloud (y fe ).
- p is the temperature term that controls the softness of the assignment while a is used to deal with the outliers.
- RPMNet (2020) uses a similar approach but with a couple of changes [YL20].
- the equation replaces the coordinates with the extracted feature values, as shown in the above equation. Additionally, a secondary network is incorporated to learn the values of a and p.
- Another innovative approach introduced in this paper involves using not only the coordinates but also the relative coordinates of neighbouring points, along with 4D point pair features (PPF), as inputs to PointNet for feature extraction.
- PPF point pair features
- DeepGMR (2020) [YEK + 20] utilizes a probabilistic formulation, which enhances its resilience to noise and outliers.
- the primary objective of this work is to represent a 3D point cloud as a Gaussian Mixture Model (GMM) and subsequently match the distributions between the source and target models to obtain the transformation matrices.
- the method consists of three main components. Firstly, it involves determining point-to-component correspondences. This step establishes the associations between points and different components of the GMM. Next, these correspondences are utilized to obtain GMM parameters and transformation matrices.
- DeepGMR aims to leverage the probabilistic representation of point clouds using GMMs to facilitate robust matching and transformation estimation.
- RGMNet (2021) [FLLW21] employs a graph-based approach for feature extraction and subsequent transformation matrix estimation.
- the network utilizes a transformer-based edge generator, which utilizes the features of the point cloud to generate edges. This is followed by a graph feature extractor that operates independently on both the source and target graphs. The extracted features are then utilized to calculate the affinity between the two graphs, leading to the creation of a correspondence matrix.
- the updated values of the transformation matrix are obtained through the singular value decomposition (SVD) of the correspondence matrix.
- the graph-based methodology adopted by RGMNet is particularly effective in addressing outliers, enhancing the robustness of the registration process.
- Predator (2021) [HGLT21] employs an encoder-decoder model for point cloud registration.
- the encoder component downsamples both point clouds using KPConv-FPN, resulting in a set of super points.
- These super points are then utilized in a graph neural network (GNN) to extract features for both point clouds.
- the extracted features from the encoder are combined and fed into the decoder, which performs an upscaling operation to generate matching scores.
- This model is similar to RGMNet, with the key difference being the employment of downscaling and upscaling operations.
- Predator enhances the algorithm's robustness to noise and outliers, improving the accuracy of point cloud registration.
- GeoTransformer or Geometric Transformer (2022) [QYW + 22] follows a similar principle to Predator, leveraging KPConv-FPN to obtain a set of super points from the input point clouds.
- GNN graph neural network
- the GeoTransformer starts with a self-attention mechanism to capture local dependencies within each point cloud. Then, cross-attention is performed between the source and target point clouds to facilitate the transfer of information and establish point correspondences. These correspondences are subsequently used to compute the transformation matrix for each pair of matched superpoints. To determine the best global transformation, the method evaluates the performance of different transformations on a global scale and selects the one that achieves optimal results.
- OMNet (2021) [XLW + 21] addresses the challenge of partial overlaps in point cloud registration by employing a mask prediction approach. The method predicts overlapping masks that indicate the regions of overlap between the source and target point clouds. By utilizing these masks, non-overlapping points are discarded, and only the relevant points are considered for calculating the transformation matrix.
- a unique aspect of OMNet is its use of regression instead of singular value decomposition (SVD) to directly compute the transformation matrix. This regression-based approach provides an alternative method for estimating the transformation parameters.
- SVD singular value decomposition
- This regression-based approach provides an alternative method for estimating the transformation parameters.
- OMNet one limitation of OMNet is that it lacks a proper exchange of information between the source and target point clouds during the feature extraction process. This may result in inaccurate mask predictions and potentially impact the overall registration accuracy [XYL + 22].
- FINet (2022) [XYL + 22] introduces a dual branch structure for handling rotation and translation separately in the point cloud registration process. This separation is motivated by the fact that translation belongs to Euclidean space, which is not directly correlated with the quaternion space typically used for representing rotation. By employing separate branches for rotation and translation, FINet effectively handles both aspects of the transformation. Additionally, FINet incorporates a multi-level feature interaction mechanism between the source and target point clouds. This approach serves as an efficient alternative to traditional cross-attention mechanisms, reducing the computational and memory requirements significantly while still enabling effective information exchange between the two point clouds. These features of FINet contribute to its robustness and efficiency in point cloud registration tasks.
- YOHO (You only hypothesize once) (2022) [WLDW22] introduces the concept of rotation equivariant descriptors for point cloud registration.
- the primary objective of using a rotation equivariant descriptor is to capture variations in the descriptor corresponding to different rotations, which can be utilized to estimate the amount of rotation between point clouds.
- YOHO achieves this by defining feature maps on icosahedral groups with 60 rotations.
- the descriptor is designed to exhibit rotational symmetry and provide consistent representations for points under different rotations.
- average pooling is applied, enabling the extraction of features that are robust to rotation.
- the equivariant descriptor of these points can be utilized to estimate the rotation between the point clouds.
- YOHO leverages this rotation information to enhance the accuracy of point cloud registration. By introducing rotation equivariant descriptors, YOHO contributes to the advancement of rotation-aware feature extraction and improves the quality of point cloud registration results.
- GMCNet Graph Matching Consensus Network (2022) [PCL21] adopts a similar approach to other methods in terms of feature extraction, correspondence matrix creation, and transformation matrix estimation.
- the key contribution of the paper lies in its focus on rotation invariant feature extraction.
- GMCNet incorporates hand-crafted rotation invariant features into the feature extraction process.
- GMCNet integrates multiscale smoothness terms derived from the geometric structures at various scales. These smoothness terms help to capture the local and global geometric characteristics of the point cloud, further improving the accuracy and reliability of the transformation matrix estimation.
- GMCNet contributes to the development of more robust and accurate point cloud registration algorithms.
- VPRNet Virtual Point Registration Network (2022) [LYLG22] introduces a unique approach to address the challenge of partial to partial registration by utilizing a Generative Adversarial Network (GAN) to generate missing points in both the source and target point clouds. This technique aims to enhance the registration process by providing complete point clouds for better alignment.
- VPRNet consists of two main phases. The first phase involves the use of a model called VPGNet, which is responsible for generating new points in the partial point clouds. By leveraging the power of GAN, VPGNet generates plausible virtual points that fill in the missing regions of the source and target point clouds. This step aids in achieving more comprehensive coverage and reducing the impact of partial overlaps during the registration process.
- the second phase of VPRNet focuses on the actual registration process. It creates a correspondence matrix between the generated virtual points and the original points in the source and target point clouds. By establishing point-to-point correspondences, Regnet then employs SVD (Singular Value Decomposition) to compute the transformation matrix that aligns the two point clouds.
- SVD Single Value Decomposition
- VRNet (2022) [LYLG22] introduces the concept of Rectified Corresponding Points (RCP) for point cloud registration.
- RCPs have the same shape as the source point cloud but share the pose of the target. By estimating the pose between the source and RCPs, the same transformation matrix can be obtained when aligning the source with the target directly. Since the source and RCPs have the same shape, registration becomes straightforward as each point in the source has a corresponding point in the RCP.
- RCPs are derived through a rectification process based on the concept of Virtual Corresponding Points (VCP) from the DeepVCP method. VRNet leverages this approach to achieve accurate and efficient point cloud registration.
- VCP Virtual Corresponding Points
- the first step of the pipeline is the construction of 3D Hologram using patients' preoperative data like CT-scans.
- This 3D Hologram eventually will be registered on the patient allowing surgeon access to internal information directly on the patient using Hololens.
- the initial phase of the project focused on constructing a 3D pre-operative Hologram using CT scans, although in other embodiments any imaging modality may be used - nuclear medicine scans, MRI, SPECT, and SPECT CT, for example. .
- CT scans were captured at different layers to isolate specific areas of interest, including the skin, tissues, bones, vessels, and lymph nodes. These segmented images from different layers were then integrated to generate a 3D mesh, serving as the source model for this project.
- Figure 3 displays different layers of the constructed 3D model. Further details regarding this process will be provided in Section 5.
- our goal is to refine our registration model to achieve better registration performance on this specific training instance.
- our initial step is to extract the 3D point cloud of the patient within the operating theatre.
- This process involves utilizing the depth camera integrated into the Hololens.
- the depth camera captures and extracts the point cloud data, which is then transmitted to the server for further processing.
- This preprocessing encompasses various tasks, including subsampling a predetermined number of points from each cloud, converting mesh data into a point cloud format in cases where the input is in mesh form, or computing point normals for different points within the point cloud.
- the calculation of normals is particularly crucial, as it generates an essential input feature vector for the registration model.
- the preprocessed point clouds are subsequently passed through the Registration model, which is deployed on the server.
- This model takes the point clouds as input and produces the rotation and translation matrices. These matrices are then applied to the 3D Hologram to achieve precise registration on the Hololens.
- the CPU in HoloLens is specifically designed to address the unique demands of augmented reality, which involve tasks like real-time environment tracking and rendering 3D holographic objects. While it may not provide the same level of raw computational power as high-end desktop CPUs, it excels in its intended functions.
- Deep Learning models typically require significantly higher computational power and memory, often relying on GPUs or TPUs for efficient execution. Therefore, it is more practical to set up the Deep Learning-based Registration model on a remote server and utilize REST API calls from HoloLens to the server for obtaining the transformation matrix for the Hologram.
- the first step of the whole process is to generate a 3D model for registration using layerwise preoperative CT scans and then combining them together after segmentation.
- image acquisition was first performed after pseudonymization of patient information. These were subsequently transferred into 3D process software for segmentation.
- the 3D model can be seen in Figure 3.
- Segmentation of images involves processing different layers of CT scan images to identify different anatomical segments using a software such as 3D Slicer, or TotalSegmentator (available before the priority date at https://totalsegmentator.com/). The focus was on the following anatomical segments:
- Segmentation of images involves processing different layers of CT scan images to identify different anatomical segments using a software called Slicer 3D.
- Threshold algorithm was used with masking limitation to include areas which has not been segmented.
- the mean Threshold value were between (-181 and -34) for this layer. No further smoothing was needed during segmentation.
- Threshold algorithm was sufficient in segmenting the muscle layer with an additional need for smoothing.
- the mean threshold value were (-4.42 and 111.82).
- the significant overlap and interference were with the intra-pelvic structures (urinary bladder, rectum, small bowel, and vessels) understandably due to the presence of smooth muscle in most of these structures and similar intensity.
- overlap was significant the proposed primary clinical application doesn't require detailed intra-pelvic anatomical accuracy. This, however, proves the challenges facing segmenting with some algorithms especially in the complex intraabdominal and intrapelvic structures.
- Threshold is an excellent tool for bone segmentation as the highest range of intensity threshold for bone always correlate to the highest intensity in the scan. There are some limitations within segmenting the bone medulla, however, the cortex can be segmented very efficiently which is more relevant clinically and is sufficient for visualization of bone contours.
- the mean thresholding values were between (122.46 and 1677). Heavy smoothing was needed in porous bone areas like the iliac wings but less required in solid bone like femur shaft.
- SPECT-CT scan is a type of nuclear medicine scan where the images or pictures from two different types of scans are combined together. Also, the complex nature of the vascular tree adds to the segmenting challenging. Best tool used was "grow from seed”. Although time consuming it yield the best results in the final hologram. Therefore, there is significant role for machine learning and advanced automated algorithms to help segmenting this layer and improve segmenting for smaller vessels. Lymph node: The segmentation of lymph nodes was carried out manually.
- This 3D model serves as a 3D Hologram within the Hololens and functions as the source point cloud for the registration method.
- the surgeon can utilize Hololens gestures to visualize distinct layers of the 3D Hologram. This visualization assists in the localization process, enabling the surgeon to identify various internal anatomical regions relevant to the surgery. In essence, these segmented layers provide valuable guidance during the surgical procedure.
- the initial step in using the application during surgery involves constructing a 3D point cloud of the patient, which serves as the target point cloud for the registration model.
- a 3D point cloud of the patient which serves as the target point cloud for the registration model.
- we intend to utilize the depth camera within the Hololens we intend to utilize the depth camera within the Hololens.
- the first approach is to use an external RGBD camera to obtain a depth image of the scene and then convert it into a 3D point cloud.
- the RGBD camera used for this purpose is the Intel Depth Camera D435. This camera operates with an error rate of less than 2% at a distance of 2 meters. While the geometry of the point cloud is clear from the front view, the side view is less accurate. For instance, in Figure 4, on the bottom right, the boundary between the phantom model and the table appears quite blurry and does not provide accurate geometry information, which affects the accuracy of the overall registration model.
- Another issue with using this sensor is, because the best operating point of this camera is around 2 meters from the object it covers a lot of area apart from just the patient. This extra information acts as noise during registration and would affect the model's performance.
- the quality of the point cloud generated through this approach is notably superior to the one generated by the RGBD camera.
- one drawback of this method is that, to create a high-quality point cloud, we need to capture images from all around the object, and there is an additional time investment required by photogrammetry for stitching all the images together. This can introduce some delay in the overall process.
- Spatial mapping provides a detailed representation of real-world surfaces in the environment around the HoloLens, enabling developers to create convincing mixed reality experiences.
- the images from this approach can be seen in Figure 6.
- the HoloLens utilizes its depth camera to construct surface mapping.
- We developed a HoloLens application that allows users to define a bounding box around the object of interest and extract the spatial mapping created by within that region.
- HoloLens 2 introduces a research mode that provides users with access to various sensors within the HoloLens, including a 'long-throw' mode depth camera.
- a 'long-throw' mode depth camera As mentioned above, we developed an additional HoloLens application that utilizes this depth camera to extract point clouds within a specified region, which can be defined using a bounding box. The images from this approach can be seen in Figure 7.
- the point cloud produced by the Hololens Depth Sensor provide the most information about the geometry of the 3D object.
- One additional advantage of using Hololens, as opposed to an external camera, is that the locations of the camera sensors for Hololens are already known, eliminating the need for recalibration.
- Hololens offers the convenience of a bounding box feature, allowing us to focus on a specific area rather than capturing all nearby objects, which may not be possible with other methods.
- Deep Learning Registration for Hololens Optimizing Registration Performance After generating the 3D point clouds, we will now delve into different methods for 3D point cloud registration. Registration involves utilizing feature matching to determine the transformation matrix between point clouds of the same object. The primary objective of this project was to explore marker-less registration, thus our focus is on Deep Learningbased algorithms. Numerous algorithms capable of achieving this task exist, as discussed in the Literature Review section. In this study, we will specifically examine two Deep Learning-based models, RPMNet and PREDATOR. We chose these two algorithms due to their performance on popular datasets and the availability of pre-trained models.
- the main objective of the deep learning registration method is to obtain the Rotation matrix, R e SO(3), and a Translation Matrix, t e IR 3 , that align the two point clouds.
- the 3D pre-operative model will be the source point cloud, and the 3D reconstructed model of the patient will be the target point cloud.
- the Registration model will provide the Rotation matrix, R, and translation matrix, t. We then apply these matrices to the 3D pre-operative model to align it with the patient.
- RPMNet A Deep Learning-Based Registration Model
- the example deep learning-based registration model e.g. RPMNET
- RPMNET The example deep learning-based registration model
- Figures 13-15 provide further details relating to parts of the deep learning-based registration model such as the feature extraction part (Figure 13), the parameter prediction part (Figure 14) and the Compute Match Matrix (Figure 15).
- RPMNet (2020) [YL20] is a Deep Learning-based Registration model.
- the core of RPMNet is inspired by RPM or Robust Point Matching (1998) [GRL + 98], which introduced the concept of generating soft assignments as a solution to handle outliers and noise.
- Equation above describes the method of creating a soft correspondence between the source point cloud (x 7 ) and the target point cloud (y fe ).
- p is the temperature term that controls the softness of the assignment, while a is used to deal with outliers.
- RPMNet deploys this method with a difference: it uses extracted features for comparison instead of just point coordinates, as shown in the following equation.
- RPMNet applies PointNet to a 10D feature descriptor constructed for each point, which is made up of x c , ⁇ zlx Ci£ ⁇ , ⁇ PPF(x c ,x £ ) ⁇ ).
- Ax c l denotes the neighboring points translated into a local frame by subtracting away the coordinates of the centroid point:
- Ax c £ x — x c .
- PPF(X C ,X £ ) are 4D point pair features (PPF) that describe the surface between the centroid point x c and each neighboring point x £ in a rotation invariant manner:
- a and /? are learned in RPMNet by passing the two point clouds through a neural network. After feature extraction, the matching matrix is calculated, illustrating the correspondence between the two matrices.
- RPMNet uses differentiable weighted Singular Value Decomposition (SVD) to calculate the transformation matrix. Weighting is done to account for the fact that not all source points might have corresponding target points. RPMNet is an iterative algorithm, so this process is repeated for the desired number of iterations. At each iteration, the source point cloud is continually updated.
- Singular Value Decomposition Singular Value Decomposition
- Loss function for RPMNet has two components, first one is a LI loss between source point cloud transformed using ground truth Transformation matrix and the one transformed using estimated Transformation matrix. RPMNet also has another component to increase the number of inliers by estimating a secondary loss on the match matrix.
- the overall loss is the weighted sum of the two losses: ⁇ total — - ⁇ reg + ⁇ inlier
- HGLT21 consists of three main components: the Encoder, Attention module, and Decoder.
- the Encoder downsamples both point clouds using KPConv-FPN, resulting in a set of super points.
- KPConv-FPN utilizes ResNets and Convolutions to aggregate point clouds into super points.
- Predator implements Graph Neural Networks (GNN) for both sets of super points to extract contextual information.
- the super points are initially connected together to form a graph using k-Nearest Neighbors (KNN) in the feature space.
- KNN k-Nearest Neighbors
- the encoder features are iteratively updated using the following equation, where ; - represents the neighboring points of x t , and h g is a linear layer followed by normalization and leaky Re LU:
- a cross-attention module is implemented to facilitate the transfer of information between both point clouds, which is crucial for learning the overlap regions and matching between the point clouds.
- the architecture used for this module is the same as in the transformer model.
- the extracted information is then combined with the previous feature values of the super points to obtain a co-contextual feature representation.
- Local context is updated using the gathered cross-point cloud information by applying another GNN with the same structure, resulting in latent feature space encodings Fx and Fy. These scores are then used to calculate the overlap score between the super points. All this extracted information is passed through the Decoder to obtain matchability scores, overlap scores, and per-point features for both point clouds. Finally, RANSAC is used to calculate the final transformation matrix.
- Predator incorporates three components into its loss function: Circle Loss, Overlap Loss, and Matchability Loss.
- Circle Loss fine-tunes the point-wise feature descriptions by leveraging the distances between corresponding points in both point clouds in feature space.
- Overlap Loss and Matchability Loss are binary loss functions designed to train the Overlap score and Matchability score extracted from the model.
- ICP or Iterative Closest Point [BM92] is one of the first methods introduced for registration.
- ICP is an iterative process where given two sets of points of the same object, we first find the correspondences between these two sets using a distance-based optimization function. Once we have the corresponding points in both sets, we use them to find the rotation matrix and translation vector. Then we use these to move the source points to their new locations. This process is continued until convergence. This method is quite prone to get stuck in local optima.
- MSE Mean Square Error
- MAE Mean Absolute Error
- MSE Mel Square Error
- MSE iy (X t - Y n —i
- X £ and V are the elements of the ground truth matrix and the corresponding extracted matrix, respectively, and n is the number of elements in the matrix.
- MAE Mean Absolute Error
- ModelNet 40 is a dataset consisting of CAD models from 40 categories. In this project, we evaluate the performance of RPM-Net, PREDATOR, RPM-Net + ICP, and PREDATOR + ICP on various categories within ModelNet 40, with a particular focus on the 'human' category due to its relevance to our specific problem. Example images from the ModelNet 40 dataset can be seen in Figure 17.
- Figure 18 shows visulations of the registration method applied to a sample 3D model from the ModelNet40 dataset for clean target point cloud.
- the target point cloud is no longer merely a transformed version of the source cloud but is also subjected to partial occlusion.
- We randomly truncate the target point cloud using a plane mimicking real-world scenarios where the patient would typically be on an operating table, granting us only a limited view of the surgical area compared to the complete 3D preoperative model.
- the RMSE and RMAE plots indicate that while the absolute error values are higher in this case, the relative performance between RPMNet and PREDATOR remains consistent, with RPMNet outperforming PREDATOR by a noticeable margin.
- RPMNET+ICP enhances the overall performance of RPMNet.
- the visibility metric is a crucial test to assess how effectively the models operate under varying degrees of visibility.
- the occlusion is simulated by intersecting a plane through the 3D point cloud at a random location and then removing a portion of it.
- RPMNet operates as an iterative algorithm, where in each iteration, we iteratively adjust the source point cloud to gradually align it with the target point cloud.
- the model is trained using simulated data, beginning with the 3D pre-operative model as our source point cloud. We then simulate a target point cloud based on the techniques discussed in the previous section. This simulated target point cloud is subsequently processed by the registration model, which provides a transformation matrix. The difference between this matrix and the ground truth serves as error feedback to train the model.
- the example flow diagram for the training pipeline of the registration model can be seen in Figure 19.
- Input STL file location The file location of the 3D point cloud file of the 3D preoperative model of the patient.
- Translation Magnitude The maximum translation that the simulated target point cloud is allowed to be away from the initial source point cloud.
- Visibility control Controls how much of the simulated target point cloud is visible.
- Noise introduction As discussed earlier, we add noise in the form of a plane through the point cloud, simulating the patient lying on a table surrounded by spherical noise scattered throughout the scene.
- RPMNet is an iterative algorithm, we can set how many iterations we want to run the model for.
- Model path specifies the model file that we want to finetune.
- the GUI offers two testing modes: users can either evaluate the model using actual data or simulate the testing with predefined parameters. Opting for simulated data mode means that the source point cloud undergoes manipulation based on the simulation parameters previously described in the Train mode.
- the example GUI can be seen in Figure 21.
- test mode In addition to the previously mentioned parameters, a distinct set of options is accessible in the test mode:
- Input Target file (3D point cloud file obtained from any 3D reconstruction method): It is the 3D point cloud file constructed from any 3D reconstruction method.
- Visualize End visualization of the registered point cloud: Visualize allows users to see the model's performance by visualizing the actual registered point cloud.
- Save Mesh Save mesh allows users to choose whether they would like to save the mesh of the registered point cloud scene.
- Capture RGB allows users to save a rendered image of the 3D registration scene.
- the goal of the registration method is to align the 3D Hologram model as precisely as possible over the patient in Hololens.
- both the patient's 3D reconstruction and the 3D Hologram model are represented as point clouds, and our objective is to ensure these point clouds overlap as closely as possible.
- One metric that aids in achieving this objective is the nearest neighbour distance.
- Evaluation Metric for Different Methods we evaluate the model using 100 random translations and rotations applied between the 3D pre-operative Hologram and the 3D reconstructed point clouds. Subsequently, we calculate the average of Nearest Neighbor Distance obtained in each configuration.
- the visualization and evaluation metric indicate that the model excels when working with point clouds extracted from the Hololens Depth Camera, achieving exceptional accuracy during registration. This outcome is especially encouraging, as it demonstrates that the Hololens possesses the capability to generate high-quality point clouds independently, negating the necessity for additional external hardware for processing.
- Fine-tuning plays a vital role in addressing subtle disparities between the source and target point clouds, resulting in a significantly improved registration process.
- this project has established a robust platform for advancing research by combining Hololens technology and Deep Learning models to offer a more convenient and informative approach to guide surgeons during surgical procedures.
- the project has achieved several noteworthy milestones:
- Model Training and Testing To improve the precision of the registration model, we generated a synthetic dataset by manipulating 3D preoperative models. This synthetic data served as the training data for our deep learning model, which we rigorously evaluated using both synthetic and real point cloud data.
- this project represents a significant step towards enhancing surgical procedures through the fusion of augmented reality and deep learning. Its potential applications extend to guiding surgeons in real-world operating scenarios, promising impactful contributions to the field.
- PCL21 Liang Pan, Zhongang Cai, and Ziwei Liu. Robust partial-to-partial point cloud registration in a full range. CoRR, abs/2111.15606, 2021.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- General Physics & Mathematics (AREA)
- Molecular Biology (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Biophysics (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Surgery (AREA)
- Human Computer Interaction (AREA)
- Robotics (AREA)
- Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
- Heart & Thoracic Surgery (AREA)
- Medical Informatics (AREA)
- Animal Behavior & Ethology (AREA)
- Public Health (AREA)
- Veterinary Medicine (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Apparatus For Radiation Diagnosis (AREA)
Abstract
A method of marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, comprising: scanning the subject using a time-of-flight sensor to generate a point cloud data set corresponding to a skin layer of the subject, obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the skin layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and displaying the three-dimensional virtual model so aligned within the live imagery of the subject.
Description
SPATIAL ANATOMY RECOGNITION TECHNOLOGY
Field
The present disclosure generally relates to spatial anatomy recognition and guidance through the use of spatial anatomy imaging technology.
More specifically, but not exclusively, the present disclosure relates to marker-less imaging for providing spatial anatomy models to locate all non-visible target body segments in the spatial environment.
Background
It is well known that medical imaging of the human anatomy has become hugely important in many facets of the medical world. Accurate imaging or scanning of the human anatomy allows healthcare professionals to diagnose and treat various medical conditions by providing an insight into internal structures such as organs, tissues, bones etc. more reliably. This aids in detecting issues such as tumors, fractures, infections etc. but can also aid in surgical procedures where avoiding other parts of the anatomy is crucial.
There is a constant drive in the medical profession to improve the efficiency and accuracy of medical and therapeutic processes. This includes the speed, efficiency and accuracy of surgery and therapies such as radiotherapy. In the last decade, there has been a great effort to bring mixed reality (MR) into the operating room to assist surgeons intraoperatively.
Two-dimensional (2D) medical visualization techniques are often insufficient for displaying complex, three-dimensional (3D) anatomical structures. Moreover, the visualization of medical data on a 2D screen during surgery is undesirable, because it requires a surgeon to continuously switch focus. The use of augmented reality (AR) has the potential to overcome these problems, for instance by using markers on target points that are aligned with the AR solution. However, placing markers for a precise holographic overlay are costly, always have to be visible within the field of view, have the potential to increase the regulatory burden, and can disrupt the surgical workflow.
Therefore, there is increasing demand for real-time imaging technology that can provide visualization of targeted non-visible anatomy without the need for placing further
equipment such as markers to the subject or patient. In some cases, there is a need for real-time imaging technology that can provide visualization of targeted non-visible anatomy which can be displayed in an augmented reality environment for a surgeon or medical professional to see during a procedure. This would allow for more informed realtime decision making.
Summary of the Disclosure
The need to develop a tool that can both aid surgeons, and medical professionals alike, to increase the ease, efficiency and accuracy of surgical and therapeutic procedures and provide precision guided interventions whilst reducing the invasiveness or discomfort on the subject is growing. Similarly, a tool that can aid in training or educational purposes for junior doctors, surgeons, interventional radiologists or even patients themselves would be massively beneficial.
As such, the present disclosure provides methods and systems for marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject. This will provide a displayed 3D virtual model of the subject to the user, the 3D virtual model being accurate relative to the subject's anatomical composition. The 3D virtual model will be provided to the user in real-time. This will allow the user to better understand the relative positioning of anatomical features within the subject e.g. bones, organs, arteries etc. such that the user can better understand the anatomy of the subject and thus, make more informed decisions during surgical and therapeutic procedures.
In a non-limiting example of the present disclosure, the method of marker-less registration of a pre-obtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject may include: scanning the subject using a time-of-flight sensor to generate a point cloud data set corresponding to a skin layer of the subject, obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the skin layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding
anatomical features of the subject in co-registration with each other; and displaying the three-dimensional virtual model aligned within the live imagery of the subject.
This would allow for more informed real-time decision making, by improving accuracy of guidance of surgery, biopsy, injection, marker placement or radiotherapy treatment. It has the further benefit of reducing the time to conduct surgery and/or radiotherapy and improving the performance of the surgical and/or radiotherapy team whilst reducing the risks to patients and so deliver patient benefits.
An aspect of the present disclosure presents a method of marker-less registration of a preobtained medical imagery three-dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, comprising: scanning the subject using a depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject; obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the anatomical layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three- dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and displaying the three-dimensional virtual model so aligned within the live imagery of the subject.
In one embodiment the camera and the depth sensor are substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween. Having the camera and the depth sensor substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween improves the success rate of co-registration between the point cloud data set and the 3D virtual model.
In one embodiment the processing comprises inputting the point cloud data set and the three dimensional virtual model into a trained neural network trained to find matching anatomical features in the two data sets. Using a trained neural network to find matching anatomical features in the point cloud data set and the three dimensional virtual model has several advantages. A trained neural network can accurately and efficiently find matching anatomical features in the two data sets, improving the success rate of coregistration between the point cloud data set and the 3D virtual model. This can improve
the accuracy of guidance during surgery or other medical procedures, and reduce the risks to patients.
In one embodiment the aligning comprises determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features. Determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features allows for accurate alignment between the two data sets.
In one embodiment the aligning further comprises determining the rotation matrix R and the translation matrix t by applying a weighted singular value decompositions (SVD) to the point cloud data set. More preferably, the applying the weighted singular value decompositions, SVD, to the point cloud data set is repeated by a predetermined number of iterations, n. Using a weighted SVD to determine the rotation and translation matrices that align the anatomical layer point cloud data set with the three-dimensional virtual model can improve the accuracy of the alignment between the two data sets. By repeating the application of the weighted SVD by a predetermined number of iterations, the alignment can be further refined, resulting in a more accurate co-registration between the point cloud data set and the 3D virtual model.
In one embodiment the scanning of the subject using the depth sensor places emphasis on segmenting the scan to highlight surface skin of the subject.
In one embodiment the one or more anatomical features are one or more of skin, organs, tissues, bones, arteries, veins, as well as the surface geography of skin, organs, and tissue planes.
In one embodiment the one or more anatomical features are one or more abnormal anatomical features, the one or more abnormal anatomical features being tumours, bone fractures, bone deformities, vascular pathologies, infections, as well as pathological lumps such as malignant tumours, skin lesions benign and malignant, bone anatomy.
In one embodiment the method further comprises repeating the above steps and updating the display to account for movement of the user about the subject, or movement of the anatomy of the subject, during the surgical or therapeutic procedure. This allows the display to track relative movement and ensure that co-registration between the point cloud data set and the corresponding anatomy in the live imagery is maintained. This means
that the 3D virtual model displayed to the user remains accurate and aligned with the subject's anatomy, even if there is movement during the procedure. This can improve the accuracy and effectiveness of the surgical or therapeutic procedure.
From another aspect the present disclosure also provides a control system for controlling an imaging system, the imaging system comprising a depth sensor and a camera, the control system comprising one or more processors collectively configured to: scan the subject using the depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject, obtain live imagery of the subject via the camera; receive the pre-obtained medical imagery three-dimensional virtual model of a subject; process the anatomical layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; align the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in coregistration with each other; and display the three-dimensional virtual model so aligned within the live imagery of the subject.
Further features and advantages will be apparent from the appended claims.
Figures
Further features and advantages of the present disclosure will become apparent from the following description, presented by way of example only, and with reference to the accompanying drawings, wherein like reference numerals refer to like parts, and wherein:
Figure 1 shows an example flow diagram of the hololens registration application pipeline, in accordance with the present disclosure;
Figure 2 shows a further example flow diagram of the marker-less hologram-guided registration method, in accordance with the present disclosure;
Figure 3 shows example software images of different layers of the human anatomy produced by a 3D Pre-Operative Hologram;
Figure 4 shows example software images showing different views of 3D Point Cloud Extraction using Intel Depth Camera D435;
Figure 5 shows example software images showing different views of 3D Point Cloud Extraction using Polycam mobile applications;
Figure 6 shows example software images showing different view of 3D Point Cloud Extraction using Hololens spatial mapping;
Figure 7 shows example software images showing different view of 3D Point Cloud Extraction using Hololens Depth Camera;
Figure 8 shows an example flow diagram for the data simulation pipeline for use in producing improved target point clouds;
Figure 9 shows example images of different configurations of target point cloud after applying different rotation and translation to the source point cloud;
Figure 10 shows example images of different configurations of target point cloud after occluding random areas;
Figure 11 shows example images of different configurations of target point cloud after adding a plane at the truncated region;
Figure 12 shows an example system diagram of a deep learning-based registration model, namely RPMNet, in accordance with the present disclosure;
Figure 13 shows an example system diagram for a feature extraction part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure;
Figure 14 shows an example system diagram for a parameter prediction part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure;
Figure 15 shows an example system diagram for a matrix part for use with the deep learning-based registration model of Figure 12, in accordance with the present disclosure;
Figure 16 shows an example system diagram of an encoder-attention-decoder model, namely PREDATOR;
Figure 17 shows example image sample sets from the ModelNet 40 dataset;
Figure 18 shows a visualization of the registration method applied to a sample 3D model from ModelNet40 dataset for clean target point cloud, in accordance with the present disclosure;
Figure 19 shows an example flow diagram of the training method for the registration model, in accordance with the present disclosure;
Figure 20 shows an example flow diagram of the testing method of the registration model, in accordance with the present disclosure;
Figure 21 shows an example graphical user interface for the training and testing of the registration model, in accordance with the present disclosure;
Figure 22 shows a comparison of images of the registration performed using the generic registration model (on the left) versus the fine-tuned registration model (on the right), in accordance with the present disclosure;
Figure 23 shows a comparison of further images of the registration performed using the generic registration model (on the left) versus the fine-tuned registration model (on the right), in accordance with the present disclosure;
Figure 24 shows a visualization of the registration applied to data simulated using SPL/NAC Brain Atlas [Opeteb];
Figure 25 shows a visualization of the registration applied to data simulated using SPL Head and Neck Atlas [Opetea];
Figure 26 shows a block diagram of a computer system for use with the 3D marker-less hologram-guided registration method, in accordance with the present disclosure;
Figure 27 shows a flow diagram of the operation of the automatic registration process using deep learning model, in accordance with the present disclosure.
Detailed Description
The present disclosure seeks to provide methods and systems which may use 3D scene scanning technology to generate a point cloud of a subject's anatomy, and in particular in some embodiments a skin layer of the anatomy, the subject being a patient about to undergo some sort of surgical or therapeutic procedure. Pre-obtained medical imagery (e.g. CT/MR, 3D ultrasound etc.) of the subject may be combined into a 3D hologram model of the subject, which may then be registered with the point cloud so as to overlay
virtually the 3D hologram model with the point cloud of the subject's anatomy in a user display. The marker-less registration may be performed by using neural networks to extract predefined corresponding anatomy features in each of the anatomical model and the 3D hologram model which then provide anchor points for a matching matrix transform to map the skin point cloud to the 3D hologram model across the two data sets. The marker-less registration and overlay may be dynamically updated as the user moves and his field of view of the subject changes or the subject position changes in the spatial environment. The result is an augmented view of the subject with the 3D medical imagery hologram model virtually overlaid onto the subject within the user's field of view. Preferably the user display is head mounted on the user, although in other embodiments the display can be on a screen such as a monitor or the like, or provided for robotics and automation purposes.
In other embodiments audio sensory cues may be provided in addition to or as an alternative to the augmented view. For example, if a user's field of view strays from the intended area of anatomy, pulsing sounds may be emitted that are faster or slower depening on how far away from the intended area of anatomy the user view has strayed.
The augmented view of the subject enables the user to receive real-time visual representation of the subject in the form of a 3D medical imagery hologram model. This enables the user to make more informed decisions during the current surgical or therapeutic procedure etc. Further, it enables procedures to be executed in a more accurate and precise manner which should reduce the time taken to complete such procedures.
Moreover, in the present disclosure, the 3D medical imagery hologram model could be used for surgical training or patient education such that both students or patients can see where the steps of a surgical procedure would occur within the subjects body in a non- invasive manner, for example.
In a non-limiting example, the present disclosure seeks to provide a method that uses a deep learning algorithm and trained Al to register the Mesh generated by depth cameras to the skin segmentation topography. Therefore, the technology locates all non-visible target body segments in the spatial environment for positional and interventional adjustment automation.
In a further non-limiting example, the present disclosure may relate to a method of marker-less registration of a pre-obtained medical imagery three-dimensional virtual
model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, the method may comprise: scanning the subject using a time-of-flight sensor to generate a point cloud data set corresponding to the skin layer of the subject, obtaining live imagery of the subject via a camera; receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; processing the skin layer point cloud data set and the three-dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in coregistration with each other; and displaying the three-dimensional virtual model aligned within the live imagery of the subject.
In essence, it may comprise the following components:
• Anatomical segmentation
This may be achieved through post-acquisition processing of medical images (CT/MRI scans) converting the typical 2D images into 3D segments with two specific components.
A - Skin segment
This may provide the point cloud mesh (a set of topographic geometric points that provide an imprint of initial position of human body contour)
B - Target anatomy
This may provide the internal position of the part of body targeted by the intervention in relation to skin (the spatial position of the target anatomy is constant to the surrounding skin with any movement on X or Y axis and any global rotational movement. The deformation of the human anatomy due to respiration movement and joint movement should be considered for accurate anatomy localisation).
• Realtime body surface scanning
This may provide the second input which depends on generating a point cloud mesh of the patient body at the time of intervention. This may be achieved by using time of flight (ToF) camera or other 3D depth cameras.
The main camera may be the Azure Kinect DK depth camera which implements the Amplitude Modulated Continuous Wave (AMCW) Time-of-Flight (ToF) principle. This may provide the real-time [3D] body surface scan image.
In one embodiment the camera may be integrated into the Hololens® device and allows for seamless integration of surgical guidance using the depth camera and display on the HoloLens.
In other embodiments other cameras such as an Intel® LiDAR camera or a standalone Azure Kinect® DK depth camera may be used, and it follows the same process in the integration of the technology pipeline. Both can be used for the other clinical application in interventional radiology and radiotherapy. Other depth capturing imaging devices are of course available that are the equivalent of those mentioned above, any and all of which can be used in embodiments of the present disclosure.
There are multiple platforms that allow such post image processing. This includes, but is not limited to, open-source software as such as 3D slicer, which include segmentation Al models, and business developed segmentation platforms.
• Deep learning Al model
The Al model seeks to process the thousands/millions of geometrical points in both point clouds and produce the necessary calculation to produce an output instruction to register both inputs together through point matching. It may be referred to as the Spatial Anatomy Recognition Artificial Intelligence (SARAI) model.
Point matching is the process of finding corresponding points between two or more sets of points. It is a technique used in computer vision, particularly in fields such as robotics, image processing, and 3D modelling.
Imagine having two sets of points: one set represents a 3D reconstruction of patient anatomy, and the other set represents the actual patient in real time. Point matching aids in aligning the two sets of points to create a 3D anatomy that represents the patient.
The proposed model may use deep neural network to learn point features and can be trained to improve its precision and standardized when achieving the needed results.
This may be further used in the tracking of any changes in patient position during the intervention and rematch the 3D anatomy to the patient.
Combining the components of anatomical segmentation, real-time body surface scanning and deep learning Al models, will give outputs for the user (e.g. surgeon) which are a 3D visualization of the internal organs, tissue, arteries and veins of a patient in real-time when conducting surgery, especially operations with limited tissue distortion e.g. lymph node biopsy, soft tissues in extremity, limb, vascular and tumour surgeries; including both
primary and metastatic lesions. In particular, having a fixed point when there is tissue distortion is important for treatment of lymph nodes, as it is all deformable.
It may be also used in interventional radiotherapy planning, delivery and monitoring. Furthermore, it has potential as a training tool for junior surgeons and radiographers; and where appropriate, to explain the surgery and/or treatment to patients to help them understand the process.
Therefore, the potential use cases may be, but are not limited to:
1- In surgery it could be used by moving the first input (anatomical segmentation) to match the second input (Realtime body surface scanning). Therefore, allowing the surgeon to visualize the targeted nonvisible anatomy on the augmented reality environment rendering them the ability to see through the skin before making an incision. This enables better locating of the targeted anatomy, decrease time required for surgery, increase surgery success.
2- In radiotherapy, it may be applied in reverse by validating and calculating the needed movement of the patient body (Realtime body surface scanning) to the radiotherapy field to ensure the non-visible target anatomy (anatomical segmentation) is in the correct position.
3- Surgical, radiotherapy and radiography training.
4- Patient education.
5- Guiding injections and injectable therapies
6- Guiding Placement of fiducial markers in tissue
7- Performing Tissue biopsy
8- Integration into robotic surgery and interventions.
The present disclosure provides a Mixed Reality solution. Mixed Reality is the merging of real and virtual worlds to produce new environments where both physical and digital objects co-exist and interact in real time such as the use of holograms and mixed reality glasses.
Augmented Reality (AR) is the result of using technology to superimpose digital elements (such as sounds, images, and text) to the world we see using a tablet, smart eyeglass or a smartphone camera.
Virtual Reality (VR) is an immersive, three-dimensional, computer-generated environment which can be explored and interacted with by a person. Instead of projecting the images
and sounds on a real environment, virtual reality uses a headset to immerse users within a 360-degree environment.
The present disclosure will now be described with reference to the figures.
An example system diagram 1 is presented in Figure 1. The example system diagram relates to a pipeline for 3D imaging technology; specifically, the system diagram relates to marker-less mixed reality visualization for surgical methods. The system diagram 1 may also be referred to as 3D medical imagery hologram model 1. In this example, the 3D medical imagery hologram model 1 comprises the following: pre-operative images 10, a time-of-flight sensor (such as, but not limited to, a hololens) 11, a 3D hologram 12 and a server 13. The time-of-flight sensor 11 enables 3D scene reconstruction 110 and hence point cloud generation of the scene. The server 13 enables processing of the point cloud generated 130 and the use of a registration model 131.
In use, the 3D medical imagery hologram model 1 initially constructs a 3D hologram model 12 of a subject by using the subjects existing image data 10. The existing image data 10 may be a CT scan, or any other applicable image type. The existing image data 10 is used to develop a simulate training set for use in training a deep learning model for the present subject. Then, the user will utilize the time-of-flight sensor 11 e.g. wear the hololens, during surgery. At this point, a point cloud of the subject is generated 130 within the server 13 using a depth camera with the time-of-flight sensor 11. The point cloud is then sent to the server 13 along with the 3D hologram 12. Both the subject's point cloud and the 3D hologram 12 are preprocessed before being utilized as input data to a registration model 131. The registration model 131 performs the registration and outputs a rotation and translation matrix 132. The rotation and translation matrix 132 is applied to the 3D hologram 12 moving it in space to mirror the position of the patient scanned point cloud.
A method or flow diagram of operation of the 3D medical imagery hologram model 1 is shown in the flow diagram 2 of Figure 2. The flow diagram 2 is split into three distinct parts; before usage of the system or model 20, during initialization of the system or model 21 and during usage of the system or model 22. The before usage part 20 relates to the patient or subject's pre-existing image data. This can be image data such as CT scans that have been taken prior to use of the system or model. The pre-existing image data is then processed by segmentation to produce 3D models of internal anatomy and 3D models of the surface.
These produced 3D models are to be used in the initialization of the system or model part 21. Initially, these 3D models are used to generate a 3D point cloud of the patient's anatomy (skin layer). At the same time, the system or model is receiving a live scene with the subject or patient somewhere in the field of view. This live scene is likely provided by a time-of-flight sensor. The depth and position of the subject in the live scene is recorded. These depth and positional values are used in the 3D point cloud registration. The depth and positional recording of the live scene and the 3D point cloud of the patient's anatomy are used in conjunction to create a transformation of the 3D model to the current position of the patient. To note, HL2 in Figure 2 refers to a Hololens 2 device as the time-of-flight sensor, however, any other applicable time-of-flight sensor may be used.
The transformed 3D model, which utilizes the live depth and position values from the time- of-flight sensor and the 3D point cloud from the pre existing image data, is provided to the user during usage of the system 22. In this way, the transformed 3D model is constantly supplied to the user during use and the relative position of the time-of-flight sensor at that moment is fed back into the transformed 3D model. This ensures that the transformed 3D model is always accurately mapped to the live scene which includes the subject or patient.
An example system diagram of a registration model 12 is shown in Figure 12. The registration model 12 is applicable for use as the registration model 131 in Figure 1. The registration model 12 may be a deep learning-based registration model, such as, but not limited to, a RPMNet model. The registration model 12 contains six main parts: a rigid transform block 120, a parameter prediction block 121, a feature extraction block 122, a compute match matrix block 123, a match matrix 124 (which may be a rotation and translation matrix) and a weighted singular value decomposition block 125. The feature extraction block 122 is shown in more detail in Figure 13. The parameter prediction block 121 is shown in more detail in Figure 14. The compute match matrix 123 is shown in more detail in Figure 15.
The details of the performance of the registration model 12 are discussed later.
Figure 26 is a block diagram 26 of a typical general purpose computer system 260 that can form the processing platform for the 3D marker-less hologram-guided registration method, as described, in accordance with the present disclosure. The general-purpose computer system 260 may also be referred to as a control system 260 for controlling an imaging system. The computer system 260 comprises a central processing unit (CPU), random access memory (RAM), and input/output ports (I/O) into which data can be
received and output therefrom as is well known in the art and a network I/C. Additionally included is a camera unit 262 which may be positioned to record live surgical video, a time-of-flight (ToF) sensor 263 and a network access point 264 for sending or receiving data including the receiving of preoperative images, surgical videos etc. Although, many further appliances could be used if necessary.
The computer system 260 also includes some non-volatile storage, such as a hard disk drive, solid-state drive, or Non-Volatile Memory Express (NVMe) drive. Stored on the nonvolatile storage 261 is a number of executable computer programs together with data and data structures required for their operation or training. Overall control of the system 26 is undertaken through the execution of the Point Cloud Generation Model 2611 and the Registration Model, in conjunction with pre-existing image data 2612 and live scene image data 2614. The other data contained in the non-volatile storage 261 is the training data 2615 which could include surgical or pharmaceutical images or videos or the like. In some examples, the training data 2615 may be images from the ModelNet40 dataset.
An example flow diagram 27 of the operation of the 3D marker-less hologram-guided registration method is shown in Figure 27. In the example flow diagram 27 , the flow starts at s270 where the time-of-flight sensor e.g. the Hololens scans the surface of the subject in its field of view. This is likely to be the surface of the patient before or during a surgery or therapeutic procedure. At s271, a point cloud of the subject surface is generated from the scan of the subject surface taken in s270. Then in s272, the generated point could from s271 is sent to a registration model. The generated point cloud of the subject surface becomes the first input to the registration model. In a separate flow containing s273, a further point cloud is produced from surfaces extracted from patient preoperative images e.g. from existing subject CT scans etc. This produced further point cloud is sent to the registration model as a second input. In s274, the two flows combine as the registration model receives the first and second inputs, namely the point cloud from the scanned subject surface and the point cloud from the subject preoperative images. Thus, at s274, the registration model calculates a translation and transformation matrix which highlights the relative positional differences between the first and second input. This translation and transformation matrix is to be used to perform co-registration of the first and second inputs. Therefore, in s274 the translation and transformation matrix is outputted. In s275, the translation and transformation matrix are used to map a generated hologram (generated from the preoperative imaging) onto the real scene subject position e.g., the ToF sensor scene. This achieves co-registration.
In a further example flow diagram, the operation of the 3D marker-less hologram-guided registration method may comprise:
1. Obtaining preoperative images of a subject which are transferred into 3D segmented models and then holograms. The transferring placing special emphasis on segmenting the subject surface skin;
2. The subject surface skin segment of the preoperative images will be used to generate a point cloud. The point cloud is transferred to a registration model as a first input;
3. Next, a ToF sensor or camera e.g. a Hololens, will reconstruct the 3D scene during the real-time intervention and a point cloud will be generated from the scene reconstruction;
4. The generated point cloud from the scene reconstruction is transferred to the registration model as a second input;
5. The registration model calculates the needed translation and transformation matrix to perform co-registration between the first input and the second input. The translation and transformation matrix being calculated from the relative positional differences between the first and second input. The translation and transformation matrix is then outputted;
6. The translation and transformation matrix is applied to the hologram generated from the pre-operative imaging.
7. The translation and transformation matrix causes the hologram to be moved to match the real scene subject position. Thus, achieving co-registration.
In further example flow, in accordance with the present disclosure, relating to a method of marker-less registration of a pre-obtained medical imagery (e.g. CT scan or the like) 3D virtual model of a subject with live imagery (e.g. from a camera) of the same subject obtained before or during a therapeutic or surgical procedure on the subject.
The flow begins by scanning the subject using a time-of-flight sensor (e.g. a Hololens) to generate a point cloud that corresponds to the skin layer of a subject. At the same time, live imagery of the subject will be obtained from the camera. At this point, the pre-obtained medical imagery 3D virtual model (e.g. CT scans) of the subject will be received. The skin layer point cloud data set and the 3D virtual model skin layer will be processed to determine anatomical features of the subject which are present in both the point cloud data set and the 3D virtual model. The processing may entail inputting the point cloud data set and the 3D virtual model into a trained neural network trained to find matching anatomical features in the two data sets. These anatomical features may relate to skin surface, organs, arteries, bones etc. Next, the point cloud data set and the 3D virtual
model will be co-registered, which may be achieved by a registration model. The coregistration requires the alignment of the point cloud data set and the 3D virtual model, which is achieved by aligning the determined anatomical features of the subject which were present in both the point cloud data and the 3D virtual model. This alignment may be achieved by determining a rotation matrix R and a translation matrix t that aligns the skin layer point cloud data set with the 3D virtual model in dependence on the determined anatomical features. Finally, the 3D virtual model which is aligned with the live imagery of the subject is displayed. This may be displayed on a head
-mounted device, such as the Hololens, or other forms of screens.
It is worth noting that the rotation matrix R and the translation matrix t may be determined by applying a weighted singular value decompositions (SVD) to the point cloud data set. This weighted singular value decompositions may be iteratively applied to the point cloud data set by a predetermined number of iterations, n, to produce the rotation matrix, R, and translation matrix, t.
In some instances, the camera and the time-of-flight sensor may be substantially aligned. This may enable them to have the same or similar field of view of the subject or have a known field of view offset therebetween. This will likely improve the success rate of coregistration between the point cloud data set and the 3D virtual model.
Moreover, the one or more anatomical features of the subject may relate to anatomical features or abnormal anatomical features. The anatomical features may include one or more of organs, tissues, bones, arteries, veins or any other anatomical features. The abnormal anatomical features may include one or more of tumours, bone fractures, bone breaks, infections, or any other abnormal anatomical features.
The above steps may be continuously repeated. Continuously repeating the above steps may ensure that the display is consistently updated to account for movement of the user in relation to the subject, or for movement of the anatomy of the subject, during the surgical and therapeutic procedure.
Further details of the operation of embodiments of the present disclosure will become apparent from the following further description.
Abstract
In the last decade, there has been a great effort to bring mixed reality (MR) into the operating room to assist surgeons intraoperatively. However, progress towards this goal
is still at an early stage. Using the HoloLens, previous research has mainly focused on projecting a virtual 3D model into a patient's body or just above it to avoid obstructing a surgeon's line of sight. However, in these studies, the models were registered manually to their respective subjects by the operating surgeon. This project would examine the feasibility and application of different techniques to perform automatic registration between a HoloLens model and a surgical phantom model to assist surgical navigation through MR visualisation. For the 3D reconstruction of the phantom model, the built-in functions of the HoloLens platform will be explored and compared to the 3D structures recovered with computer vision techniques. The versatile nature of this platform makes it suitable for a plethora of surgical procedures. Key application of this mixed reality visualisation platform will be the localisation of the sentinel lymph node (SLN) for surgical guidance during melanoma resections.
Introduction
Mixed Reality (MR) represents a fusion of the real and virtual worlds, where virtual objects interact seamlessly with the real environment. It can be seen as an enhanced form of augmented reality [zHbFwS+19], where virtual elements become integrated into the physical world. Devices such as Google Glass, Magic Leap, and Microsoft’s Hololens have emerged as key supporters of mixed reality experiences, with Hololens dominating the field since its inception. Hololens enables users to place and interact with 3D virtual models in the real world, offering features such as stereo vision, low latency room mapping, head tracking, and gesture-based interaction using hand movements [BECS22]. Furthermore, Hololens allows the transfer of data, such as camera images, depth images, and camera coordinates, to other devices for additional processing. Mixed Reality has found extensive applications in various fields, including education, engineering, construction, and medicine. This project focuses specifically on the application of Mixed Reality in the medical field, particularly in surgical assistance.
Visualization tools are crucial in aiding surgeons in understanding human anatomy. However, most of these tools rely on 2D displays, which present challenges in transferring knowledge from a 2D image to a 3D patient. Switching between the screen and the patient can be inconvenient and hinder surgical accuracy and safety [AHS+21]. Mixed Reality devices like Hololens can be used to address this issue by registering 3D models onto the patient, allowing surgeons to achieve a better field of vision during surgery and improving surgical precision and safety [zHbFwS+19]. One more advantage of using Hololens for these applications is the shared experience, where the same MR visualization can be shared across multiple Hololens allowing multiple people simultaneous access to the information. This project primarily concentrates on the registration process of 3D models
of patients using Hololens. Various methods have been proposed to achieve this registration, including manual registration, point-based registration, and surface registration [BECS22]. Although commonly used, manual registration can be prone to inaccuracies. Point-based registration requires the use of external markers on the patient for the registration. This project aims to minimize errors caused by human intervention or reliance on external indicators by aligning the 3D point cloud of the segmented preoperative CT scans with the 3D point cloud representation of an actual patient through analysis and comparison of their respective surface geometrical characteristics. The objective of this project is to evaluate deep learning-based techniques for surface registration, with a focus on assessing their compatibility with the Hololens platform. The ultimate aim is to showcase precise registration capabilities on a Phantom model using Hololens 2.
In the following chapters, we will delve into the features of Hololens and their significance in the context of this project. We will examine both traditional and deep learning-based surface registration methods, shedding light on their respective advantages and constraints. Furthermore, we will offer insights into the experimental setup and the process of integrating the Deep Learning algorithm with Hololens.
Literature Review
Hololens
Hololens is a mixed-reality device that bridges the gap between the virtual and real world. It allows users to render 3D models in the real world. Additionally, it offers a user interface based on hand gestures, enabling users to manipulate the 3D models with respect to the real world. The Hololens has found its application in several fields like Education [MPRS21], Construction [HSO19], and Medical [Pal22a].
Several studies discuss its applications in Surgical Navigation, Human-Computer Interaction, Rehabilitation, and Medical Training [Pal22b].
[WPJ+22] talks about the benefits of holograms on Hololens, as they contribute to a better presentation of tumour size and locations. Surgeons can easily compare the real patient's anatomy with holographic visualization. [vHCE21] shows how Hololens can be utilized to superimpose holograms on real-world objects using a similar approach, which will be the focus of this project.
This project aims to study deep learning-based models using data extracted from Hololens. [ILH+17] demonstrates how to run deep learning models on Hololens by connecting it to
a server for request processing. [UBG+20] provides several examples of extracting data from different sensors on Hololens for processing. [LDZS18] evaluates that Hololens is capable of estimating the user's head posture accurately at low movement speeds and reconstructing the environment with high precision, particularly for flat surfaces under bright conditions.
In this project, we will utilize Hololens to register the model of the area of surgery created through CT scans on the actual patient. For this purpose, we will be working with the new Hololens 2. Hololens 2 offers several capabilities under Hololens Research mode that can be utilized for this project, such as:
1. Access to IMU sensors like the accelerometer, gyroscope, and magnetometer, which can be utilized to track the location of the headset with respect to the real world [Pol20].
2. A depth camera which uses active infrared (IR) illumination to determine depth through phase-based time-of-flight. The camera can operate in two modes. The first mode enables high -fra me- rate (45 FPS) near-depth sensing, commonly used for hand tracking. The other mode is used for lower-frame rate (1-5 FPS) far-depth sensing, currently employed for spatial mapping [Pol20].
3. The sensor streams which can either be processed or stored on the device or wirelessly transferred to another PC or the cloud for more computationally demanding tasks [Pol20] .
During the procedure, we expect the surgeon to wear the device and focus on the patient's surgical area. This will be followed by creating a 3D point cloud using the depth image generated by Hololens, which will then be registered to the 3D point cloud generated from the CT scans.
Registration
Registration is the process of finding the transformation from a source point cloud to a target point cloud of the same object [LSW17]. In this project, we aim to register 3D model generated using the preoperative CT scans over the point cloud generated using Hololens of the actual patient. There are multiple methods for carrying out registration, such as manual registration, marker-based registration, and surface registration. In the following sections, we will discuss each method in detail.
Manual Registration
Manual Registration is a straightforward approach to registration where the user has to manipulate the 3D model in a way that aligns with the patient. Several papers have proposed various methods for executing manual registration.
[PIL+18] proposed a method that uses hand gestures and voice commands to execute registration. Different hand gestures for translation and rotation were used to move the 3D model until accurate alignment between the anatomical landmarks and skin was achieved. Voice commands were implemented to switch between translation and rotation modes. [MJC+18] used a similar method for registration. [GGS+18] attempted to manually position the scapula in a holographic mode so that the surgeon could visualize hidden parts of the scapula during surgery.
[LRvA+19] attempts to combine manual registration with automated registration. Automated 3D registration generally includes two steps: generating a 3D point cloud for the patient and registering the 3D point cloud of the segmented preoperative CT scan's 3D triangular surface model. This paper proposes manually sampling points of importance to generate a 3D point cloud for the patient using a custom-made pointing device (PD). The PD consists of a notch, a handle, and a tip. The notch is equipped with a marker of known geometry, which is used for tracking. The PD is moved along the patient to sample the 3D points of interest. Once sufficient sample points have been collected, the ICP algorithm is used to register the 3D model onto the patient by calculating correspondences between the two 3D point clouds.
[NCR+20] compared three different methods of manual registration: tap-to-place, 3-point correspondence matching, and keyboard control. Tap to place involved using an air tap as a mouse click to place the 3D model on the patient. 3-point correspondence matching involved choosing two sets of 3 points from the 3D model and the patient, followed by finding the translation and rotation matrix. Keyboard control involved using a Bluetooth keyboard to place the 3D model. The results of the experiments concluded that the keyboard method was the most accurate, followed by the tap-to-place method, and finally the 3-point correspondence method.
While manual registration is simple to execute, it is often prone to human-introduced errors. Another disadvantage of these methods is the time consumption. Due to these issues, it might be inconvenient to rely solely on manual methods in surgical procedures.
Marker based registration
To address the time consumption of the manual method, marker-based methods were introduced. Hololens offers access to recorded videos and images from the front-facing camera, along with the location of the camera in the real world and the lens model of the camera. This information can then be used to identify markers using computer vision techniques and locate them using the camera coordinates. Several papers discuss the use of marker-based registration.
[FJDV18] used Vuforia's feature detection algorithm to extract features from the input image from the camera and compare them to stored features for registration. The tracking was further enhanced by placing a known RGB cylindrical object in the environment. The paper also assumed prior knowledge of the transformation between the hologram and phantom during manual registration, using this information when performing automated registration using Vuforia.
[MMGMGS+18] suggested a method to create 3D printer-generated patient-specific tools that are attached with markers. The location of the marker and the known dimensions of the patient-specific tool help in the registration process.
[AJU+18] introduces multimodality markers that can be detected using X-ray and RGB imaging devices. The key advantage of this method is that because the marker can be detected in X-ray, it can become part of the 3D model created using the segmented preoperative CT scan, eliminating the need to add it externally, which can be a source of error. [GCJ+20] followed a similar approach, using a series of adhesive optical codes on the patient's skin prior to the CT scan. These optical codes are visible on CT and MRI scans, allowing for automatic registration with good accuracy that responds to patient movement.
Use of such markers, however, introduces several problems. In particular, marker based registration interferes with the sterility required in surgical fields, and also requires lengthy processing to provide the registration. Alternative markerless techniques can thus be preferred in some scenarios.
Surface Registration
Surface registration is a method of registration without the use of any manual intervention or markers. The objective of these methods is to utilize the surface properties of two 3D point clouds for the purpose of registration. In this project, the utilization of these algorithms is feasible because of the depth camera feature provided by Hololens which can be used to create a 3D point cloud for the patient for registering it to the 3D model generated using preoperative scans. There are multiple approaches to performing 3D point cloud registration, and we will discuss them in the next section.
Traditional ICP-based Methods
ICP, or Iterative Closest Point [BM92], is one of the first methods introduced for registration. ICP is an iterative process where given two sets of points of the same object, we first find the correspondences between these two sets using a distance-based optimization function. Once we have the corresponding points in both sets, we use them to find the rotation matrix and translation vector. Then we use these to move the source points to their new locations. This process is continued until convergence. However, there are a couple of issues with this algorithm. This method is quite prone to get stuck in local optima, and it assumes that the source and target point sets have all the same points. To overcome these issues, several improved versions of the algorithm have been devised.
[ZSN03] introduces the Picky-ICP method, which offers a couple of improvements over the original ICP algorithm. First, Picky-ICP uses hierarchical point selection. We choose every 2h-th data point and run ICP until convergence, where h+ 1 is the hierarchy level. After convergence, we move to the next hierarchy. This speeds up the computation time. The second improvement is identifying points present in more than one pair and selecting only one such pair, the one with the shortest distance. This makes the method more robust to noise and outliers.
[HLY+17] introduces GF-ICP, which utilizes the geometric features of the point clouds to be registered, such as curvature, surface normal, and point cloud density. It uses these features to search for the correspondence relationship between the two point clouds and introduces the geometric features into the error function to achieve accurate registration of the two point clouds.
[YLCJ15] introduced GO-ICP to deal with the issue of ICP getting stuck in local optima. GO-ICP uses the Branch and Bound theory for finding the global optimum. Although this algorithm ensures a globally optimal solution, there is a compromise in computation time. FGR or Fast Global Registration [ZPK16] attempts to solve the compute time issue by devising a joint objective function for global registration by introducing a robust penalty factor to get rid of spurious points removing the need to recompute correspondences during optimization.
Deep Learning Methods
With the latest advancements in Deep Learning, new efforts have been made to replace and add Deep Learning methods to the 3d registration problem.
The pioneering paper that introduced deep learning with point clouds was PointNet (2017) [QSMG17]. This paper discusses the extraction of features from a given point cloud, which can be applied to various tasks such as classification and segmentation. Although PointNet does not specifically address the problem of registration, it provides valuable insights into handling unstructured data like point clouds. One of the key concepts discussed in the paper is the importance of utilizing transformation invariant features. In a given point cloud, each point has only one associated information, namely its coordinates. However, there exist N! point clouds that represent the same information for a given set of N points. To treat all these point clouds as equivalent, transformation invariant features are necessary. PointNet addresses this by employing a symmetric function that yields the same output regardless of the input order. Additionally, the paper emphasizes the combination of local and global features. Relying solely on local features across each point is insufficient as it fails to capture the global context. Therefore, the concatenation of both local and global features offers a better representation of the overall point cloud.
One of the main limitations of PointNet is its inability to capture local geometry or local structure in the coordinate space. PointNet treats each point independently, without exchanging information among neighbouring points, which is crucial for understanding the local geometry. To address this issue, PointNet++ (2017) [QYSG17] introduces a CNN- like model specifically designed for point clouds. In a CNN, nearby pixels are grouped together hierarchically to extract high-level features at each level. PointNet++ adopts a similar approach by grouping nearby point clouds using overlapping partitions. The process begins with Farthest Point Sampling (FPS) to select centroids, and then K-nearest neighbours (KNN) are used to create partitions around these centroids. Similar to CNN, these sets of points are fed into a feature extractor, which in this case is a PointNet. A shared PointNet is applied to each partition. This process of sampling and feature extraction is repeated iteratively to ultimately obtain a global representation.
PointNetLK (2019) [AGSL19] introduces a Deep Recurrent Neural Network that utilizes the features extracted by PointNet and incorporates the Lucas Kanade (LK) approach for calculating the transformation matrix. Lucas Kanade is a classical computer vision differential method used for estimating optical flow. It tackles the optical flow problem by minimizing the discrepancy between the transformed source image and the target image, which is analogous to the 3D registration problem. PointNetLK takes the source and target point clouds and applies MLP and PointNet to extract features from them. The Lucas Kanade method is then employed to determine the transformation that minimizes the distance between the source and target features. The entire process is iterative, with the
model being applied in a loop until the change in the transformation matrix obtained from the LK method falls below a specified threshold.
DCP or Deep Closest Point (2019) [WA19a] employs DGCNN (Dynamic Graph Convolutional Neural Network) instead of PointNet for feature extraction because PointNet does not consider local neighbourhood information, as discussed earlier, which is crucial for efficient feature matching, a critical step in DCP. After feature extraction, DCP utilizes a transformer with cross-attention to facilitate information exchange between the source and target. These extracted features are then used to create soft mappings between the source and target. Soft mappings are preferred over hard mappings due to their differentiability. The final step involves employing the Singular Value Decomposition (SVD) of the obtained soft mappings to derive a closed-form solution for the optimization problem. Another important concept discussed in the paper is the combination of DCP and ICP. One of the main limitations of ICP is its susceptibility to local optima based on initialization. DCP, being a global algorithm that does not require initialization, generates a reasonably accurate solution, which can be further refined using ICP. In this case, ICP will converge towards the global optimum.
PRNet or Partial Registration Network (2019) [WS19b] follows an iterative approach similar to ICP, where it first identifies corresponding points and then calculates the transformation matrix. This iterative process continues until convergence is achieved. PRNet incorporates an additional step of keypoint detection to address the issue of partial overlaps between the source and target. Keypoint detection helps filter out relevant points while eliminating outliers and noise. In terms of creating the matching matrix, PRNet utilizes Gumbel- Softmax instead of the traditional softmax employed in DCP. The advantage of using Gumbel-Softmax is that it yields sharper mappings while maintaining differentiability. PRNet also introduces a temperature term in the Gumbel-Softmax, which can be adjusted to control the sharpness of the matches. Instead of treating the temperature term as a hyperparameter, it is learned as a model variable based on the alignment of two shapes.
DeepVCP (2019) [LWZ+19] operates on source and target point clouds, as well as a prior transformation. It employs PointNet++ to extract local geometric features from the source. These features undergo a weighting stage and are filtered to focus solely on key points. To address the challenge of no corresponding points between the source and target, DeepVCP introduces a novel method called Corresponding Point Generation. This approach takes the source features and applies the prior transformation matrix to transform them. The transformed points are used to generate virtual corresponding points (VCPs) by creating 3D grid voxels around them. These VCPs are then utilized to compute the correspondence matrix, followed by applying SVD to obtain the transformation matrix.
RPM or Robust Point Matching (1998) [GRL+98] introduced the concept of generating soft assignments as a solution to handle outliers and noise.
The equation above describes the method of creating a soft correspondence between the source point cloud ( ;) and the target point cloud (yfe). p is the temperature term that controls the softness of the assignment while a is used to deal with the outliers. RPMNet (2020) uses a similar approach but with a couple of changes [YL20].
Firstly, the equation replaces the coordinates with the extracted feature values, as shown in the above equation. Additionally, a secondary network is incorporated to learn the values of a and p. Another innovative approach introduced in this paper involves using not only the coordinates but also the relative coordinates of neighbouring points, along with 4D point pair features (PPF), as inputs to PointNet for feature extraction.
DeepGMR (2020) [YEK+20] utilizes a probabilistic formulation, which enhances its resilience to noise and outliers. The primary objective of this work is to represent a 3D point cloud as a Gaussian Mixture Model (GMM) and subsequently match the distributions between the source and target models to obtain the transformation matrices. The method consists of three main components. Firstly, it involves determining point-to-component correspondences. This step establishes the associations between points and different components of the GMM. Next, these correspondences are utilized to obtain GMM parameters and transformation matrices. Overall, DeepGMR aims to leverage the probabilistic representation of point clouds using GMMs to facilitate robust matching and transformation estimation.
RGMNet (2021) [FLLW21] employs a graph-based approach for feature extraction and subsequent transformation matrix estimation. The network utilizes a transformer-based edge generator, which utilizes the features of the point cloud to generate edges. This is followed by a graph feature extractor that operates independently on both the source and target graphs. The extracted features are then utilized to calculate the affinity between the two graphs, leading to the creation of a correspondence matrix. The updated values of the transformation matrix are obtained through the singular value decomposition (SVD) of the correspondence matrix. The graph-based methodology adopted by RGMNet is particularly effective in addressing outliers, enhancing the robustness of the registration
process.
Predator (2021) [HGLT21] employs an encoder-decoder model for point cloud registration. The encoder component downsamples both point clouds using KPConv-FPN, resulting in a set of super points. These super points are then utilized in a graph neural network (GNN) to extract features for both point clouds. The extracted features from the encoder are combined and fed into the decoder, which performs an upscaling operation to generate matching scores. This model is similar to RGMNet, with the key difference being the employment of downscaling and upscaling operations. By using super points, Predator enhances the algorithm's robustness to noise and outliers, improving the accuracy of point cloud registration.
GeoTransformer or Geometric Transformer (2022) [QYW+22] follows a similar principle to Predator, leveraging KPConv-FPN to obtain a set of super points from the input point clouds. However, instead of employing a graph neural network (GNN) for feature extraction, it utilizes a Geometric Transformer. The GeoTransformer starts with a self-attention mechanism to capture local dependencies within each point cloud. Then, cross-attention is performed between the source and target point clouds to facilitate the transfer of information and establish point correspondences. These correspondences are subsequently used to compute the transformation matrix for each pair of matched superpoints. To determine the best global transformation, the method evaluates the performance of different transformations on a global scale and selects the one that achieves optimal results.
OMNet (2021) [XLW+21] addresses the challenge of partial overlaps in point cloud registration by employing a mask prediction approach. The method predicts overlapping masks that indicate the regions of overlap between the source and target point clouds. By utilizing these masks, non-overlapping points are discarded, and only the relevant points are considered for calculating the transformation matrix. A unique aspect of OMNet is its use of regression instead of singular value decomposition (SVD) to directly compute the transformation matrix. This regression-based approach provides an alternative method for estimating the transformation parameters. However, one limitation of OMNet is that it lacks a proper exchange of information between the source and target point clouds during the feature extraction process. This may result in inaccurate mask predictions and potentially impact the overall registration accuracy [XYL+22].
FINet (2022) [XYL+22] introduces a dual branch structure for handling rotation and translation separately in the point cloud registration process. This separation is motivated by the fact that translation belongs to Euclidean space, which is not directly correlated
with the quaternion space typically used for representing rotation. By employing separate branches for rotation and translation, FINet effectively handles both aspects of the transformation. Additionally, FINet incorporates a multi-level feature interaction mechanism between the source and target point clouds. This approach serves as an efficient alternative to traditional cross-attention mechanisms, reducing the computational and memory requirements significantly while still enabling effective information exchange between the two point clouds. These features of FINet contribute to its robustness and efficiency in point cloud registration tasks.
YOHO (You only hypothesize once) (2022) [WLDW22] introduces the concept of rotation equivariant descriptors for point cloud registration. The primary objective of using a rotation equivariant descriptor is to capture variations in the descriptor corresponding to different rotations, which can be utilized to estimate the amount of rotation between point clouds. YOHO achieves this by defining feature maps on icosahedral groups with 60 rotations. By employing such a framework, the descriptor is designed to exhibit rotational symmetry and provide consistent representations for points under different rotations. To make the descriptor invariant, average pooling is applied, enabling the extraction of features that are robust to rotation. Once a pair of matched points is identified, the equivariant descriptor of these points can be utilized to estimate the rotation between the point clouds. YOHO leverages this rotation information to enhance the accuracy of point cloud registration. By introducing rotation equivariant descriptors, YOHO contributes to the advancement of rotation-aware feature extraction and improves the quality of point cloud registration results.
GMCNet (Graph Matching Consensus Network) (2022) [PCL21] adopts a similar approach to other methods in terms of feature extraction, correspondence matrix creation, and transformation matrix estimation. However, the key contribution of the paper lies in its focus on rotation invariant feature extraction. To achieve rotation invariance, GMCNet incorporates hand-crafted rotation invariant features into the feature extraction process. In addition to rotation invariant features, GMCNet integrates multiscale smoothness terms derived from the geometric structures at various scales. These smoothness terms help to capture the local and global geometric characteristics of the point cloud, further improving the accuracy and reliability of the transformation matrix estimation. By combining rotation invariant features and multiscale smoothness terms, GMCNet contributes to the development of more robust and accurate point cloud registration algorithms.
VPRNet (Virtual Point Registration Network) (2022) [LYLG22] introduces a unique approach to address the challenge of partial to partial registration by utilizing a Generative Adversarial Network (GAN) to generate missing points in both the source and target point
clouds. This technique aims to enhance the registration process by providing complete point clouds for better alignment. VPRNet consists of two main phases. The first phase involves the use of a model called VPGNet, which is responsible for generating new points in the partial point clouds. By leveraging the power of GAN, VPGNet generates plausible virtual points that fill in the missing regions of the source and target point clouds. This step aids in achieving more comprehensive coverage and reducing the impact of partial overlaps during the registration process. The second phase of VPRNet referred to as Regnet, focuses on the actual registration process. It creates a correspondence matrix between the generated virtual points and the original points in the source and target point clouds. By establishing point-to-point correspondences, Regnet then employs SVD (Singular Value Decomposition) to compute the transformation matrix that aligns the two point clouds. The combination of the GAN-based point generation in VPGNet and the subsequent registration using the correspondence matrix and SVD in Regnet contributes to the effectiveness of VPRNet in addressing partial to partial registration challenges.
VRNet (2022) [LYLG22] introduces the concept of Rectified Corresponding Points (RCP) for point cloud registration. RCPs have the same shape as the source point cloud but share the pose of the target. By estimating the pose between the source and RCPs, the same transformation matrix can be obtained when aligning the source with the target directly. Since the source and RCPs have the same shape, registration becomes straightforward as each point in the source has a corresponding point in the RCP. RCPs are derived through a rectification process based on the concept of Virtual Corresponding Points (VCP) from the DeepVCP method. VRNet leverages this approach to achieve accurate and efficient point cloud registration.
Contributions
Surgical scene 3D reconstruction
We explored an innovative approach for generating a point cloud of a real-life scenario, capitalizing on the 'long-throw' depth camera and gesture recognition of Hololens to accomplish 3D point cloud reconstruction of the patient. The point cloud obtained from this method helped in obtaining more detailed information about the geometry of the patient which helps in the overall registration of the point cloud.
Creation of a customized Phantom Model
To assess the Hololens' effectiveness in 3D reconstruction and registration, we established a phantom model. This model was created using the 3D preoperative data of a real patient. Throughout this project, we employ this phantom model to capture a 3D point cloud using
the Hololens. We then use this captured point cloud to evaluate the performance of our model by conducting registration between this point cloud and the one derived from the 3D preoperative model of the patient. We chose this specific region of the human body for our initial experiments because it experiences minimal alteration due to human breathing compared to other regions. Accounting for movement in these other regions during registration could be explored in future work.
Validation of 3D Registration Models
Various registration methods have specific scenarios where they perform optimally. Several Deep Learning-based models excel in a global space, while traditional approaches like ICP operate effectively in local spaces. Our approach involves combining these diverse registration algorithms to enhance overall accuracy and performance.
Patient Specific Data Generation for fine-tuning the Registration Model
Given the absence of an existing dataset tailored to our specific use case, which revolves around point clouds of patients in the operating theatre, we have adopted an approach to enhance our model's accuracy. To achieve this, we manipulate the 3D preoperative mesh to create a simulated dataset depicting a patient lying on a bed and adding noise to simulate the presence of different objects near the patient, which we then employ for model fine-tuning.
Mixed Reality Visualization for Surgical Guidance
This project revolves around the development of an end-to-end application using Hololens and Registration method to aid surgeons during their surgeries. The application consists of several individual components, which are then integrated to form the final application. In this section, we propose a pipeline that can be utilized to efficiently execute Registration on Hololens. The example system diagrams of the pipeline can be seen in Figures 1 and 2.
In this pipeline, we first construct a 3D Hologram model using the patient's CT scans and use it to create a simulated training set to fine-tune our Deep Learning model for this specific patient. Then, when the surgeon wears the Hololens during surgery, the initial step is to create a point cloud of the patient using the Depth Camera in the Hololens. This point cloud is then sent directly to the server along with the 3D Hologram. Afterward, both the patient's point cloud and 3D Holograms undergo some preprocessing before being used as inputs to the registration model. The registration model performs the registration and returns the Rotation and Translation Matrix back to the Hololens. This transformation is subsequently applied to the Hololens to achieve the registration. This process can be
executed at constant time intervals to accommodate changes in the scene, such as if the patient is slightly moved.
3D Pre-Operative Hologram Construction using Segmentation
The first step of the pipeline is the construction of 3D Hologram using patients' preoperative data like CT-scans. This 3D Hologram eventually will be registered on the patient allowing surgeon access to internal information directly on the patient using Hololens.
The initial phase of the project focused on constructing a 3D pre-operative Hologram using CT scans, although in other embodiments any imaging modality may be used - nuclear medicine scans, MRI, SPECT, and SPECT CT, for example. . Various segmentation techniques were applied to the CT scans captured at different layers to isolate specific areas of interest, including the skin, tissues, bones, vessels, and lymph nodes. These segmented images from different layers were then integrated to generate a 3D mesh, serving as the source model for this project. Figure 3 displays different layers of the constructed 3D model. Further details regarding this process will be provided in Section 5.
Registration Performance improvement using Simulated Data
After obtaining a 3D Hologram, our goal is to refine our registration model to achieve better registration performance on this specific training instance. To accomplish this, we start by simulating the training data using various methods such as cropping the 3D Hologram and adding noise to create a target point cloud. Once this simulation process is complete, we proceed to fine-tune our trained Registration model using this simulated data.
3D Point Cloud Extraction with Hololens Depth Camera
When in use, our initial step is to extract the 3D point cloud of the patient within the operating theatre. This process involves utilizing the depth camera integrated into the Hololens. The depth camera captures and extracts the point cloud data, which is then transmitted to the server for further processing.
Preprocessing Point Clouds: Enhancing Data Quality
After obtaining both the 3D Hologram and the 3D scene point cloud, we subject them to preprocessing before they undergo registration model processing. This preprocessing encompasses various tasks, including subsampling a predetermined number of points from each cloud, converting mesh data into a point cloud format in cases where the input is in mesh form, or computing point normals for different points within the point cloud. The
calculation of normals is particularly crucial, as it generates an essential input feature vector for the registration model.
Deep Learning Registration for Hololens: Optimizing Registration Performance
The preprocessed point clouds are subsequently passed through the Registration model, which is deployed on the server. This model takes the point clouds as input and produces the rotation and translation matrices. These matrices are then applied to the 3D Hologram to achieve precise registration on the Hololens.
The CPU in HoloLens, Qualcomm Snapdragon 850 Compute Platform, is specifically designed to address the unique demands of augmented reality, which involve tasks like real-time environment tracking and rendering 3D holographic objects. While it may not provide the same level of raw computational power as high-end desktop CPUs, it excels in its intended functions. On the other hand, Deep Learning models typically require significantly higher computational power and memory, often relying on GPUs or TPUs for efficient execution. Therefore, it is more practical to set up the Deep Learning-based Registration model on a remote server and utilize REST API calls from HoloLens to the server for obtaining the transformation matrix for the Hologram.
In the following sections we will be discussing different methods which are essential in developing this pipeline.
3D Pre-Operative Hologram Construction using Segmentation
The first step of the whole process is to generate a 3D model for registration using layerwise preoperative CT scans and then combining them together after segmentation. Here image acquisition was first performed after pseudonymization of patient information. These were subsequently transferred into 3D process software for segmentation. The 3D model can be seen in Figure 3.
Segmentation of layer-wise CT scan images
Segmentation of images involves processing different layers of CT scan images to identify different anatomical segments using a software such as 3D Slicer, or TotalSegmentator (available before the priority date at https://totalsegmentator.com/). The focus was on the following anatomical segments:
Segmentation of images involves processing different layers of CT scan images to identify different anatomical segments using a software called Slicer 3D. The focus was on the following anatomical segments:
Skin: The skin was segmented using fluid filling of the air space surrounding the patient in scan images, then inverting the selected space into the patient and finish with an algorithm to only segment the shell which interface between the patient and the air surrounding.
Soft tissue: For segmentation of the soft tissue, Threshold algorithm was used with masking limitation to include areas which has not been segmented. The mean Threshold value were between (-181 and -34) for this layer. No further smoothing was needed during segmentation.
Muscles: Threshold algorithm was sufficient in segmenting the muscle layer with an additional need for smoothing. The mean threshold value were (-4.42 and 111.82). There was some overlap for the threshold range with the low intensity bone medulla this was minimized using masking and was not evident in the final 3D reconstruct. The significant overlap and interference were with the intra-pelvic structures (urinary bladder, rectum, small bowel, and vessels) understandably due to the presence of smooth muscle in most of these structures and similar intensity. Although overlap was significant the proposed primary clinical application doesn't require detailed intra-pelvic anatomical accuracy. This, however, proves the challenges facing segmenting with some algorithms especially in the complex intraabdominal and intrapelvic structures.
Bone: Threshold is an excellent tool for bone segmentation as the highest range of intensity threshold for bone always correlate to the highest intensity in the scan. There are some limitations within segmenting the bone medulla, however, the cortex can be segmented very efficiently which is more relevant clinically and is sufficient for visualization of bone contours. The mean thresholding values were between (122.46 and 1677). Heavy smoothing was needed in porous bone areas like the iliac wings but less required in solid bone like femur shaft.
Vessels: Vessels were the most challenging in segmenting. Primarily due to the nature of the SPECT-CT scan which doesn't use contrast, and this limits the algorithms and tools that can be used in segmenting. In this respect, a SPECT-CT scan is a type of nuclear medicine scan where the images or pictures from two different types of scans are combined together. Also, the complex nature of the vascular tree adds to the segmenting challenging. Best tool used was "grow from seed". Although time consuming it yield the best results in the final hologram. Therefore, there is significant role for machine learning and advanced automated algorithms to help segmenting this layer and improve segmenting for smaller vessels.
Lymph node: The segmentation of lymph nodes was carried out manually. Due to its relatively small size and critical significance, automating this task would not necessarily reduce time consumption. Moreover, manually segmenting lymph nodes ensures a high level of precision in the segmentation process. This precision could potentially pave the way for the development of automated segmentation methods in the future.
In the segmentation pipeline, most of the segmentation tasks are automated and can be accomplished using thresholding techniques. However, the localization of the lymph node requires a higher level of precision and needs a manual intervention from the surgeon. The use of TotalSegmentator, mentioned above, decreases the time needed for the segmentation process due to its ability to segment up to 104 anatomical structure
Once successful segmentation has been achieved for various layers, the segmented images are then combined to create a comprehensive 3D model. This 3D model serves as a 3D Hologram within the Hololens and functions as the source point cloud for the registration method. Following registration, the surgeon can utilize Hololens gestures to visualize distinct layers of the 3D Hologram. This visualization assists in the localization process, enabling the surgeon to identify various internal anatomical regions relevant to the surgery. In essence, these segmented layers provide valuable guidance during the surgical procedure.
3D Point Cloud Extraction with Hololens Depth Camera
The initial step in using the application during surgery involves constructing a 3D point cloud of the patient, which serves as the target point cloud for the registration model. To obtain the patient's point cloud, we intend to utilize the depth camera within the Hololens. However, there are alternative methods for extracting point clouds from the scene using external sensors or cameras. In this section, we will examine 3D reconstructions of the Phantom model generated using various methods and then compare them to the point cloud extracted from the Hololens depth camera.
Methods
3D Point Cloud Extraction Using: External RGBD Camera
The first approach is to use an external RGBD camera to obtain a depth image of the scene and then convert it into a 3D point cloud. The RGBD camera used for this purpose is the Intel Depth Camera D435. This camera operates with an error rate of less than 2% at a distance of 2 meters.
While the geometry of the point cloud is clear from the front view, the side view is less accurate. For instance, in Figure 4, on the bottom right, the boundary between the phantom model and the table appears quite blurry and does not provide accurate geometry information, which affects the accuracy of the overall registration model. Another issue with using this sensor is, because the best operating point of this camera is around 2 meters from the object it covers a lot of area apart from just the patient. This extra information acts as noise during registration and would affect the model's performance.
3D Point Cloud Extraction Using: Photogrammetry
Another approach to generate a 3D point cloud is by using photogrammetry. The images from this approach can be seen in Figure 5. In this method, multiple images are captured from various angles and then stitched together by identifying common features in different images. To execute this, we employed an application called Polycam, which records a video of the object and then uses the different frames of the image to perform photogrammetry.
The quality of the point cloud generated through this approach is notably superior to the one generated by the RGBD camera. However, one drawback of this method is that, to create a high-quality point cloud, we need to capture images from all around the object, and there is an additional time investment required by photogrammetry for stitching all the images together. This can introduce some delay in the overall process.
Another challenge associated with this method is that the absence of camera coordinates when using images can lead to discrepancies in the actual dimensions of the extracted point cloud. Camera coordinates are a critical element in stereo rectification, particularly in photogrammetry. As a result, manual rescaling of the point cloud becomes necessary to restore its original dimensions, potentially introducing additional delays to the process.
3D Point Cloud Extraction Using: Hololens Spatial Mapping
Spatial mapping provides a detailed representation of real-world surfaces in the environment around the HoloLens, enabling developers to create convincing mixed reality experiences. The images from this approach can be seen in Figure 6. The HoloLens utilizes its depth camera to construct surface mapping. We developed a HoloLens application that allows users to define a bounding box around the object of interest and extract the spatial mapping created by within that region.
3D Point Cloud Extraction Using: Hololens Depth Camera
HoloLens 2 introduces a research mode that provides users with access to various sensors within the HoloLens, including a 'long-throw' mode depth camera. As mentioned above,
we developed an additional HoloLens application that utilizes this depth camera to extract point clouds within a specified region, which can be defined using a bounding box. The images from this approach can be seen in Figure 7.
The key difference between spatial mapping and this approach is that the point cloud created using spatial mapping is optimized for easy integration into various types of mixed reality applications. However, if we opt for raw point cloud data without any optimization, we can obtain a much more detailed point cloud, which is a significant requirement for this project since deep learning approaches rely on comparing geometric features in both the source and target point clouds. Furthermore, the quality of the point clouds improves as the user moves around the observed object.
Comparison of Point extracted from each methods
We perform a comparison based on the quantity of points or vertices extracted from all the methods, with a higher point count resulting in clearer geometric features.
Method Mesh Vertices
RGBD Camera 77,754
Photogrammetry 9,974
Hololens Spatial Mapping 525
Hololens Depth Sensor 977,566
Conclusion
The point cloud produced by the Hololens Depth Sensor provide the most information about the geometry of the 3D object. One additional advantage of using Hololens, as opposed to an external camera, is that the locations of the camera sensors for Hololens are already known, eliminating the need for recalibration. Furthermore, Hololens offers the convenience of a bounding box feature, allowing us to focus on a specific area rather than capturing all nearby objects, which may not be possible with other methods.
Registration Performance improvement using Simulated Data
The benefit of this use case (being sentinel Lymph node biopsy surgery, but also more generally nuclear medicine and tumour specific probes) is that we have access to the CT scan of the patient 2 hours before the surgery. We can use the 3D segmented model created from this CT scan to create simulated training data to fine-tune the model further
so that the model performs even better in the actual scenario. In this section, we will talk about various techniques used to simulate the data.
Sequential Procedures for Generating Simulated Training Data
The flow diagram of the data simulation pipeline can be seen in Figure 8.
Transformations: Rotation and Translation
Initially, we treat the source and target point clouds as identical. Subsequently, we generate a target point cloud by applying rotations and translations within a predefined range, ensuring that these transformations fall within the bounds where RPMNet maintains acceptable accuracy. This simulation is designed to mimic a scenario where both the source and target point clouds pertain to the same object but are positioned at distinct locations. The objective is to challenge the registration algorithm to align them accurately. The example images relating to rotation and translation can be seen in Figure 9.
Handling Partial Visibility: Occlusion Simulation
After rotating and translating we deal with the partial visibility by occluding a certain part of the target point cloud. We randomly truncate the target point cloud using a plane, mimicking real-world scenarios where the patient would typically be on an operating table, granting us only a limited view of the surgical area compared to the complete 3D preoperative model. The example images relating to the occlusion simulation can be seen in Figure 10.
Introducing Realism: Adding Bed Surface
The 3d point clouds generated by the Hololens suggest that along with the patient we can also get the points corresponding to the bed where the patient lies. In order to simulate this we add a plane at the point where we truncate the point cloud. The images relating to adding the bed surface or the plane can be seen in Figure 11.
Enhancing Model Robustness: Introducing Point Jitters
To account for variations in the contour of the patient's body, we introduce perturbations to each point using Gaussian noise. The standard deviation of this noise can be adjusted to control the level of perturbation. This approach enhances the model's ability to generalize and improves registration performance.
Deep Learning Registration for Hololens: Optimizing Registration Performance
After generating the 3D point clouds, we will now delve into different methods for 3D point cloud registration. Registration involves utilizing feature matching to determine the transformation matrix between point clouds of the same object. The primary objective of this project was to explore marker-less registration, thus our focus is on Deep Learningbased algorithms. Numerous algorithms capable of achieving this task exist, as discussed in the Literature Review section. In this study, we will specifically examine two Deep Learning-based models, RPMNet and PREDATOR. We chose these two algorithms due to their performance on popular datasets and the availability of pre-trained models.
Registration Problem Formulation
Given a Source Point Cloud X = {x7 e IR3 | j
and a Target Point Cloud Y =
{yfe e IR3 | k = the main objective of the deep learning registration method is to obtain the Rotation matrix, R e SO(3), and a Translation Matrix, t e IR3, that align the two point clouds. In our use case, the 3D pre-operative model will be the source point cloud, and the 3D reconstructed model of the patient will be the target point cloud. The Registration model will provide the Rotation matrix, R, and translation matrix, t. We then apply these matrices to the 3D pre-operative model to align it with the patient.
Overview of Registration Methods
RPMNet: A Deep Learning-Based Registration Model
The example deep learning-based registration model e.g. RPMNET, can be seen in Figure 12. Figures 13-15 provide further details relating to parts of the deep learning-based registration model such as the feature extraction part (Figure 13), the parameter prediction part (Figure 14) and the Compute Match Matrix (Figure 15).
Model Description of RPMNet
RPMNet (2020) [YL20] is a Deep Learning-based Registration model. The core of RPMNet is inspired by RPM or Robust Point Matching (1998) [GRL+98], which introduced the concept of generating soft assignments as a solution to handle outliers and noise.
The equation above describes the method of creating a soft correspondence between the source point cloud (x7) and the target point cloud (yfe). p is the temperature term that controls the softness of the assignment, while a is used to deal with outliers. RPMNet deploys this method with a difference: it uses extracted features for comparison instead of just point coordinates, as shown in the following equation.
In order to extract these features, RPMNet applies PointNet to a 10D feature descriptor constructed for each point, which is made up of xc,{zlxCi£},{PPF(xc,x£)}). Axc l denotes the neighboring points translated into a local frame by subtracting away the coordinates of the centroid point:
Axc £ = x — xc.
PPF(XC,X£) are 4D point pair features (PPF) that describe the surface between the centroid point xc and each neighboring point x£ in a rotation invariant manner:
Unlike RPM (1998), a and /? are learned in RPMNet by passing the two point clouds through a neural network. After feature extraction, the matching matrix is calculated, illustrating the correspondence between the two matrices.
Now that we have the source point cloud and its corresponding points in the target point cloud, RPMNet uses differentiable weighted Singular Value Decomposition (SVD) to calculate the transformation matrix. Weighting is done to account for the fact that not all source points might have corresponding target points. RPMNet is an iterative algorithm, so this process is repeated for the desired number of iterations. At each iteration, the source point cloud is continually updated.
Loss Function in RPMNet
Loss function for RPMNet has two components, first one is a LI loss between source point cloud transformed using ground truth Transformation matrix and the one transformed using estimated Transformation matrix. RPMNet also has another component to increase the number of inliers by estimating a secondary loss on the match matrix.
The overall loss is the weighted sum of the two losses:
^total — -^reg + ^inlier
PREDATOR: Encoder-Attention-Decoder Model
An example system diagram of an encoder-attention-decoder model can be seen in Figure 16.
Model Description of PREDATOR
Predator (2021) [HGLT21] consists of three main components: the Encoder, Attention module, and Decoder. The Encoder downsamples both point clouds using KPConv-FPN, resulting in a set of super points. KPConv-FPN utilizes ResNets and Convolutions to aggregate point clouds into super points.
Next, Predator implements Graph Neural Networks (GNN) for both sets of super points to extract contextual information. To do this, the super points are initially connected together to form a graph using k-Nearest Neighbors (KNN) in the feature space. Subsequently, the encoder features are iteratively updated using the following equation, where ;- represents the neighboring points of xt, and hg is a linear layer followed by normalization and leaky Re LU:
This process is repeated two times, and all the outputs are combined to obtain a more contextual representation of the super points:
Subsequently, a cross-attention module is implemented to facilitate the transfer of information between both point clouds, which is crucial for learning the overlap regions and matching between the point clouds. The architecture used for this module is the same as in the transformer model. The extracted information is then combined with the previous feature values of the super points to obtain a co-contextual feature representation.
Local context is updated using the gathered cross-point cloud information by applying another GNN with the same structure, resulting in latent feature space encodings Fx and Fy. These scores are then used to calculate the overlap score between the super points. All this extracted information is passed through the Decoder to obtain matchability scores,
overlap scores, and per-point features for both point clouds. Finally, RANSAC is used to calculate the final transformation matrix.
Loss Function in PREDATOR
Predator incorporates three components into its loss function: Circle Loss, Overlap Loss, and Matchability Loss.
Circle Loss fine-tunes the point-wise feature descriptions by leveraging the distances between corresponding points in both point clouds in feature space.
Overlap Loss and Matchability Loss are binary loss functions designed to train the Overlap score and Matchability score extracted from the model.
Iterative Closest Point (ICP)
ICP, or Iterative Closest Point [BM92], is one of the first methods introduced for registration. ICP is an iterative process where given two sets of points of the same object, we first find the correspondences between these two sets using a distance-based optimization function. Once we have the corresponding points in both sets, we use them to find the rotation matrix and translation vector. Then we use these to move the source points to their new locations. This process is continued until convergence. This method is quite prone to get stuck in local optima.
One of the methods discussed in the literature to prevent ICP from becoming stuck in local optima is to combine it with a deep learning method that operates more effectively in a global space. Any refinements to the results obtained can then be accomplished with the assistance of ICP. In our analysis, we will incorporate ICP into both RPM-Net and PREDATOR to assess the improvements it brings.
Performance Comparison of Registration Models
Experimental Setup for Model Comparison
To determine which registration model best suits our use case, we've designed a series of experiments that progressively increase in complexity, gradually approaching the real- world scenarios we anticipate during usage.
To test these registration models, we need to create both a source point cloud and a target point cloud. This is achieved by taking a 3D model from our dataset as the source point cloud and then applying various operations such as rotation, translation, cropping, and adding noise to generate the target point cloud. Subsequently, we evaluate the
performance of all four registration models— RPMNet, RPMNet+ICP, PREDATOR, and PREDATOR+ICP— on these inputs to assess their effectiveness in different scenarios.
Each model provides the rotation and translation matrices used to align the source point cloud with the target point cloud. To evaluate their performance, we employ two key metrics for both rotation and translation matrices: Mean Square Error (MSE) and Mean Absolute Error (MAE). These metrics allow us to compare the matrices extracted from the models with the ground truth matrices used to create the target point cloud initially.
MSE (Mean Square Error) is calculated as: n
MSE = iy (Xt - Y n —i
1=1 where X£ and V are the elements of the ground truth matrix and the corresponding extracted matrix, respectively, and n is the number of elements in the matrix.
MAE (Mean Absolute Error) is calculated as:
where Xt and Yt are the elements of the ground truth matrix and the corresponding extracted matrix, respectively, and n is the number of elements in the matrix.
Dataset: ModelNet 40
ModelNet 40 is a dataset consisting of CAD models from 40 categories. In this project, we evaluate the performance of RPM-Net, PREDATOR, RPM-Net + ICP, and PREDATOR + ICP on various categories within ModelNet 40, with a particular focus on the 'human' category due to its relevance to our specific problem. Example images from the ModelNet 40 dataset can be seen in Figure 17.
Experimental Scenarios for Model Comparison
Now that we have chosen two models for our project, the aim of this section is to determine which method is best suited for our needs. Our intended scenario for using the model is an operating theater where we have only a partial view of the patient. The patient will also be surrounded by other elements in the operating theater, such as other surgeons and a hospital bed, which can affect the 3D reconstruction. To simulate this exact scenario during the testing of these models, we begin with simple experiments and gradually introduce
complexity at each step, such as adding noise, providing a partial view of the target point cloud, and introducing jitters to the points of the target point cloud. In this section, we focus on the 'human' class of the Model-Net 40, which consists of 40 point clouds representing human figures.
Clean Data Scenario
We initially test on a simple scenario in which the target and source are identical, with the only difference being that the target point cloud is a rotated and translated version of the source point cloud. All the methods gets the same two point clouds as source and target but target is just a transformed version of the source. We then take the Rotation and Translation matrix which we obtain from both these models and compare them to the ground truth.
The results in the plot show a significantly better performance for RPMNet and RPMNet + ICP as compared to Predator and Predator + ICP. Another important observation from this plot is that including ICP reduced the error further in case of RPMNet so it is a good addtion.
Method Rotation MSE Rotation MAE
RPMNet 0.0015 +/- 0.0014 0.0277 +/- 0.0104
PREDATOR 0.0249 +/- 0.0902 0.0734 +/- 0.1227
RPMNet+ICP 0.0009 +/- 0.0010 0.0205 +/- 0.0096
PREDATOR+ICP 0.1251 +/- 0.6472 0.0845 +/- 0.3092
Method Translation MSE Translation MAE
RPMNet 0.0006 +/- 0.0007 0.0163 +/- 0.0101
PREDATOR 0.0055 +/- 0.0151 0.0362 +/- 0.0477
RPMNet+ICP 0.0005 +/- 0.0008 0.0126 +/- 0.0117
PREDATOR+ICP 0.0095 +/- 0.0440 0.0332 +/- 0.0891
Figure 18 shows visulations of the registration method applied to a sample 3D model from the ModelNet40 dataset for clean target point cloud.
Clean Partial Data Scenario
In this phase, we introduce an added layer of complexity to the task. The target point cloud is no longer merely a transformed version of the source cloud but is also subjected to partial occlusion. We randomly truncate the target point cloud using a plane, mimicking real-world scenarios where the patient would typically be on an operating table, granting us only a limited view of the surgical area compared to the complete 3D preoperative model.
The RMSE and RMAE plots indicate that while the absolute error values are higher in this case, the relative performance between RPMNet and PREDATOR remains consistent, with RPMNet outperforming PREDATOR by a noticeable margin. Once again, in this scenario, RPMNET+ICP enhances the overall performance of RPMNet.
Method Rotation MSE Rotation MAE
RPMNet 0.0062 +/- 0.0058 0.0538 +/- 0.0280
PREDATOR 0.0335 +/- 0.0564 0.1112 +/- 0.1037
RPMNet+ICP 0.0056 +/- 0.0050 0.0504 +/- 0.0252
PREDATOR+ICP 0.0336 +/- 0.0600 0.0962 +/- 0.1162
Method Translation MSE Translation MAE
RPMNet 0.0080 +/- 0.0075 0.0639 +/- 0.0325
PREDATOR 0.0160 +/- 0.0209 0.0791 +/- 0.0710
RPMNet+ICP 0.0079 +/- 0.0079 0.0606 +/- 0.0351
PREDATOR+ICP 0.0146 +/- 0.0189 0.0691 +/- 0.0629
Noisy Data Scenario
In this step, we introduce noise to the target point cloud. This simulation is critical for replicating real-world conditions within an operating theater, where multiple objects besides the patient may be present, potentially affecting the accuracy of the registration process. To introduce noise, numerous noise spheres are added. The rationale behind this choice is that the algorithms take into consideration the normals of all the points in the point cloud, and spheres encompass a broad range of normals, thus providing sufficient diversity to simulate various objects that might be present in the operating theater.
Once again, in this scenario, there is an overall increase in the error for both methods. However, RPMNet continues to perform reasonably well in achieving accurate registration.
Method Rotation MSE Rotation MAE
RPMNet 0.00595 +/- 0.01082 0.05138 +/- 0.03537
PREDATOR 0.52942 +/- 1.77869 0.31908 +/- 0.61887
RPMNet+ICP 0.00565 +/- 0.00935 0.05063 +/- 0.03512
PREDATOR+ICP 0.50375 +/- 1.69696 0.30295 +/- 0.60245
Method Translation MSE Translation MAE
RPMNet 0.00589 +/- 0.00491 0.05420 +/- 0.03230
PREDATOR 0.05791 +/- 0.13001 0.13345 +/- 0.12710
Method Rotation MSE Rotation MAE
RPMNet+ICP 0.00637 +/- 0.00534 0.05688 +/- 0.03490
PREDATOR+ICP 0.06058 +/- 0.14199 0.13377 +/- 0.13698
Translation Range Experiments
From our observations, it becomes evident that the algorithm's performance diminishes notably when the source and target point clouds are significantly distant from each other. This distance-related disparity is especially pronounced when using RPMNet, where the error increases as the distance between the point clouds grows.
It's worth noting that in these experiments, we normalized the coordinates of the point clouds to a range between 0 and 1. This means that each distance unit corresponds to one dimension of the bounding box across the object in focus.
The result below suggest that the model works with a good accuracy within a distance range of 2 units but after this there is a huge jump in error. This sets a limit in terms of distance upto which the model performs efficiently.
One straightforward solution in our use case is to manually bring the hologram closer to the patient using hand gestures with Hololens. Subsequently, we can employ the deep learning method for the more intricate registration process, thereby mitigating the challenges posed by significant initial separation between the point clouds.
Angle Range Experiment
Another significant variable impacting the performance of the deep learning method is the angle between the source and target point clouds. As this angle increases, all the models experience a noticeable drop in performance. This trend is clearly depicted in the graph below, where, for all the models, an increase in the angle between the point clouds from 30 to 180 degrees leads to a substantial increase in the overall error.
In the case of RPMNet, reducing this error at lower angles (30, 60) can be achieved by increasing the number of points to around 1000, beyond which the error stabilizes. Another consideration for selecting 1000 points is the significant increase in inference time when moving to 10000 points.
For higher angles the error is significantly high even on increasing the number of points. In practical applications, addressing this issue involves a similar approach to how we deal with translation issues. Specifically, we can manually adjust the source point cloud to
ensure it falls within a certain angle range relative to the patient, thereby optimizing the registration process.
Visibility Metric Experiment
Examining the visibility metric is a crucial test to assess how effectively the models operate under varying degrees of visibility. The occlusion is simulated by intersecting a plane through the 3D point cloud at a random location and then removing a portion of it.
The results indicate that RPMNET performs remarkably well even in cases with 0.4 visibility. In typical surgical conditions, a surgeon would have approximately 0.5 visibility of the patient's body, suggesting that the model should perform effectively in practical, real- world scenarios.
Effect of Iterations Experiment
RPMNet operates as an iterative algorithm, where in each iteration, we iteratively adjust the source point cloud to gradually align it with the target point cloud. In this experiment, we investigate various iteration counts to identify the point at which performance reaches its peak. The findings indicate that the algorithm performs remarkably well even with just one iteration, even in the presence of noisy data. This suggests that an extensive number of iterations may not be necessary for the algorithm to converge effectively.
Optimized Registration Model Configuration
The performance of the model in different scenarios suggests that RPMNet outperforms PREDATOR in the scenario closest to the real situation. Apart from this in general RPMNet's accuracy increases when combined with ICP.
One of the other observations in the experiments is that the algorithm accuracy is affected by the translation and rotation between the source and the target point cloud. The accuracy goes down for larger angles and larger distances between the source and target point cloud. But one of the ways to eradicate this issue is to manually move and rotate the target point cloud close to the source point cloud.
Visibility represents another crucial factor to consider when selecting a model. Based on our experiments, we determined that RPMNet and RPMNet+ICP demonstrate strong performance up to a visibility threshold of 40%, which closely aligns with real-world scenarios. This finding underscores the suitability of these models for practical applications
Additionally, our analysis of model accuracy in relation to the number of points in the point cloud revealed that 1000 points suffice for achieving a notably high level of accuracy.
Beyond this point count, there is a trade-off as higher numbers of points result in increased inference time.
In conclusion based on the experimentation above the final model configuration is:
Model RPMNet+ICP
Number of Points 1000 Translation Limit 2 units Rotation Limit 60 degrees Visibility Limit 40%
Registration Iterations 1
Implementation of Point Cloud Registration using Hololens Pipeline
Training Pipeline
The model is trained using simulated data, beginning with the 3D pre-operative model as our source point cloud. We then simulate a target point cloud based on the techniques discussed in the previous section. This simulated target point cloud is subsequently processed by the registration model, which provides a transformation matrix. The difference between this matrix and the ground truth serves as error feedback to train the model. The example flow diagram for the training pipeline of the registration model can be seen in Figure 19.
Testing Pipeline
The model is tested by utilizing the 3D reconstructed point cloud of the scene alongside the 3D pre-operative model. These are passed through the registration model, and the resulting transformation matrix is applied to the source pre-operative point cloud to obtain the registered output point cloud. The example flow diagram for the testing pipeline of the registration model can be seen in Figure 20.
Train Mode
In Train mode, users can fine-tune a model on simulated data. Consider a scenario where we receive a 3D segmented model of a new patient, and we want to fine-tune our model for this new patient. Users can simply select the ply file of the new segmented model and set the parameters relevant to the simulated data for model fine-tuning.
The parameters relevant to the simulated data include the following:
Input STL file location: The file location of the 3D point cloud file of the 3D preoperative model of the patient.
Number of sampled points: The number of points to sample from the STL file.
Maximum allowable rotation (Rotation Magnitude): The maximum rotation allowed while creating simulated training data.
Maximum allowable translation (Translation Magnitude): The maximum translation that the simulated target point cloud is allowed to be away from the initial source point cloud.
Visibility control: Controls how much of the simulated target point cloud is visible.
Noise introduction: As discussed earlier, we add noise in the form of a plane through the point cloud, simulating the patient lying on a table surrounded by spherical noise scattered throughout the scene.
Maximum number of noise spheres: The maximum number of noise spheres allowed around the simulated target point cloud.
Number of iterations for the RPMNet algorithms: As RPMNet is an iterative algorithm, we can set how many iterations we want to run the model for.
Model file path for fine-tuning: Model path specifies the model file that we want to finetune.
Test Mode
The GUI offers two testing modes: users can either evaluate the model using actual data or simulate the testing with predefined parameters. Opting for simulated data mode means that the source point cloud undergoes manipulation based on the simulation parameters previously described in the Train mode. The example GUI can be seen in Figure 21.
In addition to the previously mentioned parameters, a distinct set of options is accessible in the test mode:
Input Target file (3D point cloud file obtained from any 3D reconstruction method): It is the 3D point cloud file constructed from any 3D reconstruction method.
Visualize (Enables visualization of the registered point cloud): Visualize allows users to see the model's performance by visualizing the actual registered point cloud.
Save Mesh: Save mesh allows users to choose whether they would like to save the mesh of the registered point cloud scene.
Capture RGB: Capture RGB allows users to save a rendered image of the 3D registration scene.
Simulated: This option allows users to switch between using simulated data or actual data.
These features make the GUI a versatile tool for training and evaluating the registration model with ease.
Results
Evaluation Metric
The goal of the registration method is to align the 3D Hologram model as precisely as possible over the patient in Hololens. In our scenario, both the patient's 3D reconstruction and the 3D Hologram model are represented as point clouds, and our objective is to ensure these point clouds overlap as closely as possible. One metric that aids in achieving this objective is the nearest neighbour distance.
In this method, for each point in the target point cloud (which is the 3D reconstructed point cloud), we search for the nearest point in the source point cloud (the 3D Hologram). We then establish pairs of points and calculate the Euclidean distance between them. The average of these distances from all the points yields the final metric.
It's important to note that this metric isn't flawless in our context because the target point cloud contains areas with no corresponding points in the source 3D Hologram, specifically the area corresponding to the surface on which Phantom model is placed. A more precise metric would involve manually assessing the distortion between the two point clouds using the measurement features available in Hololens. However, this would necessitate deploying the registration method on Hololens, which represents the next step in the project.
Registration evaluation using different methods for 3D reconstruction
In this section, we will assess the performance of our ultimate registration model using real 3D reconstructions obtained from various methods. Furthermore, to ensure that the model is applicable beyond this specific use case, we will also evaluate it using additional medical datasets.
Evaluation Metric for Different Methods
To compute the evaluation metric for various methods, we evaluate the model using 100 random translations and rotations applied between the 3D pre-operative Hologram and the 3D reconstructed point clouds. Subsequently, we calculate the average of Nearest Neighbor Distance obtained in each configuration.
Average Nearest Neighbor Distance (in
Method mm)
RGBD Camera 81.65
Photogrammetry 58.62
Hololens Spatial Mapping 116.80
Hololens Depth Camera (without Fine 36.52 tuning)
Hololens Depth Camera (with Fine tuning) 31.52
Conclusion
The visualization and evaluation metric indicate that the model excels when working with point clouds extracted from the Hololens Depth Camera, achieving exceptional accuracy during registration. This outcome is especially encouraging, as it demonstrates that the Hololens possesses the capability to generate high-quality point clouds independently, negating the necessity for additional external hardware for processing.
Comparison of Registration with a Fine-Tuned Model vs. a Non-Fine-Tuned Model
In the previous section, we noted that there is not a significant distinction between the generic registration model and the case-specific, finely adjusted registration model when considering the evaluation metric. Nevertheless, a closer scrutiny of the registered point cloud reveals that the fine-tuned registration model excels in addressing minor disparities between the registered and target point clouds, leading to a notably more accurate registration. This is shown in Figures 22 and 23, where the results of the fine-tuned model on the right effectively compensate for the subtle imperfections in the registration achieved by the generic model.
This study underscores the significance of fine-tuning the registration model with simulated data. Fine-tuning plays a vital role in addressing subtle disparities between the source and target point clouds, resulting in a significantly improved registration process.
While this method effectively removes noise from other objects in the scene by focusing on the bounding box, it comes at the cost of losing a lot of geometric detail of the object in focus, as the resulting point cloud has a reduced number of points.
The model performance on additional medical datasets can be seen in Figures 24 and 25.
Conclusions
In summary, this project has established a robust platform for advancing research by combining Hololens technology and Deep Learning models to offer a more convenient and informative approach to guide surgeons during surgical procedures. The project has achieved several noteworthy milestones:
1. Development of 3D Preoperative Models: We successfully devised a semi-automated methodology for segmenting CT scans, resulting in the creation of precise 3D preoperative models.
2. Exploration of 3D Scene Reconstruction Techniques: We systematically explored various techniques for reconstructing 3D scenes, including RGB cameras, RGBD cameras, spatial mapping, and the internal depth sensor of the Hololens. Our investigations revealed that the Hololens' internal depth sensor offers intricate scene geometry, enhanced by effective noise reduction through Hololens gestures.
3. Model Training and Testing: To improve the precision of the registration model, we generated a synthetic dataset by manipulating 3D preoperative models. This synthetic data served as the training data for our deep learning model, which we rigorously evaluated using both synthetic and real point cloud data.
4. Examination of Registration Methods: We analyzed various registration methodologies, such as RPMNet, Predator, and ICP. After extensive experimentation, we found that RPMNet+ICP consistently produces optimal results, especially when tested with the ModelNet40 dataset.
5. User-Friendly Interface: We developed an intuitively navigable GUI that enables users to train the registration model with synthetic data and subsequently conduct testing with either simulated or real 3D reconstructed point cloud data.
In summary, this project represents a significant step towards enhancing surgical procedures through the fusion of augmented reality and deep learning. Its potential applications extend to guiding surgeons in real-world operating scenarios, promising impactful contributions to the field.
Unless the context clearly requires otherwise, throughout the description and the claims, the words "comprise," "comprising," "include," "including," and the like are to be construed
in an inclusive sense, as opposed to an exclusive or exhaustive sense; that is to say, in the sense of "including, but not limited to."
The words "coupled" or "connected" or "tied", as generally used herein, refer to two or more elements or nodes that may be either directly connected, or connected by way of one or more intermediate elements. Additionally, the words "herein,” "above," "below," and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of this application. Where the context permits, words in the Detailed Description using the singular or plural number may also include the plural or singular number, respectively. The words "or" in reference to a list of two or more items, is intended to cover all of the following interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list.
It will be understood that the above list is non-exhaustive, and that the method and system described herein is applicable to many technical problem domains to which machine learning models may be applied.
Various modifications, whether by addition, substitution, or deletion will be apparent to the intended reader to provide further examples of the present disclosure, any and all of which are intended to be encompassed by the appended claims.
References
[AGSL19] Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan, and Simon Lucey. Pointnetlk: Robust & efficient point cloud registration using pointnet.
CoRR, abs/1903.05711, 2019.
[AHS+21] Christopher M. Andrews, Alexander B. Henry, Ignacio M. Soriano, Michael K. Southworth, and Jonathan R. Silva. Registration techniques for clinical applications of three-dimensional augmented reality devices. IEEE Journal of Translational Engineering in Health and Medicine, 9, 2021.
[AJU+18] Sebastian Andress, Alex Johnson, Mathias Unberath, Alexander Felix Winkler, Kevin Yu, Javad Fotouhi, Simon Weidert, Greg Osgood, and Nassir
Navab. On-the-fly augmented reality for orthopedic surgery using a multimodal fiducial. Journal of Medical Imaging, 5(2):021209-021209, 2018.
[BECS22] Manuel Birlo, P. J. Eddie Edwards, Matthew Clarkson, and Danail Stoyanov.
Utility of optical see-through head mounted displays in augmented realityassisted surgery: A systematic review. Medical Image Analysis, 77, 2022.
[BM92] P.J. Besl and Neil D. McKay. A method for registration of 3-d shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2):239-256, 1992.
[Clo] CloudCompare. Cloud-to-cloud distance.
[FJDV18] Taylor Frantz, Bart Jansen, Johnny Duerinck, and Jef Vandemeulebroucke. Augmenting microsoft's hololens with vuforia tracking for neuronavigation.
Healthcare technology letters, 5(5):221-225, 2018.
[FLLW21] Kexue Fu, Shaolei Liu, Xiaoyuan Luo, and Manning Wang. Robust point cloud registration framework based on deep graph matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8893-8902, 2021.
[GCJ+20] Jacob Gibby, Steve Cvetko, Ramin Javan, Ryan Parr, and Wendell Gibby. Use of augmented reality for image-guided spine procedures. European Spine Journal, 29: 1823-1832, 2020.
[GGS+18] Thomas M Gregory, Jules Gregory, John Sledge, Romain Allard, and Olivier Mir. Surgery guided by mixed reality: presentation of a proof of concept.
Acta orthopaedica, 89(5):480-483, 2018.
[GRL+98] Steven Gold, Anand Rangarajan, Chien-Ping Lu, Suguna Pappu, and Eric Mjolsness. New algorithms for 2d and 3d point matching: pose estimation and correspondence. Pattern recognition, 31(8): 1019-1031, 1998.
[HGU+21] Shengyu Huang, Zan Gojcic, Mikhail Usvyatsov, Andreas Wieser, and Konrad Schindler. Predator: Registration of 3d point clouds with low overlap. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4267-4276, June 2021.
[HLY+17] Ying He, Bin Liang, Jun Yang, Shunzhi Li, and Jin He. An iterative closest points algorithm for registration of 3d laser scanner point clouds with geometric features. Sensors, 17(8), 2017.
[HS019] Yilei Huang, Samjhana Shakya, and Temitope Odeleye. Comparing the functionality between virtual reality and mixed reality for architecture and construction uses. Journal of Civil Engineering and Architecture, 13(l):409- 414, 2019.
[ILH+17] Adam Ibrahim, John B. Lanier, Brandon Huynh, John O'Donovan, and Tobias Ho llerer. Real-time object recognition on the microsoft hololens. 2017.
[LDZS18] Yang Liu, Haiwei Dong, Longyu Zhang, and Abdulmotaleb El Saddik. Technical evaluation of hololens for multimedia: A first look. IEEE MultiMedia, 25(4):8-18, 2018.
[LRvA+19] Florentin Liebmann, Simon Roner, Marco von Atzigen, Davide Scaramuzza, Reto Sutter, Jess Snedeker, Mazda Farshad, and Philipp Furnstahl. Pedicle screw navigation using surface digitization on the microsoft hololens. International journal of computer assisted radiology and surgery, 14: 1157-1165, 2019.
[LSW17] Yinlong Liu, Zhijian Song, and Manning Wang. A new robust marker-less method for automatic image-to-patient registration in image-guided neurosurgery system. Computer Assisted Surgery, 22(supl):319-325, 2017. PMID: 29094615.
[LWZ+19] Weixin Lu, Guowei Wan, Yao Zhou, Xiangyu Fu, Pengfei Yuan, and Shiyu Song. Deepvcp: An end-to-end deep neural network for point cloud registration.
In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12-21, 2019.
[LYLG22] Shikun Li, Yang Ye, Jianya Liu, and Liang Guo. Vprnet: Virtual points registration network for partial-to-partial point cloud registration. Remote Sensing, 14(11), 2022.
[MJC+18] Jonathan L McJunkin, Pawina Jiramongkolchai, Woenho Chung, Michael Southworth, Nedim Durakovic, Craig A Buchman, and Jonathan R Silva.
Development of a mixed reality platform for lateral skull base anatomy. Otology & Neurotology: Official Publication of the American Otological Society, American Neurotology Society [and] European Academy of Otology and Neurotology,
39(10):ell37, 2018.
[MMGMGS+18] Rafael Moreta-Martinez, David Garcia-Mato, Monica Garcia-Sevilla, Ruben Perez-Mananes, Jose Calvo-Haro, and Javier Pascau. Augmented reality in computer-assisted interventions based on patient-specific 3d printed reference. Healthcare technology letters, 5(5): 162-166, 2018.
[MPRS21] Christian Moro, Charlotte Phelps, Petrea Redmond, and Zane Stromberga. Hololens and mobile augmented reality in medical and health science education: A randomised controlled trial. British Journal of Educational Technology, 52(2):680-694, 2021.
[NCR+20] Nhu Q. Nguyen, Jillian Cardinell, Joel M. Ramjist, Philips Lai, Yuta Dobashi, Daipayan Guha, Dimitrios Androutsos, and Victor X.D. Yang. An augmented reality system characterization of placement accuracy in neurosurgery. Journal of Clinical Neuroscience, 72:392-396, 2020.
[Opetea] OpenAnatomy. Atlas - spl head and neck, Year of the webpage's last update or access date.
[Opeteb] OpenAnatomy. Atlas - spl nac brain, Year of the webpage's last update or access date.
[Pal22a] Arrigo Palumbo. Microsoft hololens 2 in medical and healthcare context: state of the art and future prospects. Sensors, 22(20): 7709, 2022.
[Pal22b] Arrigo Palumbo. Microsoft hololens 2 in medical and healthcare context: State of the art and future prospects. Sensors, 22(20), 2022.
[PCL21] Liang Pan, Zhongang Cai, and Ziwei Liu. Robust partial-to-partial point cloud registration in a full range. CoRR, abs/2111.15606, 2021.
[PIL+18] Philip Pratt, Matthew Ives, Graham Lawton, Jonathan Simmons, Nasko Radev, Liana Spyropoulou, and Dimitri Amiras. Through the hololens™ looking glass: augmented reality for extremity reconstruction surgery using 3d vascular models with perforating vessels. European radiology experimental, 2: 1-7, 2018.
[Pol20] Marc Polleyfeys. Microsoft hololens 2: Improved research mode to facilitate computer vision research. 2020.
[QSMG17] Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation, 2017.
[QYSG17] Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Pointnet+ + : Deep hierarchical feature learning on point sets in a metric space, 2017.
[QYW+22] Zheng Qin, Hao Yu, Changjian Wang, Yulan Guo, Yuxing Peng, and Kai Xu. Geometric transformer for fast and robust point cloud registration, 2022.
[Tre23] Laia Tremosa. Beyond ar vs. vr: What is the difference between ar vs. mr vs. vr vs. xr? 2023.
[UBG+20] Dorin Ungureanu, Federica Bogo, Silvano Galliani, Pooja Sama, Xin Duan, Casey Meekhof, Jan Sf uhmer, Thomas J Cashman, Bugra Tekin, Johannes L Sch ' onberger, et al. Hololens 2 research mode as a tool for computer vision research. arXiv preprint arXiv:2008.11239, 2020.
[vHCE21] Felix von Haxthausen, Yenjung Chen, and Floris Ernst. Superimposing holograms on real world objects using hololens 2 and its depth camera.
Current Directions in Biomedical Engineering, 7(1): 111-115, 2021.
[WLDW22] HaipingWang, Yuan Liu, Zhen Dong, andWenpingWang. You only hypothesize once: Point cloud registration with rotation-equivariant descriptors. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1630-1641, 2022.
[WPJ+22] Ryszard Wierzbicki, Maria Paw lowicz, J 'ozefa Job, Robert Balawender, Wojciech Kostarczyk, Maciej Stanuch, Krzysztof Jane, and Andrzej Skalski. 3d mixed-reality visualization of medical imaging data as a supporting tool for innovative, minimally invasive surgery for gastrointestinal tumors and systemic treatment as a new path in personalized treatment of advanced cancer diseases. Journal of Cancer Research and Clinical Oncology, 148(l):237-243, 2022.
[WS19a] Yue Wang and Justin M. Solomon. Deep closest point: Learning representations for point cloud registration, 2019.
[WS19b] Yue Wang and Justin M. Solomon. Prnet: Self-supervised learning for partial-to-partial registration. CoRR, abs/1910.12240, 2019.
[XLW+21] Hao Xu, Shuaicheng Liu, Guangfu Wang, Guanghui Liu, and Bing Zeng. Omnet: Learning overlapping mask for partial-to-partial point cloud registration,
2021.
[XYL+22] Hao Xu, Nianjin Ye, Guanghui Liu, Bing Zeng, and Shuaicheng Liu. Finet: Dual branches feature interaction for partial-to-partial point cloud registration,
2022.
[YEK+20] Wentao Yuan, Benjamin Eckart, Kihwan Kim, Varun Jampani, Dieter Fox, and Jan Kautz. Deepgmr: Learning latent gaussian mixture models for registration. CoRR, abs/2008.09088, 2020.
[YL20] Zi Jian Yew and Gim Hee Lee. RPM-net: Robust point matching using learned features. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, jun 2020.
[YLCJ15] Jiaolong Yang, Hongdong Li, Dylan Campbell, and Yunde Jia. Go-icp: A globally optimal solution to 3d icp point-set registration. IEEE transactions on pattern analysis and machine intelligence, 38(ll):2241-2254, 2015.
[zHbFwS+19] Hong zhi Hu, Xiao bo Feng, Zeng wu Shao, Mao Xie, Song Xu, Xing huo Wu, and Zhe wei Ye. Application and prospect of mixed reality technology in medical field. Current Medical Science, 39, 2019.
[ZPK16] Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Fast global registration. In Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 766-782.
Springer, 2016.
[ZSD+22] Zhiyuan Zhang, Jiadai Sun, Yuchao Dai, Bin Fan, and Mingyi He. VRNet: Learning the rectified virtual corresponding points for 3d point cloud registration. IEEE Transactions on Circuits and Systems for Video Technology,
32(8):4997-5010, aug 2022.
[ZSN03] Timo Zinser, Jochen Schmidt, and Heinrich Niemann. A refined icp algorithm for robust 3-d correspondence estimation. 2:11-695, 2003.
Claims
1. A method of marker-less registration of a pre-obtained medical imagery three- dimensional virtual model of a subject with live imagery of the same subject obtained prior to or during a surgical or therapeutic procedure performed on the subject, comprising: a. scanning the subject using a depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject; b. obtaining live imagery of the subject via a camera; c. receiving the pre-obtained medical imagery three-dimensional virtual model of a subject; d. processing the anatomical layer point cloud data set and the three- dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; e. aligning the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and f. displaying the three-dimensional virtual model so aligned within the live imagery of the subject.
2. A method according to claim 1, wherein the camera and the depth sensor are substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween.
3. A method according to claims 1 or 2, wherein the processing comprises inputting the point cloud data set and the three dimensional virtual model into a trained neural network trained to find matching anatomical features in the two data sets.
4. A method according to any of the preceding claims, wherein the aligning comprises determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features.
5. A method according to claim 4, wherein the aligning comprises determining the rotation matrix R and the translation matrix t by applying a weighted singular value decompositions, SVD, to the point cloud data set.
6. A method according to claim 5, wherein the applying the weighted singular value decompositions, SVD, to the point cloud data set is repeated by a predetermined number of iterations, n.
7. A method according to any of the preceding claims, wherein the scanning of the subject using the depth sensor places emphasis on segmenting the scan to highlight surface skin of the subject.
8. A method according to any of the preceding claims, wherein the one or more anatomical features are one or more of skin, organs, tissues, bones, arteries, veins.
9. A method according to any of the preceding claims, wherein the one or more anatomical features are one or more abnormal anatomical features, the one or more abnormal anatomical features being tumours, bone fractures, bone deformities, vascular pathologies, infections.
10. A method according to any of the preceding claims, and further comprising repeating the above steps and updating the display to account for movement of the user about the subject, or movement of the anatomy of the subject, during the surgical or therapeutic procedure.
11. A control system for controlling an imaging system, the imaging system comprising a depth sensor and a camera, the control system comprising one or more processors collectively configured to: scan the subject using the depth sensor to generate a point cloud data set corresponding to an anatomical layer of the subject, obtain live imagery of the subject via the camera; receive the pre-obtained medical imagery three-dimensional virtual model of a subject; process the anatomical layer point cloud data set and the three- dimensional virtual model to determine one or more corresponding anatomical features of the subject present in both the point cloud data set and the three dimensional virtual model; align the point cloud data set and the three-dimensional virtual model in co-registration by aligning the determined one or more corresponding anatomical features of the subject in co-registration with each other; and display the three-dimensional virtual model so aligned within the live imagery of the subject.
12. A control system according to claim 11, wherein the camera and the depth sensor are substantially aligned so as to have the same or similar field of view of the subject or have a known field of view offset therebetween.
13. A control system according to claims 11 or 12, wherein the processing comprises inputting the point cloud data set and the three dimensional virtual model into a trained neural network trained to find matching anatomical features in the two data sets.
14. A control system according to claim 11 to 13, wherein the aligning comprises determining a rotation matrix R and a translation matrix t that aligns the anatomical layer point cloud data set with the three dimensional virtual model in dependence on the determined one or more corresponding anatomical features.
15. A control system according to claim 14, wherein the aligning comprises determining the rotation matrix R and the translation matrix t by applying a weighted singular value decompositions, SVD, to the point cloud data set.
16. A control system according to claim 15, wherein the applying the weighted singular value decompositions, SVD, to the point cloud data set is repeated by a predetermined number of iterations, n.
17. A control system according to any of claims 11 to 16, wherein the one or more anatomical features are one or more of skin, organs, tissues, bones, arteries, veins.
18. A control system according to claims 11 to 17, wherein the one or more anatomical features are one or more abnormal anatomical features, the one or more abnormal anatomical features being tumours, bone fractures, bone deformities, vascular pathologies, infections.
19. A control system according to claims 11 to 19, and further comprising repeating the above steps and updating the display to account for movement of the user about the subject, or movement of the anatomy of the subject, during the surgical or therapeutic procedure.
20. An imaging system, the imaging system comprising the control system of claims 11 to 19, and further comprising: a depth sensor; and a camera; wherein the depth sensor and the camera, are each configured, in dependence on a signal from the control system, to send image data to the control system, wherein the depth sensor sends depth data indicative of a 3D scan of the subject and the camera sends image data indicative of a live scene of the subject.
21. A method or system according to any of the preceding claims, wherein the depth sensor is a sensor selected from the group comprising: a stereo sensor, a Time- of-Flight (ToF) sensor, a LiDAR sensor or a Structured Light sensor.
22. A method or system according to any of the preceding claims, wherein the anatomical layer of the subject of which the point cloud is generated is a skin layer of the subject.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2408475.8 | 2024-06-13 | ||
| GBGB2408475.8A GB202408475D0 (en) | 2024-06-13 | 2024-06-13 | Spatial anatomy recognition technology |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025257529A1 true WO2025257529A1 (en) | 2025-12-18 |
Family
ID=91961047
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/GB2025/051219 Pending WO2025257529A1 (en) | 2024-06-13 | 2025-06-04 | Mixed reality system for spatial anatomy visualization |
Country Status (2)
| Country | Link |
|---|---|
| GB (1) | GB202408475D0 (en) |
| WO (1) | WO2025257529A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116492052A (en) * | 2023-04-24 | 2023-07-28 | 中科智博(珠海)科技有限公司 | Three-dimensional visual operation navigation system based on mixed reality backbone |
| US20240013412A1 (en) * | 2021-01-04 | 2024-01-11 | Proprio, Inc. | Methods and systems for registering preoperative image data to intraoperative image data of a scene, such as a surgical scene |
| CN119741352A (en) * | 2025-03-04 | 2025-04-01 | 广东工业大学 | A liver point cloud registration method and device based on Transformer architecture |
-
2024
- 2024-06-13 GB GBGB2408475.8A patent/GB202408475D0/en not_active Ceased
-
2025
- 2025-06-04 WO PCT/GB2025/051219 patent/WO2025257529A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20240013412A1 (en) * | 2021-01-04 | 2024-01-11 | Proprio, Inc. | Methods and systems for registering preoperative image data to intraoperative image data of a scene, such as a surgical scene |
| CN116492052A (en) * | 2023-04-24 | 2023-07-28 | 中科智博(珠海)科技有限公司 | Three-dimensional visual operation navigation system based on mixed reality backbone |
| CN119741352A (en) * | 2025-03-04 | 2025-04-01 | 广东工业大学 | A liver point cloud registration method and device based on Transformer architecture |
Non-Patent Citations (44)
| Title |
|---|
| "In 2020 IEEE/CVF Conference on Computer Vision and", June 2020, IEEE, article "RPM-net: Robust point matching using learned features" |
| ADAM IBRAHIM, JOHN B. LANIER, BRANDON HUYNH, JOHN O'DONOVAN, AND TOBIAS HO ILERER, REAL-TIME OBJECT RECOGNITION ON THE MICROSOFT HOLOLENS, 2017 |
| ARRIGO PALUMBO.: " Microsoft hololens 2 in medical and healthcare context: state of the art and future prospects.", SENSORS, vol. 22, no. 20, 2022, pages 7709 |
| ARRIGO PALUMBO.: " Microsoft hololens 2 in medical and healthcare context:State of the art and future prospects", SENSORS, vol. 22, no. 20, 2022 |
| CHARLES R. QI, HAO SU, KAICHUN MO, AND LEONIDAS J. GUIBAS., DEEP LEARNING ON POINT SETS FOR 3D CLASSIFICATION AND SEGMENTATION, 2017 |
| CHARLES R. QILI YIHAO SULEONIDAS J: "Pointnet++: Deep hierarchical feature learning on point sets in a metric space", GUIBAS, 2017 |
| CHRISTIAN MORO, CHARLOTTE PHELPS, PETREA REDMOND, AND ZANE STROMBERGA.: "Hololens and mobile augmented reality in medical and health science education: A randomised controlled trial.", BRITISH JOURNAL OF EDUCATIONAL TECHNOLOGY, vol. 52, no. 2, 2021, pages 680 - 694 |
| DORIN UNGUREANU, FEDERICA BOGO, SILVANO GALLIANI, POOJA SAMA, XIN DUAN, CASEY MEEKHOF, JAN ST¨UHMER, THOMAS J CASHMAN, BUGRA TEKIN: "Hololens 2 research mode as a tool for computer vision research.", ARXIV PREPRINT ARXIV:2008.11239, 2020 |
| FELIX VON HAXTHAUSEN, YENJUNG CHEN, AND FLORIS ERNST.: "Superimposing holograms on real world objects using hololens 2 and its depth camera.", CURRENT DIRECTIONS IN BIOMEDICAL ENGINEERING, vol. 7, no. 1, 2021, pages 111 - 115 |
| FLORENTIN LIEBMANN, SIMON RONER, MARCO VON ATZIGEN, DAVIDE SCARAMUZZA, RETO SUTTER, JESS SNEDEKER, MAZDA FARSHAD, AND PHILIPP FURN: "Pedicle screw navigation using surface digitization on the microsoft hololens", INTERNATIONAL JOURNAL OF COMPUTER ASSISTED RADIOLOGY AND SURGERY, vol. 14, 2019, pages 1157 - 1165, XP036808230, DOI: 10.1007/s11548-019-01973-7 |
| HAIPINGWANG, YUAN LIU, ZHEN DONG, ANDWENPINGWANG.: "You only hypothesize once: Point cloud registration with rotation-equivariant descriptors.", PROCEEDINGS OF THE 30TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, 2022, pages 1630 - 1641 |
| HAO XUNIANJIN YEGUANGHUI LIUBING ZENGSHUAICHENG LIU, FINET: DUAL BRANCHES FEATURE INTERACTION FOR PARTIAL-TO-PARTIAL POINT CLOUD REGISTRATION, 2022 |
| HAO XUSHUAICHENG LIUGUANGFU WANGGUANGHUI LIUBING ZENG, OMNET: LEARNING OVERLAPPING MASK FOR PARTIAL-TO-PARTIAL POINT CLOUD REGISTRATION, 2021 |
| HONG ZHI HU, XIAO BO FENG, ZENG WU SHAO, MAO XIE, SONG XU, XING HUO WU, AND ZHE WEI YE.: "Application and prospect of mixed reality technology in medical field.", CURRENT MEDICAL SCIENCE, vol. 39, 2019, XP036725408, DOI: 10.1007/s11596-019-1992-8 |
| JACOB GIBBY, STEVE CVETKO, RAMIN JAVAN, RYAN PARR, AND WENDELL GIBBY.: "Use of augmented reality for image-guided spine procedures. ", JOURNAL, vol. 29, 2020, pages 1823 - 1832, XP037209858, DOI: 10.1007/s00586-020-06495-4 |
| JIAOLONG YANG, HONGDONG LI, DYLAN CAMPBELL, AND YUNDE JIA: " Go-icp: A globally optimal solution to 3d icp point-set registration", IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, vol. 38, no. 11, 2015, pages 2241 - 2254, XP011624337, DOI: 10.1109/TPAMI.2015.2513405 |
| JONATHAN L MCJUNKIN, PAWINA JIRAMONGKOLCHAI, WOENHO CHUNG, MICHAEL SOUTHWORTH, NEDIM DURAKOVIC, CRAIG A BUCHMAN, AND JONATHAN R SI: "& Neurotology: Official Publication of the American Otological Society", vol. 39, 2018, AMERICAN NEUROTOLOGY SOCIETY, article "Development of a mixed reality platform for lateral skull base anatomy.", pages: 1137 |
| KEXUE FU, SHAOLEI LIU, XIAOYUAN LUO, AND MANNING WANG: "IEEE/CVF Conference on Computer Vision and Pattern Recognition", 2021, article " Robust point cloud registration framework based on deep graph matching. In Proceedings", pages: 8893 - 8902 |
| LIANG PANZHONGANG CAIZIWEI LIU: "Robust partial-to-partial point cloud registration in a full range", CORR, 2021 |
| MANUEL BIRLO, P. J.EDDIE EDWARDS, MATTHEW CLARKSON, AND DANAIL STOYANOV.: "Utility of optical see-through head mounted displays in augmented realityassisted surgery: A systematic review", MEDICAL IMAGE ANALYSIS, vol. 77, pages 2022 |
| MAXIMILIAN WEBER ET AL: "Deep Learning-based Point Cloud Registration for Augmented Reality-guided Surgery", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 6 May 2024 (2024-05-06), XP091748914 * |
| NHU Q. NGUYEN, JILLIAN CARDINELL, JOEL M. RAMJIST, PHILIPS LAI, YUTA DOBASHI, DAIPAYAN GUHA, DIMITRIOS ANDROUTSOS, AND VICTOR X.D.: "An augmented reality system characterization of placement accuracy in neurosurgery. ", JOURNAL OF CLINICAL NEUROSCIENCE, vol. 72, 2020, pages 392 - 396 |
| P.J. BESLNEIL D. MCKAY: "A method for registration of 3-d shapes", IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, vol. 14, no. 2, 1992, pages 239 - 256 |
| PHILIP PRATTMATTHEW IVESGRAHAM LAWTONJONATHAN SIMMONSNASKO RADEVLIANA SPYROPOULOUDIMITRI AMIRAS: "Through the hololensTM looking glass: augmented reality for extremity reconstruction surgery using 3d vascular models with perforating vessels", EUROPEAN RADIOLOGY EXPERIMENTAL, vol. 2, 2018, pages 1 - 7 |
| QIAN-YI ZHOUJAESIK PARKVLADLEN KOLTUN: "In Computer Vision-ECCV 2016: 14th European Conference", 10 November 2016, SPRINGER,, article "Fast global registration", pages: 766 - 782 |
| RAFAEL MORETA-MARTINEZDAVID GARCIA-MATOMONICA GARCIA-SEVILLARUBEN PEREZ-MANANESJOSE CALVO-HAROJAVIER PASCAU: "Augmented reality in computer-assisted interventions based on patient-specific 3d printed reference", HEALTHCARE TECHNOLOGY LETTERS, vol. 5, no. 5, 2018, pages 162 - 166, XP006076217, DOI: 10.1049/htl.2018.5072 |
| RYSZARD WIERZBICKI, MARIA PAW LOWICZ, J 'OZEFA JOB, ROBERT BALAWENDER,WOJCIECH KOSTARCZYK, MACIEJ STANUCH, KRZYSZTOF JANC, AND AND: "3d mixed-reality visualization of medical imaging data as a supporting tool for innovative, minimally invasive surgery for gastrointestinal tumors and systemic treatment as a new path in personalized treatment of advanced cancer diseases", JOURNAL OF CANCER RESEARCH AND CLINICAL ONCOLOGY, vol. 148, no. 1, 2022, pages 237 - 243, XP037662861, DOI: 10.1007/s00432-021-03680-w |
| SEBASTIAN ANDRESS, ALEX JOHNSON, MATHIAS UNBERATH, ALEXANDER FELIX WINKLER, KEVIN YU, JAVAD FOTOUHI, SIMON WEIDERT, GREG OSGOOD, A: "On-the-fly augmented reality for orthopedic surgery using a multimodal fiducial", JOURNAL OF MEDICAL IMAGING, vol. 5, no. 2, 2018, pages 021209 - 021209, XP060107109, DOI: 10.1117/1.JMI.5.2.021209 |
| SHENGYU HUANG, ZAN GOJCIC, MIKHAIL USVYATSOV, ANDREAS WIESER, AND KONRAD SCHINDLER: " Predator: Registration of 3d point clouds with low overlap. ", PROCEEDINGS OF THE IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN, June 2021 (2021-06-01), pages 4267 - 4276 |
| SHIKUN LIYANG YEJIANYA LIULIANG GUO: "Vprnet: Virtual points registration network for partial-to-partial point cloud registration", REMOTE SENSING, vol. 14, no. 11, 2022 |
| STEVEN GOLDANAND RANGARAJANCHIEN-PING LUSUGUNA PAPPUERIC MJOLSNESS: "New algorithms for 2d and 3d point matching: pose estimation and correspondence", PATTERN RECOGNITION, vol. 31, no. 8, 1998, pages 1019 - 1031, XP004123268, DOI: 10.1016/S0031-3203(98)80010-1 |
| TAYLOR FRANTZ, BART JANSEN, JOHNNY DUERINCK, AND JEF VANDEMEULEBROUCKE.: "Augmenting microsoft's hololens with vuforia tracking for neuronavigation.", HEALTHCARE TECHNOLOGY LETTERS, vol. 5, no. 5, 2018, pages 221 - 225, XP006076220, DOI: 10.1049/htl.2018.5079 |
| THOMAS M GREGORY, JULES GREGORY, JOHN SLEDGE, ROMAIN ALLARD, AND OLIVIER: "Mir. Surgery guided by mixed reality: presentation of a proof of concept.", ACTA ORTHOPAEDICA, vol. 89, no. 5, 2018, pages 480 - 483 |
| WEIXIN LU, GUOWEI WAN, YAO ZHOU, XIANGYU FU, PENGFEI YUAN, AND SHIYU, IN PROCEEDINGS OF THE IEEE/CVF INTERNATIONAL CONFERENCE ON, 2019, pages 12 - 21 |
| WENTAO YUAN, BENJAMIN ECKART, KIHWAN KIM, VARUN JAMPANI, DIETER FOX,AND JAN KAUTZ.: "Deepgmr: Learning latent gaussian mixture models for registration", CORR, 2020 |
| YANG LIUHAIWEI DONGLONGYU ZHANGABDULMOTALEB EL SADDIK: "Technical evaluation of hololens for multimedia: A first look", IEEE MULTIMEDIA, vol. 25, no. 4, 2018, pages 8 - 18, XP011705921, DOI: 10.1109/MMUL.2018.2873473 |
| YASUHIRO AOKI, HUNTER GOFORTH, RANGAPRASAD ARUN SRIVATSAN, AND SIMON LUCEY.: "Pointnetlk: Robust & efficient point cloud registration using pointnet.", CORR, 2019 |
| YILEI HUANG, SAMJHANA SHAKYA, AND TEMITOPE ODELEYE.: "Comparing the functionality between virtual reality and mixed reality for architecture and construction uses", JOURNAL OF CIVIL ENGINEERING AND ARCHITECTURE, vol. 13, no. 1, 2019, pages 409 - 414 |
| YING HEBIN LIANGJUN YANGSHUNZHI LIJIN HE: "An iterative closest points algorithm for registration of 3d laser scanner point clouds with geometric features", SENSORS, vol. 17, no. 8, 2017 |
| YINLONG LIUZHIJIAN SONGMANNING WANG: "A new robust marker-less method for automatic image-to-patient registration in image-guided neurosurgery system", COMPUTER ASSISTED SURGERY, vol. 22, 2017, pages 319 - 325 |
| YUE WANGJUSTIN M. SOLOMON, DEEP CLOSEST POINT: LEARNING REPRESENTATIONS FOR POINT CLOUD REGISTRATION, 2019 |
| YUE WANGJUSTIN M. SOLOMON: "Prnet: Self-supervised learning for partial-to-partial registration", CORR, 2019 |
| ZHENG QIN, HAO YU, CHANGJIAN WANG, YULAN GUO, YUXING PENG, AND KAI XU, GEOMETRIC TRANSFORMER FOR FAST AND ROBUST POINT CLOUD REGISTRATION, 2022 |
| ZHIYUAN ZHANG, JIADAI SUN, YUCHAO DAI, BIN FAN, AND MINGYI HE.: " VRNet:Learning the rectified virtual corresponding points for 3d point cloud registration.", IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, vol. 32, no. 8, pages 4997 - 5010 |
Also Published As
| Publication number | Publication date |
|---|---|
| GB202408475D0 (en) | 2024-07-31 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12086988B2 (en) | Augmented reality patient positioning using an atlas | |
| Vercauteren et al. | Cai4cai: the rise of contextual artificial intelligence in computer-assisted interventions | |
| Wang et al. | Video see‐through augmented reality for oral and maxillofacial surgery | |
| US20230114385A1 (en) | Mri-based augmented reality assisted real-time surgery simulation and navigation | |
| US11628012B2 (en) | Patient positioning using a skeleton model | |
| Wu et al. | Ai-enhanced virtual reality in medicine: A comprehensive survey | |
| Lee et al. | Vision-based tracking system for augmented reality to localize recurrent laryngeal nerve during robotic thyroid surgery | |
| US12190522B2 (en) | Constrained object correction for a segmented image | |
| Wen et al. | In situ spatial AR surgical planning using projector-Kinect system | |
| EP3328305B1 (en) | Microscope tracking based on video analysis | |
| Halabi et al. | Virtual and augmented reality in surgery | |
| WO2016116449A1 (en) | Atlas-based determination of tumour growth direction | |
| van Doormaal et al. | Augmented reality in neurosurgery | |
| EP4128145B1 (en) | Combining angiographic information with fluoroscopic images | |
| Neri et al. | Towards patient-specific deformable registration in laparoscopic surgery | |
| Gard et al. | Image-based measurement by instrument tip tracking for tympanoplasty using digital surgical microscopy | |
| Amara et al. | A mobile ar computer-aided diagnosis: 6 dof brain tumour pose estimator using a fine-tuned efficientpose-based model | |
| WO2025257529A1 (en) | Mixed reality system for spatial anatomy visualization | |
| Nowak et al. | Intraoperative use of mixed reality (MR) in humans across surgical disciplines: a major review | |
| Chandelon et al. | Landmark-free automatic digital twin registration in robot-assisted partial nephrectomy using a generic end-to-end model | |
| Tukra et al. | AI in surgical robotics | |
| Chen | On-the-fly dense 3D surface reconstruction for geometry-aware augmented reality. | |
| Neri | Registration Methods for Surgical Augmented Reality | |
| Shrestha et al. | A novel enhanced energy function using augmented reality for a bowel: modified region and weighted factor | |
| Zampokas et al. | Augmented reality toolkit for a smart robot-assisted MIS platform |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25732150 Country of ref document: EP Kind code of ref document: A1 |