WO2025129253A1 - Training deep learning models - Google Patents

Training deep learning models Download PDF

Info

Publication number
WO2025129253A1
WO2025129253A1 PCT/AU2024/051378 AU2024051378W WO2025129253A1 WO 2025129253 A1 WO2025129253 A1 WO 2025129253A1 AU 2024051378 W AU2024051378 W AU 2024051378W WO 2025129253 A1 WO2025129253 A1 WO 2025129253A1
Authority
WO
WIPO (PCT)
Prior art keywords
machine learning
learning model
change
value
rate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/AU2024/051378
Other languages
French (fr)
Inventor
Jurgen MEJAN-FRIPP
Pierrick Bourgeat
Vincent Dore
Leo LEBRAT
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Commonwealth Scientific and Industrial Research Organization CSIRO
Original Assignee
Commonwealth Scientific and Industrial Research Organization CSIRO
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from AU2023904194A external-priority patent/AU2023904194A0/en
Application filed by Commonwealth Scientific and Industrial Research Organization CSIRO filed Critical Commonwealth Scientific and Industrial Research Organization CSIRO
Publication of WO2025129253A1 publication Critical patent/WO2025129253A1/en
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/0002Inspection of images, e.g. flaw detection
    • G06T7/0012Biomedical image inspection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/0002Inspection of images, e.g. flaw detection
    • G06T7/0012Biomedical image inspection
    • G06T7/0014Biomedical image inspection using an image reference approach
    • G06T7/0016Biomedical image inspection using an image reference approach involving temporal comparison
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H30/00ICT specially adapted for the handling or processing of medical images
    • G16H30/40ICT specially adapted for the handling or processing of medical images for processing medical images, e.g. editing
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/30ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10081Computed x-ray tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10088Magnetic resonance imaging [MRI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10104Positron emission tomography [PET]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20172Image enhancement details
    • G06T2207/20182Noise reduction or smoothing in the temporal domain; Spatio-temporal filtering
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30016Brain
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2210/00Indexing scheme for image generation or computer graphics
    • G06T2210/41Medical
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2211/00Image generation
    • G06T2211/40Computed tomography
    • G06T2211/441AI-based methods, deep learning or artificial neural networks

Definitions

  • This disclosure relates to training a machine learning model on longitudinal data. More particularly but not necessarily exclusively, this disclosure relates to training a machine learning model on longitudinal medical scan images of a patient study.
  • Machine learning models such as deep neural networks, are powerful tools in addressing challenging problems and providing predictions.
  • machine learning models can be used to determine the presence of a disease from an image or a collection of images, such as medical images. Such machine learning models are trained and can then be used to make predictions. Training involves comparing the output of the model with ground truth information (i.e., information that is known to be real or true, and is the target output of the model) and updating the machine learning model based on the comparison.
  • ground truth information i.e., information that is known to be real or true, and is the target output of the model
  • Machine learning models are trained to minimise a loss value (or a loss function) based on one or a combination of metrics based on some ground truth measurements.
  • the loss value is generally calculated from the difference between the output and the ground truth, with the goal of training being to minimise this loss such that the machine learning model produces an output that is near identical to the ground truth.
  • the loss value is indicative of the overlap between the predicted mask from the model and a manually defined mask of the structure to be segmented in an image.
  • the loss value would be formulated in a way that the overlap is maximised by minimising the loss value.
  • Other examples of loss values include the mean square error (MSE) measure between the predicted volume structure and its ground truth volume.
  • MSE mean square error
  • ground truth can often have inherent noise due to many reasons, such as the equipment used to measure the data or other errors. This can make it difficult to train a machine learning model on such ground truth data.
  • the training data is longitudinal data, the data is often measured years apart with different machines and calibration settings for each measurement. This introduces noise as there is inconsistency between each successive measurement.
  • the noise makes it difficult to train the model on ground truth data that comes from multiple sources. This is often the case when training machine learning models on medical images, where many databases of training data are collated as a significantly large amount of training data is required to accurately train a model.
  • a method for training a machine learning model on longitudinal medical scan images of a patient study comprises: calculating two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output ; and calculating a prediction using the output of the machine learning model for the medical image; calculating a rate of change between the two or more predictions for the two or more medical images; and creating a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
  • the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change as the expected rate of change can be enforced into the machine learning model during training. More specifically, training of the machine learning model leverages information that is not available during inference. This is ideal for situations with noisy ground truth data. Moreover, inference of the trained machine learning model can be performed by applying the model to a single image.
  • the method further comprises determining a polynomial function of rate of change of the biomedical factor and calculating the expected rate of change using the polynomial function.
  • calculating the expected rate of change using the polynomial function comprises applying the polynomial function to a mean value of the biomedical factor calculated by an average of the two or more predictions.
  • determining the polynomial function comprises calculating a rate of change and a mean value for each pair of medical images in the patient study.
  • the loss value further comprises a change penalty value reflecting a difference between the two or more predictions, the change penalty value penalising an unexpected temporal trend of the biomedical factor related to the expected rate of change.
  • the loss value further comprises a correction penalty value reflecting a difference between the two or more predictions and two or more corresponding ground truth values, the correction penalty value penalising the two or more predictions from deviating from the two or more corresponding ground truth values.
  • the loss value further comprises a correction penalty value reflecting a difference between the two or more outputs and an over-correction factor , the correction penalty value penalising the two or more outputs from deviating from the overcorrection factor.
  • the loss value further comprises an additional penalty value calculated by one or more of a difference between each of the two or more outputs of the machine learning model and a respective ground truth measurement; and a mean square error between each of the two or more outputs and the respective ground truth measurement.
  • the loss value is calculated by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
  • each of the medical images of the patient study has a timestamp and at least two of the two or more medical image have non-adjacent timestamps.
  • the method further comprises calculating a further two or more predictions of the biomedical factor by performing for each of a further two or more medical images of the patient study: applying the trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating a further rate of change between the further two or more predictions; and creating a further trained machine learning model by updating weights of the trained machine learning model to minimise a further loss value based of the further rate of change.
  • the method further comprises calculating another two or more predictions of the biomedical factor by performing for each of two or more medical images of another patient study different to the patient study: applying the trained machine learning model or further trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating another rate of change between the other two or more predictions; and creating another trained machine learning model by updating weights of the trained machine learning model or further trained machine learning model to minimise another loss value based on the other rate of change.
  • the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of the patient based on the output of the trained machine learning model.
  • the biomedical factor is one of: medical image quantification factor; a biological structure value; or a biomarker.
  • each of the medical images is a PET image and the method further comprises performing PET quantification by applying the trained machine learning model to a PET image of a patient.
  • the medical image quantification factor is a Centiloid or standardised uptake value ratio scaling factor; and the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of Alzheimer’s disease based on the Centiloid or standardised uptake value ratio scaling factor that is an output of the trained machine learning model.
  • the method further comprises generating one or more masks that correspond to corrections determined by the trained machine learning model.
  • the machine learning model is a neural network comprising one or more convolutional layers.
  • a system for training a machine learning model on longitudinal medical scan images of a patient study comprises: a processor configured to: calculate two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output; and calculating a prediction using the output of the machine learning model for the medical image; calculate a rate of change between the two or more predictions for the two or more medical images; and create a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
  • Fig. 1 shows a plot of calculated Centiloid values over time calculated using longitudinal PET scan images.
  • Fig. 2 shows a plot of the rate of change of the Centiloid value verses the mean value of the Centiloid value.
  • Fig. 3 illustrates a system for training a machine learning model on longitudinal medical scan images of a patient study.
  • Fig. 4 illustrates a method for training a machine learning model on longitudinal medical scan images of a patient study.
  • Fig. 6 shows the progression of the calculated rate of change of the scaled Centiloid values as the training of the machine learning model increases, in the example embodiment.
  • Fig. 7 shows the progression of the Standardized Uptake Value Ratio (SUVR) scaling factors as the optimisation processes described herein continues.
  • SUVR Standardized Uptake Value Ratio
  • Fig. 8a shows initialised target masks used to generate optimised target masks.
  • Fig. 8b shows the optimised target masks with similar SUVR scaling factors to the SUVR provided by the trained machine learning model disclosed herein.
  • Fig. 9a shows initialised reference masks used to generate optimised reference masks.
  • Fig. 9b shows the optimised reference masks with similar SUVR scaling factors to the SUVR provided by the trained machine learning model disclosed herein.
  • the disclosed system and method relate to training a machine learning model on longitudinal data to address the problem of noisy ground truth training data.
  • the disclosed system and method utilise the inherent or expected properties of the longitudinal data as a source of additional information to reduce the effects of noisy ground truth data. More specifically, longitudinal data that exhibits an expected rate of change of a factor, such as a biomedical factor, can be used to train a machine learning model by formulating the training in a way that reflects this expected rate of change.
  • the disclosed system and method uses a class of penalty values (as referred to as longitudinal constraints) in the loss value. More specifically, the penalty values exploit the expected rate of change of the longitudinal data by penalising the unexpected rate of change.
  • a simple longitudinal constraint could be to penalise negative (positive) changes between baseline and follow-up when the change is expected to be positive (negative). These changes could be biomedical factors such as the PET quantification within a mask or the volume of a segmentation mask, or other (more involved) factors.
  • Centiloid values as the expected trend is that the rate of change is positive (i.e., the Centiloid values increase over time)
  • a penalty value can be formulated to penalise training the machine learning model on negative rates of change.
  • This disclosure also provides a framework to include these constraints on a single timepoint inference model.
  • the model is trained to use the longitudinal information inherent to the longitudinal data for training (e.g., multiple images)
  • the trained model does not require the longitudinal information during inference.
  • the trained model can be used to calculate predictions on a single timepoint during inference (e.g., a single image).
  • One of the main advantages of the disclosed system and method is that it enables the inclusion of longitudinal constraints without relying on multiple timepoints during inference.
  • the proposed constraints or penalty values can also complement existing losses by ensuring longitudinal consistency in the training set, instead of solely relying on a noisy “ground truth” data.
  • the overall trajectory of expected change for a mean measure can be computed from a large population analysis. This overall trajectory can use the error between the actual change and expected change as an additional loss. Since any information can be added in the loss, additional information can also be included into the loss, such as demographics. In another example, where different trajectories are expected for different clinical diagnosis, these different trajectories can be used to define different losses for different clinical groups. Since this information is not used at inference time (in other words, it is only used to compute the loss during training), this limits the risk to overfit the model to the data.
  • the disclosed system and method are applicable to many different machine learning (specifically, deep learning) applications, where longitudinal data is available, and where there are known longitudinal constraints that can be applied based on expected trends in the longitudinal data.
  • the machine learning application may be image segmentation, where there is a model of change over time of expectations with respect to increase/decrease of volume of structure over time.
  • this information can be used in the loss function in combination with other standard losses, and in this case, penalise the volume increasing over time. This is particularly useful for problems where the ground truth might not be very reliable. For example, this is particularly useful for PET quantification, since PET images can be very noisy.
  • PET quantification is used to analyse PET images and extract quantitative information from these images.
  • PET quantification for the study of Alzheimer’s disease involves a normalization of the PET images to the Centiloid scale, which is described as the standard scale for PET images acquired using an amyloid PET tracer.
  • PET images are typically normalised using the Standardized Uptake Value Ratio or SUVR.
  • the SUVR is defined as the ratio of a region of the PET image containing specific binding (a target region where the PET tracer will bind to the target), and a region of the PET image containing non-specific binding (as referred to as a reference region).
  • amyloid imaging those are defined as a region containing amyloid, typically the neocortex, and a region containing no amyloid, typically the cerebellum.
  • the Centiloid scale was developed to harmonise the quantifications (SUVR) from all amyloid PET tracers into the same scale using a linear transform based on head-to-head paired data.
  • SUVR is typically used for most PET tracer quantification
  • the Centiloid scale is typically used as it allows to quantify different PET tracers using the same scale.
  • Fig. 1 shows a plot of calculated Centiloid values over time calculated using longitudinal PET scan images.
  • the slope of the lines represents the “rate of change” of the Centiloid values.
  • the blue lines, e.g. at 101 represent the expected trend of the Centiloid values increasing over time, while the red lines, e.g. at 102, represent the unexpected trend of the Centiloid values decreasing over time, which is a result of noise in the PET scan images. This unexpected trend of the Centiloid values is often overlooked, as the Centiloid value calculation usually occurs in the background during typical PET quantification.
  • Fig. 2 shows a plot of the rate of change of the Centiloid value verses the mean value of the Centiloid value. From the extended trend of the Centiloid values, there should not be any negative rates of change (i.e., each plot point should be above horizontal line 201). Fig. 2 also shows curve 202, which represents a calculated curve of mean value versus rate of change. Curve 202 may be obtained by calculating the mean value and rate of change for some or each piece of data in a set of longitudinal data. More specifically, curve 202 may be produced using a regression algorithm applied on the data of Fig. 2. It is expected that the data points of Fig. 2 should not deviate far from curve 202, given that curve is the expected trend of the data.
  • a training of a machine learning model can be designed to estimate a correction of the Centiloid values in such a way to penalise negative rates of changes, penalise over-correction and penalise distance from curve 202, for example. Training of the machine learning model is penalised in these ways by introducing constraints or penalty values which are formulated in a way to represent the penalisation of these factors.
  • the disclosed system and method are advantageous for combating noisy ground truth data during training, there are also other advantages.
  • the volumes of structures present in training images may not be well defined (such as volume of white matter hyperintensity which have poor inter-rater reproducibility), which is compensated during training through the disclosed method.
  • Another advantage is that the training can be performed using longitudinal data that is temporally sparse (taken far apart). This is particularly advantageous for training the machine learning model on medical images such as MRI or CT images, as data is only collected every 1 or 2 years. It is noted that the machine learning model can also be trained on less sparse data using the disclosed method, such as longitudinal data that collected every day or every week or every month, for example.
  • the disclosed system and method are not limited to applications in the medical field and medical images. More specifically, this disclosure relates more generally to training a machine learning model on longitudinal data, where the longitudinal data may include medical images. However, in other instance, the longitudinal data may include other images. For example, any model for estimating a quantification/value from an image or any other input can be trained using the disclosed system and method, as long as there is some prior knowledge of the expected changes over time.
  • the longitudinal data may not include images at all, but rather be some other form of data.
  • it may be difficult to accurately measure the exact volume of water flowing through the dam which introduces a source of noise.
  • the volume may also fluctuate unexpectedly due to flash storms, which also introduces noise.
  • the disclosed system and method can compensate for the noise in the data by utilising the expected rate of change of the data during training.
  • Fig. 3 illustrates an example system 300 for training a machine learning model on longitudinal medical scan images of a patient study.
  • Fig. 3 is one example of a configuration of system 300.
  • system 300 is not strictly limited to this configuration and this may be one possible embodiment of system 300.
  • the medical scan images, or simply medical images may be magnetic resonance (MR) image, positron emission tomography (PET) image or computed tomography (CT) image.
  • MR magnetic resonance
  • PET positron emission tomography
  • CT computed tomography
  • medical scan images or medical images may be referred to simply as images or image data, noting that system 300 is equally applicable to images other than medical images.
  • a patient study is the study of a single patient where, for example, a medical image of the patient’s brain is captured at regular time intervals, such as annually. It is noted that the machine learning model can be trained on longitudinal medical scan images of multiple patient studies, thereby being trained on data from multiple patients. This is because the penalty values indicative of an expected rate of change reduces the noise that results from inconsistencies in the data between patients. Machine learning model usually require large amounts of training data, so it is an advantage that the machine learning model can be trained on longitudinal data from multiple patient studies. In embodiments outside of the medical field, the patient study may be a study of a subject over time.
  • System 300 comprises a device 301, which may be smartphone, computer or similar device.
  • Device 301 comprises a processor 302 connected to a program memory 303 and a data memory 304.
  • device 301 is implemented in a distributed computing architecture, such as a cloud computing provider.
  • Processor 302 may receive data through a variety of interfaces, which includes memory access of volatile memory, such as cache or RAM, or non-volatile memory, such as an optical disk drive, hard disk drive, storage server or cloud storage.
  • the program memory 303 is a non-transitory computer readable medium, such as a hard drive, a solid state disk or CD-ROM.
  • Software that is, an executable program stored on program memory 303 causes processor 302 to perform methods for training a machine learning model on longitudinal medical scan images of a patient study. For example, once executed, the software causes processor 302 to calculate two or more predictions of a biomedical factor by performing for each of two or more of the medical images: (i) applying the machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output of the machine learning model for the medical image, calculate a rate of change between the two or more predictions for the two or more medical images and create a trained machine learning model by updating weights of the machine learning model to minimise a loss value.
  • the data memory 304 may store image data and retrieve the image data for later use.
  • the image data may be indicative of a two-dimensional image, which may be stored on data memory 304 as Joint Photographic Experts Group (JPEG) format, RAW image format or a similar/equivalent image format, for example.
  • JPEG Joint Photographic Experts Group
  • RAW Raster Image Writer
  • the image data may be an RGB image, a multispectral image, a hyperspectral image, an infrared image, or a two- dimensional cross-sectional.
  • the image data may be indicative of a three-dimensional image, such as a PET, MR or CT image.
  • Data memory 304 may store image data indicative of the three-dimensional image data as a single-file DICOM (Digital Imaging and Communications in Medicine), multi-file DICOM format or a similar/equivalent image format, for example.
  • the machine learning models described herein (such as the machine learning model before, during and after training) may be stored on data memory 304.
  • the data memory 304 may also store predictions calculated by processor 302 performing any one of the disclosed methods, or any other variable or data necessary to perform such methods. For example, predictions, machine learning model outputs, loss values, rates of changes are calculated when processor 302 performs the disclosed method and can be stored on data memory 304.
  • processor 302 may then store the image data on data memory 304, such as on RAM or a processor register.
  • processor 302 requests the data from the data memory 304, such as by providing a read signal together with a memory address.
  • the data memory 304 provides the data as a voltage signal on a physical bit line and processor 302 receives the image data as the input image via a memory interface.
  • System 300 may also comprise a scanning machine 305, a server 306 and a monitor 307, which are each in communication with the device 301 via an input/output (I/O) port 308.
  • the scanning machine 305 scans a patient by incrementally capturing two- dimensional cross-sectional images of the patient. The collection of cross-sectional images represents image data that is indicative of a scan and represents a 3D scan of the patient.
  • processor 302 receives the image data via the I/O port 308. Processor 302 then performs the disclosed method using the image data received from the scanning machine 305 and trains the machine learning model.
  • the image data may be created using an MRI scanning machine.
  • MRI scanning machines consists of a table that slides into a cylinder. Inside the cylinder is a magnet that creates a strong magnetic field within the cylinder. The strong magnetic field is used to align the protons in the water molecules in a patient’s body, as the protons possess a magnetic moment due to their intrinsic spin. Once the protons align with the strong magnetic field, a radio frequency signal is sent from the MRI scanning machine, which excites the protons and causes them to align against the strong magnetic field. Once the radio frequency signal is turned off, the protons relax and re-align with the strong magnetic field, thereby emitting electromagnetic radiation, which is detected by the MRI scanning machine.
  • a patient lies on a motorised table that moves horizontally, in increments, through the cylinder and captures a single MR image at that increment, which corresponds to a two- dimensional cross-sectional image (virtual “slice”) of the patient’s body.
  • Each two- dimensional cross-sectional image corresponds to a different region of the patient’s body as they move horizontally, in increments, through the cylinder.
  • the collection of two- dimensional cross-sectional images also has an associated spatial sequence, in which the cross-sectional images are ordered along.
  • the horizontal direction that the patient moves within the cylinder is indicative of a dimension in space along the spatial sequence of the two- dimensional cross-sectional images.
  • the collection of two-dimensional cross-sectional images provides a three-dimensional representation of a patient’s internals, which aids physicians in diagnosing internal diseases and conditions.
  • the two-dimensional cross-sectional image consists of black and white regions, as well as shades of grey. These regions correspond to the rate at which excited protons return to their equilibrium state after being exposed to radio frequency (equilibrium state corresponds to the spin state which aligns with the strong external magnetic field), as well as the amount of energy released, changes depending on the environment and the chemical nature of the molecules. These factors can be differently weighted in the MR image, to give more contrast to certain tissues in the image. For example, in a Tl-weighted (Tlw) image, regions of high signal representing white regions may correspond to protein-rich fluid, whereas in T2- weighted (T2w) images, these regions may correspond to water rich fluids in the patient’s body.
  • Tlw Tl-weighted
  • T2w T2- weighted
  • the image data may be created using a CT scanning machine.
  • CT scanning machines use a rotating X-ray tube and a row of detectors placed in a gantry of the machine to measure X-ray attenuations by different tissues inside the body.
  • X-ray attenuations refer to when X-ray photons, with wavelengths ranging from 10 picometers to 10 nanometres, are absorbed by the tissue when the X-ray photons pass through the body of a patient.
  • a patient lies on a motorised table that moves horizontally, in increments, through the rotating X-ray tube and captures multiple X-ray measurements from different angles at each increment.
  • the multiple X-ray measurements are then processed on a computer using reconstruction algorithms to produce tomographic two-dimensional cross-sectional images (virtual “slices”) of a body.
  • Each two-dimensional cross-sectional image corresponds to a different region of the patient’s body as they move horizontally, in increments, through the rotating X-ray tube.
  • the collection of two-dimensional cross-sectional images also has an associated spatial sequence, in which the cross-sectional images are ordered.
  • the horizontal direction that the patient moves within the cylinder is indicative of a dimension in space along the spatial sequence of the two-dimensional cross-sectional images.
  • the collection of two-dimensional cross-sectional images provides a three- dimensional representation of a patient’s internals, which aids physicians in diagnosing internal diseases and conditions.
  • the two-dimensional cross-sectional (X-ray) images consists of black and white regions, as well as shades of grey. These regions correspond to the areas of the body that have attenuated the X-rays. For example, structures such as bones readily absorb X-rays and, thus, produce high contrast on the X-ray detector. As a result, bony structures appear whiter than other tissues in the X-ray image. Conversely, X-rays travel more easily through less radiologically dense tissues such as fat and muscle, as well as through air-filled cavities such as the lungs. These structures are displayed in shades of grey in the X-ray image. Regions where little to no attenuation has occurred will show up as black regions in the X-ray image.
  • the image data is indicative of a PET scan
  • the image data may be created using a PET scanning machine.
  • a radiopharmaceutical a radioisotope attached to a drug
  • beta plus decay a radiopharmaceutical
  • a positron is emitted, and when the positron interacts with an ordinary electron present in the body, the two particles annihilate and gamma rays are emitted.
  • These gamma rays are detected by gamma cameras to form a three-dimensional image, in a similar way that an X-ray image is captured.
  • Some parts of the body absorb more of a particular radiotracer than other parts of the body.
  • diseased cells such as cancer cells
  • body absorb more of the radiotracer than healthy ones do.
  • the brain is normally a rapid user of glucose, where radiotracer may be glucose containing the radioisotope.
  • Areas where there is a higher detection of gamma ray (as referred to as “hot spots”) will appear on the PET images as hotter colours, such as yellow and red, as opposed to areas of lower gamma ray detection, which will appear as blue.
  • hot spots enable the diagnosis of disease from the PET images.
  • Processor 302 may also receive other longitudinal data from other sources.
  • system 300 may comprise a RGB camera (not shown) configured to capture RGB image data intermediately over a period of time.
  • system 300 may comprise an inertial measurement unit (IMU) which measure position and pose data of a machine intermediately over a period of time.
  • IMU inertial measurement unit
  • system 300 may comprise a set of sensors which provide measurements to processor 302 intermediately over a period of time.
  • the disclosed system and method are not limited to applications involving medical images.
  • the network of devices may also communicate directly with one another via the Internet, any other means of wireless communication or wired connections.
  • multiple devices may each capture image data of a patient. Each instance of image data from the multiple devices may then be stored on server 306 for later use.
  • a database of image data captured from multiple devices may be stored on server 306.
  • server 306 may have similar functionality to that of device 301. Server 306 may perform parts of the disclosed methods described herein.
  • program memory may comprise software, that is, an executable program stored on program memory, and causes the processor to train a machine learning model on longitudinal data.
  • Software may provide a user interface presented to the user on device 301.
  • the user interface is configured to accept input (via buttons or text fields etc.) from the user, via a touch screen or a device attached to device 301 such as a keyboard or computer mouse.
  • These devices may also include a touchpad, an externally connected touchscreen, a joystick, a button, and a dial.
  • device 301 may display multiple instances of image data, and the user may choose one of the multiple instances of image data such as medical images, by processor 302. The user may choose one of the multiple instances of image data by interacting with the touch screen or inputting the selection with a keyboard or computer mouse.
  • Processor 302 may receive or send data, such as image data, from data memory 304 as well as from the I/O port 308.
  • processor 302 sends image data from the device 301 to server 306 via I/O port 308, such as by using a Wi-Fi network according to IEEE 802.11.
  • the Wi-Fi network may be a decentralised ad-hoc network, such that no dedicated management infrastructure, such as a router, is required or a centralised network with a router or access point managing the network.
  • System 300 may further be implemented within a cloud computing environment, such as a managed group of interconnected servers hosting a dynamic number of virtual machines.
  • I/O port 308 is shown as single entity, it is to be understood that any kind of data port may be used to receive data, such as a network connection, a memory interface, a pin of the chip package of processor 302, or logical ports, such as IP sockets or parameters of functions stored on program memory 303 and executed by processor 302.
  • the parameters of functions may be stored on data memory 304 and may be handled by-value or by-reference, that is, as a pointer, in the source code.
  • Fig. 4 illustrates method 400 for training a machine learning model on longitudinal medical scan images of a patient study.
  • Fig. 4 is to be understood as a blueprint for the software program and may be implemented step-by-step, such that each step in Fig. 4 is represented by a function in a programming language, such as Python, C++ or Java.
  • the resulting source code is then compiled and stored as computer-executable instructions on program memory 303, which causes processor 302 to perform method 400.
  • processor 302 may initialise the weights of the machine learning model (which define the machine learning model). For example, processor 302 may use random weights at the beginning of the train process, which are then updated during the training process.
  • the machine learning model may be a neural network.
  • the machine learning model may be a neural network comprising one or more convolutional layers.
  • Such a machine learning is known as a convolutional neural network (CNN).
  • CNN convolutional neural network
  • the CNN is ideal for applications involving images as it accounts for the positioning and shape of objects in the image data.
  • other types of machine learning models are equally applicable here.
  • the machine learning model may be implemented as a K nearest neighbour model, decision tree and support vector machine.
  • a convolution is an integration function that expresses the amount of overlap of one function g as it is shifted over another function/.
  • a convolution acts as a blender that mixes one function with another to give reduced data space while preserving the information.
  • a CNN performs convolutions with filters (matrix / vectors) and the filters comprise learnable parameters that are used to extract low-dimensional features from an input data. They have the property to preserve the spatial or positional relationships between input data points.
  • CNNs exploit the spatially-local correlation by enforcing a local connectivity pattern between neurons of adjacent layers.
  • the CNN architecture may comprise an optimal number of convolutional layers, filter size and stride length.
  • a convolution is the step of applying the concept of a sliding window (a filter with learnable weights) over the input and producing a weighted sum (of weights and input) as the output.
  • the weighted sum is the feature space which is used as the input for the next layers.
  • each convolutional layer comprises a filter, which calculates a weighted sum of adjacent pixel values.
  • a 2 x 2 filter comprises 4 weights, which are the coefficients of the filter.
  • the filter starts at an initial position in the image data structure, multiplies each pixel value in the image data structure with the respective filter coefficient and adds the results. Finally, the filter stores the resulting number in an output pixel.
  • the output pixel value of each of the one or more convolutional layers comprises a weighted sum of input pixel values.
  • the weights in the weighted sum correspond to the coefficients of the filter.
  • the filter moves by one pixel along one direction in the data structure and repeats the calculation for the next voxel of the output image. For example, if the stride is 2, the filter will move by two pixels along one direction. That direction may be an x-dimension or a y-dimension.
  • a single convolutional layer may use multiple filters, where each filter corresponds to a different feature.
  • the corresponding feature of each filter is determined during training of the CNN. More specifically, each filter may contain weights, which are adjusted during a training process through a backpropagation process, for example.
  • the filters are used to quantitatively determine the contribution of a particular feature on the output of the CNN.
  • the filters of this convolutional filter may correspond to simple features.
  • convolutional layers that occur later in the CNN architecture may exhibit more complex or abstract features.
  • the complex or abstract features may be a combination of the simple features from the previous convolutional layer.
  • the CNN architecture may also comprise a batch normalization layer and rectifier linear unit (ReLU) activation function. Further, the output of the last convolution layer may be subjected to Global Average Pooling (GAP). Even further, a fully connected layer with Sigmoid activation may also be used on the output layer. However, other activation functions may be used. These activation functions include, but are not limited to, a binary step function, a tanh function, a ReLU function or a softmax function.
  • ReLU rectifier linear unit
  • Fig. 5 is an example CNN architecture 500.
  • This example CNN architecture 500 is designed to predict parameters from an MR image for dynamically transformation of the intensity contrast of an input MR image to adapt to a segmentation task. More specifically, the example CNN predicts the parameters for a power function transformation and a piece- wise linear transformation.
  • the example CNN architecture 500 includes three convolutional blocks 511, 512, 513 and three fully connected (FC) blocks 521, 522, (523 and 524 are collectively one FC block). This CNN architecture 500 is used to provide results for the disclosed method presented later in the disclosure.
  • processor 302 may select 401 two or more images from the longitudinal data of a patient study.
  • the two or more images come from the same patient, rather than the two or more images being from different patients.
  • the two or more images may be randomly chosen from the patient study. For example, the two or more images do not have to be images that have been captured consecutively.
  • Processor 302 then calculates 420 two or more predictions of a biomedical factor each of two or more of the medical images.
  • the biomedical factor has an expected trend that is known beforehand and is utilised to train the machine learning model.
  • processor 302 applies 421, 426 the machine learning model to the medical image to calculate 422, 427 an output, then calculates 423, 428 a prediction using the output of the machine learning model for the medical image.
  • processor 302 calculates 420 the two or more predictions using the same machine learning model and more specifically, processor 302 calculates 420 the two or more predictions using a model that has the same architecture and weights.
  • processor 302 calculates 420 two or more predictions of a factor that is not a biomedical factor.
  • the factor may be any variable that has some expected trend (e.g., expected to change over time).
  • the factor may be a physical factor such as the volume of water in a dam. It is noted that while this disclosure describes embodiments related to a biomedical factor, these embodiments are equally applicable to a factor other than a biomedical factor.
  • each prediction may be calculated in parallel, as the calculations are not necessarily dependent on each other.
  • each prediction may be calculated in parallel across many CPUs and multiple cores. More specifically, each prediction may be calculated on a single CPU efficiently by multi-platform shared-memory parallel programming, such as OpenMP, which utilises the multiple computing cores of the CPU.
  • processor 302 calculates 423, 428 the prediction by applying a function to the output of the machine learning model.
  • the function may be a polynomial, exponential or trigonometric function or combination thereof.
  • processor 302 calculates 423, 428 the prediction by multiplying the output by a factor or measurement.
  • the output may be a scaling factor to correct a specific measurement and hence, the prediction corresponds to the corrected measurement.
  • the output may simply correspond to the prediction, which can be considered as processor 302 calculating 423, 428 the prediction by multiplying the output by 1.
  • the biomedical factor may be one of medical image quantification factor, a biological structure value or a biomarker.
  • the biomedical factor may be a medical image quantification factor, which may be a Centiloid value or standardised uptake value ratio scaling factor.
  • the biomedical factor may be a biological structure value, which may be a volume of a biological structure visible in the medical image, such as the grey matter in the brain.
  • the biomarker may be blood pressure or heart rate, for example.
  • processor 302 calculates 430 a rate of change between the two or more predictions for the two or more medical images. In some embodiments, processor 302 calculates 430 the rate of change based on a difference between the two or more predictions. In other embodiments, processor 302 calculates 430 the rate of change based on a difference between the two or more predictions and a difference between a variable on which the rate of change is based. For example, if there are only two predictions, processor 302 may calculate the rate of change using a difference of the predictions divided by the difference between the timepoints of each respective medical image.
  • the calculated rate of change is indicative of a trend in the longitudinal data.
  • processor 302 may calculate the rate of change explicitly by calculating a ratio of the change in the predictions and the change in an independent variable (such as time).
  • processor 302 may also calculate the rate of change implicitly by calculating a difference between the two or more predictions, as the medical images may inherently contain change information (such as temporal information). For example, two medical images in a patient study may be taken a year apart and hence, the two medical images have inherent temporal information, given that they are temporally related. As such, the rate of change may be calculated by a difference of the two predictions of the respective medical images, as the medical images inherently contain temporal information.
  • processor 302 calculates 440 a loss value to train the machine learning model.
  • the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
  • Reflecting a difference generally means that the penalty value is indicative of the difference.
  • Reflecting a difference may mean that the penalty value is related to the difference, such as being directly proportional to the difference.
  • Reflecting a difference may also mean that processor 302 calculates the penalty value using the difference, such as by applying a mathematical function to the difference.
  • the expected rate of change may be a function of time (such as linear, polynomial, exponential, etc.) or may be a second order rate (i.e., the rate of the rate of change) and processor 302 calculates the second order rate of change.
  • the expected rate of change may be zero, indicating that the biomedical factor remains constant over time.
  • the penalty value may be formulated in such a way to penalise the rate of change deviating from zero.
  • Calculating a loss value may be referred to as applying a loss function.
  • the outcome of applying a loss function is the loss value.
  • the loss value and loss function may also be referred to as cost value and cost function, respectively.
  • the loss function may be applied to the outputs and other factors, such as the timestamps of two or more medical images.
  • the loss value may incorporate many types of additional information that is not available at inference time. For example, the difference curves based on clinical diagnostic, age or any other relevant information may be used to train the machine learning model using the loss value.
  • Processor 302 then updates 450 weights of the machine learning model to minimise a loss value, thereby creating a trained machine learning model.
  • the longitudinal data i.e., the training data
  • the longitudinal data comprises label data, such as image data that has been labelled with a class, where the label for each item of training data may be manually provided by a user.
  • Training data may also or otherwise be obtained from a database containing already labelled data, for which manually labelling the data is not needed. Training the machine learning model in this way is referred to as “supervised learning”, in the sense that the outcome of the machine learning model is validated and the weights of the machine learning model are adjusted accordingly.
  • the weights may be updated through the process of backpropagation.
  • the weights may also be updated through a gradient descent process.
  • supervised learning is not the only method that can be used to train the machine learning models.
  • “semi-supervised learning” can be used in situations where only a relatively small amount of training data can be obtained.
  • a training set may be augmented to produce more training data, which is advantageous if a large training data is difficult to obtain.
  • a training set may be augmented by randomly resizing, cropping, flipping, or rotating the image data in the training set, for example.
  • the rates of change may only be used during training to update the machine learning model, rather than at inference time as the prediction provided by the machine learning model may correspond to single timepoint measurements.
  • the rates of change may not be used at inference time to obtain the prediction from the trained machine learning model.
  • single timepoint measurements may be used at inference time to generate corresponding predictions using the trained machine learning model.
  • Method 400 is preferably repeated multiple times using two or more different images.
  • the more repetitions of method 400 the more accurate the machine learning model becomes at making a prediction during inference.
  • training the machine learning model on a range of different images or training data also generally makes the machine learning model more accurate at making a prediction during inference.
  • reference to “trained machine learning model” means that the initial machine learning model has undergone at least one iteration of method 400.
  • “trained machine learning model” may also mean that the initial machine learning model has undergone multiple iterations of method 400.
  • Other terms such as “further trained machine learning model”, “another trained machine learning model” and “fully trained machine learning model” may also be used throughout this disclosure to indicate that the machine learning model has undergone multiple training iterations.
  • processor 302 may perform method 400 multiple times using two or more different images, processor 302 may calculate a further two or more predictions of the biomedical factor by, for each of a further two or more medical images of the patient study: (i) applying the trained machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output. Processor 302 may then calculate a rate of change between the further two or more predictions; and creates a further trained machine learning model by updating weights of the trained machine learning model to minimise a further loss value based on the further rate of change. As such, the machine learning model is trained on different images of the same patient study. This process may be repeated using each permutation of the two or more images in the patient study.
  • processor 302 may further train the machine learning model using images from another patient study, which is different to the patient study used to initially trained the machine learning model. More specifically, processor 302 may calculate another two or more predictions of the biomedical factor by, for each of two or more medical images of another patient study different to the patient study: (i) applying the trained machine learning model or further trained machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output. Processor 302 may then calculate another rate of change between the other two or more predictions; and creates another trained machine learning model by updating weights of the trained machine learning model or further trained machine learning model to minimise another loss value based on the other rate of change.
  • the machine learning model may be agnostic to the order of the two or more medical images of the patient study. As such, the machine learning model is unable to learn the order as each image is processed the same way, with the only difference being how the loss is computed. It is beneficial for the machine learning model not to learn the order of the input images, as it is desired that the model learns the expected trend in the data (i.e., the expected rate of change), rather than the sequence of the images. This provides better predictions during inference.
  • each of the medical images of the patient study has a timestamp and at least two of the two or more medical image have non-adjacent timestamps.
  • Non-adjacent timestamps being non-consecutive. For example, if the patient study contained longitudinal data of a brain CT image captured every year, then the non-adjacent timestamps would be the non-consecutive years.
  • processor 302 may calculate the expected rate of change.
  • Processor 302 may calculate the expected rate of change from the longitudinal data. For example, processor 302 may calculate a rate of change for each pair of images in the patient study. Processor 302 may then average the calculated rates of change to calculate the expected rate of change. Processor 302 may also use other techniques such as regression to calculate the expected rate of change from the longitudinal data.
  • processor 302 determines a polynomial function of rate of change of the biomedical factor and calculates the expected rate of change using the polynomial function.
  • This polynomial function represents the expected trend of the longitudinal data and hence, can be used to calculate the expected rate of change of the biomedical factor based on the two or more medical images used to train the machine learning model.
  • curve 202 represents a polynomial function and is based on the longitudinal data used to train the machine learning model. It is noted that this function is not limited to a polynomial function. While preferably a polynomial function, this function may be a polynomial, exponential or trigonometric function or combination thereof.
  • processor 302 calculates the expected rate of change using the polynomial function by applying the polynomial function to a mean value of the biomedical factor calculated by an average of the two or more predictions.
  • the mean value may be calculated using a difference of the two or more predictions divided by the number of two of more predictions.
  • the polynomial function may be a rate of change verses mean value function and hence, may be a function of the mean value that, when applied to the mean value, results in the rate change corresponding to that mean value.
  • processor 302 determines the polynomial function by calculating a rate of change and a mean value for each pair of medical images in the patient study.
  • processor 302 may determine the polynomial function by calculating a rate of change and a mean value for each pair of medical images in each patient study that form the training data. Processor 302 may determine the polynomial function by determining the best representation of the training data. For example, processor 302 may use regression (e.g., curve of best fit) or other techniques, such as applying a machine learning model.
  • regression e.g., curve of best fit
  • the loss value further comprises a change penalty value reflecting a difference between the two or more predictions, where the change penalty value penalises an unexpected temporal trend of the biomedical factor related to the expected rate of change.
  • the expected temporal trend of the longitudinal data may be a positive rate of change over time.
  • the change penalty value would penalise instances of a negative rate of change occurring.
  • the change penalty value may comprise a clamp function to only keep positive or negative changes.
  • the loss value further comprises a correction penalty value reflecting a difference between the two or more predictions and two or more corresponding ground truth values.
  • the correction penalty value may penalise the two or more predictions from deviating from the two or more corresponding ground truth values.
  • the correction penalty value may comprise two penalty values reflecting the slope and intercept between the two or more predicted values and the ground truth measurements.
  • the correction penalty value may be based on a difference between the rate of change of the two or more predictions and the rate of changes of the two or more corresponding ground truth values.
  • the slope and intercept may be based on a regression line between the ground truth measurements and the two or more predictions.
  • the slope between output of the predicted values and the ground truth measurements should not deviate too far from 1.
  • over-correction may be penalised during the training of the machine learning model using the two correction penalty values that reflect this expectation. This may ensure that no bias is introduced during the training process.
  • the correction penalty value may be a sum of the differences between each output and the over-correction factor. In further examples, the correction penalty value may be a sum of the absolute value of these differences.
  • a regression line may be calculated by determining a slope and intercept based on each prediction and the corresponding ground truth value. More specifically, a plot may be made with data points corresponding to the prediction and the corresponding ground truth value, then a linear regression of the plot may be determined. In essence, a linear function may be calculated where the input to the linear function corresponds to a prediction and the output (i.e., value of the function evaluated at an input) corresponds to a ground truth value. Hence, the slope may be given as:
  • correction penalty may be based on a difference between the slope and 1, as well as the intercept and 0, such that this difference is minimised during training.
  • the loss value further comprises a correction penalty value reflecting a difference between the two or more outputs and an over-correction factor, where the correction penalty value penalises the two or more outputs from deviating from the overcorrection factor.
  • the correction penalty value may be a sum of the differences between each output and the over-correction factor.
  • the correction penalty value may be a sum of the absolute value of these differences.
  • the loss value further comprises an additional penalty value calculated by one or more of: (i) a difference between each of the two or more outputs of the machine learning model and a respective ground truth measurement, and (ii) a mean square error between each of the two or more outputs and the respective ground truth measurement.
  • the loss value may comprise additional (standard) loss terms. In essence, the loss value may comprise any number of additional loss terms.
  • processor 302 calculates the loss value by a weighted sum.
  • the loss value may comprise the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
  • processor 302 calculates the loss value by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
  • the weights of the weighted sum may be considered “learning rates”, which are predetermined before training occurs, such as by user input. In some examples, one or more of the weights may be 1.
  • the biomedical factor can then be used to aid in determining a diagnosis for the patient.
  • a clinician may apply the trained machine learning model to an image of a new patient.
  • method 400 may further comprised applying the trained machine learning model to a medical image of a patient to aid in diagnosis of the patient based on the output of the trained machine learning model.
  • PET quantification may provide quantification information such as amount of amyloid plaques in the brain, for example.
  • PET quantification may also involve processor 302 creating a three-dimensional rendering of the patient’s brain from the PET image.
  • each of the medical images is a PET image and method 400 comprises performing PET quantification by applying the trained machine learning model to a PET image of a patient.
  • the biomedical factor would be a medical image quantification factor and more specifically, the medical image quantification factor would be a Centiloid or standardised uptake value ratio scaling factor.
  • method 400 would further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of Alzheimer’s disease based on the Centiloid or standardised uptake value ratio scaling factor that is an output of the trained machine learning model.
  • the biomedical factor is a medical image quantification factor and more specifically, the biomedical factor is a Centiloid value.
  • the output of the machine learning model being trained is a scaling factor for the SUVR value, before it is transformed into Centiloid value, and the predictions are scaled Centiloid values.
  • the medical images used during training and inference will be PET images, as the SUVR and Centiloid scale relate to PET images.
  • Centiloid values are used to standardise amyloid PET images for PET quantification by normalising the PET SUVR across different amyloid PET tracers to a standard and common unit system, with each PET tracer having its own linear transform to normalise SUVR into Centiloid.
  • PET quantification that relies on PET images in the Centiloid scale may be less accurate as a result of the noise in the PET image and the variability introduced by the use of different tracers and scanners.
  • machine learning techniques can be used to determine a correction to the calculated SUVR value, before they are transformed into Centiloid value, in an effort to reduce the effect of noise in the PET images and increase the accuracy of the PET quantification.
  • the machine learning model is trained on two medical images at a time (i.e., two images from the same patient study for each training iteration).
  • processor 302 calculates two predictions of the Centiloid value (i.e., scaled Centiloid value) by applying the machine learning model to each of the two medical images, individually, calculating a respective output and calculating a respective prediction using the respective output.
  • processor 302 calculates the predictions by first calculating the SUVR value for each of the two medical images.
  • Processor 302 may calculate the SUVR value by calculating a ratio between a reference region of the respective PET image and a “hot spot” region of the respective PET scan, such as an amyloid, before transforming them into Centiloid.
  • Processor 302 may also use the tracer and/or scanner information as part of the prediction, so that there are different corrections for different tracers, which allows to better control for their different levels of noise.
  • Processor 302 then calculates the predictions by a product of the calculated SUVR value and the output of the machine learning model (i.e., the scaling factor), before transforming the scaled SUVR into Centiloid (i.e., the scaled Centiloid value).
  • Processor 302 then calculates the loss value. It is noted that in some embodiments, processor 302 calculates the loss value by applying a loss function. For example, in the example embodiment, processor 302 applies a loss function to the calculated Centiloid values, the outputs of the machine learning model and the timestamps of the two medical images. In other words, the calculated Centiloid values, the outputs of the machine learning model and the timestamps of the two medical images are inputs into the loss function. As such, processor 302 may calculate the rate of change by applying the loss function to the inputs.
  • the loss value comprises three penalty values: (i) the penalty value indicative of the rate of change, (ii) the change penalty value penalising an unexpected temporal trend of the Centiloid value, and (iii) the correction penalty value penalising the two or more outputs from deviating from the over-correction factor.
  • processor 302 calculates a rate of change and a mean value for each pair of medical images for each patient study. Processor 302 then determines a polynomial function by applying a regression algorithm to the calculated rates of change. The calculated rates of change and the determined polynomial function for this example embodiment are shown in Fig. 2. In essence, it is expected that the rate of change does not deviate from the polynomial. As such, this penalty value indicative of the rate of change penalises large distances away from the polynomial. [0136] Processor 302 calculates the expected rate of change using the polynomial function by applying the polynomial function to a mean value of the Centiloid value calculated by an average of the two or more predictions. More specifically, the mean value is given by:
  • the change penalty value penalises an unexpected temporal trend of the Centiloid value related to the expected rate of change.
  • it is expected that Centiloid value will increase over time and hence, the rate of change of the Centiloid value is expected to be positive (more than zero).
  • the change penalty value is formulated in such a way to penalise decreasing Centiloid values and negative rates of change.
  • the change penalty value is based on a “baseline/follow up” concept, where patients will have a “baseline” medical image scan performed which will be compared to a “follow up” medical image scan in the future.
  • the clamp function restricts the difference in the predictions to only be positive, thereby enforcing the extended trend of an always positive rate of change during the training.
  • the correction penalty value penalises the two or more outputs of the machine learning model from deviating from an over-correction factor.
  • the over-correction factor is 1, as the output of the machine learning model is a scaling factor and hence, the output should be ideally 1 in the absence of noise in the PET image.
  • correction penalty value abs ⁇ outpuG — 1) + abs(oi4tpitt 2 — 1)
  • Fig. 6 shows the progression of the calculated rate of change of the scaled Centiloid values as the training of the machine learning model increases. In particular, it can be shown that the calculated rates of change move towards the polynomial function as the training continues. Further, the calculated rates of change become more positive as per the expected rate of change of the Centiloid values.
  • Machine learning models are often known for being black boxes, meaning that it is unknown how they calculate a prediction from an input. For example, while the models may be able to detect the presence of a disease from image data, it may not be clear how the model makes its prediction nor may it be clear what features the model deems important when making its prediction.
  • the black box nature of machine learning model means that interpretability can be challenging. The understanding of decisions made by black-box algorithms is a challenge and assessing their fairness and unbiasedness is an important step to deployment in many industries.
  • the trained machine learning model disclosed herein may output a prediction of a biomedical factor, which may be a standardised uptake value ratio (SUVR) scaling factor (which may be simply referred to as a SUVR).
  • SUVR standardised uptake value ratio
  • the SUVR may be defined as the ratio of a region of the PET image containing specific binding (a target region where the PET tracer will bind to the target), and a region of the PET image containing non-specific binding (as referred to as a reference region).
  • a target mask may be generated which highlights the target regions in a PET image.
  • a reference mask may be generated which highlights the non-specific binding regions in a PET image.
  • the reference mask may be referred to as a Centiloid reference mask or the like.
  • generating masks that correspond to the corrections provided by the trained machine learning model may provide insight into which regions of the PET image that the trained machine learning uses in its determination of corrected SUVR.
  • the corrections provided by the trained machine learning model disclosed herein may reflect changes in either the reference or target mask, or both.
  • This idea may be leveraged to generate masks with SUVRs corresponding to the corrections provided by the trained machine learning model.
  • new masks may be optimised, for each tracer, such that the correlation with the SUVR values provided by the trained machine learning model is optimised (e.g., maximised). This may provide insight into the features deemed important to the trained machine learning model, which may not be obvious in the original PET image, for example.
  • processor 302 may generate one or more masks that correspond to corrections determined by the trained machine learning model.
  • a “mask” in the present context may refer to an image, such as a heat map, saliency map or the like, that highlights relevant regions for machine learning model decision making.
  • Processor 302 may generate the one or more masks by adapting one or more medical images (that may be input into the trained machine learning model) based on a corresponding output of the trained machine learning model. For example, processor 302 may apply an algorithm, mathematical operation or a further machine learning model to the one or more medical images based on the corresponding output to generate the one or more masks.
  • Obtaining these new masks may be formulated as an optimisation problem where, for each tracer, a new reference and target mask may be optimised across the entire dataset so that the resulting SUVR maximises the Pearson correlation coefficient with the corrected SUVRs provided by the trained machine learning model.
  • the Pearson correlation coefficient is a correlation coefficient that measures linear correlation between two sets of data.
  • This optimisation problem may be solved using a gradient descent approach or another type of optimisation technique.
  • the lost function may involve the following two losses:
  • the reference and target masks were initialised using the original Centiloid reference and target mask.
  • the masks were then smoothen using a Gaussian kernel, specifically a 4mm FWHM Gaussian kernel.
  • the masks were updated using the gradient information, and smoothed again using the same Gaussian kernel to reduce sparseness.
  • the masks were also be mirrored to generate symmetric masks.
  • the optimisation leveraged pytorch’s autograd engine for automatic computation of the gradients.
  • Fig. 7 shows the result of this optimisation process.
  • Fig. 7 shows the progression of the SUVR derived from the optimised masks as the optimisation processes continues (i.e., after a number of iterations).
  • the ratio of the Corrected SUVR provided by the trained machine learning model labelled as DeepSUVR SUVR on the x-axis
  • the SUVR determined from the optimised masks approaches 1 (as indicated by the datapoints approaching the diagonal line).
  • Fig. 8a shows the initialised target masks
  • Fig. 8b shows the optimised target masks
  • Fig. 9a shows the initialised reference masks
  • Fig. 9b shows the optimised reference masks.
  • the optimised masks still include salient features that can be seen in the initialised masks.
  • the optimisation process provides more salient features and better definition of these features.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Medical Informatics (AREA)
  • Biomedical Technology (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Evolutionary Computation (AREA)
  • General Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Public Health (AREA)
  • Mathematical Physics (AREA)
  • Computational Linguistics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biophysics (AREA)
  • Molecular Biology (AREA)
  • Radiology & Medical Imaging (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • Primary Health Care (AREA)
  • Epidemiology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Quality & Reliability (AREA)
  • Pathology (AREA)
  • Databases & Information Systems (AREA)
  • Image Analysis (AREA)

Abstract

This disclosure relates to training a machine learning model on longitudinal data. More particularly, this disclosure relates to training a machine learning model on longitudinal medical scan images of a patient study. A processor calculates two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output; and calculating a prediction using the output of the machine learning model for the medical image. The processor then calculates a rate of change between the two or more predictions for the two or more medical images. The processor then creates a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.

Description

"Training deep learning models"
Cross-Reference to Related Applications
[0001] The present application claims priority from Australian Provisional Patent Application No 2023904194 filed on 22 December 2023 and Australian Provisional Patent Application No 2024903443 filed on 23 October 2024, the contents of which are incorporated herein by reference in their entirety.
Technical Field
[0002] This disclosure relates to training a machine learning model on longitudinal data. More particularly but not necessarily exclusively, this disclosure relates to training a machine learning model on longitudinal medical scan images of a patient study.
Background
[0003] Machine learning models, such as deep neural networks, are powerful tools in addressing challenging problems and providing predictions. For example, machine learning models can be used to determine the presence of a disease from an image or a collection of images, such as medical images. Such machine learning models are trained and can then be used to make predictions. Training involves comparing the output of the model with ground truth information (i.e., information that is known to be real or true, and is the target output of the model) and updating the machine learning model based on the comparison.
[0004] Machine learning models, particularly models that utilise deep learning techniques, are trained to minimise a loss value (or a loss function) based on one or a combination of metrics based on some ground truth measurements. The loss value is generally calculated from the difference between the output and the ground truth, with the goal of training being to minimise this loss such that the machine learning model produces an output that is near identical to the ground truth. In one example, for models that are used for image segmentation, the loss value is indicative of the overlap between the predicted mask from the model and a manually defined mask of the structure to be segmented in an image. In this example, the loss value would be formulated in a way that the overlap is maximised by minimising the loss value. Other examples of loss values include the mean square error (MSE) measure between the predicted volume structure and its ground truth volume.
[0005] However, ground truth can often have inherent noise due to many reasons, such as the equipment used to measure the data or other errors. This can make it difficult to train a machine learning model on such ground truth data. In particular, when the training data is longitudinal data, the data is often measured years apart with different machines and calibration settings for each measurement. This introduces noise as there is inconsistency between each successive measurement. Moreover, the noise makes it difficult to train the model on ground truth data that comes from multiple sources. This is often the case when training machine learning models on medical images, where many databases of training data are collated as a significantly large amount of training data is required to accurately train a model.
[0006] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.
[0007] Throughout this specification the word “comprise”, or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.
Summary
[0008] A method for training a machine learning model on longitudinal medical scan images of a patient study comprises: calculating two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output ; and calculating a prediction using the output of the machine learning model for the medical image; calculating a rate of change between the two or more predictions for the two or more medical images; and creating a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
[0009] It is an advantage that the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change as the expected rate of change can be enforced into the machine learning model during training. More specifically, training of the machine learning model leverages information that is not available during inference. This is ideal for situations with noisy ground truth data. Moreover, inference of the trained machine learning model can be performed by applying the model to a single image.
[0010] In some embodiments, the method further comprises determining a polynomial function of rate of change of the biomedical factor and calculating the expected rate of change using the polynomial function.
[0011] In some embodiments, calculating the expected rate of change using the polynomial function comprises applying the polynomial function to a mean value of the biomedical factor calculated by an average of the two or more predictions.
[0012] In some embodiments, determining the polynomial function comprises calculating a rate of change and a mean value for each pair of medical images in the patient study.
[0013] In some embodiments, the loss value further comprises a change penalty value reflecting a difference between the two or more predictions, the change penalty value penalising an unexpected temporal trend of the biomedical factor related to the expected rate of change. [0014] In some embodiments, the loss value further comprises a correction penalty value reflecting a difference between the two or more predictions and two or more corresponding ground truth values, the correction penalty value penalising the two or more predictions from deviating from the two or more corresponding ground truth values.
[0015] In some embodiments, the loss value further comprises a correction penalty value reflecting a difference between the two or more outputs and an over-correction factor , the correction penalty value penalising the two or more outputs from deviating from the overcorrection factor.
[0016] In some embodiments, the loss value further comprises an additional penalty value calculated by one or more of a difference between each of the two or more outputs of the machine learning model and a respective ground truth measurement; and a mean square error between each of the two or more outputs and the respective ground truth measurement.
[0017] In some embodiments, the loss value is calculated by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
[0018] In some embodiments, each of the medical images of the patient study has a timestamp and at least two of the two or more medical image have non-adjacent timestamps.
[0019] In some embodiments, the method further comprises calculating a further two or more predictions of the biomedical factor by performing for each of a further two or more medical images of the patient study: applying the trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating a further rate of change between the further two or more predictions; and creating a further trained machine learning model by updating weights of the trained machine learning model to minimise a further loss value based of the further rate of change.
[0020] In some embodiments, the method further comprises calculating another two or more predictions of the biomedical factor by performing for each of two or more medical images of another patient study different to the patient study: applying the trained machine learning model or further trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating another rate of change between the other two or more predictions; and creating another trained machine learning model by updating weights of the trained machine learning model or further trained machine learning model to minimise another loss value based on the other rate of change.
[0021] In some embodiments, the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of the patient based on the output of the trained machine learning model.
[0022] In some embodiments, each of the medical images is one of a magnetic resonance image; positron emission tomography (PET) image; or computed tomography (CT) image.
[0023] In some embodiments, the biomedical factor is one of: medical image quantification factor; a biological structure value; or a biomarker.
[0024] In some embodiments, the medical image quantification factor is a Centiloid or standardised uptake value ratio scaling factor and the biological structure value is volume of a biological structure visible in the medical image.
[0025] In some embodiments, each of the medical images is a PET image and the method further comprises performing PET quantification by applying the trained machine learning model to a PET image of a patient.
[0026] In some embodiments, the medical image quantification factor is a Centiloid or standardised uptake value ratio scaling factor; and the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of Alzheimer’s disease based on the Centiloid or standardised uptake value ratio scaling factor that is an output of the trained machine learning model.
[0027] In some embodiments, the method further comprises generating one or more masks that correspond to corrections determined by the trained machine learning model. [0028] In some embodiments, the machine learning model is a neural network comprising one or more convolutional layers.
[0029] Software, when installed on a computer and executed by the computer, causes the computer to perform the above method.
[0030] A system for training a machine learning model on longitudinal medical scan images of a patient study comprises: a processor configured to: calculate two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output; and calculating a prediction using the output of the machine learning model for the medical image; calculate a rate of change between the two or more predictions for the two or more medical images; and create a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
[0031] Optional features provided in relation to the method, equally apply as optional features to the software and the system.
Brief Description of Drawings
[0032] An example will be described with reference to the following drawings:
[0033] Fig. 1 shows a plot of calculated Centiloid values over time calculated using longitudinal PET scan images.
[0034] Fig. 2 shows a plot of the rate of change of the Centiloid value verses the mean value of the Centiloid value. [0035] Fig. 3 illustrates a system for training a machine learning model on longitudinal medical scan images of a patient study.
[0036] Fig. 4 illustrates a method for training a machine learning model on longitudinal medical scan images of a patient study.
[0037] Fig. 5 is an example machine learning model with a convolutional neural network (CNN) architecture.
[0038] Fig. 6 shows the progression of the calculated rate of change of the scaled Centiloid values as the training of the machine learning model increases, in the example embodiment.
[0039] Fig. 7 shows the progression of the Standardized Uptake Value Ratio (SUVR) scaling factors as the optimisation processes described herein continues.
[0040] Fig. 8a shows initialised target masks used to generate optimised target masks.
[0041] Fig. 8b shows the optimised target masks with similar SUVR scaling factors to the SUVR provided by the trained machine learning model disclosed herein.
[0042] Fig. 9a shows initialised reference masks used to generate optimised reference masks.
[0043] Fig. 9b shows the optimised reference masks with similar SUVR scaling factors to the SUVR provided by the trained machine learning model disclosed herein.
Description of Embodiments
[0044] The disclosed system and method relate to training a machine learning model on longitudinal data to address the problem of noisy ground truth training data. The disclosed system and method utilise the inherent or expected properties of the longitudinal data as a source of additional information to reduce the effects of noisy ground truth data. More specifically, longitudinal data that exhibits an expected rate of change of a factor, such as a biomedical factor, can be used to train a machine learning model by formulating the training in a way that reflects this expected rate of change.
[0045] Longitudinal data includes successive measurements of the same subject over a period of time. For example, longitudinal data may be a positron emission tomography (PET) image or magnetic resonance (MR) image of the brain repeated on the same participant every year for N number of years. This is commonly performed on patients who are at risk of Alzheimer’s disease to determine the progression of the disease by studying the amount of amyloid plaques in the brain over time. As Alzheimer’s disease is a progressive disorder, it is expected that the amount of amyloid plaques in the brain increases over time, which is an expected trend observed from the longitudinal data of a patient with this disease.
[0046] To utilise the inherent or expected properties for training a machine learning model, the disclosed system and method uses a class of penalty values (as referred to as longitudinal constraints) in the loss value. More specifically, the penalty values exploit the expected rate of change of the longitudinal data by penalising the unexpected rate of change. A simple longitudinal constraint could be to penalise negative (positive) changes between baseline and follow-up when the change is expected to be positive (negative). These changes could be biomedical factors such as the PET quantification within a mask or the volume of a segmentation mask, or other (more involved) factors. For example, in relation of Centiloid values, as the expected trend is that the rate of change is positive (i.e., the Centiloid values increase over time), a penalty value can be formulated to penalise training the machine learning model on negative rates of change.
[0047] This disclosure also provides a framework to include these constraints on a single timepoint inference model. In other words, while the model is trained to use the longitudinal information inherent to the longitudinal data for training (e.g., multiple images), the trained model does not require the longitudinal information during inference. This means that the trained model can be used to calculate predictions on a single timepoint during inference (e.g., a single image). One of the main advantages of the disclosed system and method is that it enables the inclusion of longitudinal constraints without relying on multiple timepoints during inference. [0048] The proposed constraints or penalty values can also complement existing losses by ensuring longitudinal consistency in the training set, instead of solely relying on a noisy “ground truth” data. When properly trained, this also translates into better longitudinal consistency in the testing set, even though each image is inferred (tested) independently from all other timepoints. In particular, as longitudinal data can be captured over many years using machines with different calibration setting each time or completely different machines all together, or in the case of PET imaging, using different PET tracers, the proposed penalty values enable a normalisation of the longitudinal data, which in turn enables machine learning models to be trained on such noisy data.
[0049] For example, the overall trajectory of expected change for a mean measure can be computed from a large population analysis. This overall trajectory can use the error between the actual change and expected change as an additional loss. Since any information can be added in the loss, additional information can also be included into the loss, such as demographics. In another example, where different trajectories are expected for different clinical diagnosis, these different trajectories can be used to define different losses for different clinical groups. Since this information is not used at inference time (in other words, it is only used to compute the loss during training), this limits the risk to overfit the model to the data.
[0050] The disclosed system and method are applicable to many different machine learning (specifically, deep learning) applications, where longitudinal data is available, and where there are known longitudinal constraints that can be applied based on expected trends in the longitudinal data. For example, the machine learning application may be image segmentation, where there is a model of change over time of expectations with respect to increase/decrease of volume of structure over time. In another example, when training a model to predict the volume of a structure that is expected to shrink over time, this information can be used in the loss function in combination with other standard losses, and in this case, penalise the volume increasing over time. This is particularly useful for problems where the ground truth might not be very reliable. For example, this is particularly useful for PET quantification, since PET images can be very noisy. [0051] In particular, PET quantification is used to analyse PET images and extract quantitative information from these images. PET quantification for the study of Alzheimer’s disease involves a normalization of the PET images to the Centiloid scale, which is described as the standard scale for PET images acquired using an amyloid PET tracer. In a general sense, PET images are typically normalised using the Standardized Uptake Value Ratio or SUVR. The SUVR is defined as the ratio of a region of the PET image containing specific binding (a target region where the PET tracer will bind to the target), and a region of the PET image containing non-specific binding (as referred to as a reference region). In the case of amyloid imaging, those are defined as a region containing amyloid, typically the neocortex, and a region containing no amyloid, typically the cerebellum. Since there exists different PET tracers that bind to amyloid, each having different pharmacokinetic properties, the Centiloid scale was developed to harmonise the quantifications (SUVR) from all amyloid PET tracers into the same scale using a linear transform based on head-to-head paired data. In other words, while the SUVR is typically used for most PET tracer quantification, in the specific case of Amyloid PET imaging, the Centiloid scale is typically used as it allows to quantify different PET tracers using the same scale. Due to the progressive nature of Alzheimer's disease, one would expect that the unit of the Centiloid scale increases over time as the amount of amyloid deposits increases in the brain. However, by plotting the calculated Centiloid values over time, it can be shown that this is not always the observed trend in the data.
[0052] Fig. 1 shows a plot of calculated Centiloid values over time calculated using longitudinal PET scan images. The slope of the lines represents the “rate of change” of the Centiloid values. The blue lines, e.g. at 101, represent the expected trend of the Centiloid values increasing over time, while the red lines, e.g. at 102, represent the unexpected trend of the Centiloid values decreasing over time, which is a result of noise in the PET scan images. This unexpected trend of the Centiloid values is often overlooked, as the Centiloid value calculation usually occurs in the background during typical PET quantification. The noise that causes the unexpected trend in the Centiloid values may be due to noise in the PET images, noise from the reference region used to calculate the Centiloid values and changes to the PET tracer or scanner, for example. [0053] Fig. 2 shows a plot of the rate of change of the Centiloid value verses the mean value of the Centiloid value. From the extended trend of the Centiloid values, there should not be any negative rates of change (i.e., each plot point should be above horizontal line 201). Fig. 2 also shows curve 202, which represents a calculated curve of mean value versus rate of change. Curve 202 may be obtained by calculating the mean value and rate of change for some or each piece of data in a set of longitudinal data. More specifically, curve 202 may be produced using a regression algorithm applied on the data of Fig. 2. It is expected that the data points of Fig. 2 should not deviate far from curve 202, given that curve is the expected trend of the data.
[0054] Given the expected trend of the longitudinal data (in this example, that the Centiloid values are expected to increase, such that the rate of change is positive), a training of a machine learning model can be designed to estimate a correction of the Centiloid values in such a way to penalise negative rates of changes, penalise over-correction and penalise distance from curve 202, for example. Training of the machine learning model is penalised in these ways by introducing constraints or penalty values which are formulated in a way to represent the penalisation of these factors.
[0055] While the disclosed system and method are advantageous for combating noisy ground truth data during training, there are also other advantages. For example, the volumes of structures present in training images may not be well defined (such as volume of white matter hyperintensity which have poor inter-rater reproducibility), which is compensated during training through the disclosed method. Another advantage is that the training can be performed using longitudinal data that is temporally sparse (taken far apart). This is particularly advantageous for training the machine learning model on medical images such as MRI or CT images, as data is only collected every 1 or 2 years. It is noted that the machine learning model can also be trained on less sparse data using the disclosed method, such as longitudinal data that collected every day or every week or every month, for example.
[0056] While this disclosure is generally directed towards the medical images and application of machine learning models in the medical field, it is noted that the disclosed system and method are not limited to applications in the medical field and medical images. More specifically, this disclosure relates more generally to training a machine learning model on longitudinal data, where the longitudinal data may include medical images. However, in other instance, the longitudinal data may include other images. For example, any model for estimating a quantification/value from an image or any other input can be trained using the disclosed system and method, as long as there is some prior knowledge of the expected changes over time.
[0057] In other instances, the longitudinal data may not include images at all, but rather be some other form of data. For example, it may be expected that the rate of change of water flowing through a dam decreasing over time due to the predicted increase of droughts occurring in that area. However, it may be difficult to accurately measure the exact volume of water flowing through the dam, which introduces a source of noise. The volume may also fluctuate unexpectedly due to flash storms, which also introduces noise. The disclosed system and method can compensate for the noise in the data by utilising the expected rate of change of the data during training.
System for training
[0058] Fig. 3 illustrates an example system 300 for training a machine learning model on longitudinal medical scan images of a patient study. Fig. 3 is one example of a configuration of system 300. However, system 300 is not strictly limited to this configuration and this may be one possible embodiment of system 300. The medical scan images, or simply medical images, may be magnetic resonance (MR) image, positron emission tomography (PET) image or computed tomography (CT) image. Throughout this disclosure, medical scan images or medical images may be referred to simply as images or image data, noting that system 300 is equally applicable to images other than medical images.
[0059] A patient study is the study of a single patient where, for example, a medical image of the patient’s brain is captured at regular time intervals, such as annually. It is noted that the machine learning model can be trained on longitudinal medical scan images of multiple patient studies, thereby being trained on data from multiple patients. This is because the penalty values indicative of an expected rate of change reduces the noise that results from inconsistencies in the data between patients. Machine learning model usually require large amounts of training data, so it is an advantage that the machine learning model can be trained on longitudinal data from multiple patient studies. In embodiments outside of the medical field, the patient study may be a study of a subject over time.
[0060] System 300 comprises a device 301, which may be smartphone, computer or similar device. Device 301 comprises a processor 302 connected to a program memory 303 and a data memory 304. In other examples, device 301 is implemented in a distributed computing architecture, such as a cloud computing provider. Processor 302 may receive data through a variety of interfaces, which includes memory access of volatile memory, such as cache or RAM, or non-volatile memory, such as an optical disk drive, hard disk drive, storage server or cloud storage. The program memory 303 is a non-transitory computer readable medium, such as a hard drive, a solid state disk or CD-ROM.
[0061] Software, that is, an executable program stored on program memory 303 causes processor 302 to perform methods for training a machine learning model on longitudinal medical scan images of a patient study. For example, once executed, the software causes processor 302 to calculate two or more predictions of a biomedical factor by performing for each of two or more of the medical images: (i) applying the machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output of the machine learning model for the medical image, calculate a rate of change between the two or more predictions for the two or more medical images and create a trained machine learning model by updating weights of the machine learning model to minimise a loss value.
[0062] The data memory 304 may store image data and retrieve the image data for later use. The image data may be indicative of a two-dimensional image, which may be stored on data memory 304 as Joint Photographic Experts Group (JPEG) format, RAW image format or a similar/equivalent image format, for example. In some examples, the image data may be an RGB image, a multispectral image, a hyperspectral image, an infrared image, or a two- dimensional cross-sectional. The image data may be indicative of a three-dimensional image, such as a PET, MR or CT image. Data memory 304 may store image data indicative of the three-dimensional image data as a single-file DICOM (Digital Imaging and Communications in Medicine), multi-file DICOM format or a similar/equivalent image format, for example. [0063] The machine learning models described herein (such as the machine learning model before, during and after training) may be stored on data memory 304. The data memory 304 may also store predictions calculated by processor 302 performing any one of the disclosed methods, or any other variable or data necessary to perform such methods. For example, predictions, machine learning model outputs, loss values, rates of changes are calculated when processor 302 performs the disclosed method and can be stored on data memory 304.
[0064] It is to be understood that any receiving step may be preceded by processor 302 determining or computing the data that is later received. For example, processor 302 may then store the image data on data memory 304, such as on RAM or a processor register. Processor 302 then requests the data from the data memory 304, such as by providing a read signal together with a memory address. The data memory 304 provides the data as a voltage signal on a physical bit line and processor 302 receives the image data as the input image via a memory interface.
[0065] System 300 may also comprise a scanning machine 305, a server 306 and a monitor 307, which are each in communication with the device 301 via an input/output (I/O) port 308. In an example, the scanning machine 305 scans a patient by incrementally capturing two- dimensional cross-sectional images of the patient. The collection of cross-sectional images represents image data that is indicative of a scan and represents a 3D scan of the patient. Once scanning machine 305 creates the image data of the patient, processor 302 receives the image data via the I/O port 308. Processor 302 then performs the disclosed method using the image data received from the scanning machine 305 and trains the machine learning model.
[0066] If the image data is indicative of a MRI scan, then the image data may be created using an MRI scanning machine. MRI scanning machines consists of a table that slides into a cylinder. Inside the cylinder is a magnet that creates a strong magnetic field within the cylinder. The strong magnetic field is used to align the protons in the water molecules in a patient’s body, as the protons possess a magnetic moment due to their intrinsic spin. Once the protons align with the strong magnetic field, a radio frequency signal is sent from the MRI scanning machine, which excites the protons and causes them to align against the strong magnetic field. Once the radio frequency signal is turned off, the protons relax and re-align with the strong magnetic field, thereby emitting electromagnetic radiation, which is detected by the MRI scanning machine.
[0067] A patient lies on a motorised table that moves horizontally, in increments, through the cylinder and captures a single MR image at that increment, which corresponds to a two- dimensional cross-sectional image (virtual “slice”) of the patient’s body. Each two- dimensional cross-sectional image corresponds to a different region of the patient’s body as they move horizontally, in increments, through the cylinder. The collection of two- dimensional cross-sectional images also has an associated spatial sequence, in which the cross-sectional images are ordered along. The horizontal direction that the patient moves within the cylinder is indicative of a dimension in space along the spatial sequence of the two- dimensional cross-sectional images. The collection of two-dimensional cross-sectional images provides a three-dimensional representation of a patient’s internals, which aids physicians in diagnosing internal diseases and conditions.
[0068] The two-dimensional cross-sectional image consists of black and white regions, as well as shades of grey. These regions correspond to the rate at which excited protons return to their equilibrium state after being exposed to radio frequency (equilibrium state corresponds to the spin state which aligns with the strong external magnetic field), as well as the amount of energy released, changes depending on the environment and the chemical nature of the molecules. These factors can be differently weighted in the MR image, to give more contrast to certain tissues in the image. For example, in a Tl-weighted (Tlw) image, regions of high signal representing white regions may correspond to protein-rich fluid, whereas in T2- weighted (T2w) images, these regions may correspond to water rich fluids in the patient’s body.
[0069] If the image data is indicative of a CT scan, then the image data may be created using a CT scanning machine. CT scanning machines use a rotating X-ray tube and a row of detectors placed in a gantry of the machine to measure X-ray attenuations by different tissues inside the body. X-ray attenuations refer to when X-ray photons, with wavelengths ranging from 10 picometers to 10 nanometres, are absorbed by the tissue when the X-ray photons pass through the body of a patient. [0070] A patient lies on a motorised table that moves horizontally, in increments, through the rotating X-ray tube and captures multiple X-ray measurements from different angles at each increment. The multiple X-ray measurements are then processed on a computer using reconstruction algorithms to produce tomographic two-dimensional cross-sectional images (virtual “slices”) of a body. Each two-dimensional cross-sectional image corresponds to a different region of the patient’s body as they move horizontally, in increments, through the rotating X-ray tube.
[0071] Similar to MR images, the collection of two-dimensional cross-sectional images also has an associated spatial sequence, in which the cross-sectional images are ordered. The horizontal direction that the patient moves within the cylinder is indicative of a dimension in space along the spatial sequence of the two-dimensional cross-sectional images. Similar to MR images, the collection of two-dimensional cross-sectional images provides a three- dimensional representation of a patient’s internals, which aids physicians in diagnosing internal diseases and conditions.
[0072] The two-dimensional cross-sectional (X-ray) images consists of black and white regions, as well as shades of grey. These regions correspond to the areas of the body that have attenuated the X-rays. For example, structures such as bones readily absorb X-rays and, thus, produce high contrast on the X-ray detector. As a result, bony structures appear whiter than other tissues in the X-ray image. Conversely, X-rays travel more easily through less radiologically dense tissues such as fat and muscle, as well as through air-filled cavities such as the lungs. These structures are displayed in shades of grey in the X-ray image. Regions where little to no attenuation has occurred will show up as black regions in the X-ray image.
[0073] If the image data is indicative of a PET scan, then the image data may be created using a PET scanning machine. Before the patient is placed into the PET scanning machine, a radiopharmaceutical (a radioisotope attached to a drug) is injected into the body as a radiotracer. When the radiopharmaceutical undergoes beta plus decay, a positron is emitted, and when the positron interacts with an ordinary electron present in the body, the two particles annihilate and gamma rays are emitted. These gamma rays are detected by gamma cameras to form a three-dimensional image, in a similar way that an X-ray image is captured. [0074] Some parts of the body absorb more of a particular radiotracer than other parts of the body. For example, diseased cells, such as cancer cells, in a patient’s body absorb more of the radiotracer than healthy ones do. In another example, the brain is normally a rapid user of glucose, where radiotracer may be glucose containing the radioisotope. Areas where there is a higher detection of gamma ray (as referred to as “hot spots”) will appear on the PET images as hotter colours, such as yellow and red, as opposed to areas of lower gamma ray detection, which will appear as blue. As such, the size and location of these “hot spots” enable the diagnosis of disease from the PET images.
[0075] Processor 302 may also receive other longitudinal data from other sources. For example, system 300 may comprise a RGB camera (not shown) configured to capture RGB image data intermediately over a period of time. In another example, system 300 may comprise an inertial measurement unit (IMU) which measure position and pose data of a machine intermediately over a period of time. Moreover, system 300 may comprise a set of sensors which provide measurements to processor 302 intermediately over a period of time. As such, the disclosed system and method are not limited to applications involving medical images.
[0076] The image received from the scanning machine 305 may be stored on the data memory 304 before and after the disclosed method is performed by processor 302. In another example, the server 306 may store image data. Additionally, image data may be communicated to the server 306 and the server 306 may then communicate the image data to device 301. The image data may be received from a source external to system 300, such as another system located remotely to system 300. The remote system may be located within the same facility as system 300 or may be completely remote from system 300. The image data may be stored in the server 306 as a single-file DICOM or multi-file DICOM format.
[0077] The server 306 may be in communication with the scanning machine 305, or may be in communication, over a communications network. Similarly, processor 302 receives the image data via the input/output port 308 and performs the disclosed method on the image data. After processor 302 performs the disclosed method on the image data, the trained machine learning model may be communicated to the server 306 by communicating the parameters of the trained machine learning model. The server 306 may then communicate these parameters to an external system or store them.
[0078] System 300 may also comprise a monitor 307, which may be configured to display the image data from the scanning machine 305 and from the server 306. In this way, processor 302 communicates the image data from the scanning machine 305 and from the server 306 to the monitor via the input/output port 308. Monitor 307 may be further configured to the result of applying the trained machine learning model of a new medical image during inference. In this way, processor 302 communicates the inference result to the monitor via the I/O port 308.
[0079] While only one device 301 is depicted in Fig. 3, there may be a network of many devices that may communicate with the server 306. The network of devices may also communicate directly with one another via the Internet, any other means of wireless communication or wired connections. In an example, multiple devices may each capture image data of a patient. Each instance of image data from the multiple devices may then be stored on server 306 for later use. In some embodiments, a database of image data captured from multiple devices may be stored on server 306.
[0080] It is noted that server 306 may have similar functionality to that of device 301. Server 306 may perform parts of the disclosed methods described herein. For example, program memory may comprise software, that is, an executable program stored on program memory, and causes the processor to train a machine learning model on longitudinal data.
[0081] Software may provide a user interface presented to the user on device 301. The user interface is configured to accept input (via buttons or text fields etc.) from the user, via a touch screen or a device attached to device 301 such as a keyboard or computer mouse. These devices may also include a touchpad, an externally connected touchscreen, a joystick, a button, and a dial. In an example, device 301 may display multiple instances of image data, and the user may choose one of the multiple instances of image data such as medical images, by processor 302. The user may choose one of the multiple instances of image data by interacting with the touch screen or inputting the selection with a keyboard or computer mouse. [0082] Processor 302 may receive or send data, such as image data, from data memory 304 as well as from the I/O port 308. In one example, processor 302 sends image data from the device 301 to server 306 via I/O port 308, such as by using a Wi-Fi network according to IEEE 802.11. The Wi-Fi network may be a decentralised ad-hoc network, such that no dedicated management infrastructure, such as a router, is required or a centralised network with a router or access point managing the network. System 300 may further be implemented within a cloud computing environment, such as a managed group of interconnected servers hosting a dynamic number of virtual machines.
[0083] Although I/O port 308 is shown as single entity, it is to be understood that any kind of data port may be used to receive data, such as a network connection, a memory interface, a pin of the chip package of processor 302, or logical ports, such as IP sockets or parameters of functions stored on program memory 303 and executed by processor 302. The parameters of functions may be stored on data memory 304 and may be handled by-value or by-reference, that is, as a pointer, in the source code.
Method for training
[0084] Fig. 4 illustrates method 400 for training a machine learning model on longitudinal medical scan images of a patient study. Fig. 4 is to be understood as a blueprint for the software program and may be implemented step-by-step, such that each step in Fig. 4 is represented by a function in a programming language, such as Python, C++ or Java. The resulting source code is then compiled and stored as computer-executable instructions on program memory 303, which causes processor 302 to perform method 400. Before the process of training the machine learning model begins, processor 302 may initialise the weights of the machine learning model (which define the machine learning model). For example, processor 302 may use random weights at the beginning of the train process, which are then updated during the training process.
[0085] In some examples, the machine learning model may be a neural network. In further examples, the machine learning model may be a neural network comprising one or more convolutional layers. Such a machine learning is known as a convolutional neural network (CNN). The CNN is ideal for applications involving images as it accounts for the positioning and shape of objects in the image data. However, other types of machine learning models are equally applicable here. For example, the machine learning model may be implemented as a K nearest neighbour model, decision tree and support vector machine.
[0086] Mathematically, a convolution is an integration function that expresses the amount of overlap of one function g as it is shifted over another function/. Intuitively, a convolution acts as a blender that mixes one function with another to give reduced data space while preserving the information. A CNN performs convolutions with filters (matrix / vectors) and the filters comprise learnable parameters that are used to extract low-dimensional features from an input data. They have the property to preserve the spatial or positional relationships between input data points. CNNs exploit the spatially-local correlation by enforcing a local connectivity pattern between neurons of adjacent layers. In this example, the CNN architecture may comprise an optimal number of convolutional layers, filter size and stride length.
[0087] Intuitively, a convolution is the step of applying the concept of a sliding window (a filter with learnable weights) over the input and producing a weighted sum (of weights and input) as the output. The weighted sum is the feature space which is used as the input for the next layers. More specifically, each convolutional layer comprises a filter, which calculates a weighted sum of adjacent pixel values. For example, a 2 x 2 filter comprises 4 weights, which are the coefficients of the filter. The filter starts at an initial position in the image data structure, multiplies each pixel value in the image data structure with the respective filter coefficient and adds the results. Finally, the filter stores the resulting number in an output pixel. In this sense, the output pixel value of each of the one or more convolutional layers comprises a weighted sum of input pixel values. The weights in the weighted sum correspond to the coefficients of the filter. Then, the filter moves by one pixel along one direction in the data structure and repeats the calculation for the next voxel of the output image. For example, if the stride is 2, the filter will move by two pixels along one direction. That direction may be an x-dimension or a y-dimension.
[0088] A single convolutional layer may use multiple filters, where each filter corresponds to a different feature. The corresponding feature of each filter is determined during training of the CNN. More specifically, each filter may contain weights, which are adjusted during a training process through a backpropagation process, for example. The filters are used to quantitatively determine the contribution of a particular feature on the output of the CNN. For a convolutional layer that occurs at the start of a CNN architecture, the filters of this convolutional filter may correspond to simple features. However, convolutional layers that occur later in the CNN architecture may exhibit more complex or abstract features. As an example, the complex or abstract features may be a combination of the simple features from the previous convolutional layer.
[0089] The CNN architecture may also comprise a batch normalization layer and rectifier linear unit (ReLU) activation function. Further, the output of the last convolution layer may be subjected to Global Average Pooling (GAP). Even further, a fully connected layer with Sigmoid activation may also be used on the output layer. However, other activation functions may be used. These activation functions include, but are not limited to, a binary step function, a tanh function, a ReLU function or a softmax function.
[0090] Fig. 5 is an example CNN architecture 500. This example CNN architecture 500 is designed to predict parameters from an MR image for dynamically transformation of the intensity contrast of an input MR image to adapt to a segmentation task. More specifically, the example CNN predicts the parameters for a power function transformation and a piece- wise linear transformation. As can be seen in Fig. 5, the example CNN architecture 500 includes three convolutional blocks 511, 512, 513 and three fully connected (FC) blocks 521, 522, (523 and 524 are collectively one FC block). This CNN architecture 500 is used to provide results for the disclosed method presented later in the disclosure.
[0091] To begin method 400, processor 302 may select 401 two or more images from the longitudinal data of a patient study. Preferably, the two or more images come from the same patient, rather than the two or more images being from different patients. The two or more images may be randomly chosen from the patient study. For example, the two or more images do not have to be images that have been captured consecutively.
[0092] Processor 302 then calculates 420 two or more predictions of a biomedical factor each of two or more of the medical images. The biomedical factor has an expected trend that is known beforehand and is utilised to train the machine learning model. For each of the two or more of the medical images, processor 302 applies 421, 426 the machine learning model to the medical image to calculate 422, 427 an output, then calculates 423, 428 a prediction using the output of the machine learning model for the medical image. Preferably, processor 302 calculates 420 the two or more predictions using the same machine learning model and more specifically, processor 302 calculates 420 the two or more predictions using a model that has the same architecture and weights.
[0093] In some embodiments, processor 302 calculates 420 two or more predictions of a factor that is not a biomedical factor. The factor may be any variable that has some expected trend (e.g., expected to change over time). For example, the factor may be a physical factor such as the volume of water in a dam. It is noted that while this disclosure describes embodiments related to a biomedical factor, these embodiments are equally applicable to a factor other than a biomedical factor.
[0094] It is also noted that the calculation of the prediction for each of the two or more medical images does not have to be performed sequentially, in the sense that the first prediction is calculated then the second prediction is calculated and so on. The prediction for each of the two or more medical images may be calculated in parallel, as the calculations are not necessarily dependent on each other. For example, each prediction may be calculated in parallel across many CPUs and multiple cores. More specifically, each prediction may be calculated on a single CPU efficiently by multi-platform shared-memory parallel programming, such as OpenMP, which utilises the multiple computing cores of the CPU.
[0095] In some embodiments, processor 302 calculates 423, 428 the prediction by applying a function to the output of the machine learning model. For example, the function may be a polynomial, exponential or trigonometric function or combination thereof. In other embodiments, processor 302 calculates 423, 428 the prediction by multiplying the output by a factor or measurement. For example, the output may be a scaling factor to correct a specific measurement and hence, the prediction corresponds to the corrected measurement. In further embodiments, the output may simply correspond to the prediction, which can be considered as processor 302 calculating 423, 428 the prediction by multiplying the output by 1.
[0096] The biomedical factor may be one of medical image quantification factor, a biological structure value or a biomarker. For example, the biomedical factor may be a medical image quantification factor, which may be a Centiloid value or standardised uptake value ratio scaling factor. In other examples, the biomedical factor may be a biological structure value, which may be a volume of a biological structure visible in the medical image, such as the grey matter in the brain. The biomarker may be blood pressure or heart rate, for example.
[0097] After calculating 420 the two or more predictions, processor 302 calculates 430 a rate of change between the two or more predictions for the two or more medical images. In some embodiments, processor 302 calculates 430 the rate of change based on a difference between the two or more predictions. In other embodiments, processor 302 calculates 430 the rate of change based on a difference between the two or more predictions and a difference between a variable on which the rate of change is based. For example, if there are only two predictions, processor 302 may calculate the rate of change using a difference of the predictions divided by the difference between the timepoints of each respective medical image.
[0098] The calculated rate of change is indicative of a trend in the longitudinal data. As such, processor 302 may calculate the rate of change explicitly by calculating a ratio of the change in the predictions and the change in an independent variable (such as time). However, processor 302 may also calculate the rate of change implicitly by calculating a difference between the two or more predictions, as the medical images may inherently contain change information (such as temporal information). For example, two medical images in a patient study may be taken a year apart and hence, the two medical images have inherent temporal information, given that they are temporally related. As such, the rate of change may be calculated by a difference of the two predictions of the respective medical images, as the medical images inherently contain temporal information.
[0099] After calculating 430 the rate of change, processor 302 calculates 440 a loss value to train the machine learning model. In particular, the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change. Reflecting a difference generally means that the penalty value is indicative of the difference. Reflecting a difference may mean that the penalty value is related to the difference, such as being directly proportional to the difference. Reflecting a difference may also mean that processor 302 calculates the penalty value using the difference, such as by applying a mathematical function to the difference. By calculating the loss value using this penalty value, training of the machine learning model leverages information that is not available during inference.
[0100] In some examples, the expected rate of change may be a function of time (such as linear, polynomial, exponential, etc.) or may be a second order rate (i.e., the rate of the rate of change) and processor 302 calculates the second order rate of change. In other examples, the expected rate of change may be zero, indicating that the biomedical factor remains constant over time. As such, the penalty value may be formulated in such a way to penalise the rate of change deviating from zero.
[0101] Calculating a loss value may be referred to as applying a loss function. In other words, the outcome of applying a loss function is the loss value. The loss value and loss function may also be referred to as cost value and cost function, respectively. The loss function may be applied to the outputs and other factors, such as the timestamps of two or more medical images. The loss value may incorporate many types of additional information that is not available at inference time. For example, the difference curves based on clinical diagnostic, age or any other relevant information may be used to train the machine learning model using the loss value.
[0102] Processor 302 then updates 450 weights of the machine learning model to minimise a loss value, thereby creating a trained machine learning model. This process is generally referred to as “training”. The longitudinal data (i.e., the training data) comprises label data, such as image data that has been labelled with a class, where the label for each item of training data may be manually provided by a user. Training data may also or otherwise be obtained from a database containing already labelled data, for which manually labelling the data is not needed. Training the machine learning model in this way is referred to as “supervised learning”, in the sense that the outcome of the machine learning model is validated and the weights of the machine learning model are adjusted accordingly. The weights may be updated through the process of backpropagation. The weights may also be updated through a gradient descent process. [0103] It should be noted that “supervised learning” is not the only method that can be used to train the machine learning models. For example, “semi-supervised learning” can be used in situations where only a relatively small amount of training data can be obtained. In some examples, a training set may be augmented to produce more training data, which is advantageous if a large training data is difficult to obtain. A training set may be augmented by randomly resizing, cropping, flipping, or rotating the image data in the training set, for example.
[0104] It is noted that during method 400, the rates of change (e.g., the rate of change between the two or more predictions) may only be used during training to update the machine learning model, rather than at inference time as the prediction provided by the machine learning model may correspond to single timepoint measurements. As such, the rates of change may not be used at inference time to obtain the prediction from the trained machine learning model. Rather, single timepoint measurements may be used at inference time to generate corresponding predictions using the trained machine learning model.
Repeating the training method
[0105] Method 400 is preferably repeated multiple times using two or more different images. In general, the more repetitions of method 400, the more accurate the machine learning model becomes at making a prediction during inference. Moreover, training the machine learning model on a range of different images or training data also generally makes the machine learning model more accurate at making a prediction during inference. In this disclosure, reference to “trained machine learning model” means that the initial machine learning model has undergone at least one iteration of method 400. However, “trained machine learning model” may also mean that the initial machine learning model has undergone multiple iterations of method 400. Other terms such as “further trained machine learning model”, “another trained machine learning model” and “fully trained machine learning model” may also be used throughout this disclosure to indicate that the machine learning model has undergone multiple training iterations.
[0106] As processor 302 may perform method 400 multiple times using two or more different images, processor 302 may calculate a further two or more predictions of the biomedical factor by, for each of a further two or more medical images of the patient study: (i) applying the trained machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output. Processor 302 may then calculate a rate of change between the further two or more predictions; and creates a further trained machine learning model by updating weights of the trained machine learning model to minimise a further loss value based on the further rate of change. As such, the machine learning model is trained on different images of the same patient study. This process may be repeated using each permutation of the two or more images in the patient study.
[0107] In some embodiments, processor 302 may further train the machine learning model using images from another patient study, which is different to the patient study used to initially trained the machine learning model. More specifically, processor 302 may calculate another two or more predictions of the biomedical factor by, for each of two or more medical images of another patient study different to the patient study: (i) applying the trained machine learning model or further trained machine learning model to the medical image to calculate an output; and (ii) calculating a prediction using the output. Processor 302 may then calculate another rate of change between the other two or more predictions; and creates another trained machine learning model by updating weights of the trained machine learning model or further trained machine learning model to minimise another loss value based on the other rate of change.
[0108] The machine learning model may be agnostic to the order of the two or more medical images of the patient study. As such, the machine learning model is unable to learn the order as each image is processed the same way, with the only difference being how the loss is computed. It is beneficial for the machine learning model not to learn the order of the input images, as it is desired that the model learns the expected trend in the data (i.e., the expected rate of change), rather than the sequence of the images. This provides better predictions during inference.
[0109] To further re-enforce this, the order of the inputs (i.e., the two or more medical images) can be randomly switched. As such, in some embodiments, each of the medical images of the patient study has a timestamp and at least two of the two or more medical image have non-adjacent timestamps. Non-adjacent timestamps being non-consecutive. For example, if the patient study contained longitudinal data of a brain CT image captured every year, then the non-adjacent timestamps would be the non-consecutive years.
Calculating the expected rate of change
[0110] As the penalty reflects a difference between the rate of change and the expected rate of change, processor 302 may calculate the expected rate of change. Processor 302 may calculate the expected rate of change from the longitudinal data. For example, processor 302 may calculate a rate of change for each pair of images in the patient study. Processor 302 may then average the calculated rates of change to calculate the expected rate of change. Processor 302 may also use other techniques such as regression to calculate the expected rate of change from the longitudinal data.
[0111] In some embodiments, processor 302 determines a polynomial function of rate of change of the biomedical factor and calculates the expected rate of change using the polynomial function. This polynomial function represents the expected trend of the longitudinal data and hence, can be used to calculate the expected rate of change of the biomedical factor based on the two or more medical images used to train the machine learning model. For example and with reference to Fig. 2, curve 202 represents a polynomial function and is based on the longitudinal data used to train the machine learning model. It is noted that this function is not limited to a polynomial function. While preferably a polynomial function, this function may be a polynomial, exponential or trigonometric function or combination thereof.
[0112] In some embodiments, processor 302 calculates the expected rate of change using the polynomial function by applying the polynomial function to a mean value of the biomedical factor calculated by an average of the two or more predictions. Specifically, the mean value may be calculated using a difference of the two or more predictions divided by the number of two of more predictions. The polynomial function may be a rate of change verses mean value function and hence, may be a function of the mean value that, when applied to the mean value, results in the rate change corresponding to that mean value. [0113] In some embodiments, processor 302 determines the polynomial function by calculating a rate of change and a mean value for each pair of medical images in the patient study. Moreover, processor 302 may determine the polynomial function by calculating a rate of change and a mean value for each pair of medical images in each patient study that form the training data. Processor 302 may determine the polynomial function by determining the best representation of the training data. For example, processor 302 may use regression (e.g., curve of best fit) or other techniques, such as applying a machine learning model.
Further penalty values
[0114] In some embodiments, the loss value further comprises a change penalty value reflecting a difference between the two or more predictions, where the change penalty value penalises an unexpected temporal trend of the biomedical factor related to the expected rate of change. For example, the expected temporal trend of the longitudinal data may be a positive rate of change over time. As such, the change penalty value would penalise instances of a negative rate of change occurring. In some examples, the change penalty value may comprise a clamp function to only keep positive or negative changes.
[0115] In some embodiments, the loss value further comprises a correction penalty value reflecting a difference between the two or more predictions and two or more corresponding ground truth values. The correction penalty value may penalise the two or more predictions from deviating from the two or more corresponding ground truth values. For example, the correction penalty value may comprise two penalty values reflecting the slope and intercept between the two or more predicted values and the ground truth measurements. As such, the correction penalty value may be based on a difference between the rate of change of the two or more predictions and the rate of changes of the two or more corresponding ground truth values. In other examples, the slope and intercept may be based on a regression line between the ground truth measurements and the two or more predictions.
[0116] In some examples, it may be expected that the slope between output of the predicted values and the ground truth measurements should not deviate too far from 1. In some examples, it may be expected that the intercept between output of the predicted values and the ground truth measurements should not deviate too far from 0. As such, over-correction may be penalised during the training of the machine learning model using the two correction penalty values that reflect this expectation. This may ensure that no bias is introduced during the training process. In some examples, the correction penalty value may be a sum of the differences between each output and the over-correction factor. In further examples, the correction penalty value may be a sum of the absolute value of these differences.
[0117] For example, a regression line may be calculated by determining a slope and intercept based on each prediction and the corresponding ground truth value. More specifically, a plot may be made with data points corresponding to the prediction and the corresponding ground truth value, then a linear regression of the plot may be determined. In essence, a linear function may be calculated where the input to the linear function corresponds to a prediction and the output (i.e., value of the function evaluated at an input) corresponds to a ground truth value. Hence, the slope may be given as:
Figure imgf000031_0001
[0118] As it is desired that the prediction similar to the ground truth, it may be logical that the slope of this linear function should be around 1, while the intercept of the linear function should be around 0. Hence, the correction penalty may be based on a difference between the slope and 1, as well as the intercept and 0, such that this difference is minimised during training. For example, the correction penalty value may be given by: correction penalty value = abs(sZope — 1) + abs(intercept — 0)
[0119] However, it is noted that other variations of the correction penalty value may be possible.
[0120] In some embodiments, the loss value further comprises a correction penalty value reflecting a difference between the two or more outputs and an over-correction factor, where the correction penalty value penalises the two or more outputs from deviating from the overcorrection factor. In some examples, it may be known that the output of the machine learning model should not deviate too far from an over-correction factor. As such, over-correction can be penalised during the training of the machine learning model using the correction penalty value. In some examples, the correction penalty value may be a sum of the differences between each output and the over-correction factor. In further examples, the correction penalty value may be a sum of the absolute value of these differences.
[0121] In some embodiments, the loss value further comprises an additional penalty value calculated by one or more of: (i) a difference between each of the two or more outputs of the machine learning model and a respective ground truth measurement, and (ii) a mean square error between each of the two or more outputs and the respective ground truth measurement. As such, the loss value may comprise additional (standard) loss terms. In essence, the loss value may comprise any number of additional loss terms.
[0122] In some embodiments, processor 302 calculates the loss value by a weighted sum. For example, the loss value may comprise the penalty value, the change penalty value, the correction penalty value, and the additional penalty value. As such, processor 302 calculates the loss value by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value. The weights of the weighted sum may be considered “learning rates”, which are predetermined before training occurs, such as by user input. In some examples, one or more of the weights may be 1.
Aiding in diagnosis
[0123] It is noted that the biomedical factor can then be used to aid in determining a diagnosis for the patient. For example, a clinician may apply the trained machine learning model to an image of a new patient. As such, method 400 may further comprised applying the trained machine learning model to a medical image of a patient to aid in diagnosis of the patient based on the output of the trained machine learning model.
[0124] In one example, the trained machine learning model may calculate a measure, e.g., amount of amyloid plaques on the brain. The method may then apply one or more thresholds on that calculated measure to determine that the patient has Alzheimer’s pathology or that there is a low/medium/high risk of developing Alzheimer in the future. This use of the trained machine learning model may equally be possible for other medical indications, including progressive diseases, or general phenotypes.
[0125] Diagnosis and the study of Alzheimer’s disease is often performed using PET quantification of a patient’s brain PET image. The PET quantification may provide quantification information such as amount of amyloid plaques in the brain, for example. PET quantification may also involve processor 302 creating a three-dimensional rendering of the patient’s brain from the PET image. Hence, in some embodiments, each of the medical images is a PET image and method 400 comprises performing PET quantification by applying the trained machine learning model to a PET image of a patient.
[0126] As such, to aid in the diagnosis of Alzheimer’s disease, the biomedical factor would be a medical image quantification factor and more specifically, the medical image quantification factor would be a Centiloid or standardised uptake value ratio scaling factor. Moreover, to aid in the diagnosis of Alzheimer’s disease, method 400 would further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of Alzheimer’s disease based on the Centiloid or standardised uptake value ratio scaling factor that is an output of the trained machine learning model.
Example embodiment
[0127] An example embodiment of the disclosed method will now be presented for explanatory purposes but is not intended to limit the disclosed system and method to this embodiment. This example embodiment is also illustrated with experiment results of training a machine learning model using the disclosed method.
[0128] In this example embodiment, the biomedical factor is a medical image quantification factor and more specifically, the biomedical factor is a Centiloid value. The output of the machine learning model being trained is a scaling factor for the SUVR value, before it is transformed into Centiloid value, and the predictions are scaled Centiloid values. As such, the medical images used during training and inference will be PET images, as the SUVR and Centiloid scale relate to PET images. [0129] Centiloid values are used to standardise amyloid PET images for PET quantification by normalising the PET SUVR across different amyloid PET tracers to a standard and common unit system, with each PET tracer having its own linear transform to normalise SUVR into Centiloid. However, due to noise in the PET images and the use of different PET tracers/scanners which can have different levels of noise and bias, the Centiloid values are often noisy. Therefore, PET quantification that relies on PET images in the Centiloid scale may be less accurate as a result of the noise in the PET image and the variability introduced by the use of different tracers and scanners. As such, machine learning techniques can be used to determine a correction to the calculated SUVR value, before they are transformed into Centiloid value, in an effort to reduce the effect of noise in the PET images and increase the accuracy of the PET quantification.
[0130] In this example embodiment, the machine learning model is trained on two medical images at a time (i.e., two images from the same patient study for each training iteration). As such, processor 302 calculates two predictions of the Centiloid value (i.e., scaled Centiloid value) by applying the machine learning model to each of the two medical images, individually, calculating a respective output and calculating a respective prediction using the respective output.
[0131] In this embodiment, processor 302 calculates the predictions by first calculating the SUVR value for each of the two medical images. Processor 302 may calculate the SUVR value by calculating a ratio between a reference region of the respective PET image and a “hot spot” region of the respective PET scan, such as an amyloid, before transforming them into Centiloid. Processor 302 may also use the tracer and/or scanner information as part of the prediction, so that there are different corrections for different tracers, which allows to better control for their different levels of noise. Processor 302 then calculates the predictions by a product of the calculated SUVR value and the output of the machine learning model (i.e., the scaling factor), before transforming the scaled SUVR into Centiloid (i.e., the scaled Centiloid value).
[0132] Processor 302 then calculates a rate of change between the two predictions by calculating a difference between the two predictions (i.e., the scaled Centiloid value) divided by a difference between the timestamps of the two medical images, given by: prediction2 ~ prediction rate of change = — - - timestamp2 — timestamp1
[0133] Processor 302 then calculates the loss value. It is noted that in some embodiments, processor 302 calculates the loss value by applying a loss function. For example, in the example embodiment, processor 302 applies a loss function to the calculated Centiloid values, the outputs of the machine learning model and the timestamps of the two medical images. In other words, the calculated Centiloid values, the outputs of the machine learning model and the timestamps of the two medical images are inputs into the loss function. As such, processor 302 may calculate the rate of change by applying the loss function to the inputs.
[0134] In this example embodiment, the loss value comprises three penalty values: (i) the penalty value indicative of the rate of change, (ii) the change penalty value penalising an unexpected temporal trend of the Centiloid value, and (iii) the correction penalty value penalising the two or more outputs from deviating from the over-correction factor. In this embodiment, processor 302 calculates the loss value using a weighted sum of the three penalty values. For the experimental results, the loss value was calculated using: loss value = change penality value + a * correction penality value + (3
* penality value where a and P are pre-determined parameters or learning rates.
Penalty value indicative of the rate of change
[0135] To calculate this penalty value, processor 302 calculates a rate of change and a mean value for each pair of medical images for each patient study. Processor 302 then determines a polynomial function by applying a regression algorithm to the calculated rates of change. The calculated rates of change and the determined polynomial function for this example embodiment are shown in Fig. 2. In essence, it is expected that the rate of change does not deviate from the polynomial. As such, this penalty value indicative of the rate of change penalises large distances away from the polynomial. [0136] Processor 302 calculates the expected rate of change using the polynomial function by applying the polynomial function to a mean value of the Centiloid value calculated by an average of the two or more predictions. More specifically, the mean value is given by:
Figure imgf000036_0001
[0137] Processor 302 then calculates this penalty value by calculating a difference between the calculated rate of change and the expected rate of change. More specifically, processor 302 calculates the absolute value of the difference between the calculated rate of change and the expected rate of change. As such, the penalty value is given by: penalty value = abs(calculated change — f(mean value)) where (•) is the polynomial function.
Change penalty value
[0138] The change penalty value penalises an unexpected temporal trend of the Centiloid value related to the expected rate of change. In the example embodiment, it is expected that Centiloid value will increase over time and hence, the rate of change of the Centiloid value is expected to be positive (more than zero). As such, the change penalty value is formulated in such a way to penalise decreasing Centiloid values and negative rates of change. In the example embodiment, the change penalty value is based on a “baseline/follow up” concept, where patients will have a “baseline” medical image scan performed which will be compared to a “follow up” medical image scan in the future.
[0139] The change penalty value is given by: change penalty value = clamp(prediction2 — prediction^ 0) where “clamp” is the clamp function, which restricts a given value between an upper and lower bound. More specifically, the clamp function is defined by: clamp(x) = max(a, min(x, b)) G [a, b]
[0140] The clamp function restricts the difference in the predictions to only be positive, thereby enforcing the extended trend of an always positive rate of change during the training.
Correction penalty value
[0141] The correction penalty value penalises the two or more outputs of the machine learning model from deviating from an over-correction factor. In the example embodiment, the over-correction factor is 1, as the output of the machine learning model is a scaling factor and hence, the output should be ideally 1 in the absence of noise in the PET image.
[0142] The correction penalty value is given by: correction penalty value = abs^outpuG — 1) + abs(oi4tpitt2 — 1)
Results
[0143] The results of training the machine learning model are shown in Fig. 6. To produce these results, the machine learning model was trained using 5,000 PET images from 1,500 patients. The machine learning model used to generate these results was a CNN of the architecture shown in Fig. 5, noting that the input to the machine learning model is a PET image rather than an MR image. Fig. 6 shows the progression of the calculated rate of change of the scaled Centiloid values as the training of the machine learning model increases. In particular, it can be shown that the calculated rates of change move towards the polynomial function as the training continues. Further, the calculated rates of change become more positive as per the expected rate of change of the Centiloid values.
Generating masks for machine learning explainability
[0144] Machine learning models (particularly deep machine learning models) are often known for being black boxes, meaning that it is unknown how they calculate a prediction from an input. For example, while the models may be able to detect the presence of a disease from image data, it may not be clear how the model makes its prediction nor may it be clear what features the model deems important when making its prediction. The black box nature of machine learning model means that interpretability can be challenging. The understanding of decisions made by black-box algorithms is a challenge and assessing their fairness and unbiasedness is an important step to deployment in many industries.
[0145] As previously discussed, the trained machine learning model disclosed herein may output a prediction of a biomedical factor, which may be a standardised uptake value ratio (SUVR) scaling factor (which may be simply referred to as a SUVR). Recall that the SUVR may be defined as the ratio of a region of the PET image containing specific binding (a target region where the PET tracer will bind to the target), and a region of the PET image containing non-specific binding (as referred to as a reference region). As such, a target mask may be generated which highlights the target regions in a PET image. Similarly, a reference mask may be generated which highlights the non-specific binding regions in a PET image. In the case of Centiloids in a PET image, the reference mask may be referred to as a Centiloid reference mask or the like. Hence, generating masks that correspond to the corrections provided by the trained machine learning model may provide insight into which regions of the PET image that the trained machine learning uses in its determination of corrected SUVR.
[0146] It is thought that the corrections provided by the trained machine learning model disclosed herein may reflect changes in either the reference or target mask, or both. This idea may be leveraged to generate masks with SUVRs corresponding to the corrections provided by the trained machine learning model. In other words, new masks may be optimised, for each tracer, such that the correlation with the SUVR values provided by the trained machine learning model is optimised (e.g., maximised). This may provide insight into the features deemed important to the trained machine learning model, which may not be obvious in the original PET image, for example.
[0147] As such, in some embodiments, processor 302 may generate one or more masks that correspond to corrections determined by the trained machine learning model. It is noted that a “mask” in the present context may refer to an image, such as a heat map, saliency map or the like, that highlights relevant regions for machine learning model decision making. Processor 302 may generate the one or more masks by adapting one or more medical images (that may be input into the trained machine learning model) based on a corresponding output of the trained machine learning model. For example, processor 302 may apply an algorithm, mathematical operation or a further machine learning model to the one or more medical images based on the corresponding output to generate the one or more masks.
[0148] Obtaining these new masks may be formulated as an optimisation problem where, for each tracer, a new reference and target mask may be optimised across the entire dataset so that the resulting SUVR maximises the Pearson correlation coefficient with the corrected SUVRs provided by the trained machine learning model. The Pearson correlation coefficient is a correlation coefficient that measures linear correlation between two sets of data.
[0149] This optimisation problem may be solved using a gradient descent approach or another type of optimisation technique. The lost function may involve the following two losses:
1. Penalises low Pearson correlation coefficient. Using the updated masks, new SUVRs can be computed across the entire dataset. The correlation between these SUVRs and the Corrected SUVR obtained by the trained machine learning model are used to compute the Pearson coefficient R2. The resulting loss is defined as:
Lp = 1 - R2
2. Promotes binary masks. The two masks being optimised need to have continuous values to ensure that they can be derived throughout the optimisation process. To force the masks towards binarity, for each mask M, values not being either 1 or 0 are penalised as follow:
0.5 - |M - 0.5|
Figure imgf000039_0001
[0150] Experiment and results [0151] An experiment was conducted to generate reference and target masks for machine learning model explainability. The experimental details are now provided. However, it is noted that alternatives and alterations are possible in the generation of these masks.
[0152] The reference and target masks were initialised using the original Centiloid reference and target mask. The masks were then smoothen using a Gaussian kernel, specifically a 4mm FWHM Gaussian kernel. At each iteration, the masks were updated using the gradient information, and smoothed again using the same Gaussian kernel to reduce sparseness. The masks were also be mirrored to generate symmetric masks. In this experiment, the optimisation leveraged pytorch’s autograd engine for automatic computation of the gradients.
[0153] The optimisation was stopped once the Pearson loss stopped to improve anymore for 10 epochs or if the maximum number of epochs (2000) was reached. Once the final masks were obtained, the new SUVRs derived from these masks were transformed into the Corrected SUVR using a linear transform, before being transformed into Centiloids using the standard transforms.
[0154] Fig. 7 shows the result of this optimisation process. In particular, Fig. 7 shows the progression of the SUVR derived from the optimised masks as the optimisation processes continues (i.e., after a number of iterations). It can be seen in Fig. 7 that, as the optimisation process continues, the ratio of the Corrected SUVR provided by the trained machine learning model (labelled as DeepSUVR SUVR on the x-axis) and the SUVR determined from the optimised masks approaches 1 (as indicated by the datapoints approaching the diagonal line). This indicates that through the optimisation process described above, the reference and target masks are indeed optimised such that the SUVR determined from these masks correlates to the correspond SUVR provided by the trained machine learning model.
[0155] Fig. 8a shows the initialised target masks, while Fig. 8b shows the optimised target masks. Similarly, Fig. 9a shows the initialised reference masks, while Fig. 9b shows the optimised reference masks. It is noted that for both the reference and target masks, the optimised masks still include salient features that can be seen in the initialised masks. However, the optimisation process provides more salient features and better definition of these features. These results provide a better indication of the features considered important by the trained machine learning model disclosed herein in making its prediction. As such, these results aid in the explainability of the trained machine learning model. These optimised masks can also potentially replace the machine learning model for directly generating Corrected SUVRs.
[0156] It will be appreciated by persons skilled in the art that numerous variations and/or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

CLAIMS:
1. A method for training a machine learning model on longitudinal medical scan images of a patient study, the method comprises: calculating two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output; and calculating a prediction using the output of the machine learning model for the medical image; calculating a rate of change between the two or more predictions for the two or more medical images; and creating a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
2. The method of claim 1, wherein the method further comprises determining a polynomial function of rate of change of the biomedical factor and calculating the expected rate of change using the polynomial function.
3. The method of claim 2, wherein calculating the expected rate of change using the polynomial function comprises applying the polynomial function to a mean value of the biomedical factor calculated by an average of the two or more predictions.
4. The method of claim 2 or 3, wherein determining the polynomial function comprises calculating a rate of change and a mean value for each pair of medical images in the patient study.
5. The method of any one of the preceding claims, wherein the loss value further comprises a change penalty value reflecting a difference between the two or more predictions, the change penalty value penalising an unexpected temporal trend of the biomedical factor related to the expected rate of change.
6. The method of any one of the preceding claims, wherein the loss value further comprises a correction penalty value reflecting a difference between the two or more predictions and two or more corresponding ground truth values, the correction penalty value penalising the two or more predictions from deviating from the two or more corresponding ground truth values.
7. The method of any one of the preceding claims, wherein the loss value further comprises an additional penalty value calculated by one or more of: a difference between each of the two or more outputs of the machine learning model and a respective ground truth measurement; and a mean square error between each of the two or more outputs and the respective ground truth measurement.
8. The method of claim 7, wherein the loss value is calculated by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
9. The method of any one of the preceding claims, wherein each of the medical images of the patient study has a timestamp and at least two of the two or more medical images have non-adjacent timestamps.
10. The method of any one of the preceding claims, wherein the method further comprises calculating a further two or more predictions of the biomedical factor by performing for each of a further two or more medical images of the patient study: applying the trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating a further rate of change between the further two or more predictions; and creating a further trained machine learning model by updating weights of the trained machine learning model to minimise a further loss value based of the further rate of change.
11. The method of any one of the preceding claims, wherein the method further comprises calculating another two or more predictions of the biomedical factor by performing for each of two or more medical images of another patient study different to the patient study: applying the trained machine learning model or further trained machine learning model to the medical image to calculate an output; and calculating a prediction using the output; calculating another rate of change between the other two or more predictions; and creating another trained machine learning model by updating weights of the trained machine learning model or further trained machine learning model to minimise another loss value based on the other rate of change.
12. The method of any one of the preceding claims, wherein the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of the patient based on the output of the trained machine learning model.
13. The method of any one of the preceding claims, wherein each of the medical images is one of a: magnetic resonance image; positron emission tomography (PET) image; or computed tomography (CT) image.
14. The method of any one of the preceding claims, wherein the biomedical factor is one of: medical image quantification factor; a biological structure value; or a biomarker.
15. The method of claim 14, wherein the medical image quantification factor is a Centiloid or standardised uptake value ratio scaling factor and the biological structure value is volume of a biological structure visible in the medical image.
16. The method of any one of claims 11 to 15, wherein each of the medical images is a PET image and the method further comprises performing PET quantification by applying the trained machine learning model to a PET image of a patient.
17. The method of claim 15 or 16, wherein the medical image quantification factor is a Centiloid or standardised uptake value ratio scaling factor; and the method further comprises applying the trained machine learning model to a medical image of a patient to aid in diagnosis of Alzheimer’s disease based on the Centiloid or standardised uptake value ratio scaling factor that is an output of the trained machine learning model.
18. The method of any one of the preceding claims, further comprising generating one or more masks that correspond to corrections determined by the trained machine learning model.
19. Software that, when installed on a computer and executed by the computer, causes the computer to perform the method of any one of the preceding claims.
20. A system for training a machine learning model on longitudinal medical scan images of a patient study, the system comprises: a processor configured to: calculate two or more predictions of a biomedical factor by performing for each of two or more of the medical images: applying the machine learning model to the medical image to calculate an output; and calculating a prediction using the output of the machine learning model for the medical image; calculate a rate of change between the two or more predictions for the two or more medical images; and create a trained machine learning model by updating weights of the machine learning model to minimise a loss value, wherein the loss value comprises a penalty value that reflects a difference between the rate of change and an expected rate of change.
PCT/AU2024/051378 2023-12-22 2024-12-19 Training deep learning models Pending WO2025129253A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
AU2023904194A AU2023904194A0 (en) 2023-12-22 Training deep learning models
AU2023904194 2023-12-22
AU2024903443A AU2024903443A0 (en) 2024-10-23 Training deep learning models
AU2024903443 2024-10-23

Publications (1)

Publication Number Publication Date
WO2025129253A1 true WO2025129253A1 (en) 2025-06-26

Family

ID=96135966

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/AU2024/051378 Pending WO2025129253A1 (en) 2023-12-22 2024-12-19 Training deep learning models

Country Status (1)

Country Link
WO (1) WO2025129253A1 (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210090694A1 (en) * 2019-09-19 2021-03-25 Tempus Labs Data based cancer research and treatment systems and methods
WO2023102229A1 (en) * 2021-12-03 2023-06-08 University Of Pittsburgh – Of The Commonwealth System Of Higher Education Aneurysm modeling and risk prediction
WO2023164051A1 (en) * 2022-02-24 2023-08-31 Artera Inc. Systems and methods for determining cancer therapy via deep learning
US20230277118A1 (en) * 2022-02-04 2023-09-07 Spark Neuro Inc. Methods and systems for detecting and assessing cognitive impairment

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210090694A1 (en) * 2019-09-19 2021-03-25 Tempus Labs Data based cancer research and treatment systems and methods
WO2023102229A1 (en) * 2021-12-03 2023-06-08 University Of Pittsburgh – Of The Commonwealth System Of Higher Education Aneurysm modeling and risk prediction
US20230277118A1 (en) * 2022-02-04 2023-09-07 Spark Neuro Inc. Methods and systems for detecting and assessing cognitive impairment
WO2023164051A1 (en) * 2022-02-24 2023-08-31 Artera Inc. Systems and methods for determining cancer therapy via deep learning

Non-Patent Citations (8)

* Cited by examiner, † Cited by third party
Title
GIORGIO JOSEPH, JAGUST WILLIAM J., BAKER SUZANNE, LANDAU SUSAN M., TINO PETER, KOURTZI ZOE: "A robust and interpretable machine learning approach using multimodal biological data to predict future pathological tau accumulation", NATURE COMMUNICATIONS, NATURE PUBLISHING GROUP, UK, vol. 13, no. 1, UK, XP093332169, ISSN: 2041-1723, DOI: 10.1038/s41467-022-28795-7 *
HL DURGA SUPRIYA; THOMAS SWETHA MARY; S SOWMYA KAMATH: "A Multimodal Approach Integrating Convolutional and Recurrent Neural Networks for Alzheimer’s Disease Temporal Progression Prediction", 2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS (CVPRW), IEEE, 17 June 2024 (2024-06-17), pages 5207 - 5215, XP034708882, DOI: 10.1109/CVPRW63382.2024.00529 *
LEE, SEONG-WHAN ; LI, STAN Z: "SAT 2015 18th International Conference, Austin, TX, USA, September 24-27, 2015", vol. 9351 Chap.28, 18 November 2015, SPRINGER , Berlin, Heidelberg , ISBN: 3540745491, article RONNEBERGER OLAF; FISCHER PHILIPP; BROX THOMAS: "U-Net: Convolutional Networks for Biomedical Image Segmentation", pages: 234 - 241, XP047331005, 032548, DOI: 10.1007/978-3-319-24574-4_28 *
LUCKETT EMMA S., ABAKKOUY YASMINA, REINARTZ MARISKA, ADAMCZUK KATARZYNA, SCHAEVERBEKE JOLIEN, VERSTOCKT SARE, DE MEYER STEFFI, VAN: "Association of Alzheimer’s disease polygenic risk scores with amyloid accumulation in cognitively intact older adults", ALZHEIMER'S RESEARCH & THERAPY, BIOMED CENTRAL LTD, LONDON, UK, vol. 14, no. 1, London, UK , XP093332167, ISSN: 1758-9193, DOI: 10.1186/s13195-022-01079-4 *
MARKR. BATTLE;LOVENACHEDUMBARUM PILLAY;VALJ. LOWE;DAVID KNOPMAN;BRADLEY KEMP;CHRISTOPHERC. ROWE;VINCENT DORé;VICTORL. VILLEMA: "Centiloid scaling for quantification of brain amyloid with [F]flutemetamol using multiple processing methods", EJNMMI RESEARCH, BIOMED CENTRAL LTD, LONDON, UK, vol. 8, no. 1, 5 December 2018 (2018-12-05), London, UK , pages 1 - 11, XP021263425, DOI: 10.1186/s13550-018-0456-7 *
SCHALKAMP ANN-KATHRIN, HARRISON NEIL A., PEALL KATHRYN J., SANDOR CYNTHIA: "Digital outcome measures from smartwatch data relate to non-motor features of Parkinson’s disease", NPJ PARKINSON'S DISEASE, NATURE PUBLISHING GROUP, LONDON, vol. 10, no. 1, London , XP093332162, ISSN: 2373-8057, DOI: 10.1038/s41531-024-00719-w *
SHISHEGAR ROSITA, COX TIMOTHY, ROLLS DAVID, BOURGEAT PIERRICK, DORé VINCENT, LAMB FIONA, ROBERTSON JOANNE, LAWS SIMON M., PORTER : "Using imputation to provide harmonized longitudinal measures of cognition across AIBL and ADNI", SCIENTIFIC REPORTS, NATURE PUBLISHING GROUP, US, vol. 11, no. 1, US , XP093332163, ISSN: 2045-2322, DOI: 10.1038/s41598-021-02827-6 *
WU, J. : "Mining Associations between MRI Morphometry Measurements and Beta- Amyloid/tau Burden", ARIZONA STATE UNIVERSITY, 1 August 2022 (2022-08-01), XP093332164 *

Similar Documents

Publication Publication Date Title
JP7797554B2 (en) System and method for segmentation of anatomical structures in image analysis
Baheti et al. The brain tumor sequence registration (BraTS-Reg) challenge: Establishing correspondence between pre-operative and follow-up MRI scans of diffuse glioma patients
JP7757283B2 (en) Automated tumor identification and segmentation in medical images
US11696701B2 (en) Systems and methods for estimating histological features from medical images using a trained model
US11250601B2 (en) Learning-assisted multi-modality dielectric imaging
US8774481B2 (en) Atlas-assisted synthetic computed tomography using deformable image registration
US9251585B2 (en) Coregistration and analysis of multi-modal images obtained in different geometries
CN113196340B (en) Methods, non-transitory computer-readable media, and systems for imaging
EP4018371B1 (en) Systems and methods for accurate and rapid positron emission tomography using deep learning
JP7459243B2 (en) Image reconstruction by modeling image formation as one or more neural networks
Park et al. Deformable registration of CT and cone-beam CT with local intensity matching
US20130267841A1 (en) Extracting Application Dependent Extra Modal Information from an Anatomical Imaging Modality for use in Reconstruction of Functional Imaging Data
US11055559B1 (en) System, method and apparatus for assisting a determination of medical images
CN110546685A (en) Image segmentation and segmentation prediction
Carles et al. Significance of the impact of motion compensation on the variability of PET image features
CN110610527B (en) Method, device, equipment, system and computer storage medium for calculating SUV
WO2025129253A1 (en) Training deep learning models
US20250078257A1 (en) Calibration of activity concentration uptake
Gigengack et al. Motion correction in thoracic positron emission tomography
Goldsmith et al. Nonlinear tube-fitting for the analysis of anatomical and functional structures
WO2024077348A1 (en) Saliency maps for classification of images
Hahn et al. Unbiased rigid registration using transfer functions
Salman et al. Simulating the Tumor Mass Changes in PET and PET/CT Segmented Images Using Unsupervised Artificial Neural Network, HSOFM.
Jin A Quality Assurance Pipeline for Deep Learning Segmentation Models for Radiotherapy Applications
Shao et al. Neural Network Approximation Aided Kinetic Parameter Inference of Compartment Models in Dynamic PET

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24905186

Country of ref document: EP

Kind code of ref document: A1