WO2025106929A1 - Body structure segmentation using supervised and unsupervised learning - Google Patents

Body structure segmentation using supervised and unsupervised learning Download PDF

Info

Publication number
WO2025106929A1
WO2025106929A1 PCT/US2024/056293 US2024056293W WO2025106929A1 WO 2025106929 A1 WO2025106929 A1 WO 2025106929A1 US 2024056293 W US2024056293 W US 2024056293W WO 2025106929 A1 WO2025106929 A1 WO 2025106929A1
Authority
WO
WIPO (PCT)
Prior art keywords
processors
feature
resolution
dimensional image
pixel
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2024/056293
Other languages
French (fr)
Inventor
Chung Chieh Kuo
Andre Luis DE CASTRO ABREU
Giovanni CACCIAMANI
Vinay Anant DUDDALWAR
Inderbir Singh Gill
Masatomo KANEKO
Vasileios MAGOULIANITIS
Chrysostomos Loizos NIKIAS
Jintang XUE
Yijing Yang
Jiaxin Yang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Southern California USC
Original Assignee
University of Southern California USC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Southern California USC filed Critical University of Southern California USC
Publication of WO2025106929A1 publication Critical patent/WO2025106929A1/en
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/74Image or video pattern matching; Proximity measures in feature spaces
    • G06V10/75Organisation of the matching processes, e.g. simultaneous or sequential comparisons of image or video features; Coarse-fine approaches, e.g. multi-scale approaches; using context analysis; Selection of dictionaries
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10081Computed x-ray tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10088Magnetic resonance imaging [MRI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20016Hierarchical, coarse-to-fine, multiscale or multiresolution image processing; Pyramid transform
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30081Prostate
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/03Recognition of patterns in medical or anatomical images

Definitions

  • Imaging systems play a crucial role in modem medicine where they are essential for diagnosing and evaluating diseases.
  • These advanced technologies include modalities such as X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound.
  • CT computed tomography
  • MRI magnetic resonance imaging
  • ultrasound ultrasound
  • Prostate cancer for example, is reported as the second most frequent cancer among men in 2020, with an estimated almost 1.4 million new cases and 375,000 deaths worldwide.
  • TRUSGB transrectal ultrasound-guided biopsy
  • mpMRI multiparametric magnetic resonance imaging
  • Segmentation of a body structure is an important step in diagnosis and treatment planning. Segmentation allows for determining prostate boundaries for radiotherapy. Segmentation can also allow the calculation of volume and other key metrics necessary to track disease progression.
  • Radiologists also use imaging systems to detect tumors or the presence of other diseases.
  • MRI imaging systems can provide detailed images of soft tissues, which is particularly valuable for identifying cancerous growths.
  • CT scans offer high-resolution cross- sectional images that help in evaluating complex conditions, such as trauma injuries or vascular diseases.
  • radiologists can detect not only the presence of disease but also its size, shape, and location, which are vital for determining the appropriate course of treatment.
  • Some embodiments of the present disclosure relate to a method.
  • the method includes obtaining, by one or more processors, a slice of a three-dimensional image.
  • the three- dimensional image depicts a structure of an individual.
  • the slice includes a plurality of pixels each corresponding to a different location of the three-dimensional image.
  • the method also includes executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions.
  • Each feature map of the plurality of feature maps includes a plurality of pixel embeddings each corresponding to a different location of the slice.
  • the method also includes executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask.
  • the first mask includes values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual.
  • the method also includes concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution is higher than the first resolution.
  • the method also includes executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual.
  • the method also includes generating, by the one or more processors, a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
  • the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution. In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation. In some embodiments, the method also includes training, by the one or more processors, the feature extraction model using an unsupervised learning technique. [0009] In some embodiments, executing the feature extraction model includes executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings.
  • executing the feature extraction model includes performing, by the one or more processors, principal components analysis on sets of adjacent pixels of a plurality of locations of the three-dimensional image. In some embodiments, executing the feature extraction model includes executing, by the one or more processors, a transformation on a set of adjacent pixels encompassing a pixel of the plurality of pixels to generate the pixel embedding associated with the pixel.
  • the feature extraction model includes pooling, by the one or more processors, adjacent pixel embeddings to generate a lower resolution feature map.
  • the method also includes classifying, by the one or more processors, each concatenated pixel embedding based on whether the probability for the concatenated pixel embedding exceeds a threshold.
  • obtaining the slice of a three- dimensional image includes receiving, by the one or more processors, the three-dimensional image of the individual, obtaining the slice of a three-dimensional image includes segmenting, by the one or more processors, the three-dimensional image into a plurality of slices corresponding to different locations of the three-dimensional image.
  • obtaining the slice of the three-dimensional image includes obtaining the three-dimensional image depicting a prostate of the individual.
  • Executing the second machine learning model to classify each concatenated pixel embedding includes executing, by the one or more processors, the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual.
  • Generating the second mask includes generating, by the one or more processors, the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
  • Some embodiments of the present disclosure relate to a system.
  • the system includes one or more processors configured by computer-readable instructions to perform operations.
  • the operations include obtaining a slice of a three-dimensional image, the three-dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location of the three-dimensional image.
  • the operations include executing a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice.
  • the operations also include executing a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual.
  • the method also includes concatenating each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution.
  • the method also includes executes a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual.
  • the method also includes generating a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
  • the one or more processors are further configured to scale the first mask from the first resolution to the second resolution. In some embodiments, the one or more processors are configured to scale the first mask from the first resolution to the second resolution using trilinear interpolation. In some embodiments, the one or more processors are further configured to train the feature extraction model using an unsupervised learning technique. In some embodiments, the one or more processors are further configured to execute the feature extraction model to transform a set of adjacent pixels including a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings.
  • the one or more processors are configured to obtain the slice of the three-dimensional image by obtaining the three-dimensional image depicting a prostate of the individual.
  • the one or more processors are configured to execute the second machine learning model to classify each concatenated pixel embedding by executing the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual.
  • the one or more processors are configured to generate the second mask by generating the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
  • Some embodiments of the present disclosure relate to a method.
  • the method includes obtaining, by one or more processors, a slice of a three-dimensional image, the three- dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location of the three-dimensional image.
  • the method also includes executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice.
  • the method also includes executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual.
  • the method also includes concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution.
  • the method also includes executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual.
  • the method also includes presenting, by the one or more processors, a visual representation of the classifications of the concatenated pixel embeddings.
  • the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution. In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation.
  • Some embodiments of the present disclosure relate to a method.
  • the method includes obtaining, by one or more processors, a three-dimensional image.
  • the three-dimensional image depicting a structure of an individual and include a plurality of voxels.
  • the method includes performing for each voxel of a plurality of voxels; identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual.
  • the method also includes determining, by the one or more processors, one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold.
  • the method includes for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
  • the method also includes training, by the one or more processors, the feature extraction model using unsupervised learning.
  • executing the feature extraction model includes executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings.
  • executing the feature extraction model includes performing, by the one or more processors, spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
  • executing the feature extraction model further includes combining, by the one or more processors, the features of the plurality of feature maps of different resolutions.
  • determining the one or more regions of interest of the three-dimensional image includes identifying, by the one or more processors, clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
  • the method includes presenting, by the one or more processors, a first visual representation of the first set of clusters in a first color and a second visual representation of the second set of clusters in a second color.
  • the method includes concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps.
  • Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps as input.
  • the method also includes obtaining, by the one or more processors, anomaly features extracted from the three-dimensional image and concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps and the anomaly features
  • Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps and the anomaly features as input.
  • the three-dimensional image is of a plurality of three- dimensional images of different imaging technologies
  • determining the one or more regions of interest of the three-dimensional image includes determining, by the one or more processors.
  • obtaining the three-dimensional image includes obtaining, by the one or more processors, the three-dimensional image depicting a prostate of the individual.
  • Executing the first machine learning model includes executing, by the one or more processors, the first machine learning model to generate the value indicating the probability that the voxel depicts a cancerous portion of a prostate of the individual.
  • Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model to generate the classification indicating whether the region of interest is the cancerous portion of the prostate of the individual.
  • Some embodiments of the present disclosure relate to a system.
  • the system includes one or more processors configured by computer-readable instructions to perform operations.
  • the operations include obtaining a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels.
  • the method includes for each voxel of the plurality of voxels: identifying a set of adjacent voxels including the voxel of the plurality of voxels; executing a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and execute a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual.
  • the method also includes determining one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold.
  • the method also includes for each of the one or more regions of interest, executing a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
  • the one or more processors are further configured to train the feature extraction model using unsupervised learning. In some embodiments, the one or more processors are configured to execute the feature extraction model by performing spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
  • the one or more processors are configured to execute the feature extraction model further by combining the features of the plurality of feature maps of different resolutions. In some embodiments, the one or more processors are configured to determine the one or more regions of interest of the three-dimensional image by identifying clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
  • Some embodiments of the present disclosure relate to a method.
  • the method includes obtaining, by one or more processors, a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels.
  • the method also includes, for each voxel of the plurality of voxels: identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual.
  • the method also includes determining, by the one or more processors, one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold.
  • the method also includes for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and anomaly features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
  • the method also includes training, by the one or more processors, the feature extraction model using unsupervised learning.
  • executing the feature extraction model includes executing, by the one or more processors. The feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings.
  • FIG. l is a block diagram of an imaging and structure segmentation system showing the body structure segmentation system, according to some implementations.
  • FIG. 2 is flow of operations for segmenting a body structure, according to some implementations
  • FIG. 3 is a flow diagram illustrating the flow of data as it traverses the body structure segmentation system of FIG. 1, according to some embodiments;
  • FIG. 4 is flow of operations performed by a feature extractor of FIG. 3, according to some implementations;
  • FIG. 5 is an illustration of the operations of a feature extractor of FIG. 3, according to some implementations;
  • FIG. 6 is another block diagram of the imaging and structure segmentation system of FIG. 1 showing the body structure segmentation training system, according to some implementations;
  • FIG. 7 is an illustration of the operations to determine the feature extraction operators, according to some implementations.
  • FIG. 8 is an illustration of the operations to train a segmentation classifier, according to some implementations.
  • FIG. 9 is a block diagram of an imaging and identification system showing the diseased structure identification system, according to some implementations.
  • FIG. 10 is flow of operations for identifying a diseased portion of a body structure, according to some implementations.
  • FIG. 11 is a flow diagram illustrating the flow of data as it traverses the diseased structure identification system of FIG. 9, according to some implementations;
  • FIG. 12 is flow of operations performed by the feature extractor of FIG. 10, according to some implementations.
  • FIG. 13 is an illustration of the operations of a feature extractor of FIG. 10, according to some implementations.
  • FIG. 14 is another block diagram of an imaging and identification system of FIG. 9 showing the diseased structure identification training system, according to some implementations.
  • Deformable models employ flexible structures that adapt their shape to fit the contours of objects in images.
  • the models iteratively adjust their parameters to minimize an objective function that is designed to ultimately converge on the desired shape or boundary.
  • Deformable models often used in image segmentation and shape analysis are highly sensitive to initial conditions because their optimization relies on the initial settings to guide their deformation process. If the initial position is close to the desired contour, the model can effectively converge to the target shape. However, if the starting point is far from the desired boundary, the model may converge to a local minimum that does not accurately represent the true shape.
  • Deformable models may use multiple initial conditions (e.g., determined randomly) to increase the chance of finding a global optimum, leading to significant computational expense.
  • the image is modeled as a graph and formulates a minimum cut problem to optimally partition the foreground (e.g. the body structure of interest) from the background.
  • Graph cut optimization can have sensitivity to noise and may lead to oversegmentation or under-segmentation, where the model fails to accurately delineate the desired objects. Additionally, the computational cost of graph cut methods can become prohibitive for high-resolution images or large datasets. The algorithm's performance can also be limited in scenarios with occlusions or complex object shapes, leading to inaccurate segmentations. These limitations can lead to inaccurate results at high computational expense.
  • Convolutional neural networks and other deep neural networks have been used to perform image segmentation. Depending on the depth of the convolutional network and the number of kernels used in each layer, performing segmentation may still require a significant computational burden in the forward path (e.g., during inference). In addition, training the convolutional network to perform segmentation requires large sets of training data that has been labeled (e.g., the segmentation process was performed by another system or manually). Deep neural networks rely on long training times required to make hundreds of passes through the data set while adjusting the network weights based on backpropagated gradients. In addition, training may be required separately for each type of image to be segmented (e.g., each body structure).
  • the training data may be insufficient to provide accurate results or the parameters of the training process may require tuning.
  • deep neural networks may be viewed as a “black-box” and may not provide useful insight related to the reason for results. For example, if the results were inaccurate, it may be difficult to determine the reason for the error.
  • the present disclosure improves the technological field of image segmentation and more specifically segmentation of body structures by reducing the number of computations performed during the segmentation process and during the training process. In addition, the present disclosure may require significantly less training data to be collected.
  • a feature extraction method is performed on a slice of a three-dimensional image to generate a plurality of feature maps of different resolutions.
  • the features are used to indicate (e.g., to classify) structures of an individual that are depicted by the voxels of the image. For example, the features are used to indicate that the voxels are associated with a specific body structure and/or a region within the body structure.
  • Classification of a region of the slice is performed at a lower resolution and then used during classification at a higher resolution.
  • the hierarchical approach provides two improvements: (i) less training data may be required to train a classifier at a lower resolution, because this information is carried to the higher resolutions the higher resolution data will also require fewer training examples, and (ii) the hierarchy of resolutions that are provided as output allow a user to quickly identify discordance of results and take appropriate actions in real-time (e.g., rescanning the patient).
  • the feature extraction method is unsupervised leading to a further reduction in the amount of training data required and allowing for feature extraction using unlabeled training data.
  • the feature extraction technique provides for a direct computation of the weights from the training data without performing iterations (e.g., gradient descent, etc.) and may greatly reduce the segmentation computational requirements.
  • the feature extraction method performed relies on forming a sum of a product (e.g., a matrix multiplication) of a patch (e.g., a local region) of the slice of the three-dimensional image or of a feature map and a set of coefficients.
  • a product e.g., a matrix multiplication
  • a patch e.g., a local region
  • the feature extraction technique does not require the large number of kernels used by a convolutional neural network, further reducing the computational complexity of performing segmentation.
  • segmentation results may be available for the radiologist or technician in real-time allowing for rescans or adjustments to focus on a region of interest.
  • computational resource savings leads to more efficient processing and energy savings. Accuracy is improved over some segmentation methods that have more than 50 times the memory requirements and more than 150 times the computational requirements of the methods disclosed herein.
  • Segmenting a diseased portion of a body structure from healthy tissue is also an important task performed in the clinical setting.
  • the present disclosure provides systems and methods for identifying diseased portions of a body structure that have advantages (e.g., similar to those already discussed) over the traditional automated techniques and/or neural network techniques. Identification of diseased areas of a body structure (e.g., cancerous areas, kidney stones, etc.) can use similar feature extraction techniques that rely on less training data and have computational advantages in their calculation when compared to deformable models, graph cut optimization, and/or convolutional neural networks. An initial classification of voxels using the feature extraction technique may cause various regions of interest to be identified and flagged for further investigation.
  • regions may be identified for biopsy, used to automatically trigger additionally imagery to be performed, and/or to focus the attention of the radiologist.
  • visually-inspired features e.g., first order statistics, a gray level run length matrix, a gray level dependence matrix, etc.
  • FIG. 1 is a block diagram of an imaging and structure segmentation system 10 configured to obtain medical imagery of a body structure and segment (e.g., identify, highlight, etc.) the structures of interested according to some embodiments.
  • the imaging and structure segmentation system 10 may segment specific bone and/or joint structures, soft tissue (e.g., muscles, tendons, and ligaments), organs (e.g., the heart, lungs, kidneys, prostate, etc.), the nervous system, digestive tractor, or any other body structure for which segmenting a particular structure of interest from other tissue is useful.
  • the imaging and structure segmentation system 10 includes one or more client devices 20, one or more imaging systems 30, a body structure segmentation training system 150, and a body structure segmentation system 100 shown to be communicably coupled via a network 40.
  • the network 40 can include routers, switches, antennas, computers, and any other hardware required to communicate information between the components of the imaging and structure segmentation system 10 (e.g., from the imaging system 30 to the body structure segmentation system 100).
  • a portion of the network 40 can be wireless and/or a portion of the network 40 can be wired.
  • the network 40 can include one or more networks with routers to facilitate data transfer between the different networks.
  • the imaging and structure segmentation system 10 is shown to be distributed across several devices (e.g., networked computers). It is contemplated that in some embodiments, the subsystems, components, and/or functionality of the imaging and structure segmentation system 10 may be distributed on different devices.
  • the body structure segmentation system 100 may be configured within the imaging system 30 using the processors of the imaging system 30 to execute.
  • the body structure segmentation system 100 and the body structure segmentation training system 150 are configured within the same computer hardware (e.g., within a server local to the imaging system, within a remote server, or on a node within a cloud computing architecture).
  • the body structure segmentation system 100 may include a communications interface 102 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40.
  • the body structure segmentation system 100 may also include a processing circuit 104 having a processor 106 and memory 108.
  • the processor 106 may be configured to execute instructions contained on the memory 108.
  • the processor 106 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components.
  • the processor 106 may be configured to execute computer code and/or instructions stored in the memory 108 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.).
  • the processor 106 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof.
  • One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
  • the memory 108 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure.
  • the memory 108 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions.
  • the memory 108 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure.
  • the memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
  • the body structure segmentation system 100 is to acquire a three-dimensional image of a portion of the body, perform a sequence of feature extraction operations and pooling operations on one or more slices of the three-dimensional image to obtain feature maps of various resolutions and varying depths (e.g., number of features).
  • information of the feature maps is combined during a hierarchical classification procedure, where a lower resolution feature map is classified to create an array of probabilities corresponding to the probability that an area of the three- dimensional image includes a portion of the body structure. Interpolation can be performed on the probability array to increase its resolution and allow it to be included within the feature map of a higher resolution for classification at the higher resolution.
  • the classification procedure may be repeated until classification is performed at the original resolution of the input three-dimensional image.
  • the body structure segmentation system 100 may be configured to segment one or more body structures of interest including, but not limited to: glands including the prostate, thyroid, etc.; organs including the heart, the liver, etc.; bones; joints; muscles; tumors; and other body structures for which segmentation may be useful.
  • the body structure segmentation system 100 includes a segmentation coordinator 110, a feature extractor 112, feature extraction model storage 114, a pooling calculator 116, a position feature generator 118, classifiers 120, an upsampler 122, a concatenator 124, and a mask generator 126, the functionality and features of each will be described in more detail as related to FIGS. 2-5 herein.
  • the segmentation coordinator 110 may be configured to control the timing and flow of data through the other circuitry of the body structure segmentation system 100. For example, the segmentation coordinator 110 may cause the instructions or circuits to execute in a specific order to perform the functions of the body structure segmentation system 100.
  • the segmentation coordinator 110 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input.
  • the segmentation coordinator 110 may cause the instructions and/or circuits of the body structure segmentation system 100 to perform a flow of operations 300 in FIG. 2 to segment a body structure.
  • the flow of operations 300 includes obtaining (e.g., acquiring, receiving, etc.) a slice of a three-dimensional image, the three-dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location (e.g., x-y location) of the three-dimensional image in operation 302, according to some embodiments.
  • the imaging system 30 may communicate the three-dimensional image over the network 40 to begin processing (e.g., segmentation).
  • the segmentation coordinator 110 requests (e.g., from a database, from imaging system 30, etc.) one or more three-dimensional images to process (e.g., to perform batch processing).
  • the imaging system 30 may select (e.g., based on a user input, such as based on an input from a radiologist, lab technician, etc.) a three-dimensional image for segmentation to assist in the evaluation of the imagery.
  • the segmentation coordinator 110 may select a slice (e.g., cross section) of the three-dimensional image to begin processing.
  • the segmentation coordinator 110 may process the subsequent slices in depth-wise order using similar techniques.
  • the flow of operations 300 includes executing (e.g., running, performing, etc.) a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice in operation 304.
  • the feature extraction model may be executed by the feature extractor 112 using models stored in the feature extraction model storage 114.
  • the feature extraction model may depend on the feature map being processed (e.g., the stage of the feature extraction process), the body structure being segmented, the resolution of the image and/or feature map, and/or any other attribute of the original three- dimensional image or of the intermediate feature maps.
  • a voxel may refer to a three- dimensional pixel, and in some embodiments, voxel and pixel are used interchangeably.
  • the feature extractor 112 may be configured to generate a set of features from a local region (e.g., patch, area, etc.) of the slice.
  • the slice of the image for example, may be represented by an array of numeric values corresponding to the grey scale intensity of the image.
  • the feature extractor 112 obtains a feature extraction operator (e.g., a function, a linear transformation, a matrix, etc.) from the feature extraction model storage 114 and performs the feature extraction operation on a vectorized version of the local region. For each local region of the slice, the feature extractor 112 may output a set of features (e.g., an array).
  • the local region processed by the feature extractor 112 is then shifted (e.g., moved by a voxel in one direction) and the operation is repeated to obtain a set of features for the current local region of the slice.
  • the feature extractor 112 may be configured to scan the processed local region over the slice and organize the set of features output for each local region into a three-dimensional array for the slice.
  • the set of features may be stored at indices for the first and second dimension equal to the indices of a voxel (e.g., the center voxel), of the local region and each feature of the set may be indexed into the third dimension (e.g., form a pixel embedding at that location).
  • the three-dimensional array for the slice has height and width equal to that of the slice and depth equal to the number of features output by the feature extraction operator.
  • a feature at all height and width indices for a single depth (e.g., one feature index) may be referred to as a feature slice in some embodiments.
  • the feature extractor 112 may perform the same or a similar feature extraction technique on a feature map (e.g., the output of the feature extraction may be further processed by performing feature extraction on one or more feature maps).
  • the feature extractor 112 may obtain a different feature extraction operation from the feature extraction model storage 114 for each feature map.
  • the feature extraction model storage 114 may be configured to maintain (e.g., store, save, etc.) the feature extraction models and may also store the feature map that a specific feature extraction model should be used on.
  • the body structure segmentation training system may provide a map (e.g., structure, object, dictionary, or other suitable storage elements) that contains for each feature map (including the original image) a feature extraction operation.
  • pooling may be performed by pooling calculator 116.
  • pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output.
  • the voxels may be combined by taking the maximum of the pixels, by taking the average of the pixels, or any other suitable function of values of multiple voxels.
  • two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution.
  • three- by-three patches of voxels are combined resulting in an output resolution that is a third of the input.
  • rectangular patches are combined in the pooling operation. It is contemplated that the pooling calculator 116 may perform different pooling operations and/or use different sized patches based on the current feature map that is being processed or the original input image. For example, three-by-three voxel patches may be used in a first stage of feature extraction and two-by-two voxels may be used in a second stage of feature extraction.
  • the flow of operations 300 includes executing a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual in operation 306.
  • the first feature map may be a low-resolution feature mask and the machine learning model may be a classifier of the classifiers 120.
  • positional features are concatenated with the feature map (e.g., using the position feature generator 118 and the concatenator 124).
  • the probabilities of the first mask may be compared to a threshold to generate a coarse binary mask indicating regions that may depict a portion of the body structure.
  • the classifiers 120 may be configured to store one or more machine learning models.
  • the machine learning models may classify an area of the three-dimensional image as representing at least a portion of the body structure of interest based on the features of the feature map.
  • the feature map may represent a slice of the three-dimensional image and machine learning model (e.g., a classifier) may generate a probability that the features at a given x-y location represent at least a portion of the body structure of interest.
  • a single slice is used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of that slice (respective to locations of the feature map) represent at least a portion of the body structure.
  • multiple slices are used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of each of the multiple slices (respective to locations of the feature map) represent at least a portion of the body structure.
  • multiple slices are used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of one of the multiple slices (e.g., the center slice) represent at least a portion of the body structure.
  • the machine learning model used by the classifiers 120 depends on the feature map being processed. For example, three slices may be used to generate feature vectors that are in turn used as input to a first model of classifiers 120 to classify each of the three slices in a first stage of classification and a second model of classifiers 120 may be used to classify only the center slice at a second stage (e.g., the last stage) of classification. By utilizing more than one slice to generate some of the features, information from adjacent slices is incorporated into the classification; however, the classifiers may still be configured to generate a single mask representing the segmentation of the body structure at that slice.
  • the classifiers 120 may be configured to use one or more types (e.g., architectures, classes, etc.) of machine learning models to generate the classification.
  • the machine learning model may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination types of classification models.
  • the type of classification architecture may depend on the feature map being processed (e.g., the stage of the classification process), the body structure being segmented, the resolution of the image and/or feature map, and/or any other attribute of the original three-dimensional image or of the intermediate feature maps.
  • the position feature generator 118 may be configured to generate features for a feature map that correspond to a location relative to the original three-dimensional model and/or the lower resolution feature maps.
  • the position feature generator 118 may produce a three-dimensional array that includes the x, y, and z location of a voxel corresponding to the one set of features (e.g., x and y index of the feature map).
  • the x and y locations of the positional features may relate to the x and y location in the feature map because at lower resolutions the features at a specific x-y location correspond to several voxels.
  • the positional feature generator may produce the following three dimensional array: indicating the x-y location of the feature map and the z location of the slice (e.g., 10). Adding positional features may be useful as they can cause the sensitivity to increase near the center of the image (e.g., the majority of the images will have the structure of interest near the center of the image) and/or allow the classifiers 120 to use the positional information in other ways. Positional information can be encoded in a variety of ways in addition to using the x, y location of the feature map and the z location of the original slice.
  • the positions can relate to the original region used to generate the feature map. For example, if the original resolution of a slice was 512 by 512, the first position of the positional three-dimensional array may be 32.5 representing the average index of the 64 by 64 patch of voxels that are represented by the position in the lower resolution feature map.
  • positional features may be added based on their position relative to a specific reference point in the image (e.g., the center). For example, if the distance to the center of the three-dimensional image is used to add positional information by the position feature generator 118, fewer than three-dimensions can be used to encode the positional information (e.g., distance can be represented by a single number).
  • index may be used to encode positional information.
  • sines and/or cosines of the positional index can be used, potentially leading to more than three positional features (e.g., a three-dimensional array with more than three indices in the third dimension).
  • the concatenator 124 is configured to add features to a feature map.
  • the new features may be conformant to the current feature map.
  • a feature map with a resolution of N by N may require the additional features to also have a resolution of N by N before concatenation.
  • the positional features from the positional feature generator may be concatenated with the feature map using the concatenator 124.
  • the concatenator 124 may add the positional features described herein along the feature dimension (e.g., third dimension). For example, if the feature map is 8 by 8 by 45 (e.g., a resolution of 8 by 8 and a feature depth of 45), 8 by 8 positional features can be added to the depth of the feature map.
  • the flow of operations 300 includes concatenating each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution in operation 308.
  • the second resolution may be greater than the first resolution and it may be necessary to scale the first mask (e.g., the lower resolution classification result) to the resolution of the second feature map using interpolation.
  • the body structure segmentation system 100 includes the upsampler 122 to increase the resolution (e.g., scale, upsample, etc.) of a classification result (e.g., a probability mask).
  • the upsampler 122 may include various interpolation techniques including trilinear interpolation. Trilinear interpolation can estimate values within a three- dimensional grid by linearly interpolating along each of the three axes (x, y, and z). Trilinear interpolation involves using eight surrounding data points in a cube to compute the interpolated value based on their distances to the target point.
  • the upsampler 122 may perform bilinear interpolation (e.g., if the input is only two-dimensional or if it is desired that interpolation is performed within a single slice).
  • the upsampler 122 includes additional interpolation techniques including: tetrahedral interpolation, spline interpolation, radial basis function interpolation, etc. that can be used to increase the resolution of the classification result to conform with a feature map of a higher resolution.
  • the flow of operations 300 includes executing a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual in operation 310.
  • the classifiers 120 may include a second classifier that operates at the higher (e.g., second resolution) and classifies each pixel embedding (e.g., at an x-y location) of the higher resolution feature map.
  • the second classifier of the classifiers 120 has been trained to use the classification result (e.g., the probability mask) of the lower resolution classifier, for example, scaled to the higher resolution using upsampler 122 in the previous operation.
  • Operations 308-310 can be repeated a number of times for each of the various resolution feature maps generated by the feature extractor 112 in operation 304.
  • the feature extractor 112 may produce feature maps of at 4 different resolutions.
  • Operation 306 can be performed at the lowest resolution (e.g., fourth resolution), the scaled results concatenated with the third resolution in operation 308, and the third resolution classified in operation 310.
  • the results of the classification in of the third resolution can be scaled by upsampler 122 to a second resolution by repeating operation 308 and concatenated to a feature map at the second resolution before classification in the second pass of operation 310 to produce a second resolution probability map. Repeating the steps 308 and 310 one more time will result in a probability mask at the original resolution of the three-dimensional image.
  • the flow of operations 300 includes generating a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings in operation 312.
  • the second mask for example, may be a probability mask as described with reference to operation 306 or the second mask may be a binary mask.
  • a binary mask may represent a final classification of a pixel (e.g., based on the pixel embedding related to that pixel at a specific x-y location) indicating if the pixel is part of the structure being segmented.
  • the operation 312 may be performed by the mask generator 126.
  • the classifiers 120 are trained to correct errors in the lower resolution probability mask using a regression model (e.g., rather than a classification model).
  • the ground truth of the labeled training data may be subtracted from the output of a classifier 120 during the training process to allow the model to focus on the difficult to segment boundary regions of the body structure.
  • the probabilities e.g., the mask
  • This process may be repeated for any number of layers the body structure segmentation process.
  • the mask generator 126 may compare a probability mask to a threshold to determine a binary mask. For example, each pixel embedding that scores a probability that it depicts at least a portion of the structure greater than a threshold may indicate a location depicting a portion of the body structure.
  • the mask generator 126 includes logic in addition to comparing the probability to a threshold.
  • the mask generator 126 may be configured to include a type of hysteresis to cause the segmentation to generally form larger grouping of pixels and eliminate noise by performing the comparison to a second, lower threshold if any adjacent pixel exceeded the first threshold.
  • boundaries may be drawn that encompass the pixels that exceeded the threshold to generate the final mask. Additionally, the mask generator 126 may include curve simplification algorithms to reduce the effect of noise at the boundary of the mask.
  • a visual representation of the mask may be generated by the mask generator 126.
  • the visual representation may include an overlay for the original image that can be communicated to the imaging system 30, the client device 20, or any other computer system where the segmentation results are viewed.
  • multiple masks may be generated based on multiple thresholds and differently colored overlays can be generated for each threshold.
  • FIG. 3 is a flow diagram illustrating the shape of the data (e.g., representing the three- dimensional image) as it is processed by the body structure segmentation system 100, according to some embodiments.
  • the data flow and shape will be described for a configuration where the output is a binary (or probability) mask for a slice of the three-dimensional image.
  • the operations may be repeated to generate a binary (or probability) mask for each slice of the three-dimensional image.
  • extensions of the described process wherein the operations produce a mask for two or more slices after performing a single traversal through the operations illustrated in FIG. 3 should be considered within the scope of the present disclosure.
  • the operations begin by receiving a three-dimensional image including the structure of the individual.
  • One or more layers (e.g., slices) of voxels are taken from the three- dimensional image.
  • input slices 502 may include C slices, each with a resolution of Ni by Ni.
  • the feature extractor 112 is configured to produce several features from the input slices 502.
  • Features may be extracted by projecting an input vector formed from a patch of an input slice 502 onto several anchor vectors, expressed as an affine transform as: equation 1 where m is an integer from 0 to M - 1 and M is the number of anchor vectors (and features) that will be extracted from each patch or input vector.
  • a m represents the m lh anchor vector and x is the input vector obtained from a local region (e.g., patch, area, etc.) of one of the input slices.
  • the bias term b m is used to ensure that all features are positive.
  • Overlapping patches may be passed through equation 1 to generate a first feature map 504.
  • the first feature map 504 may have a resolution the same or similar to that of the input slices 502. For example, voxels near the edge of a slice may either be created using zero padding or another boundary technique or the voxels near the edge may not have a respective pixel embedding in feature map 504 (e.g., reducing the resolution by a small amount).
  • a corresponding pixel embedding is generated at the same x-y location in the feature map 504.
  • the corresponding pixel embeddings may have Ki features.
  • each slice may provide Ki/C features to the pixel embedding at the location of the feature map 504 corresponding to the location of the voxel in an input slice.
  • the center slice for which an output mask is generated may be allocated more anchor vectors and thus contribute to more features than the other slices.
  • the feature map 504 may have a size of Ni by Ni by Ki.
  • the first feature map 504 is stored (e.g., in memory) to be used in classification at the original resolution and is also subjected to a pooling operation by pooling calculator 116 as described previously.
  • the pooling operation may reduce the resolution of the first feature map 504.
  • the height of the first feature map 504 may be divided by the height of the patches of the pooling operation and the width of the first feature map 504 may be divided by the width of the patches of the pooling operation.
  • the output of the pooling operation is a pooled first feature map 506 with size of N2 by N2 by Ki.
  • the feature extractor 112 may apply a feature extraction operation to each feature set (e.g., slice) of the pooled first feature map 506.
  • the feature extractor 112 may apply equation 1 (e.g., using anchor vectors that are different than those used to transform input slices 502 to first feature map 504) to each slice of the pooled first feature map 506.
  • the anchor vectors and/or the number of anchor vectors is different for each feature set (e.g., channels) of the pooled first feature map 506.
  • Each location (e.g., x-y location) of the pooled first feature map may be used to generate a corresponding pixel embedding as described previously to generate the second feature map 508.
  • the second feature map 508 is stored (e.g., in memory) to be used in classification at the original resolution and is also subjected to a pooling operation by pooling calculator 116 as described previously. It is noted that the pooling operation performed by the pooling calculator 116 on the second feature map 508 may be different (e.g., pool a larger/smaller number of pixels from the second feature map 508 together).
  • the pooling calculator 116 may produce a pooled second feature map 510 that has the same number (e.g., K2) of features, but lower resolution (e.g., N3 by N3). Feature extraction may be repeated using the feature extractor 112.
  • Training may be performed to determine another set of anchor vectors for each feature set of the pooled second feature map 510 or previous anchor vectors may be used (e.g., during training multiple feature map layers may be combined into the same training set).
  • the result of the feature extractor 112 may be a third feature map 512 at the same N3 by N3 resolution, but with more features (e.g., K3 feature) as each of the K2 features of the pooled second feature map 510 may use one or more anchor vectors to produce one or more features of the third feature map 512.
  • the third feature map is again stored and the process of pooling and feature extraction may be repeated another time to obtain a pooled third feature map 514 and forth feature map 516 with K4 > K3 features and a resolution of N4 by N4 where N4 ⁇ N3. Any number of feature extraction and/or pooling steps may be performed, the number of pooling operations may ultimately be limited by the resolution of the original slices and the size of the patch used during pooling.
  • the encoding process of the body structure segmentation system 100 is complete after a number of feature extraction and pooling steps have been performed. For example, after four feature extraction steps and three pooling steps as shown in FIG. 3.
  • the encoding process may result in a number of feature maps (e.g., four) available for classification by the classifiers 120.
  • Each pixel embedding (e.g., a depth-wise array at a given x-y location of the feature map) corresponds to an array of the original slice of the three- dimensional image.
  • the position feature generator 118 generates feature maps related to the position of each feature embedding as described previously.
  • the positional features may be concatenated with the feature maps (e.g., feature maps 504, 508, 510, and 512) by the concatenator 124.
  • the positional features may add 3 feature sets two the pixel embeddings of the feature map, one feature set indicating the x location of the pixel embeddings, one indicating the y location of the pixel embeddings, and one indicating the z location of the pixel embeddings.
  • classification is performed starting at the lowest resolution rolling up information towards the highest resolution (e.g., the original resolution) of the three-dimensional image.
  • a coarse classification may be performed using the feature map with the largest pixel embeddings (e.g., largest number of features).
  • a coarse classifier with a larger number of features may be able to accurately classify various regions of the three-dimensional image without a large amount of training data and/or the computational burden associated with a large number of training data.
  • Information from the low-resolution feature map e.g., the fourth feature map 516) is included in the probability mask indicating the probability that a pixel embedding at a given location represents at least a portion of the body structure of the individual.
  • the lowest resolution feature map (e.g., the fourth feature map 516) may be concatenated with positional features and passed to a classifier of the classifiers 120.
  • each classifier is uniquely trained by pixel embeddings (and/or feature maps) at the respective layer within the encoding process to determine a probability mask.
  • the classifiers may include a uniquely trained classifier for each of the resolutions of feature maps (e.g., classifier 120a-d).
  • Classification at each layer may be considered a voxel-wise classification problem (e.g., each a pixel embedding at a x- y location representing a voxel of the original slice is presented to a classifier of the classifiers 120).
  • each pixel embedding of the fourth feature map 516 is concatenated with the positional features is presented layer 4 classifier 120a to determine a coarse probability map 518.
  • the classifier 120a may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination of types of classification models.
  • the coarse probability mask 518 may include a probability for each voxel of the original input slices 502 or the coarse probably map 518 may include a probability only for the center slice (e.g., the slice for which the final resolution mask will be created).
  • the layer 4 classifier 120a may include a smoothing function to increase smoothness between classification results when the individual voxels are classified independently.
  • values of the coarse probability mask 518 may be averaged with adjacent values, the median of adjacent values may be calculated, thresholds to generate a binary classification may be modified if adjacent values are above a first threshold, or additional classifiers may be trained to smooth the probability mask (e.g., soft labels). Smoothing may similarly be performed on classification outputs (e.g., probability masks) at any resolution.
  • a local refinement step is performed after the classification (e.g., after the layer 4 classifier 120a produces soft decisions for each of the input slices used to generate the feature maps). For example, a cube of the soft decisions (e.g., 3 by 3 by 3, etc.) may be selected from the probability mask. Each cube may be used to form a feature vector to train another classifier (e.g., an xgboost classifier). During segmentation of a new image, the soft decisions may be used to form a feature vector in the same way and the trained classifier may be used to generate a smoothed probability mask (e.g., smoothed soft decisions).
  • a smoothed probability mask e.g., smoothed soft decisions.
  • the smoothing classifier may be used for a single resolution (e.g., a uniquely trained smoothing classifier may be trained for each resolution) or the same smoothing classifier may be used by each resolution.
  • smoothing classifiers are cascaded to update the soft decisions using multiple iterations.
  • a median filter may be applied after multiple iterations of the smoothing classifiers.
  • the coarse probability mask 518 may be passed to the upsampler 122 (e.g., for scaling, upsampling, etc.) to generate soft decisions at the same resolution as the third feature map 512 so that information from the coarse probability mask 518 can be used by the layer 3 classifier 120b.
  • interpolation can be performed using trilinear or bilinear interpolation as described previously.
  • the third feature map 512 may be concatenated with positional features for the current resolution and the upsampled version of the coarse probability mask 518 by concatenator 124 to generate the pixel embeddings that are presented to the input of the layer 3 classifier 120b.
  • the classifier 120b may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination types of classification models to classify each pixel embedding at the third resolution as representing a portion of the structure of the individual.
  • the classification results for each pixel embedding form a probability mask at the third resolution 520.
  • the process of upsampling, concatenating, and classifying may be repeated to form the probability mask at the second resolution 522 and the probability mask at the original resolution 524.
  • the upsampled probability masks that are concatenated onto each feature map may include upsampled versions of all previous probability masks.
  • coarse probability mask 518 may be upsampled and concatenated with feature map 512
  • the upsampled version of the coarse feature map 518 may also be concatenated with the probability mask at the third resolution 520
  • the combination may be together upsampled and concatenated with he probability mask at the second resolution 522, the combination again upsampled and concatenated with the first feature map 504 and used as input to the layer 1 classifier.
  • Equation 2 The recursive upsampling operations is shown by the equation: equation 2 where Fp represents a number of upsampled probability masks at the ith layer, y 1+1 represents the probability mask after classification at the i+1 layer (e.g., the next coarser resolution), ⁇ is the concatenation operator, and ⁇ T is the upsampling operator (e.g., bilinear interpolation, etc.).
  • the final classifier (e.g., the layer 1 classifier 120d) may be configured to only output a single probability mask (e.g., not include adjacent slices) or the layer 1 classifier 120d may determine three probability masks (e.g., in order to better perform smoothing) and then discard the extra masks before giving the result for the center slice of the input slices 502.
  • a single probability mask e.g., not include adjacent slices
  • the layer 1 classifier 120d may determine three probability masks (e.g., in order to better perform smoothing) and then discard the extra masks before giving the result for the center slice of the input slices 502.
  • an entire three-dimensional image is processed slice by slice, each slice results in an output probably mask which can be concatenated together to form a final mask for the three-dimensional image.
  • the final mask may be communicated over network 40 to the imaging system 30 or the client device 20 for viewing. For example, all voxels for which the final mask is greater than a threshold probability may be highlighted using a digital color overlay on
  • the probability mask or the binary mask may include a number of visual representations.
  • One or more visual representations can be delivered to a user interface.
  • the mask can be viewed as a color overlay on a screen, a virtual object in an augmented reality environment, an object on a holographic display, etc.
  • the body structure segmentation system 100 is configured to generate a user interface (e.g., by the segmentation coordinator 110).
  • the segmentation coordinator 110 may send instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images being segmented as well as the mask and/or color overlay.
  • the segmentation coordinator 110 may include instructions to generate a local user interface on the computing hardware embodying the body structure segmentation system 100 (e.g., a desktop computer, a server, the imaging system 30, etc.).
  • the user interface may be configured to provide user interaction (e.g., in addition to viewing the results).
  • a drag-and- drop interface may be provided to move a node of the segmentation boundary, to add voxels to the segmented body structure, etc.
  • Segmentation adjustments can improve future segmentation results by providing additional training data that has been expertly labeled and the current body structure segmentation system 100 was unable to predict adequately.
  • the new training data may be stored in the body structure segmentation training system 150.
  • the body structure segmentation system 100 is configured to segment a body structure and the various regions of the body structure.
  • Each classifier may produce a classification result (e.g., a probability) for each region or a separate classifier may be trained for each region.
  • the body structure segmentation system 100 may be trained to detect the central zone, peripheral zone, transitional zone, and anterior zone of the prostate.
  • the body structure segmentation system 100 may be trained to segment the four chambers of the heart.
  • FIGS. 4 and 5 provide more detail to the feature extraction and pooling process performed by the body structure segmentation system.
  • FIG. 4 shows flow of operations 320 for generating a reduced resolution feature map by feature extraction and pooling, according to some embodiments.
  • FIG. 5 is a flow diagram illustrating the operations of the feature extraction and pooling according to some embodiments. The flow of operations 320 is described with reference to FIG. 5 using the second feature extraction operation and second pooling operation of FIG. 3 as an example (e.g., the operations that convert the pooled first feature map 506 to the pooled second feature map 510.
  • the flow of operations 320 begins by obtaining a local region (e.g., a patch, area, set of adjacent voxel s/pixels, etc.) of a slice of a first plurality of feature slices (e.g., a feature map) in operation 322.
  • a local region e.g., a patch, area, set of adjacent voxel s/pixels, etc.
  • the pooled first feature map 506 is sliced into feature slices 530-534 (e.g., one for each of the Ki features of the pooled first feature map 506) and a local region (e.g., region 536) is obtained.
  • the local region may be vectorized to obtain a vector representation of the data of the local region in operation 324.
  • the local region 536 is converted from the two-dimensional form (e.g., 3 by 3) to a single dimension array (e.g., 9 by 1) to form the vector representation 538.
  • the flow of operations 320 may include performing a feature extraction operation on the vector representation, the feature extraction operation, for example, determined by a Saab transformation of feature slices from a plurality of images in operation 326.
  • the calculation of a feature was shown in equation 1.
  • multiple anchor vectors can be combined into a matrix representation (e.g., matrix 540) and multiplied by the vector representation (e.g., the vector representation 538).
  • the matrix 540 incorporating the anchor vectors is a column of the transposed anchor vectors as shown in: y equation 3 and x is the vector representation of the local region (e.g., vector representation 538) and y includes the features.
  • y represents a portion of the pixel embedding (e.g., portion 542) for the local region of the current feature slice.
  • the features are distributed across a respective x-y location of a second plurality of feature slices of the same resolution in operation 328.
  • the portion of the pixel embedding 542 e.g., the features
  • the flow of operations 320 may include repeating operations 322-328 for each overlapping local region to complete the second plurality of feature slices in operation 330.
  • Another local region may be obtained from the same feature slice 530.
  • the next local region 546 processed may be selected by shifting the local region 536 by a number of locations in one direction (e.g., one location left, one location down, etc.). The next local region 546 may overlap with local region 536.
  • next local region 546 is vectorized (e.g., in operation 324), multiplied by the matrix defining the feature extraction operation (e.g., in operation 326), and distributed across feature maps or feature slices 544a-c to form another portion of a pixel embedding adjacent to the portion of the pixel embedding obtained from local region 536.
  • the operations 322-330 may be repeated to from a pixel embedding for each location of the feature slice 530 and complete the layer 2 feature slices 544a-c.
  • the feature map 506 may include several feature slices (e.g., feature slices 530-534) that can be used to determine new features for the output feature map 508.
  • the flow of operations 320 includes repeating operations 322-330 for each feature slice of the input feature map to obtain a plurality of feature slices for each feature slice of the input feature map.
  • local regions of the feature slice 532 are vectorized and multiplied by a second matrix 548 to determine a pixel embedding 550 for a second plurality of feature slices 552.
  • the second matrix 548 may be the same as the matrix 540 (e.g., representing the same anchor vectors, same feature extraction operation, same Saab transformation, etc.).
  • the matrix 540 may represent anchor vectors specifically trained for the feature slice 532. It should be understood that any slice may use a feature extraction operation that is trained for the feature of the slice (e.g., using training data of that feature) or may be trained for a number of features (e.g., trained using training data from several features).
  • a feature extraction operation may be performed for each slice (e.g., feature slice 530- 534) of the input feature map 506 to generate respective pluralities of feature slices of the same resolution (e.g., pluralities of feature slices 544, 552, and 554). Each of the pluralities can be concatenated together to form the output feature map 508 of the feature extractor 112.
  • the feature slice 530 may be used to generate Q features and represent the first Q features of the output feature map 508, the feature slice 532 may be used to generate S features and become features S+l to S+Q of the output feature map 508, and the final feature slice 534 of the input feature map 506 may be used to generate R features and become features K2-R+1 to K2 of the output feature map. Any feature slices of the input feature map 506 between feature slice 532 and 534 may be used to generate additional features that also become part of the output feature map 508.
  • pooling calculator 116 maps a local region of a feature map to a single location of a reduced resolution feature map.
  • Pooling may include taking the maximum, average, or any other operation of a local region of a feature (e.g., median, second maximum, etc.) of the P by P local region.
  • the local region does not need to be square nor do the local regions need to be non-overlapping.
  • a 4 by 2 local region can be shifted by 2 in each direction resulting in regions that overlap in one of the two directions with the output resolution reduced by a factor of the amount the local region is shifted (e.g., the stride, in the present example, 2).
  • FIG. 6 shows another block diagram of the imaging and structure segmentation system 10 according to some embodiments.
  • FIG. 6 shows some embodiments of an implementation of the body structure segmentation training system 150.
  • the body structure segmentation training system 150 may include a communications interface 152 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40.
  • the communications interface 152 may be used to communicate the trained machine learning models (e.g., the feature extraction models, and the classifiers and/or their parameters) to the body structure segmentation system 100.
  • the communications interface 152 may be of the same or similar type of the communications interface 102 of the body structure segmentation system 100.
  • the body structure segmentation system 100 may also include a processing circuit 154 having a processor 156 and memory 158.
  • the processor 156 may be configured to execute instructions contained on the memory 158.
  • the body structure segmentation training system 150 and the body structure segmentation system 100 are implemented on the same hardware (e.g., on hardware of the imaging system 30, on the same computer, etc.).
  • the processing circuit 154 may be the same processing circuit as the processing circuit 104, similarly the communications interface, memory and processors may be shared between the body structure segmentation system 100 and the body structure segmentation training system 150.
  • the body structure segmentation training system 150 may be implemented on separate hardware from the body structure segmentation system 100.
  • the processor 156 may be of similar type or of different type from the processor 106.
  • the processor 156 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components.
  • the processor 156 may be configured to execute computer code and/or instructions stored in the memory 158 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.).
  • the processor 156 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof.
  • graphics processing units GPUs
  • One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
  • the memory 158 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure.
  • the memory 158 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions.
  • the memory 158 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure.
  • the memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
  • the body structure segmentation training system 150 is to acquire several three-dimensional images of a portion of the body for training the feature extractor 112 (e.g., the feature extraction operations) and the classifiers of 120 of the body structure segmentation system 100.
  • the body structure segmentation training system 150 applies unsupervised learning to identify features of interest for the classifiers 120.
  • the feature extractor 112 may be trained to determine feature maps of various resolutions to be used by the classifiers 120.
  • the body structure segmentation training system 150 applies a supervised learning to train the classifiers 120 to determine if pixel embedding represents a location of the feature map that includes at least a portion of the body structure of the individual being imaged.
  • training the feature extractor 112 is performed with unsupervised learning it does not require previously segmented (e.g., labeled) images for training.
  • training using unsupervised learning allows unlabeled data to be used for training and greatly increases the available training data.
  • Previously segmented images may be provided to train the classifiers 120.
  • the segmented images are additionally used to train the feature extractor (e.g., without using the labeled segmentations) in some embodiments.
  • the body structure segmentation training system 150 applies unsupervised learning to train the classifiers 120 to determine if pixel embedding represents a location of the feature map that includes at least a portion of the body structure of the individual being imaged.
  • clustering algorithms may generate classification rules (e.g., boundaries, decisions, etc.) in the space of the pixel embeddings that can be used to determine if a location represents at least a portion of the body structure of the individual.
  • Clustering algorithms are not limited to dichotomizers but can also be used to cluster two or more expected regions of the body structure. For example, k-means clustering, hierarchical clustering, or any other suitable clustering algorithm may be used to determine pixel embeddings that represent various regions of the body structure imaged.
  • the body structure segmentation training system 150 includes a training coordinator 160, the feature extractor 112, training data storage 166, a data selector 168, the position feature generator 118, the classifiers 120, a loss calculator 172, a parameter tuner 174, and a Saab transformer 180.
  • a training coordinator 160 may be configured to control the timing and flow of data through the other circuitry of the body structure segmentation training system 150.
  • the training coordinator 160 may cause the instructions or circuits to execute in a specific order to perform the function of the body structure segmentation training system 150.
  • the training coordinator 160 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input.
  • the training coordinator 160 may cause the instructions and/or circuits of the body structure segmentation training system 150 to perform the training operations illustrated in FIGS. 7 and 8.
  • the training data storage 166 is configured to maintain (e.g., store, save, acquire, etc.) three-dimensional images of body structures that can be used to train the feature extractor 112, the classifiers 120, or both.
  • Three-dimensional images without an associated segmentation e.g., unlabeled
  • unsupervised learning e.g., of the feature extractor 112 or some classification algorithms.
  • Three-dimensional images of a body structure that has been previously segmented may be used to perform both unsupervised learning and supervised learning.
  • the data selector 168 selects the correct data to be used for a particular training algorithm.
  • the data selector 168 may select only labeled data when supervised learning is being performed to train the classifiers 120.
  • the data selector 168 may also split data into training and validation sets.
  • the validation data may be used to stop algorithm training, choose a classification algorithm, or to determine appropriate values of hyperparameters (e.g., the number of features that should be used, maximum depth of a decision tree, regularization terms, etc.)
  • the body structure segmentation training system 150 includes the Saab transformer 180.
  • the Saab transformer 180 may be used to determine the feature extraction operations of the feature extractor 112.
  • the Saab transformer 180 may be used to determine the anchor vectors used in the feature extraction operations, the number of anchor vectors to use, and/or the matrix representation of the feature extraction operations (e.g., the matrix 540 or 548).
  • the Saab transformer 180 includes a sampler 182, a DC vector calculator 184, and a principal components analysis module (PCA) 186 to perform the operations of the Saab calculation.
  • PCA principal components analysis module
  • FIG. 7 illustrates the details of the operations of the Saab transformer 180 according to some embodiments.
  • the Saab transformer receives a number of training samples from the training data storage 166.
  • the training samples may include three-dimensional images of a body structure for segmentation (e.g., three-dimensional image 560).
  • the sampler 182 may determine local regions of the three-dimensional image to be used to train the feature extraction operation.
  • the sampler 182 may choose the local regions based on one or more criteria for filtering appropriate local regions.
  • segmentation e.g., during operation of the body structure segmentation system 100
  • the feature extraction operation may be applied to local regions that satisfy the same or similar criteria.
  • the sampler 182 may select data based on the slice (e.g., the depth) to which the local region belongs and training different feature extraction operations for various layers.
  • the sampler may also disqualify regions, slices, or whole three-dimensional images based on one or more criteria. For example, a three-dimensional image may be disqualified from use in training if the signal to noise ratio is high, a slice may be disqualified from use in training if the image is subject to a blur, or any other undesirable image artifact may cause the image or slice to be disqualified.
  • multiple models may be trained based on criteria related to the three-dimensional image. For example, different feature extraction operations may be learned for three-dimensional images of differing noise levels.
  • the sampler 182 may select a local region (e.g., local region 562) and vectorize the data of the local region (e.g., 3 by 3 region) into a training vector for the Saab transformation related to the initial three-dimensional image.
  • the sampler 182 may select a local region displaced from the local region 562 by a number of voxels in any of the three directions (e.g., along the heigh, width, or depth).
  • the second local region may overlap with the local region 562 depending on the configuration of the sampler 182.
  • the sampling process may be repeated for all potential local regions of the three-dimensional image 562 and repeated again for all three-dimensional images in the training data storage 166 to form a set of training vectors 564.
  • the training vectors 564 are input to the Saab transformer 180 to calculate the anchor vectors of the feature extraction operation.
  • the anchor vectors (and thus the feature extraction operation), for example, may be used on all slices of the three-dimensional image.
  • the feature extraction operation may be used to transform the slices 502 to the feature map 504.
  • the training vectors 564 are input to the Saab transformer 180 and a DC anchor vector is determined by the DC vector calculator 184.
  • the DC vector calculator 184 may generate the DC anchor vector as: equation 4
  • the additional M - 1 anchor vectors are learned by performing principal components analysis (PCA) on the AC component of the training vectors 564.
  • PCA principal components analysis
  • the DC vector calculator 184 may subtract the DC component from each of the training vectors and the PC A calculator 186 may perform principal components analysis to calculate the principal components (e.g., the anchor vectors 566) that are used in the feature extraction operator.
  • performing PCA on the AC component of the training vectors 564 may include arranging the AC components of the training vectors 564 into a data matrix, multiplying the transpose of the data matrix by the data matrix, and calculating the eigen vectors (and in some embodiments eigen values) of the N by N result.
  • such calculations do not require the performance of iterations to determine an optimal value or training by adjusting the parameters over multiple (e.g., hundreds) of epochs of the training data, and such calculations lead to efficient feature extraction training.
  • the number of principal components selected (M - 1) to be used as anchor vectors is chosen based on the eigenvalues associated with the eigenvectors calculated during the PCA process.
  • the eigenvalues may represent the representative capability of each of the principal components.
  • the principal components that are associated with an eigenvalue of at least a threshold fraction (e.g., 10%, 20%, etc.) of the largest eigen value may be kept. In some embodiments, further processing is performed to determine the principal components to be kept.
  • the eigenvalues may be plotted in descending order and the “elbow” in the plot of the eigenvalues (e.g., the point at which the eigenvalues no longer decrease significantly) may be used as eigenvalue threshold to determine which principal components are used as anchor vectors.
  • the body structure segmentation training system 150 may use the feature extractor 112 and the pooling calculator 116 to convert slices of the three-dimensional images in the training data storage 166 into a feature map of reduced dimensionality (e.g., feature map 506).
  • the operations illustrated in FIG. 7 may be repeated to determine anchor vectors for each feature slice of the feature map of reduced resolution.
  • the sampler 182 may choose local regions 562 related to only a single feature (e.g., slice) of the feature map of reduced dimensionality and unique anchor vectors may be calculated and stored for each feature slice of the feature map of reduced resolution.
  • the process of using the feature extractor 112 with the new anchor vectors and performing pooling to calculate subsequent feature maps of reduced resolution is repeated to determine the sets of anchor vectors that are used by the feature extractor 112 at all layers of the encoding process. For example, to determine the anchor vectors used in the feature extraction operation to convert feature map 506 to feature map 508, to convert feature map 510 to feature map 512, etc.
  • a number of different feature slices can be used to determine the anchor vectors for a feature extraction operation that is used for all those feature slices. It is noted, that if more than one slice from the original image is used to generate the feature map and the same feature extraction operation is used for each slice, then the features generated from each slice may have similar information (but for a different slice) and using the same feature extraction operation may save computations in training, reduce the amount of training data required, and offer similar performance compared to performing training individually for each slice.
  • training is performed for each slice and feature slices are combined to use the same feature extraction operation if the anchor vectors satisfy a similarity criterion (e.g., pairs of anchor vectors between each set have a small angle between them, the dot product is near one, etc.)
  • a similarity criterion e.g., pairs of anchor vectors between each set have a small angle between them, the dot product is near one, etc.
  • FIG. 8 illustrates the operations to train the classifiers 120 according to some embodiments.
  • a slice 502 of the three-dimensional image (as well as adjacent slices) may be used to adjust the parameters of all the classification layers in the classifiers 120.
  • Training data slices may be provided to a labeler 190.
  • the labeler 190 may be configured to allow an expert to segment the body structure of the three-dimensional image so that it can be used in supervised training of the classifiers 120.
  • the labeler 190 may provide instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images and allow a user to select the regions of the three-dimensional image that correspond to the body structure.
  • the client device 20 e.g., JavaScript, cascading style sheets, etc.
  • labeler 190 may include instructions to generate a local user interface on the computing hardware embodying the body structure segmentation training system 150 (e.g., a desktop computer, a server, the imaging system 30, etc.). Segmented three-dimensional images may then be used to train the classifiers 120.
  • the computing hardware embodying the body structure segmentation training system 150 e.g., a desktop computer, a server, the imaging system 30, etc.
  • the features e.g., the pixel embeddings
  • the features used to determine if a location corresponds to a portion of the body structure may be calculated.
  • low resolution classification is used as part of the features to perform the higher resolution classification and the classifiers 120 are trained starting with the lowest resolution and working towards the resolution of the original three-dimensional image.
  • each slice of the training data is first passed through the sequence of feature extraction operations and pooling calculations to determine the corresponding feature maps of all resolutions. Positional feature may also be added to the feature map as show in FIG. 3.
  • a pixel embedding e.g., an array through the depth
  • the segmentation label e.g., the label as to whether or not the location represents part of the body structure
  • a number of pixel embeddings e.g., pixel embedding 570
  • the label at the respective location e.g., location 572 are generated from the three-dimensional images of the training data storage 166.
  • training of a classifier is performed iteratively.
  • the classifier 120 may determine a probability that the pixel embedding 570 represents a portion of the body structure.
  • the loss calculator 172 may compare the probability calculated by the classifier 120 to the label for the location 572 to determine a portion of a loss function.
  • the loss calculator 172 may compute the sum of the loss by comparing the classification of several pixel embeddings to the respective labels to compute a loss function.
  • the parameter tuner 174 may update the classifier 120 (e.g., parameters of the classifier, rules within the classifier, etc.) to improve the loss function with respect to the training data.
  • the image is labeled (e.g., segmented) at the original resolution and the loss calculator 172 determines a label for comparison purposes at the lower resolution.
  • the loss calculator may consider a location of the lower resolution image to be labeled as part of the body structure if any of the corresponding locations of the original resolution is labeled as part of the body structure, if a majority of the corresponding locations of the original resolution is labeled as part of the body structure, if all of the corresponding locations of the original resolution is labeled as part of the body structure, or any other suitable method for reducing the resolution of the segmented three-dimensional image.
  • the loss function may include binary cross entropy loss (e.g., for whole structure segmentation) or multi-class cross-entropy loss (e.g., for region segmentation).
  • the classifier 120 is a neural network classifier (e.g., perceptron machine, etc.) and in some embodiments, the classifier 120 includes one or more decision trees (e.g., xgboost).
  • the classifier 120a for the lowest resolution is trained as illustrated by the operations of FIG. 8. The process may be repeated as illustrated by FIG. 8 for the next higher resolution.
  • the feature maps of as previously calculated using the feature extractor 112 for the next resolution may be concatenated and with positional features.
  • the training data may be classified at the lowest resolution by the trained classifier 120a and upsampled to provide a probability map at the next higher resolution.
  • the upsampled probability map may be concatenated with the feature map and the positional features to obtain classifier inputs 568 for the next higher resolution.
  • Pixel embeddings may be classified by the classifier 120b (e.g., at the next higher resolution) and compared to the respective label by the loss calculator 172.
  • the parameter tuner 174 may then adjust the model of the classifier 120b so as to improve the loss function over all classifications of pixel embeddings of the training data at the current resolution.
  • the process may be repeated to train the classifiers 120c and 120d.
  • the classifier 120 at each resolution is a different type of classifier.
  • the classifier at the lowest resolution may be an xgboost classifier and the classifier of another resolution may be a neural network classifier.
  • the classifier 120 uses the same type of classification model for all resolutions.
  • xgboost may be used as the classifier at each resolution.
  • the process for training a classifier is repeated and the resolution of the classifier is increased with each resolution.
  • the data sets are related to the segmentation of the prostate gland.
  • the ISBI-2013 dataset consists of 60 training cases of axial T2-weighted MR 3D series, where half were obtained at 1 5T (Philips Achieva at Boston Medical Center) and the other half at 3T (Siemens TIM at Radboud University Nijmegen Medical Center). Since the ground truth segmentation includes background, peripheral zone (PZ), and central gland (CG), PZ and CG were merged into one class — prostate area.
  • the pixel spacing within each slice ranges from 0.39 mm to 0.75 mm, while the through-plane resolution ranges from 3.0 mm to 4.0 mm among different patients.
  • T2-Cube has smaller pixel spacing, especially along the z-axis. That means, higher resolution and thinner slices. Specifically, for T2-w pixel spacing ranges from 0.5 mm to 0.7 mm and the through-plane resolution (z-axis) from 3.0 mm to 4.0 mm. For T2-Cube pixel spacing ranges from 0.83 mm to 0.83 mm and the through-plane resolution (z-axis) from 1.4 mm to 4.0 mm.
  • the scanner used to acquire those images is the GE 3T with 8ch Cardiac coil.
  • the T2-Cube series is used since it may provide more accurate segmentation results because of the higher perspicuity of the images (stemming from the smaller voxel spacing).
  • the resolution of different images is regularized to the same physical resolution of 0.625 x 0.625 x 1.5mm.
  • Lanczos interpolation is used where the factor is calculated based on the original pixel spacing and through-plane resolution of each image. The through plane resolution is increased so that the segmentation in the 3D space can be more accurate.
  • contrast enhancement using CLAHE is applied.
  • PSHop segmentation algorithm disclosed herein
  • the input sequence is resized to 128 x 128.
  • a 256 x 256 centered crop around the segmented gland is resized to 128 x 128 to standardize the PSHop input.
  • DSC Dice Similarity Coefficient
  • PSHop has a more stable performance (low standard deviation in both experiments), regardless of the number of training samples.
  • This underlines one of the advantages of the current disclosure is that it is more stable even for fewer training samples, while large deep learning (DL) models fail to achieve high performance when data are scarce. That also confirms that the current statistical -based feedforward models can still perform well even with a small number of training samples.
  • Table 2 shows the benchmarking on USC-Keck dataset, since it is larger than the ISBI-2013 and thereby stronger conclusions can be drawn.
  • PSHop surpasses U-Net by large margins and maintains a small performance gap with V-Net.
  • PSHop achieves a much higher DSC score comparing to the U- Net, which is the baseline architecture in the art.
  • PSHop outperforms the other methods on the PZ (see Table 3), mainly due to the small number of training samples that PSHop has an advantage.
  • PSHop surpasses the U-Net performance by large margins.
  • V-Net generally performs better when trained with a large number of samples and U-Net performs well only on the ISBI-2013, which has many fewer samples. This is sensible because V-Net has many more trainable parameters comparing to U-Net and thus can fit better the data diversity.
  • Table 4 demonstrates the advantages of the current application when it comes to complexity and model size comparison. PSHop has an order of magnitude less parameters than the DL-based models. Also, in terms of complexity, it has 190 times fewer FLOPS than U-Net and 5269 times fewer FLOPS than V-Net.
  • PSHop maintains a very competitive standing performance-wise with other DL baseline models, outperforming in general U-Net in both tasks and having performance similar to V-Net at a significantly reduced model size and number of FLOPS.
  • the technological problem of identifying a body structure as diseased and/or determining if a region of interest of a body structure represents a diseased region may be improved using techniques similar to those used by the body structure segmentation system 100.
  • the diseased structure identification system may make use of similar unsupervised learning techniques to determine feature extraction operations that are configured to generate feature maps of various resolutions.
  • the techniques described herein may provide similar computational savings during online disease identification and during training.
  • the techniques may also provide a savings in the required amount of training data that has previously been labeled.
  • FIG. 9 is a block diagram of an imaging and disease identification system 12 configured to obtain medical imagery of a body structure and identify regions of a body structure that are indicative of a disease according to some embodiments.
  • the imaging and disease identification system 12 may detect disease in specific bone and/or joint structures, soft tissue (e.g., muscles, tendons, and ligaments), organs (e.g., the heart, lungs, kidneys, prostate, etc.), the nervous system, digestive tractor, or any other body structure for which disease may be detectable via medical imagery.
  • the imaging and disease identification system 12 includes one or more client devices 20, one or more imaging systems 30, a diseased structure identification training system 250, and a diseased structure identification system 200 shown to be communicably coupled via a network 40.
  • the network 40 can include routers, switches, antennas, computers, and any other hardware required to communicate information between the components of the imaging and disease identification system 12 (e.g., from the imaging system 30 to the diseased structure identification system 200.
  • a portion of the network 40 can be wireless and/or a portion of the network 40 can be wired.
  • the network 40 can include one or more networks with routers to facilitate data transfer between the different networks.
  • the imaging and disease identification system 12 is shown to be distributed across several devices (e.g., networked computers). It is contemplated that in some embodiments, the subsystems, components, and/or functionality of the imaging and structure segmentation system 10 may be distributed on different devices.
  • the diseased structure identification system 200 may be configured within the imaging system 30 using the processors of the imaging system 30 to execute.
  • the diseased structure identification system 200 and the diseased structure identification training system 250 are configured within the same computer hardware (e.g., within a server local to the imaging system, within a remote server, or on a node within a cloud computing architecture).
  • the diseased structure identification system 200 may include a communications interface 202 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40.
  • the diseased structure identification system 200 may also include a processing circuit 204 having a processor 206 and memory 208.
  • the processor 206 may be configured to execute instructions contained on the memory 208.
  • the processor 206 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components.
  • the processor 206 may be configured to execute computer code and/or instructions stored in the memory 208 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.).
  • the processor 206 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof.
  • One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
  • the memory 208 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure.
  • the memory 208 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions.
  • the memory 208 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure.
  • the memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
  • the diseased structure identification system 200 is to acquire a three-dimensional image of a portion of the body, perform a sequence of feature extraction operations and pooling operations on one or more sets of adjacent voxels (e.g., patches) surrounding a voxel of the three-dimensional image to obtain feature maps of various resolutions and varying depths (e.g., number of features), and to perform voxel-wise classification using the feature maps obtained from a set of adjacent voxels.
  • the feature maps are obtained for multiple imaging types (e.g., T2-weighted (T2w), apparent diffusion coefficient (ADC), and diffusion-weighted imaging (DWI).
  • the features may be combined to perform a voxel-wise classification based on the surrounding set of adjacent pixel of each of the multiple imaging types.
  • Regions of interest e.g., clusters of voxels identified as representing a diseased area
  • additional features e.g., radiomic features and/or anomaly features
  • the regions of interest may be classified as being a diseased area based on the combined feature set.
  • the diseased structure identification system 200 may be configured to identify diseased regions in one or more body structures including, but not limited to: glands including the prostate, thyroid, etc.; organs including the heart, the liver, etc.; bones; joints; muscles; tumors; and other body structures for which segmentation may be useful.
  • the diseased structure identification system 200 includes an identification coordinator 210, a feature extractor 212, feature selector 214, a region of interest (ROI) identifier 216, a region selector 218, a concatenator 124, a voxel classifier 222, and a stage 2 classifier 220, an anomaly feature collector 224 and a radiomics feature collector 226.
  • the identification coordinator 210 may be configured to control the timing and flow of data through the other circuitry of the diseased structure identification system 200.
  • the identification coordinator 210 may cause the instructions or circuits to execute in a specific order to perform the function of the diseased structure identification system 200.
  • the identification coordinator 210 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input.
  • the identification coordinator 210 may cause the instructions and/or circuits of the diseased structure identification system 200 to perform a flow of operations 400 in FIG. 10 to segment a body structure.
  • the flow of operations 400 includes obtaining (e.g., acquiring, receiving, etc.) a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels in operation 402, according to some embodiments.
  • the imaging system 30 may communicate the three- dimensional image over the network 40 to begin processing (e.g., segmentation).
  • the identification coordinator 210 requests (e.g., form a database, from imaging system 30, etc.) one or more three-dimensional images from a database to process (e.g., to perform batch processing).
  • a radiologist, lab technician, etc. selects a three-dimensional image for within which diseased regions are to be identified.
  • the flow of operations 400 includes identifying a set of adjacent voxels including the voxel (e.g., to be classified) of the plurality of voxels in operation 404.
  • a local, two-dimensional set of adjacent voxels of the three- dimensional region is obtained.
  • the identification coordinator 210 may select adjacent voxels from the three-dimensional image. For example, a 24 by 24 region may be selected.
  • the identification coordinator 210 may select overlapping sets of voxels, for example, by shifting the center of the set by one voxel as each voxel is classified.
  • the set of adjacent voxels may provide information (e.g., the adjacent pixels) that is used to determine if the voxel should be classified as being representative of a diseased area.
  • the flow of operations 400 includes executing a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions in operation 406.
  • the operation 406 may be performed by the feature extractor 212.
  • the feature extractor 212 may be configured similar to feature extractor 112.
  • To generate a set of features a local region (e.g., patch, area, etc.) of the of the set of adjacent voxels is obtained.
  • the set of the adjacent voxels may be a 24 by 24 patch and the local region used to generate the features may be a 3 by 3 patch within the set of adjacent voxels.
  • the feature extractor 212 may obtain a feature extraction operator (e.g., a function, a linear transformation, a matrix, etc.) from storage and perform the feature extraction operation on a vectorized version of the local region. For each local region of the set of the adjacent, the feature extractor 212 may output a set of features (e.g., an array) that represent the local area of the set of the adjacent voxels. Shifting the local area and again performing feature extraction with the same feature extraction operator may result in a feature map of the same or similar resolution of the set of the adjacent voxels.
  • a feature extraction operator e.g., a function, a linear transformation, a matrix, etc.
  • the three- dimensional array (e.g., feature map) for the set of adjacent voxels has height and width equal to that of the local region and depth equal to the number of features output by the feature extraction operator.
  • the resolution of the feature map generated e.g., the height and width
  • the resolution of the feature map generated is reduced by an amount that depends on the size of the smaller local region used to generate the pixel embeddings of the feature map of the local region. For example, if zero padding is not performed, features may not be extracted for voxels near the boundary of the set of adjacent voxels.
  • pooling may be performed by pooling calculator 116.
  • pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output.
  • the voxels may be combined by taking the maximum of the voxels, by taking the average of the voxels, or any other suitable function of values of multiple voxels.
  • two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution.
  • the feature extractor 212 is configured to perform feature extraction on each feature slice (e.g., slice of a feature map individually) of an input feature map. For example, feature extraction may be performed on a 11 by 11 by 9 input feature map resulting in a 9 by 9 by 56 feature map. Each of the 9 feature slices of the input feature map may provide a number of the 56 output feature maps. For example, the feature extractor may perform the operations demonstrated in FIG. 5. In some embodiments, the feature extractor 212 is configured to determine the number of features to keep for each feature slice of a feature map. For example, feature maps with low energy (e.g., low eigenvalues) may be discarded during the training phase.
  • feature slices with low energy e.g., low eigenvalues
  • the feature extractor 212 is configured to perform PC A to determine a spectrum on each feature slice of an input feature map.
  • the PCA may provide global features related to the entire set of adjacent voxels in addition to the local features obtained through the feature extraction operations.
  • global features may be extracted by:
  • the feature extractor 212 is configured to extract features from various imagery modes.
  • T2w, ADC, and DWI modes may each produce a three- dimensional image that is processed by the feature extractor 212.
  • the various imagery modes may provide important information to various tissue types and conditions.
  • the flow of operations 400 includes executing a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual in operation 408.
  • Voxel classifier 222 may perform operation 408.
  • Voxel classifier 222 may be configured to perform to use, as input, a number of features output by the feature extractor 212.
  • the voxel classifier 222 may be configured to use all the features extracted from the adjacent set of voxels. In some embodiments, a subset of the features (e.g., those with high discriminant potential) are used by the voxel classifier.
  • the voxel classifier 222 may be configured to produce a classification output (e.g., a probability or a binary decision) that the voxel included in the set of adjacent voxels (e.g., a center voxel) is part of a diseased region of the body structure.
  • the voxel classifier 222 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
  • the selection of adjacent voxels, execution of feature extraction, and classification of the voxel may be repeated for each voxel of the three-dimensional image.
  • the result of operations 404-408 for each voxel of the three-dimensional image is a probability for each voxel of the three-dimensional image that the voxel is part of a diseased region of the body structure.
  • the flow of operations 400 includes determining one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold in operation 410.
  • the ROI identifier 216 may be configured to identify regions of interest for further investigation.
  • the ROI identifier may compare the probabilities of the voxel-wise classification to a threshold and identify a region of interest anytime there are more than Nv connected voxels that exceed the probability threshold or a square or cube of size Ns for which all the voxels exceed the probability threshold.
  • the ROI identifier 216 is configured to average a number of voxels together.
  • the average may be a standard average or the average may be weighted based on a distance from the center of the region. Averaging the voxel-wise classification results may smooth the probabilities reducing the effect of any one voxel and decreasing the number of false alarms that a passed to second stage processing.
  • the ROI identifier 216 is configured to identify regions of interest in units of more than one voxel. For example, averaging probability scores may be computed for 8 by 8 regions of voxels (using the probabilities of voxel-wise identification of the voxels of the 8 by 8 region, voxels surrounding the 8 by 8 region or both). The average probability score associated with the region may be compared to a threshold to determine if the entire region should be included in a region of interest.
  • the ROI identifier 216 includes a threshold determined the diseased structure identification training system 250.
  • the threshold may be chosen to have a particular effect on the true positive rate and the false positive rate. For example, the threshold may be chosen to set the false positive rate at 5% or to set the true positive rate to 90% against the training set or an anticipated future set. The threshold may be chosen to minimize a false positive rate while constraining the true positive rate to satisfy a criteria (e.g., be greater than a certain value).
  • the flow of operations 400 includes executing, for each of the one or more regions of interest, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual in operation 412.
  • the stage 2 classifier 220 may perform operation 412.
  • the stage 2 classifier 220 may be configured to use any of statistical features of the voxel-wise classification probabilities (e.g., standard deviation of a local region, averages, etc.), the original features of the voxels making up the region of interest, and/or radiomic features (e.g., visual-based radiomic features such as first order statistics, gray level co-occurrence matrix, etc.).
  • radiomic features e.g., visual-based radiomic features such as first order statistics, gray level co-occurrence matrix, etc.
  • the difficult to determine radiomic features are only used on a few regions of interest rather than on the entire body structure.
  • additional features can be obtained by training a small CNN to identify features from an expanded receptive field around each ROI.
  • the CNN can use the probabilities from the voxel-wise classifier 222, a section of the original image, or any features used by the voxel-wise classifier as input.
  • FIG. 11 is a block diagram illustrating the data flow in the diseased structure identification system 200 according to some embodiments.
  • a number of three-dimensional images (e.g., three-dimensional images 602a-c) are obtained.
  • Each of the three-dimensional images may be of a different imaging type (e.g., modality).
  • the three-dimensional image 602a may be captured using the T2w mode of an MRI
  • the three-dimensional image 602b may be captured using the ADC mode of an MRI
  • the three-dimensional image 602c may be captured using the DWI mode of an MRI.
  • the three- dimensional images 602a-c may be aligned so a voxel location of one three-dimensional image represents the same area of the body structure scanned in each of the three-dimensional image.
  • the operations are performed for each voxel (e.g., each voxel is classified: a determination is made if the voxel represents a part of a diseased region of the body structure).
  • a set of adjacent voxels including the voxel that is to be classified is selected from same area of each of the three-dimensional images and input to the feature extractor 212.
  • the feature extractor 212 may generate a number of feature maps of various resolutions from the set of adjacent pixels.
  • a 24 by 24 set of adjacent pixels may be selected and provided to the feature extractor 212 and the feature extractor may generate a 22 by 22 by 9 feature map and a 9 by 9 by 56 feature map along with global PCA features for a total of approximately 9,000 features for each imaging type that can be used to classify a single voxel.
  • the feature extraction operation performed by feature extractor 212 is described in more detail herein with reference to FIG. 12.
  • the same feature extraction method is performed by the feature extractor 212 to the set of adjacent voxels of each of the imaging types.
  • the feature extractor 212 for each set of adjacent voxels may calculate a pixel embedding using equation 3.
  • each imaging type may have a different set of anchor vectors defining the feature extraction operation performed by the feature extractor 212.
  • the large number of features related to the set of adjacent pixels may not all contain a significant amount of discrimination potential.
  • the feature selector 214 may choose a subset of the features to use to classify each voxel.
  • the subset of features chosen by the feature selector 214 may be determined during training. All features may be ordered based on a discrimination metric.
  • the Discriminant Feature Test can be used as a metric associated with each feature’s potential to discriminate between a diseased region and a healthy region of the body structure.
  • the feature selector 214 is configured to select a fixed number (e.g., 500, 1000) of features with the highest discrimination potential.
  • the number of selected features is determined during training based on the discrimination metric.
  • the selected features from each of the imaging types are combined into a single set by the concatenator 124.
  • the voxel represented by the set of adjacent voxels may be classified by the voxel classifier 222 using the concatenated features.
  • the operations of selecting a set of adjacent voxels, performing feature extraction, feature selection, and classification may be repeated for each voxel location of the three-dimensional images 602a-c.
  • the classification may be repeated for up to as many as 33.5 million voxels to generate a probability (e.g., a soft classification) that the voxel is part of a diseased portion of the body structure.
  • a probability e.g., a soft classification
  • the ROI identifier 216 determines regions of interest where a number of nearby (e.g., adjacent voxels) have elevated probabilities (e.g., exceed a threshold). For example, the ROI identifier may determine clusters of elevated probabilities using rulebased algorithms and/or machine learning algorithms such as density-based spatial clustering of applications with noise (DBSCAN), Gaussian mixture models (GMMs), and/or ordering points to identify the clustering structure (OPTICS).
  • DBSCAN density-based spatial clustering of applications with noise
  • GMMs Gaussian mixture models
  • OTICS ordering points to identify the clustering structure
  • the ROI identifier 216 may eliminate false alarms that could be generated by any individual voxel having an associated high probability that the voxel belongs to a diseased region of the body structure.
  • the regions of interest may be passed to the region selector 218.
  • the region selector 218 may be configured to cause the collection of features for a second stage of classification. For example, the region selector 218 may obtain the features used by the voxel classifier 222 for voxels near the center of mass of the selected ROI. The region selector 218 may also be configured to cause the anomaly feature collector 224 and/or the radiomics feature collector 226 to acquire additional features for the stage 2 classification.
  • voxels averaged together by the ROI identifier 216 are used as features in the stage 2 classifier 220.
  • the averaged voxels may be acquired by the anomaly feature collector 224.
  • a weighted average of voxel classifications e.g., probabilities
  • Anomaly features may also include the mean and standard deviation of the voxel probabilities and/or the weighted average of voxel probabilities within a region near (e.g., encompassing, partially overlapping, within, etc.) the region of interest.
  • Using voxel probabilities averaged from several adjacent features may carry information from the voxel-wise classification results while allowing the stage 2 classifier to give weight to those regions that had more voxels classified as being from a diseased portion of the structure.
  • radiomics may provide several features for quantifying MRI characteristics of an ROI.
  • radiometric first order statistics may provide up to 19 features
  • 3D shape-based features include up to another 16 features
  • 2D shape-based features include up to 10 features
  • a gray level co-occurrence matrix may provide as many as 24 features
  • a gray level run length matrix can provide 16 features
  • a gray level size zone matrix can provide 16 features
  • a neighboring gray tone difference matrix can provide 5 features and/or the gray level dependence matrix may provide 14 features.
  • additional features are extracted by a convolutional neural network (CNN) to identify features from an expanded receptive field around each ROI.
  • the CNN can extract nonlinear features in addition to those obtained by performing the linear feature extraction techniques described herein.
  • the CNN can use the probabilities from the voxel-wise classifier 222, a section of the original image, or any features used by the voxel-wise classifier as input.
  • ROIs have been found using the voxel-wise classifier 222.
  • a CNN for identifying features around a region of interest does not need to accept the entire image as input and thus may not contribute excessive computational effort to the overall process.
  • the CNN is trained simultaneously with the stage 2 classifier 220.
  • the CNN may be trained simultaneously with the stage 2 classifier 220 so that it does not relearn features for information already provided by the linear feature extraction techniques, the radiomic features, or any other features that may also be provided to the stage 2 classifier 220.
  • the features used by the voxel classifier 222, the features collected by the anomaly feature collector 224, and the features collected by the radiomics feature collector 226 are combined by the concatenator 124.
  • Another feature selection process may be performed by the feature selector 214 to select the discriminant features for the stage 2 classifier 220.
  • the stage 2 classifier 220 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
  • the state 2 classifier 220 may provide a binary classification based on all the features provided including the those used by the voxel classifier 222, the anomaly features, and the radiomics.
  • the stage 2 classifier 220 may be configured to output a binary decision that indicates whether each region of interest is in the diseased portion of the structure of the individual.
  • the output 604 may be provided in several different forms.
  • the output 604 may include a listing of the regions of interest identified as part of a diseased portion of the body structure.
  • the output 604 may include a listing of the voxels included in the regions of interest identified as part of a diseased portion of the body structure.
  • the output 604 may provide the voxels associated with any identified ROIs in the form of a mask that can be overlayed on the original three-dimensional image.
  • a mask overlay may be used by a surgeon during a biopsy, may guide further imaging, etc.
  • the output of the classifier is a probability that the ROI is indicative of a diseased body structure.
  • the mask e.g., the overlay
  • the mask may be color-coded based on the probability. For example, a green overlay may indicate a ROI that was determined to be a false positive (e.g., low probability), a yellow overlay may indicate the ROI is questionable, and red overlay may indicate that the ROI was validated by the stage 2 classifier 220 (e.g., high probability).
  • the diseased structure identification system 200 is configured to generate a user interface (e.g., by the identification coordinator 210).
  • the identification coordinator 210 may send instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images being segmented as well as the mask and/or color overlay.
  • the identification coordinator 210 may include instructions to generate a local user interface on the computing hardware embodying the diseased structure identification system 200 (e.g., a desktop computer, a server, the imaging system 30, etc.).
  • the user interface may be configured to provide user interaction (e.g., in addition to viewing the results).
  • ROIs that were flagged, but ultimately labeled a false positive by the stage 2 classifier 220 may also be displayed in the user interface to provide additional data to the radiologist and receive feedback. Adjustments by the radiologist can improve results by providing additional training data that has been expertly labeled and the diseased structure identification system 200 was unable to predict.
  • the new training data may be stored in the diseased structure identification training system 250.
  • the regions of interest may be included in a number of visual representations.
  • One or more visual representations can be delivered to the user interface.
  • the regions of interest can be viewed as a color overlay on a screen, a virtual object in an augmented reality environment, an object on a holographic display, etc.
  • FIG. 12 is a flow of operations performed by the feature extractor 212 according to some embodiments.
  • FIG. 13 shows the flow of data as the feature extractor 212 processes a set of adjacent pixels according to some embodiments. The description of the feature extractor 212 will refer to both FIGS. 12 and 13 to provide details of the operations.
  • FIG. 12 shows flow of operations 420 for performing feature extraction according to some embodiments.
  • the flow of operations 420 includes obtaining a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels in operation 422. More than one three-dimensional image may be used.
  • the feature extraction algorithm may be performed on any number of related three-dimensional images (e.g., T2w, ADC, DWI modes from an MRI).
  • the feature extraction operations are different for each of the three-dimensional images.
  • the flow of operations 420 may continue with identifying a set of adjacent voxels, including a voxel for classification, of the plurality of voxels.
  • the diseased structure identification system 200 may perform vox el -wise classification, the set of adjacent voxels can provide the information used to make a determination if the voxel for classification is a portion of a diseased structure.
  • the voxel for classification 608 is shown near the center of the set of adjacent voxels 606.
  • the set of adjacent voxels 606 are provided to the feature extractor 212 and may be of any size. For example, in some embodiments, the set of adjacent voxels is square and 24 by 24.
  • the flow of operations 420 includes vectorizing a local region to generate a set of vectors for each local region of the adjacent set of voxels in operation 426.
  • the flow of operations may also include applying a feature extraction operation to each local region to generate a feature map, the feature extraction operation may be determined by a Saab transformation trained using local regions of three-dimensional training images in operation 428.
  • a local region 610 is shown.
  • the local region may be of any size, 3 by 3, for example.
  • the local region may be scanned across the set of adjacent voxels 606 and vectorized to generate a set of vectors 612 that may be used by the diseased structure identification training system 250 to generate the feature extraction operations (e.g., similar to the training performed by the body structure segmentation training system 150).
  • the set of vectors 612 may be combined with other vectors from other patients to determine the feature extraction operations of the Saab transformation.
  • Training of the Saab transformation may provide a set of anchor vectors to determine features from each of the local regions 610.
  • the feature extractor 212 may apply a feature extraction operator 230 to each local region of the set of adjacent voxels 606.
  • the feature extraction operator 230 may be represented by the anchor vectors arranged in a matrix as shown in equation 3.
  • the feature extraction operator 230 For each local region 610, the feature extraction operator 230 generates a pixel embedding of a number of features.
  • the pixel embeddings are arranged in a three-dimensional array (e.g., feature map 616) so that the x-y location of the pixel embedding corresponds with the center of the local region used to generate the pixel embedding and the depth of the three-dimensional array is the number of features of each pixel embedding.
  • a three-dimensional array e.g., feature map 616
  • the flow of operations 420 includes performing a pooling operation to each slice of the feature maps to generate a feature map of a lower resolution in operation 430.
  • the operation 430 may convert the feature map 616 (e.g., a feature map has similar resolution to the set of adjacent voxels) to the lower resolution feature map 618.
  • pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output.
  • the voxels may be combined by taking the maximum of the pixels, by taking the average of the pixels, or any other suitable function of values of multiple voxels.
  • two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution.
  • three-by-three patches of voxels are combined resulting in an output resolution that is a third of the input.
  • rectangular patches are combined in the pooling operation.
  • the flow of operations 420 includes applying a feature extraction operation to each local region of each feature slice of the lower resolution feature map 618 to generate a second feature map, the feature extraction operations determined by a Saab transformation trained using local patches of the same feature in operation 432.
  • a feature extraction operator 232 may be applied to local regions of each slice (e.g., channel) of the lower resolution feature map 618.
  • the feature extraction operator 232 may be represented by different anchor vectors (e.g., a different matrix for each of the slices).
  • each slice may provide a different number of features to the output pixel embedding (e.g., the number of anchor vectors may be different for each feature extraction operation).
  • a slice of the lower resolution feature map 618 after a feature extraction operation is applied to each local region, may generate a portion of the second feature map.
  • the first slice of the lower resolution feature map 618 may generate the portion 624a and the fourth portion of the feature map 618 may generate another portion of the second feature map 624b. All portions of the second feature map generated from slices of the lower resolution feature map 618 can be concatenated to form the second feature map 626.
  • Local regions of the lower resolution feature map 618 are also used to generate vector set 620 that can be used to train the feature extraction operator used for each of the slices (e.g., channels).
  • the set of vectors 620 generated may be used by the diseased structure identification training system 250 to generate the feature extraction operations (e.g., similar to the training performed by the body structure segmentation training system 150).
  • the set of vectors 620 may be combined with other vectors from other images (e.g., other individuals) to determine the feature extraction operations (e.g., anchor vectors) of the Saab transformation.
  • a feature extraction operation is determined for each slice (e.g., channel) of the lower resolution feature map 618 and the feature vectors 620 are grouped by which slice was used to generate them so that those vectors can be combined with vectors from the same slice of a different set of adjacent voxels 606 and/or a different three-dimensional image 602.
  • the flow of operations 420 includes generating spectra including a spectrum for each slice of the first map and each slice of the second feature map in operation 434. For example, performing a spectral PCA.
  • the PCA calculator 186 is used to generate global features for each slice of the feature maps 616 and 626.
  • the PCA calculator 186 may be used to perform eigenvalue decomposition on the slice and the spectrum may be used as features for the voxel-wise classification.
  • the flow of operations 420 includes flattening the first feature maps and the second feature map and concatenating the flattened feature maps with the spectra in operation 436.
  • the flattener 234 may vectorize the feature maps 616 and 626 so that the features may be used together by the voxel-wise classifier 220.
  • the set of adjacent voxels 606 selected for classification from the three-dimensional image is 24 by 24, the local region used to input to the feature extraction operator 230 is 3 by 3, feature extraction is performed without zero padding (and therefore local regions centered on the edge voxels are not used) resulting in a first feature map 616 that is 22 by 22 by 9, the size P of the pooling operation is two resulting a lower resolution feature map 618 that is 11 by 11 by 9, the feature extraction operator 232 uses the same size local region (e.g., 3 by 3) on all slices (e.g. channels) resulting in a second feature map 626 that is 9 by 9 by 56.
  • the local region used to input to the feature extraction operator 230 is 3 by 3
  • feature extraction is performed without zero padding (and therefore local regions centered on the edge voxels are not used) resulting in a first feature map 616 that is 22 by 22 by 9
  • the size P of the pooling operation is two resulting a lower resolution feature map 618 that is 11 by 11 by 9
  • the functionality described herein uses different parameters resulting in differently sized feature maps and a different number of features for voxel-wise classification. In some embodiments, more than two feature maps of differing resolutions are created (e.g., three or four).
  • FIG. 14 shows another block diagram of the imaging and disease identification system 12 according to some embodiments.
  • FIG. 14 shows some embodiments of an implementation of the diseased structure identification training system 250.
  • the diseased structure identification training system 250 may include a communications interface 252 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40.
  • the communications interface 252 may be used to communicate the trained machine learning models (e.g., the feature extraction models, and the classifiers) to the diseased structure identification system 200.
  • the communications interface 252 may be of the same or similar type of the communications interface 202 of the diseased structure identification system 200.
  • the diseased structure identification training system 250 may also include a processing circuit 254 having a processor 256 and memory 258.
  • the processor 256 may be configured to execute instructions contained on the memory 258.
  • the diseased structure identification training system 250 and the diseased structure identification system 200 are implemented on the same hardware (e.g., on hardware of the imaging system 30, on the same computer, etc.).
  • the processing circuit 254 may be the same processing circuit as the processing circuit 204, similarly the communications interface, memory and processors may be shared between the diseased structure identification system 200 and the diseased structure identification training system 250.
  • the diseased structure identification training system 250 may be implemented on separate hardware from the diseased structure identification system 200.
  • the processor 256 may be of similar type or of different type from the processor 206.
  • the processor 256 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components.
  • the processor 256 may be configured to execute computer code and/or instructions stored in the memory 258 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.).
  • the processor 256 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof.
  • graphics processing units GPUs
  • One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
  • the memory 258 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure.
  • the memory 258 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions.
  • the memory 258 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure.
  • the memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
  • the diseased structure identification training system 250 is to acquire several three-dimensional images of a portion of the body for training the feature extractor 212 (e.g., the feature extraction operations), the voxel classifier 222, and the stage 2 classifier 220 of the diseased structure identification system 200.
  • the diseased structure identification training system 250 applies unsupervised learning to identify features of interest for the voxel classifier 222 and the stage 2 classifier 220.
  • the feature extractor 212 may be trained to determine feature maps of various resolutions to be used by the voxel classifier 222 and the stage 2 classifier 220.
  • the diseased structure identification training system 250 applies a supervised learning to train the voxel classifier 222 to make a voxel-wise assessment (e.g., determination) of the probability that the voxel depicts a portion of a diseased body structure.
  • the diseased structure identification training system 250 applies a supervised learning to train the stage 2 classifier 220 to validate regions of interest (ROIs) indicated by several collocated voxels.
  • ROIs regions of interest
  • the region 2 classifier 220 may be trained to determine if a region of interest is a false positive or is in fact a diseased region of the body structure.
  • the training the feature extractor 212 is performed with unsupervised learning it does not require images for which diseased areas are previously labeled for training.
  • training using unsupervised learning allows unlabeled data to be used for training and greatly increases the available training data.
  • Images for which the diseased areas are labeled may be provided to train the voxel classifier 222 and the stage 2 classifier 220.
  • the segmented images are additionally used to train the feature extractor 212 (e.g., without using the labels) in some embodiments.
  • the diseased structure identification training system 250 applies unsupervised learning to train the voxel classifier 222 and the stage 2 classifier 220.
  • clustering algorithms may generate classification rules (e.g., boundaries, decisions, etc.) in the feature space that can be used to determine if a voxel or region of interest depicts a portion of a diseased body structure.
  • classification rules e.g., boundaries, decisions, etc.
  • the features of an area including a diseased body structure may naturally form a distinct cluster from those of healthy tissue.
  • the diseased structure identification training system 250 includes an identification training coordinator 260, the feature extractor 212, training data storage 266, a data selector 268, the voxel classifier 222, the stage 2 classifier 220, a loss calculator 272, a parameter tuner 274, and a Saab transformer 180.
  • several other components of the body structure segmentation system 100 or the diseased structure identification system 200 may also be included in the diseased structure identification training system 250 (e.g., the pooling calculator 116, the concatenator 124, etc.) and used to execute the feature extraction or classification during training of the voxel classifier 222, and the stage 2 classifier 220.
  • the identification training coordinator 260 may be configured to control the timing and flow of data through the other circuitry of the diseased structure identification training system 250. For example, the identification training coordinator 260 may cause the instructions or circuits to execute in a specific order to perform the function of the diseased structure identification training system 250. In some embodiments, the identification training coordinator 260 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input.
  • the training data storage 266 is configured to maintain (e.g., store, save, acquire, etc.) three-dimensional images of body structures that can be used to train the feature extractor 112, the voxel classifier 222, and the stage 2 classifier 220, or both.
  • Three-dimensional images for which diseased areas have not been labeled can be used to perform unsupervised learning (e.g., of the feature extractor or some classification algorithms).
  • Three-dimensional images of a body structure that has been previously labeled may be used to perform both unsupervised learning and supervised learning.
  • the data selector 268 selects the correct data to be used for a particular training algorithm.
  • the data selector 268 may select only labeled data when supervised learning is being performed to train the voxel classifier 222, and the stage 2 classifier 220.
  • the data selector 268 may also split data into training and validation sets.
  • the validation data may be used to stop algorithm training, choose a classification algorithm, or to determine appropriate values of hyperparameters (e.g., the number of features that should be used, maximum depth of a decision tree, regularization terms, etc.).
  • the diseased structure identification training system 250 includes the Saab transformer 180.
  • the Saab transformer 180 may be used to determine the feature extraction operations of the feature extractor 212.
  • the Saab transformer 180 may be used to determine the anchor vectors used in the feature extraction operations, the number of anchor vectors to use, and/or the matrix representation of the feature extraction operations (e.g., those used in feature extraction operation 230 and/or 232).
  • the Saab transformer 180 includes a sampler 182, a DC vector calculator 184, and a principal components analysis module (PC A) 186 to perform the operations of the Saab transform.
  • PC A principal components analysis module
  • the Saab transformer may receive a number of training samples from the training data storage 266.
  • the training samples may include three-dimensional images of a body structure (e.g., three-dimensional image 602).
  • the sampler 182 may determine local regions of the three-dimensional image to be used to train the feature extraction operation.
  • the sampler 182 may choose the local regions based on one or more criteria for filtering appropriate local regions.
  • the feature extraction operations may be applied to local regions that satisfy the same or similar criteria.
  • the sampler 182 may select data based on the slice (e.g., channel) of a feature map, and train different feature extraction operations for various slices.
  • the sampler may also disqualify regions, slices, or whole three-dimensional images based on one or more criteria. For example, a three-dimensional image may be disqualified from use in training if the signal to noise ratio is high, a slice may be disqualified from use in training if the image is subject to a blur, or any other undesirable image artifact may cause the image or slice to be disqualified.
  • the process for selecting anchor vectors during the Saab transformation is described in more detail in the description of the body structure segmentation training system 150.
  • multiple models may be trained based on criteria related to the three-dimensional image. For example, different feature extraction operations may be learned for three-dimensional images of differing noise levels.
  • features are selected by the feature selector 270 based on a Discriminant Feature Test (DFT) to quantify the discriminant power of features.
  • DFT Discriminant Feature Test
  • the minimum and maximum of the feature may be calculated, denoted by fmin and fmax, respectively.
  • the partition interval [fmin, fmax] may split into B bins in a uniform way. The edges of the bins can be used to define candidate thresholds for the evaluation of discrimination power of the feature.
  • the splitting quality of each threshold may be evaluated using the weighted entropy: where N + is the number of samples for which the feature is to the left of the threshold, N_ is the number of samples for which the feature is to the right of the threshold, tb are the candidate thresholds and choosing the minimum loss Lf tb over all threshold values.
  • the individual loss function e.g., to calculate Lf t and Lf t + ) is the entropy value, calculated by: , equation 8
  • the DFT may be calculated for each feature by the feature selector 270 during training and a number of features can be kept (e.g., those with the lowest loss). For example, the 1000 features which demonstrate the greatest discrimination potential may be kept.
  • the loss is plotted in ascending order and the elbow point in the plot is used to determine the number of features kept to train the voxel classifier 222 and/or the stage 2 classifier 220.
  • training of a classifier (e.g., the voxel classifier 222 and the stage 2 classifier 220) is performed iteratively.
  • the classifier may determine a probability that the features represent a diseased portion of the body structure.
  • the loss calculator 272 may compare the probability calculated by the classifier to the label for the voxel (e.g., the voxel from the set of adjacent voxels) that is being classified.
  • the loss calculator 272 may compute the sum of the loss by comparing several pixel embeddings to the respective labels to compute a loss function. Based on the loss function the parameter tuner 274 may update the classifier (e.g., parameters of the classifier, rules within the classifier, etc.) to improve the loss function with respect to the training data.
  • the classifiers 222 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
  • the training data includes numerous negative areas (e.g., areas that do not depict a diseased portion) of a body structure, and a small number of regions for which the disease is confirmed (positive areas).
  • training is performed using a two-step process.
  • the classifier may be trained using random region sampling from the training data. However, areas with positive areas may be oversampled. For example, by sampling with more overlap and/or by causing the random sampler to choose positive samples more often.
  • the trained classifier is used to determine the difficulty of classifying a particular training sample. Training samples can then be chosen from across the various difficulty bins in order to cause the training data to approximate a desired distribution (e.g., equally training on difficult, medium, and easy to identify training samples).
  • PI-CAI includes bpMRI data of a cohort of 1,500 patients.
  • Training data includes MRI scans across multiple clinics, such as the Radboud University Medical Center (RUMC), University Medical Centet Groningen (UMCG) and Ziekenhuis Groep Twente (ZGT). Scans from those medical centers are included both in training and testing.
  • MRI scans from the Norwegian University of Science and Technology (NTNU) is also included in the testing cohort, as an unseen institute data to training.
  • the vendors of MRI scanners were Siemens Healthineers and Philips Medical Systems.
  • the training data was divided into 1075 samples of benign or indolent PCa and 425 clinically significant prostate cancer (csPCa) cases.
  • Table 5 Comparison to deep learning models based on 1000 patients from the PI-CAI challenge
  • RadHop including only stage-1 predictions using patch size of 32x32 i.e. no refinement by stage- 2
  • DCNN deep convolutional neural network
  • U-Net deep convolutional neural network
  • AP lesion-level precision
  • RADHop has a better performance than every other method under comparison, excect U-Net that has a slightly better performance.
  • RADHop achieves a higher precision than other methods but there is still some gap with the U-Net related ones.
  • stage 2 classifier to refine the heatmap predictions of stage- 1 classification may be further improved.
  • stage-2 Given the limited access to the hidden testing cohort of PI- CAI challenge, a short performance comparison was performed to demonstrate the effectiveness of stage-2 and compare stage-1 under different patch unit sizes. Table 6 shows some experiments carried out, trying to improve stage- 1 output. It is possible to see the effectiveness of using smaller patch size, since it can capture, more easily, parts of the lesion. Also, a larger patch size can induce noise, thereby affecting the precision performance, especially for medium-sized and smaller lesion sizes. Stage-2 effectiveness and potentials are also demonstrated. In particular, AP (per lesion metric) improves from 0.356 to 0.374. This is because in stage- 2 it is possible to reduce the probability of some false positives by reclassification using the anomaly score features.
  • Instructions, modules, portions of memory, etc. described as configured to perform a function may include embodiments for which the module is configured to cause the performance of the function (or is causing the performance of the function).
  • instructions, modules, portions of memory, etc. described as configured to cause the performance of a function may include embodiments for which the module is configured to perform the function (or is performing the function).
  • references to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • General Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Mathematical Physics (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Medical Informatics (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Databases & Information Systems (AREA)
  • Multimedia (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

A method including unsupervised feature extraction is performed to determine feature maps of a various resolutions representing a three-dimensional image of a body structure. Voxel-wise classification is performed to determine if a voxel depicts a portion of a body structure, a specific subregion or zone of the body structure, and/or a diseased portion of the body structure. The method provides similar or improved performance to benchmark methods at a reduced computational and memory burden. Classifications may be validated by a second stage classifier that uses additional features.

Description

BODY STRUCTURE SEGMENTATION USING SUPERVISED AND UNSUPERVISED LEARNING
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63/600,352 filed on November 17, 2023, and U.S. Provisional Application No. 63/600,360 filed on November 17, 2023, the entirety of each of which is incorporated by reference herein.
BACKGROUND
[0002] Imaging systems play a crucial role in modem medicine where they are essential for diagnosing and evaluating diseases. These advanced technologies include modalities such as X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound. Each of these systems captures detailed images of the body's internal structures, allowing radiologists to visualize organs, tissues, and any abnormalities that may indicate disease.
[0003] Prostate cancer (PCa), for example, is reported as the second most frequent cancer among men in 2020, with an estimated almost 1.4 million new cases and 375,000 deaths worldwide. Compared to 12-core transrectal ultrasound-guided biopsy (TRUSGB), multiparametric magnetic resonance imaging (mpMRI) has been reported to reduce the detection of insignificant prostate cancer. It has become an important imaging method to detect clinically significant prostate cancer and guide biopsies.
[0004] Segmentation of a body structure (e.g., organ, prostate, etc.) is an important step in diagnosis and treatment planning. Segmentation allows for determining prostate boundaries for radiotherapy. Segmentation can also allow the calculation of volume and other key metrics necessary to track disease progression.
[0005] Radiologists also use imaging systems to detect tumors or the presence of other diseases. MRI imaging systems can provide detailed images of soft tissues, which is particularly valuable for identifying cancerous growths. CT scans offer high-resolution cross- sectional images that help in evaluating complex conditions, such as trauma injuries or vascular diseases. By leveraging these advanced imaging technologies, radiologists can detect not only the presence of disease but also its size, shape, and location, which are vital for determining the appropriate course of treatment.
[0006] The high level of expertise required to read medical imagery leads to a relatively low agreement rate among radiologists. For example, different years of experience and training can affect the calculation of various metrics related to the diseased structure and ultimately the diagnosis. Detection and segmentation of lesions affects the accuracy of targeted biopsies and the ability to minimize the radiation dose on noncancerous tissues.
SUMMARY OF THE INVENTION
[0007] Some embodiments of the present disclosure relate to a method. The method includes obtaining, by one or more processors, a slice of a three-dimensional image. The three- dimensional image depicts a structure of an individual. The slice includes a plurality of pixels each corresponding to a different location of the three-dimensional image. The method also includes executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions. Each feature map of the plurality of feature maps includes a plurality of pixel embeddings each corresponding to a different location of the slice. The method also includes executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask. The first mask includes values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual. The method also includes concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution is higher than the first resolution. The method also includes executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual. The method also includes generating, by the one or more processors, a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
[0008] In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution. In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation. In some embodiments, the method also includes training, by the one or more processors, the feature extraction model using an unsupervised learning technique. [0009] In some embodiments, executing the feature extraction model includes executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings. In some embodiments, executing the feature extraction model includes performing, by the one or more processors, principal components analysis on sets of adjacent pixels of a plurality of locations of the three-dimensional image. In some embodiments, executing the feature extraction model includes executing, by the one or more processors, a transformation on a set of adjacent pixels encompassing a pixel of the plurality of pixels to generate the pixel embedding associated with the pixel.
[0010] In some embodiments, the feature extraction model includes pooling, by the one or more processors, adjacent pixel embeddings to generate a lower resolution feature map. In some embodiments, the method also includes classifying, by the one or more processors, each concatenated pixel embedding based on whether the probability for the concatenated pixel embedding exceeds a threshold. In some embodiments, obtaining the slice of a three- dimensional image includes receiving, by the one or more processors, the three-dimensional image of the individual, obtaining the slice of a three-dimensional image includes segmenting, by the one or more processors, the three-dimensional image into a plurality of slices corresponding to different locations of the three-dimensional image.
[0011] In some embodiments, obtaining the slice of the three-dimensional image includes obtaining the three-dimensional image depicting a prostate of the individual. Executing the second machine learning model to classify each concatenated pixel embedding includes executing, by the one or more processors, the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual. Generating the second mask includes generating, by the one or more processors, the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
[0012] Some embodiments of the present disclosure relate to a system. The system includes one or more processors configured by computer-readable instructions to perform operations. The operations include obtaining a slice of a three-dimensional image, the three-dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location of the three-dimensional image. The operations include executing a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice. The operations also include executing a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual. The method also includes concatenating each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution. The method also includes executes a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual. The method also includes generating a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
[0013] In some embodiments, the one or more processors are further configured to scale the first mask from the first resolution to the second resolution. In some embodiments, the one or more processors are configured to scale the first mask from the first resolution to the second resolution using trilinear interpolation. In some embodiments, the one or more processors are further configured to train the feature extraction model using an unsupervised learning technique. In some embodiments, the one or more processors are further configured to execute the feature extraction model to transform a set of adjacent pixels including a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings.
[0014] In some embodiments the one or more processors are configured to obtain the slice of the three-dimensional image by obtaining the three-dimensional image depicting a prostate of the individual. The one or more processors are configured to execute the second machine learning model to classify each concatenated pixel embedding by executing the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual. The one or more processors are configured to generate the second mask by generating the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
[0015] Some embodiments of the present disclosure relate to a method. The method includes obtaining, by one or more processors, a slice of a three-dimensional image, the three- dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location of the three-dimensional image. The method also includes executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice. The method also includes executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual. The method also includes concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution. The method also includes executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual. The method also includes presenting, by the one or more processors, a visual representation of the classifications of the concatenated pixel embeddings.
[0016] In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution. In some embodiments, the method also includes scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation.
[0017] Some embodiments of the present disclosure relate to a method. The method includes obtaining, by one or more processors, a three-dimensional image. The three-dimensional image depicting a structure of an individual and include a plurality of voxels. The method includes performing for each voxel of a plurality of voxels; identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual. The method also includes determining, by the one or more processors, one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold. The method includes for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
[0018] In some embodiments, the method also includes training, by the one or more processors, the feature extraction model using unsupervised learning. In some embodiments, executing the feature extraction model includes executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings. In some embodiments, executing the feature extraction model includes performing, by the one or more processors, spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
[0019] In some embodiments, executing the feature extraction model further includes combining, by the one or more processors, the features of the plurality of feature maps of different resolutions. In some embodiments, determining the one or more regions of interest of the three-dimensional image includes identifying, by the one or more processors, clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
[0020] In some embodiments, determining the one or more regions of interest of the three- dimensional image includes identifying, by the one or more processors, a first set of clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a first threshold. Determining the one or more regions of interest includes identifying, by the one or more processors, a second set of clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a second threshold.
[0021] In some embodiments, the method includes presenting, by the one or more processors, a first visual representation of the first set of clusters in a first color and a second visual representation of the second set of clusters in a second color. In some embodiments, the method includes concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps. Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps as input.
[0022] In some embodiments, the method also includes obtaining, by the one or more processors, anomaly features extracted from the three-dimensional image and concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps and the anomaly features Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps and the anomaly features as input.
[0023] In some embodiments, the three-dimensional image is of a plurality of three- dimensional images of different imaging technologies, determining the one or more regions of interest of the three-dimensional image includes determining, by the one or more processors. The one or more regions of interest based feature maps extracted from each of the plurality of three-dimensional images.
[0024] In some embodiments, obtaining the three-dimensional image includes obtaining, by the one or more processors, the three-dimensional image depicting a prostate of the individual. Executing the first machine learning model includes executing, by the one or more processors, the first machine learning model to generate the value indicating the probability that the voxel depicts a cancerous portion of a prostate of the individual. Executing the second machine learning model includes executing, by the one or more processors, the second machine learning model to generate the classification indicating whether the region of interest is the cancerous portion of the prostate of the individual.
[0025] Some embodiments of the present disclosure relate to a system. The system includes one or more processors configured by computer-readable instructions to perform operations. The operations include obtaining a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels. The method includes for each voxel of the plurality of voxels: identifying a set of adjacent voxels including the voxel of the plurality of voxels; executing a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and execute a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual. The method also includes determining one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold. The method also includes for each of the one or more regions of interest, executing a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
[0026] In some embodiments, the one or more processors are further configured to train the feature extraction model using unsupervised learning. In some embodiments, the one or more processors are configured to execute the feature extraction model by performing spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
[0027] In some embodiments, the one or more processors are configured to execute the feature extraction model further by combining the features of the plurality of feature maps of different resolutions. In some embodiments, the one or more processors are configured to determine the one or more regions of interest of the three-dimensional image by identifying clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
[0028] Some embodiments of the present disclosure relate to a method. The method includes obtaining, by one or more processors, a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels. The method also includes, for each voxel of the plurality of voxels: identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual. The method also includes determining, by the one or more processors, one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold. The method also includes for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and anomaly features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
[0029] In some embodiments, the method also includes training, by the one or more processors, the feature extraction model using unsupervised learning. In some embodiments, executing the feature extraction model includes executing, by the one or more processors. The feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings.
BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Various objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the detailed description taken in conjunction with the accompanying drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and/or structurally similar elements.
[0031] FIG. l is a block diagram of an imaging and structure segmentation system showing the body structure segmentation system, according to some implementations;
[0032] FIG. 2 is flow of operations for segmenting a body structure, according to some implementations;
[0033] FIG. 3 is a flow diagram illustrating the flow of data as it traverses the body structure segmentation system of FIG. 1, according to some embodiments;
[0034] FIG. 4 is flow of operations performed by a feature extractor of FIG. 3, according to some implementations; [0035] FIG. 5 is an illustration of the operations of a feature extractor of FIG. 3, according to some implementations;
[0036] FIG. 6 is another block diagram of the imaging and structure segmentation system of FIG. 1 showing the body structure segmentation training system, according to some implementations;
[0037] FIG. 7 is an illustration of the operations to determine the feature extraction operators, according to some implementations;
[0038] FIG. 8 is an illustration of the operations to train a segmentation classifier, according to some implementations;
[0039] FIG. 9 is a block diagram of an imaging and identification system showing the diseased structure identification system, according to some implementations;
[0040] FIG. 10 is flow of operations for identifying a diseased portion of a body structure, according to some implementations;
[0041] FIG. 11 is a flow diagram illustrating the flow of data as it traverses the diseased structure identification system of FIG. 9, according to some implementations;
[0042] FIG. 12 is flow of operations performed by the feature extractor of FIG. 10, according to some implementations;
[0043] FIG. 13 is an illustration of the operations of a feature extractor of FIG. 10, according to some implementations; and
[0044] FIG. 14 is another block diagram of an imaging and identification system of FIG. 9 showing the diseased structure identification training system, according to some implementations.
DETAILED DESCRIPTION
[0045] In the following description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the description, drawings, and claims are not limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented here. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are explicitly contemplated and made part of this disclosure.
Overview
[0046] Computing systems that perform conventional methods of body structure segmentation, such as using deformable models and graph cut optimization have several disadvantages. Deformable models, for example, employ flexible structures that adapt their shape to fit the contours of objects in images. The models iteratively adjust their parameters to minimize an objective function that is designed to ultimately converge on the desired shape or boundary. Deformable models often used in image segmentation and shape analysis are highly sensitive to initial conditions because their optimization relies on the initial settings to guide their deformation process. If the initial position is close to the desired contour, the model can effectively converge to the target shape. However, if the starting point is far from the desired boundary, the model may converge to a local minimum that does not accurately represent the true shape. This sensitivity arises from the complexity (e.g., non-convexity) of the optimization function, where small changes in initial conditions can lead to significantly different outcomes. Deformable models may use multiple initial conditions (e.g., determined randomly) to increase the chance of finding a global optimum, leading to significant computational expense.
[0047] In graph cut optimization, the image is modeled as a graph and formulates a minimum cut problem to optimally partition the foreground (e.g. the body structure of interest) from the background. Graph cut optimization can have sensitivity to noise and may lead to oversegmentation or under-segmentation, where the model fails to accurately delineate the desired objects. Additionally, the computational cost of graph cut methods can become prohibitive for high-resolution images or large datasets. The algorithm's performance can also be limited in scenarios with occlusions or complex object shapes, leading to inaccurate segmentations. These limitations can lead to inaccurate results at high computational expense.
[0048] Convolutional neural networks and other deep neural networks have been used to perform image segmentation. Depending on the depth of the convolutional network and the number of kernels used in each layer, performing segmentation may still require a significant computational burden in the forward path (e.g., during inference). In addition, training the convolutional network to perform segmentation requires large sets of training data that has been labeled (e.g., the segmentation process was performed by another system or manually). Deep neural networks rely on long training times required to make hundreds of passes through the data set while adjusting the network weights based on backpropagated gradients. In addition, training may be required separately for each type of image to be segmented (e.g., each body structure). The training data may be insufficient to provide accurate results or the parameters of the training process may require tuning. In addition, deep neural networks may be viewed as a “black-box” and may not provide useful insight related to the reason for results. For example, if the results were inaccurate, it may be difficult to determine the reason for the error.
[0049] The present disclosure improves the technological field of image segmentation and more specifically segmentation of body structures by reducing the number of computations performed during the segmentation process and during the training process. In addition, the present disclosure may require significantly less training data to be collected. In some embodiments of the present disclosure, a feature extraction method is performed on a slice of a three-dimensional image to generate a plurality of feature maps of different resolutions. The features are used to indicate (e.g., to classify) structures of an individual that are depicted by the voxels of the image. For example, the features are used to indicate that the voxels are associated with a specific body structure and/or a region within the body structure. Classification of a region of the slice is performed at a lower resolution and then used during classification at a higher resolution. The hierarchical approach provides two improvements: (i) less training data may be required to train a classifier at a lower resolution, because this information is carried to the higher resolutions the higher resolution data will also require fewer training examples, and (ii) the hierarchy of resolutions that are provided as output allow a user to quickly identify discordance of results and take appropriate actions in real-time (e.g., rescanning the patient). In some embodiments, the feature extraction method is unsupervised leading to a further reduction in the amount of training data required and allowing for feature extraction using unlabeled training data. In addition, in some embodiments, the feature extraction technique provides for a direct computation of the weights from the training data without performing iterations (e.g., gradient descent, etc.) and may greatly reduce the segmentation computational requirements. [0050] The feature extraction method performed relies on forming a sum of a product (e.g., a matrix multiplication) of a patch (e.g., a local region) of the slice of the three-dimensional image or of a feature map and a set of coefficients. Such operations can be performed quickly in modem computer architectures and do not require the iteration of deformable models or graph cut optimization. In addition, the feature extraction technique does not require the large number of kernels used by a convolutional neural network, further reducing the computational complexity of performing segmentation. Advantageously, segmentation results may be available for the radiologist or technician in real-time allowing for rescans or adjustments to focus on a region of interest. In addition, computational resource savings leads to more efficient processing and energy savings. Accuracy is improved over some segmentation methods that have more than 50 times the memory requirements and more than 150 times the computational requirements of the methods disclosed herein.
[0051] Segmenting a diseased portion of a body structure from healthy tissue is also an important task performed in the clinical setting. The present disclosure provides systems and methods for identifying diseased portions of a body structure that have advantages (e.g., similar to those already discussed) over the traditional automated techniques and/or neural network techniques. Identification of diseased areas of a body structure (e.g., cancerous areas, kidney stones, etc.) can use similar feature extraction techniques that rely on less training data and have computational advantages in their calculation when compared to deformable models, graph cut optimization, and/or convolutional neural networks. An initial classification of voxels using the feature extraction technique may cause various regions of interest to be identified and flagged for further investigation. For example, regions may be identified for biopsy, used to automatically trigger additionally imagery to be performed, and/or to focus the attention of the radiologist. Advantageously, performing a coarse screening to find a few potentially diseased regions and then calculating visually-inspired features (e.g., first order statistics, a gray level run length matrix, a gray level dependence matrix, etc.) can significantly reduce the number of false positives by combining more traditional visually-inspired methods with machine learning, and significantly reduce the number of calculations that are performed by limiting the extraction of visually-inspired features to only a few regions of interest.
Body Structure Segmentation System
[0052] FIG. 1 is a block diagram of an imaging and structure segmentation system 10 configured to obtain medical imagery of a body structure and segment (e.g., identify, highlight, etc.) the structures of interested according to some embodiments. The imaging and structure segmentation system 10 may segment specific bone and/or joint structures, soft tissue (e.g., muscles, tendons, and ligaments), organs (e.g., the heart, lungs, kidneys, prostate, etc.), the nervous system, digestive tractor, or any other body structure for which segmenting a particular structure of interest from other tissue is useful. The imaging and structure segmentation system 10 includes one or more client devices 20, one or more imaging systems 30, a body structure segmentation training system 150, and a body structure segmentation system 100 shown to be communicably coupled via a network 40. The network 40 can include routers, switches, antennas, computers, and any other hardware required to communicate information between the components of the imaging and structure segmentation system 10 (e.g., from the imaging system 30 to the body structure segmentation system 100). A portion of the network 40 can be wireless and/or a portion of the network 40 can be wired. The network 40 can include one or more networks with routers to facilitate data transfer between the different networks.
[0053] The imaging and structure segmentation system 10 is shown to be distributed across several devices (e.g., networked computers). It is contemplated that in some embodiments, the subsystems, components, and/or functionality of the imaging and structure segmentation system 10 may be distributed on different devices. For example, the body structure segmentation system 100 may be configured within the imaging system 30 using the processors of the imaging system 30 to execute. In some embodiments, the body structure segmentation system 100 and the body structure segmentation training system 150 are configured within the same computer hardware (e.g., within a server local to the imaging system, within a remote server, or on a node within a cloud computing architecture).
[0054] The body structure segmentation system 100 may include a communications interface 102 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40. The body structure segmentation system 100 may also include a processing circuit 104 having a processor 106 and memory 108. For example, the processor 106 may be configured to execute instructions contained on the memory 108.
[0055] The processor 106 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components. The processor 106 may be configured to execute computer code and/or instructions stored in the memory 108 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.). The processor 106 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof. One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
[0056] The memory 108 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure. The memory 108 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions. The memory 108 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure. The memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
[0057] According to some embodiments, the body structure segmentation system 100 is to acquire a three-dimensional image of a portion of the body, perform a sequence of feature extraction operations and pooling operations on one or more slices of the three-dimensional image to obtain feature maps of various resolutions and varying depths (e.g., number of features). In some embodiments, information of the feature maps is combined during a hierarchical classification procedure, where a lower resolution feature map is classified to create an array of probabilities corresponding to the probability that an area of the three- dimensional image includes a portion of the body structure. Interpolation can be performed on the probability array to increase its resolution and allow it to be included within the feature map of a higher resolution for classification at the higher resolution. The classification procedure may be repeated until classification is performed at the original resolution of the input three-dimensional image. The body structure segmentation system 100 may be configured to segment one or more body structures of interest including, but not limited to: glands including the prostate, thyroid, etc.; organs including the heart, the liver, etc.; bones; joints; muscles; tumors; and other body structures for which segmentation may be useful.
[0058] In some embodiments, the body structure segmentation system 100 includes a segmentation coordinator 110, a feature extractor 112, feature extraction model storage 114, a pooling calculator 116, a position feature generator 118, classifiers 120, an upsampler 122, a concatenator 124, and a mask generator 126, the functionality and features of each will be described in more detail as related to FIGS. 2-5 herein. The segmentation coordinator 110 may be configured to control the timing and flow of data through the other circuitry of the body structure segmentation system 100. For example, the segmentation coordinator 110 may cause the instructions or circuits to execute in a specific order to perform the functions of the body structure segmentation system 100. In some embodiments, the segmentation coordinator 110 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input. For example, the segmentation coordinator 110 may cause the instructions and/or circuits of the body structure segmentation system 100 to perform a flow of operations 300 in FIG. 2 to segment a body structure.
[0059] With reference to FIG. 2, the flow of operations 300 includes obtaining (e.g., acquiring, receiving, etc.) a slice of a three-dimensional image, the three-dimensional image depicting a structure of an individual and the slice including a plurality of pixels each corresponding to a different location (e.g., x-y location) of the three-dimensional image in operation 302, according to some embodiments. For example, the imaging system 30 may communicate the three-dimensional image over the network 40 to begin processing (e.g., segmentation). In some embodiments, the segmentation coordinator 110 requests (e.g., from a database, from imaging system 30, etc.) one or more three-dimensional images to process (e.g., to perform batch processing). In some embodiments, the imaging system 30 may select (e.g., based on a user input, such as based on an input from a radiologist, lab technician, etc.) a three-dimensional image for segmentation to assist in the evaluation of the imagery. The segmentation coordinator 110 may select a slice (e.g., cross section) of the three-dimensional image to begin processing. The segmentation coordinator 110 may process the subsequent slices in depth-wise order using similar techniques.
[0060] In some embodiments, the flow of operations 300 includes executing (e.g., running, performing, etc.) a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps including a plurality of pixel embeddings each corresponding to a different location of the slice in operation 304. The feature extraction model, for example, may be executed by the feature extractor 112 using models stored in the feature extraction model storage 114. The feature extraction model may depend on the feature map being processed (e.g., the stage of the feature extraction process), the body structure being segmented, the resolution of the image and/or feature map, and/or any other attribute of the original three- dimensional image or of the intermediate feature maps. A voxel may refer to a three- dimensional pixel, and in some embodiments, voxel and pixel are used interchangeably.
[0061] The feature extractor 112 may be configured to generate a set of features from a local region (e.g., patch, area, etc.) of the slice. The slice of the image, for example, may be represented by an array of numeric values corresponding to the grey scale intensity of the image. In some embodiments, the feature extractor 112 obtains a feature extraction operator (e.g., a function, a linear transformation, a matrix, etc.) from the feature extraction model storage 114 and performs the feature extraction operation on a vectorized version of the local region. For each local region of the slice, the feature extractor 112 may output a set of features (e.g., an array). The local region processed by the feature extractor 112 is then shifted (e.g., moved by a voxel in one direction) and the operation is repeated to obtain a set of features for the current local region of the slice. The feature extractor 112 may be configured to scan the processed local region over the slice and organize the set of features output for each local region into a three-dimensional array for the slice. The set of features may be stored at indices for the first and second dimension equal to the indices of a voxel (e.g., the center voxel), of the local region and each feature of the set may be indexed into the third dimension (e.g., form a pixel embedding at that location). In some embodiments, the three-dimensional array for the slice has height and width equal to that of the slice and depth equal to the number of features output by the feature extraction operator. A feature at all height and width indices for a single depth (e.g., one feature index) may be referred to as a feature slice in some embodiments.
[0062] In some embodiments, the feature extractor 112 may perform the same or a similar feature extraction technique on a feature map (e.g., the output of the feature extraction may be further processed by performing feature extraction on one or more feature maps). The feature extractor 112 may obtain a different feature extraction operation from the feature extraction model storage 114 for each feature map. The feature extraction model storage 114 may be configured to maintain (e.g., store, save, etc.) the feature extraction models and may also store the feature map that a specific feature extraction model should be used on. For example, during training the body structure segmentation training system may provide a map (e.g., structure, object, dictionary, or other suitable storage elements) that contains for each feature map (including the original image) a feature extraction operation.
[0063] To generate feature maps of various resolutions a pooling operation may be performed. For example, pooling may be performed by pooling calculator 116. In some embodiments, pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output. For example, the voxels may be combined by taking the maximum of the pixels, by taking the average of the pixels, or any other suitable function of values of multiple voxels. In some embodiments, two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution. In some embodiments, three- by-three patches of voxels are combined resulting in an output resolution that is a third of the input. In some embodiments, rectangular patches are combined in the pooling operation. It is contemplated that the pooling calculator 116 may perform different pooling operations and/or use different sized patches based on the current feature map that is being processed or the original input image. For example, three-by-three voxel patches may be used in a first stage of feature extraction and two-by-two voxels may be used in a second stage of feature extraction.
[0064] In some embodiments, the flow of operations 300 includes executing a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask including values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual in operation 306. The first feature map, for example, may be a low-resolution feature mask and the machine learning model may be a classifier of the classifiers 120. In some embodiments, positional features are concatenated with the feature map (e.g., using the position feature generator 118 and the concatenator 124). In some embodiments, the probabilities of the first mask may be compared to a threshold to generate a coarse binary mask indicating regions that may depict a portion of the body structure.
[0065] The classifiers 120 may be configured to store one or more machine learning models. The machine learning models may classify an area of the three-dimensional image as representing at least a portion of the body structure of interest based on the features of the feature map. For example, the feature map may represent a slice of the three-dimensional image and machine learning model (e.g., a classifier) may generate a probability that the features at a given x-y location represent at least a portion of the body structure of interest.
[0066] In some embodiments, a single slice is used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of that slice (respective to locations of the feature map) represent at least a portion of the body structure. In some embodiments, multiple slices are used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of each of the multiple slices (respective to locations of the feature map) represent at least a portion of the body structure. In some embodiments, multiple slices are used to generate the feature map and a machine learning model of the classifiers 120 is used to generate a probability that the voxels of one of the multiple slices (e.g., the center slice) represent at least a portion of the body structure.
[0067] In some embodiments, the machine learning model used by the classifiers 120 depends on the feature map being processed. For example, three slices may be used to generate feature vectors that are in turn used as input to a first model of classifiers 120 to classify each of the three slices in a first stage of classification and a second model of classifiers 120 may be used to classify only the center slice at a second stage (e.g., the last stage) of classification. By utilizing more than one slice to generate some of the features, information from adjacent slices is incorporated into the classification; however, the classifiers may still be configured to generate a single mask representing the segmentation of the body structure at that slice.
[0068] The classifiers 120 may be configured to use one or more types (e.g., architectures, classes, etc.) of machine learning models to generate the classification. For example, the machine learning model may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination types of classification models. The type of classification architecture may depend on the feature map being processed (e.g., the stage of the classification process), the body structure being segmented, the resolution of the image and/or feature map, and/or any other attribute of the original three-dimensional image or of the intermediate feature maps.
[0069] The position feature generator 118 may be configured to generate features for a feature map that correspond to a location relative to the original three-dimensional model and/or the lower resolution feature maps. For example, the position feature generator 118 may produce a three-dimensional array that includes the x, y, and z location of a voxel corresponding to the one set of features (e.g., x and y index of the feature map). In some embodiments, the x and y locations of the positional features may relate to the x and y location in the feature map because at lower resolutions the features at a specific x-y location correspond to several voxels. Consider a feature map that has been reduced to an 8 by 8 resolution, each of the x-y indexes including a number of features in the 3rd dimension. The positional feature generator may produce the following three dimensional array:
Figure imgf000022_0001
indicating the x-y location of the feature map and the z location of the slice (e.g., 10). Adding positional features may be useful as they can cause the sensitivity to increase near the center of the image (e.g., the majority of the images will have the structure of interest near the center of the image) and/or allow the classifiers 120 to use the positional information in other ways. Positional information can be encoded in a variety of ways in addition to using the x, y location of the feature map and the z location of the original slice. The positions can relate to the original region used to generate the feature map. For example, if the original resolution of a slice was 512 by 512, the first position of the positional three-dimensional array may be 32.5 representing the average index of the 64 by 64 patch of voxels that are represented by the position in the lower resolution feature map. Alternatively or additionally, positional features may be added based on their position relative to a specific reference point in the image (e.g., the center). For example, if the distance to the center of the three-dimensional image is used to add positional information by the position feature generator 118, fewer than three-dimensions can be used to encode the positional information (e.g., distance can be represented by a single number). Alternatively or additionally, other functions of the index may be used to encode positional information. For example, sines and/or cosines of the positional index can be used, potentially leading to more than three positional features (e.g., a three-dimensional array with more than three indices in the third dimension).
[0070] In some embodiments, the concatenator 124 is configured to add features to a feature map. The new features may be conformant to the current feature map. For example, a feature map with a resolution of N by N may require the additional features to also have a resolution of N by N before concatenation. The positional features from the positional feature generator may be concatenated with the feature map using the concatenator 124. The concatenator 124 may add the positional features described herein along the feature dimension (e.g., third dimension). For example, if the feature map is 8 by 8 by 45 (e.g., a resolution of 8 by 8 and a feature depth of 45), 8 by 8 positional features can be added to the depth of the feature map.
[0071] In some embodiments, the flow of operations 300 includes concatenating each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution in operation 308. The second resolution may be greater than the first resolution and it may be necessary to scale the first mask (e.g., the lower resolution classification result) to the resolution of the second feature map using interpolation.
[0072] In some embodiments, the body structure segmentation system 100 includes the upsampler 122 to increase the resolution (e.g., scale, upsample, etc.) of a classification result (e.g., a probability mask). The upsampler 122 may include various interpolation techniques including trilinear interpolation. Trilinear interpolation can estimate values within a three- dimensional grid by linearly interpolating along each of the three axes (x, y, and z). Trilinear interpolation involves using eight surrounding data points in a cube to compute the interpolated value based on their distances to the target point. Similarly, the upsampler 122 may perform bilinear interpolation (e.g., if the input is only two-dimensional or if it is desired that interpolation is performed within a single slice). In some embodiments, the upsampler 122 includes additional interpolation techniques including: tetrahedral interpolation, spline interpolation, radial basis function interpolation, etc. that can be used to increase the resolution of the classification result to conform with a feature map of a higher resolution.
[0073] In some embodiments, the flow of operations 300 includes executing a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual in operation 310. The classifiers 120 may include a second classifier that operates at the higher (e.g., second resolution) and classifies each pixel embedding (e.g., at an x-y location) of the higher resolution feature map. In some embodiments, the second classifier of the classifiers 120 has been trained to use the classification result (e.g., the probability mask) of the lower resolution classifier, for example, scaled to the higher resolution using upsampler 122 in the previous operation.
[0074] Operations 308-310 can be repeated a number of times for each of the various resolution feature maps generated by the feature extractor 112 in operation 304. For example, the feature extractor 112 may produce feature maps of at 4 different resolutions. Operation 306 can be performed at the lowest resolution (e.g., fourth resolution), the scaled results concatenated with the third resolution in operation 308, and the third resolution classified in operation 310. The results of the classification in of the third resolution can be scaled by upsampler 122 to a second resolution by repeating operation 308 and concatenated to a feature map at the second resolution before classification in the second pass of operation 310 to produce a second resolution probability map. Repeating the steps 308 and 310 one more time will result in a probability mask at the original resolution of the three-dimensional image.
[0075] In some embodiments, the flow of operations 300 includes generating a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings in operation 312. The second mask, for example, may be a probability mask as described with reference to operation 306 or the second mask may be a binary mask. A binary mask may represent a final classification of a pixel (e.g., based on the pixel embedding related to that pixel at a specific x-y location) indicating if the pixel is part of the structure being segmented. The operation 312 may be performed by the mask generator 126.
[0076] In some embodiments, the classifiers 120 are trained to correct errors in the lower resolution probability mask using a regression model (e.g., rather than a classification model). The ground truth of the labeled training data may be subtracted from the output of a classifier 120 during the training process to allow the model to focus on the difficult to segment boundary regions of the body structure. The probabilities (e.g., the mask) may be upsampled to the higher resolution where the next of the classifiers 120 is trained to correct errors in the previous probabilities. This process may be repeated for any number of layers the body structure segmentation process.
[0077] The mask generator 126 may compare a probability mask to a threshold to determine a binary mask. For example, each pixel embedding that scores a probability that it depicts at least a portion of the structure greater than a threshold may indicate a location depicting a portion of the body structure. In some embodiments, the mask generator 126 includes logic in addition to comparing the probability to a threshold. For example, the mask generator 126 may be configured to include a type of hysteresis to cause the segmentation to generally form larger grouping of pixels and eliminate noise by performing the comparison to a second, lower threshold if any adjacent pixel exceeded the first threshold. In some embodiments, boundaries may be drawn that encompass the pixels that exceeded the threshold to generate the final mask. Additionally, the mask generator 126 may include curve simplification algorithms to reduce the effect of noise at the boundary of the mask.
[0078] A visual representation of the mask may be generated by the mask generator 126. The visual representation may include an overlay for the original image that can be communicated to the imaging system 30, the client device 20, or any other computer system where the segmentation results are viewed. In some embodiments, multiple masks may be generated based on multiple thresholds and differently colored overlays can be generated for each threshold.
[0079] FIG. 3 is a flow diagram illustrating the shape of the data (e.g., representing the three- dimensional image) as it is processed by the body structure segmentation system 100, according to some embodiments. To facilitate description, the data flow and shape will be described for a configuration where the output is a binary (or probability) mask for a slice of the three-dimensional image. The operations may be repeated to generate a binary (or probability) mask for each slice of the three-dimensional image. One skilled in the art will understand that extensions of the described process wherein the operations produce a mask for two or more slices after performing a single traversal through the operations illustrated in FIG. 3 should be considered within the scope of the present disclosure.
[0080] The operations begin by receiving a three-dimensional image including the structure of the individual. One or more layers (e.g., slices) of voxels are taken from the three- dimensional image. For example, input slices 502 may include C slices, each with a resolution of Ni by Ni. The feature extractor 112, is configured to produce several features from the input slices 502. Features may be extracted by projecting an input vector formed from a patch of an input slice 502 onto several anchor vectors, expressed as an affine transform as:
Figure imgf000026_0001
equation 1 where m is an integer from 0 to M - 1 and M is the number of anchor vectors (and features) that will be extracted from each patch or input vector. In equation 1, am represents the mlh anchor vector and x is the input vector obtained from a local region (e.g., patch, area, etc.) of one of the input slices. In some embodiments, the bias term bm is used to ensure that all features are positive.
[0081] Overlapping patches may be passed through equation 1 to generate a first feature map 504. The first feature map 504 may have a resolution the same or similar to that of the input slices 502. For example, voxels near the edge of a slice may either be created using zero padding or another boundary technique or the voxels near the edge may not have a respective pixel embedding in feature map 504 (e.g., reducing the resolution by a small amount). For each voxel of input slices 502 a corresponding pixel embedding is generated at the same x-y location in the feature map 504. The corresponding pixel embeddings may have Ki features. For example, each slice may provide Ki/C features to the pixel embedding at the location of the feature map 504 corresponding to the location of the voxel in an input slice. As another example, the center slice for which an output mask is generated may be allocated more anchor vectors and thus contribute to more features than the other slices. The feature map 504 may have a size of Ni by Ni by Ki. The operations for determining the anchor vectors will be described in more detail with reference to the body structure segmentation and training system 150 and the details of the operations of feature extractor 112 described with reference to FIGS. 4 and 5.
[0082] The first feature map 504 is stored (e.g., in memory) to be used in classification at the original resolution and is also subjected to a pooling operation by pooling calculator 116 as described previously. The pooling operation may reduce the resolution of the first feature map 504. For example, the height of the first feature map 504 may be divided by the height of the patches of the pooling operation and the width of the first feature map 504 may be divided by the width of the patches of the pooling operation. The output of the pooling operation is a pooled first feature map 506 with size of N2 by N2 by Ki.
[0083] The feature extractor 112 may apply a feature extraction operation to each feature set (e.g., slice) of the pooled first feature map 506. For example, the feature extractor 112 may apply equation 1 (e.g., using anchor vectors that are different than those used to transform input slices 502 to first feature map 504) to each slice of the pooled first feature map 506. In some embodiments, the anchor vectors and/or the number of anchor vectors is different for each feature set (e.g., channels) of the pooled first feature map 506. Each location (e.g., x-y location) of the pooled first feature map may be used to generate a corresponding pixel embedding as described previously to generate the second feature map 508.
[0084] The second feature map 508 is stored (e.g., in memory) to be used in classification at the original resolution and is also subjected to a pooling operation by pooling calculator 116 as described previously. It is noted that the pooling operation performed by the pooling calculator 116 on the second feature map 508 may be different (e.g., pool a larger/smaller number of pixels from the second feature map 508 together). The pooling calculator 116 may produce a pooled second feature map 510 that has the same number (e.g., K2) of features, but lower resolution (e.g., N3 by N3). Feature extraction may be repeated using the feature extractor 112. Training may be performed to determine another set of anchor vectors for each feature set of the pooled second feature map 510 or previous anchor vectors may be used (e.g., during training multiple feature map layers may be combined into the same training set). The result of the feature extractor 112 may be a third feature map 512 at the same N3 by N3 resolution, but with more features (e.g., K3 feature) as each of the K2 features of the pooled second feature map 510 may use one or more anchor vectors to produce one or more features of the third feature map 512. The third feature map is again stored and the process of pooling and feature extraction may be repeated another time to obtain a pooled third feature map 514 and forth feature map 516 with K4 > K3 features and a resolution of N4 by N4 where N4 < N3. Any number of feature extraction and/or pooling steps may be performed, the number of pooling operations may ultimately be limited by the resolution of the original slices and the size of the patch used during pooling.
[0085] In some embodiments, the encoding process of the body structure segmentation system 100 is complete after a number of feature extraction and pooling steps have been performed. For example, after four feature extraction steps and three pooling steps as shown in FIG. 3. The encoding process may result in a number of feature maps (e.g., four) available for classification by the classifiers 120. Each pixel embedding (e.g., a depth-wise array at a given x-y location of the feature map) corresponds to an array of the original slice of the three- dimensional image. [0086] In some embodiments, the position feature generator 118 generates feature maps related to the position of each feature embedding as described previously. The positional features may be concatenated with the feature maps (e.g., feature maps 504, 508, 510, and 512) by the concatenator 124. For example, the positional features may add 3 feature sets two the pixel embeddings of the feature map, one feature set indicating the x location of the pixel embeddings, one indicating the y location of the pixel embeddings, and one indicating the z location of the pixel embeddings.
[0087] In some embodiments, classification is performed starting at the lowest resolution rolling up information towards the highest resolution (e.g., the original resolution) of the three-dimensional image. A coarse classification may be performed using the feature map with the largest pixel embeddings (e.g., largest number of features). Advantageously, a coarse classifier with a larger number of features may be able to accurately classify various regions of the three-dimensional image without a large amount of training data and/or the computational burden associated with a large number of training data. Information from the low-resolution feature map (e.g., the fourth feature map 516) is included in the probability mask indicating the probability that a pixel embedding at a given location represents at least a portion of the body structure of the individual.
[0088] Referring still to FIG. 3, the lowest resolution feature map (e.g., the fourth feature map 516) may be concatenated with positional features and passed to a classifier of the classifiers 120. In some embodiments, each classifier is uniquely trained by pixel embeddings (and/or feature maps) at the respective layer within the encoding process to determine a probability mask. For example, the classifiers may include a uniquely trained classifier for each of the resolutions of feature maps (e.g., classifier 120a-d). Classification at each layer may be considered a voxel-wise classification problem (e.g., each a pixel embedding at a x- y location representing a voxel of the original slice is presented to a classifier of the classifiers 120).
[0089] In some embodiments, each pixel embedding of the fourth feature map 516 is concatenated with the positional features is presented layer 4 classifier 120a to determine a coarse probability map 518. The classifier 120a may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination of types of classification models. The coarse probability mask 518 may include a probability for each voxel of the original input slices 502 or the coarse probably map 518 may include a probability only for the center slice (e.g., the slice for which the final resolution mask will be created). The layer 4 classifier 120a may include a smoothing function to increase smoothness between classification results when the individual voxels are classified independently. For example, values of the coarse probability mask 518 may be averaged with adjacent values, the median of adjacent values may be calculated, thresholds to generate a binary classification may be modified if adjacent values are above a first threshold, or additional classifiers may be trained to smooth the probability mask (e.g., soft labels). Smoothing may similarly be performed on classification outputs (e.g., probability masks) at any resolution.
[0090] In some embodiments, a local refinement step (e.g., smoothing) is performed after the classification (e.g., after the layer 4 classifier 120a produces soft decisions for each of the input slices used to generate the feature maps). For example, a cube of the soft decisions (e.g., 3 by 3 by 3, etc.) may be selected from the probability mask. Each cube may be used to form a feature vector to train another classifier (e.g., an xgboost classifier). During segmentation of a new image, the soft decisions may be used to form a feature vector in the same way and the trained classifier may be used to generate a smoothed probability mask (e.g., smoothed soft decisions). The smoothing classifier may be used for a single resolution (e.g., a uniquely trained smoothing classifier may be trained for each resolution) or the same smoothing classifier may be used by each resolution. In some embodiments, smoothing classifiers are cascaded to update the soft decisions using multiple iterations. In addition, a median filter may be applied after multiple iterations of the smoothing classifiers.
[0091] The coarse probability mask 518 may be passed to the upsampler 122 (e.g., for scaling, upsampling, etc.) to generate soft decisions at the same resolution as the third feature map 512 so that information from the coarse probability mask 518 can be used by the layer 3 classifier 120b. For example, interpolation can be performed using trilinear or bilinear interpolation as described previously. The third feature map 512 may be concatenated with positional features for the current resolution and the upsampled version of the coarse probability mask 518 by concatenator 124 to generate the pixel embeddings that are presented to the input of the layer 3 classifier 120b.
[0092] In some embodiments, the classifier 120b may use logistic regression, decision trees, gradient boosting, xgboost, support vector machines, or any combination types of classification models to classify each pixel embedding at the third resolution as representing a portion of the structure of the individual. The classification results for each pixel embedding form a probability mask at the third resolution 520. The process of upsampling, concatenating, and classifying may be repeated to form the probability mask at the second resolution 522 and the probability mask at the original resolution 524. The upsampled probability masks that are concatenated onto each feature map may include upsampled versions of all previous probability masks. For example, coarse probability mask 518 may be upsampled and concatenated with feature map 512, the upsampled version of the coarse feature map 518 may also be concatenated with the probability mask at the third resolution 520, the combination may be together upsampled and concatenated with he probability mask at the second resolution 522, the combination again upsampled and concatenated with the first feature map 504 and used as input to the layer 1 classifier. The recursive upsampling operations is shown by the equation:
Figure imgf000030_0001
equation 2 where Fp represents a number of upsampled probability masks at the ith layer, y 1+1 represents the probability mask after classification at the i+1 layer (e.g., the next coarser resolution), © is the concatenation operator, and {■} T is the upsampling operator (e.g., bilinear interpolation, etc.).
[0093] The final classifier (e.g., the layer 1 classifier 120d) may be configured to only output a single probability mask (e.g., not include adjacent slices) or the layer 1 classifier 120d may determine three probability masks (e.g., in order to better perform smoothing) and then discard the extra masks before giving the result for the center slice of the input slices 502. In some embodiments, an entire three-dimensional image is processed slice by slice, each slice results in an output probably mask which can be concatenated together to form a final mask for the three-dimensional image. The final mask may be communicated over network 40 to the imaging system 30 or the client device 20 for viewing. For example, all voxels for which the final mask is greater than a threshold probability may be highlighted using a digital color overlay on the image or any slice of the image that is currently being viewed.
[0094] In some embodiments, the probability mask or the binary mask may include a number of visual representations. One or more visual representations can be delivered to a user interface. For example, the mask can be viewed as a color overlay on a screen, a virtual object in an augmented reality environment, an object on a holographic display, etc.
[0095] In some embodiments, the body structure segmentation system 100 is configured to generate a user interface (e.g., by the segmentation coordinator 110). The segmentation coordinator 110 may send instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images being segmented as well as the mask and/or color overlay. Additionally or alternatively, the segmentation coordinator 110 may include instructions to generate a local user interface on the computing hardware embodying the body structure segmentation system 100 (e.g., a desktop computer, a server, the imaging system 30, etc.). The user interface may be configured to provide user interaction (e.g., in addition to viewing the results). For example, if a radiologist is not fully satisfied with the segmentation result and desires to edit the prediction results a drag-and- drop interface may be provided to move a node of the segmentation boundary, to add voxels to the segmented body structure, etc. Segmentation adjustments can improve future segmentation results by providing additional training data that has been expertly labeled and the current body structure segmentation system 100 was unable to predict adequately. The new training data may be stored in the body structure segmentation training system 150.
[0096] In some embodiments, the body structure segmentation system 100 is configured to segment a body structure and the various regions of the body structure. Each classifier may produce a classification result (e.g., a probability) for each region or a separate classifier may be trained for each region. For example, the body structure segmentation system 100 may be trained to detect the central zone, peripheral zone, transitional zone, and anterior zone of the prostate. As another example, the body structure segmentation system 100 may be trained to segment the four chambers of the heart.
[0097] FIGS. 4 and 5 provide more detail to the feature extraction and pooling process performed by the body structure segmentation system. FIG. 4 shows flow of operations 320 for generating a reduced resolution feature map by feature extraction and pooling, according to some embodiments. FIG. 5 is a flow diagram illustrating the operations of the feature extraction and pooling according to some embodiments. The flow of operations 320 is described with reference to FIG. 5 using the second feature extraction operation and second pooling operation of FIG. 3 as an example (e.g., the operations that convert the pooled first feature map 506 to the pooled second feature map 510.
[0098] In some embodiments, the flow of operations 320 begins by obtaining a local region (e.g., a patch, area, set of adjacent voxel s/pixels, etc.) of a slice of a first plurality of feature slices (e.g., a feature map) in operation 322. For example, the pooled first feature map 506 is sliced into feature slices 530-534 (e.g., one for each of the Ki features of the pooled first feature map 506) and a local region (e.g., region 536) is obtained. The local region may be vectorized to obtain a vector representation of the data of the local region in operation 324. As show in FIG. 5 the local region 536 is converted from the two-dimensional form (e.g., 3 by 3) to a single dimension array (e.g., 9 by 1) to form the vector representation 538.
[0099] The flow of operations 320 may include performing a feature extraction operation on the vector representation, the feature extraction operation, for example, determined by a Saab transformation of feature slices from a plurality of images in operation 326. The calculation of a feature was shown in equation 1. In some embodiments, multiple anchor vectors can be combined into a matrix representation (e.g., matrix 540) and multiplied by the vector representation (e.g., the vector representation 538). For example, the matrix 540 incorporating the anchor vectors is a column of the transposed anchor vectors as shown in: y equation 3
Figure imgf000032_0001
and x is the vector representation of the local region (e.g., vector representation 538) and y includes the features. For example, y represents a portion of the pixel embedding (e.g., portion 542) for the local region of the current feature slice.
[0100] In some embodiments, the features (e.g., portion of the pixel embedding) are distributed across a respective x-y location of a second plurality of feature slices of the same resolution in operation 328. For example, the portion of the pixel embedding 542 (e.g., the features) may be distributed to the same pixel of each feature slice of the plurality of feature slices 544a-c.
[0101] The flow of operations 320 may include repeating operations 322-328 for each overlapping local region to complete the second plurality of feature slices in operation 330. Another local region may be obtained from the same feature slice 530. For example, the next local region 546 processed may be selected by shifting the local region 536 by a number of locations in one direction (e.g., one location left, one location down, etc.). The next local region 546 may overlap with local region 536. In some embodiments, the next local region 546 is vectorized (e.g., in operation 324), multiplied by the matrix defining the feature extraction operation (e.g., in operation 326), and distributed across feature maps or feature slices 544a-c to form another portion of a pixel embedding adjacent to the portion of the pixel embedding obtained from local region 536. The operations 322-330 may be repeated to from a pixel embedding for each location of the feature slice 530 and complete the layer 2 feature slices 544a-c.
[0102] The feature map 506 may include several feature slices (e.g., feature slices 530-534) that can be used to determine new features for the output feature map 508. In some embodiments, the flow of operations 320 includes repeating operations 322-330 for each feature slice of the input feature map to obtain a plurality of feature slices for each feature slice of the input feature map. For example, and as shown in FIG. 5, local regions of the feature slice 532 are vectorized and multiplied by a second matrix 548 to determine a pixel embedding 550 for a second plurality of feature slices 552. The second matrix 548 may be the same as the matrix 540 (e.g., representing the same anchor vectors, same feature extraction operation, same Saab transformation, etc.). Or in some embodiments, the matrix 540 may represent anchor vectors specifically trained for the feature slice 532. It should be understood that any slice may use a feature extraction operation that is trained for the feature of the slice (e.g., using training data of that feature) or may be trained for a number of features (e.g., trained using training data from several features).
[0103] A feature extraction operation may be performed for each slice (e.g., feature slice 530- 534) of the input feature map 506 to generate respective pluralities of feature slices of the same resolution (e.g., pluralities of feature slices 544, 552, and 554). Each of the pluralities can be concatenated together to form the output feature map 508 of the feature extractor 112. For example, the feature slice 530 may be used to generate Q features and represent the first Q features of the output feature map 508, the feature slice 532 may be used to generate S features and become features S+l to S+Q of the output feature map 508, and the final feature slice 534 of the input feature map 506 may be used to generate R features and become features K2-R+1 to K2 of the output feature map. Any feature slices of the input feature map 506 between feature slice 532 and 534 may be used to generate additional features that also become part of the output feature map 508.
[0104] As shown in FIG. 3 the output of the feature extractor 112 is passed through pooling calculator 116 to generate a reduced resolution feature map. Referring to FIG. 5, the pooling operation of the pooling calculator 116 is described in more detail. In general, pooling maps a local region of a feature map to a single location of a reduced resolution feature map. For example, the output feature map 508 of resolution N2 by N2 can be reduced to N3 by N3 in feature map 510, where in N3 = N2/P by performing pooling on non-overlapping local regions of the feature map. Pooling may include taking the maximum, average, or any other operation of a local region of a feature (e.g., median, second maximum, etc.) of the P by P local region. It is noted that the local region does not need to be square nor do the local regions need to be non-overlapping. For example, a 4 by 2 local region can be shifted by 2 in each direction resulting in regions that overlap in one of the two directions with the output resolution reduced by a factor of the amount the local region is shifted (e.g., the stride, in the present example, 2).
Body Structure Segmentation Training System
[0105] FIG. 6 shows another block diagram of the imaging and structure segmentation system 10 according to some embodiments. FIG. 6 shows some embodiments of an implementation of the body structure segmentation training system 150. The body structure segmentation training system 150 may include a communications interface 152 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40. For example, the communications interface 152 may be used to communicate the trained machine learning models (e.g., the feature extraction models, and the classifiers and/or their parameters) to the body structure segmentation system 100. The communications interface 152 may be of the same or similar type of the communications interface 102 of the body structure segmentation system 100. The body structure segmentation system 100 may also include a processing circuit 154 having a processor 156 and memory 158. For example, the processor 156 may be configured to execute instructions contained on the memory 158. In some embodiments, the body structure segmentation training system 150 and the body structure segmentation system 100 are implemented on the same hardware (e.g., on hardware of the imaging system 30, on the same computer, etc.). For example, the processing circuit 154 may be the same processing circuit as the processing circuit 104, similarly the communications interface, memory and processors may be shared between the body structure segmentation system 100 and the body structure segmentation training system 150.
[0106] In some embodiments, the body structure segmentation training system 150 may be implemented on separate hardware from the body structure segmentation system 100. The processor 156 may be of similar type or of different type from the processor 106. The processor 156 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components. The processor 156 may be configured to execute computer code and/or instructions stored in the memory 158 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.). The processor 156 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof. One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
[0107] The memory 158 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure. The memory 158 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions. The memory 158 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure. The memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
[0108] According to some embodiments, the body structure segmentation training system 150 is to acquire several three-dimensional images of a portion of the body for training the feature extractor 112 (e.g., the feature extraction operations) and the classifiers of 120 of the body structure segmentation system 100. In some embodiments, the body structure segmentation training system 150 applies unsupervised learning to identify features of interest for the classifiers 120. For example, the feature extractor 112 may be trained to determine feature maps of various resolutions to be used by the classifiers 120. In some embodiments, the body structure segmentation training system 150 applies a supervised learning to train the classifiers 120 to determine if pixel embedding represents a location of the feature map that includes at least a portion of the body structure of the individual being imaged. It is noted that because training the feature extractor 112 is performed with unsupervised learning it does not require previously segmented (e.g., labeled) images for training. Advantageously, training using unsupervised learning allows unlabeled data to be used for training and greatly increases the available training data. Previously segmented images may be provided to train the classifiers 120. The segmented images are additionally used to train the feature extractor (e.g., without using the labeled segmentations) in some embodiments.
[0109] In some embodiments, the body structure segmentation training system 150 applies unsupervised learning to train the classifiers 120 to determine if pixel embedding represents a location of the feature map that includes at least a portion of the body structure of the individual being imaged. For example, clustering algorithms may generate classification rules (e.g., boundaries, decisions, etc.) in the space of the pixel embeddings that can be used to determine if a location represents at least a portion of the body structure of the individual. Clustering algorithms are not limited to dichotomizers but can also be used to cluster two or more expected regions of the body structure. For example, k-means clustering, hierarchical clustering, or any other suitable clustering algorithm may be used to determine pixel embeddings that represent various regions of the body structure imaged.
[0110] In some embodiments, the body structure segmentation training system 150 includes a training coordinator 160, the feature extractor 112, training data storage 166, a data selector 168, the position feature generator 118, the classifiers 120, a loss calculator 172, a parameter tuner 174, and a Saab transformer 180. In addition, several other components of the body structure segmentation system 100 may also be included in the body structure segmentation training system 150 (e.g., the pooling calculator 116, the upsampler 122, the concatenator 124, etc.) and used to execute the segmentation model during training of the classifiers 120. The training coordinator 160 may be configured to control the timing and flow of data through the other circuitry of the body structure segmentation training system 150. For example, the training coordinator 160 may cause the instructions or circuits to execute in a specific order to perform the function of the body structure segmentation training system 150. In some embodiments, the training coordinator 160 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input. For example, the training coordinator 160 may cause the instructions and/or circuits of the body structure segmentation training system 150 to perform the training operations illustrated in FIGS. 7 and 8.
[OHl] In some embodiments, the training data storage 166 is configured to maintain (e.g., store, save, acquire, etc.) three-dimensional images of body structures that can be used to train the feature extractor 112, the classifiers 120, or both. Three-dimensional images without an associated segmentation (e.g., unlabeled) can be used to perform unsupervised learning (e.g., of the feature extractor 112 or some classification algorithms). Three-dimensional images of a body structure that has been previously segmented may be used to perform both unsupervised learning and supervised learning. In some embodiments, the data selector 168 selects the correct data to be used for a particular training algorithm. For example, the data selector 168 may select only labeled data when supervised learning is being performed to train the classifiers 120. The data selector 168 may also split data into training and validation sets. The validation data may be used to stop algorithm training, choose a classification algorithm, or to determine appropriate values of hyperparameters (e.g., the number of features that should be used, maximum depth of a decision tree, regularization terms, etc.)
[0112] In some embodiments, the body structure segmentation training system 150 includes the Saab transformer 180. The Saab transformer 180 may be used to determine the feature extraction operations of the feature extractor 112. For example, the Saab transformer 180 may be used to determine the anchor vectors used in the feature extraction operations, the number of anchor vectors to use, and/or the matrix representation of the feature extraction operations (e.g., the matrix 540 or 548). In some embodiments, the Saab transformer 180 includes a sampler 182, a DC vector calculator 184, and a principal components analysis module (PCA) 186 to perform the operations of the Saab calculation.
[0113] FIG. 7 illustrates the details of the operations of the Saab transformer 180 according to some embodiments. The Saab transformer receives a number of training samples from the training data storage 166. The training samples, for example, may include three-dimensional images of a body structure for segmentation (e.g., three-dimensional image 560). The sampler 182 may determine local regions of the three-dimensional image to be used to train the feature extraction operation. The sampler 182 may choose the local regions based on one or more criteria for filtering appropriate local regions. During segmentation (e.g., during operation of the body structure segmentation system 100) the feature extraction operation may be applied to local regions that satisfy the same or similar criteria. For example, the sampler 182 may select data based on the slice (e.g., the depth) to which the local region belongs and training different feature extraction operations for various layers. The sampler may also disqualify regions, slices, or whole three-dimensional images based on one or more criteria. For example, a three-dimensional image may be disqualified from use in training if the signal to noise ratio is high, a slice may be disqualified from use in training if the image is subject to a blur, or any other undesirable image artifact may cause the image or slice to be disqualified.
[0114] In some embodiments, multiple models may be trained based on criteria related to the three-dimensional image. For example, different feature extraction operations may be learned for three-dimensional images of differing noise levels.
[0115] The sampler 182 may select a local region (e.g., local region 562) and vectorize the data of the local region (e.g., 3 by 3 region) into a training vector for the Saab transformation related to the initial three-dimensional image. The sampler 182 may select a local region displaced from the local region 562 by a number of voxels in any of the three directions (e.g., along the heigh, width, or depth). The second local region may overlap with the local region 562 depending on the configuration of the sampler 182. The sampling process may be repeated for all potential local regions of the three-dimensional image 562 and repeated again for all three-dimensional images in the training data storage 166 to form a set of training vectors 564. The training vectors 564 are input to the Saab transformer 180 to calculate the anchor vectors of the feature extraction operation. The anchor vectors (and thus the feature extraction operation), for example, may be used on all slices of the three-dimensional image. Referring back to FIG. 3, the feature extraction operation may be used to transform the slices 502 to the feature map 504.
[0116] In some embodiments, the training vectors 564 are input to the Saab transformer 180 and a DC anchor vector is determined by the DC vector calculator 184. The DC vector calculator 184 may generate the DC anchor vector as: equation 4
Figure imgf000038_0001
The additional, M - 1 anchor vectors (the AC anchor vectors) are determined by first subtracting off the DC component (e.g., subtracting the projection of a training vector onto the DC anchor vector) from the from all of the training vectors 564 by: xAC = x — xDC, where xDC = x • aQ, where • represents the proj ection, in the present case the dot product because aQ is normalized. In some embodiments, the additional M - 1 anchor vectors are learned by performing principal components analysis (PCA) on the AC component of the training vectors 564. For example, the DC vector calculator 184 may subtract the DC component from each of the training vectors and the PC A calculator 186 may perform principal components analysis to calculate the principal components (e.g., the anchor vectors 566) that are used in the feature extraction operator. One skilled in the art will recognize that performing PCA on the AC component of the training vectors 564 may include arranging the AC components of the training vectors 564 into a data matrix, multiplying the transpose of the data matrix by the data matrix, and calculating the eigen vectors (and in some embodiments eigen values) of the N by N result. Advantageously, such calculations do not require the performance of iterations to determine an optimal value or training by adjusting the parameters over multiple (e.g., hundreds) of epochs of the training data, and such calculations lead to efficient feature extraction training.
[0117] In some embodiments, the number of principal components selected (M - 1) to be used as anchor vectors is chosen based on the eigenvalues associated with the eigenvectors calculated during the PCA process. The eigenvalues may represent the representative capability of each of the principal components. The principal components that are associated with an eigenvalue of at least a threshold fraction (e.g., 10%, 20%, etc.) of the largest eigen value may be kept. In some embodiments, further processing is performed to determine the principal components to be kept. For example, the eigenvalues may be plotted in descending order and the “elbow” in the plot of the eigenvalues (e.g., the point at which the eigenvalues no longer decrease significantly) may be used as eigenvalue threshold to determine which principal components are used as anchor vectors.
[0118] After determining the anchor vectors for the slices of the three-dimensional image (e.g., the feature extraction operation used in the first layer of FIG. 3), the body structure segmentation training system 150 may use the feature extractor 112 and the pooling calculator 116 to convert slices of the three-dimensional images in the training data storage 166 into a feature map of reduced dimensionality (e.g., feature map 506). The operations illustrated in FIG. 7 may be repeated to determine anchor vectors for each feature slice of the feature map of reduced resolution. For example, the sampler 182 may choose local regions 562 related to only a single feature (e.g., slice) of the feature map of reduced dimensionality and unique anchor vectors may be calculated and stored for each feature slice of the feature map of reduced resolution. The process of using the feature extractor 112 with the new anchor vectors and performing pooling to calculate subsequent feature maps of reduced resolution is repeated to determine the sets of anchor vectors that are used by the feature extractor 112 at all layers of the encoding process. For example, to determine the anchor vectors used in the feature extraction operation to convert feature map 506 to feature map 508, to convert feature map 510 to feature map 512, etc.
[0119] In some embodiments, a number of different feature slices (e.g., the first, fourth, and seventh feature slice of the feature map after the first feature extraction operation is performed) can be used to determine the anchor vectors for a feature extraction operation that is used for all those feature slices. It is noted, that if more than one slice from the original image is used to generate the feature map and the same feature extraction operation is used for each slice, then the features generated from each slice may have similar information (but for a different slice) and using the same feature extraction operation may save computations in training, reduce the amount of training data required, and offer similar performance compared to performing training individually for each slice. In some embodiments, training is performed for each slice and feature slices are combined to use the same feature extraction operation if the anchor vectors satisfy a similarity criterion (e.g., pairs of anchor vectors between each set have a small angle between them, the dot product is near one, etc.)
[0120] FIG. 8 illustrates the operations to train the classifiers 120 according to some embodiments. A slice 502 of the three-dimensional image (as well as adjacent slices) may be used to adjust the parameters of all the classification layers in the classifiers 120. Training data slices may be provided to a labeler 190. The labeler 190 may be configured to allow an expert to segment the body structure of the three-dimensional image so that it can be used in supervised training of the classifiers 120. For example, the labeler 190 may provide instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images and allow a user to select the regions of the three-dimensional image that correspond to the body structure. Additionally or alternatively, labeler 190 may include instructions to generate a local user interface on the computing hardware embodying the body structure segmentation training system 150 (e.g., a desktop computer, a server, the imaging system 30, etc.). Segmented three-dimensional images may then be used to train the classifiers 120.
[0121] To train the classifiers, the features (e.g., the pixel embeddings) used to determine if a location corresponds to a portion of the body structure may be calculated. In some embodiments, low resolution classification is used as part of the features to perform the higher resolution classification and the classifiers 120 are trained starting with the lowest resolution and working towards the resolution of the original three-dimensional image.
[0122] As shown in FIG. 8 each slice of the training data is first passed through the sequence of feature extraction operations and pooling calculations to determine the corresponding feature maps of all resolutions. Positional feature may also be added to the feature map as show in FIG. 3. A pixel embedding (e.g., an array through the depth) of the feature map concatenated with the positional features and probability map (and for resolutions other than the lowest resolution also concatenated with the upsampled probability maps of previous classifiers) is used by training data for a classifier trainer along with the segmentation label (e.g., the label as to whether or not the location represents part of the body structure) at the respective location 572 on the slice 502. A number of pixel embeddings (e.g., pixel embedding 570) and the label at the respective location (e.g., location 572) are generated from the three-dimensional images of the training data storage 166.
[0123] In some embodiments, training of a classifier (e.g., classifier 120a) is performed iteratively. For example, the classifier 120 may determine a probability that the pixel embedding 570 represents a portion of the body structure. The loss calculator 172 may compare the probability calculated by the classifier 120 to the label for the location 572 to determine a portion of a loss function. For example, the loss calculator 172 may compute the sum of the loss by comparing the classification of several pixel embeddings to the respective labels to compute a loss function. Based on the loss function the parameter tuner 174 may update the classifier 120 (e.g., parameters of the classifier, rules within the classifier, etc.) to improve the loss function with respect to the training data. In some embodiments, the image is labeled (e.g., segmented) at the original resolution and the loss calculator 172 determines a label for comparison purposes at the lower resolution. For example, the loss calculator may consider a location of the lower resolution image to be labeled as part of the body structure if any of the corresponding locations of the original resolution is labeled as part of the body structure, if a majority of the corresponding locations of the original resolution is labeled as part of the body structure, if all of the corresponding locations of the original resolution is labeled as part of the body structure, or any other suitable method for reducing the resolution of the segmented three-dimensional image.
[0124] In some embodiments, the loss function may include binary cross entropy loss (e.g., for whole structure segmentation) or multi-class cross-entropy loss (e.g., for region segmentation). In some embodiments, the classifier 120 is a neural network classifier (e.g., perceptron machine, etc.) and in some embodiments, the classifier 120 includes one or more decision trees (e.g., xgboost).
[0125] In some embodiments, the classifier 120a for the lowest resolution is trained as illustrated by the operations of FIG. 8. The process may be repeated as illustrated by FIG. 8 for the next higher resolution. For example, the feature maps of as previously calculated using the feature extractor 112 for the next resolution may be concatenated and with positional features. In addition, the training data may be classified at the lowest resolution by the trained classifier 120a and upsampled to provide a probability map at the next higher resolution. The upsampled probability map may be concatenated with the feature map and the positional features to obtain classifier inputs 568 for the next higher resolution. Pixel embeddings (e.g., pixel embedding 570) may be classified by the classifier 120b (e.g., at the next higher resolution) and compared to the respective label by the loss calculator 172. The parameter tuner 174 may then adjust the model of the classifier 120b so as to improve the loss function over all classifications of pixel embeddings of the training data at the current resolution. The process may be repeated to train the classifiers 120c and 120d.
[0126] In some embodiments, the classifier 120 at each resolution is a different type of classifier. For example, the classifier at the lowest resolution may be an xgboost classifier and the classifier of another resolution may be a neural network classifier. In some embodiments, the classifier 120 uses the same type of classification model for all resolutions. For example, xgboost may be used as the classifier at each resolution.
[0127] In some embodiments, the process for training a classifier is repeated and the resolution of the classifier is increased with each resolution.
Experimental Results of the Body Structure Identification System
[0128] To demonstrate the effectiveness of the methods of the current application experiments based on one public MR image database NCI-ISBI 2013 Challenge and one private inhouse USC-Keck dataset were conducted. The data sets are related to the segmentation of the prostate gland. The ISBI-2013 dataset consists of 60 training cases of axial T2-weighted MR 3D series, where half were obtained at 1 5T (Philips Achieva at Boston Medical Center) and the other half at 3T (Siemens TIM at Radboud University Nijmegen Medical Center). Since the ground truth segmentation includes background, peripheral zone (PZ), and central gland (CG), PZ and CG were merged into one class — prostate area. The pixel spacing within each slice ranges from 0.39 mm to 0.75 mm, while the through-plane resolution ranges from 3.0 mm to 4.0 mm among different patients.
[0129] The USC-Keck dataset consists of a cohort of 260 patients collected in the Keck Medicine School in the University of Southern California. Besides T2-w 3D series, for each patient T2-Cube series is also available. It is known that T2-Cube has smaller pixel spacing, especially along the z-axis. That means, higher resolution and thinner slices. Specifically, for T2-w pixel spacing ranges from 0.5 mm to 0.7 mm and the through-plane resolution (z-axis) from 3.0 mm to 4.0 mm. For T2-Cube pixel spacing ranges from 0.83 mm to 0.83 mm and the through-plane resolution (z-axis) from 1.4 mm to 4.0 mm. The scanner used to acquire those images is the GE 3T with 8ch Cardiac coil. In the experimental analysis with USC- Keck data, the T2-Cube series is used since it may provide more accurate segmentation results because of the higher perspicuity of the images (stemming from the smaller voxel spacing).
[0130] For both datasets, the resolution of different images is regularized to the same physical resolution of 0.625 x 0.625 x 1.5mm. Lanczos interpolation is used where the factor is calculated based on the original pixel spacing and through-plane resolution of each image. The through plane resolution is increased so that the segmentation in the 3D space can be more accurate. To reduce the artifacts while acquiring the images, contrast enhancement using CLAHE is applied. Finally, before feeding the segmentation algorithm disclosed herein (“PSHop”) input (e.g., the cl input, for the whole gland segmentation task the input sequence is resized to 128 x 128. For the zonal segmentation task, a 256 x 256 centered crop around the segmented gland is resized to 128 x 128 to standardize the PSHop input.
[0131] To quantitatively evaluate the performance, Dice Similarity Coefficient (DSC) is used, expressed as: , equation 5
Figure imgf000043_0001
in a binary scenario, where X and Y represent the ground truth and the predicted segmentation mask, respectively. DSC is used in evaluating segmentation tasks for medical images. It measures the ratio of the intersection of two binary sets to the averaged cardinality.
[0132] For both datasets the same experimental settings are applied. ForISBI-2013 and USC- Keck 5-fold cross validation is applied on the 60 and 260 training images, respectively and the mean and standard deviation of the evaluation scores are calculated. That is, for ISBI- 2013 48 sequences are used for training PSHop from each fold and the rest 12 for validation. For USC-Keck dataset 206 sequences are used for training and the rest 54 for validation. [0133] For benchmarking PSHop with other DL-based methods the same experiment setting is conducted for both datasets, and V-Net and U-Net based architectures are trained. Besides segmentation performance comparison using the DSC metric, The model size and complexity of the models are also compared, since one advantage and motivation of the present disclosure is to offer a lightweight solution comparing to other methods.
[0134] The DSC scores of ISBI-2013 and USC Keck datasets obtained by the proposed PSHop method for the prostate gland segmentation task are summarized in Table 1. The averaged validation performance over the 5-folds is reported and the standard deviation as well (shown in parenthesis). PSHop has a very competitive performance among the two baseline DL-based works. In particular, PSHop outperforms V-Net on ISBI-2013 dataset and U-Net on the USC-Keck one. ISBI-2013 has considerably fewer patients than USC-Keck. Therefore, V-Net performs better when it is training with a lot of training samples, while U- Net has a higher performance with fewer training samples. That is also evident from the high standard deviation on ISBI data from V-Net. Another important observation is that PSHop has a more stable performance (low standard deviation in both experiments), regardless of the number of training samples. This underlines one of the advantages of the current disclosure is that it is more stable even for fewer training samples, while large deep learning (DL) models fail to achieve high performance when data are scarce. That also confirms that the current statistical -based feedforward models can still perform well even with a small number of training samples.
Figure imgf000044_0001
Table 1 : Whole Gland Segmentation Results
[0135] Referring now to zonal segmentation performance, Table 2 shows the benchmarking on USC-Keck dataset, since it is larger than the ISBI-2013 and thereby stronger conclusions can be drawn. For TZ, PSHop surpasses U-Net by large margins and maintains a small performance gap with V-Net. PSHop achieves a much higher DSC score comparing to the U- Net, which is the baseline architecture in the art. For the smaller ISBI-2013 dataset PSHop outperforms the other methods on the PZ (see Table 3), mainly due to the small number of training samples that PSHop has an advantage. On the TZ, PSHop surpasses the U-Net performance by large margins.
Figure imgf000045_0001
Table 2: Zonal Segmentation using USC-Keck Data
Figure imgf000045_0002
Table 3: Zonal Segmentation using ISBI-2013 Data
[0136] In the above comparisons, V-Net generally performs better when trained with a large number of samples and U-Net performs well only on the ISBI-2013, which has many fewer samples. This is sensible because V-Net has many more trainable parameters comparing to U-Net and thus can fit better the data diversity. Table 4 demonstrates the advantages of the current application when it comes to complexity and model size comparison. PSHop has an order of magnitude less parameters than the DL-based models. Also, in terms of complexity, it has 190 times fewer FLOPS than U-Net and 5269 times fewer FLOPS than V-Net. These comparisons stress the tremendous advantages of the present disclosure for application deployment.
Figure imgf000045_0003
Table 4: Zonal Segmentation using ISBI-2013 Data [0137] Overall, PSHop maintains a very competitive standing performance-wise with other DL baseline models, outperforming in general U-Net in both tasks and having performance similar to V-Net at a significantly reduced model size and number of FLOPS.
Diseased Structure Identification System
[0138] The technological problem of identifying a body structure as diseased and/or determining if a region of interest of a body structure represents a diseased region may be improved using techniques similar to those used by the body structure segmentation system 100. For example, the diseased structure identification system may make use of similar unsupervised learning techniques to determine feature extraction operations that are configured to generate feature maps of various resolutions. The techniques described herein may provide similar computational savings during online disease identification and during training. The techniques may also provide a savings in the required amount of training data that has previously been labeled.
[0139] FIG. 9 is a block diagram of an imaging and disease identification system 12 configured to obtain medical imagery of a body structure and identify regions of a body structure that are indicative of a disease according to some embodiments. The imaging and disease identification system 12 may detect disease in specific bone and/or joint structures, soft tissue (e.g., muscles, tendons, and ligaments), organs (e.g., the heart, lungs, kidneys, prostate, etc.), the nervous system, digestive tractor, or any other body structure for which disease may be detectable via medical imagery. The imaging and disease identification system 12 includes one or more client devices 20, one or more imaging systems 30, a diseased structure identification training system 250, and a diseased structure identification system 200 shown to be communicably coupled via a network 40. The network 40 can include routers, switches, antennas, computers, and any other hardware required to communicate information between the components of the imaging and disease identification system 12 (e.g., from the imaging system 30 to the diseased structure identification system 200. A portion of the network 40 can be wireless and/or a portion of the network 40 can be wired. The network 40 can include one or more networks with routers to facilitate data transfer between the different networks.
[0140] The imaging and disease identification system 12 is shown to be distributed across several devices (e.g., networked computers). It is contemplated that in some embodiments, the subsystems, components, and/or functionality of the imaging and structure segmentation system 10 may be distributed on different devices. For example, the diseased structure identification system 200 may be configured within the imaging system 30 using the processors of the imaging system 30 to execute. In some embodiments, the diseased structure identification system 200 and the diseased structure identification training system 250 are configured within the same computer hardware (e.g., within a server local to the imaging system, within a remote server, or on a node within a cloud computing architecture).
[0141] The diseased structure identification system 200 may include a communications interface 202 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40. The diseased structure identification system 200 may also include a processing circuit 204 having a processor 206 and memory 208. For example, the processor 206 may be configured to execute instructions contained on the memory 208.
[0142] The processor 206 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components. The processor 206 may be configured to execute computer code and/or instructions stored in the memory 208 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.). The processor 206 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof. One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
[0143] The memory 208 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure. The memory 208 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions. The memory 208 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure. The memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
[0144] According to some embodiments, the diseased structure identification system 200 is to acquire a three-dimensional image of a portion of the body, perform a sequence of feature extraction operations and pooling operations on one or more sets of adjacent voxels (e.g., patches) surrounding a voxel of the three-dimensional image to obtain feature maps of various resolutions and varying depths (e.g., number of features), and to perform voxel-wise classification using the feature maps obtained from a set of adjacent voxels. In some embodiments, the feature maps are obtained for multiple imaging types (e.g., T2-weighted (T2w), apparent diffusion coefficient (ADC), and diffusion-weighted imaging (DWI). The features may be combined to perform a voxel-wise classification based on the surrounding set of adjacent pixel of each of the multiple imaging types. Regions of interest (e.g., clusters of voxels identified as representing a diseased area) may be determined. In some embodiments, additional features (e.g., radiomic features and/or anomaly features) are obtained for the regions of interest. The regions of interest may be classified as being a diseased area based on the combined feature set. The diseased structure identification system 200 may be configured to identify diseased regions in one or more body structures including, but not limited to: glands including the prostate, thyroid, etc.; organs including the heart, the liver, etc.; bones; joints; muscles; tumors; and other body structures for which segmentation may be useful.
[0145] In some embodiments, the diseased structure identification system 200 includes an identification coordinator 210, a feature extractor 212, feature selector 214, a region of interest (ROI) identifier 216, a region selector 218, a concatenator 124, a voxel classifier 222, and a stage 2 classifier 220, an anomaly feature collector 224 and a radiomics feature collector 226. The functionality and features of each will be described in more detail as related to FIGS. 10-13 herein. The identification coordinator 210 may be configured to control the timing and flow of data through the other circuitry of the diseased structure identification system 200. For example, the identification coordinator 210 may cause the instructions or circuits to execute in a specific order to perform the function of the diseased structure identification system 200. In some embodiments, the identification coordinator 210 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input. For example, the identification coordinator 210 may cause the instructions and/or circuits of the diseased structure identification system 200 to perform a flow of operations 400 in FIG. 10 to segment a body structure.
[0146] With reference to FIG. 10, the flow of operations 400 includes obtaining (e.g., acquiring, receiving, etc.) a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels in operation 402, according to some embodiments. For example, the imaging system 30 may communicate the three- dimensional image over the network 40 to begin processing (e.g., segmentation). In some embodiments, the identification coordinator 210 requests (e.g., form a database, from imaging system 30, etc.) one or more three-dimensional images from a database to process (e.g., to perform batch processing). In some embodiments, a radiologist, lab technician, etc. selects a three-dimensional image for within which diseased regions are to be identified.
[0147] In some embodiments, the flow of operations 400 includes identifying a set of adjacent voxels including the voxel (e.g., to be classified) of the plurality of voxels in operation 404. In operation 404 a local, two-dimensional set of adjacent voxels of the three- dimensional region is obtained. The identification coordinator 210 may select adjacent voxels from the three-dimensional image. For example, a 24 by 24 region may be selected. The identification coordinator 210 may select overlapping sets of voxels, for example, by shifting the center of the set by one voxel as each voxel is classified. The set of adjacent voxels may provide information (e.g., the adjacent pixels) that is used to determine if the voxel should be classified as being representative of a diseased area.
[0148] In some embodiments, the flow of operations 400 includes executing a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions in operation 406. The operation 406 may be performed by the feature extractor 212. The feature extractor 212 may be configured similar to feature extractor 112. To generate a set of features a local region (e.g., patch, area, etc.) of the of the set of adjacent voxels is obtained. For example, the set of the adjacent voxels may be a 24 by 24 patch and the local region used to generate the features may be a 3 by 3 patch within the set of adjacent voxels. The feature extractor 212 may obtain a feature extraction operator (e.g., a function, a linear transformation, a matrix, etc.) from storage and perform the feature extraction operation on a vectorized version of the local region. For each local region of the set of the adjacent, the feature extractor 212 may output a set of features (e.g., an array) that represent the local area of the set of the adjacent voxels. Shifting the local area and again performing feature extraction with the same feature extraction operator may result in a feature map of the same or similar resolution of the set of the adjacent voxels. In some embodiments, the three- dimensional array (e.g., feature map) for the set of adjacent voxels has height and width equal to that of the local region and depth equal to the number of features output by the feature extraction operator. In some embodiments, the resolution of the feature map generated (e.g., the height and width) is reduced by an amount that depends on the size of the smaller local region used to generate the pixel embeddings of the feature map of the local region. For example, if zero padding is not performed, features may not be extracted for voxels near the boundary of the set of adjacent voxels.
[0149] To generate feature maps of various resolutions a pooling operation may be performed. For example, pooling may be performed by pooling calculator 116. In some embodiments, pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output. For example, the voxels may be combined by taking the maximum of the voxels, by taking the average of the voxels, or any other suitable function of values of multiple voxels. In some embodiments, two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution.
[0150] In some embodiments, the feature extractor 212 is configured to perform feature extraction on each feature slice (e.g., slice of a feature map individually) of an input feature map. For example, feature extraction may be performed on a 11 by 11 by 9 input feature map resulting in a 9 by 9 by 56 feature map. Each of the 9 feature slices of the input feature map may provide a number of the 56 output feature maps. For example, the feature extractor may perform the operations demonstrated in FIG. 5. In some embodiments, the feature extractor 212 is configured to determine the number of features to keep for each feature slice of a feature map. For example, feature maps with low energy (e.g., low eigenvalues) may be discarded during the training phase.
[0151] In some embodiments, the feature extractor 212 is configured to perform PC A to determine a spectrum on each feature slice of an input feature map. The PCA may provide global features related to the entire set of adjacent voxels in addition to the local features obtained through the feature extraction operations. For each slice of the feature map global features may be extracted by:
GKi = PCA(Si X Si), equation 6. [0152] In some embodiments, the feature extractor 212 is configured to extract features from various imagery modes. For example, T2w, ADC, and DWI modes may each produce a three- dimensional image that is processed by the feature extractor 212. The various imagery modes may provide important information to various tissue types and conditions.
[0153] In some embodiments, the flow of operations 400 includes executing a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual in operation 408. Voxel classifier 222 may perform operation 408. Voxel classifier 222 may be configured to perform to use, as input, a number of features output by the feature extractor 212. The voxel classifier 222 may be configured to use all the features extracted from the adjacent set of voxels. In some embodiments, a subset of the features (e.g., those with high discriminant potential) are used by the voxel classifier. For example, if the feature extraction algorithm converted a 24 by 24 set of adjacent voxels to a 22 by 22 by 9 feature map and a 9 by 9 by 56 feature map the result is nearly 9,000 local features. The 1000 samples that have the most discriminative potential may be used to train the voxel classifier. The voxel classifier 222 may be configured to produce a classification output (e.g., a probability or a binary decision) that the voxel included in the set of adjacent voxels (e.g., a center voxel) is part of a diseased region of the body structure. The voxel classifier 222 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
[0154] In some embodiments, the selection of adjacent voxels, execution of feature extraction, and classification of the voxel (e.g., operations 404-408) may be repeated for each voxel of the three-dimensional image. The result of operations 404-408 for each voxel of the three-dimensional image is a probability for each voxel of the three-dimensional image that the voxel is part of a diseased region of the body structure.
[0155] In some embodiments, the flow of operations 400 includes determining one or more regions of interest of the three-dimensional image each including one or more voxels for which a value of each of the one or more voxels exceeds a threshold in operation 410. The ROI identifier 216 may be configured to identify regions of interest for further investigation. The ROI identifier may compare the probabilities of the voxel-wise classification to a threshold and identify a region of interest anytime there are more than Nv connected voxels that exceed the probability threshold or a square or cube of size Ns for which all the voxels exceed the probability threshold.
[0156] In some embodiments, the ROI identifier 216 is configured to average a number of voxels together. The average may be a standard average or the average may be weighted based on a distance from the center of the region. Averaging the voxel-wise classification results may smooth the probabilities reducing the effect of any one voxel and decreasing the number of false alarms that a passed to second stage processing.
[0157] In some embodiments, the ROI identifier 216 is configured to identify regions of interest in units of more than one voxel. For example, averaging probability scores may be computed for 8 by 8 regions of voxels (using the probabilities of voxel-wise identification of the voxels of the 8 by 8 region, voxels surrounding the 8 by 8 region or both). The average probability score associated with the region may be compared to a threshold to determine if the entire region should be included in a region of interest.
[0158] In some embodiments, the ROI identifier 216 includes a threshold determined the diseased structure identification training system 250. The threshold may be chosen to have a particular effect on the true positive rate and the false positive rate. For example, the threshold may be chosen to set the false positive rate at 5% or to set the true positive rate to 90% against the training set or an anticipated future set. The threshold may be chosen to minimize a false positive rate while constraining the true positive rate to satisfy a criteria (e.g., be greater than a certain value).
[0159] In some embodiments, the flow of operations 400 includes executing, for each of the one or more regions of interest, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual in operation 412. The stage 2 classifier 220 may perform operation 412. The stage 2 classifier 220 may be configured to use any of statistical features of the voxel-wise classification probabilities (e.g., standard deviation of a local region, averages, etc.), the original features of the voxels making up the region of interest, and/or radiomic features (e.g., visual-based radiomic features such as first order statistics, gray level co-occurrence matrix, etc.). Advantageously, the difficult to determine radiomic features are only used on a few regions of interest rather than on the entire body structure. In some embodiments, additional features can be obtained by training a small CNN to identify features from an expanded receptive field around each ROI. For example, the CNN can use the probabilities from the voxel-wise classifier 222, a section of the original image, or any features used by the voxel-wise classifier as input.
[0160] FIG. 11 is a block diagram illustrating the data flow in the diseased structure identification system 200 according to some embodiments. In some embodiments, a number of three-dimensional images (e.g., three-dimensional images 602a-c) are obtained. Each of the three-dimensional images may be of a different imaging type (e.g., modality). For example, the three-dimensional image 602a may be captured using the T2w mode of an MRI, the three-dimensional image 602b may be captured using the ADC mode of an MRI, and the three-dimensional image 602c may be captured using the DWI mode of an MRI. The three- dimensional images 602a-c may be aligned so a voxel location of one three-dimensional image represents the same area of the body structure scanned in each of the three-dimensional image.
[0161] In some embodiments, the operations are performed for each voxel (e.g., each voxel is classified: a determination is made if the voxel represents a part of a diseased region of the body structure). A set of adjacent voxels including the voxel that is to be classified is selected from same area of each of the three-dimensional images and input to the feature extractor 212. The feature extractor 212 may generate a number of feature maps of various resolutions from the set of adjacent pixels. For example, a 24 by 24 set of adjacent pixels may be selected and provided to the feature extractor 212 and the feature extractor may generate a 22 by 22 by 9 feature map and a 9 by 9 by 56 feature map along with global PCA features for a total of approximately 9,000 features for each imaging type that can be used to classify a single voxel. The feature extraction operation performed by feature extractor 212 is described in more detail herein with reference to FIG. 12.
[0162] In some embodiments, the same feature extraction method is performed by the feature extractor 212 to the set of adjacent voxels of each of the imaging types. For example, the feature extractor 212 for each set of adjacent voxels may calculate a pixel embedding using equation 3. However, each imaging type may have a different set of anchor vectors defining the feature extraction operation performed by the feature extractor 212.
[0163] The large number of features related to the set of adjacent pixels may not all contain a significant amount of discrimination potential. The feature selector 214 may choose a subset of the features to use to classify each voxel. The subset of features chosen by the feature selector 214 may be determined during training. All features may be ordered based on a discrimination metric. For example, the Discriminant Feature Test (DFT) can be used as a metric associated with each feature’s potential to discriminate between a diseased region and a healthy region of the body structure. In some embodiments, the feature selector 214 is configured to select a fixed number (e.g., 500, 1000) of features with the highest discrimination potential. In some embodiments, the number of selected features is determined during training based on the discrimination metric.
[0164] In some embodiments, the selected features from each of the imaging types are combined into a single set by the concatenator 124. The voxel represented by the set of adjacent voxels may be classified by the voxel classifier 222 using the concatenated features. The operations of selecting a set of adjacent voxels, performing feature extraction, feature selection, and classification may be repeated for each voxel location of the three-dimensional images 602a-c. For example, if each imaging type has a resolution of 512 by 512 by 128 slices, the classification may be repeated for up to as many as 33.5 million voxels to generate a probability (e.g., a soft classification) that the voxel is part of a diseased portion of the body structure.
[0165] In some embodiments, the ROI identifier 216 determines regions of interest where a number of nearby (e.g., adjacent voxels) have elevated probabilities (e.g., exceed a threshold). For example, the ROI identifier may determine clusters of elevated probabilities using rulebased algorithms and/or machine learning algorithms such as density-based spatial clustering of applications with noise (DBSCAN), Gaussian mixture models (GMMs), and/or ordering points to identify the clustering structure (OPTICS). The ROI identifier 216 may eliminate false alarms that could be generated by any individual voxel having an associated high probability that the voxel belongs to a diseased region of the body structure. The regions of interest may be passed to the region selector 218. The region selector 218 may be configured to cause the collection of features for a second stage of classification. For example, the region selector 218 may obtain the features used by the voxel classifier 222 for voxels near the center of mass of the selected ROI. The region selector 218 may also be configured to cause the anomaly feature collector 224 and/or the radiomics feature collector 226 to acquire additional features for the stage 2 classification.
[0166] In some embodiments, voxels averaged together by the ROI identifier 216 are used as features in the stage 2 classifier 220. The averaged voxels may be acquired by the anomaly feature collector 224. For example, a weighted average of voxel classifications (e.g., probabilities) may be input as features to the stage 2 classifier 220. Anomaly features may also include the mean and standard deviation of the voxel probabilities and/or the weighted average of voxel probabilities within a region near (e.g., encompassing, partially overlapping, within, etc.) the region of interest. Using voxel probabilities averaged from several adjacent features may carry information from the voxel-wise classification results while allowing the stage 2 classifier to give weight to those regions that had more voxels classified as being from a diseased portion of the structure.
[0167] In some embodiments, additional visual-based features are extracted for the ROIs. For example, radiomics may provide several features for quantifying MRI characteristics of an ROI. For classification using radiomics features from all categories of radiomics are considered. For example, radiometric first order statistics may provide up to 19 features, 3D shape-based features include up to another 16 features, 2D shape-based features include up to 10 features, a gray level co-occurrence matrix may provide as many as 24 features, a gray level run length matrix can provide 16 features, a gray level size zone matrix can provide 16 features, a neighboring gray tone difference matrix can provide 5 features and/or the gray level dependence matrix may provide 14 features.
[0168] In some embodiments, additional features are extracted by a convolutional neural network (CNN) to identify features from an expanded receptive field around each ROI. The CNN can extract nonlinear features in addition to those obtained by performing the linear feature extraction techniques described herein. For example, the CNN can use the probabilities from the voxel-wise classifier 222, a section of the original image, or any features used by the voxel-wise classifier as input. Advantageously, because ROIs have been found using the voxel-wise classifier 222. A CNN for identifying features around a region of interest does not need to accept the entire image as input and thus may not contribute excessive computational effort to the overall process. In some embodiments, the CNN is trained simultaneously with the stage 2 classifier 220. For example, the CNN may be trained simultaneously with the stage 2 classifier 220 so that it does not relearn features for information already provided by the linear feature extraction techniques, the radiomic features, or any other features that may also be provided to the stage 2 classifier 220.
[0169] In some embodiments, the features used by the voxel classifier 222, the features collected by the anomaly feature collector 224, and the features collected by the radiomics feature collector 226 are combined by the concatenator 124. Another feature selection process may be performed by the feature selector 214 to select the discriminant features for the stage 2 classifier 220. The stage 2 classifier 220 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
[0170] The state 2 classifier 220 may provide a binary classification based on all the features provided including the those used by the voxel classifier 222, the anomaly features, and the radiomics. The stage 2 classifier 220 may be configured to output a binary decision that indicates whether each region of interest is in the diseased portion of the structure of the individual. The output 604 may be provided in several different forms. The output 604 may include a listing of the regions of interest identified as part of a diseased portion of the body structure. The output 604 may include a listing of the voxels included in the regions of interest identified as part of a diseased portion of the body structure. The output 604 may provide the voxels associated with any identified ROIs in the form of a mask that can be overlayed on the original three-dimensional image. For example, a mask overlay may be used by a surgeon during a biopsy, may guide further imaging, etc. In some embodiments, the output of the classifier is a probability that the ROI is indicative of a diseased body structure. The mask (e.g., the overlay) may be color-coded based on the probability. For example, a green overlay may indicate a ROI that was determined to be a false positive (e.g., low probability), a yellow overlay may indicate the ROI is questionable, and red overlay may indicate that the ROI was validated by the stage 2 classifier 220 (e.g., high probability).
[0171] In some embodiments, the diseased structure identification system 200 is configured to generate a user interface (e.g., by the identification coordinator 210). The identification coordinator 210 may send instructions to the client device 20 (e.g., JavaScript, cascading style sheets, etc.) to generate the user interface including the original images being segmented as well as the mask and/or color overlay. Additionally or alternatively, the identification coordinator 210 may include instructions to generate a local user interface on the computing hardware embodying the diseased structure identification system 200 (e.g., a desktop computer, a server, the imaging system 30, etc.). The user interface may be configured to provide user interaction (e.g., in addition to viewing the results). For example, if a radiologist is not fully satisfied with the identification result the radiologist can add voxels to a ROI and/or provide additional regions of interest, etc. In some embodiments, ROIs that were flagged, but ultimately labeled a false positive by the stage 2 classifier 220 may also be displayed in the user interface to provide additional data to the radiologist and receive feedback. Adjustments by the radiologist can improve results by providing additional training data that has been expertly labeled and the diseased structure identification system 200 was unable to predict. The new training data may be stored in the diseased structure identification training system 250.
[0172] In some embodiments, the regions of interest may be included in a number of visual representations. One or more visual representations can be delivered to the user interface. For example, the regions of interest can be viewed as a color overlay on a screen, a virtual object in an augmented reality environment, an object on a holographic display, etc.
[0173] FIG. 12 is a flow of operations performed by the feature extractor 212 according to some embodiments. FIG. 13 shows the flow of data as the feature extractor 212 processes a set of adjacent pixels according to some embodiments. The description of the feature extractor 212 will refer to both FIGS. 12 and 13 to provide details of the operations.
[0174] FIG. 12 shows flow of operations 420 for performing feature extraction according to some embodiments. In some embodiments, the flow of operations 420 includes obtaining a three-dimensional image, the three-dimensional image depicting a structure of an individual and including a plurality of voxels in operation 422. More than one three-dimensional image may be used. The feature extraction algorithm may be performed on any number of related three-dimensional images (e.g., T2w, ADC, DWI modes from an MRI). In some embodiments, the feature extraction operations are different for each of the three-dimensional images. The flow of operations 420 may continue with identifying a set of adjacent voxels, including a voxel for classification, of the plurality of voxels. The diseased structure identification system 200 may perform vox el -wise classification, the set of adjacent voxels can provide the information used to make a determination if the voxel for classification is a portion of a diseased structure. With reference to FIG. 13, the voxel for classification 608 is shown near the center of the set of adjacent voxels 606. The set of adjacent voxels 606 are provided to the feature extractor 212 and may be of any size. For example, in some embodiments, the set of adjacent voxels is square and 24 by 24.
[0175] In some embodiments, the flow of operations 420 includes vectorizing a local region to generate a set of vectors for each local region of the adjacent set of voxels in operation 426. The flow of operations may also include applying a feature extraction operation to each local region to generate a feature map, the feature extraction operation may be determined by a Saab transformation trained using local regions of three-dimensional training images in operation 428. With reference to FIG. 13 a local region 610 is shown. The local region may be of any size, 3 by 3, for example. The local region may be scanned across the set of adjacent voxels 606 and vectorized to generate a set of vectors 612 that may be used by the diseased structure identification training system 250 to generate the feature extraction operations (e.g., similar to the training performed by the body structure segmentation training system 150). For example, the set of vectors 612 may be combined with other vectors from other patients to determine the feature extraction operations of the Saab transformation.
[0176] Training of the Saab transformation may provide a set of anchor vectors to determine features from each of the local regions 610. The feature extractor 212 may apply a feature extraction operator 230 to each local region of the set of adjacent voxels 606. For example, the feature extraction operator 230 may be represented by the anchor vectors arranged in a matrix as shown in equation 3. For each local region 610, the feature extraction operator 230 generates a pixel embedding of a number of features. In some embodiments, the pixel embeddings are arranged in a three-dimensional array (e.g., feature map 616) so that the x-y location of the pixel embedding corresponds with the center of the local region used to generate the pixel embedding and the depth of the three-dimensional array is the number of features of each pixel embedding.
[0177] In some embodiments, the flow of operations 420 includes performing a pooling operation to each slice of the feature maps to generate a feature map of a lower resolution in operation 430. For example, the operation 430 may convert the feature map 616 (e.g., a feature map has similar resolution to the set of adjacent voxels) to the lower resolution feature map 618. In some embodiments, pooling refers to combining a number of adjacent voxels (or pixels) into a single voxel in a non-overlapping manner, thus reducing the overall resolution of the output. For example, the voxels may be combined by taking the maximum of the pixels, by taking the average of the pixels, or any other suitable function of values of multiple voxels. In some embodiments, two-by-two patches of voxels in a slice may be combined by pooling resulting in an output resolution that is half that of the input resolution. In some embodiments, three-by-three patches of voxels are combined resulting in an output resolution that is a third of the input. In some embodiments, rectangular patches are combined in the pooling operation.
[0178] In some embodiments, the flow of operations 420 includes applying a feature extraction operation to each local region of each feature slice of the lower resolution feature map 618 to generate a second feature map, the feature extraction operations determined by a Saab transformation trained using local patches of the same feature in operation 432. With reference to FIG. 13, a feature extraction operator 232 may be applied to local regions of each slice (e.g., channel) of the lower resolution feature map 618. The feature extraction operator 232 may be represented by different anchor vectors (e.g., a different matrix for each of the slices). Additionally, each slice may provide a different number of features to the output pixel embedding (e.g., the number of anchor vectors may be different for each feature extraction operation).
[0179] A slice of the lower resolution feature map 618, after a feature extraction operation is applied to each local region, may generate a portion of the second feature map. For example, the first slice of the lower resolution feature map 618 may generate the portion 624a and the fourth portion of the feature map 618 may generate another portion of the second feature map 624b. All portions of the second feature map generated from slices of the lower resolution feature map 618 can be concatenated to form the second feature map 626.
[0180] Local regions of the lower resolution feature map 618 are also used to generate vector set 620 that can be used to train the feature extraction operator used for each of the slices (e.g., channels). The set of vectors 620 generated may be used by the diseased structure identification training system 250 to generate the feature extraction operations (e.g., similar to the training performed by the body structure segmentation training system 150). For example, the set of vectors 620 may be combined with other vectors from other images (e.g., other individuals) to determine the feature extraction operations (e.g., anchor vectors) of the Saab transformation. In some embodiments, a feature extraction operation is determined for each slice (e.g., channel) of the lower resolution feature map 618 and the feature vectors 620 are grouped by which slice was used to generate them so that those vectors can be combined with vectors from the same slice of a different set of adjacent voxels 606 and/or a different three-dimensional image 602.
[0181] In some embodiments, the flow of operations 420 includes generating spectra including a spectrum for each slice of the first map and each slice of the second feature map in operation 434. For example, performing a spectral PCA. With reference to FIG. 13, the PCA calculator 186 is used to generate global features for each slice of the feature maps 616 and 626. For example, the PCA calculator 186 may be used to perform eigenvalue decomposition on the slice and the spectrum may be used as features for the voxel-wise classification.
[0182] In some embodiments, the flow of operations 420 includes flattening the first feature maps and the second feature map and concatenating the flattened feature maps with the spectra in operation 436. The flattener 234 may vectorize the feature maps 616 and 626 so that the features may be used together by the voxel-wise classifier 220.
[0183] In some embodiments, the set of adjacent voxels 606 selected for classification from the three-dimensional image is 24 by 24, the local region used to input to the feature extraction operator 230 is 3 by 3, feature extraction is performed without zero padding (and therefore local regions centered on the edge voxels are not used) resulting in a first feature map 616 that is 22 by 22 by 9, the size P of the pooling operation is two resulting a lower resolution feature map 618 that is 11 by 11 by 9, the feature extraction operator 232 uses the same size local region (e.g., 3 by 3) on all slices (e.g. channels) resulting in a second feature map 626 that is 9 by 9 by 56. In some embodiments, the functionality described herein uses different parameters resulting in differently sized feature maps and a different number of features for voxel-wise classification. In some embodiments, more than two feature maps of differing resolutions are created (e.g., three or four).
Diseased Structure Identification Training System
[0184] FIG. 14 shows another block diagram of the imaging and disease identification system 12 according to some embodiments. FIG. 14 shows some embodiments of an implementation of the diseased structure identification training system 250. The diseased structure identification training system 250 may include a communications interface 252 to facilitate communication of data (e.g., information, images, etc.) to other devices and/or systems on the network 40. For example, the communications interface 252 may be used to communicate the trained machine learning models (e.g., the feature extraction models, and the classifiers) to the diseased structure identification system 200. The communications interface 252 may be of the same or similar type of the communications interface 202 of the diseased structure identification system 200. The diseased structure identification training system 250 may also include a processing circuit 254 having a processor 256 and memory 258. For example, the processor 256 may be configured to execute instructions contained on the memory 258. In some embodiments, the diseased structure identification training system 250 and the diseased structure identification system 200 are implemented on the same hardware (e.g., on hardware of the imaging system 30, on the same computer, etc.). For example, the processing circuit 254 may be the same processing circuit as the processing circuit 204, similarly the communications interface, memory and processors may be shared between the diseased structure identification system 200 and the diseased structure identification training system 250.
[0185] In some embodiments, the diseased structure identification training system 250 may be implemented on separate hardware from the diseased structure identification system 200. The processor 256 may be of similar type or of different type from the processor 206. The processor 256 may be one or more of general purpose or specific purpose processors, application specific integrated circuits (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components, or other suitable processing components. The processor 256 may be configured to execute computer code and/or instructions stored in the memory 258 or received from other computer readable media (e.g., CDROM, network storage, a remote server, etc.). The processor 256 may be configured in various computer architectures, such as graphics processing units (GPUs), distributed computing architectures, cloud server architectures, client-server architectures, or various combinations thereof. One or more first processors can be implemented by a first device, such as an edge device, and one or more second processors can be implemented by a second device, such as a server or other device that is communicatively coupled with the first device and may have greater processor and/or memory resources.
[0186] The memory 258 may include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and/or computer code for completing and/or facilitating the various processes described in the present disclosure. The memory 258 may include random access memory (RAM), read-only memory (ROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and/or computer instructions. The memory 258 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in the present disclosure. The memories may be communicably connected to the processors and can include computer code for executing (e.g., by the processors) one or more processes described herein.
[0187] According to some embodiments, the diseased structure identification training system 250 is to acquire several three-dimensional images of a portion of the body for training the feature extractor 212 (e.g., the feature extraction operations), the voxel classifier 222, and the stage 2 classifier 220 of the diseased structure identification system 200. In some embodiments, the diseased structure identification training system 250 applies unsupervised learning to identify features of interest for the voxel classifier 222 and the stage 2 classifier 220. For example, the feature extractor 212 may be trained to determine feature maps of various resolutions to be used by the voxel classifier 222 and the stage 2 classifier 220. In some embodiments, the diseased structure identification training system 250 applies a supervised learning to train the voxel classifier 222 to make a voxel-wise assessment (e.g., determination) of the probability that the voxel depicts a portion of a diseased body structure. In some embodiments, the diseased structure identification training system 250 applies a supervised learning to train the stage 2 classifier 220 to validate regions of interest (ROIs) indicated by several collocated voxels. For example, the region 2 classifier 220 may be trained to determine if a region of interest is a false positive or is in fact a diseased region of the body structure. It is noted that because the training the feature extractor 212 is performed with unsupervised learning it does not require images for which diseased areas are previously labeled for training. Advantageously, training using unsupervised learning allows unlabeled data to be used for training and greatly increases the available training data. Images for which the diseased areas are labeled may be provided to train the voxel classifier 222 and the stage 2 classifier 220. The segmented images are additionally used to train the feature extractor 212 (e.g., without using the labels) in some embodiments.
[0188] In some embodiments, the diseased structure identification training system 250 applies unsupervised learning to train the voxel classifier 222 and the stage 2 classifier 220. For example, clustering algorithms may generate classification rules (e.g., boundaries, decisions, etc.) in the feature space that can be used to determine if a voxel or region of interest depicts a portion of a diseased body structure. For example, the features of an area including a diseased body structure may naturally form a distinct cluster from those of healthy tissue.
[0189] In some embodiments, the diseased structure identification training system 250 includes an identification training coordinator 260, the feature extractor 212, training data storage 266, a data selector 268, the voxel classifier 222, the stage 2 classifier 220, a loss calculator 272, a parameter tuner 274, and a Saab transformer 180. In addition, several other components of the body structure segmentation system 100 or the diseased structure identification system 200 may also be included in the diseased structure identification training system 250 (e.g., the pooling calculator 116, the concatenator 124, etc.) and used to execute the feature extraction or classification during training of the voxel classifier 222, and the stage 2 classifier 220. The identification training coordinator 260 may be configured to control the timing and flow of data through the other circuitry of the diseased structure identification training system 250. For example, the identification training coordinator 260 may cause the instructions or circuits to execute in a specific order to perform the function of the diseased structure identification training system 250. In some embodiments, the identification training coordinator 260 routes the information and/or outputs of other instructions that are dependent on the information or use the information as an input.
[0190] In some embodiments, the training data storage 266 is configured to maintain (e.g., store, save, acquire, etc.) three-dimensional images of body structures that can be used to train the feature extractor 112, the voxel classifier 222, and the stage 2 classifier 220, or both. Three-dimensional images for which diseased areas have not been labeled can be used to perform unsupervised learning (e.g., of the feature extractor or some classification algorithms). Three-dimensional images of a body structure that has been previously labeled may be used to perform both unsupervised learning and supervised learning. In some embodiments, the data selector 268 selects the correct data to be used for a particular training algorithm. For example, the data selector 268 may select only labeled data when supervised learning is being performed to train the voxel classifier 222, and the stage 2 classifier 220. The data selector 268 may also split data into training and validation sets. The validation data may be used to stop algorithm training, choose a classification algorithm, or to determine appropriate values of hyperparameters (e.g., the number of features that should be used, maximum depth of a decision tree, regularization terms, etc.).
[0191] In some embodiments, the diseased structure identification training system 250 includes the Saab transformer 180. The Saab transformer 180 may be used to determine the feature extraction operations of the feature extractor 212. For example, the Saab transformer 180 may be used to determine the anchor vectors used in the feature extraction operations, the number of anchor vectors to use, and/or the matrix representation of the feature extraction operations (e.g., those used in feature extraction operation 230 and/or 232). In some embodiments, the Saab transformer 180 includes a sampler 182, a DC vector calculator 184, and a principal components analysis module (PC A) 186 to perform the operations of the Saab transform.
[0192] The Saab transformer may receive a number of training samples from the training data storage 266. The training samples, for example, may include three-dimensional images of a body structure (e.g., three-dimensional image 602). The sampler 182 may determine local regions of the three-dimensional image to be used to train the feature extraction operation. The sampler 182 may choose the local regions based on one or more criteria for filtering appropriate local regions. During classification of diseased regions of a body structure the feature extraction operations may be applied to local regions that satisfy the same or similar criteria. For example, the sampler 182 may select data based on the slice (e.g., channel) of a feature map, and train different feature extraction operations for various slices. The sampler may also disqualify regions, slices, or whole three-dimensional images based on one or more criteria. For example, a three-dimensional image may be disqualified from use in training if the signal to noise ratio is high, a slice may be disqualified from use in training if the image is subject to a blur, or any other undesirable image artifact may cause the image or slice to be disqualified. The process for selecting anchor vectors during the Saab transformation is described in more detail in the description of the body structure segmentation training system 150.
[0193] In some embodiments, multiple models may be trained based on criteria related to the three-dimensional image. For example, different feature extraction operations may be learned for three-dimensional images of differing noise levels.
[0194] In some embodiments, features are selected by the feature selector 270 based on a Discriminant Feature Test (DFT) to quantify the discriminant power of features. For a given feature the minimum and maximum of the feature may be calculated, denoted by fmin and fmax, respectively. The partition interval [fmin, fmax] may split into B bins in a uniform way. The edges of the bins can be used to define candidate thresholds for the evaluation of discrimination power of the feature. The splitting quality of each threshold may be evaluated using the weighted entropy:
Figure imgf000064_0001
where N+ is the number of samples for which the feature is to the left of the threshold, N_ is the number of samples for which the feature is to the right of the threshold, tb are the candidate thresholds and choosing the minimum loss Lf tb over all threshold values. The individual loss function (e.g., to calculate Lf t and Lf t +) is the entropy value, calculated by: , equation 8
Figure imgf000065_0001
[0195] The DFT may be calculated for each feature by the feature selector 270 during training and a number of features can be kept (e.g., those with the lowest loss). For example, the 1000 features which demonstrate the greatest discrimination potential may be kept. In some embodiments, the loss is plotted in ascending order and the elbow point in the plot is used to determine the number of features kept to train the voxel classifier 222 and/or the stage 2 classifier 220.
[0196] After features have been determined and calculated for the training data in training data storage 266 it is possible to train the voxel classifier 222 and the stage 2 classifier 220. In some embodiments, training of a classifier (e.g., the voxel classifier 222 and the stage 2 classifier 220) is performed iteratively. For example, the classifier may determine a probability that the features represent a diseased portion of the body structure. The loss calculator 272 may compare the probability calculated by the classifier to the label for the voxel (e.g., the voxel from the set of adjacent voxels) that is being classified. For example, the loss calculator 272 may compute the sum of the loss by comparing several pixel embeddings to the respective labels to compute a loss function. Based on the loss function the parameter tuner 274 may update the classifier (e.g., parameters of the classifier, rules within the classifier, etc.) to improve the loss function with respect to the training data. The classifiers 222 may be any type of machine learning model including a neural network (e.g., a multilayer perceptron), a decision tree, or multiple decision trees (e.g., xgboost).
[0197] In some embodiments, the training data includes numerous negative areas (e.g., areas that do not depict a diseased portion) of a body structure, and a small number of regions for which the disease is confirmed (positive areas). To mitigate this imbalance in the training data, training is performed using a two-step process. In the first step, the classifier may be trained using random region sampling from the training data. However, areas with positive areas may be oversampled. For example, by sampling with more overlap and/or by causing the random sampler to choose positive samples more often. After training in the first step, the trained classifier is used to determine the difficulty of classifying a particular training sample. Training samples can then be chosen from across the various difficulty bins in order to cause the training data to approximate a desired distribution (e.g., equally training on difficult, medium, and easy to identify training samples). Experiment Results Diseased Structure Identification System
[0198] Experiments were conducted to demonstrate the effectiveness of the disclosed methodology. The experiments were based on currently the largest publicly available dataset, PI-CAI, which includes bpMRI data of a cohort of 1,500 patients. For the testing phase, a hidden cohort of 1000 patients is also available. Training data includes MRI scans across multiple clinics, such as the Radboud University Medical Center (RUMC), University Medical Centet Groningen (UMCG) and Ziekenhuis Groep Twente (ZGT). Scans from those medical centers are included both in training and testing. MRI scans from the Norwegian University of Science and Technology (NTNU) is also included in the testing cohort, as an unseen institute data to training. The vendors of MRI scanners were Siemens Healthineers and Philips Medical Systems. The training data was divided into 1075 samples of benign or indolent PCa and 425 clinically significant prostate cancer (csPCa) cases.
[0199] For measuring the performance of the proposed pipeline, as well as benchmarking with other existing works, the same metrics as in the PI-CAI challenge were adopted. Two metrics were used to evaluate the methods performance in detecting csPCa. For lesion-level detection the average precision (AP), where each segmented lesion has a floating-point probability for being clinically significant is used. For patient-level performance, Area Under Receiver Operating Characteristic (AUROC) is used, predicting for each patient a probability of having csPCa. For PI-CAI challenge the average between the two metrics were used to rank different methods. For the per-lesion performance, to consider a hit with the TP lesion, the prediction must have an Intersection over Union (loU) at least 0.1.
[0200] For benchmarking with other methods, the online hidden validation and testing cohort was used. For ablation study and intermediate results between stage 1 and 2, another local validation cohort was randomly chosen, consisting of 300 patients. For a fair comparison, we report results with baseline methods from PI-CAI challenge trained solely based on the PI- CAI publicly available cohort, without using external private datasets. The performance of the methods against both metrics and comparisons are available in Table 5.
Figure imgf000067_0001
Table 5: Comparison to deep learning models based on 1000 patients from the PI-CAI challenge
[0201] RadHop including only stage-1 predictions using patch size of 32x32 (i.e. no refinement by stage- 2), demonstrated a competitive standing among other methods using deep convolutional neural network (DCNN) architectures, such as U-Net or SwinTransformer. The evaluation in such a large patient cohort can approach the actual performance when the algorithm of the present disclosure is used in a real clinical setting. Both the patient level prediction metric (i.e. AUROC) and lesion-level precision (i.e. AP), surpass the performance of nnDetection and SwinTransformer models. In patient-level, RADHop has a better performance than every other method under comparison, excect U-Net that has a slightly better performance. In lesion level metric, RADHop achieves a higher precision than other methods but there is still some gap with the U-Net related ones. However, with the stage 2 classifier to refine the heatmap predictions of stage- 1 classification may be further improved.
[0202] Given the limited access to the hidden testing cohort of PI- CAI challenge, a short performance comparison was performed to demonstrate the effectiveness of stage-2 and compare stage-1 under different patch unit sizes. Table 6 shows some experiments carried out, trying to improve stage- 1 output. It is possible to see the effectiveness of using smaller patch size, since it can capture, more easily, parts of the lesion. Also, a larger patch size can induce noise, thereby affecting the precision performance, especially for medium-sized and smaller lesion sizes. Stage-2 effectiveness and potentials are also demonstrated. In particular, AP (per lesion metric) improves from 0.356 to 0.374. This is because in stage- 2 it is possible to reduce the probability of some false positives by reclassification using the anomaly score features.
Figure imgf000068_0001
Table 6: Comparison of various configurations of the current application
[0203] An important aspect of models comparison, besides performance, is the model size and the Floating Point Operations (FLOPS). These two factors are consequential when algorithms are implemented on certain platforms, especially with limited computation resources. Table 7 shows the comparisons with nnU-Met and U-Net architectures (popular choices among several methods in literature for PCa detection). One can see the tiny model size of RADHop when compared to DLbased architecture. Also, the computational complexity measured in FLOPS is order of magnitudes less. Specifically, nnU-Net requires 79 times the number operations require by RADHop for inference. Simple U-Net requires 34 times the number of operations. These findings highlight the tremendous deployment advantages of present disclosure method when compared with DL models. Also, this stresses the ability for the present disclosure to offer lightweight and environment friendly solutions. Additionally, RADHop has a transparent feedforward pipeline, where each module can be explained to physicians.
Figure imgf000068_0002
Table 7: Comparison of model size and computation complexity
[0204] Instructions, modules, portions of memory, etc. described as configured to perform a function (or described as performing the function) may include embodiments for which the module is configured to cause the performance of the function (or is causing the performance of the function). Similarly, instructions, modules, portions of memory, etc. described as configured to cause the performance of a function (or described as causing the performance of a function) may include embodiments for which the module is configured to perform the function (or is performing the function).
[0205] While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order. The separation of various system components does not require separation in all implementations, and the described program components can be included in a single hardware or software product.
[0206] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. Any implementation disclosed herein may be combined with any other implementation or embodiment.
[0207] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items.
[0208] The foregoing implementations are illustrative rather than limiting of the described systems and methods. Scope of the systems and methods described herein is thus indicated by the appended claims, rather than the foregoing description, and changes that come within the meaning and range of equivalency of the claims are embraced therein.

Claims

WHAT IS CLAIMED IS:
1. A method, comprising: obtaining, by one or more processors, a slice of a three-dimensional image, the three- dimensional image depicting a structure of an individual and the slice comprising a plurality of pixels each corresponding to a different location of the three-dimensional image; executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps comprising a plurality of pixel embeddings each corresponding to a different location of the slice; executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask comprising values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual; concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution; executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual; and generating, by the one or more processors, a second mask indicating locations of the slice of the three-dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
2. The method of claim 1, further comprising: scaling, by the one or more processors, the first mask from the first resolution to the second resolution.
3. The method of claim 2, further comprising: scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation.
4. The method of claim 1, further comprising: training, by the one or more processors, the feature extraction model using an unsupervised learning technique.
5. The method of claim 1, wherein executing the feature extraction model comprises: executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels comprising a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings.
6. The method of claim 5, wherein executing the feature extraction model comprises: performing, by the one or more processors, principal components analysis on sets of adjacent pixels of a plurality of locations of the three-dimensional image.
7. The method of claim 1, wherein executing the feature extraction model comprises: executing, by the one or more processors, a transformation on a set of adjacent pixels encompassing a pixel of the plurality of pixels to generate the pixel embedding associated with the pixel.
8. The method of claim 7, wherein executing the feature extraction model comprises: pooling, by the one or more processors, adjacent pixel embeddings to generate a lower resolution feature map.
9. The method of claim 1, comprising: classifying, by the one or more processors, each concatenated pixel embedding based on whether the probability for the concatenated pixel embedding exceeds a threshold.
10. The method of claim 1, wherein obtaining the slice of a three-dimensional image comprises: receiving, by the one or more processors, the three-dimensional image of the individual; and segmenting, by the one or more processors, the three-dimensional image into a plurality of slices corresponding to different locations of the three-dimensional image.
11. The method of claim 1, wherein obtaining the slice of the three-dimensional image comprises obtaining the three-dimensional image depicting a prostate of the individual; wherein executing the second machine learning model to classify each concatenated pixel embedding comprises executing, by the one or more processors, the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual; and wherein generating the second mask comprises generating, by the one or more processors, the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
12. A system, comprising: one or more processors configured by computer-readable instructions to: obtain a slice of a three-dimensional image, the three-dimensional image depicting a structure of an individual and the slice comprising a plurality of pixels each corresponding to a different location of the three-dimensional image; execute a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps comprising a plurality of pixel embeddings each corresponding to a different location of the slice; execute a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask comprising values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual; concatenate each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution; execute a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual; and generate a second mask indicating locations of the slice of the three- dimensional image that depict the structure of the individual based on the classifications of the concatenated pixel embeddings.
13. The system of claim 12, wherein the one or more processors are further configured to: scale the first mask from the first resolution to the second resolution.
14. The system of claim 13, wherein the one or more processors are configured to: scale the first mask from the first resolution to the second resolution using trilinear interpolation.
15. The system of claim 12, wherein the one or more processors are further configured to: train the feature extraction model using an unsupervised learning technique.
16. The system of claim 12, wherein the one or more processors are further configured to: execute the feature extraction model to transform a set of adjacent pixels comprising a pixel of the plurality of pixels of the slice to a corresponding pixel embedding of the plurality of pixel embeddings.
17. The system of claim 12, wherein the one or more processors are configured to obtain the slice of the three-dimensional image by obtaining the three-dimensional image depicting a prostate of the individual; wherein the one or more processors are configured to execute the second machine learning model to classify each concatenated pixel embedding by executing the second machine learning model to generate probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the prostate of the individual; and wherein the one or more processors are configured to generate the second mask by generating the second mask to indicate locations of the slice of the three-dimensional image that depict the prostate of the individual based on the classifications of the concatenated pixel embeddings.
18. A method, comprising: obtaining, by one or more processors, a slice of a three-dimensional image, the three- dimensional image depicting a structure of an individual and the slice comprising a plurality of pixels each corresponding to a different location of the three-dimensional image; executing, by the one or more processors, a feature extraction model using the slice of the three-dimensional image to generate a plurality of feature maps of different resolutions, each feature map of the plurality of feature maps comprising a plurality of pixel embeddings each corresponding to a different location of the slice; executing, by the one or more processors, a first machine learning model using a first feature map of a first resolution to generate a first mask, the first mask comprising values indicating probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map depict at least a portion of the structure of the individual; concatenating, by the one or more processors, each pixel embedding of a second feature map of a second resolution with the first mask scaled to the second resolution, the second resolution higher than the first resolution; executing, by the one or more processors, a second machine learning model to classify each concatenated pixel embedding based at least on the values indicating the probabilities that corresponding pixel embeddings of the plurality of pixel embeddings of the first feature map of the first resolution depict at least a portion of the structure of the individual; and presenting, by the one or more processors, a visual representation of the classifications of the concatenated pixel embeddings.
19. The method of claim 18, further comprising: scaling, by the one or more processors, the first mask from the first resolution to the second resolution.
20. The method of claim 19, further comprising: scaling, by the one or more processors, the first mask from the first resolution to the second resolution using trilinear interpolation.
21. A method, comprising: obtaining, by one or more processors, a three-dimensional image, the three- dimensional image depicting a structure of an individual and comprising a plurality of voxels; for each voxel of the plurality of voxels: identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual; determining, by the one or more processors, one or more regions of interest of the three-dimensional image each comprising one or more voxels for which a value of each of the one or more voxels exceeds a threshold; and for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
22. The method of claim 21, further comprising: training, by the one or more processors, the feature extraction model using unsupervised learning.
23. The method of claim 21, wherein executing the feature extraction model comprises: executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings.
24. The method of claim 21, wherein executing the feature extraction model comprises: performing, by the one or more processors, spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
25. The method of claim 24, wherein executing the feature extraction model further comprises: combining, by the one or more processors, the features of the plurality of feature maps of different resolutions.
26. The method of claim 21, wherein determining the one or more regions of interest of the three-dimensional image comprises: identifying, by the one or more processors, clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
27. The method of claim 21, wherein determining the one or more regions of interest of the three-dimensional image comprises: identifying, by the one or more processors, a first set of clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a first threshold; and identifying, by the one or more processors, a second set of clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a second threshold.
28. The method of claim 27, further comprising: presenting, by the one or more processors, a first visual representation of the first set of clusters in a first color and a second visual representation of the second set of clusters in a second color.
29. The method of claim 21, further comprising: concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps, and wherein executing the second machine learning model comprises executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps as input.
30. The method of claim 21 further comprising: obtaining, by the one or more processors, anomaly features extracted from the three- dimensional image; and concatenating, by the one or more processors, the radiometric features of the region of interest with the one or more feature maps and the anomaly features, wherein executing the second machine learning model comprises executing, by the one or more processors, the second machine learning model using the concatenation of the radiometric features of the region of interest with the one or more feature maps and the anomaly features as input.
31. The method of claim 21, wherein the three-dimensional image is of a plurality of three-dimensional images of different imaging technologies; and wherein determining the one or more regions of interest of the three-dimensional image comprises determining, by the one or more processors, the one or more regions of interest based feature maps extracted from each of the plurality of three-dimensional images.
32. The method of claim 21, wherein obtaining the three-dimensional image comprises obtaining, by the one or more processors, the three-dimensional image depicting a prostate of the individual; wherein executing the first machine learning model comprises executing, by the one or more processors, the first machine learning model to generate the value indicating the probability that the voxel depicts a cancerous portion of a prostate of the individual; and wherein executing the second machine learning model comprises executing, by the one or more processors, the second machine learning model to generate the classification indicating whether the region of interest is the cancerous portion of the prostate of the individual.
33. A system, comprising: one or more processors configured by computer-readable instructions to: obtain a three-dimensional image, the three-dimensional image depicting a structure of an individual and comprising a plurality of voxels; for each voxel of the plurality of voxels: identify a set of adjacent voxels including the voxel of the plurality of voxels; execute a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and execute a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual; determine one or more regions of interest of the three-dimensional image each comprising one or more voxels for which a value of each of the one or more voxels exceeds a threshold; and for each of the one or more regions of interest, execute a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and radiometric features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
34. The system of claim 32, wherein the one or more processors are further configured to train the feature extraction model using unsupervised learning.
35. The system of claim 32, wherein the one or more processors are configured to execute the feature extraction model by: perform spectral principal components analysis on a feature map of the plurality of feature maps to generate features of the voxel of the plurality of voxels.
36. The system of claim 35, wherein the one or more processors are configured to execute the feature extraction model further by: combining the features of the plurality of feature maps of different resolutions.
37. The system of claim 32, wherein the one or more processors are configured to determine the one or more regions of interest of the three-dimensional image by: identifying clusters of voxels for which the value indicating the probability that the voxel depicts a diseased portion of the structure exceeds a threshold.
38. A method, comprising: obtaining, by one or more processors, a three-dimensional image, the three- dimensional image depicting a structure of an individual and comprising a plurality of voxels; for each voxel of the plurality of voxels: identifying, by the one or more processors, a set of adjacent voxels including the voxel of the plurality of voxels; executing, by the one or more processors, a feature extraction model using the set of the adjacent voxels to generate a plurality of feature maps of different resolutions; and executing, by the one or more processors, a first machine learning model using the plurality of feature maps to generate a value indicating a probability that the voxel depicts a diseased portion of the structure of the individual; determining, by the one or more processors, one or more regions of interest of the three-dimensional image each comprising one or more voxels for which a value of each of the one or more voxels exceeds a threshold; and for each of the one or more regions of interest, executing, by the one or more processors, a second machine learning model using one or more feature maps of the one or more voxels of the region of interest and anomaly features of the region of interest to generate a classification indicating whether the region of interest is the diseased portion of the structure of the individual.
39. The method of claim 38, further comprising: training, by the one or more processors, the feature extraction model using unsupervised learning.
40. The method of claim 39, wherein executing the feature extraction model comprises: executing, by the one or more processors, the feature extraction model to transform a set of adjacent pixels including a pixel of plurality of pixels to a corresponding pixel embedding of a plurality of pixel embeddings.
PCT/US2024/056293 2023-11-17 2024-11-15 Body structure segmentation using supervised and unsupervised learning Pending WO2025106929A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202363600352P 2023-11-17 2023-11-17
US202363600360P 2023-11-17 2023-11-17
US63/600,352 2023-11-17
US63/600,360 2023-11-17

Publications (1)

Publication Number Publication Date
WO2025106929A1 true WO2025106929A1 (en) 2025-05-22

Family

ID=95743447

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2024/056293 Pending WO2025106929A1 (en) 2023-11-17 2024-11-15 Body structure segmentation using supervised and unsupervised learning

Country Status (1)

Country Link
WO (1) WO2025106929A1 (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20220019810A1 (en) * 2020-07-14 2022-01-20 The Chamberlain Group, Inc. Object Monitoring System and Methods
US20230048231A1 (en) * 2021-08-11 2023-02-16 GE Precision Healthcare LLC Method and systems for aliasing artifact reduction in computed tomography imaging
US20230281961A1 (en) * 2022-03-07 2023-09-07 Hamidreza FAZLALI System and method for 3d object detection using multi-resolution features recovery using panoptic segmentation information
US20230362331A1 (en) * 2017-04-04 2023-11-09 Snap Inc. Generating an image mask using machine learning

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20230362331A1 (en) * 2017-04-04 2023-11-09 Snap Inc. Generating an image mask using machine learning
US20220019810A1 (en) * 2020-07-14 2022-01-20 The Chamberlain Group, Inc. Object Monitoring System and Methods
US20230048231A1 (en) * 2021-08-11 2023-02-16 GE Precision Healthcare LLC Method and systems for aliasing artifact reduction in computed tomography imaging
US20230281961A1 (en) * 2022-03-07 2023-09-07 Hamidreza FAZLALI System and method for 3d object detection using multi-resolution features recovery using panoptic segmentation information

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
vol. 47, 10 October 2019, SPRINGER, article REN YINHAO; ZHU ZHE; LI YINGZHOU; KONG DEHAN; HOU RUI; GRIMM LARS J.; MARKS JEFFERY R.; LO JOSEPH Y.: "Mask Embedding for Realistic High-Resolution Medical Image Synthesis", pages: 422 - 430, XP047522371, DOI: 10.1007/978-3-030-32226-7_47 *
ZHOU YUTING, LIN XIN, LUO SHI, DING SIXIAN, XIAO LUYANG, REN CHAO: "Multi-Scene Mask Detection Based on Multi-Scale Residual and Complementary Attention Mechanism", SENSORS, MDPI, CH, vol. 23, no. 21, CH , pages 8851, XP093317948, ISSN: 1424-8220, DOI: 10.3390/s23218851 *

Similar Documents

Publication Publication Date Title
US12361543B2 (en) Automated detection of tumors based on image processing
Iqbal et al. Deep learning model integrating features and novel classifiers fusion for brain tumor segmentation
Chato et al. Machine learning and deep learning techniques to predict overall survival of brain tumor patients using MRI images
US11379985B2 (en) System and computer-implemented method for segmenting an image
Hille et al. Joint liver and hepatic lesion segmentation in MRI using a hybrid CNN with transformer layers
Mendes et al. Lung CT image synthesis using GANs
Birenbaum et al. Multi-view longitudinal CNN for multiple sclerosis lesion segmentation
Guan et al. Breast cancer detection using transfer learning in convolutional neural networks
US20200160997A1 (en) Method for detection and diagnosis of lung and pancreatic cancers from imaging scans
Sadeghibakhi et al. Multiple sclerosis lesions segmentation using attention-based CNNs in FLAIR images
CN101517614A (en) Advanced computer-aided diagnosis of lung nodules
Zamanidoost et al. OMS-CNN: optimized multi-scale CNN for lung nodule detection based on faster R-CNN
Zhou et al. Automatic segmentation of 3D prostate MR images with iterative localization refinement
Dorgham et al. U-NetCTS: U-Net deep neural network for fully automatic segmentation of 3D CT DICOM volume
Bharathi et al. Combination of hand-crafted and unsupervised learned features for ischemic stroke lesion detection from Magnetic Resonance Images
Puch et al. Global planar convolutions for improved context aggregation in brain tumor segmentation
Islam et al. Fully convolutional network with hypercolumn features for brain tumor segmentation
El Badaoui et al. 3D CATBraTS: Channel attention transformer for brain tumour semantic segmentation
Zhou et al. HAUNet-3D: a novel hierarchical attention 3D UNet for lung nodule segmentation
Tripathi et al. A deep learning-oriented approach for lung ct augmentation: Leveraging u-net and gan architecture
Banerjee et al. A CADe system for gliomas in brain MRI using convolutional neural networks
Carmo et al. Extended 2D consensus hippocampus segmentation
Ganasala et al. Semiautomatic and automatic brain tumor segmentation methods: performance comparison
Prasanalakshmi et al. Automated classification and segmentation of brain tumor images using backboned resUNet framework
Benrabha et al. Automatic ROI detection and classification of the achilles tendon ultrasound images

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24892398

Country of ref document: EP

Kind code of ref document: A1