EP4713896A1 - Method for determining sub-visible particles in a test sample - Google Patents

Method for determining sub-visible particles in a test sample

Info

Publication number
EP4713896A1
EP4713896A1 EP24726284.3A EP24726284A EP4713896A1 EP 4713896 A1 EP4713896 A1 EP 4713896A1 EP 24726284 A EP24726284 A EP 24726284A EP 4713896 A1 EP4713896 A1 EP 4713896A1
Authority
EP
European Patent Office
Prior art keywords
sample
sub
images
image
clusters
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24726284.3A
Other languages
German (de)
French (fr)
Inventor
Abhijeet SATWEKAR
Puthan Veettil Sharfudheen
Nikila Varshini Easwaran
Varshini Kuppusami
Sriharsha SRIPADA
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ares Trading SA
Original Assignee
Ares Trading SA
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ares Trading SA filed Critical Ares Trading SA
Publication of EP4713896A1 publication Critical patent/EP4713896A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/60Type of objects
    • G06V20/69Microscopic objects, e.g. biological cells or cellular parts
    • G06V20/695Preprocessing, e.g. image segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/096Transfer learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/60Type of objects
    • G06V20/69Microscopic objects, e.g. biological cells or cellular parts
    • G06V20/698Matching; Classification

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • General Physics & Mathematics (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Evolutionary Computation (AREA)
  • Multimedia (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Artificial Intelligence (AREA)
  • Data Mining & Analysis (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Image Analysis (AREA)

Abstract

There is provided a method for determining particulate matter in an analytical sample. In particular the method can be used for quality control and batch release as part of a process for manufacturing biological active ingredients such as drug substance and drug product.

Description

METHOD FOR DETERMINING SUB-VISIBLE PARTICLES IN A TEST SAMPLE.
Field of the Invention.
The invention is directed to methods of determining sub-visible particles in a test sample, such as a sample from a manufacturing process for producing a biologic product.
Background
In manufacturing a biologic medicinal product the process includes a stepwise process which includes the isolation and purification of the biologic medicinal product (drug substance). In most such processes, the drug substance is provided in a solution. The drug substance is further prepared to reside in a stable formulation containing the drug substance (drug product). The drug product often times is also a formulated solution. These solutions containing the drug substance or drug product need to comply with certain quality standards, such as being free from certain amounts of particles, including sub-visible particles. These particles could for example be agglomerates of proteinaceous material or fibers.
Particle classification is a topic of high importance, as health authorities are seeking accuracy and consistency in particle classification. Particle classification is a tedious and time-consuming process when done manually. Further, when done manually it is more prone to errors in identifying such particles. Therefore, there is a need to improve accuracy in identifying the presence of particulate materials, including sub-visible particulate matter, in a biologic sample and there is a need to reduce the time to analyze such biologic sample in a manufacturing process for producing biologic medicinal products.
Summary of the Invention.
The present invention provides a solution to improve accuracy and time consumption for analyzing a biologic sample for the presence of particulate matter. An automated the process of particle classification based on a hybrid technique that uses deep learning and image processing is provided by the present invention. The method of the present invention includes the extraction of morphological features from images using a transfer learning based Convolutional Neural network (CNN) for example lnception-V3. These additional features made a significant impact in the model performance. Morphological features identified by the domain experts along with the features extracted from the deep learning model are being fed to a machine learning based classifier (Decision Tree). The particles can be classified into their respective classes with an accuracy of 88%. Through a user interface (Ul), users are given the provision to modify or correct any misclassifications made by the model. Once a significant amount of data has been collected the retraining pipeline can be triggered to improve the model performance.
In one embodiment the method provides a computer implemented method to determine sub-visible particles in a sample comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on image similarity feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, and f. Evaluating the statistical values to determine the distribution of sub-visible particles in a sample.
In another embodiment the method provides a computer implemented method for batch release based on subvisible particles in a sample from a batch comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample image, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on morphological feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, f. Comparing the statistical value for one sample with a reference sample, and g. Releasing the batch from which the sample is obtained if the sample is within a predetermined confidence level compared to the reference sample
In yet another embodiment there is provided a non-transitory computer readable medium comprising machine readable instructions arranged, when executed by one or more processors, to cause the one or more processors to carry out the method of the current invention.
Drawings
Figure 1. Morphological conditions to filter images into different classes.
Figure 2: Decision Tree Architecture
Figure 3. Retraining Pipeline
Figure 4: Application hosting in AWS.
Figure 5: Model Performance for Transfer learning (based on CNN) Figure 6: Model Performance for ensemble learning
Figure 7: Model Performance for Transfer learning (based on CNN) with deep learning features.
Figure 8: Model Performance for Random Forest
Figure 9: Model Performance for AdaBoost
Figure 10: Model Performance for K nearest Neighbor
Figure 11: Model Performance for Decision tree classifier
Figure 12: Comparison study between two experiments
Figure 13: Graphical Representation of Descriptive Statistics for Air and Oil
Fig 14: Distribution Comparison between Morpholical features for Air and Oil
Fig 15: Distribution of Circularity vs Aspect Ratio for all sub-visible particles
Fig 16: Visual Representation of Clusters from both experiments
Fig 17: Number of samples present in each cluster
Detailed Description.
The present invention provides a solution and improvement to the existing methods of determining the presence of particulate matter in samples that are analyzed with imaging technologies. In particular for samples that are obtained from manufacturing processes for the production of biologic medicinal products, containing for example drug substance or drug product. The methods herein provided are computer implement methods.
In one embodiment the method provides a computer implemented method to determine sub-visible particles in a sample comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on image similarity feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, and f. Evaluating the statistical values to determine the distribution of sub-visible particles in a sample. In another embodiment the method provides a computer implemented method for batch release based on subvisible particles in a sample from a batch comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample image, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on morphological feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, f. Comparing the statistical value for one sample with a reference sample, and g. Releasing the batch from which the sample is obtained if the sample is within a predetermined confidence level compared to the reference sample
The method of the present invention can be used for the sample that is selected from a sample that is a drug substance sample or a drug product sample.
Exemplary morphological features that can be extracted in the method of the present invention include ECD, area, perimeter, circularity, max feret diameter, aspect ratio, intensity, x- position, y-position, time (%), time (min) and any combination thereof.
In the methods of the present invention the images are classified in major categories, such major categories can be "air and oil", "dark-proteinaceous", "fibres" and "proteinaceaous". The method further can provide in each major category sub-clusters of images wherein the number of sub-clusters is from 2 to 10, such as in certain embodiments from 3 to 5 sub-clusters, or the number of sub-clusters in certain embodiments is 3.
A convolutional neural network can be used in the method of the present invention. Any such convolutional neural network, such as for example Inception V3 can be used.
In yet another embodiment there is provided a non-transitory computer readable medium comprising machine readable instructions arranged, when executed by one or more processors, to cause the one or more processors to carry out the method of the current invention.
Examples
The microscopic Flow Imaging (MFI) dataset belonging to different biopharmaceutical sub- visible particles like fibers, air and oil, proteinaceous, dark and proteinaceous were used for the study. Table 1: Datasets
Air and oil Fiber Proteinaceous Dark and
Proteinaceous
Number of 2173 2235 16363 5223 samples
Complexity YTF YTF YTF YTF
Training Set 1197 2037 11130 3876
Validation Set 462 111 462 185
Test Set 514 874 4771 1662
Data (Preparation, standardization, alignments, annotations, pre-processing)
The bio particle raw image data and morphological features obtained from microscopic flow image devices. The raw image obtained from the MFI device had no information regarding the classes (unlabeled). However, the report generated by the device consisted of certain morphological features of the particle. A python-based classification script was used to filter the images based on morphological conditions (figure 1). Later, the filtered images based on morphological conditions are validated by domain experts and checked for correctness. Manually validated and corrected data were used as ground truth for training deep learning based model.
The classification results obtained from the analysis done on initial morphological features were not reliable on all situations. Hence, a deep learning-based approach for feature extraction was used to obtain additional unique features for different bioparticles which are critical to improving the performance of Al-based particle classification. The more relevant the features, it is better to train the model for classification. A total of 2065 features were obtained that includes 17 features extracted from the MFI device and 2048 features extracted from the Deep learning model. This morphological feature information with respect to each image was stored in a .csv file. These features were used to train the classification model and hence improve performance.
The size of the images obtained was inconsistent. Thus, the MFI images were standardized and resized to 299 to ensure that there is no inconsistency across the dataset.
Deep Learning Model
A deep learning approach was explored to develop an artificial intelligence-based algorithm. The models were developed using Keras and TensorFlow framework, along with other libraries such as OpenCV, NumPy, scikit-learn, pandas, matplotlib etc. A Convolutional Neural Network (CNN) based architecture was used to extract features from the last but one layer of the architecture. A convolutional neural network is a feed-forward neural network that is generally used to analyze visual images by processing data with a grid-like topology. A CNN has multiple hidden layers that help in extracting information from an image. The four important layers in CNN are the convolution layer, ReLU layer, pooling layer, and fully connected layer.
In our study, inception_v3 architecture for extraction of features from images was considered. Total of 2065 features were extracted along with morphological features. Inception v3 is an image recognition model that has been shown to attain greater than 78.1% accuracy on the ImageNet dataset. These features improve the training and accuracy of the model.
Baseline Architecture
Inception V3
Several architectures were explored by iteratively testing and fine tuning the hyper parameters as showcased in Table 1. The performance of the various models was evaluated against a fixed subset of the data consisting of all the four sub-particles. Here, a transfer learning based CNN approach has been used to extract the features. Inception-v3 has been frequently applied in image recognition. The model is made up of symmetric and asymmetric building blocks, including convolutions, average pooling, max pooling, concatenations, dropouts, and fully connected layers. The number of parameters of Inception-v3 is fewer than half that of AlexNet (60,000,000) and fewer than one fourth that of VGGNet (140,000,000); additionally, the number of floating-point computations of the whole Inception-v3 network is approximately 5,000,000,000, which is much larger than that of Inception-vl (approximately 1,500,000,000). These characteristics make Inception-v3 more practical, that is, it can be easily implemented in a common server to provide a rapid response service. The image size input into Inception-v3 was 299 x 299. A total of 2048 features were extracted from the last but one layer of the network. The first layers of any neural network are basically responsible for identifying low-level features, such as edges, colors, and blobs, but the last layers are usually very specific to the task that's trained for.
Table 2: Parameters of baseline architecture
Layer Output Shape Number of Parameters kerasjayer (None, 2048) 21802784 dropout (None, 2048) 0
Dense (None, 4) 8196
Total params: 21,810,980
Trainable params: 21,776,548
Non-trainable params: 34,432 Decision Tree
A decision tree is a non-parametric supervised learning algorithm, which is utilized for both classification and regression tasks. It has a hierarchical, tree structure, which consists of a root node, branches, internal nodes and leaf nodes. Decision tree learning employs a divide and conquer strategy by conducting a greedy search to identify the optimal split points within a tree. This process of splitting is then repeated in a top-down, recursive manner until all, or the majority of records have been classified under specific class labels. To reduce complexity and prevent overfitting, pruning is usually employed; this is a process, which removes branches that split on features with low importance. The model's fit can then be evaluated through the process of cross-validation. While there are multiple ways to select the best attribute at each node, two methods, information gain, and Gini impurity, act as a popular splitting criterion for decision tree models. They help to evaluate the quality of each test condition and how well it will be able to classify samples into a class. Entropy is a concept that measures the impurity of the sample values. Its values can fall between 0 and 1. Information gain represents the difference in entropy before and after a split on a given attribute. The attribute with the highest information gain will produce the best split as it's doing the best job at classifying the training data according to its target classification. Gini impurity is the probability of incorrectly classifying random data points in the dataset if it were labeled based on the class distribution of the dataset. Here, 2065 features that combine the initial morphological along with the features obtained from the deep learning model is fed into the machine learning classifier.
Retraining pipeline
Machine learning and deep learning models are everywhere around us in modern organizations. Every industry has appropriate machine learning and deep learning applications, from banking to healthcare to education to manufacturing, construction, and beyond. One of the biggest challenges in all these ML and DL projects in different industries is model improvement. Continuous training is an aspect of machine learning operations that automatically and continuously retrains machine learning models to adapt to changes in the data before it is redeployed. As soon as the machine learning model is deployed in production, the performance of your model degrades. This is because the model is sensitive to changes in the real world, and user behavior keeps changing with time. Although all machine learning models decay, the speed of decay varies with time. This is mostly caused by data drift. Data drift (covariate shift) is a change in the statistical distribution of production data from the baseline data used to train or build the model. Thus, making it important to monitor and retrain the machine-learning models in production.
In this study, a retraining pipeline has been developed as shown in Figure.1 that provides the user the provision to select the data to be used to retrain the model. The data consists of images and a .csv file with information about the class of each image that is verified by a domain expert. Once the data for retraining is selected, the images are placed into the respective folders. Features are extracted using a deep learning-based approach. The features are then fed into a machine learning model for retraining purposes after which the model is evaluated against a test set. The performance of the model in terms of accuracy is returned once the retraining process is completed, which serves as a metric that helps the user decide whether to save or discard the model.
User interface and set-up in the cloud
The modular software application was built on a scalable and open architecture system comprising of an independent User Interface (Ul) module built on React. The server-side logic uses java technology stack with Sprint-boot, and an algorithm module built on python. These modules were integrated with REST APIs to allow seamless exchange within the internal services and the application was hosted on a AWS cloud environment. Ul module provided the accessibility of the users to the application by modern browsers (Chrome, Microsoft Explorer, Firefox etc). The Ul layer communication is driven securely over SSL72technology. The access was permitted through internet-facing URL and accessible to geo-fenced locations/countries. The authentication of users was enabled with SAML 2.0 (Security Assertion Markup Language 2.0) and SSO (Single-Sign-On) based access management. Authorization of the users was done through the IdP (Identity Provider) and OAuth (Open Authorization) technology. A root administrator was created to provide access to required user base by prior configuring their unique identification information (Id and email). User Interfaces built over React JS provided the following functionalities of file upload from browser, Ul-based validations, and display of results and reports. Data management and organization within the entire application was built on MySQL (Amazon RDS) as the database. Upload of the raw data files in *.zip folder format which consists of *.png and *.csv file formats was built and managed with Java/J2EE and Spring Boot as middleware. Scripts were set to check on data conformity of the uploaded files before importing them into the application. Amazon Elastic Compute Cloud (Amazon EC2) was used for building the Web application hosting, back-up and recovery. Storage of source data on application logs, model training and predictions was built using the AWS S3 (Simple Storage Service). The developed baseline architecture was deployed within the algorithm module using python as the programming language. This contained Artificial Intelligence, Machine learning and Neural network-based logic and services. NGINX was used for Web Server and Load balancing.
Concept Design and Project Execution
Understanding the datasets and exploration studies on Al
Exp 1: Transfer learning (based on CNN)
The pre-trained models like inception_v3, ResNetSO, ResNetlOl models, pre-trained weights, models and checkpoints were collected form tensorflow hub. The image data set is used directly to train the pre-trained models with all layers freezed (these hidden layers are not trained), but the last layer is only trainable. Resnet_vl_50 was the model used to build image classification model. The Training information is displayed in table 3. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 5.
Table 3: Training information Exp 2: Ensemble learning algorithms
Machine learning algorithms with morphological features Different ensemble learning algorithms such as decision tree, bagging - random forest, boosting - adaboost, gradient boost were developed for classification of data into 4 classes such as air and oil, Proteineceous, dark and proteineceous, fiber. Ensemble learning is a general meta approach to machine learning that seeks better predictive performance by combining the predictions from multiple models. These models are known as weak learners. The intuition is that when you combine several weak learners, they can become strong learners. Each weak learner is fitted on the training set and provides predictions obtained. The final prediction result is computed by combining the results from all the weak learners. Ensemble learning techniques have been proven to yield better performance on machine learning problems. The performance of the model is showcased in figure 6.
Exp 3: Transfer learning (based on CNN) with morphological and deep learning features
Transfer learning (based on CNN) with morphological features extracted from MFI device and additional deep learning features extracted using image processing techniques. The basic premise of transfer learning is simple: take a model trained on a large dataset and transfer its knowledge to a smaller dataset. The addition of other morphologiocal feature in deep learning model training showcased a improvement in model performance. The Training information is displayed in table 4. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 7.
Table 4: Training information
Classification of particles into different classes was explored on multiple machine learning algorithms like decision tree, random forest, Adaboost and K - nearest neighbor
Exp 4: Random Forest
Machine learning algorithm (Random Forest) with morphological features identified by domain experts and deep learning features extracted using image processing techniques. The random forest is a classification algorithm consisting of many decisions trees. It uses bagging and feature randomness when building each individual tree to try to create an uncorrelated forest of trees whose prediction by committee is more accurate than that of any individual tree. The Training information is displayed in table 4. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 8.
Exp 5: AdaBoost
Machine learning algorithm (AdaBoost) with morphological features identified by domain experts and deep learning features extracted using image processing techniques. AdaBoost algorithm, short for Adaptive Boosting, is a Boosting technique used as an Ensemble Method in Machine Learning. It is called Adaptive Boosting as the weights are re-assigned to each instance, with higher weights assigned to incorrectly classified instances. . The Training information is displayed in table 4. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 9.
Exp 6: K nearest Neighbor
Machine learning clustering algorithm K nearest Neighbor (KNN) with morphological features identified from MFI devices and deep learning features extracted using image processing techniques. KNN, is a non-parametric, supervised learning classifier, which uses proximity to make classifications or predictions about the grouping of an individual data point. The Training information is displayed in table 4. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 10.
Final Architecture
A deep learning-based approach (Inception V3) was implemented to extract additional features apart from morphological features obtained from MFI devices. A total of 2065 features that combines the morphological features and deep learning features were fed to machine learning classifier (Decision tree) to classify the sub visible particles into respective classes. The Training information is displayed in table 4. The classification report is scaled between 0 - 1. The performance of the model is showcased in figure 11 .
Comparison study
The presence of drug aggregates and sub-visible particles in therapeutic protein products has increasingly become a field of concern for both the pharmaceutical industry and regulatory agencies Aggregates in the micron range have been implicated in adverse reactions and/or reduction in the efficacy of therapeutic products. The sub-visible particles may vary in size, structure, and by many other features in different experiments. Different experiments include mechanical agitation (leading to exposure to hydrophobic air/water interfaces), chemical alteration, and/or temperature extremes. Protein aggregates can also be generated through protein nucleation around nano/micro-scale contaminants in the product such as silica particles shed from containers, fibers shed from filters, or metal particles shed from production equipment. Hence, a comparison study between the samples from two experiments helps researchers to identify new or unique samples and gain comprehensive knowledge. Here, the comparison study consists of two parts as shown in Figure 12.
Statistical Analysis
A statistical analysis has been done on the initial morphological features obtained. The statistical analysis takes morphological features in .csv format obtained from the MFI device for both experiments as input. The .csv files consists of the morphological features and inference results that are user corrected. It includes the classification results that gives the count of the sub-particles, quantity in percentage and its concentration with respect to the volume for each of the experiments as showcased in Table 5. A data table (evinced in Table 4) is generated that provides a descriptive statistic (Minimum, Maximum, Mean and Standard Deviation) of the morphological features (ECD, Area, Perimeter, Circularity, Max Feret Diameter and Aspect Ratio) for each of the sub-particles. Descriptive statistics are very important because by simply presenting the raw data it would be hard to visualize what the data was showing, especially if there was a lot of it. Descriptive statistics, therefore, enables to present the data in a more meaningful way, which allows simpler interpretation of the data. It also provides information on how much percent a study deviates from the other study in terms of these four metrics and highlights the top three deviations in each sub-particle as shown in Table ?.
Table 5. Classification Result for an experiment
Classification Results
Table 6. Descriptive Statistical Summary
Data Table Table 7: Deviation in descriptive statistics
Summary
| ±10% to ±30% HH 0%
A summary on particle size distribution for each sub-particle is generated that provides information regarding the count of particles with respect to size of the features in macrons. The size range fall under four categories (< 5mm, 5 - 10mm, >=10mm and >=25mm). The comparison study also provides a graphical representation of the descriptive statistics, the distribution of the morphological features, the comparison of distribution against different combinations of morphological features, and particle size distribution for each sub-particle and all together. A few of the comparisons are showcased in Figure 13,14 and 15.
Image Analysis
The image analysis includes clustering of the features extracted from images for each particle. Clustering of the images based the features and intensity helps to identify the any new or different molecules. It also helps to identify the outliers, or any misclassifications present in each of the sub- visible particles. It also visualizes the images belonging to each cluster and a graphical representation that gives a count on the number of samples present in each of the clusters. The image visualizations are as represented in Figure 16 and the graphical representations are showcased in Figure 17 Model performance and management
The Al models are built from the training data and validated on the test data. The validation of the models defines its performance in terms of its quality. Therefore, the CBA: Definition of criteria on model management and performance monitoring has a high relevance for the use of Al models in the routine.
Table 8: Model Performance Metrics

Claims

Claims
1. A computer implemented method to determine sub-visible particles in a sample comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on image similarity feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, and f. Evaluating the statistical values to determine the distribution of sub-visible particles in a sample.
2. A computer implemented method for batch release based on subvisible particles in a sample from a batch comprising: a. Obtaining one or more data files, such as raw data formats, and/or images of a sample from an analytical equipment, wherein the data files represent image information from an analyzed sample image, b. Extracting morphological feature information from the data files and/or images from the analyzed sample, c. Classifying each image of the analyzed sample using a trained convolutional neural network in major categories, d. Clustering of the images within the major categories into sub-clusters based on morphological feature information from the sample images, e. Applying a statistical analysis comparing within the major categories the sub-clusters on morphological features and/or image similarity to obtain a statistical value for one sample, f. Comparing the statistical value for one sample with a reference sample, and g. Releasing the batch from which the sample is obtained if the sample is within a predetermined confidence level compared to the reference sample.
3. The method of any of the preceding claims, wherein the sample is selected from a sample that is a drug substance sample or a drug product sample.
4. The method of any preceding claim, wherein the morphological feature information is selected from ECD, area, perimeter, circularity, max feret diameter, aspect ratio, intensity, x- position, y-position, time (%), time (min) and any combination thereof.
5. The method of any preceding claim, wherein the major categories are selected from "air and oil", "dark-proteinaceous", "fibres" and "proteinaceaous".
6. The method of any of the preceding claims, wherein the number of sub-clusters is from 2 to
7. The method of claim 6, wherein the number of sub-clusters is from 3 to 5.
8. The method of any of claims 6 or 7 , wherein the number of sub-clusters is 3.
9. The method of any of the preceding claims, wherein the convolutional neural network is Inception V3.
10. A non-transitory computer readable medium comprising machine readable instructions arranged, when executed by one or more processors, to cause the one or more processors to carry out the method of any of claims 1 to 9.
EP24726284.3A 2023-05-17 2024-05-17 Method for determining sub-visible particles in a test sample Pending EP4713896A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23174084.6A EP4465260A1 (en) 2023-05-17 2023-05-17 Method for determining sub-visible particles in a test sample
PCT/EP2024/063814 WO2024236195A1 (en) 2023-05-17 2024-05-17 Method for determining sub-visible particles in a test sample

Publications (1)

Publication Number Publication Date
EP4713896A1 true EP4713896A1 (en) 2026-03-25

Family

ID=86386736

Family Applications (2)

Application Number Title Priority Date Filing Date
EP23174084.6A Withdrawn EP4465260A1 (en) 2023-05-17 2023-05-17 Method for determining sub-visible particles in a test sample
EP24726284.3A Pending EP4713896A1 (en) 2023-05-17 2024-05-17 Method for determining sub-visible particles in a test sample

Family Applications Before (1)

Application Number Title Priority Date Filing Date
EP23174084.6A Withdrawn EP4465260A1 (en) 2023-05-17 2023-05-17 Method for determining sub-visible particles in a test sample

Country Status (5)

Country Link
EP (2) EP4465260A1 (en)
KR (1) KR20260006684A (en)
CN (1) CN121195288A (en)
AU (1) AU2024271024A1 (en)
WO (1) WO2024236195A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3766000A2 (en) * 2018-03-16 2021-01-20 The United States of America as represented by the Secretary of the Department of Health and Human Services Using machine learning and/or neural networks to validate stem cells and their derivatives for use in cell therapy, drug discovery, and diagnostics
US20210303818A1 (en) * 2018-07-31 2021-09-30 The Regents Of The University Of Colorado, A Body Corporate Systems And Methods For Applying Machine Learning to Analyze Microcopy Images in High-Throughput Systems
JP2024517592A (en) * 2021-04-09 2024-04-23 コリオリス ファーマ リサーチ ゲーエムベーハー FIM-CNN for detection of live cells and/or particulate impurities

Also Published As

Publication number Publication date
AU2024271024A1 (en) 2026-01-15
KR20260006684A (en) 2026-01-13
EP4465260A1 (en) 2024-11-20
WO2024236195A1 (en) 2024-11-21
CN121195288A (en) 2025-12-23

Similar Documents

Publication Publication Date Title
Oppong et al. A Novel Computer Vision Model for Medicinal Plant Identification Using Log‐Gabor Filters and Deep Learning Algorithms
Chen et al. Big data deep learning: challenges and perspectives
CN110929029A (en) Text classification method and system based on graph convolution neural network
Rath et al. A comparative analysis of SVM and ELM classification on software reliability prediction model
Chander et al. Data clustering using unsupervised machine learning
Liong et al. Automatic traditional Chinese painting classification: A benchmarking analysis
Hcini et al. Hyperparameter optimization in customized convolutional neural network for blood cells classification
Ahmed et al. Early detection of fetal health status based on cardiotocography using artificial intelligence
Guo et al. LWheatNet: a lightweight convolutional neural network with mixed attention mechanism for wheat seed classification
Lasso et al. Towards an alert system for coffee diseases and pests in a smart farming approach based on semi-supervised learning and graph similarity
Zhao et al. An efficient class-dependent learning label approach using feature selection to improve multi-label classification algorithms
KR20210062265A (en) System and method for predicting life cycle based on EMR data of companion animals
Lin et al. A neuronal morphology classification approach based on deep residual neural networks
Manoranjitham An artificial intelligence ensemble model for paddy leaf disease diagnosis utilizing deep transfer learning
EP4465260A1 (en) Method for determining sub-visible particles in a test sample
Manoharan et al. Ensemble Model for Educational Data Mining Based on Synthetic Minority Oversampling Technique
Pitz et al. Implementing clustering and classification approaches for big data with MATLAB
Zhang et al. VtNet: A neural network with variable importance assessment
Reddy et al. THE STUDY OF SUPERVISED CLASSIFICATION TECHNIQUES IN MACHINE LEARNING USING KERAS.
Perera Efficient algorithms to improve feature selection accuracy in high dimensional data.
Barrera-Hernandez et al. Towards abalone differentiation through machine learning
Solanki et al. Assimilate Machine Learning Algorithms in Big Data Analytics
Oppong et al. Research Article ANovelComputerVisionModelforMedicinalPlantIdentification Using Log-Gabor Filters and Deep Learning Algorithms
Santiago Sánchez et al. Identifying Genomic Relationships in Cancer Drug Response
Sawai et al. Intelligent computing systems for diagnosing plant diseases

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251215

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR