WO2025233124A1 - Sound and/or vibration predictive maintenance - Google Patents

Sound and/or vibration predictive maintenance

Info

Publication number
WO2025233124A1
WO2025233124A1 PCT/EP2025/061070 EP2025061070W WO2025233124A1 WO 2025233124 A1 WO2025233124 A1 WO 2025233124A1 EP 2025061070 W EP2025061070 W EP 2025061070W WO 2025233124 A1 WO2025233124 A1 WO 2025233124A1
Authority
WO
WIPO (PCT)
Prior art keywords
sound
analyzer module
component
trained
sensor
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/EP2025/061070
Other languages
French (fr)
Inventor
Artur Hahn
Georgii KOLOKOLNIKOV
Ingmar Graesslin
Irina Waechter-Stehle
Nicole Schadewaldt
Simon Wehle
Nils Thorben GESSERT
André GOOSSEN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Koninklijke Philips NV
Original Assignee
Koninklijke Philips NV
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Koninklijke Philips NV filed Critical Koninklijke Philips NV
Publication of WO2025233124A1 publication Critical patent/WO2025233124A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B23/00Testing or monitoring of control systems or parts thereof
    • G05B23/02Electric testing or monitoring
    • G05B23/0205Electric testing or monitoring by means of a monitoring system capable of detecting and responding to faults
    • G05B23/0259Electric testing or monitoring by means of a monitoring system capable of detecting and responding to faults characterized by the response to fault detection
    • G05B23/0283Predictive maintenance, e.g. involving the monitoring of a system and, based on the monitoring results, taking decisions on the maintenance schedule of the monitored system; Estimating remaining useful life [RUL]
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B23/00Testing or monitoring of control systems or parts thereof
    • G05B23/02Electric testing or monitoring
    • G05B23/0205Electric testing or monitoring by means of a monitoring system capable of detecting and responding to faults
    • G05B23/0218Electric testing or monitoring by means of a monitoring system capable of detecting and responding to faults characterised by the fault detection method dealing with either existing or incipient faults
    • G05B23/0224Process history based detection method, e.g. whereby history implies the availability of large amounts of data
    • G05B23/024Quantitative history assessment, e.g. mathematical relationships between available data; Functions therefor; Principal component analysis [PCA]; Partial least square [PLS]; Statistical classifiers, e.g. Bayesian networks, linear regression or correlation analysis; Neural networks
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B2219/00Program-control systems
    • G05B2219/20Pc systems
    • G05B2219/26Pc applications
    • G05B2219/2652Medical scanner
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B2219/00Program-control systems
    • G05B2219/30Nc systems
    • G05B2219/37Measurements
    • G05B2219/37337Noise, acoustic emission, sound

Definitions

  • the following generally relates to predictive maintenance of an apparatus, and, more particularly, to sound and/or vibration based predictive maintenance, including sound and/or vibration based predictive maintenance of a healthcare apparatus, such as a medical imaging scanner, a personal care apparatus, and/or other apparatus.
  • a healthcare apparatus such as a medical imaging scanner, a personal care apparatus, and/or other apparatus.
  • Electrical and/or mechanical based apparatuses include electrical and/or mechanical components that can fail.
  • the failure of individual components in some of these apparatuses often is not directly noticeable, even where the failure leads to results other than the results expected from the apparatus, e.g., degraded image quality with a medical imaging scanner.
  • Some components of these apparatuses may produce sounds and/or vibrations during operation, including characteristic sounds and/or vibrations during normal, healthy operation. However, when individual components begin to function abnormally or malfunction, the sound and/or vibrational behavior of certain components is expected to change from the characteristic sounds and/or vibrations of normal, healthy operation to other sounds and/or vibrations.
  • this approach(s) allows for continual automated monitoring of a health of an apparatus during regular operation. As such, this approach(s) may mitigate at least a shortcoming of existing maintenance approaches.
  • the approach(s) described herein can monitor sounds and/or vibrations with sufficient detail for successful predictive maintenance, including the beginning of a failure, which can mitigate subsequent more significant failure and, hence, increased repair cost, system damage, and/or system down time.
  • an apparatus in one aspect, includes at least one component that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning.
  • the apparatus further includes at least one sensor disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound.
  • the apparatus further includes a memory storing instructions for a trained artificial intelligence-based analyzer module trained to identify the characteristic sound in the signal using audio analysis.
  • the apparatus further includes a controller with a processor configured to execute the instructions to process signals from the at least one sensor to predict a health state of the apparatus based on whether the characteristic sound is identified in the signal.
  • the controller is configured to determine the health state is not normal, healthy functioning and at least invokes transmission of a notification in not identifying the characteristic sound in the signal.
  • a computer-implemented method includes training an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning.
  • the computer- implemented method further includes sensing, with a sensor, sound produced by the component, the sub-system, or the system during operation of the apparatus.
  • the computer-implemented method further includes generating, with the sensor, a signal indicative of the sensed sound.
  • the computer-implemented method further includes performing audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound.
  • the computer-implemented method further includes transmitting a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.
  • a computer readable medium is encoded with computer executable instructions that cause a processor to: train an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning, receive a signal from a sensor of an apparatus, wherein the apparatus includes a component, a sub-system, or a system that produces a characteristic sound during operation of the apparatus, and the signal sensor is configured to sense sounds of the component, the sub-system, or the system, including the characteristic sound, perform audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound, and transmit a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.
  • the invention may take form in various components and arrangements of components, and in various steps and arrangements of steps.
  • the drawings are only for the purpose of illustrating the embodiments and are not to be construed as limiting the invention.
  • FIG. 1 diagrammatically illustrates an example system including an apparatus with one or more systems, one or more sensors, a controller, and memory with instructions and data, in accordance with one or more embodiments herein.
  • FIG. 2 diagrammatically illustrates an example system of the apparatus including one or more sub-systems, in accordance with one or more embodiments herein.
  • FIG. 3 diagrammatically illustrates an example sub-system system of the system of the apparatus including one or more components, in accordance with one or more embodiments herein.
  • FIG. 4 diagrammatically illustrates an example configuration of the components and the sensors, in accordance with one or more embodiments herein.
  • FIG. 5 diagrammatically illustrates another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
  • FIG. 6 diagrammatically illustrates yet another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
  • FIG. 7 diagrammatically illustrates still another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
  • FIG. 8 diagrammatically illustrates a variation in which one or more sensors are inside a system of the apparatus, in accordance with one or more embodiments herein.
  • FIG. 9 diagrammatically illustrates a variation in which one or more sensors are inside a sub-system of the system of the apparatus, in accordance with one or more embodiments herein.
  • FIG. 10 diagrammatically illustrates a variation in which one or more sensors are on the outside of the apparatus, in accordance with one or more embodiments herein.
  • FIG. 11 diagrammatically illustrates a variation in which one or more sensors are located remote from the apparatus, in accordance with one or more embodiments herein.
  • FIG. 12 illustrates an example analyzer module, in accordance with an embodiment(s) herein.
  • FIG. 13 illustrates a variation of the analyzer module, in accordance with an embodiment(s) herein.
  • FIG. 14 illustrates another variation of the analyzer module, in accordance with an embodiment s) herein.
  • FIG. 15 illustrates an example method including training the analyzer module using supervised training in-factory and on site, in accordance with an embodiment s) herein.
  • FIG. 16 illustrates an example method including training the analyzer module using supervised training on site, in accordance with an embodiment s) herein.
  • FIG. 17 illustrates an example method employing the analyzer module trained with the supervised training, in accordance with an embodiment s) herein.
  • FIG. 18 illustrates an example method including training the analyzer module using unsupervised training in-factory and on site, in accordance with an embodiment s) herein.
  • FIG. 19 illustrates an example method including training the analyzer module using unsupervised training on site, in accordance with an embodiment s) herein.
  • FIG. 20 illustrates an example method employing the analyzer module trained with the unsupervised training, in accordance with an embodiment s) herein.
  • the following describes a sound and/or vibration predictive maintenance approach(s).
  • This approach employs sound and/or vibration sensors to sense sound and/or vibration in connection with components that produce characteristic sound and/or vibration during normal, healthy operation and other sound and/or vibration otherwise.
  • artificial intelligence Al
  • models designed and trained for sound and/or vibrational analysis are employed to process signals from the sensors and predict abnormal behavior.
  • the predictive maintenance can run under human control and/or autonomously without human interaction, producing notification when relevant, and, optionally, reports on component s) and/or apparatus health, including, in some instances, sound/vibration landscape information.
  • FIG. 1 diagrammatically illustrates an example system 100.
  • the system 100 includes an apparatus 102.
  • suitable apparatuses include apparatuses with one or more elements (e.g., electrical and/or mechanical) that produce characteristic sounds and/or vibrations during normal, healthy operation, and other sounds and/or vibrations outside of normal, healthy operation.
  • elements e.g., electrical and/or mechanical
  • Such apparatuses include medical imaging scanners such as Magnetic Resonance (MR), Computer Tomography (CT), Positron Emission Tomography (PET), Single Photon Emission Computed Tomography (SPECT), X-ray, Ultrasound (US), etc., personal care / hygienic devices such as electric toothbrushes, shavers, etc., and/or other apparatuses that include such elements.
  • MR Magnetic Resonance
  • CT Computer Tomography
  • PET Positron Emission Tomography
  • SPECT Single Photon Emission Computed Tomography
  • US Ultrasound
  • personal care / hygienic devices such as electric
  • the apparatus 102 includes a system 104i, ..., 104i, ..., and a system 104N (where N is an integer greater to or equal to one, and 1 ⁇ I ⁇ N), collectively referred to herein as systems or at least one system 104.
  • One or more of the systems 104 includes a set of subsystems, and one or more of the sub-systems includes a set of components.
  • FIG. 2 diagrammatically illustrates an embodiment in which at least the system 104i includes a subsystem 106i, . . ., a sub-system 106i, . . .
  • FIG. 3 diagrammatically illustrates an embodiment in which at least the sub-system 106i includes a component 108i, . . ., a component 108i, . . . and a component 108M (where M is an integer greater to or equal to one, and 1 ⁇ I ⁇ M), collectively referred to herein as components or at least one component 108.
  • an MR scanner apparatus generally includes the following systems: a gantry (housing the acquisition portion), an operator console (e.g., with a processor, memory, application software, etc.), and a reconstructor (e.g., with a graphics and/or other processor).
  • the gantry system may include the following sub-systems: a main magnet, a gradient coil, a gradient amplifier, an RF amplifier, RF electronics, a data acquisition system, a cooling system, etc.
  • a sub-system such as the cooling system includes cold head that produces characteristic sound (e.g., “chirping”) during normal, healthy operation as the components therein expand and contract as helium gas enters and is compressed.
  • the apparatus 102 further includes a sensor 110i, . . ., a sensor 110i, . . . and a sensor 110L (where L is an integer greater to or equal to one, and 1 ⁇ I ⁇ L), collectively referred to herein as sensors or at least one sensor 110.
  • the sensors 110 include acoustic sensors such as one or more sound sensors and/or one or more vibration sensors.
  • a sound sensor measures sound pressure / electromagnetic radiation in the audible range of the electromagnetic spectrum - 20 Hertz (Hz) to 20 kHz.
  • a suitable sound sensor can be dedicated to sensing sound of one or more components or also further configured for another purpose, e.g., communication between a patient in an examination room and a technologist at the console.
  • a non-limiting example of a sound sensor is a microphone.
  • Vibration sensors include displacement sensors, velocity sensors, and/or acceleration (Piezoelectric, micro-electromechanical systems (MEMS), etc.) sensors. In general, these sensors can measure vibration with a frequency from a few Hz to a few thousand Hz.
  • a sensor configured to sense sound may be placed such that it is not in physical contact with a component
  • a corresponding vibration sensor may have to be in physical contact with the component, sub-system, system, and/or apparatus to sense vibration of the component.
  • FIGS. 4-7 diagrammatically illustrates non -limiting configurations of the sensors 110 in connection with the components 108.
  • the sensor 110i is located within proximity to sense a sound wave 402 produced by the component 108i and is configured to sense the sound wave 402.
  • Such configuration may include low, band and/or high pass filtering as needed to sense a particular sound, signal conditioning, signal processing, etc.
  • more than one of the sensors 110 i.e., the sensor 110i, ... the sensor 110i
  • Such redundancy may be employed to confirm a detected component failure, identify a sensor that may not be functioning properly, provide backup in case a sensor fails, etc.
  • the embodiment disclosed in FIG. 6 is substantially similar to the embodiment disclosed in FIG. 5 except that the sensor 110i is configured to sense a sound wave 602 produced by the component 1081 and the sensor 110i is configured to sense a different sound wave 604 produced by the component 108i.
  • the sensor 110i is within proximity to sense a sound wave 702 produced by the component 108i, . . ., and a sound wave 704 produced by a component 108i.
  • the configuration of one or more of the components 108 and one or more of the sensors 110 can include one or more of one-to-one, one-to-many, many-to-many, and many -to- one, where one or more of the components 108 can be configured to sense one or more sound waves from each one or more of the sensors 110.
  • Another embodiment includes a combination of FIGS. 4-7.
  • FIGS. 8 and 9 diagrammatically illustrates variations of the placement of the sensors 110, including one or more of the sensors 110 inside of one or more of the systems 104, and, additionally, or alternatively, one or more of the sensors 110 inside of one or more of the sub-systems 106.
  • the system 104i includes one or more of the subsystems (i.e., the sub-system 106i, ..., the sub-system 106i, ... the sub-system 106K) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L).
  • the subsystems i.e., the sub-system 106i, ..., the sub-system 106i, ... the sub-system 106K
  • the sensors 110 i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L.
  • the sub-system 106i includes one or more of the components 108 (i.e., the component 108i, . . ., the component 108i, . . . and the component 108M) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L).
  • the components 108 i.e., the component 108i, . . ., the component 108i, . . . and the component 108M
  • the sensors 110 i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L.
  • Another embodiment includes a combination of FIGS. 1, 8 and/or 9.
  • FIG. 10 diagrammatically illustrates an embodiment in which at least the sensor 110i is disposed outside of the apparatus 102, e.g., on an outside surface of the apparatus 102.
  • FIG. 11 diagrammatically illustrates an embodiment in which at least the sensor 110i is located remote from the apparatus 102 and supported by a sensor support 1102 (a portable stand, a wall bracket, etc.). Another example combines FIGS. 1, 8 and 9.
  • the apparatus 102 further includes a computer readable storage medium (“memory”) 112, which includes non-transitory medium (e.g., a storage cell, device, etc.) and excludes transitory medium (i.e., signals, carrier waves, and the like).
  • the computer readable storage medium 112 is encoded with computer instructions (“instructions”) 114 and configured to store data 116.
  • the apparatus 102 further includes a controller 118 with a processor such as a micro-processing unit (MPU), a central processing unit (CPU), a graphics processing unit (GPU), etc.
  • the controller 118 is configured to execute the computer instructions 114 and read/write the data 116.
  • the computer instructions 114 include instructions for processing signals from the sensors 110 to determine a health state of the apparatus 102, e.g., by determining a health of one or more of at least one of the components 108, at least one of the systems 104, and/or at least one of the sub-systems 106.
  • the computer instructions 114 include artificial intelligence (Al) based instructions trained to analyze sensed information (i.e., sound and/or vibration) by one or more of the sensors 110 during operation and determine whether the apparatus 102 is functioning in a normal, healthy state or malfunctioning.
  • the Al is trained at least to learn the sounds and/or vibrations of the apparatus 102 (i.e., of one or more particular components) during regular operation, and/or distinguish sound and/or vibration from persistent and/or transient background noise, including human voice.
  • the controller 118 invokes transmission of a notification indicating a component may not be functioning properly / is likely malfunctioning.
  • the notification can also include a time stamp of when the component was determined to be malfunctioning and/or other information such as information about the signal and/or the signal itself, an expected signal, a difference therebetween, and/or decision criteria used to evaluate the difference, etc.
  • the notification may include an identification of the component, etc.
  • Other information includes an identification of a user(s) of the apparatus 102, the type of examination, parameters used for the examination, etc.
  • the controller 118 estimates a location of the component in the apparatus based on the sensor(s) sensing the signal and a spatial mapping of the sensors 110 relative to the components 108 in the data 116. Additionally, or alternatively, the controller 118 identifies an issue with the component based on the signal and a mapping between sounds and issues. In one instance, the controller 118 additionally retrieves a description of the issue and/or a possible solution to the issue based on the signal and descriptions and/or solutions to issues. In these instances, the data 116 may include multiple different sound snippets for healthy operation of a same component, but based on different uses of the apparatus 102. The estimated location, the identified issue, and/or the retrieved description and/or the possible solution can be provided with the notification.
  • a notification can be transmitted to a system and/or sub-subsystem of the apparatus 102, and/or to a computing device remote from the apparatus 102 such as a workstation, a cloud service, a smartphone, etc. where the notification can be displayed in human readable form and/or stored.
  • a user with suitable authorization e.g., personnel of the manufacturer, etc.
  • the notification can be archived and/or otherwise stored in a database, log files, in the data 116, etc.
  • FIG. 12 diagrammatically illustrates a non -limiting example of the instructions 114.
  • the instructions 114 include instructions for an analyzer module 1202.
  • the analyzer module 1202 receives, as input, one or more signals from the sensors 110.
  • the analyzer module 1202 is configured to process the one or more signals to predict a health state of the apparatus 102, the systems 104, the sub-systems 106, and/or the components 108.
  • the analyzer module 1202 is configured to process the one or more signals using artificial intelligence (Al) to detect outliers/changes in the sound landscape and log information about the timing of events, the audio changes, and/or an estimate of where the changes originated, e.g., based on the positions of the sensors 110.
  • Al artificial intelligence
  • the Al can be trained via unsupervised and/or supervised training.
  • Suitable Al include, but are not limited to, one or more of: a Neural Network (NN) such as a Recurrent Neural Network (RNN) (e.g., Long Short Term Memory (LSTM), etc.), a Variational Autoencoder (VAE), a transformer (e.g., encoders/decoders, foundation models, large language models (LLMs),etc.) with an attention mechanism (e.g., selfattention, multi-head attention, etc.), a sliding window and other temporal window analysis approach, and/or other approach(s).
  • NN Neural Network
  • RNN Recurrent Neural Network
  • VAE Variational Autoencoder
  • a transformer e.g., encoders/decoders, foundation models, large language models (LLMs),etc.
  • an attention mechanism e.g., selfattention, multi-head attention, etc.
  • a sliding window and other temporal window analysis approach e.g., selfattention, multi-head attention, etc.
  • an output from a previous step is fed back as an input to a current step in one or more hidden layers, and a hidden state stores the previous input to the network.
  • the LSTM reads and writes information important in predicting the output, with less focus on the information that is not important in predicting the output.
  • a transformer is a deep learning model that uses attention to process data.
  • the input is tokenized into individual tokens, which are encoded via an embedding layer.
  • a positional encoding vector is added to each embedding, and the embeddings then go through a multi-head self-attention layer of an encoder.
  • An add and normalize step performs a layer normalization and adds the original embeddings via a skip connection.
  • the added and normalized embeddings are processed by a multilayer perceptron consisting of multiple fully connected layers with a nonlinear activation function in between.
  • Another add and normalize step performs a layer normalization and adds the processed embeddings via a skip connection.
  • the output embeddings are then fed to a decoder, which has a processing architecture that is similar to the encoder, except it includes a masked multi-head self-attention layer and processes the output of the encoder.
  • the encoder extracts relevant information from the input and the decoder makes a prediction based thereon.
  • an X second sound snippet can be sensed and stored.
  • the transformer can take a first sub-set (i.e., all or less than X seconds) of the snippet and predict how a second snippet, a third snippet, . . . will look N seconds later. This involves embedding the sound snippets with an appropriately trained embedding model to tokenize the sound input.
  • the predicted snippet can be compared with a sensed snippet to determine any deviation of the prediction from the sense snippet.
  • the deviation may be readily noticeable because the transformer has metrics of the hidden states so any sound snippet would be transformed to a latent space vector where encoded vectors are compared with a scalar product to determine similarity.
  • the deviation can be logged, output as a graph, etc.
  • a user and/or the analyzer module 1202 can determine whether the deviation is within an acceptable tolerance or outside of the tolerance and indicative of background noise.
  • a user and/or the analyzer module 1202 can then determine whether the apparatus 102 is operating in a normal, healthy state or possibly malfunctioning, based on criteria such as a predetermined threshold.
  • the analyzer module 1202 would subtract the predicted feature vector from the feature vector of the measurement, and the result will contain a latent space representation of the unexpected deviation.
  • the decoder decodes the difference feature vector and analyzes the result.
  • a trained classifier is directly used on the difference feature vector.
  • the analyzer module 1202 can be trained using supervised learning with examples of “acceptable” noise, which is expected to show some consistent characteristics in the high-dimensional feature space. In general, a component failure and/or unacceptable functioning would produce difference vectors in feature space that are distinct from noise. If the scalar product threshold triggers the difference vector analysis, an outlier detection can be used to discriminate problematic operation signals from the “acceptable noise” distribution, which was learned during classifier training.
  • an autoregressive, generative model is employed to learn sound continuations from a given operation sound (e.g., once an apparatus has started running, predict how the sound should continue).
  • Transformers with attention mechanisms are utilized on the frequency spectra (or some transformation thereof), self-attention is utilized for sound prediction, and cross attention is utilized to determine similarity.
  • similarity between the predicted and the recorded sounds are monitored. If deviations failing a predetermined threshold arise, the deviations can be reported, e.g., via a notification and/or a report.
  • Training can be achieved by recording healthy apparatus operations and training the autoregressive model therewith. Unhealthy apparatus operation sounds can also be included with a corresponding label, including different failure and deterioration modes, if available.
  • the apparatus can be calibrated at the customer site to adapt the generative model to the “dialect” of the customer’s apparatus, e.g., via fine-tuning, i.e., adjusting a subset of transformer model parameters with an appended training to the pre-trained model from factory.
  • calibration scans can be run to adjust the parameters of the underlying neural networks to produce the recorded sounds, through back-propagation on the pre-configured network from factory training.
  • a foundation model and a LLM are forms of generative Al that are trained (e.g., self-supervised learning, semi-supervised learning, etc.) on large amounts of data and are adaptable to a wide range of tasks such as predicting, classifying, etc.
  • a sliding window analysis allows for the construction of models that can automatically learn and improve over time. They can be used to develop models that are able to automatically learn and improve from data, without the need for manual intervention.
  • One algorithm works by first dividing the data into a series of smaller windows, and then training a model on each window. The model is then used to predict the output for the next window, and the process is repeated. As the model is trained on more data, it becomes more accurate at predicting the output for future windows.
  • the analyzer module 1202 is configured to process the signal(s) based on Mel Spectrograms (e.g., visualization of sound on the Mel scale, which is a logarithmic transformation of a signal’s frequency) and Mel Frequency Cepstral Coefficients (MFCCs).
  • Mel Spectrograms e.g., visualization of sound on the Mel scale, which is a logarithmic transformation of a signal’s frequency
  • MFCCs Mel Frequency Cepstral Coefficients
  • this is achieved by converting from frequencies in Hertz to the Mel scale, taking the logarithm of Mel representation of audio, taking logarithmic magnitude and using Discrete Cosine Transformation (DCT), aggregating a spectrum over Mel frequencies as opposed to time (MFCCs), and applying an image analysis neural network framework (e.g., visual attention, and convolutional neural networks) to extract latent features of the analyzed sound and perform classification.
  • DCT Discrete Cosine Transformation
  • MFCCs e.
  • the FCC includes a two dimensional (2-D) matrix where time over short-time frequency components are represented in short time frames of that time.
  • the 2- D matrix can be linearized into a vector, which is a representation of the corresponding sound snippet.
  • the vector can be fed through a neural network or a transformer with attention mechanisms where certain different frequency components in different parts of the spectrum and/or in different times would be re-occurring and disappearing over time.
  • an X-axis of the 2-D matrix represents a timeframe and a Y axis of the 2-D matrix represents frequencies within a short time snippet.
  • the neural network or the transformer connects to different parts of the 2-D matrix to produce new snippets, new spectrograms, to continue the sound.
  • the sensors 110 are installed with the apparatus 102 in-factory.
  • the analyzer module 1202 can be trained at the factory and then installed at a customer site. The analyzer module 1202 can then be further trained at the customer site to update the analyzer module 1202 for the customer site.
  • the analyzer module 1202 is trained at the factory, installed at a customer site, and then further trained at the customer site to update the analyzer module 1202 for the customer site.
  • the sensors 110 are installed at the customer site and the analyzer module 1202 is trained at a customer site. It is to be appreciated that an installed sensor can later be re-located, e.g., after training where the analyzer module 1202 is re-trained based on the new location.
  • the instructions 114 and hence the analyzer module 1202 is part of the apparatus 102.
  • the analyzer module 1202 is located remote from the apparatus 102, e.g., a “cloud” based service, part of another computing system (e.g., local and/or remote), etc.
  • the analyzer module 1202 can operate with or without human surveillance, and produce notifications and/or reports for personnel (e.g., a user, staff, technical support, and/or vendor) when a relevant event has been detected.
  • personnel e.g., a user, staff, technical support, and/or vendor
  • the analyzer module 1202 is configured to detect abnormal functioning and produce notifications.
  • the analyzer module 1202 is further configured to describe the problem in a notification.
  • FIG. 13 diagrammatically illustrates a variation of the instructions 114 discussed in connection with FIG. 12.
  • the instructions 114 further include a report generating module 1302.
  • the report generating module 1302 is configured to generate a report with one or more of the notifications, the estimated location of the component, the identified issue of the component, the retrieved description of and/or the solution to the issue, and/or other information.
  • the report generating module 1302 can be part of a system and/or sub-system of the apparatus 102, a report generating system external to the apparatus 102, etc.
  • relevant features are extracted from the signal, e.g., via MFCCs spectrum and/or a one or multidimensional timeseries analysis.
  • the analyzer module 1202 passes the extracted features as an embedding to a LLM trained on a vast base of technical / service information to map the input to a particular malfunction.
  • a description of the malfunction is obtained from the mapping.
  • the report includes the particular malfunction along with the description thereof.
  • the report can be generated as described in US 63/604,918, filed on 1 December 2023, and entitled “maintenance service system for medical imaging systems,” which is incorporated herein by reference in its entirety.
  • the maintenance service system generates a maintenance report with information about required service actions that are tailored to a specific user and/or user group, such as an end user, e.g., by tailoring a report based on a profile of a specific individual and/or experience for self-help in medical equipment fixing/repair.
  • the analyzer module 1202 is trained to learn healthy and faulty apparatuses for different faults and with labels for problems. In this manner, for example, the analyzer module 1202 learns a particular sound and/or vibration corresponds to a loose screw, another sound and/or vibration corresponds to a bearing break, etc. With this approach, the sound spectrum is processed with a classifier that classifies the sound spectrum as a particular problem. In another instance, the analyzer module 1202 may provide a list of possible problems, each with probability, likelihood, confidence internal, etc. The data 116 and/or other storage accessible to the analyzer module 1202 can store descriptions for each of the problems, which are used for describing a problem.
  • the data 116 and/or other storage accessible to the analyzer module 1202 can store information that maps sounds to normal and/or abnormal functioning. For normal sounds, the information can be further mapped to different available uses of the apparatus 102.
  • the power requirements may be different for two different uses, and the sound and/or vibration of a gradient amplifier may be different depending on the power such that in one use the sound and/or vibration from the gradient amplifier is as expected and represents normal, healthy functioning, where that same sound and/or vibration from the same gradient amplifier in a different use either is not expected or indicates an unhealthy apparatus.
  • the sound and/or vibration will be different between an inversion recovery sequence and diffusion imaging sequence.
  • FIG. 14 diagrammatically illustrates another variation of the instructions 114.
  • the instructions 114 further include a de-identification module 1402.
  • the de-identification module 1402 is configured to evaluate the signal and determine whether the signal includes sound other than the desired sound of the systems 104, sub-systems 106 and/or components 108, e.g., human voice, machinery other than the apparatus 102, etc. In response to determining there is such sound in the signal, the de-identification module 1402 removes the sound, e.g., via filtering and/or otherwise.
  • a model is trained by overlaying different vocal recordings onto “clean” recorded machine sounds, e.g., acquired as discussed herein.
  • the combined track is the input to the model and the separate, original tracks are the output.
  • the splitter model can be prepended with a classifier that recognizes whether human speech is present in the recording. If human speech is determined to be present, the splitter can be triggered, otherwise the recording is just processed as described herein.
  • Known or other techniques can be employed. For example, known Al models such as those used to split a music track into an instrumental track and a vocal track using specifically trained neural network based models can be employed to spit human sounds from recorded machine sounds.
  • Another variation includes a combination of the variations of FIGS. 13, 14, and/or other variations.
  • the sensors 110 are employed to sense sounds while the apparatus 102 is “off’ and the analyzer module 1202 analyzes the sounds to determine a background sound landscape.
  • Different background sound landscapes can be determined for different time periods throughout a day to capture time dependent background noise such as a nearby entity with machines that produce sound sensed by the sensor 110 where the entity operates the machines only one shift each day.
  • background sound can be analyzed before, as part of and/or after a calibration of the analyzer module 1202 to determine if there is any background sound not in the background sound landscape that may impact the calibration.
  • reference sensors placed farther away from the machine to monitor can be used to determine background sounds through comparison of the signals from the main monitoring sensors with the reference sensor signals. Background sounds are expected to be found in the signals from both sensor types, but should be more prominent in the reference sensor signal.
  • the analyzer module 1202 may include Al trained under unsupervised and/or supervised training. The following describes such training and use of the trained Al for predictive maintenance.
  • FIG. 15 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is assembled in-factory, including installation of the sensors 110.
  • the apparatus 102 is confirmed to be a healthy functioning apparatus.
  • the apparatus 102 is operated in a defined use of operation.
  • the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation.
  • the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110.
  • the apparatus 102 is installed at a customer site and confirmed to be a healthy functioning apparatus.
  • the apparatus 102 is operated in one or more different operational modes.
  • the sensors 110 sense characteristic sounds of the components 108 during normal, healthy operation.
  • the analyzer module 1202 is updated, calibrated, fine-tuned, etc. based on the sounds sensed at the customer site. In one instance, this includes adjusting the model to the local sound landscape, echo characteristics, background sounds, etc. This can be achieved with predefined calibration runs, iterating through different operational modes of the machine, and/or during regular usage where the apparatus is functioning correctly.
  • FIG. 16 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is installed at a customer site.
  • the apparatus 102 is confirmed to be a healthy functioning apparatus.
  • the apparatus 102 is operated in a defined use of operation.
  • the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation.
  • the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110.
  • FIG. 17 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 15, FIG. 16, and/or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is set up for operating in a particular mode.
  • the analyzer module 1202 predicts a sound snippet for the particular mode.
  • the sensors 110 sense sounds of the apparatus 102 during the procedure and generate signals indicative thereof.
  • the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and/or otherwise, e.g., via noise cancelation, suppression, etc.
  • at least some of the pre-processing is omitted, e.g., the removal of human voice.
  • the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise.
  • a predetermined criteria e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise.
  • the analyzer module 1202 in response to the difference satisfying the predetermined criteria, the analyzer module 1202 generates a notification, as described herein and/or otherwise.
  • the notification can simply indicate a particular component may be malfunctioning or include additional information.
  • a report is also generated, as described herein and/or otherwise, with at least some of the information transmitted in and/or with the notification.
  • FIG. 18 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is assembled in-factory, including installation of the sensors 110.
  • the apparatus 102 is confirmed to be a healthy functioning apparatus.
  • the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202.
  • the training data can be acquired in-factory and/or from one or more customer sites.
  • the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training).
  • N classes for N failure modes included in training Each of the N failure modes is assigned a probability of presence by passing the output of the last neural network layer (with N neurons) through a softmax-function (i.e., a normalized exponential function,) and/or similar function to attain a probability.
  • the apparatus 102 is installed at a customer site and steps 1804-1808 are repeated as needed to update the analyzer module 1202 for the customer site.
  • FIG. 19 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is installed at a customer site.
  • the apparatus 102 is confirmed to be a healthy functioning apparatus.
  • the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202.
  • the training data can be acquired in-factory and/or from one or more customer sites.
  • the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training).
  • the analyzer module 1202 is updated as needed at the customer site.
  • FIG. 20 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 18, FIG. 19, and/or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
  • the apparatus 102 is set up for operating in a particular mode.
  • the sensors 110 sense sounds of the apparatus 102 during the procedure.
  • the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and/or otherwise, e.g., via noise cancelation, suppression, etc.
  • at least some of the pre-processing is omitted, e.g., the removal of human voice.
  • the analyzer module 1202 predicts a sound snippet for the particular mode.
  • the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise.
  • a predetermined criteria e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise.
  • the analyzer module 1202 generates a notification, as described herein and/or otherwise.
  • the notification can simply indicate a particular component may be malfunctioning or include additional information.
  • a report is also generated, as described herein and/or otherwise, with at least some of the information transmitted in and/or with the notification.
  • the sound profiles and timings of different failure modes are labeled, and, optionally, the origin location of the failure is given.
  • an incremental report is compiled by the system (e.g., once a month, etc.), with information about the functioning of the apparatus 102 and events detected in the past.
  • the collected information can be stored, e.g., by a maintaining vendor automatically in a database (e.g., possibly a vector database) and/or otherwise.
  • statistics are determined from the information (e.g., in the database) about long-term machine behavior in the field. This may also include determining more subtle precursors of approaching failure modes, or more generally attain a detailed understanding of typical weak spots of the apparatus 102 (e.g., what usually breaks first, resolved and sorted for different operational modes in the database).
  • update the analyzer module 1202 based on the statistics.
  • post-market surveillance and predictive maintenance enables automated and scalable gathering of statistics about machine failure modes in the field, filterable by different operational modes, with the availability of prior “historical” data on the machine’s functioning sound. This can boost proactive product improvement and the quality of future developments.
  • an LLM is first trained unsupervised to attain a semantic understanding of different objects, interactions, failure modes, and general physics/technical knowledge. Then, supervised learning is used to fine-tune the model to predict machine failure modes from sounds with an appropriately trained embedding model to transform the sound recordings to the latent vector domain of the LLM, as described herein. In one instance, the general knowledge of the LLM (from unsupervised training) will assist the model to recognize new failure modes, which have not been seen during supervised training.
  • a computer program may be stored/distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Automation & Control Theory (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Mathematical Physics (AREA)
  • Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)

Abstract

An apparatus (102) includes at least one component (108) that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning, at least one sensor (110) disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound, a memory (112) storing instructions (114) for a trained artificial intelligence-based analyzer module (1202) trained to identify the characteristic sound in the signal using audio analysis, and a controller (118) with a processor configured to execute the instructions to process signals from the at least one sensor (110) to predict a health state of the apparatus based on whether the characteristic sound is identified. The controller determines the health state is not normal, healthy functioning and invokes transmission of a notification in response to not identifying the characteristic sound in the signal.

Description

SOUND AND/OR VIBRATION PREDICTIVE MAINTENANCE
TECHNICAL FIELD
The following generally relates to predictive maintenance of an apparatus, and, more particularly, to sound and/or vibration based predictive maintenance, including sound and/or vibration based predictive maintenance of a healthcare apparatus, such as a medical imaging scanner, a personal care apparatus, and/or other apparatus.
BACKGROUND
Electrical and/or mechanical based apparatuses (e.g., a medical imaging scanner, an electric toothbrush, etc.) include electrical and/or mechanical components that can fail. The failure of individual components in some of these apparatuses often is not directly noticeable, even where the failure leads to results other than the results expected from the apparatus, e.g., degraded image quality with a medical imaging scanner. Some components of these apparatuses may produce sounds and/or vibrations during operation, including characteristic sounds and/or vibrations during normal, healthy operation. However, when individual components begin to function abnormally or malfunction, the sound and/or vibrational behavior of certain components is expected to change from the characteristic sounds and/or vibrations of normal, healthy operation to other sounds and/or vibrations.
With a medical imaging scanner, personnel, e.g., a technologist performing an imaging procedure, a radiologist, etc., generally are not trained to detect such failures. Moreover, job function and/or work shift rotation may prevent personnel from observing a medical imaging scanner with the requisite detail to detect such failures. Moreover, the source of a problem, e.g., which component changed its functioning, is usually unknown until service personnel have investigated and possibly disassembled the apparatus. Often, impactful failures (e.g., dangerous and/or costly) are preceded by smaller parts and functions, e.g., bearings or pumps, beginning to fail, which lead to subsequent more significant failure that can increase repair cost, system damage, and/or system down time. Unfortunately, sounds and/or vibrations of such apparatuses are not monitored with sufficient detail for useful predictive maintenance.
As such, there is an unresolved need for an improved sound and/or vibration predictive maintenance approach(s).
SUMMARY
Aspects described herein address the above-referenced problems and/or others. The following describes a sound and/or vibration predictive maintenance approach(s). In one instance, this approach(s) allows for continual automated monitoring of a health of an apparatus during regular operation. As such, this approach(s) may mitigate at least a shortcoming of existing maintenance approaches. For example, the approach(s) described herein can monitor sounds and/or vibrations with sufficient detail for successful predictive maintenance, including the beginning of a failure, which can mitigate subsequent more significant failure and, hence, increased repair cost, system damage, and/or system down time.
In one aspect, an apparatus includes at least one component that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning. The apparatus further includes at least one sensor disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound. The apparatus further includes a memory storing instructions for a trained artificial intelligence-based analyzer module trained to identify the characteristic sound in the signal using audio analysis. The apparatus further includes a controller with a processor configured to execute the instructions to process signals from the at least one sensor to predict a health state of the apparatus based on whether the characteristic sound is identified in the signal. The controller is configured to determine the health state is not normal, healthy functioning and at least invokes transmission of a notification in not identifying the characteristic sound in the signal.
In another aspect, a computer-implemented method includes training an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning. The computer- implemented method further includes sensing, with a sensor, sound produced by the component, the sub-system, or the system during operation of the apparatus. The computer-implemented method further includes generating, with the sensor, a signal indicative of the sensed sound. The computer-implemented method further includes performing audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound. The computer-implemented method further includes transmitting a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.
In another aspect, a computer readable medium is encoded with computer executable instructions that cause a processor to: train an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning, receive a signal from a sensor of an apparatus, wherein the apparatus includes a component, a sub-system, or a system that produces a characteristic sound during operation of the apparatus, and the signal sensor is configured to sense sounds of the component, the sub-system, or the system, including the characteristic sound, perform audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound, and transmit a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.
Those skilled in the art will recognize still other aspects of the present application upon reading and understanding the attached description.
BRIEF DESCRIPTION OF THE DRAWINGS
The invention may take form in various components and arrangements of components, and in various steps and arrangements of steps. The drawings are only for the purpose of illustrating the embodiments and are not to be construed as limiting the invention.
FIG. 1 diagrammatically illustrates an example system including an apparatus with one or more systems, one or more sensors, a controller, and memory with instructions and data, in accordance with one or more embodiments herein.
FIG. 2 diagrammatically illustrates an example system of the apparatus including one or more sub-systems, in accordance with one or more embodiments herein.
FIG. 3 diagrammatically illustrates an example sub-system system of the system of the apparatus including one or more components, in accordance with one or more embodiments herein.
FIG. 4 diagrammatically illustrates an example configuration of the components and the sensors, in accordance with one or more embodiments herein.
FIG. 5 diagrammatically illustrates another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
FIG. 6 diagrammatically illustrates yet another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
FIG. 7 diagrammatically illustrates still another example configuration of the components and the sensors, in accordance with one or more embodiments herein.
FIG. 8 diagrammatically illustrates a variation in which one or more sensors are inside a system of the apparatus, in accordance with one or more embodiments herein. FIG. 9 diagrammatically illustrates a variation in which one or more sensors are inside a sub-system of the system of the apparatus, in accordance with one or more embodiments herein.
FIG. 10 diagrammatically illustrates a variation in which one or more sensors are on the outside of the apparatus, in accordance with one or more embodiments herein.
FIG. 11 diagrammatically illustrates a variation in which one or more sensors are located remote from the apparatus, in accordance with one or more embodiments herein.
FIG. 12 illustrates an example analyzer module, in accordance with an embodiment(s) herein.
FIG. 13 illustrates a variation of the analyzer module, in accordance with an embodiment(s) herein.
FIG. 14 illustrates another variation of the analyzer module, in accordance with an embodiment s) herein.
FIG. 15 illustrates an example method including training the analyzer module using supervised training in-factory and on site, in accordance with an embodiment s) herein.
FIG. 16 illustrates an example method including training the analyzer module using supervised training on site, in accordance with an embodiment s) herein.
FIG. 17 illustrates an example method employing the analyzer module trained with the supervised training, in accordance with an embodiment s) herein.
FIG. 18 illustrates an example method including training the analyzer module using unsupervised training in-factory and on site, in accordance with an embodiment s) herein.
FIG. 19 illustrates an example method including training the analyzer module using unsupervised training on site, in accordance with an embodiment s) herein.
FIG. 20 illustrates an example method employing the analyzer module trained with the unsupervised training, in accordance with an embodiment s) herein.
DESCRIPTION OF EMBODIMENTS
The following describes a sound and/or vibration predictive maintenance approach(s). This approach employs sound and/or vibration sensors to sense sound and/or vibration in connection with components that produce characteristic sound and/or vibration during normal, healthy operation and other sound and/or vibration otherwise. In one instance, artificial intelligence (Al) and/or models designed and trained for sound and/or vibrational analysis are employed to process signals from the sensors and predict abnormal behavior. The predictive maintenance can run under human control and/or autonomously without human interaction, producing notification when relevant, and, optionally, reports on component s) and/or apparatus health, including, in some instances, sound/vibration landscape information.
FIG. 1 diagrammatically illustrates an example system 100. The system 100 includes an apparatus 102. Examples of suitable apparatuses include apparatuses with one or more elements (e.g., electrical and/or mechanical) that produce characteristic sounds and/or vibrations during normal, healthy operation, and other sounds and/or vibrations outside of normal, healthy operation. Examples of such apparatuses include medical imaging scanners such as Magnetic Resonance (MR), Computer Tomography (CT), Positron Emission Tomography (PET), Single Photon Emission Computed Tomography (SPECT), X-ray, Ultrasound (US), etc., personal care / hygienic devices such as electric toothbrushes, shavers, etc., and/or other apparatuses that include such elements.
The apparatus 102 includes a system 104i, ..., 104i, ..., and a system 104N (where N is an integer greater to or equal to one, and 1 < I < N), collectively referred to herein as systems or at least one system 104. One or more of the systems 104 includes a set of subsystems, and one or more of the sub-systems includes a set of components. For example, FIG. 2 diagrammatically illustrates an embodiment in which at least the system 104i includes a subsystem 106i, . . ., a sub-system 106i, . . . and 106K (where K is an integer greater to or equal to one, and 1 < I < K), collectively referred to herein as sub-systems or at least one sub-system 106, and FIG. 3 diagrammatically illustrates an embodiment in which at least the sub-system 106i includes a component 108i, . . ., a component 108i, . . . and a component 108M (where M is an integer greater to or equal to one, and 1 < I < M), collectively referred to herein as components or at least one component 108.
By way of non-limiting example, an MR scanner apparatus generally includes the following systems: a gantry (housing the acquisition portion), an operator console (e.g., with a processor, memory, application software, etc.), and a reconstructor (e.g., with a graphics and/or other processor). The gantry system may include the following sub-systems: a main magnet, a gradient coil, a gradient amplifier, an RF amplifier, RF electronics, a data acquisition system, a cooling system, etc. A sub-system such as the cooling system includes cold head that produces characteristic sound (e.g., “chirping”) during normal, healthy operation as the components therein expand and contract as helium gas enters and is compressed. Examples of other components in apparatuses in general that may produce characteristic sound and/or vibration include control electronics, a motor, a gear, a belt, a drive shaft, a pump, a compressor, an MRI bore component with characteristic vibrations in response to Lorentz forces, a moving part inside of a rotating CT gantry, etc. The apparatus 102 further includes a sensor 110i, . . ., a sensor 110i, . . . and a sensor 110L (where L is an integer greater to or equal to one, and 1 < I < L), collectively referred to herein as sensors or at least one sensor 110. The sensors 110 include acoustic sensors such as one or more sound sensors and/or one or more vibration sensors. In general, a sound sensor measures sound pressure / electromagnetic radiation in the audible range of the electromagnetic spectrum - 20 Hertz (Hz) to 20 kHz. A suitable sound sensor can be dedicated to sensing sound of one or more components or also further configured for another purpose, e.g., communication between a patient in an examination room and a technologist at the console. A non-limiting example of a sound sensor is a microphone. Vibration sensors include displacement sensors, velocity sensors, and/or acceleration (Piezoelectric, micro-electromechanical systems (MEMS), etc.) sensors. In general, these sensors can measure vibration with a frequency from a few Hz to a few thousand Hz.
For clarity and brevity, the following describes the sensors 110 in connection with sensing sound. However, one of ordinary skill in the art would readily understand differences in placement, etc. For example, whereas a sensor configured to sense sound may be placed such that it is not in physical contact with a component, a corresponding vibration sensor may have to be in physical contact with the component, sub-system, system, and/or apparatus to sense vibration of the component.
FIGS. 4-7 diagrammatically illustrates non -limiting configurations of the sensors 110 in connection with the components 108. In FIG. 4, the sensor 110i is located within proximity to sense a sound wave 402 produced by the component 108i and is configured to sense the sound wave 402. Such configuration may include low, band and/or high pass filtering as needed to sense a particular sound, signal conditioning, signal processing, etc. In FIG. 5, more than one of the sensors 110 (i.e., the sensor 110i, ... the sensor 110i) are configured to sense the same sound wave 402 produced by the component 108i. Such redundancy may be employed to confirm a detected component failure, identify a sensor that may not be functioning properly, provide backup in case a sensor fails, etc.
The embodiment disclosed in FIG. 6 is substantially similar to the embodiment disclosed in FIG. 5 except that the sensor 110i is configured to sense a sound wave 602 produced by the component 1081 and the sensor 110i is configured to sense a different sound wave 604 produced by the component 108i. In FIG. 7, the sensor 110i is within proximity to sense a sound wave 702 produced by the component 108i, . . ., and a sound wave 704 produced by a component 108i. In general, the configuration of one or more of the components 108 and one or more of the sensors 110 can include one or more of one-to-one, one-to-many, many-to-many, and many -to- one, where one or more of the components 108 can be configured to sense one or more sound waves from each one or more of the sensors 110. Another embodiment includes a combination of FIGS. 4-7.
FIGS. 8 and 9 diagrammatically illustrates variations of the placement of the sensors 110, including one or more of the sensors 110 inside of one or more of the systems 104, and, additionally, or alternatively, one or more of the sensors 110 inside of one or more of the sub-systems 106. Initially referring to FIG. 8, the system 104i includes one or more of the subsystems (i.e., the sub-system 106i, ..., the sub-system 106i, ... the sub-system 106K) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L). Next, in FIG. 9, the sub-system 106i includes one or more of the components 108 (i.e., the component 108i, . . ., the component 108i, . . . and the component 108M) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L). Another embodiment includes a combination of FIGS. 1, 8 and/or 9.
In FIG 1, the sensors 110 are shown inside of the apparatus 102. In another instance, at least one of the sensors 110 can be disposed otherwise, with the at least one of the sensors 110 located within a proximity where it can sense a sound and/or vibration of a component it is intended to monitor. For example, FIG. 10 diagrammatically illustrates an embodiment in which at least the sensor 110i is disposed outside of the apparatus 102, e.g., on an outside surface of the apparatus 102. In another example, FIG. 11 diagrammatically illustrates an embodiment in which at least the sensor 110i is located remote from the apparatus 102 and supported by a sensor support 1102 (a portable stand, a wall bracket, etc.). Another example combines FIGS. 1, 8 and 9.
Returning to FIG. 1, the apparatus 102 further includes a computer readable storage medium (“memory”) 112, which includes non-transitory medium (e.g., a storage cell, device, etc.) and excludes transitory medium (i.e., signals, carrier waves, and the like). The computer readable storage medium 112 is encoded with computer instructions (“instructions”) 114 and configured to store data 116. The apparatus 102 further includes a controller 118 with a processor such as a micro-processing unit (MPU), a central processing unit (CPU), a graphics processing unit (GPU), etc. The controller 118 is configured to execute the computer instructions 114 and read/write the data 116.
The computer instructions 114 include instructions for processing signals from the sensors 110 to determine a health state of the apparatus 102, e.g., by determining a health of one or more of at least one of the components 108, at least one of the systems 104, and/or at least one of the sub-systems 106. As described in greater detail below, in one instance the computer instructions 114 include artificial intelligence (Al) based instructions trained to analyze sensed information (i.e., sound and/or vibration) by one or more of the sensors 110 during operation and determine whether the apparatus 102 is functioning in a normal, healthy state or malfunctioning. The Al is trained at least to learn the sounds and/or vibrations of the apparatus 102 (i.e., of one or more particular components) during regular operation, and/or distinguish sound and/or vibration from persistent and/or transient background noise, including human voice.
In one instance, where it is determined that a component is not functioning in a normal, healthy state, the controller 118 invokes transmission of a notification indicating a component may not be functioning properly / is likely malfunctioning. The notification can also include a time stamp of when the component was determined to be malfunctioning and/or other information such as information about the signal and/or the signal itself, an expected signal, a difference therebetween, and/or decision criteria used to evaluate the difference, etc. In another instance, the notification may include an identification of the component, etc. Other information includes an identification of a user(s) of the apparatus 102, the type of examination, parameters used for the examination, etc.
In one instance, the controller 118 estimates a location of the component in the apparatus based on the sensor(s) sensing the signal and a spatial mapping of the sensors 110 relative to the components 108 in the data 116. Additionally, or alternatively, the controller 118 identifies an issue with the component based on the signal and a mapping between sounds and issues. In one instance, the controller 118 additionally retrieves a description of the issue and/or a possible solution to the issue based on the signal and descriptions and/or solutions to issues. In these instances, the data 116 may include multiple different sound snippets for healthy operation of a same component, but based on different uses of the apparatus 102. The estimated location, the identified issue, and/or the retrieved description and/or the possible solution can be provided with the notification.
A notification can be transmitted to a system and/or sub-subsystem of the apparatus 102, and/or to a computing device remote from the apparatus 102 such as a workstation, a cloud service, a smartphone, etc. where the notification can be displayed in human readable form and/or stored. A user with suitable authorization (e.g., personnel of the manufacturer, etc.) to read the notification can read the notification, e.g., based on their level of authorization, which can be from full access, to varying degrees of partial access, to no access. The notification, including any information therein and/or therewith, can be archived and/or otherwise stored in a database, log files, in the data 116, etc. FIG. 12 diagrammatically illustrates a non -limiting example of the instructions 114. In this example, the instructions 114 include instructions for an analyzer module 1202. The analyzer module 1202 receives, as input, one or more signals from the sensors 110. The analyzer module 1202 is configured to process the one or more signals to predict a health state of the apparatus 102, the systems 104, the sub-systems 106, and/or the components 108. In one instance, the analyzer module 1202 is configured to process the one or more signals using artificial intelligence (Al) to detect outliers/changes in the sound landscape and log information about the timing of events, the audio changes, and/or an estimate of where the changes originated, e.g., based on the positions of the sensors 110. The Al can be trained via unsupervised and/or supervised training.
Examples of suitable Al include, but are not limited to, one or more of: a Neural Network (NN) such as a Recurrent Neural Network (RNN) (e.g., Long Short Term Memory (LSTM), etc.), a Variational Autoencoder (VAE), a transformer (e.g., encoders/decoders, foundation models, large language models (LLMs),etc.) with an attention mechanism (e.g., selfattention, multi-head attention, etc.), a sliding window and other temporal window analysis approach, and/or other approach(s). In general, with a traditional NN the inputs and the outputs are independent of each other. With a RNN, an output from a previous step is fed back as an input to a current step in one or more hidden layers, and a hidden state stores the previous input to the network. The LSTM reads and writes information important in predicting the output, with less focus on the information that is not important in predicting the output.
In general, a transformer is a deep learning model that uses attention to process data. With one transformer, the input is tokenized into individual tokens, which are encoded via an embedding layer. A positional encoding vector is added to each embedding, and the embeddings then go through a multi-head self-attention layer of an encoder. An add and normalize step performs a layer normalization and adds the original embeddings via a skip connection. The added and normalized embeddings are processed by a multilayer perceptron consisting of multiple fully connected layers with a nonlinear activation function in between. Another add and normalize step performs a layer normalization and adds the processed embeddings via a skip connection. The output embeddings are then fed to a decoder, which has a processing architecture that is similar to the encoder, except it includes a masked multi-head self-attention layer and processes the output of the encoder. In this configuration, the encoder extracts relevant information from the input and the decoder makes a prediction based thereon.
With one multimodal transformer model, e.g., an X second sound snippet can be sensed and stored. The transformer can take a first sub-set (i.e., all or less than X seconds) of the snippet and predict how a second snippet, a third snippet, . . . will look N seconds later. This involves embedding the sound snippets with an appropriately trained embedding model to tokenize the sound input. The predicted snippet can be compared with a sensed snippet to determine any deviation of the prediction from the sense snippet. The deviation may be readily noticeable because the transformer has metrics of the hidden states so any sound snippet would be transformed to a latent space vector where encoded vectors are compared with a scalar product to determine similarity. The deviation can be logged, output as a graph, etc. A user and/or the analyzer module 1202 can determine whether the deviation is within an acceptable tolerance or outside of the tolerance and indicative of background noise. A user and/or the analyzer module 1202 can then determine whether the apparatus 102 is operating in a normal, healthy state or possibly malfunctioning, based on criteria such as a predetermined threshold.
In one instance, if the scalar product is too small (e.g., below a learned threshold, considering natural, “healthy” variance), the analyzer module 1202 would subtract the predicted feature vector from the feature vector of the measurement, and the result will contain a latent space representation of the unexpected deviation. In one instance, the decoder decodes the difference feature vector and analyzes the result. In another instance, a trained classifier is directly used on the difference feature vector.
For either approach, the analyzer module 1202 can be trained using supervised learning with examples of “acceptable” noise, which is expected to show some consistent characteristics in the high-dimensional feature space. In general, a component failure and/or unacceptable functioning would produce difference vectors in feature space that are distinct from noise. If the scalar product threshold triggers the difference vector analysis, an outlier detection can be used to discriminate problematic operation signals from the “acceptable noise” distribution, which was learned during classifier training.
With another transformer model, an autoregressive, generative model is employed to learn sound continuations from a given operation sound (e.g., once an apparatus has started running, predict how the sound should continue). Transformers with attention mechanisms are utilized on the frequency spectra (or some transformation thereof), self-attention is utilized for sound prediction, and cross attention is utilized to determine similarity. During operation, similarity between the predicted and the recorded sounds are monitored. If deviations failing a predetermined threshold arise, the deviations can be reported, e.g., via a notification and/or a report.
Training can be achieved by recording healthy apparatus operations and training the autoregressive model therewith. Unhealthy apparatus operation sounds can also be included with a corresponding label, including different failure and deterioration modes, if available. To adapt a generative sound model from factory training to the in-house sound landscape of a customer, the apparatus can be calibrated at the customer site to adapt the generative model to the “dialect” of the customer’s apparatus, e.g., via fine-tuning, i.e., adjusting a subset of transformer model parameters with an appended training to the pre-trained model from factory. During the healthy state, calibration scans can be run to adjust the parameters of the underlying neural networks to produce the recorded sounds, through back-propagation on the pre-configured network from factory training.
In general, a foundation model and a LLM are forms of generative Al that are trained (e.g., self-supervised learning, semi-supervised learning, etc.) on large amounts of data and are adaptable to a wide range of tasks such as predicting, classifying, etc. In general, a sliding window analysis allows for the construction of models that can automatically learn and improve over time. They can be used to develop models that are able to automatically learn and improve from data, without the need for manual intervention. One algorithm works by first dividing the data into a series of smaller windows, and then training a model on each window. The model is then used to predict the output for the next window, and the process is repeated. As the model is trained on more data, it becomes more accurate at predicting the output for future windows.
In another instance, the analyzer module 1202 is configured to process the signal(s) based on Mel Spectrograms (e.g., visualization of sound on the Mel scale, which is a logarithmic transformation of a signal’s frequency) and Mel Frequency Cepstral Coefficients (MFCCs). In one instance, this is achieved by converting from frequencies in Hertz to the Mel scale, taking the logarithm of Mel representation of audio, taking logarithmic magnitude and using Discrete Cosine Transformation (DCT), aggregating a spectrum over Mel frequencies as opposed to time (MFCCs), and applying an image analysis neural network framework (e.g., visual attention, and convolutional neural networks) to extract latent features of the analyzed sound and perform classification. In general, any approach that analyzes the spectrograms are contemplated herein, e.g., where sound is transformed to frequency components to determine which frequency components are present in one sample that are absent from another sample.
In one instance, the FCC includes a two dimensional (2-D) matrix where time over short-time frequency components are represented in short time frames of that time. The 2- D matrix can be linearized into a vector, which is a representation of the corresponding sound snippet. The vector can be fed through a neural network or a transformer with attention mechanisms where certain different frequency components in different parts of the spectrum and/or in different times would be re-occurring and disappearing over time. In one instance, an X-axis of the 2-D matrix represents a timeframe and a Y axis of the 2-D matrix represents frequencies within a short time snippet. With a retention mechanism, the neural network or the transformer connects to different parts of the 2-D matrix to produce new snippets, new spectrograms, to continue the sound.
In one instance, the sensors 110 are installed with the apparatus 102 in-factory. In this instance, the analyzer module 1202 can be trained at the factory and then installed at a customer site. The analyzer module 1202 can then be further trained at the customer site to update the analyzer module 1202 for the customer site. Alternatively, the analyzer module 1202 is trained at the factory, installed at a customer site, and then further trained at the customer site to update the analyzer module 1202 for the customer site. In another instance, the sensors 110 are installed at the customer site and the analyzer module 1202 is trained at a customer site. It is to be appreciated that an installed sensor can later be re-located, e.g., after training where the analyzer module 1202 is re-trained based on the new location.
In the illustrated embodiment, the instructions 114 and hence the analyzer module 1202 is part of the apparatus 102. In another instance, the analyzer module 1202 is located remote from the apparatus 102, e.g., a “cloud” based service, part of another computing system (e.g., local and/or remote), etc. The analyzer module 1202 can operate with or without human surveillance, and produce notifications and/or reports for personnel (e.g., a user, staff, technical support, and/or vendor) when a relevant event has been detected. In one instance, the analyzer module 1202 is configured to detect abnormal functioning and produce notifications. In another instance, the analyzer module 1202 is further configured to describe the problem in a notification.
FIG. 13 diagrammatically illustrates a variation of the instructions 114 discussed in connection with FIG. 12. In this variation, the instructions 114 further include a report generating module 1302. The report generating module 1302 is configured to generate a report with one or more of the notifications, the estimated location of the component, the identified issue of the component, the retrieved description of and/or the solution to the issue, and/or other information. The report generating module 1302 can be part of a system and/or sub-system of the apparatus 102, a report generating system external to the apparatus 102, etc.
In one instance, to determine a description of the failure, relevant features are extracted from the signal, e.g., via MFCCs spectrum and/or a one or multidimensional timeseries analysis. The analyzer module 1202 passes the extracted features as an embedding to a LLM trained on a vast base of technical / service information to map the input to a particular malfunction. A description of the malfunction is obtained from the mapping. In this instance, the report includes the particular malfunction along with the description thereof.
In another instance, the report can be generated as described in US 63/604,918, filed on 1 December 2023, and entitled “maintenance service system for medical imaging systems,” which is incorporated herein by reference in its entirety. As described in 63/604,918, the maintenance service system generates a maintenance report with information about required service actions that are tailored to a specific user and/or user group, such as an end user, e.g., by tailoring a report based on a profile of a specific individual and/or experience for self-help in medical equipment fixing/repair.
In one instance, for describing a problem and generating a report, the analyzer module 1202 is trained to learn healthy and faulty apparatuses for different faults and with labels for problems. In this manner, for example, the analyzer module 1202 learns a particular sound and/or vibration corresponds to a loose screw, another sound and/or vibration corresponds to a bearing break, etc. With this approach, the sound spectrum is processed with a classifier that classifies the sound spectrum as a particular problem. In another instance, the analyzer module 1202 may provide a list of possible problems, each with probability, likelihood, confidence internal, etc. The data 116 and/or other storage accessible to the analyzer module 1202 can store descriptions for each of the problems, which are used for describing a problem.
Additionally, or alternatively, the data 116 and/or other storage accessible to the analyzer module 1202 can store information that maps sounds to normal and/or abnormal functioning. For normal sounds, the information can be further mapped to different available uses of the apparatus 102. For example, the power requirements may be different for two different uses, and the sound and/or vibration of a gradient amplifier may be different depending on the power such that in one use the sound and/or vibration from the gradient amplifier is as expected and represents normal, healthy functioning, where that same sound and/or vibration from the same gradient amplifier in a different use either is not expected or indicates an unhealthy apparatus. By way of another example, with an MR scanner, the sound and/or vibration will be different between an inversion recovery sequence and diffusion imaging sequence.
FIG. 14 diagrammatically illustrates another variation of the instructions 114. In this variation, the instructions 114 further include a de-identification module 1402. In one instance, the de-identification module 1402 is configured to evaluate the signal and determine whether the signal includes sound other than the desired sound of the systems 104, sub-systems 106 and/or components 108, e.g., human voice, machinery other than the apparatus 102, etc. In response to determining there is such sound in the signal, the de-identification module 1402 removes the sound, e.g., via filtering and/or otherwise. In one instance, a model is trained by overlaying different vocal recordings onto “clean” recorded machine sounds, e.g., acquired as discussed herein. The combined track is the input to the model and the separate, original tracks are the output. The splitter model can be prepended with a classifier that recognizes whether human speech is present in the recording. If human speech is determined to be present, the splitter can be triggered, otherwise the recording is just processed as described herein. Known or other techniques can be employed. For example, known Al models such as those used to split a music track into an instrumental track and a vocal track using specifically trained neural network based models can be employed to spit human sounds from recorded machine sounds. Another variation includes a combination of the variations of FIGS. 13, 14, and/or other variations.
Additionally, or alternatively, known and/or other noise cancellation and/or suppression approaches can be employed. In one instance, the sensors 110 are employed to sense sounds while the apparatus 102 is “off’ and the analyzer module 1202 analyzes the sounds to determine a background sound landscape. Different background sound landscapes can be determined for different time periods throughout a day to capture time dependent background noise such as a nearby entity with machines that produce sound sensed by the sensor 110 where the entity operates the machines only one shift each day. In one instance, background sound can be analyzed before, as part of and/or after a calibration of the analyzer module 1202 to determine if there is any background sound not in the background sound landscape that may impact the calibration. This can be performed in-factory and/or at a customer site, and can be performed more than once, e.g., one or more times in the future to update the background sound landscape, if needed. Additionally, or alternatively, reference sensors placed farther away from the machine to monitor can be used to determine background sounds through comparison of the signals from the main monitoring sensors with the reference sensor signals. Background sounds are expected to be found in the signals from both sensor types, but should be more prominent in the reference sensor signal.
As briefly discussed above, the analyzer module 1202 may include Al trained under unsupervised and/or supervised training. The following describes such training and use of the trained Al for predictive maintenance.
FIG. 15 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 1502, the apparatus 102 is assembled in-factory, including installation of the sensors 110. At 1504, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1506, the apparatus 102 is operated in a defined use of operation. At 1508, the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation. At 1510, the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110. At 1512, the apparatus 102 is installed at a customer site and confirmed to be a healthy functioning apparatus.
At 1514, the apparatus 102 is operated in one or more different operational modes. At 1516, the sensors 110 sense characteristic sounds of the components 108 during normal, healthy operation. At 1518, the analyzer module 1202 is updated, calibrated, fine-tuned, etc. based on the sounds sensed at the customer site. In one instance, this includes adjusting the model to the local sound landscape, echo characteristics, background sounds, etc. This can be achieved with predefined calibration runs, iterating through different operational modes of the machine, and/or during regular usage where the apparatus is functioning correctly.
FIG. 16 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 1602 the apparatus 102 is installed at a customer site. At 1604, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1606, the apparatus 102 is operated in a defined use of operation. At 1608, the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation. At 1610, the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110.
FIG. 17 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 15, FIG. 16, and/or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 1702, the apparatus 102 is set up for operating in a particular mode. At 1704, the analyzer module 1202 predicts a sound snippet for the particular mode. At 1706, the sensors 110 sense sounds of the apparatus 102 during the procedure and generate signals indicative thereof. In one instance, the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and/or otherwise, e.g., via noise cancelation, suppression, etc. In another instance, at least some of the pre-processing is omitted, e.g., the removal of human voice.
At 1708, the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise.
At 1710, in response to the difference satisfying the predetermined criteria, the analyzer module 1202 generates a notification, as described herein and/or otherwise. For example, the notification can simply indicate a particular component may be malfunctioning or include additional information. Furthermore, in one instance, a report is also generated, as described herein and/or otherwise, with at least some of the information transmitted in and/or with the notification.
FIG. 18 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 1802, the apparatus 102 is assembled in-factory, including installation of the sensors 110. At 1804, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1806, the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202. The training data can be acquired in-factory and/or from one or more customer sites. At 1808, the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training).
This can include projecting out or computing differences between expected and measured sound data in latent space, akin to the algebraic transformations and similarity determinations used in attention mechanisms, with a neural network to link the latent space information with failure/issue states (N classes for N failure modes included in training). Each of the N failure modes is assigned a probability of presence by passing the output of the last neural network layer (with N neurons) through a softmax-function (i.e., a normalized exponential function,) and/or similar function to attain a probability. At 1810, the apparatus 102 is installed at a customer site and steps 1804-1808 are repeated as needed to update the analyzer module 1202 for the customer site.
FIG. 19 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 1902, the apparatus 102 is installed at a customer site. At 1904, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1906, the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202. The training data can be acquired in-factory and/or from one or more customer sites. At 1908, the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training). At 1910, the analyzer module 1202 is updated as needed at the customer site.
FIG. 20 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 18, FIG. 19, and/or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and/or one or more additional acts may be included.
At 2002, the apparatus 102 is set up for operating in a particular mode. At 2004, the sensors 110 sense sounds of the apparatus 102 during the procedure. In one instance, the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and/or otherwise, e.g., via noise cancelation, suppression, etc. In another instance, at least some of the pre-processing is omitted, e.g., the removal of human voice. At 2006, the analyzer module 1202 predicts a sound snippet for the particular mode.
At 2008, the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined/1 earned threshold or if latent space components associated with certain failure modes arise. At 2010, in response to the difference satisfying the predetermined criteria, the analyzer module 1202 generates a notification, as described herein and/or otherwise. For example, the notification can simply indicate a particular component may be malfunctioning or include additional information. Furthermore, in one instance, a report is also generated, as described herein and/or otherwise, with at least some of the information transmitted in and/or with the notification. At 2012, the sound profiles and timings of different failure modes are labeled, and, optionally, the origin location of the failure is given.
At 2014, an incremental report is compiled by the system (e.g., once a month, etc.), with information about the functioning of the apparatus 102 and events detected in the past. The collected information can be stored, e.g., by a maintaining vendor automatically in a database (e.g., possibly a vector database) and/or otherwise. At 2016, statistics are determined from the information (e.g., in the database) about long-term machine behavior in the field. This may also include determining more subtle precursors of approaching failure modes, or more generally attain a detailed understanding of typical weak spots of the apparatus 102 (e.g., what usually breaks first, resolved and sorted for different operational modes in the database). At 2018, where relevant, update the analyzer module 1202 based on the statistics.
In one instance, post-market surveillance and predictive maintenance enables automated and scalable gathering of statistics about machine failure modes in the field, filterable by different operational modes, with the availability of prior “historical” data on the machine’s functioning sound. This can boost proactive product improvement and the quality of future developments.
In another embodiment, an LLM is first trained unsupervised to attain a semantic understanding of different objects, interactions, failure modes, and general physics/technical knowledge. Then, supervised learning is used to fine-tune the model to predict machine failure modes from sounds with an appropriately trained embedding model to transform the sound recordings to the latent vector domain of the LLM, as described herein. In one instance, the general knowledge of the LLM (from unsupervised training) will assist the model to recognize new failure modes, which have not been seen during supervised training.
The above is implemented by way of computer readable instructions, encoded, or embedded on the computer readable storage medium, which, when executed by a processor, cause the processor to carry out the described acts or functions. Additionally, or alternatively, the above is carried out by a signal, carrier wave or other transitory medium, which is not computer readable storage medium. While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
The word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to an advantage.
A computer program may be stored/distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.

Claims

Claim 1. An apparatus (102), comprising: at least one component (108) that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning; at least one sensor (110) disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound; a memory (112) storing instructions (114) for a trained artificial intelligencebased analyzer module (1202) trained to identify the characteristic sound in the signal using audio analysis; and a controller (118) with a processor configured to execute the instructions to process signals from the at least one sensor (110) to predict a health state of the apparatus based on whether the characteristic sound is identified in the signal, wherein the controller is configured to determine the health state is not normal, healthy functioning and at least invokes transmission of a notification in response to not identifying the characteristic sound in the signal.
Claim 2. The apparatus of claim 1, wherein at least one sensor further senses environmental background sound, and the trained artificial intelligence-based analyzer module is further trained to distinguish the characteristic sound from the environmental background sound that is additionally sensed by the at least one sensor.
Claim 3. The apparatus of any of claims 1 to 2, wherein the trained artificial intelligencebased analyzer module is further trained to identify different failure modes, and the controller at least invokes transmission of the notification in response to the instructions predicting not normal, healthy functioning based on the different failure modes.
Claim 4. The apparatus of any of claims 1 to 3, wherein the memory further includes a deidentification module (1402) configured to remove human voice from sound sensed by the at least one sensor.
Claim 5. The apparatus of any of claims 1 to 4, wherein the trained artificial intelligencebased analyzer module is further trained to estimate a spatial location of the at least one component in the apparatus based on a spatial location of the at least one sensor, and the notification includes the estimated spatial location of the at least one component.
Claim 6. The apparatus of claim 5, wherein the trained artificial intelligence-based analyzer module is further trained to predict a type of the at least one component, and notification further includes the predicted type of the at least one component.
Claim 7. The apparatus of claim 6, wherein the controller is further configured to invoke generation of a report that includes the predicted type of the at least one component, the spatial location of the at least one component, and a description of the not normal, healthy functioning.
Claim 8. The apparatus of any of claims 1 to 7, wherein the artificial intelligence-based analyzer module is trained during assembly in one environment and installed and employed in a different environment.
Claim 9. The apparatus of claim 8, wherein the trained artificial intelligence-based analyzer module is updated in the different environment based on the sound landscape of the different environment.
Claim 10. The apparatus of any of claims 1 to 7, wherein the trained artificial intelligencebased analyzer module is trained and employed in a same environment.
Claim 11. A computer-implemented method, comprising: training an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning; sensing, with a sensor, sound produced by the component, the sub-system, or the system during operation of the apparatus; generating, with the sensor, a signal indicative of the sensed sound; performing audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound; and transmitting a notification in response to a result of the audio analysis not identifying the characteristic sound in the signal.
Claim 12. The computer-implemented method of claim 11, further comprising: training the analyzer module to distinguish the characteristic sound from the environmental background sound that is additionally sensed by the at least one sensor.
Claim 13. The computer-implemented method of any of claims 11 to 12, further comprising: training the analyzer module with different failure modes to identify the different failure modes.
Claim 14. The computer-implemented method of any of claims 11 to 13, further comprising: training the analyzer module during assembly with a first sound landscape of the assembly environment.
Claim 15. The computer-implemented method of claim 14, further comprising: updating the training of the analyzer module with a second sound landscape of a use environment, which is different from the assembly environment.
Claim 16. A computer readable medium encoded with computer executable instructions, which, when executed by a processor, causes the processor to: train an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning; receive a signal from a sensor of an apparatus, wherein the apparatus includes a component, a sub-system, or a system that produces a characteristic sound during operation of the apparatus, and the signal sensor is configured to sense sounds of the component, the subsystem, or the system, including the characteristic sound; perform audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound; and transmit a notification in response to a result of the audio analysis not identifying the characteristic sound in the signal.
Claim 17. The computer readable medium of claim 16, wherein the instructions further cause the processor to: train the analyzer module to distinguish the characteristic sound from the environmental background sound that is additionally sensed by the at least one sensor.
Claim 18. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: train the analyzer module with different failure modes to identify the different failure modes.
Claim 19. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: train the analyzer module during assembly with a first sound landscape of the assembly environment.
Claim 20. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: update the training of the analyzer module with a second sound landscape of a use environment, which is different from the assembly environment.
PCT/EP2025/061070 2024-05-10 2025-04-23 Sound and/or vibration predictive maintenance Pending WO2025233124A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202463645235P 2024-05-10 2024-05-10
US63/645,235 2024-05-10

Publications (1)

Publication Number Publication Date
WO2025233124A1 true WO2025233124A1 (en) 2025-11-13

Family

ID=95519066

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2025/061070 Pending WO2025233124A1 (en) 2024-05-10 2025-04-23 Sound and/or vibration predictive maintenance

Country Status (1)

Country Link
WO (1) WO2025233124A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110121969A1 (en) * 2009-11-25 2011-05-26 Unisyn Medical Technologies, Inc. Remote maintenance of medical imaging devices
US20200233397A1 (en) * 2019-01-23 2020-07-23 New York University System, method and computer-accessible medium for machine condition monitoring
US20200409653A1 (en) * 2018-02-28 2020-12-31 Robert Bosch Gmbh Intelligent Audio Analytic Apparatus (IAAA) and Method for Space System
US20210335062A1 (en) * 2020-04-23 2021-10-28 Zoox, Inc. Predicting vehicle health
US20220254366A1 (en) * 2021-02-09 2022-08-11 International Business Machines Corporation Anomalous sound detection with timbre separation

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110121969A1 (en) * 2009-11-25 2011-05-26 Unisyn Medical Technologies, Inc. Remote maintenance of medical imaging devices
US20200409653A1 (en) * 2018-02-28 2020-12-31 Robert Bosch Gmbh Intelligent Audio Analytic Apparatus (IAAA) and Method for Space System
US20200233397A1 (en) * 2019-01-23 2020-07-23 New York University System, method and computer-accessible medium for machine condition monitoring
US20210335062A1 (en) * 2020-04-23 2021-10-28 Zoox, Inc. Predicting vehicle health
US20220254366A1 (en) * 2021-02-09 2022-08-11 International Business Machines Corporation Anomalous sound detection with timbre separation

Similar Documents

Publication Publication Date Title
JP7167084B2 (en) Anomaly detection system, anomaly detection method, anomaly detection program, and learned model generation method
US12510403B2 (en) Systems and methods for monitoring of mechanical and electrical machines
US12197304B2 (en) Anomaly detection using multiple detection models
CN118801234B (en) A high voltage control cabinet with partial discharge detection device
US11853047B2 (en) Sensor-agnostic mechanical machine fault identification
US11188065B2 (en) System and method for automated fault diagnosis and prognosis for rotating equipment
JP7304545B2 (en) Anomaly prediction system and anomaly prediction method
US9443201B2 (en) Systems and methods for learning of normal sensor signatures, condition monitoring and diagnosis
US12619226B2 (en) Anomaly detection based on normal behavior modeling
US8291264B2 (en) Method and system for failure prediction with an agent
US10565080B2 (en) Discriminative hidden kalman filters for classification of streaming sensor data in condition monitoring
US8710976B2 (en) Automated incorporation of expert feedback into a monitoring system
EP3759558B1 (en) Intelligent audio analytic apparatus (iaaa) and method for space system
US12487879B2 (en) Anomaly detection on dynamic sensor data
US12455214B1 (en) Systems and methods for anomalous sound detection
JP2022068872A (en) Modular general-purpose automated anomalous data synthesizer for rotary plant
CN117129247A (en) A method for diagnosing subway door faults
Yin et al. A new Wasserstein distance-and cumulative sum-dependent health indicator and its application in prediction of remaining useful life of bearing
Xu et al. A WOA-SVMD and multi-scale CNN-transformer method for fault diagnosis of motor bearing
Martin-del-Campo et al. Algorithmic performance constraints for wind turbine condition monitoring via convolutional sparse coding with dictionary learning
EP4128088A1 (en) Training an artificial intelligence module for industrial applications
EP4529361A1 (en) Monitoring of x-ray tubes
Wang et al. Algorithm optimization based on intelligent management of computer electrical equipment: A comprehensive method for PT power monitoring and remote fault indicator
EP4585161A1 (en) Localising power supply disturbances
HK40076987A (en) Sensor-agnostic mechanical machine fault identification

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25721231

Country of ref document: EP

Kind code of ref document: A1