WO2026008772A1 - Apparatus and method for anomaly detection, monitoring system and computing cloud - Google Patents

Apparatus and method for anomaly detection, monitoring system and computing cloud

Info

Publication number
WO2026008772A1
WO2026008772A1 PCT/EP2025/068995 EP2025068995W WO2026008772A1 WO 2026008772 A1 WO2026008772 A1 WO 2026008772A1 EP 2025068995 W EP2025068995 W EP 2025068995W WO 2026008772 A1 WO2026008772 A1 WO 2026008772A1
Authority
WO
WIPO (PCT)
Prior art keywords
scene
sensor
measurement data
trained
different types
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/EP2025/068995
Other languages
French (fr)
Inventor
Bi WANG
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Europe BV United Kingdom Branch
Sony Group Corp
Original Assignee
Sony Europe Ltd
Sony Group Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Europe Ltd, Sony Group Corp filed Critical Sony Europe Ltd
Publication of WO2026008772A1 publication Critical patent/WO2026008772A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/52Surveillance or monitoring of activities, e.g. for recognising suspicious objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/809Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of classification results, e.g. where the classifiers operate on the same input data
    • G06V10/811Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of classification results, e.g. where the classifiers operate on the same input data the classifiers operating on different input data, e.g. multi-modal recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks

Definitions

  • the present disclosure relates to anomaly detection.
  • examples of the present disclosure relate to an apparatus and a method for anomaly detection, a monitoring system and a computing cloud.
  • the present disclosure provides an apparatus for anomaly detection.
  • the apparatus comprises processing circuitry configured to receive measurement data of at least two different types of sensors monitoring a scene. Further, the processing circuitry is configured to generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model. The processing circuitry is configured to determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model. In addition, the processing circuitry is configured to perform a predefined action if it is determined that an anomaly occurs in the scene.
  • the present disclosure provides a monitoring system.
  • the monitoring system comprises the apparatus for anomaly detection according to the first aspect and at least two sensors of different sensor type configured to monitor the scene and generate the measurement data.
  • the apparatus is communicatively coupled to the at least two sensors.
  • the present disclosure provides a computing cloud.
  • the computing cloud comprises the apparatus for anomaly detection according to the first aspect and interface circuitry.
  • the interface circuitry is configured to communicatively couple to at least two sensors of different sensor type monitoring the scene. Additionally, the interface circuitry is configured to receive the measurement data from the at least two sensors.
  • the present disclosure provides a method for anomaly detection.
  • the method comprises receiving measurement data of at least two different types of sensors monitoring a scene.
  • the method comprises generating semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model.
  • the method further comprises determining whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model.
  • the method comprises performing a predefined action if it is determined that an anomaly occurs in the scene.
  • the present disclosure provides a non-transitory machine- readable medium having stored thereon a program having a program code for performing the method according to the fourth aspect, when the program is executed on a processor or a programmable hardware.
  • the present disclosure provides a program having a program code for performing the method according to the fourth aspect, when the program is executed on a processor or a programmable hardware.
  • Fig. 1 illustrates an example of an apparatus for anomaly detection
  • Fig. 2 illustrates an example of a monitoring system
  • Fig. 3 illustrates an example of a computing cloud
  • Fig. 4 illustrates a flowchart of an example of a method for anomaly detection.
  • Fig. 1 schematically illustrates an apparatus 100 for anomaly (abnormality) detection.
  • the apparatus comprises processing circuitry 110.
  • the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC) a neuromorphic processor or a field programmable gate array (FPGA).
  • DSP digital signal processor
  • ASIC application specific integrated circuit
  • SoC system-on-a-chip
  • FPGA field programmable gate array
  • the processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and/or non-volatile memory.
  • the apparatus 100 may comprise memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.
  • the processing circuitry 110 is configured to receive and process measurement data 101 of at least two different types of sensors monitoring a scene.
  • the measurement data 101 may be received from various sources.
  • the measurement data 101 may at least in part be received directly from the sensors monitoring the scene.
  • the measurement data 101 may at least in part be received from an intermediate device such as a server, a cloud storage or the like which is communicatively coupled between the apparatus 100 and the sensors monitoring the scene.
  • the apparatus 100 may optionally comprise a receiver or a transceiver (not illustrated in Fig. 1) which is coupled to the processing circuitry and configured for (e.g., wireless or wired) reception of the measurement data 101.
  • the measurement data 101 originate from at least two sensors capturing different types of information in the scene or using different types of measurement principles for capturing the scene.
  • the at least two sensors may comprise one or more of an RGB sensor configured to capture a color image of the scene, an infrared sensor configured to capture a thermal image of the scene, a depth sensor configured to capture distances to objects in the scene, a Time-of-Flight (ToF) sensor configured to capture distances to objects in the scene based on the ToF of emitted light, a radar sensor configured to capture distances to objects in the scene using radar technology and an audio sensor such as a microphone configured to capture sound in the scene.
  • Processing the measurement data 101 of different types of sensors provides a richer and more accurate understanding of the scene as the measurement data of the various types of sensors may complement each other such that one type of sensor may overcome the shortcomings of another type of sensor.
  • the scene is a specific area or environment that is monitored by the at least two sensors.
  • the scene may be manifold.
  • the scene may be a private space (like a home, an apartment or a gated community), a public space (a like park, a healthcare facility, a commercial building, means of public transport, a transportation hub), an industrial space (like a factory, a manufacturing plant or a construction site) or a natural space (like a forest, a park, a farm, a greenhouse or an orchard).
  • the processing circuitry 110 is further configured to generate semantic tokens 123 indicating features extracted from the measurement data 101 using an encoder 121 of a trained transformer-based multimodal machine-learning model 120.
  • the trained transformer-based multimodal machine-learning model 120 is a data structure and/or set of rules representing a statistical model that the processing circuitry 110 uses for anomaly detection.
  • the trained transformer-based multimodal machine-learning model 120 is a machinelearning model trained to process multiple types of data inputs, such as data inputs from multiple different types of sensors, utilizing a transformer framework.
  • the trained transformer-based multimodal machine-learning model 120 comprises the encoder 121 and, optionally, a decoder, where each component includes layers that implement self-attention mechanisms to capture and analyze dependencies within and across the diverse data modalities.
  • This architecture enables the simultaneous processing of heterogeneous data sources by converting them into the semantic tokens 123, i.e., a unified representation.
  • the encoder 121 may use multi-head attention mechanisms and positional encodings to generate the semantic tokens 123.
  • the semantic tokens 123 may provide a continuous representation of the scene.
  • the semantic tokens 123 are an integrated data representation of the combined multimodal inputs (i.e., the measurement data 101) to the transformer-based multimodal machine-learning model 120.
  • the semantic tokens 123 are meaningful units of information derived from the measurement data 101.
  • the semantic token 123 represent understandable and useful features extracted from the measurement data 101 by the encoder 121.
  • the features are specific characteristics or attributes identified within the measurement data 101.
  • the features may be one or more of objects, colors, shapes, movements, distances, three-dimensional shapes, sounds, speech patterns, etc. identified in the measurement data 101.
  • the processing circuitry 100 is further configured to determine whether an anomaly occurs in the scene by analyzing the semantic tokens 123 using a self-attention mechanism 122 of the trained transformer-based multimodal machine-learning model 120.
  • the self-attention mechanism 122 uses the integrated data representation to identify anomalies in the scene.
  • the anomaly is an unusual or unexpected event or behavior in the monitored scene.
  • the anomaly may be anything that deviates from normal patterns or behaviors such as unauthorized entry, sudden movements (e.g., falling), or abnormal sounds.
  • the selfattention mechanism 122 is a core component of the transformer architecture. It allows the trained transformer-based multimodal machine-learning model 120 to weigh the importance of different parts of the input measurement data 101 (e.g.
  • Selfattention helps the trained transformer-based multimodal machine-learning model 120 focus on the most relevant features while considering the context provided by the entire input measurement data 101.
  • the processing circuitry 110 uses the self-attention mechanism to analyze the semantic tokens 123, determining how different features relate to each other to identify patterns indicative of anomalies.
  • the self-attention mechanism examines the relationships and dependencies between different tokens 123 to understand the overall context of the scene. By analyzing the tokens 123, the self-attention mechanism determines whether the patterns observed in the scene indicate an anomaly.
  • the self-attention mechanism helps the trained transformer-based multimodal machine-learning model 120 focus on critical features and their interactions, enabling it to detect deviations from normal behavior effectively.
  • an RGB camera may be used to capture visual information of a room in an apartment
  • a depth sensor may be used to measure distances to objects in the room
  • an audio sensor e.g., a microphone
  • the measurements of the sensors are provided to the processing circuitry 110 as the measurement data 101.
  • the processing circuitry 110 uses the encoder 121 of the trained transformer-based multimodal machinelearning model 120 to generate semantic tokens 123 such as "person_detected,” “distance_to_person_3_meters,” and “loud_noise_detected.”
  • the selfattention mechanism 122 of the trained transformer-based multimodal machine-learning model 120 is used by the processing circuitry 110 to analyze these tokens, examining how they relate to each other.
  • the processing circuitry 110 determines whether an anomaly is occurring.
  • an RGB camera may be used to capture visual information of a working area including machinery and workers
  • a depth sensor may be used to measure the distance between objects and detect proximity of workers to machinery
  • an audio sensor e.g., a microphone
  • the measurements of the sensors are provided to the processing circuitry 110 as the measurement data 101.
  • the processing circuitry 110 uses the encoder 121 of the trained transformer-based multimodal machine-learning model 120 to generate semantic tokens 123 such as "work- er_near_machine,” “machine_operating,” “unusual_movement” and “dis- tress_call_detected.”
  • the self-attention mechanism 122 of the trained transformer-based multimodal machine-learning model 120 is used by the processing circuitry 110 to analyze these tokens, examining how they relate to each other. For instance, it might identify that the combination of "worker_near_machine,” “unusual_movement” and “dis- tress_call_detected,” indicates a possible accident. Based on this analysis, the processing circuitry 110 determines whether an anomaly is occurring.
  • the processing circuitry 110 is configured to perform a predefined action if it is determined that an anomaly occurs in the scene.
  • the processing circuitry 110 is configured to carry out specific actions when an anomaly is detected.
  • These actions are predefined, meaning they are set up in advance based on the desired response to different types of anomalies.
  • the predefined action may be manifold.
  • the predefined action may be one or more of causing output of a notification (e.g., a message or an e- mail) about the anomaly on a user terminal (e.g., a mobile phone or a tablet-computer of a user) and causing output of an alarm (e.g., activating of an audio alarm and/or a visual alarm). Accordingly, one or more user or people in the scene may be informed about the anomaly.
  • the predefined action may further be one or more of causing recording of the scene (e.g., activating recording of the scene by a surveillance camera) or causing notification of a security service (e.g., causing output of a notification to a user terminal of the security service).
  • the predefined action may further be one or more of causing the machine in the scene to change its operation (e.g., to immediately shut down) or causing notification of a security service (e.g., causing output of a notification to a user terminal of the security service).
  • the apparatus 100 harnesses the power of transformer-based multimodal machine-learning model for abnormality detection across various input formats. Unlike traditional rule-based approaches, the apparatus 100 is built upon multimodal scene understanding principles, allowing it to analyze, e.g., dynamic visual data comprehensively, regardless of lighting conditions or input format. For instance, in a surveillance scenario with low-light conditions, the apparatus 100 is able to seamlessly integrate ToF or depth inputs alongside RGB images to enhance scene understanding and abnormality detection.
  • the apparatus 100 is able detect suspicious behavior, such as loitering or unauthorized access, even in dimly lit environments. Similarly, in industrial settings, the apparatus 100 can identify equipment malfunctions or safety hazards based on deviations from normal operating conditions, utilizing RGB inputs, depth inputs, audio inputs, etc..
  • the apparatus 100 uses the transformer-based multimodal machine-learning model's ability to enrich the content of examples by leveraging multimodal scene understanding techniques. Instead of being limited to predefined abnormal scenarios, the transformer-based multimodal machine-learning model can adapt to various input formats and make decisions based on a holistic understanding of the scene. This versatility extends its applicability to diverse environments and scenarios, enhancing safety, security, and efficiency across different applications.
  • the apparatus 100 offers a versatile and adaptable solution for abnormality detection across multiple input formats. By leveraging transformer-based multimodal scene understanding instead of traditional rule-based approaches, the apparatus 100 is able to comprehensively analyze data (such as visual data, depth data, audio data, etc.) from different sources and identify abnormalities with unprecedented accuracy and efficiency, regardless of scene conditions (e.g., light conditions or noise conditions) or input modality.
  • data such as visual data, depth data, audio data, etc.
  • the transformer-based multimodal machine-learning model 120 may be trained in various ways. By training the transformer-based multimodal machine-learning model 120 with a large set of training data and associated training content information, the machine-learning model "learns" how to generate semantic tokens and determine whether an anomaly occurs in the scene so that anomaly detection can be performed using the trained transformerbased multimodal machine-learning model 120.
  • the transformer-based multimodal machine-learning model 120 may be trained using training input data such as diverse datasets encompassing various environmental conditions input modalities such that the transformer-based multimodal machine-learning model 120 learns to recognize patterns and anomalies beyond predefined scenarios.
  • the transformer-based multimodal machine-learning model 120 may be trained to detect anomalies in one or more of different types of anomalies, different types of scenes and different types of scene settings.
  • anomalies refers to the various forms in which deviations from normal or expected patterns, behaviors, or conditions can manifest.
  • the different types of anomalies may comprise visual anomalies (i.e. , unusual or unexpected patterns detected in visual data such as images or videos) like unauthorized personnel in restricted areas, loitering, unexpected objects in a scene, equipment malfunctions, sudden changes in the visual appearance of a monitored area.
  • the different types of anomalies may comprise behavioral anomalies (i.e., irregular or unexpected behaviors or movement patterns, particularly in monitored individuals or entities) such as erratic movements, deviations from expected routes, prolonged inactivity in an active area, signs of distress or emergency behaviors.
  • the different types of anomalies may comprise environmental anomalies (i.e., deviations from normal environmental conditions) such as rapid environmental changes that could indicate a problem, presence of fire, sudden presence of fluids like water.
  • the different types of anomalies may further comprise audio anomalies (i.e., unusual sounds or noise patterns that deviate from the normal acoustic environment) such as alarms or sirens, machinery malfunction noises, verbal distress calls or sudden loud noises in typically quiet areas.
  • Each type of anomaly has specific characteristics that the trained transformer-based multimodal machine-learning model 120 is capable of recognizing and addressing.
  • Different types of scenes refers to the range of environments, contexts, or settings monitored for anomaly detection.
  • the different types of scenes refers may comprise residential areas (like homes, apartments or gated communities), industrial environments (like factories, manufacturing plants or construction sites), public areas (like parks, healthcare facilities, commercial buildings, means of public transport, transportation hubs) or natural environments (like forests, parks, farms, greenhouses or orchards).
  • Each type of scene has specific features and potential anomalies that the trained transformer-based multimodal machine-learning model 120 is capable of recognizing and addressing.
  • Different types of scenes refers to the various conditions and contexts under which the same scene or environment is observed. For example, this includes variations due to time of day, weather, and seasonal changes, which can affect the appearance, dynamics, and characteristics of the scene.
  • the different scene settings introduce diverse visual and con- textual challenges that the trained transformer-based multimodal machine-learning model 120 is able to account for.
  • the training input data may comprise measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene.
  • associated training content information such as related semantic tokens and anomaly detection results may be used.
  • Exemplary data sets for training the transformer-based multimodal machine-learning model 120 may be found at https://github.com/drmuskangarg/Multimodal-datasets.
  • the transformer-based multimodal machine-learning model 120 may be trained using a training method called "supervised learning".
  • supervised learning the transformer-based multimodal machine-learning model 120 is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value.
  • the transformer-based multimodal machine-learning model 120 "learns" which output value to provide based on an input sample that is similar to the samples provided during the training.
  • a training sample may comprise measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene settings as input data and the related semantic tokens and anomaly detection results as desired output data.
  • semi-supervised learning may be used.
  • semi-supervised learning some of the training samples lack a corresponding desired output value.
  • Supervised learning may be based on a supervised learning algorithm (e.g., a classification algorithm or a similarity learning algorithm).
  • Classification algorithms may be used as the desired outputs of the trained transformer-based multimodal machine-learning model 120 are restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values (e.g., certain types of anomalies).
  • Similarity learning algorithms are similar to classification algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are.
  • unsupervised learning may be used to train the transformer-based multimodal machine-learning model 120.
  • unsupervised learn-ing (only) input data are supplied and an unsupervised learning algorithm is used to find structure in the input data such as measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene.
  • Reinforcement learning is a third group of machine-learning algorithms.
  • reinforcement learning may be used to train the transformer-based multimodal machinelearning model 120.
  • one or more software actors (called “software agents”) are trained to take actions in an environment. Based on the taken actions, a reward is calculated.
  • Reinforcement learning is based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
  • Feature learning may be used.
  • the transformer-based multimodal machine-learning model 120 may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component.
  • Feature learning algorithms which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before per-forming classification or predictions.
  • Feature learning may be based on principal components analysis or cluster analysis, for example.
  • the transformer-based multimodal machine-learning model 120 may be a 4M (Massively Multimodal Masked Modeling) model, which is trained using the training procedure described in https://github.com/apple/ml-4m/blob/main/README_TRAINING.md.
  • 4M Massively Multimodal Masked Modeling
  • the semantic tokens 123 are generated from the multimodal measurement data 101 using the encoder 121 of the trained transformer-based multimodal machine-learning model.
  • the encoder 121 may comprise modalityspecific encoder portions (sections, modules).
  • the encoder 121 may comprise individual encoder portions for the different types of sensors.
  • the processing circuitry 110 may be configured to generate the semantic tokens 123 using the respective encoder portion for the measurement data 101 of the different types of sensors.
  • the encoder 121 may comprise individual encoder portions for measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a ToF sensor, a radar sensor and an audio sensor.
  • Using separate encoder portions allows to account for the specific characteristics of each modality (i.e. , the different types of sensor data).
  • the tokens output from these modality-specific encoder portions may be combined or fused by the trained transformer-based multimodal machine-learning model 120 in a way that allows the self-attention mechanism 122 of the trained transformer-based multimodal machinelearning model 120 to leverage information from all available data types. For example, concatenation, attention-based fusion, and more complex integration techniques may be used.
  • the apparatus 100 may be used in various ways and implementations for anomaly detection. In the following, two non-limiting examples will be given with reference to Fig. 2 and Fig. 3.
  • Fig. 2 schematically illustrates a monitoring system 200.
  • the monitoring system 200 comprises the apparatus 100 for anomaly detection as described above.
  • the monitoring system 200 comprises at least two sensors 210 and 220.
  • the at least two sensors 210 and 220 are of different sensor type.
  • the at least two sensors 210 and 220 may comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
  • the at least two sensors 210 and 220 configured to monitor (capture) the scene and generate the measurement data 101.
  • the apparatus 100 is communicatively coupled to the at least two sensors 210 and 220 and configured to receive the measurement data 101 from the at least two sensors 210 and 220 for further processing.
  • the monitoring system 200 provides a compact system with improved anomaly detection.
  • the monitoring system 200 may, e.g., be used for monitoring residential areas or industrial areas.
  • Fig. 3 schematically illustrates a computing cloud 300.
  • the computing cloud 300 comprises the apparatus 100 for anomaly detection as described above.
  • the computing cloud 300 comprises interface circuitry 310 coupled to the apparatus 100.
  • the interface circuitry 310 is configured to communicatively couple to at least two sensors 320 and 330 of different sensor type monitoring the scene (e.g., via the Internet).
  • the at least two sensors 320 and 330 are external to the computing cloud 300 (i.e., not part of the computing cloud 300).
  • the at least two sensors 320 and 330 may comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
  • the interface circuitry 310 is further configured to receive the measurement data 101 from the at least two sensors 320 and 330.
  • the apparatus 100 is configured to receive the measurement data 101 from the interface circuitry 310 for further processing.
  • the computing cloud 300 allows to provide improved anomaly detection as a cloud service.
  • Fig. 4 illustrates a flowchart of a method 400 for anomaly detection.
  • the method 400 comprises receiving 402 measurement data of at least two different types of sensors monitoring a scene (e.g., measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of- flight sensor, a radar sensor and an audio sensor).
  • the method 400 comprises generating 404 semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model.
  • the method 400 further comprises determining 406 whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformerbased multimodal machine-learning model.
  • the method 400 comprises performing 408 a predefined action if it is determined that an anomaly occurs in the scene.
  • the method 400 provides improved anomaly detection.
  • receiving 402 the measurement data and generating 404 the semantic tokens may, e.g., be performed at an edge device.
  • Determining 406 whether an anomaly occurs in the scene may be performed at a computing cloud communicatively coupled to the edge device.
  • the edge device is a local device processing data at the periphery (“edge”) of a network.
  • the local processing of the measurement data at the edge device allows to enhance privacy and security as the amount of sensitive information transmitted over the network to the computing cloud is reduced.
  • the network infrastructure is used more efficiently.
  • the method 400 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above.
  • the proposed technology is able to have multiple formats such as, e.g., RGB, depth, infrared etc. simultaneously as input.
  • the proposed technology is not a traditional rule based method where only predefined situations can be detected and reported. For example, a fire detection system will only detect whether there is a fire in the frame but can’t do anything else. Since the proposed technology is based on scene understanding and multimodality, it can process various abnormal detection scenarios with different input formats which makes it universally helpful to any kind of places, from living room to factory, from good light condition to no light condition. Since the proposed technology can not only take RGB data as input for abnormality detection, it also has the ability of privacy preserving anomaly detection compared to traditional methods.
  • An apparatus for anomaly detection comprising processing circuitry configured to: receive measurement data of at least two different types of sensors monitoring a scene; generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model; and perform a predefined action if it is determined that an anomaly occurs in the scene.
  • the predefined action is one or more of causing output of a notification about the anomaly on a user terminal, causing output of an alarm, causing recording of the scene, causing notification of a security service and causing a machine in the scene to change its operation.
  • the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scene settings.
  • the encoder comprises individual encoder portions for the different types of sensors, and wherein the processing circuitry is configured to generate the semantic tokens using the respective encoder portion for the measurement data of the different types of sensors.
  • the encoder comprises individual encoder portions for measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
  • a monitoring system comprising: the apparatus for anomaly detection according to any one of (1) to (7); and at least two sensors of different sensor type configured to monitor the scene and generate the measurement data, wherein the apparatus is communicatively coupled to the at least two sensors.
  • the at least two sensors comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
  • a computing cloud comprising: the apparatus for anomaly detection according to any one of (1) to (7); and interface circuitry configured to: communicatively couple to at least two sensors of different sensor type monitoring the scene; and receive the measurement data from the at least two sensors.
  • a method for anomaly detection comprising: receiving measurement data of at least two different types of sensors monitoring a scene; generating semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determining whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machinelearning model; and performing a predefined action if it is determined that an anomaly occurs in the scene.
  • the measurement data comprises measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
  • Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component.
  • steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components.
  • Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machineexecutable, processor-executable or computer-executable programs and instructions.
  • Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example.
  • Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
  • FPLAs field programmable logic arrays
  • F)PGAs field) programmable gate arrays
  • GPU graphics processor units
  • ASICs application-specific integrated circuits
  • ICs integrated circuits
  • SoCs system-on-a-chip
  • aspects described in relation to a device or system should also be understood as a description of the corresponding method.
  • a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method.
  • aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Evolutionary Computation (AREA)
  • Multimedia (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • Artificial Intelligence (AREA)
  • Health & Medical Sciences (AREA)
  • Testing And Monitoring For Control Systems (AREA)

Abstract

Provided is an apparatus for anomaly detection. The apparatus includes processing circuitry configured to receive measurement data of at least two different types of sensors monitoring a scene. Further, the processing circuitry is configured to generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model. The processing circuitry is configured to determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model. In addition, the processing circuitry is configured to perform a predefined action if it is determined that an anomaly occurs in the scene.

Description

APPARATUS AND METHOD FOR ANOMALY DETECTION, MONITORING SYSTEM AND COMPUTING CLOUD
Field
The present disclosure relates to anomaly detection. In particular, examples of the present disclosure relate to an apparatus and a method for anomaly detection, a monitoring system and a computing cloud.
Background
Conventional monitoring systems primarily rely on rule-based approaches, leading to several disadvantages and shortcomings. One significant limitation is the inability to effectively detect abnormalities beyond predefined scenarios. Traditional systems are often designed with specific rules tailored to detect limited events such as fire or human falls, thereby lacking flexibility in recognizing diverse abnormal occurrences. This narrow focus restricts the system's applicability to only those scenarios for which it has been explicitly programmed, rendering it inadequate for detecting anomalies in unfamiliar or evolving environments. Consequently, existing monitoring systems face challenges in adapting to new contexts and addressing emerging abnormal events outside their predefined scope.
Hence, there may be a demand for improved anomaly detection.
Summary
This demand is met by an apparatus and a method for anomaly detection, a monitoring system, a computing cloud, a non-transitory machine-readable medium and a program in accordance with the independent claims. Advantageous embodiments are defined by the dependent claims.
According to a first aspect, the present disclosure provides an apparatus for anomaly detection. The apparatus comprises processing circuitry configured to receive measurement data of at least two different types of sensors monitoring a scene. Further, the processing circuitry is configured to generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model. The processing circuitry is configured to determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model. In addition, the processing circuitry is configured to perform a predefined action if it is determined that an anomaly occurs in the scene.
According to a second aspect, the present disclosure provides a monitoring system. The monitoring system comprises the apparatus for anomaly detection according to the first aspect and at least two sensors of different sensor type configured to monitor the scene and generate the measurement data. The apparatus is communicatively coupled to the at least two sensors.
According to a third aspect, the present disclosure provides a computing cloud. The computing cloud comprises the apparatus for anomaly detection according to the first aspect and interface circuitry. The interface circuitry is configured to communicatively couple to at least two sensors of different sensor type monitoring the scene. Additionally, the interface circuitry is configured to receive the measurement data from the at least two sensors.
According to a fourth aspect, the present disclosure provides a method for anomaly detection. The method comprises receiving measurement data of at least two different types of sensors monitoring a scene. In addition, the method comprises generating semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model. The method further comprises determining whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model. Additionally, the method comprises performing a predefined action if it is determined that an anomaly occurs in the scene.
According to a fifth aspect, the present disclosure provides a non-transitory machine- readable medium having stored thereon a program having a program code for performing the method according to the fourth aspect, when the program is executed on a processor or a programmable hardware.
According to a sixth aspect, the present disclosure provides a program having a program code for performing the method according to the fourth aspect, when the program is executed on a processor or a programmable hardware.
Brief description of the Figures Some examples of apparatuses and/or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
Fig. 1 illustrates an example of an apparatus for anomaly detection;
Fig. 2 illustrates an example of a monitoring system;
Fig. 3 illustrates an example of a computing cloud; and
Fig. 4 illustrates a flowchart of an example of a method for anomaly detection.
Detailed Description
Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
Throughout the description of the figures same or similar reference numerals refer to same or similar elements and/or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and/or areas in the figures may also be exaggerated for clarification.
When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and/or B" may be used. This applies equivalently to combinations of more than two elements.
If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and/or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and/or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and/or a group thereof.
Fig. 1 schematically illustrates an apparatus 100 for anomaly (abnormality) detection.
The apparatus comprises processing circuitry 110. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC) a neuromorphic processor or a field programmable gate array (FPGA). The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and/or non-volatile memory. For example, the apparatus 100 may comprise memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.
The processing circuitry 110 is configured to receive and process measurement data 101 of at least two different types of sensors monitoring a scene. The measurement data 101 may be received from various sources. For example, the measurement data 101 may at least in part be received directly from the sensors monitoring the scene. Alternatively or additionally, the measurement data 101 may at least in part be received from an intermediate device such as a server, a cloud storage or the like which is communicatively coupled between the apparatus 100 and the sensors monitoring the scene. The apparatus 100 may optionally comprise a receiver or a transceiver (not illustrated in Fig. 1) which is coupled to the processing circuitry and configured for (e.g., wireless or wired) reception of the measurement data 101.
The measurement data 101 originate from at least two sensors capturing different types of information in the scene or using different types of measurement principles for capturing the scene. For example, the at least two sensors may comprise one or more of an RGB sensor configured to capture a color image of the scene, an infrared sensor configured to capture a thermal image of the scene, a depth sensor configured to capture distances to objects in the scene, a Time-of-Flight (ToF) sensor configured to capture distances to objects in the scene based on the ToF of emitted light, a radar sensor configured to capture distances to objects in the scene using radar technology and an audio sensor such as a microphone configured to capture sound in the scene. Processing the measurement data 101 of different types of sensors provides a richer and more accurate understanding of the scene as the measurement data of the various types of sensors may complement each other such that one type of sensor may overcome the shortcomings of another type of sensor.
The scene is a specific area or environment that is monitored by the at least two sensors. The scene may be manifold. For example, the scene may be a private space (like a home, an apartment or a gated community), a public space (a like park, a healthcare facility, a commercial building, means of public transport, a transportation hub), an industrial space (like a factory, a manufacturing plant or a construction site) or a natural space (like a forest, a park, a farm, a greenhouse or an orchard).
The processing circuitry 110 is further configured to generate semantic tokens 123 indicating features extracted from the measurement data 101 using an encoder 121 of a trained transformer-based multimodal machine-learning model 120. The trained transformer-based multimodal machine-learning model 120 is a data structure and/or set of rules representing a statistical model that the processing circuitry 110 uses for anomaly detection. In particular, the trained transformer-based multimodal machine-learning model 120 is a machinelearning model trained to process multiple types of data inputs, such as data inputs from multiple different types of sensors, utilizing a transformer framework. The trained transformer-based multimodal machine-learning model 120 comprises the encoder 121 and, optionally, a decoder, where each component includes layers that implement self-attention mechanisms to capture and analyze dependencies within and across the diverse data modalities. This architecture enables the simultaneous processing of heterogeneous data sources by converting them into the semantic tokens 123, i.e., a unified representation. For example, the encoder 121 may use multi-head attention mechanisms and positional encodings to generate the semantic tokens 123.
The semantic tokens 123 may provide a continuous representation of the scene. In other words, the semantic tokens 123 are an integrated data representation of the combined multimodal inputs (i.e., the measurement data 101) to the transformer-based multimodal machine-learning model 120. The semantic tokens 123 are meaningful units of information derived from the measurement data 101. In particular, the semantic token 123 represent understandable and useful features extracted from the measurement data 101 by the encoder 121. The features are specific characteristics or attributes identified within the measurement data 101. For example, the features may be one or more of objects, colors, shapes, movements, distances, three-dimensional shapes, sounds, speech patterns, etc. identified in the measurement data 101. The processing circuitry 100 is further configured to determine whether an anomaly occurs in the scene by analyzing the semantic tokens 123 using a self-attention mechanism 122 of the trained transformer-based multimodal machine-learning model 120. In other words, the self-attention mechanism 122 uses the integrated data representation to identify anomalies in the scene. The anomaly is an unusual or unexpected event or behavior in the monitored scene. The anomaly may be anything that deviates from normal patterns or behaviors such as unauthorized entry, sudden movements (e.g., falling), or abnormal sounds. The selfattention mechanism 122 is a core component of the transformer architecture. It allows the trained transformer-based multimodal machine-learning model 120 to weigh the importance of different parts of the input measurement data 101 (e.g. weigh the importance of the individual semantic tokens 123), capturing dependencies and relationships between them. Selfattention helps the trained transformer-based multimodal machine-learning model 120 focus on the most relevant features while considering the context provided by the entire input measurement data 101. The processing circuitry 110 uses the self-attention mechanism to analyze the semantic tokens 123, determining how different features relate to each other to identify patterns indicative of anomalies. The self-attention mechanism examines the relationships and dependencies between different tokens 123 to understand the overall context of the scene. By analyzing the tokens 123, the self-attention mechanism determines whether the patterns observed in the scene indicate an anomaly. The self-attention mechanism helps the trained transformer-based multimodal machine-learning model 120 focus on critical features and their interactions, enabling it to detect deviations from normal behavior effectively.
For example, an RGB camera may be used to capture visual information of a room in an apartment, a depth sensor may be used to measure distances to objects in the room and an audio sensor (e.g., a microphone) may be used to capture sounds in the room. The measurements of the sensors are provided to the processing circuitry 110 as the measurement data 101. Using the encoder 121 of the trained transformer-based multimodal machinelearning model 120, the processing circuitry 110 generates semantic tokens 123 such as "person_detected," "distance_to_person_3_meters," and "loud_noise_detected." The selfattention mechanism 122 of the trained transformer-based multimodal machine-learning model 120 is used by the processing circuitry 110 to analyze these tokens, examining how they relate to each other. For instance, it might identify that the combination of "per- son_detected" and "loud_noise_detected" at an unusual time indicates a potential break-in. Based on this analysis, the processing circuitry 110 determines whether an anomaly is occurring. In another example, an RGB camera may be used to capture visual information of a working area including machinery and workers, a depth sensor may be used to measure the distance between objects and detect proximity of workers to machinery and an audio sensor (e.g., a microphone) may be used to capture sounds in the work area. The measurements of the sensors are provided to the processing circuitry 110 as the measurement data 101. Using the encoder 121 of the trained transformer-based multimodal machine-learning model 120, the processing circuitry 110 generates semantic tokens 123 such as "work- er_near_machine," "machine_operating," "unusual_movement" and "dis- tress_call_detected." The self-attention mechanism 122 of the trained transformer-based multimodal machine-learning model 120 is used by the processing circuitry 110 to analyze these tokens, examining how they relate to each other. For instance, it might identify that the combination of "worker_near_machine," "unusual_movement" and "dis- tress_call_detected," indicates a possible accident. Based on this analysis, the processing circuitry 110 determines whether an anomaly is occurring.
In addition, the processing circuitry 110 is configured to perform a predefined action if it is determined that an anomaly occurs in the scene. In other words, the processing circuitry 110 is configured to carry out specific actions when an anomaly is detected. These actions are predefined, meaning they are set up in advance based on the desired response to different types of anomalies. The predefined action may be manifold. For example, the predefined action may be one or more of causing output of a notification (e.g., a message or an e- mail) about the anomaly on a user terminal (e.g., a mobile phone or a tablet-computer of a user) and causing output of an alarm (e.g., activating of an audio alarm and/or a visual alarm). Accordingly, one or more user or people in the scene may be informed about the anomaly.
In the above example of monitoring the room in the apartment, the predefined action may further be one or more of causing recording of the scene (e.g., activating recording of the scene by a surveillance camera) or causing notification of a security service (e.g., causing output of a notification to a user terminal of the security service).
In the above example of monitoring the work area, the predefined action may further be one or more of causing the machine in the scene to change its operation (e.g., to immediately shut down) or causing notification of a security service (e.g., causing output of a notification to a user terminal of the security service). The apparatus 100 harnesses the power of transformer-based multimodal machine-learning model for abnormality detection across various input formats. Unlike traditional rule-based approaches, the apparatus 100 is built upon multimodal scene understanding principles, allowing it to analyze, e.g., dynamic visual data comprehensively, regardless of lighting conditions or input format. For instance, in a surveillance scenario with low-light conditions, the apparatus 100 is able to seamlessly integrate ToF or depth inputs alongside RGB images to enhance scene understanding and abnormality detection. The apparatus 100 is able detect suspicious behavior, such as loitering or unauthorized access, even in dimly lit environments. Similarly, in industrial settings, the apparatus 100 can identify equipment malfunctions or safety hazards based on deviations from normal operating conditions, utilizing RGB inputs, depth inputs, audio inputs, etc..
The apparatus 100 uses the transformer-based multimodal machine-learning model's ability to enrich the content of examples by leveraging multimodal scene understanding techniques. Instead of being limited to predefined abnormal scenarios, the transformer-based multimodal machine-learning model can adapt to various input formats and make decisions based on a holistic understanding of the scene. This versatility extends its applicability to diverse environments and scenarios, enhancing safety, security, and efficiency across different applications.
The apparatus 100 offers a versatile and adaptable solution for abnormality detection across multiple input formats. By leveraging transformer-based multimodal scene understanding instead of traditional rule-based approaches, the apparatus 100 is able to comprehensively analyze data (such as visual data, depth data, audio data, etc.) from different sources and identify abnormalities with unprecedented accuracy and efficiency, regardless of scene conditions (e.g., light conditions or noise conditions) or input modality.
The transformer-based multimodal machine-learning model 120 may be trained in various ways. By training the transformer-based multimodal machine-learning model 120 with a large set of training data and associated training content information, the machine-learning model "learns" how to generate semantic tokens and determine whether an anomaly occurs in the scene so that anomaly detection can be performed using the trained transformerbased multimodal machine-learning model 120.
The transformer-based multimodal machine-learning model 120 may be trained using training input data such as diverse datasets encompassing various environmental conditions input modalities such that the transformer-based multimodal machine-learning model 120 learns to recognize patterns and anomalies beyond predefined scenarios. In particular, the transformer-based multimodal machine-learning model 120 may be trained to detect anomalies in one or more of different types of anomalies, different types of scenes and different types of scene settings.
Different types of anomalies refers to the various forms in which deviations from normal or expected patterns, behaviors, or conditions can manifest. For example, the different types of anomalies may comprise visual anomalies (i.e. , unusual or unexpected patterns detected in visual data such as images or videos) like unauthorized personnel in restricted areas, loitering, unexpected objects in a scene, equipment malfunctions, sudden changes in the visual appearance of a monitored area. Alternatively or additionally, the different types of anomalies may comprise behavioral anomalies (i.e., irregular or unexpected behaviors or movement patterns, particularly in monitored individuals or entities) such as erratic movements, deviations from expected routes, prolonged inactivity in an active area, signs of distress or emergency behaviors. Further alternatively or additionally, the different types of anomalies may comprise environmental anomalies (i.e., deviations from normal environmental conditions) such as rapid environmental changes that could indicate a problem, presence of fire, sudden presence of fluids like water. The different types of anomalies may further comprise audio anomalies (i.e., unusual sounds or noise patterns that deviate from the normal acoustic environment) such as alarms or sirens, machinery malfunction noises, verbal distress calls or sudden loud noises in typically quiet areas. Each type of anomaly has specific characteristics that the trained transformer-based multimodal machine-learning model 120 is capable of recognizing and addressing.
Different types of scenes refers to the range of environments, contexts, or settings monitored for anomaly detection. For example, the different types of scenes refers may comprise residential areas (like homes, apartments or gated communities), industrial environments (like factories, manufacturing plants or construction sites), public areas (like parks, healthcare facilities, commercial buildings, means of public transport, transportation hubs) or natural environments (like forests, parks, farms, greenhouses or orchards). Each type of scene has specific features and potential anomalies that the trained transformer-based multimodal machine-learning model 120 is capable of recognizing and addressing.
Different types of scenes refers to the various conditions and contexts under which the same scene or environment is observed. For example, this includes variations due to time of day, weather, and seasonal changes, which can affect the appearance, dynamics, and characteristics of the scene. The different scene settings introduce diverse visual and con- textual challenges that the trained transformer-based multimodal machine-learning model 120 is able to account for.
In order to train the transformer-based multimodal machine-learning model 120 to detect different types of anomalies, detect anomalies in different types of scenes and/or detect anomalies in different types of scene settings, the training input data may comprise measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene. Furthermore, associated training content information such as related semantic tokens and anomaly detection results may be used. Exemplary data sets for training the transformer-based multimodal machine-learning model 120 may be found at https://github.com/drmuskangarg/Multimodal-datasets.
For example, the transformer-based multimodal machine-learning model 120 may be trained using a training method called "supervised learning". In supervised learning, the transformer-based multimodal machine-learning model 120 is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value. By specifying both training samples and desired output values, the transformer-based multimodal machine-learning model 120 "learns" which output value to provide based on an input sample that is similar to the samples provided during the training. For example, a training sample may comprise measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene settings as input data and the related semantic tokens and anomaly detection results as desired output data.
Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g., a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained transformer-based multimodal machine-learning model 120 are restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values (e.g., certain types of anomalies). Similarity learning algorithms are similar to classification algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are.
Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the transformer-based multimodal machine-learning model 120. In unsupervised learn- ing, (only) input data are supplied and an unsupervised learning algorithm is used to find structure in the input data such as measurement data of different types of sensors for at least one of different types of anomalies, different types of scenes and different types of scene.
Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the transformer-based multimodal machinelearning model 120. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the transformer-based multimodal machine-learning model 120 may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before per-forming classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
According to examples, the transformer-based multimodal machine-learning model 120 may be a 4M (Massively Multimodal Masked Modeling) model, which is trained using the training procedure described in https://github.com/apple/ml-4m/blob/main/README_TRAINING.md.
As described above, the semantic tokens 123 are generated from the multimodal measurement data 101 using the encoder 121 of the trained transformer-based multimodal machine-learning model. According to examples, the encoder 121 may comprise modalityspecific encoder portions (sections, modules). In other words, the encoder 121 may comprise individual encoder portions for the different types of sensors. Accordingly, the processing circuitry 110 may be configured to generate the semantic tokens 123 using the respective encoder portion for the measurement data 101 of the different types of sensors. For example, the encoder 121 may comprise individual encoder portions for measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a ToF sensor, a radar sensor and an audio sensor. Using separate encoder portions allows to account for the specific characteristics of each modality (i.e. , the different types of sensor data). The tokens output from these modality-specific encoder portions may be combined or fused by the trained transformer-based multimodal machine-learning model 120 in a way that allows the self-attention mechanism 122 of the trained transformer-based multimodal machinelearning model 120 to leverage information from all available data types. For example, concatenation, attention-based fusion, and more complex integration techniques may be used.
The apparatus 100 may be used in various ways and implementations for anomaly detection. In the following, two non-limiting examples will be given with reference to Fig. 2 and Fig. 3.
Fig. 2 schematically illustrates a monitoring system 200. The monitoring system 200 comprises the apparatus 100 for anomaly detection as described above. Furthermore, the monitoring system 200 comprises at least two sensors 210 and 220. The at least two sensors 210 and 220 are of different sensor type. For example, the at least two sensors 210 and 220 may comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor. The at least two sensors 210 and 220 configured to monitor (capture) the scene and generate the measurement data 101. The apparatus 100 is communicatively coupled to the at least two sensors 210 and 220 and configured to receive the measurement data 101 from the at least two sensors 210 and 220 for further processing.
The monitoring system 200 provides a compact system with improved anomaly detection. The monitoring system 200 may, e.g., be used for monitoring residential areas or industrial areas.
Fig. 3 schematically illustrates a computing cloud 300. The computing cloud 300 comprises the apparatus 100 for anomaly detection as described above. Furthermore, the computing cloud 300 comprises interface circuitry 310 coupled to the apparatus 100. The interface circuitry 310 is configured to communicatively couple to at least two sensors 320 and 330 of different sensor type monitoring the scene (e.g., via the Internet). The at least two sensors 320 and 330 are external to the computing cloud 300 (i.e., not part of the computing cloud 300). For example, the at least two sensors 320 and 330 may comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor. The interface circuitry 310 is further configured to receive the measurement data 101 from the at least two sensors 320 and 330. The apparatus 100 is configured to receive the measurement data 101 from the interface circuitry 310 for further processing. The computing cloud 300 allows to provide improved anomaly detection as a cloud service.
For further highlighting the anomaly detection described above, Fig. 4 illustrates a flowchart of a method 400 for anomaly detection. The method 400 comprises receiving 402 measurement data of at least two different types of sensors monitoring a scene (e.g., measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of- flight sensor, a radar sensor and an audio sensor). In addition, the method 400 comprises generating 404 semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model. The method 400 further comprises determining 406 whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformerbased multimodal machine-learning model. Additionally, the method 400 comprises performing 408 a predefined action if it is determined that an anomaly occurs in the scene.
Analogously to what is described above, the method 400 provides improved anomaly detection.
As described above with reference to Fig. 2 and Fig. 3, the proposed technology may be used in various ways and implementations for anomaly detection. In some examples of the method 400, receiving 402 the measurement data and generating 404 the semantic tokens may, e.g., be performed at an edge device. Determining 406 whether an anomaly occurs in the scene may be performed at a computing cloud communicatively coupled to the edge device. Compared to a centralized network element like the computing cloud, the edge device is a local device processing data at the periphery (“edge”) of a network. The local processing of the measurement data at the edge device allows to enhance privacy and security as the amount of sensitive information transmitted over the network to the computing cloud is reduced. Furthermore, by only transmitting the semantic tokens rather than the measurement data, the network infrastructure is used more efficiently.
More details and aspects of the method 400 are explained in connection with the proposed technique or one or more examples described above (e.g., Fig. 1 to Fig. 3). The method 400 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above.
For example, as described above, the proposed technology is able to have multiple formats such as, e.g., RGB, depth, infrared etc. simultaneously as input. The proposed technology is not a traditional rule based method where only predefined situations can be detected and reported. For example, a fire detection system will only detect whether there is a fire in the frame but can’t do anything else. Since the proposed technology is based on scene understanding and multimodality, it can process various abnormal detection scenarios with different input formats which makes it universally helpful to any kind of places, from living room to factory, from good light condition to no light condition. Since the proposed technology can not only take RGB data as input for abnormality detection, it also has the ability of privacy preserving anomaly detection compared to traditional methods.
The following examples pertain to further embodiments:
(1) An apparatus for anomaly detection, the apparatus comprising processing circuitry configured to: receive measurement data of at least two different types of sensors monitoring a scene; generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model; and perform a predefined action if it is determined that an anomaly occurs in the scene.
(2) The apparatus of (1), wherein the predefined action is one or more of causing output of a notification about the anomaly on a user terminal, causing output of an alarm, causing recording of the scene, causing notification of a security service and causing a machine in the scene to change its operation.
(3) The apparatus of (1) or (2), wherein the trained transformer-based multimodal machine-learning model is trained to detect different types of anomalies.
(4) The apparatus of any one of (1) to (3), wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scenes.
(5) The apparatus of any one of (1) to (4), wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scene settings. (6) The apparatus of any one of (1) to (5), wherein the encoder comprises individual encoder portions for the different types of sensors, and wherein the processing circuitry is configured to generate the semantic tokens using the respective encoder portion for the measurement data of the different types of sensors.
(7) The apparatus of (6), wherein the encoder comprises individual encoder portions for measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
(8) A monitoring system comprising: the apparatus for anomaly detection according to any one of (1) to (7); and at least two sensors of different sensor type configured to monitor the scene and generate the measurement data, wherein the apparatus is communicatively coupled to the at least two sensors.
(9) The monitoring system of (8), wherein the at least two sensors comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
(10) A computing cloud, comprising: the apparatus for anomaly detection according to any one of (1) to (7); and interface circuitry configured to: communicatively couple to at least two sensors of different sensor type monitoring the scene; and receive the measurement data from the at least two sensors.
(11) A method for anomaly detection, comprising: receiving measurement data of at least two different types of sensors monitoring a scene; generating semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determining whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machinelearning model; and performing a predefined action if it is determined that an anomaly occurs in the scene.
(12) The method of (11), wherein receiving the measurement data and generating the semantic tokens is performed at an edge device, and wherein determining whether an anomaly occurs in the scene is performed at a computing cloud communicatively coupled to the edge device.
(13) The method of (11) or (12), wherein the measurement data comprises measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
(14) The method of any one of (11) to (13), wherein the predefined action is one or more of causing output of a notification about the anomaly on a user terminal, causing output of an alarm and causing a machine in the scene to change its operation.
(15) The method of any one of (11) to (14), wherein the trained transformer-based multimodal machine-learning model is trained to detect different types of anomalies.
(16) The method of any one of (11) to (15), wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scenes.
(17) The method of any one of (11) to (16), wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scene settings.
(18) The method of any one of (11) to (17), wherein the encoder comprises individual encoder portions for the different types of sensors, and wherein the semantic tokens are generated using the respective encoder portion for the measurement data of the different types of sensors.
(19) A non-transitory machine-readable medium having stored thereon a program having a program code for performing the method according to any one of (11) to (18), when the program is executed on a processor or a programmable hardware.
(20) A program having a program code for performing the method according to any one of (11) to (18), when the program is executed on a processor or a programmable hardware.
The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example. Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and/or contain machineexecutable, processor-executable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and/or be broken up into several sub-steps, -functions, -processes or -operations.
If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.

Claims

Claims What is claimed is:
1. An apparatus for anomaly detection, the apparatus comprising processing circuitry configured to: receive measurement data of at least two different types of sensors monitoring a scene; generate semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determine whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machine-learning model; and perform a predefined action if it is determined that an anomaly occurs in the scene.
2. The apparatus of claim 1 , wherein the predefined action is one or more of causing output of a notification about the anomaly on a user terminal, causing output of an alarm, causing recording of the scene, causing notification of a security service and causing a machine in the scene to change its operation.
3. The apparatus of claim 1, wherein the trained transformer-based multimodal machine-learning model is trained to detect different types of anomalies.
4. The apparatus of claim 1, wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scenes.
5. The apparatus of claim 1, wherein the trained transformer-based multimodal machine-learning model is trained to detect anomalies in different types of scene settings.
6. The apparatus of claim 1, wherein the encoder comprises individual encoder portions for the different types of sensors, and wherein the processing circuitry is configured to generate the semantic tokens using the respective encoder portion for the measurement data of the different types of sensors.
7. The apparatus of claim 6, wherein the encoder comprises individual encoder portions for measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
8. A monitoring system comprising: the apparatus for anomaly detection according to claim 1; and at least two sensors of different sensor type configured to monitor the scene and generate the measurement data, wherein the apparatus is communicatively coupled to the at least two sensors.
9. The monitoring system of claim 8, wherein the at least two sensors comprise one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
10. A computing cloud, comprising: the apparatus for anomaly detection according to claim 1; and interface circuitry configured to: communicatively couple to at least two sensors of different sensor type monitoring the scene; and receive the measurement data from the at least two sensors.
11. A method for anomaly detection, comprising: receiving measurement data of at least two different types of sensors monitoring a scene; generating semantic tokens indicating features extracted from the measurement data using an encoder of a trained transformer-based multimodal machine-learning model; determining whether an anomaly occurs in the scene by analyzing the semantic tokens using a self-attention mechanism of the trained transformer-based multimodal machinelearning model; and performing a predefined action if it is determined that an anomaly occurs in the scene.
12. The method of claim 11, wherein receiving the measurement data and generating the semantic tokens is performed at an edge device, and wherein determining whether an anomaly occurs in the scene is performed at a computing cloud communicatively coupled to the edge device.
13. The method of claim 11 , wherein the measurement data comprises measurement data of one or more of an RGB sensor, an infrared sensor, a depth sensor, a time-of-flight sensor, a radar sensor and an audio sensor.
14. The method of claim 11, wherein the predefined action is one or more of causing output of a notification about the anomaly on a user terminal, causing output of an alarm and causing a machine in the scene to change its operation.
15. The method of claim 11, wherein the trained transformer-based multimodal machinelearning model is trained to detect different types of anomalies.
16. The method of claim 11, wherein the trained transformer-based multimodal machinelearning model is trained to detect anomalies in different types of scenes.
17. The method of claim 11, wherein the trained transformer-based multimodal machinelearning model is trained to detect anomalies in different types of scene settings.
18. The method of claim 11 , wherein the encoder comprises individual encoder portions for the different types of sensors, and wherein the semantic tokens are generated using the respective encoder portion for the measurement data of the different types of sensors.
19. A non-transitory machine-readable medium having stored thereon a program having a program code for performing the method according to claim 11 , when the program is executed on a processor or a programmable hardware.
20. A program having a program code for performing the method according to claim 11, when the program is executed on a processor or a programmable hardware.
PCT/EP2025/068995 2024-07-05 2025-07-03 Apparatus and method for anomaly detection, monitoring system and computing cloud Pending WO2026008772A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP24186875.1 2024-07-05
EP24186875 2024-07-05

Publications (1)

Publication Number Publication Date
WO2026008772A1 true WO2026008772A1 (en) 2026-01-08

Family

ID=91853247

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2025/068995 Pending WO2026008772A1 (en) 2024-07-05 2025-07-03 Apparatus and method for anomaly detection, monitoring system and computing cloud

Country Status (1)

Country Link
WO (1) WO2026008772A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11921824B1 (en) * 2021-03-29 2024-03-05 Amazon Technologies, Inc. Sensor data fusion using cross-modal transformer
US20240212350A1 (en) * 2022-06-07 2024-06-27 Sri International Spatial-temporal anomaly and event detection using night vision sensors

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11921824B1 (en) * 2021-03-29 2024-03-05 Amazon Technologies, Inc. Sensor data fusion using cross-modal transformer
US20240212350A1 (en) * 2022-06-07 2024-06-27 Sri International Spatial-temporal anomaly and event detection using night vision sensors

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
WEI DONGLAI ET AL: "MSAF: Multimodal Supervise-Attention Enhanced Fusion for Video Anomaly Detection", IEEE SIGNAL PROCESSING LETTERS, vol. 29, 2 November 2022 (2022-11-02), USA, pages 2178 - 2182, XP093177021, ISSN: 1070-9908, DOI: 10.1109/LSP.2022.3216500 *

Similar Documents

Publication Publication Date Title
Rodríguez et al. Multi-agent information fusion system to manage data from a WSN in a residential home
CN118801234B (en) A high voltage control cabinet with partial discharge detection device
US11631306B2 (en) Methods and system for monitoring an environment
CN117726162A (en) A community risk level assessment method and system based on multi-modal data fusion
Sami et al. Forecasting failure rate of IoT devices: A deep learning way to predictive maintenance
US20250068885A1 (en) Integrated Multimodal Neural Network Platform for Generating Content based on Scalable Sensor Data
US12299988B2 (en) Concept for detecting an anomaly in input data
CN118038619A (en) A multimodal federated learning method and system for intelligent fire monitoring and early warning
WO2026011846A1 (en) Dual-layer perception-based distribution network cable environment perception method and system
Arslan et al. Sound based alarming based video surveillance system design
CN114647554A (en) Performance data monitoring method and device of distributed management cluster
Zou et al. Geoai for disaster response
CN116597501B (en) Video analytics algorithms and edge devices
WO2026008772A1 (en) Apparatus and method for anomaly detection, monitoring system and computing cloud
CN118570942B (en) A building intelligent security monitoring system
CN114387391A (en) Safety monitoring method, device, computer equipment and medium for substation equipment
CN116597603B (en) An intelligent fire alarm system and its control method
US11985447B1 (en) Analytics pipeline management systems for local analytics devices and remote server analytics engines
US20240395125A1 (en) Bandwidth and power-optimized hybrid high-resolution/low-resolution sensor method for a predictive analytics system
US12499400B2 (en) Sensor input and response normalization system for enterprise protection
US20260127958A1 (en) Operational information of a fire device
Nagendra et al. Disaster Monitoring and Intelligent Evacuation Planning System using Edge Computing and Simulation
CN120973845B (en) Police risk multistage early warning system based on large model cooperative computing
Alamgir Performance analysis internet of things based on sensor and data analytics
US20250384510A1 (en) Creating and updating a data variable in a computer-aided dispatch system

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25739841

Country of ref document: EP

Kind code of ref document: A1