EP4189673A1 - Computerimplementiertes verfahren und computerprogramm zum maschinellen lernen einer robustheit eines akustischen klassifikators, akustisches klassifikationssystem für automatisiert betreibbare fahrsysteme und automatisiert betreibbares fahrsystem - Google Patents
Computerimplementiertes verfahren und computerprogramm zum maschinellen lernen einer robustheit eines akustischen klassifikators, akustisches klassifikationssystem für automatisiert betreibbare fahrsysteme und automatisiert betreibbares fahrsystemInfo
- Publication number
- EP4189673A1 EP4189673A1 EP21742385.4A EP21742385A EP4189673A1 EP 4189673 A1 EP4189673 A1 EP 4189673A1 EP 21742385 A EP21742385 A EP 21742385A EP 4189673 A1 EP4189673 A1 EP 4189673A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- acoustic
- classifier
- driving system
- interference
- acoustic classifier
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
Definitions
- the invention relates to a computer-implemented method and a computer program for machine learning of a robustness of an acoustic classifier, an acoustic classification system for driving systems that can be operated in an automated manner, and a driving system that can be operated in an automated manner.
- DE 10 2020 205 825.3 generally discloses a system for detecting, avoiding and protecting against fraud by ADAS functions.
- the control system disclosed there is set up and intended for use in a motor vehicle, based on environmental data obtained from at least one environmental sensor and/or signal receiver assigned to the motor vehicle: lanes, roadway boundaries, roadway markings, other motor vehicles, traffic signs, light signals (systems) and/or other objects in an area in front of, to the side of and/or behind the motor vehicle.
- the environment sensor and/or signal receiver is set up to provide the control system with the environment data reflecting the area in front of, to the side of and/or behind the motor vehicle.
- the control system is at least set up and intended to assign the environmental data provided to at least one traffic category using a machine learning classifier, each of the at least one traffic category being one of several categories of potential driving situations, and the machine learning system being previously known environmental data has been trained with already assigned traffic categories. If the at least one traffic category was assigned incorrectly to the provided environment data, a correction signal is received which correctively indicates which at least one traffic category the provided environment data is correctly assigned to, the correction signal preferably originating from a user input.
- the machine learning classifier is based on the provided environmental data and the corrected at least one traffic category trained. The motor vehicle is controlled accordingly to the corrected at least one traffic category.
- DE 10 2020 205 825.3 discloses a front camera, rear camera, side camera, a radar sensor, a lidar sensor, an ultrasonic sensor and/or an inertial sensor as surroundings sensors.
- driving systems with AD/ADAS functions should also be able to record, analyze and evaluate acoustic signals outside the driving system.
- a human driver also uses this sense of hearing to a not inconsiderable extent, for example to determine the arrival and location of an emergency vehicle.
- the acoustic assessment of a human driver about the road condition for example wetness due to a changed background noise, should be taken over by an automated driving system.
- noise is recorded, analyzed and evaluated in the vehicle interior. Examples are voice commands from the driver, rattling noises from the driving system or noises that indicate the condition of the driver and the occupants.
- the invention was based on the object, on the one hand, of making the acoustic sensors of the driving system robust against all types of attacks and, on the other hand, of improving the general ability of my generalization of the recognition performance and classification performance of the acoustic sensor.
- the invention provides a computer-implemented method for machine learning a robustness of an acoustic classifier.
- a driving system is automatically controlled depending on classifications and/or localizations of the acoustic classifier.
- the procedure includes the steps:
- the invention provides a computer program for machine learning a robustness of an acoustic classifier.
- the program includes program instructions that cause a computer to execute a method according to the invention when the program is run on the computer.
- the program instructions are written, for example, in an object-oriented programming language, such as C++.
- the invention provides an acoustic classification system for driving systems that can be operated automatically, for classifying and/or localizing acoustic events in the exterior and/or interior of the driving system.
- the acoustic classification system includes an acoustic sensor and an acoustic classifier, wherein the acoustic classifier has learned, according to a method according to the invention, to classify and/or localize acoustic events in a robust manner against disturbances.
- the invention provides a driving system that can be operated automatically, comprising an acoustic classification system according to the invention, a control unit for automated driving and actuators for longitudinal and/or lateral guidance of the driving system.
- the control device determines regulation and/or control signals and provides these to the actuators. Disturbances are added to the first input data of the acoustic classifier in the form of signals from a loudspeaker arranged outside the driving system, a carrier signal from a loudspeaker arranged inside the driving system and/or from driving system parts that produce noise.
- Sound-producing driving system parts include, for example, an infected water pump that produces sounds to perform a targeted attack.
- Machine learning is a technology that teaches computers and other data processing devices to perform tasks by learning from data, rather than being programmed to do the tasks.
- RASES can be used to increase the robustness against any interfering signals, including noise or attacks. Attacks include deception.
- Increasing robustness against noise includes making an acoustic classifier robust against overfitting by RASES. RASES thus provides an improved, generalized acoustic recognition system that more reliably and correctly recognizes acoustic signals that have not been trained before, in particular noise signals.
- An acoustic classifier is an artificial intelligence comprising software and/or hardware components that can be trained and/or trained to recognize, classify and/or localize sounds and/or speech.
- the Acoustic signals are classified, for example, into the categories of rescue vehicle, falling branch, children playing, deer crossing, grinding noises.
- the acoustic classifier evaluates a continuous data stream from the driving system acoustic sensor.
- the acoustic classifier classifies overlapping time signals, for example the last 1s every 0.2ms.
- the acoustic classifier represents a sense of hearing for a driving system. For example, the acoustic classifier determines the arrival and/or the position of an emergency vehicle depending on a siren signal. This determination is made available as a signal to a control unit of the driving system, for example an ADAS/AD domain ECU, ie an electronic control unit for assisted or automated/autonomous driving. Depending on the classification and/or localization of the acoustic classifier, the control device determines control and/or regulation signals for actuators for longitudinal and/or lateral guidance of the driving system in order to automatically control the driving system.
- a control unit of the driving system for example an ADAS/AD domain ECU, ie an electronic control unit for assisted or automated/autonomous driving.
- the control device determines control and/or regulation signals for actuators for longitudinal and/or lateral guidance of the driving system in order to automatically control the driving system.
- the software components of the acoustic classifier are available, for example, as program commands in the programming language Python or TensorFlow.
- the analysis of the first input data is carried out, for example, with the Python program package LibROSA, which includes routines for music and audio analysis.
- the hardware components include GPUs and/or tensor processing units with a microarchitecture for parallelized processing of tasks and execution of matrix multiplications. This makes the training and use of a trained artificial intelligence more efficient.
- the driving system includes cars, commercial vehicles, trucks, buses, people movers, robots such as industrial robots, drones, rail vehicles, ships and airplanes.
- the driving system includes technical equipment for operating the driving system in accordance with SAE J3016 levels 1 to 5.
- the driving system is a road vehicle with an automation level SAE J3016 levels 2+ to 5.
- the first input data includes acoustic signals from the driving system acoustic sensor. Compared to other acoustic sensors, the driving system acoustic sensor is particularly suitable for automotive use.
- the driving system acoustic sensor when used outside of the driving system, includes a protective grille to protect against the ingress of foreign bodies, an acoustically permeable, hydrophobic and/or lipophobic membrane to protect against splash water and grease, and a flow bypass to prevent fluids or to guide foreign bodies out of the sensor.
- the driving system acoustic sensor is also used in the interior of the driving system.
- the disturbances for deception detection, avoidance and/or protection correspond to signals that a disturber, ie an attacker, calculates and plays back in order to deceive the acoustic classifier. Sound and speech recognition are vulnerable.
- the basic idea for deceiving an acoustic classifier is that a loudspeaker is used through which interference signals are played back.
- the classifier is supposed to be deceived by these interference signals, so that the original/actual event is not recognized or another desired event is recognized, although in reality no such acoustic event has occurred.
- Existing loudspeakers in the vehicle can be used for this purpose, for example infotainment or mobile phones, or loudspeakers can be set up in a targeted manner at the desired location, for example a residential area, the edge of a forest or a bus stop.
- spurious signals are either integrated into carrier signals or exist as a separate signal.
- An example of integration into carrier signals is introducing the interference into music.
- the modified music signal is then uploaded to a popular platform, such as YouTube or Spotify, and played back via the driving system's infotainment system.
- a large number of attacks are carried out in which the acoustic classifiers are fooled into recognizing an event when in reality there is no event. This can cause significant damage to a mass of users/customers.
- Another case is that of an inconspicuous interfering signal, which for humans is only a faint noise is recognizable, is played, whereby the acoustic classifier events are given before or the detection of actually happening events is prevented ver.
- Such an attack is dangerous because the human ear cannot detect the interference signals. As a result, the occupants would not notice the ongoing attack or only after the driving system had already initiated reactionary measures, such as braking if a branch fell or children were playing.
- the human ear cannot detect the present attack because either the volume of the interference signal is too low or the interference is only applied to certain frequencies which are masked by neighboring louder frequencies for the human ear.
- the invention includes untargeted and targeted attacks.
- an untargeted attack the attacker's goal is to introduce a perturbation to get the acoustic classifier to predict a class other than the correct one. It does not matter which class is predicted instead of the correct class, in contrast to a targeted attack, where the attacker wants to ensure that a specific target class is predicted instead of the correct class.
- RASES prevents this attack in that the acoustic classifier learns, through the method according to the invention, to be robust against targeted or naturally occurring disturbances and to carry out the classification of the actual, real acoustic event correctly.
- the acoustic classifier learns, through the method according to the invention, to be robust against targeted or naturally occurring disturbances and to carry out the classification of the actual, real acoustic event correctly.
- the attacker has to calculate the interference depending on the other acoustic signals in the target environment, for example residential area, edge of the forest, interior or busy street.
- exemplary signals can be accepted, which reflect the real situation as best as possible, and the generation of the interference signal can be carried out for several of these signals, for example 1000 to 100,000 exemplary signals. This allows the attacker to ensure that the calculated Noise actually deceives the acoustic classifier, regardless of any other acoustic signals.
- gradient-based methods can be used to optimize the jamming signal depending on the classification of the system.
- One method is, for example, the projected gradient descent method, abbreviated PGDM, in which a step in the positive direction of the gradient of a loss function of the acoustic classifier, also called loss function, is repeatedly carried out as a function of the input data.
- PGDM projected gradient descent method
- Corresponding attack methods are disclosed in Section 2.2 of https://arxiv.org/pdf/1611.01236.pdf.
- the attacker has no information about the acoustic classifier used, it is initially not possible to use gradient-based methods because the necessary gradients cannot be calculated. In order to still be able to use these methods, the attacker can try to obtain information about the acoustic classifier used.
- an attacker can train a system that is as identical as possible, preferably on similar training data. Then this system can be used to calculate an interfering signal.
- this interference signal can also be used to deceive the acoustic classifier that is actually being attacked.
- Techniques also exist to ensure that a transmittable jamming signal is found. For example, several substitute models can be trained on different data, which are incorporated by the loss function used in order to calculate a uniform interference signal for all models.
- model stealing attacks which have the purpose of obtaining information about an artificial intelligence.
- an attacker In order to be able to carry this out, an attacker only needs the input data of the acoustic can change the classifier, for example play a test signal, and then be able to observe the output values of the acoustic classifier.
- queries By cleverly combining different input values and testing, also called queries, how the acoustic classifier reacts to them, such attacks can collect information about how the acoustic classifier works and how it can be deceived.
- model stealing is disclosed in https://arxiv.org/pdf/1802.05351.pdf.
- Types of attack are also known as pure black-box attacks without gradient information, which also do not replicate/retrain the system locally. Instead, clever decisions are made based on the current value of the loess function as to how the current disturbance must be changed in order to fool the artificial intelligence, see https://arxiv.org/pdf/1712.04248.pdf.
- the method according to the invention also achieves and/or increases robustness against this type of attack.
- RASES prevents any of these attacks and ensures the correct functionality of the acoustic classifier even though such interference signals are present and an attack is attempted.
- this is achieved in that, during the training, disturbances are obtained as a function of the first input signals for deception detection, avoidance and/or protection and/or for improving a recognition and/or classification performance of the acoustic classifier, and these disturbances are also trained, wherein an audibility of the disturbances is reduced iteratively or successively.
- certain hyper parameters of the acoustic classifier are determined as best as possible, for example the initial maximum strength of the interference signal or target sequence. Depending on these parameters, there are various changes in the robustness and accuracy of the resulting acoustic classifier after training is complete.
- Deception and/or attacks with the aim of attacking the outward-facing acoustic sensors have the following effect, for example: • Non-recognition and/or incorrect localization of noise sources to be recognized,
- Sources of noise related to the exterior include:
- Deception and/or attacks with the attack target of the inward-facing acoustic sensors have the following effect, for example:
- RASES makes the acoustic classifier robust against these illusions by expanding the training of the acoustic classifier with these disturbances. RASES thus makes an acoustic classifier for the exterior and interior robust.
- Interior noise sources include:
- the attacks also include a target other than the ego driving system, for example a system that is connected to the driving system in some way, for example cloud storage, similar to the introduction of computer viruses, trojans, worms.
- a target other than the ego driving system for example a system that is connected to the driving system in some way, for example cloud storage, similar to the introduction of computer viruses, trojans, worms.
- Augmenting the training of the acoustic classifier with these perturbations further increases the fundamental ability of the acoustic classifier's ability to generalize, since "accidental manipulation" and deliberate attacks correspond in some ways exactly to the ability to generalize. This increases the recognition and/or classification performance.
- the acoustic classifier learns to defend itself against the attacks described above defend.
- the resulting combinations represent extended or augmented training data for the acoustic classifier.
- the machine learning of these combinations is a so-called adversarial training, i.e. the augmentation of the first input data with interference signals, which an attacker would use to fool the acoustic classifier.
- the disturbances are recalculated during the training for each input signal and are always adapted to the current parameters of the acoustic classifier.
- the interference signals are added to the original data, but the ground truth class is not changed.
- Adversarial training is disclosed in https://arxiv.org/pdf/1706.06083.pdf.
- Batches are groups of input data of equal size.
- the training can be carried out per batch. When all batches have gone through the artificial intelligence once, an epoch is complete. An epoch denotes a complete run through of all input data.
- the number of training epochs and batches is a parameter for training the artificial intelligence. For example, each batch consists of 50% original and 50% corrupted data. However, other distributions are also conceivable, e.g.: 20% original, 40% attack method 1, for example gradient-based, 40% attack method 2, for example model stealing.
- the adversarial training is used conceptually with further augmentation strategies, for example with spectrogram augmentation, see https://arxiv.org/pdf/1904.08779.pdf.
- the original signal can also be overlaid with further realistic noise signals in order to be able to reflect a real scenario even better and thus further increase the accuracy of the acoustic classifier under non-optimal conditions.
- a loss function is minimized while complying with the condition that the interference is smaller than a predetermined interference.
- the loss function also called the combined loss function, includes as first part the disturbances and as a second part a loss function of the acoustic classifier extended with the disturbances.
- the extended loss function is minimized by an interferer's intended classification of the acoustic classifier.
- PGDM can also be used for an attack in the audio sector.
- this method does not work well with the increased non-linearities that are caused by pre-processing and the possible massive use of recurrent layers in the acoustic classifier. It is therefore often not possible, particularly in the case of long sequences, for example speech recognition, to find a suitable disturbance which is inaudible to a human being.
- x means: vector with raw, first input data, d: generic disturbance,
- a y class predicted by the acoustic classifier
- t target class of the attacker
- the target class is the class that the attacker will ensure to be predicted by the acoustic classifier instead of the correct class.
- the main goal is to minimize the difference between the magnitude of the interference and the magnitude of the first input data, so that the interference is not audible to a human when it is added to the input data.
- the acoustic classifier must be successfully deceived and the targeted class, or sequence of acoustic units, is predicted.
- this optimization problem is very difficult to solve with methods based on normal gradients, since according to one aspect of the invention the classification function f(-) is represented by an artificial neural network which is very strongly non-linear.
- L loss function
- cc tradeoff parameters e: maximum allowed disturbance.
- the first part of the combined loss function causes a disturbance d with the lowest possible strength to be found and the second part causes the disturbance found to also successfully disturb the acoustic classifier.
- Successful disruption is ensured by minimizing the value of the acoustic classifier's loss function L(•), thereby ensuring that it tends to zero.
- the parameter a acts as an opportunity to set the tradeoff between successful disruption and imperceptibility and can therefore be adapted to the given circumstances and to the objective.
- the presence of the necessary condition provides an additional constraint to ensure that the interference is evenly distributed across the input signal and does not have a very high outlier in some regions that would be heard by humans, even though the first term of the loess function , which is the squared ⁇ 2 -norm of the perturbation, is small.
- further terms are added, which say, for example, that the interference should be added mainly on frequencies that are not audible to a human.
- this optimization problem is solved with gradient descent. Therefore, the combined loss function is minimized until a perturbation d is found that successfully perturbs the acoustic classifier and causes it to predict the target class t.
- a higher value is initially used for the maximum strength e with which the disturbance can be heard by a human.
- the attacker's maximum allowed strength e is reduced and the optimization continued. This process continues iteratively until a predetermined number of iterations has been completed. Consequently, during the optimization, the audibility is reduced more and more, but the deceptive character of the disturbance remains, so that the acoustic classifier is still correctly deceived.
- the first input data includes raw data from the driving system acoustic sensor, filtered raw data and/or a representation of the raw data in a time-frequency range.
- raw acoustic signals can be used as input data for the acoustic classifier without pre-processing, but this currently results in lower classification accuracies.
- the raw data are filtered with low-pass or band-pass filters in order to specifically blind or amplify noises depending on the situation.
- the representation of the raw data in the time-frequency domain is based, for example, on pre-processing the raw data with a short-time Fourier transformation, whereby different window types (Hann, Blackman) with different parameters (window width, hop distance) are used.
- Window types Hann, Blackman
- window width window width
- hop distance a short-time Fourier transformation
- the result is a time-frequency picture in which the energy is displayed in different frequencies over time. If more than one driving system acoustic sensor is evaluated, there is a signal for each sensor which is transformed independently. In this case, therefore, there are several time-frequency images, analogous to an RGB image in which three color channels are then present).
- the pre-processing can contain noise reduction in order to improve the signal quality of the acoustic signals.
- noise reduction in order to improve the signal quality of the acoustic signals.
- mechanisms can be used which exploit the different propagation times of acoustic waves to the individual sensors, for example beamforming or source separation. These methods can themselves be based on artificial intelligence.
- It is also possible to remove noise from the time signals for example using a denoising autoencoder or Wiener filter, before these signals are transformed into the time-frequency domain.
- algorithmic, statistical methods that weight the time-frequency features and try to assign a low weight to features with low speech energy.
- Raw acoustic signals can differ significantly even though they reflect the same context, such as noise or speech. For example, the current emotional state of a speaker leads to differently emphasized signals.
- the pre-processing generates features first, which have a higher Have invariance to such different signals of the same basic event.
- the method according to the invention is extended such that the attacker no longer adds the interference signal to the original input data. Instead, the interference signal is added to a representation in the time-frequency domain. It is also possible to add the interference signal to any other representation after the individual steps in the pre-processing.
- the first input data includes a representation of raw data from the driving system acoustic sensor in a time-frequency range. Masking adds the interference at low-energy frequencies.
- the masking restricts the features that the used attacker is allowed to attack during training. As a result, the attacker can only add the interference signal to a subset of all available features during training. According to the invention, the masking is used to prevent attacks on relevant features with high speech energy during training. As a result, during training, the attacker can only add the interference signal to features that receive little information about the existing speech energy.
- the jamming signal must be added on low-energy frequencies.
- the acoustic classifier can be improved more efficiently and effectively against general real-world attacks compared to the case of normal adversarial training attacking the raw speech signal.
- the acoustic classifier is thus specifically trained to utilize frequencies with high energy and to be more robust against interference from less important frequencies.
- the masking can be transferred analogously to noise detection, in that only features that are not relevant to the respective acoustic event may be disturbed by the attacker.
- the acoustic classifier will learn during training not to use the disturbed features and rely on the remaining features. Since these are particularly relevant and meaningful with regard to the existing acoustic events, the existing language, the robustness increases further because the acoustic classifier learns to make its decision mainly on the basis of these features.
- the acoustic sensor is arranged in the interior of the driving system when used, and the acoustic classifier is robust against disturbing noises from
- Noise from damage to your own driving system including rattling, squeaking, grinding, fire noise,
- the acoustic sensor is arranged outside of the driving system when it is used, and the acoustic classifier is robust against disturbing noises from
- Control commands to the driving system including opening of trunk, doors, identification of the driver.
- the acoustic classifier includes an artificial neural network for noise/speech recognition.
- the artificial neural network includes layers of convolutional networks, recurrent layers, fully connected layers and/or an encoder-decoder structure.
- Convolutional networks include filter layers, also called kernels, to minimize dimensions of respective input data, and discretization layers, for example maxpooling kernels, to further reduce dimensions of respective input data. Using these layers, new features are extracted from the input data. Contextual sequence information is evaluated by means of recurrent layers, comprising GRU, BGRU, LSTM and BLSTM. Finally, fully connected layers can be used to output the final probabilities per event class.
- An encoder-decoder structure defines an encoded context/summary vector. An encoder-decoder structure is advantageous for speech recognition. Batch normalization or sequence normalization layers are used as additional components to speed up training and increase generalization.
- RASES is independent of the specific network architecture and the existing hyper parameters, such as regularization, batch size, number of epochs, activations, classes, further data augmentation and/or dropout, and optimization settings, such as loss function, optimizer, LR schedule.
- an attack on an acoustic classifier is prevented or at least made more difficult by the invention.
- the acoustic classifier can therefore not be deceived by an attacker and also works correctly when there is an interference signal which is actually intended to deceive the acoustic classifier.
- the invention increases the generalizability and thus the recognition rates under any interference. This improves the robustness against natural disturbances, such as street noise or conversations. This is particularly relevant as acoustic classifiers operate under widely varying environments and high robustness against unknown noise types/sounds is required.
- the improvements are made possible by the fact that RASES teaches the acoustic classifier to rely on features that are representative of the relevant acoustic energy in the input data.
- the acoustic classifier focuses on features that are meaningful and extracts information from important features. noisysy features are used less, making the acoustic classifier less sensitive to various perturbations, natural and adversarial.
- a further advantage of the invention is that the increase in robustness is carried out by a synthetic augmentation of the training data. It is not necessary to record new data in reality, which depict all possible interference signals. On the one hand, this is hardly possible and, on the other hand, it requires greater effort to record as representative a quantity of noise signals as possible.
- RASES can be extended to include regression models, which are used, for example, for localization/distance estimation. It is possible that an attacker can also fool such artificial intelligences. A simple increase in robustness is possible with the help of RASES, since in this case the original data set can also be augmented with specially generated interference signals. The RASES concept can therefore be transferred to all acoustic artificial intelligences that are learned using training data.
- FIG. 1 shows a schematic representation of a normal training course of an artificial intelligence
- 4 shows an exemplary embodiment of a mask
- 5 shows an embodiment of an acoustic classifier for speech recognition
- FIG. 6 shows a schematic representation of exemplary access points of an attacker.
- the existing training data is shown to an artificial intelligence, for example an artificial neural network, and the loss function is minimized. This process is performed iteratively over multiple epochs of the training data. As a result, the artificial intelligence learns to correctly classify the existing data.
- an artificial intelligence for example an artificial neural network
- the original training data are augmented. This is done by an attacker who specifically calculates an interference signal S, which leads to the current acoustic classifier AK being deceived. An iterative attack is used for this.
- the optimization-based method according to the invention is used to attack the acoustic classifier AK.
- This introduces a combined loss function, which expresses how well the current interference signal S deceives the acoustic classifier and how audible this interference signal S is for humans.
- This combined loss function is then solved using Gradient Descent. Typically, the focus is first on finding a valid interference signal S, even if this is clearly audible to a human.
- the strength of this interference signal S is then reduced, resulting in a valid interference signal that is not recognizable to humans.
- the original data is expanded with the resulting interference signals.
- the resulting augmented training data is any combination of original and challenged/perturbed data. on this data a normal training iteration is then performed to minimize the loss function and thereby robustly train the acoustic classifier.
- V1 Provision of first input signals by means of a driving system acoustic sensor for the acoustic classifier AK,
- V2 Obtaining disturbances S as a function of the first input signals for deception detection, avoidance and/or protection and/or for improving a recognition and/or classification performance of the acoustic classifier AK, the audibility of the disturbances being reduced,
- V3 obtaining second input data from an addition of the first input data and the disturbances
- FIG. 3 shows an exemplary transformation in the time-frequency domain of the sentence: "The seven units to be offered for sale have a work force of about twenty thousand.”
- FIG. 3 shows an exemplary representation of FBank features. With Fourier transformation, a signal in the time domain is broken down into its frequencies. The acoustic events are separated into time frames and a Fourier transform is applied to each time frame. The frequency axis is then displayed logarithmically and the amplitudes in decibels. A spectrogram results. In order to obtain a Mel spectrogram as shown in Fig. 3, the frequency scale f of the spectrogram is transformed to Mel scale m according to, for example,
- FIG. 4 shows a masking according to the invention of the mel spectrogram from FIG. 3, the data from FIG. 3 having been compared with a noise image.
- 5 shows the structure of a system for speech recognition.
- the time signal x is pre-processed so that a time-frequency representation F results.
- This is used as input data for an acoustic model.
- This model is trained data-driven and represented by deep artificial neural networks, called DNN, or a mix of DNN and Hidden Markov Models. It outputs a sequence of probabilities of acoustic units comprising letters, phonemes, parts of words, which is combined to form the resulting total words and the word sequence being searched for.
- the network architecture of the acoustic model includes layers of a convolutional network, fully connected layers and recurrent layers. Only the number of output classes is typically significantly larger in order to cover all relevant acoustic units, for example 80-2000. Special loess functions, such as Connectionist Temporal Classification, see https://www.cs.toronto.edu/ ⁇ graves/icml_2006.pdf, are also used.
- the composition is performed using a decoder, which searches for the most probable sequence through the sequence of probability vectors of the acoustic units.
- a beam search decoder is often used with various options, for example with regard to beam width and/or weighting.
- additional a priori information about the formalisms of the processed language can be used. This includes a lexicon that contains legal words and a language model that expresses grammatical dependencies, including probabilities of the next word depending on the previous one.
- the language model can be represented by its own artificial intelligence or by simple probability tables and manually formed decision rules.
- the invention can be applied not only to systems that use this structure, but to all speech recognizers/noise recognizers that are learned from data. Consequently, RASES also applies in this case independently of various hyperparameters of the learned artificial intelligence.
- RASES also applies in this case independently of various hyperparameters of the learned artificial intelligence.
- the attacker attacks before pre-processing the raw data. According to the invention, this is simulated in that the interference signal S is added to the original input data.
- the attacker attacks after preprocessing for example the interference signal is added to a representation in the time-frequency domain.
- the attacker can also add the disruption to each point in the preprocessing, ie between Abs and FBANK, for example, during training.
- V1 -V5 method steps AK acoustic classifier S disturbance x time signal
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Soundproofing, Sound Blocking, And Sound Damping (AREA)
- Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE102020209446.2A DE102020209446A1 (de) | 2020-07-27 | 2020-07-27 | Computerimplementiertes Verfahren und Computerprogramm zum maschinellen Lernen einer Robustheit eines akustischen Klassifikators, akustisches Klassifikationssystem für automatisiert betreibbare Fahrsysteme und automatisiert betreibbares Fahrsystem |
| PCT/EP2021/069321 WO2022023008A1 (de) | 2020-07-27 | 2021-07-12 | Computerimplementiertes verfahren und computerprogramm zum maschinellen lernen einer robustheit eines akustischen klassifikators, akustisches klassifikationssystem für automatisiert betreibbare fahrsysteme und automatisiert betreibbares fahrsystem |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4189673A1 true EP4189673A1 (de) | 2023-06-07 |
Family
ID=76943009
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP21742385.4A Withdrawn EP4189673A1 (de) | 2020-07-27 | 2021-07-12 | Computerimplementiertes verfahren und computerprogramm zum maschinellen lernen einer robustheit eines akustischen klassifikators, akustisches klassifikationssystem für automatisiert betreibbare fahrsysteme und automatisiert betreibbares fahrsystem |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4189673A1 (de) |
| DE (1) | DE102020209446A1 (de) |
| WO (1) | WO2022023008A1 (de) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| AT525938B1 (de) * | 2022-02-24 | 2024-08-15 | Avl List Gmbh | Prüfstandsystem zum Testen eines Fahrerassistenzsystems mit einem Hörschall-Sensor |
| CN115910104B (zh) * | 2022-12-07 | 2026-04-17 | 讯飞智元信息科技有限公司 | 伪造语音检测方法、装置、电子设备和存储介质 |
| CN117993307B (zh) * | 2024-04-07 | 2024-06-14 | 中国海洋大学 | 基于深度学习的地球系统模拟结果一致性评估方法 |
| CN118366472B (zh) * | 2024-04-26 | 2025-01-28 | 东莞野松电子工业有限公司 | 一种音频多模态分类方法、系统及计算机设备 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE102018210489B4 (de) * | 2018-06-27 | 2022-02-24 | Zf Friedrichshafen Ag | Verfahren zum Montieren eines Gehäuses für Akustiksensoren eines Fahrzeuges zum Detektieren von Schallwellen eines akustischen Signals außerhalb des Fahrzeuges auf einem Fahrzeugdach an einer Position einer Dachantenne |
| US11231905B2 (en) * | 2019-03-27 | 2022-01-25 | Intel Corporation | Vehicle with external audio speaker and microphone |
| DE102020205825A1 (de) | 2020-05-08 | 2021-11-11 | Zf Friedrichshafen Ag | System zur Täuschungserkennung, -vermeidung und -schutz von ADAS Funktionen |
-
2020
- 2020-07-27 DE DE102020209446.2A patent/DE102020209446A1/de not_active Ceased
-
2021
- 2021-07-12 WO PCT/EP2021/069321 patent/WO2022023008A1/de not_active Ceased
- 2021-07-12 EP EP21742385.4A patent/EP4189673A1/de not_active Withdrawn
Also Published As
| Publication number | Publication date |
|---|---|
| DE102020209446A1 (de) | 2022-01-27 |
| WO2022023008A1 (de) | 2022-02-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4189673A1 (de) | Computerimplementiertes verfahren und computerprogramm zum maschinellen lernen einer robustheit eines akustischen klassifikators, akustisches klassifikationssystem für automatisiert betreibbare fahrsysteme und automatisiert betreibbares fahrsystem | |
| DE102020205786B4 (de) | Spracherkennung unter verwendung von nlu (natural language understanding)-bezogenem wissen über tiefe vorwärtsgerichtete neuronale netze | |
| EP3938807B1 (de) | Verfahren zur detektion von hindernisobjekten durch ein sensor system mittels neuronaler netze | |
| DE60023517T2 (de) | Klassifizierung von schallquellen | |
| DE60123161T2 (de) | Verfahren und Vorrichtung zur Spracherkennung in einer Umgebung mit variablerem Rauschpegel | |
| DE102018218586A1 (de) | Verfahren, Vorrichtung und Computerprogramm zum Erzeugen robuster automatisch lernender Systeme und Testen trainierter automatisch lernender Systeme | |
| DE102017112992A1 (de) | Trainingsalgorithmus zur kollisionsvermeidung unter verwenden von auditiven daten | |
| DE102016118902A1 (de) | Kollisionsvermeidung mit durch Kartendaten erweiterten auditorischen Daten | |
| DE112017004397T5 (de) | System und Verfahren zur Einstufung von hybriden Spracherkennungsergebnissen mit neuronalen Netzwerken | |
| DE102015109832A1 (de) | Objektklassifizierung für Fahrzeugradarsysteme | |
| DE112019000340T5 (de) | Epistemische und aleatorische tiefe plastizität auf grundlage vontonrückmeldungen | |
| DE102014118450A1 (de) | Audiobasiertes System und Verfahren zur Klassifikation von fahrzeuginternem Kontext | |
| DE102019205543A1 (de) | Verfahren zum Klassifizieren zeitlich aufeinanderfolgender digitaler Audiodaten | |
| DE102020131657A1 (de) | Diagnostizieren eines Wahrnehmungssystems auf der Grundlage der Szenenkontinuität | |
| DE202019105282U1 (de) | Vorrichtung zum Optimieren eines System für das maschinelle Lernen | |
| DE60133537T2 (de) | Automatisches umtrainieren eines spracherkennungssystems | |
| DE102017209585A1 (de) | System und verfahren zur selektiven verstärkung eines akustischen signals | |
| CN118982989A (zh) | 一种基于听觉调制机制和对比学习的单通道语音分离方法及装置 | |
| CN116778963A (zh) | 一种基于自注意力机制的汽车鸣笛识别方法 | |
| DE102025118736A1 (de) | Assistenzsystem zum Steuern einer Fahrzeugkomponente unter Verwendung eines Large-Language-Moduls | |
| DE102019218058B4 (de) | Vorrichtung und Verfahren zum Erkennen von Rückwärtsfahrmanövern | |
| Andric et al. | Ground surveillance radar target classification based on fuzzy logic approach | |
| DE102022119711A1 (de) | Verfahren, System und Computerprogrammprodukt zur Überprüfung von Datensätzen für das Testen und Trainieren eines Fahrerassistenzsystems (ADAS) und/oder eines automatisierten Fahrsystems (ADS) | |
| DE102018117205A1 (de) | Verfahren zum Informieren eines Insassen eines Kraftfahrzeugs über eine Verkehrssituation mittels einer Sprachinformation; Steuereinrichtung; Fahrerassistenzsystem; sowie Computerprogrammprodukt | |
| DE102022203422A1 (de) | Test einer automatischen Fahrsteuerfunktion mittels semi-realer Verkehrsdaten |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20230127 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20250201 |