CA3175103C - Method for sampling-rate recognition of pure voice data, apparatus, and system - Google Patents

Method for sampling-rate recognition of pure voice data, apparatus, and system

Info

Publication number
CA3175103C
CA3175103C CA3175103A CA3175103A CA3175103C CA 3175103 C CA3175103 C CA 3175103C CA 3175103 A CA3175103 A CA 3175103A CA 3175103 A CA3175103 A CA 3175103A CA 3175103 C CA3175103 C CA 3175103C
Authority
CA
Canada
Prior art keywords
voice data
frequency
pure
postulated
sampling rate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
CA3175103A
Other languages
French (fr)
Other versions
CA3175103A1 (en
Inventor
Bingbing Liu
Fei BAO
Kewei Wu
Ruyi LIU
Yang Che
Original Assignee
10353744 Canada Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 10353744 Canada Ltd filed Critical 10353744 Canada Ltd
Publication of CA3175103A1 publication Critical patent/CA3175103A1/en
Application granted granted Critical
Publication of CA3175103C publication Critical patent/CA3175103C/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/18Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Telephonic Communication Services (AREA)
  • Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)
  • Time-Division Multiplex Systems (AREA)

Abstract

A method of recognizing sampling-rate of pure voice data, as well as its apparatus and system are disclosed. The method includes: performing Fourier transform on pure voice data so as to obtain frequency domain data; according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.

Description

METHOD FOR SAMPLING-RATE RECOGNITION OF PURE VOICE DATA, APPARATUS, AND SYSTEM BACKGROUND OF THE INVENTION Technical Field
[0001] The present invention relates to the technical field of voice processing, and more particularly to a method of sampling-rate recognition for pure voice data, as well as its apparatus and system. Description of Related Art
[0002] Asampling rate defines the number of times for extracting discrete clips per second from a continuous signal and in turn the number of samples, for describing the quality or tone of a voice file, as a standard for evaluating sound cards and voice files in terms of quality.
[0003] There are cases where the object to be processed is pure voice data. The term voice data refers to voice data providing no sampling rate information. Without information of the sampling rate, the assumed sampling rate set for processing may be incorrect and leads to significant deviation during voice processing, rendering poor voice output and degraded user experience to users of voice products. SUMMARY OF THE INVENTION
[0004] In order to address the foregoing issue of the prior art, the present invention provides a method of sampling-rate recognition for pure voice data, as well as its apparatus and system. The method can recognize the sampling rate of pure voice data, so as to reduce occurrence of bad packets during network transmission of voice data packets, and 1 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 improve robustness of voice processing devices by providing additional functions for verification of the sampling rate such as a voice communication engine and common audio processing software and for reminding.
[0005] Embodiments of the present invention provide the following specific technical schemes.
[0006] In a first aspect, the present invention provides a method of sampling-rate recognition for pure voice data, which comprises: [0007] performing Fourier transform on pure voice data so as to obtain frequency domain data; [0008] according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; [0009] acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and [0010] comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
[0011] Preferably, the a priori frequency has a range between 200Hz and 4000Hz.
[0012] Preferably, the step of comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a mostsimilar result as an actual sampling rate comprises: [0013] computing Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
[0014] Preferably, after the step of performing Fourier transform on pure voice data so as to obtain frequency domain data, the method further comprises: [0015] normalizing the frequency domain data. 2 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12
[0016] Preferably, the pure voice (fata are pure voice data containing voice clips; [0017] and the pure voice data containing the voice clips are acquired through: [0018] receiving and analyzing voice data; [0019] if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; and [0020] according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips.
[0021] Preferably, the method further comprises: [0022] decoding the pure voice data according to the actual sampling rate.
[0023] In a second aspect, the present invention provides an apparatus of sampling-rate recognition for pure voice data, which comprises: [0024] a conversing module, for performing Fourier transform on pure voice data so as to obtain frequency domain data; [0025] an acquiring module, for according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; and for acquiring high-frequency cutoff frequency points of the frequency band information; [0026] a computing module, for according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and [0027] a processing module, for comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a mostsimilar result as an actual sampling rate.
[0028] Preferably, the a priori frequency has a range between 200Hz and 4000Hz.
[0029] The processing module is specifically for: [0030] computing Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with 3 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 the smallest Euclidean distance as the actual sampling rate.
[0031] Preferably, the conversing module is further for, after performing Fourier transform on pinevoice data so as to obtain frequency domain data, normalizing the frequency domain data.
[0032] Preferably, the pure voice data are pure voice data containing voice clips; [0033] and the apparatus further comprises: [0034] a rccciving module, for receiving voice data; [0035] an analyzing module, for analyzing the voice data; [0036] the conversing module is further for if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; [0037] the processing module is further for according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips.
[0038] Preferably, the apparatus further comprises:
[0039] A decoding module, for decoding the pure voice data according to the actual sampling rate.
[0040] In a third aspect, the present invention provides a computer system, which comprises: [0041] one or more processors; and [0042] memory associating with the one or more processors, the memory storing a program instruction, the program instruction when read and executed by the one or more processors performing operations of: [0043] performing Fourier transform on pure voice data so as to obtain frequency domain data; [0044] according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; 4 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 [0045] acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and [0046] comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
[0047] Embodiments of the present invention provide the following beneficial effects:
[0048] Given the a priori property that human voice has a bandwidth range between 200Hz and 4000Hz in the frequency domain, the present invention compares different postulated frequencies of pure voice data, and determines the actual sampling rate according to similarity as found in the comparison, thereby automatically pre-determining the size of the sampling rate of the pure voice data, and preventing poor voice processing effects caused by unknowing the sampling rate, so as to reduce occurrence of bad packets during network transmission of voice data packets, and improve robustness of voice processing devices by providing additional functions for verification of the sampling rate such as a voice communication engine and common audio processing software and for reminding. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To better illustrate the technical schemes as disclosed in the embodiments of the present invention, accompanying drawings referred in the description of the embodiments below are introduced briefly. It is apparent that the accompanying drawings as recited in the following description merely provide a part of possible embodiments of the present invention, and people of ordinary skill in the art would be able to obtain more drawings according to those provided herein without paying creative efforts, wherein:
[0050] FIG. 1 is a flowchart of a method of sampling-rate recognition for pure voice data according to Embodiment 1 of the present invention; 5 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12
[0051] FIG. 2 is a structural diagram of an apparatus of sampling-rate recognition for pure voice data according to Embodiment 2 of the present invention; and
[0052] FIG. 3 is a structural diagram of a computer system according to Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] To make the foregoing objectives, features, and advantages of the present invention clearer and more understandable, the following description will be directed to some embodiments as depicted in the accompanying drawings to detail the technical schemes disclosed in these embodiments. It is, however, to be understood that the embodiments referred herein are only a part of all possible embodiments and thus not exhaustive. Based on the embodiments of the present invention, all the other embodiments can be conceived without creative labor by people of ordinary skill in the art, and all these and other embodiments shall be encompassed in the scope of the present invention.
[0054] Embodiment 1
[0055] When network exchange happens between two fixed-line telephones, if a PCM voice packet received at one side with its sampling rate information lost, the received voice cannot be processed accurately and this tends to prevent the voice from being played as it is intended to be.
[0056] This problem may be solved by preliminarily recognizing the sampling rate of the voice, and such a solution can in turn reduce bad packets among voice data packets under network transmission.
[0057] Based on this, as shown in FIG. 1, the present invention provides a method of sampling¬ rate recognition for pure voice data, which can be applied in an audio device for the audio 6 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 device to perform the processing operation as described below.
[0058] In a step Sil, the audio device receives and pre-processes voice data so as to obtain pure voice data.
[0059] In the present embodiment, the pure voice data are pure voice data containing voice clips. The step comprises: [0060] 1. analyzing the voice data, [0061] which specifically refers to acquiring and analyzing information related to the voice data, wherein the related information includes data packets, the sampling rate, etc.; [0062] 2. if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; and [0063] 3. according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain die pure voice data containing the voice clips.
[0064] The step of receiving and pre-processing voice data so as to obtain pure voice data may be alternatively realized through the following steps: [0065] 1. analyzing the voice data; and [0066] 2. if the voice data do not contain the sampling rate information, filtering the voice data so as to obtain pure voice data containing voice clips.
[0067] The filtering process is to separate noise and mute from effective voice data, thereby enhancing the voice data.
[0068] The step S12 is about performing Fourier transform on pure voice data so as to obtain 7 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 frequency domain data.
[0069] By means of Fourier transform, the pure voice data in the time domain can be converted into voice data in the frequency domain.
[0070] The step S13 involves normalizing the frequency domain data.
[0071] The step S14 is implemented by, according to received a priori thresholddata, processing the frequency domain data so as to obtain frequency band information.
[0072] Therein, the a priori threshold data are obtained by processing a priori information.
[0073] Given that the human voice has a bandwidth in the frequency domain ranging between 200Hz and 4000Hz, this information is introduced as the a priori information.
[0074] The step S15 is about acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing a postulated frequency corresponding to the high-frequency cutoff frequency point.
[0075] The step SI6 involves comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a mostsimilar result as an actual sampling rate.
[0076] Specifically, the step may comprise: [0077] computing Euclidean distances between different postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
[0078] The smallest Euclidean distance means that the very postulated frequency has the greatest 8 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 similarity to the a priori frequency, and the sampling rate corresponding to that postulated frequency is one closest to the actual value. In this way, the actual sampling rate of the pure voice data can be determined.
[0079] Therein, for computing the Euclidean distances between different postulated frequencies and the a priori frequency, several frequency points may be selected from the a priori frequency for computing, such as 200Hz and 4000 Hz.
[0080] In the step SI7, the pure voice data are decoded according to actual sampling rate and then played.
[0081] Thereby, bad packets having their sampling rate information lost in the telephone exchange network can be decoded and displayed according to the recognized sampling rate as normal, so as to ensure good quality and user experience of voice calls.
[0082] In order to prove the effects of the method described above, experiments were conducted against audio files having different sampling rates, and the results of recognition are shown below:
[0083] Table 1. Test Accuracy of Audio Files Having Different Sampling Rates
[0084] Experimental Condition Determination Accuracy Number of Determinations Pcm file 8k sampling 85% 100 Pcm file 16k sampling 81% 100 Pcm file 32k sampling 88% 100
[0085] It is learned from the results of the experiments that the disclosed scheme had excellent 9 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 recognition accuracy, higher than 80%.
[0086] Additionally, for communication engines such as those used in IPcall services, to prevent sampling rates from being wrongly set by incorrect manual operation, the disclosed method may be used to make pre-determination and enable timely warning so as to minimize risks of failure and loss. For commonly used audio processing software products, the present invention helps warn users of wrongly set sampling rates, thereby saving users from waste of their working and private time and futile operation.
[0087] Embodiment 2
[0088] As shown in FIG. 2, the present invention furflier provides an apparatus of sampling-rate recognition for pure voice data, comprising: [0089] a conversing module 21, for performing Fourier transform on pure voice data so as to obtain frequency domain data; [0090] an acquiring module 22, for according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; and for acquiring high-frequency cutoff frequency points of the frequency band information; [0091] a computing module 23, for according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and [0092] a processing module 24, for comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
[0093] Preferably, the a priori frequency has a range between 200Hz and 4000Hz.
[0094] Preferably, the processing module 24 is specifically for computing Euclidean distances between different postulated frequencies and the a priori frequency, and taking the 10 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
[0095] Preferably, the conversing module 21 is further for, after the step of performing Fourier transform on pure voice data so as to obtain frequency domain data, normalizing the frequency domain data.
[0096] Preferably, the pure voice data are pure voice data containing voice clips; [0097] and the apparatus further comprises: [0098] a receiving module 25, for receiving voice data; and [0099] an analyzing module 26, for analyzing the voice data.
[0100] The conversing module 21 is further for, if the voice data do not contain sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data.
[0101] The processing module 24 is further for, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips.
[0102] Preferably, the apparatus further comprises: [0103] a decoding module 27, for decoding the pure voice data according to actual sampling rate.
[0104] Embodiment 3
[0105] The present invention also provides a computer system, which comprises: [0106] one or more processors; and [0107] memory associating with the one or more processors, the memory storing a program instruction, the program instruction when read and executed by the one or more 11 Date ReQue/Date Received 2022-09-12CA 03175103 2022-09-12 processors performing operations of: [0108] performing Fourier transform on pure voice data so as to obtain frequency domain data; [0109] according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; [0110] acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and [0111] comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
[0112] FIG. 3 exemplarily shows a structure of a computer system. It comprises a processor 32, a video display adapter 34, a disc driver 36, an input/output interface 38, a network interface 310, and memory 312. The processor 32, the video display adapter 34, the disc driver 36, the input/output interface 38, the network interface 310, and the memory 312 may be communicated with each other through a communication bus 314.
[0113] Therein, the processor 32 may be realized using a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits to execute the relevant program, thereby implementing the technical schemes provided by the present invention.
[0114] The memory 312 may be a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 312 may be used to control the operating system 316 running in the computer system 30 to control the lower grade basic output/input system (BIOS) 318 in the computer system. Furthermore, it may store a webpage browser 320, a data storage and management system 322 or the like. In short, if software or firmware is used to realize the technical schemes of the present invention, the related program codes are stored in the memory 312 for the 12 Date ReQue/Date Received 2022-09-12processor 32 to call.
[0115] The input/output interface 38 is used to connect an input/output module to enable input and output of information. The input/output module may be set in the device as a component (not shown) or may be installed outside the device to provide required functions. Therein, the input device may include a keyboard, a mouse, a touch screen, a microphone, and various sensors, while the output device may include a display device, a loud speaker, a vibrator, an indicator lamp, or the like.
[0116] The network interface 310 is used to connect a communication module (not shown) to enable communication between the disclosed device and other devices. Therein, the communication module may enable communication in a wired way (through such as USB and cables) or in a wireless way (through such as mobile networks, Wi-Fi, Bluetooth™, etc.).
[0117] The communication bus 314 comprises a path, along which information can be transmitted among various components (such as the processor 32, the video display adapter 34, the disc driver 36, the input/output interface 38, the network interface 310, and the memory 312 of the device.
[0118] Additionally, the computer system may further obtain the information of exact receiving conditions from a virtual resource object receiving condition information database and use the information for determination of conditions, among others.
[0119] It is to be noted that while the device described previously only shows die processor 32, the video display adapter 34, the disc driver 36, the input/output interface 38, the network interface 310, the memory 312, and the communication bus 314, in actual implementations, the device may further comprise other components required for normal operation. 13 Date ReQue/Date Received 2024-04-01CA 03175103 2022-09-12
[0120] Through the foregoing description of the embodiments, people skilled in the art would understand that the present invention may be implemented using software along with a necessary general-purpose hardware platform. Based on such understanding, the technical schemes of the present invention may essentially or with its parts having contribution to the prior art embodied in the form of a software product. The computer software product may be stored in a storage medium, such as a ROM/RAM, a magnetic disk, or an optical disc, and contain several instructions that make a computer device (such as a personal computer,a cloud server end, or a network device) perform the method as described in any embodiment or any part of the embodiments of the present invention.
[0121] While some preferred embodiments of the present invention have been described, people skilled in the art, with the knowledge of the basic creative concepts, might devise additional modifications and variations based on these embodiments. Hence, the appended claims are intended to encompass these preferred embodiments and all modifications and variations falling within their scopes. Moreover, the computer system and the apparatus of sampling-rate recognition for pure voice data stem from the same concept as the method for sampling-rate recognition of pure voice data as described in the previous embodiment, and since details of the specific implementation can be found in the embodiment related to the method, no repeated description is made herein.
[0122] It is obvious that people skilled in the art might make various modificationsandvariations to the present invention without departing from the spirit and scope of the present invention. Provided that these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is intended to encompass all modifications and variations. 14 Date ReQue/Date Received 2022-09-12

Claims (19)

  1. Claims: 1. An apparatus for recognizing a pure sampling-rate of pure voice data the pure sampling-rate being the rate which the pure voice data is sampled at, the apparatus comprising: a conversing module for performing Fourier transform on pure voice data so as to obtain frequency domain data; an acquiring module for, according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information, and acquiring highfrequency cutoff frequency points of the frequency band information; a computing module for according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; a processing module for comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
  2. 2. The apparatus of claim 1, wherein the a priori frequency has a range between 200Hz and 4000Hz.
  3. 3. The apparatus of claim 1, wherein the processing module is further configured for: computing Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
  4. 4. The apparatus of any one of claims 1 to 3, wherein the apparatus normalizes the frequency domain data.
  5. 5. The apparatus of any one of claims 1 to 3, wherein the pure voice data are pure voice data containing voice clips and the pure voice data containing the voice clips are acquired through: a receiving module for receiving voice data; and 15 Date ReQue/Date Received 2024-04-01an analyzing module for analyzing the voice data; wherein if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips.
  6. 6. The apparatus of any one of claims 1 to 3, further comprising: a decoding module for decoding the pure voice data according to the actual sampling rate.
  7. 7. The apparatus of any one of claims 1 to 6 wherein the apparatus is used to make pre¬ determination and enable timely warning to minimize failure and loss.
  8. 8. An system for recognizing a pure sampling-rate of pure voice data the pure sampling-rate being the rate which the pure voice data is sampled at, the system comprising: a conversing module for performing Fourier transform on pure voice data so as to obtain frequency domain data; an acquiring module for, according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information, and acquiring highfrequency cutoff frequency points of the frequency band information; a computing module for according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; a processing module for comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
  9. 9. Thesystem of claim 8, wherein the a priori frequency has a range between 200Hz and 4000Hz.
  10. 10. The system of claim 8, wherein the processing module is further configured for: 16 Date ReQue/Date Received 2024-04-01computing Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
  11. 11. The system of any one of claims 8 to 10, wherein the system normalizes the frequency domain data.
  12. 12. The system of any one of claims 8 to 10, wherein the pure voice data are pure voice data containing voice clips and the pure voice data containing the voice clips are acquired through: a receiving module for receiving voice data; and an analyzing module for analyzing the voice data; wherein if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips.
  13. 13. The system of any one of claims 8 to 10, further comprising: a decoding module for decoding the pure voice data according to the actual sampling rate.
  14. 14. The system of any one of claims 8 to 13 wherein the system is used to make pre-determination and enable timely warning to minimize failure and loss.
  15. 15. A method of recognizing a pure sampling-rate of pure voicedata, the pure sampling-rate being the rate the pure voice data is sampled at, the method comprising: performing Fourier transform on pure voice data so as to obtain frequency domain data; according to received a priori threshold data, processing the frequency domain data so as to obtain frequency band information; 17 Date ReQue/Date Received 2024-04-01acquiring high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, computing postulated frequencies corresponding to the high-frequency cutoff frequency points; and comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate.
  16. 16. The method of claim 15, wherein the a priori frequency has a range between 200Hz and 4000Hz.
  17. 17. The method of claim 15, wherein the step of comparing the postulated frequencies with an a priori frequency, and taking the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate comprises: computing Euclidean distances between die postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate.
  18. 18. The method of any one of claims 15 to 17, wherein after the step of performing Fourier transform on pure voice data so as to obtain frequency domain data, the method further comprises: normalizing the frequency domain data. 19. The method of any one of claims 15 to 17, wherein the pure voice data are pure voice data containing voice clips; and the pure voice data containing the voice clips are acquired through: receiving voice data and analyzing the voice data; if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; 18 Date ReQue/Date Received 2024-04-01according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips. 20. The method of any one of claims 15 to 17, further comprising: decoding the pure voice data according to the actual sampling rate. 21. The method of any one of claims 15 to 20 wherein the method is used to make pre¬ determination and enable timely warning to minimize failure and loss. 22. A computer equipment for recognizing a pure sampling-rate of pure voice data the pure sampling-rate being the rate which the pure voice data is sampled at comprising: a computer readable physical memory; a processor communicatively coupled to the memory, a computer program stored on the memory and operable on the processor, wherein the processor executes the computer program configured to: perform Fourier transform on pure voice data so as to obtain frequency domain data; according to received a priori threshold data, process the frequency domain data so as to obtain frequency band information; acquire high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, compute postulated frequencies corresponding to the high-frequency cutoff frequency points; and compare the postulated frequencies with an a priori frequency, and take the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate. 23. The equipment of claim 22, wherein the a priori frequency has a range between 200Hz and 4000Hz.
  19. 19 Date ReQue/Date Received 2024-04-0124. The equipment of claim 22, wherein the program is further configured to: compute Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate. 25. The equipment of any one of claims 22 to 24, wherein program is further configured to normalize the frequency domain data. 26. The equipment of any one of claims 22 to 24, wherein the pure voice data are pure voice data containing voice clips; and program is further configured to acquire the pure voice data containing the voice clips by: receiving voice data and analyzing the voice data; if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips. 27. The equipment of any one of claims 22 to 24, wherein the program is further configured to decode the pure voice data according to the actual sampling rate. 28. The equipment of any one of claims 22 to 27 wherein the equipment is used to make pre¬ determination and enable timely warning to minimize failure and loss. 29. A computer readable physical memory having stored thereon a computer program for recognizing a pure sampling-rate of pure voice data the pure sampling-rate being the rate which the pure voice data is sampled at the computer program when executed by a computer is configured to: perform Fourier transform on pure voice data so as to obtain frequency domain data; 20 Date ReQue/Date Received 2024-04-01according to received a priori threshold data, process the frequency domain data so as to obtain frequency band information; acquire high-frequency cutoff frequency points of the frequency band information, and according to predetermined different sampling rates, compute postulated frequencies corresponding to the high-frequency cutoff frequency points; and compare the postulated frequencies with an a priori frequency, and take the sampling rate corresponding to the postulated frequency with a most-similar result as an actual sampling rate. 30. The memory of claim 29, wherein the a priori frequency has a range between 200Hz and 4000Hz. 31. The memory of claim 29, wherein the program is further configured to: compute Euclidean distances between the postulated frequencies and the a priori frequency, and taking the sampling rate corresponding to the postulated frequency with the smallest Euclidean distance as the actual sampling rate. 32. The memory of any one of claims 29 to 31, wherein program is further configured to normalize the frequency domain data. 33. The memory of any one of claims 29 to 31, wherein the pure voice data are pure voice data containing voice clips; and program is further configured to acquire the pure voice data containing the voice clips by: receiving voice data and analyzing the voice data; if the voice data do not contain the sampling rate information, performing Fourier transform on the voice data so as to obtain energy of the voice data; according to a predetermined energy threshold, acquiring the voice data corresponding to energy greater than the energy threshold so as to obtain the pure voice data containing the voice clips. 21 Date ReQue/Date Received 2024-04-0134. The memory of any one of claims 29 to 31, wherein the program is further configured to decode the pure voice data according to the actual sampling rate. 35. The memory of any one of claims 29 to 34 wherein the memory is used to make pre¬ determination and enable timely warning to minimize failure and loss. Date ReQue/Date Received 2024-04-01
CA3175103A 2020-03-10 2020-06-19 Method for sampling-rate recognition of pure voice data, apparatus, and system Active CA3175103C (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
CN202010160577.4 2020-03-10
CN202010160577.4A CN111354365B (en) 2020-03-10 2020-03-10 A pure voice data sampling rate recognition method, device and system
PCT/CN2020/097008 WO2021179470A1 (en) 2020-03-10 2020-06-19 Method, device and system for recognizing sampling rate of pure voice data

Publications (2)

Publication Number Publication Date
CA3175103A1 CA3175103A1 (en) 2021-09-16
CA3175103C true CA3175103C (en) 2025-03-11

Family

ID=71196071

Family Applications (1)

Application Number Title Priority Date Filing Date
CA3175103A Active CA3175103C (en) 2020-03-10 2020-06-19 Method for sampling-rate recognition of pure voice data, apparatus, and system

Country Status (3)

Country Link
CN (1) CN111354365B (en)
CA (1) CA3175103C (en)
WO (1) WO2021179470A1 (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113447713B (en) * 2021-06-25 2023-03-07 南京丰道电力科技有限公司 Fourier-based fast high-precision power system frequency measurement method and device
CN115579019A (en) * 2022-09-06 2023-01-06 平安科技(深圳)有限公司 Optimization training method, device, computer equipment and medium for speech classification model

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7046857B2 (en) * 1997-07-31 2006-05-16 The Regents Of The University Of California Apparatus and methods for image and signal processing
CN101320560A (en) * 2008-07-01 2008-12-10 上海大学 A Method of Improving Recognition Rate Using Sampling Rate Conversion in Speech Recognition System
CN101582264A (en) * 2009-06-12 2009-11-18 瑞声声学科技(深圳)有限公司 Method and voice collecting system for speech enhancement
JP2012002858A (en) * 2010-06-14 2012-01-05 Pioneer Electronic Corp Time scaling method, pitch shift method, audio data processing apparatus and program
CN102332266B (en) * 2010-07-13 2013-04-24 炬力集成电路设计有限公司 Audio data encoding method and device
EP2981956B1 (en) * 2013-04-05 2022-11-30 Dolby International AB Audio processing system
CN103745726B (en) * 2013-11-07 2016-08-17 中国电子科技集团公司第四十一研究所 A kind of adaptive variable sampling rate audio sample method
CN105513590A (en) * 2015-11-23 2016-04-20 百度在线网络技术(北京)有限公司 Voice recognition method and device
US10249307B2 (en) * 2016-06-27 2019-04-02 Qualcomm Incorporated Audio decoding using intermediate sampling rate
CN107833581B (en) * 2017-10-20 2021-04-13 广州酷狗计算机科技有限公司 Method, device and readable storage medium for extracting fundamental tone frequency of sound

Also Published As

Publication number Publication date
CA3175103A1 (en) 2021-09-16
CN111354365B (en) 2023-10-31
WO2021179470A1 (en) 2021-09-16
CN111354365A (en) 2020-06-30

Similar Documents

Publication Publication Date Title
US11210461B2 (en) Real-time privacy filter
US10339956B2 (en) Method and apparatus for detecting audio signal according to frequency domain energy
US8826210B2 (en) Visualization interface of continuous waveform multi-speaker identification
EP1901285A2 (en) Voice Authentication Apparatus
CN111916109B (en) Audio classification method and device based on characteristics and computing equipment
WO2018014673A1 (en) Method and device for howling detection
CN111739542A (en) Method, device and equipment for detecting characteristic sound
CN111031329A (en) Method, apparatus and computer storage medium for managing audio data
US20150325252A1 (en) Method and device for eliminating noise, and mobile terminal
CA3175103A1 (en) Method for sampling-rate recognition of pure voice data, apparatus, and system
CN105791602B (en) Sound quality testing method and system
CN110444194B (en) Voice detection method and device
CN112687293B (en) Intelligent agent training method and system based on machine learning and data mining
CN111627453B (en) Public security voice information management method, device, equipment and computer storage medium
CN110189763B (en) Sound wave configuration method and device and terminal equipment
CN107154996B (en) Incoming call interception method and device, storage medium and terminal
CN111046366A (en) User identity identification method and device and electronic equipment
CN113658581B (en) Acoustic model training, speech processing methods, devices, equipment and storage media
CN117636878A (en) Voice information processing method and device and nonvolatile storage medium
CN116434774A (en) Speech recognition method and related device
WO2023173966A1 (en) Speech identification method, terminal device, and computer readable storage medium
CN115879841A (en) Data processing method and device, electronic equipment and storage medium
CN113316074B (en) Howling detection method and device and electronic equipment
CN115394304A (en) Voiceprint determination method, apparatus, system, device and storage medium
CN113990304A (en) Voice activity detection method and device, computer readable storage medium and equipment

Legal Events

Date Code Title Description
EEER Examination request

Effective date: 20220912

MFA Maintenance fee for application paid

Free format text: FEE DESCRIPTION TEXT: MF (APPLICATION, 5TH ANNIV.) - STANDARD

Year of fee payment: 5

U00 Fee paid

Free format text: ST27 STATUS EVENT CODE: A-2-2-U10-U00-U101 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: MAINTENANCE REQUEST RECEIVED

Effective date: 20241220

U11 Full renewal or maintenance fee paid

Free format text: ST27 STATUS EVENT CODE: A-2-2-U10-U11-U102 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: MAINTENANCE FEE PAYMENT DETERMINED COMPLIANT

Effective date: 20241221

Free format text: ST27 STATUS EVENT CODE: A-2-2-U10-U11-U102 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: MAINTENANCE FEE PAYMENT PAID IN FULL

Effective date: 20241221

D22 Grant of ip right intended

Free format text: ST27 STATUS EVENT CODE: A-2-4-D10-D22-D143 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: PRE-GRANT

Effective date: 20250218

Q17 Modified document published

Free format text: ST27 STATUS EVENT CODE: A-4-4-Q10-Q17-Q103 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: DOCUMENT PUBLISHED

Effective date: 20250310

F11 Ip right granted following substantive examination

Free format text: ST27 STATUS EVENT CODE: A-4-4-F10-F11-X000 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: GRANT BY ISSUANCE

Effective date: 20250311

W00 Other event occurred

Free format text: ST27 STATUS EVENT CODE: A-4-4-W10-W00-W111 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: CORRESPONDENT DETERMINED COMPLIANT

Effective date: 20250526

MPN Maintenance fee for patent paid

Free format text: FEE DESCRIPTION TEXT: MF (PATENT, 6TH ANNIV.) - STANDARD

Year of fee payment: 6

U00 Fee paid

Free format text: ST27 STATUS EVENT CODE: A-4-4-U10-U00-U101 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: MAINTENANCE REQUEST RECEIVED

Effective date: 20251229

U11 Full renewal or maintenance fee paid

Free format text: ST27 STATUS EVENT CODE: A-4-4-U10-U11-U102 (AS PROVIDED BY THE NATIONAL OFFICE); EVENT TEXT: MAINTENANCE FEE PAYMENT PAID IN FULL

Effective date: 20251229