WO2019210557A1 - 语音质检方法、装置、计算机设备及存储介质 - Google Patents

语音质检方法、装置、计算机设备及存储介质 Download PDF

Info

Publication number
WO2019210557A1
WO2019210557A1 PCT/CN2018/092322 CN2018092322W WO2019210557A1 WO 2019210557 A1 WO2019210557 A1 WO 2019210557A1 CN 2018092322 W CN2018092322 W CN 2018092322W WO 2019210557 A1 WO2019210557 A1 WO 2019210557A1
Authority
WO
WIPO (PCT)
Prior art keywords
quality inspection
call
text
word data
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/092322
Other languages
English (en)
French (fr)
Inventor
张政
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2019210557A1 publication Critical patent/WO2019210557A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04MTELEPHONIC COMMUNICATION
    • H04M3/00Automatic or semi-automatic exchanges
    • H04M3/42Systems providing special services or facilities to subscribers
    • H04M3/50Centralised arrangements for answering calls; Centralised arrangements for recording messages for absent or busy subscribers ; Centralised arrangements for recording messages
    • H04M3/51Centralised call answering arrangements requiring operator intervention, e.g. call or contact centers for telemarketing
    • H04M3/5175Call or contact centers supervision arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/22Matching criteria, e.g. proximity measures
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04MTELEPHONIC COMMUNICATION
    • H04M3/00Automatic or semi-automatic exchanges
    • H04M3/22Arrangements for supervision, monitoring or testing
    • H04M3/2227Quality of service monitoring
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems

Definitions

  • the present application relates to the field of data processing, and in particular, to a voice quality inspection method, apparatus, computer equipment, and storage medium.
  • Call Center is a complete integrated information service system based on CTI (Computer Telephony Integration) technology, which fully utilizes multiple functions of communication network and computer network to integrate with enterprises.
  • CTI Computer Telephony Integration
  • the call center can provide users with multiple services efficiently and at high speed.
  • enterprises can obtain customer's business consultation, problem feedback, product installation or maintenance acceptance, complaint acceptance, accepting customer's opinions and suggestions through call center call-in mode, and can also conduct market research, telephone shopping, and membership through call-out mode. Feedback, telemarketing, old customer return visits, care, sales and after-sales tracking services.
  • the embodiments of the present application provide a voice quality inspection method, apparatus, computer equipment, and storage medium to solve the problem of low efficiency and accuracy of voice quality inspection.
  • a voice quality inspection method includes the following steps:
  • the quality inspection report that the quality inspection is unqualified is output.
  • a voice quality inspection device comprising:
  • a call recording file obtaining module configured to obtain a call recording file, where the call recording file includes a file identifier
  • a quality check word data obtaining module configured to obtain corresponding quality test word data based on the file identifier
  • a call text conversion module configured to convert the call recording file into a call text by using a voice recognition algorithm based on the quality check word data
  • a matching degree obtaining module configured to acquire a matching degree of the call text based on a preset quality inspection template
  • the quality inspection report output module is configured to output a quality inspection report that fails the quality inspection if the matching degree of the call text does not exceed a preset threshold.
  • a computer device comprising a memory, a processor, and computer readable instructions stored in the memory and operative on the processor, the processor executing the computer readable instructions to:
  • the quality inspection report that the quality inspection is unqualified is output.
  • One or more non-transitory readable storage mediums storing computer readable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps:
  • the quality inspection report that the quality inspection is unqualified is output.
  • FIG. 1 is a flowchart of an implementation of a voice quality inspection method in Embodiment 1 of the present application
  • FIG. 2 is a flowchart of another implementation of a voice quality check method in Embodiment 1 of the present application.
  • step S30 is a flowchart of an implementation of step S30 in Embodiment 1 of the present application.
  • step S32 is a flowchart of an implementation of step S32 in Embodiment 1 of the present application.
  • FIG. 5 is a flowchart of another implementation of a voice quality check method according to Embodiment 1 of the present application.
  • FIG. 6 is a flowchart of an implementation of step S72 in Embodiment 1 of the present application.
  • FIG. 7 is a schematic diagram of a voice quality inspection apparatus provided in Embodiment 2 of the present application.
  • FIG. 8 is a schematic diagram of a computer device provided in Embodiment 4 of the present application.
  • Fig. 1 is a flow chart showing the voice quality inspection method in this embodiment.
  • the voice quality inspection method is applied to a terminal or a system to solve the problem of low efficiency and accuracy of voice quality inspection.
  • it can be applied to a communication terminal or system including a customer service center, a call center, and the like.
  • the voice quality inspection method includes the following steps:
  • S10 Acquire a call recording file, and the call recording file includes a file identifier.
  • the call recording file refers to the recording file when the agent communicates with the customer.
  • the call recording files are saved in a call recording database, and then the corresponding call recording file is obtained through the call recording database.
  • the file identifier refers to the identifier set according to the type of the call recording file. You can set different identifiers for the call recording file according to the service type or call type. For example, if the file identifier is set according to the service type, different file identifiers can be set according to different services such as "finance management", "insurance” or "savings".
  • the call recording file is obtained from the call recording database, and the corresponding call recording file may be obtained according to at least one of the agent ID, the time interval, or the file identifier.
  • the agent ID refers to an account in the system or the terminal, and is used to identify different agents.
  • the corresponding call recording file is obtained from the call recording database according to the time interval (X-X X day) and the agent ID (Zhang San). If the business handled by Zhang San is a wealth management business, the file name of the corresponding call recording file is “Financial Management”.
  • the quality test word data refers to the Chinese word data used for quality inspection of the call recording file, which can be established internally or by connecting to the big data platform.
  • the quality test word data includes different types of data, and optionally, different quality test word data may be set according to the service type or the call type. For example, different quality word data is set according to different businesses such as "finance management", "insurance” or "savings".
  • the setting of the data type of the quality inspection word corresponds to the setting of the file identification. For example, if the file identifier is set according to the service type, the quality check word data is also set according to the corresponding service type.
  • the corresponding quality word data is obtained according to the specific file identifier of the call recording file.
  • the quality test word data corresponding to “financial wealth” is obtained.
  • the obtained data can be made more accurate, and the accuracy of the subsequent call recording file conversion is improved.
  • the speech recognition algorithm refers to an algorithm that recognizes speech as text.
  • the call text refers to the text of the call content reflected in text form.
  • a Hidden Markov Model (HMM) algorithm may be used to implement the pair.
  • DTW Dynamic Time Warping
  • DNN Deep Neural Network
  • the call recording file is subjected to noise filtering/noise reduction, sentence segmentation, and sentence conversion processing, thereby converting the call recording file into corresponding word elements, such as pinyin. Matching the converted word elements with the reference words in the corresponding QC data, and combining the matched words to obtain the call text.
  • the speech segmentation may be performed by Voice Activity Detection (VAD), and the two sentences are converted into “n ⁇ h ⁇ o” and “w ⁇ sh ⁇ zh ⁇ ng s ⁇ n”; and then matched with the quality test word data, according to Chinese word segmentation algorithm rules, such as forward maximum matching method, inverse maximum matching method, minimum matching method or maximum matching method, etc., match the two sentences "Hello” and "I am Zhang San", and then carry out two sentences. Combine into a call text.
  • VAD Voice Activity Detection
  • the voice quality inspection time can be saved, and the efficiency of the voice quality check can be improved.
  • the quality inspection template is a text template for checking the quality of the agent service set for a specific service. Matching is the degree to which the call text matches the words in the QC template.
  • the words in the call text are matched with the words in the preset quality check template, and the proportion of the matching words in the total number of words in the preset quality check template is calculated as the matching degree of the call text.
  • the default quality inspection template is “Hello, I am very happy to serve you”
  • the words in the call text that match the quality inspection template are “Hello”
  • the total number of words in the quality inspection template is 9, in the call text.
  • the ratio is calculated by using a single word as a minimum unit.
  • the degree of specification of the agent in the business can be verified, which helps to present the voice quality inspection result more intuitively, and also improves the efficiency of the voice quality inspection.
  • the preset threshold is a lower limit for setting the matching degree, and can be set according to actual needs.
  • the quality inspection report is a concrete manifestation of the quality of the voice quality inspection. Specifically, the quality inspection report may include the quality inspection result, the matching degree of the call text, and the matching specific words, etc., wherein the quality inspection result includes the quality inspection and the quality inspection report. The quality inspection failed.
  • the matching degree of the call text is compared with a preset threshold, and if the matching degree is less than or equal to the preset threshold, the quality inspection report that the quality inspection is unqualified is output.
  • the preset threshold is 80%, and if the matching degree of the call text is only 70%, the quality inspection report that the quality inspection is unqualified is output.
  • the call recording file is obtained, the call recording file includes a file identifier, and the corresponding quality check word data is obtained based on the file identifier, and the call recording file is converted into the call text by using a voice recognition algorithm based on the quality check word data.
  • the matching degree of the call text is obtained based on the preset quality inspection template. If the matching degree of the call text does not exceed the preset threshold, the quality inspection report that the quality inspection is unqualified is output.
  • the call recording file is converted into the call text by the quality check word data, the conversion precision is improved, and then the test is performed according to the preset quality check template, thereby finding the problematic call recording file, saving the voice quality inspection time and improving the voice. The efficiency of quality inspection.
  • the voice quality check method further includes a process of updating the quality test word data before the step S20, as shown in FIG. 2, specifically including the following steps:
  • S61 Acquire quality inspection word update data, and the quality inspection word update data includes quality inspection word data identifier.
  • the quality check word update data refers to the embodiment of the data change of the corresponding quality test word data.
  • the embodiment is updated by means of connecting to a big data platform.
  • the quality inspection word data identifier is used to identify what type of quality inspection word data corresponds to the quality inspection word update data.
  • the corresponding quality inspection word data is updated according to the quality inspection word data identifier. It can be understood that the words in the different types of quality test word data will change, and the frequency of use of the corresponding words will be constantly updated. In this embodiment, updating the quality test word data before step S20 can ensure the real-time and validity of the quality test word data, and also improve the accuracy of converting the subsequent call recording file into the call text.
  • the big data platform can be connected.
  • the corresponding big data platform has data update
  • the corresponding update data can be synchronized to the corresponding quality inspection word data in real time, and the quality inspection word data can be real-time.
  • the update further ensures the accuracy of the quality test word data, and improves the conversion accuracy of the subsequent call recording file into the call text.
  • the quality test word data is connected to the insurance-related big data platform.
  • the relevant big data platform has data updates, for example, adding a word "safety protection”, or the word "storing” is increased by 10 times.
  • the corresponding quality inspection word data is updated according to the corresponding quality inspection word data identifier “insurance”.
  • the quality check word data may be updated at intervals, for example, every 10 minutes, and may be set according to actual needs, which is not limited in the embodiment of the present application.
  • the words of the quality test word data can be updated in time, enriching the word data, and the conversion accuracy when converting the call record into the call text can be effectively improved.
  • the voice recognition algorithm is used to convert the call recording file into a call text, as shown in FIG. 3, which specifically includes the following steps:
  • the target pinyin element refers to converting the call recording file into an element composed of pinyin, including syllables and tones.
  • a syllable is the smallest unit of phonetic structure composed of a combination of phonemes. It is the basic unit of the phonetic system in which the auditory can distinguish clearly.
  • Each syllable consists of two parts: the initial and the final. For example, the syllable of "Zhang" is "zhang", and the syllable of "Zhang San” contains two syllables for "zhang san”. Tone is a property used in Chinese Pinyin to distinguish the height and elevation of a sound, usually consisting of four sounds.
  • the call recording file is first subjected to a VAD (Voice Boundary Detection) operation, and the silence period is recognized from the sound stream of the call recording file and framing is performed accordingly. After framing, the sound stream is divided into small segments. Then, the MFCC (Mel-Frequency Cepstral Coefficients) acoustic feature extraction is performed, wherein the MFCC is a cepstrum parameter extracted in the Mel scale frequency domain, which may include two parts: the Mel frequency conversion and the cepstrum analysis.
  • the Mel scale is a nonlinear frequency scale based on the sensory judgment of the human ear on the pitch changes of the pitch, and the relationship with the sound frequency is as follows:
  • m is the Mel scale and f is the sound frequency in Hz.
  • Cepstrum analysis refers to the process of performing Fourier transform on the time domain signal, then taking the logarithm, and then performing the inverse Fourier transform. After the step of extracting the acoustic features is completed, the extracted acoustic features are matched by the HMM algorithm, the DTW algorithm or the DNN algorithm based on the DNN algorithm to obtain the target pinyin elements.
  • S32 Match the target pinyin elements based on the quality test word data, and convert the corresponding target pinyin elements into the target text data.
  • the target pinyin element is matched with the pinyin of the word in the corresponding quality word data, and the matched words are combined to form the target text data.
  • the call recording file identified as insurance is "Hello, I am happy to serve you", and convert it into the target pinyin elements of "n ⁇ n h ⁇ o" and "h ⁇ n g ⁇ o x ⁇ ng wèi n ⁇ n f ⁇ wù", these target pinyin elements Matching the pinyin of the words in the quality test word data corresponding to the insurance, according to the Chinese word segmentation algorithm rules, such as the forward maximum matching method, the inverse maximum matching method, the minimum matching method or the maximum matching method, etc., matching "Hello" Target text data for "and very happy to serve you”.
  • the obtained target text data is combined by adding appropriate punctuation marks to form a call text output.
  • the input of punctuation may be based on statistics of the frequency of use of punctuation after the last word in a sentence according to the big data platform, and input the punctuation symbol with the highest frequency of use. For example, the last word “Hello, I am happy to serve you” is “Service”. If the most frequently used punctuation is based on the statistics of the Big Data Platform, the “Service” is connected to the next sentence by a period.
  • the call recording file can be effectively converted into a call by converting the call recording file into a target pinyin element including a syllable and a tone, and then matching the pinyin of the word in the corresponding quality word data. Text that improves the accuracy of the conversion.
  • the quality word data includes common word data and business word data.
  • the common word data refers to the words used in the ordinary scene
  • the business word data refers to the words used corresponding to the specific business. It can be understood that the business word data includes fewer words than ordinary word data.
  • the target pinyin elements are matched based on the quality test word data, and the corresponding target pinyin elements are converted into the target text data, as shown in FIG. 4, specifically including the following steps:
  • S321 Match the target pinyin elements based on the business word data, and convert the sub-pinyin elements into the business text data according to the corresponding reference pinyin elements in the business word data.
  • S322 Matching the remaining sub-pinyin elements in the target pinyin element based on the common word data, and converting the remaining sub-pinyin elements into normal text data according to the corresponding reference pinyin elements in the common word data.
  • the reference pinyin element refers to a pinyin element corresponding to a word in the quality test word data.
  • the reference pinyin element corresponding to the word "insurance" in the quality word data is "b ⁇ o xi ⁇ n".
  • the sub-pinyin element refers to the corresponding pinyin element after each sentence has been transformed after the call recording file is segmented by the statement.
  • the target pinyin element is first matched with the business word data corresponding to the file identifier of the call recording file, and the sub-pinyin elements that can be matched are converted into corresponding words to obtain business text data; and then, the remaining unmatched The sub-pinyin elements are matched with the common word data, and the remaining sub-pinyin elements are converted into corresponding words to obtain ordinary text data. Finally, the two parts of text data (business text data and normal text data) are combined to obtain the target text data.
  • the call recording file whose file is marked as “insurance” is “our company has a product security guarantee”
  • the target pinyin element of the conversion is “w ⁇ men g ⁇ ng s ⁇ y ⁇ u y ⁇ gè ch ⁇ n p ⁇ n ⁇ n x ⁇ n b ⁇ o”.
  • the word matching when the word matching is performed, if there are a plurality of words corresponding to the reference pinyin elements matching the business word data or the common word data, the words with the highest frequency of use are output. For example, “g ⁇ ng s ⁇ ” can be matched with “public and private” and “company”. If “company” is used more frequently than “public and private”, the matching word is “company”.
  • the first matching with the business word data is matched with the common word data, and finally merged into the target text data.
  • the voice quality inspection method before the step S40, the voice quality inspection method further includes a quality check on the sensitive words, as shown in FIG. 5, specifically including the following steps:
  • sensitive word data refers to words that are prohibited from appearing during the business communication with the client, such as some impolite words.
  • the sensitive term data can be obtained and updated through a corresponding sensitive term database.
  • the words in the call text are matched with the sensitive word data, and if the words in the call text match any of the words in the sensitive word data, the quality inspection report with the unqualified quality check is output.
  • the matching of the sensitive word data is prioritized, and the quality inspection of the sensitive words is performed, and the setting of the quality inspection mode is ensured.
  • the reasonableness of the voice quality check, and the corresponding call text after the quality test report of the sensitive word quality test output is not required to obtain the matching degree of the call text, and the data processing efficiency of the voice quality check method is also improved. .
  • the quality inspection report that fails the quality check is output, as shown in FIG. 6 , and specifically includes the following steps:
  • S721 Acquire a sentence corresponding to a word in the call text that matches a word of the sensitive word data.
  • S723 Output the call recording segment and the quality inspection report that the quality inspection is unqualified.
  • the call recording segment refers to a plurality of recorded segments obtained by segmenting the call recording file. Specifically, the sentence corresponding to the word matching the sensitive word in the call text is obtained, and then the corresponding call recording segment in the call recording file is determined according to the sentence. Optionally, in the process of converting the call recording file into the call text, each sentence after the segmentation of the sentence is correspondingly identified, for example, if “very happy to serve you” is the second sentence, then corresponding The identifier, optionally, may be identified by means of numbers, letters or Chinese characters. After determining the sentence containing the sensitive words, the corresponding call recording segment can be obtained according to the identifier of the sentence. Finally, the call recording segment corresponding to the sentence with the sensitive words and the quality inspection report with the unqualified quality check are output.
  • the sentence that may contain the sensitive words is determined by the sensitive word data, and the corresponding call recording segment is obtained, and finally outputted together with the quality inspection report that fails the quality inspection, which can facilitate subsequent recording of the corresponding call.
  • the segment is directly reviewed for the second time to ensure the accuracy of the voice quality inspection and further improve the efficiency of the voice quality inspection.
  • the quality inspection template includes at least one article template.
  • the term template refers to a standard word or sentence when the customer communicates with different business settings.
  • the term template may include: a necessary term template, a selection term template or a template.
  • the necessary clause template refers to the words or sentences that must appear in the process of business communication;
  • the selection clause template refers to at least one clause template in the process of business communication is a word or sentence that must appear;
  • the template before and after refers to Words or sentences that must appear in a certain order in the process of business communication.
  • only one clause template may be used for voice quality inspection according to the actual quality inspection requirement, or a plurality of clause templates may be used together for voice quality inspection.
  • the matching degree of the call text is obtained based on the preset quality inspection template, and specifically includes: calculating the matching degree P of the call text by using the following formula:
  • the matching ratio of the corresponding clause template i is neutralized, and ⁇ i is the weight corresponding to the clause template i.
  • the weight ⁇ i in the quality inspection template may be preset according to actual conditions. For example, in a voice quality check, there is no requirement for the template before and after, and the weight ⁇ i corresponding to the template template may be set to 0. Or, in the current voice quality check, the requirement of a certain clause template is relatively high, and the weight of the clause template can be increased accordingly. That is, in different voice quality inspection processes, it can be flexibly set by adjusting the weight of each clause template.
  • the words in the call text are matched with the words in the necessary clause template, and the ratio of the number of matching words to the total number of words in the necessary clause template is obtained as the call text and the necessary The matching ratio of the clause template;
  • the quality check of the selection clause template is performed, the words in the call text are respectively matched with the words of each clause in the selection clause template, and the number of matching words is respectively obtained as the total number of words in each clause
  • the proportion of the highest proportion is used as the matching ratio between the call text and the selection clause;
  • the quality check of the before and after clause templates is performed, the words in the call text are matched with the words of the preceding and following terms in the previous and subsequent clause templates in the order of priority. Gets the proportion of matching words in the total number of words in the template before and after, as the matching ratio of the call text and the template before and after.
  • the matching ratio calculation is performed, the calculation is performed in a single word as a minimum unit.
  • the setting can be flexibly adjusted by adjusting the weight of each clause template. It can provide data support for the judgment of subsequent matching degree and improve the efficiency of voice quality inspection.
  • Fig. 7 is a view showing a voice quality inspection apparatus corresponding to the voice quality inspection method in the first embodiment.
  • the voice quality inspection apparatus includes a call recording file acquisition module 10, a quality inspection word data acquisition module 20, a call text conversion module 30, a matching degree acquisition module 40, and a quality inspection report output module 50.
  • the steps of the call recording file acquisition module 10, the quality check word data acquisition module 20, the call text conversion module 30, the matching degree acquisition module 40, and the quality inspection report output module 50 are the same as those of the voice quality inspection method in the first embodiment.
  • One-to-one correspondence, in order to avoid redundancy, this embodiment will not be described in detail.
  • the call recording file obtaining module 10 is configured to obtain a call recording file, and the call recording file includes a file identifier.
  • the quality word data obtaining module 20 is configured to obtain corresponding quality word data based on the file identifier.
  • the call text conversion module 30 is configured to convert the call recording file into the call text by using a voice recognition algorithm based on the quality check word data.
  • the matching degree obtaining module 40 is configured to obtain the matching degree of the call text based on the preset quality inspection template.
  • the quality inspection report output module 50 is configured to output a quality inspection report that fails the quality inspection if the matching degree of the call text does not exceed the preset threshold.
  • the voice quality inspection device further includes a quality inspection word update module 60.
  • the quality check term update module 60 further includes a QA update data acquisition unit 61 and a QC data update unit 62.
  • the quality check word update data obtaining unit 61 is configured to obtain the quality check word update data, and the quality check word update data includes the quality check word data identifier.
  • the quality inspection word data updating unit 62 is configured to update the corresponding quality inspection word data based on the quality inspection word data identifier.
  • the call text conversion module 30 further includes a target pinyin element conversion unit 31, a target text data conversion unit 32, and a call text output unit 33.
  • the target pinyin element conversion unit 31 is configured to convert the call recording file into the target pinyin element by using a voice recognition algorithm.
  • the target text data conversion unit 32 is configured to match the target pinyin elements based on the quality check word data, and convert the corresponding target pinyin elements into the target text data.
  • the call text output unit 33 is configured to output the target text data as the call text.
  • the target text data conversion unit 32 further includes a business text data conversion sub-unit 321, a normal text conversion sub-unit 322, and a target text data merge sub-unit 323.
  • the business text data conversion sub-unit 321 is configured to match the target pinyin elements based on the business word data, and convert the sub-pinyin elements into the business text data according to the corresponding reference pinyin elements in the business word data.
  • the normal text data conversion sub-unit 322 is configured to match the remaining sub-pinyin elements in the target pinyin element based on the common word data, and convert the remaining sub-pinyin elements into normal text data according to the corresponding reference pinyin elements in the common word data.
  • the target text data merging sub-unit 323 is configured to merge the business text data and the normal text data to obtain the target text data.
  • the voice quality inspection device further includes a sensitive word quality inspection module 70.
  • the sensitive word quality checking module 70 further includes a sensitive word matching unit 71 and a quality inspection report output unit 72.
  • the sensitive word matching unit 71 is configured to match the call text with the sensitive word data.
  • the quality inspection report output unit 72 is configured to output a quality inspection report that fails the quality inspection if the words in the text of the call match any of the words in the sensitive word data.
  • the quality inspection report output unit 72 further includes: a sensitive sentence acquisition sub-unit 721, a call segment acquisition sub-unit 722, and a call segment output sub-unit 723.
  • the sensitive sentence obtaining sub-unit 721 is configured to obtain a sentence corresponding to a word in the call text that matches the word of the sensitive word data.
  • the call segment acquisition sub-unit 722 is configured to obtain a corresponding call recording segment in the call recording file according to the sentence.
  • the call segment output sub-unit 723 is configured to output a call recording segment and a quality inspection report that fails the quality check.
  • the matching degree obtaining module 40 is further configured to calculate the matching degree P of the call text by using the following formula:
  • n is the number of clause templates in the quality inspection template
  • i is the corresponding clause template
  • i 1, 2, 3, ..., n
  • C i is the matching ratio of the call text and the corresponding clause template i
  • ⁇ i is the weight corresponding to the clause template i.
  • the embodiment provides one or more non-volatile readable storage media having computer readable instructions stored thereon, the computer readable instructions being stored by one or more When the processors are executed, one or more processors are executed to perform the voice quality inspection method in the above embodiment. To avoid repetition, details are not described herein again.
  • the computer readable instructions are executed by one or more processors, such that one or more processors perform the functions of the various modules/units in the voice quality inspection apparatus of the above-described embodiments. To avoid repetition, details are not described herein again.
  • non-volatile readable storage medium may include any entity or device capable of carrying the computer readable instructions, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, a computer memory, a read only Memory (ROM, Read-Only Memory), Random Access Memory (RAM), electrical carrier signals, and telecommunications signals.
  • FIG. 8 is a schematic diagram of a computer device according to an embodiment of the present application.
  • computer device 80 of this embodiment includes a processor 81, a memory 82, and computer readable instructions 83 stored in memory 82 and executable on processor 81.
  • the processor 81 executes the steps of the voice quality inspection method in the first embodiment, such as steps S10 to S50 shown in FIG. 1, when the computer readable instructions 83 are executed.
  • the processor 81 implements the functions of the modules/units in the various apparatus embodiments described above when the computer readable instructions 83 are executed, such as the functions of the modules 10 through 70 shown in FIG.

Landscapes

  • Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Signal Processing (AREA)
  • Theoretical Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Biology (AREA)
  • Evolutionary Computation (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Artificial Intelligence (AREA)
  • Quality & Reliability (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Business, Economics & Management (AREA)
  • Marketing (AREA)
  • Machine Translation (AREA)

Abstract

一种语音质检方法、装置、计算机设备及存储介质,该语音质检方法包括:获取通话录音文件,所述通话录音文件包括文件标识(S10);基于所述文件标识获取对应的质检词语数据(S20);基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本(S30);基于预设的质检模板,获取所述通话文本的匹配度(S40);若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告(S50)。本方法通过质检词语数据将通话录音文件转换为通话文本,提高了转换的准确率,再根据预设的质检模板进行检验,从而查找出有问题的通话录音文件,节约了语音质检的时间,提高语音质检的效率。

Description

语音质检方法、装置、计算机设备及存储介质
本申请以2018年5月3日提交的申请号为201810412704.8,名称为“语音质检方法、装置、计算机设备及存储介质”的中国发明专利申请为基础,并要求其优先权。
技术领域
本申请涉及数据处理领域,尤其涉及一种语音质检方法、装置、计算机设备及存储介质。
背景技术
目前,呼叫中心(Call Center)是一种基于CTI(Computer Telephony Integration,计算机电话集成)技术,并充分利用通信网和计算机网络的多项功能集成与企业连为一体的完整的综合信息服务系统。呼叫中心能够有效、高速地为用户提供多种服务。其中,企业可以通过呼叫中心呼入模式获取客户的业务咨询、问题反馈、产品安装或维修受理、投诉受理、受理客户提出意见和建议等,也可以通过呼出模式主动进行市场调查、电话购物、会员回馈、电话营销、老客户回访、关怀、销售及售后跟踪服务等。
为了提高呼叫中心的服务质量,目前企业需要对坐席的通话录音文件进行质检。目前传统的做法是通过投入大量人力来对通话录音文件进行抽检。然而,由于通话录音文件的数量非常庞大,因此很多企业只能通过抽查的方式来对坐席的服务质量进行评价,无法反映所有坐席的服务质量。并且由于抽查的基数太大,样本数量也就很大,抽查人员的质检质量也无法得到保证。
发明内容
有鉴于此,本申请实施例提供一种语音质检方法、装置、计算机设备及存储介质,以解决语音质检的效率和准确性不高的问题。
一种语音质检方法,包括以下步骤:
获取通话录音文件,所述通话录音文件包括文件标识;
基于所述文件标识获取对应的质检词语数据;
基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
基于预设的质检模板,获取所述通话文本的匹配度;
若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
一种语音质检装置,包括:
通话录音文件获取模块,用于获取通话录音文件,所述通话录音文件包括文件标识;
质检词语数据获取模块,用于基于所述文件标识获取对应的质检词语数据;
通话文本转换模块,用于基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
匹配度获取模块,用于基于预设的质检模板,获取所述通话文本的匹配度;
质检报告输出模块,用于若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
一种计算机设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现如下步骤:
获取通话录音文件,所述通话录音文件包括文件标识;
基于所述文件标识获取对应的质检词语数据;
基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
基于预设的质检模板,获取所述通话文本的匹配度;
若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
一个或多个存储有计算机可读指令的非易失性可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行如下步骤:
获取通话录音文件,所述通话录音文件包括文件标识;
基于所述文件标识获取对应的质检词语数据;
基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
基于预设的质检模板,获取所述通话文本的匹配度;
若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
本申请的一个或多个实施例的细节在下面的附图和描述中提出,本申请的其他特征和优点将从说明书、附图以及权利要求变得明显。
附图说明
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本申请实施例1中的语音质检方法的一实现流程图;
图2为本申请实施例1中语音质检方法的另一实现流程图;
图3为本申请实施例1中步骤S30的一实现流程图;
图4为本申请实施例1中步骤S32的一实现流程图;
图5为本申请实施例1中语音质检方法的另一实现流程图;
图6为本申请实施例1中步骤S72的一实现流程图;
图7为本申请实施例2中提供的语音质检装置的一示意图;
图8为本申请实施例4中提供的计算机设备的一示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
实施例1
图1示出本实施例中语音质检方法的流程图。该语音质检方法应用在一个终端或系统中,以解决语音质检的效率和准确性不高的问题。特别地,可以应用在包括客服中心、呼叫中心等的通信终端或者系统中。如图1所示,该语音质检方法包括如下步骤:
S10:获取通话录音文件,通话录音文件包括文件标识。
其中,通话录音文件是指坐席与客户进行业务沟通时的录音文件。可选地,通话录音文件都保存在一个通话录音数据库中,再通过通话录音数据库来获取对应的通话录音文件。文件标识是指根据通话录音文件类型的不同而设置的标识,可以根据业务类型或者通话类型来为通话录音文件设置不同的标识。例如,根据业务类型来设置文件标识的话可以根据“理财”、“保险”或“储蓄”等不同的业务设定不同的文件标识。
可选地,从通话录音数据库中获取通话录音文件,可以根据坐席ID、时间区间或文件标识中的至少一项来获取对应的通话录音文件。其中,坐席ID是指坐席在系统或者终端中的账号,用于标识不同的坐席。
例如,当需要对X月X日坐席张三的通话录音文件进行质检时,根据时间区间(X月X日)和坐席ID(张三)从通话录音数据库中获取对应的通话录音文件。若坐席张三处理的业务是理财业务,则相应的通话录音文件的文件标识为“理财”。
通过获取通话录音文件的文件标识,可以为后续获取相应的词语数据进行语音识别作准备。
S20:基于文件标识获取对应的质检词语数据。
其中,质检词语数据是指用于对通话录音文件进行质检时用到的汉字词语数据,可以通过内部自主建立,也可以通过连接大数据平台获取。质检词语数据包括不同类型的数据,可选地,可以根据业务类型或者通话类型设置不同的质检词语数据。例如,根据“理财”、“保险”或“储蓄”这些不同的业务设置不同的质检词语数据。其中,质检词语数据类型的设置和文件标识的设置相对应。例如,若文件标识是根据业务类型来设置,则质检词语数据也是根据对应的业务类型来设置。
具体地,根据通话录音文件的具体的文件标识获取对应的质检词语数据。
例如,若通话录音的文件标识为“理财”时,则获取“理财”对应的质检词语数据。
通过获取文件标识相应的质检词语数据,可以使获取的数据更加准确,提高后续通话录音文件转换的准确率。
S30:基于质检词语数据,采用语音识别算法将通话录音文件转换为通话文本。
其中,语音识别算法是指将语音识别成文字的算法。而通话文本是指以文字形式体现的通话内容记录文本。具体地,可以采用隐马尔可夫模型(Hidden Markov Model,简称HMM)算法、动态时间归整(Dynamic Time Warping,简称DTW)算法或者基于深层神经网络(Deep Neural Network,简称DNN)算法来实现对通话录音文件的语音识别。
具体地,将通话录音文件进行噪音过滤/降噪、语句分段和语句转换处理,从而将通话录音文件转化为对应的词语元素,例如拼音。将转换后的词语元素与对应的质检词语数据中的基准词语进行匹配,将匹配后的词语组合起来得到通话文本。
例如,将通话录音文件中的“你好,我是张三”通过语句分段分为两个语句,即“你好”和“我是张三”两个句子。可选地,可以通过语音边界检测(Voice Activity Detection,简称VAD)进行语句分段,将两个句子转换为“nǐ hǎo”和“wǒ shì zhāng sān”;再与质检词语数据进行匹配,根据中文分词算法规则,例如正向最大匹配法、逆向最大匹配法、最小匹配法或者最大匹配法等等,匹配出“你好”和“我是张三”两个句子,再将两个句子进行合并组合成通话文本。
通过将通话录音文件转换为通话文本,直接对通话文本进行质检,可以节约语音质检时间,提高语音质检的效率。
S40:基于预设的质检模板,获取通话文本的匹配度。
其中,质检模板为针对具体的业务设定的用于检验坐席服务质量的文字模板。匹配度是指通话文本与质检模板中词语的匹配程度。
具体地,将通话文本中的词语与预设的质检模板中的词语进行匹配,计算匹配词语占预 设的质检模板中的词语总数的比例,作为通话文本的匹配度。
例如,预设的质检模板为“您好,很高兴为您服务”,通话文本中与质检模板匹配的词语为“您好”,而质检模板中的词语总数9,通话文本中与质检模板相匹配的词语总数为2,则匹配度=2/9=22%。
可选地,当计算通话文本的匹配度时,是以单个字为最小单位来计算比例的。
通过计算通话文本与预设质检模板的匹配度,可以检验坐席在业务中的规范程度,有助于更直观地呈现语音质检结果,也提高了语音质检的效率。
S50:若通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
其中,预设阈值是对匹配度进行设定的一个下限值,可以根据实际需要进行设定。质检报告是对语音质检结果的一种具体体现,具体地,质检报告可以包括质检结果、通话文本的匹配度和匹配的具体词语等内容,其中,质检结果包括质检合格和质检不合格。
具体地,将通话文本的匹配度与预设阈值进行比较,若匹配度小于或等于预设阈值时,则输出质检不合格的质检报告。
例如,预设阈值为80%,若通话文本的匹配度只有70%,则输出质检不合格的质检报告。
通过匹配度与预设阈值的比较,可以快速判断出质检不合格的通话录音文件。
在图1对应的实施例中,获取通话录音文件,通话录音文件包括文件标识,基于文件标识获取对应的质检词语数据,基于质检词语数据采用语音识别算法将通话录音文件转换为通话文本,基于预设的质检模板获取通话文本的匹配度,若通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。通过质检词语数据将通话录音文件转换为通话文本,提高了转换的精度,再根据预设的质检模板进行检验,从而查找出有问题的通话录音文件,节约了语音质检时间,提高语音质检的效率。
在一个具体实施方式中,该语音质检方法在步骤S20之前,还包括了对质检词语数据更新的过程,如图2所示,具体包括以下步骤:
S61:获取质检词语更新数据,质检词语更新数据包括质检词语数据标识。
其中,质检词语更新数据是指对相应的质检词语数据的数据变动的体现。例如,对质检词语数据中词语的增加、删除或修改的更新,或者对质检词语数据中部分词语的使用频次的更新等。可选地,可以连接内部数据库进行更新,也可以连接大数据平台进行更新。优选地,本实施例采用连接大数据平台的方式进行更新。质检词语数据标识是用于标识该质检词语更新数据对应的是何种类型的质检词语数据。
S62:基于质检词语数据标识更新对应的质检词语数据。
在获取质检词语更新数据之后,根据质检词语数据标识更新对应的质检词语数据。可以 理解的是,不同类型的质检词语数据中的词语均是会产生变动的,相对应的词语的使用频次也是会不断更新。在这个实施方式中,在步骤S20之前对质检词语数据进行更新可以保证质检词语数据的实时性和有效性,也提高了后续通话录音文件转换为通话文本的准确率。
在一个具体实施方式中,可以连接大数据平台,当相应的大数据平台有数据更新时,就可以实时将对应的更新数据同步到对应的质检词语数据中,可以对质检词语数据实现实时更新,进一步保证了质检词语数据的准确性,提高了后续通话录音文件转换为通话文本的转换准确率。
例如,某个质检词语数据的类型为“保险”时,将该质检词语数据连接与保险相关的大数据平台。当相关的大数据平台有数据更新时,例如增加一个词语“安心保”,或者“提存”这个词语的使用频次增加了10次。在获取到该对应的质检词语更新数据之后,根据对应的质检词语数据标识“保险”对相应的质检词语数据进行更新。
可选地,也可以设定每隔一段时间对质检词语数据进行更新,例如每10分钟更新一次,可以根据实际需要进行设定,本申请实施例不做限制。
在图2对应的实施例中,通过获取质检词语更新数据,可以使质检词语数据的词语及时地进行更新,丰富词语数据,可以有效地提高将通话录音转换为通话文本时的转换准确率。
在一个具体的实施方式中,基于质检词语数据,采用语音识别算法将通话录音文件转换为通话文本,如图3所示,具体包括以下步骤:
S31:采用语音识别算法将通话录音文件转化为目标拼音元素。
其中,目标拼音元素是指将通话录音文件转化成由拼音组成的元素,包括音节和声调。音节是音位组合构成的最小的语音结构单位,是听觉可以区分清楚的语音的基本单位,每个音节由声母、韵母两个部分组成。例如,“张”的音节为“zhang”,“张三”的音节为“zhang san”包含了两个音节。声调是汉语拼音中用于区分声音的高低和升降的属性,通常包括四声。
具体地,首先将通话录音文件进行VAD(语音边界检测)操作,从通话录音文件的声音流中识别出静音期并据此进行分帧。分帧后,声音流被分成多个小段。然后进行MFCC(Mel-Frequency Cepstral Coefficients)声学特征提取,其中,MFCC是在梅尔刻度频率域提取出来的倒谱参数,可以包括梅尔频率转化和倒谱分析两部分。梅尔刻度是一种基于人耳对等距的音高变化的感官判断而定的非线性频率刻度,与声音频率的关系如下:
Figure PCTCN2018092322-appb-000001
其中,m为梅尔刻度,f为声音频率,单位为Hz。
倒谱分析是指对时域信号做傅里叶变换,然后取对数,然后再进行反傅里叶变换的过程。 在完成声学特征提取的步骤之后,对提取的声学特征通过HMM算法、DTW算法或者基于DNN算法进行拼音元素的匹配,得到目标拼音元素。
S32:基于质检词语数据对目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据。
具体地,将目标拼音元素与对应的质检词语数据中词语的拼音进行匹配,将匹配后的词语组合在一起形成目标文本数据。
例如,文件标识为保险的通话录音文件为“您好,很高兴为您服务”,将其转化为“nín hǎo”和“hěn gāo xìng wèi nín fú wù”的目标拼音元素,将这些目标拼音元素与保险相对应的质检词语数据中的词语的拼音进行匹配,根据中文分词算法规则,例如正向最大匹配法、逆向最大匹配法、最小匹配法或者最大匹配法等等,匹配出“您好”和“很高兴为您服务”的目标文本数据。
S33:将目标文本数据输出为通话文本。
具体地,将得到的目标文本数据通过加入适当的标点符号组合在一起,形成通话文本输出。
可选地,标点符号的输入可以根据大数据平台对一个句子中最后一个词语后面的标点符号使用频率的统计,输入使用频率最高的标点符号。例如,“您好,很高兴为您服务”最后一个词语为“服务”,若根据大数据平台的统计使用频率最高的标点符号为句号,则与“服务”通过句号与下一个句子相连。
在图3对应的实施例中,通过将通话录音文件转化为包括音节和声调的目标拼音元素,再与对应的质检词语数据中词语的拼音进行匹配,可以有效地将通话录音文件转换为通话文本,提高了转换的准确率。
在一个具体实施方式中,质检词语数据包括普通词语数据和业务词语数据。
其中,普通词语数据是指普通场景下使用到的词语,业务词语数据是指对应于具体的业务而使用到的词语。可以理解的是,业务词语数据包括的词语比普通词语数据少。
在这个实施方式中,基于质检词语数据对目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据,如图4所示,具体包括以下步骤:
S321:基于业务词语数据对目标拼音元素进行匹配,将子拼音元素根据业务词语数据中对应的基准拼音元素转化为业务文本数据。
S322:基于普通词语数据对目标拼音元素中剩余的子拼音元素进行匹配,将剩余的子拼音元素根据普通词语数据中对应的基准拼音元素转化为普通文本数据。
S323:合并业务文本数据和普通文本数据,得到目标文本数据。
其中,基准拼音元素是指质检词语数据中的词语相对应的拼音元素。例如,质检词语数据中词语“保险”相对应的基准拼音元素为“bǎo xiǎn”。子拼音元素是指通话录音文件经过语句分段之后每一个句子经过转化后对应的拼音元素。
具体地,将目标拼音元素首先与通话录音文件的文件标识相对应的业务词语数据进行匹配,将能够匹配上的子拼音元素转化为对应的词语,得到业务文本数据;然后,将剩余未匹配的子拼音元素与普通词语数据进行匹配,将剩余的子拼音元素转换为对应的词语,得到普通文本数据。最后将两部分的文本数据(业务文本数据和普通文本数据)合并,得到目标文本数据。
例如,文件标识为“保险”的通话录音文件为“我们公司有一个产品安心保”,转化的目标拼音元素为“wǒ men gōng sī yǒu yī gè chǎn pǐn ān xīn bǎo”。首先将其与“保险”相对应的业务词语数据进行匹配,子拼音元素“ān xīn bǎo”匹配出“安心保”,随即将子拼音元素“ān xīn bǎo”转换为“安心保”,得到业务文本数据;然后将剩余的子拼音元素“wǒ men gōng sī yǒu yī gè chǎn pǐn”与普通词语数据进行匹配,匹配出“我们公司有一个产品”,随即将该子拼音元素“wǒ men gōng sī yǒu yī gè chǎn pǐn”转换为“我们公司有一个产品”,得到普通文本数据;最后将两部分文本数据进行合并,得到“我们公司有一个产品安心保”的目标文本数据。
可选地,当进行词语匹配时,若与业务词语数据或者普通词语数据匹配的基准拼音元素对应的词语为多个时,则输出使用频次最高的词语。例如“gōng sī”可以与“公私”和“公司”进行匹配,若“公司”比“公私”的使用频次高,则输出匹配的词语为“公司”。
在图4对应的实施例中,通过首先与业务词语数据进行匹配,再与普通词语数据进行匹配,最后合并成目标文本数据。通过设置业务词语数据并进行优先匹配可以有效提高通话录音文件转换为通话文本时的准确率。
在一具体实施方式中,在步骤S40之前,该语音质检方法还包括对敏感词语的质检,如图5所示,具体包括以下步骤:
S71:将通话文本与敏感词语数据进行匹配。
S72:若通话文本中的词语与敏感词语数据中的任一词语匹配,则输出质检不合格的质检报告。
其中,敏感词语数据是指坐席在与客户进行业务沟通过程中禁止出现的词语,例如一些不礼貌的词语。可选地,敏感词语数据可以通过对应的敏感词语数据库来进行数据的获取和更新。
具体地,将通话文本中的词语与敏感词语数据进行匹配,如果通话文本中的词语与敏感 词语数据中的任一词语匹配时,则输出质检不合格的质检报告。
在图5对应的实施例中,在基于预设的质检模板获取通话文本匹配度之前优先进行敏感词语数据的匹配,进行敏感词语的质检,通过分层次的质检方式的设置,保证了语音质检的合理性,并且在敏感词语质检输出不合格的质检报告之后对应的通话文本就不需要再进行通话文本的匹配度的获取,也提高了该语音质检方法的数据处理效率。
在一具体实施方式中,若通话文本中的词语与敏感词语数据的任一词语匹配时,则输出质检不合格的质检报告,如图6所示,具体还包括以下步骤:
S721:获取通话文本中与敏感词语数据的词语匹配的词语对应的句子。
S722:根据句子获取通话录音文件中对应的通话录音片段。
S723:输出通话录音片段和质检不合格的质检报告。
其中,通话录音片段是指对通话录音文件进行语句分段后得到的多个录音片段。具体地,获取通话文本中和敏感词语匹配的词语对应的句子,然后根据该句子确定通话录音文件中对应的通话录音片段。可选地,在将通话录音文件转换为通话文本的过程中,对语句分段后的每个句子进行相应的标识,例如,若“很高兴为您服务”为第二句,则进行相应的标识,可选地,可以采用数字、字母或汉字等方式进行标识。这样在确定含有敏感词语的句子后,就可以根据该句子的标识来获取对应的通话录音片段。最后,输出有敏感词语的句子相对应的通话录音片段和质检不合格的质检报告。
在图6对应的实施例中,通过敏感词语数据确定可能含有敏感词语的句子,并获取相应的通话录音片段,最后与质检不合格的质检报告一起输出,可以方便后续对相应的通话录音片段直接进行二次审核,保证语音质检的准确率并进一步提高语音质检的效率。
在一个具体的实施方式中,质检模板包括至少一项条款模板。
其中,条款模板是指针对不同业务设置的与客户沟通时的标准词语或句子,具体地,条款模板可以包括:必要条款模板、选择条款模板或前后条款模板。其中,必要条款模板是指在业务沟通过程中必须要出现的词语或句子;选择条款模板是指在业务沟通过程中至少有一项条款模板是必须要出现的词语或句子;前后条款模板是指在业务沟通过程中必须按照一定的先后顺序出现的词语或句子。可选地,在语音质检过程中,可以根据实际质检需要只采用一项条款模板进行语音质检,也可以采用多项条款模板共同进行语音质检。
在这个实施方式中,基于预设的质检模板,获取通话文本的匹配度,具体包括:采用以下公式计算通话文本的匹配度P:
Figure PCTCN2018092322-appb-000002
其中,质检模板包括至少一项条款模板,n为质检模板中条款模板的数量,i为相应的 条款模板(i=1,2,3,...,n),C i为通话文本中和对应的条款模板i的匹配比例,ω i为条款模板i对应的权重。
可选地,质检模板中权重ω i可以根据实际情况进行预设,例如某次语音质检中对前后条款模板没有要求,则前后条款模板对应的权重ω i可以设为0。或者在当次语音质检中对某个条款模板的要求比较高,则可以相应地提高该条款模板的权重。即在不同的语音质检过程中,可以通过对每一条款模板的权重进行调整来灵活进行设置。
具体地,在进行必要条款模板的质检时,将通话文本中的词语与必要条款模板中的词语进行匹配,获取匹配词语的个数占必要条款模板中词语总数的比例,作为通话文本和必要条款模板的匹配比例;在进行选择条款模板的质检时,将通话文本中的词语分别与选择条款模板中的每个条款的词语进行匹配,分别获取匹配词语的个数占每个条款词语总数的比例,将其中最高的比例作为通话文本与选择条款的匹配比例;在进行前后条款模板的质检时,将通话文本中的词语按照先后顺序与前后条款模板中的前后条款的词语进行匹配,获取匹配词语占前后条款模板中的词语总数的比例,作为通话文本和前后条款模板的匹配比例。可选地,进行匹配比例计算时,以单字作为最小单位进行计算。
例如,若对应的匹配比例C i为:必要条款模板80%,选择条款模板90%,前后条款模板70%;权重ω i为:必要条款模板60%,选择条款模板20%,前后条款模板20%,则匹配度P=80%*60%+90%*20%+70*20%=80%。
需要说明的是,可以根据不同业务来设定不同的质检模板和对应的权重,本申请实施例不做具体的限定。
在本申请实施例中,通过设置不同的质检模板,再通过公式计算匹配度,可以通过对每一条款模板的权重进行调整来灵活进行设置。可以为后续匹配度的判断提供数据支持,提高语音质检的效率。
应理解,上述实施例中各步骤的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本申请实施例的实施过程构成任何限定。
实施例2
图7示出与实施例1中语音质检方法一一对应的语音质检装置的示意图。如图7所示,该语音质检装置包括通话录音文件获取模块10、质检词语数据获取模块20、通话文本转换模块30、匹配度获取模块40和质检报告输出模块50。其中,通话录音文件获取模块10、质检词语数据获取模块20、通话文本转换模块30、匹配度获取模块40和质检报告输出模块50的实现功能与实施例1中语音质检方法对应的步骤一一对应,为避免赘述,本实施例不一一详述。
通话录音文件获取模块10,用于获取通话录音文件,通话录音文件包括文件标识。
质检词语数据获取模块20,用于基于文件标识获取对应的质检词语数据。
通话文本转换模块30,用于基于质检词语数据,采用语音识别算法将通话录音文件转换为通话文本。
匹配度获取模块40,用于基于预设的质检模板,获取通话文本的匹配度。
质检报告输出模块50,用于若通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
进一步地,该语音质检装置还包括质检词语更新模块60。可选地,质检词语更新模块60还包括质检词语更新数据获取单元61和质检词语数据更新单元62。
质检词语更新数据获取单元61,用于获取质检词语更新数据,质检词语更新数据包括质检词语数据标识。
质检词语数据更新单元62,用于基于质检词语数据标识更新对应的质检词语数据。
优选地,通话文本转换模块30还包括目标拼音元素转化单元31、目标文本数据转化单元32和通话文本输出单元33。
目标拼音元素转化单元31,用于采用语音识别算法将通话录音文件转化为目标拼音元素。
目标文本数据转化单元32,用于基于质检词语数据对目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据。
通话文本输出单元33,用于将目标文本数据输出为通话文本。
优选地,目标文本数据转化单元32还包括业务文本数据转化子单元321、普通文本转化子单元322和目标文本数据合并子单元323。
业务文本数据转化子单元321,用于基于业务词语数据对目标拼音元素进行匹配,将子拼音元素根据业务词语数据中对应的基准拼音元素转化为业务文本数据。
普通文本数据转化子单元322,用于基于普通词语数据对目标拼音元素中剩余的子拼音元素进行匹配,将剩余的子拼音元素根据普通词语数据中对应的基准拼音元素转化为普通文本数据。
目标文本数据合并子单元323,用于合并业务文本数据和普通文本数据,得到目标文本数据。
进一步地,该语音质检装置还包括敏感词语质检模块70。优选地,敏感词语质检模块70还包括敏感词语匹配单元71和质检报告输出单元72。
敏感词语匹配单元71,用于将通话文本与敏感词语数据进行匹配。
质检报告输出单元72,用于若通话文本中的词语与敏感词语数据中的任一词语匹配,则输出质检不合格的质检报告。
可选地,质检报告输出单元72还包括:敏感句子获取子单元721、通话片段获取子单元722和通话片段输出子单元723。
敏感句子获取子单元721,用于获取通话文本中与敏感词语数据的词语匹配的词语对应的句子。
通话片段获取子单元722,用于根据句子获取通话录音文件中对应的通话录音片段。
通话片段输出子单元723,用于输出通话录音片段和质检不合格的质检报告。
优选地,匹配度获取模块40还用于采用以下公式计算通话文本的匹配度P:
Figure PCTCN2018092322-appb-000003
其中,n为质检模板中条款模板的数量,i为相应的条款模板,i=1,2,3,...,n,C i为通话文本和对应的条款模板i的匹配比例,ω i为条款模板i对应的权重。
实施例3
本实施例提供一个或多个存储有计算机可读指令的非易失性可读存储介质,该非易失性可读存储介质上存储有计算机可读指令,该计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行上述实施例中语音质检方法,为避免重复,这里不再赘述。或者,该计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行上述实施例中语音质检装置中各模块/单元的功能,为避免重复,这里不再赘述。
可以理解地,所述非易失性可读存储介质可以包括:能够携带所述计算机可读指令的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、电载波信号和电信信号等。
实施例4
图8是本申请一实施例提供的计算机设备的示意图。如图8所示,该实施例的计算机设备80包括:处理器81、存储器82以及存储在存储器82中并可在处理器81上运行的计算机可读指令83。处理器81执行计算机可读指令83时实现上述实施例1中语音质检方法的步骤,例如图1所示的步骤S10至S50。或者,处理器81执行计算机可读指令83时实现上述各装置实施例中各模块/单元的功能,例如图7所示模块10至70的功能。
所属领域的技术人员可以清楚地了解到,为了描述的方便和简洁,仅以上述各功能单元、模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能单元、模块完成,即将所述装置的内部结构划分成不同的功能单元或模块,以完成以上描述的全部 或者部分功能。
以上所述实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围,均应包含在本申请的保护范围之内。

Claims (20)

  1. 一种语音质检方法,其特征在于,包括以下步骤:
    获取通话录音文件,所述通话录音文件包括文件标识;
    基于所述文件标识获取对应的质检词语数据;
    基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
    基于预设的质检模板,获取所述通话文本的匹配度;
    若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
  2. 如权利要求1所述的语音质检方法,其特征在于,在所述基于所述文件标识获取对应的质检词语数据的步骤之前,所述语音质检方法还包括:
    获取质检词语更新数据,所述质检词语更新数据包括质检词语数据标识;
    基于所述质检词语数据标识更新对应的质检词语数据。
  3. 如权利要求1所述的语音质检方法,其特征在于,所述基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本,包括:
    采用语音识别算法将所述通话录音文件转化为目标拼音元素;
    基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据;
    将所述目标文本数据输出为通话文本。
  4. 如权利要求3所述的语音质检方法,其特征在于,所述质检词语数据包括普通词语数据和业务词语数据;
    所述基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据,包括:
    基于所述业务词语数据对所述目标拼音元素进行匹配,将子拼音元素根据业务词语数据中对应的基准拼音元素转化为业务文本数据;
    基于所述普通词语数据对目标拼音元素中剩余的子拼音元素进行匹配,将剩余的子拼音元素根据普通词语数据中对应的基准拼音元素转化为普通文本数据;
    合并所述业务文本数据和所述普通文本数据,得到目标文本数据。
  5. 如权利要求1所述的语音质检方法,其特征在于,在所述基于预设的质检模板,获取所述通话文本的匹配度的步骤之前,所述语音质检方法还包括:
    将所述通话文本与敏感词语数据进行匹配;
    若所述通话文本中的词语与所述敏感词语数据中的任一词语匹配,则输出质检不合格的质检报告。
  6. 如权利要求5所述的语音质检方法,其特征在于,所述若所述通话文本中的词语与所述敏感词语数据的任一词语匹配时,则输出质检不合格的质检报告,还包括:
    获取所述通话文本中与所述敏感词语数据的词语匹配的词语对应的句子;
    根据所述句子获取通话录音文件中对应的通话录音片段;
    输出所述通话录音片段和质检不合格的质检报告。
  7. 如权利要求1所述的语音质检方法,其特征在于,所述质检模板包括至少一项条款模板;
    所述基于预设的质检模板,获取所述通话文本的匹配度,包括:
    采用以下公式计算所述通话文本的匹配度P:
    Figure PCTCN2018092322-appb-100001
    其中,n为质检模板中条款模板的数量,i为相应的条款模板,i=1,2,3,...,n,Ci为所述通话文本和对应的条款模板i的匹配比例,ωi为条款模板i对应的权重。
  8. 一种语音质检装置,其特征在于,包括:
    通话录音文件获取模块,用于获取通话录音文件,所述通话录音文件包括文件标识;
    质检词语数据获取模块,用于基于所述文件标识获取对应的质检词语数据;
    通话文本转换模块,用于基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
    匹配度获取模块,用于基于预设的质检模板,获取所述通话文本的匹配度;
    质检报告输出模块,用于若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
  9. 一种计算机设备,包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机可读指令,其特征在于,其特征在于,所述处理器执行所述计算机可读指令时实现如下步骤:
    获取通话录音文件,所述通话录音文件包括文件标识;
    基于所述文件标识获取对应的质检词语数据;
    基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
    基于预设的质检模板,获取所述通话文本的匹配度;
    若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
  10. 如权利要求9所述的计算机设备,其特征在于,在所述基于所述文件标识获取对应 的质检词语数据的步骤之前,所述处理器执行所述计算机可读指令时还实现如下步骤:
    获取质检词语更新数据,所述质检词语更新数据包括质检词语数据标识;
    基于所述质检词语数据标识更新对应的质检词语数据。
  11. 如权利要求9所述的计算机设备,其特征在于,所述基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本,包括:
    采用语音识别算法将所述通话录音文件转化为目标拼音元素;
    基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据;
    将所述目标文本数据输出为通话文本。
  12. 如权利要求11所述的计算机设备,其特征在于,所述质检词语数据包括普通词语数据和业务词语数据;
    所述基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据,包括:
    基于所述业务词语数据对所述目标拼音元素进行匹配,将子拼音元素根据业务词语数据中对应的基准拼音元素转化为业务文本数据;
    基于所述普通词语数据对目标拼音元素中剩余的子拼音元素进行匹配,将剩余的子拼音元素根据普通词语数据中对应的基准拼音元素转化为普通文本数据;
    合并所述业务文本数据和所述普通文本数据,得到目标文本数据。
  13. 如权利要求9所述的计算机设备,其特征在于,在所述基于预设的质检模板,获取所述通话文本的匹配度的步骤之前,所述处理器执行所述计算机可读指令时还实现如下步骤:
    将所述通话文本与敏感词语数据进行匹配;
    若所述通话文本中的词语与所述敏感词语数据中的任一词语匹配,则输出质检不合格的质检报告。
  14. 如权利要求13所述的计算机设备,其特征在于,所述若所述通话文本中的词语与所述敏感词语数据的任一词语匹配时,则输出质检不合格的质检报告,还包括:
    获取所述通话文本中与所述敏感词语数据的词语匹配的词语对应的句子;
    根据所述句子获取通话录音文件中对应的通话录音片段;
    输出所述通话录音片段和质检不合格的质检报告。
  15. 一个或多个存储有计算机可读指令的非易失性可读存储介质,其特征在于,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行如下步骤:
    获取通话录音文件,所述通话录音文件包括文件标识;
    基于所述文件标识获取对应的质检词语数据;
    基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本;
    基于预设的质检模板,获取所述通话文本的匹配度;
    若所述通话文本的匹配度未超过预设阈值,则输出质检不合格的质检报告。
  16. 如权利要求15所述的非易失性可读存储介质,其特征在于,在所述基于所述文件标识获取对应的质检词语数据的步骤之前,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器还执行如下步骤:
    获取质检词语更新数据,所述质检词语更新数据包括质检词语数据标识;
    基于所述质检词语数据标识更新对应的质检词语数据。
  17. 如权利要求15所述的非易失性可读存储介质,其特征在于,所述基于所述质检词语数据,采用语音识别算法将所述通话录音文件转换为通话文本,包括:
    采用语音识别算法将所述通话录音文件转化为目标拼音元素;
    基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据;
    将所述目标文本数据输出为通话文本。
  18. 如权利要求17所述的非易失性可读存储介质,其特征在于,所述质检词语数据包括普通词语数据和业务词语数据;
    所述基于所述质检词语数据对所述目标拼音元素进行匹配,将对应的目标拼音元素转化为目标文本数据,包括:
    基于所述业务词语数据对所述目标拼音元素进行匹配,将子拼音元素根据业务词语数据中对应的基准拼音元素转化为业务文本数据;
    基于所述普通词语数据对目标拼音元素中剩余的子拼音元素进行匹配,将剩余的子拼音元素根据普通词语数据中对应的基准拼音元素转化为普通文本数据;
    合并所述业务文本数据和所述普通文本数据,得到目标文本数据。
  19. 如权利要求15所述的非易失性可读存储介质,其特征在于,在所述基于预设的质检模板,获取所述通话文本的匹配度的步骤之前,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器还执行如下步骤:
    将所述通话文本与敏感词语数据进行匹配;
    若所述通话文本中的词语与所述敏感词语数据中的任一词语匹配,则输出质检不合格的质检报告。
  20. 如权利要求19所述的非易失性可读存储介质,其特征在于,所述若所述通话文本中的词语与所述敏感词语数据的任一词语匹配时,则输出质检不合格的质检报告,还包括:
    获取所述通话文本中与所述敏感词语数据的词语匹配的词语对应的句子;
    根据所述句子获取通话录音文件中对应的通话录音片段;
    输出所述通话录音片段和质检不合格的质检报告。
PCT/CN2018/092322 2018-05-03 2018-06-22 语音质检方法、装置、计算机设备及存储介质 Ceased WO2019210557A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201810412704.8 2018-05-03
CN201810412704.8A CN108737667B (zh) 2018-05-03 2018-05-03 语音质检方法、装置、计算机设备及存储介质

Publications (1)

Publication Number Publication Date
WO2019210557A1 true WO2019210557A1 (zh) 2019-11-07

Family

ID=63936911

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/092322 Ceased WO2019210557A1 (zh) 2018-05-03 2018-06-22 语音质检方法、装置、计算机设备及存储介质

Country Status (2)

Country Link
CN (1) CN108737667B (zh)
WO (1) WO2019210557A1 (zh)

Cited By (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111508529A (zh) * 2020-04-16 2020-08-07 深圳航天科创实业有限公司 一种动态可扩展的语音质检评分方法
CN111523317A (zh) * 2020-03-09 2020-08-11 平安科技(深圳)有限公司 语音质检方法、装置、电子设备及介质
CN111783447A (zh) * 2020-05-28 2020-10-16 中国平安财产保险股份有限公司 基于ngram距离的敏感词检测方法、装置、设备及存储介质
CN111814481A (zh) * 2020-08-24 2020-10-23 深圳市欢太科技有限公司 购物意图识别方法、装置、终端设备及存储介质
CN112037819A (zh) * 2020-09-03 2020-12-04 阳光保险集团股份有限公司 一种基于语义的语音质检方法和装置
CN112069796A (zh) * 2020-09-03 2020-12-11 阳光保险集团股份有限公司 一种语音质检方法、装置,电子设备及存储介质
CN113033160A (zh) * 2019-12-09 2021-06-25 阿里巴巴集团控股有限公司 对话的意图分类方法及设备和生成意图分类模型的方法
CN113255361A (zh) * 2021-05-19 2021-08-13 平安科技(深圳)有限公司 语音内容的自动检测方法、装置、设备以及存储介质
CN114203200A (zh) * 2021-11-17 2022-03-18 南京苏宁软件技术有限公司 一种语音质检方法、装置、计算机设备和存储介质
CN114694656A (zh) * 2022-04-09 2022-07-01 亿玛创新网络(天津)有限公司 一种音频违禁词过滤方法、装置、电子设备及存储介质
CN115118810A (zh) * 2021-03-22 2022-09-27 奇安信科技集团股份有限公司 通话内容回溯方法、装置、电子设备及存储介质
CN115223592A (zh) * 2022-07-13 2022-10-21 深圳壹账通智能科技有限公司 语音会话质检方法、装置、设备及存储介质
CN115334201A (zh) * 2022-08-08 2022-11-11 平安银行股份有限公司 有效通话的筛选方法及其系统、计算机设备
CN115396549A (zh) * 2021-05-25 2022-11-25 中国联合网络通信集团有限公司 通话违规业务处理方法、装置及电子设备
CN121284161A (zh) * 2025-12-09 2026-01-06 广州讯鸿网络技术有限公司 基于语音转写与情感识别的智能质检系统及方法

Families Citing this family (30)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109508402A (zh) * 2018-11-15 2019-03-22 上海指旺信息科技有限公司 违规用语检测方法及装置
CN109327632A (zh) * 2018-11-23 2019-02-12 深圳前海微众银行股份有限公司 客服录音的智能质检系统、方法及计算机可读存储介质
CN109767335A (zh) * 2018-12-15 2019-05-17 深圳壹账通智能科技有限公司 双录质检方法、装置、计算机设备及存储介质
CN111355838A (zh) * 2018-12-21 2020-06-30 西安中兴新软件有限责任公司 一种语音通话识别方法、装置及存储介质
CN111383658B (zh) * 2018-12-29 2023-06-09 广州市百果园信息技术有限公司 音频信号的对齐方法和装置
CN109729383B (zh) * 2019-01-04 2021-11-02 深圳壹账通智能科技有限公司 双录视频质量检测方法、装置、计算机设备和存储介质
CN109902937B (zh) * 2019-01-31 2024-07-19 平安科技(深圳)有限公司 任务数据的质检方法、装置、计算机设备及存储介质
CN109902957B (zh) * 2019-02-28 2022-12-09 腾讯科技(深圳)有限公司 一种数据处理方法和装置
CN110189770B (zh) * 2019-06-18 2021-06-25 北京达佳互联信息技术有限公司 语音数据处理方法、装置、终端、服务器及介质
CN110414790A (zh) * 2019-06-28 2019-11-05 深圳追一科技有限公司 确定质检效果的方法、装置、设备及存储介质
CN110364183A (zh) * 2019-07-09 2019-10-22 深圳壹账通智能科技有限公司 语音质检的方法、装置、计算机设备和存储介质
CN110334241B (zh) * 2019-07-10 2023-08-25 深圳前海微众银行股份有限公司 客服录音的质检方法、装置、设备及计算机可读存储介质
CN110600056A (zh) * 2019-08-09 2019-12-20 深圳市云之音科技有限公司 语音质检方法及装置
CN110784603A (zh) * 2019-10-18 2020-02-11 深圳供电局有限公司 一种离线质检用智能语音分析方法及系统
CN110931014A (zh) * 2019-12-13 2020-03-27 集奥聚合(北京)人工智能科技有限公司 基于正则匹配规则的语音识别方法及装置
CN110956956A (zh) * 2019-12-13 2020-04-03 集奥聚合(北京)人工智能科技有限公司 基于策略规则的语音识别方法及装置
CN110970026A (zh) * 2019-12-17 2020-04-07 用友网络科技股份有限公司 语音交互匹配方法、计算机设备以及计算机可读存储介质
CN111210842B (zh) * 2019-12-27 2023-04-28 中移(杭州)信息技术有限公司 语音质检方法、装置、终端及计算机可读存储介质
CN111294468A (zh) * 2020-02-07 2020-06-16 普强时代(珠海横琴)信息技术有限公司 一种客服中心呼叫用语音质检分析系统
CN111368130B (zh) * 2020-02-26 2025-01-17 深圳前海微众银行股份有限公司 客服录音的质检方法、装置、设备及存储介质
CN111723204B (zh) * 2020-06-15 2021-04-02 龙马智芯(珠海横琴)科技有限公司 语音质检区域的校正方法、装置、校正设备及存储介质
CN111883115B (zh) * 2020-06-17 2022-01-28 马上消费金融股份有限公司 语音流程质检的方法及装置
CN111800545B (zh) * 2020-06-24 2022-05-24 Oppo(重庆)智能科技有限公司 终端通话状态检测方法、装置、终端及存储介质
CN112311937A (zh) * 2020-09-25 2021-02-02 厦门天聪智能软件有限公司 一种基于sip协议抓包和语音识别的客服实时质检方法和系统
CN112634903B (zh) * 2020-12-15 2023-09-29 平安科技(深圳)有限公司 业务语音的质检方法、装置、设备及存储介质
CN112633727A (zh) * 2020-12-29 2021-04-09 清华大学 质量监控方法、装置、电子设备和存储介质
CN112885332A (zh) * 2021-01-08 2021-06-01 天讯瑞达通信技术有限公司 一种语音质检方法、系统及存储介质
CN113223532B (zh) * 2021-04-30 2024-03-05 平安科技(深圳)有限公司 客服通话的质检方法、装置、计算机设备及存储介质
CN113657773B (zh) * 2021-08-19 2023-08-29 中国平安人寿保险股份有限公司 话术质检方法、装置、电子设备及存储介质
CN115914463A (zh) * 2022-11-30 2023-04-04 阿里巴巴(中国)有限公司 风险检测方法及装置和电子设备

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104485105A (zh) * 2014-12-31 2015-04-01 中国科学院深圳先进技术研究院 一种电子病历生成方法和电子病历系统
CN105141787A (zh) * 2015-08-14 2015-12-09 上海银天下科技有限公司 服务录音的合规检查方法及装置
CN105187674A (zh) * 2015-08-14 2015-12-23 上海银天下科技有限公司 服务录音的合规检查方法及装置
CN107093431A (zh) * 2016-02-18 2017-08-25 中国移动通信集团辽宁有限公司 一种对服务质量进行质检的方法及装置
CN107886231A (zh) * 2017-11-03 2018-04-06 广州杰赛科技股份有限公司 客服的服务质量评价方法与系统

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2004252668A (ja) * 2003-02-19 2004-09-09 Fujitsu Ltd コンタクトセンタ運用管理プログラム、装置および方法
CN103164403B (zh) * 2011-12-08 2016-03-16 深圳市北科瑞声科技有限公司 视频索引数据的生成方法和系统
CN102880630A (zh) * 2012-06-26 2013-01-16 华为技术有限公司 质检的处理方法和设备
JP6180022B2 (ja) * 2013-09-27 2017-08-16 株式会社日本総合研究所 コールセンタ応答制御システム及びその応答制御方法
CN105184315B (zh) * 2015-08-26 2019-03-12 北京中电普华信息技术有限公司 一种质检处理方法及系统
CN105489222B (zh) * 2015-12-11 2018-03-09 百度在线网络技术(北京)有限公司 语音识别方法和装置
CN106486119B (zh) * 2016-10-20 2019-09-20 海信集团有限公司 一种识别语音信息的方法和装置
CN106453971B (zh) * 2016-11-22 2019-04-16 广东电网有限责任公司佛山供电局 呼叫中心质检语音的获取方法和呼叫中心质检系统
CN107547759B (zh) * 2017-08-22 2020-01-03 深圳市融壹买信息科技有限公司 一种对客服人员通话的质检方法及装置

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104485105A (zh) * 2014-12-31 2015-04-01 中国科学院深圳先进技术研究院 一种电子病历生成方法和电子病历系统
CN105141787A (zh) * 2015-08-14 2015-12-09 上海银天下科技有限公司 服务录音的合规检查方法及装置
CN105187674A (zh) * 2015-08-14 2015-12-23 上海银天下科技有限公司 服务录音的合规检查方法及装置
CN107093431A (zh) * 2016-02-18 2017-08-25 中国移动通信集团辽宁有限公司 一种对服务质量进行质检的方法及装置
CN107886231A (zh) * 2017-11-03 2018-04-06 广州杰赛科技股份有限公司 客服的服务质量评价方法与系统

Cited By (19)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113033160A (zh) * 2019-12-09 2021-06-25 阿里巴巴集团控股有限公司 对话的意图分类方法及设备和生成意图分类模型的方法
CN111523317A (zh) * 2020-03-09 2020-08-11 平安科技(深圳)有限公司 语音质检方法、装置、电子设备及介质
CN111523317B (zh) * 2020-03-09 2023-04-07 平安科技(深圳)有限公司 语音质检方法、装置、电子设备及介质
CN111508529A (zh) * 2020-04-16 2020-08-07 深圳航天科创实业有限公司 一种动态可扩展的语音质检评分方法
CN111783447A (zh) * 2020-05-28 2020-10-16 中国平安财产保险股份有限公司 基于ngram距离的敏感词检测方法、装置、设备及存储介质
CN111814481A (zh) * 2020-08-24 2020-10-23 深圳市欢太科技有限公司 购物意图识别方法、装置、终端设备及存储介质
CN111814481B (zh) * 2020-08-24 2023-11-14 深圳市欢太科技有限公司 购物意图识别方法、装置、终端设备及存储介质
CN112069796A (zh) * 2020-09-03 2020-12-11 阳光保险集团股份有限公司 一种语音质检方法、装置,电子设备及存储介质
CN112037819A (zh) * 2020-09-03 2020-12-04 阳光保险集团股份有限公司 一种基于语义的语音质检方法和装置
CN112069796B (zh) * 2020-09-03 2023-08-04 阳光保险集团股份有限公司 一种语音质检方法、装置,电子设备及存储介质
CN115118810A (zh) * 2021-03-22 2022-09-27 奇安信科技集团股份有限公司 通话内容回溯方法、装置、电子设备及存储介质
CN113255361A (zh) * 2021-05-19 2021-08-13 平安科技(深圳)有限公司 语音内容的自动检测方法、装置、设备以及存储介质
CN113255361B (zh) * 2021-05-19 2023-12-22 平安科技(深圳)有限公司 语音内容的自动检测方法、装置、设备以及存储介质
CN115396549A (zh) * 2021-05-25 2022-11-25 中国联合网络通信集团有限公司 通话违规业务处理方法、装置及电子设备
CN114203200A (zh) * 2021-11-17 2022-03-18 南京苏宁软件技术有限公司 一种语音质检方法、装置、计算机设备和存储介质
CN114694656A (zh) * 2022-04-09 2022-07-01 亿玛创新网络(天津)有限公司 一种音频违禁词过滤方法、装置、电子设备及存储介质
CN115223592A (zh) * 2022-07-13 2022-10-21 深圳壹账通智能科技有限公司 语音会话质检方法、装置、设备及存储介质
CN115334201A (zh) * 2022-08-08 2022-11-11 平安银行股份有限公司 有效通话的筛选方法及其系统、计算机设备
CN121284161A (zh) * 2025-12-09 2026-01-06 广州讯鸿网络技术有限公司 基于语音转写与情感识别的智能质检系统及方法

Also Published As

Publication number Publication date
CN108737667A (zh) 2018-11-02
CN108737667B (zh) 2021-09-10

Similar Documents

Publication Publication Date Title
WO2019210557A1 (zh) 语音质检方法、装置、计算机设备及存储介质
JP6714607B2 (ja) 音声を要約するための方法、コンピュータ・プログラムおよびコンピュータ・システム
CN109599093B (zh) 智能质检的关键词检测方法、装置、设备及可读存储介质
US11693988B2 (en) Use of ASR confidence to improve reliability of automatic audio redaction
WO2021164147A1 (zh) 基于人工智能的服务评价方法、装置、设备及存储介质
CN109087670B (zh) 情绪分析方法、系统、服务器及存储介质
WO2019037205A1 (zh) 语音欺诈识别方法、装置、终端设备及存储介质
CN111785275A (zh) 语音识别方法及装置
WO2020228173A1 (zh) 违规话术检测方法、装置、设备及计算机可读存储介质
CN111226274A (zh) 自动阻止音频流中包含的敏感数据
WO2019037382A1 (zh) 基于情绪识别的语音质检方法、装置、设备及存储介质
US10522135B2 (en) System and method for segmenting audio files for transcription
KR20220081120A (ko) 인공 지능 콜센터 시스템 및 그 시스템 기반의 서비스 제공 방법
CN110852075B (zh) 自动添加标点符号的语音转写方法、装置及可读存储介质
Kopparapu Non-linguistic analysis of call center conversations
US20230335153A1 (en) Systems and methods for classification and rating of calls based on voice and text analysis
US20180350390A1 (en) System and method for validating and correcting transcriptions of audio files
WO2019210556A1 (zh) 通话预约方法、坐席请假处理方法、装置、设备及介质
US20180342240A1 (en) System and method for assessing audio files for transcription services
KR102610360B1 (ko) 발화 보이스에 대한 레이블링 방법, 그리고 이를 구현하기 위한 장치
CN116631412A (zh) 一种通过声纹匹配判断语音机器人的方法
CN113689886B (zh) 语音数据情感检测方法、装置、电子设备和存储介质
Pandharipande et al. A novel approach to identify problematic call center conversations
CN117275458B (zh) 智能客服的语音生成方法、装置、设备及存储介质
US11621015B2 (en) Learning speech data generating apparatus, learning speech data generating method, and program

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18917262

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 18917262

Country of ref document: EP

Kind code of ref document: A1