WO2020147407A1 - 一种会议记录生成方法、装置、存储介质及计算机设备 - Google Patents

一种会议记录生成方法、装置、存储介质及计算机设备 Download PDF

Info

Publication number
WO2020147407A1
WO2020147407A1 PCT/CN2019/118256 CN2019118256W WO2020147407A1 WO 2020147407 A1 WO2020147407 A1 WO 2020147407A1 CN 2019118256 W CN2019118256 W CN 2019118256W WO 2020147407 A1 WO2020147407 A1 WO 2020147407A1
Authority
WO
WIPO (PCT)
Prior art keywords
speech
segments
voice
categories
segment
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/118256
Other languages
English (en)
French (fr)
Inventor
吴欢
田甜
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2020147407A1 publication Critical patent/WO2020147407A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/04Segmentation; Word boundary detection
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques

Definitions

  • This application relates to the field of artificial intelligence technology, and in particular to a method and device for generating meeting records.
  • the recorder records and sorts out the speech content of each speaker in the meeting to form the meeting record.
  • the meeting time is relatively long and there are more content to be recorded, manually sorting the meeting records is time-consuming, laborious and inefficient.
  • the embodiments of the present application provide a method and device for generating meeting records to solve the problems of time-consuming, labor-intensive and low efficiency in manually sorting meeting records in the prior art.
  • an embodiment of the present application provides a method for generating conference records, the method includes: obtaining conference voice; dividing the conference voice to obtain N voice fragments, where N is a natural number greater than or equal to 2; N voice segments are clustered to obtain M categories of voice segments, where M is a natural number greater than or equal to 2, and M ⁇ N.
  • the M categories of voice segments have a one-to-one correspondence with M speakers; determine all The speaker corresponding to each of the M categories of voice segments; determine the content of each of the M speakers according to the M categories of voice segments; according to the M speech segments The speech content of each speaker in the person generates meeting minutes.
  • an embodiment of the present application provides an apparatus for generating conference records.
  • the apparatus includes: an acquisition unit, configured to acquire conference voice; and a segmentation unit, configured to divide the conference voice to obtain N voice segments. Is a natural number greater than or equal to 2; the clustering unit is used to cluster the N speech fragments to obtain M categories of speech fragments, where M is a natural number greater than or equal to 2, and M ⁇ N, the M categories of The voice fragments have a one-to-one correspondence with the M speakers; the first determining unit is used to determine the speaker corresponding to each category of the M voice fragments; the second determining unit is used to The speech fragments of the M categories determine the speech content of each of the M speakers; the generating unit is configured to generate meeting records according to the speech content of each of the M speakers.
  • an embodiment of the present application provides a storage medium that includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned method for generating meeting minutes.
  • an embodiment of the present application provides a computer device, including a memory and a processor, the memory is used to store information including program instructions, the processor is used to control the execution of the program instructions, and the program instructions are executed by the processor.
  • the steps of the above-mentioned method for generating meeting records are realized.
  • the conference speech is divided to obtain N speech fragments, and the N speech fragments are clustered to obtain the speech fragments of M categories, and the speaker corresponding to the speech fragments of each category is determined;
  • the speech fragments of each category determine the speech content of M speakers;
  • meeting records are generated according to the speech content of each speaker, which solves the problem of time-consuming, labor-intensive and low efficiency in manually sorting the meeting records in the prior art, and achieves the intelligent analysis of the meeting The content of the speech, the effect of efficiently sorting out meeting minutes.
  • Fig. 1 is a flowchart of an optional method for generating meeting records according to an embodiment of the present application
  • Fig. 2 is a schematic diagram of an optional meeting record generating device according to an embodiment of the present application.
  • Fig. 3 is a schematic diagram of an optional computer device provided by an embodiment of the present application.
  • Fig. 1 is a flowchart of an optional method for generating meeting records according to an embodiment of the present application. As shown in Fig. 1, the method includes:
  • Step S102 Obtain the conference voice.
  • step S104 the conference speech is divided to obtain N speech fragments, where N is a natural number greater than or equal to 2.
  • Step S106 Cluster the N speech segments to obtain M categories of speech segments, where M is a natural number greater than or equal to 2, and M ⁇ N, and the M categories of speech segments have a one-to-one correspondence with M speakers.
  • Step S108 Determine the speaker corresponding to each of the M categories of voice segments.
  • Step S110 Determine the speech content of each of the M speakers according to the M categories of speech segments.
  • step S112 a meeting record is generated according to the content of each of the M speakers.
  • the meeting may include two situations:
  • the first situation a situation where all participants are present. For example, a department held a meeting and everyone gathered in the same conference room for a meeting.
  • the second situation the situation of holding a meeting with the help of some application software.
  • a certain department held a meeting, some people gathered in the same meeting room, and some other people were on business trips outside the city, and participated in the meeting through WeChat, QQ or other applications.
  • a company held a meeting. There were 3 participants, namely, the department manager in Beijing, the department manager in Shanghai, and the department manager in Shenzhen. These people are located in Beijing, Shanghai, and Shenzhen. These three people Meetings in different cities via WeChat, QQ or other applications.
  • the conference voice may be a voice generated during a conference in any of the above methods.
  • the conference voice can be recorded on-site during a meeting. For example, several people gather for a meeting, and one of them uses a mobile phone, recorder, voice recorder or other recording device to record the voice generated during the meeting to obtain the conference voice; the conference voice can also be It is obtained by recording the voice generated during a meeting through instant messaging software. For example, several people have a meeting through WeChat/QQ, and one of them uses a mobile phone, voice recorder, voice recorder or other recording device to record the WeChat voice/QQ generated during the meeting Voice to get the conference voice.
  • the speaker refers to the person who speaks during the meeting.
  • the number of speakers is less than or equal to the number of participants. If all participants have spoken, the number of speakers is equal to the number of participants. The number of participants; if only some participants have spoken, then the number of speakers is less than the number of participants.
  • the voice generated during the meeting is recorded to obtain the conference voice.
  • Category 2 includes 1000 speech segments, which are speech segment P(2,1), speech segment P(2,2),..., Speech segment P(2, 1000), these 1000 speech segments correspond to the same speaker;
  • category 3 includes 2000 speech segments, which are speech segment P(3,1), speech segment P(3,2),...
  • Voice segment P(3, 2000) these 2000 voice segments correspond to the same speaker. Then, respectively determine the speaker corresponding to each category of speech fragments. For example, suppose it is determined that the 3000 speech fragments included in category 1 correspond to speaker A, the 1000 speech fragments included in category 2 correspond to speaker B, and category 3 includes 2000 speech fragments correspond to speaker C, as shown in Table 1.
  • Speech segment P(1,1), the speech segment P(1,2),..., the speech segment P(1,3000 determine the speech content of the speaker A
  • the speech segment P(2,1), the speech segment P(2,2),..., speech segment P(2,1000) determine the content of speaker B's speech
  • speech segment P(3,1), speech segment P(3,2),..., speech segment P(3, 2000) determines the content of speaker C's speech.
  • the conference speech is divided to obtain N speech fragments, and the N speech fragments are clustered to obtain the speech fragments of M categories, and the speaker corresponding to the speech fragments of each category is determined;
  • the speech fragments of each category determine the speech content of M speakers;
  • meeting records are generated according to the speech content of each speaker, which solves the problem of time-consuming, labor-intensive and low efficiency in manually sorting the meeting records in the prior art, and achieves the intelligent analysis of the meeting The content of the speech, the effect of efficiently sorting out meeting minutes.
  • the speaker list includes the information of each speaker in the M speakers; the matching instruction is received, and the matching instruction is an instruction issued by the user to instruct each of the L text fragments to match the speaker; according to The matching instruction determines the speaker corresponding to each of the M categories of voice segments.
  • the selected voice segment is voice segment P(2, 1);
  • the voice segment selected from category 3 is voice segment P(3, 1). Convert these 3 speech fragments respectively to obtain text fragment F(1,1), text fragment F(2,1), text fragment F(3,1), these three text fragments and the above three speech fragments
  • Table 2 select one speech segment from each of the three types of speech segments shown in Table 1
  • the selected speech segment from category 1 is the speech segment P(1, 1); from category 2
  • the selected voice segment is voice segment P(2, 1);
  • the voice segment selected from category 3 is voice segment P(3, 1). Convert these 3 speech fragments respectively to obtain text fragment F(1,1), text fragment F(2,1), text fragment F(3,1), these three text fragments and the above three speech fragments
  • Table 2 Convert these 3 speech fragments respectively to obtain text fragment F(1,1), text fragment F(2,1),
  • Speech fragment Text fragments converted from speech fragments Speech fragment P(1,1) Text fragment F(1,1) Speech fragment P(2, 1) Text fragment F(2,1) Speech fragment P(3,1) Text fragment F(3,1)
  • the speaker list includes the information of each of the 3 speakers.
  • the spokesperson’s information may include the name and position of the spokesperson.
  • the user can be the moderator of the meeting or other participants.
  • the matching instruction is an instruction for instructing to match each of the three text segments with the speaker.
  • the matching instruction instructs to match the text segment with the speaker according to Table 3.
  • Text fragment The speaker corresponding to the text fragment Text fragment F(1,1) A Text fragment F(2,1) B Text fragment F(3,1) C
  • the text fragment F(1,1) is converted from the speech fragments in category 1, and all the speech fragments in category 1 correspond to the same speaker
  • the text fragment F(1,1) corresponds to the speaker A is the speaker corresponding to all speech fragments in category 1, that is, all speech fragments in category 1 are spoken by speaker A; in the same way, since the text fragment F(2, 1) is the speech in category 2 All speech fragments in category 2 are produced by speaker B.
  • the text fragment F(3, 1) is converted from speech fragments in category 3, all speech fragments in category 3 They are all made by speaker C, and the corresponding relationship between voice clips and speakers is shown in Table 4.
  • the speaker list includes the information of each speaker in the M speakers; the matching instruction is received, and the matching instruction is the instruction issued by the user to instruct each of the Z voice segments to match the speaker; according to The matching instruction determines the speaker corresponding to each of the M categories of voice segments.
  • At least one speech segment is selected from the speech segments of each of the M categories of speech segments. Specifically, at least one speech segment can be randomly selected from each of the M categories of speech segments.
  • the user can issue a matching instruction, and the matching instruction is an instruction used to instruct to match each of the 6 voice segments with the speaker.
  • the matching mode is shown in Table 5.
  • Speech fragment The speaker corresponding to the speech clip Speech fragment F (1, 32), speech fragment F (1, 450) A Speech fragment F (2, 100), speech fragment F (2, 400) B Speech fragment F (3,900), speech fragment F (3,600) C
  • the speech fragment F(1, 32) and the speech fragment F(1, 450) are the speech fragments in category 1, and all the speech fragments in the category 1 correspond to the same speaker, the speech fragment F(1, 32).
  • Speaker A corresponding to speech segment F (1, 450) is the speaker corresponding to all speech segments in category 1, that is, all speech segments in category 1 are spoken by speaker A; the same applies, Since the speech fragment F(2,100) and the speech fragment F(2,400) are the speech fragments in category 2, all the speech fragments in the category 2 are issued by speaker B; similarly, since the speech fragment F(3 , 900), voice segment F (3, 600) is a voice segment in category 3. All voice segments in category 3 are made by speaker C. The corresponding relationship between voice segments and speakers is shown in Table 4. .
  • the voice segments corresponding to the same speaker are clustered together according to the clustering algorithm, and then one or more voice segments are randomly selected from each category, and the selected voice segment is played to the user.
  • the user corresponds the speech fragment to the speaker; or converts the selected speech fragment into a text fragment, and shows the text fragment to the user.
  • the user is asked to correspond the text fragment to the speaker. It is very simple and convenient, and does not need to know the speaker’s information in advance. Voiceprint features or other sound-related features.
  • S1 randomly select M voice segments from N voice segments, and use the selected M voice segments as the clustering centers of M categories
  • S2 calculate the i-th voice segment among the remaining NM voice segments The distance between i speech fragments and each cluster center in M cluster centers, and classify the i-th speech fragment into the category corresponding to the cluster center with the closest distance to the i-th speech fragment, i in turn Take a natural number between 1 and NM
  • S3 After the classification of M speech fragments is completed, recalculate the cluster centers of the M categories according to the speech fragments included in each of the M categories, and update the clustering centers of the M categories For cluster centers, execute S2 and S3 in a loop until the distance between two adjacent cluster centers of each of the M categories is within the preset distance.
  • the K-means algorithm may be used to cluster the speech segments.
  • M is the number of speakers, which can be provided by the meeting host or other participants.
  • K-means algorithm is a typical distance-based clustering algorithm. It uses distance as an evaluation index of similarity, that is, it is considered that the closer the distance between two objects, the greater the similarity.
  • the algorithm considers that clusters are composed of objects close to each other, so it takes as the final goal to obtain compact and independent clusters.
  • the center of initially represents a cluster. In each iteration, the algorithm re-assigns each object to the nearest cluster according to its distance from the center of each cluster for each remaining object in the data set.
  • step S2 For the i-th speech segment of the remaining NM speech segments, calculate the distance between the i-th speech segment and each of the M cluster centers, which can be calculated by voiceprint features,
  • the specific process can be: extract the voiceprint feature of the i-th speech segment (speech segment to be clustered); extract the voiceprint feature of each of the M cluster centers; combine the voice of the i-th speech segment
  • the similarity between the pattern feature and the voiceprint feature of each of the M cluster centers is calculated, and the calculated similarity is used as the distance between the i-th speech segment and the cluster center.
  • the voiceprint features extracted in the embodiment of the present application may be prosodic features. Tone color, tone intensity, pitch, etc., are collectively referred to as prosodic features of speech, also known as supersegment features. Tone intensity shows changes in the intensity of the voice, such as the stress and light tone, and the pitch expresses the tone and intonation of the voice.
  • the voiceprint features of the speech fragments are extracted, and the voice fragments are clustered through the voiceprint features, and the voice fragments with high similarity of the voiceprint features are grouped together as the voice fragments uttered by the same speaker.
  • the conference voice is divided to obtain N voice fragments, including: determining the silent fragment in the conference voice; removing the silent fragment in the conference voice; and dividing the conference voice after the silent fragment is removed according to the silent fragment to obtain W Long speech fragments, W is a natural number greater than or equal to 2, W ⁇ N; extract the acoustic characteristics of each long speech fragment in the W long speech fragments; compare the acoustic characteristics of each long speech fragment in the W long speech fragments
  • Entropy analysis According to the results of the relative entropy analysis, the W long speech segments are segmented to obtain N speech segments.
  • perform relative entropy analysis on the acoustic characteristics of each long speech segment; segmenting the long speech segment according to the result of the relative entropy analysis includes: framing the long speech segment to obtain the speech frame of the long speech segment, Extract the acoustic features of the speech frame, perform relative entropy analysis on the acoustic features, determine the maximum value of the relative entropy, and determine whether the duration of the long speech segment is greater than the preset duration; if the duration of the long speech segment is greater than the preset duration, the relative entropy The long speech segment is segmented at the maximum value.
  • relative entropy also known as KL divergence (Kullback–Leibler divergence)
  • KL divergence KL divergence
  • FIG. 2 is a schematic diagram of an optional meeting record generating device according to an embodiment of the present application.
  • the device is used to execute the above meeting record generating method.
  • the device includes: an acquisition unit 10, a segmentation unit 20, and a gathering unit.
  • the acquiring unit 10 is used to acquire conference voice.
  • the dividing unit 20 is configured to divide the conference speech to obtain N speech segments, where N is a natural number greater than or equal to 2.
  • the clustering unit 30 is used to cluster the N speech fragments to obtain M categories of speech fragments, where M is a natural number greater than or equal to 2, and M ⁇ N.
  • M is a natural number greater than or equal to 2
  • M ⁇ N M ⁇ N
  • the first determining unit 40 is configured to determine the speaker corresponding to each category of voice segments in the M categories of voice segments.
  • the second determining unit 50 is configured to determine the speech content of each of the M speakers according to the speech segments of the M categories.
  • the generating unit 60 is configured to generate meeting records according to the content of each of the M speakers.
  • the conference speech is divided to obtain N speech fragments, and the N speech fragments are clustered to obtain the speech fragments of M categories, and the speaker corresponding to the speech fragments of each category is determined;
  • the speech fragments of each category determine the speech content of M speakers;
  • meeting records are generated according to the speech content of each speaker, which solves the problem of time-consuming, labor-intensive and low efficiency in manually sorting the meeting records in the prior art, and achieves the intelligent analysis of the meeting The content of the speech, the effect of efficiently sorting out meeting minutes.
  • the first determining unit 40 includes: a first selecting subunit, a first displaying subunit, a first receiving subunit, and a first determining subunit.
  • the first selection subunit is used to select at least one voice segment from each of the M categories of voice segments to convert into a text segment to obtain L text segments, L is a natural number, and L ⁇ M.
  • the first display subunit is used to display L text fragments and a list of speakers to the user.
  • the list of speakers includes information about each of the M speakers.
  • the first receiving subunit is configured to receive a matching instruction, and the matching instruction is an instruction issued by the user to instruct each of the L text fragments to be matched with the speaker.
  • the first determining subunit is used to determine the speaker corresponding to each of the M categories of voice segments according to the matching instruction.
  • the first determining unit 40 includes: a second selecting subunit, a second displaying subunit, a second receiving subunit, and a second determining subunit.
  • the second selection subunit is used to select at least one voice segment from each of the M categories of voice segments to obtain Z voice segments, Z is a natural number, and Z ⁇ M.
  • the second display subunit is used to play the selected Z voice clips to the user and display the speaker list.
  • the speaker list includes the information of each of the M speakers.
  • the second receiving subunit is used to receive a matching instruction, and the matching instruction is an instruction issued by the user for instructing to match each of the Z voice segments with the speaker.
  • the second determining subunit is used to determine the speaker corresponding to each of the M categories of voice segments according to the matching instruction.
  • the clustering unit is configured to perform the following steps: S1: randomly select M voice segments from N voice segments, and use the selected M voice segments as cluster centers of M categories.
  • S2 For the i-th speech segment in the remaining NM speech segments, calculate the distance between the i-th speech segment and each cluster center in the M cluster centers, and classify the i-th speech segment into In the category corresponding to the cluster center closest to the i-th speech segment, i takes a natural number between 1 and NM in turn.
  • S3 After the classification of the M speech fragments is completed, recalculate the cluster centers of the M categories according to the speech fragments included in each of the M categories, and update the cluster centers of the M categories. Repeat S2 and S3 until the distance between the two adjacent cluster centers of each of the M categories is within the preset distance.
  • the segmentation unit 20 includes: a third determination subunit, a removal subunit, a segmentation subunit, an extraction subunit, a relative entropy analysis subunit, and a molecular cutting unit.
  • the third determining subunit is used to determine the silent segment in the conference voice.
  • the removal subunit is used to remove silent segments in the conference voice.
  • the segmentation subunit is used to segment the conference speech after the silence segment is removed according to the silence segment to obtain W long speech segments, where W is a natural number greater than or equal to 2, and W ⁇ N.
  • the extraction subunit is used to extract the acoustic features of each long speech segment in the W long speech segments.
  • the relative entropy analysis subunit is used to perform relative entropy analysis on the acoustic characteristics of each long speech segment in the W long speech segments.
  • the molecular unit is used to segment W long speech fragments according to the result of relative entropy analysis to obtain N speech fragments.
  • an embodiment of the present application provides a storage medium, the storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to perform the following steps: obtain the conference voice; divide the conference voice to obtain N voices Fragment, N is a natural number greater than or equal to 2; N speech fragments are clustered to obtain M categories of speech fragments, M is a natural number greater than or equal to 2, M ⁇ N, M categories of speech fragments speak with M respectively People have a one-to-one correspondence; determine the speaker corresponding to each of the M categories of voice segments; determine the content of each of the M speakers according to the M categories of voice segments; according to M The speech content of each speaker in the speakers generates meeting minutes.
  • the device where the storage medium is controlled further executes the following steps: select at least one voice segment from each of the M categories of voice segments to convert into a text segment to obtain L text segments, L is a natural number, L ⁇ M; to show the user L text fragments and a list of speakers, the speaker list includes the information of each speaker in the M speakers; to receive a matching instruction, the matching instruction is sent by the user to indicate An instruction for matching each text segment in the L text segments with the speaker; according to the matching instruction, the speaker corresponding to each of the M categories of speech segments is determined.
  • the device where the storage medium is controlled further executes the following steps: select at least one voice segment from each of the M categories of voice segments to obtain Z voice segments, where Z is a natural number, Z ⁇ M; Play the selected Z voice clips to the user and display the speaker list, the speaker list includes the information of each speaker in the M speakers; the matching instruction is received, and the matching instruction is sent by the user to indicate Instructions for matching each voice segment in the Z voice segments with the speaker; according to the matching instruction, determine the speaker corresponding to each of the M categories of voice segments.
  • the device where the storage medium is controlled further executes the following steps: S1: randomly select M voice segments from N voice segments, and use the selected M voice segments as cluster centers of M categories; S2 : For the i-th speech segment in the remaining NM speech segments, calculate the distance between the i-th speech segment and each cluster center in the M cluster centers, and classify the i-th speech segment into and In the category corresponding to the nearest cluster center of the i-th speech segment, i takes a natural number between 1 and NM in turn; S3: After the classification of M speech segments is completed, according to the speech included in each of the M categories The segment recalculates the cluster centers of M categories and updates the cluster centers of M categories. S2 and S3 are executed in a loop until the distance between the two adjacent cluster centers of each category in the M categories is within the preset distance Inside.
  • the device where the storage medium is controlled also executes the following steps: determine the silent segment in the conference voice; remove the silent segment in the conference voice; divide the conference voice after the silent segment is removed according to the silent segment to obtain W Long speech fragments, W is a natural number greater than or equal to 2, W ⁇ N; extract the acoustic characteristics of each long speech fragment in the W long speech fragments; compare the acoustic characteristics of each long speech fragment in the W long speech fragments
  • Entropy analysis According to the results of the relative entropy analysis, the W long speech segments are segmented to obtain N speech segments.
  • an embodiment of the present application provides a computer device, including a memory and a processor, the memory is used to store information including program instructions, the processor is used to control the execution of the program instructions, and the program instructions are loaded and executed by the processor to achieve the following Steps: Obtain the conference speech; divide the conference speech to obtain N speech fragments, where N is a natural number greater than or equal to 2; cluster the N speech fragments to obtain M categories of speech fragments, and M is a natural number greater than or equal to 2 , M ⁇ N, M categories of speech fragments have a one-to-one correspondence with M speakers; determine the speaker corresponding to each category of the M categories of speech fragments; determine according to the M categories of speech fragments The speech content of each of the M speakers; meeting records are generated based on the speech content of each of the M speakers.
  • the following steps are also implemented: select at least one voice segment from each of the M categories of voice segments to convert into a text segment to obtain L text segments, L is a natural number, L ⁇ M; to show the user L text fragments and the speaker list, the speaker list includes the information of each speaker in the M speakers; to receive a matching instruction, the matching instruction is sent by the user to indicate An instruction for matching each text segment in the L text segments with the speaker; according to the matching instruction, the speaker corresponding to each of the M categories of speech segments is determined.
  • the following steps are also implemented: select at least one voice segment from each of the M categories of voice segments to obtain Z voice segments, where Z is a natural number, Z ⁇ M; Play the selected Z voice clips to the user and display the speaker list, the speaker list includes the information of each speaker in the M speakers; the matching instruction is received, and the matching instruction is sent by the user to indicate Instructions for matching each voice segment in the Z voice segments with the speaker; according to the matching instruction, determine the speaker corresponding to each of the M categories of voice segments.
  • S1 randomly select M voice segments from N voice segments, and use the selected M voice segments as cluster centers of M categories
  • S2 For the i-th speech segment in the remaining NM speech segments, calculate the distance between the i-th speech segment and each cluster center in the M cluster centers, and classify the i-th speech segment into and In the category corresponding to the nearest cluster center of the i-th speech segment, i takes a natural number between 1 and NM in turn
  • S3 After the classification of M speech segments is completed, according to the speech included in each of the M categories The segment recalculates the cluster centers of M categories and updates the cluster centers of M categories.
  • S2 and S3 are executed in a loop until the distance between the two adjacent cluster centers of each category in the M categories is within the preset distance Inside.
  • the following steps are also implemented: determine the silent segment in the conference voice; remove the silent segment in the conference voice; divide the conference voice after the silent segment is removed according to the silent segment to obtain W Long speech fragments, W is a natural number greater than or equal to 2, W ⁇ N; extract the acoustic characteristics of each long speech fragment in the W long speech fragments; compare the acoustic characteristics of each long speech fragment in the W long speech fragments
  • Entropy analysis According to the results of the relative entropy analysis, the W long speech segments are segmented to obtain N speech segments.
  • FIG. 3 is a schematic diagram of a computer device provided by an embodiment of the present application.
  • the computer device 50 of this embodiment includes: a processor 51, a memory 52, and a computer program 53 stored in the memory 52 and running on the processor 51.
  • the computer program 53 is executed by the processor 51, To implement the method for generating meeting records in the embodiment, in order to avoid repetition, it will not be repeated here.
  • the computer program is executed by the processor 51, the function of each model/unit in the meeting record generating device in the embodiment is realized. In order to avoid repetition, it will not be repeated here.
  • the computer device 50 may be a computing device such as a desktop computer, a notebook, a palmtop computer, and a cloud server.
  • the computer equipment may include, but is not limited to, a processor 51 and a memory 52.
  • FIG. 3 is only an example of the computer device 50, and does not constitute a limitation on the computer device 50, and may include more or less components than shown, or combine some components, or different components.
  • computer equipment may also include input and output devices, network access devices, buses, and so on.
  • the so-called processor 51 can be a central processing unit (Central Processing Unit, CPU), other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc.
  • the general-purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
  • the memory 52 may be an internal storage unit of the computer device 50, such as a hard disk or memory of the computer device 50.
  • the memory 52 may also be an external storage device of the computer device 50, for example, a plug-in hard disk equipped on the computer device 50, a smart memory card (Smart Media (SMC), a secure digital (SD) card, and a flash memory card (Flash Card) etc.
  • the memory 52 may also include both the internal storage unit of the computer device 50 and the external storage device.
  • the memory 52 is used to store computer programs and other programs and data required by computer devices.
  • the memory 52 may also be used to temporarily store data that has been or will be output.
  • the disclosed system, device, and method may be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the unit is only a logical function division, and there may be other divisions in actual implementation, for example, multiple units or components may be combined Or it can be integrated into another system, or some features can be ignored or not implemented.
  • the displayed or discussed mutual coupling or direct coupling or communication connection may be indirect coupling or communication connection through some interfaces, devices or units, and may be in electrical, mechanical, or other forms.
  • the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Theoretical Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Telephonic Communication Services (AREA)
  • Catching Or Destruction (AREA)

Abstract

一种会议记录生成方法、装置、存储介质及计算机设备,涉及人工智能技术领域,该方法包括:获取会议语音(S102);将会议语音进行分割,得到N个语音片段,N为大于等于2的自然数(S104);将N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,M个类别的语音片段分别与M个发言人具有一一对应关系(S106);确定M个类别的语音片段中每个类别的语音片段对应的发言人(S108);根据M个类别的语音片段确定M个发言人中每个发言人的发言内容(S110);根据M个发言人中每个发言人的发言内容生成会议记录(S112)。上述方法能够解决现有技术中人工整理会议记录费时费力、效率低的问题。

Description

一种会议记录生成方法、装置、存储介质及计算机设备
本申请要求于2019年01月16日提交中国专利局、申请号为201910038460.6、申请名称为“一种会议记录生成方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
【技术领域】
本申请涉及人工智能技术领域,尤其涉及一种会议记录生成方法和装置。
【背景技术】
在会议过程中,由记录人员把会议的各个发言人的发言内容记录并整理,形成会议记录。当会议时间比较长,需要记录的内容比较多的时候,人工整理会议记录费时费力、效率低。
【申请内容】
有鉴于此,本申请实施例提供了一种会议记录生成方法和装置,用以解决现有技术中人工整理会议记录费时费力、效率低的问题。
一方面,本申请实施例提供了一种会议记录生成方法,所述方法包括:获取会议语音;将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;根据所述M个发言人中每个发言人的发言内容生成会议记录。
一方面,本申请实施例提供了一种会议记录生成装置,所述装置包括:获取单元,用于获取会议语音;分割单元,用于将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;聚类单元,用于将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;第一确定单元,用于确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;第二确定单元,用于根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;生成单元,用于根据所述M个发言人中每个发言人的发言内容生成会议记录。
一方面,本申请实施例提供了一种存储介质,所述存储介质包括存储的程序,其中,在所述程序运行时控制所述存储介质所在设备执行上述的会议记录生成方法。
一方面,本申请实施例提供了一种计算机设备,包括存储器和处理器,所述存储器用于存储包括程序指令的信息,所述处理器用于控制程序指令的执行,所述程序指令被处理器加载并执行时实现上述的会议记录生成方法的步骤。
在本申请实施例中,将会议语音进行分割,得到N个语音片段,将N个语音片段进行聚类,得到M个类别的语音片段,确定每个类别的语音片段对应的发言人;根据M个类别的语音片段确定M个发言人的发言内容;根据各个发言人的发言内容,生成会议记录,解决了现有技术中人工整理会议记录费时费力、效率低的问题,达到了智能分析会议上的发言内容,高效整理出会议记录的效果。
【附图说明】
为了更清楚地说明本申请实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其它的附图。
图1是根据本申请实施例一种可选的会议记录生成方法的流程图;
图2是根据本申请实施例一种可选的会议记录生成装置的示意图;
图3是本申请实施例提供的一种可选的计算机设备的示意图。
【具体实施方式】
为了更好的理解本申请的技术方案,下面结合附图对本申请实施例进行详细描述。
应当明确,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其它实施例,都属于本申请保护的范围。
在本申请实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本申请。在本申请实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。
应当理解,本文中使用的术语“和/或”仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。另外,本文中 字符“/”,一般表示前后关联对象是一种“或”的关系。
图1是根据本申请实施例一种可选的会议记录生成方法的流程图,如图1所示,该方法包括:
步骤S102,获取会议语音。
步骤S104,将会议语音进行分割,得到N个语音片段,N为大于等于2的自然数。
步骤S106,将N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,M个类别的语音片段分别与M个发言人具有一一对应关系。
步骤S108,确定M个类别的语音片段中每个类别的语音片段对应的发言人。
步骤S110,根据M个类别的语音片段确定M个发言人中每个发言人的发言内容。
步骤S112,根据M个发言人中每个发言人的发言内容生成会议记录。
在本申请实施例中,会议可以包括两种情况:
第一种情况:所有参会人员都到场的情况。例如,某部门举行了一次会议,所有人聚在同一个会议室开会。
第二种情况:借助于某些应用软件进行开会的情况。例如,某部门举行了一次会议,一些人聚在同一个会议室,另一些人在外地出差,通过微信、QQ或其他应用软件参加会议。再例如,某公司举行了一次会议,参会人员有3个,分别为北京的部门经理、上海的部门经理、深圳的部门经理,这些人的地理位置分别位于北京、上海、深圳,这三个人在不同的城市通过微信、QQ或其他应用软件开会。
在本申请实施例中,会议语音可以为以上任意一种方式开会的过程中产生的语音。会议语音可以是在开会时现场录制的,例如,若干人聚在一起开会,其中一个人用手机、录音机、录音笔或者其他录音设备录制会议过程中产生的语音得到会议语音;会议语音也可以是通过录制通过即时通讯软件开会的过程中产生的语音得到的,例如,若干人通过微信/QQ开会,其中一个人用手机、录音机、录音笔或者其他录音设备录制会议过程中产生的微信语音/QQ语音得到会议语音。
在本申请实施例中,发言人指的是在会议过程中发言的人,发言人的数量小于或等于参会人员的数量,如果所有参会人员都发言了,那么发言人的数量等于参会人员的数量;如果只有部分参会人员发言了,那么发言人的数量小于参会人员的数量。
下面举一个具体的例子对本申请实施例提供的会议记录生成方法进行说明。
例如,若干个人开会,录制会议过程中产生的语音,得到会议语音,假设会议语音为20分钟,将会议语音进行分割,例如得到6000(N=6000)个语音片段,将这6000个语音片段进行聚类,得到3(M=3)个类别的语音片段,其中,类别1包括3000个语音片段,分别为语音片段P(1,1)、语音片段P(1,2)、……、语音片段P(1,3000),这3000个语音片段对应同一个发言人;类别2包括1000个语音片段,分别为语音片段P(2,1)、语音片段P(2,2)、……、语音片段P(2,1000),这1000个语音片段对应同一个发言人;类别3包括2000个语音片段,分别为语音片段P(3,1)、语音片段P(3,2)、……、语音片段P(3,2000),这2000个语音片段对应同一个发言人。然后,分别确定每个类别的语音片段对应的发言人,例如,假设确定出类别1包括的3000个语音片段对应发言人甲,类别2包括的1000个语音片段对应发言人乙,类别3包括的2000个语音片段对应发言人丙,如表1所示。根据语音片段P(1,1)、语音片段P(1,2)、……、语音片段P(1,3000)确定发言人甲的发言内容;根据语音片段P(2,1)、语音片段P(2,2)、……、语音片段P(2,1000)确定发言人乙的发言内容;根据语音片段P(3,1)、语音片段P(3,2)、……、语音片段P(3,2000)确定发言人丙的发言内容。根据发言人甲、发言人乙和发言人丙的发言内容生成会议记录。
表1
Figure PCTCN2019118256-appb-000001
在本申请实施例中,将会议语音进行分割,得到N个语音片段,将N个语音片段进行聚类,得到M个类别的语音片段,确定每个类别的语音片段对应的发言人;根据M个类别的语音片段确定M个发言人的发言内容;根据各个发言人的发言内容,生成会议记录,解决了现有技术中人工整理会议记录费时费力、效率低的问题,达到了智能分析会议上的发言内容,高效整理出会议记录的效果。
确定每个类别的语音片段对应的发言人,具体方法可以有多种,下面举出几种。
方法一:
从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;向用户展示L个文本片段和发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将L个文本片段中每个文本片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,具体地,可以从M个类别的语音片段中每个类别的语音片段中各随机选择至少一个语音片段转换成文本片段。
例如,从表1所示的3个类别的语音片段中每个类别的语音片段中各选择一个语音片段,从类别1中选出的语音片段为语音片段P(1,1);从类别2中选出的语音片段为语音片段P(2,1);从类别3中选出的语音片段为语音片段P(3,1)。将这3个语音片段分别进行转换,得到文本片段F(1,1)、文本片段F(2,1)、文本片段F(3,1),这三个文本片段与上述三个语音片段之间的对应关系如表2所示。
表2
语音片段 语音片段转化得到的文本片段
语音片段P(1,1) 文本片段F(1,1)
语音片段P(2,1) 文本片段F(2,1)
语音片段P(3,1) 文本片段F(3,1)
向用户展示这3(L=3)个文本片段和发言人列表,发言人列表包括3个发言人中每个发言人的信息。发言人的信息可以包括发言人的姓名、职位等。
用户可以是会议的主持人,可以是其他参会人员。
用户看到这3个文本片段后,即可知道与文本片段相对应的是哪位参 会人员的发言。例如,假设有一个文本片段的内容是:“大家好,我是今天会议的主持人。”当用户看到这个文本片段后,即可知道这个文本片段对应的是会议主持人的发言。用户可发出匹配指令,匹配指令为用于指示将3个文本片段中每个文本片段与发言人进行匹配的指令,例如,匹配指令指示按照表3将文本片段与发言人进行匹配。
表3
文本片段 文本片段对应的发言人
文本片段F(1,1)
文本片段F(2,1)
文本片段F(3,1)
由于文本片段F(1,1)是类别1中的语音片段转换得到的,而类别1中的所有语音片段对应的是同一个发言人,因此,文本片段F(1,1)对应的发言人甲即为类别1中的所有语音片段对应的发言人,即,类别1中的所有语音片段都是发言人甲发出的;同理,由于文本片段F(2,1)是类别2中的语音片段转换得到的,类别2中的所有语音片段都是发言人乙发出的;同理,由于文本片段F(3,1)是类别3中的语音片段转换得到的,类别3中的所有语音片段都是发言人丙发出的,语音片段与发言人之间的对应关系如表4所示。
表4
Figure PCTCN2019118256-appb-000002
方法二:
从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;向用户播放选择出的Z个语音片段并展示发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将Z个语音片段中每个语音片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
从M个类别的语音片段中每个类别的语音片段中各选择至少一个语 音片段,具体地,可以从M个类别的语音片段中每个类别的语音片段中各随机选择至少一个语音片段。
假设从类别1中随机选择了2个语音片段,分别为语音片段F(1,32)、语音片段F(1,450);从类别2中随机选择了2个语音片段,分别为语音片段F(2,100)、语音片段F(2,400);从类别3中随机选择了2个语音片段,分别为语音片段F(3,900)、语音片段F(3,600)。
向用户播放这6(Z=6)个语音片段,用户听到这6个语音片段后,根据声音的音色能够轻松识别每个语音片段是哪位参会人员的发言。用户可发出匹配指令,匹配指令为用于指示将6个语音片段中每个语音片段与发言人进行匹配的指令,匹配方式如表5所示。
表5
语音片段 语音片段对应的发言人
语音片段F(1,32)、语音片段F(1,450)
语音片段F(2,100)、语音片段F(2,400)
语音片段F(3,900)、语音片段F(3,600)
由于语音片段F(1,32)、语音片段F(1,450)是类别1中的语音片段,而类别1中的所有语音片段对应的是同一个发言人,因此,语音片段F(1,32)、语音片段F(1,450)对应的发言人甲即为类别1中的所有语音片段对应的发言人,即,类别1中的所有语音片段都是发言人甲发出的;同理,由于语音片段F(2,100)、语音片段F(2,400)是类别2中的语音片段,类别2中的所有语音片段都是发言人乙发出的;同理,由于语音片段F(3,900)、语音片段F(3,600)是类别3中的语音片段,类别3中的所有语音片段都是发言人丙发出的,语音片段与发言人之间的对应关系如表4所示。
在本申请实施例中,通过根据聚类算法将同一个发言人对应的语音片段聚类到一起,然后从每个类别中随机选择一个或多个语音片段,向用户播放选择的语音片段,请用户将语音片段与发言人进行对应;或者将选择出的语音片段转换成文本片段,向用户展示文本片段,请用户将文本片段与发言人进行对应,非常简单方便,不需要事先知道发言人的声纹特征或其他声音相关的特征。
将N个语音片段进行聚类的具体过程如下:
S1:从N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;S2:对于剩余的N-M个语音片段中的第i个语音片段,计算第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将第i个语音片段归类到与第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;S3:当M个语音片段归类完成之后,根据M个类别中每个类别包括的语音片段重新计算M 个类别的聚类中心,并更新M个类别的聚类中心,循环执行S2和S3,直到M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
在本申请实施例中,可采用K-means算法对语音片段进行聚类。M即为发言人的数量,该数量可以由会议主持人或其他参会人员提供。
K-means算法是很典型的基于距离的聚类算法,采用距离作为相似性的评价指标,即认为两个对象的距离越近,其相似度就越大。该算法认为簇是由距离靠近的对象组成的,因此把得到紧凑且独立的簇作为最终目标。初始类聚类中心点的选取对聚类结果具有较大的影响,因为在该算法第一步中是随机的选取任意k(在本申请实施例中,k=M)个对象作为初始聚类的中心,初始地代表一个簇。该算法在每次迭代中对数据集中剩余的每个对象,根据其与各个簇中心的距离将每个对象重新赋给最近的簇。当考察完所有数据对象后,一次迭代运算完成,新的聚类中心被计算出来。如果在一次迭代前后,新的质心与原质心相等或小于指定阈值,算法结束。在本申请实施例中,当M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内时,循环结束,得到聚类结果。
上述步骤S2:对于剩余的N-M个语音片段中的第i个语音片段,计算第i个语音片段与M个聚类中心中每个聚类中心之间的距离,可以通过声纹特征进行计算,具体过程可以是:提取第i个语音片段(待聚类的语音片段)的声纹特征;提取M个聚类中心中的每个聚类中心的声纹特征;将第i个语音片段的声纹特征与M个聚类中心中的每个聚类中心的声纹特征进行相似度计算,将计算出的相似度作为第i个语音片段与聚类中心之间的距离。
由于每个人与发音有关的解剖学结构不同,并且受社会经济状况、教育水平、出生地等影响,不同人的声纹特征不完全相同。本申请实施例中提取的声纹特征可以为韵律特征。音色、音强、音高等,总称为语音的韵律特征,又称超音段特征。音强显示语音的重音、轻音等强弱变化,音高表现语音的字调与语调。
在本申请实施例中,提取语音片段的声纹特征,通过声纹特征对语音片段进行聚类,将声纹特征相似度高的语音片段聚在一起,作为同一个发言人发出的语音片段,在这个过程中,并需要预先知道发言人的声纹特征,更不需要预先存储发言人的声纹特征,保护了发言人的隐私,安全性高,用户体验好。
可选地,将会议语音进行分割,得到N个语音片段,包括:确定会议语音中的静音片段;去除会议语音中的静音片段;根据静音片段对去除静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;提取W个长语音片段中每一个长语音片段的声学特征;对W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;根据相对熵分析的结果对W个长语音片段进行切分,得到N个语音片段。
可选地,对每一个长语音段的声学特征进行相对熵分析;根据相对熵分析的结果对长语音段进行切分,包括:对长语音段进行分帧,得到长语音段的语音帧,提取语音帧的声学特征,对声学特征进行相对熵分析,确定相对熵的最大值处,判断长语音段的时长是否大于预设时长;如果长语音段的时长大于预设时长,在相对熵的最大值处对长语音段进行切分。
在概率论或信息论中,相对熵(relative entropy),又称KL散度(Kullback–Leibler divergence),是描述两个概率分布P和Q差异的一种方法。它是非对称的,这意味着D(P||Q)≠D(Q||P)。特别的,在信息论中,D(P||Q)表示当用概率分布Q来拟合真实分布P时,产生的信息损耗,其中P表示真实分布,Q表示P的拟合分布。
对一个离散随机变量的两个概率分布P和Q来说,它们的KL散度定义为:D(P||Q)=∑P(i)lnP(i)/Q(i),对于连续的随机变量,定义类似。
图2是根据本申请实施例一种可选的会议记录生成装置的示意图,该装置用于执行上述会议记录生成方法,如图2所示,该装置包括:获取单元10、分割单元20、聚类单元30、第一确定单元40、第二确定单元50、生成单元60。
获取单元10,用于获取会议语音。
分割单元20,用于将会议语音进行分割,得到N个语音片段,N为大于等于2的自然数。
聚类单元30,用于将N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,M个类别的语音片段分别与M个发言人具有一一对应关系。
第一确定单元40,用于确定M个类别的语音片段中每个类别的语音片段对应的发言人。
第二确定单元50,用于根据M个类别的语音片段确定M个发言人中每个发言人的发言内容。
生成单元60,用于根据M个发言人中每个发言人的发言内容生成会议记录。
在本申请实施例中,将会议语音进行分割,得到N个语音片段,将N个语音片段进行聚类,得到M个类别的语音片段,确定每个类别的语音片段对应的发言人;根据M个类别的语音片段确定M个发言人的发言内容;根据各个发言人的发言内容,生成会议记录,解决了现有技术中人工整理会议记录费时费力、效率低的问题,达到了智能分析会议上的发言内容,高效整理出会议记录的效果。
可选地,第一确定单元40包括:第一选择子单元、第一展示子单元、第一接收子单元、第一确定子单元。第一选择子单元,用于从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M。第一展示子单元,用 于向用户展示L个文本片段和发言人列表,发言人列表包括M个发言人中每个发言人的信息。第一接收子单元,用于接收匹配指令,匹配指令为用户发出的用于指示将L个文本片段中每个文本片段与发言人进行匹配的指令。第一确定子单元,用于根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,第一确定单元40包括:第二选择子单元、第二展示子单元、第二接收子单元、第二确定子单元。第二选择子单元,用于从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M。第二展示子单元,用于向用户播放选择出的Z个语音片段并展示发言人列表,发言人列表包括M个发言人中每个发言人的信息。第二接收子单元,用于接收匹配指令,匹配指令为用户发出的用于指示将Z个语音片段中每个语音片段与发言人进行匹配的指令。第二确定子单元,用于根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,聚类单元用于执行以下步骤:S1:从N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心。S2:对于剩余的N-M个语音片段中的第i个语音片段,计算第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将第i个语音片段归类到与第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数。S3:当M个语音片段归类完成之后,根据M个类别中每个类别包括的语音片段重新计算M个类别的聚类中心,并更新M个类别的聚类中心。循环执行S2和S3,直到M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
可选地,分割单元20包括:第三确定子单元、去除子单元、分割子单元、提取子单元、相对熵分析子单元、切分子单元。第三确定子单元,用于确定会议语音中的静音片段。去除子单元,用于去除会议语音中的静音片段。分割子单元,用于根据静音片段对去除静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N。提取子单元,用于提取W个长语音片段中每一个长语音片段的声学特征。相对熵分析子单元,用于对W个长语音片段中每一个长语音片段的声学特征进行相对熵分析。切分子单元,用于根据相对熵分析的结果对W个长语音片段进行切分,得到N个语音片段。
一方面,本申请实施例提供了一种存储介质,存储介质包括存储的程序,其中,在程序运行时控制存储介质所在设备执行以下步骤:获取会议语音;将会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;将N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,M个类别的语音片段分别与M个发言人具有一一对应关系;确定M个类别的语音片段中每个类别的语音片段对应的发言人;根据M个类别的语音片段确定M个发言人中每个发言人的发言内容; 根据M个发言人中每个发言人的发言内容生成会议记录。
可选地,在程序运行时控制存储介质所在设备还执行以下步骤:从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;向用户展示L个文本片段和发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将L个文本片段中每个文本片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,在程序运行时控制存储介质所在设备还执行以下步骤:从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;向用户播放选择出的Z个语音片段并展示发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将Z个语音片段中每个语音片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,在程序运行时控制存储介质所在设备还执行以下步骤:S1:从N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;S2:对于剩余的N-M个语音片段中的第i个语音片段,计算第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将第i个语音片段归类到与第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;S3:当M个语音片段归类完成之后,根据M个类别中每个类别包括的语音片段重新计算M个类别的聚类中心,并更新M个类别的聚类中心,循环执行S2和S3,直到M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
可选地,在程序运行时控制存储介质所在设备还执行以下步骤:确定会议语音中的静音片段;去除会议语音中的静音片段;根据静音片段对去除静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;提取W个长语音片段中每一个长语音片段的声学特征;对W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;根据相对熵分析的结果对W个长语音片段进行切分,得到N个语音片段。
一方面,本申请实施例提供了一种计算机设备,包括存储器和处理器,存储器用于存储包括程序指令的信息,处理器用于控制程序指令的执行,程序指令被处理器加载并执行时实现以下步骤:获取会议语音;将会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;将N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,M个类别的语音片段分别与M个发言人具有一一对应关系;确定M个类别的语音片段中每个类别的语音片段对应的发言人;根据M个类别的语音片段确定M个发言人中每个发言人的发言内容;根据M个发言 人中每个发言人的发言内容生成会议记录。
可选地,程序指令被处理器加载并执行时还实现以下步骤:从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;向用户展示L个文本片段和发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将L个文本片段中每个文本片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,程序指令被处理器加载并执行时还实现以下步骤:从M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;向用户播放选择出的Z个语音片段并展示发言人列表,发言人列表包括M个发言人中每个发言人的信息;接收匹配指令,匹配指令为用户发出的用于指示将Z个语音片段中每个语音片段与发言人进行匹配的指令;根据匹配指令确定M个类别的语音片段中每个类别的语音片段对应的发言人。
可选地,程序指令被处理器加载并执行时还实现以下步骤:S1:从N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;S2:对于剩余的N-M个语音片段中的第i个语音片段,计算第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将第i个语音片段归类到与第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;S3:当M个语音片段归类完成之后,根据M个类别中每个类别包括的语音片段重新计算M个类别的聚类中心,并更新M个类别的聚类中心,循环执行S2和S3,直到M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
可选地,程序指令被处理器加载并执行时还实现以下步骤:确定会议语音中的静音片段;去除会议语音中的静音片段;根据静音片段对去除静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;提取W个长语音片段中每一个长语音片段的声学特征;对W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;根据相对熵分析的结果对W个长语音片段进行切分,得到N个语音片段。
图3是本申请实施例提供的一种计算机设备的示意图。如图3所示,该实施例的计算机设备50包括:处理器51、存储器52以及存储在存储器52中并可在处理器51上运行的计算机程序53,该计算机程序53被处理器51执行时实现实施例中的会议记录生成方法,为避免重复,此处不一一赘述。或者,该计算机程序被处理器51执行时实现实施例中会议记录生成装置中各模型/单元的功能,为避免重复,此处不一一赘述。
计算机设备50可以是桌上型计算机、笔记本、掌上电脑及云端服务器等计算设备。计算机设备可包括,但不仅限于,处理器51、存储 器52。本领域技术人员可以理解,图3仅仅是计算机设备50的示例,并不构成对计算机设备50的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件,例如计算机设备还可以包括输入输出设备、网络接入设备、总线等。
所称处理器51可以是中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
存储器52可以是计算机设备50的内部存储单元,例如计算机设备50的硬盘或内存。存储器52也可以是计算机设备50的外部存储设备,例如计算机设备50上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,存储器52还可以既包括计算机设备50的内部存储单元也包括外部存储设备。存储器52用于存储计算机程序以及计算机设备所需的其他程序和数据。存储器52还可以用于暂时地存储已经输出或者将要输出的数据。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统,装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统,装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如,多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
以上所述仅为本申请的较佳实施例而已,并不用以限制本申请,凡在本申请的精神和原则之内,所做的任何修改、等同替换、改进等,均应包含在本申请保护的范围之内。

Claims (20)

  1. 一种会议记录生成方法,其特征在于,所述方法包括:
    获取会议语音;
    将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;
    将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;
    确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;
    根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;
    根据所述M个发言人中每个发言人的发言内容生成会议记录。
  2. 根据权利要求1所述的方法,其特征在于,所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;
    向用户展示所述L个文本片段和发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述L个文本片段中每个文本片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  3. 根据权利要求1所述的方法,其特征在于,所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;
    向用户播放选择出的所述Z个语音片段并展示发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述Z个语音片段中每个语音片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  4. 根据权利要求1所述的方法,其特征在于,所述将所述N个语音片段进行聚类,包括:
    S1:从所述N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;
    S2:对于剩余的N-M个语音片段中的第i个语音片段,计算所述第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将所述第i个语音片段归类到与所述第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;
    S3:当所述M个语音片段归类完成之后,根据所述M个类别中每个类别包括的语音片段重新计算所述M个类别的聚类中心,并更新所述M个类别的聚类中心,
    循环执行S2和S3,直到所述M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
  5. 根据权利要求1至4任一项所述的方法,其特征在于,所述将所述会议语音进行分割,得到N个语音片段,包括:
    确定所述会议语音中的静音片段;
    去除所述会议语音中的静音片段;
    根据所述静音片段对去除所述静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;
    提取所述W个长语音片段中每一个长语音片段的声学特征;
    对所述W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;
    根据相对熵分析的结果对所述W个长语音片段进行切分,得到所述N个语音片段。
  6. 一种会议记录生成装置,其特征在于,所述装置包括:
    获取单元,用于获取会议语音;
    分割单元,用于将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;
    聚类单元,用于将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;
    第一确定单元,用于确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;
    第二确定单元,用于根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;
    生成单元,用于根据所述M个发言人中每个发言人的发言内容生成会议记录。
  7. 根据权利要求6所述的装置,其特征在于,所述第一确定单元包括:
    第一选择子单元,用于从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;
    第一展示子单元,用于向用户展示所述L个文本片段和发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    第一接收子单元,用于接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述L个文本片段中每个文本片段与发言人进行匹配的指令;
    第一确定子单元,用于根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  8. 根据权利要求6所述的装置,其特征在于,所述第一确定单元包括:
    第二选择子单元,用于从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;
    第二展示子单元,用于向用户播放选择出的所述Z个语音片段并展示发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    第二接收子单元,用于接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述Z个语音片段中每个语音片段与发言人进行匹配的指令;
    第二确定子单元,用于根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  9. 根据权利要求6所述的装置,其特征在于,所述聚类单元用于执行以下步骤:
    S1:从所述N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;
    S2:对于剩余的N-M个语音片段中的第i个语音片段,计算所述第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将所述第i个语音片段归类到与所述第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;
    S3:当所述M个语音片段归类完成之后,根据所述M个类别中每个类别包括的语音片段重新计算所述M个类别的聚类中心,并更新所述M个类别的聚类中心,
    循环执行S2和S3,直到所述M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
  10. 根据权利要求6至9任一项所述的装置,其特征在于,所述分割单元包括:
    第三确定子单元,用于确定所述会议语音中的静音片段;
    去除子单元,用于去除所述会议语音中的静音片段;
    分割子单元,用于根据所述静音片段对去除所述静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;
    提取子单元,用于提取所述W个长语音片段中每一个长语音片段的声学特征;
    相对熵分析子单元,用于对所述W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;
    切分子单元,用于根据相对熵分析的结果对所述W个长语音片段进行切分,得到所述N个语音片段。
  11. 一种存储介质,其特征在于,所述存储介质包括存储的程序,其中,在所述程序运行时控制所述存储介质所在设备执行以下步骤:
    获取会议语音;
    将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;
    将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;
    确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;
    根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;
    根据所述M个发言人中每个发言人的发言内容生成会议记录。
  12. 根据权利要求11所述的存储介质,其特征在于,在所述程序运行时控制所述存储介质所在设备执行所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;
    向用户展示所述L个文本片段和发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述L个文本片段中每个文本片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  13. 根据权利要求11所述的存储介质,其特征在于,在所述程序运行时控制所述存储介质所在设备执行所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;
    向用户播放选择出的所述Z个语音片段并展示发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述Z个语音片段中每个语音片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音 片段对应的发言人。
  14. 根据权利要求11所述的存储介质,其特征在于,在所述程序运行时控制所述存储介质所在设备执行所述将所述N个语音片段进行聚类,包括:
    S1:从所述N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;
    S2:对于剩余的N-M个语音片段中的第i个语音片段,计算所述第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将所述第i个语音片段归类到与所述第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;
    S3:当所述M个语音片段归类完成之后,根据所述M个类别中每个类别包括的语音片段重新计算所述M个类别的聚类中心,并更新所述M个类别的聚类中心,
    循环执行S2和S3,直到所述M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
  15. 根据权利要求11至14任一项所述的存储介质,其特征在于,在所述程序运行时控制所述存储介质所在设备执行所述将所述会议语音进行分割,得到N个语音片段,包括:
    确定所述会议语音中的静音片段;
    去除所述会议语音中的静音片段;
    根据所述静音片段对去除所述静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;
    提取所述W个长语音片段中每一个长语音片段的声学特征;
    对所述W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;
    根据相对熵分析的结果对所述W个长语音片段进行切分,得到所述N个语音片段。
  16. 一种计算机设备,包括存储器和处理器,所述存储器用于存储包括程序指令的信息,所述处理器用于控制程序指令的执行,其特征在于,所述程序指令被处理器加载并执行时实现以下步骤:
    获取会议语音;
    将所述会议语音进行分割,得到N个语音片段,N为大于等于2的自然数;
    将所述N个语音片段进行聚类,得到M个类别的语音片段,M为大于等于2的自然数,M≤N,所述M个类别的语音片段分别与M个发言人具有一一对应关系;
    确定所述M个类别的语音片段中每个类别的语音片段对应的发言人;
    根据所述M个类别的语音片段确定所述M个发言人中每个发言人的发言内容;
    根据所述M个发言人中每个发言人的发言内容生成会议记录。
  17. 根据权利要求16所述的计算机设备,其特征在于,所述程序指令被处理器加载并执行时实现所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段转换成文本片段,得到L个文本片段,L为自然数,L≥M;
    向用户展示所述L个文本片段和发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述L个文本片段中每个文本片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  18. 根据权利要求16所述的计算机设备,其特征在于,所述程序指令被处理器加载并执行时实现所述确定所述M个类别的语音片段中每个类别的语音片段对应的发言人,包括:
    从所述M个类别的语音片段中每个类别的语音片段中各选择至少一个语音片段,得到Z个语音片段,Z为自然数,Z≥M;
    向用户播放选择出的所述Z个语音片段并展示发言人列表,所述发言人列表包括所述M个发言人中每个发言人的信息;
    接收匹配指令,所述匹配指令为所述用户发出的用于指示将所述Z个语音片段中每个语音片段与发言人进行匹配的指令;
    根据所述匹配指令确定所述M个类别的语音片段中每个类别的语音片段对应的发言人。
  19. 根据权利要求16所述的计算机设备,其特征在于,所述程序指令被处理器加载并执行时实现所述将所述N个语音片段进行聚类,包括:
    S1:从所述N个语音片段中随机选择M个语音片段,将选择的M个语音片段作为M个类别的聚类中心;
    S2:对于剩余的N-M个语音片段中的第i个语音片段,计算所述第i个语音片段与M个聚类中心中每个聚类中心之间的距离,并将所述第i个语音片段归类到与所述第i个语音片段距离最近的聚类中心对应的类别中,i依次取1至N-M之间的自然数;
    S3:当所述M个语音片段归类完成之后,根据所述M个类别中每个类别包括的语音片段重新计算所述M个类别的聚类中心,并更新所述M个类别的聚类中心,
    循环执行S2和S3,直到所述M个类别中每个类别的相邻两次聚类中心的距离在预设距离之内。
  20. 根据权利要求16至19任一项所述的计算机设备,其特征在于,所述程序指令被处理器加载并执行时实现所述将所述会议语音进行分割,得到N个语音片段,包括:
    确定所述会议语音中的静音片段;
    去除所述会议语音中的静音片段;
    根据所述静音片段对去除所述静音片段后的会议语音进行分割,得到W个长语音片段,W为大于等于2的自然数,W<N;
    提取所述W个长语音片段中每一个长语音片段的声学特征;
    对所述W个长语音片段中每一个长语音片段的声学特征进行相对熵分析;
    根据相对熵分析的结果对所述W个长语音片段进行切分,得到所述N个语音片段。
PCT/CN2019/118256 2019-01-16 2019-11-14 一种会议记录生成方法、装置、存储介质及计算机设备 Ceased WO2020147407A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910038460.6 2019-01-16
CN201910038460.6A CN109767757A (zh) 2019-01-16 2019-01-16 一种会议记录生成方法和装置

Publications (1)

Publication Number Publication Date
WO2020147407A1 true WO2020147407A1 (zh) 2020-07-23

Family

ID=66452786

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/118256 Ceased WO2020147407A1 (zh) 2019-01-16 2019-11-14 一种会议记录生成方法、装置、存储介质及计算机设备

Country Status (2)

Country Link
CN (1) CN109767757A (zh)
WO (1) WO2020147407A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4044179B1 (en) * 2020-09-27 2024-11-13 Comac Beijing Aircraft Technology Research Institute On-board information assisting system and method

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109767757A (zh) * 2019-01-16 2019-05-17 平安科技(深圳)有限公司 一种会议记录生成方法和装置
CN110265032A (zh) * 2019-06-05 2019-09-20 平安科技(深圳)有限公司 会议数据分析处理方法、装置、计算机设备和存储介质
CN110543559A (zh) * 2019-06-28 2019-12-06 谭浩 生成访谈报告的方法、计算机可读存储介质和终端设备
CN110335612A (zh) * 2019-07-11 2019-10-15 招商局金融科技有限公司 基于语音识别的会议记录生成方法、装置及存储介质
CN110675858A (zh) * 2019-08-29 2020-01-10 平安科技(深圳)有限公司 基于情绪识别的终端控制方法和装置
CN110930984A (zh) * 2019-12-04 2020-03-27 北京搜狗科技发展有限公司 一种语音处理方法、装置和电子设备
CN113963694B (zh) * 2020-07-20 2024-11-15 中移(苏州)软件技术有限公司 一种语音识别方法、语音识别装置、电子设备及存储介质
CN111933144A (zh) * 2020-10-09 2020-11-13 融智通科技(北京)股份有限公司 后创建声纹的会议语音转写方法、装置及存储介质
CN112562682A (zh) * 2020-12-02 2021-03-26 携程计算机技术(上海)有限公司 基于多人通话的身份识别方法、系统、设备及存储介质
CN114792522B (zh) * 2021-01-26 2026-01-02 阿里巴巴集团控股有限公司 音频信号处理、会议记录与呈现方法、设备、系统及介质
CN113674755B (zh) * 2021-08-19 2024-04-02 北京百度网讯科技有限公司 语音处理方法、装置、电子设备和介质
CN114708850B (zh) * 2022-02-24 2025-08-12 厦门快商通科技股份有限公司 一种交互式的语音分割与聚类方法、装置以及设备
DE202022101429U1 (de) 2022-03-17 2022-04-06 Waseem Ahmad Intelligentes System zur Erstellung von Sitzungsprotokollen mit Hilfe von künstlicher Intelligenz und maschinellem Lernen

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103530432A (zh) * 2013-09-24 2014-01-22 华南理工大学 一种具有语音提取功能的会议记录器及语音提取方法
CN103559882A (zh) * 2013-10-14 2014-02-05 华南理工大学 一种基于说话人分割的会议主持人语音提取方法
CN105810207A (zh) * 2014-12-30 2016-07-27 富泰华工业(深圳)有限公司 会议记录装置及其自动生成会议记录的方法
US20180075860A1 (en) * 2016-09-14 2018-03-15 Nuance Communications, Inc. Method for Microphone Selection and Multi-Talker Segmentation with Ambient Automated Speech Recognition (ASR)
CN109767757A (zh) * 2019-01-16 2019-05-17 平安科技(深圳)有限公司 一种会议记录生成方法和装置

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102185702A (zh) * 2011-04-27 2011-09-14 华东师范大学 智能会议系统终端控制器及其运作方法和应用方法
WO2016022588A1 (en) * 2014-08-04 2016-02-11 Flagler Llc Voice tallying system
CN106487757A (zh) * 2015-08-28 2017-03-08 华为技术有限公司 进行语音会议的方法、会议客户端和系统
CN107689225B (zh) * 2017-09-29 2019-11-19 福建实达电脑设备有限公司 一种自动生成会议记录的方法
CN108986826A (zh) * 2018-08-14 2018-12-11 中国平安人寿保险股份有限公司 自动生成会议记录的方法、电子装置及可读存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103530432A (zh) * 2013-09-24 2014-01-22 华南理工大学 一种具有语音提取功能的会议记录器及语音提取方法
CN103559882A (zh) * 2013-10-14 2014-02-05 华南理工大学 一种基于说话人分割的会议主持人语音提取方法
CN105810207A (zh) * 2014-12-30 2016-07-27 富泰华工业(深圳)有限公司 会议记录装置及其自动生成会议记录的方法
US20180075860A1 (en) * 2016-09-14 2018-03-15 Nuance Communications, Inc. Method for Microphone Selection and Multi-Talker Segmentation with Ambient Automated Speech Recognition (ASR)
CN109767757A (zh) * 2019-01-16 2019-05-17 平安科技(深圳)有限公司 一种会议记录生成方法和装置

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4044179B1 (en) * 2020-09-27 2024-11-13 Comac Beijing Aircraft Technology Research Institute On-board information assisting system and method

Also Published As

Publication number Publication date
CN109767757A (zh) 2019-05-17

Similar Documents

Publication Publication Date Title
WO2020147407A1 (zh) 一种会议记录生成方法、装置、存储介质及计算机设备
US11417343B2 (en) Automatic speaker identification in calls using multiple speaker-identification parameters
US8554562B2 (en) Method and system for speaker diarization
CN112309365B (zh) 语音合成模型的训练方法、装置、存储介质以及电子设备
CN110853646B (zh) 会议发言角色的区分方法、装置、设备及可读存储介质
CN109960743A (zh) 会议内容区分方法、装置、计算机设备及存储介质
US20200322399A1 (en) Automatic speaker identification in calls
CN110111808B (zh) 音频信号处理方法及相关产品
US20160284354A1 (en) Speech summarization program
CN108597525B (zh) 语音声纹建模方法及装置
CN108766418A (zh) 语音端点识别方法、装置及设备
TW202041037A (zh) 影片編輯方法及裝置
CN112148922A (zh) 会议记录方法、装置、数据处理设备及可读存储介质
WO2021159902A1 (zh) 年龄识别方法、装置、设备及计算机可读存储介质
WO2018113243A1 (zh) 语音分割的方法、装置、设备及计算机存储介质
CN110634472A (zh) 一种语音识别方法、服务器及计算机可读存储介质
US20230238002A1 (en) Signal processing device, signal processing method and program
WO2021237923A1 (zh) 智能配音方法、装置、计算机设备和存储介质
CN113889081A (zh) 语音识别方法、介质、装置和计算设备
WO2021072893A1 (zh) 一种声纹聚类方法、装置、处理设备以及计算机存储介质
CN109410956A (zh) 一种音频数据的对象识别方法、装置、设备及存储介质
WO2021196390A1 (zh) 声纹数据生成方法、装置、计算机装置及存储介质
CN111754982A (zh) 语音通话的噪声消除方法、装置、电子设备及存储介质
CN111785291A (zh) 语音分离方法和语音分离装置
CN110807370B (zh) 一种基于多模态的会议发言人身份无感确认方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19909703

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19909703

Country of ref document: EP

Kind code of ref document: A1