WO2024029152A1 - 区切り記号挿入装置及び音声認識システム - Google Patents

区切り記号挿入装置及び音声認識システム Download PDF

Info

Publication number
WO2024029152A1
WO2024029152A1 PCT/JP2023/017568 JP2023017568W WO2024029152A1 WO 2024029152 A1 WO2024029152 A1 WO 2024029152A1 JP 2023017568 W JP2023017568 W JP 2023017568W WO 2024029152 A1 WO2024029152 A1 WO 2024029152A1
Authority
WO
WIPO (PCT)
Prior art keywords
delimiter
time
interword
likelihood
word
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2023/017568
Other languages
English (en)
French (fr)
Inventor
謙吾 竹谷
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Docomo Inc
Original Assignee
NTT Docomo Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NTT Docomo Inc filed Critical NTT Docomo Inc
Priority to JP2024538828A priority Critical patent/JP7809817B2/ja
Publication of WO2024029152A1 publication Critical patent/WO2024029152A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/166Editing, e.g. inserting or deleting
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/04Segmentation; Word boundary detection

Definitions

  • the present invention relates to a delimiter insertion device and a speech recognition system.
  • Patent Document 1 discloses a technique in which punctuation marks are inserted into text by an engine trained using training data in the form of text with punctuation marks added based on statistics.
  • delimiters such as punctuation marks are inserted at positions appropriate for the sentence.
  • the insertion position of the delimiter may not be incorrect in the sentence, but may be different from the speaker's intention.
  • the present invention has been made in view of the above problems, and an object of the present invention is to insert a delimiter at the position intended by the speaker in text obtained by speech recognition processing of spoken voice. shall be.
  • a delimiter insertion device that inserts a delimiter that separates sentences after a word included in a text obtained by voice recognition of spoken voice. and an interword time acquisition unit that obtains an interword time that is the length of time until the next word is uttered in each word included in the uttered speech, and a delimiter insertion model and an interword time based on the delimiter insertion model and the interword time.
  • the delimiter insertion unit inserts a delimiter into a target text that is a text obtained by speech recognition of spoken speech, and the delimiter insertion model is configured to at least insert a delimiter removal sentence that does not include a delimiter.
  • the delimiter insertion model is configured to at least insert a delimiter removal sentence that does not include a delimiter.
  • This model is generated by machine learning using training data, and is based on delimiter inference information that is obtained by inputting the target text as a delimiter removed sentence into a delimiter insertion model and adjusted according to the interword time. and a delimiter insertion unit for inserting a delimiter into the target text.
  • the interword time in the uttered speech is obtained.
  • the interword time reflects the speaker's intention at the time of utterance.
  • delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
  • the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
  • FIG. 1 is a block diagram showing the functional configuration of a delimiter insertion device according to the present embodiment.
  • FIG. 2 is a hardware block diagram of a delimiter insertion device. It is a figure explaining the problem solved by the delimiter insertion device of this embodiment. It is a figure which shows the acquisition process of target text.
  • FIG. 3 is a diagram showing a first example of learning data used for machine learning of a delimiter insertion model.
  • FIG. 3 is a diagram showing a first example of the configuration of a delimiter insertion model.
  • FIG. 7 is a diagram illustrating an example of adjustment rule information that is referred to in order to adjust delimiter guess information based on interword time.
  • FIG. 7 is a diagram illustrating an example of adjustment processing of delimiter guess information.
  • FIG. 3 is a diagram illustrating an example of inter-word time correction processing.
  • FIG. 7 is a diagram showing a second example of learning data used for machine learning of the delimiter insertion model.
  • FIG. 7 is a diagram illustrating a second example of the configuration of a delimiter insertion model.
  • FIG. 7 is a diagram showing a third example of learning data used for machine learning of the delimiter insertion model.
  • FIG. 1 is a functional block diagram showing an example of the configuration of a speech recognition system according to the present embodiment.
  • 3 is a flowchart showing processing details of a delimiter insertion method in the delimiter insertion device.
  • FIG. 2 is a diagram showing the configuration of a delimiter insertion program.
  • FIG. 1 is a diagram showing the functional configuration of a delimiter insertion device according to this embodiment.
  • the delimiter insertion device of this embodiment is a device that inserts a delimiter for delimiting a sentence after a word included in a text obtained by voice recognition of spoken voice.
  • the delimiter insertion device 10 inserts delimiters such as a period, a comma, and a question mark, which are inserted into an English sentence, into a text consisting of an English sentence.
  • delimiters such as a period, a comma, and a question mark
  • the delimiter insertion device 10 may be a device that inserts delimiters such as periods and commas into sentences in Japanese, or a device that inserts delimiters in that language into sentences in other languages. good.
  • the delimiter insertion device 10 functionally includes a text acquisition section 11, an interword time acquisition section 12, a delimiter insertion section 13, and an output section 14.
  • Each of these functional units 11 to 14 may be configured in one device, or may be configured in a distributed manner in a plurality of devices.
  • each functional block may be realized using one physically or logically coupled device, or may be realized using two or more physically or logically separated devices directly or indirectly (e.g. , wired, wireless, etc.) and may be realized using a plurality of these devices.
  • the functional block may be realized by combining software with the one device or the plurality of devices.
  • Functions include judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, These include, but are not limited to, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assigning. I can't.
  • a functional block (configuration unit) that performs transmission is called a transmitting unit or a transmitter. In either case, as described above, the implementation method is not particularly limited.
  • the delimiter insertion device 10 in one embodiment of the present invention may function as a computer.
  • FIG. 2 is a diagram showing an example of the hardware configuration of the delimiter insertion device 10 according to this embodiment.
  • the delimiter insertion device 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
  • the word “apparatus” can be read as a circuit, a device, a unit, etc.
  • the hardware configuration of the delimiter insertion device 10 may be configured to include one or more of the devices shown in the figure, or may be configured without including some of the devices.
  • Each function in the delimiter insertion device 10 is achieved by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, so that the processor 1001 performs calculations, and the communication by the communication device 1004 and the memory 1002 and This is achieved by controlling reading and/or writing of data in the storage 1003.
  • the processor 1001 for example, operates an operating system to control the entire computer.
  • the processor 1001 may be configured with a central processing unit (CPU) that includes interfaces with peripheral devices, a control device, an arithmetic device, registers, and the like.
  • CPU central processing unit
  • each of the functional units 11 to 14 shown in FIG. 1 may be implemented by the processor 1001.
  • the processor 1001 reads programs (program codes), software modules, and data from the storage 1003 and/or the communication device 1004 to the memory 1002, and executes various processes in accordance with these.
  • the program a program that causes a computer to execute at least part of the operations described in the above embodiments is used.
  • each of the functional units 11 to 15 of the delimiter insertion device 10 may be realized by a control program stored in the memory 1002 and operated on the processor 1001.
  • Processor 1001 may be implemented with one or more chips. Note that the program may be transmitted from a network via a telecommunications line.
  • the memory 1002 is a computer-readable recording medium, and includes at least one of ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. may be done.
  • Memory 1002 may be called a register, cache, main memory, or the like.
  • the memory 1002 can store executable programs (program codes), software modules, and the like to implement the pseudo data generation method and sentence generation method according to an embodiment of the present invention.
  • the storage 1003 is a computer-readable recording medium, such as an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, or a magneto-optical disk (for example, a compact disk, a digital versatile disk, or a Blu-ray disk). (registered trademark) disk), smart card, flash memory (eg, card, stick, key drive), floppy disk, magnetic strip, etc.
  • Storage 1003 may also be called an auxiliary storage device.
  • the storage medium mentioned above may be, for example, a database including memory 1002 and/or storage 1003, a server, or other suitable medium.
  • the communication device 1004 is hardware (transmission/reception device) for communicating between computers via a wired and/or wireless network, and is also referred to as a network device, network controller, network card, communication module, etc., for example.
  • the input device 1005 is an input device (eg, keyboard, mouse, microphone, switch, button, sensor, etc.) that accepts input from the outside.
  • the output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that performs output to the outside. Note that the input device 1005 and the output device 1006 may have an integrated configuration (for example, a touch panel).
  • each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information.
  • the bus 1007 may be configured as a single bus or may be configured as different buses between devices.
  • the delimiter insertion device 10 also uses hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), and a field programmable gate array (FPGA). A part or all of each functional block may be realized by the hardware. For example, processor 1001 may be implemented with at least one of these hardware.
  • DSP digital signal processor
  • ASIC application specific integrated circuit
  • PLD programmable logic device
  • FPGA field programmable gate array
  • the problem solved by the delimiter insertion device 10 of this embodiment will be explained with reference to FIG. 3.
  • the utterances sp11 and sp21 shown in FIG. 3 are uttered by speakers with different intentions, respectively.
  • the speech recognition results sr11 and sr21 obtained by speech recognition of the uttered speech sp11 and sp21 are the same text "I know it's been there forever".
  • the interword time until the next word after the word "know” is uttered is 1.2 seconds.
  • the interword time until the next word after the word "know” is uttered is 0.1 seconds.
  • the delimiter is inserted into the text obtained as a speech recognition result at a position that is appropriate for a sentence.
  • the speech recognition results sr12 and sr22 obtained by inserting the delimiter are the same even though the original speech sounds are different.
  • the speech recognition result sr22 has a period after the word "forever”, as intended by the speaker of the uttered speech sp21.
  • the speech recognition result sr12 has a period after the word "forever.” The position of this period is different from the position intended by the speaker.
  • the delimiter insertion device 10 of this embodiment inserts a delimiter using the interword time in the uttered speech, so the word "know” and the word A period is inserted after each word "forever” to obtain the speech recognition result sr13.
  • the delimiter insertion device 10 inserts a period after the word "forever” into the speech recognition result sr21 obtained based on the uttered speech sp21 to obtain a speech recognition result sr23.
  • the speech recognition results sr13 and sr23 have delimiters at the positions intended by the speakers of the uttered speech sp11 and sp21.
  • the text acquisition unit 11 acquires target text, which is the text into which a delimiter is to be inserted.
  • the target text is text obtained by voice recognition of spoken voice.
  • the interword time acquisition unit 12 acquires the interword time, which is the length of time until the next word is uttered in each word included in the uttered voice.
  • FIG. 4 is a diagram showing an example of acquiring the target text and interword time.
  • the text acquisition unit 11 acquires the target text tx1 "I know it's been there forever" based on the speech sp3.
  • the text acquisition unit 11 may acquire the target text by performing voice recognition of the uttered voice using a well-known voice recognition processing technique and other techniques.
  • the interword time acquisition unit 12 acquires the length of time until the next word is uttered in each word included in the target text as the interword time it.
  • the inter-word time acquisition unit 12 may acquire the silent time during speech recognition of the uttered speech sp3 as the inter-word time it.
  • the silent time is, for example, a time when the volume is less than a predetermined level.
  • the interword time acquisition unit 12 sequentially acquires voice recognition results every time voice recognition is performed from a voice recognition engine used for voice recognition of uttered speech, and sets the time interval for acquiring and updating the voice recognition results between words.
  • the update time interval may be obtained as the interword time it by regarding it as a pseudo silent time between the occurrences of .
  • the delimiter insertion unit 13 inserts delimiters into the target text based on the delimiter insertion model and the interword time. Specifically, the delimiter insertion unit 13 inserts a delimiter into the target text based on delimiter guess information obtained by inputting the target text into a delimiter insertion model.
  • the delimiter guess information includes information adjusted according to interword time.
  • the delimiter insertion model receives at least a delimiter removed sentence, which is a sentence that does not include a delimiter, as input, and outputs delimiter guess information indicating the delimiter to be inserted after each word included in the delimiter removed sentence. Further, the delimiter insertion model is generated by machine learning using learning data including a pair of a delimiter-removed sentence and a delimiter-included sentence, which is a sentence including the delimiter.
  • FIG. 5 is a diagram showing a first example of learning data used for machine learning of the delimiter insertion model.
  • FIG. 6 is a diagram showing a first example of the configuration of a delimiter insertion model.
  • the learning data td1 which is an example of learning data used for machine learning of the delimiter insertion model md1, consists of a pair of a delimiter-removed sentence id1 and a delimiter-added sentence od1.
  • the delimiter-containing sentence od1 includes a word string making up the sentence and a delimiter label that is a label indicating a delimiter inserted after each word.
  • Labels indicating delimiters are schematically illustrated in FIG. 5 and the like as follows. ⁇ O>...No delimiter ⁇ P>...Period, full stop ⁇ C>...Comma, comma ⁇ Q>...Question mark, question mark
  • the delimiter removed sentence id1 is a sentence with the delimiter label removed from the delimiter-containing sentence od1. There may be.
  • the delimiter removed sentence id1 is input to the delimiter insertion model md1 in the learning process, and the output obtained from the delimiter insertion model md1 is combined with the delimiter-containing sentence od1, which is the training data. Based on the error, the weights, parameters, etc. that make up the delimiter insertion model md1 are updated.
  • the trained delimiter insertion model md1 outputs delimiter guess information dp1 in response to the input of the delimiter removed sentence sd1.
  • the delimiter insertion model md1 may be a model that includes a neural network. More specifically, the delimiter insertion model md1 may be configured as a sequence labeling model that solves a sequence labeling task of predicting a delimiter to be inserted after each word included in an input sentence.
  • the delimiter insertion model md1 which is a model that includes a trained neural network, can be read or referenced by a computer, and can be regarded as a program that causes the computer to perform a predetermined process and realize a predetermined function.
  • the trained delimiter insertion model md1 of this embodiment is used in a computer equipped with a CPU and memory. Specifically, the CPU of the computer assigns learned weights corresponding to each layer to the input data input to the input layer of the neural network according to instructions from the learned delimiter insertion model md1 stored in memory. It operates to perform calculations based on coefficients (parameters), response functions, etc., and output the results (probabilities) from the output layer.
  • the delimiter guess information dp1 includes the symbol insertion likelihood, which is the likelihood of various delimiters that can be inserted after each word included in the delimiter removed sentence sd1, and the fact that no delimiter is inserted after each word. Contains the unsymbol likelihood, which is the likelihood for . Then, based on the maximum likelihood of the symbol insertion likelihood and the no symbol likelihood, insert one of a plurality of types of delimiters after each word, or insert no delimiter. This is determined (labeled).
  • the delimiter guess information dp1 illustrated in FIG. 6 includes the symbol insertion likelihood and symbol-free likelihood of each delimiter regarding the word "I", as described below. ⁇ O>: 90%, ⁇ C>: 5%, ⁇ P>: 2%, ⁇ Q>: 3% Therefore, since the symbol-less likelihood of not inserting a delimiter (label ⁇ O>) is the maximum, the word "I" is labeled with the label ⁇ O> without a delimiter.
  • the delimiter guess information dp1 includes the likelihood of symbol insertion and the likelihood of no symbol for each delimiter regarding the word "know”, as shown below.
  • the delimiter insertion unit 13 may adjust the delimiter inference information output from the delimiter insertion model based on the interword time, as an example of adjusting the delimiter inference information based on the interword time. Specifically, the delimiter insertion unit 13 may adjust the likelihood of symbol insertion and the likelihood of no symbol included in the delimiter estimation information based on the interword time.
  • the delimiter insertion unit 13 adjusts to increase the likelihood of no symbol for one word included in the target text when the interword time in one of the words is the first time, or / and adjusting to lower the symbol insertion likelihood of at least one kind of delimiter of the one word out of the plurality of kinds of delimiters, and the first one whose interword time in the one word is longer than the first time. 2, adjustment is made to increase the likelihood of symbol insertion of at least one of the plurality of types of punctuation marks for the one word, and/or no symbol for the one word. It may be adjusted to lower the likelihood.
  • the delimiter insertion unit 13 inserts one of the plurality of delimiters after the one word based on the maximum likelihood of the adjusted symbol insertion likelihood and the no symbol likelihood. or leave the delimiter uninserted.
  • the delimiter insertion unit 13 may, for example, adjust the likelihood by referring to adjustment rule information.
  • FIG. 7 is a diagram illustrating an example of adjustment rule information that is referred to in order to adjust delimiter guess information based on interword time.
  • the adjustment rule information may be stored in a storage device that is accessible to the delimiter insertion section 13, or may be provided as a table inside the delimiter insertion section 13.
  • the adjustment rule information is such that, for each range of interword time, an interword time category indicating the length of the interword time and likelihood adjustment information indicating the contents of the likelihood adjustment are associated. It is information. For example, if the interword time (x) is 0.1 seconds or less, the interword time category is "none", and “increase the no symbol likelihood by 50%" is performed as a likelihood adjustment. This is stipulated as an adjustment rule. Also, for example, if the interword time (x) is longer than 0.5 seconds and less than 1.0 seconds, the interword time category is "medium” and the likelihood of symbol insertion is increased by 50%. It is stipulated as an adjustment rule that this is carried out as a likelihood adjustment.
  • FIG. 8 is a diagram illustrating an example of adjustment processing of delimiter guess information.
  • the delimiter guess information dp21 shown in FIG. 8 shows a part of the delimiter guess information before the adjustment process that is output from the delimiter insertion model md1.
  • the delimiter guess information dp21 includes a symbol insertion likelihood lh21 and a symbol-free likelihood of each delimiter regarding the word "know.” According to the pre-adjustment delimiter guess information dp21, the likelihood of not inserting a delimiter (label ⁇ O>) is the highest, so the label ⁇ O> without a delimiter for the word "know" is O> is labeled.
  • the delimiter insertion unit 13 adjusts each likelihood of the delimiter guess information dp21 based on the interword time it2 of the word "know". Specifically, since the interword time it2 of the word "know" acquired by the interword time acquisition unit 12 is 0.9 seconds, the delimiter insertion unit 13 refers to the adjustment rule information (FIG. 7). Then, the likelihood adjustment information "Increase the likelihood of symbol insertion for periods, commas, and question marks by 50%" associated with the interword time of 0.9 seconds is obtained, and the symbol insertion likelihood for commas, periods, and question marks lh21 is obtained. is adjusted according to the obtained likelihood adjustment information.
  • the delimiter insertion unit 13 increases each value of the symbol insertion likelihood lh21 by 50% to obtain adjusted delimiter guess information dp22.
  • the symbol insertion likelihood of a period (label ⁇ P>) is the maximum, so the delimiter insertion unit 13 inserts the label ⁇ P> of the delimiter "period" into the word "know". Label. Then, the delimiter insertion unit 13 inserts a period after the word "know" included in the target text based on the labeled label ⁇ P>.
  • the longer the interword time of one word included in the target text the higher the likelihood of symbol insertion and/or the lower the likelihood of no symbol, so that the longer the interword time of one word
  • the shorter the interword time the lower the likelihood of symbol insertion and/or the higher the likelihood of no symbol. reflected in the likelihood.
  • the delimiter is inserted or not inserted, so the delimiter is inserted at an appropriate position according to the speaker's intention. It is possible to obtain the text.
  • the delimiter insertion unit 13 may correct the interword time for use in adjusting the delimiter guess information according to predetermined conditions.
  • FIG. 9 is a diagram illustrating an example of inter-word time correction processing.
  • the delimiter insertion unit 13 inserts the interword time of each of all the words. It may be corrected to shorten it to a given degree. Then, the delimiter insertion unit 13 may adjust the likelihood of symbol insertion and/or the likelihood of no symbol based on the corrected interword time, which is the corrected interword time.
  • the target text tx31 shown in FIG. 9 includes an interword time it31 before correction.
  • the interword time of more than half of all the words included in the target text is equal to or greater than the interword time corresponding to the interword time category "small"
  • the content of the correction process is set in advance to reduce the interword time by 0.5 seconds.
  • the delimiter insertion unit 13 determines that the interword time of all words among the interword time it31 of words included in the target text tx31 is equal to or greater than the interword time corresponding to the interword time category "small". determine something. Then, as shown in the target text tx32, the delimiter insertion unit 13 subtracts 0.5 seconds from each interword time it31 to obtain a corrected interword time it32, which is the corrected interword time. The delimiter insertion unit 13 adjusts the likelihood of symbol insertion and/or the likelihood of no symbol of each word included in the target text tx32 based on the corrected interword time it32.
  • the delimiter may be inserted excessively in a position that the speaker did not intend.
  • the interword time is longer than a predetermined value, a symbol is created based on the corrected interword time that is corrected to shorten the interword time. Since the insertion likelihood and/or the symbol-less likelihood are adjusted, it is possible to obtain a text in which the delimiter is inserted at an appropriate position that suitably reflects the speaker's intention.
  • FIG. 10 is a diagram showing a second example of learning data used for machine learning of the delimiter insertion model.
  • FIG. 11 is a diagram showing a second example of the configuration of the delimiter insertion model.
  • the learning data td2 which is an example of learning data used for machine learning of the delimiter insertion model md2, consists of a pair of a delimiter-removed sentence id2 and a delimiter-added sentence od2.
  • the delimiter-containing sentence od2 like the delimiter-containing sentence od1 described with reference to FIG. 5, includes a word string composing the sentence and a delimiter label indicating the delimiter to be inserted after each word.
  • the delimiter-removed sentence id2 includes a sentence obtained by removing the delimiter label from the delimiter-containing sentence od2, and interword time it4 associated with each word constituting the sentence.
  • the delimiter removed sentence id2 is input to the delimiter insertion model md2 in the learning process, and the output obtained from the delimiter insertion model md2 is combined with the delimiter-containing sentence od2, which is the training data. Based on the error, the weights, parameters, etc. that make up the delimiter insertion model md2 are updated.
  • the delimiter insertion model md2 may be a model including a neural network. More specifically, the delimiter insertion model md2 may be configured as a sequence labeling model that solves a sequence labeling task of predicting a delimiter to be inserted after each word included in an input sentence.
  • the trained delimiter insertion model md2 outputs delimiter guess information dp2 in response to the input of the delimiter removed sentence sd2.
  • the delimiter removed sentence sd2 includes an interword time it5 associated with each word forming the delimiter removed sentence sd2.
  • the delimiter guess information dp2 includes the likelihood of symbol insertion and the likelihood of no symbol for each word included in the delimiter removed sentence sd2.
  • the delimiter insertion model md2 which is a model that includes a trained neural network, can be read or referenced by a computer, and can be regarded as a program that causes the computer to perform a predetermined process and realize a predetermined function.
  • the trained delimiter insertion model md2 of this embodiment is used in a computer equipped with a CPU and memory. Specifically, the CPU of the computer assigns learned weights corresponding to each layer to the input data input to the input layer of the neural network according to instructions from the learned delimiter insertion model md2 stored in memory. It operates to perform calculations based on coefficients (parameters), response functions, etc., and output the results (probabilities) from the output layer.
  • the symbol insertion likelihood and symbol-free likelihood included in the delimiter guess information dp1 output from the delimiter insertion model md1 are adjusted based on the interword time.
  • the delimiter removed sentence id2 including the interword time it4 is used as input as learning data in machine learning, and in response to the input of the delimiter removed sentence sd2 including the interword time it5. Since the delimiter guess information dp2 is output, the symbol insertion likelihood and the no symbol likelihood included in the delimiter guess information dp2 are values adjusted according to the interword time by calculations in the delimiter insertion model md2. It is.
  • the likelihood of symbol insertion and the likelihood of no symbol for each delimiter regarding the word "know” are calculated as follows. ⁇ O>: 10%, ⁇ C>: 20%, ⁇ P>: 70%, ⁇ Q>: 0%
  • the maximum likelihood of no delimiter was calculated by the delimiter insertion model md2. For each likelihood, the symbol insertion likelihood of inserting a period (label ⁇ P>) is the largest.
  • the delimiter insertion model md2 is generated by machine learning using delimiter-removed sentences with interword times associated with each word as training data, and delimiter-removed sentences with interword times associated with each word are generated by machine learning. Since it is input to the delimiter insertion model md2, by inputting the target text in which interword time is associated with each word to the delimiter insertion model, a separate It is possible to obtain delimiter guess information adjusted by the interword time without performing adjustment processing. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to easily obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
  • FIG. 12 is a diagram showing a third example of learning data used for machine learning of the delimiter insertion model.
  • the learning data td3 consists of a pair of a delimiter-removed sentence id3 and a delimiter-added sentence od3.
  • the delimiter-containing sentence od3 like the delimiter-containing sentences od1 and od2 described with reference to FIGS. 5 and 10, has a delimiter label indicating the word string that makes up the sentence and the delimiter inserted after each word. including.
  • the delimiter-removed sentence id3 includes a sentence from which the delimiter label has been removed from the delimiter-containing sentence od3, interword times associated with each word constituting the sentence, and situation information st3.
  • the situation information st3 is information indicating the situation when the uttered voice is uttered, and may be added as a tag to the delimiter removed sentence id3.
  • the situation information st3 may have the following variations depending on the situation when the uttered voice is uttered. Meeting: ⁇ Meeting> Lecture: ⁇ Lecture> Chatting: ⁇ Chatting>
  • the delimiter removed sentence id3 is input to the delimiter insertion model in the learning process, and the output obtained from the delimiter insertion model and the delimiter-containing sentence that is the training data are combined. Based on the error with od3, the weights, parameters, etc. that make up the delimiter insertion model are updated.
  • the delimiter insertion model which has been trained by machine learning using learning data td3, generates delimiter inference information in response to the input of a delimiter removed sentence (target text) that includes situation information and interword time associated with each word. Output.
  • the output delimiter estimation information includes the likelihood of symbol insertion and the likelihood of no symbol for each word included in the delimiter removed sentence, as well as a delimiter based on the maximum likelihood or a label of no delimiter inserted.
  • the delimiter insertion unit 13 inserts a delimiter according to the label after each word based on the delimiter guess information obtained by inputting the target text associated with the situation information into the delimiter insertion model. or decide not to insert the delimiter.
  • a delimiter insertion model is generated by machine learning using delimiter removal sentences associated with situation information as learning data, and delimiter removal sentences including situation information are used as input to the delimiter insertion model. Therefore, by inputting the target text associated with situation information into the delimiter insertion model, it is possible to obtain delimiter estimation information that takes into account the utterance tendency according to the situation when the uttered voice is uttered. Become.
  • the output unit 14 outputs the delimiter inserted text, which is the target text into which the delimiter has been inserted by the delimiter insertion unit 13.
  • the mode of output is not limited, and the output unit 14 may display the delimiter insertion text on a predetermined display means, store the delimiter insertion text in a predetermined storage means, or display the delimiter insertion text on a predetermined device. You may also send delimited text to .
  • FIG. 13 is a functional block diagram showing an example of the configuration of the speech recognition system of this embodiment.
  • the speech recognition system 20 includes a delimiter insertion device 10, and includes a speech recognition result acquisition section 21 and a speech recognition result output section 22.
  • the speech recognition result acquisition unit 21 obtains the text resulting from speech recognition by the speech recognition engine and the interword time of each word in the text as a first speech recognition result.
  • the text acquisition unit 11 of the delimiter insertion device 10 acquires the text included in the first speech recognition result as the target text.
  • the interword time acquisition unit 12 of the delimiter insertion device 10 acquires the interword time included in the first speech recognition result.
  • the delimiter insertion unit 13 of the delimiter insertion device 10 performs a process of inserting a delimiter using the text included in the first speech recognition result as the target text.
  • the speech recognition result output unit 22 outputs the target text into which the delimiter has been inserted by the delimiter insertion unit 13 as a second speech recognition result.
  • the speech recognition system 20 based on the first speech recognition result obtained from a speech recognition engine that does not take interword time into consideration in speech recognition processing, the text included in the first speech recognition result is set as the target text, Since it is possible to obtain delimiter estimation information adjusted according to the interword time included in the first speech recognition result, the second speech recognition that includes the target text with a delimiter inserted at a position that matches the speaker's intention can be performed. It becomes possible to obtain results.
  • FIG. 14 is a flowchart showing the processing details of the delimiter insertion method in the delimiter insertion device 10.
  • step S1 the text acquisition unit 11 acquires target text, which is the text obtained by voice recognition of the uttered voice and is the text into which a delimiter is to be inserted.
  • step S2 the interword time acquisition unit 12 acquires the interword time of each word in the uttered speech.
  • step S3 the delimiter insertion unit 13 inputs the target text into the delimiter insertion model.
  • step S4 the delimiter insertion unit 13 acquires delimiter guess information.
  • the delimiter estimation information is based on the symbol insertion likelihood and no symbol likelihood of each word adjusted according to the interword time, and determines whether any delimiter is inserted or no delimiter is inserted for each word. Contains a label indicating that.
  • step S5 the delimiter insertion unit 13 inserts a delimiter into the target text based on the delimiter estimation information.
  • step S6 the output unit 14 outputs the delimiter inserted text, which is the target text in which the delimiter has been inserted.
  • FIG. 15 is a diagram showing the configuration of the delimiter insertion program.
  • the delimiter insertion program P1 includes a main module m10 that centrally controls the delimiter insertion process in the delimiter insertion device 10, a text acquisition module m11, an interword time acquisition module m12, a delimiter insertion module m13, and an output module m14. It consists of Each of the modules m11 to m14 implements the functions of the text acquisition section 11, the interword time acquisition section 12, the delimiter insertion section 13, and the output section 14.
  • the delimiter insertion program P1 may be transmitted via a transmission medium such as a communication line, or may be stored in a recording medium M1 as shown in FIG. .
  • the interword time in the uttered voice is acquired.
  • the interword time reflects the speaker's intention at the time of utterance.
  • delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
  • the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention.
  • by inserting a delimiter into the target text based on the delimiter estimation information it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
  • a delimiter insertion device is a delimiter insertion device that inserts a delimiter for delimiting a sentence after a word included in a text obtained by voice recognition of an uttered voice, an interword time acquisition unit that acquires an interword time that is the length of time until the next word in each word included in the utterance is uttered;
  • a delimiter insertion unit that inserts a delimiter into a target text that is a text obtained by voice recognition of spoken speech, the delimiter insertion model including at least a delimiter removed sentence that does not include the delimiter.
  • output delimiter guess information indicating a delimiter to be inserted after each word included in the delimiter-removed sentence, and pair the delimiter-removed sentence with a delimiter-containing sentence that is a sentence including the delimiter.
  • the model is generated by machine learning using learning data containing the delimiter, and is obtained by inputting the target text as the delimiter removed sentence to the delimiter insertion model, and is adjusted according to the interword time. and a delimiter insertion unit that inserts a delimiter into the target text based on the delimiter estimation information.
  • the interword time in the uttered speech is obtained.
  • the interword time reflects the speaker's intention at the time of utterance.
  • delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
  • the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
  • the delimiter insertion unit may insert the delimiter estimation information output from the delimiter insertion model into the delimiter estimation information outputted from the delimiter insertion model. Adjustments may also be made based on time.
  • the delimiter estimation information includes a plurality of types of delimiters after each word included in the delimiter removed sentence.
  • the delimiter insertion section includes a symbol insertion likelihood, which is the likelihood of inserting each word, and a symbol-free likelihood, which is the likelihood of not inserting a delimiter after each word. If the interword time in one of the words included in the text is a first time, the one word is adjusted to increase the likelihood of no symbol, and/or the plurality of types are adjusted.
  • the symbol insertion likelihood of at least one of the delimiters of the one word to reduce the symbol insertion likelihood, and the interword time in the one word is longer than the first time; , the likelihood of at least one of a plurality of types of punctuation marks for the one word is adjusted to increase the likelihood of the symbol being inserted, and/or the likelihood of the one word not having the symbol is adjusted. and insert one of a plurality of types of delimiters after the one word based on the maximum likelihood of the symbol insertion likelihood and the no symbol likelihood. Alternatively, the delimiter may not be inserted.
  • the delimiter is inserted or not inserted, so the delimiter is inserted at an appropriate position according to the speaker's intention. It is possible to obtain the text.
  • the delimiter inserting section includes a delimiter inserting unit configured to insert a delimiter between words of one or more of all words included in the target text.
  • a delimiter inserting unit configured to insert a delimiter between words of one or more of all words included in the target text.
  • the delimiter may be inserted excessively in a position that the speaker did not intend.
  • the symbol insertion likelihood and/or since the likelihood without symbols is adjusted, it is possible to obtain a text in which delimiters are inserted at appropriate positions that suitably reflect the speaker's intention.
  • the delimiter insertion model further includes an interword time of each word included in the delimiter removed sentence as an input.
  • the delimiter is generated by machine learning using learning data consisting of a pair of the delimiter-removed sentence and the delimiter-included sentence whose interword time is associated with each word, and the delimiter is adjusted by the interword time.
  • the delimiter insertion unit may output guess information and input the target text in which each word is associated with the interword time into the delimiter insertion model.
  • a delimiter insertion model is generated by machine learning in which delimiter removal sentences with interword time associated with each word are used as training data, and delimiter removal including interword time of each word is generated by machine learning. Since a sentence is taken as the input of the delimiter insertion model, by inputting the target text with an associated interword time for each word into the delimiter insertion model, a separate It becomes possible to obtain delimiter guess information adjusted by the interword time without performing adjustment processing. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to easily obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
  • the delimiter estimation information includes a plurality of types of delimiters after each word included in the delimiter removed sentence. It is also possible to include a symbol insertion likelihood, which is the likelihood of inserting each word, and a symbol-free likelihood, which is the likelihood of not inserting a delimiter after each word.
  • delimiter guessing information including the symbol insertion likelihood and symbol-absence likelihood adjusted by the interword time.
  • the delimiter insertion model is based on a situation when the speech sound corresponding to the delimiter removed sentence is uttered.
  • the machine further includes as input situation information indicating that the interword time is associated with each word, and the machine uses learning data consisting of a pair of the delimiter-removed sentence and the delimiter-included sentence, to which the interword time is associated with each word and the situation information is associated.
  • the delimiter insertion unit may input the target text associated with the situation information to the delimiter insertion model, which is generated by learning.
  • a delimiter insertion model is generated by machine learning using delimiter removal sentences associated with situation information as learning data, and delimiter removal sentences including situation information are input to the delimiter insertion model. Therefore, by inputting the target text associated with situation information into the delimiter insertion model, it is possible to obtain delimiter estimation information that takes into account the tendency of speech according to the situation when the speech is uttered. becomes possible.
  • the speech recognition system includes the delimiter insertion device according to any one of the first to seventh aspects, and includes a text as a result of speech recognition by the speech recognition engine and each word in the text.
  • the interword time acquisition unit of the delimiter insertion device includes a speech recognition result acquisition unit that obtains the interword time as a first speech recognition result, and a speech recognition result output unit, and the interword time acquisition unit of the delimiter insertion device Obtaining the interword time included in the recognition result, the delimiter insertion unit of the delimiter insertion device inserts a delimiter using the text included in the first speech recognition result as the target text, and
  • the speech recognition result output section may output the target text into which the delimiter has been inserted by the delimiter insertion section as the second speech recognition result.
  • the text included in the first speech recognition result is set as the target text. Since it is possible to obtain delimiter estimation information adjusted according to the interword time included in the first speech recognition result, the second speech recognition result includes the target text with the delimiter inserted at a position that matches the speaker's intention. It becomes possible to obtain.
  • the notification of information may include physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, It may be implemented using broadcast information (MIB (Master Information Block), SIB (System Information Block)), other signals, or a combination thereof.
  • RRC signaling may be called an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
  • LTE Long Term Evolution
  • LTE-A Long Term Evolution-Advanced
  • SUPER 3G IMT-Advanced
  • 4G 5G
  • FRA Full Radio Access
  • W-CDMA Wideband Code Division Multiple Access
  • GSM registered trademark
  • CDMA2000 Code Division Multiple Access 2000
  • UMB Universal Mobile Broadband
  • IEEE 802.11 Wi-Fi
  • IEEE 802.16 WiMAX
  • IEEE 802.20 UWB (Ultra-WideBand)
  • the present invention may be applied to systems utilizing Bluetooth (registered trademark), other suitable systems, and/or next-generation systems extended based thereon.
  • a combination of a plurality of systems may be applied (for example, a combination of at least one of LTE and LTE-A and 5G).
  • the specific operations performed by the base station in this disclosure may be performed by its upper node.
  • various operations performed for communication with a terminal are performed by the base station and other network nodes other than the base station (e.g., MME or It is clear that this could be done by at least one of the following: (conceivable, but not limited to) S-GW, etc.).
  • MME mobile phone
  • S-GW network node
  • Information can be output from the upper layer (or lower layer) to the lower layer (or upper layer). It may be input/output via multiple network nodes.
  • the input/output information may be stored in a specific location (for example, memory) or may be managed in a management table. Information etc. to be input/output may be overwritten, updated, or additionally written. The output information etc. may be deleted. The input information etc. may be transmitted to other devices.
  • Judgment may be made using a value expressed by 1 bit (0 or 1), a truth value (Boolean: true or false), or a comparison of numerical values (for example, a predetermined value). (comparison with a value).
  • notification of prescribed information is not limited to being done explicitly, but may also be done implicitly (for example, not notifying the prescribed information). Good too.
  • Software includes instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, whether referred to as software, firmware, middleware, microcode, hardware description language, or by any other name. , should be broadly construed to mean an application, software application, software package, routine, subroutine, object, executable, thread of execution, procedure, function, etc.
  • software, instructions, etc. may be sent and received via a transmission medium.
  • a transmission medium For example, if the software uses wired technologies such as coaxial cable, fiber optic cable, twisted pair and digital subscriber line (DSL) and/or wireless technologies such as infrared, radio and microwave to When transmitted from a remote source, these wired and/or wireless technologies are included within the definition of transmission medium.
  • wired technologies such as coaxial cable, fiber optic cable, twisted pair and digital subscriber line (DSL) and/or wireless technologies such as infrared, radio and microwave
  • data, instructions, commands, information, signals, bits, symbols, chips, etc. which may be referred to throughout the above description, may refer to voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, light fields or photons, or any of these. It may also be represented by a combination of
  • system and “network” are used interchangeably.
  • radio resources may be indicated by an index.
  • determining may encompass a wide variety of operations.
  • “Judgment” and “decision” include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, search, and inquiry. (e.g., searching in a table, database, or other data structure), and regarding an ascertaining as a “judgment” or “decision.”
  • judgment and “decision” refer to receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and access.
  • (accessing) may include considering something as a “judgment” or “decision.”
  • judgment and “decision” refer to resolving, selecting, choosing, establishing, comparing, etc. as “judgment” and “decision”. may be included.
  • judgment and “decision” may include regarding some action as having been “judged” or “determined.”
  • judgment (decision) may be read as “assuming", “expecting", “considering”, etc.
  • the phrase “based on” does not mean “based only on” unless explicitly stated otherwise. In other words, the phrase “based on” means both “based only on” and “based at least on.”
  • any reference to the elements herein, such as “first”, “second”, etc., does not generally limit the amount or order of those elements. These designations may be used herein as a convenient way of distinguishing between two or more elements. Thus, reference to a first and second element does not imply that only two elements may be employed therein or that the first element must precede the second element in any way.
  • a and B are different may mean “A and B are different from each other.” Note that the term may also mean that "A and B are each different from C”. Terms such as “separate” and “coupled” may also be interpreted similarly to “different.”
  • SYMBOLS 10 Delimiter insertion device, 11... Target text acquisition unit, 12... Interword time acquisition unit, 13... Delimiter insertion unit, 14... Output unit, 20... Speech recognition system, 21... Speech recognition result acquisition unit, 22... Speech recognition result output unit, M1...recording medium, m10...main module, m11...target text acquisition module, m12...interword time acquisition module, m13...delimiter insertion module, m14...output module, md1, md2...symbol insertion model .

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Machine Translation (AREA)

Abstract

区切り記号挿入装置は、発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する語間時間取得部と、区切り記号挿入モデル及び語間時間に基づいて、発話音声の音声認識により得られたテキストである対象テキストに区切り記号を挿入する区切り記号挿入部とを備える。区切り記号挿入モデルは、区切り記号除去文の入力に応じて区切り記号を示す区切り記号推測情報を出力する。区切り記号挿入部は、対象テキストを区切り記号挿入モデルに入力することにより得られた区切り記号推測情報に基づいて、対象テキストに区切り記号を挿入する。

Description

区切り記号挿入装置及び音声認識システム
 本発明は、区切り記号挿入装置及び音声認識システムに関する。
 音声認識により得られたテキストに句読点を挿入する技術が知られている。例えば、特許文献1には、統計に基づいて句読点がつけられたテキスト形態の訓練データにより訓練されたエンジンにより、テキストに句読点が挿入される技術が開示されている。
特開2012-508903号公報
 一般的な音声認識エンジンでは、単語列からなるテキストを発話音声から取得した後に、文として妥当な位置に句読点等の区切り記号(delimiter)が挿入される。テキスト情報のみの参照により区切り記号が挿入された場合に、その区切り記号の挿入位置が、文としては誤りではないものの、発話者の意図とは異なる場合があった。
 そこで、本発明は、上記問題点に鑑みてなされたものであり、発話音声の音声認識処理により得られたテキストに対して、発話者の意図した通りの位置に区切り記号を挿入することを目的とする。
 上記課題を解決するために、本開示の一側面に係る区切り記号挿入装置は、発話音声の音声認識により得られたテキストに含まれる単語の後に、文を区切る区切り記号を挿入する区切り記号挿入装置であって、発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する語間時間取得部と、区切り記号挿入モデル及び語間時間に基づいて、発話音声の音声認識により得られたテキストである対象テキストに区切り記号を挿入する区切り記号挿入部であって、区切り記号挿入モデルは、区切り記号を含まない文である区切り記号除去文を少なくとも入力とし、区切り記号除去文に含まれる各単語の後に挿入される区切り記号を示す区切り記号推測情報を出力し、区切り記号除去文と区切り記号を含む文である区切り記号入り文とのペアを含む学習データを用いた機械学習により生成されるモデルであり、対象テキストを区切り記号除去文として区切り記号挿入モデルに入力することにより得られ、語間時間に応じて調整された区切り記号推測情報に基づいて、対象テキストに区切り記号を挿入する、区切り記号挿入部と、を備える。
 上記の側面によれば、発話音声における語間時間が取得される。語間時間には、発話時における発話者の意図が反映されている。そして、対象テキストを区切り記号挿入モデルに入力することにより得られ且つ語間情報に応じて調整された区切り記号推測情報が取得される。ここで取得された区切り記号推測情報は、発話音声における各単語の語間時間に応じて調整された情報であるので、発話者の意図が反映された区切り記号を示す。そして、区切り記号推測情報に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを得ることが可能となる。
 発話音声の音声認識処理により得られたテキストに対して、発話者の意図した通りの位置に区切り記号を挿入することが可能となる。
本実施形態の区切り記号挿入装置の機能的構成を示すブロック図である。 区切り記号挿入装置のハードブロック図である。 本実施形態の区切り記号挿入装置により解決される課題を説明する図である。 対象テキストの取得処理を示す図である。 区切り記号挿入モデルの機会学習に用いられる学習データの第1の例を示す図である。 区切り記号挿入モデルの構成の第1の例を示す図である。 区切り記号推測情報を語間時間に基づいて調整するために参照される調整ルール情報の例を示す図である。 区切り記号推測情報の調整処理の例を示す図である。 語間時間の補正処理の例を示す図である。 区切り記号挿入モデルの機会学習に用いられる学習データの第2の例を示す図である。 区切り記号挿入モデルの構成の第2の例を示す図である。 区切り記号挿入モデルの機会学習に用いられる学習データの第3の例を示す図である。 本実施形態の音声認識システムの構成の例を示す機能ブロック図である。 区切り記号挿入装置における区切り記号挿入方法の処理内容を示すフローチャートである。 区切り記号挿入プログラムの構成を示す図である。
 本発明に係る区切り記号挿入装置及び音声認識システムの実施形態について図面を参照して説明する。なお、可能な場合には、同一の部分には同一の符号を付して、重複する説明を省略する。
 図1は、本実施形態に係る区切り記号挿入装置の機能的構成を示す図である。本実施形態の区切り記号挿入装置は、発話音声の音声認識により得られたテキストに含まれる単語の後に、文を区切る区切り記号(delimiter)を挿入する装置である。
 本実施形態では、区切り記号挿入装置10が、英文に挿入される区切り記号であるピリオド(period)、カンマ(comma)、疑問符(question mark)等を、英語の文からなるテキストに挿入する場合の例を説明するが、この例には限定されない。区切り記号挿入装置10が、日本語の文に句点及び読点等の区切り記号を挿入する装置であってもよいし、その他の言語の文に、その言語における区切り記号を挿入する装置であってもよい。
 区切り記号挿入装置10は、図1に示すように、機能的には、テキスト取得部11、語間時間取得部12、区切り記号挿入部13及び出力部14を備える。これらの各機能部11~14は、一つの装置に構成されてもよいし、複数の装置に分散されて構成されてもよい。
 なお、図1に示したブロック図は、機能単位のブロックを示している。これらの機能ブロック(構成部)は、ハードウェア及びソフトウェアの少なくとも一方の任意の組み合わせによって実現される。また、各機能ブロックの実現方法は特に限定されない。すなわち、各機能ブロックは、物理的又は論理的に結合した1つの装置を用いて実現されてもよいし、物理的又は論理的に分離した2つ以上の装置を直接的又は間接的に(例えば、有線、無線などを用いて)接続し、これら複数の装置を用いて実現されてもよい。機能ブロックは、上記1つの装置又は上記複数の装置にソフトウェアを組み合わせて実現されてもよい。
 機能には、判断、決定、判定、計算、算出、処理、導出、調査、探索、確認、受信、送信、出力、アクセス、解決、選択、選定、確立、比較、想定、期待、見做し、報知(broadcasting)、通知(notifying)、通信(communicating)、転送(forwarding)、構成(configuring)、再構成(reconfiguring)、割り当て(allocating、mapping)、割り振り(assigning)などがあるが、これらに限られない。たとえば、送信を機能させる機能ブロック(構成部)は、送信部(transmitting unit)や送信機(transmitter)と呼称される。いずれも、上述したとおり、実現方法は特に限定されない。
 例えば、本発明の一実施の形態における区切り記号挿入装置10は、コンピュータとして機能してもよい。図2は、本実施形態に係る区切り記号挿入装置10のハードウェア構成の一例を示す図である。区切り記号挿入装置10は、物理的には、プロセッサ1001、メモリ1002、ストレージ1003、通信装置1004、入力装置1005、出力装置1006、バス1007などを含むコンピュータ装置として構成されてもよい。
 なお、以下の説明では、「装置」という文言は、回路、デバイス、ユニットなどに読み替えることができる。区切り記号挿入装置10のハードウェア構成は、図に示した各装置を1つ又は複数含むように構成されてもよいし、一部の装置を含まずに構成されてもよい。
 区切り記号挿入装置10における各機能は、プロセッサ1001、メモリ1002などのハードウェア上に所定のソフトウェア(プログラム)を読み込ませることで、プロセッサ1001が演算を行い、通信装置1004による通信や、メモリ1002及びストレージ1003におけるデータの読み出し及び/又は書き込みを制御することで実現される。
 プロセッサ1001は、例えば、オペレーティングシステムを動作させてコンピュータ全体を制御する。プロセッサ1001は、周辺装置とのインターフェース、制御装置、演算装置、レジスタなどを含む中央処理装置(CPU:Central Processing Unit)で構成されてもよい。例えば、図1に示した各機能部11~14などは、プロセッサ1001で実現されてもよい。
 また、プロセッサ1001は、プログラム(プログラムコード)、ソフトウェアモジュールやデータを、ストレージ1003及び/又は通信装置1004からメモリ1002に読み出し、これらに従って各種の処理を実行する。プログラムとしては、上述の実施の形態で説明した動作の少なくとも一部をコンピュータに実行させるプログラムが用いられる。例えば、区切り記号挿入装置10の各機能部11~15は、メモリ1002に格納され、プロセッサ1001で動作する制御プログラムによって実現されてもよい。上述の各種処理は、1つのプロセッサ1001で実行される旨を説明してきたが、2以上のプロセッサ1001により同時又は逐次に実行されてもよい。プロセッサ1001は、1以上のチップで実装されてもよい。なお、プログラムは、電気通信回線を介してネットワークから送信されても良い。
 メモリ1002は、コンピュータ読み取り可能な記録媒体であり、例えば、ROM(Read Only Memory)、EPROM(Erasable Programmable ROM)、EEPROM(Electrically Erasable Programmable ROM)、RAM(Random Access Memory)などの少なくとも1つで構成されてもよい。メモリ1002は、レジスタ、キャッシュ、メインメモリ(主記憶装置)などと呼ばれてもよい。メモリ1002は、本発明の一実施の形態に係る疑似データ生成方法及び文生成方法を実施するために実行可能なプログラム(プログラムコード)、ソフトウェアモジュールなどを保存することができる。
 ストレージ1003は、コンピュータ読み取り可能な記録媒体であり、例えば、CD-ROM(Compact Disc ROM)などの光ディスク、ハードディスクドライブ、フレキシブルディスク、光磁気ディスク(例えば、コンパクトディスク、デジタル多用途ディスク、Blu-ray(登録商標)ディスク)、スマートカード、フラッシュメモリ(例えば、カード、スティック、キードライブ)、フロッピー(登録商標)ディスク、磁気ストリップなどの少なくとも1つで構成されてもよい。ストレージ1003は、補助記憶装置と呼ばれてもよい。上述の記憶媒体は、例えば、メモリ1002及び/又はストレージ1003を含むデータベース、サーバその他の適切な媒体であってもよい。
 通信装置1004は、有線及び/又は無線ネットワークを介してコンピュータ間の通信を行うためのハードウェア(送受信デバイス)であり、例えばネットワークデバイス、ネットワークコントローラ、ネットワークカード、通信モジュールなどともいう。
 入力装置1005は、外部からの入力を受け付ける入力デバイス(例えば、キーボード、マウス、マイクロフォン、スイッチ、ボタン、センサなど)である。出力装置1006は、外部への出力を実施する出力デバイス(例えば、ディスプレイ、スピーカー、LEDランプなど)である。なお、入力装置1005及び出力装置1006は、一体となった構成(例えば、タッチパネル)であってもよい。
 また、プロセッサ1001やメモリ1002などの各装置は、情報を通信するためのバス1007で接続される。バス1007は、単一のバスで構成されてもよいし、装置間で異なるバスで構成されてもよい。
 また、区切り記号挿入装置10は、マイクロプロセッサ、デジタル信号プロセッサ(DSP:Digital Signal Processor)、ASIC(Application Specific Integrated Circuit)、PLD(Programmable Logic Device)、FPGA(Field Programmable Gate Array)などのハードウェアを含んで構成されてもよく、当該ハードウェアにより、各機能ブロックの一部又は全てが実現されてもよい。例えば、プロセッサ1001は、これらのハードウェアの少なくとも1つで実装されてもよい。
 図3を参照して、本実施形態の区切り記号挿入装置10により解決される課題を説明する。図3に示される発話音声sp11,sp21はそれぞれ、異なる意図により発話者により発話されたものである。発話音声sp11,sp21の音声認識により得られた音声認識結果sr11,sr21は、同一のテキスト「I know it’s been there forever」である。発話音声sp11において、単語「know」の次の単語が発話されるまでの語間時間は1.2秒である。一方、発話音声sp21において、単語「know」の次の単語が発話されるまでの語間時間は0.1秒である。
 従来の区切り記号を挿入する技術では、音声認識結果として得られたテキストに、文として妥当な位置に区切り記号が挿入されるので、音声認識結果sr11,sr21のそれぞれに、従来の区切り記号挿入技術により区切り記号を挿入して得られた音声認識結果sr12,sr22は、元となる発話音声が異なるにも関わらず同一である。
 音声認識結果sr22は、発話音声sp21の発話者の意図のとおりに、単語「forever」の後にピリオドを有する。これに対して、発話音声sp11が単語「know」の後にピリオドが挿入されることを意図されたものであるにも関わらず、音声認識結果sr12は、単語「forever」の後にピリオドを有する。このピリオドの位置は、発話者の意図する位置とは異なる。
 本実施形態の区切り記号挿入装置10は、発話音声における語間時間を用いて区切り記号を挿入するので、発話音声sp11に基づいて得られた音声認識結果sr11に対して、単語「know」及び単語「forever」のそれぞれの後の位置にピリオドを挿入して音声認識結果sr13を得る。一方、区切り記号挿入装置10は、発話音声sp21に基づいて得られた音声認識結果sr21に対して、単語「forever」の後の位置にピリオドを挿入して音声認識結果sr23を得る。音声認識結果sr13,sr23は、発話音声sp11,sp21の発話者の意図した通りの位置に区切り記号を有する。
 次に、区切り記号挿入装置10の各機能部について説明する。テキスト取得部11は、区切り記号を挿入する対象のテキストである対象テキストを取得する。対象テキストは、発話音声の音声認識により得られたテキストである。語間時間取得部12は、発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する。
 図4は、対象テキスト及び語間時間の取得の例を示す図である。テキスト取得部11は、発話音声sp3に基づいて対象テキストtx1「I know it’s been there forever」を取得する。テキスト取得部11は、周知の音声認識処理技術及びその他の技術により、発話音声の音声認識を行うことにより、対象テキストを取得してもよい。
 語間時間取得部12は、対象テキストに含まれる各単語における、次の単語が発話されるまでの時間の長さを語間時間itとして取得する。語間時間取得部12は、発話音声sp3の音声認識時における無音時間を語間時間itとして取得してもよい。無音時間は、例えば、音量が所定の程度未満である時間である。また、語間時間取得部12は、発話音声の音声認識に用いられる音声認識エンジンから、音声認識の度に逐次的に音声認識結果を取得し、音声認識結果の取得及び更新の時間間隔を単語の発生の間の擬似的な無音時間とみなして、更新時間間隔を語間時間itとして取得してもよい。
 区切り記号挿入部13は、区切り記号挿入モデル及び語間時間に基づいて、対象テキストに区切り記号を挿入する。具体的には、区切り記号挿入部13は、対象テキストを区切り記号挿入モデルに入力することにより得られる区切り記号推測情報に基づいて、対象テキストに区切り記号を挿入する。区切り記号推測情報は、語間時間に応じて調整された情報を含む。
 区切り記号挿入モデルは、区切り記号を含まない文である区切り記号除去文を少なくとも入力とし、区切り記号除去文に含まれる各単語の後に挿入される区切り記号を示す区切り記号推測情報を出力する。また、区切り記号挿入モデルは、区切り記号除去文と区切り記号を含む文である区切り記号入り文とのペアを含む学習データを用いた機械学習により生成される。
 図5は、区切り記号挿入モデルの機会学習に用いられる学習データの第1の例を示す図である。図6は、区切り記号挿入モデルの構成の第1の例を示す図である。
 図5に示されるように、区切り記号挿入モデルmd1の機会学習に用いられる学習データの一例である学習データtd1は、区切り記号除去文id1と区切り記号入り文od1とのペアからなる。
 区切り記号入り文od1は、文を構成する単語列及び各単語の後に挿入される区切り記号を示すラベルである区切り記号ラベルを含む。区切り記号を示すラベルは、図5等において、模式的に以下のように図示される。
<O>…区切り記号なし
<P>…ピリオド、句点
<C>…カンマ、読点
<Q>…クエスチョンマーク、疑問符
区切り記号除去文id1は、区切り記号入り文od1から区切り記号ラベルを除去した文であってもよい。
 区切り記号挿入モデルmd1の機会学習では、学習過程の区切り記号挿入モデルmd1に区切り記号除去文id1を入力し、区切り記号挿入モデルmd1から得られた出力と教師データである区切り記号入り文od1との誤差に基づいて、区切り記号挿入モデルmd1を構成する重み及びパラメータ等が更新される。
 図6に示されるように、学習済みの区切り記号挿入モデルmd1は、区切り記号除去文sd1の入力に応じて、区切り記号推測情報dp1を出力する。
 区切り記号挿入モデルmd1は、ニューラルネットワークを含んで構成されるモデルであってもよい。さらに具体的には、区切り記号挿入モデルmd1は、入力された文に含まれる各単語の後に挿入される区切り記号を予測する系列ラベリングタスクを解決する系列ラベリングモデルとして構成されてもよい。
 学習済みのニューラルネットワークを含むモデルである区切り記号挿入モデルmd1は、コンピュータにより読み込まれ又は参照され、コンピュータに所定の処理を実行させ及びコンピュータに所定の機能を実現させるプログラムとして捉えることができる。
 即ち、本実施形態の学習済みの区切り記号挿入モデルmd1は、CPU及びメモリを備えるコンピュータにおいて用いられる。具体的には、コンピュータのCPUが、メモリに記憶された学習済みの区切り記号挿入モデルmd1からの指令に従って、ニューラルネットワークの入力層に入力された入力データに対し、各層に対応する学習済みの重み付け係数(パラメタ)と応答関数等に基づく演算を行い、出力層から結果(確率)を出力するよう動作する。
 区切り記号推測情報dp1は、区切り記号除去文sd1に含まれる各単語の後に挿入されうる各種の区切り記号の尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号なし尤度を含む。そして、記号挿入尤度及び記号なし尤度のうちの最大の尤度に基づいて、各単語の後に、複数の種類のうちのいずれかの区切り記号を挿入すること又は区切り記号を未挿入とすることが決定(ラベリング)される。
 図6に例示される区切り記号推測情報dp1は、以下のとおり、単語「I」に関する各区切り記号の記号挿入尤度及び記号なし尤度を含む。
<O>:90%、<C>:5%、<P>:2%、<Q>:3%
従って、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であるので、単語「I」に対して区切り記号なしのラベル<O>がラベリングされる。
 同様に、区切り記号推測情報dp1において、以下のとおり、単語「know」に関する各区切り記号の記号挿入尤度及び記号なし尤度を含む。
<O>:60%、<C>:5%、<P>:30%、<Q>:5%
従って、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であるので、単語「know」に対して区切り記号なしのラベル<O>がラベリングされる。
 区切り記号挿入部13は、区切り記号推測情報の語間時間による調整の一例として、区切り記号挿入モデルから出力された区切り記号推測情報を、語間時間に基づいて調整してもよい。具体的には、区切り記号挿入部13は、区切り記号推測情報に含まれる記号挿入尤度及び記号なし尤度を、語間時間に基づいて調整してもよい。
 区切り記号挿入部13は、対象テキストに含まれる単語のうちの一の単語における語間時間が第1の時間である場合に、当該一の単語の記号なし尤度を上げるように調整し、又は/及び、複数の種類の区切り記号のうちの少なくとも一種の当該一の単語の区切り記号の記号挿入尤度を下げるように調整し、当該一の単語における語間時間が第1の時間より長い第2の時間である場合に、複数の種類の区切り記号のうちの少なくとも一種の当該一の単語の区切り記号の記号挿入尤度を上げるように調整し、又は/及び、当該一の単語の記号なし尤度を下げるように調整してもよい。
 そして、区切り記号挿入部13は、調整された記号挿入尤度及び記号なし尤度のうちの最大の尤度に基づいて、当該一の単語の後に、複数の種類のうちのいずれかの区切り記号を挿入し、又は区切り記号を未挿入とする。
 このような記号挿入尤度及び記号なし尤度の調整のために、区切り記号挿入部13は、例えば、調整ルール情報を参照して尤度を調整してもよい。図7は、区切り記号推測情報を語間時間に基づいて調整するために参照される調整ルール情報の例を示す図である。調整ルール情報は、区切り記号挿入部13がアクセス可能な記憶手段に記憶されていてもよいし、区切り記号挿入部13の内部にテーブルとして備えられていてもよい。
 図7に示されるように、調整ルール情報は、語間時間の範囲ごとに、語間時間の長さを表す語間時間カテゴリ及び尤度の調整の内容を示す尤度調整情報が関連付けられた情報である。例えば、語間時間(x)が0.1秒以下である場合には、語間時間カテゴリは「無」であり、「記号なし尤度を50%上げる」ことが尤度の調整として実施されることが調整ルールとして規定されている。また、例えば、語間時間(x)が0.5秒より長く且つ1.0秒以下である場合には、語間時間カテゴリは「中」であり、「記号挿入尤度を50%上げる」ことが尤度の調整として実施されることが調整ルールとして規定されている。
 図8は、区切り記号推測情報の調整処理の例を示す図である。図8に示される区切り記号推測情報dp21は、区切り記号挿入モデルmd1から出力された調整処理前の区切り記号推測情報の一部を示す。
 区切り記号推測情報dp21は、単語「know」に関する各区切り記号の記号挿入尤度lh21及び記号なし尤度を含む。調整前の区切り記号推測情報dp21に従うならば、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であるので、単語「know」に対して区切り記号なしのラベル<O>がラベリングされる。
 区切り記号挿入部13は、単語「know」の語間時間it2に基づいて、区切り記号推測情報dp21の各尤度を調整する。具体的には、語間時間取得部12により取得された単語「know」の語間時間it2は0.9秒であるので、区切り記号挿入部13は、調整ルール情報(図7)を参照して、語間時間0.9秒に関連付けられた尤度調整情報「ピリオド、カンマ、クエスチョンマークの記号挿入尤度を50%上げる」を取得し、カンマ、ピリオド及びクエスチョンマークの記号挿入尤度lh21を、取得した尤度調整情報に従って調整する。
 そして、区切り記号挿入部13は、記号挿入尤度lh21のそれぞれの値を50%増加させて、調整後の区切り記号推測情報dp22を得る。区切り記号推測情報dp22において、ピリオド(ラベル<P>)の記号挿入尤度が最大であるので、区切り記号挿入部13は、単語「know」に対して区切り記号「ピリオド」のラベル<P>をラベリングする。そして、区切り記号挿入部13は、対象テキストに含まれる単語「know」の後に、ラベリングされたラベル<P>に基づいて、ピリオドを挿入する。
 このように、対象テキストに含まれる一の単語の語間時間が長いほど、記号挿入尤度が上げられ、又は/及び、記号なし尤度が下げられるように相対的に調整され、一の単語の語間時間が短いほど、記号挿入尤度が下げられ、又は/及び、記号なし尤度が上げられるように相対的に調整されるので、発話者の意図が、記号挿入尤度及び記号なし尤度に反映される。そして、調整された記号挿入尤度及び記号なし尤度に基づいて、区切り記号が挿入され又は区切り記号が未挿入とされるので、発話者の意図に沿う適切な位置に区切り記号が挿入されたテキストを得ることが可能となる。
 区切り記号挿入部13は、区切り記号推測情報の調整に用いるための語間時間を所定の条件に従って補正してもよい。図9は、語間時間の補正処理の例を示す図である。区切り記号挿入部13は、対象テキストに含まれる全単語のうちの一以上の単語の語間時間の長さの程度が所与の程度より長い場合に、全単語のそれぞれの前記語間時間を所与の程度で短くなるように補正してもよい。そして、区切り記号挿入部13は、補正された語間時間である補正語間時間に基づいて、記号挿入尤度及び/又は記号なし尤度を調整してもよい。
 図9に示される対象テキストtx31は、補正前の語間時間it31を含む。ここでは、一例として、対象テキストに含まれる全ての単語のうち、半数以上の単語の語間時間が、語間時間カテゴリ「小」に対応する語間時間以上である場合に、全ての単語の語間時間を0.5秒減ずることが、補正処理の内容として予め設定されていることとする。
 この場合において、区切り記号挿入部13は、対象テキストtx31に含まれる単語の語間時間it31のうちの全ての単語の語間時間が、語間時間カテゴリ「小」に対応する語間時間以上であることを判定する。そして、区切り記号挿入部13は、対象テキストtx32に示されるように、語間時間it31をそれぞれ0.5秒減じて、補正後の語間時間である補正語間時間it32を得る。区切り記号挿入部13は、補正語間時間it32に基づいて、対象テキストtx32に含まれる各単語の記号挿入尤度及び/又は記号なし尤度を調整する。
 発話者の発話が、全体的に語間時間が長くなるような傾向を有する場合には、区切り記号の挿入後のテキストにおいて、発話者が意図しない位置に過剰に区切り記号が挿入される可能性があるところ、上記の語間時間の補正処理によれば、語間時間の長さが所定の程度より長い場合に、語間時間が短くなるように補正された補正語間時間に基づいて記号挿入尤度及び/又は記号なし尤度が調整されるので、発話者の意図が好適に反映された適切な位置に区切り記号が挿入されたテキストを得ることが可能となる。
 次に、区切り記号推測情報の語間時間に応じた調整処理の第2の例を説明する。図10は、区切り記号挿入モデルの機会学習に用いられる学習データの第2の例を示す図である。図11は、区切り記号挿入モデルの構成の第2の例を示す図である。
 図10に示されるように、区切り記号挿入モデルmd2の機会学習に用いられる学習データの一例である学習データtd2は、区切り記号除去文id2と区切り記号入り文od2とのペアからなる。
 区切り記号入り文od2は、図5を参照して説明した区切り記号入り文od1と同様に、文を構成する単語列及び各単語の後に挿入される区切り記号を示す区切り記号ラベルを含む。区切り記号除去文id2は、区切り記号入り文od2から区切り記号ラベルが除去された文、及び、当該文を構成する各単語にそれぞれ関連付けられた語間時間it4を含む。
 区切り記号挿入モデルmd2の機会学習では、学習過程の区切り記号挿入モデルmd2に区切り記号除去文id2を入力し、区切り記号挿入モデルmd2から得られた出力と教師データである区切り記号入り文od2との誤差に基づいて、区切り記号挿入モデルmd2を構成する重み及びパラメータ等が更新される。
 区切り記号挿入モデルmd2は、ニューラルネットワークを含んで構成されるモデルであってもよい。さらに具体的には、区切り記号挿入モデルmd2は、入力された文に含まれる各単語の後に挿入される区切り記号を予測する系列ラベリングタスクを解決する系列ラベリングモデルとして構成されてもよい。
 図11に示されるように、学習済みの区切り記号挿入モデルmd2は、区切り記号除去文sd2の入力に応じて、区切り記号推測情報dp2を出力する。区切り記号除去文sd2は、区切り記号除去文sd2を構成する各単語に関連付けられた語間時間it5を含む。区切り記号推測情報dp2は、区切り記号除去文sd2に含まれる各単語に関する記号挿入尤度及び記号なし尤度を含む。
 学習済みのニューラルネットワークを含むモデルである区切り記号挿入モデルmd2は、コンピュータにより読み込まれ又は参照され、コンピュータに所定の処理を実行させ及びコンピュータに所定の機能を実現させるプログラムとして捉えることができる。
 即ち、本実施形態の学習済みの区切り記号挿入モデルmd2は、CPU及びメモリを備えるコンピュータにおいて用いられる。具体的には、コンピュータのCPUが、メモリに記憶された学習済みの区切り記号挿入モデルmd2からの指令に従って、ニューラルネットワークの入力層に入力された入力データに対し、各層に対応する学習済みの重み付け係数(パラメタ)と応答関数等に基づく演算を行い、出力層から結果(確率)を出力するよう動作する。
 図6を参照して説明した例では、区切り記号挿入モデルmd1から出力された区切り記号推測情報dp1に含まれる記号挿入尤度及び記号なし尤度が、語間時間に基づいて調整される。これに対して、区切り記号挿入モデルmd2では、語間時間it4を含む区切り記号除去文id2が機会学習における学習データとして入力に用いられ、語間時間it5を含む区切り記号除去文sd2の入力に応じて区切り記号推測情報dp2を出力するので、区切り記号推測情報dp2に含まれる記号挿入尤度及び記号なし尤度は、区切り記号挿入モデルmd2における演算により、語間時間に応じた調整がされた値である。
 区切り記号挿入モデルmd2において、以下のような、単語「I」に関する各区切り記号の記号挿入尤度及び記号なし尤度が算出される。
<O>:90%、<C>:5%、<P>:2%、<Q>:3%
従って、区切り記号推測情報dp2は、単語「I」に関する各区切り記号の記号挿入尤度及び記号なし尤度のうちの最大の尤度を有する区切り記号なしのラベル<O>が、単語「I」に対してラベリングされたことを示す情報(I=<O>)を含む。そして、区切り記号挿入部13は、対象テキストに含まれる単語「I」の後に、ラベリングされたラベル<O>に基づいて、区切り記号を未挿入とすることを決定する。
 また、区切り記号挿入モデルmd2において、以下のような、単語「know」に関する各区切り記号の記号挿入尤度及び記号なし尤度が算出される。
<O>:10%、<C>:20%、<P>:70%、<Q>:0%
図6に例示した区切り記号推測情報dp1では、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であったのに対して、区切り記号挿入モデルmd2により算出された各尤度においては、ピリオドを挿入すること(ラベル<P>)の記号挿入尤度が最大である。従って、区切り記号推測情報dp2は、ピリオドを挿入することが単語「know」に対してラベリングされたことを示す情報(know=<P>)を含む。そして、区切り記号挿入部13は、対象テキストに含まれる単語「know」の後に、ラベリングされたラベル<P>に基づいて、ピリオドを挿入する。
 このように、語間時間が各単語に関連付けられた区切り記号除去文が学習データとして用いられた機会学習により区切り記号挿入モデルmd2が生成され、各単語の語間時間を含む区切り記号除去文が区切り記号挿入モデルmd2の入力とされるので、各単語に語間時間が関連付けられた対象テキストを区切り記号挿入モデルに入力することにより、区切り記号挿入モデルからの出力に対する語間時間に基づく別途の調整処理を行うことなく、語間時間により調整された区切り記号推測情報を得ることが可能となる。そして、区切り記号推測情報に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを容易に得ることが可能となる。
 次に、区切り記号挿入モデルの他の例について説明する。図12は、区切り記号挿入モデルの機会学習に用いられる学習データの第3の例を示す図である。図12に示される例では、学習データtd3は、区切り記号除去文id3と区切り記号入り文od3とのペアからなる。
 区切り記号入り文od3は、図5及び図10を参照して説明した区切り記号入り文od1,od2と同様に、文を構成する単語列及び各単語の後に挿入される区切り記号を示す区切り記号ラベルを含む。区切り記号除去文id3は、区切り記号入り文od3から区切り記号ラベルが除去された文、及び、当該文を構成する各単語にそれぞれ関連付けられた語間時間に加えて、シチュエーション情報st3を含む。シチュエーション情報st3は、発話音声が発話されたときの状況を示す情報であって、タグとして区切り記号除去文id3に付与されてもよい。
 シチュエーション情報st3は、以下のような、発話音声が発話されたときの状況に応じたバリエーションを有してもよい。
会議:<Meeting>
講演:<Lecture>
雑談:<Chatting>
 学習データtd3を用いた区切り記号挿入モデルの機会学習では、学習過程の区切り記号挿入モデルに区切り記号除去文id3を入力し、区切り記号挿入モデルから得られた出力と教師データである区切り記号入り文od3との誤差に基づいて、区切り記号挿入モデルを構成する重み及びパラメータ等が更新される。
 学習データtd3を用いた機械学習による学習済みの区切り記号挿入モデルは、シチュエーション情報及び各単語に関連付けられた語間時間を含む区切り記号除去文(対象テキスト)の入力に応じて、区切り記号推測情報を出力する。出力された区切り記号推測情報は、区切り記号除去文に含まれる各単語に関する記号挿入尤度及び記号なし尤度、並びに、最大の尤度に基づく区切り記号又は区切り記号未挿入のラベルを含む。
 区切り記号挿入部13は、シチュエーション情報が関連付けられた対象テキストを区切り記号挿入モデルに入力することに応じて得られた区切り記号推測情報に基づいて、各単語の後にラベルの応じた区切り記号を挿入し、又は、区切り記号を未挿入とすることを決定する。
 このように、シチュエーション情報が関連付けられた区切り記号除去文が学習データとして用いられた機会学習により区切り記号挿入モデルが生成され、シチュエーション情報を含む区切り記号除去文が区切り記号挿入モデルの入力とされるので、シチュエーション情報が関連付けられた対象テキストを区切り記号挿入モデルに入力することにより、発話音声が発話されたときの状況に応じた発話の傾向が考慮された区切り記号推測情報を得ることが可能となる。
 再び図1を参照して、出力部14は、区切り記号挿入部13により区切り記号が挿入された対象テキストである区切り記号挿入テキストを出力する。出力の態様は限定されず、出力部14は、区切り記号挿入テキストを所定の表示手段に表示させてもよいし、所定の記憶手段に区切り記号挿入テキストを記憶させてもよいし、所定の装置に区切り記号挿入テキストを送信してもよい。
 図13は、本実施形態の音声認識システムの構成の例を示す機能ブロック図である。図13に示されるように、音声認識システム20は、区切り記号挿入装置10を含んで構成され、音声認識結果取得部21及び音声認識結果出力部22を備える。
 音声認識結果取得部21は、音声認識エンジンにより音声認識された結果のテキスト及び当該テキストにおける各単語の語間時間を第1の音声認識結果として取得する。
 区切り記号挿入装置10のテキスト取得部11は、第1の音声認識結果に含まれるテキストを対象テキストとして取得する。区切り記号挿入装置10の語間時間取得部12は、第1の音声認識結果に含まれる語間時間を取得する。
 区切り記号挿入装置10の区切り記号挿入部13は、第1の音声認識結果に含まれるテキストを対象テキストとして、区切り記号を挿入する処理を実施する。
 音声認識結果出力部22は、区切り記号挿入部13により区切り記号が挿入された対象テキストを、第2の音声認識結果として出力する。
 音声認識システム20によれば、音声認識処理において語間時間が考慮されない音声認識エンジンから得られた第1の音声認識結果に基づいて、第1の音声認識結果に含まれるテキストを対象テキストとして、第1の音声認識結果に含まれる語間時間に応じて調整された区切り記号推測情報を取得できるので、発話者の意図に沿う位置に区切り記号が挿入された対象テキストを含む第2の音声認識結果を得ることが可能となる。
 図14は、区切り記号挿入装置10における区切り記号挿入方法の処理内容を示すフローチャートである。
 ステップS1において、テキスト取得部11は、発話音声の音声認識により得られたテキストであって、区切り記号を挿入する対象のテキストである対象テキストを取得する。
 ステップS2において、語間時間取得部12は、発話音声における各単語の語間時間を取得する。
 ステップS3において、区切り記号挿入部13は、区切り記号挿入モデルに対象テキストを入力する。
 ステップS4において、区切り記号挿入部13は、区切り記号推測情報を取得する。区切り記号推測情報は、語間時間に応じた調整された各単語の記号挿入尤度及び記号なし尤度に基づく、各単語に対するいずれかの区切り記号を挿入すること又は区切り記号を未挿入とすることを示すラベルを含む。
 ステップS5において、区切り記号挿入部13は、区切り記号推測情報に基づいて、対象テキストに区切り記号を挿入する。
 ステップS6において、出力部14は、区切り記号が挿入された対象テキストである区切り記号挿入テキストを出力する。
 次に、図15を参照して、コンピュータを、本実施形態の区切り記号挿入装置10として機能させるための区切り記号挿入プログラムについて説明する。図15は、区切り記号挿入プログラムの構成を示す図である。区切り記号挿入プログラムP1は、区切り記号挿入装置10における区切り記号挿入処理を統括的に制御するメインモジュールm10、テキスト取得モジュールm11、語間時間取得モジュールm12、区切り記号挿入モジュールm13及び出力モジュールm14を備えて構成される。そして、各モジュールm11~m14により、テキスト取得部11、語間時間取得部12、区切り記号挿入部13及び出力部14のための各機能が実現される。
 なお、区切り記号挿入プログラムP1は、通信回線等の伝送媒体を介して伝送される態様であってもよいし、図15に示されるように、記録媒体M1に記憶される態様であってもよい。
 以上説明した本実施形態の区切り記号挿入装置10、区切り記号挿入方法、区切り記号挿入プログラムP1によれば、発話音声における語間時間が取得される。語間時間には、発話時における発話者の意図が反映されている。そして、対象テキストを区切り記号挿入モデルに入力することにより得られ且つ語間情報に応じて調整された区切り記号推測情報が取得される。ここで取得された区切り記号推測情報は、発話音声における各単語の語間時間に応じて調整された情報であるので、発話者の意図が反映された区切り記号を示す。そして、区切り記号推測情報に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを得ることが可能となる。
 本開示に係る発明は、例えば、以下のように把握される。
 本開示の第1の一側面に係る区切り記号挿入装置は、発話音声の音声認識により得られたテキストに含まれる単語の後に、文を区切る区切り記号を挿入する区切り記号挿入装置であって、前記発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する語間時間取得部と、区切り記号挿入モデル及び前記語間時間に基づいて、前記発話音声の音声認識により得られたテキストである対象テキストに区切り記号を挿入する区切り記号挿入部であって、前記区切り記号挿入モデルは、前記区切り記号を含まない文である区切り記号除去文を少なくとも入力とし、前記区切り記号除去文に含まれる各単語の後に挿入される区切り記号を示す区切り記号推測情報を出力し、前記区切り記号除去文と区切り記号を含む文である区切り記号入り文とのペアを含む学習データを用いた機械学習により生成されるモデルであり、前記対象テキストを前記区切り記号除去文として前記区切り記号挿入モデルに入力することにより得られ、前記語間時間に応じて調整された前記区切り記号推測情報に基づいて、前記対象テキストに区切り記号を挿入する、区切り記号挿入部と、を備える。
 上記の側面によれば、発話音声における語間時間が取得される。語間時間には、発話時における発話者の意図が反映されている。そして、対象テキストを区切り記号挿入モデルに入力することにより得られ且つ語間情報に応じて調整された区切り記号推測情報が取得される。ここで取得された区切り記号推測情報は、発話音声における各単語の語間時間に応じて調整された情報であるので、発話者の意図が反映された区切り記号を示す。そして、区切り記号推測情報に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを得ることが可能となる。
 第2の側面に係る区切り記号挿入装置では、第1の側面に係る区切り記号挿入装置において、前記区切り記号挿入部は、前記区切り記号挿入モデルから出力された前記区切り記号推測情報を、前記語間時間に基づいて調整することとしてもよい。
 上記の側面によれば、区切り記号推測情報に、語間時間に表された発話者の意図を確実に反映させることが可能となる。
 第3の側面に係る区切り記号挿入装置では、第2の側面に係る区切り記号挿入装置において、前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号なし尤度を含み、前記区切り記号挿入部は、前記対象テキストに含まれる単語のうちの一の単語における前記語間時間が第1の時間である場合に、前記一の単語の前記記号なし尤度を上げるように調整し、又は/及び、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を下げるように調整し、前記一の単語における前記語間時間が前記第1の時間より長い第2の時間である場合に、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を上げるように調整し、又は/及び、前記一の単語の前記記号なし尤度を下げるように調整し、前記記号挿入尤度及び前記記号なし尤度のうちの最大の尤度に基づいて、前記一の単語の後に、複数の種類のうちのいずれかの区切り記号を挿入し、又は区切り記号を未挿入とすることとしてもよい。
 上記の側面によれば、対象テキストに含まれる一の単語の語間時間が長いほど、記号挿入尤度が上げられ、又は/及び、記号なし尤度が下げられるように相対的に調整され、一の単語の語間時間が短いほど、記号挿入尤度が下げられ、又は/及び、記号なし尤度が上げられるように相対的に調整されるので、発話者の意図が、記号挿入尤度及び記号なし尤度に反映される。そして、調整された記号挿入尤度及び記号なし尤度に基づいて、区切り記号が挿入され又は区切り記号が未挿入とされるので、発話者の意図に沿う適切な位置に区切り記号が挿入されたテキストを得ることが可能となる。
 第4の側面に係る区切り記号挿入装置では、第3の側面に係る区切り記号挿入装置において、前記区切り記号挿入部は、前記対象テキストに含まれる全単語のうちの一以上の単語の前記語間時間の長さの程度が、所与の程度より長い場合に、前記全単語のそれぞれの前記語間時間を所与の程度で短くなるように補正した補正語間時間に基づいて、前記記号挿入尤度及び/又は前記記号なし尤度を調整することとしてもよい。
 発話者の発話が、全体的に語間時間が長くなるような傾向を有する場合には、区切り記号の挿入後のテキストにおいて、発話者が意図しない位置に過剰に区切り記号が挿入される可能性があるところ、上記の側面によれば、語間時間の長さが所定の程度より長い場合に、語間時間が短くなるように補正された補正語間時間に基づいて記号挿入尤度及び/又は記号なし尤度が調整されるので、発話者の意図が好適に反映された適切な位置に区切り記号が挿入されたテキストを得ることが可能となる。
 第5の側面に係る区切り記号挿入装置では、第1の側面に係る区切り記号挿入装置において、前記区切り記号挿入モデルは、前記区切り記号除去文に含まれる各単語の語間時間を入力として更に含み、前記語間時間が各単語に関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、前記語間時間により調整された前記区切り記号推測情報を出力し、前記区切り記号挿入部は、各単語に前記語間時間が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力することとしてもよい。
 上記の側面によれば、語間時間が各単語に関連付けられた区切り記号除去文が学習データとして用いられた機会学習により区切り記号挿入モデルが生成され、各単語の語間時間を含む区切り記号除去文が区切り記号挿入モデルの入力とされるので、各単語に語間時間が関連付けられた対象テキストを区切り記号挿入モデルに入力することにより、区切り記号挿入モデルからの出力に対する語間時間に基づく別途の調整処理を行うことなく、語間時間により調整された区切り記号推測情報を得ることが可能となる。そして、区切り記号推測情報に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを容易に得ることが可能となる。
 第6の側面に係る区切り記号挿入装置では、第5の側面に係る区切り記号挿入装置において、前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号無し尤度を含むこととしてもよい。
 上記の側面によれば、各単語に関して、語間時間により調整された記号挿入尤度及び記号なし尤度を含む区切り記号推測情報を得ることが可能となる。調整された記号挿入尤度及び記号なし尤度に基づいて対象テキストに区切り記号が挿入されることにより、発話者の意図に沿う位置に区切り記号が挿入されたテキストを容易に得ることが可能となる。
 第7の側面に係る区切り記号挿入装置では、第5または6の側面に係る区切り記号挿入装置において、前記区切り記号挿入モデルは、前記区切り記号除去文に対応する発話音声が発話されたときの状況を示すシチュエーション情報を入力として更に含み、前記語間時間が各単語に関連付けられると共に前記シチュエーション情報が関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、前記区切り記号挿入部は、前記シチュエーション情報が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力することとしてもよい。
 上記の側面によれば、シチュエーション情報が関連付けられた区切り記号除去文が学習データとして用いられた機会学習により区切り記号挿入モデルが生成され、シチュエーション情報を含む区切り記号除去文が区切り記号挿入モデルの入力とされるので、シチュエーション情報が関連付けられた対象テキストを区切り記号挿入モデルに入力することにより、発話音声が発話されたときの状況に応じた発話の傾向が考慮された区切り記号推測情報を得ることが可能となる。
 第1の側面に係る音声認識システムでは、第1~7の側面のいずれか一つの側面に係る区切り記号挿入装置を含み、音声認識エンジンにより音声認識された結果のテキスト及び該テキストにおける各単語の前記語間時間を第1の音声認識結果として取得する音声認識結果取得部と、音声認識結果出力部と、を備え、前記区切り記号挿入装置の前記語間時間取得部は、前記第1の音声認識結果に含まれる前記語間時間を取得し、前記区切り記号挿入装置の前記区切り記号挿入部は、前記第1の音声認識結果に含まれるテキストを前記対象テキストとして、区切り記号を挿入し、前記音声認識結果出力部は、前記区切り記号挿入部により区切り記号が挿入された前記対象テキストを、第2の音声認識結果として出力することとしてもよい。
 上記の側面によれば、音声認識処理において語間時間が考慮されない音声認識エンジンから得られた第1の音声認識結果に基づいて、第1の音声認識結果に含まれるテキストを対象テキストとして、第1の音声認識結果に含まれる語間時間に応じて調整された区切り記号推測情報を取得できるので、発話者の意図に沿う位置に区切り記号が挿入された対象テキストを含む第2の音声認識結果を得ることが可能となる。
 以上、本実施形態について詳細に説明したが、当業者にとっては、本実施形態が本明細書中に説明した実施形態に限定されるものではないということは明らかである。本実施形態は、特許請求の範囲の記載により定まる本発明の趣旨及び範囲を逸脱することなく修正及び変更態様として実施することができる。したがって、本明細書の記載は、例示説明を目的とするものであり、本実施形態に対して何ら制限的な意味を有するものではない。
 情報の通知は、本開示において説明した態様/実施形態に限られず、他の方法を用いて行われてもよい。例えば、情報の通知は、物理レイヤシグナリング(例えば、DCI(Downlink Control Information)、UCI(Uplink Control Information))、上位レイヤシグナリング(例えば、RRC(Radio Resource Control)シグナリング、MAC(Medium Access Control)シグナリング、報知情報(MIB(Master Information Block)、SIB(System Information Block)))、その他の信号又はこれらの組み合わせによって実施されてもよい。また、RRCシグナリングは、RRCメッセージと呼ばれてもよく、例えば、RRC接続セットアップ(RRC Connection Setup)メッセージ、RRC接続再構成(RRC Connection Reconfiguration)メッセージなどであってもよい。
 本明細書で説明した各態様/実施形態は、LTE(Long Term Evolution)、LTE-A(LTE-Advanced)、SUPER 3G、IMT-Advanced、4G、5G、FRA(Future Radio Access)、W-CDMA(登録商標)、GSM(登録商標)、CDMA2000、UMB(Ultra Mobile Broadband)、IEEE 802.11(Wi-Fi)、IEEE 802.16(WiMAX)、IEEE 802.20、UWB(Ultra-WideBand)、Bluetooth(登録商標)、その他の適切なシステムを利用するシステム及び/又はこれらに基づいて拡張された次世代システムに適用されてもよい。また、複数のシステムが組み合わされて(例えば、LTE及びLTE-Aの少なくとも一方と5Gとの組み合わせ等)適用されてもよい。
 本明細書で説明した各態様/実施形態の処理手順、シーケンス、フローチャートなどは、矛盾の無い限り、順序を入れ替えてもよい。例えば、本明細書で説明した方法については、例示的な順序で様々なステップの要素を提示しており、提示した特定の順序に限定されない。
 本開示において基地局によって行われるとした特定動作は、場合によってはその上位ノード(upper node)によって行われることもある。基地局を有する1つ又は複数のネットワークノード(network nodes)からなるネットワークにおいて、端末との通信のために行われる様々な動作は、基地局及び基地局以外の他のネットワークノード(例えば、MME又はS-GWなどが考えられるが、これらに限られない)の少なくとも1つによって行われ得ることは明らかである。上記において基地局以外の他のネットワークノードが1つである場合を例示したが、複数の他のネットワークノードの組み合わせ(例えば、MME及びS-GW)であってもよい。
 情報等(※「情報、信号」の項目参照)は、上位レイヤ(又は下位レイヤ)から下位レイヤ(又は上位レイヤ)へ出力され得る。複数のネットワークノードを介して入出力されてもよい。
 入出力された情報等は特定の場所(例えば、メモリ)に保存されてもよいし、管理テーブルで管理してもよい。入出力される情報等は、上書き、更新、または追記され得る。出力された情報等は削除されてもよい。入力された情報等は他の装置へ送信されてもよい。
 判定は、1ビットで表される値(0か1か)によって行われてもよいし、真偽値(Boolean:trueまたはfalse)によって行われてもよいし、数値の比較(例えば、所定の値との比較)によって行われてもよい。
 本開示において説明した各態様/実施形態は単独で用いてもよいし、組み合わせて用いてもよいし、実行に伴って切り替えて用いてもよい。また、所定の情報の通知(例えば、「Xであること」の通知)は、明示的に行うものに限られず、暗黙的(例えば、当該所定の情報の通知を行わない)ことによって行われてもよい。
 以上、本開示について詳細に説明したが、当業者にとっては、本開示が本開示中に説明した実施形態に限定されるものではないということは明らかである。本開示は、請求の範囲の記載により定まる本開示の趣旨及び範囲を逸脱することなく修正及び変更態様として実施することができる。したがって、本開示の記載は、例示説明を目的とするものであり、本開示に対して何ら制限的な意味を有するものではない。
 ソフトウェアは、ソフトウェア、ファームウェア、ミドルウェア、マイクロコード、ハードウェア記述言語と呼ばれるか、他の名称で呼ばれるかを問わず、命令、命令セット、コード、コードセグメント、プログラムコード、プログラム、サブプログラム、ソフトウェアモジュール、アプリケーション、ソフトウェアアプリケーション、ソフトウェアパッケージ、ルーチン、サブルーチン、オブジェクト、実行可能ファイル、実行スレッド、手順、機能などを意味するよう広く解釈されるべきである。
 また、ソフトウェア、命令などは、伝送媒体を介して送受信されてもよい。例えば、ソフトウェアが、同軸ケーブル、光ファイバケーブル、ツイストペア及びデジタル加入者回線(DSL)などの有線技術及び/又は赤外線、無線及びマイクロ波などの無線技術を使用してウェブサイト、サーバ、又は他のリモートソースから送信される場合、これらの有線技術及び/又は無線技術は、伝送媒体の定義内に含まれる。
 本開示において説明した情報、信号などは、様々な異なる技術のいずれかを使用して表されてもよい。例えば、上記の説明全体に渡って言及され得るデータ、命令、コマンド、情報、信号、ビット、シンボル、チップなどは、電圧、電流、電磁波、磁界若しくは磁性粒子、光場若しくは光子、又はこれらの任意の組み合わせによって表されてもよい。
 なお、本開示において説明した用語及び/又は本明細書の理解に必要な用語については、同一の又は類似する意味を有する用語と置き換えてもよい。
 本明細書で使用する「システム」および「ネットワーク」という用語は、互換的に使用される。
 また、本明細書で説明した情報、パラメータなどは、絶対値で表されてもよいし、所定の値からの相対値で表されてもよいし、対応する別の情報で表されてもよい。例えば、無線リソースはインデックスによって指示されるものであってもよい。
 上述したパラメータに使用する名称はいかなる点においても限定的な名称ではない。さらに、これらのパラメータを使用する数式等は、本開示で明示的に開示したものと異なる場合もある。様々なチャネル(例えば、PUCCH、PDCCHなど)及び情報要素は、あらゆる好適な名称によって識別できるので、これらの様々なチャネル及び情報要素に割り当てている様々な名称は、いかなる点においても限定的な名称ではない。
 本開示で使用する「判断(determining)」、「決定(determining)」という用語は、多種多様な動作を包含する場合がある。「判断」、「決定」は、例えば、判定(judging)、計算(calculating)、算出(computing)、処理(processing)、導出(deriving)、調査(investigating)、探索(looking up、search、inquiry)(例えば、テーブル、データベース又は別のデータ構造での探索)、確認(ascertaining)した事を「判断」「決定」したとみなす事などを含み得る。また、「判断」、「決定」は、受信(receiving)(例えば、情報を受信すること)、送信(transmitting)(例えば、情報を送信すること)、入力(input)、出力(output)、アクセス(accessing)(例えば、メモリ中のデータにアクセスすること)した事を「判断」「決定」したとみなす事などを含み得る。また、「判断」、「決定」は、解決(resolving)、選択(selecting)、選定(choosing)、確立(establishing)、比較(comparing)などした事を「判断」「決定」したとみなす事を含み得る。つまり、「判断」「決定」は、何らかの動作を「判断」「決定」したとみなす事を含み得る。また、「判断(決定)」は、「想定する(assuming)」、「期待する(expecting)」、「みなす(considering)」などで読み替えられてもよい。
 本開示で使用する「に基づいて」という記載は、別段に明記されていない限り、「のみに基づいて」を意味しない。言い換えれば、「に基づいて」という記載は、「のみに基づいて」と「に少なくとも基づいて」の両方を意味する。
 本明細書で「第1の」、「第2の」などの呼称を使用した場合においては、その要素へのいかなる参照も、それらの要素の量または順序を全般的に限定するものではない。これらの呼称は、2つ以上の要素間を区別する便利な方法として本明細書で使用され得る。したがって、第1および第2の要素への参照は、2つの要素のみがそこで採用され得ること、または何らかの形で第1の要素が第2の要素に先行しなければならないことを意味しない。
 「含む(include)」、「含んでいる(including)」、およびそれらの変形が、本明細書あるいは特許請求の範囲で使用されている限り、これら用語は、用語「備える(comprising)」と同様に、包括的であることが意図される。さらに、本明細書あるいは特許請求の範囲において使用されている用語「または(or)」は、排他的論理和ではないことが意図される。
 本開示において、例えば、英語でのa, an及びtheのように、翻訳により冠詞が追加された場合、本開示は、これらの冠詞の後に続く名詞が複数形であることを含んでもよい。
 本開示において、「AとBが異なる」という用語は、「AとBが互いに異なる」ことを意味してもよい。なお、当該用語は、「AとBがそれぞれCと異なる」ことを意味してもよい。「離れる」、「結合される」などの用語も、「異なる」と同様に解釈されてもよい。
 10…区切り記号挿入装置、11…対象テキスト取得部、12…語間時間取得部、13…区切り記号挿入部、14…出力部、20…音声認識システム、21…音声認識結果取得部、22…音声認識結果出力部、M1…記録媒体、m10…メインモジュール、m11…対象テキスト取得モジュール、m12…語間時間取得モジュール、m13…区切り記号挿入モジュール、m14…出力モジュール、md1,md2…記号挿入モデル。

Claims (8)

  1.  発話音声の音声認識により得られたテキストに含まれる単語の後に、文を区切る区切り記号を挿入する区切り記号挿入装置であって、
     前記発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する語間時間取得部と、
     区切り記号挿入モデル及び前記語間時間に基づいて、前記発話音声の音声認識により得られたテキストである対象テキストに区切り記号を挿入する区切り記号挿入部であって、
      前記区切り記号挿入モデルは、
       前記区切り記号を含まない文である区切り記号除去文を少なくとも入力とし、
       前記区切り記号除去文に含まれる各単語の後に挿入される区切り記号を示す区切り記号推測情報を出力し、
       前記区切り記号除去文と区切り記号を含む文である区切り記号入り文とのペアを含む学習データを用いた機械学習により生成されるモデルであり、
      前記対象テキストを前記区切り記号除去文として前記区切り記号挿入モデルに入力することにより得られ、前記語間時間に応じて調整された前記区切り記号推測情報に基づいて、前記対象テキストに区切り記号を挿入する、区切り記号挿入部と、
     を備える区切り記号挿入装置。
  2.  前記区切り記号挿入部は、前記区切り記号挿入モデルから出力された前記区切り記号推測情報を、前記語間時間に基づいて調整する、
     請求項1に記載の区切り記号挿入装置。
  3.  前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号なし尤度を含み、
     前記区切り記号挿入部は、
     前記対象テキストに含まれる単語のうちの一の単語における前記語間時間が第1の時間である場合に、前記一の単語の前記記号なし尤度を上げるように調整し、又は/及び、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を下げるように調整し、
     前記一の単語における前記語間時間が前記第1の時間より長い第2の時間である場合に、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を上げるように調整し、又は/及び、前記一の単語の前記記号なし尤度を下げるように調整し、
     前記記号挿入尤度及び前記記号なし尤度のうちの最大の尤度に基づいて、前記一の単語の後に、複数の種類のうちのいずれかの区切り記号を挿入し、又は区切り記号を未挿入とする、
     請求項2に記載の区切り記号挿入装置。
  4.  前記区切り記号挿入部は、
     前記対象テキストに含まれる全単語のうちの一以上の単語の前記語間時間の長さの程度が、所与の程度より長い場合に、前記全単語のそれぞれの前記語間時間を所与の程度で短くなるように補正した補正語間時間に基づいて、前記記号挿入尤度及び/又は前記記号なし尤度を調整する、
     請求項3に記載の区切り記号挿入装置。
  5.  前記区切り記号挿入モデルは、
      前記区切り記号除去文に含まれる各単語の語間時間を入力として更に含み、
      前記語間時間が各単語に関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、
      前記語間時間により調整された前記区切り記号推測情報を出力し、
     前記区切り記号挿入部は、各単語に前記語間時間が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力する、
     請求項1に記載の区切り記号挿入装置。
  6.  前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号無し尤度を含む、
     請求項5に記載された区切り記号挿入装置。
  7.  前記区切り記号挿入モデルは、
      前記区切り記号除去文に対応する発話音声が発話されたときの状況を示すシチュエーション情報を入力として更に含み、
      前記語間時間が各単語に関連付けられると共に前記シチュエーション情報が関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、
     前記区切り記号挿入部は、前記シチュエーション情報が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力する、
     請求項5または6に記載の区切り記号挿入装置。
  8.  請求項1に記載の区切り記号挿入装置を含む音声認識システムであって、
     音声認識エンジンにより音声認識された結果のテキスト及び該テキストにおける各単語の前記語間時間を第1の音声認識結果として取得する音声認識結果取得部と、
     音声認識結果出力部と、を備え、
     前記区切り記号挿入装置の前記語間時間取得部は、前記第1の音声認識結果に含まれる前記語間時間を取得し、
     前記区切り記号挿入装置の前記区切り記号挿入部は、前記第1の音声認識結果に含まれるテキストを前記対象テキストとして、区切り記号を挿入し、
     前記音声認識結果出力部は、前記区切り記号挿入部により区切り記号が挿入された前記対象テキストを、第2の音声認識結果として出力する、
     を備える音声認識システム。
     
PCT/JP2023/017568 2022-08-05 2023-05-10 区切り記号挿入装置及び音声認識システム Ceased WO2024029152A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2024538828A JP7809817B2 (ja) 2022-08-05 2023-05-10 区切り記号挿入装置及び音声認識システム

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2022-125544 2022-08-05
JP2022125544 2022-08-05

Publications (1)

Publication Number Publication Date
WO2024029152A1 true WO2024029152A1 (ja) 2024-02-08

Family

ID=89849072

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2023/017568 Ceased WO2024029152A1 (ja) 2022-08-05 2023-05-10 区切り記号挿入装置及び音声認識システム

Country Status (2)

Country Link
JP (1) JP7809817B2 (ja)
WO (1) WO2024029152A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2009101837A1 (ja) * 2008-02-13 2009-08-20 Nec Corporation 記号挿入装置および記号挿入方法
JP2015219480A (ja) * 2014-05-21 2015-12-07 日本電信電話株式会社 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム
CN112927679A (zh) * 2021-02-07 2021-06-08 虫洞创新平台(深圳)有限公司 一种语音识别中添加标点符号的方法及语音识别装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2009101837A1 (ja) * 2008-02-13 2009-08-20 Nec Corporation 記号挿入装置および記号挿入方法
JP2015219480A (ja) * 2014-05-21 2015-12-07 日本電信電話株式会社 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム
CN112927679A (zh) * 2021-02-07 2021-06-08 虫洞创新平台(深圳)有限公司 一种语音识别中添加标点符号的方法及语音识别装置

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
TILK OTTOKAR, ALUMÄE TANEL: "LSTM for punctuation restoration in speech transcripts", INTERSPEECH 2015, ISCA, ISCA, 1 January 2015 (2015-01-01), ISCA, pages 683 - 687, XP093136514, DOI: 10.21437/Interspeech.2015-240 *

Also Published As

Publication number Publication date
JP7809817B2 (ja) 2026-02-02
JPWO2024029152A1 (ja) 2024-02-08

Similar Documents

Publication Publication Date Title
US8880398B1 (en) Localized speech recognition with offload
WO2020054451A1 (ja) 対話装置
US11663420B2 (en) Dialogue system
JP6745402B2 (ja) 質問推定装置
WO2024029152A1 (ja) 区切り記号挿入装置及び音声認識システム
JP7682862B2 (ja) 句点削除モデル学習装置、句点削除モデル及び判定装置
WO2020070943A1 (ja) パターン認識装置及び学習済みモデル
JPWO2019187463A1 (ja) 対話サーバ
JP7087095B2 (ja) 対話情報生成装置
JP6584622B1 (ja) 文章マッチングシステム
JPWO2019220791A1 (ja) 対話装置
JP7608603B2 (ja) 音声認識装置
JP6545855B1 (ja) 文章マッチングシステム
JP6960049B2 (ja) 対話装置
JP6895580B2 (ja) 対話システム
JP2022164001A (ja) 単言語変換装置
JP2024168531A (ja) 文生成モデル生成装置、文生成モデル及び文生成装置
WO2024203390A1 (ja) 音声認識誤り訂正装置
JP7093844B2 (ja) 対話システム
JP2024108744A (ja) 埋め込み表現生成システム
WO2026078751A1 (ja) 情報処理装置および情報処理方法
WO2024241734A1 (ja) 文脈文決定システム、機械翻訳装置および学習装置
WO2026078749A1 (ja) 情報処理装置および情報処理方法
JP7601649B2 (ja) 音声合成調整装置
US12260184B2 (en) Translation device

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23849715

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2024538828

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 18875751

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23849715

Country of ref document: EP

Kind code of ref document: A1