WO2024029152A1 - 区切り記号挿入装置及び音声認識システム - Google Patents
区切り記号挿入装置及び音声認識システム Download PDFInfo
- Publication number
- WO2024029152A1 WO2024029152A1 PCT/JP2023/017568 JP2023017568W WO2024029152A1 WO 2024029152 A1 WO2024029152 A1 WO 2024029152A1 JP 2023017568 W JP2023017568 W JP 2023017568W WO 2024029152 A1 WO2024029152 A1 WO 2024029152A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- delimiter
- time
- interword
- likelihood
- word
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/04—Segmentation; Word boundary detection
Definitions
- the present invention relates to a delimiter insertion device and a speech recognition system.
- Patent Document 1 discloses a technique in which punctuation marks are inserted into text by an engine trained using training data in the form of text with punctuation marks added based on statistics.
- delimiters such as punctuation marks are inserted at positions appropriate for the sentence.
- the insertion position of the delimiter may not be incorrect in the sentence, but may be different from the speaker's intention.
- the present invention has been made in view of the above problems, and an object of the present invention is to insert a delimiter at the position intended by the speaker in text obtained by speech recognition processing of spoken voice. shall be.
- a delimiter insertion device that inserts a delimiter that separates sentences after a word included in a text obtained by voice recognition of spoken voice. and an interword time acquisition unit that obtains an interword time that is the length of time until the next word is uttered in each word included in the uttered speech, and a delimiter insertion model and an interword time based on the delimiter insertion model and the interword time.
- the delimiter insertion unit inserts a delimiter into a target text that is a text obtained by speech recognition of spoken speech, and the delimiter insertion model is configured to at least insert a delimiter removal sentence that does not include a delimiter.
- the delimiter insertion model is configured to at least insert a delimiter removal sentence that does not include a delimiter.
- This model is generated by machine learning using training data, and is based on delimiter inference information that is obtained by inputting the target text as a delimiter removed sentence into a delimiter insertion model and adjusted according to the interword time. and a delimiter insertion unit for inserting a delimiter into the target text.
- the interword time in the uttered speech is obtained.
- the interword time reflects the speaker's intention at the time of utterance.
- delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
- the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
- FIG. 1 is a block diagram showing the functional configuration of a delimiter insertion device according to the present embodiment.
- FIG. 2 is a hardware block diagram of a delimiter insertion device. It is a figure explaining the problem solved by the delimiter insertion device of this embodiment. It is a figure which shows the acquisition process of target text.
- FIG. 3 is a diagram showing a first example of learning data used for machine learning of a delimiter insertion model.
- FIG. 3 is a diagram showing a first example of the configuration of a delimiter insertion model.
- FIG. 7 is a diagram illustrating an example of adjustment rule information that is referred to in order to adjust delimiter guess information based on interword time.
- FIG. 7 is a diagram illustrating an example of adjustment processing of delimiter guess information.
- FIG. 3 is a diagram illustrating an example of inter-word time correction processing.
- FIG. 7 is a diagram showing a second example of learning data used for machine learning of the delimiter insertion model.
- FIG. 7 is a diagram illustrating a second example of the configuration of a delimiter insertion model.
- FIG. 7 is a diagram showing a third example of learning data used for machine learning of the delimiter insertion model.
- FIG. 1 is a functional block diagram showing an example of the configuration of a speech recognition system according to the present embodiment.
- 3 is a flowchart showing processing details of a delimiter insertion method in the delimiter insertion device.
- FIG. 2 is a diagram showing the configuration of a delimiter insertion program.
- FIG. 1 is a diagram showing the functional configuration of a delimiter insertion device according to this embodiment.
- the delimiter insertion device of this embodiment is a device that inserts a delimiter for delimiting a sentence after a word included in a text obtained by voice recognition of spoken voice.
- the delimiter insertion device 10 inserts delimiters such as a period, a comma, and a question mark, which are inserted into an English sentence, into a text consisting of an English sentence.
- delimiters such as a period, a comma, and a question mark
- the delimiter insertion device 10 may be a device that inserts delimiters such as periods and commas into sentences in Japanese, or a device that inserts delimiters in that language into sentences in other languages. good.
- the delimiter insertion device 10 functionally includes a text acquisition section 11, an interword time acquisition section 12, a delimiter insertion section 13, and an output section 14.
- Each of these functional units 11 to 14 may be configured in one device, or may be configured in a distributed manner in a plurality of devices.
- each functional block may be realized using one physically or logically coupled device, or may be realized using two or more physically or logically separated devices directly or indirectly (e.g. , wired, wireless, etc.) and may be realized using a plurality of these devices.
- the functional block may be realized by combining software with the one device or the plurality of devices.
- Functions include judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, These include, but are not limited to, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assigning. I can't.
- a functional block (configuration unit) that performs transmission is called a transmitting unit or a transmitter. In either case, as described above, the implementation method is not particularly limited.
- the delimiter insertion device 10 in one embodiment of the present invention may function as a computer.
- FIG. 2 is a diagram showing an example of the hardware configuration of the delimiter insertion device 10 according to this embodiment.
- the delimiter insertion device 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
- the word “apparatus” can be read as a circuit, a device, a unit, etc.
- the hardware configuration of the delimiter insertion device 10 may be configured to include one or more of the devices shown in the figure, or may be configured without including some of the devices.
- Each function in the delimiter insertion device 10 is achieved by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, so that the processor 1001 performs calculations, and the communication by the communication device 1004 and the memory 1002 and This is achieved by controlling reading and/or writing of data in the storage 1003.
- the processor 1001 for example, operates an operating system to control the entire computer.
- the processor 1001 may be configured with a central processing unit (CPU) that includes interfaces with peripheral devices, a control device, an arithmetic device, registers, and the like.
- CPU central processing unit
- each of the functional units 11 to 14 shown in FIG. 1 may be implemented by the processor 1001.
- the processor 1001 reads programs (program codes), software modules, and data from the storage 1003 and/or the communication device 1004 to the memory 1002, and executes various processes in accordance with these.
- the program a program that causes a computer to execute at least part of the operations described in the above embodiments is used.
- each of the functional units 11 to 15 of the delimiter insertion device 10 may be realized by a control program stored in the memory 1002 and operated on the processor 1001.
- Processor 1001 may be implemented with one or more chips. Note that the program may be transmitted from a network via a telecommunications line.
- the memory 1002 is a computer-readable recording medium, and includes at least one of ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. may be done.
- Memory 1002 may be called a register, cache, main memory, or the like.
- the memory 1002 can store executable programs (program codes), software modules, and the like to implement the pseudo data generation method and sentence generation method according to an embodiment of the present invention.
- the storage 1003 is a computer-readable recording medium, such as an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, or a magneto-optical disk (for example, a compact disk, a digital versatile disk, or a Blu-ray disk). (registered trademark) disk), smart card, flash memory (eg, card, stick, key drive), floppy disk, magnetic strip, etc.
- Storage 1003 may also be called an auxiliary storage device.
- the storage medium mentioned above may be, for example, a database including memory 1002 and/or storage 1003, a server, or other suitable medium.
- the communication device 1004 is hardware (transmission/reception device) for communicating between computers via a wired and/or wireless network, and is also referred to as a network device, network controller, network card, communication module, etc., for example.
- the input device 1005 is an input device (eg, keyboard, mouse, microphone, switch, button, sensor, etc.) that accepts input from the outside.
- the output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that performs output to the outside. Note that the input device 1005 and the output device 1006 may have an integrated configuration (for example, a touch panel).
- each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information.
- the bus 1007 may be configured as a single bus or may be configured as different buses between devices.
- the delimiter insertion device 10 also uses hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), and a field programmable gate array (FPGA). A part or all of each functional block may be realized by the hardware. For example, processor 1001 may be implemented with at least one of these hardware.
- DSP digital signal processor
- ASIC application specific integrated circuit
- PLD programmable logic device
- FPGA field programmable gate array
- the problem solved by the delimiter insertion device 10 of this embodiment will be explained with reference to FIG. 3.
- the utterances sp11 and sp21 shown in FIG. 3 are uttered by speakers with different intentions, respectively.
- the speech recognition results sr11 and sr21 obtained by speech recognition of the uttered speech sp11 and sp21 are the same text "I know it's been there forever".
- the interword time until the next word after the word "know” is uttered is 1.2 seconds.
- the interword time until the next word after the word "know” is uttered is 0.1 seconds.
- the delimiter is inserted into the text obtained as a speech recognition result at a position that is appropriate for a sentence.
- the speech recognition results sr12 and sr22 obtained by inserting the delimiter are the same even though the original speech sounds are different.
- the speech recognition result sr22 has a period after the word "forever”, as intended by the speaker of the uttered speech sp21.
- the speech recognition result sr12 has a period after the word "forever.” The position of this period is different from the position intended by the speaker.
- the delimiter insertion device 10 of this embodiment inserts a delimiter using the interword time in the uttered speech, so the word "know” and the word A period is inserted after each word "forever” to obtain the speech recognition result sr13.
- the delimiter insertion device 10 inserts a period after the word "forever” into the speech recognition result sr21 obtained based on the uttered speech sp21 to obtain a speech recognition result sr23.
- the speech recognition results sr13 and sr23 have delimiters at the positions intended by the speakers of the uttered speech sp11 and sp21.
- the text acquisition unit 11 acquires target text, which is the text into which a delimiter is to be inserted.
- the target text is text obtained by voice recognition of spoken voice.
- the interword time acquisition unit 12 acquires the interword time, which is the length of time until the next word is uttered in each word included in the uttered voice.
- FIG. 4 is a diagram showing an example of acquiring the target text and interword time.
- the text acquisition unit 11 acquires the target text tx1 "I know it's been there forever" based on the speech sp3.
- the text acquisition unit 11 may acquire the target text by performing voice recognition of the uttered voice using a well-known voice recognition processing technique and other techniques.
- the interword time acquisition unit 12 acquires the length of time until the next word is uttered in each word included in the target text as the interword time it.
- the inter-word time acquisition unit 12 may acquire the silent time during speech recognition of the uttered speech sp3 as the inter-word time it.
- the silent time is, for example, a time when the volume is less than a predetermined level.
- the interword time acquisition unit 12 sequentially acquires voice recognition results every time voice recognition is performed from a voice recognition engine used for voice recognition of uttered speech, and sets the time interval for acquiring and updating the voice recognition results between words.
- the update time interval may be obtained as the interword time it by regarding it as a pseudo silent time between the occurrences of .
- the delimiter insertion unit 13 inserts delimiters into the target text based on the delimiter insertion model and the interword time. Specifically, the delimiter insertion unit 13 inserts a delimiter into the target text based on delimiter guess information obtained by inputting the target text into a delimiter insertion model.
- the delimiter guess information includes information adjusted according to interword time.
- the delimiter insertion model receives at least a delimiter removed sentence, which is a sentence that does not include a delimiter, as input, and outputs delimiter guess information indicating the delimiter to be inserted after each word included in the delimiter removed sentence. Further, the delimiter insertion model is generated by machine learning using learning data including a pair of a delimiter-removed sentence and a delimiter-included sentence, which is a sentence including the delimiter.
- FIG. 5 is a diagram showing a first example of learning data used for machine learning of the delimiter insertion model.
- FIG. 6 is a diagram showing a first example of the configuration of a delimiter insertion model.
- the learning data td1 which is an example of learning data used for machine learning of the delimiter insertion model md1, consists of a pair of a delimiter-removed sentence id1 and a delimiter-added sentence od1.
- the delimiter-containing sentence od1 includes a word string making up the sentence and a delimiter label that is a label indicating a delimiter inserted after each word.
- Labels indicating delimiters are schematically illustrated in FIG. 5 and the like as follows. ⁇ O>...No delimiter ⁇ P>...Period, full stop ⁇ C>...Comma, comma ⁇ Q>...Question mark, question mark
- the delimiter removed sentence id1 is a sentence with the delimiter label removed from the delimiter-containing sentence od1. There may be.
- the delimiter removed sentence id1 is input to the delimiter insertion model md1 in the learning process, and the output obtained from the delimiter insertion model md1 is combined with the delimiter-containing sentence od1, which is the training data. Based on the error, the weights, parameters, etc. that make up the delimiter insertion model md1 are updated.
- the trained delimiter insertion model md1 outputs delimiter guess information dp1 in response to the input of the delimiter removed sentence sd1.
- the delimiter insertion model md1 may be a model that includes a neural network. More specifically, the delimiter insertion model md1 may be configured as a sequence labeling model that solves a sequence labeling task of predicting a delimiter to be inserted after each word included in an input sentence.
- the delimiter insertion model md1 which is a model that includes a trained neural network, can be read or referenced by a computer, and can be regarded as a program that causes the computer to perform a predetermined process and realize a predetermined function.
- the trained delimiter insertion model md1 of this embodiment is used in a computer equipped with a CPU and memory. Specifically, the CPU of the computer assigns learned weights corresponding to each layer to the input data input to the input layer of the neural network according to instructions from the learned delimiter insertion model md1 stored in memory. It operates to perform calculations based on coefficients (parameters), response functions, etc., and output the results (probabilities) from the output layer.
- the delimiter guess information dp1 includes the symbol insertion likelihood, which is the likelihood of various delimiters that can be inserted after each word included in the delimiter removed sentence sd1, and the fact that no delimiter is inserted after each word. Contains the unsymbol likelihood, which is the likelihood for . Then, based on the maximum likelihood of the symbol insertion likelihood and the no symbol likelihood, insert one of a plurality of types of delimiters after each word, or insert no delimiter. This is determined (labeled).
- the delimiter guess information dp1 illustrated in FIG. 6 includes the symbol insertion likelihood and symbol-free likelihood of each delimiter regarding the word "I", as described below. ⁇ O>: 90%, ⁇ C>: 5%, ⁇ P>: 2%, ⁇ Q>: 3% Therefore, since the symbol-less likelihood of not inserting a delimiter (label ⁇ O>) is the maximum, the word "I" is labeled with the label ⁇ O> without a delimiter.
- the delimiter guess information dp1 includes the likelihood of symbol insertion and the likelihood of no symbol for each delimiter regarding the word "know”, as shown below.
- the delimiter insertion unit 13 may adjust the delimiter inference information output from the delimiter insertion model based on the interword time, as an example of adjusting the delimiter inference information based on the interword time. Specifically, the delimiter insertion unit 13 may adjust the likelihood of symbol insertion and the likelihood of no symbol included in the delimiter estimation information based on the interword time.
- the delimiter insertion unit 13 adjusts to increase the likelihood of no symbol for one word included in the target text when the interword time in one of the words is the first time, or / and adjusting to lower the symbol insertion likelihood of at least one kind of delimiter of the one word out of the plurality of kinds of delimiters, and the first one whose interword time in the one word is longer than the first time. 2, adjustment is made to increase the likelihood of symbol insertion of at least one of the plurality of types of punctuation marks for the one word, and/or no symbol for the one word. It may be adjusted to lower the likelihood.
- the delimiter insertion unit 13 inserts one of the plurality of delimiters after the one word based on the maximum likelihood of the adjusted symbol insertion likelihood and the no symbol likelihood. or leave the delimiter uninserted.
- the delimiter insertion unit 13 may, for example, adjust the likelihood by referring to adjustment rule information.
- FIG. 7 is a diagram illustrating an example of adjustment rule information that is referred to in order to adjust delimiter guess information based on interword time.
- the adjustment rule information may be stored in a storage device that is accessible to the delimiter insertion section 13, or may be provided as a table inside the delimiter insertion section 13.
- the adjustment rule information is such that, for each range of interword time, an interword time category indicating the length of the interword time and likelihood adjustment information indicating the contents of the likelihood adjustment are associated. It is information. For example, if the interword time (x) is 0.1 seconds or less, the interword time category is "none", and “increase the no symbol likelihood by 50%" is performed as a likelihood adjustment. This is stipulated as an adjustment rule. Also, for example, if the interword time (x) is longer than 0.5 seconds and less than 1.0 seconds, the interword time category is "medium” and the likelihood of symbol insertion is increased by 50%. It is stipulated as an adjustment rule that this is carried out as a likelihood adjustment.
- FIG. 8 is a diagram illustrating an example of adjustment processing of delimiter guess information.
- the delimiter guess information dp21 shown in FIG. 8 shows a part of the delimiter guess information before the adjustment process that is output from the delimiter insertion model md1.
- the delimiter guess information dp21 includes a symbol insertion likelihood lh21 and a symbol-free likelihood of each delimiter regarding the word "know.” According to the pre-adjustment delimiter guess information dp21, the likelihood of not inserting a delimiter (label ⁇ O>) is the highest, so the label ⁇ O> without a delimiter for the word "know" is O> is labeled.
- the delimiter insertion unit 13 adjusts each likelihood of the delimiter guess information dp21 based on the interword time it2 of the word "know". Specifically, since the interword time it2 of the word "know" acquired by the interword time acquisition unit 12 is 0.9 seconds, the delimiter insertion unit 13 refers to the adjustment rule information (FIG. 7). Then, the likelihood adjustment information "Increase the likelihood of symbol insertion for periods, commas, and question marks by 50%" associated with the interword time of 0.9 seconds is obtained, and the symbol insertion likelihood for commas, periods, and question marks lh21 is obtained. is adjusted according to the obtained likelihood adjustment information.
- the delimiter insertion unit 13 increases each value of the symbol insertion likelihood lh21 by 50% to obtain adjusted delimiter guess information dp22.
- the symbol insertion likelihood of a period (label ⁇ P>) is the maximum, so the delimiter insertion unit 13 inserts the label ⁇ P> of the delimiter "period" into the word "know". Label. Then, the delimiter insertion unit 13 inserts a period after the word "know" included in the target text based on the labeled label ⁇ P>.
- the longer the interword time of one word included in the target text the higher the likelihood of symbol insertion and/or the lower the likelihood of no symbol, so that the longer the interword time of one word
- the shorter the interword time the lower the likelihood of symbol insertion and/or the higher the likelihood of no symbol. reflected in the likelihood.
- the delimiter is inserted or not inserted, so the delimiter is inserted at an appropriate position according to the speaker's intention. It is possible to obtain the text.
- the delimiter insertion unit 13 may correct the interword time for use in adjusting the delimiter guess information according to predetermined conditions.
- FIG. 9 is a diagram illustrating an example of inter-word time correction processing.
- the delimiter insertion unit 13 inserts the interword time of each of all the words. It may be corrected to shorten it to a given degree. Then, the delimiter insertion unit 13 may adjust the likelihood of symbol insertion and/or the likelihood of no symbol based on the corrected interword time, which is the corrected interword time.
- the target text tx31 shown in FIG. 9 includes an interword time it31 before correction.
- the interword time of more than half of all the words included in the target text is equal to or greater than the interword time corresponding to the interword time category "small"
- the content of the correction process is set in advance to reduce the interword time by 0.5 seconds.
- the delimiter insertion unit 13 determines that the interword time of all words among the interword time it31 of words included in the target text tx31 is equal to or greater than the interword time corresponding to the interword time category "small". determine something. Then, as shown in the target text tx32, the delimiter insertion unit 13 subtracts 0.5 seconds from each interword time it31 to obtain a corrected interword time it32, which is the corrected interword time. The delimiter insertion unit 13 adjusts the likelihood of symbol insertion and/or the likelihood of no symbol of each word included in the target text tx32 based on the corrected interword time it32.
- the delimiter may be inserted excessively in a position that the speaker did not intend.
- the interword time is longer than a predetermined value, a symbol is created based on the corrected interword time that is corrected to shorten the interword time. Since the insertion likelihood and/or the symbol-less likelihood are adjusted, it is possible to obtain a text in which the delimiter is inserted at an appropriate position that suitably reflects the speaker's intention.
- FIG. 10 is a diagram showing a second example of learning data used for machine learning of the delimiter insertion model.
- FIG. 11 is a diagram showing a second example of the configuration of the delimiter insertion model.
- the learning data td2 which is an example of learning data used for machine learning of the delimiter insertion model md2, consists of a pair of a delimiter-removed sentence id2 and a delimiter-added sentence od2.
- the delimiter-containing sentence od2 like the delimiter-containing sentence od1 described with reference to FIG. 5, includes a word string composing the sentence and a delimiter label indicating the delimiter to be inserted after each word.
- the delimiter-removed sentence id2 includes a sentence obtained by removing the delimiter label from the delimiter-containing sentence od2, and interword time it4 associated with each word constituting the sentence.
- the delimiter removed sentence id2 is input to the delimiter insertion model md2 in the learning process, and the output obtained from the delimiter insertion model md2 is combined with the delimiter-containing sentence od2, which is the training data. Based on the error, the weights, parameters, etc. that make up the delimiter insertion model md2 are updated.
- the delimiter insertion model md2 may be a model including a neural network. More specifically, the delimiter insertion model md2 may be configured as a sequence labeling model that solves a sequence labeling task of predicting a delimiter to be inserted after each word included in an input sentence.
- the trained delimiter insertion model md2 outputs delimiter guess information dp2 in response to the input of the delimiter removed sentence sd2.
- the delimiter removed sentence sd2 includes an interword time it5 associated with each word forming the delimiter removed sentence sd2.
- the delimiter guess information dp2 includes the likelihood of symbol insertion and the likelihood of no symbol for each word included in the delimiter removed sentence sd2.
- the delimiter insertion model md2 which is a model that includes a trained neural network, can be read or referenced by a computer, and can be regarded as a program that causes the computer to perform a predetermined process and realize a predetermined function.
- the trained delimiter insertion model md2 of this embodiment is used in a computer equipped with a CPU and memory. Specifically, the CPU of the computer assigns learned weights corresponding to each layer to the input data input to the input layer of the neural network according to instructions from the learned delimiter insertion model md2 stored in memory. It operates to perform calculations based on coefficients (parameters), response functions, etc., and output the results (probabilities) from the output layer.
- the symbol insertion likelihood and symbol-free likelihood included in the delimiter guess information dp1 output from the delimiter insertion model md1 are adjusted based on the interword time.
- the delimiter removed sentence id2 including the interword time it4 is used as input as learning data in machine learning, and in response to the input of the delimiter removed sentence sd2 including the interword time it5. Since the delimiter guess information dp2 is output, the symbol insertion likelihood and the no symbol likelihood included in the delimiter guess information dp2 are values adjusted according to the interword time by calculations in the delimiter insertion model md2. It is.
- the likelihood of symbol insertion and the likelihood of no symbol for each delimiter regarding the word "know” are calculated as follows. ⁇ O>: 10%, ⁇ C>: 20%, ⁇ P>: 70%, ⁇ Q>: 0%
- the maximum likelihood of no delimiter was calculated by the delimiter insertion model md2. For each likelihood, the symbol insertion likelihood of inserting a period (label ⁇ P>) is the largest.
- the delimiter insertion model md2 is generated by machine learning using delimiter-removed sentences with interword times associated with each word as training data, and delimiter-removed sentences with interword times associated with each word are generated by machine learning. Since it is input to the delimiter insertion model md2, by inputting the target text in which interword time is associated with each word to the delimiter insertion model, a separate It is possible to obtain delimiter guess information adjusted by the interword time without performing adjustment processing. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to easily obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
- FIG. 12 is a diagram showing a third example of learning data used for machine learning of the delimiter insertion model.
- the learning data td3 consists of a pair of a delimiter-removed sentence id3 and a delimiter-added sentence od3.
- the delimiter-containing sentence od3 like the delimiter-containing sentences od1 and od2 described with reference to FIGS. 5 and 10, has a delimiter label indicating the word string that makes up the sentence and the delimiter inserted after each word. including.
- the delimiter-removed sentence id3 includes a sentence from which the delimiter label has been removed from the delimiter-containing sentence od3, interword times associated with each word constituting the sentence, and situation information st3.
- the situation information st3 is information indicating the situation when the uttered voice is uttered, and may be added as a tag to the delimiter removed sentence id3.
- the situation information st3 may have the following variations depending on the situation when the uttered voice is uttered. Meeting: ⁇ Meeting> Lecture: ⁇ Lecture> Chatting: ⁇ Chatting>
- the delimiter removed sentence id3 is input to the delimiter insertion model in the learning process, and the output obtained from the delimiter insertion model and the delimiter-containing sentence that is the training data are combined. Based on the error with od3, the weights, parameters, etc. that make up the delimiter insertion model are updated.
- the delimiter insertion model which has been trained by machine learning using learning data td3, generates delimiter inference information in response to the input of a delimiter removed sentence (target text) that includes situation information and interword time associated with each word. Output.
- the output delimiter estimation information includes the likelihood of symbol insertion and the likelihood of no symbol for each word included in the delimiter removed sentence, as well as a delimiter based on the maximum likelihood or a label of no delimiter inserted.
- the delimiter insertion unit 13 inserts a delimiter according to the label after each word based on the delimiter guess information obtained by inputting the target text associated with the situation information into the delimiter insertion model. or decide not to insert the delimiter.
- a delimiter insertion model is generated by machine learning using delimiter removal sentences associated with situation information as learning data, and delimiter removal sentences including situation information are used as input to the delimiter insertion model. Therefore, by inputting the target text associated with situation information into the delimiter insertion model, it is possible to obtain delimiter estimation information that takes into account the utterance tendency according to the situation when the uttered voice is uttered. Become.
- the output unit 14 outputs the delimiter inserted text, which is the target text into which the delimiter has been inserted by the delimiter insertion unit 13.
- the mode of output is not limited, and the output unit 14 may display the delimiter insertion text on a predetermined display means, store the delimiter insertion text in a predetermined storage means, or display the delimiter insertion text on a predetermined device. You may also send delimited text to .
- FIG. 13 is a functional block diagram showing an example of the configuration of the speech recognition system of this embodiment.
- the speech recognition system 20 includes a delimiter insertion device 10, and includes a speech recognition result acquisition section 21 and a speech recognition result output section 22.
- the speech recognition result acquisition unit 21 obtains the text resulting from speech recognition by the speech recognition engine and the interword time of each word in the text as a first speech recognition result.
- the text acquisition unit 11 of the delimiter insertion device 10 acquires the text included in the first speech recognition result as the target text.
- the interword time acquisition unit 12 of the delimiter insertion device 10 acquires the interword time included in the first speech recognition result.
- the delimiter insertion unit 13 of the delimiter insertion device 10 performs a process of inserting a delimiter using the text included in the first speech recognition result as the target text.
- the speech recognition result output unit 22 outputs the target text into which the delimiter has been inserted by the delimiter insertion unit 13 as a second speech recognition result.
- the speech recognition system 20 based on the first speech recognition result obtained from a speech recognition engine that does not take interword time into consideration in speech recognition processing, the text included in the first speech recognition result is set as the target text, Since it is possible to obtain delimiter estimation information adjusted according to the interword time included in the first speech recognition result, the second speech recognition that includes the target text with a delimiter inserted at a position that matches the speaker's intention can be performed. It becomes possible to obtain results.
- FIG. 14 is a flowchart showing the processing details of the delimiter insertion method in the delimiter insertion device 10.
- step S1 the text acquisition unit 11 acquires target text, which is the text obtained by voice recognition of the uttered voice and is the text into which a delimiter is to be inserted.
- step S2 the interword time acquisition unit 12 acquires the interword time of each word in the uttered speech.
- step S3 the delimiter insertion unit 13 inputs the target text into the delimiter insertion model.
- step S4 the delimiter insertion unit 13 acquires delimiter guess information.
- the delimiter estimation information is based on the symbol insertion likelihood and no symbol likelihood of each word adjusted according to the interword time, and determines whether any delimiter is inserted or no delimiter is inserted for each word. Contains a label indicating that.
- step S5 the delimiter insertion unit 13 inserts a delimiter into the target text based on the delimiter estimation information.
- step S6 the output unit 14 outputs the delimiter inserted text, which is the target text in which the delimiter has been inserted.
- FIG. 15 is a diagram showing the configuration of the delimiter insertion program.
- the delimiter insertion program P1 includes a main module m10 that centrally controls the delimiter insertion process in the delimiter insertion device 10, a text acquisition module m11, an interword time acquisition module m12, a delimiter insertion module m13, and an output module m14. It consists of Each of the modules m11 to m14 implements the functions of the text acquisition section 11, the interword time acquisition section 12, the delimiter insertion section 13, and the output section 14.
- the delimiter insertion program P1 may be transmitted via a transmission medium such as a communication line, or may be stored in a recording medium M1 as shown in FIG. .
- the interword time in the uttered voice is acquired.
- the interword time reflects the speaker's intention at the time of utterance.
- delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
- the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention.
- by inserting a delimiter into the target text based on the delimiter estimation information it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
- a delimiter insertion device is a delimiter insertion device that inserts a delimiter for delimiting a sentence after a word included in a text obtained by voice recognition of an uttered voice, an interword time acquisition unit that acquires an interword time that is the length of time until the next word in each word included in the utterance is uttered;
- a delimiter insertion unit that inserts a delimiter into a target text that is a text obtained by voice recognition of spoken speech, the delimiter insertion model including at least a delimiter removed sentence that does not include the delimiter.
- output delimiter guess information indicating a delimiter to be inserted after each word included in the delimiter-removed sentence, and pair the delimiter-removed sentence with a delimiter-containing sentence that is a sentence including the delimiter.
- the model is generated by machine learning using learning data containing the delimiter, and is obtained by inputting the target text as the delimiter removed sentence to the delimiter insertion model, and is adjusted according to the interword time. and a delimiter insertion unit that inserts a delimiter into the target text based on the delimiter estimation information.
- the interword time in the uttered speech is obtained.
- the interword time reflects the speaker's intention at the time of utterance.
- delimiter guess information obtained by inputting the target text into the delimiter insertion model and adjusted according to the interword information is obtained.
- the delimiter estimation information acquired here is information adjusted according to the interword time of each word in the uttered voice, and therefore indicates a delimiter that reflects the speaker's intention. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
- the delimiter insertion unit may insert the delimiter estimation information output from the delimiter insertion model into the delimiter estimation information outputted from the delimiter insertion model. Adjustments may also be made based on time.
- the delimiter estimation information includes a plurality of types of delimiters after each word included in the delimiter removed sentence.
- the delimiter insertion section includes a symbol insertion likelihood, which is the likelihood of inserting each word, and a symbol-free likelihood, which is the likelihood of not inserting a delimiter after each word. If the interword time in one of the words included in the text is a first time, the one word is adjusted to increase the likelihood of no symbol, and/or the plurality of types are adjusted.
- the symbol insertion likelihood of at least one of the delimiters of the one word to reduce the symbol insertion likelihood, and the interword time in the one word is longer than the first time; , the likelihood of at least one of a plurality of types of punctuation marks for the one word is adjusted to increase the likelihood of the symbol being inserted, and/or the likelihood of the one word not having the symbol is adjusted. and insert one of a plurality of types of delimiters after the one word based on the maximum likelihood of the symbol insertion likelihood and the no symbol likelihood. Alternatively, the delimiter may not be inserted.
- the delimiter is inserted or not inserted, so the delimiter is inserted at an appropriate position according to the speaker's intention. It is possible to obtain the text.
- the delimiter inserting section includes a delimiter inserting unit configured to insert a delimiter between words of one or more of all words included in the target text.
- a delimiter inserting unit configured to insert a delimiter between words of one or more of all words included in the target text.
- the delimiter may be inserted excessively in a position that the speaker did not intend.
- the symbol insertion likelihood and/or since the likelihood without symbols is adjusted, it is possible to obtain a text in which delimiters are inserted at appropriate positions that suitably reflect the speaker's intention.
- the delimiter insertion model further includes an interword time of each word included in the delimiter removed sentence as an input.
- the delimiter is generated by machine learning using learning data consisting of a pair of the delimiter-removed sentence and the delimiter-included sentence whose interword time is associated with each word, and the delimiter is adjusted by the interword time.
- the delimiter insertion unit may output guess information and input the target text in which each word is associated with the interword time into the delimiter insertion model.
- a delimiter insertion model is generated by machine learning in which delimiter removal sentences with interword time associated with each word are used as training data, and delimiter removal including interword time of each word is generated by machine learning. Since a sentence is taken as the input of the delimiter insertion model, by inputting the target text with an associated interword time for each word into the delimiter insertion model, a separate It becomes possible to obtain delimiter guess information adjusted by the interword time without performing adjustment processing. Then, by inserting a delimiter into the target text based on the delimiter estimation information, it becomes possible to easily obtain a text in which the delimiter is inserted at a position that meets the speaker's intention.
- the delimiter estimation information includes a plurality of types of delimiters after each word included in the delimiter removed sentence. It is also possible to include a symbol insertion likelihood, which is the likelihood of inserting each word, and a symbol-free likelihood, which is the likelihood of not inserting a delimiter after each word.
- delimiter guessing information including the symbol insertion likelihood and symbol-absence likelihood adjusted by the interword time.
- the delimiter insertion model is based on a situation when the speech sound corresponding to the delimiter removed sentence is uttered.
- the machine further includes as input situation information indicating that the interword time is associated with each word, and the machine uses learning data consisting of a pair of the delimiter-removed sentence and the delimiter-included sentence, to which the interword time is associated with each word and the situation information is associated.
- the delimiter insertion unit may input the target text associated with the situation information to the delimiter insertion model, which is generated by learning.
- a delimiter insertion model is generated by machine learning using delimiter removal sentences associated with situation information as learning data, and delimiter removal sentences including situation information are input to the delimiter insertion model. Therefore, by inputting the target text associated with situation information into the delimiter insertion model, it is possible to obtain delimiter estimation information that takes into account the tendency of speech according to the situation when the speech is uttered. becomes possible.
- the speech recognition system includes the delimiter insertion device according to any one of the first to seventh aspects, and includes a text as a result of speech recognition by the speech recognition engine and each word in the text.
- the interword time acquisition unit of the delimiter insertion device includes a speech recognition result acquisition unit that obtains the interword time as a first speech recognition result, and a speech recognition result output unit, and the interword time acquisition unit of the delimiter insertion device Obtaining the interword time included in the recognition result, the delimiter insertion unit of the delimiter insertion device inserts a delimiter using the text included in the first speech recognition result as the target text, and
- the speech recognition result output section may output the target text into which the delimiter has been inserted by the delimiter insertion section as the second speech recognition result.
- the text included in the first speech recognition result is set as the target text. Since it is possible to obtain delimiter estimation information adjusted according to the interword time included in the first speech recognition result, the second speech recognition result includes the target text with the delimiter inserted at a position that matches the speaker's intention. It becomes possible to obtain.
- the notification of information may include physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, It may be implemented using broadcast information (MIB (Master Information Block), SIB (System Information Block)), other signals, or a combination thereof.
- RRC signaling may be called an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
- LTE Long Term Evolution
- LTE-A Long Term Evolution-Advanced
- SUPER 3G IMT-Advanced
- 4G 5G
- FRA Full Radio Access
- W-CDMA Wideband Code Division Multiple Access
- GSM registered trademark
- CDMA2000 Code Division Multiple Access 2000
- UMB Universal Mobile Broadband
- IEEE 802.11 Wi-Fi
- IEEE 802.16 WiMAX
- IEEE 802.20 UWB (Ultra-WideBand)
- the present invention may be applied to systems utilizing Bluetooth (registered trademark), other suitable systems, and/or next-generation systems extended based thereon.
- a combination of a plurality of systems may be applied (for example, a combination of at least one of LTE and LTE-A and 5G).
- the specific operations performed by the base station in this disclosure may be performed by its upper node.
- various operations performed for communication with a terminal are performed by the base station and other network nodes other than the base station (e.g., MME or It is clear that this could be done by at least one of the following: (conceivable, but not limited to) S-GW, etc.).
- MME mobile phone
- S-GW network node
- Information can be output from the upper layer (or lower layer) to the lower layer (or upper layer). It may be input/output via multiple network nodes.
- the input/output information may be stored in a specific location (for example, memory) or may be managed in a management table. Information etc. to be input/output may be overwritten, updated, or additionally written. The output information etc. may be deleted. The input information etc. may be transmitted to other devices.
- Judgment may be made using a value expressed by 1 bit (0 or 1), a truth value (Boolean: true or false), or a comparison of numerical values (for example, a predetermined value). (comparison with a value).
- notification of prescribed information is not limited to being done explicitly, but may also be done implicitly (for example, not notifying the prescribed information). Good too.
- Software includes instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, whether referred to as software, firmware, middleware, microcode, hardware description language, or by any other name. , should be broadly construed to mean an application, software application, software package, routine, subroutine, object, executable, thread of execution, procedure, function, etc.
- software, instructions, etc. may be sent and received via a transmission medium.
- a transmission medium For example, if the software uses wired technologies such as coaxial cable, fiber optic cable, twisted pair and digital subscriber line (DSL) and/or wireless technologies such as infrared, radio and microwave to When transmitted from a remote source, these wired and/or wireless technologies are included within the definition of transmission medium.
- wired technologies such as coaxial cable, fiber optic cable, twisted pair and digital subscriber line (DSL) and/or wireless technologies such as infrared, radio and microwave
- data, instructions, commands, information, signals, bits, symbols, chips, etc. which may be referred to throughout the above description, may refer to voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, light fields or photons, or any of these. It may also be represented by a combination of
- system and “network” are used interchangeably.
- radio resources may be indicated by an index.
- determining may encompass a wide variety of operations.
- “Judgment” and “decision” include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, search, and inquiry. (e.g., searching in a table, database, or other data structure), and regarding an ascertaining as a “judgment” or “decision.”
- judgment and “decision” refer to receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and access.
- (accessing) may include considering something as a “judgment” or “decision.”
- judgment and “decision” refer to resolving, selecting, choosing, establishing, comparing, etc. as “judgment” and “decision”. may be included.
- judgment and “decision” may include regarding some action as having been “judged” or “determined.”
- judgment (decision) may be read as “assuming", “expecting", “considering”, etc.
- the phrase “based on” does not mean “based only on” unless explicitly stated otherwise. In other words, the phrase “based on” means both “based only on” and “based at least on.”
- any reference to the elements herein, such as “first”, “second”, etc., does not generally limit the amount or order of those elements. These designations may be used herein as a convenient way of distinguishing between two or more elements. Thus, reference to a first and second element does not imply that only two elements may be employed therein or that the first element must precede the second element in any way.
- a and B are different may mean “A and B are different from each other.” Note that the term may also mean that "A and B are each different from C”. Terms such as “separate” and “coupled” may also be interpreted similarly to “different.”
- SYMBOLS 10 Delimiter insertion device, 11... Target text acquisition unit, 12... Interword time acquisition unit, 13... Delimiter insertion unit, 14... Output unit, 20... Speech recognition system, 21... Speech recognition result acquisition unit, 22... Speech recognition result output unit, M1...recording medium, m10...main module, m11...target text acquisition module, m12...interword time acquisition module, m13...delimiter insertion module, m14...output module, md1, md2...symbol insertion model .
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Theoretical Computer Science (AREA)
- Artificial Intelligence (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Machine Translation (AREA)
Abstract
Description
<O>…区切り記号なし
<P>…ピリオド、句点
<C>…カンマ、読点
<Q>…クエスチョンマーク、疑問符
区切り記号除去文id1は、区切り記号入り文od1から区切り記号ラベルを除去した文であってもよい。
<O>:90%、<C>:5%、<P>:2%、<Q>:3%
従って、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であるので、単語「I」に対して区切り記号なしのラベル<O>がラベリングされる。
<O>:60%、<C>:5%、<P>:30%、<Q>:5%
従って、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であるので、単語「know」に対して区切り記号なしのラベル<O>がラベリングされる。
<O>:90%、<C>:5%、<P>:2%、<Q>:3%
従って、区切り記号推測情報dp2は、単語「I」に関する各区切り記号の記号挿入尤度及び記号なし尤度のうちの最大の尤度を有する区切り記号なしのラベル<O>が、単語「I」に対してラベリングされたことを示す情報(I=<O>)を含む。そして、区切り記号挿入部13は、対象テキストに含まれる単語「I」の後に、ラベリングされたラベル<O>に基づいて、区切り記号を未挿入とすることを決定する。
<O>:10%、<C>:20%、<P>:70%、<Q>:0%
図6に例示した区切り記号推測情報dp1では、区切り記号を未挿入とすること(ラベル<O>)の記号なし尤度が最大であったのに対して、区切り記号挿入モデルmd2により算出された各尤度においては、ピリオドを挿入すること(ラベル<P>)の記号挿入尤度が最大である。従って、区切り記号推測情報dp2は、ピリオドを挿入することが単語「know」に対してラベリングされたことを示す情報(know=<P>)を含む。そして、区切り記号挿入部13は、対象テキストに含まれる単語「know」の後に、ラベリングされたラベル<P>に基づいて、ピリオドを挿入する。
会議:<Meeting>
講演:<Lecture>
雑談:<Chatting>
Claims (8)
- 発話音声の音声認識により得られたテキストに含まれる単語の後に、文を区切る区切り記号を挿入する区切り記号挿入装置であって、
前記発話音声に含まれる各単語における次の単語が発話されるまでの時間の長さである語間時間を取得する語間時間取得部と、
区切り記号挿入モデル及び前記語間時間に基づいて、前記発話音声の音声認識により得られたテキストである対象テキストに区切り記号を挿入する区切り記号挿入部であって、
前記区切り記号挿入モデルは、
前記区切り記号を含まない文である区切り記号除去文を少なくとも入力とし、
前記区切り記号除去文に含まれる各単語の後に挿入される区切り記号を示す区切り記号推測情報を出力し、
前記区切り記号除去文と区切り記号を含む文である区切り記号入り文とのペアを含む学習データを用いた機械学習により生成されるモデルであり、
前記対象テキストを前記区切り記号除去文として前記区切り記号挿入モデルに入力することにより得られ、前記語間時間に応じて調整された前記区切り記号推測情報に基づいて、前記対象テキストに区切り記号を挿入する、区切り記号挿入部と、
を備える区切り記号挿入装置。 - 前記区切り記号挿入部は、前記区切り記号挿入モデルから出力された前記区切り記号推測情報を、前記語間時間に基づいて調整する、
請求項1に記載の区切り記号挿入装置。 - 前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号なし尤度を含み、
前記区切り記号挿入部は、
前記対象テキストに含まれる単語のうちの一の単語における前記語間時間が第1の時間である場合に、前記一の単語の前記記号なし尤度を上げるように調整し、又は/及び、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を下げるように調整し、
前記一の単語における前記語間時間が前記第1の時間より長い第2の時間である場合に、複数の種類の区切り記号のうちの少なくとも一種の前記一の単語の区切り記号の前記記号挿入尤度を上げるように調整し、又は/及び、前記一の単語の前記記号なし尤度を下げるように調整し、
前記記号挿入尤度及び前記記号なし尤度のうちの最大の尤度に基づいて、前記一の単語の後に、複数の種類のうちのいずれかの区切り記号を挿入し、又は区切り記号を未挿入とする、
請求項2に記載の区切り記号挿入装置。 - 前記区切り記号挿入部は、
前記対象テキストに含まれる全単語のうちの一以上の単語の前記語間時間の長さの程度が、所与の程度より長い場合に、前記全単語のそれぞれの前記語間時間を所与の程度で短くなるように補正した補正語間時間に基づいて、前記記号挿入尤度及び/又は前記記号なし尤度を調整する、
請求項3に記載の区切り記号挿入装置。 - 前記区切り記号挿入モデルは、
前記区切り記号除去文に含まれる各単語の語間時間を入力として更に含み、
前記語間時間が各単語に関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、
前記語間時間により調整された前記区切り記号推測情報を出力し、
前記区切り記号挿入部は、各単語に前記語間時間が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力する、
請求項1に記載の区切り記号挿入装置。 - 前記区切り記号推測情報は、前記区切り記号除去文に含まれる各単語の後に複数の種類の区切り記号のそれぞれを挿入することに関する尤度である記号挿入尤度、及び、各単語の後に区切り記号を未挿入とすることに関する尤度である記号無し尤度を含む、
請求項5に記載された区切り記号挿入装置。 - 前記区切り記号挿入モデルは、
前記区切り記号除去文に対応する発話音声が発話されたときの状況を示すシチュエーション情報を入力として更に含み、
前記語間時間が各単語に関連付けられると共に前記シチュエーション情報が関連付けられた前記区切り記号除去文と前記区切り記号入り文とのペアからなる学習データを用いた機械学習により生成され、
前記区切り記号挿入部は、前記シチュエーション情報が関連付けられた前記対象テキストを前記区切り記号挿入モデルに入力する、
請求項5または6に記載の区切り記号挿入装置。 - 請求項1に記載の区切り記号挿入装置を含む音声認識システムであって、
音声認識エンジンにより音声認識された結果のテキスト及び該テキストにおける各単語の前記語間時間を第1の音声認識結果として取得する音声認識結果取得部と、
音声認識結果出力部と、を備え、
前記区切り記号挿入装置の前記語間時間取得部は、前記第1の音声認識結果に含まれる前記語間時間を取得し、
前記区切り記号挿入装置の前記区切り記号挿入部は、前記第1の音声認識結果に含まれるテキストを前記対象テキストとして、区切り記号を挿入し、
前記音声認識結果出力部は、前記区切り記号挿入部により区切り記号が挿入された前記対象テキストを、第2の音声認識結果として出力する、
を備える音声認識システム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2024538828A JP7809817B2 (ja) | 2022-08-05 | 2023-05-10 | 区切り記号挿入装置及び音声認識システム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2022-125544 | 2022-08-05 | ||
| JP2022125544 | 2022-08-05 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024029152A1 true WO2024029152A1 (ja) | 2024-02-08 |
Family
ID=89849072
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2023/017568 Ceased WO2024029152A1 (ja) | 2022-08-05 | 2023-05-10 | 区切り記号挿入装置及び音声認識システム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP7809817B2 (ja) |
| WO (1) | WO2024029152A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2009101837A1 (ja) * | 2008-02-13 | 2009-08-20 | Nec Corporation | 記号挿入装置および記号挿入方法 |
| JP2015219480A (ja) * | 2014-05-21 | 2015-12-07 | 日本電信電話株式会社 | 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム |
| CN112927679A (zh) * | 2021-02-07 | 2021-06-08 | 虫洞创新平台(深圳)有限公司 | 一种语音识别中添加标点符号的方法及语音识别装置 |
-
2023
- 2023-05-10 JP JP2024538828A patent/JP7809817B2/ja active Active
- 2023-05-10 WO PCT/JP2023/017568 patent/WO2024029152A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2009101837A1 (ja) * | 2008-02-13 | 2009-08-20 | Nec Corporation | 記号挿入装置および記号挿入方法 |
| JP2015219480A (ja) * | 2014-05-21 | 2015-12-07 | 日本電信電話株式会社 | 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム |
| CN112927679A (zh) * | 2021-02-07 | 2021-06-08 | 虫洞创新平台(深圳)有限公司 | 一种语音识别中添加标点符号的方法及语音识别装置 |
Non-Patent Citations (1)
| Title |
|---|
| TILK OTTOKAR, ALUMÄE TANEL: "LSTM for punctuation restoration in speech transcripts", INTERSPEECH 2015, ISCA, ISCA, 1 January 2015 (2015-01-01), ISCA, pages 683 - 687, XP093136514, DOI: 10.21437/Interspeech.2015-240 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7809817B2 (ja) | 2026-02-02 |
| JPWO2024029152A1 (ja) | 2024-02-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8880398B1 (en) | Localized speech recognition with offload | |
| WO2020054451A1 (ja) | 対話装置 | |
| US11663420B2 (en) | Dialogue system | |
| JP6745402B2 (ja) | 質問推定装置 | |
| WO2024029152A1 (ja) | 区切り記号挿入装置及び音声認識システム | |
| JP7682862B2 (ja) | 句点削除モデル学習装置、句点削除モデル及び判定装置 | |
| WO2020070943A1 (ja) | パターン認識装置及び学習済みモデル | |
| JPWO2019187463A1 (ja) | 対話サーバ | |
| JP7087095B2 (ja) | 対話情報生成装置 | |
| JP6584622B1 (ja) | 文章マッチングシステム | |
| JPWO2019220791A1 (ja) | 対話装置 | |
| JP7608603B2 (ja) | 音声認識装置 | |
| JP6545855B1 (ja) | 文章マッチングシステム | |
| JP6960049B2 (ja) | 対話装置 | |
| JP6895580B2 (ja) | 対話システム | |
| JP2022164001A (ja) | 単言語変換装置 | |
| JP2024168531A (ja) | 文生成モデル生成装置、文生成モデル及び文生成装置 | |
| WO2024203390A1 (ja) | 音声認識誤り訂正装置 | |
| JP7093844B2 (ja) | 対話システム | |
| JP2024108744A (ja) | 埋め込み表現生成システム | |
| WO2026078751A1 (ja) | 情報処理装置および情報処理方法 | |
| WO2024241734A1 (ja) | 文脈文決定システム、機械翻訳装置および学習装置 | |
| WO2026078749A1 (ja) | 情報処理装置および情報処理方法 | |
| JP7601649B2 (ja) | 音声合成調整装置 | |
| US12260184B2 (en) | Translation device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23849715 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024538828 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18875751 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23849715 Country of ref document: EP Kind code of ref document: A1 |